Wikitech labswiki https://wikitech.wikimedia.org/wiki/Main_Page MediaWiki 1.47.0-wmf.16 first-letter Media Special Talk User User talk Wikitech Wikitech talk File File talk MediaWiki MediaWiki talk Template Template talk Help Help talk Category Category talk Obsolete Obsolete talk OfficeIT OfficeIT talk Tool Tool talk Nova Resource Nova Resource Talk Heira Heira Talk TimedText TimedText talk Module Module talk Deployments 0 4108 2450672 2450645 2026-08-23T05:02:57Z ScheduleDeploymentBot 37566 Add [[gerrit:1328310]] to Monday, August 24 UTC morning backport window 2450672 wikitext text/x-wiki {{Navigation MediaWiki deployment}} This page tracks '''upcoming''' '''deployments''' of software to the [[:m:Special:SiteMatrix|Wikimedia Foundation servers]]. == Getting started == Ensure you joined the {{irc|wikimedia-operations}} IRC channel as all deployment-related communications happen there. If you need help, contact [[:mw:Wikimedia Release Engineering Team|Release Engineering]] on IRC at {{irc|wikimedia-releng}}; and ping Tyler (<code>thcipriani</code>). * '''MediaWiki is deployed weekly''' through the [[/Train|Deployment Train]]. Other services follow their own schedule. * '''Times are pinned to San Francisco''', thus the UTC time changes in March and November per [[:en:Daylight saving time in the United States|DST]]. * '''Prefer regular [[Backport windows]]''' over adding new windows. To request deployment of a config change or backport, add your username and Gerrit URL to one of the backport windows on this page. You must be online in #wikimedia-operations on IRC during your deployment and install [[WikimediaDebug]] ahead of time. The #wikimedia-operations channel requires you to [[:m:IRC/Instructions#Register your nickname, identify, and enforce|register your nickname]] before you can join. ** You can use the '''backport scheduling tool''' to more easily edit this page: <div style="text-align: center; margin: 1em 0">{{Clickable button 2|:toollabs:schedule-deployment|Schedule a backport|class=mw-ui-progressive}}</div> * Tasks that meet [[/Inclusion criteria|Inclusion criteria]] '''require their own windows''', which includes long-running tasks. '''Schedule more time''' than you think you need to account for delays and set backs, we recommend one hour for most tasks. **To create or modify a recurring deploy window, send a patchset to [[:gitlab:repos/releng/release/-/blob/main/make-deployment-calendar/deployments-calendar.yaml|deployments-calendar.yaml file]] in <code>repos/releng/release.git</code>. **To create an one-off window, simply edit this page accordingly ** '''Announce''' changes to the [[mail:ops|ops mailing list]] ahead of time if you anticipate or are uncertain about noticeable impacts to database load, HTTP caching, or the introduction of new cookies. ** '''Announce''' deployments of major features to the community via [[:m:Tech/News/Next|Tech News]] and/or via other [[:mw:Wikimedia_Product_Guidance/Communication_channels|Product communication channels]]. * '''Something went wrong?''' See [[Incident response]]. Is there a user-impacting problem? Communicate in the {{irc|wikimedia-operations}} IRC channel. If there is a Phabricator task, ensure [[:phab:tag/wikimedia-incident/|#Wikimedia-Incident]] is tagged, and consider setting the [[:mw:Phabricator/Project_management#Priority_levels|Unbreak Now]] priority. __TOC__ {{anchor|Next Week|Near Term|Near term|Near-term}}{{clear}} [[Category:Deployment]] {{Note|content=Subscribe in Google Calendar via <code>wikimedia.org_rudis09ii2mm5fk4hgdjeh1u64@group.calendar.google.com</code>.<br>This may not include one-off windows. '''If there are differences, then the wiki page is canonical and correct'''.}} ==Week of August 24== ==={{Deployment_day|date=2026-08-23}}=== {{Deployment calendar event card |when=2026-08-23 00:00 SF |length=24 |window=No deploys all day! See [[Deployments/Emergencies]] if things are broken. |who= |what=No Deploys }} ==={{Deployment_day|date=2026-08-24}}=== {{Deployment calendar event card |when=2026-08-24 00:00 SF |length=1 |window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}} |what={{ircnick|Hide_on_rosie|Hide_on_rosie}} {{deploy|type=config|gerrit=1307270|title=InitialiseSettings: change enwiki extendedconfirmed autopromote settings|status=}} - {{phabricator|T431060}} {{ircnick|anzx|anzx}} {{deploy|type=config|gerrit=1328310|title=Lift IP cap for Mapudungun editathon on 2026-08-29|status=}} - {{phabricator|T435523}} {{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-24 03:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-24 06:00 SF |length=1 |window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Lucas_WMDE|Lucas}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}} |what={{ircnick|MichaelG_WMF|Michael Grosse (dev @ WMF Growth team)}} {{deploy|type=config|gerrit=1325527|title=Growth: Drop now unused config for benefits block|status=}} - {{phabricator|T433783}} {{ircnick|sadiya_wmde28|Sadiya Halimat Mohammed}} {{deploy|type=config|gerrit=1326848|title=Disable Wikidata Bridge on Catalan Wikipedia|status=}} - {{phabricator|T433713}} {{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-24 07:30 SF |length=0.5 |window=Test Kitchen Experiment Deployment Window |who=Test Kitchen |what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]]. }} {{Deployment calendar event card |when=2026-08-24 08:30 SF |length=0.5 |window=Wikimedia Portals Update |who={{ircnick|jan_drewniak|Jan Drewniak}} |what=Weekly window for the portals page: https://www.wikipedia.org/ }} {{Deployment calendar event card |when=2026-08-24 10:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. {{ircnick|bd808}} * Shellbox updates to pick up new Pygments release ([[phab:T421653|T421653]]) }} {{Deployment calendar event card |when=2026-08-24 10:00 SF |length=0.5 |window=Wikidata Query Service weekly deploy |who={{ircnick|ryankemper|Ryan}} |what=... }} {{Deployment calendar event card |when=2026-08-24 13:00 SF |length=1 |window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-24 14:00 SF |length=2 |window=Weekly Security deployment window |who={{ircnick|alexsanford|Alex}}, {{ircnick|Reedy|Sam}}, {{ircnick|sbassett|Scott}}, {{ircnick|Maryum|Maryum}}, {{ircnick|manfredi|Manfredi}} |what=Held deployment window for Security-team related deploys. }} {{Deployment calendar event card |when=2026-08-24 16:00 SF |length=1 |window=Readers deployment window |who=Readers |what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start }} {{Deployment calendar event card |when=2026-08-24 19:00 SF |length=1 |window=Automatic branching of MediaWiki, extensions, skins, and vendor – see [[Heterogeneous deployment/Train deploys]] |who=N/A |what=Branch <code>wmf/1.47.0-wmf.17</code> }} {{Deployment calendar event card |when=2026-08-24 20:00 SF |length=1 |window=Automatic deployment of MediaWiki, extensions, skins, and vendor to testwikis only – see [[Heterogeneous deployment/Train deploys]] |who=N/A |what=Deploy <code>wmf/1.47.0-wmf.17</code> to testwikis }} {{Deployment calendar event card |when=2026-08-24 21:00 SF |length=1 |window=Automatic removal of all obsolete MediaWiki versions from the deployment and bare metal servers (except the most-recent obsolete version) |who=N/A |what=Runs <code>scap clean auto</code> }} {{Deployment calendar event card |when=2026-08-24 23:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-24 23:00 SF |length=0.5 |window=Primary database switchover |who={{ircnick|marostegui|Manuel Arostegui}}, {{ircnick|Amir1|Amir}}, {{ircnick|federico3|Federico Ceratto}} |what=Held deployment window for database primary masters maintenance }} ==={{Deployment_day|date=2026-08-25}}=== {{Deployment calendar event card |when=2026-08-25 00:00 SF |length=1 |window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-25 01:00 SF |length=2 |window=MediaWiki train - Utc-0 Version |who={{ircnick|hashar|Antoine}}, {{ircnick|andre|Andre}} |what=[[mw:MediaWiki 1.47/Roadmap#Schedule for the deployments|1.47 schedule]] {{DeployOneWeekMini|1.47.0-wmf.16->1.47.0-wmf.17|1.47.0-wmf.16|1.47.0-wmf.16}} * group0 to [[mw:MediaWiki_1.47/wmf.17|1.47.0-wmf.17]] * '''Blockers: {{phabricator|T430836}}''' }} {{Deployment calendar event card |when=2026-08-25 03:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-25 05:00 SF |length=1 |window=Mobileapps/RESTBase/Wikifeeds |who=Content Transform Team |what=Content transform team node services (mobileapps/wikifeeds) }} {{Deployment calendar event card |when=2026-08-25 06:00 SF |length=1 |window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Lucas_WMDE|Lucas}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-25 07:00 SF |length=0.5 |window=Test Kitchen UI Deployment Window |who=Experimentation Platform Team |what=Deployment of Test Kitchen UI (fka MPIC) }} {{Deployment calendar event card |when=2026-08-25 07:30 SF |length=0.5 |window=Test Kitchen Experiment Deployment Window |who=Test Kitchen |what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]]. }} {{Deployment calendar event card |when=2026-08-25 08:00 SF |length=1 |window=SRE Collaboration Services office hours |who={{ircnick|jelto|Jelto}}, {{ircnick|arnoldokoth|Arnold}}, {{ircnick|mutante|Daniel}}, {{ircnick|arnaudb|Arnaud}} |what=Services including Gerrit, Phorge (Phabricator), GitLab }} {{Deployment calendar event card |when=2026-08-25 09:00 SF |length=1 |window=[[Puppet request window]]<br/><small>'''(Max 6 patches)'''</small> |who={{ircnick|jhathaway|JHathaway}}, {{ircnick|rzl|Reuven}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to Puppet change'' }} {{Deployment calendar event card |when=2026-08-25 10:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-25 13:00 SF |length=1 |window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-25 14:00 SF |length=1 |window=Readers deployment window |who=Readers |what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start }} {{Deployment calendar event card |when=2026-08-25 23:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} ==={{Deployment_day|date=2026-08-26}}=== {{Deployment calendar event card |when=2026-08-26 00:00 SF |length=1 |window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-26 01:00 SF |length=2 |window=MediaWiki train - Utc-0 Version |who={{ircnick|hashar|Antoine}}, {{ircnick|andre|Andre}} |what=[[mw:MediaWiki 1.47/Roadmap#Schedule for the deployments|1.47 schedule]] {{DeployOneWeekMini|1.47.0-wmf.17|1.47.0-wmf.16->1.47.0-wmf.17|1.47.0-wmf.16}} * group1 to [[mw:MediaWiki_1.47/wmf.17|1.47.0-wmf.17]] * '''Blockers: {{phabricator|T430836}}''' }} {{Deployment calendar event card |when=2026-08-26 03:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-26 04:00 SF |length=1 |window=[[mw:Services|Services]] – [[Citoid]] / [[Zotero]] |who=Marielle ({{ircnick|mvolz}}) |what=See [[mw:Citoid|Citoid]] }} {{Deployment calendar event card |when=2026-08-26 06:00 SF |length=1 |window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-26 07:00 SF |length=1 |window=Wikifunctions Services UTC Afternoon |who=Abstract Wikipedia team (Africa, Europe, Eastern Americas) |what=Wikifunctions back-end k8s services }} {{Deployment calendar event card |when=2026-08-26 07:30 SF |length=0.5 |window=Test Kitchen Experiment Deployment Window |who=Test Kitchen |what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]]. }} {{Deployment calendar event card |when=2026-08-26 10:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-26 13:00 SF |length=1 |window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-26 14:00 SF |length=1 |window=Wikifunctions Services UTC Late |who=Abstract Wikipedia team (North and South America) |what=Wikifunctions back-end k8s services }} {{Deployment calendar event card |when=2026-08-26 15:00 SF |length=1 |window=Readers deployment window |who=Readers |what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start }} {{Deployment calendar event card |when=2026-08-26 23:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-26 23:00 SF |length=0.5 |window=Primary database switchover |who={{ircnick|marostegui|Manuel Arostegui}}, {{ircnick|Amir1|Amir}}, {{ircnick|federico3|Federico Ceratto}} |what=Held deployment window for database primary masters maintenance }} ==={{Deployment_day|date=2026-08-27}}=== {{Deployment calendar event card |when=2026-08-27 00:00 SF |length=1 |window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-27 01:00 SF |length=2 |window=MediaWiki train - Utc-0 Version |who={{ircnick|hashar|Antoine}}, {{ircnick|andre|Andre}} |what=[[mw:MediaWiki 1.47/Roadmap#Schedule for the deployments|1.47 schedule]] {{DeployOneWeekMini|1.47.0-wmf.17|1.47.0-wmf.17|1.47.0-wmf.16->1.47.0-wmf.17}} * group2 to [[mw:MediaWiki_1.47/wmf.17|1.47.0-wmf.17]] * '''Blockers: {{phabricator|T430836}}''' }} {{Deployment calendar event card |when=2026-08-27 03:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-27 05:00 SF |length=1 |window=Mobileapps/RESTBase/Wikifeeds |who=Content Transform Team |what=Content transform team node services (mobileapps/wikifeeds) }} {{Deployment calendar event card |when=2026-08-27 06:00 SF |length=1 |window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Lucas_WMDE|Lucas}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-27 07:30 SF |length=0.5 |window=Test Kitchen Experiment Deployment Window |who=Test Kitchen |what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]]. }} {{Deployment calendar event card |when=2026-08-27 09:00 SF |length=1 |window=[[Puppet request window]]<br/><small>'''(Max 6 patches)'''</small> |who={{ircnick|jhathaway|JHathaway}}, {{ircnick|rzl|Reuven}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to Puppet change'' }} {{Deployment calendar event card |when=2026-08-27 10:00 SF |length=1 |window=Cloud Services/Technical Documentation weekly deploy (Toolhub, Developer portal, Striker) |who={{ircnick|bd808}} |what=... }} {{Deployment calendar event card |when=2026-08-27 10:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-27 13:00 SF |length=1 |window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-27 14:00 SF |length=1 |window=Readers deployment window |who=Readers |what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start }} {{Deployment calendar event card |when=2026-08-27 23:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} ==={{Deployment_day|date=2026-08-28}}=== {{Deployment calendar event card |when=2026-08-28 00:00 SF |length=24 |window=No deploys all day! See [[Deployments/Emergencies]] if things are broken. |who= |what=No Deploys }} {{Deployment calendar event card |when=2026-08-28 04:00 SF |length=0.5 |window=GitLab version upgrades |who={{ircnick|jelto|Jelto}}, {{ircnick|arnoldokoth|Arnold}}, {{ircnick|mutante|Daniel}}, {{ircnick|arnaudb|Arnaud}} |what=GitLab version upgrades }} ==={{Deployment_day|date=2026-08-29}}=== {{Deployment calendar event card |when=2026-08-29 00:00 SF |length=24 |window=No deploys all day! See [[Deployments/Emergencies]] if things are broken. |who= |what=No Deploys }} ==Week of August 31== ==={{Deployment_day|date=2026-08-30}}=== {{Deployment calendar event card |when=2026-08-30 00:00 SF |length=24 |window=No deploys all day! See [[Deployments/Emergencies]] if things are broken. |who= |what=No Deploys }} ==={{Deployment_day|date=2026-08-31}}=== {{Deployment calendar event card |when=2026-08-31 00:00 SF |length=1 |window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-31 03:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-31 06:00 SF |length=1 |window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Lucas_WMDE|Lucas}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-31 07:30 SF |length=0.5 |window=Test Kitchen Experiment Deployment Window |who=Test Kitchen |what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]]. }} {{Deployment calendar event card |when=2026-08-31 08:30 SF |length=0.5 |window=Wikimedia Portals Update |who={{ircnick|jan_drewniak|Jan Drewniak}} |what=Weekly window for the portals page: https://www.wikipedia.org/ }} {{Deployment calendar event card |when=2026-08-31 10:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-31 10:00 SF |length=0.5 |window=Wikidata Query Service weekly deploy |who={{ircnick|ryankemper|Ryan}} |what=... }} {{Deployment calendar event card |when=2026-08-31 13:00 SF |length=1 |window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-31 14:00 SF |length=2 |window=Weekly Security deployment window |who={{ircnick|alexsanford|Alex}}, {{ircnick|Reedy|Sam}}, {{ircnick|sbassett|Scott}}, {{ircnick|Maryum|Maryum}}, {{ircnick|manfredi|Manfredi}} |what=Held deployment window for Security-team related deploys. }} {{Deployment calendar event card |when=2026-08-31 16:00 SF |length=1 |window=Readers deployment window |who=Readers |what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start }} {{Deployment calendar event card |when=2026-08-31 19:00 SF |length=1 |window=Automatic branching of MediaWiki, extensions, skins, and vendor – see [[Heterogeneous deployment/Train deploys]] |who=N/A |what=Branch <code>wmf/1.47.0-wmf.18</code> }} {{Deployment calendar event card |when=2026-08-31 20:00 SF |length=1 |window=Automatic deployment of MediaWiki, extensions, skins, and vendor to testwikis only – see [[Heterogeneous deployment/Train deploys]] |who=N/A |what=Deploy <code>wmf/1.47.0-wmf.18</code> to testwikis }} {{Deployment calendar event card |when=2026-08-31 21:00 SF |length=1 |window=Automatic removal of all obsolete MediaWiki versions from the deployment and bare metal servers (except the most-recent obsolete version) |who=N/A |what=Runs <code>scap clean auto</code> }} {{Deployment calendar event card |when=2026-08-31 23:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-31 23:00 SF |length=0.5 |window=Primary database switchover |who={{ircnick|marostegui|Manuel Arostegui}}, {{ircnick|cezmunsta|Ceri Williams}}, {{ircnick|federico3|Federico Ceratto}} |what=Held deployment window for database primary masters maintenance }} ==={{Deployment_day|date=2026-09-01}}=== {{Deployment calendar event card |when=2026-09-01 00:00 SF |length=1 |window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-09-01 03:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-09-01 05:00 SF |length=1 |window=Mobileapps/RESTBase/Wikifeeds |who=Content Transform Team |what=Content transform team node services (mobileapps/wikifeeds) }} {{Deployment calendar event card |when=2026-09-01 06:00 SF |length=1 |window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Lucas_WMDE|Lucas}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-09-01 07:00 SF |length=0.5 |window=Test Kitchen UI Deployment Window |who=Experimentation Platform Team |what=Deployment of Test Kitchen UI (fka MPIC) }} {{Deployment calendar event card |when=2026-09-01 07:30 SF |length=0.5 |window=Test Kitchen Experiment Deployment Window |who=Test Kitchen |what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]]. }} {{Deployment calendar event card |when=2026-09-01 08:00 SF |length=1 |window=SRE Collaboration Services office hours |who={{ircnick|jelto|Jelto}}, {{ircnick|arnoldokoth|Arnold}}, {{ircnick|mutante|Daniel}}, {{ircnick|arnaudb|Arnaud}} |what=Services including Gerrit, Phorge (Phabricator), GitLab }} {{Deployment calendar event card |when=2026-09-01 09:00 SF |length=1 |window=[[Puppet request window]]<br/><small>'''(Max 6 patches)'''</small> |who={{ircnick|jhathaway|JHathaway}}, {{ircnick|rzl|Reuven}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to Puppet change'' }} {{Deployment calendar event card |when=2026-09-01 10:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-09-01 11:00 SF |length=2 |window=MediaWiki train - Utc-7+Utc-0 Version |who={{ircnick|dancy|Ahmon}}, {{ircnick|hashar|Antoine}} |what=[[mw:MediaWiki 1.47/Roadmap#Schedule for the deployments|1.47 schedule]] {{DeployOneWeekMini|1.47.0-wmf.17->1.47.0-wmf.18|1.47.0-wmf.17|1.47.0-wmf.17}} * group0 to [[mw:MediaWiki_1.47/wmf.18|1.47.0-wmf.18]] * '''Blockers: {{phabricator|T430837}}''' }} {{Deployment calendar event card |when=2026-09-01 13:00 SF |length=1 |window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-09-01 14:00 SF |length=1 |window=Readers deployment window |who=Readers |what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start }} {{Deployment calendar event card |when=2026-09-01 23:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} ==={{Deployment_day|date=2026-09-02}}=== {{Deployment calendar event card |when=2026-09-02 00:00 SF |length=1 |window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-09-02 01:00 SF |length=2 |window=MediaWiki train - Utc-7+Utc-0 Version (secondary timeslot) |who={{ircnick|dancy|Ahmon}}, {{ircnick|hashar|Antoine}} |what=[[mw:MediaWiki 1.47/Roadmap#Schedule for the deployments|1.47 schedule]] {{DeployOneWeekMini|1.47.0-wmf.18|1.47.0-wmf.17->1.47.0-wmf.18|1.47.0-wmf.17}} * group1 to [[mw:MediaWiki_1.47/wmf.18|1.47.0-wmf.18]] * '''Blockers: {{phabricator|T430837}}''' }} {{Deployment calendar event card |when=2026-09-02 03:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-09-02 04:00 SF |length=1 |window=[[mw:Services|Services]] – [[Citoid]] / [[Zotero]] |who=Marielle ({{ircnick|mvolz}}) |what=See [[mw:Citoid|Citoid]] }} {{Deployment calendar event card |when=2026-09-02 06:00 SF |length=1 |window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-09-02 07:00 SF |length=1 |window=Wikifunctions Services UTC Afternoon |who=Abstract Wikipedia team (Africa, Europe, Eastern Americas) |what=Wikifunctions back-end k8s services }} {{Deployment calendar event card |when=2026-09-02 07:30 SF |length=0.5 |window=Test Kitchen Experiment Deployment Window |who=Test Kitchen |what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]]. }} {{Deployment calendar event card |when=2026-09-02 10:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-09-02 11:00 SF |length=2 |window=MediaWiki train - Utc-7+Utc-0 Version |who={{ircnick|dancy|Ahmon}}, {{ircnick|hashar|Antoine}} |what=[[mw:MediaWiki 1.47/Roadmap#Schedule for the deployments|1.47 schedule]] {{DeployOneWeekMini|1.47.0-wmf.18|1.47.0-wmf.17->1.47.0-wmf.18|1.47.0-wmf.17}} * group1 to [[mw:MediaWiki_1.47/wmf.18|1.47.0-wmf.18]] * '''Blockers: {{phabricator|T430837}}''' }} {{Deployment calendar event card |when=2026-09-02 13:00 SF |length=1 |window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-09-02 14:00 SF |length=1 |window=Wikifunctions Services UTC Late |who=Abstract Wikipedia team (North and South America) |what=Wikifunctions back-end k8s services }} {{Deployment calendar event card |when=2026-09-02 15:00 SF |length=1 |window=Readers deployment window |who=Readers |what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start }} {{Deployment calendar event card |when=2026-09-02 23:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-09-02 23:00 SF |length=0.5 |window=Primary database switchover |who={{ircnick|marostegui|Manuel Arostegui}}, {{ircnick|cezmunsta|Ceri Williams}}, {{ircnick|federico3|Federico Ceratto}} |what=Held deployment window for database primary masters maintenance }} ==={{Deployment_day|date=2026-09-03}}=== {{Deployment calendar event card |when=2026-09-03 00:00 SF |length=1 |window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-09-03 01:00 SF |length=2 |window=MediaWiki train - Utc-7+Utc-0 Version (secondary timeslot) |who={{ircnick|dancy|Ahmon}}, {{ircnick|hashar|Antoine}} |what=[[mw:MediaWiki 1.47/Roadmap#Schedule for the deployments|1.47 schedule]] {{DeployOneWeekMini|1.47.0-wmf.18|1.47.0-wmf.18|1.47.0-wmf.17->1.47.0-wmf.18}} * group2 to [[mw:MediaWiki_1.47/wmf.18|1.47.0-wmf.18]] * '''Blockers: {{phabricator|T430837}}''' }} {{Deployment calendar event card |when=2026-09-03 03:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-09-03 05:00 SF |length=1 |window=Mobileapps/RESTBase/Wikifeeds |who=Content Transform Team |what=Content transform team node services (mobileapps/wikifeeds) }} {{Deployment calendar event card |when=2026-09-03 06:00 SF |length=1 |window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Lucas_WMDE|Lucas}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-09-03 07:30 SF |length=0.5 |window=Test Kitchen Experiment Deployment Window |who=Test Kitchen |what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]]. }} {{Deployment calendar event card |when=2026-09-03 09:00 SF |length=1 |window=[[Puppet request window]]<br/><small>'''(Max 6 patches)'''</small> |who={{ircnick|jhathaway|JHathaway}}, {{ircnick|rzl|Reuven}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to Puppet change'' }} {{Deployment calendar event card |when=2026-09-03 10:00 SF |length=1 |window=Cloud Services/Technical Documentation weekly deploy (Toolhub, Developer portal, Striker) |who={{ircnick|bd808}} |what=... }} {{Deployment calendar event card |when=2026-09-03 10:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-09-03 11:00 SF |length=2 |window=MediaWiki train - Utc-7+Utc-0 Version |who={{ircnick|dancy|Ahmon}}, {{ircnick|hashar|Antoine}} |what=[[mw:MediaWiki 1.47/Roadmap#Schedule for the deployments|1.47 schedule]] {{DeployOneWeekMini|1.47.0-wmf.18|1.47.0-wmf.18|1.47.0-wmf.17->1.47.0-wmf.18}} * group2 to [[mw:MediaWiki_1.47/wmf.18|1.47.0-wmf.18]] * '''Blockers: {{phabricator|T430837}}''' }} {{Deployment calendar event card |when=2026-09-03 13:00 SF |length=1 |window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-09-03 14:00 SF |length=1 |window=Readers deployment window |who=Readers |what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start }} {{Deployment calendar event card |when=2026-09-03 23:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} ==={{Deployment_day|date=2026-09-04}}=== {{Deployment calendar event card |when=2026-09-04 00:00 SF |length=24 |window=No deploys all day! See [[Deployments/Emergencies]] if things are broken. |who= |what=No Deploys }} {{Deployment calendar event card |when=2026-09-04 04:00 SF |length=0.5 |window=GitLab version upgrades |who={{ircnick|jelto|Jelto}}, {{ircnick|arnoldokoth|Arnold}}, {{ircnick|mutante|Daniel}}, {{ircnick|arnaudb|Arnaud}} |what=GitLab version upgrades }} ==={{Deployment_day|date=2026-09-05}}=== {{Deployment calendar event card |when=2026-09-05 00:00 SF |length=24 |window=No deploys all day! See [[Deployments/Emergencies]] if things are broken. |who= |what=No Deploys }} kk5yu61vvbi2a8eiosxokdikam12k4u 2450673 2450672 2026-08-23T05:04:18Z A826 34037 /* {{Deployment_day|date=2026-08-24}} */ 2450673 wikitext text/x-wiki {{Navigation MediaWiki deployment}} This page tracks '''upcoming''' '''deployments''' of software to the [[:m:Special:SiteMatrix|Wikimedia Foundation servers]]. == Getting started == Ensure you joined the {{irc|wikimedia-operations}} IRC channel as all deployment-related communications happen there. If you need help, contact [[:mw:Wikimedia Release Engineering Team|Release Engineering]] on IRC at {{irc|wikimedia-releng}}; and ping Tyler (<code>thcipriani</code>). * '''MediaWiki is deployed weekly''' through the [[/Train|Deployment Train]]. Other services follow their own schedule. * '''Times are pinned to San Francisco''', thus the UTC time changes in March and November per [[:en:Daylight saving time in the United States|DST]]. * '''Prefer regular [[Backport windows]]''' over adding new windows. To request deployment of a config change or backport, add your username and Gerrit URL to one of the backport windows on this page. You must be online in #wikimedia-operations on IRC during your deployment and install [[WikimediaDebug]] ahead of time. The #wikimedia-operations channel requires you to [[:m:IRC/Instructions#Register your nickname, identify, and enforce|register your nickname]] before you can join. ** You can use the '''backport scheduling tool''' to more easily edit this page: <div style="text-align: center; margin: 1em 0">{{Clickable button 2|:toollabs:schedule-deployment|Schedule a backport|class=mw-ui-progressive}}</div> * Tasks that meet [[/Inclusion criteria|Inclusion criteria]] '''require their own windows''', which includes long-running tasks. '''Schedule more time''' than you think you need to account for delays and set backs, we recommend one hour for most tasks. **To create or modify a recurring deploy window, send a patchset to [[:gitlab:repos/releng/release/-/blob/main/make-deployment-calendar/deployments-calendar.yaml|deployments-calendar.yaml file]] in <code>repos/releng/release.git</code>. **To create an one-off window, simply edit this page accordingly ** '''Announce''' changes to the [[mail:ops|ops mailing list]] ahead of time if you anticipate or are uncertain about noticeable impacts to database load, HTTP caching, or the introduction of new cookies. ** '''Announce''' deployments of major features to the community via [[:m:Tech/News/Next|Tech News]] and/or via other [[:mw:Wikimedia_Product_Guidance/Communication_channels|Product communication channels]]. * '''Something went wrong?''' See [[Incident response]]. Is there a user-impacting problem? Communicate in the {{irc|wikimedia-operations}} IRC channel. If there is a Phabricator task, ensure [[:phab:tag/wikimedia-incident/|#Wikimedia-Incident]] is tagged, and consider setting the [[:mw:Phabricator/Project_management#Priority_levels|Unbreak Now]] priority. __TOC__ {{anchor|Next Week|Near Term|Near term|Near-term}}{{clear}} [[Category:Deployment]] {{Note|content=Subscribe in Google Calendar via <code>wikimedia.org_rudis09ii2mm5fk4hgdjeh1u64@group.calendar.google.com</code>.<br>This may not include one-off windows. '''If there are differences, then the wiki page is canonical and correct'''.}} ==Week of August 24== ==={{Deployment_day|date=2026-08-23}}=== {{Deployment calendar event card |when=2026-08-23 00:00 SF |length=24 |window=No deploys all day! See [[Deployments/Emergencies]] if things are broken. |who= |what=No Deploys }} ==={{Deployment_day|date=2026-08-24}}=== {{Deployment calendar event card |when=2026-08-24 00:00 SF |length=1 |window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}} |what={{ircnick|Hide_on_rosie|Hide_on_rosie}} {{deploy|type=config|gerrit=1307270|title=InitialiseSettings: change enwiki extendedconfirmed autopromote settings|status=}} - {{phabricator|T431060}} {{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-24 03:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-24 06:00 SF |length=1 |window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Lucas_WMDE|Lucas}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}} |what={{ircnick|MichaelG_WMF|Michael Grosse (dev @ WMF Growth team)}} {{deploy|type=config|gerrit=1325527|title=Growth: Drop now unused config for benefits block|status=}} - {{phabricator|T433783}} {{ircnick|sadiya_wmde28|Sadiya Halimat Mohammed}} {{deploy|type=config|gerrit=1326848|title=Disable Wikidata Bridge on Catalan Wikipedia|status=}} - {{phabricator|T433713}} {{ircnick|anzx|anzx}} {{deploy|type=config|gerrit=1328310|title=Lift IP cap for Mapudungun editathon on 2026-08-29|status=}} - {{phabricator|T435523}} {{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-24 07:30 SF |length=0.5 |window=Test Kitchen Experiment Deployment Window |who=Test Kitchen |what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]]. }} {{Deployment calendar event card |when=2026-08-24 08:30 SF |length=0.5 |window=Wikimedia Portals Update |who={{ircnick|jan_drewniak|Jan Drewniak}} |what=Weekly window for the portals page: https://www.wikipedia.org/ }} {{Deployment calendar event card |when=2026-08-24 10:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. {{ircnick|bd808}} * Shellbox updates to pick up new Pygments release ([[phab:T421653|T421653]]) }} {{Deployment calendar event card |when=2026-08-24 10:00 SF |length=0.5 |window=Wikidata Query Service weekly deploy |who={{ircnick|ryankemper|Ryan}} |what=... }} {{Deployment calendar event card |when=2026-08-24 13:00 SF |length=1 |window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-24 14:00 SF |length=2 |window=Weekly Security deployment window |who={{ircnick|alexsanford|Alex}}, {{ircnick|Reedy|Sam}}, {{ircnick|sbassett|Scott}}, {{ircnick|Maryum|Maryum}}, {{ircnick|manfredi|Manfredi}} |what=Held deployment window for Security-team related deploys. }} {{Deployment calendar event card |when=2026-08-24 16:00 SF |length=1 |window=Readers deployment window |who=Readers |what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start }} {{Deployment calendar event card |when=2026-08-24 19:00 SF |length=1 |window=Automatic branching of MediaWiki, extensions, skins, and vendor – see [[Heterogeneous deployment/Train deploys]] |who=N/A |what=Branch <code>wmf/1.47.0-wmf.17</code> }} {{Deployment calendar event card |when=2026-08-24 20:00 SF |length=1 |window=Automatic deployment of MediaWiki, extensions, skins, and vendor to testwikis only – see [[Heterogeneous deployment/Train deploys]] |who=N/A |what=Deploy <code>wmf/1.47.0-wmf.17</code> to testwikis }} {{Deployment calendar event card |when=2026-08-24 21:00 SF |length=1 |window=Automatic removal of all obsolete MediaWiki versions from the deployment and bare metal servers (except the most-recent obsolete version) |who=N/A |what=Runs <code>scap clean auto</code> }} {{Deployment calendar event card |when=2026-08-24 23:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-24 23:00 SF |length=0.5 |window=Primary database switchover |who={{ircnick|marostegui|Manuel Arostegui}}, {{ircnick|Amir1|Amir}}, {{ircnick|federico3|Federico Ceratto}} |what=Held deployment window for database primary masters maintenance }} ==={{Deployment_day|date=2026-08-25}}=== {{Deployment calendar event card |when=2026-08-25 00:00 SF |length=1 |window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-25 01:00 SF |length=2 |window=MediaWiki train - Utc-0 Version |who={{ircnick|hashar|Antoine}}, {{ircnick|andre|Andre}} |what=[[mw:MediaWiki 1.47/Roadmap#Schedule for the deployments|1.47 schedule]] {{DeployOneWeekMini|1.47.0-wmf.16->1.47.0-wmf.17|1.47.0-wmf.16|1.47.0-wmf.16}} * group0 to [[mw:MediaWiki_1.47/wmf.17|1.47.0-wmf.17]] * '''Blockers: {{phabricator|T430836}}''' }} {{Deployment calendar event card |when=2026-08-25 03:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-25 05:00 SF |length=1 |window=Mobileapps/RESTBase/Wikifeeds |who=Content Transform Team |what=Content transform team node services (mobileapps/wikifeeds) }} {{Deployment calendar event card |when=2026-08-25 06:00 SF |length=1 |window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Lucas_WMDE|Lucas}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-25 07:00 SF |length=0.5 |window=Test Kitchen UI Deployment Window |who=Experimentation Platform Team |what=Deployment of Test Kitchen UI (fka MPIC) }} {{Deployment calendar event card |when=2026-08-25 07:30 SF |length=0.5 |window=Test Kitchen Experiment Deployment Window |who=Test Kitchen |what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]]. }} {{Deployment calendar event card |when=2026-08-25 08:00 SF |length=1 |window=SRE Collaboration Services office hours |who={{ircnick|jelto|Jelto}}, {{ircnick|arnoldokoth|Arnold}}, {{ircnick|mutante|Daniel}}, {{ircnick|arnaudb|Arnaud}} |what=Services including Gerrit, Phorge (Phabricator), GitLab }} {{Deployment calendar event card |when=2026-08-25 09:00 SF |length=1 |window=[[Puppet request window]]<br/><small>'''(Max 6 patches)'''</small> |who={{ircnick|jhathaway|JHathaway}}, {{ircnick|rzl|Reuven}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to Puppet change'' }} {{Deployment calendar event card |when=2026-08-25 10:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-25 13:00 SF |length=1 |window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-25 14:00 SF |length=1 |window=Readers deployment window |who=Readers |what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start }} {{Deployment calendar event card |when=2026-08-25 23:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} ==={{Deployment_day|date=2026-08-26}}=== {{Deployment calendar event card |when=2026-08-26 00:00 SF |length=1 |window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-26 01:00 SF |length=2 |window=MediaWiki train - Utc-0 Version |who={{ircnick|hashar|Antoine}}, {{ircnick|andre|Andre}} |what=[[mw:MediaWiki 1.47/Roadmap#Schedule for the deployments|1.47 schedule]] {{DeployOneWeekMini|1.47.0-wmf.17|1.47.0-wmf.16->1.47.0-wmf.17|1.47.0-wmf.16}} * group1 to [[mw:MediaWiki_1.47/wmf.17|1.47.0-wmf.17]] * '''Blockers: {{phabricator|T430836}}''' }} {{Deployment calendar event card |when=2026-08-26 03:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-26 04:00 SF |length=1 |window=[[mw:Services|Services]] – [[Citoid]] / [[Zotero]] |who=Marielle ({{ircnick|mvolz}}) |what=See [[mw:Citoid|Citoid]] }} {{Deployment calendar event card |when=2026-08-26 06:00 SF |length=1 |window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-26 07:00 SF |length=1 |window=Wikifunctions Services UTC Afternoon |who=Abstract Wikipedia team (Africa, Europe, Eastern Americas) |what=Wikifunctions back-end k8s services }} {{Deployment calendar event card |when=2026-08-26 07:30 SF |length=0.5 |window=Test Kitchen Experiment Deployment Window |who=Test Kitchen |what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]]. }} {{Deployment calendar event card |when=2026-08-26 10:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-26 13:00 SF |length=1 |window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-26 14:00 SF |length=1 |window=Wikifunctions Services UTC Late |who=Abstract Wikipedia team (North and South America) |what=Wikifunctions back-end k8s services }} {{Deployment calendar event card |when=2026-08-26 15:00 SF |length=1 |window=Readers deployment window |who=Readers |what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start }} {{Deployment calendar event card |when=2026-08-26 23:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-26 23:00 SF |length=0.5 |window=Primary database switchover |who={{ircnick|marostegui|Manuel Arostegui}}, {{ircnick|Amir1|Amir}}, {{ircnick|federico3|Federico Ceratto}} |what=Held deployment window for database primary masters maintenance }} ==={{Deployment_day|date=2026-08-27}}=== {{Deployment calendar event card |when=2026-08-27 00:00 SF |length=1 |window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-27 01:00 SF |length=2 |window=MediaWiki train - Utc-0 Version |who={{ircnick|hashar|Antoine}}, {{ircnick|andre|Andre}} |what=[[mw:MediaWiki 1.47/Roadmap#Schedule for the deployments|1.47 schedule]] {{DeployOneWeekMini|1.47.0-wmf.17|1.47.0-wmf.17|1.47.0-wmf.16->1.47.0-wmf.17}} * group2 to [[mw:MediaWiki_1.47/wmf.17|1.47.0-wmf.17]] * '''Blockers: {{phabricator|T430836}}''' }} {{Deployment calendar event card |when=2026-08-27 03:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-27 05:00 SF |length=1 |window=Mobileapps/RESTBase/Wikifeeds |who=Content Transform Team |what=Content transform team node services (mobileapps/wikifeeds) }} {{Deployment calendar event card |when=2026-08-27 06:00 SF |length=1 |window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Lucas_WMDE|Lucas}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-27 07:30 SF |length=0.5 |window=Test Kitchen Experiment Deployment Window |who=Test Kitchen |what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]]. }} {{Deployment calendar event card |when=2026-08-27 09:00 SF |length=1 |window=[[Puppet request window]]<br/><small>'''(Max 6 patches)'''</small> |who={{ircnick|jhathaway|JHathaway}}, {{ircnick|rzl|Reuven}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to Puppet change'' }} {{Deployment calendar event card |when=2026-08-27 10:00 SF |length=1 |window=Cloud Services/Technical Documentation weekly deploy (Toolhub, Developer portal, Striker) |who={{ircnick|bd808}} |what=... }} {{Deployment calendar event card |when=2026-08-27 10:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-27 13:00 SF |length=1 |window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-27 14:00 SF |length=1 |window=Readers deployment window |who=Readers |what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start }} {{Deployment calendar event card |when=2026-08-27 23:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} ==={{Deployment_day|date=2026-08-28}}=== {{Deployment calendar event card |when=2026-08-28 00:00 SF |length=24 |window=No deploys all day! See [[Deployments/Emergencies]] if things are broken. |who= |what=No Deploys }} {{Deployment calendar event card |when=2026-08-28 04:00 SF |length=0.5 |window=GitLab version upgrades |who={{ircnick|jelto|Jelto}}, {{ircnick|arnoldokoth|Arnold}}, {{ircnick|mutante|Daniel}}, {{ircnick|arnaudb|Arnaud}} |what=GitLab version upgrades }} ==={{Deployment_day|date=2026-08-29}}=== {{Deployment calendar event card |when=2026-08-29 00:00 SF |length=24 |window=No deploys all day! See [[Deployments/Emergencies]] if things are broken. |who= |what=No Deploys }} ==Week of August 31== ==={{Deployment_day|date=2026-08-30}}=== {{Deployment calendar event card |when=2026-08-30 00:00 SF |length=24 |window=No deploys all day! See [[Deployments/Emergencies]] if things are broken. |who= |what=No Deploys }} ==={{Deployment_day|date=2026-08-31}}=== {{Deployment calendar event card |when=2026-08-31 00:00 SF |length=1 |window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-31 03:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-31 06:00 SF |length=1 |window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Lucas_WMDE|Lucas}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-31 07:30 SF |length=0.5 |window=Test Kitchen Experiment Deployment Window |who=Test Kitchen |what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]]. }} {{Deployment calendar event card |when=2026-08-31 08:30 SF |length=0.5 |window=Wikimedia Portals Update |who={{ircnick|jan_drewniak|Jan Drewniak}} |what=Weekly window for the portals page: https://www.wikipedia.org/ }} {{Deployment calendar event card |when=2026-08-31 10:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-31 10:00 SF |length=0.5 |window=Wikidata Query Service weekly deploy |who={{ircnick|ryankemper|Ryan}} |what=... }} {{Deployment calendar event card |when=2026-08-31 13:00 SF |length=1 |window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-31 14:00 SF |length=2 |window=Weekly Security deployment window |who={{ircnick|alexsanford|Alex}}, {{ircnick|Reedy|Sam}}, {{ircnick|sbassett|Scott}}, {{ircnick|Maryum|Maryum}}, {{ircnick|manfredi|Manfredi}} |what=Held deployment window for Security-team related deploys. }} {{Deployment calendar event card |when=2026-08-31 16:00 SF |length=1 |window=Readers deployment window |who=Readers |what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start }} {{Deployment calendar event card |when=2026-08-31 19:00 SF |length=1 |window=Automatic branching of MediaWiki, extensions, skins, and vendor – see [[Heterogeneous deployment/Train deploys]] |who=N/A |what=Branch <code>wmf/1.47.0-wmf.18</code> }} {{Deployment calendar event card |when=2026-08-31 20:00 SF |length=1 |window=Automatic deployment of MediaWiki, extensions, skins, and vendor to testwikis only – see [[Heterogeneous deployment/Train deploys]] |who=N/A |what=Deploy <code>wmf/1.47.0-wmf.18</code> to testwikis }} {{Deployment calendar event card |when=2026-08-31 21:00 SF |length=1 |window=Automatic removal of all obsolete MediaWiki versions from the deployment and bare metal servers (except the most-recent obsolete version) |who=N/A |what=Runs <code>scap clean auto</code> }} {{Deployment calendar event card |when=2026-08-31 23:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-31 23:00 SF |length=0.5 |window=Primary database switchover |who={{ircnick|marostegui|Manuel Arostegui}}, {{ircnick|cezmunsta|Ceri Williams}}, {{ircnick|federico3|Federico Ceratto}} |what=Held deployment window for database primary masters maintenance }} ==={{Deployment_day|date=2026-09-01}}=== {{Deployment calendar event card |when=2026-09-01 00:00 SF |length=1 |window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-09-01 03:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-09-01 05:00 SF |length=1 |window=Mobileapps/RESTBase/Wikifeeds |who=Content Transform Team |what=Content transform team node services (mobileapps/wikifeeds) }} {{Deployment calendar event card |when=2026-09-01 06:00 SF |length=1 |window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Lucas_WMDE|Lucas}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-09-01 07:00 SF |length=0.5 |window=Test Kitchen UI Deployment Window |who=Experimentation Platform Team |what=Deployment of Test Kitchen UI (fka MPIC) }} {{Deployment calendar event card |when=2026-09-01 07:30 SF |length=0.5 |window=Test Kitchen Experiment Deployment Window |who=Test Kitchen |what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]]. }} {{Deployment calendar event card |when=2026-09-01 08:00 SF |length=1 |window=SRE Collaboration Services office hours |who={{ircnick|jelto|Jelto}}, {{ircnick|arnoldokoth|Arnold}}, {{ircnick|mutante|Daniel}}, {{ircnick|arnaudb|Arnaud}} |what=Services including Gerrit, Phorge (Phabricator), GitLab }} {{Deployment calendar event card |when=2026-09-01 09:00 SF |length=1 |window=[[Puppet request window]]<br/><small>'''(Max 6 patches)'''</small> |who={{ircnick|jhathaway|JHathaway}}, {{ircnick|rzl|Reuven}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to Puppet change'' }} {{Deployment calendar event card |when=2026-09-01 10:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-09-01 11:00 SF |length=2 |window=MediaWiki train - Utc-7+Utc-0 Version |who={{ircnick|dancy|Ahmon}}, {{ircnick|hashar|Antoine}} |what=[[mw:MediaWiki 1.47/Roadmap#Schedule for the deployments|1.47 schedule]] {{DeployOneWeekMini|1.47.0-wmf.17->1.47.0-wmf.18|1.47.0-wmf.17|1.47.0-wmf.17}} * group0 to [[mw:MediaWiki_1.47/wmf.18|1.47.0-wmf.18]] * '''Blockers: {{phabricator|T430837}}''' }} {{Deployment calendar event card |when=2026-09-01 13:00 SF |length=1 |window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-09-01 14:00 SF |length=1 |window=Readers deployment window |who=Readers |what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start }} {{Deployment calendar event card |when=2026-09-01 23:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} ==={{Deployment_day|date=2026-09-02}}=== {{Deployment calendar event card |when=2026-09-02 00:00 SF |length=1 |window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-09-02 01:00 SF |length=2 |window=MediaWiki train - Utc-7+Utc-0 Version (secondary timeslot) |who={{ircnick|dancy|Ahmon}}, {{ircnick|hashar|Antoine}} |what=[[mw:MediaWiki 1.47/Roadmap#Schedule for the deployments|1.47 schedule]] {{DeployOneWeekMini|1.47.0-wmf.18|1.47.0-wmf.17->1.47.0-wmf.18|1.47.0-wmf.17}} * group1 to [[mw:MediaWiki_1.47/wmf.18|1.47.0-wmf.18]] * '''Blockers: {{phabricator|T430837}}''' }} {{Deployment calendar event card |when=2026-09-02 03:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-09-02 04:00 SF |length=1 |window=[[mw:Services|Services]] – [[Citoid]] / [[Zotero]] |who=Marielle ({{ircnick|mvolz}}) |what=See [[mw:Citoid|Citoid]] }} {{Deployment calendar event card |when=2026-09-02 06:00 SF |length=1 |window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-09-02 07:00 SF |length=1 |window=Wikifunctions Services UTC Afternoon |who=Abstract Wikipedia team (Africa, Europe, Eastern Americas) |what=Wikifunctions back-end k8s services }} {{Deployment calendar event card |when=2026-09-02 07:30 SF |length=0.5 |window=Test Kitchen Experiment Deployment Window |who=Test Kitchen |what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]]. }} {{Deployment calendar event card |when=2026-09-02 10:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-09-02 11:00 SF |length=2 |window=MediaWiki train - Utc-7+Utc-0 Version |who={{ircnick|dancy|Ahmon}}, {{ircnick|hashar|Antoine}} |what=[[mw:MediaWiki 1.47/Roadmap#Schedule for the deployments|1.47 schedule]] {{DeployOneWeekMini|1.47.0-wmf.18|1.47.0-wmf.17->1.47.0-wmf.18|1.47.0-wmf.17}} * group1 to [[mw:MediaWiki_1.47/wmf.18|1.47.0-wmf.18]] * '''Blockers: {{phabricator|T430837}}''' }} {{Deployment calendar event card |when=2026-09-02 13:00 SF |length=1 |window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-09-02 14:00 SF |length=1 |window=Wikifunctions Services UTC Late |who=Abstract Wikipedia team (North and South America) |what=Wikifunctions back-end k8s services }} {{Deployment calendar event card |when=2026-09-02 15:00 SF |length=1 |window=Readers deployment window |who=Readers |what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start }} {{Deployment calendar event card |when=2026-09-02 23:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-09-02 23:00 SF |length=0.5 |window=Primary database switchover |who={{ircnick|marostegui|Manuel Arostegui}}, {{ircnick|cezmunsta|Ceri Williams}}, {{ircnick|federico3|Federico Ceratto}} |what=Held deployment window for database primary masters maintenance }} ==={{Deployment_day|date=2026-09-03}}=== {{Deployment calendar event card |when=2026-09-03 00:00 SF |length=1 |window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-09-03 01:00 SF |length=2 |window=MediaWiki train - Utc-7+Utc-0 Version (secondary timeslot) |who={{ircnick|dancy|Ahmon}}, {{ircnick|hashar|Antoine}} |what=[[mw:MediaWiki 1.47/Roadmap#Schedule for the deployments|1.47 schedule]] {{DeployOneWeekMini|1.47.0-wmf.18|1.47.0-wmf.18|1.47.0-wmf.17->1.47.0-wmf.18}} * group2 to [[mw:MediaWiki_1.47/wmf.18|1.47.0-wmf.18]] * '''Blockers: {{phabricator|T430837}}''' }} {{Deployment calendar event card |when=2026-09-03 03:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-09-03 05:00 SF |length=1 |window=Mobileapps/RESTBase/Wikifeeds |who=Content Transform Team |what=Content transform team node services (mobileapps/wikifeeds) }} {{Deployment calendar event card |when=2026-09-03 06:00 SF |length=1 |window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Lucas_WMDE|Lucas}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-09-03 07:30 SF |length=0.5 |window=Test Kitchen Experiment Deployment Window |who=Test Kitchen |what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]]. }} {{Deployment calendar event card |when=2026-09-03 09:00 SF |length=1 |window=[[Puppet request window]]<br/><small>'''(Max 6 patches)'''</small> |who={{ircnick|jhathaway|JHathaway}}, {{ircnick|rzl|Reuven}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to Puppet change'' }} {{Deployment calendar event card |when=2026-09-03 10:00 SF |length=1 |window=Cloud Services/Technical Documentation weekly deploy (Toolhub, Developer portal, Striker) |who={{ircnick|bd808}} |what=... }} {{Deployment calendar event card |when=2026-09-03 10:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-09-03 11:00 SF |length=2 |window=MediaWiki train - Utc-7+Utc-0 Version |who={{ircnick|dancy|Ahmon}}, {{ircnick|hashar|Antoine}} |what=[[mw:MediaWiki 1.47/Roadmap#Schedule for the deployments|1.47 schedule]] {{DeployOneWeekMini|1.47.0-wmf.18|1.47.0-wmf.18|1.47.0-wmf.17->1.47.0-wmf.18}} * group2 to [[mw:MediaWiki_1.47/wmf.18|1.47.0-wmf.18]] * '''Blockers: {{phabricator|T430837}}''' }} {{Deployment calendar event card |when=2026-09-03 13:00 SF |length=1 |window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-09-03 14:00 SF |length=1 |window=Readers deployment window |who=Readers |what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start }} {{Deployment calendar event card |when=2026-09-03 23:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} ==={{Deployment_day|date=2026-09-04}}=== {{Deployment calendar event card |when=2026-09-04 00:00 SF |length=24 |window=No deploys all day! See [[Deployments/Emergencies]] if things are broken. |who= |what=No Deploys }} {{Deployment calendar event card |when=2026-09-04 04:00 SF |length=0.5 |window=GitLab version upgrades |who={{ircnick|jelto|Jelto}}, {{ircnick|arnoldokoth|Arnold}}, {{ircnick|mutante|Daniel}}, {{ircnick|arnaudb|Arnaud}} |what=GitLab version upgrades }} ==={{Deployment_day|date=2026-09-05}}=== {{Deployment calendar event card |when=2026-09-05 00:00 SF |length=24 |window=No deploys all day! See [[Deployments/Emergencies]] if things are broken. |who= |what=No Deploys }} kh1t5etkx3hvi9hiivki7uv0r8lpty5 2450674 2450673 2026-08-23T05:04:28Z ScheduleDeploymentBot 37566 Add [[gerrit:1328315]] to Monday, August 24 UTC afternoon backport window 2450674 wikitext text/x-wiki {{Navigation MediaWiki deployment}} This page tracks '''upcoming''' '''deployments''' of software to the [[:m:Special:SiteMatrix|Wikimedia Foundation servers]]. == Getting started == Ensure you joined the {{irc|wikimedia-operations}} IRC channel as all deployment-related communications happen there. If you need help, contact [[:mw:Wikimedia Release Engineering Team|Release Engineering]] on IRC at {{irc|wikimedia-releng}}; and ping Tyler (<code>thcipriani</code>). * '''MediaWiki is deployed weekly''' through the [[/Train|Deployment Train]]. Other services follow their own schedule. * '''Times are pinned to San Francisco''', thus the UTC time changes in March and November per [[:en:Daylight saving time in the United States|DST]]. * '''Prefer regular [[Backport windows]]''' over adding new windows. To request deployment of a config change or backport, add your username and Gerrit URL to one of the backport windows on this page. You must be online in #wikimedia-operations on IRC during your deployment and install [[WikimediaDebug]] ahead of time. The #wikimedia-operations channel requires you to [[:m:IRC/Instructions#Register your nickname, identify, and enforce|register your nickname]] before you can join. ** You can use the '''backport scheduling tool''' to more easily edit this page: <div style="text-align: center; margin: 1em 0">{{Clickable button 2|:toollabs:schedule-deployment|Schedule a backport|class=mw-ui-progressive}}</div> * Tasks that meet [[/Inclusion criteria|Inclusion criteria]] '''require their own windows''', which includes long-running tasks. '''Schedule more time''' than you think you need to account for delays and set backs, we recommend one hour for most tasks. **To create or modify a recurring deploy window, send a patchset to [[:gitlab:repos/releng/release/-/blob/main/make-deployment-calendar/deployments-calendar.yaml|deployments-calendar.yaml file]] in <code>repos/releng/release.git</code>. **To create an one-off window, simply edit this page accordingly ** '''Announce''' changes to the [[mail:ops|ops mailing list]] ahead of time if you anticipate or are uncertain about noticeable impacts to database load, HTTP caching, or the introduction of new cookies. ** '''Announce''' deployments of major features to the community via [[:m:Tech/News/Next|Tech News]] and/or via other [[:mw:Wikimedia_Product_Guidance/Communication_channels|Product communication channels]]. * '''Something went wrong?''' See [[Incident response]]. Is there a user-impacting problem? Communicate in the {{irc|wikimedia-operations}} IRC channel. If there is a Phabricator task, ensure [[:phab:tag/wikimedia-incident/|#Wikimedia-Incident]] is tagged, and consider setting the [[:mw:Phabricator/Project_management#Priority_levels|Unbreak Now]] priority. __TOC__ {{anchor|Next Week|Near Term|Near term|Near-term}}{{clear}} [[Category:Deployment]] {{Note|content=Subscribe in Google Calendar via <code>wikimedia.org_rudis09ii2mm5fk4hgdjeh1u64@group.calendar.google.com</code>.<br>This may not include one-off windows. '''If there are differences, then the wiki page is canonical and correct'''.}} ==Week of August 24== ==={{Deployment_day|date=2026-08-23}}=== {{Deployment calendar event card |when=2026-08-23 00:00 SF |length=24 |window=No deploys all day! See [[Deployments/Emergencies]] if things are broken. |who= |what=No Deploys }} ==={{Deployment_day|date=2026-08-24}}=== {{Deployment calendar event card |when=2026-08-24 00:00 SF |length=1 |window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}} |what={{ircnick|Hide_on_rosie|Hide_on_rosie}} {{deploy|type=config|gerrit=1307270|title=InitialiseSettings: change enwiki extendedconfirmed autopromote settings|status=}} - {{phabricator|T431060}} {{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-24 03:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-24 06:00 SF |length=1 |window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Lucas_WMDE|Lucas}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}} |what={{ircnick|MichaelG_WMF|Michael Grosse (dev @ WMF Growth team)}} {{deploy|type=config|gerrit=1325527|title=Growth: Drop now unused config for benefits block|status=}} - {{phabricator|T433783}} {{ircnick|sadiya_wmde28|Sadiya Halimat Mohammed}} {{deploy|type=config|gerrit=1326848|title=Disable Wikidata Bridge on Catalan Wikipedia|status=}} - {{phabricator|T433713}} {{ircnick|anzx|anzx}} {{deploy|type=config|gerrit=1328310|title=Lift IP cap for Mapudungun editathon on 2026-08-29|status=}} - {{phabricator|T435523}} {{deploy|type=config|gerrit=1328315|title=arwikiquote: update wordmark|status=}} - {{phabricator|T435505}} {{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-24 07:30 SF |length=0.5 |window=Test Kitchen Experiment Deployment Window |who=Test Kitchen |what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]]. }} {{Deployment calendar event card |when=2026-08-24 08:30 SF |length=0.5 |window=Wikimedia Portals Update |who={{ircnick|jan_drewniak|Jan Drewniak}} |what=Weekly window for the portals page: https://www.wikipedia.org/ }} {{Deployment calendar event card |when=2026-08-24 10:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. {{ircnick|bd808}} * Shellbox updates to pick up new Pygments release ([[phab:T421653|T421653]]) }} {{Deployment calendar event card |when=2026-08-24 10:00 SF |length=0.5 |window=Wikidata Query Service weekly deploy |who={{ircnick|ryankemper|Ryan}} |what=... }} {{Deployment calendar event card |when=2026-08-24 13:00 SF |length=1 |window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-24 14:00 SF |length=2 |window=Weekly Security deployment window |who={{ircnick|alexsanford|Alex}}, {{ircnick|Reedy|Sam}}, {{ircnick|sbassett|Scott}}, {{ircnick|Maryum|Maryum}}, {{ircnick|manfredi|Manfredi}} |what=Held deployment window for Security-team related deploys. }} {{Deployment calendar event card |when=2026-08-24 16:00 SF |length=1 |window=Readers deployment window |who=Readers |what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start }} {{Deployment calendar event card |when=2026-08-24 19:00 SF |length=1 |window=Automatic branching of MediaWiki, extensions, skins, and vendor – see [[Heterogeneous deployment/Train deploys]] |who=N/A |what=Branch <code>wmf/1.47.0-wmf.17</code> }} {{Deployment calendar event card |when=2026-08-24 20:00 SF |length=1 |window=Automatic deployment of MediaWiki, extensions, skins, and vendor to testwikis only – see [[Heterogeneous deployment/Train deploys]] |who=N/A |what=Deploy <code>wmf/1.47.0-wmf.17</code> to testwikis }} {{Deployment calendar event card |when=2026-08-24 21:00 SF |length=1 |window=Automatic removal of all obsolete MediaWiki versions from the deployment and bare metal servers (except the most-recent obsolete version) |who=N/A |what=Runs <code>scap clean auto</code> }} {{Deployment calendar event card |when=2026-08-24 23:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-24 23:00 SF |length=0.5 |window=Primary database switchover |who={{ircnick|marostegui|Manuel Arostegui}}, {{ircnick|Amir1|Amir}}, {{ircnick|federico3|Federico Ceratto}} |what=Held deployment window for database primary masters maintenance }} ==={{Deployment_day|date=2026-08-25}}=== {{Deployment calendar event card |when=2026-08-25 00:00 SF |length=1 |window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-25 01:00 SF |length=2 |window=MediaWiki train - Utc-0 Version |who={{ircnick|hashar|Antoine}}, {{ircnick|andre|Andre}} |what=[[mw:MediaWiki 1.47/Roadmap#Schedule for the deployments|1.47 schedule]] {{DeployOneWeekMini|1.47.0-wmf.16->1.47.0-wmf.17|1.47.0-wmf.16|1.47.0-wmf.16}} * group0 to [[mw:MediaWiki_1.47/wmf.17|1.47.0-wmf.17]] * '''Blockers: {{phabricator|T430836}}''' }} {{Deployment calendar event card |when=2026-08-25 03:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-25 05:00 SF |length=1 |window=Mobileapps/RESTBase/Wikifeeds |who=Content Transform Team |what=Content transform team node services (mobileapps/wikifeeds) }} {{Deployment calendar event card |when=2026-08-25 06:00 SF |length=1 |window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Lucas_WMDE|Lucas}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-25 07:00 SF |length=0.5 |window=Test Kitchen UI Deployment Window |who=Experimentation Platform Team |what=Deployment of Test Kitchen UI (fka MPIC) }} {{Deployment calendar event card |when=2026-08-25 07:30 SF |length=0.5 |window=Test Kitchen Experiment Deployment Window |who=Test Kitchen |what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]]. }} {{Deployment calendar event card |when=2026-08-25 08:00 SF |length=1 |window=SRE Collaboration Services office hours |who={{ircnick|jelto|Jelto}}, {{ircnick|arnoldokoth|Arnold}}, {{ircnick|mutante|Daniel}}, {{ircnick|arnaudb|Arnaud}} |what=Services including Gerrit, Phorge (Phabricator), GitLab }} {{Deployment calendar event card |when=2026-08-25 09:00 SF |length=1 |window=[[Puppet request window]]<br/><small>'''(Max 6 patches)'''</small> |who={{ircnick|jhathaway|JHathaway}}, {{ircnick|rzl|Reuven}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to Puppet change'' }} {{Deployment calendar event card |when=2026-08-25 10:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-25 13:00 SF |length=1 |window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-25 14:00 SF |length=1 |window=Readers deployment window |who=Readers |what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start }} {{Deployment calendar event card |when=2026-08-25 23:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} ==={{Deployment_day|date=2026-08-26}}=== {{Deployment calendar event card |when=2026-08-26 00:00 SF |length=1 |window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-26 01:00 SF |length=2 |window=MediaWiki train - Utc-0 Version |who={{ircnick|hashar|Antoine}}, {{ircnick|andre|Andre}} |what=[[mw:MediaWiki 1.47/Roadmap#Schedule for the deployments|1.47 schedule]] {{DeployOneWeekMini|1.47.0-wmf.17|1.47.0-wmf.16->1.47.0-wmf.17|1.47.0-wmf.16}} * group1 to [[mw:MediaWiki_1.47/wmf.17|1.47.0-wmf.17]] * '''Blockers: {{phabricator|T430836}}''' }} {{Deployment calendar event card |when=2026-08-26 03:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-26 04:00 SF |length=1 |window=[[mw:Services|Services]] – [[Citoid]] / [[Zotero]] |who=Marielle ({{ircnick|mvolz}}) |what=See [[mw:Citoid|Citoid]] }} {{Deployment calendar event card |when=2026-08-26 06:00 SF |length=1 |window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-26 07:00 SF |length=1 |window=Wikifunctions Services UTC Afternoon |who=Abstract Wikipedia team (Africa, Europe, Eastern Americas) |what=Wikifunctions back-end k8s services }} {{Deployment calendar event card |when=2026-08-26 07:30 SF |length=0.5 |window=Test Kitchen Experiment Deployment Window |who=Test Kitchen |what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]]. }} {{Deployment calendar event card |when=2026-08-26 10:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-26 13:00 SF |length=1 |window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-26 14:00 SF |length=1 |window=Wikifunctions Services UTC Late |who=Abstract Wikipedia team (North and South America) |what=Wikifunctions back-end k8s services }} {{Deployment calendar event card |when=2026-08-26 15:00 SF |length=1 |window=Readers deployment window |who=Readers |what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start }} {{Deployment calendar event card |when=2026-08-26 23:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-26 23:00 SF |length=0.5 |window=Primary database switchover |who={{ircnick|marostegui|Manuel Arostegui}}, {{ircnick|Amir1|Amir}}, {{ircnick|federico3|Federico Ceratto}} |what=Held deployment window for database primary masters maintenance }} ==={{Deployment_day|date=2026-08-27}}=== {{Deployment calendar event card |when=2026-08-27 00:00 SF |length=1 |window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-27 01:00 SF |length=2 |window=MediaWiki train - Utc-0 Version |who={{ircnick|hashar|Antoine}}, {{ircnick|andre|Andre}} |what=[[mw:MediaWiki 1.47/Roadmap#Schedule for the deployments|1.47 schedule]] {{DeployOneWeekMini|1.47.0-wmf.17|1.47.0-wmf.17|1.47.0-wmf.16->1.47.0-wmf.17}} * group2 to [[mw:MediaWiki_1.47/wmf.17|1.47.0-wmf.17]] * '''Blockers: {{phabricator|T430836}}''' }} {{Deployment calendar event card |when=2026-08-27 03:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-27 05:00 SF |length=1 |window=Mobileapps/RESTBase/Wikifeeds |who=Content Transform Team |what=Content transform team node services (mobileapps/wikifeeds) }} {{Deployment calendar event card |when=2026-08-27 06:00 SF |length=1 |window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Lucas_WMDE|Lucas}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-27 07:30 SF |length=0.5 |window=Test Kitchen Experiment Deployment Window |who=Test Kitchen |what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]]. }} {{Deployment calendar event card |when=2026-08-27 09:00 SF |length=1 |window=[[Puppet request window]]<br/><small>'''(Max 6 patches)'''</small> |who={{ircnick|jhathaway|JHathaway}}, {{ircnick|rzl|Reuven}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to Puppet change'' }} {{Deployment calendar event card |when=2026-08-27 10:00 SF |length=1 |window=Cloud Services/Technical Documentation weekly deploy (Toolhub, Developer portal, Striker) |who={{ircnick|bd808}} |what=... }} {{Deployment calendar event card |when=2026-08-27 10:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-27 13:00 SF |length=1 |window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-27 14:00 SF |length=1 |window=Readers deployment window |who=Readers |what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start }} {{Deployment calendar event card |when=2026-08-27 23:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} ==={{Deployment_day|date=2026-08-28}}=== {{Deployment calendar event card |when=2026-08-28 00:00 SF |length=24 |window=No deploys all day! See [[Deployments/Emergencies]] if things are broken. |who= |what=No Deploys }} {{Deployment calendar event card |when=2026-08-28 04:00 SF |length=0.5 |window=GitLab version upgrades |who={{ircnick|jelto|Jelto}}, {{ircnick|arnoldokoth|Arnold}}, {{ircnick|mutante|Daniel}}, {{ircnick|arnaudb|Arnaud}} |what=GitLab version upgrades }} ==={{Deployment_day|date=2026-08-29}}=== {{Deployment calendar event card |when=2026-08-29 00:00 SF |length=24 |window=No deploys all day! See [[Deployments/Emergencies]] if things are broken. |who= |what=No Deploys }} ==Week of August 31== ==={{Deployment_day|date=2026-08-30}}=== {{Deployment calendar event card |when=2026-08-30 00:00 SF |length=24 |window=No deploys all day! See [[Deployments/Emergencies]] if things are broken. |who= |what=No Deploys }} ==={{Deployment_day|date=2026-08-31}}=== {{Deployment calendar event card |when=2026-08-31 00:00 SF |length=1 |window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-31 03:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-31 06:00 SF |length=1 |window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Lucas_WMDE|Lucas}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-31 07:30 SF |length=0.5 |window=Test Kitchen Experiment Deployment Window |who=Test Kitchen |what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]]. }} {{Deployment calendar event card |when=2026-08-31 08:30 SF |length=0.5 |window=Wikimedia Portals Update |who={{ircnick|jan_drewniak|Jan Drewniak}} |what=Weekly window for the portals page: https://www.wikipedia.org/ }} {{Deployment calendar event card |when=2026-08-31 10:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-31 10:00 SF |length=0.5 |window=Wikidata Query Service weekly deploy |who={{ircnick|ryankemper|Ryan}} |what=... }} {{Deployment calendar event card |when=2026-08-31 13:00 SF |length=1 |window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-31 14:00 SF |length=2 |window=Weekly Security deployment window |who={{ircnick|alexsanford|Alex}}, {{ircnick|Reedy|Sam}}, {{ircnick|sbassett|Scott}}, {{ircnick|Maryum|Maryum}}, {{ircnick|manfredi|Manfredi}} |what=Held deployment window for Security-team related deploys. }} {{Deployment calendar event card |when=2026-08-31 16:00 SF |length=1 |window=Readers deployment window |who=Readers |what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start }} {{Deployment calendar event card |when=2026-08-31 19:00 SF |length=1 |window=Automatic branching of MediaWiki, extensions, skins, and vendor – see [[Heterogeneous deployment/Train deploys]] |who=N/A |what=Branch <code>wmf/1.47.0-wmf.18</code> }} {{Deployment calendar event card |when=2026-08-31 20:00 SF |length=1 |window=Automatic deployment of MediaWiki, extensions, skins, and vendor to testwikis only – see [[Heterogeneous deployment/Train deploys]] |who=N/A |what=Deploy <code>wmf/1.47.0-wmf.18</code> to testwikis }} {{Deployment calendar event card |when=2026-08-31 21:00 SF |length=1 |window=Automatic removal of all obsolete MediaWiki versions from the deployment and bare metal servers (except the most-recent obsolete version) |who=N/A |what=Runs <code>scap clean auto</code> }} {{Deployment calendar event card |when=2026-08-31 23:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-31 23:00 SF |length=0.5 |window=Primary database switchover |who={{ircnick|marostegui|Manuel Arostegui}}, {{ircnick|cezmunsta|Ceri Williams}}, {{ircnick|federico3|Federico Ceratto}} |what=Held deployment window for database primary masters maintenance }} ==={{Deployment_day|date=2026-09-01}}=== {{Deployment calendar event card |when=2026-09-01 00:00 SF |length=1 |window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-09-01 03:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-09-01 05:00 SF |length=1 |window=Mobileapps/RESTBase/Wikifeeds |who=Content Transform Team |what=Content transform team node services (mobileapps/wikifeeds) }} {{Deployment calendar event card |when=2026-09-01 06:00 SF |length=1 |window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Lucas_WMDE|Lucas}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-09-01 07:00 SF |length=0.5 |window=Test Kitchen UI Deployment Window |who=Experimentation Platform Team |what=Deployment of Test Kitchen UI (fka MPIC) }} {{Deployment calendar event card |when=2026-09-01 07:30 SF |length=0.5 |window=Test Kitchen Experiment Deployment Window |who=Test Kitchen |what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]]. }} {{Deployment calendar event card |when=2026-09-01 08:00 SF |length=1 |window=SRE Collaboration Services office hours |who={{ircnick|jelto|Jelto}}, {{ircnick|arnoldokoth|Arnold}}, {{ircnick|mutante|Daniel}}, {{ircnick|arnaudb|Arnaud}} |what=Services including Gerrit, Phorge (Phabricator), GitLab }} {{Deployment calendar event card |when=2026-09-01 09:00 SF |length=1 |window=[[Puppet request window]]<br/><small>'''(Max 6 patches)'''</small> |who={{ircnick|jhathaway|JHathaway}}, {{ircnick|rzl|Reuven}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to Puppet change'' }} {{Deployment calendar event card |when=2026-09-01 10:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-09-01 11:00 SF |length=2 |window=MediaWiki train - Utc-7+Utc-0 Version |who={{ircnick|dancy|Ahmon}}, {{ircnick|hashar|Antoine}} |what=[[mw:MediaWiki 1.47/Roadmap#Schedule for the deployments|1.47 schedule]] {{DeployOneWeekMini|1.47.0-wmf.17->1.47.0-wmf.18|1.47.0-wmf.17|1.47.0-wmf.17}} * group0 to [[mw:MediaWiki_1.47/wmf.18|1.47.0-wmf.18]] * '''Blockers: {{phabricator|T430837}}''' }} {{Deployment calendar event card |when=2026-09-01 13:00 SF |length=1 |window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-09-01 14:00 SF |length=1 |window=Readers deployment window |who=Readers |what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start }} {{Deployment calendar event card |when=2026-09-01 23:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} ==={{Deployment_day|date=2026-09-02}}=== {{Deployment calendar event card |when=2026-09-02 00:00 SF |length=1 |window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-09-02 01:00 SF |length=2 |window=MediaWiki train - Utc-7+Utc-0 Version (secondary timeslot) |who={{ircnick|dancy|Ahmon}}, {{ircnick|hashar|Antoine}} |what=[[mw:MediaWiki 1.47/Roadmap#Schedule for the deployments|1.47 schedule]] {{DeployOneWeekMini|1.47.0-wmf.18|1.47.0-wmf.17->1.47.0-wmf.18|1.47.0-wmf.17}} * group1 to [[mw:MediaWiki_1.47/wmf.18|1.47.0-wmf.18]] * '''Blockers: {{phabricator|T430837}}''' }} {{Deployment calendar event card |when=2026-09-02 03:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-09-02 04:00 SF |length=1 |window=[[mw:Services|Services]] – [[Citoid]] / [[Zotero]] |who=Marielle ({{ircnick|mvolz}}) |what=See [[mw:Citoid|Citoid]] }} {{Deployment calendar event card |when=2026-09-02 06:00 SF |length=1 |window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-09-02 07:00 SF |length=1 |window=Wikifunctions Services UTC Afternoon |who=Abstract Wikipedia team (Africa, Europe, Eastern Americas) |what=Wikifunctions back-end k8s services }} {{Deployment calendar event card |when=2026-09-02 07:30 SF |length=0.5 |window=Test Kitchen Experiment Deployment Window |who=Test Kitchen |what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]]. }} {{Deployment calendar event card |when=2026-09-02 10:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-09-02 11:00 SF |length=2 |window=MediaWiki train - Utc-7+Utc-0 Version |who={{ircnick|dancy|Ahmon}}, {{ircnick|hashar|Antoine}} |what=[[mw:MediaWiki 1.47/Roadmap#Schedule for the deployments|1.47 schedule]] {{DeployOneWeekMini|1.47.0-wmf.18|1.47.0-wmf.17->1.47.0-wmf.18|1.47.0-wmf.17}} * group1 to [[mw:MediaWiki_1.47/wmf.18|1.47.0-wmf.18]] * '''Blockers: {{phabricator|T430837}}''' }} {{Deployment calendar event card |when=2026-09-02 13:00 SF |length=1 |window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-09-02 14:00 SF |length=1 |window=Wikifunctions Services UTC Late |who=Abstract Wikipedia team (North and South America) |what=Wikifunctions back-end k8s services }} {{Deployment calendar event card |when=2026-09-02 15:00 SF |length=1 |window=Readers deployment window |who=Readers |what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start }} {{Deployment calendar event card |when=2026-09-02 23:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-09-02 23:00 SF |length=0.5 |window=Primary database switchover |who={{ircnick|marostegui|Manuel Arostegui}}, {{ircnick|cezmunsta|Ceri Williams}}, {{ircnick|federico3|Federico Ceratto}} |what=Held deployment window for database primary masters maintenance }} ==={{Deployment_day|date=2026-09-03}}=== {{Deployment calendar event card |when=2026-09-03 00:00 SF |length=1 |window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-09-03 01:00 SF |length=2 |window=MediaWiki train - Utc-7+Utc-0 Version (secondary timeslot) |who={{ircnick|dancy|Ahmon}}, {{ircnick|hashar|Antoine}} |what=[[mw:MediaWiki 1.47/Roadmap#Schedule for the deployments|1.47 schedule]] {{DeployOneWeekMini|1.47.0-wmf.18|1.47.0-wmf.18|1.47.0-wmf.17->1.47.0-wmf.18}} * group2 to [[mw:MediaWiki_1.47/wmf.18|1.47.0-wmf.18]] * '''Blockers: {{phabricator|T430837}}''' }} {{Deployment calendar event card |when=2026-09-03 03:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-09-03 05:00 SF |length=1 |window=Mobileapps/RESTBase/Wikifeeds |who=Content Transform Team |what=Content transform team node services (mobileapps/wikifeeds) }} {{Deployment calendar event card |when=2026-09-03 06:00 SF |length=1 |window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Lucas_WMDE|Lucas}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-09-03 07:30 SF |length=0.5 |window=Test Kitchen Experiment Deployment Window |who=Test Kitchen |what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]]. }} {{Deployment calendar event card |when=2026-09-03 09:00 SF |length=1 |window=[[Puppet request window]]<br/><small>'''(Max 6 patches)'''</small> |who={{ircnick|jhathaway|JHathaway}}, {{ircnick|rzl|Reuven}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to Puppet change'' }} {{Deployment calendar event card |when=2026-09-03 10:00 SF |length=1 |window=Cloud Services/Technical Documentation weekly deploy (Toolhub, Developer portal, Striker) |who={{ircnick|bd808}} |what=... }} {{Deployment calendar event card |when=2026-09-03 10:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-09-03 11:00 SF |length=2 |window=MediaWiki train - Utc-7+Utc-0 Version |who={{ircnick|dancy|Ahmon}}, {{ircnick|hashar|Antoine}} |what=[[mw:MediaWiki 1.47/Roadmap#Schedule for the deployments|1.47 schedule]] {{DeployOneWeekMini|1.47.0-wmf.18|1.47.0-wmf.18|1.47.0-wmf.17->1.47.0-wmf.18}} * group2 to [[mw:MediaWiki_1.47/wmf.18|1.47.0-wmf.18]] * '''Blockers: {{phabricator|T430837}}''' }} {{Deployment calendar event card |when=2026-09-03 13:00 SF |length=1 |window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-09-03 14:00 SF |length=1 |window=Readers deployment window |who=Readers |what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start }} {{Deployment calendar event card |when=2026-09-03 23:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} ==={{Deployment_day|date=2026-09-04}}=== {{Deployment calendar event card |when=2026-09-04 00:00 SF |length=24 |window=No deploys all day! See [[Deployments/Emergencies]] if things are broken. |who= |what=No Deploys }} {{Deployment calendar event card |when=2026-09-04 04:00 SF |length=0.5 |window=GitLab version upgrades |who={{ircnick|jelto|Jelto}}, {{ircnick|arnoldokoth|Arnold}}, {{ircnick|mutante|Daniel}}, {{ircnick|arnaudb|Arnaud}} |what=GitLab version upgrades }} ==={{Deployment_day|date=2026-09-05}}=== {{Deployment calendar event card |when=2026-09-05 00:00 SF |length=24 |window=No deploys all day! See [[Deployments/Emergencies]] if things are broken. |who= |what=No Deploys }} s20ysg3a4frtz7mfl84d5ivow8355ki PartMan 0 4235 2450664 2449599 2026-08-22T19:08:56Z Quiddity 1884 fix "here" in/as link label; {{clear}} for image float 2450664 wikitext text/x-wiki This page is about partitioning with preseed. See [[Ubuntu installer]] for other parts of the installation process and preseed options. == Background Info == Debian preseed's partman options are an incomprehensible automatic partitioning language.<nowiki>{{fact}}</nowiki> You can now test your new partman recipes in a virtual environment with [https://gitlab.wikimedia.org/repos/sre/preseed-test preseed-test], which is quicker and easier. == Standard recipes / disk layout == Most hosts are installed via ''standard'' partman recipes found in modules/install_server/files/autoinstall/partman/ . There are two main categories of standard recipes: software and hardware RAID, plus VMs for which including <code>partman/flat.cfg virtual.cfg</code> in <code>netboot.cfg</code> is sufficient. The netboot.cfg file is generated by puppet, from configuration located in <code>modules/profile/data/profile/installserver/preseed.yaml</code>. This was done to avoid typos creep through the file. === Software raid === The host has a number <code>N</code> of (equal size) devices and will be installed with the chosen raid level <code>L</code> (0, 1, 10) plus LVM on top using volume group <code>vg0</code>. The need to be added to preseed.yaml in the following manner for a a group of servers called fooserver1001-server1004: <pre> 'fooserver100[1-4]': - partman/standard.cfg - partman/raidL-Ndev.cfg </pre> A special case is software raid0 for which a common recipe snippet needs to be included: <pre> 'barserver*': - partman/standard.cfg - partman/raid0.cfg - partman/raid0-Ndev.cfg </pre> === Hardware raid === In this case the host uses on board hardware raid and typically a single block device is presented to the operating system. The host will have LVM on top of the hardware raid, similar to software raid. Intended usage of the recipe is: <pre> 'fooserver*': - partman/standard.cfg - partman/hwraid-1dev.cfg </pre> === LVM layout === The standard recipes provide a <code>/</code> partition of 80GB, <code>swap</code> 1GB and a <code>/srv</code> taking up to 80% of the remaining space on the VG. The space is not allocated all at installation time to provide some leeway in emergencies and a little bit of headroom in capacity planning. The filesystems can be expanded online at any time with e.g. <code>lvextend --resizefs</code>. See also [[Monitoring/Disk space#Standard partman recipes with LVM|documentation on extending LVM]]. == Non-standard recipes == While a majority of cases are covered by standard recipes, some special cases remain. These are typically storage hosts of some kind and their recipes live in <code>partman/custom/</code>. New hosts/use cases should strive to use the standard recipes if at all possible, if not consider whether a custom recipe can be adapted to your use case. == Help Files == These files are in the Debian packages, and give some explanation * [[PartManAuto]] * [[PartManAutoRaid]] == Important note about RAID == From https://www.debian.org/releases/squeeze/example-preseed.txt When you use multiraid, you specify '''the partitions layout that each disk will get'''. Later on you'll be able to define them via partman-auto-raid/recipe. <pre> ## Partitioning using RAID # The method should be set to "raid". #d-i partman-auto/method string raid # Specify the disks to be partitioned. They will all get the same layout, # so this will only work if the disks are the same size. #d-i partman-auto/disk string /dev/sda /dev/sdb # Next you need to specify the physical partitions that will be used. #d-i partman-auto/expert_recipe string \ # multiraid :: \ # 1000 5000 4000 raid \ # $primary{ } method{ raid } \ # . \ # 64 512 300% raid \ # method{ raid } \ # . \ # 500 10000 1000000000 raid \ # method{ raid } \ # </pre> In the above example, partman is instructed to create three partitions on each disk, so you'll probably get: sda1/sda2/sda3 and sdb1/sdb2/sdb3 to play with in partman-auto-raid/recipe (to create raid arrays). == Important note about how partman calculates the final size of the partitions == From https://www.bishnet.net/tim/blog/2015/01/29/understanding-partman-autoexpert_recipe/ I recently created the mw-raid1-lvm partman recipe for new app-servers, ending up with the following config: <pre> # Define physical partitions d-i partman-auto/expert_recipe string multiraid :: \ 1000 1000 -1 raid \ method{ raid } \ $lvmignore{ } \ . \ 10000 71000 600000000 ext4 \ method{ format } \ format{ } \ use_filesystem{ } \ filesystem{ ext4 } \ lv_name{ root } \ $defaultignore{ } \ $lvmok{ } \ mountpoint{ / } \ . \ 10000 30000 200000000 ext4 \ method{ format } \ format{ } \ use_filesystem{ } \ filesystem{ ext4 } \ lv_name{ srv } \ $defaultignore{ } \ $lvmok{ } \ mountpoint{ /srv } \ . \ 10000 27500 200000000 ext4 \ lv_name{ placeholder } \ $defaultignore{ } \ $lvmok{ } \ . # Parameters are: # <raidtype> <devcount> <sparecount> <fstype> <mountpoint> \ # <devices> <sparedevices> d-i partman-auto-raid/recipe string \ 1 2 0 lvm - \ /dev/sda1#/dev/sdb1 \ . </pre> In this case each host offers two 1TB disks, and I wanted to have a single software RAID1 device and LVM Volumes on top of it. The hard part was how to find the right values for min/priority/max in the expert-recipe. After a lot of unsuccessful tests with numbers that I thought to have some sense, I found the link at the beginning of this section that enlightened me. Let's start with the end result of the above example: <pre> elukey@mw1321:~$ df -h Filesystem Size Used Avail Use% Mounted on udev 10M 0 10M 0% /dev tmpfs 13G 9.1M 13G 1% /run /dev/dm-0 563G 7.8G 527G 2% / tmpfs 32G 0 32G 0% /dev/shm tmpfs 5.0M 0 5.0M 0% /run/lock tmpfs 32G 0 32G 0% /sys/fs/cgroup tmpfs 1.0G 0 1.0G 0% /var/lib/nginx /dev/mapper/mw1321--vg-srv 191G 8.0G 173G 5% /srv elukey@mw1321:~$ sudo lvs LV VG Attr LSize Pool Origin Data% Meta% Move Log Cpy%Sync Convert placeholder mw1321-vg -wi-a----- 166.04g root mw1321-vg -wi-ao---- 571.66g srv mw1321-vg -wi-ao---- 193.69g elukey@mw1321:~$ sudo pvs PV VG Fmt Attr PSize PFree /dev/md0 mw1321-vg lvm2 a-- 931.38g 0 </pre> What I originally wanted was a LVM volume for root with 600GB, a /srv one with 200GB and 200GB free (the placeholder volume, that is required since partman will allocate all the remaining space in the PV to the last volume defined). The end result is pretty close, but I had to do the following calculations: <pre> 10000 71000 600000000 ext4 # root 10000 30000 200000000 ext4 # srv 10000 27500 200000000 ext4 # placeholder </pre> Each LVM volume starts with a minimum size of 10G (the first value), so ~970GB are left to use (remember that the underlying PV is a sw RAID1 of 1TB usable space). In order to achieve my goal, the following GBs are needed: * 590 for root (~61% of the 970GB) * 190 for srv (~20% of the 970GB) * 190 for placeholder (~17.5% of the 970GB - this value was a mistake since the most correct one should have been 20%, but the last volume defined should get the remaining space anyway. In this case the srv volume gets a bit more space). These are exactly the values that you can obtain subtracting the priority value (the middle one) with the minimum one (the first one on the left). Please note that this is my current understanding of how Partman works, do not trust it blindly. == Config language == <code><owner> <question name> <question type> <value></code> <var><owner></var>: "d-i" which stands for Debian Installer. <var><question name></var>: partman-auto, partman-auto-raid and partman-auto-lvm are the packages that handle automatic partitioning of various types, and some of the questions (prompts) issued by them are described here: [http://d-i.alioth.debian.org/manual/en.i386/apbs04.html#preseed-partman] <var><question type></var> says what sort of value to expect (e.g. string, boolean, select (for a menu)...) <var><question value></var> this is where you put the answer that you would otherwise be entering interactively NOTE that after the question type you can put only one space or tab; any other whitespace will get stuffed in at the front of the value, which you probably don't want. == Dissecting a Semi-working configuration == (by semi-working, I mean that you must hit <enter> twice, but not answer any meaningful questions.) <pre> # Automatic software RAID 1 with LVM partitioning d-i partman-auto/method string raid </pre> This sentence makes partman know that it will be making a raid <pre> # Use the first two disks d-i partman-auto/disk string /dev/sda /dev/sdb </pre> after <code>d-i partman-auto/disk</code> you must put the list of hard drives that are going to be used <pre> # Define physical partitions d-i partman-auto/expert_recipe string \ multiraid :: \ 400000 1000 9500000 raid \ $primary{ } method{ raid } \ . \ 4000 1200 4100 linux-swap \ $primary{ } method{ swap } format{ } \ . </pre> An interesting thing about partman is that you can put everything on one line, or you can break lines using the character "\". Use a "." in between hard drives. * <code>$primary{ }</code> is needed to make the partition a primary partition. * <code>method { }</code> is used to tell it what type to format. You can use swap, raid, or format. * <code>format { }</code> tells partman to format the partition. Don't put this statement in a section that you will be using for the raid. use_filesystem{ } makes partman use a file system (don't know why this isn't done by the filesystem command) * <code>filesystem{ X }</code> use ext3, murderfs, xfs, etc in here to tell it what filesystem to run * <code>mountpoint{ X }</code> use things like /, /mnt/sda3, etc Note that the sizes listed must fit on the physical disks - using 950000 1000 995000 will work fine on a 1T disk but fail on a 250G disk. The characteristics of 'fail' are that it throws you back to the screen that asks "Guided whole disk, guided xxx, or manual?". <pre> # Parameters are: # <raidtype> <devcount> <sparecount> <fstype> <mountpoint> \ # <devices> <sparedevices> d-i partman-auto-raid/recipe string \ 1 2 0 ext3 / \ /dev/sda1#/dev/sdb1 \ . </pre> This snippet is telling us to make a raid 1, with 2 devices, 0 spares, ext3 filesystem, mounted at /, and across /dev/sda1 and /dev/sda2 <pre> d-i partman-md/confirm boolean true </pre> Theoretically now it confirms stuff automatically. It doesn't. Partman lies <pre> d-i partman-md/device_remove_md boolean true </pre> This lets partman remove any existing raids <pre> d-i partman/confirm_write_new_label boolean true d-i partman/choose_partition select finish d-i partman/confirm boolean true d-i partman-lvm/device_remove_lvm boolean true d-i mdadm/boot_degraded boolean true </pre> Most of these possibly do what they appear to do. Boot degraded is important - if we have a disk failure we'd still like the system to boot, just warn us. == Dissecting another config == <pre> # Application server specific configuration # Implementation specific hack: d-i partman-auto/init_automatically_partition select 20some_device__________/var/lib/partman/devices/=dev=sda d-i partman-auto/method string regular d-i partman-auto/disk string /dev/sda d-i partman/choose_partition select Finish partitioning and write changes to disk d-i partman/confirm boolean true # Note, expert_recipe wants to fill up the entire disk # See http://d-i.alioth.debian.org/svn/debian-installer/installer/doc/devel/partman-auto-recipe.txt d-i partman-auto/expert_recipe string apache : 3000 5000 8000 ext3 $primary{ } $bootable{ } method{ format } format{ } use_filesystem{ } filesystem{ ext3 } mountpoint{ / } . 1000 1000 1000 linux-swap method{ swap } format{ } . 64 1000 10000000 jfs method{ format } format{ } use_fil esystem{ } filesystem{ jfs } mountpoint{ /a } . d-i partman-auto/choose_recipe apache # Preseeding of other packages fontconfig fontconfig/enable_bitmaps boolean true </pre> These are preseed file entries that get fed to prompts issued by the installer. Syntax: <code><owner> <question name> <question type> <value></code> <var><owner></var>: "d-i" which stands for Debian Installer. <var><question name></var>: partman-auto, partman-auto-raid and partman-auto-lvm are the packages that handle automatic partitioning of various types, and some of the questions (prompts) issued by them are described here: [http://d-i.alioth.debian.org/manual/en.i386/apbs04.html#preseed-partman] <var><question type></var> says what sort of value to expect (e.g. string, boolean, select (for a menu)...) <var><question value></var> this is where you put the answer that you would otherwise be entering interactively NOTE that after the question type you can put only one space or tab; any other whitespace will get stuffed in at the front of the value, which you probably don't want. The first few lines of the sample config file should be obvious: select the device, choose the partition method and confirm. After that an "expert recipe" for partioning is defined, and given the name "apache". This allows it to be chosen later from the list of recipes. Let's look at the particular recipe: it says *3000 5000 8000 ext3 $primary{ } $bootable{ } method{ format } format{ } use_filesystem{ } filesystem{ ext3 } mountpoint{ / } *:3000: minumum size of partition in mb *:5000: priority if it and other listed partitions are vying for space on the disk (this is compared with the priorities of the other partitions) *:8000: maximum size of partition in mb (80GB; this is for 80GB disks, which the apaches all have.) *:ext3: filesystem type *:$primary{ }: this is a primary, not logical partition *:$bootable{ }: this is a bootable partition *:method{ format }: set to format to format the partition, to "keep" to not format, and to "swap" for swap partitions *:format{ }: also needed so the partition will be formatted *:use_filesystem{ }: this partition will have a filesystem on it (it won't be swap, lvm, etc) *:filesystem{ ext3 }: what filesystem it gets *:mountpoint{ / }: where it's mounted The other two lines in the recipe should now be obvious. *1000 1000 1000 linux-swap method{ swap } format{ } *64 1000 10000000 jfs method{ format } format{ } use_filesystem{ } filesystem{ jfs } mountpoint{ /a } == Things that peter has found out == LVM: I learned these bits when making mailman.cfg, so I'll be dragging things out of there. (One of) The (many, many) weird thing(s) about partman is that ''partman-auto/expert_recipe'' is used for defining partitions on disks and logical volumes. How could this ever not be confusing!?!? I wanted to put some logical volumes on top of raid1, so here is what I did: I started with a nice <pre> d-i partman-auto/expert_recipe string \ multiraid :: \ 5000 8000 16000 raid \ $primary{ } $lvmignore{ } method{ raid } \ . \ </pre> and <pre> d-i partman-auto-raid/recipe string \ 1 2 0 ext3 / \ /dev/sda1#/dev/sdb1 \ . \ </pre> Which yields a slice off of both sda and sdb 16GB large, raided together, ext3 formatted, mounted at /. That's some sweet root partition. Next thing is in the ''partman-auto/expert-recipe'' I added <pre> 64 1000 10000000 raid \ $primary{ } $lvmignore { } method{ raid } \ . \ </pre> And in ''partman-auto-raid/recipe'' I added <pre> 1 2 0 lvm - \ /dev/sda2#/dev/sdb2 \ . </pre> Where the first part is basically like "for sda/sdb, grab all remaining physical disk and partition them for some raid" and the second part is like "get those two partitions, raid them together, and let's use it for some LVM" So now is the part where we make some LVM partitions. This all happens in ''partman-auto/expert_recipe''. A lot of it seems wildly redundant, but so it goes. First, let's create some swap on the lvm, so we can grow it later if need be. For this, we use <pre> 4000 4000 4000 linux-swap \ $defaultignore{ } $lvmok{ } \ lv_name{ swap } method{ swap } format{ } \ . \ </pre> Which is (relatively) high priority. It probably would have been smarter to have used 100% as the upper bound, so that it was equal to the ram on the box. But this is fine. Still allows for performance tuning. And now to create an xfs formatted logical volume with a mountpoint we put in <pre> 100 1000 300000 xfs \ $defaultignore{ } $lvmok{ } \ lv_name{ mailman } method{ format } format{ } \ use_filesystem{ } filesystem{ xfs } \ mountpoint{ /var/lib/mailman } \ . \ </pre> Relatively self explanitory options. A name is a name. ''method{ format } format{ } use_filesystem{ } filesystem{ xfs }'' seems like a rather obtuse way of saying "no really make me some xfs..." and the ''mountpoint'' is the mountpoint. We could have more of these as well. For the sake of completeness, I'll include the other logical volume in mailserver.cfg. <pre> 100 1000 60000 xfs \ $defaultignore{ } $lvmok{ } \ lv_name{ exim } method{ format } format{ } \ use_filesystem{ } filesystem{ xfs } \ mountpoint{ /var/spool/exim4 } \ . </pre> You'll notice that these are the same priority, so if there isn't enough space, the ratio of the sizes will stay the same (after the swap is made). Another couple of options use/notes: <pre> d-i partman-auto-lvm/guided_size string 80% </pre> That option is something we are fond of at WMF. It allows us to grow some partitions over time. Basically, it means that when making the logical volumes, to only use 80% of the available disk space. The means that if we need more swap later, or if we're running out of room on one partition, we can use some of the space. Another take away is that basically in ''partman-auto/expert_recipe'' <pre> $primary{ } $lvmignore{ } </pre> means "make a physical partition" whereas <pre> $defaultignore{ } $lvmok{ } </pre> means "these are logical volumes." This is what Peter knows. == Preserving Data == When you want to reimage without losing data, you can use the <code>reuse-parts.cfg</code> partman recipe, along with a specific disk configuration. These recipes can be found in: [[gerrit:plugins/gitiles/operations/puppet/+/refs/heads/production/modules/install_server/files/autoinstall/partman/custom/]] If you have a particularly important volume, or if you are doing the first of a large set of hosts, it may be helpful to add a manual verification step, by using the <code>reuse-parts-test.cfg</code> recipe, instead. === Verifying before applying === If you also associate your host with the [https://gerrit.wikimedia.org/r/plugins/gitiles/operations/puppet/+/refs/heads/production/modules/install_server/files/autoinstall/reuse-parts-test.cfg reuse-parts-test.cfg] recipe, you can inspect the results of the recipe via the serial console before it is applied. What you are looking for is a <code>K</code> on the partition/mount point you need to preserve, as well as any LVM PVs/VGs/LVs that comprise that partition. [[File:Reuse-parts recipe.png|left|thumb|500x500px]] {{clear}} == Other Helpful Documents == * [https://salsa.debian.org/installer-team/debian-installer Debian installer repo] and [https://salsa.debian.org/installer-team/debian-installer/blob/master/doc/devel/partman-auto-recipe.txt some docs] [[Category:How-To]] * [https://www.bishnet.net/tim/blog/2015/01/29/understanding-partman-autoexpert%20recipe/ Blog post with a good explanation how the priority is calculated] 4yp38ibqsrgfu6u82ovk2y6yxd169a4 Server Admin Log 0 7919 2450648 2450643 2026-08-22T16:32:58Z Stashbot 7414 arlolra@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply 2450648 wikitext text/x-wiki == 2026-08-22 == * 16:32 arlolra@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 35s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-21 == * 20:36 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in1001.wikimedia.org with reason: [[phab:T434750|T434750]] * 20:34 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in2001.wikimedia.org with reason: [[phab:T434750|T434750]] * 20:33 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out1001.wikimedia.org with reason: [[phab:T434750|T434750]] * 20:25 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out2001.wikimedia.org with reason: [[phab:T434750|T434750]] * 19:37 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host krb1002.eqiad.wmnet with OS bookworm * 19:00 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 18:59 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 18:51 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 18:51 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 18:35 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:35 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:27 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:27 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:16 bking@cumin2003: START - Cookbook sre.hosts.reimage for host krb1002.eqiad.wmnet with OS bookworm * 17:35 sukhe@dns1004: END - running authdns-update * 17:33 sukhe@dns1004: START - running authdns-update * 17:32 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns5004.wikimedia.org [reason: resolved authdns-update issues] * 17:32 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:32 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: force HEAD to {{Gerrit|be26e30ae101}} - sukhe@cumin1003" * 17:32 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: force HEAD to {{Gerrit|be26e30ae101}} - sukhe@cumin1003" * 17:28 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 17:28 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: service=authdns-update,name=dns5004.wikimedia.org [reason: resolving authdns-update issues] * 17:28 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:28 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: force HEAD to {{Gerrit|be26e30ae101}} - sukhe@cumin1003" * 17:28 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: force HEAD to {{Gerrit|be26e30ae101}} - sukhe@cumin1003" * 17:24 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 17:24 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.netbox (exit_code=97) * 17:23 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 17:16 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 17:12 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 17:10 sukhe@dns1004: END - running authdns-update * 17:08 sukhe@dns1004: START - running authdns-update * 17:08 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=dns5004.wikimedia.org [reason: resolving authdns-update issues] * 17:07 sukhe@dns1004: FAIL - running authdns-update * 17:05 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 17:05 sukhe@dns1004: START - running authdns-update * 17:01 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 16:59 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 16:56 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 16:53 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=dns5004.* [reason: trixie upgrade] * 16:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns5004.wikimedia.org * 16:52 cdobbins@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns5004.wikimedia.org * 16:44 cmooney@dns3003: END - running authdns-update * 16:41 cmooney@dns3003: START - running authdns-update * 16:41 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:41 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on eqsin<->codfw arelion - cmooney@cumin1003" * 16:37 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on eqsin<->codfw arelion - cmooney@cumin1003" * 16:33 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:11 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 16:08 sukhe@dns1004: END - running authdns-update * 16:08 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 16:06 sukhe@dns1004: START - running authdns-update * 16:04 cmooney@dns3003: END - running authdns-update * 16:02 cmooney@dns3003: START - running authdns-update * 16:00 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:00 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on eqord<->codfw arelion - cmooney@cumin1003" * 15:56 cdobbins@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host dns5004.wikimedia.org with OS trixie * 15:55 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on eqord<->codfw arelion - cmooney@cumin1003" * 15:53 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:51 cmooney@cumin1003: END (ERROR) - Cookbook sre.dns.netbox (exit_code=97) * 15:51 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:38 andrew@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudcephosd1042.eqiad.wmnet with OS bookworm * 15:18 andrew@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudcephosd1042.eqiad.wmnet with reason: host reimage * 15:17 dancy@deploy1003: Finished deploy [gerrit/gerrit@2cc11cc]: Deploying https://gerrit.wikimedia.org/r/c/operations/software/gerrit/+/1327669 ([[phab:T434726|T434726]]) (duration: 00m 14s) * 15:17 dancy@deploy1003: Started deploy [gerrit/gerrit@2cc11cc]: Deploying https://gerrit.wikimedia.org/r/c/operations/software/gerrit/+/1327669 ([[phab:T434726|T434726]]) * 15:13 andrew@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudcephosd1042.eqiad.wmnet with reason: host reimage * 15:09 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns5004.wikimedia.org with reason: host reimage * 15:05 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns5004.wikimedia.org with reason: host reimage * 14:53 andrew@cumin2003: START - Cookbook sre.hosts.reimage for host cloudcephosd1042.eqiad.wmnet with OS bookworm * 14:30 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns5004.wikimedia.org with OS trixie * 14:29 cdobbins@cumin1003: conftool action : set/pooled=no; selector: name=dns5004.* [reason: trixie upgrade] * 14:21 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:21 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove entries for cr2-eqord - cmooney@cumin1003" * 14:21 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove entries for cr2-eqord - cmooney@cumin1003" * 14:13 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 14:11 moritzm: imported openjdk 8u504-ga-1~deb12u1 for bookworm-wikimedia (backport of the latest Java 8 security fixes for bookworm) * 13:25 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "sync cr2-eqord router offline - cmooney@cumin1003" * 13:23 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "sync cr2-eqord router offline - cmooney@cumin1003" * 13:14 hashar@deploy1003: Finished deploy [integration/docroot@2d5ff9b]: opensource: add PersonalDashboard docs to MW components - [[phab:T435392|T435392]] (duration: 00m 15s) * 13:14 hashar@deploy1003: Started deploy [integration/docroot@2d5ff9b]: opensource: add PersonalDashboard docs to MW components - [[phab:T435392|T435392]] * 12:16 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:16 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: [[phab:T431682|T431682]] - filippo@cumin1003" * 12:16 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: [[phab:T431682|T431682]] - filippo@cumin1003" * 12:11 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2006.wikimedia.org with OS trixie * 12:00 kevinbazira@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 11:58 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 11:43 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2006.wikimedia.org with reason: host reimage * 11:41 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 11:38 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2006.wikimedia.org with reason: host reimage * 11:20 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2006.wikimedia.org with OS trixie * 11:11 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2005.wikimedia.org with OS trixie * 10:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2005.wikimedia.org with reason: host reimage * 10:53 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2005.wikimedia.org with reason: host reimage * 10:43 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-codfw * 10:43 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2011.codfw.wmnet * 10:43 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2011.codfw.wmnet * 10:40 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2011.codfw.wmnet * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2011.codfw.wmnet * 10:34 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2010.codfw.wmnet * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2010.codfw.wmnet * 10:33 fnegri@cumin1003: END (PASS) - Cookbook sre.wikireplicas.add-wiki (exit_code=0) for database bolwiki ([[phab:T429954|T429954]]) * 10:33 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2005.wikimedia.org with OS trixie * 10:30 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2010.codfw.wmnet * 10:25 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2010.codfw.wmnet * 10:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2009.codfw.wmnet * 10:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2009.codfw.wmnet * 10:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1006.wikimedia.org with OS trixie * 10:18 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2009.codfw.wmnet * 10:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2009.codfw.wmnet * 10:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2008.codfw.wmnet * 10:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2008.codfw.wmnet * 10:06 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2008.codfw.wmnet * 10:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1006.wikimedia.org with reason: host reimage * 10:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2008.codfw.wmnet * 10:01 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2007.codfw.wmnet * 10:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2007.codfw.wmnet * 09:57 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1006.wikimedia.org with reason: host reimage * 09:56 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2007.codfw.wmnet * 09:51 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2007.codfw.wmnet * 09:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 09:51 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 09:46 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2006.codfw.wmnet * 09:46 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1006.wikimedia.org with OS trixie * 09:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1005.wikimedia.org with OS trixie * 09:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2006.codfw.wmnet * 09:41 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2005.codfw.wmnet * 09:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2005.codfw.wmnet * 09:38 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2013.codfw.wmnet * 09:36 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2005.codfw.wmnet * 09:35 fnegri@cumin1003: START - Cookbook sre.wikireplicas.add-wiki for database bolwiki ([[phab:T429954|T429954]]) * 09:35 fnegri@cumin1003: END (PASS) - Cookbook sre.wikireplicas.add-wiki (exit_code=0) for database minwikiquote ([[phab:T429946|T429946]]) * 09:35 fnegri@cumin1003: START - Cookbook sre.wikireplicas.add-wiki for database minwikiquote ([[phab:T429946|T429946]]) * 09:32 blake@cumin1003: START - Cookbook sre.hosts.reboot-single for host rdb2013.codfw.wmnet * 09:30 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2011.codfw.wmnet * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1005.wikimedia.org with reason: host reimage * 09:26 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2005.codfw.wmnet * 09:26 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2004.codfw.wmnet * 09:26 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2004.codfw.wmnet * 09:24 blake@cumin1003: START - Cookbook sre.hosts.reboot-single for host rdb2011.codfw.wmnet * 09:22 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1005.wikimedia.org with reason: host reimage * 09:21 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2004.codfw.wmnet * 09:16 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1015.eqiad.wmnet * 09:16 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2004.codfw.wmnet * 09:16 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2003.codfw.wmnet * 09:16 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2003.codfw.wmnet * 09:13 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on cr[1-2]-eqiad,pfw1-eqiad with reason: upgrade pfw1a-eqiad and pfw1b-eqiad pair * 09:12 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 09:11 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 09:11 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 09:11 blake@cumin1003: START - Cookbook sre.hosts.reboot-single for host rdb1015.eqiad.wmnet * 09:10 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2003.codfw.wmnet * 09:09 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1013.eqiad.wmnet * 09:07 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1005.wikimedia.org with OS trixie * 09:03 blake@cumin1003: START - Cookbook sre.hosts.reboot-single for host rdb1013.eqiad.wmnet * 09:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2003.codfw.wmnet * 09:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2002.codfw.wmnet * 09:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2002.codfw.wmnet * 08:54 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2002.codfw.wmnet * 08:49 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2002.codfw.wmnet * 08:49 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2001.codfw.wmnet * 08:49 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2001.codfw.wmnet * 08:48 jmm@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts netmon2002.wikimedia.org * 08:47 jmm@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts netmon2002.wikimedia.org * 08:44 jmm@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts netmon2002.wikimedia.org * 08:44 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon2002.wikimedia.org * 08:43 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2001.codfw.wmnet * 08:36 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon2002.wikimedia.org * 08:34 jmm@dns1004: END - running authdns-update * 08:33 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2001.codfw.wmnet * 08:33 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-codfw * 08:31 jmm@dns1004: START - running authdns-update * 07:48 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327679{{!}}Block: Disable flaky API test (T435272 T389028)]], [[gerrit:1327678{{!}}API: wfDebugLog for thumberror]] (duration: 15m 34s) * 07:41 krinkle@deploy1003: krinkle: Continuing with deployment * 07:37 krinkle@deploy1003: krinkle: Backport for [[gerrit:1327679{{!}}Block: Disable flaky API test (T435272 T389028)]], [[gerrit:1327678{{!}}API: wfDebugLog for thumberror]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:33 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1327679{{!}}Block: Disable flaky API test (T435272 T389028)]], [[gerrit:1327678{{!}}API: wfDebugLog for thumberror]] * 07:25 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 07:24 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 07:18 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 07:18 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 07:15 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 07:14 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 07:14 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 07:14 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 07:13 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 07:03 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1283: Pool back * 06:42 jmm@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts netmon2002.wikimedia.org * 06:35 moritzm: powercycling netmon2002 * 06:18 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1283: Pool back * 06:17 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1283 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96213 and previous config saved to /var/cache/conftool/dbconfig/20260821-061743-marostegui.json * 04:59 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324963{{!}}Add Produnto to extension-list (T421436)]], [[gerrit:1324964{{!}}Enable Produnto on Beta (T421436)]] (duration: 34m 48s) * 04:45 tstarling@deploy1003: tstarling: Continuing with deployment * 04:44 tstarling@deploy1003: tstarling: Backport for [[gerrit:1324963{{!}}Add Produnto to extension-list (T421436)]], [[gerrit:1324964{{!}}Enable Produnto on Beta (T421436)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 04:24 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1324963{{!}}Add Produnto to extension-list (T421436)]], [[gerrit:1324964{{!}}Enable Produnto on Beta (T421436)]] * 04:21 arlolra@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 04:20 arlolra@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 04:20 arlolra@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 04:20 arlolra@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 41s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-20 == * 23:43 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327652{{!}}RunSingleJob: Add ProfilingContext::init() (T435422)]] (duration: 12m 06s) * 23:38 krinkle@deploy1003: krinkle: Continuing with deployment * 23:33 krinkle@deploy1003: krinkle: Backport for [[gerrit:1327652{{!}}RunSingleJob: Add ProfilingContext::init() (T435422)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:31 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1327652{{!}}RunSingleJob: Add ProfilingContext::init() (T435422)]] * 22:15 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1054.eqiad.wmnet with OS trixie * 22:14 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 22:14 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 21:58 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1054.eqiad.wmnet with reason: host reimage * 21:51 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1054.eqiad.wmnet with reason: host reimage * 21:36 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1054.eqiad.wmnet with OS trixie * 21:36 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:35 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327219{{!}}RunSingleJob: Define MW_ENTRY_POINT for flamegraph sample attribution (T435422)]] (duration: 08m 30s) * 21:31 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:31 krinkle@deploy1003: krinkle: Continuing with deployment * 21:31 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1054 * 21:31 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1054 * 21:29 krinkle@deploy1003: krinkle: Backport for [[gerrit:1327219{{!}}RunSingleJob: Define MW_ENTRY_POINT for flamegraph sample attribution (T435422)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:27 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1327219{{!}}RunSingleJob: Define MW_ENTRY_POINT for flamegraph sample attribution (T435422)]] * 21:17 maryum: Deployed security fix for [[phab:T433020|T433020]] * 20:59 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324752{{!}}InitialiseSettings: Enable 2FA banners on remaining private wikis (T428103)]], [[gerrit:1325920{{!}}Remove sending email to legal team about rejected requests (T374053)]] (duration: 07m 18s) * 20:54 reedy@deploy1003: neriah, reedy: Continuing with deployment * 20:54 reedy@deploy1003: neriah, reedy: Backport for [[gerrit:1324752{{!}}InitialiseSettings: Enable 2FA banners on remaining private wikis (T428103)]], [[gerrit:1325920{{!}}Remove sending email to legal team about rejected requests (T374053)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:51 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324752{{!}}InitialiseSettings: Enable 2FA banners on remaining private wikis (T428103)]], [[gerrit:1325920{{!}}Remove sending email to legal team about rejected requests (T374053)]] * 20:24 reedy@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.15,1.47.0-wmf.16,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/med * 20:23 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324752{{!}}InitialiseSettings: Enable 2FA banners on remaining private wikis (T428103)]], [[gerrit:1325920{{!}}Remove sending email to legal team about rejected requests (T374053)]] * 20:14 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 20:10 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 20:09 cdanis@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "bug fixes & UX fixes - cdanis@cumin1003" * 20:09 cdanis@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: bug fixes & UX fixes - cdanis@cumin1003 * 20:08 cdanis@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: bug fixes & UX fixes - cdanis@cumin1003 * 20:08 cdanis@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "bug fixes & UX fixes - cdanis@cumin1003" * 19:24 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327598{{!}}Make \Omicron non upright (like \Chi) (T434428)]], [[gerrit:1327596{{!}}Render overline of \bar with stretchy=false (T435456)]] (duration: 18m 54s) * 19:20 krinkle@deploy1003: krinkle: Continuing with deployment * 19:07 krinkle@deploy1003: krinkle: Backport for [[gerrit:1327598{{!}}Make \Omicron non upright (like \Chi) (T434428)]], [[gerrit:1327596{{!}}Render overline of \bar with stretchy=false (T435456)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:05 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1327598{{!}}Make \Omicron non upright (like \Chi) (T434428)]], [[gerrit:1327596{{!}}Render overline of \bar with stretchy=false (T435456)]] * 18:51 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327614{{!}}Avoid casting fpxmax to string (T318419)]] (duration: 07m 28s) * 18:50 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:46 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:46 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 18:45 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327614{{!}}Avoid casting fpxmax to string (T318419)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:43 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327614{{!}}Avoid casting fpxmax to string (T318419)]] * 18:37 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:37 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:37 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:36 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:07 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 18:05 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 18:01 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 18:01 sukhe@dns1004: END - running authdns-update * 17:59 sukhe@dns1004: START - running authdns-update * 17:58 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 17:57 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=dns6002.* [reason: depooling for trixie upgrade] * 17:56 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns6002.wikimedia.org * 17:56 cdobbins@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns6002.wikimedia.org * 17:51 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 17:51 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 17:34 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host stat1011.eqiad.wmnet with OS bookworm * 17:31 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns6002.wikimedia.org with OS trixie * 17:30 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 17:30 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 17:30 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 17:30 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 17:29 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 17:29 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 17:29 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 17:29 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:29 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:27 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:24 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327590{{!}}Make sure fpsmax is an int value (T318419)]] (duration: 08m 37s) * 17:20 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 17:17 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327590{{!}}Make sure fpsmax is an int value (T318419)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:16 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 17:15 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327590{{!}}Make sure fpsmax is an int value (T318419)]] * 16:51 swfrench-wmf: disable-puppet on A:cp for ATS Lua change - [[phab:T427666|T427666]] * 16:51 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db2901.codfw.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 16:43 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 16:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on stat1011.eqiad.wmnet with reason: host reimage * 16:39 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns6002.wikimedia.org with reason: host reimage * 16:36 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on stat1011.eqiad.wmnet with reason: host reimage * 16:36 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db2901.codfw.wmnet * 16:34 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns6002.wikimedia.org with reason: host reimage * 16:15 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns6002.wikimedia.org with OS trixie * 16:14 cdobbins@cumin1003: conftool action : set/pooled=no; selector: name=dns6002.* [reason: depooling for trixie upgrade] * 16:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host stat1011 * 16:10 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host stat1011 * 16:09 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host stat1011 * 16:09 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) stat1011.eqiad.wmnet 14.36.64.10.in-addr.arpa 4.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 bking@cumin2003: START - Cookbook sre.dns.wipe-cache stat1011.eqiad.wmnet 14.36.64.10.in-addr.arpa 4.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:09 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host stat1011 - bking@cumin2003" * 16:09 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host stat1011 - bking@cumin2003" * 16:05 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host stat1011 * 16:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host stat1011.eqiad.wmnet with OS bookworm * 16:00 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 15:59 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 15:56 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 15:55 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 15:37 jayme@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:35 jayme@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 15:35 jayme@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:32 jayme@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:32 jayme@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:30 jayme@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 15:30 jayme@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:28 jayme@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:28 jayme@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 15:28 fceratto@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host db1903.eqiad.wmnet * 15:28 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1903.eqiad.wmnet with OS trixie * 15:26 jayme@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 15:26 jayme@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 15:24 jayme@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 15:24 jayme@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 15:21 jayme@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 15:21 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 15:19 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 15:19 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 15:17 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 15:14 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1903.eqiad.wmnet with reason: host reimage * 15:07 fceratto@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1903.eqiad.wmnet with reason: host reimage * 14:54 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db1903.eqiad.wmnet with OS trixie * 14:53 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1903.eqiad.wmnet - fceratto@cumin1003" * 14:53 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1903.eqiad.wmnet - fceratto@cumin1003" * 14:53 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1903.eqiad.wmnet on all recursors * 14:53 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1903.eqiad.wmnet on all recursors * 14:53 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:53 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1903.eqiad.wmnet - fceratto@cumin1003" * 14:53 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1903.eqiad.wmnet - fceratto@cumin1003" * 14:49 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 14:49 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1903.eqiad.wmnet * 14:33 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=cp1100.* * 14:27 topranks: reconfigure eqiad<->codfw bgp settings * 14:22 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:22 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update entries used on new transport backup eqiad codfw - cmooney@cumin1003" * 14:19 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update entries used on new transport backup eqiad codfw - cmooney@cumin1003" * 14:14 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 14:14 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 14:13 moritzm: installing util-linux security updates * 14:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-staging-worker * 14:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2003.codfw.wmnet * 14:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2003.codfw.wmnet * 14:08 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2003.codfw.wmnet * 14:06 moritzm: installing libheif security updates * 13:58 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2003.codfw.wmnet * 13:58 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2002.codfw.wmnet * 13:58 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2002.codfw.wmnet * 13:56 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:56 fnegri@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for clouddb1025.eqiad.wmnet * 13:56 fnegri@cumin1003: START - Cookbook sre.hosts.remove-downtime for clouddb1025.eqiad.wmnet * 13:56 Lucas_WMDE: UTC afternoon backport+config window done * 13:53 moritzm: installing apr-util security updates * 13:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2002.codfw.wmnet * 13:50 fnegri@cumin1003: conftool action : set/weight=100; selector: name=clouddb1025.eqiad.wmnet * 13:49 fnegri@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1025.eqiad.wmnet * 13:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2002.codfw.wmnet * 13:41 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2001.codfw.wmnet * 13:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2001.codfw.wmnet * 13:41 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'. * 13:38 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'. * 13:38 fnegri@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on clouddb1025.eqiad.wmnet with reason: Removing s6 from clouddb1025 * 13:34 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2001.codfw.wmnet * 13:31 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'. * 13:29 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'. * 13:28 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet * 13:26 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host stat1009.eqiad.wmnet with OS bookworm * 13:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2001.codfw.wmnet * 13:24 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-staging-worker * 13:23 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1002.eqiad.wmnet * 13:21 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host stat1010.eqiad.wmnet with OS bookworm * 13:20 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1002.eqiad.wmnet * 13:20 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1001.eqiad.wmnet * 13:17 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1001.eqiad.wmnet * 13:16 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2001.codfw.wmnet * 13:13 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327511{{!}}UIC: Fix page:page instead of page:other in instrumentation]] (duration: 07m 00s) * 13:13 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2001.codfw.wmnet * 13:12 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2002.codfw.wmnet * 13:09 mszwarc@deploy1003: mszwarc: Continuing with deployment * 13:08 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1327511{{!}}UIC: Fix page:page instead of page:other in instrumentation]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2002.codfw.wmnet * 13:07 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2002.codfw.wmnet * 13:06 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1327511{{!}}UIC: Fix page:page instead of page:other in instrumentation]] * 13:04 jmm@dns1004: END - running authdns-update * 13:03 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2002.codfw.wmnet * 13:03 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2001.codfw.wmnet * 13:02 jmm@dns1004: START - running authdns-update * 13:01 cmooney@dns3003: END - running authdns-update * 13:00 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2001.codfw.wmnet * 12:59 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2001.codfw.wmnet * 12:59 cmooney@dns3003: START - running authdns-update * 12:57 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2001.codfw.wmnet * 12:56 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2002.codfw.wmnet * 12:55 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:55 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on drmrs<->eqiad GTT vpls - cmooney@cumin1003" * 12:54 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on drmrs<->eqiad GTT vpls - cmooney@cumin1003" * 12:54 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2002.codfw.wmnet * 12:54 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2003.codfw.wmnet * 12:50 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2003.codfw.wmnet * 12:49 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 12:48 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2001.codfw.wmnet * 12:46 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2001.codfw.wmnet * 12:46 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2002.codfw.wmnet * 12:43 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2002.codfw.wmnet * 12:43 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2003.codfw.wmnet * 12:42 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=1) for new host db1902.eqiad.wmnet * 12:42 fceratto@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host db1902.eqiad.wmnet with OS trixie * 12:41 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2003.codfw.wmnet * 12:40 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1003.eqiad.wmnet * 12:38 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1003.eqiad.wmnet * 12:37 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1002.eqiad.wmnet * 12:35 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1002.eqiad.wmnet * 12:35 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1001.eqiad.wmnet * 12:34 cmooney@dns3003: END - running authdns-update * 12:33 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1001.eqiad.wmnet * 12:32 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on stat1009.eqiad.wmnet with reason: host reimage * 12:31 cmooney@dns3003: START - running authdns-update * 12:31 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:31 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on drmrs<->eqiad cct - cmooney@cumin1003" * 12:28 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on drmrs<->eqiad cct - cmooney@cumin1003" * 12:28 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1902.eqiad.wmnet with reason: host reimage * 12:25 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on stat1009.eqiad.wmnet with reason: host reimage * 12:25 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 12:24 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on stat1010.eqiad.wmnet with reason: host reimage * 12:22 fceratto@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1902.eqiad.wmnet with reason: host reimage * 12:21 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on stat1010.eqiad.wmnet with reason: host reimage * 12:14 elukey: move the Docker Registry's /v2/dev/.* prefix to its dedicated S3 backend - [[phab:T432829|T432829]] * 12:12 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db1902.eqiad.wmnet with OS trixie * 12:09 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1902.eqiad.wmnet - fceratto@cumin1003" * 12:09 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1902.eqiad.wmnet - fceratto@cumin1003" * 12:09 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1902.eqiad.wmnet on all recursors * 12:09 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1902.eqiad.wmnet on all recursors * 12:09 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:08 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1902.eqiad.wmnet - fceratto@cumin1003" * 12:08 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1902.eqiad.wmnet - fceratto@cumin1003" * 12:08 tgr_: [[phab:T413390|T413390]] running CentralAuth:FixRenamedUserGlobalEditCount --wiki=metawiki --since=20250901000000 --until=20260301000000 --fix * 12:04 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1009.eqiad.wmnet with OS bookworm * 12:04 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1010.eqiad.wmnet with OS bookworm * 12:01 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 12:01 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1902.eqiad.wmnet * 12:00 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host stat1010.eqiad.wmnet with OS bookworm * 11:57 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1902.eqiad.wmnet * 11:57 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:57 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1902.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 11:57 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1902.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 11:51 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'. * 11:49 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'. * 11:48 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'. * 11:46 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'. * 11:39 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 11:37 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327518{{!}}Enable thumb.wikimedia.org on cswiki and fawiki (T427465)]] (duration: 10m 40s) * 11:35 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1902.eqiad.wmnet * 11:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1010.eqiad.wmnet with OS bookworm * 11:33 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 11:30 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327518{{!}}Enable thumb.wikimedia.org on cswiki and fawiki (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:26 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327518{{!}}Enable thumb.wikimedia.org on cswiki and fawiki (T427465)]] * 11:24 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host stat1010.eqiad.wmnet with OS bookworm * 11:07 fceratto@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host db1901.eqiad.wmnet * 11:07 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1901.eqiad.wmnet with OS trixie * 10:53 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1901.eqiad.wmnet with reason: host reimage * 10:47 tappof: bump space for prometheus k8s-aux in codfw * 10:47 tappof: bump space for prometheus k8s-dse in eqiad * 10:47 fceratto@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1901.eqiad.wmnet with reason: host reimage * 10:35 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db1901.eqiad.wmnet with OS trixie * 10:32 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:32 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:32 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1901.eqiad.wmnet on all recursors * 10:32 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1901.eqiad.wmnet on all recursors * 10:31 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:31 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:31 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:27 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:27 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1901.eqiad.wmnet * 10:23 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1010.eqiad.wmnet with OS bookworm * 10:20 fceratto@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts db1901.eqiad.wmnet * 10:20 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 10:18 blake@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 10:17 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:16 blake@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 10:13 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1901.eqiad.wmnet * 09:23 jelto@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'. * 09:22 jelto@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'. * 09:22 jelto: update cert-manager to 1.19.6 on wikikube staging-eqiad - [[phab:T427402|T427402]] * 09:20 moritzm: imported squid 7.6-2.1for trixie-wikimedia/main [[phab:T427282|T427282]] * 09:08 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 09:08 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 09:08 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 09:07 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 09:04 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 09:04 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:23 slyngshede@dns1004: END - running authdns-update * 08:21 slyngshede@dns1004: START - running authdns-update * 08:18 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.16 refs [[phab:T430835|T430835]] * 06:27 aokoth@dns1004: END - running authdns-update * 06:25 aokoth@dns1004: START - running authdns-update * 06:22 brennen@deploy1003: Finished deploy [phabricator/deployment@6b9b6ff]: deploy phab1005 for [[phab:T435087|T435087]] (duration: 00m 39s) * 06:21 brennen@deploy1003: Started deploy [phabricator/deployment@6b9b6ff]: deploy phab1005 for [[phab:T435087|T435087]] * 06:20 brennen@deploy1003: Finished deploy [phabricator/deployment@6b9b6ff]: deploy phab1004 for to pick up config values for [[phab:T435087|T435087]] (duration: 01m 46s) * 06:18 brennen@deploy1003: Started deploy [phabricator/deployment@6b9b6ff]: deploy phab1004 for to pick up config values for [[phab:T435087|T435087]] * 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 49s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-19 == * 23:19 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327197{{!}}Enable thumb.wikimedia.org on mediawiki.org (T427465)]] (duration: 10m 50s) * 23:18 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:16 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 23:15 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 23:10 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327197{{!}}Enable thumb.wikimedia.org on mediawiki.org (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:10 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:09 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 23:09 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:09 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 23:08 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327197{{!}}Enable thumb.wikimedia.org on mediawiki.org (T427465)]] * 22:58 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327201{{!}}Enable ReadingLists for all logged in users on test wiki (T435258)]] (duration: 11m 20s) * 22:50 jdlrobson@deploy1003: jdlrobson: Continuing with deployment * 22:49 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1327201{{!}}Enable ReadingLists for all logged in users on test wiki (T435258)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:46 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1327201{{!}}Enable ReadingLists for all logged in users on test wiki (T435258)]] * 22:42 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327162{{!}}Article: Split subjectpageheader by model and disable for wikitext]], [[gerrit:1327169{{!}}Make uppercase greek letters normal (non-italic) font (T434686 T434428)]], [[gerrit:1327176{{!}}Skin: Avoid DB lookup for pagecategorieslink message (T347123)]] (duration: 37m 52s) * 22:29 krinkle@deploy1003: krinkle: Continuing with deployment * 22:25 krinkle@deploy1003: krinkle: Backport for [[gerrit:1327162{{!}}Article: Split subjectpageheader by model and disable for wikitext]], [[gerrit:1327169{{!}}Make uppercase greek letters normal (non-italic) font (T434686 T434428)]], [[gerrit:1327176{{!}}Skin: Avoid DB lookup for pagecategorieslink message (T347123)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:04 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1327162{{!}}Article: Split subjectpageheader by model and disable for wikitext]], [[gerrit:1327169{{!}}Make uppercase greek letters normal (non-italic) font (T434686 T434428)]], [[gerrit:1327176{{!}}Skin: Avoid DB lookup for pagecategorieslink message (T347123)]] * 22:04 krinkle@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: awaiting CI (duration: 03m 06s) * 22:01 krinkle@deploy1003: Locking from deployment [ALL REPOSITORIES]: awaiting CI * 22:00 krinkle@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: awaiting CI (duration: 00m 01s) * 22:00 krinkle@deploy1003: Locking from deployment [ALL REPOSITORIES]: awaiting CI * 21:34 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2001.codfw.wmnet * 21:28 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2001.codfw.wmnet * 21:22 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327178{{!}}AccountRecovery: Notify the email address of the on file of the request (T425799)]] (duration: 47m 02s) * 21:13 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm * 21:09 catrope@deploy1003: catrope: Continuing with deployment * 20:55 catrope@deploy1003: catrope: Backport for [[gerrit:1327178{{!}}AccountRecovery: Notify the email address of the on file of the request (T425799)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:35 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1327178{{!}}AccountRecovery: Notify the email address of the on file of the request (T425799)]] * 20:31 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327128{{!}}Parsoid DataAccess: convert Parsoid fragment markers to/from strip tags (T432547)]] (duration: 07m 30s) * 20:27 catrope@deploy1003: catrope, arlolra: Continuing with deployment * 20:26 catrope@deploy1003: catrope, arlolra: Backport for [[gerrit:1327128{{!}}Parsoid DataAccess: convert Parsoid fragment markers to/from strip tags (T432547)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:24 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1327128{{!}}Parsoid DataAccess: convert Parsoid fragment markers to/from strip tags (T432547)]] * 20:23 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage * 20:17 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage * 20:15 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325878{{!}}[arwiki] Enable restricted user page editing and grant edit permissions (T434878)]] (duration: 08m 46s) * 20:11 catrope@deploy1003: catrope, gergesshamon: Continuing with deployment * 20:08 catrope@deploy1003: catrope, gergesshamon: Backport for [[gerrit:1325878{{!}}[arwiki] Enable restricted user page editing and grant edit permissions (T434878)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:06 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1325878{{!}}[arwiki] Enable restricted user page editing and grant edit permissions (T434878)]] * 19:59 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm * 19:56 eevans@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cassandra-dev2001.codfw.wmnet with OS bookworm * 19:56 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm * 19:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2207.codfw.wmnet with reason: Maintenance * 18:47 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319920{{!}}Allow setting a separate thumbUrl in production (T427465)]], [[gerrit:1327167{{!}}Fix wmgThumbUrl config (T427465)]] (duration: 18m 53s) * 18:43 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 18:30 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1319920{{!}}Allow setting a separate thumbUrl in production (T427465)]], [[gerrit:1327167{{!}}Fix wmgThumbUrl config (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:28 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1319920{{!}}Allow setting a separate thumbUrl in production (T427465)]], [[gerrit:1327167{{!}}Fix wmgThumbUrl config (T427465)]] * 18:26 sukhe@dns1004: END - running authdns-update * 18:24 sukhe@dns1004: START - running authdns-update * 18:09 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1319920{{!}}Allow setting a separate thumbUrl in production (T427465)]] * 18:03 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-eqiad and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 17:56 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-codfw and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 17:50 cmooney@dns3003: END - running authdns-update * 17:42 dancy@deploy1003: Installation of scap version "4.283.0" completed for 3 hosts * 17:41 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling reboot on A:durum and A:durum * 17:40 cmooney@dns3003: START - running authdns-update * 17:40 dancy@deploy1003: Installing scap version "4.283.0" for 3 host(s) * 17:38 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:37 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on GTT VPLS - cmooney@cumin1003" * 17:37 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --verbose --scoreLessThan=0.7 --exceptDatasetChecksums=[[phab:T434319|T434319]]-enwiki-models.txt # [[phab:T434319|T434319]] * 17:32 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on GTT VPLS - cmooney@cumin1003" * 17:29 sbassett: Deployed security fix for [[phab:T435210|T435210]] (wmf.16) * 17:26 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 17:22 sbassett: Deployed security fix for [[phab:T435210|T435210]] (wmf.15) * 17:00 sukhe@dns1004: END - running authdns-update * 16:58 sukhe@dns1004: START - running authdns-update * 16:53 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica-esams and A:liberica * 16:41 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica-esams and A:liberica * 16:41 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-eqiad and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 16:41 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-codfw and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 16:41 cjd91: sudo -i cookbook sre.cdn.roll-upgrade-ats --query 'A:cp-codfw' --task-id [[phab:T434478|T434478]] --reason '9.2.15 upgrade' * 16:41 cjd91: sudo -i cookbook sre.cdn.roll-upgrade-ats --query 'A:cp-eqiad' --task-id [[phab:T434478|T434478]] --reason '9.2.15 upgrade' * 16:40 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and A:durum * 16:28 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326881{{!}}Echo: Start using virtual domains (T380385)]] (duration: 13m 12s) * 16:23 urbanecm@deploy1003: urbanecm: Continuing with deployment * 16:21 urandom: Completed sessionstore Cassandra/JVM upgrade — [[phab:T435154|T435154]] * 16:21 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching sessionstore[2005-2006].codfw.wmnet,sessionstore[1005-1006].eqiad.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 16:19 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1326881{{!}}Echo: Start using virtual domains (T380385)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:15 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:15 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->codfw - cmooney@cumin1003" * 16:14 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1326881{{!}}Echo: Start using virtual domains (T380385)]] * 16:14 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327123{{!}}Revert^2 "Migrate database access to virtual domains" (T435305)]], [[gerrit:1327124{{!}}Pass the mapped domain of virtual-echo-shared to the push NameTableStores (T435305)]] (duration: 07m 42s) * 16:13 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching sessionstore[2005-2006].codfw.wmnet,sessionstore[1005-1006].eqiad.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 16:11 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->codfw - cmooney@cumin1003" * 16:10 urbanecm@deploy1003: urbanecm: Continuing with deployment * 16:08 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1327123{{!}}Revert^2 "Migrate database access to virtual domains" (T435305)]], [[gerrit:1327124{{!}}Pass the mapped domain of virtual-echo-shared to the push NameTableStores (T435305)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:06 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 16:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 16:06 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1327123{{!}}Revert^2 "Migrate database access to virtual domains" (T435305)]], [[gerrit:1327124{{!}}Pass the mapped domain of virtual-echo-shared to the push NameTableStores (T435305)]] * 16:06 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:03 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching sessionstore1004.eqiad.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 16:01 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching sessionstore1004.eqiad.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 15:56 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching sessionstore2004.codfw.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 15:54 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching sessionstore2004.codfw.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 15:53 urandom: beginning sessionstore Cassandra/JVM upgrade — [[phab:T435154|T435154]] * 15:52 urandom: beginning sessionstore Cassandra/JVM upgrade — [[phab:T432944|T432944]] * 15:51 cmooney@dns3003: END - running authdns-update * 15:49 cmooney@dns3003: START - running authdns-update * 15:48 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:48 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->codfw - cmooney@cumin1003" * 15:45 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->codfw - cmooney@cumin1003" * 15:44 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 15:44 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:43 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 15:42 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:42 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:38 cmooney@dns3003: END - running authdns-update * 15:36 cmooney@dns3003: START - running authdns-update * 15:36 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:36 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->eqsin - cmooney@cumin1003" * 15:34 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327138{{!}}Enable redis lock manager everywhere (T366938)]] (duration: 08m 36s) * 15:33 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->eqsin - cmooney@cumin1003" * 15:30 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:29 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 15:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1147.eqiad.wmnet with OS bookworm * 15:28 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327138{{!}}Enable redis lock manager everywhere (T366938)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:25 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327138{{!}}Enable redis lock manager everywhere (T366938)]] * 15:24 jmm@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host krb1002.eqiad.wmnet * 15:19 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2207.codfw.wmnet with reason: Host crashed * 15:17 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327098{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]], [[gerrit:1327101{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]] (duration: 07m 13s) * 15:12 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 15:12 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1327098{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]], [[gerrit:1327101{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:10 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1327098{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]], [[gerrit:1327101{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]] * 15:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2010.codfw.wmnet with OS trixie * 15:05 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 15:05 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1147.eqiad.wmnet with reason: host reimage * 14:59 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 14:59 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:58 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1147.eqiad.wmnet with reason: host reimage * 14:55 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:55 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:49 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 14:48 cmooney@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host durum1001.eqiad.wmnet * 14:46 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:46 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Delete db2902 ipv6 addr - fceratto@cumin1003" * 14:46 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Delete db2902 ipv6 addr - fceratto@cumin1003" * 14:43 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1147.eqiad.wmnet with OS bookworm * 14:42 tgr@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327094{{!}}SpecialMWOAuthListConsumers: Handle newFromMWUser returning null in addNavigationSubtitle (T435167)]] (duration: 19m 25s) * 14:42 cmooney@cumin1003: START - Cookbook sre.hosts.reboot-single for host durum1001.eqiad.wmnet * 14:42 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 14:41 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 14:38 tgr@deploy1003: tgr: Continuing with deployment * 14:36 tgr@deploy1003: tgr: Backport for [[gerrit:1327094{{!}}SpecialMWOAuthListConsumers: Handle newFromMWUser returning null in addNavigationSubtitle (T435167)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:28 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 14:28 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:28 cmooney@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host durum3005.esams.wmnet * 14:25 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 14:23 cmooney@cumin1003: START - Cookbook sre.hosts.reboot-single for host durum3005.esams.wmnet * 14:23 tgr@deploy1003: Started scap sync-world: Backport for [[gerrit:1327094{{!}}SpecialMWOAuthListConsumers: Handle newFromMWUser returning null in addNavigationSubtitle (T435167)]] * 14:18 gengh@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:17 topranks: disable puppet on hosts running BIRD BGP to test merge of patch to systemd healtchcheck service * 14:17 gengh@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:17 gengh@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:16 gengh@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:16 elukey: upgrade spicerack on cumin1003 and cumin2003 * 14:16 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:15 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314962{{!}}static: add new dir bimi/ for BIMI SVG and PEM file (T311685)]] (duration: 10m 00s) * 14:15 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:11 kharlan@deploy1003: kharlan, sukhe: Continuing with deployment * 14:11 gengh@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:09 gengh@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:09 gengh@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:08 kharlan@deploy1003: kharlan, sukhe: Backport for [[gerrit:1314962{{!}}static: add new dir bimi/ for BIMI SVG and PEM file (T311685)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:06 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 14:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:06 gengh@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:05 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1314962{{!}}static: add new dir bimi/ for BIMI SVG and PEM file (T311685)]] * 14:05 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host krb1002.eqiad.wmnet * 14:05 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:04 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:04 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326342{{!}}Revert^2 "wmf-config/ProductionServices: set URL for urldownloader to service record"]] (duration: 07m 40s) * 13:59 kharlan@deploy1003: kharlan, sukhe: Continuing with deployment * 13:58 kharlan@deploy1003: kharlan, sukhe: Backport for [[gerrit:1326342{{!}}Revert^2 "wmf-config/ProductionServices: set URL for urldownloader to service record"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:57 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host stat1008.eqiad.wmnet with OS bookworm * 13:56 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1326342{{!}}Revert^2 "wmf-config/ProductionServices: set URL for urldownloader to service record"]] * 13:56 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 13:54 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325532{{!}}srwiki: Allow bureaucrats to add and remove event-organizer group (T434748)]] (duration: 14m 56s) * 13:54 swfrench@dns1004: END - running authdns-update * 13:52 swfrench@dns1004: START - running authdns-update * 13:51 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:51 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 13:48 kharlan@deploy1003: kharlan, danielyepezgarces: Continuing with deployment * 13:45 swfrench@cumin2003: conftool action : set/pooled=yes; selector: name=wikikube-worker2330.codfw.wmnet * 13:44 swfrench-wmf: finished etcd-main codfw -> eqiad switchover - [[phab:T435103|T435103]] * 13:44 kharlan@deploy1003: kharlan, danielyepezgarces: Backport for [[gerrit:1325532{{!}}srwiki: Allow bureaucrats to add and remove event-organizer group (T434748)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:44 swfrench@cumin2003: conftool action : set/pooled=no; selector: name=wikikube-worker2330.codfw.wmnet * 13:41 swfrench@dns1004: END - running authdns-update * 13:39 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1325532{{!}}srwiki: Allow bureaucrats to add and remove event-organizer group (T434748)]] * 13:39 swfrench@dns1004: START - running authdns-update * 13:37 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327111{{!}}Special:AbuseReview: Add "no further action needed" review action (T435020)]], [[gerrit:1327110{{!}}AbuseReview: Take the review verdict as a REST path parameter (T435020)]] (duration: 31m 43s) * 13:31 swfrench-wmf: starting etcd-main codfw -> eqiad switchover - [[phab:T435103|T435103]] * 13:28 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:28 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 13:25 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=97) rolling reboot on A:durum and A:durum * 13:24 kharlan@deploy1003: kharlan: Continuing with deployment * 13:24 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1147 * 13:24 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1147 * 13:23 kharlan@deploy1003: kharlan: Backport for [[gerrit:1327111{{!}}Special:AbuseReview: Add "no further action needed" review action (T435020)]], [[gerrit:1327110{{!}}AbuseReview: Take the review verdict as a REST path parameter (T435020)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:16 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:16 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 13:12 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:12 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 13:10 cdobbins@cumin1003: END (ERROR) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=97) Rolling upgrade of ATS on A:cp-codfw and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 13:10 cdobbins@cumin1003: END (ERROR) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=97) Rolling upgrade of ATS on A:cp-eqiad and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 13:06 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1327111{{!}}Special:AbuseReview: Add "no further action needed" review action (T435020)]], [[gerrit:1327110{{!}}AbuseReview: Take the review verdict as a REST path parameter (T435020)]] * 13:06 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-eqiad and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 13:05 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-codfw and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 13:05 cjd91: sudo -i cookbook sre.cdn.roll-upgrade-ats --query 'A:cp-eqiad' --task-id [[phab:T434478|T434478]] --reason '9.2.15 upgrade' * 13:03 swfrench@dns1004: END - running authdns-update * 13:01 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:01 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 13:01 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host ncmonitor1001.eqiad.wmnet * 13:00 swfrench@dns1004: START - running authdns-update * 12:59 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 12:59 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 12:59 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 12:59 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 12:57 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and A:durum * 12:53 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on db2902.codfw.wmnet with reason: Cloning * 12:48 cmooney@dns3003: END - running authdns-update * 12:48 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327096{{!}}Switch to redis lock manager on s4 and s8 (T366938)]] (duration: 09m 11s) * 12:46 cmooney@dns3003: START - running authdns-update * 12:45 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:45 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->eqord cct - cmooney@cumin1003" * 12:45 elukey: move the /v2/releng.* prefix on the Docker Registry to its new s3 backend - [[phab:T432829|T432829]] * 12:43 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 12:42 jelto@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 12:42 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->eqord cct - cmooney@cumin1003" * 12:42 jelto@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 12:41 jelto: update cert-manager to 1.19.6 on wikikube staging-codfw - [[phab:T427402|T427402]] * 12:40 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327096{{!}}Switch to redis lock manager on s4 and s8 (T366938)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:38 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327096{{!}}Switch to redis lock manager on s4 and s8 (T366938)]] * 12:38 blake@deploy1003: Finished scap sync-world: non-build deploy for [[phab:T417800|T417800]] (duration: 03m 52s) * 12:36 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 12:35 blake@deploy1003: Started scap sync-world: non-build deploy for [[phab:T417800|T417800]] * 12:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host krb2002.codfw.wmnet * 11:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host krb2002.codfw.wmnet * 11:49 moritzm: installing kerberos security updates * 11:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on stat1008.eqiad.wmnet with reason: host reimage * 11:44 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on stat1008.eqiad.wmnet with reason: host reimage * 11:31 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327084{{!}}Revert "Migrate database access to virtual domains" (T435305)]] (duration: 11m 02s) * 11:29 gkyziridis@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:29 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:29 kart_: Updated MinT to 2026-06-04-131507-production ([[phab:T321316|T321316]]) * 11:28 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/machinetranslation: apply * 11:28 gkyziridis@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:26 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 11:24 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:24 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:23 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:23 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:23 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/machinetranslation: apply * 11:22 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327084{{!}}Revert "Migrate database access to virtual domains" (T435305)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:21 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/machinetranslation: apply * 11:21 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:21 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:20 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327084{{!}}Revert "Migrate database access to virtual domains" (T435305)]] * 11:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1008.eqiad.wmnet with OS bookworm * 11:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet * 11:17 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/machinetranslation: apply * 11:13 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/machinetranslation: apply * 11:12 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:12 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:10 kartik@deploy1003: helmfile [staging] START helmfile.d/services/machinetranslation: apply * 11:08 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-cron: apply * 11:08 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/mw-cron: apply * 11:08 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply * 11:08 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply * 11:07 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet * 11:07 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host stat1008.eqiad.wmnet with OS bookworm * 11:06 moritzm: upgrading the new trixie URL downloaders to Squid 7.6 [[phab:T427282|T427282]] * 11:01 gkyziridis@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin2002.codfw.wmnet * 10:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin2002.codfw.wmnet * 10:45 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1169.eqiad.wmnet with OS bookworm * 10:42 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1185.eqiad.wmnet with OS bookworm * 10:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1169.eqiad.wmnet with reason: host reimage * 10:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1185.eqiad.wmnet with reason: host reimage * 10:14 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1169.eqiad.wmnet with reason: host reimage * 10:14 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1185.eqiad.wmnet with reason: host reimage * 10:11 jmm@cumin2003: END (PASS) - Cookbook sre.netbox.restart-reboot (exit_code=0) rolling reboot on A:netbox * 10:06 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1008.eqiad.wmnet with OS bookworm * 10:04 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 10:04 mpostoronca@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321579{{!}}Register the mediawiki.wikimedia_antiabuse.content_policy_score stream (T432848)]] (duration: 08m 53s) * 10:03 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 10:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 10:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 10:00 mpostoronca@deploy1003: mpostoronca: Continuing with deployment * 09:59 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1185.eqiad.wmnet with OS bookworm * 09:59 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1169.eqiad.wmnet with OS bookworm * 09:59 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.convert-disks (exit_code=0) for host ms-be1065 * 09:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1065.eqiad.wmnet with OS trixie * 09:58 mpostoronca@deploy1003: mpostoronca: Backport for [[gerrit:1321579{{!}}Register the mediawiki.wikimedia_antiabuse.content_policy_score stream (T432848)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:55 jmm@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netbox.discovery.wmnet. on all recursors * 09:55 mpostoronca@deploy1003: Started scap sync-world: Backport for [[gerrit:1321579{{!}}Register the mediawiki.wikimedia_antiabuse.content_policy_score stream (T432848)]] * 09:55 jmm@cumin2003: START - Cookbook sre.dns.wipe-cache netbox.discovery.wmnet. on all recursors * 09:52 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw2001.wikimedia.org with OS trixie * 09:51 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 09:51 jmm@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netbox.discovery.wmnet. on all recursors * 09:51 jmm@cumin2003: START - Cookbook sre.dns.wipe-cache netbox.discovery.wmnet. on all recursors * 09:51 jmm@cumin2003: START - Cookbook sre.netbox.restart-reboot rolling reboot on A:netbox * 09:46 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 09:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1056.eqiad.wmnet with OS trixie * 09:44 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 09:39 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 09:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 09:36 topranks: make HE transport circuits from magru live * 09:36 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 09:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb1003.eqiad.wmnet * 09:33 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage * 09:31 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb1003.eqiad.wmnet * 09:30 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 09:28 arnaudb@dns1006: END - running authdns-update * 09:27 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage * 09:27 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb2003.codfw.wmnet * 09:26 arnaudb@dns1006: START - running authdns-update * 09:24 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1056.eqiad.wmnet with reason: host reimage * 09:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb2003.codfw.wmnet * 09:20 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.convert-disks (exit_code=0) for host ms-be1068 * 09:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1068.eqiad.wmnet with OS trixie * 09:20 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 09:19 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "cloudvirt1057 - filippo@cumin1003" * 09:19 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "cloudvirt1057 - filippo@cumin1003" * 09:18 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1056.eqiad.wmnet with reason: host reimage * 09:18 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1057.eqiad.wmnet with OS trixie * 09:18 filippo@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 09:18 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 09:15 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 09:14 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 09:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host irc1003.wikimedia.org * 09:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:13 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1065.eqiad.wmnet with OS trixie * 09:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 09:09 moritzm: installing Postgresql security updates * 09:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host irc1003.wikimedia.org * 09:07 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 09:07 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:07 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1173.eqiad.wmnet with OS bookworm * 09:06 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw2001.wikimedia.org with OS trixie * 09:03 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.convert-disks (exit_code=0) for host ms-be1064 * 09:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1064.eqiad.wmnet with OS trixie * 09:03 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1056.eqiad.wmnet with OS trixie * 09:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1057.eqiad.wmnet with reason: host reimage * 09:01 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 08:58 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 08:56 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1057.eqiad.wmnet with reason: host reimage * 08:54 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1208.eqiad.wmnet with OS bookworm * 08:53 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2902.codfw.wmnet with OS trixie * 08:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 08:50 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1174.eqiad.wmnet with OS bookworm * 08:46 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1172.eqiad.wmnet with OS bookworm * 08:45 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1173.eqiad.wmnet with reason: host reimage * 08:41 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 08:40 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1057.eqiad.wmnet with OS trixie * 08:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1057.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 08:39 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1207.eqiad.wmnet with OS bookworm * 08:38 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2902.codfw.wmnet with reason: host reimage * 08:37 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 08:34 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1068.eqiad.wmnet with OS trixie * 08:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1208.eqiad.wmnet with reason: host reimage * 08:31 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1057.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 08:29 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 08:28 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1222.eqiad.wmnet onto db1276.eqiad.wmnet * 08:28 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1222: Pool db1222.eqiad.wmnet in after cloning * 08:28 fceratto@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2902.codfw.wmnet with reason: host reimage * 08:27 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1055.eqiad.wmnet with OS trixie * 08:27 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 08:26 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1174.eqiad.wmnet with reason: host reimage * 08:25 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 08:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-misc2002.codfw.wmnet * 08:23 topranks: reboot pfw1-codfw firewall pair to upgrade JunOS [[phab:T434865|T434865]] * 08:22 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1172.eqiad.wmnet with reason: host reimage * 08:20 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1064.eqiad.wmnet with OS trixie * 08:20 mvernon@cumin2003: START - Cookbook sre.swift.convert-disks for host ms-be1065 * 08:18 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1207.eqiad.wmnet with reason: host reimage * 08:17 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1174.eqiad.wmnet with reason: host reimage * 08:17 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1173.eqiad.wmnet with reason: host reimage * 08:17 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1172.eqiad.wmnet with reason: host reimage * 08:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host mc-misc2002.codfw.wmnet * 08:15 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1208.eqiad.wmnet with reason: host reimage * 08:15 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1207.eqiad.wmnet with reason: host reimage * 08:14 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db2902.codfw.wmnet with OS trixie * 08:14 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.16 refs [[phab:T430835|T430835]] * 08:13 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db2902.codfw.wmnet * 08:13 fceratto@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host db2902.codfw.wmnet with OS trixie * 08:10 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr[1-2]-codfw with reason: upgrade pfw1a-codfw and pfw1b-codfw pair * 08:09 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1055.eqiad.wmnet with reason: host reimage * 08:07 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on pfw1-codfw with reason: upgrade pfw1a-codfw and pfw1b-codfw pair * 08:03 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1055.eqiad.wmnet with reason: host reimage * 08:02 arnaudb@dns1006: END - running authdns-update * 08:02 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1208.eqiad.wmnet with OS bookworm * 08:02 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1207.eqiad.wmnet with OS bookworm * 08:01 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1174.eqiad.wmnet with OS bookworm * 08:01 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1173.eqiad.wmnet with OS bookworm * 08:01 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1172.eqiad.wmnet with OS bookworm * 07:59 arnaudb@dns1006: START - running authdns-update * 07:58 arnaudb@dns1006: START - running authdns-update * 07:48 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1055.eqiad.wmnet with OS trixie * 07:45 moritzm: extend the disk of ldap-rw2001 by 80G [[phab:T331699|T331699]] * 07:42 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1222: Pool db1222.eqiad.wmnet in after cloning * 07:36 mvernon@cumin2003: START - Cookbook sre.swift.convert-disks for host ms-be1068 * 07:35 mvernon@cumin2003: START - Cookbook sre.swift.convert-disks for host ms-be1064 * 07:17 moritzm: installing imagemagick security updates * 07:14 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1277: Pool back * 07:14 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon1003.wikimedia.org * 07:07 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon1003.wikimedia.org * 07:03 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1280: Pool back * 07:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon2002.wikimedia.org * 06:55 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon2002.wikimedia.org * 06:54 moritzm: installing php8.2 security updates * 06:51 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1284: Pool back * 06:49 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1222: Depool db1222.eqiad.wmnet to then clone it to db1276.eqiad.wmnet - marostegui@cumin1003 * 06:49 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1222: Depool db1222.eqiad.wmnet to then clone it to db1276.eqiad.wmnet - marostegui@cumin1003 * 06:49 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1222.eqiad.wmnet onto db1276.eqiad.wmnet * 06:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd1005.eqiad.wmnet * 06:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd1005.eqiad.wmnet * 06:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd1004.eqiad.wmnet * 06:36 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2209: db2209 repool * 06:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd1004.eqiad.wmnet * 06:32 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast3007.wikimedia.org * 06:29 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1277: Pool back * 06:28 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1277 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96190 and previous config saved to /var/cache/conftool/dbconfig/20260819-062815-marostegui.json * 06:26 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast3007.wikimedia.org * 06:22 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul1001.eqiad.wmnet * 06:18 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul1001.eqiad.wmnet * 06:18 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul1003.eqiad.wmnet * 06:18 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1280: Pool back * 06:17 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1284 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96186 and previous config saved to /var/cache/conftool/dbconfig/20260819-061743-marostegui.json * 06:14 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul1003.eqiad.wmnet * 06:14 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul1002.eqiad.wmnet * 06:10 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul1002.eqiad.wmnet * 06:10 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2048.codfw.wmnet * 06:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2048.codfw.wmnet * 06:06 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1284: Pool back * 06:06 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1284 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96184 and previous config saved to /var/cache/conftool/dbconfig/20260819-060621-marostegui.json * 06:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2048.codfw.wmnet * 05:59 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2048.codfw.wmnet * 05:51 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2209: db2209 repool * 03:16 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1186.eqiad.wmnet with OS bookworm * 02:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1186.eqiad.wmnet with reason: host reimage * 02:46 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1186.eqiad.wmnet with reason: host reimage * 02:46 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2207 [[phab:T435270|T435270]]', diff saved to https://phabricator.wikimedia.org/P96181 and previous config saved to /var/cache/conftool/dbconfig/20260819-024627-marostegui.json * 02:44 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2204 to s2 primary [[phab:T435270|T435270]]', diff saved to https://phabricator.wikimedia.org/P96180 and previous config saved to /var/cache/conftool/dbconfig/20260819-024403-marostegui.json * 02:43 marostegui: Starting s2 codfw failover from db2207 to db2204 - [[phab:T435270|T435270]] * 02:39 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2204 with weight 0 [[phab:T435270|T435270]]', diff saved to https://phabricator.wikimedia.org/P96179 and previous config saved to /var/cache/conftool/dbconfig/20260819-023951-marostegui.json * 02:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s2 [[phab:T435270|T435270]] * 02:32 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1186.eqiad.wmnet with OS bookworm * 02:29 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-worker1186.eqiad.wmnet with OS bookworm * 02:18 denisse@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2207: Depooling replica * 02:18 denisse@cumin1003: START - Cookbook sre.mysql.depool depool db2207: Depooling replica * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 48s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-18 == * 23:55 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1264845{{!}}Remove unused/redundant wgMFNoindexPages=true setting (T255458)]] (duration: 09m 42s) * 23:51 krinkle@deploy1003: krinkle: Continuing with deployment * 23:48 krinkle@deploy1003: krinkle: Backport for [[gerrit:1264845{{!}}Remove unused/redundant wgMFNoindexPages=true setting (T255458)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:45 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1264845{{!}}Remove unused/redundant wgMFNoindexPages=true setting (T255458)]] * 23:38 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326944{{!}}Retire filebackend lock manager in favour of the default one (T366938)]] (duration: 08m 55s) * 23:34 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 23:31 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326944{{!}}Retire filebackend lock manager in favour of the default one (T366938)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:29 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326944{{!}}Retire filebackend lock manager in favour of the default one (T366938)]] * 23:27 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1170.eqiad.wmnet with OS bookworm * 23:21 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1205.eqiad.wmnet with OS bookworm * 23:20 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1171.eqiad.wmnet with OS bookworm * 23:15 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1206.eqiad.wmnet with OS bookworm * 23:05 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1170.eqiad.wmnet with reason: host reimage * 23:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1205.eqiad.wmnet with reason: host reimage * 22:57 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1171.eqiad.wmnet with reason: host reimage * 22:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1206.eqiad.wmnet with reason: host reimage * 22:53 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1205.eqiad.wmnet with reason: host reimage * 22:51 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1171.eqiad.wmnet with reason: host reimage * 22:51 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1170.eqiad.wmnet with reason: host reimage * 22:50 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1206.eqiad.wmnet with reason: host reimage * 22:36 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1206.eqiad.wmnet with OS bookworm * 22:35 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1205.eqiad.wmnet with OS bookworm * 22:35 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1186.eqiad.wmnet with OS bookworm * 22:35 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1171.eqiad.wmnet with OS bookworm * 22:35 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1170.eqiad.wmnet with OS bookworm * 22:33 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-worker1194.eqiad.wmnet with OS bookworm * 22:22 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326923{{!}}Enable redis lock manager on s6 (T366938)]] (duration: 11m 52s) * 22:18 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 22:13 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326923{{!}}Enable redis lock manager on s6 (T366938)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:10 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326923{{!}}Enable redis lock manager on s6 (T366938)]] * 22:04 sbassett: Deployed security fix for [[phab:T435234|T435234]] (wmf.16) * 21:54 sbassett: Deployed security fix for [[phab:T435234|T435234]] (wmf.15) * 21:38 caro@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326925{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326926{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326929{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]], [[gerrit:1326928{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]] (duration: 0 * 21:34 caro@deploy1003: caro: Continuing with deployment * 21:33 caro@deploy1003: caro: Backport for [[gerrit:1326925{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326926{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326929{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]], [[gerrit:1326928{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]] synced to the testservers (see h * 21:31 caro@deploy1003: Started scap sync-world: Backport for [[gerrit:1326925{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326926{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326929{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]], [[gerrit:1326928{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]] * 21:24 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1204.eqiad.wmnet with reason: 1204 datanode repair [[phab:T434494|T434494]] * 21:02 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326896{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]], [[gerrit:1326897{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]] (duration: 13m 56s) * 20:58 krinkle@deploy1003: krinkle: Continuing with deployment * 20:50 krinkle@deploy1003: krinkle: Backport for [[gerrit:1326896{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]], [[gerrit:1326897{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:49 ryankemper: `an-launcher1003` terminated process group `666809` (`rest_backfill_phase1.sh`) ~20 mins ago with `sudo kill -TERM -- -666809` after its local spark driver (`--driver-memory 64g`) repeatedly exhausted memory on the 32 GB VM and caused SSH to intermittently flap; host recovered to 27 GB available memory * 20:48 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1326896{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]], [[gerrit:1326897{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]] * 20:35 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 20:33 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 20:31 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 20:31 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326870{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]], [[gerrit:1326871{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]] (duration: 07m 35s) * 20:28 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 20:26 kemayo@deploy1003: kemayo: Continuing with deployment * 20:26 ryankemper: `an-launcher1003` confirmed the host is flapping because of memory thrash. chasing down the source of the thrash * 20:25 kemayo@deploy1003: kemayo: Backport for [[gerrit:1326870{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]], [[gerrit:1326871{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:23 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1326870{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]], [[gerrit:1326871{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]] * 20:23 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 20:20 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 20:14 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1168.eqiad.wmnet with OS bookworm * 20:08 zabe: zabe@deploy1003:~$ mwscript extensions/WikimediaMaintenance/maintenance/fixFileRevisionArchiveNameDrift.php enwiki # [[phab:T428406|T428406]] * 20:08 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1204.eqiad.wmnet with OS bookworm * 20:05 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1167.eqiad.wmnet with OS bookworm * 20:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1166.eqiad.wmnet with OS bookworm * 19:54 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1203.eqiad.wmnet with OS bookworm * 19:53 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326907{{!}}Revert "Disable redis lock manager on testwiki"]] (duration: 11m 05s) * 19:50 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1168.eqiad.wmnet with reason: host reimage * 19:47 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1204.eqiad.wmnet with reason: host reimage * 19:46 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 19:44 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326907{{!}}Revert "Disable redis lock manager on testwiki"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:42 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326907{{!}}Revert "Disable redis lock manager on testwiki"]] * 19:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1167.eqiad.wmnet with reason: host reimage * 19:37 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1166.eqiad.wmnet with reason: host reimage * 19:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1203.eqiad.wmnet with reason: host reimage * 19:32 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1167.eqiad.wmnet with reason: host reimage * 19:32 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1168.eqiad.wmnet with reason: host reimage * 19:32 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1166.eqiad.wmnet with reason: host reimage * 19:31 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1204.eqiad.wmnet with reason: host reimage * 19:31 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1203.eqiad.wmnet with reason: host reimage * 19:26 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324751{{!}}InitialiseSettings: Enable 2FA enforcement on more private wikis (T428103)]], [[gerrit:1326875{{!}}Add banner notifying of upcoming 2FA enforcement (T420792)]] (duration: 31m 46s) * 19:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1204.eqiad.wmnet with OS bookworm * 19:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1203.eqiad.wmnet with OS bookworm * 19:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1168.eqiad.wmnet with OS bookworm * 19:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1167.eqiad.wmnet with OS bookworm * 19:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1166.eqiad.wmnet with OS bookworm * 19:15 denisse: rebooting kafkamon2003.codfw.wmnet - [[phab:T435162|T435162]] * 19:14 denisse: rebooting kafkamon1003.eqiad.wmnet [[phab:T435162|T435162]] * 19:13 reedy@deploy1003: reedy: Continuing with deployment * 19:12 reedy@deploy1003: reedy: Backport for [[gerrit:1324751{{!}}InitialiseSettings: Enable 2FA enforcement on more private wikis (T428103)]], [[gerrit:1326875{{!}}Add banner notifying of upcoming 2FA enforcement (T420792)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:54 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324751{{!}}InitialiseSettings: Enable 2FA enforcement on more private wikis (T428103)]], [[gerrit:1326875{{!}}Add banner notifying of upcoming 2FA enforcement (T420792)]] * 18:50 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 18:44 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.16 refs [[phab:T430835|T430835]] * 18:34 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 18:31 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 18:22 aklapper@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326852{{!}}CategoryTree: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]], [[gerrit:1326853{{!}}CategoryViewer: Allow null $html in the CategoryViewerGenerateLink hook (T435161)]], [[gerrit:1326865{{!}}Flow: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]] (duration: 09m 57s) * 18:18 aklapper@deploy1003: jforrester, aklapper: Continuing with deployment * 18:17 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 18:14 aklapper@deploy1003: jforrester, aklapper: Backport for [[gerrit:1326852{{!}}CategoryTree: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]], [[gerrit:1326853{{!}}CategoryViewer: Allow null $html in the CategoryViewerGenerateLink hook (T435161)]], [[gerrit:1326865{{!}}Flow: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki * 18:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1156.eqiad.wmnet with OS bookworm * 18:12 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 18:12 aklapper@deploy1003: Started scap sync-world: Backport for [[gerrit:1326852{{!}}CategoryTree: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]], [[gerrit:1326853{{!}}CategoryViewer: Allow null $html in the CategoryViewerGenerateLink hook (T435161)]], [[gerrit:1326865{{!}}Flow: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]] * 18:11 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 18:08 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1146.eqiad.wmnet with OS bookworm * 18:07 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1177.eqiad.wmnet with OS bookworm * 18:00 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326380{{!}}Introduce main lock manager service (T366938 T427999)]] (duration: 11m 25s) * 17:58 ladsgroup@deploy1003: ladsgroup: Rolling back deployment * 17:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1202.eqiad.wmnet with OS bookworm * 17:55 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1201.eqiad.wmnet with OS bookworm * 17:50 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326380{{!}}Introduce main lock manager service (T366938 T427999)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:48 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326380{{!}}Introduce main lock manager service (T366938 T427999)]] * 17:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1156.eqiad.wmnet with reason: host reimage * 17:46 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1146.eqiad.wmnet with reason: host reimage * 17:45 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-esams and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 17:42 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 17:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1177.eqiad.wmnet with reason: host reimage * 17:38 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1202.eqiad.wmnet with reason: host reimage * 17:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1201.eqiad.wmnet with reason: host reimage * 17:30 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1156.eqiad.wmnet with reason: host reimage * 17:29 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1177.eqiad.wmnet with reason: host reimage * 17:28 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1146.eqiad.wmnet with reason: host reimage * 17:28 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1202.eqiad.wmnet with reason: host reimage * 17:27 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1201.eqiad.wmnet with reason: host reimage * 17:25 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 17:21 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 17:14 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1202.eqiad.wmnet with OS bookworm * 17:14 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1201.eqiad.wmnet with OS bookworm * 17:14 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1177.eqiad.wmnet with OS bookworm * 17:14 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1156.eqiad.wmnet with OS bookworm * 17:14 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1146.eqiad.wmnet with OS bookworm * 17:09 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326885{{!}}w/deployment-info.php: Handle new file format (T434726)]] (duration: 07m 15s) * 17:08 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 17:05 dancy@deploy1003: dancy: Continuing with deployment * 17:04 dancy@deploy1003: dancy: Backport for [[gerrit:1326885{{!}}w/deployment-info.php: Handle new file format (T434726)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:02 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1326885{{!}}w/deployment-info.php: Handle new file format (T434726)]] * 16:50 dancy@deploy1003: Finished scap sync-world: Testing [[phab:T434726|T434726]] (duration: 06m 40s) * 16:43 dancy@deploy1003: Started scap sync-world: Testing [[phab:T434726|T434726]] * 16:43 dancy@deploy1003: Installation of scap version "4.282.0" completed for 3 hosts * 16:41 dancy@deploy1003: Installing scap version "4.282.0" for 3 host(s) * 16:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1145.eqiad.wmnet with OS bookworm * 16:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1200.eqiad.wmnet with OS bookworm * 16:32 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1199.eqiad.wmnet with OS bookworm * 16:18 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1145.eqiad.wmnet with reason: host reimage * 16:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1200.eqiad.wmnet with reason: host reimage * 16:09 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1199.eqiad.wmnet with reason: host reimage * 16:05 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-esams and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 16:05 cjd91: sudo -i cookbook sre.cdn.roll-upgrade-ats --query 'A:cp-esams' --task-id [[phab:T434478|T434478]] --reason '9.2.15 upgrade' * 16:03 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1200.eqiad.wmnet with reason: host reimage * 16:02 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1145.eqiad.wmnet with reason: host reimage * 16:02 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1199.eqiad.wmnet with reason: host reimage * 15:50 moritzm: installing zip security updates * 15:48 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1200.eqiad.wmnet with OS bookworm * 15:47 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1199.eqiad.wmnet with OS bookworm * 15:47 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1145.eqiad.wmnet with OS bookworm * 15:41 topranks: bounce PIC 0/0 on cr1-magru to set port to 40G * 15:37 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1236.eqiad.wmnet with OS bookworm * 15:29 aikochou@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'ores-legacy' for release 'main' . * 15:26 aikochou@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'ores-legacy' for release 'main' . * 15:20 aikochou@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'ores-legacy' for release 'main' . * 15:14 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2002.codfw.wmnet * 15:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1236.eqiad.wmnet with reason: host reimage * 15:12 moritzm: failover ganeti master in codfw to ganeti2047 * 15:09 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1236.eqiad.wmnet with reason: host reimage * 15:09 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2044.codfw.wmnet * 15:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2002.codfw.wmnet * 15:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2044.codfw.wmnet * 15:04 brennen@deploy1003: Finished deploy [phabricator/deployment@6b9b6ff]: deploy phab1004 for [[phab:T435213|T435213]] (duration: 01m 01s) * 15:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2044.codfw.wmnet * 15:03 brennen@deploy1003: Started deploy [phabricator/deployment@6b9b6ff]: deploy phab1004 for [[phab:T435213|T435213]] * 15:03 brennen@deploy1003: Finished deploy [phabricator/deployment@6b9b6ff]: deploy phab2003 for [[phab:T435213|T435213]] (duration: 00m 57s) * 15:02 brennen@deploy1003: Started deploy [phabricator/deployment@6b9b6ff]: deploy phab2003 for [[phab:T435213|T435213]] * 14:57 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply * 14:57 arnaudb@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on phab2003.codfw.wmnet,phab[1004-1006].eqiad.wmnet with reason: maintenance * 14:56 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2044.codfw.wmnet * 14:55 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply * 14:53 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1236.eqiad.wmnet with OS bookworm * 14:51 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2043.codfw.wmnet * 14:51 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2043.codfw.wmnet * 14:45 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2043.codfw.wmnet * 14:33 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2043.codfw.wmnet * 14:22 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2042.codfw.wmnet * 14:22 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2042.codfw.wmnet * 14:21 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db2902.codfw.wmnet with OS trixie * 14:16 elukey: uploaded spicerack_13.2.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia * 14:16 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2042.codfw.wmnet * 14:04 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2042.codfw.wmnet * 14:04 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2041.codfw.wmnet * 14:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2041.codfw.wmnet * 13:59 phuedx@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: apply * 13:59 phuedx@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-main: apply * 13:59 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326838{{!}}Enable Suggested Investigations on hewiki (T435146)]] (duration: 11m 50s) * 13:58 phuedx@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: apply * 13:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2041.codfw.wmnet * 13:57 phuedx@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-main: apply * 13:57 phuedx@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-main: apply * 13:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard1003.eqiad.wmnet * 13:57 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-main: apply * 13:55 phuedx@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-logging-external: apply * 13:55 phuedx@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-logging-external: apply * 13:54 phuedx@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-logging-external: apply * 13:54 stran@deploy1003: stran: Continuing with deployment * 13:54 phuedx@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-logging-external: apply * 13:54 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-logging-external: apply * 13:54 phuedx@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-logging-external: apply * 13:53 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-logging-external: apply * 13:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard1003.eqiad.wmnet * 13:53 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2041.codfw.wmnet * 13:51 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db2902.codfw.wmnet - fceratto@cumin1003" * 13:51 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db2902.codfw.wmnet - fceratto@cumin1003" * 13:51 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2026.codfw.wmnet * 13:51 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard2003.codfw.wmnet * 13:50 phuedx@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: apply * 13:50 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-eqiad * 13:50 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp1001.eqiad.wmnet * 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp1001.eqiad.wmnet * 13:50 phuedx@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: apply * 13:49 phuedx@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: apply * 13:49 stran@deploy1003: stran: Backport for [[gerrit:1326838{{!}}Enable Suggested Investigations on hewiki (T435146)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:48 phuedx@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: apply * 13:48 phuedx@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics: apply * 13:47 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard2003.codfw.wmnet * 13:47 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics: apply * 13:47 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1326838{{!}}Enable Suggested Investigations on hewiki (T435146)]] * 13:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp1001.eqiad.wmnet * 13:44 cdanis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 13:43 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp1001.eqiad.wmnet * 13:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1315-1327].eqiad.wmnet * 13:43 cdanis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 13:43 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1315-1327].eqiad.wmnet * 13:40 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki1001.eqiad.wmnet * 13:35 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1315-1327].eqiad.wmnet * 13:34 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host rpki1001.eqiad.wmnet * 13:27 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1315-1327].eqiad.wmnet * 13:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1302-1314].eqiad.wmnet * 13:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1302-1314].eqiad.wmnet * 13:21 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326775{{!}}SI: Instrument case update on first edit (T435048)]] (duration: 07m 12s) * 13:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1302-1314].eqiad.wmnet * 13:17 stran@deploy1003: stran: Continuing with deployment * 13:16 stran@deploy1003: stran: Backport for [[gerrit:1326775{{!}}SI: Instrument case update on first edit (T435048)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:14 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1326775{{!}}SI: Instrument case update on first edit (T435048)]] * 13:13 moritzm: installing util-linux security updates * 13:11 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1302-1314].eqiad.wmnet * 13:10 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326232{{!}}prv: Enable parsoid rendering for 5 wikis (T435115)]] (duration: 08m 17s) * 13:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1288-1289,1291-1301].eqiad.wmnet * 13:10 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply * 13:10 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1288-1289,1291-1301].eqiad.wmnet * 13:10 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply * 13:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki2003.codfw.wmnet * 13:06 jgiannelos@deploy1003: jgiannelos: Continuing with deployment * 13:05 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host rpki2003.codfw.wmnet * 13:04 jgiannelos@deploy1003: jgiannelos: Backport for [[gerrit:1326232{{!}}prv: Enable parsoid rendering for 5 wikis (T435115)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1288-1289,1291-1301].eqiad.wmnet * 13:02 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1326232{{!}}prv: Enable parsoid rendering for 5 wikis (T435115)]] * 12:53 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1288-1289,1291-1301].eqiad.wmnet * 12:53 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1273,1275-1287].eqiad.wmnet * 12:53 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1273,1275-1287].eqiad.wmnet * 12:52 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt1002.wikimedia.org * 12:51 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db2902.codfw.wmnet on all recursors * 12:51 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db2902.codfw.wmnet on all recursors * 12:51 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:51 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db2902.codfw.wmnet - fceratto@cumin1003" * 12:51 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db2902.codfw.wmnet - fceratto@cumin1003" * 12:46 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt1002.wikimedia.org * 12:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt2002.wikimedia.org * 12:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1273,1275-1287].eqiad.wmnet * 12:42 dhinus: repooled clouddb1032 that was currently <nowiki>{</nowiki>"weight": 0, "pooled": "inactive"<nowiki>}</nowiki> for both s4 and s6 * 12:41 dhinus: also depooled clouddb1017 (forgot it in the previous list) * 12:41 fnegri@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet * 12:40 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 12:40 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db2902.codfw.wmnet * 12:40 fnegri@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032.eqiad.wmnet * 12:40 fnegri@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032 * 12:39 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1017.eqiad.wmnet * 12:39 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt2002.wikimedia.org * 12:38 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1020.eqiad.wmnet * 12:38 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1018.eqiad.wmnet * 12:38 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet * 12:37 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1014.eqiad.wmnet * 12:37 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1013.eqiad.wmnet * 12:37 dhinus: depool again clouddb10[13,14,16,18,20] that were repooled by the cookbook sre.mysql.multiinstance_reboot * 12:36 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1273,1275-1287].eqiad.wmnet * 12:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1248-1261].eqiad.wmnet * 12:36 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1248-1261].eqiad.wmnet * 12:35 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2026.codfw.wmnet * 12:34 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 12:34 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 12:34 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 12:34 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 12:29 lucaswerkmeister-wmde@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 12:28 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2026.codfw.wmnet * 12:28 lucaswerkmeister-wmde@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 12:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1248-1261].eqiad.wmnet * 12:27 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.addnode (exit_code=0) for new host ganeti2046.codfw.wmnet to cluster codfw and group A * 12:26 moritzm: readded ganeti2046 to the codfw cluster following firmware update and reimage [[phab:T434681|T434681]] * 12:23 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply * 12:23 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply * 12:23 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply * 12:22 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply * 12:22 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply * 12:22 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply * 12:21 jmm@cumin2003: START - Cookbook sre.ganeti.addnode for new host ganeti2046.codfw.wmnet to cluster codfw and group A * 12:21 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1169.eqiad.wmnet onto db1283.eqiad.wmnet * 12:21 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1169: Pool db1169.eqiad.wmnet in after cloning * 12:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1248-1261].eqiad.wmnet * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1149-1153,1158,1240-1247].eqiad.wmnet * 12:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1149-1153,1158,1240-1247].eqiad.wmnet * 12:11 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1149-1153,1158,1240-1247].eqiad.wmnet * 12:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1168.eqiad.wmnet onto db1282.eqiad.wmnet * 12:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1168: Pool db1168.eqiad.wmnet in after cloning * 12:06 jmm@cumin2003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-test-eqiad * 12:05 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2026.codfw.wmnet * 12:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1149-1153,1158,1240-1247].eqiad.wmnet * 12:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1128-1134,1142-1148].eqiad.wmnet * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2046.codfw.wmnet * 12:01 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1128-1134,1142-1148].eqiad.wmnet * 11:57 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2025.codfw.wmnet * 11:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2025.codfw.wmnet * 11:54 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2046.codfw.wmnet * 11:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1128-1134,1142-1148].eqiad.wmnet * 11:50 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2025.codfw.wmnet * 11:46 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1128-1134,1142-1148].eqiad.wmnet * 11:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1114-1127].eqiad.wmnet * 11:45 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1114-1127].eqiad.wmnet * 11:39 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db2901.codfw.wmnet * 11:39 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db2901.codfw.wmnet * 11:36 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2025.codfw.wmnet * 11:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1114-1127].eqiad.wmnet * 11:35 fceratto@cumin1003: END (ERROR) - Cookbook sre.ganeti.makevm (exit_code=93) for new host db1901.eqiad.wmnet * 11:35 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1169: Pool db1169.eqiad.wmnet in after cloning * 11:35 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 11:30 jmm@cumin2003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-test-eqiad * 11:27 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1114-1127].eqiad.wmnet * 11:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1076-1081,1084-1087,1093-1095,1113].eqiad.wmnet * 11:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1076-1081,1084-1087,1093-1095,1113].eqiad.wmnet * 11:25 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1168: Pool db1168.eqiad.wmnet in after cloning * 11:24 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1194.eqiad.wmnet with OS bookworm * 11:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1181.eqiad.wmnet with OS bookworm * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2050.codfw.wmnet * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2050.codfw.wmnet * 11:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1076-1081,1084-1087,1093-1095,1113].eqiad.wmnet * 11:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2050.codfw.wmnet * 11:14 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2004.codfw.wmnet * 11:10 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2050.codfw.wmnet * 11:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1076-1081,1084-1087,1093-1095,1113].eqiad.wmnet * 11:09 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1286: Pool back * 11:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1045-1050,1056-1057,1064-1066,1073-1075].eqiad.wmnet * 11:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1045-1050,1056-1057,1064-1066,1073-1075].eqiad.wmnet * 11:08 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2004.codfw.wmnet * 11:07 moritzm: installing PHP 8.4 security updates * 11:06 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on an-worker1194.eqiad.wmnet with reason: host reimage * 11:06 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1194.eqiad.wmnet with reason: host reimage * 11:05 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2049.codfw.wmnet * 11:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2049.codfw.wmnet * 11:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid1003.eqiad.wmnet * 11:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1045-1050,1056-1057,1064-1066,1073-1075].eqiad.wmnet * 10:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1181.eqiad.wmnet with reason: host reimage * 10:59 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2049.codfw.wmnet * 10:59 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid1003.eqiad.wmnet * 10:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid2003.codfw.wmnet * 10:54 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1181.eqiad.wmnet with reason: host reimage * 10:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid2003.codfw.wmnet * 10:51 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1901.eqiad.wmnet on all recursors * 10:51 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1901.eqiad.wmnet on all recursors * 10:51 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:51 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:51 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:50 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2049.codfw.wmnet * 10:50 blake@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:50 blake@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:49 blake@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:48 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1045-1050,1056-1057,1064-1066,1073-1075].eqiad.wmnet * 10:48 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1044].eqiad.wmnet * 10:48 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1044].eqiad.wmnet * 10:47 blake@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:46 blake@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:46 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2047.codfw.wmnet * 10:46 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:46 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2047.codfw.wmnet * 10:46 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1901.eqiad.wmnet * 10:46 blake@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:44 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1168: Depool db1168.eqiad.wmnet to then clone it to db1282.eqiad.wmnet - marostegui@cumin1003 * 10:44 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1168: Depool db1168.eqiad.wmnet to then clone it to db1282.eqiad.wmnet - marostegui@cumin1003 * 10:44 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1168.eqiad.wmnet onto db1282.eqiad.wmnet * 10:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2047.codfw.wmnet * 10:40 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1044].eqiad.wmnet * 10:37 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2047.codfw.wmnet * 10:37 fceratto@cumin1003: END (ERROR) - Cookbook sre.ganeti.makevm (exit_code=93) for new host db1901.eqiad.wmnet * 10:36 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 10:34 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:33 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2032.codfw.wmnet * 10:33 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2032.codfw.wmnet * 10:32 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1044].eqiad.wmnet * 10:32 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-eqiad * 10:27 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2032.codfw.wmnet * 10:24 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1286: Pool back * 10:24 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1286 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96163 and previous config saved to /var/cache/conftool/dbconfig/20260818-102431-marostegui.json * 10:22 moritzm: installing Django security updates * 10:20 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul2002.codfw.wmnet * 10:20 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul2001.codfw.wmnet * 10:17 blake@deploy1003: Finished scap sync-world: no-build deployment for [[phab:T417800|T417800]] (duration: 04m 40s) * 10:16 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul2002.codfw.wmnet * 10:16 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul2001.codfw.wmnet * 10:15 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2032.codfw.wmnet * 10:14 blake@deploy1003: Started scap sync-world: no-build deployment for [[phab:T417800|T417800]] * 10:12 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul2003.codfw.wmnet * 10:12 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy2001.codfw.wmnet * 10:12 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy3001.esams.wmnet * 10:12 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy1001.eqiad.wmnet * 10:08 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2031.codfw.wmnet * 10:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2031.codfw.wmnet * 10:08 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul2003.codfw.wmnet * 10:08 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy2001.codfw.wmnet * 10:08 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy3001.esams.wmnet * 10:08 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy1001.eqiad.wmnet * 10:07 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy1002.eqiad.wmnet * 10:05 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy2002.codfw.wmnet * 10:05 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy3002.esams.wmnet * 10:04 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy1002.eqiad.wmnet * 10:03 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy4003.ulsfo.wmnet * 10:03 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy5003.eqsin.wmnet * 10:02 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2031.codfw.wmnet * 10:02 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1169: Depool db1169.eqiad.wmnet to then clone it to db1283.eqiad.wmnet - marostegui@cumin1003 * 10:01 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy2002.codfw.wmnet * 10:01 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy4004.ulsfo.wmnet * 10:01 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy3002.esams.wmnet * 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1169: Depool db1169.eqiad.wmnet to then clone it to db1283.eqiad.wmnet - marostegui@cumin1003 * 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1169.eqiad.wmnet onto db1283.eqiad.wmnet * 10:01 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy5004.eqsin.wmnet * 09:59 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy4003.ulsfo.wmnet * 09:59 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy4004.ulsfo.wmnet * 09:59 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy5003.eqsin.wmnet * 09:59 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy7001.magru.wmnet * 09:59 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy5004.eqsin.wmnet * 09:58 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy7002.magru.wmnet * 09:57 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2031.codfw.wmnet * 09:57 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy6001.drmrs.wmnet * 09:57 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy6002.drmrs.wmnet * 09:54 filippo@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for 10 hosts * 09:54 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast4006.wikimedia.org * 09:53 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy6001.drmrs.wmnet * 09:53 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy6002.drmrs.wmnet * 09:53 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host releases1003.eqiad.wmnet * 09:53 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy7001.magru.wmnet * 09:53 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host people2004.codfw.wmnet * 09:52 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host people1005.eqiad.wmnet * 09:52 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy7002.magru.wmnet * 09:50 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host releases2003.codfw.wmnet * 09:49 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host releases2003.codfw.wmnet * 09:49 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host releases1003.eqiad.wmnet * 09:49 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host people2004.codfw.wmnet * 09:48 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host people1005.eqiad.wmnet * 09:46 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2040.codfw.wmnet * 09:46 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2040.codfw.wmnet * 09:43 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1901.eqiad.wmnet on all recursors * 09:42 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1901.eqiad.wmnet on all recursors * 09:42 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:42 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:42 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2040.codfw.wmnet * 09:40 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp2005.wikimedia.org * 09:36 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp2005.wikimedia.org * 09:31 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:31 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1901.eqiad.wmnet * 09:31 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1901.eqiad.wmnet * 09:31 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:31 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1901.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 09:31 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1901.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 09:29 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2040.codfw.wmnet * 09:29 slyngshede@dns1004: END - running authdns-update * 09:28 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2039.codfw.wmnet * 09:28 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2039.codfw.wmnet * 09:27 slyngshede@dns1004: START - running authdns-update * 09:26 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test2005.wikimedia.org * 09:22 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2039.codfw.wmnet * 09:22 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test2005.wikimedia.org * 09:22 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp1005.wikimedia.org * 09:21 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:19 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2039.codfw.wmnet * 09:19 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2038.codfw.wmnet * 09:18 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1287: Pool back * 09:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2038.codfw.wmnet * 09:18 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp1005.wikimedia.org * 09:18 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test1005.wikimedia.org * 09:17 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1901.eqiad.wmnet * 09:15 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db1901.eqiad.wmnet * 09:15 fceratto@cumin1003: END (ERROR) - Cookbook sre.dns.netbox (exit_code=97) * 09:14 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test1005.wikimedia.org * 09:13 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2038.codfw.wmnet * 09:13 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:13 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1901.eqiad.wmnet * 09:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast4006.wikimedia.org * 09:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1288: Pool back * 09:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast5005.wikimedia.org * 09:03 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2038.codfw.wmnet * 08:57 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2037.codfw.wmnet * 08:57 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast5005.wikimedia.org * 08:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2037.codfw.wmnet * 08:56 filippo@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 10 hosts * 08:52 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2037.codfw.wmnet * 08:51 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1289: Pool back * 08:35 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudidp2001-dev.codfw.wmnet * 08:34 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1002-dev.eqiad.wmnet * 08:33 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1287: Pool back * 08:33 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1287 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96150 and previous config saved to /var/cache/conftool/dbconfig/20260818-083311-marostegui.json * 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1001-dev.eqiad.wmnet * 08:31 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudidp2001-dev.codfw.wmnet * 08:30 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1002-dev.eqiad.wmnet * 08:30 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2037.codfw.wmnet * 08:29 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1001-dev.eqiad.wmnet * 08:28 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2036.codfw.wmnet * 08:28 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2036.codfw.wmnet * 08:25 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1288: Pool back * 08:23 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 08:23 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1181.eqiad.wmnet with OS bookworm * 08:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2036.codfw.wmnet * 08:22 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1288 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96148 and previous config saved to /var/cache/conftool/dbconfig/20260818-082234-marostegui.json * 08:20 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2036.codfw.wmnet * 08:18 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2035.codfw.wmnet * 08:18 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudvirt1057.eqiad.wmnet with OS trixie * 08:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2035.codfw.wmnet * 08:13 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2035.codfw.wmnet * 08:06 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2035.codfw.wmnet * 08:05 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1289: Pool back * 08:05 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1289 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96145 and previous config saved to /var/cache/conftool/dbconfig/20260818-080531-marostegui.json * 07:54 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 07:51 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326439{{!}}Fix "mathjax_ignore" handling around forcemathmode attribute (T434686)]] (duration: 13m 15s) * 07:48 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 07:47 krinkle@deploy1003: krinkle: Continuing with deployment * 07:44 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ganeti2046.codfw.wmnet with OS bookworm * 07:40 krinkle@deploy1003: krinkle: Backport for [[gerrit:1326439{{!}}Fix "mathjax_ignore" handling around forcemathmode attribute (T434686)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:39 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 07:38 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1326439{{!}}Fix "mathjax_ignore" handling around forcemathmode attribute (T434686)]] * 07:34 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 07:31 samwilson@deploy1003: Finished scap sync-world: Backport for [[gerrit:701016{{!}}InitialiseSettings and -labs: Remove redundant feature flag $wgWikisourceEnableOcr (T285311)]] (duration: 07m 47s) * 07:30 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1048.eqiad.wmnet * 07:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1048.eqiad.wmnet * 07:29 XioNoX: add gnmic 0.47.0 to bookworm and trixie reprepro * 07:28 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ganeti2046.codfw.wmnet with reason: host reimage * 07:27 samwilson@deploy1003: samwilson: Continuing with deployment * 07:25 samwilson@deploy1003: samwilson: Backport for [[gerrit:701016{{!}}InitialiseSettings and -labs: Remove redundant feature flag $wgWikisourceEnableOcr (T285311)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:25 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 07:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1048.eqiad.wmnet * 07:24 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ganeti2046.codfw.wmnet with reason: host reimage * 07:23 samwilson@deploy1003: Started scap sync-world: Backport for [[gerrit:701016{{!}}InitialiseSettings and -labs: Remove redundant feature flag $wgWikisourceEnableOcr (T285311)]] * 07:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1198.eqiad.wmnet with OS bookworm * 07:18 samwilson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326450{{!}}InitialiseSettings.php: Enable Bulk OCR on pawikisource (T434648)]] (duration: 12m 17s) * 07:17 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 07:11 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1048.eqiad.wmnet * 07:11 samwilson@deploy1003: samwilson: Continuing with deployment * 07:11 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ganeti2046.codfw.wmnet with OS bookworm * 07:10 samwilson@deploy1003: samwilson: Backport for [[gerrit:1326450{{!}}InitialiseSettings.php: Enable Bulk OCR on pawikisource (T434648)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1057.eqiad.wmnet with OS trixie * 07:05 samwilson@deploy1003: Started scap sync-world: Backport for [[gerrit:1326450{{!}}InitialiseSettings.php: Enable Bulk OCR on pawikisource (T434648)]] * 07:02 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1198.eqiad.wmnet with reason: host reimage * 07:01 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1056.eqiad.wmnet with OS trixie * 07:01 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1056.eqiad.wmnet with OS trixie * 07:00 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1056.eqiad.wmnet with OS trixie * 07:00 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1056.eqiad.wmnet with OS trixie * 06:59 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudvirt1055.eqiad.wmnet with OS trixie * 06:58 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1198.eqiad.wmnet with reason: host reimage * 06:53 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1055.eqiad.wmnet with OS trixie * 06:52 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudvirt1054.eqiad.wmnet with OS trixie * 06:44 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1194.eqiad.wmnet with OS bookworm * 06:44 XioNoX: upgrade eqsin gnmic to 0.47.0 * 06:43 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1198.eqiad.wmnet with OS bookworm * 06:41 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1054.eqiad.wmnet with OS trixie * 06:09 arnaudb@cumin1003: END (PASS) - Cookbook sre.gerrit.restart-gerrit (exit_code=0) Restarting Gerrit on gerrit2002 * 06:06 arnaudb@cumin1003: START - Cookbook sre.gerrit.restart-gerrit Restarting Gerrit on gerrit2002 * 06:06 arnaudb@cumin1003: END (PASS) - Cookbook sre.gerrit.restart-gerrit (exit_code=0) Restarting Gerrit on gerrit1003 * 06:04 arnaudb@cumin1003: START - Cookbook sre.gerrit.restart-gerrit Restarting Gerrit on gerrit1003 * 06:02 arnaudb@cumin1003: END (PASS) - Cookbook sre.gerrit.restart-gerrit (exit_code=0) Restarting Gerrit on gerrit2003 * 06:00 arnaudb@cumin1003: START - Cookbook sre.gerrit.restart-gerrit Restarting Gerrit on gerrit2003 * 05:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1155.eqiad.wmnet with OS bookworm * 05:38 arnaudb: updating prometheusBearerToken on gerrit * 05:28 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1155.eqiad.wmnet with reason: host reimage * 05:23 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1155.eqiad.wmnet with reason: host reimage * 05:06 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1155.eqiad.wmnet with OS bookworm * 04:57 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1144.eqiad.wmnet with OS bookworm * 04:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1144.eqiad.wmnet with reason: host reimage * 04:29 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1144.eqiad.wmnet with reason: host reimage * 04:14 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1144.eqiad.wmnet with OS bookworm * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.13 (duration: 02m 23s) * 03:45 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1197.eqiad.wmnet with OS bookworm * 03:41 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1165.eqiad.wmnet with OS bookworm * 03:38 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.16 refs [[phab:T430835|T430835]] (duration: 34m 43s) * 03:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1164.eqiad.wmnet with OS bookworm * 03:35 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1196.eqiad.wmnet with OS bookworm * 03:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1163.eqiad.wmnet with OS bookworm * 03:22 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1197.eqiad.wmnet with reason: host reimage * 03:18 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1165.eqiad.wmnet with reason: host reimage * 03:15 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1196.eqiad.wmnet with reason: host reimage * 03:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1164.eqiad.wmnet with reason: host reimage * 03:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1163.eqiad.wmnet with reason: host reimage * 03:05 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1165.eqiad.wmnet with reason: host reimage * 03:04 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1164.eqiad.wmnet with reason: host reimage * 03:04 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1197.eqiad.wmnet with reason: host reimage * 03:04 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1196.eqiad.wmnet with reason: host reimage * 03:04 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1163.eqiad.wmnet with reason: host reimage * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.16 refs [[phab:T430835|T430835]] * 02:50 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1197.eqiad.wmnet with OS bookworm * 02:49 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1196.eqiad.wmnet with OS bookworm * 02:49 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1165.eqiad.wmnet with OS bookworm * 02:49 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1164.eqiad.wmnet with OS bookworm * 02:48 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1163.eqiad.wmnet with OS bookworm * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 46s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:15 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-codfw: Set storage compatability to NONE — [[phab:T433026|T433026]] - eevans@cumin1003 * 00:38 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-codfw: Set storage compatability to NONE — [[phab:T433026|T433026]] - eevans@cumin1003 * 00:11 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326373{{!}}PersonalDashboard: add newly renamed *ReviewChangesMlModel setting (T422148)]] (duration: 07m 07s) * 00:09 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-eqiad: Set storage compatability to NONE — [[phab:T433026|T433026]] - eevans@cumin1003 * 00:07 musikanimal@deploy1003: musikanimal: Continuing with deployment * 00:06 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1326373{{!}}PersonalDashboard: add newly renamed *ReviewChangesMlModel setting (T422148)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:04 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1326373{{!}}PersonalDashboard: add newly renamed *ReviewChangesMlModel setting (T422148)]] == 2026-08-17 == * 23:30 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-eqiad: Set storage compatability to NONE — [[phab:T433026|T433026]] - eevans@cumin1003 * 23:05 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-codfw: Set storage compatability to UPGRADING — [[phab:T433026|T433026]] - eevans@cumin1003 * 22:28 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-codfw: Set storage compatability to UPGRADING — [[phab:T433026|T433026]] - eevans@cumin1003 * 21:46 logmsgbot: jforrester Deployed security patch for [[phab:T435085|T435085]] * 21:39 swfrench@deploy1003: mwscript-k8s job started: purgeList.php # [[phab:T432412|T432412]] * 21:37 maryum: Undeploy security fix for [[phab:T433020|T433020]] * 21:23 maryum: Deployed security fix for [[phab:T433020|T433020]] * 21:21 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 21:21 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 21:14 maryum: Deployed security fix for [[phab:T434967|T434967]] * 20:58 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-eqiad: Set storage compatability to UPGRADING — [[phab:T433026|T433026]] - eevans@cumin1003 * 20:40 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324320{{!}}[itwiki/slwiki/tgwiki] Remove temporary Wikipedia 25 logos permanently (already reverted) (T414265 T414320 T415307)]] (duration: 06m 55s) * 20:40 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 20:39 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 20:36 cjming@deploy1003: cjming, superpes: Continuing with deployment * 20:35 cjming@deploy1003: cjming, superpes: Backport for [[gerrit:1324320{{!}}[itwiki/slwiki/tgwiki] Remove temporary Wikipedia 25 logos permanently (already reverted) (T414265 T414320 T415307)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:33 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1324320{{!}}[itwiki/slwiki/tgwiki] Remove temporary Wikipedia 25 logos permanently (already reverted) (T414265 T414320 T415307)]] * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ttmserver-test: apply * 20:30 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326361{{!}}Remove escaped paths in app site association file (T432412)]] (duration: 13m 28s) * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ttmserver-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-toolhub-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-toolhub-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-toolhub-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-toolhub-test: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-toolhub: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-toolhub: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-toolhub: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-toolhub: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-test: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-test: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 20:27 inflatador: bking@deploy1003 `charlie --services_dir dse-k8s-services -s opensearch-* -e dse-k8s-* apply` [[phab:T435125|T435125]] * 20:27 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-apifeatureusage-test: apply * 20:27 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-apifeatureusage-test: apply * 20:27 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-apifeatureusage-test: apply * 20:27 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-apifeatureusage-test: apply * 20:27 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-apifeatureusage: apply * 20:26 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-apifeatureusage: apply * 20:26 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-apifeatureusage: apply * 20:26 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-apifeatureusage: apply * 20:26 cjming@deploy1003: cjming, tsev: Continuing with deployment * 20:24 inflatador: bking@deploy1003 `charlie --services_dir dse-k8s-services -s opensearch-* -e dse-k8s-*` [[phab:T435125|T435125]] * 20:21 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-eqiad: Set storage compatability to UPGRADING — [[phab:T433026|T433026]] - eevans@cumin1003 * 20:19 cjming@deploy1003: cjming, tsev: Backport for [[gerrit:1326361{{!}}Remove escaped paths in app site association file (T432412)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:17 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1326361{{!}}Remove escaped paths in app site association file (T432412)]] * 20:15 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326077{{!}}InstrumentConstructiveEdits: anchor all runs to the nearest `interval` (T431493)]] (duration: 06m 25s) * 20:11 cjming@deploy1003: cjming: Continuing with deployment * 20:10 cjming@deploy1003: cjming: Backport for [[gerrit:1326077{{!}}InstrumentConstructiveEdits: anchor all runs to the nearest `interval` (T431493)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:08 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1326077{{!}}InstrumentConstructiveEdits: anchor all runs to the nearest `interval` (T431493)]] * 20:06 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1160.eqiad.wmnet with OS bookworm * 20:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1162.eqiad.wmnet with OS bookworm * 19:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1184.eqiad.wmnet with OS bookworm * 19:55 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1161.eqiad.wmnet with OS bookworm * 19:49 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1195.eqiad.wmnet with OS bookworm * 19:49 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-test: apply * 19:49 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-test: apply * 19:44 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1160.eqiad.wmnet with reason: host reimage * 19:41 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-codfw: Upgrade to Java 17 — [[phab:T433026|T433026]] - eevans@cumin1003 * 19:39 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1162.eqiad.wmnet with reason: host reimage * 19:36 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-eqsin and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 19:36 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-test: apply * 19:36 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1184.eqiad.wmnet with reason: host reimage * 19:32 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1161.eqiad.wmnet with reason: host reimage * 19:29 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1195.eqiad.wmnet with reason: host reimage * 19:26 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1161.eqiad.wmnet with reason: host reimage * 19:26 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1160.eqiad.wmnet with reason: host reimage * 19:26 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1184.eqiad.wmnet with reason: host reimage * 19:26 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1162.eqiad.wmnet with reason: host reimage * 19:25 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1195.eqiad.wmnet with reason: host reimage * 19:19 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-test: apply * 19:11 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1195.eqiad.wmnet with OS bookworm * 19:10 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1194.eqiad.wmnet with OS bookworm * 19:10 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1184.eqiad.wmnet with OS bookworm * 19:10 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1162.eqiad.wmnet with OS bookworm * 19:10 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1161.eqiad.wmnet with OS bookworm * 19:10 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1160.eqiad.wmnet with OS bookworm * 19:10 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 19:09 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 19:03 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-codfw: Upgrade to Java 17 — [[phab:T433026|T433026]] - eevans@cumin1003 * 18:58 dancy@deploy1003: Finished scap sync-world: testing [[phab:T375514|T375514]] (duration: 03m 13s) * 18:55 dancy@deploy1003: Started scap sync-world: testing [[phab:T375514|T375514]] * 18:55 dwisehaupt@dns1006: END - running authdns-update * 18:54 dancy@deploy1003: Installation of scap version "4.281.1" completed for 3 hosts * 18:54 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2009.codfw.wmnet * 18:53 dwisehaupt@dns1006: START - running authdns-update * 18:52 dancy@deploy1003: Installing scap version "4.281.1" for 3 host(s) * 18:47 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2009.codfw.wmnet * 18:41 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2008.codfw.wmnet * 18:34 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2008.codfw.wmnet * 18:30 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2007.codfw.wmnet * 18:23 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2007.codfw.wmnet * 18:16 dwisehaupt@dns1005: END - running authdns-update * 18:14 dwisehaupt@dns1005: START - running authdns-update * 18:04 swfrench@deploy1003: Finished scap sync-world: Deploy "Point Test Wiki to new docroot" - [[phab:T432412|T432412]] (duration: 26m 03s) * 18:00 swfrench@deploy1003: swfrench: Continuing with deployment * 17:51 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1057.eqiad.wmnet with OS trixie * 17:47 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1180.eqiad.wmnet with OS bookworm * 17:39 swfrench@deploy1003: swfrench: Deploy "Point Test Wiki to new docroot" - [[phab:T432412|T432412]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:39 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-eqsin and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 17:38 swfrench@deploy1003: Started scap sync-world: Deploy "Point Test Wiki to new docroot" - [[phab:T432412|T432412]] * 17:36 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1159.eqiad.wmnet with OS bookworm * 17:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1158.eqiad.wmnet with OS bookworm * 17:30 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-eqiad: Upgrade to Java 17 — [[phab:T433026|T433026]] - eevans@cumin1003 * 17:30 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-ulsfo or A:cp-drmrs and A:cp - 9.2.15 upgrade ([[phab:T434620|T434620]]) * 17:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1157.eqiad.wmnet with OS bookworm * 17:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1180.eqiad.wmnet with reason: host reimage * 17:22 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 17:21 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1193.eqiad.wmnet with OS bookworm * 17:21 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 17:21 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 17:20 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 17:20 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 17:20 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1183.eqiad.wmnet with OS bookworm * 17:18 swfrench@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 17:18 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 17:17 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1180.eqiad.wmnet with reason: host reimage * 17:17 swfrench@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 17:17 swfrench@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 17:16 swfrench@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 17:16 swfrench@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 17:15 swfrench@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 17:15 swfrench@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 17:15 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1192.eqiad.wmnet with OS bookworm * 17:14 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1159.eqiad.wmnet with reason: host reimage * 17:14 swfrench@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 17:13 swfrench@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 17:12 swfrench@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 17:09 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1158.eqiad.wmnet with reason: host reimage * 17:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1157.eqiad.wmnet with reason: host reimage * 17:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1180 * 17:02 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1180 * 17:01 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica-eqsin and A:liberica * 17:01 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1193.eqiad.wmnet with reason: host reimage * 16:57 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1183.eqiad.wmnet with reason: host reimage * 16:56 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1054.eqiad.wmnet with OS trixie * 16:55 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1056.eqiad.wmnet with OS trixie * 16:54 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1193.eqiad.wmnet with reason: host reimage * 16:54 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1192.eqiad.wmnet with reason: host reimage * 16:51 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-eqiad: Upgrade to Java 17 — [[phab:T433026|T433026]] - eevans@cumin1003 * 16:50 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1180 * 16:50 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1180.eqiad.wmnet 17.36.64.10.in-addr.arpa 7.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:50 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1180.eqiad.wmnet 17.36.64.10.in-addr.arpa 7.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:50 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:50 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1180 - btullis@cumin1003" * 16:50 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1180 - btullis@cumin1003" * 16:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1158.eqiad.wmnet with reason: host reimage * 16:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1157.eqiad.wmnet with reason: host reimage * 16:49 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica-eqsin and A:liberica * 16:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1183.eqiad.wmnet with reason: host reimage * 16:48 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1159.eqiad.wmnet with reason: host reimage * 16:47 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1192.eqiad.wmnet with reason: host reimage * 16:46 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns4004.wikimedia.org * 16:46 sukhe@dns1004: END - running authdns-update * 16:44 sukhe@dns1004: START - running authdns-update * 16:44 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns4004.wikimedia.org,service=authdns-update * 16:43 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns4004.wikimedia.org with OS trixie * 16:39 btullis@cumin1003: START - Cookbook sre.dns.netbox * 16:39 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1180 * 16:39 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1193.eqiad.wmnet with OS bookworm * 16:39 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1180.eqiad.wmnet with OS bookworm * 16:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1192.eqiad.wmnet with OS bookworm * 16:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1183.eqiad.wmnet with OS bookworm * 16:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1159.eqiad.wmnet with OS bookworm * 16:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1158.eqiad.wmnet with OS bookworm * 16:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1157.eqiad.wmnet with OS bookworm * 16:31 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1057.eqiad.wmnet with OS trixie * 16:30 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1057.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 16:29 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica-drmrs and A:liberica * 16:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1154.eqiad.wmnet with OS bookworm * 16:23 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1057.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 16:23 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1055.eqiad.wmnet with OS trixie * 16:22 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1057 * 16:22 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1057 * 16:21 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:21 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1057] - vriley@cumin1003" * 16:21 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1057] - vriley@cumin1003" * 16:19 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica-drmrs and A:liberica * 16:17 vriley@cumin1003: START - Cookbook sre.dns.netbox * 16:16 phuedx@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics-external: apply * 16:15 phuedx@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics-external: apply * 16:13 phuedx@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics-external: apply * 16:12 phuedx@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics-external: apply * 16:11 phuedx@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics-external: apply * 16:09 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics-external: apply * 16:09 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1176.eqiad.wmnet with OS bookworm * 16:07 btullis@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 16:06 btullis@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 16:05 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1191.eqiad.wmnet with OS bookworm * 16:03 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1154.eqiad.wmnet with reason: host reimage * 16:03 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching aqs[2002-2012].codfw.wmnet,aqs[1017-1027].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433026|T433026]] - eevans@cumin1003 * 15:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1190.eqiad.wmnet with OS bookworm * 15:59 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1154.eqiad.wmnet with reason: host reimage * 15:55 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-debug: apply * 15:55 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-debug: apply * 15:55 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-debug: apply * 15:55 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/mw-debug: apply * 15:53 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns4004.wikimedia.org with reason: host reimage * 15:50 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns4004.wikimedia.org with reason: host reimage * 15:46 btullis@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 15:46 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1176.eqiad.wmnet with reason: host reimage * 15:45 btullis@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 15:43 btullis@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 15:42 btullis@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 15:42 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1191.eqiad.wmnet with reason: host reimage * 15:39 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1190.eqiad.wmnet with reason: host reimage * 15:36 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1054.eqiad.wmnet with OS trixie * 15:35 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:35 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1056.eqiad.wmnet with OS trixie * 15:35 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:34 moritzm: failover Ganeti master in eqiad to ganeti1046 * 15:34 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1176.eqiad.wmnet with reason: host reimage * 15:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1191.eqiad.wmnet with reason: host reimage * 15:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1190.eqiad.wmnet with reason: host reimage * 15:31 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns2006.wikimedia.org * 15:31 sukhe@dns1004: END - running authdns-update * 15:29 sukhe@dns1004: START - running authdns-update * 15:29 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns2006.wikimedia.org,service=authdns-update * 15:29 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns2006.wikimedia.org * 15:29 sukhe@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns2006.wikimedia.org * 15:26 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica-ulsfo and A:liberica * 15:26 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns1006.wikimedia.org * 15:26 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:25 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns2006.wikimedia.org with OS trixie * 15:25 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1056 * 15:25 sukhe@dns1004: END - running authdns-update * 15:23 sukhe@dns1004: START - running authdns-update * 15:23 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns1006.wikimedia.org,service=authdns-update * 15:23 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns1006.wikimedia.org * 15:23 sukhe@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns1006.wikimedia.org * 15:20 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1056 * 15:20 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:20 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1056~] - vriley@cumin1003" * 15:19 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1056~] - vriley@cumin1003" * 15:19 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns1006.wikimedia.org with OS trixie * 15:19 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns4004.wikimedia.org with OS trixie * 15:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1190.eqiad.wmnet with OS bookworm * 15:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1191.eqiad.wmnet with OS bookworm * 15:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1176.eqiad.wmnet with OS bookworm * 15:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1154.eqiad.wmnet with OS bookworm * 15:16 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica-ulsfo and A:liberica * 15:13 vriley@cumin1003: START - Cookbook sre.dns.netbox * 15:13 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host dns4004.wikimedia.org with OS trixie * 15:12 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1188.eqiad.wmnet with OS bookworm * 15:09 taavi@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318203{{!}}Undeploy WP25EasterEggs (II) (T418134)]] (duration: 08m 56s) * 15:06 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica-magru and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 15:06 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1051.eqiad.wmnet with OS trixie * 15:06 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 15:05 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 15:05 taavi@deploy1003: taavi: Continuing with deployment * 15:04 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-ncredir (exit_code=0) rolling reboot on A:ncredir and A:ncredir * 15:04 taavi@deploy1003: taavi: Backport for [[gerrit:1318203{{!}}Undeploy WP25EasterEggs (II) (T418134)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:03 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1055.eqiad.wmnet with OS trixie * 15:02 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:00 taavi@deploy1003: Started scap sync-world: Backport for [[gerrit:1318203{{!}}Undeploy WP25EasterEggs (II) (T418134)]] * 14:59 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1047.eqiad.wmnet * 14:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1047.eqiad.wmnet * 14:58 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy (exit_code=0) rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 14:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-codfw * 14:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp2001.codfw.wmnet * 14:57 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp2001.codfw.wmnet * 14:57 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica-magru and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 14:57 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:56 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1055 * 14:56 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1055 * 14:55 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns2006.wikimedia.org with reason: host reimage * 14:55 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:55 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1055] - vriley@cumin1003" * 14:55 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1055] - vriley@cumin1003" * 14:54 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1047.eqiad.wmnet * 14:52 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1188.eqiad.wmnet with reason: host reimage * 14:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp2001.codfw.wmnet * 14:51 vriley@cumin1003: START - Cookbook sre.dns.netbox * 14:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp2001.codfw.wmnet * 14:50 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2016.codfw.wmnet * 14:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2016.codfw.wmnet * 14:50 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1189.eqiad.wmnet with OS bookworm * 14:48 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1188.eqiad.wmnet with reason: host reimage * 14:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1051.eqiad.wmnet with reason: host reimage * 14:46 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1182.eqiad.wmnet with OS bookworm * 14:45 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief2002.codfw.wmnet * 14:44 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns1006.wikimedia.org with reason: host reimage * 14:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2016.codfw.wmnet * 14:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1143.eqiad.wmnet with OS bookworm * 14:43 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2016.codfw.wmnet * 14:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2318-2331].codfw.wmnet * 14:43 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2318-2331].codfw.wmnet * 14:42 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching aqs[2002-2012].codfw.wmnet,aqs[1017-1027].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433026|T433026]] - eevans@cumin1003 * 14:41 cgoubert@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326306{{!}}Add placeholder $wmgRedisLockPassword (T366938 T427999)]] (duration: 06m 56s) * 14:41 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief2002.codfw.wmnet * 14:40 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief1002.eqiad.wmnet * 14:39 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1051.eqiad.wmnet with reason: host reimage * 14:38 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns2006.wikimedia.org with reason: host reimage * 14:37 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns1006.wikimedia.org with reason: host reimage * 14:37 cgoubert@deploy1003: cgoubert: Continuing with deployment * 14:36 cgoubert@deploy1003: cgoubert: Backport for [[gerrit:1326306{{!}}Add placeholder $wmgRedisLockPassword (T366938 T427999)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:36 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief1002.eqiad.wmnet * 14:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2318-2331].codfw.wmnet * 14:34 cgoubert@deploy1003: Started scap sync-world: Backport for [[gerrit:1326306{{!}}Add placeholder $wmgRedisLockPassword (T366938 T427999)]] * 14:29 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2318-2331].codfw.wmnet * 14:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2304-2317].codfw.wmnet * 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2304-2317].codfw.wmnet * 14:28 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1054 * 14:27 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1054 * 14:27 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:27 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1054] - vriley@cumin1003" * 14:27 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1054] - vriley@cumin1003" * 14:26 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1189.eqiad.wmnet with reason: host reimage * 14:26 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test2001.codfw.wmnet * 14:25 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test1001.eqiad.wmnet * 14:25 claime: Deploying wmgRedisLockPassword - [[phab:T366938|T366938]] [[phab:T427999|T427999]] * 14:24 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1051.eqiad.wmnet with OS trixie * 14:23 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 14:22 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1182.eqiad.wmnet with reason: host reimage * 14:22 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1051.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:22 vriley@cumin1003: START - Cookbook sre.dns.netbox * 14:22 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test2001.codfw.wmnet * 14:21 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test1001.eqiad.wmnet * 14:21 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 14:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2304-2317].codfw.wmnet * 14:19 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns4004.wikimedia.org with OS trixie * 14:19 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns2006.wikimedia.org with OS trixie * 14:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1143.eqiad.wmnet with reason: host reimage * 14:19 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns1006.wikimedia.org with OS trixie * 14:17 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1189.eqiad.wmnet with reason: host reimage * 14:15 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1182.eqiad.wmnet with reason: host reimage * 14:14 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1143.eqiad.wmnet with reason: host reimage * 14:13 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1051.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:12 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2304-2317].codfw.wmnet * 14:12 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2290-2303].codfw.wmnet * 14:12 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1049.eqiad.wmnet with OS trixie * 14:12 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2290-2303].codfw.wmnet * 14:12 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1051 * 14:11 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1051 * 14:11 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1170.eqiad.wmnet onto db1284.eqiad.wmnet * 14:11 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1170: Pool db1170.eqiad.wmnet in after cloning * 14:09 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 14:09 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:09 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1051] - vriley@cumin1003" * 14:09 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1051] - vriley@cumin1003" * 14:06 klausman@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:04 vriley@cumin1003: START - Cookbook sre.dns.netbox * 14:04 klausman@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2290-2303].codfw.wmnet * 14:02 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1189.eqiad.wmnet with OS bookworm * 14:02 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1188.eqiad.wmnet with OS bookworm * 14:01 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 14:00 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1182.eqiad.wmnet with OS bookworm * 14:00 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1143.eqiad.wmnet with OS bookworm * 13:59 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 13:57 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply * 13:57 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply * 13:56 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply * 13:56 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-ulsfo or A:cp-drmrs and A:cp - 9.2.15 upgrade ([[phab:T434620|T434620]]) * 13:56 cjd91: sudo -i cookbook sre.cdn.roll-upgrade-ats --query 'A:cp-ulsfo or A:cp-drmrs' --task-id [[phab:T434620|T434620]] --reason '9.2.15 upgrade' * 13:56 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply * 13:55 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply * 13:55 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply * 13:54 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 13:54 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 13:54 phuedx@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:53 phuedx@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics-external: apply * 13:52 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1049.eqiad.wmnet with reason: host reimage * 13:51 phuedx@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2290-2303].codfw.wmnet * 13:50 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1047.eqiad.wmnet * 13:50 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2276-2289].codfw.wmnet * 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2276-2289].codfw.wmnet * 13:49 phuedx@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics-external: apply * 13:49 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1046.eqiad.wmnet * 13:49 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1049.eqiad.wmnet with reason: host reimage * 13:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1046.eqiad.wmnet * 13:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-ncredir rolling reboot on A:ncredir and A:ncredir * 13:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 13:46 phuedx@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:44 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics-external: apply * 13:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1046.eqiad.wmnet * 13:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2276-2289].codfw.wmnet * 13:36 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1046.eqiad.wmnet * 13:36 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1045.eqiad.wmnet * 13:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1045.eqiad.wmnet * 13:34 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2276-2289].codfw.wmnet * 13:34 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1049.eqiad.wmnet with OS trixie * 13:34 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2262-2275].codfw.wmnet * 13:34 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2262-2275].codfw.wmnet * 13:32 Lucas_WMDE: UTC afternoon backport+config window done * 13:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1045.eqiad.wmnet * 13:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2262-2275].codfw.wmnet * 13:26 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1049.eqiad.wmnet with OS trixie * 13:26 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1049.eqiad.wmnet with OS trixie * 13:26 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1170: Pool db1170.eqiad.wmnet in after cloning * 13:23 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1049.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 13:23 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1045.eqiad.wmnet * 13:21 atsukoito: manually done sudo -i docker-registryctl --debug delete-tags 'docker-registry.discovery.wmnet/repos/data-engineering/airflow-dags:airflow-3.3.0-py3.11-2026-08-17-*' to remove incorrect tags * 13:19 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311141{{!}}viwiki: Set `noindex,nofollow` for User and User talk (T432311)]] (duration: 11m 11s) * 13:15 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2262-2275].codfw.wmnet * 13:14 lucaswerkmeister-wmde@deploy1003: ndkdd, lucaswerkmeister-wmde: Continuing with deployment * 13:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2248-2261].codfw.wmnet * 13:14 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2248-2261].codfw.wmnet * 13:13 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1049.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 13:12 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1049 * 13:11 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1049 * 13:10 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:10 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1049] - vriley@cumin1003" * 13:10 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1049] - vriley@cumin1003" * 13:10 lucaswerkmeister-wmde@deploy1003: ndkdd, lucaswerkmeister-wmde: Backport for [[gerrit:1311141{{!}}viwiki: Set `noindex,nofollow` for User and User talk (T432311)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1037.eqiad.wmnet * 13:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1037.eqiad.wmnet * 13:08 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1311141{{!}}viwiki: Set `noindex,nofollow` for User and User talk (T432311)]] * 13:06 vriley@cumin1003: START - Cookbook sre.dns.netbox * 13:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2248-2261].codfw.wmnet * 13:00 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1037.eqiad.wmnet * 12:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2248-2261].codfw.wmnet * 12:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2204-2215,2242-2243].codfw.wmnet * 12:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2204-2215,2242-2243].codfw.wmnet * 12:49 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324832{{!}}Migrate $wgFlaggedRevsTags from flaggedrevs.php to ext-FlaggedRevs.php]] (duration: 14m 02s) * 12:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint2001.codfw.wmnet * 12:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2204-2215,2242-2243].codfw.wmnet * 12:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint2001.codfw.wmnet * 12:41 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1037.eqiad.wmnet * 12:40 ladsgroup@deploy1003: Rolling back deployment * 12:37 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1324832{{!}}Migrate $wgFlaggedRevsTags from flaggedrevs.php to ext-FlaggedRevs.php]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:37 seanleong-wmde: Finished populateSitesTable for [bolwiki] ([[[phab:T429955|T429955]]]) * 12:35 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1324832{{!}}Migrate $wgFlaggedRevsTags from flaggedrevs.php to ext-FlaggedRevs.php]] * 12:35 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2204-2215,2242-2243].codfw.wmnet * 12:34 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2190-2203].codfw.wmnet * 12:34 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2190-2203].codfw.wmnet * 12:32 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1028.eqiad.wmnet * 12:32 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1028.eqiad.wmnet * 12:31 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint1001.eqiad.wmnet * 12:30 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1170: Depool db1170.eqiad.wmnet to then clone it to db1284.eqiad.wmnet - marostegui@cumin1003 * 12:28 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint1001.eqiad.wmnet * 12:28 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1170: Depool db1170.eqiad.wmnet to then clone it to db1284.eqiad.wmnet - marostegui@cumin1003 * 12:27 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1170.eqiad.wmnet onto db1284.eqiad.wmnet * 12:27 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db2901.codfw.wmnet * 12:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2190-2203].codfw.wmnet * 12:27 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 12:26 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1028.eqiad.wmnet * 12:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2190-2203].codfw.wmnet * 12:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2172-2179,2184-2189].codfw.wmnet * 12:18 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2172-2179,2184-2189].codfw.wmnet * 12:13 seanleong-wmde@deploy1003: mwscript-k8s job started: foreachwikiindblist wikidataclient extensions/Wikibase/lib/maintenance/populateSitesTable.php --force-protocol https # [[phab:T429955|T429955]] * 12:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2172-2179,2184-2189].codfw.wmnet * 12:09 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1028.eqiad.wmnet * 12:03 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 12:02 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db2901.codfw.wmnet * 12:02 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db2901.codfw.wmnet * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1027.eqiad.wmnet * 12:02 fceratto@cumin1003: END (ERROR) - Cookbook sre.dns.netbox (exit_code=97) * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1027.eqiad.wmnet * 12:02 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 12:02 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db2901.codfw.wmnet * 12:01 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2172-2179,2184-2189].codfw.wmnet * 12:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2158-2171].codfw.wmnet * 12:01 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2158-2171].codfw.wmnet * 11:58 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db1902.eqiad.wmnet * 11:58 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 11:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1027.eqiad.wmnet * 11:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2158-2171].codfw.wmnet * 11:51 jayme: updated calico to v3.30.7 on wikikube eqiad - [[phab:T427400|T427400]] * 11:50 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1027.eqiad.wmnet * 11:45 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2158-2171].codfw.wmnet * 11:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2144-2157].codfw.wmnet * 11:44 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2144-2157].codfw.wmnet * 11:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1058.eqiad.wmnet * 11:43 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1058.eqiad.wmnet * 11:43 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 11:43 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 11:43 marostegui@cumin1003: Removing db1153 from zarcillo [[phab:T434638|T434638]] * 11:42 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1153.eqiad.wmnet * 11:42 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:42 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1153.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 11:42 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1172.eqiad.wmnet onto db1286.eqiad.wmnet * 11:42 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1172: Pool db1172.eqiad.wmnet in after cloning * 11:42 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1153.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 11:41 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1902.eqiad.wmnet * 11:41 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=97) for new host db2901.codfw.wmnet * 11:41 fceratto@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host db2901.codfw.wmnet with OS trixie * 11:38 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'. * 11:38 marostegui@cumin1003: START - Cookbook sre.dns.netbox * 11:37 marostegui@dns1004: END - running authdns-update * 11:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1058.eqiad.wmnet * 11:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2144-2157].codfw.wmnet * 11:35 marostegui@dns1004: START - running authdns-update * 11:32 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1153.eqiad.wmnet * 11:32 marostegui@cumin1003: START - Cookbook sre.mysql.decommission * 11:28 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326249{{!}}ImagePage: move TOC element below file link (T332644)]] (duration: 09m 56s) * 11:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2144-2157].codfw.wmnet * 11:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2130-2143].codfw.wmnet * 11:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2130-2143].codfw.wmnet * 11:26 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1058.eqiad.wmnet * 11:23 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 11:22 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'. * 11:22 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'. * 11:22 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'. * 11:22 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326249{{!}}ImagePage: move TOC element below file link (T332644)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1057.eqiad.wmnet * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1057.eqiad.wmnet * 11:21 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 11:20 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 11:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2130-2143].codfw.wmnet * 11:19 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326249{{!}}ImagePage: move TOC element below file link (T332644)]] * 11:18 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply * 11:17 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply * 11:17 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply * 11:16 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 11:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1057.eqiad.wmnet * 11:15 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 11:14 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 11:13 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 11:13 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db2901.codfw.wmnet with OS trixie * 11:12 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db2901.codfw.wmnet - fceratto@cumin1003" * 11:12 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db2901.codfw.wmnet - fceratto@cumin1003" * 11:12 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db2901.codfw.wmnet on all recursors * 11:12 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db2901.codfw.wmnet on all recursors * 11:12 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:12 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db2901.codfw.wmnet - fceratto@cumin1003" * 11:12 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2130-2143].codfw.wmnet * 11:11 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2107-2115,2124-2129].codfw.wmnet * 11:11 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2107-2115,2124-2129].codfw.wmnet * 11:11 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 11:11 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 11:11 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply * 11:11 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 11:10 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 11:06 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db2901.codfw.wmnet - fceratto@cumin1003" * 11:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2107-2115,2124-2129].codfw.wmnet * 10:57 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1172: Pool db1172.eqiad.wmnet in after cloning * 10:54 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2107-2115,2124-2129].codfw.wmnet * 10:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2078,2087-2095,2102-2106].codfw.wmnet * 10:54 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2078,2087-2095,2102-2106].codfw.wmnet * 10:47 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply * 10:46 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply * 10:46 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply * 10:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2078,2087-2095,2102-2106].codfw.wmnet * 10:45 blake@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply * 10:38 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1057.eqiad.wmnet * 10:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1056.eqiad.wmnet * 10:38 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1056.eqiad.wmnet * 10:37 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2078,2087-2095,2102-2106].codfw.wmnet * 10:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2061-2062,2064-2065,2067-2077].codfw.wmnet * 10:36 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2061-2062,2064-2065,2067-2077].codfw.wmnet * 10:32 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1056.eqiad.wmnet * 10:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2061-2062,2064-2065,2067-2077].codfw.wmnet * 10:25 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:25 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db2901.codfw.wmnet * 10:25 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db1902.eqiad.wmnet * 10:25 fceratto@cumin1003: END (ERROR) - Cookbook sre.dns.netbox (exit_code=97) * 10:24 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:24 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1902.eqiad.wmnet * 10:24 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=97) for new host db1902.eqiad.wmnet * 10:24 fceratto@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host db1902.eqiad.wmnet with OS trixie * 10:24 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=93) for new host db1903.eqiad.wmnet * 10:24 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 10:20 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1056.eqiad.wmnet * 10:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2061-2062,2064-2065,2067-2077].codfw.wmnet * 10:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2038-2039,2041-2042,2044,2046,2049-2051,2055-2060].codfw.wmnet * 10:18 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2038-2039,2041-2042,2044,2046,2049-2051,2055-2060].codfw.wmnet * 10:15 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1055.eqiad.wmnet * 10:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1055.eqiad.wmnet * 10:13 Amir1: mwscript-k8s --dblist=all -- purgeUserOptions.php --login-age 5 uls-preferences ([[phab:T406724|T406724]]) * 10:11 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:10 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1903.eqiad.wmnet on all recursors * 10:10 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1903.eqiad.wmnet on all recursors * 10:10 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2038-2039,2041-2042,2044,2046,2049-2051,2055-2060].codfw.wmnet * 10:10 moritzm: installing unzip security updates * 10:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1055.eqiad.wmnet * 10:08 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=97) for new host db1901.eqiad.wmnet * 10:08 fceratto@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host db1901.eqiad.wmnet with OS trixie * 10:08 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:08 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 10:08 fceratto@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1903.eqiad.wmnet - fceratto@cumin1003" * 10:07 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db1902.eqiad.wmnet with OS trixie * 10:07 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1902.eqiad.wmnet - fceratto@cumin1003" * 10:07 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1902.eqiad.wmnet - fceratto@cumin1003" * 10:04 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply * 10:04 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326227{{!}}Enable desktop/native lazy loading everywhere (T148047)]] (duration: 07m 13s) * 10:03 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1902.eqiad.wmnet on all recursors * 10:03 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1902.eqiad.wmnet on all recursors * 10:03 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:03 blake@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply * 10:01 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1903.eqiad.wmnet - fceratto@cumin1003" * 10:01 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2038-2039,2041-2042,2044,2046,2049-2051,2055-2060].codfw.wmnet * 10:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2002,2005-2006,2011-2015,2017-2018,2033-2037].codfw.wmnet * 10:01 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2002,2005-2006,2011-2015,2017-2018,2033-2037].codfw.wmnet * 10:00 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:00 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 09:59 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 09:59 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1055.eqiad.wmnet * 09:58 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326227{{!}}Enable desktop/native lazy loading everywhere (T148047)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:56 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326227{{!}}Enable desktop/native lazy loading everywhere (T148047)]] * 09:53 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2002,2005-2006,2011-2015,2017-2018,2033-2037].codfw.wmnet * 09:50 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:48 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1903.eqiad.wmnet * 09:48 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:48 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1902.eqiad.wmnet * 09:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 09:44 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2002,2005-2006,2011-2015,2017-2018,2033-2037].codfw.wmnet * 09:43 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-codfw * 09:43 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1174.eqiad.wmnet onto db1288.eqiad.wmnet * 09:42 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1174: Pool db1174.eqiad.wmnet in after cloning * 09:40 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1172: Depool db1172.eqiad.wmnet to then clone it to db1286.eqiad.wmnet - marostegui@cumin1003 * 09:39 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1172: Depool db1172.eqiad.wmnet to then clone it to db1286.eqiad.wmnet - marostegui@cumin1003 * 09:39 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1172.eqiad.wmnet onto db1286.eqiad.wmnet * 09:33 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1175.eqiad.wmnet onto db1289.eqiad.wmnet * 09:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1175: Pool db1175.eqiad.wmnet in after cloning * 09:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1201.eqiad.wmnet onto db1287.eqiad.wmnet * 09:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1201: Pool db1201.eqiad.wmnet in after cloning * 09:28 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db1901.eqiad.wmnet with OS trixie * 09:27 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:27 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:27 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1901.eqiad.wmnet on all recursors * 09:27 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1901.eqiad.wmnet on all recursors * 09:26 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:26 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:26 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:15 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:15 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1901.eqiad.wmnet * 09:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw1001.wikimedia.org with OS trixie * 08:57 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1174: Pool db1174.eqiad.wmnet in after cloning * 08:54 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1054.eqiad.wmnet * 08:54 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1054.eqiad.wmnet * 08:48 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1054.eqiad.wmnet * 08:47 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1175: Pool db1175.eqiad.wmnet in after cloning * 08:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 08:46 marostegui@cumin1003: Removing db1152 from zarcillo [[phab:T434480|T434480]] * 08:46 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1152.eqiad.wmnet * 08:46 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:46 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1152.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 08:46 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1201: Pool db1201.eqiad.wmnet in after cloning * 08:46 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1152.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 08:46 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1054.eqiad.wmnet * 08:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1053.eqiad.wmnet * 08:43 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1053.eqiad.wmnet * 08:42 marostegui@cumin1003: START - Cookbook sre.dns.netbox * 08:38 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage * 08:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1053.eqiad.wmnet * 08:36 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1152.eqiad.wmnet * 08:36 marostegui@cumin1003: START - Cookbook sre.mysql.decommission * 08:35 phuedx: UTC morning backport window done * 08:35 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1053.eqiad.wmnet * 08:34 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1035.eqiad.wmnet * 08:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1035.eqiad.wmnet * 08:34 phuedx@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324725{{!}}EventStreamConfig: Mark product_metrics.web_base and .web_base_with_ip as Test Kitchen streams (T429898 T430322)]], [[gerrit:1313923{{!}}EventStreamConfig: Remove unused web_ui_scroll* streams (T415370)]], [[gerrit:1325546{{!}}EventStreamConfig: Remove Watchlist click stream (T434790)]] (duration: 12m 42s) * 08:33 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1281: Pool back * 08:32 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage * 08:31 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on db2209.codfw.wmnet with reason: Maintenance * 08:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2209: Maintenance needed * 08:30 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2209: Maintenance needed * 08:26 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1035.eqiad.wmnet * 08:26 phuedx@deploy1003: bearloga, phuedx: Continuing with deployment * 08:23 phuedx@deploy1003: bearloga, phuedx: Backport for [[gerrit:1324725{{!}}EventStreamConfig: Mark product_metrics.web_base and .web_base_with_ip as Test Kitchen streams (T429898 T430322)]], [[gerrit:1313923{{!}}EventStreamConfig: Remove unused web_ui_scroll* streams (T415370)]], [[gerrit:1325546{{!}}EventStreamConfig: Remove Watchlist click stream (T434790)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug * 08:21 phuedx@deploy1003: Started scap sync-world: Backport for [[gerrit:1324725{{!}}EventStreamConfig: Mark product_metrics.web_base and .web_base_with_ip as Test Kitchen streams (T429898 T430322)]], [[gerrit:1313923{{!}}EventStreamConfig: Remove unused web_ui_scroll* streams (T415370)]], [[gerrit:1325546{{!}}EventStreamConfig: Remove Watchlist click stream (T434790)]] * 08:19 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw1001.wikimedia.org with OS trixie * 08:18 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1035.eqiad.wmnet * 08:17 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1032.eqiad.wmnet * 08:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1032.eqiad.wmnet * 08:16 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1279: Pool back * 08:15 phuedx@deploy1003: Finished scap sync-world: Backport for [[gerrit:1216721{{!}}viwikivoyage: enable relatedarticle and pop-up (T405724)]] (duration: 39m 12s) * 08:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1032.eqiad.wmnet * 08:09 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1032.eqiad.wmnet * 08:08 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1031.eqiad.wmnet * 08:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1031.eqiad.wmnet * 08:03 godog: switch production to use dumps-nfs.w.o - [[phab:T432212|T432212]] * 08:02 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1031.eqiad.wmnet * 08:02 phuedx@deploy1003: nvdtn19, phuedx: Continuing with deployment * 08:00 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1201: Depool db1201.eqiad.wmnet to then clone it to db1287.eqiad.wmnet - marostegui@cumin1003 * 08:00 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1201: Depool db1201.eqiad.wmnet to then clone it to db1287.eqiad.wmnet - marostegui@cumin1003 * 08:00 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1201.eqiad.wmnet onto db1287.eqiad.wmnet * 08:00 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1057.eqiad.wmnet * 08:00 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:00 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1057.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:59 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1057.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:59 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1275: Pool back * 07:55 filippo@cumin1003: START - Cookbook sre.dns.netbox * 07:55 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1031.eqiad.wmnet * 07:52 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1030.eqiad.wmnet * 07:52 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1030.eqiad.wmnet * 07:52 phuedx@deploy1003: nvdtn19, phuedx: Backport for [[gerrit:1216721{{!}}viwikivoyage: enable relatedarticle and pop-up (T405724)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:51 tappof: bump space for prometheus k8s-aux in eqiad * 07:50 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1057.eqiad.wmnet * 07:48 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1281: Pool back * 07:47 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1281 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96108 and previous config saved to /var/cache/conftool/dbconfig/20260817-074749-marostegui.json * 07:46 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1030.eqiad.wmnet * 07:42 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1030.eqiad.wmnet * 07:41 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1174: Depool db1174.eqiad.wmnet to then clone it to db1288.eqiad.wmnet - marostegui@cumin1003 * 07:41 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1174: Depool db1174.eqiad.wmnet to then clone it to db1288.eqiad.wmnet - marostegui@cumin1003 * 07:41 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1174.eqiad.wmnet onto db1288.eqiad.wmnet * 07:40 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1029.eqiad.wmnet * 07:40 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm2001.wikimedia.org * 07:40 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1029.eqiad.wmnet * 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1056.eqiad.wmnet * 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1056.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:38 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1056.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:36 phuedx@deploy1003: Started scap sync-world: Backport for [[gerrit:1216721{{!}}viwikivoyage: enable relatedarticle and pop-up (T405724)]] * 07:36 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm2001.wikimedia.org * 07:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1279 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96104 and previous config saved to /var/cache/conftool/dbconfig/20260817-073542-marostegui.json * 07:34 filippo@cumin1003: START - Cookbook sre.dns.netbox * 07:34 slyngshede@dns1004: END - running authdns-update * 07:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1029.eqiad.wmnet * 07:33 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm-test1001.wikimedia.org * 07:32 slyngshede@dns1004: START - running authdns-update * 07:31 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1029.eqiad.wmnet * 07:31 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1279: Pool back * 07:30 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1279 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96102 and previous config saved to /var/cache/conftool/dbconfig/20260817-073038-marostegui.json * 07:29 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm-test1001.wikimedia.org * 07:29 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm1001.wikimedia.org * 07:28 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1056.eqiad.wmnet * 07:28 moritzm: extend the disk of ldap-rw1001 by 80G [[phab:T331699|T331699]] * 07:28 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1055.eqiad.wmnet * 07:28 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:28 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1055.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:27 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1055.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:26 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1044.eqiad.wmnet * 07:26 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1044.eqiad.wmnet * 07:25 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm1001.wikimedia.org * 07:24 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1175: Depool db1175.eqiad.wmnet to then clone it to db1289.eqiad.wmnet - marostegui@cumin1003 * 07:24 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1175: Depool db1175.eqiad.wmnet to then clone it to db1289.eqiad.wmnet - marostegui@cumin1003 * 07:24 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1175.eqiad.wmnet onto db1289.eqiad.wmnet * 07:22 filippo@cumin1003: START - Cookbook sre.dns.netbox * 07:20 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1044.eqiad.wmnet * 07:16 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1055.eqiad.wmnet * 07:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1054.eqiad.wmnet * 07:15 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:15 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1054.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:15 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1044.eqiad.wmnet * 07:15 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1054.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:13 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1275: Pool back * 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1275 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96098 and previous config saved to /var/cache/conftool/dbconfig/20260817-071225-marostegui.json * 07:10 filippo@cumin1003: START - Cookbook sre.dns.netbox * 07:05 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1043.eqiad.wmnet * 07:05 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1054.eqiad.wmnet * 07:05 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1051.eqiad.wmnet * 07:05 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:05 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1051.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:05 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1043.eqiad.wmnet * 07:04 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1051.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 06:59 filippo@cumin1003: START - Cookbook sre.dns.netbox * 06:59 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin1001.eqiad.wmnet * 06:59 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1043.eqiad.wmnet * 06:59 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin2001.codfw.wmnet * 06:55 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin2001.codfw.wmnet * 06:55 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1051.eqiad.wmnet * 06:54 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1049.eqiad.wmnet * 06:54 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:54 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1049.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 06:54 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1049.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 06:54 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin1001.eqiad.wmnet * 06:53 moritzm: installing apr-util security updates * 06:52 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1043.eqiad.wmnet * 06:49 filippo@cumin1003: START - Cookbook sre.dns.netbox * 06:41 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1049.eqiad.wmnet * 06:13 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit2003.wikimedia.org * 06:13 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet * 06:07 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet * 06:06 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit2003.wikimedia.org * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 47s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-16 == * 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 01m 03s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-15 == * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 41s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-14 == * 15:38 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-staging-master-eqiad * 15:38 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster1005.eqiad.wmnet * 15:38 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster1005.eqiad.wmnet * 15:35 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sretest2009.codfw.wmnet * 15:33 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster1005.eqiad.wmnet * 15:33 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster1005.eqiad.wmnet * 15:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster1004.eqiad.wmnet * 15:32 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster1004.eqiad.wmnet * 15:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host sretest2009.codfw.wmnet * 15:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster1004.eqiad.wmnet * 15:27 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster1004.eqiad.wmnet * 15:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster1003.eqiad.wmnet * 15:27 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster1003.eqiad.wmnet * 15:24 dancy@deploy1003: Finished scap sync-world: testing (duration: 03m 23s) * 15:22 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster1003.eqiad.wmnet * 15:22 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster1003.eqiad.wmnet * 15:22 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-staging-master-eqiad * 15:20 dancy@deploy1003: Started scap sync-world: testing * 15:20 dancy@deploy1003: Installation of scap version "4.280.2" completed for 3 hosts * 15:18 dancy@deploy1003: Installing scap version "4.280.2" for 3 host(s) * 15:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sretest2006.codfw.wmnet * 14:54 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host sretest2006.codfw.wmnet * 14:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sretest2003.codfw.wmnet * 14:39 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host sretest2003.codfw.wmnet * 13:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt-staging2001.codfw.wmnet * 13:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt-staging2001.codfw.wmnet * 13:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-staging-master-codfw * 13:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster2005.codfw.wmnet * 13:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster2005.codfw.wmnet * 13:05 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox-dev2003.codfw.wmnet * 13:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster2005.codfw.wmnet * 13:04 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster2005.codfw.wmnet * 13:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster2004.codfw.wmnet * 13:04 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster2004.codfw.wmnet * 13:01 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netbox-dev2003.codfw.wmnet * 12:59 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster2004.codfw.wmnet * 12:59 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster2004.codfw.wmnet * 12:59 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster2003.codfw.wmnet * 12:59 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster2003.codfw.wmnet * 12:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster2003.codfw.wmnet * 12:54 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster2003.codfw.wmnet * 12:54 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-staging-master-codfw * 12:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-staging-worker-eqiad * 12:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage1006.eqiad.wmnet * 12:52 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage1006.eqiad.wmnet * 12:46 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage1006.eqiad.wmnet * 12:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw1001.wikimedia.org with OS trixie * 12:45 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage1006.eqiad.wmnet * 12:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage1005.eqiad.wmnet * 12:45 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage1005.eqiad.wmnet * 12:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage1005.eqiad.wmnet * 12:36 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/ratelimit: apply * 12:35 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/ratelimit: apply * 12:35 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:35 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:33 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage1005.eqiad.wmnet * 12:33 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage1004.eqiad.wmnet * 12:33 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage1004.eqiad.wmnet * 12:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage1004.eqiad.wmnet * 12:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage1004.eqiad.wmnet * 12:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage1003.eqiad.wmnet * 12:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage1003.eqiad.wmnet * 12:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage * 12:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage1003.eqiad.wmnet * 12:17 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage * 12:14 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage1003.eqiad.wmnet * 12:14 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-staging-worker-eqiad * 12:03 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw1001.wikimedia.org with OS trixie * 12:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cuminunpriv1001.eqiad.wmnet * 11:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cuminunpriv1001.eqiad.wmnet * 11:27 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1004.wikimedia.org * 11:24 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1153 from dbctl [[phab:T434638|T434638]]', diff saved to https://phabricator.wikimedia.org/P96097 and previous config saved to /var/cache/conftool/dbconfig/20260814-112449-marostegui.json * 11:21 aokoth@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1004.wikimedia.org * 11:20 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 11:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-staging-worker-codfw * 11:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2004.codfw.wmnet * 11:17 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2004.codfw.wmnet * 11:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2004.codfw.wmnet * 11:10 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2004.codfw.wmnet * 11:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2003.codfw.wmnet * 11:10 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2003.codfw.wmnet * 11:03 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-worker1181.eqiad.wmnet with OS bookworm * 11:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2003.codfw.wmnet * 11:03 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2003.codfw.wmnet * 11:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2002.codfw.wmnet * 11:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2002.codfw.wmnet * 10:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1235.eqiad.wmnet with OS bookworm * 10:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock2003.codfw.wmnet * 10:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock2003.codfw.wmnet with OS trixie * 10:56 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2002.codfw.wmnet * 10:54 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1153.eqiad.wmnet with OS bookworm * 10:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2002.codfw.wmnet * 10:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2001.codfw.wmnet * 10:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2001.codfw.wmnet * 10:50 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 10:49 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1187.eqiad.wmnet with OS bookworm * 10:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2001.codfw.wmnet * 10:44 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2001.codfw.wmnet * 10:44 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-staging-worker-codfw * 10:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock2003.codfw.wmnet with reason: host reimage * 10:38 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock2003.codfw.wmnet with reason: host reimage * 10:35 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1153.eqiad.wmnet with reason: host reimage * 10:32 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1235.eqiad.wmnet with reason: host reimage * 10:29 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1187.eqiad.wmnet with reason: host reimage * 10:24 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1153.eqiad.wmnet with reason: host reimage * 10:23 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1235.eqiad.wmnet with reason: host reimage * 10:21 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1187.eqiad.wmnet with reason: host reimage * 10:16 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock2003.codfw.wmnet with OS trixie * 10:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install1005.wikimedia.org * 10:12 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 10:12 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 10:12 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 10:12 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 10:12 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:12 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 10:12 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 10:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install1005.wikimedia.org * 10:07 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1235.eqiad.wmnet with OS bookworm * 10:07 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1187.eqiad.wmnet with OS bookworm * 10:07 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1181.eqiad.wmnet with OS bookworm * 10:07 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1153.eqiad.wmnet with OS bookworm * 10:07 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install2005.wikimedia.org * 10:03 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 10:03 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2003.codfw.wmnet * 10:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1232.eqiad.wmnet with OS bookworm * 10:00 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install2005.wikimedia.org * 10:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install3004.wikimedia.org * 09:58 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 09:58 fceratto@cumin1003: Removing db1151 from zarcillo [[phab:T434538|T434538]] * 09:56 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1231.eqiad.wmnet with OS bookworm * 09:56 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 09:53 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 09:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install3004.wikimedia.org * 09:50 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1152.eqiad.wmnet with OS bookworm * 09:49 Dreamy_Jazz: `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260808000000" --end-timestamp="20260812120000" --sleep="5" --batch-size="50"` for [[phab:T434688|T434688]] * 09:48 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host centrallog2002.codfw.wmnet * 09:48 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install4004.wikimedia.org * 09:42 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1232.eqiad.wmnet with reason: host reimage * 09:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install4004.wikimedia.org * 09:41 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host centrallog2002.codfw.wmnet * 09:39 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install5004.wikimedia.org * 09:36 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1232.eqiad.wmnet with reason: host reimage * 09:36 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'. * 09:34 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'. * 09:33 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1231.eqiad.wmnet with reason: host reimage * 09:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install5004.wikimedia.org * 09:32 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host centrallog1002.eqiad.wmnet * 09:30 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install6003.wikimedia.org * 09:30 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1231.eqiad.wmnet with reason: host reimage * 09:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1152.eqiad.wmnet with reason: host reimage * 09:25 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host centrallog1002.eqiad.wmnet * 09:25 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1152.eqiad.wmnet with reason: host reimage * 09:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install6003.wikimedia.org * 09:22 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1232 * 09:22 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1232 * 09:22 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1232 * 09:22 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1232.eqiad.wmnet 25.53.64.10.in-addr.arpa 5.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:22 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host titan1001.eqiad.wmnet * 09:22 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1232.eqiad.wmnet 25.53.64.10.in-addr.arpa 5.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:22 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:22 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1232 - btullis@cumin1003" * 09:22 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1232 - btullis@cumin1003" * 09:17 btullis@cumin1003: START - Cookbook sre.dns.netbox * 09:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install7002.wikimedia.org * 09:17 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1232 * 09:16 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1231 * 09:16 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1231 * 09:14 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1231 * 09:14 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1231.eqiad.wmnet 24.53.64.10.in-addr.arpa 4.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host titan1001.eqiad.wmnet * 09:14 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1231.eqiad.wmnet 24.53.64.10.in-addr.arpa 4.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:14 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1231 - btullis@cumin1003" * 09:14 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1231 - btullis@cumin1003" * 09:10 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install7002.wikimedia.org * 09:09 btullis@cumin1003: START - Cookbook sre.dns.netbox * 09:08 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1231 * 09:08 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1152 * 09:08 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1152 * 09:06 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1152 * 09:06 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1152.eqiad.wmnet 16.53.64.10.in-addr.arpa 6.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:06 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1152.eqiad.wmnet 16.53.64.10.in-addr.arpa 6.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:06 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:06 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1152 - btullis@cumin1003" * 09:06 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1152 - btullis@cumin1003" * 09:03 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1151.eqiad.wmnet * 09:03 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:03 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1151.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 08:55 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host titan2001.codfw.wmnet * 08:55 btullis@cumin1003: START - Cookbook sre.dns.netbox * 08:54 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1151.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 08:54 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1152 * 08:53 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1232.eqiad.wmnet with OS bookworm * 08:53 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1231.eqiad.wmnet with OS bookworm * 08:53 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1152.eqiad.wmnet with OS bookworm * 08:51 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2209: Pool back * 08:51 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1201.eqiad.wmnet * 08:50 btullis@cumin1003: START - Cookbook sre.hosts.remove-downtime for an-worker1201.eqiad.wmnet * 08:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1229.eqiad.wmnet with OS bookworm * 08:47 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host titan2001.codfw.wmnet * 08:46 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping2004.codfw.wmnet * 08:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ping2004.codfw.wmnet * 08:40 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host an-worker1230.eqiad.wmnet with OS bookworm * 08:40 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 08:34 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1151.eqiad.wmnet * 08:34 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 08:31 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host titan1002.eqiad.wmnet * 08:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1229.eqiad.wmnet with reason: host reimage * 08:25 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host titan1002.eqiad.wmnet * 08:25 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1229.eqiad.wmnet with reason: host reimage * 08:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping1004.eqiad.wmnet * 08:21 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ping1004.eqiad.wmnet * 08:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1230.eqiad.wmnet with reason: host reimage * 08:12 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1230.eqiad.wmnet with reason: host reimage * 08:11 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1229 * 08:11 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1229 * 08:11 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1229 * 08:11 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1229.eqiad.wmnet 22.53.64.10.in-addr.arpa 2.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:11 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1229.eqiad.wmnet 22.53.64.10.in-addr.arpa 2.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:11 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:11 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1229 - btullis@cumin1003" * 08:11 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1229 - btullis@cumin1003" * 08:10 btullis@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1201.eqiad.wmnet with reason: Fixing a disk * 08:07 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host titan2002.codfw.wmnet * 08:05 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2209: Pool back * 08:05 btullis@cumin1003: START - Cookbook sre.dns.netbox * 08:00 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host titan2002.codfw.wmnet * 07:59 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1229 * 07:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1230 * 07:58 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1230 * 07:55 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1230 * 07:55 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1230.eqiad.wmnet 23.53.64.10.in-addr.arpa 3.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:55 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1230.eqiad.wmnet 23.53.64.10.in-addr.arpa 3.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:55 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:55 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1230 - btullis@cumin1003" * 07:55 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1230 - btullis@cumin1003" * 07:51 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host kubestagemaster2005.codfw.wmnet with OS trixie * 07:48 btullis@cumin1003: START - Cookbook sre.dns.netbox * 07:41 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1230 * 07:41 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1229.eqiad.wmnet with OS bookworm * 07:41 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1230.eqiad.wmnet with OS bookworm * 07:39 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1152 from dbctl [[phab:T434480|T434480]]', diff saved to https://phabricator.wikimedia.org/P96090 and previous config saved to /var/cache/conftool/dbconfig/20260814-073941-marostegui.json * 07:29 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on kubestagemaster2005.codfw.wmnet with reason: host reimage * 07:23 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on kubestagemaster2005.codfw.wmnet with reason: host reimage * 07:04 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host kubestagemaster2005.codfw.wmnet with OS trixie * 06:53 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1228.eqiad.wmnet with OS bookworm * 06:44 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1227.eqiad.wmnet with OS bookworm * 06:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1209.eqiad.wmnet with OS bookworm * 06:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1175.eqiad.wmnet with OS bookworm * 06:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1228.eqiad.wmnet with reason: host reimage * 06:27 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1228.eqiad.wmnet with reason: host reimage * 06:25 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1227.eqiad.wmnet with reason: host reimage * 06:21 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1227.eqiad.wmnet with reason: host reimage * 06:18 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1209.eqiad.wmnet with reason: host reimage * 06:14 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1209.eqiad.wmnet with reason: host reimage * 06:14 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1228 * 06:14 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1228 * 06:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1175.eqiad.wmnet with reason: host reimage * 06:12 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1228 * 06:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1228.eqiad.wmnet 20.53.64.10.in-addr.arpa 0.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:12 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1228.eqiad.wmnet 20.53.64.10.in-addr.arpa 0.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1228 - ryankemper@cumin2003" * 06:12 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1228 - ryankemper@cumin2003" * 06:09 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1175.eqiad.wmnet with reason: host reimage * 06:07 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 06:07 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1228 * 06:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1227 * 06:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1227 * 06:06 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1227 * 06:06 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1227.eqiad.wmnet 19.53.64.10.in-addr.arpa 9.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:06 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1227.eqiad.wmnet 19.53.64.10.in-addr.arpa 9.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:06 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:06 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1227 - ryankemper@cumin2003" * 06:06 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1227 - ryankemper@cumin2003" * 06:00 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 06:00 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1227 * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1209 * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1209 * 06:00 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1209 * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1209.eqiad.wmnet 15.53.64.10.in-addr.arpa 5.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:00 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1209.eqiad.wmnet 15.53.64.10.in-addr.arpa 5.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1209 - ryankemper@cumin2003" * 06:00 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1209 - ryankemper@cumin2003" * 05:54 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 05:54 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1209 * 05:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1175 * 05:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1175 * 05:52 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1175 * 05:52 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1175.eqiad.wmnet 17.53.64.10.in-addr.arpa 7.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 05:52 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1175.eqiad.wmnet 17.53.64.10.in-addr.arpa 7.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 05:52 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 05:52 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1175 - ryankemper@cumin2003" * 05:52 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1175 - ryankemper@cumin2003" * 05:49 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1228.eqiad.wmnet with OS bookworm * 05:49 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1227.eqiad.wmnet with OS bookworm * 05:48 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1209.eqiad.wmnet with OS bookworm * 05:47 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 05:47 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1175 * 05:47 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1175.eqiad.wmnet with OS bookworm * 05:09 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host apifeatureusage1001.eqiad.wmnet with OS bookworm * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 03s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:10 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 01:07 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 01:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 01:02 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 00:59 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1226.eqiad.wmnet with OS bookworm * 00:47 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1225.eqiad.wmnet with OS bookworm * 00:41 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1224.eqiad.wmnet with OS bookworm * 00:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1226.eqiad.wmnet with reason: host reimage * 00:31 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1226.eqiad.wmnet with reason: host reimage * 00:28 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1225.eqiad.wmnet with reason: host reimage * 00:25 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1225.eqiad.wmnet with reason: host reimage * 00:19 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1224.eqiad.wmnet with reason: host reimage * 00:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1226 * 00:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1226 * 00:17 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1226 * 00:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1226.eqiad.wmnet 23.36.64.10.in-addr.arpa 3.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:17 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1226.eqiad.wmnet 23.36.64.10.in-addr.arpa 3.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 00:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1226 - ryankemper@cumin2003" * 00:17 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1226 - ryankemper@cumin2003" * 00:16 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1224.eqiad.wmnet with reason: host reimage * 00:12 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 00:12 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1226 * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1225 * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1225 * 00:10 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1225 * 00:10 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1225.eqiad.wmnet 22.36.64.10.in-addr.arpa 2.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:10 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1225.eqiad.wmnet 22.36.64.10.in-addr.arpa 2.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:10 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 00:10 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1225 - ryankemper@cumin2003" * 00:10 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1225 - ryankemper@cumin2003" * 00:03 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 00:02 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1225 * 00:02 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1224 * 00:02 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1224 * 00:00 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1224 * 00:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1224.eqiad.wmnet 21.36.64.10.in-addr.arpa 1.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:00 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1224.eqiad.wmnet 21.36.64.10.in-addr.arpa 1.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 00:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1224 - ryankemper@cumin2003" * 00:00 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1224 - ryankemper@cumin2003" == 2026-08-13 == * 23:54 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1226.eqiad.wmnet with OS bookworm * 23:53 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1225.eqiad.wmnet with OS bookworm * 23:52 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 23:51 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1224 * 23:51 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1224.eqiad.wmnet with OS bookworm * 23:47 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1223.eqiad.wmnet with OS bookworm * 23:29 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1223.eqiad.wmnet with reason: host reimage * 23:24 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1223.eqiad.wmnet with reason: host reimage * 23:19 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325525{{!}}ve.ui.CodeMirror.less: ensure normal font style]] (duration: 11m 40s) * 23:16 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260807000000" --end-timestamp="20260808000000" --sleep="5" --batch-size="50"` for [[phab:T434688|T434688]] * 23:13 musikanimal@deploy1003: musikanimal: Continuing with deployment * 23:11 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1325525{{!}}ve.ui.CodeMirror.less: ensure normal font style]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1223 * 23:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1223 * 23:08 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1325525{{!}}ve.ui.CodeMirror.less: ensure normal font style]] * 23:07 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1223 * 23:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1223.eqiad.wmnet 20.36.64.10.in-addr.arpa 0.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:07 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1223.eqiad.wmnet 20.36.64.10.in-addr.arpa 0.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 23:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1223 - ryankemper@cumin2003" * 23:03 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1223 - ryankemper@cumin2003" * 23:01 sbassett: Deployed security updates for [[phab:T430596|T430596]], [[phab:T120386|T120386]] * 22:55 bking@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host kubestagemaster2005.codfw.wmnet * 22:55 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host kubestagemaster2005.codfw.wmnet with OS bookworm * 22:54 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 22:53 ryankemper@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 22:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on kubestagemaster2005.codfw.wmnet with reason: host reimage * 22:49 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1222.eqiad.wmnet with OS bookworm * 22:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on kubestagemaster2005.codfw.wmnet with reason: host reimage * 22:29 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1222.eqiad.wmnet with reason: host reimage * 22:26 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1222.eqiad.wmnet with reason: host reimage * 22:23 bking@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host aux-k8s-etcd2003.codfw.wmnet * 22:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd2003.codfw.wmnet with OS bookworm * 22:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host kubestagemaster2005.codfw.wmnet with OS bookworm * 22:23 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM kubestagemaster2005.codfw.wmnet - bking@cumin2003" * 22:23 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM kubestagemaster2005.codfw.wmnet - bking@cumin2003" * 22:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) kubestagemaster2005.codfw.wmnet on all recursors * 22:22 bking@cumin2003: START - Cookbook sre.dns.wipe-cache kubestagemaster2005.codfw.wmnet on all recursors * 22:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:22 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM kubestagemaster2005.codfw.wmnet - bking@cumin2003" * 22:22 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM kubestagemaster2005.codfw.wmnet - bking@cumin2003" * 22:17 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 22:13 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1223 * 22:12 bking@cumin2003: START - Cookbook sre.dns.netbox * 22:12 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host kubestagemaster2005.codfw.wmnet * 22:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1222 * 22:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1222 * 22:11 sbassett: Deployed security updates for [[phab:T429244|T429244]], [[phab:T434039|T434039]] * 22:11 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1222 * 22:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1222.eqiad.wmnet 19.36.64.10.in-addr.arpa 9.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:11 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1222.eqiad.wmnet 19.36.64.10.in-addr.arpa 9.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1222 - ryankemper@cumin2003" * 22:07 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1222 - ryankemper@cumin2003" * 22:05 robh@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-wdqs2001.codfw.wmnet with reason: updating firmware * 22:01 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 22:01 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1223.eqiad.wmnet with OS bookworm * 22:01 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1222 * 22:01 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1222.eqiad.wmnet with OS bookworm * 22:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1212.eqiad.wmnet with OS bookworm * 21:54 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324295{{!}}Revert "Lazily reject pre-fix parser-cache entries for noreferrer/noopener links" (T429090)]] (duration: 06m 42s) * 21:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd2003.codfw.wmnet with reason: host reimage * 21:53 bking@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host dse-k8s-etcd2001.codfw.wmnet * 21:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dse-k8s-etcd2001.codfw.wmnet with OS bookworm * 21:50 sbassett@deploy1003: sbassett, kharlan: Continuing with deployment * 21:49 sbassett@deploy1003: sbassett, kharlan: Backport for [[gerrit:1324295{{!}}Revert "Lazily reject pre-fix parser-cache entries for noreferrer/noopener links" (T429090)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:48 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aux-k8s-etcd2003.codfw.wmnet with reason: host reimage * 21:47 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1324295{{!}}Revert "Lazily reject pre-fix parser-cache entries for noreferrer/noopener links" (T429090)]] * 21:42 maryum: Deployed security patch for [[phab:T434549|T434549]] * 21:39 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1212.eqiad.wmnet with reason: host reimage * 21:34 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1212.eqiad.wmnet with reason: host reimage * 21:33 bking@cumin2003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd2003.codfw.wmnet with OS bookworm * 21:32 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM aux-k8s-etcd2003.codfw.wmnet - bking@cumin2003" * 21:32 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM aux-k8s-etcd2003.codfw.wmnet - bking@cumin2003" * 21:32 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) aux-k8s-etcd2003.codfw.wmnet on all recursors * 21:32 bking@cumin2003: START - Cookbook sre.dns.wipe-cache aux-k8s-etcd2003.codfw.wmnet on all recursors * 21:32 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:32 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM aux-k8s-etcd2003.codfw.wmnet - bking@cumin2003" * 21:31 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM aux-k8s-etcd2003.codfw.wmnet - bking@cumin2003" * 21:28 maryum: Deployed security patch for [[phab:T434619|T434619]] * 21:26 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:26 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host aux-k8s-etcd2003.codfw.wmnet * 21:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-etcd2001.codfw.wmnet with reason: host reimage * 21:20 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1212 * 21:20 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1212 * 21:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1212.eqiad.wmnet with OS bookworm * 21:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on dse-k8s-etcd2001.codfw.wmnet with reason: host reimage * 21:04 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325553{{!}}InstrumentConstructiveEdits: exclude mw-reverted as well (T431493)]] (duration: 06m 25s) * 20:59 kemayo@deploy1003: kemayo: Continuing with deployment * 20:59 kemayo@deploy1003: kemayo: Backport for [[gerrit:1325553{{!}}InstrumentConstructiveEdits: exclude mw-reverted as well (T431493)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host dse-k8s-etcd2001.codfw.wmnet with OS bookworm * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 20:57 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 20:57 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1325553{{!}}InstrumentConstructiveEdits: exclude mw-reverted as well (T431493)]] * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-etcd2001.codfw.wmnet on all recursors * 20:57 bking@cumin2003: START - Cookbook sre.dns.wipe-cache dse-k8s-etcd2001.codfw.wmnet on all recursors * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 20:57 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 20:54 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320163{{!}}Add configurable RestTermsOfServiceUrl (T428147)]] (duration: 21m 39s) * 20:53 bking@cumin2003: START - Cookbook sre.dns.netbox * 20:53 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host dse-k8s-etcd2001.codfw.wmnet * 20:50 samtar@deploy1003: samtar, milazg: Continuing with deployment * 20:35 samtar@deploy1003: samtar, milazg: Backport for [[gerrit:1320163{{!}}Add configurable RestTermsOfServiceUrl (T428147)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:33 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1320163{{!}}Add configurable RestTermsOfServiceUrl (T428147)]] * 20:30 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325549{{!}}Deploy PRV to several LC wikis (T423785)]] (duration: 06m 57s) * 20:26 arlolra@deploy1003: arlolra: Continuing with deployment * 20:25 arlolra@deploy1003: arlolra: Backport for [[gerrit:1325549{{!}}Deploy PRV to several LC wikis (T423785)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:24 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 20:23 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1325549{{!}}Deploy PRV to several LC wikis (T423785)]] * 20:21 ariel@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319906{{!}}Remove boilerplate language from wmf-rest and wmf-math API modules (T433736)]] (duration: 13m 54s) * 20:14 ariel@deploy1003: ariel: Continuing with deployment * 20:11 ariel@deploy1003: ariel: Backport for [[gerrit:1319906{{!}}Remove boilerplate language from wmf-rest and wmf-math API modules (T433736)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 ariel@deploy1003: Started scap sync-world: Backport for [[gerrit:1319906{{!}}Remove boilerplate language from wmf-rest and wmf-math API modules (T433736)]] * 19:58 robh@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-wdqs2001.codfw.wmnet with reason: updating firmware * 19:54 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325551{{!}}Render the focused module view as a full-screen page (T433896)]] (duration: 30m 37s) * 19:52 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching aqs[2001,1016]*: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 19:44 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching aqs[2001,1016]*: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 19:42 musikanimal@deploy1003: musikanimal: Continuing with deployment * 19:41 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1325551{{!}}Render the focused module view as a full-screen page (T433896)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:34 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: quash java safepoint logspam - bking@cumin2003 - [[phab:T434685|T434685]] * 19:34 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 19:34 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 19:24 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1325551{{!}}Render the focused module view as a full-screen page (T433896)]] * 19:20 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 19:20 jhancock@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin1003" * 19:18 jhancock@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin1003" * 19:14 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325548{{!}}Enable image lazy loading on desktop in group1 (T148047)]] (duration: 07m 43s) * 19:10 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 19:10 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1325548{{!}}Enable image lazy loading on desktop in group1 (T148047)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:07 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1325548{{!}}Enable image lazy loading on desktop in group1 (T148047)]] * 19:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1166.eqiad.wmnet onto db1280.eqiad.wmnet * 19:03 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 19:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1166: Pool db1166.eqiad.wmnet in after cloning * 18:59 jhancock@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 18:58 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1234.eqiad.wmnet with OS bookworm * 18:52 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 18:51 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:49 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:46 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 18:46 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 18:44 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=dns3004.* * 18:39 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1221.eqiad.wmnet with OS bookworm * 18:36 inflatador: [bking@ganeti2048] ~$ sudo gnt-instance replace-disks -n ganeti2030.codfw.wmnet aux-k8s-worker2002.codfw.wmnet [[phab:T434681|T434681]] * 18:35 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 18:29 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1220.eqiad.wmnet with OS bookworm * 18:29 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1234.eqiad.wmnet with reason: host reimage * 18:27 bking@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host dse-k8s-etcd2001.codfw.wmnet * 18:27 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-etcd2001.codfw.wmnet on all recursors * 18:27 bking@cumin2003: START - Cookbook sre.dns.wipe-cache dse-k8s-etcd2001.codfw.wmnet on all recursors * 18:27 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:27 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 18:27 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 18:25 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1234.eqiad.wmnet with reason: host reimage * 18:25 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 18:19 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1221.eqiad.wmnet with reason: host reimage * 18:17 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1166: Pool db1166.eqiad.wmnet in after cloning * 18:15 brennen@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 18:15 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1221.eqiad.wmnet with reason: host reimage * 18:14 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:12 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:12 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 18:12 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-etcd2001.codfw.wmnet on all recursors * 18:12 bking@cumin2003: START - Cookbook sre.dns.wipe-cache dse-k8s-etcd2001.codfw.wmnet on all recursors * 18:12 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:12 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 18:12 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 18:10 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: quash java safepoint logspam - bking@cumin2003 - [[phab:T434685|T434685]] * 18:10 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1234 * 18:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1234 * 18:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1220.eqiad.wmnet with reason: host reimage * 18:07 brennen: 1.47.0-wmf.15 train status ([[phab:T430834|T430834]]) - no current blockers, rolling to all wikis * 18:07 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1234 * 18:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1234.eqiad.wmnet 10.36.64.10.in-addr.arpa 0.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:07 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:07 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1234.eqiad.wmnet 10.36.64.10.in-addr.arpa 0.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1234 - ryankemper@cumin2003" * 18:07 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1234 - ryankemper@cumin2003" * 18:05 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1220.eqiad.wmnet with reason: host reimage * 18:04 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host dse-k8s-etcd2001.codfw.wmnet * 18:04 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 18:03 dancy@deploy1003: Installation of scap version "4.280.1" completed for 3 hosts * 18:02 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:02 inflatador: bking@dse-k8s-etcd2002 etcdctl member remove $<nowiki>{</nowiki>UUID of dse-k8s-etcd2001<nowiki>}</nowiki> [[phab:T434681|T434681]] [[phab:T434793|T434793]] * 18:01 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 18:01 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1234 * 18:01 dancy@deploy1003: Installing scap version "4.280.1" for 3 host(s) * 18:01 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1221 * 18:01 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1221 * 18:00 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1221 * 18:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1221.eqiad.wmnet 18.36.64.10.in-addr.arpa 8.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:00 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1221.eqiad.wmnet 18.36.64.10.in-addr.arpa 8.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1221 - ryankemper@cumin2003" * 17:59 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:58 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 17:58 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:57 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 17:57 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:56 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1221 - ryankemper@cumin2003" * 17:53 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 17:52 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 17:52 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 17:51 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1221 * 17:51 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1220 * 17:51 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1220 * 17:51 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:51 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:51 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1220 * 17:51 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1220.eqiad.wmnet 11.36.64.10.in-addr.arpa 1.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:51 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1220.eqiad.wmnet 11.36.64.10.in-addr.arpa 1.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:51 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1220 - ryankemper@cumin2003" * 17:50 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1220 - ryankemper@cumin2003" * 17:47 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:47 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 17:46 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1234.eqiad.wmnet with OS bookworm * 17:46 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 17:45 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1221.eqiad.wmnet with OS bookworm * 17:45 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1220 * 17:45 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1220.eqiad.wmnet with OS bookworm * 17:43 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 17:41 bking@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host dse-k8s-etcd2001.codfw.wmnet * 17:41 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host dse-k8s-etcd2001.codfw.wmnet * 17:40 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:40 inflatador: bking@ganeti2048] `sudo gnt-instance remove --force --ignore-failures --shutdown-timeout=0` on non-DRBD VMs [[phab:T434681|T434681]] * 17:40 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:39 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:38 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:36 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:32 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:32 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:28 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 17:26 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:26 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:24 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:23 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:21 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1218.eqiad.wmnet with OS bookworm * 17:20 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:20 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:19 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1179.eqiad.wmnet with OS bookworm * 17:18 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1150.eqiad.wmnet with OS bookworm * 17:18 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns3004.wikimedia.org with OS trixie * 17:15 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:14 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:13 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:12 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:12 swfrench@deploy1003: Finished scap sync-world: Helmfile-only deployment for mediawiki chart bump - [[phab:T427666|T427666]] (duration: 03m 03s) * 17:09 swfrench@deploy1003: Started scap sync-world: Helmfile-only deployment for mediawiki chart bump - [[phab:T427666|T427666]] * 17:02 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:01 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:01 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1218.eqiad.wmnet with reason: host reimage * 17:01 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:01 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:00 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324810{{!}}deployment-info.php: Report dbname and branch for the requested wiki (T434726)]] (duration: 06m 52s) * 16:58 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 16:57 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1179.eqiad.wmnet with reason: host reimage * 16:56 dancy@deploy1003: dancy: Continuing with deployment * 16:56 dancy@deploy1003: dancy: Backport for [[gerrit:1324810{{!}}deployment-info.php: Report dbname and branch for the requested wiki (T434726)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1150.eqiad.wmnet with reason: host reimage * 16:53 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324810{{!}}deployment-info.php: Report dbname and branch for the requested wiki (T434726)]] * 16:51 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1166: Depool db1166.eqiad.wmnet to then clone it to db1280.eqiad.wmnet - cwilliams@cumin1003 * 16:50 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1166: Depool db1166.eqiad.wmnet to then clone it to db1280.eqiad.wmnet - cwilliams@cumin1003 * 16:50 cwilliams@cumin1003: START - Cookbook sre.mysql.clone of db1166.eqiad.wmnet onto db1280.eqiad.wmnet * 16:49 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1218.eqiad.wmnet with reason: host reimage * 16:48 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1179.eqiad.wmnet with reason: host reimage * 16:47 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1150.eqiad.wmnet with reason: host reimage * 16:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.upgrade (exit_code=0) for 1 hosts * 16:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2220: Upgrade of db2220.codfw.wmnet completed * 16:38 swfrench@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 16:38 swfrench-wmf: kubectl delete node kubestagemaster2005.codfw.wmnet - [[phab:T434681|T434681]] * 16:34 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1218 * 16:34 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1218 * 16:34 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1218.eqiad.wmnet with OS bookworm * 16:34 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1179 * 16:34 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1179 * 16:33 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1179.eqiad.wmnet with OS bookworm * 16:32 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1150 * 16:32 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1150 * 16:31 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1150.eqiad.wmnet with OS bookworm * 16:24 dancy@deploy1003: Installation of scap version "4.280.0" completed for 3 hosts * 16:22 dancy@deploy1003: Installing scap version "4.280.0" for 3 host(s) * 16:20 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: quash java safepoint logspam - bking@cumin2003 - [[phab:T434685|T434685]] * 16:14 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns3004.wikimedia.org with reason: host reimage * 16:08 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 16:07 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 16:07 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns3004.wikimedia.org with reason: host reimage * 16:05 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 16:04 swfrench@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 16:00 swfrench@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 15:59 swfrench@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 15:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Upgrade of db2220.codfw.wmnet completed * 15:48 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2220: Upgrading db2220.codfw.wmnet * 15:48 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2220: Upgrading db2220.codfw.wmnet * 15:48 cwilliams@cumin1003: START - Cookbook sre.mysql.upgrade for 1 hosts * 15:46 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns3004.wikimedia.org with OS trixie * 15:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2220 [[phab:T434802|T434802]]', diff saved to https://phabricator.wikimedia.org/P96079 and previous config saved to /var/cache/conftool/dbconfig/20260813-154624-cwilliams.json * 15:45 cdobbins@cumin1003: conftool action : set/pooled=no; selector: name=dns3004.* * 15:44 cjd91: depooling dns3004 to reimage to trixie * 15:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2159 to s7 primary [[phab:T434802|T434802]]', diff saved to https://phabricator.wikimedia.org/P96078 and previous config saved to /var/cache/conftool/dbconfig/20260813-154405-cwilliams.json * 15:43 cezmunsta: Starting s7 codfw failover from db2220 to db2159 - [[phab:T434802|T434802]] * 15:41 cgoubert@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: Dragonfly supernodes reboot (duration: 09m 42s) * 15:41 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dragonfly-supernode2001.codfw.wmnet * 15:39 inflatador: bking@ganeti2048] ~$ sudo gnt-node failover -f --ignore-consistency ganeti2046.codfw.wmnet [[phab:T434681|T434681]] * 15:39 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 15:39 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 15:39 swfrench@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 15:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2159 with weight 0 [[phab:T434802|T434802]]', diff saved to https://phabricator.wikimedia.org/P96077 and previous config saved to /var/cache/conftool/dbconfig/20260813-153806-cwilliams.json * 15:37 swfrench@dns1004: END - running authdns-update * 15:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 30 hosts with reason: Primary switchover s7 [[phab:T434802|T434802]] * 15:37 cgoubert@cumin2003: START - Cookbook sre.hosts.reboot-single for host dragonfly-supernode2001.codfw.wmnet * 15:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dragonfly-supernode1001.eqiad.wmnet * 15:35 swfrench@dns1004: START - running authdns-update * 15:32 cgoubert@cumin2003: START - Cookbook sre.hosts.reboot-single for host dragonfly-supernode1001.eqiad.wmnet * 15:32 cgoubert@deploy1003: Locking from deployment [ALL REPOSITORIES]: Dragonfly supernodes reboot * 15:30 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-master-codfw * 15:30 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2005.codfw.wmnet * 15:30 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2005.codfw.wmnet * 15:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1169.eqiad.wmnet onto db1277.eqiad.wmnet * 15:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1169: Pool db1169.eqiad.wmnet in after cloning * 15:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2005.codfw.wmnet * 15:23 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2005.codfw.wmnet * 15:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2004.codfw.wmnet * 15:23 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2004.codfw.wmnet * 15:18 inflatador: bking@ganeti2048 sudo gnt-node failover -f ganeti2046.codfw.wmnet [[phab:T434681|T434681]] * 15:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2004.codfw.wmnet * 15:17 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2004.codfw.wmnet * 15:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2003.codfw.wmnet * 15:17 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2003.codfw.wmnet * 15:15 cdobbins@dns1004: END - running authdns-update * 15:13 cdobbins@dns1004: START - running authdns-update * 15:10 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1217.eqiad.wmnet with OS bookworm * 15:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2003.codfw.wmnet * 15:10 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2003.codfw.wmnet * 15:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2002.codfw.wmnet * 15:10 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2002.codfw.wmnet * 15:10 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: quash java safepoint logspam - bking@cumin2003 - [[phab:T434685|T434685]] * 15:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1216.eqiad.wmnet with OS bookworm * 15:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2002.codfw.wmnet * 15:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2002.codfw.wmnet * 15:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2001.codfw.wmnet * 15:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2001.codfw.wmnet * 15:01 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1042.eqiad.wmnet * 15:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1042.eqiad.wmnet * 15:00 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325455{{!}}Api: Use correct query when continue prop=categories (T433922)]] (duration: 09m 47s) * 14:59 bking@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host dse-k8s-etcd2004.codfw.wmnet * 14:58 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-etcd2004.codfw.wmnet on all recursors * 14:58 bking@cumin2003: START - Cookbook sre.dns.wipe-cache dse-k8s-etcd2004.codfw.wmnet on all recursors * 14:58 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:58 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM dse-k8s-etcd2004.codfw.wmnet - bking@cumin2003" * 14:58 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM dse-k8s-etcd2004.codfw.wmnet - bking@cumin2003" * 14:55 zabe@deploy1003: zabe: Continuing with deployment * 14:53 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-etcd2004.codfw.wmnet on all recursors * 14:53 bking@cumin2003: START - Cookbook sre.dns.wipe-cache dse-k8s-etcd2004.codfw.wmnet on all recursors * 14:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:53 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2004.codfw.wmnet - bking@cumin2003" * 14:53 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2004.codfw.wmnet - bking@cumin2003" * 14:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2001.codfw.wmnet * 14:52 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2001.codfw.wmnet * 14:52 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-master-codfw * 14:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2001.codfw.wmnet * 14:52 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2001.codfw.wmnet * 14:52 zabe@deploy1003: zabe: Backport for [[gerrit:1325455{{!}}Api: Use correct query when continue prop=categories (T433922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:50 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1325455{{!}}Api: Use correct query when continue prop=categories (T433922)]] * 14:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1217.eqiad.wmnet with reason: host reimage * 14:48 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:48 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host dse-k8s-etcd2004.codfw.wmnet * 14:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-master-eqiad * 14:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1006.eqiad.wmnet * 14:44 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1006.eqiad.wmnet * 14:44 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1216.eqiad.wmnet with reason: host reimage * 14:40 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1217.eqiad.wmnet with reason: host reimage * 14:39 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1169: Pool db1169.eqiad.wmnet in after cloning * 14:39 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1216.eqiad.wmnet with reason: host reimage * 14:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl1005.eqiad.wmnet * 14:32 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl1005.eqiad.wmnet * 14:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1004.eqiad.wmnet * 14:32 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1004.eqiad.wmnet * 14:31 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus2008.codfw.wmnet * 14:31 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor1003.eqiad.wmnet * 14:29 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1042.eqiad.wmnet * 14:27 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor1003.eqiad.wmnet * 14:26 moritzm: installing Django security updates * 14:26 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling reboot on A:wikidough * 14:25 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1217.eqiad.wmnet with OS bookworm * 14:25 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1216.eqiad.wmnet with OS bookworm * 14:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl1004.eqiad.wmnet * 14:24 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl1004.eqiad.wmnet * 14:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1003.eqiad.wmnet * 14:24 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1003.eqiad.wmnet * 14:23 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus2008.codfw.wmnet * 14:23 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus1008.eqiad.wmnet * 14:23 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor-dev2001.codfw.wmnet * 14:22 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus2006.codfw.wmnet * 14:19 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor-dev2001.codfw.wmnet * 14:18 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1042.eqiad.wmnet * 14:17 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.* * 14:16 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl1003.eqiad.wmnet * 14:16 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl1003.eqiad.wmnet * 14:16 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1002.eqiad.wmnet * 14:16 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1002.eqiad.wmnet * 14:15 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus1008.eqiad.wmnet * 14:14 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus1006.eqiad.wmnet * 14:12 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus2006.codfw.wmnet * 14:12 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor2003.codfw.wmnet * 14:11 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus2007.codfw.wmnet * 14:11 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1041.eqiad.wmnet * 14:11 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1041.eqiad.wmnet * 14:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl1002.eqiad.wmnet * 14:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl1002.eqiad.wmnet * 14:09 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-master-eqiad * 14:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on apifeatureusage1001.eqiad.wmnet with reason: host reimage * 14:08 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1215.eqiad.wmnet with OS bookworm * 14:08 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor2003.codfw.wmnet * 14:08 jayme: updated calico to v3.30.7 on wikikube codfw [[phab:T427400|T427400]] * 14:07 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1149.eqiad.wmnet with OS bookworm * 14:06 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'. * 14:06 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1041.eqiad.wmnet * 14:04 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus1006.eqiad.wmnet * 14:03 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus2007.codfw.wmnet * 14:03 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus2005.codfw.wmnet * 14:03 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus1007.eqiad.wmnet * 14:03 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1002.eqiad.wmnet * 14:02 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on apifeatureusage1001.eqiad.wmnet with reason: host reimage * 14:02 cgoubert@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=helm-charts.*,name=eqiad * 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host chartmuseum1001.eqiad.wmnet * 14:00 moritzm: installing libxml2 security updates * 13:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1214.eqiad.wmnet with OS bookworm * 13:59 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns5003.* * 13:59 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1041.eqiad.wmnet * 13:59 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow1002.eqiad.wmnet * 13:58 cmooney@dns3003: END - running authdns-update * 13:57 cgoubert@cumin2003: START - Cookbook sre.hosts.reboot-single for host chartmuseum1001.eqiad.wmnet * 13:57 cgoubert@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=helm-charts.*,name=eqiad * 13:57 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: cloudelastic cluster restart - bking@cumin2003 * 13:57 cgoubert@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=helm-charts.*,name=codfw * 13:56 cmooney@dns3003: START - running authdns-update * 13:56 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'. * 13:56 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns5003.*,service=authdns-update * 13:55 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host chartmuseum2001.codfw.wmnet * 13:55 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus1007.eqiad.wmnet * 13:55 cmooney@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dns5003.wikimedia.org * 13:55 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus1005.eqiad.wmnet * 13:51 cgoubert@cumin2003: START - Cookbook sre.hosts.reboot-single for host chartmuseum2001.codfw.wmnet * 13:51 cgoubert@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=helm-charts.*,name=codfw * 13:51 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus2005.codfw.wmnet * 13:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host apifeatureusage1001.eqiad.wmnet with OS bookworm * 13:50 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus7002.magru.wmnet * 13:50 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts lvs1015.eqiad.wmnet * 13:50 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:50 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1015.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:49 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1015.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:49 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325480{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] (duration: 06m 39s) * 13:47 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1215.eqiad.wmnet with reason: host reimage * 13:46 cmooney@cumin1003: START - Cookbook sre.hosts.reboot-single for host dns5003.wikimedia.org * 13:46 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1003.eqiad.wmnet * 13:45 cmooney@cumin1003: conftool action : set/pooled=no; selector: name=dns5003.* * 13:45 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1040.eqiad.wmnet * 13:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1040.eqiad.wmnet * 13:45 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 13:45 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus1005.eqiad.wmnet * 13:45 stran@deploy1003: stran: Continuing with deployment * 13:44 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus7002.magru.wmnet * 13:44 stran@deploy1003: stran: Backport for [[gerrit:1325480{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:44 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus6002.drmrs.wmnet * 13:43 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260805000000" --end-timestamp="20260806000000" --sleep="5" --batch-size="10"` for [[phab:T434688|T434688]] * 13:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1149.eqiad.wmnet with reason: host reimage * 13:42 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1325480{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] * 13:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow1003.eqiad.wmnet * 13:40 sukhe@cumin1003: START - Cookbook sre.hosts.decommission for hosts lvs1015.eqiad.wmnet * 13:40 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts lvs1014.eqiad.wmnet * 13:40 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:40 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1014.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:40 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1040.eqiad.wmnet * 13:40 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1014.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:39 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1214.eqiad.wmnet with reason: host reimage * 13:38 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2004.codfw.wmnet * 13:38 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus6002.drmrs.wmnet * 13:38 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus5003.eqsin.wmnet * 13:35 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 13:35 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1149.eqiad.wmnet with reason: host reimage * 13:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1215.eqiad.wmnet with reason: host reimage * 13:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow2004.codfw.wmnet * 13:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1214.eqiad.wmnet with reason: host reimage * 13:33 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1040.eqiad.wmnet * 13:31 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus5003.eqsin.wmnet * 13:31 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: cloudelastic cluster restart - bking@cumin2003 * 13:31 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus4003.ulsfo.wmnet * 13:30 sukhe@cumin1003: START - Cookbook sre.hosts.decommission for hosts lvs1014.eqiad.wmnet * 13:30 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts lvs1013.eqiad.wmnet * 13:30 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:30 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1013.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:30 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1013.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:27 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1039.eqiad.wmnet * 13:27 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1039.eqiad.wmnet * 13:26 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1167.eqiad.wmnet onto db1281.eqiad.wmnet * 13:26 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1167: Pool db1167.eqiad.wmnet in after cloning * 13:25 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus4003.ulsfo.wmnet * 13:25 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260802000000" --end-timestamp="20260803000000" --sleep="5" --batch-size="10"` for [[phab:T434688|T434688]] * 13:24 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 13:24 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus3004.esams.wmnet * 13:24 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1274: New host * 13:24 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325476{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] (duration: 07m 19s) * 13:24 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260804000000" --end-timestamp="20260805000000" --sleep="5" --batch-size="10"` for [[phab:T434688|T434688]] * 13:24 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260803000000" --end-timestamp="20260804000000" --sleep="5" --batch-size="10"` for [[phab:T434688|T434688]] * 13:23 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2003.codfw.wmnet * 13:22 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1039.eqiad.wmnet * 13:20 sukhe@cumin1003: START - Cookbook sre.hosts.decommission for hosts lvs1013.eqiad.wmnet * 13:20 stran@deploy1003: stran: Continuing with deployment * 13:19 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow2003.codfw.wmnet * 13:19 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling reboot on A:wikidough * 13:19 stran@deploy1003: stran: Backport for [[gerrit:1325476{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:18 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus3004.esams.wmnet * 13:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1215.eqiad.wmnet with OS bookworm * 13:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1214.eqiad.wmnet with OS bookworm * 13:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1149.eqiad.wmnet with OS bookworm * 13:17 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1325476{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] * 13:11 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324731{{!}}prv: Enable parsoid rendering for 5 wikisource wikis]] (duration: 07m 26s) * 13:11 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1039.eqiad.wmnet * 13:11 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1036.eqiad.wmnet * 13:11 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1036.eqiad.wmnet * 13:07 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow3004.esams.wmnet * 13:07 jgiannelos@deploy1003: jgiannelos: Continuing with deployment * 13:06 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1233.eqiad.wmnet with OS bookworm * 13:06 jgiannelos@deploy1003: jgiannelos: Backport for [[gerrit:1324731{{!}}prv: Enable parsoid rendering for 5 wikisource wikis]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:04 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1324731{{!}}prv: Enable parsoid rendering for 5 wikisource wikis]] * 13:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow3004.esams.wmnet * 13:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1036.eqiad.wmnet * 13:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1210.eqiad.wmnet with OS bookworm * 13:01 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1036.eqiad.wmnet * 13:00 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1052.eqiad.wmnet * 13:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1052.eqiad.wmnet * 12:58 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow4003.ulsfo.wmnet * 12:55 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1211.eqiad.wmnet with OS bookworm * 12:55 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1052.eqiad.wmnet * 12:52 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow4003.ulsfo.wmnet * 12:51 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1052.eqiad.wmnet * 12:49 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1169: Depool db1169.eqiad.wmnet to then clone it to db1277.eqiad.wmnet - cwilliams@cumin1003 * 12:46 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1169: Depool db1169.eqiad.wmnet to then clone it to db1277.eqiad.wmnet - cwilliams@cumin1003 * 12:46 cwilliams@cumin1003: START - Cookbook sre.mysql.clone of db1169.eqiad.wmnet onto db1277.eqiad.wmnet * 12:45 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:45 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns record for deleted IP reservations eqsin lvs vlan ints - cmooney@cumin1003" * 12:44 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns record for deleted IP reservations eqsin lvs vlan ints - cmooney@cumin1003" * 12:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1051.eqiad.wmnet * 12:43 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1051.eqiad.wmnet * 12:42 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1210.eqiad.wmnet with reason: host reimage * 12:41 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1167: Pool db1167.eqiad.wmnet in after cloning * 12:40 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 12:39 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1274: New host * 12:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Added db1274', diff saved to https://phabricator.wikimedia.org/P96062 and previous config saved to /var/cache/conftool/dbconfig/20260813-123907-cwilliams.json * 12:38 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast7002.wikimedia.org * 12:38 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1233.eqiad.wmnet with reason: host reimage * 12:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1051.eqiad.wmnet * 12:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1211.eqiad.wmnet with reason: host reimage * 12:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1233.eqiad.wmnet with reason: host reimage * 12:32 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1051.eqiad.wmnet * 12:32 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast7002.wikimedia.org * 12:32 marostegui: Drop SecurePoll tables from closed wikis [[phab:T423128|T423128]] * 12:32 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow5003.eqsin.wmnet * 12:31 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1050.eqiad.wmnet * 12:31 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1050.eqiad.wmnet * 12:29 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1210.eqiad.wmnet with reason: host reimage * 12:29 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1211.eqiad.wmnet with reason: host reimage * 12:29 moritzm: installing Wireshark security updates * 12:26 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow5003.eqsin.wmnet * 12:26 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1050.eqiad.wmnet * 12:24 cmooney@dns3003: END - running authdns-update * 12:21 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow6001.drmrs.wmnet * 12:21 cmooney@dns3003: START - running authdns-update * 12:20 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1050.eqiad.wmnet * 12:20 cgoubert@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host rdb-lock2003.codfw.wmnet * 12:20 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:20 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update netbox dns entries for expanded public1-603-eqsin subnet - cmooney@cumin1003" * 12:20 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update netbox dns entries for expanded public1-603-eqsin subnet - cmooney@cumin1003" * 12:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 12:20 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 12:19 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1049.eqiad.wmnet * 12:19 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1049.eqiad.wmnet * 12:19 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 12:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host moss-be1003.eqiad.wmnet * 12:17 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow6001.drmrs.wmnet * 12:17 moritzm: installin curl security updates * 12:15 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 12:15 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1233.eqiad.wmnet with OS bookworm * 12:15 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1210.eqiad.wmnet with OS bookworm * 12:15 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1211.eqiad.wmnet with OS bookworm * 12:13 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1049.eqiad.wmnet * 12:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow7002.magru.wmnet * 12:12 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-worker1178.eqiad.wmnet with OS bookworm * 12:10 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host moss-be1003.eqiad.wmnet * 12:10 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 12:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be1006.eqiad.wmnet * 12:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 12:10 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 12:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 12:10 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 12:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow7002.magru.wmnet * 12:08 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1049.eqiad.wmnet * 12:05 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 12:05 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2003.codfw.wmnet * 12:04 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'. * 12:04 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:03 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be1006.eqiad.wmnet * 12:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be1005.eqiad.wmnet * 12:01 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1038.eqiad.wmnet * 12:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1038.eqiad.wmnet * 12:01 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1213.eqiad.wmnet with OS bookworm * 11:56 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be1005.eqiad.wmnet * 11:55 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be1004.eqiad.wmnet * 11:54 cgoubert@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host rdb-lock2003.codfw.wmnet * 11:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 11:53 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 11:53 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:53 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 11:53 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 11:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1038.eqiad.wmnet * 11:49 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be1004.eqiad.wmnet * 11:44 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:44 moritzm: installing Linux 5.10.262 on Bullseye hosts * 11:41 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1038.eqiad.wmnet * 11:40 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 11:40 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 11:40 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 11:40 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:40 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 11:40 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 11:38 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1213.eqiad.wmnet with reason: host reimage * 11:36 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1167: Depool db1167.eqiad.wmnet to then clone it to db1281.eqiad.wmnet - marostegui@cumin1003 * 11:35 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 11:35 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2003.codfw.wmnet * 11:35 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1167: Depool db1167.eqiad.wmnet to then clone it to db1281.eqiad.wmnet - marostegui@cumin1003 * 11:35 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1167.eqiad.wmnet onto db1281.eqiad.wmnet * 11:34 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 22 hosts with reason: Cloning * 11:34 moritzm: remove ganeti3005 from esams03 cluster, hardware issues [[phab:T434646|T434646]] * 11:32 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1213.eqiad.wmnet with reason: host reimage * 11:28 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1034.eqiad.wmnet * 11:28 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1034.eqiad.wmnet * 11:22 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1034.eqiad.wmnet * 11:19 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1034.eqiad.wmnet * 11:17 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1213.eqiad.wmnet with OS bookworm * 11:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1165.eqiad.wmnet onto db1279.eqiad.wmnet * 11:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1165: Pool db1165.eqiad.wmnet in after cloning * 11:07 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply * 10:57 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply * 10:54 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply * 10:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1274.eqiad.wmnet with reason: Enabling notifications and pooling * 10:45 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1033.eqiad.wmnet * 10:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1033.eqiad.wmnet * 10:44 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply * 10:43 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'. * 10:42 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'. * 10:42 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'. * 10:40 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db[1216,1225,1239-1240].eqiad.wmnet with reason: reboot * 10:39 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1033.eqiad.wmnet * 10:38 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1151 from dbctl [[phab:T434538|T434538]]', diff saved to https://phabricator.wikimedia.org/P96055 and previous config saved to /var/cache/conftool/dbconfig/20260813-103828-marostegui.json * 10:35 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1033.eqiad.wmnet * 10:27 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1165: Pool db1165.eqiad.wmnet in after cloning * 10:24 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1151.eqiad.wmnet with OS bookworm * 10:15 moritzm: installing bind9 security updates (client-side tools/libs only) * 10:07 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-debug: apply * 10:06 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-debug: apply * 10:02 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 7 hosts with reason: reboot * 10:01 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-debug: apply * 10:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.upgrade (exit_code=0) for 1 hosts * 10:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2214: Upgrade of db2214.codfw.wmnet completed * 10:01 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-debug: apply * 10:00 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-debug: apply * 10:00 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-debug: apply * 09:59 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.decommission (exit_code=99) * 09:59 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 09:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1151.eqiad.wmnet with reason: host reimage * 09:59 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 8 hosts * 09:59 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 8 hosts * 09:55 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1151.eqiad.wmnet with reason: host reimage * 09:46 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260801000000" --end-timestamp="20260802000000" --sleep="3" --batch-size="5"` for [[phab:T434688|T434688]] * 09:45 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 8 hosts with reason: reboot * 09:45 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 22 hosts with reason: Cloning * 09:43 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1165: Depool db1165.eqiad.wmnet to then clone it to db1279.eqiad.wmnet - marostegui@cumin1003 * 09:42 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1165: Depool db1165.eqiad.wmnet to then clone it to db1279.eqiad.wmnet - marostegui@cumin1003 * 09:42 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1165.eqiad.wmnet onto db1279.eqiad.wmnet * 09:41 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=testwiki --start-timestamp="20260311000000" --end-timestamp="20260805000000" --sleep="5" --batch-size="2"` for [[phab:T434688|T434688]] * 09:40 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for backupmon1001.eqiad.wmnet * 09:40 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for backupmon1001.eqiad.wmnet * 09:40 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1151 * 09:40 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1151 * 09:37 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325408{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]], [[gerrit:1325407{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]] (duration: 06m 57s) * 09:36 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1151 * 09:36 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1151.eqiad.wmnet 13.36.64.10.in-addr.arpa 3.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:36 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1151.eqiad.wmnet 13.36.64.10.in-addr.arpa 3.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:36 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:36 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1151 - btullis@cumin1003" * 09:36 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on backupmon1001.eqiad.wmnet with reason: reboot * 09:36 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1151 - btullis@cumin1003" * 09:35 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host moss-be2003.codfw.wmnet * 09:33 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 7 hosts * 09:33 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 7 hosts * 09:33 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 09:32 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1325408{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]], [[gerrit:1325407{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1161.eqiad.wmnet onto db1275.eqiad.wmnet * 09:31 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 09:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1161: Pool db1161.eqiad.wmnet in after cloning * 09:30 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1325408{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]], [[gerrit:1325407{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]] * 09:29 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 09:27 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 09:27 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host moss-be2003.codfw.wmnet * 09:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be2006.codfw.wmnet * 09:25 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2004.codfw.wmnet * 09:22 hashar@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 09:21 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be2006.codfw.wmnet * 09:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be2005.codfw.wmnet * 09:19 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2004.codfw.wmnet * 09:19 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2003.codfw.wmnet * 09:18 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 7 hosts with reason: reboot * 09:18 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 7 hosts * 09:18 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 7 hosts * 09:16 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2214: Upgrade of db2214.codfw.wmnet completed * 09:14 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be2005.codfw.wmnet * 09:13 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be2004.codfw.wmnet * 09:12 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2003.codfw.wmnet * 09:12 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2002.codfw.wmnet * 09:09 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2214: Upgrading db2214.codfw.wmnet * 09:09 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2214: Upgrading db2214.codfw.wmnet * 09:09 cwilliams@cumin1003: START - Cookbook sre.mysql.upgrade for 1 hosts * 09:07 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be2004.codfw.wmnet * 09:06 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2002.codfw.wmnet * 09:05 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1004.eqiad.wmnet * 09:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-cluster (exit_code=0) * 09:03 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 7 hosts with reason: reboot * 09:01 btullis@cumin1003: START - Cookbook sre.dns.netbox * 09:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2214 [[phab:T434754|T434754]]', diff saved to https://phabricator.wikimedia.org/P96045 and previous config saved to /var/cache/conftool/dbconfig/20260813-090001-cwilliams.json * 08:59 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1004.eqiad.wmnet * 08:59 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1003.eqiad.wmnet * 08:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2229 to s6 primary [[phab:T434754|T434754]]', diff saved to https://phabricator.wikimedia.org/P96044 and previous config saved to /var/cache/conftool/dbconfig/20260813-085752-cwilliams.json * 08:57 cezmunsta: Starting s6 codfw failover from db2214 to db2229 - [[phab:T434754|T434754]] * 08:54 hashar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325416{{!}}Revert "REST: Enable `GET /lexemes/<nowiki>{</nowiki>lexeme_id<nowiki>}</nowiki>` by default" (T434712)]] (duration: 07m 22s) * 08:53 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1003.eqiad.wmnet * 08:53 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1002.eqiad.wmnet * 08:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2229 with weight 0 [[phab:T434754|T434754]]', diff saved to https://phabricator.wikimedia.org/P96043 and previous config saved to /var/cache/conftool/dbconfig/20260813-085151-cwilliams.json * 08:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 23 hosts with reason: Primary switchover s6 [[phab:T434754|T434754]] * 08:51 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts2002.codfw.wmnet * 08:50 hashar@deploy1003: hashar: Continuing with deployment * 08:49 hashar@deploy1003: hashar: Backport for [[gerrit:1325416{{!}}Revert "REST: Enable `GET /lexemes/<nowiki>{</nowiki>lexeme_id<nowiki>}</nowiki>` by default" (T434712)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:47 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet * 08:47 hashar@deploy1003: Started scap sync-world: Backport for [[gerrit:1325416{{!}}Revert "REST: Enable `GET /lexemes/<nowiki>{</nowiki>lexeme_id<nowiki>}</nowiki>` by default" (T434712)]] * 08:47 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1002.eqiad.wmnet * 08:46 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1161: Pool db1161.eqiad.wmnet in after cloning * 08:46 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host stewards2001.codfw.wmnet * 08:45 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit1003.wikimedia.org * 08:45 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host stewards1001.eqiad.wmnet * 08:45 Emperor: roll-restart apus frontends in codfw * 08:45 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-cluster * 08:44 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts2002.codfw.wmnet * 08:44 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host doc2003.codfw.wmnet * 08:43 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet * 08:43 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host phab2003.codfw.wmnet * 08:42 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host stewards2001.codfw.wmnet * 08:42 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host etherpad2002.codfw.wmnet * 08:41 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host stewards1001.eqiad.wmnet * 08:41 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1273: Pool in s7 * 08:41 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1003.wikimedia.org * 08:40 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host doc2003.codfw.wmnet * 08:40 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit1003.wikimedia.org * 08:39 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host etherpad1004.eqiad.wmnet * 08:39 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host doc1004.eqiad.wmnet * 08:39 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit2002.wikimedia.org * 08:38 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host etherpad2002.codfw.wmnet * 08:37 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host phab2003.codfw.wmnet * 08:36 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lists1004.wikimedia.org * 08:35 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host etherpad1004.eqiad.wmnet * 08:35 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host doc1004.eqiad.wmnet * 08:34 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-cluster (exit_code=0) * 08:34 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1003.wikimedia.org * 08:34 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host planet2003.codfw.wmnet * 08:34 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2003.wikimedia.org * 08:33 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host planet1003.eqiad.wmnet * 08:33 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit2002.wikimedia.org * 08:32 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 22 hosts with reason: Cloning * 08:32 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2002.wikimedia.org * 08:31 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aphlict1002.eqiad.wmnet * 08:30 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host planet2003.codfw.wmnet * 08:29 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host planet1003.eqiad.wmnet * 08:29 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host lists1004.wikimedia.org * 08:28 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lists2001.wikimedia.org * 08:28 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2003.wikimedia.org * 08:27 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host aphlict1002.eqiad.wmnet * 08:27 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aphlict2001.codfw.wmnet * 08:26 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2002.wikimedia.org * 08:23 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host aphlict2001.codfw.wmnet * 08:23 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1151 * 08:22 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1151.eqiad.wmnet with OS bookworm * 08:22 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host lists2001.wikimedia.org * 08:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1161: Depool db1161.eqiad.wmnet to then clone it to db1275.eqiad.wmnet - marostegui@cumin1003 * 08:20 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1161: Depool db1161.eqiad.wmnet to then clone it to db1275.eqiad.wmnet - marostegui@cumin1003 * 08:20 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1161.eqiad.wmnet onto db1275.eqiad.wmnet * 08:15 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-cluster * 08:15 Emperor: roll-restart apus frontends in eqiad * 07:56 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1273: Pool in s7 * 07:56 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1273 to dbctl [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P96035 and previous config saved to /var/cache/conftool/dbconfig/20260813-075611-marostegui.json * 07:38 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db1273.eqiad.wmnet with reason: Reboot * 07:31 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.sanitize-wiki (exit_code=97) Managing sanitization for wikis testwiki in section s3 * 07:24 marostegui@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis testwiki in section s3 * 07:19 jayme@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on kubestagemaster2005.codfw.wmnet with reason: downtime because of hardware failure and no DRBD * 05:42 arnaudb@dns1006: END - running authdns-update * 05:40 arnaudb@dns1006: START - running authdns-update * 05:27 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 04:06 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 02:29 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1151 * 02:29 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1151 * 02:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1219.eqiad.wmnet with OS bookworm * 02:18 ryankemper: [[phab:T434494|T434494]] `ryankemper@deploy1003:~$ echo 'https://stats.wikimedia.org/' {{!}} mwscript-k8s --attach -- purgeList.php` (default page got cached during yesterday's `an-web1001` reimage) * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 46s) * 02:03 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1219.eqiad.wmnet with reason: host reimage * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 02:00 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1219.eqiad.wmnet with reason: host reimage * 01:46 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1219.eqiad.wmnet with OS bookworm * 01:03 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1142.eqiad.wmnet with OS bookworm * 00:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1142.eqiad.wmnet with reason: host reimage * 00:34 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1142.eqiad.wmnet with reason: host reimage * 00:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1142.eqiad.wmnet with OS bookworm * 00:16 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1178 * 00:16 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1178 == 2026-08-12 == * 23:18 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324409{{!}}Remove $wmg = $wg hacks in Collection (T119117)]] (duration: 06m 43s) * 23:14 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 23:13 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1324409{{!}}Remove $wmg = $wg hacks in Collection (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:11 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1324409{{!}}Remove $wmg = $wg hacks in Collection (T119117)]] * 22:59 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 22:49 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260801000000" --end-timestamp="20260802000000" --sleep=2 --batch-size=10` * 22:45 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=testwiki --start-timestamp="20200801010101" --end-timestamp="20260816010101" --sleep=15 --batch-size=5` * 22:40 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=testwiki --start-timestamp="20200101010101" --end-timestamp="20260816010101" --sleep=60` * 22:35 Dreamy_Jazz: Running `mwscript WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260101000000" --end-timestamp="20260102000000" --sleep=10` * 22:21 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324817{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]], [[gerrit:1324818{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]] (duration: 45m 29s) * 22:17 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 21:59 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 21:40 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1324817{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]], [[gerrit:1324818{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:39 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 21:36 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1324817{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]], [[gerrit:1324818{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]] * 21:32 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:31 vriley@cumin1003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:30 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:30 vriley@cumin1003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:17 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:14 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:14 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:11 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:10 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-codfw: Set storage compatability to NONE — [[phab:T433028|T433028]] - eevans@cumin1003 * 21:10 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1006 * 21:09 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1006 * 21:05 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324277{{!}}Improve Math preference labels for SVG/MathJax/MathML (T433891)]] (duration: 31m 42s) * 20:58 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1178.eqiad.wmnet with OS bookworm * 20:54 krinkle@deploy1003: krinkle: Continuing with deployment * 20:51 krinkle@deploy1003: krinkle: Backport for [[gerrit:1324277{{!}}Improve Math preference labels for SVG/MathJax/MathML (T433891)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:39 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 20:38 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:38 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp5022.eqsin.wmnet with OS trixie * 20:38 cdobbins@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - cdobbins@cumin1003" * 20:37 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:36 cdobbins@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - cdobbins@cumin1003" * 20:34 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1178.eqiad.wmnet with reason: host reimage * 20:34 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1324277{{!}}Improve Math preference labels for SVG/MathJax/MathML (T433891)]] * 20:33 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 20:28 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1178.eqiad.wmnet with reason: host reimage * 20:13 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1178.eqiad.wmnet with OS bookworm * 20:11 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-worker1178.eqiad.wmnet with OS bookworm * 20:11 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1178.eqiad.wmnet with OS bookworm * 20:09 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-codfw: Set storage compatability to NONE — [[phab:T433028|T433028]] - eevans@cumin1003 * 20:09 Dreamy_Jazz: Evening UTC backport window done * 20:08 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324794{{!}}WikimediaAntiAbuse: Enable logging channel (T431292)]] (duration: 06m 48s) * 20:08 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage * 20:05 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage * 20:04 dreamyjazz@deploy1003: kharlan, dreamyjazz: Continuing with deployment * 20:04 dreamyjazz@deploy1003: kharlan, dreamyjazz: Backport for [[gerrit:1324794{{!}}WikimediaAntiAbuse: Enable logging channel (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:01 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1324794{{!}}WikimediaAntiAbuse: Enable logging channel (T431292)]] * 19:55 brennen@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 19:47 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 19:35 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 19:35 cdobbins@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cp5022.eqsin.wmnet with OS trixie * 19:32 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:30 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:29 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:26 vriley@cumin1003: START - Cookbook sre.dns.netbox * 19:23 brennen@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 19:19 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-eqiad: Set storage compatability to NONE — [[phab:T433028|T433028]] - eevans@cumin1003 * 19:10 brennen@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 19:09 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 19:09 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 19:00 Amir1: data migrated on wikishared ([[phab:T426102|T426102]]) * 18:57 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321224{{!}}Rename ce_worklist_articles table to ce_invitation_list_articles (T426102)]] (duration: 06m 50s) * 18:53 ladsgroup@deploy1003: ladsgroup, daimona: Continuing with deployment * 18:53 ladsgroup@deploy1003: ladsgroup, daimona: Backport for [[gerrit:1321224{{!}}Rename ce_worklist_articles table to ce_invitation_list_articles (T426102)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:51 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1321224{{!}}Rename ce_worklist_articles table to ce_invitation_list_articles (T426102)]] * 18:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2192: Security update * 18:31 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 18:25 Amir1: ce_invitation_list_articles created as empty on wikishared ([[phab:T426102|T426102]]) * 18:21 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 18:21 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 18:18 Amir1: migrated testwiki entries from ce_worklist_articles to ce_invitation_list_articles ([[phab:T426102|T426102]]) * 18:18 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-eqiad: Set storage compatability to NONE — [[phab:T433028|T433028]] - eevans@cumin1003 * 18:11 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 18:08 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling reboot on A:durum-eqsin and A:durum * 18:07 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Set storage compatability to UPGRADING — [[phab:T433028|T433028]] - eevans@cumin1003 * 18:05 jhancock@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022'] * 17:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2192: Security update * 17:55 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:55 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum-eqsin and A:durum * 17:53 jhancock@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['cp5022'] * 17:47 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:47 jhancock@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['cp5022'] * 17:42 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:41 jhancock@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['cp5022'] * 17:36 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:35 jhancock@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022'] * 17:31 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-magru and not (P<nowiki>{</nowiki>cp7001*<nowiki>}</nowiki> or P<nowiki>{</nowiki>cp7009*<nowiki>}</nowiki>) and A:cp - 9.2.15 upgrade ([[phab:T434620|T434620]]) * 17:28 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:28 jhancock@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['cp5022'] * 17:22 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2192.codfw.wmnet with reason: Maintenance * 17:11 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host apifeatureusage2001.codfw.wmnet with OS bookworm * 17:04 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: apply * 17:03 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-main: apply * 17:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2192 [[phab:T434635|T434635]]', diff saved to https://phabricator.wikimedia.org/P96030 and previous config saved to /var/cache/conftool/dbconfig/20260812-170338-cwilliams.json * 17:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2213 to s5 primary [[phab:T434635|T434635]]', diff saved to https://phabricator.wikimedia.org/P96029 and previous config saved to /var/cache/conftool/dbconfig/20260812-170152-cwilliams.json * 17:01 cezmunsta: Starting s5 codfw failover from db2192 to db2213 - [[phab:T434635|T434635]] * 16:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2213 with weight 0 [[phab:T434635|T434635]]', diff saved to https://phabricator.wikimedia.org/P96028 and previous config saved to /var/cache/conftool/dbconfig/20260812-165544-cwilliams.json * 16:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 27 hosts with reason: Primary switchover s5 [[phab:T434635|T434635]] * 16:53 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: apply * 16:52 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-main: apply * 16:44 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-main: apply * 16:44 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-main: apply * 16:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1272: New host * 16:40 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324765{{!}}WikimediaAntiAbuse: Enable PersonalInfoFlagNotifications (T431292)]] (duration: 07m 02s) * 16:40 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-logging-external: apply * 16:39 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-logging-external: apply * 16:38 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-logging-external: apply * 16:37 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-logging-external: apply * 16:36 kharlan@deploy1003: kharlan: Continuing with deployment * 16:35 kharlan@deploy1003: kharlan: Backport for [[gerrit:1324765{{!}}WikimediaAntiAbuse: Enable PersonalInfoFlagNotifications (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:33 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1324765{{!}}WikimediaAntiAbuse: Enable PersonalInfoFlagNotifications (T431292)]] * 16:27 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-logging-external: apply * 16:27 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-logging-external: apply * 16:18 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 16:18 jhancock@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022'] * 16:17 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 16:16 jhancock@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['cp5022'] * 16:15 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 16:12 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Set storage compatability to UPGRADING — [[phab:T433028|T433028]] - eevans@cumin1003 * 16:10 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: apply * 16:10 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: apply * 16:08 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: apply * 16:08 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: apply * 16:08 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics: apply * 16:07 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics: apply * 16:02 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-magru and not (P<nowiki>{</nowiki>cp7001*<nowiki>}</nowiki> or P<nowiki>{</nowiki>cp7009*<nowiki>}</nowiki>) and A:cp - 9.2.15 upgrade ([[phab:T434620|T434620]]) * 15:57 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1272: New host * 15:52 jmm@cumin2003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti2046.codfw.wmnet * 15:52 jmm@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host ganeti2046.codfw.wmnet * 15:42 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:37 mforns@deploy1003: Finished deploy [analytics/refinery@49c336c] (thin): Regular analytics weekly train THIN [analytics/refinery@49c336cd] (duration: 01m 59s) * 15:37 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324733{{!}}WikimediaAntiAbuse: Enable personal info tag display on enwiki (T431292)]] (duration: 08m 12s) * 15:35 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:35 mforns@deploy1003: Started deploy [analytics/refinery@49c336c] (thin): Regular analytics weekly train THIN [analytics/refinery@49c336cd] * 15:34 mforns@deploy1003: Finished deploy [analytics/refinery@49c336c]: Regular analytics weekly train [analytics/refinery@49c336cd] (duration: 04m 20s) * 15:33 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[2024,1031]*.wmnet: Set storage compatability to UPGRADING — [[phab:T433028|T433028]] - eevans@cumin1003 * 15:33 dreamyjazz@deploy1003: kharlan, dreamyjazz: Continuing with deployment * 15:31 dreamyjazz@deploy1003: kharlan, dreamyjazz: Backport for [[gerrit:1324733{{!}}WikimediaAntiAbuse: Enable personal info tag display on enwiki (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:30 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitize-wiki (exit_code=99) Checking sanitization for wikis testwiki in section s3 * 15:30 mforns@deploy1003: Started deploy [analytics/refinery@49c336c]: Regular analytics weekly train [analytics/refinery@49c336cd] * 15:30 mforns@deploy1003: Finished deploy [analytics/refinery@49c336c] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@49c336cd] (duration: 00m 32s) * 15:29 mforns@deploy1003: Started deploy [analytics/refinery@49c336c] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@49c336cd] * 15:29 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1324733{{!}}WikimediaAntiAbuse: Enable personal info tag display on enwiki (T431292)]] * 15:27 brennen@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324700{{!}}EventDetailsParticipantsModule: populate cache with non-local users (T434597)]] (duration: 06m 38s) * 15:23 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[2024,1031]*.wmnet: Set storage compatability to UPGRADING — [[phab:T433028|T433028]] - eevans@cumin1003 * 15:23 brennen@deploy1003: brennen, daimona: Continuing with deployment * 15:23 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:22 brennen@deploy1003: brennen, daimona: Backport for [[gerrit:1324700{{!}}EventDetailsParticipantsModule: populate cache with non-local users (T434597)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:22 cgoubert@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host rdb-lock2003.codfw.wmnet * 15:21 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 15:21 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 15:21 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:21 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:21 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:20 brennen@deploy1003: Started scap sync-world: Backport for [[gerrit:1324700{{!}}EventDetailsParticipantsModule: populate cache with non-local users (T434597)]] * 15:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1178.eqiad.wmnet with OS bookworm * 15:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock1003.eqiad.wmnet * 15:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock1003.eqiad.wmnet with OS trixie * 15:16 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 15:16 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 15:16 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 15:16 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:16 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:16 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:12 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 15:12 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2003.codfw.wmnet * 15:11 cgoubert@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host rdb-lock2003.codfw.wmnet * 15:11 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 15:11 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 15:11 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:11 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:11 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:07 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324338{{!}}InitialiseSettings: Enable 2FA warnings on more private wikis (T428103)]] (duration: 07m 02s) * 15:04 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 15:04 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock1003.eqiad.wmnet with reason: host reimage * 15:03 reedy@deploy1003: reedy: Continuing with deployment * 15:02 reedy@deploy1003: reedy: Backport for [[gerrit:1324338{{!}}InitialiseSettings: Enable 2FA warnings on more private wikis (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:02 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:00 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324338{{!}}InitialiseSettings: Enable 2FA warnings on more private wikis (T428103)]] * 14:57 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 14:57 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2003.codfw.wmnet * 14:57 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock1003.eqiad.wmnet with reason: host reimage * 14:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock2002.codfw.wmnet * 14:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock2002.codfw.wmnet with OS trixie * 14:56 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 14:56 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 14:56 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 14:55 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 14:55 moritzm: powercycle ganeti2046 * 14:47 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock1003.eqiad.wmnet with OS trixie * 14:46 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1003.eqiad.wmnet - cgoubert@cumin2003" * 14:46 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1003.eqiad.wmnet - cgoubert@cumin2003" * 14:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock1003.eqiad.wmnet on all recursors * 14:45 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock1003.eqiad.wmnet on all recursors * 14:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1003.eqiad.wmnet - cgoubert@cumin2003" * 14:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1272.eqiad.wmnet with reason: Enabling notifications * 14:44 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1003.eqiad.wmnet - cgoubert@cumin2003" * 14:44 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324719{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324720{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324722{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0 (T434187)]] (duration: 11m 02s) * 14:39 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 14:39 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock1003.eqiad.wmnet * 14:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock2002.codfw.wmnet with reason: host reimage * 14:37 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 14:37 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock1002.eqiad.wmnet * 14:37 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock1002.eqiad.wmnet with OS trixie * 14:37 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1324719{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324720{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324722{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0 (T434187)]] synced to the testservers (see https://wikitech. * 14:33 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock2002.codfw.wmnet with reason: host reimage * 14:33 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1324719{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324720{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324722{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0 (T434187)]] * 14:32 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1018.eqiad.wmnet with OS bookworm * 14:32 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2046.codfw.wmnet * 14:31 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1020.eqiad.wmnet with OS bookworm * 14:27 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2046.codfw.wmnet * 14:25 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2045.codfw.wmnet * 14:25 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2045.codfw.wmnet * 14:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock1002.eqiad.wmnet with reason: host reimage * 14:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1019.eqiad.wmnet with OS bookworm * 14:22 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Checking sanitization for wikis testwiki in section s3 * 14:20 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2045.codfw.wmnet * 14:18 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock1002.eqiad.wmnet with reason: host reimage * 14:17 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:16 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324707{{!}}Backport all changes from wmf/1.47.0-wmf.15]] (duration: 40m 51s) * 14:16 moritzm: installing Linux 6.1.180 on Bookworm hosts * 14:15 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock2002.codfw.wmnet with OS trixie * 14:14 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2002.codfw.wmnet - cgoubert@cumin2003" * 14:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2002.codfw.wmnet - cgoubert@cumin2003" * 14:14 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:14 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2045.codfw.wmnet * 14:14 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2002.codfw.wmnet on all recursors * 14:14 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2002.codfw.wmnet on all recursors * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2002.codfw.wmnet - cgoubert@cumin2003" * 14:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2002.codfw.wmnet - cgoubert@cumin2003" * 14:12 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2030.codfw.wmnet * 14:12 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2030.codfw.wmnet * 14:11 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on apifeatureusage2001.codfw.wmnet with reason: host reimage * 14:09 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:08 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:07 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 14:06 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2030.codfw.wmnet * 14:06 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:06 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock1002.eqiad.wmnet with OS trixie * 14:05 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 14:05 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2002.codfw.wmnet * 14:05 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1002.eqiad.wmnet - cgoubert@cumin2003" * 14:05 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1002.eqiad.wmnet - cgoubert@cumin2003" * 14:05 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock1002.eqiad.wmnet on all recursors * 14:05 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock1002.eqiad.wmnet on all recursors * 14:05 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:05 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1002.eqiad.wmnet - cgoubert@cumin2003" * 14:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:04 kharlan@deploy1003: kharlan: Continuing with deployment * 14:04 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock2001.codfw.wmnet * 14:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock2001.codfw.wmnet with OS trixie * 14:02 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1002.eqiad.wmnet - cgoubert@cumin2003" * 14:02 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on apifeatureusage2001.codfw.wmnet with reason: host reimage * 14:01 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2030.codfw.wmnet * 13:59 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2029.codfw.wmnet * 13:58 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2029.codfw.wmnet * 13:58 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 13:58 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock1002.eqiad.wmnet * 13:56 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1019.eqiad.wmnet with reason: host reimage * 13:54 btullis@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on archiva1002.wikimedia.org with reason: Upgrading in-place * 13:53 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock1001.eqiad.wmnet * 13:53 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock1001.eqiad.wmnet with OS trixie * 13:53 kharlan@deploy1003: kharlan: Backport for [[gerrit:1324707{{!}}Backport all changes from wmf/1.47.0-wmf.15]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:52 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2029.codfw.wmnet * 13:52 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1020.eqiad.wmnet with reason: host reimage * 13:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1019.eqiad.wmnet with reason: host reimage * 13:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1020.eqiad.wmnet with reason: host reimage * 13:49 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock2001.codfw.wmnet with reason: host reimage * 13:48 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2029.codfw.wmnet * 13:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Configuring db1272 for s3 pooling', diff saved to https://phabricator.wikimedia.org/P96021 and previous config saved to /var/cache/conftool/dbconfig/20260812-134732-cwilliams.json * 13:44 bking@cumin2003: START - Cookbook sre.hosts.reimage for host apifeatureusage2001.codfw.wmnet with OS bookworm * 13:43 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock2001.codfw.wmnet with reason: host reimage * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2028.codfw.wmnet * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2028.codfw.wmnet * 13:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1018.eqiad.wmnet with reason: host reimage * 13:38 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock1001.eqiad.wmnet with reason: host reimage * 13:36 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1018.eqiad.wmnet with reason: host reimage * 13:35 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1324707{{!}}Backport all changes from wmf/1.47.0-wmf.15]] * 13:35 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2028.codfw.wmnet * 13:32 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1275938{{!}}Enable campaignEvents on bdwikimedia (T424016)]] (duration: 07m 35s) * 13:32 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1020.eqiad.wmnet with OS bookworm * 13:32 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1019.eqiad.wmnet with OS bookworm * 13:32 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock1001.eqiad.wmnet with reason: host reimage * 13:28 kharlan@deploy1003: kharlan, yahya: Continuing with deployment * 13:27 kharlan@deploy1003: kharlan, yahya: Backport for [[gerrit:1275938{{!}}Enable campaignEvents on bdwikimedia (T424016)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:27 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2028.codfw.wmnet * 13:26 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 13:26 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 13:25 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1275938{{!}}Enable campaignEvents on bdwikimedia (T424016)]] * 13:25 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitize-wiki (exit_code=99) Managing sanitization for wikis testwiki in section s3 * 13:24 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock2001.codfw.wmnet with OS trixie * 13:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2001.codfw.wmnet - cgoubert@cumin2003" * 13:24 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2001.codfw.wmnet - cgoubert@cumin2003" * 13:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2001.codfw.wmnet on all recursors * 13:23 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2001.codfw.wmnet on all recursors * 13:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2001.codfw.wmnet - cgoubert@cumin2003" * 13:23 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2001.codfw.wmnet - cgoubert@cumin2003" * 13:23 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324713{{!}}thwiki: reinstate temporary wiki25 logos (T431094)]] (duration: 07m 13s) * 13:20 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:20 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1018.eqiad.wmnet with OS bookworm * 13:20 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2027.codfw.wmnet * 13:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2027.codfw.wmnet * 13:19 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics-external: apply * 13:18 kharlan@deploy1003: anzx, kharlan: Continuing with deployment * 13:18 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:18 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock1001.eqiad.wmnet with OS trixie * 13:17 kharlan@deploy1003: anzx, kharlan: Backport for [[gerrit:1324713{{!}}thwiki: reinstate temporary wiki25 logos (T431094)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1001.eqiad.wmnet - cgoubert@cumin2003" * 13:17 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 13:17 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1001.eqiad.wmnet - cgoubert@cumin2003" * 13:17 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2001.codfw.wmnet * 13:17 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics-external: apply * 13:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock1001.eqiad.wmnet on all recursors * 13:17 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock1001.eqiad.wmnet on all recursors * 13:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1001.eqiad.wmnet - cgoubert@cumin2003" * 13:17 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1001.eqiad.wmnet - cgoubert@cumin2003" * 13:15 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1324713{{!}}thwiki: reinstate temporary wiki25 logos (T431094)]] * 13:15 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:15 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics-external: apply * 13:15 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:14 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics-external: apply * 13:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2027.codfw.wmnet * 13:12 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 13:12 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock1001.eqiad.wmnet * 13:12 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2027.codfw.wmnet * 13:08 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1158.eqiad.wmnet onto db1273.eqiad.wmnet * 13:07 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1158: Pool db1158.eqiad.wmnet in after cloning * 13:02 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti6002.drmrs.wmnet * 13:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti6002.drmrs.wmnet * 12:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1015.eqiad.wmnet with OS bookworm * 12:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti6002.drmrs.wmnet * 12:44 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1017.eqiad.wmnet with OS bookworm * 12:36 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 12:35 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 12:34 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 12:33 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 12:31 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti6002.drmrs.wmnet * 12:24 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1017.eqiad.wmnet with reason: host reimage * 12:22 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1158: Pool db1158.eqiad.wmnet in after cloning * 12:18 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1017.eqiad.wmnet with reason: host reimage * 12:11 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1015.eqiad.wmnet with reason: host reimage * 12:07 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1015.eqiad.wmnet with reason: host reimage * 12:04 moritzm: failover ganeti master in drmrs02 to ganeti6004 * 12:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1009.eqiad.wmnet with OS bookworm * 12:01 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1017.eqiad.wmnet with OS bookworm * 12:00 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti6004.drmrs.wmnet * 12:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti6004.drmrs.wmnet * 11:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti6004.drmrs.wmnet * 11:53 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1015.eqiad.wmnet with OS bookworm * 11:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1159.eqiad.wmnet onto db1274.eqiad.wmnet * 11:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Pool db1159.eqiad.wmnet in after cloning * 11:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1016.eqiad.wmnet with OS bookworm * 11:46 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti6004.drmrs.wmnet * 11:45 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti6001.drmrs.wmnet * 11:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti6001.drmrs.wmnet * 11:43 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324306{{!}}WikimediaAntiAbuse: Enable personal info for enwiki with no display (T431292)]] (duration: 10m 26s) * 11:42 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1015.eqiad.wmnet with OS bookworm * 11:39 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 11:38 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti6001.drmrs.wmnet * 11:34 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1324306{{!}}WikimediaAntiAbuse: Enable personal info for enwiki with no display (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:33 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2212: Security update * 11:33 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti6001.drmrs.wmnet * 11:32 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1324306{{!}}WikimediaAntiAbuse: Enable personal info for enwiki with no display (T431292)]] * 11:22 moritzm: failover ganeti master in drmrs01 to ganeti6003 * 11:20 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:20 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:18 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 22 hosts with reason: Cloning * 11:17 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti6003.drmrs.wmnet * 11:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti6003.drmrs.wmnet * 11:17 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1009.eqiad.wmnet with reason: host reimage * 11:17 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:16 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:14 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1016.eqiad.wmnet with reason: host reimage * 11:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti6003.drmrs.wmnet * 11:11 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1009.eqiad.wmnet with reason: host reimage * 11:10 jmm@cumin2003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti3005.esams.wmnet * 11:10 jmm@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host ganeti3005.esams.wmnet * 11:09 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 11:08 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1158: Depool db1158.eqiad.wmnet to then clone it to db1273.eqiad.wmnet - marostegui@cumin1003 * 11:07 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1016.eqiad.wmnet with reason: host reimage * 11:07 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1158: Depool db1158.eqiad.wmnet to then clone it to db1273.eqiad.wmnet - marostegui@cumin1003 * 11:07 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1158.eqiad.wmnet onto db1273.eqiad.wmnet * 11:06 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:05 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:05 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1159: Pool db1159.eqiad.wmnet in after cloning * 11:04 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 20 hosts with reason: Cloning * 11:02 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti6003.drmrs.wmnet * 11:00 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis testwiki in section s3 * 10:54 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1009.eqiad.wmnet with OS bookworm * 10:51 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1015.eqiad.wmnet with OS bookworm * 10:50 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1016.eqiad.wmnet with OS bookworm * 10:48 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2212: Security update * 10:45 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitize-wiki (exit_code=99) Managing sanitization for wikis testwiki in section s3 * 10:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1013.eqiad.wmnet with OS bookworm * 10:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1014.eqiad.wmnet with OS bookworm * 10:38 cwilliams@cumin1003: START - Cookbook sre.mysql.clone of db1159.eqiad.wmnet onto db1274.eqiad.wmnet * 10:33 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db1274.eqiad.wmnet * 10:33 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db1274.eqiad.wmnet * 10:31 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 396993 * 10:29 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 396993 * 10:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1159: Clone source for db1274 * 10:24 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1159: Clone source for db1274 * 10:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1013.eqiad.wmnet with reason: host reimage * 10:18 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1013.eqiad.wmnet with reason: host reimage * 10:13 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2212.codfw.wmnet with reason: Maintenance * 10:13 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 10:12 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 10:12 blake@deploy1003: Stopping before sync operations * 10:11 blake@deploy1003: Started scap sync-world: Non-deployment scap run to populate new release values for [[phab:T427668|T427668]] * 10:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2212 [[phab:T434644|T434644]]', diff saved to https://phabricator.wikimedia.org/P96003 and previous config saved to /var/cache/conftool/dbconfig/20260812-101053-cwilliams.json * 10:09 moritzm: powercycle ganeti3005 * 10:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2203 to s1 primary [[phab:T434644|T434644]]', diff saved to https://phabricator.wikimedia.org/P96002 and previous config saved to /var/cache/conftool/dbconfig/20260812-100849-cwilliams.json * 10:08 cezmunsta: Starting s1 codfw failover from db2212 to db2203 - [[phab:T434644|T434644]] * 10:03 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1013.eqiad.wmnet with OS bookworm * 10:02 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1013.eqiad.wmnet with OS bookworm * 10:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2203 with weight 0 [[phab:T434644|T434644]]', diff saved to https://phabricator.wikimedia.org/P96001 and previous config saved to /var/cache/conftool/dbconfig/20260812-100134-cwilliams.json * 10:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 32 hosts with reason: Primary switchover s1 [[phab:T434644|T434644]] * 09:53 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1014.eqiad.wmnet with reason: host reimage * 09:50 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 09:50 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti3005.esams.wmnet * 09:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1014.eqiad.wmnet with reason: host reimage * 09:41 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1278: Pool in x1 * 09:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-launcher1003.eqiad.wmnet with OS bookworm * 09:37 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti3005.esams.wmnet * 09:37 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1013.eqiad.wmnet with OS bookworm * 09:34 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-presto1013.eqiad.wmnet with OS bookworm * 09:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1014.eqiad.wmnet with OS bookworm * 09:29 moritzm: failover ganeti master in esams to ganeti3008 * 09:26 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti3006.esams.wmnet * 09:26 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti3006.esams.wmnet * 09:24 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1012.eqiad.wmnet with OS bookworm * 09:23 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1009.eqiad.wmnet with OS bookworm * 09:23 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 09:18 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti3006.esams.wmnet * 09:16 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti3006.esams.wmnet * 09:14 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: fix regexp escaping bug - oblivian@cumin1003" * 09:14 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: fix regexp escaping bug - oblivian@cumin1003 * 09:13 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: fix regexp escaping bug - oblivian@cumin1003 * 09:13 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: fix regexp escaping bug - oblivian@cumin1003" * 09:03 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-launcher1003.eqiad.wmnet with reason: host reimage * 08:58 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-launcher1003.eqiad.wmnet with reason: host reimage * 08:55 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1278: Pool in x1 * 08:55 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1278 to dbctl [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95996 and previous config saved to /var/cache/conftool/dbconfig/20260812-085521-marostegui.json * 08:51 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1012.eqiad.wmnet with reason: host reimage * 08:45 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis testwiki in section s3 * 08:43 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1009.eqiad.wmnet with OS bookworm * 08:42 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1012.eqiad.wmnet with reason: host reimage * 08:41 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-launcher1003.eqiad.wmnet with OS bookworm * 08:40 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1013.eqiad.wmnet with OS bookworm * 08:38 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms3', diff saved to https://phabricator.wikimedia.org/P95995 and previous config saved to /var/cache/conftool/dbconfig/20260812-083816-marostegui.json * 08:38 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-master1003.eqiad.wmnet with OS bookworm * 08:37 marostegui: Failover ms3 [[phab:T434288|T434288]] * 08:37 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1268 to dbctl [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95994 and previous config saved to /var/cache/conftool/dbconfig/20260812-083722-marostegui.json * 08:35 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-web1001.eqiad.wmnet with OS bookworm * 08:32 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db2252.codfw.wmnet,db[1153,1268].eqiad.wmnet with reason: Switching over ms3 * 08:28 marostegui@cumin1003: dbctl commit (dc=all): 'Depool ms3 [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95993 and previous config saved to /var/cache/conftool/dbconfig/20260812-082852-marostegui.json * 08:25 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1012.eqiad.wmnet with OS bookworm * 08:25 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1011.eqiad.wmnet with OS bookworm * 08:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-master1003.eqiad.wmnet with reason: host reimage * 08:07 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-master1003.eqiad.wmnet with reason: host reimage * 08:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-web1001.eqiad.wmnet with reason: host reimage * 07:58 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-web1001.eqiad.wmnet with reason: host reimage * 07:50 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1003.eqiad.wmnet with OS bookworm * 07:38 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1011.eqiad.wmnet with reason: host reimage * 07:38 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-master1003.eqiad.wmnet with OS bookworm * 07:35 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti3007.esams.wmnet * 07:35 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti3007.esams.wmnet * 07:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1011.eqiad.wmnet with reason: host reimage * 07:27 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti3007.esams.wmnet * 07:25 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti3007.esams.wmnet * 07:25 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti3008.esams.wmnet * 07:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti3008.esams.wmnet * 07:22 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-web1001.eqiad.wmnet with OS bookworm * 07:19 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1003.eqiad.wmnet with OS bookworm * 07:18 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-master1003.eqiad.wmnet * 07:18 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host an-master1003.eqiad.wmnet * 07:17 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1011.eqiad.wmnet with OS bookworm * 07:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 07:16 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti3008.esams.wmnet * 07:15 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1009.eqiad.wmnet with OS bookworm * 07:14 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host an-master1003.eqiad.wmnet * 07:13 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-master1003.eqiad.wmnet * 07:13 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-master1003.eqiad.wmnet * 07:12 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-master1003.eqiad.wmnet * 07:11 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti3008.esams.wmnet * 07:07 arnaudb@dns1006: END - running authdns-update * 07:07 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti5007.eqsin.wmnet * 07:07 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti5007.eqsin.wmnet * 07:05 arnaudb@dns1006: START - running authdns-update * 06:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti5007.eqsin.wmnet * 06:54 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti5007.eqsin.wmnet * 06:38 moritzm: failover ganeti master in eqsin to ganeti5004 * 06:36 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti5006.eqsin.wmnet * 06:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti5006.eqsin.wmnet * 06:28 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti5006.eqsin.wmnet * 06:23 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti5006.eqsin.wmnet * 06:20 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti5005.eqsin.wmnet * 06:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti5005.eqsin.wmnet * 06:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti5005.eqsin.wmnet * 06:06 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti5005.eqsin.wmnet * 06:03 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti5004.eqsin.wmnet * 06:03 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti5004.eqsin.wmnet * 05:55 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti5004.eqsin.wmnet * 05:53 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti5004.eqsin.wmnet * 04:40 ryankemper: [[phab:T434494|T434494]] reimaged `an-tool1008.eqiad.wmnet` to bookworm; yarn.wikimedia.org is back up * 04:16 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-tool1008.eqiad.wmnet with OS bookworm * 03:58 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-tool1008.eqiad.wmnet with reason: host reimage * 03:53 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-tool1008.eqiad.wmnet with reason: host reimage * 03:41 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-tool1008.eqiad.wmnet with OS bookworm * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 45s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 00:25 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324427{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]], [[gerrit:1324429{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0]], [[gerrit:1324428{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]] (duration: 07m 55s) * 00:21 kemayo@deploy1003: kemayo: Continuing with deployment * 00:19 kemayo@deploy1003: kemayo: Backport for [[gerrit:1324427{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]], [[gerrit:1324429{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0]], [[gerrit:1324428{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:17 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1324427{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]], [[gerrit:1324429{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0]], [[gerrit:1324428{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]] == 2026-08-11 == * 21:37 sbassett: Deployed security fix for [[phab:T434521|T434521]] (wmf.15) * 21:29 sbassett: Deployed security fix for [[phab:T434521|T434521]] (wmf.14) * 21:19 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324370{{!}}Phase 4 of legal footer deployment (T432796)]], [[gerrit:1319804{{!}}Disable wgMFCustomSiteModules on English Wikipedia (T375538)]] (duration: 15m 26s) * 21:15 jdlrobson@deploy1003: jdlrobson: Continuing with deployment * 21:06 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1324370{{!}}Phase 4 of legal footer deployment (T432796)]], [[gerrit:1319804{{!}}Disable wgMFCustomSiteModules on English Wikipedia (T375538)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:03 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1324370{{!}}Phase 4 of legal footer deployment (T432796)]], [[gerrit:1319804{{!}}Disable wgMFCustomSiteModules on English Wikipedia (T375538)]] * 20:59 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 20:50 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324384{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]], [[gerrit:1324385{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]] (duration: 06m 58s) * 20:46 kemayo@deploy1003: kemayo: Continuing with deployment * 20:45 kemayo@deploy1003: kemayo: Backport for [[gerrit:1324384{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]], [[gerrit:1324385{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:43 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1324384{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]], [[gerrit:1324385{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]] * 20:42 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 20:42 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324386{{!}}build: Updating js-yaml to 3.15.1, 4.3.1]] (duration: 07m 36s) * 20:38 kemayo@deploy1003: kemayo: Continuing with deployment * 20:37 jhancock@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 20:37 kemayo@deploy1003: kemayo: Backport for [[gerrit:1324386{{!}}build: Updating js-yaml to 3.15.1, 4.3.1]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:35 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1324386{{!}}build: Updating js-yaml to 3.15.1, 4.3.1]] * 20:18 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 20:15 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 20:15 jhancock@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin1003" * 20:14 jhancock@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin1003" * 19:59 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 19:54 jhancock@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 19:10 brennen@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] (duration: 06m 41s) * 19:04 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns6001.wikimedia.org * 19:04 sukhe@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns6001.wikimedia.org * 19:04 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns5003.wikimedia.org * 19:04 sukhe@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns5003.wikimedia.org * 19:03 brennen@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 18:59 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns5003.wikimedia.org with OS trixie * 18:55 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns6001.wikimedia.org with OS trixie * 18:19 brett@cumin2002: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on P<nowiki>{</nowiki>cp7009.magru.wmnet<nowiki>}</nowiki> and A:cp - 9.2.15 Upgrade () * 18:14 brett@cumin2002: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on P<nowiki>{</nowiki>cp7009.magru.wmnet<nowiki>}</nowiki> and A:cp - 9.2.15 Upgrade () * 18:13 brennen@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 18:12 brett@cumin2002: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 9.2.15 Upgrade () * 18:09 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns5003.wikimedia.org with reason: host reimage * 18:06 brett@cumin2002: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 9.2.15 Upgrade () * 18:06 brennen: 1.47.0-wmf.15 train status ([[phab:T430834|T430834]]) - no current blockers, rolling to group0 * 18:05 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns5003.wikimedia.org with reason: host reimage * 18:05 brett: import trafficserver-9.2.15~deb13+wmf1 into trixie-wikimedia ([[phab:T434478|T434478]]) * 18:01 ladsgroup@cumin1003: END (PASS) - Cookbook sre.mysql.sanitarium_restart (exit_code=0) * 17:58 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns6001.wikimedia.org with reason: host reimage * 17:53 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324369{{!}}Enable desktop lazy loading on group0 (T148047)]] (duration: 07m 31s) * 17:52 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns6001.wikimedia.org with reason: host reimage * 17:49 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 17:49 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitarium_restart (exit_code=99) * 17:49 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 17:49 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7001.magru.wmnet * 17:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti7001.magru.wmnet * 17:48 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 17:47 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1324369{{!}}Enable desktop lazy loading on group0 (T148047)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:45 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1324369{{!}}Enable desktop lazy loading on group0 (T148047)]] * 17:39 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti7001.magru.wmnet * 17:36 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns5003.wikimedia.org with OS trixie * 17:34 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns6001.wikimedia.org with OS trixie * 17:31 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324361{{!}}Move FR config from IS.php to a dedicated file]], [[gerrit:1324363{{!}}Remove $wmg = $wg hacks in CentralAuth (T119117)]] (duration: 12m 23s) * 17:26 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 17:23 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1324361{{!}}Move FR config from IS.php to a dedicated file]], [[gerrit:1324363{{!}}Remove $wmg = $wg hacks in CentralAuth (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:19 sukhe: sudo cumin "A:cp-magru" "run-puppet-agent --enable 'merging CR 1324355'" * 17:18 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1324361{{!}}Move FR config from IS.php to a dedicated file]], [[gerrit:1324363{{!}}Remove $wmg = $wg hacks in CentralAuth (T119117)]] * 17:11 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-master1004.eqiad.wmnet with OS bookworm * 17:11 sukhe: sukhe@cp7005:~$ sudo puppet agent -tv * 17:02 sukhe: sudo cumin "A:cp-magru" "disable-puppet 'merging CR 1324355'" * 16:54 sukhe@dns1004: END - running authdns-update * 16:53 sukhe@dns1004: START - running authdns-update * 16:53 sukhe@dns1004: FAIL - running authdns-update * 16:51 sukhe@dns1004: START - running authdns-update * 16:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-master1004.eqiad.wmnet with reason: host reimage * 16:44 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-master1004.eqiad.wmnet with reason: host reimage * 16:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1157.eqiad.wmnet onto db1272.eqiad.wmnet * 16:40 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1157: Pool db1157.eqiad.wmnet in after cloning * 16:38 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 16:31 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324356{{!}}InitialiseSettings: Fix wgOATHAuthEnforce2FAForAll]] (duration: 06m 52s) * 16:30 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 16:28 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 16:27 reedy@deploy1003: reedy: Continuing with deployment * 16:26 reedy@deploy1003: reedy: Backport for [[gerrit:1324356{{!}}InitialiseSettings: Fix wgOATHAuthEnforce2FAForAll]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:24 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324356{{!}}InitialiseSettings: Fix wgOATHAuthEnforce2FAForAll]] * 16:13 sukhe: restart ntpsec.serviceon dns7001 * 16:09 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324335{{!}}InitialiseSettings: Enable 2FA enforcement on various private wikis (T428103)]] (duration: 06m 40s) * 16:08 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 16:06 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2204: Security update * 16:04 reedy@deploy1003: reedy: Continuing with deployment * 16:04 reedy@deploy1003: reedy: Backport for [[gerrit:1324335{{!}}InitialiseSettings: Enable 2FA enforcement on various private wikis (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:02 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7001.magru.wmnet * 16:02 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324335{{!}}InitialiseSettings: Enable 2FA enforcement on various private wikis (T428103)]] * 16:01 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 15:55 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1157: Pool db1157.eqiad.wmnet in after cloning * 15:54 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 15:54 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 15:41 dancy@deploy1003: Finished scap sync-world: Testing (duration: 06m 28s) * 15:40 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-master1004.eqiad.wmnet with OS bookworm * 15:40 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 15:35 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti4008.ulsfo.wmnet * 15:35 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti4008.ulsfo.wmnet * 15:34 dancy@deploy1003: Started scap sync-world: Testing * 15:34 dancy@deploy1003: Installation of scap version "4.279.0" completed for 3 hosts * 15:34 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-master1004.eqiad.wmnet with OS bookworm * 15:34 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 15:33 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-master1004.eqiad.wmnet with OS bookworm * 15:32 dancy@deploy1003: Installing scap version "4.279.0" for 3 host(s) * 15:32 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324339{{!}}Add /w/deployment-info.php entrypoint]] (duration: 07m 25s) * 15:30 moritzm: failover ganeti master in magru to ganeti7004 * 15:29 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti4008.ulsfo.wmnet * 15:28 dancy@deploy1003: dancy: Continuing with deployment * 15:28 tappof: remove 2026-05 swift log archives from centrallog to free some space ([[phab:T434502|T434502]]) * 15:27 dancy@deploy1003: dancy: Backport for [[gerrit:1324339{{!}}Add /w/deployment-info.php entrypoint]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:25 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324339{{!}}Add /w/deployment-info.php entrypoint]] * 15:20 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2204: Security update * 15:18 dancy@deploy1003: Installation of scap version "4.278.0" completed for 3 hosts * 15:18 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7004.magru.wmnet * 15:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti7004.magru.wmnet * 15:16 dancy@deploy1003: Installing scap version "4.278.0" for 3 host(s) * 15:14 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2204.codfw.wmnet with reason: Maintenance * 15:12 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 15:11 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-master1004.eqiad.wmnet with OS bookworm * 15:11 hashar: Restarting CI Jenkins on contint1003 due to Java upgrade. * 15:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2204 [[phab:T434565|T434565]]', diff saved to https://phabricator.wikimedia.org/P95984 and previous config saved to /var/cache/conftool/dbconfig/20260811-151126-cwilliams.json * 15:10 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti4008.ulsfo.wmnet * 15:10 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti7004.magru.wmnet * 15:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2207 to s2 primary [[phab:T434565|T434565]]', diff saved to https://phabricator.wikimedia.org/P95983 and previous config saved to /var/cache/conftool/dbconfig/20260811-150905-cwilliams.json * 15:08 cezmunsta: Starting s2 codfw failover from db2204 to db2207 - [[phab:T434565|T434565]] * 15:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2207 with weight 0 [[phab:T434565|T434565]]', diff saved to https://phabricator.wikimedia.org/P95982 and previous config saved to /var/cache/conftool/dbconfig/20260811-150402-cwilliams.json * 15:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s2 [[phab:T434565|T434565]] * 14:55 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1010.eqiad.wmnet with OS bookworm * 14:49 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns4003.wikimedia.org with OS trixie * 14:47 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1010.eqiad.wmnet with OS bookworm * 14:47 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7001.wikimedia.org with OS trixie * 14:44 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-presto1010.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:41 btullis@cumin1003: START - Cookbook sre.hosts.provision for host an-presto1010.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:40 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-presto1010.eqiad.wmnet with OS bookworm * 14:39 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 14:39 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-presto1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:36 btullis@cumin1003: START - Cookbook sre.hosts.provision for host an-presto1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:32 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1009.eqiad.wmnet with OS bookworm * 14:32 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 14:31 cwilliams@cumin1003: START - Cookbook sre.mysql.clone of db1157.eqiad.wmnet onto db1272.eqiad.wmnet * 14:30 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7004.magru.wmnet * 14:28 moritzm: failover ganeti master in ulsfo to ganeti4005 * 14:23 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7003.magru.wmnet * 14:23 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti7003.magru.wmnet * 14:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-master1004.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:22 btullis@cumin1003: START - Cookbook sre.hosts.provision for host an-master1004.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:21 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-master1004.eqiad.wmnet with OS bookworm * 14:19 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti4007.ulsfo.wmnet * 14:19 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti4007.ulsfo.wmnet * 14:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1010.eqiad.wmnet with OS bookworm * 14:17 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1008.eqiad.wmnet with OS bookworm * 14:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti7003.magru.wmnet * 14:12 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7003.magru.wmnet * 14:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti4007.ulsfo.wmnet * 14:11 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7002.magru.wmnet * 14:11 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti7002.magru.wmnet * 14:09 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324318{{!}}Revert "wmf-config/ProductionServices: set URL for urldownloader to service record" (T429175)]] (duration: 06m 46s) * 14:09 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7001.wikimedia.org with reason: host reimage * 14:06 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti4007.ulsfo.wmnet * 14:05 kharlan@deploy1003: kharlan: Continuing with deployment * 14:05 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti4006.ulsfo.wmnet * 14:04 jayme: updated calico to v3.30.7 on staging-eqiad - [[phab:T427400|T427400]] * 14:04 kharlan@deploy1003: kharlan: Backport for [[gerrit:1324318{{!}}Revert "wmf-config/ProductionServices: set URL for urldownloader to service record" (T429175)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti4006.ulsfo.wmnet * 14:03 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns4003.wikimedia.org with reason: host reimage * 14:03 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7001.wikimedia.org with reason: host reimage * 14:02 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti7002.magru.wmnet * 14:02 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1324318{{!}}Revert "wmf-config/ProductionServices: set URL for urldownloader to service record" (T429175)]] * 14:02 btullis@dns1004: FAIL - running authdns-update * 14:00 btullis@dns1004: START - running authdns-update * 13:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1008.eqiad.wmnet with reason: host reimage * 13:59 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'. * 13:59 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1313985{{!}}wmf-config/ProductionServices: set URL for urldownloader to service record (T429175)]] (duration: 25m 06s) * 13:58 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7002.magru.wmnet * 13:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti4006.ulsfo.wmnet * 13:57 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns4003.wikimedia.org with reason: host reimage * 13:56 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'. * 13:56 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7001.magru.wmnet * 13:56 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1008.eqiad.wmnet with reason: host reimage * 13:55 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'. * 13:55 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'. * 13:55 kharlan@deploy1003: kharlan, sukhe: Continuing with deployment * 13:53 marostegui: Failover ms2 [[phab:T434288|T434288]] * 13:52 marostegui: Failover ms1 [[phab:T434288|T434288]] * 13:52 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7001.magru.wmnet * 13:51 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti4006.ulsfo.wmnet * 13:48 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti4005.ulsfo.wmnet * 13:48 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti4005.ulsfo.wmnet * 13:44 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti4005.ulsfo.wmnet * 13:40 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1008.eqiad.wmnet with OS bookworm * 13:39 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns4003.wikimedia.org with OS trixie * 13:38 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns7001.wikimedia.org with OS trixie * 13:38 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1008.eqiad.wmnet with OS bookworm * 13:37 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti4005.ulsfo.wmnet * 13:36 kharlan@deploy1003: kharlan, sukhe: Backport for [[gerrit:1313985{{!}}wmf-config/ProductionServices: set URL for urldownloader to service record (T429175)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:34 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1313985{{!}}wmf-config/ProductionServices: set URL for urldownloader to service record (T429175)]] * 13:29 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2034.codfw.wmnet * 13:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2034.codfw.wmnet * 13:25 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.clone (exit_code=99) of db1157.eqiad.wmnet onto db1272.eqiad.wmnet * 13:25 cwilliams@cumin1003: START - Cookbook sre.mysql.clone of db1157.eqiad.wmnet onto db1272.eqiad.wmnet * 13:21 urbanecm@deploy1003: mwscript-k8s job started: namespaceDupes.php --wiki=frwiktionary --fix # [[phab:T415716|T415716]] * 13:21 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2034.codfw.wmnet * 13:20 urbanecm@deploy1003: mwscript-k8s job started: namespaceDupes.php --wiki=frwiktionary # [[phab:T415716|T415716]] * 13:19 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1323350{{!}}[tgwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T415307)]], [[gerrit:1322961{{!}}[slwiki] Revert temporary logo for Wikipedia 25 (Vector legacy + Vector 2022) (T414265)]], [[gerrit:1323827{{!}}[itwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T414320)]] (duration: 08m 00s) * 13:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-coord1003.eqiad.wmnet with OS bookworm * 13:15 urbanecm@deploy1003: urbanecm, superpes: Continuing with deployment * 13:13 urbanecm@deploy1003: urbanecm, superpes: Backport for [[gerrit:1323350{{!}}[tgwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T415307)]], [[gerrit:1322961{{!}}[slwiki] Revert temporary logo for Wikipedia 25 (Vector legacy + Vector 2022) (T414265)]], [[gerrit:1323827{{!}}[itwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T414320)]] synced to the testservers (see https://wiki * 13:11 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1323350{{!}}[tgwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T415307)]], [[gerrit:1322961{{!}}[slwiki] Revert temporary logo for Wikipedia 25 (Vector legacy + Vector 2022) (T414265)]], [[gerrit:1323827{{!}}[itwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T414320)]] * 13:11 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1323779{{!}}[ukwiki] Remove reviewer usergroup (T434252)]], [[gerrit:1323312{{!}}[frwiktionary] Add new Schème namespace and its talk (T415716)]] (duration: 06m 49s) * 13:10 marostegui@dns1004: END - running authdns-update * 13:08 marostegui@dns1004: START - running authdns-update * 13:07 marostegui@cumin1003: dbctl commit (dc=all): 'Repool ms2 [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95980 and previous config saved to /var/cache/conftool/dbconfig/20260811-130725-marostegui.json * 13:06 urbanecm@deploy1003: urbanecm, superpes: Continuing with deployment * 13:06 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1266 to dbctl [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95979 and previous config saved to /var/cache/conftool/dbconfig/20260811-130627-marostegui.json * 13:06 urbanecm@deploy1003: urbanecm, superpes: Backport for [[gerrit:1323779{{!}}[ukwiki] Remove reviewer usergroup (T434252)]], [[gerrit:1323312{{!}}[frwiktionary] Add new Schème namespace and its talk (T415716)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:04 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1323779{{!}}[ukwiki] Remove reviewer usergroup (T434252)]], [[gerrit:1323312{{!}}[frwiktionary] Add new Schème namespace and its talk (T415716)]] * 12:59 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db2253.codfw.wmnet,db[1151,1266].eqiad.wmnet with reason: Switching over ms2 * 12:54 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1157: Using as clone source * 12:53 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1157: Using as clone source * 12:51 marostegui@cumin1003: dbctl commit (dc=all): 'Depool ms2 [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95977 and previous config saved to /var/cache/conftool/dbconfig/20260811-125129-marostegui.json * 12:47 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 12:46 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 12:45 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 12:44 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'recommendation-api-ng' for release 'main' . * 12:44 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'recommendation-api-ng' for release 'main' . * 12:43 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'recommendation-api-ng' for release 'main' . * 12:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-coord1003.eqiad.wmnet with reason: host reimage * 12:43 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'ores-legacy' for release 'main' . * 12:42 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'ores-legacy' for release 'main' . * 12:42 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2165: Security update * 12:40 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-coord1003.eqiad.wmnet with reason: host reimage * 12:39 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'ores-legacy' for release 'main' . * 12:38 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' . * 12:38 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' . * 12:37 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' . * 12:34 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2034.codfw.wmnet * 12:30 jmm@cumin2002: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti-test2001.codfw.wmnet * 12:30 jmm@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host ganeti-test2001.codfw.wmnet * 12:25 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1179.eqiad.wmnet onto db1278.eqiad.wmnet * 12:25 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1179: Pool db1179.eqiad.wmnet in after cloning * 12:23 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-coord1003.eqiad.wmnet with OS bookworm * 12:22 moritzm: failover ganeti master in codfw/routed to ganeti2033 * 12:22 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2033.codfw.wmnet * 12:22 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2033.codfw.wmnet * 12:19 jmm@cumin2002: START - Cookbook sre.hosts.reboot-single for host ganeti-test2001.codfw.wmnet * 12:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 12:18 jmm@cumin2002: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti-test2001.codfw.wmnet * 12:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1008.eqiad.wmnet with OS bookworm * 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2033.codfw.wmnet * 12:07 moritzm: failover ganeti master in ganeti/test to ganeti-test2003 * 12:04 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 12:03 jmm@cumin2003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti4005.ulsfo.wmnet * 12:03 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti4005.ulsfo.wmnet * 12:00 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324283{{!}}Use maximum compression level in SqlBlobStore and SqlBagOStuff (T428377)]] (duration: 11m 37s) * 11:57 jmm@cumin2002: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti-test2002.codfw.wmnet * 11:57 jmm@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti-test2002.codfw.wmnet * 11:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2165: Security update * 11:54 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 11:52 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1324283{{!}}Use maximum compression level in SqlBlobStore and SqlBagOStuff (T428377)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:51 jmm@cumin2002: START - Cookbook sre.hosts.reboot-single for host ganeti-test2002.codfw.wmnet * 11:50 jmm@cumin2002: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti-test2002.codfw.wmnet * 11:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2165.codfw.wmnet with reason: Maintenance * 11:48 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@050d19e] (releasing): [[phab:T434186|T434186]] (duration: 01m 14s) * 11:48 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1324283{{!}}Use maximum compression level in SqlBlobStore and SqlBagOStuff (T428377)]] * 11:47 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@050d19e] (releasing): [[phab:T434186|T434186]] * 11:44 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@050d19e] (releasing): test jenkins deploy for [[phab:T434186|T434186]] (duration: 01m 08s) * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2165 [[phab:T434514|T434514]]', diff saved to https://phabricator.wikimedia.org/P95969 and previous config saved to /var/cache/conftool/dbconfig/20260811-114352-cwilliams.json * 11:43 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@050d19e] (releasing): test jenkins deploy for [[phab:T434186|T434186]] * 11:42 jmm@cumin2003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti-test2003.codfw.wmnet * 11:42 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti-test2003.codfw.wmnet * 11:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2161 to s8 primary [[phab:T434514|T434514]]', diff saved to https://phabricator.wikimedia.org/P95968 and previous config saved to /var/cache/conftool/dbconfig/20260811-114136-cwilliams.json * 11:40 cezmunsta: Starting s8 codfw failover from db2165 to db2161 - [[phab:T434514|T434514]] * 11:40 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1179: Pool db1179.eqiad.wmnet in after cloning * 11:36 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti-test2003.codfw.wmnet * 11:36 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti-test2003.codfw.wmnet * 11:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2161 with weight 0 [[phab:T434514|T434514]]', diff saved to https://phabricator.wikimedia.org/P95966 and previous config saved to /var/cache/conftool/dbconfig/20260811-113449-cwilliams.json * 11:34 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 25 hosts with reason: Primary switchover s8 [[phab:T434514|T434514]] * 11:29 moritzm: installing Python 3.11 security updates * 11:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-presto1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 11:26 btullis@cumin1003: START - Cookbook sre.hosts.provision for host an-presto1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 11:23 btullis@dns1004: END - running authdns-update * 11:21 btullis@dns1004: START - running authdns-update * 11:20 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-presto1008.eqiad.wmnet with OS bookworm * 11:20 moritzm: installing curl security updates * 11:11 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1007.eqiad.wmnet with OS bookworm * 10:45 tappof: bump space for prometheus k8s-dse in codfw * 10:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1007.eqiad.wmnet with reason: host reimage * 10:38 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1007.eqiad.wmnet with reason: host reimage * 10:37 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-coord1004.eqiad.wmnet with OS bookworm * 10:35 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1008.eqiad.wmnet with OS bookworm * 10:34 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1179: Depool db1179.eqiad.wmnet to then clone it to db1278.eqiad.wmnet - marostegui@cumin1003 * 10:34 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1008.eqiad.wmnet with OS bookworm * 10:25 fceratto@cumin1003: dbctl commit (dc=all): 'Remove db1177 [[phab:T433474|T433474]]', diff saved to https://phabricator.wikimedia.org/P95964 and previous config saved to /var/cache/conftool/dbconfig/20260811-102527-fceratto.json * 10:22 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1008.eqiad.wmnet with OS bookworm * 10:22 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1007.eqiad.wmnet with OS bookworm * 10:21 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1006.eqiad.wmnet with OS bookworm * 10:20 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 10:18 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1179: Depool db1179.eqiad.wmnet to then clone it to db1278.eqiad.wmnet - marostegui@cumin1003 * 10:18 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1179.eqiad.wmnet onto db1278.eqiad.wmnet * 10:17 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 10:17 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 10:14 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 10:09 blake@deploy1003: Stopping before sync operations * 10:09 blake@deploy1003: Started scap sync-world: Non-deployment run to populate release values for [[phab:T427668|T427668]] * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 10:04 fceratto@cumin1003: Removing db1177 from zarcillo [[phab:T433474|T433474]] * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1177.eqiad.wmnet * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1177.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:03 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1177.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:00 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1006.eqiad.wmnet with reason: host reimage * 09:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-coord1004.eqiad.wmnet with reason: host reimage * 09:57 marostegui: Failover m1 from db1164 to db1213 - [[phab:T434493|T434493]] * 09:57 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1006.eqiad.wmnet with reason: host reimage * 09:55 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:54 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2232].codfw.wmnet,db[1164,1213,1217].eqiad.wmnet with reason: Primary switchover m1 [[phab:T434493|T434493]] * 09:52 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-coord1004.eqiad.wmnet with reason: host reimage * 09:49 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1213.eqiad.wmnet with OS trixie * 09:49 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1177.eqiad.wmnet * 09:41 moritzm: installing Linux 6.12.101 on Trixie hosts * 09:40 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1006.eqiad.wmnet with OS bookworm * 09:35 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-coord1004.eqiad.wmnet with OS bookworm * 09:28 moritzm: installing node-tar security updates * 09:27 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1213.eqiad.wmnet with reason: host reimage * 09:22 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1213.eqiad.wmnet with reason: host reimage * 09:09 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1177: Decommission * 09:08 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db1177: Decommission * 09:08 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 09:08 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.decommission (exit_code=99) * 09:06 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1213.eqiad.wmnet with OS trixie * 09:06 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 09:05 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1213.eqiad.wmnet with reason: Reimage * 08:53 marostegui@dns1004: END - running authdns-update * 08:51 marostegui@dns1004: START - running authdns-update * 08:48 marostegui: Switchover ms1 master in eqiad [[phab:T434288|T434288]] * 08:48 marostegui@cumin1003: dbctl commit (dc=all): 'Repool ms1 [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95962 and previous config saved to /var/cache/conftool/dbconfig/20260811-084804-marostegui.json * 08:40 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1267 to dbctl [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95961 and previous config saved to /var/cache/conftool/dbconfig/20260811-084054-marostegui.json * 08:29 marostegui: Failover m1 from db1213 to db1164 - [[phab:T434043|T434043]] * 08:25 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2232].codfw.wmnet,db[1164,1213,1217].eqiad.wmnet with reason: Primary switchover m1 [[phab:T434043|T434043]] * 08:22 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db2251.codfw.wmnet,db[1152,1267].eqiad.wmnet with reason: Switching over ms1 * 08:22 marostegui@cumin1003: dbctl commit (dc=all): 'Depool ms1 [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95960 and previous config saved to /var/cache/conftool/dbconfig/20260811-082201-marostegui.json * 08:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: Switching over ms1 * 08:20 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.parsercache (exit_code=99) * 08:20 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 08:20 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1152: Switching over ms1 * 08:19 slyngshede@dns1004: END - running authdns-update * 08:18 moritzm: installing openjdk-21 security updates * 08:17 slyngshede@dns1004: START - running authdns-update * 08:16 moritzm: imported jenkins 2.568.2 to thirdparty/jenkins for trixie-wikimedia * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.12 (duration: 02m 26s) * 03:36 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] (duration: 33m 33s) * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 35s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-10 == * 14:54 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1323973{{!}}mmv.bootstrap: Fix getUrlParam to account for TIFF lossy/lossless param (T434333)]] (duration: 11m 24s) * 14:50 krinkle@deploy1003: krinkle: Continuing with deployment * 14:45 krinkle@deploy1003: krinkle: Backport for [[gerrit:1323973{{!}}mmv.bootstrap: Fix getUrlParam to account for TIFF lossy/lossless param (T434333)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:43 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1323973{{!}}mmv.bootstrap: Fix getUrlParam to account for TIFF lossy/lossless param (T434333)]] * 14:07 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1323967{{!}}updateIsActiveFlagForMentees: Commit the final partial batch (T432959)]] (duration: 10m 33s) * 13:56 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1323967{{!}}updateIsActiveFlagForMentees: Commit the final partial batch (T432959)]] * 13:45 dani@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply * 13:45 dani@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply * 13:45 dani@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply * 13:45 dani@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply * 13:45 dani@deploy1003: helmfile [staging] DONE helmfile.d/services/miscweb: apply * 13:44 dani@deploy1003: helmfile [staging] START helmfile.d/services/miscweb: apply * 13:38 wmde-fisch@deploy1003: Finished scap sync-world: Backport for [[gerrit:1323939{{!}}Enable sub-references on more group2 wikis (batch3) (T432731)]] (duration: 33m 21s) * 13:25 wmde-fisch@deploy1003: wmde-fisch: Continuing with deployment * 13:22 wmde-fisch@deploy1003: wmde-fisch: Backport for [[gerrit:1323939{{!}}Enable sub-references on more group2 wikis (batch3) (T432731)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:05 wmde-fisch@deploy1003: Started scap sync-world: Backport for [[gerrit:1323939{{!}}Enable sub-references on more group2 wikis (batch3) (T432731)]] * 07:57 hashar@deploy1003: Finished deploy [integration/docroot@7772132]: update build dependencies (duration: 00m 13s) * 07:57 hashar@deploy1003: Started deploy [integration/docroot@7772132]: update build dependencies * 07:35 _joe_: restarting squid on urldownloader1006 * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 48s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-09 == * 16:01 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:01 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:01 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:00 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 36s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-08 == * 05:31 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9] (wcqs): [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) (duration: 02m 36s) * 05:28 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9] (wcqs): [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) * 04:56 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 04:55 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 04:47 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) (duration: 19m 22s) * 04:28 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) * 04:19 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) (duration: 00m 06s) * 04:18 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) * 04:17 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) (duration: 00m 28s) * 04:16 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) * 03:52 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 03:52 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 34s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-07 == * 23:30 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:29 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 22:45 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 22:43 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 22:41 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 22:41 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 20:54 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 20:32 andrewbogott: restarting puppetserver service on puppetserver* for [[phab:T434339|T434339]] * 19:52 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:45 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 19:32 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:25 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:22 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 19:21 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 18:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:41 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 18:35 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 18:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 18:22 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 18:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 18:16 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 18:12 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:09 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:08 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:07 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:04 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:01 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:00 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:00 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 17:59 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 17:25 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 17:14 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 17:13 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 17:13 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 17:13 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:54 maryum: Deployed security fix for [[phab:T434278|T434278]] * 16:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 16:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 16:27 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --verbose --scoreLessThan=0.7 --exceptDatasetChecksums=[[phab:T434319|T434319]]-enwiki-models.txt # [[phab:T434319|T434319]] * 16:06 cdobbins@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp5022.eqsin.wmnet with OS trixie * 15:13 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 14:19 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 14:17 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 13:50 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1156.eqiad.wmnet onto db1271.eqiad.wmnet * 13:50 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1271: Pool db1271.eqiad.wmnet in after cloning * 13:02 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1271: Pool db1271.eqiad.wmnet in after cloning * 12:19 jayme: updated calico to v3.30.7 on staging-codfw - [[phab:T427400|T427400]] * 12:09 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 12:06 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 12:06 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 12:05 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 12:02 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1156: Pool db1156.eqiad.wmnet in after cloning * 11:38 bjensen: sudo -i reprepro -C main include trixie-wikimedia $<nowiki>{</nowiki>HOME<nowiki>}</nowiki>/httpbb/trixie/httpbb_$<nowiki>{</nowiki>VERSION?<nowiki>}</nowiki>-1+deb13u1_amd64.changes #[[phab:T434052|T434052]] * 11:35 bjensen: sudo -i reprepro -C main include bookworm-wikimedia $<nowiki>{</nowiki>HOME<nowiki>}</nowiki>/httpbb/bookworm/httpbb_$<nowiki>{</nowiki>VERSION?<nowiki>}</nowiki>-1_amd64.changes #[[phab:T434052|T434052]] * 11:30 marostegui@cumin1003: dbctl commit (dc=all): 'Adding db1271 to dbctl', diff saved to https://phabricator.wikimedia.org/P95945 and previous config saved to /var/cache/conftool/dbconfig/20260807-113006-marostegui.json * 11:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1156: Pool db1156.eqiad.wmnet in after cloning * 10:23 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 10:22 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 10:22 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 10:21 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 10:20 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 10:20 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 10:19 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 10:18 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 10:06 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on 21 hosts with reason: cloning * 10:01 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1156: Depool db1156.eqiad.wmnet to then clone it to db1271.eqiad.wmnet - marostegui@cumin1003 * 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1156: Depool db1156.eqiad.wmnet to then clone it to db1271.eqiad.wmnet - marostegui@cumin1003 * 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1156.eqiad.wmnet onto db1271.eqiad.wmnet * 09:15 jynus: started stress testing db1245 dbs [[phab:T431115|T431115]] * 08:19 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:18 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:16 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:14 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:13 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:10 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:06 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:05 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:00 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 10 days, 0:00:00 on ml-serve1015.eqiad.wmnet with reason: Downtime to get full picture of current BIOS settings beyond what Redfish shows * 08:00 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 07:54 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 07:54 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 07:53 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:52 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:51 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:50 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:49 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:48 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:47 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:45 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:45 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:41 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:38 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:37 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 06:35 jayme: updated istio to 1.29.4 on wikikube eqiad - [[phab:T427401|T427401]] * 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1178.eqiad.wmnet * 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1178.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 06:06 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1178.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 05:55 marostegui@cumin1003: START - Cookbook sre.dns.netbox * 05:49 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1178.eqiad.wmnet * 05:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 05:46 marostegui@cumin1003: Removing db1178 from zarcillo [[phab:T433471|T433471]] * 05:45 marostegui@cumin1003: START - Cookbook sre.mysql.decommission * 02:42 denisse: Extended volume on prometheus2008 for the disk space alert as per https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_host_running_out_of_space * 02:37 denisse: Extended volume on prometheus2007 tor the disk space alert as per https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_host_running_out_of_space * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 56s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-06 == * 21:39 maryum: Deploy security patch for [[phab:T433070|T433070]] * 21:29 maryum: Deploy security patch for [[phab:T434189|T434189]] * 20:48 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] (duration: 08m 12s) * 20:44 aude@deploy1003: lmora, aude, anzx: Continuing with deployment * 20:41 aude@deploy1003: lmora, aude, anzx: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be * 20:41 ebernhardson@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:41 ebernhardson@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 20:40 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] * 20:37 ebernhardson@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:37 ebernhardson@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 20:32 ebernhardson@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:32 ebernhardson@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 20:31 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] (duration: 06m 41s) * 20:27 cjming@deploy1003: cjming, ebernhardson, chlod: Continuing with deployment * 20:26 cjming@deploy1003: cjming, ebernhardson, chlod: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:24 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] * 20:18 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] (duration: 09m 22s) * 20:14 cjming@deploy1003: cjming, tsev: Continuing with deployment * 20:11 cjming@deploy1003: cjming, tsev: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:09 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] * 19:41 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply * 19:40 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply * 19:31 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 19:31 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 19:00 cdobbins@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cp5022.eqsin.wmnet with OS trixie * 18:25 ladsgroup@deploy1003: Finished scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) (duration: 06m 08s) * 18:19 ladsgroup@deploy1003: Started scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) * 18:18 ladsgroup@deploy1003: Stopping before sync operations * 18:17 ladsgroup@deploy1003: Started scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) * 17:55 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 16:50 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 16:35 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1001.eqiad.wmnet with OS bookworm * 16:19 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1002.eqiad.wmnet with reason: host reimage * 16:16 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1002.eqiad.wmnet with reason: host reimage * 16:05 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1001.eqiad.wmnet with reason: host reimage * 16:00 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1001.eqiad.wmnet with reason: host reimage * 15:57 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 15:43 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm * 15:29 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1001.eqiad.wmnet with OS bookworm * 15:29 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:58 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm * 14:57 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-drmrs ([[phab:T428495|T428495]]) * 14:55 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-drmrs ([[phab:T428495|T428495]]) * 14:55 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-ui1001.eqiad.wmnet with OS bookworm * 14:54 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-presto1001.eqiad.wmnet with OS bookworm * 14:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-magru ([[phab:T428495|T428495]]) * 14:49 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-magru ([[phab:T428495|T428495]]) * 14:48 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1001.eqiad.wmnet with OS bookworm * 14:46 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-esams ([[phab:T428495|T428495]]) * 14:44 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-esams ([[phab:T428495|T428495]]) * 14:43 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:42 brouberol@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:42 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:42 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 14:40 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 14:40 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:38 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-ui1001.eqiad.wmnet with reason: host reimage * 14:34 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-presto1001.eqiad.wmnet with reason: host reimage * 14:28 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-ui1001.eqiad.wmnet with reason: host reimage * 14:27 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-presto1001.eqiad.wmnet with reason: host reimage * 14:23 sukhe: sudo cumin -b2 'A:cp-text' "run-puppet-agent --enable 'merging CR 1290731'": [[phab:T425441|T425441]] * 14:18 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo for hosts in the wikimedia.org domain - [[phab:T428495|T428495]] * 14:16 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-presto1001.eqiad.wmnet with OS bookworm * 14:14 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-ui1001.eqiad.wmnet with OS bookworm * 14:12 sukhe: sudo cumin 'A:cp-text' "disable-puppet 'merging CR 1290731'": [[phab:T425441|T425441]] * 14:11 swfrench-wmf: restarted navtiming on webperf1003 - [[phab:T428495|T428495]] * 14:04 swfrench-wmf: begin rolling restart of confd in drmrs, eqiad, esams, magru - [[phab:T428495|T428495]] * 14:04 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm * 14:04 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:02 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-client1002.eqiad.wmnet with OS bookworm * 13:58 swfrench-wmf: authdns update to direct eqiad-associated etcd clients back to eqiad - [[phab:T428495|T428495]] * 13:58 swfrench@dns1004: END - running authdns-update * 13:56 swfrench@dns1004: START - running authdns-update * 13:49 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:44 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:31 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 13:29 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 13:26 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 13:23 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 13:22 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 13:19 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 13:18 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 13:18 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 13:17 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 13:16 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 13:13 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 13:11 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 13:09 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 13:06 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 13:06 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-client1002.eqiad.wmnet with OS bookworm * 13:05 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revision-models' for release 'main' . * 13:05 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:05 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revision-models' for release 'main' . * 13:04 brouberol@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-test-client1002.eqiad.wmnet with OS bookworm * 13:04 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 13:03 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 13:02 aikochou@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:00 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'readability' for release 'main' . * 12:59 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'readability' for release 'main' . * 12:58 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 12:57 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'logo-detection' for release 'main' . * 12:57 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'logo-detection' for release 'main' . * 12:57 aikochou@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 12:55 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:54 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:53 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 12:53 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:50 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 12:48 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 12:46 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'article-models' for release 'main' . * 12:45 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'article-models' for release 'main' . * 12:41 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . * 12:39 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . * 12:38 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-client1002.eqiad.wmnet with OS bookworm * 12:12 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply * 12:12 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply * 12:09 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:08 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 11:58 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2187: Security update * 11:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:24 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:16 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 11:15 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 11:10 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2187: Security update * 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2187.codfw.wmnet with reason: Maintenance * 10:56 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 10:56 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 10:56 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 10:56 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 10:54 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 10:53 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 10:09 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2187: Security update * 10:07 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2187: Security update * 09:39 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms2', diff saved to https://phabricator.wikimedia.org/P95929 and previous config saved to /var/cache/conftool/dbconfig/20260806-093908-marostegui.json * 09:36 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1178 from dbctl [[phab:T433471|T433471]]', diff saved to https://phabricator.wikimedia.org/P95928 and previous config saved to /var/cache/conftool/dbconfig/20260806-093632-marostegui.json * 09:33 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 09:31 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 09:30 topranks: bounce cr3-eqsin<->cr2-eqiad bgp session to disable no-prepend command * 09:20 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2253.codfw.wmnet,db1151.eqiad.wmnet with reason: cloning * 09:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1151: Cloning * 09:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:19 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 09:19 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1151: Cloning * 09:10 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 09:09 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup2003.codfw.wmnet * 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup2003.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 09:06 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup2003.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 09:03 klausman@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:02 jynus@cumin1003: START - Cookbook sre.dns.netbox * 09:02 klausman@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 08:57 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup2003.codfw.wmnet * 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup1003.eqiad.wmnet * 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:54 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms3', diff saved to https://phabricator.wikimedia.org/P95925 and previous config saved to /var/cache/conftool/dbconfig/20260806-085422-marostegui.json * 08:53 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:46 jynus@cumin1003: START - Cookbook sre.dns.netbox * 08:39 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup1003.eqiad.wmnet * 08:29 XioNoX: push pfw policy - [[phab:T434115|T434115]] * 08:14 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 08:00 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 08:00 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:58 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revision-models' for release 'main' . * 07:56 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 07:54 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'readability' for release 'main' . * 07:53 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'logo-detection' for release 'main' . * 07:51 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'llm' for release 'main' . * 07:48 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . * 07:37 jayme: updated istio to 1.29.4 on wikikube codfw - [[phab:T427401|T427401]] * 07:08 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2252.codfw.wmnet,db1153.eqiad.wmnet with reason: cloning * 07:07 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1153: Cloning * 07:07 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1153: Cloning * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 40s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-05 == * 23:24 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1009.eqiad.wmnet with OS bookworm * 23:03 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1009.eqiad.wmnet with reason: host reimage * 22:59 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1009.eqiad.wmnet with reason: host reimage * 22:43 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1009.eqiad.wmnet with OS bookworm * 22:38 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1009.eqiad.wmnet * 22:34 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1009.eqiad.wmnet * 22:25 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1008.eqiad.wmnet with OS bookworm * 22:04 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1008.eqiad.wmnet with reason: host reimage * 22:00 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1008.eqiad.wmnet with reason: host reimage * 21:48 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:47 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:46 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:44 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1008.eqiad.wmnet with OS bookworm * 21:43 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:41 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1008.eqiad.wmnet * 21:36 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1008.eqiad.wmnet * 21:14 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:12 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad * 21:12 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad * 21:10 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=eqiad * 21:08 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:07 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:07 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=eqiad * 21:04 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:03 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1006 * 21:02 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1006 * 21:00 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:56 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:56 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:55 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 20:55 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 20:55 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1007.eqiad.wmnet with OS bookworm * 20:51 vriley@cumin1003: START - Cookbook sre.dns.netbox * 20:43 ebernhardson: [[phab:T434008|T434008]]: changing cloudelastic:9643 from auto_expand_replicas to number_of_replicas * 20:34 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1007.eqiad.wmnet with reason: host reimage * 20:27 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1007.eqiad.wmnet with reason: host reimage * 20:24 cjming: end of UTC late backport window * 20:23 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] (duration: 06m 26s) * 20:18 cjming@deploy1003: cjming: Continuing with deployment * 20:18 cjming@deploy1003: cjming: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:16 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] * 20:12 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1007.eqiad.wmnet with OS bookworm * 20:12 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] (duration: 08m 41s) * 20:08 swfrench@cumin2002: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host conf1007.eqiad.wmnet with OS bookworm * 20:08 jforrester@deploy1003: jforrester: Continuing with deployment * 20:07 jforrester@deploy1003: jforrester: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:03 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] * 19:51 inflatador: [bking@puppetserver1001] ~$ sudo puppetserver ca sign --certname an-worker1189.eqiad.wmnet [[phab:T434142|T434142]] * 19:47 bking@cumin2003: DONE (FAIL) - Cookbook sre.puppet.renew-cert (exit_code=99) for an-worker1189.eqiad.wmnet: Renew puppet certificate - bking@cumin2003 * 19:46 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:30 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1007.eqiad.wmnet with OS trixie * 19:30 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 19:29 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 19:20 swfrench-wmf: silenced EtcdRelicationDown 0cb709a9-f244-4f1e-971f-{{Gerrit|440ec65e7fd7}} - [[phab:T428495|T428495]] * 19:13 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1007.eqiad.wmnet with OS bookworm * 19:12 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1007.eqiad.wmnet with reason: host reimage * 19:09 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1007.eqiad.wmnet * 19:07 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1007.eqiad.wmnet with reason: host reimage * 19:03 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1007.eqiad.wmnet * 18:52 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1007.eqiad.wmnet with OS trixie * 18:52 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1007.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:35 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 18:34 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 18:34 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 18:30 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1007.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:28 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:28 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1007] - vriley@cumin1003" * 18:27 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1007] - vriley@cumin1003" * 18:23 vriley@cumin1003: START - Cookbook sre.dns.netbox * 18:22 vriley@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 18:22 robh@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:19 vriley@cumin1003: START - Cookbook sre.dns.netbox * 18:13 robh@cumin2002: START - Cookbook sre.hosts.provision for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:31 jasmine@cumin2002: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-main-eqiad * 17:12 mutante: LDAP - added vwalters to group ciadmin - [[phab:T433615|T433615]] * 16:58 aokoth@deploy1003: Finished deploy [phabricator/deployment@e2ebca5]: Deploy Phab (duration: 00m 34s) * 16:57 aokoth@deploy1003: Started deploy [phabricator/deployment@e2ebca5]: Deploy Phab * 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad * 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=eqiad * 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad * 16:53 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:41 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2187.codfw.wmnet * 16:41 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2187.codfw.wmnet * 16:41 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker2187.codfw.wmnet * 16:41 cgoubert@cumin2003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker2187.codfw.wmnet * 16:40 jasmine@cumin2002: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-main-eqiad * 16:40 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:34 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:25 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-magru and A:liberica ([[phab:T428495|T428495]]) * 16:23 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-magru and A:liberica ([[phab:T428495|T428495]]) * 16:20 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-drmrs and A:liberica ([[phab:T428495|T428495]]) * 16:19 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-drmrs and A:liberica ([[phab:T428495|T428495]]) * 16:18 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-esams and A:liberica ([[phab:T428495|T428495]]) * 16:16 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-esams and A:liberica ([[phab:T428495|T428495]]) * 16:06 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1159.eqiad.wmnet * 16:06 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1159.eqiad.wmnet * 16:06 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1159.eqiad.wmnet * 16:05 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] (duration: 09m 11s) * 15:58 reedy@deploy1003: reedy: Continuing with deployment * 15:58 reedy@deploy1003: reedy: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:56 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] * 15:54 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1159.eqiad.wmnet with OS trixie * 15:38 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:33 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1159.eqiad.wmnet with reason: host reimage * 15:32 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:27 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1159.eqiad.wmnet with reason: host reimage * 15:10 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1159 * 15:10 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1159 * 15:00 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo for hosts in the wikimedia.org domain - [[phab:T428495|T428495]] * 14:55 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS trixie * 14:54 swfrench-wmf: restarted navtiming on webperf1003 - [[phab:T428495|T428495]] * 14:52 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1159 * 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1159.eqiad.wmnet 129.48.64.10.in-addr.arpa 9.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:52 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1159.eqiad.wmnet 129.48.64.10.in-addr.arpa 9.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1159 - jayme@cumin1003" * 14:52 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1159 - jayme@cumin1003" * 14:49 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:48 jayme@cumin1003: START - Cookbook sre.dns.netbox * 14:47 swfrench-wmf: begin rolling restart of confd in drmrs, eqiad, esams, magru - [[phab:T428495|T428495]] * 14:47 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1159 * 14:46 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:46 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:46 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1159.eqiad.wmnet with OS trixie * 14:45 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:44 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:44 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1159.eqiad.wmnet * 14:43 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:43 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1159.eqiad.wmnet * 14:43 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:43 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1159.eqiad.wmnet * 14:43 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:43 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1157.eqiad.wmnet * 14:43 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1157.eqiad.wmnet * 14:43 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1157.eqiad.wmnet * 14:42 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:42 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:42 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:41 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:41 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:41 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:41 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1003.eqiad.wmnet with OS bookworm * 14:39 swfrench-wmf: authdns update to direct eqiad-associated etcd clients to codfw - [[phab:T428495|T428495]] * 14:39 swfrench@dns1004: END - running authdns-update * 14:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 14:37 swfrench@dns1004: START - running authdns-update * 14:37 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:37 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:35 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:35 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 14:28 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:27 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1157.eqiad.wmnet with OS trixie * 14:27 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:27 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:27 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 14:26 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 14:26 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:26 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 14:26 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 14:26 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:26 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host search-loader1002.eqiad.wmnet with OS trixie * 14:19 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:19 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:15 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1003.eqiad.wmnet with reason: host reimage * 14:14 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:14 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:13 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 * 14:13 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host mc2046 * 14:13 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS trixie * 14:11 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1003.eqiad.wmnet with reason: host reimage * 14:10 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:09 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:09 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:09 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:08 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:08 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1157.eqiad.wmnet with reason: host reimage * 14:08 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:04 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 14:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on search-loader1002.eqiad.wmnet with reason: host reimage * 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=eqiad * 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=eqiad * 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=eqiad * 14:00 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:59 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:58 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1157.eqiad.wmnet with reason: host reimage * 13:57 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on search-loader1002.eqiad.wmnet with reason: host reimage * 13:54 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1003.eqiad.wmnet with OS bookworm * 13:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host search-loader1002.eqiad.wmnet with OS trixie * 13:43 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1157 * 13:42 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1157 * 13:40 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1157 * 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1157.eqiad.wmnet 183.32.64.10.in-addr.arpa 3.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1157.eqiad.wmnet 183.32.64.10.in-addr.arpa 3.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1157 - jayme@cumin1003" * 13:39 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1157 - jayme@cumin1003" * 13:39 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] (duration: 07m 00s) * 13:35 jayme@cumin1003: START - Cookbook sre.dns.netbox * 13:35 reedy@deploy1003: reedy: Continuing with deployment * 13:34 reedy@deploy1003: reedy: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:32 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] * 13:23 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1157 * 13:22 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1157.eqiad.wmnet with OS trixie * 13:22 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1157.eqiad.wmnet * 13:22 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1157.eqiad.wmnet * 13:21 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1157.eqiad.wmnet * 13:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1156.eqiad.wmnet * 13:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1156.eqiad.wmnet * 13:15 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1156.eqiad.wmnet * 13:01 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1156.eqiad.wmnet with OS trixie * 12:42 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1156.eqiad.wmnet with reason: host reimage * 12:38 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1156.eqiad.wmnet with reason: host reimage * 12:32 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 12:31 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 12:30 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 12:28 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 12:26 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 12:24 topranks: update bgp confed settings in eqsin * 12:22 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1156 * 12:22 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1156 * 12:22 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 12:19 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1156 * 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1156.eqiad.wmnet 110.32.64.10.in-addr.arpa 0.1.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:19 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1156.eqiad.wmnet 110.32.64.10.in-addr.arpa 0.1.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1156 - jayme@cumin1003" * 12:19 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1156 - jayme@cumin1003" * 12:17 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:14 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS trixie * 12:09 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:06 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:04 jayme@cumin1003: START - Cookbook sre.dns.netbox * 12:04 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 12:02 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'article-models' for release 'main' . * 12:01 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1156 * 12:01 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1156.eqiad.wmnet with OS trixie * 11:59 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1156.eqiad.wmnet * 11:59 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1156.eqiad.wmnet * 11:59 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1156.eqiad.wmnet * 11:57 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 11:53 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 11:53 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:52 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:52 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:50 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:50 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:50 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:49 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:48 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:47 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:47 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:45 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:45 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:44 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:44 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:44 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:43 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:42 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:38 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:35 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 * 11:35 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host mc2046 * 11:34 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS trixie * 11:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:27 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:21 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:21 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:18 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:18 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:18 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:18 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:13 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:13 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:09 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:08 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:07 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:06 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:06 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:05 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:05 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:04 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 11:04 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:24 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:24 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:17 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:16 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1155.eqiad.wmnet * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1155.eqiad.wmnet * 10:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1155.eqiad.wmnet * 10:14 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:14 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:11 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:11 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:10 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:09 aikochou@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop: sync * 10:09 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:09 aikochou@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop: sync * 10:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:07 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:05 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:05 aikochou@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop: sync * 10:05 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:05 aikochou@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop: sync * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:04 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:04 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:04 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1155.eqiad.wmnet with OS trixie * 09:52 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms1', diff saved to https://phabricator.wikimedia.org/P95918 and previous config saved to /var/cache/conftool/dbconfig/20260805-095212-marostegui.json * 09:44 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1152: after cloning * 09:44 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.parsercache (exit_code=99) * 09:44 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 09:44 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1152: after cloning * 09:43 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1155.eqiad.wmnet with reason: host reimage * 09:40 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1155.eqiad.wmnet with reason: host reimage * 09:32 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 09:32 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:31 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 09:31 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:27 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1155 * 09:27 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1155 * 09:25 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 09:24 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 09:24 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 09:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:23 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 09:23 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 09:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:22 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 09:22 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 09:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:20 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 09:20 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:17 XioNoX: push pfw policies - [[phab:T434038|T434038]] * 09:14 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1155 * 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1155.eqiad.wmnet 109.32.64.10.in-addr.arpa 9.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1155.eqiad.wmnet 109.32.64.10.in-addr.arpa 9.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1155 - jayme@cumin1003" * 09:14 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1155 - jayme@cumin1003" * 09:10 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 09:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:09 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2251.codfw.wmnet,db1152.eqiad.wmnet with reason: cloning * 09:09 jayme@cumin1003: START - Cookbook sre.dns.netbox * 09:08 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 09:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: Cloning * 09:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:05 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 09:05 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1152: Cloning * 08:38 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1155 * 08:37 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1155.eqiad.wmnet with OS trixie * 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1171.eqiad.wmnet * 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1171.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:29 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95913 and previous config saved to /var/cache/conftool/dbconfig/20260805-082908-ladsgroup.json * 08:27 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1171.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:22 jynus@cumin1003: START - Cookbook sre.dns.netbox * 08:18 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249', diff saved to https://phabricator.wikimedia.org/P95912 and previous config saved to /var/cache/conftool/dbconfig/20260805-081823-ladsgroup.json * 08:17 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1171.eqiad.wmnet * 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1150.eqiad.wmnet * 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1150.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:15 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1150.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:15 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 08:14 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1155.eqiad.wmnet * 08:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1155.eqiad.wmnet * 08:14 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1155.eqiad.wmnet * 08:11 jynus@cumin1003: START - Cookbook sre.dns.netbox * 08:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249', diff saved to https://phabricator.wikimedia.org/P95911 and previous config saved to /var/cache/conftool/dbconfig/20260805-080737-ladsgroup.json * 08:05 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1150.eqiad.wmnet * 08:02 marostegui: Depool clouddb1020 (s5,s8) [[phab:T434048|T434048]] * 08:02 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1020.eqiad.wmnet,service=s8 * 08:02 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1020.eqiad.wmnet,service=s5 * 08:02 marostegui: Depool clouddb1018 (s2,s7) [[phab:T434048|T434048]] * 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1018.eqiad.wmnet,service=s7 * 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1018.eqiad.wmnet,service=s2 * 08:01 marostegui: Depool clouddb1017 (s1) [[phab:T434048|T434048]] * 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1017.eqiad.wmnet,service=s1 * 07:59 marostegui: Depool clouddb1016 (s5,s8) [[phab:T434048|T434048]] * 07:59 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s8 * 07:59 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s5 * 07:57 marostegui: Depool clouddb1015 (s4,s6) [[phab:T434048|T434048]] * 07:57 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s6 * 07:57 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s4 * 07:56 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95910 and previous config saved to /var/cache/conftool/dbconfig/20260805-075650-ladsgroup.json * 07:54 marostegui: Depool clouddb1014 (s2,s7) [[phab:T434048|T434048]] * 07:54 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1014.eqiad.wmnet,service=s7 * 07:54 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1014.eqiad.wmnet,service=s2 * 07:53 marostegui: Depool clouddb1013:s1 [[phab:T434048|T434048]] * 07:53 marostegui: Depool clouddb1013:s1 [[phab:T409557|T409557]] * 07:53 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1013.eqiad.wmnet,service=s1 * 07:25 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95909 and previous config saved to /var/cache/conftool/dbconfig/20260805-072529-ladsgroup.json * 07:24 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2249.codfw.wmnet with reason: Maintenance * 07:24 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95908 and previous config saved to /var/cache/conftool/dbconfig/20260805-072426-ladsgroup.json * 07:21 slyngshede@dns1004: END - running authdns-update * 07:19 slyngshede@dns1004: START - running authdns-update * 07:13 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231', diff saved to https://phabricator.wikimedia.org/P95906 and previous config saved to /var/cache/conftool/dbconfig/20260805-071340-ladsgroup.json * 07:02 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231', diff saved to https://phabricator.wikimedia.org/P95905 and previous config saved to /var/cache/conftool/dbconfig/20260805-070253-ladsgroup.json * 06:52 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95904 and previous config saved to /var/cache/conftool/dbconfig/20260805-065206-ladsgroup.json * 06:45 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 06:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95903 and previous config saved to /var/cache/conftool/dbconfig/20260805-062240-ladsgroup.json * 06:21 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2231.codfw.wmnet with reason: Maintenance * 06:21 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95902 and previous config saved to /var/cache/conftool/dbconfig/20260805-062137-ladsgroup.json * 06:10 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215', diff saved to https://phabricator.wikimedia.org/P95901 and previous config saved to /var/cache/conftool/dbconfig/20260805-061051-ladsgroup.json * 06:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215', diff saved to https://phabricator.wikimedia.org/P95900 and previous config saved to /var/cache/conftool/dbconfig/20260805-060004-ladsgroup.json * 05:49 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95899 and previous config saved to /var/cache/conftool/dbconfig/20260805-054918-ladsgroup.json * 05:19 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95898 and previous config saved to /var/cache/conftool/dbconfig/20260805-051939-ladsgroup.json * 05:18 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2215.codfw.wmnet with reason: Maintenance * 04:30 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2201.codfw.wmnet with reason: Maintenance * 03:40 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2197.codfw.wmnet with reason: Maintenance * 03:40 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95897 and previous config saved to /var/cache/conftool/dbconfig/20260805-034036-ladsgroup.json * 03:29 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196', diff saved to https://phabricator.wikimedia.org/P95896 and previous config saved to /var/cache/conftool/dbconfig/20260805-032948-ladsgroup.json * 03:19 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196', diff saved to https://phabricator.wikimedia.org/P95895 and previous config saved to /var/cache/conftool/dbconfig/20260805-031902-ladsgroup.json * 03:08 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95894 and previous config saved to /var/cache/conftool/dbconfig/20260805-030815-ladsgroup.json * 02:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95893 and previous config saved to /var/cache/conftool/dbconfig/20260805-023413-ladsgroup.json * 02:33 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2196.codfw.wmnet with reason: Maintenance * 02:33 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95892 and previous config saved to /var/cache/conftool/dbconfig/20260805-023310-ladsgroup.json * 02:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186', diff saved to https://phabricator.wikimedia.org/P95891 and previous config saved to /var/cache/conftool/dbconfig/20260805-022223-ladsgroup.json * 02:11 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186', diff saved to https://phabricator.wikimedia.org/P95890 and previous config saved to /var/cache/conftool/dbconfig/20260805-021137-ladsgroup.json * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 02:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95889 and previous config saved to /var/cache/conftool/dbconfig/20260805-020051-ladsgroup.json * 01:30 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95888 and previous config saved to /var/cache/conftool/dbconfig/20260805-013029-ladsgroup.json * 01:29 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2186.codfw.wmnet with reason: Maintenance * 00:34 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on dbstore1009.eqiad.wmnet with reason: Maintenance * 00:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95887 and previous config saved to /var/cache/conftool/dbconfig/20260805-003408-ladsgroup.json * 00:23 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264', diff saved to https://phabricator.wikimedia.org/P95886 and previous config saved to /var/cache/conftool/dbconfig/20260805-002322-ladsgroup.json * 00:12 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264', diff saved to https://phabricator.wikimedia.org/P95885 and previous config saved to /var/cache/conftool/dbconfig/20260805-001235-ladsgroup.json * 00:01 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95884 and previous config saved to /var/cache/conftool/dbconfig/20260805-000148-ladsgroup.json == 2026-08-04 == * 23:45 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95883 and previous config saved to /var/cache/conftool/dbconfig/20260804-234508-ladsgroup.json * 23:44 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1264.eqiad.wmnet with reason: Maintenance * 23:44 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95882 and previous config saved to /var/cache/conftool/dbconfig/20260804-234405-ladsgroup.json * 23:33 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237', diff saved to https://phabricator.wikimedia.org/P95881 and previous config saved to /var/cache/conftool/dbconfig/20260804-233317-ladsgroup.json * 23:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237', diff saved to https://phabricator.wikimedia.org/P95880 and previous config saved to /var/cache/conftool/dbconfig/20260804-232230-ladsgroup.json * 23:11 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95879 and previous config saved to /var/cache/conftool/dbconfig/20260804-231144-ladsgroup.json * 22:23 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95878 and previous config saved to /var/cache/conftool/dbconfig/20260804-222345-ladsgroup.json * 22:23 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1237.eqiad.wmnet with reason: Maintenance * 21:13 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1225.eqiad.wmnet with reason: Maintenance * 20:40 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] (duration: 24m 40s) * 20:33 samtar@deploy1003: samtar, kineticpelagic: Continuing with deployment * 20:28 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS bookworm * 20:21 samtar@deploy1003: samtar, kineticpelagic: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:15 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] * 20:13 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 20:09 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 20:00 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1216.eqiad.wmnet with reason: Maintenance * 20:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95877 and previous config saved to /var/cache/conftool/dbconfig/20260804-195957-ladsgroup.json * 19:51 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 19:50 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) mc2046.codfw.wmnet 120.16.192.10.in-addr.arpa 0.2.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:50 jhancock@cumin2002: START - Cookbook sre.dns.wipe-cache mc2046.codfw.wmnet 120.16.192.10.in-addr.arpa 0.2.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host mc2046 - jhancock@cumin2002" * 19:50 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host mc2046 - jhancock@cumin2002" * 19:49 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203', diff saved to https://phabricator.wikimedia.org/P95876 and previous config saved to /var/cache/conftool/dbconfig/20260804-194911-ladsgroup.json * 19:46 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 19:45 jhancock@cumin2002: START - Cookbook sre.hosts.move-vlan for host mc2046 * 19:45 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS bookworm * 19:38 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203', diff saved to https://phabricator.wikimedia.org/P95875 and previous config saved to /var/cache/conftool/dbconfig/20260804-193825-ladsgroup.json * 19:27 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95874 and previous config saved to /var/cache/conftool/dbconfig/20260804-192738-ladsgroup.json * 19:02 mutante: gerrit ssh -p 29418 gerrit.wikimedia.org gerrit index changes {{Gerrit|1320979}} * 18:20 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 18:18 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 18:14 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 18:14 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 18:13 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 18:10 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 18:08 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 18:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95872 and previous config saved to /var/cache/conftool/dbconfig/20260804-180721-ladsgroup.json * 18:07 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 18:06 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1203.eqiad.wmnet with reason: Maintenance * 18:06 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95871 and previous config saved to /var/cache/conftool/dbconfig/20260804-180618-ladsgroup.json * 17:55 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179', diff saved to https://phabricator.wikimedia.org/P95870 and previous config saved to /var/cache/conftool/dbconfig/20260804-175531-ladsgroup.json * 17:55 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1154.eqiad.wmnet * 17:55 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1154.eqiad.wmnet * 17:55 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1154.eqiad.wmnet * 17:50 swfrench@deploy1003: Finished scap sync-world: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] (duration: 04m 05s) * 17:48 swfrench@deploy1003: swfrench: Continuing with deployment * 17:46 swfrench@deploy1003: swfrench: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:45 swfrench@deploy1003: Started scap sync-world: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] * 17:44 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179', diff saved to https://phabricator.wikimedia.org/P95869 and previous config saved to /var/cache/conftool/dbconfig/20260804-174445-ladsgroup.json * 17:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95868 and previous config saved to /var/cache/conftool/dbconfig/20260804-173359-ladsgroup.json * 17:33 swfrench@deploy1003: Finished scap sync-world: Pick up new PHP production image (duration: 28m 32s) * 17:28 aokoth@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on phab1005.eqiad.wmnet with reason: Puppet Failure * 17:05 swfrench@deploy1003: Started scap sync-world: Pick up new PHP production image * 17:00 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 17:00 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 16:54 cgoubert@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on wikikube-worker2187.codfw.wmnet with reason: Hardware issue * 16:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2187.codfw.wmnet * 16:52 mutante: gerrit2003:/var/log/apache2# ln -s /srv/gerrit/site_path/review_site/logs/ gerrit * 16:52 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2187.codfw.wmnet * 16:48 mutante: gerrit2003 - moving old apache logfiles older than 60 days from /var/log/apache2 to /srv/gerrit/site_path/review_site/logs/old/ * 16:33 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 16:32 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 16:29 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 16:29 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 16:28 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 16:28 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 16:27 dzahn@cumin1003: END (PASS) - Cookbook sre.gerrit.restart-gerrit (exit_code=0) Restarting Gerrit on gerrit2003 * 16:27 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 16:27 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95867 and previous config saved to /var/cache/conftool/dbconfig/20260804-162736-ladsgroup.json * 16:27 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 16:26 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1179.eqiad.wmnet with reason: Maintenance * 16:26 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:25 mutante: restarting gerrit - dropped outdated RSA host key * 16:25 dzahn@cumin1003: START - Cookbook sre.gerrit.restart-gerrit Restarting Gerrit on gerrit2003 * 16:24 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95866 and previous config saved to /var/cache/conftool/dbconfig/20260804-162424-ladsgroup.json * 16:24 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 16:23 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 16:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95865 and previous config saved to /var/cache/conftool/dbconfig/20260804-162236-ladsgroup.json * 16:21 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1179.eqiad.wmnet with reason: Maintenance * 16:17 swfrench-wmf: reprepro include php8.3_8.3.33-1+wmf11u1 into component/php83 for bullseye-wikimedia * 16:17 swfrench-wmf: reprepro include php8.3_8.3.33-1+wmf12u1 into component/php83 for bookworm-wikimedia * 16:11 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply * 16:10 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply * 16:10 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mobileapps: apply * 16:09 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mobileapps: apply * 16:09 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply * 16:08 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply * 16:08 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:08 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:07 aokoth@cumin1003: END (PASS) - Cookbook sre.vrts.upgrade (exit_code=0) on VRTS host vrts1003.eqiad.wmnet * 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:05 aokoth@cumin1003: START - Cookbook sre.vrts.upgrade on VRTS host vrts1003.eqiad.wmnet * 16:04 mutante: gerrit2002/gerrit1003/gerrit2003 - rm /etc/gerrit/ssh_host_rsa_key * 15:59 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:59 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:59 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:59 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:56 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 15:55 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:55 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:55 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:49 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 15:49 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:44 Raine: add php8.5 packages to component/php85 - [[phab:T432983|T432983]] * 15:39 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:33 aaron@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 15:33 aaron@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 15:29 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:19 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:19 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:16 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:16 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1154.eqiad.wmnet with OS trixie * 15:16 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:15 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:15 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:06 brennen@deploy1003: Finished deploy [phabricator/deployment@56f4ffd]: deploy phab1004 for [[phab:T433981|T433981]] (duration: 00m 43s) * 15:05 brennen@deploy1003: Started deploy [phabricator/deployment@56f4ffd]: deploy phab1004 for [[phab:T433981|T433981]] * 15:05 aaron@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 15:04 aaron@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 15:02 brennen@deploy1003: Finished deploy [phabricator/deployment@56f4ffd]: deploy phab2003 for [[phab:T433981|T433981]] (duration: 00m 51s) * 15:01 brennen@deploy1003: Started deploy [phabricator/deployment@56f4ffd]: deploy phab2003 for [[phab:T433981|T433981]] * 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1004.eqiad.wmnet with reason: deployment * 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1005.eqiad.wmnet with reason: deployment * 14:58 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab2003.codfw.wmnet with reason: deployment * 14:55 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1154.eqiad.wmnet with reason: host reimage * 14:51 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1154.eqiad.wmnet with reason: host reimage * 14:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 14:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 14:38 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync * 14:38 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync * 14:38 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync * 14:37 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync * 14:37 ottomata: roll restart eventgate-main to pick up stream config change - [[phab:T433507|T433507]] * 14:37 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-main: sync * 14:36 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-main: sync * 14:36 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1154 * 14:36 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1154 * 14:34 otto@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] (duration: 08m 39s) * 14:34 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1154 * 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1154.eqiad.wmnet 108.32.64.10.in-addr.arpa 8.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:34 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1154.eqiad.wmnet 108.32.64.10.in-addr.arpa 8.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1154 - jayme@cumin1003" * 14:34 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1154 - jayme@cumin1003" * 14:30 otto@deploy1003: otto: Continuing with deployment * 14:30 jayme@cumin1003: START - Cookbook sre.dns.netbox * 14:28 otto@deploy1003: otto: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:26 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1154 * 14:26 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1154.eqiad.wmnet with OS trixie * 14:26 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1154.eqiad.wmnet * 14:26 otto@deploy1003: Started scap sync-world: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] * 14:26 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1154.eqiad.wmnet * 14:26 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1154.eqiad.wmnet * 14:17 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 14:16 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 14:15 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 14:14 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 14:13 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 14:13 swfrench@dns1004: END - running authdns-update * 14:13 Msz2001: Finished deployments for UTC afternoon backport window * 14:13 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 14:13 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] (duration: 07m 58s) * 14:11 swfrench@dns1004: START - running authdns-update * 14:08 mszwarc@deploy1003: javiermonton, mszwarc, mpostoronca: Continuing with deployment * 14:07 mszwarc@deploy1003: javiermonton, mszwarc, mpostoronca: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] synced to the testser * 14:05 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] * 14:03 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 13:49 swfrench@cumin2002: conftool action : set/pooled=yes; selector: name=wikikube-worker2330.codfw.wmnet * 13:49 swfrench@cumin2002: conftool action : set/pooled=no; selector: name=wikikube-worker2330.codfw.wmnet * 13:48 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] (duration: 09m 19s) * 13:45 swfrench@dns1004: END - running authdns-update * 13:44 mszwarc@deploy1003: mszwarc, jforrester: Continuing with deployment * 13:43 swfrench@dns1004: START - running authdns-update * 13:41 mszwarc@deploy1003: mszwarc, jforrester: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:38 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] * 13:33 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 13:33 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1154.eqiad.wmnet * 13:32 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 13:32 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 13:31 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 13:31 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:31 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:29 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1154.eqiad.wmnet * 13:28 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1154.eqiad.wmnet * 13:28 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1154.eqiad.wmnet * 13:28 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1141.eqiad.wmnet * 13:28 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1141.eqiad.wmnet * 13:28 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1141.eqiad.wmnet * 13:22 otto@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply * 13:22 otto@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply * 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1096.eqiad.wmnet with OS trixie * 13:05 swfrench@dns1004: END - running authdns-update * 13:03 swfrench@dns1004: START - running authdns-update * 12:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 12:43 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 1:00:00 on db1171.eqiad.wmnet with reason: decom * 12:42 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 1:00:00 on db1150.eqiad.wmnet with reason: decom * 12:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 12:38 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1164,1217].eqiad.wmnet with reason: cloning * 12:33 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2096.codfw.wmnet with OS trixie * 12:22 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1096.eqiad.wmnet with OS trixie * 12:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2096.codfw.wmnet with reason: host reimage * 12:14 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1141.eqiad.wmnet with OS trixie * 12:10 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2096.codfw.wmnet with reason: host reimage * 12:10 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1289.eqiad.wmnet * 12:05 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1289.eqiad.wmnet * 12:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1288.eqiad.wmnet * 11:59 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1288.eqiad.wmnet * 11:59 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1287.eqiad.wmnet * 11:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1097.eqiad.wmnet with OS trixie * 11:54 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1287.eqiad.wmnet * 11:54 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1286.eqiad.wmnet * 11:53 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1141.eqiad.wmnet with reason: host reimage * 11:51 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2096.codfw.wmnet with OS trixie * 11:49 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1141.eqiad.wmnet with reason: host reimage * 11:48 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1286.eqiad.wmnet * 11:48 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1284.eqiad.wmnet * 11:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2095.codfw.wmnet with OS trixie * 11:43 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1284.eqiad.wmnet * 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1283.eqiad.wmnet * 11:42 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on ml-serve1015.eqiad.wmnet with reason: Downtime to get full picture of current BIOS settings beyond what Redfish shows * 11:39 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad * 11:39 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:37 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1283.eqiad.wmnet * 11:37 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1282.eqiad.wmnet * 11:37 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad * 11:37 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:33 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1141 * 11:33 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1141 * 11:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 11:32 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1141 * 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1141.eqiad.wmnet 156.48.64.10.in-addr.arpa 6.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:32 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1141.eqiad.wmnet 156.48.64.10.in-addr.arpa 6.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1141 - jayme@cumin1003" * 11:32 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1141 - jayme@cumin1003" * 11:32 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1282.eqiad.wmnet * 11:32 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1281.eqiad.wmnet * 11:32 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad * 11:32 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:29 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 11:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1097.eqiad.wmnet with reason: host reimage * 11:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2095.codfw.wmnet with OS trixie * 11:26 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1281.eqiad.wmnet * 11:26 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1280.eqiad.wmnet * 11:25 jayme@cumin1003: START - Cookbook sre.dns.netbox * 11:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1097.eqiad.wmnet with reason: host reimage * 11:22 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1141 * 11:21 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1141.eqiad.wmnet with OS trixie * 11:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1280.eqiad.wmnet * 11:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1279.eqiad.wmnet * 11:20 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin with reason: upgrade new Nokia swtiches in eqsin to SR Linux v26 * 11:17 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1141.eqiad.wmnet * 11:16 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1141.eqiad.wmnet * 11:16 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1141.eqiad.wmnet * 11:16 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be2095.codfw.wmnet with OS trixie * 11:15 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1279.eqiad.wmnet * 11:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1278.eqiad.wmnet * 11:14 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1139.eqiad.wmnet * 11:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1139.eqiad.wmnet * 11:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1070.eqiad.wmnet with OS trixie * 11:13 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1140.eqiad.wmnet * 11:13 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1140.eqiad.wmnet * 11:13 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1140.eqiad.wmnet * 11:09 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1278.eqiad.wmnet * 11:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1071.eqiad.wmnet with OS trixie * 11:05 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1097.eqiad.wmnet with OS trixie * 11:04 mvernon@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be1097.eqiad.wmnet with OS trixie * 11:02 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1140.eqiad.wmnet with OS trixie * 11:02 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1097.eqiad.wmnet with OS trixie * 11:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1096.eqiad.wmnet with OS trixie * 11:00 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1139.eqiad.wmnet * 11:00 jayme@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1139.eqiad.wmnet with OS trixie * 10:56 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1069.eqiad.wmnet with OS trixie * 10:56 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 10:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1070.eqiad.wmnet with reason: host reimage * 10:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 10:45 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1071.eqiad.wmnet with reason: host reimage * 10:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 10:41 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1096.eqiad.wmnet with OS trixie * 10:41 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1140.eqiad.wmnet with reason: host reimage * 10:39 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1071.eqiad.wmnet with reason: host reimage * 10:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1070.eqiad.wmnet with reason: host reimage * 10:38 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1139.eqiad.wmnet with reason: host reimage * 10:37 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1140.eqiad.wmnet with reason: host reimage * 10:35 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1069.eqiad.wmnet with reason: host reimage * 10:33 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1095.eqiad.wmnet with OS trixie * 10:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 10:33 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1139.eqiad.wmnet with reason: host reimage * 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1069.eqiad.wmnet with reason: host reimage * 10:23 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1140 * 10:23 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1140 * 10:23 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 10:22 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1140 * 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1140.eqiad.wmnet 155.48.64.10.in-addr.arpa 5.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:21 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1071.eqiad.wmnet with OS trixie * 10:21 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1140.eqiad.wmnet 155.48.64.10.in-addr.arpa 5.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1140 - jayme@cumin1003" * 10:21 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1140 - jayme@cumin1003" * 10:21 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1071 * 10:21 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1070.eqiad.wmnet with OS trixie * 10:21 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1070 * 10:20 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 10:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1095.eqiad.wmnet with OS trixie * 10:17 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1139 * 10:17 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1139 * 10:17 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be1095.eqiad.wmnet with OS trixie * 10:15 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1139 * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1139.eqiad.wmnet 194.32.64.10.in-addr.arpa 4.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:15 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1139.eqiad.wmnet 194.32.64.10.in-addr.arpa 4.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1139 - jayme@cumin1003" * 10:15 jayme@cumin1003: START - Cookbook sre.dns.netbox * 10:15 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1139 - jayme@cumin1003" * 10:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2095.codfw.wmnet with OS trixie * 10:13 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1069.eqiad.wmnet with OS trixie * 10:12 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1069 * 10:11 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1140 * 10:11 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1140.eqiad.wmnet with OS trixie * 10:11 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1140.eqiad.wmnet * 10:10 jayme@cumin1003: START - Cookbook sre.dns.netbox * 10:10 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1139 * 10:10 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1140.eqiad.wmnet * 10:10 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1140.eqiad.wmnet * 10:10 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1139.eqiad.wmnet with OS trixie * 10:09 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1139.eqiad.wmnet * 10:08 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1139.eqiad.wmnet * 10:08 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1139.eqiad.wmnet * 10:01 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2094.codfw.wmnet with OS trixie * 09:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 09:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 09:53 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 09:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 09:44 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:44 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2094.codfw.wmnet with reason: host reimage * 09:34 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2094.codfw.wmnet with reason: host reimage * 09:34 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1071 * 09:33 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1070 * 09:33 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1095.eqiad.wmnet with OS trixie * 09:27 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1069 * 09:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:23 brouberol@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM archiva1002.wikimedia.org * 09:20 brouberol@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM archiva1002.wikimedia.org * 09:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1277.eqiad.wmnet * 09:13 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2094.codfw.wmnet with OS trixie * 09:13 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 09:12 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 09:12 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:12 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1277.eqiad.wmnet * 09:12 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1276.eqiad.wmnet * 09:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1094.eqiad.wmnet with OS trixie * 09:06 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1276.eqiad.wmnet * 09:06 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1275.eqiad.wmnet * 09:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2093.codfw.wmnet with OS trixie * 09:01 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1275.eqiad.wmnet * 09:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1274.eqiad.wmnet * 08:56 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1274.eqiad.wmnet * 08:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1273.eqiad.wmnet * 08:50 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1273.eqiad.wmnet * 08:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1272.eqiad.wmnet * 08:49 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:49 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1094.eqiad.wmnet with reason: host reimage * 08:45 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1272.eqiad.wmnet * 08:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1094.eqiad.wmnet with reason: host reimage * 08:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2093.codfw.wmnet with reason: host reimage * 08:38 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:38 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2093.codfw.wmnet with reason: host reimage * 08:35 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 08:34 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 08:29 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:28 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:26 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1271.eqiad.wmnet * 08:23 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1094.eqiad.wmnet with OS trixie * 08:21 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 08:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1271.eqiad.wmnet * 08:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1270.eqiad.wmnet * 08:15 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1270.eqiad.wmnet * 08:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1269.eqiad.wmnet * 08:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2093.codfw.wmnet with OS trixie * 08:09 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1269.eqiad.wmnet * 08:09 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1268.eqiad.wmnet * 08:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2092.codfw.wmnet with OS trixie * 08:04 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1268.eqiad.wmnet * 08:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1267.eqiad.wmnet * 07:59 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1267.eqiad.wmnet * 07:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1093.eqiad.wmnet with OS trixie * 07:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1266.eqiad.wmnet * 07:56 jynus: running extra backups to test db1285 [[phab:T433826|T433826]] * 07:51 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1266.eqiad.wmnet * 07:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2092.codfw.wmnet with reason: host reimage * 07:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1093.eqiad.wmnet with reason: host reimage * 07:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2092.codfw.wmnet with reason: host reimage * 07:32 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1093.eqiad.wmnet with reason: host reimage * 07:29 jynus: running extra backups to test db1265 [[phab:T433825|T433825]] * 07:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2092.codfw.wmnet with OS trixie * 07:11 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1093.eqiad.wmnet with OS trixie * 06:50 slyngshede@dns1004: END - running authdns-update * 06:48 slyngshede@dns1004: START - running authdns-update * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.11 (duration: 02m 29s) * 03:38 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] (duration: 32m 57s) * 03:23 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 03:22 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 03:05 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 32s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 00:45 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] (duration: 06m 20s) * 00:41 cjming@deploy1003: cjming: Continuing with deployment * 00:41 cjming@deploy1003: cjming: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:39 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] == 2026-08-03 == * 23:58 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply * 23:57 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply * 23:29 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cp5021.eqsin.wmnet * 23:29 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cp5021.eqsin.wmnet * 23:27 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cp5021.eqsin.wmnet * 23:26 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cp5021.eqsin.wmnet * 23:18 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 23:17 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 22:56 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: sync * 22:56 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: sync * 22:36 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 22:36 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 22:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host search-loader2002.codfw.wmnet with OS trixie * 21:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on search-loader2002.codfw.wmnet with reason: host reimage * 21:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on search-loader2002.codfw.wmnet with reason: host reimage * 21:42 dancy@deploy1003: Stopping before sync operations * 21:41 dancy@deploy1003: Started scap sync-world: testing * 21:39 dancy@deploy1003: Installation of scap version "4.277.0" completed for 3 hosts * 21:37 dancy@deploy1003: Installing scap version "4.277.0" for 3 host(s) * 21:37 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] (duration: 06m 13s) * 21:33 dancy@deploy1003: dancy: Continuing with deployment * 21:32 dancy@deploy1003: dancy: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:31 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] * 21:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host search-loader2002.codfw.wmnet with OS trixie * 21:03 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] (duration: 06m 34s) * 20:59 dancy@deploy1003: dancy: Continuing with deployment * 20:58 dancy@deploy1003: dancy: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:56 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] * 20:52 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] (duration: 06m 23s) * 20:48 cjming@deploy1003: cjming: Continuing with deployment * 20:47 cjming@deploy1003: cjming: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:46 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] * 20:42 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] (duration: 07m 36s) * 20:38 arlolra@deploy1003: arlolra: Continuing with deployment * 20:36 arlolra@deploy1003: arlolra: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:34 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] * 20:16 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] (duration: 08m 26s) * 20:12 krinkle@deploy1003: krinkle: Continuing with deployment * 20:09 krinkle@deploy1003: krinkle: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] * 19:45 jasmine@cumin2002: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-main-codfw * 18:58 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] (duration: 09m 23s) * 18:53 krinkle@deploy1003: krinkle: Continuing with deployment * 18:53 jasmine@cumin2002: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-main-codfw * 18:50 krinkle@deploy1003: krinkle: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:48 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] * 18:37 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] (duration: 10m 13s) * 18:34 dzahn@cumin2002: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host codesearch2001.codfw.wmnet * 18:34 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host codesearch2001.codfw.wmnet with OS trixie * 18:33 krinkle@deploy1003: krinkle: Continuing with deployment * 18:29 krinkle@deploy1003: krinkle: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:27 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] * 18:19 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 18:18 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on codesearch2001.codfw.wmnet with reason: host reimage * 18:14 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 18:14 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:12 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on codesearch2001.codfw.wmnet with reason: host reimage * 18:11 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 18:11 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 18:10 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 18:02 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 18:02 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 18:01 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 18:01 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 17:55 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host codesearch2001.codfw.wmnet with OS trixie * 17:54 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:54 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) codesearch2001.codfw.wmnet on all recursors * 17:53 dzahn@cumin2002: START - Cookbook sre.dns.wipe-cache codesearch2001.codfw.wmnet on all recursors * 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:48 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:41 dzahn@cumin2002: START - Cookbook sre.dns.netbox * 17:41 dzahn@cumin2002: START - Cookbook sre.ganeti.makevm for new host codesearch2001.codfw.wmnet * 17:37 dzahn@cumin2002: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host codesearch1001.eqiad.wmnet * 17:37 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host codesearch1001.eqiad.wmnet with OS trixie * 17:24 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on codesearch1001.eqiad.wmnet with reason: host reimage * 17:17 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on codesearch1001.eqiad.wmnet with reason: host reimage * 17:08 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host codesearch1001.eqiad.wmnet with OS trixie * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 17:06 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:06 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) codesearch1001.eqiad.wmnet on all recursors * 17:06 dzahn@cumin2002: START - Cookbook sre.dns.wipe-cache codesearch1001.eqiad.wmnet on all recursors * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 17:05 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 17:04 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:04 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 16:58 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 16:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2091.codfw.wmnet with OS trixie * 16:54 ebernhardson@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 16:54 ebernhardson@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 16:49 ebernhardson@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 16:49 ebernhardson@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 16:46 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1092.eqiad.wmnet with OS trixie * 16:43 ebernhardson@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 16:43 ebernhardson@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 16:43 dzahn@cumin2002: START - Cookbook sre.dns.netbox * 16:43 dzahn@cumin2002: START - Cookbook sre.ganeti.makevm for new host codesearch1001.eqiad.wmnet * 16:41 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 16:41 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 16:40 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2091.codfw.wmnet with reason: host reimage * 16:37 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 16:35 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2091.codfw.wmnet with reason: host reimage * 16:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1092.eqiad.wmnet with reason: host reimage * 16:24 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 16:24 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 16:23 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1092.eqiad.wmnet with reason: host reimage * 16:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2091.codfw.wmnet with OS trixie * 16:03 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1092.eqiad.wmnet with OS trixie * 16:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2090.codfw.wmnet with OS trixie * 15:51 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 15:51 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 15:51 jiji@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 15:50 jiji@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 15:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2090.codfw.wmnet with reason: host reimage * 15:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2090.codfw.wmnet with reason: host reimage * 15:33 jhathaway@dns1004: END - running authdns-update * 15:31 jhathaway@dns1004: START - running authdns-update * 15:26 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1091.eqiad.wmnet with OS trixie * 15:25 dancy@deploy1003: Installation of scap version "4.276.1" completed for 3 hosts * 15:23 dancy@deploy1003: Installing scap version "4.276.1" for 3 host(s) * 15:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2090.codfw.wmnet with OS trixie * 15:12 marostegui@cumin1003: dbctl commit (dc=all): 'Repool db2245, db2246, db2247 and db2248 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95857 and previous config saved to /var/cache/conftool/dbconfig/20260803-151212-marostegui.json * 15:09 dancy@deploy1003: Started scap sync-world: testing * 15:09 dancy@deploy1003: Installation of scap version "4.277.0" completed for 3 hosts * 15:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1091.eqiad.wmnet with reason: host reimage * 15:07 dancy@deploy1003: Installing scap version "4.277.0" for 3 host(s) * 15:03 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1091.eqiad.wmnet with reason: host reimage * 14:49 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1091.eqiad.wmnet with OS trixie * 14:33 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2089.codfw.wmnet with OS trixie * 14:29 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 14:27 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 14:18 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 14:16 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 14:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2089.codfw.wmnet with reason: host reimage * 14:10 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 14:10 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 14:09 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2089.codfw.wmnet with reason: host reimage * 13:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2089.codfw.wmnet with OS trixie * 13:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2088.codfw.wmnet with OS trixie * 13:40 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1090.eqiad.wmnet with OS trixie * 13:22 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1090.eqiad.wmnet with reason: host reimage * 13:22 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] (duration: 14m 34s) * 13:19 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1090.eqiad.wmnet with reason: host reimage * 13:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2088.codfw.wmnet with reason: host reimage * 13:16 aude@deploy1003: aude, mhorsey: Continuing with deployment * 13:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2088.codfw.wmnet with reason: host reimage * 13:12 aude@deploy1003: aude, mhorsey: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] * 13:05 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1090.eqiad.wmnet with OS trixie * 12:58 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2088.codfw.wmnet with OS trixie * 12:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db[2245-2247].codfw.wmnet * 12:49 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2247: Rebooting db2247.codfw.wmnet * 12:49 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2247: Rebooting db2247.codfw.wmnet * 12:42 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2246: Rebooting db2246.codfw.wmnet * 12:42 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2246: Rebooting db2246.codfw.wmnet * 12:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2087.codfw.wmnet with OS trixie * 12:37 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1089.eqiad.wmnet with OS trixie * 12:34 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2245: Rebooting db2245.codfw.wmnet * 12:34 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2245: Rebooting db2245.codfw.wmnet * 12:34 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db[2245-2247].codfw.wmnet * 12:32 kamila@deploy1003: Finished scap sync-world: rebuild after base image update (duration: 30m 26s) * 12:28 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 12:22 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2087.codfw.wmnet with reason: host reimage * 12:19 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1089.eqiad.wmnet with reason: host reimage * 12:14 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2087.codfw.wmnet with reason: host reimage * 12:14 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1089.eqiad.wmnet with reason: host reimage * 12:03 kamila@deploy1003: Started scap sync-world: rebuild after base image update * 12:00 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1089.eqiad.wmnet with OS trixie * 12:00 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2087.codfw.wmnet with OS trixie * 11:35 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db[2245-2248].codfw.wmnet * 11:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db[2245-2248].codfw.wmnet * 11:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2086.codfw.wmnet with OS trixie * 11:26 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 11:26 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 11:25 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 11:25 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 11:24 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1088.eqiad.wmnet with OS trixie * 11:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db[2245-2248].codfw.wmnet with reason: Checking network * 11:21 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 11:20 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 11:19 marostegui@dns1004: END - running authdns-update * 11:17 marostegui@dns1004: START - running authdns-update * 11:10 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:10 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 11:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2086.codfw.wmnet with reason: host reimage * 11:09 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:08 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 11:08 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:07 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 11:07 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:07 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 11:06 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop: apply * 11:06 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop: apply * 11:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1088.eqiad.wmnet with reason: host reimage * 11:05 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop: apply * 11:04 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop: apply * 11:04 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop: apply * 11:04 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop: apply * 11:02 marostegui@dns1004: END - running authdns-update * 11:02 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2086.codfw.wmnet with reason: host reimage * 11:01 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1088.eqiad.wmnet with reason: host reimage * 11:00 marostegui@dns1004: START - running authdns-update * 10:53 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] (duration: 10m 57s) * 10:51 cmooney@dns3003: END - running authdns-update * 10:49 cmooney@dns3003: START - running authdns-update * 10:47 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1088.eqiad.wmnet with OS trixie * 10:47 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2086.codfw.wmnet with OS trixie * 10:47 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 10:46 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:46 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:46 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new reverse ranges for eqsin CR switch links - cmooney@cumin1003" * 10:46 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new reverse ranges for eqsin CR switch links - cmooney@cumin1003" * 10:42 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] * 10:41 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 10:36 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2245, db2246 and db2247 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95855 and previous config saved to /var/cache/conftool/dbconfig/20260803-103652-marostegui.json * 10:35 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2248 from s4 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95854 and previous config saved to /var/cache/conftool/dbconfig/20260803-103535-marostegui.json * 10:27 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 10:27 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 10:26 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 10:24 kart_: cxserver: Add referencePunctuation config ([[phab:T97231|T97231]]) * 10:24 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 10:23 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:23 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:23 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:22 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:22 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply * 10:21 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply * 10:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2085.codfw.wmnet with OS trixie * 10:20 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply * 10:20 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply * 10:18 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply * 10:18 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply * 10:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1087.eqiad.wmnet with OS trixie * 09:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2085.codfw.wmnet with reason: host reimage * 09:43 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1087.eqiad.wmnet with reason: host reimage * 09:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2085.codfw.wmnet with reason: host reimage * 09:40 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1087.eqiad.wmnet with reason: host reimage * 09:26 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1087.eqiad.wmnet with OS trixie * 09:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2085.codfw.wmnet with OS trixie * 09:13 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2084.codfw.wmnet with OS trixie * 09:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1086.eqiad.wmnet with OS trixie * 08:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2084.codfw.wmnet with reason: host reimage * 08:50 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2084.codfw.wmnet with reason: host reimage * 08:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1086.eqiad.wmnet with reason: host reimage * 08:39 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1086.eqiad.wmnet with reason: host reimage * 08:38 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:38 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:37 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 08:37 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 08:35 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2084.codfw.wmnet with OS trixie * 08:34 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 08:34 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:27 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1086.eqiad.wmnet with OS trixie * 08:09 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1218: Repool after a crash * 08:07 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2083.codfw.wmnet with OS trixie * 08:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1085.eqiad.wmnet with OS trixie * 07:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2083.codfw.wmnet with reason: host reimage * 07:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1085.eqiad.wmnet with reason: host reimage * 07:40 kart_: Updated cxsever to 2026-07-16-140518-production ([[phab:T97231|T97231]]) * 07:39 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply * 07:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2083.codfw.wmnet with reason: host reimage * 07:38 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply * 07:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1085.eqiad.wmnet with reason: host reimage * 07:37 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] (duration: 32m 40s) * 07:33 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply * 07:33 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply * 07:25 jdlrobson@deploy1003: jdlrobson: Continuing with deployment * 07:24 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2083.codfw.wmnet with OS trixie * 07:24 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1085.eqiad.wmnet with OS trixie * 07:23 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1218: Repool after a crash * 07:21 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:09 marostegui: Drop renamed tables [[phab:T425074|T425074]] * 07:04 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] * 06:55 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply * 06:54 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 46s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-02 == * 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 01m 03s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-01 == * 03:30 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:30 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:30 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:30 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 34s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-31 == * 17:41 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 17:41 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 17:40 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 17:40 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 15:33 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2195: Testing * 15:02 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:02 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 15:02 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 14:48 pt1979@cumin2002: START - Cookbook sre.dns.netbox * 14:47 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2195: Testing * 14:22 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 14:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2195: Testing * 14:21 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 14:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2195.codfw.wmnet with reason: Testing * 14:16 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 14:04 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 14:04 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 13:30 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1048.eqiad.wmnet with OS trixie * 13:22 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2195: Testing * 13:22 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 13:19 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2195: Testing * 13:18 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 13:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2195: Testing * 13:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 13:05 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 13:05 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 13:04 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 13:04 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 12:50 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lswtest-d8-eqiad * 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:53 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:42 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 6515 * 11:37 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 6515 * 11:28 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:27 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:07 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2082.codfw.wmnet with OS trixie * 10:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2082.codfw.wmnet with reason: host reimage * 10:42 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2082.codfw.wmnet with reason: host reimage * 10:28 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2082.codfw.wmnet with OS trixie * 10:02 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 09:52 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 09:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts * 09:16 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts * 08:57 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 08:46 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:42 gkyziridis@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 08:37 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 08:37 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 08:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 08:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 08:11 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:11 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:08 filippo@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudvirt1048 * 08:07 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 08:07 filippo@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudvirt1048 * 08:06 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 08:01 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 08:00 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 07:19 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1048.eqiad.wmnet with reason: host reimage * 07:13 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1048.eqiad.wmnet with reason: host reimage * 07:11 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 07:11 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 07:09 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 07:09 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 06:57 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:56 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:48 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:48 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:44 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie * 06:34 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1048.eqiad.wmnet with OS trixie * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 54s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 00:57 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] (duration: 11m 04s) * 00:53 dreamyjazz@deploy1003: dreamyjazz, jforrester: Continuing with deployment * 00:48 dreamyjazz@deploy1003: dreamyjazz, jforrester: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:46 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] == 2026-07-30 == * 21:37 dancy@deploy1003: Installation of scap version "4.276.1" completed for 3 hosts * 21:35 dancy@deploy1003: Installing scap version "4.276.1" for 3 host(s) * 21:24 dancy@deploy1003: Installation of scap version "4.276.0" completed for 3 hosts * 21:22 dancy@deploy1003: Installing scap version "4.276.0" for 3 host(s) * 21:15 maryum: Deployed security fix for [[phab:T430601|T430601]] * 20:13 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] (duration: 09m 20s) * 20:07 arlolra@deploy1003: osleger, arlolra: Continuing with deployment * 20:05 arlolra@deploy1003: osleger, arlolra: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:03 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] * 19:29 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:29 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:25 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service * 19:24 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 19:24 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:24 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:24 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 19:23 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service * 19:20 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1084.eqiad.wmnet with OS trixie * 18:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1084.eqiad.wmnet with reason: host reimage * 18:52 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1084.eqiad.wmnet with reason: host reimage * 18:41 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 18:40 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 18:39 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1084.eqiad.wmnet with OS trixie * 18:25 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 18:15 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 17:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1083.eqiad.wmnet with OS trixie * 17:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1048: Maintenance * 17:36 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new security plugin settings - bking@cumin2003 - [[phab:T350516|T350516]] * 17:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1083.eqiad.wmnet with reason: host reimage * 17:26 inflatador: bking@apt1002 `reprepro --noskipold --component thirdparty/opensearch3 update trixie-wikimedia` [[phab:T433624|T433624]] * 17:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1083.eqiad.wmnet with reason: host reimage * 17:23 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 17:20 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 17:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2081.codfw.wmnet with OS trixie * 17:11 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new security plugin settings - bking@cumin2003 - [[phab:T350516|T350516]] * 17:10 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1083.eqiad.wmnet with OS trixie * 16:55 root@cumin1003: START - Cookbook sre.mysql.pool pool es1048: Maintenance * 16:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2081.codfw.wmnet with reason: host reimage * 16:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1048 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95833 and previous config saved to /var/cache/conftool/dbconfig/20260730-165053-cwilliams.json * 16:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1048.eqiad.wmnet with reason: Maintenance * 16:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1040: Maintenance * 16:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2081.codfw.wmnet with reason: host reimage * 16:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1082.eqiad.wmnet with OS trixie * 16:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2081.codfw.wmnet with OS trixie * 16:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2097.codfw.wmnet with OS trixie * 16:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1082.eqiad.wmnet with reason: host reimage * 16:08 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1082.eqiad.wmnet with reason: host reimage * 16:08 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 16:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1047: Maintenance * 16:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2080.codfw.wmnet with OS trixie * 16:04 root@cumin1003: START - Cookbook sre.mysql.pool pool es1040: Maintenance * 16:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1040: Maintenance * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new logging settings - bking@cumin2003 - [[phab:T324335|T324335]] * 15:58 root@cumin1003: START - Cookbook sre.mysql.pool pool es1040: Maintenance * 15:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1040 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95827 and previous config saved to /var/cache/conftool/dbconfig/20260730-155324-cwilliams.json * 15:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1040.eqiad.wmnet with reason: Maintenance * 15:50 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1082.eqiad.wmnet with OS trixie * 15:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2048: Maintenance * 15:44 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 24s) * 15:43 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2080.codfw.wmnet with reason: host reimage * 15:38 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new logging settings - bking@cumin2003 - [[phab:T324335|T324335]] * 15:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2080.codfw.wmnet with reason: host reimage * 15:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 15:30 mvernon@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be2097.codfw.wmnet with OS trixie * 15:23 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1081.eqiad.wmnet with OS trixie * 15:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2098.codfw.wmnet with OS trixie * 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - mvernon@cumin2003" * 15:18 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be2097.codfw.wmnet with OS trixie * 15:18 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - mvernon@cumin2003" * 15:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 15:17 root@cumin1003: START - Cookbook sre.mysql.pool pool es1047: Maintenance * 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2080.codfw.wmnet with OS trixie * 15:13 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be2097.codfw.wmnet with OS trixie * 15:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1047 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95820 and previous config saved to /var/cache/conftool/dbconfig/20260730-151200-cwilliams.json * 15:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1047.eqiad.wmnet with reason: Maintenance * 15:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: Maintenance * 15:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 15:04 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 15:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1081.eqiad.wmnet with reason: host reimage * 15:00 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1081.eqiad.wmnet with reason: host reimage * 15:00 root@cumin1003: START - Cookbook sre.mysql.pool pool es2048: Maintenance * 15:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 14:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2079.codfw.wmnet with OS trixie * 14:56 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 14:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2048 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95816 and previous config saved to /var/cache/conftool/dbconfig/20260730-145510-cwilliams.json * 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2048.codfw.wmnet with reason: Maintenance * 14:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2040: Maintenance * 14:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 14:51 tchin@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] (duration: 06m 48s) * 14:47 tchin@deploy1003: jforrester, tchin: Continuing with deployment * 14:47 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 14:47 tchin@deploy1003: jforrester, tchin: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:45 tchin@deploy1003: Started scap sync-world: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] * 14:42 sukhe@puppetserver1001: conftool action : set/weight=1; selector: cluster=urldownloader,service=squid * 14:42 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader,service=squid * 14:42 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1081.eqiad.wmnet with OS trixie * 14:39 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 14:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2079.codfw.wmnet with reason: host reimage * 14:36 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2098.codfw.wmnet with OS trixie * 14:32 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2079.codfw.wmnet with reason: host reimage * 14:30 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] (duration: 06m 31s) * 14:27 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 14:26 mszwarc@deploy1003: mszwarc: Continuing with deployment * 14:25 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:25 root@cumin1003: START - Cookbook sre.mysql.pool pool es1038: Maintenance * 14:25 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1038: Maintenance * 14:23 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] * 14:21 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] (duration: 11m 19s) * 14:20 root@cumin1003: START - Cookbook sre.mysql.pool pool es1038: Maintenance * 14:14 stran@deploy1003: stran: Continuing with deployment * 14:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1038 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95810 and previous config saved to /var/cache/conftool/dbconfig/20260730-141439-cwilliams.json * 14:14 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1038.eqiad.wmnet with reason: Maintenance * 14:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1036: Maintenance * 14:13 stran@deploy1003: stran: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2079.codfw.wmnet with OS trixie * 14:09 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] * 14:08 root@cumin1003: START - Cookbook sre.mysql.pool pool es2040: Maintenance * 14:08 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2040: Maintenance * 14:03 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] (duration: 31m 41s) * 14:03 root@cumin1003: START - Cookbook sre.mysql.pool pool es2040: Maintenance * 14:03 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2040 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95806 and previous config saved to /var/cache/conftool/dbconfig/20260730-135643-cwilliams.json * 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2040.codfw.wmnet with reason: Maintenance * 13:56 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:56 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2038: Maintenance * 13:55 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:52 lucaswerkmeister-wmde@deploy1003: migr, lucaswerkmeister-wmde: Continuing with deployment * 13:49 lucaswerkmeister-wmde@deploy1003: migr, lucaswerkmeister-wmde: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:49 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:48 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 13:45 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 13:32 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:32 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] * 13:28 root@cumin1003: START - Cookbook sre.mysql.pool pool es1036: Maintenance * 13:28 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1036: Maintenance * 13:22 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:22 root@cumin1003: START - Cookbook sre.mysql.pool pool es1036: Maintenance * 13:20 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1036 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95800 and previous config saved to /var/cache/conftool/dbconfig/20260730-131727-cwilliams.json * 13:17 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1036.eqiad.wmnet with reason: Maintenance * 13:17 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] (duration: 10m 31s) * 13:16 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2022\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 13:13 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, stran: Continuing with deployment * 13:10 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:10 root@cumin1003: START - Cookbook sre.mysql.pool pool es2038: Maintenance * 13:10 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2038: Maintenance * 13:08 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, stran: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie * 13:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2047: Maintenance * 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048 cloud-private - filippo@cumin1003" * 13:07 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048 cloud-private - filippo@cumin1003" * 13:06 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] * 13:04 root@cumin1003: START - Cookbook sre.mysql.pool pool es2038: Maintenance * 13:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2078.codfw.wmnet with OS trixie * 13:01 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2038 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95797 and previous config saved to /var/cache/conftool/dbconfig/20260730-125919-cwilliams.json * 12:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2038.codfw.wmnet with reason: Maintenance * 12:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2078.codfw.wmnet with reason: host reimage * 12:37 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2078.codfw.wmnet with reason: host reimage * 12:37 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] (duration: 06m 51s) * 12:33 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 12:32 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 12:32 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 12:32 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:30 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] * 12:19 root@cumin1003: START - Cookbook sre.mysql.pool pool es2047: Maintenance * 12:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2078.codfw.wmnet with OS trixie * 12:18 dcausse@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 12:18 dcausse@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 12:15 dcausse@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 12:14 dcausse@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 12:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2047 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95793 and previous config saved to /var/cache/conftool/dbconfig/20260730-121404-cwilliams.json * 12:13 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2047.codfw.wmnet with reason: Maintenance * 12:13 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2036: Maintenance * 12:05 ayounsi@dns1004: END - running authdns-update * 12:02 ayounsi@dns1004: START - running authdns-update * 11:51 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2077.codfw.wmnet with OS trixie * 11:48 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:46 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1080.eqiad.wmnet with OS trixie * 11:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1226: Maintenance * 11:41 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2077.codfw.wmnet with reason: host reimage * 11:28 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2077.codfw.wmnet with reason: host reimage * 11:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1080.eqiad.wmnet with reason: host reimage * 11:27 root@cumin1003: START - Cookbook sre.mysql.pool pool es2036: Maintenance * 11:27 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2036: Maintenance * 11:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1080.eqiad.wmnet with reason: host reimage * 11:21 root@cumin1003: START - Cookbook sre.mysql.pool pool es2036: Maintenance * 11:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2036 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95786 and previous config saved to /var/cache/conftool/dbconfig/20260730-111633-cwilliams.json * 11:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2036.codfw.wmnet with reason: Maintenance * 11:08 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2077.codfw.wmnet with OS trixie * 11:07 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 11:03 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie * 11:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1226: Maintenance * 10:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1226 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95783 and previous config saved to /var/cache/conftool/dbconfig/20260730-104801-cwilliams.json * 10:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1226.eqiad.wmnet with reason: Maintenance * 10:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1214: Maintenance * 10:27 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2035: Maintenance * 10:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2076.codfw.wmnet with OS trixie * 10:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1214: Maintenance * 09:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1214 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95775 and previous config saved to /var/cache/conftool/dbconfig/20260730-095451-cwilliams.json * 09:54 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1214.eqiad.wmnet with reason: Maintenance * 09:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1209: Maintenance * 09:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2076.codfw.wmnet with reason: host reimage * 09:42 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool es2035: Maintenance * 09:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.netbox.update-extras (exit_code=0) rolling restart_daemons on A:netbox * 09:41 ayounsi@cumin1003: START - Cookbook sre.netbox.update-extras rolling restart_daemons on A:netbox * 09:40 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2035: Maintenance * 09:39 ayounsi@cumin1003: END (PASS) - Cookbook sre.netbox.update-extras (exit_code=0) rolling restart_daemons on A:netbox-canary * 09:39 ayounsi@cumin1003: START - Cookbook sre.netbox.update-extras rolling restart_daemons on A:netbox-canary * 09:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2076.codfw.wmnet with reason: host reimage * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 09:34 root@cumin1003: START - Cookbook sre.mysql.pool pool es2035: Maintenance * 09:32 lucaswerkmeister-wmde@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 09:32 lucaswerkmeister-wmde@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 09:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2035 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95771 and previous config saved to /var/cache/conftool/dbconfig/20260730-092910-cwilliams.json * 09:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2035.codfw.wmnet with reason: Maintenance * 09:19 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2076.codfw.wmnet with OS trixie * 09:18 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 09:17 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie * 09:07 root@cumin1003: START - Cookbook sre.mysql.pool pool db1209: Maintenance * 09:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 23 hosts * 09:04 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Remove cable label from interfaces descriptions - ayounsi@cumin1003 * 09:04 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:02 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Remove cable label from interfaces descriptions - ayounsi@cumin1003 * 09:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1209 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95767 and previous config saved to /var/cache/conftool/dbconfig/20260730-090133-cwilliams.json * 09:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1209.eqiad.wmnet with reason: Maintenance * 09:01 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1192: Maintenance * 08:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1252: Maintenance * 08:57 jayme@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 08:56 jayme@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 08:53 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 23 hosts * 08:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:51 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:50 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1263: Maintenance * 08:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:23 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 08:15 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie * 08:14 root@cumin1003: START - Cookbook sre.mysql.pool pool db1192: Maintenance * 08:13 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2075.codfw.wmnet with OS trixie * 08:12 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1252: Maintenance * 08:11 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1252: Maintenance * 08:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1252: Maintenance * 08:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1192 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95754 and previous config saved to /var/cache/conftool/dbconfig/20260730-080611-cwilliams.json * 08:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1192.eqiad.wmnet with reason: Maintenance * 08:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1178: Maintenance * 08:05 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS bullseye * 07:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1263: Maintenance * 07:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1263 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95751 and previous config saved to /var/cache/conftool/dbconfig/20260730-075106-cwilliams.json * 07:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[1260-1262].eqiad.wmnet with reason: Maintenance * 07:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2075.codfw.wmnet with reason: host reimage * 07:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1263.eqiad.wmnet with reason: Maintenance * 07:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2075.codfw.wmnet with reason: host reimage * 07:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance * 07:38 dcausse@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:38 dcausse@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 07:35 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db1252', diff saved to https://phabricator.wikimedia.org/P95748 and previous config saved to /var/cache/conftool/dbconfig/20260730-073510-marostegui.json * 07:26 klausman@dns2004: END - running authdns-update * 07:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 07:25 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2075.codfw.wmnet with OS trixie * 07:24 klausman@dns2004: START - running authdns-update * 07:23 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host an-test-master1003.eqiad.wmnet * 07:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db1178: Maintenance * 07:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1178 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95746 and previous config saved to /var/cache/conftool/dbconfig/20260730-071112-cwilliams.json * 07:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1178.eqiad.wmnet with reason: Maintenance * 07:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1177: Maintenance * 06:24 root@cumin1003: START - Cookbook sre.mysql.pool pool db1177: Maintenance * 06:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1177 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95741 and previous config saved to /var/cache/conftool/dbconfig/20260730-061736-cwilliams.json * 06:17 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1177.eqiad.wmnet with reason: Maintenance * 06:17 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1172: Maintenance * 05:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1218.eqiad.wmnet with reason: crashed * 05:41 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db1217 it crashed', diff saved to https://phabricator.wikimedia.org/P95737 and previous config saved to /var/cache/conftool/dbconfig/20260730-054111-marostegui.json * 05:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95736 and previous config saved to /var/cache/conftool/dbconfig/20260730-053422-cwilliams.json * 05:30 root@cumin1003: START - Cookbook sre.mysql.pool pool db1172: Maintenance * 05:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95734 and previous config saved to /var/cache/conftool/dbconfig/20260730-052414-cwilliams.json * 05:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1172 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95733 and previous config saved to /var/cache/conftool/dbconfig/20260730-052354-cwilliams.json * 05:23 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1172.eqiad.wmnet with reason: Maintenance * 05:23 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1167: Maintenance * 05:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95731 and previous config saved to /var/cache/conftool/dbconfig/20260730-051406-cwilliams.json * 05:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95729 and previous config saved to /var/cache/conftool/dbconfig/20260730-050358-cwilliams.json * 04:47 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 04:47 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 04:47 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 04:35 root@cumin1003: START - Cookbook sre.mysql.pool pool db1167: Maintenance * 04:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1167 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95726 and previous config saved to /var/cache/conftool/dbconfig/20260730-042923-cwilliams.json * 04:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 04:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1167.eqiad.wmnet with reason: Maintenance * 04:22 pt1979@cumin2002: START - Cookbook sre.dns.netbox * 04:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95725 and previous config saved to /var/cache/conftool/dbconfig/20260730-040337-cwilliams.json * 04:03 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:38 brett@cumin2002: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool eqsin [reason: Switch upgrade maintenance window complete, [[phab:T433097|T433097]]] * 01:38 brett@cumin2002: START - Cookbook sre.dns.admin DNS admin: pool eqsin [reason: Switch upgrade maintenance window complete, [[phab:T433097|T433097]]] == 2026-07-29 == * 23:57 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin,mr1-eqsin IPv6,mr1-eqsin.oob,mr1-eqsin.oob IPv6 with reason: connection issue * 22:54 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 22:53 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 22:53 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 22:53 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:25 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2022.codfw.wmnet, repooling source-only afterwards * 22:20 brett@cumin2002: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool eqsin [reason: Switch upgrade maintenance window, [[phab:T433097|T433097]]] * 22:20 brett@cumin2002: START - Cookbook sre.dns.admin DNS admin: depool eqsin [reason: Switch upgrade maintenance window, [[phab:T433097|T433097]]] * 22:01 apine@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 22:00 apine@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 21:59 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 21:58 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 21:58 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 21:58 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 21:32 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1048.eqiad.wmnet with OS trixie * 21:25 pt1979@cumin2002: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be2097.codfw.wmnet with OS bullseye * 21:16 zabe@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki=metawiki 'Mental Health Resource Center' 'Safety Resource Center/Mental Health' Zabe --reason 'per request [[:phab:T433118{{!}}T433118]]' * 21:12 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2022.codfw.wmnet, repooling source-only afterwards * 21:12 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] (duration: 12m 53s) * 21:12 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2015\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 21:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1253: Maintenance * 21:08 aaron@deploy1003: aaron: Continuing with deployment * 21:01 aaron@deploy1003: aaron: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:59 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] * 20:52 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] (duration: 21m 57s) * 20:48 aaron@deploy1003: aaron: Continuing with deployment * 20:32 aaron@deploy1003: aaron: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:30 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] * 20:24 root@cumin1003: START - Cookbook sre.mysql.pool pool db1253: Maintenance * 20:19 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] (duration: 08m 07s) * 20:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1253 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95719 and previous config saved to /var/cache/conftool/dbconfig/20260729-201810-cwilliams.json * 20:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1253.eqiad.wmnet with reason: Maintenance * 20:17 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1231: Maintenance * 20:15 aaron@deploy1003: bpirkle, aaron: Continuing with deployment * 20:13 aaron@deploy1003: bpirkle, aaron: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:12 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie * 20:11 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] * 20:11 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1048.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:09 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1048.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:09 pt1979@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 20:04 pt1979@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 19:47 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 19:43 pt1979@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye * 19:41 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 19:37 zabe: zabe@deploy1003:~$ mwscript-k8s --comment='[[phab:T433529|T433529]]' --follow -- resetAuthenticationThrottle.php --wiki=aawiki --signup --ip=89.36.114.94 * 19:36 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] (duration: 06m 49s) * 19:32 zabe@deploy1003: zabe: Continuing with deployment * 19:31 zabe@deploy1003: zabe: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:31 root@cumin1003: START - Cookbook sre.mysql.pool pool db1231: Maintenance * 19:29 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] * 19:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1231 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95714 and previous config saved to /var/cache/conftool/dbconfig/20260729-192454-cwilliams.json * 19:24 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1231.eqiad.wmnet with reason: Maintenance * 19:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1227: Maintenance * 19:22 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 19:22 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 19:21 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:21 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1048] - vriley@cumin1003" * 19:21 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1048] - vriley@cumin1003" * 19:19 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 19:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1251: Maintenance * 19:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95711 and previous config saved to /var/cache/conftool/dbconfig/20260729-191756-cwilliams.json * 19:16 vriley@cumin1003: START - Cookbook sre.dns.netbox * 19:11 dduvall: rolling back wmf.13 to group0 due to [[phab:T433457|T433457]] (cc [[phab:T430832|T430832]]) * 19:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95709 and previous config saved to /var/cache/conftool/dbconfig/20260729-190748-cwilliams.json * 19:01 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2022.codfw.wmnet with OS bookworm * 19:01 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2021.codfw.wmnet, repooling source-only afterwards * 18:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95707 and previous config saved to /var/cache/conftool/dbconfig/20260729-185740-cwilliams.json * 18:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95704 and previous config saved to /var/cache/conftool/dbconfig/20260729-184732-cwilliams.json * 18:37 root@cumin1003: START - Cookbook sre.mysql.pool pool db1227: Maintenance * 18:34 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2022.codfw.wmnet with reason: host reimage * 18:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1227 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95701 and previous config saved to /var/cache/conftool/dbconfig/20260729-183117-cwilliams.json * 18:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1227.eqiad.wmnet with reason: Maintenance * 18:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1202: Maintenance * 18:30 root@cumin1003: START - Cookbook sre.mysql.pool pool db1251: Maintenance * 18:27 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2022.codfw.wmnet with reason: host reimage * 18:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1251 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95698 and previous config saved to /var/cache/conftool/dbconfig/20260729-182428-cwilliams.json * 18:24 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lvs2014.codfw.wmnet * 18:24 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for lvs2014.codfw.wmnet * 18:24 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1251.eqiad.wmnet with reason: Maintenance * 18:23 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1235: Maintenance * 18:22 brett@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 18:19 brett@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 18:19 brett@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 18:17 brett@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 18:17 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 18:16 mutante: removing jenkins during the train - living on the edge - no, just kidding, jenkins has migrated to dedicated machines, nothing should happen * 18:15 brett@cumin2002: END (ERROR) - Cookbook sre.loadbalancer.restart-pybal (exit_code=97) rolling-restart of pybal on P<nowiki>{</nowiki>lvs2014.codfw.wmnet<nowiki>}</nowiki> and A:lvs ([[phab:T428495|T428495]]) * 18:15 mutante: CI: contint1002/contint2002: apt-get remove --purge jenkins - jenkins be gone - [[phab:T418521|T418521]] * 18:13 brett@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on P<nowiki>{</nowiki>lvs2014.codfw.wmnet<nowiki>}</nowiki> and A:lvs ([[phab:T428495|T428495]]) * 18:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2022 * 18:08 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2022 * 18:03 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T428495|T428495]] * 18:03 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2022 * 18:02 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2022.codfw.wmnet 211.48.192.10.in-addr.arpa 1.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:02 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2022.codfw.wmnet 211.48.192.10.in-addr.arpa 1.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:02 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:02 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2022 - bking@cumin2003" * 18:02 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2022 - bking@cumin2003" * 17:57 bking@cumin2003: START - Cookbook sre.dns.netbox * 17:56 brett@cumin2002: END (FAIL) - Cookbook sre.loadbalancer.restart-pybal (exit_code=1) rolling-restart of pybal on A:lvs-codfw and A:lvs ([[phab:T428495|T428495]]) * 17:55 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo - [[phab:T428495|T428495]] * 17:54 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2022 * 17:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2022.codfw.wmnet with OS bookworm * 17:50 brett@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on A:lvs-codfw and A:lvs ([[phab:T428495|T428495]]) * 17:47 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2021.codfw.wmnet, repooling source-only afterwards * 17:47 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 14s) * 17:47 swfrench-wmf: authdns-update to direct codfw, eqsin, ulsfo etcd clients back to codfw - [[phab:T428495|T428495]] * 17:47 swfrench@dns1004: END - running authdns-update * 17:47 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 17:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95692 and previous config saved to /var/cache/conftool/dbconfig/20260729-174713-cwilliams.json * 17:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance * 17:46 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1249: Maintenance * 17:45 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2015\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 17:45 swfrench@dns1004: START - running authdns-update * 17:44 root@cumin1003: START - Cookbook sre.mysql.pool pool db1202: Maintenance * 17:41 akhatun: Deployed refinery using scap, then deployed onto hdfs * 17:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1202 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95688 and previous config saved to /var/cache/conftool/dbconfig/20260729-173759-cwilliams.json * 17:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1202.eqiad.wmnet with reason: Maintenance * 17:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1194: Maintenance * 17:37 root@cumin1003: START - Cookbook sre.mysql.pool pool db1235: Maintenance * 17:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1230: Maintenance * 17:30 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1235 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95684 and previous config saved to /var/cache/conftool/dbconfig/20260729-173051-cwilliams.json * 17:30 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1235.eqiad.wmnet with reason: Maintenance * 17:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1234: Maintenance * 17:26 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (thin): Regular analytics weekly train THIN [analytics/refinery@56695674] (duration: 02m 02s) * 17:24 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (thin): Regular analytics weekly train THIN [analytics/refinery@56695674] * 17:23 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567]: Regular analytics weekly train [analytics/refinery@56695674] (duration: 06m 20s) * 17:20 dancy@deploy1003: Finished scap sync-world: Testing delay_messageblobstore_purge: true (duration: 06m 29s) * 17:17 akhatun@deploy1003: Started deploy [analytics/refinery@5669567]: Regular analytics weekly train [analytics/refinery@56695674] * 17:17 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] (duration: 00m 22s) * 17:16 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] * 17:13 dancy@deploy1003: Started scap sync-world: Testing delay_messageblobstore_purge: true * 17:05 mutante: CI: contint1002/contint2002 - restarted httpd to be extra sure all is cleaned up - https://integration.wikimedia.org/ci/ is up and running [[phab:T418521|T418521]] * 17:04 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 17:03 mutante: CI: contint1002/contint2002 - rm /etc/apache2/jenkins_proxy - removing legacy jenkins proxy config - jenkins is on new dedicated machines and uses jenkins_proxy_ext config [[phab:T418521|T418521]] * 17:02 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] (duration: 36m 25s) * 17:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1249: Maintenance * 16:59 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 16:54 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2015.codfw.wmnet, repooling source-only afterwards * 16:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1249 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95674 and previous config saved to /var/cache/conftool/dbconfig/20260729-165339-cwilliams.json * 16:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1249.eqiad.wmnet with reason: Maintenance * 16:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1248: Maintenance * 16:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1194: Maintenance * 16:47 swfrench-wmf: silenced EtcdReplicationDown 57b2b421-1cc9-4e38-9276-{{Gerrit|94f223fd231c}} - [[phab:T428495|T428495]] * 16:46 tchin@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/eventstreams-internal: apply * 16:46 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye * 16:46 tchin@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/eventstreams-internal: apply * 16:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1230: Maintenance * 16:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1194 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95669 and previous config saved to /var/cache/conftool/dbconfig/20260729-164422-cwilliams.json * 16:44 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 16:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1194.eqiad.wmnet with reason: Maintenance * 16:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1191: Maintenance * 16:43 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host an-test-master1003.eqiad.wmnet * 16:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db1234: Maintenance * 16:43 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 16:43 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Rolling back deployment * 16:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host an-test-master1004.eqiad.wmnet * 16:41 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 16:40 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 16:40 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 16:39 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 16:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1230 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95667 and previous config saved to /var/cache/conftool/dbconfig/20260729-163932-cwilliams.json * 16:39 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 16:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1230.eqiad.wmnet with reason: Maintenance * 16:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1207: Maintenance * 16:38 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 16:37 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host an-test-master1004.eqiad.wmnet * 16:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1234 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95664 and previous config saved to /var/cache/conftool/dbconfig/20260729-163719-cwilliams.json * 16:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1234.eqiad.wmnet with reason: Maintenance * 16:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1079.eqiad.wmnet with OS trixie * 16:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1232: Maintenance * 16:34 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1259: Maintenance * 16:28 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] (duration: 06m 57s) * 16:28 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:26 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] * 16:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1051 hosts * 16:21 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] * 16:20 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2006.codfw.wmnet with OS bookworm * 16:19 akhatun: Deploying Refinery at {{Gerrit|56695674}} as part of weekly train * 16:18 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1079.eqiad.wmnet with reason: host reimage * 16:16 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] (duration: 15m 36s) * 16:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2021.codfw.wmnet with OS bookworm * 16:14 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1079.eqiad.wmnet with reason: host reimage * 16:12 topranks: hot-swap line card in FPC0 on cr1-eqiad with replacement MPC10E from Juniper [[phab:T426343|T426343]] * 16:10 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Continuing with deployment * 16:07 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db1248: Maintenance * 16:01 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] * 16:00 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 16:00 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 15:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1248 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95651 and previous config saved to /var/cache/conftool/dbconfig/20260729-155956-cwilliams.json * 15:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1248.eqiad.wmnet with reason: Maintenance * 15:59 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2006.codfw.wmnet with reason: host reimage * 15:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1247: Maintenance * 15:59 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2074.codfw.wmnet with OS trixie * 15:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1191: Maintenance * 15:57 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:55 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1079.eqiad.wmnet with OS trixie * 15:55 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2006.codfw.wmnet with reason: host reimage * 15:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1207: Maintenance * 15:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1191 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95646 and previous config saved to /var/cache/conftool/dbconfig/20260729-155104-cwilliams.json * 15:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1191.eqiad.wmnet with reason: Maintenance * 15:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1181: Maintenance * 15:49 root@cumin1003: START - Cookbook sre.mysql.pool pool db1232: Maintenance * 15:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2021.codfw.wmnet with reason: host reimage * 15:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1207 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95643 and previous config saved to /var/cache/conftool/dbconfig/20260729-154735-cwilliams.json * 15:47 root@cumin1003: START - Cookbook sre.mysql.pool pool db1259: Maintenance * 15:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1207.eqiad.wmnet with reason: Maintenance * 15:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1200: Maintenance * 15:46 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] (duration: 31m 59s) * 15:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 15:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1232 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95640 and previous config saved to /var/cache/conftool/dbconfig/20260729-154330-cwilliams.json * 15:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1232.eqiad.wmnet with reason: Maintenance * 15:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1219: Maintenance * 15:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2015.codfw.wmnet, repooling source-only afterwards * 15:41 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 18s) * 15:41 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1259 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95638 and previous config saved to /var/cache/conftool/dbconfig/20260729-154107-cwilliams.json * 15:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1259.eqiad.wmnet with reason: Maintenance * 15:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1254: Maintenance * 15:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2015.codfw.wmnet with OS bookworm * 15:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2021.codfw.wmnet with reason: host reimage * 15:36 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2006.codfw.wmnet with OS bookworm * 15:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 15:35 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Continuing with deployment * 15:33 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2074.codfw.wmnet with OS trixie * 15:33 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 15:32 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:29 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:28 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be2074.codfw.wmnet with OS trixie * 15:28 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2006.codfw.wmnet * 15:26 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1078.eqiad.wmnet with OS trixie * 15:25 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 15:25 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:22 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2006.codfw.wmnet * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2021 * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2021 * 15:19 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2021 * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2021.codfw.wmnet 210.48.192.10.in-addr.arpa 0.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:19 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2021.codfw.wmnet 210.48.192.10.in-addr.arpa 0.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2021 - bking@cumin2003" * 15:19 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2021 - bking@cumin2003" * 15:14 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] * 15:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2015.codfw.wmnet with reason: host reimage * 15:11 root@cumin1003: START - Cookbook sre.mysql.pool pool db1247: Maintenance * 15:11 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on ml-serve2004.codfw.wmnet with reason: [[phab:T433478|T433478]] * 15:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2015.codfw.wmnet with reason: host reimage * 15:10 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on ml-serve2002.codfw.wmnet with reason: [[phab:T433476|T433476]] * 15:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 15:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1247 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95625 and previous config saved to /var/cache/conftool/dbconfig/20260729-150459-cwilliams.json * 15:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1247.eqiad.wmnet with reason: Maintenance * 15:04 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:04 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1244: Maintenance * 15:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1078.eqiad.wmnet with reason: host reimage * 15:03 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1005.wikimedia.org * 15:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db1181: Maintenance * 15:01 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2098.codfw.wmnet with OS bullseye * 15:00 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye * 15:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1200: Maintenance * 14:59 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 14:59 root@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1285.eqiad.wmnet with OS trixie * 14:59 Amir1: mwscript-k8s -- extensions/TimedMediaHandler/maintenance/requeueTranscodes.php --wiki=commonswiki --key '360p.mpeg4.mov' --throttle --video --missing ([[phab:T358266|T358266]]) * 14:58 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1005.wikimedia.org * 14:58 jhancock@cumin2002: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['ms-be2098'] * 14:58 jhancock@cumin2002: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['ms-be2098'] * 14:58 jhancock@cumin2002: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['ms-be2097'] * 14:58 jhancock@cumin2002: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['ms-be2097'] * 14:58 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1078.eqiad.wmnet with reason: host reimage * 14:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1181 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95621 and previous config saved to /var/cache/conftool/dbconfig/20260729-145629-cwilliams.json * 14:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1181.eqiad.wmnet with reason: Maintenance * 14:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1174: Maintenance * 14:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1219: Maintenance * 14:55 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader1006.wikimedia.org on all recursors * 14:55 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader1006.wikimedia.org on all recursors * 14:55 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader1005.wikimedia.org on all recursors * 14:55 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader1005.wikimedia.org on all recursors * 14:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1200 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95618 and previous config saved to /var/cache/conftool/dbconfig/20260729-145336-cwilliams.json * 14:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1254: Maintenance * 14:53 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1200.eqiad.wmnet with reason: Maintenance * 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1185: Maintenance * 14:52 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2021 * 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2015 * 14:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2015 * 14:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1219 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95616 and previous config saved to /var/cache/conftool/dbconfig/20260729-144946-cwilliams.json * 14:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1219.eqiad.wmnet with reason: Maintenance * 14:49 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1218: Maintenance * 14:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2021.codfw.wmnet with OS bookworm * 14:48 dancy@deploy1003: Finished deploy [zuul/deploy@22703a6]: Deploying https://gerrit.wikimedia.org/r/c/integration/zuul/+/1311501 ([[phab:T432491|T432491]]) (duration: 00m 15s) * 14:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2015.codfw.wmnet with OS bookworm * 14:48 dancy@deploy1003: Started deploy [zuul/deploy@22703a6]: Deploying https://gerrit.wikimedia.org/r/c/integration/zuul/+/1311501 ([[phab:T432491|T432491]]) * 14:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1254 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95613 and previous config saved to /var/cache/conftool/dbconfig/20260729-144729-cwilliams.json * 14:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1254.eqiad.wmnet with reason: Maintenance * 14:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1233: Maintenance * 14:46 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2013\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 14:46 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2014\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 14:46 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:45 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:44 root@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1285.eqiad.wmnet with reason: host reimage * 14:43 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:42 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:41 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:40 root@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1285.eqiad.wmnet with reason: host reimage * 14:39 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1078.eqiad.wmnet with OS trixie * 14:39 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2074.codfw.wmnet with OS trixie * 14:32 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:32 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:32 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:31 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2005.codfw.wmnet with OS bookworm * 14:30 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:30 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:29 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:29 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:27 root@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host db1285 * 14:27 root@cumin1003: START - Cookbook sre.hosts.move-vlan for host db1285 * 14:27 root@cumin1003: START - Cookbook sre.hosts.reimage for host db1285.eqiad.wmnet with OS trixie * 14:24 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 14:24 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:24 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:24 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:23 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:22 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:22 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:21 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:17 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:16 root@cumin1003: START - Cookbook sre.mysql.pool pool db1244: Maintenance * 14:15 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:15 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add asw1-604 loopback ipv4 - pt1979@cumin2002" * 14:15 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add asw1-604 loopback ipv4 - pt1979@cumin2002" * 14:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:12 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 14:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95599 and previous config saved to /var/cache/conftool/dbconfig/20260729-141014-cwilliams.json * 14:10 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 14:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1244.eqiad.wmnet with reason: Maintenance * 14:10 pt1979@cumin2002: START - Cookbook sre.dns.netbox * 14:10 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 14:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1243: Maintenance * 14:09 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2005.codfw.wmnet with reason: host reimage * 14:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db1174: Maintenance * 14:08 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad * 14:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db1185: Maintenance * 14:06 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 14:05 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2005.codfw.wmnet with reason: host reimage * 14:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1174 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95595 and previous config saved to /var/cache/conftool/dbconfig/20260729-140309-cwilliams.json * 14:03 sukhe@dns1004: END - running authdns-update * 14:03 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1174.eqiad.wmnet with reason: Maintenance * 14:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1170: Maintenance * 14:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db1218: Maintenance * 14:01 sukhe@dns1004: START - running authdns-update * 14:00 sukhe@dns1004: START - running authdns-update * 13:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db1233: Maintenance * 13:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1185 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95592 and previous config saved to /var/cache/conftool/dbconfig/20260729-135925-cwilliams.json * 13:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1185.eqiad.wmnet with reason: Maintenance * 13:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1161: Maintenance * 13:58 sukhe@puppetserver1001: conftool action : set/pooled=true; selector: dnsdisc=urldownloader * 13:58 root@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1265.eqiad.wmnet with OS trixie * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1218 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95590 and previous config saved to /var/cache/conftool/dbconfig/20260729-135621-cwilliams.json * 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1218.eqiad.wmnet with reason: Maintenance * 13:55 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1206: Maintenance * 13:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2073.codfw.wmnet with OS trixie * 13:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1233 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95587 and previous config saved to /var/cache/conftool/dbconfig/20260729-135335-cwilliams.json * 13:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1233.eqiad.wmnet with reason: Maintenance * 13:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1229: Maintenance * 13:50 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/kartotherian: apply * 13:50 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service * 13:49 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:49 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/kartotherian: apply * 13:48 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 13:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1077.eqiad.wmnet with OS trixie * 13:47 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 13:46 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2005.codfw.wmnet with OS bookworm * 13:44 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 13:44 root@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1265.eqiad.wmnet with reason: host reimage * 13:40 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] (duration: 09m 22s) * 13:39 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:38 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:36 root@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1265.eqiad.wmnet with reason: host reimage * 13:35 stran@deploy1003: stran: Continuing with deployment * 13:33 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 13:32 stran@deploy1003: stran: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified t * 13:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2073.codfw.wmnet with reason: host reimage * 13:30 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ml-build1001.eqiad.wmnet * 13:30 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] * 13:29 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad * 13:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1077.eqiad.wmnet with reason: host reimage * 13:27 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 13:27 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:27 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad * 13:26 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2073.codfw.wmnet with reason: host reimage * 13:26 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] (duration: 07m 56s) * 13:25 klausman@cumin1003: START - Cookbook sre.hosts.reboot-single for host ml-build1001.eqiad.wmnet * 13:24 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 13:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>ml-serve101[2-5].eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 13:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1015.eqiad.wmnet * 13:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1015.eqiad.wmnet * 13:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1077.eqiad.wmnet with reason: host reimage * 13:23 root@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host db1265 * 13:23 root@cumin1003: START - Cookbook sre.hosts.move-vlan for host db1265 * 13:23 root@cumin1003: START - Cookbook sre.hosts.reimage for host db1265.eqiad.wmnet with OS trixie * 13:23 root@cumin1003: START - Cookbook sre.mysql.pool pool db1243: Maintenance * 13:22 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 13:22 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2005.codfw.wmnet * 13:22 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 13:22 samtar@deploy1003: dreamrimmer, samtar: Continuing with deployment * 13:20 samtar@deploy1003: dreamrimmer, samtar: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts an-test-master[1001-1002].eqiad.wmnet * 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-master[1001-1002].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 13:18 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1015.eqiad.wmnet * 13:18 sukhe@cumin1003: END (ERROR) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=97) for role: url_downloader@eqiad * 13:18 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 13:18 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] * 13:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95574 and previous config saved to /var/cache/conftool/dbconfig/20260729-131638-cwilliams.json * 13:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1243.eqiad.wmnet with reason: Maintenance * 13:16 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2005.codfw.wmnet * 13:16 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1242: Maintenance * 13:14 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] (duration: 07m 00s) * 13:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db1170: Maintenance * 13:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1015.eqiad.wmnet * 13:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1014.eqiad.wmnet * 13:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1014.eqiad.wmnet * 13:12 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1223: Maintenance * 13:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db1161: Maintenance * 13:10 samtar@deploy1003: anzx, samtar: Continuing with deployment * 13:09 samtar@deploy1003: anzx, samtar: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db1206: Maintenance * 13:08 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 13:07 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] * 13:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1170 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95566 and previous config saved to /var/cache/conftool/dbconfig/20260729-130730-cwilliams.json * 13:07 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1170.eqiad.wmnet with reason: Maintenance * 13:07 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:07 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt IPs new switches - cmooney@cumin1003" * 13:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1158: Maintenance * 13:06 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1014.eqiad.wmnet * 13:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1161 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95564 and previous config saved to /var/cache/conftool/dbconfig/20260729-130616-cwilliams.json * 13:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 13:06 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1077.eqiad.wmnet with OS trixie * 13:05 root@cumin1003: START - Cookbook sre.mysql.pool pool db1229: Maintenance * 13:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1161.eqiad.wmnet with reason: Maintenance * 13:05 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt IPs new switches - cmooney@cumin1003" * 13:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2073.codfw.wmnet with OS trixie * 13:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Maintenance * 13:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1206 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95562 and previous config saved to /var/cache/conftool/dbconfig/20260729-130258-cwilliams.json * 13:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1206.eqiad.wmnet with reason: Maintenance * 13:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1196: Maintenance * 13:01 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 13:01 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:00 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 13:00 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 12:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1229 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95559 and previous config saved to /var/cache/conftool/dbconfig/20260729-125950-cwilliams.json * 12:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1229.eqiad.wmnet with reason: Maintenance * 12:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1222: Maintenance * 12:57 sukhe: sudo cumin 'A:lvs and (A:eqiad or A:codfw)' 'disable-puppet "adding new service urldownloader"': [[phab:T429175|T429175]] * 12:56 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1014.eqiad.wmnet * 12:56 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1013.eqiad.wmnet * 12:56 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1013.eqiad.wmnet * 12:50 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1013.eqiad.wmnet * 12:50 sukhe: sudo cumin 'O:url_downloader' 'run-puppet-agent --enable "merging CR 1313948"': [[phab:T429175|T429175]] * 12:48 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-master[1001-1002].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 12:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1013.eqiad.wmnet * 12:45 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1012.eqiad.wmnet * 12:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1012.eqiad.wmnet * 12:45 sukhe: sudo cumin 'O:url_downloader' 'disable-puppet "merging CR 1313948"': [[phab:T429175|T429175]] * 12:44 btullis@cumin1003: START - Cookbook sre.dns.netbox * 12:40 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test2001.codfw.wmnet * 12:40 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test2001.codfw.wmnet * 12:38 ayounsi@dns1004: END - running authdns-update * 12:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1012.eqiad.wmnet * 12:35 ayounsi@dns1004: START - running authdns-update * 12:34 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts an-test-master[1001-1002].eqiad.wmnet * 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts an-test-coord1001.eqiad.wmnet * 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-coord1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 12:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1012.eqiad.wmnet * 12:32 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>ml-serve101[2-5].eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 12:29 root@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Maintenance * 12:25 root@cumin1003: START - Cookbook sre.mysql.pool pool db1223: Maintenance * 12:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95544 and previous config saved to /var/cache/conftool/dbconfig/20260729-122254-cwilliams.json * 12:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1242.eqiad.wmnet with reason: Maintenance * 12:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1241: Maintenance * 12:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1051 hosts * 12:20 root@cumin1003: START - Cookbook sre.mysql.pool pool db1158: Maintenance * 12:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1223 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95540 and previous config saved to /var/cache/conftool/dbconfig/20260729-121937-cwilliams.json * 12:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1223.eqiad.wmnet with reason: Maintenance * 12:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1212: Maintenance * 12:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db1159: Maintenance * 12:17 elukey@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: sync * 12:15 elukey@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: sync * 12:15 root@cumin1003: START - Cookbook sre.mysql.pool pool db1196: Maintenance * 12:14 Daimona: Creating new DB tables for the CampaignEvents extension in x1.testwiki, x1.test2wiki, x1.officewiki, and x1.wikishared # [[phab:T429339|T429339]] * 12:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db1222: Maintenance * 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95535 and previous config saved to /var/cache/conftool/dbconfig/20260729-121211-cwilliams.json * 12:12 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 12:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1158.eqiad.wmnet with reason: Maintenance * 12:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1159 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95534 and previous config saved to /var/cache/conftool/dbconfig/20260729-121146-cwilliams.json * 12:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1159.eqiad.wmnet with reason: Maintenance * 12:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1196 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95533 and previous config saved to /var/cache/conftool/dbconfig/20260729-120847-cwilliams.json * 12:08 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 12:08 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1196.eqiad.wmnet with reason: Maintenance * 12:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1195: Maintenance * 12:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1222 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95530 and previous config saved to /var/cache/conftool/dbconfig/20260729-120424-cwilliams.json * 12:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1222.eqiad.wmnet with reason: Maintenance * 12:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1098 hosts * 12:00 marostegui: Rename tables [[phab:T425074|T425074]] * 12:00 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-coord1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 11:58 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1197: Maintenance * 11:55 btullis@cumin1003: START - Cookbook sre.dns.netbox * 11:52 elukey@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: sync * 11:51 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:51 elukey@deploy1003: helmfile [codfw] START helmfile.d/services/proton: sync * 11:51 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:50 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts an-test-coord1001.eqiad.wmnet * 11:50 elukey@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: sync * 11:49 elukey@deploy1003: helmfile [staging] START helmfile.d/services/proton: sync * 11:35 root@cumin1003: START - Cookbook sre.mysql.pool pool db1241: Maintenance * 11:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db1212: Maintenance * 11:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1241 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95520 and previous config saved to /var/cache/conftool/dbconfig/20260729-112918-cwilliams.json * 11:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1241.eqiad.wmnet with reason: Maintenance * 11:29 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1238: Maintenance * 11:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1212 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95517 and previous config saved to /var/cache/conftool/dbconfig/20260729-112727-cwilliams.json * 11:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 11:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1212.eqiad.wmnet with reason: Maintenance * 11:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1198: Maintenance * 11:23 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:22 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:21 root@cumin1003: START - Cookbook sre.mysql.pool pool db1195: Maintenance * 11:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1195 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95514 and previous config saved to /var/cache/conftool/dbconfig/20260729-111450-cwilliams.json * 11:14 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1195.eqiad.wmnet with reason: Maintenance * 11:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1186: Maintenance * 11:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 11:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 11:05 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:54 marostegui: Dropping renamed tables [[phab:T425066|T425066]] * 10:41 root@cumin1003: START - Cookbook sre.mysql.pool pool db1238: Maintenance * 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1198: Maintenance * 10:39 Amir1: ran https://phabricator.wikimedia.org/T432509#12149723 in production ([[phab:T432509|T432509]]) * 10:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db1197: Maintenance * 10:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1238 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95501 and previous config saved to /var/cache/conftool/dbconfig/20260729-103532-cwilliams.json * 10:35 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1238.eqiad.wmnet with reason: Maintenance * 10:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1221: Maintenance * 10:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1198 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95499 and previous config saved to /var/cache/conftool/dbconfig/20260729-103330-cwilliams.json * 10:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1198.eqiad.wmnet with reason: Maintenance * 10:33 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1175: Maintenance * 10:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1197 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95496 and previous config saved to /var/cache/conftool/dbconfig/20260729-103217-cwilliams.json * 10:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1197.eqiad.wmnet with reason: Maintenance * 10:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1188: Maintenance * 10:27 root@cumin1003: START - Cookbook sre.mysql.pool pool db1186: Maintenance * 10:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1186 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95493 and previous config saved to /var/cache/conftool/dbconfig/20260729-102111-cwilliams.json * 10:21 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1186.eqiad.wmnet with reason: Maintenance * 10:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 10:14 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 09:53 XioNoX: reboot cr2-magru - [[phab:T431750|T431750]] * 09:52 XioNoX: drain cr2-magru - [[phab:T431750|T431750]] * 09:48 root@cumin1003: START - Cookbook sre.mysql.pool pool db1221: Maintenance * 09:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zookeeper-test1002.eqiad.wmnet * 09:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1188: Maintenance * 09:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1175: Maintenance * 09:44 btullis@dns1004: END - running authdns-update * 09:42 btullis@dns1004: START - running authdns-update * 09:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1221 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95483 and previous config saved to /var/cache/conftool/dbconfig/20260729-094200-cwilliams.json * 09:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 7 hosts with reason: Maintenance * 09:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1221.eqiad.wmnet with reason: Maintenance * 09:41 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host zookeeper-test1002.eqiad.wmnet * 09:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1199: Maintenance * 09:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1188 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95481 and previous config saved to /var/cache/conftool/dbconfig/20260729-093917-cwilliams.json * 09:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1188.eqiad.wmnet with reason: Maintenance * 09:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1182: Maintenance * 09:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1175 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95479 and previous config saved to /var/cache/conftool/dbconfig/20260729-093842-cwilliams.json * 09:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1175.eqiad.wmnet with reason: Maintenance * 09:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1166: Maintenance * 09:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1169: Maintenance * 09:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1033.eqiad.wmnet,service=s8 * 09:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1033.eqiad.wmnet,service=s5 * 09:33 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1033.eqiad.wmnet,service=s8 * 09:33 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1033.eqiad.wmnet,service=s5 * 09:21 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:21 XioNoX: reboot cr1-magru - [[phab:T431750|T431750]] * 09:17 XioNoX: drain cr1-magru - [[phab:T431750|T431750]] * 09:15 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm1001.wikimedia.org * 09:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply * 09:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply * 09:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 09:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 09:11 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr2-magru,cr2-magru IPv6,cr2-magru.mgmt with reason: router upgrade * 09:11 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm1001.wikimedia.org * 09:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp1005.wikimedia.org * 09:07 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp1005.wikimedia.org * 09:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp2005.wikimedia.org * 09:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp2005.wikimedia.org * 09:00 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr1-magru,cr1-magru IPv6,cr1-magru.mgmt with reason: router upgrade * 09:00 marostegui: Dropping renamed tables [[phab:T426341|T426341]] * 08:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1199: Maintenance * 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1182: Maintenance * 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1166: Maintenance * 08:47 root@cumin1003: START - Cookbook sre.mysql.pool pool db1169: Maintenance * 08:46 ayounsi@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 1:00:00 on cr1-magru,cr1-magru IPv6,cr1-magru.mgmt with reason: router upgrade * 08:45 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2072.codfw.wmnet with OS trixie * 08:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1199 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95464 and previous config saved to /var/cache/conftool/dbconfig/20260729-084534-cwilliams.json * 08:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1199.eqiad.wmnet with reason: Maintenance * 08:45 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 08:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1190: Maintenance * 08:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1182 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95462 and previous config saved to /var/cache/conftool/dbconfig/20260729-084436-cwilliams.json * 08:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1182.eqiad.wmnet with reason: Maintenance * 08:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1156: Maintenance * 08:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1166 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95460 and previous config saved to /var/cache/conftool/dbconfig/20260729-084400-cwilliams.json * 08:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1166.eqiad.wmnet with reason: Maintenance * 08:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1157: Maintenance * 08:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95458 and previous config saved to /var/cache/conftool/dbconfig/20260729-084147-cwilliams.json * 08:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1169.eqiad.wmnet with reason: Maintenance * 08:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1163: Maintenance * 08:30 btullis@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 11 hosts with reason: Replacing the namenodes * 08:23 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2072.codfw.wmnet with reason: host reimage * 08:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1098 hosts * 08:19 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2072.codfw.wmnet with reason: host reimage * 07:58 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2072.codfw.wmnet with OS trixie * 07:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1190: Maintenance * 07:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1156: Maintenance * 07:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1157: Maintenance * 07:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1163: Maintenance * 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1190 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95444 and previous config saved to /var/cache/conftool/dbconfig/20260729-074930-cwilliams.json * 07:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1190.eqiad.wmnet with reason: Maintenance * 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1157 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95443 and previous config saved to /var/cache/conftool/dbconfig/20260729-074914-cwilliams.json * 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1156 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95442 and previous config saved to /var/cache/conftool/dbconfig/20260729-074906-cwilliams.json * 07:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1157.eqiad.wmnet with reason: Maintenance * 07:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 07:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1156.eqiad.wmnet with reason: Maintenance * 07:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1163 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95441 and previous config saved to /var/cache/conftool/dbconfig/20260729-074652-cwilliams.json * 07:46 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1163.eqiad.wmnet with reason: Maintenance * 07:46 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2034.codfw.wmnet * 07:42 ayounsi@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2034.codfw.wmnet * 07:42 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2232.codfw.wmnet with OS trixie * 07:34 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:34 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:33 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:31 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:19 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2232.codfw.wmnet with reason: host reimage * 07:15 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2232.codfw.wmnet with reason: host reimage * 06:58 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db2232.codfw.wmnet with OS trixie * 06:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[2160,2232].codfw.wmnet with reason: Reimage * 06:26 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1164.eqiad.wmnet with OS trixie * 06:05 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1164.eqiad.wmnet with reason: host reimage * 06:01 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1164.eqiad.wmnet with reason: host reimage * 05:47 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1164.eqiad.wmnet with OS trixie * 05:46 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1164.eqiad.wmnet with reason: Reimage == 2026-07-28 == * 22:50 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1138.eqiad.wmnet * 22:50 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1138.eqiad.wmnet * 22:49 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1138.eqiad.wmnet * 22:11 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2014.codfw.wmnet, repooling source-only afterwards * 22:08 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2013.codfw.wmnet, repooling source-only afterwards * 22:03 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 20:58 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] (duration: 08m 19s) * 20:55 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2014.codfw.wmnet, repooling source-only afterwards * 20:55 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2013.codfw.wmnet, repooling source-only afterwards * 20:54 arlolra@deploy1003: arlolra: Continuing with deployment * 20:54 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 14s) * 20:54 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 20:53 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 30s) * 20:53 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 20:52 arlolra@deploy1003: arlolra: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:51 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:50 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] * 20:49 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:43 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 20:34 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] (duration: 06m 54s) * 20:34 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:34 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:31 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:31 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:30 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:30 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:30 arlolra@deploy1003: arlolra: Continuing with deployment * 20:29 arlolra@deploy1003: arlolra: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:27 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] * 20:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2014.codfw.wmnet with OS bookworm * 20:21 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 20:21 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:20 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 20:19 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:19 swfrench-wmf: switched etcd-mirror replication from conf2005 to conf2004 - [[phab:T428495|T428495]] * 20:17 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:17 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:15 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] (duration: 08m 26s) * 20:12 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:11 arlolra@deploy1003: anzx, arlolra: Continuing with deployment * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2013.codfw.wmnet with OS bookworm * 20:09 arlolra@deploy1003: anzx, arlolra: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] * 20:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2216: Maintenance * 19:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2014.codfw.wmnet with reason: host reimage * 19:57 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:54 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:54 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:53 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:52 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2014.codfw.wmnet with reason: host reimage * 19:49 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2013.codfw.wmnet with reason: host reimage * 19:42 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 19:41 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:41 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2013.codfw.wmnet with reason: host reimage * 19:39 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:39 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:39 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-eqiad: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 19:38 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2014 * 19:33 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2014 * 19:29 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2014.codfw.wmnet with OS bookworm * 19:28 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:27 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1006 * 19:26 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2012\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 19:26 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1006 * 19:26 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:26 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 19:25 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2013 * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2013 * 19:21 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2013 * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2013.codfw.wmnet 84.0.192.10.in-addr.arpa 4.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:21 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2013.codfw.wmnet 84.0.192.10.in-addr.arpa 4.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2013 - bking@cumin2003" * 19:21 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2013 - bking@cumin2003" * 19:21 vriley@cumin1003: START - Cookbook sre.dns.netbox * 19:20 root@cumin1003: START - Cookbook sre.mysql.pool pool db2216: Maintenance * 19:13 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2216 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95435 and previous config saved to /var/cache/conftool/dbconfig/20260728-191343-cwilliams.json * 19:13 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2216.codfw.wmnet with reason: Maintenance * 19:13 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2203: Maintenance * 19:06 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1005.eqiad.wmnet with OS trixie * 19:06 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 19:06 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 18:46 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 18:45 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:45 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:43 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:40 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:36 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-eqiad: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 18:35 dancy@deploy1003: Installation of scap version "4.275.0" completed for 3 hosts * 18:33 dancy@deploy1003: Installing scap version "4.275.0" for 3 host(s) * 18:32 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:32 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2097.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:30 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns3003.wikimedia.org [reason: pool for all services after reimaging] * 18:29 sukhe@dns1004: END - running authdns-update * 18:27 sukhe@dns1004: START - running authdns-update * 18:27 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns3003.wikimedia.org,service=authdns-update [reason: pool authdns-update after reimaging] * 18:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db2203: Maintenance * 18:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2203 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95430 and previous config saved to /var/cache/conftool/dbconfig/20260728-181958-cwilliams.json * 18:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2203.codfw.wmnet with reason: Maintenance * 18:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2188: Maintenance * 18:18 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2097.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:17 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-be2098 * 18:17 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host ms-be2098 * 18:17 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-be2097 * 18:16 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host ms-be2097 * 18:15 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:15 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding ms-be2097-8 to codfw - jhancock@cumin2002" * 18:15 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding ms-be2097-8 to codfw - jhancock@cumin2002" * 18:10 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 18:08 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage * 18:05 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns3003.wikimedia.org with OS trixie * 18:03 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage * 17:56 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-codfw: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 17:45 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie * 17:45 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1005.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:41 sukhe@dns1004: END - running authdns-update * 17:39 sukhe@dns1004: START - running authdns-update * 17:36 sukhe@puppetserver1001: conftool action : set/weight=1; selector: cluster=urldownloader,service=squid * 17:36 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1005.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:35 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader,service=squid * 17:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 17:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1005 * 17:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 17:34 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1005 * 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1005] - vriley@cumin1003" * 17:34 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1005] - vriley@cumin1003" * 17:32 root@cumin1003: START - Cookbook sre.mysql.pool pool db2188: Maintenance * 17:29 vriley@cumin1003: START - Cookbook sre.dns.netbox * 17:29 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 17:26 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2188 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95425 and previous config saved to /var/cache/conftool/dbconfig/20260728-172609-cwilliams.json * 17:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2188.codfw.wmnet with reason: Maintenance * 17:25 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2176: Maintenance * 17:19 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1138.eqiad.wmnet with OS trixie * 17:18 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1005 * 17:18 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1005 * 17:18 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:15 vriley@cumin1003: START - Cookbook sre.dns.netbox * 17:13 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns3003.wikimedia.org with reason: host reimage * 17:07 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns3003.wikimedia.org with reason: host reimage * 17:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-eqiad * 17:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1015.eqiad.wmnet * 17:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1015.eqiad.wmnet * 16:59 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1138.eqiad.wmnet with reason: host reimage * 16:55 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-codfw: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 16:54 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1138.eqiad.wmnet with reason: host reimage * 16:53 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1015.eqiad.wmnet * 16:43 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns3003.wikimedia.org with OS trixie * 16:43 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1015.eqiad.wmnet * 16:43 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1014.eqiad.wmnet * 16:43 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1014.eqiad.wmnet * 16:43 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=dns3003.wikimedia.org [reason: depooling for reimage to trixie] * 16:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 290 hosts * 16:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db2176: Maintenance * 16:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1138 * 16:38 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1138 * 16:37 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1138 * 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1138.eqiad.wmnet 193.32.64.10.in-addr.arpa 3.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:37 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1138.eqiad.wmnet 193.32.64.10.in-addr.arpa 3.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1138 - jiji@cumin1003" * 16:37 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1138 - jiji@cumin1003" * 16:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1014.eqiad.wmnet * 16:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2176 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95420 and previous config saved to /var/cache/conftool/dbconfig/20260728-163235-cwilliams.json * 16:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2176.codfw.wmnet with reason: Maintenance * 16:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1014.eqiad.wmnet * 16:32 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1013.eqiad.wmnet * 16:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1013.eqiad.wmnet * 16:32 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2174: Maintenance * 16:28 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2012.codfw.wmnet, repooling source-only afterwards * 16:25 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1013.eqiad.wmnet * 16:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1013.eqiad.wmnet * 16:20 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1012.eqiad.wmnet * 16:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1012.eqiad.wmnet * 16:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1012.eqiad.wmnet * 16:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1012.eqiad.wmnet * 16:03 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1011.eqiad.wmnet * 16:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1011.eqiad.wmnet * 16:00 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2004.codfw.wmnet with OS bookworm * 15:59 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1011.eqiad.wmnet * 15:56 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:55 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 15:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:54 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1011.eqiad.wmnet * 15:54 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1010.eqiad.wmnet * 15:54 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1010.eqiad.wmnet * 15:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:50 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 15:49 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1010.eqiad.wmnet * 15:48 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 15:48 jiji@cumin1003: START - Cookbook sre.dns.netbox * 15:46 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2248: Maintenance * 15:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db2174: Maintenance * 15:44 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1010.eqiad.wmnet * 15:44 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1009.eqiad.wmnet * 15:44 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1009.eqiad.wmnet * 15:42 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1138 * 15:41 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1138.eqiad.wmnet with OS trixie * 15:39 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1009.eqiad.wmnet * 15:39 robh@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on arclamp2001.codfw.wmnet with reason: ram upgrade * 15:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2174 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95413 and previous config saved to /var/cache/conftool/dbconfig/20260728-153844-cwilliams.json * 15:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2174.codfw.wmnet with reason: Maintenance * 15:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2173: Maintenance * 15:37 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 15:35 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 15:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1009.eqiad.wmnet * 15:34 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1008.eqiad.wmnet * 15:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1008.eqiad.wmnet * 15:31 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 15:31 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 15:29 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1008.eqiad.wmnet * 15:27 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 290 hosts * 15:25 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2013 * 15:25 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2195: Maintenance * 15:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1008.eqiad.wmnet * 15:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1007.eqiad.wmnet * 15:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1007.eqiad.wmnet * 15:22 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2004.codfw.wmnet with reason: host reimage * 15:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2013.codfw.wmnet with OS bookworm * 15:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 15:19 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2004.codfw.wmnet with reason: host reimage * 15:19 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 15:17 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1007.eqiad.wmnet * 15:12 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1007.eqiad.wmnet * 15:12 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1006.eqiad.wmnet * 15:12 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1006.eqiad.wmnet * 15:11 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1138.eqiad.wmnet * 15:11 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 15:11 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1138.eqiad.wmnet * 15:11 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1138.eqiad.wmnet * 15:10 brennen@deploy1003: Finished deploy [phabricator/deployment@f8b349f]: deploy phab1004 for [[phab:T433382|T433382]] (duration: 00m 43s) * 15:10 brennen@deploy1003: Started deploy [phabricator/deployment@f8b349f]: deploy phab1004 for [[phab:T433382|T433382]] * 15:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts * 15:09 brennen@deploy1003: Finished deploy [phabricator/deployment@f8b349f]: deploy phab2003 for [[phab:T433382|T433382]] (duration: 00m 55s) * 15:08 brennen@deploy1003: Started deploy [phabricator/deployment@f8b349f]: deploy phab2003 for [[phab:T433382|T433382]] * 15:07 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2012.codfw.wmnet, repooling source-only afterwards * 15:07 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts * 15:06 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1137.eqiad.wmnet * 15:06 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1137.eqiad.wmnet * 15:06 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1137.eqiad.wmnet * 15:05 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1006.eqiad.wmnet * 15:05 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1022\.eqiad\.wmnet,dc=eqiad,cluster=wdqs\-main,service=wdqs\-main * 15:01 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab2003.codfw.wmnet with reason: deployment * 15:01 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1005.eqiad.wmnet with reason: deployment * 15:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1006.eqiad.wmnet * 15:00 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1006.eqiad.wmnet with reason: deployment * 15:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1005.eqiad.wmnet * 15:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1005.eqiad.wmnet * 14:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db2248: Maintenance * 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1004.eqiad.wmnet with reason: deployment * 14:59 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2004.codfw.wmnet with OS bookworm * 14:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1005.eqiad.wmnet * 14:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2248 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95403 and previous config saved to /var/cache/conftool/dbconfig/20260728-145532-cwilliams.json * 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2245-2247].codfw.wmnet with reason: Maintenance * 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2248.codfw.wmnet with reason: Maintenance * 14:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2240: Maintenance * 14:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2173: Maintenance * 14:51 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts * 14:50 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1005.eqiad.wmnet * 14:50 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1004.eqiad.wmnet * 14:50 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1004.eqiad.wmnet * 14:49 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts * 14:45 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2004.codfw.wmnet * 14:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2173 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95399 and previous config saved to /var/cache/conftool/dbconfig/20260728-144453-cwilliams.json * 14:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2173.codfw.wmnet with reason: Maintenance * 14:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2170: Maintenance * 14:44 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1004.eqiad.wmnet * 14:39 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2004.codfw.wmnet * 14:38 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1004.eqiad.wmnet * 14:38 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet * 14:38 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet * 14:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db2195: Maintenance * 14:36 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1022.eqiad.wmnet, repooling source-only afterwards * 14:36 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 14:33 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet * 14:33 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2222: Maintenance * 14:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2195 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95394 and previous config saved to /var/cache/conftool/dbconfig/20260728-143218-cwilliams.json * 14:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2195.codfw.wmnet with reason: Maintenance * 14:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2181: Maintenance * 14:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 14:30 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 14:25 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 14:25 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 14:25 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 14:23 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet * 14:23 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1002.eqiad.wmnet * 14:23 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1002.eqiad.wmnet * 14:23 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 19s) * 14:23 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:18 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1002.eqiad.wmnet * 14:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2012.codfw.wmnet with OS bookworm * 14:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1002.eqiad.wmnet * 14:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1001.eqiad.wmnet * 14:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1001.eqiad.wmnet * 14:11 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 14:11 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 14:08 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1001.eqiad.wmnet * 14:07 elukey@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'. * 14:07 elukey@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'. * 14:06 elukey@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'. * 14:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db2240: Maintenance * 14:06 XioNoX: un-drain cr2-esams - [[phab:T431751|T431751]] * 14:05 elukey@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'. * 14:02 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1001.eqiad.wmnet * 14:02 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-eqiad * 14:01 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 14:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2240 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95384 and previous config saved to /var/cache/conftool/dbconfig/20260728-140011-cwilliams.json * 14:00 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2240.codfw.wmnet with reason: Maintenance * 13:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2237: Maintenance * 13:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db2170: Maintenance * 13:56 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 13:55 XioNoX: reboot cr2-esams - [[phab:T431751|T431751]] * 13:52 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 13:51 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr2-esams,cr2-esams IPv6,cr2-esams.mgmt with reason: router upgrade * 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 13:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2170 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95381 and previous config saved to /var/cache/conftool/dbconfig/20260728-135043-cwilliams.json * 13:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2170.codfw.wmnet with reason: Maintenance * 13:50 XioNoX: drain cr2-esams - [[phab:T431751|T431751]] * 13:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2153: Maintenance * 13:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2012.codfw.wmnet with reason: host reimage * 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2222: Maintenance * 13:45 sukhe: restart pybal on A:lvs-codfw * 13:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2012.codfw.wmnet with reason: host reimage * 13:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db2181: Maintenance * 13:44 btullis@dns1004: END - running authdns-update * 13:42 sukhe: restart pybal on lvs2014 * 13:42 btullis@dns1004: START - running authdns-update * 13:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2222 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95376 and previous config saved to /var/cache/conftool/dbconfig/20260728-133948-cwilliams.json * 13:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2222.codfw.wmnet with reason: Maintenance * 13:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2221: Maintenance * 13:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2181 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95374 and previous config saved to /var/cache/conftool/dbconfig/20260728-133857-cwilliams.json * 13:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2181.codfw.wmnet with reason: Maintenance * 13:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2167: Maintenance * 13:30 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T428495|T428495]] * 13:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 13:29 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 13:29 ayounsi@cumin1003: END (FAIL) - Cookbook sre.dns.admin (exit_code=99) DNS admin: depool esams [reason: router upgrade, [[phab:T431749|T431749]]] * 13:28 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: router upgrade, [[phab:T431749|T431749]]] * 13:27 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1022.eqiad.wmnet, repooling source-only afterwards * 13:27 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo - [[phab:T428495|T428495]] * 13:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2012 * 13:27 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2012 * 13:21 lucaswerkmeister-wmde@deploy1003: mwscript-k8s job started: cleanupTitles bolwiki # [[phab:T429951|T429951]] * 13:21 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] (duration: 07m 19s) * 13:20 swfrench-wmf: authdns-update to direct codfw, eqsin, ulsfo etcd clients to eqiad - [[phab:T428495|T428495]] * 13:18 swfrench@dns1004: END - running authdns-update * 13:17 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, anzx: Continuing with deployment * 13:16 swfrench@dns1004: START - running authdns-update * 13:16 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2012 * 13:16 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2012.codfw.wmnet 57.48.192.10.in-addr.arpa 7.5.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:16 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2012.codfw.wmnet 57.48.192.10.in-addr.arpa 7.5.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:16 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:16 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, anzx: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:14 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 13:14 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] * 13:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db2237: Maintenance * 13:13 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2237: Maintenance * 13:13 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:12 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:12 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback IPV6 for asw1-604 - pt1979@cumin2003" * 13:12 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:12 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback IPV6 for asw1-604 - pt1979@cumin2003" * 13:11 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 20s) * 13:11 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 13:10 esanders@deploy1003: Finished scap sync-world: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] (duration: 08m 11s) * 13:08 pt1979@cumin2003: START - Cookbook sre.dns.netbox * 13:07 root@cumin1003: START - Cookbook sre.mysql.pool pool db2237: Maintenance * 13:06 esanders@deploy1003: esanders: Continuing with deployment * 13:04 esanders@deploy1003: esanders: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2153: Maintenance * 13:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2153: Maintenance * 13:02 esanders@deploy1003: Started scap sync-world: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] * 13:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2237 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95362 and previous config saved to /var/cache/conftool/dbconfig/20260728-130107-cwilliams.json * 13:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2237.codfw.wmnet with reason: Maintenance * 13:00 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2236: Maintenance * 12:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2153: Maintenance * 12:52 root@cumin1003: START - Cookbook sre.mysql.pool pool db2221: Maintenance * 12:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2153 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95358 and previous config saved to /var/cache/conftool/dbconfig/20260728-125214-cwilliams.json * 12:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2153.codfw.wmnet with reason: Maintenance * 12:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2167: Maintenance * 12:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-codfw * 12:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2011.codfw.wmnet * 12:51 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2011.codfw.wmnet * 12:49 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:49 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback for asw1-603 - pt1979@cumin2003" * 12:48 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback for asw1-603 - pt1979@cumin2003" * 12:46 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2011.codfw.wmnet * 12:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2221 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95357 and previous config saved to /var/cache/conftool/dbconfig/20260728-124601-cwilliams.json * 12:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2221.codfw.wmnet with reason: Maintenance * 12:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2218: Maintenance * 12:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2167 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95354 and previous config saved to /var/cache/conftool/dbconfig/20260728-124457-cwilliams.json * 12:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2167.codfw.wmnet with reason: Maintenance * 12:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2166: Maintenance * 12:42 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 12:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2011.codfw.wmnet * 12:41 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2010.codfw.wmnet * 12:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2010.codfw.wmnet * 12:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2010.codfw.wmnet * 12:34 pt1979@cumin2003: START - Cookbook sre.dns.netbox * 12:32 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 12:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2010.codfw.wmnet * 12:31 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2009.codfw.wmnet * 12:31 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2009.codfw.wmnet * 12:27 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2009.codfw.wmnet * 12:22 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2009.codfw.wmnet * 12:22 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2008.codfw.wmnet * 12:21 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2008.codfw.wmnet * 12:16 pt1979@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-604-eqsin * 12:16 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2008.codfw.wmnet * 12:16 pt1979@cumin1003: START - Cookbook sre.network.tls for network device asw1-604-eqsin * 12:14 root@cumin1003: START - Cookbook sre.mysql.pool pool db2236: Maintenance * 12:14 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2236: Maintenance * 12:12 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1137.eqiad.wmnet with OS trixie * 12:11 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2008.codfw.wmnet * 12:11 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2007.codfw.wmnet * 12:11 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2007.codfw.wmnet * 12:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db2236: Maintenance * 12:06 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2007.codfw.wmnet * 12:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2236 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95348 and previous config saved to /var/cache/conftool/dbconfig/20260728-120253-cwilliams.json * 12:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2236.codfw.wmnet with reason: Maintenance * 12:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2007.codfw.wmnet * 12:01 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 12:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 11:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2218: Maintenance * 11:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db2166: Maintenance * 11:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2219: Maintenance * 11:56 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2006.codfw.wmnet * 11:52 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1137.eqiad.wmnet with reason: host reimage * 11:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2218 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95344 and previous config saved to /var/cache/conftool/dbconfig/20260728-115155-cwilliams.json * 11:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2218.codfw.wmnet with reason: Maintenance * 11:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2208: Maintenance * 11:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2166 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95342 and previous config saved to /var/cache/conftool/dbconfig/20260728-115119-cwilliams.json * 11:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2166.codfw.wmnet with reason: Maintenance * 11:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2164: Maintenance * 11:47 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1137.eqiad.wmnet with reason: host reimage * 11:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2006.codfw.wmnet * 11:45 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2005.codfw.wmnet * 11:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2005.codfw.wmnet * 11:40 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2005.codfw.wmnet * 11:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2005.codfw.wmnet * 11:35 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2004.codfw.wmnet * 11:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2004.codfw.wmnet * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1137 * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1137 * 11:30 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1137 * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1137.eqiad.wmnet 192.32.64.10.in-addr.arpa 2.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:30 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1137.eqiad.wmnet 192.32.64.10.in-addr.arpa 2.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1137 - jiji@cumin1003" * 11:25 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2004.codfw.wmnet * 11:19 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2004.codfw.wmnet * 11:19 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2003.codfw.wmnet * 11:19 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2003.codfw.wmnet * 11:14 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2003.codfw.wmnet * 11:11 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2219: Maintenance * 11:10 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2219: Maintenance * 11:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2219: Maintenance * 11:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2003.codfw.wmnet * 11:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2164: Maintenance * 11:03 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2002.codfw.wmnet * 11:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2002.codfw.wmnet * 11:03 root@cumin1003: START - Cookbook sre.mysql.pool pool db2208: Maintenance * 10:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2164 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95332 and previous config saved to /var/cache/conftool/dbconfig/20260728-105749-cwilliams.json * 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2164.codfw.wmnet with reason: Maintenance * 10:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2208 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95331 and previous config saved to /var/cache/conftool/dbconfig/20260728-105711-cwilliams.json * 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2208.codfw.wmnet with reason: Maintenance * 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2219 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95330 and previous config saved to /var/cache/conftool/dbconfig/20260728-105652-cwilliams.json * 10:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2219.codfw.wmnet with reason: Maintenance * 10:53 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1137 - jiji@cumin1003" * 10:52 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2002.codfw.wmnet * 10:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2002.codfw.wmnet * 10:47 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2001.codfw.wmnet * 10:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2001.codfw.wmnet * 10:39 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2001.codfw.wmnet * 10:35 jiji@cumin1003: START - Cookbook sre.dns.netbox * 10:34 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] (duration: 09m 31s) * 10:34 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1137 * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2001.codfw.wmnet * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-codfw * 10:34 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1137.eqiad.wmnet with OS trixie * 10:28 jforrester@deploy1003: jforrester: Continuing with deployment * 10:27 jforrester@deploy1003: jforrester: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:25 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] * 10:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 10:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 10:21 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 10:20 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1137.eqiad.wmnet * 10:20 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 10:20 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1137.eqiad.wmnet * 10:20 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1137.eqiad.wmnet * 10:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2163: Maintenance * 09:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-staging-worker * 09:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2003.codfw.wmnet * 09:37 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2003.codfw.wmnet * 09:32 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 09:31 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2003.codfw.wmnet * 09:30 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 09:30 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2163: Maintenance * 09:30 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 09:30 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 09:30 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:22 klausman@cumin1003: END (ERROR) - Cookbook sre.ganeti.reboot-vm (exit_code=97) for VM ml-serve-ctrl2001.codfw.wmnet * 09:22 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2001.codfw.wmnet * 09:22 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d8-eqiad * 09:22 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d8-eqiad * 09:21 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2003.codfw.wmnet * 09:20 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2002.codfw.wmnet * 09:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2002.codfw.wmnet * 09:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f2-codfw * 09:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f2-codfw * 09:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e4-codfw * 09:18 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2163: Maintenance * 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e4-codfw * 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-codfw * 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-codfw * 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e5-codfw * 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e5-codfw * 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f4-codfw * 09:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2210: Maintenance * 09:16 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f4-codfw * 09:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2182: Maintenance * 09:14 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2002.codfw.wmnet * 09:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db2163: Maintenance * 09:11 XioNoX: rebooting cr2-drmrs - [[phab:T431749|T431749]] * 09:10 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr2-drmrs,cr2-drmrs IPv6,cr2-drmrs.mgmt with reason: router upgrade * 09:06 XioNoX: draining cr2-drmrs - [[phab:T431749|T431749]] * 09:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2163 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95320 and previous config saved to /var/cache/conftool/dbconfig/20260728-090638-cwilliams.json * 09:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2163.codfw.wmnet with reason: Maintenance * 09:06 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2161: Maintenance * 09:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2002.codfw.wmnet * 09:04 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2001.codfw.wmnet * 09:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2001.codfw.wmnet * 08:57 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2001.codfw.wmnet * 08:48 XioNoX: un-drain cr1-drmrs - [[phab:T431749|T431749]] * 08:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2001.codfw.wmnet * 08:47 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-staging-worker * 08:42 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:35 XioNoX: rebooting cr1-drmrs - [[phab:T431749|T431749]] * 08:33 XioNoX: draining cr1-drmrs - [[phab:T431749|T431749]] * 08:31 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2210: Maintenance * 08:29 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2182: Maintenance * 08:21 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2182: Maintenance * 08:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db2161: Maintenance * 08:16 root@cumin1003: START - Cookbook sre.mysql.pool pool db2182: Maintenance * 08:12 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2210: Maintenance * 08:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2161 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95309 and previous config saved to /var/cache/conftool/dbconfig/20260728-081044-cwilliams.json * 08:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2161.codfw.wmnet with reason: Maintenance * 08:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2154: Maintenance * 08:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2182 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95307 and previous config saved to /var/cache/conftool/dbconfig/20260728-080947-cwilliams.json * 08:09 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2182.codfw.wmnet with reason: Maintenance * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2168: Maintenance * 08:06 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr1-drmrs,cr1-drmrs IPv6,cr1-drmrs.mgmt with reason: router upgrade * 08:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db2210: Maintenance * 08:05 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 08:05 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 08:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2210 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95305 and previous config saved to /var/cache/conftool/dbconfig/20260728-080008-cwilliams.json * 08:00 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2210.codfw.wmnet with reason: Maintenance * 07:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2206: Maintenance * 07:50 gkyziridis@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 07:50 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 07:22 root@cumin1003: START - Cookbook sre.mysql.pool pool db2154: Maintenance * 07:22 root@cumin1003: START - Cookbook sre.mysql.pool pool db2168: Maintenance * 07:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2154 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95295 and previous config saved to /var/cache/conftool/dbconfig/20260728-071640-cwilliams.json * 07:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2154.codfw.wmnet with reason: Maintenance * 07:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95294 and previous config saved to /var/cache/conftool/dbconfig/20260728-071604-cwilliams.json * 07:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2168.codfw.wmnet with reason: Maintenance * 07:08 root@cumin1003: START - Cookbook sre.mysql.pool pool db2206: Maintenance * 07:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2206 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95292 and previous config saved to /var/cache/conftool/dbconfig/20260728-070219-cwilliams.json * 07:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2206.codfw.wmnet with reason: Maintenance * 06:44 marostegui: Failover m5 from db1164 to db1228 - [[phab:T432967|T432967]] * 06:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2235].codfw.wmnet,db[1164,1217,1228].eqiad.wmnet with reason: m5 master switch [[phab:T432967|T432967]] * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.10 (duration: 02m 34s) * 03:39 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] (duration: 36m 06s) * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 02:57 dzahn@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1004.eqiad.wmnet with OS trixie * 02:57 dzahn@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - dzahn@cumin1003" * 02:55 dzahn@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - dzahn@cumin1003" * 02:37 dzahn@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1004.eqiad.wmnet with reason: host reimage * 02:31 dzahn@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1004.eqiad.wmnet with reason: host reimage * 02:16 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie * 02:15 dzahn@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host zuul1004.eqiad.wmnet with OS trixie * 01:43 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie * 01:43 dzahn@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1004.eqiad.wmnet with OS trixie * 01:25 pt1979@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-603-eqsin * 01:24 pt1979@cumin1003: START - Cookbook sre.network.tls for network device asw1-603-eqsin * 01:12 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 01:12 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt for new switches in eqsin - pt1979@cumin2003" * 01:12 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt for new switches in eqsin - pt1979@cumin2003" * 01:08 pt1979@cumin2003: START - Cookbook sre.dns.netbox * 00:48 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 00:47 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 00:47 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 00:47 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 00:26 mutante: attempting reimage with trixie on zuul1004 re-purposed physical hardware - dcops reported install issue - host was in busybox shell ([[phab:T427353|T427353]]) * 00:24 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie == 2026-07-27 == * 23:50 Amir1: mass deleting vp8 transcodes * 23:28 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:27 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1004.eqiad.wmnet with OS bullseye * 23:26 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 23:25 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:25 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 22:39 maryum: Deploy security fix for [[phab:T432877|T432877]] * 22:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1022.eqiad.wmnet with OS bookworm * 22:37 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS bullseye * 22:32 sbassett: Deployed security fix for [[phab:T432789|T432789]] * 22:22 sbassett: Deployed security patch for [[phab:T431819|T431819]] * 22:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1022.eqiad.wmnet with reason: host reimage * 22:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1022.eqiad.wmnet with reason: host reimage * 22:01 RScout-WMF: Deployed security fix for [[phab:T431819|T431819]] * 22:00 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2012 * 21:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2012.codfw.wmnet with OS bookworm * 21:55 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2011\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 21:45 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1022 * 21:45 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1022 * 21:44 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1022 * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1022.eqiad.wmnet 239.48.64.10.in-addr.arpa 9.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:44 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1022.eqiad.wmnet 239.48.64.10.in-addr.arpa 9.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:41 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:41 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 21:34 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS bookworm * 21:31 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:22 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:21 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:19 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:17 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1004 * 21:16 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1004 * 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1004] - vriley@cumin1003" * 21:15 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1004] - vriley@cumin1003" * 21:11 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:10 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2011.codfw.wmnet, repooling source-only afterwards * 21:05 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:01 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1022 * 20:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1022.eqiad.wmnet with OS bookworm * 20:53 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1021.eqiad.wmnet, repooling source-only afterwards * 20:51 mutante: zuul1001 - re-enabled puppet - revert "cherry-picked" gerrit:1314120 - [[phab:T431003|T431003]] * 20:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Maintenance * 20:15 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] (duration: 08m 03s) * 20:11 sbisson@deploy1003: sbisson: Continuing with deployment * 20:09 sbisson@deploy1003: sbisson: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] * 19:47 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 46s) * 19:47 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Maintenance * 19:27 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] (duration: 12m 26s) * 19:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2228 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95285 and previous config saved to /var/cache/conftool/dbconfig/20260727-192711-cwilliams.json * 19:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2228.codfw.wmnet with reason: Maintenance * 19:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2223: Maintenance * 19:23 krinkle@deploy1003: krinkle: Continuing with deployment * 19:16 krinkle@deploy1003: krinkle: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:15 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] * 19:12 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2238: Maintenance * 18:58 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 18:57 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 18:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2227: Maintenance * 18:57 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-experimental: apply * 18:55 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-experimental: apply * 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1021.eqiad.wmnet with OS bookworm * 18:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2011.codfw.wmnet with OS bookworm * 18:40 root@cumin1003: START - Cookbook sre.mysql.pool pool db2223: Maintenance * 18:39 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] (duration: 07m 05s) * 18:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2223 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95275 and previous config saved to /var/cache/conftool/dbconfig/20260727-183500-cwilliams.json * 18:34 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2223.codfw.wmnet with reason: Maintenance * 18:34 musikanimal@deploy1003: musikanimal: Continuing with deployment * 18:34 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2213: Maintenance * 18:33 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:32 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] * 18:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db2238: Maintenance * 18:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2011.codfw.wmnet with reason: host reimage * 18:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2238 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95271 and previous config saved to /var/cache/conftool/dbconfig/20260727-181944-cwilliams.json * 18:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2238.codfw.wmnet with reason: Maintenance * 18:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2226: Maintenance * 18:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1021.eqiad.wmnet with reason: host reimage * 18:14 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2011.codfw.wmnet with reason: host reimage * 18:12 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1021.eqiad.wmnet with reason: host reimage * 18:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db2227: Maintenance * 18:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2227 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95265 and previous config saved to /var/cache/conftool/dbconfig/20260727-180256-cwilliams.json * 18:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2227.codfw.wmnet with reason: Maintenance * 18:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2194: Maintenance * 17:57 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2011 * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2011 * 17:56 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2011 * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2011.codfw.wmnet 37.32.192.10.in-addr.arpa 7.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:56 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2011.codfw.wmnet 37.32.192.10.in-addr.arpa 7.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2011 - bking@cumin2003" * 17:56 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2011 - bking@cumin2003" * 17:52 bking@cumin2003: START - Cookbook sre.dns.netbox * 17:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2011 * 17:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1021 * 17:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1021 * 17:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2011.codfw.wmnet with OS bookworm * 17:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1021.eqiad.wmnet with OS bookworm * 17:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Maintenance * 17:38 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2010\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 17:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2213 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95260 and previous config saved to /var/cache/conftool/dbconfig/20260727-173740-cwilliams.json * 17:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2213.codfw.wmnet with reason: Maintenance * 17:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2211: Maintenance * 17:36 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1020\.eqiad\.wmnet,dc=eqiad,cluster=wdqs\-main,service=wdqs\-main * 17:32 root@cumin1003: START - Cookbook sre.mysql.pool pool db2226: Maintenance * 17:31 taavi@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] (duration: 06m 33s) * 17:27 taavi@deploy1003: taavi: Continuing with deployment * 17:27 taavi@deploy1003: taavi: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:26 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2226 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95256 and previous config saved to /var/cache/conftool/dbconfig/20260727-172636-cwilliams.json * 17:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2226.codfw.wmnet with reason: Maintenance * 17:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2225: Maintenance * 17:25 taavi@deploy1003: Started scap sync-world: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] * 17:13 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 17:11 root@cumin1003: START - Cookbook sre.mysql.pool pool db2194: Maintenance * 17:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2194 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95248 and previous config saved to /var/cache/conftool/dbconfig/20260727-170453-cwilliams.json * 17:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2194.codfw.wmnet with reason: Maintenance * 17:04 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2190: Maintenance * 16:52 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 16:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2211: Maintenance * 16:40 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95242 and previous config saved to /var/cache/conftool/dbconfig/20260727-164015-cwilliams.json * 16:40 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2211.codfw.wmnet with reason: Maintenance * 16:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2178: Maintenance * 16:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db2225: Maintenance * 16:39 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2172: Maintenance * 16:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 16:38 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 16:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2225 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95238 and previous config saved to /var/cache/conftool/dbconfig/20260727-163307-cwilliams.json * 16:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2225.codfw.wmnet with reason: Maintenance * 16:32 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2189: Maintenance * 16:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db2190: Maintenance * 16:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2190 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95230 and previous config saved to /var/cache/conftool/dbconfig/20260727-160602-cwilliams.json * 16:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2190.codfw.wmnet with reason: Maintenance * 15:53 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2177: Maintenance * 15:53 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2172: Maintenance * 15:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2178: Maintenance * 15:51 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2172: Maintenance * 15:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2178 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95224 and previous config saved to /var/cache/conftool/dbconfig/20260727-154559-cwilliams.json * 15:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2172: Maintenance * 15:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2178.codfw.wmnet with reason: Maintenance * 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2171: Maintenance * 15:44 root@cumin1003: START - Cookbook sre.mysql.pool pool db2189: Maintenance * 15:43 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:41 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2172 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95222 and previous config saved to /var/cache/conftool/dbconfig/20260727-153927-cwilliams.json * 15:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2172.codfw.wmnet with reason: Maintenance * 15:38 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2189 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95220 and previous config saved to /var/cache/conftool/dbconfig/20260727-153833-cwilliams.json * 15:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2189.codfw.wmnet with reason: Maintenance * 15:34 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:32 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 15:32 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 15:31 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:29 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:26 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 15:22 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] (duration: 07m 00s) * 15:21 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2155: Maintenance * 15:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2175: Maintenance * 15:18 zabe@deploy1003: zabe: Continuing with deployment * 15:17 zabe@deploy1003: zabe: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:15 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 15:15 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:15 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] * 15:15 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 06s) * 15:15 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:12 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2010.codfw.wmnet with OS bookworm * 15:08 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2177: Maintenance * 15:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2177: Maintenance * 14:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2177: Maintenance * 14:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2171: Maintenance * 14:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2171 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95209 and previous config saved to /var/cache/conftool/dbconfig/20260727-145236-cwilliams.json * 14:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2171.codfw.wmnet with reason: Maintenance * 14:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2177 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95208 and previous config saved to /var/cache/conftool/dbconfig/20260727-145206-cwilliams.json * 14:52 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2157: Maintenance * 14:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2177.codfw.wmnet with reason: Maintenance * 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1020.eqiad.wmnet with OS bookworm * 14:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2156: Maintenance * 14:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2010.codfw.wmnet with reason: host reimage * 14:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2010.codfw.wmnet with reason: host reimage * 14:41 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[1031,2024]*: Upgrade Cassandra to 5.0.8 (canary) - eevans@cumin1003 * 14:34 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2155: Maintenance * 14:33 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2175: Maintenance * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2010 * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2010 * 14:24 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2010 * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2010.codfw.wmnet 94.16.192.10.in-addr.arpa 4.9.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2010.codfw.wmnet 94.16.192.10.in-addr.arpa 4.9.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2010 - bking@cumin2003" * 14:24 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2010 - bking@cumin2003" * 14:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1020.eqiad.wmnet with reason: host reimage * 14:23 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[1031,2024]*: Upgrade Cassandra to 5.0.8 (canary) - eevans@cumin1003 * 14:20 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 14:20 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 14:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1020.eqiad.wmnet with reason: host reimage * 14:17 sukhe: sudo gnt-instance reboot urldownloader1005.wikimedia.org * 14:16 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:15 jelto@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:08 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2155: Maintenance * 14:05 root@cumin1003: START - Cookbook sre.mysql.pool pool db2157: Maintenance * 14:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2175: Maintenance * 14:03 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 11 hosts * 14:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db2155: Maintenance * 14:01 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 11 hosts * 14:01 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1136.eqiad.wmnet * 14:01 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1136.eqiad.wmnet * 14:01 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1136.eqiad.wmnet * 14:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db2156: Maintenance * 13:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2157 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95194 and previous config saved to /var/cache/conftool/dbconfig/20260727-135943-cwilliams.json * 13:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2157.codfw.wmnet with reason: Maintenance * 13:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db2175: Maintenance * 13:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 13:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 13:57 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2010 * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2155 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95193 and previous config saved to /var/cache/conftool/dbconfig/20260727-135613-cwilliams.json * 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2155.codfw.wmnet with reason: Maintenance * 13:55 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1020 * 13:55 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1020 * 13:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2156 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95192 and previous config saved to /var/cache/conftool/dbconfig/20260727-135413-cwilliams.json * 13:54 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2156.codfw.wmnet with reason: Maintenance * 13:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2175 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95191 and previous config saved to /var/cache/conftool/dbconfig/20260727-135300-cwilliams.json * 13:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2010.codfw.wmnet with OS bookworm * 13:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2175.codfw.wmnet with reason: Maintenance * 13:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1020.eqiad.wmnet with OS bookworm * 13:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 34 hosts * 13:46 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 34 hosts * 13:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1201: Maintenance * 13:27 Lucas_WMDE: UTC afternoon backport+config window doen * 13:18 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] (duration: 11m 57s) * 13:14 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, sihe: Continuing with deployment * 13:08 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, sihe: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:07 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool ulsfo [reason: router upgrade finished, [[phab:T431752|T431752]]] * 13:07 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool ulsfo [reason: router upgrade finished, [[phab:T431752|T431752]]] * 13:06 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] * 13:03 XioNoX: repool cr4-ulsfo - [[phab:T431752|T431752]] * 12:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db1201: Maintenance * 12:48 gkyziridis@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1201 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95186 and previous config saved to /var/cache/conftool/dbconfig/20260727-124404-cwilliams.json * 12:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1201.eqiad.wmnet with reason: Maintenance * 12:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1187: Maintenance * 12:30 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader.eqiad.wikimedia.org on all recursors * 12:30 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader.eqiad.wikimedia.org on all recursors * 12:30 sukhe@dns1004: END - running authdns-update * 12:30 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] (duration: 09m 32s) * 12:28 sukhe@dns1004: START - running authdns-update * 12:25 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 12:22 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:20 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] * 12:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts * 12:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts * 12:13 XioNoX: rebooting cr4-ulsfo for upgrade - [[phab:T431752|T431752]] * 12:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: es1038 repool * 12:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 38 hosts * 12:08 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 38 hosts * 11:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1187: Maintenance * 11:53 urbanecm@deploy1003: mwscript-k8s job started: foreachwikiindblist growthexperiments GrowthExperiments:cleanMentorList # [[phab:T431804|T431804]] * 11:50 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr4-ulsfo,cr4-ulsfo IPv6,cr4-ulsfo.mgmt with reason: router upgrade * 11:50 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] (duration: 11m 07s) * 11:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1187 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95178 and previous config saved to /var/cache/conftool/dbconfig/20260727-114844-cwilliams.json * 11:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1187.eqiad.wmnet with reason: Maintenance * 11:43 urbanecm@deploy1003: urbanecm: Continuing with deployment * 11:42 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:39 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] * 11:37 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 11:36 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 11:36 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 11:35 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 11:29 XioNoX: start draining cr4-ulsfo - [[phab:T431752|T431752]] * 11:29 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 11:29 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 11:28 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1035: testing * 11:28 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1035: testing * 11:27 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1035: testing * 11:27 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1035: testing * 11:26 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool ulsfo [reason: router upgrade, [[phab:T431752|T431752]]] * 11:26 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1038: es1038 repool * 11:26 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool ulsfo [reason: router upgrade, [[phab:T431752|T431752]]] * 11:26 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1038: testing * 11:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1264: Maintenance * 11:24 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1038: testing * 11:23 marostegui@cumin1003: dbctl commit (dc=all): 'Repool es1050 as master', diff saved to https://phabricator.wikimedia.org/P95170 and previous config saved to /var/cache/conftool/dbconfig/20260727-112326-marostegui.json * 11:23 marostegui@cumin1003: dbctl commit (dc=all): 'Repool es1050', diff saved to https://phabricator.wikimedia.org/P95169 and previous config saved to /var/cache/conftool/dbconfig/20260727-112302-marostegui.json * 11:22 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1050: testing * 11:22 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1050: testing * 11:20 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 11:18 blake@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 11:18 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 11:12 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 11:11 blake@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 11:09 blake@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 11:09 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 11:09 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 11:08 blake@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 11:05 blake@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 11:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 11:02 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 10:50 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 10:43 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:39 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply * 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1264: Maintenance * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply * 10:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply * 10:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 10:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 10:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 10:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1264 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95164 and previous config saved to /var/cache/conftool/dbconfig/20260727-103204-cwilliams.json * 10:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1264.eqiad.wmnet with reason: Maintenance * 10:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 10:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 10:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 10:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 10:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 10:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 10:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 10:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1237: Maintenance * 10:04 elukey: restart burrow main-eqiad on kafkamon2003 to clear some errors on kafka-main1008 * 09:58 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1136.eqiad.wmnet with OS trixie * 09:39 elukey: restart burrow-main-eqiad.service on kafkamon1003 to see if a recurrent kafka error on kafka-main1008 goes away * 09:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1237: Maintenance * 09:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1136.eqiad.wmnet with reason: host reimage * 09:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1237 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95159 and previous config saved to /var/cache/conftool/dbconfig/20260727-093328-cwilliams.json * 09:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1237.eqiad.wmnet with reason: Maintenance * 09:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1136.eqiad.wmnet with reason: host reimage * 09:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1203: Maintenance * 09:17 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1136 * 09:17 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1136 * 09:04 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1136 * 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1136.eqiad.wmnet 191.32.64.10.in-addr.arpa 1.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:04 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1136.eqiad.wmnet 191.32.64.10.in-addr.arpa 1.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1136 - jiji@cumin1003" * 09:04 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1136 - jiji@cumin1003" * 08:52 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 08:52 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 08:52 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 08:51 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 08:50 jiji@cumin1003: START - Cookbook sre.dns.netbox * 08:47 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1136 * 08:46 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1136.eqiad.wmnet with OS trixie * 08:44 marostegui: Rename tables on s3 [[phab:T425066|T425066]] * 08:43 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1136.eqiad.wmnet * 08:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db1203: Maintenance * 08:43 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1136.eqiad.wmnet * 08:43 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1136.eqiad.wmnet * 08:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1203 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95154 and previous config saved to /var/cache/conftool/dbconfig/20260727-083703-cwilliams.json * 08:36 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1203.eqiad.wmnet with reason: Maintenance * 08:16 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1179: Maintenance * 07:44 phuedx: UTC morning backport window done * 07:37 phuedx@deploy1003: Finished scap sync-world: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] (duration: 32m 33s) * 07:28 root@cumin1003: START - Cookbook sre.mysql.pool pool db1179: Maintenance * 07:26 marostegui: Rename tables on s3 [[phab:T426341|T426341]] * 07:25 phuedx@deploy1003: phuedx: Continuing with deployment * 07:22 marostegui: Drop tables in akwiki nawiki pihwiki - growthexperiments_* [[phab:T428885|T428885]] * 07:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95149 and previous config saved to /var/cache/conftool/dbconfig/20260727-072234-cwilliams.json * 07:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1179.eqiad.wmnet with reason: Maintenance * 07:20 phuedx@deploy1003: phuedx: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:16 ryankemper: [[phab:T430880|T430880]] [WDQS] Reimaged `wdqs1018` and `wdqs1019` to Bookworm, restored data using test-cookbook change {{Gerrit|1317128}}, and repooled both; 25/36 hosts complete * 07:04 phuedx@deploy1003: Started scap sync-world: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] * 06:57 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1019.eqiad.wmnet * 06:56 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1018.eqiad.wmnet * 06:40 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1020.eqiad.wmnet with reason: Cloning * 06:35 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db1228.eqiad.wmnet with reason: Rebooting * 06:29 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:29 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:25 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:25 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:25 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1019.eqiad.wmnet, repooling source-only afterwards * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1018.eqiad.wmnet, repooling source-only afterwards * 04:51 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1019.eqiad.wmnet, repooling source-only afterwards * 04:51 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1018.eqiad.wmnet, repooling source-only afterwards * 04:48 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s) * 04:48 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 04:48 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s) * 04:48 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 36s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-26 == * 14:59 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:59 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:59 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:59 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1019.eqiad.wmnet with OS bookworm * 01:05 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1018.eqiad.wmnet with OS bookworm * 00:43 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1019.eqiad.wmnet with reason: host reimage * 00:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1018.eqiad.wmnet with reason: host reimage * 00:34 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1019.eqiad.wmnet with reason: host reimage * 00:33 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1018.eqiad.wmnet with reason: host reimage * 00:16 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 00:16 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 00:15 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 00:15 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1019 * 00:11 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1019 * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1018 * 00:11 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1018 * 00:08 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1019.eqiad.wmnet with OS bookworm * 00:08 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1018.eqiad.wmnet with OS bookworm == 2026-07-25 == * 22:06 ryankemper: [[phab:T430880|T430880]] [WDQS] Repooled `wdqs1017` and `wdqs2024` after reimaging to bookworm, scap deploying, and data xfering * 22:04 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2024.codfw.wmnet * 22:03 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1017.eqiad.wmnet * 21:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1017.eqiad.wmnet, repooling source-only afterwards * 21:06 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2024.codfw.wmnet, repooling source-only afterwards * 20:52 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:52 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:52 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:52 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 20:18 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1017.eqiad.wmnet, repooling source-only afterwards * 20:18 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2024.codfw.wmnet, repooling source-only afterwards * 20:15 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:15 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:15 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:15 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 19:57 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s) * 19:57 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 19:57 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 07s) * 19:57 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 19:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2024.codfw.wmnet with OS bookworm * 19:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1017.eqiad.wmnet with OS bookworm * 19:02 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2024.codfw.wmnet with reason: host reimage * 18:58 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1017.eqiad.wmnet with reason: host reimage * 18:53 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2024.codfw.wmnet with reason: host reimage * 18:52 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1017.eqiad.wmnet with reason: host reimage * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2024 * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2024 * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1017 * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1017 * 18:27 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2024 * 18:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2024.codfw.wmnet 58.16.192.10.in-addr.arpa 8.5.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:26 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2024.codfw.wmnet 58.16.192.10.in-addr.arpa 8.5.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:24 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1017 * 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1017.eqiad.wmnet 238.48.64.10.in-addr.arpa 8.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:24 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1017.eqiad.wmnet 238.48.64.10.in-addr.arpa 8.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1017 - ryankemper@cumin2003" * 18:24 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1017 - ryankemper@cumin2003" * 18:23 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 18:18 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 18:17 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1017 * 18:17 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2024 * 18:14 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1017.eqiad.wmnet with OS bookworm * 18:14 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2024.codfw.wmnet with OS bookworm * 18:05 ryankemper: [WDQS] [[phab:T430880|T430880]] Reimaged `wdqs1016` and `wdqs2023` to Bookworm with `--move-vlan`, restored main and scholarly data, validated postflights, and repooled both hosts. Confirmed PyBal rebuilt both backends with their new addresses * 17:45 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2023.codfw.wmnet * 17:43 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1016.eqiad.wmnet * 06:35 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1016.eqiad.wmnet, repooling source-only afterwards * 06:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2023.codfw.wmnet, repooling source-only afterwards * 05:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2023.codfw.wmnet, repooling source-only afterwards * 05:19 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1016.eqiad.wmnet, repooling source-only afterwards * 05:07 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 07s) * 05:07 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 05:06 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 06s) * 05:06 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 03:27 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2023.codfw.wmnet with OS bookworm * 02:59 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2023.codfw.wmnet with reason: host reimage * 02:56 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2023.codfw.wmnet with reason: host reimage * 02:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2023 * 02:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2023 * 02:30 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2023 * 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2023.codfw.wmnet 35.0.192.10.in-addr.arpa 5.3.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 02:30 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2023.codfw.wmnet 35.0.192.10.in-addr.arpa 5.3.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2023 - ryankemper@cumin2003" * 02:30 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2023 - ryankemper@cumin2003" * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 26s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:15 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1016.eqiad.wmnet with OS bookworm * 00:49 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1016.eqiad.wmnet with reason: host reimage * 00:43 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1016.eqiad.wmnet with reason: host reimage * 00:31 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 00:27 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1016 * 00:27 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1016 * 00:27 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2023 * 00:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1016.eqiad.wmnet with OS bookworm * 00:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2023.codfw.wmnet with OS bookworm * 00:11 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs1014.eqiad.wmnet and wdqs2008.codfw.wmnet after Bookworm reimage, transfer, and postflight; wdqs2008 is serving, while wdqs1014 will remain outside of service until a pybal restart next monday == 2026-07-24 == * 23:54 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1014.eqiad.wmnet * 23:54 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2008.codfw.wmnet * 23:43 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2010.codfw.wmnet with OS trixie * 23:08 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 23:03 jhathaway@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 22:33 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 22:13 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 22:13 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 22:13 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 22:13 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:00 jhathaway@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 21:53 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 21:53 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie * 21:51 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 21:47 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie * 21:43 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 21:39 jhathaway@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 21:38 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 17:21 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1135.eqiad.wmnet * 17:21 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1135.eqiad.wmnet * 17:21 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1135.eqiad.wmnet * 16:34 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 16:34 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 16:34 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 16:34 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 16:33 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 16:33 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 16:28 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:28 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:28 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:28 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2008.codfw.wmnet, repooling source-only afterwards * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1014.eqiad.wmnet, repooling source-only afterwards * 15:56 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1135.eqiad.wmnet with OS trixie * 15:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 40 hosts * 15:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 40 hosts * 15:37 topranks: upgrade SR-Linux OS on lswtest-d8-eqiad * 15:36 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1135.eqiad.wmnet with reason: host reimage * 15:33 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 6 hosts with reason: upgrade lswtest-d8-eqiad * 15:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1135.eqiad.wmnet with reason: host reimage * 15:30 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc-gp2006.codfw.wmnet with OS bookworm * 15:15 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1135 * 15:15 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1135 * 15:13 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc-gp2006.codfw.wmnet with reason: host reimage * 15:08 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc-gp2006.codfw.wmnet with reason: host reimage * 14:49 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm * 14:48 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host mc-gp2006.codfw.wmnet with OS bookworm * 14:34 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] (duration: 41m 12s) * 14:32 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1135 * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1135.eqiad.wmnet 177.32.64.10.in-addr.arpa 7.7.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:32 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1135.eqiad.wmnet 177.32.64.10.in-addr.arpa 7.7.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1135 - jiji@cumin1003" * 14:32 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1135 - jiji@cumin1003" * 14:29 krinkle@deploy1003: krinkle: Continuing with deployment * 14:29 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm * 14:27 jiji@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host mc-gp2006.codfw.wmnet with OS bookworm * 14:26 jiji@cumin1003: START - Cookbook sre.dns.netbox * 14:15 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1135 * 14:14 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1135.eqiad.wmnet with OS trixie * 14:14 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1135.eqiad.wmnet * 14:13 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1135.eqiad.wmnet * 14:13 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1135.eqiad.wmnet * 14:10 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1072.eqiad.wmnet * 14:10 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1072.eqiad.wmnet * 14:10 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1072.eqiad.wmnet * 14:10 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1072.eqiad.wmnet * 14:09 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1071.eqiad.wmnet * 14:09 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1071.eqiad.wmnet * 14:09 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1071.eqiad.wmnet * 14:09 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1071.eqiad.wmnet * 13:58 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 13:58 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:58 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:57 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:55 krinkle@deploy1003: krinkle: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:53 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] * 13:45 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:45 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push new IPs for mc-gp2006 - cmooney@cumin1003" * 13:45 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push new IPs for mc-gp2006 - cmooney@cumin1003" * 13:44 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) mc-gp2006.codfw.wmnet on all recursors * 13:44 cmooney@cumin1003: START - Cookbook sre.dns.wipe-cache mc-gp2006.codfw.wmnet on all recursors * 13:42 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm * 13:41 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:30 papaul: reboot mr1-eqsin for maintenance * 13:24 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb[1029-1031].eqiad.wmnet * 13:10 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb[1029-1031].eqiad.wmnet * 11:33 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 7 hosts * 11:11 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 7 hosts * 10:56 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 7 hosts * 10:47 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 7 hosts * 10:44 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:44 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:41 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:41 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:35 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 8 hosts * 10:34 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:33 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:32 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:32 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:31 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:31 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:30 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 8 hosts * 10:24 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie * 10:19 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:18 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 16 hosts * 10:17 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2001.codfw.wmnet * 10:13 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2001.codfw.wmnet * 10:12 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2001.codfw.wmnet * 10:02 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2001.codfw.wmnet * 10:02 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2002.codfw.wmnet * 09:57 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2002.codfw.wmnet * 09:56 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2002.codfw.wmnet * 09:51 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2002.codfw.wmnet * 09:51 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1002.eqiad.wmnet * 09:47 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1002.eqiad.wmnet * 09:47 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1001.eqiad.wmnet * 09:44 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1001.eqiad.wmnet * 09:34 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2003.codfw.wmnet * 09:32 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2003.codfw.wmnet * 09:32 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2002.codfw.wmnet * 09:30 brouberol@dns1004: END - running authdns-update * 09:29 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2002.codfw.wmnet * 09:29 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2001.codfw.wmnet * 09:27 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 9 hosts * 09:27 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2001.codfw.wmnet * 09:27 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2001.codfw.wmnet * 09:26 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 9 hosts * 09:26 brouberol@dns1004: START - running authdns-update * 09:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 57 hosts * 09:24 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2001.codfw.wmnet * 09:24 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2002.codfw.wmnet * 09:22 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2002.codfw.wmnet * 09:21 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 57 hosts * 09:20 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2003.codfw.wmnet * 09:19 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 16 hosts * 09:16 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2003.codfw.wmnet * 09:16 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1003.eqiad.wmnet * 09:15 urbanecm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 09:15 urbanecm@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 09:13 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1003.eqiad.wmnet * 09:13 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1002.eqiad.wmnet * 09:11 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1002.eqiad.wmnet * 09:11 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1001.eqiad.wmnet * 09:07 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1001.eqiad.wmnet * 08:32 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:24 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:16 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 08:16 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 08:07 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:07 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:07 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 08:02 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 08:01 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:59 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:57 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 07:57 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 06:46 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1025.eqiad.wmnet with reason: Cloning * 06:46 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s4 * 06:45 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s6 * 06:44 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1019.eqiad.wmnet,service=s6 * 06:44 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1019.eqiad.wmnet,service=s4 * 03:40 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:40 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:40 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:40 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 03:37 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:37 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:37 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:36 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:49 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on mr1-eqsin,mr1-eqsin IPv6 with reason: connection issue * 02:38 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on cr[2-3]-eqsin.mgmt,ps1-[603-604]-eqsin with reason: connection issue * 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 27s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-23 == * 23:27 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin.oob,mr1-eqsin.oob IPv6 with reason: switch refresh * 22:21 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Setting storage compatibility to NONE - eevans@cumin1003 * 22:01 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Setting storage compatibility to NONE - eevans@cumin1003 * 21:29 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1014.eqiad.wmnet, repooling source-only afterwards * 21:28 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 46s) * 21:28 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 21:19 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Setting storage compatibility to UPGRADING - eevans@cumin1003 * 21:00 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Setting storage compatibility to UPGRADING - eevans@cumin1003 * 20:17 dani@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] (duration: 11m 57s) * 20:13 dani@deploy1003: dani: Continuing with deployment * 20:07 dani@deploy1003: dani: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:05 dani@deploy1003: Started scap sync-world: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] * 19:24 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:24 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating the rest of the ipv6 dns records. - jhancock@cumin2002" * 19:24 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating the rest of the ipv6 dns records. - jhancock@cumin2002" * 19:14 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 19:05 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wdqs1014.eqiad.wmnet with OS bookworm * 19:04 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.noop (exit_code=99) * 19:04 cwilliams@cumin1003: START - Cookbook sre.mysql.noop * 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1014.eqiad.wmnet with reason: host reimage * 18:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1014.eqiad.wmnet with reason: host reimage * 18:30 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2008.codfw.wmnet, repooling source-only afterwards * 18:28 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 19s) * 18:28 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1014 * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1014 * 18:22 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1014 * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1014.eqiad.wmnet 188.32.64.10.in-addr.arpa 8.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:22 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1014.eqiad.wmnet 188.32.64.10.in-addr.arpa 8.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1014 - bking@cumin2003" * 18:21 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1014 - bking@cumin2003" * 18:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2215: Maintenance * 18:18 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 18:15 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:15 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 18:06 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:06 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 18:05 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 18:04 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2052: codfw rack B8 re-pool after maintenance * 17:54 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 17:54 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:54 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 17:32 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2215: Maintenance * 17:29 cmooney@dns3003: END - running authdns-update * 17:27 cmooney@dns3003: START - running authdns-update * 17:23 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 17:22 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:18 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool es2052: codfw rack B8 re-pool after maintenance * 17:18 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2189: codfw rack B8 re-pool after maintenance * 17:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2215.codfw.wmnet with reason: Maintenance * 17:17 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 17:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2215 [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95126 and previous config saved to /var/cache/conftool/dbconfig/20260723-170903-cwilliams.json * 17:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2191 to x1 primary [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95125 and previous config saved to /var/cache/conftool/dbconfig/20260723-170612-cwilliams.json * 17:05 cezmunsta: Starting x1 codfw failover from db2215 to db2191 - [[phab:T432986|T432986]] * 16:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2191 with weight 0 [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95123 and previous config saved to /var/cache/conftool/dbconfig/20260723-165831-cwilliams.json * 16:58 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 16 hosts with reason: Primary switchover x1 [[phab:T432986|T432986]] * 16:36 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 138128 * 16:35 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 138128 * 16:33 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2189: codfw rack B8 re-pool after maintenance * 16:33 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2164: codfw rack B8 re-pool after maintenance * 16:28 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1072.eqiad.wmnet * 16:27 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1072.eqiad.wmnet with OS trixie * 16:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2249: Maintenance * 16:06 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker1072.eqiad.wmnet with reason: host reimage * 16:06 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1072.eqiad.wmnet with reason: host reimage * 15:50 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1072 * 15:50 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1072 * 15:49 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1072.eqiad.wmnet with OS trixie * 15:48 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2164: codfw rack B8 re-pool after maintenance * 15:48 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] (duration: 06m 37s) * 15:48 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2163: codfw rack B8 re-pool after maintenance * 15:45 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 15:43 musikanimal@deploy1003: musikanimal: Continuing with deployment * 15:43 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:41 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] * 15:36 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1072.eqiad.wmnet * 15:35 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1072.eqiad.wmnet * 15:35 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1072.eqiad.wmnet * 15:34 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:34 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push any outstanding updates - cmooney@cumin1003" * 15:34 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push any outstanding updates - cmooney@cumin1003" * 15:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db2249: Maintenance * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 15:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:26 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:21 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:21 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 15:21 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:21 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 15:20 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 15:19 cmooney@dns2004: END - running authdns-update * 15:17 cmooney@dns2004: START - running authdns-update * 15:14 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns2004.wikimedia.org * 15:12 brouberol@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 15:12 brouberol@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 15:12 klausman@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ml-serve1001.eqiad.wmnet with OS trixie * 15:11 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1071.eqiad.wmnet * 15:11 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1071.eqiad.wmnet with OS trixie * 15:10 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wdqs2008.codfw.wmnet with OS bookworm * 15:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2249.codfw.wmnet with reason: Maintenance * 15:08 brouberol@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 15:08 brouberol@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 15:08 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 15:08 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 15:06 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns1004.wikimedia.org * 15:02 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2002.codfw.wmnet * 15:02 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2002.codfw.wmnet * 15:02 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2163: codfw rack B8 re-pool after maintenance * 15:01 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 15:01 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 14:59 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test2001.codfw.wmnet * 14:57 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test2001.codfw.wmnet * 14:56 ryankemper: [WDQS] [[phab:T430880|T430880]] Reimaged `wdqs2016` to Bookworm, xferred scholarly_articles from `wdqs2024`, validated updater/readiness/federation, and repooled * 14:51 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2016.codfw.wmnet * 14:51 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1001.eqiad.wmnet with reason: host reimage * 14:48 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1071.eqiad.wmnet with reason: host reimage * 14:47 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1001.eqiad.wmnet with reason: host reimage * 14:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2008.codfw.wmnet with reason: host reimage * 14:43 topranks: reboot lsw1-b8-codw to upgrade JunOS [[phab:T430929|T430929]] * 14:41 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1071.eqiad.wmnet with reason: host reimage * 14:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2008.codfw.wmnet with reason: host reimage * 14:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2231: Maintenance * 14:30 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1001.eqiad.wmnet with OS trixie * 14:25 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2002.codfw.wmnet * 14:23 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1071 * 14:23 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1071 * 14:23 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 14:22 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore scholarly data after Bookworm reimage) xfer scholarly_articles from wdqs2024.codfw.wmnet -> wdqs2016.codfw.wmnet, repooling source-only afterwards * 14:22 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2052: codfw rack B8 depool for maintenance * 14:21 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool es2052: codfw rack B8 depool for maintenance * 14:21 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2249: codfw rack B8 depool for maintenance * 14:21 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1071 * 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1071.eqiad.wmnet 166.48.64.10.in-addr.arpa 6.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:21 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1071.eqiad.wmnet 166.48.64.10.in-addr.arpa 6.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1071 - jiji@cumin1003" * 14:21 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1071 - jiji@cumin1003" * 14:21 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2249: codfw rack B8 depool for maintenance * 14:21 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2189: codfw rack B8 depool for maintenance * 14:20 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2002.codfw.wmnet * 14:20 cmooney@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2050.codfw.wmnet * 14:20 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2189: codfw rack B8 depool for maintenance * 14:20 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2164: codfw rack B8 depool for maintenance * 14:20 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2164: codfw rack B8 depool for maintenance * 14:19 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2163: codfw rack B8 depool for maintenance * 14:19 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:19 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1014 * 14:19 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2163: codfw rack B8 depool for maintenance * 14:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1014.eqiad.wmnet with OS bookworm * 14:17 cmooney@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2050.codfw.wmnet * 14:16 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 14:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2008 * 14:14 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2008 * 14:14 cmooney@cumin1003: conftool action : set/pooled=no; selector: name=dns2004.wikimedia.org * 14:14 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2008.codfw.wmnet with OS bookworm * 14:13 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 14:12 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:10 topranks: depool dns2004 before lsw1-b8-codfw switch maintenance [[phab:T430929|T430929]] * 14:10 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b8-codfw,lsw1-b8-codfw IPv6,lsw1-b8-codfw.mgmt,ssw1-a[1,8]-codfw with reason: lsw1-b8-codfw JunOS upgrade * 14:07 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 30 hosts with reason: lsw1-b8-codfw JunOS upgrade * 14:06 elukey: upload python3-docker-report 0.0.19 to apt.wikimedia.org for bookworm and trixie * 13:59 jiji@cumin1003: START - Cookbook sre.dns.netbox * 13:58 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1071 * 13:57 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:57 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:57 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:57 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1071.eqiad.wmnet with OS trixie * 13:55 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1071.eqiad.wmnet * 13:55 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1071.eqiad.wmnet * 13:55 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1071.eqiad.wmnet * 13:53 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:52 logmsgbot: kharlan Deployed security patch for [[phab:T432948|T432948]] * 13:51 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:51 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 13:51 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db2231: Maintenance * 13:50 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:50 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:50 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:50 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:49 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:49 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:49 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2231 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95097 and previous config saved to /var/cache/conftool/dbconfig/20260723-134436-cwilliams.json * 13:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2231.codfw.wmnet with reason: Maintenance * 13:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:39 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:38 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 13:38 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] (duration: 09m 07s) * 13:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1037 hosts * 13:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2196: Maintenance * 13:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:33 kharlan@deploy1003: kharlan, emc-wmf: Continuing with deployment * 13:31 kharlan@deploy1003: kharlan, emc-wmf: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:30 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:28 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] * 13:17 hashar@deploy1003: Finished deploy [integration/docroot@2199146]: build: License GPL2.0+ / updating npm dependencies (duration: 00m 14s) * 13:17 hashar@deploy1003: Started deploy [integration/docroot@2199146]: build: License GPL2.0+ / updating npm dependencies * 13:14 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service * 13:07 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 12:58 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2207: Repooling * 12:49 root@cumin1003: START - Cookbook sre.mysql.pool pool db2196: Maintenance * 12:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2196 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95087 and previous config saved to /var/cache/conftool/dbconfig/20260723-123952-cwilliams.json * 12:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2196.codfw.wmnet with reason: Maintenance * 12:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2191: Maintenance * 12:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:13 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:13 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: Repooling * 12:12 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2207: Repooling * 12:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: Repooling * 11:56 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2235.codfw.wmnet with OS trixie * 11:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db2191: Maintenance * 11:46 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1070.eqiad.wmnet * 11:46 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1070.eqiad.wmnet * 11:46 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1070.eqiad.wmnet * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2191 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95080 and previous config saved to /var/cache/conftool/dbconfig/20260723-114308-cwilliams.json * 11:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2191.codfw.wmnet with reason: Maintenance * 11:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2186: Maintenance * 11:35 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 46375 * 11:34 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 46375 * 11:33 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2235.codfw.wmnet with reason: host reimage * 11:28 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2235.codfw.wmnet with reason: host reimage * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c7-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c7-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c6-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c6-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c5-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c5-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c4-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c4-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c3-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c3-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c2-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c2-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d7-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d7-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d4-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d4-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d3-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d2-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d2-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d8-eqiad * 11:23 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d8-eqiad * 11:23 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d1-eqiad * 11:23 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d1-eqiad * 11:12 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db2235.codfw.wmnet with OS trixie * 11:11 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:11 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[2160,2235].codfw.wmnet with reason: Upgrading * 10:56 root@cumin1003: START - Cookbook sre.mysql.pool pool db2186: Maintenance * 10:54 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1037: testing * 10:53 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1037: testing * 10:53 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1037: testing * 10:53 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1037: testing * 10:52 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: testing * 10:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2186 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95072 and previous config saved to /var/cache/conftool/dbconfig/20260723-104956-cwilliams.json * 10:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2186.codfw.wmnet with reason: Maintenance * 10:43 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1054: testing * 10:41 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1070.eqiad.wmnet with OS trixie * 10:30 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1037 hosts * 10:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 10:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 10:20 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1070.eqiad.wmnet with reason: host reimage * 10:16 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1070.eqiad.wmnet with reason: host reimage * 10:06 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1038: testing * 10:05 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1038: testing * 10:05 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1038: testing * 10:02 marostegui@dns1004: END - running authdns-update * 10:00 marostegui@dns1004: START - running authdns-update * 09:58 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: testing * 09:57 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1054: testing * 09:57 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1070 * 09:57 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1070 * 09:57 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1054: testing * 09:57 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1054: testing * 09:56 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1055: testing * 09:56 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1070 * 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1070.eqiad.wmnet 165.48.64.10.in-addr.arpa 5.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:56 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1070.eqiad.wmnet 165.48.64.10.in-addr.arpa 5.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1070 - jiji@cumin1003" * 09:56 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1070 - jiji@cumin1003" * 09:47 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for 1035 hosts * 09:45 jiji@cumin1003: START - Cookbook sre.dns.netbox * 09:42 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1070 * 09:42 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1070.eqiad.wmnet with OS trixie * 09:42 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1070.eqiad.wmnet * 09:41 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1070.eqiad.wmnet * 09:41 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1070.eqiad.wmnet * 09:27 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es2051: testing * 09:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: testing * 09:12 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es2051: testing * 09:11 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1055: testing * 09:09 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:09 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1055: testing * 09:09 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1055: testing * 08:50 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:50 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:50 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 08:49 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 08:49 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 08:49 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:46 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1069.eqiad.wmnet * 08:46 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1069.eqiad.wmnet * 08:46 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1069.eqiad.wmnet * 08:39 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 08:38 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 08:38 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 08:37 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 08:35 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:10 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1069.eqiad.wmnet with OS trixie * 07:49 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1069.eqiad.wmnet with reason: host reimage * 07:45 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1069.eqiad.wmnet with reason: host reimage * 07:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1035 hosts * 07:33 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 07:32 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 07:29 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1069 * 07:29 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1069 * 07:26 jiji@deploy1003: Finished scap sync-world: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules (duration: 06m 01s) * 07:25 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1069 * 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1069.eqiad.wmnet 164.48.64.10.in-addr.arpa 4.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:25 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1069.eqiad.wmnet 164.48.64.10.in-addr.arpa 4.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1069 - jiji@cumin1003" * 07:25 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1069 - jiji@cumin1003" * 07:25 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 07:25 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 07:24 jiji@deploy1003: jiji: Continuing with deployment * 07:22 jiji@deploy1003: jiji: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:21 jiji@deploy1003: Started scap sync-world: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules * 07:21 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1031.eqiad.wmnet,service=s7 * 07:20 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1031.eqiad.wmnet,service=s2 * 07:20 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1031.eqiad.wmnet,service=s7 * 07:20 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1031.eqiad.wmnet,service=s2 * 07:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts * 07:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts * 07:19 jiji@cumin1003: START - Cookbook sre.dns.netbox * 07:19 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1069 * 07:19 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1069.eqiad.wmnet with OS trixie * 07:19 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1069.eqiad.wmnet * 07:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 45 hosts * 07:17 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1069.eqiad.wmnet * 07:17 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1069.eqiad.wmnet * 07:14 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 45 hosts * 07:13 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 06:16 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs2007 after successful Bookworm reimage, data transfer, and postflight validation; wdqs1013 also passed postflights and is enabled in conftool, but remains out of IPVS pending a rolling pybal restart to clear its stale pre-VLAN-move address * 05:58 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2007.codfw.wmnet * 05:58 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1013.eqiad.wmnet * 05:54 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore scholarly data after Bookworm reimage) xfer scholarly_articles from wdqs2024.codfw.wmnet -> wdqs2016.codfw.wmnet, repooling source-only afterwards == 2026-07-22 == * 23:34 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Apply upgrade to JVM17 - eevans@cumin1003 * 23:14 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Apply upgrade to JVM17 - eevans@cumin1003 * 22:06 ryankemper: [WDQS] Added requestctl per-IP ratelimit `wdqs_heavy_sparql_bots_jul_2026_ratelimit` (chronic heavy-query bot tier driving deadlock-remediation restarts); pruned superseded `wdqs_2026_05_11_worobot` * 21:51 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] (duration: 11m 52s) * 21:44 sbassett@deploy1003: sbassett: Continuing with deployment * 21:43 sbassett@deploy1003: sbassett: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:39 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] * 20:38 dani@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] (duration: 32m 51s) * 20:38 ryankemper: [WDQS] Pruned obsolete requestctl action+pattern `wdqs_20260715_p2003_ring_ja3n` (actor rotated JA3Ns; rule inert) * 20:26 dani@deploy1003: dani, vadymts1: Continuing with deployment * 20:24 dani@deploy1003: dani, vadymts1: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:14 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2244: Testing * 20:06 dani@deploy1003: Started scap sync-world: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] * 19:56 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1013.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2016.codfw.wmnet with OS bookworm * 19:40 mutante: gerrit - one more service restart is needed - restarting * 19:29 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2244: Testing * 19:27 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2244: Testing * 19:27 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2244: Testing * 19:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2016.codfw.wmnet with reason: host reimage * 19:15 dancy@deploy1003: Finished deploy [zuul/deploy@d92e238]: Freshening Zuul installation (duration: 00m 15s) * 19:14 dancy@deploy1003: Started deploy [zuul/deploy@d92e238]: Freshening Zuul installation * 19:11 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2016.codfw.wmnet with reason: host reimage * 18:54 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1013.eqiad.wmnet, repooling source-only afterwards * 18:52 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2016 * 18:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2016 * 18:51 dancy@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 18:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2016.codfw.wmnet with OS bookworm * 18:39 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 11s) * 18:39 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 18:36 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 18:30 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417] (thin): Regular analytics weekly train THIN [analytics/refinery@2a25417d] (duration: 02m 09s) * 18:28 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417] (thin): Regular analytics weekly train THIN [analytics/refinery@2a25417d] * 18:28 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417]: Regular analytics weekly train [analytics/refinery@2a25417d] (duration: 04m 31s) * 18:27 dduvall: deploying https://gerrit.wikimedia.org/r/c/integration/config/+/1314025 (4 jobs updated) * 18:23 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417]: Regular analytics weekly train [analytics/refinery@2a25417d] * 18:22 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@2a25417d] (duration: 01m 59s) * 18:20 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@2a25417d] * 17:56 Raine: deployment server switchover => deploy1003 is primary now * 17:55 kamila@deploy1003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 28m 24s) * 17:54 mutante: restarting gerrit for maintenance * 17:29 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1013.eqiad.wmnet with OS bookworm * 17:27 kamila@deploy1003: Started scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] * 17:20 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] (duration: 22m 50s) * 17:12 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1023.eqiad.wmnet -> wdqs1024.eqiad.wmnet, repooling source-only afterwards * 17:04 Raine: point deployment.eqiad.wmnet to deploy1003 * 17:04 kamila@dns7001: END - running authdns-update * 17:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1013.eqiad.wmnet with reason: host reimage * 17:02 kamila@dns7001: START - running authdns-update * 17:01 kamila@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on releases2003.codfw.wmnet,releases1003.eqiad.wmnet with reason: Deployment server switchover * 17:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1013.eqiad.wmnet with reason: host reimage * 16:58 kamila@deploy2003: Locking from deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] * 16:57 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] (duration: 02m 33s) * 16:55 kamila@deploy2003: Locking from deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] * 16:55 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2003 - [[phab:T240266|T240266]] (duration: 00m 11s) * 16:54 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2003 - [[phab:T240266|T240266]] * 16:40 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1013 * 16:40 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1013 * 16:39 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1013 * 16:39 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1013.eqiad.wmnet 105.32.64.10.in-addr.arpa 5.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:39 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1013.eqiad.wmnet 105.32.64.10.in-addr.arpa 5.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:39 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:39 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1013 - bking@cumin2003" * 16:39 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1013 - bking@cumin2003" * 16:34 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:34 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1013 * 16:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1013.eqiad.wmnet with OS bookworm * 16:28 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1023.eqiad.wmnet -> wdqs1024.eqiad.wmnet, repooling source-only afterwards * 16:27 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-scholarly,name=eqiad * 16:27 eevans@deploy2003: helmfile [eqiad] DONE helmfile.d/services/linked-artifacts: apply * 16:26 eevans@deploy2003: helmfile [eqiad] START helmfile.d/services/linked-artifacts: apply * 16:26 eevans@deploy2003: helmfile [codfw] DONE helmfile.d/services/linked-artifacts: apply * 16:26 eevans@deploy2003: helmfile [codfw] START helmfile.d/services/linked-artifacts: apply * 16:25 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 29s) * 16:25 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 16:24 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 16:21 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 16:18 eevans@deploy2003: helmfile [codfw] DONE helmfile.d/services/linked-artifacts: apply * 16:18 eevans@deploy2003: helmfile [codfw] START helmfile.d/services/linked-artifacts: apply * 16:08 eevans@deploy2003: helmfile [staging] DONE helmfile.d/services/linked-artifacts: apply * 16:07 eevans@deploy2003: helmfile [staging] START helmfile.d/services/linked-artifacts: apply * 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 16:01 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 15:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1024.eqiad.wmnet with OS bookworm * 15:49 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] (duration: 00m 10s) * 15:49 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] * 15:48 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] (duration: 00m 15s) * 15:48 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] * 15:47 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] (duration: 00m 10s) * 15:47 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] * 15:46 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:42 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:42 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:40 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:37 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 15:37 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:36 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:36 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:36 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 15:33 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 15:33 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:31 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:28 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 15:27 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:27 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1024.eqiad.wmnet with reason: host reimage * 15:23 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:23 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:23 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:20 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1068.eqiad.wmnet * 15:20 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1068.eqiad.wmnet * 15:20 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1068.eqiad.wmnet * 15:20 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 15:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1024.eqiad.wmnet with reason: host reimage * 15:11 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:55 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 14:52 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wdqs1024.eqiad.wmnet with OS bookworm * 14:50 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] (duration: 00m 09s) * 14:50 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] * 14:49 jiji@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 14:49 jiji@deploy2003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 14:49 jiji@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 14:48 jiji@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 14:45 ecarg@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:45 ecarg@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:44 ecarg@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:44 ecarg@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:43 ecarg@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:43 ecarg@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:41 sukhe: ipvsadm --delete-service --tcp-service 10.2.1.55:8087: lvs2014 and lvs2013 * 14:39 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:39 sukhe: ipvsadm --delete-service --tcp-service 10.2.2.55:8087: [[phab:T432445|T432445]] * 14:38 ecarg@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:38 ecarg@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:37 ecarg@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:37 ecarg@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:36 ecarg@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:34 ecarg@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts datahubsearch1001.eqiad.wmnet * 14:32 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:32 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 14:31 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 14:31 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 14:28 sukhe: sudo cumin 'A:lvs-low-traffic-codfw' 'systemctl restart pybal': lvs2013 * 14:26 sukhe: sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal': lvs2014 * 14:26 sukhe: sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal' * 14:24 sukhe: restart pybal on lvs1019 * 14:24 sukhe: restart pybal on lvs1020 * 14:19 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for 1036 hosts * 14:17 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:04 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] (duration: 09m 28s) * 13:59 kharlan@deploy2003: dreamyjazz, kharlan: Continuing with deployment * 13:58 bking@cumin2003: START - Cookbook sre.hosts.decommission for hosts datahubsearch1001.eqiad.wmnet * 13:57 kharlan@deploy2003: dreamyjazz, kharlan: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:55 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] * 13:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts datahubsearch[1002-1003].eqiad.wmnet * 13:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:53 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch[1002-1003].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 13:52 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch[1002-1003].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 13:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 13:42 stran@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] (duration: 07m 30s) * 13:42 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:40 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: Test * 13:38 stran@deploy2003: dragoniez, stran: Continuing with deployment * 13:37 stran@deploy2003: dragoniez, stran: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:35 bking@cumin2003: START - Cookbook sre.hosts.decommission for hosts datahubsearch[1002-1003].eqiad.wmnet * 13:35 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024'] * 13:35 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 13:35 stran@deploy2003: Started scap sync-world: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] * 13:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 13:28 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024'] * 13:26 sukhe@dns1004: END - running authdns-update * 13:25 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 13:25 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 13:24 sukhe@dns1004: START - running authdns-update * 13:22 stran@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] (duration: 08m 20s) * 13:21 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 13:20 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 13:19 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 13:19 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 13:19 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 13:18 stran@deploy2003: stran: Continuing with deployment * 13:16 stran@deploy2003: stran: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:14 stran@deploy2003: Started scap sync-world: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] * 13:13 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 13:13 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 13:11 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 13:11 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 13:08 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 12:55 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool es2051: Test * 12:55 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: Test * 12:54 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool es2051: Test * 12:43 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1036 hosts * 12:41 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 12:40 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1048.eqiad.wmnet * 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 12:39 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 12:38 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 12:37 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 12:37 brouberol@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 12:36 brouberol@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 12:36 brouberol@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 12:36 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 12:36 elukey@cumin1003: DONE (PASS) - Cookbook sre.puppet.renew-cert (exit_code=0) for crm2001.codfw.wmnet: Renew puppet certificate - elukey@cumin1003 * 12:35 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:35 brouberol@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 12:34 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 12:31 brouberol@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 12:30 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1048.eqiad.wmnet * 12:30 brouberol@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 12:28 brouberol@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 12:27 brouberol@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 12:20 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1068.eqiad.wmnet with OS trixie * 12:01 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] (duration: 13m 19s) * 11:58 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1068.eqiad.wmnet with reason: host reimage * 11:52 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1068.eqiad.wmnet with reason: host reimage * 11:51 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 11:49 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:47 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] * 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2252: Security updates * 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:43 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 11:42 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2252: Security updates * 11:42 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply * 11:40 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply * 11:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 11:37 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2252.codfw.wmnet with OS trixie * 11:34 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1068 * 11:34 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1068 * 11:26 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1068 * 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1068.eqiad.wmnet 46.48.64.10.in-addr.arpa 6.4.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:26 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1068.eqiad.wmnet 46.48.64.10.in-addr.arpa 6.4.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1068 - jiji@cumin1003" * 11:26 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1068 - jiji@cumin1003" * 11:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2252.codfw.wmnet with reason: host reimage * 11:17 jiji@cumin1003: START - Cookbook sre.dns.netbox * 11:17 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1068 * 11:17 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1068.eqiad.wmnet with OS trixie * 11:17 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2252.codfw.wmnet with reason: host reimage * 11:15 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1068.eqiad.wmnet * 11:15 mvolz@deploy2003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:15 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1068.eqiad.wmnet * 11:15 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1068.eqiad.wmnet * 11:14 mvolz@deploy2003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:13 mvolz@deploy2003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:13 mvolz@deploy2003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:12 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] (duration: 11m 05s) * 11:11 mvolz@deploy2003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:10 mvolz@deploy2003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:07 dreamyjazz@deploy2003: dreamyjazz, kharlan: Continuing with deployment * 11:03 dreamyjazz@deploy2003: dreamyjazz, kharlan: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:03 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2252.codfw.wmnet with OS trixie * 11:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2252: Upgrading db2252.codfw.wmnet * 11:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:02 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 11:02 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2252: Upgrading db2252.codfw.wmnet * 11:02 cwilliams@cumin1003: dbmaint on ms3@codfw [[phab:T432321|T432321]] * 11:01 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 11:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db1153.eqiad.wmnet with reason: Security updates * 11:01 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] * 11:00 fnegri@deploy2003: helmfile [eqiad] DONE helmfile.d/services/toolhub: apply * 10:58 fnegri@deploy2003: helmfile [eqiad] START helmfile.d/services/toolhub: apply * 10:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1151: Security updates * 10:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:57 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1151: Security updates * 10:55 fnegri@deploy2003: helmfile [codfw] DONE helmfile.d/services/toolhub: apply * 10:54 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] (duration: 08m 38s) * 10:53 fnegri@deploy2003: helmfile [codfw] START helmfile.d/services/toolhub: apply * 10:53 fnegri@deploy2003: helmfile [staging] DONE helmfile.d/services/toolhub: apply * 10:52 fnegri@deploy2003: helmfile [staging] START helmfile.d/services/toolhub: apply * 10:50 zabe@deploy2003: zabe: Continuing with deployment * 10:47 zabe@deploy2003: zabe: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:45 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] * 10:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1151: Security updates * 10:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:42 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:42 root@cumin1003: START - Cookbook sre.mysql.depool depool db1151: Security updates * 10:38 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] (duration: 12m 47s) * 10:34 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 10:34 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 10:33 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2253.codfw.wmnet with OS trixie * 10:28 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:26 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] * 10:18 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2253.codfw.wmnet with reason: host reimage * 10:13 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2253.codfw.wmnet with reason: host reimage * 10:00 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2253.codfw.wmnet with OS trixie * 09:58 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db1151.eqiad.wmnet with reason: Security updates * 09:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2253: Upgrading db2253.codfw.wmnet * 09:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:57 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 09:56 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2253: Upgrading db2253.codfw.wmnet * 09:56 cwilliams@cumin1003: dbmaint on ms2@codfw [[phab:T432321|T432321]] * 09:56 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 09:36 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: UI improvement; support url shortener - oblivian@cumin1003" * 09:36 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: UI improvement; support url shortener - oblivian@cumin1003 * 09:35 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: UI improvement; support url shortener - oblivian@cumin1003 * 09:35 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: UI improvement; support url shortener - oblivian@cumin1003" * 09:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1152: Security updates * 09:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:26 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db1152: Security updates * 09:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: Security updates * 09:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:11 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:11 root@cumin1003: START - Cookbook sre.mysql.depool depool db1152: Security updates * 09:10 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1018.eqiad.wmnet with reason: Cloning * 09:09 Dreamy_Jazz: Deployed patch for [[phab:T432453|T432453]] and [[phab:T432454|T432454]] * 09:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 09:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2251.codfw.wmnet with OS trixie * 08:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2251.codfw.wmnet with reason: host reimage * 08:45 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2251.codfw.wmnet with reason: host reimage * 08:40 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] (duration: 12m 26s) * 08:38 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1030.eqiad.wmnet,service=s1 * 08:36 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 08:31 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2251.codfw.wmnet with OS trixie * 08:30 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:28 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] * 08:25 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] (duration: 07m 59s) * 08:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2251: Upgrading db2251.codfw.wmnet * 08:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:22 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 08:22 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2251: Upgrading db2251.codfw.wmnet * 08:20 urbanecm@deploy2003: urbanecm: Continuing with deployment * 08:20 cwilliams@cumin1003: dbmaint on ms1@codfw [[phab:T432321|T432321]] * 08:20 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 08:19 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:17 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] * 08:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade * 08:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade * 08:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db2251.codfw.wmnet,db1152.eqiad.wmnet with reason: OS upgrade * 08:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade * 08:13 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade * 08:11 Dreamy_Jazz: Created cusi_signal, cusi_case, and cusi_user on ukwiki and enwikivoyage in extension1 * 08:11 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade * 08:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade * 08:04 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1030.eqiad.wmnet,service=s1 * 08:04 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1030.eqiad.wmnet,service=s1 * 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply * 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply * 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply * 07:51 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply * 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 07:47 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 07:47 phuedx: End of UTC morning backport window * 07:43 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 07:43 phuedx@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] (duration: 13m 44s) * 07:43 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 07:39 phuedx@deploy2003: phuedx: Continuing with deployment * 07:31 phuedx@deploy2003: phuedx: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:29 phuedx@deploy2003: Started scap sync-world: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] * 07:24 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Turnilo import support - oblivian@cumin1003" * 07:24 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import support - oblivian@cumin1003 * 07:23 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import support - oblivian@cumin1003 * 07:23 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Turnilo import support - oblivian@cumin1003" * 06:42 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs2020 after successful Bookworm reimage, data transfer, and postflight validation * 06:42 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2020.codfw.wmnet * 05:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Managing sanitization for wikis bolwiki in section s5 * 05:25 marostegui@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis bolwiki in section s5 * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 41s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-21 == * 22:50 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2019.codfw.wmnet -> wdqs2020.codfw.wmnet, repooling source-only afterwards * 22:47 cwhite: force reboot arclamp2001 - appears to have run out of memory and gone unresponsive * 22:24 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 01m 26s) * 22:24 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 22:23 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 22:22 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024'] * 22:11 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 21:54 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs1024'] * 21:54 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 21:53 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs1024'] * 21:53 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 21:49 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2019.codfw.wmnet -> wdqs2020.codfw.wmnet, repooling source-only afterwards * 20:57 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] (duration: 09m 10s) * 20:55 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1024.eqiad.wmnet with OS bookworm * 20:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2020.codfw.wmnet with OS bookworm * 20:52 krinkle@deploy2003: krinkle: Continuing with deployment * 20:49 krinkle@deploy2003: krinkle: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:47 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] * 20:45 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] (duration: 05m 42s) * 20:44 krinkle@deploy2003: krinkle: Rolling back deployment * 20:41 krinkle@deploy2003: krinkle: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:39 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] * 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2003.codfw.wmnet * 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1003.eqiad.wmnet * 20:33 dani@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] (duration: 11m 15s) * 20:33 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2003.codfw.wmnet * 20:33 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1003.eqiad.wmnet * 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2020.codfw.wmnet with reason: host reimage * 20:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1002.eqiad.wmnet * 20:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2002.codfw.wmnet * 20:30 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 20:29 dani@deploy2003: dani: Continuing with deployment * 20:29 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2020.codfw.wmnet with reason: host reimage * 20:26 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1002.eqiad.wmnet * 20:26 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2002.codfw.wmnet * 20:24 dani@deploy2003: dani: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2001.codfw.wmnet * 20:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1001.eqiad.wmnet * 20:22 dani@deploy2003: Started scap sync-world: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] * 20:22 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 20:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2001.codfw.wmnet * 20:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1001.eqiad.wmnet * 20:14 sbisson@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] (duration: 09m 01s) * 20:11 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2020.codfw.wmnet with OS bookworm * 20:10 sbisson@deploy2003: sbisson: Continuing with deployment * 20:07 sbisson@deploy2003: sbisson: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:05 sbisson@deploy2003: Started scap sync-world: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] * 20:03 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] (duration: 07m 04s) * 20:01 mutante: Gerrit - tomorrow a new SSH host key will appear - it will be {{Gerrit|ed25519}} and has already been added to wmf-laptop. you can verify it here: https://wikitech.wikimedia.org/wiki/Help:SSH_Fingerprints/gerrit.wikimedia.org:29418 ([[phab:T240266|T240266]]) * 19:59 zabe@deploy2003: zabe: Continuing with deployment * 19:58 zabe@deploy2003: zabe: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:56 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] * 19:52 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] (duration: 07m 25s) * 19:48 zabe@deploy2003: zabe: Continuing with deployment * 19:47 zabe@deploy2003: zabe: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:45 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] * 19:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 19:32 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024'] * 19:27 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 19:26 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024'] * 19:26 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 19:24 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024'] * 19:12 ryankemper: [wdqs] [[phab:T430880|T430880]] Repooled `wdqs-scholarly` discovery in `eqiad` after validating `wdqs1023` end-to-end; `wdqs1024` remains disabled pending reimage recovery * 19:11 ryankemper: [wdqs] [[phab:T430880|T430880]] Repooled wdqs1012.eqiad.wmnet after successful Bookworm reimage, data transfer, service checks, readiness probe, and cross-graph federation query validation * 19:10 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 19:10 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1012.eqiad.wmnet * 19:08 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 18:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for deploy1003.eqiad.wmnet * 18:57 kamila@cumin1003: START - Cookbook sre.hosts.remove-downtime for deploy1003.eqiad.wmnet * 18:37 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:37 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding urldownloader service IPs - sukhe@cumin1003" * 18:37 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding urldownloader service IPs - sukhe@cumin1003" * 18:32 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 18:32 dancy@deploy2003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 18:30 sukhe@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 18:27 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 18:24 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1024.eqiad.wmnet with OS bookworm * 18:20 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host deploy1003.eqiad.wmnet with OS bookworm * 18:09 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deploy1003 reimage (duration: 121m 16s) * 18:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 18:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1180: Security updates * 17:55 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wcqs2003.codfw.wmnet * 17:48 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wcqs2003.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1155.eqiad.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1155.eqiad.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2224.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2224.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2217.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2217.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2193.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2193.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2180.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2180.codfw.wmnet * 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1168.eqiad.wmnet * 17:36 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1168.eqiad.wmnet * 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2169.codfw.wmnet * 17:36 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2169.codfw.wmnet * 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1165.eqiad.wmnet * 17:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1165.eqiad.wmnet * 17:35 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2158.codfw.wmnet * 17:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2158.codfw.wmnet * 17:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wcqs1003.eqiad.wmnet * 17:17 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1180: Security updates * 17:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1180.eqiad.wmnet * 17:16 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1180.eqiad.wmnet * 17:15 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp3073.* * 17:13 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wcqs1003.eqiad.wmnet * 17:11 brett@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp3073.esams.wmnet with OS trixie * 17:11 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 17:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1024 * 17:04 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1024 * 17:03 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 17:00 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 16:59 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2242: codfw rack B7 depool for maintenance * 16:59 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 16:43 brett@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp3073.esams.wmnet with reason: host reimage * 16:42 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 16:39 brett@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cp3073.esams.wmnet with reason: host reimage * 16:32 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on deploy1003.eqiad.wmnet with reason: host reimage * 16:27 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on deploy1003.eqiad.wmnet with reason: host reimage * 16:14 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2242: codfw rack B7 depool for maintenance * 16:14 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: codfw rack B7 depool for maintenance * 16:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1023.eqiad.wmnet with OS bookworm * 16:13 brett@cumin2002: START - Cookbook sre.hosts.reimage for host cp3073.esams.wmnet with OS trixie * 16:08 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host deploy1003.eqiad.wmnet with OS bookworm * 16:08 kamila@deploy2003: Locking from deployment [MediaWiki]: deploy1003 reimage * 16:03 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1012.eqiad.wmnet with OS bookworm * 15:48 inflatador: bking@apt1002 `sudo reprepro copy bookworm-wikimedia bullseye-wikimedia jvmquake` [[phab:T430880|T430880]] * 15:39 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp3073.* * 15:39 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 15:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:34 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1311536{{!}}Set $wgMathInternalRestbaseURL explicitly (take 2) (T349582)]] * 15:29 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:29 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2228: codfw rack B7 depool for maintenance * 15:29 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2229: codfw rack B7 depool for maintenance * 15:27 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1180: Security update * 15:25 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Security update * 15:21 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 15:21 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 15:19 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 24s) * 15:19 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:14 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 15:14 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db1180: Security update * 15:13 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-eqiad * 14:48 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-eqiad * 14:44 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2229: codfw rack B7 depool for maintenance * 14:44 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc2017: codfw rack B7 depool for maintenance * 14:44 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:43 cmooney@cumin2003: START - Cookbook sre.mysql.parsercache * 14:43 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool pc2017: codfw rack B7 depool for maintenance * 14:43 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2003.codfw.wmnet * 14:43 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2003.codfw.wmnet * 14:42 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2009.codfw.wmnet * 14:42 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2009.codfw.wmnet * 14:41 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:41 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:40 cmooney@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 29 hosts * 14:40 cmooney@cumin1003: START - Cookbook sre.hosts.remove-downtime for 29 hosts * 14:35 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 14:34 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 14:32 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] (duration: 07m 56s) * 14:29 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ssw1-a[1,8]-codfw with reason: lsw1-b7-codfw JunOS upgrade * 14:28 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 14:28 elukey@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 14:26 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:24 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] * 14:23 topranks: reboot lsw1-b7-codfw to upgrade JunOS (affects all hosts in rack) [[phab:T430928|T430928]] * 14:18 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2003.codfw.wmnet * 14:14 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2009.codfw.wmnet * 14:14 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Security update * 14:13 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2242: codfw rack B7 depool for maintenance * 14:13 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2242: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2228: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2228: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2229: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2229: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc2017: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.parsercache * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool pc2017: codfw rack B7 depool for maintenance * 14:08 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2003.codfw.wmnet * 14:07 cmooney@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on aux-k8s-etcd2004.codfw.wmnet,ml-etcd2001.codfw.wmnet with reason: lsw1-b7-codfw JunOS upgrade * 14:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95005 and previous config saved to /var/cache/conftool/dbconfig/20260721-140620-cwilliams.json * 14:05 cmooney@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2049.codfw.wmnet * 14:05 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-scholarly,name=eqiad * 14:04 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2009.codfw.wmnet * 14:04 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply * 14:04 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply * 14:03 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:03 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2001.codfw.wmnet * 14:03 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2001.codfw.wmnet * 14:02 cmooney@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2049.codfw.wmnet * 14:00 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:00 Dreamy_Jazz: Created cusi_case, cusi_signal, and cusi_user on svwiki, dewiki, jawiki, eswiki * 13:59 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b7-codfw,lsw1-b7-codfw IPv6,lsw1-b7-codfw.mgmt,ssw1-a[1,8]-codfw.mgmt with reason: lsw1-b7-codfw JunOS upgrade * 13:57 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs1023.eqiad.wmnet, repooling source-only afterwards * 13:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 29 hosts with reason: lsw1-b7-codfw JunOS upgrade * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224', diff saved to https://phabricator.wikimedia.org/P95003 and previous config saved to /var/cache/conftool/dbconfig/20260721-135613-cwilliams.json * 13:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1012.eqiad.wmnet with reason: host reimage * 13:53 cmooney@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 1:00:00 on 30 hosts with reason: lsw1-b7-codfw JunOS upgrade * 13:51 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1012.eqiad.wmnet with reason: host reimage * 13:48 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 13:48 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 13:46 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224', diff saved to https://phabricator.wikimedia.org/P95001 and previous config saved to /var/cache/conftool/dbconfig/20260721-134605-cwilliams.json * 13:46 cmooney@cumin1003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti2033.codfw.wmnet * 13:46 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 13:45 elukey: move the Docker Registry's /v2/wikimedia/machinelearning.* prefix to the ml S3 backend - [[phab:T428022|T428022]] * 13:45 cmooney@cumin1003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti2033.codfw.wmnet * 13:45 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 13:43 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:40 jiji@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 13:40 jiji@deploy2003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 13:39 jiji@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 13:39 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 13:38 cmooney@cumin1003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2032.codfw.wmnet * 13:38 jiji@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 13:38 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 13:37 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2032.codfw.wmnet * 13:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95000 and previous config saved to /var/cache/conftool/dbconfig/20260721-133557-cwilliams.json * 13:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1012 * 13:33 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1012 * 13:33 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1012.eqiad.wmnet with OS bookworm * 13:30 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:30 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:28 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94999 and previous config saved to /var/cache/conftool/dbconfig/20260721-132855-cwilliams.json * 13:28 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2224.codfw.wmnet with reason: Maintenance * 13:28 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94998 and previous config saved to /var/cache/conftool/dbconfig/20260721-132826-cwilliams.json * 13:28 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 13:23 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] (duration: 07m 50s) * 13:20 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 13:18 kharlan@deploy2003: kharlan: Continuing with deployment * 13:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217', diff saved to https://phabricator.wikimedia.org/P94996 and previous config saved to /var/cache/conftool/dbconfig/20260721-131817-cwilliams.json * 13:17 brouberol@dns1004: END - running authdns-update * 13:17 kharlan@deploy2003: kharlan: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:15 brouberol@dns1004: START - running authdns-update * 13:15 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] * 13:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94995 and previous config saved to /var/cache/conftool/dbconfig/20260721-131411-cwilliams.json * 13:13 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs1023.eqiad.wmnet, repooling source-only afterwards * 13:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217', diff saved to https://phabricator.wikimedia.org/P94994 and previous config saved to /var/cache/conftool/dbconfig/20260721-130809-cwilliams.json * 13:07 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 13:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180', diff saved to https://phabricator.wikimedia.org/P94993 and previous config saved to /var/cache/conftool/dbconfig/20260721-130404-cwilliams.json * 13:03 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:03 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:02 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 13:02 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 12:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94992 and previous config saved to /var/cache/conftool/dbconfig/20260721-125801-cwilliams.json * 12:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180', diff saved to https://phabricator.wikimedia.org/P94991 and previous config saved to /var/cache/conftool/dbconfig/20260721-125356-cwilliams.json * 12:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94990 and previous config saved to /var/cache/conftool/dbconfig/20260721-125049-cwilliams.json * 12:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2217.codfw.wmnet with reason: Maintenance * 12:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94989 and previous config saved to /var/cache/conftool/dbconfig/20260721-125017-cwilliams.json * 12:48 elukey: bmc cold reboot for lvs1013 and lvs1015 - [[phab:T426180|T426180]] * 12:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94988 and previous config saved to /var/cache/conftool/dbconfig/20260721-124348-cwilliams.json * 12:40 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193', diff saved to https://phabricator.wikimedia.org/P94987 and previous config saved to /var/cache/conftool/dbconfig/20260721-124009-cwilliams.json * 12:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts * 12:33 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts * 12:33 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts * 12:32 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts * 12:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:30 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193', diff saved to https://phabricator.wikimedia.org/P94986 and previous config saved to /var/cache/conftool/dbconfig/20260721-123001-cwilliams.json * 12:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94985 and previous config saved to /var/cache/conftool/dbconfig/20260721-121953-cwilliams.json * 12:17 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs2007.codfw.wmnet with OS bookworm * 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94983 and previous config saved to /var/cache/conftool/dbconfig/20260721-121257-cwilliams.json * 12:12 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2193.codfw.wmnet with reason: Maintenance * 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94982 and previous config saved to /var/cache/conftool/dbconfig/20260721-121239-cwilliams.json * 12:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180', diff saved to https://phabricator.wikimedia.org/P94980 and previous config saved to /var/cache/conftool/dbconfig/20260721-120231-cwilliams.json * 11:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180', diff saved to https://phabricator.wikimedia.org/P94979 and previous config saved to /var/cache/conftool/dbconfig/20260721-115223-cwilliams.json * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94978 and previous config saved to /var/cache/conftool/dbconfig/20260721-114333-cwilliams.json * 11:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1180.eqiad.wmnet with reason: Maintenance * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94977 and previous config saved to /var/cache/conftool/dbconfig/20260721-114305-cwilliams.json * 11:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94976 and previous config saved to /var/cache/conftool/dbconfig/20260721-114215-cwilliams.json * 11:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94975 and previous config saved to /var/cache/conftool/dbconfig/20260721-113530-cwilliams.json * 11:35 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2180.codfw.wmnet with reason: Maintenance * 11:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94974 and previous config saved to /var/cache/conftool/dbconfig/20260721-113501-cwilliams.json * 11:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168', diff saved to https://phabricator.wikimedia.org/P94973 and previous config saved to /var/cache/conftool/dbconfig/20260721-113258-cwilliams.json * 11:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169', diff saved to https://phabricator.wikimedia.org/P94972 and previous config saved to /var/cache/conftool/dbconfig/20260721-112453-cwilliams.json * 11:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168', diff saved to https://phabricator.wikimedia.org/P94971 and previous config saved to /var/cache/conftool/dbconfig/20260721-112250-cwilliams.json * 11:21 XioNoX: put eqiad-drmrs Arelion link in service * 11:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169', diff saved to https://phabricator.wikimedia.org/P94970 and previous config saved to /var/cache/conftool/dbconfig/20260721-111446-cwilliams.json * 11:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94969 and previous config saved to /var/cache/conftool/dbconfig/20260721-111242-cwilliams.json * 11:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 11:10 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 11:07 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1093 hosts * 11:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94968 and previous config saved to /var/cache/conftool/dbconfig/20260721-110548-cwilliams.json * 11:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1168.eqiad.wmnet with reason: Maintenance * 11:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94967 and previous config saved to /var/cache/conftool/dbconfig/20260721-110520-cwilliams.json * 11:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94966 and previous config saved to /var/cache/conftool/dbconfig/20260721-110439-cwilliams.json * 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94964 and previous config saved to /var/cache/conftool/dbconfig/20260721-105632-cwilliams.json * 10:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2169.codfw.wmnet with reason: Maintenance * 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94963 and previous config saved to /var/cache/conftool/dbconfig/20260721-105603-cwilliams.json * 10:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165', diff saved to https://phabricator.wikimedia.org/P94962 and previous config saved to /var/cache/conftool/dbconfig/20260721-105512-cwilliams.json * 10:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158', diff saved to https://phabricator.wikimedia.org/P94961 and previous config saved to /var/cache/conftool/dbconfig/20260721-104555-cwilliams.json * 10:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165', diff saved to https://phabricator.wikimedia.org/P94960 and previous config saved to /var/cache/conftool/dbconfig/20260721-104504-cwilliams.json * 10:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158', diff saved to https://phabricator.wikimedia.org/P94959 and previous config saved to /var/cache/conftool/dbconfig/20260721-103547-cwilliams.json * 10:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94958 and previous config saved to /var/cache/conftool/dbconfig/20260721-103456-cwilliams.json * 10:29 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2229: Upgraded kernel * 10:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94956 and previous config saved to /var/cache/conftool/dbconfig/20260721-102757-cwilliams.json * 10:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on an-redacteddb1001.eqiad.wmnet,clouddb[1015,1025,1028].eqiad.wmnet,db1155.eqiad.wmnet with reason: Maintenance * 10:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1165.eqiad.wmnet with reason: Maintenance * 10:25 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94955 and previous config saved to /var/cache/conftool/dbconfig/20260721-102539-cwilliams.json * 10:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94954 and previous config saved to /var/cache/conftool/dbconfig/20260721-101848-cwilliams.json * 10:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2158.codfw.wmnet with reason: Maintenance * 09:43 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2229: Upgraded kernel * 09:42 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2229.codfw.wmnet * 09:42 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2229.codfw.wmnet * 09:23 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db2229.codfw.wmnet * 09:23 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2229.codfw.wmnet * 08:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2229 [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94948 and previous config saved to /var/cache/conftool/dbconfig/20260721-085724-cwilliams.json * 08:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2214 to s6 primary [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94947 and previous config saved to /var/cache/conftool/dbconfig/20260721-085442-cwilliams.json * 08:53 cezmunsta: Starting s6 codfw failover from db2229 to db2214 - [[phab:T430964|T430964]] * 08:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2214 with weight 0 [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94946 and previous config saved to /var/cache/conftool/dbconfig/20260721-084613-cwilliams.json * 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 22 hosts with reason: Primary switchover s6 [[phab:T430964|T430964]] * 08:32 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1017.eqiad.wmnet,service=s1 * 08:08 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Add subrated circuit rate to interface descriptions - CR1312476 - ayounsi@cumin1003 * 08:06 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Add subrated circuit rate to interface descriptions - CR1312476 - ayounsi@cumin1003 * 07:58 reedy@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] (duration: 12m 55s) * 07:51 reedy@deploy2003: reedy, neriah: Continuing with deployment * 07:51 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 07:51 reedy@deploy2003: reedy, neriah: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:48 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1093 hosts * 07:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm2001.wikimedia.org * 07:45 reedy@deploy2003: Started scap sync-world: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] * 07:43 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 07:42 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm2001.wikimedia.org * 07:23 elukey: upgrade libtiff6 packages on zuul* trixie hosts for security upgrades * 07:22 elukey: upgrade libtiff6 packages on Wikikube trixie workers for security upgrades * 07:14 elukey@deploy2003: helmfile [codfw] DONE helmfile.d/services/proton: sync * 07:13 elukey@deploy2003: helmfile [codfw] START helmfile.d/services/proton: sync * 07:11 elukey@deploy2003: helmfile [eqiad] DONE helmfile.d/services/proton: sync * 07:10 elukey@deploy2003: helmfile [eqiad] START helmfile.d/services/proton: sync * 07:09 elukey@deploy2003: helmfile [staging] DONE helmfile.d/services/proton: sync * 07:08 elukey@deploy2003: helmfile [staging] START helmfile.d/services/proton: sync * 06:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1023.eqiad.wmnet with reason: host reimage * 06:46 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1023.eqiad.wmnet with reason: host reimage * 06:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 05:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Haproxy-only mode support - oblivian@cumin1003" * 05:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Haproxy-only mode support - oblivian@cumin1003 * 05:42 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Haproxy-only mode support - oblivian@cumin1003 * 05:42 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Haproxy-only mode support - oblivian@cumin1003" * 05:38 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1017.eqiad.wmnet with reason: Cloning * 05:37 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1017.eqiad.wmnet,service=s1 * 05:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1029.eqiad.wmnet,service=s8 * 05:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1029.eqiad.wmnet,service=s5 * 05:32 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:30 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 05:11 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet * 05:04 aokoth@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet * 05:00 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 04:56 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 04:01 mwpresync@deploy2003: Pruned MediaWiki: 1.47.0-wmf.9 (duration: 01m 08s) * 03:41 mwpresync@deploy2003: Finished scap sync-world: testwikis to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] (duration: 36m 30s) * 03:05 mwpresync@deploy2003: Started scap sync-world: testwikis to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 03:01 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:01 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:00 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:00 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:36 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:36 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:36 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:35 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:16 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 47s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 00:56 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm == 2026-07-20 == * 23:38 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 23:07 Amir1: deleting echo notifications from 2015 on group1 wikis * 23:07 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] (duration: 14m 16s) * 23:01 ladsgroup@deploy2003: ladsgroup: Continuing with deployment * 23:00 ladsgroup@deploy2003: ladsgroup: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:53 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] * 22:46 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2007.codfw.wmnet, repooling source-only afterwards * 22:39 maryum: Deployed security fixes for several security bugs * 21:42 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 21:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2007.codfw.wmnet, repooling source-only afterwards * 21:37 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 21:37 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 21:34 sbassett: Deployed security fix for [[phab:T432424|T432424]] * 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs2020.codfw.wmnet * 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1023.eqiad.wmnet * 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1011.eqiad.wmnet * 21:32 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 17s) * 21:32 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 21:27 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 21:13 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2007.codfw.wmnet with reason: host reimage * 21:08 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-internal-main,name=codfw * 21:06 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2007.codfw.wmnet with reason: host reimage * 20:59 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service * 20:58 sukhe: pybal restart for IP changes around wdqs-main hosts * 20:57 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 20:46 ryankemper@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-internal-main,name=codfw * 20:45 ebernhardson@deploy2003: Finished deploy [search/mjolnir/deploy@d4dc3b8]: Update for opensearch 2.x compat (duration: 00m 34s) * 20:44 ebernhardson@deploy2003: Started deploy [search/mjolnir/deploy@d4dc3b8]: Update for opensearch 2.x compat * 20:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2007 * 20:44 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2007 * 20:43 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2007 * 20:43 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2007.codfw.wmnet 156.16.192.10.in-addr.arpa 6.5.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:42 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2007.codfw.wmnet 156.16.192.10.in-addr.arpa 6.5.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:42 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2007 - bking@cumin2003" * 20:41 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2007 - bking@cumin2003" * 20:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94944 and previous config saved to /var/cache/conftool/dbconfig/20260720-203333-cwilliams.json * 20:32 arlolra@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] (duration: 15m 07s) * 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2020.codfw.wmnet * 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1023.eqiad.wmnet * 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1011.eqiad.wmnet * 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs2020.codfw.wmnet * 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1023.eqiad.wmnet * 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1011.eqiad.wmnet * 20:25 arlolra@deploy2003: arlolra, cscott: Continuing with deployment * 20:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257', diff saved to https://phabricator.wikimedia.org/P94943 and previous config saved to /var/cache/conftool/dbconfig/20260720-202325-cwilliams.json * 20:21 arlolra@deploy2003: arlolra, cscott: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:17 arlolra@deploy2003: Started scap sync-world: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] * 20:13 bking@cumin2003: START - Cookbook sre.dns.netbox * 20:13 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257', diff saved to https://phabricator.wikimedia.org/P94942 and previous config saved to /var/cache/conftool/dbconfig/20260720-201318-cwilliams.json * 20:13 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 20:10 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 20:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2007 * 20:04 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2007.codfw.wmnet with OS bookworm * 20:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94941 and previous config saved to /var/cache/conftool/dbconfig/20260720-200310-cwilliams.json * 19:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94940 and previous config saved to /var/cache/conftool/dbconfig/20260720-195633-cwilliams.json * 19:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1257.eqiad.wmnet with reason: Maintenance * 19:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94939 and previous config saved to /var/cache/conftool/dbconfig/20260720-195605-cwilliams.json * 19:51 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 19:50 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 19:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256', diff saved to https://phabricator.wikimedia.org/P94938 and previous config saved to /var/cache/conftool/dbconfig/20260720-194558-cwilliams.json * 19:44 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 19:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 19:41 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wdqs1011.eqiad.wmnet with OS bookworm * 19:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256', diff saved to https://phabricator.wikimedia.org/P94937 and previous config saved to /var/cache/conftool/dbconfig/20260720-193550-cwilliams.json * 19:25 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94936 and previous config saved to /var/cache/conftool/dbconfig/20260720-192542-cwilliams.json * 19:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94935 and previous config saved to /var/cache/conftool/dbconfig/20260720-191856-cwilliams.json * 19:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1256.eqiad.wmnet with reason: Maintenance * 19:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94934 and previous config saved to /var/cache/conftool/dbconfig/20260720-191839-cwilliams.json * 19:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255', diff saved to https://phabricator.wikimedia.org/P94933 and previous config saved to /var/cache/conftool/dbconfig/20260720-190831-cwilliams.json * 18:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255', diff saved to https://phabricator.wikimedia.org/P94932 and previous config saved to /var/cache/conftool/dbconfig/20260720-185824-cwilliams.json * 18:50 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 18:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94931 and previous config saved to /var/cache/conftool/dbconfig/20260720-184816-cwilliams.json * 18:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94930 and previous config saved to /var/cache/conftool/dbconfig/20260720-184224-cwilliams.json * 18:42 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1255.eqiad.wmnet with reason: Maintenance * 18:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94929 and previous config saved to /var/cache/conftool/dbconfig/20260720-184153-cwilliams.json * 18:39 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 18:39 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 16s) * 18:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 18:38 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 59m 26s) * 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 18:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211', diff saved to https://phabricator.wikimedia.org/P94928 and previous config saved to /var/cache/conftool/dbconfig/20260720-183145-cwilliams.json * 18:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211', diff saved to https://phabricator.wikimedia.org/P94927 and previous config saved to /var/cache/conftool/dbconfig/20260720-182137-cwilliams.json * 18:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94926 and previous config saved to /var/cache/conftool/dbconfig/20260720-181129-cwilliams.json * 18:09 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_codfw * 18:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2057.codfw.wmnet * 18:08 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_codfw * 18:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2058.codfw.wmnet * 18:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94925 and previous config saved to /var/cache/conftool/dbconfig/20260720-180452-cwilliams.json * 18:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on clouddb[1016,1020,1022-1023].eqiad.wmnet,db1154.eqiad.wmnet with reason: Maintenance * 18:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1211.eqiad.wmnet with reason: Maintenance * 18:02 sukhe: armed keyholder on acmechief1002.eqiad.wmnet and acmechief2002.codfw.wmnet (active host) * 18:01 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief2002.codfw.wmnet * 17:57 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief2002.codfw.wmnet * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs2020'] * 17:52 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief1002.eqiad.wmnet * 17:50 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 17:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1011.eqiad.wmnet with reason: host reimage * 17:48 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief1002.eqiad.wmnet * 17:47 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test2001.codfw.wmnet * 17:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94924 and previous config saved to /var/cache/conftool/dbconfig/20260720-174717-cwilliams.json * 17:46 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 17:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1011.eqiad.wmnet with reason: host reimage * 17:43 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs2020.codfw.wmnet with OS bookworm * 17:43 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test2001.codfw.wmnet * 17:43 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test1001.eqiad.wmnet * 17:39 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test1001.eqiad.wmnet * 17:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 17:38 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:38 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 17:37 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 17:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244', diff saved to https://phabricator.wikimedia.org/P94923 and previous config saved to /var/cache/conftool/dbconfig/20260720-173709-cwilliams.json * 17:35 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 17:31 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 20m 40s) * 17:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2055.codfw.wmnet * 17:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2056.codfw.wmnet * 17:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1011 * 17:27 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1011 * 17:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1011.eqiad.wmnet with OS bookworm * 17:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244', diff saved to https://phabricator.wikimedia.org/P94922 and previous config saved to /var/cache/conftool/dbconfig/20260720-172701-cwilliams.json * 17:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94921 and previous config saved to /var/cache/conftool/dbconfig/20260720-171653-cwilliams.json * 17:11 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 17:11 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 13m 03s) * 17:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94920 and previous config saved to /var/cache/conftool/dbconfig/20260720-171012-cwilliams.json * 17:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2244.codfw.wmnet with reason: Maintenance * 17:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94919 and previous config saved to /var/cache/conftool/dbconfig/20260720-170941-cwilliams.json * 16:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243', diff saved to https://phabricator.wikimedia.org/P94918 and previous config saved to /var/cache/conftool/dbconfig/20260720-165933-cwilliams.json * 16:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 16:58 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2053.codfw.wmnet * 16:51 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2054.codfw.wmnet * 16:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243', diff saved to https://phabricator.wikimedia.org/P94917 and previous config saved to /var/cache/conftool/dbconfig/20260720-164926-cwilliams.json * 16:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94916 and previous config saved to /var/cache/conftool/dbconfig/20260720-163918-cwilliams.json * 16:35 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 16:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94915 and previous config saved to /var/cache/conftool/dbconfig/20260720-163140-cwilliams.json * 16:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2243.codfw.wmnet with reason: Maintenance * 16:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94914 and previous config saved to /var/cache/conftool/dbconfig/20260720-163111-cwilliams.json * 16:27 btullis@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 16:27 btullis@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 16:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2020 * 16:23 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2020 * 16:21 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2020 * 16:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2020.codfw.wmnet 85.0.192.10.in-addr.arpa 5.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:21 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2020.codfw.wmnet 85.0.192.10.in-addr.arpa 5.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242', diff saved to https://phabricator.wikimedia.org/P94913 and previous config saved to /var/cache/conftool/dbconfig/20260720-162103-cwilliams.json * 16:19 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 16:18 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 16:18 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:18 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 16:17 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:17 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for netbox accounting errors - jhancock@cumin2002" * 16:17 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for netbox accounting errors - jhancock@cumin2002" * 16:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2051.codfw.wmnet * 16:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2052.codfw.wmnet * 16:11 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 16:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242', diff saved to https://phabricator.wikimedia.org/P94912 and previous config saved to /var/cache/conftool/dbconfig/20260720-161055-cwilliams.json * 16:09 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 16:08 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 16:06 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 16:06 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 16:06 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2020 * 16:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2020.codfw.wmnet with OS bookworm * 16:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94911 and previous config saved to /var/cache/conftool/dbconfig/20260720-160047-cwilliams.json * 15:58 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2019.codfw.wmnet, repooling source-only afterwards * 15:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94909 and previous config saved to /var/cache/conftool/dbconfig/20260720-155353-cwilliams.json * 15:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2242.codfw.wmnet with reason: Maintenance * 15:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94908 and previous config saved to /var/cache/conftool/dbconfig/20260720-154433-cwilliams.json * 15:35 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2049.codfw.wmnet * 15:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162', diff saved to https://phabricator.wikimedia.org/P94907 and previous config saved to /var/cache/conftool/dbconfig/20260720-153425-cwilliams.json * 15:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2050.codfw.wmnet * 15:28 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162', diff saved to https://phabricator.wikimedia.org/P94906 and previous config saved to /var/cache/conftool/dbconfig/20260720-152418-cwilliams.json * 15:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1023 * 15:14 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1023 * 15:14 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 15:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94905 and previous config saved to /var/cache/conftool/dbconfig/20260720-151407-cwilliams.json * 15:13 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] (duration: 41m 16s) * 15:08 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 15:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94902 and previous config saved to /var/cache/conftool/dbconfig/20260720-150729-cwilliams.json * 15:07 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2162.codfw.wmnet with reason: Maintenance * 15:05 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2027.codfw.wmnet, repooling source-only afterwards * 15:00 urbanecm@deploy2003: vadymts1, migr, urbanecm: Continuing with deployment * 14:59 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:58 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 07s) * 14:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:58 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 13s) * 14:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:57 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2019.codfw.wmnet, repooling source-only afterwards * 14:57 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2047.codfw.wmnet * 14:55 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2048.codfw.wmnet * 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2019.codfw.wmnet with OS bookworm * 14:49 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 14:47 urbanecm@deploy2003: vadymts1, migr, urbanecm: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:44 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool magru [reason: BGP issues in lvs7003 resolved after liberica restart, no task ID specified] * 14:44 sukhe@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool magru [reason: BGP issues in lvs7003 resolved after liberica restart, no task ID specified] * 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:41 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:39 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:39 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:33 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool magru [reason: no reason specified, no task ID specified] * 14:33 sukhe@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool magru [reason: no reason specified, no task ID specified] * 14:31 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] * 14:24 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:24 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:24 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:24 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2019.codfw.wmnet with reason: host reimage * 14:22 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2027.codfw.wmnet, repooling source-only afterwards * 14:19 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2019.codfw.wmnet with reason: host reimage * 14:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2027.codfw.wmnet with OS bookworm * 14:16 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2046.codfw.wmnet * 14:16 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2045.codfw.wmnet * 14:08 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:08 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:08 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:08 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:07 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:06 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:06 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:06 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:05 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2019 * 14:00 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2019 * 13:56 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1015.eqiad.wmnet * 13:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2027.codfw.wmnet with reason: host reimage * 13:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2071.codfw.wmnet with OS trixie * 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:51 sukhe@cumin1003: END (ERROR) - Cookbook sre.loadbalancer.admin (exit_code=97) rebooting A:liberica and P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica and P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:51 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1015.eqiad.wmnet * 13:50 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1014.eqiad.wmnet * 13:50 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1076.eqiad.wmnet with OS trixie * 13:50 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2027.codfw.wmnet with reason: host reimage * 13:45 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1014.eqiad.wmnet * 13:44 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1013.eqiad.wmnet * 13:39 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 13:38 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1013.eqiad.wmnet * 13:37 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2044.codfw.wmnet * 13:37 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2043.codfw.wmnet * 13:36 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2019 * 13:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2019.codfw.wmnet 156.32.192.10.in-addr.arpa 6.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:36 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2019.codfw.wmnet 156.32.192.10.in-addr.arpa 6.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:36 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2019 - bking@cumin2003" * 13:36 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2019 - bking@cumin2003" * 13:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry2005.codfw.wmnet * 13:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2071.codfw.wmnet with reason: host reimage * 13:31 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:31 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2019 * 13:31 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry2005.codfw.wmnet * 13:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry2004.codfw.wmnet * 13:30 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2019.codfw.wmnet with OS bookworm * 13:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2027 * 13:30 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2027 * 13:30 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2027.codfw.wmnet with OS bookworm * 13:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1076.eqiad.wmnet with reason: host reimage * 13:29 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_codfw * 13:28 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_codfw * 13:26 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry2004.codfw.wmnet * 13:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry1005.eqiad.wmnet * 13:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2071.codfw.wmnet with reason: host reimage * 13:22 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1076.eqiad.wmnet with reason: host reimage * 13:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry1005.eqiad.wmnet * 13:21 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry1004.eqiad.wmnet * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry1004.eqiad.wmnet * 13:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts * 13:13 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts * 13:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts * 13:12 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts * 13:03 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1076.eqiad.wmnet with OS trixie * 13:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2071.codfw.wmnet with OS trixie * 12:55 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:54 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:53 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:46 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 7 hosts * 12:42 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 7 hosts * 12:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts * 12:42 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts * 12:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2070.codfw.wmnet with OS trixie * 12:36 ozge@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:35 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1075.eqiad.wmnet with OS trixie * 12:32 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts * 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts * 12:22 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin1001.eqiad.wmnet * 12:19 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin1001.eqiad.wmnet * 12:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2070.codfw.wmnet with reason: host reimage * 12:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin2001.codfw.wmnet * 12:14 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1075.eqiad.wmnet with reason: host reimage * 12:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2070.codfw.wmnet with reason: host reimage * 12:10 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1075.eqiad.wmnet with reason: host reimage * 12:09 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin2001.codfw.wmnet * 11:17 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1074.eqiad.wmnet with OS trixie * 11:17 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 11:16 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 11:14 ozge@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 6 hosts * 11:09 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 6 hosts * 11:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 324 hosts * 10:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1074.eqiad.wmnet with reason: host reimage * 10:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2069.codfw.wmnet with OS trixie * 10:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1074.eqiad.wmnet with reason: host reimage * 10:30 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2069.codfw.wmnet with reason: host reimage * 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1074.eqiad.wmnet with OS trixie * 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2069.codfw.wmnet with reason: host reimage * 10:06 blake@deploy2003: Stopping before sync operations * 10:06 blake@deploy2003: Started scap sync-world: Non-deployment scap run to populate new release values * 10:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2069.codfw.wmnet with OS trixie * 10:00 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1073.eqiad.wmnet with OS trixie * 09:56 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 324 hosts * 09:39 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 09:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 8 hosts * 09:38 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1073.eqiad.wmnet with reason: host reimage * 09:37 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 8 hosts * 09:34 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1073.eqiad.wmnet with reason: host reimage * 09:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2068.codfw.wmnet with OS trixie * 09:16 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1073.eqiad.wmnet with OS trixie * 09:13 blake@deploy2003: sync-world aborted: Non-deployment scap run to populate new release values (duration: 00m 02s) * 09:13 blake@deploy2003: Started scap sync-world: Non-deployment scap run to populate new release values * 08:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2068.codfw.wmnet with reason: host reimage * 08:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2068.codfw.wmnet with reason: host reimage * 08:50 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 08:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2068.codfw.wmnet with OS trixie * 08:15 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1072.eqiad.wmnet with OS trixie * 07:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2067.codfw.wmnet with OS trixie * 07:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1072.eqiad.wmnet with reason: host reimage * 07:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1072.eqiad.wmnet with reason: host reimage * 07:45 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 07:45 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 07:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2067.codfw.wmnet with reason: host reimage * 07:35 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2067.codfw.wmnet with reason: host reimage * 07:30 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 07:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1072.eqiad.wmnet with OS trixie * 07:30 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 07:17 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 07:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2067.codfw.wmnet with OS trixie * 05:51 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:50 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:25 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:25 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on db2207.codfw.wmnet with reason: Host down * 04:28 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 07m 02s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-18 == * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 29s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 00:11 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2018.codfw.wmnet, repooling source-only afterwards == 2026-07-17 == * 23:53 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2026.codfw.wmnet, repooling source-only afterwards * 23:09 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2018.codfw.wmnet, repooling source-only afterwards * 23:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2026.codfw.wmnet, repooling source-only afterwards * 22:11 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2018.codfw.wmnet with OS bookworm * 22:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2026.codfw.wmnet with OS bookworm * 21:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2018.codfw.wmnet with reason: host reimage * 21:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2018.codfw.wmnet with reason: host reimage * 21:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2026.codfw.wmnet with reason: host reimage * 21:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2026.codfw.wmnet with reason: host reimage * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2018 * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2018 * 21:26 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2018 * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2018.codfw.wmnet 155.32.192.10.in-addr.arpa 5.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:26 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2018.codfw.wmnet 155.32.192.10.in-addr.arpa 5.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2018 - bking@cumin2003" * 21:26 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2018 - bking@cumin2003" * 21:14 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:13 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2018 * 21:13 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2018.codfw.wmnet with OS bookworm * 21:12 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2026 * 21:12 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2026 * 21:12 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2026.codfw.wmnet with OS bookworm * 21:05 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1022.eqiad.wmnet -> wdqs1026.eqiad.wmnet, repooling source-only afterwards * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs2017.codfw.wmnet, repooling source-only afterwards * 20:11 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs2017.codfw.wmnet, repooling source-only afterwards * 20:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2017.codfw.wmnet with OS bookworm * 20:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1022.eqiad.wmnet -> wdqs1026.eqiad.wmnet, repooling source-only afterwards * 20:06 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1026.eqiad.wmnet with OS bookworm * 19:55 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 19:55 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:55 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 09s) * 19:55 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:50 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 08s) * 19:50 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:50 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 10m 03s) * 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2017.codfw.wmnet with reason: host reimage * 19:40 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:40 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 15s) * 19:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1026.eqiad.wmnet with reason: host reimage * 19:37 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 16s) * 19:37 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:34 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2017.codfw.wmnet with reason: host reimage * 19:34 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1026.eqiad.wmnet with reason: host reimage * 19:33 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 19:33 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:16 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1026.eqiad.wmnet with OS bookworm * 19:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2017 * 19:16 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2017 * 19:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2017.codfw.wmnet with OS bookworm * 18:30 bking@dns1004: END - running authdns-update * 18:28 bking@dns1004: START - running authdns-update * 18:16 kamila@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1264.eqiad.wmnet * 18:16 kamila@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1264.eqiad.wmnet * 18:16 kamila@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1264.eqiad.wmnet * 17:49 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 17:46 dzahn@dns1006: END - running authdns-update * 17:44 dzahn@dns1006: START - running authdns-update * 17:44 dzahn@dns1006: END - running authdns-update * 17:42 dzahn@dns1006: START - running authdns-update * 17:28 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 17:21 kamila@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 17:01 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1264 * 17:01 kamila@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1264 * 17:01 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 17:01 kamila@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1264.eqiad.wmnet * 17:01 kamila@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1264.eqiad.wmnet * 17:01 kamila@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1264.eqiad.wmnet * 16:42 reedy@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] (duration: 10m 29s) * 16:34 reedy@deploy2003: reedy, hartman: Continuing with deployment * 16:33 reedy@deploy2003: reedy, hartman: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:31 reedy@deploy2003: Started scap sync-world: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] * 16:26 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 16:10 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in2001.wikimedia.org with reason: [[phab:T431659|T431659]] * 16:07 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in1001.wikimedia.org with reason: [[phab:T431659|T431659]] * 16:05 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 16:01 kamila@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 16:00 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out2001.wikimedia.org with reason: [[phab:T431659|T431659]] * 15:41 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 15:41 kamila@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 15:35 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out1001.wikimedia.org with reason: [[phab:T431659|T431659]] * 15:14 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1339.eqiad.wmnet * 15:13 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1339.eqiad.wmnet * 15:13 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1339.eqiad.wmnet * 14:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1339.eqiad.wmnet with OS trixie * 14:50 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:49 kamila@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:49 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:33 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage * 14:27 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage * 14:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1339 * 14:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1339 * 14:14 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1339 * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1339.eqiad.wmnet 156.32.64.10.in-addr.arpa 6.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:14 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1339.eqiad.wmnet 156.32.64.10.in-addr.arpa 6.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1339 - cgoubert@cumin2003" * 14:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1339 - cgoubert@cumin2003" * 14:09 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 14:06 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1339 * 14:06 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie * 14:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1339.eqiad.wmnet * 14:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1339.eqiad.wmnet * 14:02 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1339.eqiad.wmnet * 13:45 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb1013.eqiad.wmnet * 13:39 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb1013.eqiad.wmnet * 13:27 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:24 blake@dns1004: END - running authdns-update * 13:22 blake@dns1004: START - running authdns-update * 13:20 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 13:11 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2014.codfw.wmnet * 13:06 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb2014.codfw.wmnet * 13:06 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2012.codfw.wmnet * 13:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 13:03 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 13:01 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 15 hosts * 13:01 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb2012.codfw.wmnet * 13:01 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1016.eqiad.wmnet * 13:00 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 15 hosts * 12:55 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb1016.eqiad.wmnet * 12:55 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1014.eqiad.wmnet * 12:49 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb1014.eqiad.wmnet * 12:32 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:32 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:31 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:31 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1338.eqiad.wmnet * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1338.eqiad.wmnet * 12:18 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1338.eqiad.wmnet * 12:17 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:15 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:14 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:13 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1338.eqiad.wmnet with OS trixie * 12:01 klausman@deploy2003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 11:59 klausman@deploy2003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 11:56 klausman@deploy2003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 11:54 klausman@deploy2003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 11:53 klausman@deploy2003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 11:51 klausman@deploy2003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 11:42 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1338.eqiad.wmnet with reason: host reimage * 11:38 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1338.eqiad.wmnet with reason: host reimage * 11:31 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2230.codfw.wmnet * 11:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1338 * 11:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1338 * 11:25 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1338 * 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1338.eqiad.wmnet 155.32.64.10.in-addr.arpa 5.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:25 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1338.eqiad.wmnet 155.32.64.10.in-addr.arpa 5.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1338 - cgoubert@cumin2003" * 11:25 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1338 - cgoubert@cumin2003" * 11:23 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2230.codfw.wmnet * 11:20 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 11:20 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1338 * 11:20 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1338.eqiad.wmnet with OS trixie * 11:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1338.eqiad.wmnet * 11:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1338.eqiad.wmnet * 11:19 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1338.eqiad.wmnet * 11:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1337.eqiad.wmnet * 11:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1337.eqiad.wmnet * 11:17 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1337.eqiad.wmnet * 11:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1337.eqiad.wmnet with OS trixie * 10:51 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[2001-2002].codfw.wmnet * 10:50 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1337.eqiad.wmnet with reason: host reimage * 10:40 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 10:39 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:39 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1337.eqiad.wmnet with reason: host reimage * 10:39 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1001-1003].eqiad.wmnet * 10:34 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:34 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:30 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:28 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1001-1003].eqiad.wmnet * 10:27 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1337 * 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1337 * 10:26 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1337 * 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1337.eqiad.wmnet 154.32.64.10.in-addr.arpa 4.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:26 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1337.eqiad.wmnet 154.32.64.10.in-addr.arpa 4.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1337 - cgoubert@cumin2003" * 10:26 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1337 - cgoubert@cumin2003" * 10:21 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 10:18 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1337 * 10:17 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1337.eqiad.wmnet with OS trixie * 10:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1337.eqiad.wmnet * 10:16 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db1176.eqiad.wmnet * 10:16 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1337.eqiad.wmnet * 10:16 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1337.eqiad.wmnet * 10:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1336.eqiad.wmnet * 10:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1336.eqiad.wmnet * 10:15 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1336.eqiad.wmnet * 10:11 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db1176.eqiad.wmnet * 10:10 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db1176.eqiad.wmnet * 10:09 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db1176.eqiad.wmnet * 10:05 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts (check the cookbook's logs for more details.) * 10:03 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts (check the cookbook's logs for more details.) * 09:58 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1336.eqiad.wmnet with OS trixie * 09:47 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts (check the cookbook's logs for more details.) * 09:47 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts (check the cookbook's logs for more details.) * 09:45 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host acmechief-test2001.codfw.wmnet,acmechief-test1001.eqiad.wmnet,an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet,db-test[2001-2002].codfw.wmnet,db-test[1 * 09:40 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host acmechief-test2001.codfw.wmnet,acmechief-test1001.eqiad.wmnet,an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet,db-test[2001-2002].codfw.wmnet,db-test[1001-1003].eqiad.wmn * 09:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1336.eqiad.wmnet with reason: host reimage * 09:33 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1336.eqiad.wmnet with reason: host reimage * 09:29 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet * 09:29 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet * 09:28 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:26 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:21 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 09:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1336 * 09:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1336 * 09:19 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 09:14 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1336 * 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1336.eqiad.wmnet 152.32.64.10.in-addr.arpa 2.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1336.eqiad.wmnet 152.32.64.10.in-addr.arpa 2.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1336 - cgoubert@cumin2003" * 09:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1336 - cgoubert@cumin2003" * 09:11 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:10 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 09:09 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:09 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1336 * 09:09 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1336.eqiad.wmnet with OS trixie * 09:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1336.eqiad.wmnet * 09:08 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1336.eqiad.wmnet * 09:08 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1336.eqiad.wmnet * 09:06 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1335.eqiad.wmnet * 09:06 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1335.eqiad.wmnet * 09:06 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1335.eqiad.wmnet * 09:04 elukey: uploaded spicerack_13.1.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia * 08:55 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wikikube-worker-exp2001.codfw.wmnet * 08:54 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host testreduce1002.eqiad.wmnet * 08:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1335.eqiad.wmnet with OS trixie * 08:51 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host wikikube-worker-exp2001.codfw.wmnet * 08:51 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wikikube-worker-exp1001.eqiad.wmnet * 08:50 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host testreduce1002.eqiad.wmnet * 08:45 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host wikikube-worker-exp1001.eqiad.wmnet * 08:34 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1335.eqiad.wmnet with reason: host reimage * 08:30 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1335.eqiad.wmnet with reason: host reimage * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1335 * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1335 * 08:18 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1335 * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1335.eqiad.wmnet 150.32.64.10.in-addr.arpa 0.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:18 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1335.eqiad.wmnet 150.32.64.10.in-addr.arpa 0.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1335 - cgoubert@cumin2003" * 08:18 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1335 - cgoubert@cumin2003" * 08:14 elukey@cumin1003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 08:14 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges * 08:13 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 08:10 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1335 * 08:10 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1335.eqiad.wmnet with OS trixie * 08:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1335.eqiad.wmnet * 08:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1335.eqiad.wmnet * 08:09 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1335.eqiad.wmnet * 08:06 elukey@cumin1003: END (FAIL) - Cookbook sre.puppet.disable-merges (exit_code=99) * 08:05 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges * 08:03 elukey@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin1003.eqiad.wmnet * 07:57 elukey@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin1003.eqiad.wmnet * 07:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetdb1003.eqiad.wmnet * 07:46 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetdb1003.eqiad.wmnet * 07:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetdb2003.codfw.wmnet * 07:37 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetdb2003.codfw.wmnet * 07:37 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1001.eqiad.wmnet * 07:28 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver1001.eqiad.wmnet * 07:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet * 07:19 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet * 07:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2002.codfw.wmnet * 07:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver2002.codfw.wmnet * 07:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet * 07:05 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet * 07:04 btullis@cumin1003: END (FAIL) - Cookbook sre.hadoop.reboot-workers (exit_code=99) for Hadoop analytics cluster * 07:04 elukey@cumin1003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 07:04 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges * 06:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox1003.eqiad.wmnet * 06:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox1003.eqiad.wmnet * 02:46 ryankemper@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:46 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:44 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:37 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:37 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-internal-scholarly,name=eqiad * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 49s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 01:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore wdqs1025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling source-only afterwards * 01:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore wdqs1027 after Bookworm reimage) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs1027.eqiad.wmnet, repooling both afterwards * 00:55 urbanecm@deploy2003: helmfile [codfw] DONE helmfile.d/services/linkrecommendation: apply * 00:54 urbanecm@deploy2003: helmfile [eqiad] DONE helmfile.d/services/linkrecommendation: apply * 00:54 urbanecm@deploy2003: helmfile [staging] DONE helmfile.d/services/linkrecommendation: apply * 00:54 urbanecm@deploy2003: helmfile [codfw] START helmfile.d/services/linkrecommendation: apply * 00:53 urbanecm@deploy2003: helmfile [staging] START helmfile.d/services/linkrecommendation: apply * 00:52 urbanecm@deploy2003: helmfile [eqiad] START helmfile.d/services/linkrecommendation: apply * 00:23 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore wdqs1025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling source-only afterwards * 00:23 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore wdqs1027 after Bookworm reimage) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs1027.eqiad.wmnet, repooling both afterwards * 00:14 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1274.eqiad.wmnet * 00:14 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1274.eqiad.wmnet * 00:14 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1274.eqiad.wmnet * 00:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1027.eqiad.wmnet with OS bookworm * 00:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1025.eqiad.wmnet with OS bookworm * 00:04 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1274.eqiad.wmnet with OS trixie == 2026-07-16 == * 23:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], xfer to freshly reimaged/scap-deployed wdqs2025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs2025.codfw.wmnet, repooling source-only afterwards * 23:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1027.eqiad.wmnet with reason: host reimage * 23:47 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1025.eqiad.wmnet with reason: host reimage * 23:43 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1274.eqiad.wmnet with reason: host reimage * 23:41 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1025.eqiad.wmnet with reason: host reimage * 23:39 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1027.eqiad.wmnet with reason: host reimage * 23:38 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1274.eqiad.wmnet with reason: host reimage * 23:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1025 * 23:23 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1025 * 23:22 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1027 * 23:22 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1027 * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1274 * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1274 * 23:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1025.eqiad.wmnet with OS bookworm * 23:19 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1274 * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1274.eqiad.wmnet 145.48.64.10.in-addr.arpa 5.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:19 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1274.eqiad.wmnet 145.48.64.10.in-addr.arpa 5.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1274 - swfrench@cumin1003" * 23:19 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1274 - swfrench@cumin1003" * 23:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1027.eqiad.wmnet with OS bookworm * 23:14 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 23:14 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1274 * 23:13 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1274.eqiad.wmnet with OS trixie * 23:13 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1274.eqiad.wmnet * 23:12 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1274.eqiad.wmnet * 23:12 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1274.eqiad.wmnet * 23:12 ryankemper: [[phab:T430880|T430880]] depooled dnsdisc of wdqs-internal-scholarly-eqiad bc we only have 1 host there * 23:09 ryankemper@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-internal-scholarly,name=eqiad * 23:08 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1272.eqiad.wmnet * 23:08 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1272.eqiad.wmnet * 23:08 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1272.eqiad.wmnet * 23:01 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], xfer to freshly reimaged/scap-deployed wdqs2025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs2025.codfw.wmnet, repooling source-only afterwards * 22:57 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1272.eqiad.wmnet with OS trixie * 22:56 Amir1: deleting echo notifications from 2015 in group0 * 22:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2025.codfw.wmnet with OS bookworm * 22:35 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1272.eqiad.wmnet with reason: host reimage * 22:32 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 27s) * 22:32 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 22:28 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1269.eqiad.wmnet * 22:28 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1269.eqiad.wmnet * 22:28 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1269.eqiad.wmnet * 22:27 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1272.eqiad.wmnet with reason: host reimage * 22:26 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] (duration: 08m 51s) * 22:22 ladsgroup@deploy2003: ladsgroup, urbanecm: Continuing with deployment * 22:19 ladsgroup@deploy2003: ladsgroup, urbanecm: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:17 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] * 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2025.codfw.wmnet with reason: host reimage * 22:06 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1272 * 22:06 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1272 * 22:05 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1272 * 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1272.eqiad.wmnet 127.48.64.10.in-addr.arpa 7.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:05 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1272.eqiad.wmnet 127.48.64.10.in-addr.arpa 7.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1272 - swfrench@cumin1003" * 22:05 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1272 - swfrench@cumin1003" * 22:03 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2025.codfw.wmnet with reason: host reimage * 22:01 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 22:00 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1272 * 22:00 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1272.eqiad.wmnet with OS trixie * 22:00 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1272.eqiad.wmnet * 21:59 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1272.eqiad.wmnet * 21:59 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1272.eqiad.wmnet * 21:55 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1271.eqiad.wmnet * 21:55 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1271.eqiad.wmnet * 21:55 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1271.eqiad.wmnet * 21:46 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1271.eqiad.wmnet with OS trixie * 21:45 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] (duration: 06m 31s) * 21:40 sbassett@deploy2003: sbassett: Continuing with deployment * 21:40 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2025 * 21:40 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2025 * 21:40 sbassett@deploy2003: sbassett: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:38 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] * 21:37 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2025 * 21:37 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2025.codfw.wmnet 220.48.192.10.in-addr.arpa 0.2.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:37 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2025.codfw.wmnet 220.48.192.10.in-addr.arpa 0.2.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:37 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:37 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2025 - bking@cumin2003" * 21:37 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2025 - bking@cumin2003" * 21:30 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] (duration: 08m 19s) * 21:26 sbassett@deploy2003: sbassett: Continuing with deployment * 21:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1269.eqiad.wmnet with OS trixie * 21:24 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1271.eqiad.wmnet with reason: host reimage * 21:23 sbassett@deploy2003: sbassett: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:22 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:22 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] * 21:20 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2025 * 21:19 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2025.codfw.wmnet with OS bookworm * 21:17 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1271.eqiad.wmnet with reason: host reimage * 21:04 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1269.eqiad.wmnet with reason: host reimage * 21:00 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1269.eqiad.wmnet with reason: host reimage * 20:56 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1271 * 20:55 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1271 * 20:54 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1271 * 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1271.eqiad.wmnet 126.48.64.10.in-addr.arpa 6.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:54 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1271.eqiad.wmnet 126.48.64.10.in-addr.arpa 6.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1271 - swfrench@cumin1003" * 20:54 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1271 - swfrench@cumin1003" * 20:51 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:51 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:51 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:50 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 20:49 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 20:49 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1268.eqiad.wmnet * 20:49 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1268.eqiad.wmnet * 20:49 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1268.eqiad.wmnet * 20:48 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1271 * 20:48 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1271.eqiad.wmnet with OS trixie * 20:47 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1271.eqiad.wmnet * 20:46 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1271.eqiad.wmnet * 20:46 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1271.eqiad.wmnet * 20:41 aude@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] (duration: 07m 34s) * 20:39 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1269 * 20:39 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1269 * 20:36 aude@deploy2003: aude: Continuing with deployment * 20:35 aude@deploy2003: aude: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:33 aude@deploy2003: Started scap sync-world: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] * 20:26 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on db2207.codfw.wmnet with reason: Host down * 20:22 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-video: apply * 20:21 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-video: apply * 20:20 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-timeline: apply * 20:20 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-timeline: apply * 20:20 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-syntaxhighlight: apply * 20:19 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-syntaxhighlight: apply * 20:19 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-media: apply * 20:18 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-media: apply * 20:18 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-constraints: apply * 20:17 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-constraints: apply * 20:17 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox: apply * 20:16 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox: apply * 20:13 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1269 * 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1269.eqiad.wmnet 80.32.64.10.in-addr.arpa 0.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:13 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1269.eqiad.wmnet 80.32.64.10.in-addr.arpa 0.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1269 - kamila@cumin1003" * 20:13 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1269 - kamila@cumin1003" * 20:09 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-flink-codfw cluster: Roll restart of jvm daemons. * 20:07 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 20:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 20:03 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-flink-codfw cluster: Roll restart of jvm daemons. * 20:03 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2207 [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94893 and previous config saved to /var/cache/conftool/dbconfig/20260716-200257-marostegui.json * 20:01 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2204 to s2 primary [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94892 and previous config saved to /var/cache/conftool/dbconfig/20260716-200157-marostegui.json * 20:00 marostegui: Starting emergency s2 codfw failover from db2207 to db2204 - [[phab:T432396|T432396]] * 19:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1035.eqiad.wmnet * 19:56 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2204 with weight 0 [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94891 and previous config saved to /var/cache/conftool/dbconfig/20260716-195628-marostegui.json * 19:55 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 26 hosts with reason: Primary switchover s2 [[phab:T432396|T432396]] * 19:54 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1035.eqiad.wmnet * 19:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1034.eqiad.wmnet * 19:48 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1034.eqiad.wmnet * 19:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1033.eqiad.wmnet * 19:43 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-video: apply * 19:43 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1033.eqiad.wmnet * 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1032.eqiad.wmnet * 19:42 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-video: apply * 19:42 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-timeline: apply * 19:41 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-timeline: apply * 19:41 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-syntaxhighlight: apply * 19:41 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-syntaxhighlight: apply * 19:40 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-media: apply * 19:40 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-media: apply * 19:39 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-constraints: apply * 19:36 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-constraints: apply * 19:36 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox: apply * 19:35 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1032.eqiad.wmnet * 19:35 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1031.eqiad.wmnet * 19:35 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox: apply * 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-video: apply * 19:33 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-video: apply * 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-timeline: apply * 19:33 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-timeline: apply * 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-syntaxhighlight: apply * 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-syntaxhighlight: apply * 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-media: apply * 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-media: apply * 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-constraints: apply * 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-constraints: apply * 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox: apply * 19:31 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox: apply * 19:27 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1031.eqiad.wmnet * 19:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1030.eqiad.wmnet * 19:23 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1001.eqiad.wmnet, repooling source-only afterwards * 19:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1030.eqiad.wmnet * 19:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1029.eqiad.wmnet * 19:17 kamila@cumin1003: START - Cookbook sre.dns.netbox * 19:12 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1029.eqiad.wmnet * 19:06 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1269 * 19:05 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1269.eqiad.wmnet with OS trixie * 19:03 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1269.eqiad.wmnet * 19:03 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1269.eqiad.wmnet * 19:03 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1269.eqiad.wmnet * 18:55 dancy@deploy2003: Finished scap sync-world: testing [[phab:T428971|T428971]] (duration: 02m 41s) * 18:53 dancy@deploy2003: Started scap sync-world: testing [[phab:T428971|T428971]] * 18:31 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1268.eqiad.wmnet with OS trixie * 18:18 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 18:16 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1267.eqiad.wmnet * 18:16 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1267.eqiad.wmnet * 18:16 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1267.eqiad.wmnet * 18:09 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1268.eqiad.wmnet with reason: host reimage * 18:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1001.eqiad.wmnet, repooling source-only afterwards * 18:06 swfrench@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] (duration: 07m 34s) * 18:06 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 23s) * 18:06 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 18:06 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1268.eqiad.wmnet with reason: host reimage * 18:03 bd808@deploy2003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 18:02 bd808@deploy2003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 18:02 swfrench@deploy2003: jiji, swfrench: Continuing with deployment * 18:02 bd808@deploy2003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 18:02 bd808@deploy2003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 18:01 bd808@deploy2003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 18:01 swfrench@deploy2003: jiji, swfrench: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:01 bd808@deploy2003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:59 swfrench@deploy2003: Started scap sync-world: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] * 17:45 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1268 * 17:45 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1268 * 17:44 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1267.eqiad.wmnet with OS trixie * 17:43 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter2006.codfw.wmnet * 17:39 swfrench@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter2006.codfw.wmnet * 17:35 swfrench@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] (duration: 07m 27s) * 17:34 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1268 * 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1268.eqiad.wmnet 78.32.64.10.in-addr.arpa 8.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:34 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1268.eqiad.wmnet 78.32.64.10.in-addr.arpa 8.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1268 - kamila@cumin1003" * 17:34 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1268 - kamila@cumin1003" * 17:31 swfrench@deploy2003: jiji, swfrench: Continuing with deployment * 17:29 swfrench@deploy2003: jiji, swfrench: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:28 kamila@cumin1003: START - Cookbook sre.dns.netbox * 17:28 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1268 * 17:28 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1268.eqiad.wmnet with OS trixie * 17:27 swfrench@deploy2003: Started scap sync-world: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] * 17:23 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1267.eqiad.wmnet with reason: host reimage * 17:18 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1267.eqiad.wmnet with reason: host reimage * 17:18 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1268.eqiad.wmnet * 17:17 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1268.eqiad.wmnet * 17:17 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1268.eqiad.wmnet * 17:12 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter2005.codfw.wmnet * 17:11 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1270.eqiad.wmnet * 17:11 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1270.eqiad.wmnet * 17:11 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1270.eqiad.wmnet * 17:09 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter2005.codfw.wmnet * 17:08 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:08 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update reverse dns for moved arelion cct cr2-eqiad - cmooney@cumin1003" * 17:08 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update reverse dns for moved arelion cct cr2-eqiad - cmooney@cumin1003" * 17:08 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] (duration: 07m 34s) * 17:04 jiji@deploy2003: jiji: Continuing with deployment * 17:03 jiji@deploy2003: jiji: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 17:00 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] * 17:00 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2185.codfw.wmnet with OS trixie * 16:59 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:58 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1270.eqiad.wmnet with OS trixie * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1267 * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1267 * 16:57 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1267 * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1267.eqiad.wmnet 77.32.64.10.in-addr.arpa 7.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:57 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1267.eqiad.wmnet 77.32.64.10.in-addr.arpa 7.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1267 - kamila@cumin1003" * 16:56 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1267 - kamila@cumin1003" * 16:56 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_eqsin * 16:56 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5032.eqsin.wmnet * 16:52 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_esams * 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3073.esams.wmnet * 16:50 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_esams * 16:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3081.esams.wmnet * 16:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1266.eqiad.wmnet * 16:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1266.eqiad.wmnet * 16:45 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1266.eqiad.wmnet * 16:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2185.codfw.wmnet with reason: host reimage * 16:41 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_eqiad * 16:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1114.eqiad.wmnet * 16:41 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_eqiad * 16:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1115.eqiad.wmnet * 16:39 kamila@cumin1003: START - Cookbook sre.dns.netbox * 16:39 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1267 * 16:39 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2185.codfw.wmnet with reason: host reimage * 16:38 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1267.eqiad.wmnet with OS trixie * 16:38 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1267.eqiad.wmnet * 16:38 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1270.eqiad.wmnet with reason: host reimage * 16:37 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1267.eqiad.wmnet * 16:37 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1267.eqiad.wmnet * 16:31 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1270.eqiad.wmnet with reason: host reimage * 16:24 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1264.eqiad.wmnet * 16:24 kamila@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 16:24 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 16:23 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 16:21 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2162: switch maintenance completed codfw rack b6 * 16:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2185.codfw.wmnet with OS trixie * 16:19 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 16:16 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter1007.eqiad.wmnet * 16:15 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_eqsin * 16:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5024.eqsin.wmnet * 16:13 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5031.eqsin.wmnet * 16:13 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3072.esams.wmnet * 16:12 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter1007.eqiad.wmnet * 16:11 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] (duration: 09m 47s) * 16:10 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1266.eqiad.wmnet with OS trixie * 16:10 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1270 * 16:10 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1270 * 16:09 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1270 * 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1270.eqiad.wmnet 125.48.64.10.in-addr.arpa 5.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1270.eqiad.wmnet 125.48.64.10.in-addr.arpa 5.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1270 - swfrench@cumin1003" * 16:09 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1270 - swfrench@cumin1003" * 16:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3080.esams.wmnet * 16:07 jiji@deploy2003: jiji: Continuing with deployment * 16:06 jiji@deploy2003: jiji: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:04 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 16:04 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1265.eqiad.wmnet * 16:03 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1265.eqiad.wmnet * 16:03 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1265.eqiad.wmnet * 16:03 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1270 * 16:03 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1270.eqiad.wmnet with OS trixie * 16:02 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1270.eqiad.wmnet * 16:02 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] * 16:01 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1270.eqiad.wmnet * 16:01 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1270.eqiad.wmnet * 16:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1113.eqiad.wmnet * 16:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1112.eqiad.wmnet * 15:49 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1266.eqiad.wmnet with reason: host reimage * 15:47 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter1006.eqiad.wmnet * 15:45 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1265.eqiad.wmnet with OS trixie * 15:44 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1266.eqiad.wmnet with reason: host reimage * 15:43 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter1006.eqiad.wmnet * 15:42 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] (duration: 09m 46s) * 15:37 jiji@deploy2003: jiji: Continuing with deployment * 15:36 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2162: switch maintenance completed codfw rack b6 * 15:36 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2161: switch maintenance completed codfw rack b6 * 15:34 jiji@deploy2003: jiji: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:32 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5023.eqsin.wmnet * 15:32 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] * 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5030.eqsin.wmnet * 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3071.esams.wmnet * 15:27 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3079.esams.wmnet * 15:25 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1265.eqiad.wmnet with reason: host reimage * 15:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1266 * 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1266 * 15:21 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1110.eqiad.wmnet * 15:20 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1111.eqiad.wmnet * 15:16 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1265.eqiad.wmnet with reason: host reimage * 15:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1001.eqiad.wmnet with OS bookworm * 15:15 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1266 * 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1266.eqiad.wmnet 76.32.64.10.in-addr.arpa 6.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:15 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1266.eqiad.wmnet 76.32.64.10.in-addr.arpa 6.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1266 - kamila@cumin1003" * 15:15 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1266 - kamila@cumin1003" * 15:07 kamila@cumin1003: START - Cookbook sre.dns.netbox * 15:04 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1266 * 15:04 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1264 * 15:04 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1264 * 15:04 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1266.eqiad.wmnet with OS trixie * 15:03 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1264 * 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1264.eqiad.wmnet 74.32.64.10.in-addr.arpa 4.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:03 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1264.eqiad.wmnet 74.32.64.10.in-addr.arpa 4.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1264 - kamila@cumin1003" * 15:03 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1264 - kamila@cumin1003" * 15:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-eqiad * 15:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp1001.eqiad.wmnet * 15:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp1001.eqiad.wmnet * 15:01 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp1001.eqiad.wmnet * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp1001.eqiad.wmnet * 15:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1376-1384].eqiad.wmnet * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1376-1384].eqiad.wmnet * 14:59 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1002.eqiad.wmnet * 14:59 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1266.eqiad.wmnet * 14:58 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1266.eqiad.wmnet * 14:58 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1266.eqiad.wmnet * 14:58 kamila@cumin1003: START - Cookbook sre.dns.netbox * 14:57 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1264 * 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1265 * 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1265 * 14:57 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1310596{{!}}Set $wgMathInternalRestbaseURL explicitly (T349582)]] * 14:57 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1265 * 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1265.eqiad.wmnet 75.32.64.10.in-addr.arpa 5.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:56 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1265.eqiad.wmnet 75.32.64.10.in-addr.arpa 5.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:56 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:56 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1265 - kamila@cumin1003" * 14:56 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1265 - kamila@cumin1003" * 14:53 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1002.eqiad.wmnet * 14:53 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1376-1384].eqiad.wmnet * 14:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1001.eqiad.wmnet with reason: host reimage * 14:51 kamila@cumin1003: START - Cookbook sre.dns.netbox * 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 14:50 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:50 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2161: switch maintenance completed codfw rack b6 * 14:50 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5021.eqsin.wmnet * 14:50 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1265 * 14:49 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:49 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:49 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1265.eqiad.wmnet with OS trixie * 14:49 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5029.eqsin.wmnet * 14:49 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1265.eqiad.wmnet * 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3070.esams.wmnet * 14:48 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1264.eqiad.wmnet * 14:48 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1001.eqiad.wmnet with reason: host reimage * 14:48 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1376-1384].eqiad.wmnet * 14:48 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1265.eqiad.wmnet * 14:47 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1265.eqiad.wmnet * 14:47 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1264.eqiad.wmnet * 14:47 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1264.eqiad.wmnet * 14:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:47 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3078.esams.wmnet * 14:44 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1263.eqiad.wmnet * 14:44 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1263.eqiad.wmnet * 14:44 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1263.eqiad.wmnet * 14:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1108.eqiad.wmnet * 14:40 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1109.eqiad.wmnet * 14:40 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:35 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:34 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:34 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2006.codfw.wmnet * 14:34 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-flink-eqiad cluster: Roll restart of jvm daemons. * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf2002.codfw.wmnet * 14:31 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1002.eqiad.wmnet * 14:29 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2006.codfw.wmnet * 14:27 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:27 kamila@deploy2003: Finished scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] (duration: 02m 57s) * 14:27 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-flink-eqiad cluster: Roll restart of jvm daemons. * 14:26 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf2002.codfw.wmnet * 14:26 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf2001.codfw.wmnet * 14:25 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1002.eqiad.wmnet * 14:25 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1001.eqiad.wmnet * 14:25 kamila@deploy2003: Started scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] * 14:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:21 kamila@deploy2003: sync-world aborted: Test deployment to check rsync is working - [[phab:T432108|T432108]] (duration: 00m 36s) * 14:21 topranks: reboot lsw1-b6-codfw to upgrade JunOS [[phab:T430922|T430922]] * 14:21 kamila@deploy2003: Started scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] * 14:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1001.eqiad.wmnet with OS bookworm * 14:20 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b6-codfw,lsw1-b6-codfw IPv6,lsw1-b6-codfw.mgmt,ssw1-a[1,8]-codfw with reason: lsw1-b6-codfw JunOS upgrade * 14:20 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf2001.codfw.wmnet * 14:19 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1001.eqiad.wmnet * 14:19 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 26 hosts with reason: lsw1-b6-codfw JunOS upgrade * 14:14 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 14:13 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc2022: switch maintenance codfw rack b6 * 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:12 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.parsercache * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool pc2022: switch maintenance codfw rack b6 * 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2251: switch maintenance codfw rack b6 * 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.parsercache * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2251: switch maintenance codfw rack b6 * 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2162: switch maintenance codfw rack b6 * 14:12 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1263.eqiad.wmnet with OS trixie * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2162: switch maintenance codfw rack b6 * 14:11 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2161: switch maintenance codfw rack b6 * 14:11 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2161: switch maintenance codfw rack b6 * 14:08 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5020.eqsin.wmnet * 14:07 btullis@cumin1003: START - Cookbook sre.hadoop.reboot-workers for Hadoop analytics cluster * 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3069.esams.wmnet * 14:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5028.eqsin.wmnet * 14:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1338-1347].eqiad.wmnet * 14:06 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1338-1347].eqiad.wmnet * 14:05 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3077.esams.wmnet * 14:02 topranks: beginning depools for lsw1-b6-codfw maintenance [[phab:T430922|T430922]] * 14:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1106.eqiad.wmnet * 14:00 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-misc1002.eqiad.wmnet * 13:59 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1338-1347].eqiad.wmnet * 13:59 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1107.eqiad.wmnet * 13:56 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-codfw * 13:55 sfaci@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply * 13:54 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-misc1002.eqiad.wmnet * 13:54 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-misc1001.eqiad.wmnet * 13:54 sfaci@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply * 13:50 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1263.eqiad.wmnet with reason: host reimage * 13:50 sfaci@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 13:49 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1338-1347].eqiad.wmnet * 13:49 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-misc1001.eqiad.wmnet * 13:49 sfaci@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 13:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:49 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:45 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1263.eqiad.wmnet with reason: host reimage * 13:40 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:40 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-eqiad * 13:35 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:34 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:33 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-reboot (exit_code=0) rolling reboot on A:dnsbox and (A:eqsin or A:drmrs or A:magru) and not (P<nowiki>{</nowiki>dns5003*<nowiki>}</nowiki> or P<nowiki>{</nowiki>dns7002*<nowiki>}</nowiki>) and (A:dnsbox) * 13:33 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns7001.wikimedia.org * 13:27 sukhe@dns1004: END - running authdns-update * 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5019.eqsin.wmnet * 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3076.esams.wmnet * 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3068.esams.wmnet * 13:25 sukhe@dns1004: START - running authdns-update * 13:24 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5027.eqsin.wmnet * 13:24 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1263 * 13:24 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1263 * 13:23 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1263 * 13:23 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:23 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1104.eqiad.wmnet * 13:21 kamila@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:21 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:20 kamila@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:20 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:20 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:20 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1263 - kamila@cumin1003" * 13:20 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1263 - kamila@cumin1003" * 13:19 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1105.eqiad.wmnet * 13:19 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:19 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:18 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:18 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns7001.wikimedia.org * 13:16 cdobbins@cumin2003: conftool action : set/pooled=yes; selector: name=dns7002.* * 13:14 cdobbins@dns1004: END - running authdns-update * 13:13 sbisson@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] (duration: 08m 03s) * 13:13 cdobbins@dns1004: START - running authdns-update * 13:12 kamila@cumin1003: START - Cookbook sre.dns.netbox * 13:12 cdobbins@cumin2003: conftool action : set/pooled=yes; selector: name=dns7002.*,service=authdns-update * 13:12 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1263 * 13:11 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1263.eqiad.wmnet with OS trixie * 13:11 cdobbins@cumin2003: conftool action : set/pooled=no; selector: name=dns7002.* * 13:11 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1263.eqiad.wmnet * 13:10 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1263.eqiad.wmnet * 13:10 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1263.eqiad.wmnet * 13:09 sbisson@deploy2003: sbisson: Continuing with deployment * 13:08 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:07 sbisson@deploy2003: sbisson: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:05 sbisson@deploy2003: Started scap sync-world: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] * 13:03 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns6002.wikimedia.org * 13:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1298-1307].eqiad.wmnet * 13:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1298-1307].eqiad.wmnet * 12:59 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1262.eqiad.wmnet * 12:59 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1262.eqiad.wmnet * 12:59 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1262.eqiad.wmnet * 12:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1298-1307].eqiad.wmnet * 12:49 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns6002.wikimedia.org * 12:46 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1298-1307].eqiad.wmnet * 12:46 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:46 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3075.esams.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3067.esams.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5018.eqsin.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5026.eqsin.wmnet * 12:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1102.eqiad.wmnet * 12:39 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1103.eqiad.wmnet * 12:35 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:34 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns6001.wikimedia.org * 12:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:28 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:18 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns6001.wikimedia.org * 12:14 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1267-1276].eqiad.wmnet * 12:13 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1267-1276].eqiad.wmnet * 12:04 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1267-1276].eqiad.wmnet * 12:03 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns5004.wikimedia.org * 12:02 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1100.eqiad.wmnet * 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3066.esams.wmnet * 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3074.esams.wmnet * 12:01 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:01 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-codfw * 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5017.eqsin.wmnet * 12:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5025.eqsin.wmnet * 12:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1101.eqiad.wmnet * 11:59 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1267-1276].eqiad.wmnet * 11:58 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:58 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:54 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns5004.wikimedia.org * 11:54 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and (A:eqsin or A:drmrs or A:magru) and not (P<nowiki>{</nowiki>dns5003*<nowiki>}</nowiki> or P<nowiki>{</nowiki>dns7002*<nowiki>}</nowiki>) and (A:dnsbox) * 11:54 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:53 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:53 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-eqiad * 11:51 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:51 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:50 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:50 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_eqiad * 11:50 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:50 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_eqiad * 11:49 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_esams * 11:49 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_esams * 11:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_eqsin * 11:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_eqsin * 11:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:44 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-eqiad * 11:43 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-codfw * 11:42 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:41 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:24 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-eqiad * 11:23 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-codfw * 11:22 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:15 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:14 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:09 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2066.codfw.wmnet with OS trixie * 11:05 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1151-1160].eqiad.wmnet * 11:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1151-1160].eqiad.wmnet * 10:59 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1068.eqiad.wmnet with OS trixie * 10:55 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.major-upgrade (exit_code=99) * 10:55 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 10:54 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1151-1160].eqiad.wmnet * 10:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2066.codfw.wmnet with reason: host reimage * 10:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1151-1160].eqiad.wmnet * 10:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:42 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2066.codfw.wmnet with reason: host reimage * 10:39 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:37 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:36 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:23 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:22 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2066.codfw.wmnet with OS trixie * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:07 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2065.codfw.wmnet with OS trixie * 10:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:06 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 10:06 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 10:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 10:03 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:03 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 10:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 09:59 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 09:57 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 09:57 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 09:52 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 09:47 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:46 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:46 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2065.codfw.wmnet with reason: host reimage * 09:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2065.codfw.wmnet with reason: host reimage * 09:40 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:39 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 09:39 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:39 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:39 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:37 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:29 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox2003.codfw.wmnet * 09:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox2003.codfw.wmnet * 09:25 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:25 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:24 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:24 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:24 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:21 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:20 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2065.codfw.wmnet with OS trixie * 09:13 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1068.eqiad.wmnet with OS trixie * 09:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2064.codfw.wmnet with OS trixie * 09:08 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 09:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 09:07 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2162: Repooling after switchover * 09:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1067.eqiad.wmnet with OS trixie * 09:00 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:59 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 08:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 08:57 tappof: bump space for prometheus k8s-dse in eqiad * 08:56 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping2004.codfw.wmnet * 08:52 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host ping2004.codfw.wmnet * 08:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 08:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping1004.eqiad.wmnet * 08:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:51 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:49 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 08:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host ping1004.eqiad.wmnet * 08:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2064.codfw.wmnet with reason: host reimage * 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2064.codfw.wmnet with reason: host reimage * 08:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1067.eqiad.wmnet with reason: host reimage * 08:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:33 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1067.eqiad.wmnet with reason: host reimage * 08:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:21 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2162: Repooling after switchover * 08:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2064.codfw.wmnet with OS trixie * 08:16 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1067.eqiad.wmnet with OS trixie * 08:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:15 cgoubert@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-eqiad * 08:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2062.codfw.wmnet with OS trixie * 08:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1066.eqiad.wmnet with OS trixie * 08:02 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2162: Repooling after switchover * 07:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2162: Repooling after switchover * 07:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2162 [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94870 and previous config saved to /var/cache/conftool/dbconfig/20260716-075530-cwilliams.json * 07:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2241 to x3 primary [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94869 and previous config saved to /var/cache/conftool/dbconfig/20260716-075314-cwilliams.json * 07:52 cezmunsta: Starting x3 codfw failover from db2162 to db2241 - [[phab:T430925|T430925]] * 07:50 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:50 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:47 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 07:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2241 with weight 0 [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94868 and previous config saved to /var/cache/conftool/dbconfig/20260716-074507-cwilliams.json * 07:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 18 hosts with reason: Primary switchover x3 [[phab:T430925|T430925]] * 07:43 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1066.eqiad.wmnet with reason: host reimage * 07:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:dse-k8s-worker-eqiad * 07:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1028.eqiad.wmnet * 07:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1028.eqiad.wmnet * 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 07:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1066.eqiad.wmnet with reason: host reimage * 07:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1028.eqiad.wmnet * 07:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1028.eqiad.wmnet * 07:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1027.eqiad.wmnet * 07:35 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1027.eqiad.wmnet * 07:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1027.eqiad.wmnet * 07:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1027.eqiad.wmnet * 07:28 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1026.eqiad.wmnet * 07:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1026.eqiad.wmnet * 07:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast2003.wikimedia.org * 07:21 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1026.eqiad.wmnet * 07:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1066.eqiad.wmnet with OS trixie * 07:19 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast2003.wikimedia.org * 07:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2062.codfw.wmnet with OS trixie * 06:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1026.eqiad.wmnet * 06:51 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1025.eqiad.wmnet * 06:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1025.eqiad.wmnet * 06:47 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 06:47 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 06:44 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1025.eqiad.wmnet * 06:14 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1025.eqiad.wmnet * 06:14 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1024.eqiad.wmnet * 06:14 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1024.eqiad.wmnet * 06:07 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1024.eqiad.wmnet * 05:37 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1024.eqiad.wmnet * 05:37 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1023.eqiad.wmnet * 05:37 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1023.eqiad.wmnet * 05:26 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1023.eqiad.wmnet * 04:56 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1023.eqiad.wmnet * 04:56 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1022.eqiad.wmnet * 04:56 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1022.eqiad.wmnet * 04:49 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1022.eqiad.wmnet * 04:19 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1022.eqiad.wmnet * 04:19 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1021.eqiad.wmnet * 04:19 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1021.eqiad.wmnet * 04:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1021.eqiad.wmnet * 03:38 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1021.eqiad.wmnet * 03:38 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1020.eqiad.wmnet * 03:38 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1020.eqiad.wmnet * 03:20 btullis@cumin1003: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1020.eqiad.wmnet * 03:18 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1020.eqiad.wmnet * 03:18 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1019.eqiad.wmnet * 03:18 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1019.eqiad.wmnet * 03:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1019.eqiad.wmnet * 02:41 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1019.eqiad.wmnet * 02:41 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1018.eqiad.wmnet * 02:41 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1018.eqiad.wmnet * 02:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling both afterwards * 02:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2003.codfw.wmnet -> wcqs2001.codfw.wmnet, repooling both afterwards * 02:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1018.eqiad.wmnet * 02:30 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1018.eqiad.wmnet * 02:30 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1014.eqiad.wmnet * 02:30 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1014.eqiad.wmnet * 02:24 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1014.eqiad.wmnet * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 01:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1014.eqiad.wmnet * 01:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1013.eqiad.wmnet * 01:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1013.eqiad.wmnet * 01:47 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1013.eqiad.wmnet * 01:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2003.codfw.wmnet -> wcqs2001.codfw.wmnet, repooling both afterwards * 01:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling both afterwards * 01:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1013.eqiad.wmnet * 01:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1012.eqiad.wmnet * 01:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1012.eqiad.wmnet * 01:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1012.eqiad.wmnet * 01:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1012.eqiad.wmnet * 01:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1011.eqiad.wmnet * 01:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1011.eqiad.wmnet * 01:08 ryankemper@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] scap deploy post bookworm reimage (duration: 00m 23s) * 01:08 ryankemper@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] scap deploy post bookworm reimage * 01:08 ryankemper@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): scap deploy post bookworm reimage (duration: 00m 46s) * 01:07 ryankemper@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): scap deploy post bookworm reimage * 01:04 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1011.eqiad.wmnet * 01:04 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1011.eqiad.wmnet * 01:04 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1010.eqiad.wmnet * 01:04 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1010.eqiad.wmnet * 00:57 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1010.eqiad.wmnet * 00:57 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1010.eqiad.wmnet * 00:57 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1009.eqiad.wmnet * 00:57 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1009.eqiad.wmnet * 00:50 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1009.eqiad.wmnet * 00:20 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1009.eqiad.wmnet * 00:20 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1008.eqiad.wmnet * 00:20 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1008.eqiad.wmnet * 00:13 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1008.eqiad.wmnet == 2026-07-15 == * 23:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2001.codfw.wmnet with OS bookworm * 23:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1008.eqiad.wmnet * 23:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1007.eqiad.wmnet * 23:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1007.eqiad.wmnet * 23:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1007.eqiad.wmnet * 23:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1007.eqiad.wmnet * 23:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1006.eqiad.wmnet * 23:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1006.eqiad.wmnet * 23:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1006.eqiad.wmnet * 23:29 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1006.eqiad.wmnet * 23:28 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1005.eqiad.wmnet * 23:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1005.eqiad.wmnet * 23:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1002.eqiad.wmnet with OS bookworm * 23:21 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1005.eqiad.wmnet * 23:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2001.codfw.wmnet with reason: host reimage * 23:15 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host datahubsearch1001.eqiad.wmnet with OS bookworm * 23:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2001.codfw.wmnet with reason: host reimage * 23:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 23:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 22:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 22:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1005.eqiad.wmnet * 22:51 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1004.eqiad.wmnet * 22:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1004.eqiad.wmnet * 22:45 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1004.eqiad.wmnet * 22:44 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 22:44 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS trixie * 22:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host datahubsearch1001.eqiad.wmnet with OS bookworm * 22:34 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host datahubsearch1001.eqiad.wmnet with OS bookworm * 22:16 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on datahubsearch[1002-1003].eqiad.wmnet with reason: Using datahubsearch1001 to test bookworm reimages * 22:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1004.eqiad.wmnet * 22:15 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1003.eqiad.wmnet * 22:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1003.eqiad.wmnet * 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 22:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1003.eqiad.wmnet * 22:08 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1003.eqiad.wmnet * 22:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1002.eqiad.wmnet * 22:08 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1002.eqiad.wmnet * 22:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host datahubsearch1001.eqiad.wmnet with OS bookworm * 22:05 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 22:02 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm * 22:01 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on datahubsearch[1001-1003].eqiad.wmnet with reason: Using datahubsearch1001 to test bookworm reimages * 22:01 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1002.eqiad.wmnet * 22:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1002.eqiad.wmnet * 22:00 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1001.eqiad.wmnet * 22:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1001.eqiad.wmnet * 21:53 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1001.eqiad.wmnet * 21:52 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 21:50 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 21:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS trixie * 21:50 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS bookworm * 21:43 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 21:38 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:30 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wcqs1002'] * 21:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:29 lerickson@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 21:29 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:29 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:29 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS bookworm * 21:28 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 21:28 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm * 21:23 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1001.eqiad.wmnet * 21:23 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:23 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:22 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 21:20 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:18 swfrench-wmf: reprepro include php8.3_8.3.32-1+wmf11u2 into component/php83 for bullseye-wikimedia * 21:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:16 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:15 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-druid-public cluster: Roll restart of jvm daemons. * 21:08 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:05 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1001.eqiad.wmnet * 21:05 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1001.eqiad.wmnet * 21:04 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-druid-public cluster: Roll restart of jvm daemons. * 21:02 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 21:01 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 21:01 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 21:00 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 20:59 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1001.eqiad.wmnet * 20:59 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1001.eqiad.wmnet * 20:59 btullis@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:dse-k8s-worker-eqiad * 20:55 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 20:55 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 20:45 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm * 20:21 jhathaway: puppet is re-enabled, have fun, but not too much fun! * 20:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 20:17 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs2001'] * 20:12 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs2001'] * 20:11 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs2001'] * 20:09 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:08 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 20:05 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:05 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 20:04 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs2001'] * 20:03 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 20:03 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm * 20:02 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:02 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 20:01 jhathaway: disabling puppet fleet wide to roll out kafka patch * 19:55 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host relforge1010.eqiad.wmnet * 19:52 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 19:52 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 19:48 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 19:48 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 19:48 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 19:47 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 19:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 19:45 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 19:45 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 19:44 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1010.eqiad.wmnet * 19:38 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1262.eqiad.wmnet with OS trixie * 19:17 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1262.eqiad.wmnet with reason: host reimage * 19:11 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1262.eqiad.wmnet with reason: host reimage * 18:59 cdobbins@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS trixie * 18:54 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 18:53 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 18:52 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1262 * 18:52 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1262 * 18:51 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1262 * 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1262.eqiad.wmnet 72.32.64.10.in-addr.arpa 2.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:51 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1262.eqiad.wmnet 72.32.64.10.in-addr.arpa 2.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1262 - kamila@cumin1003" * 18:51 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1262 - kamila@cumin1003" * 18:46 kamila@cumin1003: START - Cookbook sre.dns.netbox * 18:46 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1262 * 18:46 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 18:46 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ncmonitor1001.eqiad.wmnet * 18:46 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 18:45 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1262.eqiad.wmnet with OS trixie * 18:45 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 18:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1262.eqiad.wmnet * 18:44 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1262.eqiad.wmnet * 18:44 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1262.eqiad.wmnet * 18:42 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host ncmonitor1001.eqiad.wmnet * 18:29 topranks: pull power on cr1-eqiad to install new switch-control boards [[phab:T426343|T426343]] * 18:29 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs[1018-1020].eqiad.wmnet with reason: line card install in cr1-eqiad * 18:27 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 14 hosts with reason: linecard install in cr1-eqad * 18:22 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_ulsfo * 18:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4052.ulsfo.wmnet * 18:19 cdobbins@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 18:15 cdobbins@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 18:14 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_drmrs * 18:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6016.drmrs.wmnet * 18:12 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_ulsfo * 18:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4044.ulsfo.wmnet * 18:10 sukhe@cumin1003: END (ERROR) - Cookbook sre.cdn.roll-reboot (exit_code=97) rolling reboot on A:cp-upload_drmrs * 18:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2241: Security update * 17:56 topranks: start draining traffic on cr1-eqiad ahead of line card installation [[phab:T426343|T426343]] * 17:47 cdobbins@cumin2003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie * 17:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4051.ulsfo.wmnet * 17:40 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:39 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 17:34 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6007.drmrs.wmnet * 17:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6015.drmrs.wmnet * 17:32 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:31 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 17:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4043.ulsfo.wmnet * 17:27 lerickson@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:25 lerickson@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 17:22 lerickson@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-codfw * 17:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp2001.codfw.wmnet * 17:22 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 17:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp2001.codfw.wmnet * 17:22 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 17:19 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2241: Security update * 17:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2241.codfw.wmnet * 17:17 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2241.codfw.wmnet * 17:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp2001.codfw.wmnet * 17:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp2001.codfw.wmnet * 17:15 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2366-2374].codfw.wmnet * 17:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2366-2374].codfw.wmnet * 17:10 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply * 17:10 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply * 17:08 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2366-2374].codfw.wmnet * 17:06 sukhe: sre.dns.roll-reboot to resume later * 17:06 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-reboot (exit_code=97) rolling reboot on A:dnsbox and not (A:ulsfo or A:magru) and (A:dnsbox) * 17:06 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns5003.wikimedia.org * 17:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2241: Security update * 17:03 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2241: Security update * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply * 17:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2366-2374].codfw.wmnet * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply * 17:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2357-2365].codfw.wmnet * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 17:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2357-2365].codfw.wmnet * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply * 16:55 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2357-2365].codfw.wmnet * 16:55 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 16:53 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6006.drmrs.wmnet * 16:52 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6014.drmrs.wmnet * 16:52 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 16:51 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 16:50 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2357-2365].codfw.wmnet * 16:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4042.ulsfo.wmnet * 16:50 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2347-2356].codfw.wmnet * 16:50 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2347-2356].codfw.wmnet * 16:49 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns5003.wikimedia.org * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply * 16:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4050.ulsfo.wmnet * 16:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2347-2356].codfw.wmnet * 16:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2347-2356].codfw.wmnet * 16:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2337-2346].codfw.wmnet * 16:36 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2337-2346].codfw.wmnet * 16:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:dse-k8s-worker-codfw * 16:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2003.codfw.wmnet * 16:35 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2003.codfw.wmnet * 16:34 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns3004.wikimedia.org * 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply * 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply * 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply * 16:30 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply * 16:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2003.codfw.wmnet * 16:29 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2337-2346].codfw.wmnet * 16:24 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2003.codfw.wmnet * 16:24 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2002.codfw.wmnet * 16:24 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2002.codfw.wmnet * 16:23 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns3004.wikimedia.org * 16:23 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2337-2346].codfw.wmnet * 16:23 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2327-2336].codfw.wmnet * 16:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2327-2336].codfw.wmnet * 16:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2002.codfw.wmnet * 16:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2327-2336].codfw.wmnet * 16:12 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2002.codfw.wmnet * 16:12 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2001.codfw.wmnet * 16:12 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2001.codfw.wmnet * 16:12 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1065.eqiad.wmnet with OS trixie * 16:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6005.drmrs.wmnet * 16:11 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6013.drmrs.wmnet * 16:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4041.ulsfo.wmnet * 16:08 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns3003.wikimedia.org * 16:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2327-2336].codfw.wmnet * 16:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2317-2326].codfw.wmnet * 16:06 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2317-2326].codfw.wmnet * 16:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2001.codfw.wmnet * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply * 16:03 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4049.ulsfo.wmnet * 16:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2001.codfw.wmnet * 16:00 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test2001.codfw.wmnet * 16:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test2001.codfw.wmnet * 16:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2063.codfw.wmnet with OS trixie * 15:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2317-2326].codfw.wmnet * 15:57 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns3003.wikimedia.org * 15:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test2001.codfw.wmnet * 15:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test2001.codfw.wmnet * 15:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2004.codfw.wmnet * 15:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2004.codfw.wmnet * 15:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2317-2326].codfw.wmnet * 15:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2307-2316].codfw.wmnet * 15:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2307-2316].codfw.wmnet * 15:49 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2004.codfw.wmnet * 15:48 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2004.codfw.wmnet * 15:48 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2003.codfw.wmnet * 15:48 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2003.codfw.wmnet * 15:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 15:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2307-2316].codfw.wmnet * 15:42 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2003.codfw.wmnet * 15:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 15:42 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2003.codfw.wmnet * 15:42 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2002.codfw.wmnet * 15:42 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2002.codfw.wmnet * 15:42 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2006.wikimedia.org * 15:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2063.codfw.wmnet with reason: host reimage * 15:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2307-2316].codfw.wmnet * 15:37 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2297-2306].codfw.wmnet * 15:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2297-2306].codfw.wmnet * 15:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2002.codfw.wmnet * 15:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2002.codfw.wmnet * 15:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2001.codfw.wmnet * 15:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2001.codfw.wmnet * 15:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2063.codfw.wmnet with reason: host reimage * 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6004.drmrs.wmnet * 15:31 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2001.codfw.wmnet * 15:31 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2001.codfw.wmnet * 15:31 btullis@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:dse-k8s-worker-codfw * 15:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6012.drmrs.wmnet * 15:28 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2006.wikimedia.org * 15:27 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-analytics cluster: Roll restart of jvm daemons. * 15:27 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2297-2306].codfw.wmnet * 15:27 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4040.ulsfo.wmnet * 15:24 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1065.eqiad.wmnet with OS trixie * 15:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4048.ulsfo.wmnet * 15:21 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-analytics cluster: Roll restart of jvm daemons. * 15:21 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2297-2306].codfw.wmnet * 15:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2287-2296].codfw.wmnet * 15:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2287-2296].codfw.wmnet * 15:20 btullis@cumin1003: END (PASS) - Cookbook sre.druid.reboot-workers (exit_code=0) for Druid public cluster: Reboot Druid nodes * 15:18 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 15:17 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm * 15:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2063.codfw.wmnet with OS trixie * 15:13 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2005.wikimedia.org * 15:11 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2287-2296].codfw.wmnet * 15:11 btullis@cumin1003: END (PASS) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=0) rolling reboot on A:cephosd-eqiad * 15:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1064.eqiad.wmnet with OS trixie * 15:05 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2062.codfw.wmnet with OS trixie * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2287-2296].codfw.wmnet * 15:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2277-2286].codfw.wmnet * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2277-2286].codfw.wmnet * 14:59 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2005.wikimedia.org * 14:57 brouberol@cumin1003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-jumbo-eqiad * 14:52 btullis@cumin1003: END (PASS) - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas (exit_code=0) rolling reboot on A:schema-codfw * 14:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6003.drmrs.wmnet * 14:50 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:50 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host relforge1009.eqiad.wmnet * 14:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2277-2286].codfw.wmnet * 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6011.drmrs.wmnet * 14:47 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 14:46 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>ml-serve1001.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 14:46 klausman@cumin1003: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) pool for host ml-serve1001.eqiad.wmnet * 14:46 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 14:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1001.eqiad.wmnet * 14:45 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4039.ulsfo.wmnet * 14:44 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1009.eqiad.wmnet * 14:44 btullis@cumin1003: START - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas rolling reboot on A:schema-codfw * 14:44 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2004.wikimedia.org * 14:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2277-2286].codfw.wmnet * 14:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2267-2276].codfw.wmnet * 14:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2267-2276].codfw.wmnet * 14:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 14:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4047.ulsfo.wmnet * 14:40 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1001.eqiad.wmnet * 14:38 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 14:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 14:36 topranks: disconnect power on cr2-eqiad to shut down device for switch fabric replacement [[phab:T426343|T426343]] * 14:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2267-2276].codfw.wmnet * 14:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 14:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1001.eqiad.wmnet * 14:35 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>ml-serve1001.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 14:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 14:34 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 14:33 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:33 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:30 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2004.wikimedia.org * 14:29 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2267-2276].codfw.wmnet * 14:29 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2257-2266].codfw.wmnet * 14:29 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2257-2266].codfw.wmnet * 14:24 btullis@cumin1003: END (PASS) - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas (exit_code=0) rolling reboot on A:schema-eqiad * 14:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2257-2266].codfw.wmnet * 14:20 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:20 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:19 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:17 jforrester@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2257-2266].codfw.wmnet * 14:16 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:16 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2062.codfw.wmnet with OS trixie * 14:15 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1064.eqiad.wmnet with OS trixie * 14:15 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1006.wikimedia.org * 14:15 btullis@cumin1003: START - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas rolling reboot on A:schema-eqiad * 14:14 topranks: switch routing-engine on cr2-eqiad resetting all interfaces [[phab:T417873|T417873]] * 14:11 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:11 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:10 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm * 14:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6002.drmrs.wmnet * 14:09 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6010.drmrs.wmnet * 14:06 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1006.wikimedia.org * 14:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:05 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 14:05 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4038.ulsfo.wmnet * 14:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4046.ulsfo.wmnet * 14:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:00 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on cr1-eqiad with reason: switch upgrade and line card install * 14:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:59 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:57 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:57 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:55 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-eqiad * 13:55 btullis@cumin1003: START - Cookbook sre.druid.reboot-workers for Druid public cluster: Reboot Druid nodes * 13:53 topranks: switch routing-engine on cr2-eqiad resetting all interfaces [[phab:T417873|T417873]] * 13:51 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1005.wikimedia.org * 13:50 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:49 brouberol@cumin1003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-test-eqiad * 13:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:44 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2197-2206].codfw.wmnet * 13:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2197-2206].codfw.wmnet * 13:36 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1005.wikimedia.org * 13:35 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2197-2206].codfw.wmnet * 13:30 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2197-2206].codfw.wmnet * 13:28 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6001.drmrs.wmnet * 13:28 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6009.drmrs.wmnet * 13:28 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2187-2196].codfw.wmnet * 13:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2187-2196].codfw.wmnet * 13:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 13:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4037.ulsfo.wmnet * 13:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2001 * 13:22 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2001 * 13:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4045.ulsfo.wmnet * 13:21 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1004.wikimedia.org * 13:19 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on lvs[1018-1020].eqiad.wmnet with reason: switch upgrade and line card install * 13:18 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2009.codfw.wmnet * 13:18 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2001 * 13:18 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2001.codfw.wmnet 26.16.192.10.in-addr.arpa 6.2.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:17 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2001.codfw.wmnet 26.16.192.10.in-addr.arpa 6.2.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:17 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:17 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2001 - bking@cumin2003" * 13:17 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2001 - bking@cumin2003" * 13:17 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2009.codfw.wmnet * 13:17 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_drmrs * 13:17 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2187-2196].codfw.wmnet * 13:17 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_drmrs * 13:17 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on 15 hosts with reason: switch upgrade and line card install * 13:17 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:15 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 13:13 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:13 brouberol@cumin1003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-jumbo-eqiad * 13:13 brouberol@cumin1003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-test-eqiad * 13:13 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1004.wikimedia.org * 13:13 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and not (A:ulsfo or A:magru) and (A:dnsbox) * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:12 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_ulsfo * 13:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:12 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_ulsfo * 13:11 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2187-2196].codfw.wmnet * 13:11 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 13:11 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 13:06 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 13:05 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling source-only afterwards * 13:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2001 * 13:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2009.codfw.wmnet with OS trixie * 13:03 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:03 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling source-only afterwards * 13:01 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 15s) * 13:01 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 13:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 12:57 btullis@cumin1003: END (PASS) - Cookbook sre.druid.reboot-workers (exit_code=0) for Druid analytics cluster: Reboot Druid nodes * 12:54 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 12:54 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2163-2172].codfw.wmnet * 12:54 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2163-2172].codfw.wmnet * 12:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2163-2172].codfw.wmnet * 12:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2009.codfw.wmnet with reason: host reimage * 12:41 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2163-2172].codfw.wmnet * 12:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2153-2162].codfw.wmnet * 12:40 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2153-2162].codfw.wmnet * 12:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2009.codfw.wmnet with reason: host reimage * 12:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2153-2162].codfw.wmnet * 12:29 btullis@cumin1003: END (PASS) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=0) rolling reboot on A:cephosd-codfw * 12:25 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2153-2162].codfw.wmnet * 12:25 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2143-2152].codfw.wmnet * 12:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2143-2152].codfw.wmnet * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2009 * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2009 * 12:22 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2009 * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2009.codfw.wmnet 139.0.192.10.in-addr.arpa 9.3.1.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:22 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2009.codfw.wmnet 139.0.192.10.in-addr.arpa 9.3.1.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2009 - mvernon@cumin2003" * 12:22 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2009 - mvernon@cumin2003" * 12:16 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 12:15 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 12:15 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 12:15 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2009 * 12:15 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 12:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2009.codfw.wmnet with OS trixie * 12:15 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 12:14 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 12:14 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2143-2152].codfw.wmnet * 12:13 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 12:12 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2010.codfw.wmnet * 12:11 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2010.codfw.wmnet * 12:10 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 12:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2143-2152].codfw.wmnet * 12:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2133-2142].codfw.wmnet * 12:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2133-2142].codfw.wmnet * 12:02 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 11:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2133-2142].codfw.wmnet * 11:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2133-2142].codfw.wmnet * 11:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:49 mvolz@deploy2003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:49 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-codfw * 11:48 mvolz@deploy2003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:47 btullis@cumin1003: START - Cookbook sre.druid.reboot-workers for Druid analytics cluster: Reboot Druid nodes * 11:46 mvolz@deploy2003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:46 mvolz@deploy2003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:45 mvolz@deploy2003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:44 mvolz@deploy2003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:40 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] (duration: 11m 38s) * 11:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2010.codfw.wmnet with OS trixie * 11:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2105-2114].codfw.wmnet * 11:36 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2105-2114].codfw.wmnet * 11:36 krinkle@deploy2003: physikerwelt, krinkle: Continuing with deployment * 11:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1018: Security updates * 11:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:36 root@cumin1003: START - Cookbook sre.mysql.parsercache * 11:36 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1018: Security updates * 11:31 krinkle@deploy2003: physikerwelt, krinkle: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:29 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] * 11:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2105-2114].codfw.wmnet * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2105-2114].codfw.wmnet * 11:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2010.codfw.wmnet with reason: host reimage * 11:12 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2010.codfw.wmnet with reason: host reimage * 11:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1018: Security updates * 11:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:10 root@cumin1003: START - Cookbook sre.mysql.parsercache * 11:10 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1018: Security updates * 11:09 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1009.eqiad.wmnet with OS trixie * 11:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow7002.magru.wmnet * 11:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 11:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 11:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=tegola-vector-tiles,name=eqiad * 11:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=kartotherian,name=eqiad * 11:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow7002.magru.wmnet * 10:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2010 * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2010 * 10:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 10:54 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 10:54 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2010 * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2010.codfw.wmnet 76.16.192.10.in-addr.arpa 6.7.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:54 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2010.codfw.wmnet 76.16.192.10.in-addr.arpa 6.7.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2010 - mvernon@cumin2003" * 10:54 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2010 - mvernon@cumin2003" * 10:49 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 10:49 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2010 * 10:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1009.eqiad.wmnet with reason: host reimage * 10:49 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2010.codfw.wmnet with OS trixie * 10:46 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2011.codfw.wmnet * 10:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1011.eqiad.wmnet * 10:44 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2011.codfw.wmnet * 10:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 10:44 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 10:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1009.eqiad.wmnet with reason: host reimage * 10:44 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow6001.drmrs.wmnet * 10:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1017: Security updates * 10:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:39 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1017: Security updates * 10:39 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow6001.drmrs.wmnet * 10:38 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1011.eqiad.wmnet * 10:35 cgoubert@deploy2003: Finished deploy [restbase/deploy@06301bd]: Deploying {{Gerrit|1306088}} {{Gerrit|1308347}} - [[phab:T429944|T429944]] [[phab:T428279|T428279]] (duration: 28m 34s) * 10:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1012.eqiad.wmnet * 10:35 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 10:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow5003.eqsin.wmnet * 10:34 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2011.codfw.wmnet with OS trixie * 10:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1009.eqiad.wmnet with OS trixie * 10:28 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1012.eqiad.wmnet * 10:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1013.eqiad.wmnet * 10:27 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:27 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow5003.eqsin.wmnet * 10:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:26 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow4003.ulsfo.wmnet * 10:25 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 10:25 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 10:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow4003.ulsfo.wmnet * 10:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1013.eqiad.wmnet * 10:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1014.eqiad.wmnet * 10:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:15 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2011.codfw.wmnet with reason: host reimage * 10:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1017: Security updates * 10:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:14 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:14 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1017: Security updates * 10:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow3004.esams.wmnet * 10:11 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2011.codfw.wmnet with reason: host reimage * 10:10 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:10 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1014.eqiad.wmnet * 10:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki2003.codfw.wmnet * 10:09 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 10:09 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 10:09 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow3004.esams.wmnet * 10:08 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2004.codfw.wmnet * 10:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1010.eqiad.wmnet with OS trixie * 10:07 cgoubert@deploy2003: Started deploy [restbase/deploy@06301bd]: Deploying {{Gerrit|1306088}} {{Gerrit|1308347}} - [[phab:T429944|T429944]] [[phab:T428279|T428279]] * 10:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host rpki2003.codfw.wmnet * 10:04 topranks: push out config change to BGP_outfilter on core routers [[phab:T431849|T431849]] * 10:02 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow2004.codfw.wmnet * 09:59 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 09:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2003.codfw.wmnet * 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2011 * 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2011 * 09:53 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 09:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:52 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2011 * 09:52 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2011.codfw.wmnet 36.32.192.10.in-addr.arpa 6.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:52 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2011.codfw.wmnet 36.32.192.10.in-addr.arpa 6.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:51 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:51 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2011 - mvernon@cumin2003" * 09:51 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2011 - mvernon@cumin2003" * 09:51 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow2003.codfw.wmnet * 09:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1003.eqiad.wmnet * 09:49 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 09:49 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 09:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1010.eqiad.wmnet with reason: host reimage * 09:47 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 09:47 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 09:47 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 09:47 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2011 * 09:46 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2011.codfw.wmnet with OS trixie * 09:44 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow1003.eqiad.wmnet * 09:44 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2012.codfw.wmnet * 09:44 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1002.eqiad.wmnet * 09:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1010.eqiad.wmnet with reason: host reimage * 09:43 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2012.codfw.wmnet * 09:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Security updates * 09:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:43 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:43 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Security updates * 09:42 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:40 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow1002.eqiad.wmnet * 09:40 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 09:37 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki1001.eqiad.wmnet * 09:36 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 09:36 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 09:33 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host rpki1001.eqiad.wmnet * 09:32 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:32 cgoubert@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-codfw * 09:31 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=kartotherian,name=eqiad * 09:31 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola-vector-tiles,name=eqiad * 09:31 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 09:31 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2012.codfw.wmnet with OS trixie * 09:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1010.eqiad.wmnet with OS trixie * 09:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1011.eqiad.wmnet with OS trixie * 09:21 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Security updates * 09:21 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:21 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:21 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Security updates * 09:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2012.codfw.wmnet with reason: host reimage * 09:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1011.eqiad.wmnet with reason: host reimage * 09:08 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2012.codfw.wmnet with reason: host reimage * 09:05 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1011.eqiad.wmnet with reason: host reimage * 08:55 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:52 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1011.eqiad.wmnet with OS trixie * 08:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1022: Security updates * 08:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2012 * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2012 * 08:50 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1022: Security updates * 08:50 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2012 * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2012.codfw.wmnet 44.48.192.10.in-addr.arpa 4.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:50 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2012.codfw.wmnet 44.48.192.10.in-addr.arpa 4.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2012 - mvernon@cumin2003" * 08:50 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2012 - mvernon@cumin2003" * 08:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1012.eqiad.wmnet with OS trixie * 08:44 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 08:44 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2012 * 08:43 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2012.codfw.wmnet with OS trixie * 08:42 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2013.codfw.wmnet * 08:41 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2013.codfw.wmnet * 08:35 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 08:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host krb1002.eqiad.wmnet * 08:30 elukey@dns1004: END - running authdns-update * 08:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1012.eqiad.wmnet with reason: host reimage * 08:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Security updates * 08:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:28 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:28 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Security updates * 08:27 elukey@dns1004: START - running authdns-update * 08:26 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 08:26 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host krb1002.eqiad.wmnet * 08:22 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1012.eqiad.wmnet with reason: host reimage * 08:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host krb2002.codfw.wmnet * 08:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast6003.wikimedia.org * 08:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2013.codfw.wmnet with OS trixie * 08:13 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast6003.wikimedia.org * 08:12 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast3007.wikimedia.org * 08:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host krb2002.codfw.wmnet * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Security updates * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:09 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:09 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Security updates * 08:07 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1012.eqiad.wmnet with OS trixie * 08:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast3007.wikimedia.org * 08:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast5005.wikimedia.org * 07:58 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast5005.wikimedia.org * 07:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1013.eqiad.wmnet with OS trixie * 07:53 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2013.codfw.wmnet with reason: host reimage * 07:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1021: Security updates * 07:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:53 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:53 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1021: Security updates * 07:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast1004.wikimedia.org * 07:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2013.codfw.wmnet with reason: host reimage * 07:46 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast1004.wikimedia.org * 07:40 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1013.eqiad.wmnet with reason: host reimage * 07:36 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1013.eqiad.wmnet with reason: host reimage * 07:31 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2013 * 07:31 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2013 * 07:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1021: Security updates * 07:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:30 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:30 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1021: Security updates * 07:24 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2013 * 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2013.codfw.wmnet 87.0.192.10.in-addr.arpa 7.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:24 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2013.codfw.wmnet 87.0.192.10.in-addr.arpa 7.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2013 - mvernon@cumin2003" * 07:24 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2013 - mvernon@cumin2003" * 07:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1013.eqiad.wmnet with OS trixie * 07:19 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 07:19 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2013 * 07:19 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2013.codfw.wmnet with OS trixie * 07:13 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] (duration: 07m 48s) * 07:09 kharlan@deploy2003: kharlan: Continuing with deployment * 07:08 kharlan@deploy2003: kharlan: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:06 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 01:15 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 01:14 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply == 2026-07-14 == * 22:51 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_magru * 22:51 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7016.magru.wmnet * 22:46 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_magru * 22:46 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7008.magru.wmnet * 22:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7015.magru.wmnet * 22:04 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7007.magru.wmnet * 21:29 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7014.magru.wmnet * 21:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7006.magru.wmnet * 21:13 dzahn@cumin2002: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 0:15:00 on gerrit.wikimedia.org with reason: reboot * 21:11 mutante: gerrit2003 (gerrit.wikimedia.org) - reboot for maintenance * 21:11 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on gerrit2003.wikimedia.org with reason: reboot * 20:56 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:56 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:56 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:55 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 20:48 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7013.magru.wmnet * 20:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7005.magru.wmnet * 20:41 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host phab1005.eqiad.wmnet with OS trixie * 20:28 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] (duration: 06m 47s) * 20:24 sbassett@deploy2003: sbassett: Continuing with deployment * 20:23 sbassett@deploy2003: sbassett: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:23 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on phab1005.eqiad.wmnet with reason: host reimage * 20:21 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] * 20:20 aokoth@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on phab1005.eqiad.wmnet with reason: host reimage * 20:12 jhuneidi@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] (duration: 07m 42s) * 20:07 jhuneidi@deploy2003: jhuneidi, priyankar22: Continuing with deployment * 20:06 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7012.magru.wmnet * 20:06 jhuneidi@deploy2003: jhuneidi, priyankar22: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:04 jhuneidi@deploy2003: Started scap sync-world: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] * 20:02 aokoth@cumin1003: START - Cookbook sre.hosts.reimage for host phab1005.eqiad.wmnet with OS trixie * 20:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7004.magru.wmnet * 20:00 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet * 19:57 aokoth@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet * 19:24 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7011.magru.wmnet * 19:19 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7003.magru.wmnet * 19:11 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] (duration: 08m 33s) * 19:07 jforrester@deploy2003: jforrester: Continuing with deployment * 19:04 jforrester@deploy2003: jforrester: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:02 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] * 18:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7010.magru.wmnet * 18:38 mutante: rotating phabricator-gerrit bot token (its-phabricator) * 18:18 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 17:44 swfrench@deploy2003: Finished scap sync-world: Deployment to pick up new production image (duration: 31m 44s) * 17:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7002.magru.wmnet * 17:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7009.magru.wmnet * 17:32 swfrench@deploy2003: swfrench: Continuing with deployment * 17:29 swfrench@deploy2003: swfrench: Deployment to pick up new production image synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:17 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2035: repooling after rack b5 maintenance * 17:16 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool es2035: repooling after rack b5 maintenance * 17:16 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2188: repooling after rack b5 maintenance * 17:12 swfrench@deploy2003: Started scap sync-world: Deployment to pick up new production image * 17:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7001.magru.wmnet * 16:57 swfrench-wmf: reprepro include php8.3_8.3.32-1+wmf12u2 into component/php83 for bookworm-wikimedia * 16:50 sukhe: pool cp2046 * 16:47 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4039.ulsfo.wmnet * 16:44 sukhe: sudo cumin -b31 "A:cp" "run-puppet-agent" * 16:33 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on contint1003.wikimedia.org with reason: reboot * 16:32 mutante: contint1003 - main CI server - rebooting * 16:31 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2188: repooling after rack b5 maintenance * 16:31 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2178: repooling after rack b5 maintenance * 16:29 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 16:28 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 16:28 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 16:28 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 16:18 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2014.codfw.wmnet * 16:18 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 16:17 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2014.codfw.wmnet * 16:10 mvernon@cumin1003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-thanos-proxies (exit_code=0) rolling restart_daemons on A:thanos-fe * 16:09 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 16:07 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp4039.ulsfo.wmnet * 16:07 mvernon@cumin1003: START - Cookbook sre.swift.roll-restart-reboot-swift-thanos-proxies rolling restart_daemons on A:thanos-fe * 16:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2014.codfw.wmnet with OS trixie * 15:56 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1014.eqiad.wmnet with OS trixie * 15:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2014.codfw.wmnet with reason: host reimage * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2014 * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2014 * 15:28 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2014 * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2014.codfw.wmnet 194.16.192.10.in-addr.arpa 4.9.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:28 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2014.codfw.wmnet 194.16.192.10.in-addr.arpa 4.9.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2014 - mvernon@cumin2003" * 15:28 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2014 - mvernon@cumin2003" * 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Apply title-related policies when selecting the name of the entity - kamila@cumin1003" * 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Apply title-related policies when selecting the name of the entity - kamila@cumin1003 * 15:22 kamila@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Apply title-related policies when selecting the name of the entity - kamila@cumin1003 * 15:22 kamila@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Apply title-related policies when selecting the name of the entity - kamila@cumin1003" * 15:20 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 15:20 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2014 * 15:20 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2014.codfw.wmnet with OS trixie * 15:19 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1014.eqiad.wmnet with OS trixie * 15:01 dancy@deploy2003: Installation of scap version "4.274.1" completed for 3 hosts * 15:00 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2177: repooling after rack b5 maintenance * 15:00 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2159: repooling after rack b5 maintenance * 14:59 dancy@deploy2003: Installing scap version "4.274.1" for 3 host(s) * 14:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2015.codfw.wmnet with OS trixie * 14:54 seanleong-wmde: Finished populateSitesTable for isvwiki ([[phab:T429939|T429939]]) * 14:53 javiermonton@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] (duration: 07m 35s) * 14:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1015.eqiad.wmnet with OS trixie * 14:49 javiermonton@deploy2003: javiermonton: Continuing with deployment * 14:48 javiermonton@deploy2003: javiermonton: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:46 javiermonton@deploy2003: Started scap sync-world: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] * 14:42 otto@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 14:41 otto@deploy2003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 14:41 otto@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 14:40 otto@deploy2003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 14:40 otto@deploy2003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 14:39 otto@deploy2003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 14:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2015.codfw.wmnet with reason: host reimage * 14:34 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1015.eqiad.wmnet with reason: host reimage * 14:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2015.codfw.wmnet with reason: host reimage * 14:30 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1015.eqiad.wmnet with reason: host reimage * 14:30 seanleong-wmde@deploy2003: mwscript-k8s job started: foreachwikiindblist wikidataclient extensions/Wikibase/lib/maintenance/populateSitesTable.php --force-protocol https # [[phab:T429939|T429939]] * 14:24 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling reboot on A:durum and not (A:durum-eqiad or A:durum-codfw or A:durum-esams) and A:durum * 14:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2015.codfw.wmnet with OS trixie * 14:15 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2016.codfw.wmnet with OS trixie * 14:14 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2159: repooling after rack b5 maintenance * 14:14 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1015.eqiad.wmnet with OS trixie * 14:12 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1016.eqiad.wmnet with OS trixie * 14:12 cmooney@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=pki,name=codfw * 14:12 sbisson@deploy2003: helmfile [codfw] DONE helmfile.d/services/cxserver: sync * 14:11 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2002.codfw.wmnet * 14:11 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2002.codfw.wmnet * 14:11 sbisson@deploy2003: helmfile [codfw] START helmfile.d/services/cxserver: sync * 14:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1005.wikimedia.org * 14:07 sbisson@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cxserver: sync * 14:07 sbisson@deploy2003: helmfile [eqiad] START helmfile.d/services/cxserver: sync * 14:05 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1005.wikimedia.org * 14:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader2005.wikimedia.org * 14:02 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2003.codfw.wmnet * 14:02 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2003.codfw.wmnet * 14:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=tegola-vector-tiles,name=codfw * 14:00 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=kartotherian,name=codfw * 14:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader2005.wikimedia.org * 13:58 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2016.codfw.wmnet with reason: host reimage * 13:57 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-ncredir (exit_code=0) rolling reboot on A:ncredir and A:ncredir * 13:57 sbisson@deploy2003: helmfile [staging] DONE helmfile.d/services/cxserver: sync * 13:56 sbisson@deploy2003: helmfile [staging] START helmfile.d/services/cxserver: sync * 13:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1016.eqiad.wmnet with reason: host reimage * 13:52 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:52 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:51 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2016.codfw.wmnet with reason: host reimage * 13:50 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1016.eqiad.wmnet with reason: host reimage * 13:49 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy (exit_code=0) rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 13:49 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling reboot on A:wikidough * 13:46 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-tcp-proxy (exit_code=0) rolling reboot on A:tcpproxy and A:tcpproxy * 13:43 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and not (A:durum-eqiad or A:durum-codfw or A:durum-esams) and A:durum * 13:42 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=97) rolling reboot on A:durum and A:durum * 13:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2011.codfw.wmnet * 13:36 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-reboot (exit_code=0) rolling reboot on A:dnsbox and A:ulsfo and (A:dnsbox) * 13:36 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns4004.wikimedia.org * 13:34 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1016.eqiad.wmnet with OS trixie * 13:34 topranks: reboot lsw1-b5-codfw to upgrade JunOS [[phab:T430918|T430918]] * 13:34 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2016.codfw.wmnet with OS trixie * 13:32 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2002.codfw.wmnet * 13:31 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2011.codfw.wmnet * 13:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2012.codfw.wmnet * 13:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2017.codfw.wmnet with OS trixie * 13:25 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1017.eqiad.wmnet with OS trixie * 13:24 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2012.codfw.wmnet * 13:22 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2002.codfw.wmnet * 13:22 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:22 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:22 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns4004.wikimedia.org * 13:21 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2005.codfw.wmnet * 13:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2013.codfw.wmnet * 13:18 elukey@dns1004: END - running authdns-update * 13:17 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2005.codfw.wmnet * 13:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2004.codfw.wmnet * 13:16 elukey@dns1004: START - running authdns-update * 13:16 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1046: es1046 after reimage * 13:14 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1029.eqiad.wmnet,service=s8 * 13:14 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1029.eqiad.wmnet,service=s5 * 13:13 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1029.eqiad.wmnet,service=s5 * 13:13 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1029.eqiad.wmnet,service=s8 * 13:13 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2004.codfw.wmnet * 13:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2013.codfw.wmnet * 13:11 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2014.codfw.wmnet * 13:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2003.codfw.wmnet * 13:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2017.codfw.wmnet with reason: host reimage * 13:07 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:07 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns4003.wikimedia.org * 13:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2003.codfw.wmnet * 13:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1067.eqiad.wmnet * 13:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1067.eqiad.wmnet * 13:06 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1067.eqiad.wmnet * 13:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm-test1001.wikimedia.org * 13:05 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2017.codfw.wmnet with reason: host reimage * 13:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1017.eqiad.wmnet with reason: host reimage * 13:04 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2014.codfw.wmnet * 13:03 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:02 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2188: codfw rack B5 depool for maintenance * 13:02 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_magru * 13:01 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2188: codfw rack B5 depool for maintenance * 13:01 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_magru * 13:01 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2178: codfw rack B5 depool for maintenance * 13:01 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm-test1001.wikimedia.org * 13:01 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2178: codfw rack B5 depool for maintenance * 13:01 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2177: codfw rack B5 depool for maintenance * 13:00 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2177: codfw rack B5 depool for maintenance * 12:59 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1068.eqiad.wmnet * 12:59 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1068.eqiad.wmnet * 12:58 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola-vector-tiles,name=codfw * 12:58 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2159: codfw rack B5 depool for maintenance * 12:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1017.eqiad.wmnet with reason: host reimage * 12:58 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola,name=codfw * 12:57 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=kartotherian,name=codfw * 12:57 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2159: codfw rack B5 depool for maintenance * 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 30 hosts with reason: lsw1-b5-codfw JunOS upgrade * 12:55 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lsw1-b5-codfw,lsw1-b5-codfw IPv6,lsw1-b5-codfw.mgmt,ssw1-a[1,8]-codfw.mgmt with reason: switch upgade lsw1-b5-codfw * 12:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet * 12:49 cmooney@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=pki,name=codfw * 12:49 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1067.eqiad.wmnet with OS trixie * 12:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet * 12:48 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2017.codfw.wmnet with OS trixie * 12:47 topranks: depool codfw pki in dns discovery ahead of lsw1-b5-codfw maintenance [[phab:T430918|T430918]] * 12:47 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns4003.wikimedia.org * 12:47 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and A:ulsfo and (A:dnsbox) * 12:47 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2018.codfw.wmnet * 12:45 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2018.codfw.wmnet * 12:45 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and A:durum * 12:45 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-tcp-proxy rolling reboot on A:tcpproxy and A:tcpproxy * 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host pki-root1002.eqiad.wmnet * 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1009.eqiad.wmnet * 12:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1009.eqiad.wmnet * 12:44 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 12:43 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-ncredir rolling reboot on A:ncredir and A:ncredir * 12:43 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling reboot on A:wikidough * 12:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2018.codfw.wmnet with OS trixie * 12:42 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1017.eqiad.wmnet with OS trixie * 12:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader2006.wikimedia.org * 12:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1018.eqiad.wmnet with OS trixie * 12:39 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1009.eqiad.wmnet * 12:38 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host pki-root1002.eqiad.wmnet * 12:38 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1009.eqiad.wmnet * 12:38 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1008.eqiad.wmnet * 12:38 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1008.eqiad.wmnet * 12:35 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader2006.wikimedia.org * 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1006.wikimedia.org * 12:33 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1008.eqiad.wmnet * 12:30 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1046: es1046 after reimage * 12:29 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host es1046.eqiad.wmnet with OS trixie * 12:29 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1006.wikimedia.org * 12:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test2005.wikimedia.org * 12:28 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1008.eqiad.wmnet * 12:27 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1007.eqiad.wmnet * 12:27 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1007.eqiad.wmnet * 12:27 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1067.eqiad.wmnet with reason: host reimage * 12:25 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1068.eqiad.wmnet with reason: vacuum overlarge container dbs * 12:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2018.codfw.wmnet with reason: host reimage * 12:24 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test2005.wikimedia.org * 12:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test1005.wikimedia.org * 12:23 Amir1: mwscript-k8s --follow --dblist=ores -- extensions/ORES/maintenance/PurgeScoreCache.php --model goodfaith --old ([[phab:T431159|T431159]]) * 12:22 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1007.eqiad.wmnet * 12:22 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test1005.wikimedia.org * 12:22 atsukoito: restarting pybal on lvs2013 `low-traffic` for https://gerrit.wikimedia.org/r/1310535 * 12:22 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1007.eqiad.wmnet * 12:21 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1006.eqiad.wmnet * 12:21 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1006.eqiad.wmnet * 12:20 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1018.eqiad.wmnet with reason: host reimage * 12:19 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2018.codfw.wmnet with reason: host reimage * 12:18 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1067.eqiad.wmnet with reason: host reimage * 12:16 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1006.eqiad.wmnet * 12:15 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1006.eqiad.wmnet * 12:15 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1005.eqiad.wmnet * 12:15 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1005.eqiad.wmnet * 12:15 atsukoito: restarting pybal on lvs2014 for https://gerrit.wikimedia.org/r/1310535 * 12:12 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1018.eqiad.wmnet with reason: host reimage * 12:11 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1005.eqiad.wmnet * 12:11 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1005.eqiad.wmnet * 12:10 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1004.eqiad.wmnet * 12:10 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1004.eqiad.wmnet * 12:09 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on es1046.eqiad.wmnet with reason: host reimage * 12:08 atsukoito: restarting pybal on lvs1019 `low-traffic` for https://gerrit.wikimedia.org/r/1310535 * 12:06 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1004.eqiad.wmnet * 12:06 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1004.eqiad.wmnet * 12:06 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1003.eqiad.wmnet * 12:06 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1003.eqiad.wmnet * 12:05 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on es1046.eqiad.wmnet with reason: host reimage * 12:05 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "set ml-serve1001 back to active state - cmooney@cumin1003" * 12:04 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "set ml-serve1001 back to active state - cmooney@cumin1003" * 12:04 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:02 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1003.eqiad.wmnet * 12:01 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1003.eqiad.wmnet * 12:01 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1002.eqiad.wmnet * 12:01 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1002.eqiad.wmnet * 12:01 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:59 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2018.codfw.wmnet with OS trixie * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1067 * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1067 * 11:59 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1067 * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1067.eqiad.wmnet 17.48.64.10.in-addr.arpa 7.1.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:59 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1067.eqiad.wmnet 17.48.64.10.in-addr.arpa 7.1.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1067 - blake@cumin1003" * 11:58 atsukoito: restarting pybal on lvs1018 `high-traffic2` for https://gerrit.wikimedia.org/r/1310535 * 11:57 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1002.eqiad.wmnet * 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2019.codfw.wmnet with OS trixie * 11:56 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1002.eqiad.wmnet * 11:56 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 11:56 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1018.eqiad.wmnet with OS trixie * 11:54 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 11:54 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:54 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1019.eqiad.wmnet with OS trixie * 11:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:49 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:49 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:49 aikochou@deploy2003: helmfile [codfw] DONE helmfile.d/services/changeprop: sync * 11:48 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host es1046.eqiad.wmnet with OS trixie * 11:48 aikochou@deploy2003: helmfile [codfw] START helmfile.d/services/changeprop: sync * 11:48 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310535 * 11:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1046: Reimage to Trixie * 11:44 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1046: Reimage to Trixie * 11:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5:00:00 on es1046.eqiad.wmnet with reason: Reimage to Trixie * 11:42 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:42 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:42 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 11:42 aikochou@deploy2003: helmfile [eqiad] DONE helmfile.d/services/changeprop: sync * 11:41 aikochou@deploy2003: helmfile [eqiad] START helmfile.d/services/changeprop: sync * 11:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2019.codfw.wmnet with reason: host reimage * 11:36 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] (duration: 09m 41s) * 11:36 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:36 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:35 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:35 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:32 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1019.eqiad.wmnet with reason: host reimage * 11:32 jforrester@deploy2003: jforrester, gengh: Continuing with deployment * 11:29 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2019.codfw.wmnet with reason: host reimage * 11:28 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1019.eqiad.wmnet with reason: host reimage * 11:28 jforrester@deploy2003: jforrester, gengh: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:26 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] * 11:20 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2003.codfw.wmnet * 11:20 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:19 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2003.codfw.wmnet * 11:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:12 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1019.eqiad.wmnet with OS trixie * 11:10 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2019.codfw.wmnet with OS trixie * 11:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1020.eqiad.wmnet with OS trixie * 11:10 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1067 - blake@cumin1003" * 11:09 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] (duration: 12m 12s) * 11:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2020.codfw.wmnet with OS trixie * 11:03 kharlan@deploy2003: kharlan: Continuing with deployment * 11:01 blake@cumin1003: START - Cookbook sre.dns.netbox * 11:01 kharlan@deploy2003: kharlan: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:57 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] * 10:55 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] (duration: 31m 40s) * 10:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1020.eqiad.wmnet with reason: host reimage * 10:52 marostegui@dns1004: START - running authdns-update * 10:49 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2020.codfw.wmnet with reason: host reimage * 10:49 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:48 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1020.eqiad.wmnet with reason: host reimage * 10:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2020.codfw.wmnet with reason: host reimage * 10:43 kharlan@deploy2003: kharlan: Continuing with deployment * 10:42 kharlan@deploy2003: kharlan: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:32 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1020.eqiad.wmnet with OS trixie * 10:29 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2159: Repooling after switchover * 10:29 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1067 * 10:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1021.eqiad.wmnet with OS trixie * 10:27 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1067.eqiad.wmnet with OS trixie * 10:27 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:27 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1067.eqiad.wmnet * 10:27 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:26 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1067.eqiad.wmnet * 10:26 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1067.eqiad.wmnet * 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2020.codfw.wmnet with OS trixie * 10:26 blake@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1055.eqiad.wmnet * 10:26 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1055.eqiad.wmnet * 10:26 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1055.eqiad.wmnet * 10:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2021.codfw.wmnet with OS trixie * 10:24 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] * 10:11 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1055.eqiad.wmnet with OS trixie * 10:09 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1021.eqiad.wmnet with reason: host reimage * 10:05 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2021.codfw.wmnet with reason: host reimage * 10:03 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310129 revert * 10:02 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1021.eqiad.wmnet with reason: host reimage * 10:01 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2021.codfw.wmnet with reason: host reimage * 09:58 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310129 * 09:50 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1055.eqiad.wmnet with reason: host reimage * 09:45 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1055.eqiad.wmnet with reason: host reimage * 09:45 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1021.eqiad.wmnet with OS trixie * 09:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1022.eqiad.wmnet with OS trixie * 09:44 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2159: Repooling after switchover * 09:44 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2021.codfw.wmnet with OS trixie * 09:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2022.codfw.wmnet with OS trixie * 09:31 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2159.codfw.wmnet * 09:28 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1055 * 09:28 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1055 * 09:27 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on ms-fe1022.eqiad.wmnet with reason: host reimage * 09:27 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1022.eqiad.wmnet with reason: host reimage * 09:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2022.codfw.wmnet with reason: host reimage * 09:21 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2022.codfw.wmnet with reason: host reimage * 09:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2159: Rebooting db2159.codfw.wmnet * 09:20 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2159: Rebooting db2159.codfw.wmnet * 09:18 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 09:18 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 09:18 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 09:17 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 09:16 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2159.codfw.wmnet * 09:13 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b] (thin): Regular analytics weekly train THIN [analytics/refinery@ad6e05b8] (duration: 02m 07s) * 09:11 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b] (thin): Regular analytics weekly train THIN [analytics/refinery@ad6e05b8] * 09:10 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1022.eqiad.wmnet with OS trixie * 09:07 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1023.eqiad.wmnet with OS trixie * 09:06 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b]: Regular analytics weekly train [analytics/refinery@ad6e05b8] (duration: 04m 49s) * 09:04 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2022.codfw.wmnet with OS trixie * 09:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2023.codfw.wmnet with OS trixie * 09:01 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b]: Regular analytics weekly train [analytics/refinery@ad6e05b8] * 09:01 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@ad6e05b8] (duration: 02m 01s) * 09:00 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1055 * 09:00 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1055.eqiad.wmnet 50.32.64.10.in-addr.arpa 0.5.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:00 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1055.eqiad.wmnet 50.32.64.10.in-addr.arpa 0.5.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:00 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:00 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1055 - blake@cumin1003" * 09:00 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1055 - blake@cumin1003" * 08:59 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@ad6e05b8] * 08:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2159 [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94811 and previous config saved to /var/cache/conftool/dbconfig/20260714-085624-cwilliams.json * 08:55 blake@cumin1003: START - Cookbook sre.dns.netbox * 08:55 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1055 * 08:54 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1055.eqiad.wmnet with OS trixie * 08:54 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1055.eqiad.wmnet * 08:53 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1055.eqiad.wmnet * 08:53 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1055.eqiad.wmnet * 08:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2220 to s7 primary [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94810 and previous config saved to /var/cache/conftool/dbconfig/20260714-085239-cwilliams.json * 08:51 cezmunsta: Starting s7 codfw failover from db2159 to db2220 - [[phab:T430920|T430920]] * 08:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1023.eqiad.wmnet with reason: host reimage * 08:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2220 with weight 0 [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94809 and previous config saved to /var/cache/conftool/dbconfig/20260714-084553-cwilliams.json * 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s7 [[phab:T430920|T430920]] * 08:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2023.codfw.wmnet with reason: host reimage * 08:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1023.eqiad.wmnet with reason: host reimage * 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2023.codfw.wmnet with reason: host reimage * 08:34 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:34 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:29 marostegui@dns1004: END - running authdns-update * 08:29 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox-dev2003.codfw.wmnet * 08:27 marostegui@dns1004: START - running authdns-update * 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker2*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2009.codfw.wmnet * 08:26 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2009.codfw.wmnet * 08:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1023.eqiad.wmnet with OS trixie * 08:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox-dev2003.codfw.wmnet * 08:24 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:24 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:24 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2023.codfw.wmnet with OS trixie * 08:24 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1029.eqiad.wmnet with reason: reboot * 08:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1027.eqiad.wmnet with reason: reboot * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:21 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2009.codfw.wmnet * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:20 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2009.codfw.wmnet * 08:20 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2008.codfw.wmnet * 08:20 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2008.codfw.wmnet * 08:15 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2008.codfw.wmnet * 08:14 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2008.codfw.wmnet * 08:14 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2007.codfw.wmnet * 08:14 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2007.codfw.wmnet * 08:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1024.eqiad.wmnet with OS trixie * 08:12 elukey@cumin1003: END (PASS) - Cookbook sre.pki.restart-reboot (exit_code=0) rolling reboot on P<nowiki>{</nowiki>pki*<nowiki>}</nowiki> and (A:pki) * 08:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 08:10 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 08:09 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2007.codfw.wmnet * 08:08 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2007.codfw.wmnet * 08:08 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2006.codfw.wmnet * 08:08 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2006.codfw.wmnet * 08:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2024.codfw.wmnet with OS trixie * 08:03 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2006.codfw.wmnet * 08:02 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2006.codfw.wmnet * 08:02 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2005.codfw.wmnet * 08:02 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2005.codfw.wmnet * 07:58 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2005.codfw.wmnet * 07:58 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2005.codfw.wmnet * 07:57 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2004.codfw.wmnet * 07:57 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2004.codfw.wmnet * 07:54 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki.discovery.wmnet. on all recursors * 07:54 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache pki.discovery.wmnet. on all recursors * 07:53 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2004.codfw.wmnet * 07:53 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2004.codfw.wmnet * 07:53 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2003.codfw.wmnet * 07:53 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2003.codfw.wmnet * 07:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1024.eqiad.wmnet with reason: host reimage * 07:49 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki.discovery.wmnet. on all recursors * 07:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2003.codfw.wmnet * 07:49 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache pki.discovery.wmnet. on all recursors * 07:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2024.codfw.wmnet with reason: host reimage * 07:48 elukey@cumin1003: START - Cookbook sre.pki.restart-reboot rolling reboot on P<nowiki>{</nowiki>pki*<nowiki>}</nowiki> and (A:pki) * 07:46 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1024.eqiad.wmnet with reason: host reimage * 07:45 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2024.codfw.wmnet with reason: host reimage * 07:45 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2003.codfw.wmnet * 07:44 elukey@cumin1003: END (PASS) - Cookbook sre.misc-clusters.restart-reboot-config-master (exit_code=0) rolling reboot on P<nowiki>{</nowiki>config-master*<nowiki>}</nowiki> and (A:config-master or A:config-master-eqiad or A:config-master-codfw) * 07:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2002.codfw.wmnet * 07:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2002.codfw.wmnet * 07:39 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2002.codfw.wmnet * 07:39 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) config-master.discovery.wmnet. on all recursors * 07:39 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache config-master.discovery.wmnet. on all recursors * 07:39 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2002.codfw.wmnet * 07:39 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker2*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl200*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl2003.codfw.wmnet * 07:36 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl2003.codfw.wmnet * 07:35 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) config-master.discovery.wmnet. on all recursors * 07:35 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache config-master.discovery.wmnet. on all recursors * 07:34 elukey@cumin1003: START - Cookbook sre.misc-clusters.restart-reboot-config-master rolling reboot on P<nowiki>{</nowiki>config-master*<nowiki>}</nowiki> and (A:config-master or A:config-master-eqiad or A:config-master-codfw) * 07:31 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl2003.codfw.wmnet * 07:31 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl2003.codfw.wmnet * 07:31 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl2002.codfw.wmnet * 07:31 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl2002.codfw.wmnet * 07:29 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1024.eqiad.wmnet with OS trixie * 07:28 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2024.codfw.wmnet with OS trixie * 07:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl2002.codfw.wmnet * 07:26 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl2002.codfw.wmnet * 07:26 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl200*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 07:26 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 07:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 06:50 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lists1004.wikimedia.org * 06:44 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host lists1004.wikimedia.org * 06:25 marostegui@dns1004: END - running authdns-update * 06:23 marostegui@dns1004: START - running authdns-update * 06:22 marostegui@dns1004: END - running authdns-update * 06:20 marostegui@dns1004: START - running authdns-update * 06:17 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1026.eqiad.wmnet with reason: reboot * 06:04 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: sync * 06:04 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: sync * 06:03 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync * 06:03 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync * 06:02 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync * 06:01 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync * 06:01 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync * 06:00 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync * 05:59 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:59 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:40 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:39 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:26 marostegui@dns1004: END - running authdns-update * 05:24 marostegui@dns1004: START - running authdns-update * 05:24 marostegui@dns1004: START - running authdns-update * 05:13 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1004.wikimedia.org * 05:07 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1004.wikimedia.org * 04:01 mwpresync@deploy2003: Pruned MediaWiki: 1.47.0-wmf.8 (duration: 01m 07s) * 03:39 mwpresync@deploy2003: Finished scap sync-world: testwikis to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] (duration: 36m 01s) * 03:03 mwpresync@deploy2003: Started scap sync-world: testwikis to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 29s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-13 == * 23:33 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1064.eqiad.wmnet * 23:33 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1064.eqiad.wmnet * 23:08 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1064.eqiad.wmnet with reason: vacuum overlarge container dbs * 23:06 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1069.eqiad.wmnet * 23:06 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1069.eqiad.wmnet * 22:34 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1069.eqiad.wmnet with reason: vacuum overlarge container dbs * 21:18 maryum: Deployed security fix for [[phab:T321092|T321092]] * 20:28 swfrench-wmf: reprepro include etcd-mirror_0.0.12-1+deb13u1 into main for trixie-wikimedia - [[phab:T424266|T424266]] * 20:26 swfrench-wmf: reprepro include etcd-mirror_0.0.12-1+deb12u1 into main for bookworm-wikimedia - [[phab:T428495|T428495]] * 20:23 dancy@deploy2003: Finished scap sync-world: Testing [[phab:T431635|T431635]] (duration: 03m 36s) * 20:19 dancy@deploy2003: Started scap sync-world: Testing [[phab:T431635|T431635]] * 20:18 dancy@deploy2003: Installation of scap version "4.274.0" completed for 3 hosts * 20:16 dancy@deploy2003: Installing scap version "4.274.0" for 3 host(s) * 20:12 kemayo@deploy2003: Finished scap sync-world: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] (duration: 08m 25s) * 20:07 kemayo@deploy2003: soda, esanders, kemayo: Continuing with deployment * 20:05 kemayo@deploy2003: soda, esanders, kemayo: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there * 20:04 kemayo@deploy2003: Started scap sync-world: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] * 18:22 cwhite: lvextend vg0/srv +500g on centrallog hosts * 18:19 cdobbins@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS trixie * 17:46 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1071.eqiad.wmnet * 17:46 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1071.eqiad.wmnet * 17:13 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1071.eqiad.wmnet with reason: vacuum overlarge container dbs * 17:07 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1065.eqiad.wmnet * 17:07 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1065.eqiad.wmnet * 17:06 dzahn@dns1006: END - running authdns-update * 17:04 dzahn@dns1006: START - running authdns-update * 17:01 dzahn@dns1006: END - running authdns-update * 16:59 dzahn@dns1006: START - running authdns-update * 16:51 dancy@deploy2003: Finished scap sync-world: testing [[phab:T428971|T428971]] (duration: 03m 37s) * 16:47 dancy@deploy2003: Started scap sync-world: testing [[phab:T428971|T428971]] * 16:45 atsukoito: restarting pybal on lvs1019 to flush IP address for `cirrussearch1122.eqiad.wmnet` after moving the vlan [[phab:T431311|T431311]] * 16:42 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:42 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:42 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:42 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:42 Amir1: mwscript-k8s --follow --dblist=ores -- extensions/ORES/maintenance/PurgeScoreCache.php --model damaging --old ([[phab:T431159|T431159]]) * 16:34 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Pool test * 16:34 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 16:34 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 16:34 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Pool test * 16:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Depool test * 16:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 16:33 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 16:33 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Depool test * 16:31 dancy@deploy2003: Installation of scap version "4.273.0" completed for 159 hosts * 16:29 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1065.eqiad.wmnet with reason: vacuum overlarge container dbs * 16:27 dancy@deploy2003: Installing scap version "4.273.0" for 159 host(s) * 16:27 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics-external: sync * 16:27 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics-external: sync * 16:26 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics-external: sync * 16:26 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics-external: sync * 16:22 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync * 16:21 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync * 16:21 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: sync * 16:21 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: sync * 16:19 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync * 16:19 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync * 16:18 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync * 16:17 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync * 15:59 atsukoito: restarting pybal on lvs1018 for https://gerrit.wikimedia.org/r/1310117 * 15:55 aikochou@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 15:50 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310117 * 15:46 aikochou@deploy2003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 15:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host kafka-logging1006.eqiad.wmnet * 15:43 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host kafka-logging1006.eqiad.wmnet * 15:41 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host ganeti-test[2001-2003].codfw.wmnet * 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host ganeti-test[2001-2003].codfw.wmnet * 15:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host netbox1003.eqiad.wmnet * 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host netbox1003.eqiad.wmnet * 15:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host netbox2003.codfw.wmnet * 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host netbox2003.codfw.wmnet * 15:36 sukhe: restart pybal on lvs1020 * 15:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet * 15:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet * 15:08 btullis@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:06 btullis@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 15:01 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:01 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:35 cdobbins@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 14:34 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:33 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:33 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:32 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:29 cdobbins@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 14:28 marostegui@dns1004: END - running authdns-update * 14:28 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:27 marostegui@dns1004: START - running authdns-update * 14:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1023.eqiad.wmnet with reason: reboot * 14:18 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2009.codfw.wmnet with OS trixie * 14:14 swfrench-wmf: start rolling run-puppet-agent on A:cp for ATS config change - [[phab:T428909|T428909]] [[phab:T431838|T431838]] * 14:05 swfrench-wmf: disable-puppet on A:cp for ATS config change - [[phab:T428909|T428909]] [[phab:T431838|T431838]] * 14:05 cdobbins@cumin2002: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie * 14:02 marostegui@dns1004: END - running authdns-update * 14:00 marostegui@dns1004: START - running authdns-update * 14:00 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1070.eqiad.wmnet * 14:00 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1070.eqiad.wmnet * 13:58 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2009.codfw.wmnet with reason: host reimage * 13:57 cdobbins@cumin2002: conftool action : set/pooled=no; selector: name=dns7002.* * 13:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2009.codfw.wmnet with reason: host reimage * 13:48 rscout@deploy2003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply * 13:48 rscout@deploy2003: helmfile [eqiad] START helmfile.d/services/miscweb: apply * 13:48 rscout@deploy2003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply * 13:47 rscout@deploy2003: helmfile [codfw] START helmfile.d/services/miscweb: apply * 13:40 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:33 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2009.codfw.wmnet with OS trixie * 13:30 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1070.eqiad.wmnet with reason: vacuum overlarge container dbs * 13:28 aude@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] (duration: 11m 12s) * 13:23 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:22 aude@deploy2003: aikochou, javiermonton, aude, gkm563: Continuing with deployment * 13:22 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:19 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:19 aude@deploy2003: aikochou, javiermonton, aude, gkm563: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] synced to the testservers * 13:17 aude@deploy2003: Started scap sync-world: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] * 13:01 ladsgroup@deploy2003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 13:01 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:00 ladsgroup@deploy2003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 12:59 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:52 ladsgroup@deploy2003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 12:51 ladsgroup@deploy2003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 12:48 atsuko@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 12:48 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 12:47 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2008.codfw.wmnet with OS trixie * 12:47 atsuko@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 12:47 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply * 12:47 atsuko@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:46 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 12:45 atsuko@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:45 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply * 12:45 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:44 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:43 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] (duration: 07m 02s) * 12:38 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 12:37 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:36 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] * 12:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2008.codfw.wmnet with reason: host reimage * 12:23 Msz2001: Deployed changes to private code for Suggested Investigations * 12:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2008.codfw.wmnet with reason: host reimage * 12:20 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:19 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:17 atsuko@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 12:17 atsuko@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 12:16 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:15 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] (duration: 07m 14s) * 12:10 mszwarc@deploy2003: mszwarc: Continuing with deployment * 12:09 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:07 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] * 12:04 mszwarc@deploy2003: sync-world aborted: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] (duration: 00m 29s) * 12:03 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] * 12:01 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2008.codfw.wmnet with OS trixie * 12:00 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] (duration: 07m 37s) * 11:55 zabe@deploy2003: zabe: Continuing with deployment * 11:54 zabe@deploy2003: zabe: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:52 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] * 11:51 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:43 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:35 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:34 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:33 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:30 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:28 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:27 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:17 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2007.codfw.wmnet with OS trixie * 11:09 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] * 11:06 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=s8 * 11:00 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=x3 * 11:00 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=s5 * 10:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2007.codfw.wmnet with reason: host reimage * 10:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2007.codfw.wmnet with reason: host reimage * 10:51 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host dse-k8s-worker1023 * 10:50 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host dse-k8s-worker1023 * 10:44 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dse-k8s-worker1023 * 10:43 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host dse-k8s-worker1023 * 10:42 marostegui@cumin1003: dbctl commit (dc=all): 'Change x4 masters [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P94804 and previous config saved to /var/cache/conftool/dbconfig/20260713-104248-marostegui.json * 10:37 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:37 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:35 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:35 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:34 atsuko@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 10:34 atsuko@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 10:33 marostegui@cumin1003: dbctl commit (dc=all): 'Push x4 initial dbctl config [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P94803 and previous config saved to /var/cache/conftool/dbconfig/20260713-103259-marostegui.json * 10:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2007.codfw.wmnet with OS trixie * 09:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2006.codfw.wmnet with OS trixie * 09:42 marostegui@dns1004: END - running authdns-update * 09:40 marostegui@dns1004: START - running authdns-update * 09:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2006.codfw.wmnet with reason: host reimage * 09:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2006.codfw.wmnet with reason: host reimage * 09:06 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1024.eqiad.wmnet with reason: reboot * 09:06 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:01 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2006.codfw.wmnet with OS trixie * 08:44 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 08:43 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 08:43 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:42 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:42 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:42 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:41 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 08:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup2004.codfw.wmnet * 08:38 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:33 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host db1208.eqiad.wmnet * 08:30 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=x3 * 08:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2005.codfw.wmnet with OS trixie * 08:28 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup2004.codfw.wmnet * 08:28 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup2003.codfw.wmnet * 08:24 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1039: Repooling after testing * 08:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on clouddb1016.eqiad.wmnet with reason: cloning * 08:23 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s5 * 08:23 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s8 * 08:21 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 08:21 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 08:17 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup2003.codfw.wmnet * 08:17 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1004.eqiad.wmnet * 08:14 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1208.eqiad.wmnet * 08:11 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host phab1005.eqiad.wmnet * 08:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2005.codfw.wmnet with reason: host reimage * 08:07 marostegui@dns1004: END - running authdns-update * 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1004.eqiad.wmnet * 08:07 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1003.eqiad.wmnet * 08:07 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:05 marostegui@dns1004: START - running authdns-update * 08:05 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2005.codfw.wmnet with reason: host reimage * 08:05 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 08:05 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host phab1005.eqiad.wmnet * 08:05 marostegui@dns1004: START - running authdns-update * 08:05 marostegui@dns1004: START - running authdns-update * 08:05 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 08:04 marostegui@dns1004: START - running authdns-update * 08:00 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit1003.wikimedia.org * 07:58 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1003.eqiad.wmnet * 07:58 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1002-dev.eqiad.wmnet * 07:58 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:58 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:54 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1002-dev.eqiad.wmnet * 07:54 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1001-dev.eqiad.wmnet * 07:54 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit1003.wikimedia.org * 07:53 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 07:53 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:52 Msz2001: UTC morning backport+config window done * 07:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2005.codfw.wmnet with OS trixie * {{safesubst:SAL entry|1=07:50 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark (T429943}} * 07:49 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1001-dev.eqiad.wmnet * 07:46 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:46 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:45 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 07:45 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:44 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 07:44 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:43 mszwarc@deploy2003: mszwarc, danielyepezgarces, anzx: Continuing with deployment * {{safesubst:SAL entry|1=07:39 mszwarc@deploy2003: mszwarc, danielyepezgarces, anzx: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark}} * 07:39 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1039: Repooling after testing * {{safesubst:SAL entry|1=07:36 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark (T429943)}} * 07:35 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] (duration: 30m 03s) * 07:25 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit2002.wikimedia.org * 07:22 mszwarc@deploy2003: mszwarc: Continuing with deployment * 07:21 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:19 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit2002.wikimedia.org * 07:15 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aphlict1002.eqiad.wmnet * 07:11 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host aphlict1002.eqiad.wmnet * 07:08 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2003.wikimedia.org * 07:05 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] * 07:02 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2003.wikimedia.org * 07:02 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2002.wikimedia.org * 06:55 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2002.wikimedia.org * 06:55 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1003.wikimedia.org * 06:49 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1003.wikimedia.org * 06:34 marostegui: Drop m5 ipoid database [[phab:T431007|T431007]] * 06:29 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1027.eqiad.wmnet with reason: reboot * 06:24 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1028.eqiad.wmnet with reason: reboot * 06:21 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1025.eqiad.wmnet with reason: reboot * 06:17 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1022.eqiad.wmnet with reason: reboot * 06:03 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on dbproxy[2005-2008].codfw.wmnet with reason: reboot * 05:37 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1217,1228].eqiad.wmnet with reason: cloning * 05:11 marostegui: Drop users_to_rename table [[phab:T431842|T431842]] * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-12 == * 16:01 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2209 [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94792 and previous config saved to /var/cache/conftool/dbconfig/20260712-160124-marostegui.json * 15:58 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2205 to s3 primary [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94791 and previous config saved to /var/cache/conftool/dbconfig/20260712-155853-marostegui.json * 15:58 marostegui: Starting s3 codfw emergency failover from db2209 to db2205 - [[phab:T431950|T431950]] * 15:51 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2205 with weight 0 [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94790 and previous config saved to /var/cache/conftool/dbconfig/20260712-155135-marostegui.json * 15:51 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Primary switchover s3 [[phab:T431950|T431950]] * 02:01 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 01m 17s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-11 == * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 26s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-10 == * 19:12 jhathaway@dns1004: END - running authdns-update * 19:10 jhathaway@dns1004: START - running authdns-update * 18:23 mutante: vrts2002 rebooting (not the active host) * 18:21 mutante: lists2001, phab2003 - rebooting (not the active hosts) * 18:16 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on A:lvs-high-traffic2-codfw * 18:15 swfrench@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on A:lvs-high-traffic2-codfw * 17:15 mutante: [doc1004:~] $ sudo systemctl start rsync-doc-host-data-sync ([[phab:T431856|T431856]]) * 17:09 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1004.eqiad.wmnet * 17:08 jhathaway@dns1004: END - running authdns-update * 17:07 jhathaway@dns1004: START - running authdns-update * 17:06 jhathaway: depooling puppetserver1002, cause of errors is still unknown * 17:03 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1004.eqiad.wmnet * 16:57 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1003.eqiad.wmnet * 16:51 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1003.eqiad.wmnet * 16:48 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 16:48 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2004.codfw.wmnet * 16:42 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2004.codfw.wmnet * 16:41 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2003.codfw.wmnet * 16:35 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2003.codfw.wmnet * 16:33 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2002.codfw.wmnet * 16:27 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2002.codfw.wmnet * 16:25 mutante: gitlab-runners (production) rebooting cluster one by one * 16:17 mutante: etherpad1004/etherpad2002 - (etherpad.wikimedia.org) - rebooting * 16:13 mutante: doc1004/doc2003 (doc.wikimedia.org backends) - rebooting * 16:02 mutante: releases1003/releases2003 (releases.wikimedia.org backends) - rebooting for maintenance * 15:26 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2007-dev.codfw.wmnet * 15:19 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2007-dev.codfw.wmnet * 15:14 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host cloudcephosd2007-dev.codfw.wmnet * 15:14 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2007-dev.codfw.wmnet * 15:14 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host cloudcephosd2006-dev.codfw.wmnet * 15:07 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2006-dev.codfw.wmnet * 15:07 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2005-dev.codfw.wmnet * 14:59 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2005-dev.codfw.wmnet * 14:59 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2004-dev.codfw.wmnet * 14:53 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2004-dev.codfw.wmnet * 14:53 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2007-dev.codfw.wmnet * 14:51 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1054.eqiad.wmnet * 14:51 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1054.eqiad.wmnet * 14:51 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1054.eqiad.wmnet * 14:47 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2007-dev.codfw.wmnet * 14:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2006-dev.codfw.wmnet * 14:41 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2006-dev.codfw.wmnet * 14:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2005-dev.codfw.wmnet * 14:37 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2005-dev.codfw.wmnet * 14:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2005-dev.codfw.wmnet * 14:29 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2005-dev.codfw.wmnet * 14:29 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2006-dev.codfw.wmnet * 14:21 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2006-dev.codfw.wmnet * 14:21 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2010-dev.codfw.wmnet * 14:15 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2010-dev.codfw.wmnet * 14:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudgw2004-dev.codfw.wmnet * 14:10 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1054.eqiad.wmnet with OS trixie * 14:09 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudgw2004-dev.codfw.wmnet * 14:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudgw2003-dev.codfw.wmnet * 14:02 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudgw2003-dev.codfw.wmnet * 14:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2004-dev.codfw.wmnet * 13:53 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2004-dev.codfw.wmnet * 13:53 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2003-dev.codfw.wmnet * 13:48 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:44 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2003-dev.codfw.wmnet * 13:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2002-dev.codfw.wmnet * 13:42 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:41 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:41 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:37 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2002-dev.codfw.wmnet * 13:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudidp2001-dev.codfw.wmnet * 13:33 blake@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker1054.eqiad.wmnet with reason: host reimage * 13:33 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudidp2001-dev.codfw.wmnet * 13:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudnet2006-dev.codfw.wmnet * 13:26 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudnet2006-dev.codfw.wmnet * 13:26 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudnet2005-dev.codfw.wmnet * 13:23 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1054.eqiad.wmnet with reason: host reimage * 13:18 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudnet2005-dev.codfw.wmnet * 13:18 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudservices2005-dev.codfw.wmnet * 13:12 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudservices2005-dev.codfw.wmnet * 13:11 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudservices2004-dev.codfw.wmnet * 13:08 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudservices2004-dev.codfw.wmnet * 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudweb2002-dev.wikimedia.org * 13:05 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 13:05 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1054 * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1054 * 13:04 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1054 * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1054.eqiad.wmnet 49.32.64.10.in-addr.arpa 9.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:04 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1054.eqiad.wmnet 49.32.64.10.in-addr.arpa 9.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1054 - blake@cumin1003" * 13:04 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1054 - blake@cumin1003" * 13:01 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudweb2002-dev.wikimedia.org * 13:00 blake@cumin1003: START - Cookbook sre.dns.netbox * 12:59 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1054 * 12:57 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1054.eqiad.wmnet with OS trixie * 12:57 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1054.eqiad.wmnet * 12:56 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1054.eqiad.wmnet * 12:56 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1054.eqiad.wmnet * 12:47 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS trixie * 12:44 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:39 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 12:39 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 12:38 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:37 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:14 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:10 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:08 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:07 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:00 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:00 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:51 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:49 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:48 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:47 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:44 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:32 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2001.codfw.wmnet * 11:32 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1053.eqiad.wmnet * 11:32 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2001.codfw.wmnet * 11:32 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1053.eqiad.wmnet * 11:32 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1053.eqiad.wmnet * 11:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker2001.codfw.wmnet * 11:31 cgoubert@cumin1003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker2001.codfw.wmnet * 11:31 cgoubert@cumin1003: END (FAIL) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=1) rolling reimage on P<nowiki>{</nowiki>wikikube-worker2001*<nowiki>}</nowiki> and (A:wikikube-master-codfw or A:wikikube-worker-codfw) * 11:30 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:30 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:21 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 18 hosts with reason: reboot & upgrade * 11:20 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker2001.codfw.wmnet with OS trixie * 11:16 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:15 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:14 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:14 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 11 hosts * 11:14 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 11 hosts * 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:08 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 11:02 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1053.eqiad.wmnet with OS trixie * 11:01 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:58 cgoubert@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 10:57 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:38 cgoubert@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker2001.codfw.wmnet with OS trixie * 10:38 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2001.codfw.wmnet * 10:38 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2001.codfw.wmnet * 10:38 cgoubert@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on P<nowiki>{</nowiki>wikikube-worker2001*<nowiki>}</nowiki> and (A:wikikube-master-codfw or A:wikikube-worker-codfw) * 10:35 topranks: adjust IBGP outbound policy on lsw1-e2-codfw [[phab:T423430|T423430]] towards ssw1-e1-codfw * 10:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 cgoubert@cumin1003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:27 cgoubert@cumin1003: END (FAIL) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=1) rolling reimage on A:wikikube-worker-codfw * 10:27 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker2001.codfw.wmnet with OS bookworm * 10:25 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 10:24 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:24 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:15 cgoubert@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 10:11 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 10:11 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 10:08 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:08 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:07 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 10:06 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 10:00 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:55 cgoubert@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker2001.codfw.wmnet with OS bookworm * 09:55 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2005-2006,2011-2012].codfw.wmnet * 09:55 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2005-2006,2011-2012].codfw.wmnet * 09:51 cgoubert@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on A:wikikube-worker-codfw * 09:41 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1053.eqiad.wmnet with reason: host reimage * 09:37 topranks: apply new IBGP outbound policy on lsw1-e2-codfw [[phab:T423430|T423430]] * 09:36 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:36 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1053.eqiad.wmnet with reason: host reimage * 09:16 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1053 * 09:16 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1053 * 09:15 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1053 * 09:15 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1053.eqiad.wmnet 48.32.64.10.in-addr.arpa 8.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:15 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1053.eqiad.wmnet 48.32.64.10.in-addr.arpa 8.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:15 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:15 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1053 - blake@cumin1003" * 09:15 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1053 - blake@cumin1003" * 09:11 blake@cumin1003: START - Cookbook sre.dns.netbox * 09:11 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1053 * 09:08 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1053.eqiad.wmnet with OS trixie * 09:08 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1053.eqiad.wmnet * 09:08 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1053.eqiad.wmnet * 09:08 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1053.eqiad.wmnet * 09:04 brouberol@dns1004: END - running authdns-update * 09:03 brouberol@dns1004: START - running authdns-update * 08:41 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e] (thin): Regular analytics weekly train THIN [analytics/refinery@1abf22ea] (duration: 02m 11s) * 08:38 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e] (thin): Regular analytics weekly train THIN [analytics/refinery@1abf22ea] * 08:38 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e]: Regular analytics weekly train [analytics/refinery@1abf22ea] (duration: 05m 17s) * 08:38 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:34 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:33 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e]: Regular analytics weekly train [analytics/refinery@1abf22ea] * 08:32 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@1abf22ea] (duration: 02m 03s) * 08:30 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@1abf22ea] * 08:30 JavierMonton: Deploying Refinery at {{Gerrit|1abf22ea}} for changes 1308121/T427068 1306491/T430020 and {{Gerrit|1308190}} * 08:29 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:29 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 08:24 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:24 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 08:18 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:18 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 08:00 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db[2183-2184].codfw.wmnet * 08:00 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for db[2183-2184].codfw.wmnet * 07:52 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:52 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 07:49 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 11 hosts with reason: reboot & upgrade * 07:47 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:47 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 07:44 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:44 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 07:23 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 10 hosts * 07:23 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 10 hosts * 06:45 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 10 hosts with reason: reboot & upgrade * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 41s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-09 == * 23:33 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] (duration: 13m 26s) * 23:29 ladsgroup@deploy2003: ladsgroup, jdlrobson: Continuing with deployment * 23:22 ladsgroup@deploy2003: ladsgroup, jdlrobson: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:20 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] * 22:57 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1165.eqiad.wmnet * 22:56 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1165.eqiad.wmnet * 22:56 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1165.eqiad.wmnet * 22:45 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1165.eqiad.wmnet with OS trixie * 22:38 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 22:37 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 22:37 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 22:37 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:37 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 22:25 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1165.eqiad.wmnet with reason: host reimage * 22:17 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1165.eqiad.wmnet with reason: host reimage * 22:13 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 22:12 rzl@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 22:04 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 22:04 rzl@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1165 * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1165 * 22:02 jasmine@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1165 * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1165.eqiad.wmnet 115.48.64.10.in-addr.arpa 5.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:02 jasmine@cumin2002: START - Cookbook sre.dns.wipe-cache wikikube-worker1165.eqiad.wmnet 115.48.64.10.in-addr.arpa 5.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1165 - jasmine@cumin2002" * 22:02 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1165 - jasmine@cumin2002" * 22:02 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 21:57 jasmine@cumin2002: START - Cookbook sre.dns.netbox * 21:55 jasmine@cumin2002: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1165 * 21:54 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-worker1165.eqiad.wmnet with OS trixie * 21:54 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 21:54 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1165.eqiad.wmnet * 21:53 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 21:53 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1165.eqiad.wmnet * 21:53 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1165.eqiad.wmnet * 21:53 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 21:47 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 21:45 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 21:43 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 21:43 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 21:42 maryum: Deploy fix for [[phab:T431684|T431684]] * 21:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 21:27 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] (duration: 34m 14s) * 21:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs1002 * 21:23 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs1002 * 21:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS trixie * 21:22 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 21:20 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 22s) * 21:20 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 21:16 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 21:16 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 21:15 ladsgroup@deploy2003: ladsgroup: Continuing with deployment * 21:13 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2002.codfw.wmnet with OS bookworm * 21:11 ladsgroup@deploy2003: ladsgroup: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:08 ladsgroup@cumin1003: END (PASS) - Cookbook sre.wikireplicas.update-views (exit_code=0) * 21:07 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 6 hosts with reason: reboots * 20:54 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecycle work - bking@cumin2003 * 20:53 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:53 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] * 20:51 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) * 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 20:47 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecycle work - bking@cumin2003 * 20:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 20:41 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:41 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) * 20:40 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host relforge1008.eqiad.wmnet * 20:40 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1009.eqiad.wmnet with OS trixie * 20:33 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:32 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:32 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:31 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:31 ladsgroup@cumin1003: END (PASS) - Cookbook sre.wikireplicas.update-views (exit_code=0) * 20:29 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1008.eqiad.wmnet * 20:24 rzl@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 20:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2002.codfw.wmnet with OS bookworm * 20:23 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host relforge1008.eqiad.wmnet * 20:23 rzl@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 20:23 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1008.eqiad.wmnet * 20:22 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:22 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:21 bking@cumin2003: END (ERROR) - Cookbook sre.elasticsearch.rolling-operation (exit_code=97) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:21 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1009.eqiad.wmnet with reason: host reimage * 20:16 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:15 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1009.eqiad.wmnet with reason: host reimage * 20:12 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) * 20:02 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 19:55 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1009.eqiad.wmnet with OS trixie * 19:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 19:43 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 19:30 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 19:28 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 19:27 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 19:25 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 18:42 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 18:41 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 18:16 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 18:15 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 17:45 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for doh5004.wikimedia.org * 17:45 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for doh5004.wikimedia.org * 17:38 ladsgroup@deploy2003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 17:35 ladsgroup@deploy2003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 17:29 ladsgroup@deploy2003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 17:26 ladsgroup@deploy2003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 17:09 mutante: zuul[12]00[123] - rebooting for maintenance * 17:09 ebernhardson: start full in-place reindex of eqiad cirrussearch cluster * 17:08 dzahn@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-cluster (exit_code=99) * 17:08 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-cluster * 17:03 ebernhardson: start full in-place reindex of codfw cirrussearch cluster * 16:59 mutante: stewards1001/stewards2001 - reboot for maintenance * 16:54 ebernhardson: start full in-place reindex of cloudelastic cluster * 16:53 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 16:52 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply * 16:49 mutante: planet1003/planet2003 - rebooting * 16:47 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on doh5004.wikimedia.org with reason: random high load, investigating * 15:55 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 15:54 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 15:51 jynus: restarting backupmon1001 * 15:49 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 14 hosts * 15:49 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 14 hosts * 15:47 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backupmon1001.eqiad.wmnet with reason: restart * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:06 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 15:06 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 14:59 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply * 14:58 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply * 14:51 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 14 hosts * 14:51 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 14 hosts * 14:49 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 6 hosts with reason: reboot & upgrade * 14:48 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet * 14:48 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet * 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:42 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1052.eqiad.wmnet * 14:42 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1052.eqiad.wmnet * 14:42 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1052.eqiad.wmnet * 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:31 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:31 elukey@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: sync * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:30 elukey@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: sync * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:28 elukey@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: sync * 14:28 elukey@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: sync * 14:26 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:20 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1052.eqiad.wmnet with OS trixie * 14:19 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 6 hosts with reason: reboot & upgrade * 14:18 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:15 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:15 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:13 elukey: update druid indexation job for webrequest_sampled_live - [[phab:T427068|T427068]] * 14:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:09 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:09 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for papaul - jhancock@cumin2002" * 14:09 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for papaul - jhancock@cumin2002" * 14:07 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:07 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:04 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 14:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cuminunpriv1001.eqiad.wmnet * 13:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb1003.eqiad.wmnet * 13:59 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1052.eqiad.wmnet with reason: host reimage * 13:57 moritzm: installing requests security updates * 13:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cuminunpriv1001.eqiad.wmnet * 13:55 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb1003.eqiad.wmnet * 13:53 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1052.eqiad.wmnet with reason: host reimage * 13:50 moritzm: installing python-cryptography security updates * 13:47 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb2003.codfw.wmnet * 13:44 Msz2001: UTC afternoon config+backport window is done * 13:44 Msz2001: Updated `logging` on `metawiki` to fix log performers, [[phab:T431176|T431176]]#12105297 * 13:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb2003.codfw.wmnet * 13:43 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 13:43 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt1002.wikimedia.org * 13:41 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] (duration: 07m 30s) * 13:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt1002.wikimedia.org * 13:37 mszwarc@deploy2003: mszwarc: Continuing with deployment * 13:36 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1052 * 13:36 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1052 * 13:35 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:35 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1052 * 13:35 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1052.eqiad.wmnet 47.32.64.10.in-addr.arpa 7.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:35 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1052.eqiad.wmnet 47.32.64.10.in-addr.arpa 7.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:35 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:35 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1052 - blake@cumin1003" * 13:35 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1052 - blake@cumin1003" * 13:34 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] * 13:31 blake@cumin1003: START - Cookbook sre.dns.netbox * 13:31 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1052 * 13:30 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1052.eqiad.wmnet with OS trixie * 13:30 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1052.eqiad.wmnet * 13:29 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1052.eqiad.wmnet * 13:29 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1052.eqiad.wmnet * 13:17 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] (duration: 11m 26s) * 13:13 jforrester@deploy2003: jforrester: Continuing with deployment * 13:08 jforrester@deploy2003: jforrester: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:06 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] * 12:54 cgoubert@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply * 12:54 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:52 cgoubert@deploy2003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply * 12:45 cgoubert@deploy2003: helmfile [codfw] DONE helmfile.d/services/mobileapps: apply * 12:44 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:44 cgoubert@deploy2003: helmfile [codfw] START helmfile.d/services/mobileapps: apply * 12:43 cgoubert@deploy2003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 12:43 cgoubert@deploy2003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 12:42 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast4006.wikimedia.org * 12:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt2002.wikimedia.org * 12:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast7002.wikimedia.org * 12:18 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast4006.wikimedia.org * 12:18 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host ml-serve1004 * 12:18 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host ml-serve1004 * 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt2002.wikimedia.org * 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast7002.wikimedia.org * 12:10 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backup[2003,2014].codfw.wmnet with reason: reboot & upgrade * 12:10 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt-staging2001.codfw.wmnet * 12:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid1003.eqiad.wmnet * 12:06 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt-staging2001.codfw.wmnet * 12:05 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid1003.eqiad.wmnet * 12:03 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backup[1003,1014].eqiad.wmnet with reason: reboot & upgrade * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid2003.codfw.wmnet * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host irc1003.wikimedia.org * 11:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid2003.codfw.wmnet * 11:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host irc1003.wikimedia.org * 11:55 jmm@dns1004: END - running authdns-update * 11:53 jmm@dns1004: START - running authdns-update * 11:50 jmm@dns1004: END - running authdns-update * 11:48 jmm@dns1004: START - running authdns-update * 11:27 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host irc2003.wikimedia.org * 11:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host irc2003.wikimedia.org * 11:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint2001.codfw.wmnet * 11:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint1001.eqiad.wmnet * 11:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint2001.codfw.wmnet * 11:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint1001.eqiad.wmnet * 11:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-rw2001.wikimedia.org * 11:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-rw1001.wikimedia.org * 11:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-rw2001.wikimedia.org * 11:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-rw1001.wikimedia.org * 11:03 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon1003.wikimedia.org * 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2005.codfw.wmnet * 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2005.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 10:59 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2005.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 10:57 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon1003.wikimedia.org * 10:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon2002.wikimedia.org * 10:55 jmm@cumin2003: START - Cookbook sre.dns.netbox * 10:51 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon2002.wikimedia.org * 10:51 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:50 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2005.codfw.wmnet * 10:41 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:40 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host ml-serve1003 * 10:40 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host ml-serve1003 * 10:39 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2033.codfw.wmnet * 10:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install2005.wikimedia.org * 10:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install1005.wikimedia.org * 10:35 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1004.eqiad.wmnet with OS bookworm * 10:31 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install1005.wikimedia.org * 10:31 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install2005.wikimedia.org * 10:30 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install4004.wikimedia.org * 10:30 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install3004.wikimedia.org * 10:29 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install3004.wikimedia.org * 10:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install4004.wikimedia.org * 10:23 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 10:21 moritzm: failover Ganeti master in codfw/routed to ganeti2034 [[phab:T430928|T430928]] * 10:19 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.addnode (exit_code=0) for new host ganeti2031.codfw.wmnet to cluster codfw and group B * 10:19 moritzm: readded ganeti2031 to the codfw Ganeti cluster [[phab:T430910|T430910]] * 10:18 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1004.eqiad.wmnet with reason: host reimage * 10:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install5004.wikimedia.org * 10:18 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1003 * 10:18 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1003 * 10:17 jmm@cumin2003: START - Cookbook sre.ganeti.addnode for new host ganeti2031.codfw.wmnet to cluster codfw and group B * 10:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install6003.wikimedia.org * 10:16 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install5004.wikimedia.org * 10:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install6003.wikimedia.org * 10:15 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1004.eqiad.wmnet with reason: host reimage * 10:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1001.eqiad.wmnet * 10:14 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 10:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2008.wikimedia.org * 10:00 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ml-serve1004 * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1004 * 09:57 jmm@cumin2003: START - Cookbook sre.dns.netbox * 09:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install7002.wikimedia.org * 09:57 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1004 * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ml-serve1004.eqiad.wmnet 50.48.64.10.in-addr.arpa 0.5.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:57 klausman@cumin1003: START - Cookbook sre.dns.wipe-cache ml-serve1004.eqiad.wmnet 50.48.64.10.in-addr.arpa 0.5.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1004 - klausman@cumin1003" * 09:56 klausman@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1004 - klausman@cumin1003" * 09:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-coord1001.eqiad.wmnet * 09:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 09:55 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow7002.magru.wmnet * 09:52 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-coord1001.eqiad.wmnet * 09:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 09:52 klausman@cumin1003: START - Cookbook sre.dns.netbox * 09:50 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install7002.wikimedia.org * 09:50 klausman@cumin1003: START - Cookbook sre.hosts.move-vlan for host ml-serve1004 * 09:50 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1004.eqiad.wmnet with OS bookworm * 09:50 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1003.eqiad.wmnet with OS bookworm * 09:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1001.eqiad.wmnet * 09:49 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 09:49 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2008.wikimedia.org * 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2007.codfw.wmnet * 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2007.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 09:49 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow7002.magru.wmnet * 09:49 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2007.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 09:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard1003.eqiad.wmnet * 09:39 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard2003.codfw.wmnet * 09:39 jmm@cumin2003: START - Cookbook sre.dns.netbox * 09:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard1003.eqiad.wmnet * 09:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor1003.eqiad.wmnet * 09:35 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard2003.codfw.wmnet * 09:34 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2007.codfw.wmnet * 09:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor1003.eqiad.wmnet * 09:33 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor-dev2001.codfw.wmnet * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor2003.codfw.wmnet * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sretest1006.eqiad.wmnet * 09:27 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 09:25 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor-dev2001.codfw.wmnet * 09:25 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor2003.codfw.wmnet * 09:23 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] (duration: 06m 27s) * 09:23 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2205: codfw rack B4 repool after maintenance * 09:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host sretest1006.eqiad.wmnet * 09:23 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2204: codfw rack B4 repool after maintenance * 09:19 urbanecm@deploy2003: urbanecm: Continuing with deployment * 09:19 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:18 jmm@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 6 hosts with reason: reboot * 09:17 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] * 09:08 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ml-serve1003 * 09:08 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1003 * 09:07 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1003 * 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ml-serve1003.eqiad.wmnet 81.32.64.10.in-addr.arpa 1.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:07 klausman@cumin1003: START - Cookbook sre.dns.wipe-cache ml-serve1003.eqiad.wmnet 81.32.64.10.in-addr.arpa 1.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1003 - klausman@cumin1003" * 09:06 klausman@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1003 - klausman@cumin1003" * 08:58 klausman@cumin1003: START - Cookbook sre.dns.netbox * 08:57 klausman@cumin1003: START - Cookbook sre.hosts.move-vlan for host ml-serve1003 * 08:57 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1003.eqiad.wmnet with OS bookworm * 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=0) rolling reimage on P<nowiki>{</nowiki>ml-serve1003.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet * 08:55 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet * 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1003.eqiad.wmnet with OS bookworm * 08:39 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 08:38 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool db2205: codfw rack B4 repool after maintenance * 08:37 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool db2204: codfw rack B4 repool after maintenance * 08:36 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 08:35 hashar@deploy2003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 08:32 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:32 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:31 hashar@deploy2003: Rolling back deployment * 08:26 moritzm: failover Ganeti master in codfw to ganeti2048 [[phab:T430928|T430928]] * 08:16 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1003.eqiad.wmnet with OS bookworm * 08:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2004.codfw.wmnet * 08:16 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet * 08:16 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet * 08:16 klausman@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on P<nowiki>{</nowiki>ml-serve1003.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 08:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2002.codfw.wmnet * 08:15 XioNoX: lsw1-b4-codfw> request system reboot - [[phab:T430910|T430910]] * 08:15 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b4-codfw,lsw1-b4-codfw IPv6,lsw1-b4-codfw.mgmt with reason: Switch maintenance * 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for codfw rack B4 * 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:10 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2004.codfw.wmnet * 08:10 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2002.codfw.wmnet * 08:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2205: codfw rack B4 depool for maintenance * 08:08 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool db2205: codfw rack B4 depool for maintenance * 08:08 jmm@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin2003.codfw.wmnet * 08:08 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2204: codfw rack B4 depool for maintenance * 08:08 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool db2204: codfw rack B4 depool for maintenance * 08:08 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 27 hosts with reason: codfw rack B4 depool for maintenance * 08:03 jmm@cumin2002: START - Cookbook sre.hosts.reboot-single for host cumin2003.codfw.wmnet * 07:56 ayounsi@cumin1003: START - Cookbook sre.network.depool-rack with action 'depool' for codfw rack B4 * 07:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1008.eqiad.wmnet with OS trixie * 07:49 wmde-fisch@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] (duration: 08m 36s) * 07:44 wmde-fisch@deploy2003: wmde-fisch: Continuing with deployment * 07:43 wmde-fisch@deploy2003: wmde-fisch: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:41 wmde-fisch@deploy2003: Started scap sync-world: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] * 07:35 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1008.eqiad.wmnet with reason: host reimage * 07:31 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1008.eqiad.wmnet with reason: host reimage * 07:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1008.eqiad.wmnet with OS trixie * 07:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 07:00 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 06:59 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1008.eqiad.wmnet with OS trixie * 06:57 Emperor: rebalance thanos swift rings after previous re-image of thanos-fe1004 to trixie * 06:47 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1008.eqiad.wmnet with OS trixie * 04:10 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 14 days, 0:00:00 on cp6008.drmrs.wmnet with reason: Hardware failure - [[phab:T431651|T431651]] * 03:55 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp6008.* * 03:29 ryankemper: [[phab:T431311|T431311]] Repooled eqiad cirrussearch clusters (`chi/omega/psi`) following completion of OpenSearch 2.19 migration * 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad * 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=eqiad * 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 31s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-08 == * 23:52 Amir1: ladsgroup@deploy2003:~$ mwscript-k8s --follow -- extensions/ORES/maintenance/PurgeScoreCache.php --wiki=simplewiki --model damaging --old ([[phab:T431159|T431159]]) * 23:46 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 23:46 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing PTR for 2001:df2:e500:fe08::1 - cmooney@cumin1003" * 23:46 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing PTR for 2001:df2:e500:fe08::1 - cmooney@cumin1003" * 23:40 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 23:16 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 23:15 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 22:42 rzl@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 22:40 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] (duration: 12m 55s) * 22:40 rzl@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 22:37 rzl@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 22:36 rzl@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 22:35 rzl@deploy2003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 22:34 urbanecm@deploy2003: urbanecm: Continuing with deployment * 22:33 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:33 rzl@deploy2003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 22:32 rzl@deploy2003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 22:30 rzl@deploy2003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 22:30 rzl@deploy2003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 22:29 rzl@deploy2003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 22:27 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] * 22:26 rzl@deploy2003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 22:22 rzl@deploy2003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 22:21 rzl@deploy2003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 22:19 rzl@deploy2003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 22:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 22:17 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 22:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 22:13 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 22:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 22:13 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 22:09 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 22:06 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 22:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1094.eqiad.wmnet with OS trixie * 22:01 urbanecm: Make https://test.wikipedia.org/w/index.php?title=MediaWiki:GrowthExperimentsSuggestedEdits.json&diff=prev&oldid=750552 with GrowthExperiments disabled (via mw-experimental), then run `\MediaWiki\MediaWikiServices::getInstance()->get('CommunityConfiguration.ProviderFactory')->newProvider('GrowthSuggestedEdits')->getStore()->invalidate()` ([[phab:T431625|T431625]]) * 21:56 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d2-codfw * 21:55 urbanecm@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 21:55 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d2-codfw * 21:55 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c4-codfw * 21:55 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c4-codfw * 21:55 urbanecm@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2002 * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2002 * 21:54 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2002 * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2002.codfw.wmnet 50.32.192.10.in-addr.arpa 0.5.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:54 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2002.codfw.wmnet 50.32.192.10.in-addr.arpa 0.5.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2002 - bking@cumin2003" * 21:54 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2002 - bking@cumin2003" * 21:49 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:49 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2002 * 21:49 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2002.codfw.wmnet with OS trixie * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1094.eqiad.wmnet with reason: host reimage * 21:42 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 21:39 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 21:37 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1094.eqiad.wmnet with reason: host reimage * 21:36 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 21:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 21:29 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 21:27 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 21:22 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1094.eqiad.wmnet with OS trixie * 21:21 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host restbase2039.codfw.wmnet with OS bullseye * 21:21 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin2002" * 21:21 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin2002" * 21:04 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on restbase2039.codfw.wmnet with reason: host reimage * 21:00 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on restbase2039.codfw.wmnet with reason: host reimage * 20:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1073.eqiad.wmnet with OS trixie * 20:48 mutante: deploy2003 - kill 1102 (stunnel4) ; systemctl start stunnel4 ([[phab:T418262|T418262]]) * 20:42 cjming@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] (duration: 33m 02s) * 20:42 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host restbase2039.codfw.wmnet with OS bullseye * 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1073.eqiad.wmnet with reason: host reimage * 20:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1073.eqiad.wmnet with reason: host reimage * 20:30 cjming@deploy2003: cjming: Continuing with deployment * 20:28 cjming@deploy2003: cjming: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1098.eqiad.wmnet with OS trixie * 20:13 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1073.eqiad.wmnet with OS trixie * 20:09 cjming@deploy2003: Started scap sync-world: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] * 20:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1098.eqiad.wmnet with reason: host reimage * 19:56 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1098.eqiad.wmnet with reason: host reimage * 19:55 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d4-codfw * 19:54 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d4-codfw * 19:54 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c1-codfw * 19:54 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c1-codfw * 19:52 mutante: restarting gerrit on gerrit.wikimedia.org (gerrit2003) * 19:48 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2331.codfw.wmnet * 19:48 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2331.codfw.wmnet * 19:48 mutante: restarting gerrit on gerrit-replica.wikimedia.org (gerrit1003) * 19:46 mutante: restarting gerrit on gerrit-spare.wikimedia.org (gerrit2002) * 19:43 jasmine@cumin2002: conftool action : set/pooled=yes; selector: name=wikikube-worker2331.codfw.wmnet,cluster=kubernetes,service=kubesvc * 19:43 jasmine@cumin2002: conftool action : set/weight=10; selector: name=wikikube-worker2331.codfw.wmnet,cluster=kubernetes,service=kubesvc * 19:40 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1098.eqiad.wmnet with OS trixie * 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d5-codfw * 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d5-codfw * 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c7-codfw * 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c7-codfw * 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c5-codfw * 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c5-codfw * 19:30 jasmine_: ran homer on lsw1-d8-codfw, adding wikikube-worker2331 to cluster * 19:29 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1100.eqiad.wmnet with OS trixie * 19:20 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d8-codfw * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d8-codfw * 19:19 mutante: gerrit - replacing private key for registerEmail verification * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-magru * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device cr2-magru * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d7-codfw * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d7-codfw * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d3-codfw * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d1-codfw * 19:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d1-codfw * 19:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c2-codfw * 19:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c2-codfw * 19:11 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-codfw * 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-magru * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device cr1-magru * 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d8-codfw * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d8-codfw * 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d6-codfw * 19:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1100.eqiad.wmnet with reason: host reimage * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d6-codfw * 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c6-codfw * 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c6-codfw * 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c3-codfw * 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c3-codfw * 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b4-magru * 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device asw1-b4-magru * 19:08 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b3-magru * 19:08 cmooney@cumin1003: START - Cookbook sre.network.tls for network device asw1-b3-magru * 19:05 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1100.eqiad.wmnet with reason: host reimage * 19:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1122.eqiad.wmnet with OS trixie * 19:00 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 18:59 topranks: rolling out update to BGP ACL on Nokia Switches eqiad, codfw & ulsfo [[phab:T425703|T425703]] * 18:58 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 18:57 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 18:55 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 18:53 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 18:52 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 18:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1100.eqiad.wmnet with OS trixie * 18:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1068.eqiad.wmnet with OS trixie * 18:47 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1102.eqiad.wmnet with OS trixie * 18:47 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 18:46 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1122.eqiad.wmnet with reason: host reimage * 18:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1122.eqiad.wmnet with reason: host reimage * 18:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1068.eqiad.wmnet with reason: host reimage * 18:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1122 * 18:26 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1122 * 18:25 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1122 * 18:25 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1122.eqiad.wmnet 31.48.64.10.in-addr.arpa 1.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:25 bking@cumin2003: START - Cookbook sre.dns.wipe-cache cirrussearch1122.eqiad.wmnet 31.48.64.10.in-addr.arpa 1.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:25 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:25 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1122 - bking@cumin2003" * 18:25 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1122 - bking@cumin2003" * 18:21 rzl@deploy2003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 18:21 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1068.eqiad.wmnet with reason: host reimage * 18:21 rzl@deploy2003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 18:21 rzl@deploy2003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 18:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 18:19 rzl@deploy2003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 18:19 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:18 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1122 * 18:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1122.eqiad.wmnet with OS trixie * 18:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 18:15 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 18:13 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 18:13 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 18:10 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 18:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1068.eqiad.wmnet with OS trixie * 18:01 kamila@deploy2003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 18m 29s) * 18:00 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:55 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:42 kamila@deploy2003: Started scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] * 17:42 kamila@deploy2003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 19m 50s) * 17:42 kamila@deploy2003: Rolling back deployment * 17:35 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:31 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet * 17:18 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet * 17:16 kamila@deploy1003: Unlocked for deployment [MediaWiki]: switching deployment server (duration: 22m 07s) * 17:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 17:11 kamila@dns1005: END - running authdns-update * 17:09 kamila@dns1005: START - running authdns-update * 17:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 17:04 jasmine@cumin2002: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1164.eqiad.wmnet * 17:04 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1164.eqiad.wmnet * 17:04 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1164.eqiad.wmnet * 16:56 kamila@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on releases2003.codfw.wmnet,releases1003.eqiad.wmnet with reason: Deployment server switchover * 16:54 kamila@deploy1003: Locking from deployment [MediaWiki]: switching deployment server * 16:53 kamila@deploy1003: Unlocked for deployment [MediaWiki]: switching deployment server (duration: 04m 02s) * 16:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie * 16:49 kamila@deploy1003: Locking from deployment [MediaWiki]: switching deployment server * 16:46 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1095.eqiad.wmnet with OS trixie * 16:45 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1093.eqiad.wmnet with OS trixie * 16:43 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1164.eqiad.wmnet with OS trixie * 16:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1095.eqiad.wmnet with reason: host reimage * 16:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 16:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 16:23 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1164.eqiad.wmnet with reason: host reimage * 16:18 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on cirrussearch1093.eqiad.wmnet with reason: host reimage * 16:16 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1164.eqiad.wmnet with reason: host reimage * 16:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1095.eqiad.wmnet with reason: host reimage * 16:09 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 16:09 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 16:08 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1093.eqiad.wmnet with reason: host reimage * 15:59 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Pool test * 15:59 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:59 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 15:59 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Pool test * 15:58 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Depool test * 15:58 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:58 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 15:58 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Depool test * 15:57 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1164 * 15:57 jasmine@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1164 * 15:57 jasmine@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1164 * 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1164.eqiad.wmnet 114.48.64.10.in-addr.arpa 4.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:56 jasmine@cumin2002: START - Cookbook sre.dns.wipe-cache wikikube-worker1164.eqiad.wmnet 114.48.64.10.in-addr.arpa 4.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1164 - jasmine@cumin2002" * 15:56 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1164 - jasmine@cumin2002" * 15:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1093.eqiad.wmnet with OS trixie * 15:51 jasmine@cumin2002: START - Cookbook sre.dns.netbox * 15:51 jasmine@cumin2002: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1164 * 15:50 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-worker1164.eqiad.wmnet with OS trixie * 15:50 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1164.eqiad.wmnet * 15:50 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1164.eqiad.wmnet * 15:50 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1164.eqiad.wmnet * 15:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1095.eqiad.wmnet with OS trixie * 15:42 jasmine@cumin2002: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1164.eqiad.wmnet * 15:42 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1164.eqiad.wmnet * 15:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:42 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1164.eqiad.wmnet * 15:42 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1164.eqiad.wmnet * 15:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 15:39 elukey@cumin1003: START - Cookbook sre.hosts.provision for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 15:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1007.eqiad.wmnet with OS trixie * 15:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Pool test * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1007.eqiad.wmnet with reason: host reimage * 15:15 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 15:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet * 15:15 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 15:15 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1007.eqiad.wmnet with reason: host reimage * 15:15 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:14 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Pool test * 15:14 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet * 15:14 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2228: Depool test * 15:14 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db2228: Depool test * 15:10 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 15:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:08 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 15:06 blake@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 15:06 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet * 15:06 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 15:06 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 15:05 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet * 15:05 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 15:04 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:04 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 15:04 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:03 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 15:03 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 15:03 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:03 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T430909|T430909]] * 15:03 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:03 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 15:03 swfrench-wmf: restarted eqsin, codfw confds - [[phab:T430909|T430909]] * 15:03 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test1002.eqiad.wmnet * 15:02 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet * 14:59 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:59 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:55 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1007.eqiad.wmnet with OS trixie * 14:52 swfrench-wmf: restarted ulsfo confds, confirmed now connected to codfw backends except those using wikimedia.org SRV record - [[phab:T430909|T430909]] * 14:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:49 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:41 moritzm: uninstalling dhcpcd-base from trixie hosts which still have it installed [[phab:T414341|T414341]] * 14:40 sukhe: sudo cumin -b1 -s120 "P<nowiki>{</nowiki>lvs2011*<nowiki>}</nowiki> or P<nowiki>{</nowiki>lvs2012*<nowiki>}</nowiki>" "systemctl restart pybal.service" * 14:39 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:39 mvernon@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host thanos-be1007.eqiad.wmnet with OS trixie * 14:37 sukhe: restart pybal on lvs2013 to revert back to conf2004 * 14:35 sukhe: restart pybal on lvs2014 to revert back to conf2004 * 14:34 swfrench-wmf: switched codfw, eqsin, ulsfo etcd client SRV records back to codfw - [[phab:T430909|T430909]] * 14:32 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1002.eqiad.wmnet * 14:32 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet * 14:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1007.eqiad.wmnet with OS trixie * 14:31 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:31 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:31 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:30 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:30 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Pool test * 14:30 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:29 swfrench@dns1004: END - running authdns-update * 14:29 moritzm: installing jackson-core security updates * 14:27 swfrench@dns1004: START - running authdns-update * 14:22 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:22 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:22 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1119.eqiad.wmnet with OS trixie * 14:22 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:21 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:20 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:20 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 14:20 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:19 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 14:19 moritzm: installing librabbitmq security updates * 14:19 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1002.eqiad.wmnet * 14:18 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet * 14:16 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:15 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:15 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:14 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Pool test * 14:14 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox) * 14:14 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet * 14:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 14:08 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 14:07 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet * 14:05 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1118.eqiad.wmnet with OS trixie * 14:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test1001.eqiad.wmnet * 14:00 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-worker@eqiad * 14:00 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 13:59 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 13:58 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 13:57 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1119.eqiad.wmnet with reason: host reimage * 13:54 moritzm: installing libcap2 security updates * 13:53 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1119.eqiad.wmnet with reason: host reimage * 13:52 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet * 13:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) * 13:52 fceratto@cumin1003: START - Cookbook sre.mysql.depool * 13:50 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-worker@eqiad * 13:50 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1051.eqiad.wmnet * 13:50 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1051.eqiad.wmnet * 13:50 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1051.eqiad.wmnet * 13:49 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 13:45 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1081.eqiad.wmnet with OS trixie * 13:41 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1119 * 13:41 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1119 * 13:40 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1119 * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1119.eqiad.wmnet 97.32.64.10.in-addr.arpa 7.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1119.eqiad.wmnet 97.32.64.10.in-addr.arpa 7.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1119 - atsuko@cumin1003" * 13:40 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1119 - atsuko@cumin1003" * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1118.eqiad.wmnet with reason: host reimage * 13:39 moritzm: installing krb5 security updates * 13:37 Lucas_WMDE: UTC afternoon backport+config window done * 13:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1006.eqiad.wmnet with OS trixie * 13:36 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1118.eqiad.wmnet with reason: host reimage * 13:36 atsuko@cumin1003: START - Cookbook sre.dns.netbox * 13:35 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] (duration: 07m 46s) * 13:34 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1119 * 13:34 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1119.eqiad.wmnet with OS trixie * 13:30 sbisson@deploy1003: sbisson: Continuing with deployment * 13:30 moritzm: installing openssh security updates * 13:30 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling restart_daemons on A:wikidough * 13:29 sbisson@deploy1003: sbisson: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:27 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] * 13:26 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1051.eqiad.wmnet with OS trixie * 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1081.eqiad.wmnet with reason: host reimage * 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1118 * 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1118 * 13:22 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] (duration: 12m 12s) * 13:21 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1081.eqiad.wmnet with reason: host reimage * 13:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1006.eqiad.wmnet with reason: host reimage * 13:18 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1118 * 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1118.eqiad.wmnet 90.32.64.10.in-addr.arpa 0.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:18 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1118.eqiad.wmnet 90.32.64.10.in-addr.arpa 0.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1118 - atsuko@cumin1003" * 13:18 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1118 - atsuko@cumin1003" * 13:17 stran@deploy1003: stran: Continuing with deployment * 13:16 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough * 13:15 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:13 atsuko@cumin1003: START - Cookbook sre.dns.netbox * 13:12 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1118 * 13:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1006.eqiad.wmnet with reason: host reimage * 13:12 stran@deploy1003: stran: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:12 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1118.eqiad.wmnet with OS trixie * 13:10 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] * 13:05 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-worker@codfw * 13:05 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 13:05 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1081.eqiad.wmnet with OS trixie * 13:05 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1051.eqiad.wmnet with reason: host reimage * 13:04 moritzm: installing jq security updates * 13:04 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 13:01 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1051.eqiad.wmnet with reason: host reimage * 12:58 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-worker@codfw * 12:52 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 12:50 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host thanos-be1006.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1051 * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1051 * 12:43 moritzm: installing Python 3.11 security updates * 12:43 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1051 * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1051.eqiad.wmnet 46.32.64.10.in-addr.arpa 6.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:43 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1051.eqiad.wmnet 46.32.64.10.in-addr.arpa 6.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1051 - blake@cumin1003" * 12:43 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1051 - blake@cumin1003" * 12:38 blake@cumin1003: START - Cookbook sre.dns.netbox * 12:38 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1051 * 12:38 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1051.eqiad.wmnet with OS trixie * 12:37 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1051.eqiad.wmnet * 12:36 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1051.eqiad.wmnet * 12:36 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1051.eqiad.wmnet * 12:34 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1006.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 12:34 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1006.eqiad.wmnet with OS trixie * 12:27 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 12:27 mvernon@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host thanos-be1006.eqiad.wmnet with OS trixie * 12:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:02 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 12:01 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1006.eqiad.wmnet with OS trixie * 11:43 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 11:38 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1076.eqiad.wmnet with OS trixie * 11:26 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1075.eqiad.wmnet with OS trixie * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2047.codfw.wmnet * 11:19 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2047.codfw.wmnet * 11:19 moritzm: temporarily remove ganeti2031 from codfw cluster [[phab:T430910|T430910]] * 11:08 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1076.eqiad.wmnet with reason: host reimage * 11:08 moritzm: installing Linux 6.1.176 on Bookworm servers * 11:03 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1076.eqiad.wmnet with reason: host reimage * 11:00 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1075.eqiad.wmnet with reason: host reimage * 10:56 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1075.eqiad.wmnet with reason: host reimage * 10:47 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1076.eqiad.wmnet with OS trixie * 10:46 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1074.eqiad.wmnet with OS trixie * 10:45 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1005.eqiad.wmnet with OS trixie * 10:40 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1075.eqiad.wmnet with OS trixie * 10:32 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2031.codfw.wmnet * 10:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1005.eqiad.wmnet with reason: host reimage * 10:25 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1005.eqiad.wmnet with reason: host reimage * 10:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1074.eqiad.wmnet with reason: host reimage * 10:17 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1074.eqiad.wmnet with reason: host reimage * 10:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1005.eqiad.wmnet with OS trixie * 10:12 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet * 10:04 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 10:01 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet * 10:01 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1074.eqiad.wmnet with OS trixie * 10:01 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 09:43 cgoubert@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/aux-k8s-services/redioscope: apply * 09:43 cgoubert@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/aux-k8s-services/redioscope: apply * 09:43 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply * 09:35 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply * 09:34 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 09:34 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 09:33 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 41 days, 15:00:00 on db2252.codfw.wmnet with reason: Test * 09:32 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 09:32 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 09:31 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: codfw rack B3 pool after maintenance * 09:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 09:07 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 09:07 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 09:02 ladsgroup@cumin1003: END (PASS) - Cookbook sre.mysql.sanitarium_restart (exit_code=0) * 08:57 topranks: merge patch to shift eqiad <-> esams traffic onto new 40G circuit * 08:54 hashar@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1004.eqiad.wmnet with OS trixie * 08:50 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 08:50 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitarium_restart (exit_code=99) * 08:50 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 08:45 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool es2051: codfw rack B3 pool after maintenance * 08:44 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:44 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:43 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2007.codfw.wmnet * 08:43 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2007.codfw.wmnet * 08:42 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2031.codfw.wmnet * 08:41 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2031.codfw.wmnet * 08:40 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2031.codfw.wmnet * 08:38 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:38 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:35 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.sanitize-wiki (exit_code=97) Managing sanitization for wikis minwikiquote in section s3 * 08:33 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis minwikiquote in section s3 * 08:32 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Checking sanitization for wikis minwikiquote in section s5 * 08:30 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Checking sanitization for wikis minwikiquote in section s5 * 08:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Managing sanitization for wikis minwikiquote in section s5 * 08:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1004.eqiad.wmnet with reason: host reimage * 08:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1004.eqiad.wmnet with reason: host reimage * 08:23 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:22 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis minwikiquote in section s5 * 08:19 XioNoX: lsw1-b3-codfw> request system reboot - [[phab:T430909|T430909]] * 08:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Checking sanitization for wikis minwikiquote in section s5 * 08:17 hashar@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Checking sanitization for wikis minwikiquote in section s5 * 08:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for codfw rack B3 * 08:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2007.codfw.wmnet * 08:15 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lsw1-b3-codfw,lsw1-b3-codfw IPv6,lsw1-b3-codfw.mgmt with reason: Switch maintenance * 08:15 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2007.codfw.wmnet * 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:07 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:06 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: codfw rack B3 depool for maintenance * 08:05 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool es2051: codfw rack B3 depool for maintenance * 08:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1004.eqiad.wmnet with OS trixie * 08:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1005.eqiad.wmnet with OS trixie * 08:03 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 21 hosts with reason: codfw rack B3 depool for maintenance * 07:56 ayounsi@cumin1003: START - Cookbook sre.network.depool-rack with action 'depool' for codfw rack B3 * 07:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1005.eqiad.wmnet with reason: host reimage * 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1005.eqiad.wmnet with reason: host reimage * 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1005.eqiad.wmnet with OS bookworm * 07:29 moritzm: installing gnutls28 security updates * 07:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1005.eqiad.wmnet with OS trixie * 07:13 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1125.eqiad.wmnet with OS trixie * 07:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1005.eqiad.wmnet with reason: host reimage * 07:07 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aux-k8s-etcd1005.eqiad.wmnet with reason: host reimage * 06:56 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1005.eqiad.wmnet with OS bookworm * 06:54 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1125.eqiad.wmnet with reason: host reimage * 06:52 elukey: upgrade all trixie hosts to pywmflib 3.1 - [[phab:T430552|T430552]] * 06:50 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1125.eqiad.wmnet with reason: host reimage * 06:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 06:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 06:38 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1125.eqiad.wmnet with OS trixie * 05:42 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1107.eqiad.wmnet with OS trixie * 05:35 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1124.eqiad.wmnet with OS trixie * 05:31 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1101.eqiad.wmnet with OS trixie * 05:21 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1107.eqiad.wmnet with reason: host reimage * 05:17 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1124.eqiad.wmnet with reason: host reimage * 05:13 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1107.eqiad.wmnet with reason: host reimage * 05:13 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1101.eqiad.wmnet with reason: host reimage * 05:11 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1124.eqiad.wmnet with reason: host reimage * 05:10 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1101.eqiad.wmnet with reason: host reimage * 04:58 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1124.eqiad.wmnet with OS trixie * 04:56 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1107.eqiad.wmnet with OS trixie * 04:55 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1101.eqiad.wmnet with OS trixie * 02:27 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] (duration: 08m 14s) * 02:22 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 02:21 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 02:19 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] * 01:59 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] (duration: 09m 46s) * 01:55 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 01:51 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 01:49 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] * 01:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1099.eqiad.wmnet with OS trixie * 00:57 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1110.eqiad.wmnet with OS trixie * 00:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1099.eqiad.wmnet with reason: host reimage * 00:41 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1099.eqiad.wmnet with reason: host reimage * 00:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1110.eqiad.wmnet with reason: host reimage * 00:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1110.eqiad.wmnet with reason: host reimage * 00:26 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1099.eqiad.wmnet with OS trixie * 00:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1110.eqiad.wmnet with OS trixie == 2026-07-07 == * 22:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1097.eqiad.wmnet with OS trixie * 22:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1097.eqiad.wmnet with reason: host reimage * 22:24 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1097.eqiad.wmnet with reason: host reimage * 22:14 hashar: Restarting Gerrit on gerrit2002 and gerrit1003 (replicas) * 22:09 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1097.eqiad.wmnet with OS trixie * 22:07 hashar: Restarting Gerrit on gerrit2003 * 21:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 21:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 21:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 21:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 21:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1108.eqiad.wmnet with OS trixie * 20:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1091.eqiad.wmnet with OS trixie * 20:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1090.eqiad.wmnet with OS trixie * 20:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1108.eqiad.wmnet with reason: host reimage * 20:36 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1091.eqiad.wmnet with reason: host reimage * 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1090.eqiad.wmnet with reason: host reimage * 20:33 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1091.eqiad.wmnet with reason: host reimage * 20:30 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1108.eqiad.wmnet with reason: host reimage * 20:30 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1006.eqiad.wmnet * 20:30 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1090.eqiad.wmnet with reason: host reimage * 20:30 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1006.eqiad.wmnet * 20:27 jasmine_: "homer lsw1-c2-eqiad* commit "Added new stacked control plane wikikube-ctrl1006"" * 20:22 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] (duration: 07m 29s) * 20:20 jasmine_: "homer "cr*eqiad*" commit "Added new stacked control plane wikikube-ctrl1006"" * 20:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1091.eqiad.wmnet with OS trixie * 20:17 arlolra@deploy1003: arlolra: Continuing with deployment * 20:16 arlolra@deploy1003: arlolra: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:16 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1090.eqiad.wmnet with OS trixie * 20:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1108.eqiad.wmnet with OS trixie * 20:14 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] * 20:09 cwhite: remove 2026-04 swift log archives from centrallog2002 to free some space * 20:01 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=93) for host cirrussearch1108.eqiad.wmnet with OS trixie * 19:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1108.eqiad.wmnet with OS trixie * 19:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1090.eqiad.wmnet with OS trixie * 19:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1109.eqiad.wmnet with OS trixie * 19:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1092.eqiad.wmnet with OS trixie * 19:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1123.eqiad.wmnet with OS trixie * 19:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1109.eqiad.wmnet with reason: host reimage * 19:19 cdobbins@cumin2002: conftool action : set/pooled=yes; selector: name=dns7002.* * 19:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1092.eqiad.wmnet with reason: host reimage * 19:17 jasmine@dns1004: END - running authdns-update * 19:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1109.eqiad.wmnet with reason: host reimage * 19:15 jasmine@dns1004: START - running authdns-update * 19:15 cdobbins@dns1004: END - running authdns-update * 19:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1123.eqiad.wmnet with reason: host reimage * 19:13 cdobbins@dns1004: START - running authdns-update * 19:12 cdobbins@cumin2002: conftool action : set/pooled=yes; selector: name=dns7002.*,service=authdns-update * 19:11 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1092.eqiad.wmnet with reason: host reimage * 19:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1123.eqiad.wmnet with reason: host reimage * 18:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1123.eqiad.wmnet with OS trixie * 18:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1109.eqiad.wmnet with OS trixie * 18:56 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1092.eqiad.wmnet with OS trixie * 18:52 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 18:49 swfrench@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 18:40 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 18:38 swfrench@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 18:11 swfrench-wmf: restarted eqsin, codfw confds - [[phab:T430909|T430909]] * 18:01 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T430909|T430909]] * 17:59 swfrench-wmf: restarted ulsfo confds, confirmed now connected to eqiad backends - [[phab:T430909|T430909]] * 17:52 sukhe: restart pybal on lvs2011 to switch from conf2004 to conf1008: [[phab:T430909|T430909]] * 17:51 sukhe: restart pybal on lvs2012 to switch from conf2004 to conf1008 [puppet re-enabled there]: [[phab:T430909|T430909]] * 17:46 sukhe: restart pybal on lvs2013 to switch from conf2004 to conf1008: [[phab:T430909|T430909]] * 17:44 swfrench-wmf: switched codfw, eqsin, ulsfo etcd client SRV records to eqiad - [[phab:T430909|T430909]] * 17:43 swfrench@dns1004: END - running authdns-update * 17:40 swfrench@dns1004: START - running authdns-update * 17:40 sukhe: restart pybal on lvs2014 to switch from conf2004 to conf1008: [[phab:T430909|T430909]] * 17:21 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1003.eqiad.wmnet * 17:15 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1003.eqiad.wmnet * 17:14 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1002.eqiad.wmnet * 17:06 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1002.eqiad.wmnet * 17:06 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-low-traffic-codfw' 'systemctl restart pybal.service' # lvs2013, [[phab:T416623|T416623]] * 17:04 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1001.eqiad.wmnet * 17:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1111.eqiad.wmnet with OS trixie * 17:00 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal.service' # lvs2014, [[phab:T416623|T416623]] * 16:58 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1001.eqiad.wmnet * 16:58 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS bookworm * 16:55 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-low-traffic-eqiad' 'systemctl restart pybal.service' # lvs1019, [[phab:T416623|T416623]] * 16:53 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-secondary-eqiad' 'systemctl restart pybal.service' # lvs1020, [[phab:T416623|T416623]] * 16:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1111.eqiad.wmnet with reason: host reimage * 16:40 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1111.eqiad.wmnet with reason: host reimage * 16:38 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.peering (exit_code=99) with action 'configure' for AS: 47794 * 16:35 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 47794 * 16:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1111.eqiad.wmnet with OS trixie * 16:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1006.eqiad.wmnet with OS trixie * 16:06 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 16:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1006.eqiad.wmnet with reason: host reimage * 15:58 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1121.eqiad.wmnet with OS trixie * 15:58 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 15:56 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1006.eqiad.wmnet with reason: host reimage * 15:54 mutante: jenkins down in planned maintenance window * 15:42 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1037.eqiad.wmnet * 15:42 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1037.eqiad.wmnet * 15:42 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1037.eqiad.wmnet * 15:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1006.eqiad.wmnet with OS trixie * 15:34 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1121.eqiad.wmnet with reason: host reimage * 15:33 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS bookworm * 15:33 cdobbins@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host dns7002.wikimedia.org with OS trixie * 15:30 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1121.eqiad.wmnet with reason: host reimage * 15:29 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host clouddumps1001.wikimedia.org * 15:20 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1001.wikimedia.org * 15:18 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1121 * 15:18 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1121 * 15:18 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host clouddumps1002.wikimedia.org * 15:17 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1121 * 15:17 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:17 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply * 15:16 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply * 15:16 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:16 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1121 - atsuko@cumin1003" * 15:16 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1121 - atsuko@cumin1003" * 15:14 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1037.eqiad.wmnet with OS trixie * 15:11 atsuko@cumin1003: START - Cookbook sre.dns.netbox * 15:09 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org * 15:09 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1121 * 15:09 andrew@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host clouddumps1002.wikimedia.org * 15:09 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org * 15:09 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1121.eqiad.wmnet with OS trixie * 15:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1007.eqiad.wmnet with OS trixie * 15:08 andrew@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host clouddumps1002.wikimedia.org * 15:08 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org * 15:05 brennen@deploy1003: Finished deploy [phabricator/deployment@7e02037]: deploy phab1004 for [[phab:T431440|T431440]] (duration: 00m 47s) * 15:04 brennen@deploy1003: Started deploy [phabricator/deployment@7e02037]: deploy phab1004 for [[phab:T431440|T431440]] * 15:03 brennen@deploy1003: Finished deploy [phabricator/deployment@7e02037]: deploy phab2003 for [[phab:T431440|T431440]] (duration: 00m 51s) * 15:03 brennen@deploy1003: Started deploy [phabricator/deployment@7e02037]: deploy phab2003 for [[phab:T431440|T431440]] * 15:00 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71] (thin): Regular analytics weekly train THIN [analytics/refinery@7d8dc71f] (duration: 02m 10s) * 14:58 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71] (thin): Regular analytics weekly train THIN [analytics/refinery@7d8dc71f] * 14:58 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71]: Regular analytics weekly train [analytics/refinery@7d8dc71f] (duration: 04m 14s) * 14:54 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1037.eqiad.wmnet with reason: host reimage * 14:53 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71]: Regular analytics weekly train [analytics/refinery@7d8dc71f] * 14:53 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@7d8dc71f] (duration: 02m 00s) * 14:51 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@7d8dc71f] * 14:51 arnaudb@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on phab2003.codfw.wmnet,phab[1004-1006].eqiad.wmnet with reason: maintenance * 14:51 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1037.eqiad.wmnet with reason: host reimage * 14:50 JavierMonton: Deploying Refinery at {{Gerrit|7d8dc71f}} for change {{Gerrit|1308087}} / [[phab:T431318|T431318]] - update filerevision table sqoop and table * 14:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1007.eqiad.wmnet with reason: host reimage * 14:42 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1007.eqiad.wmnet with reason: host reimage * 14:40 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1083.eqiad.wmnet with OS trixie * 14:37 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:36 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:35 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-master@eqiad * 14:35 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 14:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 14:34 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1037 * 14:34 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1037 * 14:34 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:cleanMentorList.php --wiki=frwiki # [[phab:T427386|T427386]] * 14:34 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 14:34 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308112{{!}}Revert^2 "[Growth] frwiki: Deploy automated mentor list cleaner" (T427386)]] (duration: 06m 47s) * 14:34 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 14:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:33 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:32 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1037 * 14:31 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:31 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:29 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:29 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-master@eqiad * 14:29 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:29 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:29 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:28 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:27 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:27 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1308112{{!}}Revert^2 "[Growth] frwiki: Deploy automated mentor list cleaner" (T427386)]] * 14:26 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1007.eqiad.wmnet with OS trixie * 14:26 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:26 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 14:26 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:25 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:25 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:25 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:cleanMentorList.php --wiki=frwiki # [[phab:T427386|T427386]] * 14:24 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:24 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1037 - blake@cumin1003" * 14:24 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1037 - blake@cumin1003" * 14:20 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1083.eqiad.wmnet with reason: host reimage * 14:19 blake@cumin1003: START - Cookbook sre.dns.netbox * 14:19 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-master@codfw * 14:19 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 14:19 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1037 * 14:18 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1037.eqiad.wmnet with OS trixie * 14:18 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1037.eqiad.wmnet * 14:18 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 14:18 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1037.eqiad.wmnet * 14:18 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1037.eqiad.wmnet * 14:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2007.codfw.wmnet with OS trixie * 14:16 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1083.eqiad.wmnet with reason: host reimage * 14:15 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1036.eqiad.wmnet * 14:15 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1036.eqiad.wmnet * 14:14 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1036.eqiad.wmnet * 14:12 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-master@codfw * 14:11 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1120.eqiad.wmnet with OS trixie * 14:05 moritzm: installing distro-info-data updates from trixie/bookworm point releases * 14:04 fabfur: disable puppet on A:cp-text to selectively apply https://gerrit.wikimedia.org/r/c/operations/puppet/+/1308040 * 14:03 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] (duration: 27m 48s) * 14:00 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 14:00 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1083.eqiad.wmnet with OS trixie * 13:58 urbanecm@deploy1003: urbanecm: Continuing with deployment * 13:58 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:58 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1004.eqiad.wmnet with OS bookworm * 13:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2007.codfw.wmnet with reason: host reimage * 13:57 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 13:53 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1120.eqiad.wmnet with reason: host reimage * 13:50 moritzm: installing Linux 5.10.259 on Bullseye hosts * 13:47 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply * 13:47 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply * 13:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2007.codfw.wmnet with reason: host reimage * 13:46 cgoubert@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/aux-k8s-services/redioscope: apply * 13:46 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1120.eqiad.wmnet with reason: host reimage * 13:46 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:46 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:45 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:44 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:44 cgoubert@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/aux-k8s-services/redioscope: apply * 13:40 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 13:39 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:38 moritzm: installing e2fsprogs updates from Trixie point release * 13:35 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] * 13:33 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1120.eqiad.wmnet with OS trixie * 13:33 topranks: reset cr3-eqsin configuration so traffic uses it again after upgrade * 13:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1088.eqiad.wmnet with OS trixie * 13:32 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie * 13:32 cdobbins@cumin1003: conftool action : set/pooled=no; selector: name=dns7002.* * 13:29 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2007.codfw.wmnet with OS trixie * 13:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1004.eqiad.wmnet with reason: host reimage * 13:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2006.codfw.wmnet with OS trixie * 13:18 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1036.eqiad.wmnet with OS trixie * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aux-k8s-etcd1004.eqiad.wmnet with reason: host reimage * 13:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 13:16 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 13:15 jayme: Istio is being upgraded from 1.24.2 to 1.29.4 on wikikube staging eqiad and codfw - [[phab:T427401|T427401]] * 13:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1087.eqiad.wmnet with OS trixie * 13:14 topranks: reboot cr3-eqsin to install new JunOS and set PIC 0/0/0 to 100G * 13:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1088.eqiad.wmnet with reason: host reimage * 13:13 jmm@dns1004: END - running authdns-update * 13:12 jmm@dns1004: START - running authdns-update * 13:09 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1088.eqiad.wmnet with reason: host reimage * 13:07 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1082.eqiad.wmnet with OS trixie * 13:07 jmm@dns1004: END - running authdns-update * 13:06 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1004.eqiad.wmnet with OS bookworm * 13:05 jmm@dns1004: START - running authdns-update * 13:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2006.codfw.wmnet with reason: host reimage * 12:58 topranks: load updated JunOS on cr3-eqsin [[phab:T429386|T429386]] * 12:58 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1036.eqiad.wmnet with reason: host reimage * 12:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2001.codfw.wmnet * 12:57 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2006.codfw.wmnet with reason: host reimage * 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr1-codfw,cr[2-3]-eqsin,cr3-eqsin IPv6,cr3-eqsin.mgmt with reason: upgrade JunOS cr3-eqsin * 12:56 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lvs[5004-5006].eqsin.wmnet with reason: upgrade JunOS cr3-eqsin * 12:55 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 12:55 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 12:53 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1087.eqiad.wmnet with reason: host reimage * 12:53 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 12:52 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1088.eqiad.wmnet with OS trixie * 12:52 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1002.eqiad.wmnet * 12:52 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:51 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2001.codfw.wmnet * 12:49 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1036.eqiad.wmnet with reason: host reimage * 12:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 12:48 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1087.eqiad.wmnet with reason: host reimage * 12:44 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1082.eqiad.wmnet with reason: host reimage * 12:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1002.eqiad.wmnet * 12:42 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 12:42 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 12:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 12:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 12:39 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:39 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: move dumps-nfs IP to the shared one - filippo@cumin1003" * 12:39 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: move dumps-nfs IP to the shared one - filippo@cumin1003" * 12:39 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2006.codfw.wmnet with OS trixie * 12:38 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1082.eqiad.wmnet with reason: host reimage * 12:36 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:33 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:32 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1036 * 12:32 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1036 * 12:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2005.codfw.wmnet with OS trixie * 12:32 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1087.eqiad.wmnet with OS trixie * 12:30 jmm@dns1004: END - running authdns-update * 12:29 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1036 * 12:29 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1036.eqiad.wmnet 21.32.64.10.in-addr.arpa 1.2.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:29 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1036.eqiad.wmnet 21.32.64.10.in-addr.arpa 1.2.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:29 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:29 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1036 - blake@cumin1003" * 12:29 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1036 - blake@cumin1003" * 12:28 jmm@dns1004: START - running authdns-update * 12:26 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:26 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:23 blake@cumin1003: START - Cookbook sre.dns.netbox * 12:23 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1036 * 12:23 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1036.eqiad.wmnet with OS trixie * 12:22 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1036.eqiad.wmnet * 12:22 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1082.eqiad.wmnet with OS trixie * 12:22 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1036.eqiad.wmnet * 12:22 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1036.eqiad.wmnet * 12:21 marostegui: Restart mariadb@s7 on db1155 to pick up new filters - [[phab:T431124|T431124]] * 12:21 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 21 hosts with reason: restarting for replication filter * 12:20 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:19 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:14 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2005.codfw.wmnet with reason: host reimage * 12:14 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:08 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:07 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:07 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2005.codfw.wmnet with reason: host reimage * 12:06 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:06 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:06 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:05 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:05 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:04 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-master-eqiad * 12:04 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl1002.eqiad.wmnet * 12:04 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl1002.eqiad.wmnet * 12:04 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:04 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:03 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:03 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:03 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:03 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 11:59 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl1002.eqiad.wmnet * 11:59 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl1002.eqiad.wmnet * 11:59 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl1001.eqiad.wmnet * 11:59 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl1001.eqiad.wmnet * 11:56 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl1001.eqiad.wmnet * 11:56 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl1001.eqiad.wmnet * 11:56 klausman@cumin2002: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-master-eqiad * 11:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2005.codfw.wmnet with OS trixie * 11:49 blake@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on wikikube-worker1160.eqiad.wmnet with reason: Verifying matchers for silence * 11:42 topranks: cr3-eqsin, begin traffic drain to reset PIC and upgrade JunOS [[phab:T429386|T429386]] * 11:41 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs[5004-5006].eqsin.wmnet with reason: upgrade JunOS cr3-eqsin * 11:39 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr1-codfw,cr[2-3]-eqsin,cr3-eqsin IPv6,cr3-eqsin.mgmt with reason: upgrade JunOS cr3-eqsin * 11:36 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=thanos-fe2004.codfw.wmnet * 11:35 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1086.eqiad.wmnet with OS trixie * 11:35 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=thanos-fe2004.codfw.wmnet * 11:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1085.eqiad.wmnet with OS trixie * 11:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1086.eqiad.wmnet with reason: host reimage * 11:10 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1085.eqiad.wmnet with reason: host reimage * 11:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2004.codfw.wmnet with OS trixie * 11:03 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1086.eqiad.wmnet with reason: host reimage * 11:02 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1085.eqiad.wmnet with reason: host reimage * 10:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2004.codfw.wmnet with reason: host reimage * 10:48 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:46 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1086.eqiad.wmnet with OS trixie * 10:46 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1085.eqiad.wmnet with OS trixie * 10:44 cgoubert@deploy1003: Finished deploy [restbase/deploy@2fc37d4]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] (duration: 16m 44s) * 10:43 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2004.codfw.wmnet with reason: host reimage * 10:35 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:27 cgoubert@deploy1003: Started deploy [restbase/deploy@2fc37d4]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] * 10:27 cgoubert@deploy1003: Finished deploy [restbase/deploy@8a25036]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] (duration: 00m 45s) * 10:26 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1117.eqiad.wmnet with OS trixie * 10:26 cgoubert@deploy1003: Started deploy [restbase/deploy@8a25036]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] * 10:26 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host thanos-fe2004 * 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host thanos-fe2004 * 10:22 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1116.eqiad.wmnet with OS trixie * 10:21 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host thanos-fe2004 * 10:21 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) thanos-fe2004.codfw.wmnet 157.32.192.10.in-addr.arpa 7.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:20 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache thanos-fe2004.codfw.wmnet 157.32.192.10.in-addr.arpa 7.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:20 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:20 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host thanos-fe2004 - mvernon@cumin2003" * 10:20 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host thanos-fe2004 - mvernon@cumin2003" * 10:15 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2252: Repooling after reboot * 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:15 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 10:15 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2252: Repooling after reboot * 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1153.eqiad.wmnet * 10:14 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1153.eqiad.wmnet * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 10:14 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 10:12 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 10:12 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host thanos-fe2004 * 10:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2004.codfw.wmnet with OS trixie * 10:07 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1117.eqiad.wmnet with reason: host reimage * 10:03 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1116.eqiad.wmnet with reason: host reimage * 09:58 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:58 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1117.eqiad.wmnet with reason: host reimage * 09:57 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1116.eqiad.wmnet with reason: host reimage * 09:49 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 41 days, 15:00:00 on db2252.codfw.wmnet with reason: Security updates * 09:45 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1117.eqiad.wmnet with OS trixie * 09:45 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1116.eqiad.wmnet with OS trixie * 09:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1153: Security updates * 09:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:28 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:28 root@cumin1003: START - Cookbook sre.mysql.depool depool db1153: Security updates * 09:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1016: Security updates * 09:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:21 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:21 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1016: Security updates * 09:14 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:14 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 08:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1016: Security updates * 08:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:56 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:56 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1016: Security updates * 08:50 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:50 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:45 filippo@dns1006: END - running authdns-update * 08:43 filippo@dns1006: START - running authdns-update * 08:42 godog: switch dumps-nfs address to be shared with rsync/http - [[phab:T411248|T411248]] * 08:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1016: Security updates * 08:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:40 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:40 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1016: Security updates * 08:29 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host cirrussearch1111.eqiad.wmnet * 08:29 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:27 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:27 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:25 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1015: Security updates * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:09 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:09 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1015: Security updates * 07:42 Msz2001: Deployed private patch for Suggested Ivestigations * 07:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1015: Security updates * 07:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:41 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:41 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1015: Security updates * 07:40 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 07:11 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fingerprint warnings - oblivian@cumin1003" * 07:11 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fingerprint warnings - oblivian@cumin1003 * 07:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1024: Security updates * 07:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:11 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:11 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1024: Security updates * 07:10 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fingerprint warnings - oblivian@cumin1003 * 07:10 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fingerprint warnings - oblivian@cumin1003" * 06:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host cirrussearch1111.eqiad.wmnet * 06:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 06:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1024: Security updates * 06:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 06:48 root@cumin1003: START - Cookbook sre.mysql.parsercache * 06:48 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1024: Security updates * 06:42 moritzm: install nginx security updates * 06:31 root@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool pc1024: Security updates * 06:21 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1024: Security updates * 06:19 moritzm: installing php8.2 security updates * 06:15 moritzm: installing php8.4 security updates * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.7 (duration: 02m 38s) * 03:40 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] (duration: 37m 04s) * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 51s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-06 == * 23:30 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] (duration: 09m 39s) * 23:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1078.eqiad.wmnet with OS trixie * 23:26 jdlrobson@deploy1003: jdlrobson, bwang: Continuing with deployment * 23:22 jdlrobson@deploy1003: jdlrobson, bwang: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug) * 23:21 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] * 23:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1078.eqiad.wmnet with reason: host reimage * 23:06 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1078.eqiad.wmnet with reason: host reimage * 22:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1078.eqiad.wmnet with OS trixie * 22:29 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on cirrussearch1114.eqiad.wmnet with reason: reimage on hold until restore completes * 22:22 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on cirrussearch[1079,1115].eqiad.wmnet with reason: reimage on hold until restore completes * 21:18 maryum: Deployed security fix for [[phab:T428006|T428006]] * 20:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1079.eqiad.wmnet with OS trixie * 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1077.eqiad.wmnet with OS trixie * 20:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1115.eqiad.wmnet with OS trixie * 20:25 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1079.eqiad.wmnet with reason: host reimage * 20:21 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1079.eqiad.wmnet with reason: host reimage * 20:15 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] (duration: 08m 14s) * 20:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1077.eqiad.wmnet with reason: host reimage * 20:10 krinkle@deploy1003: krinkle, pushpaktiwari: Continuing with deployment * 20:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1115.eqiad.wmnet with reason: host reimage * 20:08 krinkle@deploy1003: krinkle, pushpaktiwari: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1077.eqiad.wmnet with reason: host reimage * 20:06 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] * 20:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1079.eqiad.wmnet with OS trixie * 20:04 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1115.eqiad.wmnet with reason: host reimage * 19:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1077.eqiad.wmnet with OS trixie * 19:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1115.eqiad.wmnet with OS trixie * 19:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 19:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 18:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1114.eqiad.wmnet with OS trixie * 18:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1114.eqiad.wmnet with reason: host reimage * 18:35 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1114.eqiad.wmnet with reason: host reimage * 18:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1112.eqiad.wmnet with OS trixie * 18:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1114.eqiad.wmnet with OS trixie * 18:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1072.eqiad.wmnet with OS trixie * 18:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1112.eqiad.wmnet with reason: host reimage * 18:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1112.eqiad.wmnet with reason: host reimage * 17:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1072.eqiad.wmnet with reason: host reimage * 17:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1112.eqiad.wmnet with OS trixie * 17:55 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1072.eqiad.wmnet with reason: host reimage * 17:39 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1072.eqiad.wmnet with OS trixie * 17:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1071.eqiad.wmnet with OS trixie * 17:18 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1070.eqiad.wmnet with OS trixie * 17:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1084.eqiad.wmnet with OS trixie * 16:54 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1071.eqiad.wmnet with reason: host reimage * 16:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1084.eqiad.wmnet with reason: host reimage * 16:51 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1070.eqiad.wmnet with reason: host reimage * 16:49 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1084.eqiad.wmnet with reason: host reimage * 16:38 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1071.eqiad.wmnet with OS trixie * 16:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1096.eqiad.wmnet with OS trixie * 16:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1070.eqiad.wmnet with OS trixie * 16:33 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1084.eqiad.wmnet with OS trixie * 16:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1089.eqiad.wmnet with OS trixie * 16:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1103.eqiad.wmnet with OS trixie * 16:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1096.eqiad.wmnet with reason: host reimage * 16:14 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1096.eqiad.wmnet with reason: host reimage * 16:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1089.eqiad.wmnet with reason: host reimage * 16:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1103.eqiad.wmnet with reason: host reimage * 16:02 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1003.eqiad.wmnet with OS bookworm * 16:01 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1089.eqiad.wmnet with reason: host reimage * 16:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1103.eqiad.wmnet with reason: host reimage * 15:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1096.eqiad.wmnet with OS trixie * 15:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1080.eqiad.wmnet with OS trixie * 15:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1089.eqiad.wmnet with OS trixie * 15:45 dancy@deploy1003: Installation of scap version "4.272.0" completed for 158 hosts * 15:43 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1103.eqiad.wmnet with OS trixie * 15:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1113.eqiad.wmnet with OS trixie * 15:41 dancy@deploy1003: Installing scap version "4.272.0" for 158 host(s) * 15:40 klausman@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 15:39 klausman@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 15:38 klausman@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 15:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1069.eqiad.wmnet with OS trixie * 15:37 klausman@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 15:36 klausman@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 15:34 klausman@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 15:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1080.eqiad.wmnet with reason: host reimage * 15:27 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1080.eqiad.wmnet with reason: host reimage * 15:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1113.eqiad.wmnet with reason: host reimage * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1069.eqiad.wmnet with reason: host reimage * 15:18 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1113.eqiad.wmnet with reason: host reimage * 15:16 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1069.eqiad.wmnet with reason: host reimage * 15:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:11 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1080.eqiad.wmnet with OS trixie * 15:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1113.eqiad.wmnet with OS trixie * 15:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1003.eqiad.wmnet with reason: host reimage * 14:47 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1003.eqiad.wmnet with OS bookworm * 14:33 elukey: rolled out spicerack on all cumin nodes - [[phab:T429699|T429699]] * 14:32 elukey: upgrade all bookworm hosts to pywmflib 3.1 - [[phab:T430552|T430552]] * 14:14 marostegui: Setup x4 eqiad topology [[phab:T404715|T404715]] * 14:13 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 14:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2230.codfw.wmnet * 14:07 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2230.codfw.wmnet * 13:59 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[2001-2002].codfw.wmnet * 13:51 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 13:45 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.major-upgrade (exit_code=97) * 13:45 cwilliams@cumin1003: dbmaint on s4@codfw [[phab:T429893|T429893]] * 13:45 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 13:42 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-master-codfw * 13:42 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl2002.codfw.wmnet * 13:42 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl2002.codfw.wmnet * 13:38 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl2002.codfw.wmnet * 13:38 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl2002.codfw.wmnet * 13:38 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl2001.codfw.wmnet * 13:38 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl2001.codfw.wmnet * 13:35 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl2001.codfw.wmnet * 13:35 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl2001.codfw.wmnet * 13:35 klausman@cumin2002: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-master-codfw * 12:30 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] (duration: 25m 11s) * 12:24 krinkle@deploy1003: krinkle: Continuing with deployment * 12:10 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2048.codfw.wmnet * 12:09 krinkle@deploy1003: krinkle: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:08 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2048.codfw.wmnet * 12:05 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] * 11:57 moritzm: installing curl security updates * 11:49 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:31 moritzm: installing nano security updates * 11:07 moritzm: failover Ganeti master in codfw to ganeti2032 [[phab:T430909|T430909]] * 11:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:04 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest1005.eqiad.wmnet with OS trixie * 11:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:50 jmm@dns1004: END - running authdns-update * 10:47 jmm@dns1004: START - running authdns-update * 10:47 jmm@dns1004: START - running authdns-update * 10:46 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:44 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest1005.eqiad.wmnet with reason: host reimage * 10:38 elukey@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest1005.eqiad.wmnet with reason: host reimage * 10:31 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:31 marostegui: Setup x4 codfw topology [[phab:T404715|T404715]] * 10:31 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 10:24 elukey: spicerack 13.0.0 deployed on cumin2002 * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 10:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 10:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 10:21 elukey@cumin2002: START - Cookbook sre.hosts.reimage for host sretest1005.eqiad.wmnet with OS trixie * 10:20 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:19 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:17 elukey: uploaded spicerack_13.0.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia * 09:54 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:52 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:20 elukey: upgrade all bullseye hosts to pywmflib 3.1 - [[phab:T430552|T430552]] * 09:10 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1015.eqiad.wmnet,service=s4 * 09:10 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1015.eqiad.wmnet,service=s6 * 09:07 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 08:58 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:56 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 08:56 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 08:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 08:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 08:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin2002.codfw.wmnet * 08:06 godog: remove cloudvirt1046, cloudvirt1062, cloudvirt1074, cloudvirt1075 from maintenance aggregate and put them in network-ovs - [[phab:T424802|T424802]] * 08:00 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin2002.codfw.wmnet * 07:58 hashar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] (duration: 32m 53s) * 07:57 fabfur: repooled cp4038 * 07:57 fabfur@cumin1003: conftool action : set/pooled=yes; selector: name=cp4038.* * 07:53 moritzm: installing pyjwt security updates * 07:47 moritzm: installing openjpeg2 security updates * 07:45 hashar@deploy1003: vadymts1, hashar: Continuing with deployment * 07:43 hashar@deploy1003: vadymts1, hashar: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:38 moritzm: installing python-urllib3 security updates * 07:37 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 07:30 fabfur: depooled cp4038 to investigate on possible maxmind failure * 07:30 fabfur@cumin1003: conftool action : set/pooled=no; selector: name=cp4038.* * 07:30 fabfur@cumin1003: conftool action : set/pooled=yes; selector: name=cp4038.* * 07:29 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 07:25 hashar@deploy1003: Started scap sync-world: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] * 06:13 moritzm: installing Linux 6.12.95 on trixie hosts * 05:20 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s6 * 05:20 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s4 * 05:19 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1015.eqiad.wmnet with reason: cloning * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 08s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-05 == * 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 01m 08s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-04 == * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 58s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-03 == * 17:08 topranks: revert protocol preference changes on cr3-ulsfo after upgrade * 16:53 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on cr2-eqord with reason: upgrade JunOS cr3-ulsfo * 16:53 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on cr4-ulsfo with reason: upgrade JunOS cr3-ulsfo * 16:48 topranks: reboot cr3-ulsfo to upgrade JunOS and reset linecard [[phab:T424839|T424839]] * 15:52 topranks: adjust outbound BGP policies on cr3-ulsfo to drain router of traffic [[phab:T424839|T424839]] * 15:45 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on lvs[4008-4010].ulsfo.wmnet with reason: upgrade JunOS cr3-ulsfo * 15:44 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on asw1-[22-23]-ulsfo,cr3-ulsfo,cr3-ulsfo IPv6,cr3-ulsfo.mgmt with reason: upgrade JunOS cr3-ulsfo * 15:36 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 15:35 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 15:35 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 14:40 cmooney@dns3003: END - running authdns-update * 14:26 cmooney@dns3003: START - running authdns-update * 14:26 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:26 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to ulsfo - cmooney@cumin1003" * 14:19 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to ulsfo - cmooney@cumin1003" * 14:16 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:38 sukhe@dns1004: END - running authdns-update * 13:35 sukhe@dns1004: START - running authdns-update * 13:26 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 13:26 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 13:26 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet * 13:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 13:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 13:16 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host sretest1005.eqiad.wmnet * 13:16 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 13:16 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 13:15 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 13:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:14 moritzm: imported samplicator 1.3.8rc1-1+deb13u1 to trixie-wikimedia/main [[phab:T337208|T337208]] * 13:13 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:07 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 13:07 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 13:02 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:02 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:58 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:57 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:57 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:53 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet * 12:50 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 12:47 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:41 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:40 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:39 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:32 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet * 12:26 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet * 12:23 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2005.wikimedia.org * 12:19 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2005.wikimedia.org * 12:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 12:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup[2004-2007].codfw.wmnet * 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[2004-2007].codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin2003" * 12:15 jynus@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[2004-2007].codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin2003" * 12:09 jynus@cumin2003: START - Cookbook sre.dns.netbox * 11:58 jynus@cumin2003: START - Cookbook sre.hosts.decommission for hosts backup[2004-2007].codfw.wmnet * 10:40 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup[1004-1007].eqiad.wmnet * 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[1004-1007].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 10:01 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[1004-1007].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 09:52 jynus@cumin1003: START - Cookbook sre.dns.netbox * 09:39 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:36 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup[1004-1007].eqiad.wmnet * 09:36 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:25 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 09:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 09:16 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 09:05 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:04 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:00 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:59 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:57 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:55 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:50 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 08:50 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 08:49 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 08:49 atsukoito: depooling cirrussearch in codfw because of regression after upgrade [[phab:T431091|T431091]] * 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts mirror1001.wikimedia.org * 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: mirror1001.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 08:29 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: mirror1001.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 08:18 jmm@cumin2003: START - Cookbook sre.dns.netbox * 08:11 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts mirror1001.wikimedia.org * 06:15 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 18s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-02 == * 22:55 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host contint1003.wikimedia.org with OS trixie * 22:29 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on contint1003.wikimedia.org with reason: host reimage * 22:23 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on contint1003.wikimedia.org with reason: host reimage * 22:05 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host contint1003.wikimedia.org with OS trixie * 22:03 mutante: contint1003 (zuul.wikimedia.org) - reimaging because of [[phab:T430510|T430510]]#12067628 [[phab:T418521|T418521]] * 22:03 dzahn@cumin2002: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on zuul.wikimedia.org with reason: reimage * 21:39 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 18s) * 21:39 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 21:20 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1003.eqiad.wmnet, repooling source-only afterwards * 21:19 sbassett: Deployed security fix for [[phab:T428829|T428829]] * 20:58 cmooney@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Release v0.11.2 update for new Aerleon - cmooney@cumin1003 * 20:55 cmooney@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Release v0.11.2 update for new Aerleon - cmooney@cumin1003 * 20:40 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] (duration: 12m 35s) * 20:36 arlolra@deploy1003: cscott, arlolra: Continuing with deployment * 20:35 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 20s) * 20:35 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 20:33 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host contint2003.wikimedia.org with OS trixie * 20:31 arlolra@deploy1003: cscott, arlolra: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Cha * 20:28 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] * 20:17 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] (duration: 08m 13s) * 20:14 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on contint2003.wikimedia.org with reason: host reimage * 20:13 sbassett@deploy1003: sbassett: Continuing with deployment * 20:11 sbassett@deploy1003: sbassett: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:09 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] * 20:08 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 20:08 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 20:08 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on contint2003.wikimedia.org with reason: host reimage * 20:05 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1003.eqiad.wmnet, repooling source-only afterwards * 19:49 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host contint2003.wikimedia.org with OS trixie * 19:48 mutante: contint2003 - reimaging because of [[phab:T430510|T430510]]#12067628 [[phab:T418521|T418521]] * 18:39 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 18:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 18:13 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2002.codfw.wmnet -> wcqs2003.codfw.wmnet, repooling source-only afterwards * 17:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1003.eqiad.wmnet with OS bookworm * 17:52 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1005.eqiad.wmnet * 17:52 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1005.eqiad.wmnet * 17:51 jasmine@cumin2002: conftool action : set/pooled=yes:weight=10; selector: name=wikikube-ctrl1005.eqiad.wmnet * 17:48 jasmine_: homer "cr*eqiad*" commit "Added new stacked control plane wikikube-ctrl1005" * 17:44 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply * 17:44 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply * 17:31 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] (duration: 09m 33s) * 17:26 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 17:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1003.eqiad.wmnet with reason: host reimage * 17:23 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:21 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] * 17:18 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1003.eqiad.wmnet with reason: host reimage * 17:16 rscout@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply * 17:16 rscout@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply * 17:16 rscout@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply * 17:15 rscout@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply * 17:12 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on wcqs[2002-2003].codfw.wmnet,wcqs1002.eqiad.wmnet with reason: reimaging hosts * 17:08 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 17:08 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 17:08 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 17:07 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 17:05 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 17:05 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 17:03 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "running to make sure all updates are synced - cmooney@cumin1003" * 17:03 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "running to make sure all updates are synced - cmooney@cumin1003" * 17:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs1003 * 17:00 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs1003 * 17:00 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 17:00 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1003.eqiad.wmnet with OS bookworm * 16:58 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Re-running - btullis@cumin1003" * 16:58 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Re-running - btullis@cumin1003" * 16:58 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2002.codfw.wmnet -> wcqs2003.codfw.wmnet, repooling source-only afterwards * 16:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-master1004.eqiad.wmnet with OS bookworm * 16:58 btullis@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 16:57 tappof: bump space for prometheus k8s-aux in eqiad * 16:55 cmooney@dns3003: END - running authdns-update * 16:55 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:55 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to eqsin - cmooney@cumin1003" * 16:55 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to eqsin - cmooney@cumin1003" * 16:53 cmooney@dns3003: START - running authdns-update * 16:52 ryankemper: [ml-serve-eqiad] Cleared out 1302 failed (Evicted) pods: `kubectl -n llm delete pods --field-selector=status.phase=Failed`, freeing calico-kube-controllers from OOM crashloop (evictions were caused by disk pressure) * 16:49 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 16:46 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:39 rzl@dns1004: END - running authdns-update * 16:37 rzl@dns1004: START - running authdns-update * 16:36 rzl@dns1004: START - running authdns-update * 16:35 rzl@deploy1003: Finished scap sync-world: [[phab:T416623|T416623]] (duration: 10m 19s) * 16:34 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 16:33 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-master1004.eqiad.wmnet with reason: host reimage * 16:30 rzl@deploy1003: rzl: Continuing with deployment * 16:28 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-master1004.eqiad.wmnet with reason: host reimage * 16:26 rzl@deploy1003: rzl: [[phab:T416623|T416623]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:25 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 16:25 rzl@deploy1003: Started scap sync-world: [[phab:T416623|T416623]] * 16:25 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 16:24 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 16:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: sync * 16:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: sync * 16:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-master1004.eqiad.wmnet with OS bookworm * 16:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-master1003.eqiad.wmnet with OS bookworm * 16:11 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 16:11 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 16:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Security updates * 16:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 16:08 root@cumin1003: START - Cookbook sre.mysql.parsercache * 16:08 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Security updates * 15:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-master1003.eqiad.wmnet with reason: host reimage * 15:54 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:54 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:54 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:54 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-master1003.eqiad.wmnet with reason: host reimage * 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Security updates * 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:45 root@cumin1003: START - Cookbook sre.mysql.parsercache * 15:45 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Security updates * 15:42 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-master1003.eqiad.wmnet with OS bookworm * 15:24 moritzm: installing busybox updates from bookworm point release * 15:20 moritzm: installing busybox updates from trixie point release * 15:15 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1021: Security updates * 15:15 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:15 root@cumin1003: START - Cookbook sre.mysql.parsercache * 15:15 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1021: Security updates * 15:13 moritzm: installing giflib security updates * 15:08 moritzm: installing Tomcat security updates * 14:57 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 14:56 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 14:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:53 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Unblock taavi - oblivian@cumin1003" * 14:53 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Unblock taavi - oblivian@cumin1003 * 14:53 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1021: Security updates * 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:53 root@cumin1003: START - Cookbook sre.mysql.parsercache * 14:53 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1021: Security updates * 14:53 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Unblock taavi - oblivian@cumin1003 * 14:52 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Unblock taavi - oblivian@cumin1003" * 14:46 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94711 and previous config saved to /var/cache/conftool/dbconfig/20260702-144644-fceratto.json * 14:36 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205', diff saved to https://phabricator.wikimedia.org/P94709 and previous config saved to /var/cache/conftool/dbconfig/20260702-143636-fceratto.json * 14:32 moritzm: installing libdbi-perl security updates * 14:26 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205', diff saved to https://phabricator.wikimedia.org/P94708 and previous config saved to /var/cache/conftool/dbconfig/20260702-142628-fceratto.json * 14:16 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94707 and previous config saved to /var/cache/conftool/dbconfig/20260702-141621-fceratto.json * 14:12 moritzm: installing rsync security updates * 14:11 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox) * 14:10 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94706 and previous config saved to /var/cache/conftool/dbconfig/20260702-140959-fceratto.json * 14:09 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2205.codfw.wmnet with reason: Maintenance * 14:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2205: Repooling after switchover * 14:07 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-test-master1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 14:06 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 14:06 Tran: Deployed patch for [[phab:T427287|T427287]] * 14:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:59 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2205: Repooling after switchover * 13:59 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2205: Repooling after switchover * 13:59 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:55 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2205: Repooling after switchover * 13:55 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2205 [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94704 and previous config saved to /var/cache/conftool/dbconfig/20260702-135505-fceratto.json * 13:54 moritzm: installing sed security updates * 13:53 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:52 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2209 to s3 primary [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94703 and previous config saved to /var/cache/conftool/dbconfig/20260702-135235-fceratto.json * 13:52 federico3: Starting s3 codfw failover from db2205 to db2209 - [[phab:T430912|T430912]] * 13:51 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:51 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 13:48 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:47 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2209 with weight 0 [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94702 and previous config saved to /var/cache/conftool/dbconfig/20260702-134719-fceratto.json * 13:47 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Primary switchover s3 [[phab:T430912|T430912]] * 13:44 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:44 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:44 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:40 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 13:38 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 13:37 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 13:36 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 13:36 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:34 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 13:30 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:29 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:29 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:27 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:26 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:25 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling restart_daemons on A:wikidough * 13:23 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 13:22 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 13:17 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 13:17 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns1004.wikimedia.org * 13:12 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:11 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart (exit_code=97) rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough * 13:11 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=97) rolling restart_daemons on A:wikidough * 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough * 13:09 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] (duration: 07m 20s) * 13:05 aude@deploy1003: jdrewniak, aude: Continuing with deployment * 13:04 aude@deploy1003: jdrewniak, aude: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:02 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] * 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts wdqs-categories1001.eqiad.wmnet * 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: wdqs-categories1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 12:10 jmm@dns1004: END - running authdns-update * 12:07 jmm@dns1004: START - running authdns-update * 11:51 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: wdqs-categories1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 11:44 btullis@cumin1003: START - Cookbook sre.dns.netbox * 11:42 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 11:42 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 11:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet * 11:39 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts wdqs-categories1001.eqiad.wmnet * 11:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet * 11:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet * 11:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet * 11:29 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 11:29 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 10:57 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2214: Repooling * 10:49 jmm@dns1004: END - running authdns-update * 10:47 jmm@dns1004: START - running authdns-update * 10:31 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94698 and previous config saved to /var/cache/conftool/dbconfig/20260702-103146-fceratto.json * 10:21 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213', diff saved to https://phabricator.wikimedia.org/P94696 and previous config saved to /var/cache/conftool/dbconfig/20260702-102137-fceratto.json * 10:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:19 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb1017.eqiad.wmnet * 10:18 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 10:18 fceratto@cumin1003: Removing es1033 from zarcillo [[phab:T408772|T408772]] * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts es1033.eqiad.wmnet * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: es1033.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:14 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: es1033.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:13 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb1017.eqiad.wmnet * 10:12 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2214.codfw.wmnet * 10:12 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2214.codfw.wmnet * 10:12 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2214: Repooling * 10:11 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213', diff saved to https://phabricator.wikimedia.org/P94693 and previous config saved to /var/cache/conftool/dbconfig/20260702-101130-fceratto.json * 10:10 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:10 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:03 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts es1033.eqiad.wmnet * 10:03 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 10:01 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94691 and previous config saved to /var/cache/conftool/dbconfig/20260702-100122-fceratto.json * 09:55 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94690 and previous config saved to /var/cache/conftool/dbconfig/20260702-095529-fceratto.json * 09:55 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2213.codfw.wmnet with reason: Maintenance * 09:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 09:53 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2213: Repooling after switchover * 09:51 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover * 09:44 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2213: Repooling after switchover * 09:39 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover * 09:39 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2213 [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94688 and previous config saved to /var/cache/conftool/dbconfig/20260702-093859-fceratto.json * 09:36 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2192 to s5 primary [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94687 and previous config saved to /var/cache/conftool/dbconfig/20260702-093650-fceratto.json * 09:36 federico3: Starting s5 codfw failover from db2213 to db2192 - [[phab:T430923|T430923]] * 09:30 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94686 and previous config saved to /var/cache/conftool/dbconfig/20260702-093004-fceratto.json * 09:24 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2192 with weight 0 [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94685 and previous config saved to /var/cache/conftool/dbconfig/20260702-092455-fceratto.json * 09:24 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 23 hosts with reason: Primary switchover s5 [[phab:T430923|T430923]] * 09:19 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220', diff saved to https://phabricator.wikimedia.org/P94684 and previous config saved to /var/cache/conftool/dbconfig/20260702-091957-fceratto.json * 09:16 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] (duration: 06m 57s) * 09:13 moritzm: installing libgcrypt20 security updates * 09:12 kharlan@deploy1003: kharlan: Continuing with deployment * 09:11 kharlan@deploy1003: kharlan: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:09 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220', diff saved to https://phabricator.wikimedia.org/P94683 and previous config saved to /var/cache/conftool/dbconfig/20260702-090950-fceratto.json * 09:09 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] * 09:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 09:01 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] (duration: 07m 07s) * 08:59 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94682 and previous config saved to /var/cache/conftool/dbconfig/20260702-085942-fceratto.json * 08:57 kharlan@deploy1003: kharlan: Continuing with deployment * 08:56 kharlan@deploy1003: kharlan: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:54 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] * 08:52 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:52 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:52 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94681 and previous config saved to /var/cache/conftool/dbconfig/20260702-085237-fceratto.json * 08:52 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2220.codfw.wmnet with reason: Maintenance * 08:43 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:40 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 08:25 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] (duration: 11m 44s) * 08:21 cscott@deploy1003: cscott: Continuing with deployment * 08:16 cscott@deploy1003: cscott: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:14 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] * 08:08 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 08:08 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1244: Migration of db1244.eqiad.wmnet completed * 08:02 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:02 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:01 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] (duration: 18m 58s) * 08:01 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:59 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 07:59 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:59 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:59 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2006.wikimedia.org * 07:58 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:57 cscott@deploy1003: cscott: Continuing with deployment * 07:56 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:56 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:56 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:55 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:55 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:55 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:54 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2006.wikimedia.org * 07:54 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:54 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 07:44 cscott@deploy1003: cscott: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:44 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2005.wikimedia.org * 07:44 moritzm: installing node-lodash security updates * 07:42 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] * 07:39 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2005.wikimedia.org * 07:30 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] (duration: 07m 28s) * 07:26 cscott@deploy1003: ssastry, cscott: Continuing with deployment * 07:25 cscott@deploy1003: ssastry, cscott: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:23 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1244: Migration of db1244.eqiad.wmnet completed * 07:22 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] * 07:16 wmde-fisch@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] (duration: 06m 55s) * 07:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1244.eqiad.wmnet with OS trixie * 07:11 wmde-fisch@deploy1003: wmde-fisch: Continuing with deployment * 07:11 wmde-fisch@deploy1003: wmde-fisch: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:09 wmde-fisch@deploy1003: Started scap sync-world: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] * 06:54 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1244.eqiad.wmnet with reason: host reimage * 06:50 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1244.eqiad.wmnet with reason: host reimage * 06:38 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1250.eqiad.wmnet with OS trixie * 06:34 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db1244.eqiad.wmnet with OS trixie * 06:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1244: Upgrading db1244.eqiad.wmnet * 06:25 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1244: Upgrading db1244.eqiad.wmnet * 06:25 cwilliams@cumin1003: dbmaint on s4@eqiad [[phab:T429893|T429893]] * 06:25 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 06:15 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1250.eqiad.wmnet with reason: host reimage * 06:14 cwilliams@dns1006: END - running authdns-update * 06:12 cwilliams@dns1006: START - running authdns-update * 06:11 cwilliams@dns1006: END - running authdns-update * 06:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db1244 [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94676 and previous config saved to /var/cache/conftool/dbconfig/20260702-061059-cwilliams.json * 06:09 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1250.eqiad.wmnet with reason: host reimage * 06:09 cwilliams@dns1006: START - running authdns-update * 06:08 aokoth@cumin1003: END (PASS) - Cookbook sre.vrts.upgrade (exit_code=0) on VRTS host vrts1003.eqiad.wmnet * 06:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db1160 to s4 primary and set section read-write [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94675 and previous config saved to /var/cache/conftool/dbconfig/20260702-060746-cwilliams.json * 06:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Set s4 eqiad as read-only for maintenance - [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94674 and previous config saved to /var/cache/conftool/dbconfig/20260702-060704-cwilliams.json * 06:06 cezmunsta: Starting s4 eqiad failover from db1244 to db1160 - [[phab:T430817|T430817]] * 06:04 aokoth@cumin1003: START - Cookbook sre.vrts.upgrade on VRTS host vrts1003.eqiad.wmnet * 05:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db1160 with weight 0 [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94673 and previous config saved to /var/cache/conftool/dbconfig/20260702-055927-cwilliams.json * 05:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 40 hosts with reason: Primary switchover s4 [[phab:T430817|T430817]] * 05:55 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1250.eqiad.wmnet with OS trixie * 05:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on db1250.eqiad.wmnet with reason: m3 master switchover [[phab:T430158|T430158]] * 05:39 marostegui: Failover m3 (phabricator) from db1250 to db1228 - [[phab:T430158|T430158]] * 05:32 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2234].codfw.wmnet,db[1217,1228,1250].eqiad.wmnet with reason: m3 master switchover [[phab:T430158|T430158]] * 04:45 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] (duration: 09m 08s) * 04:41 tstarling@deploy1003: tstarling, reedy: Continuing with deployment * 04:38 tstarling@deploy1003: tstarling, reedy: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 04:36 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 59s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:16 ryankemper: [[phab:T429844|T429844]] [opensearch] completed `cirrussearch2111` reimage; all codfw search clusters are green, all nodes now report `OpenSearch 2.19.5`, and the temporary chi voting exclusion has been removed * 00:57 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2111.codfw.wmnet with OS trixie * 00:29 ryankemper: [[phab:T429844|T429844]] [opensearch] depooled codfw search-omega/search-psi discovery records to match existing codfw search depool during OpenSearch 2.19 migration * 00:29 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2111.codfw.wmnet with reason: host reimage * 00:29 ryankemper@cumin2002: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 00:29 ryankemper@cumin2002: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 00:22 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2111.codfw.wmnet with reason: host reimage * 00:01 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2111.codfw.wmnet with OS trixie * 00:00 ryankemper: [[phab:T429844|T429844]] [opensearch] chi cluster recovered after stopping `opensearch_1@production-search-codfw` on `cirrussearch2111` == 2026-07-01 == * 23:59 ryankemper: [[phab:T429844|T429844]] [opensearch] stopped `opensearch_1@production-search-codfw` on `cirrussearch2111` after chi cluster-manager election churn following `voting_config_exclusions` POST; hoping this triggers a re-election * 23:52 cscott@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 23:51 cscott@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 23:51 cscott@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 23:50 cscott@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2003.codfw.wmnet with OS bookworm * 22:29 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 22:13 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 22:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2084.codfw.wmnet with OS trixie * 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2003.codfw.wmnet with reason: host reimage * 22:03 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 22:01 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2003.codfw.wmnet with reason: host reimage * 21:50 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 21:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2084.codfw.wmnet with reason: host reimage * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2003 * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2003 * 21:42 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2003 * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2003.codfw.wmnet 45.48.192.10.in-addr.arpa 5.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:42 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2003.codfw.wmnet 45.48.192.10.in-addr.arpa 5.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2003 - bking@cumin2003" * 21:42 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2003 - bking@cumin2003" * 21:36 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2084.codfw.wmnet with reason: host reimage * 21:35 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:34 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2003 * 21:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2003.codfw.wmnet with OS bookworm * 21:19 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2084.codfw.wmnet with OS trixie * 21:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2081.codfw.wmnet with OS trixie * 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2108.codfw.wmnet with OS trixie * 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2081.codfw.wmnet with reason: host reimage * 20:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2081.codfw.wmnet with reason: host reimage * 20:28 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2081.codfw.wmnet with OS trixie * 20:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2108.codfw.wmnet with reason: host reimage * 20:19 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2108.codfw.wmnet with reason: host reimage * 19:59 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2108.codfw.wmnet with OS trixie * 19:46 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2093.codfw.wmnet with OS trixie * 19:44 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 19:44 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jasmine@cumin2002" * 19:43 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jasmine@cumin2002" * 19:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2080.codfw.wmnet with OS trixie * 19:28 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 19:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2093.codfw.wmnet with reason: host reimage * 19:18 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 19:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2093.codfw.wmnet with reason: host reimage * 19:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2080.codfw.wmnet with reason: host reimage * 19:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2080.codfw.wmnet with reason: host reimage * 18:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2093.codfw.wmnet with OS trixie * 18:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2080.codfw.wmnet with OS trixie * 18:27 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 18:18 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] (duration: 09m 15s) * 18:13 jgiannelos@deploy1003: jgiannelos, neriah: Continuing with deployment * 18:11 jgiannelos@deploy1003: jgiannelos, neriah: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:09 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] * 17:40 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 16:58 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 30 hosts * 16:57 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for 30 hosts * 16:52 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2202.codfw.wmnet * 16:52 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2202.codfw.wmnet * 16:51 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt * 16:51 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt * 16:51 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lvs2012.codfw.wmnet * 16:51 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for lvs2012.codfw.wmnet * 16:49 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2076.codfw.wmnet with OS trixie * 16:49 brett: Start pybal on lvs2012 - [[phab:T429861|T429861]] * 16:49 pt1979@cumin1003: END (ERROR) - Cookbook sre.hosts.remove-downtime (exit_code=97) for 59 hosts * 16:48 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for 59 hosts * 16:42 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2061.codfw.wmnet with OS trixie * 16:30 dancy@deploy1003: Installation of scap version "4.271.0" completed for 2 hosts * 16:28 dancy@deploy1003: Installing scap version "4.271.0" for 2 host(s) * 16:23 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2076.codfw.wmnet with reason: host reimage * 16:19 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2061.codfw.wmnet with reason: host reimage * 16:18 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2076.codfw.wmnet with reason: host reimage * 16:18 jasmine@dns1004: END - running authdns-update * 16:16 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host restbase2039.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 16:16 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host restbase2039.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 16:16 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2061.codfw.wmnet with reason: host reimage * 16:15 jasmine@dns1004: START - running authdns-update * 16:14 jasmine@dns1004: END - running authdns-update * 16:12 jasmine@dns1004: START - running authdns-update * 16:07 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2202.codfw.wmnet with reason: maintenance * 16:06 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt with reason: Junos upograde * 16:00 papaul: ongoing maintenance on lsw1-b2-codfw * 16:00 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2076.codfw.wmnet with OS trixie * 15:59 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt * 15:59 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt * 15:57 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2061.codfw.wmnet with OS trixie * 15:55 pt1979@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2042,2046].codfw.wmnet * 15:55 pt1979@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2042,2046].codfw.wmnet * 15:51 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 15:51 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2220: Repooling after switchover * 15:50 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 15:50 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 15:48 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2092.codfw.wmnet with OS trixie * 15:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 15:40 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 15:38 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 15:37 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 15:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply * 15:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply * 15:32 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 15:32 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 15:30 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 15:29 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 15:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 15:25 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:22 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs2012.codfw.wmnet with reason: Rack B2 maintenance - [[phab:T429861|T429861]] * 15:21 brett: Stopping pybal on lvs2012 in preparation for codfw rack b2 maintenance - [[phab:T429861|T429861]] * 15:20 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2092.codfw.wmnet with reason: host reimage * 15:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:12 _joe_: restarted manually alertmanager-irc-relay * 15:12 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2092.codfw.wmnet with reason: host reimage * 15:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:12 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt with reason: Junos upograde * 15:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover * 15:07 pt1979@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2042,2046].codfw.wmnet * 15:06 pt1979@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2042,2046].codfw.wmnet * 15:02 papaul: ongoing maintenance on lsw1-a8-codfw * 14:31 topranks: POWERING DOWN CR1-EQIAD for line card installation [[phab:T426343|T426343]] * 14:31 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] (duration: 08m 57s) * 14:29 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:26 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 14:24 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:22 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] * 14:22 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover * 14:16 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:15 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover * 14:14 topranks: re-enable routing-engine graceful-failover on cr1-eqiad [[phab:T417873|T417873]] * 14:13 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:13 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2220: Repooling after switchover * 14:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:12 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:12 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:11 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:08 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] (duration: 10m 01s) * 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:07 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2220 [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94664 and previous config saved to /var/cache/conftool/dbconfig/20260701-140729-fceratto.json * 14:06 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:06 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:06 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:05 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2159 to s7 primary [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94663 and previous config saved to /var/cache/conftool/dbconfig/20260701-140503-fceratto.json * 14:04 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:04 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 14:04 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 14:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:04 dreamyjazz@deploy1003: anzx, dreamyjazz: Continuing with deployment * 14:04 federico3: Starting s7 codfw failover from db2220 to db2159 - [[phab:T430826|T430826]] * 14:03 jmm@dns1004: END - running authdns-update * 14:03 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:03 topranks: flipping cr1-eqiad active routing-enginer back to RE0 [[phab:T417873|T417873]] * 14:03 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudsw1-c8-eqiad,cloudsw1-d5-eqiad with reason: router upgrades eqiad * 14:01 jmm@dns1004: START - running authdns-update * 14:00 dreamyjazz@deploy1003: anzx, dreamyjazz: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:59 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2159 with weight 0 [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94662 and previous config saved to /var/cache/conftool/dbconfig/20260701-135906-fceratto.json * 13:58 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] * 13:57 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s7 [[phab:T430826|T430826]] * 13:56 topranks: reboot routing-enginer RE0 on cr1-eqiad [[phab:T417873|T417873]] * 13:48 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1006.wikimedia.org * 13:44 atsuko@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cirrussearch2092.codfw.wmnet with OS trixie * 13:43 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1006.wikimedia.org * 13:41 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2092.codfw.wmnet with OS trixie * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1005.wikimedia.org * 13:37 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1005.wikimedia.org * 13:37 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on pfw1-eqiad with reason: router upgrades eqiad * 13:35 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on lvs[1017-1020].eqiad.wmnet with reason: router upgrades eqiad * 13:34 caro@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] (duration: 07m 59s) * 13:30 caro@deploy1003: caro: Continuing with deployment * 13:28 caro@deploy1003: caro: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:27 topranks: route-engine failover cr1-eqiad * 13:26 caro@deploy1003: Started scap sync-world: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] * 13:15 topranks: rebooting routing-engine 1 on cr1-eqiad [[phab:T417873|T417873]] * 13:13 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] (duration: 08m 29s) * 13:13 moritzm: installing qemu security updates * 13:11 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 13:11 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 13:09 jgiannelos@deploy1003: jgiannelos: Continuing with deployment * 13:08 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 13:07 jgiannelos@deploy1003: jgiannelos: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:06 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 13:06 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2214.codfw.wmnet with reason: Maintenance * 13:05 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2214: Repooling after switchover * 13:05 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] * 13:04 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2214: Repooling after switchover * 13:04 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2214 [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94660 and previous config saved to /var/cache/conftool/dbconfig/20260701-130413-fceratto.json * 13:01 moritzm: installing python3.13 security updates * 13:00 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2229 to s6 primary [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94659 and previous config saved to /var/cache/conftool/dbconfig/20260701-125959-fceratto.json * 12:59 federico3: Starting s6 codfw failover from db2214 to db2229 - [[phab:T430814|T430814]] * 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on 13 hosts with reason: router upgrade and line card install * 12:51 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2229 with weight 0 [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94658 and previous config saved to /var/cache/conftool/dbconfig/20260701-125149-fceratto.json * 12:51 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 21 hosts with reason: Primary switchover s6 [[phab:T430814|T430814]] * 12:50 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2189.codfw.wmnet * 12:50 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2189.codfw.wmnet * 12:42 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2100.codfw.wmnet with OS trixie * 12:38 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2083.codfw.wmnet with OS trixie * 12:19 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2083.codfw.wmnet with reason: host reimage * 12:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 12:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2240: Migration of db2240.codfw.wmnet completed * 12:14 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2100.codfw.wmnet with reason: host reimage * 12:09 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2083.codfw.wmnet with reason: host reimage * 12:09 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2100.codfw.wmnet with reason: host reimage * 12:00 topranks: drain traffic on cr1-eqiad to allow for line card install and JunOS upgrade [[phab:T426343|T426343]] * 11:52 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2083.codfw.wmnet with OS trixie * 11:50 cmooney@dns2005: END - running authdns-update * 11:49 cmooney@dns2005: START - running authdns-update * 11:48 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2100.codfw.wmnet with OS trixie * 11:40 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/zotero: apply * 11:40 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/zotero: apply * 11:36 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/zotero: apply * 11:36 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/zotero: apply * 11:31 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2240: Migration of db2240.codfw.wmnet completed * 11:30 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply * 11:28 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply * 11:27 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:27 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:27 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:27 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:27 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:23 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2240.codfw.wmnet with OS trixie * 11:20 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:20 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:17 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:16 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:16 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:15 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2086.codfw.wmnet with OS trixie * 11:14 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2106.codfw.wmnet with OS trixie * 11:14 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:13 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:12 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:09 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2115.codfw.wmnet with OS trixie * 11:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2240.codfw.wmnet with reason: host reimage * 11:00 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2240.codfw.wmnet with reason: host reimage * 10:53 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2106.codfw.wmnet with reason: host reimage * 10:49 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2086.codfw.wmnet with reason: host reimage * 10:44 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2115.codfw.wmnet with reason: host reimage * 10:44 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2240.codfw.wmnet with OS trixie * 10:44 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2086.codfw.wmnet with reason: host reimage * 10:42 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2106.codfw.wmnet with reason: host reimage * 10:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2240: Upgrading db2240.codfw.wmnet * 10:41 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2240: Upgrading db2240.codfw.wmnet * 10:41 cwilliams@cumin1003: dbmaint on s4@codfw [[phab:T429893|T429893]] * 10:40 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 10:39 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2115.codfw.wmnet with reason: host reimage * 10:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2240 [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94653 and previous config saved to /var/cache/conftool/dbconfig/20260701-102658-cwilliams.json * 10:26 moritzm: installing nginx security updates * 10:26 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2086.codfw.wmnet with OS trixie * 10:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2179 to s4 primary [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94652 and previous config saved to /var/cache/conftool/dbconfig/20260701-102356-cwilliams.json * 10:23 cezmunsta: Starting s4 codfw failover from db2240 to db2179 - [[phab:T430127|T430127]] * 10:23 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2106.codfw.wmnet with OS trixie * 10:20 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2115.codfw.wmnet with OS trixie * 10:15 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2179 with weight 0 [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94651 and previous config saved to /var/cache/conftool/dbconfig/20260701-101531-cwilliams.json * 10:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 40 hosts with reason: Primary switchover s4 [[phab:T430127|T430127]] * 09:56 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template (take 2) - oblivian@cumin1003" * 09:56 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template (take 2) - oblivian@cumin1003 * 09:55 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template (take 2) - oblivian@cumin1003 * 09:55 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template (take 2) - oblivian@cumin1003" * 09:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:39 mszwarc@deploy1003: Synchronized private/SuggestedInvestigationsSignals/SuggestedInvestigationsSignal4n.php: Update SI signal 4n (duration: 06m 08s) * 09:21 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 09:21 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 09:14 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 09:14 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 09:02 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 09:02 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 08:54 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 08:38 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 08:38 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 08:36 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 08:21 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 08:21 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] (duration: 36m 11s) * 08:15 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 08:09 mszwarc@deploy1003: mszwarc, abi: Continuing with deployment * 08:03 mszwarc@deploy1003: mszwarc, abi: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:55 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 07:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 07:45 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] * 07:30 aqu@deploy1003: Finished deploy [analytics/refinery@410f205]: Regular analytics weekly train 2nd try [analytics/refinery@410f2050] (duration: 00m 22s) * 07:29 aqu@deploy1003: Started deploy [analytics/refinery@410f205]: Regular analytics weekly train 2nd try [analytics/refinery@410f2050] * 07:28 aqu@deploy1003: Finished deploy [analytics/refinery@410f205] (thin): Regular analytics weekly train THIN [analytics/refinery@410f2050] (duration: 01m 59s) * 07:28 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] (duration: 07m 19s) * 07:26 aqu@deploy1003: Started deploy [analytics/refinery@410f205] (thin): Regular analytics weekly train THIN [analytics/refinery@410f2050] * 07:26 aqu@deploy1003: Finished deploy [analytics/refinery@410f205]: Regular analytics weekly train [analytics/refinery@410f2050] (duration: 04m 32s) * 07:24 mszwarc@deploy1003: wmde-fisch, mszwarc: Continuing with deployment * 07:23 mszwarc@deploy1003: wmde-fisch, mszwarc: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:21 aqu@deploy1003: Started deploy [analytics/refinery@410f205]: Regular analytics weekly train [analytics/refinery@410f2050] * 07:21 aqu@deploy1003: Finished deploy [analytics/refinery@410f205] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@410f2050] (duration: 02m 01s) * 07:20 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] * 07:19 aqu@deploy1003: Started deploy [analytics/refinery@410f205] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@410f2050] * 07:13 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] (duration: 09m 13s) * 07:09 mszwarc@deploy1003: mszwarc, chlod, revi: Continuing with deployment * 07:06 mszwarc@deploy1003: mszwarc, chlod, revi: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:04 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] * 06:55 elukey: upgrade all trixie hosts to pywmflib 3.0 - [[phab:T430552|T430552]] * 06:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:43 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:43 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:42 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:42 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:41 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:41 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:35 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:35 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:34 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:34 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:31 jmm@cumin2003: DONE (PASS) - Cookbook sre.idm.logout (exit_code=0) Logging Niharika29 out of all services on: 2453 hosts * 06:30 oblivian@cumin1003: END (FAIL) - Cookbook sre.deploy.hiddenparma (exit_code=99) Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:30 oblivian@cumin1003: END (FAIL) - Cookbook sre.deploy.python-code (exit_code=99) hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:30 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:30 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:01 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2109.codfw.wmnet with OS trixie * 05:45 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on es1039.eqiad.wmnet with reason: issues * 05:41 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1027.eqiad.wmnet * 05:40 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2068.codfw.wmnet with OS trixie * 05:40 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2109.codfw.wmnet with reason: host reimage * 05:40 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1027.eqiad.wmnet,service=s2 * 05:40 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1027.eqiad.wmnet,service=s7 * 05:36 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2109.codfw.wmnet with reason: host reimage * 05:20 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2068.codfw.wmnet with reason: host reimage * 05:16 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2109.codfw.wmnet with OS trixie * 05:15 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2068.codfw.wmnet with reason: host reimage * 05:09 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2067.codfw.wmnet with OS trixie * 04:56 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2068.codfw.wmnet with OS trixie * 04:49 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2067.codfw.wmnet with reason: host reimage * 04:45 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2067.codfw.wmnet with reason: host reimage * 04:27 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2067.codfw.wmnet with OS trixie * 03:47 slyngshede@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1039.eqiad.wmnet with reason: Hardware crash * 03:21 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2107.codfw.wmnet with OS trixie * 02:59 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2107.codfw.wmnet with reason: host reimage * 02:55 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2085.codfw.wmnet with OS trixie * 02:51 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2072.codfw.wmnet with OS trixie * 02:51 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2107.codfw.wmnet with reason: host reimage * 02:35 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2085.codfw.wmnet with reason: host reimage * 02:31 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2107.codfw.wmnet with OS trixie * 02:30 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2072.codfw.wmnet with reason: host reimage * 02:26 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2085.codfw.wmnet with reason: host reimage * 02:22 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2072.codfw.wmnet with reason: host reimage * 02:09 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2085.codfw.wmnet with OS trixie * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 54s) * 02:03 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2072.codfw.wmnet with OS trixie * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es7 eqiad back to read-write - [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94649 and previous config saved to /var/cache/conftool/dbconfig/20260701-010716-ladsgroup.json * 01:05 ladsgroup@dns1004: END - running authdns-update * 01:05 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depool es1039 [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94648 and previous config saved to /var/cache/conftool/dbconfig/20260701-010551-ladsgroup.json * 01:03 ladsgroup@dns1004: START - running authdns-update * 01:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Promote es1035 to es7 primary [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94647 and previous config saved to /var/cache/conftool/dbconfig/20260701-010002-ladsgroup.json * 00:58 Amir1: Starting es7 eqiad failover from es1039 to es1035 - [[phab:T430765|T430765]] * 00:53 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es1035 with weight 0 [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94646 and previous config saved to /var/cache/conftool/dbconfig/20260701-005329-ladsgroup.json * 00:53 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 9 hosts with reason: Primary switchover es7 [[phab:T430765|T430765]] * 00:42 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es7 eqiad as read-only for maintenance - [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94645 and previous config saved to /var/cache/conftool/dbconfig/20260701-004221-ladsgroup.json * 00:20 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2102.codfw.wmnet with OS trixie * 00:15 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2103.codfw.wmnet with OS trixie * 00:05 dr0ptp4kt: DEPLOYED Refinery at {{Gerrit|4e7a2b32}} for changes: pageview allowlist {{Gerrit|1305158}} (+min.wikiquote) {{Gerrit|1305162}} (+bol.wikipedia), {{Gerrit|1305156}} (+isv.wikipedia); {{Gerrit|1305980}} (pv allowlist -api.wikimedia, sqoop +isvwiki); sqoop {{Gerrit|1295064}} (+globalimagelinks) {{Gerrit|1295069}} (+filerevision) using scap, then deployed onto HDFS (manual copyToLocal required additionally) == Other archives == See [[Server Admin Log/Archives]]. <noinclude> [[Category:SAL]] [[Category:Operations]] </noinclude> 7t6fkxzhh8vpodb7lo0lknukvh5v7wy 2450649 2450648 2026-08-22T16:33:26Z Stashbot 7414 arlolra@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply 2450649 wikitext text/x-wiki == 2026-08-22 == * 16:33 arlolra@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:32 arlolra@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 35s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-21 == * 20:36 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in1001.wikimedia.org with reason: [[phab:T434750|T434750]] * 20:34 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in2001.wikimedia.org with reason: [[phab:T434750|T434750]] * 20:33 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out1001.wikimedia.org with reason: [[phab:T434750|T434750]] * 20:25 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out2001.wikimedia.org with reason: [[phab:T434750|T434750]] * 19:37 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host krb1002.eqiad.wmnet with OS bookworm * 19:00 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 18:59 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 18:51 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 18:51 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 18:35 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:35 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:27 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:27 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:16 bking@cumin2003: START - Cookbook sre.hosts.reimage for host krb1002.eqiad.wmnet with OS bookworm * 17:35 sukhe@dns1004: END - running authdns-update * 17:33 sukhe@dns1004: START - running authdns-update * 17:32 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns5004.wikimedia.org [reason: resolved authdns-update issues] * 17:32 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:32 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: force HEAD to {{Gerrit|be26e30ae101}} - sukhe@cumin1003" * 17:32 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: force HEAD to {{Gerrit|be26e30ae101}} - sukhe@cumin1003" * 17:28 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 17:28 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: service=authdns-update,name=dns5004.wikimedia.org [reason: resolving authdns-update issues] * 17:28 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:28 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: force HEAD to {{Gerrit|be26e30ae101}} - sukhe@cumin1003" * 17:28 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: force HEAD to {{Gerrit|be26e30ae101}} - sukhe@cumin1003" * 17:24 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 17:24 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.netbox (exit_code=97) * 17:23 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 17:16 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 17:12 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 17:10 sukhe@dns1004: END - running authdns-update * 17:08 sukhe@dns1004: START - running authdns-update * 17:08 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=dns5004.wikimedia.org [reason: resolving authdns-update issues] * 17:07 sukhe@dns1004: FAIL - running authdns-update * 17:05 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 17:05 sukhe@dns1004: START - running authdns-update * 17:01 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 16:59 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 16:56 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 16:53 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=dns5004.* [reason: trixie upgrade] * 16:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns5004.wikimedia.org * 16:52 cdobbins@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns5004.wikimedia.org * 16:44 cmooney@dns3003: END - running authdns-update * 16:41 cmooney@dns3003: START - running authdns-update * 16:41 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:41 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on eqsin<->codfw arelion - cmooney@cumin1003" * 16:37 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on eqsin<->codfw arelion - cmooney@cumin1003" * 16:33 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:11 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 16:08 sukhe@dns1004: END - running authdns-update * 16:08 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 16:06 sukhe@dns1004: START - running authdns-update * 16:04 cmooney@dns3003: END - running authdns-update * 16:02 cmooney@dns3003: START - running authdns-update * 16:00 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:00 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on eqord<->codfw arelion - cmooney@cumin1003" * 15:56 cdobbins@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host dns5004.wikimedia.org with OS trixie * 15:55 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on eqord<->codfw arelion - cmooney@cumin1003" * 15:53 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:51 cmooney@cumin1003: END (ERROR) - Cookbook sre.dns.netbox (exit_code=97) * 15:51 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:38 andrew@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudcephosd1042.eqiad.wmnet with OS bookworm * 15:18 andrew@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudcephosd1042.eqiad.wmnet with reason: host reimage * 15:17 dancy@deploy1003: Finished deploy [gerrit/gerrit@2cc11cc]: Deploying https://gerrit.wikimedia.org/r/c/operations/software/gerrit/+/1327669 ([[phab:T434726|T434726]]) (duration: 00m 14s) * 15:17 dancy@deploy1003: Started deploy [gerrit/gerrit@2cc11cc]: Deploying https://gerrit.wikimedia.org/r/c/operations/software/gerrit/+/1327669 ([[phab:T434726|T434726]]) * 15:13 andrew@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudcephosd1042.eqiad.wmnet with reason: host reimage * 15:09 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns5004.wikimedia.org with reason: host reimage * 15:05 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns5004.wikimedia.org with reason: host reimage * 14:53 andrew@cumin2003: START - Cookbook sre.hosts.reimage for host cloudcephosd1042.eqiad.wmnet with OS bookworm * 14:30 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns5004.wikimedia.org with OS trixie * 14:29 cdobbins@cumin1003: conftool action : set/pooled=no; selector: name=dns5004.* [reason: trixie upgrade] * 14:21 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:21 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove entries for cr2-eqord - cmooney@cumin1003" * 14:21 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove entries for cr2-eqord - cmooney@cumin1003" * 14:13 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 14:11 moritzm: imported openjdk 8u504-ga-1~deb12u1 for bookworm-wikimedia (backport of the latest Java 8 security fixes for bookworm) * 13:25 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "sync cr2-eqord router offline - cmooney@cumin1003" * 13:23 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "sync cr2-eqord router offline - cmooney@cumin1003" * 13:14 hashar@deploy1003: Finished deploy [integration/docroot@2d5ff9b]: opensource: add PersonalDashboard docs to MW components - [[phab:T435392|T435392]] (duration: 00m 15s) * 13:14 hashar@deploy1003: Started deploy [integration/docroot@2d5ff9b]: opensource: add PersonalDashboard docs to MW components - [[phab:T435392|T435392]] * 12:16 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:16 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: [[phab:T431682|T431682]] - filippo@cumin1003" * 12:16 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: [[phab:T431682|T431682]] - filippo@cumin1003" * 12:11 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2006.wikimedia.org with OS trixie * 12:00 kevinbazira@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 11:58 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 11:43 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2006.wikimedia.org with reason: host reimage * 11:41 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 11:38 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2006.wikimedia.org with reason: host reimage * 11:20 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2006.wikimedia.org with OS trixie * 11:11 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2005.wikimedia.org with OS trixie * 10:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2005.wikimedia.org with reason: host reimage * 10:53 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2005.wikimedia.org with reason: host reimage * 10:43 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-codfw * 10:43 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2011.codfw.wmnet * 10:43 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2011.codfw.wmnet * 10:40 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2011.codfw.wmnet * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2011.codfw.wmnet * 10:34 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2010.codfw.wmnet * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2010.codfw.wmnet * 10:33 fnegri@cumin1003: END (PASS) - Cookbook sre.wikireplicas.add-wiki (exit_code=0) for database bolwiki ([[phab:T429954|T429954]]) * 10:33 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2005.wikimedia.org with OS trixie * 10:30 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2010.codfw.wmnet * 10:25 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2010.codfw.wmnet * 10:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2009.codfw.wmnet * 10:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2009.codfw.wmnet * 10:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1006.wikimedia.org with OS trixie * 10:18 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2009.codfw.wmnet * 10:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2009.codfw.wmnet * 10:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2008.codfw.wmnet * 10:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2008.codfw.wmnet * 10:06 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2008.codfw.wmnet * 10:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1006.wikimedia.org with reason: host reimage * 10:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2008.codfw.wmnet * 10:01 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2007.codfw.wmnet * 10:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2007.codfw.wmnet * 09:57 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1006.wikimedia.org with reason: host reimage * 09:56 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2007.codfw.wmnet * 09:51 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2007.codfw.wmnet * 09:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 09:51 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 09:46 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2006.codfw.wmnet * 09:46 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1006.wikimedia.org with OS trixie * 09:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1005.wikimedia.org with OS trixie * 09:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2006.codfw.wmnet * 09:41 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2005.codfw.wmnet * 09:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2005.codfw.wmnet * 09:38 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2013.codfw.wmnet * 09:36 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2005.codfw.wmnet * 09:35 fnegri@cumin1003: START - Cookbook sre.wikireplicas.add-wiki for database bolwiki ([[phab:T429954|T429954]]) * 09:35 fnegri@cumin1003: END (PASS) - Cookbook sre.wikireplicas.add-wiki (exit_code=0) for database minwikiquote ([[phab:T429946|T429946]]) * 09:35 fnegri@cumin1003: START - Cookbook sre.wikireplicas.add-wiki for database minwikiquote ([[phab:T429946|T429946]]) * 09:32 blake@cumin1003: START - Cookbook sre.hosts.reboot-single for host rdb2013.codfw.wmnet * 09:30 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2011.codfw.wmnet * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1005.wikimedia.org with reason: host reimage * 09:26 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2005.codfw.wmnet * 09:26 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2004.codfw.wmnet * 09:26 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2004.codfw.wmnet * 09:24 blake@cumin1003: START - Cookbook sre.hosts.reboot-single for host rdb2011.codfw.wmnet * 09:22 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1005.wikimedia.org with reason: host reimage * 09:21 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2004.codfw.wmnet * 09:16 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1015.eqiad.wmnet * 09:16 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2004.codfw.wmnet * 09:16 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2003.codfw.wmnet * 09:16 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2003.codfw.wmnet * 09:13 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on cr[1-2]-eqiad,pfw1-eqiad with reason: upgrade pfw1a-eqiad and pfw1b-eqiad pair * 09:12 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 09:11 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 09:11 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 09:11 blake@cumin1003: START - Cookbook sre.hosts.reboot-single for host rdb1015.eqiad.wmnet * 09:10 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2003.codfw.wmnet * 09:09 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1013.eqiad.wmnet * 09:07 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1005.wikimedia.org with OS trixie * 09:03 blake@cumin1003: START - Cookbook sre.hosts.reboot-single for host rdb1013.eqiad.wmnet * 09:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2003.codfw.wmnet * 09:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2002.codfw.wmnet * 09:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2002.codfw.wmnet * 08:54 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2002.codfw.wmnet * 08:49 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2002.codfw.wmnet * 08:49 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2001.codfw.wmnet * 08:49 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2001.codfw.wmnet * 08:48 jmm@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts netmon2002.wikimedia.org * 08:47 jmm@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts netmon2002.wikimedia.org * 08:44 jmm@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts netmon2002.wikimedia.org * 08:44 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon2002.wikimedia.org * 08:43 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2001.codfw.wmnet * 08:36 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon2002.wikimedia.org * 08:34 jmm@dns1004: END - running authdns-update * 08:33 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2001.codfw.wmnet * 08:33 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-codfw * 08:31 jmm@dns1004: START - running authdns-update * 07:48 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327679{{!}}Block: Disable flaky API test (T435272 T389028)]], [[gerrit:1327678{{!}}API: wfDebugLog for thumberror]] (duration: 15m 34s) * 07:41 krinkle@deploy1003: krinkle: Continuing with deployment * 07:37 krinkle@deploy1003: krinkle: Backport for [[gerrit:1327679{{!}}Block: Disable flaky API test (T435272 T389028)]], [[gerrit:1327678{{!}}API: wfDebugLog for thumberror]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:33 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1327679{{!}}Block: Disable flaky API test (T435272 T389028)]], [[gerrit:1327678{{!}}API: wfDebugLog for thumberror]] * 07:25 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 07:24 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 07:18 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 07:18 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 07:15 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 07:14 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 07:14 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 07:14 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 07:13 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 07:03 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1283: Pool back * 06:42 jmm@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts netmon2002.wikimedia.org * 06:35 moritzm: powercycling netmon2002 * 06:18 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1283: Pool back * 06:17 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1283 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96213 and previous config saved to /var/cache/conftool/dbconfig/20260821-061743-marostegui.json * 04:59 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324963{{!}}Add Produnto to extension-list (T421436)]], [[gerrit:1324964{{!}}Enable Produnto on Beta (T421436)]] (duration: 34m 48s) * 04:45 tstarling@deploy1003: tstarling: Continuing with deployment * 04:44 tstarling@deploy1003: tstarling: Backport for [[gerrit:1324963{{!}}Add Produnto to extension-list (T421436)]], [[gerrit:1324964{{!}}Enable Produnto on Beta (T421436)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 04:24 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1324963{{!}}Add Produnto to extension-list (T421436)]], [[gerrit:1324964{{!}}Enable Produnto on Beta (T421436)]] * 04:21 arlolra@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 04:20 arlolra@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 04:20 arlolra@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 04:20 arlolra@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 41s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-20 == * 23:43 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327652{{!}}RunSingleJob: Add ProfilingContext::init() (T435422)]] (duration: 12m 06s) * 23:38 krinkle@deploy1003: krinkle: Continuing with deployment * 23:33 krinkle@deploy1003: krinkle: Backport for [[gerrit:1327652{{!}}RunSingleJob: Add ProfilingContext::init() (T435422)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:31 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1327652{{!}}RunSingleJob: Add ProfilingContext::init() (T435422)]] * 22:15 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1054.eqiad.wmnet with OS trixie * 22:14 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 22:14 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 21:58 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1054.eqiad.wmnet with reason: host reimage * 21:51 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1054.eqiad.wmnet with reason: host reimage * 21:36 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1054.eqiad.wmnet with OS trixie * 21:36 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:35 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327219{{!}}RunSingleJob: Define MW_ENTRY_POINT for flamegraph sample attribution (T435422)]] (duration: 08m 30s) * 21:31 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:31 krinkle@deploy1003: krinkle: Continuing with deployment * 21:31 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1054 * 21:31 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1054 * 21:29 krinkle@deploy1003: krinkle: Backport for [[gerrit:1327219{{!}}RunSingleJob: Define MW_ENTRY_POINT for flamegraph sample attribution (T435422)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:27 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1327219{{!}}RunSingleJob: Define MW_ENTRY_POINT for flamegraph sample attribution (T435422)]] * 21:17 maryum: Deployed security fix for [[phab:T433020|T433020]] * 20:59 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324752{{!}}InitialiseSettings: Enable 2FA banners on remaining private wikis (T428103)]], [[gerrit:1325920{{!}}Remove sending email to legal team about rejected requests (T374053)]] (duration: 07m 18s) * 20:54 reedy@deploy1003: neriah, reedy: Continuing with deployment * 20:54 reedy@deploy1003: neriah, reedy: Backport for [[gerrit:1324752{{!}}InitialiseSettings: Enable 2FA banners on remaining private wikis (T428103)]], [[gerrit:1325920{{!}}Remove sending email to legal team about rejected requests (T374053)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:51 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324752{{!}}InitialiseSettings: Enable 2FA banners on remaining private wikis (T428103)]], [[gerrit:1325920{{!}}Remove sending email to legal team about rejected requests (T374053)]] * 20:24 reedy@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.15,1.47.0-wmf.16,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/med * 20:23 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324752{{!}}InitialiseSettings: Enable 2FA banners on remaining private wikis (T428103)]], [[gerrit:1325920{{!}}Remove sending email to legal team about rejected requests (T374053)]] * 20:14 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 20:10 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 20:09 cdanis@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "bug fixes & UX fixes - cdanis@cumin1003" * 20:09 cdanis@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: bug fixes & UX fixes - cdanis@cumin1003 * 20:08 cdanis@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: bug fixes & UX fixes - cdanis@cumin1003 * 20:08 cdanis@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "bug fixes & UX fixes - cdanis@cumin1003" * 19:24 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327598{{!}}Make \Omicron non upright (like \Chi) (T434428)]], [[gerrit:1327596{{!}}Render overline of \bar with stretchy=false (T435456)]] (duration: 18m 54s) * 19:20 krinkle@deploy1003: krinkle: Continuing with deployment * 19:07 krinkle@deploy1003: krinkle: Backport for [[gerrit:1327598{{!}}Make \Omicron non upright (like \Chi) (T434428)]], [[gerrit:1327596{{!}}Render overline of \bar with stretchy=false (T435456)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:05 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1327598{{!}}Make \Omicron non upright (like \Chi) (T434428)]], [[gerrit:1327596{{!}}Render overline of \bar with stretchy=false (T435456)]] * 18:51 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327614{{!}}Avoid casting fpxmax to string (T318419)]] (duration: 07m 28s) * 18:50 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:46 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:46 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 18:45 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327614{{!}}Avoid casting fpxmax to string (T318419)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:43 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327614{{!}}Avoid casting fpxmax to string (T318419)]] * 18:37 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:37 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:37 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:36 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:07 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 18:05 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 18:01 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 18:01 sukhe@dns1004: END - running authdns-update * 17:59 sukhe@dns1004: START - running authdns-update * 17:58 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 17:57 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=dns6002.* [reason: depooling for trixie upgrade] * 17:56 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns6002.wikimedia.org * 17:56 cdobbins@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns6002.wikimedia.org * 17:51 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 17:51 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 17:34 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host stat1011.eqiad.wmnet with OS bookworm * 17:31 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns6002.wikimedia.org with OS trixie * 17:30 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 17:30 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 17:30 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 17:30 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 17:29 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 17:29 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 17:29 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 17:29 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:29 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:27 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:24 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327590{{!}}Make sure fpsmax is an int value (T318419)]] (duration: 08m 37s) * 17:20 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 17:17 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327590{{!}}Make sure fpsmax is an int value (T318419)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:16 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 17:15 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327590{{!}}Make sure fpsmax is an int value (T318419)]] * 16:51 swfrench-wmf: disable-puppet on A:cp for ATS Lua change - [[phab:T427666|T427666]] * 16:51 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db2901.codfw.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 16:43 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 16:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on stat1011.eqiad.wmnet with reason: host reimage * 16:39 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns6002.wikimedia.org with reason: host reimage * 16:36 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on stat1011.eqiad.wmnet with reason: host reimage * 16:36 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db2901.codfw.wmnet * 16:34 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns6002.wikimedia.org with reason: host reimage * 16:15 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns6002.wikimedia.org with OS trixie * 16:14 cdobbins@cumin1003: conftool action : set/pooled=no; selector: name=dns6002.* [reason: depooling for trixie upgrade] * 16:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host stat1011 * 16:10 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host stat1011 * 16:09 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host stat1011 * 16:09 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) stat1011.eqiad.wmnet 14.36.64.10.in-addr.arpa 4.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 bking@cumin2003: START - Cookbook sre.dns.wipe-cache stat1011.eqiad.wmnet 14.36.64.10.in-addr.arpa 4.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:09 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host stat1011 - bking@cumin2003" * 16:09 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host stat1011 - bking@cumin2003" * 16:05 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host stat1011 * 16:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host stat1011.eqiad.wmnet with OS bookworm * 16:00 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 15:59 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 15:56 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 15:55 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 15:37 jayme@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:35 jayme@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 15:35 jayme@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:32 jayme@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:32 jayme@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:30 jayme@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 15:30 jayme@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:28 jayme@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:28 jayme@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 15:28 fceratto@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host db1903.eqiad.wmnet * 15:28 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1903.eqiad.wmnet with OS trixie * 15:26 jayme@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 15:26 jayme@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 15:24 jayme@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 15:24 jayme@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 15:21 jayme@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 15:21 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 15:19 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 15:19 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 15:17 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 15:14 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1903.eqiad.wmnet with reason: host reimage * 15:07 fceratto@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1903.eqiad.wmnet with reason: host reimage * 14:54 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db1903.eqiad.wmnet with OS trixie * 14:53 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1903.eqiad.wmnet - fceratto@cumin1003" * 14:53 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1903.eqiad.wmnet - fceratto@cumin1003" * 14:53 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1903.eqiad.wmnet on all recursors * 14:53 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1903.eqiad.wmnet on all recursors * 14:53 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:53 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1903.eqiad.wmnet - fceratto@cumin1003" * 14:53 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1903.eqiad.wmnet - fceratto@cumin1003" * 14:49 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 14:49 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1903.eqiad.wmnet * 14:33 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=cp1100.* * 14:27 topranks: reconfigure eqiad<->codfw bgp settings * 14:22 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:22 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update entries used on new transport backup eqiad codfw - cmooney@cumin1003" * 14:19 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update entries used on new transport backup eqiad codfw - cmooney@cumin1003" * 14:14 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 14:14 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 14:13 moritzm: installing util-linux security updates * 14:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-staging-worker * 14:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2003.codfw.wmnet * 14:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2003.codfw.wmnet * 14:08 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2003.codfw.wmnet * 14:06 moritzm: installing libheif security updates * 13:58 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2003.codfw.wmnet * 13:58 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2002.codfw.wmnet * 13:58 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2002.codfw.wmnet * 13:56 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:56 fnegri@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for clouddb1025.eqiad.wmnet * 13:56 fnegri@cumin1003: START - Cookbook sre.hosts.remove-downtime for clouddb1025.eqiad.wmnet * 13:56 Lucas_WMDE: UTC afternoon backport+config window done * 13:53 moritzm: installing apr-util security updates * 13:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2002.codfw.wmnet * 13:50 fnegri@cumin1003: conftool action : set/weight=100; selector: name=clouddb1025.eqiad.wmnet * 13:49 fnegri@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1025.eqiad.wmnet * 13:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2002.codfw.wmnet * 13:41 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2001.codfw.wmnet * 13:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2001.codfw.wmnet * 13:41 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'. * 13:38 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'. * 13:38 fnegri@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on clouddb1025.eqiad.wmnet with reason: Removing s6 from clouddb1025 * 13:34 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2001.codfw.wmnet * 13:31 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'. * 13:29 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'. * 13:28 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet * 13:26 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host stat1009.eqiad.wmnet with OS bookworm * 13:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2001.codfw.wmnet * 13:24 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-staging-worker * 13:23 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1002.eqiad.wmnet * 13:21 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host stat1010.eqiad.wmnet with OS bookworm * 13:20 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1002.eqiad.wmnet * 13:20 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1001.eqiad.wmnet * 13:17 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1001.eqiad.wmnet * 13:16 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2001.codfw.wmnet * 13:13 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327511{{!}}UIC: Fix page:page instead of page:other in instrumentation]] (duration: 07m 00s) * 13:13 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2001.codfw.wmnet * 13:12 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2002.codfw.wmnet * 13:09 mszwarc@deploy1003: mszwarc: Continuing with deployment * 13:08 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1327511{{!}}UIC: Fix page:page instead of page:other in instrumentation]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2002.codfw.wmnet * 13:07 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2002.codfw.wmnet * 13:06 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1327511{{!}}UIC: Fix page:page instead of page:other in instrumentation]] * 13:04 jmm@dns1004: END - running authdns-update * 13:03 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2002.codfw.wmnet * 13:03 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2001.codfw.wmnet * 13:02 jmm@dns1004: START - running authdns-update * 13:01 cmooney@dns3003: END - running authdns-update * 13:00 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2001.codfw.wmnet * 12:59 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2001.codfw.wmnet * 12:59 cmooney@dns3003: START - running authdns-update * 12:57 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2001.codfw.wmnet * 12:56 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2002.codfw.wmnet * 12:55 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:55 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on drmrs<->eqiad GTT vpls - cmooney@cumin1003" * 12:54 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on drmrs<->eqiad GTT vpls - cmooney@cumin1003" * 12:54 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2002.codfw.wmnet * 12:54 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2003.codfw.wmnet * 12:50 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2003.codfw.wmnet * 12:49 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 12:48 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2001.codfw.wmnet * 12:46 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2001.codfw.wmnet * 12:46 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2002.codfw.wmnet * 12:43 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2002.codfw.wmnet * 12:43 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2003.codfw.wmnet * 12:42 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=1) for new host db1902.eqiad.wmnet * 12:42 fceratto@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host db1902.eqiad.wmnet with OS trixie * 12:41 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2003.codfw.wmnet * 12:40 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1003.eqiad.wmnet * 12:38 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1003.eqiad.wmnet * 12:37 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1002.eqiad.wmnet * 12:35 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1002.eqiad.wmnet * 12:35 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1001.eqiad.wmnet * 12:34 cmooney@dns3003: END - running authdns-update * 12:33 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1001.eqiad.wmnet * 12:32 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on stat1009.eqiad.wmnet with reason: host reimage * 12:31 cmooney@dns3003: START - running authdns-update * 12:31 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:31 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on drmrs<->eqiad cct - cmooney@cumin1003" * 12:28 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on drmrs<->eqiad cct - cmooney@cumin1003" * 12:28 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1902.eqiad.wmnet with reason: host reimage * 12:25 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on stat1009.eqiad.wmnet with reason: host reimage * 12:25 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 12:24 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on stat1010.eqiad.wmnet with reason: host reimage * 12:22 fceratto@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1902.eqiad.wmnet with reason: host reimage * 12:21 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on stat1010.eqiad.wmnet with reason: host reimage * 12:14 elukey: move the Docker Registry's /v2/dev/.* prefix to its dedicated S3 backend - [[phab:T432829|T432829]] * 12:12 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db1902.eqiad.wmnet with OS trixie * 12:09 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1902.eqiad.wmnet - fceratto@cumin1003" * 12:09 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1902.eqiad.wmnet - fceratto@cumin1003" * 12:09 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1902.eqiad.wmnet on all recursors * 12:09 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1902.eqiad.wmnet on all recursors * 12:09 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:08 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1902.eqiad.wmnet - fceratto@cumin1003" * 12:08 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1902.eqiad.wmnet - fceratto@cumin1003" * 12:08 tgr_: [[phab:T413390|T413390]] running CentralAuth:FixRenamedUserGlobalEditCount --wiki=metawiki --since=20250901000000 --until=20260301000000 --fix * 12:04 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1009.eqiad.wmnet with OS bookworm * 12:04 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1010.eqiad.wmnet with OS bookworm * 12:01 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 12:01 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1902.eqiad.wmnet * 12:00 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host stat1010.eqiad.wmnet with OS bookworm * 11:57 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1902.eqiad.wmnet * 11:57 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:57 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1902.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 11:57 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1902.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 11:51 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'. * 11:49 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'. * 11:48 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'. * 11:46 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'. * 11:39 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 11:37 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327518{{!}}Enable thumb.wikimedia.org on cswiki and fawiki (T427465)]] (duration: 10m 40s) * 11:35 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1902.eqiad.wmnet * 11:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1010.eqiad.wmnet with OS bookworm * 11:33 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 11:30 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327518{{!}}Enable thumb.wikimedia.org on cswiki and fawiki (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:26 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327518{{!}}Enable thumb.wikimedia.org on cswiki and fawiki (T427465)]] * 11:24 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host stat1010.eqiad.wmnet with OS bookworm * 11:07 fceratto@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host db1901.eqiad.wmnet * 11:07 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1901.eqiad.wmnet with OS trixie * 10:53 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1901.eqiad.wmnet with reason: host reimage * 10:47 tappof: bump space for prometheus k8s-aux in codfw * 10:47 tappof: bump space for prometheus k8s-dse in eqiad * 10:47 fceratto@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1901.eqiad.wmnet with reason: host reimage * 10:35 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db1901.eqiad.wmnet with OS trixie * 10:32 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:32 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:32 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1901.eqiad.wmnet on all recursors * 10:32 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1901.eqiad.wmnet on all recursors * 10:31 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:31 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:31 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:27 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:27 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1901.eqiad.wmnet * 10:23 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1010.eqiad.wmnet with OS bookworm * 10:20 fceratto@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts db1901.eqiad.wmnet * 10:20 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 10:18 blake@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 10:17 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:16 blake@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 10:13 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1901.eqiad.wmnet * 09:23 jelto@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'. * 09:22 jelto@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'. * 09:22 jelto: update cert-manager to 1.19.6 on wikikube staging-eqiad - [[phab:T427402|T427402]] * 09:20 moritzm: imported squid 7.6-2.1for trixie-wikimedia/main [[phab:T427282|T427282]] * 09:08 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 09:08 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 09:08 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 09:07 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 09:04 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 09:04 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:23 slyngshede@dns1004: END - running authdns-update * 08:21 slyngshede@dns1004: START - running authdns-update * 08:18 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.16 refs [[phab:T430835|T430835]] * 06:27 aokoth@dns1004: END - running authdns-update * 06:25 aokoth@dns1004: START - running authdns-update * 06:22 brennen@deploy1003: Finished deploy [phabricator/deployment@6b9b6ff]: deploy phab1005 for [[phab:T435087|T435087]] (duration: 00m 39s) * 06:21 brennen@deploy1003: Started deploy [phabricator/deployment@6b9b6ff]: deploy phab1005 for [[phab:T435087|T435087]] * 06:20 brennen@deploy1003: Finished deploy [phabricator/deployment@6b9b6ff]: deploy phab1004 for to pick up config values for [[phab:T435087|T435087]] (duration: 01m 46s) * 06:18 brennen@deploy1003: Started deploy [phabricator/deployment@6b9b6ff]: deploy phab1004 for to pick up config values for [[phab:T435087|T435087]] * 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 49s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-19 == * 23:19 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327197{{!}}Enable thumb.wikimedia.org on mediawiki.org (T427465)]] (duration: 10m 50s) * 23:18 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:16 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 23:15 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 23:10 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327197{{!}}Enable thumb.wikimedia.org on mediawiki.org (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:10 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:09 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 23:09 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:09 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 23:08 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327197{{!}}Enable thumb.wikimedia.org on mediawiki.org (T427465)]] * 22:58 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327201{{!}}Enable ReadingLists for all logged in users on test wiki (T435258)]] (duration: 11m 20s) * 22:50 jdlrobson@deploy1003: jdlrobson: Continuing with deployment * 22:49 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1327201{{!}}Enable ReadingLists for all logged in users on test wiki (T435258)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:46 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1327201{{!}}Enable ReadingLists for all logged in users on test wiki (T435258)]] * 22:42 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327162{{!}}Article: Split subjectpageheader by model and disable for wikitext]], [[gerrit:1327169{{!}}Make uppercase greek letters normal (non-italic) font (T434686 T434428)]], [[gerrit:1327176{{!}}Skin: Avoid DB lookup for pagecategorieslink message (T347123)]] (duration: 37m 52s) * 22:29 krinkle@deploy1003: krinkle: Continuing with deployment * 22:25 krinkle@deploy1003: krinkle: Backport for [[gerrit:1327162{{!}}Article: Split subjectpageheader by model and disable for wikitext]], [[gerrit:1327169{{!}}Make uppercase greek letters normal (non-italic) font (T434686 T434428)]], [[gerrit:1327176{{!}}Skin: Avoid DB lookup for pagecategorieslink message (T347123)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:04 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1327162{{!}}Article: Split subjectpageheader by model and disable for wikitext]], [[gerrit:1327169{{!}}Make uppercase greek letters normal (non-italic) font (T434686 T434428)]], [[gerrit:1327176{{!}}Skin: Avoid DB lookup for pagecategorieslink message (T347123)]] * 22:04 krinkle@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: awaiting CI (duration: 03m 06s) * 22:01 krinkle@deploy1003: Locking from deployment [ALL REPOSITORIES]: awaiting CI * 22:00 krinkle@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: awaiting CI (duration: 00m 01s) * 22:00 krinkle@deploy1003: Locking from deployment [ALL REPOSITORIES]: awaiting CI * 21:34 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2001.codfw.wmnet * 21:28 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2001.codfw.wmnet * 21:22 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327178{{!}}AccountRecovery: Notify the email address of the on file of the request (T425799)]] (duration: 47m 02s) * 21:13 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm * 21:09 catrope@deploy1003: catrope: Continuing with deployment * 20:55 catrope@deploy1003: catrope: Backport for [[gerrit:1327178{{!}}AccountRecovery: Notify the email address of the on file of the request (T425799)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:35 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1327178{{!}}AccountRecovery: Notify the email address of the on file of the request (T425799)]] * 20:31 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327128{{!}}Parsoid DataAccess: convert Parsoid fragment markers to/from strip tags (T432547)]] (duration: 07m 30s) * 20:27 catrope@deploy1003: catrope, arlolra: Continuing with deployment * 20:26 catrope@deploy1003: catrope, arlolra: Backport for [[gerrit:1327128{{!}}Parsoid DataAccess: convert Parsoid fragment markers to/from strip tags (T432547)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:24 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1327128{{!}}Parsoid DataAccess: convert Parsoid fragment markers to/from strip tags (T432547)]] * 20:23 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage * 20:17 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage * 20:15 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325878{{!}}[arwiki] Enable restricted user page editing and grant edit permissions (T434878)]] (duration: 08m 46s) * 20:11 catrope@deploy1003: catrope, gergesshamon: Continuing with deployment * 20:08 catrope@deploy1003: catrope, gergesshamon: Backport for [[gerrit:1325878{{!}}[arwiki] Enable restricted user page editing and grant edit permissions (T434878)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:06 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1325878{{!}}[arwiki] Enable restricted user page editing and grant edit permissions (T434878)]] * 19:59 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm * 19:56 eevans@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cassandra-dev2001.codfw.wmnet with OS bookworm * 19:56 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm * 19:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2207.codfw.wmnet with reason: Maintenance * 18:47 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319920{{!}}Allow setting a separate thumbUrl in production (T427465)]], [[gerrit:1327167{{!}}Fix wmgThumbUrl config (T427465)]] (duration: 18m 53s) * 18:43 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 18:30 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1319920{{!}}Allow setting a separate thumbUrl in production (T427465)]], [[gerrit:1327167{{!}}Fix wmgThumbUrl config (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:28 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1319920{{!}}Allow setting a separate thumbUrl in production (T427465)]], [[gerrit:1327167{{!}}Fix wmgThumbUrl config (T427465)]] * 18:26 sukhe@dns1004: END - running authdns-update * 18:24 sukhe@dns1004: START - running authdns-update * 18:09 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1319920{{!}}Allow setting a separate thumbUrl in production (T427465)]] * 18:03 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-eqiad and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 17:56 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-codfw and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 17:50 cmooney@dns3003: END - running authdns-update * 17:42 dancy@deploy1003: Installation of scap version "4.283.0" completed for 3 hosts * 17:41 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling reboot on A:durum and A:durum * 17:40 cmooney@dns3003: START - running authdns-update * 17:40 dancy@deploy1003: Installing scap version "4.283.0" for 3 host(s) * 17:38 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:37 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on GTT VPLS - cmooney@cumin1003" * 17:37 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --verbose --scoreLessThan=0.7 --exceptDatasetChecksums=[[phab:T434319|T434319]]-enwiki-models.txt # [[phab:T434319|T434319]] * 17:32 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on GTT VPLS - cmooney@cumin1003" * 17:29 sbassett: Deployed security fix for [[phab:T435210|T435210]] (wmf.16) * 17:26 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 17:22 sbassett: Deployed security fix for [[phab:T435210|T435210]] (wmf.15) * 17:00 sukhe@dns1004: END - running authdns-update * 16:58 sukhe@dns1004: START - running authdns-update * 16:53 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica-esams and A:liberica * 16:41 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica-esams and A:liberica * 16:41 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-eqiad and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 16:41 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-codfw and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 16:41 cjd91: sudo -i cookbook sre.cdn.roll-upgrade-ats --query 'A:cp-codfw' --task-id [[phab:T434478|T434478]] --reason '9.2.15 upgrade' * 16:41 cjd91: sudo -i cookbook sre.cdn.roll-upgrade-ats --query 'A:cp-eqiad' --task-id [[phab:T434478|T434478]] --reason '9.2.15 upgrade' * 16:40 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and A:durum * 16:28 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326881{{!}}Echo: Start using virtual domains (T380385)]] (duration: 13m 12s) * 16:23 urbanecm@deploy1003: urbanecm: Continuing with deployment * 16:21 urandom: Completed sessionstore Cassandra/JVM upgrade — [[phab:T435154|T435154]] * 16:21 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching sessionstore[2005-2006].codfw.wmnet,sessionstore[1005-1006].eqiad.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 16:19 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1326881{{!}}Echo: Start using virtual domains (T380385)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:15 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:15 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->codfw - cmooney@cumin1003" * 16:14 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1326881{{!}}Echo: Start using virtual domains (T380385)]] * 16:14 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327123{{!}}Revert^2 "Migrate database access to virtual domains" (T435305)]], [[gerrit:1327124{{!}}Pass the mapped domain of virtual-echo-shared to the push NameTableStores (T435305)]] (duration: 07m 42s) * 16:13 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching sessionstore[2005-2006].codfw.wmnet,sessionstore[1005-1006].eqiad.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 16:11 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->codfw - cmooney@cumin1003" * 16:10 urbanecm@deploy1003: urbanecm: Continuing with deployment * 16:08 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1327123{{!}}Revert^2 "Migrate database access to virtual domains" (T435305)]], [[gerrit:1327124{{!}}Pass the mapped domain of virtual-echo-shared to the push NameTableStores (T435305)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:06 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 16:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 16:06 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1327123{{!}}Revert^2 "Migrate database access to virtual domains" (T435305)]], [[gerrit:1327124{{!}}Pass the mapped domain of virtual-echo-shared to the push NameTableStores (T435305)]] * 16:06 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:03 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching sessionstore1004.eqiad.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 16:01 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching sessionstore1004.eqiad.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 15:56 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching sessionstore2004.codfw.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 15:54 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching sessionstore2004.codfw.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 15:53 urandom: beginning sessionstore Cassandra/JVM upgrade — [[phab:T435154|T435154]] * 15:52 urandom: beginning sessionstore Cassandra/JVM upgrade — [[phab:T432944|T432944]] * 15:51 cmooney@dns3003: END - running authdns-update * 15:49 cmooney@dns3003: START - running authdns-update * 15:48 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:48 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->codfw - cmooney@cumin1003" * 15:45 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->codfw - cmooney@cumin1003" * 15:44 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 15:44 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:43 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 15:42 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:42 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:38 cmooney@dns3003: END - running authdns-update * 15:36 cmooney@dns3003: START - running authdns-update * 15:36 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:36 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->eqsin - cmooney@cumin1003" * 15:34 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327138{{!}}Enable redis lock manager everywhere (T366938)]] (duration: 08m 36s) * 15:33 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->eqsin - cmooney@cumin1003" * 15:30 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:29 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 15:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1147.eqiad.wmnet with OS bookworm * 15:28 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327138{{!}}Enable redis lock manager everywhere (T366938)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:25 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327138{{!}}Enable redis lock manager everywhere (T366938)]] * 15:24 jmm@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host krb1002.eqiad.wmnet * 15:19 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2207.codfw.wmnet with reason: Host crashed * 15:17 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327098{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]], [[gerrit:1327101{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]] (duration: 07m 13s) * 15:12 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 15:12 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1327098{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]], [[gerrit:1327101{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:10 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1327098{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]], [[gerrit:1327101{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]] * 15:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2010.codfw.wmnet with OS trixie * 15:05 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 15:05 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1147.eqiad.wmnet with reason: host reimage * 14:59 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 14:59 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:58 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1147.eqiad.wmnet with reason: host reimage * 14:55 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:55 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:49 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 14:48 cmooney@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host durum1001.eqiad.wmnet * 14:46 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:46 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Delete db2902 ipv6 addr - fceratto@cumin1003" * 14:46 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Delete db2902 ipv6 addr - fceratto@cumin1003" * 14:43 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1147.eqiad.wmnet with OS bookworm * 14:42 tgr@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327094{{!}}SpecialMWOAuthListConsumers: Handle newFromMWUser returning null in addNavigationSubtitle (T435167)]] (duration: 19m 25s) * 14:42 cmooney@cumin1003: START - Cookbook sre.hosts.reboot-single for host durum1001.eqiad.wmnet * 14:42 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 14:41 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 14:38 tgr@deploy1003: tgr: Continuing with deployment * 14:36 tgr@deploy1003: tgr: Backport for [[gerrit:1327094{{!}}SpecialMWOAuthListConsumers: Handle newFromMWUser returning null in addNavigationSubtitle (T435167)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:28 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 14:28 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:28 cmooney@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host durum3005.esams.wmnet * 14:25 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 14:23 cmooney@cumin1003: START - Cookbook sre.hosts.reboot-single for host durum3005.esams.wmnet * 14:23 tgr@deploy1003: Started scap sync-world: Backport for [[gerrit:1327094{{!}}SpecialMWOAuthListConsumers: Handle newFromMWUser returning null in addNavigationSubtitle (T435167)]] * 14:18 gengh@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:17 topranks: disable puppet on hosts running BIRD BGP to test merge of patch to systemd healtchcheck service * 14:17 gengh@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:17 gengh@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:16 gengh@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:16 elukey: upgrade spicerack on cumin1003 and cumin2003 * 14:16 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:15 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314962{{!}}static: add new dir bimi/ for BIMI SVG and PEM file (T311685)]] (duration: 10m 00s) * 14:15 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:11 kharlan@deploy1003: kharlan, sukhe: Continuing with deployment * 14:11 gengh@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:09 gengh@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:09 gengh@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:08 kharlan@deploy1003: kharlan, sukhe: Backport for [[gerrit:1314962{{!}}static: add new dir bimi/ for BIMI SVG and PEM file (T311685)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:06 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 14:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:06 gengh@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:05 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1314962{{!}}static: add new dir bimi/ for BIMI SVG and PEM file (T311685)]] * 14:05 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host krb1002.eqiad.wmnet * 14:05 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:04 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:04 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326342{{!}}Revert^2 "wmf-config/ProductionServices: set URL for urldownloader to service record"]] (duration: 07m 40s) * 13:59 kharlan@deploy1003: kharlan, sukhe: Continuing with deployment * 13:58 kharlan@deploy1003: kharlan, sukhe: Backport for [[gerrit:1326342{{!}}Revert^2 "wmf-config/ProductionServices: set URL for urldownloader to service record"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:57 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host stat1008.eqiad.wmnet with OS bookworm * 13:56 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1326342{{!}}Revert^2 "wmf-config/ProductionServices: set URL for urldownloader to service record"]] * 13:56 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 13:54 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325532{{!}}srwiki: Allow bureaucrats to add and remove event-organizer group (T434748)]] (duration: 14m 56s) * 13:54 swfrench@dns1004: END - running authdns-update * 13:52 swfrench@dns1004: START - running authdns-update * 13:51 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:51 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 13:48 kharlan@deploy1003: kharlan, danielyepezgarces: Continuing with deployment * 13:45 swfrench@cumin2003: conftool action : set/pooled=yes; selector: name=wikikube-worker2330.codfw.wmnet * 13:44 swfrench-wmf: finished etcd-main codfw -> eqiad switchover - [[phab:T435103|T435103]] * 13:44 kharlan@deploy1003: kharlan, danielyepezgarces: Backport for [[gerrit:1325532{{!}}srwiki: Allow bureaucrats to add and remove event-organizer group (T434748)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:44 swfrench@cumin2003: conftool action : set/pooled=no; selector: name=wikikube-worker2330.codfw.wmnet * 13:41 swfrench@dns1004: END - running authdns-update * 13:39 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1325532{{!}}srwiki: Allow bureaucrats to add and remove event-organizer group (T434748)]] * 13:39 swfrench@dns1004: START - running authdns-update * 13:37 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327111{{!}}Special:AbuseReview: Add "no further action needed" review action (T435020)]], [[gerrit:1327110{{!}}AbuseReview: Take the review verdict as a REST path parameter (T435020)]] (duration: 31m 43s) * 13:31 swfrench-wmf: starting etcd-main codfw -> eqiad switchover - [[phab:T435103|T435103]] * 13:28 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:28 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 13:25 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=97) rolling reboot on A:durum and A:durum * 13:24 kharlan@deploy1003: kharlan: Continuing with deployment * 13:24 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1147 * 13:24 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1147 * 13:23 kharlan@deploy1003: kharlan: Backport for [[gerrit:1327111{{!}}Special:AbuseReview: Add "no further action needed" review action (T435020)]], [[gerrit:1327110{{!}}AbuseReview: Take the review verdict as a REST path parameter (T435020)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:16 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:16 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 13:12 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:12 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 13:10 cdobbins@cumin1003: END (ERROR) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=97) Rolling upgrade of ATS on A:cp-codfw and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 13:10 cdobbins@cumin1003: END (ERROR) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=97) Rolling upgrade of ATS on A:cp-eqiad and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 13:06 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1327111{{!}}Special:AbuseReview: Add "no further action needed" review action (T435020)]], [[gerrit:1327110{{!}}AbuseReview: Take the review verdict as a REST path parameter (T435020)]] * 13:06 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-eqiad and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 13:05 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-codfw and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 13:05 cjd91: sudo -i cookbook sre.cdn.roll-upgrade-ats --query 'A:cp-eqiad' --task-id [[phab:T434478|T434478]] --reason '9.2.15 upgrade' * 13:03 swfrench@dns1004: END - running authdns-update * 13:01 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:01 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 13:01 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host ncmonitor1001.eqiad.wmnet * 13:00 swfrench@dns1004: START - running authdns-update * 12:59 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 12:59 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 12:59 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 12:59 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 12:57 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and A:durum * 12:53 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on db2902.codfw.wmnet with reason: Cloning * 12:48 cmooney@dns3003: END - running authdns-update * 12:48 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327096{{!}}Switch to redis lock manager on s4 and s8 (T366938)]] (duration: 09m 11s) * 12:46 cmooney@dns3003: START - running authdns-update * 12:45 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:45 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->eqord cct - cmooney@cumin1003" * 12:45 elukey: move the /v2/releng.* prefix on the Docker Registry to its new s3 backend - [[phab:T432829|T432829]] * 12:43 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 12:42 jelto@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 12:42 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->eqord cct - cmooney@cumin1003" * 12:42 jelto@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 12:41 jelto: update cert-manager to 1.19.6 on wikikube staging-codfw - [[phab:T427402|T427402]] * 12:40 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327096{{!}}Switch to redis lock manager on s4 and s8 (T366938)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:38 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327096{{!}}Switch to redis lock manager on s4 and s8 (T366938)]] * 12:38 blake@deploy1003: Finished scap sync-world: non-build deploy for [[phab:T417800|T417800]] (duration: 03m 52s) * 12:36 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 12:35 blake@deploy1003: Started scap sync-world: non-build deploy for [[phab:T417800|T417800]] * 12:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host krb2002.codfw.wmnet * 11:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host krb2002.codfw.wmnet * 11:49 moritzm: installing kerberos security updates * 11:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on stat1008.eqiad.wmnet with reason: host reimage * 11:44 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on stat1008.eqiad.wmnet with reason: host reimage * 11:31 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327084{{!}}Revert "Migrate database access to virtual domains" (T435305)]] (duration: 11m 02s) * 11:29 gkyziridis@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:29 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:29 kart_: Updated MinT to 2026-06-04-131507-production ([[phab:T321316|T321316]]) * 11:28 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/machinetranslation: apply * 11:28 gkyziridis@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:26 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 11:24 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:24 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:23 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:23 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:23 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/machinetranslation: apply * 11:22 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327084{{!}}Revert "Migrate database access to virtual domains" (T435305)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:21 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/machinetranslation: apply * 11:21 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:21 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:20 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327084{{!}}Revert "Migrate database access to virtual domains" (T435305)]] * 11:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1008.eqiad.wmnet with OS bookworm * 11:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet * 11:17 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/machinetranslation: apply * 11:13 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/machinetranslation: apply * 11:12 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:12 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:10 kartik@deploy1003: helmfile [staging] START helmfile.d/services/machinetranslation: apply * 11:08 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-cron: apply * 11:08 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/mw-cron: apply * 11:08 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply * 11:08 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply * 11:07 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet * 11:07 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host stat1008.eqiad.wmnet with OS bookworm * 11:06 moritzm: upgrading the new trixie URL downloaders to Squid 7.6 [[phab:T427282|T427282]] * 11:01 gkyziridis@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin2002.codfw.wmnet * 10:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin2002.codfw.wmnet * 10:45 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1169.eqiad.wmnet with OS bookworm * 10:42 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1185.eqiad.wmnet with OS bookworm * 10:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1169.eqiad.wmnet with reason: host reimage * 10:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1185.eqiad.wmnet with reason: host reimage * 10:14 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1169.eqiad.wmnet with reason: host reimage * 10:14 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1185.eqiad.wmnet with reason: host reimage * 10:11 jmm@cumin2003: END (PASS) - Cookbook sre.netbox.restart-reboot (exit_code=0) rolling reboot on A:netbox * 10:06 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1008.eqiad.wmnet with OS bookworm * 10:04 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 10:04 mpostoronca@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321579{{!}}Register the mediawiki.wikimedia_antiabuse.content_policy_score stream (T432848)]] (duration: 08m 53s) * 10:03 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 10:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 10:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 10:00 mpostoronca@deploy1003: mpostoronca: Continuing with deployment * 09:59 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1185.eqiad.wmnet with OS bookworm * 09:59 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1169.eqiad.wmnet with OS bookworm * 09:59 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.convert-disks (exit_code=0) for host ms-be1065 * 09:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1065.eqiad.wmnet with OS trixie * 09:58 mpostoronca@deploy1003: mpostoronca: Backport for [[gerrit:1321579{{!}}Register the mediawiki.wikimedia_antiabuse.content_policy_score stream (T432848)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:55 jmm@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netbox.discovery.wmnet. on all recursors * 09:55 mpostoronca@deploy1003: Started scap sync-world: Backport for [[gerrit:1321579{{!}}Register the mediawiki.wikimedia_antiabuse.content_policy_score stream (T432848)]] * 09:55 jmm@cumin2003: START - Cookbook sre.dns.wipe-cache netbox.discovery.wmnet. on all recursors * 09:52 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw2001.wikimedia.org with OS trixie * 09:51 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 09:51 jmm@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netbox.discovery.wmnet. on all recursors * 09:51 jmm@cumin2003: START - Cookbook sre.dns.wipe-cache netbox.discovery.wmnet. on all recursors * 09:51 jmm@cumin2003: START - Cookbook sre.netbox.restart-reboot rolling reboot on A:netbox * 09:46 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 09:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1056.eqiad.wmnet with OS trixie * 09:44 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 09:39 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 09:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 09:36 topranks: make HE transport circuits from magru live * 09:36 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 09:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb1003.eqiad.wmnet * 09:33 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage * 09:31 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb1003.eqiad.wmnet * 09:30 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 09:28 arnaudb@dns1006: END - running authdns-update * 09:27 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage * 09:27 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb2003.codfw.wmnet * 09:26 arnaudb@dns1006: START - running authdns-update * 09:24 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1056.eqiad.wmnet with reason: host reimage * 09:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb2003.codfw.wmnet * 09:20 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.convert-disks (exit_code=0) for host ms-be1068 * 09:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1068.eqiad.wmnet with OS trixie * 09:20 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 09:19 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "cloudvirt1057 - filippo@cumin1003" * 09:19 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "cloudvirt1057 - filippo@cumin1003" * 09:18 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1056.eqiad.wmnet with reason: host reimage * 09:18 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1057.eqiad.wmnet with OS trixie * 09:18 filippo@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 09:18 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 09:15 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 09:14 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 09:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host irc1003.wikimedia.org * 09:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:13 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1065.eqiad.wmnet with OS trixie * 09:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 09:09 moritzm: installing Postgresql security updates * 09:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host irc1003.wikimedia.org * 09:07 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 09:07 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:07 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1173.eqiad.wmnet with OS bookworm * 09:06 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw2001.wikimedia.org with OS trixie * 09:03 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.convert-disks (exit_code=0) for host ms-be1064 * 09:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1064.eqiad.wmnet with OS trixie * 09:03 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1056.eqiad.wmnet with OS trixie * 09:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1057.eqiad.wmnet with reason: host reimage * 09:01 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 08:58 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 08:56 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1057.eqiad.wmnet with reason: host reimage * 08:54 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1208.eqiad.wmnet with OS bookworm * 08:53 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2902.codfw.wmnet with OS trixie * 08:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 08:50 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1174.eqiad.wmnet with OS bookworm * 08:46 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1172.eqiad.wmnet with OS bookworm * 08:45 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1173.eqiad.wmnet with reason: host reimage * 08:41 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 08:40 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1057.eqiad.wmnet with OS trixie * 08:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1057.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 08:39 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1207.eqiad.wmnet with OS bookworm * 08:38 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2902.codfw.wmnet with reason: host reimage * 08:37 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 08:34 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1068.eqiad.wmnet with OS trixie * 08:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1208.eqiad.wmnet with reason: host reimage * 08:31 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1057.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 08:29 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 08:28 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1222.eqiad.wmnet onto db1276.eqiad.wmnet * 08:28 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1222: Pool db1222.eqiad.wmnet in after cloning * 08:28 fceratto@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2902.codfw.wmnet with reason: host reimage * 08:27 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1055.eqiad.wmnet with OS trixie * 08:27 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 08:26 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1174.eqiad.wmnet with reason: host reimage * 08:25 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 08:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-misc2002.codfw.wmnet * 08:23 topranks: reboot pfw1-codfw firewall pair to upgrade JunOS [[phab:T434865|T434865]] * 08:22 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1172.eqiad.wmnet with reason: host reimage * 08:20 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1064.eqiad.wmnet with OS trixie * 08:20 mvernon@cumin2003: START - Cookbook sre.swift.convert-disks for host ms-be1065 * 08:18 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1207.eqiad.wmnet with reason: host reimage * 08:17 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1174.eqiad.wmnet with reason: host reimage * 08:17 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1173.eqiad.wmnet with reason: host reimage * 08:17 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1172.eqiad.wmnet with reason: host reimage * 08:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host mc-misc2002.codfw.wmnet * 08:15 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1208.eqiad.wmnet with reason: host reimage * 08:15 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1207.eqiad.wmnet with reason: host reimage * 08:14 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db2902.codfw.wmnet with OS trixie * 08:14 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.16 refs [[phab:T430835|T430835]] * 08:13 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db2902.codfw.wmnet * 08:13 fceratto@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host db2902.codfw.wmnet with OS trixie * 08:10 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr[1-2]-codfw with reason: upgrade pfw1a-codfw and pfw1b-codfw pair * 08:09 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1055.eqiad.wmnet with reason: host reimage * 08:07 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on pfw1-codfw with reason: upgrade pfw1a-codfw and pfw1b-codfw pair * 08:03 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1055.eqiad.wmnet with reason: host reimage * 08:02 arnaudb@dns1006: END - running authdns-update * 08:02 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1208.eqiad.wmnet with OS bookworm * 08:02 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1207.eqiad.wmnet with OS bookworm * 08:01 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1174.eqiad.wmnet with OS bookworm * 08:01 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1173.eqiad.wmnet with OS bookworm * 08:01 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1172.eqiad.wmnet with OS bookworm * 07:59 arnaudb@dns1006: START - running authdns-update * 07:58 arnaudb@dns1006: START - running authdns-update * 07:48 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1055.eqiad.wmnet with OS trixie * 07:45 moritzm: extend the disk of ldap-rw2001 by 80G [[phab:T331699|T331699]] * 07:42 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1222: Pool db1222.eqiad.wmnet in after cloning * 07:36 mvernon@cumin2003: START - Cookbook sre.swift.convert-disks for host ms-be1068 * 07:35 mvernon@cumin2003: START - Cookbook sre.swift.convert-disks for host ms-be1064 * 07:17 moritzm: installing imagemagick security updates * 07:14 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1277: Pool back * 07:14 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon1003.wikimedia.org * 07:07 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon1003.wikimedia.org * 07:03 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1280: Pool back * 07:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon2002.wikimedia.org * 06:55 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon2002.wikimedia.org * 06:54 moritzm: installing php8.2 security updates * 06:51 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1284: Pool back * 06:49 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1222: Depool db1222.eqiad.wmnet to then clone it to db1276.eqiad.wmnet - marostegui@cumin1003 * 06:49 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1222: Depool db1222.eqiad.wmnet to then clone it to db1276.eqiad.wmnet - marostegui@cumin1003 * 06:49 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1222.eqiad.wmnet onto db1276.eqiad.wmnet * 06:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd1005.eqiad.wmnet * 06:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd1005.eqiad.wmnet * 06:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd1004.eqiad.wmnet * 06:36 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2209: db2209 repool * 06:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd1004.eqiad.wmnet * 06:32 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast3007.wikimedia.org * 06:29 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1277: Pool back * 06:28 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1277 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96190 and previous config saved to /var/cache/conftool/dbconfig/20260819-062815-marostegui.json * 06:26 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast3007.wikimedia.org * 06:22 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul1001.eqiad.wmnet * 06:18 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul1001.eqiad.wmnet * 06:18 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul1003.eqiad.wmnet * 06:18 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1280: Pool back * 06:17 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1284 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96186 and previous config saved to /var/cache/conftool/dbconfig/20260819-061743-marostegui.json * 06:14 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul1003.eqiad.wmnet * 06:14 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul1002.eqiad.wmnet * 06:10 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul1002.eqiad.wmnet * 06:10 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2048.codfw.wmnet * 06:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2048.codfw.wmnet * 06:06 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1284: Pool back * 06:06 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1284 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96184 and previous config saved to /var/cache/conftool/dbconfig/20260819-060621-marostegui.json * 06:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2048.codfw.wmnet * 05:59 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2048.codfw.wmnet * 05:51 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2209: db2209 repool * 03:16 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1186.eqiad.wmnet with OS bookworm * 02:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1186.eqiad.wmnet with reason: host reimage * 02:46 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1186.eqiad.wmnet with reason: host reimage * 02:46 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2207 [[phab:T435270|T435270]]', diff saved to https://phabricator.wikimedia.org/P96181 and previous config saved to /var/cache/conftool/dbconfig/20260819-024627-marostegui.json * 02:44 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2204 to s2 primary [[phab:T435270|T435270]]', diff saved to https://phabricator.wikimedia.org/P96180 and previous config saved to /var/cache/conftool/dbconfig/20260819-024403-marostegui.json * 02:43 marostegui: Starting s2 codfw failover from db2207 to db2204 - [[phab:T435270|T435270]] * 02:39 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2204 with weight 0 [[phab:T435270|T435270]]', diff saved to https://phabricator.wikimedia.org/P96179 and previous config saved to /var/cache/conftool/dbconfig/20260819-023951-marostegui.json * 02:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s2 [[phab:T435270|T435270]] * 02:32 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1186.eqiad.wmnet with OS bookworm * 02:29 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-worker1186.eqiad.wmnet with OS bookworm * 02:18 denisse@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2207: Depooling replica * 02:18 denisse@cumin1003: START - Cookbook sre.mysql.depool depool db2207: Depooling replica * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 48s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-18 == * 23:55 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1264845{{!}}Remove unused/redundant wgMFNoindexPages=true setting (T255458)]] (duration: 09m 42s) * 23:51 krinkle@deploy1003: krinkle: Continuing with deployment * 23:48 krinkle@deploy1003: krinkle: Backport for [[gerrit:1264845{{!}}Remove unused/redundant wgMFNoindexPages=true setting (T255458)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:45 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1264845{{!}}Remove unused/redundant wgMFNoindexPages=true setting (T255458)]] * 23:38 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326944{{!}}Retire filebackend lock manager in favour of the default one (T366938)]] (duration: 08m 55s) * 23:34 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 23:31 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326944{{!}}Retire filebackend lock manager in favour of the default one (T366938)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:29 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326944{{!}}Retire filebackend lock manager in favour of the default one (T366938)]] * 23:27 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1170.eqiad.wmnet with OS bookworm * 23:21 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1205.eqiad.wmnet with OS bookworm * 23:20 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1171.eqiad.wmnet with OS bookworm * 23:15 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1206.eqiad.wmnet with OS bookworm * 23:05 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1170.eqiad.wmnet with reason: host reimage * 23:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1205.eqiad.wmnet with reason: host reimage * 22:57 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1171.eqiad.wmnet with reason: host reimage * 22:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1206.eqiad.wmnet with reason: host reimage * 22:53 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1205.eqiad.wmnet with reason: host reimage * 22:51 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1171.eqiad.wmnet with reason: host reimage * 22:51 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1170.eqiad.wmnet with reason: host reimage * 22:50 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1206.eqiad.wmnet with reason: host reimage * 22:36 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1206.eqiad.wmnet with OS bookworm * 22:35 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1205.eqiad.wmnet with OS bookworm * 22:35 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1186.eqiad.wmnet with OS bookworm * 22:35 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1171.eqiad.wmnet with OS bookworm * 22:35 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1170.eqiad.wmnet with OS bookworm * 22:33 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-worker1194.eqiad.wmnet with OS bookworm * 22:22 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326923{{!}}Enable redis lock manager on s6 (T366938)]] (duration: 11m 52s) * 22:18 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 22:13 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326923{{!}}Enable redis lock manager on s6 (T366938)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:10 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326923{{!}}Enable redis lock manager on s6 (T366938)]] * 22:04 sbassett: Deployed security fix for [[phab:T435234|T435234]] (wmf.16) * 21:54 sbassett: Deployed security fix for [[phab:T435234|T435234]] (wmf.15) * 21:38 caro@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326925{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326926{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326929{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]], [[gerrit:1326928{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]] (duration: 0 * 21:34 caro@deploy1003: caro: Continuing with deployment * 21:33 caro@deploy1003: caro: Backport for [[gerrit:1326925{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326926{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326929{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]], [[gerrit:1326928{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]] synced to the testservers (see h * 21:31 caro@deploy1003: Started scap sync-world: Backport for [[gerrit:1326925{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326926{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326929{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]], [[gerrit:1326928{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]] * 21:24 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1204.eqiad.wmnet with reason: 1204 datanode repair [[phab:T434494|T434494]] * 21:02 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326896{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]], [[gerrit:1326897{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]] (duration: 13m 56s) * 20:58 krinkle@deploy1003: krinkle: Continuing with deployment * 20:50 krinkle@deploy1003: krinkle: Backport for [[gerrit:1326896{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]], [[gerrit:1326897{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:49 ryankemper: `an-launcher1003` terminated process group `666809` (`rest_backfill_phase1.sh`) ~20 mins ago with `sudo kill -TERM -- -666809` after its local spark driver (`--driver-memory 64g`) repeatedly exhausted memory on the 32 GB VM and caused SSH to intermittently flap; host recovered to 27 GB available memory * 20:48 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1326896{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]], [[gerrit:1326897{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]] * 20:35 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 20:33 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 20:31 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 20:31 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326870{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]], [[gerrit:1326871{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]] (duration: 07m 35s) * 20:28 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 20:26 kemayo@deploy1003: kemayo: Continuing with deployment * 20:26 ryankemper: `an-launcher1003` confirmed the host is flapping because of memory thrash. chasing down the source of the thrash * 20:25 kemayo@deploy1003: kemayo: Backport for [[gerrit:1326870{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]], [[gerrit:1326871{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:23 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1326870{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]], [[gerrit:1326871{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]] * 20:23 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 20:20 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 20:14 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1168.eqiad.wmnet with OS bookworm * 20:08 zabe: zabe@deploy1003:~$ mwscript extensions/WikimediaMaintenance/maintenance/fixFileRevisionArchiveNameDrift.php enwiki # [[phab:T428406|T428406]] * 20:08 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1204.eqiad.wmnet with OS bookworm * 20:05 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1167.eqiad.wmnet with OS bookworm * 20:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1166.eqiad.wmnet with OS bookworm * 19:54 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1203.eqiad.wmnet with OS bookworm * 19:53 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326907{{!}}Revert "Disable redis lock manager on testwiki"]] (duration: 11m 05s) * 19:50 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1168.eqiad.wmnet with reason: host reimage * 19:47 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1204.eqiad.wmnet with reason: host reimage * 19:46 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 19:44 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326907{{!}}Revert "Disable redis lock manager on testwiki"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:42 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326907{{!}}Revert "Disable redis lock manager on testwiki"]] * 19:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1167.eqiad.wmnet with reason: host reimage * 19:37 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1166.eqiad.wmnet with reason: host reimage * 19:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1203.eqiad.wmnet with reason: host reimage * 19:32 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1167.eqiad.wmnet with reason: host reimage * 19:32 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1168.eqiad.wmnet with reason: host reimage * 19:32 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1166.eqiad.wmnet with reason: host reimage * 19:31 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1204.eqiad.wmnet with reason: host reimage * 19:31 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1203.eqiad.wmnet with reason: host reimage * 19:26 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324751{{!}}InitialiseSettings: Enable 2FA enforcement on more private wikis (T428103)]], [[gerrit:1326875{{!}}Add banner notifying of upcoming 2FA enforcement (T420792)]] (duration: 31m 46s) * 19:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1204.eqiad.wmnet with OS bookworm * 19:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1203.eqiad.wmnet with OS bookworm * 19:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1168.eqiad.wmnet with OS bookworm * 19:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1167.eqiad.wmnet with OS bookworm * 19:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1166.eqiad.wmnet with OS bookworm * 19:15 denisse: rebooting kafkamon2003.codfw.wmnet - [[phab:T435162|T435162]] * 19:14 denisse: rebooting kafkamon1003.eqiad.wmnet [[phab:T435162|T435162]] * 19:13 reedy@deploy1003: reedy: Continuing with deployment * 19:12 reedy@deploy1003: reedy: Backport for [[gerrit:1324751{{!}}InitialiseSettings: Enable 2FA enforcement on more private wikis (T428103)]], [[gerrit:1326875{{!}}Add banner notifying of upcoming 2FA enforcement (T420792)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:54 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324751{{!}}InitialiseSettings: Enable 2FA enforcement on more private wikis (T428103)]], [[gerrit:1326875{{!}}Add banner notifying of upcoming 2FA enforcement (T420792)]] * 18:50 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 18:44 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.16 refs [[phab:T430835|T430835]] * 18:34 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 18:31 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 18:22 aklapper@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326852{{!}}CategoryTree: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]], [[gerrit:1326853{{!}}CategoryViewer: Allow null $html in the CategoryViewerGenerateLink hook (T435161)]], [[gerrit:1326865{{!}}Flow: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]] (duration: 09m 57s) * 18:18 aklapper@deploy1003: jforrester, aklapper: Continuing with deployment * 18:17 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 18:14 aklapper@deploy1003: jforrester, aklapper: Backport for [[gerrit:1326852{{!}}CategoryTree: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]], [[gerrit:1326853{{!}}CategoryViewer: Allow null $html in the CategoryViewerGenerateLink hook (T435161)]], [[gerrit:1326865{{!}}Flow: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki * 18:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1156.eqiad.wmnet with OS bookworm * 18:12 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 18:12 aklapper@deploy1003: Started scap sync-world: Backport for [[gerrit:1326852{{!}}CategoryTree: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]], [[gerrit:1326853{{!}}CategoryViewer: Allow null $html in the CategoryViewerGenerateLink hook (T435161)]], [[gerrit:1326865{{!}}Flow: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]] * 18:11 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 18:08 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1146.eqiad.wmnet with OS bookworm * 18:07 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1177.eqiad.wmnet with OS bookworm * 18:00 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326380{{!}}Introduce main lock manager service (T366938 T427999)]] (duration: 11m 25s) * 17:58 ladsgroup@deploy1003: ladsgroup: Rolling back deployment * 17:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1202.eqiad.wmnet with OS bookworm * 17:55 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1201.eqiad.wmnet with OS bookworm * 17:50 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326380{{!}}Introduce main lock manager service (T366938 T427999)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:48 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326380{{!}}Introduce main lock manager service (T366938 T427999)]] * 17:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1156.eqiad.wmnet with reason: host reimage * 17:46 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1146.eqiad.wmnet with reason: host reimage * 17:45 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-esams and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 17:42 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 17:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1177.eqiad.wmnet with reason: host reimage * 17:38 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1202.eqiad.wmnet with reason: host reimage * 17:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1201.eqiad.wmnet with reason: host reimage * 17:30 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1156.eqiad.wmnet with reason: host reimage * 17:29 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1177.eqiad.wmnet with reason: host reimage * 17:28 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1146.eqiad.wmnet with reason: host reimage * 17:28 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1202.eqiad.wmnet with reason: host reimage * 17:27 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1201.eqiad.wmnet with reason: host reimage * 17:25 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 17:21 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 17:14 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1202.eqiad.wmnet with OS bookworm * 17:14 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1201.eqiad.wmnet with OS bookworm * 17:14 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1177.eqiad.wmnet with OS bookworm * 17:14 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1156.eqiad.wmnet with OS bookworm * 17:14 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1146.eqiad.wmnet with OS bookworm * 17:09 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326885{{!}}w/deployment-info.php: Handle new file format (T434726)]] (duration: 07m 15s) * 17:08 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 17:05 dancy@deploy1003: dancy: Continuing with deployment * 17:04 dancy@deploy1003: dancy: Backport for [[gerrit:1326885{{!}}w/deployment-info.php: Handle new file format (T434726)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:02 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1326885{{!}}w/deployment-info.php: Handle new file format (T434726)]] * 16:50 dancy@deploy1003: Finished scap sync-world: Testing [[phab:T434726|T434726]] (duration: 06m 40s) * 16:43 dancy@deploy1003: Started scap sync-world: Testing [[phab:T434726|T434726]] * 16:43 dancy@deploy1003: Installation of scap version "4.282.0" completed for 3 hosts * 16:41 dancy@deploy1003: Installing scap version "4.282.0" for 3 host(s) * 16:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1145.eqiad.wmnet with OS bookworm * 16:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1200.eqiad.wmnet with OS bookworm * 16:32 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1199.eqiad.wmnet with OS bookworm * 16:18 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1145.eqiad.wmnet with reason: host reimage * 16:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1200.eqiad.wmnet with reason: host reimage * 16:09 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1199.eqiad.wmnet with reason: host reimage * 16:05 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-esams and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 16:05 cjd91: sudo -i cookbook sre.cdn.roll-upgrade-ats --query 'A:cp-esams' --task-id [[phab:T434478|T434478]] --reason '9.2.15 upgrade' * 16:03 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1200.eqiad.wmnet with reason: host reimage * 16:02 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1145.eqiad.wmnet with reason: host reimage * 16:02 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1199.eqiad.wmnet with reason: host reimage * 15:50 moritzm: installing zip security updates * 15:48 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1200.eqiad.wmnet with OS bookworm * 15:47 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1199.eqiad.wmnet with OS bookworm * 15:47 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1145.eqiad.wmnet with OS bookworm * 15:41 topranks: bounce PIC 0/0 on cr1-magru to set port to 40G * 15:37 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1236.eqiad.wmnet with OS bookworm * 15:29 aikochou@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'ores-legacy' for release 'main' . * 15:26 aikochou@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'ores-legacy' for release 'main' . * 15:20 aikochou@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'ores-legacy' for release 'main' . * 15:14 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2002.codfw.wmnet * 15:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1236.eqiad.wmnet with reason: host reimage * 15:12 moritzm: failover ganeti master in codfw to ganeti2047 * 15:09 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1236.eqiad.wmnet with reason: host reimage * 15:09 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2044.codfw.wmnet * 15:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2002.codfw.wmnet * 15:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2044.codfw.wmnet * 15:04 brennen@deploy1003: Finished deploy [phabricator/deployment@6b9b6ff]: deploy phab1004 for [[phab:T435213|T435213]] (duration: 01m 01s) * 15:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2044.codfw.wmnet * 15:03 brennen@deploy1003: Started deploy [phabricator/deployment@6b9b6ff]: deploy phab1004 for [[phab:T435213|T435213]] * 15:03 brennen@deploy1003: Finished deploy [phabricator/deployment@6b9b6ff]: deploy phab2003 for [[phab:T435213|T435213]] (duration: 00m 57s) * 15:02 brennen@deploy1003: Started deploy [phabricator/deployment@6b9b6ff]: deploy phab2003 for [[phab:T435213|T435213]] * 14:57 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply * 14:57 arnaudb@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on phab2003.codfw.wmnet,phab[1004-1006].eqiad.wmnet with reason: maintenance * 14:56 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2044.codfw.wmnet * 14:55 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply * 14:53 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1236.eqiad.wmnet with OS bookworm * 14:51 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2043.codfw.wmnet * 14:51 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2043.codfw.wmnet * 14:45 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2043.codfw.wmnet * 14:33 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2043.codfw.wmnet * 14:22 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2042.codfw.wmnet * 14:22 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2042.codfw.wmnet * 14:21 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db2902.codfw.wmnet with OS trixie * 14:16 elukey: uploaded spicerack_13.2.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia * 14:16 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2042.codfw.wmnet * 14:04 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2042.codfw.wmnet * 14:04 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2041.codfw.wmnet * 14:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2041.codfw.wmnet * 13:59 phuedx@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: apply * 13:59 phuedx@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-main: apply * 13:59 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326838{{!}}Enable Suggested Investigations on hewiki (T435146)]] (duration: 11m 50s) * 13:58 phuedx@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: apply * 13:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2041.codfw.wmnet * 13:57 phuedx@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-main: apply * 13:57 phuedx@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-main: apply * 13:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard1003.eqiad.wmnet * 13:57 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-main: apply * 13:55 phuedx@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-logging-external: apply * 13:55 phuedx@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-logging-external: apply * 13:54 phuedx@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-logging-external: apply * 13:54 stran@deploy1003: stran: Continuing with deployment * 13:54 phuedx@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-logging-external: apply * 13:54 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-logging-external: apply * 13:54 phuedx@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-logging-external: apply * 13:53 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-logging-external: apply * 13:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard1003.eqiad.wmnet * 13:53 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2041.codfw.wmnet * 13:51 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db2902.codfw.wmnet - fceratto@cumin1003" * 13:51 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db2902.codfw.wmnet - fceratto@cumin1003" * 13:51 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2026.codfw.wmnet * 13:51 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard2003.codfw.wmnet * 13:50 phuedx@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: apply * 13:50 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-eqiad * 13:50 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp1001.eqiad.wmnet * 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp1001.eqiad.wmnet * 13:50 phuedx@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: apply * 13:49 phuedx@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: apply * 13:49 stran@deploy1003: stran: Backport for [[gerrit:1326838{{!}}Enable Suggested Investigations on hewiki (T435146)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:48 phuedx@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: apply * 13:48 phuedx@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics: apply * 13:47 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard2003.codfw.wmnet * 13:47 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics: apply * 13:47 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1326838{{!}}Enable Suggested Investigations on hewiki (T435146)]] * 13:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp1001.eqiad.wmnet * 13:44 cdanis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 13:43 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp1001.eqiad.wmnet * 13:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1315-1327].eqiad.wmnet * 13:43 cdanis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 13:43 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1315-1327].eqiad.wmnet * 13:40 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki1001.eqiad.wmnet * 13:35 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1315-1327].eqiad.wmnet * 13:34 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host rpki1001.eqiad.wmnet * 13:27 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1315-1327].eqiad.wmnet * 13:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1302-1314].eqiad.wmnet * 13:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1302-1314].eqiad.wmnet * 13:21 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326775{{!}}SI: Instrument case update on first edit (T435048)]] (duration: 07m 12s) * 13:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1302-1314].eqiad.wmnet * 13:17 stran@deploy1003: stran: Continuing with deployment * 13:16 stran@deploy1003: stran: Backport for [[gerrit:1326775{{!}}SI: Instrument case update on first edit (T435048)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:14 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1326775{{!}}SI: Instrument case update on first edit (T435048)]] * 13:13 moritzm: installing util-linux security updates * 13:11 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1302-1314].eqiad.wmnet * 13:10 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326232{{!}}prv: Enable parsoid rendering for 5 wikis (T435115)]] (duration: 08m 17s) * 13:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1288-1289,1291-1301].eqiad.wmnet * 13:10 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply * 13:10 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1288-1289,1291-1301].eqiad.wmnet * 13:10 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply * 13:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki2003.codfw.wmnet * 13:06 jgiannelos@deploy1003: jgiannelos: Continuing with deployment * 13:05 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host rpki2003.codfw.wmnet * 13:04 jgiannelos@deploy1003: jgiannelos: Backport for [[gerrit:1326232{{!}}prv: Enable parsoid rendering for 5 wikis (T435115)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1288-1289,1291-1301].eqiad.wmnet * 13:02 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1326232{{!}}prv: Enable parsoid rendering for 5 wikis (T435115)]] * 12:53 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1288-1289,1291-1301].eqiad.wmnet * 12:53 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1273,1275-1287].eqiad.wmnet * 12:53 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1273,1275-1287].eqiad.wmnet * 12:52 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt1002.wikimedia.org * 12:51 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db2902.codfw.wmnet on all recursors * 12:51 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db2902.codfw.wmnet on all recursors * 12:51 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:51 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db2902.codfw.wmnet - fceratto@cumin1003" * 12:51 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db2902.codfw.wmnet - fceratto@cumin1003" * 12:46 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt1002.wikimedia.org * 12:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt2002.wikimedia.org * 12:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1273,1275-1287].eqiad.wmnet * 12:42 dhinus: repooled clouddb1032 that was currently <nowiki>{</nowiki>"weight": 0, "pooled": "inactive"<nowiki>}</nowiki> for both s4 and s6 * 12:41 dhinus: also depooled clouddb1017 (forgot it in the previous list) * 12:41 fnegri@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet * 12:40 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 12:40 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db2902.codfw.wmnet * 12:40 fnegri@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032.eqiad.wmnet * 12:40 fnegri@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032 * 12:39 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1017.eqiad.wmnet * 12:39 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt2002.wikimedia.org * 12:38 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1020.eqiad.wmnet * 12:38 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1018.eqiad.wmnet * 12:38 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet * 12:37 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1014.eqiad.wmnet * 12:37 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1013.eqiad.wmnet * 12:37 dhinus: depool again clouddb10[13,14,16,18,20] that were repooled by the cookbook sre.mysql.multiinstance_reboot * 12:36 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1273,1275-1287].eqiad.wmnet * 12:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1248-1261].eqiad.wmnet * 12:36 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1248-1261].eqiad.wmnet * 12:35 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2026.codfw.wmnet * 12:34 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 12:34 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 12:34 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 12:34 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 12:29 lucaswerkmeister-wmde@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 12:28 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2026.codfw.wmnet * 12:28 lucaswerkmeister-wmde@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 12:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1248-1261].eqiad.wmnet * 12:27 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.addnode (exit_code=0) for new host ganeti2046.codfw.wmnet to cluster codfw and group A * 12:26 moritzm: readded ganeti2046 to the codfw cluster following firmware update and reimage [[phab:T434681|T434681]] * 12:23 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply * 12:23 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply * 12:23 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply * 12:22 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply * 12:22 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply * 12:22 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply * 12:21 jmm@cumin2003: START - Cookbook sre.ganeti.addnode for new host ganeti2046.codfw.wmnet to cluster codfw and group A * 12:21 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1169.eqiad.wmnet onto db1283.eqiad.wmnet * 12:21 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1169: Pool db1169.eqiad.wmnet in after cloning * 12:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1248-1261].eqiad.wmnet * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1149-1153,1158,1240-1247].eqiad.wmnet * 12:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1149-1153,1158,1240-1247].eqiad.wmnet * 12:11 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1149-1153,1158,1240-1247].eqiad.wmnet * 12:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1168.eqiad.wmnet onto db1282.eqiad.wmnet * 12:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1168: Pool db1168.eqiad.wmnet in after cloning * 12:06 jmm@cumin2003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-test-eqiad * 12:05 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2026.codfw.wmnet * 12:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1149-1153,1158,1240-1247].eqiad.wmnet * 12:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1128-1134,1142-1148].eqiad.wmnet * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2046.codfw.wmnet * 12:01 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1128-1134,1142-1148].eqiad.wmnet * 11:57 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2025.codfw.wmnet * 11:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2025.codfw.wmnet * 11:54 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2046.codfw.wmnet * 11:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1128-1134,1142-1148].eqiad.wmnet * 11:50 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2025.codfw.wmnet * 11:46 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1128-1134,1142-1148].eqiad.wmnet * 11:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1114-1127].eqiad.wmnet * 11:45 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1114-1127].eqiad.wmnet * 11:39 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db2901.codfw.wmnet * 11:39 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db2901.codfw.wmnet * 11:36 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2025.codfw.wmnet * 11:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1114-1127].eqiad.wmnet * 11:35 fceratto@cumin1003: END (ERROR) - Cookbook sre.ganeti.makevm (exit_code=93) for new host db1901.eqiad.wmnet * 11:35 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1169: Pool db1169.eqiad.wmnet in after cloning * 11:35 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 11:30 jmm@cumin2003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-test-eqiad * 11:27 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1114-1127].eqiad.wmnet * 11:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1076-1081,1084-1087,1093-1095,1113].eqiad.wmnet * 11:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1076-1081,1084-1087,1093-1095,1113].eqiad.wmnet * 11:25 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1168: Pool db1168.eqiad.wmnet in after cloning * 11:24 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1194.eqiad.wmnet with OS bookworm * 11:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1181.eqiad.wmnet with OS bookworm * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2050.codfw.wmnet * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2050.codfw.wmnet * 11:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1076-1081,1084-1087,1093-1095,1113].eqiad.wmnet * 11:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2050.codfw.wmnet * 11:14 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2004.codfw.wmnet * 11:10 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2050.codfw.wmnet * 11:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1076-1081,1084-1087,1093-1095,1113].eqiad.wmnet * 11:09 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1286: Pool back * 11:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1045-1050,1056-1057,1064-1066,1073-1075].eqiad.wmnet * 11:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1045-1050,1056-1057,1064-1066,1073-1075].eqiad.wmnet * 11:08 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2004.codfw.wmnet * 11:07 moritzm: installing PHP 8.4 security updates * 11:06 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on an-worker1194.eqiad.wmnet with reason: host reimage * 11:06 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1194.eqiad.wmnet with reason: host reimage * 11:05 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2049.codfw.wmnet * 11:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2049.codfw.wmnet * 11:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid1003.eqiad.wmnet * 11:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1045-1050,1056-1057,1064-1066,1073-1075].eqiad.wmnet * 10:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1181.eqiad.wmnet with reason: host reimage * 10:59 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2049.codfw.wmnet * 10:59 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid1003.eqiad.wmnet * 10:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid2003.codfw.wmnet * 10:54 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1181.eqiad.wmnet with reason: host reimage * 10:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid2003.codfw.wmnet * 10:51 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1901.eqiad.wmnet on all recursors * 10:51 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1901.eqiad.wmnet on all recursors * 10:51 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:51 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:51 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:50 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2049.codfw.wmnet * 10:50 blake@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:50 blake@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:49 blake@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:48 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1045-1050,1056-1057,1064-1066,1073-1075].eqiad.wmnet * 10:48 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1044].eqiad.wmnet * 10:48 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1044].eqiad.wmnet * 10:47 blake@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:46 blake@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:46 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2047.codfw.wmnet * 10:46 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:46 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2047.codfw.wmnet * 10:46 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1901.eqiad.wmnet * 10:46 blake@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:44 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1168: Depool db1168.eqiad.wmnet to then clone it to db1282.eqiad.wmnet - marostegui@cumin1003 * 10:44 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1168: Depool db1168.eqiad.wmnet to then clone it to db1282.eqiad.wmnet - marostegui@cumin1003 * 10:44 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1168.eqiad.wmnet onto db1282.eqiad.wmnet * 10:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2047.codfw.wmnet * 10:40 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1044].eqiad.wmnet * 10:37 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2047.codfw.wmnet * 10:37 fceratto@cumin1003: END (ERROR) - Cookbook sre.ganeti.makevm (exit_code=93) for new host db1901.eqiad.wmnet * 10:36 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 10:34 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:33 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2032.codfw.wmnet * 10:33 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2032.codfw.wmnet * 10:32 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1044].eqiad.wmnet * 10:32 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-eqiad * 10:27 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2032.codfw.wmnet * 10:24 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1286: Pool back * 10:24 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1286 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96163 and previous config saved to /var/cache/conftool/dbconfig/20260818-102431-marostegui.json * 10:22 moritzm: installing Django security updates * 10:20 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul2002.codfw.wmnet * 10:20 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul2001.codfw.wmnet * 10:17 blake@deploy1003: Finished scap sync-world: no-build deployment for [[phab:T417800|T417800]] (duration: 04m 40s) * 10:16 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul2002.codfw.wmnet * 10:16 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul2001.codfw.wmnet * 10:15 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2032.codfw.wmnet * 10:14 blake@deploy1003: Started scap sync-world: no-build deployment for [[phab:T417800|T417800]] * 10:12 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul2003.codfw.wmnet * 10:12 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy2001.codfw.wmnet * 10:12 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy3001.esams.wmnet * 10:12 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy1001.eqiad.wmnet * 10:08 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2031.codfw.wmnet * 10:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2031.codfw.wmnet * 10:08 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul2003.codfw.wmnet * 10:08 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy2001.codfw.wmnet * 10:08 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy3001.esams.wmnet * 10:08 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy1001.eqiad.wmnet * 10:07 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy1002.eqiad.wmnet * 10:05 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy2002.codfw.wmnet * 10:05 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy3002.esams.wmnet * 10:04 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy1002.eqiad.wmnet * 10:03 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy4003.ulsfo.wmnet * 10:03 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy5003.eqsin.wmnet * 10:02 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2031.codfw.wmnet * 10:02 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1169: Depool db1169.eqiad.wmnet to then clone it to db1283.eqiad.wmnet - marostegui@cumin1003 * 10:01 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy2002.codfw.wmnet * 10:01 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy4004.ulsfo.wmnet * 10:01 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy3002.esams.wmnet * 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1169: Depool db1169.eqiad.wmnet to then clone it to db1283.eqiad.wmnet - marostegui@cumin1003 * 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1169.eqiad.wmnet onto db1283.eqiad.wmnet * 10:01 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy5004.eqsin.wmnet * 09:59 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy4003.ulsfo.wmnet * 09:59 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy4004.ulsfo.wmnet * 09:59 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy5003.eqsin.wmnet * 09:59 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy7001.magru.wmnet * 09:59 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy5004.eqsin.wmnet * 09:58 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy7002.magru.wmnet * 09:57 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2031.codfw.wmnet * 09:57 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy6001.drmrs.wmnet * 09:57 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy6002.drmrs.wmnet * 09:54 filippo@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for 10 hosts * 09:54 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast4006.wikimedia.org * 09:53 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy6001.drmrs.wmnet * 09:53 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy6002.drmrs.wmnet * 09:53 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host releases1003.eqiad.wmnet * 09:53 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy7001.magru.wmnet * 09:53 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host people2004.codfw.wmnet * 09:52 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host people1005.eqiad.wmnet * 09:52 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy7002.magru.wmnet * 09:50 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host releases2003.codfw.wmnet * 09:49 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host releases2003.codfw.wmnet * 09:49 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host releases1003.eqiad.wmnet * 09:49 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host people2004.codfw.wmnet * 09:48 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host people1005.eqiad.wmnet * 09:46 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2040.codfw.wmnet * 09:46 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2040.codfw.wmnet * 09:43 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1901.eqiad.wmnet on all recursors * 09:42 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1901.eqiad.wmnet on all recursors * 09:42 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:42 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:42 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2040.codfw.wmnet * 09:40 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp2005.wikimedia.org * 09:36 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp2005.wikimedia.org * 09:31 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:31 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1901.eqiad.wmnet * 09:31 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1901.eqiad.wmnet * 09:31 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:31 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1901.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 09:31 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1901.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 09:29 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2040.codfw.wmnet * 09:29 slyngshede@dns1004: END - running authdns-update * 09:28 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2039.codfw.wmnet * 09:28 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2039.codfw.wmnet * 09:27 slyngshede@dns1004: START - running authdns-update * 09:26 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test2005.wikimedia.org * 09:22 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2039.codfw.wmnet * 09:22 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test2005.wikimedia.org * 09:22 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp1005.wikimedia.org * 09:21 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:19 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2039.codfw.wmnet * 09:19 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2038.codfw.wmnet * 09:18 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1287: Pool back * 09:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2038.codfw.wmnet * 09:18 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp1005.wikimedia.org * 09:18 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test1005.wikimedia.org * 09:17 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1901.eqiad.wmnet * 09:15 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db1901.eqiad.wmnet * 09:15 fceratto@cumin1003: END (ERROR) - Cookbook sre.dns.netbox (exit_code=97) * 09:14 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test1005.wikimedia.org * 09:13 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2038.codfw.wmnet * 09:13 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:13 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1901.eqiad.wmnet * 09:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast4006.wikimedia.org * 09:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1288: Pool back * 09:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast5005.wikimedia.org * 09:03 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2038.codfw.wmnet * 08:57 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2037.codfw.wmnet * 08:57 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast5005.wikimedia.org * 08:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2037.codfw.wmnet * 08:56 filippo@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 10 hosts * 08:52 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2037.codfw.wmnet * 08:51 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1289: Pool back * 08:35 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudidp2001-dev.codfw.wmnet * 08:34 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1002-dev.eqiad.wmnet * 08:33 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1287: Pool back * 08:33 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1287 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96150 and previous config saved to /var/cache/conftool/dbconfig/20260818-083311-marostegui.json * 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1001-dev.eqiad.wmnet * 08:31 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudidp2001-dev.codfw.wmnet * 08:30 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1002-dev.eqiad.wmnet * 08:30 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2037.codfw.wmnet * 08:29 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1001-dev.eqiad.wmnet * 08:28 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2036.codfw.wmnet * 08:28 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2036.codfw.wmnet * 08:25 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1288: Pool back * 08:23 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 08:23 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1181.eqiad.wmnet with OS bookworm * 08:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2036.codfw.wmnet * 08:22 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1288 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96148 and previous config saved to /var/cache/conftool/dbconfig/20260818-082234-marostegui.json * 08:20 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2036.codfw.wmnet * 08:18 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2035.codfw.wmnet * 08:18 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudvirt1057.eqiad.wmnet with OS trixie * 08:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2035.codfw.wmnet * 08:13 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2035.codfw.wmnet * 08:06 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2035.codfw.wmnet * 08:05 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1289: Pool back * 08:05 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1289 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96145 and previous config saved to /var/cache/conftool/dbconfig/20260818-080531-marostegui.json * 07:54 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 07:51 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326439{{!}}Fix "mathjax_ignore" handling around forcemathmode attribute (T434686)]] (duration: 13m 15s) * 07:48 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 07:47 krinkle@deploy1003: krinkle: Continuing with deployment * 07:44 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ganeti2046.codfw.wmnet with OS bookworm * 07:40 krinkle@deploy1003: krinkle: Backport for [[gerrit:1326439{{!}}Fix "mathjax_ignore" handling around forcemathmode attribute (T434686)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:39 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 07:38 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1326439{{!}}Fix "mathjax_ignore" handling around forcemathmode attribute (T434686)]] * 07:34 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 07:31 samwilson@deploy1003: Finished scap sync-world: Backport for [[gerrit:701016{{!}}InitialiseSettings and -labs: Remove redundant feature flag $wgWikisourceEnableOcr (T285311)]] (duration: 07m 47s) * 07:30 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1048.eqiad.wmnet * 07:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1048.eqiad.wmnet * 07:29 XioNoX: add gnmic 0.47.0 to bookworm and trixie reprepro * 07:28 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ganeti2046.codfw.wmnet with reason: host reimage * 07:27 samwilson@deploy1003: samwilson: Continuing with deployment * 07:25 samwilson@deploy1003: samwilson: Backport for [[gerrit:701016{{!}}InitialiseSettings and -labs: Remove redundant feature flag $wgWikisourceEnableOcr (T285311)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:25 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 07:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1048.eqiad.wmnet * 07:24 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ganeti2046.codfw.wmnet with reason: host reimage * 07:23 samwilson@deploy1003: Started scap sync-world: Backport for [[gerrit:701016{{!}}InitialiseSettings and -labs: Remove redundant feature flag $wgWikisourceEnableOcr (T285311)]] * 07:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1198.eqiad.wmnet with OS bookworm * 07:18 samwilson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326450{{!}}InitialiseSettings.php: Enable Bulk OCR on pawikisource (T434648)]] (duration: 12m 17s) * 07:17 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 07:11 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1048.eqiad.wmnet * 07:11 samwilson@deploy1003: samwilson: Continuing with deployment * 07:11 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ganeti2046.codfw.wmnet with OS bookworm * 07:10 samwilson@deploy1003: samwilson: Backport for [[gerrit:1326450{{!}}InitialiseSettings.php: Enable Bulk OCR on pawikisource (T434648)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1057.eqiad.wmnet with OS trixie * 07:05 samwilson@deploy1003: Started scap sync-world: Backport for [[gerrit:1326450{{!}}InitialiseSettings.php: Enable Bulk OCR on pawikisource (T434648)]] * 07:02 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1198.eqiad.wmnet with reason: host reimage * 07:01 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1056.eqiad.wmnet with OS trixie * 07:01 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1056.eqiad.wmnet with OS trixie * 07:00 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1056.eqiad.wmnet with OS trixie * 07:00 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1056.eqiad.wmnet with OS trixie * 06:59 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudvirt1055.eqiad.wmnet with OS trixie * 06:58 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1198.eqiad.wmnet with reason: host reimage * 06:53 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1055.eqiad.wmnet with OS trixie * 06:52 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudvirt1054.eqiad.wmnet with OS trixie * 06:44 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1194.eqiad.wmnet with OS bookworm * 06:44 XioNoX: upgrade eqsin gnmic to 0.47.0 * 06:43 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1198.eqiad.wmnet with OS bookworm * 06:41 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1054.eqiad.wmnet with OS trixie * 06:09 arnaudb@cumin1003: END (PASS) - Cookbook sre.gerrit.restart-gerrit (exit_code=0) Restarting Gerrit on gerrit2002 * 06:06 arnaudb@cumin1003: START - Cookbook sre.gerrit.restart-gerrit Restarting Gerrit on gerrit2002 * 06:06 arnaudb@cumin1003: END (PASS) - Cookbook sre.gerrit.restart-gerrit (exit_code=0) Restarting Gerrit on gerrit1003 * 06:04 arnaudb@cumin1003: START - Cookbook sre.gerrit.restart-gerrit Restarting Gerrit on gerrit1003 * 06:02 arnaudb@cumin1003: END (PASS) - Cookbook sre.gerrit.restart-gerrit (exit_code=0) Restarting Gerrit on gerrit2003 * 06:00 arnaudb@cumin1003: START - Cookbook sre.gerrit.restart-gerrit Restarting Gerrit on gerrit2003 * 05:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1155.eqiad.wmnet with OS bookworm * 05:38 arnaudb: updating prometheusBearerToken on gerrit * 05:28 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1155.eqiad.wmnet with reason: host reimage * 05:23 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1155.eqiad.wmnet with reason: host reimage * 05:06 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1155.eqiad.wmnet with OS bookworm * 04:57 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1144.eqiad.wmnet with OS bookworm * 04:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1144.eqiad.wmnet with reason: host reimage * 04:29 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1144.eqiad.wmnet with reason: host reimage * 04:14 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1144.eqiad.wmnet with OS bookworm * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.13 (duration: 02m 23s) * 03:45 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1197.eqiad.wmnet with OS bookworm * 03:41 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1165.eqiad.wmnet with OS bookworm * 03:38 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.16 refs [[phab:T430835|T430835]] (duration: 34m 43s) * 03:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1164.eqiad.wmnet with OS bookworm * 03:35 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1196.eqiad.wmnet with OS bookworm * 03:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1163.eqiad.wmnet with OS bookworm * 03:22 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1197.eqiad.wmnet with reason: host reimage * 03:18 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1165.eqiad.wmnet with reason: host reimage * 03:15 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1196.eqiad.wmnet with reason: host reimage * 03:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1164.eqiad.wmnet with reason: host reimage * 03:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1163.eqiad.wmnet with reason: host reimage * 03:05 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1165.eqiad.wmnet with reason: host reimage * 03:04 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1164.eqiad.wmnet with reason: host reimage * 03:04 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1197.eqiad.wmnet with reason: host reimage * 03:04 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1196.eqiad.wmnet with reason: host reimage * 03:04 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1163.eqiad.wmnet with reason: host reimage * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.16 refs [[phab:T430835|T430835]] * 02:50 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1197.eqiad.wmnet with OS bookworm * 02:49 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1196.eqiad.wmnet with OS bookworm * 02:49 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1165.eqiad.wmnet with OS bookworm * 02:49 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1164.eqiad.wmnet with OS bookworm * 02:48 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1163.eqiad.wmnet with OS bookworm * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 46s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:15 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-codfw: Set storage compatability to NONE — [[phab:T433026|T433026]] - eevans@cumin1003 * 00:38 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-codfw: Set storage compatability to NONE — [[phab:T433026|T433026]] - eevans@cumin1003 * 00:11 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326373{{!}}PersonalDashboard: add newly renamed *ReviewChangesMlModel setting (T422148)]] (duration: 07m 07s) * 00:09 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-eqiad: Set storage compatability to NONE — [[phab:T433026|T433026]] - eevans@cumin1003 * 00:07 musikanimal@deploy1003: musikanimal: Continuing with deployment * 00:06 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1326373{{!}}PersonalDashboard: add newly renamed *ReviewChangesMlModel setting (T422148)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:04 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1326373{{!}}PersonalDashboard: add newly renamed *ReviewChangesMlModel setting (T422148)]] == 2026-08-17 == * 23:30 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-eqiad: Set storage compatability to NONE — [[phab:T433026|T433026]] - eevans@cumin1003 * 23:05 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-codfw: Set storage compatability to UPGRADING — [[phab:T433026|T433026]] - eevans@cumin1003 * 22:28 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-codfw: Set storage compatability to UPGRADING — [[phab:T433026|T433026]] - eevans@cumin1003 * 21:46 logmsgbot: jforrester Deployed security patch for [[phab:T435085|T435085]] * 21:39 swfrench@deploy1003: mwscript-k8s job started: purgeList.php # [[phab:T432412|T432412]] * 21:37 maryum: Undeploy security fix for [[phab:T433020|T433020]] * 21:23 maryum: Deployed security fix for [[phab:T433020|T433020]] * 21:21 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 21:21 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 21:14 maryum: Deployed security fix for [[phab:T434967|T434967]] * 20:58 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-eqiad: Set storage compatability to UPGRADING — [[phab:T433026|T433026]] - eevans@cumin1003 * 20:40 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324320{{!}}[itwiki/slwiki/tgwiki] Remove temporary Wikipedia 25 logos permanently (already reverted) (T414265 T414320 T415307)]] (duration: 06m 55s) * 20:40 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 20:39 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 20:36 cjming@deploy1003: cjming, superpes: Continuing with deployment * 20:35 cjming@deploy1003: cjming, superpes: Backport for [[gerrit:1324320{{!}}[itwiki/slwiki/tgwiki] Remove temporary Wikipedia 25 logos permanently (already reverted) (T414265 T414320 T415307)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:33 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1324320{{!}}[itwiki/slwiki/tgwiki] Remove temporary Wikipedia 25 logos permanently (already reverted) (T414265 T414320 T415307)]] * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ttmserver-test: apply * 20:30 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326361{{!}}Remove escaped paths in app site association file (T432412)]] (duration: 13m 28s) * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ttmserver-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-toolhub-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-toolhub-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-toolhub-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-toolhub-test: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-toolhub: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-toolhub: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-toolhub: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-toolhub: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-test: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-test: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 20:27 inflatador: bking@deploy1003 `charlie --services_dir dse-k8s-services -s opensearch-* -e dse-k8s-* apply` [[phab:T435125|T435125]] * 20:27 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-apifeatureusage-test: apply * 20:27 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-apifeatureusage-test: apply * 20:27 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-apifeatureusage-test: apply * 20:27 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-apifeatureusage-test: apply * 20:27 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-apifeatureusage: apply * 20:26 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-apifeatureusage: apply * 20:26 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-apifeatureusage: apply * 20:26 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-apifeatureusage: apply * 20:26 cjming@deploy1003: cjming, tsev: Continuing with deployment * 20:24 inflatador: bking@deploy1003 `charlie --services_dir dse-k8s-services -s opensearch-* -e dse-k8s-*` [[phab:T435125|T435125]] * 20:21 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-eqiad: Set storage compatability to UPGRADING — [[phab:T433026|T433026]] - eevans@cumin1003 * 20:19 cjming@deploy1003: cjming, tsev: Backport for [[gerrit:1326361{{!}}Remove escaped paths in app site association file (T432412)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:17 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1326361{{!}}Remove escaped paths in app site association file (T432412)]] * 20:15 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326077{{!}}InstrumentConstructiveEdits: anchor all runs to the nearest `interval` (T431493)]] (duration: 06m 25s) * 20:11 cjming@deploy1003: cjming: Continuing with deployment * 20:10 cjming@deploy1003: cjming: Backport for [[gerrit:1326077{{!}}InstrumentConstructiveEdits: anchor all runs to the nearest `interval` (T431493)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:08 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1326077{{!}}InstrumentConstructiveEdits: anchor all runs to the nearest `interval` (T431493)]] * 20:06 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1160.eqiad.wmnet with OS bookworm * 20:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1162.eqiad.wmnet with OS bookworm * 19:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1184.eqiad.wmnet with OS bookworm * 19:55 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1161.eqiad.wmnet with OS bookworm * 19:49 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1195.eqiad.wmnet with OS bookworm * 19:49 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-test: apply * 19:49 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-test: apply * 19:44 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1160.eqiad.wmnet with reason: host reimage * 19:41 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-codfw: Upgrade to Java 17 — [[phab:T433026|T433026]] - eevans@cumin1003 * 19:39 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1162.eqiad.wmnet with reason: host reimage * 19:36 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-eqsin and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 19:36 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-test: apply * 19:36 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1184.eqiad.wmnet with reason: host reimage * 19:32 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1161.eqiad.wmnet with reason: host reimage * 19:29 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1195.eqiad.wmnet with reason: host reimage * 19:26 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1161.eqiad.wmnet with reason: host reimage * 19:26 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1160.eqiad.wmnet with reason: host reimage * 19:26 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1184.eqiad.wmnet with reason: host reimage * 19:26 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1162.eqiad.wmnet with reason: host reimage * 19:25 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1195.eqiad.wmnet with reason: host reimage * 19:19 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-test: apply * 19:11 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1195.eqiad.wmnet with OS bookworm * 19:10 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1194.eqiad.wmnet with OS bookworm * 19:10 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1184.eqiad.wmnet with OS bookworm * 19:10 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1162.eqiad.wmnet with OS bookworm * 19:10 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1161.eqiad.wmnet with OS bookworm * 19:10 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1160.eqiad.wmnet with OS bookworm * 19:10 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 19:09 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 19:03 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-codfw: Upgrade to Java 17 — [[phab:T433026|T433026]] - eevans@cumin1003 * 18:58 dancy@deploy1003: Finished scap sync-world: testing [[phab:T375514|T375514]] (duration: 03m 13s) * 18:55 dancy@deploy1003: Started scap sync-world: testing [[phab:T375514|T375514]] * 18:55 dwisehaupt@dns1006: END - running authdns-update * 18:54 dancy@deploy1003: Installation of scap version "4.281.1" completed for 3 hosts * 18:54 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2009.codfw.wmnet * 18:53 dwisehaupt@dns1006: START - running authdns-update * 18:52 dancy@deploy1003: Installing scap version "4.281.1" for 3 host(s) * 18:47 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2009.codfw.wmnet * 18:41 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2008.codfw.wmnet * 18:34 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2008.codfw.wmnet * 18:30 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2007.codfw.wmnet * 18:23 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2007.codfw.wmnet * 18:16 dwisehaupt@dns1005: END - running authdns-update * 18:14 dwisehaupt@dns1005: START - running authdns-update * 18:04 swfrench@deploy1003: Finished scap sync-world: Deploy "Point Test Wiki to new docroot" - [[phab:T432412|T432412]] (duration: 26m 03s) * 18:00 swfrench@deploy1003: swfrench: Continuing with deployment * 17:51 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1057.eqiad.wmnet with OS trixie * 17:47 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1180.eqiad.wmnet with OS bookworm * 17:39 swfrench@deploy1003: swfrench: Deploy "Point Test Wiki to new docroot" - [[phab:T432412|T432412]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:39 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-eqsin and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 17:38 swfrench@deploy1003: Started scap sync-world: Deploy "Point Test Wiki to new docroot" - [[phab:T432412|T432412]] * 17:36 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1159.eqiad.wmnet with OS bookworm * 17:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1158.eqiad.wmnet with OS bookworm * 17:30 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-eqiad: Upgrade to Java 17 — [[phab:T433026|T433026]] - eevans@cumin1003 * 17:30 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-ulsfo or A:cp-drmrs and A:cp - 9.2.15 upgrade ([[phab:T434620|T434620]]) * 17:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1157.eqiad.wmnet with OS bookworm * 17:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1180.eqiad.wmnet with reason: host reimage * 17:22 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 17:21 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1193.eqiad.wmnet with OS bookworm * 17:21 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 17:21 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 17:20 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 17:20 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 17:20 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1183.eqiad.wmnet with OS bookworm * 17:18 swfrench@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 17:18 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 17:17 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1180.eqiad.wmnet with reason: host reimage * 17:17 swfrench@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 17:17 swfrench@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 17:16 swfrench@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 17:16 swfrench@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 17:15 swfrench@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 17:15 swfrench@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 17:15 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1192.eqiad.wmnet with OS bookworm * 17:14 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1159.eqiad.wmnet with reason: host reimage * 17:14 swfrench@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 17:13 swfrench@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 17:12 swfrench@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 17:09 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1158.eqiad.wmnet with reason: host reimage * 17:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1157.eqiad.wmnet with reason: host reimage * 17:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1180 * 17:02 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1180 * 17:01 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica-eqsin and A:liberica * 17:01 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1193.eqiad.wmnet with reason: host reimage * 16:57 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1183.eqiad.wmnet with reason: host reimage * 16:56 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1054.eqiad.wmnet with OS trixie * 16:55 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1056.eqiad.wmnet with OS trixie * 16:54 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1193.eqiad.wmnet with reason: host reimage * 16:54 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1192.eqiad.wmnet with reason: host reimage * 16:51 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-eqiad: Upgrade to Java 17 — [[phab:T433026|T433026]] - eevans@cumin1003 * 16:50 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1180 * 16:50 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1180.eqiad.wmnet 17.36.64.10.in-addr.arpa 7.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:50 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1180.eqiad.wmnet 17.36.64.10.in-addr.arpa 7.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:50 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:50 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1180 - btullis@cumin1003" * 16:50 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1180 - btullis@cumin1003" * 16:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1158.eqiad.wmnet with reason: host reimage * 16:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1157.eqiad.wmnet with reason: host reimage * 16:49 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica-eqsin and A:liberica * 16:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1183.eqiad.wmnet with reason: host reimage * 16:48 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1159.eqiad.wmnet with reason: host reimage * 16:47 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1192.eqiad.wmnet with reason: host reimage * 16:46 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns4004.wikimedia.org * 16:46 sukhe@dns1004: END - running authdns-update * 16:44 sukhe@dns1004: START - running authdns-update * 16:44 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns4004.wikimedia.org,service=authdns-update * 16:43 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns4004.wikimedia.org with OS trixie * 16:39 btullis@cumin1003: START - Cookbook sre.dns.netbox * 16:39 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1180 * 16:39 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1193.eqiad.wmnet with OS bookworm * 16:39 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1180.eqiad.wmnet with OS bookworm * 16:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1192.eqiad.wmnet with OS bookworm * 16:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1183.eqiad.wmnet with OS bookworm * 16:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1159.eqiad.wmnet with OS bookworm * 16:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1158.eqiad.wmnet with OS bookworm * 16:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1157.eqiad.wmnet with OS bookworm * 16:31 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1057.eqiad.wmnet with OS trixie * 16:30 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1057.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 16:29 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica-drmrs and A:liberica * 16:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1154.eqiad.wmnet with OS bookworm * 16:23 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1057.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 16:23 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1055.eqiad.wmnet with OS trixie * 16:22 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1057 * 16:22 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1057 * 16:21 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:21 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1057] - vriley@cumin1003" * 16:21 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1057] - vriley@cumin1003" * 16:19 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica-drmrs and A:liberica * 16:17 vriley@cumin1003: START - Cookbook sre.dns.netbox * 16:16 phuedx@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics-external: apply * 16:15 phuedx@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics-external: apply * 16:13 phuedx@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics-external: apply * 16:12 phuedx@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics-external: apply * 16:11 phuedx@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics-external: apply * 16:09 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics-external: apply * 16:09 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1176.eqiad.wmnet with OS bookworm * 16:07 btullis@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 16:06 btullis@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 16:05 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1191.eqiad.wmnet with OS bookworm * 16:03 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1154.eqiad.wmnet with reason: host reimage * 16:03 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching aqs[2002-2012].codfw.wmnet,aqs[1017-1027].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433026|T433026]] - eevans@cumin1003 * 15:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1190.eqiad.wmnet with OS bookworm * 15:59 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1154.eqiad.wmnet with reason: host reimage * 15:55 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-debug: apply * 15:55 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-debug: apply * 15:55 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-debug: apply * 15:55 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/mw-debug: apply * 15:53 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns4004.wikimedia.org with reason: host reimage * 15:50 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns4004.wikimedia.org with reason: host reimage * 15:46 btullis@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 15:46 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1176.eqiad.wmnet with reason: host reimage * 15:45 btullis@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 15:43 btullis@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 15:42 btullis@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 15:42 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1191.eqiad.wmnet with reason: host reimage * 15:39 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1190.eqiad.wmnet with reason: host reimage * 15:36 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1054.eqiad.wmnet with OS trixie * 15:35 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:35 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1056.eqiad.wmnet with OS trixie * 15:35 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:34 moritzm: failover Ganeti master in eqiad to ganeti1046 * 15:34 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1176.eqiad.wmnet with reason: host reimage * 15:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1191.eqiad.wmnet with reason: host reimage * 15:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1190.eqiad.wmnet with reason: host reimage * 15:31 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns2006.wikimedia.org * 15:31 sukhe@dns1004: END - running authdns-update * 15:29 sukhe@dns1004: START - running authdns-update * 15:29 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns2006.wikimedia.org,service=authdns-update * 15:29 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns2006.wikimedia.org * 15:29 sukhe@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns2006.wikimedia.org * 15:26 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica-ulsfo and A:liberica * 15:26 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns1006.wikimedia.org * 15:26 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:25 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns2006.wikimedia.org with OS trixie * 15:25 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1056 * 15:25 sukhe@dns1004: END - running authdns-update * 15:23 sukhe@dns1004: START - running authdns-update * 15:23 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns1006.wikimedia.org,service=authdns-update * 15:23 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns1006.wikimedia.org * 15:23 sukhe@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns1006.wikimedia.org * 15:20 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1056 * 15:20 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:20 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1056~] - vriley@cumin1003" * 15:19 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1056~] - vriley@cumin1003" * 15:19 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns1006.wikimedia.org with OS trixie * 15:19 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns4004.wikimedia.org with OS trixie * 15:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1190.eqiad.wmnet with OS bookworm * 15:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1191.eqiad.wmnet with OS bookworm * 15:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1176.eqiad.wmnet with OS bookworm * 15:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1154.eqiad.wmnet with OS bookworm * 15:16 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica-ulsfo and A:liberica * 15:13 vriley@cumin1003: START - Cookbook sre.dns.netbox * 15:13 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host dns4004.wikimedia.org with OS trixie * 15:12 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1188.eqiad.wmnet with OS bookworm * 15:09 taavi@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318203{{!}}Undeploy WP25EasterEggs (II) (T418134)]] (duration: 08m 56s) * 15:06 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica-magru and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 15:06 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1051.eqiad.wmnet with OS trixie * 15:06 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 15:05 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 15:05 taavi@deploy1003: taavi: Continuing with deployment * 15:04 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-ncredir (exit_code=0) rolling reboot on A:ncredir and A:ncredir * 15:04 taavi@deploy1003: taavi: Backport for [[gerrit:1318203{{!}}Undeploy WP25EasterEggs (II) (T418134)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:03 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1055.eqiad.wmnet with OS trixie * 15:02 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:00 taavi@deploy1003: Started scap sync-world: Backport for [[gerrit:1318203{{!}}Undeploy WP25EasterEggs (II) (T418134)]] * 14:59 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1047.eqiad.wmnet * 14:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1047.eqiad.wmnet * 14:58 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy (exit_code=0) rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 14:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-codfw * 14:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp2001.codfw.wmnet * 14:57 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp2001.codfw.wmnet * 14:57 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica-magru and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 14:57 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:56 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1055 * 14:56 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1055 * 14:55 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns2006.wikimedia.org with reason: host reimage * 14:55 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:55 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1055] - vriley@cumin1003" * 14:55 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1055] - vriley@cumin1003" * 14:54 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1047.eqiad.wmnet * 14:52 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1188.eqiad.wmnet with reason: host reimage * 14:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp2001.codfw.wmnet * 14:51 vriley@cumin1003: START - Cookbook sre.dns.netbox * 14:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp2001.codfw.wmnet * 14:50 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2016.codfw.wmnet * 14:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2016.codfw.wmnet * 14:50 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1189.eqiad.wmnet with OS bookworm * 14:48 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1188.eqiad.wmnet with reason: host reimage * 14:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1051.eqiad.wmnet with reason: host reimage * 14:46 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1182.eqiad.wmnet with OS bookworm * 14:45 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief2002.codfw.wmnet * 14:44 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns1006.wikimedia.org with reason: host reimage * 14:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2016.codfw.wmnet * 14:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1143.eqiad.wmnet with OS bookworm * 14:43 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2016.codfw.wmnet * 14:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2318-2331].codfw.wmnet * 14:43 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2318-2331].codfw.wmnet * 14:42 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching aqs[2002-2012].codfw.wmnet,aqs[1017-1027].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433026|T433026]] - eevans@cumin1003 * 14:41 cgoubert@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326306{{!}}Add placeholder $wmgRedisLockPassword (T366938 T427999)]] (duration: 06m 56s) * 14:41 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief2002.codfw.wmnet * 14:40 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief1002.eqiad.wmnet * 14:39 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1051.eqiad.wmnet with reason: host reimage * 14:38 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns2006.wikimedia.org with reason: host reimage * 14:37 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns1006.wikimedia.org with reason: host reimage * 14:37 cgoubert@deploy1003: cgoubert: Continuing with deployment * 14:36 cgoubert@deploy1003: cgoubert: Backport for [[gerrit:1326306{{!}}Add placeholder $wmgRedisLockPassword (T366938 T427999)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:36 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief1002.eqiad.wmnet * 14:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2318-2331].codfw.wmnet * 14:34 cgoubert@deploy1003: Started scap sync-world: Backport for [[gerrit:1326306{{!}}Add placeholder $wmgRedisLockPassword (T366938 T427999)]] * 14:29 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2318-2331].codfw.wmnet * 14:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2304-2317].codfw.wmnet * 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2304-2317].codfw.wmnet * 14:28 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1054 * 14:27 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1054 * 14:27 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:27 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1054] - vriley@cumin1003" * 14:27 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1054] - vriley@cumin1003" * 14:26 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1189.eqiad.wmnet with reason: host reimage * 14:26 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test2001.codfw.wmnet * 14:25 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test1001.eqiad.wmnet * 14:25 claime: Deploying wmgRedisLockPassword - [[phab:T366938|T366938]] [[phab:T427999|T427999]] * 14:24 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1051.eqiad.wmnet with OS trixie * 14:23 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 14:22 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1182.eqiad.wmnet with reason: host reimage * 14:22 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1051.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:22 vriley@cumin1003: START - Cookbook sre.dns.netbox * 14:22 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test2001.codfw.wmnet * 14:21 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test1001.eqiad.wmnet * 14:21 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 14:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2304-2317].codfw.wmnet * 14:19 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns4004.wikimedia.org with OS trixie * 14:19 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns2006.wikimedia.org with OS trixie * 14:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1143.eqiad.wmnet with reason: host reimage * 14:19 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns1006.wikimedia.org with OS trixie * 14:17 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1189.eqiad.wmnet with reason: host reimage * 14:15 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1182.eqiad.wmnet with reason: host reimage * 14:14 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1143.eqiad.wmnet with reason: host reimage * 14:13 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1051.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:12 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2304-2317].codfw.wmnet * 14:12 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2290-2303].codfw.wmnet * 14:12 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1049.eqiad.wmnet with OS trixie * 14:12 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2290-2303].codfw.wmnet * 14:12 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1051 * 14:11 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1051 * 14:11 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1170.eqiad.wmnet onto db1284.eqiad.wmnet * 14:11 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1170: Pool db1170.eqiad.wmnet in after cloning * 14:09 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 14:09 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:09 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1051] - vriley@cumin1003" * 14:09 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1051] - vriley@cumin1003" * 14:06 klausman@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:04 vriley@cumin1003: START - Cookbook sre.dns.netbox * 14:04 klausman@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2290-2303].codfw.wmnet * 14:02 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1189.eqiad.wmnet with OS bookworm * 14:02 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1188.eqiad.wmnet with OS bookworm * 14:01 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 14:00 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1182.eqiad.wmnet with OS bookworm * 14:00 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1143.eqiad.wmnet with OS bookworm * 13:59 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 13:57 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply * 13:57 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply * 13:56 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply * 13:56 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-ulsfo or A:cp-drmrs and A:cp - 9.2.15 upgrade ([[phab:T434620|T434620]]) * 13:56 cjd91: sudo -i cookbook sre.cdn.roll-upgrade-ats --query 'A:cp-ulsfo or A:cp-drmrs' --task-id [[phab:T434620|T434620]] --reason '9.2.15 upgrade' * 13:56 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply * 13:55 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply * 13:55 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply * 13:54 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 13:54 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 13:54 phuedx@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:53 phuedx@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics-external: apply * 13:52 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1049.eqiad.wmnet with reason: host reimage * 13:51 phuedx@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2290-2303].codfw.wmnet * 13:50 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1047.eqiad.wmnet * 13:50 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2276-2289].codfw.wmnet * 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2276-2289].codfw.wmnet * 13:49 phuedx@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics-external: apply * 13:49 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1046.eqiad.wmnet * 13:49 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1049.eqiad.wmnet with reason: host reimage * 13:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1046.eqiad.wmnet * 13:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-ncredir rolling reboot on A:ncredir and A:ncredir * 13:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 13:46 phuedx@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:44 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics-external: apply * 13:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1046.eqiad.wmnet * 13:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2276-2289].codfw.wmnet * 13:36 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1046.eqiad.wmnet * 13:36 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1045.eqiad.wmnet * 13:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1045.eqiad.wmnet * 13:34 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2276-2289].codfw.wmnet * 13:34 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1049.eqiad.wmnet with OS trixie * 13:34 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2262-2275].codfw.wmnet * 13:34 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2262-2275].codfw.wmnet * 13:32 Lucas_WMDE: UTC afternoon backport+config window done * 13:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1045.eqiad.wmnet * 13:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2262-2275].codfw.wmnet * 13:26 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1049.eqiad.wmnet with OS trixie * 13:26 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1049.eqiad.wmnet with OS trixie * 13:26 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1170: Pool db1170.eqiad.wmnet in after cloning * 13:23 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1049.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 13:23 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1045.eqiad.wmnet * 13:21 atsukoito: manually done sudo -i docker-registryctl --debug delete-tags 'docker-registry.discovery.wmnet/repos/data-engineering/airflow-dags:airflow-3.3.0-py3.11-2026-08-17-*' to remove incorrect tags * 13:19 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311141{{!}}viwiki: Set `noindex,nofollow` for User and User talk (T432311)]] (duration: 11m 11s) * 13:15 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2262-2275].codfw.wmnet * 13:14 lucaswerkmeister-wmde@deploy1003: ndkdd, lucaswerkmeister-wmde: Continuing with deployment * 13:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2248-2261].codfw.wmnet * 13:14 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2248-2261].codfw.wmnet * 13:13 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1049.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 13:12 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1049 * 13:11 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1049 * 13:10 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:10 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1049] - vriley@cumin1003" * 13:10 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1049] - vriley@cumin1003" * 13:10 lucaswerkmeister-wmde@deploy1003: ndkdd, lucaswerkmeister-wmde: Backport for [[gerrit:1311141{{!}}viwiki: Set `noindex,nofollow` for User and User talk (T432311)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1037.eqiad.wmnet * 13:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1037.eqiad.wmnet * 13:08 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1311141{{!}}viwiki: Set `noindex,nofollow` for User and User talk (T432311)]] * 13:06 vriley@cumin1003: START - Cookbook sre.dns.netbox * 13:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2248-2261].codfw.wmnet * 13:00 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1037.eqiad.wmnet * 12:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2248-2261].codfw.wmnet * 12:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2204-2215,2242-2243].codfw.wmnet * 12:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2204-2215,2242-2243].codfw.wmnet * 12:49 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324832{{!}}Migrate $wgFlaggedRevsTags from flaggedrevs.php to ext-FlaggedRevs.php]] (duration: 14m 02s) * 12:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint2001.codfw.wmnet * 12:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2204-2215,2242-2243].codfw.wmnet * 12:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint2001.codfw.wmnet * 12:41 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1037.eqiad.wmnet * 12:40 ladsgroup@deploy1003: Rolling back deployment * 12:37 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1324832{{!}}Migrate $wgFlaggedRevsTags from flaggedrevs.php to ext-FlaggedRevs.php]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:37 seanleong-wmde: Finished populateSitesTable for [bolwiki] ([[[phab:T429955|T429955]]]) * 12:35 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1324832{{!}}Migrate $wgFlaggedRevsTags from flaggedrevs.php to ext-FlaggedRevs.php]] * 12:35 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2204-2215,2242-2243].codfw.wmnet * 12:34 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2190-2203].codfw.wmnet * 12:34 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2190-2203].codfw.wmnet * 12:32 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1028.eqiad.wmnet * 12:32 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1028.eqiad.wmnet * 12:31 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint1001.eqiad.wmnet * 12:30 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1170: Depool db1170.eqiad.wmnet to then clone it to db1284.eqiad.wmnet - marostegui@cumin1003 * 12:28 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint1001.eqiad.wmnet * 12:28 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1170: Depool db1170.eqiad.wmnet to then clone it to db1284.eqiad.wmnet - marostegui@cumin1003 * 12:27 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1170.eqiad.wmnet onto db1284.eqiad.wmnet * 12:27 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db2901.codfw.wmnet * 12:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2190-2203].codfw.wmnet * 12:27 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 12:26 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1028.eqiad.wmnet * 12:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2190-2203].codfw.wmnet * 12:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2172-2179,2184-2189].codfw.wmnet * 12:18 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2172-2179,2184-2189].codfw.wmnet * 12:13 seanleong-wmde@deploy1003: mwscript-k8s job started: foreachwikiindblist wikidataclient extensions/Wikibase/lib/maintenance/populateSitesTable.php --force-protocol https # [[phab:T429955|T429955]] * 12:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2172-2179,2184-2189].codfw.wmnet * 12:09 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1028.eqiad.wmnet * 12:03 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 12:02 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db2901.codfw.wmnet * 12:02 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db2901.codfw.wmnet * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1027.eqiad.wmnet * 12:02 fceratto@cumin1003: END (ERROR) - Cookbook sre.dns.netbox (exit_code=97) * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1027.eqiad.wmnet * 12:02 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 12:02 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db2901.codfw.wmnet * 12:01 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2172-2179,2184-2189].codfw.wmnet * 12:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2158-2171].codfw.wmnet * 12:01 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2158-2171].codfw.wmnet * 11:58 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db1902.eqiad.wmnet * 11:58 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 11:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1027.eqiad.wmnet * 11:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2158-2171].codfw.wmnet * 11:51 jayme: updated calico to v3.30.7 on wikikube eqiad - [[phab:T427400|T427400]] * 11:50 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1027.eqiad.wmnet * 11:45 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2158-2171].codfw.wmnet * 11:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2144-2157].codfw.wmnet * 11:44 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2144-2157].codfw.wmnet * 11:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1058.eqiad.wmnet * 11:43 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1058.eqiad.wmnet * 11:43 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 11:43 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 11:43 marostegui@cumin1003: Removing db1153 from zarcillo [[phab:T434638|T434638]] * 11:42 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1153.eqiad.wmnet * 11:42 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:42 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1153.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 11:42 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1172.eqiad.wmnet onto db1286.eqiad.wmnet * 11:42 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1172: Pool db1172.eqiad.wmnet in after cloning * 11:42 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1153.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 11:41 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1902.eqiad.wmnet * 11:41 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=97) for new host db2901.codfw.wmnet * 11:41 fceratto@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host db2901.codfw.wmnet with OS trixie * 11:38 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'. * 11:38 marostegui@cumin1003: START - Cookbook sre.dns.netbox * 11:37 marostegui@dns1004: END - running authdns-update * 11:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1058.eqiad.wmnet * 11:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2144-2157].codfw.wmnet * 11:35 marostegui@dns1004: START - running authdns-update * 11:32 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1153.eqiad.wmnet * 11:32 marostegui@cumin1003: START - Cookbook sre.mysql.decommission * 11:28 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326249{{!}}ImagePage: move TOC element below file link (T332644)]] (duration: 09m 56s) * 11:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2144-2157].codfw.wmnet * 11:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2130-2143].codfw.wmnet * 11:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2130-2143].codfw.wmnet * 11:26 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1058.eqiad.wmnet * 11:23 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 11:22 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'. * 11:22 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'. * 11:22 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'. * 11:22 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326249{{!}}ImagePage: move TOC element below file link (T332644)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1057.eqiad.wmnet * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1057.eqiad.wmnet * 11:21 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 11:20 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 11:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2130-2143].codfw.wmnet * 11:19 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326249{{!}}ImagePage: move TOC element below file link (T332644)]] * 11:18 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply * 11:17 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply * 11:17 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply * 11:16 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 11:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1057.eqiad.wmnet * 11:15 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 11:14 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 11:13 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 11:13 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db2901.codfw.wmnet with OS trixie * 11:12 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db2901.codfw.wmnet - fceratto@cumin1003" * 11:12 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db2901.codfw.wmnet - fceratto@cumin1003" * 11:12 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db2901.codfw.wmnet on all recursors * 11:12 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db2901.codfw.wmnet on all recursors * 11:12 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:12 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db2901.codfw.wmnet - fceratto@cumin1003" * 11:12 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2130-2143].codfw.wmnet * 11:11 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2107-2115,2124-2129].codfw.wmnet * 11:11 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2107-2115,2124-2129].codfw.wmnet * 11:11 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 11:11 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 11:11 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply * 11:11 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 11:10 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 11:06 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db2901.codfw.wmnet - fceratto@cumin1003" * 11:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2107-2115,2124-2129].codfw.wmnet * 10:57 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1172: Pool db1172.eqiad.wmnet in after cloning * 10:54 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2107-2115,2124-2129].codfw.wmnet * 10:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2078,2087-2095,2102-2106].codfw.wmnet * 10:54 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2078,2087-2095,2102-2106].codfw.wmnet * 10:47 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply * 10:46 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply * 10:46 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply * 10:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2078,2087-2095,2102-2106].codfw.wmnet * 10:45 blake@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply * 10:38 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1057.eqiad.wmnet * 10:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1056.eqiad.wmnet * 10:38 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1056.eqiad.wmnet * 10:37 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2078,2087-2095,2102-2106].codfw.wmnet * 10:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2061-2062,2064-2065,2067-2077].codfw.wmnet * 10:36 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2061-2062,2064-2065,2067-2077].codfw.wmnet * 10:32 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1056.eqiad.wmnet * 10:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2061-2062,2064-2065,2067-2077].codfw.wmnet * 10:25 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:25 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db2901.codfw.wmnet * 10:25 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db1902.eqiad.wmnet * 10:25 fceratto@cumin1003: END (ERROR) - Cookbook sre.dns.netbox (exit_code=97) * 10:24 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:24 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1902.eqiad.wmnet * 10:24 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=97) for new host db1902.eqiad.wmnet * 10:24 fceratto@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host db1902.eqiad.wmnet with OS trixie * 10:24 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=93) for new host db1903.eqiad.wmnet * 10:24 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 10:20 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1056.eqiad.wmnet * 10:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2061-2062,2064-2065,2067-2077].codfw.wmnet * 10:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2038-2039,2041-2042,2044,2046,2049-2051,2055-2060].codfw.wmnet * 10:18 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2038-2039,2041-2042,2044,2046,2049-2051,2055-2060].codfw.wmnet * 10:15 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1055.eqiad.wmnet * 10:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1055.eqiad.wmnet * 10:13 Amir1: mwscript-k8s --dblist=all -- purgeUserOptions.php --login-age 5 uls-preferences ([[phab:T406724|T406724]]) * 10:11 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:10 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1903.eqiad.wmnet on all recursors * 10:10 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1903.eqiad.wmnet on all recursors * 10:10 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2038-2039,2041-2042,2044,2046,2049-2051,2055-2060].codfw.wmnet * 10:10 moritzm: installing unzip security updates * 10:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1055.eqiad.wmnet * 10:08 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=97) for new host db1901.eqiad.wmnet * 10:08 fceratto@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host db1901.eqiad.wmnet with OS trixie * 10:08 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:08 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 10:08 fceratto@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1903.eqiad.wmnet - fceratto@cumin1003" * 10:07 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db1902.eqiad.wmnet with OS trixie * 10:07 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1902.eqiad.wmnet - fceratto@cumin1003" * 10:07 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1902.eqiad.wmnet - fceratto@cumin1003" * 10:04 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply * 10:04 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326227{{!}}Enable desktop/native lazy loading everywhere (T148047)]] (duration: 07m 13s) * 10:03 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1902.eqiad.wmnet on all recursors * 10:03 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1902.eqiad.wmnet on all recursors * 10:03 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:03 blake@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply * 10:01 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1903.eqiad.wmnet - fceratto@cumin1003" * 10:01 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2038-2039,2041-2042,2044,2046,2049-2051,2055-2060].codfw.wmnet * 10:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2002,2005-2006,2011-2015,2017-2018,2033-2037].codfw.wmnet * 10:01 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2002,2005-2006,2011-2015,2017-2018,2033-2037].codfw.wmnet * 10:00 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:00 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 09:59 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 09:59 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1055.eqiad.wmnet * 09:58 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326227{{!}}Enable desktop/native lazy loading everywhere (T148047)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:56 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326227{{!}}Enable desktop/native lazy loading everywhere (T148047)]] * 09:53 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2002,2005-2006,2011-2015,2017-2018,2033-2037].codfw.wmnet * 09:50 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:48 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1903.eqiad.wmnet * 09:48 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:48 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1902.eqiad.wmnet * 09:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 09:44 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2002,2005-2006,2011-2015,2017-2018,2033-2037].codfw.wmnet * 09:43 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-codfw * 09:43 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1174.eqiad.wmnet onto db1288.eqiad.wmnet * 09:42 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1174: Pool db1174.eqiad.wmnet in after cloning * 09:40 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1172: Depool db1172.eqiad.wmnet to then clone it to db1286.eqiad.wmnet - marostegui@cumin1003 * 09:39 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1172: Depool db1172.eqiad.wmnet to then clone it to db1286.eqiad.wmnet - marostegui@cumin1003 * 09:39 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1172.eqiad.wmnet onto db1286.eqiad.wmnet * 09:33 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1175.eqiad.wmnet onto db1289.eqiad.wmnet * 09:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1175: Pool db1175.eqiad.wmnet in after cloning * 09:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1201.eqiad.wmnet onto db1287.eqiad.wmnet * 09:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1201: Pool db1201.eqiad.wmnet in after cloning * 09:28 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db1901.eqiad.wmnet with OS trixie * 09:27 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:27 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:27 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1901.eqiad.wmnet on all recursors * 09:27 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1901.eqiad.wmnet on all recursors * 09:26 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:26 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:26 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:15 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:15 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1901.eqiad.wmnet * 09:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw1001.wikimedia.org with OS trixie * 08:57 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1174: Pool db1174.eqiad.wmnet in after cloning * 08:54 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1054.eqiad.wmnet * 08:54 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1054.eqiad.wmnet * 08:48 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1054.eqiad.wmnet * 08:47 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1175: Pool db1175.eqiad.wmnet in after cloning * 08:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 08:46 marostegui@cumin1003: Removing db1152 from zarcillo [[phab:T434480|T434480]] * 08:46 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1152.eqiad.wmnet * 08:46 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:46 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1152.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 08:46 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1201: Pool db1201.eqiad.wmnet in after cloning * 08:46 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1152.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 08:46 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1054.eqiad.wmnet * 08:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1053.eqiad.wmnet * 08:43 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1053.eqiad.wmnet * 08:42 marostegui@cumin1003: START - Cookbook sre.dns.netbox * 08:38 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage * 08:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1053.eqiad.wmnet * 08:36 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1152.eqiad.wmnet * 08:36 marostegui@cumin1003: START - Cookbook sre.mysql.decommission * 08:35 phuedx: UTC morning backport window done * 08:35 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1053.eqiad.wmnet * 08:34 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1035.eqiad.wmnet * 08:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1035.eqiad.wmnet * 08:34 phuedx@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324725{{!}}EventStreamConfig: Mark product_metrics.web_base and .web_base_with_ip as Test Kitchen streams (T429898 T430322)]], [[gerrit:1313923{{!}}EventStreamConfig: Remove unused web_ui_scroll* streams (T415370)]], [[gerrit:1325546{{!}}EventStreamConfig: Remove Watchlist click stream (T434790)]] (duration: 12m 42s) * 08:33 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1281: Pool back * 08:32 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage * 08:31 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on db2209.codfw.wmnet with reason: Maintenance * 08:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2209: Maintenance needed * 08:30 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2209: Maintenance needed * 08:26 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1035.eqiad.wmnet * 08:26 phuedx@deploy1003: bearloga, phuedx: Continuing with deployment * 08:23 phuedx@deploy1003: bearloga, phuedx: Backport for [[gerrit:1324725{{!}}EventStreamConfig: Mark product_metrics.web_base and .web_base_with_ip as Test Kitchen streams (T429898 T430322)]], [[gerrit:1313923{{!}}EventStreamConfig: Remove unused web_ui_scroll* streams (T415370)]], [[gerrit:1325546{{!}}EventStreamConfig: Remove Watchlist click stream (T434790)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug * 08:21 phuedx@deploy1003: Started scap sync-world: Backport for [[gerrit:1324725{{!}}EventStreamConfig: Mark product_metrics.web_base and .web_base_with_ip as Test Kitchen streams (T429898 T430322)]], [[gerrit:1313923{{!}}EventStreamConfig: Remove unused web_ui_scroll* streams (T415370)]], [[gerrit:1325546{{!}}EventStreamConfig: Remove Watchlist click stream (T434790)]] * 08:19 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw1001.wikimedia.org with OS trixie * 08:18 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1035.eqiad.wmnet * 08:17 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1032.eqiad.wmnet * 08:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1032.eqiad.wmnet * 08:16 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1279: Pool back * 08:15 phuedx@deploy1003: Finished scap sync-world: Backport for [[gerrit:1216721{{!}}viwikivoyage: enable relatedarticle and pop-up (T405724)]] (duration: 39m 12s) * 08:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1032.eqiad.wmnet * 08:09 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1032.eqiad.wmnet * 08:08 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1031.eqiad.wmnet * 08:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1031.eqiad.wmnet * 08:03 godog: switch production to use dumps-nfs.w.o - [[phab:T432212|T432212]] * 08:02 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1031.eqiad.wmnet * 08:02 phuedx@deploy1003: nvdtn19, phuedx: Continuing with deployment * 08:00 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1201: Depool db1201.eqiad.wmnet to then clone it to db1287.eqiad.wmnet - marostegui@cumin1003 * 08:00 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1201: Depool db1201.eqiad.wmnet to then clone it to db1287.eqiad.wmnet - marostegui@cumin1003 * 08:00 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1201.eqiad.wmnet onto db1287.eqiad.wmnet * 08:00 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1057.eqiad.wmnet * 08:00 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:00 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1057.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:59 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1057.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:59 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1275: Pool back * 07:55 filippo@cumin1003: START - Cookbook sre.dns.netbox * 07:55 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1031.eqiad.wmnet * 07:52 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1030.eqiad.wmnet * 07:52 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1030.eqiad.wmnet * 07:52 phuedx@deploy1003: nvdtn19, phuedx: Backport for [[gerrit:1216721{{!}}viwikivoyage: enable relatedarticle and pop-up (T405724)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:51 tappof: bump space for prometheus k8s-aux in eqiad * 07:50 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1057.eqiad.wmnet * 07:48 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1281: Pool back * 07:47 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1281 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96108 and previous config saved to /var/cache/conftool/dbconfig/20260817-074749-marostegui.json * 07:46 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1030.eqiad.wmnet * 07:42 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1030.eqiad.wmnet * 07:41 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1174: Depool db1174.eqiad.wmnet to then clone it to db1288.eqiad.wmnet - marostegui@cumin1003 * 07:41 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1174: Depool db1174.eqiad.wmnet to then clone it to db1288.eqiad.wmnet - marostegui@cumin1003 * 07:41 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1174.eqiad.wmnet onto db1288.eqiad.wmnet * 07:40 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1029.eqiad.wmnet * 07:40 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm2001.wikimedia.org * 07:40 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1029.eqiad.wmnet * 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1056.eqiad.wmnet * 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1056.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:38 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1056.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:36 phuedx@deploy1003: Started scap sync-world: Backport for [[gerrit:1216721{{!}}viwikivoyage: enable relatedarticle and pop-up (T405724)]] * 07:36 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm2001.wikimedia.org * 07:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1279 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96104 and previous config saved to /var/cache/conftool/dbconfig/20260817-073542-marostegui.json * 07:34 filippo@cumin1003: START - Cookbook sre.dns.netbox * 07:34 slyngshede@dns1004: END - running authdns-update * 07:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1029.eqiad.wmnet * 07:33 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm-test1001.wikimedia.org * 07:32 slyngshede@dns1004: START - running authdns-update * 07:31 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1029.eqiad.wmnet * 07:31 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1279: Pool back * 07:30 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1279 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96102 and previous config saved to /var/cache/conftool/dbconfig/20260817-073038-marostegui.json * 07:29 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm-test1001.wikimedia.org * 07:29 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm1001.wikimedia.org * 07:28 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1056.eqiad.wmnet * 07:28 moritzm: extend the disk of ldap-rw1001 by 80G [[phab:T331699|T331699]] * 07:28 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1055.eqiad.wmnet * 07:28 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:28 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1055.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:27 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1055.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:26 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1044.eqiad.wmnet * 07:26 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1044.eqiad.wmnet * 07:25 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm1001.wikimedia.org * 07:24 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1175: Depool db1175.eqiad.wmnet to then clone it to db1289.eqiad.wmnet - marostegui@cumin1003 * 07:24 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1175: Depool db1175.eqiad.wmnet to then clone it to db1289.eqiad.wmnet - marostegui@cumin1003 * 07:24 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1175.eqiad.wmnet onto db1289.eqiad.wmnet * 07:22 filippo@cumin1003: START - Cookbook sre.dns.netbox * 07:20 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1044.eqiad.wmnet * 07:16 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1055.eqiad.wmnet * 07:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1054.eqiad.wmnet * 07:15 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:15 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1054.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:15 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1044.eqiad.wmnet * 07:15 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1054.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:13 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1275: Pool back * 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1275 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96098 and previous config saved to /var/cache/conftool/dbconfig/20260817-071225-marostegui.json * 07:10 filippo@cumin1003: START - Cookbook sre.dns.netbox * 07:05 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1043.eqiad.wmnet * 07:05 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1054.eqiad.wmnet * 07:05 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1051.eqiad.wmnet * 07:05 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:05 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1051.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:05 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1043.eqiad.wmnet * 07:04 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1051.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 06:59 filippo@cumin1003: START - Cookbook sre.dns.netbox * 06:59 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin1001.eqiad.wmnet * 06:59 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1043.eqiad.wmnet * 06:59 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin2001.codfw.wmnet * 06:55 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin2001.codfw.wmnet * 06:55 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1051.eqiad.wmnet * 06:54 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1049.eqiad.wmnet * 06:54 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:54 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1049.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 06:54 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1049.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 06:54 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin1001.eqiad.wmnet * 06:53 moritzm: installing apr-util security updates * 06:52 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1043.eqiad.wmnet * 06:49 filippo@cumin1003: START - Cookbook sre.dns.netbox * 06:41 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1049.eqiad.wmnet * 06:13 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit2003.wikimedia.org * 06:13 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet * 06:07 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet * 06:06 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit2003.wikimedia.org * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 47s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-16 == * 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 01m 03s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-15 == * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 41s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-14 == * 15:38 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-staging-master-eqiad * 15:38 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster1005.eqiad.wmnet * 15:38 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster1005.eqiad.wmnet * 15:35 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sretest2009.codfw.wmnet * 15:33 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster1005.eqiad.wmnet * 15:33 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster1005.eqiad.wmnet * 15:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster1004.eqiad.wmnet * 15:32 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster1004.eqiad.wmnet * 15:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host sretest2009.codfw.wmnet * 15:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster1004.eqiad.wmnet * 15:27 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster1004.eqiad.wmnet * 15:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster1003.eqiad.wmnet * 15:27 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster1003.eqiad.wmnet * 15:24 dancy@deploy1003: Finished scap sync-world: testing (duration: 03m 23s) * 15:22 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster1003.eqiad.wmnet * 15:22 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster1003.eqiad.wmnet * 15:22 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-staging-master-eqiad * 15:20 dancy@deploy1003: Started scap sync-world: testing * 15:20 dancy@deploy1003: Installation of scap version "4.280.2" completed for 3 hosts * 15:18 dancy@deploy1003: Installing scap version "4.280.2" for 3 host(s) * 15:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sretest2006.codfw.wmnet * 14:54 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host sretest2006.codfw.wmnet * 14:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sretest2003.codfw.wmnet * 14:39 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host sretest2003.codfw.wmnet * 13:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt-staging2001.codfw.wmnet * 13:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt-staging2001.codfw.wmnet * 13:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-staging-master-codfw * 13:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster2005.codfw.wmnet * 13:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster2005.codfw.wmnet * 13:05 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox-dev2003.codfw.wmnet * 13:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster2005.codfw.wmnet * 13:04 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster2005.codfw.wmnet * 13:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster2004.codfw.wmnet * 13:04 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster2004.codfw.wmnet * 13:01 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netbox-dev2003.codfw.wmnet * 12:59 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster2004.codfw.wmnet * 12:59 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster2004.codfw.wmnet * 12:59 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster2003.codfw.wmnet * 12:59 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster2003.codfw.wmnet * 12:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster2003.codfw.wmnet * 12:54 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster2003.codfw.wmnet * 12:54 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-staging-master-codfw * 12:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-staging-worker-eqiad * 12:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage1006.eqiad.wmnet * 12:52 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage1006.eqiad.wmnet * 12:46 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage1006.eqiad.wmnet * 12:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw1001.wikimedia.org with OS trixie * 12:45 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage1006.eqiad.wmnet * 12:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage1005.eqiad.wmnet * 12:45 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage1005.eqiad.wmnet * 12:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage1005.eqiad.wmnet * 12:36 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/ratelimit: apply * 12:35 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/ratelimit: apply * 12:35 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:35 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:33 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage1005.eqiad.wmnet * 12:33 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage1004.eqiad.wmnet * 12:33 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage1004.eqiad.wmnet * 12:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage1004.eqiad.wmnet * 12:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage1004.eqiad.wmnet * 12:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage1003.eqiad.wmnet * 12:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage1003.eqiad.wmnet * 12:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage * 12:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage1003.eqiad.wmnet * 12:17 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage * 12:14 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage1003.eqiad.wmnet * 12:14 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-staging-worker-eqiad * 12:03 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw1001.wikimedia.org with OS trixie * 12:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cuminunpriv1001.eqiad.wmnet * 11:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cuminunpriv1001.eqiad.wmnet * 11:27 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1004.wikimedia.org * 11:24 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1153 from dbctl [[phab:T434638|T434638]]', diff saved to https://phabricator.wikimedia.org/P96097 and previous config saved to /var/cache/conftool/dbconfig/20260814-112449-marostegui.json * 11:21 aokoth@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1004.wikimedia.org * 11:20 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 11:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-staging-worker-codfw * 11:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2004.codfw.wmnet * 11:17 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2004.codfw.wmnet * 11:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2004.codfw.wmnet * 11:10 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2004.codfw.wmnet * 11:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2003.codfw.wmnet * 11:10 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2003.codfw.wmnet * 11:03 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-worker1181.eqiad.wmnet with OS bookworm * 11:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2003.codfw.wmnet * 11:03 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2003.codfw.wmnet * 11:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2002.codfw.wmnet * 11:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2002.codfw.wmnet * 10:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1235.eqiad.wmnet with OS bookworm * 10:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock2003.codfw.wmnet * 10:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock2003.codfw.wmnet with OS trixie * 10:56 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2002.codfw.wmnet * 10:54 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1153.eqiad.wmnet with OS bookworm * 10:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2002.codfw.wmnet * 10:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2001.codfw.wmnet * 10:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2001.codfw.wmnet * 10:50 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 10:49 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1187.eqiad.wmnet with OS bookworm * 10:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2001.codfw.wmnet * 10:44 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2001.codfw.wmnet * 10:44 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-staging-worker-codfw * 10:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock2003.codfw.wmnet with reason: host reimage * 10:38 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock2003.codfw.wmnet with reason: host reimage * 10:35 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1153.eqiad.wmnet with reason: host reimage * 10:32 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1235.eqiad.wmnet with reason: host reimage * 10:29 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1187.eqiad.wmnet with reason: host reimage * 10:24 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1153.eqiad.wmnet with reason: host reimage * 10:23 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1235.eqiad.wmnet with reason: host reimage * 10:21 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1187.eqiad.wmnet with reason: host reimage * 10:16 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock2003.codfw.wmnet with OS trixie * 10:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install1005.wikimedia.org * 10:12 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 10:12 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 10:12 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 10:12 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 10:12 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:12 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 10:12 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 10:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install1005.wikimedia.org * 10:07 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1235.eqiad.wmnet with OS bookworm * 10:07 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1187.eqiad.wmnet with OS bookworm * 10:07 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1181.eqiad.wmnet with OS bookworm * 10:07 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1153.eqiad.wmnet with OS bookworm * 10:07 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install2005.wikimedia.org * 10:03 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 10:03 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2003.codfw.wmnet * 10:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1232.eqiad.wmnet with OS bookworm * 10:00 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install2005.wikimedia.org * 10:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install3004.wikimedia.org * 09:58 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 09:58 fceratto@cumin1003: Removing db1151 from zarcillo [[phab:T434538|T434538]] * 09:56 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1231.eqiad.wmnet with OS bookworm * 09:56 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 09:53 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 09:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install3004.wikimedia.org * 09:50 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1152.eqiad.wmnet with OS bookworm * 09:49 Dreamy_Jazz: `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260808000000" --end-timestamp="20260812120000" --sleep="5" --batch-size="50"` for [[phab:T434688|T434688]] * 09:48 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host centrallog2002.codfw.wmnet * 09:48 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install4004.wikimedia.org * 09:42 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1232.eqiad.wmnet with reason: host reimage * 09:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install4004.wikimedia.org * 09:41 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host centrallog2002.codfw.wmnet * 09:39 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install5004.wikimedia.org * 09:36 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1232.eqiad.wmnet with reason: host reimage * 09:36 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'. * 09:34 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'. * 09:33 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1231.eqiad.wmnet with reason: host reimage * 09:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install5004.wikimedia.org * 09:32 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host centrallog1002.eqiad.wmnet * 09:30 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install6003.wikimedia.org * 09:30 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1231.eqiad.wmnet with reason: host reimage * 09:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1152.eqiad.wmnet with reason: host reimage * 09:25 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host centrallog1002.eqiad.wmnet * 09:25 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1152.eqiad.wmnet with reason: host reimage * 09:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install6003.wikimedia.org * 09:22 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1232 * 09:22 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1232 * 09:22 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1232 * 09:22 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1232.eqiad.wmnet 25.53.64.10.in-addr.arpa 5.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:22 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host titan1001.eqiad.wmnet * 09:22 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1232.eqiad.wmnet 25.53.64.10.in-addr.arpa 5.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:22 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:22 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1232 - btullis@cumin1003" * 09:22 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1232 - btullis@cumin1003" * 09:17 btullis@cumin1003: START - Cookbook sre.dns.netbox * 09:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install7002.wikimedia.org * 09:17 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1232 * 09:16 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1231 * 09:16 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1231 * 09:14 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1231 * 09:14 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1231.eqiad.wmnet 24.53.64.10.in-addr.arpa 4.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host titan1001.eqiad.wmnet * 09:14 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1231.eqiad.wmnet 24.53.64.10.in-addr.arpa 4.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:14 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1231 - btullis@cumin1003" * 09:14 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1231 - btullis@cumin1003" * 09:10 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install7002.wikimedia.org * 09:09 btullis@cumin1003: START - Cookbook sre.dns.netbox * 09:08 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1231 * 09:08 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1152 * 09:08 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1152 * 09:06 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1152 * 09:06 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1152.eqiad.wmnet 16.53.64.10.in-addr.arpa 6.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:06 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1152.eqiad.wmnet 16.53.64.10.in-addr.arpa 6.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:06 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:06 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1152 - btullis@cumin1003" * 09:06 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1152 - btullis@cumin1003" * 09:03 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1151.eqiad.wmnet * 09:03 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:03 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1151.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 08:55 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host titan2001.codfw.wmnet * 08:55 btullis@cumin1003: START - Cookbook sre.dns.netbox * 08:54 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1151.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 08:54 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1152 * 08:53 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1232.eqiad.wmnet with OS bookworm * 08:53 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1231.eqiad.wmnet with OS bookworm * 08:53 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1152.eqiad.wmnet with OS bookworm * 08:51 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2209: Pool back * 08:51 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1201.eqiad.wmnet * 08:50 btullis@cumin1003: START - Cookbook sre.hosts.remove-downtime for an-worker1201.eqiad.wmnet * 08:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1229.eqiad.wmnet with OS bookworm * 08:47 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host titan2001.codfw.wmnet * 08:46 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping2004.codfw.wmnet * 08:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ping2004.codfw.wmnet * 08:40 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host an-worker1230.eqiad.wmnet with OS bookworm * 08:40 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 08:34 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1151.eqiad.wmnet * 08:34 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 08:31 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host titan1002.eqiad.wmnet * 08:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1229.eqiad.wmnet with reason: host reimage * 08:25 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host titan1002.eqiad.wmnet * 08:25 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1229.eqiad.wmnet with reason: host reimage * 08:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping1004.eqiad.wmnet * 08:21 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ping1004.eqiad.wmnet * 08:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1230.eqiad.wmnet with reason: host reimage * 08:12 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1230.eqiad.wmnet with reason: host reimage * 08:11 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1229 * 08:11 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1229 * 08:11 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1229 * 08:11 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1229.eqiad.wmnet 22.53.64.10.in-addr.arpa 2.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:11 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1229.eqiad.wmnet 22.53.64.10.in-addr.arpa 2.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:11 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:11 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1229 - btullis@cumin1003" * 08:11 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1229 - btullis@cumin1003" * 08:10 btullis@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1201.eqiad.wmnet with reason: Fixing a disk * 08:07 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host titan2002.codfw.wmnet * 08:05 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2209: Pool back * 08:05 btullis@cumin1003: START - Cookbook sre.dns.netbox * 08:00 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host titan2002.codfw.wmnet * 07:59 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1229 * 07:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1230 * 07:58 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1230 * 07:55 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1230 * 07:55 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1230.eqiad.wmnet 23.53.64.10.in-addr.arpa 3.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:55 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1230.eqiad.wmnet 23.53.64.10.in-addr.arpa 3.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:55 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:55 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1230 - btullis@cumin1003" * 07:55 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1230 - btullis@cumin1003" * 07:51 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host kubestagemaster2005.codfw.wmnet with OS trixie * 07:48 btullis@cumin1003: START - Cookbook sre.dns.netbox * 07:41 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1230 * 07:41 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1229.eqiad.wmnet with OS bookworm * 07:41 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1230.eqiad.wmnet with OS bookworm * 07:39 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1152 from dbctl [[phab:T434480|T434480]]', diff saved to https://phabricator.wikimedia.org/P96090 and previous config saved to /var/cache/conftool/dbconfig/20260814-073941-marostegui.json * 07:29 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on kubestagemaster2005.codfw.wmnet with reason: host reimage * 07:23 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on kubestagemaster2005.codfw.wmnet with reason: host reimage * 07:04 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host kubestagemaster2005.codfw.wmnet with OS trixie * 06:53 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1228.eqiad.wmnet with OS bookworm * 06:44 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1227.eqiad.wmnet with OS bookworm * 06:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1209.eqiad.wmnet with OS bookworm * 06:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1175.eqiad.wmnet with OS bookworm * 06:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1228.eqiad.wmnet with reason: host reimage * 06:27 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1228.eqiad.wmnet with reason: host reimage * 06:25 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1227.eqiad.wmnet with reason: host reimage * 06:21 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1227.eqiad.wmnet with reason: host reimage * 06:18 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1209.eqiad.wmnet with reason: host reimage * 06:14 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1209.eqiad.wmnet with reason: host reimage * 06:14 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1228 * 06:14 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1228 * 06:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1175.eqiad.wmnet with reason: host reimage * 06:12 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1228 * 06:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1228.eqiad.wmnet 20.53.64.10.in-addr.arpa 0.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:12 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1228.eqiad.wmnet 20.53.64.10.in-addr.arpa 0.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1228 - ryankemper@cumin2003" * 06:12 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1228 - ryankemper@cumin2003" * 06:09 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1175.eqiad.wmnet with reason: host reimage * 06:07 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 06:07 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1228 * 06:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1227 * 06:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1227 * 06:06 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1227 * 06:06 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1227.eqiad.wmnet 19.53.64.10.in-addr.arpa 9.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:06 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1227.eqiad.wmnet 19.53.64.10.in-addr.arpa 9.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:06 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:06 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1227 - ryankemper@cumin2003" * 06:06 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1227 - ryankemper@cumin2003" * 06:00 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 06:00 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1227 * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1209 * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1209 * 06:00 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1209 * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1209.eqiad.wmnet 15.53.64.10.in-addr.arpa 5.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:00 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1209.eqiad.wmnet 15.53.64.10.in-addr.arpa 5.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1209 - ryankemper@cumin2003" * 06:00 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1209 - ryankemper@cumin2003" * 05:54 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 05:54 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1209 * 05:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1175 * 05:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1175 * 05:52 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1175 * 05:52 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1175.eqiad.wmnet 17.53.64.10.in-addr.arpa 7.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 05:52 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1175.eqiad.wmnet 17.53.64.10.in-addr.arpa 7.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 05:52 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 05:52 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1175 - ryankemper@cumin2003" * 05:52 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1175 - ryankemper@cumin2003" * 05:49 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1228.eqiad.wmnet with OS bookworm * 05:49 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1227.eqiad.wmnet with OS bookworm * 05:48 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1209.eqiad.wmnet with OS bookworm * 05:47 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 05:47 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1175 * 05:47 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1175.eqiad.wmnet with OS bookworm * 05:09 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host apifeatureusage1001.eqiad.wmnet with OS bookworm * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 03s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:10 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 01:07 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 01:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 01:02 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 00:59 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1226.eqiad.wmnet with OS bookworm * 00:47 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1225.eqiad.wmnet with OS bookworm * 00:41 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1224.eqiad.wmnet with OS bookworm * 00:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1226.eqiad.wmnet with reason: host reimage * 00:31 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1226.eqiad.wmnet with reason: host reimage * 00:28 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1225.eqiad.wmnet with reason: host reimage * 00:25 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1225.eqiad.wmnet with reason: host reimage * 00:19 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1224.eqiad.wmnet with reason: host reimage * 00:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1226 * 00:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1226 * 00:17 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1226 * 00:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1226.eqiad.wmnet 23.36.64.10.in-addr.arpa 3.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:17 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1226.eqiad.wmnet 23.36.64.10.in-addr.arpa 3.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 00:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1226 - ryankemper@cumin2003" * 00:17 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1226 - ryankemper@cumin2003" * 00:16 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1224.eqiad.wmnet with reason: host reimage * 00:12 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 00:12 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1226 * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1225 * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1225 * 00:10 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1225 * 00:10 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1225.eqiad.wmnet 22.36.64.10.in-addr.arpa 2.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:10 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1225.eqiad.wmnet 22.36.64.10.in-addr.arpa 2.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:10 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 00:10 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1225 - ryankemper@cumin2003" * 00:10 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1225 - ryankemper@cumin2003" * 00:03 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 00:02 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1225 * 00:02 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1224 * 00:02 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1224 * 00:00 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1224 * 00:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1224.eqiad.wmnet 21.36.64.10.in-addr.arpa 1.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:00 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1224.eqiad.wmnet 21.36.64.10.in-addr.arpa 1.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 00:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1224 - ryankemper@cumin2003" * 00:00 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1224 - ryankemper@cumin2003" == 2026-08-13 == * 23:54 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1226.eqiad.wmnet with OS bookworm * 23:53 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1225.eqiad.wmnet with OS bookworm * 23:52 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 23:51 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1224 * 23:51 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1224.eqiad.wmnet with OS bookworm * 23:47 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1223.eqiad.wmnet with OS bookworm * 23:29 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1223.eqiad.wmnet with reason: host reimage * 23:24 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1223.eqiad.wmnet with reason: host reimage * 23:19 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325525{{!}}ve.ui.CodeMirror.less: ensure normal font style]] (duration: 11m 40s) * 23:16 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260807000000" --end-timestamp="20260808000000" --sleep="5" --batch-size="50"` for [[phab:T434688|T434688]] * 23:13 musikanimal@deploy1003: musikanimal: Continuing with deployment * 23:11 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1325525{{!}}ve.ui.CodeMirror.less: ensure normal font style]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1223 * 23:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1223 * 23:08 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1325525{{!}}ve.ui.CodeMirror.less: ensure normal font style]] * 23:07 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1223 * 23:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1223.eqiad.wmnet 20.36.64.10.in-addr.arpa 0.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:07 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1223.eqiad.wmnet 20.36.64.10.in-addr.arpa 0.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 23:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1223 - ryankemper@cumin2003" * 23:03 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1223 - ryankemper@cumin2003" * 23:01 sbassett: Deployed security updates for [[phab:T430596|T430596]], [[phab:T120386|T120386]] * 22:55 bking@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host kubestagemaster2005.codfw.wmnet * 22:55 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host kubestagemaster2005.codfw.wmnet with OS bookworm * 22:54 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 22:53 ryankemper@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 22:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on kubestagemaster2005.codfw.wmnet with reason: host reimage * 22:49 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1222.eqiad.wmnet with OS bookworm * 22:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on kubestagemaster2005.codfw.wmnet with reason: host reimage * 22:29 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1222.eqiad.wmnet with reason: host reimage * 22:26 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1222.eqiad.wmnet with reason: host reimage * 22:23 bking@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host aux-k8s-etcd2003.codfw.wmnet * 22:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd2003.codfw.wmnet with OS bookworm * 22:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host kubestagemaster2005.codfw.wmnet with OS bookworm * 22:23 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM kubestagemaster2005.codfw.wmnet - bking@cumin2003" * 22:23 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM kubestagemaster2005.codfw.wmnet - bking@cumin2003" * 22:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) kubestagemaster2005.codfw.wmnet on all recursors * 22:22 bking@cumin2003: START - Cookbook sre.dns.wipe-cache kubestagemaster2005.codfw.wmnet on all recursors * 22:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:22 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM kubestagemaster2005.codfw.wmnet - bking@cumin2003" * 22:22 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM kubestagemaster2005.codfw.wmnet - bking@cumin2003" * 22:17 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 22:13 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1223 * 22:12 bking@cumin2003: START - Cookbook sre.dns.netbox * 22:12 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host kubestagemaster2005.codfw.wmnet * 22:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1222 * 22:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1222 * 22:11 sbassett: Deployed security updates for [[phab:T429244|T429244]], [[phab:T434039|T434039]] * 22:11 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1222 * 22:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1222.eqiad.wmnet 19.36.64.10.in-addr.arpa 9.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:11 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1222.eqiad.wmnet 19.36.64.10.in-addr.arpa 9.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1222 - ryankemper@cumin2003" * 22:07 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1222 - ryankemper@cumin2003" * 22:05 robh@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-wdqs2001.codfw.wmnet with reason: updating firmware * 22:01 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 22:01 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1223.eqiad.wmnet with OS bookworm * 22:01 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1222 * 22:01 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1222.eqiad.wmnet with OS bookworm * 22:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1212.eqiad.wmnet with OS bookworm * 21:54 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324295{{!}}Revert "Lazily reject pre-fix parser-cache entries for noreferrer/noopener links" (T429090)]] (duration: 06m 42s) * 21:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd2003.codfw.wmnet with reason: host reimage * 21:53 bking@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host dse-k8s-etcd2001.codfw.wmnet * 21:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dse-k8s-etcd2001.codfw.wmnet with OS bookworm * 21:50 sbassett@deploy1003: sbassett, kharlan: Continuing with deployment * 21:49 sbassett@deploy1003: sbassett, kharlan: Backport for [[gerrit:1324295{{!}}Revert "Lazily reject pre-fix parser-cache entries for noreferrer/noopener links" (T429090)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:48 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aux-k8s-etcd2003.codfw.wmnet with reason: host reimage * 21:47 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1324295{{!}}Revert "Lazily reject pre-fix parser-cache entries for noreferrer/noopener links" (T429090)]] * 21:42 maryum: Deployed security patch for [[phab:T434549|T434549]] * 21:39 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1212.eqiad.wmnet with reason: host reimage * 21:34 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1212.eqiad.wmnet with reason: host reimage * 21:33 bking@cumin2003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd2003.codfw.wmnet with OS bookworm * 21:32 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM aux-k8s-etcd2003.codfw.wmnet - bking@cumin2003" * 21:32 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM aux-k8s-etcd2003.codfw.wmnet - bking@cumin2003" * 21:32 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) aux-k8s-etcd2003.codfw.wmnet on all recursors * 21:32 bking@cumin2003: START - Cookbook sre.dns.wipe-cache aux-k8s-etcd2003.codfw.wmnet on all recursors * 21:32 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:32 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM aux-k8s-etcd2003.codfw.wmnet - bking@cumin2003" * 21:31 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM aux-k8s-etcd2003.codfw.wmnet - bking@cumin2003" * 21:28 maryum: Deployed security patch for [[phab:T434619|T434619]] * 21:26 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:26 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host aux-k8s-etcd2003.codfw.wmnet * 21:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-etcd2001.codfw.wmnet with reason: host reimage * 21:20 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1212 * 21:20 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1212 * 21:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1212.eqiad.wmnet with OS bookworm * 21:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on dse-k8s-etcd2001.codfw.wmnet with reason: host reimage * 21:04 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325553{{!}}InstrumentConstructiveEdits: exclude mw-reverted as well (T431493)]] (duration: 06m 25s) * 20:59 kemayo@deploy1003: kemayo: Continuing with deployment * 20:59 kemayo@deploy1003: kemayo: Backport for [[gerrit:1325553{{!}}InstrumentConstructiveEdits: exclude mw-reverted as well (T431493)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host dse-k8s-etcd2001.codfw.wmnet with OS bookworm * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 20:57 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 20:57 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1325553{{!}}InstrumentConstructiveEdits: exclude mw-reverted as well (T431493)]] * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-etcd2001.codfw.wmnet on all recursors * 20:57 bking@cumin2003: START - Cookbook sre.dns.wipe-cache dse-k8s-etcd2001.codfw.wmnet on all recursors * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 20:57 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 20:54 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320163{{!}}Add configurable RestTermsOfServiceUrl (T428147)]] (duration: 21m 39s) * 20:53 bking@cumin2003: START - Cookbook sre.dns.netbox * 20:53 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host dse-k8s-etcd2001.codfw.wmnet * 20:50 samtar@deploy1003: samtar, milazg: Continuing with deployment * 20:35 samtar@deploy1003: samtar, milazg: Backport for [[gerrit:1320163{{!}}Add configurable RestTermsOfServiceUrl (T428147)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:33 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1320163{{!}}Add configurable RestTermsOfServiceUrl (T428147)]] * 20:30 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325549{{!}}Deploy PRV to several LC wikis (T423785)]] (duration: 06m 57s) * 20:26 arlolra@deploy1003: arlolra: Continuing with deployment * 20:25 arlolra@deploy1003: arlolra: Backport for [[gerrit:1325549{{!}}Deploy PRV to several LC wikis (T423785)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:24 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 20:23 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1325549{{!}}Deploy PRV to several LC wikis (T423785)]] * 20:21 ariel@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319906{{!}}Remove boilerplate language from wmf-rest and wmf-math API modules (T433736)]] (duration: 13m 54s) * 20:14 ariel@deploy1003: ariel: Continuing with deployment * 20:11 ariel@deploy1003: ariel: Backport for [[gerrit:1319906{{!}}Remove boilerplate language from wmf-rest and wmf-math API modules (T433736)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 ariel@deploy1003: Started scap sync-world: Backport for [[gerrit:1319906{{!}}Remove boilerplate language from wmf-rest and wmf-math API modules (T433736)]] * 19:58 robh@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-wdqs2001.codfw.wmnet with reason: updating firmware * 19:54 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325551{{!}}Render the focused module view as a full-screen page (T433896)]] (duration: 30m 37s) * 19:52 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching aqs[2001,1016]*: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 19:44 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching aqs[2001,1016]*: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 19:42 musikanimal@deploy1003: musikanimal: Continuing with deployment * 19:41 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1325551{{!}}Render the focused module view as a full-screen page (T433896)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:34 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: quash java safepoint logspam - bking@cumin2003 - [[phab:T434685|T434685]] * 19:34 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 19:34 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 19:24 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1325551{{!}}Render the focused module view as a full-screen page (T433896)]] * 19:20 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 19:20 jhancock@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin1003" * 19:18 jhancock@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin1003" * 19:14 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325548{{!}}Enable image lazy loading on desktop in group1 (T148047)]] (duration: 07m 43s) * 19:10 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 19:10 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1325548{{!}}Enable image lazy loading on desktop in group1 (T148047)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:07 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1325548{{!}}Enable image lazy loading on desktop in group1 (T148047)]] * 19:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1166.eqiad.wmnet onto db1280.eqiad.wmnet * 19:03 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 19:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1166: Pool db1166.eqiad.wmnet in after cloning * 18:59 jhancock@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 18:58 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1234.eqiad.wmnet with OS bookworm * 18:52 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 18:51 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:49 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:46 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 18:46 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 18:44 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=dns3004.* * 18:39 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1221.eqiad.wmnet with OS bookworm * 18:36 inflatador: [bking@ganeti2048] ~$ sudo gnt-instance replace-disks -n ganeti2030.codfw.wmnet aux-k8s-worker2002.codfw.wmnet [[phab:T434681|T434681]] * 18:35 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 18:29 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1220.eqiad.wmnet with OS bookworm * 18:29 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1234.eqiad.wmnet with reason: host reimage * 18:27 bking@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host dse-k8s-etcd2001.codfw.wmnet * 18:27 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-etcd2001.codfw.wmnet on all recursors * 18:27 bking@cumin2003: START - Cookbook sre.dns.wipe-cache dse-k8s-etcd2001.codfw.wmnet on all recursors * 18:27 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:27 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 18:27 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 18:25 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1234.eqiad.wmnet with reason: host reimage * 18:25 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 18:19 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1221.eqiad.wmnet with reason: host reimage * 18:17 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1166: Pool db1166.eqiad.wmnet in after cloning * 18:15 brennen@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 18:15 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1221.eqiad.wmnet with reason: host reimage * 18:14 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:12 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:12 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 18:12 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-etcd2001.codfw.wmnet on all recursors * 18:12 bking@cumin2003: START - Cookbook sre.dns.wipe-cache dse-k8s-etcd2001.codfw.wmnet on all recursors * 18:12 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:12 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 18:12 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 18:10 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: quash java safepoint logspam - bking@cumin2003 - [[phab:T434685|T434685]] * 18:10 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1234 * 18:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1234 * 18:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1220.eqiad.wmnet with reason: host reimage * 18:07 brennen: 1.47.0-wmf.15 train status ([[phab:T430834|T430834]]) - no current blockers, rolling to all wikis * 18:07 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1234 * 18:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1234.eqiad.wmnet 10.36.64.10.in-addr.arpa 0.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:07 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:07 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1234.eqiad.wmnet 10.36.64.10.in-addr.arpa 0.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1234 - ryankemper@cumin2003" * 18:07 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1234 - ryankemper@cumin2003" * 18:05 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1220.eqiad.wmnet with reason: host reimage * 18:04 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host dse-k8s-etcd2001.codfw.wmnet * 18:04 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 18:03 dancy@deploy1003: Installation of scap version "4.280.1" completed for 3 hosts * 18:02 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:02 inflatador: bking@dse-k8s-etcd2002 etcdctl member remove $<nowiki>{</nowiki>UUID of dse-k8s-etcd2001<nowiki>}</nowiki> [[phab:T434681|T434681]] [[phab:T434793|T434793]] * 18:01 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 18:01 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1234 * 18:01 dancy@deploy1003: Installing scap version "4.280.1" for 3 host(s) * 18:01 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1221 * 18:01 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1221 * 18:00 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1221 * 18:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1221.eqiad.wmnet 18.36.64.10.in-addr.arpa 8.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:00 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1221.eqiad.wmnet 18.36.64.10.in-addr.arpa 8.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1221 - ryankemper@cumin2003" * 17:59 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:58 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 17:58 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:57 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 17:57 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:56 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1221 - ryankemper@cumin2003" * 17:53 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 17:52 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 17:52 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 17:51 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1221 * 17:51 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1220 * 17:51 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1220 * 17:51 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:51 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:51 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1220 * 17:51 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1220.eqiad.wmnet 11.36.64.10.in-addr.arpa 1.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:51 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1220.eqiad.wmnet 11.36.64.10.in-addr.arpa 1.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:51 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1220 - ryankemper@cumin2003" * 17:50 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1220 - ryankemper@cumin2003" * 17:47 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:47 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 17:46 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1234.eqiad.wmnet with OS bookworm * 17:46 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 17:45 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1221.eqiad.wmnet with OS bookworm * 17:45 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1220 * 17:45 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1220.eqiad.wmnet with OS bookworm * 17:43 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 17:41 bking@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host dse-k8s-etcd2001.codfw.wmnet * 17:41 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host dse-k8s-etcd2001.codfw.wmnet * 17:40 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:40 inflatador: bking@ganeti2048] `sudo gnt-instance remove --force --ignore-failures --shutdown-timeout=0` on non-DRBD VMs [[phab:T434681|T434681]] * 17:40 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:39 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:38 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:36 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:32 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:32 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:28 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 17:26 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:26 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:24 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:23 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:21 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1218.eqiad.wmnet with OS bookworm * 17:20 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:20 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:19 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1179.eqiad.wmnet with OS bookworm * 17:18 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1150.eqiad.wmnet with OS bookworm * 17:18 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns3004.wikimedia.org with OS trixie * 17:15 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:14 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:13 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:12 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:12 swfrench@deploy1003: Finished scap sync-world: Helmfile-only deployment for mediawiki chart bump - [[phab:T427666|T427666]] (duration: 03m 03s) * 17:09 swfrench@deploy1003: Started scap sync-world: Helmfile-only deployment for mediawiki chart bump - [[phab:T427666|T427666]] * 17:02 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:01 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:01 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1218.eqiad.wmnet with reason: host reimage * 17:01 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:01 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:00 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324810{{!}}deployment-info.php: Report dbname and branch for the requested wiki (T434726)]] (duration: 06m 52s) * 16:58 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 16:57 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1179.eqiad.wmnet with reason: host reimage * 16:56 dancy@deploy1003: dancy: Continuing with deployment * 16:56 dancy@deploy1003: dancy: Backport for [[gerrit:1324810{{!}}deployment-info.php: Report dbname and branch for the requested wiki (T434726)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1150.eqiad.wmnet with reason: host reimage * 16:53 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324810{{!}}deployment-info.php: Report dbname and branch for the requested wiki (T434726)]] * 16:51 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1166: Depool db1166.eqiad.wmnet to then clone it to db1280.eqiad.wmnet - cwilliams@cumin1003 * 16:50 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1166: Depool db1166.eqiad.wmnet to then clone it to db1280.eqiad.wmnet - cwilliams@cumin1003 * 16:50 cwilliams@cumin1003: START - Cookbook sre.mysql.clone of db1166.eqiad.wmnet onto db1280.eqiad.wmnet * 16:49 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1218.eqiad.wmnet with reason: host reimage * 16:48 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1179.eqiad.wmnet with reason: host reimage * 16:47 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1150.eqiad.wmnet with reason: host reimage * 16:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.upgrade (exit_code=0) for 1 hosts * 16:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2220: Upgrade of db2220.codfw.wmnet completed * 16:38 swfrench@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 16:38 swfrench-wmf: kubectl delete node kubestagemaster2005.codfw.wmnet - [[phab:T434681|T434681]] * 16:34 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1218 * 16:34 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1218 * 16:34 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1218.eqiad.wmnet with OS bookworm * 16:34 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1179 * 16:34 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1179 * 16:33 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1179.eqiad.wmnet with OS bookworm * 16:32 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1150 * 16:32 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1150 * 16:31 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1150.eqiad.wmnet with OS bookworm * 16:24 dancy@deploy1003: Installation of scap version "4.280.0" completed for 3 hosts * 16:22 dancy@deploy1003: Installing scap version "4.280.0" for 3 host(s) * 16:20 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: quash java safepoint logspam - bking@cumin2003 - [[phab:T434685|T434685]] * 16:14 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns3004.wikimedia.org with reason: host reimage * 16:08 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 16:07 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 16:07 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns3004.wikimedia.org with reason: host reimage * 16:05 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 16:04 swfrench@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 16:00 swfrench@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 15:59 swfrench@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 15:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Upgrade of db2220.codfw.wmnet completed * 15:48 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2220: Upgrading db2220.codfw.wmnet * 15:48 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2220: Upgrading db2220.codfw.wmnet * 15:48 cwilliams@cumin1003: START - Cookbook sre.mysql.upgrade for 1 hosts * 15:46 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns3004.wikimedia.org with OS trixie * 15:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2220 [[phab:T434802|T434802]]', diff saved to https://phabricator.wikimedia.org/P96079 and previous config saved to /var/cache/conftool/dbconfig/20260813-154624-cwilliams.json * 15:45 cdobbins@cumin1003: conftool action : set/pooled=no; selector: name=dns3004.* * 15:44 cjd91: depooling dns3004 to reimage to trixie * 15:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2159 to s7 primary [[phab:T434802|T434802]]', diff saved to https://phabricator.wikimedia.org/P96078 and previous config saved to /var/cache/conftool/dbconfig/20260813-154405-cwilliams.json * 15:43 cezmunsta: Starting s7 codfw failover from db2220 to db2159 - [[phab:T434802|T434802]] * 15:41 cgoubert@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: Dragonfly supernodes reboot (duration: 09m 42s) * 15:41 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dragonfly-supernode2001.codfw.wmnet * 15:39 inflatador: bking@ganeti2048] ~$ sudo gnt-node failover -f --ignore-consistency ganeti2046.codfw.wmnet [[phab:T434681|T434681]] * 15:39 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 15:39 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 15:39 swfrench@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 15:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2159 with weight 0 [[phab:T434802|T434802]]', diff saved to https://phabricator.wikimedia.org/P96077 and previous config saved to /var/cache/conftool/dbconfig/20260813-153806-cwilliams.json * 15:37 swfrench@dns1004: END - running authdns-update * 15:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 30 hosts with reason: Primary switchover s7 [[phab:T434802|T434802]] * 15:37 cgoubert@cumin2003: START - Cookbook sre.hosts.reboot-single for host dragonfly-supernode2001.codfw.wmnet * 15:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dragonfly-supernode1001.eqiad.wmnet * 15:35 swfrench@dns1004: START - running authdns-update * 15:32 cgoubert@cumin2003: START - Cookbook sre.hosts.reboot-single for host dragonfly-supernode1001.eqiad.wmnet * 15:32 cgoubert@deploy1003: Locking from deployment [ALL REPOSITORIES]: Dragonfly supernodes reboot * 15:30 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-master-codfw * 15:30 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2005.codfw.wmnet * 15:30 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2005.codfw.wmnet * 15:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1169.eqiad.wmnet onto db1277.eqiad.wmnet * 15:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1169: Pool db1169.eqiad.wmnet in after cloning * 15:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2005.codfw.wmnet * 15:23 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2005.codfw.wmnet * 15:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2004.codfw.wmnet * 15:23 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2004.codfw.wmnet * 15:18 inflatador: bking@ganeti2048 sudo gnt-node failover -f ganeti2046.codfw.wmnet [[phab:T434681|T434681]] * 15:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2004.codfw.wmnet * 15:17 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2004.codfw.wmnet * 15:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2003.codfw.wmnet * 15:17 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2003.codfw.wmnet * 15:15 cdobbins@dns1004: END - running authdns-update * 15:13 cdobbins@dns1004: START - running authdns-update * 15:10 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1217.eqiad.wmnet with OS bookworm * 15:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2003.codfw.wmnet * 15:10 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2003.codfw.wmnet * 15:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2002.codfw.wmnet * 15:10 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2002.codfw.wmnet * 15:10 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: quash java safepoint logspam - bking@cumin2003 - [[phab:T434685|T434685]] * 15:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1216.eqiad.wmnet with OS bookworm * 15:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2002.codfw.wmnet * 15:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2002.codfw.wmnet * 15:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2001.codfw.wmnet * 15:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2001.codfw.wmnet * 15:01 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1042.eqiad.wmnet * 15:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1042.eqiad.wmnet * 15:00 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325455{{!}}Api: Use correct query when continue prop=categories (T433922)]] (duration: 09m 47s) * 14:59 bking@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host dse-k8s-etcd2004.codfw.wmnet * 14:58 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-etcd2004.codfw.wmnet on all recursors * 14:58 bking@cumin2003: START - Cookbook sre.dns.wipe-cache dse-k8s-etcd2004.codfw.wmnet on all recursors * 14:58 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:58 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM dse-k8s-etcd2004.codfw.wmnet - bking@cumin2003" * 14:58 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM dse-k8s-etcd2004.codfw.wmnet - bking@cumin2003" * 14:55 zabe@deploy1003: zabe: Continuing with deployment * 14:53 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-etcd2004.codfw.wmnet on all recursors * 14:53 bking@cumin2003: START - Cookbook sre.dns.wipe-cache dse-k8s-etcd2004.codfw.wmnet on all recursors * 14:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:53 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2004.codfw.wmnet - bking@cumin2003" * 14:53 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2004.codfw.wmnet - bking@cumin2003" * 14:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2001.codfw.wmnet * 14:52 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2001.codfw.wmnet * 14:52 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-master-codfw * 14:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2001.codfw.wmnet * 14:52 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2001.codfw.wmnet * 14:52 zabe@deploy1003: zabe: Backport for [[gerrit:1325455{{!}}Api: Use correct query when continue prop=categories (T433922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:50 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1325455{{!}}Api: Use correct query when continue prop=categories (T433922)]] * 14:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1217.eqiad.wmnet with reason: host reimage * 14:48 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:48 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host dse-k8s-etcd2004.codfw.wmnet * 14:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-master-eqiad * 14:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1006.eqiad.wmnet * 14:44 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1006.eqiad.wmnet * 14:44 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1216.eqiad.wmnet with reason: host reimage * 14:40 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1217.eqiad.wmnet with reason: host reimage * 14:39 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1169: Pool db1169.eqiad.wmnet in after cloning * 14:39 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1216.eqiad.wmnet with reason: host reimage * 14:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl1005.eqiad.wmnet * 14:32 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl1005.eqiad.wmnet * 14:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1004.eqiad.wmnet * 14:32 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1004.eqiad.wmnet * 14:31 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus2008.codfw.wmnet * 14:31 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor1003.eqiad.wmnet * 14:29 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1042.eqiad.wmnet * 14:27 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor1003.eqiad.wmnet * 14:26 moritzm: installing Django security updates * 14:26 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling reboot on A:wikidough * 14:25 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1217.eqiad.wmnet with OS bookworm * 14:25 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1216.eqiad.wmnet with OS bookworm * 14:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl1004.eqiad.wmnet * 14:24 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl1004.eqiad.wmnet * 14:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1003.eqiad.wmnet * 14:24 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1003.eqiad.wmnet * 14:23 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus2008.codfw.wmnet * 14:23 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus1008.eqiad.wmnet * 14:23 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor-dev2001.codfw.wmnet * 14:22 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus2006.codfw.wmnet * 14:19 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor-dev2001.codfw.wmnet * 14:18 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1042.eqiad.wmnet * 14:17 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.* * 14:16 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl1003.eqiad.wmnet * 14:16 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl1003.eqiad.wmnet * 14:16 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1002.eqiad.wmnet * 14:16 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1002.eqiad.wmnet * 14:15 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus1008.eqiad.wmnet * 14:14 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus1006.eqiad.wmnet * 14:12 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus2006.codfw.wmnet * 14:12 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor2003.codfw.wmnet * 14:11 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus2007.codfw.wmnet * 14:11 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1041.eqiad.wmnet * 14:11 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1041.eqiad.wmnet * 14:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl1002.eqiad.wmnet * 14:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl1002.eqiad.wmnet * 14:09 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-master-eqiad * 14:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on apifeatureusage1001.eqiad.wmnet with reason: host reimage * 14:08 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1215.eqiad.wmnet with OS bookworm * 14:08 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor2003.codfw.wmnet * 14:08 jayme: updated calico to v3.30.7 on wikikube codfw [[phab:T427400|T427400]] * 14:07 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1149.eqiad.wmnet with OS bookworm * 14:06 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'. * 14:06 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1041.eqiad.wmnet * 14:04 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus1006.eqiad.wmnet * 14:03 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus2007.codfw.wmnet * 14:03 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus2005.codfw.wmnet * 14:03 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus1007.eqiad.wmnet * 14:03 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1002.eqiad.wmnet * 14:02 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on apifeatureusage1001.eqiad.wmnet with reason: host reimage * 14:02 cgoubert@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=helm-charts.*,name=eqiad * 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host chartmuseum1001.eqiad.wmnet * 14:00 moritzm: installing libxml2 security updates * 13:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1214.eqiad.wmnet with OS bookworm * 13:59 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns5003.* * 13:59 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1041.eqiad.wmnet * 13:59 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow1002.eqiad.wmnet * 13:58 cmooney@dns3003: END - running authdns-update * 13:57 cgoubert@cumin2003: START - Cookbook sre.hosts.reboot-single for host chartmuseum1001.eqiad.wmnet * 13:57 cgoubert@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=helm-charts.*,name=eqiad * 13:57 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: cloudelastic cluster restart - bking@cumin2003 * 13:57 cgoubert@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=helm-charts.*,name=codfw * 13:56 cmooney@dns3003: START - running authdns-update * 13:56 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'. * 13:56 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns5003.*,service=authdns-update * 13:55 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host chartmuseum2001.codfw.wmnet * 13:55 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus1007.eqiad.wmnet * 13:55 cmooney@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dns5003.wikimedia.org * 13:55 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus1005.eqiad.wmnet * 13:51 cgoubert@cumin2003: START - Cookbook sre.hosts.reboot-single for host chartmuseum2001.codfw.wmnet * 13:51 cgoubert@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=helm-charts.*,name=codfw * 13:51 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus2005.codfw.wmnet * 13:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host apifeatureusage1001.eqiad.wmnet with OS bookworm * 13:50 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus7002.magru.wmnet * 13:50 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts lvs1015.eqiad.wmnet * 13:50 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:50 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1015.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:49 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1015.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:49 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325480{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] (duration: 06m 39s) * 13:47 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1215.eqiad.wmnet with reason: host reimage * 13:46 cmooney@cumin1003: START - Cookbook sre.hosts.reboot-single for host dns5003.wikimedia.org * 13:46 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1003.eqiad.wmnet * 13:45 cmooney@cumin1003: conftool action : set/pooled=no; selector: name=dns5003.* * 13:45 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1040.eqiad.wmnet * 13:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1040.eqiad.wmnet * 13:45 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 13:45 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus1005.eqiad.wmnet * 13:45 stran@deploy1003: stran: Continuing with deployment * 13:44 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus7002.magru.wmnet * 13:44 stran@deploy1003: stran: Backport for [[gerrit:1325480{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:44 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus6002.drmrs.wmnet * 13:43 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260805000000" --end-timestamp="20260806000000" --sleep="5" --batch-size="10"` for [[phab:T434688|T434688]] * 13:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1149.eqiad.wmnet with reason: host reimage * 13:42 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1325480{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] * 13:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow1003.eqiad.wmnet * 13:40 sukhe@cumin1003: START - Cookbook sre.hosts.decommission for hosts lvs1015.eqiad.wmnet * 13:40 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts lvs1014.eqiad.wmnet * 13:40 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:40 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1014.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:40 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1040.eqiad.wmnet * 13:40 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1014.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:39 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1214.eqiad.wmnet with reason: host reimage * 13:38 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2004.codfw.wmnet * 13:38 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus6002.drmrs.wmnet * 13:38 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus5003.eqsin.wmnet * 13:35 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 13:35 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1149.eqiad.wmnet with reason: host reimage * 13:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1215.eqiad.wmnet with reason: host reimage * 13:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow2004.codfw.wmnet * 13:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1214.eqiad.wmnet with reason: host reimage * 13:33 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1040.eqiad.wmnet * 13:31 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus5003.eqsin.wmnet * 13:31 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: cloudelastic cluster restart - bking@cumin2003 * 13:31 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus4003.ulsfo.wmnet * 13:30 sukhe@cumin1003: START - Cookbook sre.hosts.decommission for hosts lvs1014.eqiad.wmnet * 13:30 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts lvs1013.eqiad.wmnet * 13:30 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:30 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1013.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:30 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1013.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:27 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1039.eqiad.wmnet * 13:27 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1039.eqiad.wmnet * 13:26 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1167.eqiad.wmnet onto db1281.eqiad.wmnet * 13:26 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1167: Pool db1167.eqiad.wmnet in after cloning * 13:25 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus4003.ulsfo.wmnet * 13:25 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260802000000" --end-timestamp="20260803000000" --sleep="5" --batch-size="10"` for [[phab:T434688|T434688]] * 13:24 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 13:24 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus3004.esams.wmnet * 13:24 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1274: New host * 13:24 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325476{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] (duration: 07m 19s) * 13:24 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260804000000" --end-timestamp="20260805000000" --sleep="5" --batch-size="10"` for [[phab:T434688|T434688]] * 13:24 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260803000000" --end-timestamp="20260804000000" --sleep="5" --batch-size="10"` for [[phab:T434688|T434688]] * 13:23 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2003.codfw.wmnet * 13:22 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1039.eqiad.wmnet * 13:20 sukhe@cumin1003: START - Cookbook sre.hosts.decommission for hosts lvs1013.eqiad.wmnet * 13:20 stran@deploy1003: stran: Continuing with deployment * 13:19 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow2003.codfw.wmnet * 13:19 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling reboot on A:wikidough * 13:19 stran@deploy1003: stran: Backport for [[gerrit:1325476{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:18 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus3004.esams.wmnet * 13:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1215.eqiad.wmnet with OS bookworm * 13:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1214.eqiad.wmnet with OS bookworm * 13:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1149.eqiad.wmnet with OS bookworm * 13:17 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1325476{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] * 13:11 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324731{{!}}prv: Enable parsoid rendering for 5 wikisource wikis]] (duration: 07m 26s) * 13:11 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1039.eqiad.wmnet * 13:11 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1036.eqiad.wmnet * 13:11 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1036.eqiad.wmnet * 13:07 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow3004.esams.wmnet * 13:07 jgiannelos@deploy1003: jgiannelos: Continuing with deployment * 13:06 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1233.eqiad.wmnet with OS bookworm * 13:06 jgiannelos@deploy1003: jgiannelos: Backport for [[gerrit:1324731{{!}}prv: Enable parsoid rendering for 5 wikisource wikis]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:04 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1324731{{!}}prv: Enable parsoid rendering for 5 wikisource wikis]] * 13:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow3004.esams.wmnet * 13:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1036.eqiad.wmnet * 13:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1210.eqiad.wmnet with OS bookworm * 13:01 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1036.eqiad.wmnet * 13:00 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1052.eqiad.wmnet * 13:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1052.eqiad.wmnet * 12:58 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow4003.ulsfo.wmnet * 12:55 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1211.eqiad.wmnet with OS bookworm * 12:55 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1052.eqiad.wmnet * 12:52 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow4003.ulsfo.wmnet * 12:51 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1052.eqiad.wmnet * 12:49 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1169: Depool db1169.eqiad.wmnet to then clone it to db1277.eqiad.wmnet - cwilliams@cumin1003 * 12:46 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1169: Depool db1169.eqiad.wmnet to then clone it to db1277.eqiad.wmnet - cwilliams@cumin1003 * 12:46 cwilliams@cumin1003: START - Cookbook sre.mysql.clone of db1169.eqiad.wmnet onto db1277.eqiad.wmnet * 12:45 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:45 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns record for deleted IP reservations eqsin lvs vlan ints - cmooney@cumin1003" * 12:44 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns record for deleted IP reservations eqsin lvs vlan ints - cmooney@cumin1003" * 12:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1051.eqiad.wmnet * 12:43 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1051.eqiad.wmnet * 12:42 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1210.eqiad.wmnet with reason: host reimage * 12:41 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1167: Pool db1167.eqiad.wmnet in after cloning * 12:40 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 12:39 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1274: New host * 12:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Added db1274', diff saved to https://phabricator.wikimedia.org/P96062 and previous config saved to /var/cache/conftool/dbconfig/20260813-123907-cwilliams.json * 12:38 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast7002.wikimedia.org * 12:38 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1233.eqiad.wmnet with reason: host reimage * 12:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1051.eqiad.wmnet * 12:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1211.eqiad.wmnet with reason: host reimage * 12:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1233.eqiad.wmnet with reason: host reimage * 12:32 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1051.eqiad.wmnet * 12:32 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast7002.wikimedia.org * 12:32 marostegui: Drop SecurePoll tables from closed wikis [[phab:T423128|T423128]] * 12:32 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow5003.eqsin.wmnet * 12:31 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1050.eqiad.wmnet * 12:31 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1050.eqiad.wmnet * 12:29 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1210.eqiad.wmnet with reason: host reimage * 12:29 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1211.eqiad.wmnet with reason: host reimage * 12:29 moritzm: installing Wireshark security updates * 12:26 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow5003.eqsin.wmnet * 12:26 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1050.eqiad.wmnet * 12:24 cmooney@dns3003: END - running authdns-update * 12:21 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow6001.drmrs.wmnet * 12:21 cmooney@dns3003: START - running authdns-update * 12:20 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1050.eqiad.wmnet * 12:20 cgoubert@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host rdb-lock2003.codfw.wmnet * 12:20 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:20 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update netbox dns entries for expanded public1-603-eqsin subnet - cmooney@cumin1003" * 12:20 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update netbox dns entries for expanded public1-603-eqsin subnet - cmooney@cumin1003" * 12:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 12:20 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 12:19 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1049.eqiad.wmnet * 12:19 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1049.eqiad.wmnet * 12:19 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 12:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host moss-be1003.eqiad.wmnet * 12:17 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow6001.drmrs.wmnet * 12:17 moritzm: installin curl security updates * 12:15 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 12:15 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1233.eqiad.wmnet with OS bookworm * 12:15 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1210.eqiad.wmnet with OS bookworm * 12:15 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1211.eqiad.wmnet with OS bookworm * 12:13 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1049.eqiad.wmnet * 12:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow7002.magru.wmnet * 12:12 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-worker1178.eqiad.wmnet with OS bookworm * 12:10 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host moss-be1003.eqiad.wmnet * 12:10 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 12:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be1006.eqiad.wmnet * 12:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 12:10 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 12:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 12:10 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 12:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow7002.magru.wmnet * 12:08 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1049.eqiad.wmnet * 12:05 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 12:05 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2003.codfw.wmnet * 12:04 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'. * 12:04 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:03 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be1006.eqiad.wmnet * 12:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be1005.eqiad.wmnet * 12:01 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1038.eqiad.wmnet * 12:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1038.eqiad.wmnet * 12:01 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1213.eqiad.wmnet with OS bookworm * 11:56 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be1005.eqiad.wmnet * 11:55 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be1004.eqiad.wmnet * 11:54 cgoubert@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host rdb-lock2003.codfw.wmnet * 11:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 11:53 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 11:53 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:53 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 11:53 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 11:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1038.eqiad.wmnet * 11:49 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be1004.eqiad.wmnet * 11:44 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:44 moritzm: installing Linux 5.10.262 on Bullseye hosts * 11:41 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1038.eqiad.wmnet * 11:40 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 11:40 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 11:40 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 11:40 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:40 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 11:40 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 11:38 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1213.eqiad.wmnet with reason: host reimage * 11:36 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1167: Depool db1167.eqiad.wmnet to then clone it to db1281.eqiad.wmnet - marostegui@cumin1003 * 11:35 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 11:35 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2003.codfw.wmnet * 11:35 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1167: Depool db1167.eqiad.wmnet to then clone it to db1281.eqiad.wmnet - marostegui@cumin1003 * 11:35 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1167.eqiad.wmnet onto db1281.eqiad.wmnet * 11:34 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 22 hosts with reason: Cloning * 11:34 moritzm: remove ganeti3005 from esams03 cluster, hardware issues [[phab:T434646|T434646]] * 11:32 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1213.eqiad.wmnet with reason: host reimage * 11:28 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1034.eqiad.wmnet * 11:28 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1034.eqiad.wmnet * 11:22 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1034.eqiad.wmnet * 11:19 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1034.eqiad.wmnet * 11:17 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1213.eqiad.wmnet with OS bookworm * 11:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1165.eqiad.wmnet onto db1279.eqiad.wmnet * 11:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1165: Pool db1165.eqiad.wmnet in after cloning * 11:07 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply * 10:57 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply * 10:54 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply * 10:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1274.eqiad.wmnet with reason: Enabling notifications and pooling * 10:45 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1033.eqiad.wmnet * 10:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1033.eqiad.wmnet * 10:44 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply * 10:43 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'. * 10:42 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'. * 10:42 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'. * 10:40 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db[1216,1225,1239-1240].eqiad.wmnet with reason: reboot * 10:39 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1033.eqiad.wmnet * 10:38 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1151 from dbctl [[phab:T434538|T434538]]', diff saved to https://phabricator.wikimedia.org/P96055 and previous config saved to /var/cache/conftool/dbconfig/20260813-103828-marostegui.json * 10:35 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1033.eqiad.wmnet * 10:27 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1165: Pool db1165.eqiad.wmnet in after cloning * 10:24 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1151.eqiad.wmnet with OS bookworm * 10:15 moritzm: installing bind9 security updates (client-side tools/libs only) * 10:07 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-debug: apply * 10:06 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-debug: apply * 10:02 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 7 hosts with reason: reboot * 10:01 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-debug: apply * 10:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.upgrade (exit_code=0) for 1 hosts * 10:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2214: Upgrade of db2214.codfw.wmnet completed * 10:01 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-debug: apply * 10:00 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-debug: apply * 10:00 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-debug: apply * 09:59 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.decommission (exit_code=99) * 09:59 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 09:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1151.eqiad.wmnet with reason: host reimage * 09:59 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 8 hosts * 09:59 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 8 hosts * 09:55 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1151.eqiad.wmnet with reason: host reimage * 09:46 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260801000000" --end-timestamp="20260802000000" --sleep="3" --batch-size="5"` for [[phab:T434688|T434688]] * 09:45 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 8 hosts with reason: reboot * 09:45 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 22 hosts with reason: Cloning * 09:43 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1165: Depool db1165.eqiad.wmnet to then clone it to db1279.eqiad.wmnet - marostegui@cumin1003 * 09:42 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1165: Depool db1165.eqiad.wmnet to then clone it to db1279.eqiad.wmnet - marostegui@cumin1003 * 09:42 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1165.eqiad.wmnet onto db1279.eqiad.wmnet * 09:41 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=testwiki --start-timestamp="20260311000000" --end-timestamp="20260805000000" --sleep="5" --batch-size="2"` for [[phab:T434688|T434688]] * 09:40 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for backupmon1001.eqiad.wmnet * 09:40 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for backupmon1001.eqiad.wmnet * 09:40 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1151 * 09:40 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1151 * 09:37 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325408{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]], [[gerrit:1325407{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]] (duration: 06m 57s) * 09:36 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1151 * 09:36 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1151.eqiad.wmnet 13.36.64.10.in-addr.arpa 3.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:36 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1151.eqiad.wmnet 13.36.64.10.in-addr.arpa 3.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:36 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:36 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1151 - btullis@cumin1003" * 09:36 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on backupmon1001.eqiad.wmnet with reason: reboot * 09:36 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1151 - btullis@cumin1003" * 09:35 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host moss-be2003.codfw.wmnet * 09:33 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 7 hosts * 09:33 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 7 hosts * 09:33 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 09:32 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1325408{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]], [[gerrit:1325407{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1161.eqiad.wmnet onto db1275.eqiad.wmnet * 09:31 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 09:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1161: Pool db1161.eqiad.wmnet in after cloning * 09:30 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1325408{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]], [[gerrit:1325407{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]] * 09:29 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 09:27 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 09:27 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host moss-be2003.codfw.wmnet * 09:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be2006.codfw.wmnet * 09:25 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2004.codfw.wmnet * 09:22 hashar@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 09:21 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be2006.codfw.wmnet * 09:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be2005.codfw.wmnet * 09:19 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2004.codfw.wmnet * 09:19 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2003.codfw.wmnet * 09:18 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 7 hosts with reason: reboot * 09:18 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 7 hosts * 09:18 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 7 hosts * 09:16 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2214: Upgrade of db2214.codfw.wmnet completed * 09:14 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be2005.codfw.wmnet * 09:13 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be2004.codfw.wmnet * 09:12 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2003.codfw.wmnet * 09:12 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2002.codfw.wmnet * 09:09 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2214: Upgrading db2214.codfw.wmnet * 09:09 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2214: Upgrading db2214.codfw.wmnet * 09:09 cwilliams@cumin1003: START - Cookbook sre.mysql.upgrade for 1 hosts * 09:07 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be2004.codfw.wmnet * 09:06 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2002.codfw.wmnet * 09:05 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1004.eqiad.wmnet * 09:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-cluster (exit_code=0) * 09:03 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 7 hosts with reason: reboot * 09:01 btullis@cumin1003: START - Cookbook sre.dns.netbox * 09:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2214 [[phab:T434754|T434754]]', diff saved to https://phabricator.wikimedia.org/P96045 and previous config saved to /var/cache/conftool/dbconfig/20260813-090001-cwilliams.json * 08:59 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1004.eqiad.wmnet * 08:59 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1003.eqiad.wmnet * 08:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2229 to s6 primary [[phab:T434754|T434754]]', diff saved to https://phabricator.wikimedia.org/P96044 and previous config saved to /var/cache/conftool/dbconfig/20260813-085752-cwilliams.json * 08:57 cezmunsta: Starting s6 codfw failover from db2214 to db2229 - [[phab:T434754|T434754]] * 08:54 hashar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325416{{!}}Revert "REST: Enable `GET /lexemes/<nowiki>{</nowiki>lexeme_id<nowiki>}</nowiki>` by default" (T434712)]] (duration: 07m 22s) * 08:53 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1003.eqiad.wmnet * 08:53 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1002.eqiad.wmnet * 08:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2229 with weight 0 [[phab:T434754|T434754]]', diff saved to https://phabricator.wikimedia.org/P96043 and previous config saved to /var/cache/conftool/dbconfig/20260813-085151-cwilliams.json * 08:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 23 hosts with reason: Primary switchover s6 [[phab:T434754|T434754]] * 08:51 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts2002.codfw.wmnet * 08:50 hashar@deploy1003: hashar: Continuing with deployment * 08:49 hashar@deploy1003: hashar: Backport for [[gerrit:1325416{{!}}Revert "REST: Enable `GET /lexemes/<nowiki>{</nowiki>lexeme_id<nowiki>}</nowiki>` by default" (T434712)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:47 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet * 08:47 hashar@deploy1003: Started scap sync-world: Backport for [[gerrit:1325416{{!}}Revert "REST: Enable `GET /lexemes/<nowiki>{</nowiki>lexeme_id<nowiki>}</nowiki>` by default" (T434712)]] * 08:47 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1002.eqiad.wmnet * 08:46 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1161: Pool db1161.eqiad.wmnet in after cloning * 08:46 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host stewards2001.codfw.wmnet * 08:45 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit1003.wikimedia.org * 08:45 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host stewards1001.eqiad.wmnet * 08:45 Emperor: roll-restart apus frontends in codfw * 08:45 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-cluster * 08:44 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts2002.codfw.wmnet * 08:44 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host doc2003.codfw.wmnet * 08:43 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet * 08:43 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host phab2003.codfw.wmnet * 08:42 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host stewards2001.codfw.wmnet * 08:42 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host etherpad2002.codfw.wmnet * 08:41 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host stewards1001.eqiad.wmnet * 08:41 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1273: Pool in s7 * 08:41 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1003.wikimedia.org * 08:40 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host doc2003.codfw.wmnet * 08:40 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit1003.wikimedia.org * 08:39 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host etherpad1004.eqiad.wmnet * 08:39 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host doc1004.eqiad.wmnet * 08:39 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit2002.wikimedia.org * 08:38 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host etherpad2002.codfw.wmnet * 08:37 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host phab2003.codfw.wmnet * 08:36 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lists1004.wikimedia.org * 08:35 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host etherpad1004.eqiad.wmnet * 08:35 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host doc1004.eqiad.wmnet * 08:34 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-cluster (exit_code=0) * 08:34 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1003.wikimedia.org * 08:34 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host planet2003.codfw.wmnet * 08:34 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2003.wikimedia.org * 08:33 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host planet1003.eqiad.wmnet * 08:33 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit2002.wikimedia.org * 08:32 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 22 hosts with reason: Cloning * 08:32 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2002.wikimedia.org * 08:31 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aphlict1002.eqiad.wmnet * 08:30 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host planet2003.codfw.wmnet * 08:29 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host planet1003.eqiad.wmnet * 08:29 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host lists1004.wikimedia.org * 08:28 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lists2001.wikimedia.org * 08:28 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2003.wikimedia.org * 08:27 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host aphlict1002.eqiad.wmnet * 08:27 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aphlict2001.codfw.wmnet * 08:26 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2002.wikimedia.org * 08:23 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host aphlict2001.codfw.wmnet * 08:23 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1151 * 08:22 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1151.eqiad.wmnet with OS bookworm * 08:22 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host lists2001.wikimedia.org * 08:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1161: Depool db1161.eqiad.wmnet to then clone it to db1275.eqiad.wmnet - marostegui@cumin1003 * 08:20 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1161: Depool db1161.eqiad.wmnet to then clone it to db1275.eqiad.wmnet - marostegui@cumin1003 * 08:20 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1161.eqiad.wmnet onto db1275.eqiad.wmnet * 08:15 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-cluster * 08:15 Emperor: roll-restart apus frontends in eqiad * 07:56 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1273: Pool in s7 * 07:56 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1273 to dbctl [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P96035 and previous config saved to /var/cache/conftool/dbconfig/20260813-075611-marostegui.json * 07:38 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db1273.eqiad.wmnet with reason: Reboot * 07:31 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.sanitize-wiki (exit_code=97) Managing sanitization for wikis testwiki in section s3 * 07:24 marostegui@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis testwiki in section s3 * 07:19 jayme@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on kubestagemaster2005.codfw.wmnet with reason: downtime because of hardware failure and no DRBD * 05:42 arnaudb@dns1006: END - running authdns-update * 05:40 arnaudb@dns1006: START - running authdns-update * 05:27 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 04:06 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 02:29 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1151 * 02:29 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1151 * 02:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1219.eqiad.wmnet with OS bookworm * 02:18 ryankemper: [[phab:T434494|T434494]] `ryankemper@deploy1003:~$ echo 'https://stats.wikimedia.org/' {{!}} mwscript-k8s --attach -- purgeList.php` (default page got cached during yesterday's `an-web1001` reimage) * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 46s) * 02:03 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1219.eqiad.wmnet with reason: host reimage * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 02:00 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1219.eqiad.wmnet with reason: host reimage * 01:46 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1219.eqiad.wmnet with OS bookworm * 01:03 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1142.eqiad.wmnet with OS bookworm * 00:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1142.eqiad.wmnet with reason: host reimage * 00:34 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1142.eqiad.wmnet with reason: host reimage * 00:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1142.eqiad.wmnet with OS bookworm * 00:16 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1178 * 00:16 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1178 == 2026-08-12 == * 23:18 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324409{{!}}Remove $wmg = $wg hacks in Collection (T119117)]] (duration: 06m 43s) * 23:14 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 23:13 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1324409{{!}}Remove $wmg = $wg hacks in Collection (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:11 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1324409{{!}}Remove $wmg = $wg hacks in Collection (T119117)]] * 22:59 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 22:49 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260801000000" --end-timestamp="20260802000000" --sleep=2 --batch-size=10` * 22:45 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=testwiki --start-timestamp="20200801010101" --end-timestamp="20260816010101" --sleep=15 --batch-size=5` * 22:40 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=testwiki --start-timestamp="20200101010101" --end-timestamp="20260816010101" --sleep=60` * 22:35 Dreamy_Jazz: Running `mwscript WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260101000000" --end-timestamp="20260102000000" --sleep=10` * 22:21 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324817{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]], [[gerrit:1324818{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]] (duration: 45m 29s) * 22:17 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 21:59 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 21:40 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1324817{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]], [[gerrit:1324818{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:39 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 21:36 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1324817{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]], [[gerrit:1324818{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]] * 21:32 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:31 vriley@cumin1003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:30 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:30 vriley@cumin1003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:17 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:14 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:14 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:11 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:10 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-codfw: Set storage compatability to NONE — [[phab:T433028|T433028]] - eevans@cumin1003 * 21:10 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1006 * 21:09 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1006 * 21:05 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324277{{!}}Improve Math preference labels for SVG/MathJax/MathML (T433891)]] (duration: 31m 42s) * 20:58 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1178.eqiad.wmnet with OS bookworm * 20:54 krinkle@deploy1003: krinkle: Continuing with deployment * 20:51 krinkle@deploy1003: krinkle: Backport for [[gerrit:1324277{{!}}Improve Math preference labels for SVG/MathJax/MathML (T433891)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:39 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 20:38 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:38 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp5022.eqsin.wmnet with OS trixie * 20:38 cdobbins@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - cdobbins@cumin1003" * 20:37 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:36 cdobbins@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - cdobbins@cumin1003" * 20:34 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1178.eqiad.wmnet with reason: host reimage * 20:34 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1324277{{!}}Improve Math preference labels for SVG/MathJax/MathML (T433891)]] * 20:33 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 20:28 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1178.eqiad.wmnet with reason: host reimage * 20:13 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1178.eqiad.wmnet with OS bookworm * 20:11 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-worker1178.eqiad.wmnet with OS bookworm * 20:11 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1178.eqiad.wmnet with OS bookworm * 20:09 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-codfw: Set storage compatability to NONE — [[phab:T433028|T433028]] - eevans@cumin1003 * 20:09 Dreamy_Jazz: Evening UTC backport window done * 20:08 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324794{{!}}WikimediaAntiAbuse: Enable logging channel (T431292)]] (duration: 06m 48s) * 20:08 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage * 20:05 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage * 20:04 dreamyjazz@deploy1003: kharlan, dreamyjazz: Continuing with deployment * 20:04 dreamyjazz@deploy1003: kharlan, dreamyjazz: Backport for [[gerrit:1324794{{!}}WikimediaAntiAbuse: Enable logging channel (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:01 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1324794{{!}}WikimediaAntiAbuse: Enable logging channel (T431292)]] * 19:55 brennen@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 19:47 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 19:35 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 19:35 cdobbins@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cp5022.eqsin.wmnet with OS trixie * 19:32 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:30 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:29 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:26 vriley@cumin1003: START - Cookbook sre.dns.netbox * 19:23 brennen@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 19:19 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-eqiad: Set storage compatability to NONE — [[phab:T433028|T433028]] - eevans@cumin1003 * 19:10 brennen@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 19:09 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 19:09 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 19:00 Amir1: data migrated on wikishared ([[phab:T426102|T426102]]) * 18:57 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321224{{!}}Rename ce_worklist_articles table to ce_invitation_list_articles (T426102)]] (duration: 06m 50s) * 18:53 ladsgroup@deploy1003: ladsgroup, daimona: Continuing with deployment * 18:53 ladsgroup@deploy1003: ladsgroup, daimona: Backport for [[gerrit:1321224{{!}}Rename ce_worklist_articles table to ce_invitation_list_articles (T426102)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:51 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1321224{{!}}Rename ce_worklist_articles table to ce_invitation_list_articles (T426102)]] * 18:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2192: Security update * 18:31 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 18:25 Amir1: ce_invitation_list_articles created as empty on wikishared ([[phab:T426102|T426102]]) * 18:21 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 18:21 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 18:18 Amir1: migrated testwiki entries from ce_worklist_articles to ce_invitation_list_articles ([[phab:T426102|T426102]]) * 18:18 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-eqiad: Set storage compatability to NONE — [[phab:T433028|T433028]] - eevans@cumin1003 * 18:11 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 18:08 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling reboot on A:durum-eqsin and A:durum * 18:07 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Set storage compatability to UPGRADING — [[phab:T433028|T433028]] - eevans@cumin1003 * 18:05 jhancock@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022'] * 17:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2192: Security update * 17:55 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:55 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum-eqsin and A:durum * 17:53 jhancock@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['cp5022'] * 17:47 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:47 jhancock@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['cp5022'] * 17:42 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:41 jhancock@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['cp5022'] * 17:36 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:35 jhancock@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022'] * 17:31 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-magru and not (P<nowiki>{</nowiki>cp7001*<nowiki>}</nowiki> or P<nowiki>{</nowiki>cp7009*<nowiki>}</nowiki>) and A:cp - 9.2.15 upgrade ([[phab:T434620|T434620]]) * 17:28 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:28 jhancock@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['cp5022'] * 17:22 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2192.codfw.wmnet with reason: Maintenance * 17:11 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host apifeatureusage2001.codfw.wmnet with OS bookworm * 17:04 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: apply * 17:03 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-main: apply * 17:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2192 [[phab:T434635|T434635]]', diff saved to https://phabricator.wikimedia.org/P96030 and previous config saved to /var/cache/conftool/dbconfig/20260812-170338-cwilliams.json * 17:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2213 to s5 primary [[phab:T434635|T434635]]', diff saved to https://phabricator.wikimedia.org/P96029 and previous config saved to /var/cache/conftool/dbconfig/20260812-170152-cwilliams.json * 17:01 cezmunsta: Starting s5 codfw failover from db2192 to db2213 - [[phab:T434635|T434635]] * 16:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2213 with weight 0 [[phab:T434635|T434635]]', diff saved to https://phabricator.wikimedia.org/P96028 and previous config saved to /var/cache/conftool/dbconfig/20260812-165544-cwilliams.json * 16:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 27 hosts with reason: Primary switchover s5 [[phab:T434635|T434635]] * 16:53 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: apply * 16:52 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-main: apply * 16:44 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-main: apply * 16:44 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-main: apply * 16:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1272: New host * 16:40 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324765{{!}}WikimediaAntiAbuse: Enable PersonalInfoFlagNotifications (T431292)]] (duration: 07m 02s) * 16:40 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-logging-external: apply * 16:39 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-logging-external: apply * 16:38 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-logging-external: apply * 16:37 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-logging-external: apply * 16:36 kharlan@deploy1003: kharlan: Continuing with deployment * 16:35 kharlan@deploy1003: kharlan: Backport for [[gerrit:1324765{{!}}WikimediaAntiAbuse: Enable PersonalInfoFlagNotifications (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:33 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1324765{{!}}WikimediaAntiAbuse: Enable PersonalInfoFlagNotifications (T431292)]] * 16:27 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-logging-external: apply * 16:27 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-logging-external: apply * 16:18 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 16:18 jhancock@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022'] * 16:17 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 16:16 jhancock@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['cp5022'] * 16:15 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 16:12 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Set storage compatability to UPGRADING — [[phab:T433028|T433028]] - eevans@cumin1003 * 16:10 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: apply * 16:10 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: apply * 16:08 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: apply * 16:08 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: apply * 16:08 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics: apply * 16:07 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics: apply * 16:02 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-magru and not (P<nowiki>{</nowiki>cp7001*<nowiki>}</nowiki> or P<nowiki>{</nowiki>cp7009*<nowiki>}</nowiki>) and A:cp - 9.2.15 upgrade ([[phab:T434620|T434620]]) * 15:57 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1272: New host * 15:52 jmm@cumin2003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti2046.codfw.wmnet * 15:52 jmm@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host ganeti2046.codfw.wmnet * 15:42 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:37 mforns@deploy1003: Finished deploy [analytics/refinery@49c336c] (thin): Regular analytics weekly train THIN [analytics/refinery@49c336cd] (duration: 01m 59s) * 15:37 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324733{{!}}WikimediaAntiAbuse: Enable personal info tag display on enwiki (T431292)]] (duration: 08m 12s) * 15:35 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:35 mforns@deploy1003: Started deploy [analytics/refinery@49c336c] (thin): Regular analytics weekly train THIN [analytics/refinery@49c336cd] * 15:34 mforns@deploy1003: Finished deploy [analytics/refinery@49c336c]: Regular analytics weekly train [analytics/refinery@49c336cd] (duration: 04m 20s) * 15:33 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[2024,1031]*.wmnet: Set storage compatability to UPGRADING — [[phab:T433028|T433028]] - eevans@cumin1003 * 15:33 dreamyjazz@deploy1003: kharlan, dreamyjazz: Continuing with deployment * 15:31 dreamyjazz@deploy1003: kharlan, dreamyjazz: Backport for [[gerrit:1324733{{!}}WikimediaAntiAbuse: Enable personal info tag display on enwiki (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:30 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitize-wiki (exit_code=99) Checking sanitization for wikis testwiki in section s3 * 15:30 mforns@deploy1003: Started deploy [analytics/refinery@49c336c]: Regular analytics weekly train [analytics/refinery@49c336cd] * 15:30 mforns@deploy1003: Finished deploy [analytics/refinery@49c336c] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@49c336cd] (duration: 00m 32s) * 15:29 mforns@deploy1003: Started deploy [analytics/refinery@49c336c] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@49c336cd] * 15:29 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1324733{{!}}WikimediaAntiAbuse: Enable personal info tag display on enwiki (T431292)]] * 15:27 brennen@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324700{{!}}EventDetailsParticipantsModule: populate cache with non-local users (T434597)]] (duration: 06m 38s) * 15:23 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[2024,1031]*.wmnet: Set storage compatability to UPGRADING — [[phab:T433028|T433028]] - eevans@cumin1003 * 15:23 brennen@deploy1003: brennen, daimona: Continuing with deployment * 15:23 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:22 brennen@deploy1003: brennen, daimona: Backport for [[gerrit:1324700{{!}}EventDetailsParticipantsModule: populate cache with non-local users (T434597)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:22 cgoubert@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host rdb-lock2003.codfw.wmnet * 15:21 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 15:21 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 15:21 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:21 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:21 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:20 brennen@deploy1003: Started scap sync-world: Backport for [[gerrit:1324700{{!}}EventDetailsParticipantsModule: populate cache with non-local users (T434597)]] * 15:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1178.eqiad.wmnet with OS bookworm * 15:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock1003.eqiad.wmnet * 15:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock1003.eqiad.wmnet with OS trixie * 15:16 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 15:16 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 15:16 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 15:16 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:16 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:16 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:12 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 15:12 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2003.codfw.wmnet * 15:11 cgoubert@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host rdb-lock2003.codfw.wmnet * 15:11 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 15:11 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 15:11 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:11 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:11 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:07 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324338{{!}}InitialiseSettings: Enable 2FA warnings on more private wikis (T428103)]] (duration: 07m 02s) * 15:04 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 15:04 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock1003.eqiad.wmnet with reason: host reimage * 15:03 reedy@deploy1003: reedy: Continuing with deployment * 15:02 reedy@deploy1003: reedy: Backport for [[gerrit:1324338{{!}}InitialiseSettings: Enable 2FA warnings on more private wikis (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:02 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:00 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324338{{!}}InitialiseSettings: Enable 2FA warnings on more private wikis (T428103)]] * 14:57 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 14:57 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2003.codfw.wmnet * 14:57 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock1003.eqiad.wmnet with reason: host reimage * 14:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock2002.codfw.wmnet * 14:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock2002.codfw.wmnet with OS trixie * 14:56 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 14:56 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 14:56 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 14:55 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 14:55 moritzm: powercycle ganeti2046 * 14:47 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock1003.eqiad.wmnet with OS trixie * 14:46 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1003.eqiad.wmnet - cgoubert@cumin2003" * 14:46 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1003.eqiad.wmnet - cgoubert@cumin2003" * 14:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock1003.eqiad.wmnet on all recursors * 14:45 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock1003.eqiad.wmnet on all recursors * 14:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1003.eqiad.wmnet - cgoubert@cumin2003" * 14:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1272.eqiad.wmnet with reason: Enabling notifications * 14:44 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1003.eqiad.wmnet - cgoubert@cumin2003" * 14:44 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324719{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324720{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324722{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0 (T434187)]] (duration: 11m 02s) * 14:39 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 14:39 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock1003.eqiad.wmnet * 14:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock2002.codfw.wmnet with reason: host reimage * 14:37 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 14:37 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock1002.eqiad.wmnet * 14:37 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock1002.eqiad.wmnet with OS trixie * 14:37 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1324719{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324720{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324722{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0 (T434187)]] synced to the testservers (see https://wikitech. * 14:33 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock2002.codfw.wmnet with reason: host reimage * 14:33 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1324719{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324720{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324722{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0 (T434187)]] * 14:32 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1018.eqiad.wmnet with OS bookworm * 14:32 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2046.codfw.wmnet * 14:31 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1020.eqiad.wmnet with OS bookworm * 14:27 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2046.codfw.wmnet * 14:25 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2045.codfw.wmnet * 14:25 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2045.codfw.wmnet * 14:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock1002.eqiad.wmnet with reason: host reimage * 14:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1019.eqiad.wmnet with OS bookworm * 14:22 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Checking sanitization for wikis testwiki in section s3 * 14:20 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2045.codfw.wmnet * 14:18 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock1002.eqiad.wmnet with reason: host reimage * 14:17 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:16 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324707{{!}}Backport all changes from wmf/1.47.0-wmf.15]] (duration: 40m 51s) * 14:16 moritzm: installing Linux 6.1.180 on Bookworm hosts * 14:15 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock2002.codfw.wmnet with OS trixie * 14:14 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2002.codfw.wmnet - cgoubert@cumin2003" * 14:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2002.codfw.wmnet - cgoubert@cumin2003" * 14:14 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:14 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2045.codfw.wmnet * 14:14 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2002.codfw.wmnet on all recursors * 14:14 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2002.codfw.wmnet on all recursors * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2002.codfw.wmnet - cgoubert@cumin2003" * 14:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2002.codfw.wmnet - cgoubert@cumin2003" * 14:12 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2030.codfw.wmnet * 14:12 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2030.codfw.wmnet * 14:11 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on apifeatureusage2001.codfw.wmnet with reason: host reimage * 14:09 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:08 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:07 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 14:06 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2030.codfw.wmnet * 14:06 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:06 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock1002.eqiad.wmnet with OS trixie * 14:05 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 14:05 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2002.codfw.wmnet * 14:05 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1002.eqiad.wmnet - cgoubert@cumin2003" * 14:05 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1002.eqiad.wmnet - cgoubert@cumin2003" * 14:05 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock1002.eqiad.wmnet on all recursors * 14:05 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock1002.eqiad.wmnet on all recursors * 14:05 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:05 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1002.eqiad.wmnet - cgoubert@cumin2003" * 14:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:04 kharlan@deploy1003: kharlan: Continuing with deployment * 14:04 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock2001.codfw.wmnet * 14:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock2001.codfw.wmnet with OS trixie * 14:02 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1002.eqiad.wmnet - cgoubert@cumin2003" * 14:02 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on apifeatureusage2001.codfw.wmnet with reason: host reimage * 14:01 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2030.codfw.wmnet * 13:59 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2029.codfw.wmnet * 13:58 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2029.codfw.wmnet * 13:58 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 13:58 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock1002.eqiad.wmnet * 13:56 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1019.eqiad.wmnet with reason: host reimage * 13:54 btullis@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on archiva1002.wikimedia.org with reason: Upgrading in-place * 13:53 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock1001.eqiad.wmnet * 13:53 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock1001.eqiad.wmnet with OS trixie * 13:53 kharlan@deploy1003: kharlan: Backport for [[gerrit:1324707{{!}}Backport all changes from wmf/1.47.0-wmf.15]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:52 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2029.codfw.wmnet * 13:52 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1020.eqiad.wmnet with reason: host reimage * 13:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1019.eqiad.wmnet with reason: host reimage * 13:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1020.eqiad.wmnet with reason: host reimage * 13:49 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock2001.codfw.wmnet with reason: host reimage * 13:48 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2029.codfw.wmnet * 13:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Configuring db1272 for s3 pooling', diff saved to https://phabricator.wikimedia.org/P96021 and previous config saved to /var/cache/conftool/dbconfig/20260812-134732-cwilliams.json * 13:44 bking@cumin2003: START - Cookbook sre.hosts.reimage for host apifeatureusage2001.codfw.wmnet with OS bookworm * 13:43 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock2001.codfw.wmnet with reason: host reimage * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2028.codfw.wmnet * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2028.codfw.wmnet * 13:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1018.eqiad.wmnet with reason: host reimage * 13:38 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock1001.eqiad.wmnet with reason: host reimage * 13:36 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1018.eqiad.wmnet with reason: host reimage * 13:35 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1324707{{!}}Backport all changes from wmf/1.47.0-wmf.15]] * 13:35 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2028.codfw.wmnet * 13:32 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1275938{{!}}Enable campaignEvents on bdwikimedia (T424016)]] (duration: 07m 35s) * 13:32 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1020.eqiad.wmnet with OS bookworm * 13:32 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1019.eqiad.wmnet with OS bookworm * 13:32 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock1001.eqiad.wmnet with reason: host reimage * 13:28 kharlan@deploy1003: kharlan, yahya: Continuing with deployment * 13:27 kharlan@deploy1003: kharlan, yahya: Backport for [[gerrit:1275938{{!}}Enable campaignEvents on bdwikimedia (T424016)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:27 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2028.codfw.wmnet * 13:26 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 13:26 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 13:25 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1275938{{!}}Enable campaignEvents on bdwikimedia (T424016)]] * 13:25 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitize-wiki (exit_code=99) Managing sanitization for wikis testwiki in section s3 * 13:24 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock2001.codfw.wmnet with OS trixie * 13:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2001.codfw.wmnet - cgoubert@cumin2003" * 13:24 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2001.codfw.wmnet - cgoubert@cumin2003" * 13:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2001.codfw.wmnet on all recursors * 13:23 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2001.codfw.wmnet on all recursors * 13:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2001.codfw.wmnet - cgoubert@cumin2003" * 13:23 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2001.codfw.wmnet - cgoubert@cumin2003" * 13:23 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324713{{!}}thwiki: reinstate temporary wiki25 logos (T431094)]] (duration: 07m 13s) * 13:20 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:20 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1018.eqiad.wmnet with OS bookworm * 13:20 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2027.codfw.wmnet * 13:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2027.codfw.wmnet * 13:19 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics-external: apply * 13:18 kharlan@deploy1003: anzx, kharlan: Continuing with deployment * 13:18 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:18 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock1001.eqiad.wmnet with OS trixie * 13:17 kharlan@deploy1003: anzx, kharlan: Backport for [[gerrit:1324713{{!}}thwiki: reinstate temporary wiki25 logos (T431094)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1001.eqiad.wmnet - cgoubert@cumin2003" * 13:17 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 13:17 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1001.eqiad.wmnet - cgoubert@cumin2003" * 13:17 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2001.codfw.wmnet * 13:17 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics-external: apply * 13:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock1001.eqiad.wmnet on all recursors * 13:17 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock1001.eqiad.wmnet on all recursors * 13:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1001.eqiad.wmnet - cgoubert@cumin2003" * 13:17 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1001.eqiad.wmnet - cgoubert@cumin2003" * 13:15 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1324713{{!}}thwiki: reinstate temporary wiki25 logos (T431094)]] * 13:15 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:15 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics-external: apply * 13:15 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:14 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics-external: apply * 13:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2027.codfw.wmnet * 13:12 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 13:12 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock1001.eqiad.wmnet * 13:12 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2027.codfw.wmnet * 13:08 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1158.eqiad.wmnet onto db1273.eqiad.wmnet * 13:07 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1158: Pool db1158.eqiad.wmnet in after cloning * 13:02 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti6002.drmrs.wmnet * 13:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti6002.drmrs.wmnet * 12:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1015.eqiad.wmnet with OS bookworm * 12:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti6002.drmrs.wmnet * 12:44 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1017.eqiad.wmnet with OS bookworm * 12:36 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 12:35 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 12:34 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 12:33 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 12:31 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti6002.drmrs.wmnet * 12:24 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1017.eqiad.wmnet with reason: host reimage * 12:22 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1158: Pool db1158.eqiad.wmnet in after cloning * 12:18 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1017.eqiad.wmnet with reason: host reimage * 12:11 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1015.eqiad.wmnet with reason: host reimage * 12:07 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1015.eqiad.wmnet with reason: host reimage * 12:04 moritzm: failover ganeti master in drmrs02 to ganeti6004 * 12:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1009.eqiad.wmnet with OS bookworm * 12:01 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1017.eqiad.wmnet with OS bookworm * 12:00 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti6004.drmrs.wmnet * 12:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti6004.drmrs.wmnet * 11:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti6004.drmrs.wmnet * 11:53 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1015.eqiad.wmnet with OS bookworm * 11:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1159.eqiad.wmnet onto db1274.eqiad.wmnet * 11:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Pool db1159.eqiad.wmnet in after cloning * 11:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1016.eqiad.wmnet with OS bookworm * 11:46 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti6004.drmrs.wmnet * 11:45 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti6001.drmrs.wmnet * 11:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti6001.drmrs.wmnet * 11:43 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324306{{!}}WikimediaAntiAbuse: Enable personal info for enwiki with no display (T431292)]] (duration: 10m 26s) * 11:42 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1015.eqiad.wmnet with OS bookworm * 11:39 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 11:38 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti6001.drmrs.wmnet * 11:34 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1324306{{!}}WikimediaAntiAbuse: Enable personal info for enwiki with no display (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:33 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2212: Security update * 11:33 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti6001.drmrs.wmnet * 11:32 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1324306{{!}}WikimediaAntiAbuse: Enable personal info for enwiki with no display (T431292)]] * 11:22 moritzm: failover ganeti master in drmrs01 to ganeti6003 * 11:20 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:20 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:18 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 22 hosts with reason: Cloning * 11:17 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti6003.drmrs.wmnet * 11:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti6003.drmrs.wmnet * 11:17 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1009.eqiad.wmnet with reason: host reimage * 11:17 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:16 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:14 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1016.eqiad.wmnet with reason: host reimage * 11:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti6003.drmrs.wmnet * 11:11 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1009.eqiad.wmnet with reason: host reimage * 11:10 jmm@cumin2003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti3005.esams.wmnet * 11:10 jmm@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host ganeti3005.esams.wmnet * 11:09 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 11:08 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1158: Depool db1158.eqiad.wmnet to then clone it to db1273.eqiad.wmnet - marostegui@cumin1003 * 11:07 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1016.eqiad.wmnet with reason: host reimage * 11:07 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1158: Depool db1158.eqiad.wmnet to then clone it to db1273.eqiad.wmnet - marostegui@cumin1003 * 11:07 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1158.eqiad.wmnet onto db1273.eqiad.wmnet * 11:06 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:05 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:05 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1159: Pool db1159.eqiad.wmnet in after cloning * 11:04 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 20 hosts with reason: Cloning * 11:02 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti6003.drmrs.wmnet * 11:00 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis testwiki in section s3 * 10:54 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1009.eqiad.wmnet with OS bookworm * 10:51 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1015.eqiad.wmnet with OS bookworm * 10:50 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1016.eqiad.wmnet with OS bookworm * 10:48 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2212: Security update * 10:45 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitize-wiki (exit_code=99) Managing sanitization for wikis testwiki in section s3 * 10:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1013.eqiad.wmnet with OS bookworm * 10:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1014.eqiad.wmnet with OS bookworm * 10:38 cwilliams@cumin1003: START - Cookbook sre.mysql.clone of db1159.eqiad.wmnet onto db1274.eqiad.wmnet * 10:33 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db1274.eqiad.wmnet * 10:33 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db1274.eqiad.wmnet * 10:31 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 396993 * 10:29 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 396993 * 10:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1159: Clone source for db1274 * 10:24 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1159: Clone source for db1274 * 10:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1013.eqiad.wmnet with reason: host reimage * 10:18 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1013.eqiad.wmnet with reason: host reimage * 10:13 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2212.codfw.wmnet with reason: Maintenance * 10:13 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 10:12 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 10:12 blake@deploy1003: Stopping before sync operations * 10:11 blake@deploy1003: Started scap sync-world: Non-deployment scap run to populate new release values for [[phab:T427668|T427668]] * 10:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2212 [[phab:T434644|T434644]]', diff saved to https://phabricator.wikimedia.org/P96003 and previous config saved to /var/cache/conftool/dbconfig/20260812-101053-cwilliams.json * 10:09 moritzm: powercycle ganeti3005 * 10:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2203 to s1 primary [[phab:T434644|T434644]]', diff saved to https://phabricator.wikimedia.org/P96002 and previous config saved to /var/cache/conftool/dbconfig/20260812-100849-cwilliams.json * 10:08 cezmunsta: Starting s1 codfw failover from db2212 to db2203 - [[phab:T434644|T434644]] * 10:03 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1013.eqiad.wmnet with OS bookworm * 10:02 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1013.eqiad.wmnet with OS bookworm * 10:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2203 with weight 0 [[phab:T434644|T434644]]', diff saved to https://phabricator.wikimedia.org/P96001 and previous config saved to /var/cache/conftool/dbconfig/20260812-100134-cwilliams.json * 10:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 32 hosts with reason: Primary switchover s1 [[phab:T434644|T434644]] * 09:53 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1014.eqiad.wmnet with reason: host reimage * 09:50 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 09:50 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti3005.esams.wmnet * 09:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1014.eqiad.wmnet with reason: host reimage * 09:41 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1278: Pool in x1 * 09:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-launcher1003.eqiad.wmnet with OS bookworm * 09:37 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti3005.esams.wmnet * 09:37 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1013.eqiad.wmnet with OS bookworm * 09:34 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-presto1013.eqiad.wmnet with OS bookworm * 09:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1014.eqiad.wmnet with OS bookworm * 09:29 moritzm: failover ganeti master in esams to ganeti3008 * 09:26 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti3006.esams.wmnet * 09:26 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti3006.esams.wmnet * 09:24 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1012.eqiad.wmnet with OS bookworm * 09:23 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1009.eqiad.wmnet with OS bookworm * 09:23 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 09:18 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti3006.esams.wmnet * 09:16 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti3006.esams.wmnet * 09:14 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: fix regexp escaping bug - oblivian@cumin1003" * 09:14 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: fix regexp escaping bug - oblivian@cumin1003 * 09:13 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: fix regexp escaping bug - oblivian@cumin1003 * 09:13 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: fix regexp escaping bug - oblivian@cumin1003" * 09:03 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-launcher1003.eqiad.wmnet with reason: host reimage * 08:58 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-launcher1003.eqiad.wmnet with reason: host reimage * 08:55 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1278: Pool in x1 * 08:55 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1278 to dbctl [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95996 and previous config saved to /var/cache/conftool/dbconfig/20260812-085521-marostegui.json * 08:51 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1012.eqiad.wmnet with reason: host reimage * 08:45 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis testwiki in section s3 * 08:43 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1009.eqiad.wmnet with OS bookworm * 08:42 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1012.eqiad.wmnet with reason: host reimage * 08:41 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-launcher1003.eqiad.wmnet with OS bookworm * 08:40 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1013.eqiad.wmnet with OS bookworm * 08:38 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms3', diff saved to https://phabricator.wikimedia.org/P95995 and previous config saved to /var/cache/conftool/dbconfig/20260812-083816-marostegui.json * 08:38 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-master1003.eqiad.wmnet with OS bookworm * 08:37 marostegui: Failover ms3 [[phab:T434288|T434288]] * 08:37 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1268 to dbctl [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95994 and previous config saved to /var/cache/conftool/dbconfig/20260812-083722-marostegui.json * 08:35 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-web1001.eqiad.wmnet with OS bookworm * 08:32 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db2252.codfw.wmnet,db[1153,1268].eqiad.wmnet with reason: Switching over ms3 * 08:28 marostegui@cumin1003: dbctl commit (dc=all): 'Depool ms3 [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95993 and previous config saved to /var/cache/conftool/dbconfig/20260812-082852-marostegui.json * 08:25 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1012.eqiad.wmnet with OS bookworm * 08:25 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1011.eqiad.wmnet with OS bookworm * 08:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-master1003.eqiad.wmnet with reason: host reimage * 08:07 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-master1003.eqiad.wmnet with reason: host reimage * 08:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-web1001.eqiad.wmnet with reason: host reimage * 07:58 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-web1001.eqiad.wmnet with reason: host reimage * 07:50 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1003.eqiad.wmnet with OS bookworm * 07:38 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1011.eqiad.wmnet with reason: host reimage * 07:38 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-master1003.eqiad.wmnet with OS bookworm * 07:35 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti3007.esams.wmnet * 07:35 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti3007.esams.wmnet * 07:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1011.eqiad.wmnet with reason: host reimage * 07:27 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti3007.esams.wmnet * 07:25 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti3007.esams.wmnet * 07:25 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti3008.esams.wmnet * 07:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti3008.esams.wmnet * 07:22 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-web1001.eqiad.wmnet with OS bookworm * 07:19 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1003.eqiad.wmnet with OS bookworm * 07:18 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-master1003.eqiad.wmnet * 07:18 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host an-master1003.eqiad.wmnet * 07:17 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1011.eqiad.wmnet with OS bookworm * 07:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 07:16 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti3008.esams.wmnet * 07:15 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1009.eqiad.wmnet with OS bookworm * 07:14 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host an-master1003.eqiad.wmnet * 07:13 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-master1003.eqiad.wmnet * 07:13 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-master1003.eqiad.wmnet * 07:12 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-master1003.eqiad.wmnet * 07:11 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti3008.esams.wmnet * 07:07 arnaudb@dns1006: END - running authdns-update * 07:07 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti5007.eqsin.wmnet * 07:07 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti5007.eqsin.wmnet * 07:05 arnaudb@dns1006: START - running authdns-update * 06:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti5007.eqsin.wmnet * 06:54 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti5007.eqsin.wmnet * 06:38 moritzm: failover ganeti master in eqsin to ganeti5004 * 06:36 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti5006.eqsin.wmnet * 06:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti5006.eqsin.wmnet * 06:28 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti5006.eqsin.wmnet * 06:23 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti5006.eqsin.wmnet * 06:20 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti5005.eqsin.wmnet * 06:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti5005.eqsin.wmnet * 06:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti5005.eqsin.wmnet * 06:06 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti5005.eqsin.wmnet * 06:03 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti5004.eqsin.wmnet * 06:03 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti5004.eqsin.wmnet * 05:55 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti5004.eqsin.wmnet * 05:53 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti5004.eqsin.wmnet * 04:40 ryankemper: [[phab:T434494|T434494]] reimaged `an-tool1008.eqiad.wmnet` to bookworm; yarn.wikimedia.org is back up * 04:16 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-tool1008.eqiad.wmnet with OS bookworm * 03:58 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-tool1008.eqiad.wmnet with reason: host reimage * 03:53 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-tool1008.eqiad.wmnet with reason: host reimage * 03:41 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-tool1008.eqiad.wmnet with OS bookworm * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 45s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 00:25 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324427{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]], [[gerrit:1324429{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0]], [[gerrit:1324428{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]] (duration: 07m 55s) * 00:21 kemayo@deploy1003: kemayo: Continuing with deployment * 00:19 kemayo@deploy1003: kemayo: Backport for [[gerrit:1324427{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]], [[gerrit:1324429{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0]], [[gerrit:1324428{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:17 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1324427{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]], [[gerrit:1324429{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0]], [[gerrit:1324428{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]] == 2026-08-11 == * 21:37 sbassett: Deployed security fix for [[phab:T434521|T434521]] (wmf.15) * 21:29 sbassett: Deployed security fix for [[phab:T434521|T434521]] (wmf.14) * 21:19 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324370{{!}}Phase 4 of legal footer deployment (T432796)]], [[gerrit:1319804{{!}}Disable wgMFCustomSiteModules on English Wikipedia (T375538)]] (duration: 15m 26s) * 21:15 jdlrobson@deploy1003: jdlrobson: Continuing with deployment * 21:06 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1324370{{!}}Phase 4 of legal footer deployment (T432796)]], [[gerrit:1319804{{!}}Disable wgMFCustomSiteModules on English Wikipedia (T375538)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:03 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1324370{{!}}Phase 4 of legal footer deployment (T432796)]], [[gerrit:1319804{{!}}Disable wgMFCustomSiteModules on English Wikipedia (T375538)]] * 20:59 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 20:50 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324384{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]], [[gerrit:1324385{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]] (duration: 06m 58s) * 20:46 kemayo@deploy1003: kemayo: Continuing with deployment * 20:45 kemayo@deploy1003: kemayo: Backport for [[gerrit:1324384{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]], [[gerrit:1324385{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:43 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1324384{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]], [[gerrit:1324385{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]] * 20:42 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 20:42 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324386{{!}}build: Updating js-yaml to 3.15.1, 4.3.1]] (duration: 07m 36s) * 20:38 kemayo@deploy1003: kemayo: Continuing with deployment * 20:37 jhancock@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 20:37 kemayo@deploy1003: kemayo: Backport for [[gerrit:1324386{{!}}build: Updating js-yaml to 3.15.1, 4.3.1]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:35 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1324386{{!}}build: Updating js-yaml to 3.15.1, 4.3.1]] * 20:18 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 20:15 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 20:15 jhancock@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin1003" * 20:14 jhancock@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin1003" * 19:59 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 19:54 jhancock@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 19:10 brennen@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] (duration: 06m 41s) * 19:04 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns6001.wikimedia.org * 19:04 sukhe@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns6001.wikimedia.org * 19:04 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns5003.wikimedia.org * 19:04 sukhe@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns5003.wikimedia.org * 19:03 brennen@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 18:59 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns5003.wikimedia.org with OS trixie * 18:55 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns6001.wikimedia.org with OS trixie * 18:19 brett@cumin2002: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on P<nowiki>{</nowiki>cp7009.magru.wmnet<nowiki>}</nowiki> and A:cp - 9.2.15 Upgrade () * 18:14 brett@cumin2002: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on P<nowiki>{</nowiki>cp7009.magru.wmnet<nowiki>}</nowiki> and A:cp - 9.2.15 Upgrade () * 18:13 brennen@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 18:12 brett@cumin2002: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 9.2.15 Upgrade () * 18:09 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns5003.wikimedia.org with reason: host reimage * 18:06 brett@cumin2002: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 9.2.15 Upgrade () * 18:06 brennen: 1.47.0-wmf.15 train status ([[phab:T430834|T430834]]) - no current blockers, rolling to group0 * 18:05 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns5003.wikimedia.org with reason: host reimage * 18:05 brett: import trafficserver-9.2.15~deb13+wmf1 into trixie-wikimedia ([[phab:T434478|T434478]]) * 18:01 ladsgroup@cumin1003: END (PASS) - Cookbook sre.mysql.sanitarium_restart (exit_code=0) * 17:58 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns6001.wikimedia.org with reason: host reimage * 17:53 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324369{{!}}Enable desktop lazy loading on group0 (T148047)]] (duration: 07m 31s) * 17:52 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns6001.wikimedia.org with reason: host reimage * 17:49 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 17:49 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitarium_restart (exit_code=99) * 17:49 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 17:49 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7001.magru.wmnet * 17:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti7001.magru.wmnet * 17:48 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 17:47 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1324369{{!}}Enable desktop lazy loading on group0 (T148047)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:45 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1324369{{!}}Enable desktop lazy loading on group0 (T148047)]] * 17:39 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti7001.magru.wmnet * 17:36 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns5003.wikimedia.org with OS trixie * 17:34 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns6001.wikimedia.org with OS trixie * 17:31 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324361{{!}}Move FR config from IS.php to a dedicated file]], [[gerrit:1324363{{!}}Remove $wmg = $wg hacks in CentralAuth (T119117)]] (duration: 12m 23s) * 17:26 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 17:23 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1324361{{!}}Move FR config from IS.php to a dedicated file]], [[gerrit:1324363{{!}}Remove $wmg = $wg hacks in CentralAuth (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:19 sukhe: sudo cumin "A:cp-magru" "run-puppet-agent --enable 'merging CR 1324355'" * 17:18 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1324361{{!}}Move FR config from IS.php to a dedicated file]], [[gerrit:1324363{{!}}Remove $wmg = $wg hacks in CentralAuth (T119117)]] * 17:11 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-master1004.eqiad.wmnet with OS bookworm * 17:11 sukhe: sukhe@cp7005:~$ sudo puppet agent -tv * 17:02 sukhe: sudo cumin "A:cp-magru" "disable-puppet 'merging CR 1324355'" * 16:54 sukhe@dns1004: END - running authdns-update * 16:53 sukhe@dns1004: START - running authdns-update * 16:53 sukhe@dns1004: FAIL - running authdns-update * 16:51 sukhe@dns1004: START - running authdns-update * 16:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-master1004.eqiad.wmnet with reason: host reimage * 16:44 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-master1004.eqiad.wmnet with reason: host reimage * 16:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1157.eqiad.wmnet onto db1272.eqiad.wmnet * 16:40 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1157: Pool db1157.eqiad.wmnet in after cloning * 16:38 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 16:31 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324356{{!}}InitialiseSettings: Fix wgOATHAuthEnforce2FAForAll]] (duration: 06m 52s) * 16:30 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 16:28 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 16:27 reedy@deploy1003: reedy: Continuing with deployment * 16:26 reedy@deploy1003: reedy: Backport for [[gerrit:1324356{{!}}InitialiseSettings: Fix wgOATHAuthEnforce2FAForAll]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:24 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324356{{!}}InitialiseSettings: Fix wgOATHAuthEnforce2FAForAll]] * 16:13 sukhe: restart ntpsec.serviceon dns7001 * 16:09 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324335{{!}}InitialiseSettings: Enable 2FA enforcement on various private wikis (T428103)]] (duration: 06m 40s) * 16:08 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 16:06 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2204: Security update * 16:04 reedy@deploy1003: reedy: Continuing with deployment * 16:04 reedy@deploy1003: reedy: Backport for [[gerrit:1324335{{!}}InitialiseSettings: Enable 2FA enforcement on various private wikis (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:02 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7001.magru.wmnet * 16:02 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324335{{!}}InitialiseSettings: Enable 2FA enforcement on various private wikis (T428103)]] * 16:01 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 15:55 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1157: Pool db1157.eqiad.wmnet in after cloning * 15:54 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 15:54 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 15:41 dancy@deploy1003: Finished scap sync-world: Testing (duration: 06m 28s) * 15:40 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-master1004.eqiad.wmnet with OS bookworm * 15:40 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 15:35 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti4008.ulsfo.wmnet * 15:35 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti4008.ulsfo.wmnet * 15:34 dancy@deploy1003: Started scap sync-world: Testing * 15:34 dancy@deploy1003: Installation of scap version "4.279.0" completed for 3 hosts * 15:34 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-master1004.eqiad.wmnet with OS bookworm * 15:34 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 15:33 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-master1004.eqiad.wmnet with OS bookworm * 15:32 dancy@deploy1003: Installing scap version "4.279.0" for 3 host(s) * 15:32 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324339{{!}}Add /w/deployment-info.php entrypoint]] (duration: 07m 25s) * 15:30 moritzm: failover ganeti master in magru to ganeti7004 * 15:29 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti4008.ulsfo.wmnet * 15:28 dancy@deploy1003: dancy: Continuing with deployment * 15:28 tappof: remove 2026-05 swift log archives from centrallog to free some space ([[phab:T434502|T434502]]) * 15:27 dancy@deploy1003: dancy: Backport for [[gerrit:1324339{{!}}Add /w/deployment-info.php entrypoint]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:25 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324339{{!}}Add /w/deployment-info.php entrypoint]] * 15:20 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2204: Security update * 15:18 dancy@deploy1003: Installation of scap version "4.278.0" completed for 3 hosts * 15:18 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7004.magru.wmnet * 15:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti7004.magru.wmnet * 15:16 dancy@deploy1003: Installing scap version "4.278.0" for 3 host(s) * 15:14 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2204.codfw.wmnet with reason: Maintenance * 15:12 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 15:11 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-master1004.eqiad.wmnet with OS bookworm * 15:11 hashar: Restarting CI Jenkins on contint1003 due to Java upgrade. * 15:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2204 [[phab:T434565|T434565]]', diff saved to https://phabricator.wikimedia.org/P95984 and previous config saved to /var/cache/conftool/dbconfig/20260811-151126-cwilliams.json * 15:10 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti4008.ulsfo.wmnet * 15:10 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti7004.magru.wmnet * 15:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2207 to s2 primary [[phab:T434565|T434565]]', diff saved to https://phabricator.wikimedia.org/P95983 and previous config saved to /var/cache/conftool/dbconfig/20260811-150905-cwilliams.json * 15:08 cezmunsta: Starting s2 codfw failover from db2204 to db2207 - [[phab:T434565|T434565]] * 15:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2207 with weight 0 [[phab:T434565|T434565]]', diff saved to https://phabricator.wikimedia.org/P95982 and previous config saved to /var/cache/conftool/dbconfig/20260811-150402-cwilliams.json * 15:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s2 [[phab:T434565|T434565]] * 14:55 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1010.eqiad.wmnet with OS bookworm * 14:49 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns4003.wikimedia.org with OS trixie * 14:47 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1010.eqiad.wmnet with OS bookworm * 14:47 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7001.wikimedia.org with OS trixie * 14:44 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-presto1010.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:41 btullis@cumin1003: START - Cookbook sre.hosts.provision for host an-presto1010.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:40 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-presto1010.eqiad.wmnet with OS bookworm * 14:39 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 14:39 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-presto1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:36 btullis@cumin1003: START - Cookbook sre.hosts.provision for host an-presto1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:32 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1009.eqiad.wmnet with OS bookworm * 14:32 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 14:31 cwilliams@cumin1003: START - Cookbook sre.mysql.clone of db1157.eqiad.wmnet onto db1272.eqiad.wmnet * 14:30 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7004.magru.wmnet * 14:28 moritzm: failover ganeti master in ulsfo to ganeti4005 * 14:23 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7003.magru.wmnet * 14:23 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti7003.magru.wmnet * 14:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-master1004.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:22 btullis@cumin1003: START - Cookbook sre.hosts.provision for host an-master1004.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:21 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-master1004.eqiad.wmnet with OS bookworm * 14:19 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti4007.ulsfo.wmnet * 14:19 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti4007.ulsfo.wmnet * 14:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1010.eqiad.wmnet with OS bookworm * 14:17 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1008.eqiad.wmnet with OS bookworm * 14:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti7003.magru.wmnet * 14:12 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7003.magru.wmnet * 14:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti4007.ulsfo.wmnet * 14:11 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7002.magru.wmnet * 14:11 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti7002.magru.wmnet * 14:09 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324318{{!}}Revert "wmf-config/ProductionServices: set URL for urldownloader to service record" (T429175)]] (duration: 06m 46s) * 14:09 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7001.wikimedia.org with reason: host reimage * 14:06 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti4007.ulsfo.wmnet * 14:05 kharlan@deploy1003: kharlan: Continuing with deployment * 14:05 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti4006.ulsfo.wmnet * 14:04 jayme: updated calico to v3.30.7 on staging-eqiad - [[phab:T427400|T427400]] * 14:04 kharlan@deploy1003: kharlan: Backport for [[gerrit:1324318{{!}}Revert "wmf-config/ProductionServices: set URL for urldownloader to service record" (T429175)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti4006.ulsfo.wmnet * 14:03 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns4003.wikimedia.org with reason: host reimage * 14:03 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7001.wikimedia.org with reason: host reimage * 14:02 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti7002.magru.wmnet * 14:02 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1324318{{!}}Revert "wmf-config/ProductionServices: set URL for urldownloader to service record" (T429175)]] * 14:02 btullis@dns1004: FAIL - running authdns-update * 14:00 btullis@dns1004: START - running authdns-update * 13:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1008.eqiad.wmnet with reason: host reimage * 13:59 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'. * 13:59 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1313985{{!}}wmf-config/ProductionServices: set URL for urldownloader to service record (T429175)]] (duration: 25m 06s) * 13:58 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7002.magru.wmnet * 13:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti4006.ulsfo.wmnet * 13:57 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns4003.wikimedia.org with reason: host reimage * 13:56 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'. * 13:56 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7001.magru.wmnet * 13:56 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1008.eqiad.wmnet with reason: host reimage * 13:55 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'. * 13:55 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'. * 13:55 kharlan@deploy1003: kharlan, sukhe: Continuing with deployment * 13:53 marostegui: Failover ms2 [[phab:T434288|T434288]] * 13:52 marostegui: Failover ms1 [[phab:T434288|T434288]] * 13:52 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7001.magru.wmnet * 13:51 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti4006.ulsfo.wmnet * 13:48 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti4005.ulsfo.wmnet * 13:48 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti4005.ulsfo.wmnet * 13:44 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti4005.ulsfo.wmnet * 13:40 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1008.eqiad.wmnet with OS bookworm * 13:39 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns4003.wikimedia.org with OS trixie * 13:38 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns7001.wikimedia.org with OS trixie * 13:38 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1008.eqiad.wmnet with OS bookworm * 13:37 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti4005.ulsfo.wmnet * 13:36 kharlan@deploy1003: kharlan, sukhe: Backport for [[gerrit:1313985{{!}}wmf-config/ProductionServices: set URL for urldownloader to service record (T429175)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:34 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1313985{{!}}wmf-config/ProductionServices: set URL for urldownloader to service record (T429175)]] * 13:29 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2034.codfw.wmnet * 13:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2034.codfw.wmnet * 13:25 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.clone (exit_code=99) of db1157.eqiad.wmnet onto db1272.eqiad.wmnet * 13:25 cwilliams@cumin1003: START - Cookbook sre.mysql.clone of db1157.eqiad.wmnet onto db1272.eqiad.wmnet * 13:21 urbanecm@deploy1003: mwscript-k8s job started: namespaceDupes.php --wiki=frwiktionary --fix # [[phab:T415716|T415716]] * 13:21 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2034.codfw.wmnet * 13:20 urbanecm@deploy1003: mwscript-k8s job started: namespaceDupes.php --wiki=frwiktionary # [[phab:T415716|T415716]] * 13:19 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1323350{{!}}[tgwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T415307)]], [[gerrit:1322961{{!}}[slwiki] Revert temporary logo for Wikipedia 25 (Vector legacy + Vector 2022) (T414265)]], [[gerrit:1323827{{!}}[itwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T414320)]] (duration: 08m 00s) * 13:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-coord1003.eqiad.wmnet with OS bookworm * 13:15 urbanecm@deploy1003: urbanecm, superpes: Continuing with deployment * 13:13 urbanecm@deploy1003: urbanecm, superpes: Backport for [[gerrit:1323350{{!}}[tgwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T415307)]], [[gerrit:1322961{{!}}[slwiki] Revert temporary logo for Wikipedia 25 (Vector legacy + Vector 2022) (T414265)]], [[gerrit:1323827{{!}}[itwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T414320)]] synced to the testservers (see https://wiki * 13:11 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1323350{{!}}[tgwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T415307)]], [[gerrit:1322961{{!}}[slwiki] Revert temporary logo for Wikipedia 25 (Vector legacy + Vector 2022) (T414265)]], [[gerrit:1323827{{!}}[itwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T414320)]] * 13:11 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1323779{{!}}[ukwiki] Remove reviewer usergroup (T434252)]], [[gerrit:1323312{{!}}[frwiktionary] Add new Schème namespace and its talk (T415716)]] (duration: 06m 49s) * 13:10 marostegui@dns1004: END - running authdns-update * 13:08 marostegui@dns1004: START - running authdns-update * 13:07 marostegui@cumin1003: dbctl commit (dc=all): 'Repool ms2 [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95980 and previous config saved to /var/cache/conftool/dbconfig/20260811-130725-marostegui.json * 13:06 urbanecm@deploy1003: urbanecm, superpes: Continuing with deployment * 13:06 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1266 to dbctl [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95979 and previous config saved to /var/cache/conftool/dbconfig/20260811-130627-marostegui.json * 13:06 urbanecm@deploy1003: urbanecm, superpes: Backport for [[gerrit:1323779{{!}}[ukwiki] Remove reviewer usergroup (T434252)]], [[gerrit:1323312{{!}}[frwiktionary] Add new Schème namespace and its talk (T415716)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:04 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1323779{{!}}[ukwiki] Remove reviewer usergroup (T434252)]], [[gerrit:1323312{{!}}[frwiktionary] Add new Schème namespace and its talk (T415716)]] * 12:59 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db2253.codfw.wmnet,db[1151,1266].eqiad.wmnet with reason: Switching over ms2 * 12:54 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1157: Using as clone source * 12:53 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1157: Using as clone source * 12:51 marostegui@cumin1003: dbctl commit (dc=all): 'Depool ms2 [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95977 and previous config saved to /var/cache/conftool/dbconfig/20260811-125129-marostegui.json * 12:47 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 12:46 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 12:45 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 12:44 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'recommendation-api-ng' for release 'main' . * 12:44 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'recommendation-api-ng' for release 'main' . * 12:43 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'recommendation-api-ng' for release 'main' . * 12:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-coord1003.eqiad.wmnet with reason: host reimage * 12:43 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'ores-legacy' for release 'main' . * 12:42 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'ores-legacy' for release 'main' . * 12:42 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2165: Security update * 12:40 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-coord1003.eqiad.wmnet with reason: host reimage * 12:39 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'ores-legacy' for release 'main' . * 12:38 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' . * 12:38 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' . * 12:37 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' . * 12:34 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2034.codfw.wmnet * 12:30 jmm@cumin2002: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti-test2001.codfw.wmnet * 12:30 jmm@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host ganeti-test2001.codfw.wmnet * 12:25 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1179.eqiad.wmnet onto db1278.eqiad.wmnet * 12:25 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1179: Pool db1179.eqiad.wmnet in after cloning * 12:23 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-coord1003.eqiad.wmnet with OS bookworm * 12:22 moritzm: failover ganeti master in codfw/routed to ganeti2033 * 12:22 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2033.codfw.wmnet * 12:22 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2033.codfw.wmnet * 12:19 jmm@cumin2002: START - Cookbook sre.hosts.reboot-single for host ganeti-test2001.codfw.wmnet * 12:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 12:18 jmm@cumin2002: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti-test2001.codfw.wmnet * 12:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1008.eqiad.wmnet with OS bookworm * 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2033.codfw.wmnet * 12:07 moritzm: failover ganeti master in ganeti/test to ganeti-test2003 * 12:04 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 12:03 jmm@cumin2003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti4005.ulsfo.wmnet * 12:03 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti4005.ulsfo.wmnet * 12:00 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324283{{!}}Use maximum compression level in SqlBlobStore and SqlBagOStuff (T428377)]] (duration: 11m 37s) * 11:57 jmm@cumin2002: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti-test2002.codfw.wmnet * 11:57 jmm@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti-test2002.codfw.wmnet * 11:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2165: Security update * 11:54 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 11:52 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1324283{{!}}Use maximum compression level in SqlBlobStore and SqlBagOStuff (T428377)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:51 jmm@cumin2002: START - Cookbook sre.hosts.reboot-single for host ganeti-test2002.codfw.wmnet * 11:50 jmm@cumin2002: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti-test2002.codfw.wmnet * 11:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2165.codfw.wmnet with reason: Maintenance * 11:48 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@050d19e] (releasing): [[phab:T434186|T434186]] (duration: 01m 14s) * 11:48 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1324283{{!}}Use maximum compression level in SqlBlobStore and SqlBagOStuff (T428377)]] * 11:47 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@050d19e] (releasing): [[phab:T434186|T434186]] * 11:44 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@050d19e] (releasing): test jenkins deploy for [[phab:T434186|T434186]] (duration: 01m 08s) * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2165 [[phab:T434514|T434514]]', diff saved to https://phabricator.wikimedia.org/P95969 and previous config saved to /var/cache/conftool/dbconfig/20260811-114352-cwilliams.json * 11:43 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@050d19e] (releasing): test jenkins deploy for [[phab:T434186|T434186]] * 11:42 jmm@cumin2003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti-test2003.codfw.wmnet * 11:42 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti-test2003.codfw.wmnet * 11:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2161 to s8 primary [[phab:T434514|T434514]]', diff saved to https://phabricator.wikimedia.org/P95968 and previous config saved to /var/cache/conftool/dbconfig/20260811-114136-cwilliams.json * 11:40 cezmunsta: Starting s8 codfw failover from db2165 to db2161 - [[phab:T434514|T434514]] * 11:40 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1179: Pool db1179.eqiad.wmnet in after cloning * 11:36 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti-test2003.codfw.wmnet * 11:36 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti-test2003.codfw.wmnet * 11:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2161 with weight 0 [[phab:T434514|T434514]]', diff saved to https://phabricator.wikimedia.org/P95966 and previous config saved to /var/cache/conftool/dbconfig/20260811-113449-cwilliams.json * 11:34 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 25 hosts with reason: Primary switchover s8 [[phab:T434514|T434514]] * 11:29 moritzm: installing Python 3.11 security updates * 11:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-presto1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 11:26 btullis@cumin1003: START - Cookbook sre.hosts.provision for host an-presto1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 11:23 btullis@dns1004: END - running authdns-update * 11:21 btullis@dns1004: START - running authdns-update * 11:20 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-presto1008.eqiad.wmnet with OS bookworm * 11:20 moritzm: installing curl security updates * 11:11 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1007.eqiad.wmnet with OS bookworm * 10:45 tappof: bump space for prometheus k8s-dse in codfw * 10:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1007.eqiad.wmnet with reason: host reimage * 10:38 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1007.eqiad.wmnet with reason: host reimage * 10:37 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-coord1004.eqiad.wmnet with OS bookworm * 10:35 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1008.eqiad.wmnet with OS bookworm * 10:34 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1179: Depool db1179.eqiad.wmnet to then clone it to db1278.eqiad.wmnet - marostegui@cumin1003 * 10:34 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1008.eqiad.wmnet with OS bookworm * 10:25 fceratto@cumin1003: dbctl commit (dc=all): 'Remove db1177 [[phab:T433474|T433474]]', diff saved to https://phabricator.wikimedia.org/P95964 and previous config saved to /var/cache/conftool/dbconfig/20260811-102527-fceratto.json * 10:22 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1008.eqiad.wmnet with OS bookworm * 10:22 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1007.eqiad.wmnet with OS bookworm * 10:21 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1006.eqiad.wmnet with OS bookworm * 10:20 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 10:18 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1179: Depool db1179.eqiad.wmnet to then clone it to db1278.eqiad.wmnet - marostegui@cumin1003 * 10:18 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1179.eqiad.wmnet onto db1278.eqiad.wmnet * 10:17 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 10:17 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 10:14 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 10:09 blake@deploy1003: Stopping before sync operations * 10:09 blake@deploy1003: Started scap sync-world: Non-deployment run to populate release values for [[phab:T427668|T427668]] * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 10:04 fceratto@cumin1003: Removing db1177 from zarcillo [[phab:T433474|T433474]] * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1177.eqiad.wmnet * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1177.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:03 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1177.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:00 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1006.eqiad.wmnet with reason: host reimage * 09:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-coord1004.eqiad.wmnet with reason: host reimage * 09:57 marostegui: Failover m1 from db1164 to db1213 - [[phab:T434493|T434493]] * 09:57 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1006.eqiad.wmnet with reason: host reimage * 09:55 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:54 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2232].codfw.wmnet,db[1164,1213,1217].eqiad.wmnet with reason: Primary switchover m1 [[phab:T434493|T434493]] * 09:52 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-coord1004.eqiad.wmnet with reason: host reimage * 09:49 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1213.eqiad.wmnet with OS trixie * 09:49 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1177.eqiad.wmnet * 09:41 moritzm: installing Linux 6.12.101 on Trixie hosts * 09:40 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1006.eqiad.wmnet with OS bookworm * 09:35 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-coord1004.eqiad.wmnet with OS bookworm * 09:28 moritzm: installing node-tar security updates * 09:27 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1213.eqiad.wmnet with reason: host reimage * 09:22 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1213.eqiad.wmnet with reason: host reimage * 09:09 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1177: Decommission * 09:08 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db1177: Decommission * 09:08 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 09:08 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.decommission (exit_code=99) * 09:06 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1213.eqiad.wmnet with OS trixie * 09:06 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 09:05 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1213.eqiad.wmnet with reason: Reimage * 08:53 marostegui@dns1004: END - running authdns-update * 08:51 marostegui@dns1004: START - running authdns-update * 08:48 marostegui: Switchover ms1 master in eqiad [[phab:T434288|T434288]] * 08:48 marostegui@cumin1003: dbctl commit (dc=all): 'Repool ms1 [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95962 and previous config saved to /var/cache/conftool/dbconfig/20260811-084804-marostegui.json * 08:40 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1267 to dbctl [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95961 and previous config saved to /var/cache/conftool/dbconfig/20260811-084054-marostegui.json * 08:29 marostegui: Failover m1 from db1213 to db1164 - [[phab:T434043|T434043]] * 08:25 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2232].codfw.wmnet,db[1164,1213,1217].eqiad.wmnet with reason: Primary switchover m1 [[phab:T434043|T434043]] * 08:22 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db2251.codfw.wmnet,db[1152,1267].eqiad.wmnet with reason: Switching over ms1 * 08:22 marostegui@cumin1003: dbctl commit (dc=all): 'Depool ms1 [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95960 and previous config saved to /var/cache/conftool/dbconfig/20260811-082201-marostegui.json * 08:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: Switching over ms1 * 08:20 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.parsercache (exit_code=99) * 08:20 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 08:20 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1152: Switching over ms1 * 08:19 slyngshede@dns1004: END - running authdns-update * 08:18 moritzm: installing openjdk-21 security updates * 08:17 slyngshede@dns1004: START - running authdns-update * 08:16 moritzm: imported jenkins 2.568.2 to thirdparty/jenkins for trixie-wikimedia * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.12 (duration: 02m 26s) * 03:36 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] (duration: 33m 33s) * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 35s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-10 == * 14:54 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1323973{{!}}mmv.bootstrap: Fix getUrlParam to account for TIFF lossy/lossless param (T434333)]] (duration: 11m 24s) * 14:50 krinkle@deploy1003: krinkle: Continuing with deployment * 14:45 krinkle@deploy1003: krinkle: Backport for [[gerrit:1323973{{!}}mmv.bootstrap: Fix getUrlParam to account for TIFF lossy/lossless param (T434333)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:43 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1323973{{!}}mmv.bootstrap: Fix getUrlParam to account for TIFF lossy/lossless param (T434333)]] * 14:07 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1323967{{!}}updateIsActiveFlagForMentees: Commit the final partial batch (T432959)]] (duration: 10m 33s) * 13:56 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1323967{{!}}updateIsActiveFlagForMentees: Commit the final partial batch (T432959)]] * 13:45 dani@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply * 13:45 dani@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply * 13:45 dani@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply * 13:45 dani@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply * 13:45 dani@deploy1003: helmfile [staging] DONE helmfile.d/services/miscweb: apply * 13:44 dani@deploy1003: helmfile [staging] START helmfile.d/services/miscweb: apply * 13:38 wmde-fisch@deploy1003: Finished scap sync-world: Backport for [[gerrit:1323939{{!}}Enable sub-references on more group2 wikis (batch3) (T432731)]] (duration: 33m 21s) * 13:25 wmde-fisch@deploy1003: wmde-fisch: Continuing with deployment * 13:22 wmde-fisch@deploy1003: wmde-fisch: Backport for [[gerrit:1323939{{!}}Enable sub-references on more group2 wikis (batch3) (T432731)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:05 wmde-fisch@deploy1003: Started scap sync-world: Backport for [[gerrit:1323939{{!}}Enable sub-references on more group2 wikis (batch3) (T432731)]] * 07:57 hashar@deploy1003: Finished deploy [integration/docroot@7772132]: update build dependencies (duration: 00m 13s) * 07:57 hashar@deploy1003: Started deploy [integration/docroot@7772132]: update build dependencies * 07:35 _joe_: restarting squid on urldownloader1006 * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 48s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-09 == * 16:01 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:01 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:01 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:00 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 36s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-08 == * 05:31 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9] (wcqs): [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) (duration: 02m 36s) * 05:28 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9] (wcqs): [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) * 04:56 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 04:55 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 04:47 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) (duration: 19m 22s) * 04:28 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) * 04:19 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) (duration: 00m 06s) * 04:18 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) * 04:17 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) (duration: 00m 28s) * 04:16 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) * 03:52 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 03:52 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 34s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-07 == * 23:30 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:29 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 22:45 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 22:43 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 22:41 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 22:41 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 20:54 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 20:32 andrewbogott: restarting puppetserver service on puppetserver* for [[phab:T434339|T434339]] * 19:52 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:45 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 19:32 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:25 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:22 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 19:21 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 18:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:41 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 18:35 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 18:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 18:22 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 18:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 18:16 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 18:12 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:09 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:08 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:07 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:04 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:01 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:00 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:00 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 17:59 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 17:25 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 17:14 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 17:13 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 17:13 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 17:13 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:54 maryum: Deployed security fix for [[phab:T434278|T434278]] * 16:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 16:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 16:27 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --verbose --scoreLessThan=0.7 --exceptDatasetChecksums=[[phab:T434319|T434319]]-enwiki-models.txt # [[phab:T434319|T434319]] * 16:06 cdobbins@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp5022.eqsin.wmnet with OS trixie * 15:13 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 14:19 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 14:17 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 13:50 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1156.eqiad.wmnet onto db1271.eqiad.wmnet * 13:50 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1271: Pool db1271.eqiad.wmnet in after cloning * 13:02 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1271: Pool db1271.eqiad.wmnet in after cloning * 12:19 jayme: updated calico to v3.30.7 on staging-codfw - [[phab:T427400|T427400]] * 12:09 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 12:06 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 12:06 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 12:05 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 12:02 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1156: Pool db1156.eqiad.wmnet in after cloning * 11:38 bjensen: sudo -i reprepro -C main include trixie-wikimedia $<nowiki>{</nowiki>HOME<nowiki>}</nowiki>/httpbb/trixie/httpbb_$<nowiki>{</nowiki>VERSION?<nowiki>}</nowiki>-1+deb13u1_amd64.changes #[[phab:T434052|T434052]] * 11:35 bjensen: sudo -i reprepro -C main include bookworm-wikimedia $<nowiki>{</nowiki>HOME<nowiki>}</nowiki>/httpbb/bookworm/httpbb_$<nowiki>{</nowiki>VERSION?<nowiki>}</nowiki>-1_amd64.changes #[[phab:T434052|T434052]] * 11:30 marostegui@cumin1003: dbctl commit (dc=all): 'Adding db1271 to dbctl', diff saved to https://phabricator.wikimedia.org/P95945 and previous config saved to /var/cache/conftool/dbconfig/20260807-113006-marostegui.json * 11:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1156: Pool db1156.eqiad.wmnet in after cloning * 10:23 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 10:22 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 10:22 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 10:21 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 10:20 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 10:20 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 10:19 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 10:18 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 10:06 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on 21 hosts with reason: cloning * 10:01 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1156: Depool db1156.eqiad.wmnet to then clone it to db1271.eqiad.wmnet - marostegui@cumin1003 * 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1156: Depool db1156.eqiad.wmnet to then clone it to db1271.eqiad.wmnet - marostegui@cumin1003 * 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1156.eqiad.wmnet onto db1271.eqiad.wmnet * 09:15 jynus: started stress testing db1245 dbs [[phab:T431115|T431115]] * 08:19 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:18 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:16 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:14 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:13 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:10 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:06 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:05 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:00 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 10 days, 0:00:00 on ml-serve1015.eqiad.wmnet with reason: Downtime to get full picture of current BIOS settings beyond what Redfish shows * 08:00 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 07:54 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 07:54 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 07:53 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:52 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:51 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:50 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:49 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:48 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:47 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:45 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:45 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:41 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:38 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:37 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 06:35 jayme: updated istio to 1.29.4 on wikikube eqiad - [[phab:T427401|T427401]] * 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1178.eqiad.wmnet * 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1178.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 06:06 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1178.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 05:55 marostegui@cumin1003: START - Cookbook sre.dns.netbox * 05:49 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1178.eqiad.wmnet * 05:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 05:46 marostegui@cumin1003: Removing db1178 from zarcillo [[phab:T433471|T433471]] * 05:45 marostegui@cumin1003: START - Cookbook sre.mysql.decommission * 02:42 denisse: Extended volume on prometheus2008 for the disk space alert as per https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_host_running_out_of_space * 02:37 denisse: Extended volume on prometheus2007 tor the disk space alert as per https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_host_running_out_of_space * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 56s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-06 == * 21:39 maryum: Deploy security patch for [[phab:T433070|T433070]] * 21:29 maryum: Deploy security patch for [[phab:T434189|T434189]] * 20:48 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] (duration: 08m 12s) * 20:44 aude@deploy1003: lmora, aude, anzx: Continuing with deployment * 20:41 aude@deploy1003: lmora, aude, anzx: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be * 20:41 ebernhardson@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:41 ebernhardson@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 20:40 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] * 20:37 ebernhardson@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:37 ebernhardson@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 20:32 ebernhardson@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:32 ebernhardson@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 20:31 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] (duration: 06m 41s) * 20:27 cjming@deploy1003: cjming, ebernhardson, chlod: Continuing with deployment * 20:26 cjming@deploy1003: cjming, ebernhardson, chlod: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:24 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] * 20:18 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] (duration: 09m 22s) * 20:14 cjming@deploy1003: cjming, tsev: Continuing with deployment * 20:11 cjming@deploy1003: cjming, tsev: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:09 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] * 19:41 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply * 19:40 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply * 19:31 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 19:31 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 19:00 cdobbins@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cp5022.eqsin.wmnet with OS trixie * 18:25 ladsgroup@deploy1003: Finished scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) (duration: 06m 08s) * 18:19 ladsgroup@deploy1003: Started scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) * 18:18 ladsgroup@deploy1003: Stopping before sync operations * 18:17 ladsgroup@deploy1003: Started scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) * 17:55 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 16:50 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 16:35 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1001.eqiad.wmnet with OS bookworm * 16:19 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1002.eqiad.wmnet with reason: host reimage * 16:16 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1002.eqiad.wmnet with reason: host reimage * 16:05 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1001.eqiad.wmnet with reason: host reimage * 16:00 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1001.eqiad.wmnet with reason: host reimage * 15:57 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 15:43 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm * 15:29 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1001.eqiad.wmnet with OS bookworm * 15:29 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:58 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm * 14:57 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-drmrs ([[phab:T428495|T428495]]) * 14:55 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-drmrs ([[phab:T428495|T428495]]) * 14:55 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-ui1001.eqiad.wmnet with OS bookworm * 14:54 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-presto1001.eqiad.wmnet with OS bookworm * 14:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-magru ([[phab:T428495|T428495]]) * 14:49 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-magru ([[phab:T428495|T428495]]) * 14:48 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1001.eqiad.wmnet with OS bookworm * 14:46 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-esams ([[phab:T428495|T428495]]) * 14:44 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-esams ([[phab:T428495|T428495]]) * 14:43 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:42 brouberol@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:42 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:42 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 14:40 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 14:40 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:38 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-ui1001.eqiad.wmnet with reason: host reimage * 14:34 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-presto1001.eqiad.wmnet with reason: host reimage * 14:28 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-ui1001.eqiad.wmnet with reason: host reimage * 14:27 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-presto1001.eqiad.wmnet with reason: host reimage * 14:23 sukhe: sudo cumin -b2 'A:cp-text' "run-puppet-agent --enable 'merging CR 1290731'": [[phab:T425441|T425441]] * 14:18 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo for hosts in the wikimedia.org domain - [[phab:T428495|T428495]] * 14:16 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-presto1001.eqiad.wmnet with OS bookworm * 14:14 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-ui1001.eqiad.wmnet with OS bookworm * 14:12 sukhe: sudo cumin 'A:cp-text' "disable-puppet 'merging CR 1290731'": [[phab:T425441|T425441]] * 14:11 swfrench-wmf: restarted navtiming on webperf1003 - [[phab:T428495|T428495]] * 14:04 swfrench-wmf: begin rolling restart of confd in drmrs, eqiad, esams, magru - [[phab:T428495|T428495]] * 14:04 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm * 14:04 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:02 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-client1002.eqiad.wmnet with OS bookworm * 13:58 swfrench-wmf: authdns update to direct eqiad-associated etcd clients back to eqiad - [[phab:T428495|T428495]] * 13:58 swfrench@dns1004: END - running authdns-update * 13:56 swfrench@dns1004: START - running authdns-update * 13:49 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:44 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:31 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 13:29 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 13:26 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 13:23 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 13:22 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 13:19 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 13:18 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 13:18 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 13:17 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 13:16 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 13:13 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 13:11 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 13:09 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 13:06 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 13:06 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-client1002.eqiad.wmnet with OS bookworm * 13:05 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revision-models' for release 'main' . * 13:05 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:05 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revision-models' for release 'main' . * 13:04 brouberol@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-test-client1002.eqiad.wmnet with OS bookworm * 13:04 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 13:03 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 13:02 aikochou@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:00 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'readability' for release 'main' . * 12:59 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'readability' for release 'main' . * 12:58 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 12:57 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'logo-detection' for release 'main' . * 12:57 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'logo-detection' for release 'main' . * 12:57 aikochou@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 12:55 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:54 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:53 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 12:53 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:50 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 12:48 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 12:46 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'article-models' for release 'main' . * 12:45 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'article-models' for release 'main' . * 12:41 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . * 12:39 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . * 12:38 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-client1002.eqiad.wmnet with OS bookworm * 12:12 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply * 12:12 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply * 12:09 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:08 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 11:58 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2187: Security update * 11:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:24 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:16 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 11:15 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 11:10 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2187: Security update * 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2187.codfw.wmnet with reason: Maintenance * 10:56 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 10:56 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 10:56 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 10:56 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 10:54 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 10:53 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 10:09 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2187: Security update * 10:07 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2187: Security update * 09:39 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms2', diff saved to https://phabricator.wikimedia.org/P95929 and previous config saved to /var/cache/conftool/dbconfig/20260806-093908-marostegui.json * 09:36 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1178 from dbctl [[phab:T433471|T433471]]', diff saved to https://phabricator.wikimedia.org/P95928 and previous config saved to /var/cache/conftool/dbconfig/20260806-093632-marostegui.json * 09:33 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 09:31 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 09:30 topranks: bounce cr3-eqsin<->cr2-eqiad bgp session to disable no-prepend command * 09:20 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2253.codfw.wmnet,db1151.eqiad.wmnet with reason: cloning * 09:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1151: Cloning * 09:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:19 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 09:19 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1151: Cloning * 09:10 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 09:09 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup2003.codfw.wmnet * 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup2003.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 09:06 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup2003.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 09:03 klausman@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:02 jynus@cumin1003: START - Cookbook sre.dns.netbox * 09:02 klausman@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 08:57 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup2003.codfw.wmnet * 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup1003.eqiad.wmnet * 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:54 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms3', diff saved to https://phabricator.wikimedia.org/P95925 and previous config saved to /var/cache/conftool/dbconfig/20260806-085422-marostegui.json * 08:53 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:46 jynus@cumin1003: START - Cookbook sre.dns.netbox * 08:39 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup1003.eqiad.wmnet * 08:29 XioNoX: push pfw policy - [[phab:T434115|T434115]] * 08:14 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 08:00 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 08:00 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:58 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revision-models' for release 'main' . * 07:56 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 07:54 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'readability' for release 'main' . * 07:53 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'logo-detection' for release 'main' . * 07:51 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'llm' for release 'main' . * 07:48 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . * 07:37 jayme: updated istio to 1.29.4 on wikikube codfw - [[phab:T427401|T427401]] * 07:08 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2252.codfw.wmnet,db1153.eqiad.wmnet with reason: cloning * 07:07 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1153: Cloning * 07:07 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1153: Cloning * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 40s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-05 == * 23:24 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1009.eqiad.wmnet with OS bookworm * 23:03 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1009.eqiad.wmnet with reason: host reimage * 22:59 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1009.eqiad.wmnet with reason: host reimage * 22:43 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1009.eqiad.wmnet with OS bookworm * 22:38 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1009.eqiad.wmnet * 22:34 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1009.eqiad.wmnet * 22:25 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1008.eqiad.wmnet with OS bookworm * 22:04 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1008.eqiad.wmnet with reason: host reimage * 22:00 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1008.eqiad.wmnet with reason: host reimage * 21:48 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:47 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:46 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:44 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1008.eqiad.wmnet with OS bookworm * 21:43 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:41 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1008.eqiad.wmnet * 21:36 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1008.eqiad.wmnet * 21:14 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:12 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad * 21:12 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad * 21:10 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=eqiad * 21:08 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:07 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:07 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=eqiad * 21:04 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:03 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1006 * 21:02 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1006 * 21:00 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:56 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:56 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:55 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 20:55 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 20:55 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1007.eqiad.wmnet with OS bookworm * 20:51 vriley@cumin1003: START - Cookbook sre.dns.netbox * 20:43 ebernhardson: [[phab:T434008|T434008]]: changing cloudelastic:9643 from auto_expand_replicas to number_of_replicas * 20:34 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1007.eqiad.wmnet with reason: host reimage * 20:27 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1007.eqiad.wmnet with reason: host reimage * 20:24 cjming: end of UTC late backport window * 20:23 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] (duration: 06m 26s) * 20:18 cjming@deploy1003: cjming: Continuing with deployment * 20:18 cjming@deploy1003: cjming: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:16 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] * 20:12 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1007.eqiad.wmnet with OS bookworm * 20:12 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] (duration: 08m 41s) * 20:08 swfrench@cumin2002: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host conf1007.eqiad.wmnet with OS bookworm * 20:08 jforrester@deploy1003: jforrester: Continuing with deployment * 20:07 jforrester@deploy1003: jforrester: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:03 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] * 19:51 inflatador: [bking@puppetserver1001] ~$ sudo puppetserver ca sign --certname an-worker1189.eqiad.wmnet [[phab:T434142|T434142]] * 19:47 bking@cumin2003: DONE (FAIL) - Cookbook sre.puppet.renew-cert (exit_code=99) for an-worker1189.eqiad.wmnet: Renew puppet certificate - bking@cumin2003 * 19:46 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:30 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1007.eqiad.wmnet with OS trixie * 19:30 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 19:29 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 19:20 swfrench-wmf: silenced EtcdRelicationDown 0cb709a9-f244-4f1e-971f-{{Gerrit|440ec65e7fd7}} - [[phab:T428495|T428495]] * 19:13 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1007.eqiad.wmnet with OS bookworm * 19:12 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1007.eqiad.wmnet with reason: host reimage * 19:09 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1007.eqiad.wmnet * 19:07 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1007.eqiad.wmnet with reason: host reimage * 19:03 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1007.eqiad.wmnet * 18:52 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1007.eqiad.wmnet with OS trixie * 18:52 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1007.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:35 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 18:34 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 18:34 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 18:30 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1007.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:28 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:28 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1007] - vriley@cumin1003" * 18:27 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1007] - vriley@cumin1003" * 18:23 vriley@cumin1003: START - Cookbook sre.dns.netbox * 18:22 vriley@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 18:22 robh@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:19 vriley@cumin1003: START - Cookbook sre.dns.netbox * 18:13 robh@cumin2002: START - Cookbook sre.hosts.provision for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:31 jasmine@cumin2002: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-main-eqiad * 17:12 mutante: LDAP - added vwalters to group ciadmin - [[phab:T433615|T433615]] * 16:58 aokoth@deploy1003: Finished deploy [phabricator/deployment@e2ebca5]: Deploy Phab (duration: 00m 34s) * 16:57 aokoth@deploy1003: Started deploy [phabricator/deployment@e2ebca5]: Deploy Phab * 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad * 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=eqiad * 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad * 16:53 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:41 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2187.codfw.wmnet * 16:41 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2187.codfw.wmnet * 16:41 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker2187.codfw.wmnet * 16:41 cgoubert@cumin2003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker2187.codfw.wmnet * 16:40 jasmine@cumin2002: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-main-eqiad * 16:40 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:34 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:25 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-magru and A:liberica ([[phab:T428495|T428495]]) * 16:23 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-magru and A:liberica ([[phab:T428495|T428495]]) * 16:20 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-drmrs and A:liberica ([[phab:T428495|T428495]]) * 16:19 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-drmrs and A:liberica ([[phab:T428495|T428495]]) * 16:18 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-esams and A:liberica ([[phab:T428495|T428495]]) * 16:16 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-esams and A:liberica ([[phab:T428495|T428495]]) * 16:06 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1159.eqiad.wmnet * 16:06 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1159.eqiad.wmnet * 16:06 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1159.eqiad.wmnet * 16:05 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] (duration: 09m 11s) * 15:58 reedy@deploy1003: reedy: Continuing with deployment * 15:58 reedy@deploy1003: reedy: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:56 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] * 15:54 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1159.eqiad.wmnet with OS trixie * 15:38 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:33 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1159.eqiad.wmnet with reason: host reimage * 15:32 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:27 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1159.eqiad.wmnet with reason: host reimage * 15:10 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1159 * 15:10 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1159 * 15:00 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo for hosts in the wikimedia.org domain - [[phab:T428495|T428495]] * 14:55 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS trixie * 14:54 swfrench-wmf: restarted navtiming on webperf1003 - [[phab:T428495|T428495]] * 14:52 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1159 * 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1159.eqiad.wmnet 129.48.64.10.in-addr.arpa 9.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:52 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1159.eqiad.wmnet 129.48.64.10.in-addr.arpa 9.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1159 - jayme@cumin1003" * 14:52 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1159 - jayme@cumin1003" * 14:49 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:48 jayme@cumin1003: START - Cookbook sre.dns.netbox * 14:47 swfrench-wmf: begin rolling restart of confd in drmrs, eqiad, esams, magru - [[phab:T428495|T428495]] * 14:47 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1159 * 14:46 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:46 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:46 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1159.eqiad.wmnet with OS trixie * 14:45 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:44 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:44 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1159.eqiad.wmnet * 14:43 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:43 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1159.eqiad.wmnet * 14:43 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:43 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1159.eqiad.wmnet * 14:43 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:43 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1157.eqiad.wmnet * 14:43 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1157.eqiad.wmnet * 14:43 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1157.eqiad.wmnet * 14:42 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:42 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:42 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:41 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:41 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:41 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:41 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1003.eqiad.wmnet with OS bookworm * 14:39 swfrench-wmf: authdns update to direct eqiad-associated etcd clients to codfw - [[phab:T428495|T428495]] * 14:39 swfrench@dns1004: END - running authdns-update * 14:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 14:37 swfrench@dns1004: START - running authdns-update * 14:37 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:37 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:35 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:35 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 14:28 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:27 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1157.eqiad.wmnet with OS trixie * 14:27 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:27 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:27 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 14:26 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 14:26 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:26 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 14:26 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 14:26 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:26 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host search-loader1002.eqiad.wmnet with OS trixie * 14:19 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:19 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:15 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1003.eqiad.wmnet with reason: host reimage * 14:14 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:14 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:13 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 * 14:13 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host mc2046 * 14:13 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS trixie * 14:11 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1003.eqiad.wmnet with reason: host reimage * 14:10 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:09 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:09 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:09 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:08 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:08 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1157.eqiad.wmnet with reason: host reimage * 14:08 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:04 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 14:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on search-loader1002.eqiad.wmnet with reason: host reimage * 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=eqiad * 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=eqiad * 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=eqiad * 14:00 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:59 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:58 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1157.eqiad.wmnet with reason: host reimage * 13:57 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on search-loader1002.eqiad.wmnet with reason: host reimage * 13:54 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1003.eqiad.wmnet with OS bookworm * 13:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host search-loader1002.eqiad.wmnet with OS trixie * 13:43 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1157 * 13:42 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1157 * 13:40 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1157 * 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1157.eqiad.wmnet 183.32.64.10.in-addr.arpa 3.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1157.eqiad.wmnet 183.32.64.10.in-addr.arpa 3.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1157 - jayme@cumin1003" * 13:39 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1157 - jayme@cumin1003" * 13:39 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] (duration: 07m 00s) * 13:35 jayme@cumin1003: START - Cookbook sre.dns.netbox * 13:35 reedy@deploy1003: reedy: Continuing with deployment * 13:34 reedy@deploy1003: reedy: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:32 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] * 13:23 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1157 * 13:22 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1157.eqiad.wmnet with OS trixie * 13:22 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1157.eqiad.wmnet * 13:22 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1157.eqiad.wmnet * 13:21 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1157.eqiad.wmnet * 13:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1156.eqiad.wmnet * 13:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1156.eqiad.wmnet * 13:15 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1156.eqiad.wmnet * 13:01 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1156.eqiad.wmnet with OS trixie * 12:42 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1156.eqiad.wmnet with reason: host reimage * 12:38 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1156.eqiad.wmnet with reason: host reimage * 12:32 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 12:31 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 12:30 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 12:28 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 12:26 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 12:24 topranks: update bgp confed settings in eqsin * 12:22 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1156 * 12:22 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1156 * 12:22 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 12:19 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1156 * 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1156.eqiad.wmnet 110.32.64.10.in-addr.arpa 0.1.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:19 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1156.eqiad.wmnet 110.32.64.10.in-addr.arpa 0.1.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1156 - jayme@cumin1003" * 12:19 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1156 - jayme@cumin1003" * 12:17 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:14 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS trixie * 12:09 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:06 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:04 jayme@cumin1003: START - Cookbook sre.dns.netbox * 12:04 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 12:02 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'article-models' for release 'main' . * 12:01 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1156 * 12:01 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1156.eqiad.wmnet with OS trixie * 11:59 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1156.eqiad.wmnet * 11:59 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1156.eqiad.wmnet * 11:59 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1156.eqiad.wmnet * 11:57 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 11:53 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 11:53 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:52 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:52 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:50 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:50 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:50 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:49 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:48 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:47 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:47 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:45 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:45 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:44 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:44 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:44 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:43 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:42 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:38 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:35 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 * 11:35 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host mc2046 * 11:34 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS trixie * 11:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:27 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:21 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:21 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:18 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:18 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:18 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:18 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:13 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:13 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:09 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:08 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:07 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:06 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:06 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:05 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:05 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:04 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 11:04 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:24 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:24 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:17 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:16 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1155.eqiad.wmnet * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1155.eqiad.wmnet * 10:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1155.eqiad.wmnet * 10:14 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:14 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:11 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:11 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:10 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:09 aikochou@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop: sync * 10:09 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:09 aikochou@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop: sync * 10:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:07 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:05 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:05 aikochou@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop: sync * 10:05 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:05 aikochou@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop: sync * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:04 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:04 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:04 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1155.eqiad.wmnet with OS trixie * 09:52 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms1', diff saved to https://phabricator.wikimedia.org/P95918 and previous config saved to /var/cache/conftool/dbconfig/20260805-095212-marostegui.json * 09:44 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1152: after cloning * 09:44 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.parsercache (exit_code=99) * 09:44 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 09:44 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1152: after cloning * 09:43 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1155.eqiad.wmnet with reason: host reimage * 09:40 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1155.eqiad.wmnet with reason: host reimage * 09:32 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 09:32 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:31 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 09:31 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:27 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1155 * 09:27 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1155 * 09:25 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 09:24 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 09:24 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 09:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:23 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 09:23 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 09:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:22 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 09:22 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 09:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:20 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 09:20 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:17 XioNoX: push pfw policies - [[phab:T434038|T434038]] * 09:14 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1155 * 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1155.eqiad.wmnet 109.32.64.10.in-addr.arpa 9.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1155.eqiad.wmnet 109.32.64.10.in-addr.arpa 9.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1155 - jayme@cumin1003" * 09:14 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1155 - jayme@cumin1003" * 09:10 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 09:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:09 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2251.codfw.wmnet,db1152.eqiad.wmnet with reason: cloning * 09:09 jayme@cumin1003: START - Cookbook sre.dns.netbox * 09:08 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 09:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: Cloning * 09:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:05 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 09:05 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1152: Cloning * 08:38 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1155 * 08:37 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1155.eqiad.wmnet with OS trixie * 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1171.eqiad.wmnet * 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1171.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:29 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95913 and previous config saved to /var/cache/conftool/dbconfig/20260805-082908-ladsgroup.json * 08:27 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1171.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:22 jynus@cumin1003: START - Cookbook sre.dns.netbox * 08:18 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249', diff saved to https://phabricator.wikimedia.org/P95912 and previous config saved to /var/cache/conftool/dbconfig/20260805-081823-ladsgroup.json * 08:17 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1171.eqiad.wmnet * 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1150.eqiad.wmnet * 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1150.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:15 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1150.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:15 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 08:14 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1155.eqiad.wmnet * 08:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1155.eqiad.wmnet * 08:14 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1155.eqiad.wmnet * 08:11 jynus@cumin1003: START - Cookbook sre.dns.netbox * 08:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249', diff saved to https://phabricator.wikimedia.org/P95911 and previous config saved to /var/cache/conftool/dbconfig/20260805-080737-ladsgroup.json * 08:05 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1150.eqiad.wmnet * 08:02 marostegui: Depool clouddb1020 (s5,s8) [[phab:T434048|T434048]] * 08:02 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1020.eqiad.wmnet,service=s8 * 08:02 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1020.eqiad.wmnet,service=s5 * 08:02 marostegui: Depool clouddb1018 (s2,s7) [[phab:T434048|T434048]] * 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1018.eqiad.wmnet,service=s7 * 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1018.eqiad.wmnet,service=s2 * 08:01 marostegui: Depool clouddb1017 (s1) [[phab:T434048|T434048]] * 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1017.eqiad.wmnet,service=s1 * 07:59 marostegui: Depool clouddb1016 (s5,s8) [[phab:T434048|T434048]] * 07:59 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s8 * 07:59 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s5 * 07:57 marostegui: Depool clouddb1015 (s4,s6) [[phab:T434048|T434048]] * 07:57 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s6 * 07:57 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s4 * 07:56 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95910 and previous config saved to /var/cache/conftool/dbconfig/20260805-075650-ladsgroup.json * 07:54 marostegui: Depool clouddb1014 (s2,s7) [[phab:T434048|T434048]] * 07:54 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1014.eqiad.wmnet,service=s7 * 07:54 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1014.eqiad.wmnet,service=s2 * 07:53 marostegui: Depool clouddb1013:s1 [[phab:T434048|T434048]] * 07:53 marostegui: Depool clouddb1013:s1 [[phab:T409557|T409557]] * 07:53 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1013.eqiad.wmnet,service=s1 * 07:25 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95909 and previous config saved to /var/cache/conftool/dbconfig/20260805-072529-ladsgroup.json * 07:24 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2249.codfw.wmnet with reason: Maintenance * 07:24 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95908 and previous config saved to /var/cache/conftool/dbconfig/20260805-072426-ladsgroup.json * 07:21 slyngshede@dns1004: END - running authdns-update * 07:19 slyngshede@dns1004: START - running authdns-update * 07:13 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231', diff saved to https://phabricator.wikimedia.org/P95906 and previous config saved to /var/cache/conftool/dbconfig/20260805-071340-ladsgroup.json * 07:02 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231', diff saved to https://phabricator.wikimedia.org/P95905 and previous config saved to /var/cache/conftool/dbconfig/20260805-070253-ladsgroup.json * 06:52 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95904 and previous config saved to /var/cache/conftool/dbconfig/20260805-065206-ladsgroup.json * 06:45 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 06:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95903 and previous config saved to /var/cache/conftool/dbconfig/20260805-062240-ladsgroup.json * 06:21 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2231.codfw.wmnet with reason: Maintenance * 06:21 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95902 and previous config saved to /var/cache/conftool/dbconfig/20260805-062137-ladsgroup.json * 06:10 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215', diff saved to https://phabricator.wikimedia.org/P95901 and previous config saved to /var/cache/conftool/dbconfig/20260805-061051-ladsgroup.json * 06:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215', diff saved to https://phabricator.wikimedia.org/P95900 and previous config saved to /var/cache/conftool/dbconfig/20260805-060004-ladsgroup.json * 05:49 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95899 and previous config saved to /var/cache/conftool/dbconfig/20260805-054918-ladsgroup.json * 05:19 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95898 and previous config saved to /var/cache/conftool/dbconfig/20260805-051939-ladsgroup.json * 05:18 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2215.codfw.wmnet with reason: Maintenance * 04:30 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2201.codfw.wmnet with reason: Maintenance * 03:40 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2197.codfw.wmnet with reason: Maintenance * 03:40 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95897 and previous config saved to /var/cache/conftool/dbconfig/20260805-034036-ladsgroup.json * 03:29 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196', diff saved to https://phabricator.wikimedia.org/P95896 and previous config saved to /var/cache/conftool/dbconfig/20260805-032948-ladsgroup.json * 03:19 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196', diff saved to https://phabricator.wikimedia.org/P95895 and previous config saved to /var/cache/conftool/dbconfig/20260805-031902-ladsgroup.json * 03:08 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95894 and previous config saved to /var/cache/conftool/dbconfig/20260805-030815-ladsgroup.json * 02:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95893 and previous config saved to /var/cache/conftool/dbconfig/20260805-023413-ladsgroup.json * 02:33 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2196.codfw.wmnet with reason: Maintenance * 02:33 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95892 and previous config saved to /var/cache/conftool/dbconfig/20260805-023310-ladsgroup.json * 02:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186', diff saved to https://phabricator.wikimedia.org/P95891 and previous config saved to /var/cache/conftool/dbconfig/20260805-022223-ladsgroup.json * 02:11 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186', diff saved to https://phabricator.wikimedia.org/P95890 and previous config saved to /var/cache/conftool/dbconfig/20260805-021137-ladsgroup.json * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 02:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95889 and previous config saved to /var/cache/conftool/dbconfig/20260805-020051-ladsgroup.json * 01:30 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95888 and previous config saved to /var/cache/conftool/dbconfig/20260805-013029-ladsgroup.json * 01:29 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2186.codfw.wmnet with reason: Maintenance * 00:34 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on dbstore1009.eqiad.wmnet with reason: Maintenance * 00:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95887 and previous config saved to /var/cache/conftool/dbconfig/20260805-003408-ladsgroup.json * 00:23 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264', diff saved to https://phabricator.wikimedia.org/P95886 and previous config saved to /var/cache/conftool/dbconfig/20260805-002322-ladsgroup.json * 00:12 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264', diff saved to https://phabricator.wikimedia.org/P95885 and previous config saved to /var/cache/conftool/dbconfig/20260805-001235-ladsgroup.json * 00:01 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95884 and previous config saved to /var/cache/conftool/dbconfig/20260805-000148-ladsgroup.json == 2026-08-04 == * 23:45 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95883 and previous config saved to /var/cache/conftool/dbconfig/20260804-234508-ladsgroup.json * 23:44 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1264.eqiad.wmnet with reason: Maintenance * 23:44 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95882 and previous config saved to /var/cache/conftool/dbconfig/20260804-234405-ladsgroup.json * 23:33 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237', diff saved to https://phabricator.wikimedia.org/P95881 and previous config saved to /var/cache/conftool/dbconfig/20260804-233317-ladsgroup.json * 23:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237', diff saved to https://phabricator.wikimedia.org/P95880 and previous config saved to /var/cache/conftool/dbconfig/20260804-232230-ladsgroup.json * 23:11 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95879 and previous config saved to /var/cache/conftool/dbconfig/20260804-231144-ladsgroup.json * 22:23 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95878 and previous config saved to /var/cache/conftool/dbconfig/20260804-222345-ladsgroup.json * 22:23 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1237.eqiad.wmnet with reason: Maintenance * 21:13 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1225.eqiad.wmnet with reason: Maintenance * 20:40 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] (duration: 24m 40s) * 20:33 samtar@deploy1003: samtar, kineticpelagic: Continuing with deployment * 20:28 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS bookworm * 20:21 samtar@deploy1003: samtar, kineticpelagic: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:15 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] * 20:13 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 20:09 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 20:00 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1216.eqiad.wmnet with reason: Maintenance * 20:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95877 and previous config saved to /var/cache/conftool/dbconfig/20260804-195957-ladsgroup.json * 19:51 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 19:50 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) mc2046.codfw.wmnet 120.16.192.10.in-addr.arpa 0.2.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:50 jhancock@cumin2002: START - Cookbook sre.dns.wipe-cache mc2046.codfw.wmnet 120.16.192.10.in-addr.arpa 0.2.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host mc2046 - jhancock@cumin2002" * 19:50 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host mc2046 - jhancock@cumin2002" * 19:49 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203', diff saved to https://phabricator.wikimedia.org/P95876 and previous config saved to /var/cache/conftool/dbconfig/20260804-194911-ladsgroup.json * 19:46 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 19:45 jhancock@cumin2002: START - Cookbook sre.hosts.move-vlan for host mc2046 * 19:45 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS bookworm * 19:38 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203', diff saved to https://phabricator.wikimedia.org/P95875 and previous config saved to /var/cache/conftool/dbconfig/20260804-193825-ladsgroup.json * 19:27 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95874 and previous config saved to /var/cache/conftool/dbconfig/20260804-192738-ladsgroup.json * 19:02 mutante: gerrit ssh -p 29418 gerrit.wikimedia.org gerrit index changes {{Gerrit|1320979}} * 18:20 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 18:18 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 18:14 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 18:14 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 18:13 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 18:10 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 18:08 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 18:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95872 and previous config saved to /var/cache/conftool/dbconfig/20260804-180721-ladsgroup.json * 18:07 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 18:06 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1203.eqiad.wmnet with reason: Maintenance * 18:06 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95871 and previous config saved to /var/cache/conftool/dbconfig/20260804-180618-ladsgroup.json * 17:55 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179', diff saved to https://phabricator.wikimedia.org/P95870 and previous config saved to /var/cache/conftool/dbconfig/20260804-175531-ladsgroup.json * 17:55 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1154.eqiad.wmnet * 17:55 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1154.eqiad.wmnet * 17:55 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1154.eqiad.wmnet * 17:50 swfrench@deploy1003: Finished scap sync-world: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] (duration: 04m 05s) * 17:48 swfrench@deploy1003: swfrench: Continuing with deployment * 17:46 swfrench@deploy1003: swfrench: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:45 swfrench@deploy1003: Started scap sync-world: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] * 17:44 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179', diff saved to https://phabricator.wikimedia.org/P95869 and previous config saved to /var/cache/conftool/dbconfig/20260804-174445-ladsgroup.json * 17:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95868 and previous config saved to /var/cache/conftool/dbconfig/20260804-173359-ladsgroup.json * 17:33 swfrench@deploy1003: Finished scap sync-world: Pick up new PHP production image (duration: 28m 32s) * 17:28 aokoth@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on phab1005.eqiad.wmnet with reason: Puppet Failure * 17:05 swfrench@deploy1003: Started scap sync-world: Pick up new PHP production image * 17:00 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 17:00 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 16:54 cgoubert@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on wikikube-worker2187.codfw.wmnet with reason: Hardware issue * 16:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2187.codfw.wmnet * 16:52 mutante: gerrit2003:/var/log/apache2# ln -s /srv/gerrit/site_path/review_site/logs/ gerrit * 16:52 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2187.codfw.wmnet * 16:48 mutante: gerrit2003 - moving old apache logfiles older than 60 days from /var/log/apache2 to /srv/gerrit/site_path/review_site/logs/old/ * 16:33 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 16:32 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 16:29 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 16:29 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 16:28 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 16:28 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 16:27 dzahn@cumin1003: END (PASS) - Cookbook sre.gerrit.restart-gerrit (exit_code=0) Restarting Gerrit on gerrit2003 * 16:27 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 16:27 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95867 and previous config saved to /var/cache/conftool/dbconfig/20260804-162736-ladsgroup.json * 16:27 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 16:26 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1179.eqiad.wmnet with reason: Maintenance * 16:26 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:25 mutante: restarting gerrit - dropped outdated RSA host key * 16:25 dzahn@cumin1003: START - Cookbook sre.gerrit.restart-gerrit Restarting Gerrit on gerrit2003 * 16:24 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95866 and previous config saved to /var/cache/conftool/dbconfig/20260804-162424-ladsgroup.json * 16:24 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 16:23 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 16:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95865 and previous config saved to /var/cache/conftool/dbconfig/20260804-162236-ladsgroup.json * 16:21 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1179.eqiad.wmnet with reason: Maintenance * 16:17 swfrench-wmf: reprepro include php8.3_8.3.33-1+wmf11u1 into component/php83 for bullseye-wikimedia * 16:17 swfrench-wmf: reprepro include php8.3_8.3.33-1+wmf12u1 into component/php83 for bookworm-wikimedia * 16:11 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply * 16:10 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply * 16:10 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mobileapps: apply * 16:09 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mobileapps: apply * 16:09 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply * 16:08 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply * 16:08 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:08 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:07 aokoth@cumin1003: END (PASS) - Cookbook sre.vrts.upgrade (exit_code=0) on VRTS host vrts1003.eqiad.wmnet * 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:05 aokoth@cumin1003: START - Cookbook sre.vrts.upgrade on VRTS host vrts1003.eqiad.wmnet * 16:04 mutante: gerrit2002/gerrit1003/gerrit2003 - rm /etc/gerrit/ssh_host_rsa_key * 15:59 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:59 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:59 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:59 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:56 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 15:55 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:55 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:55 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:49 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 15:49 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:44 Raine: add php8.5 packages to component/php85 - [[phab:T432983|T432983]] * 15:39 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:33 aaron@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 15:33 aaron@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 15:29 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:19 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:19 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:16 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:16 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1154.eqiad.wmnet with OS trixie * 15:16 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:15 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:15 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:06 brennen@deploy1003: Finished deploy [phabricator/deployment@56f4ffd]: deploy phab1004 for [[phab:T433981|T433981]] (duration: 00m 43s) * 15:05 brennen@deploy1003: Started deploy [phabricator/deployment@56f4ffd]: deploy phab1004 for [[phab:T433981|T433981]] * 15:05 aaron@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 15:04 aaron@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 15:02 brennen@deploy1003: Finished deploy [phabricator/deployment@56f4ffd]: deploy phab2003 for [[phab:T433981|T433981]] (duration: 00m 51s) * 15:01 brennen@deploy1003: Started deploy [phabricator/deployment@56f4ffd]: deploy phab2003 for [[phab:T433981|T433981]] * 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1004.eqiad.wmnet with reason: deployment * 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1005.eqiad.wmnet with reason: deployment * 14:58 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab2003.codfw.wmnet with reason: deployment * 14:55 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1154.eqiad.wmnet with reason: host reimage * 14:51 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1154.eqiad.wmnet with reason: host reimage * 14:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 14:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 14:38 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync * 14:38 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync * 14:38 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync * 14:37 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync * 14:37 ottomata: roll restart eventgate-main to pick up stream config change - [[phab:T433507|T433507]] * 14:37 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-main: sync * 14:36 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-main: sync * 14:36 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1154 * 14:36 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1154 * 14:34 otto@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] (duration: 08m 39s) * 14:34 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1154 * 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1154.eqiad.wmnet 108.32.64.10.in-addr.arpa 8.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:34 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1154.eqiad.wmnet 108.32.64.10.in-addr.arpa 8.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1154 - jayme@cumin1003" * 14:34 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1154 - jayme@cumin1003" * 14:30 otto@deploy1003: otto: Continuing with deployment * 14:30 jayme@cumin1003: START - Cookbook sre.dns.netbox * 14:28 otto@deploy1003: otto: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:26 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1154 * 14:26 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1154.eqiad.wmnet with OS trixie * 14:26 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1154.eqiad.wmnet * 14:26 otto@deploy1003: Started scap sync-world: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] * 14:26 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1154.eqiad.wmnet * 14:26 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1154.eqiad.wmnet * 14:17 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 14:16 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 14:15 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 14:14 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 14:13 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 14:13 swfrench@dns1004: END - running authdns-update * 14:13 Msz2001: Finished deployments for UTC afternoon backport window * 14:13 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 14:13 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] (duration: 07m 58s) * 14:11 swfrench@dns1004: START - running authdns-update * 14:08 mszwarc@deploy1003: javiermonton, mszwarc, mpostoronca: Continuing with deployment * 14:07 mszwarc@deploy1003: javiermonton, mszwarc, mpostoronca: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] synced to the testser * 14:05 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] * 14:03 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 13:49 swfrench@cumin2002: conftool action : set/pooled=yes; selector: name=wikikube-worker2330.codfw.wmnet * 13:49 swfrench@cumin2002: conftool action : set/pooled=no; selector: name=wikikube-worker2330.codfw.wmnet * 13:48 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] (duration: 09m 19s) * 13:45 swfrench@dns1004: END - running authdns-update * 13:44 mszwarc@deploy1003: mszwarc, jforrester: Continuing with deployment * 13:43 swfrench@dns1004: START - running authdns-update * 13:41 mszwarc@deploy1003: mszwarc, jforrester: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:38 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] * 13:33 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 13:33 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1154.eqiad.wmnet * 13:32 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 13:32 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 13:31 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 13:31 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:31 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:29 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1154.eqiad.wmnet * 13:28 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1154.eqiad.wmnet * 13:28 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1154.eqiad.wmnet * 13:28 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1141.eqiad.wmnet * 13:28 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1141.eqiad.wmnet * 13:28 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1141.eqiad.wmnet * 13:22 otto@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply * 13:22 otto@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply * 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1096.eqiad.wmnet with OS trixie * 13:05 swfrench@dns1004: END - running authdns-update * 13:03 swfrench@dns1004: START - running authdns-update * 12:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 12:43 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 1:00:00 on db1171.eqiad.wmnet with reason: decom * 12:42 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 1:00:00 on db1150.eqiad.wmnet with reason: decom * 12:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 12:38 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1164,1217].eqiad.wmnet with reason: cloning * 12:33 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2096.codfw.wmnet with OS trixie * 12:22 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1096.eqiad.wmnet with OS trixie * 12:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2096.codfw.wmnet with reason: host reimage * 12:14 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1141.eqiad.wmnet with OS trixie * 12:10 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2096.codfw.wmnet with reason: host reimage * 12:10 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1289.eqiad.wmnet * 12:05 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1289.eqiad.wmnet * 12:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1288.eqiad.wmnet * 11:59 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1288.eqiad.wmnet * 11:59 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1287.eqiad.wmnet * 11:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1097.eqiad.wmnet with OS trixie * 11:54 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1287.eqiad.wmnet * 11:54 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1286.eqiad.wmnet * 11:53 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1141.eqiad.wmnet with reason: host reimage * 11:51 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2096.codfw.wmnet with OS trixie * 11:49 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1141.eqiad.wmnet with reason: host reimage * 11:48 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1286.eqiad.wmnet * 11:48 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1284.eqiad.wmnet * 11:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2095.codfw.wmnet with OS trixie * 11:43 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1284.eqiad.wmnet * 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1283.eqiad.wmnet * 11:42 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on ml-serve1015.eqiad.wmnet with reason: Downtime to get full picture of current BIOS settings beyond what Redfish shows * 11:39 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad * 11:39 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:37 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1283.eqiad.wmnet * 11:37 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1282.eqiad.wmnet * 11:37 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad * 11:37 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:33 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1141 * 11:33 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1141 * 11:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 11:32 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1141 * 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1141.eqiad.wmnet 156.48.64.10.in-addr.arpa 6.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:32 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1141.eqiad.wmnet 156.48.64.10.in-addr.arpa 6.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1141 - jayme@cumin1003" * 11:32 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1141 - jayme@cumin1003" * 11:32 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1282.eqiad.wmnet * 11:32 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1281.eqiad.wmnet * 11:32 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad * 11:32 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:29 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 11:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1097.eqiad.wmnet with reason: host reimage * 11:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2095.codfw.wmnet with OS trixie * 11:26 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1281.eqiad.wmnet * 11:26 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1280.eqiad.wmnet * 11:25 jayme@cumin1003: START - Cookbook sre.dns.netbox * 11:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1097.eqiad.wmnet with reason: host reimage * 11:22 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1141 * 11:21 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1141.eqiad.wmnet with OS trixie * 11:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1280.eqiad.wmnet * 11:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1279.eqiad.wmnet * 11:20 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin with reason: upgrade new Nokia swtiches in eqsin to SR Linux v26 * 11:17 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1141.eqiad.wmnet * 11:16 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1141.eqiad.wmnet * 11:16 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1141.eqiad.wmnet * 11:16 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be2095.codfw.wmnet with OS trixie * 11:15 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1279.eqiad.wmnet * 11:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1278.eqiad.wmnet * 11:14 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1139.eqiad.wmnet * 11:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1139.eqiad.wmnet * 11:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1070.eqiad.wmnet with OS trixie * 11:13 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1140.eqiad.wmnet * 11:13 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1140.eqiad.wmnet * 11:13 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1140.eqiad.wmnet * 11:09 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1278.eqiad.wmnet * 11:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1071.eqiad.wmnet with OS trixie * 11:05 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1097.eqiad.wmnet with OS trixie * 11:04 mvernon@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be1097.eqiad.wmnet with OS trixie * 11:02 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1140.eqiad.wmnet with OS trixie * 11:02 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1097.eqiad.wmnet with OS trixie * 11:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1096.eqiad.wmnet with OS trixie * 11:00 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1139.eqiad.wmnet * 11:00 jayme@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1139.eqiad.wmnet with OS trixie * 10:56 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1069.eqiad.wmnet with OS trixie * 10:56 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 10:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1070.eqiad.wmnet with reason: host reimage * 10:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 10:45 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1071.eqiad.wmnet with reason: host reimage * 10:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 10:41 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1096.eqiad.wmnet with OS trixie * 10:41 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1140.eqiad.wmnet with reason: host reimage * 10:39 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1071.eqiad.wmnet with reason: host reimage * 10:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1070.eqiad.wmnet with reason: host reimage * 10:38 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1139.eqiad.wmnet with reason: host reimage * 10:37 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1140.eqiad.wmnet with reason: host reimage * 10:35 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1069.eqiad.wmnet with reason: host reimage * 10:33 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1095.eqiad.wmnet with OS trixie * 10:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 10:33 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1139.eqiad.wmnet with reason: host reimage * 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1069.eqiad.wmnet with reason: host reimage * 10:23 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1140 * 10:23 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1140 * 10:23 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 10:22 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1140 * 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1140.eqiad.wmnet 155.48.64.10.in-addr.arpa 5.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:21 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1071.eqiad.wmnet with OS trixie * 10:21 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1140.eqiad.wmnet 155.48.64.10.in-addr.arpa 5.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1140 - jayme@cumin1003" * 10:21 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1140 - jayme@cumin1003" * 10:21 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1071 * 10:21 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1070.eqiad.wmnet with OS trixie * 10:21 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1070 * 10:20 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 10:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1095.eqiad.wmnet with OS trixie * 10:17 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1139 * 10:17 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1139 * 10:17 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be1095.eqiad.wmnet with OS trixie * 10:15 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1139 * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1139.eqiad.wmnet 194.32.64.10.in-addr.arpa 4.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:15 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1139.eqiad.wmnet 194.32.64.10.in-addr.arpa 4.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1139 - jayme@cumin1003" * 10:15 jayme@cumin1003: START - Cookbook sre.dns.netbox * 10:15 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1139 - jayme@cumin1003" * 10:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2095.codfw.wmnet with OS trixie * 10:13 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1069.eqiad.wmnet with OS trixie * 10:12 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1069 * 10:11 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1140 * 10:11 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1140.eqiad.wmnet with OS trixie * 10:11 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1140.eqiad.wmnet * 10:10 jayme@cumin1003: START - Cookbook sre.dns.netbox * 10:10 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1139 * 10:10 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1140.eqiad.wmnet * 10:10 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1140.eqiad.wmnet * 10:10 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1139.eqiad.wmnet with OS trixie * 10:09 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1139.eqiad.wmnet * 10:08 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1139.eqiad.wmnet * 10:08 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1139.eqiad.wmnet * 10:01 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2094.codfw.wmnet with OS trixie * 09:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 09:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 09:53 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 09:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 09:44 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:44 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2094.codfw.wmnet with reason: host reimage * 09:34 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2094.codfw.wmnet with reason: host reimage * 09:34 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1071 * 09:33 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1070 * 09:33 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1095.eqiad.wmnet with OS trixie * 09:27 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1069 * 09:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:23 brouberol@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM archiva1002.wikimedia.org * 09:20 brouberol@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM archiva1002.wikimedia.org * 09:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1277.eqiad.wmnet * 09:13 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2094.codfw.wmnet with OS trixie * 09:13 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 09:12 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 09:12 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:12 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1277.eqiad.wmnet * 09:12 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1276.eqiad.wmnet * 09:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1094.eqiad.wmnet with OS trixie * 09:06 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1276.eqiad.wmnet * 09:06 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1275.eqiad.wmnet * 09:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2093.codfw.wmnet with OS trixie * 09:01 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1275.eqiad.wmnet * 09:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1274.eqiad.wmnet * 08:56 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1274.eqiad.wmnet * 08:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1273.eqiad.wmnet * 08:50 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1273.eqiad.wmnet * 08:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1272.eqiad.wmnet * 08:49 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:49 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1094.eqiad.wmnet with reason: host reimage * 08:45 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1272.eqiad.wmnet * 08:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1094.eqiad.wmnet with reason: host reimage * 08:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2093.codfw.wmnet with reason: host reimage * 08:38 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:38 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2093.codfw.wmnet with reason: host reimage * 08:35 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 08:34 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 08:29 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:28 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:26 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1271.eqiad.wmnet * 08:23 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1094.eqiad.wmnet with OS trixie * 08:21 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 08:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1271.eqiad.wmnet * 08:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1270.eqiad.wmnet * 08:15 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1270.eqiad.wmnet * 08:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1269.eqiad.wmnet * 08:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2093.codfw.wmnet with OS trixie * 08:09 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1269.eqiad.wmnet * 08:09 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1268.eqiad.wmnet * 08:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2092.codfw.wmnet with OS trixie * 08:04 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1268.eqiad.wmnet * 08:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1267.eqiad.wmnet * 07:59 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1267.eqiad.wmnet * 07:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1093.eqiad.wmnet with OS trixie * 07:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1266.eqiad.wmnet * 07:56 jynus: running extra backups to test db1285 [[phab:T433826|T433826]] * 07:51 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1266.eqiad.wmnet * 07:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2092.codfw.wmnet with reason: host reimage * 07:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1093.eqiad.wmnet with reason: host reimage * 07:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2092.codfw.wmnet with reason: host reimage * 07:32 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1093.eqiad.wmnet with reason: host reimage * 07:29 jynus: running extra backups to test db1265 [[phab:T433825|T433825]] * 07:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2092.codfw.wmnet with OS trixie * 07:11 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1093.eqiad.wmnet with OS trixie * 06:50 slyngshede@dns1004: END - running authdns-update * 06:48 slyngshede@dns1004: START - running authdns-update * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.11 (duration: 02m 29s) * 03:38 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] (duration: 32m 57s) * 03:23 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 03:22 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 03:05 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 32s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 00:45 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] (duration: 06m 20s) * 00:41 cjming@deploy1003: cjming: Continuing with deployment * 00:41 cjming@deploy1003: cjming: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:39 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] == 2026-08-03 == * 23:58 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply * 23:57 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply * 23:29 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cp5021.eqsin.wmnet * 23:29 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cp5021.eqsin.wmnet * 23:27 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cp5021.eqsin.wmnet * 23:26 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cp5021.eqsin.wmnet * 23:18 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 23:17 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 22:56 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: sync * 22:56 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: sync * 22:36 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 22:36 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 22:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host search-loader2002.codfw.wmnet with OS trixie * 21:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on search-loader2002.codfw.wmnet with reason: host reimage * 21:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on search-loader2002.codfw.wmnet with reason: host reimage * 21:42 dancy@deploy1003: Stopping before sync operations * 21:41 dancy@deploy1003: Started scap sync-world: testing * 21:39 dancy@deploy1003: Installation of scap version "4.277.0" completed for 3 hosts * 21:37 dancy@deploy1003: Installing scap version "4.277.0" for 3 host(s) * 21:37 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] (duration: 06m 13s) * 21:33 dancy@deploy1003: dancy: Continuing with deployment * 21:32 dancy@deploy1003: dancy: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:31 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] * 21:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host search-loader2002.codfw.wmnet with OS trixie * 21:03 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] (duration: 06m 34s) * 20:59 dancy@deploy1003: dancy: Continuing with deployment * 20:58 dancy@deploy1003: dancy: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:56 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] * 20:52 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] (duration: 06m 23s) * 20:48 cjming@deploy1003: cjming: Continuing with deployment * 20:47 cjming@deploy1003: cjming: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:46 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] * 20:42 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] (duration: 07m 36s) * 20:38 arlolra@deploy1003: arlolra: Continuing with deployment * 20:36 arlolra@deploy1003: arlolra: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:34 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] * 20:16 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] (duration: 08m 26s) * 20:12 krinkle@deploy1003: krinkle: Continuing with deployment * 20:09 krinkle@deploy1003: krinkle: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] * 19:45 jasmine@cumin2002: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-main-codfw * 18:58 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] (duration: 09m 23s) * 18:53 krinkle@deploy1003: krinkle: Continuing with deployment * 18:53 jasmine@cumin2002: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-main-codfw * 18:50 krinkle@deploy1003: krinkle: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:48 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] * 18:37 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] (duration: 10m 13s) * 18:34 dzahn@cumin2002: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host codesearch2001.codfw.wmnet * 18:34 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host codesearch2001.codfw.wmnet with OS trixie * 18:33 krinkle@deploy1003: krinkle: Continuing with deployment * 18:29 krinkle@deploy1003: krinkle: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:27 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] * 18:19 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 18:18 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on codesearch2001.codfw.wmnet with reason: host reimage * 18:14 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 18:14 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:12 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on codesearch2001.codfw.wmnet with reason: host reimage * 18:11 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 18:11 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 18:10 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 18:02 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 18:02 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 18:01 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 18:01 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 17:55 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host codesearch2001.codfw.wmnet with OS trixie * 17:54 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:54 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) codesearch2001.codfw.wmnet on all recursors * 17:53 dzahn@cumin2002: START - Cookbook sre.dns.wipe-cache codesearch2001.codfw.wmnet on all recursors * 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:48 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:41 dzahn@cumin2002: START - Cookbook sre.dns.netbox * 17:41 dzahn@cumin2002: START - Cookbook sre.ganeti.makevm for new host codesearch2001.codfw.wmnet * 17:37 dzahn@cumin2002: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host codesearch1001.eqiad.wmnet * 17:37 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host codesearch1001.eqiad.wmnet with OS trixie * 17:24 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on codesearch1001.eqiad.wmnet with reason: host reimage * 17:17 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on codesearch1001.eqiad.wmnet with reason: host reimage * 17:08 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host codesearch1001.eqiad.wmnet with OS trixie * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 17:06 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:06 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) codesearch1001.eqiad.wmnet on all recursors * 17:06 dzahn@cumin2002: START - Cookbook sre.dns.wipe-cache codesearch1001.eqiad.wmnet on all recursors * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 17:05 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 17:04 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:04 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 16:58 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 16:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2091.codfw.wmnet with OS trixie * 16:54 ebernhardson@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 16:54 ebernhardson@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 16:49 ebernhardson@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 16:49 ebernhardson@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 16:46 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1092.eqiad.wmnet with OS trixie * 16:43 ebernhardson@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 16:43 ebernhardson@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 16:43 dzahn@cumin2002: START - Cookbook sre.dns.netbox * 16:43 dzahn@cumin2002: START - Cookbook sre.ganeti.makevm for new host codesearch1001.eqiad.wmnet * 16:41 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 16:41 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 16:40 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2091.codfw.wmnet with reason: host reimage * 16:37 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 16:35 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2091.codfw.wmnet with reason: host reimage * 16:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1092.eqiad.wmnet with reason: host reimage * 16:24 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 16:24 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 16:23 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1092.eqiad.wmnet with reason: host reimage * 16:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2091.codfw.wmnet with OS trixie * 16:03 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1092.eqiad.wmnet with OS trixie * 16:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2090.codfw.wmnet with OS trixie * 15:51 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 15:51 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 15:51 jiji@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 15:50 jiji@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 15:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2090.codfw.wmnet with reason: host reimage * 15:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2090.codfw.wmnet with reason: host reimage * 15:33 jhathaway@dns1004: END - running authdns-update * 15:31 jhathaway@dns1004: START - running authdns-update * 15:26 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1091.eqiad.wmnet with OS trixie * 15:25 dancy@deploy1003: Installation of scap version "4.276.1" completed for 3 hosts * 15:23 dancy@deploy1003: Installing scap version "4.276.1" for 3 host(s) * 15:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2090.codfw.wmnet with OS trixie * 15:12 marostegui@cumin1003: dbctl commit (dc=all): 'Repool db2245, db2246, db2247 and db2248 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95857 and previous config saved to /var/cache/conftool/dbconfig/20260803-151212-marostegui.json * 15:09 dancy@deploy1003: Started scap sync-world: testing * 15:09 dancy@deploy1003: Installation of scap version "4.277.0" completed for 3 hosts * 15:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1091.eqiad.wmnet with reason: host reimage * 15:07 dancy@deploy1003: Installing scap version "4.277.0" for 3 host(s) * 15:03 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1091.eqiad.wmnet with reason: host reimage * 14:49 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1091.eqiad.wmnet with OS trixie * 14:33 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2089.codfw.wmnet with OS trixie * 14:29 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 14:27 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 14:18 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 14:16 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 14:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2089.codfw.wmnet with reason: host reimage * 14:10 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 14:10 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 14:09 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2089.codfw.wmnet with reason: host reimage * 13:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2089.codfw.wmnet with OS trixie * 13:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2088.codfw.wmnet with OS trixie * 13:40 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1090.eqiad.wmnet with OS trixie * 13:22 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1090.eqiad.wmnet with reason: host reimage * 13:22 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] (duration: 14m 34s) * 13:19 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1090.eqiad.wmnet with reason: host reimage * 13:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2088.codfw.wmnet with reason: host reimage * 13:16 aude@deploy1003: aude, mhorsey: Continuing with deployment * 13:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2088.codfw.wmnet with reason: host reimage * 13:12 aude@deploy1003: aude, mhorsey: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] * 13:05 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1090.eqiad.wmnet with OS trixie * 12:58 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2088.codfw.wmnet with OS trixie * 12:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db[2245-2247].codfw.wmnet * 12:49 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2247: Rebooting db2247.codfw.wmnet * 12:49 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2247: Rebooting db2247.codfw.wmnet * 12:42 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2246: Rebooting db2246.codfw.wmnet * 12:42 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2246: Rebooting db2246.codfw.wmnet * 12:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2087.codfw.wmnet with OS trixie * 12:37 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1089.eqiad.wmnet with OS trixie * 12:34 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2245: Rebooting db2245.codfw.wmnet * 12:34 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2245: Rebooting db2245.codfw.wmnet * 12:34 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db[2245-2247].codfw.wmnet * 12:32 kamila@deploy1003: Finished scap sync-world: rebuild after base image update (duration: 30m 26s) * 12:28 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 12:22 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2087.codfw.wmnet with reason: host reimage * 12:19 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1089.eqiad.wmnet with reason: host reimage * 12:14 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2087.codfw.wmnet with reason: host reimage * 12:14 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1089.eqiad.wmnet with reason: host reimage * 12:03 kamila@deploy1003: Started scap sync-world: rebuild after base image update * 12:00 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1089.eqiad.wmnet with OS trixie * 12:00 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2087.codfw.wmnet with OS trixie * 11:35 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db[2245-2248].codfw.wmnet * 11:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db[2245-2248].codfw.wmnet * 11:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2086.codfw.wmnet with OS trixie * 11:26 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 11:26 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 11:25 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 11:25 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 11:24 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1088.eqiad.wmnet with OS trixie * 11:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db[2245-2248].codfw.wmnet with reason: Checking network * 11:21 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 11:20 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 11:19 marostegui@dns1004: END - running authdns-update * 11:17 marostegui@dns1004: START - running authdns-update * 11:10 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:10 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 11:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2086.codfw.wmnet with reason: host reimage * 11:09 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:08 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 11:08 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:07 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 11:07 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:07 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 11:06 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop: apply * 11:06 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop: apply * 11:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1088.eqiad.wmnet with reason: host reimage * 11:05 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop: apply * 11:04 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop: apply * 11:04 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop: apply * 11:04 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop: apply * 11:02 marostegui@dns1004: END - running authdns-update * 11:02 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2086.codfw.wmnet with reason: host reimage * 11:01 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1088.eqiad.wmnet with reason: host reimage * 11:00 marostegui@dns1004: START - running authdns-update * 10:53 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] (duration: 10m 57s) * 10:51 cmooney@dns3003: END - running authdns-update * 10:49 cmooney@dns3003: START - running authdns-update * 10:47 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1088.eqiad.wmnet with OS trixie * 10:47 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2086.codfw.wmnet with OS trixie * 10:47 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 10:46 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:46 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:46 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new reverse ranges for eqsin CR switch links - cmooney@cumin1003" * 10:46 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new reverse ranges for eqsin CR switch links - cmooney@cumin1003" * 10:42 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] * 10:41 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 10:36 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2245, db2246 and db2247 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95855 and previous config saved to /var/cache/conftool/dbconfig/20260803-103652-marostegui.json * 10:35 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2248 from s4 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95854 and previous config saved to /var/cache/conftool/dbconfig/20260803-103535-marostegui.json * 10:27 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 10:27 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 10:26 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 10:24 kart_: cxserver: Add referencePunctuation config ([[phab:T97231|T97231]]) * 10:24 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 10:23 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:23 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:23 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:22 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:22 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply * 10:21 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply * 10:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2085.codfw.wmnet with OS trixie * 10:20 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply * 10:20 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply * 10:18 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply * 10:18 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply * 10:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1087.eqiad.wmnet with OS trixie * 09:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2085.codfw.wmnet with reason: host reimage * 09:43 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1087.eqiad.wmnet with reason: host reimage * 09:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2085.codfw.wmnet with reason: host reimage * 09:40 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1087.eqiad.wmnet with reason: host reimage * 09:26 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1087.eqiad.wmnet with OS trixie * 09:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2085.codfw.wmnet with OS trixie * 09:13 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2084.codfw.wmnet with OS trixie * 09:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1086.eqiad.wmnet with OS trixie * 08:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2084.codfw.wmnet with reason: host reimage * 08:50 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2084.codfw.wmnet with reason: host reimage * 08:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1086.eqiad.wmnet with reason: host reimage * 08:39 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1086.eqiad.wmnet with reason: host reimage * 08:38 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:38 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:37 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 08:37 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 08:35 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2084.codfw.wmnet with OS trixie * 08:34 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 08:34 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:27 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1086.eqiad.wmnet with OS trixie * 08:09 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1218: Repool after a crash * 08:07 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2083.codfw.wmnet with OS trixie * 08:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1085.eqiad.wmnet with OS trixie * 07:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2083.codfw.wmnet with reason: host reimage * 07:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1085.eqiad.wmnet with reason: host reimage * 07:40 kart_: Updated cxsever to 2026-07-16-140518-production ([[phab:T97231|T97231]]) * 07:39 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply * 07:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2083.codfw.wmnet with reason: host reimage * 07:38 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply * 07:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1085.eqiad.wmnet with reason: host reimage * 07:37 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] (duration: 32m 40s) * 07:33 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply * 07:33 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply * 07:25 jdlrobson@deploy1003: jdlrobson: Continuing with deployment * 07:24 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2083.codfw.wmnet with OS trixie * 07:24 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1085.eqiad.wmnet with OS trixie * 07:23 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1218: Repool after a crash * 07:21 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:09 marostegui: Drop renamed tables [[phab:T425074|T425074]] * 07:04 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] * 06:55 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply * 06:54 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 46s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-02 == * 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 01m 03s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-01 == * 03:30 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:30 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:30 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:30 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 34s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-31 == * 17:41 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 17:41 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 17:40 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 17:40 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 15:33 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2195: Testing * 15:02 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:02 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 15:02 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 14:48 pt1979@cumin2002: START - Cookbook sre.dns.netbox * 14:47 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2195: Testing * 14:22 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 14:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2195: Testing * 14:21 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 14:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2195.codfw.wmnet with reason: Testing * 14:16 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 14:04 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 14:04 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 13:30 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1048.eqiad.wmnet with OS trixie * 13:22 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2195: Testing * 13:22 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 13:19 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2195: Testing * 13:18 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 13:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2195: Testing * 13:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 13:05 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 13:05 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 13:04 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 13:04 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 12:50 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lswtest-d8-eqiad * 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:53 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:42 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 6515 * 11:37 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 6515 * 11:28 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:27 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:07 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2082.codfw.wmnet with OS trixie * 10:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2082.codfw.wmnet with reason: host reimage * 10:42 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2082.codfw.wmnet with reason: host reimage * 10:28 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2082.codfw.wmnet with OS trixie * 10:02 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 09:52 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 09:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts * 09:16 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts * 08:57 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 08:46 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:42 gkyziridis@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 08:37 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 08:37 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 08:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 08:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 08:11 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:11 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:08 filippo@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudvirt1048 * 08:07 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 08:07 filippo@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudvirt1048 * 08:06 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 08:01 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 08:00 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 07:19 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1048.eqiad.wmnet with reason: host reimage * 07:13 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1048.eqiad.wmnet with reason: host reimage * 07:11 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 07:11 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 07:09 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 07:09 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 06:57 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:56 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:48 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:48 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:44 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie * 06:34 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1048.eqiad.wmnet with OS trixie * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 54s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 00:57 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] (duration: 11m 04s) * 00:53 dreamyjazz@deploy1003: dreamyjazz, jforrester: Continuing with deployment * 00:48 dreamyjazz@deploy1003: dreamyjazz, jforrester: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:46 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] == 2026-07-30 == * 21:37 dancy@deploy1003: Installation of scap version "4.276.1" completed for 3 hosts * 21:35 dancy@deploy1003: Installing scap version "4.276.1" for 3 host(s) * 21:24 dancy@deploy1003: Installation of scap version "4.276.0" completed for 3 hosts * 21:22 dancy@deploy1003: Installing scap version "4.276.0" for 3 host(s) * 21:15 maryum: Deployed security fix for [[phab:T430601|T430601]] * 20:13 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] (duration: 09m 20s) * 20:07 arlolra@deploy1003: osleger, arlolra: Continuing with deployment * 20:05 arlolra@deploy1003: osleger, arlolra: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:03 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] * 19:29 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:29 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:25 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service * 19:24 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 19:24 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:24 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:24 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 19:23 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service * 19:20 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1084.eqiad.wmnet with OS trixie * 18:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1084.eqiad.wmnet with reason: host reimage * 18:52 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1084.eqiad.wmnet with reason: host reimage * 18:41 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 18:40 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 18:39 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1084.eqiad.wmnet with OS trixie * 18:25 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 18:15 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 17:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1083.eqiad.wmnet with OS trixie * 17:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1048: Maintenance * 17:36 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new security plugin settings - bking@cumin2003 - [[phab:T350516|T350516]] * 17:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1083.eqiad.wmnet with reason: host reimage * 17:26 inflatador: bking@apt1002 `reprepro --noskipold --component thirdparty/opensearch3 update trixie-wikimedia` [[phab:T433624|T433624]] * 17:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1083.eqiad.wmnet with reason: host reimage * 17:23 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 17:20 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 17:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2081.codfw.wmnet with OS trixie * 17:11 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new security plugin settings - bking@cumin2003 - [[phab:T350516|T350516]] * 17:10 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1083.eqiad.wmnet with OS trixie * 16:55 root@cumin1003: START - Cookbook sre.mysql.pool pool es1048: Maintenance * 16:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2081.codfw.wmnet with reason: host reimage * 16:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1048 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95833 and previous config saved to /var/cache/conftool/dbconfig/20260730-165053-cwilliams.json * 16:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1048.eqiad.wmnet with reason: Maintenance * 16:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1040: Maintenance * 16:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2081.codfw.wmnet with reason: host reimage * 16:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1082.eqiad.wmnet with OS trixie * 16:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2081.codfw.wmnet with OS trixie * 16:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2097.codfw.wmnet with OS trixie * 16:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1082.eqiad.wmnet with reason: host reimage * 16:08 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1082.eqiad.wmnet with reason: host reimage * 16:08 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 16:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1047: Maintenance * 16:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2080.codfw.wmnet with OS trixie * 16:04 root@cumin1003: START - Cookbook sre.mysql.pool pool es1040: Maintenance * 16:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1040: Maintenance * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new logging settings - bking@cumin2003 - [[phab:T324335|T324335]] * 15:58 root@cumin1003: START - Cookbook sre.mysql.pool pool es1040: Maintenance * 15:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1040 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95827 and previous config saved to /var/cache/conftool/dbconfig/20260730-155324-cwilliams.json * 15:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1040.eqiad.wmnet with reason: Maintenance * 15:50 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1082.eqiad.wmnet with OS trixie * 15:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2048: Maintenance * 15:44 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 24s) * 15:43 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2080.codfw.wmnet with reason: host reimage * 15:38 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new logging settings - bking@cumin2003 - [[phab:T324335|T324335]] * 15:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2080.codfw.wmnet with reason: host reimage * 15:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 15:30 mvernon@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be2097.codfw.wmnet with OS trixie * 15:23 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1081.eqiad.wmnet with OS trixie * 15:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2098.codfw.wmnet with OS trixie * 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - mvernon@cumin2003" * 15:18 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be2097.codfw.wmnet with OS trixie * 15:18 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - mvernon@cumin2003" * 15:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 15:17 root@cumin1003: START - Cookbook sre.mysql.pool pool es1047: Maintenance * 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2080.codfw.wmnet with OS trixie * 15:13 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be2097.codfw.wmnet with OS trixie * 15:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1047 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95820 and previous config saved to /var/cache/conftool/dbconfig/20260730-151200-cwilliams.json * 15:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1047.eqiad.wmnet with reason: Maintenance * 15:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: Maintenance * 15:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 15:04 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 15:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1081.eqiad.wmnet with reason: host reimage * 15:00 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1081.eqiad.wmnet with reason: host reimage * 15:00 root@cumin1003: START - Cookbook sre.mysql.pool pool es2048: Maintenance * 15:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 14:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2079.codfw.wmnet with OS trixie * 14:56 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 14:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2048 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95816 and previous config saved to /var/cache/conftool/dbconfig/20260730-145510-cwilliams.json * 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2048.codfw.wmnet with reason: Maintenance * 14:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2040: Maintenance * 14:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 14:51 tchin@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] (duration: 06m 48s) * 14:47 tchin@deploy1003: jforrester, tchin: Continuing with deployment * 14:47 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 14:47 tchin@deploy1003: jforrester, tchin: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:45 tchin@deploy1003: Started scap sync-world: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] * 14:42 sukhe@puppetserver1001: conftool action : set/weight=1; selector: cluster=urldownloader,service=squid * 14:42 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader,service=squid * 14:42 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1081.eqiad.wmnet with OS trixie * 14:39 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 14:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2079.codfw.wmnet with reason: host reimage * 14:36 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2098.codfw.wmnet with OS trixie * 14:32 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2079.codfw.wmnet with reason: host reimage * 14:30 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] (duration: 06m 31s) * 14:27 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 14:26 mszwarc@deploy1003: mszwarc: Continuing with deployment * 14:25 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:25 root@cumin1003: START - Cookbook sre.mysql.pool pool es1038: Maintenance * 14:25 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1038: Maintenance * 14:23 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] * 14:21 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] (duration: 11m 19s) * 14:20 root@cumin1003: START - Cookbook sre.mysql.pool pool es1038: Maintenance * 14:14 stran@deploy1003: stran: Continuing with deployment * 14:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1038 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95810 and previous config saved to /var/cache/conftool/dbconfig/20260730-141439-cwilliams.json * 14:14 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1038.eqiad.wmnet with reason: Maintenance * 14:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1036: Maintenance * 14:13 stran@deploy1003: stran: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2079.codfw.wmnet with OS trixie * 14:09 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] * 14:08 root@cumin1003: START - Cookbook sre.mysql.pool pool es2040: Maintenance * 14:08 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2040: Maintenance * 14:03 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] (duration: 31m 41s) * 14:03 root@cumin1003: START - Cookbook sre.mysql.pool pool es2040: Maintenance * 14:03 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2040 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95806 and previous config saved to /var/cache/conftool/dbconfig/20260730-135643-cwilliams.json * 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2040.codfw.wmnet with reason: Maintenance * 13:56 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:56 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2038: Maintenance * 13:55 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:52 lucaswerkmeister-wmde@deploy1003: migr, lucaswerkmeister-wmde: Continuing with deployment * 13:49 lucaswerkmeister-wmde@deploy1003: migr, lucaswerkmeister-wmde: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:49 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:48 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 13:45 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 13:32 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:32 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] * 13:28 root@cumin1003: START - Cookbook sre.mysql.pool pool es1036: Maintenance * 13:28 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1036: Maintenance * 13:22 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:22 root@cumin1003: START - Cookbook sre.mysql.pool pool es1036: Maintenance * 13:20 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1036 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95800 and previous config saved to /var/cache/conftool/dbconfig/20260730-131727-cwilliams.json * 13:17 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1036.eqiad.wmnet with reason: Maintenance * 13:17 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] (duration: 10m 31s) * 13:16 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2022\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 13:13 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, stran: Continuing with deployment * 13:10 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:10 root@cumin1003: START - Cookbook sre.mysql.pool pool es2038: Maintenance * 13:10 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2038: Maintenance * 13:08 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, stran: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie * 13:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2047: Maintenance * 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048 cloud-private - filippo@cumin1003" * 13:07 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048 cloud-private - filippo@cumin1003" * 13:06 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] * 13:04 root@cumin1003: START - Cookbook sre.mysql.pool pool es2038: Maintenance * 13:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2078.codfw.wmnet with OS trixie * 13:01 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2038 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95797 and previous config saved to /var/cache/conftool/dbconfig/20260730-125919-cwilliams.json * 12:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2038.codfw.wmnet with reason: Maintenance * 12:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2078.codfw.wmnet with reason: host reimage * 12:37 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2078.codfw.wmnet with reason: host reimage * 12:37 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] (duration: 06m 51s) * 12:33 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 12:32 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 12:32 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 12:32 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:30 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] * 12:19 root@cumin1003: START - Cookbook sre.mysql.pool pool es2047: Maintenance * 12:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2078.codfw.wmnet with OS trixie * 12:18 dcausse@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 12:18 dcausse@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 12:15 dcausse@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 12:14 dcausse@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 12:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2047 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95793 and previous config saved to /var/cache/conftool/dbconfig/20260730-121404-cwilliams.json * 12:13 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2047.codfw.wmnet with reason: Maintenance * 12:13 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2036: Maintenance * 12:05 ayounsi@dns1004: END - running authdns-update * 12:02 ayounsi@dns1004: START - running authdns-update * 11:51 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2077.codfw.wmnet with OS trixie * 11:48 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:46 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1080.eqiad.wmnet with OS trixie * 11:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1226: Maintenance * 11:41 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2077.codfw.wmnet with reason: host reimage * 11:28 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2077.codfw.wmnet with reason: host reimage * 11:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1080.eqiad.wmnet with reason: host reimage * 11:27 root@cumin1003: START - Cookbook sre.mysql.pool pool es2036: Maintenance * 11:27 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2036: Maintenance * 11:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1080.eqiad.wmnet with reason: host reimage * 11:21 root@cumin1003: START - Cookbook sre.mysql.pool pool es2036: Maintenance * 11:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2036 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95786 and previous config saved to /var/cache/conftool/dbconfig/20260730-111633-cwilliams.json * 11:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2036.codfw.wmnet with reason: Maintenance * 11:08 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2077.codfw.wmnet with OS trixie * 11:07 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 11:03 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie * 11:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1226: Maintenance * 10:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1226 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95783 and previous config saved to /var/cache/conftool/dbconfig/20260730-104801-cwilliams.json * 10:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1226.eqiad.wmnet with reason: Maintenance * 10:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1214: Maintenance * 10:27 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2035: Maintenance * 10:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2076.codfw.wmnet with OS trixie * 10:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1214: Maintenance * 09:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1214 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95775 and previous config saved to /var/cache/conftool/dbconfig/20260730-095451-cwilliams.json * 09:54 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1214.eqiad.wmnet with reason: Maintenance * 09:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1209: Maintenance * 09:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2076.codfw.wmnet with reason: host reimage * 09:42 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool es2035: Maintenance * 09:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.netbox.update-extras (exit_code=0) rolling restart_daemons on A:netbox * 09:41 ayounsi@cumin1003: START - Cookbook sre.netbox.update-extras rolling restart_daemons on A:netbox * 09:40 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2035: Maintenance * 09:39 ayounsi@cumin1003: END (PASS) - Cookbook sre.netbox.update-extras (exit_code=0) rolling restart_daemons on A:netbox-canary * 09:39 ayounsi@cumin1003: START - Cookbook sre.netbox.update-extras rolling restart_daemons on A:netbox-canary * 09:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2076.codfw.wmnet with reason: host reimage * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 09:34 root@cumin1003: START - Cookbook sre.mysql.pool pool es2035: Maintenance * 09:32 lucaswerkmeister-wmde@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 09:32 lucaswerkmeister-wmde@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 09:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2035 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95771 and previous config saved to /var/cache/conftool/dbconfig/20260730-092910-cwilliams.json * 09:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2035.codfw.wmnet with reason: Maintenance * 09:19 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2076.codfw.wmnet with OS trixie * 09:18 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 09:17 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie * 09:07 root@cumin1003: START - Cookbook sre.mysql.pool pool db1209: Maintenance * 09:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 23 hosts * 09:04 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Remove cable label from interfaces descriptions - ayounsi@cumin1003 * 09:04 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:02 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Remove cable label from interfaces descriptions - ayounsi@cumin1003 * 09:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1209 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95767 and previous config saved to /var/cache/conftool/dbconfig/20260730-090133-cwilliams.json * 09:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1209.eqiad.wmnet with reason: Maintenance * 09:01 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1192: Maintenance * 08:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1252: Maintenance * 08:57 jayme@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 08:56 jayme@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 08:53 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 23 hosts * 08:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:51 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:50 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1263: Maintenance * 08:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:23 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 08:15 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie * 08:14 root@cumin1003: START - Cookbook sre.mysql.pool pool db1192: Maintenance * 08:13 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2075.codfw.wmnet with OS trixie * 08:12 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1252: Maintenance * 08:11 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1252: Maintenance * 08:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1252: Maintenance * 08:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1192 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95754 and previous config saved to /var/cache/conftool/dbconfig/20260730-080611-cwilliams.json * 08:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1192.eqiad.wmnet with reason: Maintenance * 08:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1178: Maintenance * 08:05 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS bullseye * 07:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1263: Maintenance * 07:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1263 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95751 and previous config saved to /var/cache/conftool/dbconfig/20260730-075106-cwilliams.json * 07:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[1260-1262].eqiad.wmnet with reason: Maintenance * 07:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2075.codfw.wmnet with reason: host reimage * 07:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1263.eqiad.wmnet with reason: Maintenance * 07:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2075.codfw.wmnet with reason: host reimage * 07:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance * 07:38 dcausse@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:38 dcausse@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 07:35 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db1252', diff saved to https://phabricator.wikimedia.org/P95748 and previous config saved to /var/cache/conftool/dbconfig/20260730-073510-marostegui.json * 07:26 klausman@dns2004: END - running authdns-update * 07:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 07:25 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2075.codfw.wmnet with OS trixie * 07:24 klausman@dns2004: START - running authdns-update * 07:23 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host an-test-master1003.eqiad.wmnet * 07:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db1178: Maintenance * 07:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1178 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95746 and previous config saved to /var/cache/conftool/dbconfig/20260730-071112-cwilliams.json * 07:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1178.eqiad.wmnet with reason: Maintenance * 07:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1177: Maintenance * 06:24 root@cumin1003: START - Cookbook sre.mysql.pool pool db1177: Maintenance * 06:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1177 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95741 and previous config saved to /var/cache/conftool/dbconfig/20260730-061736-cwilliams.json * 06:17 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1177.eqiad.wmnet with reason: Maintenance * 06:17 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1172: Maintenance * 05:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1218.eqiad.wmnet with reason: crashed * 05:41 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db1217 it crashed', diff saved to https://phabricator.wikimedia.org/P95737 and previous config saved to /var/cache/conftool/dbconfig/20260730-054111-marostegui.json * 05:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95736 and previous config saved to /var/cache/conftool/dbconfig/20260730-053422-cwilliams.json * 05:30 root@cumin1003: START - Cookbook sre.mysql.pool pool db1172: Maintenance * 05:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95734 and previous config saved to /var/cache/conftool/dbconfig/20260730-052414-cwilliams.json * 05:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1172 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95733 and previous config saved to /var/cache/conftool/dbconfig/20260730-052354-cwilliams.json * 05:23 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1172.eqiad.wmnet with reason: Maintenance * 05:23 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1167: Maintenance * 05:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95731 and previous config saved to /var/cache/conftool/dbconfig/20260730-051406-cwilliams.json * 05:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95729 and previous config saved to /var/cache/conftool/dbconfig/20260730-050358-cwilliams.json * 04:47 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 04:47 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 04:47 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 04:35 root@cumin1003: START - Cookbook sre.mysql.pool pool db1167: Maintenance * 04:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1167 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95726 and previous config saved to /var/cache/conftool/dbconfig/20260730-042923-cwilliams.json * 04:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 04:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1167.eqiad.wmnet with reason: Maintenance * 04:22 pt1979@cumin2002: START - Cookbook sre.dns.netbox * 04:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95725 and previous config saved to /var/cache/conftool/dbconfig/20260730-040337-cwilliams.json * 04:03 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:38 brett@cumin2002: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool eqsin [reason: Switch upgrade maintenance window complete, [[phab:T433097|T433097]]] * 01:38 brett@cumin2002: START - Cookbook sre.dns.admin DNS admin: pool eqsin [reason: Switch upgrade maintenance window complete, [[phab:T433097|T433097]]] == 2026-07-29 == * 23:57 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin,mr1-eqsin IPv6,mr1-eqsin.oob,mr1-eqsin.oob IPv6 with reason: connection issue * 22:54 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 22:53 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 22:53 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 22:53 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:25 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2022.codfw.wmnet, repooling source-only afterwards * 22:20 brett@cumin2002: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool eqsin [reason: Switch upgrade maintenance window, [[phab:T433097|T433097]]] * 22:20 brett@cumin2002: START - Cookbook sre.dns.admin DNS admin: depool eqsin [reason: Switch upgrade maintenance window, [[phab:T433097|T433097]]] * 22:01 apine@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 22:00 apine@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 21:59 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 21:58 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 21:58 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 21:58 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 21:32 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1048.eqiad.wmnet with OS trixie * 21:25 pt1979@cumin2002: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be2097.codfw.wmnet with OS bullseye * 21:16 zabe@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki=metawiki 'Mental Health Resource Center' 'Safety Resource Center/Mental Health' Zabe --reason 'per request [[:phab:T433118{{!}}T433118]]' * 21:12 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2022.codfw.wmnet, repooling source-only afterwards * 21:12 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] (duration: 12m 53s) * 21:12 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2015\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 21:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1253: Maintenance * 21:08 aaron@deploy1003: aaron: Continuing with deployment * 21:01 aaron@deploy1003: aaron: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:59 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] * 20:52 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] (duration: 21m 57s) * 20:48 aaron@deploy1003: aaron: Continuing with deployment * 20:32 aaron@deploy1003: aaron: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:30 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] * 20:24 root@cumin1003: START - Cookbook sre.mysql.pool pool db1253: Maintenance * 20:19 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] (duration: 08m 07s) * 20:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1253 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95719 and previous config saved to /var/cache/conftool/dbconfig/20260729-201810-cwilliams.json * 20:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1253.eqiad.wmnet with reason: Maintenance * 20:17 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1231: Maintenance * 20:15 aaron@deploy1003: bpirkle, aaron: Continuing with deployment * 20:13 aaron@deploy1003: bpirkle, aaron: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:12 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie * 20:11 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] * 20:11 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1048.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:09 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1048.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:09 pt1979@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 20:04 pt1979@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 19:47 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 19:43 pt1979@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye * 19:41 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 19:37 zabe: zabe@deploy1003:~$ mwscript-k8s --comment='[[phab:T433529|T433529]]' --follow -- resetAuthenticationThrottle.php --wiki=aawiki --signup --ip=89.36.114.94 * 19:36 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] (duration: 06m 49s) * 19:32 zabe@deploy1003: zabe: Continuing with deployment * 19:31 zabe@deploy1003: zabe: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:31 root@cumin1003: START - Cookbook sre.mysql.pool pool db1231: Maintenance * 19:29 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] * 19:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1231 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95714 and previous config saved to /var/cache/conftool/dbconfig/20260729-192454-cwilliams.json * 19:24 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1231.eqiad.wmnet with reason: Maintenance * 19:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1227: Maintenance * 19:22 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 19:22 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 19:21 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:21 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1048] - vriley@cumin1003" * 19:21 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1048] - vriley@cumin1003" * 19:19 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 19:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1251: Maintenance * 19:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95711 and previous config saved to /var/cache/conftool/dbconfig/20260729-191756-cwilliams.json * 19:16 vriley@cumin1003: START - Cookbook sre.dns.netbox * 19:11 dduvall: rolling back wmf.13 to group0 due to [[phab:T433457|T433457]] (cc [[phab:T430832|T430832]]) * 19:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95709 and previous config saved to /var/cache/conftool/dbconfig/20260729-190748-cwilliams.json * 19:01 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2022.codfw.wmnet with OS bookworm * 19:01 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2021.codfw.wmnet, repooling source-only afterwards * 18:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95707 and previous config saved to /var/cache/conftool/dbconfig/20260729-185740-cwilliams.json * 18:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95704 and previous config saved to /var/cache/conftool/dbconfig/20260729-184732-cwilliams.json * 18:37 root@cumin1003: START - Cookbook sre.mysql.pool pool db1227: Maintenance * 18:34 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2022.codfw.wmnet with reason: host reimage * 18:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1227 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95701 and previous config saved to /var/cache/conftool/dbconfig/20260729-183117-cwilliams.json * 18:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1227.eqiad.wmnet with reason: Maintenance * 18:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1202: Maintenance * 18:30 root@cumin1003: START - Cookbook sre.mysql.pool pool db1251: Maintenance * 18:27 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2022.codfw.wmnet with reason: host reimage * 18:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1251 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95698 and previous config saved to /var/cache/conftool/dbconfig/20260729-182428-cwilliams.json * 18:24 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lvs2014.codfw.wmnet * 18:24 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for lvs2014.codfw.wmnet * 18:24 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1251.eqiad.wmnet with reason: Maintenance * 18:23 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1235: Maintenance * 18:22 brett@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 18:19 brett@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 18:19 brett@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 18:17 brett@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 18:17 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 18:16 mutante: removing jenkins during the train - living on the edge - no, just kidding, jenkins has migrated to dedicated machines, nothing should happen * 18:15 brett@cumin2002: END (ERROR) - Cookbook sre.loadbalancer.restart-pybal (exit_code=97) rolling-restart of pybal on P<nowiki>{</nowiki>lvs2014.codfw.wmnet<nowiki>}</nowiki> and A:lvs ([[phab:T428495|T428495]]) * 18:15 mutante: CI: contint1002/contint2002: apt-get remove --purge jenkins - jenkins be gone - [[phab:T418521|T418521]] * 18:13 brett@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on P<nowiki>{</nowiki>lvs2014.codfw.wmnet<nowiki>}</nowiki> and A:lvs ([[phab:T428495|T428495]]) * 18:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2022 * 18:08 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2022 * 18:03 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T428495|T428495]] * 18:03 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2022 * 18:02 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2022.codfw.wmnet 211.48.192.10.in-addr.arpa 1.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:02 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2022.codfw.wmnet 211.48.192.10.in-addr.arpa 1.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:02 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:02 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2022 - bking@cumin2003" * 18:02 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2022 - bking@cumin2003" * 17:57 bking@cumin2003: START - Cookbook sre.dns.netbox * 17:56 brett@cumin2002: END (FAIL) - Cookbook sre.loadbalancer.restart-pybal (exit_code=1) rolling-restart of pybal on A:lvs-codfw and A:lvs ([[phab:T428495|T428495]]) * 17:55 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo - [[phab:T428495|T428495]] * 17:54 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2022 * 17:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2022.codfw.wmnet with OS bookworm * 17:50 brett@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on A:lvs-codfw and A:lvs ([[phab:T428495|T428495]]) * 17:47 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2021.codfw.wmnet, repooling source-only afterwards * 17:47 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 14s) * 17:47 swfrench-wmf: authdns-update to direct codfw, eqsin, ulsfo etcd clients back to codfw - [[phab:T428495|T428495]] * 17:47 swfrench@dns1004: END - running authdns-update * 17:47 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 17:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95692 and previous config saved to /var/cache/conftool/dbconfig/20260729-174713-cwilliams.json * 17:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance * 17:46 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1249: Maintenance * 17:45 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2015\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 17:45 swfrench@dns1004: START - running authdns-update * 17:44 root@cumin1003: START - Cookbook sre.mysql.pool pool db1202: Maintenance * 17:41 akhatun: Deployed refinery using scap, then deployed onto hdfs * 17:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1202 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95688 and previous config saved to /var/cache/conftool/dbconfig/20260729-173759-cwilliams.json * 17:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1202.eqiad.wmnet with reason: Maintenance * 17:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1194: Maintenance * 17:37 root@cumin1003: START - Cookbook sre.mysql.pool pool db1235: Maintenance * 17:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1230: Maintenance * 17:30 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1235 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95684 and previous config saved to /var/cache/conftool/dbconfig/20260729-173051-cwilliams.json * 17:30 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1235.eqiad.wmnet with reason: Maintenance * 17:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1234: Maintenance * 17:26 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (thin): Regular analytics weekly train THIN [analytics/refinery@56695674] (duration: 02m 02s) * 17:24 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (thin): Regular analytics weekly train THIN [analytics/refinery@56695674] * 17:23 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567]: Regular analytics weekly train [analytics/refinery@56695674] (duration: 06m 20s) * 17:20 dancy@deploy1003: Finished scap sync-world: Testing delay_messageblobstore_purge: true (duration: 06m 29s) * 17:17 akhatun@deploy1003: Started deploy [analytics/refinery@5669567]: Regular analytics weekly train [analytics/refinery@56695674] * 17:17 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] (duration: 00m 22s) * 17:16 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] * 17:13 dancy@deploy1003: Started scap sync-world: Testing delay_messageblobstore_purge: true * 17:05 mutante: CI: contint1002/contint2002 - restarted httpd to be extra sure all is cleaned up - https://integration.wikimedia.org/ci/ is up and running [[phab:T418521|T418521]] * 17:04 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 17:03 mutante: CI: contint1002/contint2002 - rm /etc/apache2/jenkins_proxy - removing legacy jenkins proxy config - jenkins is on new dedicated machines and uses jenkins_proxy_ext config [[phab:T418521|T418521]] * 17:02 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] (duration: 36m 25s) * 17:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1249: Maintenance * 16:59 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 16:54 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2015.codfw.wmnet, repooling source-only afterwards * 16:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1249 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95674 and previous config saved to /var/cache/conftool/dbconfig/20260729-165339-cwilliams.json * 16:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1249.eqiad.wmnet with reason: Maintenance * 16:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1248: Maintenance * 16:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1194: Maintenance * 16:47 swfrench-wmf: silenced EtcdReplicationDown 57b2b421-1cc9-4e38-9276-{{Gerrit|94f223fd231c}} - [[phab:T428495|T428495]] * 16:46 tchin@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/eventstreams-internal: apply * 16:46 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye * 16:46 tchin@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/eventstreams-internal: apply * 16:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1230: Maintenance * 16:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1194 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95669 and previous config saved to /var/cache/conftool/dbconfig/20260729-164422-cwilliams.json * 16:44 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 16:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1194.eqiad.wmnet with reason: Maintenance * 16:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1191: Maintenance * 16:43 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host an-test-master1003.eqiad.wmnet * 16:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db1234: Maintenance * 16:43 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 16:43 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Rolling back deployment * 16:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host an-test-master1004.eqiad.wmnet * 16:41 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 16:40 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 16:40 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 16:39 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 16:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1230 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95667 and previous config saved to /var/cache/conftool/dbconfig/20260729-163932-cwilliams.json * 16:39 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 16:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1230.eqiad.wmnet with reason: Maintenance * 16:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1207: Maintenance * 16:38 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 16:37 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host an-test-master1004.eqiad.wmnet * 16:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1234 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95664 and previous config saved to /var/cache/conftool/dbconfig/20260729-163719-cwilliams.json * 16:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1234.eqiad.wmnet with reason: Maintenance * 16:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1079.eqiad.wmnet with OS trixie * 16:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1232: Maintenance * 16:34 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1259: Maintenance * 16:28 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] (duration: 06m 57s) * 16:28 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:26 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] * 16:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1051 hosts * 16:21 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] * 16:20 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2006.codfw.wmnet with OS bookworm * 16:19 akhatun: Deploying Refinery at {{Gerrit|56695674}} as part of weekly train * 16:18 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1079.eqiad.wmnet with reason: host reimage * 16:16 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] (duration: 15m 36s) * 16:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2021.codfw.wmnet with OS bookworm * 16:14 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1079.eqiad.wmnet with reason: host reimage * 16:12 topranks: hot-swap line card in FPC0 on cr1-eqiad with replacement MPC10E from Juniper [[phab:T426343|T426343]] * 16:10 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Continuing with deployment * 16:07 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db1248: Maintenance * 16:01 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] * 16:00 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 16:00 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 15:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1248 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95651 and previous config saved to /var/cache/conftool/dbconfig/20260729-155956-cwilliams.json * 15:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1248.eqiad.wmnet with reason: Maintenance * 15:59 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2006.codfw.wmnet with reason: host reimage * 15:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1247: Maintenance * 15:59 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2074.codfw.wmnet with OS trixie * 15:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1191: Maintenance * 15:57 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:55 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1079.eqiad.wmnet with OS trixie * 15:55 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2006.codfw.wmnet with reason: host reimage * 15:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1207: Maintenance * 15:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1191 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95646 and previous config saved to /var/cache/conftool/dbconfig/20260729-155104-cwilliams.json * 15:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1191.eqiad.wmnet with reason: Maintenance * 15:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1181: Maintenance * 15:49 root@cumin1003: START - Cookbook sre.mysql.pool pool db1232: Maintenance * 15:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2021.codfw.wmnet with reason: host reimage * 15:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1207 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95643 and previous config saved to /var/cache/conftool/dbconfig/20260729-154735-cwilliams.json * 15:47 root@cumin1003: START - Cookbook sre.mysql.pool pool db1259: Maintenance * 15:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1207.eqiad.wmnet with reason: Maintenance * 15:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1200: Maintenance * 15:46 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] (duration: 31m 59s) * 15:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 15:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1232 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95640 and previous config saved to /var/cache/conftool/dbconfig/20260729-154330-cwilliams.json * 15:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1232.eqiad.wmnet with reason: Maintenance * 15:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1219: Maintenance * 15:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2015.codfw.wmnet, repooling source-only afterwards * 15:41 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 18s) * 15:41 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1259 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95638 and previous config saved to /var/cache/conftool/dbconfig/20260729-154107-cwilliams.json * 15:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1259.eqiad.wmnet with reason: Maintenance * 15:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1254: Maintenance * 15:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2015.codfw.wmnet with OS bookworm * 15:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2021.codfw.wmnet with reason: host reimage * 15:36 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2006.codfw.wmnet with OS bookworm * 15:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 15:35 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Continuing with deployment * 15:33 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2074.codfw.wmnet with OS trixie * 15:33 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 15:32 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:29 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:28 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be2074.codfw.wmnet with OS trixie * 15:28 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2006.codfw.wmnet * 15:26 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1078.eqiad.wmnet with OS trixie * 15:25 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 15:25 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:22 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2006.codfw.wmnet * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2021 * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2021 * 15:19 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2021 * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2021.codfw.wmnet 210.48.192.10.in-addr.arpa 0.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:19 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2021.codfw.wmnet 210.48.192.10.in-addr.arpa 0.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2021 - bking@cumin2003" * 15:19 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2021 - bking@cumin2003" * 15:14 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] * 15:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2015.codfw.wmnet with reason: host reimage * 15:11 root@cumin1003: START - Cookbook sre.mysql.pool pool db1247: Maintenance * 15:11 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on ml-serve2004.codfw.wmnet with reason: [[phab:T433478|T433478]] * 15:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2015.codfw.wmnet with reason: host reimage * 15:10 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on ml-serve2002.codfw.wmnet with reason: [[phab:T433476|T433476]] * 15:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 15:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1247 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95625 and previous config saved to /var/cache/conftool/dbconfig/20260729-150459-cwilliams.json * 15:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1247.eqiad.wmnet with reason: Maintenance * 15:04 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:04 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1244: Maintenance * 15:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1078.eqiad.wmnet with reason: host reimage * 15:03 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1005.wikimedia.org * 15:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db1181: Maintenance * 15:01 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2098.codfw.wmnet with OS bullseye * 15:00 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye * 15:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1200: Maintenance * 14:59 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 14:59 root@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1285.eqiad.wmnet with OS trixie * 14:59 Amir1: mwscript-k8s -- extensions/TimedMediaHandler/maintenance/requeueTranscodes.php --wiki=commonswiki --key '360p.mpeg4.mov' --throttle --video --missing ([[phab:T358266|T358266]]) * 14:58 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1005.wikimedia.org * 14:58 jhancock@cumin2002: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['ms-be2098'] * 14:58 jhancock@cumin2002: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['ms-be2098'] * 14:58 jhancock@cumin2002: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['ms-be2097'] * 14:58 jhancock@cumin2002: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['ms-be2097'] * 14:58 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1078.eqiad.wmnet with reason: host reimage * 14:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1181 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95621 and previous config saved to /var/cache/conftool/dbconfig/20260729-145629-cwilliams.json * 14:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1181.eqiad.wmnet with reason: Maintenance * 14:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1174: Maintenance * 14:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1219: Maintenance * 14:55 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader1006.wikimedia.org on all recursors * 14:55 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader1006.wikimedia.org on all recursors * 14:55 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader1005.wikimedia.org on all recursors * 14:55 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader1005.wikimedia.org on all recursors * 14:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1200 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95618 and previous config saved to /var/cache/conftool/dbconfig/20260729-145336-cwilliams.json * 14:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1254: Maintenance * 14:53 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1200.eqiad.wmnet with reason: Maintenance * 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1185: Maintenance * 14:52 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2021 * 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2015 * 14:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2015 * 14:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1219 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95616 and previous config saved to /var/cache/conftool/dbconfig/20260729-144946-cwilliams.json * 14:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1219.eqiad.wmnet with reason: Maintenance * 14:49 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1218: Maintenance * 14:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2021.codfw.wmnet with OS bookworm * 14:48 dancy@deploy1003: Finished deploy [zuul/deploy@22703a6]: Deploying https://gerrit.wikimedia.org/r/c/integration/zuul/+/1311501 ([[phab:T432491|T432491]]) (duration: 00m 15s) * 14:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2015.codfw.wmnet with OS bookworm * 14:48 dancy@deploy1003: Started deploy [zuul/deploy@22703a6]: Deploying https://gerrit.wikimedia.org/r/c/integration/zuul/+/1311501 ([[phab:T432491|T432491]]) * 14:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1254 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95613 and previous config saved to /var/cache/conftool/dbconfig/20260729-144729-cwilliams.json * 14:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1254.eqiad.wmnet with reason: Maintenance * 14:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1233: Maintenance * 14:46 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2013\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 14:46 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2014\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 14:46 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:45 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:44 root@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1285.eqiad.wmnet with reason: host reimage * 14:43 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:42 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:41 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:40 root@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1285.eqiad.wmnet with reason: host reimage * 14:39 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1078.eqiad.wmnet with OS trixie * 14:39 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2074.codfw.wmnet with OS trixie * 14:32 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:32 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:32 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:31 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2005.codfw.wmnet with OS bookworm * 14:30 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:30 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:29 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:29 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:27 root@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host db1285 * 14:27 root@cumin1003: START - Cookbook sre.hosts.move-vlan for host db1285 * 14:27 root@cumin1003: START - Cookbook sre.hosts.reimage for host db1285.eqiad.wmnet with OS trixie * 14:24 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 14:24 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:24 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:24 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:23 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:22 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:22 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:21 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:17 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:16 root@cumin1003: START - Cookbook sre.mysql.pool pool db1244: Maintenance * 14:15 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:15 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add asw1-604 loopback ipv4 - pt1979@cumin2002" * 14:15 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add asw1-604 loopback ipv4 - pt1979@cumin2002" * 14:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:12 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 14:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95599 and previous config saved to /var/cache/conftool/dbconfig/20260729-141014-cwilliams.json * 14:10 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 14:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1244.eqiad.wmnet with reason: Maintenance * 14:10 pt1979@cumin2002: START - Cookbook sre.dns.netbox * 14:10 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 14:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1243: Maintenance * 14:09 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2005.codfw.wmnet with reason: host reimage * 14:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db1174: Maintenance * 14:08 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad * 14:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db1185: Maintenance * 14:06 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 14:05 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2005.codfw.wmnet with reason: host reimage * 14:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1174 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95595 and previous config saved to /var/cache/conftool/dbconfig/20260729-140309-cwilliams.json * 14:03 sukhe@dns1004: END - running authdns-update * 14:03 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1174.eqiad.wmnet with reason: Maintenance * 14:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1170: Maintenance * 14:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db1218: Maintenance * 14:01 sukhe@dns1004: START - running authdns-update * 14:00 sukhe@dns1004: START - running authdns-update * 13:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db1233: Maintenance * 13:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1185 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95592 and previous config saved to /var/cache/conftool/dbconfig/20260729-135925-cwilliams.json * 13:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1185.eqiad.wmnet with reason: Maintenance * 13:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1161: Maintenance * 13:58 sukhe@puppetserver1001: conftool action : set/pooled=true; selector: dnsdisc=urldownloader * 13:58 root@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1265.eqiad.wmnet with OS trixie * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1218 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95590 and previous config saved to /var/cache/conftool/dbconfig/20260729-135621-cwilliams.json * 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1218.eqiad.wmnet with reason: Maintenance * 13:55 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1206: Maintenance * 13:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2073.codfw.wmnet with OS trixie * 13:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1233 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95587 and previous config saved to /var/cache/conftool/dbconfig/20260729-135335-cwilliams.json * 13:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1233.eqiad.wmnet with reason: Maintenance * 13:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1229: Maintenance * 13:50 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/kartotherian: apply * 13:50 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service * 13:49 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:49 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/kartotherian: apply * 13:48 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 13:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1077.eqiad.wmnet with OS trixie * 13:47 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 13:46 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2005.codfw.wmnet with OS bookworm * 13:44 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 13:44 root@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1265.eqiad.wmnet with reason: host reimage * 13:40 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] (duration: 09m 22s) * 13:39 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:38 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:36 root@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1265.eqiad.wmnet with reason: host reimage * 13:35 stran@deploy1003: stran: Continuing with deployment * 13:33 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 13:32 stran@deploy1003: stran: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified t * 13:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2073.codfw.wmnet with reason: host reimage * 13:30 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ml-build1001.eqiad.wmnet * 13:30 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] * 13:29 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad * 13:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1077.eqiad.wmnet with reason: host reimage * 13:27 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 13:27 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:27 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad * 13:26 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2073.codfw.wmnet with reason: host reimage * 13:26 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] (duration: 07m 56s) * 13:25 klausman@cumin1003: START - Cookbook sre.hosts.reboot-single for host ml-build1001.eqiad.wmnet * 13:24 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 13:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>ml-serve101[2-5].eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 13:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1015.eqiad.wmnet * 13:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1015.eqiad.wmnet * 13:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1077.eqiad.wmnet with reason: host reimage * 13:23 root@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host db1265 * 13:23 root@cumin1003: START - Cookbook sre.hosts.move-vlan for host db1265 * 13:23 root@cumin1003: START - Cookbook sre.hosts.reimage for host db1265.eqiad.wmnet with OS trixie * 13:23 root@cumin1003: START - Cookbook sre.mysql.pool pool db1243: Maintenance * 13:22 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 13:22 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2005.codfw.wmnet * 13:22 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 13:22 samtar@deploy1003: dreamrimmer, samtar: Continuing with deployment * 13:20 samtar@deploy1003: dreamrimmer, samtar: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts an-test-master[1001-1002].eqiad.wmnet * 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-master[1001-1002].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 13:18 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1015.eqiad.wmnet * 13:18 sukhe@cumin1003: END (ERROR) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=97) for role: url_downloader@eqiad * 13:18 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 13:18 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] * 13:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95574 and previous config saved to /var/cache/conftool/dbconfig/20260729-131638-cwilliams.json * 13:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1243.eqiad.wmnet with reason: Maintenance * 13:16 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2005.codfw.wmnet * 13:16 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1242: Maintenance * 13:14 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] (duration: 07m 00s) * 13:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db1170: Maintenance * 13:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1015.eqiad.wmnet * 13:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1014.eqiad.wmnet * 13:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1014.eqiad.wmnet * 13:12 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1223: Maintenance * 13:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db1161: Maintenance * 13:10 samtar@deploy1003: anzx, samtar: Continuing with deployment * 13:09 samtar@deploy1003: anzx, samtar: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db1206: Maintenance * 13:08 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 13:07 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] * 13:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1170 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95566 and previous config saved to /var/cache/conftool/dbconfig/20260729-130730-cwilliams.json * 13:07 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1170.eqiad.wmnet with reason: Maintenance * 13:07 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:07 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt IPs new switches - cmooney@cumin1003" * 13:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1158: Maintenance * 13:06 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1014.eqiad.wmnet * 13:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1161 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95564 and previous config saved to /var/cache/conftool/dbconfig/20260729-130616-cwilliams.json * 13:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 13:06 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1077.eqiad.wmnet with OS trixie * 13:05 root@cumin1003: START - Cookbook sre.mysql.pool pool db1229: Maintenance * 13:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1161.eqiad.wmnet with reason: Maintenance * 13:05 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt IPs new switches - cmooney@cumin1003" * 13:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2073.codfw.wmnet with OS trixie * 13:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Maintenance * 13:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1206 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95562 and previous config saved to /var/cache/conftool/dbconfig/20260729-130258-cwilliams.json * 13:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1206.eqiad.wmnet with reason: Maintenance * 13:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1196: Maintenance * 13:01 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 13:01 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:00 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 13:00 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 12:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1229 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95559 and previous config saved to /var/cache/conftool/dbconfig/20260729-125950-cwilliams.json * 12:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1229.eqiad.wmnet with reason: Maintenance * 12:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1222: Maintenance * 12:57 sukhe: sudo cumin 'A:lvs and (A:eqiad or A:codfw)' 'disable-puppet "adding new service urldownloader"': [[phab:T429175|T429175]] * 12:56 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1014.eqiad.wmnet * 12:56 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1013.eqiad.wmnet * 12:56 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1013.eqiad.wmnet * 12:50 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1013.eqiad.wmnet * 12:50 sukhe: sudo cumin 'O:url_downloader' 'run-puppet-agent --enable "merging CR 1313948"': [[phab:T429175|T429175]] * 12:48 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-master[1001-1002].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 12:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1013.eqiad.wmnet * 12:45 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1012.eqiad.wmnet * 12:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1012.eqiad.wmnet * 12:45 sukhe: sudo cumin 'O:url_downloader' 'disable-puppet "merging CR 1313948"': [[phab:T429175|T429175]] * 12:44 btullis@cumin1003: START - Cookbook sre.dns.netbox * 12:40 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test2001.codfw.wmnet * 12:40 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test2001.codfw.wmnet * 12:38 ayounsi@dns1004: END - running authdns-update * 12:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1012.eqiad.wmnet * 12:35 ayounsi@dns1004: START - running authdns-update * 12:34 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts an-test-master[1001-1002].eqiad.wmnet * 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts an-test-coord1001.eqiad.wmnet * 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-coord1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 12:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1012.eqiad.wmnet * 12:32 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>ml-serve101[2-5].eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 12:29 root@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Maintenance * 12:25 root@cumin1003: START - Cookbook sre.mysql.pool pool db1223: Maintenance * 12:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95544 and previous config saved to /var/cache/conftool/dbconfig/20260729-122254-cwilliams.json * 12:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1242.eqiad.wmnet with reason: Maintenance * 12:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1241: Maintenance * 12:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1051 hosts * 12:20 root@cumin1003: START - Cookbook sre.mysql.pool pool db1158: Maintenance * 12:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1223 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95540 and previous config saved to /var/cache/conftool/dbconfig/20260729-121937-cwilliams.json * 12:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1223.eqiad.wmnet with reason: Maintenance * 12:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1212: Maintenance * 12:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db1159: Maintenance * 12:17 elukey@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: sync * 12:15 elukey@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: sync * 12:15 root@cumin1003: START - Cookbook sre.mysql.pool pool db1196: Maintenance * 12:14 Daimona: Creating new DB tables for the CampaignEvents extension in x1.testwiki, x1.test2wiki, x1.officewiki, and x1.wikishared # [[phab:T429339|T429339]] * 12:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db1222: Maintenance * 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95535 and previous config saved to /var/cache/conftool/dbconfig/20260729-121211-cwilliams.json * 12:12 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 12:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1158.eqiad.wmnet with reason: Maintenance * 12:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1159 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95534 and previous config saved to /var/cache/conftool/dbconfig/20260729-121146-cwilliams.json * 12:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1159.eqiad.wmnet with reason: Maintenance * 12:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1196 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95533 and previous config saved to /var/cache/conftool/dbconfig/20260729-120847-cwilliams.json * 12:08 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 12:08 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1196.eqiad.wmnet with reason: Maintenance * 12:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1195: Maintenance * 12:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1222 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95530 and previous config saved to /var/cache/conftool/dbconfig/20260729-120424-cwilliams.json * 12:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1222.eqiad.wmnet with reason: Maintenance * 12:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1098 hosts * 12:00 marostegui: Rename tables [[phab:T425074|T425074]] * 12:00 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-coord1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 11:58 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1197: Maintenance * 11:55 btullis@cumin1003: START - Cookbook sre.dns.netbox * 11:52 elukey@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: sync * 11:51 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:51 elukey@deploy1003: helmfile [codfw] START helmfile.d/services/proton: sync * 11:51 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:50 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts an-test-coord1001.eqiad.wmnet * 11:50 elukey@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: sync * 11:49 elukey@deploy1003: helmfile [staging] START helmfile.d/services/proton: sync * 11:35 root@cumin1003: START - Cookbook sre.mysql.pool pool db1241: Maintenance * 11:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db1212: Maintenance * 11:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1241 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95520 and previous config saved to /var/cache/conftool/dbconfig/20260729-112918-cwilliams.json * 11:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1241.eqiad.wmnet with reason: Maintenance * 11:29 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1238: Maintenance * 11:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1212 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95517 and previous config saved to /var/cache/conftool/dbconfig/20260729-112727-cwilliams.json * 11:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 11:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1212.eqiad.wmnet with reason: Maintenance * 11:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1198: Maintenance * 11:23 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:22 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:21 root@cumin1003: START - Cookbook sre.mysql.pool pool db1195: Maintenance * 11:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1195 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95514 and previous config saved to /var/cache/conftool/dbconfig/20260729-111450-cwilliams.json * 11:14 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1195.eqiad.wmnet with reason: Maintenance * 11:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1186: Maintenance * 11:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 11:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 11:05 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:54 marostegui: Dropping renamed tables [[phab:T425066|T425066]] * 10:41 root@cumin1003: START - Cookbook sre.mysql.pool pool db1238: Maintenance * 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1198: Maintenance * 10:39 Amir1: ran https://phabricator.wikimedia.org/T432509#12149723 in production ([[phab:T432509|T432509]]) * 10:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db1197: Maintenance * 10:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1238 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95501 and previous config saved to /var/cache/conftool/dbconfig/20260729-103532-cwilliams.json * 10:35 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1238.eqiad.wmnet with reason: Maintenance * 10:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1221: Maintenance * 10:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1198 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95499 and previous config saved to /var/cache/conftool/dbconfig/20260729-103330-cwilliams.json * 10:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1198.eqiad.wmnet with reason: Maintenance * 10:33 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1175: Maintenance * 10:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1197 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95496 and previous config saved to /var/cache/conftool/dbconfig/20260729-103217-cwilliams.json * 10:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1197.eqiad.wmnet with reason: Maintenance * 10:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1188: Maintenance * 10:27 root@cumin1003: START - Cookbook sre.mysql.pool pool db1186: Maintenance * 10:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1186 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95493 and previous config saved to /var/cache/conftool/dbconfig/20260729-102111-cwilliams.json * 10:21 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1186.eqiad.wmnet with reason: Maintenance * 10:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 10:14 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 09:53 XioNoX: reboot cr2-magru - [[phab:T431750|T431750]] * 09:52 XioNoX: drain cr2-magru - [[phab:T431750|T431750]] * 09:48 root@cumin1003: START - Cookbook sre.mysql.pool pool db1221: Maintenance * 09:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zookeeper-test1002.eqiad.wmnet * 09:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1188: Maintenance * 09:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1175: Maintenance * 09:44 btullis@dns1004: END - running authdns-update * 09:42 btullis@dns1004: START - running authdns-update * 09:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1221 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95483 and previous config saved to /var/cache/conftool/dbconfig/20260729-094200-cwilliams.json * 09:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 7 hosts with reason: Maintenance * 09:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1221.eqiad.wmnet with reason: Maintenance * 09:41 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host zookeeper-test1002.eqiad.wmnet * 09:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1199: Maintenance * 09:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1188 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95481 and previous config saved to /var/cache/conftool/dbconfig/20260729-093917-cwilliams.json * 09:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1188.eqiad.wmnet with reason: Maintenance * 09:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1182: Maintenance * 09:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1175 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95479 and previous config saved to /var/cache/conftool/dbconfig/20260729-093842-cwilliams.json * 09:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1175.eqiad.wmnet with reason: Maintenance * 09:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1166: Maintenance * 09:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1169: Maintenance * 09:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1033.eqiad.wmnet,service=s8 * 09:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1033.eqiad.wmnet,service=s5 * 09:33 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1033.eqiad.wmnet,service=s8 * 09:33 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1033.eqiad.wmnet,service=s5 * 09:21 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:21 XioNoX: reboot cr1-magru - [[phab:T431750|T431750]] * 09:17 XioNoX: drain cr1-magru - [[phab:T431750|T431750]] * 09:15 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm1001.wikimedia.org * 09:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply * 09:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply * 09:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 09:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 09:11 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr2-magru,cr2-magru IPv6,cr2-magru.mgmt with reason: router upgrade * 09:11 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm1001.wikimedia.org * 09:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp1005.wikimedia.org * 09:07 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp1005.wikimedia.org * 09:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp2005.wikimedia.org * 09:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp2005.wikimedia.org * 09:00 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr1-magru,cr1-magru IPv6,cr1-magru.mgmt with reason: router upgrade * 09:00 marostegui: Dropping renamed tables [[phab:T426341|T426341]] * 08:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1199: Maintenance * 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1182: Maintenance * 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1166: Maintenance * 08:47 root@cumin1003: START - Cookbook sre.mysql.pool pool db1169: Maintenance * 08:46 ayounsi@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 1:00:00 on cr1-magru,cr1-magru IPv6,cr1-magru.mgmt with reason: router upgrade * 08:45 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2072.codfw.wmnet with OS trixie * 08:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1199 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95464 and previous config saved to /var/cache/conftool/dbconfig/20260729-084534-cwilliams.json * 08:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1199.eqiad.wmnet with reason: Maintenance * 08:45 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 08:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1190: Maintenance * 08:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1182 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95462 and previous config saved to /var/cache/conftool/dbconfig/20260729-084436-cwilliams.json * 08:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1182.eqiad.wmnet with reason: Maintenance * 08:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1156: Maintenance * 08:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1166 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95460 and previous config saved to /var/cache/conftool/dbconfig/20260729-084400-cwilliams.json * 08:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1166.eqiad.wmnet with reason: Maintenance * 08:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1157: Maintenance * 08:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95458 and previous config saved to /var/cache/conftool/dbconfig/20260729-084147-cwilliams.json * 08:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1169.eqiad.wmnet with reason: Maintenance * 08:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1163: Maintenance * 08:30 btullis@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 11 hosts with reason: Replacing the namenodes * 08:23 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2072.codfw.wmnet with reason: host reimage * 08:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1098 hosts * 08:19 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2072.codfw.wmnet with reason: host reimage * 07:58 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2072.codfw.wmnet with OS trixie * 07:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1190: Maintenance * 07:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1156: Maintenance * 07:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1157: Maintenance * 07:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1163: Maintenance * 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1190 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95444 and previous config saved to /var/cache/conftool/dbconfig/20260729-074930-cwilliams.json * 07:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1190.eqiad.wmnet with reason: Maintenance * 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1157 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95443 and previous config saved to /var/cache/conftool/dbconfig/20260729-074914-cwilliams.json * 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1156 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95442 and previous config saved to /var/cache/conftool/dbconfig/20260729-074906-cwilliams.json * 07:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1157.eqiad.wmnet with reason: Maintenance * 07:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 07:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1156.eqiad.wmnet with reason: Maintenance * 07:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1163 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95441 and previous config saved to /var/cache/conftool/dbconfig/20260729-074652-cwilliams.json * 07:46 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1163.eqiad.wmnet with reason: Maintenance * 07:46 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2034.codfw.wmnet * 07:42 ayounsi@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2034.codfw.wmnet * 07:42 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2232.codfw.wmnet with OS trixie * 07:34 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:34 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:33 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:31 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:19 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2232.codfw.wmnet with reason: host reimage * 07:15 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2232.codfw.wmnet with reason: host reimage * 06:58 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db2232.codfw.wmnet with OS trixie * 06:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[2160,2232].codfw.wmnet with reason: Reimage * 06:26 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1164.eqiad.wmnet with OS trixie * 06:05 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1164.eqiad.wmnet with reason: host reimage * 06:01 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1164.eqiad.wmnet with reason: host reimage * 05:47 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1164.eqiad.wmnet with OS trixie * 05:46 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1164.eqiad.wmnet with reason: Reimage == 2026-07-28 == * 22:50 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1138.eqiad.wmnet * 22:50 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1138.eqiad.wmnet * 22:49 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1138.eqiad.wmnet * 22:11 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2014.codfw.wmnet, repooling source-only afterwards * 22:08 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2013.codfw.wmnet, repooling source-only afterwards * 22:03 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 20:58 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] (duration: 08m 19s) * 20:55 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2014.codfw.wmnet, repooling source-only afterwards * 20:55 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2013.codfw.wmnet, repooling source-only afterwards * 20:54 arlolra@deploy1003: arlolra: Continuing with deployment * 20:54 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 14s) * 20:54 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 20:53 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 30s) * 20:53 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 20:52 arlolra@deploy1003: arlolra: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:51 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:50 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] * 20:49 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:43 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 20:34 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] (duration: 06m 54s) * 20:34 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:34 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:31 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:31 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:30 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:30 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:30 arlolra@deploy1003: arlolra: Continuing with deployment * 20:29 arlolra@deploy1003: arlolra: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:27 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] * 20:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2014.codfw.wmnet with OS bookworm * 20:21 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 20:21 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:20 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 20:19 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:19 swfrench-wmf: switched etcd-mirror replication from conf2005 to conf2004 - [[phab:T428495|T428495]] * 20:17 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:17 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:15 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] (duration: 08m 26s) * 20:12 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:11 arlolra@deploy1003: anzx, arlolra: Continuing with deployment * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2013.codfw.wmnet with OS bookworm * 20:09 arlolra@deploy1003: anzx, arlolra: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] * 20:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2216: Maintenance * 19:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2014.codfw.wmnet with reason: host reimage * 19:57 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:54 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:54 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:53 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:52 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2014.codfw.wmnet with reason: host reimage * 19:49 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2013.codfw.wmnet with reason: host reimage * 19:42 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 19:41 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:41 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2013.codfw.wmnet with reason: host reimage * 19:39 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:39 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:39 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-eqiad: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 19:38 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2014 * 19:33 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2014 * 19:29 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2014.codfw.wmnet with OS bookworm * 19:28 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:27 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1006 * 19:26 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2012\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 19:26 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1006 * 19:26 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:26 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 19:25 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2013 * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2013 * 19:21 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2013 * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2013.codfw.wmnet 84.0.192.10.in-addr.arpa 4.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:21 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2013.codfw.wmnet 84.0.192.10.in-addr.arpa 4.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2013 - bking@cumin2003" * 19:21 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2013 - bking@cumin2003" * 19:21 vriley@cumin1003: START - Cookbook sre.dns.netbox * 19:20 root@cumin1003: START - Cookbook sre.mysql.pool pool db2216: Maintenance * 19:13 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2216 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95435 and previous config saved to /var/cache/conftool/dbconfig/20260728-191343-cwilliams.json * 19:13 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2216.codfw.wmnet with reason: Maintenance * 19:13 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2203: Maintenance * 19:06 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1005.eqiad.wmnet with OS trixie * 19:06 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 19:06 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 18:46 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 18:45 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:45 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:43 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:40 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:36 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-eqiad: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 18:35 dancy@deploy1003: Installation of scap version "4.275.0" completed for 3 hosts * 18:33 dancy@deploy1003: Installing scap version "4.275.0" for 3 host(s) * 18:32 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:32 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2097.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:30 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns3003.wikimedia.org [reason: pool for all services after reimaging] * 18:29 sukhe@dns1004: END - running authdns-update * 18:27 sukhe@dns1004: START - running authdns-update * 18:27 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns3003.wikimedia.org,service=authdns-update [reason: pool authdns-update after reimaging] * 18:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db2203: Maintenance * 18:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2203 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95430 and previous config saved to /var/cache/conftool/dbconfig/20260728-181958-cwilliams.json * 18:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2203.codfw.wmnet with reason: Maintenance * 18:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2188: Maintenance * 18:18 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2097.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:17 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-be2098 * 18:17 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host ms-be2098 * 18:17 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-be2097 * 18:16 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host ms-be2097 * 18:15 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:15 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding ms-be2097-8 to codfw - jhancock@cumin2002" * 18:15 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding ms-be2097-8 to codfw - jhancock@cumin2002" * 18:10 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 18:08 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage * 18:05 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns3003.wikimedia.org with OS trixie * 18:03 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage * 17:56 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-codfw: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 17:45 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie * 17:45 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1005.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:41 sukhe@dns1004: END - running authdns-update * 17:39 sukhe@dns1004: START - running authdns-update * 17:36 sukhe@puppetserver1001: conftool action : set/weight=1; selector: cluster=urldownloader,service=squid * 17:36 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1005.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:35 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader,service=squid * 17:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 17:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1005 * 17:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 17:34 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1005 * 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1005] - vriley@cumin1003" * 17:34 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1005] - vriley@cumin1003" * 17:32 root@cumin1003: START - Cookbook sre.mysql.pool pool db2188: Maintenance * 17:29 vriley@cumin1003: START - Cookbook sre.dns.netbox * 17:29 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 17:26 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2188 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95425 and previous config saved to /var/cache/conftool/dbconfig/20260728-172609-cwilliams.json * 17:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2188.codfw.wmnet with reason: Maintenance * 17:25 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2176: Maintenance * 17:19 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1138.eqiad.wmnet with OS trixie * 17:18 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1005 * 17:18 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1005 * 17:18 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:15 vriley@cumin1003: START - Cookbook sre.dns.netbox * 17:13 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns3003.wikimedia.org with reason: host reimage * 17:07 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns3003.wikimedia.org with reason: host reimage * 17:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-eqiad * 17:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1015.eqiad.wmnet * 17:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1015.eqiad.wmnet * 16:59 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1138.eqiad.wmnet with reason: host reimage * 16:55 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-codfw: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 16:54 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1138.eqiad.wmnet with reason: host reimage * 16:53 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1015.eqiad.wmnet * 16:43 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns3003.wikimedia.org with OS trixie * 16:43 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1015.eqiad.wmnet * 16:43 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1014.eqiad.wmnet * 16:43 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1014.eqiad.wmnet * 16:43 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=dns3003.wikimedia.org [reason: depooling for reimage to trixie] * 16:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 290 hosts * 16:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db2176: Maintenance * 16:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1138 * 16:38 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1138 * 16:37 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1138 * 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1138.eqiad.wmnet 193.32.64.10.in-addr.arpa 3.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:37 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1138.eqiad.wmnet 193.32.64.10.in-addr.arpa 3.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1138 - jiji@cumin1003" * 16:37 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1138 - jiji@cumin1003" * 16:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1014.eqiad.wmnet * 16:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2176 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95420 and previous config saved to /var/cache/conftool/dbconfig/20260728-163235-cwilliams.json * 16:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2176.codfw.wmnet with reason: Maintenance * 16:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1014.eqiad.wmnet * 16:32 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1013.eqiad.wmnet * 16:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1013.eqiad.wmnet * 16:32 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2174: Maintenance * 16:28 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2012.codfw.wmnet, repooling source-only afterwards * 16:25 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1013.eqiad.wmnet * 16:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1013.eqiad.wmnet * 16:20 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1012.eqiad.wmnet * 16:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1012.eqiad.wmnet * 16:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1012.eqiad.wmnet * 16:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1012.eqiad.wmnet * 16:03 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1011.eqiad.wmnet * 16:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1011.eqiad.wmnet * 16:00 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2004.codfw.wmnet with OS bookworm * 15:59 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1011.eqiad.wmnet * 15:56 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:55 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 15:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:54 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1011.eqiad.wmnet * 15:54 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1010.eqiad.wmnet * 15:54 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1010.eqiad.wmnet * 15:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:50 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 15:49 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1010.eqiad.wmnet * 15:48 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 15:48 jiji@cumin1003: START - Cookbook sre.dns.netbox * 15:46 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2248: Maintenance * 15:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db2174: Maintenance * 15:44 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1010.eqiad.wmnet * 15:44 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1009.eqiad.wmnet * 15:44 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1009.eqiad.wmnet * 15:42 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1138 * 15:41 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1138.eqiad.wmnet with OS trixie * 15:39 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1009.eqiad.wmnet * 15:39 robh@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on arclamp2001.codfw.wmnet with reason: ram upgrade * 15:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2174 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95413 and previous config saved to /var/cache/conftool/dbconfig/20260728-153844-cwilliams.json * 15:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2174.codfw.wmnet with reason: Maintenance * 15:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2173: Maintenance * 15:37 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 15:35 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 15:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1009.eqiad.wmnet * 15:34 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1008.eqiad.wmnet * 15:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1008.eqiad.wmnet * 15:31 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 15:31 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 15:29 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1008.eqiad.wmnet * 15:27 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 290 hosts * 15:25 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2013 * 15:25 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2195: Maintenance * 15:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1008.eqiad.wmnet * 15:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1007.eqiad.wmnet * 15:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1007.eqiad.wmnet * 15:22 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2004.codfw.wmnet with reason: host reimage * 15:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2013.codfw.wmnet with OS bookworm * 15:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 15:19 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2004.codfw.wmnet with reason: host reimage * 15:19 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 15:17 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1007.eqiad.wmnet * 15:12 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1007.eqiad.wmnet * 15:12 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1006.eqiad.wmnet * 15:12 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1006.eqiad.wmnet * 15:11 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1138.eqiad.wmnet * 15:11 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 15:11 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1138.eqiad.wmnet * 15:11 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1138.eqiad.wmnet * 15:10 brennen@deploy1003: Finished deploy [phabricator/deployment@f8b349f]: deploy phab1004 for [[phab:T433382|T433382]] (duration: 00m 43s) * 15:10 brennen@deploy1003: Started deploy [phabricator/deployment@f8b349f]: deploy phab1004 for [[phab:T433382|T433382]] * 15:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts * 15:09 brennen@deploy1003: Finished deploy [phabricator/deployment@f8b349f]: deploy phab2003 for [[phab:T433382|T433382]] (duration: 00m 55s) * 15:08 brennen@deploy1003: Started deploy [phabricator/deployment@f8b349f]: deploy phab2003 for [[phab:T433382|T433382]] * 15:07 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2012.codfw.wmnet, repooling source-only afterwards * 15:07 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts * 15:06 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1137.eqiad.wmnet * 15:06 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1137.eqiad.wmnet * 15:06 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1137.eqiad.wmnet * 15:05 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1006.eqiad.wmnet * 15:05 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1022\.eqiad\.wmnet,dc=eqiad,cluster=wdqs\-main,service=wdqs\-main * 15:01 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab2003.codfw.wmnet with reason: deployment * 15:01 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1005.eqiad.wmnet with reason: deployment * 15:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1006.eqiad.wmnet * 15:00 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1006.eqiad.wmnet with reason: deployment * 15:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1005.eqiad.wmnet * 15:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1005.eqiad.wmnet * 14:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db2248: Maintenance * 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1004.eqiad.wmnet with reason: deployment * 14:59 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2004.codfw.wmnet with OS bookworm * 14:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1005.eqiad.wmnet * 14:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2248 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95403 and previous config saved to /var/cache/conftool/dbconfig/20260728-145532-cwilliams.json * 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2245-2247].codfw.wmnet with reason: Maintenance * 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2248.codfw.wmnet with reason: Maintenance * 14:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2240: Maintenance * 14:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2173: Maintenance * 14:51 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts * 14:50 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1005.eqiad.wmnet * 14:50 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1004.eqiad.wmnet * 14:50 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1004.eqiad.wmnet * 14:49 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts * 14:45 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2004.codfw.wmnet * 14:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2173 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95399 and previous config saved to /var/cache/conftool/dbconfig/20260728-144453-cwilliams.json * 14:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2173.codfw.wmnet with reason: Maintenance * 14:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2170: Maintenance * 14:44 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1004.eqiad.wmnet * 14:39 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2004.codfw.wmnet * 14:38 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1004.eqiad.wmnet * 14:38 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet * 14:38 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet * 14:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db2195: Maintenance * 14:36 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1022.eqiad.wmnet, repooling source-only afterwards * 14:36 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 14:33 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet * 14:33 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2222: Maintenance * 14:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2195 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95394 and previous config saved to /var/cache/conftool/dbconfig/20260728-143218-cwilliams.json * 14:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2195.codfw.wmnet with reason: Maintenance * 14:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2181: Maintenance * 14:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 14:30 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 14:25 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 14:25 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 14:25 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 14:23 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet * 14:23 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1002.eqiad.wmnet * 14:23 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1002.eqiad.wmnet * 14:23 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 19s) * 14:23 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:18 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1002.eqiad.wmnet * 14:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2012.codfw.wmnet with OS bookworm * 14:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1002.eqiad.wmnet * 14:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1001.eqiad.wmnet * 14:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1001.eqiad.wmnet * 14:11 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 14:11 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 14:08 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1001.eqiad.wmnet * 14:07 elukey@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'. * 14:07 elukey@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'. * 14:06 elukey@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'. * 14:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db2240: Maintenance * 14:06 XioNoX: un-drain cr2-esams - [[phab:T431751|T431751]] * 14:05 elukey@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'. * 14:02 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1001.eqiad.wmnet * 14:02 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-eqiad * 14:01 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 14:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2240 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95384 and previous config saved to /var/cache/conftool/dbconfig/20260728-140011-cwilliams.json * 14:00 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2240.codfw.wmnet with reason: Maintenance * 13:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2237: Maintenance * 13:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db2170: Maintenance * 13:56 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 13:55 XioNoX: reboot cr2-esams - [[phab:T431751|T431751]] * 13:52 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 13:51 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr2-esams,cr2-esams IPv6,cr2-esams.mgmt with reason: router upgrade * 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 13:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2170 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95381 and previous config saved to /var/cache/conftool/dbconfig/20260728-135043-cwilliams.json * 13:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2170.codfw.wmnet with reason: Maintenance * 13:50 XioNoX: drain cr2-esams - [[phab:T431751|T431751]] * 13:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2153: Maintenance * 13:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2012.codfw.wmnet with reason: host reimage * 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2222: Maintenance * 13:45 sukhe: restart pybal on A:lvs-codfw * 13:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2012.codfw.wmnet with reason: host reimage * 13:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db2181: Maintenance * 13:44 btullis@dns1004: END - running authdns-update * 13:42 sukhe: restart pybal on lvs2014 * 13:42 btullis@dns1004: START - running authdns-update * 13:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2222 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95376 and previous config saved to /var/cache/conftool/dbconfig/20260728-133948-cwilliams.json * 13:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2222.codfw.wmnet with reason: Maintenance * 13:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2221: Maintenance * 13:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2181 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95374 and previous config saved to /var/cache/conftool/dbconfig/20260728-133857-cwilliams.json * 13:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2181.codfw.wmnet with reason: Maintenance * 13:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2167: Maintenance * 13:30 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T428495|T428495]] * 13:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 13:29 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 13:29 ayounsi@cumin1003: END (FAIL) - Cookbook sre.dns.admin (exit_code=99) DNS admin: depool esams [reason: router upgrade, [[phab:T431749|T431749]]] * 13:28 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: router upgrade, [[phab:T431749|T431749]]] * 13:27 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1022.eqiad.wmnet, repooling source-only afterwards * 13:27 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo - [[phab:T428495|T428495]] * 13:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2012 * 13:27 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2012 * 13:21 lucaswerkmeister-wmde@deploy1003: mwscript-k8s job started: cleanupTitles bolwiki # [[phab:T429951|T429951]] * 13:21 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] (duration: 07m 19s) * 13:20 swfrench-wmf: authdns-update to direct codfw, eqsin, ulsfo etcd clients to eqiad - [[phab:T428495|T428495]] * 13:18 swfrench@dns1004: END - running authdns-update * 13:17 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, anzx: Continuing with deployment * 13:16 swfrench@dns1004: START - running authdns-update * 13:16 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2012 * 13:16 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2012.codfw.wmnet 57.48.192.10.in-addr.arpa 7.5.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:16 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2012.codfw.wmnet 57.48.192.10.in-addr.arpa 7.5.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:16 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:16 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, anzx: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:14 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 13:14 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] * 13:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db2237: Maintenance * 13:13 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2237: Maintenance * 13:13 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:12 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:12 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback IPV6 for asw1-604 - pt1979@cumin2003" * 13:12 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:12 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback IPV6 for asw1-604 - pt1979@cumin2003" * 13:11 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 20s) * 13:11 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 13:10 esanders@deploy1003: Finished scap sync-world: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] (duration: 08m 11s) * 13:08 pt1979@cumin2003: START - Cookbook sre.dns.netbox * 13:07 root@cumin1003: START - Cookbook sre.mysql.pool pool db2237: Maintenance * 13:06 esanders@deploy1003: esanders: Continuing with deployment * 13:04 esanders@deploy1003: esanders: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2153: Maintenance * 13:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2153: Maintenance * 13:02 esanders@deploy1003: Started scap sync-world: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] * 13:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2237 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95362 and previous config saved to /var/cache/conftool/dbconfig/20260728-130107-cwilliams.json * 13:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2237.codfw.wmnet with reason: Maintenance * 13:00 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2236: Maintenance * 12:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2153: Maintenance * 12:52 root@cumin1003: START - Cookbook sre.mysql.pool pool db2221: Maintenance * 12:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2153 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95358 and previous config saved to /var/cache/conftool/dbconfig/20260728-125214-cwilliams.json * 12:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2153.codfw.wmnet with reason: Maintenance * 12:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2167: Maintenance * 12:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-codfw * 12:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2011.codfw.wmnet * 12:51 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2011.codfw.wmnet * 12:49 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:49 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback for asw1-603 - pt1979@cumin2003" * 12:48 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback for asw1-603 - pt1979@cumin2003" * 12:46 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2011.codfw.wmnet * 12:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2221 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95357 and previous config saved to /var/cache/conftool/dbconfig/20260728-124601-cwilliams.json * 12:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2221.codfw.wmnet with reason: Maintenance * 12:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2218: Maintenance * 12:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2167 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95354 and previous config saved to /var/cache/conftool/dbconfig/20260728-124457-cwilliams.json * 12:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2167.codfw.wmnet with reason: Maintenance * 12:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2166: Maintenance * 12:42 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 12:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2011.codfw.wmnet * 12:41 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2010.codfw.wmnet * 12:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2010.codfw.wmnet * 12:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2010.codfw.wmnet * 12:34 pt1979@cumin2003: START - Cookbook sre.dns.netbox * 12:32 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 12:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2010.codfw.wmnet * 12:31 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2009.codfw.wmnet * 12:31 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2009.codfw.wmnet * 12:27 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2009.codfw.wmnet * 12:22 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2009.codfw.wmnet * 12:22 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2008.codfw.wmnet * 12:21 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2008.codfw.wmnet * 12:16 pt1979@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-604-eqsin * 12:16 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2008.codfw.wmnet * 12:16 pt1979@cumin1003: START - Cookbook sre.network.tls for network device asw1-604-eqsin * 12:14 root@cumin1003: START - Cookbook sre.mysql.pool pool db2236: Maintenance * 12:14 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2236: Maintenance * 12:12 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1137.eqiad.wmnet with OS trixie * 12:11 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2008.codfw.wmnet * 12:11 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2007.codfw.wmnet * 12:11 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2007.codfw.wmnet * 12:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db2236: Maintenance * 12:06 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2007.codfw.wmnet * 12:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2236 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95348 and previous config saved to /var/cache/conftool/dbconfig/20260728-120253-cwilliams.json * 12:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2236.codfw.wmnet with reason: Maintenance * 12:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2007.codfw.wmnet * 12:01 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 12:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 11:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2218: Maintenance * 11:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db2166: Maintenance * 11:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2219: Maintenance * 11:56 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2006.codfw.wmnet * 11:52 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1137.eqiad.wmnet with reason: host reimage * 11:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2218 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95344 and previous config saved to /var/cache/conftool/dbconfig/20260728-115155-cwilliams.json * 11:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2218.codfw.wmnet with reason: Maintenance * 11:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2208: Maintenance * 11:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2166 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95342 and previous config saved to /var/cache/conftool/dbconfig/20260728-115119-cwilliams.json * 11:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2166.codfw.wmnet with reason: Maintenance * 11:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2164: Maintenance * 11:47 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1137.eqiad.wmnet with reason: host reimage * 11:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2006.codfw.wmnet * 11:45 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2005.codfw.wmnet * 11:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2005.codfw.wmnet * 11:40 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2005.codfw.wmnet * 11:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2005.codfw.wmnet * 11:35 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2004.codfw.wmnet * 11:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2004.codfw.wmnet * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1137 * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1137 * 11:30 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1137 * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1137.eqiad.wmnet 192.32.64.10.in-addr.arpa 2.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:30 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1137.eqiad.wmnet 192.32.64.10.in-addr.arpa 2.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1137 - jiji@cumin1003" * 11:25 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2004.codfw.wmnet * 11:19 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2004.codfw.wmnet * 11:19 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2003.codfw.wmnet * 11:19 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2003.codfw.wmnet * 11:14 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2003.codfw.wmnet * 11:11 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2219: Maintenance * 11:10 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2219: Maintenance * 11:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2219: Maintenance * 11:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2003.codfw.wmnet * 11:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2164: Maintenance * 11:03 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2002.codfw.wmnet * 11:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2002.codfw.wmnet * 11:03 root@cumin1003: START - Cookbook sre.mysql.pool pool db2208: Maintenance * 10:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2164 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95332 and previous config saved to /var/cache/conftool/dbconfig/20260728-105749-cwilliams.json * 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2164.codfw.wmnet with reason: Maintenance * 10:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2208 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95331 and previous config saved to /var/cache/conftool/dbconfig/20260728-105711-cwilliams.json * 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2208.codfw.wmnet with reason: Maintenance * 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2219 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95330 and previous config saved to /var/cache/conftool/dbconfig/20260728-105652-cwilliams.json * 10:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2219.codfw.wmnet with reason: Maintenance * 10:53 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1137 - jiji@cumin1003" * 10:52 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2002.codfw.wmnet * 10:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2002.codfw.wmnet * 10:47 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2001.codfw.wmnet * 10:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2001.codfw.wmnet * 10:39 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2001.codfw.wmnet * 10:35 jiji@cumin1003: START - Cookbook sre.dns.netbox * 10:34 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] (duration: 09m 31s) * 10:34 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1137 * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2001.codfw.wmnet * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-codfw * 10:34 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1137.eqiad.wmnet with OS trixie * 10:28 jforrester@deploy1003: jforrester: Continuing with deployment * 10:27 jforrester@deploy1003: jforrester: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:25 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] * 10:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 10:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 10:21 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 10:20 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1137.eqiad.wmnet * 10:20 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 10:20 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1137.eqiad.wmnet * 10:20 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1137.eqiad.wmnet * 10:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2163: Maintenance * 09:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-staging-worker * 09:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2003.codfw.wmnet * 09:37 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2003.codfw.wmnet * 09:32 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 09:31 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2003.codfw.wmnet * 09:30 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 09:30 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2163: Maintenance * 09:30 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 09:30 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 09:30 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:22 klausman@cumin1003: END (ERROR) - Cookbook sre.ganeti.reboot-vm (exit_code=97) for VM ml-serve-ctrl2001.codfw.wmnet * 09:22 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2001.codfw.wmnet * 09:22 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d8-eqiad * 09:22 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d8-eqiad * 09:21 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2003.codfw.wmnet * 09:20 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2002.codfw.wmnet * 09:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2002.codfw.wmnet * 09:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f2-codfw * 09:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f2-codfw * 09:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e4-codfw * 09:18 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2163: Maintenance * 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e4-codfw * 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-codfw * 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-codfw * 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e5-codfw * 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e5-codfw * 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f4-codfw * 09:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2210: Maintenance * 09:16 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f4-codfw * 09:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2182: Maintenance * 09:14 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2002.codfw.wmnet * 09:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db2163: Maintenance * 09:11 XioNoX: rebooting cr2-drmrs - [[phab:T431749|T431749]] * 09:10 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr2-drmrs,cr2-drmrs IPv6,cr2-drmrs.mgmt with reason: router upgrade * 09:06 XioNoX: draining cr2-drmrs - [[phab:T431749|T431749]] * 09:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2163 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95320 and previous config saved to /var/cache/conftool/dbconfig/20260728-090638-cwilliams.json * 09:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2163.codfw.wmnet with reason: Maintenance * 09:06 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2161: Maintenance * 09:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2002.codfw.wmnet * 09:04 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2001.codfw.wmnet * 09:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2001.codfw.wmnet * 08:57 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2001.codfw.wmnet * 08:48 XioNoX: un-drain cr1-drmrs - [[phab:T431749|T431749]] * 08:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2001.codfw.wmnet * 08:47 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-staging-worker * 08:42 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:35 XioNoX: rebooting cr1-drmrs - [[phab:T431749|T431749]] * 08:33 XioNoX: draining cr1-drmrs - [[phab:T431749|T431749]] * 08:31 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2210: Maintenance * 08:29 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2182: Maintenance * 08:21 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2182: Maintenance * 08:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db2161: Maintenance * 08:16 root@cumin1003: START - Cookbook sre.mysql.pool pool db2182: Maintenance * 08:12 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2210: Maintenance * 08:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2161 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95309 and previous config saved to /var/cache/conftool/dbconfig/20260728-081044-cwilliams.json * 08:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2161.codfw.wmnet with reason: Maintenance * 08:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2154: Maintenance * 08:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2182 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95307 and previous config saved to /var/cache/conftool/dbconfig/20260728-080947-cwilliams.json * 08:09 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2182.codfw.wmnet with reason: Maintenance * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2168: Maintenance * 08:06 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr1-drmrs,cr1-drmrs IPv6,cr1-drmrs.mgmt with reason: router upgrade * 08:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db2210: Maintenance * 08:05 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 08:05 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 08:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2210 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95305 and previous config saved to /var/cache/conftool/dbconfig/20260728-080008-cwilliams.json * 08:00 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2210.codfw.wmnet with reason: Maintenance * 07:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2206: Maintenance * 07:50 gkyziridis@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 07:50 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 07:22 root@cumin1003: START - Cookbook sre.mysql.pool pool db2154: Maintenance * 07:22 root@cumin1003: START - Cookbook sre.mysql.pool pool db2168: Maintenance * 07:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2154 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95295 and previous config saved to /var/cache/conftool/dbconfig/20260728-071640-cwilliams.json * 07:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2154.codfw.wmnet with reason: Maintenance * 07:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95294 and previous config saved to /var/cache/conftool/dbconfig/20260728-071604-cwilliams.json * 07:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2168.codfw.wmnet with reason: Maintenance * 07:08 root@cumin1003: START - Cookbook sre.mysql.pool pool db2206: Maintenance * 07:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2206 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95292 and previous config saved to /var/cache/conftool/dbconfig/20260728-070219-cwilliams.json * 07:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2206.codfw.wmnet with reason: Maintenance * 06:44 marostegui: Failover m5 from db1164 to db1228 - [[phab:T432967|T432967]] * 06:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2235].codfw.wmnet,db[1164,1217,1228].eqiad.wmnet with reason: m5 master switch [[phab:T432967|T432967]] * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.10 (duration: 02m 34s) * 03:39 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] (duration: 36m 06s) * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 02:57 dzahn@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1004.eqiad.wmnet with OS trixie * 02:57 dzahn@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - dzahn@cumin1003" * 02:55 dzahn@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - dzahn@cumin1003" * 02:37 dzahn@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1004.eqiad.wmnet with reason: host reimage * 02:31 dzahn@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1004.eqiad.wmnet with reason: host reimage * 02:16 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie * 02:15 dzahn@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host zuul1004.eqiad.wmnet with OS trixie * 01:43 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie * 01:43 dzahn@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1004.eqiad.wmnet with OS trixie * 01:25 pt1979@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-603-eqsin * 01:24 pt1979@cumin1003: START - Cookbook sre.network.tls for network device asw1-603-eqsin * 01:12 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 01:12 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt for new switches in eqsin - pt1979@cumin2003" * 01:12 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt for new switches in eqsin - pt1979@cumin2003" * 01:08 pt1979@cumin2003: START - Cookbook sre.dns.netbox * 00:48 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 00:47 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 00:47 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 00:47 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 00:26 mutante: attempting reimage with trixie on zuul1004 re-purposed physical hardware - dcops reported install issue - host was in busybox shell ([[phab:T427353|T427353]]) * 00:24 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie == 2026-07-27 == * 23:50 Amir1: mass deleting vp8 transcodes * 23:28 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:27 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1004.eqiad.wmnet with OS bullseye * 23:26 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 23:25 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:25 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 22:39 maryum: Deploy security fix for [[phab:T432877|T432877]] * 22:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1022.eqiad.wmnet with OS bookworm * 22:37 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS bullseye * 22:32 sbassett: Deployed security fix for [[phab:T432789|T432789]] * 22:22 sbassett: Deployed security patch for [[phab:T431819|T431819]] * 22:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1022.eqiad.wmnet with reason: host reimage * 22:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1022.eqiad.wmnet with reason: host reimage * 22:01 RScout-WMF: Deployed security fix for [[phab:T431819|T431819]] * 22:00 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2012 * 21:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2012.codfw.wmnet with OS bookworm * 21:55 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2011\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 21:45 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1022 * 21:45 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1022 * 21:44 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1022 * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1022.eqiad.wmnet 239.48.64.10.in-addr.arpa 9.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:44 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1022.eqiad.wmnet 239.48.64.10.in-addr.arpa 9.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:41 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:41 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 21:34 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS bookworm * 21:31 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:22 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:21 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:19 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:17 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1004 * 21:16 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1004 * 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1004] - vriley@cumin1003" * 21:15 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1004] - vriley@cumin1003" * 21:11 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:10 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2011.codfw.wmnet, repooling source-only afterwards * 21:05 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:01 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1022 * 20:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1022.eqiad.wmnet with OS bookworm * 20:53 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1021.eqiad.wmnet, repooling source-only afterwards * 20:51 mutante: zuul1001 - re-enabled puppet - revert "cherry-picked" gerrit:1314120 - [[phab:T431003|T431003]] * 20:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Maintenance * 20:15 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] (duration: 08m 03s) * 20:11 sbisson@deploy1003: sbisson: Continuing with deployment * 20:09 sbisson@deploy1003: sbisson: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] * 19:47 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 46s) * 19:47 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Maintenance * 19:27 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] (duration: 12m 26s) * 19:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2228 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95285 and previous config saved to /var/cache/conftool/dbconfig/20260727-192711-cwilliams.json * 19:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2228.codfw.wmnet with reason: Maintenance * 19:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2223: Maintenance * 19:23 krinkle@deploy1003: krinkle: Continuing with deployment * 19:16 krinkle@deploy1003: krinkle: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:15 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] * 19:12 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2238: Maintenance * 18:58 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 18:57 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 18:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2227: Maintenance * 18:57 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-experimental: apply * 18:55 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-experimental: apply * 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1021.eqiad.wmnet with OS bookworm * 18:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2011.codfw.wmnet with OS bookworm * 18:40 root@cumin1003: START - Cookbook sre.mysql.pool pool db2223: Maintenance * 18:39 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] (duration: 07m 05s) * 18:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2223 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95275 and previous config saved to /var/cache/conftool/dbconfig/20260727-183500-cwilliams.json * 18:34 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2223.codfw.wmnet with reason: Maintenance * 18:34 musikanimal@deploy1003: musikanimal: Continuing with deployment * 18:34 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2213: Maintenance * 18:33 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:32 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] * 18:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db2238: Maintenance * 18:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2011.codfw.wmnet with reason: host reimage * 18:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2238 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95271 and previous config saved to /var/cache/conftool/dbconfig/20260727-181944-cwilliams.json * 18:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2238.codfw.wmnet with reason: Maintenance * 18:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2226: Maintenance * 18:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1021.eqiad.wmnet with reason: host reimage * 18:14 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2011.codfw.wmnet with reason: host reimage * 18:12 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1021.eqiad.wmnet with reason: host reimage * 18:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db2227: Maintenance * 18:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2227 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95265 and previous config saved to /var/cache/conftool/dbconfig/20260727-180256-cwilliams.json * 18:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2227.codfw.wmnet with reason: Maintenance * 18:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2194: Maintenance * 17:57 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2011 * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2011 * 17:56 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2011 * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2011.codfw.wmnet 37.32.192.10.in-addr.arpa 7.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:56 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2011.codfw.wmnet 37.32.192.10.in-addr.arpa 7.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2011 - bking@cumin2003" * 17:56 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2011 - bking@cumin2003" * 17:52 bking@cumin2003: START - Cookbook sre.dns.netbox * 17:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2011 * 17:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1021 * 17:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1021 * 17:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2011.codfw.wmnet with OS bookworm * 17:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1021.eqiad.wmnet with OS bookworm * 17:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Maintenance * 17:38 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2010\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 17:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2213 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95260 and previous config saved to /var/cache/conftool/dbconfig/20260727-173740-cwilliams.json * 17:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2213.codfw.wmnet with reason: Maintenance * 17:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2211: Maintenance * 17:36 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1020\.eqiad\.wmnet,dc=eqiad,cluster=wdqs\-main,service=wdqs\-main * 17:32 root@cumin1003: START - Cookbook sre.mysql.pool pool db2226: Maintenance * 17:31 taavi@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] (duration: 06m 33s) * 17:27 taavi@deploy1003: taavi: Continuing with deployment * 17:27 taavi@deploy1003: taavi: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:26 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2226 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95256 and previous config saved to /var/cache/conftool/dbconfig/20260727-172636-cwilliams.json * 17:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2226.codfw.wmnet with reason: Maintenance * 17:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2225: Maintenance * 17:25 taavi@deploy1003: Started scap sync-world: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] * 17:13 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 17:11 root@cumin1003: START - Cookbook sre.mysql.pool pool db2194: Maintenance * 17:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2194 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95248 and previous config saved to /var/cache/conftool/dbconfig/20260727-170453-cwilliams.json * 17:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2194.codfw.wmnet with reason: Maintenance * 17:04 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2190: Maintenance * 16:52 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 16:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2211: Maintenance * 16:40 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95242 and previous config saved to /var/cache/conftool/dbconfig/20260727-164015-cwilliams.json * 16:40 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2211.codfw.wmnet with reason: Maintenance * 16:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2178: Maintenance * 16:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db2225: Maintenance * 16:39 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2172: Maintenance * 16:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 16:38 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 16:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2225 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95238 and previous config saved to /var/cache/conftool/dbconfig/20260727-163307-cwilliams.json * 16:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2225.codfw.wmnet with reason: Maintenance * 16:32 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2189: Maintenance * 16:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db2190: Maintenance * 16:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2190 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95230 and previous config saved to /var/cache/conftool/dbconfig/20260727-160602-cwilliams.json * 16:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2190.codfw.wmnet with reason: Maintenance * 15:53 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2177: Maintenance * 15:53 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2172: Maintenance * 15:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2178: Maintenance * 15:51 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2172: Maintenance * 15:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2178 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95224 and previous config saved to /var/cache/conftool/dbconfig/20260727-154559-cwilliams.json * 15:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2172: Maintenance * 15:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2178.codfw.wmnet with reason: Maintenance * 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2171: Maintenance * 15:44 root@cumin1003: START - Cookbook sre.mysql.pool pool db2189: Maintenance * 15:43 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:41 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2172 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95222 and previous config saved to /var/cache/conftool/dbconfig/20260727-153927-cwilliams.json * 15:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2172.codfw.wmnet with reason: Maintenance * 15:38 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2189 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95220 and previous config saved to /var/cache/conftool/dbconfig/20260727-153833-cwilliams.json * 15:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2189.codfw.wmnet with reason: Maintenance * 15:34 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:32 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 15:32 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 15:31 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:29 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:26 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 15:22 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] (duration: 07m 00s) * 15:21 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2155: Maintenance * 15:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2175: Maintenance * 15:18 zabe@deploy1003: zabe: Continuing with deployment * 15:17 zabe@deploy1003: zabe: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:15 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 15:15 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:15 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] * 15:15 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 06s) * 15:15 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:12 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2010.codfw.wmnet with OS bookworm * 15:08 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2177: Maintenance * 15:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2177: Maintenance * 14:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2177: Maintenance * 14:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2171: Maintenance * 14:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2171 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95209 and previous config saved to /var/cache/conftool/dbconfig/20260727-145236-cwilliams.json * 14:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2171.codfw.wmnet with reason: Maintenance * 14:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2177 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95208 and previous config saved to /var/cache/conftool/dbconfig/20260727-145206-cwilliams.json * 14:52 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2157: Maintenance * 14:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2177.codfw.wmnet with reason: Maintenance * 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1020.eqiad.wmnet with OS bookworm * 14:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2156: Maintenance * 14:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2010.codfw.wmnet with reason: host reimage * 14:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2010.codfw.wmnet with reason: host reimage * 14:41 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[1031,2024]*: Upgrade Cassandra to 5.0.8 (canary) - eevans@cumin1003 * 14:34 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2155: Maintenance * 14:33 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2175: Maintenance * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2010 * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2010 * 14:24 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2010 * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2010.codfw.wmnet 94.16.192.10.in-addr.arpa 4.9.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2010.codfw.wmnet 94.16.192.10.in-addr.arpa 4.9.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2010 - bking@cumin2003" * 14:24 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2010 - bking@cumin2003" * 14:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1020.eqiad.wmnet with reason: host reimage * 14:23 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[1031,2024]*: Upgrade Cassandra to 5.0.8 (canary) - eevans@cumin1003 * 14:20 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 14:20 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 14:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1020.eqiad.wmnet with reason: host reimage * 14:17 sukhe: sudo gnt-instance reboot urldownloader1005.wikimedia.org * 14:16 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:15 jelto@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:08 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2155: Maintenance * 14:05 root@cumin1003: START - Cookbook sre.mysql.pool pool db2157: Maintenance * 14:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2175: Maintenance * 14:03 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 11 hosts * 14:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db2155: Maintenance * 14:01 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 11 hosts * 14:01 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1136.eqiad.wmnet * 14:01 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1136.eqiad.wmnet * 14:01 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1136.eqiad.wmnet * 14:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db2156: Maintenance * 13:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2157 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95194 and previous config saved to /var/cache/conftool/dbconfig/20260727-135943-cwilliams.json * 13:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2157.codfw.wmnet with reason: Maintenance * 13:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db2175: Maintenance * 13:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 13:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 13:57 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2010 * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2155 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95193 and previous config saved to /var/cache/conftool/dbconfig/20260727-135613-cwilliams.json * 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2155.codfw.wmnet with reason: Maintenance * 13:55 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1020 * 13:55 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1020 * 13:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2156 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95192 and previous config saved to /var/cache/conftool/dbconfig/20260727-135413-cwilliams.json * 13:54 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2156.codfw.wmnet with reason: Maintenance * 13:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2175 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95191 and previous config saved to /var/cache/conftool/dbconfig/20260727-135300-cwilliams.json * 13:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2010.codfw.wmnet with OS bookworm * 13:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2175.codfw.wmnet with reason: Maintenance * 13:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1020.eqiad.wmnet with OS bookworm * 13:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 34 hosts * 13:46 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 34 hosts * 13:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1201: Maintenance * 13:27 Lucas_WMDE: UTC afternoon backport+config window doen * 13:18 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] (duration: 11m 57s) * 13:14 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, sihe: Continuing with deployment * 13:08 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, sihe: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:07 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool ulsfo [reason: router upgrade finished, [[phab:T431752|T431752]]] * 13:07 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool ulsfo [reason: router upgrade finished, [[phab:T431752|T431752]]] * 13:06 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] * 13:03 XioNoX: repool cr4-ulsfo - [[phab:T431752|T431752]] * 12:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db1201: Maintenance * 12:48 gkyziridis@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1201 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95186 and previous config saved to /var/cache/conftool/dbconfig/20260727-124404-cwilliams.json * 12:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1201.eqiad.wmnet with reason: Maintenance * 12:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1187: Maintenance * 12:30 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader.eqiad.wikimedia.org on all recursors * 12:30 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader.eqiad.wikimedia.org on all recursors * 12:30 sukhe@dns1004: END - running authdns-update * 12:30 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] (duration: 09m 32s) * 12:28 sukhe@dns1004: START - running authdns-update * 12:25 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 12:22 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:20 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] * 12:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts * 12:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts * 12:13 XioNoX: rebooting cr4-ulsfo for upgrade - [[phab:T431752|T431752]] * 12:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: es1038 repool * 12:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 38 hosts * 12:08 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 38 hosts * 11:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1187: Maintenance * 11:53 urbanecm@deploy1003: mwscript-k8s job started: foreachwikiindblist growthexperiments GrowthExperiments:cleanMentorList # [[phab:T431804|T431804]] * 11:50 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr4-ulsfo,cr4-ulsfo IPv6,cr4-ulsfo.mgmt with reason: router upgrade * 11:50 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] (duration: 11m 07s) * 11:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1187 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95178 and previous config saved to /var/cache/conftool/dbconfig/20260727-114844-cwilliams.json * 11:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1187.eqiad.wmnet with reason: Maintenance * 11:43 urbanecm@deploy1003: urbanecm: Continuing with deployment * 11:42 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:39 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] * 11:37 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 11:36 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 11:36 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 11:35 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 11:29 XioNoX: start draining cr4-ulsfo - [[phab:T431752|T431752]] * 11:29 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 11:29 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 11:28 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1035: testing * 11:28 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1035: testing * 11:27 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1035: testing * 11:27 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1035: testing * 11:26 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool ulsfo [reason: router upgrade, [[phab:T431752|T431752]]] * 11:26 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1038: es1038 repool * 11:26 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool ulsfo [reason: router upgrade, [[phab:T431752|T431752]]] * 11:26 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1038: testing * 11:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1264: Maintenance * 11:24 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1038: testing * 11:23 marostegui@cumin1003: dbctl commit (dc=all): 'Repool es1050 as master', diff saved to https://phabricator.wikimedia.org/P95170 and previous config saved to /var/cache/conftool/dbconfig/20260727-112326-marostegui.json * 11:23 marostegui@cumin1003: dbctl commit (dc=all): 'Repool es1050', diff saved to https://phabricator.wikimedia.org/P95169 and previous config saved to /var/cache/conftool/dbconfig/20260727-112302-marostegui.json * 11:22 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1050: testing * 11:22 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1050: testing * 11:20 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 11:18 blake@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 11:18 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 11:12 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 11:11 blake@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 11:09 blake@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 11:09 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 11:09 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 11:08 blake@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 11:05 blake@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 11:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 11:02 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 10:50 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 10:43 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:39 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply * 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1264: Maintenance * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply * 10:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply * 10:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 10:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 10:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 10:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1264 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95164 and previous config saved to /var/cache/conftool/dbconfig/20260727-103204-cwilliams.json * 10:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1264.eqiad.wmnet with reason: Maintenance * 10:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 10:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 10:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 10:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 10:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 10:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 10:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 10:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1237: Maintenance * 10:04 elukey: restart burrow main-eqiad on kafkamon2003 to clear some errors on kafka-main1008 * 09:58 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1136.eqiad.wmnet with OS trixie * 09:39 elukey: restart burrow-main-eqiad.service on kafkamon1003 to see if a recurrent kafka error on kafka-main1008 goes away * 09:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1237: Maintenance * 09:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1136.eqiad.wmnet with reason: host reimage * 09:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1237 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95159 and previous config saved to /var/cache/conftool/dbconfig/20260727-093328-cwilliams.json * 09:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1237.eqiad.wmnet with reason: Maintenance * 09:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1136.eqiad.wmnet with reason: host reimage * 09:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1203: Maintenance * 09:17 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1136 * 09:17 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1136 * 09:04 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1136 * 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1136.eqiad.wmnet 191.32.64.10.in-addr.arpa 1.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:04 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1136.eqiad.wmnet 191.32.64.10.in-addr.arpa 1.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1136 - jiji@cumin1003" * 09:04 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1136 - jiji@cumin1003" * 08:52 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 08:52 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 08:52 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 08:51 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 08:50 jiji@cumin1003: START - Cookbook sre.dns.netbox * 08:47 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1136 * 08:46 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1136.eqiad.wmnet with OS trixie * 08:44 marostegui: Rename tables on s3 [[phab:T425066|T425066]] * 08:43 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1136.eqiad.wmnet * 08:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db1203: Maintenance * 08:43 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1136.eqiad.wmnet * 08:43 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1136.eqiad.wmnet * 08:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1203 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95154 and previous config saved to /var/cache/conftool/dbconfig/20260727-083703-cwilliams.json * 08:36 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1203.eqiad.wmnet with reason: Maintenance * 08:16 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1179: Maintenance * 07:44 phuedx: UTC morning backport window done * 07:37 phuedx@deploy1003: Finished scap sync-world: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] (duration: 32m 33s) * 07:28 root@cumin1003: START - Cookbook sre.mysql.pool pool db1179: Maintenance * 07:26 marostegui: Rename tables on s3 [[phab:T426341|T426341]] * 07:25 phuedx@deploy1003: phuedx: Continuing with deployment * 07:22 marostegui: Drop tables in akwiki nawiki pihwiki - growthexperiments_* [[phab:T428885|T428885]] * 07:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95149 and previous config saved to /var/cache/conftool/dbconfig/20260727-072234-cwilliams.json * 07:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1179.eqiad.wmnet with reason: Maintenance * 07:20 phuedx@deploy1003: phuedx: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:16 ryankemper: [[phab:T430880|T430880]] [WDQS] Reimaged `wdqs1018` and `wdqs1019` to Bookworm, restored data using test-cookbook change {{Gerrit|1317128}}, and repooled both; 25/36 hosts complete * 07:04 phuedx@deploy1003: Started scap sync-world: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] * 06:57 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1019.eqiad.wmnet * 06:56 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1018.eqiad.wmnet * 06:40 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1020.eqiad.wmnet with reason: Cloning * 06:35 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db1228.eqiad.wmnet with reason: Rebooting * 06:29 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:29 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:25 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:25 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:25 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1019.eqiad.wmnet, repooling source-only afterwards * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1018.eqiad.wmnet, repooling source-only afterwards * 04:51 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1019.eqiad.wmnet, repooling source-only afterwards * 04:51 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1018.eqiad.wmnet, repooling source-only afterwards * 04:48 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s) * 04:48 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 04:48 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s) * 04:48 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 36s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-26 == * 14:59 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:59 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:59 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:59 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1019.eqiad.wmnet with OS bookworm * 01:05 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1018.eqiad.wmnet with OS bookworm * 00:43 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1019.eqiad.wmnet with reason: host reimage * 00:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1018.eqiad.wmnet with reason: host reimage * 00:34 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1019.eqiad.wmnet with reason: host reimage * 00:33 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1018.eqiad.wmnet with reason: host reimage * 00:16 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 00:16 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 00:15 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 00:15 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1019 * 00:11 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1019 * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1018 * 00:11 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1018 * 00:08 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1019.eqiad.wmnet with OS bookworm * 00:08 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1018.eqiad.wmnet with OS bookworm == 2026-07-25 == * 22:06 ryankemper: [[phab:T430880|T430880]] [WDQS] Repooled `wdqs1017` and `wdqs2024` after reimaging to bookworm, scap deploying, and data xfering * 22:04 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2024.codfw.wmnet * 22:03 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1017.eqiad.wmnet * 21:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1017.eqiad.wmnet, repooling source-only afterwards * 21:06 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2024.codfw.wmnet, repooling source-only afterwards * 20:52 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:52 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:52 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:52 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 20:18 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1017.eqiad.wmnet, repooling source-only afterwards * 20:18 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2024.codfw.wmnet, repooling source-only afterwards * 20:15 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:15 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:15 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:15 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 19:57 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s) * 19:57 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 19:57 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 07s) * 19:57 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 19:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2024.codfw.wmnet with OS bookworm * 19:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1017.eqiad.wmnet with OS bookworm * 19:02 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2024.codfw.wmnet with reason: host reimage * 18:58 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1017.eqiad.wmnet with reason: host reimage * 18:53 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2024.codfw.wmnet with reason: host reimage * 18:52 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1017.eqiad.wmnet with reason: host reimage * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2024 * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2024 * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1017 * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1017 * 18:27 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2024 * 18:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2024.codfw.wmnet 58.16.192.10.in-addr.arpa 8.5.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:26 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2024.codfw.wmnet 58.16.192.10.in-addr.arpa 8.5.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:24 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1017 * 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1017.eqiad.wmnet 238.48.64.10.in-addr.arpa 8.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:24 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1017.eqiad.wmnet 238.48.64.10.in-addr.arpa 8.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1017 - ryankemper@cumin2003" * 18:24 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1017 - ryankemper@cumin2003" * 18:23 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 18:18 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 18:17 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1017 * 18:17 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2024 * 18:14 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1017.eqiad.wmnet with OS bookworm * 18:14 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2024.codfw.wmnet with OS bookworm * 18:05 ryankemper: [WDQS] [[phab:T430880|T430880]] Reimaged `wdqs1016` and `wdqs2023` to Bookworm with `--move-vlan`, restored main and scholarly data, validated postflights, and repooled both hosts. Confirmed PyBal rebuilt both backends with their new addresses * 17:45 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2023.codfw.wmnet * 17:43 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1016.eqiad.wmnet * 06:35 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1016.eqiad.wmnet, repooling source-only afterwards * 06:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2023.codfw.wmnet, repooling source-only afterwards * 05:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2023.codfw.wmnet, repooling source-only afterwards * 05:19 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1016.eqiad.wmnet, repooling source-only afterwards * 05:07 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 07s) * 05:07 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 05:06 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 06s) * 05:06 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 03:27 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2023.codfw.wmnet with OS bookworm * 02:59 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2023.codfw.wmnet with reason: host reimage * 02:56 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2023.codfw.wmnet with reason: host reimage * 02:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2023 * 02:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2023 * 02:30 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2023 * 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2023.codfw.wmnet 35.0.192.10.in-addr.arpa 5.3.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 02:30 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2023.codfw.wmnet 35.0.192.10.in-addr.arpa 5.3.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2023 - ryankemper@cumin2003" * 02:30 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2023 - ryankemper@cumin2003" * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 26s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:15 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1016.eqiad.wmnet with OS bookworm * 00:49 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1016.eqiad.wmnet with reason: host reimage * 00:43 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1016.eqiad.wmnet with reason: host reimage * 00:31 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 00:27 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1016 * 00:27 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1016 * 00:27 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2023 * 00:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1016.eqiad.wmnet with OS bookworm * 00:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2023.codfw.wmnet with OS bookworm * 00:11 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs1014.eqiad.wmnet and wdqs2008.codfw.wmnet after Bookworm reimage, transfer, and postflight; wdqs2008 is serving, while wdqs1014 will remain outside of service until a pybal restart next monday == 2026-07-24 == * 23:54 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1014.eqiad.wmnet * 23:54 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2008.codfw.wmnet * 23:43 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2010.codfw.wmnet with OS trixie * 23:08 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 23:03 jhathaway@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 22:33 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 22:13 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 22:13 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 22:13 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 22:13 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:00 jhathaway@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 21:53 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 21:53 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie * 21:51 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 21:47 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie * 21:43 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 21:39 jhathaway@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 21:38 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 17:21 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1135.eqiad.wmnet * 17:21 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1135.eqiad.wmnet * 17:21 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1135.eqiad.wmnet * 16:34 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 16:34 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 16:34 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 16:34 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 16:33 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 16:33 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 16:28 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:28 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:28 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:28 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2008.codfw.wmnet, repooling source-only afterwards * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1014.eqiad.wmnet, repooling source-only afterwards * 15:56 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1135.eqiad.wmnet with OS trixie * 15:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 40 hosts * 15:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 40 hosts * 15:37 topranks: upgrade SR-Linux OS on lswtest-d8-eqiad * 15:36 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1135.eqiad.wmnet with reason: host reimage * 15:33 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 6 hosts with reason: upgrade lswtest-d8-eqiad * 15:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1135.eqiad.wmnet with reason: host reimage * 15:30 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc-gp2006.codfw.wmnet with OS bookworm * 15:15 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1135 * 15:15 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1135 * 15:13 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc-gp2006.codfw.wmnet with reason: host reimage * 15:08 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc-gp2006.codfw.wmnet with reason: host reimage * 14:49 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm * 14:48 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host mc-gp2006.codfw.wmnet with OS bookworm * 14:34 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] (duration: 41m 12s) * 14:32 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1135 * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1135.eqiad.wmnet 177.32.64.10.in-addr.arpa 7.7.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:32 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1135.eqiad.wmnet 177.32.64.10.in-addr.arpa 7.7.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1135 - jiji@cumin1003" * 14:32 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1135 - jiji@cumin1003" * 14:29 krinkle@deploy1003: krinkle: Continuing with deployment * 14:29 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm * 14:27 jiji@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host mc-gp2006.codfw.wmnet with OS bookworm * 14:26 jiji@cumin1003: START - Cookbook sre.dns.netbox * 14:15 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1135 * 14:14 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1135.eqiad.wmnet with OS trixie * 14:14 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1135.eqiad.wmnet * 14:13 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1135.eqiad.wmnet * 14:13 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1135.eqiad.wmnet * 14:10 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1072.eqiad.wmnet * 14:10 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1072.eqiad.wmnet * 14:10 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1072.eqiad.wmnet * 14:10 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1072.eqiad.wmnet * 14:09 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1071.eqiad.wmnet * 14:09 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1071.eqiad.wmnet * 14:09 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1071.eqiad.wmnet * 14:09 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1071.eqiad.wmnet * 13:58 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 13:58 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:58 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:57 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:55 krinkle@deploy1003: krinkle: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:53 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] * 13:45 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:45 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push new IPs for mc-gp2006 - cmooney@cumin1003" * 13:45 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push new IPs for mc-gp2006 - cmooney@cumin1003" * 13:44 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) mc-gp2006.codfw.wmnet on all recursors * 13:44 cmooney@cumin1003: START - Cookbook sre.dns.wipe-cache mc-gp2006.codfw.wmnet on all recursors * 13:42 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm * 13:41 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:30 papaul: reboot mr1-eqsin for maintenance * 13:24 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb[1029-1031].eqiad.wmnet * 13:10 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb[1029-1031].eqiad.wmnet * 11:33 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 7 hosts * 11:11 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 7 hosts * 10:56 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 7 hosts * 10:47 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 7 hosts * 10:44 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:44 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:41 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:41 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:35 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 8 hosts * 10:34 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:33 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:32 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:32 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:31 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:31 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:30 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 8 hosts * 10:24 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie * 10:19 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:18 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 16 hosts * 10:17 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2001.codfw.wmnet * 10:13 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2001.codfw.wmnet * 10:12 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2001.codfw.wmnet * 10:02 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2001.codfw.wmnet * 10:02 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2002.codfw.wmnet * 09:57 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2002.codfw.wmnet * 09:56 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2002.codfw.wmnet * 09:51 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2002.codfw.wmnet * 09:51 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1002.eqiad.wmnet * 09:47 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1002.eqiad.wmnet * 09:47 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1001.eqiad.wmnet * 09:44 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1001.eqiad.wmnet * 09:34 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2003.codfw.wmnet * 09:32 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2003.codfw.wmnet * 09:32 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2002.codfw.wmnet * 09:30 brouberol@dns1004: END - running authdns-update * 09:29 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2002.codfw.wmnet * 09:29 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2001.codfw.wmnet * 09:27 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 9 hosts * 09:27 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2001.codfw.wmnet * 09:27 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2001.codfw.wmnet * 09:26 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 9 hosts * 09:26 brouberol@dns1004: START - running authdns-update * 09:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 57 hosts * 09:24 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2001.codfw.wmnet * 09:24 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2002.codfw.wmnet * 09:22 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2002.codfw.wmnet * 09:21 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 57 hosts * 09:20 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2003.codfw.wmnet * 09:19 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 16 hosts * 09:16 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2003.codfw.wmnet * 09:16 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1003.eqiad.wmnet * 09:15 urbanecm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 09:15 urbanecm@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 09:13 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1003.eqiad.wmnet * 09:13 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1002.eqiad.wmnet * 09:11 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1002.eqiad.wmnet * 09:11 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1001.eqiad.wmnet * 09:07 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1001.eqiad.wmnet * 08:32 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:24 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:16 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 08:16 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 08:07 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:07 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:07 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 08:02 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 08:01 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:59 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:57 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 07:57 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 06:46 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1025.eqiad.wmnet with reason: Cloning * 06:46 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s4 * 06:45 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s6 * 06:44 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1019.eqiad.wmnet,service=s6 * 06:44 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1019.eqiad.wmnet,service=s4 * 03:40 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:40 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:40 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:40 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 03:37 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:37 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:37 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:36 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:49 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on mr1-eqsin,mr1-eqsin IPv6 with reason: connection issue * 02:38 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on cr[2-3]-eqsin.mgmt,ps1-[603-604]-eqsin with reason: connection issue * 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 27s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-23 == * 23:27 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin.oob,mr1-eqsin.oob IPv6 with reason: switch refresh * 22:21 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Setting storage compatibility to NONE - eevans@cumin1003 * 22:01 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Setting storage compatibility to NONE - eevans@cumin1003 * 21:29 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1014.eqiad.wmnet, repooling source-only afterwards * 21:28 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 46s) * 21:28 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 21:19 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Setting storage compatibility to UPGRADING - eevans@cumin1003 * 21:00 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Setting storage compatibility to UPGRADING - eevans@cumin1003 * 20:17 dani@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] (duration: 11m 57s) * 20:13 dani@deploy1003: dani: Continuing with deployment * 20:07 dani@deploy1003: dani: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:05 dani@deploy1003: Started scap sync-world: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] * 19:24 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:24 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating the rest of the ipv6 dns records. - jhancock@cumin2002" * 19:24 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating the rest of the ipv6 dns records. - jhancock@cumin2002" * 19:14 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 19:05 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wdqs1014.eqiad.wmnet with OS bookworm * 19:04 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.noop (exit_code=99) * 19:04 cwilliams@cumin1003: START - Cookbook sre.mysql.noop * 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1014.eqiad.wmnet with reason: host reimage * 18:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1014.eqiad.wmnet with reason: host reimage * 18:30 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2008.codfw.wmnet, repooling source-only afterwards * 18:28 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 19s) * 18:28 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1014 * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1014 * 18:22 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1014 * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1014.eqiad.wmnet 188.32.64.10.in-addr.arpa 8.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:22 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1014.eqiad.wmnet 188.32.64.10.in-addr.arpa 8.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1014 - bking@cumin2003" * 18:21 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1014 - bking@cumin2003" * 18:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2215: Maintenance * 18:18 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 18:15 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:15 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 18:06 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:06 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 18:05 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 18:04 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2052: codfw rack B8 re-pool after maintenance * 17:54 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 17:54 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:54 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 17:32 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2215: Maintenance * 17:29 cmooney@dns3003: END - running authdns-update * 17:27 cmooney@dns3003: START - running authdns-update * 17:23 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 17:22 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:18 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool es2052: codfw rack B8 re-pool after maintenance * 17:18 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2189: codfw rack B8 re-pool after maintenance * 17:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2215.codfw.wmnet with reason: Maintenance * 17:17 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 17:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2215 [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95126 and previous config saved to /var/cache/conftool/dbconfig/20260723-170903-cwilliams.json * 17:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2191 to x1 primary [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95125 and previous config saved to /var/cache/conftool/dbconfig/20260723-170612-cwilliams.json * 17:05 cezmunsta: Starting x1 codfw failover from db2215 to db2191 - [[phab:T432986|T432986]] * 16:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2191 with weight 0 [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95123 and previous config saved to /var/cache/conftool/dbconfig/20260723-165831-cwilliams.json * 16:58 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 16 hosts with reason: Primary switchover x1 [[phab:T432986|T432986]] * 16:36 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 138128 * 16:35 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 138128 * 16:33 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2189: codfw rack B8 re-pool after maintenance * 16:33 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2164: codfw rack B8 re-pool after maintenance * 16:28 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1072.eqiad.wmnet * 16:27 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1072.eqiad.wmnet with OS trixie * 16:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2249: Maintenance * 16:06 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker1072.eqiad.wmnet with reason: host reimage * 16:06 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1072.eqiad.wmnet with reason: host reimage * 15:50 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1072 * 15:50 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1072 * 15:49 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1072.eqiad.wmnet with OS trixie * 15:48 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2164: codfw rack B8 re-pool after maintenance * 15:48 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] (duration: 06m 37s) * 15:48 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2163: codfw rack B8 re-pool after maintenance * 15:45 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 15:43 musikanimal@deploy1003: musikanimal: Continuing with deployment * 15:43 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:41 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] * 15:36 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1072.eqiad.wmnet * 15:35 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1072.eqiad.wmnet * 15:35 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1072.eqiad.wmnet * 15:34 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:34 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push any outstanding updates - cmooney@cumin1003" * 15:34 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push any outstanding updates - cmooney@cumin1003" * 15:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db2249: Maintenance * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 15:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:26 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:21 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:21 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 15:21 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:21 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 15:20 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 15:19 cmooney@dns2004: END - running authdns-update * 15:17 cmooney@dns2004: START - running authdns-update * 15:14 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns2004.wikimedia.org * 15:12 brouberol@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 15:12 brouberol@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 15:12 klausman@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ml-serve1001.eqiad.wmnet with OS trixie * 15:11 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1071.eqiad.wmnet * 15:11 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1071.eqiad.wmnet with OS trixie * 15:10 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wdqs2008.codfw.wmnet with OS bookworm * 15:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2249.codfw.wmnet with reason: Maintenance * 15:08 brouberol@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 15:08 brouberol@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 15:08 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 15:08 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 15:06 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns1004.wikimedia.org * 15:02 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2002.codfw.wmnet * 15:02 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2002.codfw.wmnet * 15:02 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2163: codfw rack B8 re-pool after maintenance * 15:01 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 15:01 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 14:59 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test2001.codfw.wmnet * 14:57 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test2001.codfw.wmnet * 14:56 ryankemper: [WDQS] [[phab:T430880|T430880]] Reimaged `wdqs2016` to Bookworm, xferred scholarly_articles from `wdqs2024`, validated updater/readiness/federation, and repooled * 14:51 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2016.codfw.wmnet * 14:51 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1001.eqiad.wmnet with reason: host reimage * 14:48 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1071.eqiad.wmnet with reason: host reimage * 14:47 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1001.eqiad.wmnet with reason: host reimage * 14:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2008.codfw.wmnet with reason: host reimage * 14:43 topranks: reboot lsw1-b8-codw to upgrade JunOS [[phab:T430929|T430929]] * 14:41 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1071.eqiad.wmnet with reason: host reimage * 14:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2008.codfw.wmnet with reason: host reimage * 14:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2231: Maintenance * 14:30 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1001.eqiad.wmnet with OS trixie * 14:25 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2002.codfw.wmnet * 14:23 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1071 * 14:23 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1071 * 14:23 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 14:22 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore scholarly data after Bookworm reimage) xfer scholarly_articles from wdqs2024.codfw.wmnet -> wdqs2016.codfw.wmnet, repooling source-only afterwards * 14:22 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2052: codfw rack B8 depool for maintenance * 14:21 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool es2052: codfw rack B8 depool for maintenance * 14:21 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2249: codfw rack B8 depool for maintenance * 14:21 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1071 * 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1071.eqiad.wmnet 166.48.64.10.in-addr.arpa 6.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:21 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1071.eqiad.wmnet 166.48.64.10.in-addr.arpa 6.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1071 - jiji@cumin1003" * 14:21 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1071 - jiji@cumin1003" * 14:21 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2249: codfw rack B8 depool for maintenance * 14:21 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2189: codfw rack B8 depool for maintenance * 14:20 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2002.codfw.wmnet * 14:20 cmooney@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2050.codfw.wmnet * 14:20 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2189: codfw rack B8 depool for maintenance * 14:20 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2164: codfw rack B8 depool for maintenance * 14:20 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2164: codfw rack B8 depool for maintenance * 14:19 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2163: codfw rack B8 depool for maintenance * 14:19 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:19 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1014 * 14:19 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2163: codfw rack B8 depool for maintenance * 14:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1014.eqiad.wmnet with OS bookworm * 14:17 cmooney@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2050.codfw.wmnet * 14:16 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 14:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2008 * 14:14 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2008 * 14:14 cmooney@cumin1003: conftool action : set/pooled=no; selector: name=dns2004.wikimedia.org * 14:14 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2008.codfw.wmnet with OS bookworm * 14:13 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 14:12 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:10 topranks: depool dns2004 before lsw1-b8-codfw switch maintenance [[phab:T430929|T430929]] * 14:10 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b8-codfw,lsw1-b8-codfw IPv6,lsw1-b8-codfw.mgmt,ssw1-a[1,8]-codfw with reason: lsw1-b8-codfw JunOS upgrade * 14:07 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 30 hosts with reason: lsw1-b8-codfw JunOS upgrade * 14:06 elukey: upload python3-docker-report 0.0.19 to apt.wikimedia.org for bookworm and trixie * 13:59 jiji@cumin1003: START - Cookbook sre.dns.netbox * 13:58 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1071 * 13:57 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:57 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:57 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:57 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1071.eqiad.wmnet with OS trixie * 13:55 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1071.eqiad.wmnet * 13:55 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1071.eqiad.wmnet * 13:55 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1071.eqiad.wmnet * 13:53 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:52 logmsgbot: kharlan Deployed security patch for [[phab:T432948|T432948]] * 13:51 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:51 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 13:51 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db2231: Maintenance * 13:50 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:50 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:50 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:50 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:49 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:49 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:49 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2231 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95097 and previous config saved to /var/cache/conftool/dbconfig/20260723-134436-cwilliams.json * 13:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2231.codfw.wmnet with reason: Maintenance * 13:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:39 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:38 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 13:38 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] (duration: 09m 07s) * 13:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1037 hosts * 13:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2196: Maintenance * 13:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:33 kharlan@deploy1003: kharlan, emc-wmf: Continuing with deployment * 13:31 kharlan@deploy1003: kharlan, emc-wmf: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:30 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:28 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] * 13:17 hashar@deploy1003: Finished deploy [integration/docroot@2199146]: build: License GPL2.0+ / updating npm dependencies (duration: 00m 14s) * 13:17 hashar@deploy1003: Started deploy [integration/docroot@2199146]: build: License GPL2.0+ / updating npm dependencies * 13:14 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service * 13:07 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 12:58 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2207: Repooling * 12:49 root@cumin1003: START - Cookbook sre.mysql.pool pool db2196: Maintenance * 12:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2196 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95087 and previous config saved to /var/cache/conftool/dbconfig/20260723-123952-cwilliams.json * 12:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2196.codfw.wmnet with reason: Maintenance * 12:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2191: Maintenance * 12:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:13 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:13 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: Repooling * 12:12 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2207: Repooling * 12:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: Repooling * 11:56 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2235.codfw.wmnet with OS trixie * 11:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db2191: Maintenance * 11:46 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1070.eqiad.wmnet * 11:46 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1070.eqiad.wmnet * 11:46 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1070.eqiad.wmnet * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2191 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95080 and previous config saved to /var/cache/conftool/dbconfig/20260723-114308-cwilliams.json * 11:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2191.codfw.wmnet with reason: Maintenance * 11:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2186: Maintenance * 11:35 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 46375 * 11:34 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 46375 * 11:33 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2235.codfw.wmnet with reason: host reimage * 11:28 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2235.codfw.wmnet with reason: host reimage * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c7-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c7-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c6-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c6-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c5-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c5-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c4-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c4-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c3-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c3-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c2-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c2-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d7-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d7-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d4-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d4-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d3-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d2-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d2-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d8-eqiad * 11:23 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d8-eqiad * 11:23 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d1-eqiad * 11:23 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d1-eqiad * 11:12 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db2235.codfw.wmnet with OS trixie * 11:11 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:11 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[2160,2235].codfw.wmnet with reason: Upgrading * 10:56 root@cumin1003: START - Cookbook sre.mysql.pool pool db2186: Maintenance * 10:54 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1037: testing * 10:53 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1037: testing * 10:53 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1037: testing * 10:53 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1037: testing * 10:52 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: testing * 10:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2186 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95072 and previous config saved to /var/cache/conftool/dbconfig/20260723-104956-cwilliams.json * 10:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2186.codfw.wmnet with reason: Maintenance * 10:43 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1054: testing * 10:41 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1070.eqiad.wmnet with OS trixie * 10:30 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1037 hosts * 10:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 10:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 10:20 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1070.eqiad.wmnet with reason: host reimage * 10:16 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1070.eqiad.wmnet with reason: host reimage * 10:06 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1038: testing * 10:05 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1038: testing * 10:05 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1038: testing * 10:02 marostegui@dns1004: END - running authdns-update * 10:00 marostegui@dns1004: START - running authdns-update * 09:58 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: testing * 09:57 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1054: testing * 09:57 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1070 * 09:57 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1070 * 09:57 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1054: testing * 09:57 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1054: testing * 09:56 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1055: testing * 09:56 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1070 * 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1070.eqiad.wmnet 165.48.64.10.in-addr.arpa 5.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:56 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1070.eqiad.wmnet 165.48.64.10.in-addr.arpa 5.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1070 - jiji@cumin1003" * 09:56 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1070 - jiji@cumin1003" * 09:47 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for 1035 hosts * 09:45 jiji@cumin1003: START - Cookbook sre.dns.netbox * 09:42 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1070 * 09:42 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1070.eqiad.wmnet with OS trixie * 09:42 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1070.eqiad.wmnet * 09:41 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1070.eqiad.wmnet * 09:41 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1070.eqiad.wmnet * 09:27 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es2051: testing * 09:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: testing * 09:12 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es2051: testing * 09:11 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1055: testing * 09:09 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:09 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1055: testing * 09:09 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1055: testing * 08:50 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:50 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:50 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 08:49 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 08:49 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 08:49 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:46 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1069.eqiad.wmnet * 08:46 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1069.eqiad.wmnet * 08:46 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1069.eqiad.wmnet * 08:39 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 08:38 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 08:38 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 08:37 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 08:35 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:10 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1069.eqiad.wmnet with OS trixie * 07:49 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1069.eqiad.wmnet with reason: host reimage * 07:45 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1069.eqiad.wmnet with reason: host reimage * 07:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1035 hosts * 07:33 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 07:32 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 07:29 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1069 * 07:29 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1069 * 07:26 jiji@deploy1003: Finished scap sync-world: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules (duration: 06m 01s) * 07:25 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1069 * 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1069.eqiad.wmnet 164.48.64.10.in-addr.arpa 4.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:25 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1069.eqiad.wmnet 164.48.64.10.in-addr.arpa 4.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1069 - jiji@cumin1003" * 07:25 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1069 - jiji@cumin1003" * 07:25 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 07:25 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 07:24 jiji@deploy1003: jiji: Continuing with deployment * 07:22 jiji@deploy1003: jiji: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:21 jiji@deploy1003: Started scap sync-world: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules * 07:21 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1031.eqiad.wmnet,service=s7 * 07:20 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1031.eqiad.wmnet,service=s2 * 07:20 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1031.eqiad.wmnet,service=s7 * 07:20 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1031.eqiad.wmnet,service=s2 * 07:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts * 07:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts * 07:19 jiji@cumin1003: START - Cookbook sre.dns.netbox * 07:19 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1069 * 07:19 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1069.eqiad.wmnet with OS trixie * 07:19 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1069.eqiad.wmnet * 07:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 45 hosts * 07:17 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1069.eqiad.wmnet * 07:17 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1069.eqiad.wmnet * 07:14 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 45 hosts * 07:13 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 06:16 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs2007 after successful Bookworm reimage, data transfer, and postflight validation; wdqs1013 also passed postflights and is enabled in conftool, but remains out of IPVS pending a rolling pybal restart to clear its stale pre-VLAN-move address * 05:58 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2007.codfw.wmnet * 05:58 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1013.eqiad.wmnet * 05:54 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore scholarly data after Bookworm reimage) xfer scholarly_articles from wdqs2024.codfw.wmnet -> wdqs2016.codfw.wmnet, repooling source-only afterwards == 2026-07-22 == * 23:34 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Apply upgrade to JVM17 - eevans@cumin1003 * 23:14 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Apply upgrade to JVM17 - eevans@cumin1003 * 22:06 ryankemper: [WDQS] Added requestctl per-IP ratelimit `wdqs_heavy_sparql_bots_jul_2026_ratelimit` (chronic heavy-query bot tier driving deadlock-remediation restarts); pruned superseded `wdqs_2026_05_11_worobot` * 21:51 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] (duration: 11m 52s) * 21:44 sbassett@deploy1003: sbassett: Continuing with deployment * 21:43 sbassett@deploy1003: sbassett: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:39 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] * 20:38 dani@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] (duration: 32m 51s) * 20:38 ryankemper: [WDQS] Pruned obsolete requestctl action+pattern `wdqs_20260715_p2003_ring_ja3n` (actor rotated JA3Ns; rule inert) * 20:26 dani@deploy1003: dani, vadymts1: Continuing with deployment * 20:24 dani@deploy1003: dani, vadymts1: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:14 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2244: Testing * 20:06 dani@deploy1003: Started scap sync-world: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] * 19:56 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1013.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2016.codfw.wmnet with OS bookworm * 19:40 mutante: gerrit - one more service restart is needed - restarting * 19:29 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2244: Testing * 19:27 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2244: Testing * 19:27 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2244: Testing * 19:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2016.codfw.wmnet with reason: host reimage * 19:15 dancy@deploy1003: Finished deploy [zuul/deploy@d92e238]: Freshening Zuul installation (duration: 00m 15s) * 19:14 dancy@deploy1003: Started deploy [zuul/deploy@d92e238]: Freshening Zuul installation * 19:11 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2016.codfw.wmnet with reason: host reimage * 18:54 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1013.eqiad.wmnet, repooling source-only afterwards * 18:52 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2016 * 18:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2016 * 18:51 dancy@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 18:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2016.codfw.wmnet with OS bookworm * 18:39 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 11s) * 18:39 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 18:36 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 18:30 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417] (thin): Regular analytics weekly train THIN [analytics/refinery@2a25417d] (duration: 02m 09s) * 18:28 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417] (thin): Regular analytics weekly train THIN [analytics/refinery@2a25417d] * 18:28 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417]: Regular analytics weekly train [analytics/refinery@2a25417d] (duration: 04m 31s) * 18:27 dduvall: deploying https://gerrit.wikimedia.org/r/c/integration/config/+/1314025 (4 jobs updated) * 18:23 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417]: Regular analytics weekly train [analytics/refinery@2a25417d] * 18:22 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@2a25417d] (duration: 01m 59s) * 18:20 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@2a25417d] * 17:56 Raine: deployment server switchover => deploy1003 is primary now * 17:55 kamila@deploy1003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 28m 24s) * 17:54 mutante: restarting gerrit for maintenance * 17:29 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1013.eqiad.wmnet with OS bookworm * 17:27 kamila@deploy1003: Started scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] * 17:20 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] (duration: 22m 50s) * 17:12 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1023.eqiad.wmnet -> wdqs1024.eqiad.wmnet, repooling source-only afterwards * 17:04 Raine: point deployment.eqiad.wmnet to deploy1003 * 17:04 kamila@dns7001: END - running authdns-update * 17:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1013.eqiad.wmnet with reason: host reimage * 17:02 kamila@dns7001: START - running authdns-update * 17:01 kamila@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on releases2003.codfw.wmnet,releases1003.eqiad.wmnet with reason: Deployment server switchover * 17:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1013.eqiad.wmnet with reason: host reimage * 16:58 kamila@deploy2003: Locking from deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] * 16:57 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] (duration: 02m 33s) * 16:55 kamila@deploy2003: Locking from deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] * 16:55 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2003 - [[phab:T240266|T240266]] (duration: 00m 11s) * 16:54 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2003 - [[phab:T240266|T240266]] * 16:40 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1013 * 16:40 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1013 * 16:39 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1013 * 16:39 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1013.eqiad.wmnet 105.32.64.10.in-addr.arpa 5.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:39 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1013.eqiad.wmnet 105.32.64.10.in-addr.arpa 5.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:39 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:39 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1013 - bking@cumin2003" * 16:39 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1013 - bking@cumin2003" * 16:34 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:34 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1013 * 16:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1013.eqiad.wmnet with OS bookworm * 16:28 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1023.eqiad.wmnet -> wdqs1024.eqiad.wmnet, repooling source-only afterwards * 16:27 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-scholarly,name=eqiad * 16:27 eevans@deploy2003: helmfile [eqiad] DONE helmfile.d/services/linked-artifacts: apply * 16:26 eevans@deploy2003: helmfile [eqiad] START helmfile.d/services/linked-artifacts: apply * 16:26 eevans@deploy2003: helmfile [codfw] DONE helmfile.d/services/linked-artifacts: apply * 16:26 eevans@deploy2003: helmfile [codfw] START helmfile.d/services/linked-artifacts: apply * 16:25 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 29s) * 16:25 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 16:24 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 16:21 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 16:18 eevans@deploy2003: helmfile [codfw] DONE helmfile.d/services/linked-artifacts: apply * 16:18 eevans@deploy2003: helmfile [codfw] START helmfile.d/services/linked-artifacts: apply * 16:08 eevans@deploy2003: helmfile [staging] DONE helmfile.d/services/linked-artifacts: apply * 16:07 eevans@deploy2003: helmfile [staging] START helmfile.d/services/linked-artifacts: apply * 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 16:01 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 15:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1024.eqiad.wmnet with OS bookworm * 15:49 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] (duration: 00m 10s) * 15:49 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] * 15:48 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] (duration: 00m 15s) * 15:48 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] * 15:47 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] (duration: 00m 10s) * 15:47 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] * 15:46 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:42 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:42 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:40 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:37 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 15:37 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:36 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:36 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:36 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 15:33 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 15:33 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:31 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:28 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 15:27 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:27 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1024.eqiad.wmnet with reason: host reimage * 15:23 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:23 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:23 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:20 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1068.eqiad.wmnet * 15:20 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1068.eqiad.wmnet * 15:20 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1068.eqiad.wmnet * 15:20 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 15:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1024.eqiad.wmnet with reason: host reimage * 15:11 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:55 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 14:52 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wdqs1024.eqiad.wmnet with OS bookworm * 14:50 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] (duration: 00m 09s) * 14:50 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] * 14:49 jiji@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 14:49 jiji@deploy2003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 14:49 jiji@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 14:48 jiji@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 14:45 ecarg@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:45 ecarg@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:44 ecarg@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:44 ecarg@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:43 ecarg@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:43 ecarg@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:41 sukhe: ipvsadm --delete-service --tcp-service 10.2.1.55:8087: lvs2014 and lvs2013 * 14:39 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:39 sukhe: ipvsadm --delete-service --tcp-service 10.2.2.55:8087: [[phab:T432445|T432445]] * 14:38 ecarg@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:38 ecarg@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:37 ecarg@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:37 ecarg@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:36 ecarg@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:34 ecarg@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts datahubsearch1001.eqiad.wmnet * 14:32 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:32 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 14:31 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 14:31 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 14:28 sukhe: sudo cumin 'A:lvs-low-traffic-codfw' 'systemctl restart pybal': lvs2013 * 14:26 sukhe: sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal': lvs2014 * 14:26 sukhe: sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal' * 14:24 sukhe: restart pybal on lvs1019 * 14:24 sukhe: restart pybal on lvs1020 * 14:19 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for 1036 hosts * 14:17 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:04 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] (duration: 09m 28s) * 13:59 kharlan@deploy2003: dreamyjazz, kharlan: Continuing with deployment * 13:58 bking@cumin2003: START - Cookbook sre.hosts.decommission for hosts datahubsearch1001.eqiad.wmnet * 13:57 kharlan@deploy2003: dreamyjazz, kharlan: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:55 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] * 13:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts datahubsearch[1002-1003].eqiad.wmnet * 13:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:53 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch[1002-1003].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 13:52 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch[1002-1003].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 13:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 13:42 stran@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] (duration: 07m 30s) * 13:42 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:40 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: Test * 13:38 stran@deploy2003: dragoniez, stran: Continuing with deployment * 13:37 stran@deploy2003: dragoniez, stran: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:35 bking@cumin2003: START - Cookbook sre.hosts.decommission for hosts datahubsearch[1002-1003].eqiad.wmnet * 13:35 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024'] * 13:35 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 13:35 stran@deploy2003: Started scap sync-world: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] * 13:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 13:28 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024'] * 13:26 sukhe@dns1004: END - running authdns-update * 13:25 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 13:25 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 13:24 sukhe@dns1004: START - running authdns-update * 13:22 stran@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] (duration: 08m 20s) * 13:21 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 13:20 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 13:19 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 13:19 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 13:19 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 13:18 stran@deploy2003: stran: Continuing with deployment * 13:16 stran@deploy2003: stran: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:14 stran@deploy2003: Started scap sync-world: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] * 13:13 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 13:13 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 13:11 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 13:11 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 13:08 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 12:55 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool es2051: Test * 12:55 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: Test * 12:54 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool es2051: Test * 12:43 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1036 hosts * 12:41 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 12:40 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1048.eqiad.wmnet * 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 12:39 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 12:38 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 12:37 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 12:37 brouberol@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 12:36 brouberol@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 12:36 brouberol@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 12:36 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 12:36 elukey@cumin1003: DONE (PASS) - Cookbook sre.puppet.renew-cert (exit_code=0) for crm2001.codfw.wmnet: Renew puppet certificate - elukey@cumin1003 * 12:35 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:35 brouberol@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 12:34 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 12:31 brouberol@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 12:30 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1048.eqiad.wmnet * 12:30 brouberol@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 12:28 brouberol@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 12:27 brouberol@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 12:20 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1068.eqiad.wmnet with OS trixie * 12:01 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] (duration: 13m 19s) * 11:58 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1068.eqiad.wmnet with reason: host reimage * 11:52 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1068.eqiad.wmnet with reason: host reimage * 11:51 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 11:49 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:47 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] * 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2252: Security updates * 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:43 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 11:42 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2252: Security updates * 11:42 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply * 11:40 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply * 11:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 11:37 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2252.codfw.wmnet with OS trixie * 11:34 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1068 * 11:34 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1068 * 11:26 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1068 * 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1068.eqiad.wmnet 46.48.64.10.in-addr.arpa 6.4.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:26 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1068.eqiad.wmnet 46.48.64.10.in-addr.arpa 6.4.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1068 - jiji@cumin1003" * 11:26 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1068 - jiji@cumin1003" * 11:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2252.codfw.wmnet with reason: host reimage * 11:17 jiji@cumin1003: START - Cookbook sre.dns.netbox * 11:17 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1068 * 11:17 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1068.eqiad.wmnet with OS trixie * 11:17 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2252.codfw.wmnet with reason: host reimage * 11:15 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1068.eqiad.wmnet * 11:15 mvolz@deploy2003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:15 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1068.eqiad.wmnet * 11:15 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1068.eqiad.wmnet * 11:14 mvolz@deploy2003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:13 mvolz@deploy2003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:13 mvolz@deploy2003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:12 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] (duration: 11m 05s) * 11:11 mvolz@deploy2003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:10 mvolz@deploy2003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:07 dreamyjazz@deploy2003: dreamyjazz, kharlan: Continuing with deployment * 11:03 dreamyjazz@deploy2003: dreamyjazz, kharlan: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:03 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2252.codfw.wmnet with OS trixie * 11:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2252: Upgrading db2252.codfw.wmnet * 11:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:02 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 11:02 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2252: Upgrading db2252.codfw.wmnet * 11:02 cwilliams@cumin1003: dbmaint on ms3@codfw [[phab:T432321|T432321]] * 11:01 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 11:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db1153.eqiad.wmnet with reason: Security updates * 11:01 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] * 11:00 fnegri@deploy2003: helmfile [eqiad] DONE helmfile.d/services/toolhub: apply * 10:58 fnegri@deploy2003: helmfile [eqiad] START helmfile.d/services/toolhub: apply * 10:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1151: Security updates * 10:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:57 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1151: Security updates * 10:55 fnegri@deploy2003: helmfile [codfw] DONE helmfile.d/services/toolhub: apply * 10:54 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] (duration: 08m 38s) * 10:53 fnegri@deploy2003: helmfile [codfw] START helmfile.d/services/toolhub: apply * 10:53 fnegri@deploy2003: helmfile [staging] DONE helmfile.d/services/toolhub: apply * 10:52 fnegri@deploy2003: helmfile [staging] START helmfile.d/services/toolhub: apply * 10:50 zabe@deploy2003: zabe: Continuing with deployment * 10:47 zabe@deploy2003: zabe: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:45 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] * 10:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1151: Security updates * 10:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:42 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:42 root@cumin1003: START - Cookbook sre.mysql.depool depool db1151: Security updates * 10:38 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] (duration: 12m 47s) * 10:34 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 10:34 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 10:33 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2253.codfw.wmnet with OS trixie * 10:28 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:26 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] * 10:18 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2253.codfw.wmnet with reason: host reimage * 10:13 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2253.codfw.wmnet with reason: host reimage * 10:00 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2253.codfw.wmnet with OS trixie * 09:58 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db1151.eqiad.wmnet with reason: Security updates * 09:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2253: Upgrading db2253.codfw.wmnet * 09:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:57 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 09:56 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2253: Upgrading db2253.codfw.wmnet * 09:56 cwilliams@cumin1003: dbmaint on ms2@codfw [[phab:T432321|T432321]] * 09:56 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 09:36 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: UI improvement; support url shortener - oblivian@cumin1003" * 09:36 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: UI improvement; support url shortener - oblivian@cumin1003 * 09:35 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: UI improvement; support url shortener - oblivian@cumin1003 * 09:35 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: UI improvement; support url shortener - oblivian@cumin1003" * 09:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1152: Security updates * 09:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:26 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db1152: Security updates * 09:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: Security updates * 09:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:11 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:11 root@cumin1003: START - Cookbook sre.mysql.depool depool db1152: Security updates * 09:10 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1018.eqiad.wmnet with reason: Cloning * 09:09 Dreamy_Jazz: Deployed patch for [[phab:T432453|T432453]] and [[phab:T432454|T432454]] * 09:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 09:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2251.codfw.wmnet with OS trixie * 08:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2251.codfw.wmnet with reason: host reimage * 08:45 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2251.codfw.wmnet with reason: host reimage * 08:40 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] (duration: 12m 26s) * 08:38 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1030.eqiad.wmnet,service=s1 * 08:36 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 08:31 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2251.codfw.wmnet with OS trixie * 08:30 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:28 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] * 08:25 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] (duration: 07m 59s) * 08:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2251: Upgrading db2251.codfw.wmnet * 08:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:22 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 08:22 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2251: Upgrading db2251.codfw.wmnet * 08:20 urbanecm@deploy2003: urbanecm: Continuing with deployment * 08:20 cwilliams@cumin1003: dbmaint on ms1@codfw [[phab:T432321|T432321]] * 08:20 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 08:19 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:17 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] * 08:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade * 08:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade * 08:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db2251.codfw.wmnet,db1152.eqiad.wmnet with reason: OS upgrade * 08:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade * 08:13 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade * 08:11 Dreamy_Jazz: Created cusi_signal, cusi_case, and cusi_user on ukwiki and enwikivoyage in extension1 * 08:11 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade * 08:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade * 08:04 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1030.eqiad.wmnet,service=s1 * 08:04 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1030.eqiad.wmnet,service=s1 * 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply * 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply * 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply * 07:51 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply * 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 07:47 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 07:47 phuedx: End of UTC morning backport window * 07:43 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 07:43 phuedx@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] (duration: 13m 44s) * 07:43 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 07:39 phuedx@deploy2003: phuedx: Continuing with deployment * 07:31 phuedx@deploy2003: phuedx: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:29 phuedx@deploy2003: Started scap sync-world: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] * 07:24 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Turnilo import support - oblivian@cumin1003" * 07:24 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import support - oblivian@cumin1003 * 07:23 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import support - oblivian@cumin1003 * 07:23 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Turnilo import support - oblivian@cumin1003" * 06:42 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs2020 after successful Bookworm reimage, data transfer, and postflight validation * 06:42 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2020.codfw.wmnet * 05:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Managing sanitization for wikis bolwiki in section s5 * 05:25 marostegui@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis bolwiki in section s5 * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 41s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-21 == * 22:50 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2019.codfw.wmnet -> wdqs2020.codfw.wmnet, repooling source-only afterwards * 22:47 cwhite: force reboot arclamp2001 - appears to have run out of memory and gone unresponsive * 22:24 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 01m 26s) * 22:24 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 22:23 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 22:22 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024'] * 22:11 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 21:54 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs1024'] * 21:54 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 21:53 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs1024'] * 21:53 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 21:49 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2019.codfw.wmnet -> wdqs2020.codfw.wmnet, repooling source-only afterwards * 20:57 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] (duration: 09m 10s) * 20:55 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1024.eqiad.wmnet with OS bookworm * 20:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2020.codfw.wmnet with OS bookworm * 20:52 krinkle@deploy2003: krinkle: Continuing with deployment * 20:49 krinkle@deploy2003: krinkle: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:47 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] * 20:45 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] (duration: 05m 42s) * 20:44 krinkle@deploy2003: krinkle: Rolling back deployment * 20:41 krinkle@deploy2003: krinkle: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:39 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] * 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2003.codfw.wmnet * 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1003.eqiad.wmnet * 20:33 dani@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] (duration: 11m 15s) * 20:33 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2003.codfw.wmnet * 20:33 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1003.eqiad.wmnet * 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2020.codfw.wmnet with reason: host reimage * 20:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1002.eqiad.wmnet * 20:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2002.codfw.wmnet * 20:30 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 20:29 dani@deploy2003: dani: Continuing with deployment * 20:29 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2020.codfw.wmnet with reason: host reimage * 20:26 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1002.eqiad.wmnet * 20:26 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2002.codfw.wmnet * 20:24 dani@deploy2003: dani: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2001.codfw.wmnet * 20:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1001.eqiad.wmnet * 20:22 dani@deploy2003: Started scap sync-world: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] * 20:22 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 20:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2001.codfw.wmnet * 20:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1001.eqiad.wmnet * 20:14 sbisson@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] (duration: 09m 01s) * 20:11 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2020.codfw.wmnet with OS bookworm * 20:10 sbisson@deploy2003: sbisson: Continuing with deployment * 20:07 sbisson@deploy2003: sbisson: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:05 sbisson@deploy2003: Started scap sync-world: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] * 20:03 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] (duration: 07m 04s) * 20:01 mutante: Gerrit - tomorrow a new SSH host key will appear - it will be {{Gerrit|ed25519}} and has already been added to wmf-laptop. you can verify it here: https://wikitech.wikimedia.org/wiki/Help:SSH_Fingerprints/gerrit.wikimedia.org:29418 ([[phab:T240266|T240266]]) * 19:59 zabe@deploy2003: zabe: Continuing with deployment * 19:58 zabe@deploy2003: zabe: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:56 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] * 19:52 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] (duration: 07m 25s) * 19:48 zabe@deploy2003: zabe: Continuing with deployment * 19:47 zabe@deploy2003: zabe: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:45 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] * 19:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 19:32 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024'] * 19:27 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 19:26 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024'] * 19:26 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 19:24 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024'] * 19:12 ryankemper: [wdqs] [[phab:T430880|T430880]] Repooled `wdqs-scholarly` discovery in `eqiad` after validating `wdqs1023` end-to-end; `wdqs1024` remains disabled pending reimage recovery * 19:11 ryankemper: [wdqs] [[phab:T430880|T430880]] Repooled wdqs1012.eqiad.wmnet after successful Bookworm reimage, data transfer, service checks, readiness probe, and cross-graph federation query validation * 19:10 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 19:10 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1012.eqiad.wmnet * 19:08 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 18:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for deploy1003.eqiad.wmnet * 18:57 kamila@cumin1003: START - Cookbook sre.hosts.remove-downtime for deploy1003.eqiad.wmnet * 18:37 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:37 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding urldownloader service IPs - sukhe@cumin1003" * 18:37 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding urldownloader service IPs - sukhe@cumin1003" * 18:32 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 18:32 dancy@deploy2003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 18:30 sukhe@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 18:27 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 18:24 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1024.eqiad.wmnet with OS bookworm * 18:20 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host deploy1003.eqiad.wmnet with OS bookworm * 18:09 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deploy1003 reimage (duration: 121m 16s) * 18:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 18:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1180: Security updates * 17:55 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wcqs2003.codfw.wmnet * 17:48 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wcqs2003.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1155.eqiad.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1155.eqiad.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2224.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2224.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2217.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2217.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2193.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2193.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2180.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2180.codfw.wmnet * 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1168.eqiad.wmnet * 17:36 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1168.eqiad.wmnet * 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2169.codfw.wmnet * 17:36 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2169.codfw.wmnet * 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1165.eqiad.wmnet * 17:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1165.eqiad.wmnet * 17:35 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2158.codfw.wmnet * 17:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2158.codfw.wmnet * 17:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wcqs1003.eqiad.wmnet * 17:17 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1180: Security updates * 17:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1180.eqiad.wmnet * 17:16 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1180.eqiad.wmnet * 17:15 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp3073.* * 17:13 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wcqs1003.eqiad.wmnet * 17:11 brett@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp3073.esams.wmnet with OS trixie * 17:11 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 17:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1024 * 17:04 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1024 * 17:03 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 17:00 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 16:59 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2242: codfw rack B7 depool for maintenance * 16:59 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 16:43 brett@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp3073.esams.wmnet with reason: host reimage * 16:42 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 16:39 brett@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cp3073.esams.wmnet with reason: host reimage * 16:32 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on deploy1003.eqiad.wmnet with reason: host reimage * 16:27 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on deploy1003.eqiad.wmnet with reason: host reimage * 16:14 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2242: codfw rack B7 depool for maintenance * 16:14 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: codfw rack B7 depool for maintenance * 16:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1023.eqiad.wmnet with OS bookworm * 16:13 brett@cumin2002: START - Cookbook sre.hosts.reimage for host cp3073.esams.wmnet with OS trixie * 16:08 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host deploy1003.eqiad.wmnet with OS bookworm * 16:08 kamila@deploy2003: Locking from deployment [MediaWiki]: deploy1003 reimage * 16:03 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1012.eqiad.wmnet with OS bookworm * 15:48 inflatador: bking@apt1002 `sudo reprepro copy bookworm-wikimedia bullseye-wikimedia jvmquake` [[phab:T430880|T430880]] * 15:39 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp3073.* * 15:39 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 15:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:34 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1311536{{!}}Set $wgMathInternalRestbaseURL explicitly (take 2) (T349582)]] * 15:29 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:29 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2228: codfw rack B7 depool for maintenance * 15:29 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2229: codfw rack B7 depool for maintenance * 15:27 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1180: Security update * 15:25 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Security update * 15:21 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 15:21 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 15:19 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 24s) * 15:19 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:14 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 15:14 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db1180: Security update * 15:13 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-eqiad * 14:48 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-eqiad * 14:44 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2229: codfw rack B7 depool for maintenance * 14:44 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc2017: codfw rack B7 depool for maintenance * 14:44 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:43 cmooney@cumin2003: START - Cookbook sre.mysql.parsercache * 14:43 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool pc2017: codfw rack B7 depool for maintenance * 14:43 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2003.codfw.wmnet * 14:43 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2003.codfw.wmnet * 14:42 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2009.codfw.wmnet * 14:42 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2009.codfw.wmnet * 14:41 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:41 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:40 cmooney@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 29 hosts * 14:40 cmooney@cumin1003: START - Cookbook sre.hosts.remove-downtime for 29 hosts * 14:35 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 14:34 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 14:32 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] (duration: 07m 56s) * 14:29 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ssw1-a[1,8]-codfw with reason: lsw1-b7-codfw JunOS upgrade * 14:28 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 14:28 elukey@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 14:26 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:24 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] * 14:23 topranks: reboot lsw1-b7-codfw to upgrade JunOS (affects all hosts in rack) [[phab:T430928|T430928]] * 14:18 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2003.codfw.wmnet * 14:14 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2009.codfw.wmnet * 14:14 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Security update * 14:13 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2242: codfw rack B7 depool for maintenance * 14:13 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2242: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2228: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2228: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2229: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2229: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc2017: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.parsercache * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool pc2017: codfw rack B7 depool for maintenance * 14:08 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2003.codfw.wmnet * 14:07 cmooney@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on aux-k8s-etcd2004.codfw.wmnet,ml-etcd2001.codfw.wmnet with reason: lsw1-b7-codfw JunOS upgrade * 14:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95005 and previous config saved to /var/cache/conftool/dbconfig/20260721-140620-cwilliams.json * 14:05 cmooney@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2049.codfw.wmnet * 14:05 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-scholarly,name=eqiad * 14:04 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2009.codfw.wmnet * 14:04 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply * 14:04 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply * 14:03 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:03 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2001.codfw.wmnet * 14:03 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2001.codfw.wmnet * 14:02 cmooney@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2049.codfw.wmnet * 14:00 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:00 Dreamy_Jazz: Created cusi_case, cusi_signal, and cusi_user on svwiki, dewiki, jawiki, eswiki * 13:59 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b7-codfw,lsw1-b7-codfw IPv6,lsw1-b7-codfw.mgmt,ssw1-a[1,8]-codfw.mgmt with reason: lsw1-b7-codfw JunOS upgrade * 13:57 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs1023.eqiad.wmnet, repooling source-only afterwards * 13:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 29 hosts with reason: lsw1-b7-codfw JunOS upgrade * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224', diff saved to https://phabricator.wikimedia.org/P95003 and previous config saved to /var/cache/conftool/dbconfig/20260721-135613-cwilliams.json * 13:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1012.eqiad.wmnet with reason: host reimage * 13:53 cmooney@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 1:00:00 on 30 hosts with reason: lsw1-b7-codfw JunOS upgrade * 13:51 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1012.eqiad.wmnet with reason: host reimage * 13:48 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 13:48 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 13:46 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224', diff saved to https://phabricator.wikimedia.org/P95001 and previous config saved to /var/cache/conftool/dbconfig/20260721-134605-cwilliams.json * 13:46 cmooney@cumin1003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti2033.codfw.wmnet * 13:46 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 13:45 elukey: move the Docker Registry's /v2/wikimedia/machinelearning.* prefix to the ml S3 backend - [[phab:T428022|T428022]] * 13:45 cmooney@cumin1003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti2033.codfw.wmnet * 13:45 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 13:43 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:40 jiji@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 13:40 jiji@deploy2003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 13:39 jiji@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 13:39 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 13:38 cmooney@cumin1003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2032.codfw.wmnet * 13:38 jiji@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 13:38 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 13:37 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2032.codfw.wmnet * 13:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95000 and previous config saved to /var/cache/conftool/dbconfig/20260721-133557-cwilliams.json * 13:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1012 * 13:33 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1012 * 13:33 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1012.eqiad.wmnet with OS bookworm * 13:30 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:30 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:28 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94999 and previous config saved to /var/cache/conftool/dbconfig/20260721-132855-cwilliams.json * 13:28 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2224.codfw.wmnet with reason: Maintenance * 13:28 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94998 and previous config saved to /var/cache/conftool/dbconfig/20260721-132826-cwilliams.json * 13:28 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 13:23 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] (duration: 07m 50s) * 13:20 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 13:18 kharlan@deploy2003: kharlan: Continuing with deployment * 13:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217', diff saved to https://phabricator.wikimedia.org/P94996 and previous config saved to /var/cache/conftool/dbconfig/20260721-131817-cwilliams.json * 13:17 brouberol@dns1004: END - running authdns-update * 13:17 kharlan@deploy2003: kharlan: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:15 brouberol@dns1004: START - running authdns-update * 13:15 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] * 13:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94995 and previous config saved to /var/cache/conftool/dbconfig/20260721-131411-cwilliams.json * 13:13 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs1023.eqiad.wmnet, repooling source-only afterwards * 13:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217', diff saved to https://phabricator.wikimedia.org/P94994 and previous config saved to /var/cache/conftool/dbconfig/20260721-130809-cwilliams.json * 13:07 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 13:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180', diff saved to https://phabricator.wikimedia.org/P94993 and previous config saved to /var/cache/conftool/dbconfig/20260721-130404-cwilliams.json * 13:03 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:03 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:02 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 13:02 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 12:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94992 and previous config saved to /var/cache/conftool/dbconfig/20260721-125801-cwilliams.json * 12:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180', diff saved to https://phabricator.wikimedia.org/P94991 and previous config saved to /var/cache/conftool/dbconfig/20260721-125356-cwilliams.json * 12:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94990 and previous config saved to /var/cache/conftool/dbconfig/20260721-125049-cwilliams.json * 12:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2217.codfw.wmnet with reason: Maintenance * 12:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94989 and previous config saved to /var/cache/conftool/dbconfig/20260721-125017-cwilliams.json * 12:48 elukey: bmc cold reboot for lvs1013 and lvs1015 - [[phab:T426180|T426180]] * 12:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94988 and previous config saved to /var/cache/conftool/dbconfig/20260721-124348-cwilliams.json * 12:40 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193', diff saved to https://phabricator.wikimedia.org/P94987 and previous config saved to /var/cache/conftool/dbconfig/20260721-124009-cwilliams.json * 12:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts * 12:33 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts * 12:33 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts * 12:32 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts * 12:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:30 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193', diff saved to https://phabricator.wikimedia.org/P94986 and previous config saved to /var/cache/conftool/dbconfig/20260721-123001-cwilliams.json * 12:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94985 and previous config saved to /var/cache/conftool/dbconfig/20260721-121953-cwilliams.json * 12:17 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs2007.codfw.wmnet with OS bookworm * 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94983 and previous config saved to /var/cache/conftool/dbconfig/20260721-121257-cwilliams.json * 12:12 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2193.codfw.wmnet with reason: Maintenance * 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94982 and previous config saved to /var/cache/conftool/dbconfig/20260721-121239-cwilliams.json * 12:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180', diff saved to https://phabricator.wikimedia.org/P94980 and previous config saved to /var/cache/conftool/dbconfig/20260721-120231-cwilliams.json * 11:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180', diff saved to https://phabricator.wikimedia.org/P94979 and previous config saved to /var/cache/conftool/dbconfig/20260721-115223-cwilliams.json * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94978 and previous config saved to /var/cache/conftool/dbconfig/20260721-114333-cwilliams.json * 11:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1180.eqiad.wmnet with reason: Maintenance * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94977 and previous config saved to /var/cache/conftool/dbconfig/20260721-114305-cwilliams.json * 11:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94976 and previous config saved to /var/cache/conftool/dbconfig/20260721-114215-cwilliams.json * 11:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94975 and previous config saved to /var/cache/conftool/dbconfig/20260721-113530-cwilliams.json * 11:35 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2180.codfw.wmnet with reason: Maintenance * 11:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94974 and previous config saved to /var/cache/conftool/dbconfig/20260721-113501-cwilliams.json * 11:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168', diff saved to https://phabricator.wikimedia.org/P94973 and previous config saved to /var/cache/conftool/dbconfig/20260721-113258-cwilliams.json * 11:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169', diff saved to https://phabricator.wikimedia.org/P94972 and previous config saved to /var/cache/conftool/dbconfig/20260721-112453-cwilliams.json * 11:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168', diff saved to https://phabricator.wikimedia.org/P94971 and previous config saved to /var/cache/conftool/dbconfig/20260721-112250-cwilliams.json * 11:21 XioNoX: put eqiad-drmrs Arelion link in service * 11:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169', diff saved to https://phabricator.wikimedia.org/P94970 and previous config saved to /var/cache/conftool/dbconfig/20260721-111446-cwilliams.json * 11:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94969 and previous config saved to /var/cache/conftool/dbconfig/20260721-111242-cwilliams.json * 11:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 11:10 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 11:07 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1093 hosts * 11:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94968 and previous config saved to /var/cache/conftool/dbconfig/20260721-110548-cwilliams.json * 11:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1168.eqiad.wmnet with reason: Maintenance * 11:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94967 and previous config saved to /var/cache/conftool/dbconfig/20260721-110520-cwilliams.json * 11:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94966 and previous config saved to /var/cache/conftool/dbconfig/20260721-110439-cwilliams.json * 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94964 and previous config saved to /var/cache/conftool/dbconfig/20260721-105632-cwilliams.json * 10:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2169.codfw.wmnet with reason: Maintenance * 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94963 and previous config saved to /var/cache/conftool/dbconfig/20260721-105603-cwilliams.json * 10:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165', diff saved to https://phabricator.wikimedia.org/P94962 and previous config saved to /var/cache/conftool/dbconfig/20260721-105512-cwilliams.json * 10:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158', diff saved to https://phabricator.wikimedia.org/P94961 and previous config saved to /var/cache/conftool/dbconfig/20260721-104555-cwilliams.json * 10:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165', diff saved to https://phabricator.wikimedia.org/P94960 and previous config saved to /var/cache/conftool/dbconfig/20260721-104504-cwilliams.json * 10:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158', diff saved to https://phabricator.wikimedia.org/P94959 and previous config saved to /var/cache/conftool/dbconfig/20260721-103547-cwilliams.json * 10:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94958 and previous config saved to /var/cache/conftool/dbconfig/20260721-103456-cwilliams.json * 10:29 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2229: Upgraded kernel * 10:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94956 and previous config saved to /var/cache/conftool/dbconfig/20260721-102757-cwilliams.json * 10:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on an-redacteddb1001.eqiad.wmnet,clouddb[1015,1025,1028].eqiad.wmnet,db1155.eqiad.wmnet with reason: Maintenance * 10:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1165.eqiad.wmnet with reason: Maintenance * 10:25 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94955 and previous config saved to /var/cache/conftool/dbconfig/20260721-102539-cwilliams.json * 10:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94954 and previous config saved to /var/cache/conftool/dbconfig/20260721-101848-cwilliams.json * 10:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2158.codfw.wmnet with reason: Maintenance * 09:43 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2229: Upgraded kernel * 09:42 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2229.codfw.wmnet * 09:42 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2229.codfw.wmnet * 09:23 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db2229.codfw.wmnet * 09:23 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2229.codfw.wmnet * 08:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2229 [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94948 and previous config saved to /var/cache/conftool/dbconfig/20260721-085724-cwilliams.json * 08:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2214 to s6 primary [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94947 and previous config saved to /var/cache/conftool/dbconfig/20260721-085442-cwilliams.json * 08:53 cezmunsta: Starting s6 codfw failover from db2229 to db2214 - [[phab:T430964|T430964]] * 08:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2214 with weight 0 [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94946 and previous config saved to /var/cache/conftool/dbconfig/20260721-084613-cwilliams.json * 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 22 hosts with reason: Primary switchover s6 [[phab:T430964|T430964]] * 08:32 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1017.eqiad.wmnet,service=s1 * 08:08 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Add subrated circuit rate to interface descriptions - CR1312476 - ayounsi@cumin1003 * 08:06 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Add subrated circuit rate to interface descriptions - CR1312476 - ayounsi@cumin1003 * 07:58 reedy@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] (duration: 12m 55s) * 07:51 reedy@deploy2003: reedy, neriah: Continuing with deployment * 07:51 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 07:51 reedy@deploy2003: reedy, neriah: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:48 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1093 hosts * 07:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm2001.wikimedia.org * 07:45 reedy@deploy2003: Started scap sync-world: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] * 07:43 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 07:42 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm2001.wikimedia.org * 07:23 elukey: upgrade libtiff6 packages on zuul* trixie hosts for security upgrades * 07:22 elukey: upgrade libtiff6 packages on Wikikube trixie workers for security upgrades * 07:14 elukey@deploy2003: helmfile [codfw] DONE helmfile.d/services/proton: sync * 07:13 elukey@deploy2003: helmfile [codfw] START helmfile.d/services/proton: sync * 07:11 elukey@deploy2003: helmfile [eqiad] DONE helmfile.d/services/proton: sync * 07:10 elukey@deploy2003: helmfile [eqiad] START helmfile.d/services/proton: sync * 07:09 elukey@deploy2003: helmfile [staging] DONE helmfile.d/services/proton: sync * 07:08 elukey@deploy2003: helmfile [staging] START helmfile.d/services/proton: sync * 06:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1023.eqiad.wmnet with reason: host reimage * 06:46 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1023.eqiad.wmnet with reason: host reimage * 06:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 05:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Haproxy-only mode support - oblivian@cumin1003" * 05:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Haproxy-only mode support - oblivian@cumin1003 * 05:42 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Haproxy-only mode support - oblivian@cumin1003 * 05:42 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Haproxy-only mode support - oblivian@cumin1003" * 05:38 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1017.eqiad.wmnet with reason: Cloning * 05:37 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1017.eqiad.wmnet,service=s1 * 05:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1029.eqiad.wmnet,service=s8 * 05:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1029.eqiad.wmnet,service=s5 * 05:32 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:30 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 05:11 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet * 05:04 aokoth@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet * 05:00 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 04:56 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 04:01 mwpresync@deploy2003: Pruned MediaWiki: 1.47.0-wmf.9 (duration: 01m 08s) * 03:41 mwpresync@deploy2003: Finished scap sync-world: testwikis to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] (duration: 36m 30s) * 03:05 mwpresync@deploy2003: Started scap sync-world: testwikis to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 03:01 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:01 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:00 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:00 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:36 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:36 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:36 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:35 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:16 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 47s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 00:56 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm == 2026-07-20 == * 23:38 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 23:07 Amir1: deleting echo notifications from 2015 on group1 wikis * 23:07 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] (duration: 14m 16s) * 23:01 ladsgroup@deploy2003: ladsgroup: Continuing with deployment * 23:00 ladsgroup@deploy2003: ladsgroup: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:53 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] * 22:46 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2007.codfw.wmnet, repooling source-only afterwards * 22:39 maryum: Deployed security fixes for several security bugs * 21:42 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 21:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2007.codfw.wmnet, repooling source-only afterwards * 21:37 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 21:37 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 21:34 sbassett: Deployed security fix for [[phab:T432424|T432424]] * 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs2020.codfw.wmnet * 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1023.eqiad.wmnet * 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1011.eqiad.wmnet * 21:32 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 17s) * 21:32 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 21:27 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 21:13 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2007.codfw.wmnet with reason: host reimage * 21:08 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-internal-main,name=codfw * 21:06 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2007.codfw.wmnet with reason: host reimage * 20:59 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service * 20:58 sukhe: pybal restart for IP changes around wdqs-main hosts * 20:57 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 20:46 ryankemper@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-internal-main,name=codfw * 20:45 ebernhardson@deploy2003: Finished deploy [search/mjolnir/deploy@d4dc3b8]: Update for opensearch 2.x compat (duration: 00m 34s) * 20:44 ebernhardson@deploy2003: Started deploy [search/mjolnir/deploy@d4dc3b8]: Update for opensearch 2.x compat * 20:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2007 * 20:44 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2007 * 20:43 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2007 * 20:43 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2007.codfw.wmnet 156.16.192.10.in-addr.arpa 6.5.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:42 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2007.codfw.wmnet 156.16.192.10.in-addr.arpa 6.5.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:42 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2007 - bking@cumin2003" * 20:41 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2007 - bking@cumin2003" * 20:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94944 and previous config saved to /var/cache/conftool/dbconfig/20260720-203333-cwilliams.json * 20:32 arlolra@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] (duration: 15m 07s) * 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2020.codfw.wmnet * 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1023.eqiad.wmnet * 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1011.eqiad.wmnet * 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs2020.codfw.wmnet * 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1023.eqiad.wmnet * 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1011.eqiad.wmnet * 20:25 arlolra@deploy2003: arlolra, cscott: Continuing with deployment * 20:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257', diff saved to https://phabricator.wikimedia.org/P94943 and previous config saved to /var/cache/conftool/dbconfig/20260720-202325-cwilliams.json * 20:21 arlolra@deploy2003: arlolra, cscott: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:17 arlolra@deploy2003: Started scap sync-world: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] * 20:13 bking@cumin2003: START - Cookbook sre.dns.netbox * 20:13 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257', diff saved to https://phabricator.wikimedia.org/P94942 and previous config saved to /var/cache/conftool/dbconfig/20260720-201318-cwilliams.json * 20:13 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 20:10 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 20:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2007 * 20:04 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2007.codfw.wmnet with OS bookworm * 20:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94941 and previous config saved to /var/cache/conftool/dbconfig/20260720-200310-cwilliams.json * 19:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94940 and previous config saved to /var/cache/conftool/dbconfig/20260720-195633-cwilliams.json * 19:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1257.eqiad.wmnet with reason: Maintenance * 19:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94939 and previous config saved to /var/cache/conftool/dbconfig/20260720-195605-cwilliams.json * 19:51 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 19:50 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 19:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256', diff saved to https://phabricator.wikimedia.org/P94938 and previous config saved to /var/cache/conftool/dbconfig/20260720-194558-cwilliams.json * 19:44 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 19:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 19:41 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wdqs1011.eqiad.wmnet with OS bookworm * 19:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256', diff saved to https://phabricator.wikimedia.org/P94937 and previous config saved to /var/cache/conftool/dbconfig/20260720-193550-cwilliams.json * 19:25 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94936 and previous config saved to /var/cache/conftool/dbconfig/20260720-192542-cwilliams.json * 19:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94935 and previous config saved to /var/cache/conftool/dbconfig/20260720-191856-cwilliams.json * 19:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1256.eqiad.wmnet with reason: Maintenance * 19:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94934 and previous config saved to /var/cache/conftool/dbconfig/20260720-191839-cwilliams.json * 19:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255', diff saved to https://phabricator.wikimedia.org/P94933 and previous config saved to /var/cache/conftool/dbconfig/20260720-190831-cwilliams.json * 18:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255', diff saved to https://phabricator.wikimedia.org/P94932 and previous config saved to /var/cache/conftool/dbconfig/20260720-185824-cwilliams.json * 18:50 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 18:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94931 and previous config saved to /var/cache/conftool/dbconfig/20260720-184816-cwilliams.json * 18:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94930 and previous config saved to /var/cache/conftool/dbconfig/20260720-184224-cwilliams.json * 18:42 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1255.eqiad.wmnet with reason: Maintenance * 18:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94929 and previous config saved to /var/cache/conftool/dbconfig/20260720-184153-cwilliams.json * 18:39 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 18:39 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 16s) * 18:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 18:38 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 59m 26s) * 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 18:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211', diff saved to https://phabricator.wikimedia.org/P94928 and previous config saved to /var/cache/conftool/dbconfig/20260720-183145-cwilliams.json * 18:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211', diff saved to https://phabricator.wikimedia.org/P94927 and previous config saved to /var/cache/conftool/dbconfig/20260720-182137-cwilliams.json * 18:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94926 and previous config saved to /var/cache/conftool/dbconfig/20260720-181129-cwilliams.json * 18:09 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_codfw * 18:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2057.codfw.wmnet * 18:08 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_codfw * 18:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2058.codfw.wmnet * 18:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94925 and previous config saved to /var/cache/conftool/dbconfig/20260720-180452-cwilliams.json * 18:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on clouddb[1016,1020,1022-1023].eqiad.wmnet,db1154.eqiad.wmnet with reason: Maintenance * 18:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1211.eqiad.wmnet with reason: Maintenance * 18:02 sukhe: armed keyholder on acmechief1002.eqiad.wmnet and acmechief2002.codfw.wmnet (active host) * 18:01 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief2002.codfw.wmnet * 17:57 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief2002.codfw.wmnet * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs2020'] * 17:52 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief1002.eqiad.wmnet * 17:50 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 17:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1011.eqiad.wmnet with reason: host reimage * 17:48 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief1002.eqiad.wmnet * 17:47 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test2001.codfw.wmnet * 17:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94924 and previous config saved to /var/cache/conftool/dbconfig/20260720-174717-cwilliams.json * 17:46 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 17:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1011.eqiad.wmnet with reason: host reimage * 17:43 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs2020.codfw.wmnet with OS bookworm * 17:43 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test2001.codfw.wmnet * 17:43 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test1001.eqiad.wmnet * 17:39 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test1001.eqiad.wmnet * 17:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 17:38 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:38 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 17:37 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 17:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244', diff saved to https://phabricator.wikimedia.org/P94923 and previous config saved to /var/cache/conftool/dbconfig/20260720-173709-cwilliams.json * 17:35 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 17:31 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 20m 40s) * 17:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2055.codfw.wmnet * 17:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2056.codfw.wmnet * 17:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1011 * 17:27 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1011 * 17:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1011.eqiad.wmnet with OS bookworm * 17:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244', diff saved to https://phabricator.wikimedia.org/P94922 and previous config saved to /var/cache/conftool/dbconfig/20260720-172701-cwilliams.json * 17:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94921 and previous config saved to /var/cache/conftool/dbconfig/20260720-171653-cwilliams.json * 17:11 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 17:11 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 13m 03s) * 17:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94920 and previous config saved to /var/cache/conftool/dbconfig/20260720-171012-cwilliams.json * 17:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2244.codfw.wmnet with reason: Maintenance * 17:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94919 and previous config saved to /var/cache/conftool/dbconfig/20260720-170941-cwilliams.json * 16:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243', diff saved to https://phabricator.wikimedia.org/P94918 and previous config saved to /var/cache/conftool/dbconfig/20260720-165933-cwilliams.json * 16:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 16:58 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2053.codfw.wmnet * 16:51 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2054.codfw.wmnet * 16:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243', diff saved to https://phabricator.wikimedia.org/P94917 and previous config saved to /var/cache/conftool/dbconfig/20260720-164926-cwilliams.json * 16:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94916 and previous config saved to /var/cache/conftool/dbconfig/20260720-163918-cwilliams.json * 16:35 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 16:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94915 and previous config saved to /var/cache/conftool/dbconfig/20260720-163140-cwilliams.json * 16:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2243.codfw.wmnet with reason: Maintenance * 16:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94914 and previous config saved to /var/cache/conftool/dbconfig/20260720-163111-cwilliams.json * 16:27 btullis@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 16:27 btullis@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 16:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2020 * 16:23 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2020 * 16:21 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2020 * 16:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2020.codfw.wmnet 85.0.192.10.in-addr.arpa 5.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:21 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2020.codfw.wmnet 85.0.192.10.in-addr.arpa 5.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242', diff saved to https://phabricator.wikimedia.org/P94913 and previous config saved to /var/cache/conftool/dbconfig/20260720-162103-cwilliams.json * 16:19 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 16:18 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 16:18 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:18 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 16:17 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:17 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for netbox accounting errors - jhancock@cumin2002" * 16:17 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for netbox accounting errors - jhancock@cumin2002" * 16:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2051.codfw.wmnet * 16:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2052.codfw.wmnet * 16:11 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 16:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242', diff saved to https://phabricator.wikimedia.org/P94912 and previous config saved to /var/cache/conftool/dbconfig/20260720-161055-cwilliams.json * 16:09 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 16:08 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 16:06 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 16:06 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 16:06 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2020 * 16:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2020.codfw.wmnet with OS bookworm * 16:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94911 and previous config saved to /var/cache/conftool/dbconfig/20260720-160047-cwilliams.json * 15:58 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2019.codfw.wmnet, repooling source-only afterwards * 15:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94909 and previous config saved to /var/cache/conftool/dbconfig/20260720-155353-cwilliams.json * 15:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2242.codfw.wmnet with reason: Maintenance * 15:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94908 and previous config saved to /var/cache/conftool/dbconfig/20260720-154433-cwilliams.json * 15:35 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2049.codfw.wmnet * 15:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162', diff saved to https://phabricator.wikimedia.org/P94907 and previous config saved to /var/cache/conftool/dbconfig/20260720-153425-cwilliams.json * 15:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2050.codfw.wmnet * 15:28 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162', diff saved to https://phabricator.wikimedia.org/P94906 and previous config saved to /var/cache/conftool/dbconfig/20260720-152418-cwilliams.json * 15:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1023 * 15:14 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1023 * 15:14 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 15:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94905 and previous config saved to /var/cache/conftool/dbconfig/20260720-151407-cwilliams.json * 15:13 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] (duration: 41m 16s) * 15:08 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 15:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94902 and previous config saved to /var/cache/conftool/dbconfig/20260720-150729-cwilliams.json * 15:07 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2162.codfw.wmnet with reason: Maintenance * 15:05 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2027.codfw.wmnet, repooling source-only afterwards * 15:00 urbanecm@deploy2003: vadymts1, migr, urbanecm: Continuing with deployment * 14:59 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:58 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 07s) * 14:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:58 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 13s) * 14:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:57 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2019.codfw.wmnet, repooling source-only afterwards * 14:57 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2047.codfw.wmnet * 14:55 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2048.codfw.wmnet * 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2019.codfw.wmnet with OS bookworm * 14:49 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 14:47 urbanecm@deploy2003: vadymts1, migr, urbanecm: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:44 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool magru [reason: BGP issues in lvs7003 resolved after liberica restart, no task ID specified] * 14:44 sukhe@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool magru [reason: BGP issues in lvs7003 resolved after liberica restart, no task ID specified] * 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:41 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:39 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:39 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:33 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool magru [reason: no reason specified, no task ID specified] * 14:33 sukhe@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool magru [reason: no reason specified, no task ID specified] * 14:31 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] * 14:24 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:24 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:24 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:24 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2019.codfw.wmnet with reason: host reimage * 14:22 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2027.codfw.wmnet, repooling source-only afterwards * 14:19 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2019.codfw.wmnet with reason: host reimage * 14:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2027.codfw.wmnet with OS bookworm * 14:16 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2046.codfw.wmnet * 14:16 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2045.codfw.wmnet * 14:08 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:08 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:08 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:08 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:07 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:06 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:06 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:06 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:05 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2019 * 14:00 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2019 * 13:56 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1015.eqiad.wmnet * 13:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2027.codfw.wmnet with reason: host reimage * 13:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2071.codfw.wmnet with OS trixie * 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:51 sukhe@cumin1003: END (ERROR) - Cookbook sre.loadbalancer.admin (exit_code=97) rebooting A:liberica and P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica and P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:51 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1015.eqiad.wmnet * 13:50 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1014.eqiad.wmnet * 13:50 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1076.eqiad.wmnet with OS trixie * 13:50 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2027.codfw.wmnet with reason: host reimage * 13:45 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1014.eqiad.wmnet * 13:44 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1013.eqiad.wmnet * 13:39 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 13:38 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1013.eqiad.wmnet * 13:37 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2044.codfw.wmnet * 13:37 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2043.codfw.wmnet * 13:36 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2019 * 13:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2019.codfw.wmnet 156.32.192.10.in-addr.arpa 6.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:36 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2019.codfw.wmnet 156.32.192.10.in-addr.arpa 6.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:36 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2019 - bking@cumin2003" * 13:36 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2019 - bking@cumin2003" * 13:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry2005.codfw.wmnet * 13:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2071.codfw.wmnet with reason: host reimage * 13:31 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:31 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2019 * 13:31 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry2005.codfw.wmnet * 13:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry2004.codfw.wmnet * 13:30 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2019.codfw.wmnet with OS bookworm * 13:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2027 * 13:30 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2027 * 13:30 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2027.codfw.wmnet with OS bookworm * 13:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1076.eqiad.wmnet with reason: host reimage * 13:29 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_codfw * 13:28 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_codfw * 13:26 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry2004.codfw.wmnet * 13:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry1005.eqiad.wmnet * 13:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2071.codfw.wmnet with reason: host reimage * 13:22 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1076.eqiad.wmnet with reason: host reimage * 13:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry1005.eqiad.wmnet * 13:21 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry1004.eqiad.wmnet * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry1004.eqiad.wmnet * 13:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts * 13:13 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts * 13:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts * 13:12 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts * 13:03 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1076.eqiad.wmnet with OS trixie * 13:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2071.codfw.wmnet with OS trixie * 12:55 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:54 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:53 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:46 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 7 hosts * 12:42 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 7 hosts * 12:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts * 12:42 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts * 12:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2070.codfw.wmnet with OS trixie * 12:36 ozge@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:35 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1075.eqiad.wmnet with OS trixie * 12:32 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts * 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts * 12:22 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin1001.eqiad.wmnet * 12:19 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin1001.eqiad.wmnet * 12:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2070.codfw.wmnet with reason: host reimage * 12:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin2001.codfw.wmnet * 12:14 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1075.eqiad.wmnet with reason: host reimage * 12:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2070.codfw.wmnet with reason: host reimage * 12:10 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1075.eqiad.wmnet with reason: host reimage * 12:09 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin2001.codfw.wmnet * 11:17 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1074.eqiad.wmnet with OS trixie * 11:17 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 11:16 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 11:14 ozge@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 6 hosts * 11:09 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 6 hosts * 11:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 324 hosts * 10:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1074.eqiad.wmnet with reason: host reimage * 10:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2069.codfw.wmnet with OS trixie * 10:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1074.eqiad.wmnet with reason: host reimage * 10:30 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2069.codfw.wmnet with reason: host reimage * 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1074.eqiad.wmnet with OS trixie * 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2069.codfw.wmnet with reason: host reimage * 10:06 blake@deploy2003: Stopping before sync operations * 10:06 blake@deploy2003: Started scap sync-world: Non-deployment scap run to populate new release values * 10:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2069.codfw.wmnet with OS trixie * 10:00 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1073.eqiad.wmnet with OS trixie * 09:56 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 324 hosts * 09:39 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 09:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 8 hosts * 09:38 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1073.eqiad.wmnet with reason: host reimage * 09:37 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 8 hosts * 09:34 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1073.eqiad.wmnet with reason: host reimage * 09:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2068.codfw.wmnet with OS trixie * 09:16 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1073.eqiad.wmnet with OS trixie * 09:13 blake@deploy2003: sync-world aborted: Non-deployment scap run to populate new release values (duration: 00m 02s) * 09:13 blake@deploy2003: Started scap sync-world: Non-deployment scap run to populate new release values * 08:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2068.codfw.wmnet with reason: host reimage * 08:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2068.codfw.wmnet with reason: host reimage * 08:50 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 08:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2068.codfw.wmnet with OS trixie * 08:15 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1072.eqiad.wmnet with OS trixie * 07:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2067.codfw.wmnet with OS trixie * 07:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1072.eqiad.wmnet with reason: host reimage * 07:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1072.eqiad.wmnet with reason: host reimage * 07:45 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 07:45 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 07:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2067.codfw.wmnet with reason: host reimage * 07:35 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2067.codfw.wmnet with reason: host reimage * 07:30 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 07:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1072.eqiad.wmnet with OS trixie * 07:30 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 07:17 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 07:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2067.codfw.wmnet with OS trixie * 05:51 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:50 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:25 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:25 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on db2207.codfw.wmnet with reason: Host down * 04:28 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 07m 02s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-18 == * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 29s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 00:11 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2018.codfw.wmnet, repooling source-only afterwards == 2026-07-17 == * 23:53 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2026.codfw.wmnet, repooling source-only afterwards * 23:09 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2018.codfw.wmnet, repooling source-only afterwards * 23:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2026.codfw.wmnet, repooling source-only afterwards * 22:11 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2018.codfw.wmnet with OS bookworm * 22:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2026.codfw.wmnet with OS bookworm * 21:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2018.codfw.wmnet with reason: host reimage * 21:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2018.codfw.wmnet with reason: host reimage * 21:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2026.codfw.wmnet with reason: host reimage * 21:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2026.codfw.wmnet with reason: host reimage * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2018 * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2018 * 21:26 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2018 * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2018.codfw.wmnet 155.32.192.10.in-addr.arpa 5.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:26 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2018.codfw.wmnet 155.32.192.10.in-addr.arpa 5.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2018 - bking@cumin2003" * 21:26 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2018 - bking@cumin2003" * 21:14 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:13 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2018 * 21:13 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2018.codfw.wmnet with OS bookworm * 21:12 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2026 * 21:12 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2026 * 21:12 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2026.codfw.wmnet with OS bookworm * 21:05 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1022.eqiad.wmnet -> wdqs1026.eqiad.wmnet, repooling source-only afterwards * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs2017.codfw.wmnet, repooling source-only afterwards * 20:11 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs2017.codfw.wmnet, repooling source-only afterwards * 20:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2017.codfw.wmnet with OS bookworm * 20:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1022.eqiad.wmnet -> wdqs1026.eqiad.wmnet, repooling source-only afterwards * 20:06 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1026.eqiad.wmnet with OS bookworm * 19:55 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 19:55 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:55 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 09s) * 19:55 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:50 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 08s) * 19:50 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:50 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 10m 03s) * 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2017.codfw.wmnet with reason: host reimage * 19:40 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:40 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 15s) * 19:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1026.eqiad.wmnet with reason: host reimage * 19:37 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 16s) * 19:37 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:34 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2017.codfw.wmnet with reason: host reimage * 19:34 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1026.eqiad.wmnet with reason: host reimage * 19:33 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 19:33 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:16 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1026.eqiad.wmnet with OS bookworm * 19:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2017 * 19:16 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2017 * 19:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2017.codfw.wmnet with OS bookworm * 18:30 bking@dns1004: END - running authdns-update * 18:28 bking@dns1004: START - running authdns-update * 18:16 kamila@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1264.eqiad.wmnet * 18:16 kamila@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1264.eqiad.wmnet * 18:16 kamila@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1264.eqiad.wmnet * 17:49 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 17:46 dzahn@dns1006: END - running authdns-update * 17:44 dzahn@dns1006: START - running authdns-update * 17:44 dzahn@dns1006: END - running authdns-update * 17:42 dzahn@dns1006: START - running authdns-update * 17:28 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 17:21 kamila@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 17:01 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1264 * 17:01 kamila@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1264 * 17:01 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 17:01 kamila@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1264.eqiad.wmnet * 17:01 kamila@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1264.eqiad.wmnet * 17:01 kamila@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1264.eqiad.wmnet * 16:42 reedy@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] (duration: 10m 29s) * 16:34 reedy@deploy2003: reedy, hartman: Continuing with deployment * 16:33 reedy@deploy2003: reedy, hartman: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:31 reedy@deploy2003: Started scap sync-world: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] * 16:26 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 16:10 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in2001.wikimedia.org with reason: [[phab:T431659|T431659]] * 16:07 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in1001.wikimedia.org with reason: [[phab:T431659|T431659]] * 16:05 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 16:01 kamila@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 16:00 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out2001.wikimedia.org with reason: [[phab:T431659|T431659]] * 15:41 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 15:41 kamila@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 15:35 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out1001.wikimedia.org with reason: [[phab:T431659|T431659]] * 15:14 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1339.eqiad.wmnet * 15:13 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1339.eqiad.wmnet * 15:13 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1339.eqiad.wmnet * 14:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1339.eqiad.wmnet with OS trixie * 14:50 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:49 kamila@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:49 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:33 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage * 14:27 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage * 14:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1339 * 14:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1339 * 14:14 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1339 * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1339.eqiad.wmnet 156.32.64.10.in-addr.arpa 6.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:14 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1339.eqiad.wmnet 156.32.64.10.in-addr.arpa 6.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1339 - cgoubert@cumin2003" * 14:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1339 - cgoubert@cumin2003" * 14:09 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 14:06 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1339 * 14:06 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie * 14:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1339.eqiad.wmnet * 14:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1339.eqiad.wmnet * 14:02 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1339.eqiad.wmnet * 13:45 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb1013.eqiad.wmnet * 13:39 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb1013.eqiad.wmnet * 13:27 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:24 blake@dns1004: END - running authdns-update * 13:22 blake@dns1004: START - running authdns-update * 13:20 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 13:11 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2014.codfw.wmnet * 13:06 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb2014.codfw.wmnet * 13:06 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2012.codfw.wmnet * 13:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 13:03 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 13:01 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 15 hosts * 13:01 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb2012.codfw.wmnet * 13:01 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1016.eqiad.wmnet * 13:00 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 15 hosts * 12:55 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb1016.eqiad.wmnet * 12:55 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1014.eqiad.wmnet * 12:49 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb1014.eqiad.wmnet * 12:32 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:32 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:31 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:31 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1338.eqiad.wmnet * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1338.eqiad.wmnet * 12:18 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1338.eqiad.wmnet * 12:17 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:15 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:14 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:13 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1338.eqiad.wmnet with OS trixie * 12:01 klausman@deploy2003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 11:59 klausman@deploy2003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 11:56 klausman@deploy2003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 11:54 klausman@deploy2003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 11:53 klausman@deploy2003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 11:51 klausman@deploy2003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 11:42 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1338.eqiad.wmnet with reason: host reimage * 11:38 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1338.eqiad.wmnet with reason: host reimage * 11:31 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2230.codfw.wmnet * 11:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1338 * 11:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1338 * 11:25 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1338 * 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1338.eqiad.wmnet 155.32.64.10.in-addr.arpa 5.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:25 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1338.eqiad.wmnet 155.32.64.10.in-addr.arpa 5.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1338 - cgoubert@cumin2003" * 11:25 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1338 - cgoubert@cumin2003" * 11:23 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2230.codfw.wmnet * 11:20 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 11:20 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1338 * 11:20 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1338.eqiad.wmnet with OS trixie * 11:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1338.eqiad.wmnet * 11:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1338.eqiad.wmnet * 11:19 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1338.eqiad.wmnet * 11:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1337.eqiad.wmnet * 11:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1337.eqiad.wmnet * 11:17 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1337.eqiad.wmnet * 11:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1337.eqiad.wmnet with OS trixie * 10:51 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[2001-2002].codfw.wmnet * 10:50 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1337.eqiad.wmnet with reason: host reimage * 10:40 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 10:39 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:39 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1337.eqiad.wmnet with reason: host reimage * 10:39 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1001-1003].eqiad.wmnet * 10:34 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:34 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:30 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:28 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1001-1003].eqiad.wmnet * 10:27 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1337 * 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1337 * 10:26 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1337 * 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1337.eqiad.wmnet 154.32.64.10.in-addr.arpa 4.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:26 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1337.eqiad.wmnet 154.32.64.10.in-addr.arpa 4.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1337 - cgoubert@cumin2003" * 10:26 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1337 - cgoubert@cumin2003" * 10:21 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 10:18 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1337 * 10:17 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1337.eqiad.wmnet with OS trixie * 10:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1337.eqiad.wmnet * 10:16 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db1176.eqiad.wmnet * 10:16 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1337.eqiad.wmnet * 10:16 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1337.eqiad.wmnet * 10:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1336.eqiad.wmnet * 10:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1336.eqiad.wmnet * 10:15 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1336.eqiad.wmnet * 10:11 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db1176.eqiad.wmnet * 10:10 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db1176.eqiad.wmnet * 10:09 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db1176.eqiad.wmnet * 10:05 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts (check the cookbook's logs for more details.) * 10:03 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts (check the cookbook's logs for more details.) * 09:58 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1336.eqiad.wmnet with OS trixie * 09:47 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts (check the cookbook's logs for more details.) * 09:47 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts (check the cookbook's logs for more details.) * 09:45 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host acmechief-test2001.codfw.wmnet,acmechief-test1001.eqiad.wmnet,an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet,db-test[2001-2002].codfw.wmnet,db-test[1 * 09:40 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host acmechief-test2001.codfw.wmnet,acmechief-test1001.eqiad.wmnet,an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet,db-test[2001-2002].codfw.wmnet,db-test[1001-1003].eqiad.wmn * 09:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1336.eqiad.wmnet with reason: host reimage * 09:33 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1336.eqiad.wmnet with reason: host reimage * 09:29 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet * 09:29 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet * 09:28 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:26 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:21 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 09:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1336 * 09:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1336 * 09:19 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 09:14 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1336 * 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1336.eqiad.wmnet 152.32.64.10.in-addr.arpa 2.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1336.eqiad.wmnet 152.32.64.10.in-addr.arpa 2.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1336 - cgoubert@cumin2003" * 09:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1336 - cgoubert@cumin2003" * 09:11 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:10 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 09:09 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:09 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1336 * 09:09 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1336.eqiad.wmnet with OS trixie * 09:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1336.eqiad.wmnet * 09:08 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1336.eqiad.wmnet * 09:08 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1336.eqiad.wmnet * 09:06 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1335.eqiad.wmnet * 09:06 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1335.eqiad.wmnet * 09:06 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1335.eqiad.wmnet * 09:04 elukey: uploaded spicerack_13.1.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia * 08:55 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wikikube-worker-exp2001.codfw.wmnet * 08:54 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host testreduce1002.eqiad.wmnet * 08:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1335.eqiad.wmnet with OS trixie * 08:51 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host wikikube-worker-exp2001.codfw.wmnet * 08:51 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wikikube-worker-exp1001.eqiad.wmnet * 08:50 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host testreduce1002.eqiad.wmnet * 08:45 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host wikikube-worker-exp1001.eqiad.wmnet * 08:34 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1335.eqiad.wmnet with reason: host reimage * 08:30 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1335.eqiad.wmnet with reason: host reimage * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1335 * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1335 * 08:18 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1335 * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1335.eqiad.wmnet 150.32.64.10.in-addr.arpa 0.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:18 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1335.eqiad.wmnet 150.32.64.10.in-addr.arpa 0.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1335 - cgoubert@cumin2003" * 08:18 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1335 - cgoubert@cumin2003" * 08:14 elukey@cumin1003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 08:14 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges * 08:13 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 08:10 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1335 * 08:10 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1335.eqiad.wmnet with OS trixie * 08:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1335.eqiad.wmnet * 08:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1335.eqiad.wmnet * 08:09 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1335.eqiad.wmnet * 08:06 elukey@cumin1003: END (FAIL) - Cookbook sre.puppet.disable-merges (exit_code=99) * 08:05 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges * 08:03 elukey@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin1003.eqiad.wmnet * 07:57 elukey@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin1003.eqiad.wmnet * 07:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetdb1003.eqiad.wmnet * 07:46 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetdb1003.eqiad.wmnet * 07:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetdb2003.codfw.wmnet * 07:37 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetdb2003.codfw.wmnet * 07:37 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1001.eqiad.wmnet * 07:28 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver1001.eqiad.wmnet * 07:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet * 07:19 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet * 07:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2002.codfw.wmnet * 07:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver2002.codfw.wmnet * 07:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet * 07:05 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet * 07:04 btullis@cumin1003: END (FAIL) - Cookbook sre.hadoop.reboot-workers (exit_code=99) for Hadoop analytics cluster * 07:04 elukey@cumin1003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 07:04 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges * 06:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox1003.eqiad.wmnet * 06:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox1003.eqiad.wmnet * 02:46 ryankemper@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:46 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:44 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:37 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:37 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-internal-scholarly,name=eqiad * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 49s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 01:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore wdqs1025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling source-only afterwards * 01:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore wdqs1027 after Bookworm reimage) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs1027.eqiad.wmnet, repooling both afterwards * 00:55 urbanecm@deploy2003: helmfile [codfw] DONE helmfile.d/services/linkrecommendation: apply * 00:54 urbanecm@deploy2003: helmfile [eqiad] DONE helmfile.d/services/linkrecommendation: apply * 00:54 urbanecm@deploy2003: helmfile [staging] DONE helmfile.d/services/linkrecommendation: apply * 00:54 urbanecm@deploy2003: helmfile [codfw] START helmfile.d/services/linkrecommendation: apply * 00:53 urbanecm@deploy2003: helmfile [staging] START helmfile.d/services/linkrecommendation: apply * 00:52 urbanecm@deploy2003: helmfile [eqiad] START helmfile.d/services/linkrecommendation: apply * 00:23 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore wdqs1025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling source-only afterwards * 00:23 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore wdqs1027 after Bookworm reimage) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs1027.eqiad.wmnet, repooling both afterwards * 00:14 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1274.eqiad.wmnet * 00:14 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1274.eqiad.wmnet * 00:14 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1274.eqiad.wmnet * 00:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1027.eqiad.wmnet with OS bookworm * 00:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1025.eqiad.wmnet with OS bookworm * 00:04 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1274.eqiad.wmnet with OS trixie == 2026-07-16 == * 23:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], xfer to freshly reimaged/scap-deployed wdqs2025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs2025.codfw.wmnet, repooling source-only afterwards * 23:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1027.eqiad.wmnet with reason: host reimage * 23:47 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1025.eqiad.wmnet with reason: host reimage * 23:43 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1274.eqiad.wmnet with reason: host reimage * 23:41 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1025.eqiad.wmnet with reason: host reimage * 23:39 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1027.eqiad.wmnet with reason: host reimage * 23:38 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1274.eqiad.wmnet with reason: host reimage * 23:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1025 * 23:23 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1025 * 23:22 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1027 * 23:22 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1027 * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1274 * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1274 * 23:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1025.eqiad.wmnet with OS bookworm * 23:19 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1274 * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1274.eqiad.wmnet 145.48.64.10.in-addr.arpa 5.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:19 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1274.eqiad.wmnet 145.48.64.10.in-addr.arpa 5.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1274 - swfrench@cumin1003" * 23:19 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1274 - swfrench@cumin1003" * 23:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1027.eqiad.wmnet with OS bookworm * 23:14 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 23:14 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1274 * 23:13 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1274.eqiad.wmnet with OS trixie * 23:13 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1274.eqiad.wmnet * 23:12 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1274.eqiad.wmnet * 23:12 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1274.eqiad.wmnet * 23:12 ryankemper: [[phab:T430880|T430880]] depooled dnsdisc of wdqs-internal-scholarly-eqiad bc we only have 1 host there * 23:09 ryankemper@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-internal-scholarly,name=eqiad * 23:08 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1272.eqiad.wmnet * 23:08 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1272.eqiad.wmnet * 23:08 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1272.eqiad.wmnet * 23:01 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], xfer to freshly reimaged/scap-deployed wdqs2025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs2025.codfw.wmnet, repooling source-only afterwards * 22:57 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1272.eqiad.wmnet with OS trixie * 22:56 Amir1: deleting echo notifications from 2015 in group0 * 22:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2025.codfw.wmnet with OS bookworm * 22:35 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1272.eqiad.wmnet with reason: host reimage * 22:32 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 27s) * 22:32 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 22:28 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1269.eqiad.wmnet * 22:28 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1269.eqiad.wmnet * 22:28 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1269.eqiad.wmnet * 22:27 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1272.eqiad.wmnet with reason: host reimage * 22:26 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] (duration: 08m 51s) * 22:22 ladsgroup@deploy2003: ladsgroup, urbanecm: Continuing with deployment * 22:19 ladsgroup@deploy2003: ladsgroup, urbanecm: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:17 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] * 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2025.codfw.wmnet with reason: host reimage * 22:06 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1272 * 22:06 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1272 * 22:05 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1272 * 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1272.eqiad.wmnet 127.48.64.10.in-addr.arpa 7.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:05 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1272.eqiad.wmnet 127.48.64.10.in-addr.arpa 7.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1272 - swfrench@cumin1003" * 22:05 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1272 - swfrench@cumin1003" * 22:03 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2025.codfw.wmnet with reason: host reimage * 22:01 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 22:00 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1272 * 22:00 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1272.eqiad.wmnet with OS trixie * 22:00 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1272.eqiad.wmnet * 21:59 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1272.eqiad.wmnet * 21:59 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1272.eqiad.wmnet * 21:55 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1271.eqiad.wmnet * 21:55 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1271.eqiad.wmnet * 21:55 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1271.eqiad.wmnet * 21:46 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1271.eqiad.wmnet with OS trixie * 21:45 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] (duration: 06m 31s) * 21:40 sbassett@deploy2003: sbassett: Continuing with deployment * 21:40 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2025 * 21:40 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2025 * 21:40 sbassett@deploy2003: sbassett: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:38 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] * 21:37 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2025 * 21:37 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2025.codfw.wmnet 220.48.192.10.in-addr.arpa 0.2.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:37 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2025.codfw.wmnet 220.48.192.10.in-addr.arpa 0.2.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:37 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:37 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2025 - bking@cumin2003" * 21:37 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2025 - bking@cumin2003" * 21:30 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] (duration: 08m 19s) * 21:26 sbassett@deploy2003: sbassett: Continuing with deployment * 21:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1269.eqiad.wmnet with OS trixie * 21:24 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1271.eqiad.wmnet with reason: host reimage * 21:23 sbassett@deploy2003: sbassett: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:22 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:22 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] * 21:20 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2025 * 21:19 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2025.codfw.wmnet with OS bookworm * 21:17 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1271.eqiad.wmnet with reason: host reimage * 21:04 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1269.eqiad.wmnet with reason: host reimage * 21:00 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1269.eqiad.wmnet with reason: host reimage * 20:56 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1271 * 20:55 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1271 * 20:54 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1271 * 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1271.eqiad.wmnet 126.48.64.10.in-addr.arpa 6.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:54 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1271.eqiad.wmnet 126.48.64.10.in-addr.arpa 6.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1271 - swfrench@cumin1003" * 20:54 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1271 - swfrench@cumin1003" * 20:51 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:51 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:51 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:50 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 20:49 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 20:49 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1268.eqiad.wmnet * 20:49 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1268.eqiad.wmnet * 20:49 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1268.eqiad.wmnet * 20:48 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1271 * 20:48 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1271.eqiad.wmnet with OS trixie * 20:47 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1271.eqiad.wmnet * 20:46 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1271.eqiad.wmnet * 20:46 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1271.eqiad.wmnet * 20:41 aude@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] (duration: 07m 34s) * 20:39 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1269 * 20:39 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1269 * 20:36 aude@deploy2003: aude: Continuing with deployment * 20:35 aude@deploy2003: aude: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:33 aude@deploy2003: Started scap sync-world: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] * 20:26 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on db2207.codfw.wmnet with reason: Host down * 20:22 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-video: apply * 20:21 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-video: apply * 20:20 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-timeline: apply * 20:20 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-timeline: apply * 20:20 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-syntaxhighlight: apply * 20:19 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-syntaxhighlight: apply * 20:19 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-media: apply * 20:18 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-media: apply * 20:18 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-constraints: apply * 20:17 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-constraints: apply * 20:17 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox: apply * 20:16 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox: apply * 20:13 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1269 * 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1269.eqiad.wmnet 80.32.64.10.in-addr.arpa 0.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:13 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1269.eqiad.wmnet 80.32.64.10.in-addr.arpa 0.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1269 - kamila@cumin1003" * 20:13 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1269 - kamila@cumin1003" * 20:09 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-flink-codfw cluster: Roll restart of jvm daemons. * 20:07 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 20:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 20:03 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-flink-codfw cluster: Roll restart of jvm daemons. * 20:03 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2207 [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94893 and previous config saved to /var/cache/conftool/dbconfig/20260716-200257-marostegui.json * 20:01 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2204 to s2 primary [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94892 and previous config saved to /var/cache/conftool/dbconfig/20260716-200157-marostegui.json * 20:00 marostegui: Starting emergency s2 codfw failover from db2207 to db2204 - [[phab:T432396|T432396]] * 19:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1035.eqiad.wmnet * 19:56 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2204 with weight 0 [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94891 and previous config saved to /var/cache/conftool/dbconfig/20260716-195628-marostegui.json * 19:55 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 26 hosts with reason: Primary switchover s2 [[phab:T432396|T432396]] * 19:54 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1035.eqiad.wmnet * 19:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1034.eqiad.wmnet * 19:48 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1034.eqiad.wmnet * 19:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1033.eqiad.wmnet * 19:43 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-video: apply * 19:43 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1033.eqiad.wmnet * 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1032.eqiad.wmnet * 19:42 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-video: apply * 19:42 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-timeline: apply * 19:41 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-timeline: apply * 19:41 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-syntaxhighlight: apply * 19:41 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-syntaxhighlight: apply * 19:40 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-media: apply * 19:40 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-media: apply * 19:39 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-constraints: apply * 19:36 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-constraints: apply * 19:36 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox: apply * 19:35 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1032.eqiad.wmnet * 19:35 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1031.eqiad.wmnet * 19:35 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox: apply * 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-video: apply * 19:33 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-video: apply * 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-timeline: apply * 19:33 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-timeline: apply * 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-syntaxhighlight: apply * 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-syntaxhighlight: apply * 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-media: apply * 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-media: apply * 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-constraints: apply * 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-constraints: apply * 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox: apply * 19:31 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox: apply * 19:27 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1031.eqiad.wmnet * 19:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1030.eqiad.wmnet * 19:23 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1001.eqiad.wmnet, repooling source-only afterwards * 19:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1030.eqiad.wmnet * 19:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1029.eqiad.wmnet * 19:17 kamila@cumin1003: START - Cookbook sre.dns.netbox * 19:12 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1029.eqiad.wmnet * 19:06 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1269 * 19:05 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1269.eqiad.wmnet with OS trixie * 19:03 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1269.eqiad.wmnet * 19:03 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1269.eqiad.wmnet * 19:03 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1269.eqiad.wmnet * 18:55 dancy@deploy2003: Finished scap sync-world: testing [[phab:T428971|T428971]] (duration: 02m 41s) * 18:53 dancy@deploy2003: Started scap sync-world: testing [[phab:T428971|T428971]] * 18:31 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1268.eqiad.wmnet with OS trixie * 18:18 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 18:16 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1267.eqiad.wmnet * 18:16 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1267.eqiad.wmnet * 18:16 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1267.eqiad.wmnet * 18:09 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1268.eqiad.wmnet with reason: host reimage * 18:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1001.eqiad.wmnet, repooling source-only afterwards * 18:06 swfrench@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] (duration: 07m 34s) * 18:06 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 23s) * 18:06 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 18:06 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1268.eqiad.wmnet with reason: host reimage * 18:03 bd808@deploy2003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 18:02 bd808@deploy2003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 18:02 swfrench@deploy2003: jiji, swfrench: Continuing with deployment * 18:02 bd808@deploy2003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 18:02 bd808@deploy2003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 18:01 bd808@deploy2003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 18:01 swfrench@deploy2003: jiji, swfrench: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:01 bd808@deploy2003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:59 swfrench@deploy2003: Started scap sync-world: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] * 17:45 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1268 * 17:45 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1268 * 17:44 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1267.eqiad.wmnet with OS trixie * 17:43 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter2006.codfw.wmnet * 17:39 swfrench@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter2006.codfw.wmnet * 17:35 swfrench@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] (duration: 07m 27s) * 17:34 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1268 * 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1268.eqiad.wmnet 78.32.64.10.in-addr.arpa 8.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:34 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1268.eqiad.wmnet 78.32.64.10.in-addr.arpa 8.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1268 - kamila@cumin1003" * 17:34 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1268 - kamila@cumin1003" * 17:31 swfrench@deploy2003: jiji, swfrench: Continuing with deployment * 17:29 swfrench@deploy2003: jiji, swfrench: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:28 kamila@cumin1003: START - Cookbook sre.dns.netbox * 17:28 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1268 * 17:28 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1268.eqiad.wmnet with OS trixie * 17:27 swfrench@deploy2003: Started scap sync-world: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] * 17:23 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1267.eqiad.wmnet with reason: host reimage * 17:18 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1267.eqiad.wmnet with reason: host reimage * 17:18 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1268.eqiad.wmnet * 17:17 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1268.eqiad.wmnet * 17:17 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1268.eqiad.wmnet * 17:12 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter2005.codfw.wmnet * 17:11 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1270.eqiad.wmnet * 17:11 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1270.eqiad.wmnet * 17:11 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1270.eqiad.wmnet * 17:09 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter2005.codfw.wmnet * 17:08 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:08 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update reverse dns for moved arelion cct cr2-eqiad - cmooney@cumin1003" * 17:08 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update reverse dns for moved arelion cct cr2-eqiad - cmooney@cumin1003" * 17:08 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] (duration: 07m 34s) * 17:04 jiji@deploy2003: jiji: Continuing with deployment * 17:03 jiji@deploy2003: jiji: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 17:00 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] * 17:00 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2185.codfw.wmnet with OS trixie * 16:59 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:58 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1270.eqiad.wmnet with OS trixie * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1267 * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1267 * 16:57 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1267 * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1267.eqiad.wmnet 77.32.64.10.in-addr.arpa 7.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:57 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1267.eqiad.wmnet 77.32.64.10.in-addr.arpa 7.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1267 - kamila@cumin1003" * 16:56 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1267 - kamila@cumin1003" * 16:56 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_eqsin * 16:56 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5032.eqsin.wmnet * 16:52 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_esams * 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3073.esams.wmnet * 16:50 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_esams * 16:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3081.esams.wmnet * 16:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1266.eqiad.wmnet * 16:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1266.eqiad.wmnet * 16:45 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1266.eqiad.wmnet * 16:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2185.codfw.wmnet with reason: host reimage * 16:41 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_eqiad * 16:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1114.eqiad.wmnet * 16:41 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_eqiad * 16:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1115.eqiad.wmnet * 16:39 kamila@cumin1003: START - Cookbook sre.dns.netbox * 16:39 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1267 * 16:39 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2185.codfw.wmnet with reason: host reimage * 16:38 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1267.eqiad.wmnet with OS trixie * 16:38 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1267.eqiad.wmnet * 16:38 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1270.eqiad.wmnet with reason: host reimage * 16:37 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1267.eqiad.wmnet * 16:37 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1267.eqiad.wmnet * 16:31 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1270.eqiad.wmnet with reason: host reimage * 16:24 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1264.eqiad.wmnet * 16:24 kamila@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 16:24 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 16:23 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 16:21 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2162: switch maintenance completed codfw rack b6 * 16:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2185.codfw.wmnet with OS trixie * 16:19 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 16:16 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter1007.eqiad.wmnet * 16:15 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_eqsin * 16:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5024.eqsin.wmnet * 16:13 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5031.eqsin.wmnet * 16:13 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3072.esams.wmnet * 16:12 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter1007.eqiad.wmnet * 16:11 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] (duration: 09m 47s) * 16:10 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1266.eqiad.wmnet with OS trixie * 16:10 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1270 * 16:10 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1270 * 16:09 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1270 * 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1270.eqiad.wmnet 125.48.64.10.in-addr.arpa 5.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1270.eqiad.wmnet 125.48.64.10.in-addr.arpa 5.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1270 - swfrench@cumin1003" * 16:09 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1270 - swfrench@cumin1003" * 16:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3080.esams.wmnet * 16:07 jiji@deploy2003: jiji: Continuing with deployment * 16:06 jiji@deploy2003: jiji: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:04 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 16:04 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1265.eqiad.wmnet * 16:03 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1265.eqiad.wmnet * 16:03 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1265.eqiad.wmnet * 16:03 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1270 * 16:03 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1270.eqiad.wmnet with OS trixie * 16:02 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1270.eqiad.wmnet * 16:02 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] * 16:01 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1270.eqiad.wmnet * 16:01 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1270.eqiad.wmnet * 16:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1113.eqiad.wmnet * 16:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1112.eqiad.wmnet * 15:49 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1266.eqiad.wmnet with reason: host reimage * 15:47 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter1006.eqiad.wmnet * 15:45 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1265.eqiad.wmnet with OS trixie * 15:44 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1266.eqiad.wmnet with reason: host reimage * 15:43 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter1006.eqiad.wmnet * 15:42 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] (duration: 09m 46s) * 15:37 jiji@deploy2003: jiji: Continuing with deployment * 15:36 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2162: switch maintenance completed codfw rack b6 * 15:36 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2161: switch maintenance completed codfw rack b6 * 15:34 jiji@deploy2003: jiji: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:32 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5023.eqsin.wmnet * 15:32 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] * 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5030.eqsin.wmnet * 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3071.esams.wmnet * 15:27 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3079.esams.wmnet * 15:25 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1265.eqiad.wmnet with reason: host reimage * 15:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1266 * 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1266 * 15:21 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1110.eqiad.wmnet * 15:20 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1111.eqiad.wmnet * 15:16 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1265.eqiad.wmnet with reason: host reimage * 15:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1001.eqiad.wmnet with OS bookworm * 15:15 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1266 * 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1266.eqiad.wmnet 76.32.64.10.in-addr.arpa 6.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:15 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1266.eqiad.wmnet 76.32.64.10.in-addr.arpa 6.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1266 - kamila@cumin1003" * 15:15 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1266 - kamila@cumin1003" * 15:07 kamila@cumin1003: START - Cookbook sre.dns.netbox * 15:04 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1266 * 15:04 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1264 * 15:04 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1264 * 15:04 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1266.eqiad.wmnet with OS trixie * 15:03 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1264 * 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1264.eqiad.wmnet 74.32.64.10.in-addr.arpa 4.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:03 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1264.eqiad.wmnet 74.32.64.10.in-addr.arpa 4.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1264 - kamila@cumin1003" * 15:03 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1264 - kamila@cumin1003" * 15:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-eqiad * 15:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp1001.eqiad.wmnet * 15:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp1001.eqiad.wmnet * 15:01 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp1001.eqiad.wmnet * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp1001.eqiad.wmnet * 15:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1376-1384].eqiad.wmnet * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1376-1384].eqiad.wmnet * 14:59 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1002.eqiad.wmnet * 14:59 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1266.eqiad.wmnet * 14:58 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1266.eqiad.wmnet * 14:58 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1266.eqiad.wmnet * 14:58 kamila@cumin1003: START - Cookbook sre.dns.netbox * 14:57 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1264 * 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1265 * 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1265 * 14:57 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1310596{{!}}Set $wgMathInternalRestbaseURL explicitly (T349582)]] * 14:57 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1265 * 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1265.eqiad.wmnet 75.32.64.10.in-addr.arpa 5.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:56 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1265.eqiad.wmnet 75.32.64.10.in-addr.arpa 5.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:56 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:56 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1265 - kamila@cumin1003" * 14:56 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1265 - kamila@cumin1003" * 14:53 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1002.eqiad.wmnet * 14:53 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1376-1384].eqiad.wmnet * 14:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1001.eqiad.wmnet with reason: host reimage * 14:51 kamila@cumin1003: START - Cookbook sre.dns.netbox * 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 14:50 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:50 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2161: switch maintenance completed codfw rack b6 * 14:50 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5021.eqsin.wmnet * 14:50 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1265 * 14:49 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:49 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:49 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1265.eqiad.wmnet with OS trixie * 14:49 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5029.eqsin.wmnet * 14:49 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1265.eqiad.wmnet * 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3070.esams.wmnet * 14:48 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1264.eqiad.wmnet * 14:48 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1001.eqiad.wmnet with reason: host reimage * 14:48 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1376-1384].eqiad.wmnet * 14:48 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1265.eqiad.wmnet * 14:47 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1265.eqiad.wmnet * 14:47 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1264.eqiad.wmnet * 14:47 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1264.eqiad.wmnet * 14:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:47 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3078.esams.wmnet * 14:44 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1263.eqiad.wmnet * 14:44 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1263.eqiad.wmnet * 14:44 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1263.eqiad.wmnet * 14:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1108.eqiad.wmnet * 14:40 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1109.eqiad.wmnet * 14:40 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:35 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:34 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:34 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2006.codfw.wmnet * 14:34 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-flink-eqiad cluster: Roll restart of jvm daemons. * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf2002.codfw.wmnet * 14:31 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1002.eqiad.wmnet * 14:29 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2006.codfw.wmnet * 14:27 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:27 kamila@deploy2003: Finished scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] (duration: 02m 57s) * 14:27 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-flink-eqiad cluster: Roll restart of jvm daemons. * 14:26 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf2002.codfw.wmnet * 14:26 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf2001.codfw.wmnet * 14:25 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1002.eqiad.wmnet * 14:25 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1001.eqiad.wmnet * 14:25 kamila@deploy2003: Started scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] * 14:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:21 kamila@deploy2003: sync-world aborted: Test deployment to check rsync is working - [[phab:T432108|T432108]] (duration: 00m 36s) * 14:21 topranks: reboot lsw1-b6-codfw to upgrade JunOS [[phab:T430922|T430922]] * 14:21 kamila@deploy2003: Started scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] * 14:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1001.eqiad.wmnet with OS bookworm * 14:20 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b6-codfw,lsw1-b6-codfw IPv6,lsw1-b6-codfw.mgmt,ssw1-a[1,8]-codfw with reason: lsw1-b6-codfw JunOS upgrade * 14:20 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf2001.codfw.wmnet * 14:19 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1001.eqiad.wmnet * 14:19 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 26 hosts with reason: lsw1-b6-codfw JunOS upgrade * 14:14 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 14:13 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc2022: switch maintenance codfw rack b6 * 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:12 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.parsercache * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool pc2022: switch maintenance codfw rack b6 * 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2251: switch maintenance codfw rack b6 * 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.parsercache * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2251: switch maintenance codfw rack b6 * 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2162: switch maintenance codfw rack b6 * 14:12 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1263.eqiad.wmnet with OS trixie * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2162: switch maintenance codfw rack b6 * 14:11 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2161: switch maintenance codfw rack b6 * 14:11 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2161: switch maintenance codfw rack b6 * 14:08 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5020.eqsin.wmnet * 14:07 btullis@cumin1003: START - Cookbook sre.hadoop.reboot-workers for Hadoop analytics cluster * 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3069.esams.wmnet * 14:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5028.eqsin.wmnet * 14:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1338-1347].eqiad.wmnet * 14:06 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1338-1347].eqiad.wmnet * 14:05 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3077.esams.wmnet * 14:02 topranks: beginning depools for lsw1-b6-codfw maintenance [[phab:T430922|T430922]] * 14:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1106.eqiad.wmnet * 14:00 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-misc1002.eqiad.wmnet * 13:59 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1338-1347].eqiad.wmnet * 13:59 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1107.eqiad.wmnet * 13:56 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-codfw * 13:55 sfaci@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply * 13:54 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-misc1002.eqiad.wmnet * 13:54 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-misc1001.eqiad.wmnet * 13:54 sfaci@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply * 13:50 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1263.eqiad.wmnet with reason: host reimage * 13:50 sfaci@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 13:49 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1338-1347].eqiad.wmnet * 13:49 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-misc1001.eqiad.wmnet * 13:49 sfaci@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 13:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:49 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:45 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1263.eqiad.wmnet with reason: host reimage * 13:40 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:40 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-eqiad * 13:35 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:34 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:33 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-reboot (exit_code=0) rolling reboot on A:dnsbox and (A:eqsin or A:drmrs or A:magru) and not (P<nowiki>{</nowiki>dns5003*<nowiki>}</nowiki> or P<nowiki>{</nowiki>dns7002*<nowiki>}</nowiki>) and (A:dnsbox) * 13:33 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns7001.wikimedia.org * 13:27 sukhe@dns1004: END - running authdns-update * 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5019.eqsin.wmnet * 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3076.esams.wmnet * 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3068.esams.wmnet * 13:25 sukhe@dns1004: START - running authdns-update * 13:24 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5027.eqsin.wmnet * 13:24 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1263 * 13:24 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1263 * 13:23 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1263 * 13:23 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:23 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1104.eqiad.wmnet * 13:21 kamila@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:21 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:20 kamila@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:20 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:20 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:20 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1263 - kamila@cumin1003" * 13:20 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1263 - kamila@cumin1003" * 13:19 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1105.eqiad.wmnet * 13:19 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:19 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:18 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:18 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns7001.wikimedia.org * 13:16 cdobbins@cumin2003: conftool action : set/pooled=yes; selector: name=dns7002.* * 13:14 cdobbins@dns1004: END - running authdns-update * 13:13 sbisson@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] (duration: 08m 03s) * 13:13 cdobbins@dns1004: START - running authdns-update * 13:12 kamila@cumin1003: START - Cookbook sre.dns.netbox * 13:12 cdobbins@cumin2003: conftool action : set/pooled=yes; selector: name=dns7002.*,service=authdns-update * 13:12 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1263 * 13:11 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1263.eqiad.wmnet with OS trixie * 13:11 cdobbins@cumin2003: conftool action : set/pooled=no; selector: name=dns7002.* * 13:11 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1263.eqiad.wmnet * 13:10 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1263.eqiad.wmnet * 13:10 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1263.eqiad.wmnet * 13:09 sbisson@deploy2003: sbisson: Continuing with deployment * 13:08 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:07 sbisson@deploy2003: sbisson: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:05 sbisson@deploy2003: Started scap sync-world: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] * 13:03 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns6002.wikimedia.org * 13:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1298-1307].eqiad.wmnet * 13:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1298-1307].eqiad.wmnet * 12:59 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1262.eqiad.wmnet * 12:59 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1262.eqiad.wmnet * 12:59 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1262.eqiad.wmnet * 12:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1298-1307].eqiad.wmnet * 12:49 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns6002.wikimedia.org * 12:46 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1298-1307].eqiad.wmnet * 12:46 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:46 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3075.esams.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3067.esams.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5018.eqsin.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5026.eqsin.wmnet * 12:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1102.eqiad.wmnet * 12:39 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1103.eqiad.wmnet * 12:35 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:34 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns6001.wikimedia.org * 12:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:28 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:18 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns6001.wikimedia.org * 12:14 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1267-1276].eqiad.wmnet * 12:13 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1267-1276].eqiad.wmnet * 12:04 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1267-1276].eqiad.wmnet * 12:03 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns5004.wikimedia.org * 12:02 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1100.eqiad.wmnet * 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3066.esams.wmnet * 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3074.esams.wmnet * 12:01 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:01 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-codfw * 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5017.eqsin.wmnet * 12:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5025.eqsin.wmnet * 12:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1101.eqiad.wmnet * 11:59 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1267-1276].eqiad.wmnet * 11:58 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:58 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:54 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns5004.wikimedia.org * 11:54 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and (A:eqsin or A:drmrs or A:magru) and not (P<nowiki>{</nowiki>dns5003*<nowiki>}</nowiki> or P<nowiki>{</nowiki>dns7002*<nowiki>}</nowiki>) and (A:dnsbox) * 11:54 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:53 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:53 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-eqiad * 11:51 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:51 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:50 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:50 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_eqiad * 11:50 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:50 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_eqiad * 11:49 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_esams * 11:49 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_esams * 11:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_eqsin * 11:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_eqsin * 11:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:44 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-eqiad * 11:43 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-codfw * 11:42 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:41 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:24 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-eqiad * 11:23 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-codfw * 11:22 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:15 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:14 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:09 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2066.codfw.wmnet with OS trixie * 11:05 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1151-1160].eqiad.wmnet * 11:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1151-1160].eqiad.wmnet * 10:59 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1068.eqiad.wmnet with OS trixie * 10:55 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.major-upgrade (exit_code=99) * 10:55 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 10:54 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1151-1160].eqiad.wmnet * 10:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2066.codfw.wmnet with reason: host reimage * 10:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1151-1160].eqiad.wmnet * 10:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:42 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2066.codfw.wmnet with reason: host reimage * 10:39 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:37 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:36 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:23 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:22 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2066.codfw.wmnet with OS trixie * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:07 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2065.codfw.wmnet with OS trixie * 10:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:06 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 10:06 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 10:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 10:03 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:03 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 10:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 09:59 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 09:57 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 09:57 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 09:52 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 09:47 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:46 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:46 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2065.codfw.wmnet with reason: host reimage * 09:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2065.codfw.wmnet with reason: host reimage * 09:40 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:39 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 09:39 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:39 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:39 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:37 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:29 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox2003.codfw.wmnet * 09:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox2003.codfw.wmnet * 09:25 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:25 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:24 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:24 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:24 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:21 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:20 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2065.codfw.wmnet with OS trixie * 09:13 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1068.eqiad.wmnet with OS trixie * 09:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2064.codfw.wmnet with OS trixie * 09:08 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 09:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 09:07 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2162: Repooling after switchover * 09:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1067.eqiad.wmnet with OS trixie * 09:00 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:59 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 08:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 08:57 tappof: bump space for prometheus k8s-dse in eqiad * 08:56 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping2004.codfw.wmnet * 08:52 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host ping2004.codfw.wmnet * 08:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 08:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping1004.eqiad.wmnet * 08:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:51 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:49 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 08:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host ping1004.eqiad.wmnet * 08:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2064.codfw.wmnet with reason: host reimage * 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2064.codfw.wmnet with reason: host reimage * 08:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1067.eqiad.wmnet with reason: host reimage * 08:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:33 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1067.eqiad.wmnet with reason: host reimage * 08:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:21 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2162: Repooling after switchover * 08:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2064.codfw.wmnet with OS trixie * 08:16 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1067.eqiad.wmnet with OS trixie * 08:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:15 cgoubert@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-eqiad * 08:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2062.codfw.wmnet with OS trixie * 08:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1066.eqiad.wmnet with OS trixie * 08:02 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2162: Repooling after switchover * 07:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2162: Repooling after switchover * 07:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2162 [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94870 and previous config saved to /var/cache/conftool/dbconfig/20260716-075530-cwilliams.json * 07:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2241 to x3 primary [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94869 and previous config saved to /var/cache/conftool/dbconfig/20260716-075314-cwilliams.json * 07:52 cezmunsta: Starting x3 codfw failover from db2162 to db2241 - [[phab:T430925|T430925]] * 07:50 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:50 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:47 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 07:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2241 with weight 0 [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94868 and previous config saved to /var/cache/conftool/dbconfig/20260716-074507-cwilliams.json * 07:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 18 hosts with reason: Primary switchover x3 [[phab:T430925|T430925]] * 07:43 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1066.eqiad.wmnet with reason: host reimage * 07:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:dse-k8s-worker-eqiad * 07:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1028.eqiad.wmnet * 07:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1028.eqiad.wmnet * 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 07:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1066.eqiad.wmnet with reason: host reimage * 07:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1028.eqiad.wmnet * 07:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1028.eqiad.wmnet * 07:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1027.eqiad.wmnet * 07:35 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1027.eqiad.wmnet * 07:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1027.eqiad.wmnet * 07:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1027.eqiad.wmnet * 07:28 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1026.eqiad.wmnet * 07:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1026.eqiad.wmnet * 07:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast2003.wikimedia.org * 07:21 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1026.eqiad.wmnet * 07:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1066.eqiad.wmnet with OS trixie * 07:19 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast2003.wikimedia.org * 07:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2062.codfw.wmnet with OS trixie * 06:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1026.eqiad.wmnet * 06:51 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1025.eqiad.wmnet * 06:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1025.eqiad.wmnet * 06:47 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 06:47 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 06:44 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1025.eqiad.wmnet * 06:14 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1025.eqiad.wmnet * 06:14 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1024.eqiad.wmnet * 06:14 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1024.eqiad.wmnet * 06:07 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1024.eqiad.wmnet * 05:37 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1024.eqiad.wmnet * 05:37 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1023.eqiad.wmnet * 05:37 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1023.eqiad.wmnet * 05:26 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1023.eqiad.wmnet * 04:56 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1023.eqiad.wmnet * 04:56 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1022.eqiad.wmnet * 04:56 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1022.eqiad.wmnet * 04:49 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1022.eqiad.wmnet * 04:19 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1022.eqiad.wmnet * 04:19 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1021.eqiad.wmnet * 04:19 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1021.eqiad.wmnet * 04:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1021.eqiad.wmnet * 03:38 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1021.eqiad.wmnet * 03:38 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1020.eqiad.wmnet * 03:38 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1020.eqiad.wmnet * 03:20 btullis@cumin1003: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1020.eqiad.wmnet * 03:18 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1020.eqiad.wmnet * 03:18 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1019.eqiad.wmnet * 03:18 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1019.eqiad.wmnet * 03:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1019.eqiad.wmnet * 02:41 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1019.eqiad.wmnet * 02:41 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1018.eqiad.wmnet * 02:41 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1018.eqiad.wmnet * 02:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling both afterwards * 02:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2003.codfw.wmnet -> wcqs2001.codfw.wmnet, repooling both afterwards * 02:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1018.eqiad.wmnet * 02:30 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1018.eqiad.wmnet * 02:30 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1014.eqiad.wmnet * 02:30 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1014.eqiad.wmnet * 02:24 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1014.eqiad.wmnet * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 01:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1014.eqiad.wmnet * 01:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1013.eqiad.wmnet * 01:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1013.eqiad.wmnet * 01:47 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1013.eqiad.wmnet * 01:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2003.codfw.wmnet -> wcqs2001.codfw.wmnet, repooling both afterwards * 01:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling both afterwards * 01:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1013.eqiad.wmnet * 01:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1012.eqiad.wmnet * 01:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1012.eqiad.wmnet * 01:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1012.eqiad.wmnet * 01:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1012.eqiad.wmnet * 01:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1011.eqiad.wmnet * 01:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1011.eqiad.wmnet * 01:08 ryankemper@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] scap deploy post bookworm reimage (duration: 00m 23s) * 01:08 ryankemper@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] scap deploy post bookworm reimage * 01:08 ryankemper@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): scap deploy post bookworm reimage (duration: 00m 46s) * 01:07 ryankemper@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): scap deploy post bookworm reimage * 01:04 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1011.eqiad.wmnet * 01:04 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1011.eqiad.wmnet * 01:04 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1010.eqiad.wmnet * 01:04 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1010.eqiad.wmnet * 00:57 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1010.eqiad.wmnet * 00:57 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1010.eqiad.wmnet * 00:57 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1009.eqiad.wmnet * 00:57 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1009.eqiad.wmnet * 00:50 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1009.eqiad.wmnet * 00:20 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1009.eqiad.wmnet * 00:20 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1008.eqiad.wmnet * 00:20 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1008.eqiad.wmnet * 00:13 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1008.eqiad.wmnet == 2026-07-15 == * 23:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2001.codfw.wmnet with OS bookworm * 23:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1008.eqiad.wmnet * 23:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1007.eqiad.wmnet * 23:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1007.eqiad.wmnet * 23:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1007.eqiad.wmnet * 23:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1007.eqiad.wmnet * 23:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1006.eqiad.wmnet * 23:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1006.eqiad.wmnet * 23:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1006.eqiad.wmnet * 23:29 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1006.eqiad.wmnet * 23:28 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1005.eqiad.wmnet * 23:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1005.eqiad.wmnet * 23:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1002.eqiad.wmnet with OS bookworm * 23:21 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1005.eqiad.wmnet * 23:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2001.codfw.wmnet with reason: host reimage * 23:15 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host datahubsearch1001.eqiad.wmnet with OS bookworm * 23:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2001.codfw.wmnet with reason: host reimage * 23:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 23:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 22:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 22:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1005.eqiad.wmnet * 22:51 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1004.eqiad.wmnet * 22:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1004.eqiad.wmnet * 22:45 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1004.eqiad.wmnet * 22:44 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 22:44 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS trixie * 22:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host datahubsearch1001.eqiad.wmnet with OS bookworm * 22:34 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host datahubsearch1001.eqiad.wmnet with OS bookworm * 22:16 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on datahubsearch[1002-1003].eqiad.wmnet with reason: Using datahubsearch1001 to test bookworm reimages * 22:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1004.eqiad.wmnet * 22:15 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1003.eqiad.wmnet * 22:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1003.eqiad.wmnet * 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 22:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1003.eqiad.wmnet * 22:08 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1003.eqiad.wmnet * 22:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1002.eqiad.wmnet * 22:08 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1002.eqiad.wmnet * 22:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host datahubsearch1001.eqiad.wmnet with OS bookworm * 22:05 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 22:02 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm * 22:01 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on datahubsearch[1001-1003].eqiad.wmnet with reason: Using datahubsearch1001 to test bookworm reimages * 22:01 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1002.eqiad.wmnet * 22:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1002.eqiad.wmnet * 22:00 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1001.eqiad.wmnet * 22:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1001.eqiad.wmnet * 21:53 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1001.eqiad.wmnet * 21:52 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 21:50 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 21:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS trixie * 21:50 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS bookworm * 21:43 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 21:38 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:30 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wcqs1002'] * 21:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:29 lerickson@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 21:29 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:29 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:29 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS bookworm * 21:28 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 21:28 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm * 21:23 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1001.eqiad.wmnet * 21:23 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:23 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:22 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 21:20 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:18 swfrench-wmf: reprepro include php8.3_8.3.32-1+wmf11u2 into component/php83 for bullseye-wikimedia * 21:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:16 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:15 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-druid-public cluster: Roll restart of jvm daemons. * 21:08 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:05 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1001.eqiad.wmnet * 21:05 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1001.eqiad.wmnet * 21:04 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-druid-public cluster: Roll restart of jvm daemons. * 21:02 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 21:01 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 21:01 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 21:00 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 20:59 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1001.eqiad.wmnet * 20:59 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1001.eqiad.wmnet * 20:59 btullis@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:dse-k8s-worker-eqiad * 20:55 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 20:55 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 20:45 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm * 20:21 jhathaway: puppet is re-enabled, have fun, but not too much fun! * 20:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 20:17 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs2001'] * 20:12 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs2001'] * 20:11 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs2001'] * 20:09 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:08 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 20:05 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:05 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 20:04 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs2001'] * 20:03 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 20:03 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm * 20:02 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:02 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 20:01 jhathaway: disabling puppet fleet wide to roll out kafka patch * 19:55 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host relforge1010.eqiad.wmnet * 19:52 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 19:52 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 19:48 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 19:48 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 19:48 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 19:47 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 19:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 19:45 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 19:45 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 19:44 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1010.eqiad.wmnet * 19:38 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1262.eqiad.wmnet with OS trixie * 19:17 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1262.eqiad.wmnet with reason: host reimage * 19:11 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1262.eqiad.wmnet with reason: host reimage * 18:59 cdobbins@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS trixie * 18:54 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 18:53 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 18:52 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1262 * 18:52 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1262 * 18:51 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1262 * 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1262.eqiad.wmnet 72.32.64.10.in-addr.arpa 2.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:51 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1262.eqiad.wmnet 72.32.64.10.in-addr.arpa 2.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1262 - kamila@cumin1003" * 18:51 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1262 - kamila@cumin1003" * 18:46 kamila@cumin1003: START - Cookbook sre.dns.netbox * 18:46 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1262 * 18:46 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 18:46 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ncmonitor1001.eqiad.wmnet * 18:46 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 18:45 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1262.eqiad.wmnet with OS trixie * 18:45 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 18:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1262.eqiad.wmnet * 18:44 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1262.eqiad.wmnet * 18:44 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1262.eqiad.wmnet * 18:42 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host ncmonitor1001.eqiad.wmnet * 18:29 topranks: pull power on cr1-eqiad to install new switch-control boards [[phab:T426343|T426343]] * 18:29 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs[1018-1020].eqiad.wmnet with reason: line card install in cr1-eqiad * 18:27 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 14 hosts with reason: linecard install in cr1-eqad * 18:22 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_ulsfo * 18:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4052.ulsfo.wmnet * 18:19 cdobbins@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 18:15 cdobbins@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 18:14 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_drmrs * 18:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6016.drmrs.wmnet * 18:12 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_ulsfo * 18:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4044.ulsfo.wmnet * 18:10 sukhe@cumin1003: END (ERROR) - Cookbook sre.cdn.roll-reboot (exit_code=97) rolling reboot on A:cp-upload_drmrs * 18:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2241: Security update * 17:56 topranks: start draining traffic on cr1-eqiad ahead of line card installation [[phab:T426343|T426343]] * 17:47 cdobbins@cumin2003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie * 17:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4051.ulsfo.wmnet * 17:40 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:39 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 17:34 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6007.drmrs.wmnet * 17:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6015.drmrs.wmnet * 17:32 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:31 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 17:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4043.ulsfo.wmnet * 17:27 lerickson@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:25 lerickson@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 17:22 lerickson@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-codfw * 17:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp2001.codfw.wmnet * 17:22 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 17:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp2001.codfw.wmnet * 17:22 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 17:19 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2241: Security update * 17:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2241.codfw.wmnet * 17:17 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2241.codfw.wmnet * 17:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp2001.codfw.wmnet * 17:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp2001.codfw.wmnet * 17:15 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2366-2374].codfw.wmnet * 17:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2366-2374].codfw.wmnet * 17:10 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply * 17:10 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply * 17:08 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2366-2374].codfw.wmnet * 17:06 sukhe: sre.dns.roll-reboot to resume later * 17:06 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-reboot (exit_code=97) rolling reboot on A:dnsbox and not (A:ulsfo or A:magru) and (A:dnsbox) * 17:06 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns5003.wikimedia.org * 17:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2241: Security update * 17:03 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2241: Security update * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply * 17:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2366-2374].codfw.wmnet * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply * 17:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2357-2365].codfw.wmnet * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 17:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2357-2365].codfw.wmnet * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply * 16:55 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2357-2365].codfw.wmnet * 16:55 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 16:53 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6006.drmrs.wmnet * 16:52 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6014.drmrs.wmnet * 16:52 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 16:51 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 16:50 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2357-2365].codfw.wmnet * 16:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4042.ulsfo.wmnet * 16:50 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2347-2356].codfw.wmnet * 16:50 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2347-2356].codfw.wmnet * 16:49 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns5003.wikimedia.org * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply * 16:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4050.ulsfo.wmnet * 16:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2347-2356].codfw.wmnet * 16:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2347-2356].codfw.wmnet * 16:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2337-2346].codfw.wmnet * 16:36 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2337-2346].codfw.wmnet * 16:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:dse-k8s-worker-codfw * 16:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2003.codfw.wmnet * 16:35 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2003.codfw.wmnet * 16:34 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns3004.wikimedia.org * 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply * 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply * 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply * 16:30 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply * 16:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2003.codfw.wmnet * 16:29 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2337-2346].codfw.wmnet * 16:24 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2003.codfw.wmnet * 16:24 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2002.codfw.wmnet * 16:24 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2002.codfw.wmnet * 16:23 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns3004.wikimedia.org * 16:23 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2337-2346].codfw.wmnet * 16:23 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2327-2336].codfw.wmnet * 16:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2327-2336].codfw.wmnet * 16:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2002.codfw.wmnet * 16:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2327-2336].codfw.wmnet * 16:12 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2002.codfw.wmnet * 16:12 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2001.codfw.wmnet * 16:12 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2001.codfw.wmnet * 16:12 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1065.eqiad.wmnet with OS trixie * 16:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6005.drmrs.wmnet * 16:11 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6013.drmrs.wmnet * 16:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4041.ulsfo.wmnet * 16:08 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns3003.wikimedia.org * 16:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2327-2336].codfw.wmnet * 16:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2317-2326].codfw.wmnet * 16:06 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2317-2326].codfw.wmnet * 16:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2001.codfw.wmnet * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply * 16:03 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4049.ulsfo.wmnet * 16:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2001.codfw.wmnet * 16:00 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test2001.codfw.wmnet * 16:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test2001.codfw.wmnet * 16:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2063.codfw.wmnet with OS trixie * 15:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2317-2326].codfw.wmnet * 15:57 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns3003.wikimedia.org * 15:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test2001.codfw.wmnet * 15:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test2001.codfw.wmnet * 15:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2004.codfw.wmnet * 15:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2004.codfw.wmnet * 15:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2317-2326].codfw.wmnet * 15:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2307-2316].codfw.wmnet * 15:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2307-2316].codfw.wmnet * 15:49 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2004.codfw.wmnet * 15:48 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2004.codfw.wmnet * 15:48 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2003.codfw.wmnet * 15:48 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2003.codfw.wmnet * 15:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 15:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2307-2316].codfw.wmnet * 15:42 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2003.codfw.wmnet * 15:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 15:42 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2003.codfw.wmnet * 15:42 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2002.codfw.wmnet * 15:42 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2002.codfw.wmnet * 15:42 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2006.wikimedia.org * 15:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2063.codfw.wmnet with reason: host reimage * 15:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2307-2316].codfw.wmnet * 15:37 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2297-2306].codfw.wmnet * 15:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2297-2306].codfw.wmnet * 15:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2002.codfw.wmnet * 15:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2002.codfw.wmnet * 15:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2001.codfw.wmnet * 15:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2001.codfw.wmnet * 15:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2063.codfw.wmnet with reason: host reimage * 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6004.drmrs.wmnet * 15:31 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2001.codfw.wmnet * 15:31 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2001.codfw.wmnet * 15:31 btullis@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:dse-k8s-worker-codfw * 15:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6012.drmrs.wmnet * 15:28 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2006.wikimedia.org * 15:27 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-analytics cluster: Roll restart of jvm daemons. * 15:27 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2297-2306].codfw.wmnet * 15:27 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4040.ulsfo.wmnet * 15:24 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1065.eqiad.wmnet with OS trixie * 15:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4048.ulsfo.wmnet * 15:21 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-analytics cluster: Roll restart of jvm daemons. * 15:21 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2297-2306].codfw.wmnet * 15:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2287-2296].codfw.wmnet * 15:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2287-2296].codfw.wmnet * 15:20 btullis@cumin1003: END (PASS) - Cookbook sre.druid.reboot-workers (exit_code=0) for Druid public cluster: Reboot Druid nodes * 15:18 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 15:17 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm * 15:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2063.codfw.wmnet with OS trixie * 15:13 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2005.wikimedia.org * 15:11 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2287-2296].codfw.wmnet * 15:11 btullis@cumin1003: END (PASS) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=0) rolling reboot on A:cephosd-eqiad * 15:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1064.eqiad.wmnet with OS trixie * 15:05 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2062.codfw.wmnet with OS trixie * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2287-2296].codfw.wmnet * 15:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2277-2286].codfw.wmnet * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2277-2286].codfw.wmnet * 14:59 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2005.wikimedia.org * 14:57 brouberol@cumin1003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-jumbo-eqiad * 14:52 btullis@cumin1003: END (PASS) - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas (exit_code=0) rolling reboot on A:schema-codfw * 14:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6003.drmrs.wmnet * 14:50 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:50 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host relforge1009.eqiad.wmnet * 14:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2277-2286].codfw.wmnet * 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6011.drmrs.wmnet * 14:47 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 14:46 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>ml-serve1001.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 14:46 klausman@cumin1003: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) pool for host ml-serve1001.eqiad.wmnet * 14:46 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 14:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1001.eqiad.wmnet * 14:45 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4039.ulsfo.wmnet * 14:44 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1009.eqiad.wmnet * 14:44 btullis@cumin1003: START - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas rolling reboot on A:schema-codfw * 14:44 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2004.wikimedia.org * 14:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2277-2286].codfw.wmnet * 14:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2267-2276].codfw.wmnet * 14:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2267-2276].codfw.wmnet * 14:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 14:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4047.ulsfo.wmnet * 14:40 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1001.eqiad.wmnet * 14:38 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 14:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 14:36 topranks: disconnect power on cr2-eqiad to shut down device for switch fabric replacement [[phab:T426343|T426343]] * 14:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2267-2276].codfw.wmnet * 14:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 14:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1001.eqiad.wmnet * 14:35 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>ml-serve1001.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 14:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 14:34 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 14:33 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:33 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:30 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2004.wikimedia.org * 14:29 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2267-2276].codfw.wmnet * 14:29 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2257-2266].codfw.wmnet * 14:29 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2257-2266].codfw.wmnet * 14:24 btullis@cumin1003: END (PASS) - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas (exit_code=0) rolling reboot on A:schema-eqiad * 14:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2257-2266].codfw.wmnet * 14:20 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:20 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:19 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:17 jforrester@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2257-2266].codfw.wmnet * 14:16 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:16 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2062.codfw.wmnet with OS trixie * 14:15 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1064.eqiad.wmnet with OS trixie * 14:15 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1006.wikimedia.org * 14:15 btullis@cumin1003: START - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas rolling reboot on A:schema-eqiad * 14:14 topranks: switch routing-engine on cr2-eqiad resetting all interfaces [[phab:T417873|T417873]] * 14:11 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:11 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:10 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm * 14:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6002.drmrs.wmnet * 14:09 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6010.drmrs.wmnet * 14:06 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1006.wikimedia.org * 14:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:05 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 14:05 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4038.ulsfo.wmnet * 14:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4046.ulsfo.wmnet * 14:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:00 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on cr1-eqiad with reason: switch upgrade and line card install * 14:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:59 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:57 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:57 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:55 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-eqiad * 13:55 btullis@cumin1003: START - Cookbook sre.druid.reboot-workers for Druid public cluster: Reboot Druid nodes * 13:53 topranks: switch routing-engine on cr2-eqiad resetting all interfaces [[phab:T417873|T417873]] * 13:51 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1005.wikimedia.org * 13:50 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:49 brouberol@cumin1003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-test-eqiad * 13:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:44 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2197-2206].codfw.wmnet * 13:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2197-2206].codfw.wmnet * 13:36 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1005.wikimedia.org * 13:35 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2197-2206].codfw.wmnet * 13:30 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2197-2206].codfw.wmnet * 13:28 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6001.drmrs.wmnet * 13:28 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6009.drmrs.wmnet * 13:28 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2187-2196].codfw.wmnet * 13:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2187-2196].codfw.wmnet * 13:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 13:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4037.ulsfo.wmnet * 13:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2001 * 13:22 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2001 * 13:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4045.ulsfo.wmnet * 13:21 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1004.wikimedia.org * 13:19 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on lvs[1018-1020].eqiad.wmnet with reason: switch upgrade and line card install * 13:18 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2009.codfw.wmnet * 13:18 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2001 * 13:18 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2001.codfw.wmnet 26.16.192.10.in-addr.arpa 6.2.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:17 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2001.codfw.wmnet 26.16.192.10.in-addr.arpa 6.2.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:17 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:17 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2001 - bking@cumin2003" * 13:17 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2001 - bking@cumin2003" * 13:17 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2009.codfw.wmnet * 13:17 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_drmrs * 13:17 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2187-2196].codfw.wmnet * 13:17 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_drmrs * 13:17 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on 15 hosts with reason: switch upgrade and line card install * 13:17 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:15 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 13:13 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:13 brouberol@cumin1003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-jumbo-eqiad * 13:13 brouberol@cumin1003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-test-eqiad * 13:13 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1004.wikimedia.org * 13:13 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and not (A:ulsfo or A:magru) and (A:dnsbox) * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:12 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_ulsfo * 13:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:12 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_ulsfo * 13:11 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2187-2196].codfw.wmnet * 13:11 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 13:11 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 13:06 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 13:05 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling source-only afterwards * 13:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2001 * 13:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2009.codfw.wmnet with OS trixie * 13:03 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:03 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling source-only afterwards * 13:01 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 15s) * 13:01 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 13:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 12:57 btullis@cumin1003: END (PASS) - Cookbook sre.druid.reboot-workers (exit_code=0) for Druid analytics cluster: Reboot Druid nodes * 12:54 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 12:54 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2163-2172].codfw.wmnet * 12:54 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2163-2172].codfw.wmnet * 12:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2163-2172].codfw.wmnet * 12:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2009.codfw.wmnet with reason: host reimage * 12:41 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2163-2172].codfw.wmnet * 12:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2153-2162].codfw.wmnet * 12:40 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2153-2162].codfw.wmnet * 12:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2009.codfw.wmnet with reason: host reimage * 12:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2153-2162].codfw.wmnet * 12:29 btullis@cumin1003: END (PASS) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=0) rolling reboot on A:cephosd-codfw * 12:25 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2153-2162].codfw.wmnet * 12:25 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2143-2152].codfw.wmnet * 12:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2143-2152].codfw.wmnet * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2009 * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2009 * 12:22 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2009 * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2009.codfw.wmnet 139.0.192.10.in-addr.arpa 9.3.1.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:22 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2009.codfw.wmnet 139.0.192.10.in-addr.arpa 9.3.1.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2009 - mvernon@cumin2003" * 12:22 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2009 - mvernon@cumin2003" * 12:16 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 12:15 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 12:15 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 12:15 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2009 * 12:15 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 12:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2009.codfw.wmnet with OS trixie * 12:15 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 12:14 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 12:14 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2143-2152].codfw.wmnet * 12:13 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 12:12 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2010.codfw.wmnet * 12:11 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2010.codfw.wmnet * 12:10 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 12:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2143-2152].codfw.wmnet * 12:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2133-2142].codfw.wmnet * 12:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2133-2142].codfw.wmnet * 12:02 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 11:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2133-2142].codfw.wmnet * 11:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2133-2142].codfw.wmnet * 11:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:49 mvolz@deploy2003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:49 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-codfw * 11:48 mvolz@deploy2003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:47 btullis@cumin1003: START - Cookbook sre.druid.reboot-workers for Druid analytics cluster: Reboot Druid nodes * 11:46 mvolz@deploy2003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:46 mvolz@deploy2003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:45 mvolz@deploy2003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:44 mvolz@deploy2003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:40 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] (duration: 11m 38s) * 11:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2010.codfw.wmnet with OS trixie * 11:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2105-2114].codfw.wmnet * 11:36 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2105-2114].codfw.wmnet * 11:36 krinkle@deploy2003: physikerwelt, krinkle: Continuing with deployment * 11:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1018: Security updates * 11:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:36 root@cumin1003: START - Cookbook sre.mysql.parsercache * 11:36 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1018: Security updates * 11:31 krinkle@deploy2003: physikerwelt, krinkle: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:29 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] * 11:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2105-2114].codfw.wmnet * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2105-2114].codfw.wmnet * 11:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2010.codfw.wmnet with reason: host reimage * 11:12 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2010.codfw.wmnet with reason: host reimage * 11:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1018: Security updates * 11:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:10 root@cumin1003: START - Cookbook sre.mysql.parsercache * 11:10 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1018: Security updates * 11:09 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1009.eqiad.wmnet with OS trixie * 11:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow7002.magru.wmnet * 11:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 11:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 11:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=tegola-vector-tiles,name=eqiad * 11:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=kartotherian,name=eqiad * 11:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow7002.magru.wmnet * 10:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2010 * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2010 * 10:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 10:54 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 10:54 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2010 * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2010.codfw.wmnet 76.16.192.10.in-addr.arpa 6.7.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:54 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2010.codfw.wmnet 76.16.192.10.in-addr.arpa 6.7.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2010 - mvernon@cumin2003" * 10:54 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2010 - mvernon@cumin2003" * 10:49 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 10:49 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2010 * 10:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1009.eqiad.wmnet with reason: host reimage * 10:49 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2010.codfw.wmnet with OS trixie * 10:46 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2011.codfw.wmnet * 10:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1011.eqiad.wmnet * 10:44 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2011.codfw.wmnet * 10:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 10:44 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 10:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1009.eqiad.wmnet with reason: host reimage * 10:44 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow6001.drmrs.wmnet * 10:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1017: Security updates * 10:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:39 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1017: Security updates * 10:39 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow6001.drmrs.wmnet * 10:38 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1011.eqiad.wmnet * 10:35 cgoubert@deploy2003: Finished deploy [restbase/deploy@06301bd]: Deploying {{Gerrit|1306088}} {{Gerrit|1308347}} - [[phab:T429944|T429944]] [[phab:T428279|T428279]] (duration: 28m 34s) * 10:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1012.eqiad.wmnet * 10:35 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 10:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow5003.eqsin.wmnet * 10:34 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2011.codfw.wmnet with OS trixie * 10:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1009.eqiad.wmnet with OS trixie * 10:28 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1012.eqiad.wmnet * 10:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1013.eqiad.wmnet * 10:27 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:27 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow5003.eqsin.wmnet * 10:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:26 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow4003.ulsfo.wmnet * 10:25 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 10:25 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 10:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow4003.ulsfo.wmnet * 10:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1013.eqiad.wmnet * 10:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1014.eqiad.wmnet * 10:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:15 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2011.codfw.wmnet with reason: host reimage * 10:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1017: Security updates * 10:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:14 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:14 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1017: Security updates * 10:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow3004.esams.wmnet * 10:11 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2011.codfw.wmnet with reason: host reimage * 10:10 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:10 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1014.eqiad.wmnet * 10:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki2003.codfw.wmnet * 10:09 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 10:09 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 10:09 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow3004.esams.wmnet * 10:08 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2004.codfw.wmnet * 10:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1010.eqiad.wmnet with OS trixie * 10:07 cgoubert@deploy2003: Started deploy [restbase/deploy@06301bd]: Deploying {{Gerrit|1306088}} {{Gerrit|1308347}} - [[phab:T429944|T429944]] [[phab:T428279|T428279]] * 10:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host rpki2003.codfw.wmnet * 10:04 topranks: push out config change to BGP_outfilter on core routers [[phab:T431849|T431849]] * 10:02 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow2004.codfw.wmnet * 09:59 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 09:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2003.codfw.wmnet * 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2011 * 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2011 * 09:53 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 09:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:52 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2011 * 09:52 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2011.codfw.wmnet 36.32.192.10.in-addr.arpa 6.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:52 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2011.codfw.wmnet 36.32.192.10.in-addr.arpa 6.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:51 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:51 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2011 - mvernon@cumin2003" * 09:51 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2011 - mvernon@cumin2003" * 09:51 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow2003.codfw.wmnet * 09:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1003.eqiad.wmnet * 09:49 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 09:49 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 09:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1010.eqiad.wmnet with reason: host reimage * 09:47 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 09:47 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 09:47 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 09:47 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2011 * 09:46 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2011.codfw.wmnet with OS trixie * 09:44 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow1003.eqiad.wmnet * 09:44 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2012.codfw.wmnet * 09:44 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1002.eqiad.wmnet * 09:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1010.eqiad.wmnet with reason: host reimage * 09:43 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2012.codfw.wmnet * 09:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Security updates * 09:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:43 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:43 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Security updates * 09:42 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:40 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow1002.eqiad.wmnet * 09:40 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 09:37 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki1001.eqiad.wmnet * 09:36 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 09:36 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 09:33 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host rpki1001.eqiad.wmnet * 09:32 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:32 cgoubert@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-codfw * 09:31 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=kartotherian,name=eqiad * 09:31 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola-vector-tiles,name=eqiad * 09:31 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 09:31 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2012.codfw.wmnet with OS trixie * 09:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1010.eqiad.wmnet with OS trixie * 09:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1011.eqiad.wmnet with OS trixie * 09:21 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Security updates * 09:21 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:21 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:21 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Security updates * 09:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2012.codfw.wmnet with reason: host reimage * 09:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1011.eqiad.wmnet with reason: host reimage * 09:08 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2012.codfw.wmnet with reason: host reimage * 09:05 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1011.eqiad.wmnet with reason: host reimage * 08:55 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:52 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1011.eqiad.wmnet with OS trixie * 08:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1022: Security updates * 08:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2012 * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2012 * 08:50 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1022: Security updates * 08:50 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2012 * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2012.codfw.wmnet 44.48.192.10.in-addr.arpa 4.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:50 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2012.codfw.wmnet 44.48.192.10.in-addr.arpa 4.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2012 - mvernon@cumin2003" * 08:50 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2012 - mvernon@cumin2003" * 08:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1012.eqiad.wmnet with OS trixie * 08:44 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 08:44 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2012 * 08:43 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2012.codfw.wmnet with OS trixie * 08:42 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2013.codfw.wmnet * 08:41 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2013.codfw.wmnet * 08:35 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 08:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host krb1002.eqiad.wmnet * 08:30 elukey@dns1004: END - running authdns-update * 08:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1012.eqiad.wmnet with reason: host reimage * 08:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Security updates * 08:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:28 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:28 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Security updates * 08:27 elukey@dns1004: START - running authdns-update * 08:26 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 08:26 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host krb1002.eqiad.wmnet * 08:22 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1012.eqiad.wmnet with reason: host reimage * 08:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host krb2002.codfw.wmnet * 08:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast6003.wikimedia.org * 08:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2013.codfw.wmnet with OS trixie * 08:13 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast6003.wikimedia.org * 08:12 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast3007.wikimedia.org * 08:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host krb2002.codfw.wmnet * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Security updates * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:09 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:09 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Security updates * 08:07 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1012.eqiad.wmnet with OS trixie * 08:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast3007.wikimedia.org * 08:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast5005.wikimedia.org * 07:58 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast5005.wikimedia.org * 07:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1013.eqiad.wmnet with OS trixie * 07:53 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2013.codfw.wmnet with reason: host reimage * 07:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1021: Security updates * 07:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:53 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:53 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1021: Security updates * 07:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast1004.wikimedia.org * 07:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2013.codfw.wmnet with reason: host reimage * 07:46 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast1004.wikimedia.org * 07:40 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1013.eqiad.wmnet with reason: host reimage * 07:36 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1013.eqiad.wmnet with reason: host reimage * 07:31 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2013 * 07:31 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2013 * 07:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1021: Security updates * 07:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:30 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:30 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1021: Security updates * 07:24 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2013 * 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2013.codfw.wmnet 87.0.192.10.in-addr.arpa 7.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:24 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2013.codfw.wmnet 87.0.192.10.in-addr.arpa 7.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2013 - mvernon@cumin2003" * 07:24 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2013 - mvernon@cumin2003" * 07:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1013.eqiad.wmnet with OS trixie * 07:19 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 07:19 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2013 * 07:19 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2013.codfw.wmnet with OS trixie * 07:13 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] (duration: 07m 48s) * 07:09 kharlan@deploy2003: kharlan: Continuing with deployment * 07:08 kharlan@deploy2003: kharlan: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:06 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 01:15 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 01:14 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply == 2026-07-14 == * 22:51 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_magru * 22:51 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7016.magru.wmnet * 22:46 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_magru * 22:46 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7008.magru.wmnet * 22:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7015.magru.wmnet * 22:04 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7007.magru.wmnet * 21:29 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7014.magru.wmnet * 21:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7006.magru.wmnet * 21:13 dzahn@cumin2002: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 0:15:00 on gerrit.wikimedia.org with reason: reboot * 21:11 mutante: gerrit2003 (gerrit.wikimedia.org) - reboot for maintenance * 21:11 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on gerrit2003.wikimedia.org with reason: reboot * 20:56 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:56 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:56 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:55 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 20:48 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7013.magru.wmnet * 20:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7005.magru.wmnet * 20:41 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host phab1005.eqiad.wmnet with OS trixie * 20:28 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] (duration: 06m 47s) * 20:24 sbassett@deploy2003: sbassett: Continuing with deployment * 20:23 sbassett@deploy2003: sbassett: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:23 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on phab1005.eqiad.wmnet with reason: host reimage * 20:21 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] * 20:20 aokoth@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on phab1005.eqiad.wmnet with reason: host reimage * 20:12 jhuneidi@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] (duration: 07m 42s) * 20:07 jhuneidi@deploy2003: jhuneidi, priyankar22: Continuing with deployment * 20:06 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7012.magru.wmnet * 20:06 jhuneidi@deploy2003: jhuneidi, priyankar22: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:04 jhuneidi@deploy2003: Started scap sync-world: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] * 20:02 aokoth@cumin1003: START - Cookbook sre.hosts.reimage for host phab1005.eqiad.wmnet with OS trixie * 20:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7004.magru.wmnet * 20:00 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet * 19:57 aokoth@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet * 19:24 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7011.magru.wmnet * 19:19 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7003.magru.wmnet * 19:11 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] (duration: 08m 33s) * 19:07 jforrester@deploy2003: jforrester: Continuing with deployment * 19:04 jforrester@deploy2003: jforrester: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:02 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] * 18:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7010.magru.wmnet * 18:38 mutante: rotating phabricator-gerrit bot token (its-phabricator) * 18:18 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 17:44 swfrench@deploy2003: Finished scap sync-world: Deployment to pick up new production image (duration: 31m 44s) * 17:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7002.magru.wmnet * 17:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7009.magru.wmnet * 17:32 swfrench@deploy2003: swfrench: Continuing with deployment * 17:29 swfrench@deploy2003: swfrench: Deployment to pick up new production image synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:17 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2035: repooling after rack b5 maintenance * 17:16 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool es2035: repooling after rack b5 maintenance * 17:16 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2188: repooling after rack b5 maintenance * 17:12 swfrench@deploy2003: Started scap sync-world: Deployment to pick up new production image * 17:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7001.magru.wmnet * 16:57 swfrench-wmf: reprepro include php8.3_8.3.32-1+wmf12u2 into component/php83 for bookworm-wikimedia * 16:50 sukhe: pool cp2046 * 16:47 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4039.ulsfo.wmnet * 16:44 sukhe: sudo cumin -b31 "A:cp" "run-puppet-agent" * 16:33 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on contint1003.wikimedia.org with reason: reboot * 16:32 mutante: contint1003 - main CI server - rebooting * 16:31 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2188: repooling after rack b5 maintenance * 16:31 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2178: repooling after rack b5 maintenance * 16:29 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 16:28 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 16:28 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 16:28 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 16:18 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2014.codfw.wmnet * 16:18 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 16:17 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2014.codfw.wmnet * 16:10 mvernon@cumin1003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-thanos-proxies (exit_code=0) rolling restart_daemons on A:thanos-fe * 16:09 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 16:07 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp4039.ulsfo.wmnet * 16:07 mvernon@cumin1003: START - Cookbook sre.swift.roll-restart-reboot-swift-thanos-proxies rolling restart_daemons on A:thanos-fe * 16:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2014.codfw.wmnet with OS trixie * 15:56 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1014.eqiad.wmnet with OS trixie * 15:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2014.codfw.wmnet with reason: host reimage * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2014 * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2014 * 15:28 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2014 * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2014.codfw.wmnet 194.16.192.10.in-addr.arpa 4.9.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:28 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2014.codfw.wmnet 194.16.192.10.in-addr.arpa 4.9.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2014 - mvernon@cumin2003" * 15:28 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2014 - mvernon@cumin2003" * 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Apply title-related policies when selecting the name of the entity - kamila@cumin1003" * 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Apply title-related policies when selecting the name of the entity - kamila@cumin1003 * 15:22 kamila@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Apply title-related policies when selecting the name of the entity - kamila@cumin1003 * 15:22 kamila@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Apply title-related policies when selecting the name of the entity - kamila@cumin1003" * 15:20 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 15:20 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2014 * 15:20 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2014.codfw.wmnet with OS trixie * 15:19 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1014.eqiad.wmnet with OS trixie * 15:01 dancy@deploy2003: Installation of scap version "4.274.1" completed for 3 hosts * 15:00 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2177: repooling after rack b5 maintenance * 15:00 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2159: repooling after rack b5 maintenance * 14:59 dancy@deploy2003: Installing scap version "4.274.1" for 3 host(s) * 14:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2015.codfw.wmnet with OS trixie * 14:54 seanleong-wmde: Finished populateSitesTable for isvwiki ([[phab:T429939|T429939]]) * 14:53 javiermonton@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] (duration: 07m 35s) * 14:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1015.eqiad.wmnet with OS trixie * 14:49 javiermonton@deploy2003: javiermonton: Continuing with deployment * 14:48 javiermonton@deploy2003: javiermonton: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:46 javiermonton@deploy2003: Started scap sync-world: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] * 14:42 otto@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 14:41 otto@deploy2003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 14:41 otto@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 14:40 otto@deploy2003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 14:40 otto@deploy2003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 14:39 otto@deploy2003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 14:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2015.codfw.wmnet with reason: host reimage * 14:34 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1015.eqiad.wmnet with reason: host reimage * 14:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2015.codfw.wmnet with reason: host reimage * 14:30 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1015.eqiad.wmnet with reason: host reimage * 14:30 seanleong-wmde@deploy2003: mwscript-k8s job started: foreachwikiindblist wikidataclient extensions/Wikibase/lib/maintenance/populateSitesTable.php --force-protocol https # [[phab:T429939|T429939]] * 14:24 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling reboot on A:durum and not (A:durum-eqiad or A:durum-codfw or A:durum-esams) and A:durum * 14:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2015.codfw.wmnet with OS trixie * 14:15 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2016.codfw.wmnet with OS trixie * 14:14 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2159: repooling after rack b5 maintenance * 14:14 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1015.eqiad.wmnet with OS trixie * 14:12 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1016.eqiad.wmnet with OS trixie * 14:12 cmooney@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=pki,name=codfw * 14:12 sbisson@deploy2003: helmfile [codfw] DONE helmfile.d/services/cxserver: sync * 14:11 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2002.codfw.wmnet * 14:11 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2002.codfw.wmnet * 14:11 sbisson@deploy2003: helmfile [codfw] START helmfile.d/services/cxserver: sync * 14:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1005.wikimedia.org * 14:07 sbisson@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cxserver: sync * 14:07 sbisson@deploy2003: helmfile [eqiad] START helmfile.d/services/cxserver: sync * 14:05 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1005.wikimedia.org * 14:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader2005.wikimedia.org * 14:02 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2003.codfw.wmnet * 14:02 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2003.codfw.wmnet * 14:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=tegola-vector-tiles,name=codfw * 14:00 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=kartotherian,name=codfw * 14:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader2005.wikimedia.org * 13:58 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2016.codfw.wmnet with reason: host reimage * 13:57 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-ncredir (exit_code=0) rolling reboot on A:ncredir and A:ncredir * 13:57 sbisson@deploy2003: helmfile [staging] DONE helmfile.d/services/cxserver: sync * 13:56 sbisson@deploy2003: helmfile [staging] START helmfile.d/services/cxserver: sync * 13:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1016.eqiad.wmnet with reason: host reimage * 13:52 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:52 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:51 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2016.codfw.wmnet with reason: host reimage * 13:50 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1016.eqiad.wmnet with reason: host reimage * 13:49 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy (exit_code=0) rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 13:49 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling reboot on A:wikidough * 13:46 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-tcp-proxy (exit_code=0) rolling reboot on A:tcpproxy and A:tcpproxy * 13:43 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and not (A:durum-eqiad or A:durum-codfw or A:durum-esams) and A:durum * 13:42 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=97) rolling reboot on A:durum and A:durum * 13:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2011.codfw.wmnet * 13:36 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-reboot (exit_code=0) rolling reboot on A:dnsbox and A:ulsfo and (A:dnsbox) * 13:36 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns4004.wikimedia.org * 13:34 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1016.eqiad.wmnet with OS trixie * 13:34 topranks: reboot lsw1-b5-codfw to upgrade JunOS [[phab:T430918|T430918]] * 13:34 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2016.codfw.wmnet with OS trixie * 13:32 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2002.codfw.wmnet * 13:31 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2011.codfw.wmnet * 13:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2012.codfw.wmnet * 13:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2017.codfw.wmnet with OS trixie * 13:25 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1017.eqiad.wmnet with OS trixie * 13:24 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2012.codfw.wmnet * 13:22 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2002.codfw.wmnet * 13:22 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:22 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:22 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns4004.wikimedia.org * 13:21 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2005.codfw.wmnet * 13:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2013.codfw.wmnet * 13:18 elukey@dns1004: END - running authdns-update * 13:17 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2005.codfw.wmnet * 13:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2004.codfw.wmnet * 13:16 elukey@dns1004: START - running authdns-update * 13:16 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1046: es1046 after reimage * 13:14 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1029.eqiad.wmnet,service=s8 * 13:14 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1029.eqiad.wmnet,service=s5 * 13:13 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1029.eqiad.wmnet,service=s5 * 13:13 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1029.eqiad.wmnet,service=s8 * 13:13 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2004.codfw.wmnet * 13:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2013.codfw.wmnet * 13:11 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2014.codfw.wmnet * 13:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2003.codfw.wmnet * 13:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2017.codfw.wmnet with reason: host reimage * 13:07 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:07 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns4003.wikimedia.org * 13:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2003.codfw.wmnet * 13:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1067.eqiad.wmnet * 13:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1067.eqiad.wmnet * 13:06 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1067.eqiad.wmnet * 13:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm-test1001.wikimedia.org * 13:05 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2017.codfw.wmnet with reason: host reimage * 13:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1017.eqiad.wmnet with reason: host reimage * 13:04 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2014.codfw.wmnet * 13:03 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:02 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2188: codfw rack B5 depool for maintenance * 13:02 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_magru * 13:01 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2188: codfw rack B5 depool for maintenance * 13:01 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_magru * 13:01 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2178: codfw rack B5 depool for maintenance * 13:01 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm-test1001.wikimedia.org * 13:01 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2178: codfw rack B5 depool for maintenance * 13:01 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2177: codfw rack B5 depool for maintenance * 13:00 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2177: codfw rack B5 depool for maintenance * 12:59 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1068.eqiad.wmnet * 12:59 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1068.eqiad.wmnet * 12:58 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola-vector-tiles,name=codfw * 12:58 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2159: codfw rack B5 depool for maintenance * 12:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1017.eqiad.wmnet with reason: host reimage * 12:58 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola,name=codfw * 12:57 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=kartotherian,name=codfw * 12:57 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2159: codfw rack B5 depool for maintenance * 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 30 hosts with reason: lsw1-b5-codfw JunOS upgrade * 12:55 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lsw1-b5-codfw,lsw1-b5-codfw IPv6,lsw1-b5-codfw.mgmt,ssw1-a[1,8]-codfw.mgmt with reason: switch upgade lsw1-b5-codfw * 12:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet * 12:49 cmooney@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=pki,name=codfw * 12:49 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1067.eqiad.wmnet with OS trixie * 12:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet * 12:48 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2017.codfw.wmnet with OS trixie * 12:47 topranks: depool codfw pki in dns discovery ahead of lsw1-b5-codfw maintenance [[phab:T430918|T430918]] * 12:47 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns4003.wikimedia.org * 12:47 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and A:ulsfo and (A:dnsbox) * 12:47 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2018.codfw.wmnet * 12:45 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2018.codfw.wmnet * 12:45 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and A:durum * 12:45 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-tcp-proxy rolling reboot on A:tcpproxy and A:tcpproxy * 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host pki-root1002.eqiad.wmnet * 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1009.eqiad.wmnet * 12:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1009.eqiad.wmnet * 12:44 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 12:43 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-ncredir rolling reboot on A:ncredir and A:ncredir * 12:43 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling reboot on A:wikidough * 12:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2018.codfw.wmnet with OS trixie * 12:42 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1017.eqiad.wmnet with OS trixie * 12:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader2006.wikimedia.org * 12:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1018.eqiad.wmnet with OS trixie * 12:39 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1009.eqiad.wmnet * 12:38 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host pki-root1002.eqiad.wmnet * 12:38 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1009.eqiad.wmnet * 12:38 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1008.eqiad.wmnet * 12:38 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1008.eqiad.wmnet * 12:35 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader2006.wikimedia.org * 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1006.wikimedia.org * 12:33 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1008.eqiad.wmnet * 12:30 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1046: es1046 after reimage * 12:29 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host es1046.eqiad.wmnet with OS trixie * 12:29 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1006.wikimedia.org * 12:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test2005.wikimedia.org * 12:28 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1008.eqiad.wmnet * 12:27 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1007.eqiad.wmnet * 12:27 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1007.eqiad.wmnet * 12:27 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1067.eqiad.wmnet with reason: host reimage * 12:25 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1068.eqiad.wmnet with reason: vacuum overlarge container dbs * 12:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2018.codfw.wmnet with reason: host reimage * 12:24 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test2005.wikimedia.org * 12:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test1005.wikimedia.org * 12:23 Amir1: mwscript-k8s --follow --dblist=ores -- extensions/ORES/maintenance/PurgeScoreCache.php --model goodfaith --old ([[phab:T431159|T431159]]) * 12:22 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1007.eqiad.wmnet * 12:22 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test1005.wikimedia.org * 12:22 atsukoito: restarting pybal on lvs2013 `low-traffic` for https://gerrit.wikimedia.org/r/1310535 * 12:22 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1007.eqiad.wmnet * 12:21 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1006.eqiad.wmnet * 12:21 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1006.eqiad.wmnet * 12:20 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1018.eqiad.wmnet with reason: host reimage * 12:19 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2018.codfw.wmnet with reason: host reimage * 12:18 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1067.eqiad.wmnet with reason: host reimage * 12:16 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1006.eqiad.wmnet * 12:15 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1006.eqiad.wmnet * 12:15 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1005.eqiad.wmnet * 12:15 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1005.eqiad.wmnet * 12:15 atsukoito: restarting pybal on lvs2014 for https://gerrit.wikimedia.org/r/1310535 * 12:12 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1018.eqiad.wmnet with reason: host reimage * 12:11 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1005.eqiad.wmnet * 12:11 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1005.eqiad.wmnet * 12:10 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1004.eqiad.wmnet * 12:10 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1004.eqiad.wmnet * 12:09 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on es1046.eqiad.wmnet with reason: host reimage * 12:08 atsukoito: restarting pybal on lvs1019 `low-traffic` for https://gerrit.wikimedia.org/r/1310535 * 12:06 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1004.eqiad.wmnet * 12:06 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1004.eqiad.wmnet * 12:06 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1003.eqiad.wmnet * 12:06 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1003.eqiad.wmnet * 12:05 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on es1046.eqiad.wmnet with reason: host reimage * 12:05 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "set ml-serve1001 back to active state - cmooney@cumin1003" * 12:04 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "set ml-serve1001 back to active state - cmooney@cumin1003" * 12:04 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:02 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1003.eqiad.wmnet * 12:01 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1003.eqiad.wmnet * 12:01 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1002.eqiad.wmnet * 12:01 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1002.eqiad.wmnet * 12:01 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:59 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2018.codfw.wmnet with OS trixie * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1067 * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1067 * 11:59 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1067 * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1067.eqiad.wmnet 17.48.64.10.in-addr.arpa 7.1.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:59 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1067.eqiad.wmnet 17.48.64.10.in-addr.arpa 7.1.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1067 - blake@cumin1003" * 11:58 atsukoito: restarting pybal on lvs1018 `high-traffic2` for https://gerrit.wikimedia.org/r/1310535 * 11:57 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1002.eqiad.wmnet * 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2019.codfw.wmnet with OS trixie * 11:56 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1002.eqiad.wmnet * 11:56 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 11:56 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1018.eqiad.wmnet with OS trixie * 11:54 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 11:54 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:54 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1019.eqiad.wmnet with OS trixie * 11:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:49 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:49 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:49 aikochou@deploy2003: helmfile [codfw] DONE helmfile.d/services/changeprop: sync * 11:48 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host es1046.eqiad.wmnet with OS trixie * 11:48 aikochou@deploy2003: helmfile [codfw] START helmfile.d/services/changeprop: sync * 11:48 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310535 * 11:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1046: Reimage to Trixie * 11:44 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1046: Reimage to Trixie * 11:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5:00:00 on es1046.eqiad.wmnet with reason: Reimage to Trixie * 11:42 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:42 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:42 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 11:42 aikochou@deploy2003: helmfile [eqiad] DONE helmfile.d/services/changeprop: sync * 11:41 aikochou@deploy2003: helmfile [eqiad] START helmfile.d/services/changeprop: sync * 11:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2019.codfw.wmnet with reason: host reimage * 11:36 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] (duration: 09m 41s) * 11:36 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:36 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:35 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:35 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:32 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1019.eqiad.wmnet with reason: host reimage * 11:32 jforrester@deploy2003: jforrester, gengh: Continuing with deployment * 11:29 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2019.codfw.wmnet with reason: host reimage * 11:28 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1019.eqiad.wmnet with reason: host reimage * 11:28 jforrester@deploy2003: jforrester, gengh: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:26 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] * 11:20 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2003.codfw.wmnet * 11:20 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:19 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2003.codfw.wmnet * 11:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:12 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1019.eqiad.wmnet with OS trixie * 11:10 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2019.codfw.wmnet with OS trixie * 11:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1020.eqiad.wmnet with OS trixie * 11:10 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1067 - blake@cumin1003" * 11:09 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] (duration: 12m 12s) * 11:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2020.codfw.wmnet with OS trixie * 11:03 kharlan@deploy2003: kharlan: Continuing with deployment * 11:01 blake@cumin1003: START - Cookbook sre.dns.netbox * 11:01 kharlan@deploy2003: kharlan: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:57 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] * 10:55 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] (duration: 31m 40s) * 10:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1020.eqiad.wmnet with reason: host reimage * 10:52 marostegui@dns1004: START - running authdns-update * 10:49 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2020.codfw.wmnet with reason: host reimage * 10:49 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:48 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1020.eqiad.wmnet with reason: host reimage * 10:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2020.codfw.wmnet with reason: host reimage * 10:43 kharlan@deploy2003: kharlan: Continuing with deployment * 10:42 kharlan@deploy2003: kharlan: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:32 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1020.eqiad.wmnet with OS trixie * 10:29 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2159: Repooling after switchover * 10:29 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1067 * 10:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1021.eqiad.wmnet with OS trixie * 10:27 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1067.eqiad.wmnet with OS trixie * 10:27 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:27 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1067.eqiad.wmnet * 10:27 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:26 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1067.eqiad.wmnet * 10:26 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1067.eqiad.wmnet * 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2020.codfw.wmnet with OS trixie * 10:26 blake@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1055.eqiad.wmnet * 10:26 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1055.eqiad.wmnet * 10:26 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1055.eqiad.wmnet * 10:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2021.codfw.wmnet with OS trixie * 10:24 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] * 10:11 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1055.eqiad.wmnet with OS trixie * 10:09 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1021.eqiad.wmnet with reason: host reimage * 10:05 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2021.codfw.wmnet with reason: host reimage * 10:03 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310129 revert * 10:02 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1021.eqiad.wmnet with reason: host reimage * 10:01 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2021.codfw.wmnet with reason: host reimage * 09:58 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310129 * 09:50 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1055.eqiad.wmnet with reason: host reimage * 09:45 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1055.eqiad.wmnet with reason: host reimage * 09:45 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1021.eqiad.wmnet with OS trixie * 09:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1022.eqiad.wmnet with OS trixie * 09:44 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2159: Repooling after switchover * 09:44 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2021.codfw.wmnet with OS trixie * 09:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2022.codfw.wmnet with OS trixie * 09:31 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2159.codfw.wmnet * 09:28 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1055 * 09:28 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1055 * 09:27 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on ms-fe1022.eqiad.wmnet with reason: host reimage * 09:27 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1022.eqiad.wmnet with reason: host reimage * 09:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2022.codfw.wmnet with reason: host reimage * 09:21 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2022.codfw.wmnet with reason: host reimage * 09:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2159: Rebooting db2159.codfw.wmnet * 09:20 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2159: Rebooting db2159.codfw.wmnet * 09:18 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 09:18 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 09:18 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 09:17 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 09:16 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2159.codfw.wmnet * 09:13 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b] (thin): Regular analytics weekly train THIN [analytics/refinery@ad6e05b8] (duration: 02m 07s) * 09:11 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b] (thin): Regular analytics weekly train THIN [analytics/refinery@ad6e05b8] * 09:10 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1022.eqiad.wmnet with OS trixie * 09:07 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1023.eqiad.wmnet with OS trixie * 09:06 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b]: Regular analytics weekly train [analytics/refinery@ad6e05b8] (duration: 04m 49s) * 09:04 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2022.codfw.wmnet with OS trixie * 09:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2023.codfw.wmnet with OS trixie * 09:01 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b]: Regular analytics weekly train [analytics/refinery@ad6e05b8] * 09:01 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@ad6e05b8] (duration: 02m 01s) * 09:00 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1055 * 09:00 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1055.eqiad.wmnet 50.32.64.10.in-addr.arpa 0.5.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:00 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1055.eqiad.wmnet 50.32.64.10.in-addr.arpa 0.5.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:00 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:00 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1055 - blake@cumin1003" * 09:00 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1055 - blake@cumin1003" * 08:59 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@ad6e05b8] * 08:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2159 [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94811 and previous config saved to /var/cache/conftool/dbconfig/20260714-085624-cwilliams.json * 08:55 blake@cumin1003: START - Cookbook sre.dns.netbox * 08:55 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1055 * 08:54 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1055.eqiad.wmnet with OS trixie * 08:54 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1055.eqiad.wmnet * 08:53 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1055.eqiad.wmnet * 08:53 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1055.eqiad.wmnet * 08:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2220 to s7 primary [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94810 and previous config saved to /var/cache/conftool/dbconfig/20260714-085239-cwilliams.json * 08:51 cezmunsta: Starting s7 codfw failover from db2159 to db2220 - [[phab:T430920|T430920]] * 08:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1023.eqiad.wmnet with reason: host reimage * 08:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2220 with weight 0 [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94809 and previous config saved to /var/cache/conftool/dbconfig/20260714-084553-cwilliams.json * 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s7 [[phab:T430920|T430920]] * 08:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2023.codfw.wmnet with reason: host reimage * 08:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1023.eqiad.wmnet with reason: host reimage * 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2023.codfw.wmnet with reason: host reimage * 08:34 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:34 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:29 marostegui@dns1004: END - running authdns-update * 08:29 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox-dev2003.codfw.wmnet * 08:27 marostegui@dns1004: START - running authdns-update * 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker2*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2009.codfw.wmnet * 08:26 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2009.codfw.wmnet * 08:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1023.eqiad.wmnet with OS trixie * 08:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox-dev2003.codfw.wmnet * 08:24 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:24 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:24 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2023.codfw.wmnet with OS trixie * 08:24 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1029.eqiad.wmnet with reason: reboot * 08:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1027.eqiad.wmnet with reason: reboot * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:21 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2009.codfw.wmnet * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:20 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2009.codfw.wmnet * 08:20 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2008.codfw.wmnet * 08:20 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2008.codfw.wmnet * 08:15 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2008.codfw.wmnet * 08:14 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2008.codfw.wmnet * 08:14 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2007.codfw.wmnet * 08:14 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2007.codfw.wmnet * 08:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1024.eqiad.wmnet with OS trixie * 08:12 elukey@cumin1003: END (PASS) - Cookbook sre.pki.restart-reboot (exit_code=0) rolling reboot on P<nowiki>{</nowiki>pki*<nowiki>}</nowiki> and (A:pki) * 08:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 08:10 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 08:09 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2007.codfw.wmnet * 08:08 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2007.codfw.wmnet * 08:08 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2006.codfw.wmnet * 08:08 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2006.codfw.wmnet * 08:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2024.codfw.wmnet with OS trixie * 08:03 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2006.codfw.wmnet * 08:02 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2006.codfw.wmnet * 08:02 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2005.codfw.wmnet * 08:02 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2005.codfw.wmnet * 07:58 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2005.codfw.wmnet * 07:58 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2005.codfw.wmnet * 07:57 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2004.codfw.wmnet * 07:57 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2004.codfw.wmnet * 07:54 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki.discovery.wmnet. on all recursors * 07:54 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache pki.discovery.wmnet. on all recursors * 07:53 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2004.codfw.wmnet * 07:53 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2004.codfw.wmnet * 07:53 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2003.codfw.wmnet * 07:53 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2003.codfw.wmnet * 07:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1024.eqiad.wmnet with reason: host reimage * 07:49 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki.discovery.wmnet. on all recursors * 07:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2003.codfw.wmnet * 07:49 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache pki.discovery.wmnet. on all recursors * 07:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2024.codfw.wmnet with reason: host reimage * 07:48 elukey@cumin1003: START - Cookbook sre.pki.restart-reboot rolling reboot on P<nowiki>{</nowiki>pki*<nowiki>}</nowiki> and (A:pki) * 07:46 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1024.eqiad.wmnet with reason: host reimage * 07:45 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2024.codfw.wmnet with reason: host reimage * 07:45 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2003.codfw.wmnet * 07:44 elukey@cumin1003: END (PASS) - Cookbook sre.misc-clusters.restart-reboot-config-master (exit_code=0) rolling reboot on P<nowiki>{</nowiki>config-master*<nowiki>}</nowiki> and (A:config-master or A:config-master-eqiad or A:config-master-codfw) * 07:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2002.codfw.wmnet * 07:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2002.codfw.wmnet * 07:39 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2002.codfw.wmnet * 07:39 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) config-master.discovery.wmnet. on all recursors * 07:39 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache config-master.discovery.wmnet. on all recursors * 07:39 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2002.codfw.wmnet * 07:39 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker2*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl200*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl2003.codfw.wmnet * 07:36 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl2003.codfw.wmnet * 07:35 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) config-master.discovery.wmnet. on all recursors * 07:35 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache config-master.discovery.wmnet. on all recursors * 07:34 elukey@cumin1003: START - Cookbook sre.misc-clusters.restart-reboot-config-master rolling reboot on P<nowiki>{</nowiki>config-master*<nowiki>}</nowiki> and (A:config-master or A:config-master-eqiad or A:config-master-codfw) * 07:31 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl2003.codfw.wmnet * 07:31 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl2003.codfw.wmnet * 07:31 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl2002.codfw.wmnet * 07:31 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl2002.codfw.wmnet * 07:29 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1024.eqiad.wmnet with OS trixie * 07:28 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2024.codfw.wmnet with OS trixie * 07:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl2002.codfw.wmnet * 07:26 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl2002.codfw.wmnet * 07:26 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl200*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 07:26 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 07:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 06:50 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lists1004.wikimedia.org * 06:44 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host lists1004.wikimedia.org * 06:25 marostegui@dns1004: END - running authdns-update * 06:23 marostegui@dns1004: START - running authdns-update * 06:22 marostegui@dns1004: END - running authdns-update * 06:20 marostegui@dns1004: START - running authdns-update * 06:17 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1026.eqiad.wmnet with reason: reboot * 06:04 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: sync * 06:04 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: sync * 06:03 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync * 06:03 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync * 06:02 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync * 06:01 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync * 06:01 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync * 06:00 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync * 05:59 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:59 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:40 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:39 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:26 marostegui@dns1004: END - running authdns-update * 05:24 marostegui@dns1004: START - running authdns-update * 05:24 marostegui@dns1004: START - running authdns-update * 05:13 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1004.wikimedia.org * 05:07 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1004.wikimedia.org * 04:01 mwpresync@deploy2003: Pruned MediaWiki: 1.47.0-wmf.8 (duration: 01m 07s) * 03:39 mwpresync@deploy2003: Finished scap sync-world: testwikis to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] (duration: 36m 01s) * 03:03 mwpresync@deploy2003: Started scap sync-world: testwikis to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 29s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-13 == * 23:33 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1064.eqiad.wmnet * 23:33 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1064.eqiad.wmnet * 23:08 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1064.eqiad.wmnet with reason: vacuum overlarge container dbs * 23:06 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1069.eqiad.wmnet * 23:06 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1069.eqiad.wmnet * 22:34 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1069.eqiad.wmnet with reason: vacuum overlarge container dbs * 21:18 maryum: Deployed security fix for [[phab:T321092|T321092]] * 20:28 swfrench-wmf: reprepro include etcd-mirror_0.0.12-1+deb13u1 into main for trixie-wikimedia - [[phab:T424266|T424266]] * 20:26 swfrench-wmf: reprepro include etcd-mirror_0.0.12-1+deb12u1 into main for bookworm-wikimedia - [[phab:T428495|T428495]] * 20:23 dancy@deploy2003: Finished scap sync-world: Testing [[phab:T431635|T431635]] (duration: 03m 36s) * 20:19 dancy@deploy2003: Started scap sync-world: Testing [[phab:T431635|T431635]] * 20:18 dancy@deploy2003: Installation of scap version "4.274.0" completed for 3 hosts * 20:16 dancy@deploy2003: Installing scap version "4.274.0" for 3 host(s) * 20:12 kemayo@deploy2003: Finished scap sync-world: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] (duration: 08m 25s) * 20:07 kemayo@deploy2003: soda, esanders, kemayo: Continuing with deployment * 20:05 kemayo@deploy2003: soda, esanders, kemayo: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there * 20:04 kemayo@deploy2003: Started scap sync-world: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] * 18:22 cwhite: lvextend vg0/srv +500g on centrallog hosts * 18:19 cdobbins@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS trixie * 17:46 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1071.eqiad.wmnet * 17:46 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1071.eqiad.wmnet * 17:13 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1071.eqiad.wmnet with reason: vacuum overlarge container dbs * 17:07 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1065.eqiad.wmnet * 17:07 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1065.eqiad.wmnet * 17:06 dzahn@dns1006: END - running authdns-update * 17:04 dzahn@dns1006: START - running authdns-update * 17:01 dzahn@dns1006: END - running authdns-update * 16:59 dzahn@dns1006: START - running authdns-update * 16:51 dancy@deploy2003: Finished scap sync-world: testing [[phab:T428971|T428971]] (duration: 03m 37s) * 16:47 dancy@deploy2003: Started scap sync-world: testing [[phab:T428971|T428971]] * 16:45 atsukoito: restarting pybal on lvs1019 to flush IP address for `cirrussearch1122.eqiad.wmnet` after moving the vlan [[phab:T431311|T431311]] * 16:42 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:42 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:42 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:42 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:42 Amir1: mwscript-k8s --follow --dblist=ores -- extensions/ORES/maintenance/PurgeScoreCache.php --model damaging --old ([[phab:T431159|T431159]]) * 16:34 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Pool test * 16:34 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 16:34 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 16:34 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Pool test * 16:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Depool test * 16:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 16:33 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 16:33 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Depool test * 16:31 dancy@deploy2003: Installation of scap version "4.273.0" completed for 159 hosts * 16:29 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1065.eqiad.wmnet with reason: vacuum overlarge container dbs * 16:27 dancy@deploy2003: Installing scap version "4.273.0" for 159 host(s) * 16:27 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics-external: sync * 16:27 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics-external: sync * 16:26 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics-external: sync * 16:26 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics-external: sync * 16:22 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync * 16:21 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync * 16:21 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: sync * 16:21 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: sync * 16:19 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync * 16:19 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync * 16:18 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync * 16:17 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync * 15:59 atsukoito: restarting pybal on lvs1018 for https://gerrit.wikimedia.org/r/1310117 * 15:55 aikochou@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 15:50 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310117 * 15:46 aikochou@deploy2003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 15:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host kafka-logging1006.eqiad.wmnet * 15:43 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host kafka-logging1006.eqiad.wmnet * 15:41 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host ganeti-test[2001-2003].codfw.wmnet * 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host ganeti-test[2001-2003].codfw.wmnet * 15:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host netbox1003.eqiad.wmnet * 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host netbox1003.eqiad.wmnet * 15:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host netbox2003.codfw.wmnet * 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host netbox2003.codfw.wmnet * 15:36 sukhe: restart pybal on lvs1020 * 15:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet * 15:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet * 15:08 btullis@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:06 btullis@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 15:01 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:01 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:35 cdobbins@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 14:34 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:33 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:33 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:32 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:29 cdobbins@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 14:28 marostegui@dns1004: END - running authdns-update * 14:28 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:27 marostegui@dns1004: START - running authdns-update * 14:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1023.eqiad.wmnet with reason: reboot * 14:18 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2009.codfw.wmnet with OS trixie * 14:14 swfrench-wmf: start rolling run-puppet-agent on A:cp for ATS config change - [[phab:T428909|T428909]] [[phab:T431838|T431838]] * 14:05 swfrench-wmf: disable-puppet on A:cp for ATS config change - [[phab:T428909|T428909]] [[phab:T431838|T431838]] * 14:05 cdobbins@cumin2002: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie * 14:02 marostegui@dns1004: END - running authdns-update * 14:00 marostegui@dns1004: START - running authdns-update * 14:00 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1070.eqiad.wmnet * 14:00 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1070.eqiad.wmnet * 13:58 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2009.codfw.wmnet with reason: host reimage * 13:57 cdobbins@cumin2002: conftool action : set/pooled=no; selector: name=dns7002.* * 13:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2009.codfw.wmnet with reason: host reimage * 13:48 rscout@deploy2003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply * 13:48 rscout@deploy2003: helmfile [eqiad] START helmfile.d/services/miscweb: apply * 13:48 rscout@deploy2003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply * 13:47 rscout@deploy2003: helmfile [codfw] START helmfile.d/services/miscweb: apply * 13:40 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:33 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2009.codfw.wmnet with OS trixie * 13:30 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1070.eqiad.wmnet with reason: vacuum overlarge container dbs * 13:28 aude@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] (duration: 11m 12s) * 13:23 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:22 aude@deploy2003: aikochou, javiermonton, aude, gkm563: Continuing with deployment * 13:22 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:19 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:19 aude@deploy2003: aikochou, javiermonton, aude, gkm563: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] synced to the testservers * 13:17 aude@deploy2003: Started scap sync-world: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] * 13:01 ladsgroup@deploy2003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 13:01 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:00 ladsgroup@deploy2003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 12:59 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:52 ladsgroup@deploy2003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 12:51 ladsgroup@deploy2003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 12:48 atsuko@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 12:48 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 12:47 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2008.codfw.wmnet with OS trixie * 12:47 atsuko@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 12:47 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply * 12:47 atsuko@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:46 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 12:45 atsuko@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:45 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply * 12:45 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:44 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:43 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] (duration: 07m 02s) * 12:38 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 12:37 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:36 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] * 12:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2008.codfw.wmnet with reason: host reimage * 12:23 Msz2001: Deployed changes to private code for Suggested Investigations * 12:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2008.codfw.wmnet with reason: host reimage * 12:20 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:19 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:17 atsuko@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 12:17 atsuko@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 12:16 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:15 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] (duration: 07m 14s) * 12:10 mszwarc@deploy2003: mszwarc: Continuing with deployment * 12:09 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:07 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] * 12:04 mszwarc@deploy2003: sync-world aborted: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] (duration: 00m 29s) * 12:03 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] * 12:01 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2008.codfw.wmnet with OS trixie * 12:00 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] (duration: 07m 37s) * 11:55 zabe@deploy2003: zabe: Continuing with deployment * 11:54 zabe@deploy2003: zabe: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:52 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] * 11:51 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:43 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:35 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:34 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:33 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:30 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:28 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:27 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:17 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2007.codfw.wmnet with OS trixie * 11:09 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] * 11:06 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=s8 * 11:00 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=x3 * 11:00 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=s5 * 10:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2007.codfw.wmnet with reason: host reimage * 10:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2007.codfw.wmnet with reason: host reimage * 10:51 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host dse-k8s-worker1023 * 10:50 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host dse-k8s-worker1023 * 10:44 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dse-k8s-worker1023 * 10:43 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host dse-k8s-worker1023 * 10:42 marostegui@cumin1003: dbctl commit (dc=all): 'Change x4 masters [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P94804 and previous config saved to /var/cache/conftool/dbconfig/20260713-104248-marostegui.json * 10:37 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:37 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:35 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:35 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:34 atsuko@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 10:34 atsuko@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 10:33 marostegui@cumin1003: dbctl commit (dc=all): 'Push x4 initial dbctl config [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P94803 and previous config saved to /var/cache/conftool/dbconfig/20260713-103259-marostegui.json * 10:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2007.codfw.wmnet with OS trixie * 09:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2006.codfw.wmnet with OS trixie * 09:42 marostegui@dns1004: END - running authdns-update * 09:40 marostegui@dns1004: START - running authdns-update * 09:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2006.codfw.wmnet with reason: host reimage * 09:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2006.codfw.wmnet with reason: host reimage * 09:06 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1024.eqiad.wmnet with reason: reboot * 09:06 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:01 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2006.codfw.wmnet with OS trixie * 08:44 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 08:43 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 08:43 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:42 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:42 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:42 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:41 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 08:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup2004.codfw.wmnet * 08:38 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:33 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host db1208.eqiad.wmnet * 08:30 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=x3 * 08:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2005.codfw.wmnet with OS trixie * 08:28 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup2004.codfw.wmnet * 08:28 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup2003.codfw.wmnet * 08:24 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1039: Repooling after testing * 08:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on clouddb1016.eqiad.wmnet with reason: cloning * 08:23 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s5 * 08:23 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s8 * 08:21 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 08:21 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 08:17 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup2003.codfw.wmnet * 08:17 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1004.eqiad.wmnet * 08:14 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1208.eqiad.wmnet * 08:11 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host phab1005.eqiad.wmnet * 08:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2005.codfw.wmnet with reason: host reimage * 08:07 marostegui@dns1004: END - running authdns-update * 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1004.eqiad.wmnet * 08:07 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1003.eqiad.wmnet * 08:07 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:05 marostegui@dns1004: START - running authdns-update * 08:05 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2005.codfw.wmnet with reason: host reimage * 08:05 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 08:05 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host phab1005.eqiad.wmnet * 08:05 marostegui@dns1004: START - running authdns-update * 08:05 marostegui@dns1004: START - running authdns-update * 08:05 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 08:04 marostegui@dns1004: START - running authdns-update * 08:00 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit1003.wikimedia.org * 07:58 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1003.eqiad.wmnet * 07:58 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1002-dev.eqiad.wmnet * 07:58 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:58 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:54 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1002-dev.eqiad.wmnet * 07:54 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1001-dev.eqiad.wmnet * 07:54 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit1003.wikimedia.org * 07:53 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 07:53 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:52 Msz2001: UTC morning backport+config window done * 07:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2005.codfw.wmnet with OS trixie * {{safesubst:SAL entry|1=07:50 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark (T429943}} * 07:49 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1001-dev.eqiad.wmnet * 07:46 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:46 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:45 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 07:45 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:44 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 07:44 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:43 mszwarc@deploy2003: mszwarc, danielyepezgarces, anzx: Continuing with deployment * {{safesubst:SAL entry|1=07:39 mszwarc@deploy2003: mszwarc, danielyepezgarces, anzx: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark}} * 07:39 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1039: Repooling after testing * {{safesubst:SAL entry|1=07:36 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark (T429943)}} * 07:35 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] (duration: 30m 03s) * 07:25 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit2002.wikimedia.org * 07:22 mszwarc@deploy2003: mszwarc: Continuing with deployment * 07:21 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:19 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit2002.wikimedia.org * 07:15 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aphlict1002.eqiad.wmnet * 07:11 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host aphlict1002.eqiad.wmnet * 07:08 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2003.wikimedia.org * 07:05 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] * 07:02 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2003.wikimedia.org * 07:02 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2002.wikimedia.org * 06:55 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2002.wikimedia.org * 06:55 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1003.wikimedia.org * 06:49 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1003.wikimedia.org * 06:34 marostegui: Drop m5 ipoid database [[phab:T431007|T431007]] * 06:29 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1027.eqiad.wmnet with reason: reboot * 06:24 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1028.eqiad.wmnet with reason: reboot * 06:21 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1025.eqiad.wmnet with reason: reboot * 06:17 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1022.eqiad.wmnet with reason: reboot * 06:03 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on dbproxy[2005-2008].codfw.wmnet with reason: reboot * 05:37 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1217,1228].eqiad.wmnet with reason: cloning * 05:11 marostegui: Drop users_to_rename table [[phab:T431842|T431842]] * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-12 == * 16:01 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2209 [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94792 and previous config saved to /var/cache/conftool/dbconfig/20260712-160124-marostegui.json * 15:58 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2205 to s3 primary [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94791 and previous config saved to /var/cache/conftool/dbconfig/20260712-155853-marostegui.json * 15:58 marostegui: Starting s3 codfw emergency failover from db2209 to db2205 - [[phab:T431950|T431950]] * 15:51 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2205 with weight 0 [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94790 and previous config saved to /var/cache/conftool/dbconfig/20260712-155135-marostegui.json * 15:51 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Primary switchover s3 [[phab:T431950|T431950]] * 02:01 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 01m 17s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-11 == * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 26s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-10 == * 19:12 jhathaway@dns1004: END - running authdns-update * 19:10 jhathaway@dns1004: START - running authdns-update * 18:23 mutante: vrts2002 rebooting (not the active host) * 18:21 mutante: lists2001, phab2003 - rebooting (not the active hosts) * 18:16 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on A:lvs-high-traffic2-codfw * 18:15 swfrench@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on A:lvs-high-traffic2-codfw * 17:15 mutante: [doc1004:~] $ sudo systemctl start rsync-doc-host-data-sync ([[phab:T431856|T431856]]) * 17:09 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1004.eqiad.wmnet * 17:08 jhathaway@dns1004: END - running authdns-update * 17:07 jhathaway@dns1004: START - running authdns-update * 17:06 jhathaway: depooling puppetserver1002, cause of errors is still unknown * 17:03 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1004.eqiad.wmnet * 16:57 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1003.eqiad.wmnet * 16:51 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1003.eqiad.wmnet * 16:48 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 16:48 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2004.codfw.wmnet * 16:42 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2004.codfw.wmnet * 16:41 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2003.codfw.wmnet * 16:35 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2003.codfw.wmnet * 16:33 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2002.codfw.wmnet * 16:27 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2002.codfw.wmnet * 16:25 mutante: gitlab-runners (production) rebooting cluster one by one * 16:17 mutante: etherpad1004/etherpad2002 - (etherpad.wikimedia.org) - rebooting * 16:13 mutante: doc1004/doc2003 (doc.wikimedia.org backends) - rebooting * 16:02 mutante: releases1003/releases2003 (releases.wikimedia.org backends) - rebooting for maintenance * 15:26 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2007-dev.codfw.wmnet * 15:19 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2007-dev.codfw.wmnet * 15:14 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host cloudcephosd2007-dev.codfw.wmnet * 15:14 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2007-dev.codfw.wmnet * 15:14 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host cloudcephosd2006-dev.codfw.wmnet * 15:07 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2006-dev.codfw.wmnet * 15:07 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2005-dev.codfw.wmnet * 14:59 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2005-dev.codfw.wmnet * 14:59 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2004-dev.codfw.wmnet * 14:53 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2004-dev.codfw.wmnet * 14:53 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2007-dev.codfw.wmnet * 14:51 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1054.eqiad.wmnet * 14:51 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1054.eqiad.wmnet * 14:51 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1054.eqiad.wmnet * 14:47 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2007-dev.codfw.wmnet * 14:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2006-dev.codfw.wmnet * 14:41 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2006-dev.codfw.wmnet * 14:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2005-dev.codfw.wmnet * 14:37 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2005-dev.codfw.wmnet * 14:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2005-dev.codfw.wmnet * 14:29 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2005-dev.codfw.wmnet * 14:29 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2006-dev.codfw.wmnet * 14:21 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2006-dev.codfw.wmnet * 14:21 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2010-dev.codfw.wmnet * 14:15 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2010-dev.codfw.wmnet * 14:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudgw2004-dev.codfw.wmnet * 14:10 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1054.eqiad.wmnet with OS trixie * 14:09 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudgw2004-dev.codfw.wmnet * 14:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudgw2003-dev.codfw.wmnet * 14:02 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudgw2003-dev.codfw.wmnet * 14:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2004-dev.codfw.wmnet * 13:53 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2004-dev.codfw.wmnet * 13:53 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2003-dev.codfw.wmnet * 13:48 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:44 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2003-dev.codfw.wmnet * 13:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2002-dev.codfw.wmnet * 13:42 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:41 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:41 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:37 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2002-dev.codfw.wmnet * 13:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudidp2001-dev.codfw.wmnet * 13:33 blake@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker1054.eqiad.wmnet with reason: host reimage * 13:33 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudidp2001-dev.codfw.wmnet * 13:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudnet2006-dev.codfw.wmnet * 13:26 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudnet2006-dev.codfw.wmnet * 13:26 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudnet2005-dev.codfw.wmnet * 13:23 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1054.eqiad.wmnet with reason: host reimage * 13:18 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudnet2005-dev.codfw.wmnet * 13:18 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudservices2005-dev.codfw.wmnet * 13:12 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudservices2005-dev.codfw.wmnet * 13:11 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudservices2004-dev.codfw.wmnet * 13:08 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudservices2004-dev.codfw.wmnet * 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudweb2002-dev.wikimedia.org * 13:05 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 13:05 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1054 * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1054 * 13:04 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1054 * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1054.eqiad.wmnet 49.32.64.10.in-addr.arpa 9.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:04 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1054.eqiad.wmnet 49.32.64.10.in-addr.arpa 9.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1054 - blake@cumin1003" * 13:04 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1054 - blake@cumin1003" * 13:01 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudweb2002-dev.wikimedia.org * 13:00 blake@cumin1003: START - Cookbook sre.dns.netbox * 12:59 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1054 * 12:57 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1054.eqiad.wmnet with OS trixie * 12:57 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1054.eqiad.wmnet * 12:56 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1054.eqiad.wmnet * 12:56 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1054.eqiad.wmnet * 12:47 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS trixie * 12:44 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:39 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 12:39 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 12:38 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:37 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:14 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:10 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:08 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:07 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:00 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:00 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:51 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:49 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:48 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:47 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:44 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:32 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2001.codfw.wmnet * 11:32 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1053.eqiad.wmnet * 11:32 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2001.codfw.wmnet * 11:32 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1053.eqiad.wmnet * 11:32 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1053.eqiad.wmnet * 11:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker2001.codfw.wmnet * 11:31 cgoubert@cumin1003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker2001.codfw.wmnet * 11:31 cgoubert@cumin1003: END (FAIL) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=1) rolling reimage on P<nowiki>{</nowiki>wikikube-worker2001*<nowiki>}</nowiki> and (A:wikikube-master-codfw or A:wikikube-worker-codfw) * 11:30 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:30 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:21 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 18 hosts with reason: reboot & upgrade * 11:20 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker2001.codfw.wmnet with OS trixie * 11:16 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:15 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:14 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:14 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 11 hosts * 11:14 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 11 hosts * 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:08 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 11:02 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1053.eqiad.wmnet with OS trixie * 11:01 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:58 cgoubert@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 10:57 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:38 cgoubert@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker2001.codfw.wmnet with OS trixie * 10:38 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2001.codfw.wmnet * 10:38 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2001.codfw.wmnet * 10:38 cgoubert@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on P<nowiki>{</nowiki>wikikube-worker2001*<nowiki>}</nowiki> and (A:wikikube-master-codfw or A:wikikube-worker-codfw) * 10:35 topranks: adjust IBGP outbound policy on lsw1-e2-codfw [[phab:T423430|T423430]] towards ssw1-e1-codfw * 10:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 cgoubert@cumin1003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:27 cgoubert@cumin1003: END (FAIL) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=1) rolling reimage on A:wikikube-worker-codfw * 10:27 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker2001.codfw.wmnet with OS bookworm * 10:25 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 10:24 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:24 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:15 cgoubert@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 10:11 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 10:11 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 10:08 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:08 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:07 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 10:06 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 10:00 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:55 cgoubert@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker2001.codfw.wmnet with OS bookworm * 09:55 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2005-2006,2011-2012].codfw.wmnet * 09:55 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2005-2006,2011-2012].codfw.wmnet * 09:51 cgoubert@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on A:wikikube-worker-codfw * 09:41 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1053.eqiad.wmnet with reason: host reimage * 09:37 topranks: apply new IBGP outbound policy on lsw1-e2-codfw [[phab:T423430|T423430]] * 09:36 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:36 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1053.eqiad.wmnet with reason: host reimage * 09:16 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1053 * 09:16 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1053 * 09:15 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1053 * 09:15 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1053.eqiad.wmnet 48.32.64.10.in-addr.arpa 8.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:15 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1053.eqiad.wmnet 48.32.64.10.in-addr.arpa 8.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:15 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:15 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1053 - blake@cumin1003" * 09:15 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1053 - blake@cumin1003" * 09:11 blake@cumin1003: START - Cookbook sre.dns.netbox * 09:11 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1053 * 09:08 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1053.eqiad.wmnet with OS trixie * 09:08 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1053.eqiad.wmnet * 09:08 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1053.eqiad.wmnet * 09:08 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1053.eqiad.wmnet * 09:04 brouberol@dns1004: END - running authdns-update * 09:03 brouberol@dns1004: START - running authdns-update * 08:41 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e] (thin): Regular analytics weekly train THIN [analytics/refinery@1abf22ea] (duration: 02m 11s) * 08:38 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e] (thin): Regular analytics weekly train THIN [analytics/refinery@1abf22ea] * 08:38 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e]: Regular analytics weekly train [analytics/refinery@1abf22ea] (duration: 05m 17s) * 08:38 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:34 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:33 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e]: Regular analytics weekly train [analytics/refinery@1abf22ea] * 08:32 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@1abf22ea] (duration: 02m 03s) * 08:30 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@1abf22ea] * 08:30 JavierMonton: Deploying Refinery at {{Gerrit|1abf22ea}} for changes 1308121/T427068 1306491/T430020 and {{Gerrit|1308190}} * 08:29 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:29 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 08:24 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:24 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 08:18 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:18 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 08:00 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db[2183-2184].codfw.wmnet * 08:00 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for db[2183-2184].codfw.wmnet * 07:52 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:52 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 07:49 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 11 hosts with reason: reboot & upgrade * 07:47 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:47 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 07:44 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:44 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 07:23 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 10 hosts * 07:23 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 10 hosts * 06:45 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 10 hosts with reason: reboot & upgrade * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 41s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-09 == * 23:33 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] (duration: 13m 26s) * 23:29 ladsgroup@deploy2003: ladsgroup, jdlrobson: Continuing with deployment * 23:22 ladsgroup@deploy2003: ladsgroup, jdlrobson: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:20 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] * 22:57 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1165.eqiad.wmnet * 22:56 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1165.eqiad.wmnet * 22:56 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1165.eqiad.wmnet * 22:45 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1165.eqiad.wmnet with OS trixie * 22:38 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 22:37 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 22:37 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 22:37 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:37 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 22:25 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1165.eqiad.wmnet with reason: host reimage * 22:17 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1165.eqiad.wmnet with reason: host reimage * 22:13 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 22:12 rzl@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 22:04 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 22:04 rzl@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1165 * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1165 * 22:02 jasmine@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1165 * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1165.eqiad.wmnet 115.48.64.10.in-addr.arpa 5.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:02 jasmine@cumin2002: START - Cookbook sre.dns.wipe-cache wikikube-worker1165.eqiad.wmnet 115.48.64.10.in-addr.arpa 5.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1165 - jasmine@cumin2002" * 22:02 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1165 - jasmine@cumin2002" * 22:02 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 21:57 jasmine@cumin2002: START - Cookbook sre.dns.netbox * 21:55 jasmine@cumin2002: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1165 * 21:54 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-worker1165.eqiad.wmnet with OS trixie * 21:54 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 21:54 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1165.eqiad.wmnet * 21:53 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 21:53 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1165.eqiad.wmnet * 21:53 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1165.eqiad.wmnet * 21:53 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 21:47 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 21:45 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 21:43 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 21:43 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 21:42 maryum: Deploy fix for [[phab:T431684|T431684]] * 21:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 21:27 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] (duration: 34m 14s) * 21:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs1002 * 21:23 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs1002 * 21:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS trixie * 21:22 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 21:20 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 22s) * 21:20 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 21:16 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 21:16 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 21:15 ladsgroup@deploy2003: ladsgroup: Continuing with deployment * 21:13 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2002.codfw.wmnet with OS bookworm * 21:11 ladsgroup@deploy2003: ladsgroup: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:08 ladsgroup@cumin1003: END (PASS) - Cookbook sre.wikireplicas.update-views (exit_code=0) * 21:07 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 6 hosts with reason: reboots * 20:54 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecycle work - bking@cumin2003 * 20:53 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:53 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] * 20:51 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) * 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 20:47 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecycle work - bking@cumin2003 * 20:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 20:41 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:41 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) * 20:40 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host relforge1008.eqiad.wmnet * 20:40 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1009.eqiad.wmnet with OS trixie * 20:33 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:32 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:32 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:31 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:31 ladsgroup@cumin1003: END (PASS) - Cookbook sre.wikireplicas.update-views (exit_code=0) * 20:29 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1008.eqiad.wmnet * 20:24 rzl@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 20:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2002.codfw.wmnet with OS bookworm * 20:23 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host relforge1008.eqiad.wmnet * 20:23 rzl@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 20:23 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1008.eqiad.wmnet * 20:22 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:22 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:21 bking@cumin2003: END (ERROR) - Cookbook sre.elasticsearch.rolling-operation (exit_code=97) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:21 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1009.eqiad.wmnet with reason: host reimage * 20:16 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:15 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1009.eqiad.wmnet with reason: host reimage * 20:12 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) * 20:02 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 19:55 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1009.eqiad.wmnet with OS trixie * 19:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 19:43 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 19:30 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 19:28 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 19:27 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 19:25 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 18:42 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 18:41 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 18:16 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 18:15 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 17:45 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for doh5004.wikimedia.org * 17:45 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for doh5004.wikimedia.org * 17:38 ladsgroup@deploy2003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 17:35 ladsgroup@deploy2003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 17:29 ladsgroup@deploy2003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 17:26 ladsgroup@deploy2003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 17:09 mutante: zuul[12]00[123] - rebooting for maintenance * 17:09 ebernhardson: start full in-place reindex of eqiad cirrussearch cluster * 17:08 dzahn@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-cluster (exit_code=99) * 17:08 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-cluster * 17:03 ebernhardson: start full in-place reindex of codfw cirrussearch cluster * 16:59 mutante: stewards1001/stewards2001 - reboot for maintenance * 16:54 ebernhardson: start full in-place reindex of cloudelastic cluster * 16:53 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 16:52 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply * 16:49 mutante: planet1003/planet2003 - rebooting * 16:47 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on doh5004.wikimedia.org with reason: random high load, investigating * 15:55 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 15:54 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 15:51 jynus: restarting backupmon1001 * 15:49 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 14 hosts * 15:49 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 14 hosts * 15:47 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backupmon1001.eqiad.wmnet with reason: restart * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:06 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 15:06 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 14:59 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply * 14:58 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply * 14:51 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 14 hosts * 14:51 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 14 hosts * 14:49 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 6 hosts with reason: reboot & upgrade * 14:48 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet * 14:48 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet * 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:42 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1052.eqiad.wmnet * 14:42 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1052.eqiad.wmnet * 14:42 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1052.eqiad.wmnet * 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:31 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:31 elukey@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: sync * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:30 elukey@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: sync * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:28 elukey@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: sync * 14:28 elukey@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: sync * 14:26 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:20 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1052.eqiad.wmnet with OS trixie * 14:19 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 6 hosts with reason: reboot & upgrade * 14:18 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:15 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:15 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:13 elukey: update druid indexation job for webrequest_sampled_live - [[phab:T427068|T427068]] * 14:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:09 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:09 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for papaul - jhancock@cumin2002" * 14:09 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for papaul - jhancock@cumin2002" * 14:07 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:07 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:04 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 14:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cuminunpriv1001.eqiad.wmnet * 13:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb1003.eqiad.wmnet * 13:59 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1052.eqiad.wmnet with reason: host reimage * 13:57 moritzm: installing requests security updates * 13:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cuminunpriv1001.eqiad.wmnet * 13:55 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb1003.eqiad.wmnet * 13:53 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1052.eqiad.wmnet with reason: host reimage * 13:50 moritzm: installing python-cryptography security updates * 13:47 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb2003.codfw.wmnet * 13:44 Msz2001: UTC afternoon config+backport window is done * 13:44 Msz2001: Updated `logging` on `metawiki` to fix log performers, [[phab:T431176|T431176]]#12105297 * 13:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb2003.codfw.wmnet * 13:43 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 13:43 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt1002.wikimedia.org * 13:41 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] (duration: 07m 30s) * 13:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt1002.wikimedia.org * 13:37 mszwarc@deploy2003: mszwarc: Continuing with deployment * 13:36 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1052 * 13:36 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1052 * 13:35 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:35 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1052 * 13:35 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1052.eqiad.wmnet 47.32.64.10.in-addr.arpa 7.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:35 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1052.eqiad.wmnet 47.32.64.10.in-addr.arpa 7.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:35 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:35 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1052 - blake@cumin1003" * 13:35 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1052 - blake@cumin1003" * 13:34 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] * 13:31 blake@cumin1003: START - Cookbook sre.dns.netbox * 13:31 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1052 * 13:30 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1052.eqiad.wmnet with OS trixie * 13:30 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1052.eqiad.wmnet * 13:29 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1052.eqiad.wmnet * 13:29 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1052.eqiad.wmnet * 13:17 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] (duration: 11m 26s) * 13:13 jforrester@deploy2003: jforrester: Continuing with deployment * 13:08 jforrester@deploy2003: jforrester: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:06 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] * 12:54 cgoubert@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply * 12:54 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:52 cgoubert@deploy2003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply * 12:45 cgoubert@deploy2003: helmfile [codfw] DONE helmfile.d/services/mobileapps: apply * 12:44 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:44 cgoubert@deploy2003: helmfile [codfw] START helmfile.d/services/mobileapps: apply * 12:43 cgoubert@deploy2003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 12:43 cgoubert@deploy2003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 12:42 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast4006.wikimedia.org * 12:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt2002.wikimedia.org * 12:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast7002.wikimedia.org * 12:18 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast4006.wikimedia.org * 12:18 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host ml-serve1004 * 12:18 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host ml-serve1004 * 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt2002.wikimedia.org * 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast7002.wikimedia.org * 12:10 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backup[2003,2014].codfw.wmnet with reason: reboot & upgrade * 12:10 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt-staging2001.codfw.wmnet * 12:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid1003.eqiad.wmnet * 12:06 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt-staging2001.codfw.wmnet * 12:05 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid1003.eqiad.wmnet * 12:03 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backup[1003,1014].eqiad.wmnet with reason: reboot & upgrade * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid2003.codfw.wmnet * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host irc1003.wikimedia.org * 11:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid2003.codfw.wmnet * 11:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host irc1003.wikimedia.org * 11:55 jmm@dns1004: END - running authdns-update * 11:53 jmm@dns1004: START - running authdns-update * 11:50 jmm@dns1004: END - running authdns-update * 11:48 jmm@dns1004: START - running authdns-update * 11:27 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host irc2003.wikimedia.org * 11:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host irc2003.wikimedia.org * 11:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint2001.codfw.wmnet * 11:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint1001.eqiad.wmnet * 11:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint2001.codfw.wmnet * 11:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint1001.eqiad.wmnet * 11:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-rw2001.wikimedia.org * 11:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-rw1001.wikimedia.org * 11:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-rw2001.wikimedia.org * 11:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-rw1001.wikimedia.org * 11:03 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon1003.wikimedia.org * 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2005.codfw.wmnet * 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2005.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 10:59 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2005.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 10:57 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon1003.wikimedia.org * 10:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon2002.wikimedia.org * 10:55 jmm@cumin2003: START - Cookbook sre.dns.netbox * 10:51 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon2002.wikimedia.org * 10:51 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:50 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2005.codfw.wmnet * 10:41 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:40 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host ml-serve1003 * 10:40 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host ml-serve1003 * 10:39 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2033.codfw.wmnet * 10:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install2005.wikimedia.org * 10:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install1005.wikimedia.org * 10:35 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1004.eqiad.wmnet with OS bookworm * 10:31 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install1005.wikimedia.org * 10:31 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install2005.wikimedia.org * 10:30 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install4004.wikimedia.org * 10:30 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install3004.wikimedia.org * 10:29 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install3004.wikimedia.org * 10:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install4004.wikimedia.org * 10:23 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 10:21 moritzm: failover Ganeti master in codfw/routed to ganeti2034 [[phab:T430928|T430928]] * 10:19 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.addnode (exit_code=0) for new host ganeti2031.codfw.wmnet to cluster codfw and group B * 10:19 moritzm: readded ganeti2031 to the codfw Ganeti cluster [[phab:T430910|T430910]] * 10:18 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1004.eqiad.wmnet with reason: host reimage * 10:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install5004.wikimedia.org * 10:18 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1003 * 10:18 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1003 * 10:17 jmm@cumin2003: START - Cookbook sre.ganeti.addnode for new host ganeti2031.codfw.wmnet to cluster codfw and group B * 10:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install6003.wikimedia.org * 10:16 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install5004.wikimedia.org * 10:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install6003.wikimedia.org * 10:15 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1004.eqiad.wmnet with reason: host reimage * 10:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1001.eqiad.wmnet * 10:14 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 10:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2008.wikimedia.org * 10:00 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ml-serve1004 * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1004 * 09:57 jmm@cumin2003: START - Cookbook sre.dns.netbox * 09:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install7002.wikimedia.org * 09:57 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1004 * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ml-serve1004.eqiad.wmnet 50.48.64.10.in-addr.arpa 0.5.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:57 klausman@cumin1003: START - Cookbook sre.dns.wipe-cache ml-serve1004.eqiad.wmnet 50.48.64.10.in-addr.arpa 0.5.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1004 - klausman@cumin1003" * 09:56 klausman@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1004 - klausman@cumin1003" * 09:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-coord1001.eqiad.wmnet * 09:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 09:55 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow7002.magru.wmnet * 09:52 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-coord1001.eqiad.wmnet * 09:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 09:52 klausman@cumin1003: START - Cookbook sre.dns.netbox * 09:50 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install7002.wikimedia.org * 09:50 klausman@cumin1003: START - Cookbook sre.hosts.move-vlan for host ml-serve1004 * 09:50 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1004.eqiad.wmnet with OS bookworm * 09:50 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1003.eqiad.wmnet with OS bookworm * 09:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1001.eqiad.wmnet * 09:49 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 09:49 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2008.wikimedia.org * 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2007.codfw.wmnet * 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2007.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 09:49 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow7002.magru.wmnet * 09:49 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2007.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 09:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard1003.eqiad.wmnet * 09:39 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard2003.codfw.wmnet * 09:39 jmm@cumin2003: START - Cookbook sre.dns.netbox * 09:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard1003.eqiad.wmnet * 09:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor1003.eqiad.wmnet * 09:35 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard2003.codfw.wmnet * 09:34 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2007.codfw.wmnet * 09:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor1003.eqiad.wmnet * 09:33 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor-dev2001.codfw.wmnet * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor2003.codfw.wmnet * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sretest1006.eqiad.wmnet * 09:27 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 09:25 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor-dev2001.codfw.wmnet * 09:25 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor2003.codfw.wmnet * 09:23 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] (duration: 06m 27s) * 09:23 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2205: codfw rack B4 repool after maintenance * 09:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host sretest1006.eqiad.wmnet * 09:23 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2204: codfw rack B4 repool after maintenance * 09:19 urbanecm@deploy2003: urbanecm: Continuing with deployment * 09:19 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:18 jmm@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 6 hosts with reason: reboot * 09:17 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] * 09:08 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ml-serve1003 * 09:08 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1003 * 09:07 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1003 * 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ml-serve1003.eqiad.wmnet 81.32.64.10.in-addr.arpa 1.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:07 klausman@cumin1003: START - Cookbook sre.dns.wipe-cache ml-serve1003.eqiad.wmnet 81.32.64.10.in-addr.arpa 1.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1003 - klausman@cumin1003" * 09:06 klausman@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1003 - klausman@cumin1003" * 08:58 klausman@cumin1003: START - Cookbook sre.dns.netbox * 08:57 klausman@cumin1003: START - Cookbook sre.hosts.move-vlan for host ml-serve1003 * 08:57 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1003.eqiad.wmnet with OS bookworm * 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=0) rolling reimage on P<nowiki>{</nowiki>ml-serve1003.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet * 08:55 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet * 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1003.eqiad.wmnet with OS bookworm * 08:39 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 08:38 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool db2205: codfw rack B4 repool after maintenance * 08:37 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool db2204: codfw rack B4 repool after maintenance * 08:36 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 08:35 hashar@deploy2003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 08:32 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:32 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:31 hashar@deploy2003: Rolling back deployment * 08:26 moritzm: failover Ganeti master in codfw to ganeti2048 [[phab:T430928|T430928]] * 08:16 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1003.eqiad.wmnet with OS bookworm * 08:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2004.codfw.wmnet * 08:16 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet * 08:16 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet * 08:16 klausman@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on P<nowiki>{</nowiki>ml-serve1003.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 08:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2002.codfw.wmnet * 08:15 XioNoX: lsw1-b4-codfw> request system reboot - [[phab:T430910|T430910]] * 08:15 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b4-codfw,lsw1-b4-codfw IPv6,lsw1-b4-codfw.mgmt with reason: Switch maintenance * 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for codfw rack B4 * 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:10 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2004.codfw.wmnet * 08:10 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2002.codfw.wmnet * 08:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2205: codfw rack B4 depool for maintenance * 08:08 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool db2205: codfw rack B4 depool for maintenance * 08:08 jmm@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin2003.codfw.wmnet * 08:08 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2204: codfw rack B4 depool for maintenance * 08:08 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool db2204: codfw rack B4 depool for maintenance * 08:08 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 27 hosts with reason: codfw rack B4 depool for maintenance * 08:03 jmm@cumin2002: START - Cookbook sre.hosts.reboot-single for host cumin2003.codfw.wmnet * 07:56 ayounsi@cumin1003: START - Cookbook sre.network.depool-rack with action 'depool' for codfw rack B4 * 07:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1008.eqiad.wmnet with OS trixie * 07:49 wmde-fisch@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] (duration: 08m 36s) * 07:44 wmde-fisch@deploy2003: wmde-fisch: Continuing with deployment * 07:43 wmde-fisch@deploy2003: wmde-fisch: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:41 wmde-fisch@deploy2003: Started scap sync-world: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] * 07:35 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1008.eqiad.wmnet with reason: host reimage * 07:31 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1008.eqiad.wmnet with reason: host reimage * 07:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1008.eqiad.wmnet with OS trixie * 07:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 07:00 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 06:59 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1008.eqiad.wmnet with OS trixie * 06:57 Emperor: rebalance thanos swift rings after previous re-image of thanos-fe1004 to trixie * 06:47 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1008.eqiad.wmnet with OS trixie * 04:10 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 14 days, 0:00:00 on cp6008.drmrs.wmnet with reason: Hardware failure - [[phab:T431651|T431651]] * 03:55 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp6008.* * 03:29 ryankemper: [[phab:T431311|T431311]] Repooled eqiad cirrussearch clusters (`chi/omega/psi`) following completion of OpenSearch 2.19 migration * 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad * 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=eqiad * 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 31s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-08 == * 23:52 Amir1: ladsgroup@deploy2003:~$ mwscript-k8s --follow -- extensions/ORES/maintenance/PurgeScoreCache.php --wiki=simplewiki --model damaging --old ([[phab:T431159|T431159]]) * 23:46 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 23:46 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing PTR for 2001:df2:e500:fe08::1 - cmooney@cumin1003" * 23:46 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing PTR for 2001:df2:e500:fe08::1 - cmooney@cumin1003" * 23:40 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 23:16 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 23:15 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 22:42 rzl@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 22:40 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] (duration: 12m 55s) * 22:40 rzl@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 22:37 rzl@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 22:36 rzl@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 22:35 rzl@deploy2003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 22:34 urbanecm@deploy2003: urbanecm: Continuing with deployment * 22:33 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:33 rzl@deploy2003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 22:32 rzl@deploy2003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 22:30 rzl@deploy2003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 22:30 rzl@deploy2003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 22:29 rzl@deploy2003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 22:27 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] * 22:26 rzl@deploy2003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 22:22 rzl@deploy2003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 22:21 rzl@deploy2003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 22:19 rzl@deploy2003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 22:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 22:17 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 22:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 22:13 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 22:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 22:13 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 22:09 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 22:06 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 22:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1094.eqiad.wmnet with OS trixie * 22:01 urbanecm: Make https://test.wikipedia.org/w/index.php?title=MediaWiki:GrowthExperimentsSuggestedEdits.json&diff=prev&oldid=750552 with GrowthExperiments disabled (via mw-experimental), then run `\MediaWiki\MediaWikiServices::getInstance()->get('CommunityConfiguration.ProviderFactory')->newProvider('GrowthSuggestedEdits')->getStore()->invalidate()` ([[phab:T431625|T431625]]) * 21:56 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d2-codfw * 21:55 urbanecm@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 21:55 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d2-codfw * 21:55 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c4-codfw * 21:55 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c4-codfw * 21:55 urbanecm@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2002 * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2002 * 21:54 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2002 * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2002.codfw.wmnet 50.32.192.10.in-addr.arpa 0.5.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:54 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2002.codfw.wmnet 50.32.192.10.in-addr.arpa 0.5.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2002 - bking@cumin2003" * 21:54 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2002 - bking@cumin2003" * 21:49 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:49 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2002 * 21:49 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2002.codfw.wmnet with OS trixie * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1094.eqiad.wmnet with reason: host reimage * 21:42 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 21:39 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 21:37 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1094.eqiad.wmnet with reason: host reimage * 21:36 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 21:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 21:29 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 21:27 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 21:22 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1094.eqiad.wmnet with OS trixie * 21:21 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host restbase2039.codfw.wmnet with OS bullseye * 21:21 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin2002" * 21:21 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin2002" * 21:04 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on restbase2039.codfw.wmnet with reason: host reimage * 21:00 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on restbase2039.codfw.wmnet with reason: host reimage * 20:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1073.eqiad.wmnet with OS trixie * 20:48 mutante: deploy2003 - kill 1102 (stunnel4) ; systemctl start stunnel4 ([[phab:T418262|T418262]]) * 20:42 cjming@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] (duration: 33m 02s) * 20:42 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host restbase2039.codfw.wmnet with OS bullseye * 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1073.eqiad.wmnet with reason: host reimage * 20:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1073.eqiad.wmnet with reason: host reimage * 20:30 cjming@deploy2003: cjming: Continuing with deployment * 20:28 cjming@deploy2003: cjming: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1098.eqiad.wmnet with OS trixie * 20:13 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1073.eqiad.wmnet with OS trixie * 20:09 cjming@deploy2003: Started scap sync-world: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] * 20:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1098.eqiad.wmnet with reason: host reimage * 19:56 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1098.eqiad.wmnet with reason: host reimage * 19:55 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d4-codfw * 19:54 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d4-codfw * 19:54 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c1-codfw * 19:54 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c1-codfw * 19:52 mutante: restarting gerrit on gerrit.wikimedia.org (gerrit2003) * 19:48 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2331.codfw.wmnet * 19:48 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2331.codfw.wmnet * 19:48 mutante: restarting gerrit on gerrit-replica.wikimedia.org (gerrit1003) * 19:46 mutante: restarting gerrit on gerrit-spare.wikimedia.org (gerrit2002) * 19:43 jasmine@cumin2002: conftool action : set/pooled=yes; selector: name=wikikube-worker2331.codfw.wmnet,cluster=kubernetes,service=kubesvc * 19:43 jasmine@cumin2002: conftool action : set/weight=10; selector: name=wikikube-worker2331.codfw.wmnet,cluster=kubernetes,service=kubesvc * 19:40 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1098.eqiad.wmnet with OS trixie * 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d5-codfw * 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d5-codfw * 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c7-codfw * 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c7-codfw * 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c5-codfw * 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c5-codfw * 19:30 jasmine_: ran homer on lsw1-d8-codfw, adding wikikube-worker2331 to cluster * 19:29 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1100.eqiad.wmnet with OS trixie * 19:20 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d8-codfw * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d8-codfw * 19:19 mutante: gerrit - replacing private key for registerEmail verification * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-magru * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device cr2-magru * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d7-codfw * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d7-codfw * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d3-codfw * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d1-codfw * 19:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d1-codfw * 19:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c2-codfw * 19:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c2-codfw * 19:11 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-codfw * 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-magru * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device cr1-magru * 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d8-codfw * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d8-codfw * 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d6-codfw * 19:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1100.eqiad.wmnet with reason: host reimage * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d6-codfw * 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c6-codfw * 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c6-codfw * 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c3-codfw * 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c3-codfw * 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b4-magru * 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device asw1-b4-magru * 19:08 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b3-magru * 19:08 cmooney@cumin1003: START - Cookbook sre.network.tls for network device asw1-b3-magru * 19:05 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1100.eqiad.wmnet with reason: host reimage * 19:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1122.eqiad.wmnet with OS trixie * 19:00 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 18:59 topranks: rolling out update to BGP ACL on Nokia Switches eqiad, codfw & ulsfo [[phab:T425703|T425703]] * 18:58 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 18:57 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 18:55 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 18:53 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 18:52 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 18:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1100.eqiad.wmnet with OS trixie * 18:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1068.eqiad.wmnet with OS trixie * 18:47 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1102.eqiad.wmnet with OS trixie * 18:47 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 18:46 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1122.eqiad.wmnet with reason: host reimage * 18:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1122.eqiad.wmnet with reason: host reimage * 18:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1068.eqiad.wmnet with reason: host reimage * 18:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1122 * 18:26 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1122 * 18:25 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1122 * 18:25 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1122.eqiad.wmnet 31.48.64.10.in-addr.arpa 1.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:25 bking@cumin2003: START - Cookbook sre.dns.wipe-cache cirrussearch1122.eqiad.wmnet 31.48.64.10.in-addr.arpa 1.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:25 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:25 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1122 - bking@cumin2003" * 18:25 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1122 - bking@cumin2003" * 18:21 rzl@deploy2003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 18:21 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1068.eqiad.wmnet with reason: host reimage * 18:21 rzl@deploy2003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 18:21 rzl@deploy2003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 18:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 18:19 rzl@deploy2003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 18:19 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:18 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1122 * 18:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1122.eqiad.wmnet with OS trixie * 18:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 18:15 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 18:13 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 18:13 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 18:10 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 18:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1068.eqiad.wmnet with OS trixie * 18:01 kamila@deploy2003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 18m 29s) * 18:00 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:55 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:42 kamila@deploy2003: Started scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] * 17:42 kamila@deploy2003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 19m 50s) * 17:42 kamila@deploy2003: Rolling back deployment * 17:35 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:31 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet * 17:18 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet * 17:16 kamila@deploy1003: Unlocked for deployment [MediaWiki]: switching deployment server (duration: 22m 07s) * 17:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 17:11 kamila@dns1005: END - running authdns-update * 17:09 kamila@dns1005: START - running authdns-update * 17:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 17:04 jasmine@cumin2002: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1164.eqiad.wmnet * 17:04 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1164.eqiad.wmnet * 17:04 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1164.eqiad.wmnet * 16:56 kamila@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on releases2003.codfw.wmnet,releases1003.eqiad.wmnet with reason: Deployment server switchover * 16:54 kamila@deploy1003: Locking from deployment [MediaWiki]: switching deployment server * 16:53 kamila@deploy1003: Unlocked for deployment [MediaWiki]: switching deployment server (duration: 04m 02s) * 16:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie * 16:49 kamila@deploy1003: Locking from deployment [MediaWiki]: switching deployment server * 16:46 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1095.eqiad.wmnet with OS trixie * 16:45 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1093.eqiad.wmnet with OS trixie * 16:43 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1164.eqiad.wmnet with OS trixie * 16:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1095.eqiad.wmnet with reason: host reimage * 16:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 16:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 16:23 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1164.eqiad.wmnet with reason: host reimage * 16:18 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on cirrussearch1093.eqiad.wmnet with reason: host reimage * 16:16 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1164.eqiad.wmnet with reason: host reimage * 16:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1095.eqiad.wmnet with reason: host reimage * 16:09 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 16:09 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 16:08 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1093.eqiad.wmnet with reason: host reimage * 15:59 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Pool test * 15:59 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:59 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 15:59 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Pool test * 15:58 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Depool test * 15:58 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:58 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 15:58 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Depool test * 15:57 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1164 * 15:57 jasmine@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1164 * 15:57 jasmine@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1164 * 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1164.eqiad.wmnet 114.48.64.10.in-addr.arpa 4.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:56 jasmine@cumin2002: START - Cookbook sre.dns.wipe-cache wikikube-worker1164.eqiad.wmnet 114.48.64.10.in-addr.arpa 4.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1164 - jasmine@cumin2002" * 15:56 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1164 - jasmine@cumin2002" * 15:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1093.eqiad.wmnet with OS trixie * 15:51 jasmine@cumin2002: START - Cookbook sre.dns.netbox * 15:51 jasmine@cumin2002: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1164 * 15:50 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-worker1164.eqiad.wmnet with OS trixie * 15:50 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1164.eqiad.wmnet * 15:50 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1164.eqiad.wmnet * 15:50 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1164.eqiad.wmnet * 15:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1095.eqiad.wmnet with OS trixie * 15:42 jasmine@cumin2002: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1164.eqiad.wmnet * 15:42 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1164.eqiad.wmnet * 15:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:42 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1164.eqiad.wmnet * 15:42 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1164.eqiad.wmnet * 15:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 15:39 elukey@cumin1003: START - Cookbook sre.hosts.provision for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 15:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1007.eqiad.wmnet with OS trixie * 15:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Pool test * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1007.eqiad.wmnet with reason: host reimage * 15:15 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 15:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet * 15:15 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 15:15 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1007.eqiad.wmnet with reason: host reimage * 15:15 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:14 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Pool test * 15:14 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet * 15:14 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2228: Depool test * 15:14 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db2228: Depool test * 15:10 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 15:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:08 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 15:06 blake@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 15:06 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet * 15:06 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 15:06 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 15:05 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet * 15:05 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 15:04 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:04 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 15:04 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:03 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 15:03 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 15:03 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:03 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T430909|T430909]] * 15:03 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:03 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 15:03 swfrench-wmf: restarted eqsin, codfw confds - [[phab:T430909|T430909]] * 15:03 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test1002.eqiad.wmnet * 15:02 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet * 14:59 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:59 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:55 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1007.eqiad.wmnet with OS trixie * 14:52 swfrench-wmf: restarted ulsfo confds, confirmed now connected to codfw backends except those using wikimedia.org SRV record - [[phab:T430909|T430909]] * 14:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:49 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:41 moritzm: uninstalling dhcpcd-base from trixie hosts which still have it installed [[phab:T414341|T414341]] * 14:40 sukhe: sudo cumin -b1 -s120 "P<nowiki>{</nowiki>lvs2011*<nowiki>}</nowiki> or P<nowiki>{</nowiki>lvs2012*<nowiki>}</nowiki>" "systemctl restart pybal.service" * 14:39 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:39 mvernon@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host thanos-be1007.eqiad.wmnet with OS trixie * 14:37 sukhe: restart pybal on lvs2013 to revert back to conf2004 * 14:35 sukhe: restart pybal on lvs2014 to revert back to conf2004 * 14:34 swfrench-wmf: switched codfw, eqsin, ulsfo etcd client SRV records back to codfw - [[phab:T430909|T430909]] * 14:32 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1002.eqiad.wmnet * 14:32 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet * 14:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1007.eqiad.wmnet with OS trixie * 14:31 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:31 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:31 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:30 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:30 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Pool test * 14:30 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:29 swfrench@dns1004: END - running authdns-update * 14:29 moritzm: installing jackson-core security updates * 14:27 swfrench@dns1004: START - running authdns-update * 14:22 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:22 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:22 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1119.eqiad.wmnet with OS trixie * 14:22 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:21 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:20 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:20 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 14:20 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:19 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 14:19 moritzm: installing librabbitmq security updates * 14:19 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1002.eqiad.wmnet * 14:18 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet * 14:16 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:15 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:15 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:14 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Pool test * 14:14 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox) * 14:14 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet * 14:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 14:08 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 14:07 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet * 14:05 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1118.eqiad.wmnet with OS trixie * 14:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test1001.eqiad.wmnet * 14:00 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-worker@eqiad * 14:00 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 13:59 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 13:58 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 13:57 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1119.eqiad.wmnet with reason: host reimage * 13:54 moritzm: installing libcap2 security updates * 13:53 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1119.eqiad.wmnet with reason: host reimage * 13:52 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet * 13:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) * 13:52 fceratto@cumin1003: START - Cookbook sre.mysql.depool * 13:50 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-worker@eqiad * 13:50 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1051.eqiad.wmnet * 13:50 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1051.eqiad.wmnet * 13:50 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1051.eqiad.wmnet * 13:49 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 13:45 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1081.eqiad.wmnet with OS trixie * 13:41 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1119 * 13:41 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1119 * 13:40 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1119 * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1119.eqiad.wmnet 97.32.64.10.in-addr.arpa 7.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1119.eqiad.wmnet 97.32.64.10.in-addr.arpa 7.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1119 - atsuko@cumin1003" * 13:40 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1119 - atsuko@cumin1003" * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1118.eqiad.wmnet with reason: host reimage * 13:39 moritzm: installing krb5 security updates * 13:37 Lucas_WMDE: UTC afternoon backport+config window done * 13:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1006.eqiad.wmnet with OS trixie * 13:36 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1118.eqiad.wmnet with reason: host reimage * 13:36 atsuko@cumin1003: START - Cookbook sre.dns.netbox * 13:35 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] (duration: 07m 46s) * 13:34 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1119 * 13:34 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1119.eqiad.wmnet with OS trixie * 13:30 sbisson@deploy1003: sbisson: Continuing with deployment * 13:30 moritzm: installing openssh security updates * 13:30 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling restart_daemons on A:wikidough * 13:29 sbisson@deploy1003: sbisson: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:27 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] * 13:26 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1051.eqiad.wmnet with OS trixie * 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1081.eqiad.wmnet with reason: host reimage * 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1118 * 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1118 * 13:22 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] (duration: 12m 12s) * 13:21 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1081.eqiad.wmnet with reason: host reimage * 13:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1006.eqiad.wmnet with reason: host reimage * 13:18 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1118 * 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1118.eqiad.wmnet 90.32.64.10.in-addr.arpa 0.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:18 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1118.eqiad.wmnet 90.32.64.10.in-addr.arpa 0.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1118 - atsuko@cumin1003" * 13:18 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1118 - atsuko@cumin1003" * 13:17 stran@deploy1003: stran: Continuing with deployment * 13:16 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough * 13:15 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:13 atsuko@cumin1003: START - Cookbook sre.dns.netbox * 13:12 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1118 * 13:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1006.eqiad.wmnet with reason: host reimage * 13:12 stran@deploy1003: stran: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:12 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1118.eqiad.wmnet with OS trixie * 13:10 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] * 13:05 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-worker@codfw * 13:05 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 13:05 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1081.eqiad.wmnet with OS trixie * 13:05 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1051.eqiad.wmnet with reason: host reimage * 13:04 moritzm: installing jq security updates * 13:04 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 13:01 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1051.eqiad.wmnet with reason: host reimage * 12:58 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-worker@codfw * 12:52 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 12:50 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host thanos-be1006.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1051 * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1051 * 12:43 moritzm: installing Python 3.11 security updates * 12:43 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1051 * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1051.eqiad.wmnet 46.32.64.10.in-addr.arpa 6.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:43 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1051.eqiad.wmnet 46.32.64.10.in-addr.arpa 6.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1051 - blake@cumin1003" * 12:43 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1051 - blake@cumin1003" * 12:38 blake@cumin1003: START - Cookbook sre.dns.netbox * 12:38 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1051 * 12:38 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1051.eqiad.wmnet with OS trixie * 12:37 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1051.eqiad.wmnet * 12:36 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1051.eqiad.wmnet * 12:36 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1051.eqiad.wmnet * 12:34 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1006.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 12:34 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1006.eqiad.wmnet with OS trixie * 12:27 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 12:27 mvernon@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host thanos-be1006.eqiad.wmnet with OS trixie * 12:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:02 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 12:01 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1006.eqiad.wmnet with OS trixie * 11:43 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 11:38 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1076.eqiad.wmnet with OS trixie * 11:26 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1075.eqiad.wmnet with OS trixie * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2047.codfw.wmnet * 11:19 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2047.codfw.wmnet * 11:19 moritzm: temporarily remove ganeti2031 from codfw cluster [[phab:T430910|T430910]] * 11:08 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1076.eqiad.wmnet with reason: host reimage * 11:08 moritzm: installing Linux 6.1.176 on Bookworm servers * 11:03 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1076.eqiad.wmnet with reason: host reimage * 11:00 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1075.eqiad.wmnet with reason: host reimage * 10:56 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1075.eqiad.wmnet with reason: host reimage * 10:47 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1076.eqiad.wmnet with OS trixie * 10:46 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1074.eqiad.wmnet with OS trixie * 10:45 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1005.eqiad.wmnet with OS trixie * 10:40 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1075.eqiad.wmnet with OS trixie * 10:32 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2031.codfw.wmnet * 10:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1005.eqiad.wmnet with reason: host reimage * 10:25 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1005.eqiad.wmnet with reason: host reimage * 10:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1074.eqiad.wmnet with reason: host reimage * 10:17 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1074.eqiad.wmnet with reason: host reimage * 10:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1005.eqiad.wmnet with OS trixie * 10:12 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet * 10:04 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 10:01 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet * 10:01 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1074.eqiad.wmnet with OS trixie * 10:01 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 09:43 cgoubert@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/aux-k8s-services/redioscope: apply * 09:43 cgoubert@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/aux-k8s-services/redioscope: apply * 09:43 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply * 09:35 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply * 09:34 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 09:34 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 09:33 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 41 days, 15:00:00 on db2252.codfw.wmnet with reason: Test * 09:32 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 09:32 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 09:31 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: codfw rack B3 pool after maintenance * 09:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 09:07 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 09:07 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 09:02 ladsgroup@cumin1003: END (PASS) - Cookbook sre.mysql.sanitarium_restart (exit_code=0) * 08:57 topranks: merge patch to shift eqiad <-> esams traffic onto new 40G circuit * 08:54 hashar@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1004.eqiad.wmnet with OS trixie * 08:50 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 08:50 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitarium_restart (exit_code=99) * 08:50 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 08:45 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool es2051: codfw rack B3 pool after maintenance * 08:44 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:44 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:43 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2007.codfw.wmnet * 08:43 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2007.codfw.wmnet * 08:42 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2031.codfw.wmnet * 08:41 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2031.codfw.wmnet * 08:40 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2031.codfw.wmnet * 08:38 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:38 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:35 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.sanitize-wiki (exit_code=97) Managing sanitization for wikis minwikiquote in section s3 * 08:33 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis minwikiquote in section s3 * 08:32 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Checking sanitization for wikis minwikiquote in section s5 * 08:30 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Checking sanitization for wikis minwikiquote in section s5 * 08:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Managing sanitization for wikis minwikiquote in section s5 * 08:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1004.eqiad.wmnet with reason: host reimage * 08:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1004.eqiad.wmnet with reason: host reimage * 08:23 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:22 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis minwikiquote in section s5 * 08:19 XioNoX: lsw1-b3-codfw> request system reboot - [[phab:T430909|T430909]] * 08:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Checking sanitization for wikis minwikiquote in section s5 * 08:17 hashar@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Checking sanitization for wikis minwikiquote in section s5 * 08:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for codfw rack B3 * 08:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2007.codfw.wmnet * 08:15 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lsw1-b3-codfw,lsw1-b3-codfw IPv6,lsw1-b3-codfw.mgmt with reason: Switch maintenance * 08:15 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2007.codfw.wmnet * 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:07 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:06 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: codfw rack B3 depool for maintenance * 08:05 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool es2051: codfw rack B3 depool for maintenance * 08:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1004.eqiad.wmnet with OS trixie * 08:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1005.eqiad.wmnet with OS trixie * 08:03 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 21 hosts with reason: codfw rack B3 depool for maintenance * 07:56 ayounsi@cumin1003: START - Cookbook sre.network.depool-rack with action 'depool' for codfw rack B3 * 07:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1005.eqiad.wmnet with reason: host reimage * 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1005.eqiad.wmnet with reason: host reimage * 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1005.eqiad.wmnet with OS bookworm * 07:29 moritzm: installing gnutls28 security updates * 07:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1005.eqiad.wmnet with OS trixie * 07:13 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1125.eqiad.wmnet with OS trixie * 07:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1005.eqiad.wmnet with reason: host reimage * 07:07 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aux-k8s-etcd1005.eqiad.wmnet with reason: host reimage * 06:56 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1005.eqiad.wmnet with OS bookworm * 06:54 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1125.eqiad.wmnet with reason: host reimage * 06:52 elukey: upgrade all trixie hosts to pywmflib 3.1 - [[phab:T430552|T430552]] * 06:50 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1125.eqiad.wmnet with reason: host reimage * 06:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 06:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 06:38 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1125.eqiad.wmnet with OS trixie * 05:42 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1107.eqiad.wmnet with OS trixie * 05:35 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1124.eqiad.wmnet with OS trixie * 05:31 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1101.eqiad.wmnet with OS trixie * 05:21 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1107.eqiad.wmnet with reason: host reimage * 05:17 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1124.eqiad.wmnet with reason: host reimage * 05:13 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1107.eqiad.wmnet with reason: host reimage * 05:13 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1101.eqiad.wmnet with reason: host reimage * 05:11 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1124.eqiad.wmnet with reason: host reimage * 05:10 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1101.eqiad.wmnet with reason: host reimage * 04:58 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1124.eqiad.wmnet with OS trixie * 04:56 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1107.eqiad.wmnet with OS trixie * 04:55 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1101.eqiad.wmnet with OS trixie * 02:27 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] (duration: 08m 14s) * 02:22 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 02:21 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 02:19 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] * 01:59 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] (duration: 09m 46s) * 01:55 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 01:51 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 01:49 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] * 01:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1099.eqiad.wmnet with OS trixie * 00:57 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1110.eqiad.wmnet with OS trixie * 00:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1099.eqiad.wmnet with reason: host reimage * 00:41 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1099.eqiad.wmnet with reason: host reimage * 00:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1110.eqiad.wmnet with reason: host reimage * 00:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1110.eqiad.wmnet with reason: host reimage * 00:26 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1099.eqiad.wmnet with OS trixie * 00:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1110.eqiad.wmnet with OS trixie == 2026-07-07 == * 22:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1097.eqiad.wmnet with OS trixie * 22:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1097.eqiad.wmnet with reason: host reimage * 22:24 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1097.eqiad.wmnet with reason: host reimage * 22:14 hashar: Restarting Gerrit on gerrit2002 and gerrit1003 (replicas) * 22:09 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1097.eqiad.wmnet with OS trixie * 22:07 hashar: Restarting Gerrit on gerrit2003 * 21:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 21:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 21:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 21:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 21:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1108.eqiad.wmnet with OS trixie * 20:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1091.eqiad.wmnet with OS trixie * 20:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1090.eqiad.wmnet with OS trixie * 20:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1108.eqiad.wmnet with reason: host reimage * 20:36 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1091.eqiad.wmnet with reason: host reimage * 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1090.eqiad.wmnet with reason: host reimage * 20:33 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1091.eqiad.wmnet with reason: host reimage * 20:30 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1108.eqiad.wmnet with reason: host reimage * 20:30 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1006.eqiad.wmnet * 20:30 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1090.eqiad.wmnet with reason: host reimage * 20:30 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1006.eqiad.wmnet * 20:27 jasmine_: "homer lsw1-c2-eqiad* commit "Added new stacked control plane wikikube-ctrl1006"" * 20:22 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] (duration: 07m 29s) * 20:20 jasmine_: "homer "cr*eqiad*" commit "Added new stacked control plane wikikube-ctrl1006"" * 20:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1091.eqiad.wmnet with OS trixie * 20:17 arlolra@deploy1003: arlolra: Continuing with deployment * 20:16 arlolra@deploy1003: arlolra: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:16 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1090.eqiad.wmnet with OS trixie * 20:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1108.eqiad.wmnet with OS trixie * 20:14 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] * 20:09 cwhite: remove 2026-04 swift log archives from centrallog2002 to free some space * 20:01 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=93) for host cirrussearch1108.eqiad.wmnet with OS trixie * 19:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1108.eqiad.wmnet with OS trixie * 19:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1090.eqiad.wmnet with OS trixie * 19:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1109.eqiad.wmnet with OS trixie * 19:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1092.eqiad.wmnet with OS trixie * 19:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1123.eqiad.wmnet with OS trixie * 19:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1109.eqiad.wmnet with reason: host reimage * 19:19 cdobbins@cumin2002: conftool action : set/pooled=yes; selector: name=dns7002.* * 19:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1092.eqiad.wmnet with reason: host reimage * 19:17 jasmine@dns1004: END - running authdns-update * 19:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1109.eqiad.wmnet with reason: host reimage * 19:15 jasmine@dns1004: START - running authdns-update * 19:15 cdobbins@dns1004: END - running authdns-update * 19:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1123.eqiad.wmnet with reason: host reimage * 19:13 cdobbins@dns1004: START - running authdns-update * 19:12 cdobbins@cumin2002: conftool action : set/pooled=yes; selector: name=dns7002.*,service=authdns-update * 19:11 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1092.eqiad.wmnet with reason: host reimage * 19:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1123.eqiad.wmnet with reason: host reimage * 18:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1123.eqiad.wmnet with OS trixie * 18:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1109.eqiad.wmnet with OS trixie * 18:56 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1092.eqiad.wmnet with OS trixie * 18:52 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 18:49 swfrench@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 18:40 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 18:38 swfrench@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 18:11 swfrench-wmf: restarted eqsin, codfw confds - [[phab:T430909|T430909]] * 18:01 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T430909|T430909]] * 17:59 swfrench-wmf: restarted ulsfo confds, confirmed now connected to eqiad backends - [[phab:T430909|T430909]] * 17:52 sukhe: restart pybal on lvs2011 to switch from conf2004 to conf1008: [[phab:T430909|T430909]] * 17:51 sukhe: restart pybal on lvs2012 to switch from conf2004 to conf1008 [puppet re-enabled there]: [[phab:T430909|T430909]] * 17:46 sukhe: restart pybal on lvs2013 to switch from conf2004 to conf1008: [[phab:T430909|T430909]] * 17:44 swfrench-wmf: switched codfw, eqsin, ulsfo etcd client SRV records to eqiad - [[phab:T430909|T430909]] * 17:43 swfrench@dns1004: END - running authdns-update * 17:40 swfrench@dns1004: START - running authdns-update * 17:40 sukhe: restart pybal on lvs2014 to switch from conf2004 to conf1008: [[phab:T430909|T430909]] * 17:21 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1003.eqiad.wmnet * 17:15 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1003.eqiad.wmnet * 17:14 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1002.eqiad.wmnet * 17:06 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1002.eqiad.wmnet * 17:06 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-low-traffic-codfw' 'systemctl restart pybal.service' # lvs2013, [[phab:T416623|T416623]] * 17:04 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1001.eqiad.wmnet * 17:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1111.eqiad.wmnet with OS trixie * 17:00 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal.service' # lvs2014, [[phab:T416623|T416623]] * 16:58 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1001.eqiad.wmnet * 16:58 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS bookworm * 16:55 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-low-traffic-eqiad' 'systemctl restart pybal.service' # lvs1019, [[phab:T416623|T416623]] * 16:53 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-secondary-eqiad' 'systemctl restart pybal.service' # lvs1020, [[phab:T416623|T416623]] * 16:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1111.eqiad.wmnet with reason: host reimage * 16:40 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1111.eqiad.wmnet with reason: host reimage * 16:38 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.peering (exit_code=99) with action 'configure' for AS: 47794 * 16:35 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 47794 * 16:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1111.eqiad.wmnet with OS trixie * 16:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1006.eqiad.wmnet with OS trixie * 16:06 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 16:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1006.eqiad.wmnet with reason: host reimage * 15:58 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1121.eqiad.wmnet with OS trixie * 15:58 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 15:56 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1006.eqiad.wmnet with reason: host reimage * 15:54 mutante: jenkins down in planned maintenance window * 15:42 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1037.eqiad.wmnet * 15:42 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1037.eqiad.wmnet * 15:42 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1037.eqiad.wmnet * 15:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1006.eqiad.wmnet with OS trixie * 15:34 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1121.eqiad.wmnet with reason: host reimage * 15:33 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS bookworm * 15:33 cdobbins@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host dns7002.wikimedia.org with OS trixie * 15:30 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1121.eqiad.wmnet with reason: host reimage * 15:29 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host clouddumps1001.wikimedia.org * 15:20 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1001.wikimedia.org * 15:18 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1121 * 15:18 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1121 * 15:18 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host clouddumps1002.wikimedia.org * 15:17 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1121 * 15:17 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:17 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply * 15:16 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply * 15:16 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:16 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1121 - atsuko@cumin1003" * 15:16 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1121 - atsuko@cumin1003" * 15:14 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1037.eqiad.wmnet with OS trixie * 15:11 atsuko@cumin1003: START - Cookbook sre.dns.netbox * 15:09 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org * 15:09 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1121 * 15:09 andrew@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host clouddumps1002.wikimedia.org * 15:09 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org * 15:09 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1121.eqiad.wmnet with OS trixie * 15:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1007.eqiad.wmnet with OS trixie * 15:08 andrew@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host clouddumps1002.wikimedia.org * 15:08 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org * 15:05 brennen@deploy1003: Finished deploy [phabricator/deployment@7e02037]: deploy phab1004 for [[phab:T431440|T431440]] (duration: 00m 47s) * 15:04 brennen@deploy1003: Started deploy [phabricator/deployment@7e02037]: deploy phab1004 for [[phab:T431440|T431440]] * 15:03 brennen@deploy1003: Finished deploy [phabricator/deployment@7e02037]: deploy phab2003 for [[phab:T431440|T431440]] (duration: 00m 51s) * 15:03 brennen@deploy1003: Started deploy [phabricator/deployment@7e02037]: deploy phab2003 for [[phab:T431440|T431440]] * 15:00 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71] (thin): Regular analytics weekly train THIN [analytics/refinery@7d8dc71f] (duration: 02m 10s) * 14:58 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71] (thin): Regular analytics weekly train THIN [analytics/refinery@7d8dc71f] * 14:58 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71]: Regular analytics weekly train [analytics/refinery@7d8dc71f] (duration: 04m 14s) * 14:54 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1037.eqiad.wmnet with reason: host reimage * 14:53 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71]: Regular analytics weekly train [analytics/refinery@7d8dc71f] * 14:53 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@7d8dc71f] (duration: 02m 00s) * 14:51 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@7d8dc71f] * 14:51 arnaudb@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on phab2003.codfw.wmnet,phab[1004-1006].eqiad.wmnet with reason: maintenance * 14:51 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1037.eqiad.wmnet with reason: host reimage * 14:50 JavierMonton: Deploying Refinery at {{Gerrit|7d8dc71f}} for change {{Gerrit|1308087}} / [[phab:T431318|T431318]] - update filerevision table sqoop and table * 14:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1007.eqiad.wmnet with reason: host reimage * 14:42 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1007.eqiad.wmnet with reason: host reimage * 14:40 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1083.eqiad.wmnet with OS trixie * 14:37 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:36 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:35 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-master@eqiad * 14:35 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 14:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 14:34 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1037 * 14:34 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1037 * 14:34 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:cleanMentorList.php --wiki=frwiki # [[phab:T427386|T427386]] * 14:34 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 14:34 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308112{{!}}Revert^2 "[Growth] frwiki: Deploy automated mentor list cleaner" (T427386)]] (duration: 06m 47s) * 14:34 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 14:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:33 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:32 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1037 * 14:31 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:31 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:29 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:29 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-master@eqiad * 14:29 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:29 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:29 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:28 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:27 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:27 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1308112{{!}}Revert^2 "[Growth] frwiki: Deploy automated mentor list cleaner" (T427386)]] * 14:26 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1007.eqiad.wmnet with OS trixie * 14:26 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:26 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 14:26 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:25 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:25 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:25 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:cleanMentorList.php --wiki=frwiki # [[phab:T427386|T427386]] * 14:24 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:24 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1037 - blake@cumin1003" * 14:24 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1037 - blake@cumin1003" * 14:20 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1083.eqiad.wmnet with reason: host reimage * 14:19 blake@cumin1003: START - Cookbook sre.dns.netbox * 14:19 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-master@codfw * 14:19 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 14:19 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1037 * 14:18 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1037.eqiad.wmnet with OS trixie * 14:18 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1037.eqiad.wmnet * 14:18 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 14:18 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1037.eqiad.wmnet * 14:18 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1037.eqiad.wmnet * 14:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2007.codfw.wmnet with OS trixie * 14:16 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1083.eqiad.wmnet with reason: host reimage * 14:15 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1036.eqiad.wmnet * 14:15 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1036.eqiad.wmnet * 14:14 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1036.eqiad.wmnet * 14:12 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-master@codfw * 14:11 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1120.eqiad.wmnet with OS trixie * 14:05 moritzm: installing distro-info-data updates from trixie/bookworm point releases * 14:04 fabfur: disable puppet on A:cp-text to selectively apply https://gerrit.wikimedia.org/r/c/operations/puppet/+/1308040 * 14:03 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] (duration: 27m 48s) * 14:00 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 14:00 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1083.eqiad.wmnet with OS trixie * 13:58 urbanecm@deploy1003: urbanecm: Continuing with deployment * 13:58 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:58 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1004.eqiad.wmnet with OS bookworm * 13:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2007.codfw.wmnet with reason: host reimage * 13:57 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 13:53 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1120.eqiad.wmnet with reason: host reimage * 13:50 moritzm: installing Linux 5.10.259 on Bullseye hosts * 13:47 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply * 13:47 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply * 13:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2007.codfw.wmnet with reason: host reimage * 13:46 cgoubert@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/aux-k8s-services/redioscope: apply * 13:46 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1120.eqiad.wmnet with reason: host reimage * 13:46 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:46 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:45 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:44 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:44 cgoubert@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/aux-k8s-services/redioscope: apply * 13:40 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 13:39 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:38 moritzm: installing e2fsprogs updates from Trixie point release * 13:35 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] * 13:33 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1120.eqiad.wmnet with OS trixie * 13:33 topranks: reset cr3-eqsin configuration so traffic uses it again after upgrade * 13:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1088.eqiad.wmnet with OS trixie * 13:32 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie * 13:32 cdobbins@cumin1003: conftool action : set/pooled=no; selector: name=dns7002.* * 13:29 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2007.codfw.wmnet with OS trixie * 13:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1004.eqiad.wmnet with reason: host reimage * 13:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2006.codfw.wmnet with OS trixie * 13:18 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1036.eqiad.wmnet with OS trixie * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aux-k8s-etcd1004.eqiad.wmnet with reason: host reimage * 13:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 13:16 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 13:15 jayme: Istio is being upgraded from 1.24.2 to 1.29.4 on wikikube staging eqiad and codfw - [[phab:T427401|T427401]] * 13:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1087.eqiad.wmnet with OS trixie * 13:14 topranks: reboot cr3-eqsin to install new JunOS and set PIC 0/0/0 to 100G * 13:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1088.eqiad.wmnet with reason: host reimage * 13:13 jmm@dns1004: END - running authdns-update * 13:12 jmm@dns1004: START - running authdns-update * 13:09 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1088.eqiad.wmnet with reason: host reimage * 13:07 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1082.eqiad.wmnet with OS trixie * 13:07 jmm@dns1004: END - running authdns-update * 13:06 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1004.eqiad.wmnet with OS bookworm * 13:05 jmm@dns1004: START - running authdns-update * 13:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2006.codfw.wmnet with reason: host reimage * 12:58 topranks: load updated JunOS on cr3-eqsin [[phab:T429386|T429386]] * 12:58 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1036.eqiad.wmnet with reason: host reimage * 12:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2001.codfw.wmnet * 12:57 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2006.codfw.wmnet with reason: host reimage * 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr1-codfw,cr[2-3]-eqsin,cr3-eqsin IPv6,cr3-eqsin.mgmt with reason: upgrade JunOS cr3-eqsin * 12:56 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lvs[5004-5006].eqsin.wmnet with reason: upgrade JunOS cr3-eqsin * 12:55 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 12:55 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 12:53 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1087.eqiad.wmnet with reason: host reimage * 12:53 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 12:52 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1088.eqiad.wmnet with OS trixie * 12:52 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1002.eqiad.wmnet * 12:52 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:51 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2001.codfw.wmnet * 12:49 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1036.eqiad.wmnet with reason: host reimage * 12:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 12:48 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1087.eqiad.wmnet with reason: host reimage * 12:44 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1082.eqiad.wmnet with reason: host reimage * 12:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1002.eqiad.wmnet * 12:42 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 12:42 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 12:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 12:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 12:39 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:39 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: move dumps-nfs IP to the shared one - filippo@cumin1003" * 12:39 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: move dumps-nfs IP to the shared one - filippo@cumin1003" * 12:39 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2006.codfw.wmnet with OS trixie * 12:38 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1082.eqiad.wmnet with reason: host reimage * 12:36 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:33 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:32 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1036 * 12:32 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1036 * 12:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2005.codfw.wmnet with OS trixie * 12:32 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1087.eqiad.wmnet with OS trixie * 12:30 jmm@dns1004: END - running authdns-update * 12:29 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1036 * 12:29 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1036.eqiad.wmnet 21.32.64.10.in-addr.arpa 1.2.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:29 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1036.eqiad.wmnet 21.32.64.10.in-addr.arpa 1.2.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:29 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:29 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1036 - blake@cumin1003" * 12:29 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1036 - blake@cumin1003" * 12:28 jmm@dns1004: START - running authdns-update * 12:26 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:26 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:23 blake@cumin1003: START - Cookbook sre.dns.netbox * 12:23 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1036 * 12:23 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1036.eqiad.wmnet with OS trixie * 12:22 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1036.eqiad.wmnet * 12:22 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1082.eqiad.wmnet with OS trixie * 12:22 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1036.eqiad.wmnet * 12:22 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1036.eqiad.wmnet * 12:21 marostegui: Restart mariadb@s7 on db1155 to pick up new filters - [[phab:T431124|T431124]] * 12:21 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 21 hosts with reason: restarting for replication filter * 12:20 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:19 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:14 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2005.codfw.wmnet with reason: host reimage * 12:14 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:08 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:07 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:07 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2005.codfw.wmnet with reason: host reimage * 12:06 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:06 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:06 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:05 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:05 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:04 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-master-eqiad * 12:04 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl1002.eqiad.wmnet * 12:04 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl1002.eqiad.wmnet * 12:04 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:04 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:03 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:03 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:03 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:03 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 11:59 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl1002.eqiad.wmnet * 11:59 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl1002.eqiad.wmnet * 11:59 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl1001.eqiad.wmnet * 11:59 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl1001.eqiad.wmnet * 11:56 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl1001.eqiad.wmnet * 11:56 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl1001.eqiad.wmnet * 11:56 klausman@cumin2002: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-master-eqiad * 11:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2005.codfw.wmnet with OS trixie * 11:49 blake@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on wikikube-worker1160.eqiad.wmnet with reason: Verifying matchers for silence * 11:42 topranks: cr3-eqsin, begin traffic drain to reset PIC and upgrade JunOS [[phab:T429386|T429386]] * 11:41 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs[5004-5006].eqsin.wmnet with reason: upgrade JunOS cr3-eqsin * 11:39 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr1-codfw,cr[2-3]-eqsin,cr3-eqsin IPv6,cr3-eqsin.mgmt with reason: upgrade JunOS cr3-eqsin * 11:36 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=thanos-fe2004.codfw.wmnet * 11:35 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1086.eqiad.wmnet with OS trixie * 11:35 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=thanos-fe2004.codfw.wmnet * 11:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1085.eqiad.wmnet with OS trixie * 11:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1086.eqiad.wmnet with reason: host reimage * 11:10 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1085.eqiad.wmnet with reason: host reimage * 11:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2004.codfw.wmnet with OS trixie * 11:03 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1086.eqiad.wmnet with reason: host reimage * 11:02 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1085.eqiad.wmnet with reason: host reimage * 10:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2004.codfw.wmnet with reason: host reimage * 10:48 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:46 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1086.eqiad.wmnet with OS trixie * 10:46 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1085.eqiad.wmnet with OS trixie * 10:44 cgoubert@deploy1003: Finished deploy [restbase/deploy@2fc37d4]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] (duration: 16m 44s) * 10:43 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2004.codfw.wmnet with reason: host reimage * 10:35 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:27 cgoubert@deploy1003: Started deploy [restbase/deploy@2fc37d4]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] * 10:27 cgoubert@deploy1003: Finished deploy [restbase/deploy@8a25036]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] (duration: 00m 45s) * 10:26 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1117.eqiad.wmnet with OS trixie * 10:26 cgoubert@deploy1003: Started deploy [restbase/deploy@8a25036]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] * 10:26 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host thanos-fe2004 * 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host thanos-fe2004 * 10:22 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1116.eqiad.wmnet with OS trixie * 10:21 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host thanos-fe2004 * 10:21 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) thanos-fe2004.codfw.wmnet 157.32.192.10.in-addr.arpa 7.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:20 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache thanos-fe2004.codfw.wmnet 157.32.192.10.in-addr.arpa 7.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:20 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:20 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host thanos-fe2004 - mvernon@cumin2003" * 10:20 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host thanos-fe2004 - mvernon@cumin2003" * 10:15 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2252: Repooling after reboot * 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:15 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 10:15 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2252: Repooling after reboot * 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1153.eqiad.wmnet * 10:14 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1153.eqiad.wmnet * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 10:14 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 10:12 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 10:12 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host thanos-fe2004 * 10:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2004.codfw.wmnet with OS trixie * 10:07 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1117.eqiad.wmnet with reason: host reimage * 10:03 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1116.eqiad.wmnet with reason: host reimage * 09:58 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:58 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1117.eqiad.wmnet with reason: host reimage * 09:57 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1116.eqiad.wmnet with reason: host reimage * 09:49 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 41 days, 15:00:00 on db2252.codfw.wmnet with reason: Security updates * 09:45 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1117.eqiad.wmnet with OS trixie * 09:45 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1116.eqiad.wmnet with OS trixie * 09:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1153: Security updates * 09:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:28 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:28 root@cumin1003: START - Cookbook sre.mysql.depool depool db1153: Security updates * 09:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1016: Security updates * 09:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:21 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:21 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1016: Security updates * 09:14 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:14 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 08:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1016: Security updates * 08:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:56 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:56 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1016: Security updates * 08:50 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:50 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:45 filippo@dns1006: END - running authdns-update * 08:43 filippo@dns1006: START - running authdns-update * 08:42 godog: switch dumps-nfs address to be shared with rsync/http - [[phab:T411248|T411248]] * 08:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1016: Security updates * 08:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:40 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:40 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1016: Security updates * 08:29 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host cirrussearch1111.eqiad.wmnet * 08:29 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:27 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:27 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:25 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1015: Security updates * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:09 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:09 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1015: Security updates * 07:42 Msz2001: Deployed private patch for Suggested Ivestigations * 07:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1015: Security updates * 07:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:41 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:41 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1015: Security updates * 07:40 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 07:11 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fingerprint warnings - oblivian@cumin1003" * 07:11 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fingerprint warnings - oblivian@cumin1003 * 07:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1024: Security updates * 07:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:11 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:11 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1024: Security updates * 07:10 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fingerprint warnings - oblivian@cumin1003 * 07:10 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fingerprint warnings - oblivian@cumin1003" * 06:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host cirrussearch1111.eqiad.wmnet * 06:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 06:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1024: Security updates * 06:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 06:48 root@cumin1003: START - Cookbook sre.mysql.parsercache * 06:48 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1024: Security updates * 06:42 moritzm: install nginx security updates * 06:31 root@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool pc1024: Security updates * 06:21 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1024: Security updates * 06:19 moritzm: installing php8.2 security updates * 06:15 moritzm: installing php8.4 security updates * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.7 (duration: 02m 38s) * 03:40 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] (duration: 37m 04s) * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 51s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-06 == * 23:30 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] (duration: 09m 39s) * 23:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1078.eqiad.wmnet with OS trixie * 23:26 jdlrobson@deploy1003: jdlrobson, bwang: Continuing with deployment * 23:22 jdlrobson@deploy1003: jdlrobson, bwang: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug) * 23:21 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] * 23:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1078.eqiad.wmnet with reason: host reimage * 23:06 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1078.eqiad.wmnet with reason: host reimage * 22:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1078.eqiad.wmnet with OS trixie * 22:29 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on cirrussearch1114.eqiad.wmnet with reason: reimage on hold until restore completes * 22:22 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on cirrussearch[1079,1115].eqiad.wmnet with reason: reimage on hold until restore completes * 21:18 maryum: Deployed security fix for [[phab:T428006|T428006]] * 20:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1079.eqiad.wmnet with OS trixie * 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1077.eqiad.wmnet with OS trixie * 20:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1115.eqiad.wmnet with OS trixie * 20:25 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1079.eqiad.wmnet with reason: host reimage * 20:21 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1079.eqiad.wmnet with reason: host reimage * 20:15 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] (duration: 08m 14s) * 20:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1077.eqiad.wmnet with reason: host reimage * 20:10 krinkle@deploy1003: krinkle, pushpaktiwari: Continuing with deployment * 20:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1115.eqiad.wmnet with reason: host reimage * 20:08 krinkle@deploy1003: krinkle, pushpaktiwari: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1077.eqiad.wmnet with reason: host reimage * 20:06 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] * 20:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1079.eqiad.wmnet with OS trixie * 20:04 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1115.eqiad.wmnet with reason: host reimage * 19:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1077.eqiad.wmnet with OS trixie * 19:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1115.eqiad.wmnet with OS trixie * 19:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 19:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 18:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1114.eqiad.wmnet with OS trixie * 18:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1114.eqiad.wmnet with reason: host reimage * 18:35 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1114.eqiad.wmnet with reason: host reimage * 18:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1112.eqiad.wmnet with OS trixie * 18:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1114.eqiad.wmnet with OS trixie * 18:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1072.eqiad.wmnet with OS trixie * 18:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1112.eqiad.wmnet with reason: host reimage * 18:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1112.eqiad.wmnet with reason: host reimage * 17:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1072.eqiad.wmnet with reason: host reimage * 17:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1112.eqiad.wmnet with OS trixie * 17:55 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1072.eqiad.wmnet with reason: host reimage * 17:39 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1072.eqiad.wmnet with OS trixie * 17:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1071.eqiad.wmnet with OS trixie * 17:18 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1070.eqiad.wmnet with OS trixie * 17:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1084.eqiad.wmnet with OS trixie * 16:54 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1071.eqiad.wmnet with reason: host reimage * 16:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1084.eqiad.wmnet with reason: host reimage * 16:51 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1070.eqiad.wmnet with reason: host reimage * 16:49 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1084.eqiad.wmnet with reason: host reimage * 16:38 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1071.eqiad.wmnet with OS trixie * 16:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1096.eqiad.wmnet with OS trixie * 16:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1070.eqiad.wmnet with OS trixie * 16:33 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1084.eqiad.wmnet with OS trixie * 16:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1089.eqiad.wmnet with OS trixie * 16:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1103.eqiad.wmnet with OS trixie * 16:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1096.eqiad.wmnet with reason: host reimage * 16:14 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1096.eqiad.wmnet with reason: host reimage * 16:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1089.eqiad.wmnet with reason: host reimage * 16:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1103.eqiad.wmnet with reason: host reimage * 16:02 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1003.eqiad.wmnet with OS bookworm * 16:01 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1089.eqiad.wmnet with reason: host reimage * 16:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1103.eqiad.wmnet with reason: host reimage * 15:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1096.eqiad.wmnet with OS trixie * 15:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1080.eqiad.wmnet with OS trixie * 15:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1089.eqiad.wmnet with OS trixie * 15:45 dancy@deploy1003: Installation of scap version "4.272.0" completed for 158 hosts * 15:43 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1103.eqiad.wmnet with OS trixie * 15:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1113.eqiad.wmnet with OS trixie * 15:41 dancy@deploy1003: Installing scap version "4.272.0" for 158 host(s) * 15:40 klausman@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 15:39 klausman@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 15:38 klausman@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 15:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1069.eqiad.wmnet with OS trixie * 15:37 klausman@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 15:36 klausman@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 15:34 klausman@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 15:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1080.eqiad.wmnet with reason: host reimage * 15:27 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1080.eqiad.wmnet with reason: host reimage * 15:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1113.eqiad.wmnet with reason: host reimage * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1069.eqiad.wmnet with reason: host reimage * 15:18 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1113.eqiad.wmnet with reason: host reimage * 15:16 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1069.eqiad.wmnet with reason: host reimage * 15:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:11 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1080.eqiad.wmnet with OS trixie * 15:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1113.eqiad.wmnet with OS trixie * 15:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1003.eqiad.wmnet with reason: host reimage * 14:47 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1003.eqiad.wmnet with OS bookworm * 14:33 elukey: rolled out spicerack on all cumin nodes - [[phab:T429699|T429699]] * 14:32 elukey: upgrade all bookworm hosts to pywmflib 3.1 - [[phab:T430552|T430552]] * 14:14 marostegui: Setup x4 eqiad topology [[phab:T404715|T404715]] * 14:13 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 14:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2230.codfw.wmnet * 14:07 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2230.codfw.wmnet * 13:59 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[2001-2002].codfw.wmnet * 13:51 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 13:45 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.major-upgrade (exit_code=97) * 13:45 cwilliams@cumin1003: dbmaint on s4@codfw [[phab:T429893|T429893]] * 13:45 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 13:42 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-master-codfw * 13:42 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl2002.codfw.wmnet * 13:42 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl2002.codfw.wmnet * 13:38 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl2002.codfw.wmnet * 13:38 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl2002.codfw.wmnet * 13:38 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl2001.codfw.wmnet * 13:38 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl2001.codfw.wmnet * 13:35 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl2001.codfw.wmnet * 13:35 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl2001.codfw.wmnet * 13:35 klausman@cumin2002: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-master-codfw * 12:30 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] (duration: 25m 11s) * 12:24 krinkle@deploy1003: krinkle: Continuing with deployment * 12:10 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2048.codfw.wmnet * 12:09 krinkle@deploy1003: krinkle: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:08 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2048.codfw.wmnet * 12:05 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] * 11:57 moritzm: installing curl security updates * 11:49 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:31 moritzm: installing nano security updates * 11:07 moritzm: failover Ganeti master in codfw to ganeti2032 [[phab:T430909|T430909]] * 11:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:04 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest1005.eqiad.wmnet with OS trixie * 11:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:50 jmm@dns1004: END - running authdns-update * 10:47 jmm@dns1004: START - running authdns-update * 10:47 jmm@dns1004: START - running authdns-update * 10:46 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:44 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest1005.eqiad.wmnet with reason: host reimage * 10:38 elukey@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest1005.eqiad.wmnet with reason: host reimage * 10:31 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:31 marostegui: Setup x4 codfw topology [[phab:T404715|T404715]] * 10:31 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 10:24 elukey: spicerack 13.0.0 deployed on cumin2002 * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 10:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 10:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 10:21 elukey@cumin2002: START - Cookbook sre.hosts.reimage for host sretest1005.eqiad.wmnet with OS trixie * 10:20 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:19 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:17 elukey: uploaded spicerack_13.0.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia * 09:54 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:52 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:20 elukey: upgrade all bullseye hosts to pywmflib 3.1 - [[phab:T430552|T430552]] * 09:10 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1015.eqiad.wmnet,service=s4 * 09:10 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1015.eqiad.wmnet,service=s6 * 09:07 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 08:58 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:56 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 08:56 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 08:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 08:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 08:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin2002.codfw.wmnet * 08:06 godog: remove cloudvirt1046, cloudvirt1062, cloudvirt1074, cloudvirt1075 from maintenance aggregate and put them in network-ovs - [[phab:T424802|T424802]] * 08:00 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin2002.codfw.wmnet * 07:58 hashar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] (duration: 32m 53s) * 07:57 fabfur: repooled cp4038 * 07:57 fabfur@cumin1003: conftool action : set/pooled=yes; selector: name=cp4038.* * 07:53 moritzm: installing pyjwt security updates * 07:47 moritzm: installing openjpeg2 security updates * 07:45 hashar@deploy1003: vadymts1, hashar: Continuing with deployment * 07:43 hashar@deploy1003: vadymts1, hashar: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:38 moritzm: installing python-urllib3 security updates * 07:37 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 07:30 fabfur: depooled cp4038 to investigate on possible maxmind failure * 07:30 fabfur@cumin1003: conftool action : set/pooled=no; selector: name=cp4038.* * 07:30 fabfur@cumin1003: conftool action : set/pooled=yes; selector: name=cp4038.* * 07:29 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 07:25 hashar@deploy1003: Started scap sync-world: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] * 06:13 moritzm: installing Linux 6.12.95 on trixie hosts * 05:20 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s6 * 05:20 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s4 * 05:19 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1015.eqiad.wmnet with reason: cloning * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 08s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-05 == * 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 01m 08s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-04 == * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 58s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-03 == * 17:08 topranks: revert protocol preference changes on cr3-ulsfo after upgrade * 16:53 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on cr2-eqord with reason: upgrade JunOS cr3-ulsfo * 16:53 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on cr4-ulsfo with reason: upgrade JunOS cr3-ulsfo * 16:48 topranks: reboot cr3-ulsfo to upgrade JunOS and reset linecard [[phab:T424839|T424839]] * 15:52 topranks: adjust outbound BGP policies on cr3-ulsfo to drain router of traffic [[phab:T424839|T424839]] * 15:45 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on lvs[4008-4010].ulsfo.wmnet with reason: upgrade JunOS cr3-ulsfo * 15:44 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on asw1-[22-23]-ulsfo,cr3-ulsfo,cr3-ulsfo IPv6,cr3-ulsfo.mgmt with reason: upgrade JunOS cr3-ulsfo * 15:36 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 15:35 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 15:35 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 14:40 cmooney@dns3003: END - running authdns-update * 14:26 cmooney@dns3003: START - running authdns-update * 14:26 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:26 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to ulsfo - cmooney@cumin1003" * 14:19 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to ulsfo - cmooney@cumin1003" * 14:16 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:38 sukhe@dns1004: END - running authdns-update * 13:35 sukhe@dns1004: START - running authdns-update * 13:26 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 13:26 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 13:26 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet * 13:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 13:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 13:16 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host sretest1005.eqiad.wmnet * 13:16 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 13:16 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 13:15 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 13:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:14 moritzm: imported samplicator 1.3.8rc1-1+deb13u1 to trixie-wikimedia/main [[phab:T337208|T337208]] * 13:13 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:07 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 13:07 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 13:02 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:02 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:58 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:57 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:57 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:53 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet * 12:50 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 12:47 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:41 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:40 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:39 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:32 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet * 12:26 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet * 12:23 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2005.wikimedia.org * 12:19 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2005.wikimedia.org * 12:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 12:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup[2004-2007].codfw.wmnet * 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[2004-2007].codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin2003" * 12:15 jynus@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[2004-2007].codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin2003" * 12:09 jynus@cumin2003: START - Cookbook sre.dns.netbox * 11:58 jynus@cumin2003: START - Cookbook sre.hosts.decommission for hosts backup[2004-2007].codfw.wmnet * 10:40 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup[1004-1007].eqiad.wmnet * 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[1004-1007].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 10:01 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[1004-1007].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 09:52 jynus@cumin1003: START - Cookbook sre.dns.netbox * 09:39 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:36 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup[1004-1007].eqiad.wmnet * 09:36 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:25 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 09:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 09:16 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 09:05 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:04 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:00 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:59 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:57 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:55 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:50 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 08:50 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 08:49 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 08:49 atsukoito: depooling cirrussearch in codfw because of regression after upgrade [[phab:T431091|T431091]] * 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts mirror1001.wikimedia.org * 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: mirror1001.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 08:29 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: mirror1001.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 08:18 jmm@cumin2003: START - Cookbook sre.dns.netbox * 08:11 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts mirror1001.wikimedia.org * 06:15 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 18s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-02 == * 22:55 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host contint1003.wikimedia.org with OS trixie * 22:29 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on contint1003.wikimedia.org with reason: host reimage * 22:23 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on contint1003.wikimedia.org with reason: host reimage * 22:05 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host contint1003.wikimedia.org with OS trixie * 22:03 mutante: contint1003 (zuul.wikimedia.org) - reimaging because of [[phab:T430510|T430510]]#12067628 [[phab:T418521|T418521]] * 22:03 dzahn@cumin2002: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on zuul.wikimedia.org with reason: reimage * 21:39 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 18s) * 21:39 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 21:20 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1003.eqiad.wmnet, repooling source-only afterwards * 21:19 sbassett: Deployed security fix for [[phab:T428829|T428829]] * 20:58 cmooney@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Release v0.11.2 update for new Aerleon - cmooney@cumin1003 * 20:55 cmooney@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Release v0.11.2 update for new Aerleon - cmooney@cumin1003 * 20:40 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] (duration: 12m 35s) * 20:36 arlolra@deploy1003: cscott, arlolra: Continuing with deployment * 20:35 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 20s) * 20:35 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 20:33 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host contint2003.wikimedia.org with OS trixie * 20:31 arlolra@deploy1003: cscott, arlolra: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Cha * 20:28 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] * 20:17 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] (duration: 08m 13s) * 20:14 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on contint2003.wikimedia.org with reason: host reimage * 20:13 sbassett@deploy1003: sbassett: Continuing with deployment * 20:11 sbassett@deploy1003: sbassett: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:09 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] * 20:08 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 20:08 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 20:08 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on contint2003.wikimedia.org with reason: host reimage * 20:05 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1003.eqiad.wmnet, repooling source-only afterwards * 19:49 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host contint2003.wikimedia.org with OS trixie * 19:48 mutante: contint2003 - reimaging because of [[phab:T430510|T430510]]#12067628 [[phab:T418521|T418521]] * 18:39 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 18:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 18:13 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2002.codfw.wmnet -> wcqs2003.codfw.wmnet, repooling source-only afterwards * 17:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1003.eqiad.wmnet with OS bookworm * 17:52 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1005.eqiad.wmnet * 17:52 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1005.eqiad.wmnet * 17:51 jasmine@cumin2002: conftool action : set/pooled=yes:weight=10; selector: name=wikikube-ctrl1005.eqiad.wmnet * 17:48 jasmine_: homer "cr*eqiad*" commit "Added new stacked control plane wikikube-ctrl1005" * 17:44 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply * 17:44 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply * 17:31 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] (duration: 09m 33s) * 17:26 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 17:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1003.eqiad.wmnet with reason: host reimage * 17:23 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:21 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] * 17:18 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1003.eqiad.wmnet with reason: host reimage * 17:16 rscout@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply * 17:16 rscout@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply * 17:16 rscout@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply * 17:15 rscout@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply * 17:12 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on wcqs[2002-2003].codfw.wmnet,wcqs1002.eqiad.wmnet with reason: reimaging hosts * 17:08 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 17:08 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 17:08 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 17:07 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 17:05 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 17:05 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 17:03 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "running to make sure all updates are synced - cmooney@cumin1003" * 17:03 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "running to make sure all updates are synced - cmooney@cumin1003" * 17:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs1003 * 17:00 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs1003 * 17:00 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 17:00 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1003.eqiad.wmnet with OS bookworm * 16:58 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Re-running - btullis@cumin1003" * 16:58 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Re-running - btullis@cumin1003" * 16:58 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2002.codfw.wmnet -> wcqs2003.codfw.wmnet, repooling source-only afterwards * 16:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-master1004.eqiad.wmnet with OS bookworm * 16:58 btullis@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 16:57 tappof: bump space for prometheus k8s-aux in eqiad * 16:55 cmooney@dns3003: END - running authdns-update * 16:55 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:55 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to eqsin - cmooney@cumin1003" * 16:55 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to eqsin - cmooney@cumin1003" * 16:53 cmooney@dns3003: START - running authdns-update * 16:52 ryankemper: [ml-serve-eqiad] Cleared out 1302 failed (Evicted) pods: `kubectl -n llm delete pods --field-selector=status.phase=Failed`, freeing calico-kube-controllers from OOM crashloop (evictions were caused by disk pressure) * 16:49 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 16:46 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:39 rzl@dns1004: END - running authdns-update * 16:37 rzl@dns1004: START - running authdns-update * 16:36 rzl@dns1004: START - running authdns-update * 16:35 rzl@deploy1003: Finished scap sync-world: [[phab:T416623|T416623]] (duration: 10m 19s) * 16:34 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 16:33 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-master1004.eqiad.wmnet with reason: host reimage * 16:30 rzl@deploy1003: rzl: Continuing with deployment * 16:28 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-master1004.eqiad.wmnet with reason: host reimage * 16:26 rzl@deploy1003: rzl: [[phab:T416623|T416623]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:25 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 16:25 rzl@deploy1003: Started scap sync-world: [[phab:T416623|T416623]] * 16:25 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 16:24 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 16:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: sync * 16:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: sync * 16:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-master1004.eqiad.wmnet with OS bookworm * 16:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-master1003.eqiad.wmnet with OS bookworm * 16:11 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 16:11 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 16:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Security updates * 16:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 16:08 root@cumin1003: START - Cookbook sre.mysql.parsercache * 16:08 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Security updates * 15:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-master1003.eqiad.wmnet with reason: host reimage * 15:54 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:54 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:54 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:54 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-master1003.eqiad.wmnet with reason: host reimage * 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Security updates * 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:45 root@cumin1003: START - Cookbook sre.mysql.parsercache * 15:45 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Security updates * 15:42 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-master1003.eqiad.wmnet with OS bookworm * 15:24 moritzm: installing busybox updates from bookworm point release * 15:20 moritzm: installing busybox updates from trixie point release * 15:15 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1021: Security updates * 15:15 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:15 root@cumin1003: START - Cookbook sre.mysql.parsercache * 15:15 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1021: Security updates * 15:13 moritzm: installing giflib security updates * 15:08 moritzm: installing Tomcat security updates * 14:57 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 14:56 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 14:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:53 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Unblock taavi - oblivian@cumin1003" * 14:53 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Unblock taavi - oblivian@cumin1003 * 14:53 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1021: Security updates * 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:53 root@cumin1003: START - Cookbook sre.mysql.parsercache * 14:53 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1021: Security updates * 14:53 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Unblock taavi - oblivian@cumin1003 * 14:52 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Unblock taavi - oblivian@cumin1003" * 14:46 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94711 and previous config saved to /var/cache/conftool/dbconfig/20260702-144644-fceratto.json * 14:36 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205', diff saved to https://phabricator.wikimedia.org/P94709 and previous config saved to /var/cache/conftool/dbconfig/20260702-143636-fceratto.json * 14:32 moritzm: installing libdbi-perl security updates * 14:26 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205', diff saved to https://phabricator.wikimedia.org/P94708 and previous config saved to /var/cache/conftool/dbconfig/20260702-142628-fceratto.json * 14:16 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94707 and previous config saved to /var/cache/conftool/dbconfig/20260702-141621-fceratto.json * 14:12 moritzm: installing rsync security updates * 14:11 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox) * 14:10 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94706 and previous config saved to /var/cache/conftool/dbconfig/20260702-140959-fceratto.json * 14:09 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2205.codfw.wmnet with reason: Maintenance * 14:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2205: Repooling after switchover * 14:07 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-test-master1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 14:06 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 14:06 Tran: Deployed patch for [[phab:T427287|T427287]] * 14:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:59 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2205: Repooling after switchover * 13:59 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2205: Repooling after switchover * 13:59 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:55 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2205: Repooling after switchover * 13:55 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2205 [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94704 and previous config saved to /var/cache/conftool/dbconfig/20260702-135505-fceratto.json * 13:54 moritzm: installing sed security updates * 13:53 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:52 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2209 to s3 primary [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94703 and previous config saved to /var/cache/conftool/dbconfig/20260702-135235-fceratto.json * 13:52 federico3: Starting s3 codfw failover from db2205 to db2209 - [[phab:T430912|T430912]] * 13:51 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:51 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 13:48 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:47 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2209 with weight 0 [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94702 and previous config saved to /var/cache/conftool/dbconfig/20260702-134719-fceratto.json * 13:47 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Primary switchover s3 [[phab:T430912|T430912]] * 13:44 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:44 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:44 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:40 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 13:38 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 13:37 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 13:36 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 13:36 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:34 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 13:30 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:29 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:29 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:27 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:26 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:25 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling restart_daemons on A:wikidough * 13:23 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 13:22 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 13:17 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 13:17 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns1004.wikimedia.org * 13:12 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:11 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart (exit_code=97) rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough * 13:11 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=97) rolling restart_daemons on A:wikidough * 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough * 13:09 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] (duration: 07m 20s) * 13:05 aude@deploy1003: jdrewniak, aude: Continuing with deployment * 13:04 aude@deploy1003: jdrewniak, aude: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:02 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] * 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts wdqs-categories1001.eqiad.wmnet * 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: wdqs-categories1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 12:10 jmm@dns1004: END - running authdns-update * 12:07 jmm@dns1004: START - running authdns-update * 11:51 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: wdqs-categories1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 11:44 btullis@cumin1003: START - Cookbook sre.dns.netbox * 11:42 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 11:42 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 11:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet * 11:39 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts wdqs-categories1001.eqiad.wmnet * 11:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet * 11:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet * 11:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet * 11:29 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 11:29 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 10:57 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2214: Repooling * 10:49 jmm@dns1004: END - running authdns-update * 10:47 jmm@dns1004: START - running authdns-update * 10:31 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94698 and previous config saved to /var/cache/conftool/dbconfig/20260702-103146-fceratto.json * 10:21 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213', diff saved to https://phabricator.wikimedia.org/P94696 and previous config saved to /var/cache/conftool/dbconfig/20260702-102137-fceratto.json * 10:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:19 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb1017.eqiad.wmnet * 10:18 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 10:18 fceratto@cumin1003: Removing es1033 from zarcillo [[phab:T408772|T408772]] * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts es1033.eqiad.wmnet * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: es1033.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:14 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: es1033.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:13 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb1017.eqiad.wmnet * 10:12 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2214.codfw.wmnet * 10:12 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2214.codfw.wmnet * 10:12 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2214: Repooling * 10:11 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213', diff saved to https://phabricator.wikimedia.org/P94693 and previous config saved to /var/cache/conftool/dbconfig/20260702-101130-fceratto.json * 10:10 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:10 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:03 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts es1033.eqiad.wmnet * 10:03 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 10:01 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94691 and previous config saved to /var/cache/conftool/dbconfig/20260702-100122-fceratto.json * 09:55 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94690 and previous config saved to /var/cache/conftool/dbconfig/20260702-095529-fceratto.json * 09:55 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2213.codfw.wmnet with reason: Maintenance * 09:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 09:53 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2213: Repooling after switchover * 09:51 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover * 09:44 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2213: Repooling after switchover * 09:39 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover * 09:39 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2213 [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94688 and previous config saved to /var/cache/conftool/dbconfig/20260702-093859-fceratto.json * 09:36 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2192 to s5 primary [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94687 and previous config saved to /var/cache/conftool/dbconfig/20260702-093650-fceratto.json * 09:36 federico3: Starting s5 codfw failover from db2213 to db2192 - [[phab:T430923|T430923]] * 09:30 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94686 and previous config saved to /var/cache/conftool/dbconfig/20260702-093004-fceratto.json * 09:24 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2192 with weight 0 [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94685 and previous config saved to /var/cache/conftool/dbconfig/20260702-092455-fceratto.json * 09:24 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 23 hosts with reason: Primary switchover s5 [[phab:T430923|T430923]] * 09:19 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220', diff saved to https://phabricator.wikimedia.org/P94684 and previous config saved to /var/cache/conftool/dbconfig/20260702-091957-fceratto.json * 09:16 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] (duration: 06m 57s) * 09:13 moritzm: installing libgcrypt20 security updates * 09:12 kharlan@deploy1003: kharlan: Continuing with deployment * 09:11 kharlan@deploy1003: kharlan: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:09 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220', diff saved to https://phabricator.wikimedia.org/P94683 and previous config saved to /var/cache/conftool/dbconfig/20260702-090950-fceratto.json * 09:09 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] * 09:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 09:01 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] (duration: 07m 07s) * 08:59 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94682 and previous config saved to /var/cache/conftool/dbconfig/20260702-085942-fceratto.json * 08:57 kharlan@deploy1003: kharlan: Continuing with deployment * 08:56 kharlan@deploy1003: kharlan: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:54 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] * 08:52 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:52 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:52 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94681 and previous config saved to /var/cache/conftool/dbconfig/20260702-085237-fceratto.json * 08:52 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2220.codfw.wmnet with reason: Maintenance * 08:43 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:40 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 08:25 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] (duration: 11m 44s) * 08:21 cscott@deploy1003: cscott: Continuing with deployment * 08:16 cscott@deploy1003: cscott: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:14 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] * 08:08 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 08:08 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1244: Migration of db1244.eqiad.wmnet completed * 08:02 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:02 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:01 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] (duration: 18m 58s) * 08:01 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:59 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 07:59 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:59 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:59 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2006.wikimedia.org * 07:58 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:57 cscott@deploy1003: cscott: Continuing with deployment * 07:56 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:56 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:56 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:55 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:55 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:55 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:54 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2006.wikimedia.org * 07:54 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:54 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 07:44 cscott@deploy1003: cscott: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:44 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2005.wikimedia.org * 07:44 moritzm: installing node-lodash security updates * 07:42 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] * 07:39 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2005.wikimedia.org * 07:30 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] (duration: 07m 28s) * 07:26 cscott@deploy1003: ssastry, cscott: Continuing with deployment * 07:25 cscott@deploy1003: ssastry, cscott: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:23 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1244: Migration of db1244.eqiad.wmnet completed * 07:22 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] * 07:16 wmde-fisch@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] (duration: 06m 55s) * 07:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1244.eqiad.wmnet with OS trixie * 07:11 wmde-fisch@deploy1003: wmde-fisch: Continuing with deployment * 07:11 wmde-fisch@deploy1003: wmde-fisch: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:09 wmde-fisch@deploy1003: Started scap sync-world: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] * 06:54 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1244.eqiad.wmnet with reason: host reimage * 06:50 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1244.eqiad.wmnet with reason: host reimage * 06:38 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1250.eqiad.wmnet with OS trixie * 06:34 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db1244.eqiad.wmnet with OS trixie * 06:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1244: Upgrading db1244.eqiad.wmnet * 06:25 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1244: Upgrading db1244.eqiad.wmnet * 06:25 cwilliams@cumin1003: dbmaint on s4@eqiad [[phab:T429893|T429893]] * 06:25 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 06:15 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1250.eqiad.wmnet with reason: host reimage * 06:14 cwilliams@dns1006: END - running authdns-update * 06:12 cwilliams@dns1006: START - running authdns-update * 06:11 cwilliams@dns1006: END - running authdns-update * 06:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db1244 [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94676 and previous config saved to /var/cache/conftool/dbconfig/20260702-061059-cwilliams.json * 06:09 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1250.eqiad.wmnet with reason: host reimage * 06:09 cwilliams@dns1006: START - running authdns-update * 06:08 aokoth@cumin1003: END (PASS) - Cookbook sre.vrts.upgrade (exit_code=0) on VRTS host vrts1003.eqiad.wmnet * 06:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db1160 to s4 primary and set section read-write [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94675 and previous config saved to /var/cache/conftool/dbconfig/20260702-060746-cwilliams.json * 06:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Set s4 eqiad as read-only for maintenance - [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94674 and previous config saved to /var/cache/conftool/dbconfig/20260702-060704-cwilliams.json * 06:06 cezmunsta: Starting s4 eqiad failover from db1244 to db1160 - [[phab:T430817|T430817]] * 06:04 aokoth@cumin1003: START - Cookbook sre.vrts.upgrade on VRTS host vrts1003.eqiad.wmnet * 05:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db1160 with weight 0 [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94673 and previous config saved to /var/cache/conftool/dbconfig/20260702-055927-cwilliams.json * 05:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 40 hosts with reason: Primary switchover s4 [[phab:T430817|T430817]] * 05:55 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1250.eqiad.wmnet with OS trixie * 05:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on db1250.eqiad.wmnet with reason: m3 master switchover [[phab:T430158|T430158]] * 05:39 marostegui: Failover m3 (phabricator) from db1250 to db1228 - [[phab:T430158|T430158]] * 05:32 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2234].codfw.wmnet,db[1217,1228,1250].eqiad.wmnet with reason: m3 master switchover [[phab:T430158|T430158]] * 04:45 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] (duration: 09m 08s) * 04:41 tstarling@deploy1003: tstarling, reedy: Continuing with deployment * 04:38 tstarling@deploy1003: tstarling, reedy: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 04:36 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 59s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:16 ryankemper: [[phab:T429844|T429844]] [opensearch] completed `cirrussearch2111` reimage; all codfw search clusters are green, all nodes now report `OpenSearch 2.19.5`, and the temporary chi voting exclusion has been removed * 00:57 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2111.codfw.wmnet with OS trixie * 00:29 ryankemper: [[phab:T429844|T429844]] [opensearch] depooled codfw search-omega/search-psi discovery records to match existing codfw search depool during OpenSearch 2.19 migration * 00:29 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2111.codfw.wmnet with reason: host reimage * 00:29 ryankemper@cumin2002: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 00:29 ryankemper@cumin2002: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 00:22 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2111.codfw.wmnet with reason: host reimage * 00:01 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2111.codfw.wmnet with OS trixie * 00:00 ryankemper: [[phab:T429844|T429844]] [opensearch] chi cluster recovered after stopping `opensearch_1@production-search-codfw` on `cirrussearch2111` == 2026-07-01 == * 23:59 ryankemper: [[phab:T429844|T429844]] [opensearch] stopped `opensearch_1@production-search-codfw` on `cirrussearch2111` after chi cluster-manager election churn following `voting_config_exclusions` POST; hoping this triggers a re-election * 23:52 cscott@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 23:51 cscott@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 23:51 cscott@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 23:50 cscott@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2003.codfw.wmnet with OS bookworm * 22:29 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 22:13 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 22:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2084.codfw.wmnet with OS trixie * 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2003.codfw.wmnet with reason: host reimage * 22:03 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 22:01 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2003.codfw.wmnet with reason: host reimage * 21:50 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 21:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2084.codfw.wmnet with reason: host reimage * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2003 * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2003 * 21:42 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2003 * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2003.codfw.wmnet 45.48.192.10.in-addr.arpa 5.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:42 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2003.codfw.wmnet 45.48.192.10.in-addr.arpa 5.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2003 - bking@cumin2003" * 21:42 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2003 - bking@cumin2003" * 21:36 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2084.codfw.wmnet with reason: host reimage * 21:35 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:34 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2003 * 21:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2003.codfw.wmnet with OS bookworm * 21:19 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2084.codfw.wmnet with OS trixie * 21:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2081.codfw.wmnet with OS trixie * 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2108.codfw.wmnet with OS trixie * 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2081.codfw.wmnet with reason: host reimage * 20:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2081.codfw.wmnet with reason: host reimage * 20:28 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2081.codfw.wmnet with OS trixie * 20:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2108.codfw.wmnet with reason: host reimage * 20:19 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2108.codfw.wmnet with reason: host reimage * 19:59 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2108.codfw.wmnet with OS trixie * 19:46 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2093.codfw.wmnet with OS trixie * 19:44 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 19:44 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jasmine@cumin2002" * 19:43 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jasmine@cumin2002" * 19:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2080.codfw.wmnet with OS trixie * 19:28 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 19:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2093.codfw.wmnet with reason: host reimage * 19:18 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 19:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2093.codfw.wmnet with reason: host reimage * 19:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2080.codfw.wmnet with reason: host reimage * 19:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2080.codfw.wmnet with reason: host reimage * 18:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2093.codfw.wmnet with OS trixie * 18:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2080.codfw.wmnet with OS trixie * 18:27 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 18:18 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] (duration: 09m 15s) * 18:13 jgiannelos@deploy1003: jgiannelos, neriah: Continuing with deployment * 18:11 jgiannelos@deploy1003: jgiannelos, neriah: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:09 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] * 17:40 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 16:58 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 30 hosts * 16:57 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for 30 hosts * 16:52 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2202.codfw.wmnet * 16:52 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2202.codfw.wmnet * 16:51 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt * 16:51 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt * 16:51 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lvs2012.codfw.wmnet * 16:51 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for lvs2012.codfw.wmnet * 16:49 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2076.codfw.wmnet with OS trixie * 16:49 brett: Start pybal on lvs2012 - [[phab:T429861|T429861]] * 16:49 pt1979@cumin1003: END (ERROR) - Cookbook sre.hosts.remove-downtime (exit_code=97) for 59 hosts * 16:48 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for 59 hosts * 16:42 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2061.codfw.wmnet with OS trixie * 16:30 dancy@deploy1003: Installation of scap version "4.271.0" completed for 2 hosts * 16:28 dancy@deploy1003: Installing scap version "4.271.0" for 2 host(s) * 16:23 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2076.codfw.wmnet with reason: host reimage * 16:19 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2061.codfw.wmnet with reason: host reimage * 16:18 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2076.codfw.wmnet with reason: host reimage * 16:18 jasmine@dns1004: END - running authdns-update * 16:16 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host restbase2039.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 16:16 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host restbase2039.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 16:16 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2061.codfw.wmnet with reason: host reimage * 16:15 jasmine@dns1004: START - running authdns-update * 16:14 jasmine@dns1004: END - running authdns-update * 16:12 jasmine@dns1004: START - running authdns-update * 16:07 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2202.codfw.wmnet with reason: maintenance * 16:06 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt with reason: Junos upograde * 16:00 papaul: ongoing maintenance on lsw1-b2-codfw * 16:00 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2076.codfw.wmnet with OS trixie * 15:59 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt * 15:59 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt * 15:57 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2061.codfw.wmnet with OS trixie * 15:55 pt1979@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2042,2046].codfw.wmnet * 15:55 pt1979@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2042,2046].codfw.wmnet * 15:51 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 15:51 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2220: Repooling after switchover * 15:50 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 15:50 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 15:48 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2092.codfw.wmnet with OS trixie * 15:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 15:40 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 15:38 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 15:37 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 15:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply * 15:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply * 15:32 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 15:32 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 15:30 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 15:29 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 15:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 15:25 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:22 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs2012.codfw.wmnet with reason: Rack B2 maintenance - [[phab:T429861|T429861]] * 15:21 brett: Stopping pybal on lvs2012 in preparation for codfw rack b2 maintenance - [[phab:T429861|T429861]] * 15:20 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2092.codfw.wmnet with reason: host reimage * 15:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:12 _joe_: restarted manually alertmanager-irc-relay * 15:12 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2092.codfw.wmnet with reason: host reimage * 15:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:12 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt with reason: Junos upograde * 15:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover * 15:07 pt1979@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2042,2046].codfw.wmnet * 15:06 pt1979@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2042,2046].codfw.wmnet * 15:02 papaul: ongoing maintenance on lsw1-a8-codfw * 14:31 topranks: POWERING DOWN CR1-EQIAD for line card installation [[phab:T426343|T426343]] * 14:31 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] (duration: 08m 57s) * 14:29 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:26 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 14:24 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:22 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] * 14:22 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover * 14:16 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:15 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover * 14:14 topranks: re-enable routing-engine graceful-failover on cr1-eqiad [[phab:T417873|T417873]] * 14:13 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:13 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2220: Repooling after switchover * 14:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:12 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:12 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:11 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:08 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] (duration: 10m 01s) * 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:07 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2220 [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94664 and previous config saved to /var/cache/conftool/dbconfig/20260701-140729-fceratto.json * 14:06 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:06 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:06 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:05 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2159 to s7 primary [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94663 and previous config saved to /var/cache/conftool/dbconfig/20260701-140503-fceratto.json * 14:04 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:04 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 14:04 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 14:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:04 dreamyjazz@deploy1003: anzx, dreamyjazz: Continuing with deployment * 14:04 federico3: Starting s7 codfw failover from db2220 to db2159 - [[phab:T430826|T430826]] * 14:03 jmm@dns1004: END - running authdns-update * 14:03 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:03 topranks: flipping cr1-eqiad active routing-enginer back to RE0 [[phab:T417873|T417873]] * 14:03 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudsw1-c8-eqiad,cloudsw1-d5-eqiad with reason: router upgrades eqiad * 14:01 jmm@dns1004: START - running authdns-update * 14:00 dreamyjazz@deploy1003: anzx, dreamyjazz: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:59 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2159 with weight 0 [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94662 and previous config saved to /var/cache/conftool/dbconfig/20260701-135906-fceratto.json * 13:58 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] * 13:57 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s7 [[phab:T430826|T430826]] * 13:56 topranks: reboot routing-enginer RE0 on cr1-eqiad [[phab:T417873|T417873]] * 13:48 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1006.wikimedia.org * 13:44 atsuko@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cirrussearch2092.codfw.wmnet with OS trixie * 13:43 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1006.wikimedia.org * 13:41 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2092.codfw.wmnet with OS trixie * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1005.wikimedia.org * 13:37 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1005.wikimedia.org * 13:37 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on pfw1-eqiad with reason: router upgrades eqiad * 13:35 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on lvs[1017-1020].eqiad.wmnet with reason: router upgrades eqiad * 13:34 caro@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] (duration: 07m 59s) * 13:30 caro@deploy1003: caro: Continuing with deployment * 13:28 caro@deploy1003: caro: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:27 topranks: route-engine failover cr1-eqiad * 13:26 caro@deploy1003: Started scap sync-world: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] * 13:15 topranks: rebooting routing-engine 1 on cr1-eqiad [[phab:T417873|T417873]] * 13:13 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] (duration: 08m 29s) * 13:13 moritzm: installing qemu security updates * 13:11 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 13:11 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 13:09 jgiannelos@deploy1003: jgiannelos: Continuing with deployment * 13:08 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 13:07 jgiannelos@deploy1003: jgiannelos: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:06 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 13:06 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2214.codfw.wmnet with reason: Maintenance * 13:05 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2214: Repooling after switchover * 13:05 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] * 13:04 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2214: Repooling after switchover * 13:04 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2214 [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94660 and previous config saved to /var/cache/conftool/dbconfig/20260701-130413-fceratto.json * 13:01 moritzm: installing python3.13 security updates * 13:00 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2229 to s6 primary [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94659 and previous config saved to /var/cache/conftool/dbconfig/20260701-125959-fceratto.json * 12:59 federico3: Starting s6 codfw failover from db2214 to db2229 - [[phab:T430814|T430814]] * 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on 13 hosts with reason: router upgrade and line card install * 12:51 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2229 with weight 0 [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94658 and previous config saved to /var/cache/conftool/dbconfig/20260701-125149-fceratto.json * 12:51 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 21 hosts with reason: Primary switchover s6 [[phab:T430814|T430814]] * 12:50 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2189.codfw.wmnet * 12:50 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2189.codfw.wmnet * 12:42 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2100.codfw.wmnet with OS trixie * 12:38 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2083.codfw.wmnet with OS trixie * 12:19 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2083.codfw.wmnet with reason: host reimage * 12:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 12:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2240: Migration of db2240.codfw.wmnet completed * 12:14 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2100.codfw.wmnet with reason: host reimage * 12:09 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2083.codfw.wmnet with reason: host reimage * 12:09 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2100.codfw.wmnet with reason: host reimage * 12:00 topranks: drain traffic on cr1-eqiad to allow for line card install and JunOS upgrade [[phab:T426343|T426343]] * 11:52 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2083.codfw.wmnet with OS trixie * 11:50 cmooney@dns2005: END - running authdns-update * 11:49 cmooney@dns2005: START - running authdns-update * 11:48 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2100.codfw.wmnet with OS trixie * 11:40 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/zotero: apply * 11:40 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/zotero: apply * 11:36 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/zotero: apply * 11:36 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/zotero: apply * 11:31 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2240: Migration of db2240.codfw.wmnet completed * 11:30 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply * 11:28 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply * 11:27 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:27 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:27 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:27 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:27 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:23 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2240.codfw.wmnet with OS trixie * 11:20 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:20 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:17 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:16 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:16 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:15 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2086.codfw.wmnet with OS trixie * 11:14 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2106.codfw.wmnet with OS trixie * 11:14 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:13 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:12 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:09 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2115.codfw.wmnet with OS trixie * 11:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2240.codfw.wmnet with reason: host reimage * 11:00 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2240.codfw.wmnet with reason: host reimage * 10:53 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2106.codfw.wmnet with reason: host reimage * 10:49 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2086.codfw.wmnet with reason: host reimage * 10:44 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2115.codfw.wmnet with reason: host reimage * 10:44 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2240.codfw.wmnet with OS trixie * 10:44 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2086.codfw.wmnet with reason: host reimage * 10:42 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2106.codfw.wmnet with reason: host reimage * 10:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2240: Upgrading db2240.codfw.wmnet * 10:41 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2240: Upgrading db2240.codfw.wmnet * 10:41 cwilliams@cumin1003: dbmaint on s4@codfw [[phab:T429893|T429893]] * 10:40 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 10:39 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2115.codfw.wmnet with reason: host reimage * 10:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2240 [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94653 and previous config saved to /var/cache/conftool/dbconfig/20260701-102658-cwilliams.json * 10:26 moritzm: installing nginx security updates * 10:26 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2086.codfw.wmnet with OS trixie * 10:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2179 to s4 primary [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94652 and previous config saved to /var/cache/conftool/dbconfig/20260701-102356-cwilliams.json * 10:23 cezmunsta: Starting s4 codfw failover from db2240 to db2179 - [[phab:T430127|T430127]] * 10:23 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2106.codfw.wmnet with OS trixie * 10:20 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2115.codfw.wmnet with OS trixie * 10:15 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2179 with weight 0 [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94651 and previous config saved to /var/cache/conftool/dbconfig/20260701-101531-cwilliams.json * 10:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 40 hosts with reason: Primary switchover s4 [[phab:T430127|T430127]] * 09:56 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template (take 2) - oblivian@cumin1003" * 09:56 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template (take 2) - oblivian@cumin1003 * 09:55 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template (take 2) - oblivian@cumin1003 * 09:55 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template (take 2) - oblivian@cumin1003" * 09:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:39 mszwarc@deploy1003: Synchronized private/SuggestedInvestigationsSignals/SuggestedInvestigationsSignal4n.php: Update SI signal 4n (duration: 06m 08s) * 09:21 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 09:21 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 09:14 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 09:14 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 09:02 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 09:02 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 08:54 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 08:38 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 08:38 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 08:36 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 08:21 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 08:21 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] (duration: 36m 11s) * 08:15 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 08:09 mszwarc@deploy1003: mszwarc, abi: Continuing with deployment * 08:03 mszwarc@deploy1003: mszwarc, abi: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:55 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 07:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 07:45 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] * 07:30 aqu@deploy1003: Finished deploy [analytics/refinery@410f205]: Regular analytics weekly train 2nd try [analytics/refinery@410f2050] (duration: 00m 22s) * 07:29 aqu@deploy1003: Started deploy [analytics/refinery@410f205]: Regular analytics weekly train 2nd try [analytics/refinery@410f2050] * 07:28 aqu@deploy1003: Finished deploy [analytics/refinery@410f205] (thin): Regular analytics weekly train THIN [analytics/refinery@410f2050] (duration: 01m 59s) * 07:28 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] (duration: 07m 19s) * 07:26 aqu@deploy1003: Started deploy [analytics/refinery@410f205] (thin): Regular analytics weekly train THIN [analytics/refinery@410f2050] * 07:26 aqu@deploy1003: Finished deploy [analytics/refinery@410f205]: Regular analytics weekly train [analytics/refinery@410f2050] (duration: 04m 32s) * 07:24 mszwarc@deploy1003: wmde-fisch, mszwarc: Continuing with deployment * 07:23 mszwarc@deploy1003: wmde-fisch, mszwarc: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:21 aqu@deploy1003: Started deploy [analytics/refinery@410f205]: Regular analytics weekly train [analytics/refinery@410f2050] * 07:21 aqu@deploy1003: Finished deploy [analytics/refinery@410f205] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@410f2050] (duration: 02m 01s) * 07:20 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] * 07:19 aqu@deploy1003: Started deploy [analytics/refinery@410f205] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@410f2050] * 07:13 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] (duration: 09m 13s) * 07:09 mszwarc@deploy1003: mszwarc, chlod, revi: Continuing with deployment * 07:06 mszwarc@deploy1003: mszwarc, chlod, revi: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:04 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] * 06:55 elukey: upgrade all trixie hosts to pywmflib 3.0 - [[phab:T430552|T430552]] * 06:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:43 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:43 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:42 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:42 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:41 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:41 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:35 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:35 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:34 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:34 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:31 jmm@cumin2003: DONE (PASS) - Cookbook sre.idm.logout (exit_code=0) Logging Niharika29 out of all services on: 2453 hosts * 06:30 oblivian@cumin1003: END (FAIL) - Cookbook sre.deploy.hiddenparma (exit_code=99) Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:30 oblivian@cumin1003: END (FAIL) - Cookbook sre.deploy.python-code (exit_code=99) hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:30 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:30 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:01 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2109.codfw.wmnet with OS trixie * 05:45 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on es1039.eqiad.wmnet with reason: issues * 05:41 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1027.eqiad.wmnet * 05:40 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2068.codfw.wmnet with OS trixie * 05:40 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2109.codfw.wmnet with reason: host reimage * 05:40 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1027.eqiad.wmnet,service=s2 * 05:40 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1027.eqiad.wmnet,service=s7 * 05:36 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2109.codfw.wmnet with reason: host reimage * 05:20 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2068.codfw.wmnet with reason: host reimage * 05:16 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2109.codfw.wmnet with OS trixie * 05:15 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2068.codfw.wmnet with reason: host reimage * 05:09 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2067.codfw.wmnet with OS trixie * 04:56 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2068.codfw.wmnet with OS trixie * 04:49 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2067.codfw.wmnet with reason: host reimage * 04:45 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2067.codfw.wmnet with reason: host reimage * 04:27 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2067.codfw.wmnet with OS trixie * 03:47 slyngshede@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1039.eqiad.wmnet with reason: Hardware crash * 03:21 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2107.codfw.wmnet with OS trixie * 02:59 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2107.codfw.wmnet with reason: host reimage * 02:55 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2085.codfw.wmnet with OS trixie * 02:51 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2072.codfw.wmnet with OS trixie * 02:51 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2107.codfw.wmnet with reason: host reimage * 02:35 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2085.codfw.wmnet with reason: host reimage * 02:31 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2107.codfw.wmnet with OS trixie * 02:30 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2072.codfw.wmnet with reason: host reimage * 02:26 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2085.codfw.wmnet with reason: host reimage * 02:22 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2072.codfw.wmnet with reason: host reimage * 02:09 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2085.codfw.wmnet with OS trixie * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 54s) * 02:03 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2072.codfw.wmnet with OS trixie * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es7 eqiad back to read-write - [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94649 and previous config saved to /var/cache/conftool/dbconfig/20260701-010716-ladsgroup.json * 01:05 ladsgroup@dns1004: END - running authdns-update * 01:05 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depool es1039 [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94648 and previous config saved to /var/cache/conftool/dbconfig/20260701-010551-ladsgroup.json * 01:03 ladsgroup@dns1004: START - running authdns-update * 01:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Promote es1035 to es7 primary [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94647 and previous config saved to /var/cache/conftool/dbconfig/20260701-010002-ladsgroup.json * 00:58 Amir1: Starting es7 eqiad failover from es1039 to es1035 - [[phab:T430765|T430765]] * 00:53 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es1035 with weight 0 [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94646 and previous config saved to /var/cache/conftool/dbconfig/20260701-005329-ladsgroup.json * 00:53 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 9 hosts with reason: Primary switchover es7 [[phab:T430765|T430765]] * 00:42 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es7 eqiad as read-only for maintenance - [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94645 and previous config saved to /var/cache/conftool/dbconfig/20260701-004221-ladsgroup.json * 00:20 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2102.codfw.wmnet with OS trixie * 00:15 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2103.codfw.wmnet with OS trixie * 00:05 dr0ptp4kt: DEPLOYED Refinery at {{Gerrit|4e7a2b32}} for changes: pageview allowlist {{Gerrit|1305158}} (+min.wikiquote) {{Gerrit|1305162}} (+bol.wikipedia), {{Gerrit|1305156}} (+isv.wikipedia); {{Gerrit|1305980}} (pv allowlist -api.wikimedia, sqoop +isvwiki); sqoop {{Gerrit|1295064}} (+globalimagelinks) {{Gerrit|1295069}} (+filerevision) using scap, then deployed onto HDFS (manual copyToLocal required additionally) == Other archives == See [[Server Admin Log/Archives]]. <noinclude> [[Category:SAL]] [[Category:Operations]] </noinclude> 3t0nghf031nhqxgdi7xow1i49oingak 2450650 2450649 2026-08-22T16:33:29Z Stashbot 7414 arlolra@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply 2450650 wikitext text/x-wiki == 2026-08-22 == * 16:33 arlolra@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:33 arlolra@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:32 arlolra@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 35s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-21 == * 20:36 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in1001.wikimedia.org with reason: [[phab:T434750|T434750]] * 20:34 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in2001.wikimedia.org with reason: [[phab:T434750|T434750]] * 20:33 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out1001.wikimedia.org with reason: [[phab:T434750|T434750]] * 20:25 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out2001.wikimedia.org with reason: [[phab:T434750|T434750]] * 19:37 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host krb1002.eqiad.wmnet with OS bookworm * 19:00 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 18:59 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 18:51 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 18:51 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 18:35 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:35 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:27 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:27 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:16 bking@cumin2003: START - Cookbook sre.hosts.reimage for host krb1002.eqiad.wmnet with OS bookworm * 17:35 sukhe@dns1004: END - running authdns-update * 17:33 sukhe@dns1004: START - running authdns-update * 17:32 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns5004.wikimedia.org [reason: resolved authdns-update issues] * 17:32 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:32 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: force HEAD to {{Gerrit|be26e30ae101}} - sukhe@cumin1003" * 17:32 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: force HEAD to {{Gerrit|be26e30ae101}} - sukhe@cumin1003" * 17:28 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 17:28 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: service=authdns-update,name=dns5004.wikimedia.org [reason: resolving authdns-update issues] * 17:28 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:28 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: force HEAD to {{Gerrit|be26e30ae101}} - sukhe@cumin1003" * 17:28 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: force HEAD to {{Gerrit|be26e30ae101}} - sukhe@cumin1003" * 17:24 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 17:24 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.netbox (exit_code=97) * 17:23 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 17:16 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 17:12 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 17:10 sukhe@dns1004: END - running authdns-update * 17:08 sukhe@dns1004: START - running authdns-update * 17:08 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=dns5004.wikimedia.org [reason: resolving authdns-update issues] * 17:07 sukhe@dns1004: FAIL - running authdns-update * 17:05 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 17:05 sukhe@dns1004: START - running authdns-update * 17:01 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 16:59 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 16:56 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 16:53 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=dns5004.* [reason: trixie upgrade] * 16:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns5004.wikimedia.org * 16:52 cdobbins@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns5004.wikimedia.org * 16:44 cmooney@dns3003: END - running authdns-update * 16:41 cmooney@dns3003: START - running authdns-update * 16:41 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:41 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on eqsin<->codfw arelion - cmooney@cumin1003" * 16:37 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on eqsin<->codfw arelion - cmooney@cumin1003" * 16:33 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:11 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 16:08 sukhe@dns1004: END - running authdns-update * 16:08 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 16:06 sukhe@dns1004: START - running authdns-update * 16:04 cmooney@dns3003: END - running authdns-update * 16:02 cmooney@dns3003: START - running authdns-update * 16:00 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:00 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on eqord<->codfw arelion - cmooney@cumin1003" * 15:56 cdobbins@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host dns5004.wikimedia.org with OS trixie * 15:55 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on eqord<->codfw arelion - cmooney@cumin1003" * 15:53 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:51 cmooney@cumin1003: END (ERROR) - Cookbook sre.dns.netbox (exit_code=97) * 15:51 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:38 andrew@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudcephosd1042.eqiad.wmnet with OS bookworm * 15:18 andrew@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudcephosd1042.eqiad.wmnet with reason: host reimage * 15:17 dancy@deploy1003: Finished deploy [gerrit/gerrit@2cc11cc]: Deploying https://gerrit.wikimedia.org/r/c/operations/software/gerrit/+/1327669 ([[phab:T434726|T434726]]) (duration: 00m 14s) * 15:17 dancy@deploy1003: Started deploy [gerrit/gerrit@2cc11cc]: Deploying https://gerrit.wikimedia.org/r/c/operations/software/gerrit/+/1327669 ([[phab:T434726|T434726]]) * 15:13 andrew@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudcephosd1042.eqiad.wmnet with reason: host reimage * 15:09 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns5004.wikimedia.org with reason: host reimage * 15:05 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns5004.wikimedia.org with reason: host reimage * 14:53 andrew@cumin2003: START - Cookbook sre.hosts.reimage for host cloudcephosd1042.eqiad.wmnet with OS bookworm * 14:30 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns5004.wikimedia.org with OS trixie * 14:29 cdobbins@cumin1003: conftool action : set/pooled=no; selector: name=dns5004.* [reason: trixie upgrade] * 14:21 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:21 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove entries for cr2-eqord - cmooney@cumin1003" * 14:21 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove entries for cr2-eqord - cmooney@cumin1003" * 14:13 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 14:11 moritzm: imported openjdk 8u504-ga-1~deb12u1 for bookworm-wikimedia (backport of the latest Java 8 security fixes for bookworm) * 13:25 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "sync cr2-eqord router offline - cmooney@cumin1003" * 13:23 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "sync cr2-eqord router offline - cmooney@cumin1003" * 13:14 hashar@deploy1003: Finished deploy [integration/docroot@2d5ff9b]: opensource: add PersonalDashboard docs to MW components - [[phab:T435392|T435392]] (duration: 00m 15s) * 13:14 hashar@deploy1003: Started deploy [integration/docroot@2d5ff9b]: opensource: add PersonalDashboard docs to MW components - [[phab:T435392|T435392]] * 12:16 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:16 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: [[phab:T431682|T431682]] - filippo@cumin1003" * 12:16 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: [[phab:T431682|T431682]] - filippo@cumin1003" * 12:11 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2006.wikimedia.org with OS trixie * 12:00 kevinbazira@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 11:58 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 11:43 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2006.wikimedia.org with reason: host reimage * 11:41 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 11:38 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2006.wikimedia.org with reason: host reimage * 11:20 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2006.wikimedia.org with OS trixie * 11:11 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2005.wikimedia.org with OS trixie * 10:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2005.wikimedia.org with reason: host reimage * 10:53 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2005.wikimedia.org with reason: host reimage * 10:43 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-codfw * 10:43 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2011.codfw.wmnet * 10:43 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2011.codfw.wmnet * 10:40 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2011.codfw.wmnet * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2011.codfw.wmnet * 10:34 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2010.codfw.wmnet * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2010.codfw.wmnet * 10:33 fnegri@cumin1003: END (PASS) - Cookbook sre.wikireplicas.add-wiki (exit_code=0) for database bolwiki ([[phab:T429954|T429954]]) * 10:33 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2005.wikimedia.org with OS trixie * 10:30 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2010.codfw.wmnet * 10:25 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2010.codfw.wmnet * 10:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2009.codfw.wmnet * 10:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2009.codfw.wmnet * 10:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1006.wikimedia.org with OS trixie * 10:18 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2009.codfw.wmnet * 10:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2009.codfw.wmnet * 10:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2008.codfw.wmnet * 10:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2008.codfw.wmnet * 10:06 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2008.codfw.wmnet * 10:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1006.wikimedia.org with reason: host reimage * 10:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2008.codfw.wmnet * 10:01 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2007.codfw.wmnet * 10:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2007.codfw.wmnet * 09:57 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1006.wikimedia.org with reason: host reimage * 09:56 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2007.codfw.wmnet * 09:51 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2007.codfw.wmnet * 09:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 09:51 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 09:46 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2006.codfw.wmnet * 09:46 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1006.wikimedia.org with OS trixie * 09:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1005.wikimedia.org with OS trixie * 09:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2006.codfw.wmnet * 09:41 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2005.codfw.wmnet * 09:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2005.codfw.wmnet * 09:38 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2013.codfw.wmnet * 09:36 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2005.codfw.wmnet * 09:35 fnegri@cumin1003: START - Cookbook sre.wikireplicas.add-wiki for database bolwiki ([[phab:T429954|T429954]]) * 09:35 fnegri@cumin1003: END (PASS) - Cookbook sre.wikireplicas.add-wiki (exit_code=0) for database minwikiquote ([[phab:T429946|T429946]]) * 09:35 fnegri@cumin1003: START - Cookbook sre.wikireplicas.add-wiki for database minwikiquote ([[phab:T429946|T429946]]) * 09:32 blake@cumin1003: START - Cookbook sre.hosts.reboot-single for host rdb2013.codfw.wmnet * 09:30 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2011.codfw.wmnet * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1005.wikimedia.org with reason: host reimage * 09:26 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2005.codfw.wmnet * 09:26 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2004.codfw.wmnet * 09:26 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2004.codfw.wmnet * 09:24 blake@cumin1003: START - Cookbook sre.hosts.reboot-single for host rdb2011.codfw.wmnet * 09:22 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1005.wikimedia.org with reason: host reimage * 09:21 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2004.codfw.wmnet * 09:16 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1015.eqiad.wmnet * 09:16 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2004.codfw.wmnet * 09:16 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2003.codfw.wmnet * 09:16 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2003.codfw.wmnet * 09:13 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on cr[1-2]-eqiad,pfw1-eqiad with reason: upgrade pfw1a-eqiad and pfw1b-eqiad pair * 09:12 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 09:11 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 09:11 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 09:11 blake@cumin1003: START - Cookbook sre.hosts.reboot-single for host rdb1015.eqiad.wmnet * 09:10 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2003.codfw.wmnet * 09:09 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1013.eqiad.wmnet * 09:07 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1005.wikimedia.org with OS trixie * 09:03 blake@cumin1003: START - Cookbook sre.hosts.reboot-single for host rdb1013.eqiad.wmnet * 09:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2003.codfw.wmnet * 09:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2002.codfw.wmnet * 09:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2002.codfw.wmnet * 08:54 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2002.codfw.wmnet * 08:49 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2002.codfw.wmnet * 08:49 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2001.codfw.wmnet * 08:49 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2001.codfw.wmnet * 08:48 jmm@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts netmon2002.wikimedia.org * 08:47 jmm@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts netmon2002.wikimedia.org * 08:44 jmm@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts netmon2002.wikimedia.org * 08:44 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon2002.wikimedia.org * 08:43 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2001.codfw.wmnet * 08:36 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon2002.wikimedia.org * 08:34 jmm@dns1004: END - running authdns-update * 08:33 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2001.codfw.wmnet * 08:33 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-codfw * 08:31 jmm@dns1004: START - running authdns-update * 07:48 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327679{{!}}Block: Disable flaky API test (T435272 T389028)]], [[gerrit:1327678{{!}}API: wfDebugLog for thumberror]] (duration: 15m 34s) * 07:41 krinkle@deploy1003: krinkle: Continuing with deployment * 07:37 krinkle@deploy1003: krinkle: Backport for [[gerrit:1327679{{!}}Block: Disable flaky API test (T435272 T389028)]], [[gerrit:1327678{{!}}API: wfDebugLog for thumberror]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:33 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1327679{{!}}Block: Disable flaky API test (T435272 T389028)]], [[gerrit:1327678{{!}}API: wfDebugLog for thumberror]] * 07:25 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 07:24 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 07:18 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 07:18 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 07:15 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 07:14 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 07:14 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 07:14 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 07:13 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 07:03 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1283: Pool back * 06:42 jmm@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts netmon2002.wikimedia.org * 06:35 moritzm: powercycling netmon2002 * 06:18 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1283: Pool back * 06:17 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1283 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96213 and previous config saved to /var/cache/conftool/dbconfig/20260821-061743-marostegui.json * 04:59 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324963{{!}}Add Produnto to extension-list (T421436)]], [[gerrit:1324964{{!}}Enable Produnto on Beta (T421436)]] (duration: 34m 48s) * 04:45 tstarling@deploy1003: tstarling: Continuing with deployment * 04:44 tstarling@deploy1003: tstarling: Backport for [[gerrit:1324963{{!}}Add Produnto to extension-list (T421436)]], [[gerrit:1324964{{!}}Enable Produnto on Beta (T421436)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 04:24 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1324963{{!}}Add Produnto to extension-list (T421436)]], [[gerrit:1324964{{!}}Enable Produnto on Beta (T421436)]] * 04:21 arlolra@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 04:20 arlolra@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 04:20 arlolra@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 04:20 arlolra@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 41s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-20 == * 23:43 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327652{{!}}RunSingleJob: Add ProfilingContext::init() (T435422)]] (duration: 12m 06s) * 23:38 krinkle@deploy1003: krinkle: Continuing with deployment * 23:33 krinkle@deploy1003: krinkle: Backport for [[gerrit:1327652{{!}}RunSingleJob: Add ProfilingContext::init() (T435422)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:31 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1327652{{!}}RunSingleJob: Add ProfilingContext::init() (T435422)]] * 22:15 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1054.eqiad.wmnet with OS trixie * 22:14 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 22:14 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 21:58 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1054.eqiad.wmnet with reason: host reimage * 21:51 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1054.eqiad.wmnet with reason: host reimage * 21:36 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1054.eqiad.wmnet with OS trixie * 21:36 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:35 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327219{{!}}RunSingleJob: Define MW_ENTRY_POINT for flamegraph sample attribution (T435422)]] (duration: 08m 30s) * 21:31 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:31 krinkle@deploy1003: krinkle: Continuing with deployment * 21:31 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1054 * 21:31 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1054 * 21:29 krinkle@deploy1003: krinkle: Backport for [[gerrit:1327219{{!}}RunSingleJob: Define MW_ENTRY_POINT for flamegraph sample attribution (T435422)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:27 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1327219{{!}}RunSingleJob: Define MW_ENTRY_POINT for flamegraph sample attribution (T435422)]] * 21:17 maryum: Deployed security fix for [[phab:T433020|T433020]] * 20:59 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324752{{!}}InitialiseSettings: Enable 2FA banners on remaining private wikis (T428103)]], [[gerrit:1325920{{!}}Remove sending email to legal team about rejected requests (T374053)]] (duration: 07m 18s) * 20:54 reedy@deploy1003: neriah, reedy: Continuing with deployment * 20:54 reedy@deploy1003: neriah, reedy: Backport for [[gerrit:1324752{{!}}InitialiseSettings: Enable 2FA banners on remaining private wikis (T428103)]], [[gerrit:1325920{{!}}Remove sending email to legal team about rejected requests (T374053)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:51 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324752{{!}}InitialiseSettings: Enable 2FA banners on remaining private wikis (T428103)]], [[gerrit:1325920{{!}}Remove sending email to legal team about rejected requests (T374053)]] * 20:24 reedy@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.15,1.47.0-wmf.16,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/med * 20:23 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324752{{!}}InitialiseSettings: Enable 2FA banners on remaining private wikis (T428103)]], [[gerrit:1325920{{!}}Remove sending email to legal team about rejected requests (T374053)]] * 20:14 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 20:10 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 20:09 cdanis@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "bug fixes & UX fixes - cdanis@cumin1003" * 20:09 cdanis@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: bug fixes & UX fixes - cdanis@cumin1003 * 20:08 cdanis@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: bug fixes & UX fixes - cdanis@cumin1003 * 20:08 cdanis@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "bug fixes & UX fixes - cdanis@cumin1003" * 19:24 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327598{{!}}Make \Omicron non upright (like \Chi) (T434428)]], [[gerrit:1327596{{!}}Render overline of \bar with stretchy=false (T435456)]] (duration: 18m 54s) * 19:20 krinkle@deploy1003: krinkle: Continuing with deployment * 19:07 krinkle@deploy1003: krinkle: Backport for [[gerrit:1327598{{!}}Make \Omicron non upright (like \Chi) (T434428)]], [[gerrit:1327596{{!}}Render overline of \bar with stretchy=false (T435456)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:05 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1327598{{!}}Make \Omicron non upright (like \Chi) (T434428)]], [[gerrit:1327596{{!}}Render overline of \bar with stretchy=false (T435456)]] * 18:51 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327614{{!}}Avoid casting fpxmax to string (T318419)]] (duration: 07m 28s) * 18:50 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:46 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:46 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 18:45 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327614{{!}}Avoid casting fpxmax to string (T318419)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:43 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327614{{!}}Avoid casting fpxmax to string (T318419)]] * 18:37 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:37 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:37 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:36 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:07 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 18:05 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 18:01 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 18:01 sukhe@dns1004: END - running authdns-update * 17:59 sukhe@dns1004: START - running authdns-update * 17:58 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 17:57 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=dns6002.* [reason: depooling for trixie upgrade] * 17:56 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns6002.wikimedia.org * 17:56 cdobbins@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns6002.wikimedia.org * 17:51 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 17:51 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 17:34 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host stat1011.eqiad.wmnet with OS bookworm * 17:31 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns6002.wikimedia.org with OS trixie * 17:30 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 17:30 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 17:30 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 17:30 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 17:29 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 17:29 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 17:29 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 17:29 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:29 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:27 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:24 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327590{{!}}Make sure fpsmax is an int value (T318419)]] (duration: 08m 37s) * 17:20 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 17:17 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327590{{!}}Make sure fpsmax is an int value (T318419)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:16 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 17:15 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327590{{!}}Make sure fpsmax is an int value (T318419)]] * 16:51 swfrench-wmf: disable-puppet on A:cp for ATS Lua change - [[phab:T427666|T427666]] * 16:51 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db2901.codfw.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 16:43 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 16:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on stat1011.eqiad.wmnet with reason: host reimage * 16:39 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns6002.wikimedia.org with reason: host reimage * 16:36 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on stat1011.eqiad.wmnet with reason: host reimage * 16:36 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db2901.codfw.wmnet * 16:34 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns6002.wikimedia.org with reason: host reimage * 16:15 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns6002.wikimedia.org with OS trixie * 16:14 cdobbins@cumin1003: conftool action : set/pooled=no; selector: name=dns6002.* [reason: depooling for trixie upgrade] * 16:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host stat1011 * 16:10 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host stat1011 * 16:09 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host stat1011 * 16:09 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) stat1011.eqiad.wmnet 14.36.64.10.in-addr.arpa 4.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 bking@cumin2003: START - Cookbook sre.dns.wipe-cache stat1011.eqiad.wmnet 14.36.64.10.in-addr.arpa 4.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:09 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host stat1011 - bking@cumin2003" * 16:09 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host stat1011 - bking@cumin2003" * 16:05 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host stat1011 * 16:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host stat1011.eqiad.wmnet with OS bookworm * 16:00 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 15:59 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 15:56 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 15:55 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 15:37 jayme@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:35 jayme@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 15:35 jayme@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:32 jayme@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:32 jayme@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:30 jayme@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 15:30 jayme@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:28 jayme@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:28 jayme@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 15:28 fceratto@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host db1903.eqiad.wmnet * 15:28 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1903.eqiad.wmnet with OS trixie * 15:26 jayme@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 15:26 jayme@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 15:24 jayme@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 15:24 jayme@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 15:21 jayme@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 15:21 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 15:19 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 15:19 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 15:17 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 15:14 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1903.eqiad.wmnet with reason: host reimage * 15:07 fceratto@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1903.eqiad.wmnet with reason: host reimage * 14:54 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db1903.eqiad.wmnet with OS trixie * 14:53 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1903.eqiad.wmnet - fceratto@cumin1003" * 14:53 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1903.eqiad.wmnet - fceratto@cumin1003" * 14:53 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1903.eqiad.wmnet on all recursors * 14:53 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1903.eqiad.wmnet on all recursors * 14:53 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:53 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1903.eqiad.wmnet - fceratto@cumin1003" * 14:53 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1903.eqiad.wmnet - fceratto@cumin1003" * 14:49 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 14:49 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1903.eqiad.wmnet * 14:33 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=cp1100.* * 14:27 topranks: reconfigure eqiad<->codfw bgp settings * 14:22 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:22 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update entries used on new transport backup eqiad codfw - cmooney@cumin1003" * 14:19 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update entries used on new transport backup eqiad codfw - cmooney@cumin1003" * 14:14 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 14:14 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 14:13 moritzm: installing util-linux security updates * 14:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-staging-worker * 14:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2003.codfw.wmnet * 14:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2003.codfw.wmnet * 14:08 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2003.codfw.wmnet * 14:06 moritzm: installing libheif security updates * 13:58 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2003.codfw.wmnet * 13:58 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2002.codfw.wmnet * 13:58 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2002.codfw.wmnet * 13:56 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:56 fnegri@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for clouddb1025.eqiad.wmnet * 13:56 fnegri@cumin1003: START - Cookbook sre.hosts.remove-downtime for clouddb1025.eqiad.wmnet * 13:56 Lucas_WMDE: UTC afternoon backport+config window done * 13:53 moritzm: installing apr-util security updates * 13:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2002.codfw.wmnet * 13:50 fnegri@cumin1003: conftool action : set/weight=100; selector: name=clouddb1025.eqiad.wmnet * 13:49 fnegri@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1025.eqiad.wmnet * 13:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2002.codfw.wmnet * 13:41 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2001.codfw.wmnet * 13:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2001.codfw.wmnet * 13:41 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'. * 13:38 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'. * 13:38 fnegri@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on clouddb1025.eqiad.wmnet with reason: Removing s6 from clouddb1025 * 13:34 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2001.codfw.wmnet * 13:31 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'. * 13:29 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'. * 13:28 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet * 13:26 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host stat1009.eqiad.wmnet with OS bookworm * 13:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2001.codfw.wmnet * 13:24 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-staging-worker * 13:23 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1002.eqiad.wmnet * 13:21 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host stat1010.eqiad.wmnet with OS bookworm * 13:20 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1002.eqiad.wmnet * 13:20 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1001.eqiad.wmnet * 13:17 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1001.eqiad.wmnet * 13:16 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2001.codfw.wmnet * 13:13 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327511{{!}}UIC: Fix page:page instead of page:other in instrumentation]] (duration: 07m 00s) * 13:13 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2001.codfw.wmnet * 13:12 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2002.codfw.wmnet * 13:09 mszwarc@deploy1003: mszwarc: Continuing with deployment * 13:08 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1327511{{!}}UIC: Fix page:page instead of page:other in instrumentation]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2002.codfw.wmnet * 13:07 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2002.codfw.wmnet * 13:06 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1327511{{!}}UIC: Fix page:page instead of page:other in instrumentation]] * 13:04 jmm@dns1004: END - running authdns-update * 13:03 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2002.codfw.wmnet * 13:03 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2001.codfw.wmnet * 13:02 jmm@dns1004: START - running authdns-update * 13:01 cmooney@dns3003: END - running authdns-update * 13:00 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2001.codfw.wmnet * 12:59 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2001.codfw.wmnet * 12:59 cmooney@dns3003: START - running authdns-update * 12:57 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2001.codfw.wmnet * 12:56 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2002.codfw.wmnet * 12:55 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:55 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on drmrs<->eqiad GTT vpls - cmooney@cumin1003" * 12:54 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on drmrs<->eqiad GTT vpls - cmooney@cumin1003" * 12:54 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2002.codfw.wmnet * 12:54 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2003.codfw.wmnet * 12:50 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2003.codfw.wmnet * 12:49 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 12:48 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2001.codfw.wmnet * 12:46 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2001.codfw.wmnet * 12:46 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2002.codfw.wmnet * 12:43 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2002.codfw.wmnet * 12:43 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2003.codfw.wmnet * 12:42 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=1) for new host db1902.eqiad.wmnet * 12:42 fceratto@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host db1902.eqiad.wmnet with OS trixie * 12:41 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2003.codfw.wmnet * 12:40 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1003.eqiad.wmnet * 12:38 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1003.eqiad.wmnet * 12:37 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1002.eqiad.wmnet * 12:35 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1002.eqiad.wmnet * 12:35 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1001.eqiad.wmnet * 12:34 cmooney@dns3003: END - running authdns-update * 12:33 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1001.eqiad.wmnet * 12:32 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on stat1009.eqiad.wmnet with reason: host reimage * 12:31 cmooney@dns3003: START - running authdns-update * 12:31 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:31 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on drmrs<->eqiad cct - cmooney@cumin1003" * 12:28 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on drmrs<->eqiad cct - cmooney@cumin1003" * 12:28 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1902.eqiad.wmnet with reason: host reimage * 12:25 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on stat1009.eqiad.wmnet with reason: host reimage * 12:25 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 12:24 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on stat1010.eqiad.wmnet with reason: host reimage * 12:22 fceratto@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1902.eqiad.wmnet with reason: host reimage * 12:21 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on stat1010.eqiad.wmnet with reason: host reimage * 12:14 elukey: move the Docker Registry's /v2/dev/.* prefix to its dedicated S3 backend - [[phab:T432829|T432829]] * 12:12 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db1902.eqiad.wmnet with OS trixie * 12:09 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1902.eqiad.wmnet - fceratto@cumin1003" * 12:09 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1902.eqiad.wmnet - fceratto@cumin1003" * 12:09 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1902.eqiad.wmnet on all recursors * 12:09 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1902.eqiad.wmnet on all recursors * 12:09 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:08 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1902.eqiad.wmnet - fceratto@cumin1003" * 12:08 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1902.eqiad.wmnet - fceratto@cumin1003" * 12:08 tgr_: [[phab:T413390|T413390]] running CentralAuth:FixRenamedUserGlobalEditCount --wiki=metawiki --since=20250901000000 --until=20260301000000 --fix * 12:04 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1009.eqiad.wmnet with OS bookworm * 12:04 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1010.eqiad.wmnet with OS bookworm * 12:01 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 12:01 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1902.eqiad.wmnet * 12:00 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host stat1010.eqiad.wmnet with OS bookworm * 11:57 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1902.eqiad.wmnet * 11:57 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:57 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1902.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 11:57 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1902.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 11:51 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'. * 11:49 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'. * 11:48 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'. * 11:46 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'. * 11:39 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 11:37 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327518{{!}}Enable thumb.wikimedia.org on cswiki and fawiki (T427465)]] (duration: 10m 40s) * 11:35 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1902.eqiad.wmnet * 11:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1010.eqiad.wmnet with OS bookworm * 11:33 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 11:30 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327518{{!}}Enable thumb.wikimedia.org on cswiki and fawiki (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:26 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327518{{!}}Enable thumb.wikimedia.org on cswiki and fawiki (T427465)]] * 11:24 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host stat1010.eqiad.wmnet with OS bookworm * 11:07 fceratto@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host db1901.eqiad.wmnet * 11:07 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1901.eqiad.wmnet with OS trixie * 10:53 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1901.eqiad.wmnet with reason: host reimage * 10:47 tappof: bump space for prometheus k8s-aux in codfw * 10:47 tappof: bump space for prometheus k8s-dse in eqiad * 10:47 fceratto@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1901.eqiad.wmnet with reason: host reimage * 10:35 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db1901.eqiad.wmnet with OS trixie * 10:32 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:32 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:32 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1901.eqiad.wmnet on all recursors * 10:32 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1901.eqiad.wmnet on all recursors * 10:31 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:31 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:31 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:27 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:27 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1901.eqiad.wmnet * 10:23 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1010.eqiad.wmnet with OS bookworm * 10:20 fceratto@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts db1901.eqiad.wmnet * 10:20 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 10:18 blake@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 10:17 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:16 blake@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 10:13 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1901.eqiad.wmnet * 09:23 jelto@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'. * 09:22 jelto@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'. * 09:22 jelto: update cert-manager to 1.19.6 on wikikube staging-eqiad - [[phab:T427402|T427402]] * 09:20 moritzm: imported squid 7.6-2.1for trixie-wikimedia/main [[phab:T427282|T427282]] * 09:08 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 09:08 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 09:08 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 09:07 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 09:04 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 09:04 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:23 slyngshede@dns1004: END - running authdns-update * 08:21 slyngshede@dns1004: START - running authdns-update * 08:18 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.16 refs [[phab:T430835|T430835]] * 06:27 aokoth@dns1004: END - running authdns-update * 06:25 aokoth@dns1004: START - running authdns-update * 06:22 brennen@deploy1003: Finished deploy [phabricator/deployment@6b9b6ff]: deploy phab1005 for [[phab:T435087|T435087]] (duration: 00m 39s) * 06:21 brennen@deploy1003: Started deploy [phabricator/deployment@6b9b6ff]: deploy phab1005 for [[phab:T435087|T435087]] * 06:20 brennen@deploy1003: Finished deploy [phabricator/deployment@6b9b6ff]: deploy phab1004 for to pick up config values for [[phab:T435087|T435087]] (duration: 01m 46s) * 06:18 brennen@deploy1003: Started deploy [phabricator/deployment@6b9b6ff]: deploy phab1004 for to pick up config values for [[phab:T435087|T435087]] * 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 49s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-19 == * 23:19 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327197{{!}}Enable thumb.wikimedia.org on mediawiki.org (T427465)]] (duration: 10m 50s) * 23:18 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:16 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 23:15 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 23:10 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327197{{!}}Enable thumb.wikimedia.org on mediawiki.org (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:10 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:09 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 23:09 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:09 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 23:08 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327197{{!}}Enable thumb.wikimedia.org on mediawiki.org (T427465)]] * 22:58 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327201{{!}}Enable ReadingLists for all logged in users on test wiki (T435258)]] (duration: 11m 20s) * 22:50 jdlrobson@deploy1003: jdlrobson: Continuing with deployment * 22:49 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1327201{{!}}Enable ReadingLists for all logged in users on test wiki (T435258)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:46 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1327201{{!}}Enable ReadingLists for all logged in users on test wiki (T435258)]] * 22:42 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327162{{!}}Article: Split subjectpageheader by model and disable for wikitext]], [[gerrit:1327169{{!}}Make uppercase greek letters normal (non-italic) font (T434686 T434428)]], [[gerrit:1327176{{!}}Skin: Avoid DB lookup for pagecategorieslink message (T347123)]] (duration: 37m 52s) * 22:29 krinkle@deploy1003: krinkle: Continuing with deployment * 22:25 krinkle@deploy1003: krinkle: Backport for [[gerrit:1327162{{!}}Article: Split subjectpageheader by model and disable for wikitext]], [[gerrit:1327169{{!}}Make uppercase greek letters normal (non-italic) font (T434686 T434428)]], [[gerrit:1327176{{!}}Skin: Avoid DB lookup for pagecategorieslink message (T347123)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:04 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1327162{{!}}Article: Split subjectpageheader by model and disable for wikitext]], [[gerrit:1327169{{!}}Make uppercase greek letters normal (non-italic) font (T434686 T434428)]], [[gerrit:1327176{{!}}Skin: Avoid DB lookup for pagecategorieslink message (T347123)]] * 22:04 krinkle@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: awaiting CI (duration: 03m 06s) * 22:01 krinkle@deploy1003: Locking from deployment [ALL REPOSITORIES]: awaiting CI * 22:00 krinkle@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: awaiting CI (duration: 00m 01s) * 22:00 krinkle@deploy1003: Locking from deployment [ALL REPOSITORIES]: awaiting CI * 21:34 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2001.codfw.wmnet * 21:28 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2001.codfw.wmnet * 21:22 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327178{{!}}AccountRecovery: Notify the email address of the on file of the request (T425799)]] (duration: 47m 02s) * 21:13 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm * 21:09 catrope@deploy1003: catrope: Continuing with deployment * 20:55 catrope@deploy1003: catrope: Backport for [[gerrit:1327178{{!}}AccountRecovery: Notify the email address of the on file of the request (T425799)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:35 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1327178{{!}}AccountRecovery: Notify the email address of the on file of the request (T425799)]] * 20:31 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327128{{!}}Parsoid DataAccess: convert Parsoid fragment markers to/from strip tags (T432547)]] (duration: 07m 30s) * 20:27 catrope@deploy1003: catrope, arlolra: Continuing with deployment * 20:26 catrope@deploy1003: catrope, arlolra: Backport for [[gerrit:1327128{{!}}Parsoid DataAccess: convert Parsoid fragment markers to/from strip tags (T432547)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:24 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1327128{{!}}Parsoid DataAccess: convert Parsoid fragment markers to/from strip tags (T432547)]] * 20:23 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage * 20:17 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage * 20:15 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325878{{!}}[arwiki] Enable restricted user page editing and grant edit permissions (T434878)]] (duration: 08m 46s) * 20:11 catrope@deploy1003: catrope, gergesshamon: Continuing with deployment * 20:08 catrope@deploy1003: catrope, gergesshamon: Backport for [[gerrit:1325878{{!}}[arwiki] Enable restricted user page editing and grant edit permissions (T434878)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:06 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1325878{{!}}[arwiki] Enable restricted user page editing and grant edit permissions (T434878)]] * 19:59 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm * 19:56 eevans@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cassandra-dev2001.codfw.wmnet with OS bookworm * 19:56 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm * 19:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2207.codfw.wmnet with reason: Maintenance * 18:47 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319920{{!}}Allow setting a separate thumbUrl in production (T427465)]], [[gerrit:1327167{{!}}Fix wmgThumbUrl config (T427465)]] (duration: 18m 53s) * 18:43 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 18:30 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1319920{{!}}Allow setting a separate thumbUrl in production (T427465)]], [[gerrit:1327167{{!}}Fix wmgThumbUrl config (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:28 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1319920{{!}}Allow setting a separate thumbUrl in production (T427465)]], [[gerrit:1327167{{!}}Fix wmgThumbUrl config (T427465)]] * 18:26 sukhe@dns1004: END - running authdns-update * 18:24 sukhe@dns1004: START - running authdns-update * 18:09 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1319920{{!}}Allow setting a separate thumbUrl in production (T427465)]] * 18:03 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-eqiad and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 17:56 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-codfw and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 17:50 cmooney@dns3003: END - running authdns-update * 17:42 dancy@deploy1003: Installation of scap version "4.283.0" completed for 3 hosts * 17:41 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling reboot on A:durum and A:durum * 17:40 cmooney@dns3003: START - running authdns-update * 17:40 dancy@deploy1003: Installing scap version "4.283.0" for 3 host(s) * 17:38 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:37 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on GTT VPLS - cmooney@cumin1003" * 17:37 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --verbose --scoreLessThan=0.7 --exceptDatasetChecksums=[[phab:T434319|T434319]]-enwiki-models.txt # [[phab:T434319|T434319]] * 17:32 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on GTT VPLS - cmooney@cumin1003" * 17:29 sbassett: Deployed security fix for [[phab:T435210|T435210]] (wmf.16) * 17:26 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 17:22 sbassett: Deployed security fix for [[phab:T435210|T435210]] (wmf.15) * 17:00 sukhe@dns1004: END - running authdns-update * 16:58 sukhe@dns1004: START - running authdns-update * 16:53 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica-esams and A:liberica * 16:41 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica-esams and A:liberica * 16:41 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-eqiad and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 16:41 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-codfw and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 16:41 cjd91: sudo -i cookbook sre.cdn.roll-upgrade-ats --query 'A:cp-codfw' --task-id [[phab:T434478|T434478]] --reason '9.2.15 upgrade' * 16:41 cjd91: sudo -i cookbook sre.cdn.roll-upgrade-ats --query 'A:cp-eqiad' --task-id [[phab:T434478|T434478]] --reason '9.2.15 upgrade' * 16:40 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and A:durum * 16:28 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326881{{!}}Echo: Start using virtual domains (T380385)]] (duration: 13m 12s) * 16:23 urbanecm@deploy1003: urbanecm: Continuing with deployment * 16:21 urandom: Completed sessionstore Cassandra/JVM upgrade — [[phab:T435154|T435154]] * 16:21 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching sessionstore[2005-2006].codfw.wmnet,sessionstore[1005-1006].eqiad.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 16:19 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1326881{{!}}Echo: Start using virtual domains (T380385)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:15 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:15 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->codfw - cmooney@cumin1003" * 16:14 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1326881{{!}}Echo: Start using virtual domains (T380385)]] * 16:14 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327123{{!}}Revert^2 "Migrate database access to virtual domains" (T435305)]], [[gerrit:1327124{{!}}Pass the mapped domain of virtual-echo-shared to the push NameTableStores (T435305)]] (duration: 07m 42s) * 16:13 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching sessionstore[2005-2006].codfw.wmnet,sessionstore[1005-1006].eqiad.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 16:11 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->codfw - cmooney@cumin1003" * 16:10 urbanecm@deploy1003: urbanecm: Continuing with deployment * 16:08 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1327123{{!}}Revert^2 "Migrate database access to virtual domains" (T435305)]], [[gerrit:1327124{{!}}Pass the mapped domain of virtual-echo-shared to the push NameTableStores (T435305)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:06 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 16:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 16:06 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1327123{{!}}Revert^2 "Migrate database access to virtual domains" (T435305)]], [[gerrit:1327124{{!}}Pass the mapped domain of virtual-echo-shared to the push NameTableStores (T435305)]] * 16:06 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:03 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching sessionstore1004.eqiad.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 16:01 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching sessionstore1004.eqiad.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 15:56 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching sessionstore2004.codfw.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 15:54 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching sessionstore2004.codfw.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 15:53 urandom: beginning sessionstore Cassandra/JVM upgrade — [[phab:T435154|T435154]] * 15:52 urandom: beginning sessionstore Cassandra/JVM upgrade — [[phab:T432944|T432944]] * 15:51 cmooney@dns3003: END - running authdns-update * 15:49 cmooney@dns3003: START - running authdns-update * 15:48 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:48 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->codfw - cmooney@cumin1003" * 15:45 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->codfw - cmooney@cumin1003" * 15:44 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 15:44 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:43 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 15:42 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:42 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:38 cmooney@dns3003: END - running authdns-update * 15:36 cmooney@dns3003: START - running authdns-update * 15:36 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:36 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->eqsin - cmooney@cumin1003" * 15:34 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327138{{!}}Enable redis lock manager everywhere (T366938)]] (duration: 08m 36s) * 15:33 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->eqsin - cmooney@cumin1003" * 15:30 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:29 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 15:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1147.eqiad.wmnet with OS bookworm * 15:28 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327138{{!}}Enable redis lock manager everywhere (T366938)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:25 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327138{{!}}Enable redis lock manager everywhere (T366938)]] * 15:24 jmm@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host krb1002.eqiad.wmnet * 15:19 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2207.codfw.wmnet with reason: Host crashed * 15:17 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327098{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]], [[gerrit:1327101{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]] (duration: 07m 13s) * 15:12 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 15:12 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1327098{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]], [[gerrit:1327101{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:10 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1327098{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]], [[gerrit:1327101{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]] * 15:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2010.codfw.wmnet with OS trixie * 15:05 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 15:05 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1147.eqiad.wmnet with reason: host reimage * 14:59 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 14:59 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:58 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1147.eqiad.wmnet with reason: host reimage * 14:55 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:55 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:49 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 14:48 cmooney@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host durum1001.eqiad.wmnet * 14:46 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:46 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Delete db2902 ipv6 addr - fceratto@cumin1003" * 14:46 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Delete db2902 ipv6 addr - fceratto@cumin1003" * 14:43 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1147.eqiad.wmnet with OS bookworm * 14:42 tgr@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327094{{!}}SpecialMWOAuthListConsumers: Handle newFromMWUser returning null in addNavigationSubtitle (T435167)]] (duration: 19m 25s) * 14:42 cmooney@cumin1003: START - Cookbook sre.hosts.reboot-single for host durum1001.eqiad.wmnet * 14:42 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 14:41 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 14:38 tgr@deploy1003: tgr: Continuing with deployment * 14:36 tgr@deploy1003: tgr: Backport for [[gerrit:1327094{{!}}SpecialMWOAuthListConsumers: Handle newFromMWUser returning null in addNavigationSubtitle (T435167)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:28 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 14:28 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:28 cmooney@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host durum3005.esams.wmnet * 14:25 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 14:23 cmooney@cumin1003: START - Cookbook sre.hosts.reboot-single for host durum3005.esams.wmnet * 14:23 tgr@deploy1003: Started scap sync-world: Backport for [[gerrit:1327094{{!}}SpecialMWOAuthListConsumers: Handle newFromMWUser returning null in addNavigationSubtitle (T435167)]] * 14:18 gengh@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:17 topranks: disable puppet on hosts running BIRD BGP to test merge of patch to systemd healtchcheck service * 14:17 gengh@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:17 gengh@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:16 gengh@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:16 elukey: upgrade spicerack on cumin1003 and cumin2003 * 14:16 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:15 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314962{{!}}static: add new dir bimi/ for BIMI SVG and PEM file (T311685)]] (duration: 10m 00s) * 14:15 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:11 kharlan@deploy1003: kharlan, sukhe: Continuing with deployment * 14:11 gengh@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:09 gengh@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:09 gengh@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:08 kharlan@deploy1003: kharlan, sukhe: Backport for [[gerrit:1314962{{!}}static: add new dir bimi/ for BIMI SVG and PEM file (T311685)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:06 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 14:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:06 gengh@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:05 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1314962{{!}}static: add new dir bimi/ for BIMI SVG and PEM file (T311685)]] * 14:05 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host krb1002.eqiad.wmnet * 14:05 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:04 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:04 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326342{{!}}Revert^2 "wmf-config/ProductionServices: set URL for urldownloader to service record"]] (duration: 07m 40s) * 13:59 kharlan@deploy1003: kharlan, sukhe: Continuing with deployment * 13:58 kharlan@deploy1003: kharlan, sukhe: Backport for [[gerrit:1326342{{!}}Revert^2 "wmf-config/ProductionServices: set URL for urldownloader to service record"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:57 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host stat1008.eqiad.wmnet with OS bookworm * 13:56 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1326342{{!}}Revert^2 "wmf-config/ProductionServices: set URL for urldownloader to service record"]] * 13:56 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 13:54 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325532{{!}}srwiki: Allow bureaucrats to add and remove event-organizer group (T434748)]] (duration: 14m 56s) * 13:54 swfrench@dns1004: END - running authdns-update * 13:52 swfrench@dns1004: START - running authdns-update * 13:51 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:51 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 13:48 kharlan@deploy1003: kharlan, danielyepezgarces: Continuing with deployment * 13:45 swfrench@cumin2003: conftool action : set/pooled=yes; selector: name=wikikube-worker2330.codfw.wmnet * 13:44 swfrench-wmf: finished etcd-main codfw -> eqiad switchover - [[phab:T435103|T435103]] * 13:44 kharlan@deploy1003: kharlan, danielyepezgarces: Backport for [[gerrit:1325532{{!}}srwiki: Allow bureaucrats to add and remove event-organizer group (T434748)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:44 swfrench@cumin2003: conftool action : set/pooled=no; selector: name=wikikube-worker2330.codfw.wmnet * 13:41 swfrench@dns1004: END - running authdns-update * 13:39 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1325532{{!}}srwiki: Allow bureaucrats to add and remove event-organizer group (T434748)]] * 13:39 swfrench@dns1004: START - running authdns-update * 13:37 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327111{{!}}Special:AbuseReview: Add "no further action needed" review action (T435020)]], [[gerrit:1327110{{!}}AbuseReview: Take the review verdict as a REST path parameter (T435020)]] (duration: 31m 43s) * 13:31 swfrench-wmf: starting etcd-main codfw -> eqiad switchover - [[phab:T435103|T435103]] * 13:28 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:28 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 13:25 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=97) rolling reboot on A:durum and A:durum * 13:24 kharlan@deploy1003: kharlan: Continuing with deployment * 13:24 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1147 * 13:24 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1147 * 13:23 kharlan@deploy1003: kharlan: Backport for [[gerrit:1327111{{!}}Special:AbuseReview: Add "no further action needed" review action (T435020)]], [[gerrit:1327110{{!}}AbuseReview: Take the review verdict as a REST path parameter (T435020)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:16 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:16 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 13:12 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:12 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 13:10 cdobbins@cumin1003: END (ERROR) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=97) Rolling upgrade of ATS on A:cp-codfw and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 13:10 cdobbins@cumin1003: END (ERROR) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=97) Rolling upgrade of ATS on A:cp-eqiad and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 13:06 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1327111{{!}}Special:AbuseReview: Add "no further action needed" review action (T435020)]], [[gerrit:1327110{{!}}AbuseReview: Take the review verdict as a REST path parameter (T435020)]] * 13:06 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-eqiad and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 13:05 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-codfw and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 13:05 cjd91: sudo -i cookbook sre.cdn.roll-upgrade-ats --query 'A:cp-eqiad' --task-id [[phab:T434478|T434478]] --reason '9.2.15 upgrade' * 13:03 swfrench@dns1004: END - running authdns-update * 13:01 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:01 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 13:01 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host ncmonitor1001.eqiad.wmnet * 13:00 swfrench@dns1004: START - running authdns-update * 12:59 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 12:59 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 12:59 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 12:59 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 12:57 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and A:durum * 12:53 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on db2902.codfw.wmnet with reason: Cloning * 12:48 cmooney@dns3003: END - running authdns-update * 12:48 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327096{{!}}Switch to redis lock manager on s4 and s8 (T366938)]] (duration: 09m 11s) * 12:46 cmooney@dns3003: START - running authdns-update * 12:45 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:45 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->eqord cct - cmooney@cumin1003" * 12:45 elukey: move the /v2/releng.* prefix on the Docker Registry to its new s3 backend - [[phab:T432829|T432829]] * 12:43 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 12:42 jelto@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 12:42 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->eqord cct - cmooney@cumin1003" * 12:42 jelto@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 12:41 jelto: update cert-manager to 1.19.6 on wikikube staging-codfw - [[phab:T427402|T427402]] * 12:40 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327096{{!}}Switch to redis lock manager on s4 and s8 (T366938)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:38 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327096{{!}}Switch to redis lock manager on s4 and s8 (T366938)]] * 12:38 blake@deploy1003: Finished scap sync-world: non-build deploy for [[phab:T417800|T417800]] (duration: 03m 52s) * 12:36 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 12:35 blake@deploy1003: Started scap sync-world: non-build deploy for [[phab:T417800|T417800]] * 12:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host krb2002.codfw.wmnet * 11:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host krb2002.codfw.wmnet * 11:49 moritzm: installing kerberos security updates * 11:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on stat1008.eqiad.wmnet with reason: host reimage * 11:44 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on stat1008.eqiad.wmnet with reason: host reimage * 11:31 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327084{{!}}Revert "Migrate database access to virtual domains" (T435305)]] (duration: 11m 02s) * 11:29 gkyziridis@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:29 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:29 kart_: Updated MinT to 2026-06-04-131507-production ([[phab:T321316|T321316]]) * 11:28 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/machinetranslation: apply * 11:28 gkyziridis@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:26 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 11:24 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:24 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:23 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:23 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:23 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/machinetranslation: apply * 11:22 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327084{{!}}Revert "Migrate database access to virtual domains" (T435305)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:21 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/machinetranslation: apply * 11:21 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:21 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:20 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327084{{!}}Revert "Migrate database access to virtual domains" (T435305)]] * 11:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1008.eqiad.wmnet with OS bookworm * 11:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet * 11:17 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/machinetranslation: apply * 11:13 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/machinetranslation: apply * 11:12 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:12 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:10 kartik@deploy1003: helmfile [staging] START helmfile.d/services/machinetranslation: apply * 11:08 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-cron: apply * 11:08 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/mw-cron: apply * 11:08 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply * 11:08 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply * 11:07 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet * 11:07 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host stat1008.eqiad.wmnet with OS bookworm * 11:06 moritzm: upgrading the new trixie URL downloaders to Squid 7.6 [[phab:T427282|T427282]] * 11:01 gkyziridis@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin2002.codfw.wmnet * 10:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin2002.codfw.wmnet * 10:45 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1169.eqiad.wmnet with OS bookworm * 10:42 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1185.eqiad.wmnet with OS bookworm * 10:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1169.eqiad.wmnet with reason: host reimage * 10:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1185.eqiad.wmnet with reason: host reimage * 10:14 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1169.eqiad.wmnet with reason: host reimage * 10:14 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1185.eqiad.wmnet with reason: host reimage * 10:11 jmm@cumin2003: END (PASS) - Cookbook sre.netbox.restart-reboot (exit_code=0) rolling reboot on A:netbox * 10:06 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1008.eqiad.wmnet with OS bookworm * 10:04 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 10:04 mpostoronca@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321579{{!}}Register the mediawiki.wikimedia_antiabuse.content_policy_score stream (T432848)]] (duration: 08m 53s) * 10:03 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 10:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 10:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 10:00 mpostoronca@deploy1003: mpostoronca: Continuing with deployment * 09:59 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1185.eqiad.wmnet with OS bookworm * 09:59 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1169.eqiad.wmnet with OS bookworm * 09:59 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.convert-disks (exit_code=0) for host ms-be1065 * 09:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1065.eqiad.wmnet with OS trixie * 09:58 mpostoronca@deploy1003: mpostoronca: Backport for [[gerrit:1321579{{!}}Register the mediawiki.wikimedia_antiabuse.content_policy_score stream (T432848)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:55 jmm@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netbox.discovery.wmnet. on all recursors * 09:55 mpostoronca@deploy1003: Started scap sync-world: Backport for [[gerrit:1321579{{!}}Register the mediawiki.wikimedia_antiabuse.content_policy_score stream (T432848)]] * 09:55 jmm@cumin2003: START - Cookbook sre.dns.wipe-cache netbox.discovery.wmnet. on all recursors * 09:52 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw2001.wikimedia.org with OS trixie * 09:51 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 09:51 jmm@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netbox.discovery.wmnet. on all recursors * 09:51 jmm@cumin2003: START - Cookbook sre.dns.wipe-cache netbox.discovery.wmnet. on all recursors * 09:51 jmm@cumin2003: START - Cookbook sre.netbox.restart-reboot rolling reboot on A:netbox * 09:46 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 09:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1056.eqiad.wmnet with OS trixie * 09:44 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 09:39 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 09:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 09:36 topranks: make HE transport circuits from magru live * 09:36 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 09:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb1003.eqiad.wmnet * 09:33 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage * 09:31 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb1003.eqiad.wmnet * 09:30 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 09:28 arnaudb@dns1006: END - running authdns-update * 09:27 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage * 09:27 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb2003.codfw.wmnet * 09:26 arnaudb@dns1006: START - running authdns-update * 09:24 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1056.eqiad.wmnet with reason: host reimage * 09:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb2003.codfw.wmnet * 09:20 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.convert-disks (exit_code=0) for host ms-be1068 * 09:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1068.eqiad.wmnet with OS trixie * 09:20 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 09:19 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "cloudvirt1057 - filippo@cumin1003" * 09:19 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "cloudvirt1057 - filippo@cumin1003" * 09:18 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1056.eqiad.wmnet with reason: host reimage * 09:18 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1057.eqiad.wmnet with OS trixie * 09:18 filippo@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 09:18 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 09:15 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 09:14 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 09:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host irc1003.wikimedia.org * 09:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:13 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1065.eqiad.wmnet with OS trixie * 09:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 09:09 moritzm: installing Postgresql security updates * 09:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host irc1003.wikimedia.org * 09:07 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 09:07 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:07 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1173.eqiad.wmnet with OS bookworm * 09:06 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw2001.wikimedia.org with OS trixie * 09:03 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.convert-disks (exit_code=0) for host ms-be1064 * 09:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1064.eqiad.wmnet with OS trixie * 09:03 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1056.eqiad.wmnet with OS trixie * 09:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1057.eqiad.wmnet with reason: host reimage * 09:01 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 08:58 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 08:56 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1057.eqiad.wmnet with reason: host reimage * 08:54 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1208.eqiad.wmnet with OS bookworm * 08:53 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2902.codfw.wmnet with OS trixie * 08:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 08:50 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1174.eqiad.wmnet with OS bookworm * 08:46 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1172.eqiad.wmnet with OS bookworm * 08:45 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1173.eqiad.wmnet with reason: host reimage * 08:41 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 08:40 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1057.eqiad.wmnet with OS trixie * 08:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1057.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 08:39 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1207.eqiad.wmnet with OS bookworm * 08:38 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2902.codfw.wmnet with reason: host reimage * 08:37 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 08:34 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1068.eqiad.wmnet with OS trixie * 08:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1208.eqiad.wmnet with reason: host reimage * 08:31 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1057.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 08:29 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 08:28 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1222.eqiad.wmnet onto db1276.eqiad.wmnet * 08:28 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1222: Pool db1222.eqiad.wmnet in after cloning * 08:28 fceratto@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2902.codfw.wmnet with reason: host reimage * 08:27 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1055.eqiad.wmnet with OS trixie * 08:27 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 08:26 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1174.eqiad.wmnet with reason: host reimage * 08:25 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 08:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-misc2002.codfw.wmnet * 08:23 topranks: reboot pfw1-codfw firewall pair to upgrade JunOS [[phab:T434865|T434865]] * 08:22 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1172.eqiad.wmnet with reason: host reimage * 08:20 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1064.eqiad.wmnet with OS trixie * 08:20 mvernon@cumin2003: START - Cookbook sre.swift.convert-disks for host ms-be1065 * 08:18 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1207.eqiad.wmnet with reason: host reimage * 08:17 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1174.eqiad.wmnet with reason: host reimage * 08:17 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1173.eqiad.wmnet with reason: host reimage * 08:17 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1172.eqiad.wmnet with reason: host reimage * 08:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host mc-misc2002.codfw.wmnet * 08:15 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1208.eqiad.wmnet with reason: host reimage * 08:15 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1207.eqiad.wmnet with reason: host reimage * 08:14 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db2902.codfw.wmnet with OS trixie * 08:14 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.16 refs [[phab:T430835|T430835]] * 08:13 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db2902.codfw.wmnet * 08:13 fceratto@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host db2902.codfw.wmnet with OS trixie * 08:10 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr[1-2]-codfw with reason: upgrade pfw1a-codfw and pfw1b-codfw pair * 08:09 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1055.eqiad.wmnet with reason: host reimage * 08:07 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on pfw1-codfw with reason: upgrade pfw1a-codfw and pfw1b-codfw pair * 08:03 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1055.eqiad.wmnet with reason: host reimage * 08:02 arnaudb@dns1006: END - running authdns-update * 08:02 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1208.eqiad.wmnet with OS bookworm * 08:02 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1207.eqiad.wmnet with OS bookworm * 08:01 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1174.eqiad.wmnet with OS bookworm * 08:01 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1173.eqiad.wmnet with OS bookworm * 08:01 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1172.eqiad.wmnet with OS bookworm * 07:59 arnaudb@dns1006: START - running authdns-update * 07:58 arnaudb@dns1006: START - running authdns-update * 07:48 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1055.eqiad.wmnet with OS trixie * 07:45 moritzm: extend the disk of ldap-rw2001 by 80G [[phab:T331699|T331699]] * 07:42 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1222: Pool db1222.eqiad.wmnet in after cloning * 07:36 mvernon@cumin2003: START - Cookbook sre.swift.convert-disks for host ms-be1068 * 07:35 mvernon@cumin2003: START - Cookbook sre.swift.convert-disks for host ms-be1064 * 07:17 moritzm: installing imagemagick security updates * 07:14 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1277: Pool back * 07:14 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon1003.wikimedia.org * 07:07 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon1003.wikimedia.org * 07:03 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1280: Pool back * 07:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon2002.wikimedia.org * 06:55 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon2002.wikimedia.org * 06:54 moritzm: installing php8.2 security updates * 06:51 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1284: Pool back * 06:49 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1222: Depool db1222.eqiad.wmnet to then clone it to db1276.eqiad.wmnet - marostegui@cumin1003 * 06:49 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1222: Depool db1222.eqiad.wmnet to then clone it to db1276.eqiad.wmnet - marostegui@cumin1003 * 06:49 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1222.eqiad.wmnet onto db1276.eqiad.wmnet * 06:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd1005.eqiad.wmnet * 06:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd1005.eqiad.wmnet * 06:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd1004.eqiad.wmnet * 06:36 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2209: db2209 repool * 06:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd1004.eqiad.wmnet * 06:32 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast3007.wikimedia.org * 06:29 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1277: Pool back * 06:28 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1277 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96190 and previous config saved to /var/cache/conftool/dbconfig/20260819-062815-marostegui.json * 06:26 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast3007.wikimedia.org * 06:22 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul1001.eqiad.wmnet * 06:18 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul1001.eqiad.wmnet * 06:18 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul1003.eqiad.wmnet * 06:18 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1280: Pool back * 06:17 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1284 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96186 and previous config saved to /var/cache/conftool/dbconfig/20260819-061743-marostegui.json * 06:14 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul1003.eqiad.wmnet * 06:14 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul1002.eqiad.wmnet * 06:10 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul1002.eqiad.wmnet * 06:10 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2048.codfw.wmnet * 06:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2048.codfw.wmnet * 06:06 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1284: Pool back * 06:06 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1284 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96184 and previous config saved to /var/cache/conftool/dbconfig/20260819-060621-marostegui.json * 06:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2048.codfw.wmnet * 05:59 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2048.codfw.wmnet * 05:51 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2209: db2209 repool * 03:16 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1186.eqiad.wmnet with OS bookworm * 02:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1186.eqiad.wmnet with reason: host reimage * 02:46 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1186.eqiad.wmnet with reason: host reimage * 02:46 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2207 [[phab:T435270|T435270]]', diff saved to https://phabricator.wikimedia.org/P96181 and previous config saved to /var/cache/conftool/dbconfig/20260819-024627-marostegui.json * 02:44 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2204 to s2 primary [[phab:T435270|T435270]]', diff saved to https://phabricator.wikimedia.org/P96180 and previous config saved to /var/cache/conftool/dbconfig/20260819-024403-marostegui.json * 02:43 marostegui: Starting s2 codfw failover from db2207 to db2204 - [[phab:T435270|T435270]] * 02:39 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2204 with weight 0 [[phab:T435270|T435270]]', diff saved to https://phabricator.wikimedia.org/P96179 and previous config saved to /var/cache/conftool/dbconfig/20260819-023951-marostegui.json * 02:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s2 [[phab:T435270|T435270]] * 02:32 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1186.eqiad.wmnet with OS bookworm * 02:29 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-worker1186.eqiad.wmnet with OS bookworm * 02:18 denisse@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2207: Depooling replica * 02:18 denisse@cumin1003: START - Cookbook sre.mysql.depool depool db2207: Depooling replica * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 48s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-18 == * 23:55 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1264845{{!}}Remove unused/redundant wgMFNoindexPages=true setting (T255458)]] (duration: 09m 42s) * 23:51 krinkle@deploy1003: krinkle: Continuing with deployment * 23:48 krinkle@deploy1003: krinkle: Backport for [[gerrit:1264845{{!}}Remove unused/redundant wgMFNoindexPages=true setting (T255458)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:45 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1264845{{!}}Remove unused/redundant wgMFNoindexPages=true setting (T255458)]] * 23:38 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326944{{!}}Retire filebackend lock manager in favour of the default one (T366938)]] (duration: 08m 55s) * 23:34 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 23:31 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326944{{!}}Retire filebackend lock manager in favour of the default one (T366938)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:29 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326944{{!}}Retire filebackend lock manager in favour of the default one (T366938)]] * 23:27 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1170.eqiad.wmnet with OS bookworm * 23:21 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1205.eqiad.wmnet with OS bookworm * 23:20 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1171.eqiad.wmnet with OS bookworm * 23:15 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1206.eqiad.wmnet with OS bookworm * 23:05 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1170.eqiad.wmnet with reason: host reimage * 23:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1205.eqiad.wmnet with reason: host reimage * 22:57 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1171.eqiad.wmnet with reason: host reimage * 22:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1206.eqiad.wmnet with reason: host reimage * 22:53 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1205.eqiad.wmnet with reason: host reimage * 22:51 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1171.eqiad.wmnet with reason: host reimage * 22:51 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1170.eqiad.wmnet with reason: host reimage * 22:50 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1206.eqiad.wmnet with reason: host reimage * 22:36 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1206.eqiad.wmnet with OS bookworm * 22:35 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1205.eqiad.wmnet with OS bookworm * 22:35 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1186.eqiad.wmnet with OS bookworm * 22:35 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1171.eqiad.wmnet with OS bookworm * 22:35 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1170.eqiad.wmnet with OS bookworm * 22:33 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-worker1194.eqiad.wmnet with OS bookworm * 22:22 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326923{{!}}Enable redis lock manager on s6 (T366938)]] (duration: 11m 52s) * 22:18 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 22:13 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326923{{!}}Enable redis lock manager on s6 (T366938)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:10 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326923{{!}}Enable redis lock manager on s6 (T366938)]] * 22:04 sbassett: Deployed security fix for [[phab:T435234|T435234]] (wmf.16) * 21:54 sbassett: Deployed security fix for [[phab:T435234|T435234]] (wmf.15) * 21:38 caro@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326925{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326926{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326929{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]], [[gerrit:1326928{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]] (duration: 0 * 21:34 caro@deploy1003: caro: Continuing with deployment * 21:33 caro@deploy1003: caro: Backport for [[gerrit:1326925{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326926{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326929{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]], [[gerrit:1326928{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]] synced to the testservers (see h * 21:31 caro@deploy1003: Started scap sync-world: Backport for [[gerrit:1326925{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326926{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326929{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]], [[gerrit:1326928{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]] * 21:24 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1204.eqiad.wmnet with reason: 1204 datanode repair [[phab:T434494|T434494]] * 21:02 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326896{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]], [[gerrit:1326897{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]] (duration: 13m 56s) * 20:58 krinkle@deploy1003: krinkle: Continuing with deployment * 20:50 krinkle@deploy1003: krinkle: Backport for [[gerrit:1326896{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]], [[gerrit:1326897{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:49 ryankemper: `an-launcher1003` terminated process group `666809` (`rest_backfill_phase1.sh`) ~20 mins ago with `sudo kill -TERM -- -666809` after its local spark driver (`--driver-memory 64g`) repeatedly exhausted memory on the 32 GB VM and caused SSH to intermittently flap; host recovered to 27 GB available memory * 20:48 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1326896{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]], [[gerrit:1326897{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]] * 20:35 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 20:33 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 20:31 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 20:31 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326870{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]], [[gerrit:1326871{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]] (duration: 07m 35s) * 20:28 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 20:26 kemayo@deploy1003: kemayo: Continuing with deployment * 20:26 ryankemper: `an-launcher1003` confirmed the host is flapping because of memory thrash. chasing down the source of the thrash * 20:25 kemayo@deploy1003: kemayo: Backport for [[gerrit:1326870{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]], [[gerrit:1326871{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:23 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1326870{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]], [[gerrit:1326871{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]] * 20:23 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 20:20 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 20:14 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1168.eqiad.wmnet with OS bookworm * 20:08 zabe: zabe@deploy1003:~$ mwscript extensions/WikimediaMaintenance/maintenance/fixFileRevisionArchiveNameDrift.php enwiki # [[phab:T428406|T428406]] * 20:08 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1204.eqiad.wmnet with OS bookworm * 20:05 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1167.eqiad.wmnet with OS bookworm * 20:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1166.eqiad.wmnet with OS bookworm * 19:54 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1203.eqiad.wmnet with OS bookworm * 19:53 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326907{{!}}Revert "Disable redis lock manager on testwiki"]] (duration: 11m 05s) * 19:50 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1168.eqiad.wmnet with reason: host reimage * 19:47 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1204.eqiad.wmnet with reason: host reimage * 19:46 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 19:44 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326907{{!}}Revert "Disable redis lock manager on testwiki"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:42 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326907{{!}}Revert "Disable redis lock manager on testwiki"]] * 19:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1167.eqiad.wmnet with reason: host reimage * 19:37 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1166.eqiad.wmnet with reason: host reimage * 19:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1203.eqiad.wmnet with reason: host reimage * 19:32 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1167.eqiad.wmnet with reason: host reimage * 19:32 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1168.eqiad.wmnet with reason: host reimage * 19:32 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1166.eqiad.wmnet with reason: host reimage * 19:31 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1204.eqiad.wmnet with reason: host reimage * 19:31 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1203.eqiad.wmnet with reason: host reimage * 19:26 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324751{{!}}InitialiseSettings: Enable 2FA enforcement on more private wikis (T428103)]], [[gerrit:1326875{{!}}Add banner notifying of upcoming 2FA enforcement (T420792)]] (duration: 31m 46s) * 19:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1204.eqiad.wmnet with OS bookworm * 19:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1203.eqiad.wmnet with OS bookworm * 19:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1168.eqiad.wmnet with OS bookworm * 19:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1167.eqiad.wmnet with OS bookworm * 19:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1166.eqiad.wmnet with OS bookworm * 19:15 denisse: rebooting kafkamon2003.codfw.wmnet - [[phab:T435162|T435162]] * 19:14 denisse: rebooting kafkamon1003.eqiad.wmnet [[phab:T435162|T435162]] * 19:13 reedy@deploy1003: reedy: Continuing with deployment * 19:12 reedy@deploy1003: reedy: Backport for [[gerrit:1324751{{!}}InitialiseSettings: Enable 2FA enforcement on more private wikis (T428103)]], [[gerrit:1326875{{!}}Add banner notifying of upcoming 2FA enforcement (T420792)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:54 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324751{{!}}InitialiseSettings: Enable 2FA enforcement on more private wikis (T428103)]], [[gerrit:1326875{{!}}Add banner notifying of upcoming 2FA enforcement (T420792)]] * 18:50 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 18:44 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.16 refs [[phab:T430835|T430835]] * 18:34 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 18:31 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 18:22 aklapper@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326852{{!}}CategoryTree: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]], [[gerrit:1326853{{!}}CategoryViewer: Allow null $html in the CategoryViewerGenerateLink hook (T435161)]], [[gerrit:1326865{{!}}Flow: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]] (duration: 09m 57s) * 18:18 aklapper@deploy1003: jforrester, aklapper: Continuing with deployment * 18:17 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 18:14 aklapper@deploy1003: jforrester, aklapper: Backport for [[gerrit:1326852{{!}}CategoryTree: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]], [[gerrit:1326853{{!}}CategoryViewer: Allow null $html in the CategoryViewerGenerateLink hook (T435161)]], [[gerrit:1326865{{!}}Flow: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki * 18:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1156.eqiad.wmnet with OS bookworm * 18:12 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 18:12 aklapper@deploy1003: Started scap sync-world: Backport for [[gerrit:1326852{{!}}CategoryTree: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]], [[gerrit:1326853{{!}}CategoryViewer: Allow null $html in the CategoryViewerGenerateLink hook (T435161)]], [[gerrit:1326865{{!}}Flow: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]] * 18:11 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 18:08 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1146.eqiad.wmnet with OS bookworm * 18:07 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1177.eqiad.wmnet with OS bookworm * 18:00 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326380{{!}}Introduce main lock manager service (T366938 T427999)]] (duration: 11m 25s) * 17:58 ladsgroup@deploy1003: ladsgroup: Rolling back deployment * 17:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1202.eqiad.wmnet with OS bookworm * 17:55 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1201.eqiad.wmnet with OS bookworm * 17:50 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326380{{!}}Introduce main lock manager service (T366938 T427999)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:48 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326380{{!}}Introduce main lock manager service (T366938 T427999)]] * 17:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1156.eqiad.wmnet with reason: host reimage * 17:46 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1146.eqiad.wmnet with reason: host reimage * 17:45 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-esams and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 17:42 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 17:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1177.eqiad.wmnet with reason: host reimage * 17:38 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1202.eqiad.wmnet with reason: host reimage * 17:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1201.eqiad.wmnet with reason: host reimage * 17:30 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1156.eqiad.wmnet with reason: host reimage * 17:29 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1177.eqiad.wmnet with reason: host reimage * 17:28 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1146.eqiad.wmnet with reason: host reimage * 17:28 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1202.eqiad.wmnet with reason: host reimage * 17:27 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1201.eqiad.wmnet with reason: host reimage * 17:25 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 17:21 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 17:14 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1202.eqiad.wmnet with OS bookworm * 17:14 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1201.eqiad.wmnet with OS bookworm * 17:14 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1177.eqiad.wmnet with OS bookworm * 17:14 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1156.eqiad.wmnet with OS bookworm * 17:14 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1146.eqiad.wmnet with OS bookworm * 17:09 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326885{{!}}w/deployment-info.php: Handle new file format (T434726)]] (duration: 07m 15s) * 17:08 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 17:05 dancy@deploy1003: dancy: Continuing with deployment * 17:04 dancy@deploy1003: dancy: Backport for [[gerrit:1326885{{!}}w/deployment-info.php: Handle new file format (T434726)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:02 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1326885{{!}}w/deployment-info.php: Handle new file format (T434726)]] * 16:50 dancy@deploy1003: Finished scap sync-world: Testing [[phab:T434726|T434726]] (duration: 06m 40s) * 16:43 dancy@deploy1003: Started scap sync-world: Testing [[phab:T434726|T434726]] * 16:43 dancy@deploy1003: Installation of scap version "4.282.0" completed for 3 hosts * 16:41 dancy@deploy1003: Installing scap version "4.282.0" for 3 host(s) * 16:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1145.eqiad.wmnet with OS bookworm * 16:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1200.eqiad.wmnet with OS bookworm * 16:32 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1199.eqiad.wmnet with OS bookworm * 16:18 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1145.eqiad.wmnet with reason: host reimage * 16:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1200.eqiad.wmnet with reason: host reimage * 16:09 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1199.eqiad.wmnet with reason: host reimage * 16:05 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-esams and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 16:05 cjd91: sudo -i cookbook sre.cdn.roll-upgrade-ats --query 'A:cp-esams' --task-id [[phab:T434478|T434478]] --reason '9.2.15 upgrade' * 16:03 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1200.eqiad.wmnet with reason: host reimage * 16:02 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1145.eqiad.wmnet with reason: host reimage * 16:02 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1199.eqiad.wmnet with reason: host reimage * 15:50 moritzm: installing zip security updates * 15:48 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1200.eqiad.wmnet with OS bookworm * 15:47 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1199.eqiad.wmnet with OS bookworm * 15:47 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1145.eqiad.wmnet with OS bookworm * 15:41 topranks: bounce PIC 0/0 on cr1-magru to set port to 40G * 15:37 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1236.eqiad.wmnet with OS bookworm * 15:29 aikochou@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'ores-legacy' for release 'main' . * 15:26 aikochou@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'ores-legacy' for release 'main' . * 15:20 aikochou@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'ores-legacy' for release 'main' . * 15:14 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2002.codfw.wmnet * 15:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1236.eqiad.wmnet with reason: host reimage * 15:12 moritzm: failover ganeti master in codfw to ganeti2047 * 15:09 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1236.eqiad.wmnet with reason: host reimage * 15:09 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2044.codfw.wmnet * 15:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2002.codfw.wmnet * 15:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2044.codfw.wmnet * 15:04 brennen@deploy1003: Finished deploy [phabricator/deployment@6b9b6ff]: deploy phab1004 for [[phab:T435213|T435213]] (duration: 01m 01s) * 15:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2044.codfw.wmnet * 15:03 brennen@deploy1003: Started deploy [phabricator/deployment@6b9b6ff]: deploy phab1004 for [[phab:T435213|T435213]] * 15:03 brennen@deploy1003: Finished deploy [phabricator/deployment@6b9b6ff]: deploy phab2003 for [[phab:T435213|T435213]] (duration: 00m 57s) * 15:02 brennen@deploy1003: Started deploy [phabricator/deployment@6b9b6ff]: deploy phab2003 for [[phab:T435213|T435213]] * 14:57 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply * 14:57 arnaudb@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on phab2003.codfw.wmnet,phab[1004-1006].eqiad.wmnet with reason: maintenance * 14:56 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2044.codfw.wmnet * 14:55 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply * 14:53 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1236.eqiad.wmnet with OS bookworm * 14:51 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2043.codfw.wmnet * 14:51 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2043.codfw.wmnet * 14:45 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2043.codfw.wmnet * 14:33 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2043.codfw.wmnet * 14:22 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2042.codfw.wmnet * 14:22 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2042.codfw.wmnet * 14:21 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db2902.codfw.wmnet with OS trixie * 14:16 elukey: uploaded spicerack_13.2.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia * 14:16 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2042.codfw.wmnet * 14:04 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2042.codfw.wmnet * 14:04 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2041.codfw.wmnet * 14:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2041.codfw.wmnet * 13:59 phuedx@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: apply * 13:59 phuedx@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-main: apply * 13:59 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326838{{!}}Enable Suggested Investigations on hewiki (T435146)]] (duration: 11m 50s) * 13:58 phuedx@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: apply * 13:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2041.codfw.wmnet * 13:57 phuedx@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-main: apply * 13:57 phuedx@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-main: apply * 13:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard1003.eqiad.wmnet * 13:57 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-main: apply * 13:55 phuedx@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-logging-external: apply * 13:55 phuedx@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-logging-external: apply * 13:54 phuedx@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-logging-external: apply * 13:54 stran@deploy1003: stran: Continuing with deployment * 13:54 phuedx@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-logging-external: apply * 13:54 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-logging-external: apply * 13:54 phuedx@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-logging-external: apply * 13:53 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-logging-external: apply * 13:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard1003.eqiad.wmnet * 13:53 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2041.codfw.wmnet * 13:51 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db2902.codfw.wmnet - fceratto@cumin1003" * 13:51 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db2902.codfw.wmnet - fceratto@cumin1003" * 13:51 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2026.codfw.wmnet * 13:51 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard2003.codfw.wmnet * 13:50 phuedx@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: apply * 13:50 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-eqiad * 13:50 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp1001.eqiad.wmnet * 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp1001.eqiad.wmnet * 13:50 phuedx@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: apply * 13:49 phuedx@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: apply * 13:49 stran@deploy1003: stran: Backport for [[gerrit:1326838{{!}}Enable Suggested Investigations on hewiki (T435146)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:48 phuedx@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: apply * 13:48 phuedx@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics: apply * 13:47 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard2003.codfw.wmnet * 13:47 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics: apply * 13:47 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1326838{{!}}Enable Suggested Investigations on hewiki (T435146)]] * 13:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp1001.eqiad.wmnet * 13:44 cdanis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 13:43 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp1001.eqiad.wmnet * 13:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1315-1327].eqiad.wmnet * 13:43 cdanis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 13:43 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1315-1327].eqiad.wmnet * 13:40 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki1001.eqiad.wmnet * 13:35 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1315-1327].eqiad.wmnet * 13:34 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host rpki1001.eqiad.wmnet * 13:27 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1315-1327].eqiad.wmnet * 13:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1302-1314].eqiad.wmnet * 13:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1302-1314].eqiad.wmnet * 13:21 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326775{{!}}SI: Instrument case update on first edit (T435048)]] (duration: 07m 12s) * 13:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1302-1314].eqiad.wmnet * 13:17 stran@deploy1003: stran: Continuing with deployment * 13:16 stran@deploy1003: stran: Backport for [[gerrit:1326775{{!}}SI: Instrument case update on first edit (T435048)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:14 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1326775{{!}}SI: Instrument case update on first edit (T435048)]] * 13:13 moritzm: installing util-linux security updates * 13:11 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1302-1314].eqiad.wmnet * 13:10 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326232{{!}}prv: Enable parsoid rendering for 5 wikis (T435115)]] (duration: 08m 17s) * 13:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1288-1289,1291-1301].eqiad.wmnet * 13:10 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply * 13:10 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1288-1289,1291-1301].eqiad.wmnet * 13:10 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply * 13:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki2003.codfw.wmnet * 13:06 jgiannelos@deploy1003: jgiannelos: Continuing with deployment * 13:05 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host rpki2003.codfw.wmnet * 13:04 jgiannelos@deploy1003: jgiannelos: Backport for [[gerrit:1326232{{!}}prv: Enable parsoid rendering for 5 wikis (T435115)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1288-1289,1291-1301].eqiad.wmnet * 13:02 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1326232{{!}}prv: Enable parsoid rendering for 5 wikis (T435115)]] * 12:53 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1288-1289,1291-1301].eqiad.wmnet * 12:53 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1273,1275-1287].eqiad.wmnet * 12:53 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1273,1275-1287].eqiad.wmnet * 12:52 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt1002.wikimedia.org * 12:51 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db2902.codfw.wmnet on all recursors * 12:51 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db2902.codfw.wmnet on all recursors * 12:51 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:51 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db2902.codfw.wmnet - fceratto@cumin1003" * 12:51 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db2902.codfw.wmnet - fceratto@cumin1003" * 12:46 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt1002.wikimedia.org * 12:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt2002.wikimedia.org * 12:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1273,1275-1287].eqiad.wmnet * 12:42 dhinus: repooled clouddb1032 that was currently <nowiki>{</nowiki>"weight": 0, "pooled": "inactive"<nowiki>}</nowiki> for both s4 and s6 * 12:41 dhinus: also depooled clouddb1017 (forgot it in the previous list) * 12:41 fnegri@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet * 12:40 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 12:40 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db2902.codfw.wmnet * 12:40 fnegri@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032.eqiad.wmnet * 12:40 fnegri@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032 * 12:39 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1017.eqiad.wmnet * 12:39 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt2002.wikimedia.org * 12:38 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1020.eqiad.wmnet * 12:38 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1018.eqiad.wmnet * 12:38 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet * 12:37 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1014.eqiad.wmnet * 12:37 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1013.eqiad.wmnet * 12:37 dhinus: depool again clouddb10[13,14,16,18,20] that were repooled by the cookbook sre.mysql.multiinstance_reboot * 12:36 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1273,1275-1287].eqiad.wmnet * 12:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1248-1261].eqiad.wmnet * 12:36 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1248-1261].eqiad.wmnet * 12:35 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2026.codfw.wmnet * 12:34 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 12:34 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 12:34 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 12:34 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 12:29 lucaswerkmeister-wmde@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 12:28 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2026.codfw.wmnet * 12:28 lucaswerkmeister-wmde@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 12:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1248-1261].eqiad.wmnet * 12:27 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.addnode (exit_code=0) for new host ganeti2046.codfw.wmnet to cluster codfw and group A * 12:26 moritzm: readded ganeti2046 to the codfw cluster following firmware update and reimage [[phab:T434681|T434681]] * 12:23 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply * 12:23 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply * 12:23 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply * 12:22 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply * 12:22 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply * 12:22 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply * 12:21 jmm@cumin2003: START - Cookbook sre.ganeti.addnode for new host ganeti2046.codfw.wmnet to cluster codfw and group A * 12:21 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1169.eqiad.wmnet onto db1283.eqiad.wmnet * 12:21 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1169: Pool db1169.eqiad.wmnet in after cloning * 12:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1248-1261].eqiad.wmnet * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1149-1153,1158,1240-1247].eqiad.wmnet * 12:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1149-1153,1158,1240-1247].eqiad.wmnet * 12:11 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1149-1153,1158,1240-1247].eqiad.wmnet * 12:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1168.eqiad.wmnet onto db1282.eqiad.wmnet * 12:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1168: Pool db1168.eqiad.wmnet in after cloning * 12:06 jmm@cumin2003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-test-eqiad * 12:05 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2026.codfw.wmnet * 12:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1149-1153,1158,1240-1247].eqiad.wmnet * 12:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1128-1134,1142-1148].eqiad.wmnet * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2046.codfw.wmnet * 12:01 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1128-1134,1142-1148].eqiad.wmnet * 11:57 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2025.codfw.wmnet * 11:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2025.codfw.wmnet * 11:54 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2046.codfw.wmnet * 11:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1128-1134,1142-1148].eqiad.wmnet * 11:50 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2025.codfw.wmnet * 11:46 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1128-1134,1142-1148].eqiad.wmnet * 11:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1114-1127].eqiad.wmnet * 11:45 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1114-1127].eqiad.wmnet * 11:39 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db2901.codfw.wmnet * 11:39 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db2901.codfw.wmnet * 11:36 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2025.codfw.wmnet * 11:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1114-1127].eqiad.wmnet * 11:35 fceratto@cumin1003: END (ERROR) - Cookbook sre.ganeti.makevm (exit_code=93) for new host db1901.eqiad.wmnet * 11:35 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1169: Pool db1169.eqiad.wmnet in after cloning * 11:35 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 11:30 jmm@cumin2003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-test-eqiad * 11:27 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1114-1127].eqiad.wmnet * 11:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1076-1081,1084-1087,1093-1095,1113].eqiad.wmnet * 11:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1076-1081,1084-1087,1093-1095,1113].eqiad.wmnet * 11:25 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1168: Pool db1168.eqiad.wmnet in after cloning * 11:24 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1194.eqiad.wmnet with OS bookworm * 11:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1181.eqiad.wmnet with OS bookworm * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2050.codfw.wmnet * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2050.codfw.wmnet * 11:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1076-1081,1084-1087,1093-1095,1113].eqiad.wmnet * 11:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2050.codfw.wmnet * 11:14 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2004.codfw.wmnet * 11:10 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2050.codfw.wmnet * 11:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1076-1081,1084-1087,1093-1095,1113].eqiad.wmnet * 11:09 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1286: Pool back * 11:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1045-1050,1056-1057,1064-1066,1073-1075].eqiad.wmnet * 11:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1045-1050,1056-1057,1064-1066,1073-1075].eqiad.wmnet * 11:08 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2004.codfw.wmnet * 11:07 moritzm: installing PHP 8.4 security updates * 11:06 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on an-worker1194.eqiad.wmnet with reason: host reimage * 11:06 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1194.eqiad.wmnet with reason: host reimage * 11:05 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2049.codfw.wmnet * 11:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2049.codfw.wmnet * 11:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid1003.eqiad.wmnet * 11:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1045-1050,1056-1057,1064-1066,1073-1075].eqiad.wmnet * 10:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1181.eqiad.wmnet with reason: host reimage * 10:59 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2049.codfw.wmnet * 10:59 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid1003.eqiad.wmnet * 10:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid2003.codfw.wmnet * 10:54 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1181.eqiad.wmnet with reason: host reimage * 10:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid2003.codfw.wmnet * 10:51 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1901.eqiad.wmnet on all recursors * 10:51 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1901.eqiad.wmnet on all recursors * 10:51 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:51 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:51 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:50 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2049.codfw.wmnet * 10:50 blake@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:50 blake@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:49 blake@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:48 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1045-1050,1056-1057,1064-1066,1073-1075].eqiad.wmnet * 10:48 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1044].eqiad.wmnet * 10:48 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1044].eqiad.wmnet * 10:47 blake@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:46 blake@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:46 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2047.codfw.wmnet * 10:46 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:46 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2047.codfw.wmnet * 10:46 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1901.eqiad.wmnet * 10:46 blake@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:44 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1168: Depool db1168.eqiad.wmnet to then clone it to db1282.eqiad.wmnet - marostegui@cumin1003 * 10:44 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1168: Depool db1168.eqiad.wmnet to then clone it to db1282.eqiad.wmnet - marostegui@cumin1003 * 10:44 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1168.eqiad.wmnet onto db1282.eqiad.wmnet * 10:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2047.codfw.wmnet * 10:40 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1044].eqiad.wmnet * 10:37 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2047.codfw.wmnet * 10:37 fceratto@cumin1003: END (ERROR) - Cookbook sre.ganeti.makevm (exit_code=93) for new host db1901.eqiad.wmnet * 10:36 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 10:34 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:33 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2032.codfw.wmnet * 10:33 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2032.codfw.wmnet * 10:32 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1044].eqiad.wmnet * 10:32 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-eqiad * 10:27 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2032.codfw.wmnet * 10:24 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1286: Pool back * 10:24 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1286 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96163 and previous config saved to /var/cache/conftool/dbconfig/20260818-102431-marostegui.json * 10:22 moritzm: installing Django security updates * 10:20 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul2002.codfw.wmnet * 10:20 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul2001.codfw.wmnet * 10:17 blake@deploy1003: Finished scap sync-world: no-build deployment for [[phab:T417800|T417800]] (duration: 04m 40s) * 10:16 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul2002.codfw.wmnet * 10:16 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul2001.codfw.wmnet * 10:15 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2032.codfw.wmnet * 10:14 blake@deploy1003: Started scap sync-world: no-build deployment for [[phab:T417800|T417800]] * 10:12 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul2003.codfw.wmnet * 10:12 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy2001.codfw.wmnet * 10:12 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy3001.esams.wmnet * 10:12 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy1001.eqiad.wmnet * 10:08 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2031.codfw.wmnet * 10:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2031.codfw.wmnet * 10:08 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul2003.codfw.wmnet * 10:08 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy2001.codfw.wmnet * 10:08 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy3001.esams.wmnet * 10:08 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy1001.eqiad.wmnet * 10:07 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy1002.eqiad.wmnet * 10:05 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy2002.codfw.wmnet * 10:05 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy3002.esams.wmnet * 10:04 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy1002.eqiad.wmnet * 10:03 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy4003.ulsfo.wmnet * 10:03 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy5003.eqsin.wmnet * 10:02 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2031.codfw.wmnet * 10:02 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1169: Depool db1169.eqiad.wmnet to then clone it to db1283.eqiad.wmnet - marostegui@cumin1003 * 10:01 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy2002.codfw.wmnet * 10:01 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy4004.ulsfo.wmnet * 10:01 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy3002.esams.wmnet * 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1169: Depool db1169.eqiad.wmnet to then clone it to db1283.eqiad.wmnet - marostegui@cumin1003 * 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1169.eqiad.wmnet onto db1283.eqiad.wmnet * 10:01 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy5004.eqsin.wmnet * 09:59 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy4003.ulsfo.wmnet * 09:59 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy4004.ulsfo.wmnet * 09:59 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy5003.eqsin.wmnet * 09:59 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy7001.magru.wmnet * 09:59 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy5004.eqsin.wmnet * 09:58 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy7002.magru.wmnet * 09:57 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2031.codfw.wmnet * 09:57 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy6001.drmrs.wmnet * 09:57 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy6002.drmrs.wmnet * 09:54 filippo@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for 10 hosts * 09:54 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast4006.wikimedia.org * 09:53 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy6001.drmrs.wmnet * 09:53 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy6002.drmrs.wmnet * 09:53 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host releases1003.eqiad.wmnet * 09:53 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy7001.magru.wmnet * 09:53 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host people2004.codfw.wmnet * 09:52 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host people1005.eqiad.wmnet * 09:52 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy7002.magru.wmnet * 09:50 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host releases2003.codfw.wmnet * 09:49 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host releases2003.codfw.wmnet * 09:49 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host releases1003.eqiad.wmnet * 09:49 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host people2004.codfw.wmnet * 09:48 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host people1005.eqiad.wmnet * 09:46 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2040.codfw.wmnet * 09:46 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2040.codfw.wmnet * 09:43 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1901.eqiad.wmnet on all recursors * 09:42 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1901.eqiad.wmnet on all recursors * 09:42 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:42 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:42 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2040.codfw.wmnet * 09:40 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp2005.wikimedia.org * 09:36 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp2005.wikimedia.org * 09:31 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:31 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1901.eqiad.wmnet * 09:31 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1901.eqiad.wmnet * 09:31 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:31 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1901.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 09:31 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1901.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 09:29 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2040.codfw.wmnet * 09:29 slyngshede@dns1004: END - running authdns-update * 09:28 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2039.codfw.wmnet * 09:28 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2039.codfw.wmnet * 09:27 slyngshede@dns1004: START - running authdns-update * 09:26 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test2005.wikimedia.org * 09:22 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2039.codfw.wmnet * 09:22 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test2005.wikimedia.org * 09:22 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp1005.wikimedia.org * 09:21 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:19 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2039.codfw.wmnet * 09:19 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2038.codfw.wmnet * 09:18 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1287: Pool back * 09:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2038.codfw.wmnet * 09:18 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp1005.wikimedia.org * 09:18 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test1005.wikimedia.org * 09:17 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1901.eqiad.wmnet * 09:15 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db1901.eqiad.wmnet * 09:15 fceratto@cumin1003: END (ERROR) - Cookbook sre.dns.netbox (exit_code=97) * 09:14 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test1005.wikimedia.org * 09:13 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2038.codfw.wmnet * 09:13 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:13 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1901.eqiad.wmnet * 09:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast4006.wikimedia.org * 09:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1288: Pool back * 09:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast5005.wikimedia.org * 09:03 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2038.codfw.wmnet * 08:57 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2037.codfw.wmnet * 08:57 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast5005.wikimedia.org * 08:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2037.codfw.wmnet * 08:56 filippo@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 10 hosts * 08:52 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2037.codfw.wmnet * 08:51 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1289: Pool back * 08:35 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudidp2001-dev.codfw.wmnet * 08:34 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1002-dev.eqiad.wmnet * 08:33 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1287: Pool back * 08:33 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1287 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96150 and previous config saved to /var/cache/conftool/dbconfig/20260818-083311-marostegui.json * 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1001-dev.eqiad.wmnet * 08:31 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudidp2001-dev.codfw.wmnet * 08:30 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1002-dev.eqiad.wmnet * 08:30 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2037.codfw.wmnet * 08:29 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1001-dev.eqiad.wmnet * 08:28 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2036.codfw.wmnet * 08:28 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2036.codfw.wmnet * 08:25 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1288: Pool back * 08:23 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 08:23 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1181.eqiad.wmnet with OS bookworm * 08:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2036.codfw.wmnet * 08:22 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1288 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96148 and previous config saved to /var/cache/conftool/dbconfig/20260818-082234-marostegui.json * 08:20 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2036.codfw.wmnet * 08:18 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2035.codfw.wmnet * 08:18 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudvirt1057.eqiad.wmnet with OS trixie * 08:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2035.codfw.wmnet * 08:13 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2035.codfw.wmnet * 08:06 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2035.codfw.wmnet * 08:05 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1289: Pool back * 08:05 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1289 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96145 and previous config saved to /var/cache/conftool/dbconfig/20260818-080531-marostegui.json * 07:54 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 07:51 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326439{{!}}Fix "mathjax_ignore" handling around forcemathmode attribute (T434686)]] (duration: 13m 15s) * 07:48 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 07:47 krinkle@deploy1003: krinkle: Continuing with deployment * 07:44 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ganeti2046.codfw.wmnet with OS bookworm * 07:40 krinkle@deploy1003: krinkle: Backport for [[gerrit:1326439{{!}}Fix "mathjax_ignore" handling around forcemathmode attribute (T434686)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:39 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 07:38 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1326439{{!}}Fix "mathjax_ignore" handling around forcemathmode attribute (T434686)]] * 07:34 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 07:31 samwilson@deploy1003: Finished scap sync-world: Backport for [[gerrit:701016{{!}}InitialiseSettings and -labs: Remove redundant feature flag $wgWikisourceEnableOcr (T285311)]] (duration: 07m 47s) * 07:30 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1048.eqiad.wmnet * 07:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1048.eqiad.wmnet * 07:29 XioNoX: add gnmic 0.47.0 to bookworm and trixie reprepro * 07:28 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ganeti2046.codfw.wmnet with reason: host reimage * 07:27 samwilson@deploy1003: samwilson: Continuing with deployment * 07:25 samwilson@deploy1003: samwilson: Backport for [[gerrit:701016{{!}}InitialiseSettings and -labs: Remove redundant feature flag $wgWikisourceEnableOcr (T285311)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:25 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 07:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1048.eqiad.wmnet * 07:24 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ganeti2046.codfw.wmnet with reason: host reimage * 07:23 samwilson@deploy1003: Started scap sync-world: Backport for [[gerrit:701016{{!}}InitialiseSettings and -labs: Remove redundant feature flag $wgWikisourceEnableOcr (T285311)]] * 07:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1198.eqiad.wmnet with OS bookworm * 07:18 samwilson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326450{{!}}InitialiseSettings.php: Enable Bulk OCR on pawikisource (T434648)]] (duration: 12m 17s) * 07:17 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 07:11 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1048.eqiad.wmnet * 07:11 samwilson@deploy1003: samwilson: Continuing with deployment * 07:11 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ganeti2046.codfw.wmnet with OS bookworm * 07:10 samwilson@deploy1003: samwilson: Backport for [[gerrit:1326450{{!}}InitialiseSettings.php: Enable Bulk OCR on pawikisource (T434648)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1057.eqiad.wmnet with OS trixie * 07:05 samwilson@deploy1003: Started scap sync-world: Backport for [[gerrit:1326450{{!}}InitialiseSettings.php: Enable Bulk OCR on pawikisource (T434648)]] * 07:02 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1198.eqiad.wmnet with reason: host reimage * 07:01 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1056.eqiad.wmnet with OS trixie * 07:01 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1056.eqiad.wmnet with OS trixie * 07:00 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1056.eqiad.wmnet with OS trixie * 07:00 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1056.eqiad.wmnet with OS trixie * 06:59 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudvirt1055.eqiad.wmnet with OS trixie * 06:58 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1198.eqiad.wmnet with reason: host reimage * 06:53 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1055.eqiad.wmnet with OS trixie * 06:52 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudvirt1054.eqiad.wmnet with OS trixie * 06:44 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1194.eqiad.wmnet with OS bookworm * 06:44 XioNoX: upgrade eqsin gnmic to 0.47.0 * 06:43 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1198.eqiad.wmnet with OS bookworm * 06:41 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1054.eqiad.wmnet with OS trixie * 06:09 arnaudb@cumin1003: END (PASS) - Cookbook sre.gerrit.restart-gerrit (exit_code=0) Restarting Gerrit on gerrit2002 * 06:06 arnaudb@cumin1003: START - Cookbook sre.gerrit.restart-gerrit Restarting Gerrit on gerrit2002 * 06:06 arnaudb@cumin1003: END (PASS) - Cookbook sre.gerrit.restart-gerrit (exit_code=0) Restarting Gerrit on gerrit1003 * 06:04 arnaudb@cumin1003: START - Cookbook sre.gerrit.restart-gerrit Restarting Gerrit on gerrit1003 * 06:02 arnaudb@cumin1003: END (PASS) - Cookbook sre.gerrit.restart-gerrit (exit_code=0) Restarting Gerrit on gerrit2003 * 06:00 arnaudb@cumin1003: START - Cookbook sre.gerrit.restart-gerrit Restarting Gerrit on gerrit2003 * 05:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1155.eqiad.wmnet with OS bookworm * 05:38 arnaudb: updating prometheusBearerToken on gerrit * 05:28 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1155.eqiad.wmnet with reason: host reimage * 05:23 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1155.eqiad.wmnet with reason: host reimage * 05:06 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1155.eqiad.wmnet with OS bookworm * 04:57 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1144.eqiad.wmnet with OS bookworm * 04:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1144.eqiad.wmnet with reason: host reimage * 04:29 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1144.eqiad.wmnet with reason: host reimage * 04:14 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1144.eqiad.wmnet with OS bookworm * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.13 (duration: 02m 23s) * 03:45 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1197.eqiad.wmnet with OS bookworm * 03:41 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1165.eqiad.wmnet with OS bookworm * 03:38 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.16 refs [[phab:T430835|T430835]] (duration: 34m 43s) * 03:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1164.eqiad.wmnet with OS bookworm * 03:35 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1196.eqiad.wmnet with OS bookworm * 03:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1163.eqiad.wmnet with OS bookworm * 03:22 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1197.eqiad.wmnet with reason: host reimage * 03:18 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1165.eqiad.wmnet with reason: host reimage * 03:15 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1196.eqiad.wmnet with reason: host reimage * 03:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1164.eqiad.wmnet with reason: host reimage * 03:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1163.eqiad.wmnet with reason: host reimage * 03:05 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1165.eqiad.wmnet with reason: host reimage * 03:04 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1164.eqiad.wmnet with reason: host reimage * 03:04 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1197.eqiad.wmnet with reason: host reimage * 03:04 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1196.eqiad.wmnet with reason: host reimage * 03:04 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1163.eqiad.wmnet with reason: host reimage * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.16 refs [[phab:T430835|T430835]] * 02:50 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1197.eqiad.wmnet with OS bookworm * 02:49 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1196.eqiad.wmnet with OS bookworm * 02:49 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1165.eqiad.wmnet with OS bookworm * 02:49 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1164.eqiad.wmnet with OS bookworm * 02:48 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1163.eqiad.wmnet with OS bookworm * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 46s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:15 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-codfw: Set storage compatability to NONE — [[phab:T433026|T433026]] - eevans@cumin1003 * 00:38 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-codfw: Set storage compatability to NONE — [[phab:T433026|T433026]] - eevans@cumin1003 * 00:11 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326373{{!}}PersonalDashboard: add newly renamed *ReviewChangesMlModel setting (T422148)]] (duration: 07m 07s) * 00:09 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-eqiad: Set storage compatability to NONE — [[phab:T433026|T433026]] - eevans@cumin1003 * 00:07 musikanimal@deploy1003: musikanimal: Continuing with deployment * 00:06 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1326373{{!}}PersonalDashboard: add newly renamed *ReviewChangesMlModel setting (T422148)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:04 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1326373{{!}}PersonalDashboard: add newly renamed *ReviewChangesMlModel setting (T422148)]] == 2026-08-17 == * 23:30 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-eqiad: Set storage compatability to NONE — [[phab:T433026|T433026]] - eevans@cumin1003 * 23:05 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-codfw: Set storage compatability to UPGRADING — [[phab:T433026|T433026]] - eevans@cumin1003 * 22:28 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-codfw: Set storage compatability to UPGRADING — [[phab:T433026|T433026]] - eevans@cumin1003 * 21:46 logmsgbot: jforrester Deployed security patch for [[phab:T435085|T435085]] * 21:39 swfrench@deploy1003: mwscript-k8s job started: purgeList.php # [[phab:T432412|T432412]] * 21:37 maryum: Undeploy security fix for [[phab:T433020|T433020]] * 21:23 maryum: Deployed security fix for [[phab:T433020|T433020]] * 21:21 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 21:21 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 21:14 maryum: Deployed security fix for [[phab:T434967|T434967]] * 20:58 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-eqiad: Set storage compatability to UPGRADING — [[phab:T433026|T433026]] - eevans@cumin1003 * 20:40 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324320{{!}}[itwiki/slwiki/tgwiki] Remove temporary Wikipedia 25 logos permanently (already reverted) (T414265 T414320 T415307)]] (duration: 06m 55s) * 20:40 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 20:39 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 20:36 cjming@deploy1003: cjming, superpes: Continuing with deployment * 20:35 cjming@deploy1003: cjming, superpes: Backport for [[gerrit:1324320{{!}}[itwiki/slwiki/tgwiki] Remove temporary Wikipedia 25 logos permanently (already reverted) (T414265 T414320 T415307)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:33 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1324320{{!}}[itwiki/slwiki/tgwiki] Remove temporary Wikipedia 25 logos permanently (already reverted) (T414265 T414320 T415307)]] * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ttmserver-test: apply * 20:30 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326361{{!}}Remove escaped paths in app site association file (T432412)]] (duration: 13m 28s) * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ttmserver-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-toolhub-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-toolhub-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-toolhub-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-toolhub-test: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-toolhub: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-toolhub: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-toolhub: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-toolhub: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-test: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-test: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 20:27 inflatador: bking@deploy1003 `charlie --services_dir dse-k8s-services -s opensearch-* -e dse-k8s-* apply` [[phab:T435125|T435125]] * 20:27 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-apifeatureusage-test: apply * 20:27 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-apifeatureusage-test: apply * 20:27 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-apifeatureusage-test: apply * 20:27 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-apifeatureusage-test: apply * 20:27 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-apifeatureusage: apply * 20:26 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-apifeatureusage: apply * 20:26 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-apifeatureusage: apply * 20:26 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-apifeatureusage: apply * 20:26 cjming@deploy1003: cjming, tsev: Continuing with deployment * 20:24 inflatador: bking@deploy1003 `charlie --services_dir dse-k8s-services -s opensearch-* -e dse-k8s-*` [[phab:T435125|T435125]] * 20:21 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-eqiad: Set storage compatability to UPGRADING — [[phab:T433026|T433026]] - eevans@cumin1003 * 20:19 cjming@deploy1003: cjming, tsev: Backport for [[gerrit:1326361{{!}}Remove escaped paths in app site association file (T432412)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:17 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1326361{{!}}Remove escaped paths in app site association file (T432412)]] * 20:15 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326077{{!}}InstrumentConstructiveEdits: anchor all runs to the nearest `interval` (T431493)]] (duration: 06m 25s) * 20:11 cjming@deploy1003: cjming: Continuing with deployment * 20:10 cjming@deploy1003: cjming: Backport for [[gerrit:1326077{{!}}InstrumentConstructiveEdits: anchor all runs to the nearest `interval` (T431493)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:08 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1326077{{!}}InstrumentConstructiveEdits: anchor all runs to the nearest `interval` (T431493)]] * 20:06 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1160.eqiad.wmnet with OS bookworm * 20:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1162.eqiad.wmnet with OS bookworm * 19:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1184.eqiad.wmnet with OS bookworm * 19:55 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1161.eqiad.wmnet with OS bookworm * 19:49 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1195.eqiad.wmnet with OS bookworm * 19:49 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-test: apply * 19:49 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-test: apply * 19:44 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1160.eqiad.wmnet with reason: host reimage * 19:41 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-codfw: Upgrade to Java 17 — [[phab:T433026|T433026]] - eevans@cumin1003 * 19:39 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1162.eqiad.wmnet with reason: host reimage * 19:36 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-eqsin and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 19:36 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-test: apply * 19:36 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1184.eqiad.wmnet with reason: host reimage * 19:32 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1161.eqiad.wmnet with reason: host reimage * 19:29 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1195.eqiad.wmnet with reason: host reimage * 19:26 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1161.eqiad.wmnet with reason: host reimage * 19:26 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1160.eqiad.wmnet with reason: host reimage * 19:26 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1184.eqiad.wmnet with reason: host reimage * 19:26 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1162.eqiad.wmnet with reason: host reimage * 19:25 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1195.eqiad.wmnet with reason: host reimage * 19:19 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-test: apply * 19:11 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1195.eqiad.wmnet with OS bookworm * 19:10 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1194.eqiad.wmnet with OS bookworm * 19:10 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1184.eqiad.wmnet with OS bookworm * 19:10 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1162.eqiad.wmnet with OS bookworm * 19:10 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1161.eqiad.wmnet with OS bookworm * 19:10 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1160.eqiad.wmnet with OS bookworm * 19:10 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 19:09 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 19:03 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-codfw: Upgrade to Java 17 — [[phab:T433026|T433026]] - eevans@cumin1003 * 18:58 dancy@deploy1003: Finished scap sync-world: testing [[phab:T375514|T375514]] (duration: 03m 13s) * 18:55 dancy@deploy1003: Started scap sync-world: testing [[phab:T375514|T375514]] * 18:55 dwisehaupt@dns1006: END - running authdns-update * 18:54 dancy@deploy1003: Installation of scap version "4.281.1" completed for 3 hosts * 18:54 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2009.codfw.wmnet * 18:53 dwisehaupt@dns1006: START - running authdns-update * 18:52 dancy@deploy1003: Installing scap version "4.281.1" for 3 host(s) * 18:47 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2009.codfw.wmnet * 18:41 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2008.codfw.wmnet * 18:34 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2008.codfw.wmnet * 18:30 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2007.codfw.wmnet * 18:23 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2007.codfw.wmnet * 18:16 dwisehaupt@dns1005: END - running authdns-update * 18:14 dwisehaupt@dns1005: START - running authdns-update * 18:04 swfrench@deploy1003: Finished scap sync-world: Deploy "Point Test Wiki to new docroot" - [[phab:T432412|T432412]] (duration: 26m 03s) * 18:00 swfrench@deploy1003: swfrench: Continuing with deployment * 17:51 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1057.eqiad.wmnet with OS trixie * 17:47 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1180.eqiad.wmnet with OS bookworm * 17:39 swfrench@deploy1003: swfrench: Deploy "Point Test Wiki to new docroot" - [[phab:T432412|T432412]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:39 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-eqsin and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 17:38 swfrench@deploy1003: Started scap sync-world: Deploy "Point Test Wiki to new docroot" - [[phab:T432412|T432412]] * 17:36 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1159.eqiad.wmnet with OS bookworm * 17:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1158.eqiad.wmnet with OS bookworm * 17:30 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-eqiad: Upgrade to Java 17 — [[phab:T433026|T433026]] - eevans@cumin1003 * 17:30 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-ulsfo or A:cp-drmrs and A:cp - 9.2.15 upgrade ([[phab:T434620|T434620]]) * 17:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1157.eqiad.wmnet with OS bookworm * 17:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1180.eqiad.wmnet with reason: host reimage * 17:22 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 17:21 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1193.eqiad.wmnet with OS bookworm * 17:21 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 17:21 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 17:20 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 17:20 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 17:20 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1183.eqiad.wmnet with OS bookworm * 17:18 swfrench@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 17:18 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 17:17 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1180.eqiad.wmnet with reason: host reimage * 17:17 swfrench@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 17:17 swfrench@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 17:16 swfrench@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 17:16 swfrench@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 17:15 swfrench@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 17:15 swfrench@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 17:15 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1192.eqiad.wmnet with OS bookworm * 17:14 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1159.eqiad.wmnet with reason: host reimage * 17:14 swfrench@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 17:13 swfrench@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 17:12 swfrench@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 17:09 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1158.eqiad.wmnet with reason: host reimage * 17:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1157.eqiad.wmnet with reason: host reimage * 17:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1180 * 17:02 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1180 * 17:01 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica-eqsin and A:liberica * 17:01 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1193.eqiad.wmnet with reason: host reimage * 16:57 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1183.eqiad.wmnet with reason: host reimage * 16:56 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1054.eqiad.wmnet with OS trixie * 16:55 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1056.eqiad.wmnet with OS trixie * 16:54 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1193.eqiad.wmnet with reason: host reimage * 16:54 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1192.eqiad.wmnet with reason: host reimage * 16:51 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-eqiad: Upgrade to Java 17 — [[phab:T433026|T433026]] - eevans@cumin1003 * 16:50 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1180 * 16:50 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1180.eqiad.wmnet 17.36.64.10.in-addr.arpa 7.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:50 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1180.eqiad.wmnet 17.36.64.10.in-addr.arpa 7.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:50 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:50 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1180 - btullis@cumin1003" * 16:50 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1180 - btullis@cumin1003" * 16:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1158.eqiad.wmnet with reason: host reimage * 16:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1157.eqiad.wmnet with reason: host reimage * 16:49 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica-eqsin and A:liberica * 16:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1183.eqiad.wmnet with reason: host reimage * 16:48 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1159.eqiad.wmnet with reason: host reimage * 16:47 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1192.eqiad.wmnet with reason: host reimage * 16:46 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns4004.wikimedia.org * 16:46 sukhe@dns1004: END - running authdns-update * 16:44 sukhe@dns1004: START - running authdns-update * 16:44 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns4004.wikimedia.org,service=authdns-update * 16:43 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns4004.wikimedia.org with OS trixie * 16:39 btullis@cumin1003: START - Cookbook sre.dns.netbox * 16:39 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1180 * 16:39 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1193.eqiad.wmnet with OS bookworm * 16:39 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1180.eqiad.wmnet with OS bookworm * 16:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1192.eqiad.wmnet with OS bookworm * 16:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1183.eqiad.wmnet with OS bookworm * 16:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1159.eqiad.wmnet with OS bookworm * 16:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1158.eqiad.wmnet with OS bookworm * 16:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1157.eqiad.wmnet with OS bookworm * 16:31 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1057.eqiad.wmnet with OS trixie * 16:30 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1057.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 16:29 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica-drmrs and A:liberica * 16:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1154.eqiad.wmnet with OS bookworm * 16:23 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1057.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 16:23 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1055.eqiad.wmnet with OS trixie * 16:22 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1057 * 16:22 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1057 * 16:21 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:21 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1057] - vriley@cumin1003" * 16:21 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1057] - vriley@cumin1003" * 16:19 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica-drmrs and A:liberica * 16:17 vriley@cumin1003: START - Cookbook sre.dns.netbox * 16:16 phuedx@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics-external: apply * 16:15 phuedx@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics-external: apply * 16:13 phuedx@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics-external: apply * 16:12 phuedx@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics-external: apply * 16:11 phuedx@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics-external: apply * 16:09 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics-external: apply * 16:09 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1176.eqiad.wmnet with OS bookworm * 16:07 btullis@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 16:06 btullis@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 16:05 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1191.eqiad.wmnet with OS bookworm * 16:03 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1154.eqiad.wmnet with reason: host reimage * 16:03 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching aqs[2002-2012].codfw.wmnet,aqs[1017-1027].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433026|T433026]] - eevans@cumin1003 * 15:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1190.eqiad.wmnet with OS bookworm * 15:59 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1154.eqiad.wmnet with reason: host reimage * 15:55 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-debug: apply * 15:55 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-debug: apply * 15:55 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-debug: apply * 15:55 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/mw-debug: apply * 15:53 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns4004.wikimedia.org with reason: host reimage * 15:50 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns4004.wikimedia.org with reason: host reimage * 15:46 btullis@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 15:46 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1176.eqiad.wmnet with reason: host reimage * 15:45 btullis@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 15:43 btullis@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 15:42 btullis@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 15:42 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1191.eqiad.wmnet with reason: host reimage * 15:39 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1190.eqiad.wmnet with reason: host reimage * 15:36 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1054.eqiad.wmnet with OS trixie * 15:35 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:35 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1056.eqiad.wmnet with OS trixie * 15:35 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:34 moritzm: failover Ganeti master in eqiad to ganeti1046 * 15:34 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1176.eqiad.wmnet with reason: host reimage * 15:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1191.eqiad.wmnet with reason: host reimage * 15:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1190.eqiad.wmnet with reason: host reimage * 15:31 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns2006.wikimedia.org * 15:31 sukhe@dns1004: END - running authdns-update * 15:29 sukhe@dns1004: START - running authdns-update * 15:29 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns2006.wikimedia.org,service=authdns-update * 15:29 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns2006.wikimedia.org * 15:29 sukhe@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns2006.wikimedia.org * 15:26 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica-ulsfo and A:liberica * 15:26 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns1006.wikimedia.org * 15:26 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:25 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns2006.wikimedia.org with OS trixie * 15:25 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1056 * 15:25 sukhe@dns1004: END - running authdns-update * 15:23 sukhe@dns1004: START - running authdns-update * 15:23 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns1006.wikimedia.org,service=authdns-update * 15:23 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns1006.wikimedia.org * 15:23 sukhe@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns1006.wikimedia.org * 15:20 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1056 * 15:20 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:20 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1056~] - vriley@cumin1003" * 15:19 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1056~] - vriley@cumin1003" * 15:19 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns1006.wikimedia.org with OS trixie * 15:19 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns4004.wikimedia.org with OS trixie * 15:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1190.eqiad.wmnet with OS bookworm * 15:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1191.eqiad.wmnet with OS bookworm * 15:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1176.eqiad.wmnet with OS bookworm * 15:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1154.eqiad.wmnet with OS bookworm * 15:16 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica-ulsfo and A:liberica * 15:13 vriley@cumin1003: START - Cookbook sre.dns.netbox * 15:13 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host dns4004.wikimedia.org with OS trixie * 15:12 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1188.eqiad.wmnet with OS bookworm * 15:09 taavi@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318203{{!}}Undeploy WP25EasterEggs (II) (T418134)]] (duration: 08m 56s) * 15:06 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica-magru and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 15:06 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1051.eqiad.wmnet with OS trixie * 15:06 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 15:05 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 15:05 taavi@deploy1003: taavi: Continuing with deployment * 15:04 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-ncredir (exit_code=0) rolling reboot on A:ncredir and A:ncredir * 15:04 taavi@deploy1003: taavi: Backport for [[gerrit:1318203{{!}}Undeploy WP25EasterEggs (II) (T418134)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:03 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1055.eqiad.wmnet with OS trixie * 15:02 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:00 taavi@deploy1003: Started scap sync-world: Backport for [[gerrit:1318203{{!}}Undeploy WP25EasterEggs (II) (T418134)]] * 14:59 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1047.eqiad.wmnet * 14:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1047.eqiad.wmnet * 14:58 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy (exit_code=0) rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 14:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-codfw * 14:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp2001.codfw.wmnet * 14:57 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp2001.codfw.wmnet * 14:57 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica-magru and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 14:57 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:56 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1055 * 14:56 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1055 * 14:55 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns2006.wikimedia.org with reason: host reimage * 14:55 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:55 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1055] - vriley@cumin1003" * 14:55 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1055] - vriley@cumin1003" * 14:54 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1047.eqiad.wmnet * 14:52 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1188.eqiad.wmnet with reason: host reimage * 14:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp2001.codfw.wmnet * 14:51 vriley@cumin1003: START - Cookbook sre.dns.netbox * 14:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp2001.codfw.wmnet * 14:50 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2016.codfw.wmnet * 14:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2016.codfw.wmnet * 14:50 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1189.eqiad.wmnet with OS bookworm * 14:48 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1188.eqiad.wmnet with reason: host reimage * 14:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1051.eqiad.wmnet with reason: host reimage * 14:46 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1182.eqiad.wmnet with OS bookworm * 14:45 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief2002.codfw.wmnet * 14:44 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns1006.wikimedia.org with reason: host reimage * 14:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2016.codfw.wmnet * 14:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1143.eqiad.wmnet with OS bookworm * 14:43 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2016.codfw.wmnet * 14:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2318-2331].codfw.wmnet * 14:43 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2318-2331].codfw.wmnet * 14:42 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching aqs[2002-2012].codfw.wmnet,aqs[1017-1027].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433026|T433026]] - eevans@cumin1003 * 14:41 cgoubert@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326306{{!}}Add placeholder $wmgRedisLockPassword (T366938 T427999)]] (duration: 06m 56s) * 14:41 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief2002.codfw.wmnet * 14:40 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief1002.eqiad.wmnet * 14:39 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1051.eqiad.wmnet with reason: host reimage * 14:38 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns2006.wikimedia.org with reason: host reimage * 14:37 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns1006.wikimedia.org with reason: host reimage * 14:37 cgoubert@deploy1003: cgoubert: Continuing with deployment * 14:36 cgoubert@deploy1003: cgoubert: Backport for [[gerrit:1326306{{!}}Add placeholder $wmgRedisLockPassword (T366938 T427999)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:36 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief1002.eqiad.wmnet * 14:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2318-2331].codfw.wmnet * 14:34 cgoubert@deploy1003: Started scap sync-world: Backport for [[gerrit:1326306{{!}}Add placeholder $wmgRedisLockPassword (T366938 T427999)]] * 14:29 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2318-2331].codfw.wmnet * 14:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2304-2317].codfw.wmnet * 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2304-2317].codfw.wmnet * 14:28 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1054 * 14:27 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1054 * 14:27 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:27 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1054] - vriley@cumin1003" * 14:27 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1054] - vriley@cumin1003" * 14:26 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1189.eqiad.wmnet with reason: host reimage * 14:26 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test2001.codfw.wmnet * 14:25 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test1001.eqiad.wmnet * 14:25 claime: Deploying wmgRedisLockPassword - [[phab:T366938|T366938]] [[phab:T427999|T427999]] * 14:24 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1051.eqiad.wmnet with OS trixie * 14:23 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 14:22 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1182.eqiad.wmnet with reason: host reimage * 14:22 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1051.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:22 vriley@cumin1003: START - Cookbook sre.dns.netbox * 14:22 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test2001.codfw.wmnet * 14:21 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test1001.eqiad.wmnet * 14:21 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 14:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2304-2317].codfw.wmnet * 14:19 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns4004.wikimedia.org with OS trixie * 14:19 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns2006.wikimedia.org with OS trixie * 14:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1143.eqiad.wmnet with reason: host reimage * 14:19 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns1006.wikimedia.org with OS trixie * 14:17 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1189.eqiad.wmnet with reason: host reimage * 14:15 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1182.eqiad.wmnet with reason: host reimage * 14:14 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1143.eqiad.wmnet with reason: host reimage * 14:13 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1051.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:12 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2304-2317].codfw.wmnet * 14:12 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2290-2303].codfw.wmnet * 14:12 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1049.eqiad.wmnet with OS trixie * 14:12 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2290-2303].codfw.wmnet * 14:12 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1051 * 14:11 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1051 * 14:11 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1170.eqiad.wmnet onto db1284.eqiad.wmnet * 14:11 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1170: Pool db1170.eqiad.wmnet in after cloning * 14:09 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 14:09 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:09 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1051] - vriley@cumin1003" * 14:09 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1051] - vriley@cumin1003" * 14:06 klausman@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:04 vriley@cumin1003: START - Cookbook sre.dns.netbox * 14:04 klausman@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2290-2303].codfw.wmnet * 14:02 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1189.eqiad.wmnet with OS bookworm * 14:02 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1188.eqiad.wmnet with OS bookworm * 14:01 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 14:00 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1182.eqiad.wmnet with OS bookworm * 14:00 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1143.eqiad.wmnet with OS bookworm * 13:59 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 13:57 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply * 13:57 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply * 13:56 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply * 13:56 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-ulsfo or A:cp-drmrs and A:cp - 9.2.15 upgrade ([[phab:T434620|T434620]]) * 13:56 cjd91: sudo -i cookbook sre.cdn.roll-upgrade-ats --query 'A:cp-ulsfo or A:cp-drmrs' --task-id [[phab:T434620|T434620]] --reason '9.2.15 upgrade' * 13:56 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply * 13:55 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply * 13:55 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply * 13:54 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 13:54 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 13:54 phuedx@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:53 phuedx@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics-external: apply * 13:52 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1049.eqiad.wmnet with reason: host reimage * 13:51 phuedx@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2290-2303].codfw.wmnet * 13:50 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1047.eqiad.wmnet * 13:50 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2276-2289].codfw.wmnet * 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2276-2289].codfw.wmnet * 13:49 phuedx@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics-external: apply * 13:49 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1046.eqiad.wmnet * 13:49 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1049.eqiad.wmnet with reason: host reimage * 13:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1046.eqiad.wmnet * 13:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-ncredir rolling reboot on A:ncredir and A:ncredir * 13:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 13:46 phuedx@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:44 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics-external: apply * 13:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1046.eqiad.wmnet * 13:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2276-2289].codfw.wmnet * 13:36 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1046.eqiad.wmnet * 13:36 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1045.eqiad.wmnet * 13:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1045.eqiad.wmnet * 13:34 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2276-2289].codfw.wmnet * 13:34 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1049.eqiad.wmnet with OS trixie * 13:34 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2262-2275].codfw.wmnet * 13:34 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2262-2275].codfw.wmnet * 13:32 Lucas_WMDE: UTC afternoon backport+config window done * 13:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1045.eqiad.wmnet * 13:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2262-2275].codfw.wmnet * 13:26 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1049.eqiad.wmnet with OS trixie * 13:26 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1049.eqiad.wmnet with OS trixie * 13:26 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1170: Pool db1170.eqiad.wmnet in after cloning * 13:23 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1049.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 13:23 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1045.eqiad.wmnet * 13:21 atsukoito: manually done sudo -i docker-registryctl --debug delete-tags 'docker-registry.discovery.wmnet/repos/data-engineering/airflow-dags:airflow-3.3.0-py3.11-2026-08-17-*' to remove incorrect tags * 13:19 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311141{{!}}viwiki: Set `noindex,nofollow` for User and User talk (T432311)]] (duration: 11m 11s) * 13:15 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2262-2275].codfw.wmnet * 13:14 lucaswerkmeister-wmde@deploy1003: ndkdd, lucaswerkmeister-wmde: Continuing with deployment * 13:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2248-2261].codfw.wmnet * 13:14 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2248-2261].codfw.wmnet * 13:13 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1049.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 13:12 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1049 * 13:11 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1049 * 13:10 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:10 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1049] - vriley@cumin1003" * 13:10 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1049] - vriley@cumin1003" * 13:10 lucaswerkmeister-wmde@deploy1003: ndkdd, lucaswerkmeister-wmde: Backport for [[gerrit:1311141{{!}}viwiki: Set `noindex,nofollow` for User and User talk (T432311)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1037.eqiad.wmnet * 13:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1037.eqiad.wmnet * 13:08 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1311141{{!}}viwiki: Set `noindex,nofollow` for User and User talk (T432311)]] * 13:06 vriley@cumin1003: START - Cookbook sre.dns.netbox * 13:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2248-2261].codfw.wmnet * 13:00 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1037.eqiad.wmnet * 12:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2248-2261].codfw.wmnet * 12:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2204-2215,2242-2243].codfw.wmnet * 12:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2204-2215,2242-2243].codfw.wmnet * 12:49 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324832{{!}}Migrate $wgFlaggedRevsTags from flaggedrevs.php to ext-FlaggedRevs.php]] (duration: 14m 02s) * 12:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint2001.codfw.wmnet * 12:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2204-2215,2242-2243].codfw.wmnet * 12:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint2001.codfw.wmnet * 12:41 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1037.eqiad.wmnet * 12:40 ladsgroup@deploy1003: Rolling back deployment * 12:37 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1324832{{!}}Migrate $wgFlaggedRevsTags from flaggedrevs.php to ext-FlaggedRevs.php]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:37 seanleong-wmde: Finished populateSitesTable for [bolwiki] ([[[phab:T429955|T429955]]]) * 12:35 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1324832{{!}}Migrate $wgFlaggedRevsTags from flaggedrevs.php to ext-FlaggedRevs.php]] * 12:35 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2204-2215,2242-2243].codfw.wmnet * 12:34 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2190-2203].codfw.wmnet * 12:34 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2190-2203].codfw.wmnet * 12:32 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1028.eqiad.wmnet * 12:32 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1028.eqiad.wmnet * 12:31 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint1001.eqiad.wmnet * 12:30 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1170: Depool db1170.eqiad.wmnet to then clone it to db1284.eqiad.wmnet - marostegui@cumin1003 * 12:28 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint1001.eqiad.wmnet * 12:28 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1170: Depool db1170.eqiad.wmnet to then clone it to db1284.eqiad.wmnet - marostegui@cumin1003 * 12:27 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1170.eqiad.wmnet onto db1284.eqiad.wmnet * 12:27 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db2901.codfw.wmnet * 12:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2190-2203].codfw.wmnet * 12:27 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 12:26 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1028.eqiad.wmnet * 12:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2190-2203].codfw.wmnet * 12:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2172-2179,2184-2189].codfw.wmnet * 12:18 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2172-2179,2184-2189].codfw.wmnet * 12:13 seanleong-wmde@deploy1003: mwscript-k8s job started: foreachwikiindblist wikidataclient extensions/Wikibase/lib/maintenance/populateSitesTable.php --force-protocol https # [[phab:T429955|T429955]] * 12:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2172-2179,2184-2189].codfw.wmnet * 12:09 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1028.eqiad.wmnet * 12:03 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 12:02 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db2901.codfw.wmnet * 12:02 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db2901.codfw.wmnet * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1027.eqiad.wmnet * 12:02 fceratto@cumin1003: END (ERROR) - Cookbook sre.dns.netbox (exit_code=97) * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1027.eqiad.wmnet * 12:02 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 12:02 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db2901.codfw.wmnet * 12:01 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2172-2179,2184-2189].codfw.wmnet * 12:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2158-2171].codfw.wmnet * 12:01 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2158-2171].codfw.wmnet * 11:58 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db1902.eqiad.wmnet * 11:58 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 11:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1027.eqiad.wmnet * 11:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2158-2171].codfw.wmnet * 11:51 jayme: updated calico to v3.30.7 on wikikube eqiad - [[phab:T427400|T427400]] * 11:50 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1027.eqiad.wmnet * 11:45 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2158-2171].codfw.wmnet * 11:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2144-2157].codfw.wmnet * 11:44 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2144-2157].codfw.wmnet * 11:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1058.eqiad.wmnet * 11:43 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1058.eqiad.wmnet * 11:43 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 11:43 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 11:43 marostegui@cumin1003: Removing db1153 from zarcillo [[phab:T434638|T434638]] * 11:42 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1153.eqiad.wmnet * 11:42 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:42 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1153.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 11:42 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1172.eqiad.wmnet onto db1286.eqiad.wmnet * 11:42 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1172: Pool db1172.eqiad.wmnet in after cloning * 11:42 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1153.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 11:41 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1902.eqiad.wmnet * 11:41 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=97) for new host db2901.codfw.wmnet * 11:41 fceratto@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host db2901.codfw.wmnet with OS trixie * 11:38 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'. * 11:38 marostegui@cumin1003: START - Cookbook sre.dns.netbox * 11:37 marostegui@dns1004: END - running authdns-update * 11:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1058.eqiad.wmnet * 11:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2144-2157].codfw.wmnet * 11:35 marostegui@dns1004: START - running authdns-update * 11:32 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1153.eqiad.wmnet * 11:32 marostegui@cumin1003: START - Cookbook sre.mysql.decommission * 11:28 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326249{{!}}ImagePage: move TOC element below file link (T332644)]] (duration: 09m 56s) * 11:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2144-2157].codfw.wmnet * 11:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2130-2143].codfw.wmnet * 11:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2130-2143].codfw.wmnet * 11:26 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1058.eqiad.wmnet * 11:23 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 11:22 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'. * 11:22 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'. * 11:22 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'. * 11:22 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326249{{!}}ImagePage: move TOC element below file link (T332644)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1057.eqiad.wmnet * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1057.eqiad.wmnet * 11:21 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 11:20 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 11:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2130-2143].codfw.wmnet * 11:19 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326249{{!}}ImagePage: move TOC element below file link (T332644)]] * 11:18 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply * 11:17 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply * 11:17 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply * 11:16 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 11:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1057.eqiad.wmnet * 11:15 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 11:14 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 11:13 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 11:13 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db2901.codfw.wmnet with OS trixie * 11:12 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db2901.codfw.wmnet - fceratto@cumin1003" * 11:12 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db2901.codfw.wmnet - fceratto@cumin1003" * 11:12 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db2901.codfw.wmnet on all recursors * 11:12 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db2901.codfw.wmnet on all recursors * 11:12 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:12 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db2901.codfw.wmnet - fceratto@cumin1003" * 11:12 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2130-2143].codfw.wmnet * 11:11 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2107-2115,2124-2129].codfw.wmnet * 11:11 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2107-2115,2124-2129].codfw.wmnet * 11:11 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 11:11 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 11:11 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply * 11:11 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 11:10 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 11:06 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db2901.codfw.wmnet - fceratto@cumin1003" * 11:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2107-2115,2124-2129].codfw.wmnet * 10:57 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1172: Pool db1172.eqiad.wmnet in after cloning * 10:54 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2107-2115,2124-2129].codfw.wmnet * 10:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2078,2087-2095,2102-2106].codfw.wmnet * 10:54 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2078,2087-2095,2102-2106].codfw.wmnet * 10:47 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply * 10:46 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply * 10:46 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply * 10:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2078,2087-2095,2102-2106].codfw.wmnet * 10:45 blake@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply * 10:38 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1057.eqiad.wmnet * 10:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1056.eqiad.wmnet * 10:38 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1056.eqiad.wmnet * 10:37 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2078,2087-2095,2102-2106].codfw.wmnet * 10:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2061-2062,2064-2065,2067-2077].codfw.wmnet * 10:36 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2061-2062,2064-2065,2067-2077].codfw.wmnet * 10:32 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1056.eqiad.wmnet * 10:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2061-2062,2064-2065,2067-2077].codfw.wmnet * 10:25 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:25 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db2901.codfw.wmnet * 10:25 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db1902.eqiad.wmnet * 10:25 fceratto@cumin1003: END (ERROR) - Cookbook sre.dns.netbox (exit_code=97) * 10:24 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:24 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1902.eqiad.wmnet * 10:24 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=97) for new host db1902.eqiad.wmnet * 10:24 fceratto@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host db1902.eqiad.wmnet with OS trixie * 10:24 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=93) for new host db1903.eqiad.wmnet * 10:24 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 10:20 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1056.eqiad.wmnet * 10:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2061-2062,2064-2065,2067-2077].codfw.wmnet * 10:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2038-2039,2041-2042,2044,2046,2049-2051,2055-2060].codfw.wmnet * 10:18 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2038-2039,2041-2042,2044,2046,2049-2051,2055-2060].codfw.wmnet * 10:15 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1055.eqiad.wmnet * 10:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1055.eqiad.wmnet * 10:13 Amir1: mwscript-k8s --dblist=all -- purgeUserOptions.php --login-age 5 uls-preferences ([[phab:T406724|T406724]]) * 10:11 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:10 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1903.eqiad.wmnet on all recursors * 10:10 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1903.eqiad.wmnet on all recursors * 10:10 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2038-2039,2041-2042,2044,2046,2049-2051,2055-2060].codfw.wmnet * 10:10 moritzm: installing unzip security updates * 10:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1055.eqiad.wmnet * 10:08 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=97) for new host db1901.eqiad.wmnet * 10:08 fceratto@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host db1901.eqiad.wmnet with OS trixie * 10:08 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:08 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 10:08 fceratto@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1903.eqiad.wmnet - fceratto@cumin1003" * 10:07 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db1902.eqiad.wmnet with OS trixie * 10:07 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1902.eqiad.wmnet - fceratto@cumin1003" * 10:07 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1902.eqiad.wmnet - fceratto@cumin1003" * 10:04 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply * 10:04 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326227{{!}}Enable desktop/native lazy loading everywhere (T148047)]] (duration: 07m 13s) * 10:03 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1902.eqiad.wmnet on all recursors * 10:03 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1902.eqiad.wmnet on all recursors * 10:03 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:03 blake@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply * 10:01 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1903.eqiad.wmnet - fceratto@cumin1003" * 10:01 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2038-2039,2041-2042,2044,2046,2049-2051,2055-2060].codfw.wmnet * 10:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2002,2005-2006,2011-2015,2017-2018,2033-2037].codfw.wmnet * 10:01 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2002,2005-2006,2011-2015,2017-2018,2033-2037].codfw.wmnet * 10:00 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:00 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 09:59 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 09:59 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1055.eqiad.wmnet * 09:58 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326227{{!}}Enable desktop/native lazy loading everywhere (T148047)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:56 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326227{{!}}Enable desktop/native lazy loading everywhere (T148047)]] * 09:53 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2002,2005-2006,2011-2015,2017-2018,2033-2037].codfw.wmnet * 09:50 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:48 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1903.eqiad.wmnet * 09:48 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:48 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1902.eqiad.wmnet * 09:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 09:44 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2002,2005-2006,2011-2015,2017-2018,2033-2037].codfw.wmnet * 09:43 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-codfw * 09:43 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1174.eqiad.wmnet onto db1288.eqiad.wmnet * 09:42 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1174: Pool db1174.eqiad.wmnet in after cloning * 09:40 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1172: Depool db1172.eqiad.wmnet to then clone it to db1286.eqiad.wmnet - marostegui@cumin1003 * 09:39 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1172: Depool db1172.eqiad.wmnet to then clone it to db1286.eqiad.wmnet - marostegui@cumin1003 * 09:39 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1172.eqiad.wmnet onto db1286.eqiad.wmnet * 09:33 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1175.eqiad.wmnet onto db1289.eqiad.wmnet * 09:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1175: Pool db1175.eqiad.wmnet in after cloning * 09:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1201.eqiad.wmnet onto db1287.eqiad.wmnet * 09:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1201: Pool db1201.eqiad.wmnet in after cloning * 09:28 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db1901.eqiad.wmnet with OS trixie * 09:27 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:27 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:27 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1901.eqiad.wmnet on all recursors * 09:27 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1901.eqiad.wmnet on all recursors * 09:26 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:26 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:26 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:15 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:15 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1901.eqiad.wmnet * 09:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw1001.wikimedia.org with OS trixie * 08:57 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1174: Pool db1174.eqiad.wmnet in after cloning * 08:54 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1054.eqiad.wmnet * 08:54 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1054.eqiad.wmnet * 08:48 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1054.eqiad.wmnet * 08:47 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1175: Pool db1175.eqiad.wmnet in after cloning * 08:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 08:46 marostegui@cumin1003: Removing db1152 from zarcillo [[phab:T434480|T434480]] * 08:46 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1152.eqiad.wmnet * 08:46 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:46 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1152.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 08:46 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1201: Pool db1201.eqiad.wmnet in after cloning * 08:46 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1152.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 08:46 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1054.eqiad.wmnet * 08:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1053.eqiad.wmnet * 08:43 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1053.eqiad.wmnet * 08:42 marostegui@cumin1003: START - Cookbook sre.dns.netbox * 08:38 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage * 08:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1053.eqiad.wmnet * 08:36 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1152.eqiad.wmnet * 08:36 marostegui@cumin1003: START - Cookbook sre.mysql.decommission * 08:35 phuedx: UTC morning backport window done * 08:35 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1053.eqiad.wmnet * 08:34 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1035.eqiad.wmnet * 08:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1035.eqiad.wmnet * 08:34 phuedx@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324725{{!}}EventStreamConfig: Mark product_metrics.web_base and .web_base_with_ip as Test Kitchen streams (T429898 T430322)]], [[gerrit:1313923{{!}}EventStreamConfig: Remove unused web_ui_scroll* streams (T415370)]], [[gerrit:1325546{{!}}EventStreamConfig: Remove Watchlist click stream (T434790)]] (duration: 12m 42s) * 08:33 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1281: Pool back * 08:32 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage * 08:31 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on db2209.codfw.wmnet with reason: Maintenance * 08:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2209: Maintenance needed * 08:30 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2209: Maintenance needed * 08:26 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1035.eqiad.wmnet * 08:26 phuedx@deploy1003: bearloga, phuedx: Continuing with deployment * 08:23 phuedx@deploy1003: bearloga, phuedx: Backport for [[gerrit:1324725{{!}}EventStreamConfig: Mark product_metrics.web_base and .web_base_with_ip as Test Kitchen streams (T429898 T430322)]], [[gerrit:1313923{{!}}EventStreamConfig: Remove unused web_ui_scroll* streams (T415370)]], [[gerrit:1325546{{!}}EventStreamConfig: Remove Watchlist click stream (T434790)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug * 08:21 phuedx@deploy1003: Started scap sync-world: Backport for [[gerrit:1324725{{!}}EventStreamConfig: Mark product_metrics.web_base and .web_base_with_ip as Test Kitchen streams (T429898 T430322)]], [[gerrit:1313923{{!}}EventStreamConfig: Remove unused web_ui_scroll* streams (T415370)]], [[gerrit:1325546{{!}}EventStreamConfig: Remove Watchlist click stream (T434790)]] * 08:19 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw1001.wikimedia.org with OS trixie * 08:18 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1035.eqiad.wmnet * 08:17 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1032.eqiad.wmnet * 08:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1032.eqiad.wmnet * 08:16 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1279: Pool back * 08:15 phuedx@deploy1003: Finished scap sync-world: Backport for [[gerrit:1216721{{!}}viwikivoyage: enable relatedarticle and pop-up (T405724)]] (duration: 39m 12s) * 08:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1032.eqiad.wmnet * 08:09 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1032.eqiad.wmnet * 08:08 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1031.eqiad.wmnet * 08:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1031.eqiad.wmnet * 08:03 godog: switch production to use dumps-nfs.w.o - [[phab:T432212|T432212]] * 08:02 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1031.eqiad.wmnet * 08:02 phuedx@deploy1003: nvdtn19, phuedx: Continuing with deployment * 08:00 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1201: Depool db1201.eqiad.wmnet to then clone it to db1287.eqiad.wmnet - marostegui@cumin1003 * 08:00 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1201: Depool db1201.eqiad.wmnet to then clone it to db1287.eqiad.wmnet - marostegui@cumin1003 * 08:00 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1201.eqiad.wmnet onto db1287.eqiad.wmnet * 08:00 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1057.eqiad.wmnet * 08:00 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:00 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1057.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:59 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1057.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:59 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1275: Pool back * 07:55 filippo@cumin1003: START - Cookbook sre.dns.netbox * 07:55 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1031.eqiad.wmnet * 07:52 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1030.eqiad.wmnet * 07:52 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1030.eqiad.wmnet * 07:52 phuedx@deploy1003: nvdtn19, phuedx: Backport for [[gerrit:1216721{{!}}viwikivoyage: enable relatedarticle and pop-up (T405724)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:51 tappof: bump space for prometheus k8s-aux in eqiad * 07:50 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1057.eqiad.wmnet * 07:48 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1281: Pool back * 07:47 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1281 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96108 and previous config saved to /var/cache/conftool/dbconfig/20260817-074749-marostegui.json * 07:46 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1030.eqiad.wmnet * 07:42 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1030.eqiad.wmnet * 07:41 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1174: Depool db1174.eqiad.wmnet to then clone it to db1288.eqiad.wmnet - marostegui@cumin1003 * 07:41 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1174: Depool db1174.eqiad.wmnet to then clone it to db1288.eqiad.wmnet - marostegui@cumin1003 * 07:41 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1174.eqiad.wmnet onto db1288.eqiad.wmnet * 07:40 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1029.eqiad.wmnet * 07:40 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm2001.wikimedia.org * 07:40 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1029.eqiad.wmnet * 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1056.eqiad.wmnet * 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1056.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:38 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1056.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:36 phuedx@deploy1003: Started scap sync-world: Backport for [[gerrit:1216721{{!}}viwikivoyage: enable relatedarticle and pop-up (T405724)]] * 07:36 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm2001.wikimedia.org * 07:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1279 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96104 and previous config saved to /var/cache/conftool/dbconfig/20260817-073542-marostegui.json * 07:34 filippo@cumin1003: START - Cookbook sre.dns.netbox * 07:34 slyngshede@dns1004: END - running authdns-update * 07:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1029.eqiad.wmnet * 07:33 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm-test1001.wikimedia.org * 07:32 slyngshede@dns1004: START - running authdns-update * 07:31 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1029.eqiad.wmnet * 07:31 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1279: Pool back * 07:30 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1279 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96102 and previous config saved to /var/cache/conftool/dbconfig/20260817-073038-marostegui.json * 07:29 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm-test1001.wikimedia.org * 07:29 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm1001.wikimedia.org * 07:28 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1056.eqiad.wmnet * 07:28 moritzm: extend the disk of ldap-rw1001 by 80G [[phab:T331699|T331699]] * 07:28 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1055.eqiad.wmnet * 07:28 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:28 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1055.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:27 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1055.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:26 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1044.eqiad.wmnet * 07:26 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1044.eqiad.wmnet * 07:25 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm1001.wikimedia.org * 07:24 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1175: Depool db1175.eqiad.wmnet to then clone it to db1289.eqiad.wmnet - marostegui@cumin1003 * 07:24 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1175: Depool db1175.eqiad.wmnet to then clone it to db1289.eqiad.wmnet - marostegui@cumin1003 * 07:24 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1175.eqiad.wmnet onto db1289.eqiad.wmnet * 07:22 filippo@cumin1003: START - Cookbook sre.dns.netbox * 07:20 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1044.eqiad.wmnet * 07:16 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1055.eqiad.wmnet * 07:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1054.eqiad.wmnet * 07:15 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:15 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1054.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:15 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1044.eqiad.wmnet * 07:15 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1054.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:13 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1275: Pool back * 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1275 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96098 and previous config saved to /var/cache/conftool/dbconfig/20260817-071225-marostegui.json * 07:10 filippo@cumin1003: START - Cookbook sre.dns.netbox * 07:05 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1043.eqiad.wmnet * 07:05 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1054.eqiad.wmnet * 07:05 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1051.eqiad.wmnet * 07:05 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:05 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1051.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:05 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1043.eqiad.wmnet * 07:04 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1051.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 06:59 filippo@cumin1003: START - Cookbook sre.dns.netbox * 06:59 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin1001.eqiad.wmnet * 06:59 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1043.eqiad.wmnet * 06:59 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin2001.codfw.wmnet * 06:55 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin2001.codfw.wmnet * 06:55 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1051.eqiad.wmnet * 06:54 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1049.eqiad.wmnet * 06:54 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:54 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1049.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 06:54 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1049.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 06:54 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin1001.eqiad.wmnet * 06:53 moritzm: installing apr-util security updates * 06:52 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1043.eqiad.wmnet * 06:49 filippo@cumin1003: START - Cookbook sre.dns.netbox * 06:41 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1049.eqiad.wmnet * 06:13 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit2003.wikimedia.org * 06:13 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet * 06:07 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet * 06:06 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit2003.wikimedia.org * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 47s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-16 == * 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 01m 03s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-15 == * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 41s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-14 == * 15:38 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-staging-master-eqiad * 15:38 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster1005.eqiad.wmnet * 15:38 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster1005.eqiad.wmnet * 15:35 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sretest2009.codfw.wmnet * 15:33 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster1005.eqiad.wmnet * 15:33 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster1005.eqiad.wmnet * 15:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster1004.eqiad.wmnet * 15:32 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster1004.eqiad.wmnet * 15:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host sretest2009.codfw.wmnet * 15:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster1004.eqiad.wmnet * 15:27 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster1004.eqiad.wmnet * 15:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster1003.eqiad.wmnet * 15:27 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster1003.eqiad.wmnet * 15:24 dancy@deploy1003: Finished scap sync-world: testing (duration: 03m 23s) * 15:22 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster1003.eqiad.wmnet * 15:22 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster1003.eqiad.wmnet * 15:22 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-staging-master-eqiad * 15:20 dancy@deploy1003: Started scap sync-world: testing * 15:20 dancy@deploy1003: Installation of scap version "4.280.2" completed for 3 hosts * 15:18 dancy@deploy1003: Installing scap version "4.280.2" for 3 host(s) * 15:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sretest2006.codfw.wmnet * 14:54 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host sretest2006.codfw.wmnet * 14:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sretest2003.codfw.wmnet * 14:39 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host sretest2003.codfw.wmnet * 13:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt-staging2001.codfw.wmnet * 13:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt-staging2001.codfw.wmnet * 13:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-staging-master-codfw * 13:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster2005.codfw.wmnet * 13:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster2005.codfw.wmnet * 13:05 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox-dev2003.codfw.wmnet * 13:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster2005.codfw.wmnet * 13:04 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster2005.codfw.wmnet * 13:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster2004.codfw.wmnet * 13:04 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster2004.codfw.wmnet * 13:01 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netbox-dev2003.codfw.wmnet * 12:59 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster2004.codfw.wmnet * 12:59 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster2004.codfw.wmnet * 12:59 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster2003.codfw.wmnet * 12:59 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster2003.codfw.wmnet * 12:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster2003.codfw.wmnet * 12:54 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster2003.codfw.wmnet * 12:54 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-staging-master-codfw * 12:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-staging-worker-eqiad * 12:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage1006.eqiad.wmnet * 12:52 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage1006.eqiad.wmnet * 12:46 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage1006.eqiad.wmnet * 12:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw1001.wikimedia.org with OS trixie * 12:45 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage1006.eqiad.wmnet * 12:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage1005.eqiad.wmnet * 12:45 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage1005.eqiad.wmnet * 12:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage1005.eqiad.wmnet * 12:36 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/ratelimit: apply * 12:35 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/ratelimit: apply * 12:35 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:35 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:33 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage1005.eqiad.wmnet * 12:33 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage1004.eqiad.wmnet * 12:33 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage1004.eqiad.wmnet * 12:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage1004.eqiad.wmnet * 12:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage1004.eqiad.wmnet * 12:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage1003.eqiad.wmnet * 12:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage1003.eqiad.wmnet * 12:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage * 12:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage1003.eqiad.wmnet * 12:17 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage * 12:14 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage1003.eqiad.wmnet * 12:14 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-staging-worker-eqiad * 12:03 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw1001.wikimedia.org with OS trixie * 12:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cuminunpriv1001.eqiad.wmnet * 11:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cuminunpriv1001.eqiad.wmnet * 11:27 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1004.wikimedia.org * 11:24 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1153 from dbctl [[phab:T434638|T434638]]', diff saved to https://phabricator.wikimedia.org/P96097 and previous config saved to /var/cache/conftool/dbconfig/20260814-112449-marostegui.json * 11:21 aokoth@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1004.wikimedia.org * 11:20 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 11:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-staging-worker-codfw * 11:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2004.codfw.wmnet * 11:17 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2004.codfw.wmnet * 11:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2004.codfw.wmnet * 11:10 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2004.codfw.wmnet * 11:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2003.codfw.wmnet * 11:10 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2003.codfw.wmnet * 11:03 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-worker1181.eqiad.wmnet with OS bookworm * 11:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2003.codfw.wmnet * 11:03 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2003.codfw.wmnet * 11:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2002.codfw.wmnet * 11:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2002.codfw.wmnet * 10:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1235.eqiad.wmnet with OS bookworm * 10:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock2003.codfw.wmnet * 10:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock2003.codfw.wmnet with OS trixie * 10:56 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2002.codfw.wmnet * 10:54 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1153.eqiad.wmnet with OS bookworm * 10:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2002.codfw.wmnet * 10:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2001.codfw.wmnet * 10:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2001.codfw.wmnet * 10:50 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 10:49 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1187.eqiad.wmnet with OS bookworm * 10:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2001.codfw.wmnet * 10:44 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2001.codfw.wmnet * 10:44 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-staging-worker-codfw * 10:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock2003.codfw.wmnet with reason: host reimage * 10:38 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock2003.codfw.wmnet with reason: host reimage * 10:35 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1153.eqiad.wmnet with reason: host reimage * 10:32 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1235.eqiad.wmnet with reason: host reimage * 10:29 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1187.eqiad.wmnet with reason: host reimage * 10:24 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1153.eqiad.wmnet with reason: host reimage * 10:23 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1235.eqiad.wmnet with reason: host reimage * 10:21 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1187.eqiad.wmnet with reason: host reimage * 10:16 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock2003.codfw.wmnet with OS trixie * 10:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install1005.wikimedia.org * 10:12 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 10:12 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 10:12 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 10:12 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 10:12 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:12 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 10:12 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 10:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install1005.wikimedia.org * 10:07 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1235.eqiad.wmnet with OS bookworm * 10:07 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1187.eqiad.wmnet with OS bookworm * 10:07 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1181.eqiad.wmnet with OS bookworm * 10:07 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1153.eqiad.wmnet with OS bookworm * 10:07 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install2005.wikimedia.org * 10:03 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 10:03 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2003.codfw.wmnet * 10:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1232.eqiad.wmnet with OS bookworm * 10:00 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install2005.wikimedia.org * 10:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install3004.wikimedia.org * 09:58 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 09:58 fceratto@cumin1003: Removing db1151 from zarcillo [[phab:T434538|T434538]] * 09:56 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1231.eqiad.wmnet with OS bookworm * 09:56 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 09:53 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 09:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install3004.wikimedia.org * 09:50 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1152.eqiad.wmnet with OS bookworm * 09:49 Dreamy_Jazz: `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260808000000" --end-timestamp="20260812120000" --sleep="5" --batch-size="50"` for [[phab:T434688|T434688]] * 09:48 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host centrallog2002.codfw.wmnet * 09:48 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install4004.wikimedia.org * 09:42 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1232.eqiad.wmnet with reason: host reimage * 09:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install4004.wikimedia.org * 09:41 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host centrallog2002.codfw.wmnet * 09:39 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install5004.wikimedia.org * 09:36 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1232.eqiad.wmnet with reason: host reimage * 09:36 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'. * 09:34 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'. * 09:33 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1231.eqiad.wmnet with reason: host reimage * 09:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install5004.wikimedia.org * 09:32 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host centrallog1002.eqiad.wmnet * 09:30 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install6003.wikimedia.org * 09:30 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1231.eqiad.wmnet with reason: host reimage * 09:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1152.eqiad.wmnet with reason: host reimage * 09:25 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host centrallog1002.eqiad.wmnet * 09:25 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1152.eqiad.wmnet with reason: host reimage * 09:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install6003.wikimedia.org * 09:22 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1232 * 09:22 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1232 * 09:22 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1232 * 09:22 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1232.eqiad.wmnet 25.53.64.10.in-addr.arpa 5.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:22 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host titan1001.eqiad.wmnet * 09:22 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1232.eqiad.wmnet 25.53.64.10.in-addr.arpa 5.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:22 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:22 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1232 - btullis@cumin1003" * 09:22 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1232 - btullis@cumin1003" * 09:17 btullis@cumin1003: START - Cookbook sre.dns.netbox * 09:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install7002.wikimedia.org * 09:17 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1232 * 09:16 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1231 * 09:16 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1231 * 09:14 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1231 * 09:14 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1231.eqiad.wmnet 24.53.64.10.in-addr.arpa 4.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host titan1001.eqiad.wmnet * 09:14 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1231.eqiad.wmnet 24.53.64.10.in-addr.arpa 4.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:14 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1231 - btullis@cumin1003" * 09:14 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1231 - btullis@cumin1003" * 09:10 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install7002.wikimedia.org * 09:09 btullis@cumin1003: START - Cookbook sre.dns.netbox * 09:08 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1231 * 09:08 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1152 * 09:08 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1152 * 09:06 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1152 * 09:06 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1152.eqiad.wmnet 16.53.64.10.in-addr.arpa 6.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:06 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1152.eqiad.wmnet 16.53.64.10.in-addr.arpa 6.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:06 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:06 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1152 - btullis@cumin1003" * 09:06 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1152 - btullis@cumin1003" * 09:03 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1151.eqiad.wmnet * 09:03 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:03 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1151.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 08:55 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host titan2001.codfw.wmnet * 08:55 btullis@cumin1003: START - Cookbook sre.dns.netbox * 08:54 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1151.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 08:54 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1152 * 08:53 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1232.eqiad.wmnet with OS bookworm * 08:53 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1231.eqiad.wmnet with OS bookworm * 08:53 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1152.eqiad.wmnet with OS bookworm * 08:51 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2209: Pool back * 08:51 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1201.eqiad.wmnet * 08:50 btullis@cumin1003: START - Cookbook sre.hosts.remove-downtime for an-worker1201.eqiad.wmnet * 08:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1229.eqiad.wmnet with OS bookworm * 08:47 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host titan2001.codfw.wmnet * 08:46 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping2004.codfw.wmnet * 08:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ping2004.codfw.wmnet * 08:40 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host an-worker1230.eqiad.wmnet with OS bookworm * 08:40 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 08:34 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1151.eqiad.wmnet * 08:34 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 08:31 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host titan1002.eqiad.wmnet * 08:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1229.eqiad.wmnet with reason: host reimage * 08:25 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host titan1002.eqiad.wmnet * 08:25 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1229.eqiad.wmnet with reason: host reimage * 08:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping1004.eqiad.wmnet * 08:21 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ping1004.eqiad.wmnet * 08:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1230.eqiad.wmnet with reason: host reimage * 08:12 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1230.eqiad.wmnet with reason: host reimage * 08:11 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1229 * 08:11 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1229 * 08:11 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1229 * 08:11 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1229.eqiad.wmnet 22.53.64.10.in-addr.arpa 2.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:11 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1229.eqiad.wmnet 22.53.64.10.in-addr.arpa 2.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:11 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:11 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1229 - btullis@cumin1003" * 08:11 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1229 - btullis@cumin1003" * 08:10 btullis@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1201.eqiad.wmnet with reason: Fixing a disk * 08:07 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host titan2002.codfw.wmnet * 08:05 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2209: Pool back * 08:05 btullis@cumin1003: START - Cookbook sre.dns.netbox * 08:00 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host titan2002.codfw.wmnet * 07:59 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1229 * 07:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1230 * 07:58 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1230 * 07:55 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1230 * 07:55 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1230.eqiad.wmnet 23.53.64.10.in-addr.arpa 3.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:55 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1230.eqiad.wmnet 23.53.64.10.in-addr.arpa 3.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:55 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:55 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1230 - btullis@cumin1003" * 07:55 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1230 - btullis@cumin1003" * 07:51 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host kubestagemaster2005.codfw.wmnet with OS trixie * 07:48 btullis@cumin1003: START - Cookbook sre.dns.netbox * 07:41 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1230 * 07:41 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1229.eqiad.wmnet with OS bookworm * 07:41 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1230.eqiad.wmnet with OS bookworm * 07:39 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1152 from dbctl [[phab:T434480|T434480]]', diff saved to https://phabricator.wikimedia.org/P96090 and previous config saved to /var/cache/conftool/dbconfig/20260814-073941-marostegui.json * 07:29 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on kubestagemaster2005.codfw.wmnet with reason: host reimage * 07:23 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on kubestagemaster2005.codfw.wmnet with reason: host reimage * 07:04 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host kubestagemaster2005.codfw.wmnet with OS trixie * 06:53 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1228.eqiad.wmnet with OS bookworm * 06:44 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1227.eqiad.wmnet with OS bookworm * 06:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1209.eqiad.wmnet with OS bookworm * 06:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1175.eqiad.wmnet with OS bookworm * 06:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1228.eqiad.wmnet with reason: host reimage * 06:27 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1228.eqiad.wmnet with reason: host reimage * 06:25 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1227.eqiad.wmnet with reason: host reimage * 06:21 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1227.eqiad.wmnet with reason: host reimage * 06:18 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1209.eqiad.wmnet with reason: host reimage * 06:14 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1209.eqiad.wmnet with reason: host reimage * 06:14 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1228 * 06:14 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1228 * 06:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1175.eqiad.wmnet with reason: host reimage * 06:12 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1228 * 06:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1228.eqiad.wmnet 20.53.64.10.in-addr.arpa 0.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:12 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1228.eqiad.wmnet 20.53.64.10.in-addr.arpa 0.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1228 - ryankemper@cumin2003" * 06:12 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1228 - ryankemper@cumin2003" * 06:09 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1175.eqiad.wmnet with reason: host reimage * 06:07 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 06:07 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1228 * 06:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1227 * 06:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1227 * 06:06 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1227 * 06:06 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1227.eqiad.wmnet 19.53.64.10.in-addr.arpa 9.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:06 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1227.eqiad.wmnet 19.53.64.10.in-addr.arpa 9.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:06 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:06 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1227 - ryankemper@cumin2003" * 06:06 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1227 - ryankemper@cumin2003" * 06:00 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 06:00 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1227 * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1209 * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1209 * 06:00 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1209 * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1209.eqiad.wmnet 15.53.64.10.in-addr.arpa 5.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:00 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1209.eqiad.wmnet 15.53.64.10.in-addr.arpa 5.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1209 - ryankemper@cumin2003" * 06:00 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1209 - ryankemper@cumin2003" * 05:54 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 05:54 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1209 * 05:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1175 * 05:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1175 * 05:52 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1175 * 05:52 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1175.eqiad.wmnet 17.53.64.10.in-addr.arpa 7.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 05:52 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1175.eqiad.wmnet 17.53.64.10.in-addr.arpa 7.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 05:52 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 05:52 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1175 - ryankemper@cumin2003" * 05:52 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1175 - ryankemper@cumin2003" * 05:49 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1228.eqiad.wmnet with OS bookworm * 05:49 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1227.eqiad.wmnet with OS bookworm * 05:48 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1209.eqiad.wmnet with OS bookworm * 05:47 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 05:47 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1175 * 05:47 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1175.eqiad.wmnet with OS bookworm * 05:09 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host apifeatureusage1001.eqiad.wmnet with OS bookworm * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 03s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:10 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 01:07 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 01:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 01:02 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 00:59 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1226.eqiad.wmnet with OS bookworm * 00:47 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1225.eqiad.wmnet with OS bookworm * 00:41 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1224.eqiad.wmnet with OS bookworm * 00:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1226.eqiad.wmnet with reason: host reimage * 00:31 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1226.eqiad.wmnet with reason: host reimage * 00:28 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1225.eqiad.wmnet with reason: host reimage * 00:25 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1225.eqiad.wmnet with reason: host reimage * 00:19 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1224.eqiad.wmnet with reason: host reimage * 00:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1226 * 00:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1226 * 00:17 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1226 * 00:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1226.eqiad.wmnet 23.36.64.10.in-addr.arpa 3.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:17 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1226.eqiad.wmnet 23.36.64.10.in-addr.arpa 3.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 00:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1226 - ryankemper@cumin2003" * 00:17 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1226 - ryankemper@cumin2003" * 00:16 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1224.eqiad.wmnet with reason: host reimage * 00:12 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 00:12 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1226 * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1225 * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1225 * 00:10 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1225 * 00:10 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1225.eqiad.wmnet 22.36.64.10.in-addr.arpa 2.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:10 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1225.eqiad.wmnet 22.36.64.10.in-addr.arpa 2.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:10 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 00:10 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1225 - ryankemper@cumin2003" * 00:10 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1225 - ryankemper@cumin2003" * 00:03 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 00:02 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1225 * 00:02 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1224 * 00:02 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1224 * 00:00 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1224 * 00:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1224.eqiad.wmnet 21.36.64.10.in-addr.arpa 1.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:00 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1224.eqiad.wmnet 21.36.64.10.in-addr.arpa 1.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 00:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1224 - ryankemper@cumin2003" * 00:00 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1224 - ryankemper@cumin2003" == 2026-08-13 == * 23:54 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1226.eqiad.wmnet with OS bookworm * 23:53 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1225.eqiad.wmnet with OS bookworm * 23:52 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 23:51 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1224 * 23:51 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1224.eqiad.wmnet with OS bookworm * 23:47 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1223.eqiad.wmnet with OS bookworm * 23:29 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1223.eqiad.wmnet with reason: host reimage * 23:24 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1223.eqiad.wmnet with reason: host reimage * 23:19 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325525{{!}}ve.ui.CodeMirror.less: ensure normal font style]] (duration: 11m 40s) * 23:16 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260807000000" --end-timestamp="20260808000000" --sleep="5" --batch-size="50"` for [[phab:T434688|T434688]] * 23:13 musikanimal@deploy1003: musikanimal: Continuing with deployment * 23:11 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1325525{{!}}ve.ui.CodeMirror.less: ensure normal font style]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1223 * 23:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1223 * 23:08 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1325525{{!}}ve.ui.CodeMirror.less: ensure normal font style]] * 23:07 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1223 * 23:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1223.eqiad.wmnet 20.36.64.10.in-addr.arpa 0.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:07 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1223.eqiad.wmnet 20.36.64.10.in-addr.arpa 0.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 23:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1223 - ryankemper@cumin2003" * 23:03 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1223 - ryankemper@cumin2003" * 23:01 sbassett: Deployed security updates for [[phab:T430596|T430596]], [[phab:T120386|T120386]] * 22:55 bking@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host kubestagemaster2005.codfw.wmnet * 22:55 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host kubestagemaster2005.codfw.wmnet with OS bookworm * 22:54 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 22:53 ryankemper@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 22:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on kubestagemaster2005.codfw.wmnet with reason: host reimage * 22:49 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1222.eqiad.wmnet with OS bookworm * 22:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on kubestagemaster2005.codfw.wmnet with reason: host reimage * 22:29 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1222.eqiad.wmnet with reason: host reimage * 22:26 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1222.eqiad.wmnet with reason: host reimage * 22:23 bking@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host aux-k8s-etcd2003.codfw.wmnet * 22:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd2003.codfw.wmnet with OS bookworm * 22:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host kubestagemaster2005.codfw.wmnet with OS bookworm * 22:23 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM kubestagemaster2005.codfw.wmnet - bking@cumin2003" * 22:23 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM kubestagemaster2005.codfw.wmnet - bking@cumin2003" * 22:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) kubestagemaster2005.codfw.wmnet on all recursors * 22:22 bking@cumin2003: START - Cookbook sre.dns.wipe-cache kubestagemaster2005.codfw.wmnet on all recursors * 22:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:22 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM kubestagemaster2005.codfw.wmnet - bking@cumin2003" * 22:22 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM kubestagemaster2005.codfw.wmnet - bking@cumin2003" * 22:17 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 22:13 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1223 * 22:12 bking@cumin2003: START - Cookbook sre.dns.netbox * 22:12 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host kubestagemaster2005.codfw.wmnet * 22:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1222 * 22:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1222 * 22:11 sbassett: Deployed security updates for [[phab:T429244|T429244]], [[phab:T434039|T434039]] * 22:11 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1222 * 22:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1222.eqiad.wmnet 19.36.64.10.in-addr.arpa 9.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:11 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1222.eqiad.wmnet 19.36.64.10.in-addr.arpa 9.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1222 - ryankemper@cumin2003" * 22:07 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1222 - ryankemper@cumin2003" * 22:05 robh@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-wdqs2001.codfw.wmnet with reason: updating firmware * 22:01 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 22:01 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1223.eqiad.wmnet with OS bookworm * 22:01 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1222 * 22:01 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1222.eqiad.wmnet with OS bookworm * 22:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1212.eqiad.wmnet with OS bookworm * 21:54 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324295{{!}}Revert "Lazily reject pre-fix parser-cache entries for noreferrer/noopener links" (T429090)]] (duration: 06m 42s) * 21:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd2003.codfw.wmnet with reason: host reimage * 21:53 bking@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host dse-k8s-etcd2001.codfw.wmnet * 21:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dse-k8s-etcd2001.codfw.wmnet with OS bookworm * 21:50 sbassett@deploy1003: sbassett, kharlan: Continuing with deployment * 21:49 sbassett@deploy1003: sbassett, kharlan: Backport for [[gerrit:1324295{{!}}Revert "Lazily reject pre-fix parser-cache entries for noreferrer/noopener links" (T429090)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:48 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aux-k8s-etcd2003.codfw.wmnet with reason: host reimage * 21:47 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1324295{{!}}Revert "Lazily reject pre-fix parser-cache entries for noreferrer/noopener links" (T429090)]] * 21:42 maryum: Deployed security patch for [[phab:T434549|T434549]] * 21:39 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1212.eqiad.wmnet with reason: host reimage * 21:34 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1212.eqiad.wmnet with reason: host reimage * 21:33 bking@cumin2003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd2003.codfw.wmnet with OS bookworm * 21:32 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM aux-k8s-etcd2003.codfw.wmnet - bking@cumin2003" * 21:32 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM aux-k8s-etcd2003.codfw.wmnet - bking@cumin2003" * 21:32 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) aux-k8s-etcd2003.codfw.wmnet on all recursors * 21:32 bking@cumin2003: START - Cookbook sre.dns.wipe-cache aux-k8s-etcd2003.codfw.wmnet on all recursors * 21:32 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:32 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM aux-k8s-etcd2003.codfw.wmnet - bking@cumin2003" * 21:31 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM aux-k8s-etcd2003.codfw.wmnet - bking@cumin2003" * 21:28 maryum: Deployed security patch for [[phab:T434619|T434619]] * 21:26 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:26 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host aux-k8s-etcd2003.codfw.wmnet * 21:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-etcd2001.codfw.wmnet with reason: host reimage * 21:20 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1212 * 21:20 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1212 * 21:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1212.eqiad.wmnet with OS bookworm * 21:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on dse-k8s-etcd2001.codfw.wmnet with reason: host reimage * 21:04 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325553{{!}}InstrumentConstructiveEdits: exclude mw-reverted as well (T431493)]] (duration: 06m 25s) * 20:59 kemayo@deploy1003: kemayo: Continuing with deployment * 20:59 kemayo@deploy1003: kemayo: Backport for [[gerrit:1325553{{!}}InstrumentConstructiveEdits: exclude mw-reverted as well (T431493)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host dse-k8s-etcd2001.codfw.wmnet with OS bookworm * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 20:57 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 20:57 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1325553{{!}}InstrumentConstructiveEdits: exclude mw-reverted as well (T431493)]] * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-etcd2001.codfw.wmnet on all recursors * 20:57 bking@cumin2003: START - Cookbook sre.dns.wipe-cache dse-k8s-etcd2001.codfw.wmnet on all recursors * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 20:57 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 20:54 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320163{{!}}Add configurable RestTermsOfServiceUrl (T428147)]] (duration: 21m 39s) * 20:53 bking@cumin2003: START - Cookbook sre.dns.netbox * 20:53 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host dse-k8s-etcd2001.codfw.wmnet * 20:50 samtar@deploy1003: samtar, milazg: Continuing with deployment * 20:35 samtar@deploy1003: samtar, milazg: Backport for [[gerrit:1320163{{!}}Add configurable RestTermsOfServiceUrl (T428147)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:33 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1320163{{!}}Add configurable RestTermsOfServiceUrl (T428147)]] * 20:30 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325549{{!}}Deploy PRV to several LC wikis (T423785)]] (duration: 06m 57s) * 20:26 arlolra@deploy1003: arlolra: Continuing with deployment * 20:25 arlolra@deploy1003: arlolra: Backport for [[gerrit:1325549{{!}}Deploy PRV to several LC wikis (T423785)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:24 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 20:23 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1325549{{!}}Deploy PRV to several LC wikis (T423785)]] * 20:21 ariel@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319906{{!}}Remove boilerplate language from wmf-rest and wmf-math API modules (T433736)]] (duration: 13m 54s) * 20:14 ariel@deploy1003: ariel: Continuing with deployment * 20:11 ariel@deploy1003: ariel: Backport for [[gerrit:1319906{{!}}Remove boilerplate language from wmf-rest and wmf-math API modules (T433736)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 ariel@deploy1003: Started scap sync-world: Backport for [[gerrit:1319906{{!}}Remove boilerplate language from wmf-rest and wmf-math API modules (T433736)]] * 19:58 robh@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-wdqs2001.codfw.wmnet with reason: updating firmware * 19:54 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325551{{!}}Render the focused module view as a full-screen page (T433896)]] (duration: 30m 37s) * 19:52 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching aqs[2001,1016]*: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 19:44 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching aqs[2001,1016]*: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 19:42 musikanimal@deploy1003: musikanimal: Continuing with deployment * 19:41 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1325551{{!}}Render the focused module view as a full-screen page (T433896)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:34 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: quash java safepoint logspam - bking@cumin2003 - [[phab:T434685|T434685]] * 19:34 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 19:34 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 19:24 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1325551{{!}}Render the focused module view as a full-screen page (T433896)]] * 19:20 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 19:20 jhancock@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin1003" * 19:18 jhancock@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin1003" * 19:14 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325548{{!}}Enable image lazy loading on desktop in group1 (T148047)]] (duration: 07m 43s) * 19:10 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 19:10 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1325548{{!}}Enable image lazy loading on desktop in group1 (T148047)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:07 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1325548{{!}}Enable image lazy loading on desktop in group1 (T148047)]] * 19:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1166.eqiad.wmnet onto db1280.eqiad.wmnet * 19:03 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 19:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1166: Pool db1166.eqiad.wmnet in after cloning * 18:59 jhancock@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 18:58 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1234.eqiad.wmnet with OS bookworm * 18:52 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 18:51 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:49 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:46 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 18:46 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 18:44 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=dns3004.* * 18:39 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1221.eqiad.wmnet with OS bookworm * 18:36 inflatador: [bking@ganeti2048] ~$ sudo gnt-instance replace-disks -n ganeti2030.codfw.wmnet aux-k8s-worker2002.codfw.wmnet [[phab:T434681|T434681]] * 18:35 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 18:29 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1220.eqiad.wmnet with OS bookworm * 18:29 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1234.eqiad.wmnet with reason: host reimage * 18:27 bking@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host dse-k8s-etcd2001.codfw.wmnet * 18:27 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-etcd2001.codfw.wmnet on all recursors * 18:27 bking@cumin2003: START - Cookbook sre.dns.wipe-cache dse-k8s-etcd2001.codfw.wmnet on all recursors * 18:27 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:27 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 18:27 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 18:25 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1234.eqiad.wmnet with reason: host reimage * 18:25 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 18:19 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1221.eqiad.wmnet with reason: host reimage * 18:17 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1166: Pool db1166.eqiad.wmnet in after cloning * 18:15 brennen@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 18:15 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1221.eqiad.wmnet with reason: host reimage * 18:14 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:12 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:12 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 18:12 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-etcd2001.codfw.wmnet on all recursors * 18:12 bking@cumin2003: START - Cookbook sre.dns.wipe-cache dse-k8s-etcd2001.codfw.wmnet on all recursors * 18:12 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:12 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 18:12 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 18:10 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: quash java safepoint logspam - bking@cumin2003 - [[phab:T434685|T434685]] * 18:10 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1234 * 18:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1234 * 18:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1220.eqiad.wmnet with reason: host reimage * 18:07 brennen: 1.47.0-wmf.15 train status ([[phab:T430834|T430834]]) - no current blockers, rolling to all wikis * 18:07 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1234 * 18:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1234.eqiad.wmnet 10.36.64.10.in-addr.arpa 0.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:07 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:07 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1234.eqiad.wmnet 10.36.64.10.in-addr.arpa 0.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1234 - ryankemper@cumin2003" * 18:07 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1234 - ryankemper@cumin2003" * 18:05 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1220.eqiad.wmnet with reason: host reimage * 18:04 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host dse-k8s-etcd2001.codfw.wmnet * 18:04 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 18:03 dancy@deploy1003: Installation of scap version "4.280.1" completed for 3 hosts * 18:02 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:02 inflatador: bking@dse-k8s-etcd2002 etcdctl member remove $<nowiki>{</nowiki>UUID of dse-k8s-etcd2001<nowiki>}</nowiki> [[phab:T434681|T434681]] [[phab:T434793|T434793]] * 18:01 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 18:01 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1234 * 18:01 dancy@deploy1003: Installing scap version "4.280.1" for 3 host(s) * 18:01 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1221 * 18:01 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1221 * 18:00 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1221 * 18:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1221.eqiad.wmnet 18.36.64.10.in-addr.arpa 8.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:00 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1221.eqiad.wmnet 18.36.64.10.in-addr.arpa 8.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1221 - ryankemper@cumin2003" * 17:59 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:58 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 17:58 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:57 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 17:57 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:56 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1221 - ryankemper@cumin2003" * 17:53 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 17:52 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 17:52 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 17:51 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1221 * 17:51 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1220 * 17:51 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1220 * 17:51 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:51 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:51 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1220 * 17:51 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1220.eqiad.wmnet 11.36.64.10.in-addr.arpa 1.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:51 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1220.eqiad.wmnet 11.36.64.10.in-addr.arpa 1.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:51 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1220 - ryankemper@cumin2003" * 17:50 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1220 - ryankemper@cumin2003" * 17:47 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:47 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 17:46 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1234.eqiad.wmnet with OS bookworm * 17:46 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 17:45 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1221.eqiad.wmnet with OS bookworm * 17:45 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1220 * 17:45 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1220.eqiad.wmnet with OS bookworm * 17:43 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 17:41 bking@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host dse-k8s-etcd2001.codfw.wmnet * 17:41 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host dse-k8s-etcd2001.codfw.wmnet * 17:40 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:40 inflatador: bking@ganeti2048] `sudo gnt-instance remove --force --ignore-failures --shutdown-timeout=0` on non-DRBD VMs [[phab:T434681|T434681]] * 17:40 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:39 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:38 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:36 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:32 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:32 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:28 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 17:26 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:26 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:24 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:23 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:21 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1218.eqiad.wmnet with OS bookworm * 17:20 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:20 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:19 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1179.eqiad.wmnet with OS bookworm * 17:18 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1150.eqiad.wmnet with OS bookworm * 17:18 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns3004.wikimedia.org with OS trixie * 17:15 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:14 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:13 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:12 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:12 swfrench@deploy1003: Finished scap sync-world: Helmfile-only deployment for mediawiki chart bump - [[phab:T427666|T427666]] (duration: 03m 03s) * 17:09 swfrench@deploy1003: Started scap sync-world: Helmfile-only deployment for mediawiki chart bump - [[phab:T427666|T427666]] * 17:02 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:01 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:01 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1218.eqiad.wmnet with reason: host reimage * 17:01 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:01 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:00 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324810{{!}}deployment-info.php: Report dbname and branch for the requested wiki (T434726)]] (duration: 06m 52s) * 16:58 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 16:57 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1179.eqiad.wmnet with reason: host reimage * 16:56 dancy@deploy1003: dancy: Continuing with deployment * 16:56 dancy@deploy1003: dancy: Backport for [[gerrit:1324810{{!}}deployment-info.php: Report dbname and branch for the requested wiki (T434726)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1150.eqiad.wmnet with reason: host reimage * 16:53 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324810{{!}}deployment-info.php: Report dbname and branch for the requested wiki (T434726)]] * 16:51 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1166: Depool db1166.eqiad.wmnet to then clone it to db1280.eqiad.wmnet - cwilliams@cumin1003 * 16:50 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1166: Depool db1166.eqiad.wmnet to then clone it to db1280.eqiad.wmnet - cwilliams@cumin1003 * 16:50 cwilliams@cumin1003: START - Cookbook sre.mysql.clone of db1166.eqiad.wmnet onto db1280.eqiad.wmnet * 16:49 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1218.eqiad.wmnet with reason: host reimage * 16:48 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1179.eqiad.wmnet with reason: host reimage * 16:47 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1150.eqiad.wmnet with reason: host reimage * 16:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.upgrade (exit_code=0) for 1 hosts * 16:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2220: Upgrade of db2220.codfw.wmnet completed * 16:38 swfrench@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 16:38 swfrench-wmf: kubectl delete node kubestagemaster2005.codfw.wmnet - [[phab:T434681|T434681]] * 16:34 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1218 * 16:34 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1218 * 16:34 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1218.eqiad.wmnet with OS bookworm * 16:34 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1179 * 16:34 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1179 * 16:33 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1179.eqiad.wmnet with OS bookworm * 16:32 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1150 * 16:32 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1150 * 16:31 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1150.eqiad.wmnet with OS bookworm * 16:24 dancy@deploy1003: Installation of scap version "4.280.0" completed for 3 hosts * 16:22 dancy@deploy1003: Installing scap version "4.280.0" for 3 host(s) * 16:20 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: quash java safepoint logspam - bking@cumin2003 - [[phab:T434685|T434685]] * 16:14 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns3004.wikimedia.org with reason: host reimage * 16:08 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 16:07 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 16:07 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns3004.wikimedia.org with reason: host reimage * 16:05 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 16:04 swfrench@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 16:00 swfrench@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 15:59 swfrench@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 15:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Upgrade of db2220.codfw.wmnet completed * 15:48 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2220: Upgrading db2220.codfw.wmnet * 15:48 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2220: Upgrading db2220.codfw.wmnet * 15:48 cwilliams@cumin1003: START - Cookbook sre.mysql.upgrade for 1 hosts * 15:46 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns3004.wikimedia.org with OS trixie * 15:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2220 [[phab:T434802|T434802]]', diff saved to https://phabricator.wikimedia.org/P96079 and previous config saved to /var/cache/conftool/dbconfig/20260813-154624-cwilliams.json * 15:45 cdobbins@cumin1003: conftool action : set/pooled=no; selector: name=dns3004.* * 15:44 cjd91: depooling dns3004 to reimage to trixie * 15:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2159 to s7 primary [[phab:T434802|T434802]]', diff saved to https://phabricator.wikimedia.org/P96078 and previous config saved to /var/cache/conftool/dbconfig/20260813-154405-cwilliams.json * 15:43 cezmunsta: Starting s7 codfw failover from db2220 to db2159 - [[phab:T434802|T434802]] * 15:41 cgoubert@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: Dragonfly supernodes reboot (duration: 09m 42s) * 15:41 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dragonfly-supernode2001.codfw.wmnet * 15:39 inflatador: bking@ganeti2048] ~$ sudo gnt-node failover -f --ignore-consistency ganeti2046.codfw.wmnet [[phab:T434681|T434681]] * 15:39 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 15:39 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 15:39 swfrench@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 15:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2159 with weight 0 [[phab:T434802|T434802]]', diff saved to https://phabricator.wikimedia.org/P96077 and previous config saved to /var/cache/conftool/dbconfig/20260813-153806-cwilliams.json * 15:37 swfrench@dns1004: END - running authdns-update * 15:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 30 hosts with reason: Primary switchover s7 [[phab:T434802|T434802]] * 15:37 cgoubert@cumin2003: START - Cookbook sre.hosts.reboot-single for host dragonfly-supernode2001.codfw.wmnet * 15:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dragonfly-supernode1001.eqiad.wmnet * 15:35 swfrench@dns1004: START - running authdns-update * 15:32 cgoubert@cumin2003: START - Cookbook sre.hosts.reboot-single for host dragonfly-supernode1001.eqiad.wmnet * 15:32 cgoubert@deploy1003: Locking from deployment [ALL REPOSITORIES]: Dragonfly supernodes reboot * 15:30 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-master-codfw * 15:30 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2005.codfw.wmnet * 15:30 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2005.codfw.wmnet * 15:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1169.eqiad.wmnet onto db1277.eqiad.wmnet * 15:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1169: Pool db1169.eqiad.wmnet in after cloning * 15:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2005.codfw.wmnet * 15:23 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2005.codfw.wmnet * 15:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2004.codfw.wmnet * 15:23 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2004.codfw.wmnet * 15:18 inflatador: bking@ganeti2048 sudo gnt-node failover -f ganeti2046.codfw.wmnet [[phab:T434681|T434681]] * 15:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2004.codfw.wmnet * 15:17 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2004.codfw.wmnet * 15:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2003.codfw.wmnet * 15:17 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2003.codfw.wmnet * 15:15 cdobbins@dns1004: END - running authdns-update * 15:13 cdobbins@dns1004: START - running authdns-update * 15:10 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1217.eqiad.wmnet with OS bookworm * 15:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2003.codfw.wmnet * 15:10 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2003.codfw.wmnet * 15:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2002.codfw.wmnet * 15:10 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2002.codfw.wmnet * 15:10 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: quash java safepoint logspam - bking@cumin2003 - [[phab:T434685|T434685]] * 15:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1216.eqiad.wmnet with OS bookworm * 15:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2002.codfw.wmnet * 15:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2002.codfw.wmnet * 15:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2001.codfw.wmnet * 15:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2001.codfw.wmnet * 15:01 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1042.eqiad.wmnet * 15:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1042.eqiad.wmnet * 15:00 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325455{{!}}Api: Use correct query when continue prop=categories (T433922)]] (duration: 09m 47s) * 14:59 bking@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host dse-k8s-etcd2004.codfw.wmnet * 14:58 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-etcd2004.codfw.wmnet on all recursors * 14:58 bking@cumin2003: START - Cookbook sre.dns.wipe-cache dse-k8s-etcd2004.codfw.wmnet on all recursors * 14:58 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:58 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM dse-k8s-etcd2004.codfw.wmnet - bking@cumin2003" * 14:58 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM dse-k8s-etcd2004.codfw.wmnet - bking@cumin2003" * 14:55 zabe@deploy1003: zabe: Continuing with deployment * 14:53 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-etcd2004.codfw.wmnet on all recursors * 14:53 bking@cumin2003: START - Cookbook sre.dns.wipe-cache dse-k8s-etcd2004.codfw.wmnet on all recursors * 14:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:53 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2004.codfw.wmnet - bking@cumin2003" * 14:53 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2004.codfw.wmnet - bking@cumin2003" * 14:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2001.codfw.wmnet * 14:52 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2001.codfw.wmnet * 14:52 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-master-codfw * 14:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2001.codfw.wmnet * 14:52 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2001.codfw.wmnet * 14:52 zabe@deploy1003: zabe: Backport for [[gerrit:1325455{{!}}Api: Use correct query when continue prop=categories (T433922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:50 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1325455{{!}}Api: Use correct query when continue prop=categories (T433922)]] * 14:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1217.eqiad.wmnet with reason: host reimage * 14:48 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:48 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host dse-k8s-etcd2004.codfw.wmnet * 14:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-master-eqiad * 14:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1006.eqiad.wmnet * 14:44 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1006.eqiad.wmnet * 14:44 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1216.eqiad.wmnet with reason: host reimage * 14:40 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1217.eqiad.wmnet with reason: host reimage * 14:39 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1169: Pool db1169.eqiad.wmnet in after cloning * 14:39 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1216.eqiad.wmnet with reason: host reimage * 14:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl1005.eqiad.wmnet * 14:32 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl1005.eqiad.wmnet * 14:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1004.eqiad.wmnet * 14:32 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1004.eqiad.wmnet * 14:31 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus2008.codfw.wmnet * 14:31 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor1003.eqiad.wmnet * 14:29 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1042.eqiad.wmnet * 14:27 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor1003.eqiad.wmnet * 14:26 moritzm: installing Django security updates * 14:26 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling reboot on A:wikidough * 14:25 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1217.eqiad.wmnet with OS bookworm * 14:25 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1216.eqiad.wmnet with OS bookworm * 14:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl1004.eqiad.wmnet * 14:24 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl1004.eqiad.wmnet * 14:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1003.eqiad.wmnet * 14:24 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1003.eqiad.wmnet * 14:23 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus2008.codfw.wmnet * 14:23 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus1008.eqiad.wmnet * 14:23 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor-dev2001.codfw.wmnet * 14:22 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus2006.codfw.wmnet * 14:19 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor-dev2001.codfw.wmnet * 14:18 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1042.eqiad.wmnet * 14:17 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.* * 14:16 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl1003.eqiad.wmnet * 14:16 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl1003.eqiad.wmnet * 14:16 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1002.eqiad.wmnet * 14:16 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1002.eqiad.wmnet * 14:15 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus1008.eqiad.wmnet * 14:14 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus1006.eqiad.wmnet * 14:12 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus2006.codfw.wmnet * 14:12 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor2003.codfw.wmnet * 14:11 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus2007.codfw.wmnet * 14:11 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1041.eqiad.wmnet * 14:11 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1041.eqiad.wmnet * 14:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl1002.eqiad.wmnet * 14:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl1002.eqiad.wmnet * 14:09 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-master-eqiad * 14:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on apifeatureusage1001.eqiad.wmnet with reason: host reimage * 14:08 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1215.eqiad.wmnet with OS bookworm * 14:08 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor2003.codfw.wmnet * 14:08 jayme: updated calico to v3.30.7 on wikikube codfw [[phab:T427400|T427400]] * 14:07 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1149.eqiad.wmnet with OS bookworm * 14:06 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'. * 14:06 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1041.eqiad.wmnet * 14:04 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus1006.eqiad.wmnet * 14:03 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus2007.codfw.wmnet * 14:03 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus2005.codfw.wmnet * 14:03 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus1007.eqiad.wmnet * 14:03 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1002.eqiad.wmnet * 14:02 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on apifeatureusage1001.eqiad.wmnet with reason: host reimage * 14:02 cgoubert@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=helm-charts.*,name=eqiad * 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host chartmuseum1001.eqiad.wmnet * 14:00 moritzm: installing libxml2 security updates * 13:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1214.eqiad.wmnet with OS bookworm * 13:59 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns5003.* * 13:59 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1041.eqiad.wmnet * 13:59 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow1002.eqiad.wmnet * 13:58 cmooney@dns3003: END - running authdns-update * 13:57 cgoubert@cumin2003: START - Cookbook sre.hosts.reboot-single for host chartmuseum1001.eqiad.wmnet * 13:57 cgoubert@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=helm-charts.*,name=eqiad * 13:57 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: cloudelastic cluster restart - bking@cumin2003 * 13:57 cgoubert@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=helm-charts.*,name=codfw * 13:56 cmooney@dns3003: START - running authdns-update * 13:56 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'. * 13:56 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns5003.*,service=authdns-update * 13:55 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host chartmuseum2001.codfw.wmnet * 13:55 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus1007.eqiad.wmnet * 13:55 cmooney@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dns5003.wikimedia.org * 13:55 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus1005.eqiad.wmnet * 13:51 cgoubert@cumin2003: START - Cookbook sre.hosts.reboot-single for host chartmuseum2001.codfw.wmnet * 13:51 cgoubert@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=helm-charts.*,name=codfw * 13:51 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus2005.codfw.wmnet * 13:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host apifeatureusage1001.eqiad.wmnet with OS bookworm * 13:50 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus7002.magru.wmnet * 13:50 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts lvs1015.eqiad.wmnet * 13:50 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:50 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1015.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:49 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1015.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:49 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325480{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] (duration: 06m 39s) * 13:47 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1215.eqiad.wmnet with reason: host reimage * 13:46 cmooney@cumin1003: START - Cookbook sre.hosts.reboot-single for host dns5003.wikimedia.org * 13:46 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1003.eqiad.wmnet * 13:45 cmooney@cumin1003: conftool action : set/pooled=no; selector: name=dns5003.* * 13:45 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1040.eqiad.wmnet * 13:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1040.eqiad.wmnet * 13:45 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 13:45 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus1005.eqiad.wmnet * 13:45 stran@deploy1003: stran: Continuing with deployment * 13:44 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus7002.magru.wmnet * 13:44 stran@deploy1003: stran: Backport for [[gerrit:1325480{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:44 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus6002.drmrs.wmnet * 13:43 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260805000000" --end-timestamp="20260806000000" --sleep="5" --batch-size="10"` for [[phab:T434688|T434688]] * 13:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1149.eqiad.wmnet with reason: host reimage * 13:42 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1325480{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] * 13:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow1003.eqiad.wmnet * 13:40 sukhe@cumin1003: START - Cookbook sre.hosts.decommission for hosts lvs1015.eqiad.wmnet * 13:40 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts lvs1014.eqiad.wmnet * 13:40 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:40 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1014.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:40 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1040.eqiad.wmnet * 13:40 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1014.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:39 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1214.eqiad.wmnet with reason: host reimage * 13:38 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2004.codfw.wmnet * 13:38 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus6002.drmrs.wmnet * 13:38 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus5003.eqsin.wmnet * 13:35 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 13:35 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1149.eqiad.wmnet with reason: host reimage * 13:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1215.eqiad.wmnet with reason: host reimage * 13:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow2004.codfw.wmnet * 13:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1214.eqiad.wmnet with reason: host reimage * 13:33 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1040.eqiad.wmnet * 13:31 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus5003.eqsin.wmnet * 13:31 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: cloudelastic cluster restart - bking@cumin2003 * 13:31 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus4003.ulsfo.wmnet * 13:30 sukhe@cumin1003: START - Cookbook sre.hosts.decommission for hosts lvs1014.eqiad.wmnet * 13:30 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts lvs1013.eqiad.wmnet * 13:30 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:30 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1013.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:30 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1013.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:27 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1039.eqiad.wmnet * 13:27 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1039.eqiad.wmnet * 13:26 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1167.eqiad.wmnet onto db1281.eqiad.wmnet * 13:26 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1167: Pool db1167.eqiad.wmnet in after cloning * 13:25 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus4003.ulsfo.wmnet * 13:25 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260802000000" --end-timestamp="20260803000000" --sleep="5" --batch-size="10"` for [[phab:T434688|T434688]] * 13:24 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 13:24 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus3004.esams.wmnet * 13:24 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1274: New host * 13:24 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325476{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] (duration: 07m 19s) * 13:24 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260804000000" --end-timestamp="20260805000000" --sleep="5" --batch-size="10"` for [[phab:T434688|T434688]] * 13:24 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260803000000" --end-timestamp="20260804000000" --sleep="5" --batch-size="10"` for [[phab:T434688|T434688]] * 13:23 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2003.codfw.wmnet * 13:22 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1039.eqiad.wmnet * 13:20 sukhe@cumin1003: START - Cookbook sre.hosts.decommission for hosts lvs1013.eqiad.wmnet * 13:20 stran@deploy1003: stran: Continuing with deployment * 13:19 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow2003.codfw.wmnet * 13:19 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling reboot on A:wikidough * 13:19 stran@deploy1003: stran: Backport for [[gerrit:1325476{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:18 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus3004.esams.wmnet * 13:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1215.eqiad.wmnet with OS bookworm * 13:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1214.eqiad.wmnet with OS bookworm * 13:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1149.eqiad.wmnet with OS bookworm * 13:17 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1325476{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] * 13:11 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324731{{!}}prv: Enable parsoid rendering for 5 wikisource wikis]] (duration: 07m 26s) * 13:11 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1039.eqiad.wmnet * 13:11 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1036.eqiad.wmnet * 13:11 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1036.eqiad.wmnet * 13:07 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow3004.esams.wmnet * 13:07 jgiannelos@deploy1003: jgiannelos: Continuing with deployment * 13:06 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1233.eqiad.wmnet with OS bookworm * 13:06 jgiannelos@deploy1003: jgiannelos: Backport for [[gerrit:1324731{{!}}prv: Enable parsoid rendering for 5 wikisource wikis]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:04 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1324731{{!}}prv: Enable parsoid rendering for 5 wikisource wikis]] * 13:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow3004.esams.wmnet * 13:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1036.eqiad.wmnet * 13:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1210.eqiad.wmnet with OS bookworm * 13:01 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1036.eqiad.wmnet * 13:00 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1052.eqiad.wmnet * 13:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1052.eqiad.wmnet * 12:58 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow4003.ulsfo.wmnet * 12:55 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1211.eqiad.wmnet with OS bookworm * 12:55 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1052.eqiad.wmnet * 12:52 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow4003.ulsfo.wmnet * 12:51 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1052.eqiad.wmnet * 12:49 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1169: Depool db1169.eqiad.wmnet to then clone it to db1277.eqiad.wmnet - cwilliams@cumin1003 * 12:46 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1169: Depool db1169.eqiad.wmnet to then clone it to db1277.eqiad.wmnet - cwilliams@cumin1003 * 12:46 cwilliams@cumin1003: START - Cookbook sre.mysql.clone of db1169.eqiad.wmnet onto db1277.eqiad.wmnet * 12:45 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:45 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns record for deleted IP reservations eqsin lvs vlan ints - cmooney@cumin1003" * 12:44 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns record for deleted IP reservations eqsin lvs vlan ints - cmooney@cumin1003" * 12:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1051.eqiad.wmnet * 12:43 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1051.eqiad.wmnet * 12:42 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1210.eqiad.wmnet with reason: host reimage * 12:41 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1167: Pool db1167.eqiad.wmnet in after cloning * 12:40 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 12:39 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1274: New host * 12:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Added db1274', diff saved to https://phabricator.wikimedia.org/P96062 and previous config saved to /var/cache/conftool/dbconfig/20260813-123907-cwilliams.json * 12:38 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast7002.wikimedia.org * 12:38 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1233.eqiad.wmnet with reason: host reimage * 12:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1051.eqiad.wmnet * 12:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1211.eqiad.wmnet with reason: host reimage * 12:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1233.eqiad.wmnet with reason: host reimage * 12:32 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1051.eqiad.wmnet * 12:32 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast7002.wikimedia.org * 12:32 marostegui: Drop SecurePoll tables from closed wikis [[phab:T423128|T423128]] * 12:32 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow5003.eqsin.wmnet * 12:31 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1050.eqiad.wmnet * 12:31 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1050.eqiad.wmnet * 12:29 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1210.eqiad.wmnet with reason: host reimage * 12:29 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1211.eqiad.wmnet with reason: host reimage * 12:29 moritzm: installing Wireshark security updates * 12:26 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow5003.eqsin.wmnet * 12:26 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1050.eqiad.wmnet * 12:24 cmooney@dns3003: END - running authdns-update * 12:21 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow6001.drmrs.wmnet * 12:21 cmooney@dns3003: START - running authdns-update * 12:20 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1050.eqiad.wmnet * 12:20 cgoubert@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host rdb-lock2003.codfw.wmnet * 12:20 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:20 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update netbox dns entries for expanded public1-603-eqsin subnet - cmooney@cumin1003" * 12:20 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update netbox dns entries for expanded public1-603-eqsin subnet - cmooney@cumin1003" * 12:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 12:20 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 12:19 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1049.eqiad.wmnet * 12:19 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1049.eqiad.wmnet * 12:19 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 12:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host moss-be1003.eqiad.wmnet * 12:17 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow6001.drmrs.wmnet * 12:17 moritzm: installin curl security updates * 12:15 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 12:15 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1233.eqiad.wmnet with OS bookworm * 12:15 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1210.eqiad.wmnet with OS bookworm * 12:15 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1211.eqiad.wmnet with OS bookworm * 12:13 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1049.eqiad.wmnet * 12:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow7002.magru.wmnet * 12:12 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-worker1178.eqiad.wmnet with OS bookworm * 12:10 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host moss-be1003.eqiad.wmnet * 12:10 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 12:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be1006.eqiad.wmnet * 12:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 12:10 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 12:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 12:10 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 12:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow7002.magru.wmnet * 12:08 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1049.eqiad.wmnet * 12:05 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 12:05 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2003.codfw.wmnet * 12:04 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'. * 12:04 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:03 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be1006.eqiad.wmnet * 12:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be1005.eqiad.wmnet * 12:01 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1038.eqiad.wmnet * 12:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1038.eqiad.wmnet * 12:01 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1213.eqiad.wmnet with OS bookworm * 11:56 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be1005.eqiad.wmnet * 11:55 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be1004.eqiad.wmnet * 11:54 cgoubert@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host rdb-lock2003.codfw.wmnet * 11:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 11:53 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 11:53 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:53 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 11:53 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 11:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1038.eqiad.wmnet * 11:49 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be1004.eqiad.wmnet * 11:44 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:44 moritzm: installing Linux 5.10.262 on Bullseye hosts * 11:41 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1038.eqiad.wmnet * 11:40 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 11:40 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 11:40 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 11:40 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:40 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 11:40 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 11:38 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1213.eqiad.wmnet with reason: host reimage * 11:36 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1167: Depool db1167.eqiad.wmnet to then clone it to db1281.eqiad.wmnet - marostegui@cumin1003 * 11:35 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 11:35 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2003.codfw.wmnet * 11:35 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1167: Depool db1167.eqiad.wmnet to then clone it to db1281.eqiad.wmnet - marostegui@cumin1003 * 11:35 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1167.eqiad.wmnet onto db1281.eqiad.wmnet * 11:34 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 22 hosts with reason: Cloning * 11:34 moritzm: remove ganeti3005 from esams03 cluster, hardware issues [[phab:T434646|T434646]] * 11:32 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1213.eqiad.wmnet with reason: host reimage * 11:28 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1034.eqiad.wmnet * 11:28 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1034.eqiad.wmnet * 11:22 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1034.eqiad.wmnet * 11:19 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1034.eqiad.wmnet * 11:17 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1213.eqiad.wmnet with OS bookworm * 11:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1165.eqiad.wmnet onto db1279.eqiad.wmnet * 11:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1165: Pool db1165.eqiad.wmnet in after cloning * 11:07 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply * 10:57 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply * 10:54 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply * 10:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1274.eqiad.wmnet with reason: Enabling notifications and pooling * 10:45 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1033.eqiad.wmnet * 10:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1033.eqiad.wmnet * 10:44 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply * 10:43 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'. * 10:42 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'. * 10:42 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'. * 10:40 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db[1216,1225,1239-1240].eqiad.wmnet with reason: reboot * 10:39 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1033.eqiad.wmnet * 10:38 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1151 from dbctl [[phab:T434538|T434538]]', diff saved to https://phabricator.wikimedia.org/P96055 and previous config saved to /var/cache/conftool/dbconfig/20260813-103828-marostegui.json * 10:35 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1033.eqiad.wmnet * 10:27 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1165: Pool db1165.eqiad.wmnet in after cloning * 10:24 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1151.eqiad.wmnet with OS bookworm * 10:15 moritzm: installing bind9 security updates (client-side tools/libs only) * 10:07 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-debug: apply * 10:06 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-debug: apply * 10:02 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 7 hosts with reason: reboot * 10:01 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-debug: apply * 10:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.upgrade (exit_code=0) for 1 hosts * 10:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2214: Upgrade of db2214.codfw.wmnet completed * 10:01 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-debug: apply * 10:00 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-debug: apply * 10:00 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-debug: apply * 09:59 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.decommission (exit_code=99) * 09:59 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 09:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1151.eqiad.wmnet with reason: host reimage * 09:59 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 8 hosts * 09:59 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 8 hosts * 09:55 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1151.eqiad.wmnet with reason: host reimage * 09:46 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260801000000" --end-timestamp="20260802000000" --sleep="3" --batch-size="5"` for [[phab:T434688|T434688]] * 09:45 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 8 hosts with reason: reboot * 09:45 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 22 hosts with reason: Cloning * 09:43 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1165: Depool db1165.eqiad.wmnet to then clone it to db1279.eqiad.wmnet - marostegui@cumin1003 * 09:42 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1165: Depool db1165.eqiad.wmnet to then clone it to db1279.eqiad.wmnet - marostegui@cumin1003 * 09:42 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1165.eqiad.wmnet onto db1279.eqiad.wmnet * 09:41 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=testwiki --start-timestamp="20260311000000" --end-timestamp="20260805000000" --sleep="5" --batch-size="2"` for [[phab:T434688|T434688]] * 09:40 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for backupmon1001.eqiad.wmnet * 09:40 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for backupmon1001.eqiad.wmnet * 09:40 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1151 * 09:40 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1151 * 09:37 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325408{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]], [[gerrit:1325407{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]] (duration: 06m 57s) * 09:36 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1151 * 09:36 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1151.eqiad.wmnet 13.36.64.10.in-addr.arpa 3.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:36 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1151.eqiad.wmnet 13.36.64.10.in-addr.arpa 3.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:36 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:36 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1151 - btullis@cumin1003" * 09:36 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on backupmon1001.eqiad.wmnet with reason: reboot * 09:36 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1151 - btullis@cumin1003" * 09:35 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host moss-be2003.codfw.wmnet * 09:33 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 7 hosts * 09:33 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 7 hosts * 09:33 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 09:32 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1325408{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]], [[gerrit:1325407{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1161.eqiad.wmnet onto db1275.eqiad.wmnet * 09:31 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 09:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1161: Pool db1161.eqiad.wmnet in after cloning * 09:30 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1325408{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]], [[gerrit:1325407{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]] * 09:29 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 09:27 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 09:27 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host moss-be2003.codfw.wmnet * 09:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be2006.codfw.wmnet * 09:25 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2004.codfw.wmnet * 09:22 hashar@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 09:21 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be2006.codfw.wmnet * 09:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be2005.codfw.wmnet * 09:19 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2004.codfw.wmnet * 09:19 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2003.codfw.wmnet * 09:18 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 7 hosts with reason: reboot * 09:18 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 7 hosts * 09:18 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 7 hosts * 09:16 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2214: Upgrade of db2214.codfw.wmnet completed * 09:14 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be2005.codfw.wmnet * 09:13 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be2004.codfw.wmnet * 09:12 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2003.codfw.wmnet * 09:12 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2002.codfw.wmnet * 09:09 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2214: Upgrading db2214.codfw.wmnet * 09:09 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2214: Upgrading db2214.codfw.wmnet * 09:09 cwilliams@cumin1003: START - Cookbook sre.mysql.upgrade for 1 hosts * 09:07 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be2004.codfw.wmnet * 09:06 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2002.codfw.wmnet * 09:05 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1004.eqiad.wmnet * 09:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-cluster (exit_code=0) * 09:03 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 7 hosts with reason: reboot * 09:01 btullis@cumin1003: START - Cookbook sre.dns.netbox * 09:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2214 [[phab:T434754|T434754]]', diff saved to https://phabricator.wikimedia.org/P96045 and previous config saved to /var/cache/conftool/dbconfig/20260813-090001-cwilliams.json * 08:59 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1004.eqiad.wmnet * 08:59 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1003.eqiad.wmnet * 08:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2229 to s6 primary [[phab:T434754|T434754]]', diff saved to https://phabricator.wikimedia.org/P96044 and previous config saved to /var/cache/conftool/dbconfig/20260813-085752-cwilliams.json * 08:57 cezmunsta: Starting s6 codfw failover from db2214 to db2229 - [[phab:T434754|T434754]] * 08:54 hashar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325416{{!}}Revert "REST: Enable `GET /lexemes/<nowiki>{</nowiki>lexeme_id<nowiki>}</nowiki>` by default" (T434712)]] (duration: 07m 22s) * 08:53 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1003.eqiad.wmnet * 08:53 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1002.eqiad.wmnet * 08:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2229 with weight 0 [[phab:T434754|T434754]]', diff saved to https://phabricator.wikimedia.org/P96043 and previous config saved to /var/cache/conftool/dbconfig/20260813-085151-cwilliams.json * 08:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 23 hosts with reason: Primary switchover s6 [[phab:T434754|T434754]] * 08:51 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts2002.codfw.wmnet * 08:50 hashar@deploy1003: hashar: Continuing with deployment * 08:49 hashar@deploy1003: hashar: Backport for [[gerrit:1325416{{!}}Revert "REST: Enable `GET /lexemes/<nowiki>{</nowiki>lexeme_id<nowiki>}</nowiki>` by default" (T434712)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:47 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet * 08:47 hashar@deploy1003: Started scap sync-world: Backport for [[gerrit:1325416{{!}}Revert "REST: Enable `GET /lexemes/<nowiki>{</nowiki>lexeme_id<nowiki>}</nowiki>` by default" (T434712)]] * 08:47 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1002.eqiad.wmnet * 08:46 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1161: Pool db1161.eqiad.wmnet in after cloning * 08:46 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host stewards2001.codfw.wmnet * 08:45 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit1003.wikimedia.org * 08:45 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host stewards1001.eqiad.wmnet * 08:45 Emperor: roll-restart apus frontends in codfw * 08:45 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-cluster * 08:44 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts2002.codfw.wmnet * 08:44 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host doc2003.codfw.wmnet * 08:43 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet * 08:43 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host phab2003.codfw.wmnet * 08:42 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host stewards2001.codfw.wmnet * 08:42 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host etherpad2002.codfw.wmnet * 08:41 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host stewards1001.eqiad.wmnet * 08:41 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1273: Pool in s7 * 08:41 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1003.wikimedia.org * 08:40 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host doc2003.codfw.wmnet * 08:40 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit1003.wikimedia.org * 08:39 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host etherpad1004.eqiad.wmnet * 08:39 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host doc1004.eqiad.wmnet * 08:39 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit2002.wikimedia.org * 08:38 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host etherpad2002.codfw.wmnet * 08:37 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host phab2003.codfw.wmnet * 08:36 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lists1004.wikimedia.org * 08:35 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host etherpad1004.eqiad.wmnet * 08:35 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host doc1004.eqiad.wmnet * 08:34 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-cluster (exit_code=0) * 08:34 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1003.wikimedia.org * 08:34 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host planet2003.codfw.wmnet * 08:34 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2003.wikimedia.org * 08:33 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host planet1003.eqiad.wmnet * 08:33 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit2002.wikimedia.org * 08:32 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 22 hosts with reason: Cloning * 08:32 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2002.wikimedia.org * 08:31 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aphlict1002.eqiad.wmnet * 08:30 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host planet2003.codfw.wmnet * 08:29 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host planet1003.eqiad.wmnet * 08:29 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host lists1004.wikimedia.org * 08:28 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lists2001.wikimedia.org * 08:28 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2003.wikimedia.org * 08:27 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host aphlict1002.eqiad.wmnet * 08:27 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aphlict2001.codfw.wmnet * 08:26 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2002.wikimedia.org * 08:23 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host aphlict2001.codfw.wmnet * 08:23 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1151 * 08:22 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1151.eqiad.wmnet with OS bookworm * 08:22 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host lists2001.wikimedia.org * 08:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1161: Depool db1161.eqiad.wmnet to then clone it to db1275.eqiad.wmnet - marostegui@cumin1003 * 08:20 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1161: Depool db1161.eqiad.wmnet to then clone it to db1275.eqiad.wmnet - marostegui@cumin1003 * 08:20 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1161.eqiad.wmnet onto db1275.eqiad.wmnet * 08:15 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-cluster * 08:15 Emperor: roll-restart apus frontends in eqiad * 07:56 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1273: Pool in s7 * 07:56 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1273 to dbctl [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P96035 and previous config saved to /var/cache/conftool/dbconfig/20260813-075611-marostegui.json * 07:38 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db1273.eqiad.wmnet with reason: Reboot * 07:31 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.sanitize-wiki (exit_code=97) Managing sanitization for wikis testwiki in section s3 * 07:24 marostegui@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis testwiki in section s3 * 07:19 jayme@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on kubestagemaster2005.codfw.wmnet with reason: downtime because of hardware failure and no DRBD * 05:42 arnaudb@dns1006: END - running authdns-update * 05:40 arnaudb@dns1006: START - running authdns-update * 05:27 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 04:06 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 02:29 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1151 * 02:29 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1151 * 02:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1219.eqiad.wmnet with OS bookworm * 02:18 ryankemper: [[phab:T434494|T434494]] `ryankemper@deploy1003:~$ echo 'https://stats.wikimedia.org/' {{!}} mwscript-k8s --attach -- purgeList.php` (default page got cached during yesterday's `an-web1001` reimage) * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 46s) * 02:03 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1219.eqiad.wmnet with reason: host reimage * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 02:00 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1219.eqiad.wmnet with reason: host reimage * 01:46 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1219.eqiad.wmnet with OS bookworm * 01:03 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1142.eqiad.wmnet with OS bookworm * 00:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1142.eqiad.wmnet with reason: host reimage * 00:34 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1142.eqiad.wmnet with reason: host reimage * 00:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1142.eqiad.wmnet with OS bookworm * 00:16 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1178 * 00:16 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1178 == 2026-08-12 == * 23:18 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324409{{!}}Remove $wmg = $wg hacks in Collection (T119117)]] (duration: 06m 43s) * 23:14 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 23:13 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1324409{{!}}Remove $wmg = $wg hacks in Collection (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:11 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1324409{{!}}Remove $wmg = $wg hacks in Collection (T119117)]] * 22:59 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 22:49 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260801000000" --end-timestamp="20260802000000" --sleep=2 --batch-size=10` * 22:45 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=testwiki --start-timestamp="20200801010101" --end-timestamp="20260816010101" --sleep=15 --batch-size=5` * 22:40 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=testwiki --start-timestamp="20200101010101" --end-timestamp="20260816010101" --sleep=60` * 22:35 Dreamy_Jazz: Running `mwscript WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260101000000" --end-timestamp="20260102000000" --sleep=10` * 22:21 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324817{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]], [[gerrit:1324818{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]] (duration: 45m 29s) * 22:17 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 21:59 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 21:40 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1324817{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]], [[gerrit:1324818{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:39 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 21:36 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1324817{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]], [[gerrit:1324818{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]] * 21:32 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:31 vriley@cumin1003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:30 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:30 vriley@cumin1003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:17 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:14 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:14 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:11 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:10 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-codfw: Set storage compatability to NONE — [[phab:T433028|T433028]] - eevans@cumin1003 * 21:10 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1006 * 21:09 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1006 * 21:05 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324277{{!}}Improve Math preference labels for SVG/MathJax/MathML (T433891)]] (duration: 31m 42s) * 20:58 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1178.eqiad.wmnet with OS bookworm * 20:54 krinkle@deploy1003: krinkle: Continuing with deployment * 20:51 krinkle@deploy1003: krinkle: Backport for [[gerrit:1324277{{!}}Improve Math preference labels for SVG/MathJax/MathML (T433891)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:39 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 20:38 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:38 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp5022.eqsin.wmnet with OS trixie * 20:38 cdobbins@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - cdobbins@cumin1003" * 20:37 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:36 cdobbins@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - cdobbins@cumin1003" * 20:34 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1178.eqiad.wmnet with reason: host reimage * 20:34 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1324277{{!}}Improve Math preference labels for SVG/MathJax/MathML (T433891)]] * 20:33 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 20:28 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1178.eqiad.wmnet with reason: host reimage * 20:13 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1178.eqiad.wmnet with OS bookworm * 20:11 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-worker1178.eqiad.wmnet with OS bookworm * 20:11 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1178.eqiad.wmnet with OS bookworm * 20:09 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-codfw: Set storage compatability to NONE — [[phab:T433028|T433028]] - eevans@cumin1003 * 20:09 Dreamy_Jazz: Evening UTC backport window done * 20:08 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324794{{!}}WikimediaAntiAbuse: Enable logging channel (T431292)]] (duration: 06m 48s) * 20:08 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage * 20:05 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage * 20:04 dreamyjazz@deploy1003: kharlan, dreamyjazz: Continuing with deployment * 20:04 dreamyjazz@deploy1003: kharlan, dreamyjazz: Backport for [[gerrit:1324794{{!}}WikimediaAntiAbuse: Enable logging channel (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:01 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1324794{{!}}WikimediaAntiAbuse: Enable logging channel (T431292)]] * 19:55 brennen@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 19:47 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 19:35 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 19:35 cdobbins@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cp5022.eqsin.wmnet with OS trixie * 19:32 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:30 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:29 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:26 vriley@cumin1003: START - Cookbook sre.dns.netbox * 19:23 brennen@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 19:19 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-eqiad: Set storage compatability to NONE — [[phab:T433028|T433028]] - eevans@cumin1003 * 19:10 brennen@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 19:09 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 19:09 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 19:00 Amir1: data migrated on wikishared ([[phab:T426102|T426102]]) * 18:57 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321224{{!}}Rename ce_worklist_articles table to ce_invitation_list_articles (T426102)]] (duration: 06m 50s) * 18:53 ladsgroup@deploy1003: ladsgroup, daimona: Continuing with deployment * 18:53 ladsgroup@deploy1003: ladsgroup, daimona: Backport for [[gerrit:1321224{{!}}Rename ce_worklist_articles table to ce_invitation_list_articles (T426102)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:51 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1321224{{!}}Rename ce_worklist_articles table to ce_invitation_list_articles (T426102)]] * 18:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2192: Security update * 18:31 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 18:25 Amir1: ce_invitation_list_articles created as empty on wikishared ([[phab:T426102|T426102]]) * 18:21 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 18:21 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 18:18 Amir1: migrated testwiki entries from ce_worklist_articles to ce_invitation_list_articles ([[phab:T426102|T426102]]) * 18:18 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-eqiad: Set storage compatability to NONE — [[phab:T433028|T433028]] - eevans@cumin1003 * 18:11 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 18:08 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling reboot on A:durum-eqsin and A:durum * 18:07 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Set storage compatability to UPGRADING — [[phab:T433028|T433028]] - eevans@cumin1003 * 18:05 jhancock@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022'] * 17:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2192: Security update * 17:55 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:55 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum-eqsin and A:durum * 17:53 jhancock@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['cp5022'] * 17:47 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:47 jhancock@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['cp5022'] * 17:42 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:41 jhancock@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['cp5022'] * 17:36 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:35 jhancock@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022'] * 17:31 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-magru and not (P<nowiki>{</nowiki>cp7001*<nowiki>}</nowiki> or P<nowiki>{</nowiki>cp7009*<nowiki>}</nowiki>) and A:cp - 9.2.15 upgrade ([[phab:T434620|T434620]]) * 17:28 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:28 jhancock@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['cp5022'] * 17:22 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2192.codfw.wmnet with reason: Maintenance * 17:11 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host apifeatureusage2001.codfw.wmnet with OS bookworm * 17:04 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: apply * 17:03 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-main: apply * 17:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2192 [[phab:T434635|T434635]]', diff saved to https://phabricator.wikimedia.org/P96030 and previous config saved to /var/cache/conftool/dbconfig/20260812-170338-cwilliams.json * 17:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2213 to s5 primary [[phab:T434635|T434635]]', diff saved to https://phabricator.wikimedia.org/P96029 and previous config saved to /var/cache/conftool/dbconfig/20260812-170152-cwilliams.json * 17:01 cezmunsta: Starting s5 codfw failover from db2192 to db2213 - [[phab:T434635|T434635]] * 16:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2213 with weight 0 [[phab:T434635|T434635]]', diff saved to https://phabricator.wikimedia.org/P96028 and previous config saved to /var/cache/conftool/dbconfig/20260812-165544-cwilliams.json * 16:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 27 hosts with reason: Primary switchover s5 [[phab:T434635|T434635]] * 16:53 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: apply * 16:52 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-main: apply * 16:44 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-main: apply * 16:44 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-main: apply * 16:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1272: New host * 16:40 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324765{{!}}WikimediaAntiAbuse: Enable PersonalInfoFlagNotifications (T431292)]] (duration: 07m 02s) * 16:40 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-logging-external: apply * 16:39 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-logging-external: apply * 16:38 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-logging-external: apply * 16:37 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-logging-external: apply * 16:36 kharlan@deploy1003: kharlan: Continuing with deployment * 16:35 kharlan@deploy1003: kharlan: Backport for [[gerrit:1324765{{!}}WikimediaAntiAbuse: Enable PersonalInfoFlagNotifications (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:33 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1324765{{!}}WikimediaAntiAbuse: Enable PersonalInfoFlagNotifications (T431292)]] * 16:27 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-logging-external: apply * 16:27 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-logging-external: apply * 16:18 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 16:18 jhancock@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022'] * 16:17 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 16:16 jhancock@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['cp5022'] * 16:15 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 16:12 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Set storage compatability to UPGRADING — [[phab:T433028|T433028]] - eevans@cumin1003 * 16:10 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: apply * 16:10 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: apply * 16:08 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: apply * 16:08 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: apply * 16:08 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics: apply * 16:07 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics: apply * 16:02 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-magru and not (P<nowiki>{</nowiki>cp7001*<nowiki>}</nowiki> or P<nowiki>{</nowiki>cp7009*<nowiki>}</nowiki>) and A:cp - 9.2.15 upgrade ([[phab:T434620|T434620]]) * 15:57 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1272: New host * 15:52 jmm@cumin2003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti2046.codfw.wmnet * 15:52 jmm@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host ganeti2046.codfw.wmnet * 15:42 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:37 mforns@deploy1003: Finished deploy [analytics/refinery@49c336c] (thin): Regular analytics weekly train THIN [analytics/refinery@49c336cd] (duration: 01m 59s) * 15:37 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324733{{!}}WikimediaAntiAbuse: Enable personal info tag display on enwiki (T431292)]] (duration: 08m 12s) * 15:35 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:35 mforns@deploy1003: Started deploy [analytics/refinery@49c336c] (thin): Regular analytics weekly train THIN [analytics/refinery@49c336cd] * 15:34 mforns@deploy1003: Finished deploy [analytics/refinery@49c336c]: Regular analytics weekly train [analytics/refinery@49c336cd] (duration: 04m 20s) * 15:33 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[2024,1031]*.wmnet: Set storage compatability to UPGRADING — [[phab:T433028|T433028]] - eevans@cumin1003 * 15:33 dreamyjazz@deploy1003: kharlan, dreamyjazz: Continuing with deployment * 15:31 dreamyjazz@deploy1003: kharlan, dreamyjazz: Backport for [[gerrit:1324733{{!}}WikimediaAntiAbuse: Enable personal info tag display on enwiki (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:30 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitize-wiki (exit_code=99) Checking sanitization for wikis testwiki in section s3 * 15:30 mforns@deploy1003: Started deploy [analytics/refinery@49c336c]: Regular analytics weekly train [analytics/refinery@49c336cd] * 15:30 mforns@deploy1003: Finished deploy [analytics/refinery@49c336c] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@49c336cd] (duration: 00m 32s) * 15:29 mforns@deploy1003: Started deploy [analytics/refinery@49c336c] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@49c336cd] * 15:29 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1324733{{!}}WikimediaAntiAbuse: Enable personal info tag display on enwiki (T431292)]] * 15:27 brennen@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324700{{!}}EventDetailsParticipantsModule: populate cache with non-local users (T434597)]] (duration: 06m 38s) * 15:23 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[2024,1031]*.wmnet: Set storage compatability to UPGRADING — [[phab:T433028|T433028]] - eevans@cumin1003 * 15:23 brennen@deploy1003: brennen, daimona: Continuing with deployment * 15:23 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:22 brennen@deploy1003: brennen, daimona: Backport for [[gerrit:1324700{{!}}EventDetailsParticipantsModule: populate cache with non-local users (T434597)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:22 cgoubert@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host rdb-lock2003.codfw.wmnet * 15:21 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 15:21 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 15:21 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:21 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:21 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:20 brennen@deploy1003: Started scap sync-world: Backport for [[gerrit:1324700{{!}}EventDetailsParticipantsModule: populate cache with non-local users (T434597)]] * 15:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1178.eqiad.wmnet with OS bookworm * 15:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock1003.eqiad.wmnet * 15:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock1003.eqiad.wmnet with OS trixie * 15:16 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 15:16 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 15:16 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 15:16 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:16 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:16 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:12 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 15:12 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2003.codfw.wmnet * 15:11 cgoubert@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host rdb-lock2003.codfw.wmnet * 15:11 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 15:11 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 15:11 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:11 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:11 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:07 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324338{{!}}InitialiseSettings: Enable 2FA warnings on more private wikis (T428103)]] (duration: 07m 02s) * 15:04 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 15:04 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock1003.eqiad.wmnet with reason: host reimage * 15:03 reedy@deploy1003: reedy: Continuing with deployment * 15:02 reedy@deploy1003: reedy: Backport for [[gerrit:1324338{{!}}InitialiseSettings: Enable 2FA warnings on more private wikis (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:02 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:00 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324338{{!}}InitialiseSettings: Enable 2FA warnings on more private wikis (T428103)]] * 14:57 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 14:57 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2003.codfw.wmnet * 14:57 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock1003.eqiad.wmnet with reason: host reimage * 14:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock2002.codfw.wmnet * 14:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock2002.codfw.wmnet with OS trixie * 14:56 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 14:56 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 14:56 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 14:55 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 14:55 moritzm: powercycle ganeti2046 * 14:47 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock1003.eqiad.wmnet with OS trixie * 14:46 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1003.eqiad.wmnet - cgoubert@cumin2003" * 14:46 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1003.eqiad.wmnet - cgoubert@cumin2003" * 14:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock1003.eqiad.wmnet on all recursors * 14:45 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock1003.eqiad.wmnet on all recursors * 14:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1003.eqiad.wmnet - cgoubert@cumin2003" * 14:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1272.eqiad.wmnet with reason: Enabling notifications * 14:44 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1003.eqiad.wmnet - cgoubert@cumin2003" * 14:44 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324719{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324720{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324722{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0 (T434187)]] (duration: 11m 02s) * 14:39 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 14:39 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock1003.eqiad.wmnet * 14:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock2002.codfw.wmnet with reason: host reimage * 14:37 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 14:37 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock1002.eqiad.wmnet * 14:37 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock1002.eqiad.wmnet with OS trixie * 14:37 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1324719{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324720{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324722{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0 (T434187)]] synced to the testservers (see https://wikitech. * 14:33 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock2002.codfw.wmnet with reason: host reimage * 14:33 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1324719{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324720{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324722{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0 (T434187)]] * 14:32 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1018.eqiad.wmnet with OS bookworm * 14:32 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2046.codfw.wmnet * 14:31 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1020.eqiad.wmnet with OS bookworm * 14:27 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2046.codfw.wmnet * 14:25 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2045.codfw.wmnet * 14:25 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2045.codfw.wmnet * 14:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock1002.eqiad.wmnet with reason: host reimage * 14:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1019.eqiad.wmnet with OS bookworm * 14:22 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Checking sanitization for wikis testwiki in section s3 * 14:20 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2045.codfw.wmnet * 14:18 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock1002.eqiad.wmnet with reason: host reimage * 14:17 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:16 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324707{{!}}Backport all changes from wmf/1.47.0-wmf.15]] (duration: 40m 51s) * 14:16 moritzm: installing Linux 6.1.180 on Bookworm hosts * 14:15 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock2002.codfw.wmnet with OS trixie * 14:14 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2002.codfw.wmnet - cgoubert@cumin2003" * 14:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2002.codfw.wmnet - cgoubert@cumin2003" * 14:14 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:14 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2045.codfw.wmnet * 14:14 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2002.codfw.wmnet on all recursors * 14:14 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2002.codfw.wmnet on all recursors * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2002.codfw.wmnet - cgoubert@cumin2003" * 14:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2002.codfw.wmnet - cgoubert@cumin2003" * 14:12 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2030.codfw.wmnet * 14:12 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2030.codfw.wmnet * 14:11 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on apifeatureusage2001.codfw.wmnet with reason: host reimage * 14:09 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:08 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:07 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 14:06 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2030.codfw.wmnet * 14:06 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:06 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock1002.eqiad.wmnet with OS trixie * 14:05 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 14:05 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2002.codfw.wmnet * 14:05 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1002.eqiad.wmnet - cgoubert@cumin2003" * 14:05 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1002.eqiad.wmnet - cgoubert@cumin2003" * 14:05 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock1002.eqiad.wmnet on all recursors * 14:05 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock1002.eqiad.wmnet on all recursors * 14:05 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:05 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1002.eqiad.wmnet - cgoubert@cumin2003" * 14:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:04 kharlan@deploy1003: kharlan: Continuing with deployment * 14:04 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock2001.codfw.wmnet * 14:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock2001.codfw.wmnet with OS trixie * 14:02 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1002.eqiad.wmnet - cgoubert@cumin2003" * 14:02 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on apifeatureusage2001.codfw.wmnet with reason: host reimage * 14:01 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2030.codfw.wmnet * 13:59 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2029.codfw.wmnet * 13:58 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2029.codfw.wmnet * 13:58 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 13:58 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock1002.eqiad.wmnet * 13:56 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1019.eqiad.wmnet with reason: host reimage * 13:54 btullis@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on archiva1002.wikimedia.org with reason: Upgrading in-place * 13:53 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock1001.eqiad.wmnet * 13:53 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock1001.eqiad.wmnet with OS trixie * 13:53 kharlan@deploy1003: kharlan: Backport for [[gerrit:1324707{{!}}Backport all changes from wmf/1.47.0-wmf.15]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:52 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2029.codfw.wmnet * 13:52 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1020.eqiad.wmnet with reason: host reimage * 13:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1019.eqiad.wmnet with reason: host reimage * 13:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1020.eqiad.wmnet with reason: host reimage * 13:49 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock2001.codfw.wmnet with reason: host reimage * 13:48 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2029.codfw.wmnet * 13:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Configuring db1272 for s3 pooling', diff saved to https://phabricator.wikimedia.org/P96021 and previous config saved to /var/cache/conftool/dbconfig/20260812-134732-cwilliams.json * 13:44 bking@cumin2003: START - Cookbook sre.hosts.reimage for host apifeatureusage2001.codfw.wmnet with OS bookworm * 13:43 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock2001.codfw.wmnet with reason: host reimage * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2028.codfw.wmnet * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2028.codfw.wmnet * 13:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1018.eqiad.wmnet with reason: host reimage * 13:38 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock1001.eqiad.wmnet with reason: host reimage * 13:36 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1018.eqiad.wmnet with reason: host reimage * 13:35 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1324707{{!}}Backport all changes from wmf/1.47.0-wmf.15]] * 13:35 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2028.codfw.wmnet * 13:32 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1275938{{!}}Enable campaignEvents on bdwikimedia (T424016)]] (duration: 07m 35s) * 13:32 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1020.eqiad.wmnet with OS bookworm * 13:32 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1019.eqiad.wmnet with OS bookworm * 13:32 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock1001.eqiad.wmnet with reason: host reimage * 13:28 kharlan@deploy1003: kharlan, yahya: Continuing with deployment * 13:27 kharlan@deploy1003: kharlan, yahya: Backport for [[gerrit:1275938{{!}}Enable campaignEvents on bdwikimedia (T424016)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:27 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2028.codfw.wmnet * 13:26 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 13:26 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 13:25 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1275938{{!}}Enable campaignEvents on bdwikimedia (T424016)]] * 13:25 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitize-wiki (exit_code=99) Managing sanitization for wikis testwiki in section s3 * 13:24 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock2001.codfw.wmnet with OS trixie * 13:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2001.codfw.wmnet - cgoubert@cumin2003" * 13:24 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2001.codfw.wmnet - cgoubert@cumin2003" * 13:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2001.codfw.wmnet on all recursors * 13:23 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2001.codfw.wmnet on all recursors * 13:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2001.codfw.wmnet - cgoubert@cumin2003" * 13:23 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2001.codfw.wmnet - cgoubert@cumin2003" * 13:23 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324713{{!}}thwiki: reinstate temporary wiki25 logos (T431094)]] (duration: 07m 13s) * 13:20 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:20 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1018.eqiad.wmnet with OS bookworm * 13:20 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2027.codfw.wmnet * 13:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2027.codfw.wmnet * 13:19 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics-external: apply * 13:18 kharlan@deploy1003: anzx, kharlan: Continuing with deployment * 13:18 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:18 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock1001.eqiad.wmnet with OS trixie * 13:17 kharlan@deploy1003: anzx, kharlan: Backport for [[gerrit:1324713{{!}}thwiki: reinstate temporary wiki25 logos (T431094)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1001.eqiad.wmnet - cgoubert@cumin2003" * 13:17 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 13:17 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1001.eqiad.wmnet - cgoubert@cumin2003" * 13:17 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2001.codfw.wmnet * 13:17 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics-external: apply * 13:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock1001.eqiad.wmnet on all recursors * 13:17 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock1001.eqiad.wmnet on all recursors * 13:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1001.eqiad.wmnet - cgoubert@cumin2003" * 13:17 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1001.eqiad.wmnet - cgoubert@cumin2003" * 13:15 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1324713{{!}}thwiki: reinstate temporary wiki25 logos (T431094)]] * 13:15 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:15 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics-external: apply * 13:15 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:14 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics-external: apply * 13:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2027.codfw.wmnet * 13:12 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 13:12 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock1001.eqiad.wmnet * 13:12 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2027.codfw.wmnet * 13:08 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1158.eqiad.wmnet onto db1273.eqiad.wmnet * 13:07 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1158: Pool db1158.eqiad.wmnet in after cloning * 13:02 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti6002.drmrs.wmnet * 13:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti6002.drmrs.wmnet * 12:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1015.eqiad.wmnet with OS bookworm * 12:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti6002.drmrs.wmnet * 12:44 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1017.eqiad.wmnet with OS bookworm * 12:36 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 12:35 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 12:34 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 12:33 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 12:31 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti6002.drmrs.wmnet * 12:24 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1017.eqiad.wmnet with reason: host reimage * 12:22 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1158: Pool db1158.eqiad.wmnet in after cloning * 12:18 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1017.eqiad.wmnet with reason: host reimage * 12:11 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1015.eqiad.wmnet with reason: host reimage * 12:07 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1015.eqiad.wmnet with reason: host reimage * 12:04 moritzm: failover ganeti master in drmrs02 to ganeti6004 * 12:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1009.eqiad.wmnet with OS bookworm * 12:01 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1017.eqiad.wmnet with OS bookworm * 12:00 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti6004.drmrs.wmnet * 12:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti6004.drmrs.wmnet * 11:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti6004.drmrs.wmnet * 11:53 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1015.eqiad.wmnet with OS bookworm * 11:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1159.eqiad.wmnet onto db1274.eqiad.wmnet * 11:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Pool db1159.eqiad.wmnet in after cloning * 11:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1016.eqiad.wmnet with OS bookworm * 11:46 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti6004.drmrs.wmnet * 11:45 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti6001.drmrs.wmnet * 11:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti6001.drmrs.wmnet * 11:43 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324306{{!}}WikimediaAntiAbuse: Enable personal info for enwiki with no display (T431292)]] (duration: 10m 26s) * 11:42 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1015.eqiad.wmnet with OS bookworm * 11:39 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 11:38 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti6001.drmrs.wmnet * 11:34 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1324306{{!}}WikimediaAntiAbuse: Enable personal info for enwiki with no display (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:33 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2212: Security update * 11:33 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti6001.drmrs.wmnet * 11:32 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1324306{{!}}WikimediaAntiAbuse: Enable personal info for enwiki with no display (T431292)]] * 11:22 moritzm: failover ganeti master in drmrs01 to ganeti6003 * 11:20 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:20 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:18 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 22 hosts with reason: Cloning * 11:17 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti6003.drmrs.wmnet * 11:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti6003.drmrs.wmnet * 11:17 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1009.eqiad.wmnet with reason: host reimage * 11:17 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:16 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:14 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1016.eqiad.wmnet with reason: host reimage * 11:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti6003.drmrs.wmnet * 11:11 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1009.eqiad.wmnet with reason: host reimage * 11:10 jmm@cumin2003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti3005.esams.wmnet * 11:10 jmm@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host ganeti3005.esams.wmnet * 11:09 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 11:08 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1158: Depool db1158.eqiad.wmnet to then clone it to db1273.eqiad.wmnet - marostegui@cumin1003 * 11:07 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1016.eqiad.wmnet with reason: host reimage * 11:07 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1158: Depool db1158.eqiad.wmnet to then clone it to db1273.eqiad.wmnet - marostegui@cumin1003 * 11:07 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1158.eqiad.wmnet onto db1273.eqiad.wmnet * 11:06 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:05 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:05 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1159: Pool db1159.eqiad.wmnet in after cloning * 11:04 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 20 hosts with reason: Cloning * 11:02 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti6003.drmrs.wmnet * 11:00 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis testwiki in section s3 * 10:54 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1009.eqiad.wmnet with OS bookworm * 10:51 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1015.eqiad.wmnet with OS bookworm * 10:50 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1016.eqiad.wmnet with OS bookworm * 10:48 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2212: Security update * 10:45 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitize-wiki (exit_code=99) Managing sanitization for wikis testwiki in section s3 * 10:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1013.eqiad.wmnet with OS bookworm * 10:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1014.eqiad.wmnet with OS bookworm * 10:38 cwilliams@cumin1003: START - Cookbook sre.mysql.clone of db1159.eqiad.wmnet onto db1274.eqiad.wmnet * 10:33 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db1274.eqiad.wmnet * 10:33 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db1274.eqiad.wmnet * 10:31 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 396993 * 10:29 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 396993 * 10:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1159: Clone source for db1274 * 10:24 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1159: Clone source for db1274 * 10:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1013.eqiad.wmnet with reason: host reimage * 10:18 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1013.eqiad.wmnet with reason: host reimage * 10:13 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2212.codfw.wmnet with reason: Maintenance * 10:13 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 10:12 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 10:12 blake@deploy1003: Stopping before sync operations * 10:11 blake@deploy1003: Started scap sync-world: Non-deployment scap run to populate new release values for [[phab:T427668|T427668]] * 10:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2212 [[phab:T434644|T434644]]', diff saved to https://phabricator.wikimedia.org/P96003 and previous config saved to /var/cache/conftool/dbconfig/20260812-101053-cwilliams.json * 10:09 moritzm: powercycle ganeti3005 * 10:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2203 to s1 primary [[phab:T434644|T434644]]', diff saved to https://phabricator.wikimedia.org/P96002 and previous config saved to /var/cache/conftool/dbconfig/20260812-100849-cwilliams.json * 10:08 cezmunsta: Starting s1 codfw failover from db2212 to db2203 - [[phab:T434644|T434644]] * 10:03 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1013.eqiad.wmnet with OS bookworm * 10:02 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1013.eqiad.wmnet with OS bookworm * 10:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2203 with weight 0 [[phab:T434644|T434644]]', diff saved to https://phabricator.wikimedia.org/P96001 and previous config saved to /var/cache/conftool/dbconfig/20260812-100134-cwilliams.json * 10:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 32 hosts with reason: Primary switchover s1 [[phab:T434644|T434644]] * 09:53 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1014.eqiad.wmnet with reason: host reimage * 09:50 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 09:50 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti3005.esams.wmnet * 09:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1014.eqiad.wmnet with reason: host reimage * 09:41 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1278: Pool in x1 * 09:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-launcher1003.eqiad.wmnet with OS bookworm * 09:37 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti3005.esams.wmnet * 09:37 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1013.eqiad.wmnet with OS bookworm * 09:34 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-presto1013.eqiad.wmnet with OS bookworm * 09:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1014.eqiad.wmnet with OS bookworm * 09:29 moritzm: failover ganeti master in esams to ganeti3008 * 09:26 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti3006.esams.wmnet * 09:26 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti3006.esams.wmnet * 09:24 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1012.eqiad.wmnet with OS bookworm * 09:23 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1009.eqiad.wmnet with OS bookworm * 09:23 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 09:18 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti3006.esams.wmnet * 09:16 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti3006.esams.wmnet * 09:14 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: fix regexp escaping bug - oblivian@cumin1003" * 09:14 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: fix regexp escaping bug - oblivian@cumin1003 * 09:13 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: fix regexp escaping bug - oblivian@cumin1003 * 09:13 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: fix regexp escaping bug - oblivian@cumin1003" * 09:03 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-launcher1003.eqiad.wmnet with reason: host reimage * 08:58 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-launcher1003.eqiad.wmnet with reason: host reimage * 08:55 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1278: Pool in x1 * 08:55 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1278 to dbctl [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95996 and previous config saved to /var/cache/conftool/dbconfig/20260812-085521-marostegui.json * 08:51 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1012.eqiad.wmnet with reason: host reimage * 08:45 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis testwiki in section s3 * 08:43 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1009.eqiad.wmnet with OS bookworm * 08:42 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1012.eqiad.wmnet with reason: host reimage * 08:41 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-launcher1003.eqiad.wmnet with OS bookworm * 08:40 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1013.eqiad.wmnet with OS bookworm * 08:38 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms3', diff saved to https://phabricator.wikimedia.org/P95995 and previous config saved to /var/cache/conftool/dbconfig/20260812-083816-marostegui.json * 08:38 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-master1003.eqiad.wmnet with OS bookworm * 08:37 marostegui: Failover ms3 [[phab:T434288|T434288]] * 08:37 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1268 to dbctl [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95994 and previous config saved to /var/cache/conftool/dbconfig/20260812-083722-marostegui.json * 08:35 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-web1001.eqiad.wmnet with OS bookworm * 08:32 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db2252.codfw.wmnet,db[1153,1268].eqiad.wmnet with reason: Switching over ms3 * 08:28 marostegui@cumin1003: dbctl commit (dc=all): 'Depool ms3 [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95993 and previous config saved to /var/cache/conftool/dbconfig/20260812-082852-marostegui.json * 08:25 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1012.eqiad.wmnet with OS bookworm * 08:25 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1011.eqiad.wmnet with OS bookworm * 08:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-master1003.eqiad.wmnet with reason: host reimage * 08:07 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-master1003.eqiad.wmnet with reason: host reimage * 08:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-web1001.eqiad.wmnet with reason: host reimage * 07:58 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-web1001.eqiad.wmnet with reason: host reimage * 07:50 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1003.eqiad.wmnet with OS bookworm * 07:38 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1011.eqiad.wmnet with reason: host reimage * 07:38 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-master1003.eqiad.wmnet with OS bookworm * 07:35 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti3007.esams.wmnet * 07:35 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti3007.esams.wmnet * 07:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1011.eqiad.wmnet with reason: host reimage * 07:27 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti3007.esams.wmnet * 07:25 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti3007.esams.wmnet * 07:25 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti3008.esams.wmnet * 07:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti3008.esams.wmnet * 07:22 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-web1001.eqiad.wmnet with OS bookworm * 07:19 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1003.eqiad.wmnet with OS bookworm * 07:18 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-master1003.eqiad.wmnet * 07:18 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host an-master1003.eqiad.wmnet * 07:17 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1011.eqiad.wmnet with OS bookworm * 07:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 07:16 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti3008.esams.wmnet * 07:15 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1009.eqiad.wmnet with OS bookworm * 07:14 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host an-master1003.eqiad.wmnet * 07:13 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-master1003.eqiad.wmnet * 07:13 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-master1003.eqiad.wmnet * 07:12 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-master1003.eqiad.wmnet * 07:11 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti3008.esams.wmnet * 07:07 arnaudb@dns1006: END - running authdns-update * 07:07 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti5007.eqsin.wmnet * 07:07 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti5007.eqsin.wmnet * 07:05 arnaudb@dns1006: START - running authdns-update * 06:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti5007.eqsin.wmnet * 06:54 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti5007.eqsin.wmnet * 06:38 moritzm: failover ganeti master in eqsin to ganeti5004 * 06:36 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti5006.eqsin.wmnet * 06:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti5006.eqsin.wmnet * 06:28 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti5006.eqsin.wmnet * 06:23 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti5006.eqsin.wmnet * 06:20 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti5005.eqsin.wmnet * 06:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti5005.eqsin.wmnet * 06:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti5005.eqsin.wmnet * 06:06 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti5005.eqsin.wmnet * 06:03 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti5004.eqsin.wmnet * 06:03 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti5004.eqsin.wmnet * 05:55 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti5004.eqsin.wmnet * 05:53 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti5004.eqsin.wmnet * 04:40 ryankemper: [[phab:T434494|T434494]] reimaged `an-tool1008.eqiad.wmnet` to bookworm; yarn.wikimedia.org is back up * 04:16 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-tool1008.eqiad.wmnet with OS bookworm * 03:58 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-tool1008.eqiad.wmnet with reason: host reimage * 03:53 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-tool1008.eqiad.wmnet with reason: host reimage * 03:41 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-tool1008.eqiad.wmnet with OS bookworm * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 45s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 00:25 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324427{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]], [[gerrit:1324429{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0]], [[gerrit:1324428{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]] (duration: 07m 55s) * 00:21 kemayo@deploy1003: kemayo: Continuing with deployment * 00:19 kemayo@deploy1003: kemayo: Backport for [[gerrit:1324427{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]], [[gerrit:1324429{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0]], [[gerrit:1324428{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:17 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1324427{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]], [[gerrit:1324429{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0]], [[gerrit:1324428{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]] == 2026-08-11 == * 21:37 sbassett: Deployed security fix for [[phab:T434521|T434521]] (wmf.15) * 21:29 sbassett: Deployed security fix for [[phab:T434521|T434521]] (wmf.14) * 21:19 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324370{{!}}Phase 4 of legal footer deployment (T432796)]], [[gerrit:1319804{{!}}Disable wgMFCustomSiteModules on English Wikipedia (T375538)]] (duration: 15m 26s) * 21:15 jdlrobson@deploy1003: jdlrobson: Continuing with deployment * 21:06 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1324370{{!}}Phase 4 of legal footer deployment (T432796)]], [[gerrit:1319804{{!}}Disable wgMFCustomSiteModules on English Wikipedia (T375538)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:03 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1324370{{!}}Phase 4 of legal footer deployment (T432796)]], [[gerrit:1319804{{!}}Disable wgMFCustomSiteModules on English Wikipedia (T375538)]] * 20:59 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 20:50 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324384{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]], [[gerrit:1324385{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]] (duration: 06m 58s) * 20:46 kemayo@deploy1003: kemayo: Continuing with deployment * 20:45 kemayo@deploy1003: kemayo: Backport for [[gerrit:1324384{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]], [[gerrit:1324385{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:43 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1324384{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]], [[gerrit:1324385{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]] * 20:42 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 20:42 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324386{{!}}build: Updating js-yaml to 3.15.1, 4.3.1]] (duration: 07m 36s) * 20:38 kemayo@deploy1003: kemayo: Continuing with deployment * 20:37 jhancock@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 20:37 kemayo@deploy1003: kemayo: Backport for [[gerrit:1324386{{!}}build: Updating js-yaml to 3.15.1, 4.3.1]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:35 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1324386{{!}}build: Updating js-yaml to 3.15.1, 4.3.1]] * 20:18 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 20:15 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 20:15 jhancock@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin1003" * 20:14 jhancock@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin1003" * 19:59 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 19:54 jhancock@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 19:10 brennen@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] (duration: 06m 41s) * 19:04 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns6001.wikimedia.org * 19:04 sukhe@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns6001.wikimedia.org * 19:04 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns5003.wikimedia.org * 19:04 sukhe@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns5003.wikimedia.org * 19:03 brennen@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 18:59 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns5003.wikimedia.org with OS trixie * 18:55 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns6001.wikimedia.org with OS trixie * 18:19 brett@cumin2002: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on P<nowiki>{</nowiki>cp7009.magru.wmnet<nowiki>}</nowiki> and A:cp - 9.2.15 Upgrade () * 18:14 brett@cumin2002: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on P<nowiki>{</nowiki>cp7009.magru.wmnet<nowiki>}</nowiki> and A:cp - 9.2.15 Upgrade () * 18:13 brennen@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 18:12 brett@cumin2002: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 9.2.15 Upgrade () * 18:09 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns5003.wikimedia.org with reason: host reimage * 18:06 brett@cumin2002: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 9.2.15 Upgrade () * 18:06 brennen: 1.47.0-wmf.15 train status ([[phab:T430834|T430834]]) - no current blockers, rolling to group0 * 18:05 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns5003.wikimedia.org with reason: host reimage * 18:05 brett: import trafficserver-9.2.15~deb13+wmf1 into trixie-wikimedia ([[phab:T434478|T434478]]) * 18:01 ladsgroup@cumin1003: END (PASS) - Cookbook sre.mysql.sanitarium_restart (exit_code=0) * 17:58 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns6001.wikimedia.org with reason: host reimage * 17:53 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324369{{!}}Enable desktop lazy loading on group0 (T148047)]] (duration: 07m 31s) * 17:52 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns6001.wikimedia.org with reason: host reimage * 17:49 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 17:49 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitarium_restart (exit_code=99) * 17:49 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 17:49 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7001.magru.wmnet * 17:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti7001.magru.wmnet * 17:48 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 17:47 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1324369{{!}}Enable desktop lazy loading on group0 (T148047)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:45 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1324369{{!}}Enable desktop lazy loading on group0 (T148047)]] * 17:39 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti7001.magru.wmnet * 17:36 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns5003.wikimedia.org with OS trixie * 17:34 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns6001.wikimedia.org with OS trixie * 17:31 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324361{{!}}Move FR config from IS.php to a dedicated file]], [[gerrit:1324363{{!}}Remove $wmg = $wg hacks in CentralAuth (T119117)]] (duration: 12m 23s) * 17:26 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 17:23 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1324361{{!}}Move FR config from IS.php to a dedicated file]], [[gerrit:1324363{{!}}Remove $wmg = $wg hacks in CentralAuth (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:19 sukhe: sudo cumin "A:cp-magru" "run-puppet-agent --enable 'merging CR 1324355'" * 17:18 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1324361{{!}}Move FR config from IS.php to a dedicated file]], [[gerrit:1324363{{!}}Remove $wmg = $wg hacks in CentralAuth (T119117)]] * 17:11 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-master1004.eqiad.wmnet with OS bookworm * 17:11 sukhe: sukhe@cp7005:~$ sudo puppet agent -tv * 17:02 sukhe: sudo cumin "A:cp-magru" "disable-puppet 'merging CR 1324355'" * 16:54 sukhe@dns1004: END - running authdns-update * 16:53 sukhe@dns1004: START - running authdns-update * 16:53 sukhe@dns1004: FAIL - running authdns-update * 16:51 sukhe@dns1004: START - running authdns-update * 16:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-master1004.eqiad.wmnet with reason: host reimage * 16:44 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-master1004.eqiad.wmnet with reason: host reimage * 16:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1157.eqiad.wmnet onto db1272.eqiad.wmnet * 16:40 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1157: Pool db1157.eqiad.wmnet in after cloning * 16:38 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 16:31 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324356{{!}}InitialiseSettings: Fix wgOATHAuthEnforce2FAForAll]] (duration: 06m 52s) * 16:30 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 16:28 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 16:27 reedy@deploy1003: reedy: Continuing with deployment * 16:26 reedy@deploy1003: reedy: Backport for [[gerrit:1324356{{!}}InitialiseSettings: Fix wgOATHAuthEnforce2FAForAll]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:24 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324356{{!}}InitialiseSettings: Fix wgOATHAuthEnforce2FAForAll]] * 16:13 sukhe: restart ntpsec.serviceon dns7001 * 16:09 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324335{{!}}InitialiseSettings: Enable 2FA enforcement on various private wikis (T428103)]] (duration: 06m 40s) * 16:08 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 16:06 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2204: Security update * 16:04 reedy@deploy1003: reedy: Continuing with deployment * 16:04 reedy@deploy1003: reedy: Backport for [[gerrit:1324335{{!}}InitialiseSettings: Enable 2FA enforcement on various private wikis (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:02 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7001.magru.wmnet * 16:02 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324335{{!}}InitialiseSettings: Enable 2FA enforcement on various private wikis (T428103)]] * 16:01 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 15:55 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1157: Pool db1157.eqiad.wmnet in after cloning * 15:54 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 15:54 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 15:41 dancy@deploy1003: Finished scap sync-world: Testing (duration: 06m 28s) * 15:40 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-master1004.eqiad.wmnet with OS bookworm * 15:40 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 15:35 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti4008.ulsfo.wmnet * 15:35 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti4008.ulsfo.wmnet * 15:34 dancy@deploy1003: Started scap sync-world: Testing * 15:34 dancy@deploy1003: Installation of scap version "4.279.0" completed for 3 hosts * 15:34 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-master1004.eqiad.wmnet with OS bookworm * 15:34 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 15:33 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-master1004.eqiad.wmnet with OS bookworm * 15:32 dancy@deploy1003: Installing scap version "4.279.0" for 3 host(s) * 15:32 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324339{{!}}Add /w/deployment-info.php entrypoint]] (duration: 07m 25s) * 15:30 moritzm: failover ganeti master in magru to ganeti7004 * 15:29 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti4008.ulsfo.wmnet * 15:28 dancy@deploy1003: dancy: Continuing with deployment * 15:28 tappof: remove 2026-05 swift log archives from centrallog to free some space ([[phab:T434502|T434502]]) * 15:27 dancy@deploy1003: dancy: Backport for [[gerrit:1324339{{!}}Add /w/deployment-info.php entrypoint]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:25 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324339{{!}}Add /w/deployment-info.php entrypoint]] * 15:20 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2204: Security update * 15:18 dancy@deploy1003: Installation of scap version "4.278.0" completed for 3 hosts * 15:18 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7004.magru.wmnet * 15:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti7004.magru.wmnet * 15:16 dancy@deploy1003: Installing scap version "4.278.0" for 3 host(s) * 15:14 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2204.codfw.wmnet with reason: Maintenance * 15:12 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 15:11 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-master1004.eqiad.wmnet with OS bookworm * 15:11 hashar: Restarting CI Jenkins on contint1003 due to Java upgrade. * 15:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2204 [[phab:T434565|T434565]]', diff saved to https://phabricator.wikimedia.org/P95984 and previous config saved to /var/cache/conftool/dbconfig/20260811-151126-cwilliams.json * 15:10 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti4008.ulsfo.wmnet * 15:10 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti7004.magru.wmnet * 15:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2207 to s2 primary [[phab:T434565|T434565]]', diff saved to https://phabricator.wikimedia.org/P95983 and previous config saved to /var/cache/conftool/dbconfig/20260811-150905-cwilliams.json * 15:08 cezmunsta: Starting s2 codfw failover from db2204 to db2207 - [[phab:T434565|T434565]] * 15:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2207 with weight 0 [[phab:T434565|T434565]]', diff saved to https://phabricator.wikimedia.org/P95982 and previous config saved to /var/cache/conftool/dbconfig/20260811-150402-cwilliams.json * 15:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s2 [[phab:T434565|T434565]] * 14:55 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1010.eqiad.wmnet with OS bookworm * 14:49 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns4003.wikimedia.org with OS trixie * 14:47 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1010.eqiad.wmnet with OS bookworm * 14:47 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7001.wikimedia.org with OS trixie * 14:44 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-presto1010.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:41 btullis@cumin1003: START - Cookbook sre.hosts.provision for host an-presto1010.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:40 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-presto1010.eqiad.wmnet with OS bookworm * 14:39 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 14:39 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-presto1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:36 btullis@cumin1003: START - Cookbook sre.hosts.provision for host an-presto1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:32 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1009.eqiad.wmnet with OS bookworm * 14:32 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 14:31 cwilliams@cumin1003: START - Cookbook sre.mysql.clone of db1157.eqiad.wmnet onto db1272.eqiad.wmnet * 14:30 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7004.magru.wmnet * 14:28 moritzm: failover ganeti master in ulsfo to ganeti4005 * 14:23 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7003.magru.wmnet * 14:23 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti7003.magru.wmnet * 14:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-master1004.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:22 btullis@cumin1003: START - Cookbook sre.hosts.provision for host an-master1004.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:21 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-master1004.eqiad.wmnet with OS bookworm * 14:19 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti4007.ulsfo.wmnet * 14:19 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti4007.ulsfo.wmnet * 14:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1010.eqiad.wmnet with OS bookworm * 14:17 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1008.eqiad.wmnet with OS bookworm * 14:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti7003.magru.wmnet * 14:12 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7003.magru.wmnet * 14:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti4007.ulsfo.wmnet * 14:11 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7002.magru.wmnet * 14:11 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti7002.magru.wmnet * 14:09 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324318{{!}}Revert "wmf-config/ProductionServices: set URL for urldownloader to service record" (T429175)]] (duration: 06m 46s) * 14:09 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7001.wikimedia.org with reason: host reimage * 14:06 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti4007.ulsfo.wmnet * 14:05 kharlan@deploy1003: kharlan: Continuing with deployment * 14:05 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti4006.ulsfo.wmnet * 14:04 jayme: updated calico to v3.30.7 on staging-eqiad - [[phab:T427400|T427400]] * 14:04 kharlan@deploy1003: kharlan: Backport for [[gerrit:1324318{{!}}Revert "wmf-config/ProductionServices: set URL for urldownloader to service record" (T429175)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti4006.ulsfo.wmnet * 14:03 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns4003.wikimedia.org with reason: host reimage * 14:03 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7001.wikimedia.org with reason: host reimage * 14:02 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti7002.magru.wmnet * 14:02 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1324318{{!}}Revert "wmf-config/ProductionServices: set URL for urldownloader to service record" (T429175)]] * 14:02 btullis@dns1004: FAIL - running authdns-update * 14:00 btullis@dns1004: START - running authdns-update * 13:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1008.eqiad.wmnet with reason: host reimage * 13:59 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'. * 13:59 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1313985{{!}}wmf-config/ProductionServices: set URL for urldownloader to service record (T429175)]] (duration: 25m 06s) * 13:58 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7002.magru.wmnet * 13:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti4006.ulsfo.wmnet * 13:57 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns4003.wikimedia.org with reason: host reimage * 13:56 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'. * 13:56 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7001.magru.wmnet * 13:56 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1008.eqiad.wmnet with reason: host reimage * 13:55 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'. * 13:55 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'. * 13:55 kharlan@deploy1003: kharlan, sukhe: Continuing with deployment * 13:53 marostegui: Failover ms2 [[phab:T434288|T434288]] * 13:52 marostegui: Failover ms1 [[phab:T434288|T434288]] * 13:52 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7001.magru.wmnet * 13:51 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti4006.ulsfo.wmnet * 13:48 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti4005.ulsfo.wmnet * 13:48 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti4005.ulsfo.wmnet * 13:44 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti4005.ulsfo.wmnet * 13:40 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1008.eqiad.wmnet with OS bookworm * 13:39 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns4003.wikimedia.org with OS trixie * 13:38 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns7001.wikimedia.org with OS trixie * 13:38 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1008.eqiad.wmnet with OS bookworm * 13:37 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti4005.ulsfo.wmnet * 13:36 kharlan@deploy1003: kharlan, sukhe: Backport for [[gerrit:1313985{{!}}wmf-config/ProductionServices: set URL for urldownloader to service record (T429175)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:34 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1313985{{!}}wmf-config/ProductionServices: set URL for urldownloader to service record (T429175)]] * 13:29 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2034.codfw.wmnet * 13:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2034.codfw.wmnet * 13:25 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.clone (exit_code=99) of db1157.eqiad.wmnet onto db1272.eqiad.wmnet * 13:25 cwilliams@cumin1003: START - Cookbook sre.mysql.clone of db1157.eqiad.wmnet onto db1272.eqiad.wmnet * 13:21 urbanecm@deploy1003: mwscript-k8s job started: namespaceDupes.php --wiki=frwiktionary --fix # [[phab:T415716|T415716]] * 13:21 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2034.codfw.wmnet * 13:20 urbanecm@deploy1003: mwscript-k8s job started: namespaceDupes.php --wiki=frwiktionary # [[phab:T415716|T415716]] * 13:19 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1323350{{!}}[tgwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T415307)]], [[gerrit:1322961{{!}}[slwiki] Revert temporary logo for Wikipedia 25 (Vector legacy + Vector 2022) (T414265)]], [[gerrit:1323827{{!}}[itwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T414320)]] (duration: 08m 00s) * 13:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-coord1003.eqiad.wmnet with OS bookworm * 13:15 urbanecm@deploy1003: urbanecm, superpes: Continuing with deployment * 13:13 urbanecm@deploy1003: urbanecm, superpes: Backport for [[gerrit:1323350{{!}}[tgwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T415307)]], [[gerrit:1322961{{!}}[slwiki] Revert temporary logo for Wikipedia 25 (Vector legacy + Vector 2022) (T414265)]], [[gerrit:1323827{{!}}[itwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T414320)]] synced to the testservers (see https://wiki * 13:11 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1323350{{!}}[tgwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T415307)]], [[gerrit:1322961{{!}}[slwiki] Revert temporary logo for Wikipedia 25 (Vector legacy + Vector 2022) (T414265)]], [[gerrit:1323827{{!}}[itwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T414320)]] * 13:11 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1323779{{!}}[ukwiki] Remove reviewer usergroup (T434252)]], [[gerrit:1323312{{!}}[frwiktionary] Add new Schème namespace and its talk (T415716)]] (duration: 06m 49s) * 13:10 marostegui@dns1004: END - running authdns-update * 13:08 marostegui@dns1004: START - running authdns-update * 13:07 marostegui@cumin1003: dbctl commit (dc=all): 'Repool ms2 [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95980 and previous config saved to /var/cache/conftool/dbconfig/20260811-130725-marostegui.json * 13:06 urbanecm@deploy1003: urbanecm, superpes: Continuing with deployment * 13:06 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1266 to dbctl [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95979 and previous config saved to /var/cache/conftool/dbconfig/20260811-130627-marostegui.json * 13:06 urbanecm@deploy1003: urbanecm, superpes: Backport for [[gerrit:1323779{{!}}[ukwiki] Remove reviewer usergroup (T434252)]], [[gerrit:1323312{{!}}[frwiktionary] Add new Schème namespace and its talk (T415716)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:04 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1323779{{!}}[ukwiki] Remove reviewer usergroup (T434252)]], [[gerrit:1323312{{!}}[frwiktionary] Add new Schème namespace and its talk (T415716)]] * 12:59 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db2253.codfw.wmnet,db[1151,1266].eqiad.wmnet with reason: Switching over ms2 * 12:54 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1157: Using as clone source * 12:53 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1157: Using as clone source * 12:51 marostegui@cumin1003: dbctl commit (dc=all): 'Depool ms2 [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95977 and previous config saved to /var/cache/conftool/dbconfig/20260811-125129-marostegui.json * 12:47 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 12:46 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 12:45 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 12:44 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'recommendation-api-ng' for release 'main' . * 12:44 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'recommendation-api-ng' for release 'main' . * 12:43 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'recommendation-api-ng' for release 'main' . * 12:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-coord1003.eqiad.wmnet with reason: host reimage * 12:43 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'ores-legacy' for release 'main' . * 12:42 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'ores-legacy' for release 'main' . * 12:42 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2165: Security update * 12:40 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-coord1003.eqiad.wmnet with reason: host reimage * 12:39 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'ores-legacy' for release 'main' . * 12:38 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' . * 12:38 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' . * 12:37 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' . * 12:34 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2034.codfw.wmnet * 12:30 jmm@cumin2002: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti-test2001.codfw.wmnet * 12:30 jmm@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host ganeti-test2001.codfw.wmnet * 12:25 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1179.eqiad.wmnet onto db1278.eqiad.wmnet * 12:25 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1179: Pool db1179.eqiad.wmnet in after cloning * 12:23 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-coord1003.eqiad.wmnet with OS bookworm * 12:22 moritzm: failover ganeti master in codfw/routed to ganeti2033 * 12:22 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2033.codfw.wmnet * 12:22 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2033.codfw.wmnet * 12:19 jmm@cumin2002: START - Cookbook sre.hosts.reboot-single for host ganeti-test2001.codfw.wmnet * 12:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 12:18 jmm@cumin2002: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti-test2001.codfw.wmnet * 12:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1008.eqiad.wmnet with OS bookworm * 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2033.codfw.wmnet * 12:07 moritzm: failover ganeti master in ganeti/test to ganeti-test2003 * 12:04 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 12:03 jmm@cumin2003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti4005.ulsfo.wmnet * 12:03 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti4005.ulsfo.wmnet * 12:00 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324283{{!}}Use maximum compression level in SqlBlobStore and SqlBagOStuff (T428377)]] (duration: 11m 37s) * 11:57 jmm@cumin2002: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti-test2002.codfw.wmnet * 11:57 jmm@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti-test2002.codfw.wmnet * 11:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2165: Security update * 11:54 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 11:52 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1324283{{!}}Use maximum compression level in SqlBlobStore and SqlBagOStuff (T428377)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:51 jmm@cumin2002: START - Cookbook sre.hosts.reboot-single for host ganeti-test2002.codfw.wmnet * 11:50 jmm@cumin2002: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti-test2002.codfw.wmnet * 11:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2165.codfw.wmnet with reason: Maintenance * 11:48 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@050d19e] (releasing): [[phab:T434186|T434186]] (duration: 01m 14s) * 11:48 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1324283{{!}}Use maximum compression level in SqlBlobStore and SqlBagOStuff (T428377)]] * 11:47 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@050d19e] (releasing): [[phab:T434186|T434186]] * 11:44 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@050d19e] (releasing): test jenkins deploy for [[phab:T434186|T434186]] (duration: 01m 08s) * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2165 [[phab:T434514|T434514]]', diff saved to https://phabricator.wikimedia.org/P95969 and previous config saved to /var/cache/conftool/dbconfig/20260811-114352-cwilliams.json * 11:43 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@050d19e] (releasing): test jenkins deploy for [[phab:T434186|T434186]] * 11:42 jmm@cumin2003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti-test2003.codfw.wmnet * 11:42 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti-test2003.codfw.wmnet * 11:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2161 to s8 primary [[phab:T434514|T434514]]', diff saved to https://phabricator.wikimedia.org/P95968 and previous config saved to /var/cache/conftool/dbconfig/20260811-114136-cwilliams.json * 11:40 cezmunsta: Starting s8 codfw failover from db2165 to db2161 - [[phab:T434514|T434514]] * 11:40 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1179: Pool db1179.eqiad.wmnet in after cloning * 11:36 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti-test2003.codfw.wmnet * 11:36 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti-test2003.codfw.wmnet * 11:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2161 with weight 0 [[phab:T434514|T434514]]', diff saved to https://phabricator.wikimedia.org/P95966 and previous config saved to /var/cache/conftool/dbconfig/20260811-113449-cwilliams.json * 11:34 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 25 hosts with reason: Primary switchover s8 [[phab:T434514|T434514]] * 11:29 moritzm: installing Python 3.11 security updates * 11:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-presto1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 11:26 btullis@cumin1003: START - Cookbook sre.hosts.provision for host an-presto1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 11:23 btullis@dns1004: END - running authdns-update * 11:21 btullis@dns1004: START - running authdns-update * 11:20 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-presto1008.eqiad.wmnet with OS bookworm * 11:20 moritzm: installing curl security updates * 11:11 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1007.eqiad.wmnet with OS bookworm * 10:45 tappof: bump space for prometheus k8s-dse in codfw * 10:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1007.eqiad.wmnet with reason: host reimage * 10:38 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1007.eqiad.wmnet with reason: host reimage * 10:37 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-coord1004.eqiad.wmnet with OS bookworm * 10:35 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1008.eqiad.wmnet with OS bookworm * 10:34 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1179: Depool db1179.eqiad.wmnet to then clone it to db1278.eqiad.wmnet - marostegui@cumin1003 * 10:34 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1008.eqiad.wmnet with OS bookworm * 10:25 fceratto@cumin1003: dbctl commit (dc=all): 'Remove db1177 [[phab:T433474|T433474]]', diff saved to https://phabricator.wikimedia.org/P95964 and previous config saved to /var/cache/conftool/dbconfig/20260811-102527-fceratto.json * 10:22 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1008.eqiad.wmnet with OS bookworm * 10:22 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1007.eqiad.wmnet with OS bookworm * 10:21 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1006.eqiad.wmnet with OS bookworm * 10:20 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 10:18 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1179: Depool db1179.eqiad.wmnet to then clone it to db1278.eqiad.wmnet - marostegui@cumin1003 * 10:18 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1179.eqiad.wmnet onto db1278.eqiad.wmnet * 10:17 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 10:17 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 10:14 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 10:09 blake@deploy1003: Stopping before sync operations * 10:09 blake@deploy1003: Started scap sync-world: Non-deployment run to populate release values for [[phab:T427668|T427668]] * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 10:04 fceratto@cumin1003: Removing db1177 from zarcillo [[phab:T433474|T433474]] * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1177.eqiad.wmnet * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1177.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:03 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1177.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:00 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1006.eqiad.wmnet with reason: host reimage * 09:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-coord1004.eqiad.wmnet with reason: host reimage * 09:57 marostegui: Failover m1 from db1164 to db1213 - [[phab:T434493|T434493]] * 09:57 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1006.eqiad.wmnet with reason: host reimage * 09:55 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:54 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2232].codfw.wmnet,db[1164,1213,1217].eqiad.wmnet with reason: Primary switchover m1 [[phab:T434493|T434493]] * 09:52 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-coord1004.eqiad.wmnet with reason: host reimage * 09:49 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1213.eqiad.wmnet with OS trixie * 09:49 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1177.eqiad.wmnet * 09:41 moritzm: installing Linux 6.12.101 on Trixie hosts * 09:40 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1006.eqiad.wmnet with OS bookworm * 09:35 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-coord1004.eqiad.wmnet with OS bookworm * 09:28 moritzm: installing node-tar security updates * 09:27 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1213.eqiad.wmnet with reason: host reimage * 09:22 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1213.eqiad.wmnet with reason: host reimage * 09:09 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1177: Decommission * 09:08 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db1177: Decommission * 09:08 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 09:08 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.decommission (exit_code=99) * 09:06 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1213.eqiad.wmnet with OS trixie * 09:06 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 09:05 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1213.eqiad.wmnet with reason: Reimage * 08:53 marostegui@dns1004: END - running authdns-update * 08:51 marostegui@dns1004: START - running authdns-update * 08:48 marostegui: Switchover ms1 master in eqiad [[phab:T434288|T434288]] * 08:48 marostegui@cumin1003: dbctl commit (dc=all): 'Repool ms1 [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95962 and previous config saved to /var/cache/conftool/dbconfig/20260811-084804-marostegui.json * 08:40 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1267 to dbctl [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95961 and previous config saved to /var/cache/conftool/dbconfig/20260811-084054-marostegui.json * 08:29 marostegui: Failover m1 from db1213 to db1164 - [[phab:T434043|T434043]] * 08:25 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2232].codfw.wmnet,db[1164,1213,1217].eqiad.wmnet with reason: Primary switchover m1 [[phab:T434043|T434043]] * 08:22 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db2251.codfw.wmnet,db[1152,1267].eqiad.wmnet with reason: Switching over ms1 * 08:22 marostegui@cumin1003: dbctl commit (dc=all): 'Depool ms1 [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95960 and previous config saved to /var/cache/conftool/dbconfig/20260811-082201-marostegui.json * 08:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: Switching over ms1 * 08:20 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.parsercache (exit_code=99) * 08:20 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 08:20 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1152: Switching over ms1 * 08:19 slyngshede@dns1004: END - running authdns-update * 08:18 moritzm: installing openjdk-21 security updates * 08:17 slyngshede@dns1004: START - running authdns-update * 08:16 moritzm: imported jenkins 2.568.2 to thirdparty/jenkins for trixie-wikimedia * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.12 (duration: 02m 26s) * 03:36 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] (duration: 33m 33s) * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 35s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-10 == * 14:54 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1323973{{!}}mmv.bootstrap: Fix getUrlParam to account for TIFF lossy/lossless param (T434333)]] (duration: 11m 24s) * 14:50 krinkle@deploy1003: krinkle: Continuing with deployment * 14:45 krinkle@deploy1003: krinkle: Backport for [[gerrit:1323973{{!}}mmv.bootstrap: Fix getUrlParam to account for TIFF lossy/lossless param (T434333)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:43 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1323973{{!}}mmv.bootstrap: Fix getUrlParam to account for TIFF lossy/lossless param (T434333)]] * 14:07 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1323967{{!}}updateIsActiveFlagForMentees: Commit the final partial batch (T432959)]] (duration: 10m 33s) * 13:56 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1323967{{!}}updateIsActiveFlagForMentees: Commit the final partial batch (T432959)]] * 13:45 dani@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply * 13:45 dani@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply * 13:45 dani@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply * 13:45 dani@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply * 13:45 dani@deploy1003: helmfile [staging] DONE helmfile.d/services/miscweb: apply * 13:44 dani@deploy1003: helmfile [staging] START helmfile.d/services/miscweb: apply * 13:38 wmde-fisch@deploy1003: Finished scap sync-world: Backport for [[gerrit:1323939{{!}}Enable sub-references on more group2 wikis (batch3) (T432731)]] (duration: 33m 21s) * 13:25 wmde-fisch@deploy1003: wmde-fisch: Continuing with deployment * 13:22 wmde-fisch@deploy1003: wmde-fisch: Backport for [[gerrit:1323939{{!}}Enable sub-references on more group2 wikis (batch3) (T432731)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:05 wmde-fisch@deploy1003: Started scap sync-world: Backport for [[gerrit:1323939{{!}}Enable sub-references on more group2 wikis (batch3) (T432731)]] * 07:57 hashar@deploy1003: Finished deploy [integration/docroot@7772132]: update build dependencies (duration: 00m 13s) * 07:57 hashar@deploy1003: Started deploy [integration/docroot@7772132]: update build dependencies * 07:35 _joe_: restarting squid on urldownloader1006 * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 48s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-09 == * 16:01 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:01 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:01 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:00 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 36s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-08 == * 05:31 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9] (wcqs): [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) (duration: 02m 36s) * 05:28 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9] (wcqs): [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) * 04:56 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 04:55 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 04:47 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) (duration: 19m 22s) * 04:28 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) * 04:19 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) (duration: 00m 06s) * 04:18 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) * 04:17 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) (duration: 00m 28s) * 04:16 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) * 03:52 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 03:52 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 34s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-07 == * 23:30 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:29 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 22:45 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 22:43 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 22:41 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 22:41 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 20:54 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 20:32 andrewbogott: restarting puppetserver service on puppetserver* for [[phab:T434339|T434339]] * 19:52 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:45 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 19:32 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:25 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:22 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 19:21 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 18:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:41 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 18:35 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 18:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 18:22 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 18:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 18:16 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 18:12 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:09 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:08 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:07 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:04 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:01 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:00 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:00 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 17:59 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 17:25 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 17:14 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 17:13 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 17:13 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 17:13 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:54 maryum: Deployed security fix for [[phab:T434278|T434278]] * 16:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 16:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 16:27 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --verbose --scoreLessThan=0.7 --exceptDatasetChecksums=[[phab:T434319|T434319]]-enwiki-models.txt # [[phab:T434319|T434319]] * 16:06 cdobbins@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp5022.eqsin.wmnet with OS trixie * 15:13 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 14:19 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 14:17 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 13:50 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1156.eqiad.wmnet onto db1271.eqiad.wmnet * 13:50 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1271: Pool db1271.eqiad.wmnet in after cloning * 13:02 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1271: Pool db1271.eqiad.wmnet in after cloning * 12:19 jayme: updated calico to v3.30.7 on staging-codfw - [[phab:T427400|T427400]] * 12:09 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 12:06 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 12:06 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 12:05 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 12:02 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1156: Pool db1156.eqiad.wmnet in after cloning * 11:38 bjensen: sudo -i reprepro -C main include trixie-wikimedia $<nowiki>{</nowiki>HOME<nowiki>}</nowiki>/httpbb/trixie/httpbb_$<nowiki>{</nowiki>VERSION?<nowiki>}</nowiki>-1+deb13u1_amd64.changes #[[phab:T434052|T434052]] * 11:35 bjensen: sudo -i reprepro -C main include bookworm-wikimedia $<nowiki>{</nowiki>HOME<nowiki>}</nowiki>/httpbb/bookworm/httpbb_$<nowiki>{</nowiki>VERSION?<nowiki>}</nowiki>-1_amd64.changes #[[phab:T434052|T434052]] * 11:30 marostegui@cumin1003: dbctl commit (dc=all): 'Adding db1271 to dbctl', diff saved to https://phabricator.wikimedia.org/P95945 and previous config saved to /var/cache/conftool/dbconfig/20260807-113006-marostegui.json * 11:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1156: Pool db1156.eqiad.wmnet in after cloning * 10:23 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 10:22 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 10:22 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 10:21 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 10:20 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 10:20 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 10:19 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 10:18 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 10:06 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on 21 hosts with reason: cloning * 10:01 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1156: Depool db1156.eqiad.wmnet to then clone it to db1271.eqiad.wmnet - marostegui@cumin1003 * 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1156: Depool db1156.eqiad.wmnet to then clone it to db1271.eqiad.wmnet - marostegui@cumin1003 * 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1156.eqiad.wmnet onto db1271.eqiad.wmnet * 09:15 jynus: started stress testing db1245 dbs [[phab:T431115|T431115]] * 08:19 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:18 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:16 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:14 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:13 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:10 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:06 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:05 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:00 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 10 days, 0:00:00 on ml-serve1015.eqiad.wmnet with reason: Downtime to get full picture of current BIOS settings beyond what Redfish shows * 08:00 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 07:54 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 07:54 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 07:53 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:52 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:51 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:50 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:49 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:48 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:47 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:45 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:45 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:41 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:38 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:37 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 06:35 jayme: updated istio to 1.29.4 on wikikube eqiad - [[phab:T427401|T427401]] * 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1178.eqiad.wmnet * 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1178.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 06:06 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1178.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 05:55 marostegui@cumin1003: START - Cookbook sre.dns.netbox * 05:49 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1178.eqiad.wmnet * 05:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 05:46 marostegui@cumin1003: Removing db1178 from zarcillo [[phab:T433471|T433471]] * 05:45 marostegui@cumin1003: START - Cookbook sre.mysql.decommission * 02:42 denisse: Extended volume on prometheus2008 for the disk space alert as per https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_host_running_out_of_space * 02:37 denisse: Extended volume on prometheus2007 tor the disk space alert as per https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_host_running_out_of_space * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 56s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-06 == * 21:39 maryum: Deploy security patch for [[phab:T433070|T433070]] * 21:29 maryum: Deploy security patch for [[phab:T434189|T434189]] * 20:48 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] (duration: 08m 12s) * 20:44 aude@deploy1003: lmora, aude, anzx: Continuing with deployment * 20:41 aude@deploy1003: lmora, aude, anzx: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be * 20:41 ebernhardson@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:41 ebernhardson@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 20:40 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] * 20:37 ebernhardson@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:37 ebernhardson@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 20:32 ebernhardson@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:32 ebernhardson@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 20:31 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] (duration: 06m 41s) * 20:27 cjming@deploy1003: cjming, ebernhardson, chlod: Continuing with deployment * 20:26 cjming@deploy1003: cjming, ebernhardson, chlod: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:24 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] * 20:18 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] (duration: 09m 22s) * 20:14 cjming@deploy1003: cjming, tsev: Continuing with deployment * 20:11 cjming@deploy1003: cjming, tsev: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:09 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] * 19:41 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply * 19:40 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply * 19:31 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 19:31 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 19:00 cdobbins@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cp5022.eqsin.wmnet with OS trixie * 18:25 ladsgroup@deploy1003: Finished scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) (duration: 06m 08s) * 18:19 ladsgroup@deploy1003: Started scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) * 18:18 ladsgroup@deploy1003: Stopping before sync operations * 18:17 ladsgroup@deploy1003: Started scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) * 17:55 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 16:50 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 16:35 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1001.eqiad.wmnet with OS bookworm * 16:19 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1002.eqiad.wmnet with reason: host reimage * 16:16 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1002.eqiad.wmnet with reason: host reimage * 16:05 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1001.eqiad.wmnet with reason: host reimage * 16:00 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1001.eqiad.wmnet with reason: host reimage * 15:57 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 15:43 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm * 15:29 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1001.eqiad.wmnet with OS bookworm * 15:29 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:58 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm * 14:57 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-drmrs ([[phab:T428495|T428495]]) * 14:55 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-drmrs ([[phab:T428495|T428495]]) * 14:55 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-ui1001.eqiad.wmnet with OS bookworm * 14:54 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-presto1001.eqiad.wmnet with OS bookworm * 14:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-magru ([[phab:T428495|T428495]]) * 14:49 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-magru ([[phab:T428495|T428495]]) * 14:48 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1001.eqiad.wmnet with OS bookworm * 14:46 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-esams ([[phab:T428495|T428495]]) * 14:44 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-esams ([[phab:T428495|T428495]]) * 14:43 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:42 brouberol@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:42 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:42 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 14:40 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 14:40 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:38 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-ui1001.eqiad.wmnet with reason: host reimage * 14:34 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-presto1001.eqiad.wmnet with reason: host reimage * 14:28 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-ui1001.eqiad.wmnet with reason: host reimage * 14:27 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-presto1001.eqiad.wmnet with reason: host reimage * 14:23 sukhe: sudo cumin -b2 'A:cp-text' "run-puppet-agent --enable 'merging CR 1290731'": [[phab:T425441|T425441]] * 14:18 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo for hosts in the wikimedia.org domain - [[phab:T428495|T428495]] * 14:16 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-presto1001.eqiad.wmnet with OS bookworm * 14:14 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-ui1001.eqiad.wmnet with OS bookworm * 14:12 sukhe: sudo cumin 'A:cp-text' "disable-puppet 'merging CR 1290731'": [[phab:T425441|T425441]] * 14:11 swfrench-wmf: restarted navtiming on webperf1003 - [[phab:T428495|T428495]] * 14:04 swfrench-wmf: begin rolling restart of confd in drmrs, eqiad, esams, magru - [[phab:T428495|T428495]] * 14:04 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm * 14:04 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:02 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-client1002.eqiad.wmnet with OS bookworm * 13:58 swfrench-wmf: authdns update to direct eqiad-associated etcd clients back to eqiad - [[phab:T428495|T428495]] * 13:58 swfrench@dns1004: END - running authdns-update * 13:56 swfrench@dns1004: START - running authdns-update * 13:49 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:44 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:31 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 13:29 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 13:26 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 13:23 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 13:22 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 13:19 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 13:18 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 13:18 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 13:17 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 13:16 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 13:13 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 13:11 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 13:09 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 13:06 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 13:06 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-client1002.eqiad.wmnet with OS bookworm * 13:05 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revision-models' for release 'main' . * 13:05 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:05 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revision-models' for release 'main' . * 13:04 brouberol@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-test-client1002.eqiad.wmnet with OS bookworm * 13:04 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 13:03 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 13:02 aikochou@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:00 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'readability' for release 'main' . * 12:59 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'readability' for release 'main' . * 12:58 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 12:57 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'logo-detection' for release 'main' . * 12:57 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'logo-detection' for release 'main' . * 12:57 aikochou@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 12:55 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:54 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:53 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 12:53 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:50 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 12:48 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 12:46 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'article-models' for release 'main' . * 12:45 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'article-models' for release 'main' . * 12:41 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . * 12:39 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . * 12:38 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-client1002.eqiad.wmnet with OS bookworm * 12:12 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply * 12:12 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply * 12:09 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:08 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 11:58 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2187: Security update * 11:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:24 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:16 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 11:15 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 11:10 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2187: Security update * 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2187.codfw.wmnet with reason: Maintenance * 10:56 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 10:56 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 10:56 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 10:56 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 10:54 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 10:53 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 10:09 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2187: Security update * 10:07 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2187: Security update * 09:39 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms2', diff saved to https://phabricator.wikimedia.org/P95929 and previous config saved to /var/cache/conftool/dbconfig/20260806-093908-marostegui.json * 09:36 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1178 from dbctl [[phab:T433471|T433471]]', diff saved to https://phabricator.wikimedia.org/P95928 and previous config saved to /var/cache/conftool/dbconfig/20260806-093632-marostegui.json * 09:33 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 09:31 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 09:30 topranks: bounce cr3-eqsin<->cr2-eqiad bgp session to disable no-prepend command * 09:20 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2253.codfw.wmnet,db1151.eqiad.wmnet with reason: cloning * 09:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1151: Cloning * 09:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:19 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 09:19 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1151: Cloning * 09:10 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 09:09 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup2003.codfw.wmnet * 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup2003.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 09:06 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup2003.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 09:03 klausman@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:02 jynus@cumin1003: START - Cookbook sre.dns.netbox * 09:02 klausman@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 08:57 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup2003.codfw.wmnet * 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup1003.eqiad.wmnet * 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:54 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms3', diff saved to https://phabricator.wikimedia.org/P95925 and previous config saved to /var/cache/conftool/dbconfig/20260806-085422-marostegui.json * 08:53 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:46 jynus@cumin1003: START - Cookbook sre.dns.netbox * 08:39 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup1003.eqiad.wmnet * 08:29 XioNoX: push pfw policy - [[phab:T434115|T434115]] * 08:14 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 08:00 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 08:00 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:58 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revision-models' for release 'main' . * 07:56 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 07:54 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'readability' for release 'main' . * 07:53 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'logo-detection' for release 'main' . * 07:51 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'llm' for release 'main' . * 07:48 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . * 07:37 jayme: updated istio to 1.29.4 on wikikube codfw - [[phab:T427401|T427401]] * 07:08 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2252.codfw.wmnet,db1153.eqiad.wmnet with reason: cloning * 07:07 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1153: Cloning * 07:07 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1153: Cloning * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 40s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-05 == * 23:24 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1009.eqiad.wmnet with OS bookworm * 23:03 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1009.eqiad.wmnet with reason: host reimage * 22:59 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1009.eqiad.wmnet with reason: host reimage * 22:43 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1009.eqiad.wmnet with OS bookworm * 22:38 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1009.eqiad.wmnet * 22:34 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1009.eqiad.wmnet * 22:25 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1008.eqiad.wmnet with OS bookworm * 22:04 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1008.eqiad.wmnet with reason: host reimage * 22:00 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1008.eqiad.wmnet with reason: host reimage * 21:48 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:47 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:46 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:44 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1008.eqiad.wmnet with OS bookworm * 21:43 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:41 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1008.eqiad.wmnet * 21:36 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1008.eqiad.wmnet * 21:14 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:12 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad * 21:12 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad * 21:10 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=eqiad * 21:08 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:07 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:07 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=eqiad * 21:04 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:03 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1006 * 21:02 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1006 * 21:00 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:56 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:56 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:55 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 20:55 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 20:55 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1007.eqiad.wmnet with OS bookworm * 20:51 vriley@cumin1003: START - Cookbook sre.dns.netbox * 20:43 ebernhardson: [[phab:T434008|T434008]]: changing cloudelastic:9643 from auto_expand_replicas to number_of_replicas * 20:34 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1007.eqiad.wmnet with reason: host reimage * 20:27 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1007.eqiad.wmnet with reason: host reimage * 20:24 cjming: end of UTC late backport window * 20:23 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] (duration: 06m 26s) * 20:18 cjming@deploy1003: cjming: Continuing with deployment * 20:18 cjming@deploy1003: cjming: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:16 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] * 20:12 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1007.eqiad.wmnet with OS bookworm * 20:12 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] (duration: 08m 41s) * 20:08 swfrench@cumin2002: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host conf1007.eqiad.wmnet with OS bookworm * 20:08 jforrester@deploy1003: jforrester: Continuing with deployment * 20:07 jforrester@deploy1003: jforrester: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:03 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] * 19:51 inflatador: [bking@puppetserver1001] ~$ sudo puppetserver ca sign --certname an-worker1189.eqiad.wmnet [[phab:T434142|T434142]] * 19:47 bking@cumin2003: DONE (FAIL) - Cookbook sre.puppet.renew-cert (exit_code=99) for an-worker1189.eqiad.wmnet: Renew puppet certificate - bking@cumin2003 * 19:46 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:30 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1007.eqiad.wmnet with OS trixie * 19:30 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 19:29 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 19:20 swfrench-wmf: silenced EtcdRelicationDown 0cb709a9-f244-4f1e-971f-{{Gerrit|440ec65e7fd7}} - [[phab:T428495|T428495]] * 19:13 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1007.eqiad.wmnet with OS bookworm * 19:12 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1007.eqiad.wmnet with reason: host reimage * 19:09 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1007.eqiad.wmnet * 19:07 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1007.eqiad.wmnet with reason: host reimage * 19:03 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1007.eqiad.wmnet * 18:52 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1007.eqiad.wmnet with OS trixie * 18:52 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1007.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:35 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 18:34 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 18:34 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 18:30 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1007.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:28 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:28 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1007] - vriley@cumin1003" * 18:27 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1007] - vriley@cumin1003" * 18:23 vriley@cumin1003: START - Cookbook sre.dns.netbox * 18:22 vriley@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 18:22 robh@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:19 vriley@cumin1003: START - Cookbook sre.dns.netbox * 18:13 robh@cumin2002: START - Cookbook sre.hosts.provision for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:31 jasmine@cumin2002: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-main-eqiad * 17:12 mutante: LDAP - added vwalters to group ciadmin - [[phab:T433615|T433615]] * 16:58 aokoth@deploy1003: Finished deploy [phabricator/deployment@e2ebca5]: Deploy Phab (duration: 00m 34s) * 16:57 aokoth@deploy1003: Started deploy [phabricator/deployment@e2ebca5]: Deploy Phab * 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad * 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=eqiad * 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad * 16:53 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:41 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2187.codfw.wmnet * 16:41 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2187.codfw.wmnet * 16:41 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker2187.codfw.wmnet * 16:41 cgoubert@cumin2003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker2187.codfw.wmnet * 16:40 jasmine@cumin2002: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-main-eqiad * 16:40 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:34 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:25 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-magru and A:liberica ([[phab:T428495|T428495]]) * 16:23 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-magru and A:liberica ([[phab:T428495|T428495]]) * 16:20 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-drmrs and A:liberica ([[phab:T428495|T428495]]) * 16:19 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-drmrs and A:liberica ([[phab:T428495|T428495]]) * 16:18 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-esams and A:liberica ([[phab:T428495|T428495]]) * 16:16 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-esams and A:liberica ([[phab:T428495|T428495]]) * 16:06 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1159.eqiad.wmnet * 16:06 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1159.eqiad.wmnet * 16:06 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1159.eqiad.wmnet * 16:05 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] (duration: 09m 11s) * 15:58 reedy@deploy1003: reedy: Continuing with deployment * 15:58 reedy@deploy1003: reedy: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:56 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] * 15:54 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1159.eqiad.wmnet with OS trixie * 15:38 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:33 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1159.eqiad.wmnet with reason: host reimage * 15:32 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:27 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1159.eqiad.wmnet with reason: host reimage * 15:10 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1159 * 15:10 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1159 * 15:00 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo for hosts in the wikimedia.org domain - [[phab:T428495|T428495]] * 14:55 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS trixie * 14:54 swfrench-wmf: restarted navtiming on webperf1003 - [[phab:T428495|T428495]] * 14:52 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1159 * 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1159.eqiad.wmnet 129.48.64.10.in-addr.arpa 9.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:52 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1159.eqiad.wmnet 129.48.64.10.in-addr.arpa 9.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1159 - jayme@cumin1003" * 14:52 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1159 - jayme@cumin1003" * 14:49 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:48 jayme@cumin1003: START - Cookbook sre.dns.netbox * 14:47 swfrench-wmf: begin rolling restart of confd in drmrs, eqiad, esams, magru - [[phab:T428495|T428495]] * 14:47 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1159 * 14:46 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:46 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:46 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1159.eqiad.wmnet with OS trixie * 14:45 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:44 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:44 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1159.eqiad.wmnet * 14:43 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:43 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1159.eqiad.wmnet * 14:43 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:43 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1159.eqiad.wmnet * 14:43 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:43 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1157.eqiad.wmnet * 14:43 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1157.eqiad.wmnet * 14:43 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1157.eqiad.wmnet * 14:42 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:42 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:42 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:41 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:41 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:41 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:41 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1003.eqiad.wmnet with OS bookworm * 14:39 swfrench-wmf: authdns update to direct eqiad-associated etcd clients to codfw - [[phab:T428495|T428495]] * 14:39 swfrench@dns1004: END - running authdns-update * 14:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 14:37 swfrench@dns1004: START - running authdns-update * 14:37 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:37 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:35 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:35 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 14:28 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:27 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1157.eqiad.wmnet with OS trixie * 14:27 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:27 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:27 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 14:26 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 14:26 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:26 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 14:26 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 14:26 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:26 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host search-loader1002.eqiad.wmnet with OS trixie * 14:19 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:19 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:15 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1003.eqiad.wmnet with reason: host reimage * 14:14 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:14 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:13 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 * 14:13 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host mc2046 * 14:13 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS trixie * 14:11 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1003.eqiad.wmnet with reason: host reimage * 14:10 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:09 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:09 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:09 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:08 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:08 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1157.eqiad.wmnet with reason: host reimage * 14:08 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:04 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 14:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on search-loader1002.eqiad.wmnet with reason: host reimage * 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=eqiad * 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=eqiad * 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=eqiad * 14:00 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:59 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:58 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1157.eqiad.wmnet with reason: host reimage * 13:57 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on search-loader1002.eqiad.wmnet with reason: host reimage * 13:54 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1003.eqiad.wmnet with OS bookworm * 13:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host search-loader1002.eqiad.wmnet with OS trixie * 13:43 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1157 * 13:42 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1157 * 13:40 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1157 * 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1157.eqiad.wmnet 183.32.64.10.in-addr.arpa 3.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1157.eqiad.wmnet 183.32.64.10.in-addr.arpa 3.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1157 - jayme@cumin1003" * 13:39 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1157 - jayme@cumin1003" * 13:39 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] (duration: 07m 00s) * 13:35 jayme@cumin1003: START - Cookbook sre.dns.netbox * 13:35 reedy@deploy1003: reedy: Continuing with deployment * 13:34 reedy@deploy1003: reedy: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:32 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] * 13:23 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1157 * 13:22 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1157.eqiad.wmnet with OS trixie * 13:22 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1157.eqiad.wmnet * 13:22 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1157.eqiad.wmnet * 13:21 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1157.eqiad.wmnet * 13:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1156.eqiad.wmnet * 13:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1156.eqiad.wmnet * 13:15 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1156.eqiad.wmnet * 13:01 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1156.eqiad.wmnet with OS trixie * 12:42 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1156.eqiad.wmnet with reason: host reimage * 12:38 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1156.eqiad.wmnet with reason: host reimage * 12:32 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 12:31 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 12:30 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 12:28 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 12:26 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 12:24 topranks: update bgp confed settings in eqsin * 12:22 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1156 * 12:22 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1156 * 12:22 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 12:19 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1156 * 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1156.eqiad.wmnet 110.32.64.10.in-addr.arpa 0.1.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:19 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1156.eqiad.wmnet 110.32.64.10.in-addr.arpa 0.1.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1156 - jayme@cumin1003" * 12:19 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1156 - jayme@cumin1003" * 12:17 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:14 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS trixie * 12:09 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:06 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:04 jayme@cumin1003: START - Cookbook sre.dns.netbox * 12:04 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 12:02 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'article-models' for release 'main' . * 12:01 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1156 * 12:01 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1156.eqiad.wmnet with OS trixie * 11:59 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1156.eqiad.wmnet * 11:59 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1156.eqiad.wmnet * 11:59 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1156.eqiad.wmnet * 11:57 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 11:53 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 11:53 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:52 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:52 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:50 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:50 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:50 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:49 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:48 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:47 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:47 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:45 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:45 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:44 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:44 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:44 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:43 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:42 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:38 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:35 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 * 11:35 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host mc2046 * 11:34 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS trixie * 11:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:27 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:21 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:21 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:18 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:18 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:18 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:18 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:13 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:13 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:09 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:08 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:07 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:06 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:06 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:05 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:05 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:04 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 11:04 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:24 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:24 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:17 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:16 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1155.eqiad.wmnet * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1155.eqiad.wmnet * 10:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1155.eqiad.wmnet * 10:14 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:14 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:11 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:11 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:10 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:09 aikochou@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop: sync * 10:09 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:09 aikochou@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop: sync * 10:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:07 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:05 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:05 aikochou@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop: sync * 10:05 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:05 aikochou@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop: sync * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:04 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:04 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:04 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1155.eqiad.wmnet with OS trixie * 09:52 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms1', diff saved to https://phabricator.wikimedia.org/P95918 and previous config saved to /var/cache/conftool/dbconfig/20260805-095212-marostegui.json * 09:44 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1152: after cloning * 09:44 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.parsercache (exit_code=99) * 09:44 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 09:44 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1152: after cloning * 09:43 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1155.eqiad.wmnet with reason: host reimage * 09:40 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1155.eqiad.wmnet with reason: host reimage * 09:32 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 09:32 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:31 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 09:31 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:27 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1155 * 09:27 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1155 * 09:25 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 09:24 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 09:24 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 09:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:23 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 09:23 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 09:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:22 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 09:22 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 09:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:20 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 09:20 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:17 XioNoX: push pfw policies - [[phab:T434038|T434038]] * 09:14 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1155 * 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1155.eqiad.wmnet 109.32.64.10.in-addr.arpa 9.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1155.eqiad.wmnet 109.32.64.10.in-addr.arpa 9.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1155 - jayme@cumin1003" * 09:14 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1155 - jayme@cumin1003" * 09:10 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 09:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:09 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2251.codfw.wmnet,db1152.eqiad.wmnet with reason: cloning * 09:09 jayme@cumin1003: START - Cookbook sre.dns.netbox * 09:08 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 09:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: Cloning * 09:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:05 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 09:05 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1152: Cloning * 08:38 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1155 * 08:37 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1155.eqiad.wmnet with OS trixie * 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1171.eqiad.wmnet * 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1171.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:29 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95913 and previous config saved to /var/cache/conftool/dbconfig/20260805-082908-ladsgroup.json * 08:27 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1171.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:22 jynus@cumin1003: START - Cookbook sre.dns.netbox * 08:18 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249', diff saved to https://phabricator.wikimedia.org/P95912 and previous config saved to /var/cache/conftool/dbconfig/20260805-081823-ladsgroup.json * 08:17 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1171.eqiad.wmnet * 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1150.eqiad.wmnet * 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1150.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:15 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1150.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:15 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 08:14 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1155.eqiad.wmnet * 08:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1155.eqiad.wmnet * 08:14 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1155.eqiad.wmnet * 08:11 jynus@cumin1003: START - Cookbook sre.dns.netbox * 08:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249', diff saved to https://phabricator.wikimedia.org/P95911 and previous config saved to /var/cache/conftool/dbconfig/20260805-080737-ladsgroup.json * 08:05 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1150.eqiad.wmnet * 08:02 marostegui: Depool clouddb1020 (s5,s8) [[phab:T434048|T434048]] * 08:02 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1020.eqiad.wmnet,service=s8 * 08:02 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1020.eqiad.wmnet,service=s5 * 08:02 marostegui: Depool clouddb1018 (s2,s7) [[phab:T434048|T434048]] * 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1018.eqiad.wmnet,service=s7 * 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1018.eqiad.wmnet,service=s2 * 08:01 marostegui: Depool clouddb1017 (s1) [[phab:T434048|T434048]] * 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1017.eqiad.wmnet,service=s1 * 07:59 marostegui: Depool clouddb1016 (s5,s8) [[phab:T434048|T434048]] * 07:59 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s8 * 07:59 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s5 * 07:57 marostegui: Depool clouddb1015 (s4,s6) [[phab:T434048|T434048]] * 07:57 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s6 * 07:57 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s4 * 07:56 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95910 and previous config saved to /var/cache/conftool/dbconfig/20260805-075650-ladsgroup.json * 07:54 marostegui: Depool clouddb1014 (s2,s7) [[phab:T434048|T434048]] * 07:54 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1014.eqiad.wmnet,service=s7 * 07:54 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1014.eqiad.wmnet,service=s2 * 07:53 marostegui: Depool clouddb1013:s1 [[phab:T434048|T434048]] * 07:53 marostegui: Depool clouddb1013:s1 [[phab:T409557|T409557]] * 07:53 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1013.eqiad.wmnet,service=s1 * 07:25 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95909 and previous config saved to /var/cache/conftool/dbconfig/20260805-072529-ladsgroup.json * 07:24 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2249.codfw.wmnet with reason: Maintenance * 07:24 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95908 and previous config saved to /var/cache/conftool/dbconfig/20260805-072426-ladsgroup.json * 07:21 slyngshede@dns1004: END - running authdns-update * 07:19 slyngshede@dns1004: START - running authdns-update * 07:13 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231', diff saved to https://phabricator.wikimedia.org/P95906 and previous config saved to /var/cache/conftool/dbconfig/20260805-071340-ladsgroup.json * 07:02 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231', diff saved to https://phabricator.wikimedia.org/P95905 and previous config saved to /var/cache/conftool/dbconfig/20260805-070253-ladsgroup.json * 06:52 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95904 and previous config saved to /var/cache/conftool/dbconfig/20260805-065206-ladsgroup.json * 06:45 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 06:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95903 and previous config saved to /var/cache/conftool/dbconfig/20260805-062240-ladsgroup.json * 06:21 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2231.codfw.wmnet with reason: Maintenance * 06:21 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95902 and previous config saved to /var/cache/conftool/dbconfig/20260805-062137-ladsgroup.json * 06:10 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215', diff saved to https://phabricator.wikimedia.org/P95901 and previous config saved to /var/cache/conftool/dbconfig/20260805-061051-ladsgroup.json * 06:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215', diff saved to https://phabricator.wikimedia.org/P95900 and previous config saved to /var/cache/conftool/dbconfig/20260805-060004-ladsgroup.json * 05:49 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95899 and previous config saved to /var/cache/conftool/dbconfig/20260805-054918-ladsgroup.json * 05:19 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95898 and previous config saved to /var/cache/conftool/dbconfig/20260805-051939-ladsgroup.json * 05:18 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2215.codfw.wmnet with reason: Maintenance * 04:30 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2201.codfw.wmnet with reason: Maintenance * 03:40 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2197.codfw.wmnet with reason: Maintenance * 03:40 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95897 and previous config saved to /var/cache/conftool/dbconfig/20260805-034036-ladsgroup.json * 03:29 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196', diff saved to https://phabricator.wikimedia.org/P95896 and previous config saved to /var/cache/conftool/dbconfig/20260805-032948-ladsgroup.json * 03:19 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196', diff saved to https://phabricator.wikimedia.org/P95895 and previous config saved to /var/cache/conftool/dbconfig/20260805-031902-ladsgroup.json * 03:08 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95894 and previous config saved to /var/cache/conftool/dbconfig/20260805-030815-ladsgroup.json * 02:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95893 and previous config saved to /var/cache/conftool/dbconfig/20260805-023413-ladsgroup.json * 02:33 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2196.codfw.wmnet with reason: Maintenance * 02:33 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95892 and previous config saved to /var/cache/conftool/dbconfig/20260805-023310-ladsgroup.json * 02:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186', diff saved to https://phabricator.wikimedia.org/P95891 and previous config saved to /var/cache/conftool/dbconfig/20260805-022223-ladsgroup.json * 02:11 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186', diff saved to https://phabricator.wikimedia.org/P95890 and previous config saved to /var/cache/conftool/dbconfig/20260805-021137-ladsgroup.json * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 02:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95889 and previous config saved to /var/cache/conftool/dbconfig/20260805-020051-ladsgroup.json * 01:30 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95888 and previous config saved to /var/cache/conftool/dbconfig/20260805-013029-ladsgroup.json * 01:29 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2186.codfw.wmnet with reason: Maintenance * 00:34 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on dbstore1009.eqiad.wmnet with reason: Maintenance * 00:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95887 and previous config saved to /var/cache/conftool/dbconfig/20260805-003408-ladsgroup.json * 00:23 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264', diff saved to https://phabricator.wikimedia.org/P95886 and previous config saved to /var/cache/conftool/dbconfig/20260805-002322-ladsgroup.json * 00:12 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264', diff saved to https://phabricator.wikimedia.org/P95885 and previous config saved to /var/cache/conftool/dbconfig/20260805-001235-ladsgroup.json * 00:01 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95884 and previous config saved to /var/cache/conftool/dbconfig/20260805-000148-ladsgroup.json == 2026-08-04 == * 23:45 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95883 and previous config saved to /var/cache/conftool/dbconfig/20260804-234508-ladsgroup.json * 23:44 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1264.eqiad.wmnet with reason: Maintenance * 23:44 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95882 and previous config saved to /var/cache/conftool/dbconfig/20260804-234405-ladsgroup.json * 23:33 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237', diff saved to https://phabricator.wikimedia.org/P95881 and previous config saved to /var/cache/conftool/dbconfig/20260804-233317-ladsgroup.json * 23:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237', diff saved to https://phabricator.wikimedia.org/P95880 and previous config saved to /var/cache/conftool/dbconfig/20260804-232230-ladsgroup.json * 23:11 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95879 and previous config saved to /var/cache/conftool/dbconfig/20260804-231144-ladsgroup.json * 22:23 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95878 and previous config saved to /var/cache/conftool/dbconfig/20260804-222345-ladsgroup.json * 22:23 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1237.eqiad.wmnet with reason: Maintenance * 21:13 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1225.eqiad.wmnet with reason: Maintenance * 20:40 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] (duration: 24m 40s) * 20:33 samtar@deploy1003: samtar, kineticpelagic: Continuing with deployment * 20:28 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS bookworm * 20:21 samtar@deploy1003: samtar, kineticpelagic: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:15 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] * 20:13 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 20:09 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 20:00 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1216.eqiad.wmnet with reason: Maintenance * 20:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95877 and previous config saved to /var/cache/conftool/dbconfig/20260804-195957-ladsgroup.json * 19:51 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 19:50 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) mc2046.codfw.wmnet 120.16.192.10.in-addr.arpa 0.2.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:50 jhancock@cumin2002: START - Cookbook sre.dns.wipe-cache mc2046.codfw.wmnet 120.16.192.10.in-addr.arpa 0.2.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host mc2046 - jhancock@cumin2002" * 19:50 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host mc2046 - jhancock@cumin2002" * 19:49 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203', diff saved to https://phabricator.wikimedia.org/P95876 and previous config saved to /var/cache/conftool/dbconfig/20260804-194911-ladsgroup.json * 19:46 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 19:45 jhancock@cumin2002: START - Cookbook sre.hosts.move-vlan for host mc2046 * 19:45 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS bookworm * 19:38 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203', diff saved to https://phabricator.wikimedia.org/P95875 and previous config saved to /var/cache/conftool/dbconfig/20260804-193825-ladsgroup.json * 19:27 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95874 and previous config saved to /var/cache/conftool/dbconfig/20260804-192738-ladsgroup.json * 19:02 mutante: gerrit ssh -p 29418 gerrit.wikimedia.org gerrit index changes {{Gerrit|1320979}} * 18:20 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 18:18 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 18:14 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 18:14 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 18:13 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 18:10 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 18:08 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 18:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95872 and previous config saved to /var/cache/conftool/dbconfig/20260804-180721-ladsgroup.json * 18:07 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 18:06 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1203.eqiad.wmnet with reason: Maintenance * 18:06 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95871 and previous config saved to /var/cache/conftool/dbconfig/20260804-180618-ladsgroup.json * 17:55 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179', diff saved to https://phabricator.wikimedia.org/P95870 and previous config saved to /var/cache/conftool/dbconfig/20260804-175531-ladsgroup.json * 17:55 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1154.eqiad.wmnet * 17:55 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1154.eqiad.wmnet * 17:55 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1154.eqiad.wmnet * 17:50 swfrench@deploy1003: Finished scap sync-world: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] (duration: 04m 05s) * 17:48 swfrench@deploy1003: swfrench: Continuing with deployment * 17:46 swfrench@deploy1003: swfrench: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:45 swfrench@deploy1003: Started scap sync-world: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] * 17:44 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179', diff saved to https://phabricator.wikimedia.org/P95869 and previous config saved to /var/cache/conftool/dbconfig/20260804-174445-ladsgroup.json * 17:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95868 and previous config saved to /var/cache/conftool/dbconfig/20260804-173359-ladsgroup.json * 17:33 swfrench@deploy1003: Finished scap sync-world: Pick up new PHP production image (duration: 28m 32s) * 17:28 aokoth@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on phab1005.eqiad.wmnet with reason: Puppet Failure * 17:05 swfrench@deploy1003: Started scap sync-world: Pick up new PHP production image * 17:00 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 17:00 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 16:54 cgoubert@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on wikikube-worker2187.codfw.wmnet with reason: Hardware issue * 16:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2187.codfw.wmnet * 16:52 mutante: gerrit2003:/var/log/apache2# ln -s /srv/gerrit/site_path/review_site/logs/ gerrit * 16:52 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2187.codfw.wmnet * 16:48 mutante: gerrit2003 - moving old apache logfiles older than 60 days from /var/log/apache2 to /srv/gerrit/site_path/review_site/logs/old/ * 16:33 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 16:32 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 16:29 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 16:29 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 16:28 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 16:28 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 16:27 dzahn@cumin1003: END (PASS) - Cookbook sre.gerrit.restart-gerrit (exit_code=0) Restarting Gerrit on gerrit2003 * 16:27 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 16:27 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95867 and previous config saved to /var/cache/conftool/dbconfig/20260804-162736-ladsgroup.json * 16:27 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 16:26 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1179.eqiad.wmnet with reason: Maintenance * 16:26 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:25 mutante: restarting gerrit - dropped outdated RSA host key * 16:25 dzahn@cumin1003: START - Cookbook sre.gerrit.restart-gerrit Restarting Gerrit on gerrit2003 * 16:24 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95866 and previous config saved to /var/cache/conftool/dbconfig/20260804-162424-ladsgroup.json * 16:24 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 16:23 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 16:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95865 and previous config saved to /var/cache/conftool/dbconfig/20260804-162236-ladsgroup.json * 16:21 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1179.eqiad.wmnet with reason: Maintenance * 16:17 swfrench-wmf: reprepro include php8.3_8.3.33-1+wmf11u1 into component/php83 for bullseye-wikimedia * 16:17 swfrench-wmf: reprepro include php8.3_8.3.33-1+wmf12u1 into component/php83 for bookworm-wikimedia * 16:11 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply * 16:10 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply * 16:10 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mobileapps: apply * 16:09 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mobileapps: apply * 16:09 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply * 16:08 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply * 16:08 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:08 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:07 aokoth@cumin1003: END (PASS) - Cookbook sre.vrts.upgrade (exit_code=0) on VRTS host vrts1003.eqiad.wmnet * 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:05 aokoth@cumin1003: START - Cookbook sre.vrts.upgrade on VRTS host vrts1003.eqiad.wmnet * 16:04 mutante: gerrit2002/gerrit1003/gerrit2003 - rm /etc/gerrit/ssh_host_rsa_key * 15:59 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:59 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:59 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:59 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:56 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 15:55 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:55 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:55 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:49 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 15:49 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:44 Raine: add php8.5 packages to component/php85 - [[phab:T432983|T432983]] * 15:39 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:33 aaron@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 15:33 aaron@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 15:29 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:19 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:19 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:16 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:16 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1154.eqiad.wmnet with OS trixie * 15:16 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:15 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:15 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:06 brennen@deploy1003: Finished deploy [phabricator/deployment@56f4ffd]: deploy phab1004 for [[phab:T433981|T433981]] (duration: 00m 43s) * 15:05 brennen@deploy1003: Started deploy [phabricator/deployment@56f4ffd]: deploy phab1004 for [[phab:T433981|T433981]] * 15:05 aaron@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 15:04 aaron@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 15:02 brennen@deploy1003: Finished deploy [phabricator/deployment@56f4ffd]: deploy phab2003 for [[phab:T433981|T433981]] (duration: 00m 51s) * 15:01 brennen@deploy1003: Started deploy [phabricator/deployment@56f4ffd]: deploy phab2003 for [[phab:T433981|T433981]] * 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1004.eqiad.wmnet with reason: deployment * 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1005.eqiad.wmnet with reason: deployment * 14:58 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab2003.codfw.wmnet with reason: deployment * 14:55 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1154.eqiad.wmnet with reason: host reimage * 14:51 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1154.eqiad.wmnet with reason: host reimage * 14:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 14:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 14:38 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync * 14:38 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync * 14:38 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync * 14:37 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync * 14:37 ottomata: roll restart eventgate-main to pick up stream config change - [[phab:T433507|T433507]] * 14:37 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-main: sync * 14:36 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-main: sync * 14:36 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1154 * 14:36 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1154 * 14:34 otto@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] (duration: 08m 39s) * 14:34 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1154 * 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1154.eqiad.wmnet 108.32.64.10.in-addr.arpa 8.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:34 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1154.eqiad.wmnet 108.32.64.10.in-addr.arpa 8.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1154 - jayme@cumin1003" * 14:34 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1154 - jayme@cumin1003" * 14:30 otto@deploy1003: otto: Continuing with deployment * 14:30 jayme@cumin1003: START - Cookbook sre.dns.netbox * 14:28 otto@deploy1003: otto: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:26 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1154 * 14:26 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1154.eqiad.wmnet with OS trixie * 14:26 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1154.eqiad.wmnet * 14:26 otto@deploy1003: Started scap sync-world: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] * 14:26 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1154.eqiad.wmnet * 14:26 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1154.eqiad.wmnet * 14:17 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 14:16 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 14:15 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 14:14 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 14:13 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 14:13 swfrench@dns1004: END - running authdns-update * 14:13 Msz2001: Finished deployments for UTC afternoon backport window * 14:13 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 14:13 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] (duration: 07m 58s) * 14:11 swfrench@dns1004: START - running authdns-update * 14:08 mszwarc@deploy1003: javiermonton, mszwarc, mpostoronca: Continuing with deployment * 14:07 mszwarc@deploy1003: javiermonton, mszwarc, mpostoronca: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] synced to the testser * 14:05 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] * 14:03 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 13:49 swfrench@cumin2002: conftool action : set/pooled=yes; selector: name=wikikube-worker2330.codfw.wmnet * 13:49 swfrench@cumin2002: conftool action : set/pooled=no; selector: name=wikikube-worker2330.codfw.wmnet * 13:48 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] (duration: 09m 19s) * 13:45 swfrench@dns1004: END - running authdns-update * 13:44 mszwarc@deploy1003: mszwarc, jforrester: Continuing with deployment * 13:43 swfrench@dns1004: START - running authdns-update * 13:41 mszwarc@deploy1003: mszwarc, jforrester: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:38 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] * 13:33 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 13:33 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1154.eqiad.wmnet * 13:32 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 13:32 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 13:31 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 13:31 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:31 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:29 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1154.eqiad.wmnet * 13:28 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1154.eqiad.wmnet * 13:28 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1154.eqiad.wmnet * 13:28 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1141.eqiad.wmnet * 13:28 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1141.eqiad.wmnet * 13:28 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1141.eqiad.wmnet * 13:22 otto@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply * 13:22 otto@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply * 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1096.eqiad.wmnet with OS trixie * 13:05 swfrench@dns1004: END - running authdns-update * 13:03 swfrench@dns1004: START - running authdns-update * 12:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 12:43 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 1:00:00 on db1171.eqiad.wmnet with reason: decom * 12:42 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 1:00:00 on db1150.eqiad.wmnet with reason: decom * 12:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 12:38 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1164,1217].eqiad.wmnet with reason: cloning * 12:33 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2096.codfw.wmnet with OS trixie * 12:22 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1096.eqiad.wmnet with OS trixie * 12:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2096.codfw.wmnet with reason: host reimage * 12:14 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1141.eqiad.wmnet with OS trixie * 12:10 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2096.codfw.wmnet with reason: host reimage * 12:10 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1289.eqiad.wmnet * 12:05 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1289.eqiad.wmnet * 12:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1288.eqiad.wmnet * 11:59 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1288.eqiad.wmnet * 11:59 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1287.eqiad.wmnet * 11:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1097.eqiad.wmnet with OS trixie * 11:54 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1287.eqiad.wmnet * 11:54 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1286.eqiad.wmnet * 11:53 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1141.eqiad.wmnet with reason: host reimage * 11:51 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2096.codfw.wmnet with OS trixie * 11:49 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1141.eqiad.wmnet with reason: host reimage * 11:48 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1286.eqiad.wmnet * 11:48 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1284.eqiad.wmnet * 11:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2095.codfw.wmnet with OS trixie * 11:43 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1284.eqiad.wmnet * 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1283.eqiad.wmnet * 11:42 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on ml-serve1015.eqiad.wmnet with reason: Downtime to get full picture of current BIOS settings beyond what Redfish shows * 11:39 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad * 11:39 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:37 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1283.eqiad.wmnet * 11:37 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1282.eqiad.wmnet * 11:37 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad * 11:37 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:33 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1141 * 11:33 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1141 * 11:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 11:32 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1141 * 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1141.eqiad.wmnet 156.48.64.10.in-addr.arpa 6.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:32 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1141.eqiad.wmnet 156.48.64.10.in-addr.arpa 6.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1141 - jayme@cumin1003" * 11:32 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1141 - jayme@cumin1003" * 11:32 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1282.eqiad.wmnet * 11:32 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1281.eqiad.wmnet * 11:32 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad * 11:32 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:29 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 11:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1097.eqiad.wmnet with reason: host reimage * 11:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2095.codfw.wmnet with OS trixie * 11:26 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1281.eqiad.wmnet * 11:26 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1280.eqiad.wmnet * 11:25 jayme@cumin1003: START - Cookbook sre.dns.netbox * 11:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1097.eqiad.wmnet with reason: host reimage * 11:22 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1141 * 11:21 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1141.eqiad.wmnet with OS trixie * 11:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1280.eqiad.wmnet * 11:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1279.eqiad.wmnet * 11:20 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin with reason: upgrade new Nokia swtiches in eqsin to SR Linux v26 * 11:17 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1141.eqiad.wmnet * 11:16 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1141.eqiad.wmnet * 11:16 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1141.eqiad.wmnet * 11:16 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be2095.codfw.wmnet with OS trixie * 11:15 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1279.eqiad.wmnet * 11:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1278.eqiad.wmnet * 11:14 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1139.eqiad.wmnet * 11:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1139.eqiad.wmnet * 11:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1070.eqiad.wmnet with OS trixie * 11:13 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1140.eqiad.wmnet * 11:13 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1140.eqiad.wmnet * 11:13 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1140.eqiad.wmnet * 11:09 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1278.eqiad.wmnet * 11:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1071.eqiad.wmnet with OS trixie * 11:05 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1097.eqiad.wmnet with OS trixie * 11:04 mvernon@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be1097.eqiad.wmnet with OS trixie * 11:02 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1140.eqiad.wmnet with OS trixie * 11:02 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1097.eqiad.wmnet with OS trixie * 11:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1096.eqiad.wmnet with OS trixie * 11:00 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1139.eqiad.wmnet * 11:00 jayme@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1139.eqiad.wmnet with OS trixie * 10:56 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1069.eqiad.wmnet with OS trixie * 10:56 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 10:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1070.eqiad.wmnet with reason: host reimage * 10:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 10:45 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1071.eqiad.wmnet with reason: host reimage * 10:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 10:41 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1096.eqiad.wmnet with OS trixie * 10:41 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1140.eqiad.wmnet with reason: host reimage * 10:39 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1071.eqiad.wmnet with reason: host reimage * 10:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1070.eqiad.wmnet with reason: host reimage * 10:38 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1139.eqiad.wmnet with reason: host reimage * 10:37 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1140.eqiad.wmnet with reason: host reimage * 10:35 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1069.eqiad.wmnet with reason: host reimage * 10:33 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1095.eqiad.wmnet with OS trixie * 10:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 10:33 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1139.eqiad.wmnet with reason: host reimage * 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1069.eqiad.wmnet with reason: host reimage * 10:23 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1140 * 10:23 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1140 * 10:23 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 10:22 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1140 * 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1140.eqiad.wmnet 155.48.64.10.in-addr.arpa 5.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:21 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1071.eqiad.wmnet with OS trixie * 10:21 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1140.eqiad.wmnet 155.48.64.10.in-addr.arpa 5.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1140 - jayme@cumin1003" * 10:21 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1140 - jayme@cumin1003" * 10:21 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1071 * 10:21 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1070.eqiad.wmnet with OS trixie * 10:21 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1070 * 10:20 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 10:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1095.eqiad.wmnet with OS trixie * 10:17 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1139 * 10:17 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1139 * 10:17 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be1095.eqiad.wmnet with OS trixie * 10:15 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1139 * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1139.eqiad.wmnet 194.32.64.10.in-addr.arpa 4.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:15 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1139.eqiad.wmnet 194.32.64.10.in-addr.arpa 4.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1139 - jayme@cumin1003" * 10:15 jayme@cumin1003: START - Cookbook sre.dns.netbox * 10:15 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1139 - jayme@cumin1003" * 10:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2095.codfw.wmnet with OS trixie * 10:13 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1069.eqiad.wmnet with OS trixie * 10:12 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1069 * 10:11 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1140 * 10:11 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1140.eqiad.wmnet with OS trixie * 10:11 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1140.eqiad.wmnet * 10:10 jayme@cumin1003: START - Cookbook sre.dns.netbox * 10:10 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1139 * 10:10 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1140.eqiad.wmnet * 10:10 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1140.eqiad.wmnet * 10:10 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1139.eqiad.wmnet with OS trixie * 10:09 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1139.eqiad.wmnet * 10:08 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1139.eqiad.wmnet * 10:08 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1139.eqiad.wmnet * 10:01 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2094.codfw.wmnet with OS trixie * 09:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 09:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 09:53 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 09:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 09:44 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:44 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2094.codfw.wmnet with reason: host reimage * 09:34 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2094.codfw.wmnet with reason: host reimage * 09:34 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1071 * 09:33 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1070 * 09:33 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1095.eqiad.wmnet with OS trixie * 09:27 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1069 * 09:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:23 brouberol@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM archiva1002.wikimedia.org * 09:20 brouberol@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM archiva1002.wikimedia.org * 09:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1277.eqiad.wmnet * 09:13 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2094.codfw.wmnet with OS trixie * 09:13 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 09:12 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 09:12 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:12 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1277.eqiad.wmnet * 09:12 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1276.eqiad.wmnet * 09:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1094.eqiad.wmnet with OS trixie * 09:06 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1276.eqiad.wmnet * 09:06 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1275.eqiad.wmnet * 09:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2093.codfw.wmnet with OS trixie * 09:01 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1275.eqiad.wmnet * 09:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1274.eqiad.wmnet * 08:56 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1274.eqiad.wmnet * 08:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1273.eqiad.wmnet * 08:50 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1273.eqiad.wmnet * 08:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1272.eqiad.wmnet * 08:49 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:49 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1094.eqiad.wmnet with reason: host reimage * 08:45 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1272.eqiad.wmnet * 08:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1094.eqiad.wmnet with reason: host reimage * 08:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2093.codfw.wmnet with reason: host reimage * 08:38 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:38 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2093.codfw.wmnet with reason: host reimage * 08:35 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 08:34 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 08:29 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:28 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:26 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1271.eqiad.wmnet * 08:23 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1094.eqiad.wmnet with OS trixie * 08:21 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 08:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1271.eqiad.wmnet * 08:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1270.eqiad.wmnet * 08:15 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1270.eqiad.wmnet * 08:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1269.eqiad.wmnet * 08:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2093.codfw.wmnet with OS trixie * 08:09 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1269.eqiad.wmnet * 08:09 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1268.eqiad.wmnet * 08:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2092.codfw.wmnet with OS trixie * 08:04 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1268.eqiad.wmnet * 08:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1267.eqiad.wmnet * 07:59 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1267.eqiad.wmnet * 07:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1093.eqiad.wmnet with OS trixie * 07:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1266.eqiad.wmnet * 07:56 jynus: running extra backups to test db1285 [[phab:T433826|T433826]] * 07:51 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1266.eqiad.wmnet * 07:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2092.codfw.wmnet with reason: host reimage * 07:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1093.eqiad.wmnet with reason: host reimage * 07:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2092.codfw.wmnet with reason: host reimage * 07:32 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1093.eqiad.wmnet with reason: host reimage * 07:29 jynus: running extra backups to test db1265 [[phab:T433825|T433825]] * 07:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2092.codfw.wmnet with OS trixie * 07:11 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1093.eqiad.wmnet with OS trixie * 06:50 slyngshede@dns1004: END - running authdns-update * 06:48 slyngshede@dns1004: START - running authdns-update * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.11 (duration: 02m 29s) * 03:38 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] (duration: 32m 57s) * 03:23 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 03:22 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 03:05 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 32s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 00:45 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] (duration: 06m 20s) * 00:41 cjming@deploy1003: cjming: Continuing with deployment * 00:41 cjming@deploy1003: cjming: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:39 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] == 2026-08-03 == * 23:58 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply * 23:57 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply * 23:29 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cp5021.eqsin.wmnet * 23:29 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cp5021.eqsin.wmnet * 23:27 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cp5021.eqsin.wmnet * 23:26 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cp5021.eqsin.wmnet * 23:18 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 23:17 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 22:56 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: sync * 22:56 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: sync * 22:36 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 22:36 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 22:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host search-loader2002.codfw.wmnet with OS trixie * 21:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on search-loader2002.codfw.wmnet with reason: host reimage * 21:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on search-loader2002.codfw.wmnet with reason: host reimage * 21:42 dancy@deploy1003: Stopping before sync operations * 21:41 dancy@deploy1003: Started scap sync-world: testing * 21:39 dancy@deploy1003: Installation of scap version "4.277.0" completed for 3 hosts * 21:37 dancy@deploy1003: Installing scap version "4.277.0" for 3 host(s) * 21:37 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] (duration: 06m 13s) * 21:33 dancy@deploy1003: dancy: Continuing with deployment * 21:32 dancy@deploy1003: dancy: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:31 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] * 21:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host search-loader2002.codfw.wmnet with OS trixie * 21:03 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] (duration: 06m 34s) * 20:59 dancy@deploy1003: dancy: Continuing with deployment * 20:58 dancy@deploy1003: dancy: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:56 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] * 20:52 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] (duration: 06m 23s) * 20:48 cjming@deploy1003: cjming: Continuing with deployment * 20:47 cjming@deploy1003: cjming: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:46 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] * 20:42 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] (duration: 07m 36s) * 20:38 arlolra@deploy1003: arlolra: Continuing with deployment * 20:36 arlolra@deploy1003: arlolra: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:34 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] * 20:16 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] (duration: 08m 26s) * 20:12 krinkle@deploy1003: krinkle: Continuing with deployment * 20:09 krinkle@deploy1003: krinkle: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] * 19:45 jasmine@cumin2002: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-main-codfw * 18:58 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] (duration: 09m 23s) * 18:53 krinkle@deploy1003: krinkle: Continuing with deployment * 18:53 jasmine@cumin2002: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-main-codfw * 18:50 krinkle@deploy1003: krinkle: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:48 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] * 18:37 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] (duration: 10m 13s) * 18:34 dzahn@cumin2002: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host codesearch2001.codfw.wmnet * 18:34 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host codesearch2001.codfw.wmnet with OS trixie * 18:33 krinkle@deploy1003: krinkle: Continuing with deployment * 18:29 krinkle@deploy1003: krinkle: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:27 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] * 18:19 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 18:18 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on codesearch2001.codfw.wmnet with reason: host reimage * 18:14 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 18:14 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:12 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on codesearch2001.codfw.wmnet with reason: host reimage * 18:11 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 18:11 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 18:10 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 18:02 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 18:02 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 18:01 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 18:01 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 17:55 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host codesearch2001.codfw.wmnet with OS trixie * 17:54 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:54 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) codesearch2001.codfw.wmnet on all recursors * 17:53 dzahn@cumin2002: START - Cookbook sre.dns.wipe-cache codesearch2001.codfw.wmnet on all recursors * 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:48 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:41 dzahn@cumin2002: START - Cookbook sre.dns.netbox * 17:41 dzahn@cumin2002: START - Cookbook sre.ganeti.makevm for new host codesearch2001.codfw.wmnet * 17:37 dzahn@cumin2002: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host codesearch1001.eqiad.wmnet * 17:37 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host codesearch1001.eqiad.wmnet with OS trixie * 17:24 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on codesearch1001.eqiad.wmnet with reason: host reimage * 17:17 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on codesearch1001.eqiad.wmnet with reason: host reimage * 17:08 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host codesearch1001.eqiad.wmnet with OS trixie * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 17:06 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:06 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) codesearch1001.eqiad.wmnet on all recursors * 17:06 dzahn@cumin2002: START - Cookbook sre.dns.wipe-cache codesearch1001.eqiad.wmnet on all recursors * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 17:05 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 17:04 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:04 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 16:58 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 16:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2091.codfw.wmnet with OS trixie * 16:54 ebernhardson@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 16:54 ebernhardson@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 16:49 ebernhardson@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 16:49 ebernhardson@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 16:46 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1092.eqiad.wmnet with OS trixie * 16:43 ebernhardson@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 16:43 ebernhardson@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 16:43 dzahn@cumin2002: START - Cookbook sre.dns.netbox * 16:43 dzahn@cumin2002: START - Cookbook sre.ganeti.makevm for new host codesearch1001.eqiad.wmnet * 16:41 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 16:41 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 16:40 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2091.codfw.wmnet with reason: host reimage * 16:37 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 16:35 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2091.codfw.wmnet with reason: host reimage * 16:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1092.eqiad.wmnet with reason: host reimage * 16:24 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 16:24 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 16:23 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1092.eqiad.wmnet with reason: host reimage * 16:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2091.codfw.wmnet with OS trixie * 16:03 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1092.eqiad.wmnet with OS trixie * 16:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2090.codfw.wmnet with OS trixie * 15:51 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 15:51 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 15:51 jiji@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 15:50 jiji@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 15:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2090.codfw.wmnet with reason: host reimage * 15:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2090.codfw.wmnet with reason: host reimage * 15:33 jhathaway@dns1004: END - running authdns-update * 15:31 jhathaway@dns1004: START - running authdns-update * 15:26 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1091.eqiad.wmnet with OS trixie * 15:25 dancy@deploy1003: Installation of scap version "4.276.1" completed for 3 hosts * 15:23 dancy@deploy1003: Installing scap version "4.276.1" for 3 host(s) * 15:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2090.codfw.wmnet with OS trixie * 15:12 marostegui@cumin1003: dbctl commit (dc=all): 'Repool db2245, db2246, db2247 and db2248 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95857 and previous config saved to /var/cache/conftool/dbconfig/20260803-151212-marostegui.json * 15:09 dancy@deploy1003: Started scap sync-world: testing * 15:09 dancy@deploy1003: Installation of scap version "4.277.0" completed for 3 hosts * 15:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1091.eqiad.wmnet with reason: host reimage * 15:07 dancy@deploy1003: Installing scap version "4.277.0" for 3 host(s) * 15:03 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1091.eqiad.wmnet with reason: host reimage * 14:49 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1091.eqiad.wmnet with OS trixie * 14:33 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2089.codfw.wmnet with OS trixie * 14:29 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 14:27 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 14:18 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 14:16 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 14:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2089.codfw.wmnet with reason: host reimage * 14:10 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 14:10 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 14:09 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2089.codfw.wmnet with reason: host reimage * 13:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2089.codfw.wmnet with OS trixie * 13:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2088.codfw.wmnet with OS trixie * 13:40 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1090.eqiad.wmnet with OS trixie * 13:22 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1090.eqiad.wmnet with reason: host reimage * 13:22 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] (duration: 14m 34s) * 13:19 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1090.eqiad.wmnet with reason: host reimage * 13:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2088.codfw.wmnet with reason: host reimage * 13:16 aude@deploy1003: aude, mhorsey: Continuing with deployment * 13:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2088.codfw.wmnet with reason: host reimage * 13:12 aude@deploy1003: aude, mhorsey: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] * 13:05 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1090.eqiad.wmnet with OS trixie * 12:58 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2088.codfw.wmnet with OS trixie * 12:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db[2245-2247].codfw.wmnet * 12:49 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2247: Rebooting db2247.codfw.wmnet * 12:49 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2247: Rebooting db2247.codfw.wmnet * 12:42 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2246: Rebooting db2246.codfw.wmnet * 12:42 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2246: Rebooting db2246.codfw.wmnet * 12:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2087.codfw.wmnet with OS trixie * 12:37 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1089.eqiad.wmnet with OS trixie * 12:34 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2245: Rebooting db2245.codfw.wmnet * 12:34 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2245: Rebooting db2245.codfw.wmnet * 12:34 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db[2245-2247].codfw.wmnet * 12:32 kamila@deploy1003: Finished scap sync-world: rebuild after base image update (duration: 30m 26s) * 12:28 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 12:22 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2087.codfw.wmnet with reason: host reimage * 12:19 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1089.eqiad.wmnet with reason: host reimage * 12:14 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2087.codfw.wmnet with reason: host reimage * 12:14 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1089.eqiad.wmnet with reason: host reimage * 12:03 kamila@deploy1003: Started scap sync-world: rebuild after base image update * 12:00 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1089.eqiad.wmnet with OS trixie * 12:00 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2087.codfw.wmnet with OS trixie * 11:35 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db[2245-2248].codfw.wmnet * 11:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db[2245-2248].codfw.wmnet * 11:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2086.codfw.wmnet with OS trixie * 11:26 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 11:26 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 11:25 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 11:25 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 11:24 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1088.eqiad.wmnet with OS trixie * 11:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db[2245-2248].codfw.wmnet with reason: Checking network * 11:21 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 11:20 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 11:19 marostegui@dns1004: END - running authdns-update * 11:17 marostegui@dns1004: START - running authdns-update * 11:10 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:10 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 11:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2086.codfw.wmnet with reason: host reimage * 11:09 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:08 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 11:08 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:07 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 11:07 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:07 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 11:06 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop: apply * 11:06 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop: apply * 11:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1088.eqiad.wmnet with reason: host reimage * 11:05 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop: apply * 11:04 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop: apply * 11:04 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop: apply * 11:04 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop: apply * 11:02 marostegui@dns1004: END - running authdns-update * 11:02 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2086.codfw.wmnet with reason: host reimage * 11:01 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1088.eqiad.wmnet with reason: host reimage * 11:00 marostegui@dns1004: START - running authdns-update * 10:53 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] (duration: 10m 57s) * 10:51 cmooney@dns3003: END - running authdns-update * 10:49 cmooney@dns3003: START - running authdns-update * 10:47 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1088.eqiad.wmnet with OS trixie * 10:47 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2086.codfw.wmnet with OS trixie * 10:47 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 10:46 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:46 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:46 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new reverse ranges for eqsin CR switch links - cmooney@cumin1003" * 10:46 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new reverse ranges for eqsin CR switch links - cmooney@cumin1003" * 10:42 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] * 10:41 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 10:36 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2245, db2246 and db2247 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95855 and previous config saved to /var/cache/conftool/dbconfig/20260803-103652-marostegui.json * 10:35 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2248 from s4 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95854 and previous config saved to /var/cache/conftool/dbconfig/20260803-103535-marostegui.json * 10:27 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 10:27 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 10:26 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 10:24 kart_: cxserver: Add referencePunctuation config ([[phab:T97231|T97231]]) * 10:24 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 10:23 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:23 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:23 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:22 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:22 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply * 10:21 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply * 10:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2085.codfw.wmnet with OS trixie * 10:20 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply * 10:20 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply * 10:18 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply * 10:18 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply * 10:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1087.eqiad.wmnet with OS trixie * 09:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2085.codfw.wmnet with reason: host reimage * 09:43 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1087.eqiad.wmnet with reason: host reimage * 09:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2085.codfw.wmnet with reason: host reimage * 09:40 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1087.eqiad.wmnet with reason: host reimage * 09:26 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1087.eqiad.wmnet with OS trixie * 09:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2085.codfw.wmnet with OS trixie * 09:13 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2084.codfw.wmnet with OS trixie * 09:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1086.eqiad.wmnet with OS trixie * 08:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2084.codfw.wmnet with reason: host reimage * 08:50 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2084.codfw.wmnet with reason: host reimage * 08:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1086.eqiad.wmnet with reason: host reimage * 08:39 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1086.eqiad.wmnet with reason: host reimage * 08:38 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:38 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:37 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 08:37 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 08:35 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2084.codfw.wmnet with OS trixie * 08:34 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 08:34 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:27 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1086.eqiad.wmnet with OS trixie * 08:09 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1218: Repool after a crash * 08:07 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2083.codfw.wmnet with OS trixie * 08:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1085.eqiad.wmnet with OS trixie * 07:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2083.codfw.wmnet with reason: host reimage * 07:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1085.eqiad.wmnet with reason: host reimage * 07:40 kart_: Updated cxsever to 2026-07-16-140518-production ([[phab:T97231|T97231]]) * 07:39 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply * 07:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2083.codfw.wmnet with reason: host reimage * 07:38 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply * 07:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1085.eqiad.wmnet with reason: host reimage * 07:37 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] (duration: 32m 40s) * 07:33 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply * 07:33 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply * 07:25 jdlrobson@deploy1003: jdlrobson: Continuing with deployment * 07:24 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2083.codfw.wmnet with OS trixie * 07:24 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1085.eqiad.wmnet with OS trixie * 07:23 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1218: Repool after a crash * 07:21 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:09 marostegui: Drop renamed tables [[phab:T425074|T425074]] * 07:04 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] * 06:55 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply * 06:54 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 46s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-02 == * 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 01m 03s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-01 == * 03:30 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:30 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:30 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:30 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 34s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-31 == * 17:41 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 17:41 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 17:40 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 17:40 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 15:33 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2195: Testing * 15:02 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:02 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 15:02 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 14:48 pt1979@cumin2002: START - Cookbook sre.dns.netbox * 14:47 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2195: Testing * 14:22 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 14:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2195: Testing * 14:21 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 14:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2195.codfw.wmnet with reason: Testing * 14:16 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 14:04 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 14:04 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 13:30 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1048.eqiad.wmnet with OS trixie * 13:22 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2195: Testing * 13:22 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 13:19 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2195: Testing * 13:18 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 13:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2195: Testing * 13:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 13:05 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 13:05 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 13:04 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 13:04 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 12:50 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lswtest-d8-eqiad * 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:53 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:42 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 6515 * 11:37 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 6515 * 11:28 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:27 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:07 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2082.codfw.wmnet with OS trixie * 10:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2082.codfw.wmnet with reason: host reimage * 10:42 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2082.codfw.wmnet with reason: host reimage * 10:28 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2082.codfw.wmnet with OS trixie * 10:02 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 09:52 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 09:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts * 09:16 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts * 08:57 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 08:46 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:42 gkyziridis@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 08:37 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 08:37 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 08:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 08:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 08:11 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:11 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:08 filippo@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudvirt1048 * 08:07 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 08:07 filippo@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudvirt1048 * 08:06 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 08:01 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 08:00 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 07:19 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1048.eqiad.wmnet with reason: host reimage * 07:13 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1048.eqiad.wmnet with reason: host reimage * 07:11 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 07:11 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 07:09 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 07:09 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 06:57 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:56 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:48 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:48 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:44 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie * 06:34 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1048.eqiad.wmnet with OS trixie * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 54s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 00:57 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] (duration: 11m 04s) * 00:53 dreamyjazz@deploy1003: dreamyjazz, jforrester: Continuing with deployment * 00:48 dreamyjazz@deploy1003: dreamyjazz, jforrester: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:46 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] == 2026-07-30 == * 21:37 dancy@deploy1003: Installation of scap version "4.276.1" completed for 3 hosts * 21:35 dancy@deploy1003: Installing scap version "4.276.1" for 3 host(s) * 21:24 dancy@deploy1003: Installation of scap version "4.276.0" completed for 3 hosts * 21:22 dancy@deploy1003: Installing scap version "4.276.0" for 3 host(s) * 21:15 maryum: Deployed security fix for [[phab:T430601|T430601]] * 20:13 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] (duration: 09m 20s) * 20:07 arlolra@deploy1003: osleger, arlolra: Continuing with deployment * 20:05 arlolra@deploy1003: osleger, arlolra: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:03 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] * 19:29 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:29 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:25 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service * 19:24 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 19:24 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:24 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:24 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 19:23 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service * 19:20 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1084.eqiad.wmnet with OS trixie * 18:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1084.eqiad.wmnet with reason: host reimage * 18:52 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1084.eqiad.wmnet with reason: host reimage * 18:41 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 18:40 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 18:39 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1084.eqiad.wmnet with OS trixie * 18:25 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 18:15 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 17:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1083.eqiad.wmnet with OS trixie * 17:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1048: Maintenance * 17:36 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new security plugin settings - bking@cumin2003 - [[phab:T350516|T350516]] * 17:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1083.eqiad.wmnet with reason: host reimage * 17:26 inflatador: bking@apt1002 `reprepro --noskipold --component thirdparty/opensearch3 update trixie-wikimedia` [[phab:T433624|T433624]] * 17:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1083.eqiad.wmnet with reason: host reimage * 17:23 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 17:20 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 17:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2081.codfw.wmnet with OS trixie * 17:11 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new security plugin settings - bking@cumin2003 - [[phab:T350516|T350516]] * 17:10 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1083.eqiad.wmnet with OS trixie * 16:55 root@cumin1003: START - Cookbook sre.mysql.pool pool es1048: Maintenance * 16:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2081.codfw.wmnet with reason: host reimage * 16:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1048 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95833 and previous config saved to /var/cache/conftool/dbconfig/20260730-165053-cwilliams.json * 16:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1048.eqiad.wmnet with reason: Maintenance * 16:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1040: Maintenance * 16:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2081.codfw.wmnet with reason: host reimage * 16:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1082.eqiad.wmnet with OS trixie * 16:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2081.codfw.wmnet with OS trixie * 16:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2097.codfw.wmnet with OS trixie * 16:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1082.eqiad.wmnet with reason: host reimage * 16:08 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1082.eqiad.wmnet with reason: host reimage * 16:08 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 16:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1047: Maintenance * 16:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2080.codfw.wmnet with OS trixie * 16:04 root@cumin1003: START - Cookbook sre.mysql.pool pool es1040: Maintenance * 16:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1040: Maintenance * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new logging settings - bking@cumin2003 - [[phab:T324335|T324335]] * 15:58 root@cumin1003: START - Cookbook sre.mysql.pool pool es1040: Maintenance * 15:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1040 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95827 and previous config saved to /var/cache/conftool/dbconfig/20260730-155324-cwilliams.json * 15:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1040.eqiad.wmnet with reason: Maintenance * 15:50 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1082.eqiad.wmnet with OS trixie * 15:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2048: Maintenance * 15:44 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 24s) * 15:43 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2080.codfw.wmnet with reason: host reimage * 15:38 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new logging settings - bking@cumin2003 - [[phab:T324335|T324335]] * 15:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2080.codfw.wmnet with reason: host reimage * 15:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 15:30 mvernon@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be2097.codfw.wmnet with OS trixie * 15:23 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1081.eqiad.wmnet with OS trixie * 15:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2098.codfw.wmnet with OS trixie * 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - mvernon@cumin2003" * 15:18 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be2097.codfw.wmnet with OS trixie * 15:18 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - mvernon@cumin2003" * 15:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 15:17 root@cumin1003: START - Cookbook sre.mysql.pool pool es1047: Maintenance * 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2080.codfw.wmnet with OS trixie * 15:13 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be2097.codfw.wmnet with OS trixie * 15:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1047 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95820 and previous config saved to /var/cache/conftool/dbconfig/20260730-151200-cwilliams.json * 15:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1047.eqiad.wmnet with reason: Maintenance * 15:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: Maintenance * 15:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 15:04 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 15:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1081.eqiad.wmnet with reason: host reimage * 15:00 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1081.eqiad.wmnet with reason: host reimage * 15:00 root@cumin1003: START - Cookbook sre.mysql.pool pool es2048: Maintenance * 15:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 14:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2079.codfw.wmnet with OS trixie * 14:56 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 14:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2048 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95816 and previous config saved to /var/cache/conftool/dbconfig/20260730-145510-cwilliams.json * 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2048.codfw.wmnet with reason: Maintenance * 14:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2040: Maintenance * 14:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 14:51 tchin@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] (duration: 06m 48s) * 14:47 tchin@deploy1003: jforrester, tchin: Continuing with deployment * 14:47 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 14:47 tchin@deploy1003: jforrester, tchin: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:45 tchin@deploy1003: Started scap sync-world: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] * 14:42 sukhe@puppetserver1001: conftool action : set/weight=1; selector: cluster=urldownloader,service=squid * 14:42 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader,service=squid * 14:42 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1081.eqiad.wmnet with OS trixie * 14:39 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 14:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2079.codfw.wmnet with reason: host reimage * 14:36 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2098.codfw.wmnet with OS trixie * 14:32 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2079.codfw.wmnet with reason: host reimage * 14:30 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] (duration: 06m 31s) * 14:27 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 14:26 mszwarc@deploy1003: mszwarc: Continuing with deployment * 14:25 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:25 root@cumin1003: START - Cookbook sre.mysql.pool pool es1038: Maintenance * 14:25 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1038: Maintenance * 14:23 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] * 14:21 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] (duration: 11m 19s) * 14:20 root@cumin1003: START - Cookbook sre.mysql.pool pool es1038: Maintenance * 14:14 stran@deploy1003: stran: Continuing with deployment * 14:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1038 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95810 and previous config saved to /var/cache/conftool/dbconfig/20260730-141439-cwilliams.json * 14:14 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1038.eqiad.wmnet with reason: Maintenance * 14:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1036: Maintenance * 14:13 stran@deploy1003: stran: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2079.codfw.wmnet with OS trixie * 14:09 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] * 14:08 root@cumin1003: START - Cookbook sre.mysql.pool pool es2040: Maintenance * 14:08 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2040: Maintenance * 14:03 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] (duration: 31m 41s) * 14:03 root@cumin1003: START - Cookbook sre.mysql.pool pool es2040: Maintenance * 14:03 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2040 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95806 and previous config saved to /var/cache/conftool/dbconfig/20260730-135643-cwilliams.json * 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2040.codfw.wmnet with reason: Maintenance * 13:56 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:56 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2038: Maintenance * 13:55 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:52 lucaswerkmeister-wmde@deploy1003: migr, lucaswerkmeister-wmde: Continuing with deployment * 13:49 lucaswerkmeister-wmde@deploy1003: migr, lucaswerkmeister-wmde: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:49 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:48 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 13:45 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 13:32 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:32 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] * 13:28 root@cumin1003: START - Cookbook sre.mysql.pool pool es1036: Maintenance * 13:28 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1036: Maintenance * 13:22 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:22 root@cumin1003: START - Cookbook sre.mysql.pool pool es1036: Maintenance * 13:20 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1036 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95800 and previous config saved to /var/cache/conftool/dbconfig/20260730-131727-cwilliams.json * 13:17 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1036.eqiad.wmnet with reason: Maintenance * 13:17 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] (duration: 10m 31s) * 13:16 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2022\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 13:13 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, stran: Continuing with deployment * 13:10 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:10 root@cumin1003: START - Cookbook sre.mysql.pool pool es2038: Maintenance * 13:10 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2038: Maintenance * 13:08 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, stran: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie * 13:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2047: Maintenance * 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048 cloud-private - filippo@cumin1003" * 13:07 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048 cloud-private - filippo@cumin1003" * 13:06 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] * 13:04 root@cumin1003: START - Cookbook sre.mysql.pool pool es2038: Maintenance * 13:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2078.codfw.wmnet with OS trixie * 13:01 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2038 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95797 and previous config saved to /var/cache/conftool/dbconfig/20260730-125919-cwilliams.json * 12:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2038.codfw.wmnet with reason: Maintenance * 12:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2078.codfw.wmnet with reason: host reimage * 12:37 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2078.codfw.wmnet with reason: host reimage * 12:37 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] (duration: 06m 51s) * 12:33 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 12:32 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 12:32 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 12:32 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:30 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] * 12:19 root@cumin1003: START - Cookbook sre.mysql.pool pool es2047: Maintenance * 12:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2078.codfw.wmnet with OS trixie * 12:18 dcausse@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 12:18 dcausse@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 12:15 dcausse@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 12:14 dcausse@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 12:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2047 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95793 and previous config saved to /var/cache/conftool/dbconfig/20260730-121404-cwilliams.json * 12:13 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2047.codfw.wmnet with reason: Maintenance * 12:13 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2036: Maintenance * 12:05 ayounsi@dns1004: END - running authdns-update * 12:02 ayounsi@dns1004: START - running authdns-update * 11:51 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2077.codfw.wmnet with OS trixie * 11:48 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:46 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1080.eqiad.wmnet with OS trixie * 11:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1226: Maintenance * 11:41 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2077.codfw.wmnet with reason: host reimage * 11:28 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2077.codfw.wmnet with reason: host reimage * 11:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1080.eqiad.wmnet with reason: host reimage * 11:27 root@cumin1003: START - Cookbook sre.mysql.pool pool es2036: Maintenance * 11:27 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2036: Maintenance * 11:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1080.eqiad.wmnet with reason: host reimage * 11:21 root@cumin1003: START - Cookbook sre.mysql.pool pool es2036: Maintenance * 11:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2036 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95786 and previous config saved to /var/cache/conftool/dbconfig/20260730-111633-cwilliams.json * 11:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2036.codfw.wmnet with reason: Maintenance * 11:08 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2077.codfw.wmnet with OS trixie * 11:07 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 11:03 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie * 11:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1226: Maintenance * 10:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1226 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95783 and previous config saved to /var/cache/conftool/dbconfig/20260730-104801-cwilliams.json * 10:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1226.eqiad.wmnet with reason: Maintenance * 10:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1214: Maintenance * 10:27 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2035: Maintenance * 10:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2076.codfw.wmnet with OS trixie * 10:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1214: Maintenance * 09:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1214 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95775 and previous config saved to /var/cache/conftool/dbconfig/20260730-095451-cwilliams.json * 09:54 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1214.eqiad.wmnet with reason: Maintenance * 09:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1209: Maintenance * 09:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2076.codfw.wmnet with reason: host reimage * 09:42 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool es2035: Maintenance * 09:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.netbox.update-extras (exit_code=0) rolling restart_daemons on A:netbox * 09:41 ayounsi@cumin1003: START - Cookbook sre.netbox.update-extras rolling restart_daemons on A:netbox * 09:40 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2035: Maintenance * 09:39 ayounsi@cumin1003: END (PASS) - Cookbook sre.netbox.update-extras (exit_code=0) rolling restart_daemons on A:netbox-canary * 09:39 ayounsi@cumin1003: START - Cookbook sre.netbox.update-extras rolling restart_daemons on A:netbox-canary * 09:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2076.codfw.wmnet with reason: host reimage * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 09:34 root@cumin1003: START - Cookbook sre.mysql.pool pool es2035: Maintenance * 09:32 lucaswerkmeister-wmde@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 09:32 lucaswerkmeister-wmde@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 09:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2035 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95771 and previous config saved to /var/cache/conftool/dbconfig/20260730-092910-cwilliams.json * 09:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2035.codfw.wmnet with reason: Maintenance * 09:19 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2076.codfw.wmnet with OS trixie * 09:18 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 09:17 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie * 09:07 root@cumin1003: START - Cookbook sre.mysql.pool pool db1209: Maintenance * 09:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 23 hosts * 09:04 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Remove cable label from interfaces descriptions - ayounsi@cumin1003 * 09:04 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:02 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Remove cable label from interfaces descriptions - ayounsi@cumin1003 * 09:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1209 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95767 and previous config saved to /var/cache/conftool/dbconfig/20260730-090133-cwilliams.json * 09:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1209.eqiad.wmnet with reason: Maintenance * 09:01 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1192: Maintenance * 08:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1252: Maintenance * 08:57 jayme@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 08:56 jayme@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 08:53 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 23 hosts * 08:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:51 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:50 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1263: Maintenance * 08:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:23 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 08:15 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie * 08:14 root@cumin1003: START - Cookbook sre.mysql.pool pool db1192: Maintenance * 08:13 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2075.codfw.wmnet with OS trixie * 08:12 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1252: Maintenance * 08:11 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1252: Maintenance * 08:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1252: Maintenance * 08:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1192 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95754 and previous config saved to /var/cache/conftool/dbconfig/20260730-080611-cwilliams.json * 08:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1192.eqiad.wmnet with reason: Maintenance * 08:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1178: Maintenance * 08:05 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS bullseye * 07:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1263: Maintenance * 07:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1263 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95751 and previous config saved to /var/cache/conftool/dbconfig/20260730-075106-cwilliams.json * 07:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[1260-1262].eqiad.wmnet with reason: Maintenance * 07:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2075.codfw.wmnet with reason: host reimage * 07:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1263.eqiad.wmnet with reason: Maintenance * 07:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2075.codfw.wmnet with reason: host reimage * 07:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance * 07:38 dcausse@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:38 dcausse@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 07:35 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db1252', diff saved to https://phabricator.wikimedia.org/P95748 and previous config saved to /var/cache/conftool/dbconfig/20260730-073510-marostegui.json * 07:26 klausman@dns2004: END - running authdns-update * 07:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 07:25 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2075.codfw.wmnet with OS trixie * 07:24 klausman@dns2004: START - running authdns-update * 07:23 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host an-test-master1003.eqiad.wmnet * 07:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db1178: Maintenance * 07:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1178 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95746 and previous config saved to /var/cache/conftool/dbconfig/20260730-071112-cwilliams.json * 07:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1178.eqiad.wmnet with reason: Maintenance * 07:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1177: Maintenance * 06:24 root@cumin1003: START - Cookbook sre.mysql.pool pool db1177: Maintenance * 06:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1177 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95741 and previous config saved to /var/cache/conftool/dbconfig/20260730-061736-cwilliams.json * 06:17 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1177.eqiad.wmnet with reason: Maintenance * 06:17 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1172: Maintenance * 05:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1218.eqiad.wmnet with reason: crashed * 05:41 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db1217 it crashed', diff saved to https://phabricator.wikimedia.org/P95737 and previous config saved to /var/cache/conftool/dbconfig/20260730-054111-marostegui.json * 05:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95736 and previous config saved to /var/cache/conftool/dbconfig/20260730-053422-cwilliams.json * 05:30 root@cumin1003: START - Cookbook sre.mysql.pool pool db1172: Maintenance * 05:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95734 and previous config saved to /var/cache/conftool/dbconfig/20260730-052414-cwilliams.json * 05:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1172 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95733 and previous config saved to /var/cache/conftool/dbconfig/20260730-052354-cwilliams.json * 05:23 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1172.eqiad.wmnet with reason: Maintenance * 05:23 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1167: Maintenance * 05:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95731 and previous config saved to /var/cache/conftool/dbconfig/20260730-051406-cwilliams.json * 05:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95729 and previous config saved to /var/cache/conftool/dbconfig/20260730-050358-cwilliams.json * 04:47 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 04:47 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 04:47 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 04:35 root@cumin1003: START - Cookbook sre.mysql.pool pool db1167: Maintenance * 04:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1167 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95726 and previous config saved to /var/cache/conftool/dbconfig/20260730-042923-cwilliams.json * 04:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 04:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1167.eqiad.wmnet with reason: Maintenance * 04:22 pt1979@cumin2002: START - Cookbook sre.dns.netbox * 04:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95725 and previous config saved to /var/cache/conftool/dbconfig/20260730-040337-cwilliams.json * 04:03 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:38 brett@cumin2002: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool eqsin [reason: Switch upgrade maintenance window complete, [[phab:T433097|T433097]]] * 01:38 brett@cumin2002: START - Cookbook sre.dns.admin DNS admin: pool eqsin [reason: Switch upgrade maintenance window complete, [[phab:T433097|T433097]]] == 2026-07-29 == * 23:57 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin,mr1-eqsin IPv6,mr1-eqsin.oob,mr1-eqsin.oob IPv6 with reason: connection issue * 22:54 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 22:53 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 22:53 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 22:53 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:25 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2022.codfw.wmnet, repooling source-only afterwards * 22:20 brett@cumin2002: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool eqsin [reason: Switch upgrade maintenance window, [[phab:T433097|T433097]]] * 22:20 brett@cumin2002: START - Cookbook sre.dns.admin DNS admin: depool eqsin [reason: Switch upgrade maintenance window, [[phab:T433097|T433097]]] * 22:01 apine@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 22:00 apine@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 21:59 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 21:58 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 21:58 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 21:58 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 21:32 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1048.eqiad.wmnet with OS trixie * 21:25 pt1979@cumin2002: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be2097.codfw.wmnet with OS bullseye * 21:16 zabe@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki=metawiki 'Mental Health Resource Center' 'Safety Resource Center/Mental Health' Zabe --reason 'per request [[:phab:T433118{{!}}T433118]]' * 21:12 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2022.codfw.wmnet, repooling source-only afterwards * 21:12 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] (duration: 12m 53s) * 21:12 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2015\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 21:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1253: Maintenance * 21:08 aaron@deploy1003: aaron: Continuing with deployment * 21:01 aaron@deploy1003: aaron: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:59 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] * 20:52 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] (duration: 21m 57s) * 20:48 aaron@deploy1003: aaron: Continuing with deployment * 20:32 aaron@deploy1003: aaron: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:30 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] * 20:24 root@cumin1003: START - Cookbook sre.mysql.pool pool db1253: Maintenance * 20:19 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] (duration: 08m 07s) * 20:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1253 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95719 and previous config saved to /var/cache/conftool/dbconfig/20260729-201810-cwilliams.json * 20:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1253.eqiad.wmnet with reason: Maintenance * 20:17 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1231: Maintenance * 20:15 aaron@deploy1003: bpirkle, aaron: Continuing with deployment * 20:13 aaron@deploy1003: bpirkle, aaron: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:12 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie * 20:11 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] * 20:11 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1048.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:09 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1048.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:09 pt1979@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 20:04 pt1979@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 19:47 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 19:43 pt1979@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye * 19:41 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 19:37 zabe: zabe@deploy1003:~$ mwscript-k8s --comment='[[phab:T433529|T433529]]' --follow -- resetAuthenticationThrottle.php --wiki=aawiki --signup --ip=89.36.114.94 * 19:36 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] (duration: 06m 49s) * 19:32 zabe@deploy1003: zabe: Continuing with deployment * 19:31 zabe@deploy1003: zabe: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:31 root@cumin1003: START - Cookbook sre.mysql.pool pool db1231: Maintenance * 19:29 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] * 19:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1231 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95714 and previous config saved to /var/cache/conftool/dbconfig/20260729-192454-cwilliams.json * 19:24 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1231.eqiad.wmnet with reason: Maintenance * 19:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1227: Maintenance * 19:22 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 19:22 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 19:21 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:21 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1048] - vriley@cumin1003" * 19:21 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1048] - vriley@cumin1003" * 19:19 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 19:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1251: Maintenance * 19:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95711 and previous config saved to /var/cache/conftool/dbconfig/20260729-191756-cwilliams.json * 19:16 vriley@cumin1003: START - Cookbook sre.dns.netbox * 19:11 dduvall: rolling back wmf.13 to group0 due to [[phab:T433457|T433457]] (cc [[phab:T430832|T430832]]) * 19:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95709 and previous config saved to /var/cache/conftool/dbconfig/20260729-190748-cwilliams.json * 19:01 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2022.codfw.wmnet with OS bookworm * 19:01 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2021.codfw.wmnet, repooling source-only afterwards * 18:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95707 and previous config saved to /var/cache/conftool/dbconfig/20260729-185740-cwilliams.json * 18:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95704 and previous config saved to /var/cache/conftool/dbconfig/20260729-184732-cwilliams.json * 18:37 root@cumin1003: START - Cookbook sre.mysql.pool pool db1227: Maintenance * 18:34 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2022.codfw.wmnet with reason: host reimage * 18:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1227 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95701 and previous config saved to /var/cache/conftool/dbconfig/20260729-183117-cwilliams.json * 18:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1227.eqiad.wmnet with reason: Maintenance * 18:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1202: Maintenance * 18:30 root@cumin1003: START - Cookbook sre.mysql.pool pool db1251: Maintenance * 18:27 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2022.codfw.wmnet with reason: host reimage * 18:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1251 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95698 and previous config saved to /var/cache/conftool/dbconfig/20260729-182428-cwilliams.json * 18:24 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lvs2014.codfw.wmnet * 18:24 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for lvs2014.codfw.wmnet * 18:24 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1251.eqiad.wmnet with reason: Maintenance * 18:23 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1235: Maintenance * 18:22 brett@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 18:19 brett@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 18:19 brett@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 18:17 brett@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 18:17 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 18:16 mutante: removing jenkins during the train - living on the edge - no, just kidding, jenkins has migrated to dedicated machines, nothing should happen * 18:15 brett@cumin2002: END (ERROR) - Cookbook sre.loadbalancer.restart-pybal (exit_code=97) rolling-restart of pybal on P<nowiki>{</nowiki>lvs2014.codfw.wmnet<nowiki>}</nowiki> and A:lvs ([[phab:T428495|T428495]]) * 18:15 mutante: CI: contint1002/contint2002: apt-get remove --purge jenkins - jenkins be gone - [[phab:T418521|T418521]] * 18:13 brett@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on P<nowiki>{</nowiki>lvs2014.codfw.wmnet<nowiki>}</nowiki> and A:lvs ([[phab:T428495|T428495]]) * 18:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2022 * 18:08 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2022 * 18:03 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T428495|T428495]] * 18:03 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2022 * 18:02 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2022.codfw.wmnet 211.48.192.10.in-addr.arpa 1.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:02 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2022.codfw.wmnet 211.48.192.10.in-addr.arpa 1.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:02 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:02 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2022 - bking@cumin2003" * 18:02 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2022 - bking@cumin2003" * 17:57 bking@cumin2003: START - Cookbook sre.dns.netbox * 17:56 brett@cumin2002: END (FAIL) - Cookbook sre.loadbalancer.restart-pybal (exit_code=1) rolling-restart of pybal on A:lvs-codfw and A:lvs ([[phab:T428495|T428495]]) * 17:55 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo - [[phab:T428495|T428495]] * 17:54 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2022 * 17:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2022.codfw.wmnet with OS bookworm * 17:50 brett@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on A:lvs-codfw and A:lvs ([[phab:T428495|T428495]]) * 17:47 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2021.codfw.wmnet, repooling source-only afterwards * 17:47 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 14s) * 17:47 swfrench-wmf: authdns-update to direct codfw, eqsin, ulsfo etcd clients back to codfw - [[phab:T428495|T428495]] * 17:47 swfrench@dns1004: END - running authdns-update * 17:47 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 17:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95692 and previous config saved to /var/cache/conftool/dbconfig/20260729-174713-cwilliams.json * 17:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance * 17:46 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1249: Maintenance * 17:45 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2015\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 17:45 swfrench@dns1004: START - running authdns-update * 17:44 root@cumin1003: START - Cookbook sre.mysql.pool pool db1202: Maintenance * 17:41 akhatun: Deployed refinery using scap, then deployed onto hdfs * 17:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1202 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95688 and previous config saved to /var/cache/conftool/dbconfig/20260729-173759-cwilliams.json * 17:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1202.eqiad.wmnet with reason: Maintenance * 17:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1194: Maintenance * 17:37 root@cumin1003: START - Cookbook sre.mysql.pool pool db1235: Maintenance * 17:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1230: Maintenance * 17:30 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1235 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95684 and previous config saved to /var/cache/conftool/dbconfig/20260729-173051-cwilliams.json * 17:30 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1235.eqiad.wmnet with reason: Maintenance * 17:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1234: Maintenance * 17:26 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (thin): Regular analytics weekly train THIN [analytics/refinery@56695674] (duration: 02m 02s) * 17:24 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (thin): Regular analytics weekly train THIN [analytics/refinery@56695674] * 17:23 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567]: Regular analytics weekly train [analytics/refinery@56695674] (duration: 06m 20s) * 17:20 dancy@deploy1003: Finished scap sync-world: Testing delay_messageblobstore_purge: true (duration: 06m 29s) * 17:17 akhatun@deploy1003: Started deploy [analytics/refinery@5669567]: Regular analytics weekly train [analytics/refinery@56695674] * 17:17 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] (duration: 00m 22s) * 17:16 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] * 17:13 dancy@deploy1003: Started scap sync-world: Testing delay_messageblobstore_purge: true * 17:05 mutante: CI: contint1002/contint2002 - restarted httpd to be extra sure all is cleaned up - https://integration.wikimedia.org/ci/ is up and running [[phab:T418521|T418521]] * 17:04 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 17:03 mutante: CI: contint1002/contint2002 - rm /etc/apache2/jenkins_proxy - removing legacy jenkins proxy config - jenkins is on new dedicated machines and uses jenkins_proxy_ext config [[phab:T418521|T418521]] * 17:02 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] (duration: 36m 25s) * 17:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1249: Maintenance * 16:59 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 16:54 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2015.codfw.wmnet, repooling source-only afterwards * 16:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1249 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95674 and previous config saved to /var/cache/conftool/dbconfig/20260729-165339-cwilliams.json * 16:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1249.eqiad.wmnet with reason: Maintenance * 16:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1248: Maintenance * 16:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1194: Maintenance * 16:47 swfrench-wmf: silenced EtcdReplicationDown 57b2b421-1cc9-4e38-9276-{{Gerrit|94f223fd231c}} - [[phab:T428495|T428495]] * 16:46 tchin@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/eventstreams-internal: apply * 16:46 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye * 16:46 tchin@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/eventstreams-internal: apply * 16:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1230: Maintenance * 16:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1194 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95669 and previous config saved to /var/cache/conftool/dbconfig/20260729-164422-cwilliams.json * 16:44 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 16:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1194.eqiad.wmnet with reason: Maintenance * 16:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1191: Maintenance * 16:43 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host an-test-master1003.eqiad.wmnet * 16:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db1234: Maintenance * 16:43 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 16:43 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Rolling back deployment * 16:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host an-test-master1004.eqiad.wmnet * 16:41 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 16:40 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 16:40 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 16:39 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 16:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1230 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95667 and previous config saved to /var/cache/conftool/dbconfig/20260729-163932-cwilliams.json * 16:39 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 16:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1230.eqiad.wmnet with reason: Maintenance * 16:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1207: Maintenance * 16:38 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 16:37 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host an-test-master1004.eqiad.wmnet * 16:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1234 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95664 and previous config saved to /var/cache/conftool/dbconfig/20260729-163719-cwilliams.json * 16:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1234.eqiad.wmnet with reason: Maintenance * 16:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1079.eqiad.wmnet with OS trixie * 16:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1232: Maintenance * 16:34 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1259: Maintenance * 16:28 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] (duration: 06m 57s) * 16:28 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:26 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] * 16:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1051 hosts * 16:21 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] * 16:20 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2006.codfw.wmnet with OS bookworm * 16:19 akhatun: Deploying Refinery at {{Gerrit|56695674}} as part of weekly train * 16:18 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1079.eqiad.wmnet with reason: host reimage * 16:16 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] (duration: 15m 36s) * 16:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2021.codfw.wmnet with OS bookworm * 16:14 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1079.eqiad.wmnet with reason: host reimage * 16:12 topranks: hot-swap line card in FPC0 on cr1-eqiad with replacement MPC10E from Juniper [[phab:T426343|T426343]] * 16:10 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Continuing with deployment * 16:07 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db1248: Maintenance * 16:01 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] * 16:00 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 16:00 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 15:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1248 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95651 and previous config saved to /var/cache/conftool/dbconfig/20260729-155956-cwilliams.json * 15:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1248.eqiad.wmnet with reason: Maintenance * 15:59 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2006.codfw.wmnet with reason: host reimage * 15:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1247: Maintenance * 15:59 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2074.codfw.wmnet with OS trixie * 15:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1191: Maintenance * 15:57 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:55 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1079.eqiad.wmnet with OS trixie * 15:55 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2006.codfw.wmnet with reason: host reimage * 15:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1207: Maintenance * 15:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1191 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95646 and previous config saved to /var/cache/conftool/dbconfig/20260729-155104-cwilliams.json * 15:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1191.eqiad.wmnet with reason: Maintenance * 15:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1181: Maintenance * 15:49 root@cumin1003: START - Cookbook sre.mysql.pool pool db1232: Maintenance * 15:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2021.codfw.wmnet with reason: host reimage * 15:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1207 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95643 and previous config saved to /var/cache/conftool/dbconfig/20260729-154735-cwilliams.json * 15:47 root@cumin1003: START - Cookbook sre.mysql.pool pool db1259: Maintenance * 15:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1207.eqiad.wmnet with reason: Maintenance * 15:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1200: Maintenance * 15:46 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] (duration: 31m 59s) * 15:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 15:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1232 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95640 and previous config saved to /var/cache/conftool/dbconfig/20260729-154330-cwilliams.json * 15:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1232.eqiad.wmnet with reason: Maintenance * 15:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1219: Maintenance * 15:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2015.codfw.wmnet, repooling source-only afterwards * 15:41 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 18s) * 15:41 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1259 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95638 and previous config saved to /var/cache/conftool/dbconfig/20260729-154107-cwilliams.json * 15:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1259.eqiad.wmnet with reason: Maintenance * 15:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1254: Maintenance * 15:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2015.codfw.wmnet with OS bookworm * 15:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2021.codfw.wmnet with reason: host reimage * 15:36 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2006.codfw.wmnet with OS bookworm * 15:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 15:35 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Continuing with deployment * 15:33 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2074.codfw.wmnet with OS trixie * 15:33 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 15:32 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:29 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:28 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be2074.codfw.wmnet with OS trixie * 15:28 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2006.codfw.wmnet * 15:26 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1078.eqiad.wmnet with OS trixie * 15:25 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 15:25 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:22 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2006.codfw.wmnet * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2021 * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2021 * 15:19 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2021 * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2021.codfw.wmnet 210.48.192.10.in-addr.arpa 0.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:19 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2021.codfw.wmnet 210.48.192.10.in-addr.arpa 0.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2021 - bking@cumin2003" * 15:19 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2021 - bking@cumin2003" * 15:14 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] * 15:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2015.codfw.wmnet with reason: host reimage * 15:11 root@cumin1003: START - Cookbook sre.mysql.pool pool db1247: Maintenance * 15:11 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on ml-serve2004.codfw.wmnet with reason: [[phab:T433478|T433478]] * 15:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2015.codfw.wmnet with reason: host reimage * 15:10 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on ml-serve2002.codfw.wmnet with reason: [[phab:T433476|T433476]] * 15:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 15:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1247 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95625 and previous config saved to /var/cache/conftool/dbconfig/20260729-150459-cwilliams.json * 15:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1247.eqiad.wmnet with reason: Maintenance * 15:04 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:04 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1244: Maintenance * 15:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1078.eqiad.wmnet with reason: host reimage * 15:03 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1005.wikimedia.org * 15:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db1181: Maintenance * 15:01 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2098.codfw.wmnet with OS bullseye * 15:00 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye * 15:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1200: Maintenance * 14:59 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 14:59 root@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1285.eqiad.wmnet with OS trixie * 14:59 Amir1: mwscript-k8s -- extensions/TimedMediaHandler/maintenance/requeueTranscodes.php --wiki=commonswiki --key '360p.mpeg4.mov' --throttle --video --missing ([[phab:T358266|T358266]]) * 14:58 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1005.wikimedia.org * 14:58 jhancock@cumin2002: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['ms-be2098'] * 14:58 jhancock@cumin2002: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['ms-be2098'] * 14:58 jhancock@cumin2002: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['ms-be2097'] * 14:58 jhancock@cumin2002: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['ms-be2097'] * 14:58 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1078.eqiad.wmnet with reason: host reimage * 14:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1181 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95621 and previous config saved to /var/cache/conftool/dbconfig/20260729-145629-cwilliams.json * 14:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1181.eqiad.wmnet with reason: Maintenance * 14:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1174: Maintenance * 14:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1219: Maintenance * 14:55 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader1006.wikimedia.org on all recursors * 14:55 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader1006.wikimedia.org on all recursors * 14:55 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader1005.wikimedia.org on all recursors * 14:55 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader1005.wikimedia.org on all recursors * 14:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1200 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95618 and previous config saved to /var/cache/conftool/dbconfig/20260729-145336-cwilliams.json * 14:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1254: Maintenance * 14:53 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1200.eqiad.wmnet with reason: Maintenance * 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1185: Maintenance * 14:52 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2021 * 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2015 * 14:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2015 * 14:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1219 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95616 and previous config saved to /var/cache/conftool/dbconfig/20260729-144946-cwilliams.json * 14:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1219.eqiad.wmnet with reason: Maintenance * 14:49 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1218: Maintenance * 14:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2021.codfw.wmnet with OS bookworm * 14:48 dancy@deploy1003: Finished deploy [zuul/deploy@22703a6]: Deploying https://gerrit.wikimedia.org/r/c/integration/zuul/+/1311501 ([[phab:T432491|T432491]]) (duration: 00m 15s) * 14:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2015.codfw.wmnet with OS bookworm * 14:48 dancy@deploy1003: Started deploy [zuul/deploy@22703a6]: Deploying https://gerrit.wikimedia.org/r/c/integration/zuul/+/1311501 ([[phab:T432491|T432491]]) * 14:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1254 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95613 and previous config saved to /var/cache/conftool/dbconfig/20260729-144729-cwilliams.json * 14:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1254.eqiad.wmnet with reason: Maintenance * 14:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1233: Maintenance * 14:46 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2013\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 14:46 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2014\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 14:46 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:45 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:44 root@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1285.eqiad.wmnet with reason: host reimage * 14:43 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:42 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:41 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:40 root@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1285.eqiad.wmnet with reason: host reimage * 14:39 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1078.eqiad.wmnet with OS trixie * 14:39 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2074.codfw.wmnet with OS trixie * 14:32 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:32 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:32 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:31 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2005.codfw.wmnet with OS bookworm * 14:30 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:30 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:29 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:29 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:27 root@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host db1285 * 14:27 root@cumin1003: START - Cookbook sre.hosts.move-vlan for host db1285 * 14:27 root@cumin1003: START - Cookbook sre.hosts.reimage for host db1285.eqiad.wmnet with OS trixie * 14:24 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 14:24 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:24 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:24 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:23 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:22 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:22 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:21 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:17 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:16 root@cumin1003: START - Cookbook sre.mysql.pool pool db1244: Maintenance * 14:15 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:15 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add asw1-604 loopback ipv4 - pt1979@cumin2002" * 14:15 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add asw1-604 loopback ipv4 - pt1979@cumin2002" * 14:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:12 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 14:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95599 and previous config saved to /var/cache/conftool/dbconfig/20260729-141014-cwilliams.json * 14:10 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 14:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1244.eqiad.wmnet with reason: Maintenance * 14:10 pt1979@cumin2002: START - Cookbook sre.dns.netbox * 14:10 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 14:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1243: Maintenance * 14:09 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2005.codfw.wmnet with reason: host reimage * 14:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db1174: Maintenance * 14:08 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad * 14:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db1185: Maintenance * 14:06 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 14:05 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2005.codfw.wmnet with reason: host reimage * 14:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1174 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95595 and previous config saved to /var/cache/conftool/dbconfig/20260729-140309-cwilliams.json * 14:03 sukhe@dns1004: END - running authdns-update * 14:03 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1174.eqiad.wmnet with reason: Maintenance * 14:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1170: Maintenance * 14:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db1218: Maintenance * 14:01 sukhe@dns1004: START - running authdns-update * 14:00 sukhe@dns1004: START - running authdns-update * 13:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db1233: Maintenance * 13:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1185 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95592 and previous config saved to /var/cache/conftool/dbconfig/20260729-135925-cwilliams.json * 13:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1185.eqiad.wmnet with reason: Maintenance * 13:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1161: Maintenance * 13:58 sukhe@puppetserver1001: conftool action : set/pooled=true; selector: dnsdisc=urldownloader * 13:58 root@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1265.eqiad.wmnet with OS trixie * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1218 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95590 and previous config saved to /var/cache/conftool/dbconfig/20260729-135621-cwilliams.json * 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1218.eqiad.wmnet with reason: Maintenance * 13:55 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1206: Maintenance * 13:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2073.codfw.wmnet with OS trixie * 13:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1233 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95587 and previous config saved to /var/cache/conftool/dbconfig/20260729-135335-cwilliams.json * 13:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1233.eqiad.wmnet with reason: Maintenance * 13:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1229: Maintenance * 13:50 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/kartotherian: apply * 13:50 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service * 13:49 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:49 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/kartotherian: apply * 13:48 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 13:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1077.eqiad.wmnet with OS trixie * 13:47 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 13:46 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2005.codfw.wmnet with OS bookworm * 13:44 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 13:44 root@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1265.eqiad.wmnet with reason: host reimage * 13:40 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] (duration: 09m 22s) * 13:39 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:38 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:36 root@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1265.eqiad.wmnet with reason: host reimage * 13:35 stran@deploy1003: stran: Continuing with deployment * 13:33 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 13:32 stran@deploy1003: stran: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified t * 13:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2073.codfw.wmnet with reason: host reimage * 13:30 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ml-build1001.eqiad.wmnet * 13:30 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] * 13:29 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad * 13:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1077.eqiad.wmnet with reason: host reimage * 13:27 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 13:27 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:27 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad * 13:26 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2073.codfw.wmnet with reason: host reimage * 13:26 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] (duration: 07m 56s) * 13:25 klausman@cumin1003: START - Cookbook sre.hosts.reboot-single for host ml-build1001.eqiad.wmnet * 13:24 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 13:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>ml-serve101[2-5].eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 13:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1015.eqiad.wmnet * 13:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1015.eqiad.wmnet * 13:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1077.eqiad.wmnet with reason: host reimage * 13:23 root@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host db1265 * 13:23 root@cumin1003: START - Cookbook sre.hosts.move-vlan for host db1265 * 13:23 root@cumin1003: START - Cookbook sre.hosts.reimage for host db1265.eqiad.wmnet with OS trixie * 13:23 root@cumin1003: START - Cookbook sre.mysql.pool pool db1243: Maintenance * 13:22 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 13:22 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2005.codfw.wmnet * 13:22 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 13:22 samtar@deploy1003: dreamrimmer, samtar: Continuing with deployment * 13:20 samtar@deploy1003: dreamrimmer, samtar: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts an-test-master[1001-1002].eqiad.wmnet * 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-master[1001-1002].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 13:18 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1015.eqiad.wmnet * 13:18 sukhe@cumin1003: END (ERROR) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=97) for role: url_downloader@eqiad * 13:18 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 13:18 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] * 13:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95574 and previous config saved to /var/cache/conftool/dbconfig/20260729-131638-cwilliams.json * 13:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1243.eqiad.wmnet with reason: Maintenance * 13:16 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2005.codfw.wmnet * 13:16 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1242: Maintenance * 13:14 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] (duration: 07m 00s) * 13:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db1170: Maintenance * 13:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1015.eqiad.wmnet * 13:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1014.eqiad.wmnet * 13:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1014.eqiad.wmnet * 13:12 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1223: Maintenance * 13:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db1161: Maintenance * 13:10 samtar@deploy1003: anzx, samtar: Continuing with deployment * 13:09 samtar@deploy1003: anzx, samtar: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db1206: Maintenance * 13:08 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 13:07 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] * 13:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1170 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95566 and previous config saved to /var/cache/conftool/dbconfig/20260729-130730-cwilliams.json * 13:07 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1170.eqiad.wmnet with reason: Maintenance * 13:07 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:07 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt IPs new switches - cmooney@cumin1003" * 13:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1158: Maintenance * 13:06 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1014.eqiad.wmnet * 13:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1161 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95564 and previous config saved to /var/cache/conftool/dbconfig/20260729-130616-cwilliams.json * 13:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 13:06 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1077.eqiad.wmnet with OS trixie * 13:05 root@cumin1003: START - Cookbook sre.mysql.pool pool db1229: Maintenance * 13:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1161.eqiad.wmnet with reason: Maintenance * 13:05 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt IPs new switches - cmooney@cumin1003" * 13:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2073.codfw.wmnet with OS trixie * 13:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Maintenance * 13:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1206 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95562 and previous config saved to /var/cache/conftool/dbconfig/20260729-130258-cwilliams.json * 13:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1206.eqiad.wmnet with reason: Maintenance * 13:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1196: Maintenance * 13:01 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 13:01 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:00 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 13:00 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 12:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1229 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95559 and previous config saved to /var/cache/conftool/dbconfig/20260729-125950-cwilliams.json * 12:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1229.eqiad.wmnet with reason: Maintenance * 12:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1222: Maintenance * 12:57 sukhe: sudo cumin 'A:lvs and (A:eqiad or A:codfw)' 'disable-puppet "adding new service urldownloader"': [[phab:T429175|T429175]] * 12:56 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1014.eqiad.wmnet * 12:56 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1013.eqiad.wmnet * 12:56 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1013.eqiad.wmnet * 12:50 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1013.eqiad.wmnet * 12:50 sukhe: sudo cumin 'O:url_downloader' 'run-puppet-agent --enable "merging CR 1313948"': [[phab:T429175|T429175]] * 12:48 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-master[1001-1002].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 12:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1013.eqiad.wmnet * 12:45 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1012.eqiad.wmnet * 12:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1012.eqiad.wmnet * 12:45 sukhe: sudo cumin 'O:url_downloader' 'disable-puppet "merging CR 1313948"': [[phab:T429175|T429175]] * 12:44 btullis@cumin1003: START - Cookbook sre.dns.netbox * 12:40 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test2001.codfw.wmnet * 12:40 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test2001.codfw.wmnet * 12:38 ayounsi@dns1004: END - running authdns-update * 12:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1012.eqiad.wmnet * 12:35 ayounsi@dns1004: START - running authdns-update * 12:34 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts an-test-master[1001-1002].eqiad.wmnet * 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts an-test-coord1001.eqiad.wmnet * 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-coord1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 12:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1012.eqiad.wmnet * 12:32 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>ml-serve101[2-5].eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 12:29 root@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Maintenance * 12:25 root@cumin1003: START - Cookbook sre.mysql.pool pool db1223: Maintenance * 12:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95544 and previous config saved to /var/cache/conftool/dbconfig/20260729-122254-cwilliams.json * 12:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1242.eqiad.wmnet with reason: Maintenance * 12:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1241: Maintenance * 12:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1051 hosts * 12:20 root@cumin1003: START - Cookbook sre.mysql.pool pool db1158: Maintenance * 12:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1223 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95540 and previous config saved to /var/cache/conftool/dbconfig/20260729-121937-cwilliams.json * 12:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1223.eqiad.wmnet with reason: Maintenance * 12:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1212: Maintenance * 12:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db1159: Maintenance * 12:17 elukey@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: sync * 12:15 elukey@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: sync * 12:15 root@cumin1003: START - Cookbook sre.mysql.pool pool db1196: Maintenance * 12:14 Daimona: Creating new DB tables for the CampaignEvents extension in x1.testwiki, x1.test2wiki, x1.officewiki, and x1.wikishared # [[phab:T429339|T429339]] * 12:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db1222: Maintenance * 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95535 and previous config saved to /var/cache/conftool/dbconfig/20260729-121211-cwilliams.json * 12:12 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 12:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1158.eqiad.wmnet with reason: Maintenance * 12:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1159 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95534 and previous config saved to /var/cache/conftool/dbconfig/20260729-121146-cwilliams.json * 12:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1159.eqiad.wmnet with reason: Maintenance * 12:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1196 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95533 and previous config saved to /var/cache/conftool/dbconfig/20260729-120847-cwilliams.json * 12:08 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 12:08 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1196.eqiad.wmnet with reason: Maintenance * 12:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1195: Maintenance * 12:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1222 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95530 and previous config saved to /var/cache/conftool/dbconfig/20260729-120424-cwilliams.json * 12:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1222.eqiad.wmnet with reason: Maintenance * 12:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1098 hosts * 12:00 marostegui: Rename tables [[phab:T425074|T425074]] * 12:00 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-coord1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 11:58 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1197: Maintenance * 11:55 btullis@cumin1003: START - Cookbook sre.dns.netbox * 11:52 elukey@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: sync * 11:51 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:51 elukey@deploy1003: helmfile [codfw] START helmfile.d/services/proton: sync * 11:51 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:50 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts an-test-coord1001.eqiad.wmnet * 11:50 elukey@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: sync * 11:49 elukey@deploy1003: helmfile [staging] START helmfile.d/services/proton: sync * 11:35 root@cumin1003: START - Cookbook sre.mysql.pool pool db1241: Maintenance * 11:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db1212: Maintenance * 11:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1241 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95520 and previous config saved to /var/cache/conftool/dbconfig/20260729-112918-cwilliams.json * 11:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1241.eqiad.wmnet with reason: Maintenance * 11:29 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1238: Maintenance * 11:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1212 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95517 and previous config saved to /var/cache/conftool/dbconfig/20260729-112727-cwilliams.json * 11:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 11:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1212.eqiad.wmnet with reason: Maintenance * 11:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1198: Maintenance * 11:23 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:22 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:21 root@cumin1003: START - Cookbook sre.mysql.pool pool db1195: Maintenance * 11:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1195 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95514 and previous config saved to /var/cache/conftool/dbconfig/20260729-111450-cwilliams.json * 11:14 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1195.eqiad.wmnet with reason: Maintenance * 11:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1186: Maintenance * 11:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 11:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 11:05 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:54 marostegui: Dropping renamed tables [[phab:T425066|T425066]] * 10:41 root@cumin1003: START - Cookbook sre.mysql.pool pool db1238: Maintenance * 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1198: Maintenance * 10:39 Amir1: ran https://phabricator.wikimedia.org/T432509#12149723 in production ([[phab:T432509|T432509]]) * 10:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db1197: Maintenance * 10:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1238 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95501 and previous config saved to /var/cache/conftool/dbconfig/20260729-103532-cwilliams.json * 10:35 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1238.eqiad.wmnet with reason: Maintenance * 10:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1221: Maintenance * 10:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1198 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95499 and previous config saved to /var/cache/conftool/dbconfig/20260729-103330-cwilliams.json * 10:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1198.eqiad.wmnet with reason: Maintenance * 10:33 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1175: Maintenance * 10:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1197 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95496 and previous config saved to /var/cache/conftool/dbconfig/20260729-103217-cwilliams.json * 10:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1197.eqiad.wmnet with reason: Maintenance * 10:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1188: Maintenance * 10:27 root@cumin1003: START - Cookbook sre.mysql.pool pool db1186: Maintenance * 10:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1186 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95493 and previous config saved to /var/cache/conftool/dbconfig/20260729-102111-cwilliams.json * 10:21 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1186.eqiad.wmnet with reason: Maintenance * 10:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 10:14 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 09:53 XioNoX: reboot cr2-magru - [[phab:T431750|T431750]] * 09:52 XioNoX: drain cr2-magru - [[phab:T431750|T431750]] * 09:48 root@cumin1003: START - Cookbook sre.mysql.pool pool db1221: Maintenance * 09:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zookeeper-test1002.eqiad.wmnet * 09:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1188: Maintenance * 09:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1175: Maintenance * 09:44 btullis@dns1004: END - running authdns-update * 09:42 btullis@dns1004: START - running authdns-update * 09:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1221 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95483 and previous config saved to /var/cache/conftool/dbconfig/20260729-094200-cwilliams.json * 09:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 7 hosts with reason: Maintenance * 09:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1221.eqiad.wmnet with reason: Maintenance * 09:41 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host zookeeper-test1002.eqiad.wmnet * 09:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1199: Maintenance * 09:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1188 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95481 and previous config saved to /var/cache/conftool/dbconfig/20260729-093917-cwilliams.json * 09:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1188.eqiad.wmnet with reason: Maintenance * 09:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1182: Maintenance * 09:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1175 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95479 and previous config saved to /var/cache/conftool/dbconfig/20260729-093842-cwilliams.json * 09:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1175.eqiad.wmnet with reason: Maintenance * 09:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1166: Maintenance * 09:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1169: Maintenance * 09:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1033.eqiad.wmnet,service=s8 * 09:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1033.eqiad.wmnet,service=s5 * 09:33 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1033.eqiad.wmnet,service=s8 * 09:33 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1033.eqiad.wmnet,service=s5 * 09:21 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:21 XioNoX: reboot cr1-magru - [[phab:T431750|T431750]] * 09:17 XioNoX: drain cr1-magru - [[phab:T431750|T431750]] * 09:15 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm1001.wikimedia.org * 09:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply * 09:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply * 09:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 09:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 09:11 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr2-magru,cr2-magru IPv6,cr2-magru.mgmt with reason: router upgrade * 09:11 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm1001.wikimedia.org * 09:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp1005.wikimedia.org * 09:07 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp1005.wikimedia.org * 09:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp2005.wikimedia.org * 09:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp2005.wikimedia.org * 09:00 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr1-magru,cr1-magru IPv6,cr1-magru.mgmt with reason: router upgrade * 09:00 marostegui: Dropping renamed tables [[phab:T426341|T426341]] * 08:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1199: Maintenance * 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1182: Maintenance * 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1166: Maintenance * 08:47 root@cumin1003: START - Cookbook sre.mysql.pool pool db1169: Maintenance * 08:46 ayounsi@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 1:00:00 on cr1-magru,cr1-magru IPv6,cr1-magru.mgmt with reason: router upgrade * 08:45 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2072.codfw.wmnet with OS trixie * 08:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1199 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95464 and previous config saved to /var/cache/conftool/dbconfig/20260729-084534-cwilliams.json * 08:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1199.eqiad.wmnet with reason: Maintenance * 08:45 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 08:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1190: Maintenance * 08:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1182 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95462 and previous config saved to /var/cache/conftool/dbconfig/20260729-084436-cwilliams.json * 08:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1182.eqiad.wmnet with reason: Maintenance * 08:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1156: Maintenance * 08:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1166 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95460 and previous config saved to /var/cache/conftool/dbconfig/20260729-084400-cwilliams.json * 08:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1166.eqiad.wmnet with reason: Maintenance * 08:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1157: Maintenance * 08:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95458 and previous config saved to /var/cache/conftool/dbconfig/20260729-084147-cwilliams.json * 08:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1169.eqiad.wmnet with reason: Maintenance * 08:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1163: Maintenance * 08:30 btullis@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 11 hosts with reason: Replacing the namenodes * 08:23 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2072.codfw.wmnet with reason: host reimage * 08:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1098 hosts * 08:19 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2072.codfw.wmnet with reason: host reimage * 07:58 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2072.codfw.wmnet with OS trixie * 07:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1190: Maintenance * 07:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1156: Maintenance * 07:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1157: Maintenance * 07:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1163: Maintenance * 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1190 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95444 and previous config saved to /var/cache/conftool/dbconfig/20260729-074930-cwilliams.json * 07:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1190.eqiad.wmnet with reason: Maintenance * 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1157 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95443 and previous config saved to /var/cache/conftool/dbconfig/20260729-074914-cwilliams.json * 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1156 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95442 and previous config saved to /var/cache/conftool/dbconfig/20260729-074906-cwilliams.json * 07:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1157.eqiad.wmnet with reason: Maintenance * 07:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 07:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1156.eqiad.wmnet with reason: Maintenance * 07:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1163 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95441 and previous config saved to /var/cache/conftool/dbconfig/20260729-074652-cwilliams.json * 07:46 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1163.eqiad.wmnet with reason: Maintenance * 07:46 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2034.codfw.wmnet * 07:42 ayounsi@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2034.codfw.wmnet * 07:42 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2232.codfw.wmnet with OS trixie * 07:34 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:34 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:33 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:31 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:19 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2232.codfw.wmnet with reason: host reimage * 07:15 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2232.codfw.wmnet with reason: host reimage * 06:58 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db2232.codfw.wmnet with OS trixie * 06:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[2160,2232].codfw.wmnet with reason: Reimage * 06:26 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1164.eqiad.wmnet with OS trixie * 06:05 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1164.eqiad.wmnet with reason: host reimage * 06:01 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1164.eqiad.wmnet with reason: host reimage * 05:47 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1164.eqiad.wmnet with OS trixie * 05:46 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1164.eqiad.wmnet with reason: Reimage == 2026-07-28 == * 22:50 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1138.eqiad.wmnet * 22:50 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1138.eqiad.wmnet * 22:49 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1138.eqiad.wmnet * 22:11 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2014.codfw.wmnet, repooling source-only afterwards * 22:08 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2013.codfw.wmnet, repooling source-only afterwards * 22:03 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 20:58 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] (duration: 08m 19s) * 20:55 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2014.codfw.wmnet, repooling source-only afterwards * 20:55 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2013.codfw.wmnet, repooling source-only afterwards * 20:54 arlolra@deploy1003: arlolra: Continuing with deployment * 20:54 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 14s) * 20:54 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 20:53 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 30s) * 20:53 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 20:52 arlolra@deploy1003: arlolra: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:51 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:50 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] * 20:49 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:43 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 20:34 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] (duration: 06m 54s) * 20:34 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:34 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:31 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:31 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:30 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:30 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:30 arlolra@deploy1003: arlolra: Continuing with deployment * 20:29 arlolra@deploy1003: arlolra: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:27 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] * 20:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2014.codfw.wmnet with OS bookworm * 20:21 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 20:21 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:20 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 20:19 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:19 swfrench-wmf: switched etcd-mirror replication from conf2005 to conf2004 - [[phab:T428495|T428495]] * 20:17 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:17 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:15 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] (duration: 08m 26s) * 20:12 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:11 arlolra@deploy1003: anzx, arlolra: Continuing with deployment * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2013.codfw.wmnet with OS bookworm * 20:09 arlolra@deploy1003: anzx, arlolra: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] * 20:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2216: Maintenance * 19:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2014.codfw.wmnet with reason: host reimage * 19:57 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:54 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:54 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:53 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:52 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2014.codfw.wmnet with reason: host reimage * 19:49 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2013.codfw.wmnet with reason: host reimage * 19:42 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 19:41 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:41 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2013.codfw.wmnet with reason: host reimage * 19:39 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:39 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:39 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-eqiad: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 19:38 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2014 * 19:33 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2014 * 19:29 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2014.codfw.wmnet with OS bookworm * 19:28 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:27 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1006 * 19:26 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2012\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 19:26 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1006 * 19:26 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:26 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 19:25 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2013 * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2013 * 19:21 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2013 * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2013.codfw.wmnet 84.0.192.10.in-addr.arpa 4.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:21 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2013.codfw.wmnet 84.0.192.10.in-addr.arpa 4.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2013 - bking@cumin2003" * 19:21 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2013 - bking@cumin2003" * 19:21 vriley@cumin1003: START - Cookbook sre.dns.netbox * 19:20 root@cumin1003: START - Cookbook sre.mysql.pool pool db2216: Maintenance * 19:13 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2216 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95435 and previous config saved to /var/cache/conftool/dbconfig/20260728-191343-cwilliams.json * 19:13 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2216.codfw.wmnet with reason: Maintenance * 19:13 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2203: Maintenance * 19:06 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1005.eqiad.wmnet with OS trixie * 19:06 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 19:06 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 18:46 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 18:45 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:45 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:43 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:40 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:36 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-eqiad: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 18:35 dancy@deploy1003: Installation of scap version "4.275.0" completed for 3 hosts * 18:33 dancy@deploy1003: Installing scap version "4.275.0" for 3 host(s) * 18:32 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:32 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2097.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:30 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns3003.wikimedia.org [reason: pool for all services after reimaging] * 18:29 sukhe@dns1004: END - running authdns-update * 18:27 sukhe@dns1004: START - running authdns-update * 18:27 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns3003.wikimedia.org,service=authdns-update [reason: pool authdns-update after reimaging] * 18:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db2203: Maintenance * 18:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2203 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95430 and previous config saved to /var/cache/conftool/dbconfig/20260728-181958-cwilliams.json * 18:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2203.codfw.wmnet with reason: Maintenance * 18:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2188: Maintenance * 18:18 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2097.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:17 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-be2098 * 18:17 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host ms-be2098 * 18:17 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-be2097 * 18:16 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host ms-be2097 * 18:15 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:15 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding ms-be2097-8 to codfw - jhancock@cumin2002" * 18:15 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding ms-be2097-8 to codfw - jhancock@cumin2002" * 18:10 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 18:08 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage * 18:05 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns3003.wikimedia.org with OS trixie * 18:03 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage * 17:56 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-codfw: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 17:45 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie * 17:45 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1005.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:41 sukhe@dns1004: END - running authdns-update * 17:39 sukhe@dns1004: START - running authdns-update * 17:36 sukhe@puppetserver1001: conftool action : set/weight=1; selector: cluster=urldownloader,service=squid * 17:36 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1005.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:35 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader,service=squid * 17:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 17:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1005 * 17:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 17:34 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1005 * 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1005] - vriley@cumin1003" * 17:34 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1005] - vriley@cumin1003" * 17:32 root@cumin1003: START - Cookbook sre.mysql.pool pool db2188: Maintenance * 17:29 vriley@cumin1003: START - Cookbook sre.dns.netbox * 17:29 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 17:26 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2188 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95425 and previous config saved to /var/cache/conftool/dbconfig/20260728-172609-cwilliams.json * 17:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2188.codfw.wmnet with reason: Maintenance * 17:25 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2176: Maintenance * 17:19 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1138.eqiad.wmnet with OS trixie * 17:18 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1005 * 17:18 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1005 * 17:18 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:15 vriley@cumin1003: START - Cookbook sre.dns.netbox * 17:13 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns3003.wikimedia.org with reason: host reimage * 17:07 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns3003.wikimedia.org with reason: host reimage * 17:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-eqiad * 17:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1015.eqiad.wmnet * 17:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1015.eqiad.wmnet * 16:59 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1138.eqiad.wmnet with reason: host reimage * 16:55 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-codfw: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 16:54 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1138.eqiad.wmnet with reason: host reimage * 16:53 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1015.eqiad.wmnet * 16:43 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns3003.wikimedia.org with OS trixie * 16:43 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1015.eqiad.wmnet * 16:43 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1014.eqiad.wmnet * 16:43 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1014.eqiad.wmnet * 16:43 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=dns3003.wikimedia.org [reason: depooling for reimage to trixie] * 16:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 290 hosts * 16:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db2176: Maintenance * 16:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1138 * 16:38 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1138 * 16:37 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1138 * 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1138.eqiad.wmnet 193.32.64.10.in-addr.arpa 3.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:37 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1138.eqiad.wmnet 193.32.64.10.in-addr.arpa 3.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1138 - jiji@cumin1003" * 16:37 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1138 - jiji@cumin1003" * 16:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1014.eqiad.wmnet * 16:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2176 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95420 and previous config saved to /var/cache/conftool/dbconfig/20260728-163235-cwilliams.json * 16:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2176.codfw.wmnet with reason: Maintenance * 16:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1014.eqiad.wmnet * 16:32 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1013.eqiad.wmnet * 16:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1013.eqiad.wmnet * 16:32 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2174: Maintenance * 16:28 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2012.codfw.wmnet, repooling source-only afterwards * 16:25 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1013.eqiad.wmnet * 16:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1013.eqiad.wmnet * 16:20 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1012.eqiad.wmnet * 16:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1012.eqiad.wmnet * 16:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1012.eqiad.wmnet * 16:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1012.eqiad.wmnet * 16:03 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1011.eqiad.wmnet * 16:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1011.eqiad.wmnet * 16:00 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2004.codfw.wmnet with OS bookworm * 15:59 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1011.eqiad.wmnet * 15:56 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:55 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 15:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:54 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1011.eqiad.wmnet * 15:54 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1010.eqiad.wmnet * 15:54 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1010.eqiad.wmnet * 15:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:50 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 15:49 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1010.eqiad.wmnet * 15:48 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 15:48 jiji@cumin1003: START - Cookbook sre.dns.netbox * 15:46 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2248: Maintenance * 15:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db2174: Maintenance * 15:44 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1010.eqiad.wmnet * 15:44 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1009.eqiad.wmnet * 15:44 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1009.eqiad.wmnet * 15:42 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1138 * 15:41 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1138.eqiad.wmnet with OS trixie * 15:39 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1009.eqiad.wmnet * 15:39 robh@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on arclamp2001.codfw.wmnet with reason: ram upgrade * 15:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2174 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95413 and previous config saved to /var/cache/conftool/dbconfig/20260728-153844-cwilliams.json * 15:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2174.codfw.wmnet with reason: Maintenance * 15:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2173: Maintenance * 15:37 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 15:35 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 15:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1009.eqiad.wmnet * 15:34 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1008.eqiad.wmnet * 15:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1008.eqiad.wmnet * 15:31 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 15:31 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 15:29 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1008.eqiad.wmnet * 15:27 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 290 hosts * 15:25 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2013 * 15:25 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2195: Maintenance * 15:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1008.eqiad.wmnet * 15:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1007.eqiad.wmnet * 15:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1007.eqiad.wmnet * 15:22 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2004.codfw.wmnet with reason: host reimage * 15:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2013.codfw.wmnet with OS bookworm * 15:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 15:19 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2004.codfw.wmnet with reason: host reimage * 15:19 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 15:17 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1007.eqiad.wmnet * 15:12 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1007.eqiad.wmnet * 15:12 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1006.eqiad.wmnet * 15:12 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1006.eqiad.wmnet * 15:11 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1138.eqiad.wmnet * 15:11 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 15:11 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1138.eqiad.wmnet * 15:11 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1138.eqiad.wmnet * 15:10 brennen@deploy1003: Finished deploy [phabricator/deployment@f8b349f]: deploy phab1004 for [[phab:T433382|T433382]] (duration: 00m 43s) * 15:10 brennen@deploy1003: Started deploy [phabricator/deployment@f8b349f]: deploy phab1004 for [[phab:T433382|T433382]] * 15:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts * 15:09 brennen@deploy1003: Finished deploy [phabricator/deployment@f8b349f]: deploy phab2003 for [[phab:T433382|T433382]] (duration: 00m 55s) * 15:08 brennen@deploy1003: Started deploy [phabricator/deployment@f8b349f]: deploy phab2003 for [[phab:T433382|T433382]] * 15:07 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2012.codfw.wmnet, repooling source-only afterwards * 15:07 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts * 15:06 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1137.eqiad.wmnet * 15:06 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1137.eqiad.wmnet * 15:06 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1137.eqiad.wmnet * 15:05 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1006.eqiad.wmnet * 15:05 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1022\.eqiad\.wmnet,dc=eqiad,cluster=wdqs\-main,service=wdqs\-main * 15:01 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab2003.codfw.wmnet with reason: deployment * 15:01 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1005.eqiad.wmnet with reason: deployment * 15:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1006.eqiad.wmnet * 15:00 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1006.eqiad.wmnet with reason: deployment * 15:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1005.eqiad.wmnet * 15:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1005.eqiad.wmnet * 14:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db2248: Maintenance * 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1004.eqiad.wmnet with reason: deployment * 14:59 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2004.codfw.wmnet with OS bookworm * 14:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1005.eqiad.wmnet * 14:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2248 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95403 and previous config saved to /var/cache/conftool/dbconfig/20260728-145532-cwilliams.json * 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2245-2247].codfw.wmnet with reason: Maintenance * 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2248.codfw.wmnet with reason: Maintenance * 14:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2240: Maintenance * 14:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2173: Maintenance * 14:51 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts * 14:50 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1005.eqiad.wmnet * 14:50 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1004.eqiad.wmnet * 14:50 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1004.eqiad.wmnet * 14:49 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts * 14:45 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2004.codfw.wmnet * 14:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2173 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95399 and previous config saved to /var/cache/conftool/dbconfig/20260728-144453-cwilliams.json * 14:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2173.codfw.wmnet with reason: Maintenance * 14:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2170: Maintenance * 14:44 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1004.eqiad.wmnet * 14:39 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2004.codfw.wmnet * 14:38 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1004.eqiad.wmnet * 14:38 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet * 14:38 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet * 14:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db2195: Maintenance * 14:36 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1022.eqiad.wmnet, repooling source-only afterwards * 14:36 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 14:33 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet * 14:33 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2222: Maintenance * 14:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2195 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95394 and previous config saved to /var/cache/conftool/dbconfig/20260728-143218-cwilliams.json * 14:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2195.codfw.wmnet with reason: Maintenance * 14:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2181: Maintenance * 14:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 14:30 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 14:25 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 14:25 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 14:25 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 14:23 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet * 14:23 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1002.eqiad.wmnet * 14:23 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1002.eqiad.wmnet * 14:23 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 19s) * 14:23 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:18 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1002.eqiad.wmnet * 14:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2012.codfw.wmnet with OS bookworm * 14:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1002.eqiad.wmnet * 14:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1001.eqiad.wmnet * 14:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1001.eqiad.wmnet * 14:11 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 14:11 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 14:08 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1001.eqiad.wmnet * 14:07 elukey@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'. * 14:07 elukey@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'. * 14:06 elukey@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'. * 14:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db2240: Maintenance * 14:06 XioNoX: un-drain cr2-esams - [[phab:T431751|T431751]] * 14:05 elukey@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'. * 14:02 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1001.eqiad.wmnet * 14:02 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-eqiad * 14:01 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 14:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2240 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95384 and previous config saved to /var/cache/conftool/dbconfig/20260728-140011-cwilliams.json * 14:00 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2240.codfw.wmnet with reason: Maintenance * 13:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2237: Maintenance * 13:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db2170: Maintenance * 13:56 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 13:55 XioNoX: reboot cr2-esams - [[phab:T431751|T431751]] * 13:52 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 13:51 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr2-esams,cr2-esams IPv6,cr2-esams.mgmt with reason: router upgrade * 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 13:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2170 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95381 and previous config saved to /var/cache/conftool/dbconfig/20260728-135043-cwilliams.json * 13:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2170.codfw.wmnet with reason: Maintenance * 13:50 XioNoX: drain cr2-esams - [[phab:T431751|T431751]] * 13:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2153: Maintenance * 13:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2012.codfw.wmnet with reason: host reimage * 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2222: Maintenance * 13:45 sukhe: restart pybal on A:lvs-codfw * 13:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2012.codfw.wmnet with reason: host reimage * 13:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db2181: Maintenance * 13:44 btullis@dns1004: END - running authdns-update * 13:42 sukhe: restart pybal on lvs2014 * 13:42 btullis@dns1004: START - running authdns-update * 13:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2222 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95376 and previous config saved to /var/cache/conftool/dbconfig/20260728-133948-cwilliams.json * 13:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2222.codfw.wmnet with reason: Maintenance * 13:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2221: Maintenance * 13:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2181 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95374 and previous config saved to /var/cache/conftool/dbconfig/20260728-133857-cwilliams.json * 13:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2181.codfw.wmnet with reason: Maintenance * 13:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2167: Maintenance * 13:30 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T428495|T428495]] * 13:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 13:29 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 13:29 ayounsi@cumin1003: END (FAIL) - Cookbook sre.dns.admin (exit_code=99) DNS admin: depool esams [reason: router upgrade, [[phab:T431749|T431749]]] * 13:28 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: router upgrade, [[phab:T431749|T431749]]] * 13:27 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1022.eqiad.wmnet, repooling source-only afterwards * 13:27 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo - [[phab:T428495|T428495]] * 13:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2012 * 13:27 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2012 * 13:21 lucaswerkmeister-wmde@deploy1003: mwscript-k8s job started: cleanupTitles bolwiki # [[phab:T429951|T429951]] * 13:21 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] (duration: 07m 19s) * 13:20 swfrench-wmf: authdns-update to direct codfw, eqsin, ulsfo etcd clients to eqiad - [[phab:T428495|T428495]] * 13:18 swfrench@dns1004: END - running authdns-update * 13:17 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, anzx: Continuing with deployment * 13:16 swfrench@dns1004: START - running authdns-update * 13:16 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2012 * 13:16 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2012.codfw.wmnet 57.48.192.10.in-addr.arpa 7.5.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:16 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2012.codfw.wmnet 57.48.192.10.in-addr.arpa 7.5.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:16 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:16 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, anzx: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:14 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 13:14 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] * 13:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db2237: Maintenance * 13:13 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2237: Maintenance * 13:13 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:12 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:12 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback IPV6 for asw1-604 - pt1979@cumin2003" * 13:12 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:12 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback IPV6 for asw1-604 - pt1979@cumin2003" * 13:11 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 20s) * 13:11 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 13:10 esanders@deploy1003: Finished scap sync-world: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] (duration: 08m 11s) * 13:08 pt1979@cumin2003: START - Cookbook sre.dns.netbox * 13:07 root@cumin1003: START - Cookbook sre.mysql.pool pool db2237: Maintenance * 13:06 esanders@deploy1003: esanders: Continuing with deployment * 13:04 esanders@deploy1003: esanders: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2153: Maintenance * 13:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2153: Maintenance * 13:02 esanders@deploy1003: Started scap sync-world: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] * 13:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2237 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95362 and previous config saved to /var/cache/conftool/dbconfig/20260728-130107-cwilliams.json * 13:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2237.codfw.wmnet with reason: Maintenance * 13:00 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2236: Maintenance * 12:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2153: Maintenance * 12:52 root@cumin1003: START - Cookbook sre.mysql.pool pool db2221: Maintenance * 12:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2153 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95358 and previous config saved to /var/cache/conftool/dbconfig/20260728-125214-cwilliams.json * 12:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2153.codfw.wmnet with reason: Maintenance * 12:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2167: Maintenance * 12:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-codfw * 12:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2011.codfw.wmnet * 12:51 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2011.codfw.wmnet * 12:49 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:49 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback for asw1-603 - pt1979@cumin2003" * 12:48 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback for asw1-603 - pt1979@cumin2003" * 12:46 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2011.codfw.wmnet * 12:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2221 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95357 and previous config saved to /var/cache/conftool/dbconfig/20260728-124601-cwilliams.json * 12:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2221.codfw.wmnet with reason: Maintenance * 12:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2218: Maintenance * 12:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2167 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95354 and previous config saved to /var/cache/conftool/dbconfig/20260728-124457-cwilliams.json * 12:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2167.codfw.wmnet with reason: Maintenance * 12:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2166: Maintenance * 12:42 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 12:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2011.codfw.wmnet * 12:41 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2010.codfw.wmnet * 12:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2010.codfw.wmnet * 12:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2010.codfw.wmnet * 12:34 pt1979@cumin2003: START - Cookbook sre.dns.netbox * 12:32 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 12:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2010.codfw.wmnet * 12:31 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2009.codfw.wmnet * 12:31 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2009.codfw.wmnet * 12:27 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2009.codfw.wmnet * 12:22 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2009.codfw.wmnet * 12:22 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2008.codfw.wmnet * 12:21 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2008.codfw.wmnet * 12:16 pt1979@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-604-eqsin * 12:16 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2008.codfw.wmnet * 12:16 pt1979@cumin1003: START - Cookbook sre.network.tls for network device asw1-604-eqsin * 12:14 root@cumin1003: START - Cookbook sre.mysql.pool pool db2236: Maintenance * 12:14 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2236: Maintenance * 12:12 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1137.eqiad.wmnet with OS trixie * 12:11 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2008.codfw.wmnet * 12:11 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2007.codfw.wmnet * 12:11 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2007.codfw.wmnet * 12:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db2236: Maintenance * 12:06 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2007.codfw.wmnet * 12:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2236 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95348 and previous config saved to /var/cache/conftool/dbconfig/20260728-120253-cwilliams.json * 12:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2236.codfw.wmnet with reason: Maintenance * 12:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2007.codfw.wmnet * 12:01 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 12:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 11:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2218: Maintenance * 11:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db2166: Maintenance * 11:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2219: Maintenance * 11:56 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2006.codfw.wmnet * 11:52 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1137.eqiad.wmnet with reason: host reimage * 11:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2218 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95344 and previous config saved to /var/cache/conftool/dbconfig/20260728-115155-cwilliams.json * 11:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2218.codfw.wmnet with reason: Maintenance * 11:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2208: Maintenance * 11:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2166 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95342 and previous config saved to /var/cache/conftool/dbconfig/20260728-115119-cwilliams.json * 11:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2166.codfw.wmnet with reason: Maintenance * 11:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2164: Maintenance * 11:47 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1137.eqiad.wmnet with reason: host reimage * 11:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2006.codfw.wmnet * 11:45 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2005.codfw.wmnet * 11:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2005.codfw.wmnet * 11:40 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2005.codfw.wmnet * 11:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2005.codfw.wmnet * 11:35 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2004.codfw.wmnet * 11:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2004.codfw.wmnet * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1137 * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1137 * 11:30 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1137 * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1137.eqiad.wmnet 192.32.64.10.in-addr.arpa 2.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:30 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1137.eqiad.wmnet 192.32.64.10.in-addr.arpa 2.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1137 - jiji@cumin1003" * 11:25 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2004.codfw.wmnet * 11:19 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2004.codfw.wmnet * 11:19 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2003.codfw.wmnet * 11:19 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2003.codfw.wmnet * 11:14 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2003.codfw.wmnet * 11:11 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2219: Maintenance * 11:10 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2219: Maintenance * 11:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2219: Maintenance * 11:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2003.codfw.wmnet * 11:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2164: Maintenance * 11:03 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2002.codfw.wmnet * 11:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2002.codfw.wmnet * 11:03 root@cumin1003: START - Cookbook sre.mysql.pool pool db2208: Maintenance * 10:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2164 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95332 and previous config saved to /var/cache/conftool/dbconfig/20260728-105749-cwilliams.json * 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2164.codfw.wmnet with reason: Maintenance * 10:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2208 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95331 and previous config saved to /var/cache/conftool/dbconfig/20260728-105711-cwilliams.json * 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2208.codfw.wmnet with reason: Maintenance * 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2219 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95330 and previous config saved to /var/cache/conftool/dbconfig/20260728-105652-cwilliams.json * 10:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2219.codfw.wmnet with reason: Maintenance * 10:53 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1137 - jiji@cumin1003" * 10:52 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2002.codfw.wmnet * 10:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2002.codfw.wmnet * 10:47 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2001.codfw.wmnet * 10:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2001.codfw.wmnet * 10:39 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2001.codfw.wmnet * 10:35 jiji@cumin1003: START - Cookbook sre.dns.netbox * 10:34 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] (duration: 09m 31s) * 10:34 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1137 * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2001.codfw.wmnet * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-codfw * 10:34 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1137.eqiad.wmnet with OS trixie * 10:28 jforrester@deploy1003: jforrester: Continuing with deployment * 10:27 jforrester@deploy1003: jforrester: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:25 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] * 10:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 10:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 10:21 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 10:20 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1137.eqiad.wmnet * 10:20 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 10:20 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1137.eqiad.wmnet * 10:20 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1137.eqiad.wmnet * 10:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2163: Maintenance * 09:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-staging-worker * 09:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2003.codfw.wmnet * 09:37 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2003.codfw.wmnet * 09:32 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 09:31 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2003.codfw.wmnet * 09:30 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 09:30 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2163: Maintenance * 09:30 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 09:30 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 09:30 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:22 klausman@cumin1003: END (ERROR) - Cookbook sre.ganeti.reboot-vm (exit_code=97) for VM ml-serve-ctrl2001.codfw.wmnet * 09:22 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2001.codfw.wmnet * 09:22 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d8-eqiad * 09:22 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d8-eqiad * 09:21 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2003.codfw.wmnet * 09:20 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2002.codfw.wmnet * 09:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2002.codfw.wmnet * 09:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f2-codfw * 09:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f2-codfw * 09:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e4-codfw * 09:18 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2163: Maintenance * 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e4-codfw * 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-codfw * 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-codfw * 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e5-codfw * 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e5-codfw * 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f4-codfw * 09:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2210: Maintenance * 09:16 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f4-codfw * 09:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2182: Maintenance * 09:14 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2002.codfw.wmnet * 09:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db2163: Maintenance * 09:11 XioNoX: rebooting cr2-drmrs - [[phab:T431749|T431749]] * 09:10 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr2-drmrs,cr2-drmrs IPv6,cr2-drmrs.mgmt with reason: router upgrade * 09:06 XioNoX: draining cr2-drmrs - [[phab:T431749|T431749]] * 09:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2163 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95320 and previous config saved to /var/cache/conftool/dbconfig/20260728-090638-cwilliams.json * 09:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2163.codfw.wmnet with reason: Maintenance * 09:06 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2161: Maintenance * 09:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2002.codfw.wmnet * 09:04 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2001.codfw.wmnet * 09:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2001.codfw.wmnet * 08:57 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2001.codfw.wmnet * 08:48 XioNoX: un-drain cr1-drmrs - [[phab:T431749|T431749]] * 08:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2001.codfw.wmnet * 08:47 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-staging-worker * 08:42 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:35 XioNoX: rebooting cr1-drmrs - [[phab:T431749|T431749]] * 08:33 XioNoX: draining cr1-drmrs - [[phab:T431749|T431749]] * 08:31 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2210: Maintenance * 08:29 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2182: Maintenance * 08:21 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2182: Maintenance * 08:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db2161: Maintenance * 08:16 root@cumin1003: START - Cookbook sre.mysql.pool pool db2182: Maintenance * 08:12 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2210: Maintenance * 08:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2161 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95309 and previous config saved to /var/cache/conftool/dbconfig/20260728-081044-cwilliams.json * 08:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2161.codfw.wmnet with reason: Maintenance * 08:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2154: Maintenance * 08:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2182 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95307 and previous config saved to /var/cache/conftool/dbconfig/20260728-080947-cwilliams.json * 08:09 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2182.codfw.wmnet with reason: Maintenance * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2168: Maintenance * 08:06 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr1-drmrs,cr1-drmrs IPv6,cr1-drmrs.mgmt with reason: router upgrade * 08:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db2210: Maintenance * 08:05 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 08:05 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 08:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2210 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95305 and previous config saved to /var/cache/conftool/dbconfig/20260728-080008-cwilliams.json * 08:00 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2210.codfw.wmnet with reason: Maintenance * 07:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2206: Maintenance * 07:50 gkyziridis@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 07:50 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 07:22 root@cumin1003: START - Cookbook sre.mysql.pool pool db2154: Maintenance * 07:22 root@cumin1003: START - Cookbook sre.mysql.pool pool db2168: Maintenance * 07:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2154 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95295 and previous config saved to /var/cache/conftool/dbconfig/20260728-071640-cwilliams.json * 07:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2154.codfw.wmnet with reason: Maintenance * 07:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95294 and previous config saved to /var/cache/conftool/dbconfig/20260728-071604-cwilliams.json * 07:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2168.codfw.wmnet with reason: Maintenance * 07:08 root@cumin1003: START - Cookbook sre.mysql.pool pool db2206: Maintenance * 07:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2206 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95292 and previous config saved to /var/cache/conftool/dbconfig/20260728-070219-cwilliams.json * 07:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2206.codfw.wmnet with reason: Maintenance * 06:44 marostegui: Failover m5 from db1164 to db1228 - [[phab:T432967|T432967]] * 06:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2235].codfw.wmnet,db[1164,1217,1228].eqiad.wmnet with reason: m5 master switch [[phab:T432967|T432967]] * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.10 (duration: 02m 34s) * 03:39 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] (duration: 36m 06s) * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 02:57 dzahn@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1004.eqiad.wmnet with OS trixie * 02:57 dzahn@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - dzahn@cumin1003" * 02:55 dzahn@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - dzahn@cumin1003" * 02:37 dzahn@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1004.eqiad.wmnet with reason: host reimage * 02:31 dzahn@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1004.eqiad.wmnet with reason: host reimage * 02:16 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie * 02:15 dzahn@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host zuul1004.eqiad.wmnet with OS trixie * 01:43 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie * 01:43 dzahn@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1004.eqiad.wmnet with OS trixie * 01:25 pt1979@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-603-eqsin * 01:24 pt1979@cumin1003: START - Cookbook sre.network.tls for network device asw1-603-eqsin * 01:12 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 01:12 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt for new switches in eqsin - pt1979@cumin2003" * 01:12 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt for new switches in eqsin - pt1979@cumin2003" * 01:08 pt1979@cumin2003: START - Cookbook sre.dns.netbox * 00:48 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 00:47 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 00:47 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 00:47 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 00:26 mutante: attempting reimage with trixie on zuul1004 re-purposed physical hardware - dcops reported install issue - host was in busybox shell ([[phab:T427353|T427353]]) * 00:24 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie == 2026-07-27 == * 23:50 Amir1: mass deleting vp8 transcodes * 23:28 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:27 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1004.eqiad.wmnet with OS bullseye * 23:26 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 23:25 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:25 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 22:39 maryum: Deploy security fix for [[phab:T432877|T432877]] * 22:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1022.eqiad.wmnet with OS bookworm * 22:37 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS bullseye * 22:32 sbassett: Deployed security fix for [[phab:T432789|T432789]] * 22:22 sbassett: Deployed security patch for [[phab:T431819|T431819]] * 22:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1022.eqiad.wmnet with reason: host reimage * 22:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1022.eqiad.wmnet with reason: host reimage * 22:01 RScout-WMF: Deployed security fix for [[phab:T431819|T431819]] * 22:00 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2012 * 21:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2012.codfw.wmnet with OS bookworm * 21:55 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2011\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 21:45 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1022 * 21:45 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1022 * 21:44 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1022 * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1022.eqiad.wmnet 239.48.64.10.in-addr.arpa 9.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:44 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1022.eqiad.wmnet 239.48.64.10.in-addr.arpa 9.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:41 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:41 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 21:34 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS bookworm * 21:31 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:22 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:21 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:19 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:17 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1004 * 21:16 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1004 * 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1004] - vriley@cumin1003" * 21:15 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1004] - vriley@cumin1003" * 21:11 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:10 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2011.codfw.wmnet, repooling source-only afterwards * 21:05 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:01 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1022 * 20:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1022.eqiad.wmnet with OS bookworm * 20:53 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1021.eqiad.wmnet, repooling source-only afterwards * 20:51 mutante: zuul1001 - re-enabled puppet - revert "cherry-picked" gerrit:1314120 - [[phab:T431003|T431003]] * 20:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Maintenance * 20:15 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] (duration: 08m 03s) * 20:11 sbisson@deploy1003: sbisson: Continuing with deployment * 20:09 sbisson@deploy1003: sbisson: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] * 19:47 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 46s) * 19:47 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Maintenance * 19:27 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] (duration: 12m 26s) * 19:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2228 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95285 and previous config saved to /var/cache/conftool/dbconfig/20260727-192711-cwilliams.json * 19:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2228.codfw.wmnet with reason: Maintenance * 19:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2223: Maintenance * 19:23 krinkle@deploy1003: krinkle: Continuing with deployment * 19:16 krinkle@deploy1003: krinkle: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:15 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] * 19:12 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2238: Maintenance * 18:58 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 18:57 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 18:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2227: Maintenance * 18:57 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-experimental: apply * 18:55 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-experimental: apply * 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1021.eqiad.wmnet with OS bookworm * 18:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2011.codfw.wmnet with OS bookworm * 18:40 root@cumin1003: START - Cookbook sre.mysql.pool pool db2223: Maintenance * 18:39 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] (duration: 07m 05s) * 18:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2223 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95275 and previous config saved to /var/cache/conftool/dbconfig/20260727-183500-cwilliams.json * 18:34 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2223.codfw.wmnet with reason: Maintenance * 18:34 musikanimal@deploy1003: musikanimal: Continuing with deployment * 18:34 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2213: Maintenance * 18:33 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:32 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] * 18:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db2238: Maintenance * 18:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2011.codfw.wmnet with reason: host reimage * 18:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2238 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95271 and previous config saved to /var/cache/conftool/dbconfig/20260727-181944-cwilliams.json * 18:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2238.codfw.wmnet with reason: Maintenance * 18:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2226: Maintenance * 18:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1021.eqiad.wmnet with reason: host reimage * 18:14 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2011.codfw.wmnet with reason: host reimage * 18:12 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1021.eqiad.wmnet with reason: host reimage * 18:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db2227: Maintenance * 18:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2227 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95265 and previous config saved to /var/cache/conftool/dbconfig/20260727-180256-cwilliams.json * 18:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2227.codfw.wmnet with reason: Maintenance * 18:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2194: Maintenance * 17:57 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2011 * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2011 * 17:56 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2011 * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2011.codfw.wmnet 37.32.192.10.in-addr.arpa 7.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:56 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2011.codfw.wmnet 37.32.192.10.in-addr.arpa 7.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2011 - bking@cumin2003" * 17:56 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2011 - bking@cumin2003" * 17:52 bking@cumin2003: START - Cookbook sre.dns.netbox * 17:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2011 * 17:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1021 * 17:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1021 * 17:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2011.codfw.wmnet with OS bookworm * 17:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1021.eqiad.wmnet with OS bookworm * 17:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Maintenance * 17:38 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2010\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 17:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2213 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95260 and previous config saved to /var/cache/conftool/dbconfig/20260727-173740-cwilliams.json * 17:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2213.codfw.wmnet with reason: Maintenance * 17:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2211: Maintenance * 17:36 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1020\.eqiad\.wmnet,dc=eqiad,cluster=wdqs\-main,service=wdqs\-main * 17:32 root@cumin1003: START - Cookbook sre.mysql.pool pool db2226: Maintenance * 17:31 taavi@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] (duration: 06m 33s) * 17:27 taavi@deploy1003: taavi: Continuing with deployment * 17:27 taavi@deploy1003: taavi: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:26 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2226 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95256 and previous config saved to /var/cache/conftool/dbconfig/20260727-172636-cwilliams.json * 17:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2226.codfw.wmnet with reason: Maintenance * 17:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2225: Maintenance * 17:25 taavi@deploy1003: Started scap sync-world: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] * 17:13 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 17:11 root@cumin1003: START - Cookbook sre.mysql.pool pool db2194: Maintenance * 17:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2194 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95248 and previous config saved to /var/cache/conftool/dbconfig/20260727-170453-cwilliams.json * 17:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2194.codfw.wmnet with reason: Maintenance * 17:04 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2190: Maintenance * 16:52 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 16:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2211: Maintenance * 16:40 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95242 and previous config saved to /var/cache/conftool/dbconfig/20260727-164015-cwilliams.json * 16:40 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2211.codfw.wmnet with reason: Maintenance * 16:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2178: Maintenance * 16:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db2225: Maintenance * 16:39 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2172: Maintenance * 16:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 16:38 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 16:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2225 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95238 and previous config saved to /var/cache/conftool/dbconfig/20260727-163307-cwilliams.json * 16:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2225.codfw.wmnet with reason: Maintenance * 16:32 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2189: Maintenance * 16:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db2190: Maintenance * 16:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2190 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95230 and previous config saved to /var/cache/conftool/dbconfig/20260727-160602-cwilliams.json * 16:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2190.codfw.wmnet with reason: Maintenance * 15:53 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2177: Maintenance * 15:53 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2172: Maintenance * 15:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2178: Maintenance * 15:51 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2172: Maintenance * 15:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2178 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95224 and previous config saved to /var/cache/conftool/dbconfig/20260727-154559-cwilliams.json * 15:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2172: Maintenance * 15:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2178.codfw.wmnet with reason: Maintenance * 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2171: Maintenance * 15:44 root@cumin1003: START - Cookbook sre.mysql.pool pool db2189: Maintenance * 15:43 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:41 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2172 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95222 and previous config saved to /var/cache/conftool/dbconfig/20260727-153927-cwilliams.json * 15:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2172.codfw.wmnet with reason: Maintenance * 15:38 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2189 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95220 and previous config saved to /var/cache/conftool/dbconfig/20260727-153833-cwilliams.json * 15:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2189.codfw.wmnet with reason: Maintenance * 15:34 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:32 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 15:32 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 15:31 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:29 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:26 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 15:22 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] (duration: 07m 00s) * 15:21 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2155: Maintenance * 15:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2175: Maintenance * 15:18 zabe@deploy1003: zabe: Continuing with deployment * 15:17 zabe@deploy1003: zabe: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:15 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 15:15 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:15 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] * 15:15 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 06s) * 15:15 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:12 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2010.codfw.wmnet with OS bookworm * 15:08 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2177: Maintenance * 15:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2177: Maintenance * 14:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2177: Maintenance * 14:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2171: Maintenance * 14:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2171 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95209 and previous config saved to /var/cache/conftool/dbconfig/20260727-145236-cwilliams.json * 14:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2171.codfw.wmnet with reason: Maintenance * 14:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2177 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95208 and previous config saved to /var/cache/conftool/dbconfig/20260727-145206-cwilliams.json * 14:52 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2157: Maintenance * 14:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2177.codfw.wmnet with reason: Maintenance * 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1020.eqiad.wmnet with OS bookworm * 14:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2156: Maintenance * 14:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2010.codfw.wmnet with reason: host reimage * 14:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2010.codfw.wmnet with reason: host reimage * 14:41 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[1031,2024]*: Upgrade Cassandra to 5.0.8 (canary) - eevans@cumin1003 * 14:34 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2155: Maintenance * 14:33 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2175: Maintenance * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2010 * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2010 * 14:24 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2010 * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2010.codfw.wmnet 94.16.192.10.in-addr.arpa 4.9.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2010.codfw.wmnet 94.16.192.10.in-addr.arpa 4.9.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2010 - bking@cumin2003" * 14:24 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2010 - bking@cumin2003" * 14:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1020.eqiad.wmnet with reason: host reimage * 14:23 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[1031,2024]*: Upgrade Cassandra to 5.0.8 (canary) - eevans@cumin1003 * 14:20 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 14:20 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 14:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1020.eqiad.wmnet with reason: host reimage * 14:17 sukhe: sudo gnt-instance reboot urldownloader1005.wikimedia.org * 14:16 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:15 jelto@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:08 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2155: Maintenance * 14:05 root@cumin1003: START - Cookbook sre.mysql.pool pool db2157: Maintenance * 14:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2175: Maintenance * 14:03 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 11 hosts * 14:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db2155: Maintenance * 14:01 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 11 hosts * 14:01 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1136.eqiad.wmnet * 14:01 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1136.eqiad.wmnet * 14:01 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1136.eqiad.wmnet * 14:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db2156: Maintenance * 13:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2157 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95194 and previous config saved to /var/cache/conftool/dbconfig/20260727-135943-cwilliams.json * 13:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2157.codfw.wmnet with reason: Maintenance * 13:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db2175: Maintenance * 13:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 13:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 13:57 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2010 * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2155 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95193 and previous config saved to /var/cache/conftool/dbconfig/20260727-135613-cwilliams.json * 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2155.codfw.wmnet with reason: Maintenance * 13:55 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1020 * 13:55 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1020 * 13:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2156 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95192 and previous config saved to /var/cache/conftool/dbconfig/20260727-135413-cwilliams.json * 13:54 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2156.codfw.wmnet with reason: Maintenance * 13:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2175 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95191 and previous config saved to /var/cache/conftool/dbconfig/20260727-135300-cwilliams.json * 13:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2010.codfw.wmnet with OS bookworm * 13:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2175.codfw.wmnet with reason: Maintenance * 13:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1020.eqiad.wmnet with OS bookworm * 13:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 34 hosts * 13:46 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 34 hosts * 13:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1201: Maintenance * 13:27 Lucas_WMDE: UTC afternoon backport+config window doen * 13:18 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] (duration: 11m 57s) * 13:14 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, sihe: Continuing with deployment * 13:08 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, sihe: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:07 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool ulsfo [reason: router upgrade finished, [[phab:T431752|T431752]]] * 13:07 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool ulsfo [reason: router upgrade finished, [[phab:T431752|T431752]]] * 13:06 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] * 13:03 XioNoX: repool cr4-ulsfo - [[phab:T431752|T431752]] * 12:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db1201: Maintenance * 12:48 gkyziridis@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1201 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95186 and previous config saved to /var/cache/conftool/dbconfig/20260727-124404-cwilliams.json * 12:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1201.eqiad.wmnet with reason: Maintenance * 12:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1187: Maintenance * 12:30 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader.eqiad.wikimedia.org on all recursors * 12:30 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader.eqiad.wikimedia.org on all recursors * 12:30 sukhe@dns1004: END - running authdns-update * 12:30 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] (duration: 09m 32s) * 12:28 sukhe@dns1004: START - running authdns-update * 12:25 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 12:22 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:20 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] * 12:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts * 12:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts * 12:13 XioNoX: rebooting cr4-ulsfo for upgrade - [[phab:T431752|T431752]] * 12:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: es1038 repool * 12:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 38 hosts * 12:08 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 38 hosts * 11:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1187: Maintenance * 11:53 urbanecm@deploy1003: mwscript-k8s job started: foreachwikiindblist growthexperiments GrowthExperiments:cleanMentorList # [[phab:T431804|T431804]] * 11:50 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr4-ulsfo,cr4-ulsfo IPv6,cr4-ulsfo.mgmt with reason: router upgrade * 11:50 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] (duration: 11m 07s) * 11:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1187 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95178 and previous config saved to /var/cache/conftool/dbconfig/20260727-114844-cwilliams.json * 11:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1187.eqiad.wmnet with reason: Maintenance * 11:43 urbanecm@deploy1003: urbanecm: Continuing with deployment * 11:42 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:39 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] * 11:37 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 11:36 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 11:36 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 11:35 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 11:29 XioNoX: start draining cr4-ulsfo - [[phab:T431752|T431752]] * 11:29 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 11:29 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 11:28 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1035: testing * 11:28 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1035: testing * 11:27 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1035: testing * 11:27 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1035: testing * 11:26 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool ulsfo [reason: router upgrade, [[phab:T431752|T431752]]] * 11:26 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1038: es1038 repool * 11:26 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool ulsfo [reason: router upgrade, [[phab:T431752|T431752]]] * 11:26 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1038: testing * 11:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1264: Maintenance * 11:24 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1038: testing * 11:23 marostegui@cumin1003: dbctl commit (dc=all): 'Repool es1050 as master', diff saved to https://phabricator.wikimedia.org/P95170 and previous config saved to /var/cache/conftool/dbconfig/20260727-112326-marostegui.json * 11:23 marostegui@cumin1003: dbctl commit (dc=all): 'Repool es1050', diff saved to https://phabricator.wikimedia.org/P95169 and previous config saved to /var/cache/conftool/dbconfig/20260727-112302-marostegui.json * 11:22 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1050: testing * 11:22 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1050: testing * 11:20 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 11:18 blake@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 11:18 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 11:12 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 11:11 blake@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 11:09 blake@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 11:09 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 11:09 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 11:08 blake@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 11:05 blake@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 11:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 11:02 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 10:50 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 10:43 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:39 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply * 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1264: Maintenance * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply * 10:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply * 10:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 10:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 10:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 10:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1264 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95164 and previous config saved to /var/cache/conftool/dbconfig/20260727-103204-cwilliams.json * 10:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1264.eqiad.wmnet with reason: Maintenance * 10:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 10:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 10:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 10:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 10:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 10:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 10:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 10:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1237: Maintenance * 10:04 elukey: restart burrow main-eqiad on kafkamon2003 to clear some errors on kafka-main1008 * 09:58 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1136.eqiad.wmnet with OS trixie * 09:39 elukey: restart burrow-main-eqiad.service on kafkamon1003 to see if a recurrent kafka error on kafka-main1008 goes away * 09:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1237: Maintenance * 09:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1136.eqiad.wmnet with reason: host reimage * 09:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1237 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95159 and previous config saved to /var/cache/conftool/dbconfig/20260727-093328-cwilliams.json * 09:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1237.eqiad.wmnet with reason: Maintenance * 09:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1136.eqiad.wmnet with reason: host reimage * 09:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1203: Maintenance * 09:17 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1136 * 09:17 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1136 * 09:04 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1136 * 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1136.eqiad.wmnet 191.32.64.10.in-addr.arpa 1.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:04 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1136.eqiad.wmnet 191.32.64.10.in-addr.arpa 1.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1136 - jiji@cumin1003" * 09:04 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1136 - jiji@cumin1003" * 08:52 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 08:52 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 08:52 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 08:51 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 08:50 jiji@cumin1003: START - Cookbook sre.dns.netbox * 08:47 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1136 * 08:46 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1136.eqiad.wmnet with OS trixie * 08:44 marostegui: Rename tables on s3 [[phab:T425066|T425066]] * 08:43 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1136.eqiad.wmnet * 08:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db1203: Maintenance * 08:43 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1136.eqiad.wmnet * 08:43 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1136.eqiad.wmnet * 08:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1203 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95154 and previous config saved to /var/cache/conftool/dbconfig/20260727-083703-cwilliams.json * 08:36 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1203.eqiad.wmnet with reason: Maintenance * 08:16 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1179: Maintenance * 07:44 phuedx: UTC morning backport window done * 07:37 phuedx@deploy1003: Finished scap sync-world: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] (duration: 32m 33s) * 07:28 root@cumin1003: START - Cookbook sre.mysql.pool pool db1179: Maintenance * 07:26 marostegui: Rename tables on s3 [[phab:T426341|T426341]] * 07:25 phuedx@deploy1003: phuedx: Continuing with deployment * 07:22 marostegui: Drop tables in akwiki nawiki pihwiki - growthexperiments_* [[phab:T428885|T428885]] * 07:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95149 and previous config saved to /var/cache/conftool/dbconfig/20260727-072234-cwilliams.json * 07:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1179.eqiad.wmnet with reason: Maintenance * 07:20 phuedx@deploy1003: phuedx: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:16 ryankemper: [[phab:T430880|T430880]] [WDQS] Reimaged `wdqs1018` and `wdqs1019` to Bookworm, restored data using test-cookbook change {{Gerrit|1317128}}, and repooled both; 25/36 hosts complete * 07:04 phuedx@deploy1003: Started scap sync-world: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] * 06:57 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1019.eqiad.wmnet * 06:56 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1018.eqiad.wmnet * 06:40 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1020.eqiad.wmnet with reason: Cloning * 06:35 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db1228.eqiad.wmnet with reason: Rebooting * 06:29 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:29 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:25 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:25 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:25 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1019.eqiad.wmnet, repooling source-only afterwards * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1018.eqiad.wmnet, repooling source-only afterwards * 04:51 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1019.eqiad.wmnet, repooling source-only afterwards * 04:51 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1018.eqiad.wmnet, repooling source-only afterwards * 04:48 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s) * 04:48 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 04:48 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s) * 04:48 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 36s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-26 == * 14:59 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:59 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:59 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:59 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1019.eqiad.wmnet with OS bookworm * 01:05 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1018.eqiad.wmnet with OS bookworm * 00:43 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1019.eqiad.wmnet with reason: host reimage * 00:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1018.eqiad.wmnet with reason: host reimage * 00:34 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1019.eqiad.wmnet with reason: host reimage * 00:33 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1018.eqiad.wmnet with reason: host reimage * 00:16 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 00:16 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 00:15 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 00:15 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1019 * 00:11 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1019 * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1018 * 00:11 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1018 * 00:08 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1019.eqiad.wmnet with OS bookworm * 00:08 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1018.eqiad.wmnet with OS bookworm == 2026-07-25 == * 22:06 ryankemper: [[phab:T430880|T430880]] [WDQS] Repooled `wdqs1017` and `wdqs2024` after reimaging to bookworm, scap deploying, and data xfering * 22:04 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2024.codfw.wmnet * 22:03 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1017.eqiad.wmnet * 21:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1017.eqiad.wmnet, repooling source-only afterwards * 21:06 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2024.codfw.wmnet, repooling source-only afterwards * 20:52 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:52 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:52 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:52 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 20:18 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1017.eqiad.wmnet, repooling source-only afterwards * 20:18 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2024.codfw.wmnet, repooling source-only afterwards * 20:15 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:15 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:15 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:15 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 19:57 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s) * 19:57 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 19:57 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 07s) * 19:57 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 19:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2024.codfw.wmnet with OS bookworm * 19:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1017.eqiad.wmnet with OS bookworm * 19:02 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2024.codfw.wmnet with reason: host reimage * 18:58 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1017.eqiad.wmnet with reason: host reimage * 18:53 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2024.codfw.wmnet with reason: host reimage * 18:52 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1017.eqiad.wmnet with reason: host reimage * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2024 * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2024 * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1017 * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1017 * 18:27 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2024 * 18:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2024.codfw.wmnet 58.16.192.10.in-addr.arpa 8.5.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:26 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2024.codfw.wmnet 58.16.192.10.in-addr.arpa 8.5.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:24 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1017 * 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1017.eqiad.wmnet 238.48.64.10.in-addr.arpa 8.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:24 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1017.eqiad.wmnet 238.48.64.10.in-addr.arpa 8.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1017 - ryankemper@cumin2003" * 18:24 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1017 - ryankemper@cumin2003" * 18:23 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 18:18 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 18:17 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1017 * 18:17 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2024 * 18:14 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1017.eqiad.wmnet with OS bookworm * 18:14 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2024.codfw.wmnet with OS bookworm * 18:05 ryankemper: [WDQS] [[phab:T430880|T430880]] Reimaged `wdqs1016` and `wdqs2023` to Bookworm with `--move-vlan`, restored main and scholarly data, validated postflights, and repooled both hosts. Confirmed PyBal rebuilt both backends with their new addresses * 17:45 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2023.codfw.wmnet * 17:43 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1016.eqiad.wmnet * 06:35 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1016.eqiad.wmnet, repooling source-only afterwards * 06:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2023.codfw.wmnet, repooling source-only afterwards * 05:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2023.codfw.wmnet, repooling source-only afterwards * 05:19 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1016.eqiad.wmnet, repooling source-only afterwards * 05:07 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 07s) * 05:07 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 05:06 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 06s) * 05:06 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 03:27 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2023.codfw.wmnet with OS bookworm * 02:59 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2023.codfw.wmnet with reason: host reimage * 02:56 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2023.codfw.wmnet with reason: host reimage * 02:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2023 * 02:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2023 * 02:30 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2023 * 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2023.codfw.wmnet 35.0.192.10.in-addr.arpa 5.3.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 02:30 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2023.codfw.wmnet 35.0.192.10.in-addr.arpa 5.3.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2023 - ryankemper@cumin2003" * 02:30 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2023 - ryankemper@cumin2003" * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 26s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:15 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1016.eqiad.wmnet with OS bookworm * 00:49 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1016.eqiad.wmnet with reason: host reimage * 00:43 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1016.eqiad.wmnet with reason: host reimage * 00:31 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 00:27 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1016 * 00:27 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1016 * 00:27 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2023 * 00:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1016.eqiad.wmnet with OS bookworm * 00:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2023.codfw.wmnet with OS bookworm * 00:11 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs1014.eqiad.wmnet and wdqs2008.codfw.wmnet after Bookworm reimage, transfer, and postflight; wdqs2008 is serving, while wdqs1014 will remain outside of service until a pybal restart next monday == 2026-07-24 == * 23:54 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1014.eqiad.wmnet * 23:54 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2008.codfw.wmnet * 23:43 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2010.codfw.wmnet with OS trixie * 23:08 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 23:03 jhathaway@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 22:33 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 22:13 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 22:13 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 22:13 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 22:13 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:00 jhathaway@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 21:53 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 21:53 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie * 21:51 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 21:47 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie * 21:43 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 21:39 jhathaway@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 21:38 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 17:21 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1135.eqiad.wmnet * 17:21 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1135.eqiad.wmnet * 17:21 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1135.eqiad.wmnet * 16:34 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 16:34 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 16:34 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 16:34 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 16:33 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 16:33 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 16:28 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:28 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:28 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:28 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2008.codfw.wmnet, repooling source-only afterwards * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1014.eqiad.wmnet, repooling source-only afterwards * 15:56 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1135.eqiad.wmnet with OS trixie * 15:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 40 hosts * 15:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 40 hosts * 15:37 topranks: upgrade SR-Linux OS on lswtest-d8-eqiad * 15:36 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1135.eqiad.wmnet with reason: host reimage * 15:33 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 6 hosts with reason: upgrade lswtest-d8-eqiad * 15:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1135.eqiad.wmnet with reason: host reimage * 15:30 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc-gp2006.codfw.wmnet with OS bookworm * 15:15 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1135 * 15:15 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1135 * 15:13 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc-gp2006.codfw.wmnet with reason: host reimage * 15:08 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc-gp2006.codfw.wmnet with reason: host reimage * 14:49 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm * 14:48 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host mc-gp2006.codfw.wmnet with OS bookworm * 14:34 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] (duration: 41m 12s) * 14:32 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1135 * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1135.eqiad.wmnet 177.32.64.10.in-addr.arpa 7.7.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:32 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1135.eqiad.wmnet 177.32.64.10.in-addr.arpa 7.7.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1135 - jiji@cumin1003" * 14:32 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1135 - jiji@cumin1003" * 14:29 krinkle@deploy1003: krinkle: Continuing with deployment * 14:29 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm * 14:27 jiji@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host mc-gp2006.codfw.wmnet with OS bookworm * 14:26 jiji@cumin1003: START - Cookbook sre.dns.netbox * 14:15 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1135 * 14:14 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1135.eqiad.wmnet with OS trixie * 14:14 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1135.eqiad.wmnet * 14:13 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1135.eqiad.wmnet * 14:13 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1135.eqiad.wmnet * 14:10 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1072.eqiad.wmnet * 14:10 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1072.eqiad.wmnet * 14:10 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1072.eqiad.wmnet * 14:10 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1072.eqiad.wmnet * 14:09 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1071.eqiad.wmnet * 14:09 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1071.eqiad.wmnet * 14:09 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1071.eqiad.wmnet * 14:09 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1071.eqiad.wmnet * 13:58 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 13:58 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:58 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:57 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:55 krinkle@deploy1003: krinkle: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:53 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] * 13:45 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:45 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push new IPs for mc-gp2006 - cmooney@cumin1003" * 13:45 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push new IPs for mc-gp2006 - cmooney@cumin1003" * 13:44 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) mc-gp2006.codfw.wmnet on all recursors * 13:44 cmooney@cumin1003: START - Cookbook sre.dns.wipe-cache mc-gp2006.codfw.wmnet on all recursors * 13:42 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm * 13:41 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:30 papaul: reboot mr1-eqsin for maintenance * 13:24 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb[1029-1031].eqiad.wmnet * 13:10 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb[1029-1031].eqiad.wmnet * 11:33 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 7 hosts * 11:11 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 7 hosts * 10:56 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 7 hosts * 10:47 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 7 hosts * 10:44 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:44 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:41 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:41 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:35 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 8 hosts * 10:34 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:33 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:32 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:32 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:31 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:31 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:30 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 8 hosts * 10:24 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie * 10:19 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:18 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 16 hosts * 10:17 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2001.codfw.wmnet * 10:13 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2001.codfw.wmnet * 10:12 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2001.codfw.wmnet * 10:02 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2001.codfw.wmnet * 10:02 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2002.codfw.wmnet * 09:57 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2002.codfw.wmnet * 09:56 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2002.codfw.wmnet * 09:51 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2002.codfw.wmnet * 09:51 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1002.eqiad.wmnet * 09:47 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1002.eqiad.wmnet * 09:47 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1001.eqiad.wmnet * 09:44 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1001.eqiad.wmnet * 09:34 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2003.codfw.wmnet * 09:32 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2003.codfw.wmnet * 09:32 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2002.codfw.wmnet * 09:30 brouberol@dns1004: END - running authdns-update * 09:29 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2002.codfw.wmnet * 09:29 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2001.codfw.wmnet * 09:27 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 9 hosts * 09:27 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2001.codfw.wmnet * 09:27 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2001.codfw.wmnet * 09:26 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 9 hosts * 09:26 brouberol@dns1004: START - running authdns-update * 09:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 57 hosts * 09:24 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2001.codfw.wmnet * 09:24 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2002.codfw.wmnet * 09:22 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2002.codfw.wmnet * 09:21 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 57 hosts * 09:20 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2003.codfw.wmnet * 09:19 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 16 hosts * 09:16 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2003.codfw.wmnet * 09:16 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1003.eqiad.wmnet * 09:15 urbanecm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 09:15 urbanecm@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 09:13 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1003.eqiad.wmnet * 09:13 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1002.eqiad.wmnet * 09:11 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1002.eqiad.wmnet * 09:11 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1001.eqiad.wmnet * 09:07 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1001.eqiad.wmnet * 08:32 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:24 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:16 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 08:16 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 08:07 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:07 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:07 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 08:02 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 08:01 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:59 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:57 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 07:57 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 06:46 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1025.eqiad.wmnet with reason: Cloning * 06:46 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s4 * 06:45 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s6 * 06:44 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1019.eqiad.wmnet,service=s6 * 06:44 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1019.eqiad.wmnet,service=s4 * 03:40 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:40 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:40 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:40 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 03:37 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:37 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:37 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:36 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:49 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on mr1-eqsin,mr1-eqsin IPv6 with reason: connection issue * 02:38 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on cr[2-3]-eqsin.mgmt,ps1-[603-604]-eqsin with reason: connection issue * 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 27s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-23 == * 23:27 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin.oob,mr1-eqsin.oob IPv6 with reason: switch refresh * 22:21 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Setting storage compatibility to NONE - eevans@cumin1003 * 22:01 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Setting storage compatibility to NONE - eevans@cumin1003 * 21:29 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1014.eqiad.wmnet, repooling source-only afterwards * 21:28 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 46s) * 21:28 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 21:19 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Setting storage compatibility to UPGRADING - eevans@cumin1003 * 21:00 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Setting storage compatibility to UPGRADING - eevans@cumin1003 * 20:17 dani@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] (duration: 11m 57s) * 20:13 dani@deploy1003: dani: Continuing with deployment * 20:07 dani@deploy1003: dani: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:05 dani@deploy1003: Started scap sync-world: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] * 19:24 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:24 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating the rest of the ipv6 dns records. - jhancock@cumin2002" * 19:24 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating the rest of the ipv6 dns records. - jhancock@cumin2002" * 19:14 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 19:05 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wdqs1014.eqiad.wmnet with OS bookworm * 19:04 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.noop (exit_code=99) * 19:04 cwilliams@cumin1003: START - Cookbook sre.mysql.noop * 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1014.eqiad.wmnet with reason: host reimage * 18:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1014.eqiad.wmnet with reason: host reimage * 18:30 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2008.codfw.wmnet, repooling source-only afterwards * 18:28 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 19s) * 18:28 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1014 * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1014 * 18:22 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1014 * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1014.eqiad.wmnet 188.32.64.10.in-addr.arpa 8.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:22 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1014.eqiad.wmnet 188.32.64.10.in-addr.arpa 8.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1014 - bking@cumin2003" * 18:21 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1014 - bking@cumin2003" * 18:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2215: Maintenance * 18:18 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 18:15 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:15 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 18:06 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:06 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 18:05 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 18:04 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2052: codfw rack B8 re-pool after maintenance * 17:54 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 17:54 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:54 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 17:32 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2215: Maintenance * 17:29 cmooney@dns3003: END - running authdns-update * 17:27 cmooney@dns3003: START - running authdns-update * 17:23 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 17:22 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:18 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool es2052: codfw rack B8 re-pool after maintenance * 17:18 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2189: codfw rack B8 re-pool after maintenance * 17:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2215.codfw.wmnet with reason: Maintenance * 17:17 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 17:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2215 [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95126 and previous config saved to /var/cache/conftool/dbconfig/20260723-170903-cwilliams.json * 17:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2191 to x1 primary [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95125 and previous config saved to /var/cache/conftool/dbconfig/20260723-170612-cwilliams.json * 17:05 cezmunsta: Starting x1 codfw failover from db2215 to db2191 - [[phab:T432986|T432986]] * 16:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2191 with weight 0 [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95123 and previous config saved to /var/cache/conftool/dbconfig/20260723-165831-cwilliams.json * 16:58 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 16 hosts with reason: Primary switchover x1 [[phab:T432986|T432986]] * 16:36 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 138128 * 16:35 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 138128 * 16:33 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2189: codfw rack B8 re-pool after maintenance * 16:33 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2164: codfw rack B8 re-pool after maintenance * 16:28 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1072.eqiad.wmnet * 16:27 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1072.eqiad.wmnet with OS trixie * 16:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2249: Maintenance * 16:06 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker1072.eqiad.wmnet with reason: host reimage * 16:06 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1072.eqiad.wmnet with reason: host reimage * 15:50 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1072 * 15:50 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1072 * 15:49 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1072.eqiad.wmnet with OS trixie * 15:48 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2164: codfw rack B8 re-pool after maintenance * 15:48 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] (duration: 06m 37s) * 15:48 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2163: codfw rack B8 re-pool after maintenance * 15:45 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 15:43 musikanimal@deploy1003: musikanimal: Continuing with deployment * 15:43 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:41 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] * 15:36 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1072.eqiad.wmnet * 15:35 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1072.eqiad.wmnet * 15:35 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1072.eqiad.wmnet * 15:34 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:34 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push any outstanding updates - cmooney@cumin1003" * 15:34 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push any outstanding updates - cmooney@cumin1003" * 15:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db2249: Maintenance * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 15:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:26 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:21 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:21 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 15:21 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:21 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 15:20 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 15:19 cmooney@dns2004: END - running authdns-update * 15:17 cmooney@dns2004: START - running authdns-update * 15:14 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns2004.wikimedia.org * 15:12 brouberol@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 15:12 brouberol@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 15:12 klausman@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ml-serve1001.eqiad.wmnet with OS trixie * 15:11 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1071.eqiad.wmnet * 15:11 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1071.eqiad.wmnet with OS trixie * 15:10 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wdqs2008.codfw.wmnet with OS bookworm * 15:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2249.codfw.wmnet with reason: Maintenance * 15:08 brouberol@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 15:08 brouberol@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 15:08 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 15:08 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 15:06 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns1004.wikimedia.org * 15:02 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2002.codfw.wmnet * 15:02 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2002.codfw.wmnet * 15:02 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2163: codfw rack B8 re-pool after maintenance * 15:01 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 15:01 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 14:59 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test2001.codfw.wmnet * 14:57 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test2001.codfw.wmnet * 14:56 ryankemper: [WDQS] [[phab:T430880|T430880]] Reimaged `wdqs2016` to Bookworm, xferred scholarly_articles from `wdqs2024`, validated updater/readiness/federation, and repooled * 14:51 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2016.codfw.wmnet * 14:51 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1001.eqiad.wmnet with reason: host reimage * 14:48 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1071.eqiad.wmnet with reason: host reimage * 14:47 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1001.eqiad.wmnet with reason: host reimage * 14:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2008.codfw.wmnet with reason: host reimage * 14:43 topranks: reboot lsw1-b8-codw to upgrade JunOS [[phab:T430929|T430929]] * 14:41 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1071.eqiad.wmnet with reason: host reimage * 14:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2008.codfw.wmnet with reason: host reimage * 14:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2231: Maintenance * 14:30 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1001.eqiad.wmnet with OS trixie * 14:25 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2002.codfw.wmnet * 14:23 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1071 * 14:23 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1071 * 14:23 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 14:22 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore scholarly data after Bookworm reimage) xfer scholarly_articles from wdqs2024.codfw.wmnet -> wdqs2016.codfw.wmnet, repooling source-only afterwards * 14:22 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2052: codfw rack B8 depool for maintenance * 14:21 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool es2052: codfw rack B8 depool for maintenance * 14:21 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2249: codfw rack B8 depool for maintenance * 14:21 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1071 * 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1071.eqiad.wmnet 166.48.64.10.in-addr.arpa 6.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:21 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1071.eqiad.wmnet 166.48.64.10.in-addr.arpa 6.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1071 - jiji@cumin1003" * 14:21 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1071 - jiji@cumin1003" * 14:21 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2249: codfw rack B8 depool for maintenance * 14:21 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2189: codfw rack B8 depool for maintenance * 14:20 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2002.codfw.wmnet * 14:20 cmooney@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2050.codfw.wmnet * 14:20 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2189: codfw rack B8 depool for maintenance * 14:20 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2164: codfw rack B8 depool for maintenance * 14:20 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2164: codfw rack B8 depool for maintenance * 14:19 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2163: codfw rack B8 depool for maintenance * 14:19 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:19 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1014 * 14:19 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2163: codfw rack B8 depool for maintenance * 14:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1014.eqiad.wmnet with OS bookworm * 14:17 cmooney@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2050.codfw.wmnet * 14:16 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 14:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2008 * 14:14 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2008 * 14:14 cmooney@cumin1003: conftool action : set/pooled=no; selector: name=dns2004.wikimedia.org * 14:14 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2008.codfw.wmnet with OS bookworm * 14:13 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 14:12 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:10 topranks: depool dns2004 before lsw1-b8-codfw switch maintenance [[phab:T430929|T430929]] * 14:10 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b8-codfw,lsw1-b8-codfw IPv6,lsw1-b8-codfw.mgmt,ssw1-a[1,8]-codfw with reason: lsw1-b8-codfw JunOS upgrade * 14:07 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 30 hosts with reason: lsw1-b8-codfw JunOS upgrade * 14:06 elukey: upload python3-docker-report 0.0.19 to apt.wikimedia.org for bookworm and trixie * 13:59 jiji@cumin1003: START - Cookbook sre.dns.netbox * 13:58 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1071 * 13:57 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:57 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:57 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:57 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1071.eqiad.wmnet with OS trixie * 13:55 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1071.eqiad.wmnet * 13:55 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1071.eqiad.wmnet * 13:55 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1071.eqiad.wmnet * 13:53 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:52 logmsgbot: kharlan Deployed security patch for [[phab:T432948|T432948]] * 13:51 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:51 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 13:51 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db2231: Maintenance * 13:50 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:50 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:50 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:50 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:49 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:49 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:49 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2231 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95097 and previous config saved to /var/cache/conftool/dbconfig/20260723-134436-cwilliams.json * 13:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2231.codfw.wmnet with reason: Maintenance * 13:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:39 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:38 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 13:38 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] (duration: 09m 07s) * 13:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1037 hosts * 13:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2196: Maintenance * 13:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:33 kharlan@deploy1003: kharlan, emc-wmf: Continuing with deployment * 13:31 kharlan@deploy1003: kharlan, emc-wmf: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:30 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:28 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] * 13:17 hashar@deploy1003: Finished deploy [integration/docroot@2199146]: build: License GPL2.0+ / updating npm dependencies (duration: 00m 14s) * 13:17 hashar@deploy1003: Started deploy [integration/docroot@2199146]: build: License GPL2.0+ / updating npm dependencies * 13:14 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service * 13:07 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 12:58 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2207: Repooling * 12:49 root@cumin1003: START - Cookbook sre.mysql.pool pool db2196: Maintenance * 12:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2196 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95087 and previous config saved to /var/cache/conftool/dbconfig/20260723-123952-cwilliams.json * 12:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2196.codfw.wmnet with reason: Maintenance * 12:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2191: Maintenance * 12:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:13 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:13 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: Repooling * 12:12 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2207: Repooling * 12:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: Repooling * 11:56 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2235.codfw.wmnet with OS trixie * 11:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db2191: Maintenance * 11:46 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1070.eqiad.wmnet * 11:46 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1070.eqiad.wmnet * 11:46 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1070.eqiad.wmnet * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2191 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95080 and previous config saved to /var/cache/conftool/dbconfig/20260723-114308-cwilliams.json * 11:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2191.codfw.wmnet with reason: Maintenance * 11:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2186: Maintenance * 11:35 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 46375 * 11:34 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 46375 * 11:33 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2235.codfw.wmnet with reason: host reimage * 11:28 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2235.codfw.wmnet with reason: host reimage * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c7-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c7-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c6-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c6-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c5-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c5-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c4-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c4-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c3-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c3-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c2-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c2-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d7-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d7-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d4-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d4-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d3-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d2-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d2-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d8-eqiad * 11:23 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d8-eqiad * 11:23 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d1-eqiad * 11:23 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d1-eqiad * 11:12 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db2235.codfw.wmnet with OS trixie * 11:11 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:11 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[2160,2235].codfw.wmnet with reason: Upgrading * 10:56 root@cumin1003: START - Cookbook sre.mysql.pool pool db2186: Maintenance * 10:54 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1037: testing * 10:53 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1037: testing * 10:53 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1037: testing * 10:53 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1037: testing * 10:52 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: testing * 10:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2186 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95072 and previous config saved to /var/cache/conftool/dbconfig/20260723-104956-cwilliams.json * 10:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2186.codfw.wmnet with reason: Maintenance * 10:43 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1054: testing * 10:41 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1070.eqiad.wmnet with OS trixie * 10:30 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1037 hosts * 10:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 10:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 10:20 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1070.eqiad.wmnet with reason: host reimage * 10:16 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1070.eqiad.wmnet with reason: host reimage * 10:06 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1038: testing * 10:05 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1038: testing * 10:05 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1038: testing * 10:02 marostegui@dns1004: END - running authdns-update * 10:00 marostegui@dns1004: START - running authdns-update * 09:58 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: testing * 09:57 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1054: testing * 09:57 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1070 * 09:57 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1070 * 09:57 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1054: testing * 09:57 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1054: testing * 09:56 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1055: testing * 09:56 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1070 * 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1070.eqiad.wmnet 165.48.64.10.in-addr.arpa 5.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:56 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1070.eqiad.wmnet 165.48.64.10.in-addr.arpa 5.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1070 - jiji@cumin1003" * 09:56 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1070 - jiji@cumin1003" * 09:47 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for 1035 hosts * 09:45 jiji@cumin1003: START - Cookbook sre.dns.netbox * 09:42 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1070 * 09:42 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1070.eqiad.wmnet with OS trixie * 09:42 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1070.eqiad.wmnet * 09:41 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1070.eqiad.wmnet * 09:41 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1070.eqiad.wmnet * 09:27 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es2051: testing * 09:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: testing * 09:12 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es2051: testing * 09:11 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1055: testing * 09:09 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:09 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1055: testing * 09:09 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1055: testing * 08:50 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:50 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:50 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 08:49 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 08:49 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 08:49 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:46 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1069.eqiad.wmnet * 08:46 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1069.eqiad.wmnet * 08:46 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1069.eqiad.wmnet * 08:39 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 08:38 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 08:38 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 08:37 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 08:35 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:10 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1069.eqiad.wmnet with OS trixie * 07:49 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1069.eqiad.wmnet with reason: host reimage * 07:45 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1069.eqiad.wmnet with reason: host reimage * 07:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1035 hosts * 07:33 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 07:32 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 07:29 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1069 * 07:29 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1069 * 07:26 jiji@deploy1003: Finished scap sync-world: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules (duration: 06m 01s) * 07:25 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1069 * 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1069.eqiad.wmnet 164.48.64.10.in-addr.arpa 4.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:25 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1069.eqiad.wmnet 164.48.64.10.in-addr.arpa 4.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1069 - jiji@cumin1003" * 07:25 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1069 - jiji@cumin1003" * 07:25 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 07:25 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 07:24 jiji@deploy1003: jiji: Continuing with deployment * 07:22 jiji@deploy1003: jiji: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:21 jiji@deploy1003: Started scap sync-world: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules * 07:21 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1031.eqiad.wmnet,service=s7 * 07:20 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1031.eqiad.wmnet,service=s2 * 07:20 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1031.eqiad.wmnet,service=s7 * 07:20 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1031.eqiad.wmnet,service=s2 * 07:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts * 07:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts * 07:19 jiji@cumin1003: START - Cookbook sre.dns.netbox * 07:19 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1069 * 07:19 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1069.eqiad.wmnet with OS trixie * 07:19 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1069.eqiad.wmnet * 07:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 45 hosts * 07:17 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1069.eqiad.wmnet * 07:17 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1069.eqiad.wmnet * 07:14 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 45 hosts * 07:13 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 06:16 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs2007 after successful Bookworm reimage, data transfer, and postflight validation; wdqs1013 also passed postflights and is enabled in conftool, but remains out of IPVS pending a rolling pybal restart to clear its stale pre-VLAN-move address * 05:58 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2007.codfw.wmnet * 05:58 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1013.eqiad.wmnet * 05:54 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore scholarly data after Bookworm reimage) xfer scholarly_articles from wdqs2024.codfw.wmnet -> wdqs2016.codfw.wmnet, repooling source-only afterwards == 2026-07-22 == * 23:34 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Apply upgrade to JVM17 - eevans@cumin1003 * 23:14 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Apply upgrade to JVM17 - eevans@cumin1003 * 22:06 ryankemper: [WDQS] Added requestctl per-IP ratelimit `wdqs_heavy_sparql_bots_jul_2026_ratelimit` (chronic heavy-query bot tier driving deadlock-remediation restarts); pruned superseded `wdqs_2026_05_11_worobot` * 21:51 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] (duration: 11m 52s) * 21:44 sbassett@deploy1003: sbassett: Continuing with deployment * 21:43 sbassett@deploy1003: sbassett: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:39 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] * 20:38 dani@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] (duration: 32m 51s) * 20:38 ryankemper: [WDQS] Pruned obsolete requestctl action+pattern `wdqs_20260715_p2003_ring_ja3n` (actor rotated JA3Ns; rule inert) * 20:26 dani@deploy1003: dani, vadymts1: Continuing with deployment * 20:24 dani@deploy1003: dani, vadymts1: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:14 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2244: Testing * 20:06 dani@deploy1003: Started scap sync-world: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] * 19:56 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1013.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2016.codfw.wmnet with OS bookworm * 19:40 mutante: gerrit - one more service restart is needed - restarting * 19:29 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2244: Testing * 19:27 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2244: Testing * 19:27 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2244: Testing * 19:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2016.codfw.wmnet with reason: host reimage * 19:15 dancy@deploy1003: Finished deploy [zuul/deploy@d92e238]: Freshening Zuul installation (duration: 00m 15s) * 19:14 dancy@deploy1003: Started deploy [zuul/deploy@d92e238]: Freshening Zuul installation * 19:11 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2016.codfw.wmnet with reason: host reimage * 18:54 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1013.eqiad.wmnet, repooling source-only afterwards * 18:52 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2016 * 18:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2016 * 18:51 dancy@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 18:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2016.codfw.wmnet with OS bookworm * 18:39 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 11s) * 18:39 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 18:36 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 18:30 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417] (thin): Regular analytics weekly train THIN [analytics/refinery@2a25417d] (duration: 02m 09s) * 18:28 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417] (thin): Regular analytics weekly train THIN [analytics/refinery@2a25417d] * 18:28 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417]: Regular analytics weekly train [analytics/refinery@2a25417d] (duration: 04m 31s) * 18:27 dduvall: deploying https://gerrit.wikimedia.org/r/c/integration/config/+/1314025 (4 jobs updated) * 18:23 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417]: Regular analytics weekly train [analytics/refinery@2a25417d] * 18:22 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@2a25417d] (duration: 01m 59s) * 18:20 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@2a25417d] * 17:56 Raine: deployment server switchover => deploy1003 is primary now * 17:55 kamila@deploy1003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 28m 24s) * 17:54 mutante: restarting gerrit for maintenance * 17:29 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1013.eqiad.wmnet with OS bookworm * 17:27 kamila@deploy1003: Started scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] * 17:20 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] (duration: 22m 50s) * 17:12 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1023.eqiad.wmnet -> wdqs1024.eqiad.wmnet, repooling source-only afterwards * 17:04 Raine: point deployment.eqiad.wmnet to deploy1003 * 17:04 kamila@dns7001: END - running authdns-update * 17:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1013.eqiad.wmnet with reason: host reimage * 17:02 kamila@dns7001: START - running authdns-update * 17:01 kamila@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on releases2003.codfw.wmnet,releases1003.eqiad.wmnet with reason: Deployment server switchover * 17:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1013.eqiad.wmnet with reason: host reimage * 16:58 kamila@deploy2003: Locking from deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] * 16:57 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] (duration: 02m 33s) * 16:55 kamila@deploy2003: Locking from deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] * 16:55 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2003 - [[phab:T240266|T240266]] (duration: 00m 11s) * 16:54 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2003 - [[phab:T240266|T240266]] * 16:40 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1013 * 16:40 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1013 * 16:39 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1013 * 16:39 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1013.eqiad.wmnet 105.32.64.10.in-addr.arpa 5.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:39 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1013.eqiad.wmnet 105.32.64.10.in-addr.arpa 5.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:39 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:39 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1013 - bking@cumin2003" * 16:39 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1013 - bking@cumin2003" * 16:34 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:34 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1013 * 16:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1013.eqiad.wmnet with OS bookworm * 16:28 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1023.eqiad.wmnet -> wdqs1024.eqiad.wmnet, repooling source-only afterwards * 16:27 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-scholarly,name=eqiad * 16:27 eevans@deploy2003: helmfile [eqiad] DONE helmfile.d/services/linked-artifacts: apply * 16:26 eevans@deploy2003: helmfile [eqiad] START helmfile.d/services/linked-artifacts: apply * 16:26 eevans@deploy2003: helmfile [codfw] DONE helmfile.d/services/linked-artifacts: apply * 16:26 eevans@deploy2003: helmfile [codfw] START helmfile.d/services/linked-artifacts: apply * 16:25 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 29s) * 16:25 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 16:24 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 16:21 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 16:18 eevans@deploy2003: helmfile [codfw] DONE helmfile.d/services/linked-artifacts: apply * 16:18 eevans@deploy2003: helmfile [codfw] START helmfile.d/services/linked-artifacts: apply * 16:08 eevans@deploy2003: helmfile [staging] DONE helmfile.d/services/linked-artifacts: apply * 16:07 eevans@deploy2003: helmfile [staging] START helmfile.d/services/linked-artifacts: apply * 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 16:01 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 15:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1024.eqiad.wmnet with OS bookworm * 15:49 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] (duration: 00m 10s) * 15:49 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] * 15:48 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] (duration: 00m 15s) * 15:48 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] * 15:47 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] (duration: 00m 10s) * 15:47 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] * 15:46 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:42 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:42 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:40 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:37 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 15:37 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:36 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:36 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:36 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 15:33 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 15:33 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:31 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:28 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 15:27 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:27 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1024.eqiad.wmnet with reason: host reimage * 15:23 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:23 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:23 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:20 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1068.eqiad.wmnet * 15:20 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1068.eqiad.wmnet * 15:20 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1068.eqiad.wmnet * 15:20 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 15:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1024.eqiad.wmnet with reason: host reimage * 15:11 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:55 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 14:52 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wdqs1024.eqiad.wmnet with OS bookworm * 14:50 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] (duration: 00m 09s) * 14:50 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] * 14:49 jiji@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 14:49 jiji@deploy2003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 14:49 jiji@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 14:48 jiji@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 14:45 ecarg@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:45 ecarg@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:44 ecarg@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:44 ecarg@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:43 ecarg@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:43 ecarg@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:41 sukhe: ipvsadm --delete-service --tcp-service 10.2.1.55:8087: lvs2014 and lvs2013 * 14:39 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:39 sukhe: ipvsadm --delete-service --tcp-service 10.2.2.55:8087: [[phab:T432445|T432445]] * 14:38 ecarg@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:38 ecarg@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:37 ecarg@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:37 ecarg@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:36 ecarg@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:34 ecarg@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts datahubsearch1001.eqiad.wmnet * 14:32 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:32 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 14:31 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 14:31 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 14:28 sukhe: sudo cumin 'A:lvs-low-traffic-codfw' 'systemctl restart pybal': lvs2013 * 14:26 sukhe: sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal': lvs2014 * 14:26 sukhe: sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal' * 14:24 sukhe: restart pybal on lvs1019 * 14:24 sukhe: restart pybal on lvs1020 * 14:19 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for 1036 hosts * 14:17 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:04 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] (duration: 09m 28s) * 13:59 kharlan@deploy2003: dreamyjazz, kharlan: Continuing with deployment * 13:58 bking@cumin2003: START - Cookbook sre.hosts.decommission for hosts datahubsearch1001.eqiad.wmnet * 13:57 kharlan@deploy2003: dreamyjazz, kharlan: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:55 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] * 13:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts datahubsearch[1002-1003].eqiad.wmnet * 13:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:53 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch[1002-1003].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 13:52 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch[1002-1003].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 13:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 13:42 stran@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] (duration: 07m 30s) * 13:42 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:40 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: Test * 13:38 stran@deploy2003: dragoniez, stran: Continuing with deployment * 13:37 stran@deploy2003: dragoniez, stran: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:35 bking@cumin2003: START - Cookbook sre.hosts.decommission for hosts datahubsearch[1002-1003].eqiad.wmnet * 13:35 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024'] * 13:35 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 13:35 stran@deploy2003: Started scap sync-world: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] * 13:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 13:28 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024'] * 13:26 sukhe@dns1004: END - running authdns-update * 13:25 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 13:25 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 13:24 sukhe@dns1004: START - running authdns-update * 13:22 stran@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] (duration: 08m 20s) * 13:21 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 13:20 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 13:19 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 13:19 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 13:19 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 13:18 stran@deploy2003: stran: Continuing with deployment * 13:16 stran@deploy2003: stran: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:14 stran@deploy2003: Started scap sync-world: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] * 13:13 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 13:13 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 13:11 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 13:11 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 13:08 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 12:55 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool es2051: Test * 12:55 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: Test * 12:54 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool es2051: Test * 12:43 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1036 hosts * 12:41 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 12:40 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1048.eqiad.wmnet * 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 12:39 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 12:38 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 12:37 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 12:37 brouberol@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 12:36 brouberol@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 12:36 brouberol@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 12:36 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 12:36 elukey@cumin1003: DONE (PASS) - Cookbook sre.puppet.renew-cert (exit_code=0) for crm2001.codfw.wmnet: Renew puppet certificate - elukey@cumin1003 * 12:35 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:35 brouberol@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 12:34 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 12:31 brouberol@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 12:30 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1048.eqiad.wmnet * 12:30 brouberol@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 12:28 brouberol@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 12:27 brouberol@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 12:20 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1068.eqiad.wmnet with OS trixie * 12:01 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] (duration: 13m 19s) * 11:58 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1068.eqiad.wmnet with reason: host reimage * 11:52 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1068.eqiad.wmnet with reason: host reimage * 11:51 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 11:49 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:47 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] * 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2252: Security updates * 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:43 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 11:42 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2252: Security updates * 11:42 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply * 11:40 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply * 11:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 11:37 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2252.codfw.wmnet with OS trixie * 11:34 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1068 * 11:34 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1068 * 11:26 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1068 * 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1068.eqiad.wmnet 46.48.64.10.in-addr.arpa 6.4.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:26 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1068.eqiad.wmnet 46.48.64.10.in-addr.arpa 6.4.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1068 - jiji@cumin1003" * 11:26 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1068 - jiji@cumin1003" * 11:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2252.codfw.wmnet with reason: host reimage * 11:17 jiji@cumin1003: START - Cookbook sre.dns.netbox * 11:17 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1068 * 11:17 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1068.eqiad.wmnet with OS trixie * 11:17 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2252.codfw.wmnet with reason: host reimage * 11:15 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1068.eqiad.wmnet * 11:15 mvolz@deploy2003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:15 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1068.eqiad.wmnet * 11:15 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1068.eqiad.wmnet * 11:14 mvolz@deploy2003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:13 mvolz@deploy2003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:13 mvolz@deploy2003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:12 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] (duration: 11m 05s) * 11:11 mvolz@deploy2003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:10 mvolz@deploy2003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:07 dreamyjazz@deploy2003: dreamyjazz, kharlan: Continuing with deployment * 11:03 dreamyjazz@deploy2003: dreamyjazz, kharlan: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:03 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2252.codfw.wmnet with OS trixie * 11:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2252: Upgrading db2252.codfw.wmnet * 11:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:02 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 11:02 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2252: Upgrading db2252.codfw.wmnet * 11:02 cwilliams@cumin1003: dbmaint on ms3@codfw [[phab:T432321|T432321]] * 11:01 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 11:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db1153.eqiad.wmnet with reason: Security updates * 11:01 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] * 11:00 fnegri@deploy2003: helmfile [eqiad] DONE helmfile.d/services/toolhub: apply * 10:58 fnegri@deploy2003: helmfile [eqiad] START helmfile.d/services/toolhub: apply * 10:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1151: Security updates * 10:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:57 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1151: Security updates * 10:55 fnegri@deploy2003: helmfile [codfw] DONE helmfile.d/services/toolhub: apply * 10:54 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] (duration: 08m 38s) * 10:53 fnegri@deploy2003: helmfile [codfw] START helmfile.d/services/toolhub: apply * 10:53 fnegri@deploy2003: helmfile [staging] DONE helmfile.d/services/toolhub: apply * 10:52 fnegri@deploy2003: helmfile [staging] START helmfile.d/services/toolhub: apply * 10:50 zabe@deploy2003: zabe: Continuing with deployment * 10:47 zabe@deploy2003: zabe: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:45 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] * 10:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1151: Security updates * 10:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:42 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:42 root@cumin1003: START - Cookbook sre.mysql.depool depool db1151: Security updates * 10:38 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] (duration: 12m 47s) * 10:34 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 10:34 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 10:33 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2253.codfw.wmnet with OS trixie * 10:28 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:26 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] * 10:18 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2253.codfw.wmnet with reason: host reimage * 10:13 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2253.codfw.wmnet with reason: host reimage * 10:00 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2253.codfw.wmnet with OS trixie * 09:58 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db1151.eqiad.wmnet with reason: Security updates * 09:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2253: Upgrading db2253.codfw.wmnet * 09:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:57 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 09:56 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2253: Upgrading db2253.codfw.wmnet * 09:56 cwilliams@cumin1003: dbmaint on ms2@codfw [[phab:T432321|T432321]] * 09:56 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 09:36 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: UI improvement; support url shortener - oblivian@cumin1003" * 09:36 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: UI improvement; support url shortener - oblivian@cumin1003 * 09:35 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: UI improvement; support url shortener - oblivian@cumin1003 * 09:35 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: UI improvement; support url shortener - oblivian@cumin1003" * 09:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1152: Security updates * 09:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:26 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db1152: Security updates * 09:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: Security updates * 09:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:11 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:11 root@cumin1003: START - Cookbook sre.mysql.depool depool db1152: Security updates * 09:10 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1018.eqiad.wmnet with reason: Cloning * 09:09 Dreamy_Jazz: Deployed patch for [[phab:T432453|T432453]] and [[phab:T432454|T432454]] * 09:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 09:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2251.codfw.wmnet with OS trixie * 08:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2251.codfw.wmnet with reason: host reimage * 08:45 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2251.codfw.wmnet with reason: host reimage * 08:40 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] (duration: 12m 26s) * 08:38 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1030.eqiad.wmnet,service=s1 * 08:36 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 08:31 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2251.codfw.wmnet with OS trixie * 08:30 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:28 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] * 08:25 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] (duration: 07m 59s) * 08:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2251: Upgrading db2251.codfw.wmnet * 08:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:22 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 08:22 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2251: Upgrading db2251.codfw.wmnet * 08:20 urbanecm@deploy2003: urbanecm: Continuing with deployment * 08:20 cwilliams@cumin1003: dbmaint on ms1@codfw [[phab:T432321|T432321]] * 08:20 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 08:19 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:17 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] * 08:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade * 08:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade * 08:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db2251.codfw.wmnet,db1152.eqiad.wmnet with reason: OS upgrade * 08:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade * 08:13 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade * 08:11 Dreamy_Jazz: Created cusi_signal, cusi_case, and cusi_user on ukwiki and enwikivoyage in extension1 * 08:11 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade * 08:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade * 08:04 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1030.eqiad.wmnet,service=s1 * 08:04 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1030.eqiad.wmnet,service=s1 * 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply * 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply * 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply * 07:51 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply * 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 07:47 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 07:47 phuedx: End of UTC morning backport window * 07:43 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 07:43 phuedx@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] (duration: 13m 44s) * 07:43 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 07:39 phuedx@deploy2003: phuedx: Continuing with deployment * 07:31 phuedx@deploy2003: phuedx: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:29 phuedx@deploy2003: Started scap sync-world: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] * 07:24 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Turnilo import support - oblivian@cumin1003" * 07:24 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import support - oblivian@cumin1003 * 07:23 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import support - oblivian@cumin1003 * 07:23 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Turnilo import support - oblivian@cumin1003" * 06:42 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs2020 after successful Bookworm reimage, data transfer, and postflight validation * 06:42 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2020.codfw.wmnet * 05:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Managing sanitization for wikis bolwiki in section s5 * 05:25 marostegui@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis bolwiki in section s5 * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 41s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-21 == * 22:50 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2019.codfw.wmnet -> wdqs2020.codfw.wmnet, repooling source-only afterwards * 22:47 cwhite: force reboot arclamp2001 - appears to have run out of memory and gone unresponsive * 22:24 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 01m 26s) * 22:24 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 22:23 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 22:22 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024'] * 22:11 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 21:54 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs1024'] * 21:54 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 21:53 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs1024'] * 21:53 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 21:49 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2019.codfw.wmnet -> wdqs2020.codfw.wmnet, repooling source-only afterwards * 20:57 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] (duration: 09m 10s) * 20:55 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1024.eqiad.wmnet with OS bookworm * 20:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2020.codfw.wmnet with OS bookworm * 20:52 krinkle@deploy2003: krinkle: Continuing with deployment * 20:49 krinkle@deploy2003: krinkle: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:47 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] * 20:45 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] (duration: 05m 42s) * 20:44 krinkle@deploy2003: krinkle: Rolling back deployment * 20:41 krinkle@deploy2003: krinkle: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:39 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] * 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2003.codfw.wmnet * 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1003.eqiad.wmnet * 20:33 dani@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] (duration: 11m 15s) * 20:33 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2003.codfw.wmnet * 20:33 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1003.eqiad.wmnet * 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2020.codfw.wmnet with reason: host reimage * 20:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1002.eqiad.wmnet * 20:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2002.codfw.wmnet * 20:30 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 20:29 dani@deploy2003: dani: Continuing with deployment * 20:29 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2020.codfw.wmnet with reason: host reimage * 20:26 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1002.eqiad.wmnet * 20:26 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2002.codfw.wmnet * 20:24 dani@deploy2003: dani: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2001.codfw.wmnet * 20:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1001.eqiad.wmnet * 20:22 dani@deploy2003: Started scap sync-world: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] * 20:22 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 20:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2001.codfw.wmnet * 20:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1001.eqiad.wmnet * 20:14 sbisson@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] (duration: 09m 01s) * 20:11 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2020.codfw.wmnet with OS bookworm * 20:10 sbisson@deploy2003: sbisson: Continuing with deployment * 20:07 sbisson@deploy2003: sbisson: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:05 sbisson@deploy2003: Started scap sync-world: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] * 20:03 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] (duration: 07m 04s) * 20:01 mutante: Gerrit - tomorrow a new SSH host key will appear - it will be {{Gerrit|ed25519}} and has already been added to wmf-laptop. you can verify it here: https://wikitech.wikimedia.org/wiki/Help:SSH_Fingerprints/gerrit.wikimedia.org:29418 ([[phab:T240266|T240266]]) * 19:59 zabe@deploy2003: zabe: Continuing with deployment * 19:58 zabe@deploy2003: zabe: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:56 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] * 19:52 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] (duration: 07m 25s) * 19:48 zabe@deploy2003: zabe: Continuing with deployment * 19:47 zabe@deploy2003: zabe: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:45 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] * 19:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 19:32 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024'] * 19:27 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 19:26 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024'] * 19:26 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 19:24 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024'] * 19:12 ryankemper: [wdqs] [[phab:T430880|T430880]] Repooled `wdqs-scholarly` discovery in `eqiad` after validating `wdqs1023` end-to-end; `wdqs1024` remains disabled pending reimage recovery * 19:11 ryankemper: [wdqs] [[phab:T430880|T430880]] Repooled wdqs1012.eqiad.wmnet after successful Bookworm reimage, data transfer, service checks, readiness probe, and cross-graph federation query validation * 19:10 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 19:10 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1012.eqiad.wmnet * 19:08 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 18:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for deploy1003.eqiad.wmnet * 18:57 kamila@cumin1003: START - Cookbook sre.hosts.remove-downtime for deploy1003.eqiad.wmnet * 18:37 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:37 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding urldownloader service IPs - sukhe@cumin1003" * 18:37 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding urldownloader service IPs - sukhe@cumin1003" * 18:32 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 18:32 dancy@deploy2003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 18:30 sukhe@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 18:27 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 18:24 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1024.eqiad.wmnet with OS bookworm * 18:20 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host deploy1003.eqiad.wmnet with OS bookworm * 18:09 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deploy1003 reimage (duration: 121m 16s) * 18:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 18:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1180: Security updates * 17:55 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wcqs2003.codfw.wmnet * 17:48 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wcqs2003.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1155.eqiad.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1155.eqiad.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2224.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2224.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2217.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2217.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2193.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2193.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2180.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2180.codfw.wmnet * 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1168.eqiad.wmnet * 17:36 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1168.eqiad.wmnet * 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2169.codfw.wmnet * 17:36 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2169.codfw.wmnet * 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1165.eqiad.wmnet * 17:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1165.eqiad.wmnet * 17:35 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2158.codfw.wmnet * 17:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2158.codfw.wmnet * 17:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wcqs1003.eqiad.wmnet * 17:17 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1180: Security updates * 17:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1180.eqiad.wmnet * 17:16 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1180.eqiad.wmnet * 17:15 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp3073.* * 17:13 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wcqs1003.eqiad.wmnet * 17:11 brett@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp3073.esams.wmnet with OS trixie * 17:11 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 17:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1024 * 17:04 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1024 * 17:03 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 17:00 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 16:59 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2242: codfw rack B7 depool for maintenance * 16:59 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 16:43 brett@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp3073.esams.wmnet with reason: host reimage * 16:42 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 16:39 brett@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cp3073.esams.wmnet with reason: host reimage * 16:32 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on deploy1003.eqiad.wmnet with reason: host reimage * 16:27 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on deploy1003.eqiad.wmnet with reason: host reimage * 16:14 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2242: codfw rack B7 depool for maintenance * 16:14 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: codfw rack B7 depool for maintenance * 16:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1023.eqiad.wmnet with OS bookworm * 16:13 brett@cumin2002: START - Cookbook sre.hosts.reimage for host cp3073.esams.wmnet with OS trixie * 16:08 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host deploy1003.eqiad.wmnet with OS bookworm * 16:08 kamila@deploy2003: Locking from deployment [MediaWiki]: deploy1003 reimage * 16:03 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1012.eqiad.wmnet with OS bookworm * 15:48 inflatador: bking@apt1002 `sudo reprepro copy bookworm-wikimedia bullseye-wikimedia jvmquake` [[phab:T430880|T430880]] * 15:39 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp3073.* * 15:39 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 15:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:34 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1311536{{!}}Set $wgMathInternalRestbaseURL explicitly (take 2) (T349582)]] * 15:29 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:29 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2228: codfw rack B7 depool for maintenance * 15:29 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2229: codfw rack B7 depool for maintenance * 15:27 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1180: Security update * 15:25 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Security update * 15:21 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 15:21 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 15:19 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 24s) * 15:19 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:14 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 15:14 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db1180: Security update * 15:13 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-eqiad * 14:48 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-eqiad * 14:44 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2229: codfw rack B7 depool for maintenance * 14:44 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc2017: codfw rack B7 depool for maintenance * 14:44 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:43 cmooney@cumin2003: START - Cookbook sre.mysql.parsercache * 14:43 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool pc2017: codfw rack B7 depool for maintenance * 14:43 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2003.codfw.wmnet * 14:43 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2003.codfw.wmnet * 14:42 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2009.codfw.wmnet * 14:42 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2009.codfw.wmnet * 14:41 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:41 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:40 cmooney@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 29 hosts * 14:40 cmooney@cumin1003: START - Cookbook sre.hosts.remove-downtime for 29 hosts * 14:35 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 14:34 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 14:32 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] (duration: 07m 56s) * 14:29 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ssw1-a[1,8]-codfw with reason: lsw1-b7-codfw JunOS upgrade * 14:28 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 14:28 elukey@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 14:26 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:24 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] * 14:23 topranks: reboot lsw1-b7-codfw to upgrade JunOS (affects all hosts in rack) [[phab:T430928|T430928]] * 14:18 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2003.codfw.wmnet * 14:14 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2009.codfw.wmnet * 14:14 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Security update * 14:13 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2242: codfw rack B7 depool for maintenance * 14:13 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2242: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2228: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2228: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2229: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2229: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc2017: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.parsercache * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool pc2017: codfw rack B7 depool for maintenance * 14:08 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2003.codfw.wmnet * 14:07 cmooney@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on aux-k8s-etcd2004.codfw.wmnet,ml-etcd2001.codfw.wmnet with reason: lsw1-b7-codfw JunOS upgrade * 14:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95005 and previous config saved to /var/cache/conftool/dbconfig/20260721-140620-cwilliams.json * 14:05 cmooney@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2049.codfw.wmnet * 14:05 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-scholarly,name=eqiad * 14:04 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2009.codfw.wmnet * 14:04 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply * 14:04 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply * 14:03 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:03 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2001.codfw.wmnet * 14:03 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2001.codfw.wmnet * 14:02 cmooney@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2049.codfw.wmnet * 14:00 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:00 Dreamy_Jazz: Created cusi_case, cusi_signal, and cusi_user on svwiki, dewiki, jawiki, eswiki * 13:59 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b7-codfw,lsw1-b7-codfw IPv6,lsw1-b7-codfw.mgmt,ssw1-a[1,8]-codfw.mgmt with reason: lsw1-b7-codfw JunOS upgrade * 13:57 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs1023.eqiad.wmnet, repooling source-only afterwards * 13:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 29 hosts with reason: lsw1-b7-codfw JunOS upgrade * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224', diff saved to https://phabricator.wikimedia.org/P95003 and previous config saved to /var/cache/conftool/dbconfig/20260721-135613-cwilliams.json * 13:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1012.eqiad.wmnet with reason: host reimage * 13:53 cmooney@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 1:00:00 on 30 hosts with reason: lsw1-b7-codfw JunOS upgrade * 13:51 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1012.eqiad.wmnet with reason: host reimage * 13:48 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 13:48 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 13:46 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224', diff saved to https://phabricator.wikimedia.org/P95001 and previous config saved to /var/cache/conftool/dbconfig/20260721-134605-cwilliams.json * 13:46 cmooney@cumin1003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti2033.codfw.wmnet * 13:46 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 13:45 elukey: move the Docker Registry's /v2/wikimedia/machinelearning.* prefix to the ml S3 backend - [[phab:T428022|T428022]] * 13:45 cmooney@cumin1003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti2033.codfw.wmnet * 13:45 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 13:43 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:40 jiji@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 13:40 jiji@deploy2003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 13:39 jiji@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 13:39 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 13:38 cmooney@cumin1003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2032.codfw.wmnet * 13:38 jiji@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 13:38 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 13:37 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2032.codfw.wmnet * 13:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95000 and previous config saved to /var/cache/conftool/dbconfig/20260721-133557-cwilliams.json * 13:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1012 * 13:33 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1012 * 13:33 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1012.eqiad.wmnet with OS bookworm * 13:30 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:30 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:28 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94999 and previous config saved to /var/cache/conftool/dbconfig/20260721-132855-cwilliams.json * 13:28 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2224.codfw.wmnet with reason: Maintenance * 13:28 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94998 and previous config saved to /var/cache/conftool/dbconfig/20260721-132826-cwilliams.json * 13:28 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 13:23 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] (duration: 07m 50s) * 13:20 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 13:18 kharlan@deploy2003: kharlan: Continuing with deployment * 13:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217', diff saved to https://phabricator.wikimedia.org/P94996 and previous config saved to /var/cache/conftool/dbconfig/20260721-131817-cwilliams.json * 13:17 brouberol@dns1004: END - running authdns-update * 13:17 kharlan@deploy2003: kharlan: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:15 brouberol@dns1004: START - running authdns-update * 13:15 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] * 13:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94995 and previous config saved to /var/cache/conftool/dbconfig/20260721-131411-cwilliams.json * 13:13 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs1023.eqiad.wmnet, repooling source-only afterwards * 13:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217', diff saved to https://phabricator.wikimedia.org/P94994 and previous config saved to /var/cache/conftool/dbconfig/20260721-130809-cwilliams.json * 13:07 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 13:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180', diff saved to https://phabricator.wikimedia.org/P94993 and previous config saved to /var/cache/conftool/dbconfig/20260721-130404-cwilliams.json * 13:03 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:03 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:02 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 13:02 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 12:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94992 and previous config saved to /var/cache/conftool/dbconfig/20260721-125801-cwilliams.json * 12:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180', diff saved to https://phabricator.wikimedia.org/P94991 and previous config saved to /var/cache/conftool/dbconfig/20260721-125356-cwilliams.json * 12:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94990 and previous config saved to /var/cache/conftool/dbconfig/20260721-125049-cwilliams.json * 12:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2217.codfw.wmnet with reason: Maintenance * 12:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94989 and previous config saved to /var/cache/conftool/dbconfig/20260721-125017-cwilliams.json * 12:48 elukey: bmc cold reboot for lvs1013 and lvs1015 - [[phab:T426180|T426180]] * 12:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94988 and previous config saved to /var/cache/conftool/dbconfig/20260721-124348-cwilliams.json * 12:40 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193', diff saved to https://phabricator.wikimedia.org/P94987 and previous config saved to /var/cache/conftool/dbconfig/20260721-124009-cwilliams.json * 12:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts * 12:33 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts * 12:33 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts * 12:32 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts * 12:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:30 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193', diff saved to https://phabricator.wikimedia.org/P94986 and previous config saved to /var/cache/conftool/dbconfig/20260721-123001-cwilliams.json * 12:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94985 and previous config saved to /var/cache/conftool/dbconfig/20260721-121953-cwilliams.json * 12:17 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs2007.codfw.wmnet with OS bookworm * 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94983 and previous config saved to /var/cache/conftool/dbconfig/20260721-121257-cwilliams.json * 12:12 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2193.codfw.wmnet with reason: Maintenance * 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94982 and previous config saved to /var/cache/conftool/dbconfig/20260721-121239-cwilliams.json * 12:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180', diff saved to https://phabricator.wikimedia.org/P94980 and previous config saved to /var/cache/conftool/dbconfig/20260721-120231-cwilliams.json * 11:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180', diff saved to https://phabricator.wikimedia.org/P94979 and previous config saved to /var/cache/conftool/dbconfig/20260721-115223-cwilliams.json * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94978 and previous config saved to /var/cache/conftool/dbconfig/20260721-114333-cwilliams.json * 11:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1180.eqiad.wmnet with reason: Maintenance * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94977 and previous config saved to /var/cache/conftool/dbconfig/20260721-114305-cwilliams.json * 11:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94976 and previous config saved to /var/cache/conftool/dbconfig/20260721-114215-cwilliams.json * 11:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94975 and previous config saved to /var/cache/conftool/dbconfig/20260721-113530-cwilliams.json * 11:35 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2180.codfw.wmnet with reason: Maintenance * 11:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94974 and previous config saved to /var/cache/conftool/dbconfig/20260721-113501-cwilliams.json * 11:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168', diff saved to https://phabricator.wikimedia.org/P94973 and previous config saved to /var/cache/conftool/dbconfig/20260721-113258-cwilliams.json * 11:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169', diff saved to https://phabricator.wikimedia.org/P94972 and previous config saved to /var/cache/conftool/dbconfig/20260721-112453-cwilliams.json * 11:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168', diff saved to https://phabricator.wikimedia.org/P94971 and previous config saved to /var/cache/conftool/dbconfig/20260721-112250-cwilliams.json * 11:21 XioNoX: put eqiad-drmrs Arelion link in service * 11:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169', diff saved to https://phabricator.wikimedia.org/P94970 and previous config saved to /var/cache/conftool/dbconfig/20260721-111446-cwilliams.json * 11:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94969 and previous config saved to /var/cache/conftool/dbconfig/20260721-111242-cwilliams.json * 11:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 11:10 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 11:07 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1093 hosts * 11:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94968 and previous config saved to /var/cache/conftool/dbconfig/20260721-110548-cwilliams.json * 11:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1168.eqiad.wmnet with reason: Maintenance * 11:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94967 and previous config saved to /var/cache/conftool/dbconfig/20260721-110520-cwilliams.json * 11:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94966 and previous config saved to /var/cache/conftool/dbconfig/20260721-110439-cwilliams.json * 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94964 and previous config saved to /var/cache/conftool/dbconfig/20260721-105632-cwilliams.json * 10:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2169.codfw.wmnet with reason: Maintenance * 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94963 and previous config saved to /var/cache/conftool/dbconfig/20260721-105603-cwilliams.json * 10:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165', diff saved to https://phabricator.wikimedia.org/P94962 and previous config saved to /var/cache/conftool/dbconfig/20260721-105512-cwilliams.json * 10:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158', diff saved to https://phabricator.wikimedia.org/P94961 and previous config saved to /var/cache/conftool/dbconfig/20260721-104555-cwilliams.json * 10:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165', diff saved to https://phabricator.wikimedia.org/P94960 and previous config saved to /var/cache/conftool/dbconfig/20260721-104504-cwilliams.json * 10:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158', diff saved to https://phabricator.wikimedia.org/P94959 and previous config saved to /var/cache/conftool/dbconfig/20260721-103547-cwilliams.json * 10:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94958 and previous config saved to /var/cache/conftool/dbconfig/20260721-103456-cwilliams.json * 10:29 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2229: Upgraded kernel * 10:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94956 and previous config saved to /var/cache/conftool/dbconfig/20260721-102757-cwilliams.json * 10:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on an-redacteddb1001.eqiad.wmnet,clouddb[1015,1025,1028].eqiad.wmnet,db1155.eqiad.wmnet with reason: Maintenance * 10:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1165.eqiad.wmnet with reason: Maintenance * 10:25 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94955 and previous config saved to /var/cache/conftool/dbconfig/20260721-102539-cwilliams.json * 10:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94954 and previous config saved to /var/cache/conftool/dbconfig/20260721-101848-cwilliams.json * 10:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2158.codfw.wmnet with reason: Maintenance * 09:43 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2229: Upgraded kernel * 09:42 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2229.codfw.wmnet * 09:42 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2229.codfw.wmnet * 09:23 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db2229.codfw.wmnet * 09:23 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2229.codfw.wmnet * 08:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2229 [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94948 and previous config saved to /var/cache/conftool/dbconfig/20260721-085724-cwilliams.json * 08:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2214 to s6 primary [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94947 and previous config saved to /var/cache/conftool/dbconfig/20260721-085442-cwilliams.json * 08:53 cezmunsta: Starting s6 codfw failover from db2229 to db2214 - [[phab:T430964|T430964]] * 08:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2214 with weight 0 [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94946 and previous config saved to /var/cache/conftool/dbconfig/20260721-084613-cwilliams.json * 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 22 hosts with reason: Primary switchover s6 [[phab:T430964|T430964]] * 08:32 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1017.eqiad.wmnet,service=s1 * 08:08 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Add subrated circuit rate to interface descriptions - CR1312476 - ayounsi@cumin1003 * 08:06 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Add subrated circuit rate to interface descriptions - CR1312476 - ayounsi@cumin1003 * 07:58 reedy@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] (duration: 12m 55s) * 07:51 reedy@deploy2003: reedy, neriah: Continuing with deployment * 07:51 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 07:51 reedy@deploy2003: reedy, neriah: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:48 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1093 hosts * 07:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm2001.wikimedia.org * 07:45 reedy@deploy2003: Started scap sync-world: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] * 07:43 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 07:42 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm2001.wikimedia.org * 07:23 elukey: upgrade libtiff6 packages on zuul* trixie hosts for security upgrades * 07:22 elukey: upgrade libtiff6 packages on Wikikube trixie workers for security upgrades * 07:14 elukey@deploy2003: helmfile [codfw] DONE helmfile.d/services/proton: sync * 07:13 elukey@deploy2003: helmfile [codfw] START helmfile.d/services/proton: sync * 07:11 elukey@deploy2003: helmfile [eqiad] DONE helmfile.d/services/proton: sync * 07:10 elukey@deploy2003: helmfile [eqiad] START helmfile.d/services/proton: sync * 07:09 elukey@deploy2003: helmfile [staging] DONE helmfile.d/services/proton: sync * 07:08 elukey@deploy2003: helmfile [staging] START helmfile.d/services/proton: sync * 06:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1023.eqiad.wmnet with reason: host reimage * 06:46 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1023.eqiad.wmnet with reason: host reimage * 06:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 05:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Haproxy-only mode support - oblivian@cumin1003" * 05:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Haproxy-only mode support - oblivian@cumin1003 * 05:42 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Haproxy-only mode support - oblivian@cumin1003 * 05:42 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Haproxy-only mode support - oblivian@cumin1003" * 05:38 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1017.eqiad.wmnet with reason: Cloning * 05:37 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1017.eqiad.wmnet,service=s1 * 05:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1029.eqiad.wmnet,service=s8 * 05:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1029.eqiad.wmnet,service=s5 * 05:32 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:30 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 05:11 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet * 05:04 aokoth@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet * 05:00 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 04:56 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 04:01 mwpresync@deploy2003: Pruned MediaWiki: 1.47.0-wmf.9 (duration: 01m 08s) * 03:41 mwpresync@deploy2003: Finished scap sync-world: testwikis to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] (duration: 36m 30s) * 03:05 mwpresync@deploy2003: Started scap sync-world: testwikis to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 03:01 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:01 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:00 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:00 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:36 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:36 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:36 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:35 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:16 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 47s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 00:56 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm == 2026-07-20 == * 23:38 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 23:07 Amir1: deleting echo notifications from 2015 on group1 wikis * 23:07 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] (duration: 14m 16s) * 23:01 ladsgroup@deploy2003: ladsgroup: Continuing with deployment * 23:00 ladsgroup@deploy2003: ladsgroup: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:53 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] * 22:46 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2007.codfw.wmnet, repooling source-only afterwards * 22:39 maryum: Deployed security fixes for several security bugs * 21:42 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 21:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2007.codfw.wmnet, repooling source-only afterwards * 21:37 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 21:37 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 21:34 sbassett: Deployed security fix for [[phab:T432424|T432424]] * 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs2020.codfw.wmnet * 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1023.eqiad.wmnet * 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1011.eqiad.wmnet * 21:32 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 17s) * 21:32 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 21:27 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 21:13 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2007.codfw.wmnet with reason: host reimage * 21:08 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-internal-main,name=codfw * 21:06 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2007.codfw.wmnet with reason: host reimage * 20:59 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service * 20:58 sukhe: pybal restart for IP changes around wdqs-main hosts * 20:57 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 20:46 ryankemper@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-internal-main,name=codfw * 20:45 ebernhardson@deploy2003: Finished deploy [search/mjolnir/deploy@d4dc3b8]: Update for opensearch 2.x compat (duration: 00m 34s) * 20:44 ebernhardson@deploy2003: Started deploy [search/mjolnir/deploy@d4dc3b8]: Update for opensearch 2.x compat * 20:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2007 * 20:44 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2007 * 20:43 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2007 * 20:43 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2007.codfw.wmnet 156.16.192.10.in-addr.arpa 6.5.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:42 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2007.codfw.wmnet 156.16.192.10.in-addr.arpa 6.5.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:42 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2007 - bking@cumin2003" * 20:41 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2007 - bking@cumin2003" * 20:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94944 and previous config saved to /var/cache/conftool/dbconfig/20260720-203333-cwilliams.json * 20:32 arlolra@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] (duration: 15m 07s) * 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2020.codfw.wmnet * 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1023.eqiad.wmnet * 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1011.eqiad.wmnet * 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs2020.codfw.wmnet * 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1023.eqiad.wmnet * 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1011.eqiad.wmnet * 20:25 arlolra@deploy2003: arlolra, cscott: Continuing with deployment * 20:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257', diff saved to https://phabricator.wikimedia.org/P94943 and previous config saved to /var/cache/conftool/dbconfig/20260720-202325-cwilliams.json * 20:21 arlolra@deploy2003: arlolra, cscott: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:17 arlolra@deploy2003: Started scap sync-world: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] * 20:13 bking@cumin2003: START - Cookbook sre.dns.netbox * 20:13 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257', diff saved to https://phabricator.wikimedia.org/P94942 and previous config saved to /var/cache/conftool/dbconfig/20260720-201318-cwilliams.json * 20:13 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 20:10 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 20:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2007 * 20:04 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2007.codfw.wmnet with OS bookworm * 20:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94941 and previous config saved to /var/cache/conftool/dbconfig/20260720-200310-cwilliams.json * 19:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94940 and previous config saved to /var/cache/conftool/dbconfig/20260720-195633-cwilliams.json * 19:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1257.eqiad.wmnet with reason: Maintenance * 19:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94939 and previous config saved to /var/cache/conftool/dbconfig/20260720-195605-cwilliams.json * 19:51 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 19:50 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 19:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256', diff saved to https://phabricator.wikimedia.org/P94938 and previous config saved to /var/cache/conftool/dbconfig/20260720-194558-cwilliams.json * 19:44 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 19:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 19:41 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wdqs1011.eqiad.wmnet with OS bookworm * 19:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256', diff saved to https://phabricator.wikimedia.org/P94937 and previous config saved to /var/cache/conftool/dbconfig/20260720-193550-cwilliams.json * 19:25 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94936 and previous config saved to /var/cache/conftool/dbconfig/20260720-192542-cwilliams.json * 19:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94935 and previous config saved to /var/cache/conftool/dbconfig/20260720-191856-cwilliams.json * 19:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1256.eqiad.wmnet with reason: Maintenance * 19:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94934 and previous config saved to /var/cache/conftool/dbconfig/20260720-191839-cwilliams.json * 19:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255', diff saved to https://phabricator.wikimedia.org/P94933 and previous config saved to /var/cache/conftool/dbconfig/20260720-190831-cwilliams.json * 18:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255', diff saved to https://phabricator.wikimedia.org/P94932 and previous config saved to /var/cache/conftool/dbconfig/20260720-185824-cwilliams.json * 18:50 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 18:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94931 and previous config saved to /var/cache/conftool/dbconfig/20260720-184816-cwilliams.json * 18:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94930 and previous config saved to /var/cache/conftool/dbconfig/20260720-184224-cwilliams.json * 18:42 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1255.eqiad.wmnet with reason: Maintenance * 18:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94929 and previous config saved to /var/cache/conftool/dbconfig/20260720-184153-cwilliams.json * 18:39 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 18:39 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 16s) * 18:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 18:38 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 59m 26s) * 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 18:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211', diff saved to https://phabricator.wikimedia.org/P94928 and previous config saved to /var/cache/conftool/dbconfig/20260720-183145-cwilliams.json * 18:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211', diff saved to https://phabricator.wikimedia.org/P94927 and previous config saved to /var/cache/conftool/dbconfig/20260720-182137-cwilliams.json * 18:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94926 and previous config saved to /var/cache/conftool/dbconfig/20260720-181129-cwilliams.json * 18:09 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_codfw * 18:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2057.codfw.wmnet * 18:08 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_codfw * 18:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2058.codfw.wmnet * 18:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94925 and previous config saved to /var/cache/conftool/dbconfig/20260720-180452-cwilliams.json * 18:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on clouddb[1016,1020,1022-1023].eqiad.wmnet,db1154.eqiad.wmnet with reason: Maintenance * 18:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1211.eqiad.wmnet with reason: Maintenance * 18:02 sukhe: armed keyholder on acmechief1002.eqiad.wmnet and acmechief2002.codfw.wmnet (active host) * 18:01 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief2002.codfw.wmnet * 17:57 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief2002.codfw.wmnet * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs2020'] * 17:52 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief1002.eqiad.wmnet * 17:50 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 17:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1011.eqiad.wmnet with reason: host reimage * 17:48 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief1002.eqiad.wmnet * 17:47 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test2001.codfw.wmnet * 17:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94924 and previous config saved to /var/cache/conftool/dbconfig/20260720-174717-cwilliams.json * 17:46 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 17:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1011.eqiad.wmnet with reason: host reimage * 17:43 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs2020.codfw.wmnet with OS bookworm * 17:43 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test2001.codfw.wmnet * 17:43 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test1001.eqiad.wmnet * 17:39 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test1001.eqiad.wmnet * 17:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 17:38 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:38 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 17:37 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 17:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244', diff saved to https://phabricator.wikimedia.org/P94923 and previous config saved to /var/cache/conftool/dbconfig/20260720-173709-cwilliams.json * 17:35 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 17:31 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 20m 40s) * 17:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2055.codfw.wmnet * 17:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2056.codfw.wmnet * 17:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1011 * 17:27 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1011 * 17:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1011.eqiad.wmnet with OS bookworm * 17:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244', diff saved to https://phabricator.wikimedia.org/P94922 and previous config saved to /var/cache/conftool/dbconfig/20260720-172701-cwilliams.json * 17:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94921 and previous config saved to /var/cache/conftool/dbconfig/20260720-171653-cwilliams.json * 17:11 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 17:11 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 13m 03s) * 17:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94920 and previous config saved to /var/cache/conftool/dbconfig/20260720-171012-cwilliams.json * 17:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2244.codfw.wmnet with reason: Maintenance * 17:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94919 and previous config saved to /var/cache/conftool/dbconfig/20260720-170941-cwilliams.json * 16:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243', diff saved to https://phabricator.wikimedia.org/P94918 and previous config saved to /var/cache/conftool/dbconfig/20260720-165933-cwilliams.json * 16:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 16:58 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2053.codfw.wmnet * 16:51 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2054.codfw.wmnet * 16:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243', diff saved to https://phabricator.wikimedia.org/P94917 and previous config saved to /var/cache/conftool/dbconfig/20260720-164926-cwilliams.json * 16:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94916 and previous config saved to /var/cache/conftool/dbconfig/20260720-163918-cwilliams.json * 16:35 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 16:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94915 and previous config saved to /var/cache/conftool/dbconfig/20260720-163140-cwilliams.json * 16:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2243.codfw.wmnet with reason: Maintenance * 16:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94914 and previous config saved to /var/cache/conftool/dbconfig/20260720-163111-cwilliams.json * 16:27 btullis@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 16:27 btullis@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 16:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2020 * 16:23 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2020 * 16:21 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2020 * 16:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2020.codfw.wmnet 85.0.192.10.in-addr.arpa 5.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:21 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2020.codfw.wmnet 85.0.192.10.in-addr.arpa 5.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242', diff saved to https://phabricator.wikimedia.org/P94913 and previous config saved to /var/cache/conftool/dbconfig/20260720-162103-cwilliams.json * 16:19 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 16:18 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 16:18 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:18 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 16:17 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:17 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for netbox accounting errors - jhancock@cumin2002" * 16:17 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for netbox accounting errors - jhancock@cumin2002" * 16:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2051.codfw.wmnet * 16:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2052.codfw.wmnet * 16:11 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 16:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242', diff saved to https://phabricator.wikimedia.org/P94912 and previous config saved to /var/cache/conftool/dbconfig/20260720-161055-cwilliams.json * 16:09 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 16:08 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 16:06 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 16:06 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 16:06 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2020 * 16:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2020.codfw.wmnet with OS bookworm * 16:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94911 and previous config saved to /var/cache/conftool/dbconfig/20260720-160047-cwilliams.json * 15:58 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2019.codfw.wmnet, repooling source-only afterwards * 15:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94909 and previous config saved to /var/cache/conftool/dbconfig/20260720-155353-cwilliams.json * 15:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2242.codfw.wmnet with reason: Maintenance * 15:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94908 and previous config saved to /var/cache/conftool/dbconfig/20260720-154433-cwilliams.json * 15:35 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2049.codfw.wmnet * 15:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162', diff saved to https://phabricator.wikimedia.org/P94907 and previous config saved to /var/cache/conftool/dbconfig/20260720-153425-cwilliams.json * 15:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2050.codfw.wmnet * 15:28 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162', diff saved to https://phabricator.wikimedia.org/P94906 and previous config saved to /var/cache/conftool/dbconfig/20260720-152418-cwilliams.json * 15:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1023 * 15:14 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1023 * 15:14 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 15:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94905 and previous config saved to /var/cache/conftool/dbconfig/20260720-151407-cwilliams.json * 15:13 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] (duration: 41m 16s) * 15:08 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 15:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94902 and previous config saved to /var/cache/conftool/dbconfig/20260720-150729-cwilliams.json * 15:07 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2162.codfw.wmnet with reason: Maintenance * 15:05 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2027.codfw.wmnet, repooling source-only afterwards * 15:00 urbanecm@deploy2003: vadymts1, migr, urbanecm: Continuing with deployment * 14:59 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:58 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 07s) * 14:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:58 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 13s) * 14:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:57 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2019.codfw.wmnet, repooling source-only afterwards * 14:57 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2047.codfw.wmnet * 14:55 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2048.codfw.wmnet * 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2019.codfw.wmnet with OS bookworm * 14:49 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 14:47 urbanecm@deploy2003: vadymts1, migr, urbanecm: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:44 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool magru [reason: BGP issues in lvs7003 resolved after liberica restart, no task ID specified] * 14:44 sukhe@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool magru [reason: BGP issues in lvs7003 resolved after liberica restart, no task ID specified] * 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:41 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:39 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:39 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:33 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool magru [reason: no reason specified, no task ID specified] * 14:33 sukhe@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool magru [reason: no reason specified, no task ID specified] * 14:31 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] * 14:24 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:24 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:24 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:24 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2019.codfw.wmnet with reason: host reimage * 14:22 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2027.codfw.wmnet, repooling source-only afterwards * 14:19 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2019.codfw.wmnet with reason: host reimage * 14:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2027.codfw.wmnet with OS bookworm * 14:16 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2046.codfw.wmnet * 14:16 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2045.codfw.wmnet * 14:08 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:08 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:08 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:08 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:07 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:06 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:06 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:06 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:05 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2019 * 14:00 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2019 * 13:56 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1015.eqiad.wmnet * 13:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2027.codfw.wmnet with reason: host reimage * 13:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2071.codfw.wmnet with OS trixie * 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:51 sukhe@cumin1003: END (ERROR) - Cookbook sre.loadbalancer.admin (exit_code=97) rebooting A:liberica and P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica and P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:51 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1015.eqiad.wmnet * 13:50 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1014.eqiad.wmnet * 13:50 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1076.eqiad.wmnet with OS trixie * 13:50 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2027.codfw.wmnet with reason: host reimage * 13:45 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1014.eqiad.wmnet * 13:44 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1013.eqiad.wmnet * 13:39 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 13:38 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1013.eqiad.wmnet * 13:37 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2044.codfw.wmnet * 13:37 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2043.codfw.wmnet * 13:36 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2019 * 13:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2019.codfw.wmnet 156.32.192.10.in-addr.arpa 6.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:36 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2019.codfw.wmnet 156.32.192.10.in-addr.arpa 6.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:36 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2019 - bking@cumin2003" * 13:36 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2019 - bking@cumin2003" * 13:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry2005.codfw.wmnet * 13:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2071.codfw.wmnet with reason: host reimage * 13:31 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:31 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2019 * 13:31 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry2005.codfw.wmnet * 13:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry2004.codfw.wmnet * 13:30 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2019.codfw.wmnet with OS bookworm * 13:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2027 * 13:30 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2027 * 13:30 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2027.codfw.wmnet with OS bookworm * 13:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1076.eqiad.wmnet with reason: host reimage * 13:29 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_codfw * 13:28 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_codfw * 13:26 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry2004.codfw.wmnet * 13:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry1005.eqiad.wmnet * 13:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2071.codfw.wmnet with reason: host reimage * 13:22 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1076.eqiad.wmnet with reason: host reimage * 13:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry1005.eqiad.wmnet * 13:21 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry1004.eqiad.wmnet * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry1004.eqiad.wmnet * 13:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts * 13:13 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts * 13:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts * 13:12 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts * 13:03 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1076.eqiad.wmnet with OS trixie * 13:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2071.codfw.wmnet with OS trixie * 12:55 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:54 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:53 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:46 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 7 hosts * 12:42 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 7 hosts * 12:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts * 12:42 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts * 12:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2070.codfw.wmnet with OS trixie * 12:36 ozge@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:35 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1075.eqiad.wmnet with OS trixie * 12:32 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts * 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts * 12:22 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin1001.eqiad.wmnet * 12:19 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin1001.eqiad.wmnet * 12:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2070.codfw.wmnet with reason: host reimage * 12:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin2001.codfw.wmnet * 12:14 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1075.eqiad.wmnet with reason: host reimage * 12:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2070.codfw.wmnet with reason: host reimage * 12:10 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1075.eqiad.wmnet with reason: host reimage * 12:09 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin2001.codfw.wmnet * 11:17 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1074.eqiad.wmnet with OS trixie * 11:17 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 11:16 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 11:14 ozge@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 6 hosts * 11:09 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 6 hosts * 11:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 324 hosts * 10:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1074.eqiad.wmnet with reason: host reimage * 10:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2069.codfw.wmnet with OS trixie * 10:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1074.eqiad.wmnet with reason: host reimage * 10:30 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2069.codfw.wmnet with reason: host reimage * 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1074.eqiad.wmnet with OS trixie * 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2069.codfw.wmnet with reason: host reimage * 10:06 blake@deploy2003: Stopping before sync operations * 10:06 blake@deploy2003: Started scap sync-world: Non-deployment scap run to populate new release values * 10:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2069.codfw.wmnet with OS trixie * 10:00 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1073.eqiad.wmnet with OS trixie * 09:56 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 324 hosts * 09:39 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 09:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 8 hosts * 09:38 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1073.eqiad.wmnet with reason: host reimage * 09:37 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 8 hosts * 09:34 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1073.eqiad.wmnet with reason: host reimage * 09:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2068.codfw.wmnet with OS trixie * 09:16 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1073.eqiad.wmnet with OS trixie * 09:13 blake@deploy2003: sync-world aborted: Non-deployment scap run to populate new release values (duration: 00m 02s) * 09:13 blake@deploy2003: Started scap sync-world: Non-deployment scap run to populate new release values * 08:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2068.codfw.wmnet with reason: host reimage * 08:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2068.codfw.wmnet with reason: host reimage * 08:50 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 08:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2068.codfw.wmnet with OS trixie * 08:15 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1072.eqiad.wmnet with OS trixie * 07:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2067.codfw.wmnet with OS trixie * 07:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1072.eqiad.wmnet with reason: host reimage * 07:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1072.eqiad.wmnet with reason: host reimage * 07:45 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 07:45 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 07:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2067.codfw.wmnet with reason: host reimage * 07:35 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2067.codfw.wmnet with reason: host reimage * 07:30 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 07:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1072.eqiad.wmnet with OS trixie * 07:30 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 07:17 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 07:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2067.codfw.wmnet with OS trixie * 05:51 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:50 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:25 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:25 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on db2207.codfw.wmnet with reason: Host down * 04:28 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 07m 02s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-18 == * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 29s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 00:11 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2018.codfw.wmnet, repooling source-only afterwards == 2026-07-17 == * 23:53 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2026.codfw.wmnet, repooling source-only afterwards * 23:09 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2018.codfw.wmnet, repooling source-only afterwards * 23:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2026.codfw.wmnet, repooling source-only afterwards * 22:11 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2018.codfw.wmnet with OS bookworm * 22:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2026.codfw.wmnet with OS bookworm * 21:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2018.codfw.wmnet with reason: host reimage * 21:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2018.codfw.wmnet with reason: host reimage * 21:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2026.codfw.wmnet with reason: host reimage * 21:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2026.codfw.wmnet with reason: host reimage * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2018 * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2018 * 21:26 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2018 * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2018.codfw.wmnet 155.32.192.10.in-addr.arpa 5.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:26 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2018.codfw.wmnet 155.32.192.10.in-addr.arpa 5.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2018 - bking@cumin2003" * 21:26 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2018 - bking@cumin2003" * 21:14 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:13 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2018 * 21:13 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2018.codfw.wmnet with OS bookworm * 21:12 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2026 * 21:12 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2026 * 21:12 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2026.codfw.wmnet with OS bookworm * 21:05 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1022.eqiad.wmnet -> wdqs1026.eqiad.wmnet, repooling source-only afterwards * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs2017.codfw.wmnet, repooling source-only afterwards * 20:11 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs2017.codfw.wmnet, repooling source-only afterwards * 20:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2017.codfw.wmnet with OS bookworm * 20:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1022.eqiad.wmnet -> wdqs1026.eqiad.wmnet, repooling source-only afterwards * 20:06 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1026.eqiad.wmnet with OS bookworm * 19:55 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 19:55 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:55 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 09s) * 19:55 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:50 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 08s) * 19:50 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:50 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 10m 03s) * 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2017.codfw.wmnet with reason: host reimage * 19:40 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:40 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 15s) * 19:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1026.eqiad.wmnet with reason: host reimage * 19:37 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 16s) * 19:37 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:34 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2017.codfw.wmnet with reason: host reimage * 19:34 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1026.eqiad.wmnet with reason: host reimage * 19:33 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 19:33 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:16 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1026.eqiad.wmnet with OS bookworm * 19:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2017 * 19:16 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2017 * 19:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2017.codfw.wmnet with OS bookworm * 18:30 bking@dns1004: END - running authdns-update * 18:28 bking@dns1004: START - running authdns-update * 18:16 kamila@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1264.eqiad.wmnet * 18:16 kamila@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1264.eqiad.wmnet * 18:16 kamila@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1264.eqiad.wmnet * 17:49 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 17:46 dzahn@dns1006: END - running authdns-update * 17:44 dzahn@dns1006: START - running authdns-update * 17:44 dzahn@dns1006: END - running authdns-update * 17:42 dzahn@dns1006: START - running authdns-update * 17:28 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 17:21 kamila@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 17:01 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1264 * 17:01 kamila@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1264 * 17:01 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 17:01 kamila@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1264.eqiad.wmnet * 17:01 kamila@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1264.eqiad.wmnet * 17:01 kamila@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1264.eqiad.wmnet * 16:42 reedy@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] (duration: 10m 29s) * 16:34 reedy@deploy2003: reedy, hartman: Continuing with deployment * 16:33 reedy@deploy2003: reedy, hartman: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:31 reedy@deploy2003: Started scap sync-world: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] * 16:26 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 16:10 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in2001.wikimedia.org with reason: [[phab:T431659|T431659]] * 16:07 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in1001.wikimedia.org with reason: [[phab:T431659|T431659]] * 16:05 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 16:01 kamila@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 16:00 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out2001.wikimedia.org with reason: [[phab:T431659|T431659]] * 15:41 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 15:41 kamila@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 15:35 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out1001.wikimedia.org with reason: [[phab:T431659|T431659]] * 15:14 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1339.eqiad.wmnet * 15:13 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1339.eqiad.wmnet * 15:13 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1339.eqiad.wmnet * 14:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1339.eqiad.wmnet with OS trixie * 14:50 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:49 kamila@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:49 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:33 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage * 14:27 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage * 14:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1339 * 14:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1339 * 14:14 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1339 * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1339.eqiad.wmnet 156.32.64.10.in-addr.arpa 6.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:14 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1339.eqiad.wmnet 156.32.64.10.in-addr.arpa 6.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1339 - cgoubert@cumin2003" * 14:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1339 - cgoubert@cumin2003" * 14:09 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 14:06 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1339 * 14:06 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie * 14:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1339.eqiad.wmnet * 14:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1339.eqiad.wmnet * 14:02 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1339.eqiad.wmnet * 13:45 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb1013.eqiad.wmnet * 13:39 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb1013.eqiad.wmnet * 13:27 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:24 blake@dns1004: END - running authdns-update * 13:22 blake@dns1004: START - running authdns-update * 13:20 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 13:11 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2014.codfw.wmnet * 13:06 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb2014.codfw.wmnet * 13:06 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2012.codfw.wmnet * 13:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 13:03 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 13:01 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 15 hosts * 13:01 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb2012.codfw.wmnet * 13:01 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1016.eqiad.wmnet * 13:00 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 15 hosts * 12:55 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb1016.eqiad.wmnet * 12:55 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1014.eqiad.wmnet * 12:49 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb1014.eqiad.wmnet * 12:32 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:32 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:31 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:31 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1338.eqiad.wmnet * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1338.eqiad.wmnet * 12:18 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1338.eqiad.wmnet * 12:17 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:15 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:14 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:13 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1338.eqiad.wmnet with OS trixie * 12:01 klausman@deploy2003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 11:59 klausman@deploy2003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 11:56 klausman@deploy2003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 11:54 klausman@deploy2003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 11:53 klausman@deploy2003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 11:51 klausman@deploy2003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 11:42 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1338.eqiad.wmnet with reason: host reimage * 11:38 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1338.eqiad.wmnet with reason: host reimage * 11:31 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2230.codfw.wmnet * 11:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1338 * 11:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1338 * 11:25 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1338 * 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1338.eqiad.wmnet 155.32.64.10.in-addr.arpa 5.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:25 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1338.eqiad.wmnet 155.32.64.10.in-addr.arpa 5.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1338 - cgoubert@cumin2003" * 11:25 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1338 - cgoubert@cumin2003" * 11:23 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2230.codfw.wmnet * 11:20 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 11:20 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1338 * 11:20 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1338.eqiad.wmnet with OS trixie * 11:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1338.eqiad.wmnet * 11:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1338.eqiad.wmnet * 11:19 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1338.eqiad.wmnet * 11:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1337.eqiad.wmnet * 11:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1337.eqiad.wmnet * 11:17 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1337.eqiad.wmnet * 11:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1337.eqiad.wmnet with OS trixie * 10:51 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[2001-2002].codfw.wmnet * 10:50 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1337.eqiad.wmnet with reason: host reimage * 10:40 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 10:39 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:39 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1337.eqiad.wmnet with reason: host reimage * 10:39 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1001-1003].eqiad.wmnet * 10:34 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:34 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:30 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:28 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1001-1003].eqiad.wmnet * 10:27 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1337 * 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1337 * 10:26 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1337 * 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1337.eqiad.wmnet 154.32.64.10.in-addr.arpa 4.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:26 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1337.eqiad.wmnet 154.32.64.10.in-addr.arpa 4.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1337 - cgoubert@cumin2003" * 10:26 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1337 - cgoubert@cumin2003" * 10:21 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 10:18 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1337 * 10:17 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1337.eqiad.wmnet with OS trixie * 10:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1337.eqiad.wmnet * 10:16 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db1176.eqiad.wmnet * 10:16 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1337.eqiad.wmnet * 10:16 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1337.eqiad.wmnet * 10:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1336.eqiad.wmnet * 10:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1336.eqiad.wmnet * 10:15 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1336.eqiad.wmnet * 10:11 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db1176.eqiad.wmnet * 10:10 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db1176.eqiad.wmnet * 10:09 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db1176.eqiad.wmnet * 10:05 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts (check the cookbook's logs for more details.) * 10:03 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts (check the cookbook's logs for more details.) * 09:58 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1336.eqiad.wmnet with OS trixie * 09:47 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts (check the cookbook's logs for more details.) * 09:47 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts (check the cookbook's logs for more details.) * 09:45 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host acmechief-test2001.codfw.wmnet,acmechief-test1001.eqiad.wmnet,an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet,db-test[2001-2002].codfw.wmnet,db-test[1 * 09:40 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host acmechief-test2001.codfw.wmnet,acmechief-test1001.eqiad.wmnet,an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet,db-test[2001-2002].codfw.wmnet,db-test[1001-1003].eqiad.wmn * 09:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1336.eqiad.wmnet with reason: host reimage * 09:33 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1336.eqiad.wmnet with reason: host reimage * 09:29 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet * 09:29 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet * 09:28 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:26 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:21 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 09:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1336 * 09:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1336 * 09:19 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 09:14 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1336 * 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1336.eqiad.wmnet 152.32.64.10.in-addr.arpa 2.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1336.eqiad.wmnet 152.32.64.10.in-addr.arpa 2.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1336 - cgoubert@cumin2003" * 09:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1336 - cgoubert@cumin2003" * 09:11 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:10 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 09:09 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:09 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1336 * 09:09 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1336.eqiad.wmnet with OS trixie * 09:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1336.eqiad.wmnet * 09:08 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1336.eqiad.wmnet * 09:08 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1336.eqiad.wmnet * 09:06 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1335.eqiad.wmnet * 09:06 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1335.eqiad.wmnet * 09:06 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1335.eqiad.wmnet * 09:04 elukey: uploaded spicerack_13.1.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia * 08:55 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wikikube-worker-exp2001.codfw.wmnet * 08:54 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host testreduce1002.eqiad.wmnet * 08:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1335.eqiad.wmnet with OS trixie * 08:51 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host wikikube-worker-exp2001.codfw.wmnet * 08:51 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wikikube-worker-exp1001.eqiad.wmnet * 08:50 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host testreduce1002.eqiad.wmnet * 08:45 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host wikikube-worker-exp1001.eqiad.wmnet * 08:34 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1335.eqiad.wmnet with reason: host reimage * 08:30 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1335.eqiad.wmnet with reason: host reimage * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1335 * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1335 * 08:18 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1335 * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1335.eqiad.wmnet 150.32.64.10.in-addr.arpa 0.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:18 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1335.eqiad.wmnet 150.32.64.10.in-addr.arpa 0.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1335 - cgoubert@cumin2003" * 08:18 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1335 - cgoubert@cumin2003" * 08:14 elukey@cumin1003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 08:14 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges * 08:13 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 08:10 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1335 * 08:10 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1335.eqiad.wmnet with OS trixie * 08:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1335.eqiad.wmnet * 08:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1335.eqiad.wmnet * 08:09 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1335.eqiad.wmnet * 08:06 elukey@cumin1003: END (FAIL) - Cookbook sre.puppet.disable-merges (exit_code=99) * 08:05 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges * 08:03 elukey@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin1003.eqiad.wmnet * 07:57 elukey@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin1003.eqiad.wmnet * 07:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetdb1003.eqiad.wmnet * 07:46 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetdb1003.eqiad.wmnet * 07:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetdb2003.codfw.wmnet * 07:37 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetdb2003.codfw.wmnet * 07:37 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1001.eqiad.wmnet * 07:28 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver1001.eqiad.wmnet * 07:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet * 07:19 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet * 07:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2002.codfw.wmnet * 07:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver2002.codfw.wmnet * 07:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet * 07:05 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet * 07:04 btullis@cumin1003: END (FAIL) - Cookbook sre.hadoop.reboot-workers (exit_code=99) for Hadoop analytics cluster * 07:04 elukey@cumin1003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 07:04 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges * 06:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox1003.eqiad.wmnet * 06:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox1003.eqiad.wmnet * 02:46 ryankemper@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:46 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:44 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:37 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:37 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-internal-scholarly,name=eqiad * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 49s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 01:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore wdqs1025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling source-only afterwards * 01:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore wdqs1027 after Bookworm reimage) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs1027.eqiad.wmnet, repooling both afterwards * 00:55 urbanecm@deploy2003: helmfile [codfw] DONE helmfile.d/services/linkrecommendation: apply * 00:54 urbanecm@deploy2003: helmfile [eqiad] DONE helmfile.d/services/linkrecommendation: apply * 00:54 urbanecm@deploy2003: helmfile [staging] DONE helmfile.d/services/linkrecommendation: apply * 00:54 urbanecm@deploy2003: helmfile [codfw] START helmfile.d/services/linkrecommendation: apply * 00:53 urbanecm@deploy2003: helmfile [staging] START helmfile.d/services/linkrecommendation: apply * 00:52 urbanecm@deploy2003: helmfile [eqiad] START helmfile.d/services/linkrecommendation: apply * 00:23 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore wdqs1025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling source-only afterwards * 00:23 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore wdqs1027 after Bookworm reimage) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs1027.eqiad.wmnet, repooling both afterwards * 00:14 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1274.eqiad.wmnet * 00:14 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1274.eqiad.wmnet * 00:14 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1274.eqiad.wmnet * 00:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1027.eqiad.wmnet with OS bookworm * 00:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1025.eqiad.wmnet with OS bookworm * 00:04 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1274.eqiad.wmnet with OS trixie == 2026-07-16 == * 23:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], xfer to freshly reimaged/scap-deployed wdqs2025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs2025.codfw.wmnet, repooling source-only afterwards * 23:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1027.eqiad.wmnet with reason: host reimage * 23:47 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1025.eqiad.wmnet with reason: host reimage * 23:43 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1274.eqiad.wmnet with reason: host reimage * 23:41 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1025.eqiad.wmnet with reason: host reimage * 23:39 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1027.eqiad.wmnet with reason: host reimage * 23:38 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1274.eqiad.wmnet with reason: host reimage * 23:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1025 * 23:23 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1025 * 23:22 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1027 * 23:22 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1027 * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1274 * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1274 * 23:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1025.eqiad.wmnet with OS bookworm * 23:19 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1274 * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1274.eqiad.wmnet 145.48.64.10.in-addr.arpa 5.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:19 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1274.eqiad.wmnet 145.48.64.10.in-addr.arpa 5.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1274 - swfrench@cumin1003" * 23:19 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1274 - swfrench@cumin1003" * 23:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1027.eqiad.wmnet with OS bookworm * 23:14 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 23:14 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1274 * 23:13 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1274.eqiad.wmnet with OS trixie * 23:13 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1274.eqiad.wmnet * 23:12 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1274.eqiad.wmnet * 23:12 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1274.eqiad.wmnet * 23:12 ryankemper: [[phab:T430880|T430880]] depooled dnsdisc of wdqs-internal-scholarly-eqiad bc we only have 1 host there * 23:09 ryankemper@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-internal-scholarly,name=eqiad * 23:08 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1272.eqiad.wmnet * 23:08 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1272.eqiad.wmnet * 23:08 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1272.eqiad.wmnet * 23:01 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], xfer to freshly reimaged/scap-deployed wdqs2025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs2025.codfw.wmnet, repooling source-only afterwards * 22:57 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1272.eqiad.wmnet with OS trixie * 22:56 Amir1: deleting echo notifications from 2015 in group0 * 22:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2025.codfw.wmnet with OS bookworm * 22:35 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1272.eqiad.wmnet with reason: host reimage * 22:32 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 27s) * 22:32 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 22:28 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1269.eqiad.wmnet * 22:28 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1269.eqiad.wmnet * 22:28 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1269.eqiad.wmnet * 22:27 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1272.eqiad.wmnet with reason: host reimage * 22:26 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] (duration: 08m 51s) * 22:22 ladsgroup@deploy2003: ladsgroup, urbanecm: Continuing with deployment * 22:19 ladsgroup@deploy2003: ladsgroup, urbanecm: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:17 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] * 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2025.codfw.wmnet with reason: host reimage * 22:06 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1272 * 22:06 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1272 * 22:05 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1272 * 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1272.eqiad.wmnet 127.48.64.10.in-addr.arpa 7.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:05 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1272.eqiad.wmnet 127.48.64.10.in-addr.arpa 7.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1272 - swfrench@cumin1003" * 22:05 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1272 - swfrench@cumin1003" * 22:03 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2025.codfw.wmnet with reason: host reimage * 22:01 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 22:00 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1272 * 22:00 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1272.eqiad.wmnet with OS trixie * 22:00 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1272.eqiad.wmnet * 21:59 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1272.eqiad.wmnet * 21:59 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1272.eqiad.wmnet * 21:55 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1271.eqiad.wmnet * 21:55 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1271.eqiad.wmnet * 21:55 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1271.eqiad.wmnet * 21:46 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1271.eqiad.wmnet with OS trixie * 21:45 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] (duration: 06m 31s) * 21:40 sbassett@deploy2003: sbassett: Continuing with deployment * 21:40 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2025 * 21:40 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2025 * 21:40 sbassett@deploy2003: sbassett: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:38 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] * 21:37 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2025 * 21:37 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2025.codfw.wmnet 220.48.192.10.in-addr.arpa 0.2.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:37 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2025.codfw.wmnet 220.48.192.10.in-addr.arpa 0.2.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:37 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:37 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2025 - bking@cumin2003" * 21:37 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2025 - bking@cumin2003" * 21:30 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] (duration: 08m 19s) * 21:26 sbassett@deploy2003: sbassett: Continuing with deployment * 21:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1269.eqiad.wmnet with OS trixie * 21:24 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1271.eqiad.wmnet with reason: host reimage * 21:23 sbassett@deploy2003: sbassett: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:22 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:22 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] * 21:20 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2025 * 21:19 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2025.codfw.wmnet with OS bookworm * 21:17 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1271.eqiad.wmnet with reason: host reimage * 21:04 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1269.eqiad.wmnet with reason: host reimage * 21:00 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1269.eqiad.wmnet with reason: host reimage * 20:56 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1271 * 20:55 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1271 * 20:54 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1271 * 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1271.eqiad.wmnet 126.48.64.10.in-addr.arpa 6.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:54 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1271.eqiad.wmnet 126.48.64.10.in-addr.arpa 6.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1271 - swfrench@cumin1003" * 20:54 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1271 - swfrench@cumin1003" * 20:51 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:51 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:51 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:50 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 20:49 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 20:49 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1268.eqiad.wmnet * 20:49 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1268.eqiad.wmnet * 20:49 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1268.eqiad.wmnet * 20:48 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1271 * 20:48 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1271.eqiad.wmnet with OS trixie * 20:47 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1271.eqiad.wmnet * 20:46 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1271.eqiad.wmnet * 20:46 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1271.eqiad.wmnet * 20:41 aude@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] (duration: 07m 34s) * 20:39 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1269 * 20:39 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1269 * 20:36 aude@deploy2003: aude: Continuing with deployment * 20:35 aude@deploy2003: aude: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:33 aude@deploy2003: Started scap sync-world: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] * 20:26 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on db2207.codfw.wmnet with reason: Host down * 20:22 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-video: apply * 20:21 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-video: apply * 20:20 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-timeline: apply * 20:20 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-timeline: apply * 20:20 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-syntaxhighlight: apply * 20:19 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-syntaxhighlight: apply * 20:19 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-media: apply * 20:18 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-media: apply * 20:18 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-constraints: apply * 20:17 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-constraints: apply * 20:17 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox: apply * 20:16 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox: apply * 20:13 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1269 * 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1269.eqiad.wmnet 80.32.64.10.in-addr.arpa 0.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:13 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1269.eqiad.wmnet 80.32.64.10.in-addr.arpa 0.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1269 - kamila@cumin1003" * 20:13 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1269 - kamila@cumin1003" * 20:09 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-flink-codfw cluster: Roll restart of jvm daemons. * 20:07 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 20:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 20:03 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-flink-codfw cluster: Roll restart of jvm daemons. * 20:03 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2207 [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94893 and previous config saved to /var/cache/conftool/dbconfig/20260716-200257-marostegui.json * 20:01 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2204 to s2 primary [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94892 and previous config saved to /var/cache/conftool/dbconfig/20260716-200157-marostegui.json * 20:00 marostegui: Starting emergency s2 codfw failover from db2207 to db2204 - [[phab:T432396|T432396]] * 19:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1035.eqiad.wmnet * 19:56 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2204 with weight 0 [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94891 and previous config saved to /var/cache/conftool/dbconfig/20260716-195628-marostegui.json * 19:55 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 26 hosts with reason: Primary switchover s2 [[phab:T432396|T432396]] * 19:54 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1035.eqiad.wmnet * 19:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1034.eqiad.wmnet * 19:48 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1034.eqiad.wmnet * 19:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1033.eqiad.wmnet * 19:43 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-video: apply * 19:43 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1033.eqiad.wmnet * 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1032.eqiad.wmnet * 19:42 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-video: apply * 19:42 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-timeline: apply * 19:41 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-timeline: apply * 19:41 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-syntaxhighlight: apply * 19:41 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-syntaxhighlight: apply * 19:40 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-media: apply * 19:40 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-media: apply * 19:39 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-constraints: apply * 19:36 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-constraints: apply * 19:36 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox: apply * 19:35 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1032.eqiad.wmnet * 19:35 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1031.eqiad.wmnet * 19:35 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox: apply * 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-video: apply * 19:33 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-video: apply * 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-timeline: apply * 19:33 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-timeline: apply * 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-syntaxhighlight: apply * 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-syntaxhighlight: apply * 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-media: apply * 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-media: apply * 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-constraints: apply * 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-constraints: apply * 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox: apply * 19:31 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox: apply * 19:27 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1031.eqiad.wmnet * 19:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1030.eqiad.wmnet * 19:23 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1001.eqiad.wmnet, repooling source-only afterwards * 19:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1030.eqiad.wmnet * 19:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1029.eqiad.wmnet * 19:17 kamila@cumin1003: START - Cookbook sre.dns.netbox * 19:12 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1029.eqiad.wmnet * 19:06 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1269 * 19:05 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1269.eqiad.wmnet with OS trixie * 19:03 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1269.eqiad.wmnet * 19:03 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1269.eqiad.wmnet * 19:03 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1269.eqiad.wmnet * 18:55 dancy@deploy2003: Finished scap sync-world: testing [[phab:T428971|T428971]] (duration: 02m 41s) * 18:53 dancy@deploy2003: Started scap sync-world: testing [[phab:T428971|T428971]] * 18:31 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1268.eqiad.wmnet with OS trixie * 18:18 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 18:16 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1267.eqiad.wmnet * 18:16 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1267.eqiad.wmnet * 18:16 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1267.eqiad.wmnet * 18:09 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1268.eqiad.wmnet with reason: host reimage * 18:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1001.eqiad.wmnet, repooling source-only afterwards * 18:06 swfrench@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] (duration: 07m 34s) * 18:06 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 23s) * 18:06 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 18:06 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1268.eqiad.wmnet with reason: host reimage * 18:03 bd808@deploy2003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 18:02 bd808@deploy2003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 18:02 swfrench@deploy2003: jiji, swfrench: Continuing with deployment * 18:02 bd808@deploy2003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 18:02 bd808@deploy2003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 18:01 bd808@deploy2003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 18:01 swfrench@deploy2003: jiji, swfrench: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:01 bd808@deploy2003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:59 swfrench@deploy2003: Started scap sync-world: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] * 17:45 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1268 * 17:45 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1268 * 17:44 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1267.eqiad.wmnet with OS trixie * 17:43 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter2006.codfw.wmnet * 17:39 swfrench@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter2006.codfw.wmnet * 17:35 swfrench@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] (duration: 07m 27s) * 17:34 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1268 * 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1268.eqiad.wmnet 78.32.64.10.in-addr.arpa 8.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:34 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1268.eqiad.wmnet 78.32.64.10.in-addr.arpa 8.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1268 - kamila@cumin1003" * 17:34 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1268 - kamila@cumin1003" * 17:31 swfrench@deploy2003: jiji, swfrench: Continuing with deployment * 17:29 swfrench@deploy2003: jiji, swfrench: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:28 kamila@cumin1003: START - Cookbook sre.dns.netbox * 17:28 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1268 * 17:28 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1268.eqiad.wmnet with OS trixie * 17:27 swfrench@deploy2003: Started scap sync-world: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] * 17:23 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1267.eqiad.wmnet with reason: host reimage * 17:18 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1267.eqiad.wmnet with reason: host reimage * 17:18 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1268.eqiad.wmnet * 17:17 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1268.eqiad.wmnet * 17:17 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1268.eqiad.wmnet * 17:12 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter2005.codfw.wmnet * 17:11 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1270.eqiad.wmnet * 17:11 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1270.eqiad.wmnet * 17:11 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1270.eqiad.wmnet * 17:09 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter2005.codfw.wmnet * 17:08 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:08 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update reverse dns for moved arelion cct cr2-eqiad - cmooney@cumin1003" * 17:08 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update reverse dns for moved arelion cct cr2-eqiad - cmooney@cumin1003" * 17:08 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] (duration: 07m 34s) * 17:04 jiji@deploy2003: jiji: Continuing with deployment * 17:03 jiji@deploy2003: jiji: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 17:00 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] * 17:00 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2185.codfw.wmnet with OS trixie * 16:59 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:58 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1270.eqiad.wmnet with OS trixie * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1267 * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1267 * 16:57 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1267 * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1267.eqiad.wmnet 77.32.64.10.in-addr.arpa 7.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:57 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1267.eqiad.wmnet 77.32.64.10.in-addr.arpa 7.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1267 - kamila@cumin1003" * 16:56 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1267 - kamila@cumin1003" * 16:56 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_eqsin * 16:56 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5032.eqsin.wmnet * 16:52 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_esams * 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3073.esams.wmnet * 16:50 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_esams * 16:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3081.esams.wmnet * 16:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1266.eqiad.wmnet * 16:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1266.eqiad.wmnet * 16:45 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1266.eqiad.wmnet * 16:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2185.codfw.wmnet with reason: host reimage * 16:41 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_eqiad * 16:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1114.eqiad.wmnet * 16:41 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_eqiad * 16:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1115.eqiad.wmnet * 16:39 kamila@cumin1003: START - Cookbook sre.dns.netbox * 16:39 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1267 * 16:39 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2185.codfw.wmnet with reason: host reimage * 16:38 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1267.eqiad.wmnet with OS trixie * 16:38 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1267.eqiad.wmnet * 16:38 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1270.eqiad.wmnet with reason: host reimage * 16:37 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1267.eqiad.wmnet * 16:37 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1267.eqiad.wmnet * 16:31 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1270.eqiad.wmnet with reason: host reimage * 16:24 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1264.eqiad.wmnet * 16:24 kamila@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 16:24 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 16:23 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 16:21 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2162: switch maintenance completed codfw rack b6 * 16:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2185.codfw.wmnet with OS trixie * 16:19 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 16:16 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter1007.eqiad.wmnet * 16:15 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_eqsin * 16:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5024.eqsin.wmnet * 16:13 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5031.eqsin.wmnet * 16:13 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3072.esams.wmnet * 16:12 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter1007.eqiad.wmnet * 16:11 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] (duration: 09m 47s) * 16:10 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1266.eqiad.wmnet with OS trixie * 16:10 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1270 * 16:10 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1270 * 16:09 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1270 * 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1270.eqiad.wmnet 125.48.64.10.in-addr.arpa 5.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1270.eqiad.wmnet 125.48.64.10.in-addr.arpa 5.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1270 - swfrench@cumin1003" * 16:09 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1270 - swfrench@cumin1003" * 16:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3080.esams.wmnet * 16:07 jiji@deploy2003: jiji: Continuing with deployment * 16:06 jiji@deploy2003: jiji: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:04 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 16:04 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1265.eqiad.wmnet * 16:03 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1265.eqiad.wmnet * 16:03 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1265.eqiad.wmnet * 16:03 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1270 * 16:03 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1270.eqiad.wmnet with OS trixie * 16:02 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1270.eqiad.wmnet * 16:02 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] * 16:01 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1270.eqiad.wmnet * 16:01 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1270.eqiad.wmnet * 16:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1113.eqiad.wmnet * 16:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1112.eqiad.wmnet * 15:49 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1266.eqiad.wmnet with reason: host reimage * 15:47 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter1006.eqiad.wmnet * 15:45 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1265.eqiad.wmnet with OS trixie * 15:44 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1266.eqiad.wmnet with reason: host reimage * 15:43 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter1006.eqiad.wmnet * 15:42 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] (duration: 09m 46s) * 15:37 jiji@deploy2003: jiji: Continuing with deployment * 15:36 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2162: switch maintenance completed codfw rack b6 * 15:36 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2161: switch maintenance completed codfw rack b6 * 15:34 jiji@deploy2003: jiji: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:32 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5023.eqsin.wmnet * 15:32 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] * 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5030.eqsin.wmnet * 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3071.esams.wmnet * 15:27 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3079.esams.wmnet * 15:25 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1265.eqiad.wmnet with reason: host reimage * 15:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1266 * 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1266 * 15:21 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1110.eqiad.wmnet * 15:20 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1111.eqiad.wmnet * 15:16 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1265.eqiad.wmnet with reason: host reimage * 15:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1001.eqiad.wmnet with OS bookworm * 15:15 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1266 * 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1266.eqiad.wmnet 76.32.64.10.in-addr.arpa 6.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:15 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1266.eqiad.wmnet 76.32.64.10.in-addr.arpa 6.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1266 - kamila@cumin1003" * 15:15 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1266 - kamila@cumin1003" * 15:07 kamila@cumin1003: START - Cookbook sre.dns.netbox * 15:04 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1266 * 15:04 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1264 * 15:04 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1264 * 15:04 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1266.eqiad.wmnet with OS trixie * 15:03 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1264 * 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1264.eqiad.wmnet 74.32.64.10.in-addr.arpa 4.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:03 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1264.eqiad.wmnet 74.32.64.10.in-addr.arpa 4.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1264 - kamila@cumin1003" * 15:03 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1264 - kamila@cumin1003" * 15:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-eqiad * 15:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp1001.eqiad.wmnet * 15:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp1001.eqiad.wmnet * 15:01 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp1001.eqiad.wmnet * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp1001.eqiad.wmnet * 15:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1376-1384].eqiad.wmnet * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1376-1384].eqiad.wmnet * 14:59 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1002.eqiad.wmnet * 14:59 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1266.eqiad.wmnet * 14:58 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1266.eqiad.wmnet * 14:58 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1266.eqiad.wmnet * 14:58 kamila@cumin1003: START - Cookbook sre.dns.netbox * 14:57 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1264 * 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1265 * 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1265 * 14:57 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1310596{{!}}Set $wgMathInternalRestbaseURL explicitly (T349582)]] * 14:57 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1265 * 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1265.eqiad.wmnet 75.32.64.10.in-addr.arpa 5.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:56 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1265.eqiad.wmnet 75.32.64.10.in-addr.arpa 5.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:56 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:56 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1265 - kamila@cumin1003" * 14:56 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1265 - kamila@cumin1003" * 14:53 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1002.eqiad.wmnet * 14:53 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1376-1384].eqiad.wmnet * 14:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1001.eqiad.wmnet with reason: host reimage * 14:51 kamila@cumin1003: START - Cookbook sre.dns.netbox * 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 14:50 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:50 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2161: switch maintenance completed codfw rack b6 * 14:50 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5021.eqsin.wmnet * 14:50 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1265 * 14:49 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:49 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:49 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1265.eqiad.wmnet with OS trixie * 14:49 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5029.eqsin.wmnet * 14:49 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1265.eqiad.wmnet * 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3070.esams.wmnet * 14:48 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1264.eqiad.wmnet * 14:48 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1001.eqiad.wmnet with reason: host reimage * 14:48 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1376-1384].eqiad.wmnet * 14:48 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1265.eqiad.wmnet * 14:47 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1265.eqiad.wmnet * 14:47 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1264.eqiad.wmnet * 14:47 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1264.eqiad.wmnet * 14:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:47 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3078.esams.wmnet * 14:44 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1263.eqiad.wmnet * 14:44 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1263.eqiad.wmnet * 14:44 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1263.eqiad.wmnet * 14:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1108.eqiad.wmnet * 14:40 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1109.eqiad.wmnet * 14:40 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:35 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:34 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:34 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2006.codfw.wmnet * 14:34 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-flink-eqiad cluster: Roll restart of jvm daemons. * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf2002.codfw.wmnet * 14:31 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1002.eqiad.wmnet * 14:29 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2006.codfw.wmnet * 14:27 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:27 kamila@deploy2003: Finished scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] (duration: 02m 57s) * 14:27 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-flink-eqiad cluster: Roll restart of jvm daemons. * 14:26 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf2002.codfw.wmnet * 14:26 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf2001.codfw.wmnet * 14:25 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1002.eqiad.wmnet * 14:25 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1001.eqiad.wmnet * 14:25 kamila@deploy2003: Started scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] * 14:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:21 kamila@deploy2003: sync-world aborted: Test deployment to check rsync is working - [[phab:T432108|T432108]] (duration: 00m 36s) * 14:21 topranks: reboot lsw1-b6-codfw to upgrade JunOS [[phab:T430922|T430922]] * 14:21 kamila@deploy2003: Started scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] * 14:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1001.eqiad.wmnet with OS bookworm * 14:20 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b6-codfw,lsw1-b6-codfw IPv6,lsw1-b6-codfw.mgmt,ssw1-a[1,8]-codfw with reason: lsw1-b6-codfw JunOS upgrade * 14:20 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf2001.codfw.wmnet * 14:19 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1001.eqiad.wmnet * 14:19 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 26 hosts with reason: lsw1-b6-codfw JunOS upgrade * 14:14 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 14:13 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc2022: switch maintenance codfw rack b6 * 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:12 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.parsercache * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool pc2022: switch maintenance codfw rack b6 * 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2251: switch maintenance codfw rack b6 * 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.parsercache * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2251: switch maintenance codfw rack b6 * 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2162: switch maintenance codfw rack b6 * 14:12 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1263.eqiad.wmnet with OS trixie * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2162: switch maintenance codfw rack b6 * 14:11 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2161: switch maintenance codfw rack b6 * 14:11 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2161: switch maintenance codfw rack b6 * 14:08 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5020.eqsin.wmnet * 14:07 btullis@cumin1003: START - Cookbook sre.hadoop.reboot-workers for Hadoop analytics cluster * 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3069.esams.wmnet * 14:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5028.eqsin.wmnet * 14:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1338-1347].eqiad.wmnet * 14:06 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1338-1347].eqiad.wmnet * 14:05 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3077.esams.wmnet * 14:02 topranks: beginning depools for lsw1-b6-codfw maintenance [[phab:T430922|T430922]] * 14:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1106.eqiad.wmnet * 14:00 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-misc1002.eqiad.wmnet * 13:59 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1338-1347].eqiad.wmnet * 13:59 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1107.eqiad.wmnet * 13:56 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-codfw * 13:55 sfaci@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply * 13:54 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-misc1002.eqiad.wmnet * 13:54 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-misc1001.eqiad.wmnet * 13:54 sfaci@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply * 13:50 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1263.eqiad.wmnet with reason: host reimage * 13:50 sfaci@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 13:49 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1338-1347].eqiad.wmnet * 13:49 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-misc1001.eqiad.wmnet * 13:49 sfaci@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 13:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:49 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:45 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1263.eqiad.wmnet with reason: host reimage * 13:40 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:40 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-eqiad * 13:35 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:34 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:33 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-reboot (exit_code=0) rolling reboot on A:dnsbox and (A:eqsin or A:drmrs or A:magru) and not (P<nowiki>{</nowiki>dns5003*<nowiki>}</nowiki> or P<nowiki>{</nowiki>dns7002*<nowiki>}</nowiki>) and (A:dnsbox) * 13:33 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns7001.wikimedia.org * 13:27 sukhe@dns1004: END - running authdns-update * 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5019.eqsin.wmnet * 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3076.esams.wmnet * 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3068.esams.wmnet * 13:25 sukhe@dns1004: START - running authdns-update * 13:24 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5027.eqsin.wmnet * 13:24 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1263 * 13:24 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1263 * 13:23 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1263 * 13:23 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:23 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1104.eqiad.wmnet * 13:21 kamila@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:21 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:20 kamila@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:20 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:20 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:20 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1263 - kamila@cumin1003" * 13:20 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1263 - kamila@cumin1003" * 13:19 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1105.eqiad.wmnet * 13:19 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:19 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:18 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:18 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns7001.wikimedia.org * 13:16 cdobbins@cumin2003: conftool action : set/pooled=yes; selector: name=dns7002.* * 13:14 cdobbins@dns1004: END - running authdns-update * 13:13 sbisson@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] (duration: 08m 03s) * 13:13 cdobbins@dns1004: START - running authdns-update * 13:12 kamila@cumin1003: START - Cookbook sre.dns.netbox * 13:12 cdobbins@cumin2003: conftool action : set/pooled=yes; selector: name=dns7002.*,service=authdns-update * 13:12 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1263 * 13:11 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1263.eqiad.wmnet with OS trixie * 13:11 cdobbins@cumin2003: conftool action : set/pooled=no; selector: name=dns7002.* * 13:11 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1263.eqiad.wmnet * 13:10 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1263.eqiad.wmnet * 13:10 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1263.eqiad.wmnet * 13:09 sbisson@deploy2003: sbisson: Continuing with deployment * 13:08 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:07 sbisson@deploy2003: sbisson: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:05 sbisson@deploy2003: Started scap sync-world: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] * 13:03 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns6002.wikimedia.org * 13:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1298-1307].eqiad.wmnet * 13:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1298-1307].eqiad.wmnet * 12:59 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1262.eqiad.wmnet * 12:59 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1262.eqiad.wmnet * 12:59 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1262.eqiad.wmnet * 12:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1298-1307].eqiad.wmnet * 12:49 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns6002.wikimedia.org * 12:46 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1298-1307].eqiad.wmnet * 12:46 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:46 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3075.esams.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3067.esams.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5018.eqsin.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5026.eqsin.wmnet * 12:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1102.eqiad.wmnet * 12:39 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1103.eqiad.wmnet * 12:35 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:34 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns6001.wikimedia.org * 12:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:28 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:18 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns6001.wikimedia.org * 12:14 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1267-1276].eqiad.wmnet * 12:13 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1267-1276].eqiad.wmnet * 12:04 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1267-1276].eqiad.wmnet * 12:03 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns5004.wikimedia.org * 12:02 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1100.eqiad.wmnet * 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3066.esams.wmnet * 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3074.esams.wmnet * 12:01 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:01 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-codfw * 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5017.eqsin.wmnet * 12:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5025.eqsin.wmnet * 12:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1101.eqiad.wmnet * 11:59 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1267-1276].eqiad.wmnet * 11:58 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:58 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:54 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns5004.wikimedia.org * 11:54 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and (A:eqsin or A:drmrs or A:magru) and not (P<nowiki>{</nowiki>dns5003*<nowiki>}</nowiki> or P<nowiki>{</nowiki>dns7002*<nowiki>}</nowiki>) and (A:dnsbox) * 11:54 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:53 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:53 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-eqiad * 11:51 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:51 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:50 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:50 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_eqiad * 11:50 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:50 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_eqiad * 11:49 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_esams * 11:49 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_esams * 11:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_eqsin * 11:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_eqsin * 11:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:44 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-eqiad * 11:43 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-codfw * 11:42 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:41 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:24 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-eqiad * 11:23 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-codfw * 11:22 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:15 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:14 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:09 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2066.codfw.wmnet with OS trixie * 11:05 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1151-1160].eqiad.wmnet * 11:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1151-1160].eqiad.wmnet * 10:59 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1068.eqiad.wmnet with OS trixie * 10:55 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.major-upgrade (exit_code=99) * 10:55 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 10:54 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1151-1160].eqiad.wmnet * 10:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2066.codfw.wmnet with reason: host reimage * 10:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1151-1160].eqiad.wmnet * 10:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:42 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2066.codfw.wmnet with reason: host reimage * 10:39 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:37 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:36 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:23 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:22 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2066.codfw.wmnet with OS trixie * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:07 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2065.codfw.wmnet with OS trixie * 10:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:06 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 10:06 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 10:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 10:03 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:03 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 10:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 09:59 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 09:57 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 09:57 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 09:52 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 09:47 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:46 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:46 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2065.codfw.wmnet with reason: host reimage * 09:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2065.codfw.wmnet with reason: host reimage * 09:40 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:39 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 09:39 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:39 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:39 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:37 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:29 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox2003.codfw.wmnet * 09:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox2003.codfw.wmnet * 09:25 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:25 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:24 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:24 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:24 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:21 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:20 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2065.codfw.wmnet with OS trixie * 09:13 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1068.eqiad.wmnet with OS trixie * 09:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2064.codfw.wmnet with OS trixie * 09:08 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 09:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 09:07 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2162: Repooling after switchover * 09:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1067.eqiad.wmnet with OS trixie * 09:00 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:59 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 08:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 08:57 tappof: bump space for prometheus k8s-dse in eqiad * 08:56 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping2004.codfw.wmnet * 08:52 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host ping2004.codfw.wmnet * 08:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 08:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping1004.eqiad.wmnet * 08:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:51 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:49 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 08:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host ping1004.eqiad.wmnet * 08:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2064.codfw.wmnet with reason: host reimage * 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2064.codfw.wmnet with reason: host reimage * 08:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1067.eqiad.wmnet with reason: host reimage * 08:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:33 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1067.eqiad.wmnet with reason: host reimage * 08:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:21 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2162: Repooling after switchover * 08:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2064.codfw.wmnet with OS trixie * 08:16 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1067.eqiad.wmnet with OS trixie * 08:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:15 cgoubert@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-eqiad * 08:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2062.codfw.wmnet with OS trixie * 08:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1066.eqiad.wmnet with OS trixie * 08:02 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2162: Repooling after switchover * 07:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2162: Repooling after switchover * 07:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2162 [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94870 and previous config saved to /var/cache/conftool/dbconfig/20260716-075530-cwilliams.json * 07:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2241 to x3 primary [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94869 and previous config saved to /var/cache/conftool/dbconfig/20260716-075314-cwilliams.json * 07:52 cezmunsta: Starting x3 codfw failover from db2162 to db2241 - [[phab:T430925|T430925]] * 07:50 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:50 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:47 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 07:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2241 with weight 0 [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94868 and previous config saved to /var/cache/conftool/dbconfig/20260716-074507-cwilliams.json * 07:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 18 hosts with reason: Primary switchover x3 [[phab:T430925|T430925]] * 07:43 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1066.eqiad.wmnet with reason: host reimage * 07:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:dse-k8s-worker-eqiad * 07:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1028.eqiad.wmnet * 07:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1028.eqiad.wmnet * 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 07:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1066.eqiad.wmnet with reason: host reimage * 07:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1028.eqiad.wmnet * 07:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1028.eqiad.wmnet * 07:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1027.eqiad.wmnet * 07:35 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1027.eqiad.wmnet * 07:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1027.eqiad.wmnet * 07:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1027.eqiad.wmnet * 07:28 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1026.eqiad.wmnet * 07:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1026.eqiad.wmnet * 07:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast2003.wikimedia.org * 07:21 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1026.eqiad.wmnet * 07:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1066.eqiad.wmnet with OS trixie * 07:19 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast2003.wikimedia.org * 07:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2062.codfw.wmnet with OS trixie * 06:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1026.eqiad.wmnet * 06:51 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1025.eqiad.wmnet * 06:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1025.eqiad.wmnet * 06:47 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 06:47 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 06:44 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1025.eqiad.wmnet * 06:14 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1025.eqiad.wmnet * 06:14 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1024.eqiad.wmnet * 06:14 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1024.eqiad.wmnet * 06:07 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1024.eqiad.wmnet * 05:37 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1024.eqiad.wmnet * 05:37 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1023.eqiad.wmnet * 05:37 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1023.eqiad.wmnet * 05:26 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1023.eqiad.wmnet * 04:56 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1023.eqiad.wmnet * 04:56 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1022.eqiad.wmnet * 04:56 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1022.eqiad.wmnet * 04:49 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1022.eqiad.wmnet * 04:19 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1022.eqiad.wmnet * 04:19 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1021.eqiad.wmnet * 04:19 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1021.eqiad.wmnet * 04:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1021.eqiad.wmnet * 03:38 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1021.eqiad.wmnet * 03:38 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1020.eqiad.wmnet * 03:38 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1020.eqiad.wmnet * 03:20 btullis@cumin1003: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1020.eqiad.wmnet * 03:18 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1020.eqiad.wmnet * 03:18 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1019.eqiad.wmnet * 03:18 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1019.eqiad.wmnet * 03:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1019.eqiad.wmnet * 02:41 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1019.eqiad.wmnet * 02:41 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1018.eqiad.wmnet * 02:41 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1018.eqiad.wmnet * 02:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling both afterwards * 02:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2003.codfw.wmnet -> wcqs2001.codfw.wmnet, repooling both afterwards * 02:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1018.eqiad.wmnet * 02:30 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1018.eqiad.wmnet * 02:30 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1014.eqiad.wmnet * 02:30 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1014.eqiad.wmnet * 02:24 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1014.eqiad.wmnet * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 01:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1014.eqiad.wmnet * 01:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1013.eqiad.wmnet * 01:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1013.eqiad.wmnet * 01:47 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1013.eqiad.wmnet * 01:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2003.codfw.wmnet -> wcqs2001.codfw.wmnet, repooling both afterwards * 01:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling both afterwards * 01:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1013.eqiad.wmnet * 01:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1012.eqiad.wmnet * 01:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1012.eqiad.wmnet * 01:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1012.eqiad.wmnet * 01:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1012.eqiad.wmnet * 01:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1011.eqiad.wmnet * 01:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1011.eqiad.wmnet * 01:08 ryankemper@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] scap deploy post bookworm reimage (duration: 00m 23s) * 01:08 ryankemper@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] scap deploy post bookworm reimage * 01:08 ryankemper@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): scap deploy post bookworm reimage (duration: 00m 46s) * 01:07 ryankemper@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): scap deploy post bookworm reimage * 01:04 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1011.eqiad.wmnet * 01:04 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1011.eqiad.wmnet * 01:04 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1010.eqiad.wmnet * 01:04 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1010.eqiad.wmnet * 00:57 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1010.eqiad.wmnet * 00:57 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1010.eqiad.wmnet * 00:57 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1009.eqiad.wmnet * 00:57 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1009.eqiad.wmnet * 00:50 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1009.eqiad.wmnet * 00:20 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1009.eqiad.wmnet * 00:20 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1008.eqiad.wmnet * 00:20 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1008.eqiad.wmnet * 00:13 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1008.eqiad.wmnet == 2026-07-15 == * 23:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2001.codfw.wmnet with OS bookworm * 23:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1008.eqiad.wmnet * 23:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1007.eqiad.wmnet * 23:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1007.eqiad.wmnet * 23:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1007.eqiad.wmnet * 23:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1007.eqiad.wmnet * 23:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1006.eqiad.wmnet * 23:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1006.eqiad.wmnet * 23:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1006.eqiad.wmnet * 23:29 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1006.eqiad.wmnet * 23:28 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1005.eqiad.wmnet * 23:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1005.eqiad.wmnet * 23:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1002.eqiad.wmnet with OS bookworm * 23:21 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1005.eqiad.wmnet * 23:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2001.codfw.wmnet with reason: host reimage * 23:15 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host datahubsearch1001.eqiad.wmnet with OS bookworm * 23:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2001.codfw.wmnet with reason: host reimage * 23:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 23:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 22:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 22:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1005.eqiad.wmnet * 22:51 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1004.eqiad.wmnet * 22:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1004.eqiad.wmnet * 22:45 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1004.eqiad.wmnet * 22:44 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 22:44 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS trixie * 22:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host datahubsearch1001.eqiad.wmnet with OS bookworm * 22:34 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host datahubsearch1001.eqiad.wmnet with OS bookworm * 22:16 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on datahubsearch[1002-1003].eqiad.wmnet with reason: Using datahubsearch1001 to test bookworm reimages * 22:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1004.eqiad.wmnet * 22:15 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1003.eqiad.wmnet * 22:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1003.eqiad.wmnet * 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 22:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1003.eqiad.wmnet * 22:08 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1003.eqiad.wmnet * 22:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1002.eqiad.wmnet * 22:08 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1002.eqiad.wmnet * 22:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host datahubsearch1001.eqiad.wmnet with OS bookworm * 22:05 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 22:02 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm * 22:01 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on datahubsearch[1001-1003].eqiad.wmnet with reason: Using datahubsearch1001 to test bookworm reimages * 22:01 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1002.eqiad.wmnet * 22:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1002.eqiad.wmnet * 22:00 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1001.eqiad.wmnet * 22:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1001.eqiad.wmnet * 21:53 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1001.eqiad.wmnet * 21:52 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 21:50 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 21:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS trixie * 21:50 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS bookworm * 21:43 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 21:38 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:30 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wcqs1002'] * 21:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:29 lerickson@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 21:29 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:29 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:29 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS bookworm * 21:28 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 21:28 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm * 21:23 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1001.eqiad.wmnet * 21:23 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:23 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:22 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 21:20 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:18 swfrench-wmf: reprepro include php8.3_8.3.32-1+wmf11u2 into component/php83 for bullseye-wikimedia * 21:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:16 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:15 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-druid-public cluster: Roll restart of jvm daemons. * 21:08 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:05 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1001.eqiad.wmnet * 21:05 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1001.eqiad.wmnet * 21:04 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-druid-public cluster: Roll restart of jvm daemons. * 21:02 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 21:01 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 21:01 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 21:00 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 20:59 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1001.eqiad.wmnet * 20:59 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1001.eqiad.wmnet * 20:59 btullis@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:dse-k8s-worker-eqiad * 20:55 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 20:55 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 20:45 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm * 20:21 jhathaway: puppet is re-enabled, have fun, but not too much fun! * 20:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 20:17 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs2001'] * 20:12 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs2001'] * 20:11 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs2001'] * 20:09 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:08 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 20:05 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:05 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 20:04 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs2001'] * 20:03 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 20:03 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm * 20:02 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:02 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 20:01 jhathaway: disabling puppet fleet wide to roll out kafka patch * 19:55 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host relforge1010.eqiad.wmnet * 19:52 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 19:52 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 19:48 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 19:48 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 19:48 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 19:47 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 19:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 19:45 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 19:45 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 19:44 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1010.eqiad.wmnet * 19:38 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1262.eqiad.wmnet with OS trixie * 19:17 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1262.eqiad.wmnet with reason: host reimage * 19:11 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1262.eqiad.wmnet with reason: host reimage * 18:59 cdobbins@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS trixie * 18:54 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 18:53 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 18:52 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1262 * 18:52 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1262 * 18:51 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1262 * 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1262.eqiad.wmnet 72.32.64.10.in-addr.arpa 2.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:51 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1262.eqiad.wmnet 72.32.64.10.in-addr.arpa 2.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1262 - kamila@cumin1003" * 18:51 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1262 - kamila@cumin1003" * 18:46 kamila@cumin1003: START - Cookbook sre.dns.netbox * 18:46 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1262 * 18:46 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 18:46 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ncmonitor1001.eqiad.wmnet * 18:46 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 18:45 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1262.eqiad.wmnet with OS trixie * 18:45 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 18:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1262.eqiad.wmnet * 18:44 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1262.eqiad.wmnet * 18:44 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1262.eqiad.wmnet * 18:42 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host ncmonitor1001.eqiad.wmnet * 18:29 topranks: pull power on cr1-eqiad to install new switch-control boards [[phab:T426343|T426343]] * 18:29 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs[1018-1020].eqiad.wmnet with reason: line card install in cr1-eqiad * 18:27 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 14 hosts with reason: linecard install in cr1-eqad * 18:22 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_ulsfo * 18:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4052.ulsfo.wmnet * 18:19 cdobbins@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 18:15 cdobbins@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 18:14 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_drmrs * 18:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6016.drmrs.wmnet * 18:12 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_ulsfo * 18:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4044.ulsfo.wmnet * 18:10 sukhe@cumin1003: END (ERROR) - Cookbook sre.cdn.roll-reboot (exit_code=97) rolling reboot on A:cp-upload_drmrs * 18:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2241: Security update * 17:56 topranks: start draining traffic on cr1-eqiad ahead of line card installation [[phab:T426343|T426343]] * 17:47 cdobbins@cumin2003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie * 17:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4051.ulsfo.wmnet * 17:40 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:39 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 17:34 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6007.drmrs.wmnet * 17:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6015.drmrs.wmnet * 17:32 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:31 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 17:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4043.ulsfo.wmnet * 17:27 lerickson@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:25 lerickson@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 17:22 lerickson@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-codfw * 17:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp2001.codfw.wmnet * 17:22 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 17:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp2001.codfw.wmnet * 17:22 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 17:19 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2241: Security update * 17:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2241.codfw.wmnet * 17:17 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2241.codfw.wmnet * 17:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp2001.codfw.wmnet * 17:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp2001.codfw.wmnet * 17:15 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2366-2374].codfw.wmnet * 17:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2366-2374].codfw.wmnet * 17:10 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply * 17:10 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply * 17:08 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2366-2374].codfw.wmnet * 17:06 sukhe: sre.dns.roll-reboot to resume later * 17:06 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-reboot (exit_code=97) rolling reboot on A:dnsbox and not (A:ulsfo or A:magru) and (A:dnsbox) * 17:06 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns5003.wikimedia.org * 17:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2241: Security update * 17:03 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2241: Security update * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply * 17:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2366-2374].codfw.wmnet * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply * 17:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2357-2365].codfw.wmnet * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 17:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2357-2365].codfw.wmnet * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply * 16:55 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2357-2365].codfw.wmnet * 16:55 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 16:53 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6006.drmrs.wmnet * 16:52 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6014.drmrs.wmnet * 16:52 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 16:51 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 16:50 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2357-2365].codfw.wmnet * 16:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4042.ulsfo.wmnet * 16:50 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2347-2356].codfw.wmnet * 16:50 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2347-2356].codfw.wmnet * 16:49 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns5003.wikimedia.org * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply * 16:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4050.ulsfo.wmnet * 16:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2347-2356].codfw.wmnet * 16:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2347-2356].codfw.wmnet * 16:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2337-2346].codfw.wmnet * 16:36 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2337-2346].codfw.wmnet * 16:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:dse-k8s-worker-codfw * 16:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2003.codfw.wmnet * 16:35 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2003.codfw.wmnet * 16:34 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns3004.wikimedia.org * 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply * 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply * 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply * 16:30 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply * 16:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2003.codfw.wmnet * 16:29 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2337-2346].codfw.wmnet * 16:24 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2003.codfw.wmnet * 16:24 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2002.codfw.wmnet * 16:24 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2002.codfw.wmnet * 16:23 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns3004.wikimedia.org * 16:23 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2337-2346].codfw.wmnet * 16:23 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2327-2336].codfw.wmnet * 16:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2327-2336].codfw.wmnet * 16:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2002.codfw.wmnet * 16:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2327-2336].codfw.wmnet * 16:12 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2002.codfw.wmnet * 16:12 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2001.codfw.wmnet * 16:12 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2001.codfw.wmnet * 16:12 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1065.eqiad.wmnet with OS trixie * 16:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6005.drmrs.wmnet * 16:11 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6013.drmrs.wmnet * 16:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4041.ulsfo.wmnet * 16:08 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns3003.wikimedia.org * 16:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2327-2336].codfw.wmnet * 16:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2317-2326].codfw.wmnet * 16:06 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2317-2326].codfw.wmnet * 16:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2001.codfw.wmnet * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply * 16:03 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4049.ulsfo.wmnet * 16:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2001.codfw.wmnet * 16:00 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test2001.codfw.wmnet * 16:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test2001.codfw.wmnet * 16:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2063.codfw.wmnet with OS trixie * 15:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2317-2326].codfw.wmnet * 15:57 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns3003.wikimedia.org * 15:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test2001.codfw.wmnet * 15:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test2001.codfw.wmnet * 15:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2004.codfw.wmnet * 15:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2004.codfw.wmnet * 15:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2317-2326].codfw.wmnet * 15:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2307-2316].codfw.wmnet * 15:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2307-2316].codfw.wmnet * 15:49 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2004.codfw.wmnet * 15:48 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2004.codfw.wmnet * 15:48 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2003.codfw.wmnet * 15:48 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2003.codfw.wmnet * 15:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 15:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2307-2316].codfw.wmnet * 15:42 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2003.codfw.wmnet * 15:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 15:42 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2003.codfw.wmnet * 15:42 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2002.codfw.wmnet * 15:42 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2002.codfw.wmnet * 15:42 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2006.wikimedia.org * 15:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2063.codfw.wmnet with reason: host reimage * 15:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2307-2316].codfw.wmnet * 15:37 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2297-2306].codfw.wmnet * 15:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2297-2306].codfw.wmnet * 15:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2002.codfw.wmnet * 15:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2002.codfw.wmnet * 15:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2001.codfw.wmnet * 15:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2001.codfw.wmnet * 15:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2063.codfw.wmnet with reason: host reimage * 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6004.drmrs.wmnet * 15:31 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2001.codfw.wmnet * 15:31 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2001.codfw.wmnet * 15:31 btullis@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:dse-k8s-worker-codfw * 15:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6012.drmrs.wmnet * 15:28 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2006.wikimedia.org * 15:27 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-analytics cluster: Roll restart of jvm daemons. * 15:27 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2297-2306].codfw.wmnet * 15:27 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4040.ulsfo.wmnet * 15:24 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1065.eqiad.wmnet with OS trixie * 15:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4048.ulsfo.wmnet * 15:21 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-analytics cluster: Roll restart of jvm daemons. * 15:21 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2297-2306].codfw.wmnet * 15:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2287-2296].codfw.wmnet * 15:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2287-2296].codfw.wmnet * 15:20 btullis@cumin1003: END (PASS) - Cookbook sre.druid.reboot-workers (exit_code=0) for Druid public cluster: Reboot Druid nodes * 15:18 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 15:17 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm * 15:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2063.codfw.wmnet with OS trixie * 15:13 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2005.wikimedia.org * 15:11 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2287-2296].codfw.wmnet * 15:11 btullis@cumin1003: END (PASS) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=0) rolling reboot on A:cephosd-eqiad * 15:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1064.eqiad.wmnet with OS trixie * 15:05 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2062.codfw.wmnet with OS trixie * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2287-2296].codfw.wmnet * 15:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2277-2286].codfw.wmnet * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2277-2286].codfw.wmnet * 14:59 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2005.wikimedia.org * 14:57 brouberol@cumin1003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-jumbo-eqiad * 14:52 btullis@cumin1003: END (PASS) - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas (exit_code=0) rolling reboot on A:schema-codfw * 14:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6003.drmrs.wmnet * 14:50 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:50 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host relforge1009.eqiad.wmnet * 14:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2277-2286].codfw.wmnet * 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6011.drmrs.wmnet * 14:47 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 14:46 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>ml-serve1001.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 14:46 klausman@cumin1003: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) pool for host ml-serve1001.eqiad.wmnet * 14:46 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 14:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1001.eqiad.wmnet * 14:45 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4039.ulsfo.wmnet * 14:44 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1009.eqiad.wmnet * 14:44 btullis@cumin1003: START - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas rolling reboot on A:schema-codfw * 14:44 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2004.wikimedia.org * 14:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2277-2286].codfw.wmnet * 14:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2267-2276].codfw.wmnet * 14:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2267-2276].codfw.wmnet * 14:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 14:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4047.ulsfo.wmnet * 14:40 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1001.eqiad.wmnet * 14:38 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 14:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 14:36 topranks: disconnect power on cr2-eqiad to shut down device for switch fabric replacement [[phab:T426343|T426343]] * 14:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2267-2276].codfw.wmnet * 14:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 14:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1001.eqiad.wmnet * 14:35 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>ml-serve1001.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 14:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 14:34 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 14:33 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:33 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:30 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2004.wikimedia.org * 14:29 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2267-2276].codfw.wmnet * 14:29 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2257-2266].codfw.wmnet * 14:29 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2257-2266].codfw.wmnet * 14:24 btullis@cumin1003: END (PASS) - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas (exit_code=0) rolling reboot on A:schema-eqiad * 14:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2257-2266].codfw.wmnet * 14:20 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:20 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:19 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:17 jforrester@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2257-2266].codfw.wmnet * 14:16 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:16 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2062.codfw.wmnet with OS trixie * 14:15 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1064.eqiad.wmnet with OS trixie * 14:15 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1006.wikimedia.org * 14:15 btullis@cumin1003: START - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas rolling reboot on A:schema-eqiad * 14:14 topranks: switch routing-engine on cr2-eqiad resetting all interfaces [[phab:T417873|T417873]] * 14:11 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:11 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:10 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm * 14:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6002.drmrs.wmnet * 14:09 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6010.drmrs.wmnet * 14:06 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1006.wikimedia.org * 14:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:05 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 14:05 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4038.ulsfo.wmnet * 14:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4046.ulsfo.wmnet * 14:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:00 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on cr1-eqiad with reason: switch upgrade and line card install * 14:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:59 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:57 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:57 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:55 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-eqiad * 13:55 btullis@cumin1003: START - Cookbook sre.druid.reboot-workers for Druid public cluster: Reboot Druid nodes * 13:53 topranks: switch routing-engine on cr2-eqiad resetting all interfaces [[phab:T417873|T417873]] * 13:51 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1005.wikimedia.org * 13:50 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:49 brouberol@cumin1003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-test-eqiad * 13:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:44 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2197-2206].codfw.wmnet * 13:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2197-2206].codfw.wmnet * 13:36 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1005.wikimedia.org * 13:35 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2197-2206].codfw.wmnet * 13:30 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2197-2206].codfw.wmnet * 13:28 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6001.drmrs.wmnet * 13:28 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6009.drmrs.wmnet * 13:28 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2187-2196].codfw.wmnet * 13:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2187-2196].codfw.wmnet * 13:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 13:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4037.ulsfo.wmnet * 13:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2001 * 13:22 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2001 * 13:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4045.ulsfo.wmnet * 13:21 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1004.wikimedia.org * 13:19 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on lvs[1018-1020].eqiad.wmnet with reason: switch upgrade and line card install * 13:18 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2009.codfw.wmnet * 13:18 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2001 * 13:18 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2001.codfw.wmnet 26.16.192.10.in-addr.arpa 6.2.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:17 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2001.codfw.wmnet 26.16.192.10.in-addr.arpa 6.2.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:17 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:17 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2001 - bking@cumin2003" * 13:17 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2001 - bking@cumin2003" * 13:17 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2009.codfw.wmnet * 13:17 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_drmrs * 13:17 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2187-2196].codfw.wmnet * 13:17 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_drmrs * 13:17 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on 15 hosts with reason: switch upgrade and line card install * 13:17 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:15 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 13:13 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:13 brouberol@cumin1003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-jumbo-eqiad * 13:13 brouberol@cumin1003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-test-eqiad * 13:13 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1004.wikimedia.org * 13:13 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and not (A:ulsfo or A:magru) and (A:dnsbox) * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:12 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_ulsfo * 13:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:12 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_ulsfo * 13:11 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2187-2196].codfw.wmnet * 13:11 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 13:11 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 13:06 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 13:05 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling source-only afterwards * 13:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2001 * 13:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2009.codfw.wmnet with OS trixie * 13:03 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:03 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling source-only afterwards * 13:01 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 15s) * 13:01 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 13:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 12:57 btullis@cumin1003: END (PASS) - Cookbook sre.druid.reboot-workers (exit_code=0) for Druid analytics cluster: Reboot Druid nodes * 12:54 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 12:54 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2163-2172].codfw.wmnet * 12:54 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2163-2172].codfw.wmnet * 12:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2163-2172].codfw.wmnet * 12:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2009.codfw.wmnet with reason: host reimage * 12:41 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2163-2172].codfw.wmnet * 12:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2153-2162].codfw.wmnet * 12:40 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2153-2162].codfw.wmnet * 12:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2009.codfw.wmnet with reason: host reimage * 12:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2153-2162].codfw.wmnet * 12:29 btullis@cumin1003: END (PASS) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=0) rolling reboot on A:cephosd-codfw * 12:25 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2153-2162].codfw.wmnet * 12:25 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2143-2152].codfw.wmnet * 12:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2143-2152].codfw.wmnet * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2009 * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2009 * 12:22 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2009 * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2009.codfw.wmnet 139.0.192.10.in-addr.arpa 9.3.1.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:22 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2009.codfw.wmnet 139.0.192.10.in-addr.arpa 9.3.1.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2009 - mvernon@cumin2003" * 12:22 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2009 - mvernon@cumin2003" * 12:16 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 12:15 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 12:15 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 12:15 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2009 * 12:15 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 12:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2009.codfw.wmnet with OS trixie * 12:15 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 12:14 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 12:14 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2143-2152].codfw.wmnet * 12:13 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 12:12 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2010.codfw.wmnet * 12:11 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2010.codfw.wmnet * 12:10 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 12:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2143-2152].codfw.wmnet * 12:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2133-2142].codfw.wmnet * 12:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2133-2142].codfw.wmnet * 12:02 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 11:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2133-2142].codfw.wmnet * 11:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2133-2142].codfw.wmnet * 11:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:49 mvolz@deploy2003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:49 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-codfw * 11:48 mvolz@deploy2003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:47 btullis@cumin1003: START - Cookbook sre.druid.reboot-workers for Druid analytics cluster: Reboot Druid nodes * 11:46 mvolz@deploy2003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:46 mvolz@deploy2003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:45 mvolz@deploy2003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:44 mvolz@deploy2003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:40 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] (duration: 11m 38s) * 11:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2010.codfw.wmnet with OS trixie * 11:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2105-2114].codfw.wmnet * 11:36 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2105-2114].codfw.wmnet * 11:36 krinkle@deploy2003: physikerwelt, krinkle: Continuing with deployment * 11:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1018: Security updates * 11:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:36 root@cumin1003: START - Cookbook sre.mysql.parsercache * 11:36 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1018: Security updates * 11:31 krinkle@deploy2003: physikerwelt, krinkle: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:29 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] * 11:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2105-2114].codfw.wmnet * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2105-2114].codfw.wmnet * 11:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2010.codfw.wmnet with reason: host reimage * 11:12 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2010.codfw.wmnet with reason: host reimage * 11:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1018: Security updates * 11:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:10 root@cumin1003: START - Cookbook sre.mysql.parsercache * 11:10 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1018: Security updates * 11:09 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1009.eqiad.wmnet with OS trixie * 11:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow7002.magru.wmnet * 11:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 11:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 11:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=tegola-vector-tiles,name=eqiad * 11:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=kartotherian,name=eqiad * 11:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow7002.magru.wmnet * 10:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2010 * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2010 * 10:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 10:54 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 10:54 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2010 * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2010.codfw.wmnet 76.16.192.10.in-addr.arpa 6.7.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:54 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2010.codfw.wmnet 76.16.192.10.in-addr.arpa 6.7.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2010 - mvernon@cumin2003" * 10:54 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2010 - mvernon@cumin2003" * 10:49 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 10:49 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2010 * 10:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1009.eqiad.wmnet with reason: host reimage * 10:49 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2010.codfw.wmnet with OS trixie * 10:46 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2011.codfw.wmnet * 10:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1011.eqiad.wmnet * 10:44 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2011.codfw.wmnet * 10:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 10:44 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 10:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1009.eqiad.wmnet with reason: host reimage * 10:44 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow6001.drmrs.wmnet * 10:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1017: Security updates * 10:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:39 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1017: Security updates * 10:39 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow6001.drmrs.wmnet * 10:38 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1011.eqiad.wmnet * 10:35 cgoubert@deploy2003: Finished deploy [restbase/deploy@06301bd]: Deploying {{Gerrit|1306088}} {{Gerrit|1308347}} - [[phab:T429944|T429944]] [[phab:T428279|T428279]] (duration: 28m 34s) * 10:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1012.eqiad.wmnet * 10:35 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 10:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow5003.eqsin.wmnet * 10:34 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2011.codfw.wmnet with OS trixie * 10:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1009.eqiad.wmnet with OS trixie * 10:28 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1012.eqiad.wmnet * 10:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1013.eqiad.wmnet * 10:27 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:27 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow5003.eqsin.wmnet * 10:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:26 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow4003.ulsfo.wmnet * 10:25 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 10:25 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 10:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow4003.ulsfo.wmnet * 10:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1013.eqiad.wmnet * 10:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1014.eqiad.wmnet * 10:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:15 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2011.codfw.wmnet with reason: host reimage * 10:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1017: Security updates * 10:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:14 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:14 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1017: Security updates * 10:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow3004.esams.wmnet * 10:11 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2011.codfw.wmnet with reason: host reimage * 10:10 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:10 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1014.eqiad.wmnet * 10:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki2003.codfw.wmnet * 10:09 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 10:09 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 10:09 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow3004.esams.wmnet * 10:08 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2004.codfw.wmnet * 10:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1010.eqiad.wmnet with OS trixie * 10:07 cgoubert@deploy2003: Started deploy [restbase/deploy@06301bd]: Deploying {{Gerrit|1306088}} {{Gerrit|1308347}} - [[phab:T429944|T429944]] [[phab:T428279|T428279]] * 10:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host rpki2003.codfw.wmnet * 10:04 topranks: push out config change to BGP_outfilter on core routers [[phab:T431849|T431849]] * 10:02 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow2004.codfw.wmnet * 09:59 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 09:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2003.codfw.wmnet * 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2011 * 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2011 * 09:53 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 09:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:52 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2011 * 09:52 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2011.codfw.wmnet 36.32.192.10.in-addr.arpa 6.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:52 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2011.codfw.wmnet 36.32.192.10.in-addr.arpa 6.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:51 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:51 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2011 - mvernon@cumin2003" * 09:51 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2011 - mvernon@cumin2003" * 09:51 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow2003.codfw.wmnet * 09:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1003.eqiad.wmnet * 09:49 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 09:49 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 09:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1010.eqiad.wmnet with reason: host reimage * 09:47 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 09:47 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 09:47 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 09:47 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2011 * 09:46 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2011.codfw.wmnet with OS trixie * 09:44 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow1003.eqiad.wmnet * 09:44 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2012.codfw.wmnet * 09:44 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1002.eqiad.wmnet * 09:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1010.eqiad.wmnet with reason: host reimage * 09:43 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2012.codfw.wmnet * 09:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Security updates * 09:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:43 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:43 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Security updates * 09:42 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:40 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow1002.eqiad.wmnet * 09:40 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 09:37 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki1001.eqiad.wmnet * 09:36 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 09:36 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 09:33 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host rpki1001.eqiad.wmnet * 09:32 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:32 cgoubert@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-codfw * 09:31 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=kartotherian,name=eqiad * 09:31 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola-vector-tiles,name=eqiad * 09:31 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 09:31 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2012.codfw.wmnet with OS trixie * 09:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1010.eqiad.wmnet with OS trixie * 09:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1011.eqiad.wmnet with OS trixie * 09:21 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Security updates * 09:21 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:21 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:21 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Security updates * 09:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2012.codfw.wmnet with reason: host reimage * 09:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1011.eqiad.wmnet with reason: host reimage * 09:08 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2012.codfw.wmnet with reason: host reimage * 09:05 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1011.eqiad.wmnet with reason: host reimage * 08:55 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:52 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1011.eqiad.wmnet with OS trixie * 08:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1022: Security updates * 08:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2012 * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2012 * 08:50 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1022: Security updates * 08:50 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2012 * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2012.codfw.wmnet 44.48.192.10.in-addr.arpa 4.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:50 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2012.codfw.wmnet 44.48.192.10.in-addr.arpa 4.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2012 - mvernon@cumin2003" * 08:50 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2012 - mvernon@cumin2003" * 08:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1012.eqiad.wmnet with OS trixie * 08:44 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 08:44 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2012 * 08:43 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2012.codfw.wmnet with OS trixie * 08:42 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2013.codfw.wmnet * 08:41 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2013.codfw.wmnet * 08:35 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 08:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host krb1002.eqiad.wmnet * 08:30 elukey@dns1004: END - running authdns-update * 08:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1012.eqiad.wmnet with reason: host reimage * 08:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Security updates * 08:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:28 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:28 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Security updates * 08:27 elukey@dns1004: START - running authdns-update * 08:26 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 08:26 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host krb1002.eqiad.wmnet * 08:22 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1012.eqiad.wmnet with reason: host reimage * 08:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host krb2002.codfw.wmnet * 08:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast6003.wikimedia.org * 08:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2013.codfw.wmnet with OS trixie * 08:13 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast6003.wikimedia.org * 08:12 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast3007.wikimedia.org * 08:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host krb2002.codfw.wmnet * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Security updates * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:09 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:09 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Security updates * 08:07 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1012.eqiad.wmnet with OS trixie * 08:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast3007.wikimedia.org * 08:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast5005.wikimedia.org * 07:58 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast5005.wikimedia.org * 07:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1013.eqiad.wmnet with OS trixie * 07:53 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2013.codfw.wmnet with reason: host reimage * 07:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1021: Security updates * 07:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:53 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:53 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1021: Security updates * 07:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast1004.wikimedia.org * 07:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2013.codfw.wmnet with reason: host reimage * 07:46 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast1004.wikimedia.org * 07:40 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1013.eqiad.wmnet with reason: host reimage * 07:36 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1013.eqiad.wmnet with reason: host reimage * 07:31 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2013 * 07:31 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2013 * 07:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1021: Security updates * 07:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:30 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:30 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1021: Security updates * 07:24 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2013 * 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2013.codfw.wmnet 87.0.192.10.in-addr.arpa 7.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:24 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2013.codfw.wmnet 87.0.192.10.in-addr.arpa 7.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2013 - mvernon@cumin2003" * 07:24 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2013 - mvernon@cumin2003" * 07:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1013.eqiad.wmnet with OS trixie * 07:19 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 07:19 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2013 * 07:19 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2013.codfw.wmnet with OS trixie * 07:13 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] (duration: 07m 48s) * 07:09 kharlan@deploy2003: kharlan: Continuing with deployment * 07:08 kharlan@deploy2003: kharlan: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:06 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 01:15 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 01:14 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply == 2026-07-14 == * 22:51 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_magru * 22:51 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7016.magru.wmnet * 22:46 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_magru * 22:46 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7008.magru.wmnet * 22:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7015.magru.wmnet * 22:04 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7007.magru.wmnet * 21:29 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7014.magru.wmnet * 21:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7006.magru.wmnet * 21:13 dzahn@cumin2002: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 0:15:00 on gerrit.wikimedia.org with reason: reboot * 21:11 mutante: gerrit2003 (gerrit.wikimedia.org) - reboot for maintenance * 21:11 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on gerrit2003.wikimedia.org with reason: reboot * 20:56 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:56 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:56 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:55 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 20:48 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7013.magru.wmnet * 20:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7005.magru.wmnet * 20:41 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host phab1005.eqiad.wmnet with OS trixie * 20:28 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] (duration: 06m 47s) * 20:24 sbassett@deploy2003: sbassett: Continuing with deployment * 20:23 sbassett@deploy2003: sbassett: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:23 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on phab1005.eqiad.wmnet with reason: host reimage * 20:21 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] * 20:20 aokoth@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on phab1005.eqiad.wmnet with reason: host reimage * 20:12 jhuneidi@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] (duration: 07m 42s) * 20:07 jhuneidi@deploy2003: jhuneidi, priyankar22: Continuing with deployment * 20:06 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7012.magru.wmnet * 20:06 jhuneidi@deploy2003: jhuneidi, priyankar22: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:04 jhuneidi@deploy2003: Started scap sync-world: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] * 20:02 aokoth@cumin1003: START - Cookbook sre.hosts.reimage for host phab1005.eqiad.wmnet with OS trixie * 20:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7004.magru.wmnet * 20:00 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet * 19:57 aokoth@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet * 19:24 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7011.magru.wmnet * 19:19 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7003.magru.wmnet * 19:11 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] (duration: 08m 33s) * 19:07 jforrester@deploy2003: jforrester: Continuing with deployment * 19:04 jforrester@deploy2003: jforrester: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:02 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] * 18:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7010.magru.wmnet * 18:38 mutante: rotating phabricator-gerrit bot token (its-phabricator) * 18:18 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 17:44 swfrench@deploy2003: Finished scap sync-world: Deployment to pick up new production image (duration: 31m 44s) * 17:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7002.magru.wmnet * 17:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7009.magru.wmnet * 17:32 swfrench@deploy2003: swfrench: Continuing with deployment * 17:29 swfrench@deploy2003: swfrench: Deployment to pick up new production image synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:17 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2035: repooling after rack b5 maintenance * 17:16 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool es2035: repooling after rack b5 maintenance * 17:16 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2188: repooling after rack b5 maintenance * 17:12 swfrench@deploy2003: Started scap sync-world: Deployment to pick up new production image * 17:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7001.magru.wmnet * 16:57 swfrench-wmf: reprepro include php8.3_8.3.32-1+wmf12u2 into component/php83 for bookworm-wikimedia * 16:50 sukhe: pool cp2046 * 16:47 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4039.ulsfo.wmnet * 16:44 sukhe: sudo cumin -b31 "A:cp" "run-puppet-agent" * 16:33 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on contint1003.wikimedia.org with reason: reboot * 16:32 mutante: contint1003 - main CI server - rebooting * 16:31 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2188: repooling after rack b5 maintenance * 16:31 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2178: repooling after rack b5 maintenance * 16:29 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 16:28 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 16:28 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 16:28 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 16:18 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2014.codfw.wmnet * 16:18 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 16:17 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2014.codfw.wmnet * 16:10 mvernon@cumin1003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-thanos-proxies (exit_code=0) rolling restart_daemons on A:thanos-fe * 16:09 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 16:07 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp4039.ulsfo.wmnet * 16:07 mvernon@cumin1003: START - Cookbook sre.swift.roll-restart-reboot-swift-thanos-proxies rolling restart_daemons on A:thanos-fe * 16:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2014.codfw.wmnet with OS trixie * 15:56 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1014.eqiad.wmnet with OS trixie * 15:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2014.codfw.wmnet with reason: host reimage * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2014 * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2014 * 15:28 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2014 * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2014.codfw.wmnet 194.16.192.10.in-addr.arpa 4.9.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:28 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2014.codfw.wmnet 194.16.192.10.in-addr.arpa 4.9.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2014 - mvernon@cumin2003" * 15:28 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2014 - mvernon@cumin2003" * 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Apply title-related policies when selecting the name of the entity - kamila@cumin1003" * 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Apply title-related policies when selecting the name of the entity - kamila@cumin1003 * 15:22 kamila@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Apply title-related policies when selecting the name of the entity - kamila@cumin1003 * 15:22 kamila@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Apply title-related policies when selecting the name of the entity - kamila@cumin1003" * 15:20 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 15:20 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2014 * 15:20 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2014.codfw.wmnet with OS trixie * 15:19 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1014.eqiad.wmnet with OS trixie * 15:01 dancy@deploy2003: Installation of scap version "4.274.1" completed for 3 hosts * 15:00 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2177: repooling after rack b5 maintenance * 15:00 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2159: repooling after rack b5 maintenance * 14:59 dancy@deploy2003: Installing scap version "4.274.1" for 3 host(s) * 14:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2015.codfw.wmnet with OS trixie * 14:54 seanleong-wmde: Finished populateSitesTable for isvwiki ([[phab:T429939|T429939]]) * 14:53 javiermonton@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] (duration: 07m 35s) * 14:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1015.eqiad.wmnet with OS trixie * 14:49 javiermonton@deploy2003: javiermonton: Continuing with deployment * 14:48 javiermonton@deploy2003: javiermonton: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:46 javiermonton@deploy2003: Started scap sync-world: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] * 14:42 otto@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 14:41 otto@deploy2003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 14:41 otto@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 14:40 otto@deploy2003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 14:40 otto@deploy2003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 14:39 otto@deploy2003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 14:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2015.codfw.wmnet with reason: host reimage * 14:34 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1015.eqiad.wmnet with reason: host reimage * 14:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2015.codfw.wmnet with reason: host reimage * 14:30 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1015.eqiad.wmnet with reason: host reimage * 14:30 seanleong-wmde@deploy2003: mwscript-k8s job started: foreachwikiindblist wikidataclient extensions/Wikibase/lib/maintenance/populateSitesTable.php --force-protocol https # [[phab:T429939|T429939]] * 14:24 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling reboot on A:durum and not (A:durum-eqiad or A:durum-codfw or A:durum-esams) and A:durum * 14:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2015.codfw.wmnet with OS trixie * 14:15 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2016.codfw.wmnet with OS trixie * 14:14 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2159: repooling after rack b5 maintenance * 14:14 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1015.eqiad.wmnet with OS trixie * 14:12 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1016.eqiad.wmnet with OS trixie * 14:12 cmooney@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=pki,name=codfw * 14:12 sbisson@deploy2003: helmfile [codfw] DONE helmfile.d/services/cxserver: sync * 14:11 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2002.codfw.wmnet * 14:11 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2002.codfw.wmnet * 14:11 sbisson@deploy2003: helmfile [codfw] START helmfile.d/services/cxserver: sync * 14:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1005.wikimedia.org * 14:07 sbisson@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cxserver: sync * 14:07 sbisson@deploy2003: helmfile [eqiad] START helmfile.d/services/cxserver: sync * 14:05 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1005.wikimedia.org * 14:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader2005.wikimedia.org * 14:02 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2003.codfw.wmnet * 14:02 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2003.codfw.wmnet * 14:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=tegola-vector-tiles,name=codfw * 14:00 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=kartotherian,name=codfw * 14:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader2005.wikimedia.org * 13:58 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2016.codfw.wmnet with reason: host reimage * 13:57 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-ncredir (exit_code=0) rolling reboot on A:ncredir and A:ncredir * 13:57 sbisson@deploy2003: helmfile [staging] DONE helmfile.d/services/cxserver: sync * 13:56 sbisson@deploy2003: helmfile [staging] START helmfile.d/services/cxserver: sync * 13:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1016.eqiad.wmnet with reason: host reimage * 13:52 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:52 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:51 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2016.codfw.wmnet with reason: host reimage * 13:50 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1016.eqiad.wmnet with reason: host reimage * 13:49 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy (exit_code=0) rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 13:49 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling reboot on A:wikidough * 13:46 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-tcp-proxy (exit_code=0) rolling reboot on A:tcpproxy and A:tcpproxy * 13:43 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and not (A:durum-eqiad or A:durum-codfw or A:durum-esams) and A:durum * 13:42 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=97) rolling reboot on A:durum and A:durum * 13:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2011.codfw.wmnet * 13:36 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-reboot (exit_code=0) rolling reboot on A:dnsbox and A:ulsfo and (A:dnsbox) * 13:36 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns4004.wikimedia.org * 13:34 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1016.eqiad.wmnet with OS trixie * 13:34 topranks: reboot lsw1-b5-codfw to upgrade JunOS [[phab:T430918|T430918]] * 13:34 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2016.codfw.wmnet with OS trixie * 13:32 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2002.codfw.wmnet * 13:31 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2011.codfw.wmnet * 13:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2012.codfw.wmnet * 13:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2017.codfw.wmnet with OS trixie * 13:25 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1017.eqiad.wmnet with OS trixie * 13:24 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2012.codfw.wmnet * 13:22 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2002.codfw.wmnet * 13:22 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:22 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:22 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns4004.wikimedia.org * 13:21 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2005.codfw.wmnet * 13:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2013.codfw.wmnet * 13:18 elukey@dns1004: END - running authdns-update * 13:17 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2005.codfw.wmnet * 13:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2004.codfw.wmnet * 13:16 elukey@dns1004: START - running authdns-update * 13:16 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1046: es1046 after reimage * 13:14 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1029.eqiad.wmnet,service=s8 * 13:14 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1029.eqiad.wmnet,service=s5 * 13:13 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1029.eqiad.wmnet,service=s5 * 13:13 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1029.eqiad.wmnet,service=s8 * 13:13 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2004.codfw.wmnet * 13:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2013.codfw.wmnet * 13:11 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2014.codfw.wmnet * 13:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2003.codfw.wmnet * 13:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2017.codfw.wmnet with reason: host reimage * 13:07 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:07 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns4003.wikimedia.org * 13:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2003.codfw.wmnet * 13:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1067.eqiad.wmnet * 13:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1067.eqiad.wmnet * 13:06 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1067.eqiad.wmnet * 13:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm-test1001.wikimedia.org * 13:05 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2017.codfw.wmnet with reason: host reimage * 13:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1017.eqiad.wmnet with reason: host reimage * 13:04 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2014.codfw.wmnet * 13:03 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:02 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2188: codfw rack B5 depool for maintenance * 13:02 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_magru * 13:01 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2188: codfw rack B5 depool for maintenance * 13:01 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_magru * 13:01 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2178: codfw rack B5 depool for maintenance * 13:01 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm-test1001.wikimedia.org * 13:01 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2178: codfw rack B5 depool for maintenance * 13:01 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2177: codfw rack B5 depool for maintenance * 13:00 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2177: codfw rack B5 depool for maintenance * 12:59 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1068.eqiad.wmnet * 12:59 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1068.eqiad.wmnet * 12:58 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola-vector-tiles,name=codfw * 12:58 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2159: codfw rack B5 depool for maintenance * 12:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1017.eqiad.wmnet with reason: host reimage * 12:58 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola,name=codfw * 12:57 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=kartotherian,name=codfw * 12:57 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2159: codfw rack B5 depool for maintenance * 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 30 hosts with reason: lsw1-b5-codfw JunOS upgrade * 12:55 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lsw1-b5-codfw,lsw1-b5-codfw IPv6,lsw1-b5-codfw.mgmt,ssw1-a[1,8]-codfw.mgmt with reason: switch upgade lsw1-b5-codfw * 12:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet * 12:49 cmooney@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=pki,name=codfw * 12:49 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1067.eqiad.wmnet with OS trixie * 12:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet * 12:48 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2017.codfw.wmnet with OS trixie * 12:47 topranks: depool codfw pki in dns discovery ahead of lsw1-b5-codfw maintenance [[phab:T430918|T430918]] * 12:47 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns4003.wikimedia.org * 12:47 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and A:ulsfo and (A:dnsbox) * 12:47 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2018.codfw.wmnet * 12:45 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2018.codfw.wmnet * 12:45 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and A:durum * 12:45 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-tcp-proxy rolling reboot on A:tcpproxy and A:tcpproxy * 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host pki-root1002.eqiad.wmnet * 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1009.eqiad.wmnet * 12:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1009.eqiad.wmnet * 12:44 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 12:43 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-ncredir rolling reboot on A:ncredir and A:ncredir * 12:43 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling reboot on A:wikidough * 12:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2018.codfw.wmnet with OS trixie * 12:42 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1017.eqiad.wmnet with OS trixie * 12:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader2006.wikimedia.org * 12:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1018.eqiad.wmnet with OS trixie * 12:39 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1009.eqiad.wmnet * 12:38 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host pki-root1002.eqiad.wmnet * 12:38 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1009.eqiad.wmnet * 12:38 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1008.eqiad.wmnet * 12:38 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1008.eqiad.wmnet * 12:35 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader2006.wikimedia.org * 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1006.wikimedia.org * 12:33 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1008.eqiad.wmnet * 12:30 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1046: es1046 after reimage * 12:29 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host es1046.eqiad.wmnet with OS trixie * 12:29 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1006.wikimedia.org * 12:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test2005.wikimedia.org * 12:28 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1008.eqiad.wmnet * 12:27 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1007.eqiad.wmnet * 12:27 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1007.eqiad.wmnet * 12:27 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1067.eqiad.wmnet with reason: host reimage * 12:25 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1068.eqiad.wmnet with reason: vacuum overlarge container dbs * 12:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2018.codfw.wmnet with reason: host reimage * 12:24 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test2005.wikimedia.org * 12:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test1005.wikimedia.org * 12:23 Amir1: mwscript-k8s --follow --dblist=ores -- extensions/ORES/maintenance/PurgeScoreCache.php --model goodfaith --old ([[phab:T431159|T431159]]) * 12:22 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1007.eqiad.wmnet * 12:22 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test1005.wikimedia.org * 12:22 atsukoito: restarting pybal on lvs2013 `low-traffic` for https://gerrit.wikimedia.org/r/1310535 * 12:22 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1007.eqiad.wmnet * 12:21 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1006.eqiad.wmnet * 12:21 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1006.eqiad.wmnet * 12:20 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1018.eqiad.wmnet with reason: host reimage * 12:19 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2018.codfw.wmnet with reason: host reimage * 12:18 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1067.eqiad.wmnet with reason: host reimage * 12:16 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1006.eqiad.wmnet * 12:15 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1006.eqiad.wmnet * 12:15 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1005.eqiad.wmnet * 12:15 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1005.eqiad.wmnet * 12:15 atsukoito: restarting pybal on lvs2014 for https://gerrit.wikimedia.org/r/1310535 * 12:12 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1018.eqiad.wmnet with reason: host reimage * 12:11 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1005.eqiad.wmnet * 12:11 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1005.eqiad.wmnet * 12:10 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1004.eqiad.wmnet * 12:10 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1004.eqiad.wmnet * 12:09 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on es1046.eqiad.wmnet with reason: host reimage * 12:08 atsukoito: restarting pybal on lvs1019 `low-traffic` for https://gerrit.wikimedia.org/r/1310535 * 12:06 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1004.eqiad.wmnet * 12:06 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1004.eqiad.wmnet * 12:06 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1003.eqiad.wmnet * 12:06 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1003.eqiad.wmnet * 12:05 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on es1046.eqiad.wmnet with reason: host reimage * 12:05 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "set ml-serve1001 back to active state - cmooney@cumin1003" * 12:04 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "set ml-serve1001 back to active state - cmooney@cumin1003" * 12:04 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:02 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1003.eqiad.wmnet * 12:01 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1003.eqiad.wmnet * 12:01 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1002.eqiad.wmnet * 12:01 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1002.eqiad.wmnet * 12:01 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:59 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2018.codfw.wmnet with OS trixie * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1067 * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1067 * 11:59 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1067 * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1067.eqiad.wmnet 17.48.64.10.in-addr.arpa 7.1.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:59 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1067.eqiad.wmnet 17.48.64.10.in-addr.arpa 7.1.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1067 - blake@cumin1003" * 11:58 atsukoito: restarting pybal on lvs1018 `high-traffic2` for https://gerrit.wikimedia.org/r/1310535 * 11:57 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1002.eqiad.wmnet * 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2019.codfw.wmnet with OS trixie * 11:56 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1002.eqiad.wmnet * 11:56 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 11:56 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1018.eqiad.wmnet with OS trixie * 11:54 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 11:54 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:54 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1019.eqiad.wmnet with OS trixie * 11:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:49 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:49 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:49 aikochou@deploy2003: helmfile [codfw] DONE helmfile.d/services/changeprop: sync * 11:48 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host es1046.eqiad.wmnet with OS trixie * 11:48 aikochou@deploy2003: helmfile [codfw] START helmfile.d/services/changeprop: sync * 11:48 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310535 * 11:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1046: Reimage to Trixie * 11:44 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1046: Reimage to Trixie * 11:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5:00:00 on es1046.eqiad.wmnet with reason: Reimage to Trixie * 11:42 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:42 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:42 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 11:42 aikochou@deploy2003: helmfile [eqiad] DONE helmfile.d/services/changeprop: sync * 11:41 aikochou@deploy2003: helmfile [eqiad] START helmfile.d/services/changeprop: sync * 11:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2019.codfw.wmnet with reason: host reimage * 11:36 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] (duration: 09m 41s) * 11:36 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:36 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:35 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:35 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:32 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1019.eqiad.wmnet with reason: host reimage * 11:32 jforrester@deploy2003: jforrester, gengh: Continuing with deployment * 11:29 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2019.codfw.wmnet with reason: host reimage * 11:28 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1019.eqiad.wmnet with reason: host reimage * 11:28 jforrester@deploy2003: jforrester, gengh: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:26 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] * 11:20 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2003.codfw.wmnet * 11:20 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:19 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2003.codfw.wmnet * 11:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:12 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1019.eqiad.wmnet with OS trixie * 11:10 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2019.codfw.wmnet with OS trixie * 11:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1020.eqiad.wmnet with OS trixie * 11:10 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1067 - blake@cumin1003" * 11:09 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] (duration: 12m 12s) * 11:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2020.codfw.wmnet with OS trixie * 11:03 kharlan@deploy2003: kharlan: Continuing with deployment * 11:01 blake@cumin1003: START - Cookbook sre.dns.netbox * 11:01 kharlan@deploy2003: kharlan: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:57 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] * 10:55 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] (duration: 31m 40s) * 10:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1020.eqiad.wmnet with reason: host reimage * 10:52 marostegui@dns1004: START - running authdns-update * 10:49 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2020.codfw.wmnet with reason: host reimage * 10:49 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:48 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1020.eqiad.wmnet with reason: host reimage * 10:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2020.codfw.wmnet with reason: host reimage * 10:43 kharlan@deploy2003: kharlan: Continuing with deployment * 10:42 kharlan@deploy2003: kharlan: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:32 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1020.eqiad.wmnet with OS trixie * 10:29 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2159: Repooling after switchover * 10:29 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1067 * 10:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1021.eqiad.wmnet with OS trixie * 10:27 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1067.eqiad.wmnet with OS trixie * 10:27 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:27 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1067.eqiad.wmnet * 10:27 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:26 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1067.eqiad.wmnet * 10:26 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1067.eqiad.wmnet * 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2020.codfw.wmnet with OS trixie * 10:26 blake@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1055.eqiad.wmnet * 10:26 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1055.eqiad.wmnet * 10:26 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1055.eqiad.wmnet * 10:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2021.codfw.wmnet with OS trixie * 10:24 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] * 10:11 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1055.eqiad.wmnet with OS trixie * 10:09 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1021.eqiad.wmnet with reason: host reimage * 10:05 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2021.codfw.wmnet with reason: host reimage * 10:03 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310129 revert * 10:02 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1021.eqiad.wmnet with reason: host reimage * 10:01 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2021.codfw.wmnet with reason: host reimage * 09:58 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310129 * 09:50 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1055.eqiad.wmnet with reason: host reimage * 09:45 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1055.eqiad.wmnet with reason: host reimage * 09:45 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1021.eqiad.wmnet with OS trixie * 09:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1022.eqiad.wmnet with OS trixie * 09:44 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2159: Repooling after switchover * 09:44 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2021.codfw.wmnet with OS trixie * 09:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2022.codfw.wmnet with OS trixie * 09:31 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2159.codfw.wmnet * 09:28 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1055 * 09:28 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1055 * 09:27 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on ms-fe1022.eqiad.wmnet with reason: host reimage * 09:27 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1022.eqiad.wmnet with reason: host reimage * 09:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2022.codfw.wmnet with reason: host reimage * 09:21 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2022.codfw.wmnet with reason: host reimage * 09:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2159: Rebooting db2159.codfw.wmnet * 09:20 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2159: Rebooting db2159.codfw.wmnet * 09:18 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 09:18 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 09:18 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 09:17 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 09:16 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2159.codfw.wmnet * 09:13 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b] (thin): Regular analytics weekly train THIN [analytics/refinery@ad6e05b8] (duration: 02m 07s) * 09:11 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b] (thin): Regular analytics weekly train THIN [analytics/refinery@ad6e05b8] * 09:10 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1022.eqiad.wmnet with OS trixie * 09:07 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1023.eqiad.wmnet with OS trixie * 09:06 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b]: Regular analytics weekly train [analytics/refinery@ad6e05b8] (duration: 04m 49s) * 09:04 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2022.codfw.wmnet with OS trixie * 09:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2023.codfw.wmnet with OS trixie * 09:01 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b]: Regular analytics weekly train [analytics/refinery@ad6e05b8] * 09:01 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@ad6e05b8] (duration: 02m 01s) * 09:00 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1055 * 09:00 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1055.eqiad.wmnet 50.32.64.10.in-addr.arpa 0.5.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:00 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1055.eqiad.wmnet 50.32.64.10.in-addr.arpa 0.5.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:00 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:00 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1055 - blake@cumin1003" * 09:00 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1055 - blake@cumin1003" * 08:59 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@ad6e05b8] * 08:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2159 [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94811 and previous config saved to /var/cache/conftool/dbconfig/20260714-085624-cwilliams.json * 08:55 blake@cumin1003: START - Cookbook sre.dns.netbox * 08:55 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1055 * 08:54 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1055.eqiad.wmnet with OS trixie * 08:54 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1055.eqiad.wmnet * 08:53 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1055.eqiad.wmnet * 08:53 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1055.eqiad.wmnet * 08:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2220 to s7 primary [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94810 and previous config saved to /var/cache/conftool/dbconfig/20260714-085239-cwilliams.json * 08:51 cezmunsta: Starting s7 codfw failover from db2159 to db2220 - [[phab:T430920|T430920]] * 08:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1023.eqiad.wmnet with reason: host reimage * 08:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2220 with weight 0 [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94809 and previous config saved to /var/cache/conftool/dbconfig/20260714-084553-cwilliams.json * 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s7 [[phab:T430920|T430920]] * 08:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2023.codfw.wmnet with reason: host reimage * 08:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1023.eqiad.wmnet with reason: host reimage * 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2023.codfw.wmnet with reason: host reimage * 08:34 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:34 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:29 marostegui@dns1004: END - running authdns-update * 08:29 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox-dev2003.codfw.wmnet * 08:27 marostegui@dns1004: START - running authdns-update * 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker2*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2009.codfw.wmnet * 08:26 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2009.codfw.wmnet * 08:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1023.eqiad.wmnet with OS trixie * 08:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox-dev2003.codfw.wmnet * 08:24 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:24 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:24 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2023.codfw.wmnet with OS trixie * 08:24 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1029.eqiad.wmnet with reason: reboot * 08:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1027.eqiad.wmnet with reason: reboot * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:21 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2009.codfw.wmnet * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:20 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2009.codfw.wmnet * 08:20 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2008.codfw.wmnet * 08:20 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2008.codfw.wmnet * 08:15 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2008.codfw.wmnet * 08:14 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2008.codfw.wmnet * 08:14 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2007.codfw.wmnet * 08:14 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2007.codfw.wmnet * 08:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1024.eqiad.wmnet with OS trixie * 08:12 elukey@cumin1003: END (PASS) - Cookbook sre.pki.restart-reboot (exit_code=0) rolling reboot on P<nowiki>{</nowiki>pki*<nowiki>}</nowiki> and (A:pki) * 08:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 08:10 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 08:09 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2007.codfw.wmnet * 08:08 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2007.codfw.wmnet * 08:08 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2006.codfw.wmnet * 08:08 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2006.codfw.wmnet * 08:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2024.codfw.wmnet with OS trixie * 08:03 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2006.codfw.wmnet * 08:02 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2006.codfw.wmnet * 08:02 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2005.codfw.wmnet * 08:02 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2005.codfw.wmnet * 07:58 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2005.codfw.wmnet * 07:58 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2005.codfw.wmnet * 07:57 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2004.codfw.wmnet * 07:57 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2004.codfw.wmnet * 07:54 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki.discovery.wmnet. on all recursors * 07:54 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache pki.discovery.wmnet. on all recursors * 07:53 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2004.codfw.wmnet * 07:53 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2004.codfw.wmnet * 07:53 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2003.codfw.wmnet * 07:53 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2003.codfw.wmnet * 07:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1024.eqiad.wmnet with reason: host reimage * 07:49 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki.discovery.wmnet. on all recursors * 07:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2003.codfw.wmnet * 07:49 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache pki.discovery.wmnet. on all recursors * 07:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2024.codfw.wmnet with reason: host reimage * 07:48 elukey@cumin1003: START - Cookbook sre.pki.restart-reboot rolling reboot on P<nowiki>{</nowiki>pki*<nowiki>}</nowiki> and (A:pki) * 07:46 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1024.eqiad.wmnet with reason: host reimage * 07:45 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2024.codfw.wmnet with reason: host reimage * 07:45 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2003.codfw.wmnet * 07:44 elukey@cumin1003: END (PASS) - Cookbook sre.misc-clusters.restart-reboot-config-master (exit_code=0) rolling reboot on P<nowiki>{</nowiki>config-master*<nowiki>}</nowiki> and (A:config-master or A:config-master-eqiad or A:config-master-codfw) * 07:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2002.codfw.wmnet * 07:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2002.codfw.wmnet * 07:39 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2002.codfw.wmnet * 07:39 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) config-master.discovery.wmnet. on all recursors * 07:39 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache config-master.discovery.wmnet. on all recursors * 07:39 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2002.codfw.wmnet * 07:39 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker2*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl200*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl2003.codfw.wmnet * 07:36 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl2003.codfw.wmnet * 07:35 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) config-master.discovery.wmnet. on all recursors * 07:35 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache config-master.discovery.wmnet. on all recursors * 07:34 elukey@cumin1003: START - Cookbook sre.misc-clusters.restart-reboot-config-master rolling reboot on P<nowiki>{</nowiki>config-master*<nowiki>}</nowiki> and (A:config-master or A:config-master-eqiad or A:config-master-codfw) * 07:31 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl2003.codfw.wmnet * 07:31 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl2003.codfw.wmnet * 07:31 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl2002.codfw.wmnet * 07:31 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl2002.codfw.wmnet * 07:29 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1024.eqiad.wmnet with OS trixie * 07:28 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2024.codfw.wmnet with OS trixie * 07:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl2002.codfw.wmnet * 07:26 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl2002.codfw.wmnet * 07:26 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl200*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 07:26 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 07:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 06:50 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lists1004.wikimedia.org * 06:44 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host lists1004.wikimedia.org * 06:25 marostegui@dns1004: END - running authdns-update * 06:23 marostegui@dns1004: START - running authdns-update * 06:22 marostegui@dns1004: END - running authdns-update * 06:20 marostegui@dns1004: START - running authdns-update * 06:17 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1026.eqiad.wmnet with reason: reboot * 06:04 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: sync * 06:04 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: sync * 06:03 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync * 06:03 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync * 06:02 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync * 06:01 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync * 06:01 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync * 06:00 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync * 05:59 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:59 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:40 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:39 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:26 marostegui@dns1004: END - running authdns-update * 05:24 marostegui@dns1004: START - running authdns-update * 05:24 marostegui@dns1004: START - running authdns-update * 05:13 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1004.wikimedia.org * 05:07 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1004.wikimedia.org * 04:01 mwpresync@deploy2003: Pruned MediaWiki: 1.47.0-wmf.8 (duration: 01m 07s) * 03:39 mwpresync@deploy2003: Finished scap sync-world: testwikis to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] (duration: 36m 01s) * 03:03 mwpresync@deploy2003: Started scap sync-world: testwikis to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 29s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-13 == * 23:33 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1064.eqiad.wmnet * 23:33 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1064.eqiad.wmnet * 23:08 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1064.eqiad.wmnet with reason: vacuum overlarge container dbs * 23:06 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1069.eqiad.wmnet * 23:06 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1069.eqiad.wmnet * 22:34 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1069.eqiad.wmnet with reason: vacuum overlarge container dbs * 21:18 maryum: Deployed security fix for [[phab:T321092|T321092]] * 20:28 swfrench-wmf: reprepro include etcd-mirror_0.0.12-1+deb13u1 into main for trixie-wikimedia - [[phab:T424266|T424266]] * 20:26 swfrench-wmf: reprepro include etcd-mirror_0.0.12-1+deb12u1 into main for bookworm-wikimedia - [[phab:T428495|T428495]] * 20:23 dancy@deploy2003: Finished scap sync-world: Testing [[phab:T431635|T431635]] (duration: 03m 36s) * 20:19 dancy@deploy2003: Started scap sync-world: Testing [[phab:T431635|T431635]] * 20:18 dancy@deploy2003: Installation of scap version "4.274.0" completed for 3 hosts * 20:16 dancy@deploy2003: Installing scap version "4.274.0" for 3 host(s) * 20:12 kemayo@deploy2003: Finished scap sync-world: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] (duration: 08m 25s) * 20:07 kemayo@deploy2003: soda, esanders, kemayo: Continuing with deployment * 20:05 kemayo@deploy2003: soda, esanders, kemayo: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there * 20:04 kemayo@deploy2003: Started scap sync-world: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] * 18:22 cwhite: lvextend vg0/srv +500g on centrallog hosts * 18:19 cdobbins@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS trixie * 17:46 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1071.eqiad.wmnet * 17:46 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1071.eqiad.wmnet * 17:13 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1071.eqiad.wmnet with reason: vacuum overlarge container dbs * 17:07 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1065.eqiad.wmnet * 17:07 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1065.eqiad.wmnet * 17:06 dzahn@dns1006: END - running authdns-update * 17:04 dzahn@dns1006: START - running authdns-update * 17:01 dzahn@dns1006: END - running authdns-update * 16:59 dzahn@dns1006: START - running authdns-update * 16:51 dancy@deploy2003: Finished scap sync-world: testing [[phab:T428971|T428971]] (duration: 03m 37s) * 16:47 dancy@deploy2003: Started scap sync-world: testing [[phab:T428971|T428971]] * 16:45 atsukoito: restarting pybal on lvs1019 to flush IP address for `cirrussearch1122.eqiad.wmnet` after moving the vlan [[phab:T431311|T431311]] * 16:42 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:42 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:42 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:42 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:42 Amir1: mwscript-k8s --follow --dblist=ores -- extensions/ORES/maintenance/PurgeScoreCache.php --model damaging --old ([[phab:T431159|T431159]]) * 16:34 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Pool test * 16:34 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 16:34 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 16:34 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Pool test * 16:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Depool test * 16:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 16:33 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 16:33 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Depool test * 16:31 dancy@deploy2003: Installation of scap version "4.273.0" completed for 159 hosts * 16:29 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1065.eqiad.wmnet with reason: vacuum overlarge container dbs * 16:27 dancy@deploy2003: Installing scap version "4.273.0" for 159 host(s) * 16:27 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics-external: sync * 16:27 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics-external: sync * 16:26 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics-external: sync * 16:26 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics-external: sync * 16:22 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync * 16:21 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync * 16:21 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: sync * 16:21 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: sync * 16:19 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync * 16:19 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync * 16:18 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync * 16:17 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync * 15:59 atsukoito: restarting pybal on lvs1018 for https://gerrit.wikimedia.org/r/1310117 * 15:55 aikochou@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 15:50 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310117 * 15:46 aikochou@deploy2003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 15:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host kafka-logging1006.eqiad.wmnet * 15:43 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host kafka-logging1006.eqiad.wmnet * 15:41 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host ganeti-test[2001-2003].codfw.wmnet * 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host ganeti-test[2001-2003].codfw.wmnet * 15:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host netbox1003.eqiad.wmnet * 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host netbox1003.eqiad.wmnet * 15:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host netbox2003.codfw.wmnet * 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host netbox2003.codfw.wmnet * 15:36 sukhe: restart pybal on lvs1020 * 15:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet * 15:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet * 15:08 btullis@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:06 btullis@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 15:01 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:01 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:35 cdobbins@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 14:34 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:33 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:33 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:32 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:29 cdobbins@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 14:28 marostegui@dns1004: END - running authdns-update * 14:28 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:27 marostegui@dns1004: START - running authdns-update * 14:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1023.eqiad.wmnet with reason: reboot * 14:18 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2009.codfw.wmnet with OS trixie * 14:14 swfrench-wmf: start rolling run-puppet-agent on A:cp for ATS config change - [[phab:T428909|T428909]] [[phab:T431838|T431838]] * 14:05 swfrench-wmf: disable-puppet on A:cp for ATS config change - [[phab:T428909|T428909]] [[phab:T431838|T431838]] * 14:05 cdobbins@cumin2002: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie * 14:02 marostegui@dns1004: END - running authdns-update * 14:00 marostegui@dns1004: START - running authdns-update * 14:00 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1070.eqiad.wmnet * 14:00 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1070.eqiad.wmnet * 13:58 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2009.codfw.wmnet with reason: host reimage * 13:57 cdobbins@cumin2002: conftool action : set/pooled=no; selector: name=dns7002.* * 13:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2009.codfw.wmnet with reason: host reimage * 13:48 rscout@deploy2003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply * 13:48 rscout@deploy2003: helmfile [eqiad] START helmfile.d/services/miscweb: apply * 13:48 rscout@deploy2003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply * 13:47 rscout@deploy2003: helmfile [codfw] START helmfile.d/services/miscweb: apply * 13:40 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:33 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2009.codfw.wmnet with OS trixie * 13:30 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1070.eqiad.wmnet with reason: vacuum overlarge container dbs * 13:28 aude@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] (duration: 11m 12s) * 13:23 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:22 aude@deploy2003: aikochou, javiermonton, aude, gkm563: Continuing with deployment * 13:22 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:19 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:19 aude@deploy2003: aikochou, javiermonton, aude, gkm563: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] synced to the testservers * 13:17 aude@deploy2003: Started scap sync-world: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] * 13:01 ladsgroup@deploy2003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 13:01 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:00 ladsgroup@deploy2003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 12:59 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:52 ladsgroup@deploy2003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 12:51 ladsgroup@deploy2003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 12:48 atsuko@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 12:48 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 12:47 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2008.codfw.wmnet with OS trixie * 12:47 atsuko@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 12:47 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply * 12:47 atsuko@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:46 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 12:45 atsuko@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:45 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply * 12:45 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:44 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:43 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] (duration: 07m 02s) * 12:38 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 12:37 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:36 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] * 12:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2008.codfw.wmnet with reason: host reimage * 12:23 Msz2001: Deployed changes to private code for Suggested Investigations * 12:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2008.codfw.wmnet with reason: host reimage * 12:20 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:19 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:17 atsuko@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 12:17 atsuko@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 12:16 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:15 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] (duration: 07m 14s) * 12:10 mszwarc@deploy2003: mszwarc: Continuing with deployment * 12:09 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:07 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] * 12:04 mszwarc@deploy2003: sync-world aborted: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] (duration: 00m 29s) * 12:03 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] * 12:01 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2008.codfw.wmnet with OS trixie * 12:00 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] (duration: 07m 37s) * 11:55 zabe@deploy2003: zabe: Continuing with deployment * 11:54 zabe@deploy2003: zabe: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:52 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] * 11:51 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:43 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:35 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:34 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:33 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:30 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:28 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:27 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:17 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2007.codfw.wmnet with OS trixie * 11:09 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] * 11:06 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=s8 * 11:00 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=x3 * 11:00 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=s5 * 10:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2007.codfw.wmnet with reason: host reimage * 10:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2007.codfw.wmnet with reason: host reimage * 10:51 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host dse-k8s-worker1023 * 10:50 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host dse-k8s-worker1023 * 10:44 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dse-k8s-worker1023 * 10:43 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host dse-k8s-worker1023 * 10:42 marostegui@cumin1003: dbctl commit (dc=all): 'Change x4 masters [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P94804 and previous config saved to /var/cache/conftool/dbconfig/20260713-104248-marostegui.json * 10:37 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:37 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:35 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:35 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:34 atsuko@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 10:34 atsuko@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 10:33 marostegui@cumin1003: dbctl commit (dc=all): 'Push x4 initial dbctl config [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P94803 and previous config saved to /var/cache/conftool/dbconfig/20260713-103259-marostegui.json * 10:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2007.codfw.wmnet with OS trixie * 09:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2006.codfw.wmnet with OS trixie * 09:42 marostegui@dns1004: END - running authdns-update * 09:40 marostegui@dns1004: START - running authdns-update * 09:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2006.codfw.wmnet with reason: host reimage * 09:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2006.codfw.wmnet with reason: host reimage * 09:06 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1024.eqiad.wmnet with reason: reboot * 09:06 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:01 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2006.codfw.wmnet with OS trixie * 08:44 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 08:43 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 08:43 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:42 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:42 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:42 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:41 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 08:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup2004.codfw.wmnet * 08:38 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:33 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host db1208.eqiad.wmnet * 08:30 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=x3 * 08:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2005.codfw.wmnet with OS trixie * 08:28 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup2004.codfw.wmnet * 08:28 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup2003.codfw.wmnet * 08:24 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1039: Repooling after testing * 08:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on clouddb1016.eqiad.wmnet with reason: cloning * 08:23 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s5 * 08:23 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s8 * 08:21 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 08:21 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 08:17 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup2003.codfw.wmnet * 08:17 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1004.eqiad.wmnet * 08:14 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1208.eqiad.wmnet * 08:11 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host phab1005.eqiad.wmnet * 08:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2005.codfw.wmnet with reason: host reimage * 08:07 marostegui@dns1004: END - running authdns-update * 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1004.eqiad.wmnet * 08:07 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1003.eqiad.wmnet * 08:07 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:05 marostegui@dns1004: START - running authdns-update * 08:05 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2005.codfw.wmnet with reason: host reimage * 08:05 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 08:05 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host phab1005.eqiad.wmnet * 08:05 marostegui@dns1004: START - running authdns-update * 08:05 marostegui@dns1004: START - running authdns-update * 08:05 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 08:04 marostegui@dns1004: START - running authdns-update * 08:00 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit1003.wikimedia.org * 07:58 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1003.eqiad.wmnet * 07:58 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1002-dev.eqiad.wmnet * 07:58 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:58 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:54 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1002-dev.eqiad.wmnet * 07:54 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1001-dev.eqiad.wmnet * 07:54 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit1003.wikimedia.org * 07:53 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 07:53 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:52 Msz2001: UTC morning backport+config window done * 07:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2005.codfw.wmnet with OS trixie * {{safesubst:SAL entry|1=07:50 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark (T429943}} * 07:49 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1001-dev.eqiad.wmnet * 07:46 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:46 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:45 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 07:45 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:44 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 07:44 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:43 mszwarc@deploy2003: mszwarc, danielyepezgarces, anzx: Continuing with deployment * {{safesubst:SAL entry|1=07:39 mszwarc@deploy2003: mszwarc, danielyepezgarces, anzx: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark}} * 07:39 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1039: Repooling after testing * {{safesubst:SAL entry|1=07:36 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark (T429943)}} * 07:35 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] (duration: 30m 03s) * 07:25 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit2002.wikimedia.org * 07:22 mszwarc@deploy2003: mszwarc: Continuing with deployment * 07:21 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:19 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit2002.wikimedia.org * 07:15 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aphlict1002.eqiad.wmnet * 07:11 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host aphlict1002.eqiad.wmnet * 07:08 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2003.wikimedia.org * 07:05 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] * 07:02 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2003.wikimedia.org * 07:02 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2002.wikimedia.org * 06:55 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2002.wikimedia.org * 06:55 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1003.wikimedia.org * 06:49 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1003.wikimedia.org * 06:34 marostegui: Drop m5 ipoid database [[phab:T431007|T431007]] * 06:29 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1027.eqiad.wmnet with reason: reboot * 06:24 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1028.eqiad.wmnet with reason: reboot * 06:21 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1025.eqiad.wmnet with reason: reboot * 06:17 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1022.eqiad.wmnet with reason: reboot * 06:03 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on dbproxy[2005-2008].codfw.wmnet with reason: reboot * 05:37 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1217,1228].eqiad.wmnet with reason: cloning * 05:11 marostegui: Drop users_to_rename table [[phab:T431842|T431842]] * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-12 == * 16:01 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2209 [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94792 and previous config saved to /var/cache/conftool/dbconfig/20260712-160124-marostegui.json * 15:58 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2205 to s3 primary [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94791 and previous config saved to /var/cache/conftool/dbconfig/20260712-155853-marostegui.json * 15:58 marostegui: Starting s3 codfw emergency failover from db2209 to db2205 - [[phab:T431950|T431950]] * 15:51 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2205 with weight 0 [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94790 and previous config saved to /var/cache/conftool/dbconfig/20260712-155135-marostegui.json * 15:51 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Primary switchover s3 [[phab:T431950|T431950]] * 02:01 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 01m 17s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-11 == * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 26s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-10 == * 19:12 jhathaway@dns1004: END - running authdns-update * 19:10 jhathaway@dns1004: START - running authdns-update * 18:23 mutante: vrts2002 rebooting (not the active host) * 18:21 mutante: lists2001, phab2003 - rebooting (not the active hosts) * 18:16 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on A:lvs-high-traffic2-codfw * 18:15 swfrench@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on A:lvs-high-traffic2-codfw * 17:15 mutante: [doc1004:~] $ sudo systemctl start rsync-doc-host-data-sync ([[phab:T431856|T431856]]) * 17:09 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1004.eqiad.wmnet * 17:08 jhathaway@dns1004: END - running authdns-update * 17:07 jhathaway@dns1004: START - running authdns-update * 17:06 jhathaway: depooling puppetserver1002, cause of errors is still unknown * 17:03 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1004.eqiad.wmnet * 16:57 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1003.eqiad.wmnet * 16:51 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1003.eqiad.wmnet * 16:48 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 16:48 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2004.codfw.wmnet * 16:42 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2004.codfw.wmnet * 16:41 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2003.codfw.wmnet * 16:35 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2003.codfw.wmnet * 16:33 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2002.codfw.wmnet * 16:27 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2002.codfw.wmnet * 16:25 mutante: gitlab-runners (production) rebooting cluster one by one * 16:17 mutante: etherpad1004/etherpad2002 - (etherpad.wikimedia.org) - rebooting * 16:13 mutante: doc1004/doc2003 (doc.wikimedia.org backends) - rebooting * 16:02 mutante: releases1003/releases2003 (releases.wikimedia.org backends) - rebooting for maintenance * 15:26 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2007-dev.codfw.wmnet * 15:19 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2007-dev.codfw.wmnet * 15:14 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host cloudcephosd2007-dev.codfw.wmnet * 15:14 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2007-dev.codfw.wmnet * 15:14 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host cloudcephosd2006-dev.codfw.wmnet * 15:07 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2006-dev.codfw.wmnet * 15:07 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2005-dev.codfw.wmnet * 14:59 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2005-dev.codfw.wmnet * 14:59 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2004-dev.codfw.wmnet * 14:53 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2004-dev.codfw.wmnet * 14:53 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2007-dev.codfw.wmnet * 14:51 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1054.eqiad.wmnet * 14:51 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1054.eqiad.wmnet * 14:51 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1054.eqiad.wmnet * 14:47 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2007-dev.codfw.wmnet * 14:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2006-dev.codfw.wmnet * 14:41 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2006-dev.codfw.wmnet * 14:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2005-dev.codfw.wmnet * 14:37 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2005-dev.codfw.wmnet * 14:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2005-dev.codfw.wmnet * 14:29 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2005-dev.codfw.wmnet * 14:29 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2006-dev.codfw.wmnet * 14:21 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2006-dev.codfw.wmnet * 14:21 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2010-dev.codfw.wmnet * 14:15 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2010-dev.codfw.wmnet * 14:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudgw2004-dev.codfw.wmnet * 14:10 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1054.eqiad.wmnet with OS trixie * 14:09 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudgw2004-dev.codfw.wmnet * 14:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudgw2003-dev.codfw.wmnet * 14:02 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudgw2003-dev.codfw.wmnet * 14:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2004-dev.codfw.wmnet * 13:53 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2004-dev.codfw.wmnet * 13:53 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2003-dev.codfw.wmnet * 13:48 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:44 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2003-dev.codfw.wmnet * 13:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2002-dev.codfw.wmnet * 13:42 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:41 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:41 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:37 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2002-dev.codfw.wmnet * 13:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudidp2001-dev.codfw.wmnet * 13:33 blake@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker1054.eqiad.wmnet with reason: host reimage * 13:33 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudidp2001-dev.codfw.wmnet * 13:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudnet2006-dev.codfw.wmnet * 13:26 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudnet2006-dev.codfw.wmnet * 13:26 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudnet2005-dev.codfw.wmnet * 13:23 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1054.eqiad.wmnet with reason: host reimage * 13:18 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudnet2005-dev.codfw.wmnet * 13:18 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudservices2005-dev.codfw.wmnet * 13:12 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudservices2005-dev.codfw.wmnet * 13:11 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudservices2004-dev.codfw.wmnet * 13:08 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudservices2004-dev.codfw.wmnet * 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudweb2002-dev.wikimedia.org * 13:05 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 13:05 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1054 * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1054 * 13:04 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1054 * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1054.eqiad.wmnet 49.32.64.10.in-addr.arpa 9.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:04 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1054.eqiad.wmnet 49.32.64.10.in-addr.arpa 9.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1054 - blake@cumin1003" * 13:04 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1054 - blake@cumin1003" * 13:01 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudweb2002-dev.wikimedia.org * 13:00 blake@cumin1003: START - Cookbook sre.dns.netbox * 12:59 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1054 * 12:57 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1054.eqiad.wmnet with OS trixie * 12:57 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1054.eqiad.wmnet * 12:56 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1054.eqiad.wmnet * 12:56 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1054.eqiad.wmnet * 12:47 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS trixie * 12:44 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:39 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 12:39 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 12:38 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:37 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:14 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:10 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:08 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:07 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:00 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:00 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:51 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:49 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:48 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:47 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:44 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:32 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2001.codfw.wmnet * 11:32 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1053.eqiad.wmnet * 11:32 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2001.codfw.wmnet * 11:32 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1053.eqiad.wmnet * 11:32 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1053.eqiad.wmnet * 11:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker2001.codfw.wmnet * 11:31 cgoubert@cumin1003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker2001.codfw.wmnet * 11:31 cgoubert@cumin1003: END (FAIL) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=1) rolling reimage on P<nowiki>{</nowiki>wikikube-worker2001*<nowiki>}</nowiki> and (A:wikikube-master-codfw or A:wikikube-worker-codfw) * 11:30 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:30 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:21 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 18 hosts with reason: reboot & upgrade * 11:20 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker2001.codfw.wmnet with OS trixie * 11:16 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:15 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:14 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:14 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 11 hosts * 11:14 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 11 hosts * 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:08 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 11:02 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1053.eqiad.wmnet with OS trixie * 11:01 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:58 cgoubert@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 10:57 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:38 cgoubert@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker2001.codfw.wmnet with OS trixie * 10:38 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2001.codfw.wmnet * 10:38 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2001.codfw.wmnet * 10:38 cgoubert@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on P<nowiki>{</nowiki>wikikube-worker2001*<nowiki>}</nowiki> and (A:wikikube-master-codfw or A:wikikube-worker-codfw) * 10:35 topranks: adjust IBGP outbound policy on lsw1-e2-codfw [[phab:T423430|T423430]] towards ssw1-e1-codfw * 10:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 cgoubert@cumin1003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:27 cgoubert@cumin1003: END (FAIL) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=1) rolling reimage on A:wikikube-worker-codfw * 10:27 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker2001.codfw.wmnet with OS bookworm * 10:25 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 10:24 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:24 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:15 cgoubert@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 10:11 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 10:11 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 10:08 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:08 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:07 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 10:06 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 10:00 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:55 cgoubert@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker2001.codfw.wmnet with OS bookworm * 09:55 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2005-2006,2011-2012].codfw.wmnet * 09:55 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2005-2006,2011-2012].codfw.wmnet * 09:51 cgoubert@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on A:wikikube-worker-codfw * 09:41 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1053.eqiad.wmnet with reason: host reimage * 09:37 topranks: apply new IBGP outbound policy on lsw1-e2-codfw [[phab:T423430|T423430]] * 09:36 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:36 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1053.eqiad.wmnet with reason: host reimage * 09:16 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1053 * 09:16 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1053 * 09:15 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1053 * 09:15 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1053.eqiad.wmnet 48.32.64.10.in-addr.arpa 8.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:15 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1053.eqiad.wmnet 48.32.64.10.in-addr.arpa 8.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:15 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:15 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1053 - blake@cumin1003" * 09:15 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1053 - blake@cumin1003" * 09:11 blake@cumin1003: START - Cookbook sre.dns.netbox * 09:11 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1053 * 09:08 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1053.eqiad.wmnet with OS trixie * 09:08 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1053.eqiad.wmnet * 09:08 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1053.eqiad.wmnet * 09:08 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1053.eqiad.wmnet * 09:04 brouberol@dns1004: END - running authdns-update * 09:03 brouberol@dns1004: START - running authdns-update * 08:41 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e] (thin): Regular analytics weekly train THIN [analytics/refinery@1abf22ea] (duration: 02m 11s) * 08:38 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e] (thin): Regular analytics weekly train THIN [analytics/refinery@1abf22ea] * 08:38 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e]: Regular analytics weekly train [analytics/refinery@1abf22ea] (duration: 05m 17s) * 08:38 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:34 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:33 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e]: Regular analytics weekly train [analytics/refinery@1abf22ea] * 08:32 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@1abf22ea] (duration: 02m 03s) * 08:30 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@1abf22ea] * 08:30 JavierMonton: Deploying Refinery at {{Gerrit|1abf22ea}} for changes 1308121/T427068 1306491/T430020 and {{Gerrit|1308190}} * 08:29 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:29 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 08:24 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:24 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 08:18 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:18 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 08:00 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db[2183-2184].codfw.wmnet * 08:00 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for db[2183-2184].codfw.wmnet * 07:52 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:52 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 07:49 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 11 hosts with reason: reboot & upgrade * 07:47 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:47 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 07:44 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:44 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 07:23 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 10 hosts * 07:23 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 10 hosts * 06:45 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 10 hosts with reason: reboot & upgrade * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 41s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-09 == * 23:33 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] (duration: 13m 26s) * 23:29 ladsgroup@deploy2003: ladsgroup, jdlrobson: Continuing with deployment * 23:22 ladsgroup@deploy2003: ladsgroup, jdlrobson: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:20 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] * 22:57 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1165.eqiad.wmnet * 22:56 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1165.eqiad.wmnet * 22:56 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1165.eqiad.wmnet * 22:45 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1165.eqiad.wmnet with OS trixie * 22:38 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 22:37 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 22:37 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 22:37 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:37 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 22:25 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1165.eqiad.wmnet with reason: host reimage * 22:17 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1165.eqiad.wmnet with reason: host reimage * 22:13 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 22:12 rzl@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 22:04 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 22:04 rzl@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1165 * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1165 * 22:02 jasmine@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1165 * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1165.eqiad.wmnet 115.48.64.10.in-addr.arpa 5.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:02 jasmine@cumin2002: START - Cookbook sre.dns.wipe-cache wikikube-worker1165.eqiad.wmnet 115.48.64.10.in-addr.arpa 5.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1165 - jasmine@cumin2002" * 22:02 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1165 - jasmine@cumin2002" * 22:02 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 21:57 jasmine@cumin2002: START - Cookbook sre.dns.netbox * 21:55 jasmine@cumin2002: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1165 * 21:54 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-worker1165.eqiad.wmnet with OS trixie * 21:54 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 21:54 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1165.eqiad.wmnet * 21:53 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 21:53 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1165.eqiad.wmnet * 21:53 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1165.eqiad.wmnet * 21:53 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 21:47 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 21:45 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 21:43 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 21:43 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 21:42 maryum: Deploy fix for [[phab:T431684|T431684]] * 21:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 21:27 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] (duration: 34m 14s) * 21:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs1002 * 21:23 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs1002 * 21:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS trixie * 21:22 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 21:20 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 22s) * 21:20 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 21:16 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 21:16 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 21:15 ladsgroup@deploy2003: ladsgroup: Continuing with deployment * 21:13 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2002.codfw.wmnet with OS bookworm * 21:11 ladsgroup@deploy2003: ladsgroup: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:08 ladsgroup@cumin1003: END (PASS) - Cookbook sre.wikireplicas.update-views (exit_code=0) * 21:07 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 6 hosts with reason: reboots * 20:54 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecycle work - bking@cumin2003 * 20:53 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:53 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] * 20:51 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) * 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 20:47 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecycle work - bking@cumin2003 * 20:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 20:41 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:41 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) * 20:40 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host relforge1008.eqiad.wmnet * 20:40 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1009.eqiad.wmnet with OS trixie * 20:33 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:32 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:32 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:31 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:31 ladsgroup@cumin1003: END (PASS) - Cookbook sre.wikireplicas.update-views (exit_code=0) * 20:29 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1008.eqiad.wmnet * 20:24 rzl@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 20:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2002.codfw.wmnet with OS bookworm * 20:23 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host relforge1008.eqiad.wmnet * 20:23 rzl@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 20:23 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1008.eqiad.wmnet * 20:22 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:22 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:21 bking@cumin2003: END (ERROR) - Cookbook sre.elasticsearch.rolling-operation (exit_code=97) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:21 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1009.eqiad.wmnet with reason: host reimage * 20:16 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:15 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1009.eqiad.wmnet with reason: host reimage * 20:12 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) * 20:02 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 19:55 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1009.eqiad.wmnet with OS trixie * 19:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 19:43 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 19:30 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 19:28 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 19:27 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 19:25 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 18:42 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 18:41 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 18:16 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 18:15 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 17:45 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for doh5004.wikimedia.org * 17:45 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for doh5004.wikimedia.org * 17:38 ladsgroup@deploy2003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 17:35 ladsgroup@deploy2003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 17:29 ladsgroup@deploy2003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 17:26 ladsgroup@deploy2003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 17:09 mutante: zuul[12]00[123] - rebooting for maintenance * 17:09 ebernhardson: start full in-place reindex of eqiad cirrussearch cluster * 17:08 dzahn@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-cluster (exit_code=99) * 17:08 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-cluster * 17:03 ebernhardson: start full in-place reindex of codfw cirrussearch cluster * 16:59 mutante: stewards1001/stewards2001 - reboot for maintenance * 16:54 ebernhardson: start full in-place reindex of cloudelastic cluster * 16:53 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 16:52 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply * 16:49 mutante: planet1003/planet2003 - rebooting * 16:47 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on doh5004.wikimedia.org with reason: random high load, investigating * 15:55 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 15:54 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 15:51 jynus: restarting backupmon1001 * 15:49 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 14 hosts * 15:49 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 14 hosts * 15:47 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backupmon1001.eqiad.wmnet with reason: restart * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:06 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 15:06 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 14:59 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply * 14:58 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply * 14:51 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 14 hosts * 14:51 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 14 hosts * 14:49 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 6 hosts with reason: reboot & upgrade * 14:48 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet * 14:48 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet * 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:42 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1052.eqiad.wmnet * 14:42 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1052.eqiad.wmnet * 14:42 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1052.eqiad.wmnet * 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:31 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:31 elukey@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: sync * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:30 elukey@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: sync * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:28 elukey@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: sync * 14:28 elukey@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: sync * 14:26 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:20 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1052.eqiad.wmnet with OS trixie * 14:19 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 6 hosts with reason: reboot & upgrade * 14:18 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:15 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:15 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:13 elukey: update druid indexation job for webrequest_sampled_live - [[phab:T427068|T427068]] * 14:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:09 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:09 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for papaul - jhancock@cumin2002" * 14:09 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for papaul - jhancock@cumin2002" * 14:07 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:07 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:04 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 14:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cuminunpriv1001.eqiad.wmnet * 13:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb1003.eqiad.wmnet * 13:59 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1052.eqiad.wmnet with reason: host reimage * 13:57 moritzm: installing requests security updates * 13:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cuminunpriv1001.eqiad.wmnet * 13:55 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb1003.eqiad.wmnet * 13:53 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1052.eqiad.wmnet with reason: host reimage * 13:50 moritzm: installing python-cryptography security updates * 13:47 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb2003.codfw.wmnet * 13:44 Msz2001: UTC afternoon config+backport window is done * 13:44 Msz2001: Updated `logging` on `metawiki` to fix log performers, [[phab:T431176|T431176]]#12105297 * 13:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb2003.codfw.wmnet * 13:43 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 13:43 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt1002.wikimedia.org * 13:41 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] (duration: 07m 30s) * 13:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt1002.wikimedia.org * 13:37 mszwarc@deploy2003: mszwarc: Continuing with deployment * 13:36 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1052 * 13:36 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1052 * 13:35 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:35 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1052 * 13:35 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1052.eqiad.wmnet 47.32.64.10.in-addr.arpa 7.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:35 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1052.eqiad.wmnet 47.32.64.10.in-addr.arpa 7.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:35 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:35 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1052 - blake@cumin1003" * 13:35 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1052 - blake@cumin1003" * 13:34 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] * 13:31 blake@cumin1003: START - Cookbook sre.dns.netbox * 13:31 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1052 * 13:30 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1052.eqiad.wmnet with OS trixie * 13:30 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1052.eqiad.wmnet * 13:29 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1052.eqiad.wmnet * 13:29 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1052.eqiad.wmnet * 13:17 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] (duration: 11m 26s) * 13:13 jforrester@deploy2003: jforrester: Continuing with deployment * 13:08 jforrester@deploy2003: jforrester: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:06 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] * 12:54 cgoubert@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply * 12:54 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:52 cgoubert@deploy2003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply * 12:45 cgoubert@deploy2003: helmfile [codfw] DONE helmfile.d/services/mobileapps: apply * 12:44 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:44 cgoubert@deploy2003: helmfile [codfw] START helmfile.d/services/mobileapps: apply * 12:43 cgoubert@deploy2003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 12:43 cgoubert@deploy2003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 12:42 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast4006.wikimedia.org * 12:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt2002.wikimedia.org * 12:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast7002.wikimedia.org * 12:18 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast4006.wikimedia.org * 12:18 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host ml-serve1004 * 12:18 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host ml-serve1004 * 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt2002.wikimedia.org * 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast7002.wikimedia.org * 12:10 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backup[2003,2014].codfw.wmnet with reason: reboot & upgrade * 12:10 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt-staging2001.codfw.wmnet * 12:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid1003.eqiad.wmnet * 12:06 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt-staging2001.codfw.wmnet * 12:05 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid1003.eqiad.wmnet * 12:03 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backup[1003,1014].eqiad.wmnet with reason: reboot & upgrade * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid2003.codfw.wmnet * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host irc1003.wikimedia.org * 11:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid2003.codfw.wmnet * 11:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host irc1003.wikimedia.org * 11:55 jmm@dns1004: END - running authdns-update * 11:53 jmm@dns1004: START - running authdns-update * 11:50 jmm@dns1004: END - running authdns-update * 11:48 jmm@dns1004: START - running authdns-update * 11:27 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host irc2003.wikimedia.org * 11:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host irc2003.wikimedia.org * 11:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint2001.codfw.wmnet * 11:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint1001.eqiad.wmnet * 11:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint2001.codfw.wmnet * 11:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint1001.eqiad.wmnet * 11:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-rw2001.wikimedia.org * 11:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-rw1001.wikimedia.org * 11:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-rw2001.wikimedia.org * 11:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-rw1001.wikimedia.org * 11:03 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon1003.wikimedia.org * 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2005.codfw.wmnet * 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2005.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 10:59 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2005.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 10:57 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon1003.wikimedia.org * 10:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon2002.wikimedia.org * 10:55 jmm@cumin2003: START - Cookbook sre.dns.netbox * 10:51 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon2002.wikimedia.org * 10:51 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:50 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2005.codfw.wmnet * 10:41 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:40 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host ml-serve1003 * 10:40 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host ml-serve1003 * 10:39 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2033.codfw.wmnet * 10:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install2005.wikimedia.org * 10:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install1005.wikimedia.org * 10:35 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1004.eqiad.wmnet with OS bookworm * 10:31 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install1005.wikimedia.org * 10:31 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install2005.wikimedia.org * 10:30 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install4004.wikimedia.org * 10:30 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install3004.wikimedia.org * 10:29 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install3004.wikimedia.org * 10:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install4004.wikimedia.org * 10:23 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 10:21 moritzm: failover Ganeti master in codfw/routed to ganeti2034 [[phab:T430928|T430928]] * 10:19 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.addnode (exit_code=0) for new host ganeti2031.codfw.wmnet to cluster codfw and group B * 10:19 moritzm: readded ganeti2031 to the codfw Ganeti cluster [[phab:T430910|T430910]] * 10:18 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1004.eqiad.wmnet with reason: host reimage * 10:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install5004.wikimedia.org * 10:18 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1003 * 10:18 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1003 * 10:17 jmm@cumin2003: START - Cookbook sre.ganeti.addnode for new host ganeti2031.codfw.wmnet to cluster codfw and group B * 10:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install6003.wikimedia.org * 10:16 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install5004.wikimedia.org * 10:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install6003.wikimedia.org * 10:15 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1004.eqiad.wmnet with reason: host reimage * 10:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1001.eqiad.wmnet * 10:14 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 10:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2008.wikimedia.org * 10:00 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ml-serve1004 * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1004 * 09:57 jmm@cumin2003: START - Cookbook sre.dns.netbox * 09:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install7002.wikimedia.org * 09:57 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1004 * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ml-serve1004.eqiad.wmnet 50.48.64.10.in-addr.arpa 0.5.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:57 klausman@cumin1003: START - Cookbook sre.dns.wipe-cache ml-serve1004.eqiad.wmnet 50.48.64.10.in-addr.arpa 0.5.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1004 - klausman@cumin1003" * 09:56 klausman@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1004 - klausman@cumin1003" * 09:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-coord1001.eqiad.wmnet * 09:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 09:55 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow7002.magru.wmnet * 09:52 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-coord1001.eqiad.wmnet * 09:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 09:52 klausman@cumin1003: START - Cookbook sre.dns.netbox * 09:50 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install7002.wikimedia.org * 09:50 klausman@cumin1003: START - Cookbook sre.hosts.move-vlan for host ml-serve1004 * 09:50 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1004.eqiad.wmnet with OS bookworm * 09:50 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1003.eqiad.wmnet with OS bookworm * 09:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1001.eqiad.wmnet * 09:49 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 09:49 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2008.wikimedia.org * 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2007.codfw.wmnet * 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2007.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 09:49 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow7002.magru.wmnet * 09:49 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2007.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 09:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard1003.eqiad.wmnet * 09:39 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard2003.codfw.wmnet * 09:39 jmm@cumin2003: START - Cookbook sre.dns.netbox * 09:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard1003.eqiad.wmnet * 09:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor1003.eqiad.wmnet * 09:35 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard2003.codfw.wmnet * 09:34 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2007.codfw.wmnet * 09:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor1003.eqiad.wmnet * 09:33 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor-dev2001.codfw.wmnet * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor2003.codfw.wmnet * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sretest1006.eqiad.wmnet * 09:27 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 09:25 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor-dev2001.codfw.wmnet * 09:25 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor2003.codfw.wmnet * 09:23 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] (duration: 06m 27s) * 09:23 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2205: codfw rack B4 repool after maintenance * 09:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host sretest1006.eqiad.wmnet * 09:23 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2204: codfw rack B4 repool after maintenance * 09:19 urbanecm@deploy2003: urbanecm: Continuing with deployment * 09:19 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:18 jmm@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 6 hosts with reason: reboot * 09:17 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] * 09:08 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ml-serve1003 * 09:08 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1003 * 09:07 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1003 * 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ml-serve1003.eqiad.wmnet 81.32.64.10.in-addr.arpa 1.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:07 klausman@cumin1003: START - Cookbook sre.dns.wipe-cache ml-serve1003.eqiad.wmnet 81.32.64.10.in-addr.arpa 1.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1003 - klausman@cumin1003" * 09:06 klausman@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1003 - klausman@cumin1003" * 08:58 klausman@cumin1003: START - Cookbook sre.dns.netbox * 08:57 klausman@cumin1003: START - Cookbook sre.hosts.move-vlan for host ml-serve1003 * 08:57 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1003.eqiad.wmnet with OS bookworm * 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=0) rolling reimage on P<nowiki>{</nowiki>ml-serve1003.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet * 08:55 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet * 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1003.eqiad.wmnet with OS bookworm * 08:39 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 08:38 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool db2205: codfw rack B4 repool after maintenance * 08:37 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool db2204: codfw rack B4 repool after maintenance * 08:36 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 08:35 hashar@deploy2003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 08:32 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:32 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:31 hashar@deploy2003: Rolling back deployment * 08:26 moritzm: failover Ganeti master in codfw to ganeti2048 [[phab:T430928|T430928]] * 08:16 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1003.eqiad.wmnet with OS bookworm * 08:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2004.codfw.wmnet * 08:16 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet * 08:16 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet * 08:16 klausman@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on P<nowiki>{</nowiki>ml-serve1003.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 08:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2002.codfw.wmnet * 08:15 XioNoX: lsw1-b4-codfw> request system reboot - [[phab:T430910|T430910]] * 08:15 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b4-codfw,lsw1-b4-codfw IPv6,lsw1-b4-codfw.mgmt with reason: Switch maintenance * 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for codfw rack B4 * 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:10 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2004.codfw.wmnet * 08:10 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2002.codfw.wmnet * 08:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2205: codfw rack B4 depool for maintenance * 08:08 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool db2205: codfw rack B4 depool for maintenance * 08:08 jmm@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin2003.codfw.wmnet * 08:08 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2204: codfw rack B4 depool for maintenance * 08:08 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool db2204: codfw rack B4 depool for maintenance * 08:08 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 27 hosts with reason: codfw rack B4 depool for maintenance * 08:03 jmm@cumin2002: START - Cookbook sre.hosts.reboot-single for host cumin2003.codfw.wmnet * 07:56 ayounsi@cumin1003: START - Cookbook sre.network.depool-rack with action 'depool' for codfw rack B4 * 07:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1008.eqiad.wmnet with OS trixie * 07:49 wmde-fisch@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] (duration: 08m 36s) * 07:44 wmde-fisch@deploy2003: wmde-fisch: Continuing with deployment * 07:43 wmde-fisch@deploy2003: wmde-fisch: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:41 wmde-fisch@deploy2003: Started scap sync-world: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] * 07:35 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1008.eqiad.wmnet with reason: host reimage * 07:31 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1008.eqiad.wmnet with reason: host reimage * 07:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1008.eqiad.wmnet with OS trixie * 07:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 07:00 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 06:59 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1008.eqiad.wmnet with OS trixie * 06:57 Emperor: rebalance thanos swift rings after previous re-image of thanos-fe1004 to trixie * 06:47 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1008.eqiad.wmnet with OS trixie * 04:10 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 14 days, 0:00:00 on cp6008.drmrs.wmnet with reason: Hardware failure - [[phab:T431651|T431651]] * 03:55 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp6008.* * 03:29 ryankemper: [[phab:T431311|T431311]] Repooled eqiad cirrussearch clusters (`chi/omega/psi`) following completion of OpenSearch 2.19 migration * 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad * 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=eqiad * 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 31s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-08 == * 23:52 Amir1: ladsgroup@deploy2003:~$ mwscript-k8s --follow -- extensions/ORES/maintenance/PurgeScoreCache.php --wiki=simplewiki --model damaging --old ([[phab:T431159|T431159]]) * 23:46 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 23:46 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing PTR for 2001:df2:e500:fe08::1 - cmooney@cumin1003" * 23:46 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing PTR for 2001:df2:e500:fe08::1 - cmooney@cumin1003" * 23:40 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 23:16 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 23:15 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 22:42 rzl@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 22:40 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] (duration: 12m 55s) * 22:40 rzl@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 22:37 rzl@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 22:36 rzl@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 22:35 rzl@deploy2003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 22:34 urbanecm@deploy2003: urbanecm: Continuing with deployment * 22:33 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:33 rzl@deploy2003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 22:32 rzl@deploy2003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 22:30 rzl@deploy2003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 22:30 rzl@deploy2003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 22:29 rzl@deploy2003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 22:27 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] * 22:26 rzl@deploy2003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 22:22 rzl@deploy2003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 22:21 rzl@deploy2003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 22:19 rzl@deploy2003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 22:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 22:17 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 22:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 22:13 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 22:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 22:13 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 22:09 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 22:06 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 22:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1094.eqiad.wmnet with OS trixie * 22:01 urbanecm: Make https://test.wikipedia.org/w/index.php?title=MediaWiki:GrowthExperimentsSuggestedEdits.json&diff=prev&oldid=750552 with GrowthExperiments disabled (via mw-experimental), then run `\MediaWiki\MediaWikiServices::getInstance()->get('CommunityConfiguration.ProviderFactory')->newProvider('GrowthSuggestedEdits')->getStore()->invalidate()` ([[phab:T431625|T431625]]) * 21:56 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d2-codfw * 21:55 urbanecm@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 21:55 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d2-codfw * 21:55 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c4-codfw * 21:55 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c4-codfw * 21:55 urbanecm@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2002 * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2002 * 21:54 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2002 * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2002.codfw.wmnet 50.32.192.10.in-addr.arpa 0.5.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:54 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2002.codfw.wmnet 50.32.192.10.in-addr.arpa 0.5.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2002 - bking@cumin2003" * 21:54 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2002 - bking@cumin2003" * 21:49 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:49 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2002 * 21:49 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2002.codfw.wmnet with OS trixie * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1094.eqiad.wmnet with reason: host reimage * 21:42 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 21:39 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 21:37 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1094.eqiad.wmnet with reason: host reimage * 21:36 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 21:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 21:29 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 21:27 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 21:22 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1094.eqiad.wmnet with OS trixie * 21:21 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host restbase2039.codfw.wmnet with OS bullseye * 21:21 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin2002" * 21:21 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin2002" * 21:04 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on restbase2039.codfw.wmnet with reason: host reimage * 21:00 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on restbase2039.codfw.wmnet with reason: host reimage * 20:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1073.eqiad.wmnet with OS trixie * 20:48 mutante: deploy2003 - kill 1102 (stunnel4) ; systemctl start stunnel4 ([[phab:T418262|T418262]]) * 20:42 cjming@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] (duration: 33m 02s) * 20:42 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host restbase2039.codfw.wmnet with OS bullseye * 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1073.eqiad.wmnet with reason: host reimage * 20:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1073.eqiad.wmnet with reason: host reimage * 20:30 cjming@deploy2003: cjming: Continuing with deployment * 20:28 cjming@deploy2003: cjming: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1098.eqiad.wmnet with OS trixie * 20:13 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1073.eqiad.wmnet with OS trixie * 20:09 cjming@deploy2003: Started scap sync-world: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] * 20:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1098.eqiad.wmnet with reason: host reimage * 19:56 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1098.eqiad.wmnet with reason: host reimage * 19:55 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d4-codfw * 19:54 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d4-codfw * 19:54 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c1-codfw * 19:54 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c1-codfw * 19:52 mutante: restarting gerrit on gerrit.wikimedia.org (gerrit2003) * 19:48 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2331.codfw.wmnet * 19:48 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2331.codfw.wmnet * 19:48 mutante: restarting gerrit on gerrit-replica.wikimedia.org (gerrit1003) * 19:46 mutante: restarting gerrit on gerrit-spare.wikimedia.org (gerrit2002) * 19:43 jasmine@cumin2002: conftool action : set/pooled=yes; selector: name=wikikube-worker2331.codfw.wmnet,cluster=kubernetes,service=kubesvc * 19:43 jasmine@cumin2002: conftool action : set/weight=10; selector: name=wikikube-worker2331.codfw.wmnet,cluster=kubernetes,service=kubesvc * 19:40 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1098.eqiad.wmnet with OS trixie * 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d5-codfw * 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d5-codfw * 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c7-codfw * 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c7-codfw * 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c5-codfw * 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c5-codfw * 19:30 jasmine_: ran homer on lsw1-d8-codfw, adding wikikube-worker2331 to cluster * 19:29 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1100.eqiad.wmnet with OS trixie * 19:20 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d8-codfw * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d8-codfw * 19:19 mutante: gerrit - replacing private key for registerEmail verification * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-magru * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device cr2-magru * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d7-codfw * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d7-codfw * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d3-codfw * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d1-codfw * 19:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d1-codfw * 19:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c2-codfw * 19:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c2-codfw * 19:11 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-codfw * 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-magru * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device cr1-magru * 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d8-codfw * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d8-codfw * 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d6-codfw * 19:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1100.eqiad.wmnet with reason: host reimage * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d6-codfw * 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c6-codfw * 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c6-codfw * 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c3-codfw * 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c3-codfw * 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b4-magru * 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device asw1-b4-magru * 19:08 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b3-magru * 19:08 cmooney@cumin1003: START - Cookbook sre.network.tls for network device asw1-b3-magru * 19:05 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1100.eqiad.wmnet with reason: host reimage * 19:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1122.eqiad.wmnet with OS trixie * 19:00 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 18:59 topranks: rolling out update to BGP ACL on Nokia Switches eqiad, codfw & ulsfo [[phab:T425703|T425703]] * 18:58 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 18:57 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 18:55 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 18:53 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 18:52 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 18:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1100.eqiad.wmnet with OS trixie * 18:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1068.eqiad.wmnet with OS trixie * 18:47 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1102.eqiad.wmnet with OS trixie * 18:47 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 18:46 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1122.eqiad.wmnet with reason: host reimage * 18:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1122.eqiad.wmnet with reason: host reimage * 18:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1068.eqiad.wmnet with reason: host reimage * 18:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1122 * 18:26 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1122 * 18:25 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1122 * 18:25 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1122.eqiad.wmnet 31.48.64.10.in-addr.arpa 1.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:25 bking@cumin2003: START - Cookbook sre.dns.wipe-cache cirrussearch1122.eqiad.wmnet 31.48.64.10.in-addr.arpa 1.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:25 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:25 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1122 - bking@cumin2003" * 18:25 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1122 - bking@cumin2003" * 18:21 rzl@deploy2003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 18:21 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1068.eqiad.wmnet with reason: host reimage * 18:21 rzl@deploy2003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 18:21 rzl@deploy2003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 18:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 18:19 rzl@deploy2003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 18:19 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:18 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1122 * 18:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1122.eqiad.wmnet with OS trixie * 18:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 18:15 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 18:13 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 18:13 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 18:10 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 18:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1068.eqiad.wmnet with OS trixie * 18:01 kamila@deploy2003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 18m 29s) * 18:00 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:55 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:42 kamila@deploy2003: Started scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] * 17:42 kamila@deploy2003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 19m 50s) * 17:42 kamila@deploy2003: Rolling back deployment * 17:35 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:31 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet * 17:18 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet * 17:16 kamila@deploy1003: Unlocked for deployment [MediaWiki]: switching deployment server (duration: 22m 07s) * 17:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 17:11 kamila@dns1005: END - running authdns-update * 17:09 kamila@dns1005: START - running authdns-update * 17:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 17:04 jasmine@cumin2002: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1164.eqiad.wmnet * 17:04 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1164.eqiad.wmnet * 17:04 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1164.eqiad.wmnet * 16:56 kamila@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on releases2003.codfw.wmnet,releases1003.eqiad.wmnet with reason: Deployment server switchover * 16:54 kamila@deploy1003: Locking from deployment [MediaWiki]: switching deployment server * 16:53 kamila@deploy1003: Unlocked for deployment [MediaWiki]: switching deployment server (duration: 04m 02s) * 16:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie * 16:49 kamila@deploy1003: Locking from deployment [MediaWiki]: switching deployment server * 16:46 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1095.eqiad.wmnet with OS trixie * 16:45 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1093.eqiad.wmnet with OS trixie * 16:43 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1164.eqiad.wmnet with OS trixie * 16:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1095.eqiad.wmnet with reason: host reimage * 16:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 16:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 16:23 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1164.eqiad.wmnet with reason: host reimage * 16:18 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on cirrussearch1093.eqiad.wmnet with reason: host reimage * 16:16 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1164.eqiad.wmnet with reason: host reimage * 16:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1095.eqiad.wmnet with reason: host reimage * 16:09 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 16:09 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 16:08 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1093.eqiad.wmnet with reason: host reimage * 15:59 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Pool test * 15:59 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:59 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 15:59 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Pool test * 15:58 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Depool test * 15:58 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:58 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 15:58 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Depool test * 15:57 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1164 * 15:57 jasmine@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1164 * 15:57 jasmine@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1164 * 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1164.eqiad.wmnet 114.48.64.10.in-addr.arpa 4.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:56 jasmine@cumin2002: START - Cookbook sre.dns.wipe-cache wikikube-worker1164.eqiad.wmnet 114.48.64.10.in-addr.arpa 4.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1164 - jasmine@cumin2002" * 15:56 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1164 - jasmine@cumin2002" * 15:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1093.eqiad.wmnet with OS trixie * 15:51 jasmine@cumin2002: START - Cookbook sre.dns.netbox * 15:51 jasmine@cumin2002: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1164 * 15:50 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-worker1164.eqiad.wmnet with OS trixie * 15:50 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1164.eqiad.wmnet * 15:50 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1164.eqiad.wmnet * 15:50 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1164.eqiad.wmnet * 15:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1095.eqiad.wmnet with OS trixie * 15:42 jasmine@cumin2002: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1164.eqiad.wmnet * 15:42 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1164.eqiad.wmnet * 15:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:42 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1164.eqiad.wmnet * 15:42 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1164.eqiad.wmnet * 15:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 15:39 elukey@cumin1003: START - Cookbook sre.hosts.provision for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 15:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1007.eqiad.wmnet with OS trixie * 15:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Pool test * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1007.eqiad.wmnet with reason: host reimage * 15:15 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 15:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet * 15:15 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 15:15 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1007.eqiad.wmnet with reason: host reimage * 15:15 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:14 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Pool test * 15:14 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet * 15:14 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2228: Depool test * 15:14 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db2228: Depool test * 15:10 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 15:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:08 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 15:06 blake@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 15:06 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet * 15:06 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 15:06 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 15:05 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet * 15:05 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 15:04 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:04 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 15:04 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:03 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 15:03 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 15:03 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:03 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T430909|T430909]] * 15:03 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:03 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 15:03 swfrench-wmf: restarted eqsin, codfw confds - [[phab:T430909|T430909]] * 15:03 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test1002.eqiad.wmnet * 15:02 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet * 14:59 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:59 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:55 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1007.eqiad.wmnet with OS trixie * 14:52 swfrench-wmf: restarted ulsfo confds, confirmed now connected to codfw backends except those using wikimedia.org SRV record - [[phab:T430909|T430909]] * 14:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:49 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:41 moritzm: uninstalling dhcpcd-base from trixie hosts which still have it installed [[phab:T414341|T414341]] * 14:40 sukhe: sudo cumin -b1 -s120 "P<nowiki>{</nowiki>lvs2011*<nowiki>}</nowiki> or P<nowiki>{</nowiki>lvs2012*<nowiki>}</nowiki>" "systemctl restart pybal.service" * 14:39 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:39 mvernon@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host thanos-be1007.eqiad.wmnet with OS trixie * 14:37 sukhe: restart pybal on lvs2013 to revert back to conf2004 * 14:35 sukhe: restart pybal on lvs2014 to revert back to conf2004 * 14:34 swfrench-wmf: switched codfw, eqsin, ulsfo etcd client SRV records back to codfw - [[phab:T430909|T430909]] * 14:32 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1002.eqiad.wmnet * 14:32 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet * 14:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1007.eqiad.wmnet with OS trixie * 14:31 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:31 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:31 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:30 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:30 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Pool test * 14:30 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:29 swfrench@dns1004: END - running authdns-update * 14:29 moritzm: installing jackson-core security updates * 14:27 swfrench@dns1004: START - running authdns-update * 14:22 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:22 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:22 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1119.eqiad.wmnet with OS trixie * 14:22 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:21 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:20 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:20 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 14:20 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:19 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 14:19 moritzm: installing librabbitmq security updates * 14:19 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1002.eqiad.wmnet * 14:18 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet * 14:16 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:15 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:15 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:14 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Pool test * 14:14 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox) * 14:14 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet * 14:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 14:08 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 14:07 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet * 14:05 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1118.eqiad.wmnet with OS trixie * 14:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test1001.eqiad.wmnet * 14:00 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-worker@eqiad * 14:00 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 13:59 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 13:58 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 13:57 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1119.eqiad.wmnet with reason: host reimage * 13:54 moritzm: installing libcap2 security updates * 13:53 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1119.eqiad.wmnet with reason: host reimage * 13:52 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet * 13:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) * 13:52 fceratto@cumin1003: START - Cookbook sre.mysql.depool * 13:50 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-worker@eqiad * 13:50 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1051.eqiad.wmnet * 13:50 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1051.eqiad.wmnet * 13:50 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1051.eqiad.wmnet * 13:49 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 13:45 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1081.eqiad.wmnet with OS trixie * 13:41 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1119 * 13:41 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1119 * 13:40 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1119 * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1119.eqiad.wmnet 97.32.64.10.in-addr.arpa 7.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1119.eqiad.wmnet 97.32.64.10.in-addr.arpa 7.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1119 - atsuko@cumin1003" * 13:40 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1119 - atsuko@cumin1003" * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1118.eqiad.wmnet with reason: host reimage * 13:39 moritzm: installing krb5 security updates * 13:37 Lucas_WMDE: UTC afternoon backport+config window done * 13:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1006.eqiad.wmnet with OS trixie * 13:36 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1118.eqiad.wmnet with reason: host reimage * 13:36 atsuko@cumin1003: START - Cookbook sre.dns.netbox * 13:35 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] (duration: 07m 46s) * 13:34 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1119 * 13:34 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1119.eqiad.wmnet with OS trixie * 13:30 sbisson@deploy1003: sbisson: Continuing with deployment * 13:30 moritzm: installing openssh security updates * 13:30 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling restart_daemons on A:wikidough * 13:29 sbisson@deploy1003: sbisson: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:27 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] * 13:26 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1051.eqiad.wmnet with OS trixie * 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1081.eqiad.wmnet with reason: host reimage * 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1118 * 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1118 * 13:22 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] (duration: 12m 12s) * 13:21 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1081.eqiad.wmnet with reason: host reimage * 13:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1006.eqiad.wmnet with reason: host reimage * 13:18 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1118 * 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1118.eqiad.wmnet 90.32.64.10.in-addr.arpa 0.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:18 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1118.eqiad.wmnet 90.32.64.10.in-addr.arpa 0.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1118 - atsuko@cumin1003" * 13:18 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1118 - atsuko@cumin1003" * 13:17 stran@deploy1003: stran: Continuing with deployment * 13:16 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough * 13:15 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:13 atsuko@cumin1003: START - Cookbook sre.dns.netbox * 13:12 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1118 * 13:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1006.eqiad.wmnet with reason: host reimage * 13:12 stran@deploy1003: stran: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:12 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1118.eqiad.wmnet with OS trixie * 13:10 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] * 13:05 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-worker@codfw * 13:05 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 13:05 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1081.eqiad.wmnet with OS trixie * 13:05 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1051.eqiad.wmnet with reason: host reimage * 13:04 moritzm: installing jq security updates * 13:04 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 13:01 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1051.eqiad.wmnet with reason: host reimage * 12:58 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-worker@codfw * 12:52 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 12:50 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host thanos-be1006.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1051 * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1051 * 12:43 moritzm: installing Python 3.11 security updates * 12:43 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1051 * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1051.eqiad.wmnet 46.32.64.10.in-addr.arpa 6.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:43 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1051.eqiad.wmnet 46.32.64.10.in-addr.arpa 6.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1051 - blake@cumin1003" * 12:43 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1051 - blake@cumin1003" * 12:38 blake@cumin1003: START - Cookbook sre.dns.netbox * 12:38 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1051 * 12:38 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1051.eqiad.wmnet with OS trixie * 12:37 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1051.eqiad.wmnet * 12:36 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1051.eqiad.wmnet * 12:36 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1051.eqiad.wmnet * 12:34 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1006.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 12:34 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1006.eqiad.wmnet with OS trixie * 12:27 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 12:27 mvernon@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host thanos-be1006.eqiad.wmnet with OS trixie * 12:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:02 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 12:01 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1006.eqiad.wmnet with OS trixie * 11:43 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 11:38 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1076.eqiad.wmnet with OS trixie * 11:26 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1075.eqiad.wmnet with OS trixie * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2047.codfw.wmnet * 11:19 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2047.codfw.wmnet * 11:19 moritzm: temporarily remove ganeti2031 from codfw cluster [[phab:T430910|T430910]] * 11:08 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1076.eqiad.wmnet with reason: host reimage * 11:08 moritzm: installing Linux 6.1.176 on Bookworm servers * 11:03 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1076.eqiad.wmnet with reason: host reimage * 11:00 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1075.eqiad.wmnet with reason: host reimage * 10:56 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1075.eqiad.wmnet with reason: host reimage * 10:47 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1076.eqiad.wmnet with OS trixie * 10:46 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1074.eqiad.wmnet with OS trixie * 10:45 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1005.eqiad.wmnet with OS trixie * 10:40 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1075.eqiad.wmnet with OS trixie * 10:32 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2031.codfw.wmnet * 10:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1005.eqiad.wmnet with reason: host reimage * 10:25 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1005.eqiad.wmnet with reason: host reimage * 10:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1074.eqiad.wmnet with reason: host reimage * 10:17 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1074.eqiad.wmnet with reason: host reimage * 10:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1005.eqiad.wmnet with OS trixie * 10:12 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet * 10:04 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 10:01 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet * 10:01 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1074.eqiad.wmnet with OS trixie * 10:01 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 09:43 cgoubert@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/aux-k8s-services/redioscope: apply * 09:43 cgoubert@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/aux-k8s-services/redioscope: apply * 09:43 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply * 09:35 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply * 09:34 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 09:34 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 09:33 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 41 days, 15:00:00 on db2252.codfw.wmnet with reason: Test * 09:32 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 09:32 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 09:31 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: codfw rack B3 pool after maintenance * 09:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 09:07 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 09:07 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 09:02 ladsgroup@cumin1003: END (PASS) - Cookbook sre.mysql.sanitarium_restart (exit_code=0) * 08:57 topranks: merge patch to shift eqiad <-> esams traffic onto new 40G circuit * 08:54 hashar@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1004.eqiad.wmnet with OS trixie * 08:50 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 08:50 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitarium_restart (exit_code=99) * 08:50 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 08:45 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool es2051: codfw rack B3 pool after maintenance * 08:44 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:44 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:43 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2007.codfw.wmnet * 08:43 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2007.codfw.wmnet * 08:42 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2031.codfw.wmnet * 08:41 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2031.codfw.wmnet * 08:40 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2031.codfw.wmnet * 08:38 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:38 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:35 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.sanitize-wiki (exit_code=97) Managing sanitization for wikis minwikiquote in section s3 * 08:33 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis minwikiquote in section s3 * 08:32 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Checking sanitization for wikis minwikiquote in section s5 * 08:30 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Checking sanitization for wikis minwikiquote in section s5 * 08:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Managing sanitization for wikis minwikiquote in section s5 * 08:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1004.eqiad.wmnet with reason: host reimage * 08:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1004.eqiad.wmnet with reason: host reimage * 08:23 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:22 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis minwikiquote in section s5 * 08:19 XioNoX: lsw1-b3-codfw> request system reboot - [[phab:T430909|T430909]] * 08:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Checking sanitization for wikis minwikiquote in section s5 * 08:17 hashar@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Checking sanitization for wikis minwikiquote in section s5 * 08:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for codfw rack B3 * 08:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2007.codfw.wmnet * 08:15 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lsw1-b3-codfw,lsw1-b3-codfw IPv6,lsw1-b3-codfw.mgmt with reason: Switch maintenance * 08:15 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2007.codfw.wmnet * 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:07 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:06 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: codfw rack B3 depool for maintenance * 08:05 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool es2051: codfw rack B3 depool for maintenance * 08:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1004.eqiad.wmnet with OS trixie * 08:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1005.eqiad.wmnet with OS trixie * 08:03 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 21 hosts with reason: codfw rack B3 depool for maintenance * 07:56 ayounsi@cumin1003: START - Cookbook sre.network.depool-rack with action 'depool' for codfw rack B3 * 07:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1005.eqiad.wmnet with reason: host reimage * 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1005.eqiad.wmnet with reason: host reimage * 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1005.eqiad.wmnet with OS bookworm * 07:29 moritzm: installing gnutls28 security updates * 07:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1005.eqiad.wmnet with OS trixie * 07:13 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1125.eqiad.wmnet with OS trixie * 07:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1005.eqiad.wmnet with reason: host reimage * 07:07 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aux-k8s-etcd1005.eqiad.wmnet with reason: host reimage * 06:56 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1005.eqiad.wmnet with OS bookworm * 06:54 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1125.eqiad.wmnet with reason: host reimage * 06:52 elukey: upgrade all trixie hosts to pywmflib 3.1 - [[phab:T430552|T430552]] * 06:50 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1125.eqiad.wmnet with reason: host reimage * 06:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 06:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 06:38 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1125.eqiad.wmnet with OS trixie * 05:42 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1107.eqiad.wmnet with OS trixie * 05:35 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1124.eqiad.wmnet with OS trixie * 05:31 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1101.eqiad.wmnet with OS trixie * 05:21 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1107.eqiad.wmnet with reason: host reimage * 05:17 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1124.eqiad.wmnet with reason: host reimage * 05:13 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1107.eqiad.wmnet with reason: host reimage * 05:13 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1101.eqiad.wmnet with reason: host reimage * 05:11 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1124.eqiad.wmnet with reason: host reimage * 05:10 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1101.eqiad.wmnet with reason: host reimage * 04:58 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1124.eqiad.wmnet with OS trixie * 04:56 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1107.eqiad.wmnet with OS trixie * 04:55 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1101.eqiad.wmnet with OS trixie * 02:27 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] (duration: 08m 14s) * 02:22 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 02:21 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 02:19 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] * 01:59 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] (duration: 09m 46s) * 01:55 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 01:51 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 01:49 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] * 01:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1099.eqiad.wmnet with OS trixie * 00:57 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1110.eqiad.wmnet with OS trixie * 00:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1099.eqiad.wmnet with reason: host reimage * 00:41 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1099.eqiad.wmnet with reason: host reimage * 00:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1110.eqiad.wmnet with reason: host reimage * 00:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1110.eqiad.wmnet with reason: host reimage * 00:26 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1099.eqiad.wmnet with OS trixie * 00:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1110.eqiad.wmnet with OS trixie == 2026-07-07 == * 22:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1097.eqiad.wmnet with OS trixie * 22:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1097.eqiad.wmnet with reason: host reimage * 22:24 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1097.eqiad.wmnet with reason: host reimage * 22:14 hashar: Restarting Gerrit on gerrit2002 and gerrit1003 (replicas) * 22:09 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1097.eqiad.wmnet with OS trixie * 22:07 hashar: Restarting Gerrit on gerrit2003 * 21:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 21:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 21:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 21:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 21:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1108.eqiad.wmnet with OS trixie * 20:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1091.eqiad.wmnet with OS trixie * 20:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1090.eqiad.wmnet with OS trixie * 20:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1108.eqiad.wmnet with reason: host reimage * 20:36 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1091.eqiad.wmnet with reason: host reimage * 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1090.eqiad.wmnet with reason: host reimage * 20:33 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1091.eqiad.wmnet with reason: host reimage * 20:30 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1108.eqiad.wmnet with reason: host reimage * 20:30 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1006.eqiad.wmnet * 20:30 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1090.eqiad.wmnet with reason: host reimage * 20:30 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1006.eqiad.wmnet * 20:27 jasmine_: "homer lsw1-c2-eqiad* commit "Added new stacked control plane wikikube-ctrl1006"" * 20:22 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] (duration: 07m 29s) * 20:20 jasmine_: "homer "cr*eqiad*" commit "Added new stacked control plane wikikube-ctrl1006"" * 20:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1091.eqiad.wmnet with OS trixie * 20:17 arlolra@deploy1003: arlolra: Continuing with deployment * 20:16 arlolra@deploy1003: arlolra: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:16 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1090.eqiad.wmnet with OS trixie * 20:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1108.eqiad.wmnet with OS trixie * 20:14 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] * 20:09 cwhite: remove 2026-04 swift log archives from centrallog2002 to free some space * 20:01 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=93) for host cirrussearch1108.eqiad.wmnet with OS trixie * 19:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1108.eqiad.wmnet with OS trixie * 19:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1090.eqiad.wmnet with OS trixie * 19:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1109.eqiad.wmnet with OS trixie * 19:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1092.eqiad.wmnet with OS trixie * 19:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1123.eqiad.wmnet with OS trixie * 19:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1109.eqiad.wmnet with reason: host reimage * 19:19 cdobbins@cumin2002: conftool action : set/pooled=yes; selector: name=dns7002.* * 19:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1092.eqiad.wmnet with reason: host reimage * 19:17 jasmine@dns1004: END - running authdns-update * 19:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1109.eqiad.wmnet with reason: host reimage * 19:15 jasmine@dns1004: START - running authdns-update * 19:15 cdobbins@dns1004: END - running authdns-update * 19:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1123.eqiad.wmnet with reason: host reimage * 19:13 cdobbins@dns1004: START - running authdns-update * 19:12 cdobbins@cumin2002: conftool action : set/pooled=yes; selector: name=dns7002.*,service=authdns-update * 19:11 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1092.eqiad.wmnet with reason: host reimage * 19:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1123.eqiad.wmnet with reason: host reimage * 18:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1123.eqiad.wmnet with OS trixie * 18:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1109.eqiad.wmnet with OS trixie * 18:56 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1092.eqiad.wmnet with OS trixie * 18:52 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 18:49 swfrench@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 18:40 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 18:38 swfrench@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 18:11 swfrench-wmf: restarted eqsin, codfw confds - [[phab:T430909|T430909]] * 18:01 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T430909|T430909]] * 17:59 swfrench-wmf: restarted ulsfo confds, confirmed now connected to eqiad backends - [[phab:T430909|T430909]] * 17:52 sukhe: restart pybal on lvs2011 to switch from conf2004 to conf1008: [[phab:T430909|T430909]] * 17:51 sukhe: restart pybal on lvs2012 to switch from conf2004 to conf1008 [puppet re-enabled there]: [[phab:T430909|T430909]] * 17:46 sukhe: restart pybal on lvs2013 to switch from conf2004 to conf1008: [[phab:T430909|T430909]] * 17:44 swfrench-wmf: switched codfw, eqsin, ulsfo etcd client SRV records to eqiad - [[phab:T430909|T430909]] * 17:43 swfrench@dns1004: END - running authdns-update * 17:40 swfrench@dns1004: START - running authdns-update * 17:40 sukhe: restart pybal on lvs2014 to switch from conf2004 to conf1008: [[phab:T430909|T430909]] * 17:21 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1003.eqiad.wmnet * 17:15 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1003.eqiad.wmnet * 17:14 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1002.eqiad.wmnet * 17:06 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1002.eqiad.wmnet * 17:06 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-low-traffic-codfw' 'systemctl restart pybal.service' # lvs2013, [[phab:T416623|T416623]] * 17:04 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1001.eqiad.wmnet * 17:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1111.eqiad.wmnet with OS trixie * 17:00 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal.service' # lvs2014, [[phab:T416623|T416623]] * 16:58 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1001.eqiad.wmnet * 16:58 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS bookworm * 16:55 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-low-traffic-eqiad' 'systemctl restart pybal.service' # lvs1019, [[phab:T416623|T416623]] * 16:53 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-secondary-eqiad' 'systemctl restart pybal.service' # lvs1020, [[phab:T416623|T416623]] * 16:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1111.eqiad.wmnet with reason: host reimage * 16:40 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1111.eqiad.wmnet with reason: host reimage * 16:38 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.peering (exit_code=99) with action 'configure' for AS: 47794 * 16:35 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 47794 * 16:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1111.eqiad.wmnet with OS trixie * 16:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1006.eqiad.wmnet with OS trixie * 16:06 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 16:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1006.eqiad.wmnet with reason: host reimage * 15:58 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1121.eqiad.wmnet with OS trixie * 15:58 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 15:56 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1006.eqiad.wmnet with reason: host reimage * 15:54 mutante: jenkins down in planned maintenance window * 15:42 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1037.eqiad.wmnet * 15:42 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1037.eqiad.wmnet * 15:42 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1037.eqiad.wmnet * 15:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1006.eqiad.wmnet with OS trixie * 15:34 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1121.eqiad.wmnet with reason: host reimage * 15:33 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS bookworm * 15:33 cdobbins@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host dns7002.wikimedia.org with OS trixie * 15:30 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1121.eqiad.wmnet with reason: host reimage * 15:29 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host clouddumps1001.wikimedia.org * 15:20 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1001.wikimedia.org * 15:18 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1121 * 15:18 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1121 * 15:18 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host clouddumps1002.wikimedia.org * 15:17 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1121 * 15:17 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:17 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply * 15:16 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply * 15:16 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:16 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1121 - atsuko@cumin1003" * 15:16 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1121 - atsuko@cumin1003" * 15:14 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1037.eqiad.wmnet with OS trixie * 15:11 atsuko@cumin1003: START - Cookbook sre.dns.netbox * 15:09 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org * 15:09 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1121 * 15:09 andrew@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host clouddumps1002.wikimedia.org * 15:09 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org * 15:09 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1121.eqiad.wmnet with OS trixie * 15:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1007.eqiad.wmnet with OS trixie * 15:08 andrew@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host clouddumps1002.wikimedia.org * 15:08 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org * 15:05 brennen@deploy1003: Finished deploy [phabricator/deployment@7e02037]: deploy phab1004 for [[phab:T431440|T431440]] (duration: 00m 47s) * 15:04 brennen@deploy1003: Started deploy [phabricator/deployment@7e02037]: deploy phab1004 for [[phab:T431440|T431440]] * 15:03 brennen@deploy1003: Finished deploy [phabricator/deployment@7e02037]: deploy phab2003 for [[phab:T431440|T431440]] (duration: 00m 51s) * 15:03 brennen@deploy1003: Started deploy [phabricator/deployment@7e02037]: deploy phab2003 for [[phab:T431440|T431440]] * 15:00 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71] (thin): Regular analytics weekly train THIN [analytics/refinery@7d8dc71f] (duration: 02m 10s) * 14:58 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71] (thin): Regular analytics weekly train THIN [analytics/refinery@7d8dc71f] * 14:58 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71]: Regular analytics weekly train [analytics/refinery@7d8dc71f] (duration: 04m 14s) * 14:54 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1037.eqiad.wmnet with reason: host reimage * 14:53 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71]: Regular analytics weekly train [analytics/refinery@7d8dc71f] * 14:53 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@7d8dc71f] (duration: 02m 00s) * 14:51 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@7d8dc71f] * 14:51 arnaudb@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on phab2003.codfw.wmnet,phab[1004-1006].eqiad.wmnet with reason: maintenance * 14:51 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1037.eqiad.wmnet with reason: host reimage * 14:50 JavierMonton: Deploying Refinery at {{Gerrit|7d8dc71f}} for change {{Gerrit|1308087}} / [[phab:T431318|T431318]] - update filerevision table sqoop and table * 14:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1007.eqiad.wmnet with reason: host reimage * 14:42 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1007.eqiad.wmnet with reason: host reimage * 14:40 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1083.eqiad.wmnet with OS trixie * 14:37 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:36 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:35 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-master@eqiad * 14:35 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 14:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 14:34 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1037 * 14:34 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1037 * 14:34 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:cleanMentorList.php --wiki=frwiki # [[phab:T427386|T427386]] * 14:34 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 14:34 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308112{{!}}Revert^2 "[Growth] frwiki: Deploy automated mentor list cleaner" (T427386)]] (duration: 06m 47s) * 14:34 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 14:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:33 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:32 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1037 * 14:31 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:31 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:29 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:29 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-master@eqiad * 14:29 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:29 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:29 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:28 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:27 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:27 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1308112{{!}}Revert^2 "[Growth] frwiki: Deploy automated mentor list cleaner" (T427386)]] * 14:26 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1007.eqiad.wmnet with OS trixie * 14:26 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:26 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 14:26 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:25 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:25 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:25 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:cleanMentorList.php --wiki=frwiki # [[phab:T427386|T427386]] * 14:24 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:24 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1037 - blake@cumin1003" * 14:24 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1037 - blake@cumin1003" * 14:20 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1083.eqiad.wmnet with reason: host reimage * 14:19 blake@cumin1003: START - Cookbook sre.dns.netbox * 14:19 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-master@codfw * 14:19 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 14:19 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1037 * 14:18 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1037.eqiad.wmnet with OS trixie * 14:18 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1037.eqiad.wmnet * 14:18 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 14:18 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1037.eqiad.wmnet * 14:18 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1037.eqiad.wmnet * 14:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2007.codfw.wmnet with OS trixie * 14:16 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1083.eqiad.wmnet with reason: host reimage * 14:15 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1036.eqiad.wmnet * 14:15 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1036.eqiad.wmnet * 14:14 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1036.eqiad.wmnet * 14:12 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-master@codfw * 14:11 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1120.eqiad.wmnet with OS trixie * 14:05 moritzm: installing distro-info-data updates from trixie/bookworm point releases * 14:04 fabfur: disable puppet on A:cp-text to selectively apply https://gerrit.wikimedia.org/r/c/operations/puppet/+/1308040 * 14:03 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] (duration: 27m 48s) * 14:00 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 14:00 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1083.eqiad.wmnet with OS trixie * 13:58 urbanecm@deploy1003: urbanecm: Continuing with deployment * 13:58 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:58 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1004.eqiad.wmnet with OS bookworm * 13:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2007.codfw.wmnet with reason: host reimage * 13:57 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 13:53 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1120.eqiad.wmnet with reason: host reimage * 13:50 moritzm: installing Linux 5.10.259 on Bullseye hosts * 13:47 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply * 13:47 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply * 13:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2007.codfw.wmnet with reason: host reimage * 13:46 cgoubert@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/aux-k8s-services/redioscope: apply * 13:46 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1120.eqiad.wmnet with reason: host reimage * 13:46 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:46 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:45 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:44 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:44 cgoubert@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/aux-k8s-services/redioscope: apply * 13:40 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 13:39 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:38 moritzm: installing e2fsprogs updates from Trixie point release * 13:35 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] * 13:33 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1120.eqiad.wmnet with OS trixie * 13:33 topranks: reset cr3-eqsin configuration so traffic uses it again after upgrade * 13:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1088.eqiad.wmnet with OS trixie * 13:32 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie * 13:32 cdobbins@cumin1003: conftool action : set/pooled=no; selector: name=dns7002.* * 13:29 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2007.codfw.wmnet with OS trixie * 13:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1004.eqiad.wmnet with reason: host reimage * 13:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2006.codfw.wmnet with OS trixie * 13:18 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1036.eqiad.wmnet with OS trixie * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aux-k8s-etcd1004.eqiad.wmnet with reason: host reimage * 13:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 13:16 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 13:15 jayme: Istio is being upgraded from 1.24.2 to 1.29.4 on wikikube staging eqiad and codfw - [[phab:T427401|T427401]] * 13:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1087.eqiad.wmnet with OS trixie * 13:14 topranks: reboot cr3-eqsin to install new JunOS and set PIC 0/0/0 to 100G * 13:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1088.eqiad.wmnet with reason: host reimage * 13:13 jmm@dns1004: END - running authdns-update * 13:12 jmm@dns1004: START - running authdns-update * 13:09 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1088.eqiad.wmnet with reason: host reimage * 13:07 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1082.eqiad.wmnet with OS trixie * 13:07 jmm@dns1004: END - running authdns-update * 13:06 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1004.eqiad.wmnet with OS bookworm * 13:05 jmm@dns1004: START - running authdns-update * 13:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2006.codfw.wmnet with reason: host reimage * 12:58 topranks: load updated JunOS on cr3-eqsin [[phab:T429386|T429386]] * 12:58 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1036.eqiad.wmnet with reason: host reimage * 12:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2001.codfw.wmnet * 12:57 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2006.codfw.wmnet with reason: host reimage * 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr1-codfw,cr[2-3]-eqsin,cr3-eqsin IPv6,cr3-eqsin.mgmt with reason: upgrade JunOS cr3-eqsin * 12:56 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lvs[5004-5006].eqsin.wmnet with reason: upgrade JunOS cr3-eqsin * 12:55 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 12:55 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 12:53 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1087.eqiad.wmnet with reason: host reimage * 12:53 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 12:52 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1088.eqiad.wmnet with OS trixie * 12:52 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1002.eqiad.wmnet * 12:52 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:51 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2001.codfw.wmnet * 12:49 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1036.eqiad.wmnet with reason: host reimage * 12:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 12:48 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1087.eqiad.wmnet with reason: host reimage * 12:44 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1082.eqiad.wmnet with reason: host reimage * 12:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1002.eqiad.wmnet * 12:42 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 12:42 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 12:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 12:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 12:39 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:39 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: move dumps-nfs IP to the shared one - filippo@cumin1003" * 12:39 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: move dumps-nfs IP to the shared one - filippo@cumin1003" * 12:39 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2006.codfw.wmnet with OS trixie * 12:38 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1082.eqiad.wmnet with reason: host reimage * 12:36 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:33 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:32 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1036 * 12:32 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1036 * 12:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2005.codfw.wmnet with OS trixie * 12:32 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1087.eqiad.wmnet with OS trixie * 12:30 jmm@dns1004: END - running authdns-update * 12:29 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1036 * 12:29 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1036.eqiad.wmnet 21.32.64.10.in-addr.arpa 1.2.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:29 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1036.eqiad.wmnet 21.32.64.10.in-addr.arpa 1.2.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:29 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:29 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1036 - blake@cumin1003" * 12:29 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1036 - blake@cumin1003" * 12:28 jmm@dns1004: START - running authdns-update * 12:26 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:26 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:23 blake@cumin1003: START - Cookbook sre.dns.netbox * 12:23 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1036 * 12:23 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1036.eqiad.wmnet with OS trixie * 12:22 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1036.eqiad.wmnet * 12:22 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1082.eqiad.wmnet with OS trixie * 12:22 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1036.eqiad.wmnet * 12:22 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1036.eqiad.wmnet * 12:21 marostegui: Restart mariadb@s7 on db1155 to pick up new filters - [[phab:T431124|T431124]] * 12:21 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 21 hosts with reason: restarting for replication filter * 12:20 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:19 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:14 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2005.codfw.wmnet with reason: host reimage * 12:14 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:08 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:07 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:07 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2005.codfw.wmnet with reason: host reimage * 12:06 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:06 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:06 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:05 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:05 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:04 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-master-eqiad * 12:04 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl1002.eqiad.wmnet * 12:04 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl1002.eqiad.wmnet * 12:04 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:04 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:03 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:03 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:03 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:03 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 11:59 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl1002.eqiad.wmnet * 11:59 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl1002.eqiad.wmnet * 11:59 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl1001.eqiad.wmnet * 11:59 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl1001.eqiad.wmnet * 11:56 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl1001.eqiad.wmnet * 11:56 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl1001.eqiad.wmnet * 11:56 klausman@cumin2002: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-master-eqiad * 11:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2005.codfw.wmnet with OS trixie * 11:49 blake@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on wikikube-worker1160.eqiad.wmnet with reason: Verifying matchers for silence * 11:42 topranks: cr3-eqsin, begin traffic drain to reset PIC and upgrade JunOS [[phab:T429386|T429386]] * 11:41 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs[5004-5006].eqsin.wmnet with reason: upgrade JunOS cr3-eqsin * 11:39 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr1-codfw,cr[2-3]-eqsin,cr3-eqsin IPv6,cr3-eqsin.mgmt with reason: upgrade JunOS cr3-eqsin * 11:36 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=thanos-fe2004.codfw.wmnet * 11:35 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1086.eqiad.wmnet with OS trixie * 11:35 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=thanos-fe2004.codfw.wmnet * 11:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1085.eqiad.wmnet with OS trixie * 11:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1086.eqiad.wmnet with reason: host reimage * 11:10 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1085.eqiad.wmnet with reason: host reimage * 11:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2004.codfw.wmnet with OS trixie * 11:03 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1086.eqiad.wmnet with reason: host reimage * 11:02 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1085.eqiad.wmnet with reason: host reimage * 10:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2004.codfw.wmnet with reason: host reimage * 10:48 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:46 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1086.eqiad.wmnet with OS trixie * 10:46 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1085.eqiad.wmnet with OS trixie * 10:44 cgoubert@deploy1003: Finished deploy [restbase/deploy@2fc37d4]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] (duration: 16m 44s) * 10:43 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2004.codfw.wmnet with reason: host reimage * 10:35 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:27 cgoubert@deploy1003: Started deploy [restbase/deploy@2fc37d4]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] * 10:27 cgoubert@deploy1003: Finished deploy [restbase/deploy@8a25036]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] (duration: 00m 45s) * 10:26 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1117.eqiad.wmnet with OS trixie * 10:26 cgoubert@deploy1003: Started deploy [restbase/deploy@8a25036]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] * 10:26 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host thanos-fe2004 * 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host thanos-fe2004 * 10:22 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1116.eqiad.wmnet with OS trixie * 10:21 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host thanos-fe2004 * 10:21 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) thanos-fe2004.codfw.wmnet 157.32.192.10.in-addr.arpa 7.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:20 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache thanos-fe2004.codfw.wmnet 157.32.192.10.in-addr.arpa 7.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:20 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:20 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host thanos-fe2004 - mvernon@cumin2003" * 10:20 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host thanos-fe2004 - mvernon@cumin2003" * 10:15 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2252: Repooling after reboot * 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:15 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 10:15 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2252: Repooling after reboot * 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1153.eqiad.wmnet * 10:14 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1153.eqiad.wmnet * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 10:14 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 10:12 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 10:12 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host thanos-fe2004 * 10:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2004.codfw.wmnet with OS trixie * 10:07 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1117.eqiad.wmnet with reason: host reimage * 10:03 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1116.eqiad.wmnet with reason: host reimage * 09:58 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:58 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1117.eqiad.wmnet with reason: host reimage * 09:57 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1116.eqiad.wmnet with reason: host reimage * 09:49 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 41 days, 15:00:00 on db2252.codfw.wmnet with reason: Security updates * 09:45 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1117.eqiad.wmnet with OS trixie * 09:45 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1116.eqiad.wmnet with OS trixie * 09:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1153: Security updates * 09:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:28 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:28 root@cumin1003: START - Cookbook sre.mysql.depool depool db1153: Security updates * 09:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1016: Security updates * 09:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:21 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:21 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1016: Security updates * 09:14 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:14 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 08:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1016: Security updates * 08:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:56 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:56 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1016: Security updates * 08:50 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:50 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:45 filippo@dns1006: END - running authdns-update * 08:43 filippo@dns1006: START - running authdns-update * 08:42 godog: switch dumps-nfs address to be shared with rsync/http - [[phab:T411248|T411248]] * 08:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1016: Security updates * 08:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:40 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:40 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1016: Security updates * 08:29 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host cirrussearch1111.eqiad.wmnet * 08:29 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:27 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:27 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:25 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1015: Security updates * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:09 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:09 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1015: Security updates * 07:42 Msz2001: Deployed private patch for Suggested Ivestigations * 07:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1015: Security updates * 07:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:41 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:41 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1015: Security updates * 07:40 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 07:11 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fingerprint warnings - oblivian@cumin1003" * 07:11 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fingerprint warnings - oblivian@cumin1003 * 07:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1024: Security updates * 07:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:11 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:11 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1024: Security updates * 07:10 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fingerprint warnings - oblivian@cumin1003 * 07:10 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fingerprint warnings - oblivian@cumin1003" * 06:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host cirrussearch1111.eqiad.wmnet * 06:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 06:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1024: Security updates * 06:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 06:48 root@cumin1003: START - Cookbook sre.mysql.parsercache * 06:48 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1024: Security updates * 06:42 moritzm: install nginx security updates * 06:31 root@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool pc1024: Security updates * 06:21 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1024: Security updates * 06:19 moritzm: installing php8.2 security updates * 06:15 moritzm: installing php8.4 security updates * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.7 (duration: 02m 38s) * 03:40 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] (duration: 37m 04s) * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 51s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-06 == * 23:30 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] (duration: 09m 39s) * 23:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1078.eqiad.wmnet with OS trixie * 23:26 jdlrobson@deploy1003: jdlrobson, bwang: Continuing with deployment * 23:22 jdlrobson@deploy1003: jdlrobson, bwang: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug) * 23:21 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] * 23:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1078.eqiad.wmnet with reason: host reimage * 23:06 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1078.eqiad.wmnet with reason: host reimage * 22:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1078.eqiad.wmnet with OS trixie * 22:29 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on cirrussearch1114.eqiad.wmnet with reason: reimage on hold until restore completes * 22:22 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on cirrussearch[1079,1115].eqiad.wmnet with reason: reimage on hold until restore completes * 21:18 maryum: Deployed security fix for [[phab:T428006|T428006]] * 20:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1079.eqiad.wmnet with OS trixie * 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1077.eqiad.wmnet with OS trixie * 20:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1115.eqiad.wmnet with OS trixie * 20:25 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1079.eqiad.wmnet with reason: host reimage * 20:21 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1079.eqiad.wmnet with reason: host reimage * 20:15 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] (duration: 08m 14s) * 20:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1077.eqiad.wmnet with reason: host reimage * 20:10 krinkle@deploy1003: krinkle, pushpaktiwari: Continuing with deployment * 20:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1115.eqiad.wmnet with reason: host reimage * 20:08 krinkle@deploy1003: krinkle, pushpaktiwari: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1077.eqiad.wmnet with reason: host reimage * 20:06 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] * 20:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1079.eqiad.wmnet with OS trixie * 20:04 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1115.eqiad.wmnet with reason: host reimage * 19:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1077.eqiad.wmnet with OS trixie * 19:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1115.eqiad.wmnet with OS trixie * 19:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 19:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 18:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1114.eqiad.wmnet with OS trixie * 18:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1114.eqiad.wmnet with reason: host reimage * 18:35 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1114.eqiad.wmnet with reason: host reimage * 18:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1112.eqiad.wmnet with OS trixie * 18:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1114.eqiad.wmnet with OS trixie * 18:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1072.eqiad.wmnet with OS trixie * 18:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1112.eqiad.wmnet with reason: host reimage * 18:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1112.eqiad.wmnet with reason: host reimage * 17:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1072.eqiad.wmnet with reason: host reimage * 17:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1112.eqiad.wmnet with OS trixie * 17:55 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1072.eqiad.wmnet with reason: host reimage * 17:39 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1072.eqiad.wmnet with OS trixie * 17:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1071.eqiad.wmnet with OS trixie * 17:18 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1070.eqiad.wmnet with OS trixie * 17:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1084.eqiad.wmnet with OS trixie * 16:54 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1071.eqiad.wmnet with reason: host reimage * 16:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1084.eqiad.wmnet with reason: host reimage * 16:51 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1070.eqiad.wmnet with reason: host reimage * 16:49 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1084.eqiad.wmnet with reason: host reimage * 16:38 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1071.eqiad.wmnet with OS trixie * 16:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1096.eqiad.wmnet with OS trixie * 16:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1070.eqiad.wmnet with OS trixie * 16:33 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1084.eqiad.wmnet with OS trixie * 16:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1089.eqiad.wmnet with OS trixie * 16:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1103.eqiad.wmnet with OS trixie * 16:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1096.eqiad.wmnet with reason: host reimage * 16:14 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1096.eqiad.wmnet with reason: host reimage * 16:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1089.eqiad.wmnet with reason: host reimage * 16:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1103.eqiad.wmnet with reason: host reimage * 16:02 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1003.eqiad.wmnet with OS bookworm * 16:01 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1089.eqiad.wmnet with reason: host reimage * 16:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1103.eqiad.wmnet with reason: host reimage * 15:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1096.eqiad.wmnet with OS trixie * 15:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1080.eqiad.wmnet with OS trixie * 15:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1089.eqiad.wmnet with OS trixie * 15:45 dancy@deploy1003: Installation of scap version "4.272.0" completed for 158 hosts * 15:43 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1103.eqiad.wmnet with OS trixie * 15:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1113.eqiad.wmnet with OS trixie * 15:41 dancy@deploy1003: Installing scap version "4.272.0" for 158 host(s) * 15:40 klausman@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 15:39 klausman@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 15:38 klausman@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 15:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1069.eqiad.wmnet with OS trixie * 15:37 klausman@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 15:36 klausman@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 15:34 klausman@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 15:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1080.eqiad.wmnet with reason: host reimage * 15:27 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1080.eqiad.wmnet with reason: host reimage * 15:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1113.eqiad.wmnet with reason: host reimage * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1069.eqiad.wmnet with reason: host reimage * 15:18 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1113.eqiad.wmnet with reason: host reimage * 15:16 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1069.eqiad.wmnet with reason: host reimage * 15:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:11 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1080.eqiad.wmnet with OS trixie * 15:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1113.eqiad.wmnet with OS trixie * 15:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1003.eqiad.wmnet with reason: host reimage * 14:47 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1003.eqiad.wmnet with OS bookworm * 14:33 elukey: rolled out spicerack on all cumin nodes - [[phab:T429699|T429699]] * 14:32 elukey: upgrade all bookworm hosts to pywmflib 3.1 - [[phab:T430552|T430552]] * 14:14 marostegui: Setup x4 eqiad topology [[phab:T404715|T404715]] * 14:13 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 14:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2230.codfw.wmnet * 14:07 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2230.codfw.wmnet * 13:59 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[2001-2002].codfw.wmnet * 13:51 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 13:45 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.major-upgrade (exit_code=97) * 13:45 cwilliams@cumin1003: dbmaint on s4@codfw [[phab:T429893|T429893]] * 13:45 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 13:42 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-master-codfw * 13:42 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl2002.codfw.wmnet * 13:42 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl2002.codfw.wmnet * 13:38 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl2002.codfw.wmnet * 13:38 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl2002.codfw.wmnet * 13:38 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl2001.codfw.wmnet * 13:38 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl2001.codfw.wmnet * 13:35 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl2001.codfw.wmnet * 13:35 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl2001.codfw.wmnet * 13:35 klausman@cumin2002: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-master-codfw * 12:30 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] (duration: 25m 11s) * 12:24 krinkle@deploy1003: krinkle: Continuing with deployment * 12:10 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2048.codfw.wmnet * 12:09 krinkle@deploy1003: krinkle: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:08 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2048.codfw.wmnet * 12:05 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] * 11:57 moritzm: installing curl security updates * 11:49 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:31 moritzm: installing nano security updates * 11:07 moritzm: failover Ganeti master in codfw to ganeti2032 [[phab:T430909|T430909]] * 11:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:04 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest1005.eqiad.wmnet with OS trixie * 11:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:50 jmm@dns1004: END - running authdns-update * 10:47 jmm@dns1004: START - running authdns-update * 10:47 jmm@dns1004: START - running authdns-update * 10:46 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:44 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest1005.eqiad.wmnet with reason: host reimage * 10:38 elukey@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest1005.eqiad.wmnet with reason: host reimage * 10:31 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:31 marostegui: Setup x4 codfw topology [[phab:T404715|T404715]] * 10:31 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 10:24 elukey: spicerack 13.0.0 deployed on cumin2002 * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 10:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 10:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 10:21 elukey@cumin2002: START - Cookbook sre.hosts.reimage for host sretest1005.eqiad.wmnet with OS trixie * 10:20 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:19 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:17 elukey: uploaded spicerack_13.0.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia * 09:54 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:52 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:20 elukey: upgrade all bullseye hosts to pywmflib 3.1 - [[phab:T430552|T430552]] * 09:10 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1015.eqiad.wmnet,service=s4 * 09:10 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1015.eqiad.wmnet,service=s6 * 09:07 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 08:58 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:56 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 08:56 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 08:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 08:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 08:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin2002.codfw.wmnet * 08:06 godog: remove cloudvirt1046, cloudvirt1062, cloudvirt1074, cloudvirt1075 from maintenance aggregate and put them in network-ovs - [[phab:T424802|T424802]] * 08:00 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin2002.codfw.wmnet * 07:58 hashar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] (duration: 32m 53s) * 07:57 fabfur: repooled cp4038 * 07:57 fabfur@cumin1003: conftool action : set/pooled=yes; selector: name=cp4038.* * 07:53 moritzm: installing pyjwt security updates * 07:47 moritzm: installing openjpeg2 security updates * 07:45 hashar@deploy1003: vadymts1, hashar: Continuing with deployment * 07:43 hashar@deploy1003: vadymts1, hashar: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:38 moritzm: installing python-urllib3 security updates * 07:37 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 07:30 fabfur: depooled cp4038 to investigate on possible maxmind failure * 07:30 fabfur@cumin1003: conftool action : set/pooled=no; selector: name=cp4038.* * 07:30 fabfur@cumin1003: conftool action : set/pooled=yes; selector: name=cp4038.* * 07:29 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 07:25 hashar@deploy1003: Started scap sync-world: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] * 06:13 moritzm: installing Linux 6.12.95 on trixie hosts * 05:20 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s6 * 05:20 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s4 * 05:19 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1015.eqiad.wmnet with reason: cloning * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 08s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-05 == * 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 01m 08s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-04 == * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 58s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-03 == * 17:08 topranks: revert protocol preference changes on cr3-ulsfo after upgrade * 16:53 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on cr2-eqord with reason: upgrade JunOS cr3-ulsfo * 16:53 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on cr4-ulsfo with reason: upgrade JunOS cr3-ulsfo * 16:48 topranks: reboot cr3-ulsfo to upgrade JunOS and reset linecard [[phab:T424839|T424839]] * 15:52 topranks: adjust outbound BGP policies on cr3-ulsfo to drain router of traffic [[phab:T424839|T424839]] * 15:45 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on lvs[4008-4010].ulsfo.wmnet with reason: upgrade JunOS cr3-ulsfo * 15:44 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on asw1-[22-23]-ulsfo,cr3-ulsfo,cr3-ulsfo IPv6,cr3-ulsfo.mgmt with reason: upgrade JunOS cr3-ulsfo * 15:36 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 15:35 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 15:35 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 14:40 cmooney@dns3003: END - running authdns-update * 14:26 cmooney@dns3003: START - running authdns-update * 14:26 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:26 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to ulsfo - cmooney@cumin1003" * 14:19 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to ulsfo - cmooney@cumin1003" * 14:16 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:38 sukhe@dns1004: END - running authdns-update * 13:35 sukhe@dns1004: START - running authdns-update * 13:26 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 13:26 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 13:26 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet * 13:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 13:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 13:16 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host sretest1005.eqiad.wmnet * 13:16 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 13:16 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 13:15 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 13:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:14 moritzm: imported samplicator 1.3.8rc1-1+deb13u1 to trixie-wikimedia/main [[phab:T337208|T337208]] * 13:13 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:07 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 13:07 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 13:02 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:02 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:58 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:57 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:57 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:53 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet * 12:50 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 12:47 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:41 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:40 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:39 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:32 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet * 12:26 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet * 12:23 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2005.wikimedia.org * 12:19 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2005.wikimedia.org * 12:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 12:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup[2004-2007].codfw.wmnet * 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[2004-2007].codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin2003" * 12:15 jynus@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[2004-2007].codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin2003" * 12:09 jynus@cumin2003: START - Cookbook sre.dns.netbox * 11:58 jynus@cumin2003: START - Cookbook sre.hosts.decommission for hosts backup[2004-2007].codfw.wmnet * 10:40 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup[1004-1007].eqiad.wmnet * 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[1004-1007].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 10:01 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[1004-1007].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 09:52 jynus@cumin1003: START - Cookbook sre.dns.netbox * 09:39 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:36 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup[1004-1007].eqiad.wmnet * 09:36 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:25 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 09:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 09:16 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 09:05 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:04 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:00 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:59 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:57 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:55 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:50 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 08:50 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 08:49 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 08:49 atsukoito: depooling cirrussearch in codfw because of regression after upgrade [[phab:T431091|T431091]] * 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts mirror1001.wikimedia.org * 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: mirror1001.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 08:29 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: mirror1001.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 08:18 jmm@cumin2003: START - Cookbook sre.dns.netbox * 08:11 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts mirror1001.wikimedia.org * 06:15 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 18s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-02 == * 22:55 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host contint1003.wikimedia.org with OS trixie * 22:29 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on contint1003.wikimedia.org with reason: host reimage * 22:23 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on contint1003.wikimedia.org with reason: host reimage * 22:05 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host contint1003.wikimedia.org with OS trixie * 22:03 mutante: contint1003 (zuul.wikimedia.org) - reimaging because of [[phab:T430510|T430510]]#12067628 [[phab:T418521|T418521]] * 22:03 dzahn@cumin2002: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on zuul.wikimedia.org with reason: reimage * 21:39 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 18s) * 21:39 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 21:20 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1003.eqiad.wmnet, repooling source-only afterwards * 21:19 sbassett: Deployed security fix for [[phab:T428829|T428829]] * 20:58 cmooney@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Release v0.11.2 update for new Aerleon - cmooney@cumin1003 * 20:55 cmooney@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Release v0.11.2 update for new Aerleon - cmooney@cumin1003 * 20:40 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] (duration: 12m 35s) * 20:36 arlolra@deploy1003: cscott, arlolra: Continuing with deployment * 20:35 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 20s) * 20:35 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 20:33 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host contint2003.wikimedia.org with OS trixie * 20:31 arlolra@deploy1003: cscott, arlolra: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Cha * 20:28 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] * 20:17 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] (duration: 08m 13s) * 20:14 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on contint2003.wikimedia.org with reason: host reimage * 20:13 sbassett@deploy1003: sbassett: Continuing with deployment * 20:11 sbassett@deploy1003: sbassett: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:09 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] * 20:08 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 20:08 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 20:08 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on contint2003.wikimedia.org with reason: host reimage * 20:05 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1003.eqiad.wmnet, repooling source-only afterwards * 19:49 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host contint2003.wikimedia.org with OS trixie * 19:48 mutante: contint2003 - reimaging because of [[phab:T430510|T430510]]#12067628 [[phab:T418521|T418521]] * 18:39 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 18:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 18:13 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2002.codfw.wmnet -> wcqs2003.codfw.wmnet, repooling source-only afterwards * 17:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1003.eqiad.wmnet with OS bookworm * 17:52 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1005.eqiad.wmnet * 17:52 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1005.eqiad.wmnet * 17:51 jasmine@cumin2002: conftool action : set/pooled=yes:weight=10; selector: name=wikikube-ctrl1005.eqiad.wmnet * 17:48 jasmine_: homer "cr*eqiad*" commit "Added new stacked control plane wikikube-ctrl1005" * 17:44 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply * 17:44 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply * 17:31 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] (duration: 09m 33s) * 17:26 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 17:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1003.eqiad.wmnet with reason: host reimage * 17:23 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:21 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] * 17:18 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1003.eqiad.wmnet with reason: host reimage * 17:16 rscout@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply * 17:16 rscout@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply * 17:16 rscout@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply * 17:15 rscout@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply * 17:12 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on wcqs[2002-2003].codfw.wmnet,wcqs1002.eqiad.wmnet with reason: reimaging hosts * 17:08 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 17:08 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 17:08 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 17:07 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 17:05 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 17:05 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 17:03 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "running to make sure all updates are synced - cmooney@cumin1003" * 17:03 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "running to make sure all updates are synced - cmooney@cumin1003" * 17:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs1003 * 17:00 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs1003 * 17:00 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 17:00 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1003.eqiad.wmnet with OS bookworm * 16:58 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Re-running - btullis@cumin1003" * 16:58 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Re-running - btullis@cumin1003" * 16:58 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2002.codfw.wmnet -> wcqs2003.codfw.wmnet, repooling source-only afterwards * 16:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-master1004.eqiad.wmnet with OS bookworm * 16:58 btullis@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 16:57 tappof: bump space for prometheus k8s-aux in eqiad * 16:55 cmooney@dns3003: END - running authdns-update * 16:55 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:55 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to eqsin - cmooney@cumin1003" * 16:55 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to eqsin - cmooney@cumin1003" * 16:53 cmooney@dns3003: START - running authdns-update * 16:52 ryankemper: [ml-serve-eqiad] Cleared out 1302 failed (Evicted) pods: `kubectl -n llm delete pods --field-selector=status.phase=Failed`, freeing calico-kube-controllers from OOM crashloop (evictions were caused by disk pressure) * 16:49 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 16:46 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:39 rzl@dns1004: END - running authdns-update * 16:37 rzl@dns1004: START - running authdns-update * 16:36 rzl@dns1004: START - running authdns-update * 16:35 rzl@deploy1003: Finished scap sync-world: [[phab:T416623|T416623]] (duration: 10m 19s) * 16:34 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 16:33 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-master1004.eqiad.wmnet with reason: host reimage * 16:30 rzl@deploy1003: rzl: Continuing with deployment * 16:28 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-master1004.eqiad.wmnet with reason: host reimage * 16:26 rzl@deploy1003: rzl: [[phab:T416623|T416623]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:25 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 16:25 rzl@deploy1003: Started scap sync-world: [[phab:T416623|T416623]] * 16:25 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 16:24 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 16:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: sync * 16:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: sync * 16:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-master1004.eqiad.wmnet with OS bookworm * 16:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-master1003.eqiad.wmnet with OS bookworm * 16:11 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 16:11 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 16:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Security updates * 16:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 16:08 root@cumin1003: START - Cookbook sre.mysql.parsercache * 16:08 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Security updates * 15:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-master1003.eqiad.wmnet with reason: host reimage * 15:54 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:54 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:54 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:54 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-master1003.eqiad.wmnet with reason: host reimage * 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Security updates * 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:45 root@cumin1003: START - Cookbook sre.mysql.parsercache * 15:45 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Security updates * 15:42 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-master1003.eqiad.wmnet with OS bookworm * 15:24 moritzm: installing busybox updates from bookworm point release * 15:20 moritzm: installing busybox updates from trixie point release * 15:15 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1021: Security updates * 15:15 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:15 root@cumin1003: START - Cookbook sre.mysql.parsercache * 15:15 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1021: Security updates * 15:13 moritzm: installing giflib security updates * 15:08 moritzm: installing Tomcat security updates * 14:57 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 14:56 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 14:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:53 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Unblock taavi - oblivian@cumin1003" * 14:53 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Unblock taavi - oblivian@cumin1003 * 14:53 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1021: Security updates * 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:53 root@cumin1003: START - Cookbook sre.mysql.parsercache * 14:53 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1021: Security updates * 14:53 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Unblock taavi - oblivian@cumin1003 * 14:52 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Unblock taavi - oblivian@cumin1003" * 14:46 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94711 and previous config saved to /var/cache/conftool/dbconfig/20260702-144644-fceratto.json * 14:36 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205', diff saved to https://phabricator.wikimedia.org/P94709 and previous config saved to /var/cache/conftool/dbconfig/20260702-143636-fceratto.json * 14:32 moritzm: installing libdbi-perl security updates * 14:26 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205', diff saved to https://phabricator.wikimedia.org/P94708 and previous config saved to /var/cache/conftool/dbconfig/20260702-142628-fceratto.json * 14:16 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94707 and previous config saved to /var/cache/conftool/dbconfig/20260702-141621-fceratto.json * 14:12 moritzm: installing rsync security updates * 14:11 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox) * 14:10 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94706 and previous config saved to /var/cache/conftool/dbconfig/20260702-140959-fceratto.json * 14:09 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2205.codfw.wmnet with reason: Maintenance * 14:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2205: Repooling after switchover * 14:07 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-test-master1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 14:06 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 14:06 Tran: Deployed patch for [[phab:T427287|T427287]] * 14:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:59 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2205: Repooling after switchover * 13:59 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2205: Repooling after switchover * 13:59 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:55 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2205: Repooling after switchover * 13:55 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2205 [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94704 and previous config saved to /var/cache/conftool/dbconfig/20260702-135505-fceratto.json * 13:54 moritzm: installing sed security updates * 13:53 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:52 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2209 to s3 primary [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94703 and previous config saved to /var/cache/conftool/dbconfig/20260702-135235-fceratto.json * 13:52 federico3: Starting s3 codfw failover from db2205 to db2209 - [[phab:T430912|T430912]] * 13:51 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:51 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 13:48 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:47 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2209 with weight 0 [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94702 and previous config saved to /var/cache/conftool/dbconfig/20260702-134719-fceratto.json * 13:47 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Primary switchover s3 [[phab:T430912|T430912]] * 13:44 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:44 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:44 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:40 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 13:38 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 13:37 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 13:36 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 13:36 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:34 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 13:30 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:29 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:29 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:27 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:26 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:25 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling restart_daemons on A:wikidough * 13:23 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 13:22 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 13:17 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 13:17 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns1004.wikimedia.org * 13:12 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:11 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart (exit_code=97) rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough * 13:11 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=97) rolling restart_daemons on A:wikidough * 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough * 13:09 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] (duration: 07m 20s) * 13:05 aude@deploy1003: jdrewniak, aude: Continuing with deployment * 13:04 aude@deploy1003: jdrewniak, aude: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:02 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] * 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts wdqs-categories1001.eqiad.wmnet * 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: wdqs-categories1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 12:10 jmm@dns1004: END - running authdns-update * 12:07 jmm@dns1004: START - running authdns-update * 11:51 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: wdqs-categories1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 11:44 btullis@cumin1003: START - Cookbook sre.dns.netbox * 11:42 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 11:42 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 11:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet * 11:39 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts wdqs-categories1001.eqiad.wmnet * 11:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet * 11:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet * 11:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet * 11:29 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 11:29 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 10:57 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2214: Repooling * 10:49 jmm@dns1004: END - running authdns-update * 10:47 jmm@dns1004: START - running authdns-update * 10:31 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94698 and previous config saved to /var/cache/conftool/dbconfig/20260702-103146-fceratto.json * 10:21 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213', diff saved to https://phabricator.wikimedia.org/P94696 and previous config saved to /var/cache/conftool/dbconfig/20260702-102137-fceratto.json * 10:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:19 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb1017.eqiad.wmnet * 10:18 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 10:18 fceratto@cumin1003: Removing es1033 from zarcillo [[phab:T408772|T408772]] * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts es1033.eqiad.wmnet * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: es1033.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:14 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: es1033.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:13 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb1017.eqiad.wmnet * 10:12 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2214.codfw.wmnet * 10:12 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2214.codfw.wmnet * 10:12 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2214: Repooling * 10:11 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213', diff saved to https://phabricator.wikimedia.org/P94693 and previous config saved to /var/cache/conftool/dbconfig/20260702-101130-fceratto.json * 10:10 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:10 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:03 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts es1033.eqiad.wmnet * 10:03 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 10:01 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94691 and previous config saved to /var/cache/conftool/dbconfig/20260702-100122-fceratto.json * 09:55 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94690 and previous config saved to /var/cache/conftool/dbconfig/20260702-095529-fceratto.json * 09:55 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2213.codfw.wmnet with reason: Maintenance * 09:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 09:53 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2213: Repooling after switchover * 09:51 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover * 09:44 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2213: Repooling after switchover * 09:39 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover * 09:39 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2213 [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94688 and previous config saved to /var/cache/conftool/dbconfig/20260702-093859-fceratto.json * 09:36 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2192 to s5 primary [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94687 and previous config saved to /var/cache/conftool/dbconfig/20260702-093650-fceratto.json * 09:36 federico3: Starting s5 codfw failover from db2213 to db2192 - [[phab:T430923|T430923]] * 09:30 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94686 and previous config saved to /var/cache/conftool/dbconfig/20260702-093004-fceratto.json * 09:24 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2192 with weight 0 [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94685 and previous config saved to /var/cache/conftool/dbconfig/20260702-092455-fceratto.json * 09:24 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 23 hosts with reason: Primary switchover s5 [[phab:T430923|T430923]] * 09:19 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220', diff saved to https://phabricator.wikimedia.org/P94684 and previous config saved to /var/cache/conftool/dbconfig/20260702-091957-fceratto.json * 09:16 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] (duration: 06m 57s) * 09:13 moritzm: installing libgcrypt20 security updates * 09:12 kharlan@deploy1003: kharlan: Continuing with deployment * 09:11 kharlan@deploy1003: kharlan: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:09 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220', diff saved to https://phabricator.wikimedia.org/P94683 and previous config saved to /var/cache/conftool/dbconfig/20260702-090950-fceratto.json * 09:09 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] * 09:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 09:01 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] (duration: 07m 07s) * 08:59 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94682 and previous config saved to /var/cache/conftool/dbconfig/20260702-085942-fceratto.json * 08:57 kharlan@deploy1003: kharlan: Continuing with deployment * 08:56 kharlan@deploy1003: kharlan: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:54 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] * 08:52 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:52 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:52 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94681 and previous config saved to /var/cache/conftool/dbconfig/20260702-085237-fceratto.json * 08:52 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2220.codfw.wmnet with reason: Maintenance * 08:43 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:40 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 08:25 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] (duration: 11m 44s) * 08:21 cscott@deploy1003: cscott: Continuing with deployment * 08:16 cscott@deploy1003: cscott: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:14 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] * 08:08 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 08:08 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1244: Migration of db1244.eqiad.wmnet completed * 08:02 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:02 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:01 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] (duration: 18m 58s) * 08:01 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:59 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 07:59 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:59 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:59 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2006.wikimedia.org * 07:58 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:57 cscott@deploy1003: cscott: Continuing with deployment * 07:56 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:56 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:56 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:55 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:55 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:55 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:54 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2006.wikimedia.org * 07:54 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:54 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 07:44 cscott@deploy1003: cscott: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:44 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2005.wikimedia.org * 07:44 moritzm: installing node-lodash security updates * 07:42 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] * 07:39 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2005.wikimedia.org * 07:30 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] (duration: 07m 28s) * 07:26 cscott@deploy1003: ssastry, cscott: Continuing with deployment * 07:25 cscott@deploy1003: ssastry, cscott: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:23 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1244: Migration of db1244.eqiad.wmnet completed * 07:22 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] * 07:16 wmde-fisch@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] (duration: 06m 55s) * 07:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1244.eqiad.wmnet with OS trixie * 07:11 wmde-fisch@deploy1003: wmde-fisch: Continuing with deployment * 07:11 wmde-fisch@deploy1003: wmde-fisch: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:09 wmde-fisch@deploy1003: Started scap sync-world: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] * 06:54 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1244.eqiad.wmnet with reason: host reimage * 06:50 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1244.eqiad.wmnet with reason: host reimage * 06:38 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1250.eqiad.wmnet with OS trixie * 06:34 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db1244.eqiad.wmnet with OS trixie * 06:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1244: Upgrading db1244.eqiad.wmnet * 06:25 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1244: Upgrading db1244.eqiad.wmnet * 06:25 cwilliams@cumin1003: dbmaint on s4@eqiad [[phab:T429893|T429893]] * 06:25 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 06:15 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1250.eqiad.wmnet with reason: host reimage * 06:14 cwilliams@dns1006: END - running authdns-update * 06:12 cwilliams@dns1006: START - running authdns-update * 06:11 cwilliams@dns1006: END - running authdns-update * 06:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db1244 [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94676 and previous config saved to /var/cache/conftool/dbconfig/20260702-061059-cwilliams.json * 06:09 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1250.eqiad.wmnet with reason: host reimage * 06:09 cwilliams@dns1006: START - running authdns-update * 06:08 aokoth@cumin1003: END (PASS) - Cookbook sre.vrts.upgrade (exit_code=0) on VRTS host vrts1003.eqiad.wmnet * 06:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db1160 to s4 primary and set section read-write [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94675 and previous config saved to /var/cache/conftool/dbconfig/20260702-060746-cwilliams.json * 06:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Set s4 eqiad as read-only for maintenance - [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94674 and previous config saved to /var/cache/conftool/dbconfig/20260702-060704-cwilliams.json * 06:06 cezmunsta: Starting s4 eqiad failover from db1244 to db1160 - [[phab:T430817|T430817]] * 06:04 aokoth@cumin1003: START - Cookbook sre.vrts.upgrade on VRTS host vrts1003.eqiad.wmnet * 05:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db1160 with weight 0 [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94673 and previous config saved to /var/cache/conftool/dbconfig/20260702-055927-cwilliams.json * 05:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 40 hosts with reason: Primary switchover s4 [[phab:T430817|T430817]] * 05:55 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1250.eqiad.wmnet with OS trixie * 05:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on db1250.eqiad.wmnet with reason: m3 master switchover [[phab:T430158|T430158]] * 05:39 marostegui: Failover m3 (phabricator) from db1250 to db1228 - [[phab:T430158|T430158]] * 05:32 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2234].codfw.wmnet,db[1217,1228,1250].eqiad.wmnet with reason: m3 master switchover [[phab:T430158|T430158]] * 04:45 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] (duration: 09m 08s) * 04:41 tstarling@deploy1003: tstarling, reedy: Continuing with deployment * 04:38 tstarling@deploy1003: tstarling, reedy: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 04:36 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 59s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:16 ryankemper: [[phab:T429844|T429844]] [opensearch] completed `cirrussearch2111` reimage; all codfw search clusters are green, all nodes now report `OpenSearch 2.19.5`, and the temporary chi voting exclusion has been removed * 00:57 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2111.codfw.wmnet with OS trixie * 00:29 ryankemper: [[phab:T429844|T429844]] [opensearch] depooled codfw search-omega/search-psi discovery records to match existing codfw search depool during OpenSearch 2.19 migration * 00:29 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2111.codfw.wmnet with reason: host reimage * 00:29 ryankemper@cumin2002: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 00:29 ryankemper@cumin2002: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 00:22 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2111.codfw.wmnet with reason: host reimage * 00:01 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2111.codfw.wmnet with OS trixie * 00:00 ryankemper: [[phab:T429844|T429844]] [opensearch] chi cluster recovered after stopping `opensearch_1@production-search-codfw` on `cirrussearch2111` == 2026-07-01 == * 23:59 ryankemper: [[phab:T429844|T429844]] [opensearch] stopped `opensearch_1@production-search-codfw` on `cirrussearch2111` after chi cluster-manager election churn following `voting_config_exclusions` POST; hoping this triggers a re-election * 23:52 cscott@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 23:51 cscott@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 23:51 cscott@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 23:50 cscott@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2003.codfw.wmnet with OS bookworm * 22:29 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 22:13 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 22:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2084.codfw.wmnet with OS trixie * 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2003.codfw.wmnet with reason: host reimage * 22:03 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 22:01 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2003.codfw.wmnet with reason: host reimage * 21:50 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 21:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2084.codfw.wmnet with reason: host reimage * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2003 * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2003 * 21:42 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2003 * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2003.codfw.wmnet 45.48.192.10.in-addr.arpa 5.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:42 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2003.codfw.wmnet 45.48.192.10.in-addr.arpa 5.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2003 - bking@cumin2003" * 21:42 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2003 - bking@cumin2003" * 21:36 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2084.codfw.wmnet with reason: host reimage * 21:35 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:34 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2003 * 21:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2003.codfw.wmnet with OS bookworm * 21:19 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2084.codfw.wmnet with OS trixie * 21:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2081.codfw.wmnet with OS trixie * 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2108.codfw.wmnet with OS trixie * 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2081.codfw.wmnet with reason: host reimage * 20:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2081.codfw.wmnet with reason: host reimage * 20:28 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2081.codfw.wmnet with OS trixie * 20:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2108.codfw.wmnet with reason: host reimage * 20:19 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2108.codfw.wmnet with reason: host reimage * 19:59 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2108.codfw.wmnet with OS trixie * 19:46 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2093.codfw.wmnet with OS trixie * 19:44 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 19:44 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jasmine@cumin2002" * 19:43 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jasmine@cumin2002" * 19:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2080.codfw.wmnet with OS trixie * 19:28 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 19:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2093.codfw.wmnet with reason: host reimage * 19:18 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 19:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2093.codfw.wmnet with reason: host reimage * 19:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2080.codfw.wmnet with reason: host reimage * 19:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2080.codfw.wmnet with reason: host reimage * 18:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2093.codfw.wmnet with OS trixie * 18:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2080.codfw.wmnet with OS trixie * 18:27 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 18:18 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] (duration: 09m 15s) * 18:13 jgiannelos@deploy1003: jgiannelos, neriah: Continuing with deployment * 18:11 jgiannelos@deploy1003: jgiannelos, neriah: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:09 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] * 17:40 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 16:58 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 30 hosts * 16:57 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for 30 hosts * 16:52 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2202.codfw.wmnet * 16:52 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2202.codfw.wmnet * 16:51 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt * 16:51 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt * 16:51 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lvs2012.codfw.wmnet * 16:51 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for lvs2012.codfw.wmnet * 16:49 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2076.codfw.wmnet with OS trixie * 16:49 brett: Start pybal on lvs2012 - [[phab:T429861|T429861]] * 16:49 pt1979@cumin1003: END (ERROR) - Cookbook sre.hosts.remove-downtime (exit_code=97) for 59 hosts * 16:48 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for 59 hosts * 16:42 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2061.codfw.wmnet with OS trixie * 16:30 dancy@deploy1003: Installation of scap version "4.271.0" completed for 2 hosts * 16:28 dancy@deploy1003: Installing scap version "4.271.0" for 2 host(s) * 16:23 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2076.codfw.wmnet with reason: host reimage * 16:19 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2061.codfw.wmnet with reason: host reimage * 16:18 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2076.codfw.wmnet with reason: host reimage * 16:18 jasmine@dns1004: END - running authdns-update * 16:16 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host restbase2039.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 16:16 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host restbase2039.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 16:16 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2061.codfw.wmnet with reason: host reimage * 16:15 jasmine@dns1004: START - running authdns-update * 16:14 jasmine@dns1004: END - running authdns-update * 16:12 jasmine@dns1004: START - running authdns-update * 16:07 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2202.codfw.wmnet with reason: maintenance * 16:06 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt with reason: Junos upograde * 16:00 papaul: ongoing maintenance on lsw1-b2-codfw * 16:00 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2076.codfw.wmnet with OS trixie * 15:59 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt * 15:59 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt * 15:57 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2061.codfw.wmnet with OS trixie * 15:55 pt1979@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2042,2046].codfw.wmnet * 15:55 pt1979@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2042,2046].codfw.wmnet * 15:51 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 15:51 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2220: Repooling after switchover * 15:50 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 15:50 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 15:48 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2092.codfw.wmnet with OS trixie * 15:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 15:40 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 15:38 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 15:37 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 15:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply * 15:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply * 15:32 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 15:32 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 15:30 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 15:29 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 15:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 15:25 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:22 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs2012.codfw.wmnet with reason: Rack B2 maintenance - [[phab:T429861|T429861]] * 15:21 brett: Stopping pybal on lvs2012 in preparation for codfw rack b2 maintenance - [[phab:T429861|T429861]] * 15:20 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2092.codfw.wmnet with reason: host reimage * 15:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:12 _joe_: restarted manually alertmanager-irc-relay * 15:12 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2092.codfw.wmnet with reason: host reimage * 15:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:12 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt with reason: Junos upograde * 15:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover * 15:07 pt1979@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2042,2046].codfw.wmnet * 15:06 pt1979@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2042,2046].codfw.wmnet * 15:02 papaul: ongoing maintenance on lsw1-a8-codfw * 14:31 topranks: POWERING DOWN CR1-EQIAD for line card installation [[phab:T426343|T426343]] * 14:31 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] (duration: 08m 57s) * 14:29 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:26 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 14:24 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:22 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] * 14:22 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover * 14:16 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:15 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover * 14:14 topranks: re-enable routing-engine graceful-failover on cr1-eqiad [[phab:T417873|T417873]] * 14:13 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:13 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2220: Repooling after switchover * 14:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:12 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:12 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:11 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:08 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] (duration: 10m 01s) * 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:07 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2220 [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94664 and previous config saved to /var/cache/conftool/dbconfig/20260701-140729-fceratto.json * 14:06 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:06 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:06 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:05 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2159 to s7 primary [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94663 and previous config saved to /var/cache/conftool/dbconfig/20260701-140503-fceratto.json * 14:04 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:04 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 14:04 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 14:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:04 dreamyjazz@deploy1003: anzx, dreamyjazz: Continuing with deployment * 14:04 federico3: Starting s7 codfw failover from db2220 to db2159 - [[phab:T430826|T430826]] * 14:03 jmm@dns1004: END - running authdns-update * 14:03 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:03 topranks: flipping cr1-eqiad active routing-enginer back to RE0 [[phab:T417873|T417873]] * 14:03 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudsw1-c8-eqiad,cloudsw1-d5-eqiad with reason: router upgrades eqiad * 14:01 jmm@dns1004: START - running authdns-update * 14:00 dreamyjazz@deploy1003: anzx, dreamyjazz: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:59 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2159 with weight 0 [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94662 and previous config saved to /var/cache/conftool/dbconfig/20260701-135906-fceratto.json * 13:58 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] * 13:57 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s7 [[phab:T430826|T430826]] * 13:56 topranks: reboot routing-enginer RE0 on cr1-eqiad [[phab:T417873|T417873]] * 13:48 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1006.wikimedia.org * 13:44 atsuko@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cirrussearch2092.codfw.wmnet with OS trixie * 13:43 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1006.wikimedia.org * 13:41 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2092.codfw.wmnet with OS trixie * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1005.wikimedia.org * 13:37 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1005.wikimedia.org * 13:37 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on pfw1-eqiad with reason: router upgrades eqiad * 13:35 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on lvs[1017-1020].eqiad.wmnet with reason: router upgrades eqiad * 13:34 caro@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] (duration: 07m 59s) * 13:30 caro@deploy1003: caro: Continuing with deployment * 13:28 caro@deploy1003: caro: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:27 topranks: route-engine failover cr1-eqiad * 13:26 caro@deploy1003: Started scap sync-world: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] * 13:15 topranks: rebooting routing-engine 1 on cr1-eqiad [[phab:T417873|T417873]] * 13:13 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] (duration: 08m 29s) * 13:13 moritzm: installing qemu security updates * 13:11 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 13:11 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 13:09 jgiannelos@deploy1003: jgiannelos: Continuing with deployment * 13:08 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 13:07 jgiannelos@deploy1003: jgiannelos: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:06 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 13:06 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2214.codfw.wmnet with reason: Maintenance * 13:05 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2214: Repooling after switchover * 13:05 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] * 13:04 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2214: Repooling after switchover * 13:04 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2214 [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94660 and previous config saved to /var/cache/conftool/dbconfig/20260701-130413-fceratto.json * 13:01 moritzm: installing python3.13 security updates * 13:00 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2229 to s6 primary [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94659 and previous config saved to /var/cache/conftool/dbconfig/20260701-125959-fceratto.json * 12:59 federico3: Starting s6 codfw failover from db2214 to db2229 - [[phab:T430814|T430814]] * 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on 13 hosts with reason: router upgrade and line card install * 12:51 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2229 with weight 0 [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94658 and previous config saved to /var/cache/conftool/dbconfig/20260701-125149-fceratto.json * 12:51 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 21 hosts with reason: Primary switchover s6 [[phab:T430814|T430814]] * 12:50 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2189.codfw.wmnet * 12:50 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2189.codfw.wmnet * 12:42 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2100.codfw.wmnet with OS trixie * 12:38 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2083.codfw.wmnet with OS trixie * 12:19 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2083.codfw.wmnet with reason: host reimage * 12:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 12:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2240: Migration of db2240.codfw.wmnet completed * 12:14 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2100.codfw.wmnet with reason: host reimage * 12:09 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2083.codfw.wmnet with reason: host reimage * 12:09 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2100.codfw.wmnet with reason: host reimage * 12:00 topranks: drain traffic on cr1-eqiad to allow for line card install and JunOS upgrade [[phab:T426343|T426343]] * 11:52 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2083.codfw.wmnet with OS trixie * 11:50 cmooney@dns2005: END - running authdns-update * 11:49 cmooney@dns2005: START - running authdns-update * 11:48 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2100.codfw.wmnet with OS trixie * 11:40 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/zotero: apply * 11:40 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/zotero: apply * 11:36 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/zotero: apply * 11:36 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/zotero: apply * 11:31 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2240: Migration of db2240.codfw.wmnet completed * 11:30 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply * 11:28 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply * 11:27 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:27 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:27 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:27 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:27 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:23 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2240.codfw.wmnet with OS trixie * 11:20 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:20 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:17 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:16 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:16 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:15 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2086.codfw.wmnet with OS trixie * 11:14 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2106.codfw.wmnet with OS trixie * 11:14 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:13 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:12 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:09 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2115.codfw.wmnet with OS trixie * 11:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2240.codfw.wmnet with reason: host reimage * 11:00 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2240.codfw.wmnet with reason: host reimage * 10:53 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2106.codfw.wmnet with reason: host reimage * 10:49 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2086.codfw.wmnet with reason: host reimage * 10:44 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2115.codfw.wmnet with reason: host reimage * 10:44 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2240.codfw.wmnet with OS trixie * 10:44 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2086.codfw.wmnet with reason: host reimage * 10:42 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2106.codfw.wmnet with reason: host reimage * 10:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2240: Upgrading db2240.codfw.wmnet * 10:41 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2240: Upgrading db2240.codfw.wmnet * 10:41 cwilliams@cumin1003: dbmaint on s4@codfw [[phab:T429893|T429893]] * 10:40 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 10:39 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2115.codfw.wmnet with reason: host reimage * 10:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2240 [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94653 and previous config saved to /var/cache/conftool/dbconfig/20260701-102658-cwilliams.json * 10:26 moritzm: installing nginx security updates * 10:26 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2086.codfw.wmnet with OS trixie * 10:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2179 to s4 primary [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94652 and previous config saved to /var/cache/conftool/dbconfig/20260701-102356-cwilliams.json * 10:23 cezmunsta: Starting s4 codfw failover from db2240 to db2179 - [[phab:T430127|T430127]] * 10:23 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2106.codfw.wmnet with OS trixie * 10:20 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2115.codfw.wmnet with OS trixie * 10:15 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2179 with weight 0 [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94651 and previous config saved to /var/cache/conftool/dbconfig/20260701-101531-cwilliams.json * 10:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 40 hosts with reason: Primary switchover s4 [[phab:T430127|T430127]] * 09:56 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template (take 2) - oblivian@cumin1003" * 09:56 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template (take 2) - oblivian@cumin1003 * 09:55 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template (take 2) - oblivian@cumin1003 * 09:55 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template (take 2) - oblivian@cumin1003" * 09:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:39 mszwarc@deploy1003: Synchronized private/SuggestedInvestigationsSignals/SuggestedInvestigationsSignal4n.php: Update SI signal 4n (duration: 06m 08s) * 09:21 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 09:21 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 09:14 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 09:14 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 09:02 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 09:02 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 08:54 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 08:38 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 08:38 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 08:36 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 08:21 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 08:21 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] (duration: 36m 11s) * 08:15 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 08:09 mszwarc@deploy1003: mszwarc, abi: Continuing with deployment * 08:03 mszwarc@deploy1003: mszwarc, abi: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:55 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 07:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 07:45 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] * 07:30 aqu@deploy1003: Finished deploy [analytics/refinery@410f205]: Regular analytics weekly train 2nd try [analytics/refinery@410f2050] (duration: 00m 22s) * 07:29 aqu@deploy1003: Started deploy [analytics/refinery@410f205]: Regular analytics weekly train 2nd try [analytics/refinery@410f2050] * 07:28 aqu@deploy1003: Finished deploy [analytics/refinery@410f205] (thin): Regular analytics weekly train THIN [analytics/refinery@410f2050] (duration: 01m 59s) * 07:28 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] (duration: 07m 19s) * 07:26 aqu@deploy1003: Started deploy [analytics/refinery@410f205] (thin): Regular analytics weekly train THIN [analytics/refinery@410f2050] * 07:26 aqu@deploy1003: Finished deploy [analytics/refinery@410f205]: Regular analytics weekly train [analytics/refinery@410f2050] (duration: 04m 32s) * 07:24 mszwarc@deploy1003: wmde-fisch, mszwarc: Continuing with deployment * 07:23 mszwarc@deploy1003: wmde-fisch, mszwarc: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:21 aqu@deploy1003: Started deploy [analytics/refinery@410f205]: Regular analytics weekly train [analytics/refinery@410f2050] * 07:21 aqu@deploy1003: Finished deploy [analytics/refinery@410f205] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@410f2050] (duration: 02m 01s) * 07:20 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] * 07:19 aqu@deploy1003: Started deploy [analytics/refinery@410f205] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@410f2050] * 07:13 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] (duration: 09m 13s) * 07:09 mszwarc@deploy1003: mszwarc, chlod, revi: Continuing with deployment * 07:06 mszwarc@deploy1003: mszwarc, chlod, revi: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:04 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] * 06:55 elukey: upgrade all trixie hosts to pywmflib 3.0 - [[phab:T430552|T430552]] * 06:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:43 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:43 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:42 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:42 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:41 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:41 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:35 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:35 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:34 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:34 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:31 jmm@cumin2003: DONE (PASS) - Cookbook sre.idm.logout (exit_code=0) Logging Niharika29 out of all services on: 2453 hosts * 06:30 oblivian@cumin1003: END (FAIL) - Cookbook sre.deploy.hiddenparma (exit_code=99) Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:30 oblivian@cumin1003: END (FAIL) - Cookbook sre.deploy.python-code (exit_code=99) hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:30 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:30 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:01 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2109.codfw.wmnet with OS trixie * 05:45 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on es1039.eqiad.wmnet with reason: issues * 05:41 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1027.eqiad.wmnet * 05:40 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2068.codfw.wmnet with OS trixie * 05:40 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2109.codfw.wmnet with reason: host reimage * 05:40 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1027.eqiad.wmnet,service=s2 * 05:40 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1027.eqiad.wmnet,service=s7 * 05:36 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2109.codfw.wmnet with reason: host reimage * 05:20 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2068.codfw.wmnet with reason: host reimage * 05:16 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2109.codfw.wmnet with OS trixie * 05:15 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2068.codfw.wmnet with reason: host reimage * 05:09 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2067.codfw.wmnet with OS trixie * 04:56 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2068.codfw.wmnet with OS trixie * 04:49 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2067.codfw.wmnet with reason: host reimage * 04:45 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2067.codfw.wmnet with reason: host reimage * 04:27 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2067.codfw.wmnet with OS trixie * 03:47 slyngshede@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1039.eqiad.wmnet with reason: Hardware crash * 03:21 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2107.codfw.wmnet with OS trixie * 02:59 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2107.codfw.wmnet with reason: host reimage * 02:55 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2085.codfw.wmnet with OS trixie * 02:51 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2072.codfw.wmnet with OS trixie * 02:51 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2107.codfw.wmnet with reason: host reimage * 02:35 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2085.codfw.wmnet with reason: host reimage * 02:31 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2107.codfw.wmnet with OS trixie * 02:30 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2072.codfw.wmnet with reason: host reimage * 02:26 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2085.codfw.wmnet with reason: host reimage * 02:22 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2072.codfw.wmnet with reason: host reimage * 02:09 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2085.codfw.wmnet with OS trixie * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 54s) * 02:03 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2072.codfw.wmnet with OS trixie * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es7 eqiad back to read-write - [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94649 and previous config saved to /var/cache/conftool/dbconfig/20260701-010716-ladsgroup.json * 01:05 ladsgroup@dns1004: END - running authdns-update * 01:05 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depool es1039 [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94648 and previous config saved to /var/cache/conftool/dbconfig/20260701-010551-ladsgroup.json * 01:03 ladsgroup@dns1004: START - running authdns-update * 01:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Promote es1035 to es7 primary [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94647 and previous config saved to /var/cache/conftool/dbconfig/20260701-010002-ladsgroup.json * 00:58 Amir1: Starting es7 eqiad failover from es1039 to es1035 - [[phab:T430765|T430765]] * 00:53 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es1035 with weight 0 [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94646 and previous config saved to /var/cache/conftool/dbconfig/20260701-005329-ladsgroup.json * 00:53 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 9 hosts with reason: Primary switchover es7 [[phab:T430765|T430765]] * 00:42 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es7 eqiad as read-only for maintenance - [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94645 and previous config saved to /var/cache/conftool/dbconfig/20260701-004221-ladsgroup.json * 00:20 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2102.codfw.wmnet with OS trixie * 00:15 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2103.codfw.wmnet with OS trixie * 00:05 dr0ptp4kt: DEPLOYED Refinery at {{Gerrit|4e7a2b32}} for changes: pageview allowlist {{Gerrit|1305158}} (+min.wikiquote) {{Gerrit|1305162}} (+bol.wikipedia), {{Gerrit|1305156}} (+isv.wikipedia); {{Gerrit|1305980}} (pv allowlist -api.wikimedia, sqoop +isvwiki); sqoop {{Gerrit|1295064}} (+globalimagelinks) {{Gerrit|1295069}} (+filerevision) using scap, then deployed onto HDFS (manual copyToLocal required additionally) == Other archives == See [[Server Admin Log/Archives]]. <noinclude> [[Category:SAL]] [[Category:Operations]] </noinclude> ll6zqymedp9v0kswjo595sn4166m56s 2450651 2450650 2026-08-22T16:33:58Z Stashbot 7414 arlolra@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply 2450651 wikitext text/x-wiki == 2026-08-22 == * 16:33 arlolra@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:33 arlolra@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:33 arlolra@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:32 arlolra@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 35s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-21 == * 20:36 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in1001.wikimedia.org with reason: [[phab:T434750|T434750]] * 20:34 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in2001.wikimedia.org with reason: [[phab:T434750|T434750]] * 20:33 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out1001.wikimedia.org with reason: [[phab:T434750|T434750]] * 20:25 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out2001.wikimedia.org with reason: [[phab:T434750|T434750]] * 19:37 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host krb1002.eqiad.wmnet with OS bookworm * 19:00 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 18:59 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 18:51 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 18:51 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 18:35 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:35 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:27 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:27 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:16 bking@cumin2003: START - Cookbook sre.hosts.reimage for host krb1002.eqiad.wmnet with OS bookworm * 17:35 sukhe@dns1004: END - running authdns-update * 17:33 sukhe@dns1004: START - running authdns-update * 17:32 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns5004.wikimedia.org [reason: resolved authdns-update issues] * 17:32 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:32 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: force HEAD to {{Gerrit|be26e30ae101}} - sukhe@cumin1003" * 17:32 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: force HEAD to {{Gerrit|be26e30ae101}} - sukhe@cumin1003" * 17:28 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 17:28 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: service=authdns-update,name=dns5004.wikimedia.org [reason: resolving authdns-update issues] * 17:28 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:28 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: force HEAD to {{Gerrit|be26e30ae101}} - sukhe@cumin1003" * 17:28 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: force HEAD to {{Gerrit|be26e30ae101}} - sukhe@cumin1003" * 17:24 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 17:24 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.netbox (exit_code=97) * 17:23 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 17:16 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 17:12 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 17:10 sukhe@dns1004: END - running authdns-update * 17:08 sukhe@dns1004: START - running authdns-update * 17:08 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=dns5004.wikimedia.org [reason: resolving authdns-update issues] * 17:07 sukhe@dns1004: FAIL - running authdns-update * 17:05 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 17:05 sukhe@dns1004: START - running authdns-update * 17:01 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 16:59 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 16:56 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 16:53 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=dns5004.* [reason: trixie upgrade] * 16:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns5004.wikimedia.org * 16:52 cdobbins@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns5004.wikimedia.org * 16:44 cmooney@dns3003: END - running authdns-update * 16:41 cmooney@dns3003: START - running authdns-update * 16:41 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:41 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on eqsin<->codfw arelion - cmooney@cumin1003" * 16:37 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on eqsin<->codfw arelion - cmooney@cumin1003" * 16:33 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:11 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 16:08 sukhe@dns1004: END - running authdns-update * 16:08 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 16:06 sukhe@dns1004: START - running authdns-update * 16:04 cmooney@dns3003: END - running authdns-update * 16:02 cmooney@dns3003: START - running authdns-update * 16:00 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:00 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on eqord<->codfw arelion - cmooney@cumin1003" * 15:56 cdobbins@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host dns5004.wikimedia.org with OS trixie * 15:55 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on eqord<->codfw arelion - cmooney@cumin1003" * 15:53 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:51 cmooney@cumin1003: END (ERROR) - Cookbook sre.dns.netbox (exit_code=97) * 15:51 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:38 andrew@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudcephosd1042.eqiad.wmnet with OS bookworm * 15:18 andrew@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudcephosd1042.eqiad.wmnet with reason: host reimage * 15:17 dancy@deploy1003: Finished deploy [gerrit/gerrit@2cc11cc]: Deploying https://gerrit.wikimedia.org/r/c/operations/software/gerrit/+/1327669 ([[phab:T434726|T434726]]) (duration: 00m 14s) * 15:17 dancy@deploy1003: Started deploy [gerrit/gerrit@2cc11cc]: Deploying https://gerrit.wikimedia.org/r/c/operations/software/gerrit/+/1327669 ([[phab:T434726|T434726]]) * 15:13 andrew@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudcephosd1042.eqiad.wmnet with reason: host reimage * 15:09 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns5004.wikimedia.org with reason: host reimage * 15:05 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns5004.wikimedia.org with reason: host reimage * 14:53 andrew@cumin2003: START - Cookbook sre.hosts.reimage for host cloudcephosd1042.eqiad.wmnet with OS bookworm * 14:30 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns5004.wikimedia.org with OS trixie * 14:29 cdobbins@cumin1003: conftool action : set/pooled=no; selector: name=dns5004.* [reason: trixie upgrade] * 14:21 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:21 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove entries for cr2-eqord - cmooney@cumin1003" * 14:21 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove entries for cr2-eqord - cmooney@cumin1003" * 14:13 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 14:11 moritzm: imported openjdk 8u504-ga-1~deb12u1 for bookworm-wikimedia (backport of the latest Java 8 security fixes for bookworm) * 13:25 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "sync cr2-eqord router offline - cmooney@cumin1003" * 13:23 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "sync cr2-eqord router offline - cmooney@cumin1003" * 13:14 hashar@deploy1003: Finished deploy [integration/docroot@2d5ff9b]: opensource: add PersonalDashboard docs to MW components - [[phab:T435392|T435392]] (duration: 00m 15s) * 13:14 hashar@deploy1003: Started deploy [integration/docroot@2d5ff9b]: opensource: add PersonalDashboard docs to MW components - [[phab:T435392|T435392]] * 12:16 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:16 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: [[phab:T431682|T431682]] - filippo@cumin1003" * 12:16 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: [[phab:T431682|T431682]] - filippo@cumin1003" * 12:11 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2006.wikimedia.org with OS trixie * 12:00 kevinbazira@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 11:58 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 11:43 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2006.wikimedia.org with reason: host reimage * 11:41 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 11:38 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2006.wikimedia.org with reason: host reimage * 11:20 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2006.wikimedia.org with OS trixie * 11:11 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2005.wikimedia.org with OS trixie * 10:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2005.wikimedia.org with reason: host reimage * 10:53 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2005.wikimedia.org with reason: host reimage * 10:43 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-codfw * 10:43 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2011.codfw.wmnet * 10:43 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2011.codfw.wmnet * 10:40 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2011.codfw.wmnet * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2011.codfw.wmnet * 10:34 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2010.codfw.wmnet * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2010.codfw.wmnet * 10:33 fnegri@cumin1003: END (PASS) - Cookbook sre.wikireplicas.add-wiki (exit_code=0) for database bolwiki ([[phab:T429954|T429954]]) * 10:33 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2005.wikimedia.org with OS trixie * 10:30 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2010.codfw.wmnet * 10:25 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2010.codfw.wmnet * 10:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2009.codfw.wmnet * 10:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2009.codfw.wmnet * 10:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1006.wikimedia.org with OS trixie * 10:18 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2009.codfw.wmnet * 10:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2009.codfw.wmnet * 10:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2008.codfw.wmnet * 10:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2008.codfw.wmnet * 10:06 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2008.codfw.wmnet * 10:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1006.wikimedia.org with reason: host reimage * 10:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2008.codfw.wmnet * 10:01 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2007.codfw.wmnet * 10:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2007.codfw.wmnet * 09:57 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1006.wikimedia.org with reason: host reimage * 09:56 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2007.codfw.wmnet * 09:51 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2007.codfw.wmnet * 09:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 09:51 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 09:46 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2006.codfw.wmnet * 09:46 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1006.wikimedia.org with OS trixie * 09:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1005.wikimedia.org with OS trixie * 09:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2006.codfw.wmnet * 09:41 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2005.codfw.wmnet * 09:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2005.codfw.wmnet * 09:38 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2013.codfw.wmnet * 09:36 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2005.codfw.wmnet * 09:35 fnegri@cumin1003: START - Cookbook sre.wikireplicas.add-wiki for database bolwiki ([[phab:T429954|T429954]]) * 09:35 fnegri@cumin1003: END (PASS) - Cookbook sre.wikireplicas.add-wiki (exit_code=0) for database minwikiquote ([[phab:T429946|T429946]]) * 09:35 fnegri@cumin1003: START - Cookbook sre.wikireplicas.add-wiki for database minwikiquote ([[phab:T429946|T429946]]) * 09:32 blake@cumin1003: START - Cookbook sre.hosts.reboot-single for host rdb2013.codfw.wmnet * 09:30 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2011.codfw.wmnet * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1005.wikimedia.org with reason: host reimage * 09:26 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2005.codfw.wmnet * 09:26 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2004.codfw.wmnet * 09:26 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2004.codfw.wmnet * 09:24 blake@cumin1003: START - Cookbook sre.hosts.reboot-single for host rdb2011.codfw.wmnet * 09:22 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1005.wikimedia.org with reason: host reimage * 09:21 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2004.codfw.wmnet * 09:16 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1015.eqiad.wmnet * 09:16 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2004.codfw.wmnet * 09:16 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2003.codfw.wmnet * 09:16 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2003.codfw.wmnet * 09:13 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on cr[1-2]-eqiad,pfw1-eqiad with reason: upgrade pfw1a-eqiad and pfw1b-eqiad pair * 09:12 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 09:11 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 09:11 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 09:11 blake@cumin1003: START - Cookbook sre.hosts.reboot-single for host rdb1015.eqiad.wmnet * 09:10 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2003.codfw.wmnet * 09:09 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1013.eqiad.wmnet * 09:07 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1005.wikimedia.org with OS trixie * 09:03 blake@cumin1003: START - Cookbook sre.hosts.reboot-single for host rdb1013.eqiad.wmnet * 09:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2003.codfw.wmnet * 09:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2002.codfw.wmnet * 09:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2002.codfw.wmnet * 08:54 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2002.codfw.wmnet * 08:49 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2002.codfw.wmnet * 08:49 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2001.codfw.wmnet * 08:49 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2001.codfw.wmnet * 08:48 jmm@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts netmon2002.wikimedia.org * 08:47 jmm@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts netmon2002.wikimedia.org * 08:44 jmm@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts netmon2002.wikimedia.org * 08:44 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon2002.wikimedia.org * 08:43 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2001.codfw.wmnet * 08:36 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon2002.wikimedia.org * 08:34 jmm@dns1004: END - running authdns-update * 08:33 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2001.codfw.wmnet * 08:33 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-codfw * 08:31 jmm@dns1004: START - running authdns-update * 07:48 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327679{{!}}Block: Disable flaky API test (T435272 T389028)]], [[gerrit:1327678{{!}}API: wfDebugLog for thumberror]] (duration: 15m 34s) * 07:41 krinkle@deploy1003: krinkle: Continuing with deployment * 07:37 krinkle@deploy1003: krinkle: Backport for [[gerrit:1327679{{!}}Block: Disable flaky API test (T435272 T389028)]], [[gerrit:1327678{{!}}API: wfDebugLog for thumberror]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:33 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1327679{{!}}Block: Disable flaky API test (T435272 T389028)]], [[gerrit:1327678{{!}}API: wfDebugLog for thumberror]] * 07:25 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 07:24 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 07:18 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 07:18 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 07:15 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 07:14 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 07:14 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 07:14 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 07:13 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 07:03 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1283: Pool back * 06:42 jmm@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts netmon2002.wikimedia.org * 06:35 moritzm: powercycling netmon2002 * 06:18 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1283: Pool back * 06:17 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1283 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96213 and previous config saved to /var/cache/conftool/dbconfig/20260821-061743-marostegui.json * 04:59 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324963{{!}}Add Produnto to extension-list (T421436)]], [[gerrit:1324964{{!}}Enable Produnto on Beta (T421436)]] (duration: 34m 48s) * 04:45 tstarling@deploy1003: tstarling: Continuing with deployment * 04:44 tstarling@deploy1003: tstarling: Backport for [[gerrit:1324963{{!}}Add Produnto to extension-list (T421436)]], [[gerrit:1324964{{!}}Enable Produnto on Beta (T421436)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 04:24 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1324963{{!}}Add Produnto to extension-list (T421436)]], [[gerrit:1324964{{!}}Enable Produnto on Beta (T421436)]] * 04:21 arlolra@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 04:20 arlolra@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 04:20 arlolra@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 04:20 arlolra@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 41s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-20 == * 23:43 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327652{{!}}RunSingleJob: Add ProfilingContext::init() (T435422)]] (duration: 12m 06s) * 23:38 krinkle@deploy1003: krinkle: Continuing with deployment * 23:33 krinkle@deploy1003: krinkle: Backport for [[gerrit:1327652{{!}}RunSingleJob: Add ProfilingContext::init() (T435422)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:31 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1327652{{!}}RunSingleJob: Add ProfilingContext::init() (T435422)]] * 22:15 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1054.eqiad.wmnet with OS trixie * 22:14 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 22:14 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 21:58 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1054.eqiad.wmnet with reason: host reimage * 21:51 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1054.eqiad.wmnet with reason: host reimage * 21:36 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1054.eqiad.wmnet with OS trixie * 21:36 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:35 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327219{{!}}RunSingleJob: Define MW_ENTRY_POINT for flamegraph sample attribution (T435422)]] (duration: 08m 30s) * 21:31 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:31 krinkle@deploy1003: krinkle: Continuing with deployment * 21:31 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1054 * 21:31 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1054 * 21:29 krinkle@deploy1003: krinkle: Backport for [[gerrit:1327219{{!}}RunSingleJob: Define MW_ENTRY_POINT for flamegraph sample attribution (T435422)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:27 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1327219{{!}}RunSingleJob: Define MW_ENTRY_POINT for flamegraph sample attribution (T435422)]] * 21:17 maryum: Deployed security fix for [[phab:T433020|T433020]] * 20:59 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324752{{!}}InitialiseSettings: Enable 2FA banners on remaining private wikis (T428103)]], [[gerrit:1325920{{!}}Remove sending email to legal team about rejected requests (T374053)]] (duration: 07m 18s) * 20:54 reedy@deploy1003: neriah, reedy: Continuing with deployment * 20:54 reedy@deploy1003: neriah, reedy: Backport for [[gerrit:1324752{{!}}InitialiseSettings: Enable 2FA banners on remaining private wikis (T428103)]], [[gerrit:1325920{{!}}Remove sending email to legal team about rejected requests (T374053)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:51 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324752{{!}}InitialiseSettings: Enable 2FA banners on remaining private wikis (T428103)]], [[gerrit:1325920{{!}}Remove sending email to legal team about rejected requests (T374053)]] * 20:24 reedy@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.15,1.47.0-wmf.16,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/med * 20:23 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324752{{!}}InitialiseSettings: Enable 2FA banners on remaining private wikis (T428103)]], [[gerrit:1325920{{!}}Remove sending email to legal team about rejected requests (T374053)]] * 20:14 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 20:10 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 20:09 cdanis@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "bug fixes & UX fixes - cdanis@cumin1003" * 20:09 cdanis@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: bug fixes & UX fixes - cdanis@cumin1003 * 20:08 cdanis@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: bug fixes & UX fixes - cdanis@cumin1003 * 20:08 cdanis@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "bug fixes & UX fixes - cdanis@cumin1003" * 19:24 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327598{{!}}Make \Omicron non upright (like \Chi) (T434428)]], [[gerrit:1327596{{!}}Render overline of \bar with stretchy=false (T435456)]] (duration: 18m 54s) * 19:20 krinkle@deploy1003: krinkle: Continuing with deployment * 19:07 krinkle@deploy1003: krinkle: Backport for [[gerrit:1327598{{!}}Make \Omicron non upright (like \Chi) (T434428)]], [[gerrit:1327596{{!}}Render overline of \bar with stretchy=false (T435456)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:05 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1327598{{!}}Make \Omicron non upright (like \Chi) (T434428)]], [[gerrit:1327596{{!}}Render overline of \bar with stretchy=false (T435456)]] * 18:51 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327614{{!}}Avoid casting fpxmax to string (T318419)]] (duration: 07m 28s) * 18:50 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:46 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:46 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 18:45 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327614{{!}}Avoid casting fpxmax to string (T318419)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:43 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327614{{!}}Avoid casting fpxmax to string (T318419)]] * 18:37 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:37 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:37 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:36 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:07 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 18:05 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 18:01 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 18:01 sukhe@dns1004: END - running authdns-update * 17:59 sukhe@dns1004: START - running authdns-update * 17:58 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 17:57 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=dns6002.* [reason: depooling for trixie upgrade] * 17:56 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns6002.wikimedia.org * 17:56 cdobbins@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns6002.wikimedia.org * 17:51 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 17:51 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 17:34 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host stat1011.eqiad.wmnet with OS bookworm * 17:31 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns6002.wikimedia.org with OS trixie * 17:30 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 17:30 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 17:30 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 17:30 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 17:29 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 17:29 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 17:29 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 17:29 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:29 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:27 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:24 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327590{{!}}Make sure fpsmax is an int value (T318419)]] (duration: 08m 37s) * 17:20 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 17:17 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327590{{!}}Make sure fpsmax is an int value (T318419)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:16 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 17:15 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327590{{!}}Make sure fpsmax is an int value (T318419)]] * 16:51 swfrench-wmf: disable-puppet on A:cp for ATS Lua change - [[phab:T427666|T427666]] * 16:51 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db2901.codfw.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 16:43 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 16:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on stat1011.eqiad.wmnet with reason: host reimage * 16:39 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns6002.wikimedia.org with reason: host reimage * 16:36 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on stat1011.eqiad.wmnet with reason: host reimage * 16:36 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db2901.codfw.wmnet * 16:34 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns6002.wikimedia.org with reason: host reimage * 16:15 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns6002.wikimedia.org with OS trixie * 16:14 cdobbins@cumin1003: conftool action : set/pooled=no; selector: name=dns6002.* [reason: depooling for trixie upgrade] * 16:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host stat1011 * 16:10 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host stat1011 * 16:09 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host stat1011 * 16:09 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) stat1011.eqiad.wmnet 14.36.64.10.in-addr.arpa 4.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 bking@cumin2003: START - Cookbook sre.dns.wipe-cache stat1011.eqiad.wmnet 14.36.64.10.in-addr.arpa 4.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:09 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host stat1011 - bking@cumin2003" * 16:09 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host stat1011 - bking@cumin2003" * 16:05 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host stat1011 * 16:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host stat1011.eqiad.wmnet with OS bookworm * 16:00 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 15:59 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 15:56 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 15:55 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 15:37 jayme@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:35 jayme@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 15:35 jayme@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:32 jayme@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:32 jayme@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:30 jayme@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 15:30 jayme@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:28 jayme@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:28 jayme@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 15:28 fceratto@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host db1903.eqiad.wmnet * 15:28 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1903.eqiad.wmnet with OS trixie * 15:26 jayme@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 15:26 jayme@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 15:24 jayme@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 15:24 jayme@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 15:21 jayme@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 15:21 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 15:19 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 15:19 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 15:17 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 15:14 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1903.eqiad.wmnet with reason: host reimage * 15:07 fceratto@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1903.eqiad.wmnet with reason: host reimage * 14:54 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db1903.eqiad.wmnet with OS trixie * 14:53 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1903.eqiad.wmnet - fceratto@cumin1003" * 14:53 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1903.eqiad.wmnet - fceratto@cumin1003" * 14:53 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1903.eqiad.wmnet on all recursors * 14:53 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1903.eqiad.wmnet on all recursors * 14:53 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:53 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1903.eqiad.wmnet - fceratto@cumin1003" * 14:53 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1903.eqiad.wmnet - fceratto@cumin1003" * 14:49 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 14:49 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1903.eqiad.wmnet * 14:33 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=cp1100.* * 14:27 topranks: reconfigure eqiad<->codfw bgp settings * 14:22 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:22 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update entries used on new transport backup eqiad codfw - cmooney@cumin1003" * 14:19 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update entries used on new transport backup eqiad codfw - cmooney@cumin1003" * 14:14 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 14:14 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 14:13 moritzm: installing util-linux security updates * 14:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-staging-worker * 14:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2003.codfw.wmnet * 14:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2003.codfw.wmnet * 14:08 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2003.codfw.wmnet * 14:06 moritzm: installing libheif security updates * 13:58 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2003.codfw.wmnet * 13:58 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2002.codfw.wmnet * 13:58 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2002.codfw.wmnet * 13:56 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:56 fnegri@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for clouddb1025.eqiad.wmnet * 13:56 fnegri@cumin1003: START - Cookbook sre.hosts.remove-downtime for clouddb1025.eqiad.wmnet * 13:56 Lucas_WMDE: UTC afternoon backport+config window done * 13:53 moritzm: installing apr-util security updates * 13:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2002.codfw.wmnet * 13:50 fnegri@cumin1003: conftool action : set/weight=100; selector: name=clouddb1025.eqiad.wmnet * 13:49 fnegri@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1025.eqiad.wmnet * 13:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2002.codfw.wmnet * 13:41 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2001.codfw.wmnet * 13:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2001.codfw.wmnet * 13:41 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'. * 13:38 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'. * 13:38 fnegri@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on clouddb1025.eqiad.wmnet with reason: Removing s6 from clouddb1025 * 13:34 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2001.codfw.wmnet * 13:31 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'. * 13:29 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'. * 13:28 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet * 13:26 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host stat1009.eqiad.wmnet with OS bookworm * 13:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2001.codfw.wmnet * 13:24 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-staging-worker * 13:23 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1002.eqiad.wmnet * 13:21 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host stat1010.eqiad.wmnet with OS bookworm * 13:20 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1002.eqiad.wmnet * 13:20 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1001.eqiad.wmnet * 13:17 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1001.eqiad.wmnet * 13:16 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2001.codfw.wmnet * 13:13 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327511{{!}}UIC: Fix page:page instead of page:other in instrumentation]] (duration: 07m 00s) * 13:13 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2001.codfw.wmnet * 13:12 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2002.codfw.wmnet * 13:09 mszwarc@deploy1003: mszwarc: Continuing with deployment * 13:08 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1327511{{!}}UIC: Fix page:page instead of page:other in instrumentation]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2002.codfw.wmnet * 13:07 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2002.codfw.wmnet * 13:06 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1327511{{!}}UIC: Fix page:page instead of page:other in instrumentation]] * 13:04 jmm@dns1004: END - running authdns-update * 13:03 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2002.codfw.wmnet * 13:03 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2001.codfw.wmnet * 13:02 jmm@dns1004: START - running authdns-update * 13:01 cmooney@dns3003: END - running authdns-update * 13:00 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2001.codfw.wmnet * 12:59 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2001.codfw.wmnet * 12:59 cmooney@dns3003: START - running authdns-update * 12:57 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2001.codfw.wmnet * 12:56 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2002.codfw.wmnet * 12:55 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:55 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on drmrs<->eqiad GTT vpls - cmooney@cumin1003" * 12:54 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on drmrs<->eqiad GTT vpls - cmooney@cumin1003" * 12:54 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2002.codfw.wmnet * 12:54 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2003.codfw.wmnet * 12:50 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2003.codfw.wmnet * 12:49 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 12:48 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2001.codfw.wmnet * 12:46 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2001.codfw.wmnet * 12:46 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2002.codfw.wmnet * 12:43 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2002.codfw.wmnet * 12:43 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2003.codfw.wmnet * 12:42 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=1) for new host db1902.eqiad.wmnet * 12:42 fceratto@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host db1902.eqiad.wmnet with OS trixie * 12:41 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2003.codfw.wmnet * 12:40 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1003.eqiad.wmnet * 12:38 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1003.eqiad.wmnet * 12:37 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1002.eqiad.wmnet * 12:35 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1002.eqiad.wmnet * 12:35 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1001.eqiad.wmnet * 12:34 cmooney@dns3003: END - running authdns-update * 12:33 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1001.eqiad.wmnet * 12:32 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on stat1009.eqiad.wmnet with reason: host reimage * 12:31 cmooney@dns3003: START - running authdns-update * 12:31 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:31 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on drmrs<->eqiad cct - cmooney@cumin1003" * 12:28 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on drmrs<->eqiad cct - cmooney@cumin1003" * 12:28 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1902.eqiad.wmnet with reason: host reimage * 12:25 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on stat1009.eqiad.wmnet with reason: host reimage * 12:25 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 12:24 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on stat1010.eqiad.wmnet with reason: host reimage * 12:22 fceratto@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1902.eqiad.wmnet with reason: host reimage * 12:21 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on stat1010.eqiad.wmnet with reason: host reimage * 12:14 elukey: move the Docker Registry's /v2/dev/.* prefix to its dedicated S3 backend - [[phab:T432829|T432829]] * 12:12 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db1902.eqiad.wmnet with OS trixie * 12:09 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1902.eqiad.wmnet - fceratto@cumin1003" * 12:09 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1902.eqiad.wmnet - fceratto@cumin1003" * 12:09 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1902.eqiad.wmnet on all recursors * 12:09 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1902.eqiad.wmnet on all recursors * 12:09 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:08 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1902.eqiad.wmnet - fceratto@cumin1003" * 12:08 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1902.eqiad.wmnet - fceratto@cumin1003" * 12:08 tgr_: [[phab:T413390|T413390]] running CentralAuth:FixRenamedUserGlobalEditCount --wiki=metawiki --since=20250901000000 --until=20260301000000 --fix * 12:04 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1009.eqiad.wmnet with OS bookworm * 12:04 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1010.eqiad.wmnet with OS bookworm * 12:01 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 12:01 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1902.eqiad.wmnet * 12:00 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host stat1010.eqiad.wmnet with OS bookworm * 11:57 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1902.eqiad.wmnet * 11:57 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:57 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1902.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 11:57 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1902.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 11:51 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'. * 11:49 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'. * 11:48 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'. * 11:46 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'. * 11:39 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 11:37 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327518{{!}}Enable thumb.wikimedia.org on cswiki and fawiki (T427465)]] (duration: 10m 40s) * 11:35 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1902.eqiad.wmnet * 11:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1010.eqiad.wmnet with OS bookworm * 11:33 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 11:30 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327518{{!}}Enable thumb.wikimedia.org on cswiki and fawiki (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:26 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327518{{!}}Enable thumb.wikimedia.org on cswiki and fawiki (T427465)]] * 11:24 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host stat1010.eqiad.wmnet with OS bookworm * 11:07 fceratto@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host db1901.eqiad.wmnet * 11:07 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1901.eqiad.wmnet with OS trixie * 10:53 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1901.eqiad.wmnet with reason: host reimage * 10:47 tappof: bump space for prometheus k8s-aux in codfw * 10:47 tappof: bump space for prometheus k8s-dse in eqiad * 10:47 fceratto@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1901.eqiad.wmnet with reason: host reimage * 10:35 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db1901.eqiad.wmnet with OS trixie * 10:32 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:32 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:32 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1901.eqiad.wmnet on all recursors * 10:32 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1901.eqiad.wmnet on all recursors * 10:31 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:31 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:31 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:27 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:27 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1901.eqiad.wmnet * 10:23 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1010.eqiad.wmnet with OS bookworm * 10:20 fceratto@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts db1901.eqiad.wmnet * 10:20 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 10:18 blake@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 10:17 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:16 blake@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 10:13 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1901.eqiad.wmnet * 09:23 jelto@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'. * 09:22 jelto@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'. * 09:22 jelto: update cert-manager to 1.19.6 on wikikube staging-eqiad - [[phab:T427402|T427402]] * 09:20 moritzm: imported squid 7.6-2.1for trixie-wikimedia/main [[phab:T427282|T427282]] * 09:08 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 09:08 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 09:08 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 09:07 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 09:04 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 09:04 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:23 slyngshede@dns1004: END - running authdns-update * 08:21 slyngshede@dns1004: START - running authdns-update * 08:18 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.16 refs [[phab:T430835|T430835]] * 06:27 aokoth@dns1004: END - running authdns-update * 06:25 aokoth@dns1004: START - running authdns-update * 06:22 brennen@deploy1003: Finished deploy [phabricator/deployment@6b9b6ff]: deploy phab1005 for [[phab:T435087|T435087]] (duration: 00m 39s) * 06:21 brennen@deploy1003: Started deploy [phabricator/deployment@6b9b6ff]: deploy phab1005 for [[phab:T435087|T435087]] * 06:20 brennen@deploy1003: Finished deploy [phabricator/deployment@6b9b6ff]: deploy phab1004 for to pick up config values for [[phab:T435087|T435087]] (duration: 01m 46s) * 06:18 brennen@deploy1003: Started deploy [phabricator/deployment@6b9b6ff]: deploy phab1004 for to pick up config values for [[phab:T435087|T435087]] * 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 49s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-19 == * 23:19 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327197{{!}}Enable thumb.wikimedia.org on mediawiki.org (T427465)]] (duration: 10m 50s) * 23:18 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:16 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 23:15 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 23:10 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327197{{!}}Enable thumb.wikimedia.org on mediawiki.org (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:10 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:09 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 23:09 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:09 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 23:08 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327197{{!}}Enable thumb.wikimedia.org on mediawiki.org (T427465)]] * 22:58 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327201{{!}}Enable ReadingLists for all logged in users on test wiki (T435258)]] (duration: 11m 20s) * 22:50 jdlrobson@deploy1003: jdlrobson: Continuing with deployment * 22:49 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1327201{{!}}Enable ReadingLists for all logged in users on test wiki (T435258)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:46 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1327201{{!}}Enable ReadingLists for all logged in users on test wiki (T435258)]] * 22:42 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327162{{!}}Article: Split subjectpageheader by model and disable for wikitext]], [[gerrit:1327169{{!}}Make uppercase greek letters normal (non-italic) font (T434686 T434428)]], [[gerrit:1327176{{!}}Skin: Avoid DB lookup for pagecategorieslink message (T347123)]] (duration: 37m 52s) * 22:29 krinkle@deploy1003: krinkle: Continuing with deployment * 22:25 krinkle@deploy1003: krinkle: Backport for [[gerrit:1327162{{!}}Article: Split subjectpageheader by model and disable for wikitext]], [[gerrit:1327169{{!}}Make uppercase greek letters normal (non-italic) font (T434686 T434428)]], [[gerrit:1327176{{!}}Skin: Avoid DB lookup for pagecategorieslink message (T347123)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:04 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1327162{{!}}Article: Split subjectpageheader by model and disable for wikitext]], [[gerrit:1327169{{!}}Make uppercase greek letters normal (non-italic) font (T434686 T434428)]], [[gerrit:1327176{{!}}Skin: Avoid DB lookup for pagecategorieslink message (T347123)]] * 22:04 krinkle@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: awaiting CI (duration: 03m 06s) * 22:01 krinkle@deploy1003: Locking from deployment [ALL REPOSITORIES]: awaiting CI * 22:00 krinkle@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: awaiting CI (duration: 00m 01s) * 22:00 krinkle@deploy1003: Locking from deployment [ALL REPOSITORIES]: awaiting CI * 21:34 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2001.codfw.wmnet * 21:28 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2001.codfw.wmnet * 21:22 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327178{{!}}AccountRecovery: Notify the email address of the on file of the request (T425799)]] (duration: 47m 02s) * 21:13 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm * 21:09 catrope@deploy1003: catrope: Continuing with deployment * 20:55 catrope@deploy1003: catrope: Backport for [[gerrit:1327178{{!}}AccountRecovery: Notify the email address of the on file of the request (T425799)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:35 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1327178{{!}}AccountRecovery: Notify the email address of the on file of the request (T425799)]] * 20:31 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327128{{!}}Parsoid DataAccess: convert Parsoid fragment markers to/from strip tags (T432547)]] (duration: 07m 30s) * 20:27 catrope@deploy1003: catrope, arlolra: Continuing with deployment * 20:26 catrope@deploy1003: catrope, arlolra: Backport for [[gerrit:1327128{{!}}Parsoid DataAccess: convert Parsoid fragment markers to/from strip tags (T432547)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:24 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1327128{{!}}Parsoid DataAccess: convert Parsoid fragment markers to/from strip tags (T432547)]] * 20:23 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage * 20:17 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage * 20:15 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325878{{!}}[arwiki] Enable restricted user page editing and grant edit permissions (T434878)]] (duration: 08m 46s) * 20:11 catrope@deploy1003: catrope, gergesshamon: Continuing with deployment * 20:08 catrope@deploy1003: catrope, gergesshamon: Backport for [[gerrit:1325878{{!}}[arwiki] Enable restricted user page editing and grant edit permissions (T434878)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:06 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1325878{{!}}[arwiki] Enable restricted user page editing and grant edit permissions (T434878)]] * 19:59 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm * 19:56 eevans@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cassandra-dev2001.codfw.wmnet with OS bookworm * 19:56 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm * 19:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2207.codfw.wmnet with reason: Maintenance * 18:47 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319920{{!}}Allow setting a separate thumbUrl in production (T427465)]], [[gerrit:1327167{{!}}Fix wmgThumbUrl config (T427465)]] (duration: 18m 53s) * 18:43 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 18:30 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1319920{{!}}Allow setting a separate thumbUrl in production (T427465)]], [[gerrit:1327167{{!}}Fix wmgThumbUrl config (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:28 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1319920{{!}}Allow setting a separate thumbUrl in production (T427465)]], [[gerrit:1327167{{!}}Fix wmgThumbUrl config (T427465)]] * 18:26 sukhe@dns1004: END - running authdns-update * 18:24 sukhe@dns1004: START - running authdns-update * 18:09 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1319920{{!}}Allow setting a separate thumbUrl in production (T427465)]] * 18:03 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-eqiad and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 17:56 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-codfw and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 17:50 cmooney@dns3003: END - running authdns-update * 17:42 dancy@deploy1003: Installation of scap version "4.283.0" completed for 3 hosts * 17:41 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling reboot on A:durum and A:durum * 17:40 cmooney@dns3003: START - running authdns-update * 17:40 dancy@deploy1003: Installing scap version "4.283.0" for 3 host(s) * 17:38 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:37 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on GTT VPLS - cmooney@cumin1003" * 17:37 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --verbose --scoreLessThan=0.7 --exceptDatasetChecksums=[[phab:T434319|T434319]]-enwiki-models.txt # [[phab:T434319|T434319]] * 17:32 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on GTT VPLS - cmooney@cumin1003" * 17:29 sbassett: Deployed security fix for [[phab:T435210|T435210]] (wmf.16) * 17:26 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 17:22 sbassett: Deployed security fix for [[phab:T435210|T435210]] (wmf.15) * 17:00 sukhe@dns1004: END - running authdns-update * 16:58 sukhe@dns1004: START - running authdns-update * 16:53 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica-esams and A:liberica * 16:41 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica-esams and A:liberica * 16:41 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-eqiad and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 16:41 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-codfw and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 16:41 cjd91: sudo -i cookbook sre.cdn.roll-upgrade-ats --query 'A:cp-codfw' --task-id [[phab:T434478|T434478]] --reason '9.2.15 upgrade' * 16:41 cjd91: sudo -i cookbook sre.cdn.roll-upgrade-ats --query 'A:cp-eqiad' --task-id [[phab:T434478|T434478]] --reason '9.2.15 upgrade' * 16:40 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and A:durum * 16:28 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326881{{!}}Echo: Start using virtual domains (T380385)]] (duration: 13m 12s) * 16:23 urbanecm@deploy1003: urbanecm: Continuing with deployment * 16:21 urandom: Completed sessionstore Cassandra/JVM upgrade — [[phab:T435154|T435154]] * 16:21 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching sessionstore[2005-2006].codfw.wmnet,sessionstore[1005-1006].eqiad.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 16:19 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1326881{{!}}Echo: Start using virtual domains (T380385)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:15 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:15 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->codfw - cmooney@cumin1003" * 16:14 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1326881{{!}}Echo: Start using virtual domains (T380385)]] * 16:14 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327123{{!}}Revert^2 "Migrate database access to virtual domains" (T435305)]], [[gerrit:1327124{{!}}Pass the mapped domain of virtual-echo-shared to the push NameTableStores (T435305)]] (duration: 07m 42s) * 16:13 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching sessionstore[2005-2006].codfw.wmnet,sessionstore[1005-1006].eqiad.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 16:11 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->codfw - cmooney@cumin1003" * 16:10 urbanecm@deploy1003: urbanecm: Continuing with deployment * 16:08 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1327123{{!}}Revert^2 "Migrate database access to virtual domains" (T435305)]], [[gerrit:1327124{{!}}Pass the mapped domain of virtual-echo-shared to the push NameTableStores (T435305)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:06 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 16:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 16:06 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1327123{{!}}Revert^2 "Migrate database access to virtual domains" (T435305)]], [[gerrit:1327124{{!}}Pass the mapped domain of virtual-echo-shared to the push NameTableStores (T435305)]] * 16:06 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:03 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching sessionstore1004.eqiad.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 16:01 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching sessionstore1004.eqiad.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 15:56 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching sessionstore2004.codfw.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 15:54 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching sessionstore2004.codfw.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 15:53 urandom: beginning sessionstore Cassandra/JVM upgrade — [[phab:T435154|T435154]] * 15:52 urandom: beginning sessionstore Cassandra/JVM upgrade — [[phab:T432944|T432944]] * 15:51 cmooney@dns3003: END - running authdns-update * 15:49 cmooney@dns3003: START - running authdns-update * 15:48 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:48 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->codfw - cmooney@cumin1003" * 15:45 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->codfw - cmooney@cumin1003" * 15:44 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 15:44 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:43 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 15:42 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:42 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:38 cmooney@dns3003: END - running authdns-update * 15:36 cmooney@dns3003: START - running authdns-update * 15:36 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:36 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->eqsin - cmooney@cumin1003" * 15:34 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327138{{!}}Enable redis lock manager everywhere (T366938)]] (duration: 08m 36s) * 15:33 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->eqsin - cmooney@cumin1003" * 15:30 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:29 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 15:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1147.eqiad.wmnet with OS bookworm * 15:28 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327138{{!}}Enable redis lock manager everywhere (T366938)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:25 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327138{{!}}Enable redis lock manager everywhere (T366938)]] * 15:24 jmm@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host krb1002.eqiad.wmnet * 15:19 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2207.codfw.wmnet with reason: Host crashed * 15:17 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327098{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]], [[gerrit:1327101{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]] (duration: 07m 13s) * 15:12 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 15:12 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1327098{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]], [[gerrit:1327101{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:10 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1327098{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]], [[gerrit:1327101{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]] * 15:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2010.codfw.wmnet with OS trixie * 15:05 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 15:05 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1147.eqiad.wmnet with reason: host reimage * 14:59 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 14:59 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:58 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1147.eqiad.wmnet with reason: host reimage * 14:55 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:55 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:49 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 14:48 cmooney@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host durum1001.eqiad.wmnet * 14:46 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:46 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Delete db2902 ipv6 addr - fceratto@cumin1003" * 14:46 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Delete db2902 ipv6 addr - fceratto@cumin1003" * 14:43 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1147.eqiad.wmnet with OS bookworm * 14:42 tgr@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327094{{!}}SpecialMWOAuthListConsumers: Handle newFromMWUser returning null in addNavigationSubtitle (T435167)]] (duration: 19m 25s) * 14:42 cmooney@cumin1003: START - Cookbook sre.hosts.reboot-single for host durum1001.eqiad.wmnet * 14:42 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 14:41 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 14:38 tgr@deploy1003: tgr: Continuing with deployment * 14:36 tgr@deploy1003: tgr: Backport for [[gerrit:1327094{{!}}SpecialMWOAuthListConsumers: Handle newFromMWUser returning null in addNavigationSubtitle (T435167)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:28 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 14:28 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:28 cmooney@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host durum3005.esams.wmnet * 14:25 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 14:23 cmooney@cumin1003: START - Cookbook sre.hosts.reboot-single for host durum3005.esams.wmnet * 14:23 tgr@deploy1003: Started scap sync-world: Backport for [[gerrit:1327094{{!}}SpecialMWOAuthListConsumers: Handle newFromMWUser returning null in addNavigationSubtitle (T435167)]] * 14:18 gengh@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:17 topranks: disable puppet on hosts running BIRD BGP to test merge of patch to systemd healtchcheck service * 14:17 gengh@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:17 gengh@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:16 gengh@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:16 elukey: upgrade spicerack on cumin1003 and cumin2003 * 14:16 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:15 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314962{{!}}static: add new dir bimi/ for BIMI SVG and PEM file (T311685)]] (duration: 10m 00s) * 14:15 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:11 kharlan@deploy1003: kharlan, sukhe: Continuing with deployment * 14:11 gengh@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:09 gengh@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:09 gengh@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:08 kharlan@deploy1003: kharlan, sukhe: Backport for [[gerrit:1314962{{!}}static: add new dir bimi/ for BIMI SVG and PEM file (T311685)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:06 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 14:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:06 gengh@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:05 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1314962{{!}}static: add new dir bimi/ for BIMI SVG and PEM file (T311685)]] * 14:05 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host krb1002.eqiad.wmnet * 14:05 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:04 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:04 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326342{{!}}Revert^2 "wmf-config/ProductionServices: set URL for urldownloader to service record"]] (duration: 07m 40s) * 13:59 kharlan@deploy1003: kharlan, sukhe: Continuing with deployment * 13:58 kharlan@deploy1003: kharlan, sukhe: Backport for [[gerrit:1326342{{!}}Revert^2 "wmf-config/ProductionServices: set URL for urldownloader to service record"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:57 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host stat1008.eqiad.wmnet with OS bookworm * 13:56 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1326342{{!}}Revert^2 "wmf-config/ProductionServices: set URL for urldownloader to service record"]] * 13:56 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 13:54 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325532{{!}}srwiki: Allow bureaucrats to add and remove event-organizer group (T434748)]] (duration: 14m 56s) * 13:54 swfrench@dns1004: END - running authdns-update * 13:52 swfrench@dns1004: START - running authdns-update * 13:51 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:51 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 13:48 kharlan@deploy1003: kharlan, danielyepezgarces: Continuing with deployment * 13:45 swfrench@cumin2003: conftool action : set/pooled=yes; selector: name=wikikube-worker2330.codfw.wmnet * 13:44 swfrench-wmf: finished etcd-main codfw -> eqiad switchover - [[phab:T435103|T435103]] * 13:44 kharlan@deploy1003: kharlan, danielyepezgarces: Backport for [[gerrit:1325532{{!}}srwiki: Allow bureaucrats to add and remove event-organizer group (T434748)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:44 swfrench@cumin2003: conftool action : set/pooled=no; selector: name=wikikube-worker2330.codfw.wmnet * 13:41 swfrench@dns1004: END - running authdns-update * 13:39 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1325532{{!}}srwiki: Allow bureaucrats to add and remove event-organizer group (T434748)]] * 13:39 swfrench@dns1004: START - running authdns-update * 13:37 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327111{{!}}Special:AbuseReview: Add "no further action needed" review action (T435020)]], [[gerrit:1327110{{!}}AbuseReview: Take the review verdict as a REST path parameter (T435020)]] (duration: 31m 43s) * 13:31 swfrench-wmf: starting etcd-main codfw -> eqiad switchover - [[phab:T435103|T435103]] * 13:28 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:28 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 13:25 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=97) rolling reboot on A:durum and A:durum * 13:24 kharlan@deploy1003: kharlan: Continuing with deployment * 13:24 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1147 * 13:24 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1147 * 13:23 kharlan@deploy1003: kharlan: Backport for [[gerrit:1327111{{!}}Special:AbuseReview: Add "no further action needed" review action (T435020)]], [[gerrit:1327110{{!}}AbuseReview: Take the review verdict as a REST path parameter (T435020)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:16 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:16 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 13:12 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:12 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 13:10 cdobbins@cumin1003: END (ERROR) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=97) Rolling upgrade of ATS on A:cp-codfw and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 13:10 cdobbins@cumin1003: END (ERROR) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=97) Rolling upgrade of ATS on A:cp-eqiad and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 13:06 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1327111{{!}}Special:AbuseReview: Add "no further action needed" review action (T435020)]], [[gerrit:1327110{{!}}AbuseReview: Take the review verdict as a REST path parameter (T435020)]] * 13:06 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-eqiad and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 13:05 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-codfw and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 13:05 cjd91: sudo -i cookbook sre.cdn.roll-upgrade-ats --query 'A:cp-eqiad' --task-id [[phab:T434478|T434478]] --reason '9.2.15 upgrade' * 13:03 swfrench@dns1004: END - running authdns-update * 13:01 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:01 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 13:01 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host ncmonitor1001.eqiad.wmnet * 13:00 swfrench@dns1004: START - running authdns-update * 12:59 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 12:59 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 12:59 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 12:59 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 12:57 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and A:durum * 12:53 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on db2902.codfw.wmnet with reason: Cloning * 12:48 cmooney@dns3003: END - running authdns-update * 12:48 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327096{{!}}Switch to redis lock manager on s4 and s8 (T366938)]] (duration: 09m 11s) * 12:46 cmooney@dns3003: START - running authdns-update * 12:45 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:45 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->eqord cct - cmooney@cumin1003" * 12:45 elukey: move the /v2/releng.* prefix on the Docker Registry to its new s3 backend - [[phab:T432829|T432829]] * 12:43 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 12:42 jelto@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 12:42 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->eqord cct - cmooney@cumin1003" * 12:42 jelto@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 12:41 jelto: update cert-manager to 1.19.6 on wikikube staging-codfw - [[phab:T427402|T427402]] * 12:40 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327096{{!}}Switch to redis lock manager on s4 and s8 (T366938)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:38 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327096{{!}}Switch to redis lock manager on s4 and s8 (T366938)]] * 12:38 blake@deploy1003: Finished scap sync-world: non-build deploy for [[phab:T417800|T417800]] (duration: 03m 52s) * 12:36 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 12:35 blake@deploy1003: Started scap sync-world: non-build deploy for [[phab:T417800|T417800]] * 12:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host krb2002.codfw.wmnet * 11:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host krb2002.codfw.wmnet * 11:49 moritzm: installing kerberos security updates * 11:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on stat1008.eqiad.wmnet with reason: host reimage * 11:44 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on stat1008.eqiad.wmnet with reason: host reimage * 11:31 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327084{{!}}Revert "Migrate database access to virtual domains" (T435305)]] (duration: 11m 02s) * 11:29 gkyziridis@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:29 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:29 kart_: Updated MinT to 2026-06-04-131507-production ([[phab:T321316|T321316]]) * 11:28 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/machinetranslation: apply * 11:28 gkyziridis@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:26 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 11:24 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:24 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:23 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:23 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:23 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/machinetranslation: apply * 11:22 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327084{{!}}Revert "Migrate database access to virtual domains" (T435305)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:21 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/machinetranslation: apply * 11:21 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:21 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:20 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327084{{!}}Revert "Migrate database access to virtual domains" (T435305)]] * 11:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1008.eqiad.wmnet with OS bookworm * 11:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet * 11:17 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/machinetranslation: apply * 11:13 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/machinetranslation: apply * 11:12 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:12 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:10 kartik@deploy1003: helmfile [staging] START helmfile.d/services/machinetranslation: apply * 11:08 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-cron: apply * 11:08 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/mw-cron: apply * 11:08 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply * 11:08 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply * 11:07 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet * 11:07 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host stat1008.eqiad.wmnet with OS bookworm * 11:06 moritzm: upgrading the new trixie URL downloaders to Squid 7.6 [[phab:T427282|T427282]] * 11:01 gkyziridis@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin2002.codfw.wmnet * 10:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin2002.codfw.wmnet * 10:45 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1169.eqiad.wmnet with OS bookworm * 10:42 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1185.eqiad.wmnet with OS bookworm * 10:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1169.eqiad.wmnet with reason: host reimage * 10:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1185.eqiad.wmnet with reason: host reimage * 10:14 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1169.eqiad.wmnet with reason: host reimage * 10:14 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1185.eqiad.wmnet with reason: host reimage * 10:11 jmm@cumin2003: END (PASS) - Cookbook sre.netbox.restart-reboot (exit_code=0) rolling reboot on A:netbox * 10:06 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1008.eqiad.wmnet with OS bookworm * 10:04 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 10:04 mpostoronca@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321579{{!}}Register the mediawiki.wikimedia_antiabuse.content_policy_score stream (T432848)]] (duration: 08m 53s) * 10:03 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 10:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 10:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 10:00 mpostoronca@deploy1003: mpostoronca: Continuing with deployment * 09:59 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1185.eqiad.wmnet with OS bookworm * 09:59 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1169.eqiad.wmnet with OS bookworm * 09:59 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.convert-disks (exit_code=0) for host ms-be1065 * 09:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1065.eqiad.wmnet with OS trixie * 09:58 mpostoronca@deploy1003: mpostoronca: Backport for [[gerrit:1321579{{!}}Register the mediawiki.wikimedia_antiabuse.content_policy_score stream (T432848)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:55 jmm@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netbox.discovery.wmnet. on all recursors * 09:55 mpostoronca@deploy1003: Started scap sync-world: Backport for [[gerrit:1321579{{!}}Register the mediawiki.wikimedia_antiabuse.content_policy_score stream (T432848)]] * 09:55 jmm@cumin2003: START - Cookbook sre.dns.wipe-cache netbox.discovery.wmnet. on all recursors * 09:52 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw2001.wikimedia.org with OS trixie * 09:51 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 09:51 jmm@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netbox.discovery.wmnet. on all recursors * 09:51 jmm@cumin2003: START - Cookbook sre.dns.wipe-cache netbox.discovery.wmnet. on all recursors * 09:51 jmm@cumin2003: START - Cookbook sre.netbox.restart-reboot rolling reboot on A:netbox * 09:46 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 09:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1056.eqiad.wmnet with OS trixie * 09:44 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 09:39 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 09:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 09:36 topranks: make HE transport circuits from magru live * 09:36 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 09:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb1003.eqiad.wmnet * 09:33 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage * 09:31 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb1003.eqiad.wmnet * 09:30 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 09:28 arnaudb@dns1006: END - running authdns-update * 09:27 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage * 09:27 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb2003.codfw.wmnet * 09:26 arnaudb@dns1006: START - running authdns-update * 09:24 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1056.eqiad.wmnet with reason: host reimage * 09:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb2003.codfw.wmnet * 09:20 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.convert-disks (exit_code=0) for host ms-be1068 * 09:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1068.eqiad.wmnet with OS trixie * 09:20 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 09:19 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "cloudvirt1057 - filippo@cumin1003" * 09:19 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "cloudvirt1057 - filippo@cumin1003" * 09:18 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1056.eqiad.wmnet with reason: host reimage * 09:18 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1057.eqiad.wmnet with OS trixie * 09:18 filippo@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 09:18 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 09:15 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 09:14 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 09:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host irc1003.wikimedia.org * 09:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:13 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1065.eqiad.wmnet with OS trixie * 09:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 09:09 moritzm: installing Postgresql security updates * 09:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host irc1003.wikimedia.org * 09:07 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 09:07 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:07 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1173.eqiad.wmnet with OS bookworm * 09:06 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw2001.wikimedia.org with OS trixie * 09:03 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.convert-disks (exit_code=0) for host ms-be1064 * 09:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1064.eqiad.wmnet with OS trixie * 09:03 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1056.eqiad.wmnet with OS trixie * 09:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1057.eqiad.wmnet with reason: host reimage * 09:01 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 08:58 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 08:56 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1057.eqiad.wmnet with reason: host reimage * 08:54 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1208.eqiad.wmnet with OS bookworm * 08:53 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2902.codfw.wmnet with OS trixie * 08:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 08:50 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1174.eqiad.wmnet with OS bookworm * 08:46 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1172.eqiad.wmnet with OS bookworm * 08:45 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1173.eqiad.wmnet with reason: host reimage * 08:41 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 08:40 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1057.eqiad.wmnet with OS trixie * 08:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1057.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 08:39 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1207.eqiad.wmnet with OS bookworm * 08:38 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2902.codfw.wmnet with reason: host reimage * 08:37 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 08:34 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1068.eqiad.wmnet with OS trixie * 08:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1208.eqiad.wmnet with reason: host reimage * 08:31 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1057.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 08:29 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 08:28 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1222.eqiad.wmnet onto db1276.eqiad.wmnet * 08:28 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1222: Pool db1222.eqiad.wmnet in after cloning * 08:28 fceratto@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2902.codfw.wmnet with reason: host reimage * 08:27 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1055.eqiad.wmnet with OS trixie * 08:27 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 08:26 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1174.eqiad.wmnet with reason: host reimage * 08:25 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 08:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-misc2002.codfw.wmnet * 08:23 topranks: reboot pfw1-codfw firewall pair to upgrade JunOS [[phab:T434865|T434865]] * 08:22 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1172.eqiad.wmnet with reason: host reimage * 08:20 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1064.eqiad.wmnet with OS trixie * 08:20 mvernon@cumin2003: START - Cookbook sre.swift.convert-disks for host ms-be1065 * 08:18 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1207.eqiad.wmnet with reason: host reimage * 08:17 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1174.eqiad.wmnet with reason: host reimage * 08:17 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1173.eqiad.wmnet with reason: host reimage * 08:17 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1172.eqiad.wmnet with reason: host reimage * 08:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host mc-misc2002.codfw.wmnet * 08:15 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1208.eqiad.wmnet with reason: host reimage * 08:15 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1207.eqiad.wmnet with reason: host reimage * 08:14 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db2902.codfw.wmnet with OS trixie * 08:14 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.16 refs [[phab:T430835|T430835]] * 08:13 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db2902.codfw.wmnet * 08:13 fceratto@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host db2902.codfw.wmnet with OS trixie * 08:10 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr[1-2]-codfw with reason: upgrade pfw1a-codfw and pfw1b-codfw pair * 08:09 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1055.eqiad.wmnet with reason: host reimage * 08:07 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on pfw1-codfw with reason: upgrade pfw1a-codfw and pfw1b-codfw pair * 08:03 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1055.eqiad.wmnet with reason: host reimage * 08:02 arnaudb@dns1006: END - running authdns-update * 08:02 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1208.eqiad.wmnet with OS bookworm * 08:02 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1207.eqiad.wmnet with OS bookworm * 08:01 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1174.eqiad.wmnet with OS bookworm * 08:01 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1173.eqiad.wmnet with OS bookworm * 08:01 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1172.eqiad.wmnet with OS bookworm * 07:59 arnaudb@dns1006: START - running authdns-update * 07:58 arnaudb@dns1006: START - running authdns-update * 07:48 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1055.eqiad.wmnet with OS trixie * 07:45 moritzm: extend the disk of ldap-rw2001 by 80G [[phab:T331699|T331699]] * 07:42 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1222: Pool db1222.eqiad.wmnet in after cloning * 07:36 mvernon@cumin2003: START - Cookbook sre.swift.convert-disks for host ms-be1068 * 07:35 mvernon@cumin2003: START - Cookbook sre.swift.convert-disks for host ms-be1064 * 07:17 moritzm: installing imagemagick security updates * 07:14 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1277: Pool back * 07:14 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon1003.wikimedia.org * 07:07 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon1003.wikimedia.org * 07:03 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1280: Pool back * 07:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon2002.wikimedia.org * 06:55 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon2002.wikimedia.org * 06:54 moritzm: installing php8.2 security updates * 06:51 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1284: Pool back * 06:49 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1222: Depool db1222.eqiad.wmnet to then clone it to db1276.eqiad.wmnet - marostegui@cumin1003 * 06:49 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1222: Depool db1222.eqiad.wmnet to then clone it to db1276.eqiad.wmnet - marostegui@cumin1003 * 06:49 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1222.eqiad.wmnet onto db1276.eqiad.wmnet * 06:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd1005.eqiad.wmnet * 06:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd1005.eqiad.wmnet * 06:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd1004.eqiad.wmnet * 06:36 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2209: db2209 repool * 06:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd1004.eqiad.wmnet * 06:32 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast3007.wikimedia.org * 06:29 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1277: Pool back * 06:28 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1277 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96190 and previous config saved to /var/cache/conftool/dbconfig/20260819-062815-marostegui.json * 06:26 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast3007.wikimedia.org * 06:22 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul1001.eqiad.wmnet * 06:18 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul1001.eqiad.wmnet * 06:18 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul1003.eqiad.wmnet * 06:18 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1280: Pool back * 06:17 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1284 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96186 and previous config saved to /var/cache/conftool/dbconfig/20260819-061743-marostegui.json * 06:14 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul1003.eqiad.wmnet * 06:14 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul1002.eqiad.wmnet * 06:10 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul1002.eqiad.wmnet * 06:10 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2048.codfw.wmnet * 06:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2048.codfw.wmnet * 06:06 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1284: Pool back * 06:06 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1284 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96184 and previous config saved to /var/cache/conftool/dbconfig/20260819-060621-marostegui.json * 06:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2048.codfw.wmnet * 05:59 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2048.codfw.wmnet * 05:51 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2209: db2209 repool * 03:16 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1186.eqiad.wmnet with OS bookworm * 02:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1186.eqiad.wmnet with reason: host reimage * 02:46 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1186.eqiad.wmnet with reason: host reimage * 02:46 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2207 [[phab:T435270|T435270]]', diff saved to https://phabricator.wikimedia.org/P96181 and previous config saved to /var/cache/conftool/dbconfig/20260819-024627-marostegui.json * 02:44 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2204 to s2 primary [[phab:T435270|T435270]]', diff saved to https://phabricator.wikimedia.org/P96180 and previous config saved to /var/cache/conftool/dbconfig/20260819-024403-marostegui.json * 02:43 marostegui: Starting s2 codfw failover from db2207 to db2204 - [[phab:T435270|T435270]] * 02:39 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2204 with weight 0 [[phab:T435270|T435270]]', diff saved to https://phabricator.wikimedia.org/P96179 and previous config saved to /var/cache/conftool/dbconfig/20260819-023951-marostegui.json * 02:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s2 [[phab:T435270|T435270]] * 02:32 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1186.eqiad.wmnet with OS bookworm * 02:29 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-worker1186.eqiad.wmnet with OS bookworm * 02:18 denisse@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2207: Depooling replica * 02:18 denisse@cumin1003: START - Cookbook sre.mysql.depool depool db2207: Depooling replica * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 48s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-18 == * 23:55 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1264845{{!}}Remove unused/redundant wgMFNoindexPages=true setting (T255458)]] (duration: 09m 42s) * 23:51 krinkle@deploy1003: krinkle: Continuing with deployment * 23:48 krinkle@deploy1003: krinkle: Backport for [[gerrit:1264845{{!}}Remove unused/redundant wgMFNoindexPages=true setting (T255458)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:45 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1264845{{!}}Remove unused/redundant wgMFNoindexPages=true setting (T255458)]] * 23:38 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326944{{!}}Retire filebackend lock manager in favour of the default one (T366938)]] (duration: 08m 55s) * 23:34 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 23:31 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326944{{!}}Retire filebackend lock manager in favour of the default one (T366938)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:29 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326944{{!}}Retire filebackend lock manager in favour of the default one (T366938)]] * 23:27 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1170.eqiad.wmnet with OS bookworm * 23:21 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1205.eqiad.wmnet with OS bookworm * 23:20 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1171.eqiad.wmnet with OS bookworm * 23:15 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1206.eqiad.wmnet with OS bookworm * 23:05 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1170.eqiad.wmnet with reason: host reimage * 23:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1205.eqiad.wmnet with reason: host reimage * 22:57 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1171.eqiad.wmnet with reason: host reimage * 22:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1206.eqiad.wmnet with reason: host reimage * 22:53 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1205.eqiad.wmnet with reason: host reimage * 22:51 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1171.eqiad.wmnet with reason: host reimage * 22:51 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1170.eqiad.wmnet with reason: host reimage * 22:50 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1206.eqiad.wmnet with reason: host reimage * 22:36 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1206.eqiad.wmnet with OS bookworm * 22:35 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1205.eqiad.wmnet with OS bookworm * 22:35 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1186.eqiad.wmnet with OS bookworm * 22:35 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1171.eqiad.wmnet with OS bookworm * 22:35 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1170.eqiad.wmnet with OS bookworm * 22:33 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-worker1194.eqiad.wmnet with OS bookworm * 22:22 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326923{{!}}Enable redis lock manager on s6 (T366938)]] (duration: 11m 52s) * 22:18 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 22:13 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326923{{!}}Enable redis lock manager on s6 (T366938)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:10 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326923{{!}}Enable redis lock manager on s6 (T366938)]] * 22:04 sbassett: Deployed security fix for [[phab:T435234|T435234]] (wmf.16) * 21:54 sbassett: Deployed security fix for [[phab:T435234|T435234]] (wmf.15) * 21:38 caro@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326925{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326926{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326929{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]], [[gerrit:1326928{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]] (duration: 0 * 21:34 caro@deploy1003: caro: Continuing with deployment * 21:33 caro@deploy1003: caro: Backport for [[gerrit:1326925{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326926{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326929{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]], [[gerrit:1326928{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]] synced to the testservers (see h * 21:31 caro@deploy1003: Started scap sync-world: Backport for [[gerrit:1326925{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326926{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326929{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]], [[gerrit:1326928{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]] * 21:24 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1204.eqiad.wmnet with reason: 1204 datanode repair [[phab:T434494|T434494]] * 21:02 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326896{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]], [[gerrit:1326897{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]] (duration: 13m 56s) * 20:58 krinkle@deploy1003: krinkle: Continuing with deployment * 20:50 krinkle@deploy1003: krinkle: Backport for [[gerrit:1326896{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]], [[gerrit:1326897{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:49 ryankemper: `an-launcher1003` terminated process group `666809` (`rest_backfill_phase1.sh`) ~20 mins ago with `sudo kill -TERM -- -666809` after its local spark driver (`--driver-memory 64g`) repeatedly exhausted memory on the 32 GB VM and caused SSH to intermittently flap; host recovered to 27 GB available memory * 20:48 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1326896{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]], [[gerrit:1326897{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]] * 20:35 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 20:33 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 20:31 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 20:31 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326870{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]], [[gerrit:1326871{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]] (duration: 07m 35s) * 20:28 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 20:26 kemayo@deploy1003: kemayo: Continuing with deployment * 20:26 ryankemper: `an-launcher1003` confirmed the host is flapping because of memory thrash. chasing down the source of the thrash * 20:25 kemayo@deploy1003: kemayo: Backport for [[gerrit:1326870{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]], [[gerrit:1326871{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:23 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1326870{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]], [[gerrit:1326871{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]] * 20:23 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 20:20 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 20:14 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1168.eqiad.wmnet with OS bookworm * 20:08 zabe: zabe@deploy1003:~$ mwscript extensions/WikimediaMaintenance/maintenance/fixFileRevisionArchiveNameDrift.php enwiki # [[phab:T428406|T428406]] * 20:08 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1204.eqiad.wmnet with OS bookworm * 20:05 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1167.eqiad.wmnet with OS bookworm * 20:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1166.eqiad.wmnet with OS bookworm * 19:54 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1203.eqiad.wmnet with OS bookworm * 19:53 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326907{{!}}Revert "Disable redis lock manager on testwiki"]] (duration: 11m 05s) * 19:50 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1168.eqiad.wmnet with reason: host reimage * 19:47 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1204.eqiad.wmnet with reason: host reimage * 19:46 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 19:44 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326907{{!}}Revert "Disable redis lock manager on testwiki"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:42 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326907{{!}}Revert "Disable redis lock manager on testwiki"]] * 19:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1167.eqiad.wmnet with reason: host reimage * 19:37 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1166.eqiad.wmnet with reason: host reimage * 19:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1203.eqiad.wmnet with reason: host reimage * 19:32 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1167.eqiad.wmnet with reason: host reimage * 19:32 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1168.eqiad.wmnet with reason: host reimage * 19:32 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1166.eqiad.wmnet with reason: host reimage * 19:31 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1204.eqiad.wmnet with reason: host reimage * 19:31 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1203.eqiad.wmnet with reason: host reimage * 19:26 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324751{{!}}InitialiseSettings: Enable 2FA enforcement on more private wikis (T428103)]], [[gerrit:1326875{{!}}Add banner notifying of upcoming 2FA enforcement (T420792)]] (duration: 31m 46s) * 19:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1204.eqiad.wmnet with OS bookworm * 19:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1203.eqiad.wmnet with OS bookworm * 19:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1168.eqiad.wmnet with OS bookworm * 19:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1167.eqiad.wmnet with OS bookworm * 19:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1166.eqiad.wmnet with OS bookworm * 19:15 denisse: rebooting kafkamon2003.codfw.wmnet - [[phab:T435162|T435162]] * 19:14 denisse: rebooting kafkamon1003.eqiad.wmnet [[phab:T435162|T435162]] * 19:13 reedy@deploy1003: reedy: Continuing with deployment * 19:12 reedy@deploy1003: reedy: Backport for [[gerrit:1324751{{!}}InitialiseSettings: Enable 2FA enforcement on more private wikis (T428103)]], [[gerrit:1326875{{!}}Add banner notifying of upcoming 2FA enforcement (T420792)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:54 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324751{{!}}InitialiseSettings: Enable 2FA enforcement on more private wikis (T428103)]], [[gerrit:1326875{{!}}Add banner notifying of upcoming 2FA enforcement (T420792)]] * 18:50 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 18:44 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.16 refs [[phab:T430835|T430835]] * 18:34 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 18:31 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 18:22 aklapper@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326852{{!}}CategoryTree: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]], [[gerrit:1326853{{!}}CategoryViewer: Allow null $html in the CategoryViewerGenerateLink hook (T435161)]], [[gerrit:1326865{{!}}Flow: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]] (duration: 09m 57s) * 18:18 aklapper@deploy1003: jforrester, aklapper: Continuing with deployment * 18:17 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 18:14 aklapper@deploy1003: jforrester, aklapper: Backport for [[gerrit:1326852{{!}}CategoryTree: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]], [[gerrit:1326853{{!}}CategoryViewer: Allow null $html in the CategoryViewerGenerateLink hook (T435161)]], [[gerrit:1326865{{!}}Flow: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki * 18:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1156.eqiad.wmnet with OS bookworm * 18:12 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 18:12 aklapper@deploy1003: Started scap sync-world: Backport for [[gerrit:1326852{{!}}CategoryTree: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]], [[gerrit:1326853{{!}}CategoryViewer: Allow null $html in the CategoryViewerGenerateLink hook (T435161)]], [[gerrit:1326865{{!}}Flow: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]] * 18:11 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 18:08 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1146.eqiad.wmnet with OS bookworm * 18:07 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1177.eqiad.wmnet with OS bookworm * 18:00 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326380{{!}}Introduce main lock manager service (T366938 T427999)]] (duration: 11m 25s) * 17:58 ladsgroup@deploy1003: ladsgroup: Rolling back deployment * 17:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1202.eqiad.wmnet with OS bookworm * 17:55 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1201.eqiad.wmnet with OS bookworm * 17:50 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326380{{!}}Introduce main lock manager service (T366938 T427999)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:48 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326380{{!}}Introduce main lock manager service (T366938 T427999)]] * 17:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1156.eqiad.wmnet with reason: host reimage * 17:46 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1146.eqiad.wmnet with reason: host reimage * 17:45 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-esams and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 17:42 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 17:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1177.eqiad.wmnet with reason: host reimage * 17:38 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1202.eqiad.wmnet with reason: host reimage * 17:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1201.eqiad.wmnet with reason: host reimage * 17:30 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1156.eqiad.wmnet with reason: host reimage * 17:29 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1177.eqiad.wmnet with reason: host reimage * 17:28 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1146.eqiad.wmnet with reason: host reimage * 17:28 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1202.eqiad.wmnet with reason: host reimage * 17:27 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1201.eqiad.wmnet with reason: host reimage * 17:25 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 17:21 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 17:14 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1202.eqiad.wmnet with OS bookworm * 17:14 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1201.eqiad.wmnet with OS bookworm * 17:14 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1177.eqiad.wmnet with OS bookworm * 17:14 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1156.eqiad.wmnet with OS bookworm * 17:14 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1146.eqiad.wmnet with OS bookworm * 17:09 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326885{{!}}w/deployment-info.php: Handle new file format (T434726)]] (duration: 07m 15s) * 17:08 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 17:05 dancy@deploy1003: dancy: Continuing with deployment * 17:04 dancy@deploy1003: dancy: Backport for [[gerrit:1326885{{!}}w/deployment-info.php: Handle new file format (T434726)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:02 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1326885{{!}}w/deployment-info.php: Handle new file format (T434726)]] * 16:50 dancy@deploy1003: Finished scap sync-world: Testing [[phab:T434726|T434726]] (duration: 06m 40s) * 16:43 dancy@deploy1003: Started scap sync-world: Testing [[phab:T434726|T434726]] * 16:43 dancy@deploy1003: Installation of scap version "4.282.0" completed for 3 hosts * 16:41 dancy@deploy1003: Installing scap version "4.282.0" for 3 host(s) * 16:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1145.eqiad.wmnet with OS bookworm * 16:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1200.eqiad.wmnet with OS bookworm * 16:32 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1199.eqiad.wmnet with OS bookworm * 16:18 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1145.eqiad.wmnet with reason: host reimage * 16:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1200.eqiad.wmnet with reason: host reimage * 16:09 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1199.eqiad.wmnet with reason: host reimage * 16:05 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-esams and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 16:05 cjd91: sudo -i cookbook sre.cdn.roll-upgrade-ats --query 'A:cp-esams' --task-id [[phab:T434478|T434478]] --reason '9.2.15 upgrade' * 16:03 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1200.eqiad.wmnet with reason: host reimage * 16:02 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1145.eqiad.wmnet with reason: host reimage * 16:02 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1199.eqiad.wmnet with reason: host reimage * 15:50 moritzm: installing zip security updates * 15:48 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1200.eqiad.wmnet with OS bookworm * 15:47 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1199.eqiad.wmnet with OS bookworm * 15:47 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1145.eqiad.wmnet with OS bookworm * 15:41 topranks: bounce PIC 0/0 on cr1-magru to set port to 40G * 15:37 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1236.eqiad.wmnet with OS bookworm * 15:29 aikochou@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'ores-legacy' for release 'main' . * 15:26 aikochou@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'ores-legacy' for release 'main' . * 15:20 aikochou@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'ores-legacy' for release 'main' . * 15:14 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2002.codfw.wmnet * 15:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1236.eqiad.wmnet with reason: host reimage * 15:12 moritzm: failover ganeti master in codfw to ganeti2047 * 15:09 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1236.eqiad.wmnet with reason: host reimage * 15:09 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2044.codfw.wmnet * 15:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2002.codfw.wmnet * 15:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2044.codfw.wmnet * 15:04 brennen@deploy1003: Finished deploy [phabricator/deployment@6b9b6ff]: deploy phab1004 for [[phab:T435213|T435213]] (duration: 01m 01s) * 15:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2044.codfw.wmnet * 15:03 brennen@deploy1003: Started deploy [phabricator/deployment@6b9b6ff]: deploy phab1004 for [[phab:T435213|T435213]] * 15:03 brennen@deploy1003: Finished deploy [phabricator/deployment@6b9b6ff]: deploy phab2003 for [[phab:T435213|T435213]] (duration: 00m 57s) * 15:02 brennen@deploy1003: Started deploy [phabricator/deployment@6b9b6ff]: deploy phab2003 for [[phab:T435213|T435213]] * 14:57 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply * 14:57 arnaudb@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on phab2003.codfw.wmnet,phab[1004-1006].eqiad.wmnet with reason: maintenance * 14:56 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2044.codfw.wmnet * 14:55 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply * 14:53 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1236.eqiad.wmnet with OS bookworm * 14:51 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2043.codfw.wmnet * 14:51 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2043.codfw.wmnet * 14:45 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2043.codfw.wmnet * 14:33 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2043.codfw.wmnet * 14:22 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2042.codfw.wmnet * 14:22 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2042.codfw.wmnet * 14:21 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db2902.codfw.wmnet with OS trixie * 14:16 elukey: uploaded spicerack_13.2.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia * 14:16 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2042.codfw.wmnet * 14:04 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2042.codfw.wmnet * 14:04 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2041.codfw.wmnet * 14:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2041.codfw.wmnet * 13:59 phuedx@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: apply * 13:59 phuedx@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-main: apply * 13:59 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326838{{!}}Enable Suggested Investigations on hewiki (T435146)]] (duration: 11m 50s) * 13:58 phuedx@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: apply * 13:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2041.codfw.wmnet * 13:57 phuedx@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-main: apply * 13:57 phuedx@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-main: apply * 13:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard1003.eqiad.wmnet * 13:57 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-main: apply * 13:55 phuedx@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-logging-external: apply * 13:55 phuedx@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-logging-external: apply * 13:54 phuedx@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-logging-external: apply * 13:54 stran@deploy1003: stran: Continuing with deployment * 13:54 phuedx@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-logging-external: apply * 13:54 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-logging-external: apply * 13:54 phuedx@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-logging-external: apply * 13:53 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-logging-external: apply * 13:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard1003.eqiad.wmnet * 13:53 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2041.codfw.wmnet * 13:51 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db2902.codfw.wmnet - fceratto@cumin1003" * 13:51 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db2902.codfw.wmnet - fceratto@cumin1003" * 13:51 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2026.codfw.wmnet * 13:51 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard2003.codfw.wmnet * 13:50 phuedx@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: apply * 13:50 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-eqiad * 13:50 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp1001.eqiad.wmnet * 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp1001.eqiad.wmnet * 13:50 phuedx@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: apply * 13:49 phuedx@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: apply * 13:49 stran@deploy1003: stran: Backport for [[gerrit:1326838{{!}}Enable Suggested Investigations on hewiki (T435146)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:48 phuedx@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: apply * 13:48 phuedx@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics: apply * 13:47 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard2003.codfw.wmnet * 13:47 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics: apply * 13:47 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1326838{{!}}Enable Suggested Investigations on hewiki (T435146)]] * 13:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp1001.eqiad.wmnet * 13:44 cdanis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 13:43 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp1001.eqiad.wmnet * 13:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1315-1327].eqiad.wmnet * 13:43 cdanis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 13:43 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1315-1327].eqiad.wmnet * 13:40 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki1001.eqiad.wmnet * 13:35 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1315-1327].eqiad.wmnet * 13:34 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host rpki1001.eqiad.wmnet * 13:27 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1315-1327].eqiad.wmnet * 13:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1302-1314].eqiad.wmnet * 13:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1302-1314].eqiad.wmnet * 13:21 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326775{{!}}SI: Instrument case update on first edit (T435048)]] (duration: 07m 12s) * 13:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1302-1314].eqiad.wmnet * 13:17 stran@deploy1003: stran: Continuing with deployment * 13:16 stran@deploy1003: stran: Backport for [[gerrit:1326775{{!}}SI: Instrument case update on first edit (T435048)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:14 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1326775{{!}}SI: Instrument case update on first edit (T435048)]] * 13:13 moritzm: installing util-linux security updates * 13:11 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1302-1314].eqiad.wmnet * 13:10 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326232{{!}}prv: Enable parsoid rendering for 5 wikis (T435115)]] (duration: 08m 17s) * 13:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1288-1289,1291-1301].eqiad.wmnet * 13:10 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply * 13:10 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1288-1289,1291-1301].eqiad.wmnet * 13:10 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply * 13:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki2003.codfw.wmnet * 13:06 jgiannelos@deploy1003: jgiannelos: Continuing with deployment * 13:05 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host rpki2003.codfw.wmnet * 13:04 jgiannelos@deploy1003: jgiannelos: Backport for [[gerrit:1326232{{!}}prv: Enable parsoid rendering for 5 wikis (T435115)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1288-1289,1291-1301].eqiad.wmnet * 13:02 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1326232{{!}}prv: Enable parsoid rendering for 5 wikis (T435115)]] * 12:53 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1288-1289,1291-1301].eqiad.wmnet * 12:53 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1273,1275-1287].eqiad.wmnet * 12:53 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1273,1275-1287].eqiad.wmnet * 12:52 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt1002.wikimedia.org * 12:51 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db2902.codfw.wmnet on all recursors * 12:51 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db2902.codfw.wmnet on all recursors * 12:51 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:51 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db2902.codfw.wmnet - fceratto@cumin1003" * 12:51 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db2902.codfw.wmnet - fceratto@cumin1003" * 12:46 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt1002.wikimedia.org * 12:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt2002.wikimedia.org * 12:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1273,1275-1287].eqiad.wmnet * 12:42 dhinus: repooled clouddb1032 that was currently <nowiki>{</nowiki>"weight": 0, "pooled": "inactive"<nowiki>}</nowiki> for both s4 and s6 * 12:41 dhinus: also depooled clouddb1017 (forgot it in the previous list) * 12:41 fnegri@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet * 12:40 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 12:40 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db2902.codfw.wmnet * 12:40 fnegri@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032.eqiad.wmnet * 12:40 fnegri@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032 * 12:39 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1017.eqiad.wmnet * 12:39 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt2002.wikimedia.org * 12:38 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1020.eqiad.wmnet * 12:38 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1018.eqiad.wmnet * 12:38 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet * 12:37 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1014.eqiad.wmnet * 12:37 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1013.eqiad.wmnet * 12:37 dhinus: depool again clouddb10[13,14,16,18,20] that were repooled by the cookbook sre.mysql.multiinstance_reboot * 12:36 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1273,1275-1287].eqiad.wmnet * 12:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1248-1261].eqiad.wmnet * 12:36 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1248-1261].eqiad.wmnet * 12:35 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2026.codfw.wmnet * 12:34 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 12:34 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 12:34 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 12:34 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 12:29 lucaswerkmeister-wmde@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 12:28 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2026.codfw.wmnet * 12:28 lucaswerkmeister-wmde@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 12:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1248-1261].eqiad.wmnet * 12:27 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.addnode (exit_code=0) for new host ganeti2046.codfw.wmnet to cluster codfw and group A * 12:26 moritzm: readded ganeti2046 to the codfw cluster following firmware update and reimage [[phab:T434681|T434681]] * 12:23 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply * 12:23 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply * 12:23 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply * 12:22 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply * 12:22 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply * 12:22 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply * 12:21 jmm@cumin2003: START - Cookbook sre.ganeti.addnode for new host ganeti2046.codfw.wmnet to cluster codfw and group A * 12:21 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1169.eqiad.wmnet onto db1283.eqiad.wmnet * 12:21 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1169: Pool db1169.eqiad.wmnet in after cloning * 12:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1248-1261].eqiad.wmnet * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1149-1153,1158,1240-1247].eqiad.wmnet * 12:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1149-1153,1158,1240-1247].eqiad.wmnet * 12:11 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1149-1153,1158,1240-1247].eqiad.wmnet * 12:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1168.eqiad.wmnet onto db1282.eqiad.wmnet * 12:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1168: Pool db1168.eqiad.wmnet in after cloning * 12:06 jmm@cumin2003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-test-eqiad * 12:05 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2026.codfw.wmnet * 12:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1149-1153,1158,1240-1247].eqiad.wmnet * 12:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1128-1134,1142-1148].eqiad.wmnet * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2046.codfw.wmnet * 12:01 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1128-1134,1142-1148].eqiad.wmnet * 11:57 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2025.codfw.wmnet * 11:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2025.codfw.wmnet * 11:54 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2046.codfw.wmnet * 11:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1128-1134,1142-1148].eqiad.wmnet * 11:50 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2025.codfw.wmnet * 11:46 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1128-1134,1142-1148].eqiad.wmnet * 11:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1114-1127].eqiad.wmnet * 11:45 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1114-1127].eqiad.wmnet * 11:39 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db2901.codfw.wmnet * 11:39 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db2901.codfw.wmnet * 11:36 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2025.codfw.wmnet * 11:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1114-1127].eqiad.wmnet * 11:35 fceratto@cumin1003: END (ERROR) - Cookbook sre.ganeti.makevm (exit_code=93) for new host db1901.eqiad.wmnet * 11:35 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1169: Pool db1169.eqiad.wmnet in after cloning * 11:35 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 11:30 jmm@cumin2003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-test-eqiad * 11:27 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1114-1127].eqiad.wmnet * 11:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1076-1081,1084-1087,1093-1095,1113].eqiad.wmnet * 11:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1076-1081,1084-1087,1093-1095,1113].eqiad.wmnet * 11:25 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1168: Pool db1168.eqiad.wmnet in after cloning * 11:24 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1194.eqiad.wmnet with OS bookworm * 11:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1181.eqiad.wmnet with OS bookworm * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2050.codfw.wmnet * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2050.codfw.wmnet * 11:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1076-1081,1084-1087,1093-1095,1113].eqiad.wmnet * 11:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2050.codfw.wmnet * 11:14 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2004.codfw.wmnet * 11:10 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2050.codfw.wmnet * 11:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1076-1081,1084-1087,1093-1095,1113].eqiad.wmnet * 11:09 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1286: Pool back * 11:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1045-1050,1056-1057,1064-1066,1073-1075].eqiad.wmnet * 11:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1045-1050,1056-1057,1064-1066,1073-1075].eqiad.wmnet * 11:08 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2004.codfw.wmnet * 11:07 moritzm: installing PHP 8.4 security updates * 11:06 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on an-worker1194.eqiad.wmnet with reason: host reimage * 11:06 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1194.eqiad.wmnet with reason: host reimage * 11:05 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2049.codfw.wmnet * 11:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2049.codfw.wmnet * 11:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid1003.eqiad.wmnet * 11:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1045-1050,1056-1057,1064-1066,1073-1075].eqiad.wmnet * 10:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1181.eqiad.wmnet with reason: host reimage * 10:59 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2049.codfw.wmnet * 10:59 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid1003.eqiad.wmnet * 10:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid2003.codfw.wmnet * 10:54 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1181.eqiad.wmnet with reason: host reimage * 10:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid2003.codfw.wmnet * 10:51 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1901.eqiad.wmnet on all recursors * 10:51 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1901.eqiad.wmnet on all recursors * 10:51 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:51 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:51 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:50 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2049.codfw.wmnet * 10:50 blake@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:50 blake@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:49 blake@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:48 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1045-1050,1056-1057,1064-1066,1073-1075].eqiad.wmnet * 10:48 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1044].eqiad.wmnet * 10:48 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1044].eqiad.wmnet * 10:47 blake@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:46 blake@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:46 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2047.codfw.wmnet * 10:46 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:46 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2047.codfw.wmnet * 10:46 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1901.eqiad.wmnet * 10:46 blake@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:44 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1168: Depool db1168.eqiad.wmnet to then clone it to db1282.eqiad.wmnet - marostegui@cumin1003 * 10:44 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1168: Depool db1168.eqiad.wmnet to then clone it to db1282.eqiad.wmnet - marostegui@cumin1003 * 10:44 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1168.eqiad.wmnet onto db1282.eqiad.wmnet * 10:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2047.codfw.wmnet * 10:40 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1044].eqiad.wmnet * 10:37 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2047.codfw.wmnet * 10:37 fceratto@cumin1003: END (ERROR) - Cookbook sre.ganeti.makevm (exit_code=93) for new host db1901.eqiad.wmnet * 10:36 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 10:34 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:33 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2032.codfw.wmnet * 10:33 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2032.codfw.wmnet * 10:32 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1044].eqiad.wmnet * 10:32 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-eqiad * 10:27 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2032.codfw.wmnet * 10:24 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1286: Pool back * 10:24 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1286 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96163 and previous config saved to /var/cache/conftool/dbconfig/20260818-102431-marostegui.json * 10:22 moritzm: installing Django security updates * 10:20 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul2002.codfw.wmnet * 10:20 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul2001.codfw.wmnet * 10:17 blake@deploy1003: Finished scap sync-world: no-build deployment for [[phab:T417800|T417800]] (duration: 04m 40s) * 10:16 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul2002.codfw.wmnet * 10:16 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul2001.codfw.wmnet * 10:15 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2032.codfw.wmnet * 10:14 blake@deploy1003: Started scap sync-world: no-build deployment for [[phab:T417800|T417800]] * 10:12 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul2003.codfw.wmnet * 10:12 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy2001.codfw.wmnet * 10:12 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy3001.esams.wmnet * 10:12 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy1001.eqiad.wmnet * 10:08 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2031.codfw.wmnet * 10:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2031.codfw.wmnet * 10:08 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul2003.codfw.wmnet * 10:08 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy2001.codfw.wmnet * 10:08 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy3001.esams.wmnet * 10:08 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy1001.eqiad.wmnet * 10:07 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy1002.eqiad.wmnet * 10:05 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy2002.codfw.wmnet * 10:05 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy3002.esams.wmnet * 10:04 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy1002.eqiad.wmnet * 10:03 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy4003.ulsfo.wmnet * 10:03 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy5003.eqsin.wmnet * 10:02 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2031.codfw.wmnet * 10:02 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1169: Depool db1169.eqiad.wmnet to then clone it to db1283.eqiad.wmnet - marostegui@cumin1003 * 10:01 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy2002.codfw.wmnet * 10:01 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy4004.ulsfo.wmnet * 10:01 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy3002.esams.wmnet * 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1169: Depool db1169.eqiad.wmnet to then clone it to db1283.eqiad.wmnet - marostegui@cumin1003 * 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1169.eqiad.wmnet onto db1283.eqiad.wmnet * 10:01 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy5004.eqsin.wmnet * 09:59 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy4003.ulsfo.wmnet * 09:59 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy4004.ulsfo.wmnet * 09:59 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy5003.eqsin.wmnet * 09:59 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy7001.magru.wmnet * 09:59 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy5004.eqsin.wmnet * 09:58 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy7002.magru.wmnet * 09:57 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2031.codfw.wmnet * 09:57 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy6001.drmrs.wmnet * 09:57 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy6002.drmrs.wmnet * 09:54 filippo@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for 10 hosts * 09:54 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast4006.wikimedia.org * 09:53 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy6001.drmrs.wmnet * 09:53 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy6002.drmrs.wmnet * 09:53 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host releases1003.eqiad.wmnet * 09:53 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy7001.magru.wmnet * 09:53 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host people2004.codfw.wmnet * 09:52 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host people1005.eqiad.wmnet * 09:52 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy7002.magru.wmnet * 09:50 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host releases2003.codfw.wmnet * 09:49 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host releases2003.codfw.wmnet * 09:49 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host releases1003.eqiad.wmnet * 09:49 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host people2004.codfw.wmnet * 09:48 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host people1005.eqiad.wmnet * 09:46 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2040.codfw.wmnet * 09:46 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2040.codfw.wmnet * 09:43 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1901.eqiad.wmnet on all recursors * 09:42 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1901.eqiad.wmnet on all recursors * 09:42 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:42 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:42 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2040.codfw.wmnet * 09:40 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp2005.wikimedia.org * 09:36 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp2005.wikimedia.org * 09:31 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:31 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1901.eqiad.wmnet * 09:31 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1901.eqiad.wmnet * 09:31 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:31 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1901.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 09:31 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1901.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 09:29 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2040.codfw.wmnet * 09:29 slyngshede@dns1004: END - running authdns-update * 09:28 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2039.codfw.wmnet * 09:28 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2039.codfw.wmnet * 09:27 slyngshede@dns1004: START - running authdns-update * 09:26 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test2005.wikimedia.org * 09:22 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2039.codfw.wmnet * 09:22 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test2005.wikimedia.org * 09:22 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp1005.wikimedia.org * 09:21 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:19 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2039.codfw.wmnet * 09:19 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2038.codfw.wmnet * 09:18 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1287: Pool back * 09:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2038.codfw.wmnet * 09:18 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp1005.wikimedia.org * 09:18 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test1005.wikimedia.org * 09:17 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1901.eqiad.wmnet * 09:15 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db1901.eqiad.wmnet * 09:15 fceratto@cumin1003: END (ERROR) - Cookbook sre.dns.netbox (exit_code=97) * 09:14 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test1005.wikimedia.org * 09:13 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2038.codfw.wmnet * 09:13 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:13 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1901.eqiad.wmnet * 09:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast4006.wikimedia.org * 09:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1288: Pool back * 09:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast5005.wikimedia.org * 09:03 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2038.codfw.wmnet * 08:57 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2037.codfw.wmnet * 08:57 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast5005.wikimedia.org * 08:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2037.codfw.wmnet * 08:56 filippo@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 10 hosts * 08:52 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2037.codfw.wmnet * 08:51 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1289: Pool back * 08:35 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudidp2001-dev.codfw.wmnet * 08:34 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1002-dev.eqiad.wmnet * 08:33 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1287: Pool back * 08:33 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1287 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96150 and previous config saved to /var/cache/conftool/dbconfig/20260818-083311-marostegui.json * 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1001-dev.eqiad.wmnet * 08:31 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudidp2001-dev.codfw.wmnet * 08:30 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1002-dev.eqiad.wmnet * 08:30 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2037.codfw.wmnet * 08:29 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1001-dev.eqiad.wmnet * 08:28 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2036.codfw.wmnet * 08:28 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2036.codfw.wmnet * 08:25 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1288: Pool back * 08:23 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 08:23 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1181.eqiad.wmnet with OS bookworm * 08:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2036.codfw.wmnet * 08:22 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1288 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96148 and previous config saved to /var/cache/conftool/dbconfig/20260818-082234-marostegui.json * 08:20 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2036.codfw.wmnet * 08:18 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2035.codfw.wmnet * 08:18 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudvirt1057.eqiad.wmnet with OS trixie * 08:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2035.codfw.wmnet * 08:13 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2035.codfw.wmnet * 08:06 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2035.codfw.wmnet * 08:05 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1289: Pool back * 08:05 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1289 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96145 and previous config saved to /var/cache/conftool/dbconfig/20260818-080531-marostegui.json * 07:54 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 07:51 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326439{{!}}Fix "mathjax_ignore" handling around forcemathmode attribute (T434686)]] (duration: 13m 15s) * 07:48 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 07:47 krinkle@deploy1003: krinkle: Continuing with deployment * 07:44 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ganeti2046.codfw.wmnet with OS bookworm * 07:40 krinkle@deploy1003: krinkle: Backport for [[gerrit:1326439{{!}}Fix "mathjax_ignore" handling around forcemathmode attribute (T434686)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:39 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 07:38 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1326439{{!}}Fix "mathjax_ignore" handling around forcemathmode attribute (T434686)]] * 07:34 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 07:31 samwilson@deploy1003: Finished scap sync-world: Backport for [[gerrit:701016{{!}}InitialiseSettings and -labs: Remove redundant feature flag $wgWikisourceEnableOcr (T285311)]] (duration: 07m 47s) * 07:30 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1048.eqiad.wmnet * 07:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1048.eqiad.wmnet * 07:29 XioNoX: add gnmic 0.47.0 to bookworm and trixie reprepro * 07:28 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ganeti2046.codfw.wmnet with reason: host reimage * 07:27 samwilson@deploy1003: samwilson: Continuing with deployment * 07:25 samwilson@deploy1003: samwilson: Backport for [[gerrit:701016{{!}}InitialiseSettings and -labs: Remove redundant feature flag $wgWikisourceEnableOcr (T285311)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:25 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 07:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1048.eqiad.wmnet * 07:24 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ganeti2046.codfw.wmnet with reason: host reimage * 07:23 samwilson@deploy1003: Started scap sync-world: Backport for [[gerrit:701016{{!}}InitialiseSettings and -labs: Remove redundant feature flag $wgWikisourceEnableOcr (T285311)]] * 07:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1198.eqiad.wmnet with OS bookworm * 07:18 samwilson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326450{{!}}InitialiseSettings.php: Enable Bulk OCR on pawikisource (T434648)]] (duration: 12m 17s) * 07:17 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 07:11 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1048.eqiad.wmnet * 07:11 samwilson@deploy1003: samwilson: Continuing with deployment * 07:11 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ganeti2046.codfw.wmnet with OS bookworm * 07:10 samwilson@deploy1003: samwilson: Backport for [[gerrit:1326450{{!}}InitialiseSettings.php: Enable Bulk OCR on pawikisource (T434648)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1057.eqiad.wmnet with OS trixie * 07:05 samwilson@deploy1003: Started scap sync-world: Backport for [[gerrit:1326450{{!}}InitialiseSettings.php: Enable Bulk OCR on pawikisource (T434648)]] * 07:02 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1198.eqiad.wmnet with reason: host reimage * 07:01 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1056.eqiad.wmnet with OS trixie * 07:01 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1056.eqiad.wmnet with OS trixie * 07:00 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1056.eqiad.wmnet with OS trixie * 07:00 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1056.eqiad.wmnet with OS trixie * 06:59 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudvirt1055.eqiad.wmnet with OS trixie * 06:58 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1198.eqiad.wmnet with reason: host reimage * 06:53 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1055.eqiad.wmnet with OS trixie * 06:52 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudvirt1054.eqiad.wmnet with OS trixie * 06:44 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1194.eqiad.wmnet with OS bookworm * 06:44 XioNoX: upgrade eqsin gnmic to 0.47.0 * 06:43 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1198.eqiad.wmnet with OS bookworm * 06:41 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1054.eqiad.wmnet with OS trixie * 06:09 arnaudb@cumin1003: END (PASS) - Cookbook sre.gerrit.restart-gerrit (exit_code=0) Restarting Gerrit on gerrit2002 * 06:06 arnaudb@cumin1003: START - Cookbook sre.gerrit.restart-gerrit Restarting Gerrit on gerrit2002 * 06:06 arnaudb@cumin1003: END (PASS) - Cookbook sre.gerrit.restart-gerrit (exit_code=0) Restarting Gerrit on gerrit1003 * 06:04 arnaudb@cumin1003: START - Cookbook sre.gerrit.restart-gerrit Restarting Gerrit on gerrit1003 * 06:02 arnaudb@cumin1003: END (PASS) - Cookbook sre.gerrit.restart-gerrit (exit_code=0) Restarting Gerrit on gerrit2003 * 06:00 arnaudb@cumin1003: START - Cookbook sre.gerrit.restart-gerrit Restarting Gerrit on gerrit2003 * 05:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1155.eqiad.wmnet with OS bookworm * 05:38 arnaudb: updating prometheusBearerToken on gerrit * 05:28 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1155.eqiad.wmnet with reason: host reimage * 05:23 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1155.eqiad.wmnet with reason: host reimage * 05:06 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1155.eqiad.wmnet with OS bookworm * 04:57 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1144.eqiad.wmnet with OS bookworm * 04:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1144.eqiad.wmnet with reason: host reimage * 04:29 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1144.eqiad.wmnet with reason: host reimage * 04:14 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1144.eqiad.wmnet with OS bookworm * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.13 (duration: 02m 23s) * 03:45 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1197.eqiad.wmnet with OS bookworm * 03:41 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1165.eqiad.wmnet with OS bookworm * 03:38 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.16 refs [[phab:T430835|T430835]] (duration: 34m 43s) * 03:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1164.eqiad.wmnet with OS bookworm * 03:35 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1196.eqiad.wmnet with OS bookworm * 03:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1163.eqiad.wmnet with OS bookworm * 03:22 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1197.eqiad.wmnet with reason: host reimage * 03:18 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1165.eqiad.wmnet with reason: host reimage * 03:15 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1196.eqiad.wmnet with reason: host reimage * 03:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1164.eqiad.wmnet with reason: host reimage * 03:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1163.eqiad.wmnet with reason: host reimage * 03:05 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1165.eqiad.wmnet with reason: host reimage * 03:04 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1164.eqiad.wmnet with reason: host reimage * 03:04 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1197.eqiad.wmnet with reason: host reimage * 03:04 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1196.eqiad.wmnet with reason: host reimage * 03:04 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1163.eqiad.wmnet with reason: host reimage * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.16 refs [[phab:T430835|T430835]] * 02:50 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1197.eqiad.wmnet with OS bookworm * 02:49 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1196.eqiad.wmnet with OS bookworm * 02:49 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1165.eqiad.wmnet with OS bookworm * 02:49 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1164.eqiad.wmnet with OS bookworm * 02:48 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1163.eqiad.wmnet with OS bookworm * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 46s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:15 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-codfw: Set storage compatability to NONE — [[phab:T433026|T433026]] - eevans@cumin1003 * 00:38 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-codfw: Set storage compatability to NONE — [[phab:T433026|T433026]] - eevans@cumin1003 * 00:11 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326373{{!}}PersonalDashboard: add newly renamed *ReviewChangesMlModel setting (T422148)]] (duration: 07m 07s) * 00:09 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-eqiad: Set storage compatability to NONE — [[phab:T433026|T433026]] - eevans@cumin1003 * 00:07 musikanimal@deploy1003: musikanimal: Continuing with deployment * 00:06 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1326373{{!}}PersonalDashboard: add newly renamed *ReviewChangesMlModel setting (T422148)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:04 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1326373{{!}}PersonalDashboard: add newly renamed *ReviewChangesMlModel setting (T422148)]] == 2026-08-17 == * 23:30 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-eqiad: Set storage compatability to NONE — [[phab:T433026|T433026]] - eevans@cumin1003 * 23:05 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-codfw: Set storage compatability to UPGRADING — [[phab:T433026|T433026]] - eevans@cumin1003 * 22:28 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-codfw: Set storage compatability to UPGRADING — [[phab:T433026|T433026]] - eevans@cumin1003 * 21:46 logmsgbot: jforrester Deployed security patch for [[phab:T435085|T435085]] * 21:39 swfrench@deploy1003: mwscript-k8s job started: purgeList.php # [[phab:T432412|T432412]] * 21:37 maryum: Undeploy security fix for [[phab:T433020|T433020]] * 21:23 maryum: Deployed security fix for [[phab:T433020|T433020]] * 21:21 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 21:21 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 21:14 maryum: Deployed security fix for [[phab:T434967|T434967]] * 20:58 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-eqiad: Set storage compatability to UPGRADING — [[phab:T433026|T433026]] - eevans@cumin1003 * 20:40 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324320{{!}}[itwiki/slwiki/tgwiki] Remove temporary Wikipedia 25 logos permanently (already reverted) (T414265 T414320 T415307)]] (duration: 06m 55s) * 20:40 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 20:39 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 20:36 cjming@deploy1003: cjming, superpes: Continuing with deployment * 20:35 cjming@deploy1003: cjming, superpes: Backport for [[gerrit:1324320{{!}}[itwiki/slwiki/tgwiki] Remove temporary Wikipedia 25 logos permanently (already reverted) (T414265 T414320 T415307)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:33 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1324320{{!}}[itwiki/slwiki/tgwiki] Remove temporary Wikipedia 25 logos permanently (already reverted) (T414265 T414320 T415307)]] * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ttmserver-test: apply * 20:30 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326361{{!}}Remove escaped paths in app site association file (T432412)]] (duration: 13m 28s) * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ttmserver-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-toolhub-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-toolhub-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-toolhub-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-toolhub-test: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-toolhub: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-toolhub: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-toolhub: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-toolhub: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-test: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-test: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 20:27 inflatador: bking@deploy1003 `charlie --services_dir dse-k8s-services -s opensearch-* -e dse-k8s-* apply` [[phab:T435125|T435125]] * 20:27 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-apifeatureusage-test: apply * 20:27 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-apifeatureusage-test: apply * 20:27 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-apifeatureusage-test: apply * 20:27 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-apifeatureusage-test: apply * 20:27 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-apifeatureusage: apply * 20:26 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-apifeatureusage: apply * 20:26 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-apifeatureusage: apply * 20:26 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-apifeatureusage: apply * 20:26 cjming@deploy1003: cjming, tsev: Continuing with deployment * 20:24 inflatador: bking@deploy1003 `charlie --services_dir dse-k8s-services -s opensearch-* -e dse-k8s-*` [[phab:T435125|T435125]] * 20:21 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-eqiad: Set storage compatability to UPGRADING — [[phab:T433026|T433026]] - eevans@cumin1003 * 20:19 cjming@deploy1003: cjming, tsev: Backport for [[gerrit:1326361{{!}}Remove escaped paths in app site association file (T432412)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:17 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1326361{{!}}Remove escaped paths in app site association file (T432412)]] * 20:15 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326077{{!}}InstrumentConstructiveEdits: anchor all runs to the nearest `interval` (T431493)]] (duration: 06m 25s) * 20:11 cjming@deploy1003: cjming: Continuing with deployment * 20:10 cjming@deploy1003: cjming: Backport for [[gerrit:1326077{{!}}InstrumentConstructiveEdits: anchor all runs to the nearest `interval` (T431493)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:08 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1326077{{!}}InstrumentConstructiveEdits: anchor all runs to the nearest `interval` (T431493)]] * 20:06 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1160.eqiad.wmnet with OS bookworm * 20:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1162.eqiad.wmnet with OS bookworm * 19:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1184.eqiad.wmnet with OS bookworm * 19:55 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1161.eqiad.wmnet with OS bookworm * 19:49 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1195.eqiad.wmnet with OS bookworm * 19:49 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-test: apply * 19:49 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-test: apply * 19:44 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1160.eqiad.wmnet with reason: host reimage * 19:41 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-codfw: Upgrade to Java 17 — [[phab:T433026|T433026]] - eevans@cumin1003 * 19:39 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1162.eqiad.wmnet with reason: host reimage * 19:36 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-eqsin and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 19:36 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-test: apply * 19:36 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1184.eqiad.wmnet with reason: host reimage * 19:32 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1161.eqiad.wmnet with reason: host reimage * 19:29 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1195.eqiad.wmnet with reason: host reimage * 19:26 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1161.eqiad.wmnet with reason: host reimage * 19:26 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1160.eqiad.wmnet with reason: host reimage * 19:26 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1184.eqiad.wmnet with reason: host reimage * 19:26 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1162.eqiad.wmnet with reason: host reimage * 19:25 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1195.eqiad.wmnet with reason: host reimage * 19:19 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-test: apply * 19:11 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1195.eqiad.wmnet with OS bookworm * 19:10 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1194.eqiad.wmnet with OS bookworm * 19:10 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1184.eqiad.wmnet with OS bookworm * 19:10 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1162.eqiad.wmnet with OS bookworm * 19:10 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1161.eqiad.wmnet with OS bookworm * 19:10 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1160.eqiad.wmnet with OS bookworm * 19:10 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 19:09 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 19:03 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-codfw: Upgrade to Java 17 — [[phab:T433026|T433026]] - eevans@cumin1003 * 18:58 dancy@deploy1003: Finished scap sync-world: testing [[phab:T375514|T375514]] (duration: 03m 13s) * 18:55 dancy@deploy1003: Started scap sync-world: testing [[phab:T375514|T375514]] * 18:55 dwisehaupt@dns1006: END - running authdns-update * 18:54 dancy@deploy1003: Installation of scap version "4.281.1" completed for 3 hosts * 18:54 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2009.codfw.wmnet * 18:53 dwisehaupt@dns1006: START - running authdns-update * 18:52 dancy@deploy1003: Installing scap version "4.281.1" for 3 host(s) * 18:47 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2009.codfw.wmnet * 18:41 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2008.codfw.wmnet * 18:34 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2008.codfw.wmnet * 18:30 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2007.codfw.wmnet * 18:23 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2007.codfw.wmnet * 18:16 dwisehaupt@dns1005: END - running authdns-update * 18:14 dwisehaupt@dns1005: START - running authdns-update * 18:04 swfrench@deploy1003: Finished scap sync-world: Deploy "Point Test Wiki to new docroot" - [[phab:T432412|T432412]] (duration: 26m 03s) * 18:00 swfrench@deploy1003: swfrench: Continuing with deployment * 17:51 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1057.eqiad.wmnet with OS trixie * 17:47 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1180.eqiad.wmnet with OS bookworm * 17:39 swfrench@deploy1003: swfrench: Deploy "Point Test Wiki to new docroot" - [[phab:T432412|T432412]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:39 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-eqsin and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 17:38 swfrench@deploy1003: Started scap sync-world: Deploy "Point Test Wiki to new docroot" - [[phab:T432412|T432412]] * 17:36 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1159.eqiad.wmnet with OS bookworm * 17:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1158.eqiad.wmnet with OS bookworm * 17:30 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-eqiad: Upgrade to Java 17 — [[phab:T433026|T433026]] - eevans@cumin1003 * 17:30 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-ulsfo or A:cp-drmrs and A:cp - 9.2.15 upgrade ([[phab:T434620|T434620]]) * 17:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1157.eqiad.wmnet with OS bookworm * 17:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1180.eqiad.wmnet with reason: host reimage * 17:22 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 17:21 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1193.eqiad.wmnet with OS bookworm * 17:21 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 17:21 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 17:20 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 17:20 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 17:20 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1183.eqiad.wmnet with OS bookworm * 17:18 swfrench@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 17:18 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 17:17 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1180.eqiad.wmnet with reason: host reimage * 17:17 swfrench@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 17:17 swfrench@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 17:16 swfrench@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 17:16 swfrench@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 17:15 swfrench@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 17:15 swfrench@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 17:15 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1192.eqiad.wmnet with OS bookworm * 17:14 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1159.eqiad.wmnet with reason: host reimage * 17:14 swfrench@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 17:13 swfrench@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 17:12 swfrench@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 17:09 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1158.eqiad.wmnet with reason: host reimage * 17:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1157.eqiad.wmnet with reason: host reimage * 17:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1180 * 17:02 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1180 * 17:01 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica-eqsin and A:liberica * 17:01 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1193.eqiad.wmnet with reason: host reimage * 16:57 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1183.eqiad.wmnet with reason: host reimage * 16:56 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1054.eqiad.wmnet with OS trixie * 16:55 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1056.eqiad.wmnet with OS trixie * 16:54 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1193.eqiad.wmnet with reason: host reimage * 16:54 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1192.eqiad.wmnet with reason: host reimage * 16:51 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-eqiad: Upgrade to Java 17 — [[phab:T433026|T433026]] - eevans@cumin1003 * 16:50 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1180 * 16:50 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1180.eqiad.wmnet 17.36.64.10.in-addr.arpa 7.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:50 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1180.eqiad.wmnet 17.36.64.10.in-addr.arpa 7.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:50 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:50 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1180 - btullis@cumin1003" * 16:50 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1180 - btullis@cumin1003" * 16:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1158.eqiad.wmnet with reason: host reimage * 16:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1157.eqiad.wmnet with reason: host reimage * 16:49 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica-eqsin and A:liberica * 16:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1183.eqiad.wmnet with reason: host reimage * 16:48 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1159.eqiad.wmnet with reason: host reimage * 16:47 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1192.eqiad.wmnet with reason: host reimage * 16:46 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns4004.wikimedia.org * 16:46 sukhe@dns1004: END - running authdns-update * 16:44 sukhe@dns1004: START - running authdns-update * 16:44 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns4004.wikimedia.org,service=authdns-update * 16:43 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns4004.wikimedia.org with OS trixie * 16:39 btullis@cumin1003: START - Cookbook sre.dns.netbox * 16:39 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1180 * 16:39 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1193.eqiad.wmnet with OS bookworm * 16:39 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1180.eqiad.wmnet with OS bookworm * 16:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1192.eqiad.wmnet with OS bookworm * 16:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1183.eqiad.wmnet with OS bookworm * 16:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1159.eqiad.wmnet with OS bookworm * 16:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1158.eqiad.wmnet with OS bookworm * 16:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1157.eqiad.wmnet with OS bookworm * 16:31 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1057.eqiad.wmnet with OS trixie * 16:30 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1057.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 16:29 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica-drmrs and A:liberica * 16:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1154.eqiad.wmnet with OS bookworm * 16:23 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1057.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 16:23 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1055.eqiad.wmnet with OS trixie * 16:22 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1057 * 16:22 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1057 * 16:21 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:21 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1057] - vriley@cumin1003" * 16:21 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1057] - vriley@cumin1003" * 16:19 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica-drmrs and A:liberica * 16:17 vriley@cumin1003: START - Cookbook sre.dns.netbox * 16:16 phuedx@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics-external: apply * 16:15 phuedx@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics-external: apply * 16:13 phuedx@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics-external: apply * 16:12 phuedx@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics-external: apply * 16:11 phuedx@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics-external: apply * 16:09 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics-external: apply * 16:09 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1176.eqiad.wmnet with OS bookworm * 16:07 btullis@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 16:06 btullis@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 16:05 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1191.eqiad.wmnet with OS bookworm * 16:03 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1154.eqiad.wmnet with reason: host reimage * 16:03 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching aqs[2002-2012].codfw.wmnet,aqs[1017-1027].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433026|T433026]] - eevans@cumin1003 * 15:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1190.eqiad.wmnet with OS bookworm * 15:59 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1154.eqiad.wmnet with reason: host reimage * 15:55 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-debug: apply * 15:55 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-debug: apply * 15:55 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-debug: apply * 15:55 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/mw-debug: apply * 15:53 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns4004.wikimedia.org with reason: host reimage * 15:50 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns4004.wikimedia.org with reason: host reimage * 15:46 btullis@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 15:46 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1176.eqiad.wmnet with reason: host reimage * 15:45 btullis@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 15:43 btullis@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 15:42 btullis@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 15:42 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1191.eqiad.wmnet with reason: host reimage * 15:39 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1190.eqiad.wmnet with reason: host reimage * 15:36 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1054.eqiad.wmnet with OS trixie * 15:35 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:35 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1056.eqiad.wmnet with OS trixie * 15:35 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:34 moritzm: failover Ganeti master in eqiad to ganeti1046 * 15:34 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1176.eqiad.wmnet with reason: host reimage * 15:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1191.eqiad.wmnet with reason: host reimage * 15:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1190.eqiad.wmnet with reason: host reimage * 15:31 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns2006.wikimedia.org * 15:31 sukhe@dns1004: END - running authdns-update * 15:29 sukhe@dns1004: START - running authdns-update * 15:29 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns2006.wikimedia.org,service=authdns-update * 15:29 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns2006.wikimedia.org * 15:29 sukhe@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns2006.wikimedia.org * 15:26 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica-ulsfo and A:liberica * 15:26 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns1006.wikimedia.org * 15:26 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:25 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns2006.wikimedia.org with OS trixie * 15:25 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1056 * 15:25 sukhe@dns1004: END - running authdns-update * 15:23 sukhe@dns1004: START - running authdns-update * 15:23 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns1006.wikimedia.org,service=authdns-update * 15:23 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns1006.wikimedia.org * 15:23 sukhe@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns1006.wikimedia.org * 15:20 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1056 * 15:20 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:20 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1056~] - vriley@cumin1003" * 15:19 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1056~] - vriley@cumin1003" * 15:19 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns1006.wikimedia.org with OS trixie * 15:19 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns4004.wikimedia.org with OS trixie * 15:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1190.eqiad.wmnet with OS bookworm * 15:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1191.eqiad.wmnet with OS bookworm * 15:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1176.eqiad.wmnet with OS bookworm * 15:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1154.eqiad.wmnet with OS bookworm * 15:16 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica-ulsfo and A:liberica * 15:13 vriley@cumin1003: START - Cookbook sre.dns.netbox * 15:13 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host dns4004.wikimedia.org with OS trixie * 15:12 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1188.eqiad.wmnet with OS bookworm * 15:09 taavi@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318203{{!}}Undeploy WP25EasterEggs (II) (T418134)]] (duration: 08m 56s) * 15:06 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica-magru and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 15:06 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1051.eqiad.wmnet with OS trixie * 15:06 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 15:05 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 15:05 taavi@deploy1003: taavi: Continuing with deployment * 15:04 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-ncredir (exit_code=0) rolling reboot on A:ncredir and A:ncredir * 15:04 taavi@deploy1003: taavi: Backport for [[gerrit:1318203{{!}}Undeploy WP25EasterEggs (II) (T418134)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:03 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1055.eqiad.wmnet with OS trixie * 15:02 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:00 taavi@deploy1003: Started scap sync-world: Backport for [[gerrit:1318203{{!}}Undeploy WP25EasterEggs (II) (T418134)]] * 14:59 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1047.eqiad.wmnet * 14:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1047.eqiad.wmnet * 14:58 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy (exit_code=0) rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 14:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-codfw * 14:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp2001.codfw.wmnet * 14:57 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp2001.codfw.wmnet * 14:57 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica-magru and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 14:57 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:56 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1055 * 14:56 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1055 * 14:55 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns2006.wikimedia.org with reason: host reimage * 14:55 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:55 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1055] - vriley@cumin1003" * 14:55 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1055] - vriley@cumin1003" * 14:54 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1047.eqiad.wmnet * 14:52 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1188.eqiad.wmnet with reason: host reimage * 14:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp2001.codfw.wmnet * 14:51 vriley@cumin1003: START - Cookbook sre.dns.netbox * 14:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp2001.codfw.wmnet * 14:50 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2016.codfw.wmnet * 14:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2016.codfw.wmnet * 14:50 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1189.eqiad.wmnet with OS bookworm * 14:48 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1188.eqiad.wmnet with reason: host reimage * 14:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1051.eqiad.wmnet with reason: host reimage * 14:46 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1182.eqiad.wmnet with OS bookworm * 14:45 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief2002.codfw.wmnet * 14:44 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns1006.wikimedia.org with reason: host reimage * 14:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2016.codfw.wmnet * 14:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1143.eqiad.wmnet with OS bookworm * 14:43 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2016.codfw.wmnet * 14:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2318-2331].codfw.wmnet * 14:43 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2318-2331].codfw.wmnet * 14:42 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching aqs[2002-2012].codfw.wmnet,aqs[1017-1027].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433026|T433026]] - eevans@cumin1003 * 14:41 cgoubert@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326306{{!}}Add placeholder $wmgRedisLockPassword (T366938 T427999)]] (duration: 06m 56s) * 14:41 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief2002.codfw.wmnet * 14:40 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief1002.eqiad.wmnet * 14:39 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1051.eqiad.wmnet with reason: host reimage * 14:38 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns2006.wikimedia.org with reason: host reimage * 14:37 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns1006.wikimedia.org with reason: host reimage * 14:37 cgoubert@deploy1003: cgoubert: Continuing with deployment * 14:36 cgoubert@deploy1003: cgoubert: Backport for [[gerrit:1326306{{!}}Add placeholder $wmgRedisLockPassword (T366938 T427999)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:36 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief1002.eqiad.wmnet * 14:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2318-2331].codfw.wmnet * 14:34 cgoubert@deploy1003: Started scap sync-world: Backport for [[gerrit:1326306{{!}}Add placeholder $wmgRedisLockPassword (T366938 T427999)]] * 14:29 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2318-2331].codfw.wmnet * 14:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2304-2317].codfw.wmnet * 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2304-2317].codfw.wmnet * 14:28 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1054 * 14:27 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1054 * 14:27 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:27 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1054] - vriley@cumin1003" * 14:27 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1054] - vriley@cumin1003" * 14:26 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1189.eqiad.wmnet with reason: host reimage * 14:26 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test2001.codfw.wmnet * 14:25 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test1001.eqiad.wmnet * 14:25 claime: Deploying wmgRedisLockPassword - [[phab:T366938|T366938]] [[phab:T427999|T427999]] * 14:24 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1051.eqiad.wmnet with OS trixie * 14:23 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 14:22 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1182.eqiad.wmnet with reason: host reimage * 14:22 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1051.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:22 vriley@cumin1003: START - Cookbook sre.dns.netbox * 14:22 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test2001.codfw.wmnet * 14:21 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test1001.eqiad.wmnet * 14:21 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 14:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2304-2317].codfw.wmnet * 14:19 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns4004.wikimedia.org with OS trixie * 14:19 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns2006.wikimedia.org with OS trixie * 14:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1143.eqiad.wmnet with reason: host reimage * 14:19 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns1006.wikimedia.org with OS trixie * 14:17 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1189.eqiad.wmnet with reason: host reimage * 14:15 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1182.eqiad.wmnet with reason: host reimage * 14:14 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1143.eqiad.wmnet with reason: host reimage * 14:13 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1051.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:12 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2304-2317].codfw.wmnet * 14:12 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2290-2303].codfw.wmnet * 14:12 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1049.eqiad.wmnet with OS trixie * 14:12 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2290-2303].codfw.wmnet * 14:12 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1051 * 14:11 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1051 * 14:11 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1170.eqiad.wmnet onto db1284.eqiad.wmnet * 14:11 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1170: Pool db1170.eqiad.wmnet in after cloning * 14:09 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 14:09 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:09 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1051] - vriley@cumin1003" * 14:09 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1051] - vriley@cumin1003" * 14:06 klausman@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:04 vriley@cumin1003: START - Cookbook sre.dns.netbox * 14:04 klausman@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2290-2303].codfw.wmnet * 14:02 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1189.eqiad.wmnet with OS bookworm * 14:02 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1188.eqiad.wmnet with OS bookworm * 14:01 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 14:00 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1182.eqiad.wmnet with OS bookworm * 14:00 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1143.eqiad.wmnet with OS bookworm * 13:59 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 13:57 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply * 13:57 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply * 13:56 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply * 13:56 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-ulsfo or A:cp-drmrs and A:cp - 9.2.15 upgrade ([[phab:T434620|T434620]]) * 13:56 cjd91: sudo -i cookbook sre.cdn.roll-upgrade-ats --query 'A:cp-ulsfo or A:cp-drmrs' --task-id [[phab:T434620|T434620]] --reason '9.2.15 upgrade' * 13:56 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply * 13:55 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply * 13:55 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply * 13:54 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 13:54 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 13:54 phuedx@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:53 phuedx@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics-external: apply * 13:52 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1049.eqiad.wmnet with reason: host reimage * 13:51 phuedx@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2290-2303].codfw.wmnet * 13:50 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1047.eqiad.wmnet * 13:50 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2276-2289].codfw.wmnet * 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2276-2289].codfw.wmnet * 13:49 phuedx@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics-external: apply * 13:49 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1046.eqiad.wmnet * 13:49 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1049.eqiad.wmnet with reason: host reimage * 13:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1046.eqiad.wmnet * 13:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-ncredir rolling reboot on A:ncredir and A:ncredir * 13:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 13:46 phuedx@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:44 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics-external: apply * 13:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1046.eqiad.wmnet * 13:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2276-2289].codfw.wmnet * 13:36 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1046.eqiad.wmnet * 13:36 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1045.eqiad.wmnet * 13:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1045.eqiad.wmnet * 13:34 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2276-2289].codfw.wmnet * 13:34 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1049.eqiad.wmnet with OS trixie * 13:34 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2262-2275].codfw.wmnet * 13:34 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2262-2275].codfw.wmnet * 13:32 Lucas_WMDE: UTC afternoon backport+config window done * 13:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1045.eqiad.wmnet * 13:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2262-2275].codfw.wmnet * 13:26 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1049.eqiad.wmnet with OS trixie * 13:26 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1049.eqiad.wmnet with OS trixie * 13:26 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1170: Pool db1170.eqiad.wmnet in after cloning * 13:23 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1049.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 13:23 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1045.eqiad.wmnet * 13:21 atsukoito: manually done sudo -i docker-registryctl --debug delete-tags 'docker-registry.discovery.wmnet/repos/data-engineering/airflow-dags:airflow-3.3.0-py3.11-2026-08-17-*' to remove incorrect tags * 13:19 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311141{{!}}viwiki: Set `noindex,nofollow` for User and User talk (T432311)]] (duration: 11m 11s) * 13:15 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2262-2275].codfw.wmnet * 13:14 lucaswerkmeister-wmde@deploy1003: ndkdd, lucaswerkmeister-wmde: Continuing with deployment * 13:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2248-2261].codfw.wmnet * 13:14 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2248-2261].codfw.wmnet * 13:13 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1049.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 13:12 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1049 * 13:11 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1049 * 13:10 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:10 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1049] - vriley@cumin1003" * 13:10 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1049] - vriley@cumin1003" * 13:10 lucaswerkmeister-wmde@deploy1003: ndkdd, lucaswerkmeister-wmde: Backport for [[gerrit:1311141{{!}}viwiki: Set `noindex,nofollow` for User and User talk (T432311)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1037.eqiad.wmnet * 13:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1037.eqiad.wmnet * 13:08 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1311141{{!}}viwiki: Set `noindex,nofollow` for User and User talk (T432311)]] * 13:06 vriley@cumin1003: START - Cookbook sre.dns.netbox * 13:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2248-2261].codfw.wmnet * 13:00 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1037.eqiad.wmnet * 12:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2248-2261].codfw.wmnet * 12:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2204-2215,2242-2243].codfw.wmnet * 12:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2204-2215,2242-2243].codfw.wmnet * 12:49 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324832{{!}}Migrate $wgFlaggedRevsTags from flaggedrevs.php to ext-FlaggedRevs.php]] (duration: 14m 02s) * 12:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint2001.codfw.wmnet * 12:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2204-2215,2242-2243].codfw.wmnet * 12:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint2001.codfw.wmnet * 12:41 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1037.eqiad.wmnet * 12:40 ladsgroup@deploy1003: Rolling back deployment * 12:37 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1324832{{!}}Migrate $wgFlaggedRevsTags from flaggedrevs.php to ext-FlaggedRevs.php]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:37 seanleong-wmde: Finished populateSitesTable for [bolwiki] ([[[phab:T429955|T429955]]]) * 12:35 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1324832{{!}}Migrate $wgFlaggedRevsTags from flaggedrevs.php to ext-FlaggedRevs.php]] * 12:35 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2204-2215,2242-2243].codfw.wmnet * 12:34 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2190-2203].codfw.wmnet * 12:34 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2190-2203].codfw.wmnet * 12:32 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1028.eqiad.wmnet * 12:32 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1028.eqiad.wmnet * 12:31 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint1001.eqiad.wmnet * 12:30 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1170: Depool db1170.eqiad.wmnet to then clone it to db1284.eqiad.wmnet - marostegui@cumin1003 * 12:28 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint1001.eqiad.wmnet * 12:28 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1170: Depool db1170.eqiad.wmnet to then clone it to db1284.eqiad.wmnet - marostegui@cumin1003 * 12:27 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1170.eqiad.wmnet onto db1284.eqiad.wmnet * 12:27 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db2901.codfw.wmnet * 12:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2190-2203].codfw.wmnet * 12:27 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 12:26 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1028.eqiad.wmnet * 12:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2190-2203].codfw.wmnet * 12:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2172-2179,2184-2189].codfw.wmnet * 12:18 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2172-2179,2184-2189].codfw.wmnet * 12:13 seanleong-wmde@deploy1003: mwscript-k8s job started: foreachwikiindblist wikidataclient extensions/Wikibase/lib/maintenance/populateSitesTable.php --force-protocol https # [[phab:T429955|T429955]] * 12:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2172-2179,2184-2189].codfw.wmnet * 12:09 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1028.eqiad.wmnet * 12:03 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 12:02 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db2901.codfw.wmnet * 12:02 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db2901.codfw.wmnet * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1027.eqiad.wmnet * 12:02 fceratto@cumin1003: END (ERROR) - Cookbook sre.dns.netbox (exit_code=97) * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1027.eqiad.wmnet * 12:02 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 12:02 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db2901.codfw.wmnet * 12:01 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2172-2179,2184-2189].codfw.wmnet * 12:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2158-2171].codfw.wmnet * 12:01 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2158-2171].codfw.wmnet * 11:58 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db1902.eqiad.wmnet * 11:58 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 11:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1027.eqiad.wmnet * 11:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2158-2171].codfw.wmnet * 11:51 jayme: updated calico to v3.30.7 on wikikube eqiad - [[phab:T427400|T427400]] * 11:50 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1027.eqiad.wmnet * 11:45 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2158-2171].codfw.wmnet * 11:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2144-2157].codfw.wmnet * 11:44 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2144-2157].codfw.wmnet * 11:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1058.eqiad.wmnet * 11:43 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1058.eqiad.wmnet * 11:43 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 11:43 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 11:43 marostegui@cumin1003: Removing db1153 from zarcillo [[phab:T434638|T434638]] * 11:42 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1153.eqiad.wmnet * 11:42 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:42 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1153.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 11:42 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1172.eqiad.wmnet onto db1286.eqiad.wmnet * 11:42 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1172: Pool db1172.eqiad.wmnet in after cloning * 11:42 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1153.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 11:41 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1902.eqiad.wmnet * 11:41 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=97) for new host db2901.codfw.wmnet * 11:41 fceratto@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host db2901.codfw.wmnet with OS trixie * 11:38 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'. * 11:38 marostegui@cumin1003: START - Cookbook sre.dns.netbox * 11:37 marostegui@dns1004: END - running authdns-update * 11:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1058.eqiad.wmnet * 11:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2144-2157].codfw.wmnet * 11:35 marostegui@dns1004: START - running authdns-update * 11:32 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1153.eqiad.wmnet * 11:32 marostegui@cumin1003: START - Cookbook sre.mysql.decommission * 11:28 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326249{{!}}ImagePage: move TOC element below file link (T332644)]] (duration: 09m 56s) * 11:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2144-2157].codfw.wmnet * 11:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2130-2143].codfw.wmnet * 11:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2130-2143].codfw.wmnet * 11:26 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1058.eqiad.wmnet * 11:23 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 11:22 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'. * 11:22 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'. * 11:22 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'. * 11:22 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326249{{!}}ImagePage: move TOC element below file link (T332644)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1057.eqiad.wmnet * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1057.eqiad.wmnet * 11:21 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 11:20 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 11:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2130-2143].codfw.wmnet * 11:19 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326249{{!}}ImagePage: move TOC element below file link (T332644)]] * 11:18 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply * 11:17 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply * 11:17 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply * 11:16 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 11:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1057.eqiad.wmnet * 11:15 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 11:14 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 11:13 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 11:13 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db2901.codfw.wmnet with OS trixie * 11:12 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db2901.codfw.wmnet - fceratto@cumin1003" * 11:12 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db2901.codfw.wmnet - fceratto@cumin1003" * 11:12 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db2901.codfw.wmnet on all recursors * 11:12 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db2901.codfw.wmnet on all recursors * 11:12 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:12 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db2901.codfw.wmnet - fceratto@cumin1003" * 11:12 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2130-2143].codfw.wmnet * 11:11 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2107-2115,2124-2129].codfw.wmnet * 11:11 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2107-2115,2124-2129].codfw.wmnet * 11:11 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 11:11 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 11:11 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply * 11:11 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 11:10 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 11:06 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db2901.codfw.wmnet - fceratto@cumin1003" * 11:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2107-2115,2124-2129].codfw.wmnet * 10:57 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1172: Pool db1172.eqiad.wmnet in after cloning * 10:54 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2107-2115,2124-2129].codfw.wmnet * 10:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2078,2087-2095,2102-2106].codfw.wmnet * 10:54 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2078,2087-2095,2102-2106].codfw.wmnet * 10:47 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply * 10:46 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply * 10:46 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply * 10:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2078,2087-2095,2102-2106].codfw.wmnet * 10:45 blake@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply * 10:38 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1057.eqiad.wmnet * 10:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1056.eqiad.wmnet * 10:38 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1056.eqiad.wmnet * 10:37 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2078,2087-2095,2102-2106].codfw.wmnet * 10:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2061-2062,2064-2065,2067-2077].codfw.wmnet * 10:36 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2061-2062,2064-2065,2067-2077].codfw.wmnet * 10:32 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1056.eqiad.wmnet * 10:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2061-2062,2064-2065,2067-2077].codfw.wmnet * 10:25 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:25 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db2901.codfw.wmnet * 10:25 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db1902.eqiad.wmnet * 10:25 fceratto@cumin1003: END (ERROR) - Cookbook sre.dns.netbox (exit_code=97) * 10:24 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:24 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1902.eqiad.wmnet * 10:24 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=97) for new host db1902.eqiad.wmnet * 10:24 fceratto@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host db1902.eqiad.wmnet with OS trixie * 10:24 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=93) for new host db1903.eqiad.wmnet * 10:24 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 10:20 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1056.eqiad.wmnet * 10:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2061-2062,2064-2065,2067-2077].codfw.wmnet * 10:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2038-2039,2041-2042,2044,2046,2049-2051,2055-2060].codfw.wmnet * 10:18 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2038-2039,2041-2042,2044,2046,2049-2051,2055-2060].codfw.wmnet * 10:15 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1055.eqiad.wmnet * 10:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1055.eqiad.wmnet * 10:13 Amir1: mwscript-k8s --dblist=all -- purgeUserOptions.php --login-age 5 uls-preferences ([[phab:T406724|T406724]]) * 10:11 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:10 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1903.eqiad.wmnet on all recursors * 10:10 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1903.eqiad.wmnet on all recursors * 10:10 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2038-2039,2041-2042,2044,2046,2049-2051,2055-2060].codfw.wmnet * 10:10 moritzm: installing unzip security updates * 10:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1055.eqiad.wmnet * 10:08 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=97) for new host db1901.eqiad.wmnet * 10:08 fceratto@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host db1901.eqiad.wmnet with OS trixie * 10:08 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:08 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 10:08 fceratto@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1903.eqiad.wmnet - fceratto@cumin1003" * 10:07 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db1902.eqiad.wmnet with OS trixie * 10:07 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1902.eqiad.wmnet - fceratto@cumin1003" * 10:07 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1902.eqiad.wmnet - fceratto@cumin1003" * 10:04 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply * 10:04 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326227{{!}}Enable desktop/native lazy loading everywhere (T148047)]] (duration: 07m 13s) * 10:03 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1902.eqiad.wmnet on all recursors * 10:03 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1902.eqiad.wmnet on all recursors * 10:03 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:03 blake@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply * 10:01 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1903.eqiad.wmnet - fceratto@cumin1003" * 10:01 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2038-2039,2041-2042,2044,2046,2049-2051,2055-2060].codfw.wmnet * 10:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2002,2005-2006,2011-2015,2017-2018,2033-2037].codfw.wmnet * 10:01 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2002,2005-2006,2011-2015,2017-2018,2033-2037].codfw.wmnet * 10:00 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:00 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 09:59 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 09:59 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1055.eqiad.wmnet * 09:58 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326227{{!}}Enable desktop/native lazy loading everywhere (T148047)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:56 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326227{{!}}Enable desktop/native lazy loading everywhere (T148047)]] * 09:53 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2002,2005-2006,2011-2015,2017-2018,2033-2037].codfw.wmnet * 09:50 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:48 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1903.eqiad.wmnet * 09:48 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:48 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1902.eqiad.wmnet * 09:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 09:44 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2002,2005-2006,2011-2015,2017-2018,2033-2037].codfw.wmnet * 09:43 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-codfw * 09:43 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1174.eqiad.wmnet onto db1288.eqiad.wmnet * 09:42 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1174: Pool db1174.eqiad.wmnet in after cloning * 09:40 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1172: Depool db1172.eqiad.wmnet to then clone it to db1286.eqiad.wmnet - marostegui@cumin1003 * 09:39 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1172: Depool db1172.eqiad.wmnet to then clone it to db1286.eqiad.wmnet - marostegui@cumin1003 * 09:39 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1172.eqiad.wmnet onto db1286.eqiad.wmnet * 09:33 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1175.eqiad.wmnet onto db1289.eqiad.wmnet * 09:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1175: Pool db1175.eqiad.wmnet in after cloning * 09:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1201.eqiad.wmnet onto db1287.eqiad.wmnet * 09:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1201: Pool db1201.eqiad.wmnet in after cloning * 09:28 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db1901.eqiad.wmnet with OS trixie * 09:27 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:27 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:27 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1901.eqiad.wmnet on all recursors * 09:27 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1901.eqiad.wmnet on all recursors * 09:26 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:26 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:26 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:15 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:15 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1901.eqiad.wmnet * 09:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw1001.wikimedia.org with OS trixie * 08:57 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1174: Pool db1174.eqiad.wmnet in after cloning * 08:54 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1054.eqiad.wmnet * 08:54 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1054.eqiad.wmnet * 08:48 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1054.eqiad.wmnet * 08:47 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1175: Pool db1175.eqiad.wmnet in after cloning * 08:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 08:46 marostegui@cumin1003: Removing db1152 from zarcillo [[phab:T434480|T434480]] * 08:46 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1152.eqiad.wmnet * 08:46 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:46 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1152.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 08:46 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1201: Pool db1201.eqiad.wmnet in after cloning * 08:46 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1152.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 08:46 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1054.eqiad.wmnet * 08:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1053.eqiad.wmnet * 08:43 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1053.eqiad.wmnet * 08:42 marostegui@cumin1003: START - Cookbook sre.dns.netbox * 08:38 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage * 08:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1053.eqiad.wmnet * 08:36 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1152.eqiad.wmnet * 08:36 marostegui@cumin1003: START - Cookbook sre.mysql.decommission * 08:35 phuedx: UTC morning backport window done * 08:35 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1053.eqiad.wmnet * 08:34 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1035.eqiad.wmnet * 08:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1035.eqiad.wmnet * 08:34 phuedx@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324725{{!}}EventStreamConfig: Mark product_metrics.web_base and .web_base_with_ip as Test Kitchen streams (T429898 T430322)]], [[gerrit:1313923{{!}}EventStreamConfig: Remove unused web_ui_scroll* streams (T415370)]], [[gerrit:1325546{{!}}EventStreamConfig: Remove Watchlist click stream (T434790)]] (duration: 12m 42s) * 08:33 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1281: Pool back * 08:32 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage * 08:31 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on db2209.codfw.wmnet with reason: Maintenance * 08:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2209: Maintenance needed * 08:30 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2209: Maintenance needed * 08:26 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1035.eqiad.wmnet * 08:26 phuedx@deploy1003: bearloga, phuedx: Continuing with deployment * 08:23 phuedx@deploy1003: bearloga, phuedx: Backport for [[gerrit:1324725{{!}}EventStreamConfig: Mark product_metrics.web_base and .web_base_with_ip as Test Kitchen streams (T429898 T430322)]], [[gerrit:1313923{{!}}EventStreamConfig: Remove unused web_ui_scroll* streams (T415370)]], [[gerrit:1325546{{!}}EventStreamConfig: Remove Watchlist click stream (T434790)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug * 08:21 phuedx@deploy1003: Started scap sync-world: Backport for [[gerrit:1324725{{!}}EventStreamConfig: Mark product_metrics.web_base and .web_base_with_ip as Test Kitchen streams (T429898 T430322)]], [[gerrit:1313923{{!}}EventStreamConfig: Remove unused web_ui_scroll* streams (T415370)]], [[gerrit:1325546{{!}}EventStreamConfig: Remove Watchlist click stream (T434790)]] * 08:19 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw1001.wikimedia.org with OS trixie * 08:18 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1035.eqiad.wmnet * 08:17 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1032.eqiad.wmnet * 08:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1032.eqiad.wmnet * 08:16 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1279: Pool back * 08:15 phuedx@deploy1003: Finished scap sync-world: Backport for [[gerrit:1216721{{!}}viwikivoyage: enable relatedarticle and pop-up (T405724)]] (duration: 39m 12s) * 08:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1032.eqiad.wmnet * 08:09 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1032.eqiad.wmnet * 08:08 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1031.eqiad.wmnet * 08:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1031.eqiad.wmnet * 08:03 godog: switch production to use dumps-nfs.w.o - [[phab:T432212|T432212]] * 08:02 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1031.eqiad.wmnet * 08:02 phuedx@deploy1003: nvdtn19, phuedx: Continuing with deployment * 08:00 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1201: Depool db1201.eqiad.wmnet to then clone it to db1287.eqiad.wmnet - marostegui@cumin1003 * 08:00 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1201: Depool db1201.eqiad.wmnet to then clone it to db1287.eqiad.wmnet - marostegui@cumin1003 * 08:00 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1201.eqiad.wmnet onto db1287.eqiad.wmnet * 08:00 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1057.eqiad.wmnet * 08:00 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:00 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1057.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:59 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1057.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:59 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1275: Pool back * 07:55 filippo@cumin1003: START - Cookbook sre.dns.netbox * 07:55 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1031.eqiad.wmnet * 07:52 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1030.eqiad.wmnet * 07:52 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1030.eqiad.wmnet * 07:52 phuedx@deploy1003: nvdtn19, phuedx: Backport for [[gerrit:1216721{{!}}viwikivoyage: enable relatedarticle and pop-up (T405724)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:51 tappof: bump space for prometheus k8s-aux in eqiad * 07:50 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1057.eqiad.wmnet * 07:48 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1281: Pool back * 07:47 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1281 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96108 and previous config saved to /var/cache/conftool/dbconfig/20260817-074749-marostegui.json * 07:46 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1030.eqiad.wmnet * 07:42 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1030.eqiad.wmnet * 07:41 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1174: Depool db1174.eqiad.wmnet to then clone it to db1288.eqiad.wmnet - marostegui@cumin1003 * 07:41 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1174: Depool db1174.eqiad.wmnet to then clone it to db1288.eqiad.wmnet - marostegui@cumin1003 * 07:41 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1174.eqiad.wmnet onto db1288.eqiad.wmnet * 07:40 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1029.eqiad.wmnet * 07:40 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm2001.wikimedia.org * 07:40 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1029.eqiad.wmnet * 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1056.eqiad.wmnet * 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1056.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:38 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1056.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:36 phuedx@deploy1003: Started scap sync-world: Backport for [[gerrit:1216721{{!}}viwikivoyage: enable relatedarticle and pop-up (T405724)]] * 07:36 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm2001.wikimedia.org * 07:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1279 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96104 and previous config saved to /var/cache/conftool/dbconfig/20260817-073542-marostegui.json * 07:34 filippo@cumin1003: START - Cookbook sre.dns.netbox * 07:34 slyngshede@dns1004: END - running authdns-update * 07:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1029.eqiad.wmnet * 07:33 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm-test1001.wikimedia.org * 07:32 slyngshede@dns1004: START - running authdns-update * 07:31 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1029.eqiad.wmnet * 07:31 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1279: Pool back * 07:30 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1279 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96102 and previous config saved to /var/cache/conftool/dbconfig/20260817-073038-marostegui.json * 07:29 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm-test1001.wikimedia.org * 07:29 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm1001.wikimedia.org * 07:28 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1056.eqiad.wmnet * 07:28 moritzm: extend the disk of ldap-rw1001 by 80G [[phab:T331699|T331699]] * 07:28 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1055.eqiad.wmnet * 07:28 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:28 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1055.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:27 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1055.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:26 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1044.eqiad.wmnet * 07:26 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1044.eqiad.wmnet * 07:25 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm1001.wikimedia.org * 07:24 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1175: Depool db1175.eqiad.wmnet to then clone it to db1289.eqiad.wmnet - marostegui@cumin1003 * 07:24 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1175: Depool db1175.eqiad.wmnet to then clone it to db1289.eqiad.wmnet - marostegui@cumin1003 * 07:24 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1175.eqiad.wmnet onto db1289.eqiad.wmnet * 07:22 filippo@cumin1003: START - Cookbook sre.dns.netbox * 07:20 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1044.eqiad.wmnet * 07:16 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1055.eqiad.wmnet * 07:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1054.eqiad.wmnet * 07:15 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:15 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1054.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:15 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1044.eqiad.wmnet * 07:15 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1054.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:13 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1275: Pool back * 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1275 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96098 and previous config saved to /var/cache/conftool/dbconfig/20260817-071225-marostegui.json * 07:10 filippo@cumin1003: START - Cookbook sre.dns.netbox * 07:05 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1043.eqiad.wmnet * 07:05 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1054.eqiad.wmnet * 07:05 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1051.eqiad.wmnet * 07:05 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:05 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1051.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:05 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1043.eqiad.wmnet * 07:04 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1051.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 06:59 filippo@cumin1003: START - Cookbook sre.dns.netbox * 06:59 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin1001.eqiad.wmnet * 06:59 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1043.eqiad.wmnet * 06:59 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin2001.codfw.wmnet * 06:55 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin2001.codfw.wmnet * 06:55 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1051.eqiad.wmnet * 06:54 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1049.eqiad.wmnet * 06:54 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:54 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1049.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 06:54 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1049.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 06:54 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin1001.eqiad.wmnet * 06:53 moritzm: installing apr-util security updates * 06:52 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1043.eqiad.wmnet * 06:49 filippo@cumin1003: START - Cookbook sre.dns.netbox * 06:41 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1049.eqiad.wmnet * 06:13 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit2003.wikimedia.org * 06:13 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet * 06:07 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet * 06:06 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit2003.wikimedia.org * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 47s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-16 == * 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 01m 03s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-15 == * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 41s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-14 == * 15:38 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-staging-master-eqiad * 15:38 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster1005.eqiad.wmnet * 15:38 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster1005.eqiad.wmnet * 15:35 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sretest2009.codfw.wmnet * 15:33 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster1005.eqiad.wmnet * 15:33 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster1005.eqiad.wmnet * 15:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster1004.eqiad.wmnet * 15:32 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster1004.eqiad.wmnet * 15:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host sretest2009.codfw.wmnet * 15:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster1004.eqiad.wmnet * 15:27 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster1004.eqiad.wmnet * 15:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster1003.eqiad.wmnet * 15:27 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster1003.eqiad.wmnet * 15:24 dancy@deploy1003: Finished scap sync-world: testing (duration: 03m 23s) * 15:22 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster1003.eqiad.wmnet * 15:22 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster1003.eqiad.wmnet * 15:22 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-staging-master-eqiad * 15:20 dancy@deploy1003: Started scap sync-world: testing * 15:20 dancy@deploy1003: Installation of scap version "4.280.2" completed for 3 hosts * 15:18 dancy@deploy1003: Installing scap version "4.280.2" for 3 host(s) * 15:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sretest2006.codfw.wmnet * 14:54 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host sretest2006.codfw.wmnet * 14:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sretest2003.codfw.wmnet * 14:39 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host sretest2003.codfw.wmnet * 13:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt-staging2001.codfw.wmnet * 13:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt-staging2001.codfw.wmnet * 13:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-staging-master-codfw * 13:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster2005.codfw.wmnet * 13:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster2005.codfw.wmnet * 13:05 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox-dev2003.codfw.wmnet * 13:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster2005.codfw.wmnet * 13:04 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster2005.codfw.wmnet * 13:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster2004.codfw.wmnet * 13:04 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster2004.codfw.wmnet * 13:01 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netbox-dev2003.codfw.wmnet * 12:59 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster2004.codfw.wmnet * 12:59 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster2004.codfw.wmnet * 12:59 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster2003.codfw.wmnet * 12:59 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster2003.codfw.wmnet * 12:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster2003.codfw.wmnet * 12:54 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster2003.codfw.wmnet * 12:54 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-staging-master-codfw * 12:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-staging-worker-eqiad * 12:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage1006.eqiad.wmnet * 12:52 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage1006.eqiad.wmnet * 12:46 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage1006.eqiad.wmnet * 12:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw1001.wikimedia.org with OS trixie * 12:45 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage1006.eqiad.wmnet * 12:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage1005.eqiad.wmnet * 12:45 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage1005.eqiad.wmnet * 12:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage1005.eqiad.wmnet * 12:36 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/ratelimit: apply * 12:35 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/ratelimit: apply * 12:35 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:35 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:33 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage1005.eqiad.wmnet * 12:33 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage1004.eqiad.wmnet * 12:33 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage1004.eqiad.wmnet * 12:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage1004.eqiad.wmnet * 12:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage1004.eqiad.wmnet * 12:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage1003.eqiad.wmnet * 12:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage1003.eqiad.wmnet * 12:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage * 12:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage1003.eqiad.wmnet * 12:17 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage * 12:14 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage1003.eqiad.wmnet * 12:14 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-staging-worker-eqiad * 12:03 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw1001.wikimedia.org with OS trixie * 12:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cuminunpriv1001.eqiad.wmnet * 11:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cuminunpriv1001.eqiad.wmnet * 11:27 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1004.wikimedia.org * 11:24 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1153 from dbctl [[phab:T434638|T434638]]', diff saved to https://phabricator.wikimedia.org/P96097 and previous config saved to /var/cache/conftool/dbconfig/20260814-112449-marostegui.json * 11:21 aokoth@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1004.wikimedia.org * 11:20 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 11:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-staging-worker-codfw * 11:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2004.codfw.wmnet * 11:17 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2004.codfw.wmnet * 11:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2004.codfw.wmnet * 11:10 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2004.codfw.wmnet * 11:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2003.codfw.wmnet * 11:10 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2003.codfw.wmnet * 11:03 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-worker1181.eqiad.wmnet with OS bookworm * 11:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2003.codfw.wmnet * 11:03 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2003.codfw.wmnet * 11:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2002.codfw.wmnet * 11:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2002.codfw.wmnet * 10:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1235.eqiad.wmnet with OS bookworm * 10:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock2003.codfw.wmnet * 10:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock2003.codfw.wmnet with OS trixie * 10:56 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2002.codfw.wmnet * 10:54 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1153.eqiad.wmnet with OS bookworm * 10:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2002.codfw.wmnet * 10:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2001.codfw.wmnet * 10:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2001.codfw.wmnet * 10:50 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 10:49 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1187.eqiad.wmnet with OS bookworm * 10:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2001.codfw.wmnet * 10:44 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2001.codfw.wmnet * 10:44 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-staging-worker-codfw * 10:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock2003.codfw.wmnet with reason: host reimage * 10:38 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock2003.codfw.wmnet with reason: host reimage * 10:35 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1153.eqiad.wmnet with reason: host reimage * 10:32 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1235.eqiad.wmnet with reason: host reimage * 10:29 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1187.eqiad.wmnet with reason: host reimage * 10:24 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1153.eqiad.wmnet with reason: host reimage * 10:23 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1235.eqiad.wmnet with reason: host reimage * 10:21 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1187.eqiad.wmnet with reason: host reimage * 10:16 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock2003.codfw.wmnet with OS trixie * 10:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install1005.wikimedia.org * 10:12 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 10:12 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 10:12 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 10:12 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 10:12 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:12 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 10:12 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 10:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install1005.wikimedia.org * 10:07 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1235.eqiad.wmnet with OS bookworm * 10:07 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1187.eqiad.wmnet with OS bookworm * 10:07 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1181.eqiad.wmnet with OS bookworm * 10:07 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1153.eqiad.wmnet with OS bookworm * 10:07 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install2005.wikimedia.org * 10:03 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 10:03 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2003.codfw.wmnet * 10:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1232.eqiad.wmnet with OS bookworm * 10:00 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install2005.wikimedia.org * 10:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install3004.wikimedia.org * 09:58 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 09:58 fceratto@cumin1003: Removing db1151 from zarcillo [[phab:T434538|T434538]] * 09:56 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1231.eqiad.wmnet with OS bookworm * 09:56 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 09:53 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 09:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install3004.wikimedia.org * 09:50 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1152.eqiad.wmnet with OS bookworm * 09:49 Dreamy_Jazz: `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260808000000" --end-timestamp="20260812120000" --sleep="5" --batch-size="50"` for [[phab:T434688|T434688]] * 09:48 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host centrallog2002.codfw.wmnet * 09:48 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install4004.wikimedia.org * 09:42 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1232.eqiad.wmnet with reason: host reimage * 09:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install4004.wikimedia.org * 09:41 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host centrallog2002.codfw.wmnet * 09:39 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install5004.wikimedia.org * 09:36 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1232.eqiad.wmnet with reason: host reimage * 09:36 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'. * 09:34 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'. * 09:33 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1231.eqiad.wmnet with reason: host reimage * 09:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install5004.wikimedia.org * 09:32 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host centrallog1002.eqiad.wmnet * 09:30 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install6003.wikimedia.org * 09:30 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1231.eqiad.wmnet with reason: host reimage * 09:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1152.eqiad.wmnet with reason: host reimage * 09:25 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host centrallog1002.eqiad.wmnet * 09:25 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1152.eqiad.wmnet with reason: host reimage * 09:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install6003.wikimedia.org * 09:22 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1232 * 09:22 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1232 * 09:22 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1232 * 09:22 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1232.eqiad.wmnet 25.53.64.10.in-addr.arpa 5.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:22 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host titan1001.eqiad.wmnet * 09:22 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1232.eqiad.wmnet 25.53.64.10.in-addr.arpa 5.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:22 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:22 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1232 - btullis@cumin1003" * 09:22 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1232 - btullis@cumin1003" * 09:17 btullis@cumin1003: START - Cookbook sre.dns.netbox * 09:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install7002.wikimedia.org * 09:17 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1232 * 09:16 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1231 * 09:16 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1231 * 09:14 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1231 * 09:14 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1231.eqiad.wmnet 24.53.64.10.in-addr.arpa 4.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host titan1001.eqiad.wmnet * 09:14 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1231.eqiad.wmnet 24.53.64.10.in-addr.arpa 4.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:14 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1231 - btullis@cumin1003" * 09:14 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1231 - btullis@cumin1003" * 09:10 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install7002.wikimedia.org * 09:09 btullis@cumin1003: START - Cookbook sre.dns.netbox * 09:08 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1231 * 09:08 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1152 * 09:08 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1152 * 09:06 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1152 * 09:06 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1152.eqiad.wmnet 16.53.64.10.in-addr.arpa 6.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:06 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1152.eqiad.wmnet 16.53.64.10.in-addr.arpa 6.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:06 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:06 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1152 - btullis@cumin1003" * 09:06 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1152 - btullis@cumin1003" * 09:03 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1151.eqiad.wmnet * 09:03 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:03 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1151.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 08:55 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host titan2001.codfw.wmnet * 08:55 btullis@cumin1003: START - Cookbook sre.dns.netbox * 08:54 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1151.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 08:54 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1152 * 08:53 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1232.eqiad.wmnet with OS bookworm * 08:53 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1231.eqiad.wmnet with OS bookworm * 08:53 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1152.eqiad.wmnet with OS bookworm * 08:51 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2209: Pool back * 08:51 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1201.eqiad.wmnet * 08:50 btullis@cumin1003: START - Cookbook sre.hosts.remove-downtime for an-worker1201.eqiad.wmnet * 08:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1229.eqiad.wmnet with OS bookworm * 08:47 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host titan2001.codfw.wmnet * 08:46 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping2004.codfw.wmnet * 08:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ping2004.codfw.wmnet * 08:40 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host an-worker1230.eqiad.wmnet with OS bookworm * 08:40 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 08:34 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1151.eqiad.wmnet * 08:34 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 08:31 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host titan1002.eqiad.wmnet * 08:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1229.eqiad.wmnet with reason: host reimage * 08:25 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host titan1002.eqiad.wmnet * 08:25 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1229.eqiad.wmnet with reason: host reimage * 08:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping1004.eqiad.wmnet * 08:21 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ping1004.eqiad.wmnet * 08:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1230.eqiad.wmnet with reason: host reimage * 08:12 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1230.eqiad.wmnet with reason: host reimage * 08:11 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1229 * 08:11 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1229 * 08:11 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1229 * 08:11 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1229.eqiad.wmnet 22.53.64.10.in-addr.arpa 2.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:11 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1229.eqiad.wmnet 22.53.64.10.in-addr.arpa 2.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:11 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:11 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1229 - btullis@cumin1003" * 08:11 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1229 - btullis@cumin1003" * 08:10 btullis@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1201.eqiad.wmnet with reason: Fixing a disk * 08:07 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host titan2002.codfw.wmnet * 08:05 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2209: Pool back * 08:05 btullis@cumin1003: START - Cookbook sre.dns.netbox * 08:00 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host titan2002.codfw.wmnet * 07:59 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1229 * 07:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1230 * 07:58 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1230 * 07:55 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1230 * 07:55 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1230.eqiad.wmnet 23.53.64.10.in-addr.arpa 3.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:55 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1230.eqiad.wmnet 23.53.64.10.in-addr.arpa 3.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:55 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:55 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1230 - btullis@cumin1003" * 07:55 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1230 - btullis@cumin1003" * 07:51 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host kubestagemaster2005.codfw.wmnet with OS trixie * 07:48 btullis@cumin1003: START - Cookbook sre.dns.netbox * 07:41 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1230 * 07:41 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1229.eqiad.wmnet with OS bookworm * 07:41 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1230.eqiad.wmnet with OS bookworm * 07:39 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1152 from dbctl [[phab:T434480|T434480]]', diff saved to https://phabricator.wikimedia.org/P96090 and previous config saved to /var/cache/conftool/dbconfig/20260814-073941-marostegui.json * 07:29 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on kubestagemaster2005.codfw.wmnet with reason: host reimage * 07:23 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on kubestagemaster2005.codfw.wmnet with reason: host reimage * 07:04 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host kubestagemaster2005.codfw.wmnet with OS trixie * 06:53 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1228.eqiad.wmnet with OS bookworm * 06:44 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1227.eqiad.wmnet with OS bookworm * 06:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1209.eqiad.wmnet with OS bookworm * 06:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1175.eqiad.wmnet with OS bookworm * 06:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1228.eqiad.wmnet with reason: host reimage * 06:27 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1228.eqiad.wmnet with reason: host reimage * 06:25 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1227.eqiad.wmnet with reason: host reimage * 06:21 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1227.eqiad.wmnet with reason: host reimage * 06:18 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1209.eqiad.wmnet with reason: host reimage * 06:14 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1209.eqiad.wmnet with reason: host reimage * 06:14 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1228 * 06:14 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1228 * 06:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1175.eqiad.wmnet with reason: host reimage * 06:12 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1228 * 06:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1228.eqiad.wmnet 20.53.64.10.in-addr.arpa 0.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:12 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1228.eqiad.wmnet 20.53.64.10.in-addr.arpa 0.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1228 - ryankemper@cumin2003" * 06:12 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1228 - ryankemper@cumin2003" * 06:09 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1175.eqiad.wmnet with reason: host reimage * 06:07 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 06:07 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1228 * 06:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1227 * 06:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1227 * 06:06 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1227 * 06:06 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1227.eqiad.wmnet 19.53.64.10.in-addr.arpa 9.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:06 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1227.eqiad.wmnet 19.53.64.10.in-addr.arpa 9.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:06 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:06 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1227 - ryankemper@cumin2003" * 06:06 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1227 - ryankemper@cumin2003" * 06:00 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 06:00 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1227 * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1209 * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1209 * 06:00 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1209 * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1209.eqiad.wmnet 15.53.64.10.in-addr.arpa 5.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:00 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1209.eqiad.wmnet 15.53.64.10.in-addr.arpa 5.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1209 - ryankemper@cumin2003" * 06:00 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1209 - ryankemper@cumin2003" * 05:54 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 05:54 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1209 * 05:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1175 * 05:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1175 * 05:52 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1175 * 05:52 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1175.eqiad.wmnet 17.53.64.10.in-addr.arpa 7.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 05:52 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1175.eqiad.wmnet 17.53.64.10.in-addr.arpa 7.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 05:52 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 05:52 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1175 - ryankemper@cumin2003" * 05:52 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1175 - ryankemper@cumin2003" * 05:49 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1228.eqiad.wmnet with OS bookworm * 05:49 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1227.eqiad.wmnet with OS bookworm * 05:48 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1209.eqiad.wmnet with OS bookworm * 05:47 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 05:47 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1175 * 05:47 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1175.eqiad.wmnet with OS bookworm * 05:09 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host apifeatureusage1001.eqiad.wmnet with OS bookworm * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 03s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:10 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 01:07 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 01:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 01:02 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 00:59 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1226.eqiad.wmnet with OS bookworm * 00:47 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1225.eqiad.wmnet with OS bookworm * 00:41 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1224.eqiad.wmnet with OS bookworm * 00:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1226.eqiad.wmnet with reason: host reimage * 00:31 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1226.eqiad.wmnet with reason: host reimage * 00:28 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1225.eqiad.wmnet with reason: host reimage * 00:25 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1225.eqiad.wmnet with reason: host reimage * 00:19 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1224.eqiad.wmnet with reason: host reimage * 00:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1226 * 00:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1226 * 00:17 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1226 * 00:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1226.eqiad.wmnet 23.36.64.10.in-addr.arpa 3.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:17 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1226.eqiad.wmnet 23.36.64.10.in-addr.arpa 3.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 00:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1226 - ryankemper@cumin2003" * 00:17 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1226 - ryankemper@cumin2003" * 00:16 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1224.eqiad.wmnet with reason: host reimage * 00:12 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 00:12 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1226 * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1225 * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1225 * 00:10 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1225 * 00:10 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1225.eqiad.wmnet 22.36.64.10.in-addr.arpa 2.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:10 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1225.eqiad.wmnet 22.36.64.10.in-addr.arpa 2.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:10 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 00:10 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1225 - ryankemper@cumin2003" * 00:10 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1225 - ryankemper@cumin2003" * 00:03 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 00:02 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1225 * 00:02 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1224 * 00:02 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1224 * 00:00 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1224 * 00:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1224.eqiad.wmnet 21.36.64.10.in-addr.arpa 1.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:00 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1224.eqiad.wmnet 21.36.64.10.in-addr.arpa 1.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 00:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1224 - ryankemper@cumin2003" * 00:00 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1224 - ryankemper@cumin2003" == 2026-08-13 == * 23:54 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1226.eqiad.wmnet with OS bookworm * 23:53 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1225.eqiad.wmnet with OS bookworm * 23:52 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 23:51 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1224 * 23:51 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1224.eqiad.wmnet with OS bookworm * 23:47 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1223.eqiad.wmnet with OS bookworm * 23:29 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1223.eqiad.wmnet with reason: host reimage * 23:24 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1223.eqiad.wmnet with reason: host reimage * 23:19 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325525{{!}}ve.ui.CodeMirror.less: ensure normal font style]] (duration: 11m 40s) * 23:16 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260807000000" --end-timestamp="20260808000000" --sleep="5" --batch-size="50"` for [[phab:T434688|T434688]] * 23:13 musikanimal@deploy1003: musikanimal: Continuing with deployment * 23:11 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1325525{{!}}ve.ui.CodeMirror.less: ensure normal font style]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1223 * 23:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1223 * 23:08 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1325525{{!}}ve.ui.CodeMirror.less: ensure normal font style]] * 23:07 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1223 * 23:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1223.eqiad.wmnet 20.36.64.10.in-addr.arpa 0.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:07 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1223.eqiad.wmnet 20.36.64.10.in-addr.arpa 0.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 23:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1223 - ryankemper@cumin2003" * 23:03 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1223 - ryankemper@cumin2003" * 23:01 sbassett: Deployed security updates for [[phab:T430596|T430596]], [[phab:T120386|T120386]] * 22:55 bking@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host kubestagemaster2005.codfw.wmnet * 22:55 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host kubestagemaster2005.codfw.wmnet with OS bookworm * 22:54 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 22:53 ryankemper@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 22:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on kubestagemaster2005.codfw.wmnet with reason: host reimage * 22:49 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1222.eqiad.wmnet with OS bookworm * 22:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on kubestagemaster2005.codfw.wmnet with reason: host reimage * 22:29 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1222.eqiad.wmnet with reason: host reimage * 22:26 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1222.eqiad.wmnet with reason: host reimage * 22:23 bking@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host aux-k8s-etcd2003.codfw.wmnet * 22:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd2003.codfw.wmnet with OS bookworm * 22:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host kubestagemaster2005.codfw.wmnet with OS bookworm * 22:23 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM kubestagemaster2005.codfw.wmnet - bking@cumin2003" * 22:23 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM kubestagemaster2005.codfw.wmnet - bking@cumin2003" * 22:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) kubestagemaster2005.codfw.wmnet on all recursors * 22:22 bking@cumin2003: START - Cookbook sre.dns.wipe-cache kubestagemaster2005.codfw.wmnet on all recursors * 22:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:22 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM kubestagemaster2005.codfw.wmnet - bking@cumin2003" * 22:22 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM kubestagemaster2005.codfw.wmnet - bking@cumin2003" * 22:17 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 22:13 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1223 * 22:12 bking@cumin2003: START - Cookbook sre.dns.netbox * 22:12 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host kubestagemaster2005.codfw.wmnet * 22:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1222 * 22:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1222 * 22:11 sbassett: Deployed security updates for [[phab:T429244|T429244]], [[phab:T434039|T434039]] * 22:11 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1222 * 22:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1222.eqiad.wmnet 19.36.64.10.in-addr.arpa 9.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:11 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1222.eqiad.wmnet 19.36.64.10.in-addr.arpa 9.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1222 - ryankemper@cumin2003" * 22:07 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1222 - ryankemper@cumin2003" * 22:05 robh@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-wdqs2001.codfw.wmnet with reason: updating firmware * 22:01 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 22:01 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1223.eqiad.wmnet with OS bookworm * 22:01 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1222 * 22:01 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1222.eqiad.wmnet with OS bookworm * 22:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1212.eqiad.wmnet with OS bookworm * 21:54 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324295{{!}}Revert "Lazily reject pre-fix parser-cache entries for noreferrer/noopener links" (T429090)]] (duration: 06m 42s) * 21:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd2003.codfw.wmnet with reason: host reimage * 21:53 bking@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host dse-k8s-etcd2001.codfw.wmnet * 21:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dse-k8s-etcd2001.codfw.wmnet with OS bookworm * 21:50 sbassett@deploy1003: sbassett, kharlan: Continuing with deployment * 21:49 sbassett@deploy1003: sbassett, kharlan: Backport for [[gerrit:1324295{{!}}Revert "Lazily reject pre-fix parser-cache entries for noreferrer/noopener links" (T429090)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:48 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aux-k8s-etcd2003.codfw.wmnet with reason: host reimage * 21:47 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1324295{{!}}Revert "Lazily reject pre-fix parser-cache entries for noreferrer/noopener links" (T429090)]] * 21:42 maryum: Deployed security patch for [[phab:T434549|T434549]] * 21:39 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1212.eqiad.wmnet with reason: host reimage * 21:34 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1212.eqiad.wmnet with reason: host reimage * 21:33 bking@cumin2003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd2003.codfw.wmnet with OS bookworm * 21:32 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM aux-k8s-etcd2003.codfw.wmnet - bking@cumin2003" * 21:32 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM aux-k8s-etcd2003.codfw.wmnet - bking@cumin2003" * 21:32 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) aux-k8s-etcd2003.codfw.wmnet on all recursors * 21:32 bking@cumin2003: START - Cookbook sre.dns.wipe-cache aux-k8s-etcd2003.codfw.wmnet on all recursors * 21:32 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:32 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM aux-k8s-etcd2003.codfw.wmnet - bking@cumin2003" * 21:31 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM aux-k8s-etcd2003.codfw.wmnet - bking@cumin2003" * 21:28 maryum: Deployed security patch for [[phab:T434619|T434619]] * 21:26 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:26 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host aux-k8s-etcd2003.codfw.wmnet * 21:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-etcd2001.codfw.wmnet with reason: host reimage * 21:20 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1212 * 21:20 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1212 * 21:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1212.eqiad.wmnet with OS bookworm * 21:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on dse-k8s-etcd2001.codfw.wmnet with reason: host reimage * 21:04 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325553{{!}}InstrumentConstructiveEdits: exclude mw-reverted as well (T431493)]] (duration: 06m 25s) * 20:59 kemayo@deploy1003: kemayo: Continuing with deployment * 20:59 kemayo@deploy1003: kemayo: Backport for [[gerrit:1325553{{!}}InstrumentConstructiveEdits: exclude mw-reverted as well (T431493)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host dse-k8s-etcd2001.codfw.wmnet with OS bookworm * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 20:57 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 20:57 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1325553{{!}}InstrumentConstructiveEdits: exclude mw-reverted as well (T431493)]] * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-etcd2001.codfw.wmnet on all recursors * 20:57 bking@cumin2003: START - Cookbook sre.dns.wipe-cache dse-k8s-etcd2001.codfw.wmnet on all recursors * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 20:57 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 20:54 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320163{{!}}Add configurable RestTermsOfServiceUrl (T428147)]] (duration: 21m 39s) * 20:53 bking@cumin2003: START - Cookbook sre.dns.netbox * 20:53 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host dse-k8s-etcd2001.codfw.wmnet * 20:50 samtar@deploy1003: samtar, milazg: Continuing with deployment * 20:35 samtar@deploy1003: samtar, milazg: Backport for [[gerrit:1320163{{!}}Add configurable RestTermsOfServiceUrl (T428147)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:33 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1320163{{!}}Add configurable RestTermsOfServiceUrl (T428147)]] * 20:30 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325549{{!}}Deploy PRV to several LC wikis (T423785)]] (duration: 06m 57s) * 20:26 arlolra@deploy1003: arlolra: Continuing with deployment * 20:25 arlolra@deploy1003: arlolra: Backport for [[gerrit:1325549{{!}}Deploy PRV to several LC wikis (T423785)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:24 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 20:23 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1325549{{!}}Deploy PRV to several LC wikis (T423785)]] * 20:21 ariel@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319906{{!}}Remove boilerplate language from wmf-rest and wmf-math API modules (T433736)]] (duration: 13m 54s) * 20:14 ariel@deploy1003: ariel: Continuing with deployment * 20:11 ariel@deploy1003: ariel: Backport for [[gerrit:1319906{{!}}Remove boilerplate language from wmf-rest and wmf-math API modules (T433736)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 ariel@deploy1003: Started scap sync-world: Backport for [[gerrit:1319906{{!}}Remove boilerplate language from wmf-rest and wmf-math API modules (T433736)]] * 19:58 robh@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-wdqs2001.codfw.wmnet with reason: updating firmware * 19:54 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325551{{!}}Render the focused module view as a full-screen page (T433896)]] (duration: 30m 37s) * 19:52 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching aqs[2001,1016]*: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 19:44 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching aqs[2001,1016]*: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 19:42 musikanimal@deploy1003: musikanimal: Continuing with deployment * 19:41 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1325551{{!}}Render the focused module view as a full-screen page (T433896)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:34 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: quash java safepoint logspam - bking@cumin2003 - [[phab:T434685|T434685]] * 19:34 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 19:34 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 19:24 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1325551{{!}}Render the focused module view as a full-screen page (T433896)]] * 19:20 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 19:20 jhancock@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin1003" * 19:18 jhancock@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin1003" * 19:14 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325548{{!}}Enable image lazy loading on desktop in group1 (T148047)]] (duration: 07m 43s) * 19:10 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 19:10 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1325548{{!}}Enable image lazy loading on desktop in group1 (T148047)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:07 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1325548{{!}}Enable image lazy loading on desktop in group1 (T148047)]] * 19:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1166.eqiad.wmnet onto db1280.eqiad.wmnet * 19:03 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 19:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1166: Pool db1166.eqiad.wmnet in after cloning * 18:59 jhancock@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 18:58 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1234.eqiad.wmnet with OS bookworm * 18:52 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 18:51 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:49 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:46 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 18:46 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 18:44 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=dns3004.* * 18:39 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1221.eqiad.wmnet with OS bookworm * 18:36 inflatador: [bking@ganeti2048] ~$ sudo gnt-instance replace-disks -n ganeti2030.codfw.wmnet aux-k8s-worker2002.codfw.wmnet [[phab:T434681|T434681]] * 18:35 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 18:29 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1220.eqiad.wmnet with OS bookworm * 18:29 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1234.eqiad.wmnet with reason: host reimage * 18:27 bking@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host dse-k8s-etcd2001.codfw.wmnet * 18:27 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-etcd2001.codfw.wmnet on all recursors * 18:27 bking@cumin2003: START - Cookbook sre.dns.wipe-cache dse-k8s-etcd2001.codfw.wmnet on all recursors * 18:27 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:27 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 18:27 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 18:25 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1234.eqiad.wmnet with reason: host reimage * 18:25 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 18:19 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1221.eqiad.wmnet with reason: host reimage * 18:17 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1166: Pool db1166.eqiad.wmnet in after cloning * 18:15 brennen@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 18:15 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1221.eqiad.wmnet with reason: host reimage * 18:14 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:12 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:12 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 18:12 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-etcd2001.codfw.wmnet on all recursors * 18:12 bking@cumin2003: START - Cookbook sre.dns.wipe-cache dse-k8s-etcd2001.codfw.wmnet on all recursors * 18:12 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:12 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 18:12 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 18:10 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: quash java safepoint logspam - bking@cumin2003 - [[phab:T434685|T434685]] * 18:10 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1234 * 18:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1234 * 18:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1220.eqiad.wmnet with reason: host reimage * 18:07 brennen: 1.47.0-wmf.15 train status ([[phab:T430834|T430834]]) - no current blockers, rolling to all wikis * 18:07 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1234 * 18:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1234.eqiad.wmnet 10.36.64.10.in-addr.arpa 0.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:07 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:07 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1234.eqiad.wmnet 10.36.64.10.in-addr.arpa 0.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1234 - ryankemper@cumin2003" * 18:07 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1234 - ryankemper@cumin2003" * 18:05 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1220.eqiad.wmnet with reason: host reimage * 18:04 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host dse-k8s-etcd2001.codfw.wmnet * 18:04 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 18:03 dancy@deploy1003: Installation of scap version "4.280.1" completed for 3 hosts * 18:02 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:02 inflatador: bking@dse-k8s-etcd2002 etcdctl member remove $<nowiki>{</nowiki>UUID of dse-k8s-etcd2001<nowiki>}</nowiki> [[phab:T434681|T434681]] [[phab:T434793|T434793]] * 18:01 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 18:01 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1234 * 18:01 dancy@deploy1003: Installing scap version "4.280.1" for 3 host(s) * 18:01 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1221 * 18:01 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1221 * 18:00 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1221 * 18:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1221.eqiad.wmnet 18.36.64.10.in-addr.arpa 8.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:00 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1221.eqiad.wmnet 18.36.64.10.in-addr.arpa 8.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1221 - ryankemper@cumin2003" * 17:59 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:58 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 17:58 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:57 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 17:57 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:56 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1221 - ryankemper@cumin2003" * 17:53 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 17:52 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 17:52 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 17:51 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1221 * 17:51 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1220 * 17:51 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1220 * 17:51 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:51 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:51 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1220 * 17:51 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1220.eqiad.wmnet 11.36.64.10.in-addr.arpa 1.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:51 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1220.eqiad.wmnet 11.36.64.10.in-addr.arpa 1.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:51 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1220 - ryankemper@cumin2003" * 17:50 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1220 - ryankemper@cumin2003" * 17:47 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:47 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 17:46 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1234.eqiad.wmnet with OS bookworm * 17:46 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 17:45 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1221.eqiad.wmnet with OS bookworm * 17:45 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1220 * 17:45 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1220.eqiad.wmnet with OS bookworm * 17:43 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 17:41 bking@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host dse-k8s-etcd2001.codfw.wmnet * 17:41 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host dse-k8s-etcd2001.codfw.wmnet * 17:40 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:40 inflatador: bking@ganeti2048] `sudo gnt-instance remove --force --ignore-failures --shutdown-timeout=0` on non-DRBD VMs [[phab:T434681|T434681]] * 17:40 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:39 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:38 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:36 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:32 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:32 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:28 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 17:26 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:26 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:24 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:23 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:21 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1218.eqiad.wmnet with OS bookworm * 17:20 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:20 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:19 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1179.eqiad.wmnet with OS bookworm * 17:18 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1150.eqiad.wmnet with OS bookworm * 17:18 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns3004.wikimedia.org with OS trixie * 17:15 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:14 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:13 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:12 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:12 swfrench@deploy1003: Finished scap sync-world: Helmfile-only deployment for mediawiki chart bump - [[phab:T427666|T427666]] (duration: 03m 03s) * 17:09 swfrench@deploy1003: Started scap sync-world: Helmfile-only deployment for mediawiki chart bump - [[phab:T427666|T427666]] * 17:02 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:01 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:01 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1218.eqiad.wmnet with reason: host reimage * 17:01 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:01 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:00 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324810{{!}}deployment-info.php: Report dbname and branch for the requested wiki (T434726)]] (duration: 06m 52s) * 16:58 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 16:57 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1179.eqiad.wmnet with reason: host reimage * 16:56 dancy@deploy1003: dancy: Continuing with deployment * 16:56 dancy@deploy1003: dancy: Backport for [[gerrit:1324810{{!}}deployment-info.php: Report dbname and branch for the requested wiki (T434726)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1150.eqiad.wmnet with reason: host reimage * 16:53 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324810{{!}}deployment-info.php: Report dbname and branch for the requested wiki (T434726)]] * 16:51 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1166: Depool db1166.eqiad.wmnet to then clone it to db1280.eqiad.wmnet - cwilliams@cumin1003 * 16:50 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1166: Depool db1166.eqiad.wmnet to then clone it to db1280.eqiad.wmnet - cwilliams@cumin1003 * 16:50 cwilliams@cumin1003: START - Cookbook sre.mysql.clone of db1166.eqiad.wmnet onto db1280.eqiad.wmnet * 16:49 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1218.eqiad.wmnet with reason: host reimage * 16:48 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1179.eqiad.wmnet with reason: host reimage * 16:47 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1150.eqiad.wmnet with reason: host reimage * 16:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.upgrade (exit_code=0) for 1 hosts * 16:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2220: Upgrade of db2220.codfw.wmnet completed * 16:38 swfrench@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 16:38 swfrench-wmf: kubectl delete node kubestagemaster2005.codfw.wmnet - [[phab:T434681|T434681]] * 16:34 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1218 * 16:34 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1218 * 16:34 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1218.eqiad.wmnet with OS bookworm * 16:34 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1179 * 16:34 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1179 * 16:33 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1179.eqiad.wmnet with OS bookworm * 16:32 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1150 * 16:32 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1150 * 16:31 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1150.eqiad.wmnet with OS bookworm * 16:24 dancy@deploy1003: Installation of scap version "4.280.0" completed for 3 hosts * 16:22 dancy@deploy1003: Installing scap version "4.280.0" for 3 host(s) * 16:20 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: quash java safepoint logspam - bking@cumin2003 - [[phab:T434685|T434685]] * 16:14 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns3004.wikimedia.org with reason: host reimage * 16:08 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 16:07 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 16:07 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns3004.wikimedia.org with reason: host reimage * 16:05 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 16:04 swfrench@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 16:00 swfrench@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 15:59 swfrench@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 15:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Upgrade of db2220.codfw.wmnet completed * 15:48 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2220: Upgrading db2220.codfw.wmnet * 15:48 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2220: Upgrading db2220.codfw.wmnet * 15:48 cwilliams@cumin1003: START - Cookbook sre.mysql.upgrade for 1 hosts * 15:46 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns3004.wikimedia.org with OS trixie * 15:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2220 [[phab:T434802|T434802]]', diff saved to https://phabricator.wikimedia.org/P96079 and previous config saved to /var/cache/conftool/dbconfig/20260813-154624-cwilliams.json * 15:45 cdobbins@cumin1003: conftool action : set/pooled=no; selector: name=dns3004.* * 15:44 cjd91: depooling dns3004 to reimage to trixie * 15:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2159 to s7 primary [[phab:T434802|T434802]]', diff saved to https://phabricator.wikimedia.org/P96078 and previous config saved to /var/cache/conftool/dbconfig/20260813-154405-cwilliams.json * 15:43 cezmunsta: Starting s7 codfw failover from db2220 to db2159 - [[phab:T434802|T434802]] * 15:41 cgoubert@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: Dragonfly supernodes reboot (duration: 09m 42s) * 15:41 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dragonfly-supernode2001.codfw.wmnet * 15:39 inflatador: bking@ganeti2048] ~$ sudo gnt-node failover -f --ignore-consistency ganeti2046.codfw.wmnet [[phab:T434681|T434681]] * 15:39 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 15:39 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 15:39 swfrench@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 15:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2159 with weight 0 [[phab:T434802|T434802]]', diff saved to https://phabricator.wikimedia.org/P96077 and previous config saved to /var/cache/conftool/dbconfig/20260813-153806-cwilliams.json * 15:37 swfrench@dns1004: END - running authdns-update * 15:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 30 hosts with reason: Primary switchover s7 [[phab:T434802|T434802]] * 15:37 cgoubert@cumin2003: START - Cookbook sre.hosts.reboot-single for host dragonfly-supernode2001.codfw.wmnet * 15:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dragonfly-supernode1001.eqiad.wmnet * 15:35 swfrench@dns1004: START - running authdns-update * 15:32 cgoubert@cumin2003: START - Cookbook sre.hosts.reboot-single for host dragonfly-supernode1001.eqiad.wmnet * 15:32 cgoubert@deploy1003: Locking from deployment [ALL REPOSITORIES]: Dragonfly supernodes reboot * 15:30 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-master-codfw * 15:30 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2005.codfw.wmnet * 15:30 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2005.codfw.wmnet * 15:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1169.eqiad.wmnet onto db1277.eqiad.wmnet * 15:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1169: Pool db1169.eqiad.wmnet in after cloning * 15:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2005.codfw.wmnet * 15:23 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2005.codfw.wmnet * 15:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2004.codfw.wmnet * 15:23 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2004.codfw.wmnet * 15:18 inflatador: bking@ganeti2048 sudo gnt-node failover -f ganeti2046.codfw.wmnet [[phab:T434681|T434681]] * 15:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2004.codfw.wmnet * 15:17 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2004.codfw.wmnet * 15:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2003.codfw.wmnet * 15:17 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2003.codfw.wmnet * 15:15 cdobbins@dns1004: END - running authdns-update * 15:13 cdobbins@dns1004: START - running authdns-update * 15:10 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1217.eqiad.wmnet with OS bookworm * 15:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2003.codfw.wmnet * 15:10 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2003.codfw.wmnet * 15:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2002.codfw.wmnet * 15:10 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2002.codfw.wmnet * 15:10 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: quash java safepoint logspam - bking@cumin2003 - [[phab:T434685|T434685]] * 15:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1216.eqiad.wmnet with OS bookworm * 15:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2002.codfw.wmnet * 15:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2002.codfw.wmnet * 15:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2001.codfw.wmnet * 15:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2001.codfw.wmnet * 15:01 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1042.eqiad.wmnet * 15:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1042.eqiad.wmnet * 15:00 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325455{{!}}Api: Use correct query when continue prop=categories (T433922)]] (duration: 09m 47s) * 14:59 bking@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host dse-k8s-etcd2004.codfw.wmnet * 14:58 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-etcd2004.codfw.wmnet on all recursors * 14:58 bking@cumin2003: START - Cookbook sre.dns.wipe-cache dse-k8s-etcd2004.codfw.wmnet on all recursors * 14:58 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:58 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM dse-k8s-etcd2004.codfw.wmnet - bking@cumin2003" * 14:58 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM dse-k8s-etcd2004.codfw.wmnet - bking@cumin2003" * 14:55 zabe@deploy1003: zabe: Continuing with deployment * 14:53 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-etcd2004.codfw.wmnet on all recursors * 14:53 bking@cumin2003: START - Cookbook sre.dns.wipe-cache dse-k8s-etcd2004.codfw.wmnet on all recursors * 14:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:53 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2004.codfw.wmnet - bking@cumin2003" * 14:53 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2004.codfw.wmnet - bking@cumin2003" * 14:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2001.codfw.wmnet * 14:52 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2001.codfw.wmnet * 14:52 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-master-codfw * 14:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2001.codfw.wmnet * 14:52 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2001.codfw.wmnet * 14:52 zabe@deploy1003: zabe: Backport for [[gerrit:1325455{{!}}Api: Use correct query when continue prop=categories (T433922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:50 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1325455{{!}}Api: Use correct query when continue prop=categories (T433922)]] * 14:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1217.eqiad.wmnet with reason: host reimage * 14:48 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:48 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host dse-k8s-etcd2004.codfw.wmnet * 14:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-master-eqiad * 14:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1006.eqiad.wmnet * 14:44 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1006.eqiad.wmnet * 14:44 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1216.eqiad.wmnet with reason: host reimage * 14:40 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1217.eqiad.wmnet with reason: host reimage * 14:39 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1169: Pool db1169.eqiad.wmnet in after cloning * 14:39 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1216.eqiad.wmnet with reason: host reimage * 14:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl1005.eqiad.wmnet * 14:32 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl1005.eqiad.wmnet * 14:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1004.eqiad.wmnet * 14:32 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1004.eqiad.wmnet * 14:31 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus2008.codfw.wmnet * 14:31 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor1003.eqiad.wmnet * 14:29 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1042.eqiad.wmnet * 14:27 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor1003.eqiad.wmnet * 14:26 moritzm: installing Django security updates * 14:26 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling reboot on A:wikidough * 14:25 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1217.eqiad.wmnet with OS bookworm * 14:25 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1216.eqiad.wmnet with OS bookworm * 14:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl1004.eqiad.wmnet * 14:24 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl1004.eqiad.wmnet * 14:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1003.eqiad.wmnet * 14:24 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1003.eqiad.wmnet * 14:23 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus2008.codfw.wmnet * 14:23 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus1008.eqiad.wmnet * 14:23 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor-dev2001.codfw.wmnet * 14:22 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus2006.codfw.wmnet * 14:19 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor-dev2001.codfw.wmnet * 14:18 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1042.eqiad.wmnet * 14:17 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.* * 14:16 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl1003.eqiad.wmnet * 14:16 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl1003.eqiad.wmnet * 14:16 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1002.eqiad.wmnet * 14:16 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1002.eqiad.wmnet * 14:15 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus1008.eqiad.wmnet * 14:14 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus1006.eqiad.wmnet * 14:12 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus2006.codfw.wmnet * 14:12 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor2003.codfw.wmnet * 14:11 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus2007.codfw.wmnet * 14:11 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1041.eqiad.wmnet * 14:11 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1041.eqiad.wmnet * 14:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl1002.eqiad.wmnet * 14:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl1002.eqiad.wmnet * 14:09 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-master-eqiad * 14:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on apifeatureusage1001.eqiad.wmnet with reason: host reimage * 14:08 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1215.eqiad.wmnet with OS bookworm * 14:08 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor2003.codfw.wmnet * 14:08 jayme: updated calico to v3.30.7 on wikikube codfw [[phab:T427400|T427400]] * 14:07 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1149.eqiad.wmnet with OS bookworm * 14:06 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'. * 14:06 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1041.eqiad.wmnet * 14:04 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus1006.eqiad.wmnet * 14:03 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus2007.codfw.wmnet * 14:03 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus2005.codfw.wmnet * 14:03 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus1007.eqiad.wmnet * 14:03 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1002.eqiad.wmnet * 14:02 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on apifeatureusage1001.eqiad.wmnet with reason: host reimage * 14:02 cgoubert@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=helm-charts.*,name=eqiad * 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host chartmuseum1001.eqiad.wmnet * 14:00 moritzm: installing libxml2 security updates * 13:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1214.eqiad.wmnet with OS bookworm * 13:59 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns5003.* * 13:59 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1041.eqiad.wmnet * 13:59 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow1002.eqiad.wmnet * 13:58 cmooney@dns3003: END - running authdns-update * 13:57 cgoubert@cumin2003: START - Cookbook sre.hosts.reboot-single for host chartmuseum1001.eqiad.wmnet * 13:57 cgoubert@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=helm-charts.*,name=eqiad * 13:57 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: cloudelastic cluster restart - bking@cumin2003 * 13:57 cgoubert@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=helm-charts.*,name=codfw * 13:56 cmooney@dns3003: START - running authdns-update * 13:56 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'. * 13:56 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns5003.*,service=authdns-update * 13:55 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host chartmuseum2001.codfw.wmnet * 13:55 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus1007.eqiad.wmnet * 13:55 cmooney@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dns5003.wikimedia.org * 13:55 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus1005.eqiad.wmnet * 13:51 cgoubert@cumin2003: START - Cookbook sre.hosts.reboot-single for host chartmuseum2001.codfw.wmnet * 13:51 cgoubert@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=helm-charts.*,name=codfw * 13:51 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus2005.codfw.wmnet * 13:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host apifeatureusage1001.eqiad.wmnet with OS bookworm * 13:50 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus7002.magru.wmnet * 13:50 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts lvs1015.eqiad.wmnet * 13:50 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:50 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1015.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:49 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1015.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:49 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325480{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] (duration: 06m 39s) * 13:47 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1215.eqiad.wmnet with reason: host reimage * 13:46 cmooney@cumin1003: START - Cookbook sre.hosts.reboot-single for host dns5003.wikimedia.org * 13:46 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1003.eqiad.wmnet * 13:45 cmooney@cumin1003: conftool action : set/pooled=no; selector: name=dns5003.* * 13:45 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1040.eqiad.wmnet * 13:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1040.eqiad.wmnet * 13:45 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 13:45 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus1005.eqiad.wmnet * 13:45 stran@deploy1003: stran: Continuing with deployment * 13:44 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus7002.magru.wmnet * 13:44 stran@deploy1003: stran: Backport for [[gerrit:1325480{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:44 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus6002.drmrs.wmnet * 13:43 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260805000000" --end-timestamp="20260806000000" --sleep="5" --batch-size="10"` for [[phab:T434688|T434688]] * 13:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1149.eqiad.wmnet with reason: host reimage * 13:42 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1325480{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] * 13:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow1003.eqiad.wmnet * 13:40 sukhe@cumin1003: START - Cookbook sre.hosts.decommission for hosts lvs1015.eqiad.wmnet * 13:40 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts lvs1014.eqiad.wmnet * 13:40 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:40 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1014.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:40 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1040.eqiad.wmnet * 13:40 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1014.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:39 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1214.eqiad.wmnet with reason: host reimage * 13:38 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2004.codfw.wmnet * 13:38 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus6002.drmrs.wmnet * 13:38 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus5003.eqsin.wmnet * 13:35 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 13:35 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1149.eqiad.wmnet with reason: host reimage * 13:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1215.eqiad.wmnet with reason: host reimage * 13:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow2004.codfw.wmnet * 13:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1214.eqiad.wmnet with reason: host reimage * 13:33 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1040.eqiad.wmnet * 13:31 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus5003.eqsin.wmnet * 13:31 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: cloudelastic cluster restart - bking@cumin2003 * 13:31 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus4003.ulsfo.wmnet * 13:30 sukhe@cumin1003: START - Cookbook sre.hosts.decommission for hosts lvs1014.eqiad.wmnet * 13:30 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts lvs1013.eqiad.wmnet * 13:30 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:30 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1013.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:30 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1013.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:27 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1039.eqiad.wmnet * 13:27 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1039.eqiad.wmnet * 13:26 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1167.eqiad.wmnet onto db1281.eqiad.wmnet * 13:26 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1167: Pool db1167.eqiad.wmnet in after cloning * 13:25 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus4003.ulsfo.wmnet * 13:25 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260802000000" --end-timestamp="20260803000000" --sleep="5" --batch-size="10"` for [[phab:T434688|T434688]] * 13:24 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 13:24 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus3004.esams.wmnet * 13:24 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1274: New host * 13:24 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325476{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] (duration: 07m 19s) * 13:24 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260804000000" --end-timestamp="20260805000000" --sleep="5" --batch-size="10"` for [[phab:T434688|T434688]] * 13:24 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260803000000" --end-timestamp="20260804000000" --sleep="5" --batch-size="10"` for [[phab:T434688|T434688]] * 13:23 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2003.codfw.wmnet * 13:22 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1039.eqiad.wmnet * 13:20 sukhe@cumin1003: START - Cookbook sre.hosts.decommission for hosts lvs1013.eqiad.wmnet * 13:20 stran@deploy1003: stran: Continuing with deployment * 13:19 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow2003.codfw.wmnet * 13:19 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling reboot on A:wikidough * 13:19 stran@deploy1003: stran: Backport for [[gerrit:1325476{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:18 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus3004.esams.wmnet * 13:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1215.eqiad.wmnet with OS bookworm * 13:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1214.eqiad.wmnet with OS bookworm * 13:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1149.eqiad.wmnet with OS bookworm * 13:17 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1325476{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] * 13:11 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324731{{!}}prv: Enable parsoid rendering for 5 wikisource wikis]] (duration: 07m 26s) * 13:11 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1039.eqiad.wmnet * 13:11 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1036.eqiad.wmnet * 13:11 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1036.eqiad.wmnet * 13:07 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow3004.esams.wmnet * 13:07 jgiannelos@deploy1003: jgiannelos: Continuing with deployment * 13:06 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1233.eqiad.wmnet with OS bookworm * 13:06 jgiannelos@deploy1003: jgiannelos: Backport for [[gerrit:1324731{{!}}prv: Enable parsoid rendering for 5 wikisource wikis]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:04 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1324731{{!}}prv: Enable parsoid rendering for 5 wikisource wikis]] * 13:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow3004.esams.wmnet * 13:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1036.eqiad.wmnet * 13:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1210.eqiad.wmnet with OS bookworm * 13:01 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1036.eqiad.wmnet * 13:00 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1052.eqiad.wmnet * 13:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1052.eqiad.wmnet * 12:58 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow4003.ulsfo.wmnet * 12:55 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1211.eqiad.wmnet with OS bookworm * 12:55 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1052.eqiad.wmnet * 12:52 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow4003.ulsfo.wmnet * 12:51 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1052.eqiad.wmnet * 12:49 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1169: Depool db1169.eqiad.wmnet to then clone it to db1277.eqiad.wmnet - cwilliams@cumin1003 * 12:46 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1169: Depool db1169.eqiad.wmnet to then clone it to db1277.eqiad.wmnet - cwilliams@cumin1003 * 12:46 cwilliams@cumin1003: START - Cookbook sre.mysql.clone of db1169.eqiad.wmnet onto db1277.eqiad.wmnet * 12:45 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:45 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns record for deleted IP reservations eqsin lvs vlan ints - cmooney@cumin1003" * 12:44 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns record for deleted IP reservations eqsin lvs vlan ints - cmooney@cumin1003" * 12:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1051.eqiad.wmnet * 12:43 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1051.eqiad.wmnet * 12:42 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1210.eqiad.wmnet with reason: host reimage * 12:41 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1167: Pool db1167.eqiad.wmnet in after cloning * 12:40 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 12:39 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1274: New host * 12:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Added db1274', diff saved to https://phabricator.wikimedia.org/P96062 and previous config saved to /var/cache/conftool/dbconfig/20260813-123907-cwilliams.json * 12:38 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast7002.wikimedia.org * 12:38 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1233.eqiad.wmnet with reason: host reimage * 12:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1051.eqiad.wmnet * 12:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1211.eqiad.wmnet with reason: host reimage * 12:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1233.eqiad.wmnet with reason: host reimage * 12:32 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1051.eqiad.wmnet * 12:32 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast7002.wikimedia.org * 12:32 marostegui: Drop SecurePoll tables from closed wikis [[phab:T423128|T423128]] * 12:32 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow5003.eqsin.wmnet * 12:31 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1050.eqiad.wmnet * 12:31 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1050.eqiad.wmnet * 12:29 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1210.eqiad.wmnet with reason: host reimage * 12:29 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1211.eqiad.wmnet with reason: host reimage * 12:29 moritzm: installing Wireshark security updates * 12:26 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow5003.eqsin.wmnet * 12:26 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1050.eqiad.wmnet * 12:24 cmooney@dns3003: END - running authdns-update * 12:21 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow6001.drmrs.wmnet * 12:21 cmooney@dns3003: START - running authdns-update * 12:20 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1050.eqiad.wmnet * 12:20 cgoubert@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host rdb-lock2003.codfw.wmnet * 12:20 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:20 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update netbox dns entries for expanded public1-603-eqsin subnet - cmooney@cumin1003" * 12:20 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update netbox dns entries for expanded public1-603-eqsin subnet - cmooney@cumin1003" * 12:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 12:20 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 12:19 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1049.eqiad.wmnet * 12:19 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1049.eqiad.wmnet * 12:19 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 12:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host moss-be1003.eqiad.wmnet * 12:17 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow6001.drmrs.wmnet * 12:17 moritzm: installin curl security updates * 12:15 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 12:15 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1233.eqiad.wmnet with OS bookworm * 12:15 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1210.eqiad.wmnet with OS bookworm * 12:15 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1211.eqiad.wmnet with OS bookworm * 12:13 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1049.eqiad.wmnet * 12:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow7002.magru.wmnet * 12:12 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-worker1178.eqiad.wmnet with OS bookworm * 12:10 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host moss-be1003.eqiad.wmnet * 12:10 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 12:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be1006.eqiad.wmnet * 12:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 12:10 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 12:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 12:10 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 12:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow7002.magru.wmnet * 12:08 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1049.eqiad.wmnet * 12:05 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 12:05 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2003.codfw.wmnet * 12:04 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'. * 12:04 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:03 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be1006.eqiad.wmnet * 12:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be1005.eqiad.wmnet * 12:01 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1038.eqiad.wmnet * 12:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1038.eqiad.wmnet * 12:01 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1213.eqiad.wmnet with OS bookworm * 11:56 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be1005.eqiad.wmnet * 11:55 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be1004.eqiad.wmnet * 11:54 cgoubert@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host rdb-lock2003.codfw.wmnet * 11:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 11:53 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 11:53 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:53 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 11:53 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 11:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1038.eqiad.wmnet * 11:49 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be1004.eqiad.wmnet * 11:44 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:44 moritzm: installing Linux 5.10.262 on Bullseye hosts * 11:41 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1038.eqiad.wmnet * 11:40 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 11:40 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 11:40 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 11:40 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:40 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 11:40 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 11:38 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1213.eqiad.wmnet with reason: host reimage * 11:36 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1167: Depool db1167.eqiad.wmnet to then clone it to db1281.eqiad.wmnet - marostegui@cumin1003 * 11:35 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 11:35 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2003.codfw.wmnet * 11:35 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1167: Depool db1167.eqiad.wmnet to then clone it to db1281.eqiad.wmnet - marostegui@cumin1003 * 11:35 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1167.eqiad.wmnet onto db1281.eqiad.wmnet * 11:34 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 22 hosts with reason: Cloning * 11:34 moritzm: remove ganeti3005 from esams03 cluster, hardware issues [[phab:T434646|T434646]] * 11:32 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1213.eqiad.wmnet with reason: host reimage * 11:28 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1034.eqiad.wmnet * 11:28 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1034.eqiad.wmnet * 11:22 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1034.eqiad.wmnet * 11:19 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1034.eqiad.wmnet * 11:17 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1213.eqiad.wmnet with OS bookworm * 11:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1165.eqiad.wmnet onto db1279.eqiad.wmnet * 11:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1165: Pool db1165.eqiad.wmnet in after cloning * 11:07 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply * 10:57 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply * 10:54 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply * 10:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1274.eqiad.wmnet with reason: Enabling notifications and pooling * 10:45 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1033.eqiad.wmnet * 10:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1033.eqiad.wmnet * 10:44 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply * 10:43 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'. * 10:42 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'. * 10:42 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'. * 10:40 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db[1216,1225,1239-1240].eqiad.wmnet with reason: reboot * 10:39 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1033.eqiad.wmnet * 10:38 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1151 from dbctl [[phab:T434538|T434538]]', diff saved to https://phabricator.wikimedia.org/P96055 and previous config saved to /var/cache/conftool/dbconfig/20260813-103828-marostegui.json * 10:35 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1033.eqiad.wmnet * 10:27 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1165: Pool db1165.eqiad.wmnet in after cloning * 10:24 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1151.eqiad.wmnet with OS bookworm * 10:15 moritzm: installing bind9 security updates (client-side tools/libs only) * 10:07 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-debug: apply * 10:06 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-debug: apply * 10:02 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 7 hosts with reason: reboot * 10:01 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-debug: apply * 10:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.upgrade (exit_code=0) for 1 hosts * 10:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2214: Upgrade of db2214.codfw.wmnet completed * 10:01 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-debug: apply * 10:00 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-debug: apply * 10:00 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-debug: apply * 09:59 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.decommission (exit_code=99) * 09:59 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 09:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1151.eqiad.wmnet with reason: host reimage * 09:59 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 8 hosts * 09:59 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 8 hosts * 09:55 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1151.eqiad.wmnet with reason: host reimage * 09:46 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260801000000" --end-timestamp="20260802000000" --sleep="3" --batch-size="5"` for [[phab:T434688|T434688]] * 09:45 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 8 hosts with reason: reboot * 09:45 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 22 hosts with reason: Cloning * 09:43 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1165: Depool db1165.eqiad.wmnet to then clone it to db1279.eqiad.wmnet - marostegui@cumin1003 * 09:42 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1165: Depool db1165.eqiad.wmnet to then clone it to db1279.eqiad.wmnet - marostegui@cumin1003 * 09:42 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1165.eqiad.wmnet onto db1279.eqiad.wmnet * 09:41 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=testwiki --start-timestamp="20260311000000" --end-timestamp="20260805000000" --sleep="5" --batch-size="2"` for [[phab:T434688|T434688]] * 09:40 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for backupmon1001.eqiad.wmnet * 09:40 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for backupmon1001.eqiad.wmnet * 09:40 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1151 * 09:40 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1151 * 09:37 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325408{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]], [[gerrit:1325407{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]] (duration: 06m 57s) * 09:36 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1151 * 09:36 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1151.eqiad.wmnet 13.36.64.10.in-addr.arpa 3.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:36 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1151.eqiad.wmnet 13.36.64.10.in-addr.arpa 3.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:36 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:36 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1151 - btullis@cumin1003" * 09:36 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on backupmon1001.eqiad.wmnet with reason: reboot * 09:36 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1151 - btullis@cumin1003" * 09:35 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host moss-be2003.codfw.wmnet * 09:33 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 7 hosts * 09:33 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 7 hosts * 09:33 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 09:32 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1325408{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]], [[gerrit:1325407{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1161.eqiad.wmnet onto db1275.eqiad.wmnet * 09:31 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 09:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1161: Pool db1161.eqiad.wmnet in after cloning * 09:30 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1325408{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]], [[gerrit:1325407{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]] * 09:29 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 09:27 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 09:27 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host moss-be2003.codfw.wmnet * 09:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be2006.codfw.wmnet * 09:25 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2004.codfw.wmnet * 09:22 hashar@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 09:21 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be2006.codfw.wmnet * 09:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be2005.codfw.wmnet * 09:19 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2004.codfw.wmnet * 09:19 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2003.codfw.wmnet * 09:18 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 7 hosts with reason: reboot * 09:18 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 7 hosts * 09:18 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 7 hosts * 09:16 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2214: Upgrade of db2214.codfw.wmnet completed * 09:14 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be2005.codfw.wmnet * 09:13 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be2004.codfw.wmnet * 09:12 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2003.codfw.wmnet * 09:12 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2002.codfw.wmnet * 09:09 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2214: Upgrading db2214.codfw.wmnet * 09:09 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2214: Upgrading db2214.codfw.wmnet * 09:09 cwilliams@cumin1003: START - Cookbook sre.mysql.upgrade for 1 hosts * 09:07 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be2004.codfw.wmnet * 09:06 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2002.codfw.wmnet * 09:05 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1004.eqiad.wmnet * 09:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-cluster (exit_code=0) * 09:03 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 7 hosts with reason: reboot * 09:01 btullis@cumin1003: START - Cookbook sre.dns.netbox * 09:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2214 [[phab:T434754|T434754]]', diff saved to https://phabricator.wikimedia.org/P96045 and previous config saved to /var/cache/conftool/dbconfig/20260813-090001-cwilliams.json * 08:59 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1004.eqiad.wmnet * 08:59 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1003.eqiad.wmnet * 08:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2229 to s6 primary [[phab:T434754|T434754]]', diff saved to https://phabricator.wikimedia.org/P96044 and previous config saved to /var/cache/conftool/dbconfig/20260813-085752-cwilliams.json * 08:57 cezmunsta: Starting s6 codfw failover from db2214 to db2229 - [[phab:T434754|T434754]] * 08:54 hashar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325416{{!}}Revert "REST: Enable `GET /lexemes/<nowiki>{</nowiki>lexeme_id<nowiki>}</nowiki>` by default" (T434712)]] (duration: 07m 22s) * 08:53 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1003.eqiad.wmnet * 08:53 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1002.eqiad.wmnet * 08:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2229 with weight 0 [[phab:T434754|T434754]]', diff saved to https://phabricator.wikimedia.org/P96043 and previous config saved to /var/cache/conftool/dbconfig/20260813-085151-cwilliams.json * 08:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 23 hosts with reason: Primary switchover s6 [[phab:T434754|T434754]] * 08:51 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts2002.codfw.wmnet * 08:50 hashar@deploy1003: hashar: Continuing with deployment * 08:49 hashar@deploy1003: hashar: Backport for [[gerrit:1325416{{!}}Revert "REST: Enable `GET /lexemes/<nowiki>{</nowiki>lexeme_id<nowiki>}</nowiki>` by default" (T434712)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:47 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet * 08:47 hashar@deploy1003: Started scap sync-world: Backport for [[gerrit:1325416{{!}}Revert "REST: Enable `GET /lexemes/<nowiki>{</nowiki>lexeme_id<nowiki>}</nowiki>` by default" (T434712)]] * 08:47 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1002.eqiad.wmnet * 08:46 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1161: Pool db1161.eqiad.wmnet in after cloning * 08:46 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host stewards2001.codfw.wmnet * 08:45 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit1003.wikimedia.org * 08:45 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host stewards1001.eqiad.wmnet * 08:45 Emperor: roll-restart apus frontends in codfw * 08:45 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-cluster * 08:44 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts2002.codfw.wmnet * 08:44 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host doc2003.codfw.wmnet * 08:43 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet * 08:43 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host phab2003.codfw.wmnet * 08:42 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host stewards2001.codfw.wmnet * 08:42 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host etherpad2002.codfw.wmnet * 08:41 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host stewards1001.eqiad.wmnet * 08:41 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1273: Pool in s7 * 08:41 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1003.wikimedia.org * 08:40 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host doc2003.codfw.wmnet * 08:40 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit1003.wikimedia.org * 08:39 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host etherpad1004.eqiad.wmnet * 08:39 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host doc1004.eqiad.wmnet * 08:39 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit2002.wikimedia.org * 08:38 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host etherpad2002.codfw.wmnet * 08:37 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host phab2003.codfw.wmnet * 08:36 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lists1004.wikimedia.org * 08:35 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host etherpad1004.eqiad.wmnet * 08:35 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host doc1004.eqiad.wmnet * 08:34 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-cluster (exit_code=0) * 08:34 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1003.wikimedia.org * 08:34 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host planet2003.codfw.wmnet * 08:34 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2003.wikimedia.org * 08:33 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host planet1003.eqiad.wmnet * 08:33 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit2002.wikimedia.org * 08:32 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 22 hosts with reason: Cloning * 08:32 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2002.wikimedia.org * 08:31 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aphlict1002.eqiad.wmnet * 08:30 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host planet2003.codfw.wmnet * 08:29 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host planet1003.eqiad.wmnet * 08:29 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host lists1004.wikimedia.org * 08:28 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lists2001.wikimedia.org * 08:28 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2003.wikimedia.org * 08:27 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host aphlict1002.eqiad.wmnet * 08:27 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aphlict2001.codfw.wmnet * 08:26 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2002.wikimedia.org * 08:23 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host aphlict2001.codfw.wmnet * 08:23 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1151 * 08:22 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1151.eqiad.wmnet with OS bookworm * 08:22 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host lists2001.wikimedia.org * 08:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1161: Depool db1161.eqiad.wmnet to then clone it to db1275.eqiad.wmnet - marostegui@cumin1003 * 08:20 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1161: Depool db1161.eqiad.wmnet to then clone it to db1275.eqiad.wmnet - marostegui@cumin1003 * 08:20 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1161.eqiad.wmnet onto db1275.eqiad.wmnet * 08:15 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-cluster * 08:15 Emperor: roll-restart apus frontends in eqiad * 07:56 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1273: Pool in s7 * 07:56 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1273 to dbctl [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P96035 and previous config saved to /var/cache/conftool/dbconfig/20260813-075611-marostegui.json * 07:38 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db1273.eqiad.wmnet with reason: Reboot * 07:31 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.sanitize-wiki (exit_code=97) Managing sanitization for wikis testwiki in section s3 * 07:24 marostegui@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis testwiki in section s3 * 07:19 jayme@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on kubestagemaster2005.codfw.wmnet with reason: downtime because of hardware failure and no DRBD * 05:42 arnaudb@dns1006: END - running authdns-update * 05:40 arnaudb@dns1006: START - running authdns-update * 05:27 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 04:06 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 02:29 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1151 * 02:29 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1151 * 02:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1219.eqiad.wmnet with OS bookworm * 02:18 ryankemper: [[phab:T434494|T434494]] `ryankemper@deploy1003:~$ echo 'https://stats.wikimedia.org/' {{!}} mwscript-k8s --attach -- purgeList.php` (default page got cached during yesterday's `an-web1001` reimage) * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 46s) * 02:03 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1219.eqiad.wmnet with reason: host reimage * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 02:00 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1219.eqiad.wmnet with reason: host reimage * 01:46 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1219.eqiad.wmnet with OS bookworm * 01:03 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1142.eqiad.wmnet with OS bookworm * 00:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1142.eqiad.wmnet with reason: host reimage * 00:34 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1142.eqiad.wmnet with reason: host reimage * 00:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1142.eqiad.wmnet with OS bookworm * 00:16 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1178 * 00:16 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1178 == 2026-08-12 == * 23:18 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324409{{!}}Remove $wmg = $wg hacks in Collection (T119117)]] (duration: 06m 43s) * 23:14 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 23:13 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1324409{{!}}Remove $wmg = $wg hacks in Collection (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:11 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1324409{{!}}Remove $wmg = $wg hacks in Collection (T119117)]] * 22:59 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 22:49 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260801000000" --end-timestamp="20260802000000" --sleep=2 --batch-size=10` * 22:45 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=testwiki --start-timestamp="20200801010101" --end-timestamp="20260816010101" --sleep=15 --batch-size=5` * 22:40 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=testwiki --start-timestamp="20200101010101" --end-timestamp="20260816010101" --sleep=60` * 22:35 Dreamy_Jazz: Running `mwscript WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260101000000" --end-timestamp="20260102000000" --sleep=10` * 22:21 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324817{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]], [[gerrit:1324818{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]] (duration: 45m 29s) * 22:17 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 21:59 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 21:40 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1324817{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]], [[gerrit:1324818{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:39 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 21:36 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1324817{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]], [[gerrit:1324818{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]] * 21:32 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:31 vriley@cumin1003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:30 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:30 vriley@cumin1003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:17 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:14 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:14 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:11 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:10 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-codfw: Set storage compatability to NONE — [[phab:T433028|T433028]] - eevans@cumin1003 * 21:10 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1006 * 21:09 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1006 * 21:05 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324277{{!}}Improve Math preference labels for SVG/MathJax/MathML (T433891)]] (duration: 31m 42s) * 20:58 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1178.eqiad.wmnet with OS bookworm * 20:54 krinkle@deploy1003: krinkle: Continuing with deployment * 20:51 krinkle@deploy1003: krinkle: Backport for [[gerrit:1324277{{!}}Improve Math preference labels for SVG/MathJax/MathML (T433891)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:39 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 20:38 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:38 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp5022.eqsin.wmnet with OS trixie * 20:38 cdobbins@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - cdobbins@cumin1003" * 20:37 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:36 cdobbins@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - cdobbins@cumin1003" * 20:34 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1178.eqiad.wmnet with reason: host reimage * 20:34 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1324277{{!}}Improve Math preference labels for SVG/MathJax/MathML (T433891)]] * 20:33 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 20:28 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1178.eqiad.wmnet with reason: host reimage * 20:13 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1178.eqiad.wmnet with OS bookworm * 20:11 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-worker1178.eqiad.wmnet with OS bookworm * 20:11 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1178.eqiad.wmnet with OS bookworm * 20:09 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-codfw: Set storage compatability to NONE — [[phab:T433028|T433028]] - eevans@cumin1003 * 20:09 Dreamy_Jazz: Evening UTC backport window done * 20:08 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324794{{!}}WikimediaAntiAbuse: Enable logging channel (T431292)]] (duration: 06m 48s) * 20:08 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage * 20:05 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage * 20:04 dreamyjazz@deploy1003: kharlan, dreamyjazz: Continuing with deployment * 20:04 dreamyjazz@deploy1003: kharlan, dreamyjazz: Backport for [[gerrit:1324794{{!}}WikimediaAntiAbuse: Enable logging channel (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:01 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1324794{{!}}WikimediaAntiAbuse: Enable logging channel (T431292)]] * 19:55 brennen@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 19:47 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 19:35 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 19:35 cdobbins@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cp5022.eqsin.wmnet with OS trixie * 19:32 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:30 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:29 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:26 vriley@cumin1003: START - Cookbook sre.dns.netbox * 19:23 brennen@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 19:19 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-eqiad: Set storage compatability to NONE — [[phab:T433028|T433028]] - eevans@cumin1003 * 19:10 brennen@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 19:09 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 19:09 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 19:00 Amir1: data migrated on wikishared ([[phab:T426102|T426102]]) * 18:57 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321224{{!}}Rename ce_worklist_articles table to ce_invitation_list_articles (T426102)]] (duration: 06m 50s) * 18:53 ladsgroup@deploy1003: ladsgroup, daimona: Continuing with deployment * 18:53 ladsgroup@deploy1003: ladsgroup, daimona: Backport for [[gerrit:1321224{{!}}Rename ce_worklist_articles table to ce_invitation_list_articles (T426102)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:51 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1321224{{!}}Rename ce_worklist_articles table to ce_invitation_list_articles (T426102)]] * 18:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2192: Security update * 18:31 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 18:25 Amir1: ce_invitation_list_articles created as empty on wikishared ([[phab:T426102|T426102]]) * 18:21 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 18:21 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 18:18 Amir1: migrated testwiki entries from ce_worklist_articles to ce_invitation_list_articles ([[phab:T426102|T426102]]) * 18:18 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-eqiad: Set storage compatability to NONE — [[phab:T433028|T433028]] - eevans@cumin1003 * 18:11 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 18:08 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling reboot on A:durum-eqsin and A:durum * 18:07 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Set storage compatability to UPGRADING — [[phab:T433028|T433028]] - eevans@cumin1003 * 18:05 jhancock@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022'] * 17:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2192: Security update * 17:55 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:55 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum-eqsin and A:durum * 17:53 jhancock@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['cp5022'] * 17:47 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:47 jhancock@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['cp5022'] * 17:42 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:41 jhancock@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['cp5022'] * 17:36 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:35 jhancock@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022'] * 17:31 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-magru and not (P<nowiki>{</nowiki>cp7001*<nowiki>}</nowiki> or P<nowiki>{</nowiki>cp7009*<nowiki>}</nowiki>) and A:cp - 9.2.15 upgrade ([[phab:T434620|T434620]]) * 17:28 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:28 jhancock@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['cp5022'] * 17:22 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2192.codfw.wmnet with reason: Maintenance * 17:11 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host apifeatureusage2001.codfw.wmnet with OS bookworm * 17:04 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: apply * 17:03 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-main: apply * 17:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2192 [[phab:T434635|T434635]]', diff saved to https://phabricator.wikimedia.org/P96030 and previous config saved to /var/cache/conftool/dbconfig/20260812-170338-cwilliams.json * 17:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2213 to s5 primary [[phab:T434635|T434635]]', diff saved to https://phabricator.wikimedia.org/P96029 and previous config saved to /var/cache/conftool/dbconfig/20260812-170152-cwilliams.json * 17:01 cezmunsta: Starting s5 codfw failover from db2192 to db2213 - [[phab:T434635|T434635]] * 16:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2213 with weight 0 [[phab:T434635|T434635]]', diff saved to https://phabricator.wikimedia.org/P96028 and previous config saved to /var/cache/conftool/dbconfig/20260812-165544-cwilliams.json * 16:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 27 hosts with reason: Primary switchover s5 [[phab:T434635|T434635]] * 16:53 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: apply * 16:52 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-main: apply * 16:44 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-main: apply * 16:44 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-main: apply * 16:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1272: New host * 16:40 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324765{{!}}WikimediaAntiAbuse: Enable PersonalInfoFlagNotifications (T431292)]] (duration: 07m 02s) * 16:40 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-logging-external: apply * 16:39 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-logging-external: apply * 16:38 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-logging-external: apply * 16:37 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-logging-external: apply * 16:36 kharlan@deploy1003: kharlan: Continuing with deployment * 16:35 kharlan@deploy1003: kharlan: Backport for [[gerrit:1324765{{!}}WikimediaAntiAbuse: Enable PersonalInfoFlagNotifications (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:33 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1324765{{!}}WikimediaAntiAbuse: Enable PersonalInfoFlagNotifications (T431292)]] * 16:27 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-logging-external: apply * 16:27 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-logging-external: apply * 16:18 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 16:18 jhancock@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022'] * 16:17 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 16:16 jhancock@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['cp5022'] * 16:15 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 16:12 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Set storage compatability to UPGRADING — [[phab:T433028|T433028]] - eevans@cumin1003 * 16:10 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: apply * 16:10 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: apply * 16:08 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: apply * 16:08 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: apply * 16:08 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics: apply * 16:07 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics: apply * 16:02 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-magru and not (P<nowiki>{</nowiki>cp7001*<nowiki>}</nowiki> or P<nowiki>{</nowiki>cp7009*<nowiki>}</nowiki>) and A:cp - 9.2.15 upgrade ([[phab:T434620|T434620]]) * 15:57 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1272: New host * 15:52 jmm@cumin2003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti2046.codfw.wmnet * 15:52 jmm@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host ganeti2046.codfw.wmnet * 15:42 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:37 mforns@deploy1003: Finished deploy [analytics/refinery@49c336c] (thin): Regular analytics weekly train THIN [analytics/refinery@49c336cd] (duration: 01m 59s) * 15:37 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324733{{!}}WikimediaAntiAbuse: Enable personal info tag display on enwiki (T431292)]] (duration: 08m 12s) * 15:35 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:35 mforns@deploy1003: Started deploy [analytics/refinery@49c336c] (thin): Regular analytics weekly train THIN [analytics/refinery@49c336cd] * 15:34 mforns@deploy1003: Finished deploy [analytics/refinery@49c336c]: Regular analytics weekly train [analytics/refinery@49c336cd] (duration: 04m 20s) * 15:33 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[2024,1031]*.wmnet: Set storage compatability to UPGRADING — [[phab:T433028|T433028]] - eevans@cumin1003 * 15:33 dreamyjazz@deploy1003: kharlan, dreamyjazz: Continuing with deployment * 15:31 dreamyjazz@deploy1003: kharlan, dreamyjazz: Backport for [[gerrit:1324733{{!}}WikimediaAntiAbuse: Enable personal info tag display on enwiki (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:30 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitize-wiki (exit_code=99) Checking sanitization for wikis testwiki in section s3 * 15:30 mforns@deploy1003: Started deploy [analytics/refinery@49c336c]: Regular analytics weekly train [analytics/refinery@49c336cd] * 15:30 mforns@deploy1003: Finished deploy [analytics/refinery@49c336c] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@49c336cd] (duration: 00m 32s) * 15:29 mforns@deploy1003: Started deploy [analytics/refinery@49c336c] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@49c336cd] * 15:29 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1324733{{!}}WikimediaAntiAbuse: Enable personal info tag display on enwiki (T431292)]] * 15:27 brennen@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324700{{!}}EventDetailsParticipantsModule: populate cache with non-local users (T434597)]] (duration: 06m 38s) * 15:23 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[2024,1031]*.wmnet: Set storage compatability to UPGRADING — [[phab:T433028|T433028]] - eevans@cumin1003 * 15:23 brennen@deploy1003: brennen, daimona: Continuing with deployment * 15:23 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:22 brennen@deploy1003: brennen, daimona: Backport for [[gerrit:1324700{{!}}EventDetailsParticipantsModule: populate cache with non-local users (T434597)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:22 cgoubert@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host rdb-lock2003.codfw.wmnet * 15:21 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 15:21 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 15:21 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:21 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:21 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:20 brennen@deploy1003: Started scap sync-world: Backport for [[gerrit:1324700{{!}}EventDetailsParticipantsModule: populate cache with non-local users (T434597)]] * 15:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1178.eqiad.wmnet with OS bookworm * 15:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock1003.eqiad.wmnet * 15:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock1003.eqiad.wmnet with OS trixie * 15:16 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 15:16 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 15:16 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 15:16 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:16 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:16 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:12 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 15:12 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2003.codfw.wmnet * 15:11 cgoubert@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host rdb-lock2003.codfw.wmnet * 15:11 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 15:11 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 15:11 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:11 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:11 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:07 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324338{{!}}InitialiseSettings: Enable 2FA warnings on more private wikis (T428103)]] (duration: 07m 02s) * 15:04 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 15:04 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock1003.eqiad.wmnet with reason: host reimage * 15:03 reedy@deploy1003: reedy: Continuing with deployment * 15:02 reedy@deploy1003: reedy: Backport for [[gerrit:1324338{{!}}InitialiseSettings: Enable 2FA warnings on more private wikis (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:02 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:00 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324338{{!}}InitialiseSettings: Enable 2FA warnings on more private wikis (T428103)]] * 14:57 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 14:57 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2003.codfw.wmnet * 14:57 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock1003.eqiad.wmnet with reason: host reimage * 14:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock2002.codfw.wmnet * 14:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock2002.codfw.wmnet with OS trixie * 14:56 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 14:56 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 14:56 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 14:55 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 14:55 moritzm: powercycle ganeti2046 * 14:47 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock1003.eqiad.wmnet with OS trixie * 14:46 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1003.eqiad.wmnet - cgoubert@cumin2003" * 14:46 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1003.eqiad.wmnet - cgoubert@cumin2003" * 14:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock1003.eqiad.wmnet on all recursors * 14:45 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock1003.eqiad.wmnet on all recursors * 14:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1003.eqiad.wmnet - cgoubert@cumin2003" * 14:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1272.eqiad.wmnet with reason: Enabling notifications * 14:44 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1003.eqiad.wmnet - cgoubert@cumin2003" * 14:44 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324719{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324720{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324722{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0 (T434187)]] (duration: 11m 02s) * 14:39 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 14:39 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock1003.eqiad.wmnet * 14:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock2002.codfw.wmnet with reason: host reimage * 14:37 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 14:37 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock1002.eqiad.wmnet * 14:37 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock1002.eqiad.wmnet with OS trixie * 14:37 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1324719{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324720{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324722{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0 (T434187)]] synced to the testservers (see https://wikitech. * 14:33 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock2002.codfw.wmnet with reason: host reimage * 14:33 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1324719{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324720{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324722{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0 (T434187)]] * 14:32 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1018.eqiad.wmnet with OS bookworm * 14:32 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2046.codfw.wmnet * 14:31 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1020.eqiad.wmnet with OS bookworm * 14:27 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2046.codfw.wmnet * 14:25 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2045.codfw.wmnet * 14:25 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2045.codfw.wmnet * 14:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock1002.eqiad.wmnet with reason: host reimage * 14:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1019.eqiad.wmnet with OS bookworm * 14:22 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Checking sanitization for wikis testwiki in section s3 * 14:20 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2045.codfw.wmnet * 14:18 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock1002.eqiad.wmnet with reason: host reimage * 14:17 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:16 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324707{{!}}Backport all changes from wmf/1.47.0-wmf.15]] (duration: 40m 51s) * 14:16 moritzm: installing Linux 6.1.180 on Bookworm hosts * 14:15 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock2002.codfw.wmnet with OS trixie * 14:14 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2002.codfw.wmnet - cgoubert@cumin2003" * 14:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2002.codfw.wmnet - cgoubert@cumin2003" * 14:14 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:14 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2045.codfw.wmnet * 14:14 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2002.codfw.wmnet on all recursors * 14:14 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2002.codfw.wmnet on all recursors * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2002.codfw.wmnet - cgoubert@cumin2003" * 14:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2002.codfw.wmnet - cgoubert@cumin2003" * 14:12 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2030.codfw.wmnet * 14:12 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2030.codfw.wmnet * 14:11 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on apifeatureusage2001.codfw.wmnet with reason: host reimage * 14:09 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:08 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:07 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 14:06 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2030.codfw.wmnet * 14:06 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:06 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock1002.eqiad.wmnet with OS trixie * 14:05 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 14:05 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2002.codfw.wmnet * 14:05 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1002.eqiad.wmnet - cgoubert@cumin2003" * 14:05 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1002.eqiad.wmnet - cgoubert@cumin2003" * 14:05 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock1002.eqiad.wmnet on all recursors * 14:05 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock1002.eqiad.wmnet on all recursors * 14:05 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:05 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1002.eqiad.wmnet - cgoubert@cumin2003" * 14:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:04 kharlan@deploy1003: kharlan: Continuing with deployment * 14:04 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock2001.codfw.wmnet * 14:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock2001.codfw.wmnet with OS trixie * 14:02 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1002.eqiad.wmnet - cgoubert@cumin2003" * 14:02 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on apifeatureusage2001.codfw.wmnet with reason: host reimage * 14:01 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2030.codfw.wmnet * 13:59 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2029.codfw.wmnet * 13:58 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2029.codfw.wmnet * 13:58 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 13:58 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock1002.eqiad.wmnet * 13:56 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1019.eqiad.wmnet with reason: host reimage * 13:54 btullis@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on archiva1002.wikimedia.org with reason: Upgrading in-place * 13:53 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock1001.eqiad.wmnet * 13:53 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock1001.eqiad.wmnet with OS trixie * 13:53 kharlan@deploy1003: kharlan: Backport for [[gerrit:1324707{{!}}Backport all changes from wmf/1.47.0-wmf.15]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:52 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2029.codfw.wmnet * 13:52 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1020.eqiad.wmnet with reason: host reimage * 13:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1019.eqiad.wmnet with reason: host reimage * 13:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1020.eqiad.wmnet with reason: host reimage * 13:49 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock2001.codfw.wmnet with reason: host reimage * 13:48 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2029.codfw.wmnet * 13:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Configuring db1272 for s3 pooling', diff saved to https://phabricator.wikimedia.org/P96021 and previous config saved to /var/cache/conftool/dbconfig/20260812-134732-cwilliams.json * 13:44 bking@cumin2003: START - Cookbook sre.hosts.reimage for host apifeatureusage2001.codfw.wmnet with OS bookworm * 13:43 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock2001.codfw.wmnet with reason: host reimage * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2028.codfw.wmnet * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2028.codfw.wmnet * 13:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1018.eqiad.wmnet with reason: host reimage * 13:38 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock1001.eqiad.wmnet with reason: host reimage * 13:36 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1018.eqiad.wmnet with reason: host reimage * 13:35 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1324707{{!}}Backport all changes from wmf/1.47.0-wmf.15]] * 13:35 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2028.codfw.wmnet * 13:32 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1275938{{!}}Enable campaignEvents on bdwikimedia (T424016)]] (duration: 07m 35s) * 13:32 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1020.eqiad.wmnet with OS bookworm * 13:32 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1019.eqiad.wmnet with OS bookworm * 13:32 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock1001.eqiad.wmnet with reason: host reimage * 13:28 kharlan@deploy1003: kharlan, yahya: Continuing with deployment * 13:27 kharlan@deploy1003: kharlan, yahya: Backport for [[gerrit:1275938{{!}}Enable campaignEvents on bdwikimedia (T424016)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:27 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2028.codfw.wmnet * 13:26 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 13:26 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 13:25 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1275938{{!}}Enable campaignEvents on bdwikimedia (T424016)]] * 13:25 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitize-wiki (exit_code=99) Managing sanitization for wikis testwiki in section s3 * 13:24 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock2001.codfw.wmnet with OS trixie * 13:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2001.codfw.wmnet - cgoubert@cumin2003" * 13:24 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2001.codfw.wmnet - cgoubert@cumin2003" * 13:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2001.codfw.wmnet on all recursors * 13:23 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2001.codfw.wmnet on all recursors * 13:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2001.codfw.wmnet - cgoubert@cumin2003" * 13:23 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2001.codfw.wmnet - cgoubert@cumin2003" * 13:23 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324713{{!}}thwiki: reinstate temporary wiki25 logos (T431094)]] (duration: 07m 13s) * 13:20 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:20 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1018.eqiad.wmnet with OS bookworm * 13:20 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2027.codfw.wmnet * 13:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2027.codfw.wmnet * 13:19 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics-external: apply * 13:18 kharlan@deploy1003: anzx, kharlan: Continuing with deployment * 13:18 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:18 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock1001.eqiad.wmnet with OS trixie * 13:17 kharlan@deploy1003: anzx, kharlan: Backport for [[gerrit:1324713{{!}}thwiki: reinstate temporary wiki25 logos (T431094)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1001.eqiad.wmnet - cgoubert@cumin2003" * 13:17 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 13:17 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1001.eqiad.wmnet - cgoubert@cumin2003" * 13:17 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2001.codfw.wmnet * 13:17 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics-external: apply * 13:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock1001.eqiad.wmnet on all recursors * 13:17 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock1001.eqiad.wmnet on all recursors * 13:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1001.eqiad.wmnet - cgoubert@cumin2003" * 13:17 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1001.eqiad.wmnet - cgoubert@cumin2003" * 13:15 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1324713{{!}}thwiki: reinstate temporary wiki25 logos (T431094)]] * 13:15 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:15 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics-external: apply * 13:15 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:14 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics-external: apply * 13:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2027.codfw.wmnet * 13:12 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 13:12 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock1001.eqiad.wmnet * 13:12 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2027.codfw.wmnet * 13:08 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1158.eqiad.wmnet onto db1273.eqiad.wmnet * 13:07 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1158: Pool db1158.eqiad.wmnet in after cloning * 13:02 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti6002.drmrs.wmnet * 13:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti6002.drmrs.wmnet * 12:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1015.eqiad.wmnet with OS bookworm * 12:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti6002.drmrs.wmnet * 12:44 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1017.eqiad.wmnet with OS bookworm * 12:36 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 12:35 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 12:34 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 12:33 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 12:31 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti6002.drmrs.wmnet * 12:24 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1017.eqiad.wmnet with reason: host reimage * 12:22 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1158: Pool db1158.eqiad.wmnet in after cloning * 12:18 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1017.eqiad.wmnet with reason: host reimage * 12:11 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1015.eqiad.wmnet with reason: host reimage * 12:07 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1015.eqiad.wmnet with reason: host reimage * 12:04 moritzm: failover ganeti master in drmrs02 to ganeti6004 * 12:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1009.eqiad.wmnet with OS bookworm * 12:01 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1017.eqiad.wmnet with OS bookworm * 12:00 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti6004.drmrs.wmnet * 12:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti6004.drmrs.wmnet * 11:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti6004.drmrs.wmnet * 11:53 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1015.eqiad.wmnet with OS bookworm * 11:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1159.eqiad.wmnet onto db1274.eqiad.wmnet * 11:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Pool db1159.eqiad.wmnet in after cloning * 11:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1016.eqiad.wmnet with OS bookworm * 11:46 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti6004.drmrs.wmnet * 11:45 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti6001.drmrs.wmnet * 11:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti6001.drmrs.wmnet * 11:43 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324306{{!}}WikimediaAntiAbuse: Enable personal info for enwiki with no display (T431292)]] (duration: 10m 26s) * 11:42 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1015.eqiad.wmnet with OS bookworm * 11:39 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 11:38 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti6001.drmrs.wmnet * 11:34 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1324306{{!}}WikimediaAntiAbuse: Enable personal info for enwiki with no display (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:33 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2212: Security update * 11:33 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti6001.drmrs.wmnet * 11:32 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1324306{{!}}WikimediaAntiAbuse: Enable personal info for enwiki with no display (T431292)]] * 11:22 moritzm: failover ganeti master in drmrs01 to ganeti6003 * 11:20 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:20 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:18 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 22 hosts with reason: Cloning * 11:17 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti6003.drmrs.wmnet * 11:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti6003.drmrs.wmnet * 11:17 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1009.eqiad.wmnet with reason: host reimage * 11:17 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:16 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:14 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1016.eqiad.wmnet with reason: host reimage * 11:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti6003.drmrs.wmnet * 11:11 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1009.eqiad.wmnet with reason: host reimage * 11:10 jmm@cumin2003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti3005.esams.wmnet * 11:10 jmm@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host ganeti3005.esams.wmnet * 11:09 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 11:08 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1158: Depool db1158.eqiad.wmnet to then clone it to db1273.eqiad.wmnet - marostegui@cumin1003 * 11:07 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1016.eqiad.wmnet with reason: host reimage * 11:07 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1158: Depool db1158.eqiad.wmnet to then clone it to db1273.eqiad.wmnet - marostegui@cumin1003 * 11:07 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1158.eqiad.wmnet onto db1273.eqiad.wmnet * 11:06 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:05 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:05 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1159: Pool db1159.eqiad.wmnet in after cloning * 11:04 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 20 hosts with reason: Cloning * 11:02 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti6003.drmrs.wmnet * 11:00 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis testwiki in section s3 * 10:54 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1009.eqiad.wmnet with OS bookworm * 10:51 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1015.eqiad.wmnet with OS bookworm * 10:50 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1016.eqiad.wmnet with OS bookworm * 10:48 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2212: Security update * 10:45 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitize-wiki (exit_code=99) Managing sanitization for wikis testwiki in section s3 * 10:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1013.eqiad.wmnet with OS bookworm * 10:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1014.eqiad.wmnet with OS bookworm * 10:38 cwilliams@cumin1003: START - Cookbook sre.mysql.clone of db1159.eqiad.wmnet onto db1274.eqiad.wmnet * 10:33 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db1274.eqiad.wmnet * 10:33 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db1274.eqiad.wmnet * 10:31 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 396993 * 10:29 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 396993 * 10:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1159: Clone source for db1274 * 10:24 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1159: Clone source for db1274 * 10:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1013.eqiad.wmnet with reason: host reimage * 10:18 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1013.eqiad.wmnet with reason: host reimage * 10:13 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2212.codfw.wmnet with reason: Maintenance * 10:13 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 10:12 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 10:12 blake@deploy1003: Stopping before sync operations * 10:11 blake@deploy1003: Started scap sync-world: Non-deployment scap run to populate new release values for [[phab:T427668|T427668]] * 10:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2212 [[phab:T434644|T434644]]', diff saved to https://phabricator.wikimedia.org/P96003 and previous config saved to /var/cache/conftool/dbconfig/20260812-101053-cwilliams.json * 10:09 moritzm: powercycle ganeti3005 * 10:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2203 to s1 primary [[phab:T434644|T434644]]', diff saved to https://phabricator.wikimedia.org/P96002 and previous config saved to /var/cache/conftool/dbconfig/20260812-100849-cwilliams.json * 10:08 cezmunsta: Starting s1 codfw failover from db2212 to db2203 - [[phab:T434644|T434644]] * 10:03 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1013.eqiad.wmnet with OS bookworm * 10:02 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1013.eqiad.wmnet with OS bookworm * 10:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2203 with weight 0 [[phab:T434644|T434644]]', diff saved to https://phabricator.wikimedia.org/P96001 and previous config saved to /var/cache/conftool/dbconfig/20260812-100134-cwilliams.json * 10:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 32 hosts with reason: Primary switchover s1 [[phab:T434644|T434644]] * 09:53 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1014.eqiad.wmnet with reason: host reimage * 09:50 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 09:50 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti3005.esams.wmnet * 09:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1014.eqiad.wmnet with reason: host reimage * 09:41 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1278: Pool in x1 * 09:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-launcher1003.eqiad.wmnet with OS bookworm * 09:37 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti3005.esams.wmnet * 09:37 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1013.eqiad.wmnet with OS bookworm * 09:34 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-presto1013.eqiad.wmnet with OS bookworm * 09:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1014.eqiad.wmnet with OS bookworm * 09:29 moritzm: failover ganeti master in esams to ganeti3008 * 09:26 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti3006.esams.wmnet * 09:26 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti3006.esams.wmnet * 09:24 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1012.eqiad.wmnet with OS bookworm * 09:23 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1009.eqiad.wmnet with OS bookworm * 09:23 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 09:18 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti3006.esams.wmnet * 09:16 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti3006.esams.wmnet * 09:14 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: fix regexp escaping bug - oblivian@cumin1003" * 09:14 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: fix regexp escaping bug - oblivian@cumin1003 * 09:13 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: fix regexp escaping bug - oblivian@cumin1003 * 09:13 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: fix regexp escaping bug - oblivian@cumin1003" * 09:03 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-launcher1003.eqiad.wmnet with reason: host reimage * 08:58 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-launcher1003.eqiad.wmnet with reason: host reimage * 08:55 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1278: Pool in x1 * 08:55 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1278 to dbctl [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95996 and previous config saved to /var/cache/conftool/dbconfig/20260812-085521-marostegui.json * 08:51 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1012.eqiad.wmnet with reason: host reimage * 08:45 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis testwiki in section s3 * 08:43 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1009.eqiad.wmnet with OS bookworm * 08:42 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1012.eqiad.wmnet with reason: host reimage * 08:41 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-launcher1003.eqiad.wmnet with OS bookworm * 08:40 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1013.eqiad.wmnet with OS bookworm * 08:38 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms3', diff saved to https://phabricator.wikimedia.org/P95995 and previous config saved to /var/cache/conftool/dbconfig/20260812-083816-marostegui.json * 08:38 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-master1003.eqiad.wmnet with OS bookworm * 08:37 marostegui: Failover ms3 [[phab:T434288|T434288]] * 08:37 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1268 to dbctl [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95994 and previous config saved to /var/cache/conftool/dbconfig/20260812-083722-marostegui.json * 08:35 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-web1001.eqiad.wmnet with OS bookworm * 08:32 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db2252.codfw.wmnet,db[1153,1268].eqiad.wmnet with reason: Switching over ms3 * 08:28 marostegui@cumin1003: dbctl commit (dc=all): 'Depool ms3 [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95993 and previous config saved to /var/cache/conftool/dbconfig/20260812-082852-marostegui.json * 08:25 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1012.eqiad.wmnet with OS bookworm * 08:25 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1011.eqiad.wmnet with OS bookworm * 08:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-master1003.eqiad.wmnet with reason: host reimage * 08:07 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-master1003.eqiad.wmnet with reason: host reimage * 08:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-web1001.eqiad.wmnet with reason: host reimage * 07:58 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-web1001.eqiad.wmnet with reason: host reimage * 07:50 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1003.eqiad.wmnet with OS bookworm * 07:38 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1011.eqiad.wmnet with reason: host reimage * 07:38 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-master1003.eqiad.wmnet with OS bookworm * 07:35 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti3007.esams.wmnet * 07:35 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti3007.esams.wmnet * 07:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1011.eqiad.wmnet with reason: host reimage * 07:27 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti3007.esams.wmnet * 07:25 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti3007.esams.wmnet * 07:25 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti3008.esams.wmnet * 07:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti3008.esams.wmnet * 07:22 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-web1001.eqiad.wmnet with OS bookworm * 07:19 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1003.eqiad.wmnet with OS bookworm * 07:18 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-master1003.eqiad.wmnet * 07:18 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host an-master1003.eqiad.wmnet * 07:17 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1011.eqiad.wmnet with OS bookworm * 07:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 07:16 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti3008.esams.wmnet * 07:15 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1009.eqiad.wmnet with OS bookworm * 07:14 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host an-master1003.eqiad.wmnet * 07:13 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-master1003.eqiad.wmnet * 07:13 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-master1003.eqiad.wmnet * 07:12 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-master1003.eqiad.wmnet * 07:11 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti3008.esams.wmnet * 07:07 arnaudb@dns1006: END - running authdns-update * 07:07 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti5007.eqsin.wmnet * 07:07 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti5007.eqsin.wmnet * 07:05 arnaudb@dns1006: START - running authdns-update * 06:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti5007.eqsin.wmnet * 06:54 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti5007.eqsin.wmnet * 06:38 moritzm: failover ganeti master in eqsin to ganeti5004 * 06:36 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti5006.eqsin.wmnet * 06:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti5006.eqsin.wmnet * 06:28 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti5006.eqsin.wmnet * 06:23 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti5006.eqsin.wmnet * 06:20 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti5005.eqsin.wmnet * 06:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti5005.eqsin.wmnet * 06:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti5005.eqsin.wmnet * 06:06 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti5005.eqsin.wmnet * 06:03 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti5004.eqsin.wmnet * 06:03 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti5004.eqsin.wmnet * 05:55 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti5004.eqsin.wmnet * 05:53 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti5004.eqsin.wmnet * 04:40 ryankemper: [[phab:T434494|T434494]] reimaged `an-tool1008.eqiad.wmnet` to bookworm; yarn.wikimedia.org is back up * 04:16 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-tool1008.eqiad.wmnet with OS bookworm * 03:58 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-tool1008.eqiad.wmnet with reason: host reimage * 03:53 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-tool1008.eqiad.wmnet with reason: host reimage * 03:41 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-tool1008.eqiad.wmnet with OS bookworm * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 45s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 00:25 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324427{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]], [[gerrit:1324429{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0]], [[gerrit:1324428{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]] (duration: 07m 55s) * 00:21 kemayo@deploy1003: kemayo: Continuing with deployment * 00:19 kemayo@deploy1003: kemayo: Backport for [[gerrit:1324427{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]], [[gerrit:1324429{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0]], [[gerrit:1324428{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:17 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1324427{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]], [[gerrit:1324429{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0]], [[gerrit:1324428{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]] == 2026-08-11 == * 21:37 sbassett: Deployed security fix for [[phab:T434521|T434521]] (wmf.15) * 21:29 sbassett: Deployed security fix for [[phab:T434521|T434521]] (wmf.14) * 21:19 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324370{{!}}Phase 4 of legal footer deployment (T432796)]], [[gerrit:1319804{{!}}Disable wgMFCustomSiteModules on English Wikipedia (T375538)]] (duration: 15m 26s) * 21:15 jdlrobson@deploy1003: jdlrobson: Continuing with deployment * 21:06 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1324370{{!}}Phase 4 of legal footer deployment (T432796)]], [[gerrit:1319804{{!}}Disable wgMFCustomSiteModules on English Wikipedia (T375538)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:03 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1324370{{!}}Phase 4 of legal footer deployment (T432796)]], [[gerrit:1319804{{!}}Disable wgMFCustomSiteModules on English Wikipedia (T375538)]] * 20:59 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 20:50 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324384{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]], [[gerrit:1324385{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]] (duration: 06m 58s) * 20:46 kemayo@deploy1003: kemayo: Continuing with deployment * 20:45 kemayo@deploy1003: kemayo: Backport for [[gerrit:1324384{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]], [[gerrit:1324385{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:43 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1324384{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]], [[gerrit:1324385{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]] * 20:42 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 20:42 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324386{{!}}build: Updating js-yaml to 3.15.1, 4.3.1]] (duration: 07m 36s) * 20:38 kemayo@deploy1003: kemayo: Continuing with deployment * 20:37 jhancock@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 20:37 kemayo@deploy1003: kemayo: Backport for [[gerrit:1324386{{!}}build: Updating js-yaml to 3.15.1, 4.3.1]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:35 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1324386{{!}}build: Updating js-yaml to 3.15.1, 4.3.1]] * 20:18 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 20:15 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 20:15 jhancock@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin1003" * 20:14 jhancock@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin1003" * 19:59 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 19:54 jhancock@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 19:10 brennen@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] (duration: 06m 41s) * 19:04 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns6001.wikimedia.org * 19:04 sukhe@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns6001.wikimedia.org * 19:04 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns5003.wikimedia.org * 19:04 sukhe@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns5003.wikimedia.org * 19:03 brennen@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 18:59 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns5003.wikimedia.org with OS trixie * 18:55 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns6001.wikimedia.org with OS trixie * 18:19 brett@cumin2002: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on P<nowiki>{</nowiki>cp7009.magru.wmnet<nowiki>}</nowiki> and A:cp - 9.2.15 Upgrade () * 18:14 brett@cumin2002: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on P<nowiki>{</nowiki>cp7009.magru.wmnet<nowiki>}</nowiki> and A:cp - 9.2.15 Upgrade () * 18:13 brennen@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 18:12 brett@cumin2002: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 9.2.15 Upgrade () * 18:09 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns5003.wikimedia.org with reason: host reimage * 18:06 brett@cumin2002: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 9.2.15 Upgrade () * 18:06 brennen: 1.47.0-wmf.15 train status ([[phab:T430834|T430834]]) - no current blockers, rolling to group0 * 18:05 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns5003.wikimedia.org with reason: host reimage * 18:05 brett: import trafficserver-9.2.15~deb13+wmf1 into trixie-wikimedia ([[phab:T434478|T434478]]) * 18:01 ladsgroup@cumin1003: END (PASS) - Cookbook sre.mysql.sanitarium_restart (exit_code=0) * 17:58 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns6001.wikimedia.org with reason: host reimage * 17:53 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324369{{!}}Enable desktop lazy loading on group0 (T148047)]] (duration: 07m 31s) * 17:52 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns6001.wikimedia.org with reason: host reimage * 17:49 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 17:49 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitarium_restart (exit_code=99) * 17:49 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 17:49 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7001.magru.wmnet * 17:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti7001.magru.wmnet * 17:48 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 17:47 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1324369{{!}}Enable desktop lazy loading on group0 (T148047)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:45 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1324369{{!}}Enable desktop lazy loading on group0 (T148047)]] * 17:39 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti7001.magru.wmnet * 17:36 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns5003.wikimedia.org with OS trixie * 17:34 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns6001.wikimedia.org with OS trixie * 17:31 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324361{{!}}Move FR config from IS.php to a dedicated file]], [[gerrit:1324363{{!}}Remove $wmg = $wg hacks in CentralAuth (T119117)]] (duration: 12m 23s) * 17:26 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 17:23 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1324361{{!}}Move FR config from IS.php to a dedicated file]], [[gerrit:1324363{{!}}Remove $wmg = $wg hacks in CentralAuth (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:19 sukhe: sudo cumin "A:cp-magru" "run-puppet-agent --enable 'merging CR 1324355'" * 17:18 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1324361{{!}}Move FR config from IS.php to a dedicated file]], [[gerrit:1324363{{!}}Remove $wmg = $wg hacks in CentralAuth (T119117)]] * 17:11 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-master1004.eqiad.wmnet with OS bookworm * 17:11 sukhe: sukhe@cp7005:~$ sudo puppet agent -tv * 17:02 sukhe: sudo cumin "A:cp-magru" "disable-puppet 'merging CR 1324355'" * 16:54 sukhe@dns1004: END - running authdns-update * 16:53 sukhe@dns1004: START - running authdns-update * 16:53 sukhe@dns1004: FAIL - running authdns-update * 16:51 sukhe@dns1004: START - running authdns-update * 16:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-master1004.eqiad.wmnet with reason: host reimage * 16:44 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-master1004.eqiad.wmnet with reason: host reimage * 16:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1157.eqiad.wmnet onto db1272.eqiad.wmnet * 16:40 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1157: Pool db1157.eqiad.wmnet in after cloning * 16:38 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 16:31 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324356{{!}}InitialiseSettings: Fix wgOATHAuthEnforce2FAForAll]] (duration: 06m 52s) * 16:30 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 16:28 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 16:27 reedy@deploy1003: reedy: Continuing with deployment * 16:26 reedy@deploy1003: reedy: Backport for [[gerrit:1324356{{!}}InitialiseSettings: Fix wgOATHAuthEnforce2FAForAll]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:24 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324356{{!}}InitialiseSettings: Fix wgOATHAuthEnforce2FAForAll]] * 16:13 sukhe: restart ntpsec.serviceon dns7001 * 16:09 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324335{{!}}InitialiseSettings: Enable 2FA enforcement on various private wikis (T428103)]] (duration: 06m 40s) * 16:08 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 16:06 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2204: Security update * 16:04 reedy@deploy1003: reedy: Continuing with deployment * 16:04 reedy@deploy1003: reedy: Backport for [[gerrit:1324335{{!}}InitialiseSettings: Enable 2FA enforcement on various private wikis (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:02 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7001.magru.wmnet * 16:02 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324335{{!}}InitialiseSettings: Enable 2FA enforcement on various private wikis (T428103)]] * 16:01 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 15:55 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1157: Pool db1157.eqiad.wmnet in after cloning * 15:54 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 15:54 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 15:41 dancy@deploy1003: Finished scap sync-world: Testing (duration: 06m 28s) * 15:40 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-master1004.eqiad.wmnet with OS bookworm * 15:40 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 15:35 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti4008.ulsfo.wmnet * 15:35 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti4008.ulsfo.wmnet * 15:34 dancy@deploy1003: Started scap sync-world: Testing * 15:34 dancy@deploy1003: Installation of scap version "4.279.0" completed for 3 hosts * 15:34 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-master1004.eqiad.wmnet with OS bookworm * 15:34 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 15:33 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-master1004.eqiad.wmnet with OS bookworm * 15:32 dancy@deploy1003: Installing scap version "4.279.0" for 3 host(s) * 15:32 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324339{{!}}Add /w/deployment-info.php entrypoint]] (duration: 07m 25s) * 15:30 moritzm: failover ganeti master in magru to ganeti7004 * 15:29 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti4008.ulsfo.wmnet * 15:28 dancy@deploy1003: dancy: Continuing with deployment * 15:28 tappof: remove 2026-05 swift log archives from centrallog to free some space ([[phab:T434502|T434502]]) * 15:27 dancy@deploy1003: dancy: Backport for [[gerrit:1324339{{!}}Add /w/deployment-info.php entrypoint]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:25 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324339{{!}}Add /w/deployment-info.php entrypoint]] * 15:20 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2204: Security update * 15:18 dancy@deploy1003: Installation of scap version "4.278.0" completed for 3 hosts * 15:18 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7004.magru.wmnet * 15:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti7004.magru.wmnet * 15:16 dancy@deploy1003: Installing scap version "4.278.0" for 3 host(s) * 15:14 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2204.codfw.wmnet with reason: Maintenance * 15:12 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 15:11 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-master1004.eqiad.wmnet with OS bookworm * 15:11 hashar: Restarting CI Jenkins on contint1003 due to Java upgrade. * 15:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2204 [[phab:T434565|T434565]]', diff saved to https://phabricator.wikimedia.org/P95984 and previous config saved to /var/cache/conftool/dbconfig/20260811-151126-cwilliams.json * 15:10 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti4008.ulsfo.wmnet * 15:10 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti7004.magru.wmnet * 15:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2207 to s2 primary [[phab:T434565|T434565]]', diff saved to https://phabricator.wikimedia.org/P95983 and previous config saved to /var/cache/conftool/dbconfig/20260811-150905-cwilliams.json * 15:08 cezmunsta: Starting s2 codfw failover from db2204 to db2207 - [[phab:T434565|T434565]] * 15:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2207 with weight 0 [[phab:T434565|T434565]]', diff saved to https://phabricator.wikimedia.org/P95982 and previous config saved to /var/cache/conftool/dbconfig/20260811-150402-cwilliams.json * 15:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s2 [[phab:T434565|T434565]] * 14:55 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1010.eqiad.wmnet with OS bookworm * 14:49 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns4003.wikimedia.org with OS trixie * 14:47 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1010.eqiad.wmnet with OS bookworm * 14:47 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7001.wikimedia.org with OS trixie * 14:44 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-presto1010.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:41 btullis@cumin1003: START - Cookbook sre.hosts.provision for host an-presto1010.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:40 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-presto1010.eqiad.wmnet with OS bookworm * 14:39 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 14:39 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-presto1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:36 btullis@cumin1003: START - Cookbook sre.hosts.provision for host an-presto1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:32 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1009.eqiad.wmnet with OS bookworm * 14:32 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 14:31 cwilliams@cumin1003: START - Cookbook sre.mysql.clone of db1157.eqiad.wmnet onto db1272.eqiad.wmnet * 14:30 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7004.magru.wmnet * 14:28 moritzm: failover ganeti master in ulsfo to ganeti4005 * 14:23 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7003.magru.wmnet * 14:23 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti7003.magru.wmnet * 14:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-master1004.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:22 btullis@cumin1003: START - Cookbook sre.hosts.provision for host an-master1004.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:21 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-master1004.eqiad.wmnet with OS bookworm * 14:19 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti4007.ulsfo.wmnet * 14:19 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti4007.ulsfo.wmnet * 14:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1010.eqiad.wmnet with OS bookworm * 14:17 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1008.eqiad.wmnet with OS bookworm * 14:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti7003.magru.wmnet * 14:12 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7003.magru.wmnet * 14:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti4007.ulsfo.wmnet * 14:11 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7002.magru.wmnet * 14:11 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti7002.magru.wmnet * 14:09 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324318{{!}}Revert "wmf-config/ProductionServices: set URL for urldownloader to service record" (T429175)]] (duration: 06m 46s) * 14:09 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7001.wikimedia.org with reason: host reimage * 14:06 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti4007.ulsfo.wmnet * 14:05 kharlan@deploy1003: kharlan: Continuing with deployment * 14:05 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti4006.ulsfo.wmnet * 14:04 jayme: updated calico to v3.30.7 on staging-eqiad - [[phab:T427400|T427400]] * 14:04 kharlan@deploy1003: kharlan: Backport for [[gerrit:1324318{{!}}Revert "wmf-config/ProductionServices: set URL for urldownloader to service record" (T429175)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti4006.ulsfo.wmnet * 14:03 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns4003.wikimedia.org with reason: host reimage * 14:03 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7001.wikimedia.org with reason: host reimage * 14:02 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti7002.magru.wmnet * 14:02 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1324318{{!}}Revert "wmf-config/ProductionServices: set URL for urldownloader to service record" (T429175)]] * 14:02 btullis@dns1004: FAIL - running authdns-update * 14:00 btullis@dns1004: START - running authdns-update * 13:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1008.eqiad.wmnet with reason: host reimage * 13:59 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'. * 13:59 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1313985{{!}}wmf-config/ProductionServices: set URL for urldownloader to service record (T429175)]] (duration: 25m 06s) * 13:58 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7002.magru.wmnet * 13:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti4006.ulsfo.wmnet * 13:57 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns4003.wikimedia.org with reason: host reimage * 13:56 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'. * 13:56 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7001.magru.wmnet * 13:56 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1008.eqiad.wmnet with reason: host reimage * 13:55 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'. * 13:55 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'. * 13:55 kharlan@deploy1003: kharlan, sukhe: Continuing with deployment * 13:53 marostegui: Failover ms2 [[phab:T434288|T434288]] * 13:52 marostegui: Failover ms1 [[phab:T434288|T434288]] * 13:52 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7001.magru.wmnet * 13:51 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti4006.ulsfo.wmnet * 13:48 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti4005.ulsfo.wmnet * 13:48 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti4005.ulsfo.wmnet * 13:44 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti4005.ulsfo.wmnet * 13:40 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1008.eqiad.wmnet with OS bookworm * 13:39 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns4003.wikimedia.org with OS trixie * 13:38 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns7001.wikimedia.org with OS trixie * 13:38 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1008.eqiad.wmnet with OS bookworm * 13:37 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti4005.ulsfo.wmnet * 13:36 kharlan@deploy1003: kharlan, sukhe: Backport for [[gerrit:1313985{{!}}wmf-config/ProductionServices: set URL for urldownloader to service record (T429175)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:34 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1313985{{!}}wmf-config/ProductionServices: set URL for urldownloader to service record (T429175)]] * 13:29 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2034.codfw.wmnet * 13:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2034.codfw.wmnet * 13:25 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.clone (exit_code=99) of db1157.eqiad.wmnet onto db1272.eqiad.wmnet * 13:25 cwilliams@cumin1003: START - Cookbook sre.mysql.clone of db1157.eqiad.wmnet onto db1272.eqiad.wmnet * 13:21 urbanecm@deploy1003: mwscript-k8s job started: namespaceDupes.php --wiki=frwiktionary --fix # [[phab:T415716|T415716]] * 13:21 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2034.codfw.wmnet * 13:20 urbanecm@deploy1003: mwscript-k8s job started: namespaceDupes.php --wiki=frwiktionary # [[phab:T415716|T415716]] * 13:19 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1323350{{!}}[tgwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T415307)]], [[gerrit:1322961{{!}}[slwiki] Revert temporary logo for Wikipedia 25 (Vector legacy + Vector 2022) (T414265)]], [[gerrit:1323827{{!}}[itwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T414320)]] (duration: 08m 00s) * 13:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-coord1003.eqiad.wmnet with OS bookworm * 13:15 urbanecm@deploy1003: urbanecm, superpes: Continuing with deployment * 13:13 urbanecm@deploy1003: urbanecm, superpes: Backport for [[gerrit:1323350{{!}}[tgwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T415307)]], [[gerrit:1322961{{!}}[slwiki] Revert temporary logo for Wikipedia 25 (Vector legacy + Vector 2022) (T414265)]], [[gerrit:1323827{{!}}[itwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T414320)]] synced to the testservers (see https://wiki * 13:11 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1323350{{!}}[tgwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T415307)]], [[gerrit:1322961{{!}}[slwiki] Revert temporary logo for Wikipedia 25 (Vector legacy + Vector 2022) (T414265)]], [[gerrit:1323827{{!}}[itwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T414320)]] * 13:11 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1323779{{!}}[ukwiki] Remove reviewer usergroup (T434252)]], [[gerrit:1323312{{!}}[frwiktionary] Add new Schème namespace and its talk (T415716)]] (duration: 06m 49s) * 13:10 marostegui@dns1004: END - running authdns-update * 13:08 marostegui@dns1004: START - running authdns-update * 13:07 marostegui@cumin1003: dbctl commit (dc=all): 'Repool ms2 [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95980 and previous config saved to /var/cache/conftool/dbconfig/20260811-130725-marostegui.json * 13:06 urbanecm@deploy1003: urbanecm, superpes: Continuing with deployment * 13:06 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1266 to dbctl [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95979 and previous config saved to /var/cache/conftool/dbconfig/20260811-130627-marostegui.json * 13:06 urbanecm@deploy1003: urbanecm, superpes: Backport for [[gerrit:1323779{{!}}[ukwiki] Remove reviewer usergroup (T434252)]], [[gerrit:1323312{{!}}[frwiktionary] Add new Schème namespace and its talk (T415716)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:04 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1323779{{!}}[ukwiki] Remove reviewer usergroup (T434252)]], [[gerrit:1323312{{!}}[frwiktionary] Add new Schème namespace and its talk (T415716)]] * 12:59 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db2253.codfw.wmnet,db[1151,1266].eqiad.wmnet with reason: Switching over ms2 * 12:54 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1157: Using as clone source * 12:53 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1157: Using as clone source * 12:51 marostegui@cumin1003: dbctl commit (dc=all): 'Depool ms2 [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95977 and previous config saved to /var/cache/conftool/dbconfig/20260811-125129-marostegui.json * 12:47 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 12:46 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 12:45 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 12:44 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'recommendation-api-ng' for release 'main' . * 12:44 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'recommendation-api-ng' for release 'main' . * 12:43 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'recommendation-api-ng' for release 'main' . * 12:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-coord1003.eqiad.wmnet with reason: host reimage * 12:43 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'ores-legacy' for release 'main' . * 12:42 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'ores-legacy' for release 'main' . * 12:42 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2165: Security update * 12:40 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-coord1003.eqiad.wmnet with reason: host reimage * 12:39 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'ores-legacy' for release 'main' . * 12:38 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' . * 12:38 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' . * 12:37 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' . * 12:34 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2034.codfw.wmnet * 12:30 jmm@cumin2002: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti-test2001.codfw.wmnet * 12:30 jmm@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host ganeti-test2001.codfw.wmnet * 12:25 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1179.eqiad.wmnet onto db1278.eqiad.wmnet * 12:25 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1179: Pool db1179.eqiad.wmnet in after cloning * 12:23 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-coord1003.eqiad.wmnet with OS bookworm * 12:22 moritzm: failover ganeti master in codfw/routed to ganeti2033 * 12:22 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2033.codfw.wmnet * 12:22 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2033.codfw.wmnet * 12:19 jmm@cumin2002: START - Cookbook sre.hosts.reboot-single for host ganeti-test2001.codfw.wmnet * 12:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 12:18 jmm@cumin2002: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti-test2001.codfw.wmnet * 12:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1008.eqiad.wmnet with OS bookworm * 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2033.codfw.wmnet * 12:07 moritzm: failover ganeti master in ganeti/test to ganeti-test2003 * 12:04 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 12:03 jmm@cumin2003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti4005.ulsfo.wmnet * 12:03 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti4005.ulsfo.wmnet * 12:00 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324283{{!}}Use maximum compression level in SqlBlobStore and SqlBagOStuff (T428377)]] (duration: 11m 37s) * 11:57 jmm@cumin2002: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti-test2002.codfw.wmnet * 11:57 jmm@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti-test2002.codfw.wmnet * 11:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2165: Security update * 11:54 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 11:52 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1324283{{!}}Use maximum compression level in SqlBlobStore and SqlBagOStuff (T428377)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:51 jmm@cumin2002: START - Cookbook sre.hosts.reboot-single for host ganeti-test2002.codfw.wmnet * 11:50 jmm@cumin2002: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti-test2002.codfw.wmnet * 11:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2165.codfw.wmnet with reason: Maintenance * 11:48 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@050d19e] (releasing): [[phab:T434186|T434186]] (duration: 01m 14s) * 11:48 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1324283{{!}}Use maximum compression level in SqlBlobStore and SqlBagOStuff (T428377)]] * 11:47 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@050d19e] (releasing): [[phab:T434186|T434186]] * 11:44 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@050d19e] (releasing): test jenkins deploy for [[phab:T434186|T434186]] (duration: 01m 08s) * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2165 [[phab:T434514|T434514]]', diff saved to https://phabricator.wikimedia.org/P95969 and previous config saved to /var/cache/conftool/dbconfig/20260811-114352-cwilliams.json * 11:43 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@050d19e] (releasing): test jenkins deploy for [[phab:T434186|T434186]] * 11:42 jmm@cumin2003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti-test2003.codfw.wmnet * 11:42 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti-test2003.codfw.wmnet * 11:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2161 to s8 primary [[phab:T434514|T434514]]', diff saved to https://phabricator.wikimedia.org/P95968 and previous config saved to /var/cache/conftool/dbconfig/20260811-114136-cwilliams.json * 11:40 cezmunsta: Starting s8 codfw failover from db2165 to db2161 - [[phab:T434514|T434514]] * 11:40 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1179: Pool db1179.eqiad.wmnet in after cloning * 11:36 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti-test2003.codfw.wmnet * 11:36 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti-test2003.codfw.wmnet * 11:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2161 with weight 0 [[phab:T434514|T434514]]', diff saved to https://phabricator.wikimedia.org/P95966 and previous config saved to /var/cache/conftool/dbconfig/20260811-113449-cwilliams.json * 11:34 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 25 hosts with reason: Primary switchover s8 [[phab:T434514|T434514]] * 11:29 moritzm: installing Python 3.11 security updates * 11:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-presto1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 11:26 btullis@cumin1003: START - Cookbook sre.hosts.provision for host an-presto1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 11:23 btullis@dns1004: END - running authdns-update * 11:21 btullis@dns1004: START - running authdns-update * 11:20 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-presto1008.eqiad.wmnet with OS bookworm * 11:20 moritzm: installing curl security updates * 11:11 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1007.eqiad.wmnet with OS bookworm * 10:45 tappof: bump space for prometheus k8s-dse in codfw * 10:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1007.eqiad.wmnet with reason: host reimage * 10:38 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1007.eqiad.wmnet with reason: host reimage * 10:37 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-coord1004.eqiad.wmnet with OS bookworm * 10:35 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1008.eqiad.wmnet with OS bookworm * 10:34 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1179: Depool db1179.eqiad.wmnet to then clone it to db1278.eqiad.wmnet - marostegui@cumin1003 * 10:34 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1008.eqiad.wmnet with OS bookworm * 10:25 fceratto@cumin1003: dbctl commit (dc=all): 'Remove db1177 [[phab:T433474|T433474]]', diff saved to https://phabricator.wikimedia.org/P95964 and previous config saved to /var/cache/conftool/dbconfig/20260811-102527-fceratto.json * 10:22 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1008.eqiad.wmnet with OS bookworm * 10:22 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1007.eqiad.wmnet with OS bookworm * 10:21 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1006.eqiad.wmnet with OS bookworm * 10:20 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 10:18 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1179: Depool db1179.eqiad.wmnet to then clone it to db1278.eqiad.wmnet - marostegui@cumin1003 * 10:18 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1179.eqiad.wmnet onto db1278.eqiad.wmnet * 10:17 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 10:17 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 10:14 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 10:09 blake@deploy1003: Stopping before sync operations * 10:09 blake@deploy1003: Started scap sync-world: Non-deployment run to populate release values for [[phab:T427668|T427668]] * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 10:04 fceratto@cumin1003: Removing db1177 from zarcillo [[phab:T433474|T433474]] * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1177.eqiad.wmnet * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1177.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:03 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1177.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:00 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1006.eqiad.wmnet with reason: host reimage * 09:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-coord1004.eqiad.wmnet with reason: host reimage * 09:57 marostegui: Failover m1 from db1164 to db1213 - [[phab:T434493|T434493]] * 09:57 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1006.eqiad.wmnet with reason: host reimage * 09:55 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:54 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2232].codfw.wmnet,db[1164,1213,1217].eqiad.wmnet with reason: Primary switchover m1 [[phab:T434493|T434493]] * 09:52 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-coord1004.eqiad.wmnet with reason: host reimage * 09:49 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1213.eqiad.wmnet with OS trixie * 09:49 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1177.eqiad.wmnet * 09:41 moritzm: installing Linux 6.12.101 on Trixie hosts * 09:40 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1006.eqiad.wmnet with OS bookworm * 09:35 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-coord1004.eqiad.wmnet with OS bookworm * 09:28 moritzm: installing node-tar security updates * 09:27 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1213.eqiad.wmnet with reason: host reimage * 09:22 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1213.eqiad.wmnet with reason: host reimage * 09:09 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1177: Decommission * 09:08 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db1177: Decommission * 09:08 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 09:08 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.decommission (exit_code=99) * 09:06 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1213.eqiad.wmnet with OS trixie * 09:06 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 09:05 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1213.eqiad.wmnet with reason: Reimage * 08:53 marostegui@dns1004: END - running authdns-update * 08:51 marostegui@dns1004: START - running authdns-update * 08:48 marostegui: Switchover ms1 master in eqiad [[phab:T434288|T434288]] * 08:48 marostegui@cumin1003: dbctl commit (dc=all): 'Repool ms1 [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95962 and previous config saved to /var/cache/conftool/dbconfig/20260811-084804-marostegui.json * 08:40 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1267 to dbctl [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95961 and previous config saved to /var/cache/conftool/dbconfig/20260811-084054-marostegui.json * 08:29 marostegui: Failover m1 from db1213 to db1164 - [[phab:T434043|T434043]] * 08:25 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2232].codfw.wmnet,db[1164,1213,1217].eqiad.wmnet with reason: Primary switchover m1 [[phab:T434043|T434043]] * 08:22 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db2251.codfw.wmnet,db[1152,1267].eqiad.wmnet with reason: Switching over ms1 * 08:22 marostegui@cumin1003: dbctl commit (dc=all): 'Depool ms1 [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95960 and previous config saved to /var/cache/conftool/dbconfig/20260811-082201-marostegui.json * 08:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: Switching over ms1 * 08:20 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.parsercache (exit_code=99) * 08:20 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 08:20 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1152: Switching over ms1 * 08:19 slyngshede@dns1004: END - running authdns-update * 08:18 moritzm: installing openjdk-21 security updates * 08:17 slyngshede@dns1004: START - running authdns-update * 08:16 moritzm: imported jenkins 2.568.2 to thirdparty/jenkins for trixie-wikimedia * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.12 (duration: 02m 26s) * 03:36 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] (duration: 33m 33s) * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 35s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-10 == * 14:54 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1323973{{!}}mmv.bootstrap: Fix getUrlParam to account for TIFF lossy/lossless param (T434333)]] (duration: 11m 24s) * 14:50 krinkle@deploy1003: krinkle: Continuing with deployment * 14:45 krinkle@deploy1003: krinkle: Backport for [[gerrit:1323973{{!}}mmv.bootstrap: Fix getUrlParam to account for TIFF lossy/lossless param (T434333)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:43 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1323973{{!}}mmv.bootstrap: Fix getUrlParam to account for TIFF lossy/lossless param (T434333)]] * 14:07 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1323967{{!}}updateIsActiveFlagForMentees: Commit the final partial batch (T432959)]] (duration: 10m 33s) * 13:56 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1323967{{!}}updateIsActiveFlagForMentees: Commit the final partial batch (T432959)]] * 13:45 dani@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply * 13:45 dani@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply * 13:45 dani@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply * 13:45 dani@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply * 13:45 dani@deploy1003: helmfile [staging] DONE helmfile.d/services/miscweb: apply * 13:44 dani@deploy1003: helmfile [staging] START helmfile.d/services/miscweb: apply * 13:38 wmde-fisch@deploy1003: Finished scap sync-world: Backport for [[gerrit:1323939{{!}}Enable sub-references on more group2 wikis (batch3) (T432731)]] (duration: 33m 21s) * 13:25 wmde-fisch@deploy1003: wmde-fisch: Continuing with deployment * 13:22 wmde-fisch@deploy1003: wmde-fisch: Backport for [[gerrit:1323939{{!}}Enable sub-references on more group2 wikis (batch3) (T432731)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:05 wmde-fisch@deploy1003: Started scap sync-world: Backport for [[gerrit:1323939{{!}}Enable sub-references on more group2 wikis (batch3) (T432731)]] * 07:57 hashar@deploy1003: Finished deploy [integration/docroot@7772132]: update build dependencies (duration: 00m 13s) * 07:57 hashar@deploy1003: Started deploy [integration/docroot@7772132]: update build dependencies * 07:35 _joe_: restarting squid on urldownloader1006 * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 48s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-09 == * 16:01 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:01 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:01 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:00 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 36s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-08 == * 05:31 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9] (wcqs): [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) (duration: 02m 36s) * 05:28 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9] (wcqs): [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) * 04:56 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 04:55 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 04:47 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) (duration: 19m 22s) * 04:28 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) * 04:19 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) (duration: 00m 06s) * 04:18 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) * 04:17 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) (duration: 00m 28s) * 04:16 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) * 03:52 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 03:52 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 34s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-07 == * 23:30 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:29 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 22:45 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 22:43 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 22:41 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 22:41 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 20:54 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 20:32 andrewbogott: restarting puppetserver service on puppetserver* for [[phab:T434339|T434339]] * 19:52 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:45 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 19:32 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:25 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:22 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 19:21 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 18:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:41 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 18:35 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 18:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 18:22 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 18:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 18:16 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 18:12 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:09 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:08 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:07 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:04 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:01 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:00 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:00 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 17:59 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 17:25 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 17:14 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 17:13 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 17:13 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 17:13 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:54 maryum: Deployed security fix for [[phab:T434278|T434278]] * 16:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 16:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 16:27 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --verbose --scoreLessThan=0.7 --exceptDatasetChecksums=[[phab:T434319|T434319]]-enwiki-models.txt # [[phab:T434319|T434319]] * 16:06 cdobbins@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp5022.eqsin.wmnet with OS trixie * 15:13 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 14:19 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 14:17 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 13:50 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1156.eqiad.wmnet onto db1271.eqiad.wmnet * 13:50 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1271: Pool db1271.eqiad.wmnet in after cloning * 13:02 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1271: Pool db1271.eqiad.wmnet in after cloning * 12:19 jayme: updated calico to v3.30.7 on staging-codfw - [[phab:T427400|T427400]] * 12:09 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 12:06 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 12:06 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 12:05 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 12:02 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1156: Pool db1156.eqiad.wmnet in after cloning * 11:38 bjensen: sudo -i reprepro -C main include trixie-wikimedia $<nowiki>{</nowiki>HOME<nowiki>}</nowiki>/httpbb/trixie/httpbb_$<nowiki>{</nowiki>VERSION?<nowiki>}</nowiki>-1+deb13u1_amd64.changes #[[phab:T434052|T434052]] * 11:35 bjensen: sudo -i reprepro -C main include bookworm-wikimedia $<nowiki>{</nowiki>HOME<nowiki>}</nowiki>/httpbb/bookworm/httpbb_$<nowiki>{</nowiki>VERSION?<nowiki>}</nowiki>-1_amd64.changes #[[phab:T434052|T434052]] * 11:30 marostegui@cumin1003: dbctl commit (dc=all): 'Adding db1271 to dbctl', diff saved to https://phabricator.wikimedia.org/P95945 and previous config saved to /var/cache/conftool/dbconfig/20260807-113006-marostegui.json * 11:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1156: Pool db1156.eqiad.wmnet in after cloning * 10:23 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 10:22 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 10:22 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 10:21 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 10:20 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 10:20 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 10:19 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 10:18 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 10:06 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on 21 hosts with reason: cloning * 10:01 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1156: Depool db1156.eqiad.wmnet to then clone it to db1271.eqiad.wmnet - marostegui@cumin1003 * 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1156: Depool db1156.eqiad.wmnet to then clone it to db1271.eqiad.wmnet - marostegui@cumin1003 * 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1156.eqiad.wmnet onto db1271.eqiad.wmnet * 09:15 jynus: started stress testing db1245 dbs [[phab:T431115|T431115]] * 08:19 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:18 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:16 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:14 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:13 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:10 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:06 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:05 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:00 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 10 days, 0:00:00 on ml-serve1015.eqiad.wmnet with reason: Downtime to get full picture of current BIOS settings beyond what Redfish shows * 08:00 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 07:54 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 07:54 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 07:53 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:52 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:51 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:50 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:49 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:48 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:47 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:45 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:45 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:41 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:38 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:37 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 06:35 jayme: updated istio to 1.29.4 on wikikube eqiad - [[phab:T427401|T427401]] * 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1178.eqiad.wmnet * 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1178.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 06:06 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1178.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 05:55 marostegui@cumin1003: START - Cookbook sre.dns.netbox * 05:49 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1178.eqiad.wmnet * 05:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 05:46 marostegui@cumin1003: Removing db1178 from zarcillo [[phab:T433471|T433471]] * 05:45 marostegui@cumin1003: START - Cookbook sre.mysql.decommission * 02:42 denisse: Extended volume on prometheus2008 for the disk space alert as per https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_host_running_out_of_space * 02:37 denisse: Extended volume on prometheus2007 tor the disk space alert as per https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_host_running_out_of_space * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 56s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-06 == * 21:39 maryum: Deploy security patch for [[phab:T433070|T433070]] * 21:29 maryum: Deploy security patch for [[phab:T434189|T434189]] * 20:48 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] (duration: 08m 12s) * 20:44 aude@deploy1003: lmora, aude, anzx: Continuing with deployment * 20:41 aude@deploy1003: lmora, aude, anzx: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be * 20:41 ebernhardson@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:41 ebernhardson@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 20:40 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] * 20:37 ebernhardson@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:37 ebernhardson@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 20:32 ebernhardson@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:32 ebernhardson@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 20:31 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] (duration: 06m 41s) * 20:27 cjming@deploy1003: cjming, ebernhardson, chlod: Continuing with deployment * 20:26 cjming@deploy1003: cjming, ebernhardson, chlod: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:24 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] * 20:18 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] (duration: 09m 22s) * 20:14 cjming@deploy1003: cjming, tsev: Continuing with deployment * 20:11 cjming@deploy1003: cjming, tsev: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:09 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] * 19:41 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply * 19:40 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply * 19:31 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 19:31 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 19:00 cdobbins@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cp5022.eqsin.wmnet with OS trixie * 18:25 ladsgroup@deploy1003: Finished scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) (duration: 06m 08s) * 18:19 ladsgroup@deploy1003: Started scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) * 18:18 ladsgroup@deploy1003: Stopping before sync operations * 18:17 ladsgroup@deploy1003: Started scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) * 17:55 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 16:50 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 16:35 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1001.eqiad.wmnet with OS bookworm * 16:19 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1002.eqiad.wmnet with reason: host reimage * 16:16 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1002.eqiad.wmnet with reason: host reimage * 16:05 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1001.eqiad.wmnet with reason: host reimage * 16:00 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1001.eqiad.wmnet with reason: host reimage * 15:57 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 15:43 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm * 15:29 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1001.eqiad.wmnet with OS bookworm * 15:29 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:58 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm * 14:57 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-drmrs ([[phab:T428495|T428495]]) * 14:55 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-drmrs ([[phab:T428495|T428495]]) * 14:55 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-ui1001.eqiad.wmnet with OS bookworm * 14:54 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-presto1001.eqiad.wmnet with OS bookworm * 14:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-magru ([[phab:T428495|T428495]]) * 14:49 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-magru ([[phab:T428495|T428495]]) * 14:48 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1001.eqiad.wmnet with OS bookworm * 14:46 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-esams ([[phab:T428495|T428495]]) * 14:44 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-esams ([[phab:T428495|T428495]]) * 14:43 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:42 brouberol@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:42 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:42 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 14:40 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 14:40 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:38 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-ui1001.eqiad.wmnet with reason: host reimage * 14:34 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-presto1001.eqiad.wmnet with reason: host reimage * 14:28 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-ui1001.eqiad.wmnet with reason: host reimage * 14:27 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-presto1001.eqiad.wmnet with reason: host reimage * 14:23 sukhe: sudo cumin -b2 'A:cp-text' "run-puppet-agent --enable 'merging CR 1290731'": [[phab:T425441|T425441]] * 14:18 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo for hosts in the wikimedia.org domain - [[phab:T428495|T428495]] * 14:16 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-presto1001.eqiad.wmnet with OS bookworm * 14:14 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-ui1001.eqiad.wmnet with OS bookworm * 14:12 sukhe: sudo cumin 'A:cp-text' "disable-puppet 'merging CR 1290731'": [[phab:T425441|T425441]] * 14:11 swfrench-wmf: restarted navtiming on webperf1003 - [[phab:T428495|T428495]] * 14:04 swfrench-wmf: begin rolling restart of confd in drmrs, eqiad, esams, magru - [[phab:T428495|T428495]] * 14:04 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm * 14:04 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:02 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-client1002.eqiad.wmnet with OS bookworm * 13:58 swfrench-wmf: authdns update to direct eqiad-associated etcd clients back to eqiad - [[phab:T428495|T428495]] * 13:58 swfrench@dns1004: END - running authdns-update * 13:56 swfrench@dns1004: START - running authdns-update * 13:49 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:44 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:31 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 13:29 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 13:26 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 13:23 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 13:22 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 13:19 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 13:18 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 13:18 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 13:17 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 13:16 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 13:13 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 13:11 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 13:09 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 13:06 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 13:06 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-client1002.eqiad.wmnet with OS bookworm * 13:05 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revision-models' for release 'main' . * 13:05 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:05 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revision-models' for release 'main' . * 13:04 brouberol@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-test-client1002.eqiad.wmnet with OS bookworm * 13:04 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 13:03 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 13:02 aikochou@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:00 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'readability' for release 'main' . * 12:59 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'readability' for release 'main' . * 12:58 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 12:57 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'logo-detection' for release 'main' . * 12:57 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'logo-detection' for release 'main' . * 12:57 aikochou@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 12:55 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:54 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:53 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 12:53 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:50 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 12:48 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 12:46 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'article-models' for release 'main' . * 12:45 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'article-models' for release 'main' . * 12:41 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . * 12:39 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . * 12:38 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-client1002.eqiad.wmnet with OS bookworm * 12:12 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply * 12:12 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply * 12:09 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:08 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 11:58 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2187: Security update * 11:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:24 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:16 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 11:15 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 11:10 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2187: Security update * 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2187.codfw.wmnet with reason: Maintenance * 10:56 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 10:56 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 10:56 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 10:56 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 10:54 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 10:53 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 10:09 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2187: Security update * 10:07 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2187: Security update * 09:39 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms2', diff saved to https://phabricator.wikimedia.org/P95929 and previous config saved to /var/cache/conftool/dbconfig/20260806-093908-marostegui.json * 09:36 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1178 from dbctl [[phab:T433471|T433471]]', diff saved to https://phabricator.wikimedia.org/P95928 and previous config saved to /var/cache/conftool/dbconfig/20260806-093632-marostegui.json * 09:33 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 09:31 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 09:30 topranks: bounce cr3-eqsin<->cr2-eqiad bgp session to disable no-prepend command * 09:20 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2253.codfw.wmnet,db1151.eqiad.wmnet with reason: cloning * 09:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1151: Cloning * 09:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:19 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 09:19 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1151: Cloning * 09:10 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 09:09 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup2003.codfw.wmnet * 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup2003.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 09:06 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup2003.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 09:03 klausman@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:02 jynus@cumin1003: START - Cookbook sre.dns.netbox * 09:02 klausman@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 08:57 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup2003.codfw.wmnet * 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup1003.eqiad.wmnet * 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:54 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms3', diff saved to https://phabricator.wikimedia.org/P95925 and previous config saved to /var/cache/conftool/dbconfig/20260806-085422-marostegui.json * 08:53 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:46 jynus@cumin1003: START - Cookbook sre.dns.netbox * 08:39 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup1003.eqiad.wmnet * 08:29 XioNoX: push pfw policy - [[phab:T434115|T434115]] * 08:14 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 08:00 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 08:00 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:58 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revision-models' for release 'main' . * 07:56 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 07:54 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'readability' for release 'main' . * 07:53 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'logo-detection' for release 'main' . * 07:51 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'llm' for release 'main' . * 07:48 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . * 07:37 jayme: updated istio to 1.29.4 on wikikube codfw - [[phab:T427401|T427401]] * 07:08 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2252.codfw.wmnet,db1153.eqiad.wmnet with reason: cloning * 07:07 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1153: Cloning * 07:07 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1153: Cloning * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 40s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-05 == * 23:24 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1009.eqiad.wmnet with OS bookworm * 23:03 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1009.eqiad.wmnet with reason: host reimage * 22:59 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1009.eqiad.wmnet with reason: host reimage * 22:43 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1009.eqiad.wmnet with OS bookworm * 22:38 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1009.eqiad.wmnet * 22:34 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1009.eqiad.wmnet * 22:25 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1008.eqiad.wmnet with OS bookworm * 22:04 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1008.eqiad.wmnet with reason: host reimage * 22:00 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1008.eqiad.wmnet with reason: host reimage * 21:48 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:47 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:46 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:44 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1008.eqiad.wmnet with OS bookworm * 21:43 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:41 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1008.eqiad.wmnet * 21:36 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1008.eqiad.wmnet * 21:14 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:12 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad * 21:12 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad * 21:10 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=eqiad * 21:08 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:07 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:07 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=eqiad * 21:04 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:03 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1006 * 21:02 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1006 * 21:00 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:56 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:56 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:55 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 20:55 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 20:55 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1007.eqiad.wmnet with OS bookworm * 20:51 vriley@cumin1003: START - Cookbook sre.dns.netbox * 20:43 ebernhardson: [[phab:T434008|T434008]]: changing cloudelastic:9643 from auto_expand_replicas to number_of_replicas * 20:34 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1007.eqiad.wmnet with reason: host reimage * 20:27 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1007.eqiad.wmnet with reason: host reimage * 20:24 cjming: end of UTC late backport window * 20:23 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] (duration: 06m 26s) * 20:18 cjming@deploy1003: cjming: Continuing with deployment * 20:18 cjming@deploy1003: cjming: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:16 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] * 20:12 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1007.eqiad.wmnet with OS bookworm * 20:12 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] (duration: 08m 41s) * 20:08 swfrench@cumin2002: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host conf1007.eqiad.wmnet with OS bookworm * 20:08 jforrester@deploy1003: jforrester: Continuing with deployment * 20:07 jforrester@deploy1003: jforrester: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:03 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] * 19:51 inflatador: [bking@puppetserver1001] ~$ sudo puppetserver ca sign --certname an-worker1189.eqiad.wmnet [[phab:T434142|T434142]] * 19:47 bking@cumin2003: DONE (FAIL) - Cookbook sre.puppet.renew-cert (exit_code=99) for an-worker1189.eqiad.wmnet: Renew puppet certificate - bking@cumin2003 * 19:46 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:30 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1007.eqiad.wmnet with OS trixie * 19:30 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 19:29 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 19:20 swfrench-wmf: silenced EtcdRelicationDown 0cb709a9-f244-4f1e-971f-{{Gerrit|440ec65e7fd7}} - [[phab:T428495|T428495]] * 19:13 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1007.eqiad.wmnet with OS bookworm * 19:12 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1007.eqiad.wmnet with reason: host reimage * 19:09 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1007.eqiad.wmnet * 19:07 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1007.eqiad.wmnet with reason: host reimage * 19:03 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1007.eqiad.wmnet * 18:52 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1007.eqiad.wmnet with OS trixie * 18:52 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1007.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:35 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 18:34 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 18:34 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 18:30 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1007.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:28 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:28 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1007] - vriley@cumin1003" * 18:27 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1007] - vriley@cumin1003" * 18:23 vriley@cumin1003: START - Cookbook sre.dns.netbox * 18:22 vriley@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 18:22 robh@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:19 vriley@cumin1003: START - Cookbook sre.dns.netbox * 18:13 robh@cumin2002: START - Cookbook sre.hosts.provision for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:31 jasmine@cumin2002: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-main-eqiad * 17:12 mutante: LDAP - added vwalters to group ciadmin - [[phab:T433615|T433615]] * 16:58 aokoth@deploy1003: Finished deploy [phabricator/deployment@e2ebca5]: Deploy Phab (duration: 00m 34s) * 16:57 aokoth@deploy1003: Started deploy [phabricator/deployment@e2ebca5]: Deploy Phab * 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad * 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=eqiad * 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad * 16:53 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:41 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2187.codfw.wmnet * 16:41 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2187.codfw.wmnet * 16:41 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker2187.codfw.wmnet * 16:41 cgoubert@cumin2003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker2187.codfw.wmnet * 16:40 jasmine@cumin2002: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-main-eqiad * 16:40 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:34 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:25 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-magru and A:liberica ([[phab:T428495|T428495]]) * 16:23 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-magru and A:liberica ([[phab:T428495|T428495]]) * 16:20 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-drmrs and A:liberica ([[phab:T428495|T428495]]) * 16:19 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-drmrs and A:liberica ([[phab:T428495|T428495]]) * 16:18 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-esams and A:liberica ([[phab:T428495|T428495]]) * 16:16 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-esams and A:liberica ([[phab:T428495|T428495]]) * 16:06 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1159.eqiad.wmnet * 16:06 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1159.eqiad.wmnet * 16:06 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1159.eqiad.wmnet * 16:05 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] (duration: 09m 11s) * 15:58 reedy@deploy1003: reedy: Continuing with deployment * 15:58 reedy@deploy1003: reedy: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:56 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] * 15:54 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1159.eqiad.wmnet with OS trixie * 15:38 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:33 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1159.eqiad.wmnet with reason: host reimage * 15:32 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:27 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1159.eqiad.wmnet with reason: host reimage * 15:10 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1159 * 15:10 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1159 * 15:00 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo for hosts in the wikimedia.org domain - [[phab:T428495|T428495]] * 14:55 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS trixie * 14:54 swfrench-wmf: restarted navtiming on webperf1003 - [[phab:T428495|T428495]] * 14:52 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1159 * 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1159.eqiad.wmnet 129.48.64.10.in-addr.arpa 9.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:52 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1159.eqiad.wmnet 129.48.64.10.in-addr.arpa 9.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1159 - jayme@cumin1003" * 14:52 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1159 - jayme@cumin1003" * 14:49 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:48 jayme@cumin1003: START - Cookbook sre.dns.netbox * 14:47 swfrench-wmf: begin rolling restart of confd in drmrs, eqiad, esams, magru - [[phab:T428495|T428495]] * 14:47 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1159 * 14:46 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:46 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:46 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1159.eqiad.wmnet with OS trixie * 14:45 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:44 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:44 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1159.eqiad.wmnet * 14:43 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:43 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1159.eqiad.wmnet * 14:43 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:43 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1159.eqiad.wmnet * 14:43 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:43 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1157.eqiad.wmnet * 14:43 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1157.eqiad.wmnet * 14:43 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1157.eqiad.wmnet * 14:42 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:42 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:42 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:41 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:41 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:41 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:41 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1003.eqiad.wmnet with OS bookworm * 14:39 swfrench-wmf: authdns update to direct eqiad-associated etcd clients to codfw - [[phab:T428495|T428495]] * 14:39 swfrench@dns1004: END - running authdns-update * 14:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 14:37 swfrench@dns1004: START - running authdns-update * 14:37 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:37 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:35 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:35 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 14:28 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:27 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1157.eqiad.wmnet with OS trixie * 14:27 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:27 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:27 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 14:26 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 14:26 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:26 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 14:26 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 14:26 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:26 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host search-loader1002.eqiad.wmnet with OS trixie * 14:19 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:19 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:15 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1003.eqiad.wmnet with reason: host reimage * 14:14 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:14 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:13 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 * 14:13 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host mc2046 * 14:13 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS trixie * 14:11 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1003.eqiad.wmnet with reason: host reimage * 14:10 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:09 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:09 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:09 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:08 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:08 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1157.eqiad.wmnet with reason: host reimage * 14:08 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:04 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 14:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on search-loader1002.eqiad.wmnet with reason: host reimage * 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=eqiad * 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=eqiad * 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=eqiad * 14:00 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:59 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:58 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1157.eqiad.wmnet with reason: host reimage * 13:57 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on search-loader1002.eqiad.wmnet with reason: host reimage * 13:54 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1003.eqiad.wmnet with OS bookworm * 13:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host search-loader1002.eqiad.wmnet with OS trixie * 13:43 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1157 * 13:42 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1157 * 13:40 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1157 * 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1157.eqiad.wmnet 183.32.64.10.in-addr.arpa 3.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1157.eqiad.wmnet 183.32.64.10.in-addr.arpa 3.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1157 - jayme@cumin1003" * 13:39 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1157 - jayme@cumin1003" * 13:39 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] (duration: 07m 00s) * 13:35 jayme@cumin1003: START - Cookbook sre.dns.netbox * 13:35 reedy@deploy1003: reedy: Continuing with deployment * 13:34 reedy@deploy1003: reedy: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:32 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] * 13:23 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1157 * 13:22 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1157.eqiad.wmnet with OS trixie * 13:22 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1157.eqiad.wmnet * 13:22 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1157.eqiad.wmnet * 13:21 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1157.eqiad.wmnet * 13:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1156.eqiad.wmnet * 13:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1156.eqiad.wmnet * 13:15 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1156.eqiad.wmnet * 13:01 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1156.eqiad.wmnet with OS trixie * 12:42 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1156.eqiad.wmnet with reason: host reimage * 12:38 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1156.eqiad.wmnet with reason: host reimage * 12:32 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 12:31 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 12:30 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 12:28 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 12:26 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 12:24 topranks: update bgp confed settings in eqsin * 12:22 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1156 * 12:22 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1156 * 12:22 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 12:19 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1156 * 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1156.eqiad.wmnet 110.32.64.10.in-addr.arpa 0.1.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:19 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1156.eqiad.wmnet 110.32.64.10.in-addr.arpa 0.1.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1156 - jayme@cumin1003" * 12:19 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1156 - jayme@cumin1003" * 12:17 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:14 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS trixie * 12:09 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:06 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:04 jayme@cumin1003: START - Cookbook sre.dns.netbox * 12:04 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 12:02 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'article-models' for release 'main' . * 12:01 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1156 * 12:01 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1156.eqiad.wmnet with OS trixie * 11:59 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1156.eqiad.wmnet * 11:59 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1156.eqiad.wmnet * 11:59 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1156.eqiad.wmnet * 11:57 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 11:53 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 11:53 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:52 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:52 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:50 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:50 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:50 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:49 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:48 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:47 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:47 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:45 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:45 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:44 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:44 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:44 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:43 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:42 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:38 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:35 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 * 11:35 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host mc2046 * 11:34 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS trixie * 11:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:27 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:21 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:21 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:18 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:18 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:18 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:18 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:13 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:13 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:09 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:08 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:07 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:06 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:06 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:05 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:05 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:04 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 11:04 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:24 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:24 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:17 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:16 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1155.eqiad.wmnet * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1155.eqiad.wmnet * 10:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1155.eqiad.wmnet * 10:14 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:14 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:11 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:11 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:10 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:09 aikochou@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop: sync * 10:09 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:09 aikochou@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop: sync * 10:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:07 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:05 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:05 aikochou@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop: sync * 10:05 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:05 aikochou@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop: sync * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:04 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:04 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:04 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1155.eqiad.wmnet with OS trixie * 09:52 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms1', diff saved to https://phabricator.wikimedia.org/P95918 and previous config saved to /var/cache/conftool/dbconfig/20260805-095212-marostegui.json * 09:44 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1152: after cloning * 09:44 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.parsercache (exit_code=99) * 09:44 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 09:44 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1152: after cloning * 09:43 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1155.eqiad.wmnet with reason: host reimage * 09:40 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1155.eqiad.wmnet with reason: host reimage * 09:32 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 09:32 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:31 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 09:31 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:27 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1155 * 09:27 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1155 * 09:25 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 09:24 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 09:24 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 09:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:23 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 09:23 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 09:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:22 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 09:22 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 09:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:20 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 09:20 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:17 XioNoX: push pfw policies - [[phab:T434038|T434038]] * 09:14 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1155 * 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1155.eqiad.wmnet 109.32.64.10.in-addr.arpa 9.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1155.eqiad.wmnet 109.32.64.10.in-addr.arpa 9.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1155 - jayme@cumin1003" * 09:14 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1155 - jayme@cumin1003" * 09:10 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 09:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:09 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2251.codfw.wmnet,db1152.eqiad.wmnet with reason: cloning * 09:09 jayme@cumin1003: START - Cookbook sre.dns.netbox * 09:08 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 09:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: Cloning * 09:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:05 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 09:05 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1152: Cloning * 08:38 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1155 * 08:37 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1155.eqiad.wmnet with OS trixie * 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1171.eqiad.wmnet * 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1171.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:29 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95913 and previous config saved to /var/cache/conftool/dbconfig/20260805-082908-ladsgroup.json * 08:27 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1171.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:22 jynus@cumin1003: START - Cookbook sre.dns.netbox * 08:18 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249', diff saved to https://phabricator.wikimedia.org/P95912 and previous config saved to /var/cache/conftool/dbconfig/20260805-081823-ladsgroup.json * 08:17 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1171.eqiad.wmnet * 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1150.eqiad.wmnet * 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1150.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:15 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1150.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:15 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 08:14 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1155.eqiad.wmnet * 08:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1155.eqiad.wmnet * 08:14 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1155.eqiad.wmnet * 08:11 jynus@cumin1003: START - Cookbook sre.dns.netbox * 08:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249', diff saved to https://phabricator.wikimedia.org/P95911 and previous config saved to /var/cache/conftool/dbconfig/20260805-080737-ladsgroup.json * 08:05 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1150.eqiad.wmnet * 08:02 marostegui: Depool clouddb1020 (s5,s8) [[phab:T434048|T434048]] * 08:02 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1020.eqiad.wmnet,service=s8 * 08:02 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1020.eqiad.wmnet,service=s5 * 08:02 marostegui: Depool clouddb1018 (s2,s7) [[phab:T434048|T434048]] * 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1018.eqiad.wmnet,service=s7 * 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1018.eqiad.wmnet,service=s2 * 08:01 marostegui: Depool clouddb1017 (s1) [[phab:T434048|T434048]] * 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1017.eqiad.wmnet,service=s1 * 07:59 marostegui: Depool clouddb1016 (s5,s8) [[phab:T434048|T434048]] * 07:59 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s8 * 07:59 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s5 * 07:57 marostegui: Depool clouddb1015 (s4,s6) [[phab:T434048|T434048]] * 07:57 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s6 * 07:57 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s4 * 07:56 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95910 and previous config saved to /var/cache/conftool/dbconfig/20260805-075650-ladsgroup.json * 07:54 marostegui: Depool clouddb1014 (s2,s7) [[phab:T434048|T434048]] * 07:54 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1014.eqiad.wmnet,service=s7 * 07:54 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1014.eqiad.wmnet,service=s2 * 07:53 marostegui: Depool clouddb1013:s1 [[phab:T434048|T434048]] * 07:53 marostegui: Depool clouddb1013:s1 [[phab:T409557|T409557]] * 07:53 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1013.eqiad.wmnet,service=s1 * 07:25 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95909 and previous config saved to /var/cache/conftool/dbconfig/20260805-072529-ladsgroup.json * 07:24 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2249.codfw.wmnet with reason: Maintenance * 07:24 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95908 and previous config saved to /var/cache/conftool/dbconfig/20260805-072426-ladsgroup.json * 07:21 slyngshede@dns1004: END - running authdns-update * 07:19 slyngshede@dns1004: START - running authdns-update * 07:13 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231', diff saved to https://phabricator.wikimedia.org/P95906 and previous config saved to /var/cache/conftool/dbconfig/20260805-071340-ladsgroup.json * 07:02 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231', diff saved to https://phabricator.wikimedia.org/P95905 and previous config saved to /var/cache/conftool/dbconfig/20260805-070253-ladsgroup.json * 06:52 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95904 and previous config saved to /var/cache/conftool/dbconfig/20260805-065206-ladsgroup.json * 06:45 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 06:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95903 and previous config saved to /var/cache/conftool/dbconfig/20260805-062240-ladsgroup.json * 06:21 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2231.codfw.wmnet with reason: Maintenance * 06:21 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95902 and previous config saved to /var/cache/conftool/dbconfig/20260805-062137-ladsgroup.json * 06:10 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215', diff saved to https://phabricator.wikimedia.org/P95901 and previous config saved to /var/cache/conftool/dbconfig/20260805-061051-ladsgroup.json * 06:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215', diff saved to https://phabricator.wikimedia.org/P95900 and previous config saved to /var/cache/conftool/dbconfig/20260805-060004-ladsgroup.json * 05:49 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95899 and previous config saved to /var/cache/conftool/dbconfig/20260805-054918-ladsgroup.json * 05:19 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95898 and previous config saved to /var/cache/conftool/dbconfig/20260805-051939-ladsgroup.json * 05:18 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2215.codfw.wmnet with reason: Maintenance * 04:30 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2201.codfw.wmnet with reason: Maintenance * 03:40 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2197.codfw.wmnet with reason: Maintenance * 03:40 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95897 and previous config saved to /var/cache/conftool/dbconfig/20260805-034036-ladsgroup.json * 03:29 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196', diff saved to https://phabricator.wikimedia.org/P95896 and previous config saved to /var/cache/conftool/dbconfig/20260805-032948-ladsgroup.json * 03:19 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196', diff saved to https://phabricator.wikimedia.org/P95895 and previous config saved to /var/cache/conftool/dbconfig/20260805-031902-ladsgroup.json * 03:08 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95894 and previous config saved to /var/cache/conftool/dbconfig/20260805-030815-ladsgroup.json * 02:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95893 and previous config saved to /var/cache/conftool/dbconfig/20260805-023413-ladsgroup.json * 02:33 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2196.codfw.wmnet with reason: Maintenance * 02:33 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95892 and previous config saved to /var/cache/conftool/dbconfig/20260805-023310-ladsgroup.json * 02:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186', diff saved to https://phabricator.wikimedia.org/P95891 and previous config saved to /var/cache/conftool/dbconfig/20260805-022223-ladsgroup.json * 02:11 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186', diff saved to https://phabricator.wikimedia.org/P95890 and previous config saved to /var/cache/conftool/dbconfig/20260805-021137-ladsgroup.json * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 02:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95889 and previous config saved to /var/cache/conftool/dbconfig/20260805-020051-ladsgroup.json * 01:30 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95888 and previous config saved to /var/cache/conftool/dbconfig/20260805-013029-ladsgroup.json * 01:29 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2186.codfw.wmnet with reason: Maintenance * 00:34 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on dbstore1009.eqiad.wmnet with reason: Maintenance * 00:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95887 and previous config saved to /var/cache/conftool/dbconfig/20260805-003408-ladsgroup.json * 00:23 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264', diff saved to https://phabricator.wikimedia.org/P95886 and previous config saved to /var/cache/conftool/dbconfig/20260805-002322-ladsgroup.json * 00:12 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264', diff saved to https://phabricator.wikimedia.org/P95885 and previous config saved to /var/cache/conftool/dbconfig/20260805-001235-ladsgroup.json * 00:01 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95884 and previous config saved to /var/cache/conftool/dbconfig/20260805-000148-ladsgroup.json == 2026-08-04 == * 23:45 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95883 and previous config saved to /var/cache/conftool/dbconfig/20260804-234508-ladsgroup.json * 23:44 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1264.eqiad.wmnet with reason: Maintenance * 23:44 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95882 and previous config saved to /var/cache/conftool/dbconfig/20260804-234405-ladsgroup.json * 23:33 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237', diff saved to https://phabricator.wikimedia.org/P95881 and previous config saved to /var/cache/conftool/dbconfig/20260804-233317-ladsgroup.json * 23:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237', diff saved to https://phabricator.wikimedia.org/P95880 and previous config saved to /var/cache/conftool/dbconfig/20260804-232230-ladsgroup.json * 23:11 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95879 and previous config saved to /var/cache/conftool/dbconfig/20260804-231144-ladsgroup.json * 22:23 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95878 and previous config saved to /var/cache/conftool/dbconfig/20260804-222345-ladsgroup.json * 22:23 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1237.eqiad.wmnet with reason: Maintenance * 21:13 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1225.eqiad.wmnet with reason: Maintenance * 20:40 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] (duration: 24m 40s) * 20:33 samtar@deploy1003: samtar, kineticpelagic: Continuing with deployment * 20:28 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS bookworm * 20:21 samtar@deploy1003: samtar, kineticpelagic: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:15 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] * 20:13 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 20:09 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 20:00 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1216.eqiad.wmnet with reason: Maintenance * 20:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95877 and previous config saved to /var/cache/conftool/dbconfig/20260804-195957-ladsgroup.json * 19:51 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 19:50 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) mc2046.codfw.wmnet 120.16.192.10.in-addr.arpa 0.2.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:50 jhancock@cumin2002: START - Cookbook sre.dns.wipe-cache mc2046.codfw.wmnet 120.16.192.10.in-addr.arpa 0.2.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host mc2046 - jhancock@cumin2002" * 19:50 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host mc2046 - jhancock@cumin2002" * 19:49 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203', diff saved to https://phabricator.wikimedia.org/P95876 and previous config saved to /var/cache/conftool/dbconfig/20260804-194911-ladsgroup.json * 19:46 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 19:45 jhancock@cumin2002: START - Cookbook sre.hosts.move-vlan for host mc2046 * 19:45 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS bookworm * 19:38 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203', diff saved to https://phabricator.wikimedia.org/P95875 and previous config saved to /var/cache/conftool/dbconfig/20260804-193825-ladsgroup.json * 19:27 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95874 and previous config saved to /var/cache/conftool/dbconfig/20260804-192738-ladsgroup.json * 19:02 mutante: gerrit ssh -p 29418 gerrit.wikimedia.org gerrit index changes {{Gerrit|1320979}} * 18:20 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 18:18 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 18:14 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 18:14 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 18:13 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 18:10 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 18:08 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 18:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95872 and previous config saved to /var/cache/conftool/dbconfig/20260804-180721-ladsgroup.json * 18:07 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 18:06 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1203.eqiad.wmnet with reason: Maintenance * 18:06 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95871 and previous config saved to /var/cache/conftool/dbconfig/20260804-180618-ladsgroup.json * 17:55 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179', diff saved to https://phabricator.wikimedia.org/P95870 and previous config saved to /var/cache/conftool/dbconfig/20260804-175531-ladsgroup.json * 17:55 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1154.eqiad.wmnet * 17:55 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1154.eqiad.wmnet * 17:55 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1154.eqiad.wmnet * 17:50 swfrench@deploy1003: Finished scap sync-world: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] (duration: 04m 05s) * 17:48 swfrench@deploy1003: swfrench: Continuing with deployment * 17:46 swfrench@deploy1003: swfrench: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:45 swfrench@deploy1003: Started scap sync-world: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] * 17:44 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179', diff saved to https://phabricator.wikimedia.org/P95869 and previous config saved to /var/cache/conftool/dbconfig/20260804-174445-ladsgroup.json * 17:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95868 and previous config saved to /var/cache/conftool/dbconfig/20260804-173359-ladsgroup.json * 17:33 swfrench@deploy1003: Finished scap sync-world: Pick up new PHP production image (duration: 28m 32s) * 17:28 aokoth@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on phab1005.eqiad.wmnet with reason: Puppet Failure * 17:05 swfrench@deploy1003: Started scap sync-world: Pick up new PHP production image * 17:00 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 17:00 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 16:54 cgoubert@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on wikikube-worker2187.codfw.wmnet with reason: Hardware issue * 16:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2187.codfw.wmnet * 16:52 mutante: gerrit2003:/var/log/apache2# ln -s /srv/gerrit/site_path/review_site/logs/ gerrit * 16:52 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2187.codfw.wmnet * 16:48 mutante: gerrit2003 - moving old apache logfiles older than 60 days from /var/log/apache2 to /srv/gerrit/site_path/review_site/logs/old/ * 16:33 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 16:32 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 16:29 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 16:29 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 16:28 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 16:28 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 16:27 dzahn@cumin1003: END (PASS) - Cookbook sre.gerrit.restart-gerrit (exit_code=0) Restarting Gerrit on gerrit2003 * 16:27 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 16:27 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95867 and previous config saved to /var/cache/conftool/dbconfig/20260804-162736-ladsgroup.json * 16:27 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 16:26 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1179.eqiad.wmnet with reason: Maintenance * 16:26 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:25 mutante: restarting gerrit - dropped outdated RSA host key * 16:25 dzahn@cumin1003: START - Cookbook sre.gerrit.restart-gerrit Restarting Gerrit on gerrit2003 * 16:24 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95866 and previous config saved to /var/cache/conftool/dbconfig/20260804-162424-ladsgroup.json * 16:24 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 16:23 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 16:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95865 and previous config saved to /var/cache/conftool/dbconfig/20260804-162236-ladsgroup.json * 16:21 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1179.eqiad.wmnet with reason: Maintenance * 16:17 swfrench-wmf: reprepro include php8.3_8.3.33-1+wmf11u1 into component/php83 for bullseye-wikimedia * 16:17 swfrench-wmf: reprepro include php8.3_8.3.33-1+wmf12u1 into component/php83 for bookworm-wikimedia * 16:11 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply * 16:10 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply * 16:10 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mobileapps: apply * 16:09 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mobileapps: apply * 16:09 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply * 16:08 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply * 16:08 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:08 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:07 aokoth@cumin1003: END (PASS) - Cookbook sre.vrts.upgrade (exit_code=0) on VRTS host vrts1003.eqiad.wmnet * 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:05 aokoth@cumin1003: START - Cookbook sre.vrts.upgrade on VRTS host vrts1003.eqiad.wmnet * 16:04 mutante: gerrit2002/gerrit1003/gerrit2003 - rm /etc/gerrit/ssh_host_rsa_key * 15:59 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:59 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:59 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:59 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:56 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 15:55 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:55 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:55 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:49 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 15:49 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:44 Raine: add php8.5 packages to component/php85 - [[phab:T432983|T432983]] * 15:39 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:33 aaron@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 15:33 aaron@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 15:29 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:19 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:19 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:16 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:16 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1154.eqiad.wmnet with OS trixie * 15:16 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:15 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:15 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:06 brennen@deploy1003: Finished deploy [phabricator/deployment@56f4ffd]: deploy phab1004 for [[phab:T433981|T433981]] (duration: 00m 43s) * 15:05 brennen@deploy1003: Started deploy [phabricator/deployment@56f4ffd]: deploy phab1004 for [[phab:T433981|T433981]] * 15:05 aaron@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 15:04 aaron@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 15:02 brennen@deploy1003: Finished deploy [phabricator/deployment@56f4ffd]: deploy phab2003 for [[phab:T433981|T433981]] (duration: 00m 51s) * 15:01 brennen@deploy1003: Started deploy [phabricator/deployment@56f4ffd]: deploy phab2003 for [[phab:T433981|T433981]] * 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1004.eqiad.wmnet with reason: deployment * 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1005.eqiad.wmnet with reason: deployment * 14:58 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab2003.codfw.wmnet with reason: deployment * 14:55 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1154.eqiad.wmnet with reason: host reimage * 14:51 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1154.eqiad.wmnet with reason: host reimage * 14:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 14:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 14:38 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync * 14:38 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync * 14:38 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync * 14:37 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync * 14:37 ottomata: roll restart eventgate-main to pick up stream config change - [[phab:T433507|T433507]] * 14:37 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-main: sync * 14:36 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-main: sync * 14:36 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1154 * 14:36 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1154 * 14:34 otto@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] (duration: 08m 39s) * 14:34 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1154 * 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1154.eqiad.wmnet 108.32.64.10.in-addr.arpa 8.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:34 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1154.eqiad.wmnet 108.32.64.10.in-addr.arpa 8.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1154 - jayme@cumin1003" * 14:34 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1154 - jayme@cumin1003" * 14:30 otto@deploy1003: otto: Continuing with deployment * 14:30 jayme@cumin1003: START - Cookbook sre.dns.netbox * 14:28 otto@deploy1003: otto: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:26 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1154 * 14:26 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1154.eqiad.wmnet with OS trixie * 14:26 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1154.eqiad.wmnet * 14:26 otto@deploy1003: Started scap sync-world: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] * 14:26 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1154.eqiad.wmnet * 14:26 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1154.eqiad.wmnet * 14:17 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 14:16 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 14:15 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 14:14 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 14:13 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 14:13 swfrench@dns1004: END - running authdns-update * 14:13 Msz2001: Finished deployments for UTC afternoon backport window * 14:13 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 14:13 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] (duration: 07m 58s) * 14:11 swfrench@dns1004: START - running authdns-update * 14:08 mszwarc@deploy1003: javiermonton, mszwarc, mpostoronca: Continuing with deployment * 14:07 mszwarc@deploy1003: javiermonton, mszwarc, mpostoronca: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] synced to the testser * 14:05 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] * 14:03 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 13:49 swfrench@cumin2002: conftool action : set/pooled=yes; selector: name=wikikube-worker2330.codfw.wmnet * 13:49 swfrench@cumin2002: conftool action : set/pooled=no; selector: name=wikikube-worker2330.codfw.wmnet * 13:48 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] (duration: 09m 19s) * 13:45 swfrench@dns1004: END - running authdns-update * 13:44 mszwarc@deploy1003: mszwarc, jforrester: Continuing with deployment * 13:43 swfrench@dns1004: START - running authdns-update * 13:41 mszwarc@deploy1003: mszwarc, jforrester: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:38 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] * 13:33 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 13:33 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1154.eqiad.wmnet * 13:32 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 13:32 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 13:31 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 13:31 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:31 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:29 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1154.eqiad.wmnet * 13:28 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1154.eqiad.wmnet * 13:28 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1154.eqiad.wmnet * 13:28 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1141.eqiad.wmnet * 13:28 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1141.eqiad.wmnet * 13:28 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1141.eqiad.wmnet * 13:22 otto@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply * 13:22 otto@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply * 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1096.eqiad.wmnet with OS trixie * 13:05 swfrench@dns1004: END - running authdns-update * 13:03 swfrench@dns1004: START - running authdns-update * 12:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 12:43 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 1:00:00 on db1171.eqiad.wmnet with reason: decom * 12:42 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 1:00:00 on db1150.eqiad.wmnet with reason: decom * 12:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 12:38 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1164,1217].eqiad.wmnet with reason: cloning * 12:33 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2096.codfw.wmnet with OS trixie * 12:22 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1096.eqiad.wmnet with OS trixie * 12:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2096.codfw.wmnet with reason: host reimage * 12:14 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1141.eqiad.wmnet with OS trixie * 12:10 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2096.codfw.wmnet with reason: host reimage * 12:10 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1289.eqiad.wmnet * 12:05 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1289.eqiad.wmnet * 12:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1288.eqiad.wmnet * 11:59 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1288.eqiad.wmnet * 11:59 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1287.eqiad.wmnet * 11:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1097.eqiad.wmnet with OS trixie * 11:54 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1287.eqiad.wmnet * 11:54 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1286.eqiad.wmnet * 11:53 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1141.eqiad.wmnet with reason: host reimage * 11:51 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2096.codfw.wmnet with OS trixie * 11:49 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1141.eqiad.wmnet with reason: host reimage * 11:48 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1286.eqiad.wmnet * 11:48 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1284.eqiad.wmnet * 11:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2095.codfw.wmnet with OS trixie * 11:43 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1284.eqiad.wmnet * 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1283.eqiad.wmnet * 11:42 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on ml-serve1015.eqiad.wmnet with reason: Downtime to get full picture of current BIOS settings beyond what Redfish shows * 11:39 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad * 11:39 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:37 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1283.eqiad.wmnet * 11:37 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1282.eqiad.wmnet * 11:37 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad * 11:37 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:33 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1141 * 11:33 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1141 * 11:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 11:32 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1141 * 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1141.eqiad.wmnet 156.48.64.10.in-addr.arpa 6.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:32 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1141.eqiad.wmnet 156.48.64.10.in-addr.arpa 6.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1141 - jayme@cumin1003" * 11:32 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1141 - jayme@cumin1003" * 11:32 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1282.eqiad.wmnet * 11:32 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1281.eqiad.wmnet * 11:32 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad * 11:32 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:29 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 11:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1097.eqiad.wmnet with reason: host reimage * 11:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2095.codfw.wmnet with OS trixie * 11:26 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1281.eqiad.wmnet * 11:26 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1280.eqiad.wmnet * 11:25 jayme@cumin1003: START - Cookbook sre.dns.netbox * 11:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1097.eqiad.wmnet with reason: host reimage * 11:22 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1141 * 11:21 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1141.eqiad.wmnet with OS trixie * 11:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1280.eqiad.wmnet * 11:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1279.eqiad.wmnet * 11:20 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin with reason: upgrade new Nokia swtiches in eqsin to SR Linux v26 * 11:17 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1141.eqiad.wmnet * 11:16 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1141.eqiad.wmnet * 11:16 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1141.eqiad.wmnet * 11:16 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be2095.codfw.wmnet with OS trixie * 11:15 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1279.eqiad.wmnet * 11:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1278.eqiad.wmnet * 11:14 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1139.eqiad.wmnet * 11:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1139.eqiad.wmnet * 11:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1070.eqiad.wmnet with OS trixie * 11:13 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1140.eqiad.wmnet * 11:13 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1140.eqiad.wmnet * 11:13 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1140.eqiad.wmnet * 11:09 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1278.eqiad.wmnet * 11:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1071.eqiad.wmnet with OS trixie * 11:05 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1097.eqiad.wmnet with OS trixie * 11:04 mvernon@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be1097.eqiad.wmnet with OS trixie * 11:02 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1140.eqiad.wmnet with OS trixie * 11:02 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1097.eqiad.wmnet with OS trixie * 11:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1096.eqiad.wmnet with OS trixie * 11:00 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1139.eqiad.wmnet * 11:00 jayme@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1139.eqiad.wmnet with OS trixie * 10:56 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1069.eqiad.wmnet with OS trixie * 10:56 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 10:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1070.eqiad.wmnet with reason: host reimage * 10:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 10:45 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1071.eqiad.wmnet with reason: host reimage * 10:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 10:41 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1096.eqiad.wmnet with OS trixie * 10:41 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1140.eqiad.wmnet with reason: host reimage * 10:39 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1071.eqiad.wmnet with reason: host reimage * 10:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1070.eqiad.wmnet with reason: host reimage * 10:38 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1139.eqiad.wmnet with reason: host reimage * 10:37 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1140.eqiad.wmnet with reason: host reimage * 10:35 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1069.eqiad.wmnet with reason: host reimage * 10:33 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1095.eqiad.wmnet with OS trixie * 10:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 10:33 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1139.eqiad.wmnet with reason: host reimage * 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1069.eqiad.wmnet with reason: host reimage * 10:23 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1140 * 10:23 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1140 * 10:23 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 10:22 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1140 * 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1140.eqiad.wmnet 155.48.64.10.in-addr.arpa 5.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:21 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1071.eqiad.wmnet with OS trixie * 10:21 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1140.eqiad.wmnet 155.48.64.10.in-addr.arpa 5.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1140 - jayme@cumin1003" * 10:21 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1140 - jayme@cumin1003" * 10:21 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1071 * 10:21 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1070.eqiad.wmnet with OS trixie * 10:21 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1070 * 10:20 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 10:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1095.eqiad.wmnet with OS trixie * 10:17 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1139 * 10:17 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1139 * 10:17 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be1095.eqiad.wmnet with OS trixie * 10:15 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1139 * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1139.eqiad.wmnet 194.32.64.10.in-addr.arpa 4.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:15 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1139.eqiad.wmnet 194.32.64.10.in-addr.arpa 4.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1139 - jayme@cumin1003" * 10:15 jayme@cumin1003: START - Cookbook sre.dns.netbox * 10:15 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1139 - jayme@cumin1003" * 10:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2095.codfw.wmnet with OS trixie * 10:13 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1069.eqiad.wmnet with OS trixie * 10:12 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1069 * 10:11 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1140 * 10:11 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1140.eqiad.wmnet with OS trixie * 10:11 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1140.eqiad.wmnet * 10:10 jayme@cumin1003: START - Cookbook sre.dns.netbox * 10:10 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1139 * 10:10 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1140.eqiad.wmnet * 10:10 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1140.eqiad.wmnet * 10:10 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1139.eqiad.wmnet with OS trixie * 10:09 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1139.eqiad.wmnet * 10:08 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1139.eqiad.wmnet * 10:08 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1139.eqiad.wmnet * 10:01 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2094.codfw.wmnet with OS trixie * 09:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 09:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 09:53 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 09:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 09:44 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:44 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2094.codfw.wmnet with reason: host reimage * 09:34 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2094.codfw.wmnet with reason: host reimage * 09:34 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1071 * 09:33 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1070 * 09:33 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1095.eqiad.wmnet with OS trixie * 09:27 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1069 * 09:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:23 brouberol@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM archiva1002.wikimedia.org * 09:20 brouberol@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM archiva1002.wikimedia.org * 09:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1277.eqiad.wmnet * 09:13 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2094.codfw.wmnet with OS trixie * 09:13 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 09:12 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 09:12 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:12 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1277.eqiad.wmnet * 09:12 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1276.eqiad.wmnet * 09:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1094.eqiad.wmnet with OS trixie * 09:06 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1276.eqiad.wmnet * 09:06 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1275.eqiad.wmnet * 09:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2093.codfw.wmnet with OS trixie * 09:01 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1275.eqiad.wmnet * 09:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1274.eqiad.wmnet * 08:56 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1274.eqiad.wmnet * 08:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1273.eqiad.wmnet * 08:50 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1273.eqiad.wmnet * 08:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1272.eqiad.wmnet * 08:49 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:49 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1094.eqiad.wmnet with reason: host reimage * 08:45 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1272.eqiad.wmnet * 08:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1094.eqiad.wmnet with reason: host reimage * 08:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2093.codfw.wmnet with reason: host reimage * 08:38 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:38 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2093.codfw.wmnet with reason: host reimage * 08:35 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 08:34 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 08:29 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:28 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:26 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1271.eqiad.wmnet * 08:23 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1094.eqiad.wmnet with OS trixie * 08:21 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 08:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1271.eqiad.wmnet * 08:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1270.eqiad.wmnet * 08:15 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1270.eqiad.wmnet * 08:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1269.eqiad.wmnet * 08:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2093.codfw.wmnet with OS trixie * 08:09 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1269.eqiad.wmnet * 08:09 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1268.eqiad.wmnet * 08:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2092.codfw.wmnet with OS trixie * 08:04 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1268.eqiad.wmnet * 08:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1267.eqiad.wmnet * 07:59 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1267.eqiad.wmnet * 07:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1093.eqiad.wmnet with OS trixie * 07:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1266.eqiad.wmnet * 07:56 jynus: running extra backups to test db1285 [[phab:T433826|T433826]] * 07:51 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1266.eqiad.wmnet * 07:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2092.codfw.wmnet with reason: host reimage * 07:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1093.eqiad.wmnet with reason: host reimage * 07:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2092.codfw.wmnet with reason: host reimage * 07:32 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1093.eqiad.wmnet with reason: host reimage * 07:29 jynus: running extra backups to test db1265 [[phab:T433825|T433825]] * 07:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2092.codfw.wmnet with OS trixie * 07:11 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1093.eqiad.wmnet with OS trixie * 06:50 slyngshede@dns1004: END - running authdns-update * 06:48 slyngshede@dns1004: START - running authdns-update * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.11 (duration: 02m 29s) * 03:38 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] (duration: 32m 57s) * 03:23 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 03:22 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 03:05 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 32s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 00:45 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] (duration: 06m 20s) * 00:41 cjming@deploy1003: cjming: Continuing with deployment * 00:41 cjming@deploy1003: cjming: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:39 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] == 2026-08-03 == * 23:58 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply * 23:57 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply * 23:29 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cp5021.eqsin.wmnet * 23:29 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cp5021.eqsin.wmnet * 23:27 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cp5021.eqsin.wmnet * 23:26 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cp5021.eqsin.wmnet * 23:18 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 23:17 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 22:56 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: sync * 22:56 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: sync * 22:36 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 22:36 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 22:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host search-loader2002.codfw.wmnet with OS trixie * 21:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on search-loader2002.codfw.wmnet with reason: host reimage * 21:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on search-loader2002.codfw.wmnet with reason: host reimage * 21:42 dancy@deploy1003: Stopping before sync operations * 21:41 dancy@deploy1003: Started scap sync-world: testing * 21:39 dancy@deploy1003: Installation of scap version "4.277.0" completed for 3 hosts * 21:37 dancy@deploy1003: Installing scap version "4.277.0" for 3 host(s) * 21:37 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] (duration: 06m 13s) * 21:33 dancy@deploy1003: dancy: Continuing with deployment * 21:32 dancy@deploy1003: dancy: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:31 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] * 21:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host search-loader2002.codfw.wmnet with OS trixie * 21:03 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] (duration: 06m 34s) * 20:59 dancy@deploy1003: dancy: Continuing with deployment * 20:58 dancy@deploy1003: dancy: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:56 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] * 20:52 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] (duration: 06m 23s) * 20:48 cjming@deploy1003: cjming: Continuing with deployment * 20:47 cjming@deploy1003: cjming: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:46 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] * 20:42 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] (duration: 07m 36s) * 20:38 arlolra@deploy1003: arlolra: Continuing with deployment * 20:36 arlolra@deploy1003: arlolra: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:34 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] * 20:16 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] (duration: 08m 26s) * 20:12 krinkle@deploy1003: krinkle: Continuing with deployment * 20:09 krinkle@deploy1003: krinkle: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] * 19:45 jasmine@cumin2002: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-main-codfw * 18:58 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] (duration: 09m 23s) * 18:53 krinkle@deploy1003: krinkle: Continuing with deployment * 18:53 jasmine@cumin2002: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-main-codfw * 18:50 krinkle@deploy1003: krinkle: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:48 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] * 18:37 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] (duration: 10m 13s) * 18:34 dzahn@cumin2002: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host codesearch2001.codfw.wmnet * 18:34 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host codesearch2001.codfw.wmnet with OS trixie * 18:33 krinkle@deploy1003: krinkle: Continuing with deployment * 18:29 krinkle@deploy1003: krinkle: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:27 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] * 18:19 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 18:18 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on codesearch2001.codfw.wmnet with reason: host reimage * 18:14 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 18:14 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:12 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on codesearch2001.codfw.wmnet with reason: host reimage * 18:11 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 18:11 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 18:10 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 18:02 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 18:02 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 18:01 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 18:01 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 17:55 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host codesearch2001.codfw.wmnet with OS trixie * 17:54 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:54 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) codesearch2001.codfw.wmnet on all recursors * 17:53 dzahn@cumin2002: START - Cookbook sre.dns.wipe-cache codesearch2001.codfw.wmnet on all recursors * 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:48 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:41 dzahn@cumin2002: START - Cookbook sre.dns.netbox * 17:41 dzahn@cumin2002: START - Cookbook sre.ganeti.makevm for new host codesearch2001.codfw.wmnet * 17:37 dzahn@cumin2002: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host codesearch1001.eqiad.wmnet * 17:37 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host codesearch1001.eqiad.wmnet with OS trixie * 17:24 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on codesearch1001.eqiad.wmnet with reason: host reimage * 17:17 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on codesearch1001.eqiad.wmnet with reason: host reimage * 17:08 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host codesearch1001.eqiad.wmnet with OS trixie * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 17:06 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:06 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) codesearch1001.eqiad.wmnet on all recursors * 17:06 dzahn@cumin2002: START - Cookbook sre.dns.wipe-cache codesearch1001.eqiad.wmnet on all recursors * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 17:05 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 17:04 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:04 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 16:58 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 16:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2091.codfw.wmnet with OS trixie * 16:54 ebernhardson@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 16:54 ebernhardson@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 16:49 ebernhardson@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 16:49 ebernhardson@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 16:46 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1092.eqiad.wmnet with OS trixie * 16:43 ebernhardson@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 16:43 ebernhardson@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 16:43 dzahn@cumin2002: START - Cookbook sre.dns.netbox * 16:43 dzahn@cumin2002: START - Cookbook sre.ganeti.makevm for new host codesearch1001.eqiad.wmnet * 16:41 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 16:41 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 16:40 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2091.codfw.wmnet with reason: host reimage * 16:37 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 16:35 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2091.codfw.wmnet with reason: host reimage * 16:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1092.eqiad.wmnet with reason: host reimage * 16:24 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 16:24 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 16:23 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1092.eqiad.wmnet with reason: host reimage * 16:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2091.codfw.wmnet with OS trixie * 16:03 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1092.eqiad.wmnet with OS trixie * 16:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2090.codfw.wmnet with OS trixie * 15:51 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 15:51 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 15:51 jiji@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 15:50 jiji@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 15:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2090.codfw.wmnet with reason: host reimage * 15:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2090.codfw.wmnet with reason: host reimage * 15:33 jhathaway@dns1004: END - running authdns-update * 15:31 jhathaway@dns1004: START - running authdns-update * 15:26 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1091.eqiad.wmnet with OS trixie * 15:25 dancy@deploy1003: Installation of scap version "4.276.1" completed for 3 hosts * 15:23 dancy@deploy1003: Installing scap version "4.276.1" for 3 host(s) * 15:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2090.codfw.wmnet with OS trixie * 15:12 marostegui@cumin1003: dbctl commit (dc=all): 'Repool db2245, db2246, db2247 and db2248 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95857 and previous config saved to /var/cache/conftool/dbconfig/20260803-151212-marostegui.json * 15:09 dancy@deploy1003: Started scap sync-world: testing * 15:09 dancy@deploy1003: Installation of scap version "4.277.0" completed for 3 hosts * 15:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1091.eqiad.wmnet with reason: host reimage * 15:07 dancy@deploy1003: Installing scap version "4.277.0" for 3 host(s) * 15:03 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1091.eqiad.wmnet with reason: host reimage * 14:49 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1091.eqiad.wmnet with OS trixie * 14:33 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2089.codfw.wmnet with OS trixie * 14:29 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 14:27 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 14:18 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 14:16 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 14:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2089.codfw.wmnet with reason: host reimage * 14:10 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 14:10 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 14:09 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2089.codfw.wmnet with reason: host reimage * 13:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2089.codfw.wmnet with OS trixie * 13:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2088.codfw.wmnet with OS trixie * 13:40 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1090.eqiad.wmnet with OS trixie * 13:22 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1090.eqiad.wmnet with reason: host reimage * 13:22 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] (duration: 14m 34s) * 13:19 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1090.eqiad.wmnet with reason: host reimage * 13:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2088.codfw.wmnet with reason: host reimage * 13:16 aude@deploy1003: aude, mhorsey: Continuing with deployment * 13:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2088.codfw.wmnet with reason: host reimage * 13:12 aude@deploy1003: aude, mhorsey: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] * 13:05 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1090.eqiad.wmnet with OS trixie * 12:58 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2088.codfw.wmnet with OS trixie * 12:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db[2245-2247].codfw.wmnet * 12:49 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2247: Rebooting db2247.codfw.wmnet * 12:49 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2247: Rebooting db2247.codfw.wmnet * 12:42 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2246: Rebooting db2246.codfw.wmnet * 12:42 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2246: Rebooting db2246.codfw.wmnet * 12:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2087.codfw.wmnet with OS trixie * 12:37 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1089.eqiad.wmnet with OS trixie * 12:34 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2245: Rebooting db2245.codfw.wmnet * 12:34 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2245: Rebooting db2245.codfw.wmnet * 12:34 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db[2245-2247].codfw.wmnet * 12:32 kamila@deploy1003: Finished scap sync-world: rebuild after base image update (duration: 30m 26s) * 12:28 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 12:22 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2087.codfw.wmnet with reason: host reimage * 12:19 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1089.eqiad.wmnet with reason: host reimage * 12:14 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2087.codfw.wmnet with reason: host reimage * 12:14 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1089.eqiad.wmnet with reason: host reimage * 12:03 kamila@deploy1003: Started scap sync-world: rebuild after base image update * 12:00 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1089.eqiad.wmnet with OS trixie * 12:00 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2087.codfw.wmnet with OS trixie * 11:35 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db[2245-2248].codfw.wmnet * 11:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db[2245-2248].codfw.wmnet * 11:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2086.codfw.wmnet with OS trixie * 11:26 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 11:26 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 11:25 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 11:25 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 11:24 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1088.eqiad.wmnet with OS trixie * 11:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db[2245-2248].codfw.wmnet with reason: Checking network * 11:21 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 11:20 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 11:19 marostegui@dns1004: END - running authdns-update * 11:17 marostegui@dns1004: START - running authdns-update * 11:10 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:10 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 11:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2086.codfw.wmnet with reason: host reimage * 11:09 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:08 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 11:08 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:07 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 11:07 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:07 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 11:06 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop: apply * 11:06 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop: apply * 11:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1088.eqiad.wmnet with reason: host reimage * 11:05 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop: apply * 11:04 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop: apply * 11:04 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop: apply * 11:04 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop: apply * 11:02 marostegui@dns1004: END - running authdns-update * 11:02 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2086.codfw.wmnet with reason: host reimage * 11:01 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1088.eqiad.wmnet with reason: host reimage * 11:00 marostegui@dns1004: START - running authdns-update * 10:53 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] (duration: 10m 57s) * 10:51 cmooney@dns3003: END - running authdns-update * 10:49 cmooney@dns3003: START - running authdns-update * 10:47 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1088.eqiad.wmnet with OS trixie * 10:47 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2086.codfw.wmnet with OS trixie * 10:47 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 10:46 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:46 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:46 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new reverse ranges for eqsin CR switch links - cmooney@cumin1003" * 10:46 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new reverse ranges for eqsin CR switch links - cmooney@cumin1003" * 10:42 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] * 10:41 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 10:36 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2245, db2246 and db2247 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95855 and previous config saved to /var/cache/conftool/dbconfig/20260803-103652-marostegui.json * 10:35 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2248 from s4 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95854 and previous config saved to /var/cache/conftool/dbconfig/20260803-103535-marostegui.json * 10:27 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 10:27 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 10:26 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 10:24 kart_: cxserver: Add referencePunctuation config ([[phab:T97231|T97231]]) * 10:24 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 10:23 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:23 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:23 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:22 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:22 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply * 10:21 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply * 10:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2085.codfw.wmnet with OS trixie * 10:20 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply * 10:20 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply * 10:18 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply * 10:18 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply * 10:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1087.eqiad.wmnet with OS trixie * 09:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2085.codfw.wmnet with reason: host reimage * 09:43 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1087.eqiad.wmnet with reason: host reimage * 09:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2085.codfw.wmnet with reason: host reimage * 09:40 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1087.eqiad.wmnet with reason: host reimage * 09:26 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1087.eqiad.wmnet with OS trixie * 09:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2085.codfw.wmnet with OS trixie * 09:13 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2084.codfw.wmnet with OS trixie * 09:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1086.eqiad.wmnet with OS trixie * 08:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2084.codfw.wmnet with reason: host reimage * 08:50 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2084.codfw.wmnet with reason: host reimage * 08:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1086.eqiad.wmnet with reason: host reimage * 08:39 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1086.eqiad.wmnet with reason: host reimage * 08:38 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:38 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:37 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 08:37 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 08:35 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2084.codfw.wmnet with OS trixie * 08:34 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 08:34 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:27 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1086.eqiad.wmnet with OS trixie * 08:09 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1218: Repool after a crash * 08:07 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2083.codfw.wmnet with OS trixie * 08:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1085.eqiad.wmnet with OS trixie * 07:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2083.codfw.wmnet with reason: host reimage * 07:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1085.eqiad.wmnet with reason: host reimage * 07:40 kart_: Updated cxsever to 2026-07-16-140518-production ([[phab:T97231|T97231]]) * 07:39 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply * 07:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2083.codfw.wmnet with reason: host reimage * 07:38 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply * 07:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1085.eqiad.wmnet with reason: host reimage * 07:37 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] (duration: 32m 40s) * 07:33 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply * 07:33 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply * 07:25 jdlrobson@deploy1003: jdlrobson: Continuing with deployment * 07:24 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2083.codfw.wmnet with OS trixie * 07:24 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1085.eqiad.wmnet with OS trixie * 07:23 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1218: Repool after a crash * 07:21 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:09 marostegui: Drop renamed tables [[phab:T425074|T425074]] * 07:04 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] * 06:55 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply * 06:54 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 46s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-02 == * 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 01m 03s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-01 == * 03:30 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:30 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:30 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:30 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 34s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-31 == * 17:41 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 17:41 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 17:40 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 17:40 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 15:33 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2195: Testing * 15:02 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:02 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 15:02 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 14:48 pt1979@cumin2002: START - Cookbook sre.dns.netbox * 14:47 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2195: Testing * 14:22 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 14:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2195: Testing * 14:21 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 14:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2195.codfw.wmnet with reason: Testing * 14:16 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 14:04 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 14:04 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 13:30 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1048.eqiad.wmnet with OS trixie * 13:22 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2195: Testing * 13:22 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 13:19 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2195: Testing * 13:18 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 13:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2195: Testing * 13:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 13:05 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 13:05 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 13:04 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 13:04 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 12:50 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lswtest-d8-eqiad * 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:53 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:42 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 6515 * 11:37 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 6515 * 11:28 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:27 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:07 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2082.codfw.wmnet with OS trixie * 10:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2082.codfw.wmnet with reason: host reimage * 10:42 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2082.codfw.wmnet with reason: host reimage * 10:28 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2082.codfw.wmnet with OS trixie * 10:02 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 09:52 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 09:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts * 09:16 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts * 08:57 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 08:46 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:42 gkyziridis@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 08:37 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 08:37 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 08:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 08:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 08:11 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:11 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:08 filippo@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudvirt1048 * 08:07 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 08:07 filippo@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudvirt1048 * 08:06 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 08:01 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 08:00 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 07:19 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1048.eqiad.wmnet with reason: host reimage * 07:13 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1048.eqiad.wmnet with reason: host reimage * 07:11 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 07:11 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 07:09 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 07:09 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 06:57 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:56 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:48 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:48 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:44 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie * 06:34 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1048.eqiad.wmnet with OS trixie * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 54s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 00:57 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] (duration: 11m 04s) * 00:53 dreamyjazz@deploy1003: dreamyjazz, jforrester: Continuing with deployment * 00:48 dreamyjazz@deploy1003: dreamyjazz, jforrester: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:46 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] == 2026-07-30 == * 21:37 dancy@deploy1003: Installation of scap version "4.276.1" completed for 3 hosts * 21:35 dancy@deploy1003: Installing scap version "4.276.1" for 3 host(s) * 21:24 dancy@deploy1003: Installation of scap version "4.276.0" completed for 3 hosts * 21:22 dancy@deploy1003: Installing scap version "4.276.0" for 3 host(s) * 21:15 maryum: Deployed security fix for [[phab:T430601|T430601]] * 20:13 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] (duration: 09m 20s) * 20:07 arlolra@deploy1003: osleger, arlolra: Continuing with deployment * 20:05 arlolra@deploy1003: osleger, arlolra: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:03 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] * 19:29 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:29 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:25 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service * 19:24 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 19:24 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:24 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:24 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 19:23 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service * 19:20 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1084.eqiad.wmnet with OS trixie * 18:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1084.eqiad.wmnet with reason: host reimage * 18:52 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1084.eqiad.wmnet with reason: host reimage * 18:41 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 18:40 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 18:39 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1084.eqiad.wmnet with OS trixie * 18:25 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 18:15 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 17:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1083.eqiad.wmnet with OS trixie * 17:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1048: Maintenance * 17:36 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new security plugin settings - bking@cumin2003 - [[phab:T350516|T350516]] * 17:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1083.eqiad.wmnet with reason: host reimage * 17:26 inflatador: bking@apt1002 `reprepro --noskipold --component thirdparty/opensearch3 update trixie-wikimedia` [[phab:T433624|T433624]] * 17:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1083.eqiad.wmnet with reason: host reimage * 17:23 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 17:20 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 17:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2081.codfw.wmnet with OS trixie * 17:11 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new security plugin settings - bking@cumin2003 - [[phab:T350516|T350516]] * 17:10 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1083.eqiad.wmnet with OS trixie * 16:55 root@cumin1003: START - Cookbook sre.mysql.pool pool es1048: Maintenance * 16:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2081.codfw.wmnet with reason: host reimage * 16:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1048 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95833 and previous config saved to /var/cache/conftool/dbconfig/20260730-165053-cwilliams.json * 16:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1048.eqiad.wmnet with reason: Maintenance * 16:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1040: Maintenance * 16:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2081.codfw.wmnet with reason: host reimage * 16:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1082.eqiad.wmnet with OS trixie * 16:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2081.codfw.wmnet with OS trixie * 16:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2097.codfw.wmnet with OS trixie * 16:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1082.eqiad.wmnet with reason: host reimage * 16:08 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1082.eqiad.wmnet with reason: host reimage * 16:08 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 16:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1047: Maintenance * 16:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2080.codfw.wmnet with OS trixie * 16:04 root@cumin1003: START - Cookbook sre.mysql.pool pool es1040: Maintenance * 16:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1040: Maintenance * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new logging settings - bking@cumin2003 - [[phab:T324335|T324335]] * 15:58 root@cumin1003: START - Cookbook sre.mysql.pool pool es1040: Maintenance * 15:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1040 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95827 and previous config saved to /var/cache/conftool/dbconfig/20260730-155324-cwilliams.json * 15:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1040.eqiad.wmnet with reason: Maintenance * 15:50 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1082.eqiad.wmnet with OS trixie * 15:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2048: Maintenance * 15:44 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 24s) * 15:43 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2080.codfw.wmnet with reason: host reimage * 15:38 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new logging settings - bking@cumin2003 - [[phab:T324335|T324335]] * 15:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2080.codfw.wmnet with reason: host reimage * 15:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 15:30 mvernon@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be2097.codfw.wmnet with OS trixie * 15:23 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1081.eqiad.wmnet with OS trixie * 15:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2098.codfw.wmnet with OS trixie * 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - mvernon@cumin2003" * 15:18 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be2097.codfw.wmnet with OS trixie * 15:18 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - mvernon@cumin2003" * 15:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 15:17 root@cumin1003: START - Cookbook sre.mysql.pool pool es1047: Maintenance * 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2080.codfw.wmnet with OS trixie * 15:13 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be2097.codfw.wmnet with OS trixie * 15:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1047 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95820 and previous config saved to /var/cache/conftool/dbconfig/20260730-151200-cwilliams.json * 15:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1047.eqiad.wmnet with reason: Maintenance * 15:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: Maintenance * 15:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 15:04 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 15:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1081.eqiad.wmnet with reason: host reimage * 15:00 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1081.eqiad.wmnet with reason: host reimage * 15:00 root@cumin1003: START - Cookbook sre.mysql.pool pool es2048: Maintenance * 15:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 14:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2079.codfw.wmnet with OS trixie * 14:56 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 14:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2048 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95816 and previous config saved to /var/cache/conftool/dbconfig/20260730-145510-cwilliams.json * 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2048.codfw.wmnet with reason: Maintenance * 14:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2040: Maintenance * 14:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 14:51 tchin@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] (duration: 06m 48s) * 14:47 tchin@deploy1003: jforrester, tchin: Continuing with deployment * 14:47 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 14:47 tchin@deploy1003: jforrester, tchin: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:45 tchin@deploy1003: Started scap sync-world: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] * 14:42 sukhe@puppetserver1001: conftool action : set/weight=1; selector: cluster=urldownloader,service=squid * 14:42 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader,service=squid * 14:42 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1081.eqiad.wmnet with OS trixie * 14:39 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 14:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2079.codfw.wmnet with reason: host reimage * 14:36 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2098.codfw.wmnet with OS trixie * 14:32 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2079.codfw.wmnet with reason: host reimage * 14:30 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] (duration: 06m 31s) * 14:27 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 14:26 mszwarc@deploy1003: mszwarc: Continuing with deployment * 14:25 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:25 root@cumin1003: START - Cookbook sre.mysql.pool pool es1038: Maintenance * 14:25 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1038: Maintenance * 14:23 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] * 14:21 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] (duration: 11m 19s) * 14:20 root@cumin1003: START - Cookbook sre.mysql.pool pool es1038: Maintenance * 14:14 stran@deploy1003: stran: Continuing with deployment * 14:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1038 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95810 and previous config saved to /var/cache/conftool/dbconfig/20260730-141439-cwilliams.json * 14:14 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1038.eqiad.wmnet with reason: Maintenance * 14:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1036: Maintenance * 14:13 stran@deploy1003: stran: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2079.codfw.wmnet with OS trixie * 14:09 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] * 14:08 root@cumin1003: START - Cookbook sre.mysql.pool pool es2040: Maintenance * 14:08 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2040: Maintenance * 14:03 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] (duration: 31m 41s) * 14:03 root@cumin1003: START - Cookbook sre.mysql.pool pool es2040: Maintenance * 14:03 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2040 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95806 and previous config saved to /var/cache/conftool/dbconfig/20260730-135643-cwilliams.json * 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2040.codfw.wmnet with reason: Maintenance * 13:56 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:56 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2038: Maintenance * 13:55 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:52 lucaswerkmeister-wmde@deploy1003: migr, lucaswerkmeister-wmde: Continuing with deployment * 13:49 lucaswerkmeister-wmde@deploy1003: migr, lucaswerkmeister-wmde: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:49 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:48 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 13:45 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 13:32 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:32 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] * 13:28 root@cumin1003: START - Cookbook sre.mysql.pool pool es1036: Maintenance * 13:28 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1036: Maintenance * 13:22 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:22 root@cumin1003: START - Cookbook sre.mysql.pool pool es1036: Maintenance * 13:20 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1036 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95800 and previous config saved to /var/cache/conftool/dbconfig/20260730-131727-cwilliams.json * 13:17 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1036.eqiad.wmnet with reason: Maintenance * 13:17 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] (duration: 10m 31s) * 13:16 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2022\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 13:13 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, stran: Continuing with deployment * 13:10 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:10 root@cumin1003: START - Cookbook sre.mysql.pool pool es2038: Maintenance * 13:10 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2038: Maintenance * 13:08 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, stran: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie * 13:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2047: Maintenance * 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048 cloud-private - filippo@cumin1003" * 13:07 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048 cloud-private - filippo@cumin1003" * 13:06 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] * 13:04 root@cumin1003: START - Cookbook sre.mysql.pool pool es2038: Maintenance * 13:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2078.codfw.wmnet with OS trixie * 13:01 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2038 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95797 and previous config saved to /var/cache/conftool/dbconfig/20260730-125919-cwilliams.json * 12:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2038.codfw.wmnet with reason: Maintenance * 12:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2078.codfw.wmnet with reason: host reimage * 12:37 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2078.codfw.wmnet with reason: host reimage * 12:37 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] (duration: 06m 51s) * 12:33 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 12:32 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 12:32 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 12:32 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:30 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] * 12:19 root@cumin1003: START - Cookbook sre.mysql.pool pool es2047: Maintenance * 12:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2078.codfw.wmnet with OS trixie * 12:18 dcausse@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 12:18 dcausse@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 12:15 dcausse@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 12:14 dcausse@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 12:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2047 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95793 and previous config saved to /var/cache/conftool/dbconfig/20260730-121404-cwilliams.json * 12:13 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2047.codfw.wmnet with reason: Maintenance * 12:13 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2036: Maintenance * 12:05 ayounsi@dns1004: END - running authdns-update * 12:02 ayounsi@dns1004: START - running authdns-update * 11:51 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2077.codfw.wmnet with OS trixie * 11:48 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:46 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1080.eqiad.wmnet with OS trixie * 11:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1226: Maintenance * 11:41 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2077.codfw.wmnet with reason: host reimage * 11:28 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2077.codfw.wmnet with reason: host reimage * 11:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1080.eqiad.wmnet with reason: host reimage * 11:27 root@cumin1003: START - Cookbook sre.mysql.pool pool es2036: Maintenance * 11:27 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2036: Maintenance * 11:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1080.eqiad.wmnet with reason: host reimage * 11:21 root@cumin1003: START - Cookbook sre.mysql.pool pool es2036: Maintenance * 11:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2036 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95786 and previous config saved to /var/cache/conftool/dbconfig/20260730-111633-cwilliams.json * 11:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2036.codfw.wmnet with reason: Maintenance * 11:08 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2077.codfw.wmnet with OS trixie * 11:07 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 11:03 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie * 11:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1226: Maintenance * 10:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1226 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95783 and previous config saved to /var/cache/conftool/dbconfig/20260730-104801-cwilliams.json * 10:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1226.eqiad.wmnet with reason: Maintenance * 10:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1214: Maintenance * 10:27 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2035: Maintenance * 10:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2076.codfw.wmnet with OS trixie * 10:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1214: Maintenance * 09:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1214 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95775 and previous config saved to /var/cache/conftool/dbconfig/20260730-095451-cwilliams.json * 09:54 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1214.eqiad.wmnet with reason: Maintenance * 09:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1209: Maintenance * 09:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2076.codfw.wmnet with reason: host reimage * 09:42 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool es2035: Maintenance * 09:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.netbox.update-extras (exit_code=0) rolling restart_daemons on A:netbox * 09:41 ayounsi@cumin1003: START - Cookbook sre.netbox.update-extras rolling restart_daemons on A:netbox * 09:40 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2035: Maintenance * 09:39 ayounsi@cumin1003: END (PASS) - Cookbook sre.netbox.update-extras (exit_code=0) rolling restart_daemons on A:netbox-canary * 09:39 ayounsi@cumin1003: START - Cookbook sre.netbox.update-extras rolling restart_daemons on A:netbox-canary * 09:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2076.codfw.wmnet with reason: host reimage * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 09:34 root@cumin1003: START - Cookbook sre.mysql.pool pool es2035: Maintenance * 09:32 lucaswerkmeister-wmde@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 09:32 lucaswerkmeister-wmde@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 09:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2035 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95771 and previous config saved to /var/cache/conftool/dbconfig/20260730-092910-cwilliams.json * 09:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2035.codfw.wmnet with reason: Maintenance * 09:19 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2076.codfw.wmnet with OS trixie * 09:18 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 09:17 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie * 09:07 root@cumin1003: START - Cookbook sre.mysql.pool pool db1209: Maintenance * 09:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 23 hosts * 09:04 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Remove cable label from interfaces descriptions - ayounsi@cumin1003 * 09:04 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:02 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Remove cable label from interfaces descriptions - ayounsi@cumin1003 * 09:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1209 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95767 and previous config saved to /var/cache/conftool/dbconfig/20260730-090133-cwilliams.json * 09:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1209.eqiad.wmnet with reason: Maintenance * 09:01 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1192: Maintenance * 08:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1252: Maintenance * 08:57 jayme@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 08:56 jayme@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 08:53 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 23 hosts * 08:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:51 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:50 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1263: Maintenance * 08:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:23 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 08:15 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie * 08:14 root@cumin1003: START - Cookbook sre.mysql.pool pool db1192: Maintenance * 08:13 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2075.codfw.wmnet with OS trixie * 08:12 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1252: Maintenance * 08:11 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1252: Maintenance * 08:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1252: Maintenance * 08:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1192 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95754 and previous config saved to /var/cache/conftool/dbconfig/20260730-080611-cwilliams.json * 08:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1192.eqiad.wmnet with reason: Maintenance * 08:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1178: Maintenance * 08:05 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS bullseye * 07:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1263: Maintenance * 07:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1263 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95751 and previous config saved to /var/cache/conftool/dbconfig/20260730-075106-cwilliams.json * 07:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[1260-1262].eqiad.wmnet with reason: Maintenance * 07:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2075.codfw.wmnet with reason: host reimage * 07:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1263.eqiad.wmnet with reason: Maintenance * 07:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2075.codfw.wmnet with reason: host reimage * 07:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance * 07:38 dcausse@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:38 dcausse@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 07:35 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db1252', diff saved to https://phabricator.wikimedia.org/P95748 and previous config saved to /var/cache/conftool/dbconfig/20260730-073510-marostegui.json * 07:26 klausman@dns2004: END - running authdns-update * 07:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 07:25 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2075.codfw.wmnet with OS trixie * 07:24 klausman@dns2004: START - running authdns-update * 07:23 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host an-test-master1003.eqiad.wmnet * 07:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db1178: Maintenance * 07:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1178 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95746 and previous config saved to /var/cache/conftool/dbconfig/20260730-071112-cwilliams.json * 07:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1178.eqiad.wmnet with reason: Maintenance * 07:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1177: Maintenance * 06:24 root@cumin1003: START - Cookbook sre.mysql.pool pool db1177: Maintenance * 06:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1177 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95741 and previous config saved to /var/cache/conftool/dbconfig/20260730-061736-cwilliams.json * 06:17 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1177.eqiad.wmnet with reason: Maintenance * 06:17 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1172: Maintenance * 05:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1218.eqiad.wmnet with reason: crashed * 05:41 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db1217 it crashed', diff saved to https://phabricator.wikimedia.org/P95737 and previous config saved to /var/cache/conftool/dbconfig/20260730-054111-marostegui.json * 05:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95736 and previous config saved to /var/cache/conftool/dbconfig/20260730-053422-cwilliams.json * 05:30 root@cumin1003: START - Cookbook sre.mysql.pool pool db1172: Maintenance * 05:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95734 and previous config saved to /var/cache/conftool/dbconfig/20260730-052414-cwilliams.json * 05:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1172 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95733 and previous config saved to /var/cache/conftool/dbconfig/20260730-052354-cwilliams.json * 05:23 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1172.eqiad.wmnet with reason: Maintenance * 05:23 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1167: Maintenance * 05:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95731 and previous config saved to /var/cache/conftool/dbconfig/20260730-051406-cwilliams.json * 05:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95729 and previous config saved to /var/cache/conftool/dbconfig/20260730-050358-cwilliams.json * 04:47 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 04:47 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 04:47 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 04:35 root@cumin1003: START - Cookbook sre.mysql.pool pool db1167: Maintenance * 04:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1167 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95726 and previous config saved to /var/cache/conftool/dbconfig/20260730-042923-cwilliams.json * 04:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 04:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1167.eqiad.wmnet with reason: Maintenance * 04:22 pt1979@cumin2002: START - Cookbook sre.dns.netbox * 04:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95725 and previous config saved to /var/cache/conftool/dbconfig/20260730-040337-cwilliams.json * 04:03 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:38 brett@cumin2002: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool eqsin [reason: Switch upgrade maintenance window complete, [[phab:T433097|T433097]]] * 01:38 brett@cumin2002: START - Cookbook sre.dns.admin DNS admin: pool eqsin [reason: Switch upgrade maintenance window complete, [[phab:T433097|T433097]]] == 2026-07-29 == * 23:57 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin,mr1-eqsin IPv6,mr1-eqsin.oob,mr1-eqsin.oob IPv6 with reason: connection issue * 22:54 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 22:53 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 22:53 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 22:53 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:25 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2022.codfw.wmnet, repooling source-only afterwards * 22:20 brett@cumin2002: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool eqsin [reason: Switch upgrade maintenance window, [[phab:T433097|T433097]]] * 22:20 brett@cumin2002: START - Cookbook sre.dns.admin DNS admin: depool eqsin [reason: Switch upgrade maintenance window, [[phab:T433097|T433097]]] * 22:01 apine@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 22:00 apine@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 21:59 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 21:58 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 21:58 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 21:58 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 21:32 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1048.eqiad.wmnet with OS trixie * 21:25 pt1979@cumin2002: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be2097.codfw.wmnet with OS bullseye * 21:16 zabe@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki=metawiki 'Mental Health Resource Center' 'Safety Resource Center/Mental Health' Zabe --reason 'per request [[:phab:T433118{{!}}T433118]]' * 21:12 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2022.codfw.wmnet, repooling source-only afterwards * 21:12 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] (duration: 12m 53s) * 21:12 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2015\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 21:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1253: Maintenance * 21:08 aaron@deploy1003: aaron: Continuing with deployment * 21:01 aaron@deploy1003: aaron: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:59 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] * 20:52 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] (duration: 21m 57s) * 20:48 aaron@deploy1003: aaron: Continuing with deployment * 20:32 aaron@deploy1003: aaron: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:30 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] * 20:24 root@cumin1003: START - Cookbook sre.mysql.pool pool db1253: Maintenance * 20:19 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] (duration: 08m 07s) * 20:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1253 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95719 and previous config saved to /var/cache/conftool/dbconfig/20260729-201810-cwilliams.json * 20:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1253.eqiad.wmnet with reason: Maintenance * 20:17 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1231: Maintenance * 20:15 aaron@deploy1003: bpirkle, aaron: Continuing with deployment * 20:13 aaron@deploy1003: bpirkle, aaron: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:12 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie * 20:11 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] * 20:11 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1048.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:09 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1048.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:09 pt1979@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 20:04 pt1979@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 19:47 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 19:43 pt1979@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye * 19:41 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 19:37 zabe: zabe@deploy1003:~$ mwscript-k8s --comment='[[phab:T433529|T433529]]' --follow -- resetAuthenticationThrottle.php --wiki=aawiki --signup --ip=89.36.114.94 * 19:36 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] (duration: 06m 49s) * 19:32 zabe@deploy1003: zabe: Continuing with deployment * 19:31 zabe@deploy1003: zabe: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:31 root@cumin1003: START - Cookbook sre.mysql.pool pool db1231: Maintenance * 19:29 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] * 19:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1231 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95714 and previous config saved to /var/cache/conftool/dbconfig/20260729-192454-cwilliams.json * 19:24 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1231.eqiad.wmnet with reason: Maintenance * 19:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1227: Maintenance * 19:22 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 19:22 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 19:21 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:21 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1048] - vriley@cumin1003" * 19:21 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1048] - vriley@cumin1003" * 19:19 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 19:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1251: Maintenance * 19:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95711 and previous config saved to /var/cache/conftool/dbconfig/20260729-191756-cwilliams.json * 19:16 vriley@cumin1003: START - Cookbook sre.dns.netbox * 19:11 dduvall: rolling back wmf.13 to group0 due to [[phab:T433457|T433457]] (cc [[phab:T430832|T430832]]) * 19:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95709 and previous config saved to /var/cache/conftool/dbconfig/20260729-190748-cwilliams.json * 19:01 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2022.codfw.wmnet with OS bookworm * 19:01 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2021.codfw.wmnet, repooling source-only afterwards * 18:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95707 and previous config saved to /var/cache/conftool/dbconfig/20260729-185740-cwilliams.json * 18:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95704 and previous config saved to /var/cache/conftool/dbconfig/20260729-184732-cwilliams.json * 18:37 root@cumin1003: START - Cookbook sre.mysql.pool pool db1227: Maintenance * 18:34 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2022.codfw.wmnet with reason: host reimage * 18:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1227 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95701 and previous config saved to /var/cache/conftool/dbconfig/20260729-183117-cwilliams.json * 18:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1227.eqiad.wmnet with reason: Maintenance * 18:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1202: Maintenance * 18:30 root@cumin1003: START - Cookbook sre.mysql.pool pool db1251: Maintenance * 18:27 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2022.codfw.wmnet with reason: host reimage * 18:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1251 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95698 and previous config saved to /var/cache/conftool/dbconfig/20260729-182428-cwilliams.json * 18:24 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lvs2014.codfw.wmnet * 18:24 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for lvs2014.codfw.wmnet * 18:24 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1251.eqiad.wmnet with reason: Maintenance * 18:23 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1235: Maintenance * 18:22 brett@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 18:19 brett@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 18:19 brett@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 18:17 brett@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 18:17 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 18:16 mutante: removing jenkins during the train - living on the edge - no, just kidding, jenkins has migrated to dedicated machines, nothing should happen * 18:15 brett@cumin2002: END (ERROR) - Cookbook sre.loadbalancer.restart-pybal (exit_code=97) rolling-restart of pybal on P<nowiki>{</nowiki>lvs2014.codfw.wmnet<nowiki>}</nowiki> and A:lvs ([[phab:T428495|T428495]]) * 18:15 mutante: CI: contint1002/contint2002: apt-get remove --purge jenkins - jenkins be gone - [[phab:T418521|T418521]] * 18:13 brett@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on P<nowiki>{</nowiki>lvs2014.codfw.wmnet<nowiki>}</nowiki> and A:lvs ([[phab:T428495|T428495]]) * 18:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2022 * 18:08 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2022 * 18:03 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T428495|T428495]] * 18:03 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2022 * 18:02 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2022.codfw.wmnet 211.48.192.10.in-addr.arpa 1.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:02 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2022.codfw.wmnet 211.48.192.10.in-addr.arpa 1.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:02 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:02 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2022 - bking@cumin2003" * 18:02 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2022 - bking@cumin2003" * 17:57 bking@cumin2003: START - Cookbook sre.dns.netbox * 17:56 brett@cumin2002: END (FAIL) - Cookbook sre.loadbalancer.restart-pybal (exit_code=1) rolling-restart of pybal on A:lvs-codfw and A:lvs ([[phab:T428495|T428495]]) * 17:55 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo - [[phab:T428495|T428495]] * 17:54 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2022 * 17:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2022.codfw.wmnet with OS bookworm * 17:50 brett@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on A:lvs-codfw and A:lvs ([[phab:T428495|T428495]]) * 17:47 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2021.codfw.wmnet, repooling source-only afterwards * 17:47 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 14s) * 17:47 swfrench-wmf: authdns-update to direct codfw, eqsin, ulsfo etcd clients back to codfw - [[phab:T428495|T428495]] * 17:47 swfrench@dns1004: END - running authdns-update * 17:47 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 17:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95692 and previous config saved to /var/cache/conftool/dbconfig/20260729-174713-cwilliams.json * 17:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance * 17:46 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1249: Maintenance * 17:45 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2015\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 17:45 swfrench@dns1004: START - running authdns-update * 17:44 root@cumin1003: START - Cookbook sre.mysql.pool pool db1202: Maintenance * 17:41 akhatun: Deployed refinery using scap, then deployed onto hdfs * 17:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1202 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95688 and previous config saved to /var/cache/conftool/dbconfig/20260729-173759-cwilliams.json * 17:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1202.eqiad.wmnet with reason: Maintenance * 17:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1194: Maintenance * 17:37 root@cumin1003: START - Cookbook sre.mysql.pool pool db1235: Maintenance * 17:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1230: Maintenance * 17:30 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1235 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95684 and previous config saved to /var/cache/conftool/dbconfig/20260729-173051-cwilliams.json * 17:30 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1235.eqiad.wmnet with reason: Maintenance * 17:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1234: Maintenance * 17:26 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (thin): Regular analytics weekly train THIN [analytics/refinery@56695674] (duration: 02m 02s) * 17:24 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (thin): Regular analytics weekly train THIN [analytics/refinery@56695674] * 17:23 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567]: Regular analytics weekly train [analytics/refinery@56695674] (duration: 06m 20s) * 17:20 dancy@deploy1003: Finished scap sync-world: Testing delay_messageblobstore_purge: true (duration: 06m 29s) * 17:17 akhatun@deploy1003: Started deploy [analytics/refinery@5669567]: Regular analytics weekly train [analytics/refinery@56695674] * 17:17 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] (duration: 00m 22s) * 17:16 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] * 17:13 dancy@deploy1003: Started scap sync-world: Testing delay_messageblobstore_purge: true * 17:05 mutante: CI: contint1002/contint2002 - restarted httpd to be extra sure all is cleaned up - https://integration.wikimedia.org/ci/ is up and running [[phab:T418521|T418521]] * 17:04 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 17:03 mutante: CI: contint1002/contint2002 - rm /etc/apache2/jenkins_proxy - removing legacy jenkins proxy config - jenkins is on new dedicated machines and uses jenkins_proxy_ext config [[phab:T418521|T418521]] * 17:02 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] (duration: 36m 25s) * 17:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1249: Maintenance * 16:59 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 16:54 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2015.codfw.wmnet, repooling source-only afterwards * 16:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1249 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95674 and previous config saved to /var/cache/conftool/dbconfig/20260729-165339-cwilliams.json * 16:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1249.eqiad.wmnet with reason: Maintenance * 16:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1248: Maintenance * 16:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1194: Maintenance * 16:47 swfrench-wmf: silenced EtcdReplicationDown 57b2b421-1cc9-4e38-9276-{{Gerrit|94f223fd231c}} - [[phab:T428495|T428495]] * 16:46 tchin@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/eventstreams-internal: apply * 16:46 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye * 16:46 tchin@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/eventstreams-internal: apply * 16:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1230: Maintenance * 16:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1194 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95669 and previous config saved to /var/cache/conftool/dbconfig/20260729-164422-cwilliams.json * 16:44 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 16:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1194.eqiad.wmnet with reason: Maintenance * 16:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1191: Maintenance * 16:43 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host an-test-master1003.eqiad.wmnet * 16:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db1234: Maintenance * 16:43 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 16:43 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Rolling back deployment * 16:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host an-test-master1004.eqiad.wmnet * 16:41 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 16:40 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 16:40 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 16:39 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 16:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1230 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95667 and previous config saved to /var/cache/conftool/dbconfig/20260729-163932-cwilliams.json * 16:39 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 16:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1230.eqiad.wmnet with reason: Maintenance * 16:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1207: Maintenance * 16:38 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 16:37 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host an-test-master1004.eqiad.wmnet * 16:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1234 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95664 and previous config saved to /var/cache/conftool/dbconfig/20260729-163719-cwilliams.json * 16:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1234.eqiad.wmnet with reason: Maintenance * 16:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1079.eqiad.wmnet with OS trixie * 16:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1232: Maintenance * 16:34 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1259: Maintenance * 16:28 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] (duration: 06m 57s) * 16:28 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:26 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] * 16:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1051 hosts * 16:21 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] * 16:20 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2006.codfw.wmnet with OS bookworm * 16:19 akhatun: Deploying Refinery at {{Gerrit|56695674}} as part of weekly train * 16:18 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1079.eqiad.wmnet with reason: host reimage * 16:16 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] (duration: 15m 36s) * 16:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2021.codfw.wmnet with OS bookworm * 16:14 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1079.eqiad.wmnet with reason: host reimage * 16:12 topranks: hot-swap line card in FPC0 on cr1-eqiad with replacement MPC10E from Juniper [[phab:T426343|T426343]] * 16:10 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Continuing with deployment * 16:07 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db1248: Maintenance * 16:01 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] * 16:00 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 16:00 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 15:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1248 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95651 and previous config saved to /var/cache/conftool/dbconfig/20260729-155956-cwilliams.json * 15:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1248.eqiad.wmnet with reason: Maintenance * 15:59 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2006.codfw.wmnet with reason: host reimage * 15:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1247: Maintenance * 15:59 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2074.codfw.wmnet with OS trixie * 15:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1191: Maintenance * 15:57 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:55 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1079.eqiad.wmnet with OS trixie * 15:55 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2006.codfw.wmnet with reason: host reimage * 15:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1207: Maintenance * 15:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1191 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95646 and previous config saved to /var/cache/conftool/dbconfig/20260729-155104-cwilliams.json * 15:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1191.eqiad.wmnet with reason: Maintenance * 15:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1181: Maintenance * 15:49 root@cumin1003: START - Cookbook sre.mysql.pool pool db1232: Maintenance * 15:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2021.codfw.wmnet with reason: host reimage * 15:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1207 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95643 and previous config saved to /var/cache/conftool/dbconfig/20260729-154735-cwilliams.json * 15:47 root@cumin1003: START - Cookbook sre.mysql.pool pool db1259: Maintenance * 15:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1207.eqiad.wmnet with reason: Maintenance * 15:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1200: Maintenance * 15:46 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] (duration: 31m 59s) * 15:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 15:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1232 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95640 and previous config saved to /var/cache/conftool/dbconfig/20260729-154330-cwilliams.json * 15:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1232.eqiad.wmnet with reason: Maintenance * 15:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1219: Maintenance * 15:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2015.codfw.wmnet, repooling source-only afterwards * 15:41 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 18s) * 15:41 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1259 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95638 and previous config saved to /var/cache/conftool/dbconfig/20260729-154107-cwilliams.json * 15:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1259.eqiad.wmnet with reason: Maintenance * 15:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1254: Maintenance * 15:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2015.codfw.wmnet with OS bookworm * 15:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2021.codfw.wmnet with reason: host reimage * 15:36 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2006.codfw.wmnet with OS bookworm * 15:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 15:35 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Continuing with deployment * 15:33 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2074.codfw.wmnet with OS trixie * 15:33 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 15:32 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:29 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:28 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be2074.codfw.wmnet with OS trixie * 15:28 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2006.codfw.wmnet * 15:26 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1078.eqiad.wmnet with OS trixie * 15:25 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 15:25 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:22 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2006.codfw.wmnet * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2021 * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2021 * 15:19 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2021 * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2021.codfw.wmnet 210.48.192.10.in-addr.arpa 0.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:19 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2021.codfw.wmnet 210.48.192.10.in-addr.arpa 0.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2021 - bking@cumin2003" * 15:19 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2021 - bking@cumin2003" * 15:14 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] * 15:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2015.codfw.wmnet with reason: host reimage * 15:11 root@cumin1003: START - Cookbook sre.mysql.pool pool db1247: Maintenance * 15:11 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on ml-serve2004.codfw.wmnet with reason: [[phab:T433478|T433478]] * 15:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2015.codfw.wmnet with reason: host reimage * 15:10 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on ml-serve2002.codfw.wmnet with reason: [[phab:T433476|T433476]] * 15:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 15:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1247 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95625 and previous config saved to /var/cache/conftool/dbconfig/20260729-150459-cwilliams.json * 15:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1247.eqiad.wmnet with reason: Maintenance * 15:04 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:04 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1244: Maintenance * 15:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1078.eqiad.wmnet with reason: host reimage * 15:03 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1005.wikimedia.org * 15:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db1181: Maintenance * 15:01 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2098.codfw.wmnet with OS bullseye * 15:00 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye * 15:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1200: Maintenance * 14:59 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 14:59 root@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1285.eqiad.wmnet with OS trixie * 14:59 Amir1: mwscript-k8s -- extensions/TimedMediaHandler/maintenance/requeueTranscodes.php --wiki=commonswiki --key '360p.mpeg4.mov' --throttle --video --missing ([[phab:T358266|T358266]]) * 14:58 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1005.wikimedia.org * 14:58 jhancock@cumin2002: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['ms-be2098'] * 14:58 jhancock@cumin2002: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['ms-be2098'] * 14:58 jhancock@cumin2002: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['ms-be2097'] * 14:58 jhancock@cumin2002: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['ms-be2097'] * 14:58 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1078.eqiad.wmnet with reason: host reimage * 14:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1181 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95621 and previous config saved to /var/cache/conftool/dbconfig/20260729-145629-cwilliams.json * 14:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1181.eqiad.wmnet with reason: Maintenance * 14:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1174: Maintenance * 14:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1219: Maintenance * 14:55 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader1006.wikimedia.org on all recursors * 14:55 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader1006.wikimedia.org on all recursors * 14:55 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader1005.wikimedia.org on all recursors * 14:55 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader1005.wikimedia.org on all recursors * 14:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1200 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95618 and previous config saved to /var/cache/conftool/dbconfig/20260729-145336-cwilliams.json * 14:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1254: Maintenance * 14:53 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1200.eqiad.wmnet with reason: Maintenance * 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1185: Maintenance * 14:52 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2021 * 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2015 * 14:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2015 * 14:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1219 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95616 and previous config saved to /var/cache/conftool/dbconfig/20260729-144946-cwilliams.json * 14:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1219.eqiad.wmnet with reason: Maintenance * 14:49 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1218: Maintenance * 14:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2021.codfw.wmnet with OS bookworm * 14:48 dancy@deploy1003: Finished deploy [zuul/deploy@22703a6]: Deploying https://gerrit.wikimedia.org/r/c/integration/zuul/+/1311501 ([[phab:T432491|T432491]]) (duration: 00m 15s) * 14:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2015.codfw.wmnet with OS bookworm * 14:48 dancy@deploy1003: Started deploy [zuul/deploy@22703a6]: Deploying https://gerrit.wikimedia.org/r/c/integration/zuul/+/1311501 ([[phab:T432491|T432491]]) * 14:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1254 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95613 and previous config saved to /var/cache/conftool/dbconfig/20260729-144729-cwilliams.json * 14:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1254.eqiad.wmnet with reason: Maintenance * 14:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1233: Maintenance * 14:46 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2013\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 14:46 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2014\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 14:46 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:45 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:44 root@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1285.eqiad.wmnet with reason: host reimage * 14:43 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:42 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:41 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:40 root@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1285.eqiad.wmnet with reason: host reimage * 14:39 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1078.eqiad.wmnet with OS trixie * 14:39 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2074.codfw.wmnet with OS trixie * 14:32 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:32 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:32 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:31 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2005.codfw.wmnet with OS bookworm * 14:30 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:30 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:29 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:29 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:27 root@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host db1285 * 14:27 root@cumin1003: START - Cookbook sre.hosts.move-vlan for host db1285 * 14:27 root@cumin1003: START - Cookbook sre.hosts.reimage for host db1285.eqiad.wmnet with OS trixie * 14:24 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 14:24 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:24 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:24 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:23 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:22 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:22 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:21 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:17 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:16 root@cumin1003: START - Cookbook sre.mysql.pool pool db1244: Maintenance * 14:15 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:15 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add asw1-604 loopback ipv4 - pt1979@cumin2002" * 14:15 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add asw1-604 loopback ipv4 - pt1979@cumin2002" * 14:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:12 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 14:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95599 and previous config saved to /var/cache/conftool/dbconfig/20260729-141014-cwilliams.json * 14:10 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 14:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1244.eqiad.wmnet with reason: Maintenance * 14:10 pt1979@cumin2002: START - Cookbook sre.dns.netbox * 14:10 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 14:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1243: Maintenance * 14:09 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2005.codfw.wmnet with reason: host reimage * 14:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db1174: Maintenance * 14:08 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad * 14:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db1185: Maintenance * 14:06 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 14:05 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2005.codfw.wmnet with reason: host reimage * 14:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1174 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95595 and previous config saved to /var/cache/conftool/dbconfig/20260729-140309-cwilliams.json * 14:03 sukhe@dns1004: END - running authdns-update * 14:03 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1174.eqiad.wmnet with reason: Maintenance * 14:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1170: Maintenance * 14:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db1218: Maintenance * 14:01 sukhe@dns1004: START - running authdns-update * 14:00 sukhe@dns1004: START - running authdns-update * 13:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db1233: Maintenance * 13:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1185 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95592 and previous config saved to /var/cache/conftool/dbconfig/20260729-135925-cwilliams.json * 13:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1185.eqiad.wmnet with reason: Maintenance * 13:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1161: Maintenance * 13:58 sukhe@puppetserver1001: conftool action : set/pooled=true; selector: dnsdisc=urldownloader * 13:58 root@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1265.eqiad.wmnet with OS trixie * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1218 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95590 and previous config saved to /var/cache/conftool/dbconfig/20260729-135621-cwilliams.json * 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1218.eqiad.wmnet with reason: Maintenance * 13:55 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1206: Maintenance * 13:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2073.codfw.wmnet with OS trixie * 13:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1233 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95587 and previous config saved to /var/cache/conftool/dbconfig/20260729-135335-cwilliams.json * 13:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1233.eqiad.wmnet with reason: Maintenance * 13:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1229: Maintenance * 13:50 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/kartotherian: apply * 13:50 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service * 13:49 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:49 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/kartotherian: apply * 13:48 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 13:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1077.eqiad.wmnet with OS trixie * 13:47 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 13:46 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2005.codfw.wmnet with OS bookworm * 13:44 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 13:44 root@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1265.eqiad.wmnet with reason: host reimage * 13:40 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] (duration: 09m 22s) * 13:39 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:38 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:36 root@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1265.eqiad.wmnet with reason: host reimage * 13:35 stran@deploy1003: stran: Continuing with deployment * 13:33 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 13:32 stran@deploy1003: stran: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified t * 13:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2073.codfw.wmnet with reason: host reimage * 13:30 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ml-build1001.eqiad.wmnet * 13:30 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] * 13:29 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad * 13:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1077.eqiad.wmnet with reason: host reimage * 13:27 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 13:27 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:27 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad * 13:26 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2073.codfw.wmnet with reason: host reimage * 13:26 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] (duration: 07m 56s) * 13:25 klausman@cumin1003: START - Cookbook sre.hosts.reboot-single for host ml-build1001.eqiad.wmnet * 13:24 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 13:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>ml-serve101[2-5].eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 13:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1015.eqiad.wmnet * 13:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1015.eqiad.wmnet * 13:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1077.eqiad.wmnet with reason: host reimage * 13:23 root@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host db1265 * 13:23 root@cumin1003: START - Cookbook sre.hosts.move-vlan for host db1265 * 13:23 root@cumin1003: START - Cookbook sre.hosts.reimage for host db1265.eqiad.wmnet with OS trixie * 13:23 root@cumin1003: START - Cookbook sre.mysql.pool pool db1243: Maintenance * 13:22 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 13:22 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2005.codfw.wmnet * 13:22 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 13:22 samtar@deploy1003: dreamrimmer, samtar: Continuing with deployment * 13:20 samtar@deploy1003: dreamrimmer, samtar: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts an-test-master[1001-1002].eqiad.wmnet * 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-master[1001-1002].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 13:18 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1015.eqiad.wmnet * 13:18 sukhe@cumin1003: END (ERROR) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=97) for role: url_downloader@eqiad * 13:18 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 13:18 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] * 13:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95574 and previous config saved to /var/cache/conftool/dbconfig/20260729-131638-cwilliams.json * 13:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1243.eqiad.wmnet with reason: Maintenance * 13:16 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2005.codfw.wmnet * 13:16 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1242: Maintenance * 13:14 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] (duration: 07m 00s) * 13:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db1170: Maintenance * 13:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1015.eqiad.wmnet * 13:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1014.eqiad.wmnet * 13:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1014.eqiad.wmnet * 13:12 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1223: Maintenance * 13:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db1161: Maintenance * 13:10 samtar@deploy1003: anzx, samtar: Continuing with deployment * 13:09 samtar@deploy1003: anzx, samtar: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db1206: Maintenance * 13:08 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 13:07 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] * 13:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1170 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95566 and previous config saved to /var/cache/conftool/dbconfig/20260729-130730-cwilliams.json * 13:07 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1170.eqiad.wmnet with reason: Maintenance * 13:07 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:07 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt IPs new switches - cmooney@cumin1003" * 13:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1158: Maintenance * 13:06 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1014.eqiad.wmnet * 13:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1161 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95564 and previous config saved to /var/cache/conftool/dbconfig/20260729-130616-cwilliams.json * 13:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 13:06 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1077.eqiad.wmnet with OS trixie * 13:05 root@cumin1003: START - Cookbook sre.mysql.pool pool db1229: Maintenance * 13:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1161.eqiad.wmnet with reason: Maintenance * 13:05 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt IPs new switches - cmooney@cumin1003" * 13:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2073.codfw.wmnet with OS trixie * 13:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Maintenance * 13:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1206 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95562 and previous config saved to /var/cache/conftool/dbconfig/20260729-130258-cwilliams.json * 13:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1206.eqiad.wmnet with reason: Maintenance * 13:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1196: Maintenance * 13:01 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 13:01 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:00 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 13:00 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 12:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1229 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95559 and previous config saved to /var/cache/conftool/dbconfig/20260729-125950-cwilliams.json * 12:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1229.eqiad.wmnet with reason: Maintenance * 12:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1222: Maintenance * 12:57 sukhe: sudo cumin 'A:lvs and (A:eqiad or A:codfw)' 'disable-puppet "adding new service urldownloader"': [[phab:T429175|T429175]] * 12:56 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1014.eqiad.wmnet * 12:56 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1013.eqiad.wmnet * 12:56 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1013.eqiad.wmnet * 12:50 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1013.eqiad.wmnet * 12:50 sukhe: sudo cumin 'O:url_downloader' 'run-puppet-agent --enable "merging CR 1313948"': [[phab:T429175|T429175]] * 12:48 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-master[1001-1002].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 12:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1013.eqiad.wmnet * 12:45 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1012.eqiad.wmnet * 12:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1012.eqiad.wmnet * 12:45 sukhe: sudo cumin 'O:url_downloader' 'disable-puppet "merging CR 1313948"': [[phab:T429175|T429175]] * 12:44 btullis@cumin1003: START - Cookbook sre.dns.netbox * 12:40 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test2001.codfw.wmnet * 12:40 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test2001.codfw.wmnet * 12:38 ayounsi@dns1004: END - running authdns-update * 12:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1012.eqiad.wmnet * 12:35 ayounsi@dns1004: START - running authdns-update * 12:34 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts an-test-master[1001-1002].eqiad.wmnet * 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts an-test-coord1001.eqiad.wmnet * 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-coord1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 12:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1012.eqiad.wmnet * 12:32 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>ml-serve101[2-5].eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 12:29 root@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Maintenance * 12:25 root@cumin1003: START - Cookbook sre.mysql.pool pool db1223: Maintenance * 12:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95544 and previous config saved to /var/cache/conftool/dbconfig/20260729-122254-cwilliams.json * 12:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1242.eqiad.wmnet with reason: Maintenance * 12:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1241: Maintenance * 12:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1051 hosts * 12:20 root@cumin1003: START - Cookbook sre.mysql.pool pool db1158: Maintenance * 12:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1223 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95540 and previous config saved to /var/cache/conftool/dbconfig/20260729-121937-cwilliams.json * 12:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1223.eqiad.wmnet with reason: Maintenance * 12:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1212: Maintenance * 12:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db1159: Maintenance * 12:17 elukey@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: sync * 12:15 elukey@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: sync * 12:15 root@cumin1003: START - Cookbook sre.mysql.pool pool db1196: Maintenance * 12:14 Daimona: Creating new DB tables for the CampaignEvents extension in x1.testwiki, x1.test2wiki, x1.officewiki, and x1.wikishared # [[phab:T429339|T429339]] * 12:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db1222: Maintenance * 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95535 and previous config saved to /var/cache/conftool/dbconfig/20260729-121211-cwilliams.json * 12:12 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 12:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1158.eqiad.wmnet with reason: Maintenance * 12:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1159 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95534 and previous config saved to /var/cache/conftool/dbconfig/20260729-121146-cwilliams.json * 12:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1159.eqiad.wmnet with reason: Maintenance * 12:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1196 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95533 and previous config saved to /var/cache/conftool/dbconfig/20260729-120847-cwilliams.json * 12:08 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 12:08 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1196.eqiad.wmnet with reason: Maintenance * 12:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1195: Maintenance * 12:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1222 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95530 and previous config saved to /var/cache/conftool/dbconfig/20260729-120424-cwilliams.json * 12:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1222.eqiad.wmnet with reason: Maintenance * 12:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1098 hosts * 12:00 marostegui: Rename tables [[phab:T425074|T425074]] * 12:00 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-coord1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 11:58 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1197: Maintenance * 11:55 btullis@cumin1003: START - Cookbook sre.dns.netbox * 11:52 elukey@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: sync * 11:51 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:51 elukey@deploy1003: helmfile [codfw] START helmfile.d/services/proton: sync * 11:51 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:50 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts an-test-coord1001.eqiad.wmnet * 11:50 elukey@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: sync * 11:49 elukey@deploy1003: helmfile [staging] START helmfile.d/services/proton: sync * 11:35 root@cumin1003: START - Cookbook sre.mysql.pool pool db1241: Maintenance * 11:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db1212: Maintenance * 11:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1241 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95520 and previous config saved to /var/cache/conftool/dbconfig/20260729-112918-cwilliams.json * 11:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1241.eqiad.wmnet with reason: Maintenance * 11:29 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1238: Maintenance * 11:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1212 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95517 and previous config saved to /var/cache/conftool/dbconfig/20260729-112727-cwilliams.json * 11:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 11:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1212.eqiad.wmnet with reason: Maintenance * 11:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1198: Maintenance * 11:23 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:22 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:21 root@cumin1003: START - Cookbook sre.mysql.pool pool db1195: Maintenance * 11:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1195 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95514 and previous config saved to /var/cache/conftool/dbconfig/20260729-111450-cwilliams.json * 11:14 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1195.eqiad.wmnet with reason: Maintenance * 11:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1186: Maintenance * 11:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 11:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 11:05 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:54 marostegui: Dropping renamed tables [[phab:T425066|T425066]] * 10:41 root@cumin1003: START - Cookbook sre.mysql.pool pool db1238: Maintenance * 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1198: Maintenance * 10:39 Amir1: ran https://phabricator.wikimedia.org/T432509#12149723 in production ([[phab:T432509|T432509]]) * 10:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db1197: Maintenance * 10:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1238 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95501 and previous config saved to /var/cache/conftool/dbconfig/20260729-103532-cwilliams.json * 10:35 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1238.eqiad.wmnet with reason: Maintenance * 10:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1221: Maintenance * 10:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1198 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95499 and previous config saved to /var/cache/conftool/dbconfig/20260729-103330-cwilliams.json * 10:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1198.eqiad.wmnet with reason: Maintenance * 10:33 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1175: Maintenance * 10:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1197 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95496 and previous config saved to /var/cache/conftool/dbconfig/20260729-103217-cwilliams.json * 10:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1197.eqiad.wmnet with reason: Maintenance * 10:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1188: Maintenance * 10:27 root@cumin1003: START - Cookbook sre.mysql.pool pool db1186: Maintenance * 10:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1186 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95493 and previous config saved to /var/cache/conftool/dbconfig/20260729-102111-cwilliams.json * 10:21 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1186.eqiad.wmnet with reason: Maintenance * 10:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 10:14 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 09:53 XioNoX: reboot cr2-magru - [[phab:T431750|T431750]] * 09:52 XioNoX: drain cr2-magru - [[phab:T431750|T431750]] * 09:48 root@cumin1003: START - Cookbook sre.mysql.pool pool db1221: Maintenance * 09:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zookeeper-test1002.eqiad.wmnet * 09:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1188: Maintenance * 09:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1175: Maintenance * 09:44 btullis@dns1004: END - running authdns-update * 09:42 btullis@dns1004: START - running authdns-update * 09:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1221 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95483 and previous config saved to /var/cache/conftool/dbconfig/20260729-094200-cwilliams.json * 09:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 7 hosts with reason: Maintenance * 09:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1221.eqiad.wmnet with reason: Maintenance * 09:41 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host zookeeper-test1002.eqiad.wmnet * 09:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1199: Maintenance * 09:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1188 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95481 and previous config saved to /var/cache/conftool/dbconfig/20260729-093917-cwilliams.json * 09:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1188.eqiad.wmnet with reason: Maintenance * 09:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1182: Maintenance * 09:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1175 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95479 and previous config saved to /var/cache/conftool/dbconfig/20260729-093842-cwilliams.json * 09:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1175.eqiad.wmnet with reason: Maintenance * 09:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1166: Maintenance * 09:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1169: Maintenance * 09:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1033.eqiad.wmnet,service=s8 * 09:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1033.eqiad.wmnet,service=s5 * 09:33 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1033.eqiad.wmnet,service=s8 * 09:33 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1033.eqiad.wmnet,service=s5 * 09:21 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:21 XioNoX: reboot cr1-magru - [[phab:T431750|T431750]] * 09:17 XioNoX: drain cr1-magru - [[phab:T431750|T431750]] * 09:15 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm1001.wikimedia.org * 09:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply * 09:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply * 09:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 09:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 09:11 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr2-magru,cr2-magru IPv6,cr2-magru.mgmt with reason: router upgrade * 09:11 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm1001.wikimedia.org * 09:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp1005.wikimedia.org * 09:07 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp1005.wikimedia.org * 09:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp2005.wikimedia.org * 09:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp2005.wikimedia.org * 09:00 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr1-magru,cr1-magru IPv6,cr1-magru.mgmt with reason: router upgrade * 09:00 marostegui: Dropping renamed tables [[phab:T426341|T426341]] * 08:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1199: Maintenance * 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1182: Maintenance * 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1166: Maintenance * 08:47 root@cumin1003: START - Cookbook sre.mysql.pool pool db1169: Maintenance * 08:46 ayounsi@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 1:00:00 on cr1-magru,cr1-magru IPv6,cr1-magru.mgmt with reason: router upgrade * 08:45 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2072.codfw.wmnet with OS trixie * 08:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1199 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95464 and previous config saved to /var/cache/conftool/dbconfig/20260729-084534-cwilliams.json * 08:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1199.eqiad.wmnet with reason: Maintenance * 08:45 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 08:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1190: Maintenance * 08:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1182 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95462 and previous config saved to /var/cache/conftool/dbconfig/20260729-084436-cwilliams.json * 08:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1182.eqiad.wmnet with reason: Maintenance * 08:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1156: Maintenance * 08:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1166 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95460 and previous config saved to /var/cache/conftool/dbconfig/20260729-084400-cwilliams.json * 08:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1166.eqiad.wmnet with reason: Maintenance * 08:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1157: Maintenance * 08:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95458 and previous config saved to /var/cache/conftool/dbconfig/20260729-084147-cwilliams.json * 08:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1169.eqiad.wmnet with reason: Maintenance * 08:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1163: Maintenance * 08:30 btullis@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 11 hosts with reason: Replacing the namenodes * 08:23 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2072.codfw.wmnet with reason: host reimage * 08:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1098 hosts * 08:19 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2072.codfw.wmnet with reason: host reimage * 07:58 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2072.codfw.wmnet with OS trixie * 07:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1190: Maintenance * 07:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1156: Maintenance * 07:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1157: Maintenance * 07:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1163: Maintenance * 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1190 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95444 and previous config saved to /var/cache/conftool/dbconfig/20260729-074930-cwilliams.json * 07:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1190.eqiad.wmnet with reason: Maintenance * 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1157 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95443 and previous config saved to /var/cache/conftool/dbconfig/20260729-074914-cwilliams.json * 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1156 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95442 and previous config saved to /var/cache/conftool/dbconfig/20260729-074906-cwilliams.json * 07:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1157.eqiad.wmnet with reason: Maintenance * 07:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 07:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1156.eqiad.wmnet with reason: Maintenance * 07:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1163 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95441 and previous config saved to /var/cache/conftool/dbconfig/20260729-074652-cwilliams.json * 07:46 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1163.eqiad.wmnet with reason: Maintenance * 07:46 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2034.codfw.wmnet * 07:42 ayounsi@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2034.codfw.wmnet * 07:42 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2232.codfw.wmnet with OS trixie * 07:34 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:34 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:33 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:31 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:19 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2232.codfw.wmnet with reason: host reimage * 07:15 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2232.codfw.wmnet with reason: host reimage * 06:58 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db2232.codfw.wmnet with OS trixie * 06:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[2160,2232].codfw.wmnet with reason: Reimage * 06:26 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1164.eqiad.wmnet with OS trixie * 06:05 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1164.eqiad.wmnet with reason: host reimage * 06:01 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1164.eqiad.wmnet with reason: host reimage * 05:47 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1164.eqiad.wmnet with OS trixie * 05:46 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1164.eqiad.wmnet with reason: Reimage == 2026-07-28 == * 22:50 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1138.eqiad.wmnet * 22:50 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1138.eqiad.wmnet * 22:49 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1138.eqiad.wmnet * 22:11 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2014.codfw.wmnet, repooling source-only afterwards * 22:08 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2013.codfw.wmnet, repooling source-only afterwards * 22:03 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 20:58 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] (duration: 08m 19s) * 20:55 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2014.codfw.wmnet, repooling source-only afterwards * 20:55 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2013.codfw.wmnet, repooling source-only afterwards * 20:54 arlolra@deploy1003: arlolra: Continuing with deployment * 20:54 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 14s) * 20:54 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 20:53 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 30s) * 20:53 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 20:52 arlolra@deploy1003: arlolra: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:51 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:50 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] * 20:49 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:43 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 20:34 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] (duration: 06m 54s) * 20:34 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:34 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:31 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:31 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:30 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:30 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:30 arlolra@deploy1003: arlolra: Continuing with deployment * 20:29 arlolra@deploy1003: arlolra: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:27 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] * 20:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2014.codfw.wmnet with OS bookworm * 20:21 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 20:21 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:20 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 20:19 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:19 swfrench-wmf: switched etcd-mirror replication from conf2005 to conf2004 - [[phab:T428495|T428495]] * 20:17 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:17 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:15 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] (duration: 08m 26s) * 20:12 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:11 arlolra@deploy1003: anzx, arlolra: Continuing with deployment * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2013.codfw.wmnet with OS bookworm * 20:09 arlolra@deploy1003: anzx, arlolra: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] * 20:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2216: Maintenance * 19:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2014.codfw.wmnet with reason: host reimage * 19:57 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:54 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:54 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:53 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:52 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2014.codfw.wmnet with reason: host reimage * 19:49 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2013.codfw.wmnet with reason: host reimage * 19:42 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 19:41 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:41 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2013.codfw.wmnet with reason: host reimage * 19:39 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:39 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:39 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-eqiad: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 19:38 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2014 * 19:33 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2014 * 19:29 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2014.codfw.wmnet with OS bookworm * 19:28 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:27 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1006 * 19:26 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2012\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 19:26 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1006 * 19:26 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:26 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 19:25 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2013 * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2013 * 19:21 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2013 * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2013.codfw.wmnet 84.0.192.10.in-addr.arpa 4.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:21 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2013.codfw.wmnet 84.0.192.10.in-addr.arpa 4.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2013 - bking@cumin2003" * 19:21 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2013 - bking@cumin2003" * 19:21 vriley@cumin1003: START - Cookbook sre.dns.netbox * 19:20 root@cumin1003: START - Cookbook sre.mysql.pool pool db2216: Maintenance * 19:13 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2216 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95435 and previous config saved to /var/cache/conftool/dbconfig/20260728-191343-cwilliams.json * 19:13 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2216.codfw.wmnet with reason: Maintenance * 19:13 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2203: Maintenance * 19:06 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1005.eqiad.wmnet with OS trixie * 19:06 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 19:06 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 18:46 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 18:45 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:45 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:43 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:40 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:36 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-eqiad: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 18:35 dancy@deploy1003: Installation of scap version "4.275.0" completed for 3 hosts * 18:33 dancy@deploy1003: Installing scap version "4.275.0" for 3 host(s) * 18:32 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:32 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2097.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:30 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns3003.wikimedia.org [reason: pool for all services after reimaging] * 18:29 sukhe@dns1004: END - running authdns-update * 18:27 sukhe@dns1004: START - running authdns-update * 18:27 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns3003.wikimedia.org,service=authdns-update [reason: pool authdns-update after reimaging] * 18:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db2203: Maintenance * 18:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2203 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95430 and previous config saved to /var/cache/conftool/dbconfig/20260728-181958-cwilliams.json * 18:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2203.codfw.wmnet with reason: Maintenance * 18:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2188: Maintenance * 18:18 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2097.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:17 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-be2098 * 18:17 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host ms-be2098 * 18:17 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-be2097 * 18:16 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host ms-be2097 * 18:15 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:15 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding ms-be2097-8 to codfw - jhancock@cumin2002" * 18:15 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding ms-be2097-8 to codfw - jhancock@cumin2002" * 18:10 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 18:08 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage * 18:05 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns3003.wikimedia.org with OS trixie * 18:03 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage * 17:56 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-codfw: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 17:45 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie * 17:45 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1005.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:41 sukhe@dns1004: END - running authdns-update * 17:39 sukhe@dns1004: START - running authdns-update * 17:36 sukhe@puppetserver1001: conftool action : set/weight=1; selector: cluster=urldownloader,service=squid * 17:36 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1005.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:35 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader,service=squid * 17:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 17:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1005 * 17:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 17:34 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1005 * 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1005] - vriley@cumin1003" * 17:34 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1005] - vriley@cumin1003" * 17:32 root@cumin1003: START - Cookbook sre.mysql.pool pool db2188: Maintenance * 17:29 vriley@cumin1003: START - Cookbook sre.dns.netbox * 17:29 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 17:26 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2188 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95425 and previous config saved to /var/cache/conftool/dbconfig/20260728-172609-cwilliams.json * 17:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2188.codfw.wmnet with reason: Maintenance * 17:25 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2176: Maintenance * 17:19 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1138.eqiad.wmnet with OS trixie * 17:18 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1005 * 17:18 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1005 * 17:18 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:15 vriley@cumin1003: START - Cookbook sre.dns.netbox * 17:13 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns3003.wikimedia.org with reason: host reimage * 17:07 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns3003.wikimedia.org with reason: host reimage * 17:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-eqiad * 17:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1015.eqiad.wmnet * 17:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1015.eqiad.wmnet * 16:59 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1138.eqiad.wmnet with reason: host reimage * 16:55 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-codfw: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 16:54 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1138.eqiad.wmnet with reason: host reimage * 16:53 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1015.eqiad.wmnet * 16:43 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns3003.wikimedia.org with OS trixie * 16:43 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1015.eqiad.wmnet * 16:43 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1014.eqiad.wmnet * 16:43 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1014.eqiad.wmnet * 16:43 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=dns3003.wikimedia.org [reason: depooling for reimage to trixie] * 16:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 290 hosts * 16:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db2176: Maintenance * 16:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1138 * 16:38 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1138 * 16:37 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1138 * 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1138.eqiad.wmnet 193.32.64.10.in-addr.arpa 3.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:37 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1138.eqiad.wmnet 193.32.64.10.in-addr.arpa 3.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1138 - jiji@cumin1003" * 16:37 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1138 - jiji@cumin1003" * 16:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1014.eqiad.wmnet * 16:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2176 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95420 and previous config saved to /var/cache/conftool/dbconfig/20260728-163235-cwilliams.json * 16:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2176.codfw.wmnet with reason: Maintenance * 16:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1014.eqiad.wmnet * 16:32 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1013.eqiad.wmnet * 16:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1013.eqiad.wmnet * 16:32 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2174: Maintenance * 16:28 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2012.codfw.wmnet, repooling source-only afterwards * 16:25 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1013.eqiad.wmnet * 16:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1013.eqiad.wmnet * 16:20 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1012.eqiad.wmnet * 16:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1012.eqiad.wmnet * 16:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1012.eqiad.wmnet * 16:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1012.eqiad.wmnet * 16:03 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1011.eqiad.wmnet * 16:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1011.eqiad.wmnet * 16:00 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2004.codfw.wmnet with OS bookworm * 15:59 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1011.eqiad.wmnet * 15:56 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:55 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 15:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:54 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1011.eqiad.wmnet * 15:54 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1010.eqiad.wmnet * 15:54 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1010.eqiad.wmnet * 15:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:50 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 15:49 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1010.eqiad.wmnet * 15:48 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 15:48 jiji@cumin1003: START - Cookbook sre.dns.netbox * 15:46 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2248: Maintenance * 15:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db2174: Maintenance * 15:44 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1010.eqiad.wmnet * 15:44 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1009.eqiad.wmnet * 15:44 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1009.eqiad.wmnet * 15:42 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1138 * 15:41 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1138.eqiad.wmnet with OS trixie * 15:39 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1009.eqiad.wmnet * 15:39 robh@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on arclamp2001.codfw.wmnet with reason: ram upgrade * 15:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2174 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95413 and previous config saved to /var/cache/conftool/dbconfig/20260728-153844-cwilliams.json * 15:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2174.codfw.wmnet with reason: Maintenance * 15:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2173: Maintenance * 15:37 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 15:35 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 15:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1009.eqiad.wmnet * 15:34 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1008.eqiad.wmnet * 15:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1008.eqiad.wmnet * 15:31 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 15:31 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 15:29 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1008.eqiad.wmnet * 15:27 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 290 hosts * 15:25 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2013 * 15:25 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2195: Maintenance * 15:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1008.eqiad.wmnet * 15:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1007.eqiad.wmnet * 15:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1007.eqiad.wmnet * 15:22 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2004.codfw.wmnet with reason: host reimage * 15:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2013.codfw.wmnet with OS bookworm * 15:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 15:19 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2004.codfw.wmnet with reason: host reimage * 15:19 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 15:17 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1007.eqiad.wmnet * 15:12 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1007.eqiad.wmnet * 15:12 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1006.eqiad.wmnet * 15:12 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1006.eqiad.wmnet * 15:11 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1138.eqiad.wmnet * 15:11 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 15:11 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1138.eqiad.wmnet * 15:11 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1138.eqiad.wmnet * 15:10 brennen@deploy1003: Finished deploy [phabricator/deployment@f8b349f]: deploy phab1004 for [[phab:T433382|T433382]] (duration: 00m 43s) * 15:10 brennen@deploy1003: Started deploy [phabricator/deployment@f8b349f]: deploy phab1004 for [[phab:T433382|T433382]] * 15:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts * 15:09 brennen@deploy1003: Finished deploy [phabricator/deployment@f8b349f]: deploy phab2003 for [[phab:T433382|T433382]] (duration: 00m 55s) * 15:08 brennen@deploy1003: Started deploy [phabricator/deployment@f8b349f]: deploy phab2003 for [[phab:T433382|T433382]] * 15:07 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2012.codfw.wmnet, repooling source-only afterwards * 15:07 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts * 15:06 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1137.eqiad.wmnet * 15:06 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1137.eqiad.wmnet * 15:06 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1137.eqiad.wmnet * 15:05 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1006.eqiad.wmnet * 15:05 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1022\.eqiad\.wmnet,dc=eqiad,cluster=wdqs\-main,service=wdqs\-main * 15:01 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab2003.codfw.wmnet with reason: deployment * 15:01 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1005.eqiad.wmnet with reason: deployment * 15:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1006.eqiad.wmnet * 15:00 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1006.eqiad.wmnet with reason: deployment * 15:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1005.eqiad.wmnet * 15:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1005.eqiad.wmnet * 14:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db2248: Maintenance * 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1004.eqiad.wmnet with reason: deployment * 14:59 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2004.codfw.wmnet with OS bookworm * 14:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1005.eqiad.wmnet * 14:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2248 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95403 and previous config saved to /var/cache/conftool/dbconfig/20260728-145532-cwilliams.json * 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2245-2247].codfw.wmnet with reason: Maintenance * 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2248.codfw.wmnet with reason: Maintenance * 14:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2240: Maintenance * 14:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2173: Maintenance * 14:51 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts * 14:50 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1005.eqiad.wmnet * 14:50 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1004.eqiad.wmnet * 14:50 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1004.eqiad.wmnet * 14:49 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts * 14:45 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2004.codfw.wmnet * 14:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2173 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95399 and previous config saved to /var/cache/conftool/dbconfig/20260728-144453-cwilliams.json * 14:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2173.codfw.wmnet with reason: Maintenance * 14:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2170: Maintenance * 14:44 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1004.eqiad.wmnet * 14:39 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2004.codfw.wmnet * 14:38 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1004.eqiad.wmnet * 14:38 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet * 14:38 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet * 14:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db2195: Maintenance * 14:36 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1022.eqiad.wmnet, repooling source-only afterwards * 14:36 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 14:33 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet * 14:33 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2222: Maintenance * 14:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2195 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95394 and previous config saved to /var/cache/conftool/dbconfig/20260728-143218-cwilliams.json * 14:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2195.codfw.wmnet with reason: Maintenance * 14:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2181: Maintenance * 14:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 14:30 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 14:25 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 14:25 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 14:25 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 14:23 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet * 14:23 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1002.eqiad.wmnet * 14:23 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1002.eqiad.wmnet * 14:23 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 19s) * 14:23 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:18 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1002.eqiad.wmnet * 14:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2012.codfw.wmnet with OS bookworm * 14:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1002.eqiad.wmnet * 14:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1001.eqiad.wmnet * 14:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1001.eqiad.wmnet * 14:11 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 14:11 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 14:08 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1001.eqiad.wmnet * 14:07 elukey@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'. * 14:07 elukey@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'. * 14:06 elukey@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'. * 14:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db2240: Maintenance * 14:06 XioNoX: un-drain cr2-esams - [[phab:T431751|T431751]] * 14:05 elukey@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'. * 14:02 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1001.eqiad.wmnet * 14:02 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-eqiad * 14:01 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 14:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2240 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95384 and previous config saved to /var/cache/conftool/dbconfig/20260728-140011-cwilliams.json * 14:00 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2240.codfw.wmnet with reason: Maintenance * 13:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2237: Maintenance * 13:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db2170: Maintenance * 13:56 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 13:55 XioNoX: reboot cr2-esams - [[phab:T431751|T431751]] * 13:52 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 13:51 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr2-esams,cr2-esams IPv6,cr2-esams.mgmt with reason: router upgrade * 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 13:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2170 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95381 and previous config saved to /var/cache/conftool/dbconfig/20260728-135043-cwilliams.json * 13:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2170.codfw.wmnet with reason: Maintenance * 13:50 XioNoX: drain cr2-esams - [[phab:T431751|T431751]] * 13:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2153: Maintenance * 13:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2012.codfw.wmnet with reason: host reimage * 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2222: Maintenance * 13:45 sukhe: restart pybal on A:lvs-codfw * 13:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2012.codfw.wmnet with reason: host reimage * 13:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db2181: Maintenance * 13:44 btullis@dns1004: END - running authdns-update * 13:42 sukhe: restart pybal on lvs2014 * 13:42 btullis@dns1004: START - running authdns-update * 13:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2222 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95376 and previous config saved to /var/cache/conftool/dbconfig/20260728-133948-cwilliams.json * 13:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2222.codfw.wmnet with reason: Maintenance * 13:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2221: Maintenance * 13:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2181 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95374 and previous config saved to /var/cache/conftool/dbconfig/20260728-133857-cwilliams.json * 13:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2181.codfw.wmnet with reason: Maintenance * 13:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2167: Maintenance * 13:30 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T428495|T428495]] * 13:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 13:29 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 13:29 ayounsi@cumin1003: END (FAIL) - Cookbook sre.dns.admin (exit_code=99) DNS admin: depool esams [reason: router upgrade, [[phab:T431749|T431749]]] * 13:28 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: router upgrade, [[phab:T431749|T431749]]] * 13:27 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1022.eqiad.wmnet, repooling source-only afterwards * 13:27 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo - [[phab:T428495|T428495]] * 13:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2012 * 13:27 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2012 * 13:21 lucaswerkmeister-wmde@deploy1003: mwscript-k8s job started: cleanupTitles bolwiki # [[phab:T429951|T429951]] * 13:21 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] (duration: 07m 19s) * 13:20 swfrench-wmf: authdns-update to direct codfw, eqsin, ulsfo etcd clients to eqiad - [[phab:T428495|T428495]] * 13:18 swfrench@dns1004: END - running authdns-update * 13:17 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, anzx: Continuing with deployment * 13:16 swfrench@dns1004: START - running authdns-update * 13:16 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2012 * 13:16 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2012.codfw.wmnet 57.48.192.10.in-addr.arpa 7.5.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:16 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2012.codfw.wmnet 57.48.192.10.in-addr.arpa 7.5.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:16 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:16 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, anzx: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:14 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 13:14 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] * 13:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db2237: Maintenance * 13:13 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2237: Maintenance * 13:13 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:12 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:12 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback IPV6 for asw1-604 - pt1979@cumin2003" * 13:12 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:12 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback IPV6 for asw1-604 - pt1979@cumin2003" * 13:11 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 20s) * 13:11 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 13:10 esanders@deploy1003: Finished scap sync-world: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] (duration: 08m 11s) * 13:08 pt1979@cumin2003: START - Cookbook sre.dns.netbox * 13:07 root@cumin1003: START - Cookbook sre.mysql.pool pool db2237: Maintenance * 13:06 esanders@deploy1003: esanders: Continuing with deployment * 13:04 esanders@deploy1003: esanders: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2153: Maintenance * 13:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2153: Maintenance * 13:02 esanders@deploy1003: Started scap sync-world: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] * 13:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2237 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95362 and previous config saved to /var/cache/conftool/dbconfig/20260728-130107-cwilliams.json * 13:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2237.codfw.wmnet with reason: Maintenance * 13:00 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2236: Maintenance * 12:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2153: Maintenance * 12:52 root@cumin1003: START - Cookbook sre.mysql.pool pool db2221: Maintenance * 12:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2153 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95358 and previous config saved to /var/cache/conftool/dbconfig/20260728-125214-cwilliams.json * 12:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2153.codfw.wmnet with reason: Maintenance * 12:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2167: Maintenance * 12:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-codfw * 12:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2011.codfw.wmnet * 12:51 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2011.codfw.wmnet * 12:49 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:49 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback for asw1-603 - pt1979@cumin2003" * 12:48 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback for asw1-603 - pt1979@cumin2003" * 12:46 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2011.codfw.wmnet * 12:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2221 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95357 and previous config saved to /var/cache/conftool/dbconfig/20260728-124601-cwilliams.json * 12:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2221.codfw.wmnet with reason: Maintenance * 12:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2218: Maintenance * 12:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2167 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95354 and previous config saved to /var/cache/conftool/dbconfig/20260728-124457-cwilliams.json * 12:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2167.codfw.wmnet with reason: Maintenance * 12:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2166: Maintenance * 12:42 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 12:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2011.codfw.wmnet * 12:41 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2010.codfw.wmnet * 12:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2010.codfw.wmnet * 12:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2010.codfw.wmnet * 12:34 pt1979@cumin2003: START - Cookbook sre.dns.netbox * 12:32 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 12:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2010.codfw.wmnet * 12:31 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2009.codfw.wmnet * 12:31 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2009.codfw.wmnet * 12:27 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2009.codfw.wmnet * 12:22 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2009.codfw.wmnet * 12:22 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2008.codfw.wmnet * 12:21 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2008.codfw.wmnet * 12:16 pt1979@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-604-eqsin * 12:16 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2008.codfw.wmnet * 12:16 pt1979@cumin1003: START - Cookbook sre.network.tls for network device asw1-604-eqsin * 12:14 root@cumin1003: START - Cookbook sre.mysql.pool pool db2236: Maintenance * 12:14 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2236: Maintenance * 12:12 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1137.eqiad.wmnet with OS trixie * 12:11 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2008.codfw.wmnet * 12:11 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2007.codfw.wmnet * 12:11 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2007.codfw.wmnet * 12:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db2236: Maintenance * 12:06 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2007.codfw.wmnet * 12:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2236 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95348 and previous config saved to /var/cache/conftool/dbconfig/20260728-120253-cwilliams.json * 12:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2236.codfw.wmnet with reason: Maintenance * 12:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2007.codfw.wmnet * 12:01 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 12:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 11:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2218: Maintenance * 11:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db2166: Maintenance * 11:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2219: Maintenance * 11:56 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2006.codfw.wmnet * 11:52 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1137.eqiad.wmnet with reason: host reimage * 11:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2218 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95344 and previous config saved to /var/cache/conftool/dbconfig/20260728-115155-cwilliams.json * 11:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2218.codfw.wmnet with reason: Maintenance * 11:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2208: Maintenance * 11:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2166 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95342 and previous config saved to /var/cache/conftool/dbconfig/20260728-115119-cwilliams.json * 11:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2166.codfw.wmnet with reason: Maintenance * 11:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2164: Maintenance * 11:47 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1137.eqiad.wmnet with reason: host reimage * 11:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2006.codfw.wmnet * 11:45 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2005.codfw.wmnet * 11:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2005.codfw.wmnet * 11:40 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2005.codfw.wmnet * 11:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2005.codfw.wmnet * 11:35 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2004.codfw.wmnet * 11:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2004.codfw.wmnet * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1137 * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1137 * 11:30 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1137 * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1137.eqiad.wmnet 192.32.64.10.in-addr.arpa 2.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:30 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1137.eqiad.wmnet 192.32.64.10.in-addr.arpa 2.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1137 - jiji@cumin1003" * 11:25 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2004.codfw.wmnet * 11:19 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2004.codfw.wmnet * 11:19 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2003.codfw.wmnet * 11:19 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2003.codfw.wmnet * 11:14 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2003.codfw.wmnet * 11:11 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2219: Maintenance * 11:10 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2219: Maintenance * 11:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2219: Maintenance * 11:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2003.codfw.wmnet * 11:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2164: Maintenance * 11:03 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2002.codfw.wmnet * 11:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2002.codfw.wmnet * 11:03 root@cumin1003: START - Cookbook sre.mysql.pool pool db2208: Maintenance * 10:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2164 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95332 and previous config saved to /var/cache/conftool/dbconfig/20260728-105749-cwilliams.json * 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2164.codfw.wmnet with reason: Maintenance * 10:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2208 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95331 and previous config saved to /var/cache/conftool/dbconfig/20260728-105711-cwilliams.json * 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2208.codfw.wmnet with reason: Maintenance * 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2219 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95330 and previous config saved to /var/cache/conftool/dbconfig/20260728-105652-cwilliams.json * 10:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2219.codfw.wmnet with reason: Maintenance * 10:53 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1137 - jiji@cumin1003" * 10:52 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2002.codfw.wmnet * 10:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2002.codfw.wmnet * 10:47 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2001.codfw.wmnet * 10:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2001.codfw.wmnet * 10:39 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2001.codfw.wmnet * 10:35 jiji@cumin1003: START - Cookbook sre.dns.netbox * 10:34 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] (duration: 09m 31s) * 10:34 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1137 * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2001.codfw.wmnet * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-codfw * 10:34 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1137.eqiad.wmnet with OS trixie * 10:28 jforrester@deploy1003: jforrester: Continuing with deployment * 10:27 jforrester@deploy1003: jforrester: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:25 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] * 10:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 10:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 10:21 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 10:20 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1137.eqiad.wmnet * 10:20 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 10:20 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1137.eqiad.wmnet * 10:20 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1137.eqiad.wmnet * 10:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2163: Maintenance * 09:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-staging-worker * 09:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2003.codfw.wmnet * 09:37 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2003.codfw.wmnet * 09:32 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 09:31 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2003.codfw.wmnet * 09:30 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 09:30 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2163: Maintenance * 09:30 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 09:30 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 09:30 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:22 klausman@cumin1003: END (ERROR) - Cookbook sre.ganeti.reboot-vm (exit_code=97) for VM ml-serve-ctrl2001.codfw.wmnet * 09:22 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2001.codfw.wmnet * 09:22 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d8-eqiad * 09:22 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d8-eqiad * 09:21 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2003.codfw.wmnet * 09:20 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2002.codfw.wmnet * 09:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2002.codfw.wmnet * 09:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f2-codfw * 09:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f2-codfw * 09:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e4-codfw * 09:18 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2163: Maintenance * 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e4-codfw * 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-codfw * 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-codfw * 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e5-codfw * 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e5-codfw * 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f4-codfw * 09:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2210: Maintenance * 09:16 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f4-codfw * 09:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2182: Maintenance * 09:14 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2002.codfw.wmnet * 09:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db2163: Maintenance * 09:11 XioNoX: rebooting cr2-drmrs - [[phab:T431749|T431749]] * 09:10 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr2-drmrs,cr2-drmrs IPv6,cr2-drmrs.mgmt with reason: router upgrade * 09:06 XioNoX: draining cr2-drmrs - [[phab:T431749|T431749]] * 09:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2163 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95320 and previous config saved to /var/cache/conftool/dbconfig/20260728-090638-cwilliams.json * 09:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2163.codfw.wmnet with reason: Maintenance * 09:06 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2161: Maintenance * 09:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2002.codfw.wmnet * 09:04 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2001.codfw.wmnet * 09:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2001.codfw.wmnet * 08:57 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2001.codfw.wmnet * 08:48 XioNoX: un-drain cr1-drmrs - [[phab:T431749|T431749]] * 08:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2001.codfw.wmnet * 08:47 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-staging-worker * 08:42 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:35 XioNoX: rebooting cr1-drmrs - [[phab:T431749|T431749]] * 08:33 XioNoX: draining cr1-drmrs - [[phab:T431749|T431749]] * 08:31 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2210: Maintenance * 08:29 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2182: Maintenance * 08:21 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2182: Maintenance * 08:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db2161: Maintenance * 08:16 root@cumin1003: START - Cookbook sre.mysql.pool pool db2182: Maintenance * 08:12 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2210: Maintenance * 08:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2161 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95309 and previous config saved to /var/cache/conftool/dbconfig/20260728-081044-cwilliams.json * 08:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2161.codfw.wmnet with reason: Maintenance * 08:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2154: Maintenance * 08:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2182 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95307 and previous config saved to /var/cache/conftool/dbconfig/20260728-080947-cwilliams.json * 08:09 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2182.codfw.wmnet with reason: Maintenance * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2168: Maintenance * 08:06 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr1-drmrs,cr1-drmrs IPv6,cr1-drmrs.mgmt with reason: router upgrade * 08:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db2210: Maintenance * 08:05 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 08:05 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 08:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2210 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95305 and previous config saved to /var/cache/conftool/dbconfig/20260728-080008-cwilliams.json * 08:00 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2210.codfw.wmnet with reason: Maintenance * 07:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2206: Maintenance * 07:50 gkyziridis@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 07:50 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 07:22 root@cumin1003: START - Cookbook sre.mysql.pool pool db2154: Maintenance * 07:22 root@cumin1003: START - Cookbook sre.mysql.pool pool db2168: Maintenance * 07:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2154 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95295 and previous config saved to /var/cache/conftool/dbconfig/20260728-071640-cwilliams.json * 07:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2154.codfw.wmnet with reason: Maintenance * 07:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95294 and previous config saved to /var/cache/conftool/dbconfig/20260728-071604-cwilliams.json * 07:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2168.codfw.wmnet with reason: Maintenance * 07:08 root@cumin1003: START - Cookbook sre.mysql.pool pool db2206: Maintenance * 07:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2206 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95292 and previous config saved to /var/cache/conftool/dbconfig/20260728-070219-cwilliams.json * 07:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2206.codfw.wmnet with reason: Maintenance * 06:44 marostegui: Failover m5 from db1164 to db1228 - [[phab:T432967|T432967]] * 06:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2235].codfw.wmnet,db[1164,1217,1228].eqiad.wmnet with reason: m5 master switch [[phab:T432967|T432967]] * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.10 (duration: 02m 34s) * 03:39 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] (duration: 36m 06s) * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 02:57 dzahn@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1004.eqiad.wmnet with OS trixie * 02:57 dzahn@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - dzahn@cumin1003" * 02:55 dzahn@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - dzahn@cumin1003" * 02:37 dzahn@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1004.eqiad.wmnet with reason: host reimage * 02:31 dzahn@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1004.eqiad.wmnet with reason: host reimage * 02:16 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie * 02:15 dzahn@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host zuul1004.eqiad.wmnet with OS trixie * 01:43 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie * 01:43 dzahn@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1004.eqiad.wmnet with OS trixie * 01:25 pt1979@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-603-eqsin * 01:24 pt1979@cumin1003: START - Cookbook sre.network.tls for network device asw1-603-eqsin * 01:12 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 01:12 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt for new switches in eqsin - pt1979@cumin2003" * 01:12 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt for new switches in eqsin - pt1979@cumin2003" * 01:08 pt1979@cumin2003: START - Cookbook sre.dns.netbox * 00:48 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 00:47 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 00:47 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 00:47 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 00:26 mutante: attempting reimage with trixie on zuul1004 re-purposed physical hardware - dcops reported install issue - host was in busybox shell ([[phab:T427353|T427353]]) * 00:24 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie == 2026-07-27 == * 23:50 Amir1: mass deleting vp8 transcodes * 23:28 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:27 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1004.eqiad.wmnet with OS bullseye * 23:26 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 23:25 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:25 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 22:39 maryum: Deploy security fix for [[phab:T432877|T432877]] * 22:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1022.eqiad.wmnet with OS bookworm * 22:37 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS bullseye * 22:32 sbassett: Deployed security fix for [[phab:T432789|T432789]] * 22:22 sbassett: Deployed security patch for [[phab:T431819|T431819]] * 22:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1022.eqiad.wmnet with reason: host reimage * 22:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1022.eqiad.wmnet with reason: host reimage * 22:01 RScout-WMF: Deployed security fix for [[phab:T431819|T431819]] * 22:00 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2012 * 21:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2012.codfw.wmnet with OS bookworm * 21:55 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2011\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 21:45 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1022 * 21:45 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1022 * 21:44 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1022 * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1022.eqiad.wmnet 239.48.64.10.in-addr.arpa 9.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:44 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1022.eqiad.wmnet 239.48.64.10.in-addr.arpa 9.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:41 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:41 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 21:34 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS bookworm * 21:31 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:22 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:21 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:19 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:17 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1004 * 21:16 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1004 * 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1004] - vriley@cumin1003" * 21:15 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1004] - vriley@cumin1003" * 21:11 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:10 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2011.codfw.wmnet, repooling source-only afterwards * 21:05 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:01 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1022 * 20:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1022.eqiad.wmnet with OS bookworm * 20:53 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1021.eqiad.wmnet, repooling source-only afterwards * 20:51 mutante: zuul1001 - re-enabled puppet - revert "cherry-picked" gerrit:1314120 - [[phab:T431003|T431003]] * 20:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Maintenance * 20:15 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] (duration: 08m 03s) * 20:11 sbisson@deploy1003: sbisson: Continuing with deployment * 20:09 sbisson@deploy1003: sbisson: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] * 19:47 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 46s) * 19:47 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Maintenance * 19:27 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] (duration: 12m 26s) * 19:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2228 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95285 and previous config saved to /var/cache/conftool/dbconfig/20260727-192711-cwilliams.json * 19:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2228.codfw.wmnet with reason: Maintenance * 19:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2223: Maintenance * 19:23 krinkle@deploy1003: krinkle: Continuing with deployment * 19:16 krinkle@deploy1003: krinkle: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:15 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] * 19:12 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2238: Maintenance * 18:58 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 18:57 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 18:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2227: Maintenance * 18:57 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-experimental: apply * 18:55 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-experimental: apply * 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1021.eqiad.wmnet with OS bookworm * 18:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2011.codfw.wmnet with OS bookworm * 18:40 root@cumin1003: START - Cookbook sre.mysql.pool pool db2223: Maintenance * 18:39 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] (duration: 07m 05s) * 18:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2223 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95275 and previous config saved to /var/cache/conftool/dbconfig/20260727-183500-cwilliams.json * 18:34 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2223.codfw.wmnet with reason: Maintenance * 18:34 musikanimal@deploy1003: musikanimal: Continuing with deployment * 18:34 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2213: Maintenance * 18:33 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:32 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] * 18:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db2238: Maintenance * 18:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2011.codfw.wmnet with reason: host reimage * 18:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2238 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95271 and previous config saved to /var/cache/conftool/dbconfig/20260727-181944-cwilliams.json * 18:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2238.codfw.wmnet with reason: Maintenance * 18:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2226: Maintenance * 18:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1021.eqiad.wmnet with reason: host reimage * 18:14 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2011.codfw.wmnet with reason: host reimage * 18:12 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1021.eqiad.wmnet with reason: host reimage * 18:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db2227: Maintenance * 18:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2227 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95265 and previous config saved to /var/cache/conftool/dbconfig/20260727-180256-cwilliams.json * 18:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2227.codfw.wmnet with reason: Maintenance * 18:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2194: Maintenance * 17:57 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2011 * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2011 * 17:56 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2011 * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2011.codfw.wmnet 37.32.192.10.in-addr.arpa 7.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:56 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2011.codfw.wmnet 37.32.192.10.in-addr.arpa 7.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2011 - bking@cumin2003" * 17:56 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2011 - bking@cumin2003" * 17:52 bking@cumin2003: START - Cookbook sre.dns.netbox * 17:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2011 * 17:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1021 * 17:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1021 * 17:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2011.codfw.wmnet with OS bookworm * 17:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1021.eqiad.wmnet with OS bookworm * 17:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Maintenance * 17:38 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2010\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 17:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2213 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95260 and previous config saved to /var/cache/conftool/dbconfig/20260727-173740-cwilliams.json * 17:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2213.codfw.wmnet with reason: Maintenance * 17:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2211: Maintenance * 17:36 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1020\.eqiad\.wmnet,dc=eqiad,cluster=wdqs\-main,service=wdqs\-main * 17:32 root@cumin1003: START - Cookbook sre.mysql.pool pool db2226: Maintenance * 17:31 taavi@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] (duration: 06m 33s) * 17:27 taavi@deploy1003: taavi: Continuing with deployment * 17:27 taavi@deploy1003: taavi: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:26 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2226 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95256 and previous config saved to /var/cache/conftool/dbconfig/20260727-172636-cwilliams.json * 17:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2226.codfw.wmnet with reason: Maintenance * 17:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2225: Maintenance * 17:25 taavi@deploy1003: Started scap sync-world: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] * 17:13 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 17:11 root@cumin1003: START - Cookbook sre.mysql.pool pool db2194: Maintenance * 17:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2194 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95248 and previous config saved to /var/cache/conftool/dbconfig/20260727-170453-cwilliams.json * 17:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2194.codfw.wmnet with reason: Maintenance * 17:04 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2190: Maintenance * 16:52 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 16:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2211: Maintenance * 16:40 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95242 and previous config saved to /var/cache/conftool/dbconfig/20260727-164015-cwilliams.json * 16:40 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2211.codfw.wmnet with reason: Maintenance * 16:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2178: Maintenance * 16:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db2225: Maintenance * 16:39 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2172: Maintenance * 16:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 16:38 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 16:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2225 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95238 and previous config saved to /var/cache/conftool/dbconfig/20260727-163307-cwilliams.json * 16:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2225.codfw.wmnet with reason: Maintenance * 16:32 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2189: Maintenance * 16:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db2190: Maintenance * 16:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2190 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95230 and previous config saved to /var/cache/conftool/dbconfig/20260727-160602-cwilliams.json * 16:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2190.codfw.wmnet with reason: Maintenance * 15:53 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2177: Maintenance * 15:53 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2172: Maintenance * 15:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2178: Maintenance * 15:51 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2172: Maintenance * 15:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2178 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95224 and previous config saved to /var/cache/conftool/dbconfig/20260727-154559-cwilliams.json * 15:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2172: Maintenance * 15:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2178.codfw.wmnet with reason: Maintenance * 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2171: Maintenance * 15:44 root@cumin1003: START - Cookbook sre.mysql.pool pool db2189: Maintenance * 15:43 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:41 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2172 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95222 and previous config saved to /var/cache/conftool/dbconfig/20260727-153927-cwilliams.json * 15:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2172.codfw.wmnet with reason: Maintenance * 15:38 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2189 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95220 and previous config saved to /var/cache/conftool/dbconfig/20260727-153833-cwilliams.json * 15:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2189.codfw.wmnet with reason: Maintenance * 15:34 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:32 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 15:32 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 15:31 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:29 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:26 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 15:22 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] (duration: 07m 00s) * 15:21 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2155: Maintenance * 15:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2175: Maintenance * 15:18 zabe@deploy1003: zabe: Continuing with deployment * 15:17 zabe@deploy1003: zabe: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:15 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 15:15 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:15 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] * 15:15 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 06s) * 15:15 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:12 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2010.codfw.wmnet with OS bookworm * 15:08 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2177: Maintenance * 15:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2177: Maintenance * 14:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2177: Maintenance * 14:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2171: Maintenance * 14:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2171 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95209 and previous config saved to /var/cache/conftool/dbconfig/20260727-145236-cwilliams.json * 14:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2171.codfw.wmnet with reason: Maintenance * 14:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2177 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95208 and previous config saved to /var/cache/conftool/dbconfig/20260727-145206-cwilliams.json * 14:52 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2157: Maintenance * 14:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2177.codfw.wmnet with reason: Maintenance * 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1020.eqiad.wmnet with OS bookworm * 14:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2156: Maintenance * 14:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2010.codfw.wmnet with reason: host reimage * 14:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2010.codfw.wmnet with reason: host reimage * 14:41 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[1031,2024]*: Upgrade Cassandra to 5.0.8 (canary) - eevans@cumin1003 * 14:34 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2155: Maintenance * 14:33 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2175: Maintenance * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2010 * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2010 * 14:24 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2010 * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2010.codfw.wmnet 94.16.192.10.in-addr.arpa 4.9.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2010.codfw.wmnet 94.16.192.10.in-addr.arpa 4.9.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2010 - bking@cumin2003" * 14:24 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2010 - bking@cumin2003" * 14:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1020.eqiad.wmnet with reason: host reimage * 14:23 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[1031,2024]*: Upgrade Cassandra to 5.0.8 (canary) - eevans@cumin1003 * 14:20 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 14:20 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 14:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1020.eqiad.wmnet with reason: host reimage * 14:17 sukhe: sudo gnt-instance reboot urldownloader1005.wikimedia.org * 14:16 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:15 jelto@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:08 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2155: Maintenance * 14:05 root@cumin1003: START - Cookbook sre.mysql.pool pool db2157: Maintenance * 14:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2175: Maintenance * 14:03 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 11 hosts * 14:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db2155: Maintenance * 14:01 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 11 hosts * 14:01 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1136.eqiad.wmnet * 14:01 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1136.eqiad.wmnet * 14:01 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1136.eqiad.wmnet * 14:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db2156: Maintenance * 13:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2157 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95194 and previous config saved to /var/cache/conftool/dbconfig/20260727-135943-cwilliams.json * 13:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2157.codfw.wmnet with reason: Maintenance * 13:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db2175: Maintenance * 13:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 13:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 13:57 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2010 * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2155 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95193 and previous config saved to /var/cache/conftool/dbconfig/20260727-135613-cwilliams.json * 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2155.codfw.wmnet with reason: Maintenance * 13:55 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1020 * 13:55 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1020 * 13:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2156 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95192 and previous config saved to /var/cache/conftool/dbconfig/20260727-135413-cwilliams.json * 13:54 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2156.codfw.wmnet with reason: Maintenance * 13:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2175 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95191 and previous config saved to /var/cache/conftool/dbconfig/20260727-135300-cwilliams.json * 13:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2010.codfw.wmnet with OS bookworm * 13:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2175.codfw.wmnet with reason: Maintenance * 13:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1020.eqiad.wmnet with OS bookworm * 13:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 34 hosts * 13:46 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 34 hosts * 13:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1201: Maintenance * 13:27 Lucas_WMDE: UTC afternoon backport+config window doen * 13:18 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] (duration: 11m 57s) * 13:14 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, sihe: Continuing with deployment * 13:08 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, sihe: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:07 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool ulsfo [reason: router upgrade finished, [[phab:T431752|T431752]]] * 13:07 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool ulsfo [reason: router upgrade finished, [[phab:T431752|T431752]]] * 13:06 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] * 13:03 XioNoX: repool cr4-ulsfo - [[phab:T431752|T431752]] * 12:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db1201: Maintenance * 12:48 gkyziridis@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1201 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95186 and previous config saved to /var/cache/conftool/dbconfig/20260727-124404-cwilliams.json * 12:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1201.eqiad.wmnet with reason: Maintenance * 12:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1187: Maintenance * 12:30 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader.eqiad.wikimedia.org on all recursors * 12:30 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader.eqiad.wikimedia.org on all recursors * 12:30 sukhe@dns1004: END - running authdns-update * 12:30 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] (duration: 09m 32s) * 12:28 sukhe@dns1004: START - running authdns-update * 12:25 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 12:22 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:20 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] * 12:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts * 12:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts * 12:13 XioNoX: rebooting cr4-ulsfo for upgrade - [[phab:T431752|T431752]] * 12:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: es1038 repool * 12:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 38 hosts * 12:08 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 38 hosts * 11:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1187: Maintenance * 11:53 urbanecm@deploy1003: mwscript-k8s job started: foreachwikiindblist growthexperiments GrowthExperiments:cleanMentorList # [[phab:T431804|T431804]] * 11:50 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr4-ulsfo,cr4-ulsfo IPv6,cr4-ulsfo.mgmt with reason: router upgrade * 11:50 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] (duration: 11m 07s) * 11:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1187 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95178 and previous config saved to /var/cache/conftool/dbconfig/20260727-114844-cwilliams.json * 11:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1187.eqiad.wmnet with reason: Maintenance * 11:43 urbanecm@deploy1003: urbanecm: Continuing with deployment * 11:42 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:39 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] * 11:37 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 11:36 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 11:36 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 11:35 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 11:29 XioNoX: start draining cr4-ulsfo - [[phab:T431752|T431752]] * 11:29 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 11:29 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 11:28 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1035: testing * 11:28 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1035: testing * 11:27 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1035: testing * 11:27 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1035: testing * 11:26 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool ulsfo [reason: router upgrade, [[phab:T431752|T431752]]] * 11:26 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1038: es1038 repool * 11:26 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool ulsfo [reason: router upgrade, [[phab:T431752|T431752]]] * 11:26 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1038: testing * 11:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1264: Maintenance * 11:24 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1038: testing * 11:23 marostegui@cumin1003: dbctl commit (dc=all): 'Repool es1050 as master', diff saved to https://phabricator.wikimedia.org/P95170 and previous config saved to /var/cache/conftool/dbconfig/20260727-112326-marostegui.json * 11:23 marostegui@cumin1003: dbctl commit (dc=all): 'Repool es1050', diff saved to https://phabricator.wikimedia.org/P95169 and previous config saved to /var/cache/conftool/dbconfig/20260727-112302-marostegui.json * 11:22 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1050: testing * 11:22 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1050: testing * 11:20 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 11:18 blake@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 11:18 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 11:12 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 11:11 blake@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 11:09 blake@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 11:09 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 11:09 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 11:08 blake@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 11:05 blake@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 11:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 11:02 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 10:50 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 10:43 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:39 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply * 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1264: Maintenance * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply * 10:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply * 10:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 10:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 10:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 10:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1264 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95164 and previous config saved to /var/cache/conftool/dbconfig/20260727-103204-cwilliams.json * 10:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1264.eqiad.wmnet with reason: Maintenance * 10:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 10:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 10:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 10:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 10:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 10:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 10:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 10:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1237: Maintenance * 10:04 elukey: restart burrow main-eqiad on kafkamon2003 to clear some errors on kafka-main1008 * 09:58 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1136.eqiad.wmnet with OS trixie * 09:39 elukey: restart burrow-main-eqiad.service on kafkamon1003 to see if a recurrent kafka error on kafka-main1008 goes away * 09:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1237: Maintenance * 09:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1136.eqiad.wmnet with reason: host reimage * 09:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1237 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95159 and previous config saved to /var/cache/conftool/dbconfig/20260727-093328-cwilliams.json * 09:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1237.eqiad.wmnet with reason: Maintenance * 09:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1136.eqiad.wmnet with reason: host reimage * 09:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1203: Maintenance * 09:17 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1136 * 09:17 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1136 * 09:04 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1136 * 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1136.eqiad.wmnet 191.32.64.10.in-addr.arpa 1.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:04 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1136.eqiad.wmnet 191.32.64.10.in-addr.arpa 1.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1136 - jiji@cumin1003" * 09:04 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1136 - jiji@cumin1003" * 08:52 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 08:52 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 08:52 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 08:51 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 08:50 jiji@cumin1003: START - Cookbook sre.dns.netbox * 08:47 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1136 * 08:46 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1136.eqiad.wmnet with OS trixie * 08:44 marostegui: Rename tables on s3 [[phab:T425066|T425066]] * 08:43 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1136.eqiad.wmnet * 08:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db1203: Maintenance * 08:43 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1136.eqiad.wmnet * 08:43 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1136.eqiad.wmnet * 08:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1203 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95154 and previous config saved to /var/cache/conftool/dbconfig/20260727-083703-cwilliams.json * 08:36 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1203.eqiad.wmnet with reason: Maintenance * 08:16 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1179: Maintenance * 07:44 phuedx: UTC morning backport window done * 07:37 phuedx@deploy1003: Finished scap sync-world: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] (duration: 32m 33s) * 07:28 root@cumin1003: START - Cookbook sre.mysql.pool pool db1179: Maintenance * 07:26 marostegui: Rename tables on s3 [[phab:T426341|T426341]] * 07:25 phuedx@deploy1003: phuedx: Continuing with deployment * 07:22 marostegui: Drop tables in akwiki nawiki pihwiki - growthexperiments_* [[phab:T428885|T428885]] * 07:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95149 and previous config saved to /var/cache/conftool/dbconfig/20260727-072234-cwilliams.json * 07:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1179.eqiad.wmnet with reason: Maintenance * 07:20 phuedx@deploy1003: phuedx: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:16 ryankemper: [[phab:T430880|T430880]] [WDQS] Reimaged `wdqs1018` and `wdqs1019` to Bookworm, restored data using test-cookbook change {{Gerrit|1317128}}, and repooled both; 25/36 hosts complete * 07:04 phuedx@deploy1003: Started scap sync-world: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] * 06:57 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1019.eqiad.wmnet * 06:56 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1018.eqiad.wmnet * 06:40 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1020.eqiad.wmnet with reason: Cloning * 06:35 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db1228.eqiad.wmnet with reason: Rebooting * 06:29 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:29 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:25 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:25 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:25 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1019.eqiad.wmnet, repooling source-only afterwards * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1018.eqiad.wmnet, repooling source-only afterwards * 04:51 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1019.eqiad.wmnet, repooling source-only afterwards * 04:51 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1018.eqiad.wmnet, repooling source-only afterwards * 04:48 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s) * 04:48 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 04:48 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s) * 04:48 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 36s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-26 == * 14:59 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:59 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:59 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:59 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1019.eqiad.wmnet with OS bookworm * 01:05 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1018.eqiad.wmnet with OS bookworm * 00:43 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1019.eqiad.wmnet with reason: host reimage * 00:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1018.eqiad.wmnet with reason: host reimage * 00:34 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1019.eqiad.wmnet with reason: host reimage * 00:33 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1018.eqiad.wmnet with reason: host reimage * 00:16 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 00:16 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 00:15 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 00:15 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1019 * 00:11 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1019 * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1018 * 00:11 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1018 * 00:08 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1019.eqiad.wmnet with OS bookworm * 00:08 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1018.eqiad.wmnet with OS bookworm == 2026-07-25 == * 22:06 ryankemper: [[phab:T430880|T430880]] [WDQS] Repooled `wdqs1017` and `wdqs2024` after reimaging to bookworm, scap deploying, and data xfering * 22:04 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2024.codfw.wmnet * 22:03 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1017.eqiad.wmnet * 21:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1017.eqiad.wmnet, repooling source-only afterwards * 21:06 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2024.codfw.wmnet, repooling source-only afterwards * 20:52 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:52 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:52 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:52 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 20:18 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1017.eqiad.wmnet, repooling source-only afterwards * 20:18 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2024.codfw.wmnet, repooling source-only afterwards * 20:15 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:15 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:15 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:15 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 19:57 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s) * 19:57 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 19:57 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 07s) * 19:57 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 19:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2024.codfw.wmnet with OS bookworm * 19:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1017.eqiad.wmnet with OS bookworm * 19:02 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2024.codfw.wmnet with reason: host reimage * 18:58 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1017.eqiad.wmnet with reason: host reimage * 18:53 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2024.codfw.wmnet with reason: host reimage * 18:52 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1017.eqiad.wmnet with reason: host reimage * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2024 * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2024 * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1017 * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1017 * 18:27 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2024 * 18:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2024.codfw.wmnet 58.16.192.10.in-addr.arpa 8.5.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:26 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2024.codfw.wmnet 58.16.192.10.in-addr.arpa 8.5.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:24 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1017 * 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1017.eqiad.wmnet 238.48.64.10.in-addr.arpa 8.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:24 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1017.eqiad.wmnet 238.48.64.10.in-addr.arpa 8.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1017 - ryankemper@cumin2003" * 18:24 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1017 - ryankemper@cumin2003" * 18:23 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 18:18 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 18:17 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1017 * 18:17 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2024 * 18:14 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1017.eqiad.wmnet with OS bookworm * 18:14 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2024.codfw.wmnet with OS bookworm * 18:05 ryankemper: [WDQS] [[phab:T430880|T430880]] Reimaged `wdqs1016` and `wdqs2023` to Bookworm with `--move-vlan`, restored main and scholarly data, validated postflights, and repooled both hosts. Confirmed PyBal rebuilt both backends with their new addresses * 17:45 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2023.codfw.wmnet * 17:43 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1016.eqiad.wmnet * 06:35 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1016.eqiad.wmnet, repooling source-only afterwards * 06:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2023.codfw.wmnet, repooling source-only afterwards * 05:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2023.codfw.wmnet, repooling source-only afterwards * 05:19 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1016.eqiad.wmnet, repooling source-only afterwards * 05:07 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 07s) * 05:07 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 05:06 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 06s) * 05:06 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 03:27 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2023.codfw.wmnet with OS bookworm * 02:59 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2023.codfw.wmnet with reason: host reimage * 02:56 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2023.codfw.wmnet with reason: host reimage * 02:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2023 * 02:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2023 * 02:30 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2023 * 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2023.codfw.wmnet 35.0.192.10.in-addr.arpa 5.3.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 02:30 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2023.codfw.wmnet 35.0.192.10.in-addr.arpa 5.3.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2023 - ryankemper@cumin2003" * 02:30 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2023 - ryankemper@cumin2003" * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 26s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:15 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1016.eqiad.wmnet with OS bookworm * 00:49 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1016.eqiad.wmnet with reason: host reimage * 00:43 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1016.eqiad.wmnet with reason: host reimage * 00:31 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 00:27 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1016 * 00:27 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1016 * 00:27 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2023 * 00:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1016.eqiad.wmnet with OS bookworm * 00:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2023.codfw.wmnet with OS bookworm * 00:11 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs1014.eqiad.wmnet and wdqs2008.codfw.wmnet after Bookworm reimage, transfer, and postflight; wdqs2008 is serving, while wdqs1014 will remain outside of service until a pybal restart next monday == 2026-07-24 == * 23:54 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1014.eqiad.wmnet * 23:54 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2008.codfw.wmnet * 23:43 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2010.codfw.wmnet with OS trixie * 23:08 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 23:03 jhathaway@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 22:33 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 22:13 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 22:13 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 22:13 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 22:13 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:00 jhathaway@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 21:53 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 21:53 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie * 21:51 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 21:47 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie * 21:43 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 21:39 jhathaway@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 21:38 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 17:21 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1135.eqiad.wmnet * 17:21 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1135.eqiad.wmnet * 17:21 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1135.eqiad.wmnet * 16:34 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 16:34 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 16:34 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 16:34 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 16:33 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 16:33 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 16:28 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:28 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:28 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:28 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2008.codfw.wmnet, repooling source-only afterwards * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1014.eqiad.wmnet, repooling source-only afterwards * 15:56 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1135.eqiad.wmnet with OS trixie * 15:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 40 hosts * 15:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 40 hosts * 15:37 topranks: upgrade SR-Linux OS on lswtest-d8-eqiad * 15:36 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1135.eqiad.wmnet with reason: host reimage * 15:33 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 6 hosts with reason: upgrade lswtest-d8-eqiad * 15:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1135.eqiad.wmnet with reason: host reimage * 15:30 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc-gp2006.codfw.wmnet with OS bookworm * 15:15 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1135 * 15:15 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1135 * 15:13 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc-gp2006.codfw.wmnet with reason: host reimage * 15:08 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc-gp2006.codfw.wmnet with reason: host reimage * 14:49 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm * 14:48 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host mc-gp2006.codfw.wmnet with OS bookworm * 14:34 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] (duration: 41m 12s) * 14:32 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1135 * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1135.eqiad.wmnet 177.32.64.10.in-addr.arpa 7.7.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:32 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1135.eqiad.wmnet 177.32.64.10.in-addr.arpa 7.7.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1135 - jiji@cumin1003" * 14:32 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1135 - jiji@cumin1003" * 14:29 krinkle@deploy1003: krinkle: Continuing with deployment * 14:29 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm * 14:27 jiji@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host mc-gp2006.codfw.wmnet with OS bookworm * 14:26 jiji@cumin1003: START - Cookbook sre.dns.netbox * 14:15 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1135 * 14:14 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1135.eqiad.wmnet with OS trixie * 14:14 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1135.eqiad.wmnet * 14:13 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1135.eqiad.wmnet * 14:13 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1135.eqiad.wmnet * 14:10 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1072.eqiad.wmnet * 14:10 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1072.eqiad.wmnet * 14:10 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1072.eqiad.wmnet * 14:10 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1072.eqiad.wmnet * 14:09 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1071.eqiad.wmnet * 14:09 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1071.eqiad.wmnet * 14:09 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1071.eqiad.wmnet * 14:09 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1071.eqiad.wmnet * 13:58 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 13:58 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:58 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:57 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:55 krinkle@deploy1003: krinkle: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:53 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] * 13:45 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:45 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push new IPs for mc-gp2006 - cmooney@cumin1003" * 13:45 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push new IPs for mc-gp2006 - cmooney@cumin1003" * 13:44 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) mc-gp2006.codfw.wmnet on all recursors * 13:44 cmooney@cumin1003: START - Cookbook sre.dns.wipe-cache mc-gp2006.codfw.wmnet on all recursors * 13:42 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm * 13:41 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:30 papaul: reboot mr1-eqsin for maintenance * 13:24 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb[1029-1031].eqiad.wmnet * 13:10 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb[1029-1031].eqiad.wmnet * 11:33 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 7 hosts * 11:11 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 7 hosts * 10:56 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 7 hosts * 10:47 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 7 hosts * 10:44 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:44 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:41 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:41 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:35 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 8 hosts * 10:34 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:33 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:32 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:32 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:31 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:31 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:30 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 8 hosts * 10:24 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie * 10:19 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:18 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 16 hosts * 10:17 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2001.codfw.wmnet * 10:13 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2001.codfw.wmnet * 10:12 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2001.codfw.wmnet * 10:02 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2001.codfw.wmnet * 10:02 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2002.codfw.wmnet * 09:57 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2002.codfw.wmnet * 09:56 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2002.codfw.wmnet * 09:51 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2002.codfw.wmnet * 09:51 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1002.eqiad.wmnet * 09:47 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1002.eqiad.wmnet * 09:47 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1001.eqiad.wmnet * 09:44 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1001.eqiad.wmnet * 09:34 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2003.codfw.wmnet * 09:32 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2003.codfw.wmnet * 09:32 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2002.codfw.wmnet * 09:30 brouberol@dns1004: END - running authdns-update * 09:29 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2002.codfw.wmnet * 09:29 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2001.codfw.wmnet * 09:27 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 9 hosts * 09:27 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2001.codfw.wmnet * 09:27 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2001.codfw.wmnet * 09:26 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 9 hosts * 09:26 brouberol@dns1004: START - running authdns-update * 09:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 57 hosts * 09:24 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2001.codfw.wmnet * 09:24 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2002.codfw.wmnet * 09:22 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2002.codfw.wmnet * 09:21 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 57 hosts * 09:20 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2003.codfw.wmnet * 09:19 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 16 hosts * 09:16 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2003.codfw.wmnet * 09:16 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1003.eqiad.wmnet * 09:15 urbanecm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 09:15 urbanecm@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 09:13 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1003.eqiad.wmnet * 09:13 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1002.eqiad.wmnet * 09:11 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1002.eqiad.wmnet * 09:11 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1001.eqiad.wmnet * 09:07 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1001.eqiad.wmnet * 08:32 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:24 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:16 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 08:16 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 08:07 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:07 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:07 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 08:02 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 08:01 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:59 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:57 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 07:57 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 06:46 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1025.eqiad.wmnet with reason: Cloning * 06:46 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s4 * 06:45 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s6 * 06:44 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1019.eqiad.wmnet,service=s6 * 06:44 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1019.eqiad.wmnet,service=s4 * 03:40 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:40 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:40 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:40 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 03:37 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:37 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:37 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:36 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:49 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on mr1-eqsin,mr1-eqsin IPv6 with reason: connection issue * 02:38 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on cr[2-3]-eqsin.mgmt,ps1-[603-604]-eqsin with reason: connection issue * 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 27s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-23 == * 23:27 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin.oob,mr1-eqsin.oob IPv6 with reason: switch refresh * 22:21 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Setting storage compatibility to NONE - eevans@cumin1003 * 22:01 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Setting storage compatibility to NONE - eevans@cumin1003 * 21:29 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1014.eqiad.wmnet, repooling source-only afterwards * 21:28 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 46s) * 21:28 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 21:19 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Setting storage compatibility to UPGRADING - eevans@cumin1003 * 21:00 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Setting storage compatibility to UPGRADING - eevans@cumin1003 * 20:17 dani@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] (duration: 11m 57s) * 20:13 dani@deploy1003: dani: Continuing with deployment * 20:07 dani@deploy1003: dani: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:05 dani@deploy1003: Started scap sync-world: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] * 19:24 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:24 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating the rest of the ipv6 dns records. - jhancock@cumin2002" * 19:24 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating the rest of the ipv6 dns records. - jhancock@cumin2002" * 19:14 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 19:05 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wdqs1014.eqiad.wmnet with OS bookworm * 19:04 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.noop (exit_code=99) * 19:04 cwilliams@cumin1003: START - Cookbook sre.mysql.noop * 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1014.eqiad.wmnet with reason: host reimage * 18:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1014.eqiad.wmnet with reason: host reimage * 18:30 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2008.codfw.wmnet, repooling source-only afterwards * 18:28 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 19s) * 18:28 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1014 * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1014 * 18:22 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1014 * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1014.eqiad.wmnet 188.32.64.10.in-addr.arpa 8.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:22 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1014.eqiad.wmnet 188.32.64.10.in-addr.arpa 8.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1014 - bking@cumin2003" * 18:21 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1014 - bking@cumin2003" * 18:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2215: Maintenance * 18:18 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 18:15 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:15 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 18:06 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:06 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 18:05 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 18:04 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2052: codfw rack B8 re-pool after maintenance * 17:54 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 17:54 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:54 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 17:32 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2215: Maintenance * 17:29 cmooney@dns3003: END - running authdns-update * 17:27 cmooney@dns3003: START - running authdns-update * 17:23 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 17:22 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:18 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool es2052: codfw rack B8 re-pool after maintenance * 17:18 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2189: codfw rack B8 re-pool after maintenance * 17:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2215.codfw.wmnet with reason: Maintenance * 17:17 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 17:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2215 [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95126 and previous config saved to /var/cache/conftool/dbconfig/20260723-170903-cwilliams.json * 17:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2191 to x1 primary [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95125 and previous config saved to /var/cache/conftool/dbconfig/20260723-170612-cwilliams.json * 17:05 cezmunsta: Starting x1 codfw failover from db2215 to db2191 - [[phab:T432986|T432986]] * 16:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2191 with weight 0 [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95123 and previous config saved to /var/cache/conftool/dbconfig/20260723-165831-cwilliams.json * 16:58 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 16 hosts with reason: Primary switchover x1 [[phab:T432986|T432986]] * 16:36 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 138128 * 16:35 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 138128 * 16:33 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2189: codfw rack B8 re-pool after maintenance * 16:33 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2164: codfw rack B8 re-pool after maintenance * 16:28 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1072.eqiad.wmnet * 16:27 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1072.eqiad.wmnet with OS trixie * 16:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2249: Maintenance * 16:06 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker1072.eqiad.wmnet with reason: host reimage * 16:06 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1072.eqiad.wmnet with reason: host reimage * 15:50 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1072 * 15:50 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1072 * 15:49 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1072.eqiad.wmnet with OS trixie * 15:48 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2164: codfw rack B8 re-pool after maintenance * 15:48 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] (duration: 06m 37s) * 15:48 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2163: codfw rack B8 re-pool after maintenance * 15:45 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 15:43 musikanimal@deploy1003: musikanimal: Continuing with deployment * 15:43 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:41 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] * 15:36 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1072.eqiad.wmnet * 15:35 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1072.eqiad.wmnet * 15:35 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1072.eqiad.wmnet * 15:34 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:34 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push any outstanding updates - cmooney@cumin1003" * 15:34 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push any outstanding updates - cmooney@cumin1003" * 15:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db2249: Maintenance * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 15:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:26 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:21 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:21 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 15:21 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:21 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 15:20 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 15:19 cmooney@dns2004: END - running authdns-update * 15:17 cmooney@dns2004: START - running authdns-update * 15:14 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns2004.wikimedia.org * 15:12 brouberol@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 15:12 brouberol@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 15:12 klausman@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ml-serve1001.eqiad.wmnet with OS trixie * 15:11 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1071.eqiad.wmnet * 15:11 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1071.eqiad.wmnet with OS trixie * 15:10 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wdqs2008.codfw.wmnet with OS bookworm * 15:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2249.codfw.wmnet with reason: Maintenance * 15:08 brouberol@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 15:08 brouberol@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 15:08 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 15:08 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 15:06 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns1004.wikimedia.org * 15:02 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2002.codfw.wmnet * 15:02 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2002.codfw.wmnet * 15:02 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2163: codfw rack B8 re-pool after maintenance * 15:01 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 15:01 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 14:59 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test2001.codfw.wmnet * 14:57 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test2001.codfw.wmnet * 14:56 ryankemper: [WDQS] [[phab:T430880|T430880]] Reimaged `wdqs2016` to Bookworm, xferred scholarly_articles from `wdqs2024`, validated updater/readiness/federation, and repooled * 14:51 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2016.codfw.wmnet * 14:51 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1001.eqiad.wmnet with reason: host reimage * 14:48 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1071.eqiad.wmnet with reason: host reimage * 14:47 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1001.eqiad.wmnet with reason: host reimage * 14:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2008.codfw.wmnet with reason: host reimage * 14:43 topranks: reboot lsw1-b8-codw to upgrade JunOS [[phab:T430929|T430929]] * 14:41 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1071.eqiad.wmnet with reason: host reimage * 14:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2008.codfw.wmnet with reason: host reimage * 14:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2231: Maintenance * 14:30 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1001.eqiad.wmnet with OS trixie * 14:25 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2002.codfw.wmnet * 14:23 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1071 * 14:23 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1071 * 14:23 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 14:22 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore scholarly data after Bookworm reimage) xfer scholarly_articles from wdqs2024.codfw.wmnet -> wdqs2016.codfw.wmnet, repooling source-only afterwards * 14:22 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2052: codfw rack B8 depool for maintenance * 14:21 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool es2052: codfw rack B8 depool for maintenance * 14:21 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2249: codfw rack B8 depool for maintenance * 14:21 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1071 * 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1071.eqiad.wmnet 166.48.64.10.in-addr.arpa 6.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:21 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1071.eqiad.wmnet 166.48.64.10.in-addr.arpa 6.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1071 - jiji@cumin1003" * 14:21 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1071 - jiji@cumin1003" * 14:21 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2249: codfw rack B8 depool for maintenance * 14:21 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2189: codfw rack B8 depool for maintenance * 14:20 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2002.codfw.wmnet * 14:20 cmooney@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2050.codfw.wmnet * 14:20 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2189: codfw rack B8 depool for maintenance * 14:20 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2164: codfw rack B8 depool for maintenance * 14:20 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2164: codfw rack B8 depool for maintenance * 14:19 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2163: codfw rack B8 depool for maintenance * 14:19 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:19 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1014 * 14:19 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2163: codfw rack B8 depool for maintenance * 14:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1014.eqiad.wmnet with OS bookworm * 14:17 cmooney@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2050.codfw.wmnet * 14:16 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 14:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2008 * 14:14 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2008 * 14:14 cmooney@cumin1003: conftool action : set/pooled=no; selector: name=dns2004.wikimedia.org * 14:14 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2008.codfw.wmnet with OS bookworm * 14:13 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 14:12 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:10 topranks: depool dns2004 before lsw1-b8-codfw switch maintenance [[phab:T430929|T430929]] * 14:10 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b8-codfw,lsw1-b8-codfw IPv6,lsw1-b8-codfw.mgmt,ssw1-a[1,8]-codfw with reason: lsw1-b8-codfw JunOS upgrade * 14:07 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 30 hosts with reason: lsw1-b8-codfw JunOS upgrade * 14:06 elukey: upload python3-docker-report 0.0.19 to apt.wikimedia.org for bookworm and trixie * 13:59 jiji@cumin1003: START - Cookbook sre.dns.netbox * 13:58 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1071 * 13:57 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:57 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:57 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:57 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1071.eqiad.wmnet with OS trixie * 13:55 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1071.eqiad.wmnet * 13:55 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1071.eqiad.wmnet * 13:55 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1071.eqiad.wmnet * 13:53 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:52 logmsgbot: kharlan Deployed security patch for [[phab:T432948|T432948]] * 13:51 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:51 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 13:51 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db2231: Maintenance * 13:50 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:50 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:50 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:50 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:49 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:49 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:49 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2231 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95097 and previous config saved to /var/cache/conftool/dbconfig/20260723-134436-cwilliams.json * 13:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2231.codfw.wmnet with reason: Maintenance * 13:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:39 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:38 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 13:38 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] (duration: 09m 07s) * 13:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1037 hosts * 13:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2196: Maintenance * 13:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:33 kharlan@deploy1003: kharlan, emc-wmf: Continuing with deployment * 13:31 kharlan@deploy1003: kharlan, emc-wmf: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:30 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:28 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] * 13:17 hashar@deploy1003: Finished deploy [integration/docroot@2199146]: build: License GPL2.0+ / updating npm dependencies (duration: 00m 14s) * 13:17 hashar@deploy1003: Started deploy [integration/docroot@2199146]: build: License GPL2.0+ / updating npm dependencies * 13:14 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service * 13:07 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 12:58 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2207: Repooling * 12:49 root@cumin1003: START - Cookbook sre.mysql.pool pool db2196: Maintenance * 12:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2196 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95087 and previous config saved to /var/cache/conftool/dbconfig/20260723-123952-cwilliams.json * 12:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2196.codfw.wmnet with reason: Maintenance * 12:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2191: Maintenance * 12:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:13 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:13 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: Repooling * 12:12 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2207: Repooling * 12:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: Repooling * 11:56 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2235.codfw.wmnet with OS trixie * 11:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db2191: Maintenance * 11:46 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1070.eqiad.wmnet * 11:46 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1070.eqiad.wmnet * 11:46 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1070.eqiad.wmnet * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2191 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95080 and previous config saved to /var/cache/conftool/dbconfig/20260723-114308-cwilliams.json * 11:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2191.codfw.wmnet with reason: Maintenance * 11:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2186: Maintenance * 11:35 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 46375 * 11:34 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 46375 * 11:33 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2235.codfw.wmnet with reason: host reimage * 11:28 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2235.codfw.wmnet with reason: host reimage * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c7-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c7-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c6-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c6-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c5-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c5-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c4-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c4-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c3-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c3-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c2-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c2-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d7-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d7-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d4-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d4-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d3-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d2-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d2-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d8-eqiad * 11:23 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d8-eqiad * 11:23 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d1-eqiad * 11:23 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d1-eqiad * 11:12 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db2235.codfw.wmnet with OS trixie * 11:11 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:11 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[2160,2235].codfw.wmnet with reason: Upgrading * 10:56 root@cumin1003: START - Cookbook sre.mysql.pool pool db2186: Maintenance * 10:54 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1037: testing * 10:53 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1037: testing * 10:53 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1037: testing * 10:53 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1037: testing * 10:52 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: testing * 10:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2186 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95072 and previous config saved to /var/cache/conftool/dbconfig/20260723-104956-cwilliams.json * 10:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2186.codfw.wmnet with reason: Maintenance * 10:43 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1054: testing * 10:41 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1070.eqiad.wmnet with OS trixie * 10:30 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1037 hosts * 10:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 10:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 10:20 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1070.eqiad.wmnet with reason: host reimage * 10:16 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1070.eqiad.wmnet with reason: host reimage * 10:06 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1038: testing * 10:05 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1038: testing * 10:05 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1038: testing * 10:02 marostegui@dns1004: END - running authdns-update * 10:00 marostegui@dns1004: START - running authdns-update * 09:58 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: testing * 09:57 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1054: testing * 09:57 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1070 * 09:57 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1070 * 09:57 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1054: testing * 09:57 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1054: testing * 09:56 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1055: testing * 09:56 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1070 * 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1070.eqiad.wmnet 165.48.64.10.in-addr.arpa 5.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:56 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1070.eqiad.wmnet 165.48.64.10.in-addr.arpa 5.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1070 - jiji@cumin1003" * 09:56 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1070 - jiji@cumin1003" * 09:47 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for 1035 hosts * 09:45 jiji@cumin1003: START - Cookbook sre.dns.netbox * 09:42 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1070 * 09:42 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1070.eqiad.wmnet with OS trixie * 09:42 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1070.eqiad.wmnet * 09:41 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1070.eqiad.wmnet * 09:41 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1070.eqiad.wmnet * 09:27 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es2051: testing * 09:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: testing * 09:12 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es2051: testing * 09:11 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1055: testing * 09:09 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:09 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1055: testing * 09:09 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1055: testing * 08:50 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:50 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:50 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 08:49 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 08:49 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 08:49 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:46 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1069.eqiad.wmnet * 08:46 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1069.eqiad.wmnet * 08:46 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1069.eqiad.wmnet * 08:39 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 08:38 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 08:38 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 08:37 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 08:35 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:10 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1069.eqiad.wmnet with OS trixie * 07:49 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1069.eqiad.wmnet with reason: host reimage * 07:45 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1069.eqiad.wmnet with reason: host reimage * 07:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1035 hosts * 07:33 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 07:32 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 07:29 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1069 * 07:29 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1069 * 07:26 jiji@deploy1003: Finished scap sync-world: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules (duration: 06m 01s) * 07:25 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1069 * 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1069.eqiad.wmnet 164.48.64.10.in-addr.arpa 4.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:25 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1069.eqiad.wmnet 164.48.64.10.in-addr.arpa 4.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1069 - jiji@cumin1003" * 07:25 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1069 - jiji@cumin1003" * 07:25 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 07:25 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 07:24 jiji@deploy1003: jiji: Continuing with deployment * 07:22 jiji@deploy1003: jiji: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:21 jiji@deploy1003: Started scap sync-world: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules * 07:21 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1031.eqiad.wmnet,service=s7 * 07:20 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1031.eqiad.wmnet,service=s2 * 07:20 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1031.eqiad.wmnet,service=s7 * 07:20 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1031.eqiad.wmnet,service=s2 * 07:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts * 07:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts * 07:19 jiji@cumin1003: START - Cookbook sre.dns.netbox * 07:19 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1069 * 07:19 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1069.eqiad.wmnet with OS trixie * 07:19 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1069.eqiad.wmnet * 07:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 45 hosts * 07:17 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1069.eqiad.wmnet * 07:17 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1069.eqiad.wmnet * 07:14 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 45 hosts * 07:13 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 06:16 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs2007 after successful Bookworm reimage, data transfer, and postflight validation; wdqs1013 also passed postflights and is enabled in conftool, but remains out of IPVS pending a rolling pybal restart to clear its stale pre-VLAN-move address * 05:58 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2007.codfw.wmnet * 05:58 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1013.eqiad.wmnet * 05:54 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore scholarly data after Bookworm reimage) xfer scholarly_articles from wdqs2024.codfw.wmnet -> wdqs2016.codfw.wmnet, repooling source-only afterwards == 2026-07-22 == * 23:34 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Apply upgrade to JVM17 - eevans@cumin1003 * 23:14 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Apply upgrade to JVM17 - eevans@cumin1003 * 22:06 ryankemper: [WDQS] Added requestctl per-IP ratelimit `wdqs_heavy_sparql_bots_jul_2026_ratelimit` (chronic heavy-query bot tier driving deadlock-remediation restarts); pruned superseded `wdqs_2026_05_11_worobot` * 21:51 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] (duration: 11m 52s) * 21:44 sbassett@deploy1003: sbassett: Continuing with deployment * 21:43 sbassett@deploy1003: sbassett: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:39 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] * 20:38 dani@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] (duration: 32m 51s) * 20:38 ryankemper: [WDQS] Pruned obsolete requestctl action+pattern `wdqs_20260715_p2003_ring_ja3n` (actor rotated JA3Ns; rule inert) * 20:26 dani@deploy1003: dani, vadymts1: Continuing with deployment * 20:24 dani@deploy1003: dani, vadymts1: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:14 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2244: Testing * 20:06 dani@deploy1003: Started scap sync-world: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] * 19:56 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1013.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2016.codfw.wmnet with OS bookworm * 19:40 mutante: gerrit - one more service restart is needed - restarting * 19:29 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2244: Testing * 19:27 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2244: Testing * 19:27 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2244: Testing * 19:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2016.codfw.wmnet with reason: host reimage * 19:15 dancy@deploy1003: Finished deploy [zuul/deploy@d92e238]: Freshening Zuul installation (duration: 00m 15s) * 19:14 dancy@deploy1003: Started deploy [zuul/deploy@d92e238]: Freshening Zuul installation * 19:11 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2016.codfw.wmnet with reason: host reimage * 18:54 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1013.eqiad.wmnet, repooling source-only afterwards * 18:52 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2016 * 18:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2016 * 18:51 dancy@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 18:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2016.codfw.wmnet with OS bookworm * 18:39 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 11s) * 18:39 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 18:36 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 18:30 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417] (thin): Regular analytics weekly train THIN [analytics/refinery@2a25417d] (duration: 02m 09s) * 18:28 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417] (thin): Regular analytics weekly train THIN [analytics/refinery@2a25417d] * 18:28 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417]: Regular analytics weekly train [analytics/refinery@2a25417d] (duration: 04m 31s) * 18:27 dduvall: deploying https://gerrit.wikimedia.org/r/c/integration/config/+/1314025 (4 jobs updated) * 18:23 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417]: Regular analytics weekly train [analytics/refinery@2a25417d] * 18:22 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@2a25417d] (duration: 01m 59s) * 18:20 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@2a25417d] * 17:56 Raine: deployment server switchover => deploy1003 is primary now * 17:55 kamila@deploy1003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 28m 24s) * 17:54 mutante: restarting gerrit for maintenance * 17:29 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1013.eqiad.wmnet with OS bookworm * 17:27 kamila@deploy1003: Started scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] * 17:20 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] (duration: 22m 50s) * 17:12 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1023.eqiad.wmnet -> wdqs1024.eqiad.wmnet, repooling source-only afterwards * 17:04 Raine: point deployment.eqiad.wmnet to deploy1003 * 17:04 kamila@dns7001: END - running authdns-update * 17:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1013.eqiad.wmnet with reason: host reimage * 17:02 kamila@dns7001: START - running authdns-update * 17:01 kamila@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on releases2003.codfw.wmnet,releases1003.eqiad.wmnet with reason: Deployment server switchover * 17:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1013.eqiad.wmnet with reason: host reimage * 16:58 kamila@deploy2003: Locking from deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] * 16:57 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] (duration: 02m 33s) * 16:55 kamila@deploy2003: Locking from deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] * 16:55 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2003 - [[phab:T240266|T240266]] (duration: 00m 11s) * 16:54 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2003 - [[phab:T240266|T240266]] * 16:40 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1013 * 16:40 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1013 * 16:39 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1013 * 16:39 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1013.eqiad.wmnet 105.32.64.10.in-addr.arpa 5.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:39 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1013.eqiad.wmnet 105.32.64.10.in-addr.arpa 5.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:39 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:39 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1013 - bking@cumin2003" * 16:39 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1013 - bking@cumin2003" * 16:34 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:34 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1013 * 16:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1013.eqiad.wmnet with OS bookworm * 16:28 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1023.eqiad.wmnet -> wdqs1024.eqiad.wmnet, repooling source-only afterwards * 16:27 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-scholarly,name=eqiad * 16:27 eevans@deploy2003: helmfile [eqiad] DONE helmfile.d/services/linked-artifacts: apply * 16:26 eevans@deploy2003: helmfile [eqiad] START helmfile.d/services/linked-artifacts: apply * 16:26 eevans@deploy2003: helmfile [codfw] DONE helmfile.d/services/linked-artifacts: apply * 16:26 eevans@deploy2003: helmfile [codfw] START helmfile.d/services/linked-artifacts: apply * 16:25 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 29s) * 16:25 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 16:24 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 16:21 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 16:18 eevans@deploy2003: helmfile [codfw] DONE helmfile.d/services/linked-artifacts: apply * 16:18 eevans@deploy2003: helmfile [codfw] START helmfile.d/services/linked-artifacts: apply * 16:08 eevans@deploy2003: helmfile [staging] DONE helmfile.d/services/linked-artifacts: apply * 16:07 eevans@deploy2003: helmfile [staging] START helmfile.d/services/linked-artifacts: apply * 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 16:01 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 15:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1024.eqiad.wmnet with OS bookworm * 15:49 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] (duration: 00m 10s) * 15:49 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] * 15:48 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] (duration: 00m 15s) * 15:48 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] * 15:47 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] (duration: 00m 10s) * 15:47 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] * 15:46 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:42 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:42 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:40 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:37 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 15:37 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:36 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:36 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:36 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 15:33 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 15:33 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:31 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:28 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 15:27 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:27 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1024.eqiad.wmnet with reason: host reimage * 15:23 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:23 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:23 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:20 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1068.eqiad.wmnet * 15:20 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1068.eqiad.wmnet * 15:20 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1068.eqiad.wmnet * 15:20 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 15:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1024.eqiad.wmnet with reason: host reimage * 15:11 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:55 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 14:52 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wdqs1024.eqiad.wmnet with OS bookworm * 14:50 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] (duration: 00m 09s) * 14:50 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] * 14:49 jiji@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 14:49 jiji@deploy2003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 14:49 jiji@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 14:48 jiji@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 14:45 ecarg@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:45 ecarg@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:44 ecarg@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:44 ecarg@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:43 ecarg@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:43 ecarg@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:41 sukhe: ipvsadm --delete-service --tcp-service 10.2.1.55:8087: lvs2014 and lvs2013 * 14:39 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:39 sukhe: ipvsadm --delete-service --tcp-service 10.2.2.55:8087: [[phab:T432445|T432445]] * 14:38 ecarg@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:38 ecarg@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:37 ecarg@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:37 ecarg@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:36 ecarg@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:34 ecarg@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts datahubsearch1001.eqiad.wmnet * 14:32 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:32 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 14:31 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 14:31 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 14:28 sukhe: sudo cumin 'A:lvs-low-traffic-codfw' 'systemctl restart pybal': lvs2013 * 14:26 sukhe: sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal': lvs2014 * 14:26 sukhe: sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal' * 14:24 sukhe: restart pybal on lvs1019 * 14:24 sukhe: restart pybal on lvs1020 * 14:19 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for 1036 hosts * 14:17 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:04 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] (duration: 09m 28s) * 13:59 kharlan@deploy2003: dreamyjazz, kharlan: Continuing with deployment * 13:58 bking@cumin2003: START - Cookbook sre.hosts.decommission for hosts datahubsearch1001.eqiad.wmnet * 13:57 kharlan@deploy2003: dreamyjazz, kharlan: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:55 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] * 13:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts datahubsearch[1002-1003].eqiad.wmnet * 13:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:53 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch[1002-1003].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 13:52 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch[1002-1003].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 13:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 13:42 stran@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] (duration: 07m 30s) * 13:42 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:40 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: Test * 13:38 stran@deploy2003: dragoniez, stran: Continuing with deployment * 13:37 stran@deploy2003: dragoniez, stran: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:35 bking@cumin2003: START - Cookbook sre.hosts.decommission for hosts datahubsearch[1002-1003].eqiad.wmnet * 13:35 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024'] * 13:35 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 13:35 stran@deploy2003: Started scap sync-world: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] * 13:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 13:28 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024'] * 13:26 sukhe@dns1004: END - running authdns-update * 13:25 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 13:25 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 13:24 sukhe@dns1004: START - running authdns-update * 13:22 stran@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] (duration: 08m 20s) * 13:21 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 13:20 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 13:19 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 13:19 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 13:19 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 13:18 stran@deploy2003: stran: Continuing with deployment * 13:16 stran@deploy2003: stran: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:14 stran@deploy2003: Started scap sync-world: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] * 13:13 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 13:13 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 13:11 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 13:11 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 13:08 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 12:55 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool es2051: Test * 12:55 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: Test * 12:54 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool es2051: Test * 12:43 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1036 hosts * 12:41 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 12:40 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1048.eqiad.wmnet * 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 12:39 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 12:38 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 12:37 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 12:37 brouberol@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 12:36 brouberol@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 12:36 brouberol@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 12:36 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 12:36 elukey@cumin1003: DONE (PASS) - Cookbook sre.puppet.renew-cert (exit_code=0) for crm2001.codfw.wmnet: Renew puppet certificate - elukey@cumin1003 * 12:35 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:35 brouberol@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 12:34 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 12:31 brouberol@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 12:30 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1048.eqiad.wmnet * 12:30 brouberol@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 12:28 brouberol@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 12:27 brouberol@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 12:20 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1068.eqiad.wmnet with OS trixie * 12:01 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] (duration: 13m 19s) * 11:58 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1068.eqiad.wmnet with reason: host reimage * 11:52 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1068.eqiad.wmnet with reason: host reimage * 11:51 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 11:49 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:47 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] * 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2252: Security updates * 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:43 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 11:42 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2252: Security updates * 11:42 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply * 11:40 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply * 11:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 11:37 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2252.codfw.wmnet with OS trixie * 11:34 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1068 * 11:34 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1068 * 11:26 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1068 * 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1068.eqiad.wmnet 46.48.64.10.in-addr.arpa 6.4.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:26 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1068.eqiad.wmnet 46.48.64.10.in-addr.arpa 6.4.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1068 - jiji@cumin1003" * 11:26 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1068 - jiji@cumin1003" * 11:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2252.codfw.wmnet with reason: host reimage * 11:17 jiji@cumin1003: START - Cookbook sre.dns.netbox * 11:17 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1068 * 11:17 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1068.eqiad.wmnet with OS trixie * 11:17 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2252.codfw.wmnet with reason: host reimage * 11:15 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1068.eqiad.wmnet * 11:15 mvolz@deploy2003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:15 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1068.eqiad.wmnet * 11:15 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1068.eqiad.wmnet * 11:14 mvolz@deploy2003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:13 mvolz@deploy2003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:13 mvolz@deploy2003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:12 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] (duration: 11m 05s) * 11:11 mvolz@deploy2003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:10 mvolz@deploy2003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:07 dreamyjazz@deploy2003: dreamyjazz, kharlan: Continuing with deployment * 11:03 dreamyjazz@deploy2003: dreamyjazz, kharlan: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:03 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2252.codfw.wmnet with OS trixie * 11:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2252: Upgrading db2252.codfw.wmnet * 11:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:02 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 11:02 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2252: Upgrading db2252.codfw.wmnet * 11:02 cwilliams@cumin1003: dbmaint on ms3@codfw [[phab:T432321|T432321]] * 11:01 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 11:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db1153.eqiad.wmnet with reason: Security updates * 11:01 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] * 11:00 fnegri@deploy2003: helmfile [eqiad] DONE helmfile.d/services/toolhub: apply * 10:58 fnegri@deploy2003: helmfile [eqiad] START helmfile.d/services/toolhub: apply * 10:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1151: Security updates * 10:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:57 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1151: Security updates * 10:55 fnegri@deploy2003: helmfile [codfw] DONE helmfile.d/services/toolhub: apply * 10:54 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] (duration: 08m 38s) * 10:53 fnegri@deploy2003: helmfile [codfw] START helmfile.d/services/toolhub: apply * 10:53 fnegri@deploy2003: helmfile [staging] DONE helmfile.d/services/toolhub: apply * 10:52 fnegri@deploy2003: helmfile [staging] START helmfile.d/services/toolhub: apply * 10:50 zabe@deploy2003: zabe: Continuing with deployment * 10:47 zabe@deploy2003: zabe: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:45 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] * 10:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1151: Security updates * 10:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:42 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:42 root@cumin1003: START - Cookbook sre.mysql.depool depool db1151: Security updates * 10:38 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] (duration: 12m 47s) * 10:34 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 10:34 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 10:33 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2253.codfw.wmnet with OS trixie * 10:28 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:26 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] * 10:18 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2253.codfw.wmnet with reason: host reimage * 10:13 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2253.codfw.wmnet with reason: host reimage * 10:00 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2253.codfw.wmnet with OS trixie * 09:58 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db1151.eqiad.wmnet with reason: Security updates * 09:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2253: Upgrading db2253.codfw.wmnet * 09:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:57 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 09:56 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2253: Upgrading db2253.codfw.wmnet * 09:56 cwilliams@cumin1003: dbmaint on ms2@codfw [[phab:T432321|T432321]] * 09:56 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 09:36 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: UI improvement; support url shortener - oblivian@cumin1003" * 09:36 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: UI improvement; support url shortener - oblivian@cumin1003 * 09:35 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: UI improvement; support url shortener - oblivian@cumin1003 * 09:35 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: UI improvement; support url shortener - oblivian@cumin1003" * 09:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1152: Security updates * 09:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:26 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db1152: Security updates * 09:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: Security updates * 09:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:11 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:11 root@cumin1003: START - Cookbook sre.mysql.depool depool db1152: Security updates * 09:10 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1018.eqiad.wmnet with reason: Cloning * 09:09 Dreamy_Jazz: Deployed patch for [[phab:T432453|T432453]] and [[phab:T432454|T432454]] * 09:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 09:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2251.codfw.wmnet with OS trixie * 08:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2251.codfw.wmnet with reason: host reimage * 08:45 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2251.codfw.wmnet with reason: host reimage * 08:40 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] (duration: 12m 26s) * 08:38 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1030.eqiad.wmnet,service=s1 * 08:36 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 08:31 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2251.codfw.wmnet with OS trixie * 08:30 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:28 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] * 08:25 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] (duration: 07m 59s) * 08:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2251: Upgrading db2251.codfw.wmnet * 08:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:22 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 08:22 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2251: Upgrading db2251.codfw.wmnet * 08:20 urbanecm@deploy2003: urbanecm: Continuing with deployment * 08:20 cwilliams@cumin1003: dbmaint on ms1@codfw [[phab:T432321|T432321]] * 08:20 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 08:19 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:17 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] * 08:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade * 08:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade * 08:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db2251.codfw.wmnet,db1152.eqiad.wmnet with reason: OS upgrade * 08:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade * 08:13 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade * 08:11 Dreamy_Jazz: Created cusi_signal, cusi_case, and cusi_user on ukwiki and enwikivoyage in extension1 * 08:11 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade * 08:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade * 08:04 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1030.eqiad.wmnet,service=s1 * 08:04 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1030.eqiad.wmnet,service=s1 * 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply * 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply * 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply * 07:51 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply * 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 07:47 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 07:47 phuedx: End of UTC morning backport window * 07:43 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 07:43 phuedx@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] (duration: 13m 44s) * 07:43 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 07:39 phuedx@deploy2003: phuedx: Continuing with deployment * 07:31 phuedx@deploy2003: phuedx: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:29 phuedx@deploy2003: Started scap sync-world: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] * 07:24 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Turnilo import support - oblivian@cumin1003" * 07:24 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import support - oblivian@cumin1003 * 07:23 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import support - oblivian@cumin1003 * 07:23 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Turnilo import support - oblivian@cumin1003" * 06:42 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs2020 after successful Bookworm reimage, data transfer, and postflight validation * 06:42 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2020.codfw.wmnet * 05:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Managing sanitization for wikis bolwiki in section s5 * 05:25 marostegui@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis bolwiki in section s5 * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 41s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-21 == * 22:50 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2019.codfw.wmnet -> wdqs2020.codfw.wmnet, repooling source-only afterwards * 22:47 cwhite: force reboot arclamp2001 - appears to have run out of memory and gone unresponsive * 22:24 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 01m 26s) * 22:24 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 22:23 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 22:22 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024'] * 22:11 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 21:54 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs1024'] * 21:54 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 21:53 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs1024'] * 21:53 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 21:49 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2019.codfw.wmnet -> wdqs2020.codfw.wmnet, repooling source-only afterwards * 20:57 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] (duration: 09m 10s) * 20:55 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1024.eqiad.wmnet with OS bookworm * 20:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2020.codfw.wmnet with OS bookworm * 20:52 krinkle@deploy2003: krinkle: Continuing with deployment * 20:49 krinkle@deploy2003: krinkle: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:47 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] * 20:45 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] (duration: 05m 42s) * 20:44 krinkle@deploy2003: krinkle: Rolling back deployment * 20:41 krinkle@deploy2003: krinkle: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:39 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] * 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2003.codfw.wmnet * 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1003.eqiad.wmnet * 20:33 dani@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] (duration: 11m 15s) * 20:33 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2003.codfw.wmnet * 20:33 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1003.eqiad.wmnet * 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2020.codfw.wmnet with reason: host reimage * 20:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1002.eqiad.wmnet * 20:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2002.codfw.wmnet * 20:30 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 20:29 dani@deploy2003: dani: Continuing with deployment * 20:29 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2020.codfw.wmnet with reason: host reimage * 20:26 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1002.eqiad.wmnet * 20:26 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2002.codfw.wmnet * 20:24 dani@deploy2003: dani: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2001.codfw.wmnet * 20:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1001.eqiad.wmnet * 20:22 dani@deploy2003: Started scap sync-world: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] * 20:22 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 20:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2001.codfw.wmnet * 20:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1001.eqiad.wmnet * 20:14 sbisson@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] (duration: 09m 01s) * 20:11 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2020.codfw.wmnet with OS bookworm * 20:10 sbisson@deploy2003: sbisson: Continuing with deployment * 20:07 sbisson@deploy2003: sbisson: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:05 sbisson@deploy2003: Started scap sync-world: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] * 20:03 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] (duration: 07m 04s) * 20:01 mutante: Gerrit - tomorrow a new SSH host key will appear - it will be {{Gerrit|ed25519}} and has already been added to wmf-laptop. you can verify it here: https://wikitech.wikimedia.org/wiki/Help:SSH_Fingerprints/gerrit.wikimedia.org:29418 ([[phab:T240266|T240266]]) * 19:59 zabe@deploy2003: zabe: Continuing with deployment * 19:58 zabe@deploy2003: zabe: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:56 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] * 19:52 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] (duration: 07m 25s) * 19:48 zabe@deploy2003: zabe: Continuing with deployment * 19:47 zabe@deploy2003: zabe: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:45 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] * 19:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 19:32 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024'] * 19:27 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 19:26 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024'] * 19:26 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 19:24 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024'] * 19:12 ryankemper: [wdqs] [[phab:T430880|T430880]] Repooled `wdqs-scholarly` discovery in `eqiad` after validating `wdqs1023` end-to-end; `wdqs1024` remains disabled pending reimage recovery * 19:11 ryankemper: [wdqs] [[phab:T430880|T430880]] Repooled wdqs1012.eqiad.wmnet after successful Bookworm reimage, data transfer, service checks, readiness probe, and cross-graph federation query validation * 19:10 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 19:10 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1012.eqiad.wmnet * 19:08 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 18:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for deploy1003.eqiad.wmnet * 18:57 kamila@cumin1003: START - Cookbook sre.hosts.remove-downtime for deploy1003.eqiad.wmnet * 18:37 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:37 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding urldownloader service IPs - sukhe@cumin1003" * 18:37 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding urldownloader service IPs - sukhe@cumin1003" * 18:32 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 18:32 dancy@deploy2003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 18:30 sukhe@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 18:27 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 18:24 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1024.eqiad.wmnet with OS bookworm * 18:20 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host deploy1003.eqiad.wmnet with OS bookworm * 18:09 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deploy1003 reimage (duration: 121m 16s) * 18:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 18:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1180: Security updates * 17:55 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wcqs2003.codfw.wmnet * 17:48 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wcqs2003.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1155.eqiad.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1155.eqiad.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2224.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2224.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2217.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2217.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2193.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2193.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2180.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2180.codfw.wmnet * 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1168.eqiad.wmnet * 17:36 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1168.eqiad.wmnet * 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2169.codfw.wmnet * 17:36 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2169.codfw.wmnet * 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1165.eqiad.wmnet * 17:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1165.eqiad.wmnet * 17:35 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2158.codfw.wmnet * 17:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2158.codfw.wmnet * 17:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wcqs1003.eqiad.wmnet * 17:17 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1180: Security updates * 17:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1180.eqiad.wmnet * 17:16 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1180.eqiad.wmnet * 17:15 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp3073.* * 17:13 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wcqs1003.eqiad.wmnet * 17:11 brett@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp3073.esams.wmnet with OS trixie * 17:11 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 17:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1024 * 17:04 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1024 * 17:03 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 17:00 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 16:59 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2242: codfw rack B7 depool for maintenance * 16:59 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 16:43 brett@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp3073.esams.wmnet with reason: host reimage * 16:42 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 16:39 brett@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cp3073.esams.wmnet with reason: host reimage * 16:32 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on deploy1003.eqiad.wmnet with reason: host reimage * 16:27 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on deploy1003.eqiad.wmnet with reason: host reimage * 16:14 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2242: codfw rack B7 depool for maintenance * 16:14 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: codfw rack B7 depool for maintenance * 16:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1023.eqiad.wmnet with OS bookworm * 16:13 brett@cumin2002: START - Cookbook sre.hosts.reimage for host cp3073.esams.wmnet with OS trixie * 16:08 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host deploy1003.eqiad.wmnet with OS bookworm * 16:08 kamila@deploy2003: Locking from deployment [MediaWiki]: deploy1003 reimage * 16:03 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1012.eqiad.wmnet with OS bookworm * 15:48 inflatador: bking@apt1002 `sudo reprepro copy bookworm-wikimedia bullseye-wikimedia jvmquake` [[phab:T430880|T430880]] * 15:39 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp3073.* * 15:39 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 15:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:34 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1311536{{!}}Set $wgMathInternalRestbaseURL explicitly (take 2) (T349582)]] * 15:29 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:29 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2228: codfw rack B7 depool for maintenance * 15:29 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2229: codfw rack B7 depool for maintenance * 15:27 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1180: Security update * 15:25 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Security update * 15:21 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 15:21 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 15:19 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 24s) * 15:19 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:14 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 15:14 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db1180: Security update * 15:13 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-eqiad * 14:48 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-eqiad * 14:44 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2229: codfw rack B7 depool for maintenance * 14:44 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc2017: codfw rack B7 depool for maintenance * 14:44 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:43 cmooney@cumin2003: START - Cookbook sre.mysql.parsercache * 14:43 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool pc2017: codfw rack B7 depool for maintenance * 14:43 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2003.codfw.wmnet * 14:43 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2003.codfw.wmnet * 14:42 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2009.codfw.wmnet * 14:42 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2009.codfw.wmnet * 14:41 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:41 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:40 cmooney@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 29 hosts * 14:40 cmooney@cumin1003: START - Cookbook sre.hosts.remove-downtime for 29 hosts * 14:35 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 14:34 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 14:32 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] (duration: 07m 56s) * 14:29 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ssw1-a[1,8]-codfw with reason: lsw1-b7-codfw JunOS upgrade * 14:28 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 14:28 elukey@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 14:26 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:24 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] * 14:23 topranks: reboot lsw1-b7-codfw to upgrade JunOS (affects all hosts in rack) [[phab:T430928|T430928]] * 14:18 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2003.codfw.wmnet * 14:14 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2009.codfw.wmnet * 14:14 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Security update * 14:13 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2242: codfw rack B7 depool for maintenance * 14:13 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2242: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2228: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2228: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2229: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2229: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc2017: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.parsercache * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool pc2017: codfw rack B7 depool for maintenance * 14:08 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2003.codfw.wmnet * 14:07 cmooney@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on aux-k8s-etcd2004.codfw.wmnet,ml-etcd2001.codfw.wmnet with reason: lsw1-b7-codfw JunOS upgrade * 14:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95005 and previous config saved to /var/cache/conftool/dbconfig/20260721-140620-cwilliams.json * 14:05 cmooney@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2049.codfw.wmnet * 14:05 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-scholarly,name=eqiad * 14:04 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2009.codfw.wmnet * 14:04 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply * 14:04 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply * 14:03 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:03 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2001.codfw.wmnet * 14:03 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2001.codfw.wmnet * 14:02 cmooney@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2049.codfw.wmnet * 14:00 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:00 Dreamy_Jazz: Created cusi_case, cusi_signal, and cusi_user on svwiki, dewiki, jawiki, eswiki * 13:59 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b7-codfw,lsw1-b7-codfw IPv6,lsw1-b7-codfw.mgmt,ssw1-a[1,8]-codfw.mgmt with reason: lsw1-b7-codfw JunOS upgrade * 13:57 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs1023.eqiad.wmnet, repooling source-only afterwards * 13:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 29 hosts with reason: lsw1-b7-codfw JunOS upgrade * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224', diff saved to https://phabricator.wikimedia.org/P95003 and previous config saved to /var/cache/conftool/dbconfig/20260721-135613-cwilliams.json * 13:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1012.eqiad.wmnet with reason: host reimage * 13:53 cmooney@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 1:00:00 on 30 hosts with reason: lsw1-b7-codfw JunOS upgrade * 13:51 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1012.eqiad.wmnet with reason: host reimage * 13:48 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 13:48 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 13:46 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224', diff saved to https://phabricator.wikimedia.org/P95001 and previous config saved to /var/cache/conftool/dbconfig/20260721-134605-cwilliams.json * 13:46 cmooney@cumin1003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti2033.codfw.wmnet * 13:46 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 13:45 elukey: move the Docker Registry's /v2/wikimedia/machinelearning.* prefix to the ml S3 backend - [[phab:T428022|T428022]] * 13:45 cmooney@cumin1003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti2033.codfw.wmnet * 13:45 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 13:43 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:40 jiji@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 13:40 jiji@deploy2003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 13:39 jiji@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 13:39 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 13:38 cmooney@cumin1003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2032.codfw.wmnet * 13:38 jiji@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 13:38 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 13:37 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2032.codfw.wmnet * 13:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95000 and previous config saved to /var/cache/conftool/dbconfig/20260721-133557-cwilliams.json * 13:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1012 * 13:33 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1012 * 13:33 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1012.eqiad.wmnet with OS bookworm * 13:30 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:30 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:28 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94999 and previous config saved to /var/cache/conftool/dbconfig/20260721-132855-cwilliams.json * 13:28 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2224.codfw.wmnet with reason: Maintenance * 13:28 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94998 and previous config saved to /var/cache/conftool/dbconfig/20260721-132826-cwilliams.json * 13:28 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 13:23 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] (duration: 07m 50s) * 13:20 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 13:18 kharlan@deploy2003: kharlan: Continuing with deployment * 13:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217', diff saved to https://phabricator.wikimedia.org/P94996 and previous config saved to /var/cache/conftool/dbconfig/20260721-131817-cwilliams.json * 13:17 brouberol@dns1004: END - running authdns-update * 13:17 kharlan@deploy2003: kharlan: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:15 brouberol@dns1004: START - running authdns-update * 13:15 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] * 13:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94995 and previous config saved to /var/cache/conftool/dbconfig/20260721-131411-cwilliams.json * 13:13 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs1023.eqiad.wmnet, repooling source-only afterwards * 13:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217', diff saved to https://phabricator.wikimedia.org/P94994 and previous config saved to /var/cache/conftool/dbconfig/20260721-130809-cwilliams.json * 13:07 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 13:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180', diff saved to https://phabricator.wikimedia.org/P94993 and previous config saved to /var/cache/conftool/dbconfig/20260721-130404-cwilliams.json * 13:03 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:03 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:02 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 13:02 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 12:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94992 and previous config saved to /var/cache/conftool/dbconfig/20260721-125801-cwilliams.json * 12:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180', diff saved to https://phabricator.wikimedia.org/P94991 and previous config saved to /var/cache/conftool/dbconfig/20260721-125356-cwilliams.json * 12:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94990 and previous config saved to /var/cache/conftool/dbconfig/20260721-125049-cwilliams.json * 12:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2217.codfw.wmnet with reason: Maintenance * 12:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94989 and previous config saved to /var/cache/conftool/dbconfig/20260721-125017-cwilliams.json * 12:48 elukey: bmc cold reboot for lvs1013 and lvs1015 - [[phab:T426180|T426180]] * 12:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94988 and previous config saved to /var/cache/conftool/dbconfig/20260721-124348-cwilliams.json * 12:40 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193', diff saved to https://phabricator.wikimedia.org/P94987 and previous config saved to /var/cache/conftool/dbconfig/20260721-124009-cwilliams.json * 12:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts * 12:33 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts * 12:33 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts * 12:32 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts * 12:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:30 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193', diff saved to https://phabricator.wikimedia.org/P94986 and previous config saved to /var/cache/conftool/dbconfig/20260721-123001-cwilliams.json * 12:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94985 and previous config saved to /var/cache/conftool/dbconfig/20260721-121953-cwilliams.json * 12:17 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs2007.codfw.wmnet with OS bookworm * 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94983 and previous config saved to /var/cache/conftool/dbconfig/20260721-121257-cwilliams.json * 12:12 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2193.codfw.wmnet with reason: Maintenance * 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94982 and previous config saved to /var/cache/conftool/dbconfig/20260721-121239-cwilliams.json * 12:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180', diff saved to https://phabricator.wikimedia.org/P94980 and previous config saved to /var/cache/conftool/dbconfig/20260721-120231-cwilliams.json * 11:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180', diff saved to https://phabricator.wikimedia.org/P94979 and previous config saved to /var/cache/conftool/dbconfig/20260721-115223-cwilliams.json * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94978 and previous config saved to /var/cache/conftool/dbconfig/20260721-114333-cwilliams.json * 11:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1180.eqiad.wmnet with reason: Maintenance * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94977 and previous config saved to /var/cache/conftool/dbconfig/20260721-114305-cwilliams.json * 11:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94976 and previous config saved to /var/cache/conftool/dbconfig/20260721-114215-cwilliams.json * 11:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94975 and previous config saved to /var/cache/conftool/dbconfig/20260721-113530-cwilliams.json * 11:35 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2180.codfw.wmnet with reason: Maintenance * 11:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94974 and previous config saved to /var/cache/conftool/dbconfig/20260721-113501-cwilliams.json * 11:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168', diff saved to https://phabricator.wikimedia.org/P94973 and previous config saved to /var/cache/conftool/dbconfig/20260721-113258-cwilliams.json * 11:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169', diff saved to https://phabricator.wikimedia.org/P94972 and previous config saved to /var/cache/conftool/dbconfig/20260721-112453-cwilliams.json * 11:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168', diff saved to https://phabricator.wikimedia.org/P94971 and previous config saved to /var/cache/conftool/dbconfig/20260721-112250-cwilliams.json * 11:21 XioNoX: put eqiad-drmrs Arelion link in service * 11:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169', diff saved to https://phabricator.wikimedia.org/P94970 and previous config saved to /var/cache/conftool/dbconfig/20260721-111446-cwilliams.json * 11:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94969 and previous config saved to /var/cache/conftool/dbconfig/20260721-111242-cwilliams.json * 11:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 11:10 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 11:07 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1093 hosts * 11:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94968 and previous config saved to /var/cache/conftool/dbconfig/20260721-110548-cwilliams.json * 11:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1168.eqiad.wmnet with reason: Maintenance * 11:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94967 and previous config saved to /var/cache/conftool/dbconfig/20260721-110520-cwilliams.json * 11:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94966 and previous config saved to /var/cache/conftool/dbconfig/20260721-110439-cwilliams.json * 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94964 and previous config saved to /var/cache/conftool/dbconfig/20260721-105632-cwilliams.json * 10:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2169.codfw.wmnet with reason: Maintenance * 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94963 and previous config saved to /var/cache/conftool/dbconfig/20260721-105603-cwilliams.json * 10:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165', diff saved to https://phabricator.wikimedia.org/P94962 and previous config saved to /var/cache/conftool/dbconfig/20260721-105512-cwilliams.json * 10:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158', diff saved to https://phabricator.wikimedia.org/P94961 and previous config saved to /var/cache/conftool/dbconfig/20260721-104555-cwilliams.json * 10:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165', diff saved to https://phabricator.wikimedia.org/P94960 and previous config saved to /var/cache/conftool/dbconfig/20260721-104504-cwilliams.json * 10:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158', diff saved to https://phabricator.wikimedia.org/P94959 and previous config saved to /var/cache/conftool/dbconfig/20260721-103547-cwilliams.json * 10:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94958 and previous config saved to /var/cache/conftool/dbconfig/20260721-103456-cwilliams.json * 10:29 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2229: Upgraded kernel * 10:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94956 and previous config saved to /var/cache/conftool/dbconfig/20260721-102757-cwilliams.json * 10:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on an-redacteddb1001.eqiad.wmnet,clouddb[1015,1025,1028].eqiad.wmnet,db1155.eqiad.wmnet with reason: Maintenance * 10:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1165.eqiad.wmnet with reason: Maintenance * 10:25 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94955 and previous config saved to /var/cache/conftool/dbconfig/20260721-102539-cwilliams.json * 10:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94954 and previous config saved to /var/cache/conftool/dbconfig/20260721-101848-cwilliams.json * 10:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2158.codfw.wmnet with reason: Maintenance * 09:43 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2229: Upgraded kernel * 09:42 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2229.codfw.wmnet * 09:42 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2229.codfw.wmnet * 09:23 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db2229.codfw.wmnet * 09:23 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2229.codfw.wmnet * 08:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2229 [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94948 and previous config saved to /var/cache/conftool/dbconfig/20260721-085724-cwilliams.json * 08:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2214 to s6 primary [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94947 and previous config saved to /var/cache/conftool/dbconfig/20260721-085442-cwilliams.json * 08:53 cezmunsta: Starting s6 codfw failover from db2229 to db2214 - [[phab:T430964|T430964]] * 08:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2214 with weight 0 [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94946 and previous config saved to /var/cache/conftool/dbconfig/20260721-084613-cwilliams.json * 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 22 hosts with reason: Primary switchover s6 [[phab:T430964|T430964]] * 08:32 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1017.eqiad.wmnet,service=s1 * 08:08 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Add subrated circuit rate to interface descriptions - CR1312476 - ayounsi@cumin1003 * 08:06 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Add subrated circuit rate to interface descriptions - CR1312476 - ayounsi@cumin1003 * 07:58 reedy@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] (duration: 12m 55s) * 07:51 reedy@deploy2003: reedy, neriah: Continuing with deployment * 07:51 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 07:51 reedy@deploy2003: reedy, neriah: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:48 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1093 hosts * 07:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm2001.wikimedia.org * 07:45 reedy@deploy2003: Started scap sync-world: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] * 07:43 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 07:42 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm2001.wikimedia.org * 07:23 elukey: upgrade libtiff6 packages on zuul* trixie hosts for security upgrades * 07:22 elukey: upgrade libtiff6 packages on Wikikube trixie workers for security upgrades * 07:14 elukey@deploy2003: helmfile [codfw] DONE helmfile.d/services/proton: sync * 07:13 elukey@deploy2003: helmfile [codfw] START helmfile.d/services/proton: sync * 07:11 elukey@deploy2003: helmfile [eqiad] DONE helmfile.d/services/proton: sync * 07:10 elukey@deploy2003: helmfile [eqiad] START helmfile.d/services/proton: sync * 07:09 elukey@deploy2003: helmfile [staging] DONE helmfile.d/services/proton: sync * 07:08 elukey@deploy2003: helmfile [staging] START helmfile.d/services/proton: sync * 06:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1023.eqiad.wmnet with reason: host reimage * 06:46 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1023.eqiad.wmnet with reason: host reimage * 06:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 05:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Haproxy-only mode support - oblivian@cumin1003" * 05:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Haproxy-only mode support - oblivian@cumin1003 * 05:42 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Haproxy-only mode support - oblivian@cumin1003 * 05:42 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Haproxy-only mode support - oblivian@cumin1003" * 05:38 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1017.eqiad.wmnet with reason: Cloning * 05:37 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1017.eqiad.wmnet,service=s1 * 05:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1029.eqiad.wmnet,service=s8 * 05:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1029.eqiad.wmnet,service=s5 * 05:32 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:30 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 05:11 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet * 05:04 aokoth@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet * 05:00 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 04:56 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 04:01 mwpresync@deploy2003: Pruned MediaWiki: 1.47.0-wmf.9 (duration: 01m 08s) * 03:41 mwpresync@deploy2003: Finished scap sync-world: testwikis to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] (duration: 36m 30s) * 03:05 mwpresync@deploy2003: Started scap sync-world: testwikis to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 03:01 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:01 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:00 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:00 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:36 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:36 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:36 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:35 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:16 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 47s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 00:56 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm == 2026-07-20 == * 23:38 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 23:07 Amir1: deleting echo notifications from 2015 on group1 wikis * 23:07 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] (duration: 14m 16s) * 23:01 ladsgroup@deploy2003: ladsgroup: Continuing with deployment * 23:00 ladsgroup@deploy2003: ladsgroup: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:53 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] * 22:46 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2007.codfw.wmnet, repooling source-only afterwards * 22:39 maryum: Deployed security fixes for several security bugs * 21:42 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 21:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2007.codfw.wmnet, repooling source-only afterwards * 21:37 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 21:37 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 21:34 sbassett: Deployed security fix for [[phab:T432424|T432424]] * 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs2020.codfw.wmnet * 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1023.eqiad.wmnet * 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1011.eqiad.wmnet * 21:32 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 17s) * 21:32 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 21:27 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 21:13 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2007.codfw.wmnet with reason: host reimage * 21:08 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-internal-main,name=codfw * 21:06 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2007.codfw.wmnet with reason: host reimage * 20:59 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service * 20:58 sukhe: pybal restart for IP changes around wdqs-main hosts * 20:57 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 20:46 ryankemper@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-internal-main,name=codfw * 20:45 ebernhardson@deploy2003: Finished deploy [search/mjolnir/deploy@d4dc3b8]: Update for opensearch 2.x compat (duration: 00m 34s) * 20:44 ebernhardson@deploy2003: Started deploy [search/mjolnir/deploy@d4dc3b8]: Update for opensearch 2.x compat * 20:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2007 * 20:44 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2007 * 20:43 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2007 * 20:43 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2007.codfw.wmnet 156.16.192.10.in-addr.arpa 6.5.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:42 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2007.codfw.wmnet 156.16.192.10.in-addr.arpa 6.5.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:42 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2007 - bking@cumin2003" * 20:41 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2007 - bking@cumin2003" * 20:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94944 and previous config saved to /var/cache/conftool/dbconfig/20260720-203333-cwilliams.json * 20:32 arlolra@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] (duration: 15m 07s) * 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2020.codfw.wmnet * 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1023.eqiad.wmnet * 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1011.eqiad.wmnet * 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs2020.codfw.wmnet * 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1023.eqiad.wmnet * 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1011.eqiad.wmnet * 20:25 arlolra@deploy2003: arlolra, cscott: Continuing with deployment * 20:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257', diff saved to https://phabricator.wikimedia.org/P94943 and previous config saved to /var/cache/conftool/dbconfig/20260720-202325-cwilliams.json * 20:21 arlolra@deploy2003: arlolra, cscott: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:17 arlolra@deploy2003: Started scap sync-world: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] * 20:13 bking@cumin2003: START - Cookbook sre.dns.netbox * 20:13 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257', diff saved to https://phabricator.wikimedia.org/P94942 and previous config saved to /var/cache/conftool/dbconfig/20260720-201318-cwilliams.json * 20:13 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 20:10 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 20:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2007 * 20:04 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2007.codfw.wmnet with OS bookworm * 20:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94941 and previous config saved to /var/cache/conftool/dbconfig/20260720-200310-cwilliams.json * 19:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94940 and previous config saved to /var/cache/conftool/dbconfig/20260720-195633-cwilliams.json * 19:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1257.eqiad.wmnet with reason: Maintenance * 19:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94939 and previous config saved to /var/cache/conftool/dbconfig/20260720-195605-cwilliams.json * 19:51 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 19:50 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 19:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256', diff saved to https://phabricator.wikimedia.org/P94938 and previous config saved to /var/cache/conftool/dbconfig/20260720-194558-cwilliams.json * 19:44 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 19:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 19:41 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wdqs1011.eqiad.wmnet with OS bookworm * 19:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256', diff saved to https://phabricator.wikimedia.org/P94937 and previous config saved to /var/cache/conftool/dbconfig/20260720-193550-cwilliams.json * 19:25 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94936 and previous config saved to /var/cache/conftool/dbconfig/20260720-192542-cwilliams.json * 19:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94935 and previous config saved to /var/cache/conftool/dbconfig/20260720-191856-cwilliams.json * 19:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1256.eqiad.wmnet with reason: Maintenance * 19:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94934 and previous config saved to /var/cache/conftool/dbconfig/20260720-191839-cwilliams.json * 19:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255', diff saved to https://phabricator.wikimedia.org/P94933 and previous config saved to /var/cache/conftool/dbconfig/20260720-190831-cwilliams.json * 18:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255', diff saved to https://phabricator.wikimedia.org/P94932 and previous config saved to /var/cache/conftool/dbconfig/20260720-185824-cwilliams.json * 18:50 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 18:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94931 and previous config saved to /var/cache/conftool/dbconfig/20260720-184816-cwilliams.json * 18:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94930 and previous config saved to /var/cache/conftool/dbconfig/20260720-184224-cwilliams.json * 18:42 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1255.eqiad.wmnet with reason: Maintenance * 18:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94929 and previous config saved to /var/cache/conftool/dbconfig/20260720-184153-cwilliams.json * 18:39 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 18:39 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 16s) * 18:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 18:38 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 59m 26s) * 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 18:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211', diff saved to https://phabricator.wikimedia.org/P94928 and previous config saved to /var/cache/conftool/dbconfig/20260720-183145-cwilliams.json * 18:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211', diff saved to https://phabricator.wikimedia.org/P94927 and previous config saved to /var/cache/conftool/dbconfig/20260720-182137-cwilliams.json * 18:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94926 and previous config saved to /var/cache/conftool/dbconfig/20260720-181129-cwilliams.json * 18:09 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_codfw * 18:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2057.codfw.wmnet * 18:08 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_codfw * 18:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2058.codfw.wmnet * 18:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94925 and previous config saved to /var/cache/conftool/dbconfig/20260720-180452-cwilliams.json * 18:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on clouddb[1016,1020,1022-1023].eqiad.wmnet,db1154.eqiad.wmnet with reason: Maintenance * 18:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1211.eqiad.wmnet with reason: Maintenance * 18:02 sukhe: armed keyholder on acmechief1002.eqiad.wmnet and acmechief2002.codfw.wmnet (active host) * 18:01 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief2002.codfw.wmnet * 17:57 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief2002.codfw.wmnet * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs2020'] * 17:52 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief1002.eqiad.wmnet * 17:50 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 17:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1011.eqiad.wmnet with reason: host reimage * 17:48 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief1002.eqiad.wmnet * 17:47 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test2001.codfw.wmnet * 17:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94924 and previous config saved to /var/cache/conftool/dbconfig/20260720-174717-cwilliams.json * 17:46 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 17:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1011.eqiad.wmnet with reason: host reimage * 17:43 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs2020.codfw.wmnet with OS bookworm * 17:43 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test2001.codfw.wmnet * 17:43 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test1001.eqiad.wmnet * 17:39 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test1001.eqiad.wmnet * 17:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 17:38 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:38 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 17:37 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 17:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244', diff saved to https://phabricator.wikimedia.org/P94923 and previous config saved to /var/cache/conftool/dbconfig/20260720-173709-cwilliams.json * 17:35 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 17:31 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 20m 40s) * 17:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2055.codfw.wmnet * 17:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2056.codfw.wmnet * 17:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1011 * 17:27 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1011 * 17:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1011.eqiad.wmnet with OS bookworm * 17:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244', diff saved to https://phabricator.wikimedia.org/P94922 and previous config saved to /var/cache/conftool/dbconfig/20260720-172701-cwilliams.json * 17:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94921 and previous config saved to /var/cache/conftool/dbconfig/20260720-171653-cwilliams.json * 17:11 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 17:11 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 13m 03s) * 17:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94920 and previous config saved to /var/cache/conftool/dbconfig/20260720-171012-cwilliams.json * 17:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2244.codfw.wmnet with reason: Maintenance * 17:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94919 and previous config saved to /var/cache/conftool/dbconfig/20260720-170941-cwilliams.json * 16:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243', diff saved to https://phabricator.wikimedia.org/P94918 and previous config saved to /var/cache/conftool/dbconfig/20260720-165933-cwilliams.json * 16:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 16:58 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2053.codfw.wmnet * 16:51 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2054.codfw.wmnet * 16:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243', diff saved to https://phabricator.wikimedia.org/P94917 and previous config saved to /var/cache/conftool/dbconfig/20260720-164926-cwilliams.json * 16:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94916 and previous config saved to /var/cache/conftool/dbconfig/20260720-163918-cwilliams.json * 16:35 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 16:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94915 and previous config saved to /var/cache/conftool/dbconfig/20260720-163140-cwilliams.json * 16:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2243.codfw.wmnet with reason: Maintenance * 16:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94914 and previous config saved to /var/cache/conftool/dbconfig/20260720-163111-cwilliams.json * 16:27 btullis@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 16:27 btullis@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 16:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2020 * 16:23 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2020 * 16:21 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2020 * 16:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2020.codfw.wmnet 85.0.192.10.in-addr.arpa 5.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:21 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2020.codfw.wmnet 85.0.192.10.in-addr.arpa 5.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242', diff saved to https://phabricator.wikimedia.org/P94913 and previous config saved to /var/cache/conftool/dbconfig/20260720-162103-cwilliams.json * 16:19 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 16:18 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 16:18 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:18 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 16:17 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:17 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for netbox accounting errors - jhancock@cumin2002" * 16:17 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for netbox accounting errors - jhancock@cumin2002" * 16:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2051.codfw.wmnet * 16:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2052.codfw.wmnet * 16:11 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 16:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242', diff saved to https://phabricator.wikimedia.org/P94912 and previous config saved to /var/cache/conftool/dbconfig/20260720-161055-cwilliams.json * 16:09 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 16:08 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 16:06 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 16:06 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 16:06 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2020 * 16:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2020.codfw.wmnet with OS bookworm * 16:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94911 and previous config saved to /var/cache/conftool/dbconfig/20260720-160047-cwilliams.json * 15:58 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2019.codfw.wmnet, repooling source-only afterwards * 15:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94909 and previous config saved to /var/cache/conftool/dbconfig/20260720-155353-cwilliams.json * 15:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2242.codfw.wmnet with reason: Maintenance * 15:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94908 and previous config saved to /var/cache/conftool/dbconfig/20260720-154433-cwilliams.json * 15:35 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2049.codfw.wmnet * 15:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162', diff saved to https://phabricator.wikimedia.org/P94907 and previous config saved to /var/cache/conftool/dbconfig/20260720-153425-cwilliams.json * 15:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2050.codfw.wmnet * 15:28 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162', diff saved to https://phabricator.wikimedia.org/P94906 and previous config saved to /var/cache/conftool/dbconfig/20260720-152418-cwilliams.json * 15:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1023 * 15:14 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1023 * 15:14 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 15:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94905 and previous config saved to /var/cache/conftool/dbconfig/20260720-151407-cwilliams.json * 15:13 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] (duration: 41m 16s) * 15:08 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 15:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94902 and previous config saved to /var/cache/conftool/dbconfig/20260720-150729-cwilliams.json * 15:07 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2162.codfw.wmnet with reason: Maintenance * 15:05 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2027.codfw.wmnet, repooling source-only afterwards * 15:00 urbanecm@deploy2003: vadymts1, migr, urbanecm: Continuing with deployment * 14:59 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:58 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 07s) * 14:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:58 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 13s) * 14:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:57 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2019.codfw.wmnet, repooling source-only afterwards * 14:57 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2047.codfw.wmnet * 14:55 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2048.codfw.wmnet * 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2019.codfw.wmnet with OS bookworm * 14:49 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 14:47 urbanecm@deploy2003: vadymts1, migr, urbanecm: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:44 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool magru [reason: BGP issues in lvs7003 resolved after liberica restart, no task ID specified] * 14:44 sukhe@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool magru [reason: BGP issues in lvs7003 resolved after liberica restart, no task ID specified] * 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:41 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:39 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:39 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:33 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool magru [reason: no reason specified, no task ID specified] * 14:33 sukhe@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool magru [reason: no reason specified, no task ID specified] * 14:31 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] * 14:24 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:24 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:24 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:24 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2019.codfw.wmnet with reason: host reimage * 14:22 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2027.codfw.wmnet, repooling source-only afterwards * 14:19 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2019.codfw.wmnet with reason: host reimage * 14:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2027.codfw.wmnet with OS bookworm * 14:16 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2046.codfw.wmnet * 14:16 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2045.codfw.wmnet * 14:08 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:08 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:08 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:08 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:07 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:06 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:06 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:06 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:05 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2019 * 14:00 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2019 * 13:56 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1015.eqiad.wmnet * 13:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2027.codfw.wmnet with reason: host reimage * 13:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2071.codfw.wmnet with OS trixie * 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:51 sukhe@cumin1003: END (ERROR) - Cookbook sre.loadbalancer.admin (exit_code=97) rebooting A:liberica and P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica and P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:51 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1015.eqiad.wmnet * 13:50 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1014.eqiad.wmnet * 13:50 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1076.eqiad.wmnet with OS trixie * 13:50 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2027.codfw.wmnet with reason: host reimage * 13:45 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1014.eqiad.wmnet * 13:44 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1013.eqiad.wmnet * 13:39 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 13:38 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1013.eqiad.wmnet * 13:37 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2044.codfw.wmnet * 13:37 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2043.codfw.wmnet * 13:36 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2019 * 13:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2019.codfw.wmnet 156.32.192.10.in-addr.arpa 6.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:36 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2019.codfw.wmnet 156.32.192.10.in-addr.arpa 6.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:36 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2019 - bking@cumin2003" * 13:36 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2019 - bking@cumin2003" * 13:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry2005.codfw.wmnet * 13:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2071.codfw.wmnet with reason: host reimage * 13:31 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:31 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2019 * 13:31 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry2005.codfw.wmnet * 13:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry2004.codfw.wmnet * 13:30 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2019.codfw.wmnet with OS bookworm * 13:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2027 * 13:30 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2027 * 13:30 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2027.codfw.wmnet with OS bookworm * 13:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1076.eqiad.wmnet with reason: host reimage * 13:29 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_codfw * 13:28 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_codfw * 13:26 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry2004.codfw.wmnet * 13:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry1005.eqiad.wmnet * 13:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2071.codfw.wmnet with reason: host reimage * 13:22 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1076.eqiad.wmnet with reason: host reimage * 13:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry1005.eqiad.wmnet * 13:21 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry1004.eqiad.wmnet * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry1004.eqiad.wmnet * 13:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts * 13:13 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts * 13:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts * 13:12 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts * 13:03 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1076.eqiad.wmnet with OS trixie * 13:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2071.codfw.wmnet with OS trixie * 12:55 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:54 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:53 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:46 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 7 hosts * 12:42 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 7 hosts * 12:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts * 12:42 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts * 12:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2070.codfw.wmnet with OS trixie * 12:36 ozge@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:35 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1075.eqiad.wmnet with OS trixie * 12:32 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts * 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts * 12:22 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin1001.eqiad.wmnet * 12:19 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin1001.eqiad.wmnet * 12:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2070.codfw.wmnet with reason: host reimage * 12:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin2001.codfw.wmnet * 12:14 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1075.eqiad.wmnet with reason: host reimage * 12:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2070.codfw.wmnet with reason: host reimage * 12:10 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1075.eqiad.wmnet with reason: host reimage * 12:09 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin2001.codfw.wmnet * 11:17 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1074.eqiad.wmnet with OS trixie * 11:17 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 11:16 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 11:14 ozge@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 6 hosts * 11:09 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 6 hosts * 11:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 324 hosts * 10:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1074.eqiad.wmnet with reason: host reimage * 10:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2069.codfw.wmnet with OS trixie * 10:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1074.eqiad.wmnet with reason: host reimage * 10:30 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2069.codfw.wmnet with reason: host reimage * 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1074.eqiad.wmnet with OS trixie * 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2069.codfw.wmnet with reason: host reimage * 10:06 blake@deploy2003: Stopping before sync operations * 10:06 blake@deploy2003: Started scap sync-world: Non-deployment scap run to populate new release values * 10:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2069.codfw.wmnet with OS trixie * 10:00 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1073.eqiad.wmnet with OS trixie * 09:56 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 324 hosts * 09:39 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 09:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 8 hosts * 09:38 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1073.eqiad.wmnet with reason: host reimage * 09:37 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 8 hosts * 09:34 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1073.eqiad.wmnet with reason: host reimage * 09:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2068.codfw.wmnet with OS trixie * 09:16 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1073.eqiad.wmnet with OS trixie * 09:13 blake@deploy2003: sync-world aborted: Non-deployment scap run to populate new release values (duration: 00m 02s) * 09:13 blake@deploy2003: Started scap sync-world: Non-deployment scap run to populate new release values * 08:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2068.codfw.wmnet with reason: host reimage * 08:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2068.codfw.wmnet with reason: host reimage * 08:50 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 08:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2068.codfw.wmnet with OS trixie * 08:15 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1072.eqiad.wmnet with OS trixie * 07:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2067.codfw.wmnet with OS trixie * 07:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1072.eqiad.wmnet with reason: host reimage * 07:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1072.eqiad.wmnet with reason: host reimage * 07:45 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 07:45 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 07:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2067.codfw.wmnet with reason: host reimage * 07:35 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2067.codfw.wmnet with reason: host reimage * 07:30 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 07:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1072.eqiad.wmnet with OS trixie * 07:30 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 07:17 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 07:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2067.codfw.wmnet with OS trixie * 05:51 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:50 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:25 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:25 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on db2207.codfw.wmnet with reason: Host down * 04:28 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 07m 02s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-18 == * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 29s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 00:11 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2018.codfw.wmnet, repooling source-only afterwards == 2026-07-17 == * 23:53 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2026.codfw.wmnet, repooling source-only afterwards * 23:09 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2018.codfw.wmnet, repooling source-only afterwards * 23:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2026.codfw.wmnet, repooling source-only afterwards * 22:11 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2018.codfw.wmnet with OS bookworm * 22:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2026.codfw.wmnet with OS bookworm * 21:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2018.codfw.wmnet with reason: host reimage * 21:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2018.codfw.wmnet with reason: host reimage * 21:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2026.codfw.wmnet with reason: host reimage * 21:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2026.codfw.wmnet with reason: host reimage * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2018 * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2018 * 21:26 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2018 * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2018.codfw.wmnet 155.32.192.10.in-addr.arpa 5.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:26 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2018.codfw.wmnet 155.32.192.10.in-addr.arpa 5.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2018 - bking@cumin2003" * 21:26 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2018 - bking@cumin2003" * 21:14 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:13 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2018 * 21:13 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2018.codfw.wmnet with OS bookworm * 21:12 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2026 * 21:12 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2026 * 21:12 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2026.codfw.wmnet with OS bookworm * 21:05 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1022.eqiad.wmnet -> wdqs1026.eqiad.wmnet, repooling source-only afterwards * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs2017.codfw.wmnet, repooling source-only afterwards * 20:11 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs2017.codfw.wmnet, repooling source-only afterwards * 20:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2017.codfw.wmnet with OS bookworm * 20:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1022.eqiad.wmnet -> wdqs1026.eqiad.wmnet, repooling source-only afterwards * 20:06 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1026.eqiad.wmnet with OS bookworm * 19:55 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 19:55 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:55 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 09s) * 19:55 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:50 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 08s) * 19:50 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:50 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 10m 03s) * 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2017.codfw.wmnet with reason: host reimage * 19:40 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:40 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 15s) * 19:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1026.eqiad.wmnet with reason: host reimage * 19:37 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 16s) * 19:37 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:34 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2017.codfw.wmnet with reason: host reimage * 19:34 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1026.eqiad.wmnet with reason: host reimage * 19:33 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 19:33 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:16 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1026.eqiad.wmnet with OS bookworm * 19:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2017 * 19:16 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2017 * 19:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2017.codfw.wmnet with OS bookworm * 18:30 bking@dns1004: END - running authdns-update * 18:28 bking@dns1004: START - running authdns-update * 18:16 kamila@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1264.eqiad.wmnet * 18:16 kamila@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1264.eqiad.wmnet * 18:16 kamila@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1264.eqiad.wmnet * 17:49 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 17:46 dzahn@dns1006: END - running authdns-update * 17:44 dzahn@dns1006: START - running authdns-update * 17:44 dzahn@dns1006: END - running authdns-update * 17:42 dzahn@dns1006: START - running authdns-update * 17:28 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 17:21 kamila@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 17:01 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1264 * 17:01 kamila@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1264 * 17:01 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 17:01 kamila@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1264.eqiad.wmnet * 17:01 kamila@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1264.eqiad.wmnet * 17:01 kamila@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1264.eqiad.wmnet * 16:42 reedy@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] (duration: 10m 29s) * 16:34 reedy@deploy2003: reedy, hartman: Continuing with deployment * 16:33 reedy@deploy2003: reedy, hartman: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:31 reedy@deploy2003: Started scap sync-world: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] * 16:26 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 16:10 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in2001.wikimedia.org with reason: [[phab:T431659|T431659]] * 16:07 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in1001.wikimedia.org with reason: [[phab:T431659|T431659]] * 16:05 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 16:01 kamila@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 16:00 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out2001.wikimedia.org with reason: [[phab:T431659|T431659]] * 15:41 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 15:41 kamila@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 15:35 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out1001.wikimedia.org with reason: [[phab:T431659|T431659]] * 15:14 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1339.eqiad.wmnet * 15:13 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1339.eqiad.wmnet * 15:13 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1339.eqiad.wmnet * 14:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1339.eqiad.wmnet with OS trixie * 14:50 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:49 kamila@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:49 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:33 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage * 14:27 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage * 14:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1339 * 14:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1339 * 14:14 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1339 * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1339.eqiad.wmnet 156.32.64.10.in-addr.arpa 6.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:14 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1339.eqiad.wmnet 156.32.64.10.in-addr.arpa 6.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1339 - cgoubert@cumin2003" * 14:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1339 - cgoubert@cumin2003" * 14:09 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 14:06 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1339 * 14:06 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie * 14:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1339.eqiad.wmnet * 14:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1339.eqiad.wmnet * 14:02 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1339.eqiad.wmnet * 13:45 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb1013.eqiad.wmnet * 13:39 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb1013.eqiad.wmnet * 13:27 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:24 blake@dns1004: END - running authdns-update * 13:22 blake@dns1004: START - running authdns-update * 13:20 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 13:11 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2014.codfw.wmnet * 13:06 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb2014.codfw.wmnet * 13:06 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2012.codfw.wmnet * 13:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 13:03 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 13:01 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 15 hosts * 13:01 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb2012.codfw.wmnet * 13:01 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1016.eqiad.wmnet * 13:00 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 15 hosts * 12:55 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb1016.eqiad.wmnet * 12:55 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1014.eqiad.wmnet * 12:49 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb1014.eqiad.wmnet * 12:32 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:32 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:31 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:31 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1338.eqiad.wmnet * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1338.eqiad.wmnet * 12:18 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1338.eqiad.wmnet * 12:17 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:15 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:14 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:13 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1338.eqiad.wmnet with OS trixie * 12:01 klausman@deploy2003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 11:59 klausman@deploy2003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 11:56 klausman@deploy2003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 11:54 klausman@deploy2003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 11:53 klausman@deploy2003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 11:51 klausman@deploy2003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 11:42 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1338.eqiad.wmnet with reason: host reimage * 11:38 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1338.eqiad.wmnet with reason: host reimage * 11:31 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2230.codfw.wmnet * 11:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1338 * 11:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1338 * 11:25 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1338 * 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1338.eqiad.wmnet 155.32.64.10.in-addr.arpa 5.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:25 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1338.eqiad.wmnet 155.32.64.10.in-addr.arpa 5.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1338 - cgoubert@cumin2003" * 11:25 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1338 - cgoubert@cumin2003" * 11:23 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2230.codfw.wmnet * 11:20 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 11:20 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1338 * 11:20 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1338.eqiad.wmnet with OS trixie * 11:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1338.eqiad.wmnet * 11:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1338.eqiad.wmnet * 11:19 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1338.eqiad.wmnet * 11:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1337.eqiad.wmnet * 11:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1337.eqiad.wmnet * 11:17 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1337.eqiad.wmnet * 11:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1337.eqiad.wmnet with OS trixie * 10:51 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[2001-2002].codfw.wmnet * 10:50 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1337.eqiad.wmnet with reason: host reimage * 10:40 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 10:39 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:39 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1337.eqiad.wmnet with reason: host reimage * 10:39 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1001-1003].eqiad.wmnet * 10:34 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:34 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:30 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:28 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1001-1003].eqiad.wmnet * 10:27 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1337 * 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1337 * 10:26 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1337 * 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1337.eqiad.wmnet 154.32.64.10.in-addr.arpa 4.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:26 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1337.eqiad.wmnet 154.32.64.10.in-addr.arpa 4.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1337 - cgoubert@cumin2003" * 10:26 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1337 - cgoubert@cumin2003" * 10:21 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 10:18 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1337 * 10:17 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1337.eqiad.wmnet with OS trixie * 10:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1337.eqiad.wmnet * 10:16 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db1176.eqiad.wmnet * 10:16 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1337.eqiad.wmnet * 10:16 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1337.eqiad.wmnet * 10:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1336.eqiad.wmnet * 10:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1336.eqiad.wmnet * 10:15 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1336.eqiad.wmnet * 10:11 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db1176.eqiad.wmnet * 10:10 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db1176.eqiad.wmnet * 10:09 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db1176.eqiad.wmnet * 10:05 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts (check the cookbook's logs for more details.) * 10:03 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts (check the cookbook's logs for more details.) * 09:58 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1336.eqiad.wmnet with OS trixie * 09:47 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts (check the cookbook's logs for more details.) * 09:47 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts (check the cookbook's logs for more details.) * 09:45 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host acmechief-test2001.codfw.wmnet,acmechief-test1001.eqiad.wmnet,an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet,db-test[2001-2002].codfw.wmnet,db-test[1 * 09:40 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host acmechief-test2001.codfw.wmnet,acmechief-test1001.eqiad.wmnet,an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet,db-test[2001-2002].codfw.wmnet,db-test[1001-1003].eqiad.wmn * 09:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1336.eqiad.wmnet with reason: host reimage * 09:33 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1336.eqiad.wmnet with reason: host reimage * 09:29 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet * 09:29 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet * 09:28 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:26 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:21 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 09:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1336 * 09:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1336 * 09:19 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 09:14 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1336 * 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1336.eqiad.wmnet 152.32.64.10.in-addr.arpa 2.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1336.eqiad.wmnet 152.32.64.10.in-addr.arpa 2.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1336 - cgoubert@cumin2003" * 09:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1336 - cgoubert@cumin2003" * 09:11 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:10 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 09:09 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:09 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1336 * 09:09 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1336.eqiad.wmnet with OS trixie * 09:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1336.eqiad.wmnet * 09:08 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1336.eqiad.wmnet * 09:08 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1336.eqiad.wmnet * 09:06 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1335.eqiad.wmnet * 09:06 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1335.eqiad.wmnet * 09:06 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1335.eqiad.wmnet * 09:04 elukey: uploaded spicerack_13.1.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia * 08:55 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wikikube-worker-exp2001.codfw.wmnet * 08:54 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host testreduce1002.eqiad.wmnet * 08:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1335.eqiad.wmnet with OS trixie * 08:51 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host wikikube-worker-exp2001.codfw.wmnet * 08:51 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wikikube-worker-exp1001.eqiad.wmnet * 08:50 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host testreduce1002.eqiad.wmnet * 08:45 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host wikikube-worker-exp1001.eqiad.wmnet * 08:34 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1335.eqiad.wmnet with reason: host reimage * 08:30 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1335.eqiad.wmnet with reason: host reimage * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1335 * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1335 * 08:18 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1335 * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1335.eqiad.wmnet 150.32.64.10.in-addr.arpa 0.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:18 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1335.eqiad.wmnet 150.32.64.10.in-addr.arpa 0.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1335 - cgoubert@cumin2003" * 08:18 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1335 - cgoubert@cumin2003" * 08:14 elukey@cumin1003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 08:14 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges * 08:13 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 08:10 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1335 * 08:10 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1335.eqiad.wmnet with OS trixie * 08:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1335.eqiad.wmnet * 08:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1335.eqiad.wmnet * 08:09 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1335.eqiad.wmnet * 08:06 elukey@cumin1003: END (FAIL) - Cookbook sre.puppet.disable-merges (exit_code=99) * 08:05 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges * 08:03 elukey@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin1003.eqiad.wmnet * 07:57 elukey@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin1003.eqiad.wmnet * 07:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetdb1003.eqiad.wmnet * 07:46 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetdb1003.eqiad.wmnet * 07:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetdb2003.codfw.wmnet * 07:37 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetdb2003.codfw.wmnet * 07:37 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1001.eqiad.wmnet * 07:28 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver1001.eqiad.wmnet * 07:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet * 07:19 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet * 07:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2002.codfw.wmnet * 07:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver2002.codfw.wmnet * 07:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet * 07:05 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet * 07:04 btullis@cumin1003: END (FAIL) - Cookbook sre.hadoop.reboot-workers (exit_code=99) for Hadoop analytics cluster * 07:04 elukey@cumin1003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 07:04 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges * 06:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox1003.eqiad.wmnet * 06:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox1003.eqiad.wmnet * 02:46 ryankemper@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:46 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:44 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:37 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:37 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-internal-scholarly,name=eqiad * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 49s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 01:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore wdqs1025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling source-only afterwards * 01:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore wdqs1027 after Bookworm reimage) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs1027.eqiad.wmnet, repooling both afterwards * 00:55 urbanecm@deploy2003: helmfile [codfw] DONE helmfile.d/services/linkrecommendation: apply * 00:54 urbanecm@deploy2003: helmfile [eqiad] DONE helmfile.d/services/linkrecommendation: apply * 00:54 urbanecm@deploy2003: helmfile [staging] DONE helmfile.d/services/linkrecommendation: apply * 00:54 urbanecm@deploy2003: helmfile [codfw] START helmfile.d/services/linkrecommendation: apply * 00:53 urbanecm@deploy2003: helmfile [staging] START helmfile.d/services/linkrecommendation: apply * 00:52 urbanecm@deploy2003: helmfile [eqiad] START helmfile.d/services/linkrecommendation: apply * 00:23 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore wdqs1025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling source-only afterwards * 00:23 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore wdqs1027 after Bookworm reimage) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs1027.eqiad.wmnet, repooling both afterwards * 00:14 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1274.eqiad.wmnet * 00:14 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1274.eqiad.wmnet * 00:14 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1274.eqiad.wmnet * 00:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1027.eqiad.wmnet with OS bookworm * 00:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1025.eqiad.wmnet with OS bookworm * 00:04 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1274.eqiad.wmnet with OS trixie == 2026-07-16 == * 23:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], xfer to freshly reimaged/scap-deployed wdqs2025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs2025.codfw.wmnet, repooling source-only afterwards * 23:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1027.eqiad.wmnet with reason: host reimage * 23:47 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1025.eqiad.wmnet with reason: host reimage * 23:43 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1274.eqiad.wmnet with reason: host reimage * 23:41 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1025.eqiad.wmnet with reason: host reimage * 23:39 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1027.eqiad.wmnet with reason: host reimage * 23:38 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1274.eqiad.wmnet with reason: host reimage * 23:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1025 * 23:23 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1025 * 23:22 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1027 * 23:22 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1027 * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1274 * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1274 * 23:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1025.eqiad.wmnet with OS bookworm * 23:19 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1274 * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1274.eqiad.wmnet 145.48.64.10.in-addr.arpa 5.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:19 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1274.eqiad.wmnet 145.48.64.10.in-addr.arpa 5.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1274 - swfrench@cumin1003" * 23:19 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1274 - swfrench@cumin1003" * 23:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1027.eqiad.wmnet with OS bookworm * 23:14 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 23:14 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1274 * 23:13 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1274.eqiad.wmnet with OS trixie * 23:13 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1274.eqiad.wmnet * 23:12 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1274.eqiad.wmnet * 23:12 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1274.eqiad.wmnet * 23:12 ryankemper: [[phab:T430880|T430880]] depooled dnsdisc of wdqs-internal-scholarly-eqiad bc we only have 1 host there * 23:09 ryankemper@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-internal-scholarly,name=eqiad * 23:08 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1272.eqiad.wmnet * 23:08 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1272.eqiad.wmnet * 23:08 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1272.eqiad.wmnet * 23:01 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], xfer to freshly reimaged/scap-deployed wdqs2025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs2025.codfw.wmnet, repooling source-only afterwards * 22:57 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1272.eqiad.wmnet with OS trixie * 22:56 Amir1: deleting echo notifications from 2015 in group0 * 22:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2025.codfw.wmnet with OS bookworm * 22:35 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1272.eqiad.wmnet with reason: host reimage * 22:32 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 27s) * 22:32 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 22:28 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1269.eqiad.wmnet * 22:28 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1269.eqiad.wmnet * 22:28 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1269.eqiad.wmnet * 22:27 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1272.eqiad.wmnet with reason: host reimage * 22:26 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] (duration: 08m 51s) * 22:22 ladsgroup@deploy2003: ladsgroup, urbanecm: Continuing with deployment * 22:19 ladsgroup@deploy2003: ladsgroup, urbanecm: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:17 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] * 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2025.codfw.wmnet with reason: host reimage * 22:06 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1272 * 22:06 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1272 * 22:05 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1272 * 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1272.eqiad.wmnet 127.48.64.10.in-addr.arpa 7.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:05 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1272.eqiad.wmnet 127.48.64.10.in-addr.arpa 7.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1272 - swfrench@cumin1003" * 22:05 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1272 - swfrench@cumin1003" * 22:03 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2025.codfw.wmnet with reason: host reimage * 22:01 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 22:00 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1272 * 22:00 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1272.eqiad.wmnet with OS trixie * 22:00 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1272.eqiad.wmnet * 21:59 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1272.eqiad.wmnet * 21:59 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1272.eqiad.wmnet * 21:55 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1271.eqiad.wmnet * 21:55 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1271.eqiad.wmnet * 21:55 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1271.eqiad.wmnet * 21:46 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1271.eqiad.wmnet with OS trixie * 21:45 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] (duration: 06m 31s) * 21:40 sbassett@deploy2003: sbassett: Continuing with deployment * 21:40 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2025 * 21:40 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2025 * 21:40 sbassett@deploy2003: sbassett: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:38 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] * 21:37 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2025 * 21:37 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2025.codfw.wmnet 220.48.192.10.in-addr.arpa 0.2.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:37 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2025.codfw.wmnet 220.48.192.10.in-addr.arpa 0.2.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:37 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:37 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2025 - bking@cumin2003" * 21:37 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2025 - bking@cumin2003" * 21:30 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] (duration: 08m 19s) * 21:26 sbassett@deploy2003: sbassett: Continuing with deployment * 21:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1269.eqiad.wmnet with OS trixie * 21:24 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1271.eqiad.wmnet with reason: host reimage * 21:23 sbassett@deploy2003: sbassett: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:22 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:22 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] * 21:20 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2025 * 21:19 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2025.codfw.wmnet with OS bookworm * 21:17 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1271.eqiad.wmnet with reason: host reimage * 21:04 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1269.eqiad.wmnet with reason: host reimage * 21:00 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1269.eqiad.wmnet with reason: host reimage * 20:56 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1271 * 20:55 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1271 * 20:54 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1271 * 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1271.eqiad.wmnet 126.48.64.10.in-addr.arpa 6.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:54 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1271.eqiad.wmnet 126.48.64.10.in-addr.arpa 6.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1271 - swfrench@cumin1003" * 20:54 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1271 - swfrench@cumin1003" * 20:51 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:51 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:51 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:50 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 20:49 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 20:49 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1268.eqiad.wmnet * 20:49 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1268.eqiad.wmnet * 20:49 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1268.eqiad.wmnet * 20:48 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1271 * 20:48 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1271.eqiad.wmnet with OS trixie * 20:47 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1271.eqiad.wmnet * 20:46 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1271.eqiad.wmnet * 20:46 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1271.eqiad.wmnet * 20:41 aude@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] (duration: 07m 34s) * 20:39 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1269 * 20:39 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1269 * 20:36 aude@deploy2003: aude: Continuing with deployment * 20:35 aude@deploy2003: aude: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:33 aude@deploy2003: Started scap sync-world: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] * 20:26 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on db2207.codfw.wmnet with reason: Host down * 20:22 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-video: apply * 20:21 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-video: apply * 20:20 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-timeline: apply * 20:20 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-timeline: apply * 20:20 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-syntaxhighlight: apply * 20:19 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-syntaxhighlight: apply * 20:19 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-media: apply * 20:18 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-media: apply * 20:18 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-constraints: apply * 20:17 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-constraints: apply * 20:17 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox: apply * 20:16 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox: apply * 20:13 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1269 * 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1269.eqiad.wmnet 80.32.64.10.in-addr.arpa 0.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:13 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1269.eqiad.wmnet 80.32.64.10.in-addr.arpa 0.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1269 - kamila@cumin1003" * 20:13 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1269 - kamila@cumin1003" * 20:09 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-flink-codfw cluster: Roll restart of jvm daemons. * 20:07 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 20:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 20:03 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-flink-codfw cluster: Roll restart of jvm daemons. * 20:03 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2207 [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94893 and previous config saved to /var/cache/conftool/dbconfig/20260716-200257-marostegui.json * 20:01 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2204 to s2 primary [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94892 and previous config saved to /var/cache/conftool/dbconfig/20260716-200157-marostegui.json * 20:00 marostegui: Starting emergency s2 codfw failover from db2207 to db2204 - [[phab:T432396|T432396]] * 19:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1035.eqiad.wmnet * 19:56 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2204 with weight 0 [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94891 and previous config saved to /var/cache/conftool/dbconfig/20260716-195628-marostegui.json * 19:55 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 26 hosts with reason: Primary switchover s2 [[phab:T432396|T432396]] * 19:54 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1035.eqiad.wmnet * 19:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1034.eqiad.wmnet * 19:48 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1034.eqiad.wmnet * 19:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1033.eqiad.wmnet * 19:43 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-video: apply * 19:43 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1033.eqiad.wmnet * 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1032.eqiad.wmnet * 19:42 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-video: apply * 19:42 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-timeline: apply * 19:41 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-timeline: apply * 19:41 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-syntaxhighlight: apply * 19:41 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-syntaxhighlight: apply * 19:40 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-media: apply * 19:40 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-media: apply * 19:39 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-constraints: apply * 19:36 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-constraints: apply * 19:36 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox: apply * 19:35 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1032.eqiad.wmnet * 19:35 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1031.eqiad.wmnet * 19:35 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox: apply * 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-video: apply * 19:33 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-video: apply * 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-timeline: apply * 19:33 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-timeline: apply * 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-syntaxhighlight: apply * 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-syntaxhighlight: apply * 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-media: apply * 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-media: apply * 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-constraints: apply * 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-constraints: apply * 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox: apply * 19:31 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox: apply * 19:27 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1031.eqiad.wmnet * 19:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1030.eqiad.wmnet * 19:23 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1001.eqiad.wmnet, repooling source-only afterwards * 19:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1030.eqiad.wmnet * 19:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1029.eqiad.wmnet * 19:17 kamila@cumin1003: START - Cookbook sre.dns.netbox * 19:12 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1029.eqiad.wmnet * 19:06 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1269 * 19:05 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1269.eqiad.wmnet with OS trixie * 19:03 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1269.eqiad.wmnet * 19:03 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1269.eqiad.wmnet * 19:03 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1269.eqiad.wmnet * 18:55 dancy@deploy2003: Finished scap sync-world: testing [[phab:T428971|T428971]] (duration: 02m 41s) * 18:53 dancy@deploy2003: Started scap sync-world: testing [[phab:T428971|T428971]] * 18:31 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1268.eqiad.wmnet with OS trixie * 18:18 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 18:16 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1267.eqiad.wmnet * 18:16 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1267.eqiad.wmnet * 18:16 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1267.eqiad.wmnet * 18:09 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1268.eqiad.wmnet with reason: host reimage * 18:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1001.eqiad.wmnet, repooling source-only afterwards * 18:06 swfrench@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] (duration: 07m 34s) * 18:06 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 23s) * 18:06 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 18:06 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1268.eqiad.wmnet with reason: host reimage * 18:03 bd808@deploy2003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 18:02 bd808@deploy2003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 18:02 swfrench@deploy2003: jiji, swfrench: Continuing with deployment * 18:02 bd808@deploy2003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 18:02 bd808@deploy2003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 18:01 bd808@deploy2003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 18:01 swfrench@deploy2003: jiji, swfrench: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:01 bd808@deploy2003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:59 swfrench@deploy2003: Started scap sync-world: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] * 17:45 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1268 * 17:45 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1268 * 17:44 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1267.eqiad.wmnet with OS trixie * 17:43 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter2006.codfw.wmnet * 17:39 swfrench@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter2006.codfw.wmnet * 17:35 swfrench@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] (duration: 07m 27s) * 17:34 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1268 * 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1268.eqiad.wmnet 78.32.64.10.in-addr.arpa 8.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:34 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1268.eqiad.wmnet 78.32.64.10.in-addr.arpa 8.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1268 - kamila@cumin1003" * 17:34 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1268 - kamila@cumin1003" * 17:31 swfrench@deploy2003: jiji, swfrench: Continuing with deployment * 17:29 swfrench@deploy2003: jiji, swfrench: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:28 kamila@cumin1003: START - Cookbook sre.dns.netbox * 17:28 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1268 * 17:28 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1268.eqiad.wmnet with OS trixie * 17:27 swfrench@deploy2003: Started scap sync-world: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] * 17:23 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1267.eqiad.wmnet with reason: host reimage * 17:18 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1267.eqiad.wmnet with reason: host reimage * 17:18 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1268.eqiad.wmnet * 17:17 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1268.eqiad.wmnet * 17:17 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1268.eqiad.wmnet * 17:12 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter2005.codfw.wmnet * 17:11 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1270.eqiad.wmnet * 17:11 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1270.eqiad.wmnet * 17:11 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1270.eqiad.wmnet * 17:09 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter2005.codfw.wmnet * 17:08 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:08 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update reverse dns for moved arelion cct cr2-eqiad - cmooney@cumin1003" * 17:08 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update reverse dns for moved arelion cct cr2-eqiad - cmooney@cumin1003" * 17:08 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] (duration: 07m 34s) * 17:04 jiji@deploy2003: jiji: Continuing with deployment * 17:03 jiji@deploy2003: jiji: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 17:00 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] * 17:00 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2185.codfw.wmnet with OS trixie * 16:59 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:58 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1270.eqiad.wmnet with OS trixie * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1267 * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1267 * 16:57 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1267 * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1267.eqiad.wmnet 77.32.64.10.in-addr.arpa 7.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:57 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1267.eqiad.wmnet 77.32.64.10.in-addr.arpa 7.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1267 - kamila@cumin1003" * 16:56 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1267 - kamila@cumin1003" * 16:56 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_eqsin * 16:56 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5032.eqsin.wmnet * 16:52 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_esams * 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3073.esams.wmnet * 16:50 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_esams * 16:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3081.esams.wmnet * 16:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1266.eqiad.wmnet * 16:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1266.eqiad.wmnet * 16:45 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1266.eqiad.wmnet * 16:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2185.codfw.wmnet with reason: host reimage * 16:41 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_eqiad * 16:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1114.eqiad.wmnet * 16:41 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_eqiad * 16:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1115.eqiad.wmnet * 16:39 kamila@cumin1003: START - Cookbook sre.dns.netbox * 16:39 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1267 * 16:39 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2185.codfw.wmnet with reason: host reimage * 16:38 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1267.eqiad.wmnet with OS trixie * 16:38 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1267.eqiad.wmnet * 16:38 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1270.eqiad.wmnet with reason: host reimage * 16:37 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1267.eqiad.wmnet * 16:37 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1267.eqiad.wmnet * 16:31 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1270.eqiad.wmnet with reason: host reimage * 16:24 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1264.eqiad.wmnet * 16:24 kamila@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 16:24 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 16:23 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 16:21 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2162: switch maintenance completed codfw rack b6 * 16:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2185.codfw.wmnet with OS trixie * 16:19 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 16:16 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter1007.eqiad.wmnet * 16:15 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_eqsin * 16:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5024.eqsin.wmnet * 16:13 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5031.eqsin.wmnet * 16:13 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3072.esams.wmnet * 16:12 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter1007.eqiad.wmnet * 16:11 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] (duration: 09m 47s) * 16:10 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1266.eqiad.wmnet with OS trixie * 16:10 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1270 * 16:10 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1270 * 16:09 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1270 * 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1270.eqiad.wmnet 125.48.64.10.in-addr.arpa 5.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1270.eqiad.wmnet 125.48.64.10.in-addr.arpa 5.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1270 - swfrench@cumin1003" * 16:09 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1270 - swfrench@cumin1003" * 16:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3080.esams.wmnet * 16:07 jiji@deploy2003: jiji: Continuing with deployment * 16:06 jiji@deploy2003: jiji: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:04 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 16:04 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1265.eqiad.wmnet * 16:03 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1265.eqiad.wmnet * 16:03 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1265.eqiad.wmnet * 16:03 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1270 * 16:03 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1270.eqiad.wmnet with OS trixie * 16:02 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1270.eqiad.wmnet * 16:02 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] * 16:01 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1270.eqiad.wmnet * 16:01 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1270.eqiad.wmnet * 16:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1113.eqiad.wmnet * 16:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1112.eqiad.wmnet * 15:49 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1266.eqiad.wmnet with reason: host reimage * 15:47 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter1006.eqiad.wmnet * 15:45 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1265.eqiad.wmnet with OS trixie * 15:44 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1266.eqiad.wmnet with reason: host reimage * 15:43 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter1006.eqiad.wmnet * 15:42 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] (duration: 09m 46s) * 15:37 jiji@deploy2003: jiji: Continuing with deployment * 15:36 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2162: switch maintenance completed codfw rack b6 * 15:36 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2161: switch maintenance completed codfw rack b6 * 15:34 jiji@deploy2003: jiji: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:32 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5023.eqsin.wmnet * 15:32 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] * 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5030.eqsin.wmnet * 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3071.esams.wmnet * 15:27 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3079.esams.wmnet * 15:25 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1265.eqiad.wmnet with reason: host reimage * 15:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1266 * 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1266 * 15:21 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1110.eqiad.wmnet * 15:20 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1111.eqiad.wmnet * 15:16 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1265.eqiad.wmnet with reason: host reimage * 15:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1001.eqiad.wmnet with OS bookworm * 15:15 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1266 * 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1266.eqiad.wmnet 76.32.64.10.in-addr.arpa 6.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:15 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1266.eqiad.wmnet 76.32.64.10.in-addr.arpa 6.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1266 - kamila@cumin1003" * 15:15 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1266 - kamila@cumin1003" * 15:07 kamila@cumin1003: START - Cookbook sre.dns.netbox * 15:04 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1266 * 15:04 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1264 * 15:04 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1264 * 15:04 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1266.eqiad.wmnet with OS trixie * 15:03 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1264 * 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1264.eqiad.wmnet 74.32.64.10.in-addr.arpa 4.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:03 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1264.eqiad.wmnet 74.32.64.10.in-addr.arpa 4.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1264 - kamila@cumin1003" * 15:03 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1264 - kamila@cumin1003" * 15:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-eqiad * 15:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp1001.eqiad.wmnet * 15:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp1001.eqiad.wmnet * 15:01 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp1001.eqiad.wmnet * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp1001.eqiad.wmnet * 15:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1376-1384].eqiad.wmnet * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1376-1384].eqiad.wmnet * 14:59 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1002.eqiad.wmnet * 14:59 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1266.eqiad.wmnet * 14:58 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1266.eqiad.wmnet * 14:58 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1266.eqiad.wmnet * 14:58 kamila@cumin1003: START - Cookbook sre.dns.netbox * 14:57 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1264 * 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1265 * 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1265 * 14:57 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1310596{{!}}Set $wgMathInternalRestbaseURL explicitly (T349582)]] * 14:57 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1265 * 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1265.eqiad.wmnet 75.32.64.10.in-addr.arpa 5.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:56 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1265.eqiad.wmnet 75.32.64.10.in-addr.arpa 5.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:56 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:56 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1265 - kamila@cumin1003" * 14:56 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1265 - kamila@cumin1003" * 14:53 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1002.eqiad.wmnet * 14:53 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1376-1384].eqiad.wmnet * 14:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1001.eqiad.wmnet with reason: host reimage * 14:51 kamila@cumin1003: START - Cookbook sre.dns.netbox * 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 14:50 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:50 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2161: switch maintenance completed codfw rack b6 * 14:50 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5021.eqsin.wmnet * 14:50 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1265 * 14:49 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:49 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:49 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1265.eqiad.wmnet with OS trixie * 14:49 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5029.eqsin.wmnet * 14:49 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1265.eqiad.wmnet * 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3070.esams.wmnet * 14:48 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1264.eqiad.wmnet * 14:48 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1001.eqiad.wmnet with reason: host reimage * 14:48 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1376-1384].eqiad.wmnet * 14:48 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1265.eqiad.wmnet * 14:47 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1265.eqiad.wmnet * 14:47 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1264.eqiad.wmnet * 14:47 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1264.eqiad.wmnet * 14:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:47 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3078.esams.wmnet * 14:44 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1263.eqiad.wmnet * 14:44 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1263.eqiad.wmnet * 14:44 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1263.eqiad.wmnet * 14:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1108.eqiad.wmnet * 14:40 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1109.eqiad.wmnet * 14:40 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:35 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:34 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:34 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2006.codfw.wmnet * 14:34 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-flink-eqiad cluster: Roll restart of jvm daemons. * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf2002.codfw.wmnet * 14:31 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1002.eqiad.wmnet * 14:29 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2006.codfw.wmnet * 14:27 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:27 kamila@deploy2003: Finished scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] (duration: 02m 57s) * 14:27 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-flink-eqiad cluster: Roll restart of jvm daemons. * 14:26 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf2002.codfw.wmnet * 14:26 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf2001.codfw.wmnet * 14:25 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1002.eqiad.wmnet * 14:25 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1001.eqiad.wmnet * 14:25 kamila@deploy2003: Started scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] * 14:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:21 kamila@deploy2003: sync-world aborted: Test deployment to check rsync is working - [[phab:T432108|T432108]] (duration: 00m 36s) * 14:21 topranks: reboot lsw1-b6-codfw to upgrade JunOS [[phab:T430922|T430922]] * 14:21 kamila@deploy2003: Started scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] * 14:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1001.eqiad.wmnet with OS bookworm * 14:20 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b6-codfw,lsw1-b6-codfw IPv6,lsw1-b6-codfw.mgmt,ssw1-a[1,8]-codfw with reason: lsw1-b6-codfw JunOS upgrade * 14:20 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf2001.codfw.wmnet * 14:19 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1001.eqiad.wmnet * 14:19 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 26 hosts with reason: lsw1-b6-codfw JunOS upgrade * 14:14 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 14:13 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc2022: switch maintenance codfw rack b6 * 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:12 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.parsercache * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool pc2022: switch maintenance codfw rack b6 * 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2251: switch maintenance codfw rack b6 * 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.parsercache * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2251: switch maintenance codfw rack b6 * 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2162: switch maintenance codfw rack b6 * 14:12 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1263.eqiad.wmnet with OS trixie * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2162: switch maintenance codfw rack b6 * 14:11 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2161: switch maintenance codfw rack b6 * 14:11 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2161: switch maintenance codfw rack b6 * 14:08 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5020.eqsin.wmnet * 14:07 btullis@cumin1003: START - Cookbook sre.hadoop.reboot-workers for Hadoop analytics cluster * 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3069.esams.wmnet * 14:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5028.eqsin.wmnet * 14:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1338-1347].eqiad.wmnet * 14:06 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1338-1347].eqiad.wmnet * 14:05 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3077.esams.wmnet * 14:02 topranks: beginning depools for lsw1-b6-codfw maintenance [[phab:T430922|T430922]] * 14:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1106.eqiad.wmnet * 14:00 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-misc1002.eqiad.wmnet * 13:59 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1338-1347].eqiad.wmnet * 13:59 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1107.eqiad.wmnet * 13:56 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-codfw * 13:55 sfaci@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply * 13:54 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-misc1002.eqiad.wmnet * 13:54 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-misc1001.eqiad.wmnet * 13:54 sfaci@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply * 13:50 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1263.eqiad.wmnet with reason: host reimage * 13:50 sfaci@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 13:49 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1338-1347].eqiad.wmnet * 13:49 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-misc1001.eqiad.wmnet * 13:49 sfaci@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 13:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:49 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:45 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1263.eqiad.wmnet with reason: host reimage * 13:40 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:40 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-eqiad * 13:35 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:34 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:33 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-reboot (exit_code=0) rolling reboot on A:dnsbox and (A:eqsin or A:drmrs or A:magru) and not (P<nowiki>{</nowiki>dns5003*<nowiki>}</nowiki> or P<nowiki>{</nowiki>dns7002*<nowiki>}</nowiki>) and (A:dnsbox) * 13:33 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns7001.wikimedia.org * 13:27 sukhe@dns1004: END - running authdns-update * 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5019.eqsin.wmnet * 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3076.esams.wmnet * 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3068.esams.wmnet * 13:25 sukhe@dns1004: START - running authdns-update * 13:24 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5027.eqsin.wmnet * 13:24 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1263 * 13:24 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1263 * 13:23 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1263 * 13:23 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:23 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1104.eqiad.wmnet * 13:21 kamila@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:21 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:20 kamila@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:20 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:20 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:20 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1263 - kamila@cumin1003" * 13:20 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1263 - kamila@cumin1003" * 13:19 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1105.eqiad.wmnet * 13:19 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:19 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:18 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:18 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns7001.wikimedia.org * 13:16 cdobbins@cumin2003: conftool action : set/pooled=yes; selector: name=dns7002.* * 13:14 cdobbins@dns1004: END - running authdns-update * 13:13 sbisson@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] (duration: 08m 03s) * 13:13 cdobbins@dns1004: START - running authdns-update * 13:12 kamila@cumin1003: START - Cookbook sre.dns.netbox * 13:12 cdobbins@cumin2003: conftool action : set/pooled=yes; selector: name=dns7002.*,service=authdns-update * 13:12 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1263 * 13:11 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1263.eqiad.wmnet with OS trixie * 13:11 cdobbins@cumin2003: conftool action : set/pooled=no; selector: name=dns7002.* * 13:11 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1263.eqiad.wmnet * 13:10 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1263.eqiad.wmnet * 13:10 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1263.eqiad.wmnet * 13:09 sbisson@deploy2003: sbisson: Continuing with deployment * 13:08 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:07 sbisson@deploy2003: sbisson: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:05 sbisson@deploy2003: Started scap sync-world: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] * 13:03 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns6002.wikimedia.org * 13:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1298-1307].eqiad.wmnet * 13:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1298-1307].eqiad.wmnet * 12:59 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1262.eqiad.wmnet * 12:59 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1262.eqiad.wmnet * 12:59 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1262.eqiad.wmnet * 12:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1298-1307].eqiad.wmnet * 12:49 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns6002.wikimedia.org * 12:46 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1298-1307].eqiad.wmnet * 12:46 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:46 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3075.esams.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3067.esams.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5018.eqsin.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5026.eqsin.wmnet * 12:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1102.eqiad.wmnet * 12:39 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1103.eqiad.wmnet * 12:35 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:34 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns6001.wikimedia.org * 12:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:28 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:18 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns6001.wikimedia.org * 12:14 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1267-1276].eqiad.wmnet * 12:13 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1267-1276].eqiad.wmnet * 12:04 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1267-1276].eqiad.wmnet * 12:03 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns5004.wikimedia.org * 12:02 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1100.eqiad.wmnet * 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3066.esams.wmnet * 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3074.esams.wmnet * 12:01 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:01 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-codfw * 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5017.eqsin.wmnet * 12:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5025.eqsin.wmnet * 12:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1101.eqiad.wmnet * 11:59 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1267-1276].eqiad.wmnet * 11:58 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:58 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:54 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns5004.wikimedia.org * 11:54 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and (A:eqsin or A:drmrs or A:magru) and not (P<nowiki>{</nowiki>dns5003*<nowiki>}</nowiki> or P<nowiki>{</nowiki>dns7002*<nowiki>}</nowiki>) and (A:dnsbox) * 11:54 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:53 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:53 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-eqiad * 11:51 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:51 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:50 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:50 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_eqiad * 11:50 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:50 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_eqiad * 11:49 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_esams * 11:49 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_esams * 11:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_eqsin * 11:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_eqsin * 11:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:44 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-eqiad * 11:43 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-codfw * 11:42 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:41 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:24 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-eqiad * 11:23 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-codfw * 11:22 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:15 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:14 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:09 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2066.codfw.wmnet with OS trixie * 11:05 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1151-1160].eqiad.wmnet * 11:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1151-1160].eqiad.wmnet * 10:59 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1068.eqiad.wmnet with OS trixie * 10:55 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.major-upgrade (exit_code=99) * 10:55 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 10:54 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1151-1160].eqiad.wmnet * 10:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2066.codfw.wmnet with reason: host reimage * 10:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1151-1160].eqiad.wmnet * 10:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:42 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2066.codfw.wmnet with reason: host reimage * 10:39 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:37 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:36 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:23 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:22 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2066.codfw.wmnet with OS trixie * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:07 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2065.codfw.wmnet with OS trixie * 10:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:06 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 10:06 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 10:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 10:03 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:03 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 10:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 09:59 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 09:57 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 09:57 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 09:52 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 09:47 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:46 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:46 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2065.codfw.wmnet with reason: host reimage * 09:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2065.codfw.wmnet with reason: host reimage * 09:40 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:39 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 09:39 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:39 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:39 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:37 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:29 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox2003.codfw.wmnet * 09:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox2003.codfw.wmnet * 09:25 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:25 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:24 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:24 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:24 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:21 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:20 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2065.codfw.wmnet with OS trixie * 09:13 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1068.eqiad.wmnet with OS trixie * 09:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2064.codfw.wmnet with OS trixie * 09:08 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 09:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 09:07 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2162: Repooling after switchover * 09:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1067.eqiad.wmnet with OS trixie * 09:00 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:59 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 08:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 08:57 tappof: bump space for prometheus k8s-dse in eqiad * 08:56 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping2004.codfw.wmnet * 08:52 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host ping2004.codfw.wmnet * 08:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 08:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping1004.eqiad.wmnet * 08:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:51 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:49 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 08:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host ping1004.eqiad.wmnet * 08:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2064.codfw.wmnet with reason: host reimage * 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2064.codfw.wmnet with reason: host reimage * 08:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1067.eqiad.wmnet with reason: host reimage * 08:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:33 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1067.eqiad.wmnet with reason: host reimage * 08:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:21 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2162: Repooling after switchover * 08:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2064.codfw.wmnet with OS trixie * 08:16 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1067.eqiad.wmnet with OS trixie * 08:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:15 cgoubert@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-eqiad * 08:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2062.codfw.wmnet with OS trixie * 08:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1066.eqiad.wmnet with OS trixie * 08:02 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2162: Repooling after switchover * 07:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2162: Repooling after switchover * 07:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2162 [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94870 and previous config saved to /var/cache/conftool/dbconfig/20260716-075530-cwilliams.json * 07:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2241 to x3 primary [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94869 and previous config saved to /var/cache/conftool/dbconfig/20260716-075314-cwilliams.json * 07:52 cezmunsta: Starting x3 codfw failover from db2162 to db2241 - [[phab:T430925|T430925]] * 07:50 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:50 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:47 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 07:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2241 with weight 0 [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94868 and previous config saved to /var/cache/conftool/dbconfig/20260716-074507-cwilliams.json * 07:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 18 hosts with reason: Primary switchover x3 [[phab:T430925|T430925]] * 07:43 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1066.eqiad.wmnet with reason: host reimage * 07:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:dse-k8s-worker-eqiad * 07:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1028.eqiad.wmnet * 07:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1028.eqiad.wmnet * 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 07:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1066.eqiad.wmnet with reason: host reimage * 07:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1028.eqiad.wmnet * 07:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1028.eqiad.wmnet * 07:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1027.eqiad.wmnet * 07:35 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1027.eqiad.wmnet * 07:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1027.eqiad.wmnet * 07:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1027.eqiad.wmnet * 07:28 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1026.eqiad.wmnet * 07:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1026.eqiad.wmnet * 07:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast2003.wikimedia.org * 07:21 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1026.eqiad.wmnet * 07:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1066.eqiad.wmnet with OS trixie * 07:19 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast2003.wikimedia.org * 07:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2062.codfw.wmnet with OS trixie * 06:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1026.eqiad.wmnet * 06:51 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1025.eqiad.wmnet * 06:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1025.eqiad.wmnet * 06:47 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 06:47 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 06:44 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1025.eqiad.wmnet * 06:14 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1025.eqiad.wmnet * 06:14 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1024.eqiad.wmnet * 06:14 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1024.eqiad.wmnet * 06:07 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1024.eqiad.wmnet * 05:37 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1024.eqiad.wmnet * 05:37 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1023.eqiad.wmnet * 05:37 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1023.eqiad.wmnet * 05:26 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1023.eqiad.wmnet * 04:56 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1023.eqiad.wmnet * 04:56 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1022.eqiad.wmnet * 04:56 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1022.eqiad.wmnet * 04:49 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1022.eqiad.wmnet * 04:19 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1022.eqiad.wmnet * 04:19 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1021.eqiad.wmnet * 04:19 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1021.eqiad.wmnet * 04:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1021.eqiad.wmnet * 03:38 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1021.eqiad.wmnet * 03:38 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1020.eqiad.wmnet * 03:38 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1020.eqiad.wmnet * 03:20 btullis@cumin1003: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1020.eqiad.wmnet * 03:18 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1020.eqiad.wmnet * 03:18 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1019.eqiad.wmnet * 03:18 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1019.eqiad.wmnet * 03:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1019.eqiad.wmnet * 02:41 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1019.eqiad.wmnet * 02:41 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1018.eqiad.wmnet * 02:41 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1018.eqiad.wmnet * 02:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling both afterwards * 02:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2003.codfw.wmnet -> wcqs2001.codfw.wmnet, repooling both afterwards * 02:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1018.eqiad.wmnet * 02:30 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1018.eqiad.wmnet * 02:30 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1014.eqiad.wmnet * 02:30 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1014.eqiad.wmnet * 02:24 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1014.eqiad.wmnet * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 01:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1014.eqiad.wmnet * 01:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1013.eqiad.wmnet * 01:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1013.eqiad.wmnet * 01:47 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1013.eqiad.wmnet * 01:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2003.codfw.wmnet -> wcqs2001.codfw.wmnet, repooling both afterwards * 01:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling both afterwards * 01:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1013.eqiad.wmnet * 01:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1012.eqiad.wmnet * 01:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1012.eqiad.wmnet * 01:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1012.eqiad.wmnet * 01:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1012.eqiad.wmnet * 01:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1011.eqiad.wmnet * 01:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1011.eqiad.wmnet * 01:08 ryankemper@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] scap deploy post bookworm reimage (duration: 00m 23s) * 01:08 ryankemper@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] scap deploy post bookworm reimage * 01:08 ryankemper@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): scap deploy post bookworm reimage (duration: 00m 46s) * 01:07 ryankemper@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): scap deploy post bookworm reimage * 01:04 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1011.eqiad.wmnet * 01:04 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1011.eqiad.wmnet * 01:04 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1010.eqiad.wmnet * 01:04 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1010.eqiad.wmnet * 00:57 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1010.eqiad.wmnet * 00:57 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1010.eqiad.wmnet * 00:57 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1009.eqiad.wmnet * 00:57 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1009.eqiad.wmnet * 00:50 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1009.eqiad.wmnet * 00:20 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1009.eqiad.wmnet * 00:20 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1008.eqiad.wmnet * 00:20 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1008.eqiad.wmnet * 00:13 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1008.eqiad.wmnet == 2026-07-15 == * 23:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2001.codfw.wmnet with OS bookworm * 23:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1008.eqiad.wmnet * 23:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1007.eqiad.wmnet * 23:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1007.eqiad.wmnet * 23:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1007.eqiad.wmnet * 23:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1007.eqiad.wmnet * 23:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1006.eqiad.wmnet * 23:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1006.eqiad.wmnet * 23:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1006.eqiad.wmnet * 23:29 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1006.eqiad.wmnet * 23:28 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1005.eqiad.wmnet * 23:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1005.eqiad.wmnet * 23:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1002.eqiad.wmnet with OS bookworm * 23:21 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1005.eqiad.wmnet * 23:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2001.codfw.wmnet with reason: host reimage * 23:15 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host datahubsearch1001.eqiad.wmnet with OS bookworm * 23:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2001.codfw.wmnet with reason: host reimage * 23:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 23:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 22:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 22:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1005.eqiad.wmnet * 22:51 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1004.eqiad.wmnet * 22:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1004.eqiad.wmnet * 22:45 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1004.eqiad.wmnet * 22:44 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 22:44 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS trixie * 22:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host datahubsearch1001.eqiad.wmnet with OS bookworm * 22:34 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host datahubsearch1001.eqiad.wmnet with OS bookworm * 22:16 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on datahubsearch[1002-1003].eqiad.wmnet with reason: Using datahubsearch1001 to test bookworm reimages * 22:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1004.eqiad.wmnet * 22:15 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1003.eqiad.wmnet * 22:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1003.eqiad.wmnet * 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 22:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1003.eqiad.wmnet * 22:08 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1003.eqiad.wmnet * 22:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1002.eqiad.wmnet * 22:08 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1002.eqiad.wmnet * 22:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host datahubsearch1001.eqiad.wmnet with OS bookworm * 22:05 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 22:02 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm * 22:01 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on datahubsearch[1001-1003].eqiad.wmnet with reason: Using datahubsearch1001 to test bookworm reimages * 22:01 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1002.eqiad.wmnet * 22:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1002.eqiad.wmnet * 22:00 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1001.eqiad.wmnet * 22:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1001.eqiad.wmnet * 21:53 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1001.eqiad.wmnet * 21:52 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 21:50 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 21:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS trixie * 21:50 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS bookworm * 21:43 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 21:38 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:30 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wcqs1002'] * 21:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:29 lerickson@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 21:29 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:29 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:29 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS bookworm * 21:28 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 21:28 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm * 21:23 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1001.eqiad.wmnet * 21:23 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:23 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:22 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 21:20 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:18 swfrench-wmf: reprepro include php8.3_8.3.32-1+wmf11u2 into component/php83 for bullseye-wikimedia * 21:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:16 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:15 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-druid-public cluster: Roll restart of jvm daemons. * 21:08 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:05 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1001.eqiad.wmnet * 21:05 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1001.eqiad.wmnet * 21:04 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-druid-public cluster: Roll restart of jvm daemons. * 21:02 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 21:01 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 21:01 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 21:00 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 20:59 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1001.eqiad.wmnet * 20:59 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1001.eqiad.wmnet * 20:59 btullis@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:dse-k8s-worker-eqiad * 20:55 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 20:55 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 20:45 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm * 20:21 jhathaway: puppet is re-enabled, have fun, but not too much fun! * 20:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 20:17 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs2001'] * 20:12 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs2001'] * 20:11 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs2001'] * 20:09 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:08 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 20:05 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:05 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 20:04 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs2001'] * 20:03 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 20:03 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm * 20:02 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:02 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 20:01 jhathaway: disabling puppet fleet wide to roll out kafka patch * 19:55 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host relforge1010.eqiad.wmnet * 19:52 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 19:52 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 19:48 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 19:48 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 19:48 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 19:47 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 19:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 19:45 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 19:45 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 19:44 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1010.eqiad.wmnet * 19:38 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1262.eqiad.wmnet with OS trixie * 19:17 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1262.eqiad.wmnet with reason: host reimage * 19:11 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1262.eqiad.wmnet with reason: host reimage * 18:59 cdobbins@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS trixie * 18:54 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 18:53 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 18:52 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1262 * 18:52 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1262 * 18:51 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1262 * 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1262.eqiad.wmnet 72.32.64.10.in-addr.arpa 2.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:51 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1262.eqiad.wmnet 72.32.64.10.in-addr.arpa 2.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1262 - kamila@cumin1003" * 18:51 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1262 - kamila@cumin1003" * 18:46 kamila@cumin1003: START - Cookbook sre.dns.netbox * 18:46 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1262 * 18:46 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 18:46 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ncmonitor1001.eqiad.wmnet * 18:46 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 18:45 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1262.eqiad.wmnet with OS trixie * 18:45 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 18:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1262.eqiad.wmnet * 18:44 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1262.eqiad.wmnet * 18:44 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1262.eqiad.wmnet * 18:42 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host ncmonitor1001.eqiad.wmnet * 18:29 topranks: pull power on cr1-eqiad to install new switch-control boards [[phab:T426343|T426343]] * 18:29 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs[1018-1020].eqiad.wmnet with reason: line card install in cr1-eqiad * 18:27 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 14 hosts with reason: linecard install in cr1-eqad * 18:22 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_ulsfo * 18:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4052.ulsfo.wmnet * 18:19 cdobbins@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 18:15 cdobbins@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 18:14 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_drmrs * 18:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6016.drmrs.wmnet * 18:12 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_ulsfo * 18:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4044.ulsfo.wmnet * 18:10 sukhe@cumin1003: END (ERROR) - Cookbook sre.cdn.roll-reboot (exit_code=97) rolling reboot on A:cp-upload_drmrs * 18:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2241: Security update * 17:56 topranks: start draining traffic on cr1-eqiad ahead of line card installation [[phab:T426343|T426343]] * 17:47 cdobbins@cumin2003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie * 17:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4051.ulsfo.wmnet * 17:40 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:39 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 17:34 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6007.drmrs.wmnet * 17:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6015.drmrs.wmnet * 17:32 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:31 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 17:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4043.ulsfo.wmnet * 17:27 lerickson@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:25 lerickson@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 17:22 lerickson@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-codfw * 17:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp2001.codfw.wmnet * 17:22 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 17:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp2001.codfw.wmnet * 17:22 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 17:19 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2241: Security update * 17:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2241.codfw.wmnet * 17:17 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2241.codfw.wmnet * 17:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp2001.codfw.wmnet * 17:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp2001.codfw.wmnet * 17:15 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2366-2374].codfw.wmnet * 17:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2366-2374].codfw.wmnet * 17:10 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply * 17:10 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply * 17:08 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2366-2374].codfw.wmnet * 17:06 sukhe: sre.dns.roll-reboot to resume later * 17:06 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-reboot (exit_code=97) rolling reboot on A:dnsbox and not (A:ulsfo or A:magru) and (A:dnsbox) * 17:06 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns5003.wikimedia.org * 17:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2241: Security update * 17:03 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2241: Security update * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply * 17:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2366-2374].codfw.wmnet * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply * 17:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2357-2365].codfw.wmnet * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 17:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2357-2365].codfw.wmnet * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply * 16:55 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2357-2365].codfw.wmnet * 16:55 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 16:53 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6006.drmrs.wmnet * 16:52 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6014.drmrs.wmnet * 16:52 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 16:51 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 16:50 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2357-2365].codfw.wmnet * 16:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4042.ulsfo.wmnet * 16:50 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2347-2356].codfw.wmnet * 16:50 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2347-2356].codfw.wmnet * 16:49 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns5003.wikimedia.org * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply * 16:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4050.ulsfo.wmnet * 16:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2347-2356].codfw.wmnet * 16:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2347-2356].codfw.wmnet * 16:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2337-2346].codfw.wmnet * 16:36 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2337-2346].codfw.wmnet * 16:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:dse-k8s-worker-codfw * 16:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2003.codfw.wmnet * 16:35 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2003.codfw.wmnet * 16:34 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns3004.wikimedia.org * 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply * 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply * 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply * 16:30 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply * 16:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2003.codfw.wmnet * 16:29 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2337-2346].codfw.wmnet * 16:24 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2003.codfw.wmnet * 16:24 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2002.codfw.wmnet * 16:24 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2002.codfw.wmnet * 16:23 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns3004.wikimedia.org * 16:23 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2337-2346].codfw.wmnet * 16:23 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2327-2336].codfw.wmnet * 16:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2327-2336].codfw.wmnet * 16:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2002.codfw.wmnet * 16:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2327-2336].codfw.wmnet * 16:12 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2002.codfw.wmnet * 16:12 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2001.codfw.wmnet * 16:12 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2001.codfw.wmnet * 16:12 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1065.eqiad.wmnet with OS trixie * 16:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6005.drmrs.wmnet * 16:11 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6013.drmrs.wmnet * 16:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4041.ulsfo.wmnet * 16:08 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns3003.wikimedia.org * 16:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2327-2336].codfw.wmnet * 16:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2317-2326].codfw.wmnet * 16:06 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2317-2326].codfw.wmnet * 16:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2001.codfw.wmnet * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply * 16:03 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4049.ulsfo.wmnet * 16:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2001.codfw.wmnet * 16:00 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test2001.codfw.wmnet * 16:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test2001.codfw.wmnet * 16:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2063.codfw.wmnet with OS trixie * 15:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2317-2326].codfw.wmnet * 15:57 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns3003.wikimedia.org * 15:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test2001.codfw.wmnet * 15:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test2001.codfw.wmnet * 15:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2004.codfw.wmnet * 15:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2004.codfw.wmnet * 15:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2317-2326].codfw.wmnet * 15:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2307-2316].codfw.wmnet * 15:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2307-2316].codfw.wmnet * 15:49 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2004.codfw.wmnet * 15:48 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2004.codfw.wmnet * 15:48 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2003.codfw.wmnet * 15:48 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2003.codfw.wmnet * 15:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 15:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2307-2316].codfw.wmnet * 15:42 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2003.codfw.wmnet * 15:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 15:42 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2003.codfw.wmnet * 15:42 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2002.codfw.wmnet * 15:42 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2002.codfw.wmnet * 15:42 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2006.wikimedia.org * 15:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2063.codfw.wmnet with reason: host reimage * 15:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2307-2316].codfw.wmnet * 15:37 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2297-2306].codfw.wmnet * 15:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2297-2306].codfw.wmnet * 15:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2002.codfw.wmnet * 15:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2002.codfw.wmnet * 15:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2001.codfw.wmnet * 15:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2001.codfw.wmnet * 15:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2063.codfw.wmnet with reason: host reimage * 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6004.drmrs.wmnet * 15:31 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2001.codfw.wmnet * 15:31 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2001.codfw.wmnet * 15:31 btullis@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:dse-k8s-worker-codfw * 15:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6012.drmrs.wmnet * 15:28 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2006.wikimedia.org * 15:27 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-analytics cluster: Roll restart of jvm daemons. * 15:27 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2297-2306].codfw.wmnet * 15:27 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4040.ulsfo.wmnet * 15:24 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1065.eqiad.wmnet with OS trixie * 15:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4048.ulsfo.wmnet * 15:21 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-analytics cluster: Roll restart of jvm daemons. * 15:21 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2297-2306].codfw.wmnet * 15:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2287-2296].codfw.wmnet * 15:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2287-2296].codfw.wmnet * 15:20 btullis@cumin1003: END (PASS) - Cookbook sre.druid.reboot-workers (exit_code=0) for Druid public cluster: Reboot Druid nodes * 15:18 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 15:17 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm * 15:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2063.codfw.wmnet with OS trixie * 15:13 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2005.wikimedia.org * 15:11 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2287-2296].codfw.wmnet * 15:11 btullis@cumin1003: END (PASS) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=0) rolling reboot on A:cephosd-eqiad * 15:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1064.eqiad.wmnet with OS trixie * 15:05 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2062.codfw.wmnet with OS trixie * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2287-2296].codfw.wmnet * 15:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2277-2286].codfw.wmnet * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2277-2286].codfw.wmnet * 14:59 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2005.wikimedia.org * 14:57 brouberol@cumin1003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-jumbo-eqiad * 14:52 btullis@cumin1003: END (PASS) - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas (exit_code=0) rolling reboot on A:schema-codfw * 14:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6003.drmrs.wmnet * 14:50 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:50 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host relforge1009.eqiad.wmnet * 14:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2277-2286].codfw.wmnet * 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6011.drmrs.wmnet * 14:47 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 14:46 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>ml-serve1001.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 14:46 klausman@cumin1003: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) pool for host ml-serve1001.eqiad.wmnet * 14:46 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 14:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1001.eqiad.wmnet * 14:45 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4039.ulsfo.wmnet * 14:44 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1009.eqiad.wmnet * 14:44 btullis@cumin1003: START - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas rolling reboot on A:schema-codfw * 14:44 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2004.wikimedia.org * 14:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2277-2286].codfw.wmnet * 14:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2267-2276].codfw.wmnet * 14:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2267-2276].codfw.wmnet * 14:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 14:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4047.ulsfo.wmnet * 14:40 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1001.eqiad.wmnet * 14:38 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 14:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 14:36 topranks: disconnect power on cr2-eqiad to shut down device for switch fabric replacement [[phab:T426343|T426343]] * 14:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2267-2276].codfw.wmnet * 14:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 14:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1001.eqiad.wmnet * 14:35 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>ml-serve1001.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 14:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 14:34 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 14:33 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:33 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:30 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2004.wikimedia.org * 14:29 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2267-2276].codfw.wmnet * 14:29 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2257-2266].codfw.wmnet * 14:29 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2257-2266].codfw.wmnet * 14:24 btullis@cumin1003: END (PASS) - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas (exit_code=0) rolling reboot on A:schema-eqiad * 14:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2257-2266].codfw.wmnet * 14:20 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:20 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:19 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:17 jforrester@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2257-2266].codfw.wmnet * 14:16 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:16 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2062.codfw.wmnet with OS trixie * 14:15 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1064.eqiad.wmnet with OS trixie * 14:15 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1006.wikimedia.org * 14:15 btullis@cumin1003: START - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas rolling reboot on A:schema-eqiad * 14:14 topranks: switch routing-engine on cr2-eqiad resetting all interfaces [[phab:T417873|T417873]] * 14:11 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:11 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:10 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm * 14:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6002.drmrs.wmnet * 14:09 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6010.drmrs.wmnet * 14:06 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1006.wikimedia.org * 14:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:05 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 14:05 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4038.ulsfo.wmnet * 14:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4046.ulsfo.wmnet * 14:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:00 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on cr1-eqiad with reason: switch upgrade and line card install * 14:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:59 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:57 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:57 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:55 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-eqiad * 13:55 btullis@cumin1003: START - Cookbook sre.druid.reboot-workers for Druid public cluster: Reboot Druid nodes * 13:53 topranks: switch routing-engine on cr2-eqiad resetting all interfaces [[phab:T417873|T417873]] * 13:51 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1005.wikimedia.org * 13:50 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:49 brouberol@cumin1003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-test-eqiad * 13:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:44 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2197-2206].codfw.wmnet * 13:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2197-2206].codfw.wmnet * 13:36 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1005.wikimedia.org * 13:35 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2197-2206].codfw.wmnet * 13:30 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2197-2206].codfw.wmnet * 13:28 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6001.drmrs.wmnet * 13:28 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6009.drmrs.wmnet * 13:28 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2187-2196].codfw.wmnet * 13:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2187-2196].codfw.wmnet * 13:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 13:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4037.ulsfo.wmnet * 13:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2001 * 13:22 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2001 * 13:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4045.ulsfo.wmnet * 13:21 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1004.wikimedia.org * 13:19 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on lvs[1018-1020].eqiad.wmnet with reason: switch upgrade and line card install * 13:18 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2009.codfw.wmnet * 13:18 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2001 * 13:18 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2001.codfw.wmnet 26.16.192.10.in-addr.arpa 6.2.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:17 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2001.codfw.wmnet 26.16.192.10.in-addr.arpa 6.2.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:17 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:17 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2001 - bking@cumin2003" * 13:17 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2001 - bking@cumin2003" * 13:17 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2009.codfw.wmnet * 13:17 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_drmrs * 13:17 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2187-2196].codfw.wmnet * 13:17 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_drmrs * 13:17 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on 15 hosts with reason: switch upgrade and line card install * 13:17 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:15 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 13:13 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:13 brouberol@cumin1003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-jumbo-eqiad * 13:13 brouberol@cumin1003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-test-eqiad * 13:13 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1004.wikimedia.org * 13:13 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and not (A:ulsfo or A:magru) and (A:dnsbox) * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:12 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_ulsfo * 13:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:12 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_ulsfo * 13:11 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2187-2196].codfw.wmnet * 13:11 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 13:11 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 13:06 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 13:05 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling source-only afterwards * 13:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2001 * 13:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2009.codfw.wmnet with OS trixie * 13:03 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:03 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling source-only afterwards * 13:01 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 15s) * 13:01 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 13:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 12:57 btullis@cumin1003: END (PASS) - Cookbook sre.druid.reboot-workers (exit_code=0) for Druid analytics cluster: Reboot Druid nodes * 12:54 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 12:54 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2163-2172].codfw.wmnet * 12:54 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2163-2172].codfw.wmnet * 12:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2163-2172].codfw.wmnet * 12:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2009.codfw.wmnet with reason: host reimage * 12:41 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2163-2172].codfw.wmnet * 12:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2153-2162].codfw.wmnet * 12:40 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2153-2162].codfw.wmnet * 12:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2009.codfw.wmnet with reason: host reimage * 12:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2153-2162].codfw.wmnet * 12:29 btullis@cumin1003: END (PASS) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=0) rolling reboot on A:cephosd-codfw * 12:25 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2153-2162].codfw.wmnet * 12:25 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2143-2152].codfw.wmnet * 12:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2143-2152].codfw.wmnet * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2009 * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2009 * 12:22 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2009 * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2009.codfw.wmnet 139.0.192.10.in-addr.arpa 9.3.1.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:22 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2009.codfw.wmnet 139.0.192.10.in-addr.arpa 9.3.1.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2009 - mvernon@cumin2003" * 12:22 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2009 - mvernon@cumin2003" * 12:16 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 12:15 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 12:15 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 12:15 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2009 * 12:15 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 12:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2009.codfw.wmnet with OS trixie * 12:15 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 12:14 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 12:14 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2143-2152].codfw.wmnet * 12:13 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 12:12 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2010.codfw.wmnet * 12:11 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2010.codfw.wmnet * 12:10 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 12:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2143-2152].codfw.wmnet * 12:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2133-2142].codfw.wmnet * 12:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2133-2142].codfw.wmnet * 12:02 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 11:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2133-2142].codfw.wmnet * 11:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2133-2142].codfw.wmnet * 11:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:49 mvolz@deploy2003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:49 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-codfw * 11:48 mvolz@deploy2003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:47 btullis@cumin1003: START - Cookbook sre.druid.reboot-workers for Druid analytics cluster: Reboot Druid nodes * 11:46 mvolz@deploy2003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:46 mvolz@deploy2003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:45 mvolz@deploy2003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:44 mvolz@deploy2003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:40 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] (duration: 11m 38s) * 11:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2010.codfw.wmnet with OS trixie * 11:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2105-2114].codfw.wmnet * 11:36 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2105-2114].codfw.wmnet * 11:36 krinkle@deploy2003: physikerwelt, krinkle: Continuing with deployment * 11:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1018: Security updates * 11:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:36 root@cumin1003: START - Cookbook sre.mysql.parsercache * 11:36 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1018: Security updates * 11:31 krinkle@deploy2003: physikerwelt, krinkle: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:29 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] * 11:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2105-2114].codfw.wmnet * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2105-2114].codfw.wmnet * 11:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2010.codfw.wmnet with reason: host reimage * 11:12 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2010.codfw.wmnet with reason: host reimage * 11:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1018: Security updates * 11:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:10 root@cumin1003: START - Cookbook sre.mysql.parsercache * 11:10 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1018: Security updates * 11:09 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1009.eqiad.wmnet with OS trixie * 11:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow7002.magru.wmnet * 11:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 11:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 11:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=tegola-vector-tiles,name=eqiad * 11:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=kartotherian,name=eqiad * 11:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow7002.magru.wmnet * 10:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2010 * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2010 * 10:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 10:54 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 10:54 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2010 * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2010.codfw.wmnet 76.16.192.10.in-addr.arpa 6.7.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:54 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2010.codfw.wmnet 76.16.192.10.in-addr.arpa 6.7.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2010 - mvernon@cumin2003" * 10:54 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2010 - mvernon@cumin2003" * 10:49 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 10:49 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2010 * 10:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1009.eqiad.wmnet with reason: host reimage * 10:49 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2010.codfw.wmnet with OS trixie * 10:46 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2011.codfw.wmnet * 10:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1011.eqiad.wmnet * 10:44 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2011.codfw.wmnet * 10:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 10:44 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 10:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1009.eqiad.wmnet with reason: host reimage * 10:44 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow6001.drmrs.wmnet * 10:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1017: Security updates * 10:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:39 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1017: Security updates * 10:39 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow6001.drmrs.wmnet * 10:38 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1011.eqiad.wmnet * 10:35 cgoubert@deploy2003: Finished deploy [restbase/deploy@06301bd]: Deploying {{Gerrit|1306088}} {{Gerrit|1308347}} - [[phab:T429944|T429944]] [[phab:T428279|T428279]] (duration: 28m 34s) * 10:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1012.eqiad.wmnet * 10:35 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 10:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow5003.eqsin.wmnet * 10:34 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2011.codfw.wmnet with OS trixie * 10:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1009.eqiad.wmnet with OS trixie * 10:28 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1012.eqiad.wmnet * 10:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1013.eqiad.wmnet * 10:27 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:27 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow5003.eqsin.wmnet * 10:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:26 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow4003.ulsfo.wmnet * 10:25 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 10:25 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 10:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow4003.ulsfo.wmnet * 10:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1013.eqiad.wmnet * 10:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1014.eqiad.wmnet * 10:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:15 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2011.codfw.wmnet with reason: host reimage * 10:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1017: Security updates * 10:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:14 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:14 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1017: Security updates * 10:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow3004.esams.wmnet * 10:11 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2011.codfw.wmnet with reason: host reimage * 10:10 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:10 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1014.eqiad.wmnet * 10:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki2003.codfw.wmnet * 10:09 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 10:09 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 10:09 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow3004.esams.wmnet * 10:08 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2004.codfw.wmnet * 10:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1010.eqiad.wmnet with OS trixie * 10:07 cgoubert@deploy2003: Started deploy [restbase/deploy@06301bd]: Deploying {{Gerrit|1306088}} {{Gerrit|1308347}} - [[phab:T429944|T429944]] [[phab:T428279|T428279]] * 10:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host rpki2003.codfw.wmnet * 10:04 topranks: push out config change to BGP_outfilter on core routers [[phab:T431849|T431849]] * 10:02 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow2004.codfw.wmnet * 09:59 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 09:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2003.codfw.wmnet * 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2011 * 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2011 * 09:53 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 09:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:52 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2011 * 09:52 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2011.codfw.wmnet 36.32.192.10.in-addr.arpa 6.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:52 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2011.codfw.wmnet 36.32.192.10.in-addr.arpa 6.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:51 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:51 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2011 - mvernon@cumin2003" * 09:51 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2011 - mvernon@cumin2003" * 09:51 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow2003.codfw.wmnet * 09:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1003.eqiad.wmnet * 09:49 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 09:49 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 09:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1010.eqiad.wmnet with reason: host reimage * 09:47 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 09:47 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 09:47 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 09:47 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2011 * 09:46 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2011.codfw.wmnet with OS trixie * 09:44 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow1003.eqiad.wmnet * 09:44 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2012.codfw.wmnet * 09:44 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1002.eqiad.wmnet * 09:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1010.eqiad.wmnet with reason: host reimage * 09:43 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2012.codfw.wmnet * 09:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Security updates * 09:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:43 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:43 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Security updates * 09:42 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:40 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow1002.eqiad.wmnet * 09:40 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 09:37 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki1001.eqiad.wmnet * 09:36 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 09:36 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 09:33 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host rpki1001.eqiad.wmnet * 09:32 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:32 cgoubert@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-codfw * 09:31 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=kartotherian,name=eqiad * 09:31 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola-vector-tiles,name=eqiad * 09:31 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 09:31 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2012.codfw.wmnet with OS trixie * 09:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1010.eqiad.wmnet with OS trixie * 09:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1011.eqiad.wmnet with OS trixie * 09:21 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Security updates * 09:21 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:21 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:21 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Security updates * 09:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2012.codfw.wmnet with reason: host reimage * 09:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1011.eqiad.wmnet with reason: host reimage * 09:08 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2012.codfw.wmnet with reason: host reimage * 09:05 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1011.eqiad.wmnet with reason: host reimage * 08:55 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:52 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1011.eqiad.wmnet with OS trixie * 08:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1022: Security updates * 08:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2012 * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2012 * 08:50 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1022: Security updates * 08:50 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2012 * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2012.codfw.wmnet 44.48.192.10.in-addr.arpa 4.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:50 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2012.codfw.wmnet 44.48.192.10.in-addr.arpa 4.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2012 - mvernon@cumin2003" * 08:50 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2012 - mvernon@cumin2003" * 08:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1012.eqiad.wmnet with OS trixie * 08:44 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 08:44 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2012 * 08:43 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2012.codfw.wmnet with OS trixie * 08:42 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2013.codfw.wmnet * 08:41 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2013.codfw.wmnet * 08:35 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 08:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host krb1002.eqiad.wmnet * 08:30 elukey@dns1004: END - running authdns-update * 08:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1012.eqiad.wmnet with reason: host reimage * 08:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Security updates * 08:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:28 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:28 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Security updates * 08:27 elukey@dns1004: START - running authdns-update * 08:26 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 08:26 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host krb1002.eqiad.wmnet * 08:22 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1012.eqiad.wmnet with reason: host reimage * 08:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host krb2002.codfw.wmnet * 08:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast6003.wikimedia.org * 08:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2013.codfw.wmnet with OS trixie * 08:13 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast6003.wikimedia.org * 08:12 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast3007.wikimedia.org * 08:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host krb2002.codfw.wmnet * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Security updates * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:09 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:09 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Security updates * 08:07 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1012.eqiad.wmnet with OS trixie * 08:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast3007.wikimedia.org * 08:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast5005.wikimedia.org * 07:58 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast5005.wikimedia.org * 07:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1013.eqiad.wmnet with OS trixie * 07:53 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2013.codfw.wmnet with reason: host reimage * 07:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1021: Security updates * 07:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:53 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:53 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1021: Security updates * 07:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast1004.wikimedia.org * 07:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2013.codfw.wmnet with reason: host reimage * 07:46 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast1004.wikimedia.org * 07:40 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1013.eqiad.wmnet with reason: host reimage * 07:36 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1013.eqiad.wmnet with reason: host reimage * 07:31 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2013 * 07:31 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2013 * 07:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1021: Security updates * 07:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:30 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:30 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1021: Security updates * 07:24 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2013 * 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2013.codfw.wmnet 87.0.192.10.in-addr.arpa 7.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:24 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2013.codfw.wmnet 87.0.192.10.in-addr.arpa 7.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2013 - mvernon@cumin2003" * 07:24 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2013 - mvernon@cumin2003" * 07:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1013.eqiad.wmnet with OS trixie * 07:19 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 07:19 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2013 * 07:19 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2013.codfw.wmnet with OS trixie * 07:13 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] (duration: 07m 48s) * 07:09 kharlan@deploy2003: kharlan: Continuing with deployment * 07:08 kharlan@deploy2003: kharlan: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:06 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 01:15 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 01:14 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply == 2026-07-14 == * 22:51 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_magru * 22:51 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7016.magru.wmnet * 22:46 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_magru * 22:46 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7008.magru.wmnet * 22:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7015.magru.wmnet * 22:04 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7007.magru.wmnet * 21:29 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7014.magru.wmnet * 21:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7006.magru.wmnet * 21:13 dzahn@cumin2002: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 0:15:00 on gerrit.wikimedia.org with reason: reboot * 21:11 mutante: gerrit2003 (gerrit.wikimedia.org) - reboot for maintenance * 21:11 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on gerrit2003.wikimedia.org with reason: reboot * 20:56 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:56 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:56 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:55 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 20:48 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7013.magru.wmnet * 20:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7005.magru.wmnet * 20:41 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host phab1005.eqiad.wmnet with OS trixie * 20:28 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] (duration: 06m 47s) * 20:24 sbassett@deploy2003: sbassett: Continuing with deployment * 20:23 sbassett@deploy2003: sbassett: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:23 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on phab1005.eqiad.wmnet with reason: host reimage * 20:21 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] * 20:20 aokoth@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on phab1005.eqiad.wmnet with reason: host reimage * 20:12 jhuneidi@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] (duration: 07m 42s) * 20:07 jhuneidi@deploy2003: jhuneidi, priyankar22: Continuing with deployment * 20:06 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7012.magru.wmnet * 20:06 jhuneidi@deploy2003: jhuneidi, priyankar22: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:04 jhuneidi@deploy2003: Started scap sync-world: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] * 20:02 aokoth@cumin1003: START - Cookbook sre.hosts.reimage for host phab1005.eqiad.wmnet with OS trixie * 20:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7004.magru.wmnet * 20:00 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet * 19:57 aokoth@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet * 19:24 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7011.magru.wmnet * 19:19 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7003.magru.wmnet * 19:11 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] (duration: 08m 33s) * 19:07 jforrester@deploy2003: jforrester: Continuing with deployment * 19:04 jforrester@deploy2003: jforrester: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:02 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] * 18:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7010.magru.wmnet * 18:38 mutante: rotating phabricator-gerrit bot token (its-phabricator) * 18:18 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 17:44 swfrench@deploy2003: Finished scap sync-world: Deployment to pick up new production image (duration: 31m 44s) * 17:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7002.magru.wmnet * 17:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7009.magru.wmnet * 17:32 swfrench@deploy2003: swfrench: Continuing with deployment * 17:29 swfrench@deploy2003: swfrench: Deployment to pick up new production image synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:17 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2035: repooling after rack b5 maintenance * 17:16 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool es2035: repooling after rack b5 maintenance * 17:16 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2188: repooling after rack b5 maintenance * 17:12 swfrench@deploy2003: Started scap sync-world: Deployment to pick up new production image * 17:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7001.magru.wmnet * 16:57 swfrench-wmf: reprepro include php8.3_8.3.32-1+wmf12u2 into component/php83 for bookworm-wikimedia * 16:50 sukhe: pool cp2046 * 16:47 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4039.ulsfo.wmnet * 16:44 sukhe: sudo cumin -b31 "A:cp" "run-puppet-agent" * 16:33 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on contint1003.wikimedia.org with reason: reboot * 16:32 mutante: contint1003 - main CI server - rebooting * 16:31 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2188: repooling after rack b5 maintenance * 16:31 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2178: repooling after rack b5 maintenance * 16:29 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 16:28 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 16:28 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 16:28 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 16:18 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2014.codfw.wmnet * 16:18 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 16:17 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2014.codfw.wmnet * 16:10 mvernon@cumin1003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-thanos-proxies (exit_code=0) rolling restart_daemons on A:thanos-fe * 16:09 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 16:07 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp4039.ulsfo.wmnet * 16:07 mvernon@cumin1003: START - Cookbook sre.swift.roll-restart-reboot-swift-thanos-proxies rolling restart_daemons on A:thanos-fe * 16:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2014.codfw.wmnet with OS trixie * 15:56 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1014.eqiad.wmnet with OS trixie * 15:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2014.codfw.wmnet with reason: host reimage * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2014 * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2014 * 15:28 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2014 * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2014.codfw.wmnet 194.16.192.10.in-addr.arpa 4.9.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:28 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2014.codfw.wmnet 194.16.192.10.in-addr.arpa 4.9.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2014 - mvernon@cumin2003" * 15:28 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2014 - mvernon@cumin2003" * 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Apply title-related policies when selecting the name of the entity - kamila@cumin1003" * 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Apply title-related policies when selecting the name of the entity - kamila@cumin1003 * 15:22 kamila@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Apply title-related policies when selecting the name of the entity - kamila@cumin1003 * 15:22 kamila@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Apply title-related policies when selecting the name of the entity - kamila@cumin1003" * 15:20 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 15:20 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2014 * 15:20 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2014.codfw.wmnet with OS trixie * 15:19 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1014.eqiad.wmnet with OS trixie * 15:01 dancy@deploy2003: Installation of scap version "4.274.1" completed for 3 hosts * 15:00 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2177: repooling after rack b5 maintenance * 15:00 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2159: repooling after rack b5 maintenance * 14:59 dancy@deploy2003: Installing scap version "4.274.1" for 3 host(s) * 14:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2015.codfw.wmnet with OS trixie * 14:54 seanleong-wmde: Finished populateSitesTable for isvwiki ([[phab:T429939|T429939]]) * 14:53 javiermonton@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] (duration: 07m 35s) * 14:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1015.eqiad.wmnet with OS trixie * 14:49 javiermonton@deploy2003: javiermonton: Continuing with deployment * 14:48 javiermonton@deploy2003: javiermonton: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:46 javiermonton@deploy2003: Started scap sync-world: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] * 14:42 otto@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 14:41 otto@deploy2003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 14:41 otto@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 14:40 otto@deploy2003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 14:40 otto@deploy2003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 14:39 otto@deploy2003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 14:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2015.codfw.wmnet with reason: host reimage * 14:34 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1015.eqiad.wmnet with reason: host reimage * 14:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2015.codfw.wmnet with reason: host reimage * 14:30 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1015.eqiad.wmnet with reason: host reimage * 14:30 seanleong-wmde@deploy2003: mwscript-k8s job started: foreachwikiindblist wikidataclient extensions/Wikibase/lib/maintenance/populateSitesTable.php --force-protocol https # [[phab:T429939|T429939]] * 14:24 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling reboot on A:durum and not (A:durum-eqiad or A:durum-codfw or A:durum-esams) and A:durum * 14:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2015.codfw.wmnet with OS trixie * 14:15 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2016.codfw.wmnet with OS trixie * 14:14 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2159: repooling after rack b5 maintenance * 14:14 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1015.eqiad.wmnet with OS trixie * 14:12 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1016.eqiad.wmnet with OS trixie * 14:12 cmooney@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=pki,name=codfw * 14:12 sbisson@deploy2003: helmfile [codfw] DONE helmfile.d/services/cxserver: sync * 14:11 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2002.codfw.wmnet * 14:11 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2002.codfw.wmnet * 14:11 sbisson@deploy2003: helmfile [codfw] START helmfile.d/services/cxserver: sync * 14:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1005.wikimedia.org * 14:07 sbisson@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cxserver: sync * 14:07 sbisson@deploy2003: helmfile [eqiad] START helmfile.d/services/cxserver: sync * 14:05 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1005.wikimedia.org * 14:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader2005.wikimedia.org * 14:02 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2003.codfw.wmnet * 14:02 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2003.codfw.wmnet * 14:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=tegola-vector-tiles,name=codfw * 14:00 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=kartotherian,name=codfw * 14:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader2005.wikimedia.org * 13:58 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2016.codfw.wmnet with reason: host reimage * 13:57 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-ncredir (exit_code=0) rolling reboot on A:ncredir and A:ncredir * 13:57 sbisson@deploy2003: helmfile [staging] DONE helmfile.d/services/cxserver: sync * 13:56 sbisson@deploy2003: helmfile [staging] START helmfile.d/services/cxserver: sync * 13:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1016.eqiad.wmnet with reason: host reimage * 13:52 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:52 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:51 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2016.codfw.wmnet with reason: host reimage * 13:50 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1016.eqiad.wmnet with reason: host reimage * 13:49 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy (exit_code=0) rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 13:49 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling reboot on A:wikidough * 13:46 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-tcp-proxy (exit_code=0) rolling reboot on A:tcpproxy and A:tcpproxy * 13:43 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and not (A:durum-eqiad or A:durum-codfw or A:durum-esams) and A:durum * 13:42 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=97) rolling reboot on A:durum and A:durum * 13:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2011.codfw.wmnet * 13:36 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-reboot (exit_code=0) rolling reboot on A:dnsbox and A:ulsfo and (A:dnsbox) * 13:36 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns4004.wikimedia.org * 13:34 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1016.eqiad.wmnet with OS trixie * 13:34 topranks: reboot lsw1-b5-codfw to upgrade JunOS [[phab:T430918|T430918]] * 13:34 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2016.codfw.wmnet with OS trixie * 13:32 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2002.codfw.wmnet * 13:31 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2011.codfw.wmnet * 13:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2012.codfw.wmnet * 13:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2017.codfw.wmnet with OS trixie * 13:25 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1017.eqiad.wmnet with OS trixie * 13:24 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2012.codfw.wmnet * 13:22 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2002.codfw.wmnet * 13:22 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:22 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:22 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns4004.wikimedia.org * 13:21 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2005.codfw.wmnet * 13:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2013.codfw.wmnet * 13:18 elukey@dns1004: END - running authdns-update * 13:17 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2005.codfw.wmnet * 13:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2004.codfw.wmnet * 13:16 elukey@dns1004: START - running authdns-update * 13:16 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1046: es1046 after reimage * 13:14 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1029.eqiad.wmnet,service=s8 * 13:14 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1029.eqiad.wmnet,service=s5 * 13:13 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1029.eqiad.wmnet,service=s5 * 13:13 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1029.eqiad.wmnet,service=s8 * 13:13 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2004.codfw.wmnet * 13:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2013.codfw.wmnet * 13:11 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2014.codfw.wmnet * 13:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2003.codfw.wmnet * 13:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2017.codfw.wmnet with reason: host reimage * 13:07 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:07 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns4003.wikimedia.org * 13:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2003.codfw.wmnet * 13:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1067.eqiad.wmnet * 13:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1067.eqiad.wmnet * 13:06 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1067.eqiad.wmnet * 13:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm-test1001.wikimedia.org * 13:05 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2017.codfw.wmnet with reason: host reimage * 13:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1017.eqiad.wmnet with reason: host reimage * 13:04 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2014.codfw.wmnet * 13:03 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:02 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2188: codfw rack B5 depool for maintenance * 13:02 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_magru * 13:01 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2188: codfw rack B5 depool for maintenance * 13:01 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_magru * 13:01 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2178: codfw rack B5 depool for maintenance * 13:01 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm-test1001.wikimedia.org * 13:01 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2178: codfw rack B5 depool for maintenance * 13:01 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2177: codfw rack B5 depool for maintenance * 13:00 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2177: codfw rack B5 depool for maintenance * 12:59 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1068.eqiad.wmnet * 12:59 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1068.eqiad.wmnet * 12:58 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola-vector-tiles,name=codfw * 12:58 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2159: codfw rack B5 depool for maintenance * 12:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1017.eqiad.wmnet with reason: host reimage * 12:58 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola,name=codfw * 12:57 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=kartotherian,name=codfw * 12:57 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2159: codfw rack B5 depool for maintenance * 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 30 hosts with reason: lsw1-b5-codfw JunOS upgrade * 12:55 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lsw1-b5-codfw,lsw1-b5-codfw IPv6,lsw1-b5-codfw.mgmt,ssw1-a[1,8]-codfw.mgmt with reason: switch upgade lsw1-b5-codfw * 12:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet * 12:49 cmooney@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=pki,name=codfw * 12:49 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1067.eqiad.wmnet with OS trixie * 12:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet * 12:48 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2017.codfw.wmnet with OS trixie * 12:47 topranks: depool codfw pki in dns discovery ahead of lsw1-b5-codfw maintenance [[phab:T430918|T430918]] * 12:47 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns4003.wikimedia.org * 12:47 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and A:ulsfo and (A:dnsbox) * 12:47 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2018.codfw.wmnet * 12:45 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2018.codfw.wmnet * 12:45 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and A:durum * 12:45 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-tcp-proxy rolling reboot on A:tcpproxy and A:tcpproxy * 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host pki-root1002.eqiad.wmnet * 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1009.eqiad.wmnet * 12:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1009.eqiad.wmnet * 12:44 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 12:43 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-ncredir rolling reboot on A:ncredir and A:ncredir * 12:43 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling reboot on A:wikidough * 12:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2018.codfw.wmnet with OS trixie * 12:42 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1017.eqiad.wmnet with OS trixie * 12:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader2006.wikimedia.org * 12:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1018.eqiad.wmnet with OS trixie * 12:39 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1009.eqiad.wmnet * 12:38 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host pki-root1002.eqiad.wmnet * 12:38 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1009.eqiad.wmnet * 12:38 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1008.eqiad.wmnet * 12:38 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1008.eqiad.wmnet * 12:35 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader2006.wikimedia.org * 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1006.wikimedia.org * 12:33 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1008.eqiad.wmnet * 12:30 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1046: es1046 after reimage * 12:29 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host es1046.eqiad.wmnet with OS trixie * 12:29 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1006.wikimedia.org * 12:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test2005.wikimedia.org * 12:28 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1008.eqiad.wmnet * 12:27 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1007.eqiad.wmnet * 12:27 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1007.eqiad.wmnet * 12:27 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1067.eqiad.wmnet with reason: host reimage * 12:25 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1068.eqiad.wmnet with reason: vacuum overlarge container dbs * 12:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2018.codfw.wmnet with reason: host reimage * 12:24 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test2005.wikimedia.org * 12:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test1005.wikimedia.org * 12:23 Amir1: mwscript-k8s --follow --dblist=ores -- extensions/ORES/maintenance/PurgeScoreCache.php --model goodfaith --old ([[phab:T431159|T431159]]) * 12:22 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1007.eqiad.wmnet * 12:22 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test1005.wikimedia.org * 12:22 atsukoito: restarting pybal on lvs2013 `low-traffic` for https://gerrit.wikimedia.org/r/1310535 * 12:22 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1007.eqiad.wmnet * 12:21 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1006.eqiad.wmnet * 12:21 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1006.eqiad.wmnet * 12:20 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1018.eqiad.wmnet with reason: host reimage * 12:19 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2018.codfw.wmnet with reason: host reimage * 12:18 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1067.eqiad.wmnet with reason: host reimage * 12:16 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1006.eqiad.wmnet * 12:15 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1006.eqiad.wmnet * 12:15 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1005.eqiad.wmnet * 12:15 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1005.eqiad.wmnet * 12:15 atsukoito: restarting pybal on lvs2014 for https://gerrit.wikimedia.org/r/1310535 * 12:12 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1018.eqiad.wmnet with reason: host reimage * 12:11 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1005.eqiad.wmnet * 12:11 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1005.eqiad.wmnet * 12:10 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1004.eqiad.wmnet * 12:10 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1004.eqiad.wmnet * 12:09 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on es1046.eqiad.wmnet with reason: host reimage * 12:08 atsukoito: restarting pybal on lvs1019 `low-traffic` for https://gerrit.wikimedia.org/r/1310535 * 12:06 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1004.eqiad.wmnet * 12:06 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1004.eqiad.wmnet * 12:06 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1003.eqiad.wmnet * 12:06 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1003.eqiad.wmnet * 12:05 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on es1046.eqiad.wmnet with reason: host reimage * 12:05 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "set ml-serve1001 back to active state - cmooney@cumin1003" * 12:04 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "set ml-serve1001 back to active state - cmooney@cumin1003" * 12:04 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:02 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1003.eqiad.wmnet * 12:01 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1003.eqiad.wmnet * 12:01 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1002.eqiad.wmnet * 12:01 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1002.eqiad.wmnet * 12:01 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:59 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2018.codfw.wmnet with OS trixie * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1067 * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1067 * 11:59 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1067 * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1067.eqiad.wmnet 17.48.64.10.in-addr.arpa 7.1.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:59 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1067.eqiad.wmnet 17.48.64.10.in-addr.arpa 7.1.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1067 - blake@cumin1003" * 11:58 atsukoito: restarting pybal on lvs1018 `high-traffic2` for https://gerrit.wikimedia.org/r/1310535 * 11:57 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1002.eqiad.wmnet * 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2019.codfw.wmnet with OS trixie * 11:56 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1002.eqiad.wmnet * 11:56 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 11:56 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1018.eqiad.wmnet with OS trixie * 11:54 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 11:54 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:54 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1019.eqiad.wmnet with OS trixie * 11:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:49 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:49 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:49 aikochou@deploy2003: helmfile [codfw] DONE helmfile.d/services/changeprop: sync * 11:48 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host es1046.eqiad.wmnet with OS trixie * 11:48 aikochou@deploy2003: helmfile [codfw] START helmfile.d/services/changeprop: sync * 11:48 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310535 * 11:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1046: Reimage to Trixie * 11:44 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1046: Reimage to Trixie * 11:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5:00:00 on es1046.eqiad.wmnet with reason: Reimage to Trixie * 11:42 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:42 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:42 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 11:42 aikochou@deploy2003: helmfile [eqiad] DONE helmfile.d/services/changeprop: sync * 11:41 aikochou@deploy2003: helmfile [eqiad] START helmfile.d/services/changeprop: sync * 11:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2019.codfw.wmnet with reason: host reimage * 11:36 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] (duration: 09m 41s) * 11:36 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:36 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:35 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:35 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:32 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1019.eqiad.wmnet with reason: host reimage * 11:32 jforrester@deploy2003: jforrester, gengh: Continuing with deployment * 11:29 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2019.codfw.wmnet with reason: host reimage * 11:28 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1019.eqiad.wmnet with reason: host reimage * 11:28 jforrester@deploy2003: jforrester, gengh: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:26 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] * 11:20 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2003.codfw.wmnet * 11:20 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:19 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2003.codfw.wmnet * 11:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:12 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1019.eqiad.wmnet with OS trixie * 11:10 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2019.codfw.wmnet with OS trixie * 11:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1020.eqiad.wmnet with OS trixie * 11:10 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1067 - blake@cumin1003" * 11:09 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] (duration: 12m 12s) * 11:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2020.codfw.wmnet with OS trixie * 11:03 kharlan@deploy2003: kharlan: Continuing with deployment * 11:01 blake@cumin1003: START - Cookbook sre.dns.netbox * 11:01 kharlan@deploy2003: kharlan: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:57 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] * 10:55 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] (duration: 31m 40s) * 10:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1020.eqiad.wmnet with reason: host reimage * 10:52 marostegui@dns1004: START - running authdns-update * 10:49 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2020.codfw.wmnet with reason: host reimage * 10:49 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:48 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1020.eqiad.wmnet with reason: host reimage * 10:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2020.codfw.wmnet with reason: host reimage * 10:43 kharlan@deploy2003: kharlan: Continuing with deployment * 10:42 kharlan@deploy2003: kharlan: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:32 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1020.eqiad.wmnet with OS trixie * 10:29 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2159: Repooling after switchover * 10:29 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1067 * 10:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1021.eqiad.wmnet with OS trixie * 10:27 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1067.eqiad.wmnet with OS trixie * 10:27 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:27 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1067.eqiad.wmnet * 10:27 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:26 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1067.eqiad.wmnet * 10:26 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1067.eqiad.wmnet * 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2020.codfw.wmnet with OS trixie * 10:26 blake@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1055.eqiad.wmnet * 10:26 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1055.eqiad.wmnet * 10:26 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1055.eqiad.wmnet * 10:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2021.codfw.wmnet with OS trixie * 10:24 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] * 10:11 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1055.eqiad.wmnet with OS trixie * 10:09 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1021.eqiad.wmnet with reason: host reimage * 10:05 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2021.codfw.wmnet with reason: host reimage * 10:03 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310129 revert * 10:02 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1021.eqiad.wmnet with reason: host reimage * 10:01 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2021.codfw.wmnet with reason: host reimage * 09:58 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310129 * 09:50 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1055.eqiad.wmnet with reason: host reimage * 09:45 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1055.eqiad.wmnet with reason: host reimage * 09:45 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1021.eqiad.wmnet with OS trixie * 09:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1022.eqiad.wmnet with OS trixie * 09:44 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2159: Repooling after switchover * 09:44 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2021.codfw.wmnet with OS trixie * 09:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2022.codfw.wmnet with OS trixie * 09:31 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2159.codfw.wmnet * 09:28 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1055 * 09:28 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1055 * 09:27 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on ms-fe1022.eqiad.wmnet with reason: host reimage * 09:27 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1022.eqiad.wmnet with reason: host reimage * 09:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2022.codfw.wmnet with reason: host reimage * 09:21 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2022.codfw.wmnet with reason: host reimage * 09:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2159: Rebooting db2159.codfw.wmnet * 09:20 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2159: Rebooting db2159.codfw.wmnet * 09:18 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 09:18 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 09:18 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 09:17 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 09:16 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2159.codfw.wmnet * 09:13 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b] (thin): Regular analytics weekly train THIN [analytics/refinery@ad6e05b8] (duration: 02m 07s) * 09:11 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b] (thin): Regular analytics weekly train THIN [analytics/refinery@ad6e05b8] * 09:10 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1022.eqiad.wmnet with OS trixie * 09:07 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1023.eqiad.wmnet with OS trixie * 09:06 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b]: Regular analytics weekly train [analytics/refinery@ad6e05b8] (duration: 04m 49s) * 09:04 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2022.codfw.wmnet with OS trixie * 09:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2023.codfw.wmnet with OS trixie * 09:01 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b]: Regular analytics weekly train [analytics/refinery@ad6e05b8] * 09:01 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@ad6e05b8] (duration: 02m 01s) * 09:00 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1055 * 09:00 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1055.eqiad.wmnet 50.32.64.10.in-addr.arpa 0.5.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:00 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1055.eqiad.wmnet 50.32.64.10.in-addr.arpa 0.5.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:00 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:00 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1055 - blake@cumin1003" * 09:00 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1055 - blake@cumin1003" * 08:59 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@ad6e05b8] * 08:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2159 [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94811 and previous config saved to /var/cache/conftool/dbconfig/20260714-085624-cwilliams.json * 08:55 blake@cumin1003: START - Cookbook sre.dns.netbox * 08:55 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1055 * 08:54 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1055.eqiad.wmnet with OS trixie * 08:54 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1055.eqiad.wmnet * 08:53 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1055.eqiad.wmnet * 08:53 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1055.eqiad.wmnet * 08:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2220 to s7 primary [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94810 and previous config saved to /var/cache/conftool/dbconfig/20260714-085239-cwilliams.json * 08:51 cezmunsta: Starting s7 codfw failover from db2159 to db2220 - [[phab:T430920|T430920]] * 08:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1023.eqiad.wmnet with reason: host reimage * 08:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2220 with weight 0 [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94809 and previous config saved to /var/cache/conftool/dbconfig/20260714-084553-cwilliams.json * 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s7 [[phab:T430920|T430920]] * 08:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2023.codfw.wmnet with reason: host reimage * 08:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1023.eqiad.wmnet with reason: host reimage * 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2023.codfw.wmnet with reason: host reimage * 08:34 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:34 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:29 marostegui@dns1004: END - running authdns-update * 08:29 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox-dev2003.codfw.wmnet * 08:27 marostegui@dns1004: START - running authdns-update * 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker2*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2009.codfw.wmnet * 08:26 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2009.codfw.wmnet * 08:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1023.eqiad.wmnet with OS trixie * 08:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox-dev2003.codfw.wmnet * 08:24 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:24 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:24 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2023.codfw.wmnet with OS trixie * 08:24 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1029.eqiad.wmnet with reason: reboot * 08:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1027.eqiad.wmnet with reason: reboot * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:21 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2009.codfw.wmnet * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:20 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2009.codfw.wmnet * 08:20 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2008.codfw.wmnet * 08:20 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2008.codfw.wmnet * 08:15 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2008.codfw.wmnet * 08:14 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2008.codfw.wmnet * 08:14 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2007.codfw.wmnet * 08:14 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2007.codfw.wmnet * 08:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1024.eqiad.wmnet with OS trixie * 08:12 elukey@cumin1003: END (PASS) - Cookbook sre.pki.restart-reboot (exit_code=0) rolling reboot on P<nowiki>{</nowiki>pki*<nowiki>}</nowiki> and (A:pki) * 08:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 08:10 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 08:09 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2007.codfw.wmnet * 08:08 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2007.codfw.wmnet * 08:08 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2006.codfw.wmnet * 08:08 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2006.codfw.wmnet * 08:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2024.codfw.wmnet with OS trixie * 08:03 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2006.codfw.wmnet * 08:02 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2006.codfw.wmnet * 08:02 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2005.codfw.wmnet * 08:02 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2005.codfw.wmnet * 07:58 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2005.codfw.wmnet * 07:58 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2005.codfw.wmnet * 07:57 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2004.codfw.wmnet * 07:57 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2004.codfw.wmnet * 07:54 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki.discovery.wmnet. on all recursors * 07:54 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache pki.discovery.wmnet. on all recursors * 07:53 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2004.codfw.wmnet * 07:53 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2004.codfw.wmnet * 07:53 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2003.codfw.wmnet * 07:53 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2003.codfw.wmnet * 07:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1024.eqiad.wmnet with reason: host reimage * 07:49 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki.discovery.wmnet. on all recursors * 07:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2003.codfw.wmnet * 07:49 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache pki.discovery.wmnet. on all recursors * 07:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2024.codfw.wmnet with reason: host reimage * 07:48 elukey@cumin1003: START - Cookbook sre.pki.restart-reboot rolling reboot on P<nowiki>{</nowiki>pki*<nowiki>}</nowiki> and (A:pki) * 07:46 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1024.eqiad.wmnet with reason: host reimage * 07:45 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2024.codfw.wmnet with reason: host reimage * 07:45 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2003.codfw.wmnet * 07:44 elukey@cumin1003: END (PASS) - Cookbook sre.misc-clusters.restart-reboot-config-master (exit_code=0) rolling reboot on P<nowiki>{</nowiki>config-master*<nowiki>}</nowiki> and (A:config-master or A:config-master-eqiad or A:config-master-codfw) * 07:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2002.codfw.wmnet * 07:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2002.codfw.wmnet * 07:39 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2002.codfw.wmnet * 07:39 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) config-master.discovery.wmnet. on all recursors * 07:39 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache config-master.discovery.wmnet. on all recursors * 07:39 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2002.codfw.wmnet * 07:39 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker2*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl200*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl2003.codfw.wmnet * 07:36 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl2003.codfw.wmnet * 07:35 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) config-master.discovery.wmnet. on all recursors * 07:35 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache config-master.discovery.wmnet. on all recursors * 07:34 elukey@cumin1003: START - Cookbook sre.misc-clusters.restart-reboot-config-master rolling reboot on P<nowiki>{</nowiki>config-master*<nowiki>}</nowiki> and (A:config-master or A:config-master-eqiad or A:config-master-codfw) * 07:31 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl2003.codfw.wmnet * 07:31 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl2003.codfw.wmnet * 07:31 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl2002.codfw.wmnet * 07:31 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl2002.codfw.wmnet * 07:29 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1024.eqiad.wmnet with OS trixie * 07:28 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2024.codfw.wmnet with OS trixie * 07:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl2002.codfw.wmnet * 07:26 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl2002.codfw.wmnet * 07:26 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl200*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 07:26 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 07:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 06:50 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lists1004.wikimedia.org * 06:44 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host lists1004.wikimedia.org * 06:25 marostegui@dns1004: END - running authdns-update * 06:23 marostegui@dns1004: START - running authdns-update * 06:22 marostegui@dns1004: END - running authdns-update * 06:20 marostegui@dns1004: START - running authdns-update * 06:17 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1026.eqiad.wmnet with reason: reboot * 06:04 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: sync * 06:04 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: sync * 06:03 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync * 06:03 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync * 06:02 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync * 06:01 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync * 06:01 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync * 06:00 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync * 05:59 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:59 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:40 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:39 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:26 marostegui@dns1004: END - running authdns-update * 05:24 marostegui@dns1004: START - running authdns-update * 05:24 marostegui@dns1004: START - running authdns-update * 05:13 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1004.wikimedia.org * 05:07 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1004.wikimedia.org * 04:01 mwpresync@deploy2003: Pruned MediaWiki: 1.47.0-wmf.8 (duration: 01m 07s) * 03:39 mwpresync@deploy2003: Finished scap sync-world: testwikis to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] (duration: 36m 01s) * 03:03 mwpresync@deploy2003: Started scap sync-world: testwikis to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 29s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-13 == * 23:33 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1064.eqiad.wmnet * 23:33 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1064.eqiad.wmnet * 23:08 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1064.eqiad.wmnet with reason: vacuum overlarge container dbs * 23:06 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1069.eqiad.wmnet * 23:06 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1069.eqiad.wmnet * 22:34 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1069.eqiad.wmnet with reason: vacuum overlarge container dbs * 21:18 maryum: Deployed security fix for [[phab:T321092|T321092]] * 20:28 swfrench-wmf: reprepro include etcd-mirror_0.0.12-1+deb13u1 into main for trixie-wikimedia - [[phab:T424266|T424266]] * 20:26 swfrench-wmf: reprepro include etcd-mirror_0.0.12-1+deb12u1 into main for bookworm-wikimedia - [[phab:T428495|T428495]] * 20:23 dancy@deploy2003: Finished scap sync-world: Testing [[phab:T431635|T431635]] (duration: 03m 36s) * 20:19 dancy@deploy2003: Started scap sync-world: Testing [[phab:T431635|T431635]] * 20:18 dancy@deploy2003: Installation of scap version "4.274.0" completed for 3 hosts * 20:16 dancy@deploy2003: Installing scap version "4.274.0" for 3 host(s) * 20:12 kemayo@deploy2003: Finished scap sync-world: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] (duration: 08m 25s) * 20:07 kemayo@deploy2003: soda, esanders, kemayo: Continuing with deployment * 20:05 kemayo@deploy2003: soda, esanders, kemayo: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there * 20:04 kemayo@deploy2003: Started scap sync-world: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] * 18:22 cwhite: lvextend vg0/srv +500g on centrallog hosts * 18:19 cdobbins@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS trixie * 17:46 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1071.eqiad.wmnet * 17:46 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1071.eqiad.wmnet * 17:13 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1071.eqiad.wmnet with reason: vacuum overlarge container dbs * 17:07 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1065.eqiad.wmnet * 17:07 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1065.eqiad.wmnet * 17:06 dzahn@dns1006: END - running authdns-update * 17:04 dzahn@dns1006: START - running authdns-update * 17:01 dzahn@dns1006: END - running authdns-update * 16:59 dzahn@dns1006: START - running authdns-update * 16:51 dancy@deploy2003: Finished scap sync-world: testing [[phab:T428971|T428971]] (duration: 03m 37s) * 16:47 dancy@deploy2003: Started scap sync-world: testing [[phab:T428971|T428971]] * 16:45 atsukoito: restarting pybal on lvs1019 to flush IP address for `cirrussearch1122.eqiad.wmnet` after moving the vlan [[phab:T431311|T431311]] * 16:42 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:42 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:42 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:42 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:42 Amir1: mwscript-k8s --follow --dblist=ores -- extensions/ORES/maintenance/PurgeScoreCache.php --model damaging --old ([[phab:T431159|T431159]]) * 16:34 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Pool test * 16:34 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 16:34 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 16:34 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Pool test * 16:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Depool test * 16:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 16:33 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 16:33 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Depool test * 16:31 dancy@deploy2003: Installation of scap version "4.273.0" completed for 159 hosts * 16:29 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1065.eqiad.wmnet with reason: vacuum overlarge container dbs * 16:27 dancy@deploy2003: Installing scap version "4.273.0" for 159 host(s) * 16:27 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics-external: sync * 16:27 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics-external: sync * 16:26 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics-external: sync * 16:26 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics-external: sync * 16:22 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync * 16:21 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync * 16:21 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: sync * 16:21 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: sync * 16:19 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync * 16:19 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync * 16:18 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync * 16:17 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync * 15:59 atsukoito: restarting pybal on lvs1018 for https://gerrit.wikimedia.org/r/1310117 * 15:55 aikochou@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 15:50 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310117 * 15:46 aikochou@deploy2003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 15:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host kafka-logging1006.eqiad.wmnet * 15:43 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host kafka-logging1006.eqiad.wmnet * 15:41 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host ganeti-test[2001-2003].codfw.wmnet * 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host ganeti-test[2001-2003].codfw.wmnet * 15:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host netbox1003.eqiad.wmnet * 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host netbox1003.eqiad.wmnet * 15:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host netbox2003.codfw.wmnet * 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host netbox2003.codfw.wmnet * 15:36 sukhe: restart pybal on lvs1020 * 15:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet * 15:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet * 15:08 btullis@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:06 btullis@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 15:01 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:01 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:35 cdobbins@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 14:34 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:33 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:33 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:32 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:29 cdobbins@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 14:28 marostegui@dns1004: END - running authdns-update * 14:28 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:27 marostegui@dns1004: START - running authdns-update * 14:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1023.eqiad.wmnet with reason: reboot * 14:18 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2009.codfw.wmnet with OS trixie * 14:14 swfrench-wmf: start rolling run-puppet-agent on A:cp for ATS config change - [[phab:T428909|T428909]] [[phab:T431838|T431838]] * 14:05 swfrench-wmf: disable-puppet on A:cp for ATS config change - [[phab:T428909|T428909]] [[phab:T431838|T431838]] * 14:05 cdobbins@cumin2002: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie * 14:02 marostegui@dns1004: END - running authdns-update * 14:00 marostegui@dns1004: START - running authdns-update * 14:00 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1070.eqiad.wmnet * 14:00 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1070.eqiad.wmnet * 13:58 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2009.codfw.wmnet with reason: host reimage * 13:57 cdobbins@cumin2002: conftool action : set/pooled=no; selector: name=dns7002.* * 13:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2009.codfw.wmnet with reason: host reimage * 13:48 rscout@deploy2003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply * 13:48 rscout@deploy2003: helmfile [eqiad] START helmfile.d/services/miscweb: apply * 13:48 rscout@deploy2003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply * 13:47 rscout@deploy2003: helmfile [codfw] START helmfile.d/services/miscweb: apply * 13:40 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:33 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2009.codfw.wmnet with OS trixie * 13:30 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1070.eqiad.wmnet with reason: vacuum overlarge container dbs * 13:28 aude@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] (duration: 11m 12s) * 13:23 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:22 aude@deploy2003: aikochou, javiermonton, aude, gkm563: Continuing with deployment * 13:22 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:19 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:19 aude@deploy2003: aikochou, javiermonton, aude, gkm563: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] synced to the testservers * 13:17 aude@deploy2003: Started scap sync-world: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] * 13:01 ladsgroup@deploy2003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 13:01 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:00 ladsgroup@deploy2003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 12:59 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:52 ladsgroup@deploy2003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 12:51 ladsgroup@deploy2003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 12:48 atsuko@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 12:48 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 12:47 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2008.codfw.wmnet with OS trixie * 12:47 atsuko@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 12:47 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply * 12:47 atsuko@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:46 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 12:45 atsuko@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:45 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply * 12:45 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:44 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:43 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] (duration: 07m 02s) * 12:38 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 12:37 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:36 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] * 12:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2008.codfw.wmnet with reason: host reimage * 12:23 Msz2001: Deployed changes to private code for Suggested Investigations * 12:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2008.codfw.wmnet with reason: host reimage * 12:20 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:19 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:17 atsuko@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 12:17 atsuko@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 12:16 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:15 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] (duration: 07m 14s) * 12:10 mszwarc@deploy2003: mszwarc: Continuing with deployment * 12:09 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:07 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] * 12:04 mszwarc@deploy2003: sync-world aborted: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] (duration: 00m 29s) * 12:03 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] * 12:01 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2008.codfw.wmnet with OS trixie * 12:00 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] (duration: 07m 37s) * 11:55 zabe@deploy2003: zabe: Continuing with deployment * 11:54 zabe@deploy2003: zabe: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:52 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] * 11:51 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:43 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:35 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:34 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:33 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:30 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:28 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:27 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:17 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2007.codfw.wmnet with OS trixie * 11:09 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] * 11:06 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=s8 * 11:00 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=x3 * 11:00 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=s5 * 10:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2007.codfw.wmnet with reason: host reimage * 10:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2007.codfw.wmnet with reason: host reimage * 10:51 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host dse-k8s-worker1023 * 10:50 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host dse-k8s-worker1023 * 10:44 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dse-k8s-worker1023 * 10:43 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host dse-k8s-worker1023 * 10:42 marostegui@cumin1003: dbctl commit (dc=all): 'Change x4 masters [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P94804 and previous config saved to /var/cache/conftool/dbconfig/20260713-104248-marostegui.json * 10:37 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:37 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:35 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:35 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:34 atsuko@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 10:34 atsuko@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 10:33 marostegui@cumin1003: dbctl commit (dc=all): 'Push x4 initial dbctl config [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P94803 and previous config saved to /var/cache/conftool/dbconfig/20260713-103259-marostegui.json * 10:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2007.codfw.wmnet with OS trixie * 09:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2006.codfw.wmnet with OS trixie * 09:42 marostegui@dns1004: END - running authdns-update * 09:40 marostegui@dns1004: START - running authdns-update * 09:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2006.codfw.wmnet with reason: host reimage * 09:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2006.codfw.wmnet with reason: host reimage * 09:06 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1024.eqiad.wmnet with reason: reboot * 09:06 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:01 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2006.codfw.wmnet with OS trixie * 08:44 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 08:43 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 08:43 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:42 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:42 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:42 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:41 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 08:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup2004.codfw.wmnet * 08:38 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:33 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host db1208.eqiad.wmnet * 08:30 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=x3 * 08:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2005.codfw.wmnet with OS trixie * 08:28 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup2004.codfw.wmnet * 08:28 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup2003.codfw.wmnet * 08:24 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1039: Repooling after testing * 08:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on clouddb1016.eqiad.wmnet with reason: cloning * 08:23 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s5 * 08:23 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s8 * 08:21 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 08:21 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 08:17 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup2003.codfw.wmnet * 08:17 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1004.eqiad.wmnet * 08:14 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1208.eqiad.wmnet * 08:11 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host phab1005.eqiad.wmnet * 08:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2005.codfw.wmnet with reason: host reimage * 08:07 marostegui@dns1004: END - running authdns-update * 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1004.eqiad.wmnet * 08:07 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1003.eqiad.wmnet * 08:07 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:05 marostegui@dns1004: START - running authdns-update * 08:05 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2005.codfw.wmnet with reason: host reimage * 08:05 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 08:05 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host phab1005.eqiad.wmnet * 08:05 marostegui@dns1004: START - running authdns-update * 08:05 marostegui@dns1004: START - running authdns-update * 08:05 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 08:04 marostegui@dns1004: START - running authdns-update * 08:00 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit1003.wikimedia.org * 07:58 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1003.eqiad.wmnet * 07:58 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1002-dev.eqiad.wmnet * 07:58 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:58 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:54 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1002-dev.eqiad.wmnet * 07:54 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1001-dev.eqiad.wmnet * 07:54 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit1003.wikimedia.org * 07:53 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 07:53 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:52 Msz2001: UTC morning backport+config window done * 07:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2005.codfw.wmnet with OS trixie * {{safesubst:SAL entry|1=07:50 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark (T429943}} * 07:49 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1001-dev.eqiad.wmnet * 07:46 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:46 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:45 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 07:45 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:44 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 07:44 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:43 mszwarc@deploy2003: mszwarc, danielyepezgarces, anzx: Continuing with deployment * {{safesubst:SAL entry|1=07:39 mszwarc@deploy2003: mszwarc, danielyepezgarces, anzx: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark}} * 07:39 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1039: Repooling after testing * {{safesubst:SAL entry|1=07:36 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark (T429943)}} * 07:35 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] (duration: 30m 03s) * 07:25 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit2002.wikimedia.org * 07:22 mszwarc@deploy2003: mszwarc: Continuing with deployment * 07:21 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:19 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit2002.wikimedia.org * 07:15 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aphlict1002.eqiad.wmnet * 07:11 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host aphlict1002.eqiad.wmnet * 07:08 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2003.wikimedia.org * 07:05 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] * 07:02 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2003.wikimedia.org * 07:02 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2002.wikimedia.org * 06:55 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2002.wikimedia.org * 06:55 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1003.wikimedia.org * 06:49 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1003.wikimedia.org * 06:34 marostegui: Drop m5 ipoid database [[phab:T431007|T431007]] * 06:29 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1027.eqiad.wmnet with reason: reboot * 06:24 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1028.eqiad.wmnet with reason: reboot * 06:21 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1025.eqiad.wmnet with reason: reboot * 06:17 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1022.eqiad.wmnet with reason: reboot * 06:03 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on dbproxy[2005-2008].codfw.wmnet with reason: reboot * 05:37 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1217,1228].eqiad.wmnet with reason: cloning * 05:11 marostegui: Drop users_to_rename table [[phab:T431842|T431842]] * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-12 == * 16:01 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2209 [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94792 and previous config saved to /var/cache/conftool/dbconfig/20260712-160124-marostegui.json * 15:58 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2205 to s3 primary [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94791 and previous config saved to /var/cache/conftool/dbconfig/20260712-155853-marostegui.json * 15:58 marostegui: Starting s3 codfw emergency failover from db2209 to db2205 - [[phab:T431950|T431950]] * 15:51 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2205 with weight 0 [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94790 and previous config saved to /var/cache/conftool/dbconfig/20260712-155135-marostegui.json * 15:51 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Primary switchover s3 [[phab:T431950|T431950]] * 02:01 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 01m 17s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-11 == * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 26s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-10 == * 19:12 jhathaway@dns1004: END - running authdns-update * 19:10 jhathaway@dns1004: START - running authdns-update * 18:23 mutante: vrts2002 rebooting (not the active host) * 18:21 mutante: lists2001, phab2003 - rebooting (not the active hosts) * 18:16 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on A:lvs-high-traffic2-codfw * 18:15 swfrench@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on A:lvs-high-traffic2-codfw * 17:15 mutante: [doc1004:~] $ sudo systemctl start rsync-doc-host-data-sync ([[phab:T431856|T431856]]) * 17:09 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1004.eqiad.wmnet * 17:08 jhathaway@dns1004: END - running authdns-update * 17:07 jhathaway@dns1004: START - running authdns-update * 17:06 jhathaway: depooling puppetserver1002, cause of errors is still unknown * 17:03 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1004.eqiad.wmnet * 16:57 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1003.eqiad.wmnet * 16:51 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1003.eqiad.wmnet * 16:48 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 16:48 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2004.codfw.wmnet * 16:42 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2004.codfw.wmnet * 16:41 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2003.codfw.wmnet * 16:35 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2003.codfw.wmnet * 16:33 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2002.codfw.wmnet * 16:27 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2002.codfw.wmnet * 16:25 mutante: gitlab-runners (production) rebooting cluster one by one * 16:17 mutante: etherpad1004/etherpad2002 - (etherpad.wikimedia.org) - rebooting * 16:13 mutante: doc1004/doc2003 (doc.wikimedia.org backends) - rebooting * 16:02 mutante: releases1003/releases2003 (releases.wikimedia.org backends) - rebooting for maintenance * 15:26 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2007-dev.codfw.wmnet * 15:19 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2007-dev.codfw.wmnet * 15:14 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host cloudcephosd2007-dev.codfw.wmnet * 15:14 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2007-dev.codfw.wmnet * 15:14 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host cloudcephosd2006-dev.codfw.wmnet * 15:07 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2006-dev.codfw.wmnet * 15:07 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2005-dev.codfw.wmnet * 14:59 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2005-dev.codfw.wmnet * 14:59 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2004-dev.codfw.wmnet * 14:53 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2004-dev.codfw.wmnet * 14:53 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2007-dev.codfw.wmnet * 14:51 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1054.eqiad.wmnet * 14:51 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1054.eqiad.wmnet * 14:51 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1054.eqiad.wmnet * 14:47 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2007-dev.codfw.wmnet * 14:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2006-dev.codfw.wmnet * 14:41 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2006-dev.codfw.wmnet * 14:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2005-dev.codfw.wmnet * 14:37 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2005-dev.codfw.wmnet * 14:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2005-dev.codfw.wmnet * 14:29 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2005-dev.codfw.wmnet * 14:29 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2006-dev.codfw.wmnet * 14:21 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2006-dev.codfw.wmnet * 14:21 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2010-dev.codfw.wmnet * 14:15 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2010-dev.codfw.wmnet * 14:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudgw2004-dev.codfw.wmnet * 14:10 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1054.eqiad.wmnet with OS trixie * 14:09 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudgw2004-dev.codfw.wmnet * 14:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudgw2003-dev.codfw.wmnet * 14:02 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudgw2003-dev.codfw.wmnet * 14:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2004-dev.codfw.wmnet * 13:53 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2004-dev.codfw.wmnet * 13:53 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2003-dev.codfw.wmnet * 13:48 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:44 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2003-dev.codfw.wmnet * 13:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2002-dev.codfw.wmnet * 13:42 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:41 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:41 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:37 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2002-dev.codfw.wmnet * 13:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudidp2001-dev.codfw.wmnet * 13:33 blake@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker1054.eqiad.wmnet with reason: host reimage * 13:33 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudidp2001-dev.codfw.wmnet * 13:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudnet2006-dev.codfw.wmnet * 13:26 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudnet2006-dev.codfw.wmnet * 13:26 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudnet2005-dev.codfw.wmnet * 13:23 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1054.eqiad.wmnet with reason: host reimage * 13:18 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudnet2005-dev.codfw.wmnet * 13:18 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudservices2005-dev.codfw.wmnet * 13:12 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudservices2005-dev.codfw.wmnet * 13:11 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudservices2004-dev.codfw.wmnet * 13:08 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudservices2004-dev.codfw.wmnet * 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudweb2002-dev.wikimedia.org * 13:05 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 13:05 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1054 * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1054 * 13:04 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1054 * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1054.eqiad.wmnet 49.32.64.10.in-addr.arpa 9.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:04 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1054.eqiad.wmnet 49.32.64.10.in-addr.arpa 9.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1054 - blake@cumin1003" * 13:04 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1054 - blake@cumin1003" * 13:01 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudweb2002-dev.wikimedia.org * 13:00 blake@cumin1003: START - Cookbook sre.dns.netbox * 12:59 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1054 * 12:57 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1054.eqiad.wmnet with OS trixie * 12:57 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1054.eqiad.wmnet * 12:56 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1054.eqiad.wmnet * 12:56 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1054.eqiad.wmnet * 12:47 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS trixie * 12:44 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:39 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 12:39 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 12:38 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:37 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:14 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:10 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:08 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:07 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:00 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:00 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:51 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:49 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:48 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:47 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:44 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:32 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2001.codfw.wmnet * 11:32 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1053.eqiad.wmnet * 11:32 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2001.codfw.wmnet * 11:32 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1053.eqiad.wmnet * 11:32 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1053.eqiad.wmnet * 11:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker2001.codfw.wmnet * 11:31 cgoubert@cumin1003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker2001.codfw.wmnet * 11:31 cgoubert@cumin1003: END (FAIL) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=1) rolling reimage on P<nowiki>{</nowiki>wikikube-worker2001*<nowiki>}</nowiki> and (A:wikikube-master-codfw or A:wikikube-worker-codfw) * 11:30 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:30 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:21 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 18 hosts with reason: reboot & upgrade * 11:20 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker2001.codfw.wmnet with OS trixie * 11:16 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:15 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:14 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:14 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 11 hosts * 11:14 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 11 hosts * 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:08 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 11:02 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1053.eqiad.wmnet with OS trixie * 11:01 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:58 cgoubert@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 10:57 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:38 cgoubert@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker2001.codfw.wmnet with OS trixie * 10:38 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2001.codfw.wmnet * 10:38 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2001.codfw.wmnet * 10:38 cgoubert@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on P<nowiki>{</nowiki>wikikube-worker2001*<nowiki>}</nowiki> and (A:wikikube-master-codfw or A:wikikube-worker-codfw) * 10:35 topranks: adjust IBGP outbound policy on lsw1-e2-codfw [[phab:T423430|T423430]] towards ssw1-e1-codfw * 10:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 cgoubert@cumin1003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:27 cgoubert@cumin1003: END (FAIL) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=1) rolling reimage on A:wikikube-worker-codfw * 10:27 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker2001.codfw.wmnet with OS bookworm * 10:25 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 10:24 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:24 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:15 cgoubert@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 10:11 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 10:11 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 10:08 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:08 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:07 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 10:06 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 10:00 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:55 cgoubert@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker2001.codfw.wmnet with OS bookworm * 09:55 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2005-2006,2011-2012].codfw.wmnet * 09:55 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2005-2006,2011-2012].codfw.wmnet * 09:51 cgoubert@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on A:wikikube-worker-codfw * 09:41 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1053.eqiad.wmnet with reason: host reimage * 09:37 topranks: apply new IBGP outbound policy on lsw1-e2-codfw [[phab:T423430|T423430]] * 09:36 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:36 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1053.eqiad.wmnet with reason: host reimage * 09:16 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1053 * 09:16 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1053 * 09:15 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1053 * 09:15 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1053.eqiad.wmnet 48.32.64.10.in-addr.arpa 8.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:15 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1053.eqiad.wmnet 48.32.64.10.in-addr.arpa 8.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:15 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:15 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1053 - blake@cumin1003" * 09:15 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1053 - blake@cumin1003" * 09:11 blake@cumin1003: START - Cookbook sre.dns.netbox * 09:11 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1053 * 09:08 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1053.eqiad.wmnet with OS trixie * 09:08 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1053.eqiad.wmnet * 09:08 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1053.eqiad.wmnet * 09:08 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1053.eqiad.wmnet * 09:04 brouberol@dns1004: END - running authdns-update * 09:03 brouberol@dns1004: START - running authdns-update * 08:41 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e] (thin): Regular analytics weekly train THIN [analytics/refinery@1abf22ea] (duration: 02m 11s) * 08:38 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e] (thin): Regular analytics weekly train THIN [analytics/refinery@1abf22ea] * 08:38 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e]: Regular analytics weekly train [analytics/refinery@1abf22ea] (duration: 05m 17s) * 08:38 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:34 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:33 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e]: Regular analytics weekly train [analytics/refinery@1abf22ea] * 08:32 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@1abf22ea] (duration: 02m 03s) * 08:30 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@1abf22ea] * 08:30 JavierMonton: Deploying Refinery at {{Gerrit|1abf22ea}} for changes 1308121/T427068 1306491/T430020 and {{Gerrit|1308190}} * 08:29 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:29 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 08:24 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:24 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 08:18 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:18 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 08:00 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db[2183-2184].codfw.wmnet * 08:00 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for db[2183-2184].codfw.wmnet * 07:52 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:52 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 07:49 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 11 hosts with reason: reboot & upgrade * 07:47 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:47 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 07:44 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:44 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 07:23 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 10 hosts * 07:23 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 10 hosts * 06:45 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 10 hosts with reason: reboot & upgrade * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 41s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-09 == * 23:33 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] (duration: 13m 26s) * 23:29 ladsgroup@deploy2003: ladsgroup, jdlrobson: Continuing with deployment * 23:22 ladsgroup@deploy2003: ladsgroup, jdlrobson: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:20 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] * 22:57 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1165.eqiad.wmnet * 22:56 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1165.eqiad.wmnet * 22:56 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1165.eqiad.wmnet * 22:45 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1165.eqiad.wmnet with OS trixie * 22:38 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 22:37 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 22:37 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 22:37 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:37 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 22:25 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1165.eqiad.wmnet with reason: host reimage * 22:17 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1165.eqiad.wmnet with reason: host reimage * 22:13 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 22:12 rzl@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 22:04 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 22:04 rzl@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1165 * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1165 * 22:02 jasmine@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1165 * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1165.eqiad.wmnet 115.48.64.10.in-addr.arpa 5.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:02 jasmine@cumin2002: START - Cookbook sre.dns.wipe-cache wikikube-worker1165.eqiad.wmnet 115.48.64.10.in-addr.arpa 5.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1165 - jasmine@cumin2002" * 22:02 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1165 - jasmine@cumin2002" * 22:02 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 21:57 jasmine@cumin2002: START - Cookbook sre.dns.netbox * 21:55 jasmine@cumin2002: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1165 * 21:54 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-worker1165.eqiad.wmnet with OS trixie * 21:54 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 21:54 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1165.eqiad.wmnet * 21:53 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 21:53 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1165.eqiad.wmnet * 21:53 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1165.eqiad.wmnet * 21:53 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 21:47 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 21:45 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 21:43 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 21:43 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 21:42 maryum: Deploy fix for [[phab:T431684|T431684]] * 21:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 21:27 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] (duration: 34m 14s) * 21:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs1002 * 21:23 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs1002 * 21:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS trixie * 21:22 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 21:20 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 22s) * 21:20 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 21:16 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 21:16 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 21:15 ladsgroup@deploy2003: ladsgroup: Continuing with deployment * 21:13 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2002.codfw.wmnet with OS bookworm * 21:11 ladsgroup@deploy2003: ladsgroup: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:08 ladsgroup@cumin1003: END (PASS) - Cookbook sre.wikireplicas.update-views (exit_code=0) * 21:07 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 6 hosts with reason: reboots * 20:54 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecycle work - bking@cumin2003 * 20:53 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:53 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] * 20:51 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) * 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 20:47 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecycle work - bking@cumin2003 * 20:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 20:41 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:41 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) * 20:40 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host relforge1008.eqiad.wmnet * 20:40 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1009.eqiad.wmnet with OS trixie * 20:33 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:32 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:32 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:31 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:31 ladsgroup@cumin1003: END (PASS) - Cookbook sre.wikireplicas.update-views (exit_code=0) * 20:29 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1008.eqiad.wmnet * 20:24 rzl@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 20:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2002.codfw.wmnet with OS bookworm * 20:23 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host relforge1008.eqiad.wmnet * 20:23 rzl@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 20:23 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1008.eqiad.wmnet * 20:22 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:22 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:21 bking@cumin2003: END (ERROR) - Cookbook sre.elasticsearch.rolling-operation (exit_code=97) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:21 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1009.eqiad.wmnet with reason: host reimage * 20:16 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:15 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1009.eqiad.wmnet with reason: host reimage * 20:12 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) * 20:02 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 19:55 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1009.eqiad.wmnet with OS trixie * 19:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 19:43 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 19:30 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 19:28 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 19:27 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 19:25 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 18:42 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 18:41 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 18:16 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 18:15 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 17:45 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for doh5004.wikimedia.org * 17:45 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for doh5004.wikimedia.org * 17:38 ladsgroup@deploy2003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 17:35 ladsgroup@deploy2003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 17:29 ladsgroup@deploy2003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 17:26 ladsgroup@deploy2003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 17:09 mutante: zuul[12]00[123] - rebooting for maintenance * 17:09 ebernhardson: start full in-place reindex of eqiad cirrussearch cluster * 17:08 dzahn@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-cluster (exit_code=99) * 17:08 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-cluster * 17:03 ebernhardson: start full in-place reindex of codfw cirrussearch cluster * 16:59 mutante: stewards1001/stewards2001 - reboot for maintenance * 16:54 ebernhardson: start full in-place reindex of cloudelastic cluster * 16:53 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 16:52 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply * 16:49 mutante: planet1003/planet2003 - rebooting * 16:47 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on doh5004.wikimedia.org with reason: random high load, investigating * 15:55 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 15:54 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 15:51 jynus: restarting backupmon1001 * 15:49 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 14 hosts * 15:49 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 14 hosts * 15:47 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backupmon1001.eqiad.wmnet with reason: restart * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:06 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 15:06 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 14:59 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply * 14:58 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply * 14:51 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 14 hosts * 14:51 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 14 hosts * 14:49 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 6 hosts with reason: reboot & upgrade * 14:48 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet * 14:48 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet * 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:42 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1052.eqiad.wmnet * 14:42 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1052.eqiad.wmnet * 14:42 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1052.eqiad.wmnet * 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:31 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:31 elukey@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: sync * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:30 elukey@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: sync * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:28 elukey@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: sync * 14:28 elukey@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: sync * 14:26 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:20 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1052.eqiad.wmnet with OS trixie * 14:19 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 6 hosts with reason: reboot & upgrade * 14:18 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:15 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:15 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:13 elukey: update druid indexation job for webrequest_sampled_live - [[phab:T427068|T427068]] * 14:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:09 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:09 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for papaul - jhancock@cumin2002" * 14:09 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for papaul - jhancock@cumin2002" * 14:07 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:07 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:04 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 14:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cuminunpriv1001.eqiad.wmnet * 13:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb1003.eqiad.wmnet * 13:59 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1052.eqiad.wmnet with reason: host reimage * 13:57 moritzm: installing requests security updates * 13:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cuminunpriv1001.eqiad.wmnet * 13:55 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb1003.eqiad.wmnet * 13:53 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1052.eqiad.wmnet with reason: host reimage * 13:50 moritzm: installing python-cryptography security updates * 13:47 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb2003.codfw.wmnet * 13:44 Msz2001: UTC afternoon config+backport window is done * 13:44 Msz2001: Updated `logging` on `metawiki` to fix log performers, [[phab:T431176|T431176]]#12105297 * 13:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb2003.codfw.wmnet * 13:43 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 13:43 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt1002.wikimedia.org * 13:41 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] (duration: 07m 30s) * 13:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt1002.wikimedia.org * 13:37 mszwarc@deploy2003: mszwarc: Continuing with deployment * 13:36 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1052 * 13:36 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1052 * 13:35 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:35 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1052 * 13:35 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1052.eqiad.wmnet 47.32.64.10.in-addr.arpa 7.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:35 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1052.eqiad.wmnet 47.32.64.10.in-addr.arpa 7.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:35 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:35 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1052 - blake@cumin1003" * 13:35 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1052 - blake@cumin1003" * 13:34 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] * 13:31 blake@cumin1003: START - Cookbook sre.dns.netbox * 13:31 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1052 * 13:30 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1052.eqiad.wmnet with OS trixie * 13:30 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1052.eqiad.wmnet * 13:29 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1052.eqiad.wmnet * 13:29 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1052.eqiad.wmnet * 13:17 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] (duration: 11m 26s) * 13:13 jforrester@deploy2003: jforrester: Continuing with deployment * 13:08 jforrester@deploy2003: jforrester: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:06 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] * 12:54 cgoubert@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply * 12:54 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:52 cgoubert@deploy2003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply * 12:45 cgoubert@deploy2003: helmfile [codfw] DONE helmfile.d/services/mobileapps: apply * 12:44 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:44 cgoubert@deploy2003: helmfile [codfw] START helmfile.d/services/mobileapps: apply * 12:43 cgoubert@deploy2003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 12:43 cgoubert@deploy2003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 12:42 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast4006.wikimedia.org * 12:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt2002.wikimedia.org * 12:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast7002.wikimedia.org * 12:18 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast4006.wikimedia.org * 12:18 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host ml-serve1004 * 12:18 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host ml-serve1004 * 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt2002.wikimedia.org * 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast7002.wikimedia.org * 12:10 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backup[2003,2014].codfw.wmnet with reason: reboot & upgrade * 12:10 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt-staging2001.codfw.wmnet * 12:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid1003.eqiad.wmnet * 12:06 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt-staging2001.codfw.wmnet * 12:05 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid1003.eqiad.wmnet * 12:03 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backup[1003,1014].eqiad.wmnet with reason: reboot & upgrade * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid2003.codfw.wmnet * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host irc1003.wikimedia.org * 11:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid2003.codfw.wmnet * 11:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host irc1003.wikimedia.org * 11:55 jmm@dns1004: END - running authdns-update * 11:53 jmm@dns1004: START - running authdns-update * 11:50 jmm@dns1004: END - running authdns-update * 11:48 jmm@dns1004: START - running authdns-update * 11:27 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host irc2003.wikimedia.org * 11:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host irc2003.wikimedia.org * 11:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint2001.codfw.wmnet * 11:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint1001.eqiad.wmnet * 11:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint2001.codfw.wmnet * 11:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint1001.eqiad.wmnet * 11:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-rw2001.wikimedia.org * 11:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-rw1001.wikimedia.org * 11:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-rw2001.wikimedia.org * 11:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-rw1001.wikimedia.org * 11:03 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon1003.wikimedia.org * 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2005.codfw.wmnet * 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2005.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 10:59 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2005.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 10:57 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon1003.wikimedia.org * 10:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon2002.wikimedia.org * 10:55 jmm@cumin2003: START - Cookbook sre.dns.netbox * 10:51 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon2002.wikimedia.org * 10:51 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:50 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2005.codfw.wmnet * 10:41 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:40 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host ml-serve1003 * 10:40 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host ml-serve1003 * 10:39 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2033.codfw.wmnet * 10:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install2005.wikimedia.org * 10:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install1005.wikimedia.org * 10:35 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1004.eqiad.wmnet with OS bookworm * 10:31 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install1005.wikimedia.org * 10:31 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install2005.wikimedia.org * 10:30 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install4004.wikimedia.org * 10:30 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install3004.wikimedia.org * 10:29 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install3004.wikimedia.org * 10:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install4004.wikimedia.org * 10:23 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 10:21 moritzm: failover Ganeti master in codfw/routed to ganeti2034 [[phab:T430928|T430928]] * 10:19 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.addnode (exit_code=0) for new host ganeti2031.codfw.wmnet to cluster codfw and group B * 10:19 moritzm: readded ganeti2031 to the codfw Ganeti cluster [[phab:T430910|T430910]] * 10:18 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1004.eqiad.wmnet with reason: host reimage * 10:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install5004.wikimedia.org * 10:18 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1003 * 10:18 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1003 * 10:17 jmm@cumin2003: START - Cookbook sre.ganeti.addnode for new host ganeti2031.codfw.wmnet to cluster codfw and group B * 10:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install6003.wikimedia.org * 10:16 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install5004.wikimedia.org * 10:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install6003.wikimedia.org * 10:15 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1004.eqiad.wmnet with reason: host reimage * 10:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1001.eqiad.wmnet * 10:14 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 10:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2008.wikimedia.org * 10:00 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ml-serve1004 * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1004 * 09:57 jmm@cumin2003: START - Cookbook sre.dns.netbox * 09:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install7002.wikimedia.org * 09:57 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1004 * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ml-serve1004.eqiad.wmnet 50.48.64.10.in-addr.arpa 0.5.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:57 klausman@cumin1003: START - Cookbook sre.dns.wipe-cache ml-serve1004.eqiad.wmnet 50.48.64.10.in-addr.arpa 0.5.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1004 - klausman@cumin1003" * 09:56 klausman@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1004 - klausman@cumin1003" * 09:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-coord1001.eqiad.wmnet * 09:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 09:55 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow7002.magru.wmnet * 09:52 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-coord1001.eqiad.wmnet * 09:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 09:52 klausman@cumin1003: START - Cookbook sre.dns.netbox * 09:50 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install7002.wikimedia.org * 09:50 klausman@cumin1003: START - Cookbook sre.hosts.move-vlan for host ml-serve1004 * 09:50 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1004.eqiad.wmnet with OS bookworm * 09:50 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1003.eqiad.wmnet with OS bookworm * 09:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1001.eqiad.wmnet * 09:49 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 09:49 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2008.wikimedia.org * 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2007.codfw.wmnet * 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2007.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 09:49 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow7002.magru.wmnet * 09:49 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2007.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 09:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard1003.eqiad.wmnet * 09:39 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard2003.codfw.wmnet * 09:39 jmm@cumin2003: START - Cookbook sre.dns.netbox * 09:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard1003.eqiad.wmnet * 09:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor1003.eqiad.wmnet * 09:35 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard2003.codfw.wmnet * 09:34 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2007.codfw.wmnet * 09:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor1003.eqiad.wmnet * 09:33 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor-dev2001.codfw.wmnet * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor2003.codfw.wmnet * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sretest1006.eqiad.wmnet * 09:27 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 09:25 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor-dev2001.codfw.wmnet * 09:25 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor2003.codfw.wmnet * 09:23 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] (duration: 06m 27s) * 09:23 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2205: codfw rack B4 repool after maintenance * 09:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host sretest1006.eqiad.wmnet * 09:23 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2204: codfw rack B4 repool after maintenance * 09:19 urbanecm@deploy2003: urbanecm: Continuing with deployment * 09:19 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:18 jmm@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 6 hosts with reason: reboot * 09:17 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] * 09:08 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ml-serve1003 * 09:08 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1003 * 09:07 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1003 * 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ml-serve1003.eqiad.wmnet 81.32.64.10.in-addr.arpa 1.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:07 klausman@cumin1003: START - Cookbook sre.dns.wipe-cache ml-serve1003.eqiad.wmnet 81.32.64.10.in-addr.arpa 1.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1003 - klausman@cumin1003" * 09:06 klausman@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1003 - klausman@cumin1003" * 08:58 klausman@cumin1003: START - Cookbook sre.dns.netbox * 08:57 klausman@cumin1003: START - Cookbook sre.hosts.move-vlan for host ml-serve1003 * 08:57 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1003.eqiad.wmnet with OS bookworm * 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=0) rolling reimage on P<nowiki>{</nowiki>ml-serve1003.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet * 08:55 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet * 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1003.eqiad.wmnet with OS bookworm * 08:39 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 08:38 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool db2205: codfw rack B4 repool after maintenance * 08:37 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool db2204: codfw rack B4 repool after maintenance * 08:36 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 08:35 hashar@deploy2003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 08:32 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:32 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:31 hashar@deploy2003: Rolling back deployment * 08:26 moritzm: failover Ganeti master in codfw to ganeti2048 [[phab:T430928|T430928]] * 08:16 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1003.eqiad.wmnet with OS bookworm * 08:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2004.codfw.wmnet * 08:16 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet * 08:16 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet * 08:16 klausman@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on P<nowiki>{</nowiki>ml-serve1003.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 08:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2002.codfw.wmnet * 08:15 XioNoX: lsw1-b4-codfw> request system reboot - [[phab:T430910|T430910]] * 08:15 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b4-codfw,lsw1-b4-codfw IPv6,lsw1-b4-codfw.mgmt with reason: Switch maintenance * 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for codfw rack B4 * 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:10 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2004.codfw.wmnet * 08:10 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2002.codfw.wmnet * 08:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2205: codfw rack B4 depool for maintenance * 08:08 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool db2205: codfw rack B4 depool for maintenance * 08:08 jmm@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin2003.codfw.wmnet * 08:08 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2204: codfw rack B4 depool for maintenance * 08:08 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool db2204: codfw rack B4 depool for maintenance * 08:08 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 27 hosts with reason: codfw rack B4 depool for maintenance * 08:03 jmm@cumin2002: START - Cookbook sre.hosts.reboot-single for host cumin2003.codfw.wmnet * 07:56 ayounsi@cumin1003: START - Cookbook sre.network.depool-rack with action 'depool' for codfw rack B4 * 07:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1008.eqiad.wmnet with OS trixie * 07:49 wmde-fisch@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] (duration: 08m 36s) * 07:44 wmde-fisch@deploy2003: wmde-fisch: Continuing with deployment * 07:43 wmde-fisch@deploy2003: wmde-fisch: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:41 wmde-fisch@deploy2003: Started scap sync-world: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] * 07:35 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1008.eqiad.wmnet with reason: host reimage * 07:31 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1008.eqiad.wmnet with reason: host reimage * 07:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1008.eqiad.wmnet with OS trixie * 07:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 07:00 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 06:59 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1008.eqiad.wmnet with OS trixie * 06:57 Emperor: rebalance thanos swift rings after previous re-image of thanos-fe1004 to trixie * 06:47 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1008.eqiad.wmnet with OS trixie * 04:10 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 14 days, 0:00:00 on cp6008.drmrs.wmnet with reason: Hardware failure - [[phab:T431651|T431651]] * 03:55 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp6008.* * 03:29 ryankemper: [[phab:T431311|T431311]] Repooled eqiad cirrussearch clusters (`chi/omega/psi`) following completion of OpenSearch 2.19 migration * 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad * 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=eqiad * 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 31s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-08 == * 23:52 Amir1: ladsgroup@deploy2003:~$ mwscript-k8s --follow -- extensions/ORES/maintenance/PurgeScoreCache.php --wiki=simplewiki --model damaging --old ([[phab:T431159|T431159]]) * 23:46 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 23:46 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing PTR for 2001:df2:e500:fe08::1 - cmooney@cumin1003" * 23:46 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing PTR for 2001:df2:e500:fe08::1 - cmooney@cumin1003" * 23:40 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 23:16 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 23:15 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 22:42 rzl@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 22:40 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] (duration: 12m 55s) * 22:40 rzl@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 22:37 rzl@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 22:36 rzl@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 22:35 rzl@deploy2003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 22:34 urbanecm@deploy2003: urbanecm: Continuing with deployment * 22:33 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:33 rzl@deploy2003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 22:32 rzl@deploy2003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 22:30 rzl@deploy2003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 22:30 rzl@deploy2003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 22:29 rzl@deploy2003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 22:27 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] * 22:26 rzl@deploy2003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 22:22 rzl@deploy2003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 22:21 rzl@deploy2003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 22:19 rzl@deploy2003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 22:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 22:17 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 22:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 22:13 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 22:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 22:13 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 22:09 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 22:06 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 22:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1094.eqiad.wmnet with OS trixie * 22:01 urbanecm: Make https://test.wikipedia.org/w/index.php?title=MediaWiki:GrowthExperimentsSuggestedEdits.json&diff=prev&oldid=750552 with GrowthExperiments disabled (via mw-experimental), then run `\MediaWiki\MediaWikiServices::getInstance()->get('CommunityConfiguration.ProviderFactory')->newProvider('GrowthSuggestedEdits')->getStore()->invalidate()` ([[phab:T431625|T431625]]) * 21:56 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d2-codfw * 21:55 urbanecm@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 21:55 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d2-codfw * 21:55 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c4-codfw * 21:55 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c4-codfw * 21:55 urbanecm@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2002 * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2002 * 21:54 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2002 * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2002.codfw.wmnet 50.32.192.10.in-addr.arpa 0.5.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:54 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2002.codfw.wmnet 50.32.192.10.in-addr.arpa 0.5.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2002 - bking@cumin2003" * 21:54 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2002 - bking@cumin2003" * 21:49 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:49 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2002 * 21:49 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2002.codfw.wmnet with OS trixie * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1094.eqiad.wmnet with reason: host reimage * 21:42 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 21:39 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 21:37 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1094.eqiad.wmnet with reason: host reimage * 21:36 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 21:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 21:29 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 21:27 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 21:22 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1094.eqiad.wmnet with OS trixie * 21:21 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host restbase2039.codfw.wmnet with OS bullseye * 21:21 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin2002" * 21:21 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin2002" * 21:04 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on restbase2039.codfw.wmnet with reason: host reimage * 21:00 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on restbase2039.codfw.wmnet with reason: host reimage * 20:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1073.eqiad.wmnet with OS trixie * 20:48 mutante: deploy2003 - kill 1102 (stunnel4) ; systemctl start stunnel4 ([[phab:T418262|T418262]]) * 20:42 cjming@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] (duration: 33m 02s) * 20:42 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host restbase2039.codfw.wmnet with OS bullseye * 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1073.eqiad.wmnet with reason: host reimage * 20:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1073.eqiad.wmnet with reason: host reimage * 20:30 cjming@deploy2003: cjming: Continuing with deployment * 20:28 cjming@deploy2003: cjming: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1098.eqiad.wmnet with OS trixie * 20:13 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1073.eqiad.wmnet with OS trixie * 20:09 cjming@deploy2003: Started scap sync-world: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] * 20:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1098.eqiad.wmnet with reason: host reimage * 19:56 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1098.eqiad.wmnet with reason: host reimage * 19:55 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d4-codfw * 19:54 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d4-codfw * 19:54 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c1-codfw * 19:54 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c1-codfw * 19:52 mutante: restarting gerrit on gerrit.wikimedia.org (gerrit2003) * 19:48 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2331.codfw.wmnet * 19:48 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2331.codfw.wmnet * 19:48 mutante: restarting gerrit on gerrit-replica.wikimedia.org (gerrit1003) * 19:46 mutante: restarting gerrit on gerrit-spare.wikimedia.org (gerrit2002) * 19:43 jasmine@cumin2002: conftool action : set/pooled=yes; selector: name=wikikube-worker2331.codfw.wmnet,cluster=kubernetes,service=kubesvc * 19:43 jasmine@cumin2002: conftool action : set/weight=10; selector: name=wikikube-worker2331.codfw.wmnet,cluster=kubernetes,service=kubesvc * 19:40 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1098.eqiad.wmnet with OS trixie * 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d5-codfw * 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d5-codfw * 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c7-codfw * 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c7-codfw * 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c5-codfw * 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c5-codfw * 19:30 jasmine_: ran homer on lsw1-d8-codfw, adding wikikube-worker2331 to cluster * 19:29 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1100.eqiad.wmnet with OS trixie * 19:20 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d8-codfw * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d8-codfw * 19:19 mutante: gerrit - replacing private key for registerEmail verification * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-magru * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device cr2-magru * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d7-codfw * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d7-codfw * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d3-codfw * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d1-codfw * 19:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d1-codfw * 19:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c2-codfw * 19:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c2-codfw * 19:11 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-codfw * 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-magru * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device cr1-magru * 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d8-codfw * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d8-codfw * 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d6-codfw * 19:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1100.eqiad.wmnet with reason: host reimage * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d6-codfw * 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c6-codfw * 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c6-codfw * 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c3-codfw * 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c3-codfw * 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b4-magru * 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device asw1-b4-magru * 19:08 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b3-magru * 19:08 cmooney@cumin1003: START - Cookbook sre.network.tls for network device asw1-b3-magru * 19:05 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1100.eqiad.wmnet with reason: host reimage * 19:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1122.eqiad.wmnet with OS trixie * 19:00 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 18:59 topranks: rolling out update to BGP ACL on Nokia Switches eqiad, codfw & ulsfo [[phab:T425703|T425703]] * 18:58 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 18:57 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 18:55 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 18:53 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 18:52 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 18:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1100.eqiad.wmnet with OS trixie * 18:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1068.eqiad.wmnet with OS trixie * 18:47 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1102.eqiad.wmnet with OS trixie * 18:47 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 18:46 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1122.eqiad.wmnet with reason: host reimage * 18:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1122.eqiad.wmnet with reason: host reimage * 18:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1068.eqiad.wmnet with reason: host reimage * 18:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1122 * 18:26 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1122 * 18:25 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1122 * 18:25 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1122.eqiad.wmnet 31.48.64.10.in-addr.arpa 1.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:25 bking@cumin2003: START - Cookbook sre.dns.wipe-cache cirrussearch1122.eqiad.wmnet 31.48.64.10.in-addr.arpa 1.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:25 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:25 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1122 - bking@cumin2003" * 18:25 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1122 - bking@cumin2003" * 18:21 rzl@deploy2003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 18:21 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1068.eqiad.wmnet with reason: host reimage * 18:21 rzl@deploy2003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 18:21 rzl@deploy2003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 18:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 18:19 rzl@deploy2003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 18:19 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:18 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1122 * 18:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1122.eqiad.wmnet with OS trixie * 18:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 18:15 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 18:13 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 18:13 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 18:10 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 18:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1068.eqiad.wmnet with OS trixie * 18:01 kamila@deploy2003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 18m 29s) * 18:00 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:55 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:42 kamila@deploy2003: Started scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] * 17:42 kamila@deploy2003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 19m 50s) * 17:42 kamila@deploy2003: Rolling back deployment * 17:35 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:31 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet * 17:18 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet * 17:16 kamila@deploy1003: Unlocked for deployment [MediaWiki]: switching deployment server (duration: 22m 07s) * 17:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 17:11 kamila@dns1005: END - running authdns-update * 17:09 kamila@dns1005: START - running authdns-update * 17:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 17:04 jasmine@cumin2002: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1164.eqiad.wmnet * 17:04 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1164.eqiad.wmnet * 17:04 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1164.eqiad.wmnet * 16:56 kamila@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on releases2003.codfw.wmnet,releases1003.eqiad.wmnet with reason: Deployment server switchover * 16:54 kamila@deploy1003: Locking from deployment [MediaWiki]: switching deployment server * 16:53 kamila@deploy1003: Unlocked for deployment [MediaWiki]: switching deployment server (duration: 04m 02s) * 16:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie * 16:49 kamila@deploy1003: Locking from deployment [MediaWiki]: switching deployment server * 16:46 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1095.eqiad.wmnet with OS trixie * 16:45 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1093.eqiad.wmnet with OS trixie * 16:43 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1164.eqiad.wmnet with OS trixie * 16:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1095.eqiad.wmnet with reason: host reimage * 16:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 16:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 16:23 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1164.eqiad.wmnet with reason: host reimage * 16:18 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on cirrussearch1093.eqiad.wmnet with reason: host reimage * 16:16 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1164.eqiad.wmnet with reason: host reimage * 16:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1095.eqiad.wmnet with reason: host reimage * 16:09 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 16:09 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 16:08 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1093.eqiad.wmnet with reason: host reimage * 15:59 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Pool test * 15:59 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:59 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 15:59 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Pool test * 15:58 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Depool test * 15:58 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:58 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 15:58 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Depool test * 15:57 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1164 * 15:57 jasmine@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1164 * 15:57 jasmine@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1164 * 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1164.eqiad.wmnet 114.48.64.10.in-addr.arpa 4.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:56 jasmine@cumin2002: START - Cookbook sre.dns.wipe-cache wikikube-worker1164.eqiad.wmnet 114.48.64.10.in-addr.arpa 4.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1164 - jasmine@cumin2002" * 15:56 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1164 - jasmine@cumin2002" * 15:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1093.eqiad.wmnet with OS trixie * 15:51 jasmine@cumin2002: START - Cookbook sre.dns.netbox * 15:51 jasmine@cumin2002: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1164 * 15:50 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-worker1164.eqiad.wmnet with OS trixie * 15:50 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1164.eqiad.wmnet * 15:50 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1164.eqiad.wmnet * 15:50 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1164.eqiad.wmnet * 15:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1095.eqiad.wmnet with OS trixie * 15:42 jasmine@cumin2002: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1164.eqiad.wmnet * 15:42 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1164.eqiad.wmnet * 15:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:42 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1164.eqiad.wmnet * 15:42 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1164.eqiad.wmnet * 15:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 15:39 elukey@cumin1003: START - Cookbook sre.hosts.provision for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 15:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1007.eqiad.wmnet with OS trixie * 15:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Pool test * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1007.eqiad.wmnet with reason: host reimage * 15:15 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 15:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet * 15:15 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 15:15 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1007.eqiad.wmnet with reason: host reimage * 15:15 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:14 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Pool test * 15:14 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet * 15:14 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2228: Depool test * 15:14 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db2228: Depool test * 15:10 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 15:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:08 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 15:06 blake@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 15:06 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet * 15:06 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 15:06 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 15:05 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet * 15:05 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 15:04 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:04 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 15:04 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:03 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 15:03 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 15:03 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:03 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T430909|T430909]] * 15:03 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:03 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 15:03 swfrench-wmf: restarted eqsin, codfw confds - [[phab:T430909|T430909]] * 15:03 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test1002.eqiad.wmnet * 15:02 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet * 14:59 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:59 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:55 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1007.eqiad.wmnet with OS trixie * 14:52 swfrench-wmf: restarted ulsfo confds, confirmed now connected to codfw backends except those using wikimedia.org SRV record - [[phab:T430909|T430909]] * 14:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:49 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:41 moritzm: uninstalling dhcpcd-base from trixie hosts which still have it installed [[phab:T414341|T414341]] * 14:40 sukhe: sudo cumin -b1 -s120 "P<nowiki>{</nowiki>lvs2011*<nowiki>}</nowiki> or P<nowiki>{</nowiki>lvs2012*<nowiki>}</nowiki>" "systemctl restart pybal.service" * 14:39 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:39 mvernon@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host thanos-be1007.eqiad.wmnet with OS trixie * 14:37 sukhe: restart pybal on lvs2013 to revert back to conf2004 * 14:35 sukhe: restart pybal on lvs2014 to revert back to conf2004 * 14:34 swfrench-wmf: switched codfw, eqsin, ulsfo etcd client SRV records back to codfw - [[phab:T430909|T430909]] * 14:32 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1002.eqiad.wmnet * 14:32 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet * 14:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1007.eqiad.wmnet with OS trixie * 14:31 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:31 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:31 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:30 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:30 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Pool test * 14:30 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:29 swfrench@dns1004: END - running authdns-update * 14:29 moritzm: installing jackson-core security updates * 14:27 swfrench@dns1004: START - running authdns-update * 14:22 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:22 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:22 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1119.eqiad.wmnet with OS trixie * 14:22 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:21 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:20 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:20 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 14:20 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:19 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 14:19 moritzm: installing librabbitmq security updates * 14:19 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1002.eqiad.wmnet * 14:18 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet * 14:16 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:15 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:15 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:14 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Pool test * 14:14 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox) * 14:14 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet * 14:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 14:08 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 14:07 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet * 14:05 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1118.eqiad.wmnet with OS trixie * 14:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test1001.eqiad.wmnet * 14:00 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-worker@eqiad * 14:00 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 13:59 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 13:58 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 13:57 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1119.eqiad.wmnet with reason: host reimage * 13:54 moritzm: installing libcap2 security updates * 13:53 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1119.eqiad.wmnet with reason: host reimage * 13:52 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet * 13:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) * 13:52 fceratto@cumin1003: START - Cookbook sre.mysql.depool * 13:50 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-worker@eqiad * 13:50 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1051.eqiad.wmnet * 13:50 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1051.eqiad.wmnet * 13:50 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1051.eqiad.wmnet * 13:49 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 13:45 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1081.eqiad.wmnet with OS trixie * 13:41 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1119 * 13:41 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1119 * 13:40 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1119 * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1119.eqiad.wmnet 97.32.64.10.in-addr.arpa 7.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1119.eqiad.wmnet 97.32.64.10.in-addr.arpa 7.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1119 - atsuko@cumin1003" * 13:40 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1119 - atsuko@cumin1003" * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1118.eqiad.wmnet with reason: host reimage * 13:39 moritzm: installing krb5 security updates * 13:37 Lucas_WMDE: UTC afternoon backport+config window done * 13:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1006.eqiad.wmnet with OS trixie * 13:36 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1118.eqiad.wmnet with reason: host reimage * 13:36 atsuko@cumin1003: START - Cookbook sre.dns.netbox * 13:35 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] (duration: 07m 46s) * 13:34 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1119 * 13:34 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1119.eqiad.wmnet with OS trixie * 13:30 sbisson@deploy1003: sbisson: Continuing with deployment * 13:30 moritzm: installing openssh security updates * 13:30 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling restart_daemons on A:wikidough * 13:29 sbisson@deploy1003: sbisson: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:27 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] * 13:26 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1051.eqiad.wmnet with OS trixie * 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1081.eqiad.wmnet with reason: host reimage * 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1118 * 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1118 * 13:22 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] (duration: 12m 12s) * 13:21 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1081.eqiad.wmnet with reason: host reimage * 13:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1006.eqiad.wmnet with reason: host reimage * 13:18 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1118 * 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1118.eqiad.wmnet 90.32.64.10.in-addr.arpa 0.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:18 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1118.eqiad.wmnet 90.32.64.10.in-addr.arpa 0.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1118 - atsuko@cumin1003" * 13:18 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1118 - atsuko@cumin1003" * 13:17 stran@deploy1003: stran: Continuing with deployment * 13:16 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough * 13:15 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:13 atsuko@cumin1003: START - Cookbook sre.dns.netbox * 13:12 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1118 * 13:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1006.eqiad.wmnet with reason: host reimage * 13:12 stran@deploy1003: stran: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:12 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1118.eqiad.wmnet with OS trixie * 13:10 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] * 13:05 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-worker@codfw * 13:05 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 13:05 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1081.eqiad.wmnet with OS trixie * 13:05 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1051.eqiad.wmnet with reason: host reimage * 13:04 moritzm: installing jq security updates * 13:04 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 13:01 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1051.eqiad.wmnet with reason: host reimage * 12:58 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-worker@codfw * 12:52 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 12:50 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host thanos-be1006.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1051 * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1051 * 12:43 moritzm: installing Python 3.11 security updates * 12:43 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1051 * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1051.eqiad.wmnet 46.32.64.10.in-addr.arpa 6.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:43 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1051.eqiad.wmnet 46.32.64.10.in-addr.arpa 6.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1051 - blake@cumin1003" * 12:43 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1051 - blake@cumin1003" * 12:38 blake@cumin1003: START - Cookbook sre.dns.netbox * 12:38 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1051 * 12:38 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1051.eqiad.wmnet with OS trixie * 12:37 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1051.eqiad.wmnet * 12:36 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1051.eqiad.wmnet * 12:36 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1051.eqiad.wmnet * 12:34 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1006.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 12:34 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1006.eqiad.wmnet with OS trixie * 12:27 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 12:27 mvernon@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host thanos-be1006.eqiad.wmnet with OS trixie * 12:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:02 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 12:01 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1006.eqiad.wmnet with OS trixie * 11:43 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 11:38 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1076.eqiad.wmnet with OS trixie * 11:26 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1075.eqiad.wmnet with OS trixie * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2047.codfw.wmnet * 11:19 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2047.codfw.wmnet * 11:19 moritzm: temporarily remove ganeti2031 from codfw cluster [[phab:T430910|T430910]] * 11:08 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1076.eqiad.wmnet with reason: host reimage * 11:08 moritzm: installing Linux 6.1.176 on Bookworm servers * 11:03 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1076.eqiad.wmnet with reason: host reimage * 11:00 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1075.eqiad.wmnet with reason: host reimage * 10:56 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1075.eqiad.wmnet with reason: host reimage * 10:47 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1076.eqiad.wmnet with OS trixie * 10:46 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1074.eqiad.wmnet with OS trixie * 10:45 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1005.eqiad.wmnet with OS trixie * 10:40 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1075.eqiad.wmnet with OS trixie * 10:32 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2031.codfw.wmnet * 10:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1005.eqiad.wmnet with reason: host reimage * 10:25 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1005.eqiad.wmnet with reason: host reimage * 10:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1074.eqiad.wmnet with reason: host reimage * 10:17 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1074.eqiad.wmnet with reason: host reimage * 10:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1005.eqiad.wmnet with OS trixie * 10:12 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet * 10:04 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 10:01 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet * 10:01 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1074.eqiad.wmnet with OS trixie * 10:01 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 09:43 cgoubert@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/aux-k8s-services/redioscope: apply * 09:43 cgoubert@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/aux-k8s-services/redioscope: apply * 09:43 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply * 09:35 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply * 09:34 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 09:34 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 09:33 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 41 days, 15:00:00 on db2252.codfw.wmnet with reason: Test * 09:32 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 09:32 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 09:31 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: codfw rack B3 pool after maintenance * 09:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 09:07 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 09:07 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 09:02 ladsgroup@cumin1003: END (PASS) - Cookbook sre.mysql.sanitarium_restart (exit_code=0) * 08:57 topranks: merge patch to shift eqiad <-> esams traffic onto new 40G circuit * 08:54 hashar@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1004.eqiad.wmnet with OS trixie * 08:50 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 08:50 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitarium_restart (exit_code=99) * 08:50 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 08:45 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool es2051: codfw rack B3 pool after maintenance * 08:44 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:44 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:43 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2007.codfw.wmnet * 08:43 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2007.codfw.wmnet * 08:42 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2031.codfw.wmnet * 08:41 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2031.codfw.wmnet * 08:40 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2031.codfw.wmnet * 08:38 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:38 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:35 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.sanitize-wiki (exit_code=97) Managing sanitization for wikis minwikiquote in section s3 * 08:33 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis minwikiquote in section s3 * 08:32 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Checking sanitization for wikis minwikiquote in section s5 * 08:30 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Checking sanitization for wikis minwikiquote in section s5 * 08:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Managing sanitization for wikis minwikiquote in section s5 * 08:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1004.eqiad.wmnet with reason: host reimage * 08:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1004.eqiad.wmnet with reason: host reimage * 08:23 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:22 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis minwikiquote in section s5 * 08:19 XioNoX: lsw1-b3-codfw> request system reboot - [[phab:T430909|T430909]] * 08:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Checking sanitization for wikis minwikiquote in section s5 * 08:17 hashar@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Checking sanitization for wikis minwikiquote in section s5 * 08:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for codfw rack B3 * 08:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2007.codfw.wmnet * 08:15 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lsw1-b3-codfw,lsw1-b3-codfw IPv6,lsw1-b3-codfw.mgmt with reason: Switch maintenance * 08:15 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2007.codfw.wmnet * 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:07 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:06 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: codfw rack B3 depool for maintenance * 08:05 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool es2051: codfw rack B3 depool for maintenance * 08:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1004.eqiad.wmnet with OS trixie * 08:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1005.eqiad.wmnet with OS trixie * 08:03 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 21 hosts with reason: codfw rack B3 depool for maintenance * 07:56 ayounsi@cumin1003: START - Cookbook sre.network.depool-rack with action 'depool' for codfw rack B3 * 07:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1005.eqiad.wmnet with reason: host reimage * 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1005.eqiad.wmnet with reason: host reimage * 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1005.eqiad.wmnet with OS bookworm * 07:29 moritzm: installing gnutls28 security updates * 07:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1005.eqiad.wmnet with OS trixie * 07:13 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1125.eqiad.wmnet with OS trixie * 07:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1005.eqiad.wmnet with reason: host reimage * 07:07 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aux-k8s-etcd1005.eqiad.wmnet with reason: host reimage * 06:56 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1005.eqiad.wmnet with OS bookworm * 06:54 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1125.eqiad.wmnet with reason: host reimage * 06:52 elukey: upgrade all trixie hosts to pywmflib 3.1 - [[phab:T430552|T430552]] * 06:50 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1125.eqiad.wmnet with reason: host reimage * 06:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 06:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 06:38 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1125.eqiad.wmnet with OS trixie * 05:42 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1107.eqiad.wmnet with OS trixie * 05:35 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1124.eqiad.wmnet with OS trixie * 05:31 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1101.eqiad.wmnet with OS trixie * 05:21 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1107.eqiad.wmnet with reason: host reimage * 05:17 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1124.eqiad.wmnet with reason: host reimage * 05:13 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1107.eqiad.wmnet with reason: host reimage * 05:13 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1101.eqiad.wmnet with reason: host reimage * 05:11 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1124.eqiad.wmnet with reason: host reimage * 05:10 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1101.eqiad.wmnet with reason: host reimage * 04:58 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1124.eqiad.wmnet with OS trixie * 04:56 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1107.eqiad.wmnet with OS trixie * 04:55 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1101.eqiad.wmnet with OS trixie * 02:27 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] (duration: 08m 14s) * 02:22 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 02:21 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 02:19 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] * 01:59 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] (duration: 09m 46s) * 01:55 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 01:51 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 01:49 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] * 01:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1099.eqiad.wmnet with OS trixie * 00:57 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1110.eqiad.wmnet with OS trixie * 00:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1099.eqiad.wmnet with reason: host reimage * 00:41 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1099.eqiad.wmnet with reason: host reimage * 00:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1110.eqiad.wmnet with reason: host reimage * 00:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1110.eqiad.wmnet with reason: host reimage * 00:26 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1099.eqiad.wmnet with OS trixie * 00:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1110.eqiad.wmnet with OS trixie == 2026-07-07 == * 22:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1097.eqiad.wmnet with OS trixie * 22:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1097.eqiad.wmnet with reason: host reimage * 22:24 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1097.eqiad.wmnet with reason: host reimage * 22:14 hashar: Restarting Gerrit on gerrit2002 and gerrit1003 (replicas) * 22:09 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1097.eqiad.wmnet with OS trixie * 22:07 hashar: Restarting Gerrit on gerrit2003 * 21:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 21:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 21:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 21:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 21:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1108.eqiad.wmnet with OS trixie * 20:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1091.eqiad.wmnet with OS trixie * 20:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1090.eqiad.wmnet with OS trixie * 20:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1108.eqiad.wmnet with reason: host reimage * 20:36 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1091.eqiad.wmnet with reason: host reimage * 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1090.eqiad.wmnet with reason: host reimage * 20:33 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1091.eqiad.wmnet with reason: host reimage * 20:30 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1108.eqiad.wmnet with reason: host reimage * 20:30 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1006.eqiad.wmnet * 20:30 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1090.eqiad.wmnet with reason: host reimage * 20:30 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1006.eqiad.wmnet * 20:27 jasmine_: "homer lsw1-c2-eqiad* commit "Added new stacked control plane wikikube-ctrl1006"" * 20:22 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] (duration: 07m 29s) * 20:20 jasmine_: "homer "cr*eqiad*" commit "Added new stacked control plane wikikube-ctrl1006"" * 20:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1091.eqiad.wmnet with OS trixie * 20:17 arlolra@deploy1003: arlolra: Continuing with deployment * 20:16 arlolra@deploy1003: arlolra: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:16 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1090.eqiad.wmnet with OS trixie * 20:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1108.eqiad.wmnet with OS trixie * 20:14 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] * 20:09 cwhite: remove 2026-04 swift log archives from centrallog2002 to free some space * 20:01 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=93) for host cirrussearch1108.eqiad.wmnet with OS trixie * 19:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1108.eqiad.wmnet with OS trixie * 19:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1090.eqiad.wmnet with OS trixie * 19:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1109.eqiad.wmnet with OS trixie * 19:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1092.eqiad.wmnet with OS trixie * 19:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1123.eqiad.wmnet with OS trixie * 19:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1109.eqiad.wmnet with reason: host reimage * 19:19 cdobbins@cumin2002: conftool action : set/pooled=yes; selector: name=dns7002.* * 19:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1092.eqiad.wmnet with reason: host reimage * 19:17 jasmine@dns1004: END - running authdns-update * 19:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1109.eqiad.wmnet with reason: host reimage * 19:15 jasmine@dns1004: START - running authdns-update * 19:15 cdobbins@dns1004: END - running authdns-update * 19:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1123.eqiad.wmnet with reason: host reimage * 19:13 cdobbins@dns1004: START - running authdns-update * 19:12 cdobbins@cumin2002: conftool action : set/pooled=yes; selector: name=dns7002.*,service=authdns-update * 19:11 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1092.eqiad.wmnet with reason: host reimage * 19:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1123.eqiad.wmnet with reason: host reimage * 18:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1123.eqiad.wmnet with OS trixie * 18:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1109.eqiad.wmnet with OS trixie * 18:56 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1092.eqiad.wmnet with OS trixie * 18:52 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 18:49 swfrench@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 18:40 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 18:38 swfrench@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 18:11 swfrench-wmf: restarted eqsin, codfw confds - [[phab:T430909|T430909]] * 18:01 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T430909|T430909]] * 17:59 swfrench-wmf: restarted ulsfo confds, confirmed now connected to eqiad backends - [[phab:T430909|T430909]] * 17:52 sukhe: restart pybal on lvs2011 to switch from conf2004 to conf1008: [[phab:T430909|T430909]] * 17:51 sukhe: restart pybal on lvs2012 to switch from conf2004 to conf1008 [puppet re-enabled there]: [[phab:T430909|T430909]] * 17:46 sukhe: restart pybal on lvs2013 to switch from conf2004 to conf1008: [[phab:T430909|T430909]] * 17:44 swfrench-wmf: switched codfw, eqsin, ulsfo etcd client SRV records to eqiad - [[phab:T430909|T430909]] * 17:43 swfrench@dns1004: END - running authdns-update * 17:40 swfrench@dns1004: START - running authdns-update * 17:40 sukhe: restart pybal on lvs2014 to switch from conf2004 to conf1008: [[phab:T430909|T430909]] * 17:21 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1003.eqiad.wmnet * 17:15 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1003.eqiad.wmnet * 17:14 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1002.eqiad.wmnet * 17:06 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1002.eqiad.wmnet * 17:06 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-low-traffic-codfw' 'systemctl restart pybal.service' # lvs2013, [[phab:T416623|T416623]] * 17:04 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1001.eqiad.wmnet * 17:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1111.eqiad.wmnet with OS trixie * 17:00 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal.service' # lvs2014, [[phab:T416623|T416623]] * 16:58 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1001.eqiad.wmnet * 16:58 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS bookworm * 16:55 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-low-traffic-eqiad' 'systemctl restart pybal.service' # lvs1019, [[phab:T416623|T416623]] * 16:53 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-secondary-eqiad' 'systemctl restart pybal.service' # lvs1020, [[phab:T416623|T416623]] * 16:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1111.eqiad.wmnet with reason: host reimage * 16:40 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1111.eqiad.wmnet with reason: host reimage * 16:38 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.peering (exit_code=99) with action 'configure' for AS: 47794 * 16:35 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 47794 * 16:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1111.eqiad.wmnet with OS trixie * 16:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1006.eqiad.wmnet with OS trixie * 16:06 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 16:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1006.eqiad.wmnet with reason: host reimage * 15:58 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1121.eqiad.wmnet with OS trixie * 15:58 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 15:56 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1006.eqiad.wmnet with reason: host reimage * 15:54 mutante: jenkins down in planned maintenance window * 15:42 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1037.eqiad.wmnet * 15:42 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1037.eqiad.wmnet * 15:42 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1037.eqiad.wmnet * 15:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1006.eqiad.wmnet with OS trixie * 15:34 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1121.eqiad.wmnet with reason: host reimage * 15:33 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS bookworm * 15:33 cdobbins@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host dns7002.wikimedia.org with OS trixie * 15:30 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1121.eqiad.wmnet with reason: host reimage * 15:29 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host clouddumps1001.wikimedia.org * 15:20 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1001.wikimedia.org * 15:18 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1121 * 15:18 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1121 * 15:18 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host clouddumps1002.wikimedia.org * 15:17 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1121 * 15:17 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:17 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply * 15:16 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply * 15:16 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:16 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1121 - atsuko@cumin1003" * 15:16 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1121 - atsuko@cumin1003" * 15:14 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1037.eqiad.wmnet with OS trixie * 15:11 atsuko@cumin1003: START - Cookbook sre.dns.netbox * 15:09 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org * 15:09 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1121 * 15:09 andrew@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host clouddumps1002.wikimedia.org * 15:09 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org * 15:09 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1121.eqiad.wmnet with OS trixie * 15:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1007.eqiad.wmnet with OS trixie * 15:08 andrew@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host clouddumps1002.wikimedia.org * 15:08 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org * 15:05 brennen@deploy1003: Finished deploy [phabricator/deployment@7e02037]: deploy phab1004 for [[phab:T431440|T431440]] (duration: 00m 47s) * 15:04 brennen@deploy1003: Started deploy [phabricator/deployment@7e02037]: deploy phab1004 for [[phab:T431440|T431440]] * 15:03 brennen@deploy1003: Finished deploy [phabricator/deployment@7e02037]: deploy phab2003 for [[phab:T431440|T431440]] (duration: 00m 51s) * 15:03 brennen@deploy1003: Started deploy [phabricator/deployment@7e02037]: deploy phab2003 for [[phab:T431440|T431440]] * 15:00 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71] (thin): Regular analytics weekly train THIN [analytics/refinery@7d8dc71f] (duration: 02m 10s) * 14:58 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71] (thin): Regular analytics weekly train THIN [analytics/refinery@7d8dc71f] * 14:58 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71]: Regular analytics weekly train [analytics/refinery@7d8dc71f] (duration: 04m 14s) * 14:54 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1037.eqiad.wmnet with reason: host reimage * 14:53 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71]: Regular analytics weekly train [analytics/refinery@7d8dc71f] * 14:53 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@7d8dc71f] (duration: 02m 00s) * 14:51 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@7d8dc71f] * 14:51 arnaudb@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on phab2003.codfw.wmnet,phab[1004-1006].eqiad.wmnet with reason: maintenance * 14:51 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1037.eqiad.wmnet with reason: host reimage * 14:50 JavierMonton: Deploying Refinery at {{Gerrit|7d8dc71f}} for change {{Gerrit|1308087}} / [[phab:T431318|T431318]] - update filerevision table sqoop and table * 14:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1007.eqiad.wmnet with reason: host reimage * 14:42 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1007.eqiad.wmnet with reason: host reimage * 14:40 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1083.eqiad.wmnet with OS trixie * 14:37 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:36 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:35 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-master@eqiad * 14:35 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 14:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 14:34 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1037 * 14:34 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1037 * 14:34 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:cleanMentorList.php --wiki=frwiki # [[phab:T427386|T427386]] * 14:34 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 14:34 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308112{{!}}Revert^2 "[Growth] frwiki: Deploy automated mentor list cleaner" (T427386)]] (duration: 06m 47s) * 14:34 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 14:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:33 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:32 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1037 * 14:31 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:31 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:29 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:29 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-master@eqiad * 14:29 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:29 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:29 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:28 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:27 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:27 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1308112{{!}}Revert^2 "[Growth] frwiki: Deploy automated mentor list cleaner" (T427386)]] * 14:26 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1007.eqiad.wmnet with OS trixie * 14:26 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:26 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 14:26 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:25 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:25 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:25 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:cleanMentorList.php --wiki=frwiki # [[phab:T427386|T427386]] * 14:24 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:24 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1037 - blake@cumin1003" * 14:24 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1037 - blake@cumin1003" * 14:20 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1083.eqiad.wmnet with reason: host reimage * 14:19 blake@cumin1003: START - Cookbook sre.dns.netbox * 14:19 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-master@codfw * 14:19 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 14:19 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1037 * 14:18 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1037.eqiad.wmnet with OS trixie * 14:18 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1037.eqiad.wmnet * 14:18 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 14:18 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1037.eqiad.wmnet * 14:18 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1037.eqiad.wmnet * 14:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2007.codfw.wmnet with OS trixie * 14:16 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1083.eqiad.wmnet with reason: host reimage * 14:15 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1036.eqiad.wmnet * 14:15 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1036.eqiad.wmnet * 14:14 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1036.eqiad.wmnet * 14:12 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-master@codfw * 14:11 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1120.eqiad.wmnet with OS trixie * 14:05 moritzm: installing distro-info-data updates from trixie/bookworm point releases * 14:04 fabfur: disable puppet on A:cp-text to selectively apply https://gerrit.wikimedia.org/r/c/operations/puppet/+/1308040 * 14:03 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] (duration: 27m 48s) * 14:00 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 14:00 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1083.eqiad.wmnet with OS trixie * 13:58 urbanecm@deploy1003: urbanecm: Continuing with deployment * 13:58 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:58 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1004.eqiad.wmnet with OS bookworm * 13:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2007.codfw.wmnet with reason: host reimage * 13:57 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 13:53 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1120.eqiad.wmnet with reason: host reimage * 13:50 moritzm: installing Linux 5.10.259 on Bullseye hosts * 13:47 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply * 13:47 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply * 13:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2007.codfw.wmnet with reason: host reimage * 13:46 cgoubert@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/aux-k8s-services/redioscope: apply * 13:46 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1120.eqiad.wmnet with reason: host reimage * 13:46 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:46 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:45 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:44 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:44 cgoubert@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/aux-k8s-services/redioscope: apply * 13:40 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 13:39 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:38 moritzm: installing e2fsprogs updates from Trixie point release * 13:35 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] * 13:33 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1120.eqiad.wmnet with OS trixie * 13:33 topranks: reset cr3-eqsin configuration so traffic uses it again after upgrade * 13:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1088.eqiad.wmnet with OS trixie * 13:32 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie * 13:32 cdobbins@cumin1003: conftool action : set/pooled=no; selector: name=dns7002.* * 13:29 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2007.codfw.wmnet with OS trixie * 13:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1004.eqiad.wmnet with reason: host reimage * 13:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2006.codfw.wmnet with OS trixie * 13:18 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1036.eqiad.wmnet with OS trixie * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aux-k8s-etcd1004.eqiad.wmnet with reason: host reimage * 13:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 13:16 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 13:15 jayme: Istio is being upgraded from 1.24.2 to 1.29.4 on wikikube staging eqiad and codfw - [[phab:T427401|T427401]] * 13:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1087.eqiad.wmnet with OS trixie * 13:14 topranks: reboot cr3-eqsin to install new JunOS and set PIC 0/0/0 to 100G * 13:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1088.eqiad.wmnet with reason: host reimage * 13:13 jmm@dns1004: END - running authdns-update * 13:12 jmm@dns1004: START - running authdns-update * 13:09 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1088.eqiad.wmnet with reason: host reimage * 13:07 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1082.eqiad.wmnet with OS trixie * 13:07 jmm@dns1004: END - running authdns-update * 13:06 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1004.eqiad.wmnet with OS bookworm * 13:05 jmm@dns1004: START - running authdns-update * 13:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2006.codfw.wmnet with reason: host reimage * 12:58 topranks: load updated JunOS on cr3-eqsin [[phab:T429386|T429386]] * 12:58 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1036.eqiad.wmnet with reason: host reimage * 12:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2001.codfw.wmnet * 12:57 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2006.codfw.wmnet with reason: host reimage * 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr1-codfw,cr[2-3]-eqsin,cr3-eqsin IPv6,cr3-eqsin.mgmt with reason: upgrade JunOS cr3-eqsin * 12:56 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lvs[5004-5006].eqsin.wmnet with reason: upgrade JunOS cr3-eqsin * 12:55 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 12:55 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 12:53 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1087.eqiad.wmnet with reason: host reimage * 12:53 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 12:52 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1088.eqiad.wmnet with OS trixie * 12:52 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1002.eqiad.wmnet * 12:52 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:51 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2001.codfw.wmnet * 12:49 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1036.eqiad.wmnet with reason: host reimage * 12:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 12:48 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1087.eqiad.wmnet with reason: host reimage * 12:44 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1082.eqiad.wmnet with reason: host reimage * 12:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1002.eqiad.wmnet * 12:42 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 12:42 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 12:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 12:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 12:39 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:39 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: move dumps-nfs IP to the shared one - filippo@cumin1003" * 12:39 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: move dumps-nfs IP to the shared one - filippo@cumin1003" * 12:39 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2006.codfw.wmnet with OS trixie * 12:38 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1082.eqiad.wmnet with reason: host reimage * 12:36 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:33 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:32 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1036 * 12:32 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1036 * 12:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2005.codfw.wmnet with OS trixie * 12:32 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1087.eqiad.wmnet with OS trixie * 12:30 jmm@dns1004: END - running authdns-update * 12:29 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1036 * 12:29 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1036.eqiad.wmnet 21.32.64.10.in-addr.arpa 1.2.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:29 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1036.eqiad.wmnet 21.32.64.10.in-addr.arpa 1.2.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:29 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:29 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1036 - blake@cumin1003" * 12:29 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1036 - blake@cumin1003" * 12:28 jmm@dns1004: START - running authdns-update * 12:26 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:26 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:23 blake@cumin1003: START - Cookbook sre.dns.netbox * 12:23 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1036 * 12:23 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1036.eqiad.wmnet with OS trixie * 12:22 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1036.eqiad.wmnet * 12:22 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1082.eqiad.wmnet with OS trixie * 12:22 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1036.eqiad.wmnet * 12:22 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1036.eqiad.wmnet * 12:21 marostegui: Restart mariadb@s7 on db1155 to pick up new filters - [[phab:T431124|T431124]] * 12:21 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 21 hosts with reason: restarting for replication filter * 12:20 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:19 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:14 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2005.codfw.wmnet with reason: host reimage * 12:14 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:08 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:07 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:07 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2005.codfw.wmnet with reason: host reimage * 12:06 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:06 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:06 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:05 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:05 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:04 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-master-eqiad * 12:04 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl1002.eqiad.wmnet * 12:04 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl1002.eqiad.wmnet * 12:04 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:04 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:03 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:03 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:03 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:03 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 11:59 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl1002.eqiad.wmnet * 11:59 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl1002.eqiad.wmnet * 11:59 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl1001.eqiad.wmnet * 11:59 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl1001.eqiad.wmnet * 11:56 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl1001.eqiad.wmnet * 11:56 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl1001.eqiad.wmnet * 11:56 klausman@cumin2002: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-master-eqiad * 11:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2005.codfw.wmnet with OS trixie * 11:49 blake@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on wikikube-worker1160.eqiad.wmnet with reason: Verifying matchers for silence * 11:42 topranks: cr3-eqsin, begin traffic drain to reset PIC and upgrade JunOS [[phab:T429386|T429386]] * 11:41 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs[5004-5006].eqsin.wmnet with reason: upgrade JunOS cr3-eqsin * 11:39 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr1-codfw,cr[2-3]-eqsin,cr3-eqsin IPv6,cr3-eqsin.mgmt with reason: upgrade JunOS cr3-eqsin * 11:36 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=thanos-fe2004.codfw.wmnet * 11:35 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1086.eqiad.wmnet with OS trixie * 11:35 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=thanos-fe2004.codfw.wmnet * 11:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1085.eqiad.wmnet with OS trixie * 11:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1086.eqiad.wmnet with reason: host reimage * 11:10 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1085.eqiad.wmnet with reason: host reimage * 11:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2004.codfw.wmnet with OS trixie * 11:03 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1086.eqiad.wmnet with reason: host reimage * 11:02 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1085.eqiad.wmnet with reason: host reimage * 10:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2004.codfw.wmnet with reason: host reimage * 10:48 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:46 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1086.eqiad.wmnet with OS trixie * 10:46 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1085.eqiad.wmnet with OS trixie * 10:44 cgoubert@deploy1003: Finished deploy [restbase/deploy@2fc37d4]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] (duration: 16m 44s) * 10:43 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2004.codfw.wmnet with reason: host reimage * 10:35 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:27 cgoubert@deploy1003: Started deploy [restbase/deploy@2fc37d4]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] * 10:27 cgoubert@deploy1003: Finished deploy [restbase/deploy@8a25036]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] (duration: 00m 45s) * 10:26 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1117.eqiad.wmnet with OS trixie * 10:26 cgoubert@deploy1003: Started deploy [restbase/deploy@8a25036]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] * 10:26 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host thanos-fe2004 * 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host thanos-fe2004 * 10:22 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1116.eqiad.wmnet with OS trixie * 10:21 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host thanos-fe2004 * 10:21 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) thanos-fe2004.codfw.wmnet 157.32.192.10.in-addr.arpa 7.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:20 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache thanos-fe2004.codfw.wmnet 157.32.192.10.in-addr.arpa 7.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:20 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:20 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host thanos-fe2004 - mvernon@cumin2003" * 10:20 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host thanos-fe2004 - mvernon@cumin2003" * 10:15 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2252: Repooling after reboot * 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:15 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 10:15 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2252: Repooling after reboot * 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1153.eqiad.wmnet * 10:14 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1153.eqiad.wmnet * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 10:14 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 10:12 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 10:12 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host thanos-fe2004 * 10:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2004.codfw.wmnet with OS trixie * 10:07 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1117.eqiad.wmnet with reason: host reimage * 10:03 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1116.eqiad.wmnet with reason: host reimage * 09:58 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:58 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1117.eqiad.wmnet with reason: host reimage * 09:57 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1116.eqiad.wmnet with reason: host reimage * 09:49 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 41 days, 15:00:00 on db2252.codfw.wmnet with reason: Security updates * 09:45 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1117.eqiad.wmnet with OS trixie * 09:45 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1116.eqiad.wmnet with OS trixie * 09:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1153: Security updates * 09:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:28 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:28 root@cumin1003: START - Cookbook sre.mysql.depool depool db1153: Security updates * 09:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1016: Security updates * 09:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:21 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:21 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1016: Security updates * 09:14 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:14 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 08:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1016: Security updates * 08:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:56 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:56 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1016: Security updates * 08:50 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:50 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:45 filippo@dns1006: END - running authdns-update * 08:43 filippo@dns1006: START - running authdns-update * 08:42 godog: switch dumps-nfs address to be shared with rsync/http - [[phab:T411248|T411248]] * 08:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1016: Security updates * 08:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:40 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:40 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1016: Security updates * 08:29 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host cirrussearch1111.eqiad.wmnet * 08:29 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:27 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:27 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:25 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1015: Security updates * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:09 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:09 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1015: Security updates * 07:42 Msz2001: Deployed private patch for Suggested Ivestigations * 07:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1015: Security updates * 07:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:41 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:41 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1015: Security updates * 07:40 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 07:11 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fingerprint warnings - oblivian@cumin1003" * 07:11 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fingerprint warnings - oblivian@cumin1003 * 07:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1024: Security updates * 07:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:11 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:11 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1024: Security updates * 07:10 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fingerprint warnings - oblivian@cumin1003 * 07:10 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fingerprint warnings - oblivian@cumin1003" * 06:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host cirrussearch1111.eqiad.wmnet * 06:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 06:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1024: Security updates * 06:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 06:48 root@cumin1003: START - Cookbook sre.mysql.parsercache * 06:48 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1024: Security updates * 06:42 moritzm: install nginx security updates * 06:31 root@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool pc1024: Security updates * 06:21 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1024: Security updates * 06:19 moritzm: installing php8.2 security updates * 06:15 moritzm: installing php8.4 security updates * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.7 (duration: 02m 38s) * 03:40 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] (duration: 37m 04s) * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 51s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-06 == * 23:30 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] (duration: 09m 39s) * 23:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1078.eqiad.wmnet with OS trixie * 23:26 jdlrobson@deploy1003: jdlrobson, bwang: Continuing with deployment * 23:22 jdlrobson@deploy1003: jdlrobson, bwang: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug) * 23:21 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] * 23:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1078.eqiad.wmnet with reason: host reimage * 23:06 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1078.eqiad.wmnet with reason: host reimage * 22:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1078.eqiad.wmnet with OS trixie * 22:29 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on cirrussearch1114.eqiad.wmnet with reason: reimage on hold until restore completes * 22:22 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on cirrussearch[1079,1115].eqiad.wmnet with reason: reimage on hold until restore completes * 21:18 maryum: Deployed security fix for [[phab:T428006|T428006]] * 20:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1079.eqiad.wmnet with OS trixie * 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1077.eqiad.wmnet with OS trixie * 20:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1115.eqiad.wmnet with OS trixie * 20:25 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1079.eqiad.wmnet with reason: host reimage * 20:21 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1079.eqiad.wmnet with reason: host reimage * 20:15 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] (duration: 08m 14s) * 20:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1077.eqiad.wmnet with reason: host reimage * 20:10 krinkle@deploy1003: krinkle, pushpaktiwari: Continuing with deployment * 20:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1115.eqiad.wmnet with reason: host reimage * 20:08 krinkle@deploy1003: krinkle, pushpaktiwari: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1077.eqiad.wmnet with reason: host reimage * 20:06 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] * 20:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1079.eqiad.wmnet with OS trixie * 20:04 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1115.eqiad.wmnet with reason: host reimage * 19:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1077.eqiad.wmnet with OS trixie * 19:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1115.eqiad.wmnet with OS trixie * 19:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 19:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 18:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1114.eqiad.wmnet with OS trixie * 18:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1114.eqiad.wmnet with reason: host reimage * 18:35 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1114.eqiad.wmnet with reason: host reimage * 18:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1112.eqiad.wmnet with OS trixie * 18:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1114.eqiad.wmnet with OS trixie * 18:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1072.eqiad.wmnet with OS trixie * 18:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1112.eqiad.wmnet with reason: host reimage * 18:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1112.eqiad.wmnet with reason: host reimage * 17:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1072.eqiad.wmnet with reason: host reimage * 17:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1112.eqiad.wmnet with OS trixie * 17:55 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1072.eqiad.wmnet with reason: host reimage * 17:39 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1072.eqiad.wmnet with OS trixie * 17:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1071.eqiad.wmnet with OS trixie * 17:18 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1070.eqiad.wmnet with OS trixie * 17:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1084.eqiad.wmnet with OS trixie * 16:54 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1071.eqiad.wmnet with reason: host reimage * 16:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1084.eqiad.wmnet with reason: host reimage * 16:51 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1070.eqiad.wmnet with reason: host reimage * 16:49 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1084.eqiad.wmnet with reason: host reimage * 16:38 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1071.eqiad.wmnet with OS trixie * 16:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1096.eqiad.wmnet with OS trixie * 16:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1070.eqiad.wmnet with OS trixie * 16:33 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1084.eqiad.wmnet with OS trixie * 16:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1089.eqiad.wmnet with OS trixie * 16:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1103.eqiad.wmnet with OS trixie * 16:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1096.eqiad.wmnet with reason: host reimage * 16:14 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1096.eqiad.wmnet with reason: host reimage * 16:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1089.eqiad.wmnet with reason: host reimage * 16:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1103.eqiad.wmnet with reason: host reimage * 16:02 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1003.eqiad.wmnet with OS bookworm * 16:01 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1089.eqiad.wmnet with reason: host reimage * 16:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1103.eqiad.wmnet with reason: host reimage * 15:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1096.eqiad.wmnet with OS trixie * 15:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1080.eqiad.wmnet with OS trixie * 15:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1089.eqiad.wmnet with OS trixie * 15:45 dancy@deploy1003: Installation of scap version "4.272.0" completed for 158 hosts * 15:43 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1103.eqiad.wmnet with OS trixie * 15:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1113.eqiad.wmnet with OS trixie * 15:41 dancy@deploy1003: Installing scap version "4.272.0" for 158 host(s) * 15:40 klausman@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 15:39 klausman@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 15:38 klausman@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 15:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1069.eqiad.wmnet with OS trixie * 15:37 klausman@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 15:36 klausman@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 15:34 klausman@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 15:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1080.eqiad.wmnet with reason: host reimage * 15:27 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1080.eqiad.wmnet with reason: host reimage * 15:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1113.eqiad.wmnet with reason: host reimage * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1069.eqiad.wmnet with reason: host reimage * 15:18 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1113.eqiad.wmnet with reason: host reimage * 15:16 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1069.eqiad.wmnet with reason: host reimage * 15:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:11 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1080.eqiad.wmnet with OS trixie * 15:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1113.eqiad.wmnet with OS trixie * 15:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1003.eqiad.wmnet with reason: host reimage * 14:47 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1003.eqiad.wmnet with OS bookworm * 14:33 elukey: rolled out spicerack on all cumin nodes - [[phab:T429699|T429699]] * 14:32 elukey: upgrade all bookworm hosts to pywmflib 3.1 - [[phab:T430552|T430552]] * 14:14 marostegui: Setup x4 eqiad topology [[phab:T404715|T404715]] * 14:13 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 14:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2230.codfw.wmnet * 14:07 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2230.codfw.wmnet * 13:59 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[2001-2002].codfw.wmnet * 13:51 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 13:45 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.major-upgrade (exit_code=97) * 13:45 cwilliams@cumin1003: dbmaint on s4@codfw [[phab:T429893|T429893]] * 13:45 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 13:42 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-master-codfw * 13:42 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl2002.codfw.wmnet * 13:42 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl2002.codfw.wmnet * 13:38 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl2002.codfw.wmnet * 13:38 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl2002.codfw.wmnet * 13:38 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl2001.codfw.wmnet * 13:38 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl2001.codfw.wmnet * 13:35 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl2001.codfw.wmnet * 13:35 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl2001.codfw.wmnet * 13:35 klausman@cumin2002: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-master-codfw * 12:30 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] (duration: 25m 11s) * 12:24 krinkle@deploy1003: krinkle: Continuing with deployment * 12:10 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2048.codfw.wmnet * 12:09 krinkle@deploy1003: krinkle: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:08 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2048.codfw.wmnet * 12:05 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] * 11:57 moritzm: installing curl security updates * 11:49 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:31 moritzm: installing nano security updates * 11:07 moritzm: failover Ganeti master in codfw to ganeti2032 [[phab:T430909|T430909]] * 11:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:04 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest1005.eqiad.wmnet with OS trixie * 11:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:50 jmm@dns1004: END - running authdns-update * 10:47 jmm@dns1004: START - running authdns-update * 10:47 jmm@dns1004: START - running authdns-update * 10:46 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:44 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest1005.eqiad.wmnet with reason: host reimage * 10:38 elukey@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest1005.eqiad.wmnet with reason: host reimage * 10:31 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:31 marostegui: Setup x4 codfw topology [[phab:T404715|T404715]] * 10:31 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 10:24 elukey: spicerack 13.0.0 deployed on cumin2002 * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 10:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 10:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 10:21 elukey@cumin2002: START - Cookbook sre.hosts.reimage for host sretest1005.eqiad.wmnet with OS trixie * 10:20 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:19 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:17 elukey: uploaded spicerack_13.0.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia * 09:54 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:52 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:20 elukey: upgrade all bullseye hosts to pywmflib 3.1 - [[phab:T430552|T430552]] * 09:10 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1015.eqiad.wmnet,service=s4 * 09:10 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1015.eqiad.wmnet,service=s6 * 09:07 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 08:58 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:56 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 08:56 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 08:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 08:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 08:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin2002.codfw.wmnet * 08:06 godog: remove cloudvirt1046, cloudvirt1062, cloudvirt1074, cloudvirt1075 from maintenance aggregate and put them in network-ovs - [[phab:T424802|T424802]] * 08:00 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin2002.codfw.wmnet * 07:58 hashar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] (duration: 32m 53s) * 07:57 fabfur: repooled cp4038 * 07:57 fabfur@cumin1003: conftool action : set/pooled=yes; selector: name=cp4038.* * 07:53 moritzm: installing pyjwt security updates * 07:47 moritzm: installing openjpeg2 security updates * 07:45 hashar@deploy1003: vadymts1, hashar: Continuing with deployment * 07:43 hashar@deploy1003: vadymts1, hashar: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:38 moritzm: installing python-urllib3 security updates * 07:37 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 07:30 fabfur: depooled cp4038 to investigate on possible maxmind failure * 07:30 fabfur@cumin1003: conftool action : set/pooled=no; selector: name=cp4038.* * 07:30 fabfur@cumin1003: conftool action : set/pooled=yes; selector: name=cp4038.* * 07:29 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 07:25 hashar@deploy1003: Started scap sync-world: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] * 06:13 moritzm: installing Linux 6.12.95 on trixie hosts * 05:20 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s6 * 05:20 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s4 * 05:19 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1015.eqiad.wmnet with reason: cloning * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 08s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-05 == * 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 01m 08s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-04 == * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 58s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-03 == * 17:08 topranks: revert protocol preference changes on cr3-ulsfo after upgrade * 16:53 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on cr2-eqord with reason: upgrade JunOS cr3-ulsfo * 16:53 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on cr4-ulsfo with reason: upgrade JunOS cr3-ulsfo * 16:48 topranks: reboot cr3-ulsfo to upgrade JunOS and reset linecard [[phab:T424839|T424839]] * 15:52 topranks: adjust outbound BGP policies on cr3-ulsfo to drain router of traffic [[phab:T424839|T424839]] * 15:45 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on lvs[4008-4010].ulsfo.wmnet with reason: upgrade JunOS cr3-ulsfo * 15:44 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on asw1-[22-23]-ulsfo,cr3-ulsfo,cr3-ulsfo IPv6,cr3-ulsfo.mgmt with reason: upgrade JunOS cr3-ulsfo * 15:36 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 15:35 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 15:35 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 14:40 cmooney@dns3003: END - running authdns-update * 14:26 cmooney@dns3003: START - running authdns-update * 14:26 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:26 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to ulsfo - cmooney@cumin1003" * 14:19 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to ulsfo - cmooney@cumin1003" * 14:16 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:38 sukhe@dns1004: END - running authdns-update * 13:35 sukhe@dns1004: START - running authdns-update * 13:26 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 13:26 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 13:26 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet * 13:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 13:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 13:16 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host sretest1005.eqiad.wmnet * 13:16 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 13:16 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 13:15 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 13:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:14 moritzm: imported samplicator 1.3.8rc1-1+deb13u1 to trixie-wikimedia/main [[phab:T337208|T337208]] * 13:13 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:07 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 13:07 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 13:02 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:02 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:58 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:57 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:57 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:53 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet * 12:50 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 12:47 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:41 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:40 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:39 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:32 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet * 12:26 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet * 12:23 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2005.wikimedia.org * 12:19 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2005.wikimedia.org * 12:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 12:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup[2004-2007].codfw.wmnet * 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[2004-2007].codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin2003" * 12:15 jynus@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[2004-2007].codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin2003" * 12:09 jynus@cumin2003: START - Cookbook sre.dns.netbox * 11:58 jynus@cumin2003: START - Cookbook sre.hosts.decommission for hosts backup[2004-2007].codfw.wmnet * 10:40 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup[1004-1007].eqiad.wmnet * 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[1004-1007].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 10:01 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[1004-1007].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 09:52 jynus@cumin1003: START - Cookbook sre.dns.netbox * 09:39 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:36 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup[1004-1007].eqiad.wmnet * 09:36 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:25 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 09:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 09:16 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 09:05 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:04 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:00 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:59 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:57 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:55 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:50 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 08:50 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 08:49 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 08:49 atsukoito: depooling cirrussearch in codfw because of regression after upgrade [[phab:T431091|T431091]] * 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts mirror1001.wikimedia.org * 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: mirror1001.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 08:29 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: mirror1001.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 08:18 jmm@cumin2003: START - Cookbook sre.dns.netbox * 08:11 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts mirror1001.wikimedia.org * 06:15 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 18s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-02 == * 22:55 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host contint1003.wikimedia.org with OS trixie * 22:29 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on contint1003.wikimedia.org with reason: host reimage * 22:23 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on contint1003.wikimedia.org with reason: host reimage * 22:05 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host contint1003.wikimedia.org with OS trixie * 22:03 mutante: contint1003 (zuul.wikimedia.org) - reimaging because of [[phab:T430510|T430510]]#12067628 [[phab:T418521|T418521]] * 22:03 dzahn@cumin2002: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on zuul.wikimedia.org with reason: reimage * 21:39 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 18s) * 21:39 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 21:20 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1003.eqiad.wmnet, repooling source-only afterwards * 21:19 sbassett: Deployed security fix for [[phab:T428829|T428829]] * 20:58 cmooney@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Release v0.11.2 update for new Aerleon - cmooney@cumin1003 * 20:55 cmooney@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Release v0.11.2 update for new Aerleon - cmooney@cumin1003 * 20:40 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] (duration: 12m 35s) * 20:36 arlolra@deploy1003: cscott, arlolra: Continuing with deployment * 20:35 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 20s) * 20:35 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 20:33 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host contint2003.wikimedia.org with OS trixie * 20:31 arlolra@deploy1003: cscott, arlolra: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Cha * 20:28 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] * 20:17 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] (duration: 08m 13s) * 20:14 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on contint2003.wikimedia.org with reason: host reimage * 20:13 sbassett@deploy1003: sbassett: Continuing with deployment * 20:11 sbassett@deploy1003: sbassett: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:09 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] * 20:08 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 20:08 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 20:08 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on contint2003.wikimedia.org with reason: host reimage * 20:05 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1003.eqiad.wmnet, repooling source-only afterwards * 19:49 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host contint2003.wikimedia.org with OS trixie * 19:48 mutante: contint2003 - reimaging because of [[phab:T430510|T430510]]#12067628 [[phab:T418521|T418521]] * 18:39 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 18:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 18:13 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2002.codfw.wmnet -> wcqs2003.codfw.wmnet, repooling source-only afterwards * 17:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1003.eqiad.wmnet with OS bookworm * 17:52 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1005.eqiad.wmnet * 17:52 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1005.eqiad.wmnet * 17:51 jasmine@cumin2002: conftool action : set/pooled=yes:weight=10; selector: name=wikikube-ctrl1005.eqiad.wmnet * 17:48 jasmine_: homer "cr*eqiad*" commit "Added new stacked control plane wikikube-ctrl1005" * 17:44 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply * 17:44 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply * 17:31 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] (duration: 09m 33s) * 17:26 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 17:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1003.eqiad.wmnet with reason: host reimage * 17:23 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:21 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] * 17:18 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1003.eqiad.wmnet with reason: host reimage * 17:16 rscout@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply * 17:16 rscout@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply * 17:16 rscout@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply * 17:15 rscout@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply * 17:12 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on wcqs[2002-2003].codfw.wmnet,wcqs1002.eqiad.wmnet with reason: reimaging hosts * 17:08 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 17:08 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 17:08 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 17:07 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 17:05 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 17:05 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 17:03 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "running to make sure all updates are synced - cmooney@cumin1003" * 17:03 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "running to make sure all updates are synced - cmooney@cumin1003" * 17:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs1003 * 17:00 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs1003 * 17:00 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 17:00 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1003.eqiad.wmnet with OS bookworm * 16:58 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Re-running - btullis@cumin1003" * 16:58 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Re-running - btullis@cumin1003" * 16:58 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2002.codfw.wmnet -> wcqs2003.codfw.wmnet, repooling source-only afterwards * 16:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-master1004.eqiad.wmnet with OS bookworm * 16:58 btullis@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 16:57 tappof: bump space for prometheus k8s-aux in eqiad * 16:55 cmooney@dns3003: END - running authdns-update * 16:55 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:55 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to eqsin - cmooney@cumin1003" * 16:55 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to eqsin - cmooney@cumin1003" * 16:53 cmooney@dns3003: START - running authdns-update * 16:52 ryankemper: [ml-serve-eqiad] Cleared out 1302 failed (Evicted) pods: `kubectl -n llm delete pods --field-selector=status.phase=Failed`, freeing calico-kube-controllers from OOM crashloop (evictions were caused by disk pressure) * 16:49 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 16:46 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:39 rzl@dns1004: END - running authdns-update * 16:37 rzl@dns1004: START - running authdns-update * 16:36 rzl@dns1004: START - running authdns-update * 16:35 rzl@deploy1003: Finished scap sync-world: [[phab:T416623|T416623]] (duration: 10m 19s) * 16:34 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 16:33 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-master1004.eqiad.wmnet with reason: host reimage * 16:30 rzl@deploy1003: rzl: Continuing with deployment * 16:28 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-master1004.eqiad.wmnet with reason: host reimage * 16:26 rzl@deploy1003: rzl: [[phab:T416623|T416623]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:25 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 16:25 rzl@deploy1003: Started scap sync-world: [[phab:T416623|T416623]] * 16:25 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 16:24 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 16:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: sync * 16:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: sync * 16:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-master1004.eqiad.wmnet with OS bookworm * 16:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-master1003.eqiad.wmnet with OS bookworm * 16:11 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 16:11 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 16:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Security updates * 16:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 16:08 root@cumin1003: START - Cookbook sre.mysql.parsercache * 16:08 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Security updates * 15:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-master1003.eqiad.wmnet with reason: host reimage * 15:54 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:54 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:54 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:54 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-master1003.eqiad.wmnet with reason: host reimage * 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Security updates * 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:45 root@cumin1003: START - Cookbook sre.mysql.parsercache * 15:45 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Security updates * 15:42 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-master1003.eqiad.wmnet with OS bookworm * 15:24 moritzm: installing busybox updates from bookworm point release * 15:20 moritzm: installing busybox updates from trixie point release * 15:15 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1021: Security updates * 15:15 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:15 root@cumin1003: START - Cookbook sre.mysql.parsercache * 15:15 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1021: Security updates * 15:13 moritzm: installing giflib security updates * 15:08 moritzm: installing Tomcat security updates * 14:57 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 14:56 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 14:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:53 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Unblock taavi - oblivian@cumin1003" * 14:53 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Unblock taavi - oblivian@cumin1003 * 14:53 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1021: Security updates * 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:53 root@cumin1003: START - Cookbook sre.mysql.parsercache * 14:53 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1021: Security updates * 14:53 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Unblock taavi - oblivian@cumin1003 * 14:52 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Unblock taavi - oblivian@cumin1003" * 14:46 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94711 and previous config saved to /var/cache/conftool/dbconfig/20260702-144644-fceratto.json * 14:36 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205', diff saved to https://phabricator.wikimedia.org/P94709 and previous config saved to /var/cache/conftool/dbconfig/20260702-143636-fceratto.json * 14:32 moritzm: installing libdbi-perl security updates * 14:26 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205', diff saved to https://phabricator.wikimedia.org/P94708 and previous config saved to /var/cache/conftool/dbconfig/20260702-142628-fceratto.json * 14:16 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94707 and previous config saved to /var/cache/conftool/dbconfig/20260702-141621-fceratto.json * 14:12 moritzm: installing rsync security updates * 14:11 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox) * 14:10 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94706 and previous config saved to /var/cache/conftool/dbconfig/20260702-140959-fceratto.json * 14:09 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2205.codfw.wmnet with reason: Maintenance * 14:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2205: Repooling after switchover * 14:07 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-test-master1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 14:06 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 14:06 Tran: Deployed patch for [[phab:T427287|T427287]] * 14:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:59 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2205: Repooling after switchover * 13:59 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2205: Repooling after switchover * 13:59 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:55 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2205: Repooling after switchover * 13:55 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2205 [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94704 and previous config saved to /var/cache/conftool/dbconfig/20260702-135505-fceratto.json * 13:54 moritzm: installing sed security updates * 13:53 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:52 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2209 to s3 primary [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94703 and previous config saved to /var/cache/conftool/dbconfig/20260702-135235-fceratto.json * 13:52 federico3: Starting s3 codfw failover from db2205 to db2209 - [[phab:T430912|T430912]] * 13:51 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:51 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 13:48 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:47 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2209 with weight 0 [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94702 and previous config saved to /var/cache/conftool/dbconfig/20260702-134719-fceratto.json * 13:47 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Primary switchover s3 [[phab:T430912|T430912]] * 13:44 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:44 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:44 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:40 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 13:38 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 13:37 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 13:36 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 13:36 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:34 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 13:30 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:29 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:29 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:27 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:26 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:25 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling restart_daemons on A:wikidough * 13:23 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 13:22 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 13:17 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 13:17 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns1004.wikimedia.org * 13:12 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:11 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart (exit_code=97) rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough * 13:11 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=97) rolling restart_daemons on A:wikidough * 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough * 13:09 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] (duration: 07m 20s) * 13:05 aude@deploy1003: jdrewniak, aude: Continuing with deployment * 13:04 aude@deploy1003: jdrewniak, aude: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:02 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] * 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts wdqs-categories1001.eqiad.wmnet * 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: wdqs-categories1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 12:10 jmm@dns1004: END - running authdns-update * 12:07 jmm@dns1004: START - running authdns-update * 11:51 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: wdqs-categories1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 11:44 btullis@cumin1003: START - Cookbook sre.dns.netbox * 11:42 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 11:42 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 11:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet * 11:39 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts wdqs-categories1001.eqiad.wmnet * 11:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet * 11:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet * 11:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet * 11:29 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 11:29 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 10:57 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2214: Repooling * 10:49 jmm@dns1004: END - running authdns-update * 10:47 jmm@dns1004: START - running authdns-update * 10:31 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94698 and previous config saved to /var/cache/conftool/dbconfig/20260702-103146-fceratto.json * 10:21 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213', diff saved to https://phabricator.wikimedia.org/P94696 and previous config saved to /var/cache/conftool/dbconfig/20260702-102137-fceratto.json * 10:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:19 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb1017.eqiad.wmnet * 10:18 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 10:18 fceratto@cumin1003: Removing es1033 from zarcillo [[phab:T408772|T408772]] * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts es1033.eqiad.wmnet * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: es1033.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:14 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: es1033.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:13 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb1017.eqiad.wmnet * 10:12 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2214.codfw.wmnet * 10:12 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2214.codfw.wmnet * 10:12 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2214: Repooling * 10:11 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213', diff saved to https://phabricator.wikimedia.org/P94693 and previous config saved to /var/cache/conftool/dbconfig/20260702-101130-fceratto.json * 10:10 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:10 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:03 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts es1033.eqiad.wmnet * 10:03 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 10:01 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94691 and previous config saved to /var/cache/conftool/dbconfig/20260702-100122-fceratto.json * 09:55 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94690 and previous config saved to /var/cache/conftool/dbconfig/20260702-095529-fceratto.json * 09:55 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2213.codfw.wmnet with reason: Maintenance * 09:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 09:53 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2213: Repooling after switchover * 09:51 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover * 09:44 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2213: Repooling after switchover * 09:39 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover * 09:39 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2213 [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94688 and previous config saved to /var/cache/conftool/dbconfig/20260702-093859-fceratto.json * 09:36 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2192 to s5 primary [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94687 and previous config saved to /var/cache/conftool/dbconfig/20260702-093650-fceratto.json * 09:36 federico3: Starting s5 codfw failover from db2213 to db2192 - [[phab:T430923|T430923]] * 09:30 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94686 and previous config saved to /var/cache/conftool/dbconfig/20260702-093004-fceratto.json * 09:24 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2192 with weight 0 [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94685 and previous config saved to /var/cache/conftool/dbconfig/20260702-092455-fceratto.json * 09:24 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 23 hosts with reason: Primary switchover s5 [[phab:T430923|T430923]] * 09:19 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220', diff saved to https://phabricator.wikimedia.org/P94684 and previous config saved to /var/cache/conftool/dbconfig/20260702-091957-fceratto.json * 09:16 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] (duration: 06m 57s) * 09:13 moritzm: installing libgcrypt20 security updates * 09:12 kharlan@deploy1003: kharlan: Continuing with deployment * 09:11 kharlan@deploy1003: kharlan: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:09 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220', diff saved to https://phabricator.wikimedia.org/P94683 and previous config saved to /var/cache/conftool/dbconfig/20260702-090950-fceratto.json * 09:09 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] * 09:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 09:01 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] (duration: 07m 07s) * 08:59 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94682 and previous config saved to /var/cache/conftool/dbconfig/20260702-085942-fceratto.json * 08:57 kharlan@deploy1003: kharlan: Continuing with deployment * 08:56 kharlan@deploy1003: kharlan: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:54 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] * 08:52 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:52 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:52 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94681 and previous config saved to /var/cache/conftool/dbconfig/20260702-085237-fceratto.json * 08:52 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2220.codfw.wmnet with reason: Maintenance * 08:43 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:40 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 08:25 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] (duration: 11m 44s) * 08:21 cscott@deploy1003: cscott: Continuing with deployment * 08:16 cscott@deploy1003: cscott: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:14 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] * 08:08 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 08:08 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1244: Migration of db1244.eqiad.wmnet completed * 08:02 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:02 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:01 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] (duration: 18m 58s) * 08:01 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:59 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 07:59 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:59 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:59 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2006.wikimedia.org * 07:58 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:57 cscott@deploy1003: cscott: Continuing with deployment * 07:56 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:56 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:56 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:55 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:55 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:55 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:54 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2006.wikimedia.org * 07:54 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:54 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 07:44 cscott@deploy1003: cscott: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:44 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2005.wikimedia.org * 07:44 moritzm: installing node-lodash security updates * 07:42 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] * 07:39 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2005.wikimedia.org * 07:30 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] (duration: 07m 28s) * 07:26 cscott@deploy1003: ssastry, cscott: Continuing with deployment * 07:25 cscott@deploy1003: ssastry, cscott: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:23 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1244: Migration of db1244.eqiad.wmnet completed * 07:22 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] * 07:16 wmde-fisch@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] (duration: 06m 55s) * 07:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1244.eqiad.wmnet with OS trixie * 07:11 wmde-fisch@deploy1003: wmde-fisch: Continuing with deployment * 07:11 wmde-fisch@deploy1003: wmde-fisch: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:09 wmde-fisch@deploy1003: Started scap sync-world: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] * 06:54 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1244.eqiad.wmnet with reason: host reimage * 06:50 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1244.eqiad.wmnet with reason: host reimage * 06:38 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1250.eqiad.wmnet with OS trixie * 06:34 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db1244.eqiad.wmnet with OS trixie * 06:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1244: Upgrading db1244.eqiad.wmnet * 06:25 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1244: Upgrading db1244.eqiad.wmnet * 06:25 cwilliams@cumin1003: dbmaint on s4@eqiad [[phab:T429893|T429893]] * 06:25 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 06:15 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1250.eqiad.wmnet with reason: host reimage * 06:14 cwilliams@dns1006: END - running authdns-update * 06:12 cwilliams@dns1006: START - running authdns-update * 06:11 cwilliams@dns1006: END - running authdns-update * 06:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db1244 [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94676 and previous config saved to /var/cache/conftool/dbconfig/20260702-061059-cwilliams.json * 06:09 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1250.eqiad.wmnet with reason: host reimage * 06:09 cwilliams@dns1006: START - running authdns-update * 06:08 aokoth@cumin1003: END (PASS) - Cookbook sre.vrts.upgrade (exit_code=0) on VRTS host vrts1003.eqiad.wmnet * 06:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db1160 to s4 primary and set section read-write [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94675 and previous config saved to /var/cache/conftool/dbconfig/20260702-060746-cwilliams.json * 06:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Set s4 eqiad as read-only for maintenance - [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94674 and previous config saved to /var/cache/conftool/dbconfig/20260702-060704-cwilliams.json * 06:06 cezmunsta: Starting s4 eqiad failover from db1244 to db1160 - [[phab:T430817|T430817]] * 06:04 aokoth@cumin1003: START - Cookbook sre.vrts.upgrade on VRTS host vrts1003.eqiad.wmnet * 05:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db1160 with weight 0 [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94673 and previous config saved to /var/cache/conftool/dbconfig/20260702-055927-cwilliams.json * 05:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 40 hosts with reason: Primary switchover s4 [[phab:T430817|T430817]] * 05:55 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1250.eqiad.wmnet with OS trixie * 05:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on db1250.eqiad.wmnet with reason: m3 master switchover [[phab:T430158|T430158]] * 05:39 marostegui: Failover m3 (phabricator) from db1250 to db1228 - [[phab:T430158|T430158]] * 05:32 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2234].codfw.wmnet,db[1217,1228,1250].eqiad.wmnet with reason: m3 master switchover [[phab:T430158|T430158]] * 04:45 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] (duration: 09m 08s) * 04:41 tstarling@deploy1003: tstarling, reedy: Continuing with deployment * 04:38 tstarling@deploy1003: tstarling, reedy: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 04:36 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 59s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:16 ryankemper: [[phab:T429844|T429844]] [opensearch] completed `cirrussearch2111` reimage; all codfw search clusters are green, all nodes now report `OpenSearch 2.19.5`, and the temporary chi voting exclusion has been removed * 00:57 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2111.codfw.wmnet with OS trixie * 00:29 ryankemper: [[phab:T429844|T429844]] [opensearch] depooled codfw search-omega/search-psi discovery records to match existing codfw search depool during OpenSearch 2.19 migration * 00:29 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2111.codfw.wmnet with reason: host reimage * 00:29 ryankemper@cumin2002: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 00:29 ryankemper@cumin2002: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 00:22 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2111.codfw.wmnet with reason: host reimage * 00:01 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2111.codfw.wmnet with OS trixie * 00:00 ryankemper: [[phab:T429844|T429844]] [opensearch] chi cluster recovered after stopping `opensearch_1@production-search-codfw` on `cirrussearch2111` == 2026-07-01 == * 23:59 ryankemper: [[phab:T429844|T429844]] [opensearch] stopped `opensearch_1@production-search-codfw` on `cirrussearch2111` after chi cluster-manager election churn following `voting_config_exclusions` POST; hoping this triggers a re-election * 23:52 cscott@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 23:51 cscott@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 23:51 cscott@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 23:50 cscott@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2003.codfw.wmnet with OS bookworm * 22:29 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 22:13 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 22:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2084.codfw.wmnet with OS trixie * 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2003.codfw.wmnet with reason: host reimage * 22:03 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 22:01 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2003.codfw.wmnet with reason: host reimage * 21:50 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 21:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2084.codfw.wmnet with reason: host reimage * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2003 * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2003 * 21:42 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2003 * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2003.codfw.wmnet 45.48.192.10.in-addr.arpa 5.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:42 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2003.codfw.wmnet 45.48.192.10.in-addr.arpa 5.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2003 - bking@cumin2003" * 21:42 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2003 - bking@cumin2003" * 21:36 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2084.codfw.wmnet with reason: host reimage * 21:35 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:34 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2003 * 21:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2003.codfw.wmnet with OS bookworm * 21:19 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2084.codfw.wmnet with OS trixie * 21:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2081.codfw.wmnet with OS trixie * 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2108.codfw.wmnet with OS trixie * 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2081.codfw.wmnet with reason: host reimage * 20:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2081.codfw.wmnet with reason: host reimage * 20:28 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2081.codfw.wmnet with OS trixie * 20:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2108.codfw.wmnet with reason: host reimage * 20:19 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2108.codfw.wmnet with reason: host reimage * 19:59 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2108.codfw.wmnet with OS trixie * 19:46 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2093.codfw.wmnet with OS trixie * 19:44 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 19:44 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jasmine@cumin2002" * 19:43 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jasmine@cumin2002" * 19:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2080.codfw.wmnet with OS trixie * 19:28 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 19:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2093.codfw.wmnet with reason: host reimage * 19:18 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 19:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2093.codfw.wmnet with reason: host reimage * 19:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2080.codfw.wmnet with reason: host reimage * 19:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2080.codfw.wmnet with reason: host reimage * 18:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2093.codfw.wmnet with OS trixie * 18:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2080.codfw.wmnet with OS trixie * 18:27 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 18:18 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] (duration: 09m 15s) * 18:13 jgiannelos@deploy1003: jgiannelos, neriah: Continuing with deployment * 18:11 jgiannelos@deploy1003: jgiannelos, neriah: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:09 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] * 17:40 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 16:58 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 30 hosts * 16:57 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for 30 hosts * 16:52 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2202.codfw.wmnet * 16:52 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2202.codfw.wmnet * 16:51 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt * 16:51 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt * 16:51 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lvs2012.codfw.wmnet * 16:51 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for lvs2012.codfw.wmnet * 16:49 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2076.codfw.wmnet with OS trixie * 16:49 brett: Start pybal on lvs2012 - [[phab:T429861|T429861]] * 16:49 pt1979@cumin1003: END (ERROR) - Cookbook sre.hosts.remove-downtime (exit_code=97) for 59 hosts * 16:48 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for 59 hosts * 16:42 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2061.codfw.wmnet with OS trixie * 16:30 dancy@deploy1003: Installation of scap version "4.271.0" completed for 2 hosts * 16:28 dancy@deploy1003: Installing scap version "4.271.0" for 2 host(s) * 16:23 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2076.codfw.wmnet with reason: host reimage * 16:19 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2061.codfw.wmnet with reason: host reimage * 16:18 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2076.codfw.wmnet with reason: host reimage * 16:18 jasmine@dns1004: END - running authdns-update * 16:16 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host restbase2039.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 16:16 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host restbase2039.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 16:16 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2061.codfw.wmnet with reason: host reimage * 16:15 jasmine@dns1004: START - running authdns-update * 16:14 jasmine@dns1004: END - running authdns-update * 16:12 jasmine@dns1004: START - running authdns-update * 16:07 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2202.codfw.wmnet with reason: maintenance * 16:06 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt with reason: Junos upograde * 16:00 papaul: ongoing maintenance on lsw1-b2-codfw * 16:00 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2076.codfw.wmnet with OS trixie * 15:59 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt * 15:59 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt * 15:57 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2061.codfw.wmnet with OS trixie * 15:55 pt1979@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2042,2046].codfw.wmnet * 15:55 pt1979@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2042,2046].codfw.wmnet * 15:51 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 15:51 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2220: Repooling after switchover * 15:50 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 15:50 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 15:48 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2092.codfw.wmnet with OS trixie * 15:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 15:40 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 15:38 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 15:37 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 15:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply * 15:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply * 15:32 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 15:32 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 15:30 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 15:29 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 15:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 15:25 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:22 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs2012.codfw.wmnet with reason: Rack B2 maintenance - [[phab:T429861|T429861]] * 15:21 brett: Stopping pybal on lvs2012 in preparation for codfw rack b2 maintenance - [[phab:T429861|T429861]] * 15:20 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2092.codfw.wmnet with reason: host reimage * 15:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:12 _joe_: restarted manually alertmanager-irc-relay * 15:12 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2092.codfw.wmnet with reason: host reimage * 15:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:12 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt with reason: Junos upograde * 15:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover * 15:07 pt1979@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2042,2046].codfw.wmnet * 15:06 pt1979@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2042,2046].codfw.wmnet * 15:02 papaul: ongoing maintenance on lsw1-a8-codfw * 14:31 topranks: POWERING DOWN CR1-EQIAD for line card installation [[phab:T426343|T426343]] * 14:31 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] (duration: 08m 57s) * 14:29 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:26 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 14:24 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:22 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] * 14:22 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover * 14:16 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:15 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover * 14:14 topranks: re-enable routing-engine graceful-failover on cr1-eqiad [[phab:T417873|T417873]] * 14:13 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:13 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2220: Repooling after switchover * 14:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:12 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:12 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:11 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:08 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] (duration: 10m 01s) * 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:07 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2220 [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94664 and previous config saved to /var/cache/conftool/dbconfig/20260701-140729-fceratto.json * 14:06 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:06 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:06 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:05 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2159 to s7 primary [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94663 and previous config saved to /var/cache/conftool/dbconfig/20260701-140503-fceratto.json * 14:04 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:04 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 14:04 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 14:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:04 dreamyjazz@deploy1003: anzx, dreamyjazz: Continuing with deployment * 14:04 federico3: Starting s7 codfw failover from db2220 to db2159 - [[phab:T430826|T430826]] * 14:03 jmm@dns1004: END - running authdns-update * 14:03 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:03 topranks: flipping cr1-eqiad active routing-enginer back to RE0 [[phab:T417873|T417873]] * 14:03 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudsw1-c8-eqiad,cloudsw1-d5-eqiad with reason: router upgrades eqiad * 14:01 jmm@dns1004: START - running authdns-update * 14:00 dreamyjazz@deploy1003: anzx, dreamyjazz: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:59 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2159 with weight 0 [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94662 and previous config saved to /var/cache/conftool/dbconfig/20260701-135906-fceratto.json * 13:58 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] * 13:57 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s7 [[phab:T430826|T430826]] * 13:56 topranks: reboot routing-enginer RE0 on cr1-eqiad [[phab:T417873|T417873]] * 13:48 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1006.wikimedia.org * 13:44 atsuko@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cirrussearch2092.codfw.wmnet with OS trixie * 13:43 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1006.wikimedia.org * 13:41 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2092.codfw.wmnet with OS trixie * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1005.wikimedia.org * 13:37 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1005.wikimedia.org * 13:37 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on pfw1-eqiad with reason: router upgrades eqiad * 13:35 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on lvs[1017-1020].eqiad.wmnet with reason: router upgrades eqiad * 13:34 caro@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] (duration: 07m 59s) * 13:30 caro@deploy1003: caro: Continuing with deployment * 13:28 caro@deploy1003: caro: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:27 topranks: route-engine failover cr1-eqiad * 13:26 caro@deploy1003: Started scap sync-world: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] * 13:15 topranks: rebooting routing-engine 1 on cr1-eqiad [[phab:T417873|T417873]] * 13:13 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] (duration: 08m 29s) * 13:13 moritzm: installing qemu security updates * 13:11 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 13:11 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 13:09 jgiannelos@deploy1003: jgiannelos: Continuing with deployment * 13:08 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 13:07 jgiannelos@deploy1003: jgiannelos: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:06 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 13:06 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2214.codfw.wmnet with reason: Maintenance * 13:05 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2214: Repooling after switchover * 13:05 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] * 13:04 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2214: Repooling after switchover * 13:04 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2214 [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94660 and previous config saved to /var/cache/conftool/dbconfig/20260701-130413-fceratto.json * 13:01 moritzm: installing python3.13 security updates * 13:00 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2229 to s6 primary [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94659 and previous config saved to /var/cache/conftool/dbconfig/20260701-125959-fceratto.json * 12:59 federico3: Starting s6 codfw failover from db2214 to db2229 - [[phab:T430814|T430814]] * 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on 13 hosts with reason: router upgrade and line card install * 12:51 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2229 with weight 0 [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94658 and previous config saved to /var/cache/conftool/dbconfig/20260701-125149-fceratto.json * 12:51 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 21 hosts with reason: Primary switchover s6 [[phab:T430814|T430814]] * 12:50 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2189.codfw.wmnet * 12:50 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2189.codfw.wmnet * 12:42 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2100.codfw.wmnet with OS trixie * 12:38 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2083.codfw.wmnet with OS trixie * 12:19 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2083.codfw.wmnet with reason: host reimage * 12:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 12:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2240: Migration of db2240.codfw.wmnet completed * 12:14 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2100.codfw.wmnet with reason: host reimage * 12:09 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2083.codfw.wmnet with reason: host reimage * 12:09 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2100.codfw.wmnet with reason: host reimage * 12:00 topranks: drain traffic on cr1-eqiad to allow for line card install and JunOS upgrade [[phab:T426343|T426343]] * 11:52 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2083.codfw.wmnet with OS trixie * 11:50 cmooney@dns2005: END - running authdns-update * 11:49 cmooney@dns2005: START - running authdns-update * 11:48 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2100.codfw.wmnet with OS trixie * 11:40 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/zotero: apply * 11:40 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/zotero: apply * 11:36 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/zotero: apply * 11:36 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/zotero: apply * 11:31 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2240: Migration of db2240.codfw.wmnet completed * 11:30 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply * 11:28 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply * 11:27 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:27 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:27 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:27 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:27 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:23 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2240.codfw.wmnet with OS trixie * 11:20 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:20 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:17 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:16 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:16 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:15 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2086.codfw.wmnet with OS trixie * 11:14 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2106.codfw.wmnet with OS trixie * 11:14 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:13 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:12 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:09 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2115.codfw.wmnet with OS trixie * 11:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2240.codfw.wmnet with reason: host reimage * 11:00 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2240.codfw.wmnet with reason: host reimage * 10:53 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2106.codfw.wmnet with reason: host reimage * 10:49 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2086.codfw.wmnet with reason: host reimage * 10:44 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2115.codfw.wmnet with reason: host reimage * 10:44 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2240.codfw.wmnet with OS trixie * 10:44 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2086.codfw.wmnet with reason: host reimage * 10:42 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2106.codfw.wmnet with reason: host reimage * 10:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2240: Upgrading db2240.codfw.wmnet * 10:41 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2240: Upgrading db2240.codfw.wmnet * 10:41 cwilliams@cumin1003: dbmaint on s4@codfw [[phab:T429893|T429893]] * 10:40 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 10:39 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2115.codfw.wmnet with reason: host reimage * 10:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2240 [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94653 and previous config saved to /var/cache/conftool/dbconfig/20260701-102658-cwilliams.json * 10:26 moritzm: installing nginx security updates * 10:26 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2086.codfw.wmnet with OS trixie * 10:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2179 to s4 primary [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94652 and previous config saved to /var/cache/conftool/dbconfig/20260701-102356-cwilliams.json * 10:23 cezmunsta: Starting s4 codfw failover from db2240 to db2179 - [[phab:T430127|T430127]] * 10:23 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2106.codfw.wmnet with OS trixie * 10:20 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2115.codfw.wmnet with OS trixie * 10:15 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2179 with weight 0 [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94651 and previous config saved to /var/cache/conftool/dbconfig/20260701-101531-cwilliams.json * 10:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 40 hosts with reason: Primary switchover s4 [[phab:T430127|T430127]] * 09:56 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template (take 2) - oblivian@cumin1003" * 09:56 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template (take 2) - oblivian@cumin1003 * 09:55 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template (take 2) - oblivian@cumin1003 * 09:55 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template (take 2) - oblivian@cumin1003" * 09:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:39 mszwarc@deploy1003: Synchronized private/SuggestedInvestigationsSignals/SuggestedInvestigationsSignal4n.php: Update SI signal 4n (duration: 06m 08s) * 09:21 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 09:21 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 09:14 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 09:14 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 09:02 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 09:02 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 08:54 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 08:38 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 08:38 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 08:36 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 08:21 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 08:21 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] (duration: 36m 11s) * 08:15 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 08:09 mszwarc@deploy1003: mszwarc, abi: Continuing with deployment * 08:03 mszwarc@deploy1003: mszwarc, abi: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:55 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 07:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 07:45 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] * 07:30 aqu@deploy1003: Finished deploy [analytics/refinery@410f205]: Regular analytics weekly train 2nd try [analytics/refinery@410f2050] (duration: 00m 22s) * 07:29 aqu@deploy1003: Started deploy [analytics/refinery@410f205]: Regular analytics weekly train 2nd try [analytics/refinery@410f2050] * 07:28 aqu@deploy1003: Finished deploy [analytics/refinery@410f205] (thin): Regular analytics weekly train THIN [analytics/refinery@410f2050] (duration: 01m 59s) * 07:28 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] (duration: 07m 19s) * 07:26 aqu@deploy1003: Started deploy [analytics/refinery@410f205] (thin): Regular analytics weekly train THIN [analytics/refinery@410f2050] * 07:26 aqu@deploy1003: Finished deploy [analytics/refinery@410f205]: Regular analytics weekly train [analytics/refinery@410f2050] (duration: 04m 32s) * 07:24 mszwarc@deploy1003: wmde-fisch, mszwarc: Continuing with deployment * 07:23 mszwarc@deploy1003: wmde-fisch, mszwarc: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:21 aqu@deploy1003: Started deploy [analytics/refinery@410f205]: Regular analytics weekly train [analytics/refinery@410f2050] * 07:21 aqu@deploy1003: Finished deploy [analytics/refinery@410f205] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@410f2050] (duration: 02m 01s) * 07:20 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] * 07:19 aqu@deploy1003: Started deploy [analytics/refinery@410f205] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@410f2050] * 07:13 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] (duration: 09m 13s) * 07:09 mszwarc@deploy1003: mszwarc, chlod, revi: Continuing with deployment * 07:06 mszwarc@deploy1003: mszwarc, chlod, revi: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:04 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] * 06:55 elukey: upgrade all trixie hosts to pywmflib 3.0 - [[phab:T430552|T430552]] * 06:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:43 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:43 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:42 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:42 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:41 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:41 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:35 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:35 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:34 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:34 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:31 jmm@cumin2003: DONE (PASS) - Cookbook sre.idm.logout (exit_code=0) Logging Niharika29 out of all services on: 2453 hosts * 06:30 oblivian@cumin1003: END (FAIL) - Cookbook sre.deploy.hiddenparma (exit_code=99) Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:30 oblivian@cumin1003: END (FAIL) - Cookbook sre.deploy.python-code (exit_code=99) hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:30 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:30 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:01 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2109.codfw.wmnet with OS trixie * 05:45 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on es1039.eqiad.wmnet with reason: issues * 05:41 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1027.eqiad.wmnet * 05:40 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2068.codfw.wmnet with OS trixie * 05:40 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2109.codfw.wmnet with reason: host reimage * 05:40 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1027.eqiad.wmnet,service=s2 * 05:40 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1027.eqiad.wmnet,service=s7 * 05:36 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2109.codfw.wmnet with reason: host reimage * 05:20 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2068.codfw.wmnet with reason: host reimage * 05:16 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2109.codfw.wmnet with OS trixie * 05:15 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2068.codfw.wmnet with reason: host reimage * 05:09 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2067.codfw.wmnet with OS trixie * 04:56 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2068.codfw.wmnet with OS trixie * 04:49 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2067.codfw.wmnet with reason: host reimage * 04:45 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2067.codfw.wmnet with reason: host reimage * 04:27 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2067.codfw.wmnet with OS trixie * 03:47 slyngshede@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1039.eqiad.wmnet with reason: Hardware crash * 03:21 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2107.codfw.wmnet with OS trixie * 02:59 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2107.codfw.wmnet with reason: host reimage * 02:55 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2085.codfw.wmnet with OS trixie * 02:51 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2072.codfw.wmnet with OS trixie * 02:51 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2107.codfw.wmnet with reason: host reimage * 02:35 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2085.codfw.wmnet with reason: host reimage * 02:31 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2107.codfw.wmnet with OS trixie * 02:30 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2072.codfw.wmnet with reason: host reimage * 02:26 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2085.codfw.wmnet with reason: host reimage * 02:22 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2072.codfw.wmnet with reason: host reimage * 02:09 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2085.codfw.wmnet with OS trixie * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 54s) * 02:03 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2072.codfw.wmnet with OS trixie * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es7 eqiad back to read-write - [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94649 and previous config saved to /var/cache/conftool/dbconfig/20260701-010716-ladsgroup.json * 01:05 ladsgroup@dns1004: END - running authdns-update * 01:05 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depool es1039 [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94648 and previous config saved to /var/cache/conftool/dbconfig/20260701-010551-ladsgroup.json * 01:03 ladsgroup@dns1004: START - running authdns-update * 01:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Promote es1035 to es7 primary [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94647 and previous config saved to /var/cache/conftool/dbconfig/20260701-010002-ladsgroup.json * 00:58 Amir1: Starting es7 eqiad failover from es1039 to es1035 - [[phab:T430765|T430765]] * 00:53 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es1035 with weight 0 [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94646 and previous config saved to /var/cache/conftool/dbconfig/20260701-005329-ladsgroup.json * 00:53 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 9 hosts with reason: Primary switchover es7 [[phab:T430765|T430765]] * 00:42 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es7 eqiad as read-only for maintenance - [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94645 and previous config saved to /var/cache/conftool/dbconfig/20260701-004221-ladsgroup.json * 00:20 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2102.codfw.wmnet with OS trixie * 00:15 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2103.codfw.wmnet with OS trixie * 00:05 dr0ptp4kt: DEPLOYED Refinery at {{Gerrit|4e7a2b32}} for changes: pageview allowlist {{Gerrit|1305158}} (+min.wikiquote) {{Gerrit|1305162}} (+bol.wikipedia), {{Gerrit|1305156}} (+isv.wikipedia); {{Gerrit|1305980}} (pv allowlist -api.wikimedia, sqoop +isvwiki); sqoop {{Gerrit|1295064}} (+globalimagelinks) {{Gerrit|1295069}} (+filerevision) using scap, then deployed onto HDFS (manual copyToLocal required additionally) == Other archives == See [[Server Admin Log/Archives]]. <noinclude> [[Category:SAL]] [[Category:Operations]] </noinclude> 3b090r7k5l5qz6yvx4s9blqa31txrpn 2450652 2450651 2026-08-22T16:36:08Z Stashbot 7414 arlolra@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply 2450652 wikitext text/x-wiki == 2026-08-22 == * 16:36 arlolra@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:33 arlolra@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:33 arlolra@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:33 arlolra@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:32 arlolra@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 35s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-21 == * 20:36 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in1001.wikimedia.org with reason: [[phab:T434750|T434750]] * 20:34 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in2001.wikimedia.org with reason: [[phab:T434750|T434750]] * 20:33 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out1001.wikimedia.org with reason: [[phab:T434750|T434750]] * 20:25 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out2001.wikimedia.org with reason: [[phab:T434750|T434750]] * 19:37 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host krb1002.eqiad.wmnet with OS bookworm * 19:00 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 18:59 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 18:51 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 18:51 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 18:35 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:35 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:27 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:27 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:16 bking@cumin2003: START - Cookbook sre.hosts.reimage for host krb1002.eqiad.wmnet with OS bookworm * 17:35 sukhe@dns1004: END - running authdns-update * 17:33 sukhe@dns1004: START - running authdns-update * 17:32 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns5004.wikimedia.org [reason: resolved authdns-update issues] * 17:32 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:32 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: force HEAD to {{Gerrit|be26e30ae101}} - sukhe@cumin1003" * 17:32 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: force HEAD to {{Gerrit|be26e30ae101}} - sukhe@cumin1003" * 17:28 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 17:28 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: service=authdns-update,name=dns5004.wikimedia.org [reason: resolving authdns-update issues] * 17:28 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:28 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: force HEAD to {{Gerrit|be26e30ae101}} - sukhe@cumin1003" * 17:28 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: force HEAD to {{Gerrit|be26e30ae101}} - sukhe@cumin1003" * 17:24 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 17:24 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.netbox (exit_code=97) * 17:23 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 17:16 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 17:12 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 17:10 sukhe@dns1004: END - running authdns-update * 17:08 sukhe@dns1004: START - running authdns-update * 17:08 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=dns5004.wikimedia.org [reason: resolving authdns-update issues] * 17:07 sukhe@dns1004: FAIL - running authdns-update * 17:05 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 17:05 sukhe@dns1004: START - running authdns-update * 17:01 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 16:59 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 16:56 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 16:53 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=dns5004.* [reason: trixie upgrade] * 16:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns5004.wikimedia.org * 16:52 cdobbins@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns5004.wikimedia.org * 16:44 cmooney@dns3003: END - running authdns-update * 16:41 cmooney@dns3003: START - running authdns-update * 16:41 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:41 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on eqsin<->codfw arelion - cmooney@cumin1003" * 16:37 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on eqsin<->codfw arelion - cmooney@cumin1003" * 16:33 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:11 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 16:08 sukhe@dns1004: END - running authdns-update * 16:08 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 16:06 sukhe@dns1004: START - running authdns-update * 16:04 cmooney@dns3003: END - running authdns-update * 16:02 cmooney@dns3003: START - running authdns-update * 16:00 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:00 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on eqord<->codfw arelion - cmooney@cumin1003" * 15:56 cdobbins@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host dns5004.wikimedia.org with OS trixie * 15:55 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on eqord<->codfw arelion - cmooney@cumin1003" * 15:53 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:51 cmooney@cumin1003: END (ERROR) - Cookbook sre.dns.netbox (exit_code=97) * 15:51 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:38 andrew@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudcephosd1042.eqiad.wmnet with OS bookworm * 15:18 andrew@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudcephosd1042.eqiad.wmnet with reason: host reimage * 15:17 dancy@deploy1003: Finished deploy [gerrit/gerrit@2cc11cc]: Deploying https://gerrit.wikimedia.org/r/c/operations/software/gerrit/+/1327669 ([[phab:T434726|T434726]]) (duration: 00m 14s) * 15:17 dancy@deploy1003: Started deploy [gerrit/gerrit@2cc11cc]: Deploying https://gerrit.wikimedia.org/r/c/operations/software/gerrit/+/1327669 ([[phab:T434726|T434726]]) * 15:13 andrew@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudcephosd1042.eqiad.wmnet with reason: host reimage * 15:09 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns5004.wikimedia.org with reason: host reimage * 15:05 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns5004.wikimedia.org with reason: host reimage * 14:53 andrew@cumin2003: START - Cookbook sre.hosts.reimage for host cloudcephosd1042.eqiad.wmnet with OS bookworm * 14:30 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns5004.wikimedia.org with OS trixie * 14:29 cdobbins@cumin1003: conftool action : set/pooled=no; selector: name=dns5004.* [reason: trixie upgrade] * 14:21 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:21 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove entries for cr2-eqord - cmooney@cumin1003" * 14:21 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove entries for cr2-eqord - cmooney@cumin1003" * 14:13 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 14:11 moritzm: imported openjdk 8u504-ga-1~deb12u1 for bookworm-wikimedia (backport of the latest Java 8 security fixes for bookworm) * 13:25 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "sync cr2-eqord router offline - cmooney@cumin1003" * 13:23 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "sync cr2-eqord router offline - cmooney@cumin1003" * 13:14 hashar@deploy1003: Finished deploy [integration/docroot@2d5ff9b]: opensource: add PersonalDashboard docs to MW components - [[phab:T435392|T435392]] (duration: 00m 15s) * 13:14 hashar@deploy1003: Started deploy [integration/docroot@2d5ff9b]: opensource: add PersonalDashboard docs to MW components - [[phab:T435392|T435392]] * 12:16 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:16 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: [[phab:T431682|T431682]] - filippo@cumin1003" * 12:16 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: [[phab:T431682|T431682]] - filippo@cumin1003" * 12:11 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2006.wikimedia.org with OS trixie * 12:00 kevinbazira@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 11:58 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 11:43 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2006.wikimedia.org with reason: host reimage * 11:41 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 11:38 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2006.wikimedia.org with reason: host reimage * 11:20 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2006.wikimedia.org with OS trixie * 11:11 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2005.wikimedia.org with OS trixie * 10:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2005.wikimedia.org with reason: host reimage * 10:53 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2005.wikimedia.org with reason: host reimage * 10:43 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-codfw * 10:43 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2011.codfw.wmnet * 10:43 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2011.codfw.wmnet * 10:40 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2011.codfw.wmnet * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2011.codfw.wmnet * 10:34 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2010.codfw.wmnet * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2010.codfw.wmnet * 10:33 fnegri@cumin1003: END (PASS) - Cookbook sre.wikireplicas.add-wiki (exit_code=0) for database bolwiki ([[phab:T429954|T429954]]) * 10:33 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2005.wikimedia.org with OS trixie * 10:30 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2010.codfw.wmnet * 10:25 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2010.codfw.wmnet * 10:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2009.codfw.wmnet * 10:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2009.codfw.wmnet * 10:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1006.wikimedia.org with OS trixie * 10:18 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2009.codfw.wmnet * 10:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2009.codfw.wmnet * 10:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2008.codfw.wmnet * 10:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2008.codfw.wmnet * 10:06 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2008.codfw.wmnet * 10:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1006.wikimedia.org with reason: host reimage * 10:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2008.codfw.wmnet * 10:01 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2007.codfw.wmnet * 10:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2007.codfw.wmnet * 09:57 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1006.wikimedia.org with reason: host reimage * 09:56 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2007.codfw.wmnet * 09:51 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2007.codfw.wmnet * 09:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 09:51 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 09:46 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2006.codfw.wmnet * 09:46 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1006.wikimedia.org with OS trixie * 09:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1005.wikimedia.org with OS trixie * 09:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2006.codfw.wmnet * 09:41 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2005.codfw.wmnet * 09:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2005.codfw.wmnet * 09:38 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2013.codfw.wmnet * 09:36 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2005.codfw.wmnet * 09:35 fnegri@cumin1003: START - Cookbook sre.wikireplicas.add-wiki for database bolwiki ([[phab:T429954|T429954]]) * 09:35 fnegri@cumin1003: END (PASS) - Cookbook sre.wikireplicas.add-wiki (exit_code=0) for database minwikiquote ([[phab:T429946|T429946]]) * 09:35 fnegri@cumin1003: START - Cookbook sre.wikireplicas.add-wiki for database minwikiquote ([[phab:T429946|T429946]]) * 09:32 blake@cumin1003: START - Cookbook sre.hosts.reboot-single for host rdb2013.codfw.wmnet * 09:30 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2011.codfw.wmnet * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1005.wikimedia.org with reason: host reimage * 09:26 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2005.codfw.wmnet * 09:26 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2004.codfw.wmnet * 09:26 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2004.codfw.wmnet * 09:24 blake@cumin1003: START - Cookbook sre.hosts.reboot-single for host rdb2011.codfw.wmnet * 09:22 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1005.wikimedia.org with reason: host reimage * 09:21 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2004.codfw.wmnet * 09:16 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1015.eqiad.wmnet * 09:16 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2004.codfw.wmnet * 09:16 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2003.codfw.wmnet * 09:16 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2003.codfw.wmnet * 09:13 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on cr[1-2]-eqiad,pfw1-eqiad with reason: upgrade pfw1a-eqiad and pfw1b-eqiad pair * 09:12 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 09:11 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 09:11 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 09:11 blake@cumin1003: START - Cookbook sre.hosts.reboot-single for host rdb1015.eqiad.wmnet * 09:10 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2003.codfw.wmnet * 09:09 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1013.eqiad.wmnet * 09:07 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1005.wikimedia.org with OS trixie * 09:03 blake@cumin1003: START - Cookbook sre.hosts.reboot-single for host rdb1013.eqiad.wmnet * 09:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2003.codfw.wmnet * 09:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2002.codfw.wmnet * 09:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2002.codfw.wmnet * 08:54 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2002.codfw.wmnet * 08:49 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2002.codfw.wmnet * 08:49 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2001.codfw.wmnet * 08:49 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2001.codfw.wmnet * 08:48 jmm@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts netmon2002.wikimedia.org * 08:47 jmm@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts netmon2002.wikimedia.org * 08:44 jmm@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts netmon2002.wikimedia.org * 08:44 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon2002.wikimedia.org * 08:43 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2001.codfw.wmnet * 08:36 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon2002.wikimedia.org * 08:34 jmm@dns1004: END - running authdns-update * 08:33 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2001.codfw.wmnet * 08:33 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-codfw * 08:31 jmm@dns1004: START - running authdns-update * 07:48 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327679{{!}}Block: Disable flaky API test (T435272 T389028)]], [[gerrit:1327678{{!}}API: wfDebugLog for thumberror]] (duration: 15m 34s) * 07:41 krinkle@deploy1003: krinkle: Continuing with deployment * 07:37 krinkle@deploy1003: krinkle: Backport for [[gerrit:1327679{{!}}Block: Disable flaky API test (T435272 T389028)]], [[gerrit:1327678{{!}}API: wfDebugLog for thumberror]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:33 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1327679{{!}}Block: Disable flaky API test (T435272 T389028)]], [[gerrit:1327678{{!}}API: wfDebugLog for thumberror]] * 07:25 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 07:24 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 07:18 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 07:18 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 07:15 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 07:14 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 07:14 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 07:14 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 07:13 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 07:03 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1283: Pool back * 06:42 jmm@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts netmon2002.wikimedia.org * 06:35 moritzm: powercycling netmon2002 * 06:18 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1283: Pool back * 06:17 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1283 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96213 and previous config saved to /var/cache/conftool/dbconfig/20260821-061743-marostegui.json * 04:59 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324963{{!}}Add Produnto to extension-list (T421436)]], [[gerrit:1324964{{!}}Enable Produnto on Beta (T421436)]] (duration: 34m 48s) * 04:45 tstarling@deploy1003: tstarling: Continuing with deployment * 04:44 tstarling@deploy1003: tstarling: Backport for [[gerrit:1324963{{!}}Add Produnto to extension-list (T421436)]], [[gerrit:1324964{{!}}Enable Produnto on Beta (T421436)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 04:24 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1324963{{!}}Add Produnto to extension-list (T421436)]], [[gerrit:1324964{{!}}Enable Produnto on Beta (T421436)]] * 04:21 arlolra@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 04:20 arlolra@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 04:20 arlolra@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 04:20 arlolra@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 41s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-20 == * 23:43 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327652{{!}}RunSingleJob: Add ProfilingContext::init() (T435422)]] (duration: 12m 06s) * 23:38 krinkle@deploy1003: krinkle: Continuing with deployment * 23:33 krinkle@deploy1003: krinkle: Backport for [[gerrit:1327652{{!}}RunSingleJob: Add ProfilingContext::init() (T435422)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:31 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1327652{{!}}RunSingleJob: Add ProfilingContext::init() (T435422)]] * 22:15 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1054.eqiad.wmnet with OS trixie * 22:14 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 22:14 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 21:58 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1054.eqiad.wmnet with reason: host reimage * 21:51 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1054.eqiad.wmnet with reason: host reimage * 21:36 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1054.eqiad.wmnet with OS trixie * 21:36 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:35 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327219{{!}}RunSingleJob: Define MW_ENTRY_POINT for flamegraph sample attribution (T435422)]] (duration: 08m 30s) * 21:31 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:31 krinkle@deploy1003: krinkle: Continuing with deployment * 21:31 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1054 * 21:31 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1054 * 21:29 krinkle@deploy1003: krinkle: Backport for [[gerrit:1327219{{!}}RunSingleJob: Define MW_ENTRY_POINT for flamegraph sample attribution (T435422)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:27 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1327219{{!}}RunSingleJob: Define MW_ENTRY_POINT for flamegraph sample attribution (T435422)]] * 21:17 maryum: Deployed security fix for [[phab:T433020|T433020]] * 20:59 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324752{{!}}InitialiseSettings: Enable 2FA banners on remaining private wikis (T428103)]], [[gerrit:1325920{{!}}Remove sending email to legal team about rejected requests (T374053)]] (duration: 07m 18s) * 20:54 reedy@deploy1003: neriah, reedy: Continuing with deployment * 20:54 reedy@deploy1003: neriah, reedy: Backport for [[gerrit:1324752{{!}}InitialiseSettings: Enable 2FA banners on remaining private wikis (T428103)]], [[gerrit:1325920{{!}}Remove sending email to legal team about rejected requests (T374053)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:51 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324752{{!}}InitialiseSettings: Enable 2FA banners on remaining private wikis (T428103)]], [[gerrit:1325920{{!}}Remove sending email to legal team about rejected requests (T374053)]] * 20:24 reedy@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.15,1.47.0-wmf.16,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/med * 20:23 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324752{{!}}InitialiseSettings: Enable 2FA banners on remaining private wikis (T428103)]], [[gerrit:1325920{{!}}Remove sending email to legal team about rejected requests (T374053)]] * 20:14 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 20:10 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 20:09 cdanis@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "bug fixes & UX fixes - cdanis@cumin1003" * 20:09 cdanis@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: bug fixes & UX fixes - cdanis@cumin1003 * 20:08 cdanis@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: bug fixes & UX fixes - cdanis@cumin1003 * 20:08 cdanis@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "bug fixes & UX fixes - cdanis@cumin1003" * 19:24 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327598{{!}}Make \Omicron non upright (like \Chi) (T434428)]], [[gerrit:1327596{{!}}Render overline of \bar with stretchy=false (T435456)]] (duration: 18m 54s) * 19:20 krinkle@deploy1003: krinkle: Continuing with deployment * 19:07 krinkle@deploy1003: krinkle: Backport for [[gerrit:1327598{{!}}Make \Omicron non upright (like \Chi) (T434428)]], [[gerrit:1327596{{!}}Render overline of \bar with stretchy=false (T435456)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:05 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1327598{{!}}Make \Omicron non upright (like \Chi) (T434428)]], [[gerrit:1327596{{!}}Render overline of \bar with stretchy=false (T435456)]] * 18:51 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327614{{!}}Avoid casting fpxmax to string (T318419)]] (duration: 07m 28s) * 18:50 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:46 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:46 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 18:45 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327614{{!}}Avoid casting fpxmax to string (T318419)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:43 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327614{{!}}Avoid casting fpxmax to string (T318419)]] * 18:37 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:37 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:37 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:36 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:07 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 18:05 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 18:01 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 18:01 sukhe@dns1004: END - running authdns-update * 17:59 sukhe@dns1004: START - running authdns-update * 17:58 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 17:57 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=dns6002.* [reason: depooling for trixie upgrade] * 17:56 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns6002.wikimedia.org * 17:56 cdobbins@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns6002.wikimedia.org * 17:51 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 17:51 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 17:34 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host stat1011.eqiad.wmnet with OS bookworm * 17:31 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns6002.wikimedia.org with OS trixie * 17:30 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 17:30 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 17:30 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 17:30 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 17:29 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 17:29 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 17:29 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 17:29 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:29 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:27 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:24 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327590{{!}}Make sure fpsmax is an int value (T318419)]] (duration: 08m 37s) * 17:20 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 17:17 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327590{{!}}Make sure fpsmax is an int value (T318419)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:16 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 17:15 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327590{{!}}Make sure fpsmax is an int value (T318419)]] * 16:51 swfrench-wmf: disable-puppet on A:cp for ATS Lua change - [[phab:T427666|T427666]] * 16:51 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db2901.codfw.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 16:43 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 16:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on stat1011.eqiad.wmnet with reason: host reimage * 16:39 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns6002.wikimedia.org with reason: host reimage * 16:36 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on stat1011.eqiad.wmnet with reason: host reimage * 16:36 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db2901.codfw.wmnet * 16:34 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns6002.wikimedia.org with reason: host reimage * 16:15 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns6002.wikimedia.org with OS trixie * 16:14 cdobbins@cumin1003: conftool action : set/pooled=no; selector: name=dns6002.* [reason: depooling for trixie upgrade] * 16:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host stat1011 * 16:10 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host stat1011 * 16:09 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host stat1011 * 16:09 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) stat1011.eqiad.wmnet 14.36.64.10.in-addr.arpa 4.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 bking@cumin2003: START - Cookbook sre.dns.wipe-cache stat1011.eqiad.wmnet 14.36.64.10.in-addr.arpa 4.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:09 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host stat1011 - bking@cumin2003" * 16:09 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host stat1011 - bking@cumin2003" * 16:05 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host stat1011 * 16:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host stat1011.eqiad.wmnet with OS bookworm * 16:00 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 15:59 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 15:56 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 15:55 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 15:37 jayme@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:35 jayme@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 15:35 jayme@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:32 jayme@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:32 jayme@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:30 jayme@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 15:30 jayme@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:28 jayme@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:28 jayme@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 15:28 fceratto@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host db1903.eqiad.wmnet * 15:28 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1903.eqiad.wmnet with OS trixie * 15:26 jayme@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 15:26 jayme@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 15:24 jayme@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 15:24 jayme@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 15:21 jayme@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 15:21 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 15:19 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 15:19 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 15:17 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 15:14 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1903.eqiad.wmnet with reason: host reimage * 15:07 fceratto@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1903.eqiad.wmnet with reason: host reimage * 14:54 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db1903.eqiad.wmnet with OS trixie * 14:53 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1903.eqiad.wmnet - fceratto@cumin1003" * 14:53 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1903.eqiad.wmnet - fceratto@cumin1003" * 14:53 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1903.eqiad.wmnet on all recursors * 14:53 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1903.eqiad.wmnet on all recursors * 14:53 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:53 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1903.eqiad.wmnet - fceratto@cumin1003" * 14:53 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1903.eqiad.wmnet - fceratto@cumin1003" * 14:49 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 14:49 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1903.eqiad.wmnet * 14:33 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=cp1100.* * 14:27 topranks: reconfigure eqiad<->codfw bgp settings * 14:22 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:22 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update entries used on new transport backup eqiad codfw - cmooney@cumin1003" * 14:19 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update entries used on new transport backup eqiad codfw - cmooney@cumin1003" * 14:14 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 14:14 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 14:13 moritzm: installing util-linux security updates * 14:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-staging-worker * 14:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2003.codfw.wmnet * 14:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2003.codfw.wmnet * 14:08 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2003.codfw.wmnet * 14:06 moritzm: installing libheif security updates * 13:58 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2003.codfw.wmnet * 13:58 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2002.codfw.wmnet * 13:58 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2002.codfw.wmnet * 13:56 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:56 fnegri@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for clouddb1025.eqiad.wmnet * 13:56 fnegri@cumin1003: START - Cookbook sre.hosts.remove-downtime for clouddb1025.eqiad.wmnet * 13:56 Lucas_WMDE: UTC afternoon backport+config window done * 13:53 moritzm: installing apr-util security updates * 13:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2002.codfw.wmnet * 13:50 fnegri@cumin1003: conftool action : set/weight=100; selector: name=clouddb1025.eqiad.wmnet * 13:49 fnegri@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1025.eqiad.wmnet * 13:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2002.codfw.wmnet * 13:41 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2001.codfw.wmnet * 13:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2001.codfw.wmnet * 13:41 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'. * 13:38 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'. * 13:38 fnegri@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on clouddb1025.eqiad.wmnet with reason: Removing s6 from clouddb1025 * 13:34 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2001.codfw.wmnet * 13:31 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'. * 13:29 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'. * 13:28 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet * 13:26 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host stat1009.eqiad.wmnet with OS bookworm * 13:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2001.codfw.wmnet * 13:24 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-staging-worker * 13:23 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1002.eqiad.wmnet * 13:21 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host stat1010.eqiad.wmnet with OS bookworm * 13:20 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1002.eqiad.wmnet * 13:20 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1001.eqiad.wmnet * 13:17 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1001.eqiad.wmnet * 13:16 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2001.codfw.wmnet * 13:13 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327511{{!}}UIC: Fix page:page instead of page:other in instrumentation]] (duration: 07m 00s) * 13:13 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2001.codfw.wmnet * 13:12 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2002.codfw.wmnet * 13:09 mszwarc@deploy1003: mszwarc: Continuing with deployment * 13:08 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1327511{{!}}UIC: Fix page:page instead of page:other in instrumentation]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2002.codfw.wmnet * 13:07 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2002.codfw.wmnet * 13:06 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1327511{{!}}UIC: Fix page:page instead of page:other in instrumentation]] * 13:04 jmm@dns1004: END - running authdns-update * 13:03 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2002.codfw.wmnet * 13:03 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2001.codfw.wmnet * 13:02 jmm@dns1004: START - running authdns-update * 13:01 cmooney@dns3003: END - running authdns-update * 13:00 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2001.codfw.wmnet * 12:59 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2001.codfw.wmnet * 12:59 cmooney@dns3003: START - running authdns-update * 12:57 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2001.codfw.wmnet * 12:56 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2002.codfw.wmnet * 12:55 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:55 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on drmrs<->eqiad GTT vpls - cmooney@cumin1003" * 12:54 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on drmrs<->eqiad GTT vpls - cmooney@cumin1003" * 12:54 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2002.codfw.wmnet * 12:54 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2003.codfw.wmnet * 12:50 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2003.codfw.wmnet * 12:49 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 12:48 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2001.codfw.wmnet * 12:46 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2001.codfw.wmnet * 12:46 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2002.codfw.wmnet * 12:43 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2002.codfw.wmnet * 12:43 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2003.codfw.wmnet * 12:42 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=1) for new host db1902.eqiad.wmnet * 12:42 fceratto@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host db1902.eqiad.wmnet with OS trixie * 12:41 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2003.codfw.wmnet * 12:40 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1003.eqiad.wmnet * 12:38 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1003.eqiad.wmnet * 12:37 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1002.eqiad.wmnet * 12:35 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1002.eqiad.wmnet * 12:35 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1001.eqiad.wmnet * 12:34 cmooney@dns3003: END - running authdns-update * 12:33 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1001.eqiad.wmnet * 12:32 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on stat1009.eqiad.wmnet with reason: host reimage * 12:31 cmooney@dns3003: START - running authdns-update * 12:31 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:31 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on drmrs<->eqiad cct - cmooney@cumin1003" * 12:28 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on drmrs<->eqiad cct - cmooney@cumin1003" * 12:28 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1902.eqiad.wmnet with reason: host reimage * 12:25 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on stat1009.eqiad.wmnet with reason: host reimage * 12:25 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 12:24 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on stat1010.eqiad.wmnet with reason: host reimage * 12:22 fceratto@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1902.eqiad.wmnet with reason: host reimage * 12:21 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on stat1010.eqiad.wmnet with reason: host reimage * 12:14 elukey: move the Docker Registry's /v2/dev/.* prefix to its dedicated S3 backend - [[phab:T432829|T432829]] * 12:12 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db1902.eqiad.wmnet with OS trixie * 12:09 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1902.eqiad.wmnet - fceratto@cumin1003" * 12:09 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1902.eqiad.wmnet - fceratto@cumin1003" * 12:09 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1902.eqiad.wmnet on all recursors * 12:09 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1902.eqiad.wmnet on all recursors * 12:09 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:08 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1902.eqiad.wmnet - fceratto@cumin1003" * 12:08 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1902.eqiad.wmnet - fceratto@cumin1003" * 12:08 tgr_: [[phab:T413390|T413390]] running CentralAuth:FixRenamedUserGlobalEditCount --wiki=metawiki --since=20250901000000 --until=20260301000000 --fix * 12:04 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1009.eqiad.wmnet with OS bookworm * 12:04 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1010.eqiad.wmnet with OS bookworm * 12:01 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 12:01 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1902.eqiad.wmnet * 12:00 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host stat1010.eqiad.wmnet with OS bookworm * 11:57 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1902.eqiad.wmnet * 11:57 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:57 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1902.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 11:57 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1902.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 11:51 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'. * 11:49 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'. * 11:48 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'. * 11:46 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'. * 11:39 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 11:37 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327518{{!}}Enable thumb.wikimedia.org on cswiki and fawiki (T427465)]] (duration: 10m 40s) * 11:35 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1902.eqiad.wmnet * 11:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1010.eqiad.wmnet with OS bookworm * 11:33 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 11:30 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327518{{!}}Enable thumb.wikimedia.org on cswiki and fawiki (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:26 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327518{{!}}Enable thumb.wikimedia.org on cswiki and fawiki (T427465)]] * 11:24 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host stat1010.eqiad.wmnet with OS bookworm * 11:07 fceratto@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host db1901.eqiad.wmnet * 11:07 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1901.eqiad.wmnet with OS trixie * 10:53 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1901.eqiad.wmnet with reason: host reimage * 10:47 tappof: bump space for prometheus k8s-aux in codfw * 10:47 tappof: bump space for prometheus k8s-dse in eqiad * 10:47 fceratto@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1901.eqiad.wmnet with reason: host reimage * 10:35 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db1901.eqiad.wmnet with OS trixie * 10:32 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:32 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:32 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1901.eqiad.wmnet on all recursors * 10:32 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1901.eqiad.wmnet on all recursors * 10:31 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:31 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:31 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:27 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:27 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1901.eqiad.wmnet * 10:23 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1010.eqiad.wmnet with OS bookworm * 10:20 fceratto@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts db1901.eqiad.wmnet * 10:20 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 10:18 blake@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 10:17 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:16 blake@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 10:13 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1901.eqiad.wmnet * 09:23 jelto@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'. * 09:22 jelto@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'. * 09:22 jelto: update cert-manager to 1.19.6 on wikikube staging-eqiad - [[phab:T427402|T427402]] * 09:20 moritzm: imported squid 7.6-2.1for trixie-wikimedia/main [[phab:T427282|T427282]] * 09:08 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 09:08 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 09:08 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 09:07 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 09:04 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 09:04 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:23 slyngshede@dns1004: END - running authdns-update * 08:21 slyngshede@dns1004: START - running authdns-update * 08:18 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.16 refs [[phab:T430835|T430835]] * 06:27 aokoth@dns1004: END - running authdns-update * 06:25 aokoth@dns1004: START - running authdns-update * 06:22 brennen@deploy1003: Finished deploy [phabricator/deployment@6b9b6ff]: deploy phab1005 for [[phab:T435087|T435087]] (duration: 00m 39s) * 06:21 brennen@deploy1003: Started deploy [phabricator/deployment@6b9b6ff]: deploy phab1005 for [[phab:T435087|T435087]] * 06:20 brennen@deploy1003: Finished deploy [phabricator/deployment@6b9b6ff]: deploy phab1004 for to pick up config values for [[phab:T435087|T435087]] (duration: 01m 46s) * 06:18 brennen@deploy1003: Started deploy [phabricator/deployment@6b9b6ff]: deploy phab1004 for to pick up config values for [[phab:T435087|T435087]] * 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 49s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-19 == * 23:19 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327197{{!}}Enable thumb.wikimedia.org on mediawiki.org (T427465)]] (duration: 10m 50s) * 23:18 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:16 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 23:15 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 23:10 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327197{{!}}Enable thumb.wikimedia.org on mediawiki.org (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:10 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:09 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 23:09 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:09 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 23:08 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327197{{!}}Enable thumb.wikimedia.org on mediawiki.org (T427465)]] * 22:58 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327201{{!}}Enable ReadingLists for all logged in users on test wiki (T435258)]] (duration: 11m 20s) * 22:50 jdlrobson@deploy1003: jdlrobson: Continuing with deployment * 22:49 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1327201{{!}}Enable ReadingLists for all logged in users on test wiki (T435258)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:46 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1327201{{!}}Enable ReadingLists for all logged in users on test wiki (T435258)]] * 22:42 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327162{{!}}Article: Split subjectpageheader by model and disable for wikitext]], [[gerrit:1327169{{!}}Make uppercase greek letters normal (non-italic) font (T434686 T434428)]], [[gerrit:1327176{{!}}Skin: Avoid DB lookup for pagecategorieslink message (T347123)]] (duration: 37m 52s) * 22:29 krinkle@deploy1003: krinkle: Continuing with deployment * 22:25 krinkle@deploy1003: krinkle: Backport for [[gerrit:1327162{{!}}Article: Split subjectpageheader by model and disable for wikitext]], [[gerrit:1327169{{!}}Make uppercase greek letters normal (non-italic) font (T434686 T434428)]], [[gerrit:1327176{{!}}Skin: Avoid DB lookup for pagecategorieslink message (T347123)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:04 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1327162{{!}}Article: Split subjectpageheader by model and disable for wikitext]], [[gerrit:1327169{{!}}Make uppercase greek letters normal (non-italic) font (T434686 T434428)]], [[gerrit:1327176{{!}}Skin: Avoid DB lookup for pagecategorieslink message (T347123)]] * 22:04 krinkle@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: awaiting CI (duration: 03m 06s) * 22:01 krinkle@deploy1003: Locking from deployment [ALL REPOSITORIES]: awaiting CI * 22:00 krinkle@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: awaiting CI (duration: 00m 01s) * 22:00 krinkle@deploy1003: Locking from deployment [ALL REPOSITORIES]: awaiting CI * 21:34 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2001.codfw.wmnet * 21:28 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2001.codfw.wmnet * 21:22 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327178{{!}}AccountRecovery: Notify the email address of the on file of the request (T425799)]] (duration: 47m 02s) * 21:13 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm * 21:09 catrope@deploy1003: catrope: Continuing with deployment * 20:55 catrope@deploy1003: catrope: Backport for [[gerrit:1327178{{!}}AccountRecovery: Notify the email address of the on file of the request (T425799)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:35 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1327178{{!}}AccountRecovery: Notify the email address of the on file of the request (T425799)]] * 20:31 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327128{{!}}Parsoid DataAccess: convert Parsoid fragment markers to/from strip tags (T432547)]] (duration: 07m 30s) * 20:27 catrope@deploy1003: catrope, arlolra: Continuing with deployment * 20:26 catrope@deploy1003: catrope, arlolra: Backport for [[gerrit:1327128{{!}}Parsoid DataAccess: convert Parsoid fragment markers to/from strip tags (T432547)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:24 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1327128{{!}}Parsoid DataAccess: convert Parsoid fragment markers to/from strip tags (T432547)]] * 20:23 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage * 20:17 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage * 20:15 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325878{{!}}[arwiki] Enable restricted user page editing and grant edit permissions (T434878)]] (duration: 08m 46s) * 20:11 catrope@deploy1003: catrope, gergesshamon: Continuing with deployment * 20:08 catrope@deploy1003: catrope, gergesshamon: Backport for [[gerrit:1325878{{!}}[arwiki] Enable restricted user page editing and grant edit permissions (T434878)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:06 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1325878{{!}}[arwiki] Enable restricted user page editing and grant edit permissions (T434878)]] * 19:59 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm * 19:56 eevans@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cassandra-dev2001.codfw.wmnet with OS bookworm * 19:56 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm * 19:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2207.codfw.wmnet with reason: Maintenance * 18:47 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319920{{!}}Allow setting a separate thumbUrl in production (T427465)]], [[gerrit:1327167{{!}}Fix wmgThumbUrl config (T427465)]] (duration: 18m 53s) * 18:43 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 18:30 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1319920{{!}}Allow setting a separate thumbUrl in production (T427465)]], [[gerrit:1327167{{!}}Fix wmgThumbUrl config (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:28 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1319920{{!}}Allow setting a separate thumbUrl in production (T427465)]], [[gerrit:1327167{{!}}Fix wmgThumbUrl config (T427465)]] * 18:26 sukhe@dns1004: END - running authdns-update * 18:24 sukhe@dns1004: START - running authdns-update * 18:09 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1319920{{!}}Allow setting a separate thumbUrl in production (T427465)]] * 18:03 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-eqiad and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 17:56 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-codfw and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 17:50 cmooney@dns3003: END - running authdns-update * 17:42 dancy@deploy1003: Installation of scap version "4.283.0" completed for 3 hosts * 17:41 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling reboot on A:durum and A:durum * 17:40 cmooney@dns3003: START - running authdns-update * 17:40 dancy@deploy1003: Installing scap version "4.283.0" for 3 host(s) * 17:38 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:37 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on GTT VPLS - cmooney@cumin1003" * 17:37 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --verbose --scoreLessThan=0.7 --exceptDatasetChecksums=[[phab:T434319|T434319]]-enwiki-models.txt # [[phab:T434319|T434319]] * 17:32 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on GTT VPLS - cmooney@cumin1003" * 17:29 sbassett: Deployed security fix for [[phab:T435210|T435210]] (wmf.16) * 17:26 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 17:22 sbassett: Deployed security fix for [[phab:T435210|T435210]] (wmf.15) * 17:00 sukhe@dns1004: END - running authdns-update * 16:58 sukhe@dns1004: START - running authdns-update * 16:53 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica-esams and A:liberica * 16:41 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica-esams and A:liberica * 16:41 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-eqiad and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 16:41 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-codfw and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 16:41 cjd91: sudo -i cookbook sre.cdn.roll-upgrade-ats --query 'A:cp-codfw' --task-id [[phab:T434478|T434478]] --reason '9.2.15 upgrade' * 16:41 cjd91: sudo -i cookbook sre.cdn.roll-upgrade-ats --query 'A:cp-eqiad' --task-id [[phab:T434478|T434478]] --reason '9.2.15 upgrade' * 16:40 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and A:durum * 16:28 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326881{{!}}Echo: Start using virtual domains (T380385)]] (duration: 13m 12s) * 16:23 urbanecm@deploy1003: urbanecm: Continuing with deployment * 16:21 urandom: Completed sessionstore Cassandra/JVM upgrade — [[phab:T435154|T435154]] * 16:21 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching sessionstore[2005-2006].codfw.wmnet,sessionstore[1005-1006].eqiad.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 16:19 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1326881{{!}}Echo: Start using virtual domains (T380385)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:15 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:15 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->codfw - cmooney@cumin1003" * 16:14 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1326881{{!}}Echo: Start using virtual domains (T380385)]] * 16:14 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327123{{!}}Revert^2 "Migrate database access to virtual domains" (T435305)]], [[gerrit:1327124{{!}}Pass the mapped domain of virtual-echo-shared to the push NameTableStores (T435305)]] (duration: 07m 42s) * 16:13 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching sessionstore[2005-2006].codfw.wmnet,sessionstore[1005-1006].eqiad.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 16:11 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->codfw - cmooney@cumin1003" * 16:10 urbanecm@deploy1003: urbanecm: Continuing with deployment * 16:08 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1327123{{!}}Revert^2 "Migrate database access to virtual domains" (T435305)]], [[gerrit:1327124{{!}}Pass the mapped domain of virtual-echo-shared to the push NameTableStores (T435305)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:06 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 16:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 16:06 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1327123{{!}}Revert^2 "Migrate database access to virtual domains" (T435305)]], [[gerrit:1327124{{!}}Pass the mapped domain of virtual-echo-shared to the push NameTableStores (T435305)]] * 16:06 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:03 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching sessionstore1004.eqiad.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 16:01 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching sessionstore1004.eqiad.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 15:56 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching sessionstore2004.codfw.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 15:54 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching sessionstore2004.codfw.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 15:53 urandom: beginning sessionstore Cassandra/JVM upgrade — [[phab:T435154|T435154]] * 15:52 urandom: beginning sessionstore Cassandra/JVM upgrade — [[phab:T432944|T432944]] * 15:51 cmooney@dns3003: END - running authdns-update * 15:49 cmooney@dns3003: START - running authdns-update * 15:48 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:48 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->codfw - cmooney@cumin1003" * 15:45 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->codfw - cmooney@cumin1003" * 15:44 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 15:44 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:43 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 15:42 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:42 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:38 cmooney@dns3003: END - running authdns-update * 15:36 cmooney@dns3003: START - running authdns-update * 15:36 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:36 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->eqsin - cmooney@cumin1003" * 15:34 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327138{{!}}Enable redis lock manager everywhere (T366938)]] (duration: 08m 36s) * 15:33 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->eqsin - cmooney@cumin1003" * 15:30 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:29 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 15:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1147.eqiad.wmnet with OS bookworm * 15:28 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327138{{!}}Enable redis lock manager everywhere (T366938)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:25 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327138{{!}}Enable redis lock manager everywhere (T366938)]] * 15:24 jmm@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host krb1002.eqiad.wmnet * 15:19 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2207.codfw.wmnet with reason: Host crashed * 15:17 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327098{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]], [[gerrit:1327101{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]] (duration: 07m 13s) * 15:12 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 15:12 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1327098{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]], [[gerrit:1327101{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:10 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1327098{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]], [[gerrit:1327101{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]] * 15:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2010.codfw.wmnet with OS trixie * 15:05 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 15:05 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1147.eqiad.wmnet with reason: host reimage * 14:59 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 14:59 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:58 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1147.eqiad.wmnet with reason: host reimage * 14:55 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:55 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:49 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 14:48 cmooney@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host durum1001.eqiad.wmnet * 14:46 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:46 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Delete db2902 ipv6 addr - fceratto@cumin1003" * 14:46 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Delete db2902 ipv6 addr - fceratto@cumin1003" * 14:43 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1147.eqiad.wmnet with OS bookworm * 14:42 tgr@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327094{{!}}SpecialMWOAuthListConsumers: Handle newFromMWUser returning null in addNavigationSubtitle (T435167)]] (duration: 19m 25s) * 14:42 cmooney@cumin1003: START - Cookbook sre.hosts.reboot-single for host durum1001.eqiad.wmnet * 14:42 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 14:41 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 14:38 tgr@deploy1003: tgr: Continuing with deployment * 14:36 tgr@deploy1003: tgr: Backport for [[gerrit:1327094{{!}}SpecialMWOAuthListConsumers: Handle newFromMWUser returning null in addNavigationSubtitle (T435167)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:28 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 14:28 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:28 cmooney@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host durum3005.esams.wmnet * 14:25 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 14:23 cmooney@cumin1003: START - Cookbook sre.hosts.reboot-single for host durum3005.esams.wmnet * 14:23 tgr@deploy1003: Started scap sync-world: Backport for [[gerrit:1327094{{!}}SpecialMWOAuthListConsumers: Handle newFromMWUser returning null in addNavigationSubtitle (T435167)]] * 14:18 gengh@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:17 topranks: disable puppet on hosts running BIRD BGP to test merge of patch to systemd healtchcheck service * 14:17 gengh@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:17 gengh@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:16 gengh@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:16 elukey: upgrade spicerack on cumin1003 and cumin2003 * 14:16 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:15 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314962{{!}}static: add new dir bimi/ for BIMI SVG and PEM file (T311685)]] (duration: 10m 00s) * 14:15 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:11 kharlan@deploy1003: kharlan, sukhe: Continuing with deployment * 14:11 gengh@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:09 gengh@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:09 gengh@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:08 kharlan@deploy1003: kharlan, sukhe: Backport for [[gerrit:1314962{{!}}static: add new dir bimi/ for BIMI SVG and PEM file (T311685)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:06 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 14:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:06 gengh@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:05 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1314962{{!}}static: add new dir bimi/ for BIMI SVG and PEM file (T311685)]] * 14:05 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host krb1002.eqiad.wmnet * 14:05 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:04 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:04 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326342{{!}}Revert^2 "wmf-config/ProductionServices: set URL for urldownloader to service record"]] (duration: 07m 40s) * 13:59 kharlan@deploy1003: kharlan, sukhe: Continuing with deployment * 13:58 kharlan@deploy1003: kharlan, sukhe: Backport for [[gerrit:1326342{{!}}Revert^2 "wmf-config/ProductionServices: set URL for urldownloader to service record"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:57 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host stat1008.eqiad.wmnet with OS bookworm * 13:56 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1326342{{!}}Revert^2 "wmf-config/ProductionServices: set URL for urldownloader to service record"]] * 13:56 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 13:54 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325532{{!}}srwiki: Allow bureaucrats to add and remove event-organizer group (T434748)]] (duration: 14m 56s) * 13:54 swfrench@dns1004: END - running authdns-update * 13:52 swfrench@dns1004: START - running authdns-update * 13:51 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:51 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 13:48 kharlan@deploy1003: kharlan, danielyepezgarces: Continuing with deployment * 13:45 swfrench@cumin2003: conftool action : set/pooled=yes; selector: name=wikikube-worker2330.codfw.wmnet * 13:44 swfrench-wmf: finished etcd-main codfw -> eqiad switchover - [[phab:T435103|T435103]] * 13:44 kharlan@deploy1003: kharlan, danielyepezgarces: Backport for [[gerrit:1325532{{!}}srwiki: Allow bureaucrats to add and remove event-organizer group (T434748)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:44 swfrench@cumin2003: conftool action : set/pooled=no; selector: name=wikikube-worker2330.codfw.wmnet * 13:41 swfrench@dns1004: END - running authdns-update * 13:39 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1325532{{!}}srwiki: Allow bureaucrats to add and remove event-organizer group (T434748)]] * 13:39 swfrench@dns1004: START - running authdns-update * 13:37 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327111{{!}}Special:AbuseReview: Add "no further action needed" review action (T435020)]], [[gerrit:1327110{{!}}AbuseReview: Take the review verdict as a REST path parameter (T435020)]] (duration: 31m 43s) * 13:31 swfrench-wmf: starting etcd-main codfw -> eqiad switchover - [[phab:T435103|T435103]] * 13:28 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:28 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 13:25 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=97) rolling reboot on A:durum and A:durum * 13:24 kharlan@deploy1003: kharlan: Continuing with deployment * 13:24 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1147 * 13:24 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1147 * 13:23 kharlan@deploy1003: kharlan: Backport for [[gerrit:1327111{{!}}Special:AbuseReview: Add "no further action needed" review action (T435020)]], [[gerrit:1327110{{!}}AbuseReview: Take the review verdict as a REST path parameter (T435020)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:16 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:16 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 13:12 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:12 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 13:10 cdobbins@cumin1003: END (ERROR) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=97) Rolling upgrade of ATS on A:cp-codfw and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 13:10 cdobbins@cumin1003: END (ERROR) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=97) Rolling upgrade of ATS on A:cp-eqiad and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 13:06 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1327111{{!}}Special:AbuseReview: Add "no further action needed" review action (T435020)]], [[gerrit:1327110{{!}}AbuseReview: Take the review verdict as a REST path parameter (T435020)]] * 13:06 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-eqiad and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 13:05 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-codfw and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 13:05 cjd91: sudo -i cookbook sre.cdn.roll-upgrade-ats --query 'A:cp-eqiad' --task-id [[phab:T434478|T434478]] --reason '9.2.15 upgrade' * 13:03 swfrench@dns1004: END - running authdns-update * 13:01 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:01 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 13:01 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host ncmonitor1001.eqiad.wmnet * 13:00 swfrench@dns1004: START - running authdns-update * 12:59 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 12:59 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 12:59 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 12:59 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 12:57 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and A:durum * 12:53 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on db2902.codfw.wmnet with reason: Cloning * 12:48 cmooney@dns3003: END - running authdns-update * 12:48 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327096{{!}}Switch to redis lock manager on s4 and s8 (T366938)]] (duration: 09m 11s) * 12:46 cmooney@dns3003: START - running authdns-update * 12:45 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:45 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->eqord cct - cmooney@cumin1003" * 12:45 elukey: move the /v2/releng.* prefix on the Docker Registry to its new s3 backend - [[phab:T432829|T432829]] * 12:43 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 12:42 jelto@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 12:42 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->eqord cct - cmooney@cumin1003" * 12:42 jelto@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 12:41 jelto: update cert-manager to 1.19.6 on wikikube staging-codfw - [[phab:T427402|T427402]] * 12:40 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327096{{!}}Switch to redis lock manager on s4 and s8 (T366938)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:38 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327096{{!}}Switch to redis lock manager on s4 and s8 (T366938)]] * 12:38 blake@deploy1003: Finished scap sync-world: non-build deploy for [[phab:T417800|T417800]] (duration: 03m 52s) * 12:36 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 12:35 blake@deploy1003: Started scap sync-world: non-build deploy for [[phab:T417800|T417800]] * 12:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host krb2002.codfw.wmnet * 11:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host krb2002.codfw.wmnet * 11:49 moritzm: installing kerberos security updates * 11:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on stat1008.eqiad.wmnet with reason: host reimage * 11:44 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on stat1008.eqiad.wmnet with reason: host reimage * 11:31 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327084{{!}}Revert "Migrate database access to virtual domains" (T435305)]] (duration: 11m 02s) * 11:29 gkyziridis@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:29 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:29 kart_: Updated MinT to 2026-06-04-131507-production ([[phab:T321316|T321316]]) * 11:28 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/machinetranslation: apply * 11:28 gkyziridis@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:26 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 11:24 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:24 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:23 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:23 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:23 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/machinetranslation: apply * 11:22 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327084{{!}}Revert "Migrate database access to virtual domains" (T435305)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:21 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/machinetranslation: apply * 11:21 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:21 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:20 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327084{{!}}Revert "Migrate database access to virtual domains" (T435305)]] * 11:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1008.eqiad.wmnet with OS bookworm * 11:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet * 11:17 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/machinetranslation: apply * 11:13 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/machinetranslation: apply * 11:12 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:12 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:10 kartik@deploy1003: helmfile [staging] START helmfile.d/services/machinetranslation: apply * 11:08 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-cron: apply * 11:08 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/mw-cron: apply * 11:08 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply * 11:08 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply * 11:07 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet * 11:07 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host stat1008.eqiad.wmnet with OS bookworm * 11:06 moritzm: upgrading the new trixie URL downloaders to Squid 7.6 [[phab:T427282|T427282]] * 11:01 gkyziridis@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin2002.codfw.wmnet * 10:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin2002.codfw.wmnet * 10:45 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1169.eqiad.wmnet with OS bookworm * 10:42 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1185.eqiad.wmnet with OS bookworm * 10:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1169.eqiad.wmnet with reason: host reimage * 10:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1185.eqiad.wmnet with reason: host reimage * 10:14 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1169.eqiad.wmnet with reason: host reimage * 10:14 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1185.eqiad.wmnet with reason: host reimage * 10:11 jmm@cumin2003: END (PASS) - Cookbook sre.netbox.restart-reboot (exit_code=0) rolling reboot on A:netbox * 10:06 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1008.eqiad.wmnet with OS bookworm * 10:04 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 10:04 mpostoronca@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321579{{!}}Register the mediawiki.wikimedia_antiabuse.content_policy_score stream (T432848)]] (duration: 08m 53s) * 10:03 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 10:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 10:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 10:00 mpostoronca@deploy1003: mpostoronca: Continuing with deployment * 09:59 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1185.eqiad.wmnet with OS bookworm * 09:59 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1169.eqiad.wmnet with OS bookworm * 09:59 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.convert-disks (exit_code=0) for host ms-be1065 * 09:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1065.eqiad.wmnet with OS trixie * 09:58 mpostoronca@deploy1003: mpostoronca: Backport for [[gerrit:1321579{{!}}Register the mediawiki.wikimedia_antiabuse.content_policy_score stream (T432848)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:55 jmm@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netbox.discovery.wmnet. on all recursors * 09:55 mpostoronca@deploy1003: Started scap sync-world: Backport for [[gerrit:1321579{{!}}Register the mediawiki.wikimedia_antiabuse.content_policy_score stream (T432848)]] * 09:55 jmm@cumin2003: START - Cookbook sre.dns.wipe-cache netbox.discovery.wmnet. on all recursors * 09:52 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw2001.wikimedia.org with OS trixie * 09:51 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 09:51 jmm@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netbox.discovery.wmnet. on all recursors * 09:51 jmm@cumin2003: START - Cookbook sre.dns.wipe-cache netbox.discovery.wmnet. on all recursors * 09:51 jmm@cumin2003: START - Cookbook sre.netbox.restart-reboot rolling reboot on A:netbox * 09:46 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 09:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1056.eqiad.wmnet with OS trixie * 09:44 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 09:39 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 09:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 09:36 topranks: make HE transport circuits from magru live * 09:36 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 09:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb1003.eqiad.wmnet * 09:33 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage * 09:31 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb1003.eqiad.wmnet * 09:30 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 09:28 arnaudb@dns1006: END - running authdns-update * 09:27 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage * 09:27 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb2003.codfw.wmnet * 09:26 arnaudb@dns1006: START - running authdns-update * 09:24 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1056.eqiad.wmnet with reason: host reimage * 09:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb2003.codfw.wmnet * 09:20 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.convert-disks (exit_code=0) for host ms-be1068 * 09:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1068.eqiad.wmnet with OS trixie * 09:20 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 09:19 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "cloudvirt1057 - filippo@cumin1003" * 09:19 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "cloudvirt1057 - filippo@cumin1003" * 09:18 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1056.eqiad.wmnet with reason: host reimage * 09:18 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1057.eqiad.wmnet with OS trixie * 09:18 filippo@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 09:18 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 09:15 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 09:14 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 09:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host irc1003.wikimedia.org * 09:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:13 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1065.eqiad.wmnet with OS trixie * 09:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 09:09 moritzm: installing Postgresql security updates * 09:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host irc1003.wikimedia.org * 09:07 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 09:07 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:07 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1173.eqiad.wmnet with OS bookworm * 09:06 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw2001.wikimedia.org with OS trixie * 09:03 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.convert-disks (exit_code=0) for host ms-be1064 * 09:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1064.eqiad.wmnet with OS trixie * 09:03 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1056.eqiad.wmnet with OS trixie * 09:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1057.eqiad.wmnet with reason: host reimage * 09:01 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 08:58 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 08:56 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1057.eqiad.wmnet with reason: host reimage * 08:54 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1208.eqiad.wmnet with OS bookworm * 08:53 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2902.codfw.wmnet with OS trixie * 08:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 08:50 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1174.eqiad.wmnet with OS bookworm * 08:46 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1172.eqiad.wmnet with OS bookworm * 08:45 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1173.eqiad.wmnet with reason: host reimage * 08:41 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 08:40 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1057.eqiad.wmnet with OS trixie * 08:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1057.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 08:39 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1207.eqiad.wmnet with OS bookworm * 08:38 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2902.codfw.wmnet with reason: host reimage * 08:37 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 08:34 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1068.eqiad.wmnet with OS trixie * 08:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1208.eqiad.wmnet with reason: host reimage * 08:31 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1057.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 08:29 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 08:28 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1222.eqiad.wmnet onto db1276.eqiad.wmnet * 08:28 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1222: Pool db1222.eqiad.wmnet in after cloning * 08:28 fceratto@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2902.codfw.wmnet with reason: host reimage * 08:27 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1055.eqiad.wmnet with OS trixie * 08:27 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 08:26 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1174.eqiad.wmnet with reason: host reimage * 08:25 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 08:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-misc2002.codfw.wmnet * 08:23 topranks: reboot pfw1-codfw firewall pair to upgrade JunOS [[phab:T434865|T434865]] * 08:22 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1172.eqiad.wmnet with reason: host reimage * 08:20 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1064.eqiad.wmnet with OS trixie * 08:20 mvernon@cumin2003: START - Cookbook sre.swift.convert-disks for host ms-be1065 * 08:18 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1207.eqiad.wmnet with reason: host reimage * 08:17 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1174.eqiad.wmnet with reason: host reimage * 08:17 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1173.eqiad.wmnet with reason: host reimage * 08:17 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1172.eqiad.wmnet with reason: host reimage * 08:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host mc-misc2002.codfw.wmnet * 08:15 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1208.eqiad.wmnet with reason: host reimage * 08:15 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1207.eqiad.wmnet with reason: host reimage * 08:14 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db2902.codfw.wmnet with OS trixie * 08:14 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.16 refs [[phab:T430835|T430835]] * 08:13 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db2902.codfw.wmnet * 08:13 fceratto@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host db2902.codfw.wmnet with OS trixie * 08:10 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr[1-2]-codfw with reason: upgrade pfw1a-codfw and pfw1b-codfw pair * 08:09 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1055.eqiad.wmnet with reason: host reimage * 08:07 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on pfw1-codfw with reason: upgrade pfw1a-codfw and pfw1b-codfw pair * 08:03 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1055.eqiad.wmnet with reason: host reimage * 08:02 arnaudb@dns1006: END - running authdns-update * 08:02 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1208.eqiad.wmnet with OS bookworm * 08:02 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1207.eqiad.wmnet with OS bookworm * 08:01 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1174.eqiad.wmnet with OS bookworm * 08:01 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1173.eqiad.wmnet with OS bookworm * 08:01 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1172.eqiad.wmnet with OS bookworm * 07:59 arnaudb@dns1006: START - running authdns-update * 07:58 arnaudb@dns1006: START - running authdns-update * 07:48 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1055.eqiad.wmnet with OS trixie * 07:45 moritzm: extend the disk of ldap-rw2001 by 80G [[phab:T331699|T331699]] * 07:42 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1222: Pool db1222.eqiad.wmnet in after cloning * 07:36 mvernon@cumin2003: START - Cookbook sre.swift.convert-disks for host ms-be1068 * 07:35 mvernon@cumin2003: START - Cookbook sre.swift.convert-disks for host ms-be1064 * 07:17 moritzm: installing imagemagick security updates * 07:14 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1277: Pool back * 07:14 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon1003.wikimedia.org * 07:07 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon1003.wikimedia.org * 07:03 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1280: Pool back * 07:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon2002.wikimedia.org * 06:55 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon2002.wikimedia.org * 06:54 moritzm: installing php8.2 security updates * 06:51 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1284: Pool back * 06:49 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1222: Depool db1222.eqiad.wmnet to then clone it to db1276.eqiad.wmnet - marostegui@cumin1003 * 06:49 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1222: Depool db1222.eqiad.wmnet to then clone it to db1276.eqiad.wmnet - marostegui@cumin1003 * 06:49 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1222.eqiad.wmnet onto db1276.eqiad.wmnet * 06:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd1005.eqiad.wmnet * 06:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd1005.eqiad.wmnet * 06:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd1004.eqiad.wmnet * 06:36 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2209: db2209 repool * 06:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd1004.eqiad.wmnet * 06:32 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast3007.wikimedia.org * 06:29 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1277: Pool back * 06:28 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1277 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96190 and previous config saved to /var/cache/conftool/dbconfig/20260819-062815-marostegui.json * 06:26 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast3007.wikimedia.org * 06:22 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul1001.eqiad.wmnet * 06:18 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul1001.eqiad.wmnet * 06:18 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul1003.eqiad.wmnet * 06:18 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1280: Pool back * 06:17 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1284 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96186 and previous config saved to /var/cache/conftool/dbconfig/20260819-061743-marostegui.json * 06:14 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul1003.eqiad.wmnet * 06:14 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul1002.eqiad.wmnet * 06:10 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul1002.eqiad.wmnet * 06:10 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2048.codfw.wmnet * 06:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2048.codfw.wmnet * 06:06 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1284: Pool back * 06:06 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1284 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96184 and previous config saved to /var/cache/conftool/dbconfig/20260819-060621-marostegui.json * 06:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2048.codfw.wmnet * 05:59 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2048.codfw.wmnet * 05:51 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2209: db2209 repool * 03:16 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1186.eqiad.wmnet with OS bookworm * 02:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1186.eqiad.wmnet with reason: host reimage * 02:46 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1186.eqiad.wmnet with reason: host reimage * 02:46 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2207 [[phab:T435270|T435270]]', diff saved to https://phabricator.wikimedia.org/P96181 and previous config saved to /var/cache/conftool/dbconfig/20260819-024627-marostegui.json * 02:44 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2204 to s2 primary [[phab:T435270|T435270]]', diff saved to https://phabricator.wikimedia.org/P96180 and previous config saved to /var/cache/conftool/dbconfig/20260819-024403-marostegui.json * 02:43 marostegui: Starting s2 codfw failover from db2207 to db2204 - [[phab:T435270|T435270]] * 02:39 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2204 with weight 0 [[phab:T435270|T435270]]', diff saved to https://phabricator.wikimedia.org/P96179 and previous config saved to /var/cache/conftool/dbconfig/20260819-023951-marostegui.json * 02:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s2 [[phab:T435270|T435270]] * 02:32 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1186.eqiad.wmnet with OS bookworm * 02:29 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-worker1186.eqiad.wmnet with OS bookworm * 02:18 denisse@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2207: Depooling replica * 02:18 denisse@cumin1003: START - Cookbook sre.mysql.depool depool db2207: Depooling replica * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 48s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-18 == * 23:55 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1264845{{!}}Remove unused/redundant wgMFNoindexPages=true setting (T255458)]] (duration: 09m 42s) * 23:51 krinkle@deploy1003: krinkle: Continuing with deployment * 23:48 krinkle@deploy1003: krinkle: Backport for [[gerrit:1264845{{!}}Remove unused/redundant wgMFNoindexPages=true setting (T255458)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:45 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1264845{{!}}Remove unused/redundant wgMFNoindexPages=true setting (T255458)]] * 23:38 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326944{{!}}Retire filebackend lock manager in favour of the default one (T366938)]] (duration: 08m 55s) * 23:34 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 23:31 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326944{{!}}Retire filebackend lock manager in favour of the default one (T366938)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:29 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326944{{!}}Retire filebackend lock manager in favour of the default one (T366938)]] * 23:27 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1170.eqiad.wmnet with OS bookworm * 23:21 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1205.eqiad.wmnet with OS bookworm * 23:20 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1171.eqiad.wmnet with OS bookworm * 23:15 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1206.eqiad.wmnet with OS bookworm * 23:05 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1170.eqiad.wmnet with reason: host reimage * 23:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1205.eqiad.wmnet with reason: host reimage * 22:57 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1171.eqiad.wmnet with reason: host reimage * 22:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1206.eqiad.wmnet with reason: host reimage * 22:53 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1205.eqiad.wmnet with reason: host reimage * 22:51 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1171.eqiad.wmnet with reason: host reimage * 22:51 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1170.eqiad.wmnet with reason: host reimage * 22:50 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1206.eqiad.wmnet with reason: host reimage * 22:36 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1206.eqiad.wmnet with OS bookworm * 22:35 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1205.eqiad.wmnet with OS bookworm * 22:35 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1186.eqiad.wmnet with OS bookworm * 22:35 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1171.eqiad.wmnet with OS bookworm * 22:35 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1170.eqiad.wmnet with OS bookworm * 22:33 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-worker1194.eqiad.wmnet with OS bookworm * 22:22 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326923{{!}}Enable redis lock manager on s6 (T366938)]] (duration: 11m 52s) * 22:18 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 22:13 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326923{{!}}Enable redis lock manager on s6 (T366938)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:10 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326923{{!}}Enable redis lock manager on s6 (T366938)]] * 22:04 sbassett: Deployed security fix for [[phab:T435234|T435234]] (wmf.16) * 21:54 sbassett: Deployed security fix for [[phab:T435234|T435234]] (wmf.15) * 21:38 caro@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326925{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326926{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326929{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]], [[gerrit:1326928{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]] (duration: 0 * 21:34 caro@deploy1003: caro: Continuing with deployment * 21:33 caro@deploy1003: caro: Backport for [[gerrit:1326925{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326926{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326929{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]], [[gerrit:1326928{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]] synced to the testservers (see h * 21:31 caro@deploy1003: Started scap sync-world: Backport for [[gerrit:1326925{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326926{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326929{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]], [[gerrit:1326928{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]] * 21:24 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1204.eqiad.wmnet with reason: 1204 datanode repair [[phab:T434494|T434494]] * 21:02 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326896{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]], [[gerrit:1326897{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]] (duration: 13m 56s) * 20:58 krinkle@deploy1003: krinkle: Continuing with deployment * 20:50 krinkle@deploy1003: krinkle: Backport for [[gerrit:1326896{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]], [[gerrit:1326897{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:49 ryankemper: `an-launcher1003` terminated process group `666809` (`rest_backfill_phase1.sh`) ~20 mins ago with `sudo kill -TERM -- -666809` after its local spark driver (`--driver-memory 64g`) repeatedly exhausted memory on the 32 GB VM and caused SSH to intermittently flap; host recovered to 27 GB available memory * 20:48 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1326896{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]], [[gerrit:1326897{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]] * 20:35 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 20:33 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 20:31 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 20:31 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326870{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]], [[gerrit:1326871{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]] (duration: 07m 35s) * 20:28 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 20:26 kemayo@deploy1003: kemayo: Continuing with deployment * 20:26 ryankemper: `an-launcher1003` confirmed the host is flapping because of memory thrash. chasing down the source of the thrash * 20:25 kemayo@deploy1003: kemayo: Backport for [[gerrit:1326870{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]], [[gerrit:1326871{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:23 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1326870{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]], [[gerrit:1326871{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]] * 20:23 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 20:20 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 20:14 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1168.eqiad.wmnet with OS bookworm * 20:08 zabe: zabe@deploy1003:~$ mwscript extensions/WikimediaMaintenance/maintenance/fixFileRevisionArchiveNameDrift.php enwiki # [[phab:T428406|T428406]] * 20:08 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1204.eqiad.wmnet with OS bookworm * 20:05 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1167.eqiad.wmnet with OS bookworm * 20:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1166.eqiad.wmnet with OS bookworm * 19:54 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1203.eqiad.wmnet with OS bookworm * 19:53 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326907{{!}}Revert "Disable redis lock manager on testwiki"]] (duration: 11m 05s) * 19:50 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1168.eqiad.wmnet with reason: host reimage * 19:47 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1204.eqiad.wmnet with reason: host reimage * 19:46 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 19:44 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326907{{!}}Revert "Disable redis lock manager on testwiki"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:42 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326907{{!}}Revert "Disable redis lock manager on testwiki"]] * 19:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1167.eqiad.wmnet with reason: host reimage * 19:37 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1166.eqiad.wmnet with reason: host reimage * 19:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1203.eqiad.wmnet with reason: host reimage * 19:32 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1167.eqiad.wmnet with reason: host reimage * 19:32 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1168.eqiad.wmnet with reason: host reimage * 19:32 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1166.eqiad.wmnet with reason: host reimage * 19:31 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1204.eqiad.wmnet with reason: host reimage * 19:31 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1203.eqiad.wmnet with reason: host reimage * 19:26 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324751{{!}}InitialiseSettings: Enable 2FA enforcement on more private wikis (T428103)]], [[gerrit:1326875{{!}}Add banner notifying of upcoming 2FA enforcement (T420792)]] (duration: 31m 46s) * 19:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1204.eqiad.wmnet with OS bookworm * 19:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1203.eqiad.wmnet with OS bookworm * 19:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1168.eqiad.wmnet with OS bookworm * 19:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1167.eqiad.wmnet with OS bookworm * 19:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1166.eqiad.wmnet with OS bookworm * 19:15 denisse: rebooting kafkamon2003.codfw.wmnet - [[phab:T435162|T435162]] * 19:14 denisse: rebooting kafkamon1003.eqiad.wmnet [[phab:T435162|T435162]] * 19:13 reedy@deploy1003: reedy: Continuing with deployment * 19:12 reedy@deploy1003: reedy: Backport for [[gerrit:1324751{{!}}InitialiseSettings: Enable 2FA enforcement on more private wikis (T428103)]], [[gerrit:1326875{{!}}Add banner notifying of upcoming 2FA enforcement (T420792)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:54 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324751{{!}}InitialiseSettings: Enable 2FA enforcement on more private wikis (T428103)]], [[gerrit:1326875{{!}}Add banner notifying of upcoming 2FA enforcement (T420792)]] * 18:50 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 18:44 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.16 refs [[phab:T430835|T430835]] * 18:34 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 18:31 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 18:22 aklapper@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326852{{!}}CategoryTree: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]], [[gerrit:1326853{{!}}CategoryViewer: Allow null $html in the CategoryViewerGenerateLink hook (T435161)]], [[gerrit:1326865{{!}}Flow: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]] (duration: 09m 57s) * 18:18 aklapper@deploy1003: jforrester, aklapper: Continuing with deployment * 18:17 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 18:14 aklapper@deploy1003: jforrester, aklapper: Backport for [[gerrit:1326852{{!}}CategoryTree: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]], [[gerrit:1326853{{!}}CategoryViewer: Allow null $html in the CategoryViewerGenerateLink hook (T435161)]], [[gerrit:1326865{{!}}Flow: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki * 18:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1156.eqiad.wmnet with OS bookworm * 18:12 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 18:12 aklapper@deploy1003: Started scap sync-world: Backport for [[gerrit:1326852{{!}}CategoryTree: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]], [[gerrit:1326853{{!}}CategoryViewer: Allow null $html in the CategoryViewerGenerateLink hook (T435161)]], [[gerrit:1326865{{!}}Flow: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]] * 18:11 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 18:08 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1146.eqiad.wmnet with OS bookworm * 18:07 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1177.eqiad.wmnet with OS bookworm * 18:00 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326380{{!}}Introduce main lock manager service (T366938 T427999)]] (duration: 11m 25s) * 17:58 ladsgroup@deploy1003: ladsgroup: Rolling back deployment * 17:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1202.eqiad.wmnet with OS bookworm * 17:55 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1201.eqiad.wmnet with OS bookworm * 17:50 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326380{{!}}Introduce main lock manager service (T366938 T427999)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:48 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326380{{!}}Introduce main lock manager service (T366938 T427999)]] * 17:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1156.eqiad.wmnet with reason: host reimage * 17:46 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1146.eqiad.wmnet with reason: host reimage * 17:45 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-esams and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 17:42 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 17:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1177.eqiad.wmnet with reason: host reimage * 17:38 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1202.eqiad.wmnet with reason: host reimage * 17:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1201.eqiad.wmnet with reason: host reimage * 17:30 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1156.eqiad.wmnet with reason: host reimage * 17:29 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1177.eqiad.wmnet with reason: host reimage * 17:28 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1146.eqiad.wmnet with reason: host reimage * 17:28 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1202.eqiad.wmnet with reason: host reimage * 17:27 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1201.eqiad.wmnet with reason: host reimage * 17:25 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 17:21 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 17:14 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1202.eqiad.wmnet with OS bookworm * 17:14 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1201.eqiad.wmnet with OS bookworm * 17:14 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1177.eqiad.wmnet with OS bookworm * 17:14 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1156.eqiad.wmnet with OS bookworm * 17:14 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1146.eqiad.wmnet with OS bookworm * 17:09 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326885{{!}}w/deployment-info.php: Handle new file format (T434726)]] (duration: 07m 15s) * 17:08 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 17:05 dancy@deploy1003: dancy: Continuing with deployment * 17:04 dancy@deploy1003: dancy: Backport for [[gerrit:1326885{{!}}w/deployment-info.php: Handle new file format (T434726)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:02 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1326885{{!}}w/deployment-info.php: Handle new file format (T434726)]] * 16:50 dancy@deploy1003: Finished scap sync-world: Testing [[phab:T434726|T434726]] (duration: 06m 40s) * 16:43 dancy@deploy1003: Started scap sync-world: Testing [[phab:T434726|T434726]] * 16:43 dancy@deploy1003: Installation of scap version "4.282.0" completed for 3 hosts * 16:41 dancy@deploy1003: Installing scap version "4.282.0" for 3 host(s) * 16:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1145.eqiad.wmnet with OS bookworm * 16:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1200.eqiad.wmnet with OS bookworm * 16:32 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1199.eqiad.wmnet with OS bookworm * 16:18 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1145.eqiad.wmnet with reason: host reimage * 16:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1200.eqiad.wmnet with reason: host reimage * 16:09 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1199.eqiad.wmnet with reason: host reimage * 16:05 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-esams and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 16:05 cjd91: sudo -i cookbook sre.cdn.roll-upgrade-ats --query 'A:cp-esams' --task-id [[phab:T434478|T434478]] --reason '9.2.15 upgrade' * 16:03 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1200.eqiad.wmnet with reason: host reimage * 16:02 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1145.eqiad.wmnet with reason: host reimage * 16:02 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1199.eqiad.wmnet with reason: host reimage * 15:50 moritzm: installing zip security updates * 15:48 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1200.eqiad.wmnet with OS bookworm * 15:47 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1199.eqiad.wmnet with OS bookworm * 15:47 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1145.eqiad.wmnet with OS bookworm * 15:41 topranks: bounce PIC 0/0 on cr1-magru to set port to 40G * 15:37 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1236.eqiad.wmnet with OS bookworm * 15:29 aikochou@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'ores-legacy' for release 'main' . * 15:26 aikochou@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'ores-legacy' for release 'main' . * 15:20 aikochou@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'ores-legacy' for release 'main' . * 15:14 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2002.codfw.wmnet * 15:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1236.eqiad.wmnet with reason: host reimage * 15:12 moritzm: failover ganeti master in codfw to ganeti2047 * 15:09 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1236.eqiad.wmnet with reason: host reimage * 15:09 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2044.codfw.wmnet * 15:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2002.codfw.wmnet * 15:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2044.codfw.wmnet * 15:04 brennen@deploy1003: Finished deploy [phabricator/deployment@6b9b6ff]: deploy phab1004 for [[phab:T435213|T435213]] (duration: 01m 01s) * 15:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2044.codfw.wmnet * 15:03 brennen@deploy1003: Started deploy [phabricator/deployment@6b9b6ff]: deploy phab1004 for [[phab:T435213|T435213]] * 15:03 brennen@deploy1003: Finished deploy [phabricator/deployment@6b9b6ff]: deploy phab2003 for [[phab:T435213|T435213]] (duration: 00m 57s) * 15:02 brennen@deploy1003: Started deploy [phabricator/deployment@6b9b6ff]: deploy phab2003 for [[phab:T435213|T435213]] * 14:57 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply * 14:57 arnaudb@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on phab2003.codfw.wmnet,phab[1004-1006].eqiad.wmnet with reason: maintenance * 14:56 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2044.codfw.wmnet * 14:55 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply * 14:53 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1236.eqiad.wmnet with OS bookworm * 14:51 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2043.codfw.wmnet * 14:51 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2043.codfw.wmnet * 14:45 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2043.codfw.wmnet * 14:33 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2043.codfw.wmnet * 14:22 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2042.codfw.wmnet * 14:22 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2042.codfw.wmnet * 14:21 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db2902.codfw.wmnet with OS trixie * 14:16 elukey: uploaded spicerack_13.2.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia * 14:16 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2042.codfw.wmnet * 14:04 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2042.codfw.wmnet * 14:04 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2041.codfw.wmnet * 14:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2041.codfw.wmnet * 13:59 phuedx@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: apply * 13:59 phuedx@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-main: apply * 13:59 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326838{{!}}Enable Suggested Investigations on hewiki (T435146)]] (duration: 11m 50s) * 13:58 phuedx@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: apply * 13:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2041.codfw.wmnet * 13:57 phuedx@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-main: apply * 13:57 phuedx@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-main: apply * 13:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard1003.eqiad.wmnet * 13:57 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-main: apply * 13:55 phuedx@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-logging-external: apply * 13:55 phuedx@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-logging-external: apply * 13:54 phuedx@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-logging-external: apply * 13:54 stran@deploy1003: stran: Continuing with deployment * 13:54 phuedx@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-logging-external: apply * 13:54 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-logging-external: apply * 13:54 phuedx@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-logging-external: apply * 13:53 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-logging-external: apply * 13:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard1003.eqiad.wmnet * 13:53 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2041.codfw.wmnet * 13:51 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db2902.codfw.wmnet - fceratto@cumin1003" * 13:51 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db2902.codfw.wmnet - fceratto@cumin1003" * 13:51 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2026.codfw.wmnet * 13:51 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard2003.codfw.wmnet * 13:50 phuedx@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: apply * 13:50 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-eqiad * 13:50 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp1001.eqiad.wmnet * 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp1001.eqiad.wmnet * 13:50 phuedx@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: apply * 13:49 phuedx@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: apply * 13:49 stran@deploy1003: stran: Backport for [[gerrit:1326838{{!}}Enable Suggested Investigations on hewiki (T435146)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:48 phuedx@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: apply * 13:48 phuedx@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics: apply * 13:47 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard2003.codfw.wmnet * 13:47 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics: apply * 13:47 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1326838{{!}}Enable Suggested Investigations on hewiki (T435146)]] * 13:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp1001.eqiad.wmnet * 13:44 cdanis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 13:43 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp1001.eqiad.wmnet * 13:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1315-1327].eqiad.wmnet * 13:43 cdanis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 13:43 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1315-1327].eqiad.wmnet * 13:40 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki1001.eqiad.wmnet * 13:35 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1315-1327].eqiad.wmnet * 13:34 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host rpki1001.eqiad.wmnet * 13:27 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1315-1327].eqiad.wmnet * 13:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1302-1314].eqiad.wmnet * 13:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1302-1314].eqiad.wmnet * 13:21 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326775{{!}}SI: Instrument case update on first edit (T435048)]] (duration: 07m 12s) * 13:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1302-1314].eqiad.wmnet * 13:17 stran@deploy1003: stran: Continuing with deployment * 13:16 stran@deploy1003: stran: Backport for [[gerrit:1326775{{!}}SI: Instrument case update on first edit (T435048)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:14 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1326775{{!}}SI: Instrument case update on first edit (T435048)]] * 13:13 moritzm: installing util-linux security updates * 13:11 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1302-1314].eqiad.wmnet * 13:10 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326232{{!}}prv: Enable parsoid rendering for 5 wikis (T435115)]] (duration: 08m 17s) * 13:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1288-1289,1291-1301].eqiad.wmnet * 13:10 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply * 13:10 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1288-1289,1291-1301].eqiad.wmnet * 13:10 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply * 13:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki2003.codfw.wmnet * 13:06 jgiannelos@deploy1003: jgiannelos: Continuing with deployment * 13:05 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host rpki2003.codfw.wmnet * 13:04 jgiannelos@deploy1003: jgiannelos: Backport for [[gerrit:1326232{{!}}prv: Enable parsoid rendering for 5 wikis (T435115)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1288-1289,1291-1301].eqiad.wmnet * 13:02 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1326232{{!}}prv: Enable parsoid rendering for 5 wikis (T435115)]] * 12:53 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1288-1289,1291-1301].eqiad.wmnet * 12:53 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1273,1275-1287].eqiad.wmnet * 12:53 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1273,1275-1287].eqiad.wmnet * 12:52 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt1002.wikimedia.org * 12:51 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db2902.codfw.wmnet on all recursors * 12:51 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db2902.codfw.wmnet on all recursors * 12:51 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:51 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db2902.codfw.wmnet - fceratto@cumin1003" * 12:51 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db2902.codfw.wmnet - fceratto@cumin1003" * 12:46 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt1002.wikimedia.org * 12:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt2002.wikimedia.org * 12:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1273,1275-1287].eqiad.wmnet * 12:42 dhinus: repooled clouddb1032 that was currently <nowiki>{</nowiki>"weight": 0, "pooled": "inactive"<nowiki>}</nowiki> for both s4 and s6 * 12:41 dhinus: also depooled clouddb1017 (forgot it in the previous list) * 12:41 fnegri@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet * 12:40 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 12:40 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db2902.codfw.wmnet * 12:40 fnegri@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032.eqiad.wmnet * 12:40 fnegri@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032 * 12:39 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1017.eqiad.wmnet * 12:39 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt2002.wikimedia.org * 12:38 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1020.eqiad.wmnet * 12:38 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1018.eqiad.wmnet * 12:38 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet * 12:37 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1014.eqiad.wmnet * 12:37 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1013.eqiad.wmnet * 12:37 dhinus: depool again clouddb10[13,14,16,18,20] that were repooled by the cookbook sre.mysql.multiinstance_reboot * 12:36 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1273,1275-1287].eqiad.wmnet * 12:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1248-1261].eqiad.wmnet * 12:36 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1248-1261].eqiad.wmnet * 12:35 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2026.codfw.wmnet * 12:34 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 12:34 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 12:34 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 12:34 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 12:29 lucaswerkmeister-wmde@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 12:28 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2026.codfw.wmnet * 12:28 lucaswerkmeister-wmde@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 12:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1248-1261].eqiad.wmnet * 12:27 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.addnode (exit_code=0) for new host ganeti2046.codfw.wmnet to cluster codfw and group A * 12:26 moritzm: readded ganeti2046 to the codfw cluster following firmware update and reimage [[phab:T434681|T434681]] * 12:23 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply * 12:23 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply * 12:23 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply * 12:22 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply * 12:22 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply * 12:22 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply * 12:21 jmm@cumin2003: START - Cookbook sre.ganeti.addnode for new host ganeti2046.codfw.wmnet to cluster codfw and group A * 12:21 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1169.eqiad.wmnet onto db1283.eqiad.wmnet * 12:21 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1169: Pool db1169.eqiad.wmnet in after cloning * 12:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1248-1261].eqiad.wmnet * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1149-1153,1158,1240-1247].eqiad.wmnet * 12:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1149-1153,1158,1240-1247].eqiad.wmnet * 12:11 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1149-1153,1158,1240-1247].eqiad.wmnet * 12:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1168.eqiad.wmnet onto db1282.eqiad.wmnet * 12:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1168: Pool db1168.eqiad.wmnet in after cloning * 12:06 jmm@cumin2003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-test-eqiad * 12:05 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2026.codfw.wmnet * 12:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1149-1153,1158,1240-1247].eqiad.wmnet * 12:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1128-1134,1142-1148].eqiad.wmnet * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2046.codfw.wmnet * 12:01 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1128-1134,1142-1148].eqiad.wmnet * 11:57 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2025.codfw.wmnet * 11:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2025.codfw.wmnet * 11:54 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2046.codfw.wmnet * 11:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1128-1134,1142-1148].eqiad.wmnet * 11:50 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2025.codfw.wmnet * 11:46 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1128-1134,1142-1148].eqiad.wmnet * 11:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1114-1127].eqiad.wmnet * 11:45 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1114-1127].eqiad.wmnet * 11:39 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db2901.codfw.wmnet * 11:39 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db2901.codfw.wmnet * 11:36 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2025.codfw.wmnet * 11:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1114-1127].eqiad.wmnet * 11:35 fceratto@cumin1003: END (ERROR) - Cookbook sre.ganeti.makevm (exit_code=93) for new host db1901.eqiad.wmnet * 11:35 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1169: Pool db1169.eqiad.wmnet in after cloning * 11:35 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 11:30 jmm@cumin2003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-test-eqiad * 11:27 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1114-1127].eqiad.wmnet * 11:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1076-1081,1084-1087,1093-1095,1113].eqiad.wmnet * 11:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1076-1081,1084-1087,1093-1095,1113].eqiad.wmnet * 11:25 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1168: Pool db1168.eqiad.wmnet in after cloning * 11:24 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1194.eqiad.wmnet with OS bookworm * 11:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1181.eqiad.wmnet with OS bookworm * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2050.codfw.wmnet * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2050.codfw.wmnet * 11:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1076-1081,1084-1087,1093-1095,1113].eqiad.wmnet * 11:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2050.codfw.wmnet * 11:14 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2004.codfw.wmnet * 11:10 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2050.codfw.wmnet * 11:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1076-1081,1084-1087,1093-1095,1113].eqiad.wmnet * 11:09 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1286: Pool back * 11:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1045-1050,1056-1057,1064-1066,1073-1075].eqiad.wmnet * 11:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1045-1050,1056-1057,1064-1066,1073-1075].eqiad.wmnet * 11:08 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2004.codfw.wmnet * 11:07 moritzm: installing PHP 8.4 security updates * 11:06 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on an-worker1194.eqiad.wmnet with reason: host reimage * 11:06 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1194.eqiad.wmnet with reason: host reimage * 11:05 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2049.codfw.wmnet * 11:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2049.codfw.wmnet * 11:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid1003.eqiad.wmnet * 11:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1045-1050,1056-1057,1064-1066,1073-1075].eqiad.wmnet * 10:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1181.eqiad.wmnet with reason: host reimage * 10:59 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2049.codfw.wmnet * 10:59 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid1003.eqiad.wmnet * 10:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid2003.codfw.wmnet * 10:54 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1181.eqiad.wmnet with reason: host reimage * 10:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid2003.codfw.wmnet * 10:51 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1901.eqiad.wmnet on all recursors * 10:51 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1901.eqiad.wmnet on all recursors * 10:51 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:51 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:51 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:50 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2049.codfw.wmnet * 10:50 blake@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:50 blake@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:49 blake@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:48 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1045-1050,1056-1057,1064-1066,1073-1075].eqiad.wmnet * 10:48 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1044].eqiad.wmnet * 10:48 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1044].eqiad.wmnet * 10:47 blake@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:46 blake@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:46 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2047.codfw.wmnet * 10:46 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:46 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2047.codfw.wmnet * 10:46 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1901.eqiad.wmnet * 10:46 blake@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:44 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1168: Depool db1168.eqiad.wmnet to then clone it to db1282.eqiad.wmnet - marostegui@cumin1003 * 10:44 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1168: Depool db1168.eqiad.wmnet to then clone it to db1282.eqiad.wmnet - marostegui@cumin1003 * 10:44 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1168.eqiad.wmnet onto db1282.eqiad.wmnet * 10:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2047.codfw.wmnet * 10:40 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1044].eqiad.wmnet * 10:37 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2047.codfw.wmnet * 10:37 fceratto@cumin1003: END (ERROR) - Cookbook sre.ganeti.makevm (exit_code=93) for new host db1901.eqiad.wmnet * 10:36 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 10:34 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:33 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2032.codfw.wmnet * 10:33 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2032.codfw.wmnet * 10:32 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1044].eqiad.wmnet * 10:32 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-eqiad * 10:27 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2032.codfw.wmnet * 10:24 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1286: Pool back * 10:24 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1286 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96163 and previous config saved to /var/cache/conftool/dbconfig/20260818-102431-marostegui.json * 10:22 moritzm: installing Django security updates * 10:20 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul2002.codfw.wmnet * 10:20 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul2001.codfw.wmnet * 10:17 blake@deploy1003: Finished scap sync-world: no-build deployment for [[phab:T417800|T417800]] (duration: 04m 40s) * 10:16 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul2002.codfw.wmnet * 10:16 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul2001.codfw.wmnet * 10:15 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2032.codfw.wmnet * 10:14 blake@deploy1003: Started scap sync-world: no-build deployment for [[phab:T417800|T417800]] * 10:12 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul2003.codfw.wmnet * 10:12 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy2001.codfw.wmnet * 10:12 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy3001.esams.wmnet * 10:12 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy1001.eqiad.wmnet * 10:08 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2031.codfw.wmnet * 10:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2031.codfw.wmnet * 10:08 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul2003.codfw.wmnet * 10:08 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy2001.codfw.wmnet * 10:08 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy3001.esams.wmnet * 10:08 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy1001.eqiad.wmnet * 10:07 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy1002.eqiad.wmnet * 10:05 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy2002.codfw.wmnet * 10:05 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy3002.esams.wmnet * 10:04 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy1002.eqiad.wmnet * 10:03 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy4003.ulsfo.wmnet * 10:03 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy5003.eqsin.wmnet * 10:02 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2031.codfw.wmnet * 10:02 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1169: Depool db1169.eqiad.wmnet to then clone it to db1283.eqiad.wmnet - marostegui@cumin1003 * 10:01 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy2002.codfw.wmnet * 10:01 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy4004.ulsfo.wmnet * 10:01 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy3002.esams.wmnet * 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1169: Depool db1169.eqiad.wmnet to then clone it to db1283.eqiad.wmnet - marostegui@cumin1003 * 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1169.eqiad.wmnet onto db1283.eqiad.wmnet * 10:01 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy5004.eqsin.wmnet * 09:59 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy4003.ulsfo.wmnet * 09:59 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy4004.ulsfo.wmnet * 09:59 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy5003.eqsin.wmnet * 09:59 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy7001.magru.wmnet * 09:59 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy5004.eqsin.wmnet * 09:58 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy7002.magru.wmnet * 09:57 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2031.codfw.wmnet * 09:57 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy6001.drmrs.wmnet * 09:57 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy6002.drmrs.wmnet * 09:54 filippo@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for 10 hosts * 09:54 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast4006.wikimedia.org * 09:53 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy6001.drmrs.wmnet * 09:53 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy6002.drmrs.wmnet * 09:53 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host releases1003.eqiad.wmnet * 09:53 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy7001.magru.wmnet * 09:53 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host people2004.codfw.wmnet * 09:52 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host people1005.eqiad.wmnet * 09:52 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy7002.magru.wmnet * 09:50 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host releases2003.codfw.wmnet * 09:49 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host releases2003.codfw.wmnet * 09:49 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host releases1003.eqiad.wmnet * 09:49 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host people2004.codfw.wmnet * 09:48 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host people1005.eqiad.wmnet * 09:46 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2040.codfw.wmnet * 09:46 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2040.codfw.wmnet * 09:43 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1901.eqiad.wmnet on all recursors * 09:42 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1901.eqiad.wmnet on all recursors * 09:42 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:42 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:42 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2040.codfw.wmnet * 09:40 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp2005.wikimedia.org * 09:36 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp2005.wikimedia.org * 09:31 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:31 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1901.eqiad.wmnet * 09:31 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1901.eqiad.wmnet * 09:31 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:31 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1901.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 09:31 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1901.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 09:29 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2040.codfw.wmnet * 09:29 slyngshede@dns1004: END - running authdns-update * 09:28 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2039.codfw.wmnet * 09:28 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2039.codfw.wmnet * 09:27 slyngshede@dns1004: START - running authdns-update * 09:26 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test2005.wikimedia.org * 09:22 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2039.codfw.wmnet * 09:22 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test2005.wikimedia.org * 09:22 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp1005.wikimedia.org * 09:21 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:19 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2039.codfw.wmnet * 09:19 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2038.codfw.wmnet * 09:18 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1287: Pool back * 09:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2038.codfw.wmnet * 09:18 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp1005.wikimedia.org * 09:18 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test1005.wikimedia.org * 09:17 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1901.eqiad.wmnet * 09:15 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db1901.eqiad.wmnet * 09:15 fceratto@cumin1003: END (ERROR) - Cookbook sre.dns.netbox (exit_code=97) * 09:14 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test1005.wikimedia.org * 09:13 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2038.codfw.wmnet * 09:13 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:13 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1901.eqiad.wmnet * 09:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast4006.wikimedia.org * 09:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1288: Pool back * 09:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast5005.wikimedia.org * 09:03 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2038.codfw.wmnet * 08:57 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2037.codfw.wmnet * 08:57 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast5005.wikimedia.org * 08:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2037.codfw.wmnet * 08:56 filippo@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 10 hosts * 08:52 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2037.codfw.wmnet * 08:51 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1289: Pool back * 08:35 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudidp2001-dev.codfw.wmnet * 08:34 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1002-dev.eqiad.wmnet * 08:33 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1287: Pool back * 08:33 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1287 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96150 and previous config saved to /var/cache/conftool/dbconfig/20260818-083311-marostegui.json * 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1001-dev.eqiad.wmnet * 08:31 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudidp2001-dev.codfw.wmnet * 08:30 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1002-dev.eqiad.wmnet * 08:30 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2037.codfw.wmnet * 08:29 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1001-dev.eqiad.wmnet * 08:28 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2036.codfw.wmnet * 08:28 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2036.codfw.wmnet * 08:25 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1288: Pool back * 08:23 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 08:23 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1181.eqiad.wmnet with OS bookworm * 08:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2036.codfw.wmnet * 08:22 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1288 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96148 and previous config saved to /var/cache/conftool/dbconfig/20260818-082234-marostegui.json * 08:20 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2036.codfw.wmnet * 08:18 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2035.codfw.wmnet * 08:18 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudvirt1057.eqiad.wmnet with OS trixie * 08:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2035.codfw.wmnet * 08:13 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2035.codfw.wmnet * 08:06 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2035.codfw.wmnet * 08:05 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1289: Pool back * 08:05 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1289 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96145 and previous config saved to /var/cache/conftool/dbconfig/20260818-080531-marostegui.json * 07:54 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 07:51 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326439{{!}}Fix "mathjax_ignore" handling around forcemathmode attribute (T434686)]] (duration: 13m 15s) * 07:48 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 07:47 krinkle@deploy1003: krinkle: Continuing with deployment * 07:44 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ganeti2046.codfw.wmnet with OS bookworm * 07:40 krinkle@deploy1003: krinkle: Backport for [[gerrit:1326439{{!}}Fix "mathjax_ignore" handling around forcemathmode attribute (T434686)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:39 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 07:38 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1326439{{!}}Fix "mathjax_ignore" handling around forcemathmode attribute (T434686)]] * 07:34 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 07:31 samwilson@deploy1003: Finished scap sync-world: Backport for [[gerrit:701016{{!}}InitialiseSettings and -labs: Remove redundant feature flag $wgWikisourceEnableOcr (T285311)]] (duration: 07m 47s) * 07:30 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1048.eqiad.wmnet * 07:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1048.eqiad.wmnet * 07:29 XioNoX: add gnmic 0.47.0 to bookworm and trixie reprepro * 07:28 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ganeti2046.codfw.wmnet with reason: host reimage * 07:27 samwilson@deploy1003: samwilson: Continuing with deployment * 07:25 samwilson@deploy1003: samwilson: Backport for [[gerrit:701016{{!}}InitialiseSettings and -labs: Remove redundant feature flag $wgWikisourceEnableOcr (T285311)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:25 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 07:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1048.eqiad.wmnet * 07:24 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ganeti2046.codfw.wmnet with reason: host reimage * 07:23 samwilson@deploy1003: Started scap sync-world: Backport for [[gerrit:701016{{!}}InitialiseSettings and -labs: Remove redundant feature flag $wgWikisourceEnableOcr (T285311)]] * 07:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1198.eqiad.wmnet with OS bookworm * 07:18 samwilson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326450{{!}}InitialiseSettings.php: Enable Bulk OCR on pawikisource (T434648)]] (duration: 12m 17s) * 07:17 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 07:11 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1048.eqiad.wmnet * 07:11 samwilson@deploy1003: samwilson: Continuing with deployment * 07:11 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ganeti2046.codfw.wmnet with OS bookworm * 07:10 samwilson@deploy1003: samwilson: Backport for [[gerrit:1326450{{!}}InitialiseSettings.php: Enable Bulk OCR on pawikisource (T434648)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1057.eqiad.wmnet with OS trixie * 07:05 samwilson@deploy1003: Started scap sync-world: Backport for [[gerrit:1326450{{!}}InitialiseSettings.php: Enable Bulk OCR on pawikisource (T434648)]] * 07:02 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1198.eqiad.wmnet with reason: host reimage * 07:01 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1056.eqiad.wmnet with OS trixie * 07:01 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1056.eqiad.wmnet with OS trixie * 07:00 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1056.eqiad.wmnet with OS trixie * 07:00 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1056.eqiad.wmnet with OS trixie * 06:59 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudvirt1055.eqiad.wmnet with OS trixie * 06:58 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1198.eqiad.wmnet with reason: host reimage * 06:53 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1055.eqiad.wmnet with OS trixie * 06:52 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudvirt1054.eqiad.wmnet with OS trixie * 06:44 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1194.eqiad.wmnet with OS bookworm * 06:44 XioNoX: upgrade eqsin gnmic to 0.47.0 * 06:43 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1198.eqiad.wmnet with OS bookworm * 06:41 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1054.eqiad.wmnet with OS trixie * 06:09 arnaudb@cumin1003: END (PASS) - Cookbook sre.gerrit.restart-gerrit (exit_code=0) Restarting Gerrit on gerrit2002 * 06:06 arnaudb@cumin1003: START - Cookbook sre.gerrit.restart-gerrit Restarting Gerrit on gerrit2002 * 06:06 arnaudb@cumin1003: END (PASS) - Cookbook sre.gerrit.restart-gerrit (exit_code=0) Restarting Gerrit on gerrit1003 * 06:04 arnaudb@cumin1003: START - Cookbook sre.gerrit.restart-gerrit Restarting Gerrit on gerrit1003 * 06:02 arnaudb@cumin1003: END (PASS) - Cookbook sre.gerrit.restart-gerrit (exit_code=0) Restarting Gerrit on gerrit2003 * 06:00 arnaudb@cumin1003: START - Cookbook sre.gerrit.restart-gerrit Restarting Gerrit on gerrit2003 * 05:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1155.eqiad.wmnet with OS bookworm * 05:38 arnaudb: updating prometheusBearerToken on gerrit * 05:28 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1155.eqiad.wmnet with reason: host reimage * 05:23 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1155.eqiad.wmnet with reason: host reimage * 05:06 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1155.eqiad.wmnet with OS bookworm * 04:57 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1144.eqiad.wmnet with OS bookworm * 04:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1144.eqiad.wmnet with reason: host reimage * 04:29 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1144.eqiad.wmnet with reason: host reimage * 04:14 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1144.eqiad.wmnet with OS bookworm * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.13 (duration: 02m 23s) * 03:45 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1197.eqiad.wmnet with OS bookworm * 03:41 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1165.eqiad.wmnet with OS bookworm * 03:38 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.16 refs [[phab:T430835|T430835]] (duration: 34m 43s) * 03:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1164.eqiad.wmnet with OS bookworm * 03:35 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1196.eqiad.wmnet with OS bookworm * 03:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1163.eqiad.wmnet with OS bookworm * 03:22 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1197.eqiad.wmnet with reason: host reimage * 03:18 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1165.eqiad.wmnet with reason: host reimage * 03:15 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1196.eqiad.wmnet with reason: host reimage * 03:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1164.eqiad.wmnet with reason: host reimage * 03:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1163.eqiad.wmnet with reason: host reimage * 03:05 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1165.eqiad.wmnet with reason: host reimage * 03:04 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1164.eqiad.wmnet with reason: host reimage * 03:04 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1197.eqiad.wmnet with reason: host reimage * 03:04 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1196.eqiad.wmnet with reason: host reimage * 03:04 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1163.eqiad.wmnet with reason: host reimage * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.16 refs [[phab:T430835|T430835]] * 02:50 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1197.eqiad.wmnet with OS bookworm * 02:49 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1196.eqiad.wmnet with OS bookworm * 02:49 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1165.eqiad.wmnet with OS bookworm * 02:49 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1164.eqiad.wmnet with OS bookworm * 02:48 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1163.eqiad.wmnet with OS bookworm * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 46s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:15 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-codfw: Set storage compatability to NONE — [[phab:T433026|T433026]] - eevans@cumin1003 * 00:38 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-codfw: Set storage compatability to NONE — [[phab:T433026|T433026]] - eevans@cumin1003 * 00:11 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326373{{!}}PersonalDashboard: add newly renamed *ReviewChangesMlModel setting (T422148)]] (duration: 07m 07s) * 00:09 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-eqiad: Set storage compatability to NONE — [[phab:T433026|T433026]] - eevans@cumin1003 * 00:07 musikanimal@deploy1003: musikanimal: Continuing with deployment * 00:06 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1326373{{!}}PersonalDashboard: add newly renamed *ReviewChangesMlModel setting (T422148)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:04 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1326373{{!}}PersonalDashboard: add newly renamed *ReviewChangesMlModel setting (T422148)]] == 2026-08-17 == * 23:30 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-eqiad: Set storage compatability to NONE — [[phab:T433026|T433026]] - eevans@cumin1003 * 23:05 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-codfw: Set storage compatability to UPGRADING — [[phab:T433026|T433026]] - eevans@cumin1003 * 22:28 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-codfw: Set storage compatability to UPGRADING — [[phab:T433026|T433026]] - eevans@cumin1003 * 21:46 logmsgbot: jforrester Deployed security patch for [[phab:T435085|T435085]] * 21:39 swfrench@deploy1003: mwscript-k8s job started: purgeList.php # [[phab:T432412|T432412]] * 21:37 maryum: Undeploy security fix for [[phab:T433020|T433020]] * 21:23 maryum: Deployed security fix for [[phab:T433020|T433020]] * 21:21 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 21:21 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 21:14 maryum: Deployed security fix for [[phab:T434967|T434967]] * 20:58 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-eqiad: Set storage compatability to UPGRADING — [[phab:T433026|T433026]] - eevans@cumin1003 * 20:40 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324320{{!}}[itwiki/slwiki/tgwiki] Remove temporary Wikipedia 25 logos permanently (already reverted) (T414265 T414320 T415307)]] (duration: 06m 55s) * 20:40 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 20:39 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 20:36 cjming@deploy1003: cjming, superpes: Continuing with deployment * 20:35 cjming@deploy1003: cjming, superpes: Backport for [[gerrit:1324320{{!}}[itwiki/slwiki/tgwiki] Remove temporary Wikipedia 25 logos permanently (already reverted) (T414265 T414320 T415307)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:33 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1324320{{!}}[itwiki/slwiki/tgwiki] Remove temporary Wikipedia 25 logos permanently (already reverted) (T414265 T414320 T415307)]] * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ttmserver-test: apply * 20:30 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326361{{!}}Remove escaped paths in app site association file (T432412)]] (duration: 13m 28s) * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ttmserver-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-toolhub-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-toolhub-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-toolhub-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-toolhub-test: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-toolhub: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-toolhub: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-toolhub: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-toolhub: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-test: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-test: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 20:27 inflatador: bking@deploy1003 `charlie --services_dir dse-k8s-services -s opensearch-* -e dse-k8s-* apply` [[phab:T435125|T435125]] * 20:27 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-apifeatureusage-test: apply * 20:27 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-apifeatureusage-test: apply * 20:27 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-apifeatureusage-test: apply * 20:27 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-apifeatureusage-test: apply * 20:27 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-apifeatureusage: apply * 20:26 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-apifeatureusage: apply * 20:26 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-apifeatureusage: apply * 20:26 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-apifeatureusage: apply * 20:26 cjming@deploy1003: cjming, tsev: Continuing with deployment * 20:24 inflatador: bking@deploy1003 `charlie --services_dir dse-k8s-services -s opensearch-* -e dse-k8s-*` [[phab:T435125|T435125]] * 20:21 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-eqiad: Set storage compatability to UPGRADING — [[phab:T433026|T433026]] - eevans@cumin1003 * 20:19 cjming@deploy1003: cjming, tsev: Backport for [[gerrit:1326361{{!}}Remove escaped paths in app site association file (T432412)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:17 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1326361{{!}}Remove escaped paths in app site association file (T432412)]] * 20:15 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326077{{!}}InstrumentConstructiveEdits: anchor all runs to the nearest `interval` (T431493)]] (duration: 06m 25s) * 20:11 cjming@deploy1003: cjming: Continuing with deployment * 20:10 cjming@deploy1003: cjming: Backport for [[gerrit:1326077{{!}}InstrumentConstructiveEdits: anchor all runs to the nearest `interval` (T431493)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:08 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1326077{{!}}InstrumentConstructiveEdits: anchor all runs to the nearest `interval` (T431493)]] * 20:06 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1160.eqiad.wmnet with OS bookworm * 20:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1162.eqiad.wmnet with OS bookworm * 19:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1184.eqiad.wmnet with OS bookworm * 19:55 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1161.eqiad.wmnet with OS bookworm * 19:49 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1195.eqiad.wmnet with OS bookworm * 19:49 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-test: apply * 19:49 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-test: apply * 19:44 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1160.eqiad.wmnet with reason: host reimage * 19:41 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-codfw: Upgrade to Java 17 — [[phab:T433026|T433026]] - eevans@cumin1003 * 19:39 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1162.eqiad.wmnet with reason: host reimage * 19:36 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-eqsin and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 19:36 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-test: apply * 19:36 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1184.eqiad.wmnet with reason: host reimage * 19:32 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1161.eqiad.wmnet with reason: host reimage * 19:29 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1195.eqiad.wmnet with reason: host reimage * 19:26 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1161.eqiad.wmnet with reason: host reimage * 19:26 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1160.eqiad.wmnet with reason: host reimage * 19:26 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1184.eqiad.wmnet with reason: host reimage * 19:26 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1162.eqiad.wmnet with reason: host reimage * 19:25 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1195.eqiad.wmnet with reason: host reimage * 19:19 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-test: apply * 19:11 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1195.eqiad.wmnet with OS bookworm * 19:10 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1194.eqiad.wmnet with OS bookworm * 19:10 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1184.eqiad.wmnet with OS bookworm * 19:10 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1162.eqiad.wmnet with OS bookworm * 19:10 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1161.eqiad.wmnet with OS bookworm * 19:10 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1160.eqiad.wmnet with OS bookworm * 19:10 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 19:09 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 19:03 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-codfw: Upgrade to Java 17 — [[phab:T433026|T433026]] - eevans@cumin1003 * 18:58 dancy@deploy1003: Finished scap sync-world: testing [[phab:T375514|T375514]] (duration: 03m 13s) * 18:55 dancy@deploy1003: Started scap sync-world: testing [[phab:T375514|T375514]] * 18:55 dwisehaupt@dns1006: END - running authdns-update * 18:54 dancy@deploy1003: Installation of scap version "4.281.1" completed for 3 hosts * 18:54 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2009.codfw.wmnet * 18:53 dwisehaupt@dns1006: START - running authdns-update * 18:52 dancy@deploy1003: Installing scap version "4.281.1" for 3 host(s) * 18:47 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2009.codfw.wmnet * 18:41 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2008.codfw.wmnet * 18:34 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2008.codfw.wmnet * 18:30 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2007.codfw.wmnet * 18:23 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2007.codfw.wmnet * 18:16 dwisehaupt@dns1005: END - running authdns-update * 18:14 dwisehaupt@dns1005: START - running authdns-update * 18:04 swfrench@deploy1003: Finished scap sync-world: Deploy "Point Test Wiki to new docroot" - [[phab:T432412|T432412]] (duration: 26m 03s) * 18:00 swfrench@deploy1003: swfrench: Continuing with deployment * 17:51 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1057.eqiad.wmnet with OS trixie * 17:47 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1180.eqiad.wmnet with OS bookworm * 17:39 swfrench@deploy1003: swfrench: Deploy "Point Test Wiki to new docroot" - [[phab:T432412|T432412]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:39 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-eqsin and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 17:38 swfrench@deploy1003: Started scap sync-world: Deploy "Point Test Wiki to new docroot" - [[phab:T432412|T432412]] * 17:36 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1159.eqiad.wmnet with OS bookworm * 17:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1158.eqiad.wmnet with OS bookworm * 17:30 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-eqiad: Upgrade to Java 17 — [[phab:T433026|T433026]] - eevans@cumin1003 * 17:30 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-ulsfo or A:cp-drmrs and A:cp - 9.2.15 upgrade ([[phab:T434620|T434620]]) * 17:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1157.eqiad.wmnet with OS bookworm * 17:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1180.eqiad.wmnet with reason: host reimage * 17:22 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 17:21 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1193.eqiad.wmnet with OS bookworm * 17:21 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 17:21 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 17:20 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 17:20 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 17:20 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1183.eqiad.wmnet with OS bookworm * 17:18 swfrench@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 17:18 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 17:17 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1180.eqiad.wmnet with reason: host reimage * 17:17 swfrench@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 17:17 swfrench@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 17:16 swfrench@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 17:16 swfrench@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 17:15 swfrench@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 17:15 swfrench@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 17:15 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1192.eqiad.wmnet with OS bookworm * 17:14 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1159.eqiad.wmnet with reason: host reimage * 17:14 swfrench@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 17:13 swfrench@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 17:12 swfrench@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 17:09 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1158.eqiad.wmnet with reason: host reimage * 17:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1157.eqiad.wmnet with reason: host reimage * 17:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1180 * 17:02 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1180 * 17:01 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica-eqsin and A:liberica * 17:01 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1193.eqiad.wmnet with reason: host reimage * 16:57 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1183.eqiad.wmnet with reason: host reimage * 16:56 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1054.eqiad.wmnet with OS trixie * 16:55 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1056.eqiad.wmnet with OS trixie * 16:54 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1193.eqiad.wmnet with reason: host reimage * 16:54 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1192.eqiad.wmnet with reason: host reimage * 16:51 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-eqiad: Upgrade to Java 17 — [[phab:T433026|T433026]] - eevans@cumin1003 * 16:50 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1180 * 16:50 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1180.eqiad.wmnet 17.36.64.10.in-addr.arpa 7.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:50 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1180.eqiad.wmnet 17.36.64.10.in-addr.arpa 7.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:50 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:50 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1180 - btullis@cumin1003" * 16:50 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1180 - btullis@cumin1003" * 16:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1158.eqiad.wmnet with reason: host reimage * 16:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1157.eqiad.wmnet with reason: host reimage * 16:49 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica-eqsin and A:liberica * 16:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1183.eqiad.wmnet with reason: host reimage * 16:48 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1159.eqiad.wmnet with reason: host reimage * 16:47 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1192.eqiad.wmnet with reason: host reimage * 16:46 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns4004.wikimedia.org * 16:46 sukhe@dns1004: END - running authdns-update * 16:44 sukhe@dns1004: START - running authdns-update * 16:44 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns4004.wikimedia.org,service=authdns-update * 16:43 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns4004.wikimedia.org with OS trixie * 16:39 btullis@cumin1003: START - Cookbook sre.dns.netbox * 16:39 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1180 * 16:39 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1193.eqiad.wmnet with OS bookworm * 16:39 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1180.eqiad.wmnet with OS bookworm * 16:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1192.eqiad.wmnet with OS bookworm * 16:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1183.eqiad.wmnet with OS bookworm * 16:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1159.eqiad.wmnet with OS bookworm * 16:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1158.eqiad.wmnet with OS bookworm * 16:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1157.eqiad.wmnet with OS bookworm * 16:31 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1057.eqiad.wmnet with OS trixie * 16:30 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1057.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 16:29 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica-drmrs and A:liberica * 16:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1154.eqiad.wmnet with OS bookworm * 16:23 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1057.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 16:23 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1055.eqiad.wmnet with OS trixie * 16:22 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1057 * 16:22 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1057 * 16:21 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:21 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1057] - vriley@cumin1003" * 16:21 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1057] - vriley@cumin1003" * 16:19 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica-drmrs and A:liberica * 16:17 vriley@cumin1003: START - Cookbook sre.dns.netbox * 16:16 phuedx@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics-external: apply * 16:15 phuedx@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics-external: apply * 16:13 phuedx@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics-external: apply * 16:12 phuedx@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics-external: apply * 16:11 phuedx@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics-external: apply * 16:09 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics-external: apply * 16:09 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1176.eqiad.wmnet with OS bookworm * 16:07 btullis@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 16:06 btullis@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 16:05 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1191.eqiad.wmnet with OS bookworm * 16:03 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1154.eqiad.wmnet with reason: host reimage * 16:03 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching aqs[2002-2012].codfw.wmnet,aqs[1017-1027].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433026|T433026]] - eevans@cumin1003 * 15:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1190.eqiad.wmnet with OS bookworm * 15:59 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1154.eqiad.wmnet with reason: host reimage * 15:55 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-debug: apply * 15:55 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-debug: apply * 15:55 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-debug: apply * 15:55 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/mw-debug: apply * 15:53 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns4004.wikimedia.org with reason: host reimage * 15:50 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns4004.wikimedia.org with reason: host reimage * 15:46 btullis@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 15:46 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1176.eqiad.wmnet with reason: host reimage * 15:45 btullis@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 15:43 btullis@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 15:42 btullis@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 15:42 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1191.eqiad.wmnet with reason: host reimage * 15:39 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1190.eqiad.wmnet with reason: host reimage * 15:36 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1054.eqiad.wmnet with OS trixie * 15:35 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:35 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1056.eqiad.wmnet with OS trixie * 15:35 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:34 moritzm: failover Ganeti master in eqiad to ganeti1046 * 15:34 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1176.eqiad.wmnet with reason: host reimage * 15:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1191.eqiad.wmnet with reason: host reimage * 15:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1190.eqiad.wmnet with reason: host reimage * 15:31 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns2006.wikimedia.org * 15:31 sukhe@dns1004: END - running authdns-update * 15:29 sukhe@dns1004: START - running authdns-update * 15:29 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns2006.wikimedia.org,service=authdns-update * 15:29 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns2006.wikimedia.org * 15:29 sukhe@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns2006.wikimedia.org * 15:26 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica-ulsfo and A:liberica * 15:26 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns1006.wikimedia.org * 15:26 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:25 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns2006.wikimedia.org with OS trixie * 15:25 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1056 * 15:25 sukhe@dns1004: END - running authdns-update * 15:23 sukhe@dns1004: START - running authdns-update * 15:23 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns1006.wikimedia.org,service=authdns-update * 15:23 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns1006.wikimedia.org * 15:23 sukhe@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns1006.wikimedia.org * 15:20 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1056 * 15:20 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:20 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1056~] - vriley@cumin1003" * 15:19 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1056~] - vriley@cumin1003" * 15:19 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns1006.wikimedia.org with OS trixie * 15:19 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns4004.wikimedia.org with OS trixie * 15:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1190.eqiad.wmnet with OS bookworm * 15:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1191.eqiad.wmnet with OS bookworm * 15:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1176.eqiad.wmnet with OS bookworm * 15:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1154.eqiad.wmnet with OS bookworm * 15:16 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica-ulsfo and A:liberica * 15:13 vriley@cumin1003: START - Cookbook sre.dns.netbox * 15:13 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host dns4004.wikimedia.org with OS trixie * 15:12 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1188.eqiad.wmnet with OS bookworm * 15:09 taavi@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318203{{!}}Undeploy WP25EasterEggs (II) (T418134)]] (duration: 08m 56s) * 15:06 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica-magru and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 15:06 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1051.eqiad.wmnet with OS trixie * 15:06 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 15:05 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 15:05 taavi@deploy1003: taavi: Continuing with deployment * 15:04 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-ncredir (exit_code=0) rolling reboot on A:ncredir and A:ncredir * 15:04 taavi@deploy1003: taavi: Backport for [[gerrit:1318203{{!}}Undeploy WP25EasterEggs (II) (T418134)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:03 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1055.eqiad.wmnet with OS trixie * 15:02 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:00 taavi@deploy1003: Started scap sync-world: Backport for [[gerrit:1318203{{!}}Undeploy WP25EasterEggs (II) (T418134)]] * 14:59 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1047.eqiad.wmnet * 14:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1047.eqiad.wmnet * 14:58 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy (exit_code=0) rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 14:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-codfw * 14:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp2001.codfw.wmnet * 14:57 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp2001.codfw.wmnet * 14:57 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica-magru and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 14:57 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:56 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1055 * 14:56 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1055 * 14:55 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns2006.wikimedia.org with reason: host reimage * 14:55 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:55 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1055] - vriley@cumin1003" * 14:55 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1055] - vriley@cumin1003" * 14:54 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1047.eqiad.wmnet * 14:52 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1188.eqiad.wmnet with reason: host reimage * 14:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp2001.codfw.wmnet * 14:51 vriley@cumin1003: START - Cookbook sre.dns.netbox * 14:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp2001.codfw.wmnet * 14:50 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2016.codfw.wmnet * 14:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2016.codfw.wmnet * 14:50 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1189.eqiad.wmnet with OS bookworm * 14:48 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1188.eqiad.wmnet with reason: host reimage * 14:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1051.eqiad.wmnet with reason: host reimage * 14:46 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1182.eqiad.wmnet with OS bookworm * 14:45 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief2002.codfw.wmnet * 14:44 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns1006.wikimedia.org with reason: host reimage * 14:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2016.codfw.wmnet * 14:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1143.eqiad.wmnet with OS bookworm * 14:43 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2016.codfw.wmnet * 14:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2318-2331].codfw.wmnet * 14:43 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2318-2331].codfw.wmnet * 14:42 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching aqs[2002-2012].codfw.wmnet,aqs[1017-1027].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433026|T433026]] - eevans@cumin1003 * 14:41 cgoubert@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326306{{!}}Add placeholder $wmgRedisLockPassword (T366938 T427999)]] (duration: 06m 56s) * 14:41 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief2002.codfw.wmnet * 14:40 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief1002.eqiad.wmnet * 14:39 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1051.eqiad.wmnet with reason: host reimage * 14:38 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns2006.wikimedia.org with reason: host reimage * 14:37 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns1006.wikimedia.org with reason: host reimage * 14:37 cgoubert@deploy1003: cgoubert: Continuing with deployment * 14:36 cgoubert@deploy1003: cgoubert: Backport for [[gerrit:1326306{{!}}Add placeholder $wmgRedisLockPassword (T366938 T427999)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:36 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief1002.eqiad.wmnet * 14:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2318-2331].codfw.wmnet * 14:34 cgoubert@deploy1003: Started scap sync-world: Backport for [[gerrit:1326306{{!}}Add placeholder $wmgRedisLockPassword (T366938 T427999)]] * 14:29 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2318-2331].codfw.wmnet * 14:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2304-2317].codfw.wmnet * 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2304-2317].codfw.wmnet * 14:28 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1054 * 14:27 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1054 * 14:27 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:27 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1054] - vriley@cumin1003" * 14:27 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1054] - vriley@cumin1003" * 14:26 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1189.eqiad.wmnet with reason: host reimage * 14:26 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test2001.codfw.wmnet * 14:25 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test1001.eqiad.wmnet * 14:25 claime: Deploying wmgRedisLockPassword - [[phab:T366938|T366938]] [[phab:T427999|T427999]] * 14:24 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1051.eqiad.wmnet with OS trixie * 14:23 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 14:22 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1182.eqiad.wmnet with reason: host reimage * 14:22 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1051.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:22 vriley@cumin1003: START - Cookbook sre.dns.netbox * 14:22 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test2001.codfw.wmnet * 14:21 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test1001.eqiad.wmnet * 14:21 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 14:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2304-2317].codfw.wmnet * 14:19 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns4004.wikimedia.org with OS trixie * 14:19 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns2006.wikimedia.org with OS trixie * 14:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1143.eqiad.wmnet with reason: host reimage * 14:19 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns1006.wikimedia.org with OS trixie * 14:17 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1189.eqiad.wmnet with reason: host reimage * 14:15 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1182.eqiad.wmnet with reason: host reimage * 14:14 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1143.eqiad.wmnet with reason: host reimage * 14:13 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1051.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:12 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2304-2317].codfw.wmnet * 14:12 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2290-2303].codfw.wmnet * 14:12 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1049.eqiad.wmnet with OS trixie * 14:12 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2290-2303].codfw.wmnet * 14:12 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1051 * 14:11 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1051 * 14:11 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1170.eqiad.wmnet onto db1284.eqiad.wmnet * 14:11 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1170: Pool db1170.eqiad.wmnet in after cloning * 14:09 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 14:09 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:09 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1051] - vriley@cumin1003" * 14:09 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1051] - vriley@cumin1003" * 14:06 klausman@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:04 vriley@cumin1003: START - Cookbook sre.dns.netbox * 14:04 klausman@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2290-2303].codfw.wmnet * 14:02 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1189.eqiad.wmnet with OS bookworm * 14:02 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1188.eqiad.wmnet with OS bookworm * 14:01 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 14:00 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1182.eqiad.wmnet with OS bookworm * 14:00 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1143.eqiad.wmnet with OS bookworm * 13:59 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 13:57 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply * 13:57 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply * 13:56 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply * 13:56 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-ulsfo or A:cp-drmrs and A:cp - 9.2.15 upgrade ([[phab:T434620|T434620]]) * 13:56 cjd91: sudo -i cookbook sre.cdn.roll-upgrade-ats --query 'A:cp-ulsfo or A:cp-drmrs' --task-id [[phab:T434620|T434620]] --reason '9.2.15 upgrade' * 13:56 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply * 13:55 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply * 13:55 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply * 13:54 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 13:54 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 13:54 phuedx@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:53 phuedx@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics-external: apply * 13:52 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1049.eqiad.wmnet with reason: host reimage * 13:51 phuedx@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2290-2303].codfw.wmnet * 13:50 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1047.eqiad.wmnet * 13:50 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2276-2289].codfw.wmnet * 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2276-2289].codfw.wmnet * 13:49 phuedx@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics-external: apply * 13:49 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1046.eqiad.wmnet * 13:49 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1049.eqiad.wmnet with reason: host reimage * 13:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1046.eqiad.wmnet * 13:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-ncredir rolling reboot on A:ncredir and A:ncredir * 13:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 13:46 phuedx@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:44 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics-external: apply * 13:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1046.eqiad.wmnet * 13:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2276-2289].codfw.wmnet * 13:36 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1046.eqiad.wmnet * 13:36 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1045.eqiad.wmnet * 13:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1045.eqiad.wmnet * 13:34 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2276-2289].codfw.wmnet * 13:34 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1049.eqiad.wmnet with OS trixie * 13:34 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2262-2275].codfw.wmnet * 13:34 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2262-2275].codfw.wmnet * 13:32 Lucas_WMDE: UTC afternoon backport+config window done * 13:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1045.eqiad.wmnet * 13:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2262-2275].codfw.wmnet * 13:26 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1049.eqiad.wmnet with OS trixie * 13:26 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1049.eqiad.wmnet with OS trixie * 13:26 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1170: Pool db1170.eqiad.wmnet in after cloning * 13:23 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1049.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 13:23 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1045.eqiad.wmnet * 13:21 atsukoito: manually done sudo -i docker-registryctl --debug delete-tags 'docker-registry.discovery.wmnet/repos/data-engineering/airflow-dags:airflow-3.3.0-py3.11-2026-08-17-*' to remove incorrect tags * 13:19 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311141{{!}}viwiki: Set `noindex,nofollow` for User and User talk (T432311)]] (duration: 11m 11s) * 13:15 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2262-2275].codfw.wmnet * 13:14 lucaswerkmeister-wmde@deploy1003: ndkdd, lucaswerkmeister-wmde: Continuing with deployment * 13:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2248-2261].codfw.wmnet * 13:14 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2248-2261].codfw.wmnet * 13:13 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1049.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 13:12 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1049 * 13:11 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1049 * 13:10 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:10 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1049] - vriley@cumin1003" * 13:10 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1049] - vriley@cumin1003" * 13:10 lucaswerkmeister-wmde@deploy1003: ndkdd, lucaswerkmeister-wmde: Backport for [[gerrit:1311141{{!}}viwiki: Set `noindex,nofollow` for User and User talk (T432311)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1037.eqiad.wmnet * 13:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1037.eqiad.wmnet * 13:08 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1311141{{!}}viwiki: Set `noindex,nofollow` for User and User talk (T432311)]] * 13:06 vriley@cumin1003: START - Cookbook sre.dns.netbox * 13:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2248-2261].codfw.wmnet * 13:00 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1037.eqiad.wmnet * 12:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2248-2261].codfw.wmnet * 12:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2204-2215,2242-2243].codfw.wmnet * 12:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2204-2215,2242-2243].codfw.wmnet * 12:49 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324832{{!}}Migrate $wgFlaggedRevsTags from flaggedrevs.php to ext-FlaggedRevs.php]] (duration: 14m 02s) * 12:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint2001.codfw.wmnet * 12:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2204-2215,2242-2243].codfw.wmnet * 12:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint2001.codfw.wmnet * 12:41 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1037.eqiad.wmnet * 12:40 ladsgroup@deploy1003: Rolling back deployment * 12:37 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1324832{{!}}Migrate $wgFlaggedRevsTags from flaggedrevs.php to ext-FlaggedRevs.php]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:37 seanleong-wmde: Finished populateSitesTable for [bolwiki] ([[[phab:T429955|T429955]]]) * 12:35 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1324832{{!}}Migrate $wgFlaggedRevsTags from flaggedrevs.php to ext-FlaggedRevs.php]] * 12:35 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2204-2215,2242-2243].codfw.wmnet * 12:34 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2190-2203].codfw.wmnet * 12:34 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2190-2203].codfw.wmnet * 12:32 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1028.eqiad.wmnet * 12:32 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1028.eqiad.wmnet * 12:31 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint1001.eqiad.wmnet * 12:30 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1170: Depool db1170.eqiad.wmnet to then clone it to db1284.eqiad.wmnet - marostegui@cumin1003 * 12:28 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint1001.eqiad.wmnet * 12:28 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1170: Depool db1170.eqiad.wmnet to then clone it to db1284.eqiad.wmnet - marostegui@cumin1003 * 12:27 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1170.eqiad.wmnet onto db1284.eqiad.wmnet * 12:27 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db2901.codfw.wmnet * 12:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2190-2203].codfw.wmnet * 12:27 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 12:26 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1028.eqiad.wmnet * 12:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2190-2203].codfw.wmnet * 12:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2172-2179,2184-2189].codfw.wmnet * 12:18 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2172-2179,2184-2189].codfw.wmnet * 12:13 seanleong-wmde@deploy1003: mwscript-k8s job started: foreachwikiindblist wikidataclient extensions/Wikibase/lib/maintenance/populateSitesTable.php --force-protocol https # [[phab:T429955|T429955]] * 12:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2172-2179,2184-2189].codfw.wmnet * 12:09 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1028.eqiad.wmnet * 12:03 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 12:02 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db2901.codfw.wmnet * 12:02 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db2901.codfw.wmnet * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1027.eqiad.wmnet * 12:02 fceratto@cumin1003: END (ERROR) - Cookbook sre.dns.netbox (exit_code=97) * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1027.eqiad.wmnet * 12:02 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 12:02 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db2901.codfw.wmnet * 12:01 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2172-2179,2184-2189].codfw.wmnet * 12:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2158-2171].codfw.wmnet * 12:01 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2158-2171].codfw.wmnet * 11:58 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db1902.eqiad.wmnet * 11:58 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 11:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1027.eqiad.wmnet * 11:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2158-2171].codfw.wmnet * 11:51 jayme: updated calico to v3.30.7 on wikikube eqiad - [[phab:T427400|T427400]] * 11:50 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1027.eqiad.wmnet * 11:45 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2158-2171].codfw.wmnet * 11:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2144-2157].codfw.wmnet * 11:44 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2144-2157].codfw.wmnet * 11:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1058.eqiad.wmnet * 11:43 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1058.eqiad.wmnet * 11:43 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 11:43 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 11:43 marostegui@cumin1003: Removing db1153 from zarcillo [[phab:T434638|T434638]] * 11:42 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1153.eqiad.wmnet * 11:42 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:42 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1153.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 11:42 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1172.eqiad.wmnet onto db1286.eqiad.wmnet * 11:42 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1172: Pool db1172.eqiad.wmnet in after cloning * 11:42 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1153.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 11:41 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1902.eqiad.wmnet * 11:41 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=97) for new host db2901.codfw.wmnet * 11:41 fceratto@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host db2901.codfw.wmnet with OS trixie * 11:38 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'. * 11:38 marostegui@cumin1003: START - Cookbook sre.dns.netbox * 11:37 marostegui@dns1004: END - running authdns-update * 11:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1058.eqiad.wmnet * 11:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2144-2157].codfw.wmnet * 11:35 marostegui@dns1004: START - running authdns-update * 11:32 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1153.eqiad.wmnet * 11:32 marostegui@cumin1003: START - Cookbook sre.mysql.decommission * 11:28 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326249{{!}}ImagePage: move TOC element below file link (T332644)]] (duration: 09m 56s) * 11:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2144-2157].codfw.wmnet * 11:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2130-2143].codfw.wmnet * 11:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2130-2143].codfw.wmnet * 11:26 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1058.eqiad.wmnet * 11:23 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 11:22 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'. * 11:22 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'. * 11:22 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'. * 11:22 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326249{{!}}ImagePage: move TOC element below file link (T332644)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1057.eqiad.wmnet * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1057.eqiad.wmnet * 11:21 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 11:20 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 11:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2130-2143].codfw.wmnet * 11:19 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326249{{!}}ImagePage: move TOC element below file link (T332644)]] * 11:18 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply * 11:17 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply * 11:17 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply * 11:16 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 11:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1057.eqiad.wmnet * 11:15 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 11:14 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 11:13 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 11:13 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db2901.codfw.wmnet with OS trixie * 11:12 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db2901.codfw.wmnet - fceratto@cumin1003" * 11:12 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db2901.codfw.wmnet - fceratto@cumin1003" * 11:12 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db2901.codfw.wmnet on all recursors * 11:12 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db2901.codfw.wmnet on all recursors * 11:12 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:12 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db2901.codfw.wmnet - fceratto@cumin1003" * 11:12 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2130-2143].codfw.wmnet * 11:11 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2107-2115,2124-2129].codfw.wmnet * 11:11 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2107-2115,2124-2129].codfw.wmnet * 11:11 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 11:11 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 11:11 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply * 11:11 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 11:10 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 11:06 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db2901.codfw.wmnet - fceratto@cumin1003" * 11:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2107-2115,2124-2129].codfw.wmnet * 10:57 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1172: Pool db1172.eqiad.wmnet in after cloning * 10:54 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2107-2115,2124-2129].codfw.wmnet * 10:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2078,2087-2095,2102-2106].codfw.wmnet * 10:54 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2078,2087-2095,2102-2106].codfw.wmnet * 10:47 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply * 10:46 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply * 10:46 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply * 10:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2078,2087-2095,2102-2106].codfw.wmnet * 10:45 blake@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply * 10:38 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1057.eqiad.wmnet * 10:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1056.eqiad.wmnet * 10:38 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1056.eqiad.wmnet * 10:37 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2078,2087-2095,2102-2106].codfw.wmnet * 10:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2061-2062,2064-2065,2067-2077].codfw.wmnet * 10:36 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2061-2062,2064-2065,2067-2077].codfw.wmnet * 10:32 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1056.eqiad.wmnet * 10:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2061-2062,2064-2065,2067-2077].codfw.wmnet * 10:25 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:25 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db2901.codfw.wmnet * 10:25 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db1902.eqiad.wmnet * 10:25 fceratto@cumin1003: END (ERROR) - Cookbook sre.dns.netbox (exit_code=97) * 10:24 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:24 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1902.eqiad.wmnet * 10:24 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=97) for new host db1902.eqiad.wmnet * 10:24 fceratto@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host db1902.eqiad.wmnet with OS trixie * 10:24 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=93) for new host db1903.eqiad.wmnet * 10:24 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 10:20 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1056.eqiad.wmnet * 10:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2061-2062,2064-2065,2067-2077].codfw.wmnet * 10:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2038-2039,2041-2042,2044,2046,2049-2051,2055-2060].codfw.wmnet * 10:18 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2038-2039,2041-2042,2044,2046,2049-2051,2055-2060].codfw.wmnet * 10:15 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1055.eqiad.wmnet * 10:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1055.eqiad.wmnet * 10:13 Amir1: mwscript-k8s --dblist=all -- purgeUserOptions.php --login-age 5 uls-preferences ([[phab:T406724|T406724]]) * 10:11 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:10 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1903.eqiad.wmnet on all recursors * 10:10 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1903.eqiad.wmnet on all recursors * 10:10 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2038-2039,2041-2042,2044,2046,2049-2051,2055-2060].codfw.wmnet * 10:10 moritzm: installing unzip security updates * 10:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1055.eqiad.wmnet * 10:08 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=97) for new host db1901.eqiad.wmnet * 10:08 fceratto@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host db1901.eqiad.wmnet with OS trixie * 10:08 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:08 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 10:08 fceratto@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1903.eqiad.wmnet - fceratto@cumin1003" * 10:07 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db1902.eqiad.wmnet with OS trixie * 10:07 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1902.eqiad.wmnet - fceratto@cumin1003" * 10:07 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1902.eqiad.wmnet - fceratto@cumin1003" * 10:04 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply * 10:04 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326227{{!}}Enable desktop/native lazy loading everywhere (T148047)]] (duration: 07m 13s) * 10:03 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1902.eqiad.wmnet on all recursors * 10:03 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1902.eqiad.wmnet on all recursors * 10:03 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:03 blake@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply * 10:01 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1903.eqiad.wmnet - fceratto@cumin1003" * 10:01 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2038-2039,2041-2042,2044,2046,2049-2051,2055-2060].codfw.wmnet * 10:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2002,2005-2006,2011-2015,2017-2018,2033-2037].codfw.wmnet * 10:01 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2002,2005-2006,2011-2015,2017-2018,2033-2037].codfw.wmnet * 10:00 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:00 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 09:59 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 09:59 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1055.eqiad.wmnet * 09:58 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326227{{!}}Enable desktop/native lazy loading everywhere (T148047)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:56 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326227{{!}}Enable desktop/native lazy loading everywhere (T148047)]] * 09:53 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2002,2005-2006,2011-2015,2017-2018,2033-2037].codfw.wmnet * 09:50 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:48 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1903.eqiad.wmnet * 09:48 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:48 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1902.eqiad.wmnet * 09:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 09:44 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2002,2005-2006,2011-2015,2017-2018,2033-2037].codfw.wmnet * 09:43 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-codfw * 09:43 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1174.eqiad.wmnet onto db1288.eqiad.wmnet * 09:42 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1174: Pool db1174.eqiad.wmnet in after cloning * 09:40 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1172: Depool db1172.eqiad.wmnet to then clone it to db1286.eqiad.wmnet - marostegui@cumin1003 * 09:39 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1172: Depool db1172.eqiad.wmnet to then clone it to db1286.eqiad.wmnet - marostegui@cumin1003 * 09:39 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1172.eqiad.wmnet onto db1286.eqiad.wmnet * 09:33 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1175.eqiad.wmnet onto db1289.eqiad.wmnet * 09:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1175: Pool db1175.eqiad.wmnet in after cloning * 09:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1201.eqiad.wmnet onto db1287.eqiad.wmnet * 09:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1201: Pool db1201.eqiad.wmnet in after cloning * 09:28 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db1901.eqiad.wmnet with OS trixie * 09:27 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:27 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:27 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1901.eqiad.wmnet on all recursors * 09:27 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1901.eqiad.wmnet on all recursors * 09:26 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:26 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:26 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:15 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:15 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1901.eqiad.wmnet * 09:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw1001.wikimedia.org with OS trixie * 08:57 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1174: Pool db1174.eqiad.wmnet in after cloning * 08:54 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1054.eqiad.wmnet * 08:54 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1054.eqiad.wmnet * 08:48 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1054.eqiad.wmnet * 08:47 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1175: Pool db1175.eqiad.wmnet in after cloning * 08:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 08:46 marostegui@cumin1003: Removing db1152 from zarcillo [[phab:T434480|T434480]] * 08:46 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1152.eqiad.wmnet * 08:46 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:46 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1152.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 08:46 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1201: Pool db1201.eqiad.wmnet in after cloning * 08:46 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1152.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 08:46 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1054.eqiad.wmnet * 08:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1053.eqiad.wmnet * 08:43 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1053.eqiad.wmnet * 08:42 marostegui@cumin1003: START - Cookbook sre.dns.netbox * 08:38 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage * 08:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1053.eqiad.wmnet * 08:36 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1152.eqiad.wmnet * 08:36 marostegui@cumin1003: START - Cookbook sre.mysql.decommission * 08:35 phuedx: UTC morning backport window done * 08:35 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1053.eqiad.wmnet * 08:34 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1035.eqiad.wmnet * 08:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1035.eqiad.wmnet * 08:34 phuedx@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324725{{!}}EventStreamConfig: Mark product_metrics.web_base and .web_base_with_ip as Test Kitchen streams (T429898 T430322)]], [[gerrit:1313923{{!}}EventStreamConfig: Remove unused web_ui_scroll* streams (T415370)]], [[gerrit:1325546{{!}}EventStreamConfig: Remove Watchlist click stream (T434790)]] (duration: 12m 42s) * 08:33 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1281: Pool back * 08:32 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage * 08:31 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on db2209.codfw.wmnet with reason: Maintenance * 08:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2209: Maintenance needed * 08:30 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2209: Maintenance needed * 08:26 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1035.eqiad.wmnet * 08:26 phuedx@deploy1003: bearloga, phuedx: Continuing with deployment * 08:23 phuedx@deploy1003: bearloga, phuedx: Backport for [[gerrit:1324725{{!}}EventStreamConfig: Mark product_metrics.web_base and .web_base_with_ip as Test Kitchen streams (T429898 T430322)]], [[gerrit:1313923{{!}}EventStreamConfig: Remove unused web_ui_scroll* streams (T415370)]], [[gerrit:1325546{{!}}EventStreamConfig: Remove Watchlist click stream (T434790)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug * 08:21 phuedx@deploy1003: Started scap sync-world: Backport for [[gerrit:1324725{{!}}EventStreamConfig: Mark product_metrics.web_base and .web_base_with_ip as Test Kitchen streams (T429898 T430322)]], [[gerrit:1313923{{!}}EventStreamConfig: Remove unused web_ui_scroll* streams (T415370)]], [[gerrit:1325546{{!}}EventStreamConfig: Remove Watchlist click stream (T434790)]] * 08:19 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw1001.wikimedia.org with OS trixie * 08:18 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1035.eqiad.wmnet * 08:17 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1032.eqiad.wmnet * 08:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1032.eqiad.wmnet * 08:16 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1279: Pool back * 08:15 phuedx@deploy1003: Finished scap sync-world: Backport for [[gerrit:1216721{{!}}viwikivoyage: enable relatedarticle and pop-up (T405724)]] (duration: 39m 12s) * 08:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1032.eqiad.wmnet * 08:09 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1032.eqiad.wmnet * 08:08 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1031.eqiad.wmnet * 08:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1031.eqiad.wmnet * 08:03 godog: switch production to use dumps-nfs.w.o - [[phab:T432212|T432212]] * 08:02 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1031.eqiad.wmnet * 08:02 phuedx@deploy1003: nvdtn19, phuedx: Continuing with deployment * 08:00 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1201: Depool db1201.eqiad.wmnet to then clone it to db1287.eqiad.wmnet - marostegui@cumin1003 * 08:00 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1201: Depool db1201.eqiad.wmnet to then clone it to db1287.eqiad.wmnet - marostegui@cumin1003 * 08:00 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1201.eqiad.wmnet onto db1287.eqiad.wmnet * 08:00 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1057.eqiad.wmnet * 08:00 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:00 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1057.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:59 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1057.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:59 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1275: Pool back * 07:55 filippo@cumin1003: START - Cookbook sre.dns.netbox * 07:55 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1031.eqiad.wmnet * 07:52 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1030.eqiad.wmnet * 07:52 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1030.eqiad.wmnet * 07:52 phuedx@deploy1003: nvdtn19, phuedx: Backport for [[gerrit:1216721{{!}}viwikivoyage: enable relatedarticle and pop-up (T405724)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:51 tappof: bump space for prometheus k8s-aux in eqiad * 07:50 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1057.eqiad.wmnet * 07:48 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1281: Pool back * 07:47 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1281 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96108 and previous config saved to /var/cache/conftool/dbconfig/20260817-074749-marostegui.json * 07:46 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1030.eqiad.wmnet * 07:42 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1030.eqiad.wmnet * 07:41 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1174: Depool db1174.eqiad.wmnet to then clone it to db1288.eqiad.wmnet - marostegui@cumin1003 * 07:41 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1174: Depool db1174.eqiad.wmnet to then clone it to db1288.eqiad.wmnet - marostegui@cumin1003 * 07:41 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1174.eqiad.wmnet onto db1288.eqiad.wmnet * 07:40 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1029.eqiad.wmnet * 07:40 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm2001.wikimedia.org * 07:40 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1029.eqiad.wmnet * 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1056.eqiad.wmnet * 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1056.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:38 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1056.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:36 phuedx@deploy1003: Started scap sync-world: Backport for [[gerrit:1216721{{!}}viwikivoyage: enable relatedarticle and pop-up (T405724)]] * 07:36 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm2001.wikimedia.org * 07:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1279 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96104 and previous config saved to /var/cache/conftool/dbconfig/20260817-073542-marostegui.json * 07:34 filippo@cumin1003: START - Cookbook sre.dns.netbox * 07:34 slyngshede@dns1004: END - running authdns-update * 07:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1029.eqiad.wmnet * 07:33 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm-test1001.wikimedia.org * 07:32 slyngshede@dns1004: START - running authdns-update * 07:31 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1029.eqiad.wmnet * 07:31 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1279: Pool back * 07:30 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1279 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96102 and previous config saved to /var/cache/conftool/dbconfig/20260817-073038-marostegui.json * 07:29 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm-test1001.wikimedia.org * 07:29 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm1001.wikimedia.org * 07:28 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1056.eqiad.wmnet * 07:28 moritzm: extend the disk of ldap-rw1001 by 80G [[phab:T331699|T331699]] * 07:28 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1055.eqiad.wmnet * 07:28 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:28 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1055.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:27 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1055.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:26 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1044.eqiad.wmnet * 07:26 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1044.eqiad.wmnet * 07:25 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm1001.wikimedia.org * 07:24 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1175: Depool db1175.eqiad.wmnet to then clone it to db1289.eqiad.wmnet - marostegui@cumin1003 * 07:24 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1175: Depool db1175.eqiad.wmnet to then clone it to db1289.eqiad.wmnet - marostegui@cumin1003 * 07:24 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1175.eqiad.wmnet onto db1289.eqiad.wmnet * 07:22 filippo@cumin1003: START - Cookbook sre.dns.netbox * 07:20 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1044.eqiad.wmnet * 07:16 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1055.eqiad.wmnet * 07:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1054.eqiad.wmnet * 07:15 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:15 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1054.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:15 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1044.eqiad.wmnet * 07:15 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1054.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:13 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1275: Pool back * 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1275 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96098 and previous config saved to /var/cache/conftool/dbconfig/20260817-071225-marostegui.json * 07:10 filippo@cumin1003: START - Cookbook sre.dns.netbox * 07:05 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1043.eqiad.wmnet * 07:05 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1054.eqiad.wmnet * 07:05 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1051.eqiad.wmnet * 07:05 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:05 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1051.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:05 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1043.eqiad.wmnet * 07:04 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1051.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 06:59 filippo@cumin1003: START - Cookbook sre.dns.netbox * 06:59 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin1001.eqiad.wmnet * 06:59 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1043.eqiad.wmnet * 06:59 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin2001.codfw.wmnet * 06:55 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin2001.codfw.wmnet * 06:55 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1051.eqiad.wmnet * 06:54 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1049.eqiad.wmnet * 06:54 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:54 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1049.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 06:54 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1049.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 06:54 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin1001.eqiad.wmnet * 06:53 moritzm: installing apr-util security updates * 06:52 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1043.eqiad.wmnet * 06:49 filippo@cumin1003: START - Cookbook sre.dns.netbox * 06:41 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1049.eqiad.wmnet * 06:13 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit2003.wikimedia.org * 06:13 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet * 06:07 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet * 06:06 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit2003.wikimedia.org * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 47s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-16 == * 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 01m 03s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-15 == * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 41s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-14 == * 15:38 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-staging-master-eqiad * 15:38 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster1005.eqiad.wmnet * 15:38 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster1005.eqiad.wmnet * 15:35 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sretest2009.codfw.wmnet * 15:33 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster1005.eqiad.wmnet * 15:33 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster1005.eqiad.wmnet * 15:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster1004.eqiad.wmnet * 15:32 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster1004.eqiad.wmnet * 15:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host sretest2009.codfw.wmnet * 15:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster1004.eqiad.wmnet * 15:27 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster1004.eqiad.wmnet * 15:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster1003.eqiad.wmnet * 15:27 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster1003.eqiad.wmnet * 15:24 dancy@deploy1003: Finished scap sync-world: testing (duration: 03m 23s) * 15:22 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster1003.eqiad.wmnet * 15:22 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster1003.eqiad.wmnet * 15:22 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-staging-master-eqiad * 15:20 dancy@deploy1003: Started scap sync-world: testing * 15:20 dancy@deploy1003: Installation of scap version "4.280.2" completed for 3 hosts * 15:18 dancy@deploy1003: Installing scap version "4.280.2" for 3 host(s) * 15:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sretest2006.codfw.wmnet * 14:54 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host sretest2006.codfw.wmnet * 14:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sretest2003.codfw.wmnet * 14:39 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host sretest2003.codfw.wmnet * 13:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt-staging2001.codfw.wmnet * 13:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt-staging2001.codfw.wmnet * 13:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-staging-master-codfw * 13:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster2005.codfw.wmnet * 13:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster2005.codfw.wmnet * 13:05 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox-dev2003.codfw.wmnet * 13:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster2005.codfw.wmnet * 13:04 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster2005.codfw.wmnet * 13:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster2004.codfw.wmnet * 13:04 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster2004.codfw.wmnet * 13:01 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netbox-dev2003.codfw.wmnet * 12:59 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster2004.codfw.wmnet * 12:59 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster2004.codfw.wmnet * 12:59 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster2003.codfw.wmnet * 12:59 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster2003.codfw.wmnet * 12:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster2003.codfw.wmnet * 12:54 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster2003.codfw.wmnet * 12:54 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-staging-master-codfw * 12:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-staging-worker-eqiad * 12:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage1006.eqiad.wmnet * 12:52 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage1006.eqiad.wmnet * 12:46 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage1006.eqiad.wmnet * 12:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw1001.wikimedia.org with OS trixie * 12:45 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage1006.eqiad.wmnet * 12:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage1005.eqiad.wmnet * 12:45 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage1005.eqiad.wmnet * 12:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage1005.eqiad.wmnet * 12:36 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/ratelimit: apply * 12:35 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/ratelimit: apply * 12:35 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:35 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:33 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage1005.eqiad.wmnet * 12:33 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage1004.eqiad.wmnet * 12:33 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage1004.eqiad.wmnet * 12:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage1004.eqiad.wmnet * 12:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage1004.eqiad.wmnet * 12:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage1003.eqiad.wmnet * 12:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage1003.eqiad.wmnet * 12:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage * 12:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage1003.eqiad.wmnet * 12:17 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage * 12:14 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage1003.eqiad.wmnet * 12:14 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-staging-worker-eqiad * 12:03 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw1001.wikimedia.org with OS trixie * 12:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cuminunpriv1001.eqiad.wmnet * 11:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cuminunpriv1001.eqiad.wmnet * 11:27 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1004.wikimedia.org * 11:24 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1153 from dbctl [[phab:T434638|T434638]]', diff saved to https://phabricator.wikimedia.org/P96097 and previous config saved to /var/cache/conftool/dbconfig/20260814-112449-marostegui.json * 11:21 aokoth@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1004.wikimedia.org * 11:20 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 11:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-staging-worker-codfw * 11:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2004.codfw.wmnet * 11:17 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2004.codfw.wmnet * 11:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2004.codfw.wmnet * 11:10 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2004.codfw.wmnet * 11:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2003.codfw.wmnet * 11:10 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2003.codfw.wmnet * 11:03 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-worker1181.eqiad.wmnet with OS bookworm * 11:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2003.codfw.wmnet * 11:03 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2003.codfw.wmnet * 11:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2002.codfw.wmnet * 11:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2002.codfw.wmnet * 10:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1235.eqiad.wmnet with OS bookworm * 10:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock2003.codfw.wmnet * 10:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock2003.codfw.wmnet with OS trixie * 10:56 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2002.codfw.wmnet * 10:54 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1153.eqiad.wmnet with OS bookworm * 10:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2002.codfw.wmnet * 10:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2001.codfw.wmnet * 10:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2001.codfw.wmnet * 10:50 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 10:49 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1187.eqiad.wmnet with OS bookworm * 10:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2001.codfw.wmnet * 10:44 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2001.codfw.wmnet * 10:44 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-staging-worker-codfw * 10:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock2003.codfw.wmnet with reason: host reimage * 10:38 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock2003.codfw.wmnet with reason: host reimage * 10:35 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1153.eqiad.wmnet with reason: host reimage * 10:32 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1235.eqiad.wmnet with reason: host reimage * 10:29 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1187.eqiad.wmnet with reason: host reimage * 10:24 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1153.eqiad.wmnet with reason: host reimage * 10:23 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1235.eqiad.wmnet with reason: host reimage * 10:21 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1187.eqiad.wmnet with reason: host reimage * 10:16 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock2003.codfw.wmnet with OS trixie * 10:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install1005.wikimedia.org * 10:12 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 10:12 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 10:12 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 10:12 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 10:12 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:12 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 10:12 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 10:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install1005.wikimedia.org * 10:07 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1235.eqiad.wmnet with OS bookworm * 10:07 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1187.eqiad.wmnet with OS bookworm * 10:07 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1181.eqiad.wmnet with OS bookworm * 10:07 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1153.eqiad.wmnet with OS bookworm * 10:07 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install2005.wikimedia.org * 10:03 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 10:03 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2003.codfw.wmnet * 10:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1232.eqiad.wmnet with OS bookworm * 10:00 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install2005.wikimedia.org * 10:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install3004.wikimedia.org * 09:58 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 09:58 fceratto@cumin1003: Removing db1151 from zarcillo [[phab:T434538|T434538]] * 09:56 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1231.eqiad.wmnet with OS bookworm * 09:56 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 09:53 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 09:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install3004.wikimedia.org * 09:50 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1152.eqiad.wmnet with OS bookworm * 09:49 Dreamy_Jazz: `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260808000000" --end-timestamp="20260812120000" --sleep="5" --batch-size="50"` for [[phab:T434688|T434688]] * 09:48 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host centrallog2002.codfw.wmnet * 09:48 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install4004.wikimedia.org * 09:42 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1232.eqiad.wmnet with reason: host reimage * 09:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install4004.wikimedia.org * 09:41 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host centrallog2002.codfw.wmnet * 09:39 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install5004.wikimedia.org * 09:36 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1232.eqiad.wmnet with reason: host reimage * 09:36 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'. * 09:34 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'. * 09:33 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1231.eqiad.wmnet with reason: host reimage * 09:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install5004.wikimedia.org * 09:32 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host centrallog1002.eqiad.wmnet * 09:30 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install6003.wikimedia.org * 09:30 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1231.eqiad.wmnet with reason: host reimage * 09:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1152.eqiad.wmnet with reason: host reimage * 09:25 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host centrallog1002.eqiad.wmnet * 09:25 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1152.eqiad.wmnet with reason: host reimage * 09:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install6003.wikimedia.org * 09:22 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1232 * 09:22 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1232 * 09:22 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1232 * 09:22 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1232.eqiad.wmnet 25.53.64.10.in-addr.arpa 5.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:22 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host titan1001.eqiad.wmnet * 09:22 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1232.eqiad.wmnet 25.53.64.10.in-addr.arpa 5.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:22 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:22 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1232 - btullis@cumin1003" * 09:22 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1232 - btullis@cumin1003" * 09:17 btullis@cumin1003: START - Cookbook sre.dns.netbox * 09:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install7002.wikimedia.org * 09:17 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1232 * 09:16 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1231 * 09:16 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1231 * 09:14 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1231 * 09:14 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1231.eqiad.wmnet 24.53.64.10.in-addr.arpa 4.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host titan1001.eqiad.wmnet * 09:14 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1231.eqiad.wmnet 24.53.64.10.in-addr.arpa 4.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:14 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1231 - btullis@cumin1003" * 09:14 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1231 - btullis@cumin1003" * 09:10 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install7002.wikimedia.org * 09:09 btullis@cumin1003: START - Cookbook sre.dns.netbox * 09:08 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1231 * 09:08 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1152 * 09:08 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1152 * 09:06 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1152 * 09:06 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1152.eqiad.wmnet 16.53.64.10.in-addr.arpa 6.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:06 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1152.eqiad.wmnet 16.53.64.10.in-addr.arpa 6.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:06 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:06 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1152 - btullis@cumin1003" * 09:06 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1152 - btullis@cumin1003" * 09:03 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1151.eqiad.wmnet * 09:03 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:03 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1151.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 08:55 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host titan2001.codfw.wmnet * 08:55 btullis@cumin1003: START - Cookbook sre.dns.netbox * 08:54 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1151.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 08:54 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1152 * 08:53 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1232.eqiad.wmnet with OS bookworm * 08:53 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1231.eqiad.wmnet with OS bookworm * 08:53 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1152.eqiad.wmnet with OS bookworm * 08:51 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2209: Pool back * 08:51 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1201.eqiad.wmnet * 08:50 btullis@cumin1003: START - Cookbook sre.hosts.remove-downtime for an-worker1201.eqiad.wmnet * 08:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1229.eqiad.wmnet with OS bookworm * 08:47 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host titan2001.codfw.wmnet * 08:46 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping2004.codfw.wmnet * 08:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ping2004.codfw.wmnet * 08:40 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host an-worker1230.eqiad.wmnet with OS bookworm * 08:40 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 08:34 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1151.eqiad.wmnet * 08:34 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 08:31 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host titan1002.eqiad.wmnet * 08:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1229.eqiad.wmnet with reason: host reimage * 08:25 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host titan1002.eqiad.wmnet * 08:25 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1229.eqiad.wmnet with reason: host reimage * 08:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping1004.eqiad.wmnet * 08:21 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ping1004.eqiad.wmnet * 08:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1230.eqiad.wmnet with reason: host reimage * 08:12 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1230.eqiad.wmnet with reason: host reimage * 08:11 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1229 * 08:11 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1229 * 08:11 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1229 * 08:11 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1229.eqiad.wmnet 22.53.64.10.in-addr.arpa 2.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:11 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1229.eqiad.wmnet 22.53.64.10.in-addr.arpa 2.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:11 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:11 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1229 - btullis@cumin1003" * 08:11 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1229 - btullis@cumin1003" * 08:10 btullis@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1201.eqiad.wmnet with reason: Fixing a disk * 08:07 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host titan2002.codfw.wmnet * 08:05 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2209: Pool back * 08:05 btullis@cumin1003: START - Cookbook sre.dns.netbox * 08:00 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host titan2002.codfw.wmnet * 07:59 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1229 * 07:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1230 * 07:58 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1230 * 07:55 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1230 * 07:55 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1230.eqiad.wmnet 23.53.64.10.in-addr.arpa 3.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:55 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1230.eqiad.wmnet 23.53.64.10.in-addr.arpa 3.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:55 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:55 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1230 - btullis@cumin1003" * 07:55 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1230 - btullis@cumin1003" * 07:51 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host kubestagemaster2005.codfw.wmnet with OS trixie * 07:48 btullis@cumin1003: START - Cookbook sre.dns.netbox * 07:41 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1230 * 07:41 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1229.eqiad.wmnet with OS bookworm * 07:41 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1230.eqiad.wmnet with OS bookworm * 07:39 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1152 from dbctl [[phab:T434480|T434480]]', diff saved to https://phabricator.wikimedia.org/P96090 and previous config saved to /var/cache/conftool/dbconfig/20260814-073941-marostegui.json * 07:29 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on kubestagemaster2005.codfw.wmnet with reason: host reimage * 07:23 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on kubestagemaster2005.codfw.wmnet with reason: host reimage * 07:04 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host kubestagemaster2005.codfw.wmnet with OS trixie * 06:53 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1228.eqiad.wmnet with OS bookworm * 06:44 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1227.eqiad.wmnet with OS bookworm * 06:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1209.eqiad.wmnet with OS bookworm * 06:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1175.eqiad.wmnet with OS bookworm * 06:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1228.eqiad.wmnet with reason: host reimage * 06:27 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1228.eqiad.wmnet with reason: host reimage * 06:25 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1227.eqiad.wmnet with reason: host reimage * 06:21 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1227.eqiad.wmnet with reason: host reimage * 06:18 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1209.eqiad.wmnet with reason: host reimage * 06:14 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1209.eqiad.wmnet with reason: host reimage * 06:14 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1228 * 06:14 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1228 * 06:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1175.eqiad.wmnet with reason: host reimage * 06:12 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1228 * 06:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1228.eqiad.wmnet 20.53.64.10.in-addr.arpa 0.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:12 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1228.eqiad.wmnet 20.53.64.10.in-addr.arpa 0.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1228 - ryankemper@cumin2003" * 06:12 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1228 - ryankemper@cumin2003" * 06:09 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1175.eqiad.wmnet with reason: host reimage * 06:07 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 06:07 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1228 * 06:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1227 * 06:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1227 * 06:06 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1227 * 06:06 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1227.eqiad.wmnet 19.53.64.10.in-addr.arpa 9.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:06 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1227.eqiad.wmnet 19.53.64.10.in-addr.arpa 9.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:06 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:06 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1227 - ryankemper@cumin2003" * 06:06 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1227 - ryankemper@cumin2003" * 06:00 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 06:00 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1227 * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1209 * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1209 * 06:00 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1209 * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1209.eqiad.wmnet 15.53.64.10.in-addr.arpa 5.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:00 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1209.eqiad.wmnet 15.53.64.10.in-addr.arpa 5.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1209 - ryankemper@cumin2003" * 06:00 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1209 - ryankemper@cumin2003" * 05:54 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 05:54 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1209 * 05:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1175 * 05:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1175 * 05:52 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1175 * 05:52 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1175.eqiad.wmnet 17.53.64.10.in-addr.arpa 7.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 05:52 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1175.eqiad.wmnet 17.53.64.10.in-addr.arpa 7.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 05:52 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 05:52 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1175 - ryankemper@cumin2003" * 05:52 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1175 - ryankemper@cumin2003" * 05:49 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1228.eqiad.wmnet with OS bookworm * 05:49 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1227.eqiad.wmnet with OS bookworm * 05:48 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1209.eqiad.wmnet with OS bookworm * 05:47 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 05:47 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1175 * 05:47 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1175.eqiad.wmnet with OS bookworm * 05:09 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host apifeatureusage1001.eqiad.wmnet with OS bookworm * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 03s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:10 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 01:07 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 01:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 01:02 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 00:59 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1226.eqiad.wmnet with OS bookworm * 00:47 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1225.eqiad.wmnet with OS bookworm * 00:41 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1224.eqiad.wmnet with OS bookworm * 00:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1226.eqiad.wmnet with reason: host reimage * 00:31 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1226.eqiad.wmnet with reason: host reimage * 00:28 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1225.eqiad.wmnet with reason: host reimage * 00:25 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1225.eqiad.wmnet with reason: host reimage * 00:19 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1224.eqiad.wmnet with reason: host reimage * 00:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1226 * 00:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1226 * 00:17 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1226 * 00:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1226.eqiad.wmnet 23.36.64.10.in-addr.arpa 3.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:17 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1226.eqiad.wmnet 23.36.64.10.in-addr.arpa 3.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 00:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1226 - ryankemper@cumin2003" * 00:17 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1226 - ryankemper@cumin2003" * 00:16 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1224.eqiad.wmnet with reason: host reimage * 00:12 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 00:12 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1226 * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1225 * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1225 * 00:10 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1225 * 00:10 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1225.eqiad.wmnet 22.36.64.10.in-addr.arpa 2.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:10 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1225.eqiad.wmnet 22.36.64.10.in-addr.arpa 2.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:10 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 00:10 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1225 - ryankemper@cumin2003" * 00:10 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1225 - ryankemper@cumin2003" * 00:03 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 00:02 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1225 * 00:02 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1224 * 00:02 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1224 * 00:00 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1224 * 00:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1224.eqiad.wmnet 21.36.64.10.in-addr.arpa 1.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:00 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1224.eqiad.wmnet 21.36.64.10.in-addr.arpa 1.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 00:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1224 - ryankemper@cumin2003" * 00:00 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1224 - ryankemper@cumin2003" == 2026-08-13 == * 23:54 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1226.eqiad.wmnet with OS bookworm * 23:53 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1225.eqiad.wmnet with OS bookworm * 23:52 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 23:51 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1224 * 23:51 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1224.eqiad.wmnet with OS bookworm * 23:47 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1223.eqiad.wmnet with OS bookworm * 23:29 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1223.eqiad.wmnet with reason: host reimage * 23:24 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1223.eqiad.wmnet with reason: host reimage * 23:19 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325525{{!}}ve.ui.CodeMirror.less: ensure normal font style]] (duration: 11m 40s) * 23:16 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260807000000" --end-timestamp="20260808000000" --sleep="5" --batch-size="50"` for [[phab:T434688|T434688]] * 23:13 musikanimal@deploy1003: musikanimal: Continuing with deployment * 23:11 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1325525{{!}}ve.ui.CodeMirror.less: ensure normal font style]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1223 * 23:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1223 * 23:08 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1325525{{!}}ve.ui.CodeMirror.less: ensure normal font style]] * 23:07 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1223 * 23:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1223.eqiad.wmnet 20.36.64.10.in-addr.arpa 0.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:07 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1223.eqiad.wmnet 20.36.64.10.in-addr.arpa 0.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 23:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1223 - ryankemper@cumin2003" * 23:03 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1223 - ryankemper@cumin2003" * 23:01 sbassett: Deployed security updates for [[phab:T430596|T430596]], [[phab:T120386|T120386]] * 22:55 bking@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host kubestagemaster2005.codfw.wmnet * 22:55 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host kubestagemaster2005.codfw.wmnet with OS bookworm * 22:54 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 22:53 ryankemper@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 22:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on kubestagemaster2005.codfw.wmnet with reason: host reimage * 22:49 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1222.eqiad.wmnet with OS bookworm * 22:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on kubestagemaster2005.codfw.wmnet with reason: host reimage * 22:29 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1222.eqiad.wmnet with reason: host reimage * 22:26 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1222.eqiad.wmnet with reason: host reimage * 22:23 bking@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host aux-k8s-etcd2003.codfw.wmnet * 22:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd2003.codfw.wmnet with OS bookworm * 22:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host kubestagemaster2005.codfw.wmnet with OS bookworm * 22:23 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM kubestagemaster2005.codfw.wmnet - bking@cumin2003" * 22:23 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM kubestagemaster2005.codfw.wmnet - bking@cumin2003" * 22:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) kubestagemaster2005.codfw.wmnet on all recursors * 22:22 bking@cumin2003: START - Cookbook sre.dns.wipe-cache kubestagemaster2005.codfw.wmnet on all recursors * 22:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:22 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM kubestagemaster2005.codfw.wmnet - bking@cumin2003" * 22:22 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM kubestagemaster2005.codfw.wmnet - bking@cumin2003" * 22:17 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 22:13 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1223 * 22:12 bking@cumin2003: START - Cookbook sre.dns.netbox * 22:12 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host kubestagemaster2005.codfw.wmnet * 22:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1222 * 22:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1222 * 22:11 sbassett: Deployed security updates for [[phab:T429244|T429244]], [[phab:T434039|T434039]] * 22:11 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1222 * 22:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1222.eqiad.wmnet 19.36.64.10.in-addr.arpa 9.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:11 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1222.eqiad.wmnet 19.36.64.10.in-addr.arpa 9.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1222 - ryankemper@cumin2003" * 22:07 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1222 - ryankemper@cumin2003" * 22:05 robh@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-wdqs2001.codfw.wmnet with reason: updating firmware * 22:01 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 22:01 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1223.eqiad.wmnet with OS bookworm * 22:01 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1222 * 22:01 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1222.eqiad.wmnet with OS bookworm * 22:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1212.eqiad.wmnet with OS bookworm * 21:54 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324295{{!}}Revert "Lazily reject pre-fix parser-cache entries for noreferrer/noopener links" (T429090)]] (duration: 06m 42s) * 21:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd2003.codfw.wmnet with reason: host reimage * 21:53 bking@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host dse-k8s-etcd2001.codfw.wmnet * 21:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dse-k8s-etcd2001.codfw.wmnet with OS bookworm * 21:50 sbassett@deploy1003: sbassett, kharlan: Continuing with deployment * 21:49 sbassett@deploy1003: sbassett, kharlan: Backport for [[gerrit:1324295{{!}}Revert "Lazily reject pre-fix parser-cache entries for noreferrer/noopener links" (T429090)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:48 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aux-k8s-etcd2003.codfw.wmnet with reason: host reimage * 21:47 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1324295{{!}}Revert "Lazily reject pre-fix parser-cache entries for noreferrer/noopener links" (T429090)]] * 21:42 maryum: Deployed security patch for [[phab:T434549|T434549]] * 21:39 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1212.eqiad.wmnet with reason: host reimage * 21:34 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1212.eqiad.wmnet with reason: host reimage * 21:33 bking@cumin2003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd2003.codfw.wmnet with OS bookworm * 21:32 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM aux-k8s-etcd2003.codfw.wmnet - bking@cumin2003" * 21:32 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM aux-k8s-etcd2003.codfw.wmnet - bking@cumin2003" * 21:32 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) aux-k8s-etcd2003.codfw.wmnet on all recursors * 21:32 bking@cumin2003: START - Cookbook sre.dns.wipe-cache aux-k8s-etcd2003.codfw.wmnet on all recursors * 21:32 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:32 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM aux-k8s-etcd2003.codfw.wmnet - bking@cumin2003" * 21:31 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM aux-k8s-etcd2003.codfw.wmnet - bking@cumin2003" * 21:28 maryum: Deployed security patch for [[phab:T434619|T434619]] * 21:26 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:26 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host aux-k8s-etcd2003.codfw.wmnet * 21:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-etcd2001.codfw.wmnet with reason: host reimage * 21:20 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1212 * 21:20 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1212 * 21:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1212.eqiad.wmnet with OS bookworm * 21:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on dse-k8s-etcd2001.codfw.wmnet with reason: host reimage * 21:04 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325553{{!}}InstrumentConstructiveEdits: exclude mw-reverted as well (T431493)]] (duration: 06m 25s) * 20:59 kemayo@deploy1003: kemayo: Continuing with deployment * 20:59 kemayo@deploy1003: kemayo: Backport for [[gerrit:1325553{{!}}InstrumentConstructiveEdits: exclude mw-reverted as well (T431493)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host dse-k8s-etcd2001.codfw.wmnet with OS bookworm * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 20:57 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 20:57 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1325553{{!}}InstrumentConstructiveEdits: exclude mw-reverted as well (T431493)]] * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-etcd2001.codfw.wmnet on all recursors * 20:57 bking@cumin2003: START - Cookbook sre.dns.wipe-cache dse-k8s-etcd2001.codfw.wmnet on all recursors * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 20:57 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 20:54 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320163{{!}}Add configurable RestTermsOfServiceUrl (T428147)]] (duration: 21m 39s) * 20:53 bking@cumin2003: START - Cookbook sre.dns.netbox * 20:53 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host dse-k8s-etcd2001.codfw.wmnet * 20:50 samtar@deploy1003: samtar, milazg: Continuing with deployment * 20:35 samtar@deploy1003: samtar, milazg: Backport for [[gerrit:1320163{{!}}Add configurable RestTermsOfServiceUrl (T428147)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:33 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1320163{{!}}Add configurable RestTermsOfServiceUrl (T428147)]] * 20:30 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325549{{!}}Deploy PRV to several LC wikis (T423785)]] (duration: 06m 57s) * 20:26 arlolra@deploy1003: arlolra: Continuing with deployment * 20:25 arlolra@deploy1003: arlolra: Backport for [[gerrit:1325549{{!}}Deploy PRV to several LC wikis (T423785)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:24 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 20:23 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1325549{{!}}Deploy PRV to several LC wikis (T423785)]] * 20:21 ariel@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319906{{!}}Remove boilerplate language from wmf-rest and wmf-math API modules (T433736)]] (duration: 13m 54s) * 20:14 ariel@deploy1003: ariel: Continuing with deployment * 20:11 ariel@deploy1003: ariel: Backport for [[gerrit:1319906{{!}}Remove boilerplate language from wmf-rest and wmf-math API modules (T433736)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 ariel@deploy1003: Started scap sync-world: Backport for [[gerrit:1319906{{!}}Remove boilerplate language from wmf-rest and wmf-math API modules (T433736)]] * 19:58 robh@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-wdqs2001.codfw.wmnet with reason: updating firmware * 19:54 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325551{{!}}Render the focused module view as a full-screen page (T433896)]] (duration: 30m 37s) * 19:52 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching aqs[2001,1016]*: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 19:44 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching aqs[2001,1016]*: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 19:42 musikanimal@deploy1003: musikanimal: Continuing with deployment * 19:41 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1325551{{!}}Render the focused module view as a full-screen page (T433896)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:34 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: quash java safepoint logspam - bking@cumin2003 - [[phab:T434685|T434685]] * 19:34 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 19:34 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 19:24 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1325551{{!}}Render the focused module view as a full-screen page (T433896)]] * 19:20 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 19:20 jhancock@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin1003" * 19:18 jhancock@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin1003" * 19:14 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325548{{!}}Enable image lazy loading on desktop in group1 (T148047)]] (duration: 07m 43s) * 19:10 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 19:10 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1325548{{!}}Enable image lazy loading on desktop in group1 (T148047)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:07 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1325548{{!}}Enable image lazy loading on desktop in group1 (T148047)]] * 19:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1166.eqiad.wmnet onto db1280.eqiad.wmnet * 19:03 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 19:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1166: Pool db1166.eqiad.wmnet in after cloning * 18:59 jhancock@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 18:58 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1234.eqiad.wmnet with OS bookworm * 18:52 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 18:51 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:49 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:46 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 18:46 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 18:44 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=dns3004.* * 18:39 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1221.eqiad.wmnet with OS bookworm * 18:36 inflatador: [bking@ganeti2048] ~$ sudo gnt-instance replace-disks -n ganeti2030.codfw.wmnet aux-k8s-worker2002.codfw.wmnet [[phab:T434681|T434681]] * 18:35 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 18:29 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1220.eqiad.wmnet with OS bookworm * 18:29 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1234.eqiad.wmnet with reason: host reimage * 18:27 bking@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host dse-k8s-etcd2001.codfw.wmnet * 18:27 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-etcd2001.codfw.wmnet on all recursors * 18:27 bking@cumin2003: START - Cookbook sre.dns.wipe-cache dse-k8s-etcd2001.codfw.wmnet on all recursors * 18:27 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:27 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 18:27 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 18:25 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1234.eqiad.wmnet with reason: host reimage * 18:25 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 18:19 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1221.eqiad.wmnet with reason: host reimage * 18:17 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1166: Pool db1166.eqiad.wmnet in after cloning * 18:15 brennen@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 18:15 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1221.eqiad.wmnet with reason: host reimage * 18:14 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:12 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:12 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 18:12 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-etcd2001.codfw.wmnet on all recursors * 18:12 bking@cumin2003: START - Cookbook sre.dns.wipe-cache dse-k8s-etcd2001.codfw.wmnet on all recursors * 18:12 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:12 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 18:12 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 18:10 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: quash java safepoint logspam - bking@cumin2003 - [[phab:T434685|T434685]] * 18:10 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1234 * 18:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1234 * 18:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1220.eqiad.wmnet with reason: host reimage * 18:07 brennen: 1.47.0-wmf.15 train status ([[phab:T430834|T430834]]) - no current blockers, rolling to all wikis * 18:07 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1234 * 18:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1234.eqiad.wmnet 10.36.64.10.in-addr.arpa 0.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:07 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:07 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1234.eqiad.wmnet 10.36.64.10.in-addr.arpa 0.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1234 - ryankemper@cumin2003" * 18:07 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1234 - ryankemper@cumin2003" * 18:05 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1220.eqiad.wmnet with reason: host reimage * 18:04 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host dse-k8s-etcd2001.codfw.wmnet * 18:04 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 18:03 dancy@deploy1003: Installation of scap version "4.280.1" completed for 3 hosts * 18:02 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:02 inflatador: bking@dse-k8s-etcd2002 etcdctl member remove $<nowiki>{</nowiki>UUID of dse-k8s-etcd2001<nowiki>}</nowiki> [[phab:T434681|T434681]] [[phab:T434793|T434793]] * 18:01 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 18:01 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1234 * 18:01 dancy@deploy1003: Installing scap version "4.280.1" for 3 host(s) * 18:01 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1221 * 18:01 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1221 * 18:00 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1221 * 18:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1221.eqiad.wmnet 18.36.64.10.in-addr.arpa 8.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:00 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1221.eqiad.wmnet 18.36.64.10.in-addr.arpa 8.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1221 - ryankemper@cumin2003" * 17:59 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:58 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 17:58 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:57 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 17:57 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:56 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1221 - ryankemper@cumin2003" * 17:53 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 17:52 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 17:52 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 17:51 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1221 * 17:51 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1220 * 17:51 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1220 * 17:51 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:51 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:51 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1220 * 17:51 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1220.eqiad.wmnet 11.36.64.10.in-addr.arpa 1.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:51 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1220.eqiad.wmnet 11.36.64.10.in-addr.arpa 1.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:51 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1220 - ryankemper@cumin2003" * 17:50 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1220 - ryankemper@cumin2003" * 17:47 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:47 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 17:46 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1234.eqiad.wmnet with OS bookworm * 17:46 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 17:45 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1221.eqiad.wmnet with OS bookworm * 17:45 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1220 * 17:45 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1220.eqiad.wmnet with OS bookworm * 17:43 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 17:41 bking@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host dse-k8s-etcd2001.codfw.wmnet * 17:41 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host dse-k8s-etcd2001.codfw.wmnet * 17:40 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:40 inflatador: bking@ganeti2048] `sudo gnt-instance remove --force --ignore-failures --shutdown-timeout=0` on non-DRBD VMs [[phab:T434681|T434681]] * 17:40 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:39 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:38 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:36 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:32 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:32 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:28 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 17:26 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:26 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:24 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:23 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:21 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1218.eqiad.wmnet with OS bookworm * 17:20 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:20 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:19 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1179.eqiad.wmnet with OS bookworm * 17:18 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1150.eqiad.wmnet with OS bookworm * 17:18 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns3004.wikimedia.org with OS trixie * 17:15 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:14 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:13 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:12 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:12 swfrench@deploy1003: Finished scap sync-world: Helmfile-only deployment for mediawiki chart bump - [[phab:T427666|T427666]] (duration: 03m 03s) * 17:09 swfrench@deploy1003: Started scap sync-world: Helmfile-only deployment for mediawiki chart bump - [[phab:T427666|T427666]] * 17:02 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:01 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:01 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1218.eqiad.wmnet with reason: host reimage * 17:01 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:01 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:00 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324810{{!}}deployment-info.php: Report dbname and branch for the requested wiki (T434726)]] (duration: 06m 52s) * 16:58 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 16:57 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1179.eqiad.wmnet with reason: host reimage * 16:56 dancy@deploy1003: dancy: Continuing with deployment * 16:56 dancy@deploy1003: dancy: Backport for [[gerrit:1324810{{!}}deployment-info.php: Report dbname and branch for the requested wiki (T434726)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1150.eqiad.wmnet with reason: host reimage * 16:53 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324810{{!}}deployment-info.php: Report dbname and branch for the requested wiki (T434726)]] * 16:51 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1166: Depool db1166.eqiad.wmnet to then clone it to db1280.eqiad.wmnet - cwilliams@cumin1003 * 16:50 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1166: Depool db1166.eqiad.wmnet to then clone it to db1280.eqiad.wmnet - cwilliams@cumin1003 * 16:50 cwilliams@cumin1003: START - Cookbook sre.mysql.clone of db1166.eqiad.wmnet onto db1280.eqiad.wmnet * 16:49 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1218.eqiad.wmnet with reason: host reimage * 16:48 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1179.eqiad.wmnet with reason: host reimage * 16:47 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1150.eqiad.wmnet with reason: host reimage * 16:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.upgrade (exit_code=0) for 1 hosts * 16:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2220: Upgrade of db2220.codfw.wmnet completed * 16:38 swfrench@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 16:38 swfrench-wmf: kubectl delete node kubestagemaster2005.codfw.wmnet - [[phab:T434681|T434681]] * 16:34 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1218 * 16:34 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1218 * 16:34 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1218.eqiad.wmnet with OS bookworm * 16:34 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1179 * 16:34 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1179 * 16:33 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1179.eqiad.wmnet with OS bookworm * 16:32 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1150 * 16:32 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1150 * 16:31 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1150.eqiad.wmnet with OS bookworm * 16:24 dancy@deploy1003: Installation of scap version "4.280.0" completed for 3 hosts * 16:22 dancy@deploy1003: Installing scap version "4.280.0" for 3 host(s) * 16:20 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: quash java safepoint logspam - bking@cumin2003 - [[phab:T434685|T434685]] * 16:14 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns3004.wikimedia.org with reason: host reimage * 16:08 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 16:07 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 16:07 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns3004.wikimedia.org with reason: host reimage * 16:05 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 16:04 swfrench@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 16:00 swfrench@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 15:59 swfrench@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 15:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Upgrade of db2220.codfw.wmnet completed * 15:48 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2220: Upgrading db2220.codfw.wmnet * 15:48 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2220: Upgrading db2220.codfw.wmnet * 15:48 cwilliams@cumin1003: START - Cookbook sre.mysql.upgrade for 1 hosts * 15:46 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns3004.wikimedia.org with OS trixie * 15:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2220 [[phab:T434802|T434802]]', diff saved to https://phabricator.wikimedia.org/P96079 and previous config saved to /var/cache/conftool/dbconfig/20260813-154624-cwilliams.json * 15:45 cdobbins@cumin1003: conftool action : set/pooled=no; selector: name=dns3004.* * 15:44 cjd91: depooling dns3004 to reimage to trixie * 15:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2159 to s7 primary [[phab:T434802|T434802]]', diff saved to https://phabricator.wikimedia.org/P96078 and previous config saved to /var/cache/conftool/dbconfig/20260813-154405-cwilliams.json * 15:43 cezmunsta: Starting s7 codfw failover from db2220 to db2159 - [[phab:T434802|T434802]] * 15:41 cgoubert@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: Dragonfly supernodes reboot (duration: 09m 42s) * 15:41 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dragonfly-supernode2001.codfw.wmnet * 15:39 inflatador: bking@ganeti2048] ~$ sudo gnt-node failover -f --ignore-consistency ganeti2046.codfw.wmnet [[phab:T434681|T434681]] * 15:39 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 15:39 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 15:39 swfrench@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 15:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2159 with weight 0 [[phab:T434802|T434802]]', diff saved to https://phabricator.wikimedia.org/P96077 and previous config saved to /var/cache/conftool/dbconfig/20260813-153806-cwilliams.json * 15:37 swfrench@dns1004: END - running authdns-update * 15:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 30 hosts with reason: Primary switchover s7 [[phab:T434802|T434802]] * 15:37 cgoubert@cumin2003: START - Cookbook sre.hosts.reboot-single for host dragonfly-supernode2001.codfw.wmnet * 15:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dragonfly-supernode1001.eqiad.wmnet * 15:35 swfrench@dns1004: START - running authdns-update * 15:32 cgoubert@cumin2003: START - Cookbook sre.hosts.reboot-single for host dragonfly-supernode1001.eqiad.wmnet * 15:32 cgoubert@deploy1003: Locking from deployment [ALL REPOSITORIES]: Dragonfly supernodes reboot * 15:30 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-master-codfw * 15:30 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2005.codfw.wmnet * 15:30 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2005.codfw.wmnet * 15:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1169.eqiad.wmnet onto db1277.eqiad.wmnet * 15:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1169: Pool db1169.eqiad.wmnet in after cloning * 15:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2005.codfw.wmnet * 15:23 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2005.codfw.wmnet * 15:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2004.codfw.wmnet * 15:23 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2004.codfw.wmnet * 15:18 inflatador: bking@ganeti2048 sudo gnt-node failover -f ganeti2046.codfw.wmnet [[phab:T434681|T434681]] * 15:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2004.codfw.wmnet * 15:17 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2004.codfw.wmnet * 15:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2003.codfw.wmnet * 15:17 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2003.codfw.wmnet * 15:15 cdobbins@dns1004: END - running authdns-update * 15:13 cdobbins@dns1004: START - running authdns-update * 15:10 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1217.eqiad.wmnet with OS bookworm * 15:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2003.codfw.wmnet * 15:10 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2003.codfw.wmnet * 15:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2002.codfw.wmnet * 15:10 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2002.codfw.wmnet * 15:10 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: quash java safepoint logspam - bking@cumin2003 - [[phab:T434685|T434685]] * 15:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1216.eqiad.wmnet with OS bookworm * 15:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2002.codfw.wmnet * 15:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2002.codfw.wmnet * 15:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2001.codfw.wmnet * 15:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2001.codfw.wmnet * 15:01 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1042.eqiad.wmnet * 15:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1042.eqiad.wmnet * 15:00 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325455{{!}}Api: Use correct query when continue prop=categories (T433922)]] (duration: 09m 47s) * 14:59 bking@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host dse-k8s-etcd2004.codfw.wmnet * 14:58 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-etcd2004.codfw.wmnet on all recursors * 14:58 bking@cumin2003: START - Cookbook sre.dns.wipe-cache dse-k8s-etcd2004.codfw.wmnet on all recursors * 14:58 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:58 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM dse-k8s-etcd2004.codfw.wmnet - bking@cumin2003" * 14:58 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM dse-k8s-etcd2004.codfw.wmnet - bking@cumin2003" * 14:55 zabe@deploy1003: zabe: Continuing with deployment * 14:53 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-etcd2004.codfw.wmnet on all recursors * 14:53 bking@cumin2003: START - Cookbook sre.dns.wipe-cache dse-k8s-etcd2004.codfw.wmnet on all recursors * 14:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:53 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2004.codfw.wmnet - bking@cumin2003" * 14:53 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2004.codfw.wmnet - bking@cumin2003" * 14:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2001.codfw.wmnet * 14:52 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2001.codfw.wmnet * 14:52 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-master-codfw * 14:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2001.codfw.wmnet * 14:52 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2001.codfw.wmnet * 14:52 zabe@deploy1003: zabe: Backport for [[gerrit:1325455{{!}}Api: Use correct query when continue prop=categories (T433922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:50 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1325455{{!}}Api: Use correct query when continue prop=categories (T433922)]] * 14:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1217.eqiad.wmnet with reason: host reimage * 14:48 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:48 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host dse-k8s-etcd2004.codfw.wmnet * 14:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-master-eqiad * 14:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1006.eqiad.wmnet * 14:44 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1006.eqiad.wmnet * 14:44 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1216.eqiad.wmnet with reason: host reimage * 14:40 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1217.eqiad.wmnet with reason: host reimage * 14:39 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1169: Pool db1169.eqiad.wmnet in after cloning * 14:39 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1216.eqiad.wmnet with reason: host reimage * 14:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl1005.eqiad.wmnet * 14:32 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl1005.eqiad.wmnet * 14:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1004.eqiad.wmnet * 14:32 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1004.eqiad.wmnet * 14:31 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus2008.codfw.wmnet * 14:31 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor1003.eqiad.wmnet * 14:29 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1042.eqiad.wmnet * 14:27 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor1003.eqiad.wmnet * 14:26 moritzm: installing Django security updates * 14:26 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling reboot on A:wikidough * 14:25 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1217.eqiad.wmnet with OS bookworm * 14:25 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1216.eqiad.wmnet with OS bookworm * 14:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl1004.eqiad.wmnet * 14:24 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl1004.eqiad.wmnet * 14:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1003.eqiad.wmnet * 14:24 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1003.eqiad.wmnet * 14:23 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus2008.codfw.wmnet * 14:23 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus1008.eqiad.wmnet * 14:23 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor-dev2001.codfw.wmnet * 14:22 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus2006.codfw.wmnet * 14:19 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor-dev2001.codfw.wmnet * 14:18 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1042.eqiad.wmnet * 14:17 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.* * 14:16 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl1003.eqiad.wmnet * 14:16 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl1003.eqiad.wmnet * 14:16 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1002.eqiad.wmnet * 14:16 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1002.eqiad.wmnet * 14:15 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus1008.eqiad.wmnet * 14:14 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus1006.eqiad.wmnet * 14:12 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus2006.codfw.wmnet * 14:12 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor2003.codfw.wmnet * 14:11 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus2007.codfw.wmnet * 14:11 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1041.eqiad.wmnet * 14:11 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1041.eqiad.wmnet * 14:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl1002.eqiad.wmnet * 14:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl1002.eqiad.wmnet * 14:09 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-master-eqiad * 14:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on apifeatureusage1001.eqiad.wmnet with reason: host reimage * 14:08 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1215.eqiad.wmnet with OS bookworm * 14:08 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor2003.codfw.wmnet * 14:08 jayme: updated calico to v3.30.7 on wikikube codfw [[phab:T427400|T427400]] * 14:07 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1149.eqiad.wmnet with OS bookworm * 14:06 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'. * 14:06 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1041.eqiad.wmnet * 14:04 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus1006.eqiad.wmnet * 14:03 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus2007.codfw.wmnet * 14:03 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus2005.codfw.wmnet * 14:03 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus1007.eqiad.wmnet * 14:03 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1002.eqiad.wmnet * 14:02 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on apifeatureusage1001.eqiad.wmnet with reason: host reimage * 14:02 cgoubert@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=helm-charts.*,name=eqiad * 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host chartmuseum1001.eqiad.wmnet * 14:00 moritzm: installing libxml2 security updates * 13:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1214.eqiad.wmnet with OS bookworm * 13:59 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns5003.* * 13:59 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1041.eqiad.wmnet * 13:59 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow1002.eqiad.wmnet * 13:58 cmooney@dns3003: END - running authdns-update * 13:57 cgoubert@cumin2003: START - Cookbook sre.hosts.reboot-single for host chartmuseum1001.eqiad.wmnet * 13:57 cgoubert@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=helm-charts.*,name=eqiad * 13:57 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: cloudelastic cluster restart - bking@cumin2003 * 13:57 cgoubert@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=helm-charts.*,name=codfw * 13:56 cmooney@dns3003: START - running authdns-update * 13:56 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'. * 13:56 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns5003.*,service=authdns-update * 13:55 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host chartmuseum2001.codfw.wmnet * 13:55 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus1007.eqiad.wmnet * 13:55 cmooney@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dns5003.wikimedia.org * 13:55 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus1005.eqiad.wmnet * 13:51 cgoubert@cumin2003: START - Cookbook sre.hosts.reboot-single for host chartmuseum2001.codfw.wmnet * 13:51 cgoubert@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=helm-charts.*,name=codfw * 13:51 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus2005.codfw.wmnet * 13:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host apifeatureusage1001.eqiad.wmnet with OS bookworm * 13:50 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus7002.magru.wmnet * 13:50 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts lvs1015.eqiad.wmnet * 13:50 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:50 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1015.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:49 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1015.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:49 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325480{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] (duration: 06m 39s) * 13:47 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1215.eqiad.wmnet with reason: host reimage * 13:46 cmooney@cumin1003: START - Cookbook sre.hosts.reboot-single for host dns5003.wikimedia.org * 13:46 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1003.eqiad.wmnet * 13:45 cmooney@cumin1003: conftool action : set/pooled=no; selector: name=dns5003.* * 13:45 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1040.eqiad.wmnet * 13:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1040.eqiad.wmnet * 13:45 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 13:45 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus1005.eqiad.wmnet * 13:45 stran@deploy1003: stran: Continuing with deployment * 13:44 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus7002.magru.wmnet * 13:44 stran@deploy1003: stran: Backport for [[gerrit:1325480{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:44 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus6002.drmrs.wmnet * 13:43 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260805000000" --end-timestamp="20260806000000" --sleep="5" --batch-size="10"` for [[phab:T434688|T434688]] * 13:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1149.eqiad.wmnet with reason: host reimage * 13:42 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1325480{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] * 13:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow1003.eqiad.wmnet * 13:40 sukhe@cumin1003: START - Cookbook sre.hosts.decommission for hosts lvs1015.eqiad.wmnet * 13:40 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts lvs1014.eqiad.wmnet * 13:40 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:40 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1014.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:40 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1040.eqiad.wmnet * 13:40 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1014.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:39 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1214.eqiad.wmnet with reason: host reimage * 13:38 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2004.codfw.wmnet * 13:38 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus6002.drmrs.wmnet * 13:38 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus5003.eqsin.wmnet * 13:35 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 13:35 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1149.eqiad.wmnet with reason: host reimage * 13:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1215.eqiad.wmnet with reason: host reimage * 13:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow2004.codfw.wmnet * 13:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1214.eqiad.wmnet with reason: host reimage * 13:33 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1040.eqiad.wmnet * 13:31 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus5003.eqsin.wmnet * 13:31 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: cloudelastic cluster restart - bking@cumin2003 * 13:31 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus4003.ulsfo.wmnet * 13:30 sukhe@cumin1003: START - Cookbook sre.hosts.decommission for hosts lvs1014.eqiad.wmnet * 13:30 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts lvs1013.eqiad.wmnet * 13:30 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:30 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1013.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:30 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1013.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:27 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1039.eqiad.wmnet * 13:27 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1039.eqiad.wmnet * 13:26 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1167.eqiad.wmnet onto db1281.eqiad.wmnet * 13:26 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1167: Pool db1167.eqiad.wmnet in after cloning * 13:25 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus4003.ulsfo.wmnet * 13:25 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260802000000" --end-timestamp="20260803000000" --sleep="5" --batch-size="10"` for [[phab:T434688|T434688]] * 13:24 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 13:24 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus3004.esams.wmnet * 13:24 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1274: New host * 13:24 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325476{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] (duration: 07m 19s) * 13:24 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260804000000" --end-timestamp="20260805000000" --sleep="5" --batch-size="10"` for [[phab:T434688|T434688]] * 13:24 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260803000000" --end-timestamp="20260804000000" --sleep="5" --batch-size="10"` for [[phab:T434688|T434688]] * 13:23 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2003.codfw.wmnet * 13:22 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1039.eqiad.wmnet * 13:20 sukhe@cumin1003: START - Cookbook sre.hosts.decommission for hosts lvs1013.eqiad.wmnet * 13:20 stran@deploy1003: stran: Continuing with deployment * 13:19 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow2003.codfw.wmnet * 13:19 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling reboot on A:wikidough * 13:19 stran@deploy1003: stran: Backport for [[gerrit:1325476{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:18 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus3004.esams.wmnet * 13:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1215.eqiad.wmnet with OS bookworm * 13:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1214.eqiad.wmnet with OS bookworm * 13:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1149.eqiad.wmnet with OS bookworm * 13:17 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1325476{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] * 13:11 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324731{{!}}prv: Enable parsoid rendering for 5 wikisource wikis]] (duration: 07m 26s) * 13:11 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1039.eqiad.wmnet * 13:11 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1036.eqiad.wmnet * 13:11 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1036.eqiad.wmnet * 13:07 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow3004.esams.wmnet * 13:07 jgiannelos@deploy1003: jgiannelos: Continuing with deployment * 13:06 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1233.eqiad.wmnet with OS bookworm * 13:06 jgiannelos@deploy1003: jgiannelos: Backport for [[gerrit:1324731{{!}}prv: Enable parsoid rendering for 5 wikisource wikis]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:04 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1324731{{!}}prv: Enable parsoid rendering for 5 wikisource wikis]] * 13:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow3004.esams.wmnet * 13:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1036.eqiad.wmnet * 13:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1210.eqiad.wmnet with OS bookworm * 13:01 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1036.eqiad.wmnet * 13:00 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1052.eqiad.wmnet * 13:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1052.eqiad.wmnet * 12:58 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow4003.ulsfo.wmnet * 12:55 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1211.eqiad.wmnet with OS bookworm * 12:55 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1052.eqiad.wmnet * 12:52 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow4003.ulsfo.wmnet * 12:51 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1052.eqiad.wmnet * 12:49 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1169: Depool db1169.eqiad.wmnet to then clone it to db1277.eqiad.wmnet - cwilliams@cumin1003 * 12:46 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1169: Depool db1169.eqiad.wmnet to then clone it to db1277.eqiad.wmnet - cwilliams@cumin1003 * 12:46 cwilliams@cumin1003: START - Cookbook sre.mysql.clone of db1169.eqiad.wmnet onto db1277.eqiad.wmnet * 12:45 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:45 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns record for deleted IP reservations eqsin lvs vlan ints - cmooney@cumin1003" * 12:44 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns record for deleted IP reservations eqsin lvs vlan ints - cmooney@cumin1003" * 12:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1051.eqiad.wmnet * 12:43 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1051.eqiad.wmnet * 12:42 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1210.eqiad.wmnet with reason: host reimage * 12:41 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1167: Pool db1167.eqiad.wmnet in after cloning * 12:40 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 12:39 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1274: New host * 12:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Added db1274', diff saved to https://phabricator.wikimedia.org/P96062 and previous config saved to /var/cache/conftool/dbconfig/20260813-123907-cwilliams.json * 12:38 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast7002.wikimedia.org * 12:38 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1233.eqiad.wmnet with reason: host reimage * 12:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1051.eqiad.wmnet * 12:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1211.eqiad.wmnet with reason: host reimage * 12:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1233.eqiad.wmnet with reason: host reimage * 12:32 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1051.eqiad.wmnet * 12:32 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast7002.wikimedia.org * 12:32 marostegui: Drop SecurePoll tables from closed wikis [[phab:T423128|T423128]] * 12:32 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow5003.eqsin.wmnet * 12:31 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1050.eqiad.wmnet * 12:31 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1050.eqiad.wmnet * 12:29 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1210.eqiad.wmnet with reason: host reimage * 12:29 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1211.eqiad.wmnet with reason: host reimage * 12:29 moritzm: installing Wireshark security updates * 12:26 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow5003.eqsin.wmnet * 12:26 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1050.eqiad.wmnet * 12:24 cmooney@dns3003: END - running authdns-update * 12:21 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow6001.drmrs.wmnet * 12:21 cmooney@dns3003: START - running authdns-update * 12:20 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1050.eqiad.wmnet * 12:20 cgoubert@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host rdb-lock2003.codfw.wmnet * 12:20 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:20 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update netbox dns entries for expanded public1-603-eqsin subnet - cmooney@cumin1003" * 12:20 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update netbox dns entries for expanded public1-603-eqsin subnet - cmooney@cumin1003" * 12:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 12:20 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 12:19 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1049.eqiad.wmnet * 12:19 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1049.eqiad.wmnet * 12:19 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 12:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host moss-be1003.eqiad.wmnet * 12:17 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow6001.drmrs.wmnet * 12:17 moritzm: installin curl security updates * 12:15 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 12:15 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1233.eqiad.wmnet with OS bookworm * 12:15 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1210.eqiad.wmnet with OS bookworm * 12:15 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1211.eqiad.wmnet with OS bookworm * 12:13 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1049.eqiad.wmnet * 12:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow7002.magru.wmnet * 12:12 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-worker1178.eqiad.wmnet with OS bookworm * 12:10 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host moss-be1003.eqiad.wmnet * 12:10 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 12:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be1006.eqiad.wmnet * 12:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 12:10 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 12:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 12:10 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 12:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow7002.magru.wmnet * 12:08 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1049.eqiad.wmnet * 12:05 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 12:05 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2003.codfw.wmnet * 12:04 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'. * 12:04 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:03 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be1006.eqiad.wmnet * 12:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be1005.eqiad.wmnet * 12:01 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1038.eqiad.wmnet * 12:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1038.eqiad.wmnet * 12:01 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1213.eqiad.wmnet with OS bookworm * 11:56 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be1005.eqiad.wmnet * 11:55 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be1004.eqiad.wmnet * 11:54 cgoubert@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host rdb-lock2003.codfw.wmnet * 11:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 11:53 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 11:53 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:53 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 11:53 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 11:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1038.eqiad.wmnet * 11:49 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be1004.eqiad.wmnet * 11:44 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:44 moritzm: installing Linux 5.10.262 on Bullseye hosts * 11:41 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1038.eqiad.wmnet * 11:40 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 11:40 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 11:40 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 11:40 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:40 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 11:40 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 11:38 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1213.eqiad.wmnet with reason: host reimage * 11:36 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1167: Depool db1167.eqiad.wmnet to then clone it to db1281.eqiad.wmnet - marostegui@cumin1003 * 11:35 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 11:35 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2003.codfw.wmnet * 11:35 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1167: Depool db1167.eqiad.wmnet to then clone it to db1281.eqiad.wmnet - marostegui@cumin1003 * 11:35 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1167.eqiad.wmnet onto db1281.eqiad.wmnet * 11:34 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 22 hosts with reason: Cloning * 11:34 moritzm: remove ganeti3005 from esams03 cluster, hardware issues [[phab:T434646|T434646]] * 11:32 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1213.eqiad.wmnet with reason: host reimage * 11:28 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1034.eqiad.wmnet * 11:28 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1034.eqiad.wmnet * 11:22 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1034.eqiad.wmnet * 11:19 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1034.eqiad.wmnet * 11:17 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1213.eqiad.wmnet with OS bookworm * 11:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1165.eqiad.wmnet onto db1279.eqiad.wmnet * 11:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1165: Pool db1165.eqiad.wmnet in after cloning * 11:07 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply * 10:57 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply * 10:54 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply * 10:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1274.eqiad.wmnet with reason: Enabling notifications and pooling * 10:45 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1033.eqiad.wmnet * 10:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1033.eqiad.wmnet * 10:44 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply * 10:43 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'. * 10:42 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'. * 10:42 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'. * 10:40 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db[1216,1225,1239-1240].eqiad.wmnet with reason: reboot * 10:39 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1033.eqiad.wmnet * 10:38 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1151 from dbctl [[phab:T434538|T434538]]', diff saved to https://phabricator.wikimedia.org/P96055 and previous config saved to /var/cache/conftool/dbconfig/20260813-103828-marostegui.json * 10:35 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1033.eqiad.wmnet * 10:27 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1165: Pool db1165.eqiad.wmnet in after cloning * 10:24 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1151.eqiad.wmnet with OS bookworm * 10:15 moritzm: installing bind9 security updates (client-side tools/libs only) * 10:07 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-debug: apply * 10:06 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-debug: apply * 10:02 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 7 hosts with reason: reboot * 10:01 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-debug: apply * 10:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.upgrade (exit_code=0) for 1 hosts * 10:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2214: Upgrade of db2214.codfw.wmnet completed * 10:01 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-debug: apply * 10:00 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-debug: apply * 10:00 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-debug: apply * 09:59 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.decommission (exit_code=99) * 09:59 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 09:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1151.eqiad.wmnet with reason: host reimage * 09:59 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 8 hosts * 09:59 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 8 hosts * 09:55 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1151.eqiad.wmnet with reason: host reimage * 09:46 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260801000000" --end-timestamp="20260802000000" --sleep="3" --batch-size="5"` for [[phab:T434688|T434688]] * 09:45 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 8 hosts with reason: reboot * 09:45 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 22 hosts with reason: Cloning * 09:43 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1165: Depool db1165.eqiad.wmnet to then clone it to db1279.eqiad.wmnet - marostegui@cumin1003 * 09:42 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1165: Depool db1165.eqiad.wmnet to then clone it to db1279.eqiad.wmnet - marostegui@cumin1003 * 09:42 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1165.eqiad.wmnet onto db1279.eqiad.wmnet * 09:41 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=testwiki --start-timestamp="20260311000000" --end-timestamp="20260805000000" --sleep="5" --batch-size="2"` for [[phab:T434688|T434688]] * 09:40 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for backupmon1001.eqiad.wmnet * 09:40 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for backupmon1001.eqiad.wmnet * 09:40 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1151 * 09:40 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1151 * 09:37 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325408{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]], [[gerrit:1325407{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]] (duration: 06m 57s) * 09:36 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1151 * 09:36 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1151.eqiad.wmnet 13.36.64.10.in-addr.arpa 3.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:36 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1151.eqiad.wmnet 13.36.64.10.in-addr.arpa 3.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:36 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:36 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1151 - btullis@cumin1003" * 09:36 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on backupmon1001.eqiad.wmnet with reason: reboot * 09:36 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1151 - btullis@cumin1003" * 09:35 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host moss-be2003.codfw.wmnet * 09:33 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 7 hosts * 09:33 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 7 hosts * 09:33 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 09:32 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1325408{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]], [[gerrit:1325407{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1161.eqiad.wmnet onto db1275.eqiad.wmnet * 09:31 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 09:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1161: Pool db1161.eqiad.wmnet in after cloning * 09:30 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1325408{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]], [[gerrit:1325407{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]] * 09:29 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 09:27 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 09:27 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host moss-be2003.codfw.wmnet * 09:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be2006.codfw.wmnet * 09:25 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2004.codfw.wmnet * 09:22 hashar@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 09:21 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be2006.codfw.wmnet * 09:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be2005.codfw.wmnet * 09:19 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2004.codfw.wmnet * 09:19 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2003.codfw.wmnet * 09:18 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 7 hosts with reason: reboot * 09:18 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 7 hosts * 09:18 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 7 hosts * 09:16 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2214: Upgrade of db2214.codfw.wmnet completed * 09:14 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be2005.codfw.wmnet * 09:13 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be2004.codfw.wmnet * 09:12 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2003.codfw.wmnet * 09:12 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2002.codfw.wmnet * 09:09 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2214: Upgrading db2214.codfw.wmnet * 09:09 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2214: Upgrading db2214.codfw.wmnet * 09:09 cwilliams@cumin1003: START - Cookbook sre.mysql.upgrade for 1 hosts * 09:07 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be2004.codfw.wmnet * 09:06 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2002.codfw.wmnet * 09:05 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1004.eqiad.wmnet * 09:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-cluster (exit_code=0) * 09:03 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 7 hosts with reason: reboot * 09:01 btullis@cumin1003: START - Cookbook sre.dns.netbox * 09:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2214 [[phab:T434754|T434754]]', diff saved to https://phabricator.wikimedia.org/P96045 and previous config saved to /var/cache/conftool/dbconfig/20260813-090001-cwilliams.json * 08:59 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1004.eqiad.wmnet * 08:59 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1003.eqiad.wmnet * 08:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2229 to s6 primary [[phab:T434754|T434754]]', diff saved to https://phabricator.wikimedia.org/P96044 and previous config saved to /var/cache/conftool/dbconfig/20260813-085752-cwilliams.json * 08:57 cezmunsta: Starting s6 codfw failover from db2214 to db2229 - [[phab:T434754|T434754]] * 08:54 hashar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325416{{!}}Revert "REST: Enable `GET /lexemes/<nowiki>{</nowiki>lexeme_id<nowiki>}</nowiki>` by default" (T434712)]] (duration: 07m 22s) * 08:53 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1003.eqiad.wmnet * 08:53 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1002.eqiad.wmnet * 08:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2229 with weight 0 [[phab:T434754|T434754]]', diff saved to https://phabricator.wikimedia.org/P96043 and previous config saved to /var/cache/conftool/dbconfig/20260813-085151-cwilliams.json * 08:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 23 hosts with reason: Primary switchover s6 [[phab:T434754|T434754]] * 08:51 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts2002.codfw.wmnet * 08:50 hashar@deploy1003: hashar: Continuing with deployment * 08:49 hashar@deploy1003: hashar: Backport for [[gerrit:1325416{{!}}Revert "REST: Enable `GET /lexemes/<nowiki>{</nowiki>lexeme_id<nowiki>}</nowiki>` by default" (T434712)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:47 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet * 08:47 hashar@deploy1003: Started scap sync-world: Backport for [[gerrit:1325416{{!}}Revert "REST: Enable `GET /lexemes/<nowiki>{</nowiki>lexeme_id<nowiki>}</nowiki>` by default" (T434712)]] * 08:47 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1002.eqiad.wmnet * 08:46 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1161: Pool db1161.eqiad.wmnet in after cloning * 08:46 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host stewards2001.codfw.wmnet * 08:45 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit1003.wikimedia.org * 08:45 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host stewards1001.eqiad.wmnet * 08:45 Emperor: roll-restart apus frontends in codfw * 08:45 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-cluster * 08:44 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts2002.codfw.wmnet * 08:44 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host doc2003.codfw.wmnet * 08:43 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet * 08:43 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host phab2003.codfw.wmnet * 08:42 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host stewards2001.codfw.wmnet * 08:42 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host etherpad2002.codfw.wmnet * 08:41 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host stewards1001.eqiad.wmnet * 08:41 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1273: Pool in s7 * 08:41 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1003.wikimedia.org * 08:40 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host doc2003.codfw.wmnet * 08:40 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit1003.wikimedia.org * 08:39 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host etherpad1004.eqiad.wmnet * 08:39 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host doc1004.eqiad.wmnet * 08:39 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit2002.wikimedia.org * 08:38 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host etherpad2002.codfw.wmnet * 08:37 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host phab2003.codfw.wmnet * 08:36 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lists1004.wikimedia.org * 08:35 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host etherpad1004.eqiad.wmnet * 08:35 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host doc1004.eqiad.wmnet * 08:34 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-cluster (exit_code=0) * 08:34 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1003.wikimedia.org * 08:34 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host planet2003.codfw.wmnet * 08:34 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2003.wikimedia.org * 08:33 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host planet1003.eqiad.wmnet * 08:33 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit2002.wikimedia.org * 08:32 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 22 hosts with reason: Cloning * 08:32 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2002.wikimedia.org * 08:31 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aphlict1002.eqiad.wmnet * 08:30 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host planet2003.codfw.wmnet * 08:29 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host planet1003.eqiad.wmnet * 08:29 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host lists1004.wikimedia.org * 08:28 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lists2001.wikimedia.org * 08:28 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2003.wikimedia.org * 08:27 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host aphlict1002.eqiad.wmnet * 08:27 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aphlict2001.codfw.wmnet * 08:26 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2002.wikimedia.org * 08:23 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host aphlict2001.codfw.wmnet * 08:23 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1151 * 08:22 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1151.eqiad.wmnet with OS bookworm * 08:22 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host lists2001.wikimedia.org * 08:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1161: Depool db1161.eqiad.wmnet to then clone it to db1275.eqiad.wmnet - marostegui@cumin1003 * 08:20 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1161: Depool db1161.eqiad.wmnet to then clone it to db1275.eqiad.wmnet - marostegui@cumin1003 * 08:20 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1161.eqiad.wmnet onto db1275.eqiad.wmnet * 08:15 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-cluster * 08:15 Emperor: roll-restart apus frontends in eqiad * 07:56 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1273: Pool in s7 * 07:56 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1273 to dbctl [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P96035 and previous config saved to /var/cache/conftool/dbconfig/20260813-075611-marostegui.json * 07:38 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db1273.eqiad.wmnet with reason: Reboot * 07:31 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.sanitize-wiki (exit_code=97) Managing sanitization for wikis testwiki in section s3 * 07:24 marostegui@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis testwiki in section s3 * 07:19 jayme@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on kubestagemaster2005.codfw.wmnet with reason: downtime because of hardware failure and no DRBD * 05:42 arnaudb@dns1006: END - running authdns-update * 05:40 arnaudb@dns1006: START - running authdns-update * 05:27 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 04:06 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 02:29 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1151 * 02:29 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1151 * 02:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1219.eqiad.wmnet with OS bookworm * 02:18 ryankemper: [[phab:T434494|T434494]] `ryankemper@deploy1003:~$ echo 'https://stats.wikimedia.org/' {{!}} mwscript-k8s --attach -- purgeList.php` (default page got cached during yesterday's `an-web1001` reimage) * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 46s) * 02:03 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1219.eqiad.wmnet with reason: host reimage * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 02:00 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1219.eqiad.wmnet with reason: host reimage * 01:46 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1219.eqiad.wmnet with OS bookworm * 01:03 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1142.eqiad.wmnet with OS bookworm * 00:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1142.eqiad.wmnet with reason: host reimage * 00:34 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1142.eqiad.wmnet with reason: host reimage * 00:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1142.eqiad.wmnet with OS bookworm * 00:16 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1178 * 00:16 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1178 == 2026-08-12 == * 23:18 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324409{{!}}Remove $wmg = $wg hacks in Collection (T119117)]] (duration: 06m 43s) * 23:14 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 23:13 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1324409{{!}}Remove $wmg = $wg hacks in Collection (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:11 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1324409{{!}}Remove $wmg = $wg hacks in Collection (T119117)]] * 22:59 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 22:49 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260801000000" --end-timestamp="20260802000000" --sleep=2 --batch-size=10` * 22:45 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=testwiki --start-timestamp="20200801010101" --end-timestamp="20260816010101" --sleep=15 --batch-size=5` * 22:40 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=testwiki --start-timestamp="20200101010101" --end-timestamp="20260816010101" --sleep=60` * 22:35 Dreamy_Jazz: Running `mwscript WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260101000000" --end-timestamp="20260102000000" --sleep=10` * 22:21 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324817{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]], [[gerrit:1324818{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]] (duration: 45m 29s) * 22:17 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 21:59 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 21:40 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1324817{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]], [[gerrit:1324818{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:39 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 21:36 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1324817{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]], [[gerrit:1324818{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]] * 21:32 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:31 vriley@cumin1003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:30 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:30 vriley@cumin1003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:17 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:14 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:14 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:11 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:10 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-codfw: Set storage compatability to NONE — [[phab:T433028|T433028]] - eevans@cumin1003 * 21:10 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1006 * 21:09 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1006 * 21:05 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324277{{!}}Improve Math preference labels for SVG/MathJax/MathML (T433891)]] (duration: 31m 42s) * 20:58 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1178.eqiad.wmnet with OS bookworm * 20:54 krinkle@deploy1003: krinkle: Continuing with deployment * 20:51 krinkle@deploy1003: krinkle: Backport for [[gerrit:1324277{{!}}Improve Math preference labels for SVG/MathJax/MathML (T433891)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:39 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 20:38 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:38 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp5022.eqsin.wmnet with OS trixie * 20:38 cdobbins@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - cdobbins@cumin1003" * 20:37 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:36 cdobbins@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - cdobbins@cumin1003" * 20:34 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1178.eqiad.wmnet with reason: host reimage * 20:34 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1324277{{!}}Improve Math preference labels for SVG/MathJax/MathML (T433891)]] * 20:33 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 20:28 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1178.eqiad.wmnet with reason: host reimage * 20:13 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1178.eqiad.wmnet with OS bookworm * 20:11 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-worker1178.eqiad.wmnet with OS bookworm * 20:11 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1178.eqiad.wmnet with OS bookworm * 20:09 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-codfw: Set storage compatability to NONE — [[phab:T433028|T433028]] - eevans@cumin1003 * 20:09 Dreamy_Jazz: Evening UTC backport window done * 20:08 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324794{{!}}WikimediaAntiAbuse: Enable logging channel (T431292)]] (duration: 06m 48s) * 20:08 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage * 20:05 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage * 20:04 dreamyjazz@deploy1003: kharlan, dreamyjazz: Continuing with deployment * 20:04 dreamyjazz@deploy1003: kharlan, dreamyjazz: Backport for [[gerrit:1324794{{!}}WikimediaAntiAbuse: Enable logging channel (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:01 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1324794{{!}}WikimediaAntiAbuse: Enable logging channel (T431292)]] * 19:55 brennen@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 19:47 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 19:35 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 19:35 cdobbins@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cp5022.eqsin.wmnet with OS trixie * 19:32 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:30 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:29 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:26 vriley@cumin1003: START - Cookbook sre.dns.netbox * 19:23 brennen@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 19:19 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-eqiad: Set storage compatability to NONE — [[phab:T433028|T433028]] - eevans@cumin1003 * 19:10 brennen@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 19:09 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 19:09 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 19:00 Amir1: data migrated on wikishared ([[phab:T426102|T426102]]) * 18:57 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321224{{!}}Rename ce_worklist_articles table to ce_invitation_list_articles (T426102)]] (duration: 06m 50s) * 18:53 ladsgroup@deploy1003: ladsgroup, daimona: Continuing with deployment * 18:53 ladsgroup@deploy1003: ladsgroup, daimona: Backport for [[gerrit:1321224{{!}}Rename ce_worklist_articles table to ce_invitation_list_articles (T426102)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:51 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1321224{{!}}Rename ce_worklist_articles table to ce_invitation_list_articles (T426102)]] * 18:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2192: Security update * 18:31 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 18:25 Amir1: ce_invitation_list_articles created as empty on wikishared ([[phab:T426102|T426102]]) * 18:21 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 18:21 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 18:18 Amir1: migrated testwiki entries from ce_worklist_articles to ce_invitation_list_articles ([[phab:T426102|T426102]]) * 18:18 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-eqiad: Set storage compatability to NONE — [[phab:T433028|T433028]] - eevans@cumin1003 * 18:11 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 18:08 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling reboot on A:durum-eqsin and A:durum * 18:07 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Set storage compatability to UPGRADING — [[phab:T433028|T433028]] - eevans@cumin1003 * 18:05 jhancock@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022'] * 17:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2192: Security update * 17:55 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:55 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum-eqsin and A:durum * 17:53 jhancock@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['cp5022'] * 17:47 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:47 jhancock@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['cp5022'] * 17:42 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:41 jhancock@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['cp5022'] * 17:36 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:35 jhancock@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022'] * 17:31 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-magru and not (P<nowiki>{</nowiki>cp7001*<nowiki>}</nowiki> or P<nowiki>{</nowiki>cp7009*<nowiki>}</nowiki>) and A:cp - 9.2.15 upgrade ([[phab:T434620|T434620]]) * 17:28 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:28 jhancock@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['cp5022'] * 17:22 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2192.codfw.wmnet with reason: Maintenance * 17:11 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host apifeatureusage2001.codfw.wmnet with OS bookworm * 17:04 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: apply * 17:03 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-main: apply * 17:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2192 [[phab:T434635|T434635]]', diff saved to https://phabricator.wikimedia.org/P96030 and previous config saved to /var/cache/conftool/dbconfig/20260812-170338-cwilliams.json * 17:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2213 to s5 primary [[phab:T434635|T434635]]', diff saved to https://phabricator.wikimedia.org/P96029 and previous config saved to /var/cache/conftool/dbconfig/20260812-170152-cwilliams.json * 17:01 cezmunsta: Starting s5 codfw failover from db2192 to db2213 - [[phab:T434635|T434635]] * 16:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2213 with weight 0 [[phab:T434635|T434635]]', diff saved to https://phabricator.wikimedia.org/P96028 and previous config saved to /var/cache/conftool/dbconfig/20260812-165544-cwilliams.json * 16:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 27 hosts with reason: Primary switchover s5 [[phab:T434635|T434635]] * 16:53 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: apply * 16:52 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-main: apply * 16:44 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-main: apply * 16:44 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-main: apply * 16:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1272: New host * 16:40 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324765{{!}}WikimediaAntiAbuse: Enable PersonalInfoFlagNotifications (T431292)]] (duration: 07m 02s) * 16:40 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-logging-external: apply * 16:39 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-logging-external: apply * 16:38 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-logging-external: apply * 16:37 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-logging-external: apply * 16:36 kharlan@deploy1003: kharlan: Continuing with deployment * 16:35 kharlan@deploy1003: kharlan: Backport for [[gerrit:1324765{{!}}WikimediaAntiAbuse: Enable PersonalInfoFlagNotifications (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:33 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1324765{{!}}WikimediaAntiAbuse: Enable PersonalInfoFlagNotifications (T431292)]] * 16:27 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-logging-external: apply * 16:27 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-logging-external: apply * 16:18 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 16:18 jhancock@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022'] * 16:17 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 16:16 jhancock@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['cp5022'] * 16:15 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 16:12 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Set storage compatability to UPGRADING — [[phab:T433028|T433028]] - eevans@cumin1003 * 16:10 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: apply * 16:10 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: apply * 16:08 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: apply * 16:08 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: apply * 16:08 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics: apply * 16:07 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics: apply * 16:02 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-magru and not (P<nowiki>{</nowiki>cp7001*<nowiki>}</nowiki> or P<nowiki>{</nowiki>cp7009*<nowiki>}</nowiki>) and A:cp - 9.2.15 upgrade ([[phab:T434620|T434620]]) * 15:57 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1272: New host * 15:52 jmm@cumin2003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti2046.codfw.wmnet * 15:52 jmm@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host ganeti2046.codfw.wmnet * 15:42 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:37 mforns@deploy1003: Finished deploy [analytics/refinery@49c336c] (thin): Regular analytics weekly train THIN [analytics/refinery@49c336cd] (duration: 01m 59s) * 15:37 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324733{{!}}WikimediaAntiAbuse: Enable personal info tag display on enwiki (T431292)]] (duration: 08m 12s) * 15:35 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:35 mforns@deploy1003: Started deploy [analytics/refinery@49c336c] (thin): Regular analytics weekly train THIN [analytics/refinery@49c336cd] * 15:34 mforns@deploy1003: Finished deploy [analytics/refinery@49c336c]: Regular analytics weekly train [analytics/refinery@49c336cd] (duration: 04m 20s) * 15:33 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[2024,1031]*.wmnet: Set storage compatability to UPGRADING — [[phab:T433028|T433028]] - eevans@cumin1003 * 15:33 dreamyjazz@deploy1003: kharlan, dreamyjazz: Continuing with deployment * 15:31 dreamyjazz@deploy1003: kharlan, dreamyjazz: Backport for [[gerrit:1324733{{!}}WikimediaAntiAbuse: Enable personal info tag display on enwiki (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:30 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitize-wiki (exit_code=99) Checking sanitization for wikis testwiki in section s3 * 15:30 mforns@deploy1003: Started deploy [analytics/refinery@49c336c]: Regular analytics weekly train [analytics/refinery@49c336cd] * 15:30 mforns@deploy1003: Finished deploy [analytics/refinery@49c336c] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@49c336cd] (duration: 00m 32s) * 15:29 mforns@deploy1003: Started deploy [analytics/refinery@49c336c] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@49c336cd] * 15:29 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1324733{{!}}WikimediaAntiAbuse: Enable personal info tag display on enwiki (T431292)]] * 15:27 brennen@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324700{{!}}EventDetailsParticipantsModule: populate cache with non-local users (T434597)]] (duration: 06m 38s) * 15:23 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[2024,1031]*.wmnet: Set storage compatability to UPGRADING — [[phab:T433028|T433028]] - eevans@cumin1003 * 15:23 brennen@deploy1003: brennen, daimona: Continuing with deployment * 15:23 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:22 brennen@deploy1003: brennen, daimona: Backport for [[gerrit:1324700{{!}}EventDetailsParticipantsModule: populate cache with non-local users (T434597)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:22 cgoubert@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host rdb-lock2003.codfw.wmnet * 15:21 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 15:21 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 15:21 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:21 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:21 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:20 brennen@deploy1003: Started scap sync-world: Backport for [[gerrit:1324700{{!}}EventDetailsParticipantsModule: populate cache with non-local users (T434597)]] * 15:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1178.eqiad.wmnet with OS bookworm * 15:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock1003.eqiad.wmnet * 15:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock1003.eqiad.wmnet with OS trixie * 15:16 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 15:16 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 15:16 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 15:16 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:16 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:16 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:12 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 15:12 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2003.codfw.wmnet * 15:11 cgoubert@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host rdb-lock2003.codfw.wmnet * 15:11 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 15:11 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 15:11 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:11 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:11 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:07 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324338{{!}}InitialiseSettings: Enable 2FA warnings on more private wikis (T428103)]] (duration: 07m 02s) * 15:04 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 15:04 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock1003.eqiad.wmnet with reason: host reimage * 15:03 reedy@deploy1003: reedy: Continuing with deployment * 15:02 reedy@deploy1003: reedy: Backport for [[gerrit:1324338{{!}}InitialiseSettings: Enable 2FA warnings on more private wikis (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:02 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:00 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324338{{!}}InitialiseSettings: Enable 2FA warnings on more private wikis (T428103)]] * 14:57 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 14:57 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2003.codfw.wmnet * 14:57 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock1003.eqiad.wmnet with reason: host reimage * 14:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock2002.codfw.wmnet * 14:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock2002.codfw.wmnet with OS trixie * 14:56 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 14:56 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 14:56 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 14:55 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 14:55 moritzm: powercycle ganeti2046 * 14:47 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock1003.eqiad.wmnet with OS trixie * 14:46 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1003.eqiad.wmnet - cgoubert@cumin2003" * 14:46 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1003.eqiad.wmnet - cgoubert@cumin2003" * 14:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock1003.eqiad.wmnet on all recursors * 14:45 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock1003.eqiad.wmnet on all recursors * 14:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1003.eqiad.wmnet - cgoubert@cumin2003" * 14:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1272.eqiad.wmnet with reason: Enabling notifications * 14:44 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1003.eqiad.wmnet - cgoubert@cumin2003" * 14:44 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324719{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324720{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324722{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0 (T434187)]] (duration: 11m 02s) * 14:39 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 14:39 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock1003.eqiad.wmnet * 14:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock2002.codfw.wmnet with reason: host reimage * 14:37 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 14:37 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock1002.eqiad.wmnet * 14:37 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock1002.eqiad.wmnet with OS trixie * 14:37 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1324719{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324720{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324722{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0 (T434187)]] synced to the testservers (see https://wikitech. * 14:33 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock2002.codfw.wmnet with reason: host reimage * 14:33 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1324719{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324720{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324722{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0 (T434187)]] * 14:32 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1018.eqiad.wmnet with OS bookworm * 14:32 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2046.codfw.wmnet * 14:31 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1020.eqiad.wmnet with OS bookworm * 14:27 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2046.codfw.wmnet * 14:25 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2045.codfw.wmnet * 14:25 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2045.codfw.wmnet * 14:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock1002.eqiad.wmnet with reason: host reimage * 14:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1019.eqiad.wmnet with OS bookworm * 14:22 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Checking sanitization for wikis testwiki in section s3 * 14:20 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2045.codfw.wmnet * 14:18 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock1002.eqiad.wmnet with reason: host reimage * 14:17 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:16 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324707{{!}}Backport all changes from wmf/1.47.0-wmf.15]] (duration: 40m 51s) * 14:16 moritzm: installing Linux 6.1.180 on Bookworm hosts * 14:15 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock2002.codfw.wmnet with OS trixie * 14:14 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2002.codfw.wmnet - cgoubert@cumin2003" * 14:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2002.codfw.wmnet - cgoubert@cumin2003" * 14:14 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:14 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2045.codfw.wmnet * 14:14 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2002.codfw.wmnet on all recursors * 14:14 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2002.codfw.wmnet on all recursors * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2002.codfw.wmnet - cgoubert@cumin2003" * 14:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2002.codfw.wmnet - cgoubert@cumin2003" * 14:12 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2030.codfw.wmnet * 14:12 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2030.codfw.wmnet * 14:11 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on apifeatureusage2001.codfw.wmnet with reason: host reimage * 14:09 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:08 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:07 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 14:06 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2030.codfw.wmnet * 14:06 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:06 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock1002.eqiad.wmnet with OS trixie * 14:05 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 14:05 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2002.codfw.wmnet * 14:05 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1002.eqiad.wmnet - cgoubert@cumin2003" * 14:05 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1002.eqiad.wmnet - cgoubert@cumin2003" * 14:05 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock1002.eqiad.wmnet on all recursors * 14:05 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock1002.eqiad.wmnet on all recursors * 14:05 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:05 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1002.eqiad.wmnet - cgoubert@cumin2003" * 14:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:04 kharlan@deploy1003: kharlan: Continuing with deployment * 14:04 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock2001.codfw.wmnet * 14:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock2001.codfw.wmnet with OS trixie * 14:02 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1002.eqiad.wmnet - cgoubert@cumin2003" * 14:02 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on apifeatureusage2001.codfw.wmnet with reason: host reimage * 14:01 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2030.codfw.wmnet * 13:59 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2029.codfw.wmnet * 13:58 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2029.codfw.wmnet * 13:58 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 13:58 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock1002.eqiad.wmnet * 13:56 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1019.eqiad.wmnet with reason: host reimage * 13:54 btullis@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on archiva1002.wikimedia.org with reason: Upgrading in-place * 13:53 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock1001.eqiad.wmnet * 13:53 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock1001.eqiad.wmnet with OS trixie * 13:53 kharlan@deploy1003: kharlan: Backport for [[gerrit:1324707{{!}}Backport all changes from wmf/1.47.0-wmf.15]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:52 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2029.codfw.wmnet * 13:52 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1020.eqiad.wmnet with reason: host reimage * 13:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1019.eqiad.wmnet with reason: host reimage * 13:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1020.eqiad.wmnet with reason: host reimage * 13:49 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock2001.codfw.wmnet with reason: host reimage * 13:48 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2029.codfw.wmnet * 13:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Configuring db1272 for s3 pooling', diff saved to https://phabricator.wikimedia.org/P96021 and previous config saved to /var/cache/conftool/dbconfig/20260812-134732-cwilliams.json * 13:44 bking@cumin2003: START - Cookbook sre.hosts.reimage for host apifeatureusage2001.codfw.wmnet with OS bookworm * 13:43 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock2001.codfw.wmnet with reason: host reimage * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2028.codfw.wmnet * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2028.codfw.wmnet * 13:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1018.eqiad.wmnet with reason: host reimage * 13:38 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock1001.eqiad.wmnet with reason: host reimage * 13:36 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1018.eqiad.wmnet with reason: host reimage * 13:35 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1324707{{!}}Backport all changes from wmf/1.47.0-wmf.15]] * 13:35 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2028.codfw.wmnet * 13:32 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1275938{{!}}Enable campaignEvents on bdwikimedia (T424016)]] (duration: 07m 35s) * 13:32 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1020.eqiad.wmnet with OS bookworm * 13:32 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1019.eqiad.wmnet with OS bookworm * 13:32 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock1001.eqiad.wmnet with reason: host reimage * 13:28 kharlan@deploy1003: kharlan, yahya: Continuing with deployment * 13:27 kharlan@deploy1003: kharlan, yahya: Backport for [[gerrit:1275938{{!}}Enable campaignEvents on bdwikimedia (T424016)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:27 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2028.codfw.wmnet * 13:26 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 13:26 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 13:25 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1275938{{!}}Enable campaignEvents on bdwikimedia (T424016)]] * 13:25 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitize-wiki (exit_code=99) Managing sanitization for wikis testwiki in section s3 * 13:24 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock2001.codfw.wmnet with OS trixie * 13:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2001.codfw.wmnet - cgoubert@cumin2003" * 13:24 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2001.codfw.wmnet - cgoubert@cumin2003" * 13:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2001.codfw.wmnet on all recursors * 13:23 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2001.codfw.wmnet on all recursors * 13:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2001.codfw.wmnet - cgoubert@cumin2003" * 13:23 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2001.codfw.wmnet - cgoubert@cumin2003" * 13:23 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324713{{!}}thwiki: reinstate temporary wiki25 logos (T431094)]] (duration: 07m 13s) * 13:20 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:20 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1018.eqiad.wmnet with OS bookworm * 13:20 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2027.codfw.wmnet * 13:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2027.codfw.wmnet * 13:19 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics-external: apply * 13:18 kharlan@deploy1003: anzx, kharlan: Continuing with deployment * 13:18 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:18 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock1001.eqiad.wmnet with OS trixie * 13:17 kharlan@deploy1003: anzx, kharlan: Backport for [[gerrit:1324713{{!}}thwiki: reinstate temporary wiki25 logos (T431094)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1001.eqiad.wmnet - cgoubert@cumin2003" * 13:17 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 13:17 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1001.eqiad.wmnet - cgoubert@cumin2003" * 13:17 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2001.codfw.wmnet * 13:17 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics-external: apply * 13:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock1001.eqiad.wmnet on all recursors * 13:17 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock1001.eqiad.wmnet on all recursors * 13:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1001.eqiad.wmnet - cgoubert@cumin2003" * 13:17 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1001.eqiad.wmnet - cgoubert@cumin2003" * 13:15 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1324713{{!}}thwiki: reinstate temporary wiki25 logos (T431094)]] * 13:15 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:15 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics-external: apply * 13:15 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:14 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics-external: apply * 13:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2027.codfw.wmnet * 13:12 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 13:12 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock1001.eqiad.wmnet * 13:12 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2027.codfw.wmnet * 13:08 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1158.eqiad.wmnet onto db1273.eqiad.wmnet * 13:07 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1158: Pool db1158.eqiad.wmnet in after cloning * 13:02 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti6002.drmrs.wmnet * 13:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti6002.drmrs.wmnet * 12:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1015.eqiad.wmnet with OS bookworm * 12:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti6002.drmrs.wmnet * 12:44 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1017.eqiad.wmnet with OS bookworm * 12:36 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 12:35 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 12:34 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 12:33 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 12:31 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti6002.drmrs.wmnet * 12:24 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1017.eqiad.wmnet with reason: host reimage * 12:22 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1158: Pool db1158.eqiad.wmnet in after cloning * 12:18 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1017.eqiad.wmnet with reason: host reimage * 12:11 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1015.eqiad.wmnet with reason: host reimage * 12:07 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1015.eqiad.wmnet with reason: host reimage * 12:04 moritzm: failover ganeti master in drmrs02 to ganeti6004 * 12:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1009.eqiad.wmnet with OS bookworm * 12:01 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1017.eqiad.wmnet with OS bookworm * 12:00 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti6004.drmrs.wmnet * 12:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti6004.drmrs.wmnet * 11:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti6004.drmrs.wmnet * 11:53 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1015.eqiad.wmnet with OS bookworm * 11:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1159.eqiad.wmnet onto db1274.eqiad.wmnet * 11:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Pool db1159.eqiad.wmnet in after cloning * 11:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1016.eqiad.wmnet with OS bookworm * 11:46 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti6004.drmrs.wmnet * 11:45 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti6001.drmrs.wmnet * 11:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti6001.drmrs.wmnet * 11:43 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324306{{!}}WikimediaAntiAbuse: Enable personal info for enwiki with no display (T431292)]] (duration: 10m 26s) * 11:42 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1015.eqiad.wmnet with OS bookworm * 11:39 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 11:38 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti6001.drmrs.wmnet * 11:34 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1324306{{!}}WikimediaAntiAbuse: Enable personal info for enwiki with no display (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:33 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2212: Security update * 11:33 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti6001.drmrs.wmnet * 11:32 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1324306{{!}}WikimediaAntiAbuse: Enable personal info for enwiki with no display (T431292)]] * 11:22 moritzm: failover ganeti master in drmrs01 to ganeti6003 * 11:20 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:20 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:18 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 22 hosts with reason: Cloning * 11:17 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti6003.drmrs.wmnet * 11:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti6003.drmrs.wmnet * 11:17 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1009.eqiad.wmnet with reason: host reimage * 11:17 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:16 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:14 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1016.eqiad.wmnet with reason: host reimage * 11:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti6003.drmrs.wmnet * 11:11 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1009.eqiad.wmnet with reason: host reimage * 11:10 jmm@cumin2003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti3005.esams.wmnet * 11:10 jmm@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host ganeti3005.esams.wmnet * 11:09 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 11:08 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1158: Depool db1158.eqiad.wmnet to then clone it to db1273.eqiad.wmnet - marostegui@cumin1003 * 11:07 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1016.eqiad.wmnet with reason: host reimage * 11:07 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1158: Depool db1158.eqiad.wmnet to then clone it to db1273.eqiad.wmnet - marostegui@cumin1003 * 11:07 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1158.eqiad.wmnet onto db1273.eqiad.wmnet * 11:06 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:05 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:05 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1159: Pool db1159.eqiad.wmnet in after cloning * 11:04 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 20 hosts with reason: Cloning * 11:02 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti6003.drmrs.wmnet * 11:00 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis testwiki in section s3 * 10:54 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1009.eqiad.wmnet with OS bookworm * 10:51 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1015.eqiad.wmnet with OS bookworm * 10:50 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1016.eqiad.wmnet with OS bookworm * 10:48 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2212: Security update * 10:45 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitize-wiki (exit_code=99) Managing sanitization for wikis testwiki in section s3 * 10:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1013.eqiad.wmnet with OS bookworm * 10:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1014.eqiad.wmnet with OS bookworm * 10:38 cwilliams@cumin1003: START - Cookbook sre.mysql.clone of db1159.eqiad.wmnet onto db1274.eqiad.wmnet * 10:33 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db1274.eqiad.wmnet * 10:33 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db1274.eqiad.wmnet * 10:31 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 396993 * 10:29 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 396993 * 10:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1159: Clone source for db1274 * 10:24 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1159: Clone source for db1274 * 10:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1013.eqiad.wmnet with reason: host reimage * 10:18 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1013.eqiad.wmnet with reason: host reimage * 10:13 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2212.codfw.wmnet with reason: Maintenance * 10:13 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 10:12 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 10:12 blake@deploy1003: Stopping before sync operations * 10:11 blake@deploy1003: Started scap sync-world: Non-deployment scap run to populate new release values for [[phab:T427668|T427668]] * 10:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2212 [[phab:T434644|T434644]]', diff saved to https://phabricator.wikimedia.org/P96003 and previous config saved to /var/cache/conftool/dbconfig/20260812-101053-cwilliams.json * 10:09 moritzm: powercycle ganeti3005 * 10:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2203 to s1 primary [[phab:T434644|T434644]]', diff saved to https://phabricator.wikimedia.org/P96002 and previous config saved to /var/cache/conftool/dbconfig/20260812-100849-cwilliams.json * 10:08 cezmunsta: Starting s1 codfw failover from db2212 to db2203 - [[phab:T434644|T434644]] * 10:03 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1013.eqiad.wmnet with OS bookworm * 10:02 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1013.eqiad.wmnet with OS bookworm * 10:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2203 with weight 0 [[phab:T434644|T434644]]', diff saved to https://phabricator.wikimedia.org/P96001 and previous config saved to /var/cache/conftool/dbconfig/20260812-100134-cwilliams.json * 10:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 32 hosts with reason: Primary switchover s1 [[phab:T434644|T434644]] * 09:53 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1014.eqiad.wmnet with reason: host reimage * 09:50 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 09:50 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti3005.esams.wmnet * 09:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1014.eqiad.wmnet with reason: host reimage * 09:41 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1278: Pool in x1 * 09:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-launcher1003.eqiad.wmnet with OS bookworm * 09:37 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti3005.esams.wmnet * 09:37 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1013.eqiad.wmnet with OS bookworm * 09:34 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-presto1013.eqiad.wmnet with OS bookworm * 09:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1014.eqiad.wmnet with OS bookworm * 09:29 moritzm: failover ganeti master in esams to ganeti3008 * 09:26 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti3006.esams.wmnet * 09:26 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti3006.esams.wmnet * 09:24 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1012.eqiad.wmnet with OS bookworm * 09:23 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1009.eqiad.wmnet with OS bookworm * 09:23 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 09:18 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti3006.esams.wmnet * 09:16 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti3006.esams.wmnet * 09:14 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: fix regexp escaping bug - oblivian@cumin1003" * 09:14 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: fix regexp escaping bug - oblivian@cumin1003 * 09:13 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: fix regexp escaping bug - oblivian@cumin1003 * 09:13 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: fix regexp escaping bug - oblivian@cumin1003" * 09:03 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-launcher1003.eqiad.wmnet with reason: host reimage * 08:58 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-launcher1003.eqiad.wmnet with reason: host reimage * 08:55 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1278: Pool in x1 * 08:55 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1278 to dbctl [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95996 and previous config saved to /var/cache/conftool/dbconfig/20260812-085521-marostegui.json * 08:51 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1012.eqiad.wmnet with reason: host reimage * 08:45 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis testwiki in section s3 * 08:43 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1009.eqiad.wmnet with OS bookworm * 08:42 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1012.eqiad.wmnet with reason: host reimage * 08:41 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-launcher1003.eqiad.wmnet with OS bookworm * 08:40 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1013.eqiad.wmnet with OS bookworm * 08:38 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms3', diff saved to https://phabricator.wikimedia.org/P95995 and previous config saved to /var/cache/conftool/dbconfig/20260812-083816-marostegui.json * 08:38 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-master1003.eqiad.wmnet with OS bookworm * 08:37 marostegui: Failover ms3 [[phab:T434288|T434288]] * 08:37 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1268 to dbctl [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95994 and previous config saved to /var/cache/conftool/dbconfig/20260812-083722-marostegui.json * 08:35 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-web1001.eqiad.wmnet with OS bookworm * 08:32 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db2252.codfw.wmnet,db[1153,1268].eqiad.wmnet with reason: Switching over ms3 * 08:28 marostegui@cumin1003: dbctl commit (dc=all): 'Depool ms3 [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95993 and previous config saved to /var/cache/conftool/dbconfig/20260812-082852-marostegui.json * 08:25 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1012.eqiad.wmnet with OS bookworm * 08:25 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1011.eqiad.wmnet with OS bookworm * 08:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-master1003.eqiad.wmnet with reason: host reimage * 08:07 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-master1003.eqiad.wmnet with reason: host reimage * 08:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-web1001.eqiad.wmnet with reason: host reimage * 07:58 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-web1001.eqiad.wmnet with reason: host reimage * 07:50 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1003.eqiad.wmnet with OS bookworm * 07:38 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1011.eqiad.wmnet with reason: host reimage * 07:38 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-master1003.eqiad.wmnet with OS bookworm * 07:35 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti3007.esams.wmnet * 07:35 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti3007.esams.wmnet * 07:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1011.eqiad.wmnet with reason: host reimage * 07:27 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti3007.esams.wmnet * 07:25 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti3007.esams.wmnet * 07:25 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti3008.esams.wmnet * 07:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti3008.esams.wmnet * 07:22 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-web1001.eqiad.wmnet with OS bookworm * 07:19 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1003.eqiad.wmnet with OS bookworm * 07:18 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-master1003.eqiad.wmnet * 07:18 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host an-master1003.eqiad.wmnet * 07:17 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1011.eqiad.wmnet with OS bookworm * 07:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 07:16 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti3008.esams.wmnet * 07:15 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1009.eqiad.wmnet with OS bookworm * 07:14 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host an-master1003.eqiad.wmnet * 07:13 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-master1003.eqiad.wmnet * 07:13 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-master1003.eqiad.wmnet * 07:12 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-master1003.eqiad.wmnet * 07:11 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti3008.esams.wmnet * 07:07 arnaudb@dns1006: END - running authdns-update * 07:07 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti5007.eqsin.wmnet * 07:07 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti5007.eqsin.wmnet * 07:05 arnaudb@dns1006: START - running authdns-update * 06:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti5007.eqsin.wmnet * 06:54 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti5007.eqsin.wmnet * 06:38 moritzm: failover ganeti master in eqsin to ganeti5004 * 06:36 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti5006.eqsin.wmnet * 06:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti5006.eqsin.wmnet * 06:28 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti5006.eqsin.wmnet * 06:23 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti5006.eqsin.wmnet * 06:20 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti5005.eqsin.wmnet * 06:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti5005.eqsin.wmnet * 06:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti5005.eqsin.wmnet * 06:06 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti5005.eqsin.wmnet * 06:03 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti5004.eqsin.wmnet * 06:03 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti5004.eqsin.wmnet * 05:55 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti5004.eqsin.wmnet * 05:53 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti5004.eqsin.wmnet * 04:40 ryankemper: [[phab:T434494|T434494]] reimaged `an-tool1008.eqiad.wmnet` to bookworm; yarn.wikimedia.org is back up * 04:16 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-tool1008.eqiad.wmnet with OS bookworm * 03:58 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-tool1008.eqiad.wmnet with reason: host reimage * 03:53 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-tool1008.eqiad.wmnet with reason: host reimage * 03:41 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-tool1008.eqiad.wmnet with OS bookworm * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 45s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 00:25 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324427{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]], [[gerrit:1324429{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0]], [[gerrit:1324428{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]] (duration: 07m 55s) * 00:21 kemayo@deploy1003: kemayo: Continuing with deployment * 00:19 kemayo@deploy1003: kemayo: Backport for [[gerrit:1324427{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]], [[gerrit:1324429{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0]], [[gerrit:1324428{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:17 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1324427{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]], [[gerrit:1324429{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0]], [[gerrit:1324428{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]] == 2026-08-11 == * 21:37 sbassett: Deployed security fix for [[phab:T434521|T434521]] (wmf.15) * 21:29 sbassett: Deployed security fix for [[phab:T434521|T434521]] (wmf.14) * 21:19 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324370{{!}}Phase 4 of legal footer deployment (T432796)]], [[gerrit:1319804{{!}}Disable wgMFCustomSiteModules on English Wikipedia (T375538)]] (duration: 15m 26s) * 21:15 jdlrobson@deploy1003: jdlrobson: Continuing with deployment * 21:06 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1324370{{!}}Phase 4 of legal footer deployment (T432796)]], [[gerrit:1319804{{!}}Disable wgMFCustomSiteModules on English Wikipedia (T375538)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:03 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1324370{{!}}Phase 4 of legal footer deployment (T432796)]], [[gerrit:1319804{{!}}Disable wgMFCustomSiteModules on English Wikipedia (T375538)]] * 20:59 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 20:50 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324384{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]], [[gerrit:1324385{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]] (duration: 06m 58s) * 20:46 kemayo@deploy1003: kemayo: Continuing with deployment * 20:45 kemayo@deploy1003: kemayo: Backport for [[gerrit:1324384{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]], [[gerrit:1324385{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:43 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1324384{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]], [[gerrit:1324385{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]] * 20:42 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 20:42 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324386{{!}}build: Updating js-yaml to 3.15.1, 4.3.1]] (duration: 07m 36s) * 20:38 kemayo@deploy1003: kemayo: Continuing with deployment * 20:37 jhancock@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 20:37 kemayo@deploy1003: kemayo: Backport for [[gerrit:1324386{{!}}build: Updating js-yaml to 3.15.1, 4.3.1]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:35 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1324386{{!}}build: Updating js-yaml to 3.15.1, 4.3.1]] * 20:18 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 20:15 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 20:15 jhancock@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin1003" * 20:14 jhancock@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin1003" * 19:59 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 19:54 jhancock@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 19:10 brennen@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] (duration: 06m 41s) * 19:04 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns6001.wikimedia.org * 19:04 sukhe@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns6001.wikimedia.org * 19:04 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns5003.wikimedia.org * 19:04 sukhe@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns5003.wikimedia.org * 19:03 brennen@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 18:59 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns5003.wikimedia.org with OS trixie * 18:55 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns6001.wikimedia.org with OS trixie * 18:19 brett@cumin2002: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on P<nowiki>{</nowiki>cp7009.magru.wmnet<nowiki>}</nowiki> and A:cp - 9.2.15 Upgrade () * 18:14 brett@cumin2002: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on P<nowiki>{</nowiki>cp7009.magru.wmnet<nowiki>}</nowiki> and A:cp - 9.2.15 Upgrade () * 18:13 brennen@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 18:12 brett@cumin2002: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 9.2.15 Upgrade () * 18:09 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns5003.wikimedia.org with reason: host reimage * 18:06 brett@cumin2002: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 9.2.15 Upgrade () * 18:06 brennen: 1.47.0-wmf.15 train status ([[phab:T430834|T430834]]) - no current blockers, rolling to group0 * 18:05 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns5003.wikimedia.org with reason: host reimage * 18:05 brett: import trafficserver-9.2.15~deb13+wmf1 into trixie-wikimedia ([[phab:T434478|T434478]]) * 18:01 ladsgroup@cumin1003: END (PASS) - Cookbook sre.mysql.sanitarium_restart (exit_code=0) * 17:58 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns6001.wikimedia.org with reason: host reimage * 17:53 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324369{{!}}Enable desktop lazy loading on group0 (T148047)]] (duration: 07m 31s) * 17:52 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns6001.wikimedia.org with reason: host reimage * 17:49 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 17:49 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitarium_restart (exit_code=99) * 17:49 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 17:49 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7001.magru.wmnet * 17:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti7001.magru.wmnet * 17:48 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 17:47 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1324369{{!}}Enable desktop lazy loading on group0 (T148047)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:45 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1324369{{!}}Enable desktop lazy loading on group0 (T148047)]] * 17:39 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti7001.magru.wmnet * 17:36 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns5003.wikimedia.org with OS trixie * 17:34 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns6001.wikimedia.org with OS trixie * 17:31 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324361{{!}}Move FR config from IS.php to a dedicated file]], [[gerrit:1324363{{!}}Remove $wmg = $wg hacks in CentralAuth (T119117)]] (duration: 12m 23s) * 17:26 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 17:23 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1324361{{!}}Move FR config from IS.php to a dedicated file]], [[gerrit:1324363{{!}}Remove $wmg = $wg hacks in CentralAuth (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:19 sukhe: sudo cumin "A:cp-magru" "run-puppet-agent --enable 'merging CR 1324355'" * 17:18 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1324361{{!}}Move FR config from IS.php to a dedicated file]], [[gerrit:1324363{{!}}Remove $wmg = $wg hacks in CentralAuth (T119117)]] * 17:11 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-master1004.eqiad.wmnet with OS bookworm * 17:11 sukhe: sukhe@cp7005:~$ sudo puppet agent -tv * 17:02 sukhe: sudo cumin "A:cp-magru" "disable-puppet 'merging CR 1324355'" * 16:54 sukhe@dns1004: END - running authdns-update * 16:53 sukhe@dns1004: START - running authdns-update * 16:53 sukhe@dns1004: FAIL - running authdns-update * 16:51 sukhe@dns1004: START - running authdns-update * 16:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-master1004.eqiad.wmnet with reason: host reimage * 16:44 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-master1004.eqiad.wmnet with reason: host reimage * 16:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1157.eqiad.wmnet onto db1272.eqiad.wmnet * 16:40 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1157: Pool db1157.eqiad.wmnet in after cloning * 16:38 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 16:31 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324356{{!}}InitialiseSettings: Fix wgOATHAuthEnforce2FAForAll]] (duration: 06m 52s) * 16:30 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 16:28 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 16:27 reedy@deploy1003: reedy: Continuing with deployment * 16:26 reedy@deploy1003: reedy: Backport for [[gerrit:1324356{{!}}InitialiseSettings: Fix wgOATHAuthEnforce2FAForAll]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:24 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324356{{!}}InitialiseSettings: Fix wgOATHAuthEnforce2FAForAll]] * 16:13 sukhe: restart ntpsec.serviceon dns7001 * 16:09 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324335{{!}}InitialiseSettings: Enable 2FA enforcement on various private wikis (T428103)]] (duration: 06m 40s) * 16:08 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 16:06 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2204: Security update * 16:04 reedy@deploy1003: reedy: Continuing with deployment * 16:04 reedy@deploy1003: reedy: Backport for [[gerrit:1324335{{!}}InitialiseSettings: Enable 2FA enforcement on various private wikis (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:02 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7001.magru.wmnet * 16:02 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324335{{!}}InitialiseSettings: Enable 2FA enforcement on various private wikis (T428103)]] * 16:01 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 15:55 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1157: Pool db1157.eqiad.wmnet in after cloning * 15:54 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 15:54 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 15:41 dancy@deploy1003: Finished scap sync-world: Testing (duration: 06m 28s) * 15:40 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-master1004.eqiad.wmnet with OS bookworm * 15:40 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 15:35 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti4008.ulsfo.wmnet * 15:35 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti4008.ulsfo.wmnet * 15:34 dancy@deploy1003: Started scap sync-world: Testing * 15:34 dancy@deploy1003: Installation of scap version "4.279.0" completed for 3 hosts * 15:34 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-master1004.eqiad.wmnet with OS bookworm * 15:34 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 15:33 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-master1004.eqiad.wmnet with OS bookworm * 15:32 dancy@deploy1003: Installing scap version "4.279.0" for 3 host(s) * 15:32 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324339{{!}}Add /w/deployment-info.php entrypoint]] (duration: 07m 25s) * 15:30 moritzm: failover ganeti master in magru to ganeti7004 * 15:29 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti4008.ulsfo.wmnet * 15:28 dancy@deploy1003: dancy: Continuing with deployment * 15:28 tappof: remove 2026-05 swift log archives from centrallog to free some space ([[phab:T434502|T434502]]) * 15:27 dancy@deploy1003: dancy: Backport for [[gerrit:1324339{{!}}Add /w/deployment-info.php entrypoint]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:25 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324339{{!}}Add /w/deployment-info.php entrypoint]] * 15:20 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2204: Security update * 15:18 dancy@deploy1003: Installation of scap version "4.278.0" completed for 3 hosts * 15:18 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7004.magru.wmnet * 15:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti7004.magru.wmnet * 15:16 dancy@deploy1003: Installing scap version "4.278.0" for 3 host(s) * 15:14 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2204.codfw.wmnet with reason: Maintenance * 15:12 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 15:11 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-master1004.eqiad.wmnet with OS bookworm * 15:11 hashar: Restarting CI Jenkins on contint1003 due to Java upgrade. * 15:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2204 [[phab:T434565|T434565]]', diff saved to https://phabricator.wikimedia.org/P95984 and previous config saved to /var/cache/conftool/dbconfig/20260811-151126-cwilliams.json * 15:10 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti4008.ulsfo.wmnet * 15:10 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti7004.magru.wmnet * 15:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2207 to s2 primary [[phab:T434565|T434565]]', diff saved to https://phabricator.wikimedia.org/P95983 and previous config saved to /var/cache/conftool/dbconfig/20260811-150905-cwilliams.json * 15:08 cezmunsta: Starting s2 codfw failover from db2204 to db2207 - [[phab:T434565|T434565]] * 15:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2207 with weight 0 [[phab:T434565|T434565]]', diff saved to https://phabricator.wikimedia.org/P95982 and previous config saved to /var/cache/conftool/dbconfig/20260811-150402-cwilliams.json * 15:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s2 [[phab:T434565|T434565]] * 14:55 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1010.eqiad.wmnet with OS bookworm * 14:49 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns4003.wikimedia.org with OS trixie * 14:47 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1010.eqiad.wmnet with OS bookworm * 14:47 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7001.wikimedia.org with OS trixie * 14:44 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-presto1010.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:41 btullis@cumin1003: START - Cookbook sre.hosts.provision for host an-presto1010.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:40 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-presto1010.eqiad.wmnet with OS bookworm * 14:39 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 14:39 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-presto1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:36 btullis@cumin1003: START - Cookbook sre.hosts.provision for host an-presto1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:32 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1009.eqiad.wmnet with OS bookworm * 14:32 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 14:31 cwilliams@cumin1003: START - Cookbook sre.mysql.clone of db1157.eqiad.wmnet onto db1272.eqiad.wmnet * 14:30 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7004.magru.wmnet * 14:28 moritzm: failover ganeti master in ulsfo to ganeti4005 * 14:23 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7003.magru.wmnet * 14:23 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti7003.magru.wmnet * 14:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-master1004.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:22 btullis@cumin1003: START - Cookbook sre.hosts.provision for host an-master1004.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:21 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-master1004.eqiad.wmnet with OS bookworm * 14:19 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti4007.ulsfo.wmnet * 14:19 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti4007.ulsfo.wmnet * 14:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1010.eqiad.wmnet with OS bookworm * 14:17 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1008.eqiad.wmnet with OS bookworm * 14:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti7003.magru.wmnet * 14:12 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7003.magru.wmnet * 14:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti4007.ulsfo.wmnet * 14:11 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7002.magru.wmnet * 14:11 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti7002.magru.wmnet * 14:09 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324318{{!}}Revert "wmf-config/ProductionServices: set URL for urldownloader to service record" (T429175)]] (duration: 06m 46s) * 14:09 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7001.wikimedia.org with reason: host reimage * 14:06 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti4007.ulsfo.wmnet * 14:05 kharlan@deploy1003: kharlan: Continuing with deployment * 14:05 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti4006.ulsfo.wmnet * 14:04 jayme: updated calico to v3.30.7 on staging-eqiad - [[phab:T427400|T427400]] * 14:04 kharlan@deploy1003: kharlan: Backport for [[gerrit:1324318{{!}}Revert "wmf-config/ProductionServices: set URL for urldownloader to service record" (T429175)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti4006.ulsfo.wmnet * 14:03 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns4003.wikimedia.org with reason: host reimage * 14:03 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7001.wikimedia.org with reason: host reimage * 14:02 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti7002.magru.wmnet * 14:02 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1324318{{!}}Revert "wmf-config/ProductionServices: set URL for urldownloader to service record" (T429175)]] * 14:02 btullis@dns1004: FAIL - running authdns-update * 14:00 btullis@dns1004: START - running authdns-update * 13:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1008.eqiad.wmnet with reason: host reimage * 13:59 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'. * 13:59 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1313985{{!}}wmf-config/ProductionServices: set URL for urldownloader to service record (T429175)]] (duration: 25m 06s) * 13:58 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7002.magru.wmnet * 13:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti4006.ulsfo.wmnet * 13:57 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns4003.wikimedia.org with reason: host reimage * 13:56 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'. * 13:56 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7001.magru.wmnet * 13:56 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1008.eqiad.wmnet with reason: host reimage * 13:55 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'. * 13:55 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'. * 13:55 kharlan@deploy1003: kharlan, sukhe: Continuing with deployment * 13:53 marostegui: Failover ms2 [[phab:T434288|T434288]] * 13:52 marostegui: Failover ms1 [[phab:T434288|T434288]] * 13:52 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7001.magru.wmnet * 13:51 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti4006.ulsfo.wmnet * 13:48 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti4005.ulsfo.wmnet * 13:48 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti4005.ulsfo.wmnet * 13:44 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti4005.ulsfo.wmnet * 13:40 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1008.eqiad.wmnet with OS bookworm * 13:39 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns4003.wikimedia.org with OS trixie * 13:38 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns7001.wikimedia.org with OS trixie * 13:38 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1008.eqiad.wmnet with OS bookworm * 13:37 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti4005.ulsfo.wmnet * 13:36 kharlan@deploy1003: kharlan, sukhe: Backport for [[gerrit:1313985{{!}}wmf-config/ProductionServices: set URL for urldownloader to service record (T429175)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:34 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1313985{{!}}wmf-config/ProductionServices: set URL for urldownloader to service record (T429175)]] * 13:29 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2034.codfw.wmnet * 13:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2034.codfw.wmnet * 13:25 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.clone (exit_code=99) of db1157.eqiad.wmnet onto db1272.eqiad.wmnet * 13:25 cwilliams@cumin1003: START - Cookbook sre.mysql.clone of db1157.eqiad.wmnet onto db1272.eqiad.wmnet * 13:21 urbanecm@deploy1003: mwscript-k8s job started: namespaceDupes.php --wiki=frwiktionary --fix # [[phab:T415716|T415716]] * 13:21 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2034.codfw.wmnet * 13:20 urbanecm@deploy1003: mwscript-k8s job started: namespaceDupes.php --wiki=frwiktionary # [[phab:T415716|T415716]] * 13:19 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1323350{{!}}[tgwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T415307)]], [[gerrit:1322961{{!}}[slwiki] Revert temporary logo for Wikipedia 25 (Vector legacy + Vector 2022) (T414265)]], [[gerrit:1323827{{!}}[itwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T414320)]] (duration: 08m 00s) * 13:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-coord1003.eqiad.wmnet with OS bookworm * 13:15 urbanecm@deploy1003: urbanecm, superpes: Continuing with deployment * 13:13 urbanecm@deploy1003: urbanecm, superpes: Backport for [[gerrit:1323350{{!}}[tgwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T415307)]], [[gerrit:1322961{{!}}[slwiki] Revert temporary logo for Wikipedia 25 (Vector legacy + Vector 2022) (T414265)]], [[gerrit:1323827{{!}}[itwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T414320)]] synced to the testservers (see https://wiki * 13:11 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1323350{{!}}[tgwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T415307)]], [[gerrit:1322961{{!}}[slwiki] Revert temporary logo for Wikipedia 25 (Vector legacy + Vector 2022) (T414265)]], [[gerrit:1323827{{!}}[itwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T414320)]] * 13:11 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1323779{{!}}[ukwiki] Remove reviewer usergroup (T434252)]], [[gerrit:1323312{{!}}[frwiktionary] Add new Schème namespace and its talk (T415716)]] (duration: 06m 49s) * 13:10 marostegui@dns1004: END - running authdns-update * 13:08 marostegui@dns1004: START - running authdns-update * 13:07 marostegui@cumin1003: dbctl commit (dc=all): 'Repool ms2 [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95980 and previous config saved to /var/cache/conftool/dbconfig/20260811-130725-marostegui.json * 13:06 urbanecm@deploy1003: urbanecm, superpes: Continuing with deployment * 13:06 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1266 to dbctl [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95979 and previous config saved to /var/cache/conftool/dbconfig/20260811-130627-marostegui.json * 13:06 urbanecm@deploy1003: urbanecm, superpes: Backport for [[gerrit:1323779{{!}}[ukwiki] Remove reviewer usergroup (T434252)]], [[gerrit:1323312{{!}}[frwiktionary] Add new Schème namespace and its talk (T415716)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:04 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1323779{{!}}[ukwiki] Remove reviewer usergroup (T434252)]], [[gerrit:1323312{{!}}[frwiktionary] Add new Schème namespace and its talk (T415716)]] * 12:59 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db2253.codfw.wmnet,db[1151,1266].eqiad.wmnet with reason: Switching over ms2 * 12:54 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1157: Using as clone source * 12:53 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1157: Using as clone source * 12:51 marostegui@cumin1003: dbctl commit (dc=all): 'Depool ms2 [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95977 and previous config saved to /var/cache/conftool/dbconfig/20260811-125129-marostegui.json * 12:47 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 12:46 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 12:45 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 12:44 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'recommendation-api-ng' for release 'main' . * 12:44 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'recommendation-api-ng' for release 'main' . * 12:43 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'recommendation-api-ng' for release 'main' . * 12:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-coord1003.eqiad.wmnet with reason: host reimage * 12:43 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'ores-legacy' for release 'main' . * 12:42 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'ores-legacy' for release 'main' . * 12:42 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2165: Security update * 12:40 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-coord1003.eqiad.wmnet with reason: host reimage * 12:39 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'ores-legacy' for release 'main' . * 12:38 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' . * 12:38 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' . * 12:37 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' . * 12:34 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2034.codfw.wmnet * 12:30 jmm@cumin2002: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti-test2001.codfw.wmnet * 12:30 jmm@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host ganeti-test2001.codfw.wmnet * 12:25 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1179.eqiad.wmnet onto db1278.eqiad.wmnet * 12:25 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1179: Pool db1179.eqiad.wmnet in after cloning * 12:23 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-coord1003.eqiad.wmnet with OS bookworm * 12:22 moritzm: failover ganeti master in codfw/routed to ganeti2033 * 12:22 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2033.codfw.wmnet * 12:22 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2033.codfw.wmnet * 12:19 jmm@cumin2002: START - Cookbook sre.hosts.reboot-single for host ganeti-test2001.codfw.wmnet * 12:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 12:18 jmm@cumin2002: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti-test2001.codfw.wmnet * 12:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1008.eqiad.wmnet with OS bookworm * 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2033.codfw.wmnet * 12:07 moritzm: failover ganeti master in ganeti/test to ganeti-test2003 * 12:04 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 12:03 jmm@cumin2003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti4005.ulsfo.wmnet * 12:03 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti4005.ulsfo.wmnet * 12:00 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324283{{!}}Use maximum compression level in SqlBlobStore and SqlBagOStuff (T428377)]] (duration: 11m 37s) * 11:57 jmm@cumin2002: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti-test2002.codfw.wmnet * 11:57 jmm@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti-test2002.codfw.wmnet * 11:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2165: Security update * 11:54 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 11:52 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1324283{{!}}Use maximum compression level in SqlBlobStore and SqlBagOStuff (T428377)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:51 jmm@cumin2002: START - Cookbook sre.hosts.reboot-single for host ganeti-test2002.codfw.wmnet * 11:50 jmm@cumin2002: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti-test2002.codfw.wmnet * 11:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2165.codfw.wmnet with reason: Maintenance * 11:48 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@050d19e] (releasing): [[phab:T434186|T434186]] (duration: 01m 14s) * 11:48 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1324283{{!}}Use maximum compression level in SqlBlobStore and SqlBagOStuff (T428377)]] * 11:47 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@050d19e] (releasing): [[phab:T434186|T434186]] * 11:44 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@050d19e] (releasing): test jenkins deploy for [[phab:T434186|T434186]] (duration: 01m 08s) * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2165 [[phab:T434514|T434514]]', diff saved to https://phabricator.wikimedia.org/P95969 and previous config saved to /var/cache/conftool/dbconfig/20260811-114352-cwilliams.json * 11:43 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@050d19e] (releasing): test jenkins deploy for [[phab:T434186|T434186]] * 11:42 jmm@cumin2003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti-test2003.codfw.wmnet * 11:42 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti-test2003.codfw.wmnet * 11:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2161 to s8 primary [[phab:T434514|T434514]]', diff saved to https://phabricator.wikimedia.org/P95968 and previous config saved to /var/cache/conftool/dbconfig/20260811-114136-cwilliams.json * 11:40 cezmunsta: Starting s8 codfw failover from db2165 to db2161 - [[phab:T434514|T434514]] * 11:40 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1179: Pool db1179.eqiad.wmnet in after cloning * 11:36 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti-test2003.codfw.wmnet * 11:36 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti-test2003.codfw.wmnet * 11:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2161 with weight 0 [[phab:T434514|T434514]]', diff saved to https://phabricator.wikimedia.org/P95966 and previous config saved to /var/cache/conftool/dbconfig/20260811-113449-cwilliams.json * 11:34 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 25 hosts with reason: Primary switchover s8 [[phab:T434514|T434514]] * 11:29 moritzm: installing Python 3.11 security updates * 11:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-presto1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 11:26 btullis@cumin1003: START - Cookbook sre.hosts.provision for host an-presto1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 11:23 btullis@dns1004: END - running authdns-update * 11:21 btullis@dns1004: START - running authdns-update * 11:20 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-presto1008.eqiad.wmnet with OS bookworm * 11:20 moritzm: installing curl security updates * 11:11 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1007.eqiad.wmnet with OS bookworm * 10:45 tappof: bump space for prometheus k8s-dse in codfw * 10:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1007.eqiad.wmnet with reason: host reimage * 10:38 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1007.eqiad.wmnet with reason: host reimage * 10:37 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-coord1004.eqiad.wmnet with OS bookworm * 10:35 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1008.eqiad.wmnet with OS bookworm * 10:34 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1179: Depool db1179.eqiad.wmnet to then clone it to db1278.eqiad.wmnet - marostegui@cumin1003 * 10:34 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1008.eqiad.wmnet with OS bookworm * 10:25 fceratto@cumin1003: dbctl commit (dc=all): 'Remove db1177 [[phab:T433474|T433474]]', diff saved to https://phabricator.wikimedia.org/P95964 and previous config saved to /var/cache/conftool/dbconfig/20260811-102527-fceratto.json * 10:22 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1008.eqiad.wmnet with OS bookworm * 10:22 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1007.eqiad.wmnet with OS bookworm * 10:21 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1006.eqiad.wmnet with OS bookworm * 10:20 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 10:18 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1179: Depool db1179.eqiad.wmnet to then clone it to db1278.eqiad.wmnet - marostegui@cumin1003 * 10:18 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1179.eqiad.wmnet onto db1278.eqiad.wmnet * 10:17 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 10:17 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 10:14 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 10:09 blake@deploy1003: Stopping before sync operations * 10:09 blake@deploy1003: Started scap sync-world: Non-deployment run to populate release values for [[phab:T427668|T427668]] * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 10:04 fceratto@cumin1003: Removing db1177 from zarcillo [[phab:T433474|T433474]] * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1177.eqiad.wmnet * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1177.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:03 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1177.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:00 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1006.eqiad.wmnet with reason: host reimage * 09:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-coord1004.eqiad.wmnet with reason: host reimage * 09:57 marostegui: Failover m1 from db1164 to db1213 - [[phab:T434493|T434493]] * 09:57 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1006.eqiad.wmnet with reason: host reimage * 09:55 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:54 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2232].codfw.wmnet,db[1164,1213,1217].eqiad.wmnet with reason: Primary switchover m1 [[phab:T434493|T434493]] * 09:52 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-coord1004.eqiad.wmnet with reason: host reimage * 09:49 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1213.eqiad.wmnet with OS trixie * 09:49 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1177.eqiad.wmnet * 09:41 moritzm: installing Linux 6.12.101 on Trixie hosts * 09:40 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1006.eqiad.wmnet with OS bookworm * 09:35 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-coord1004.eqiad.wmnet with OS bookworm * 09:28 moritzm: installing node-tar security updates * 09:27 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1213.eqiad.wmnet with reason: host reimage * 09:22 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1213.eqiad.wmnet with reason: host reimage * 09:09 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1177: Decommission * 09:08 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db1177: Decommission * 09:08 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 09:08 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.decommission (exit_code=99) * 09:06 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1213.eqiad.wmnet with OS trixie * 09:06 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 09:05 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1213.eqiad.wmnet with reason: Reimage * 08:53 marostegui@dns1004: END - running authdns-update * 08:51 marostegui@dns1004: START - running authdns-update * 08:48 marostegui: Switchover ms1 master in eqiad [[phab:T434288|T434288]] * 08:48 marostegui@cumin1003: dbctl commit (dc=all): 'Repool ms1 [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95962 and previous config saved to /var/cache/conftool/dbconfig/20260811-084804-marostegui.json * 08:40 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1267 to dbctl [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95961 and previous config saved to /var/cache/conftool/dbconfig/20260811-084054-marostegui.json * 08:29 marostegui: Failover m1 from db1213 to db1164 - [[phab:T434043|T434043]] * 08:25 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2232].codfw.wmnet,db[1164,1213,1217].eqiad.wmnet with reason: Primary switchover m1 [[phab:T434043|T434043]] * 08:22 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db2251.codfw.wmnet,db[1152,1267].eqiad.wmnet with reason: Switching over ms1 * 08:22 marostegui@cumin1003: dbctl commit (dc=all): 'Depool ms1 [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95960 and previous config saved to /var/cache/conftool/dbconfig/20260811-082201-marostegui.json * 08:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: Switching over ms1 * 08:20 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.parsercache (exit_code=99) * 08:20 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 08:20 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1152: Switching over ms1 * 08:19 slyngshede@dns1004: END - running authdns-update * 08:18 moritzm: installing openjdk-21 security updates * 08:17 slyngshede@dns1004: START - running authdns-update * 08:16 moritzm: imported jenkins 2.568.2 to thirdparty/jenkins for trixie-wikimedia * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.12 (duration: 02m 26s) * 03:36 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] (duration: 33m 33s) * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 35s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-10 == * 14:54 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1323973{{!}}mmv.bootstrap: Fix getUrlParam to account for TIFF lossy/lossless param (T434333)]] (duration: 11m 24s) * 14:50 krinkle@deploy1003: krinkle: Continuing with deployment * 14:45 krinkle@deploy1003: krinkle: Backport for [[gerrit:1323973{{!}}mmv.bootstrap: Fix getUrlParam to account for TIFF lossy/lossless param (T434333)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:43 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1323973{{!}}mmv.bootstrap: Fix getUrlParam to account for TIFF lossy/lossless param (T434333)]] * 14:07 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1323967{{!}}updateIsActiveFlagForMentees: Commit the final partial batch (T432959)]] (duration: 10m 33s) * 13:56 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1323967{{!}}updateIsActiveFlagForMentees: Commit the final partial batch (T432959)]] * 13:45 dani@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply * 13:45 dani@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply * 13:45 dani@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply * 13:45 dani@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply * 13:45 dani@deploy1003: helmfile [staging] DONE helmfile.d/services/miscweb: apply * 13:44 dani@deploy1003: helmfile [staging] START helmfile.d/services/miscweb: apply * 13:38 wmde-fisch@deploy1003: Finished scap sync-world: Backport for [[gerrit:1323939{{!}}Enable sub-references on more group2 wikis (batch3) (T432731)]] (duration: 33m 21s) * 13:25 wmde-fisch@deploy1003: wmde-fisch: Continuing with deployment * 13:22 wmde-fisch@deploy1003: wmde-fisch: Backport for [[gerrit:1323939{{!}}Enable sub-references on more group2 wikis (batch3) (T432731)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:05 wmde-fisch@deploy1003: Started scap sync-world: Backport for [[gerrit:1323939{{!}}Enable sub-references on more group2 wikis (batch3) (T432731)]] * 07:57 hashar@deploy1003: Finished deploy [integration/docroot@7772132]: update build dependencies (duration: 00m 13s) * 07:57 hashar@deploy1003: Started deploy [integration/docroot@7772132]: update build dependencies * 07:35 _joe_: restarting squid on urldownloader1006 * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 48s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-09 == * 16:01 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:01 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:01 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:00 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 36s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-08 == * 05:31 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9] (wcqs): [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) (duration: 02m 36s) * 05:28 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9] (wcqs): [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) * 04:56 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 04:55 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 04:47 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) (duration: 19m 22s) * 04:28 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) * 04:19 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) (duration: 00m 06s) * 04:18 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) * 04:17 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) (duration: 00m 28s) * 04:16 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) * 03:52 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 03:52 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 34s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-07 == * 23:30 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:29 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 22:45 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 22:43 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 22:41 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 22:41 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 20:54 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 20:32 andrewbogott: restarting puppetserver service on puppetserver* for [[phab:T434339|T434339]] * 19:52 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:45 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 19:32 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:25 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:22 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 19:21 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 18:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:41 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 18:35 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 18:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 18:22 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 18:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 18:16 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 18:12 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:09 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:08 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:07 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:04 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:01 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:00 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:00 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 17:59 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 17:25 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 17:14 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 17:13 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 17:13 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 17:13 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:54 maryum: Deployed security fix for [[phab:T434278|T434278]] * 16:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 16:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 16:27 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --verbose --scoreLessThan=0.7 --exceptDatasetChecksums=[[phab:T434319|T434319]]-enwiki-models.txt # [[phab:T434319|T434319]] * 16:06 cdobbins@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp5022.eqsin.wmnet with OS trixie * 15:13 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 14:19 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 14:17 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 13:50 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1156.eqiad.wmnet onto db1271.eqiad.wmnet * 13:50 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1271: Pool db1271.eqiad.wmnet in after cloning * 13:02 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1271: Pool db1271.eqiad.wmnet in after cloning * 12:19 jayme: updated calico to v3.30.7 on staging-codfw - [[phab:T427400|T427400]] * 12:09 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 12:06 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 12:06 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 12:05 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 12:02 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1156: Pool db1156.eqiad.wmnet in after cloning * 11:38 bjensen: sudo -i reprepro -C main include trixie-wikimedia $<nowiki>{</nowiki>HOME<nowiki>}</nowiki>/httpbb/trixie/httpbb_$<nowiki>{</nowiki>VERSION?<nowiki>}</nowiki>-1+deb13u1_amd64.changes #[[phab:T434052|T434052]] * 11:35 bjensen: sudo -i reprepro -C main include bookworm-wikimedia $<nowiki>{</nowiki>HOME<nowiki>}</nowiki>/httpbb/bookworm/httpbb_$<nowiki>{</nowiki>VERSION?<nowiki>}</nowiki>-1_amd64.changes #[[phab:T434052|T434052]] * 11:30 marostegui@cumin1003: dbctl commit (dc=all): 'Adding db1271 to dbctl', diff saved to https://phabricator.wikimedia.org/P95945 and previous config saved to /var/cache/conftool/dbconfig/20260807-113006-marostegui.json * 11:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1156: Pool db1156.eqiad.wmnet in after cloning * 10:23 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 10:22 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 10:22 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 10:21 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 10:20 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 10:20 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 10:19 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 10:18 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 10:06 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on 21 hosts with reason: cloning * 10:01 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1156: Depool db1156.eqiad.wmnet to then clone it to db1271.eqiad.wmnet - marostegui@cumin1003 * 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1156: Depool db1156.eqiad.wmnet to then clone it to db1271.eqiad.wmnet - marostegui@cumin1003 * 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1156.eqiad.wmnet onto db1271.eqiad.wmnet * 09:15 jynus: started stress testing db1245 dbs [[phab:T431115|T431115]] * 08:19 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:18 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:16 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:14 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:13 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:10 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:06 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:05 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:00 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 10 days, 0:00:00 on ml-serve1015.eqiad.wmnet with reason: Downtime to get full picture of current BIOS settings beyond what Redfish shows * 08:00 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 07:54 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 07:54 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 07:53 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:52 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:51 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:50 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:49 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:48 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:47 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:45 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:45 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:41 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:38 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:37 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 06:35 jayme: updated istio to 1.29.4 on wikikube eqiad - [[phab:T427401|T427401]] * 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1178.eqiad.wmnet * 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1178.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 06:06 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1178.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 05:55 marostegui@cumin1003: START - Cookbook sre.dns.netbox * 05:49 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1178.eqiad.wmnet * 05:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 05:46 marostegui@cumin1003: Removing db1178 from zarcillo [[phab:T433471|T433471]] * 05:45 marostegui@cumin1003: START - Cookbook sre.mysql.decommission * 02:42 denisse: Extended volume on prometheus2008 for the disk space alert as per https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_host_running_out_of_space * 02:37 denisse: Extended volume on prometheus2007 tor the disk space alert as per https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_host_running_out_of_space * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 56s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-06 == * 21:39 maryum: Deploy security patch for [[phab:T433070|T433070]] * 21:29 maryum: Deploy security patch for [[phab:T434189|T434189]] * 20:48 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] (duration: 08m 12s) * 20:44 aude@deploy1003: lmora, aude, anzx: Continuing with deployment * 20:41 aude@deploy1003: lmora, aude, anzx: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be * 20:41 ebernhardson@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:41 ebernhardson@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 20:40 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] * 20:37 ebernhardson@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:37 ebernhardson@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 20:32 ebernhardson@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:32 ebernhardson@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 20:31 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] (duration: 06m 41s) * 20:27 cjming@deploy1003: cjming, ebernhardson, chlod: Continuing with deployment * 20:26 cjming@deploy1003: cjming, ebernhardson, chlod: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:24 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] * 20:18 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] (duration: 09m 22s) * 20:14 cjming@deploy1003: cjming, tsev: Continuing with deployment * 20:11 cjming@deploy1003: cjming, tsev: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:09 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] * 19:41 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply * 19:40 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply * 19:31 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 19:31 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 19:00 cdobbins@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cp5022.eqsin.wmnet with OS trixie * 18:25 ladsgroup@deploy1003: Finished scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) (duration: 06m 08s) * 18:19 ladsgroup@deploy1003: Started scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) * 18:18 ladsgroup@deploy1003: Stopping before sync operations * 18:17 ladsgroup@deploy1003: Started scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) * 17:55 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 16:50 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 16:35 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1001.eqiad.wmnet with OS bookworm * 16:19 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1002.eqiad.wmnet with reason: host reimage * 16:16 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1002.eqiad.wmnet with reason: host reimage * 16:05 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1001.eqiad.wmnet with reason: host reimage * 16:00 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1001.eqiad.wmnet with reason: host reimage * 15:57 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 15:43 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm * 15:29 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1001.eqiad.wmnet with OS bookworm * 15:29 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:58 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm * 14:57 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-drmrs ([[phab:T428495|T428495]]) * 14:55 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-drmrs ([[phab:T428495|T428495]]) * 14:55 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-ui1001.eqiad.wmnet with OS bookworm * 14:54 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-presto1001.eqiad.wmnet with OS bookworm * 14:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-magru ([[phab:T428495|T428495]]) * 14:49 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-magru ([[phab:T428495|T428495]]) * 14:48 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1001.eqiad.wmnet with OS bookworm * 14:46 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-esams ([[phab:T428495|T428495]]) * 14:44 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-esams ([[phab:T428495|T428495]]) * 14:43 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:42 brouberol@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:42 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:42 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 14:40 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 14:40 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:38 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-ui1001.eqiad.wmnet with reason: host reimage * 14:34 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-presto1001.eqiad.wmnet with reason: host reimage * 14:28 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-ui1001.eqiad.wmnet with reason: host reimage * 14:27 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-presto1001.eqiad.wmnet with reason: host reimage * 14:23 sukhe: sudo cumin -b2 'A:cp-text' "run-puppet-agent --enable 'merging CR 1290731'": [[phab:T425441|T425441]] * 14:18 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo for hosts in the wikimedia.org domain - [[phab:T428495|T428495]] * 14:16 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-presto1001.eqiad.wmnet with OS bookworm * 14:14 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-ui1001.eqiad.wmnet with OS bookworm * 14:12 sukhe: sudo cumin 'A:cp-text' "disable-puppet 'merging CR 1290731'": [[phab:T425441|T425441]] * 14:11 swfrench-wmf: restarted navtiming on webperf1003 - [[phab:T428495|T428495]] * 14:04 swfrench-wmf: begin rolling restart of confd in drmrs, eqiad, esams, magru - [[phab:T428495|T428495]] * 14:04 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm * 14:04 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:02 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-client1002.eqiad.wmnet with OS bookworm * 13:58 swfrench-wmf: authdns update to direct eqiad-associated etcd clients back to eqiad - [[phab:T428495|T428495]] * 13:58 swfrench@dns1004: END - running authdns-update * 13:56 swfrench@dns1004: START - running authdns-update * 13:49 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:44 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:31 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 13:29 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 13:26 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 13:23 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 13:22 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 13:19 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 13:18 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 13:18 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 13:17 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 13:16 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 13:13 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 13:11 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 13:09 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 13:06 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 13:06 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-client1002.eqiad.wmnet with OS bookworm * 13:05 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revision-models' for release 'main' . * 13:05 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:05 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revision-models' for release 'main' . * 13:04 brouberol@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-test-client1002.eqiad.wmnet with OS bookworm * 13:04 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 13:03 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 13:02 aikochou@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:00 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'readability' for release 'main' . * 12:59 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'readability' for release 'main' . * 12:58 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 12:57 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'logo-detection' for release 'main' . * 12:57 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'logo-detection' for release 'main' . * 12:57 aikochou@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 12:55 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:54 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:53 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 12:53 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:50 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 12:48 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 12:46 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'article-models' for release 'main' . * 12:45 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'article-models' for release 'main' . * 12:41 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . * 12:39 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . * 12:38 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-client1002.eqiad.wmnet with OS bookworm * 12:12 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply * 12:12 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply * 12:09 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:08 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 11:58 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2187: Security update * 11:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:24 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:16 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 11:15 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 11:10 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2187: Security update * 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2187.codfw.wmnet with reason: Maintenance * 10:56 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 10:56 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 10:56 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 10:56 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 10:54 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 10:53 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 10:09 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2187: Security update * 10:07 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2187: Security update * 09:39 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms2', diff saved to https://phabricator.wikimedia.org/P95929 and previous config saved to /var/cache/conftool/dbconfig/20260806-093908-marostegui.json * 09:36 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1178 from dbctl [[phab:T433471|T433471]]', diff saved to https://phabricator.wikimedia.org/P95928 and previous config saved to /var/cache/conftool/dbconfig/20260806-093632-marostegui.json * 09:33 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 09:31 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 09:30 topranks: bounce cr3-eqsin<->cr2-eqiad bgp session to disable no-prepend command * 09:20 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2253.codfw.wmnet,db1151.eqiad.wmnet with reason: cloning * 09:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1151: Cloning * 09:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:19 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 09:19 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1151: Cloning * 09:10 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 09:09 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup2003.codfw.wmnet * 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup2003.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 09:06 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup2003.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 09:03 klausman@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:02 jynus@cumin1003: START - Cookbook sre.dns.netbox * 09:02 klausman@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 08:57 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup2003.codfw.wmnet * 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup1003.eqiad.wmnet * 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:54 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms3', diff saved to https://phabricator.wikimedia.org/P95925 and previous config saved to /var/cache/conftool/dbconfig/20260806-085422-marostegui.json * 08:53 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:46 jynus@cumin1003: START - Cookbook sre.dns.netbox * 08:39 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup1003.eqiad.wmnet * 08:29 XioNoX: push pfw policy - [[phab:T434115|T434115]] * 08:14 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 08:00 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 08:00 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:58 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revision-models' for release 'main' . * 07:56 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 07:54 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'readability' for release 'main' . * 07:53 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'logo-detection' for release 'main' . * 07:51 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'llm' for release 'main' . * 07:48 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . * 07:37 jayme: updated istio to 1.29.4 on wikikube codfw - [[phab:T427401|T427401]] * 07:08 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2252.codfw.wmnet,db1153.eqiad.wmnet with reason: cloning * 07:07 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1153: Cloning * 07:07 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1153: Cloning * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 40s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-05 == * 23:24 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1009.eqiad.wmnet with OS bookworm * 23:03 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1009.eqiad.wmnet with reason: host reimage * 22:59 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1009.eqiad.wmnet with reason: host reimage * 22:43 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1009.eqiad.wmnet with OS bookworm * 22:38 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1009.eqiad.wmnet * 22:34 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1009.eqiad.wmnet * 22:25 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1008.eqiad.wmnet with OS bookworm * 22:04 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1008.eqiad.wmnet with reason: host reimage * 22:00 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1008.eqiad.wmnet with reason: host reimage * 21:48 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:47 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:46 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:44 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1008.eqiad.wmnet with OS bookworm * 21:43 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:41 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1008.eqiad.wmnet * 21:36 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1008.eqiad.wmnet * 21:14 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:12 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad * 21:12 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad * 21:10 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=eqiad * 21:08 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:07 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:07 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=eqiad * 21:04 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:03 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1006 * 21:02 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1006 * 21:00 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:56 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:56 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:55 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 20:55 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 20:55 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1007.eqiad.wmnet with OS bookworm * 20:51 vriley@cumin1003: START - Cookbook sre.dns.netbox * 20:43 ebernhardson: [[phab:T434008|T434008]]: changing cloudelastic:9643 from auto_expand_replicas to number_of_replicas * 20:34 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1007.eqiad.wmnet with reason: host reimage * 20:27 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1007.eqiad.wmnet with reason: host reimage * 20:24 cjming: end of UTC late backport window * 20:23 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] (duration: 06m 26s) * 20:18 cjming@deploy1003: cjming: Continuing with deployment * 20:18 cjming@deploy1003: cjming: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:16 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] * 20:12 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1007.eqiad.wmnet with OS bookworm * 20:12 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] (duration: 08m 41s) * 20:08 swfrench@cumin2002: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host conf1007.eqiad.wmnet with OS bookworm * 20:08 jforrester@deploy1003: jforrester: Continuing with deployment * 20:07 jforrester@deploy1003: jforrester: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:03 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] * 19:51 inflatador: [bking@puppetserver1001] ~$ sudo puppetserver ca sign --certname an-worker1189.eqiad.wmnet [[phab:T434142|T434142]] * 19:47 bking@cumin2003: DONE (FAIL) - Cookbook sre.puppet.renew-cert (exit_code=99) for an-worker1189.eqiad.wmnet: Renew puppet certificate - bking@cumin2003 * 19:46 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:30 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1007.eqiad.wmnet with OS trixie * 19:30 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 19:29 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 19:20 swfrench-wmf: silenced EtcdRelicationDown 0cb709a9-f244-4f1e-971f-{{Gerrit|440ec65e7fd7}} - [[phab:T428495|T428495]] * 19:13 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1007.eqiad.wmnet with OS bookworm * 19:12 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1007.eqiad.wmnet with reason: host reimage * 19:09 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1007.eqiad.wmnet * 19:07 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1007.eqiad.wmnet with reason: host reimage * 19:03 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1007.eqiad.wmnet * 18:52 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1007.eqiad.wmnet with OS trixie * 18:52 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1007.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:35 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 18:34 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 18:34 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 18:30 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1007.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:28 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:28 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1007] - vriley@cumin1003" * 18:27 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1007] - vriley@cumin1003" * 18:23 vriley@cumin1003: START - Cookbook sre.dns.netbox * 18:22 vriley@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 18:22 robh@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:19 vriley@cumin1003: START - Cookbook sre.dns.netbox * 18:13 robh@cumin2002: START - Cookbook sre.hosts.provision for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:31 jasmine@cumin2002: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-main-eqiad * 17:12 mutante: LDAP - added vwalters to group ciadmin - [[phab:T433615|T433615]] * 16:58 aokoth@deploy1003: Finished deploy [phabricator/deployment@e2ebca5]: Deploy Phab (duration: 00m 34s) * 16:57 aokoth@deploy1003: Started deploy [phabricator/deployment@e2ebca5]: Deploy Phab * 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad * 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=eqiad * 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad * 16:53 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:41 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2187.codfw.wmnet * 16:41 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2187.codfw.wmnet * 16:41 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker2187.codfw.wmnet * 16:41 cgoubert@cumin2003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker2187.codfw.wmnet * 16:40 jasmine@cumin2002: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-main-eqiad * 16:40 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:34 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:25 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-magru and A:liberica ([[phab:T428495|T428495]]) * 16:23 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-magru and A:liberica ([[phab:T428495|T428495]]) * 16:20 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-drmrs and A:liberica ([[phab:T428495|T428495]]) * 16:19 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-drmrs and A:liberica ([[phab:T428495|T428495]]) * 16:18 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-esams and A:liberica ([[phab:T428495|T428495]]) * 16:16 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-esams and A:liberica ([[phab:T428495|T428495]]) * 16:06 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1159.eqiad.wmnet * 16:06 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1159.eqiad.wmnet * 16:06 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1159.eqiad.wmnet * 16:05 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] (duration: 09m 11s) * 15:58 reedy@deploy1003: reedy: Continuing with deployment * 15:58 reedy@deploy1003: reedy: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:56 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] * 15:54 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1159.eqiad.wmnet with OS trixie * 15:38 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:33 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1159.eqiad.wmnet with reason: host reimage * 15:32 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:27 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1159.eqiad.wmnet with reason: host reimage * 15:10 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1159 * 15:10 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1159 * 15:00 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo for hosts in the wikimedia.org domain - [[phab:T428495|T428495]] * 14:55 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS trixie * 14:54 swfrench-wmf: restarted navtiming on webperf1003 - [[phab:T428495|T428495]] * 14:52 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1159 * 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1159.eqiad.wmnet 129.48.64.10.in-addr.arpa 9.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:52 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1159.eqiad.wmnet 129.48.64.10.in-addr.arpa 9.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1159 - jayme@cumin1003" * 14:52 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1159 - jayme@cumin1003" * 14:49 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:48 jayme@cumin1003: START - Cookbook sre.dns.netbox * 14:47 swfrench-wmf: begin rolling restart of confd in drmrs, eqiad, esams, magru - [[phab:T428495|T428495]] * 14:47 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1159 * 14:46 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:46 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:46 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1159.eqiad.wmnet with OS trixie * 14:45 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:44 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:44 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1159.eqiad.wmnet * 14:43 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:43 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1159.eqiad.wmnet * 14:43 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:43 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1159.eqiad.wmnet * 14:43 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:43 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1157.eqiad.wmnet * 14:43 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1157.eqiad.wmnet * 14:43 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1157.eqiad.wmnet * 14:42 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:42 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:42 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:41 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:41 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:41 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:41 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1003.eqiad.wmnet with OS bookworm * 14:39 swfrench-wmf: authdns update to direct eqiad-associated etcd clients to codfw - [[phab:T428495|T428495]] * 14:39 swfrench@dns1004: END - running authdns-update * 14:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 14:37 swfrench@dns1004: START - running authdns-update * 14:37 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:37 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:35 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:35 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 14:28 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:27 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1157.eqiad.wmnet with OS trixie * 14:27 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:27 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:27 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 14:26 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 14:26 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:26 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 14:26 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 14:26 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:26 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host search-loader1002.eqiad.wmnet with OS trixie * 14:19 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:19 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:15 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1003.eqiad.wmnet with reason: host reimage * 14:14 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:14 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:13 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 * 14:13 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host mc2046 * 14:13 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS trixie * 14:11 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1003.eqiad.wmnet with reason: host reimage * 14:10 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:09 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:09 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:09 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:08 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:08 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1157.eqiad.wmnet with reason: host reimage * 14:08 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:04 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 14:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on search-loader1002.eqiad.wmnet with reason: host reimage * 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=eqiad * 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=eqiad * 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=eqiad * 14:00 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:59 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:58 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1157.eqiad.wmnet with reason: host reimage * 13:57 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on search-loader1002.eqiad.wmnet with reason: host reimage * 13:54 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1003.eqiad.wmnet with OS bookworm * 13:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host search-loader1002.eqiad.wmnet with OS trixie * 13:43 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1157 * 13:42 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1157 * 13:40 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1157 * 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1157.eqiad.wmnet 183.32.64.10.in-addr.arpa 3.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1157.eqiad.wmnet 183.32.64.10.in-addr.arpa 3.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1157 - jayme@cumin1003" * 13:39 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1157 - jayme@cumin1003" * 13:39 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] (duration: 07m 00s) * 13:35 jayme@cumin1003: START - Cookbook sre.dns.netbox * 13:35 reedy@deploy1003: reedy: Continuing with deployment * 13:34 reedy@deploy1003: reedy: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:32 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] * 13:23 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1157 * 13:22 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1157.eqiad.wmnet with OS trixie * 13:22 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1157.eqiad.wmnet * 13:22 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1157.eqiad.wmnet * 13:21 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1157.eqiad.wmnet * 13:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1156.eqiad.wmnet * 13:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1156.eqiad.wmnet * 13:15 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1156.eqiad.wmnet * 13:01 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1156.eqiad.wmnet with OS trixie * 12:42 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1156.eqiad.wmnet with reason: host reimage * 12:38 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1156.eqiad.wmnet with reason: host reimage * 12:32 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 12:31 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 12:30 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 12:28 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 12:26 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 12:24 topranks: update bgp confed settings in eqsin * 12:22 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1156 * 12:22 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1156 * 12:22 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 12:19 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1156 * 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1156.eqiad.wmnet 110.32.64.10.in-addr.arpa 0.1.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:19 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1156.eqiad.wmnet 110.32.64.10.in-addr.arpa 0.1.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1156 - jayme@cumin1003" * 12:19 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1156 - jayme@cumin1003" * 12:17 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:14 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS trixie * 12:09 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:06 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:04 jayme@cumin1003: START - Cookbook sre.dns.netbox * 12:04 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 12:02 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'article-models' for release 'main' . * 12:01 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1156 * 12:01 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1156.eqiad.wmnet with OS trixie * 11:59 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1156.eqiad.wmnet * 11:59 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1156.eqiad.wmnet * 11:59 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1156.eqiad.wmnet * 11:57 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 11:53 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 11:53 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:52 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:52 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:50 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:50 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:50 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:49 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:48 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:47 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:47 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:45 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:45 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:44 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:44 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:44 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:43 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:42 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:38 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:35 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 * 11:35 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host mc2046 * 11:34 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS trixie * 11:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:27 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:21 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:21 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:18 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:18 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:18 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:18 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:13 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:13 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:09 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:08 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:07 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:06 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:06 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:05 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:05 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:04 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 11:04 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:24 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:24 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:17 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:16 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1155.eqiad.wmnet * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1155.eqiad.wmnet * 10:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1155.eqiad.wmnet * 10:14 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:14 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:11 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:11 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:10 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:09 aikochou@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop: sync * 10:09 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:09 aikochou@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop: sync * 10:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:07 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:05 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:05 aikochou@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop: sync * 10:05 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:05 aikochou@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop: sync * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:04 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:04 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:04 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1155.eqiad.wmnet with OS trixie * 09:52 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms1', diff saved to https://phabricator.wikimedia.org/P95918 and previous config saved to /var/cache/conftool/dbconfig/20260805-095212-marostegui.json * 09:44 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1152: after cloning * 09:44 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.parsercache (exit_code=99) * 09:44 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 09:44 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1152: after cloning * 09:43 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1155.eqiad.wmnet with reason: host reimage * 09:40 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1155.eqiad.wmnet with reason: host reimage * 09:32 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 09:32 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:31 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 09:31 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:27 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1155 * 09:27 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1155 * 09:25 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 09:24 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 09:24 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 09:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:23 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 09:23 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 09:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:22 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 09:22 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 09:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:20 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 09:20 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:17 XioNoX: push pfw policies - [[phab:T434038|T434038]] * 09:14 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1155 * 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1155.eqiad.wmnet 109.32.64.10.in-addr.arpa 9.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1155.eqiad.wmnet 109.32.64.10.in-addr.arpa 9.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1155 - jayme@cumin1003" * 09:14 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1155 - jayme@cumin1003" * 09:10 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 09:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:09 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2251.codfw.wmnet,db1152.eqiad.wmnet with reason: cloning * 09:09 jayme@cumin1003: START - Cookbook sre.dns.netbox * 09:08 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 09:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: Cloning * 09:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:05 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 09:05 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1152: Cloning * 08:38 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1155 * 08:37 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1155.eqiad.wmnet with OS trixie * 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1171.eqiad.wmnet * 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1171.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:29 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95913 and previous config saved to /var/cache/conftool/dbconfig/20260805-082908-ladsgroup.json * 08:27 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1171.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:22 jynus@cumin1003: START - Cookbook sre.dns.netbox * 08:18 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249', diff saved to https://phabricator.wikimedia.org/P95912 and previous config saved to /var/cache/conftool/dbconfig/20260805-081823-ladsgroup.json * 08:17 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1171.eqiad.wmnet * 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1150.eqiad.wmnet * 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1150.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:15 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1150.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:15 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 08:14 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1155.eqiad.wmnet * 08:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1155.eqiad.wmnet * 08:14 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1155.eqiad.wmnet * 08:11 jynus@cumin1003: START - Cookbook sre.dns.netbox * 08:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249', diff saved to https://phabricator.wikimedia.org/P95911 and previous config saved to /var/cache/conftool/dbconfig/20260805-080737-ladsgroup.json * 08:05 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1150.eqiad.wmnet * 08:02 marostegui: Depool clouddb1020 (s5,s8) [[phab:T434048|T434048]] * 08:02 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1020.eqiad.wmnet,service=s8 * 08:02 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1020.eqiad.wmnet,service=s5 * 08:02 marostegui: Depool clouddb1018 (s2,s7) [[phab:T434048|T434048]] * 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1018.eqiad.wmnet,service=s7 * 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1018.eqiad.wmnet,service=s2 * 08:01 marostegui: Depool clouddb1017 (s1) [[phab:T434048|T434048]] * 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1017.eqiad.wmnet,service=s1 * 07:59 marostegui: Depool clouddb1016 (s5,s8) [[phab:T434048|T434048]] * 07:59 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s8 * 07:59 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s5 * 07:57 marostegui: Depool clouddb1015 (s4,s6) [[phab:T434048|T434048]] * 07:57 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s6 * 07:57 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s4 * 07:56 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95910 and previous config saved to /var/cache/conftool/dbconfig/20260805-075650-ladsgroup.json * 07:54 marostegui: Depool clouddb1014 (s2,s7) [[phab:T434048|T434048]] * 07:54 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1014.eqiad.wmnet,service=s7 * 07:54 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1014.eqiad.wmnet,service=s2 * 07:53 marostegui: Depool clouddb1013:s1 [[phab:T434048|T434048]] * 07:53 marostegui: Depool clouddb1013:s1 [[phab:T409557|T409557]] * 07:53 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1013.eqiad.wmnet,service=s1 * 07:25 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95909 and previous config saved to /var/cache/conftool/dbconfig/20260805-072529-ladsgroup.json * 07:24 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2249.codfw.wmnet with reason: Maintenance * 07:24 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95908 and previous config saved to /var/cache/conftool/dbconfig/20260805-072426-ladsgroup.json * 07:21 slyngshede@dns1004: END - running authdns-update * 07:19 slyngshede@dns1004: START - running authdns-update * 07:13 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231', diff saved to https://phabricator.wikimedia.org/P95906 and previous config saved to /var/cache/conftool/dbconfig/20260805-071340-ladsgroup.json * 07:02 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231', diff saved to https://phabricator.wikimedia.org/P95905 and previous config saved to /var/cache/conftool/dbconfig/20260805-070253-ladsgroup.json * 06:52 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95904 and previous config saved to /var/cache/conftool/dbconfig/20260805-065206-ladsgroup.json * 06:45 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 06:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95903 and previous config saved to /var/cache/conftool/dbconfig/20260805-062240-ladsgroup.json * 06:21 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2231.codfw.wmnet with reason: Maintenance * 06:21 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95902 and previous config saved to /var/cache/conftool/dbconfig/20260805-062137-ladsgroup.json * 06:10 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215', diff saved to https://phabricator.wikimedia.org/P95901 and previous config saved to /var/cache/conftool/dbconfig/20260805-061051-ladsgroup.json * 06:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215', diff saved to https://phabricator.wikimedia.org/P95900 and previous config saved to /var/cache/conftool/dbconfig/20260805-060004-ladsgroup.json * 05:49 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95899 and previous config saved to /var/cache/conftool/dbconfig/20260805-054918-ladsgroup.json * 05:19 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95898 and previous config saved to /var/cache/conftool/dbconfig/20260805-051939-ladsgroup.json * 05:18 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2215.codfw.wmnet with reason: Maintenance * 04:30 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2201.codfw.wmnet with reason: Maintenance * 03:40 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2197.codfw.wmnet with reason: Maintenance * 03:40 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95897 and previous config saved to /var/cache/conftool/dbconfig/20260805-034036-ladsgroup.json * 03:29 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196', diff saved to https://phabricator.wikimedia.org/P95896 and previous config saved to /var/cache/conftool/dbconfig/20260805-032948-ladsgroup.json * 03:19 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196', diff saved to https://phabricator.wikimedia.org/P95895 and previous config saved to /var/cache/conftool/dbconfig/20260805-031902-ladsgroup.json * 03:08 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95894 and previous config saved to /var/cache/conftool/dbconfig/20260805-030815-ladsgroup.json * 02:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95893 and previous config saved to /var/cache/conftool/dbconfig/20260805-023413-ladsgroup.json * 02:33 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2196.codfw.wmnet with reason: Maintenance * 02:33 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95892 and previous config saved to /var/cache/conftool/dbconfig/20260805-023310-ladsgroup.json * 02:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186', diff saved to https://phabricator.wikimedia.org/P95891 and previous config saved to /var/cache/conftool/dbconfig/20260805-022223-ladsgroup.json * 02:11 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186', diff saved to https://phabricator.wikimedia.org/P95890 and previous config saved to /var/cache/conftool/dbconfig/20260805-021137-ladsgroup.json * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 02:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95889 and previous config saved to /var/cache/conftool/dbconfig/20260805-020051-ladsgroup.json * 01:30 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95888 and previous config saved to /var/cache/conftool/dbconfig/20260805-013029-ladsgroup.json * 01:29 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2186.codfw.wmnet with reason: Maintenance * 00:34 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on dbstore1009.eqiad.wmnet with reason: Maintenance * 00:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95887 and previous config saved to /var/cache/conftool/dbconfig/20260805-003408-ladsgroup.json * 00:23 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264', diff saved to https://phabricator.wikimedia.org/P95886 and previous config saved to /var/cache/conftool/dbconfig/20260805-002322-ladsgroup.json * 00:12 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264', diff saved to https://phabricator.wikimedia.org/P95885 and previous config saved to /var/cache/conftool/dbconfig/20260805-001235-ladsgroup.json * 00:01 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95884 and previous config saved to /var/cache/conftool/dbconfig/20260805-000148-ladsgroup.json == 2026-08-04 == * 23:45 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95883 and previous config saved to /var/cache/conftool/dbconfig/20260804-234508-ladsgroup.json * 23:44 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1264.eqiad.wmnet with reason: Maintenance * 23:44 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95882 and previous config saved to /var/cache/conftool/dbconfig/20260804-234405-ladsgroup.json * 23:33 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237', diff saved to https://phabricator.wikimedia.org/P95881 and previous config saved to /var/cache/conftool/dbconfig/20260804-233317-ladsgroup.json * 23:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237', diff saved to https://phabricator.wikimedia.org/P95880 and previous config saved to /var/cache/conftool/dbconfig/20260804-232230-ladsgroup.json * 23:11 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95879 and previous config saved to /var/cache/conftool/dbconfig/20260804-231144-ladsgroup.json * 22:23 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95878 and previous config saved to /var/cache/conftool/dbconfig/20260804-222345-ladsgroup.json * 22:23 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1237.eqiad.wmnet with reason: Maintenance * 21:13 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1225.eqiad.wmnet with reason: Maintenance * 20:40 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] (duration: 24m 40s) * 20:33 samtar@deploy1003: samtar, kineticpelagic: Continuing with deployment * 20:28 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS bookworm * 20:21 samtar@deploy1003: samtar, kineticpelagic: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:15 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] * 20:13 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 20:09 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 20:00 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1216.eqiad.wmnet with reason: Maintenance * 20:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95877 and previous config saved to /var/cache/conftool/dbconfig/20260804-195957-ladsgroup.json * 19:51 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 19:50 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) mc2046.codfw.wmnet 120.16.192.10.in-addr.arpa 0.2.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:50 jhancock@cumin2002: START - Cookbook sre.dns.wipe-cache mc2046.codfw.wmnet 120.16.192.10.in-addr.arpa 0.2.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host mc2046 - jhancock@cumin2002" * 19:50 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host mc2046 - jhancock@cumin2002" * 19:49 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203', diff saved to https://phabricator.wikimedia.org/P95876 and previous config saved to /var/cache/conftool/dbconfig/20260804-194911-ladsgroup.json * 19:46 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 19:45 jhancock@cumin2002: START - Cookbook sre.hosts.move-vlan for host mc2046 * 19:45 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS bookworm * 19:38 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203', diff saved to https://phabricator.wikimedia.org/P95875 and previous config saved to /var/cache/conftool/dbconfig/20260804-193825-ladsgroup.json * 19:27 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95874 and previous config saved to /var/cache/conftool/dbconfig/20260804-192738-ladsgroup.json * 19:02 mutante: gerrit ssh -p 29418 gerrit.wikimedia.org gerrit index changes {{Gerrit|1320979}} * 18:20 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 18:18 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 18:14 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 18:14 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 18:13 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 18:10 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 18:08 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 18:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95872 and previous config saved to /var/cache/conftool/dbconfig/20260804-180721-ladsgroup.json * 18:07 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 18:06 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1203.eqiad.wmnet with reason: Maintenance * 18:06 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95871 and previous config saved to /var/cache/conftool/dbconfig/20260804-180618-ladsgroup.json * 17:55 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179', diff saved to https://phabricator.wikimedia.org/P95870 and previous config saved to /var/cache/conftool/dbconfig/20260804-175531-ladsgroup.json * 17:55 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1154.eqiad.wmnet * 17:55 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1154.eqiad.wmnet * 17:55 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1154.eqiad.wmnet * 17:50 swfrench@deploy1003: Finished scap sync-world: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] (duration: 04m 05s) * 17:48 swfrench@deploy1003: swfrench: Continuing with deployment * 17:46 swfrench@deploy1003: swfrench: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:45 swfrench@deploy1003: Started scap sync-world: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] * 17:44 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179', diff saved to https://phabricator.wikimedia.org/P95869 and previous config saved to /var/cache/conftool/dbconfig/20260804-174445-ladsgroup.json * 17:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95868 and previous config saved to /var/cache/conftool/dbconfig/20260804-173359-ladsgroup.json * 17:33 swfrench@deploy1003: Finished scap sync-world: Pick up new PHP production image (duration: 28m 32s) * 17:28 aokoth@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on phab1005.eqiad.wmnet with reason: Puppet Failure * 17:05 swfrench@deploy1003: Started scap sync-world: Pick up new PHP production image * 17:00 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 17:00 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 16:54 cgoubert@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on wikikube-worker2187.codfw.wmnet with reason: Hardware issue * 16:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2187.codfw.wmnet * 16:52 mutante: gerrit2003:/var/log/apache2# ln -s /srv/gerrit/site_path/review_site/logs/ gerrit * 16:52 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2187.codfw.wmnet * 16:48 mutante: gerrit2003 - moving old apache logfiles older than 60 days from /var/log/apache2 to /srv/gerrit/site_path/review_site/logs/old/ * 16:33 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 16:32 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 16:29 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 16:29 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 16:28 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 16:28 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 16:27 dzahn@cumin1003: END (PASS) - Cookbook sre.gerrit.restart-gerrit (exit_code=0) Restarting Gerrit on gerrit2003 * 16:27 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 16:27 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95867 and previous config saved to /var/cache/conftool/dbconfig/20260804-162736-ladsgroup.json * 16:27 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 16:26 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1179.eqiad.wmnet with reason: Maintenance * 16:26 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:25 mutante: restarting gerrit - dropped outdated RSA host key * 16:25 dzahn@cumin1003: START - Cookbook sre.gerrit.restart-gerrit Restarting Gerrit on gerrit2003 * 16:24 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95866 and previous config saved to /var/cache/conftool/dbconfig/20260804-162424-ladsgroup.json * 16:24 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 16:23 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 16:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95865 and previous config saved to /var/cache/conftool/dbconfig/20260804-162236-ladsgroup.json * 16:21 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1179.eqiad.wmnet with reason: Maintenance * 16:17 swfrench-wmf: reprepro include php8.3_8.3.33-1+wmf11u1 into component/php83 for bullseye-wikimedia * 16:17 swfrench-wmf: reprepro include php8.3_8.3.33-1+wmf12u1 into component/php83 for bookworm-wikimedia * 16:11 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply * 16:10 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply * 16:10 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mobileapps: apply * 16:09 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mobileapps: apply * 16:09 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply * 16:08 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply * 16:08 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:08 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:07 aokoth@cumin1003: END (PASS) - Cookbook sre.vrts.upgrade (exit_code=0) on VRTS host vrts1003.eqiad.wmnet * 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:05 aokoth@cumin1003: START - Cookbook sre.vrts.upgrade on VRTS host vrts1003.eqiad.wmnet * 16:04 mutante: gerrit2002/gerrit1003/gerrit2003 - rm /etc/gerrit/ssh_host_rsa_key * 15:59 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:59 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:59 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:59 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:56 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 15:55 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:55 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:55 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:49 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 15:49 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:44 Raine: add php8.5 packages to component/php85 - [[phab:T432983|T432983]] * 15:39 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:33 aaron@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 15:33 aaron@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 15:29 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:19 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:19 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:16 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:16 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1154.eqiad.wmnet with OS trixie * 15:16 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:15 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:15 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:06 brennen@deploy1003: Finished deploy [phabricator/deployment@56f4ffd]: deploy phab1004 for [[phab:T433981|T433981]] (duration: 00m 43s) * 15:05 brennen@deploy1003: Started deploy [phabricator/deployment@56f4ffd]: deploy phab1004 for [[phab:T433981|T433981]] * 15:05 aaron@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 15:04 aaron@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 15:02 brennen@deploy1003: Finished deploy [phabricator/deployment@56f4ffd]: deploy phab2003 for [[phab:T433981|T433981]] (duration: 00m 51s) * 15:01 brennen@deploy1003: Started deploy [phabricator/deployment@56f4ffd]: deploy phab2003 for [[phab:T433981|T433981]] * 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1004.eqiad.wmnet with reason: deployment * 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1005.eqiad.wmnet with reason: deployment * 14:58 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab2003.codfw.wmnet with reason: deployment * 14:55 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1154.eqiad.wmnet with reason: host reimage * 14:51 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1154.eqiad.wmnet with reason: host reimage * 14:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 14:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 14:38 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync * 14:38 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync * 14:38 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync * 14:37 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync * 14:37 ottomata: roll restart eventgate-main to pick up stream config change - [[phab:T433507|T433507]] * 14:37 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-main: sync * 14:36 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-main: sync * 14:36 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1154 * 14:36 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1154 * 14:34 otto@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] (duration: 08m 39s) * 14:34 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1154 * 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1154.eqiad.wmnet 108.32.64.10.in-addr.arpa 8.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:34 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1154.eqiad.wmnet 108.32.64.10.in-addr.arpa 8.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1154 - jayme@cumin1003" * 14:34 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1154 - jayme@cumin1003" * 14:30 otto@deploy1003: otto: Continuing with deployment * 14:30 jayme@cumin1003: START - Cookbook sre.dns.netbox * 14:28 otto@deploy1003: otto: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:26 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1154 * 14:26 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1154.eqiad.wmnet with OS trixie * 14:26 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1154.eqiad.wmnet * 14:26 otto@deploy1003: Started scap sync-world: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] * 14:26 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1154.eqiad.wmnet * 14:26 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1154.eqiad.wmnet * 14:17 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 14:16 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 14:15 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 14:14 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 14:13 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 14:13 swfrench@dns1004: END - running authdns-update * 14:13 Msz2001: Finished deployments for UTC afternoon backport window * 14:13 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 14:13 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] (duration: 07m 58s) * 14:11 swfrench@dns1004: START - running authdns-update * 14:08 mszwarc@deploy1003: javiermonton, mszwarc, mpostoronca: Continuing with deployment * 14:07 mszwarc@deploy1003: javiermonton, mszwarc, mpostoronca: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] synced to the testser * 14:05 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] * 14:03 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 13:49 swfrench@cumin2002: conftool action : set/pooled=yes; selector: name=wikikube-worker2330.codfw.wmnet * 13:49 swfrench@cumin2002: conftool action : set/pooled=no; selector: name=wikikube-worker2330.codfw.wmnet * 13:48 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] (duration: 09m 19s) * 13:45 swfrench@dns1004: END - running authdns-update * 13:44 mszwarc@deploy1003: mszwarc, jforrester: Continuing with deployment * 13:43 swfrench@dns1004: START - running authdns-update * 13:41 mszwarc@deploy1003: mszwarc, jforrester: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:38 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] * 13:33 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 13:33 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1154.eqiad.wmnet * 13:32 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 13:32 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 13:31 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 13:31 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:31 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:29 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1154.eqiad.wmnet * 13:28 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1154.eqiad.wmnet * 13:28 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1154.eqiad.wmnet * 13:28 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1141.eqiad.wmnet * 13:28 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1141.eqiad.wmnet * 13:28 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1141.eqiad.wmnet * 13:22 otto@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply * 13:22 otto@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply * 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1096.eqiad.wmnet with OS trixie * 13:05 swfrench@dns1004: END - running authdns-update * 13:03 swfrench@dns1004: START - running authdns-update * 12:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 12:43 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 1:00:00 on db1171.eqiad.wmnet with reason: decom * 12:42 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 1:00:00 on db1150.eqiad.wmnet with reason: decom * 12:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 12:38 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1164,1217].eqiad.wmnet with reason: cloning * 12:33 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2096.codfw.wmnet with OS trixie * 12:22 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1096.eqiad.wmnet with OS trixie * 12:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2096.codfw.wmnet with reason: host reimage * 12:14 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1141.eqiad.wmnet with OS trixie * 12:10 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2096.codfw.wmnet with reason: host reimage * 12:10 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1289.eqiad.wmnet * 12:05 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1289.eqiad.wmnet * 12:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1288.eqiad.wmnet * 11:59 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1288.eqiad.wmnet * 11:59 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1287.eqiad.wmnet * 11:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1097.eqiad.wmnet with OS trixie * 11:54 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1287.eqiad.wmnet * 11:54 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1286.eqiad.wmnet * 11:53 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1141.eqiad.wmnet with reason: host reimage * 11:51 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2096.codfw.wmnet with OS trixie * 11:49 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1141.eqiad.wmnet with reason: host reimage * 11:48 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1286.eqiad.wmnet * 11:48 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1284.eqiad.wmnet * 11:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2095.codfw.wmnet with OS trixie * 11:43 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1284.eqiad.wmnet * 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1283.eqiad.wmnet * 11:42 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on ml-serve1015.eqiad.wmnet with reason: Downtime to get full picture of current BIOS settings beyond what Redfish shows * 11:39 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad * 11:39 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:37 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1283.eqiad.wmnet * 11:37 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1282.eqiad.wmnet * 11:37 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad * 11:37 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:33 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1141 * 11:33 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1141 * 11:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 11:32 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1141 * 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1141.eqiad.wmnet 156.48.64.10.in-addr.arpa 6.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:32 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1141.eqiad.wmnet 156.48.64.10.in-addr.arpa 6.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1141 - jayme@cumin1003" * 11:32 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1141 - jayme@cumin1003" * 11:32 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1282.eqiad.wmnet * 11:32 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1281.eqiad.wmnet * 11:32 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad * 11:32 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:29 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 11:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1097.eqiad.wmnet with reason: host reimage * 11:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2095.codfw.wmnet with OS trixie * 11:26 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1281.eqiad.wmnet * 11:26 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1280.eqiad.wmnet * 11:25 jayme@cumin1003: START - Cookbook sre.dns.netbox * 11:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1097.eqiad.wmnet with reason: host reimage * 11:22 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1141 * 11:21 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1141.eqiad.wmnet with OS trixie * 11:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1280.eqiad.wmnet * 11:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1279.eqiad.wmnet * 11:20 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin with reason: upgrade new Nokia swtiches in eqsin to SR Linux v26 * 11:17 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1141.eqiad.wmnet * 11:16 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1141.eqiad.wmnet * 11:16 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1141.eqiad.wmnet * 11:16 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be2095.codfw.wmnet with OS trixie * 11:15 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1279.eqiad.wmnet * 11:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1278.eqiad.wmnet * 11:14 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1139.eqiad.wmnet * 11:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1139.eqiad.wmnet * 11:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1070.eqiad.wmnet with OS trixie * 11:13 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1140.eqiad.wmnet * 11:13 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1140.eqiad.wmnet * 11:13 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1140.eqiad.wmnet * 11:09 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1278.eqiad.wmnet * 11:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1071.eqiad.wmnet with OS trixie * 11:05 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1097.eqiad.wmnet with OS trixie * 11:04 mvernon@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be1097.eqiad.wmnet with OS trixie * 11:02 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1140.eqiad.wmnet with OS trixie * 11:02 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1097.eqiad.wmnet with OS trixie * 11:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1096.eqiad.wmnet with OS trixie * 11:00 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1139.eqiad.wmnet * 11:00 jayme@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1139.eqiad.wmnet with OS trixie * 10:56 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1069.eqiad.wmnet with OS trixie * 10:56 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 10:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1070.eqiad.wmnet with reason: host reimage * 10:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 10:45 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1071.eqiad.wmnet with reason: host reimage * 10:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 10:41 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1096.eqiad.wmnet with OS trixie * 10:41 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1140.eqiad.wmnet with reason: host reimage * 10:39 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1071.eqiad.wmnet with reason: host reimage * 10:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1070.eqiad.wmnet with reason: host reimage * 10:38 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1139.eqiad.wmnet with reason: host reimage * 10:37 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1140.eqiad.wmnet with reason: host reimage * 10:35 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1069.eqiad.wmnet with reason: host reimage * 10:33 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1095.eqiad.wmnet with OS trixie * 10:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 10:33 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1139.eqiad.wmnet with reason: host reimage * 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1069.eqiad.wmnet with reason: host reimage * 10:23 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1140 * 10:23 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1140 * 10:23 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 10:22 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1140 * 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1140.eqiad.wmnet 155.48.64.10.in-addr.arpa 5.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:21 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1071.eqiad.wmnet with OS trixie * 10:21 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1140.eqiad.wmnet 155.48.64.10.in-addr.arpa 5.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1140 - jayme@cumin1003" * 10:21 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1140 - jayme@cumin1003" * 10:21 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1071 * 10:21 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1070.eqiad.wmnet with OS trixie * 10:21 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1070 * 10:20 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 10:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1095.eqiad.wmnet with OS trixie * 10:17 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1139 * 10:17 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1139 * 10:17 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be1095.eqiad.wmnet with OS trixie * 10:15 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1139 * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1139.eqiad.wmnet 194.32.64.10.in-addr.arpa 4.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:15 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1139.eqiad.wmnet 194.32.64.10.in-addr.arpa 4.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1139 - jayme@cumin1003" * 10:15 jayme@cumin1003: START - Cookbook sre.dns.netbox * 10:15 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1139 - jayme@cumin1003" * 10:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2095.codfw.wmnet with OS trixie * 10:13 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1069.eqiad.wmnet with OS trixie * 10:12 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1069 * 10:11 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1140 * 10:11 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1140.eqiad.wmnet with OS trixie * 10:11 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1140.eqiad.wmnet * 10:10 jayme@cumin1003: START - Cookbook sre.dns.netbox * 10:10 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1139 * 10:10 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1140.eqiad.wmnet * 10:10 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1140.eqiad.wmnet * 10:10 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1139.eqiad.wmnet with OS trixie * 10:09 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1139.eqiad.wmnet * 10:08 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1139.eqiad.wmnet * 10:08 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1139.eqiad.wmnet * 10:01 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2094.codfw.wmnet with OS trixie * 09:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 09:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 09:53 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 09:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 09:44 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:44 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2094.codfw.wmnet with reason: host reimage * 09:34 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2094.codfw.wmnet with reason: host reimage * 09:34 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1071 * 09:33 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1070 * 09:33 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1095.eqiad.wmnet with OS trixie * 09:27 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1069 * 09:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:23 brouberol@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM archiva1002.wikimedia.org * 09:20 brouberol@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM archiva1002.wikimedia.org * 09:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1277.eqiad.wmnet * 09:13 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2094.codfw.wmnet with OS trixie * 09:13 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 09:12 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 09:12 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:12 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1277.eqiad.wmnet * 09:12 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1276.eqiad.wmnet * 09:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1094.eqiad.wmnet with OS trixie * 09:06 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1276.eqiad.wmnet * 09:06 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1275.eqiad.wmnet * 09:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2093.codfw.wmnet with OS trixie * 09:01 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1275.eqiad.wmnet * 09:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1274.eqiad.wmnet * 08:56 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1274.eqiad.wmnet * 08:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1273.eqiad.wmnet * 08:50 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1273.eqiad.wmnet * 08:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1272.eqiad.wmnet * 08:49 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:49 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1094.eqiad.wmnet with reason: host reimage * 08:45 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1272.eqiad.wmnet * 08:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1094.eqiad.wmnet with reason: host reimage * 08:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2093.codfw.wmnet with reason: host reimage * 08:38 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:38 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2093.codfw.wmnet with reason: host reimage * 08:35 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 08:34 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 08:29 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:28 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:26 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1271.eqiad.wmnet * 08:23 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1094.eqiad.wmnet with OS trixie * 08:21 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 08:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1271.eqiad.wmnet * 08:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1270.eqiad.wmnet * 08:15 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1270.eqiad.wmnet * 08:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1269.eqiad.wmnet * 08:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2093.codfw.wmnet with OS trixie * 08:09 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1269.eqiad.wmnet * 08:09 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1268.eqiad.wmnet * 08:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2092.codfw.wmnet with OS trixie * 08:04 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1268.eqiad.wmnet * 08:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1267.eqiad.wmnet * 07:59 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1267.eqiad.wmnet * 07:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1093.eqiad.wmnet with OS trixie * 07:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1266.eqiad.wmnet * 07:56 jynus: running extra backups to test db1285 [[phab:T433826|T433826]] * 07:51 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1266.eqiad.wmnet * 07:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2092.codfw.wmnet with reason: host reimage * 07:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1093.eqiad.wmnet with reason: host reimage * 07:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2092.codfw.wmnet with reason: host reimage * 07:32 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1093.eqiad.wmnet with reason: host reimage * 07:29 jynus: running extra backups to test db1265 [[phab:T433825|T433825]] * 07:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2092.codfw.wmnet with OS trixie * 07:11 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1093.eqiad.wmnet with OS trixie * 06:50 slyngshede@dns1004: END - running authdns-update * 06:48 slyngshede@dns1004: START - running authdns-update * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.11 (duration: 02m 29s) * 03:38 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] (duration: 32m 57s) * 03:23 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 03:22 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 03:05 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 32s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 00:45 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] (duration: 06m 20s) * 00:41 cjming@deploy1003: cjming: Continuing with deployment * 00:41 cjming@deploy1003: cjming: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:39 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] == 2026-08-03 == * 23:58 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply * 23:57 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply * 23:29 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cp5021.eqsin.wmnet * 23:29 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cp5021.eqsin.wmnet * 23:27 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cp5021.eqsin.wmnet * 23:26 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cp5021.eqsin.wmnet * 23:18 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 23:17 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 22:56 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: sync * 22:56 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: sync * 22:36 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 22:36 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 22:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host search-loader2002.codfw.wmnet with OS trixie * 21:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on search-loader2002.codfw.wmnet with reason: host reimage * 21:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on search-loader2002.codfw.wmnet with reason: host reimage * 21:42 dancy@deploy1003: Stopping before sync operations * 21:41 dancy@deploy1003: Started scap sync-world: testing * 21:39 dancy@deploy1003: Installation of scap version "4.277.0" completed for 3 hosts * 21:37 dancy@deploy1003: Installing scap version "4.277.0" for 3 host(s) * 21:37 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] (duration: 06m 13s) * 21:33 dancy@deploy1003: dancy: Continuing with deployment * 21:32 dancy@deploy1003: dancy: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:31 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] * 21:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host search-loader2002.codfw.wmnet with OS trixie * 21:03 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] (duration: 06m 34s) * 20:59 dancy@deploy1003: dancy: Continuing with deployment * 20:58 dancy@deploy1003: dancy: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:56 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] * 20:52 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] (duration: 06m 23s) * 20:48 cjming@deploy1003: cjming: Continuing with deployment * 20:47 cjming@deploy1003: cjming: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:46 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] * 20:42 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] (duration: 07m 36s) * 20:38 arlolra@deploy1003: arlolra: Continuing with deployment * 20:36 arlolra@deploy1003: arlolra: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:34 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] * 20:16 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] (duration: 08m 26s) * 20:12 krinkle@deploy1003: krinkle: Continuing with deployment * 20:09 krinkle@deploy1003: krinkle: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] * 19:45 jasmine@cumin2002: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-main-codfw * 18:58 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] (duration: 09m 23s) * 18:53 krinkle@deploy1003: krinkle: Continuing with deployment * 18:53 jasmine@cumin2002: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-main-codfw * 18:50 krinkle@deploy1003: krinkle: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:48 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] * 18:37 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] (duration: 10m 13s) * 18:34 dzahn@cumin2002: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host codesearch2001.codfw.wmnet * 18:34 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host codesearch2001.codfw.wmnet with OS trixie * 18:33 krinkle@deploy1003: krinkle: Continuing with deployment * 18:29 krinkle@deploy1003: krinkle: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:27 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] * 18:19 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 18:18 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on codesearch2001.codfw.wmnet with reason: host reimage * 18:14 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 18:14 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:12 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on codesearch2001.codfw.wmnet with reason: host reimage * 18:11 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 18:11 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 18:10 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 18:02 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 18:02 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 18:01 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 18:01 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 17:55 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host codesearch2001.codfw.wmnet with OS trixie * 17:54 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:54 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) codesearch2001.codfw.wmnet on all recursors * 17:53 dzahn@cumin2002: START - Cookbook sre.dns.wipe-cache codesearch2001.codfw.wmnet on all recursors * 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:48 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:41 dzahn@cumin2002: START - Cookbook sre.dns.netbox * 17:41 dzahn@cumin2002: START - Cookbook sre.ganeti.makevm for new host codesearch2001.codfw.wmnet * 17:37 dzahn@cumin2002: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host codesearch1001.eqiad.wmnet * 17:37 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host codesearch1001.eqiad.wmnet with OS trixie * 17:24 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on codesearch1001.eqiad.wmnet with reason: host reimage * 17:17 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on codesearch1001.eqiad.wmnet with reason: host reimage * 17:08 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host codesearch1001.eqiad.wmnet with OS trixie * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 17:06 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:06 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) codesearch1001.eqiad.wmnet on all recursors * 17:06 dzahn@cumin2002: START - Cookbook sre.dns.wipe-cache codesearch1001.eqiad.wmnet on all recursors * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 17:05 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 17:04 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:04 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 16:58 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 16:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2091.codfw.wmnet with OS trixie * 16:54 ebernhardson@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 16:54 ebernhardson@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 16:49 ebernhardson@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 16:49 ebernhardson@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 16:46 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1092.eqiad.wmnet with OS trixie * 16:43 ebernhardson@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 16:43 ebernhardson@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 16:43 dzahn@cumin2002: START - Cookbook sre.dns.netbox * 16:43 dzahn@cumin2002: START - Cookbook sre.ganeti.makevm for new host codesearch1001.eqiad.wmnet * 16:41 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 16:41 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 16:40 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2091.codfw.wmnet with reason: host reimage * 16:37 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 16:35 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2091.codfw.wmnet with reason: host reimage * 16:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1092.eqiad.wmnet with reason: host reimage * 16:24 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 16:24 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 16:23 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1092.eqiad.wmnet with reason: host reimage * 16:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2091.codfw.wmnet with OS trixie * 16:03 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1092.eqiad.wmnet with OS trixie * 16:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2090.codfw.wmnet with OS trixie * 15:51 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 15:51 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 15:51 jiji@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 15:50 jiji@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 15:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2090.codfw.wmnet with reason: host reimage * 15:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2090.codfw.wmnet with reason: host reimage * 15:33 jhathaway@dns1004: END - running authdns-update * 15:31 jhathaway@dns1004: START - running authdns-update * 15:26 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1091.eqiad.wmnet with OS trixie * 15:25 dancy@deploy1003: Installation of scap version "4.276.1" completed for 3 hosts * 15:23 dancy@deploy1003: Installing scap version "4.276.1" for 3 host(s) * 15:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2090.codfw.wmnet with OS trixie * 15:12 marostegui@cumin1003: dbctl commit (dc=all): 'Repool db2245, db2246, db2247 and db2248 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95857 and previous config saved to /var/cache/conftool/dbconfig/20260803-151212-marostegui.json * 15:09 dancy@deploy1003: Started scap sync-world: testing * 15:09 dancy@deploy1003: Installation of scap version "4.277.0" completed for 3 hosts * 15:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1091.eqiad.wmnet with reason: host reimage * 15:07 dancy@deploy1003: Installing scap version "4.277.0" for 3 host(s) * 15:03 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1091.eqiad.wmnet with reason: host reimage * 14:49 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1091.eqiad.wmnet with OS trixie * 14:33 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2089.codfw.wmnet with OS trixie * 14:29 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 14:27 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 14:18 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 14:16 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 14:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2089.codfw.wmnet with reason: host reimage * 14:10 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 14:10 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 14:09 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2089.codfw.wmnet with reason: host reimage * 13:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2089.codfw.wmnet with OS trixie * 13:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2088.codfw.wmnet with OS trixie * 13:40 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1090.eqiad.wmnet with OS trixie * 13:22 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1090.eqiad.wmnet with reason: host reimage * 13:22 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] (duration: 14m 34s) * 13:19 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1090.eqiad.wmnet with reason: host reimage * 13:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2088.codfw.wmnet with reason: host reimage * 13:16 aude@deploy1003: aude, mhorsey: Continuing with deployment * 13:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2088.codfw.wmnet with reason: host reimage * 13:12 aude@deploy1003: aude, mhorsey: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] * 13:05 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1090.eqiad.wmnet with OS trixie * 12:58 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2088.codfw.wmnet with OS trixie * 12:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db[2245-2247].codfw.wmnet * 12:49 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2247: Rebooting db2247.codfw.wmnet * 12:49 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2247: Rebooting db2247.codfw.wmnet * 12:42 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2246: Rebooting db2246.codfw.wmnet * 12:42 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2246: Rebooting db2246.codfw.wmnet * 12:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2087.codfw.wmnet with OS trixie * 12:37 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1089.eqiad.wmnet with OS trixie * 12:34 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2245: Rebooting db2245.codfw.wmnet * 12:34 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2245: Rebooting db2245.codfw.wmnet * 12:34 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db[2245-2247].codfw.wmnet * 12:32 kamila@deploy1003: Finished scap sync-world: rebuild after base image update (duration: 30m 26s) * 12:28 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 12:22 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2087.codfw.wmnet with reason: host reimage * 12:19 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1089.eqiad.wmnet with reason: host reimage * 12:14 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2087.codfw.wmnet with reason: host reimage * 12:14 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1089.eqiad.wmnet with reason: host reimage * 12:03 kamila@deploy1003: Started scap sync-world: rebuild after base image update * 12:00 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1089.eqiad.wmnet with OS trixie * 12:00 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2087.codfw.wmnet with OS trixie * 11:35 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db[2245-2248].codfw.wmnet * 11:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db[2245-2248].codfw.wmnet * 11:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2086.codfw.wmnet with OS trixie * 11:26 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 11:26 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 11:25 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 11:25 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 11:24 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1088.eqiad.wmnet with OS trixie * 11:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db[2245-2248].codfw.wmnet with reason: Checking network * 11:21 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 11:20 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 11:19 marostegui@dns1004: END - running authdns-update * 11:17 marostegui@dns1004: START - running authdns-update * 11:10 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:10 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 11:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2086.codfw.wmnet with reason: host reimage * 11:09 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:08 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 11:08 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:07 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 11:07 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:07 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 11:06 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop: apply * 11:06 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop: apply * 11:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1088.eqiad.wmnet with reason: host reimage * 11:05 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop: apply * 11:04 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop: apply * 11:04 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop: apply * 11:04 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop: apply * 11:02 marostegui@dns1004: END - running authdns-update * 11:02 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2086.codfw.wmnet with reason: host reimage * 11:01 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1088.eqiad.wmnet with reason: host reimage * 11:00 marostegui@dns1004: START - running authdns-update * 10:53 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] (duration: 10m 57s) * 10:51 cmooney@dns3003: END - running authdns-update * 10:49 cmooney@dns3003: START - running authdns-update * 10:47 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1088.eqiad.wmnet with OS trixie * 10:47 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2086.codfw.wmnet with OS trixie * 10:47 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 10:46 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:46 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:46 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new reverse ranges for eqsin CR switch links - cmooney@cumin1003" * 10:46 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new reverse ranges for eqsin CR switch links - cmooney@cumin1003" * 10:42 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] * 10:41 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 10:36 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2245, db2246 and db2247 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95855 and previous config saved to /var/cache/conftool/dbconfig/20260803-103652-marostegui.json * 10:35 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2248 from s4 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95854 and previous config saved to /var/cache/conftool/dbconfig/20260803-103535-marostegui.json * 10:27 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 10:27 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 10:26 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 10:24 kart_: cxserver: Add referencePunctuation config ([[phab:T97231|T97231]]) * 10:24 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 10:23 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:23 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:23 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:22 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:22 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply * 10:21 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply * 10:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2085.codfw.wmnet with OS trixie * 10:20 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply * 10:20 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply * 10:18 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply * 10:18 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply * 10:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1087.eqiad.wmnet with OS trixie * 09:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2085.codfw.wmnet with reason: host reimage * 09:43 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1087.eqiad.wmnet with reason: host reimage * 09:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2085.codfw.wmnet with reason: host reimage * 09:40 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1087.eqiad.wmnet with reason: host reimage * 09:26 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1087.eqiad.wmnet with OS trixie * 09:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2085.codfw.wmnet with OS trixie * 09:13 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2084.codfw.wmnet with OS trixie * 09:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1086.eqiad.wmnet with OS trixie * 08:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2084.codfw.wmnet with reason: host reimage * 08:50 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2084.codfw.wmnet with reason: host reimage * 08:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1086.eqiad.wmnet with reason: host reimage * 08:39 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1086.eqiad.wmnet with reason: host reimage * 08:38 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:38 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:37 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 08:37 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 08:35 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2084.codfw.wmnet with OS trixie * 08:34 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 08:34 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:27 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1086.eqiad.wmnet with OS trixie * 08:09 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1218: Repool after a crash * 08:07 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2083.codfw.wmnet with OS trixie * 08:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1085.eqiad.wmnet with OS trixie * 07:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2083.codfw.wmnet with reason: host reimage * 07:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1085.eqiad.wmnet with reason: host reimage * 07:40 kart_: Updated cxsever to 2026-07-16-140518-production ([[phab:T97231|T97231]]) * 07:39 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply * 07:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2083.codfw.wmnet with reason: host reimage * 07:38 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply * 07:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1085.eqiad.wmnet with reason: host reimage * 07:37 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] (duration: 32m 40s) * 07:33 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply * 07:33 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply * 07:25 jdlrobson@deploy1003: jdlrobson: Continuing with deployment * 07:24 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2083.codfw.wmnet with OS trixie * 07:24 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1085.eqiad.wmnet with OS trixie * 07:23 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1218: Repool after a crash * 07:21 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:09 marostegui: Drop renamed tables [[phab:T425074|T425074]] * 07:04 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] * 06:55 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply * 06:54 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 46s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-02 == * 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 01m 03s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-01 == * 03:30 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:30 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:30 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:30 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 34s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-31 == * 17:41 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 17:41 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 17:40 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 17:40 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 15:33 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2195: Testing * 15:02 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:02 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 15:02 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 14:48 pt1979@cumin2002: START - Cookbook sre.dns.netbox * 14:47 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2195: Testing * 14:22 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 14:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2195: Testing * 14:21 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 14:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2195.codfw.wmnet with reason: Testing * 14:16 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 14:04 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 14:04 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 13:30 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1048.eqiad.wmnet with OS trixie * 13:22 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2195: Testing * 13:22 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 13:19 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2195: Testing * 13:18 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 13:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2195: Testing * 13:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 13:05 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 13:05 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 13:04 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 13:04 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 12:50 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lswtest-d8-eqiad * 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:53 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:42 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 6515 * 11:37 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 6515 * 11:28 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:27 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:07 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2082.codfw.wmnet with OS trixie * 10:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2082.codfw.wmnet with reason: host reimage * 10:42 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2082.codfw.wmnet with reason: host reimage * 10:28 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2082.codfw.wmnet with OS trixie * 10:02 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 09:52 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 09:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts * 09:16 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts * 08:57 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 08:46 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:42 gkyziridis@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 08:37 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 08:37 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 08:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 08:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 08:11 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:11 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:08 filippo@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudvirt1048 * 08:07 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 08:07 filippo@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudvirt1048 * 08:06 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 08:01 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 08:00 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 07:19 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1048.eqiad.wmnet with reason: host reimage * 07:13 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1048.eqiad.wmnet with reason: host reimage * 07:11 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 07:11 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 07:09 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 07:09 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 06:57 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:56 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:48 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:48 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:44 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie * 06:34 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1048.eqiad.wmnet with OS trixie * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 54s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 00:57 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] (duration: 11m 04s) * 00:53 dreamyjazz@deploy1003: dreamyjazz, jforrester: Continuing with deployment * 00:48 dreamyjazz@deploy1003: dreamyjazz, jforrester: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:46 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] == 2026-07-30 == * 21:37 dancy@deploy1003: Installation of scap version "4.276.1" completed for 3 hosts * 21:35 dancy@deploy1003: Installing scap version "4.276.1" for 3 host(s) * 21:24 dancy@deploy1003: Installation of scap version "4.276.0" completed for 3 hosts * 21:22 dancy@deploy1003: Installing scap version "4.276.0" for 3 host(s) * 21:15 maryum: Deployed security fix for [[phab:T430601|T430601]] * 20:13 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] (duration: 09m 20s) * 20:07 arlolra@deploy1003: osleger, arlolra: Continuing with deployment * 20:05 arlolra@deploy1003: osleger, arlolra: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:03 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] * 19:29 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:29 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:25 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service * 19:24 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 19:24 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:24 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:24 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 19:23 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service * 19:20 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1084.eqiad.wmnet with OS trixie * 18:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1084.eqiad.wmnet with reason: host reimage * 18:52 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1084.eqiad.wmnet with reason: host reimage * 18:41 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 18:40 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 18:39 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1084.eqiad.wmnet with OS trixie * 18:25 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 18:15 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 17:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1083.eqiad.wmnet with OS trixie * 17:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1048: Maintenance * 17:36 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new security plugin settings - bking@cumin2003 - [[phab:T350516|T350516]] * 17:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1083.eqiad.wmnet with reason: host reimage * 17:26 inflatador: bking@apt1002 `reprepro --noskipold --component thirdparty/opensearch3 update trixie-wikimedia` [[phab:T433624|T433624]] * 17:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1083.eqiad.wmnet with reason: host reimage * 17:23 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 17:20 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 17:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2081.codfw.wmnet with OS trixie * 17:11 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new security plugin settings - bking@cumin2003 - [[phab:T350516|T350516]] * 17:10 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1083.eqiad.wmnet with OS trixie * 16:55 root@cumin1003: START - Cookbook sre.mysql.pool pool es1048: Maintenance * 16:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2081.codfw.wmnet with reason: host reimage * 16:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1048 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95833 and previous config saved to /var/cache/conftool/dbconfig/20260730-165053-cwilliams.json * 16:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1048.eqiad.wmnet with reason: Maintenance * 16:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1040: Maintenance * 16:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2081.codfw.wmnet with reason: host reimage * 16:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1082.eqiad.wmnet with OS trixie * 16:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2081.codfw.wmnet with OS trixie * 16:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2097.codfw.wmnet with OS trixie * 16:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1082.eqiad.wmnet with reason: host reimage * 16:08 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1082.eqiad.wmnet with reason: host reimage * 16:08 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 16:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1047: Maintenance * 16:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2080.codfw.wmnet with OS trixie * 16:04 root@cumin1003: START - Cookbook sre.mysql.pool pool es1040: Maintenance * 16:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1040: Maintenance * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new logging settings - bking@cumin2003 - [[phab:T324335|T324335]] * 15:58 root@cumin1003: START - Cookbook sre.mysql.pool pool es1040: Maintenance * 15:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1040 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95827 and previous config saved to /var/cache/conftool/dbconfig/20260730-155324-cwilliams.json * 15:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1040.eqiad.wmnet with reason: Maintenance * 15:50 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1082.eqiad.wmnet with OS trixie * 15:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2048: Maintenance * 15:44 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 24s) * 15:43 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2080.codfw.wmnet with reason: host reimage * 15:38 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new logging settings - bking@cumin2003 - [[phab:T324335|T324335]] * 15:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2080.codfw.wmnet with reason: host reimage * 15:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 15:30 mvernon@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be2097.codfw.wmnet with OS trixie * 15:23 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1081.eqiad.wmnet with OS trixie * 15:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2098.codfw.wmnet with OS trixie * 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - mvernon@cumin2003" * 15:18 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be2097.codfw.wmnet with OS trixie * 15:18 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - mvernon@cumin2003" * 15:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 15:17 root@cumin1003: START - Cookbook sre.mysql.pool pool es1047: Maintenance * 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2080.codfw.wmnet with OS trixie * 15:13 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be2097.codfw.wmnet with OS trixie * 15:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1047 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95820 and previous config saved to /var/cache/conftool/dbconfig/20260730-151200-cwilliams.json * 15:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1047.eqiad.wmnet with reason: Maintenance * 15:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: Maintenance * 15:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 15:04 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 15:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1081.eqiad.wmnet with reason: host reimage * 15:00 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1081.eqiad.wmnet with reason: host reimage * 15:00 root@cumin1003: START - Cookbook sre.mysql.pool pool es2048: Maintenance * 15:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 14:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2079.codfw.wmnet with OS trixie * 14:56 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 14:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2048 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95816 and previous config saved to /var/cache/conftool/dbconfig/20260730-145510-cwilliams.json * 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2048.codfw.wmnet with reason: Maintenance * 14:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2040: Maintenance * 14:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 14:51 tchin@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] (duration: 06m 48s) * 14:47 tchin@deploy1003: jforrester, tchin: Continuing with deployment * 14:47 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 14:47 tchin@deploy1003: jforrester, tchin: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:45 tchin@deploy1003: Started scap sync-world: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] * 14:42 sukhe@puppetserver1001: conftool action : set/weight=1; selector: cluster=urldownloader,service=squid * 14:42 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader,service=squid * 14:42 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1081.eqiad.wmnet with OS trixie * 14:39 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 14:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2079.codfw.wmnet with reason: host reimage * 14:36 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2098.codfw.wmnet with OS trixie * 14:32 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2079.codfw.wmnet with reason: host reimage * 14:30 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] (duration: 06m 31s) * 14:27 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 14:26 mszwarc@deploy1003: mszwarc: Continuing with deployment * 14:25 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:25 root@cumin1003: START - Cookbook sre.mysql.pool pool es1038: Maintenance * 14:25 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1038: Maintenance * 14:23 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] * 14:21 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] (duration: 11m 19s) * 14:20 root@cumin1003: START - Cookbook sre.mysql.pool pool es1038: Maintenance * 14:14 stran@deploy1003: stran: Continuing with deployment * 14:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1038 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95810 and previous config saved to /var/cache/conftool/dbconfig/20260730-141439-cwilliams.json * 14:14 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1038.eqiad.wmnet with reason: Maintenance * 14:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1036: Maintenance * 14:13 stran@deploy1003: stran: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2079.codfw.wmnet with OS trixie * 14:09 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] * 14:08 root@cumin1003: START - Cookbook sre.mysql.pool pool es2040: Maintenance * 14:08 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2040: Maintenance * 14:03 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] (duration: 31m 41s) * 14:03 root@cumin1003: START - Cookbook sre.mysql.pool pool es2040: Maintenance * 14:03 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2040 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95806 and previous config saved to /var/cache/conftool/dbconfig/20260730-135643-cwilliams.json * 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2040.codfw.wmnet with reason: Maintenance * 13:56 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:56 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2038: Maintenance * 13:55 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:52 lucaswerkmeister-wmde@deploy1003: migr, lucaswerkmeister-wmde: Continuing with deployment * 13:49 lucaswerkmeister-wmde@deploy1003: migr, lucaswerkmeister-wmde: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:49 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:48 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 13:45 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 13:32 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:32 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] * 13:28 root@cumin1003: START - Cookbook sre.mysql.pool pool es1036: Maintenance * 13:28 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1036: Maintenance * 13:22 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:22 root@cumin1003: START - Cookbook sre.mysql.pool pool es1036: Maintenance * 13:20 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1036 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95800 and previous config saved to /var/cache/conftool/dbconfig/20260730-131727-cwilliams.json * 13:17 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1036.eqiad.wmnet with reason: Maintenance * 13:17 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] (duration: 10m 31s) * 13:16 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2022\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 13:13 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, stran: Continuing with deployment * 13:10 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:10 root@cumin1003: START - Cookbook sre.mysql.pool pool es2038: Maintenance * 13:10 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2038: Maintenance * 13:08 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, stran: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie * 13:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2047: Maintenance * 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048 cloud-private - filippo@cumin1003" * 13:07 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048 cloud-private - filippo@cumin1003" * 13:06 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] * 13:04 root@cumin1003: START - Cookbook sre.mysql.pool pool es2038: Maintenance * 13:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2078.codfw.wmnet with OS trixie * 13:01 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2038 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95797 and previous config saved to /var/cache/conftool/dbconfig/20260730-125919-cwilliams.json * 12:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2038.codfw.wmnet with reason: Maintenance * 12:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2078.codfw.wmnet with reason: host reimage * 12:37 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2078.codfw.wmnet with reason: host reimage * 12:37 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] (duration: 06m 51s) * 12:33 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 12:32 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 12:32 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 12:32 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:30 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] * 12:19 root@cumin1003: START - Cookbook sre.mysql.pool pool es2047: Maintenance * 12:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2078.codfw.wmnet with OS trixie * 12:18 dcausse@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 12:18 dcausse@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 12:15 dcausse@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 12:14 dcausse@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 12:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2047 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95793 and previous config saved to /var/cache/conftool/dbconfig/20260730-121404-cwilliams.json * 12:13 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2047.codfw.wmnet with reason: Maintenance * 12:13 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2036: Maintenance * 12:05 ayounsi@dns1004: END - running authdns-update * 12:02 ayounsi@dns1004: START - running authdns-update * 11:51 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2077.codfw.wmnet with OS trixie * 11:48 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:46 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1080.eqiad.wmnet with OS trixie * 11:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1226: Maintenance * 11:41 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2077.codfw.wmnet with reason: host reimage * 11:28 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2077.codfw.wmnet with reason: host reimage * 11:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1080.eqiad.wmnet with reason: host reimage * 11:27 root@cumin1003: START - Cookbook sre.mysql.pool pool es2036: Maintenance * 11:27 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2036: Maintenance * 11:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1080.eqiad.wmnet with reason: host reimage * 11:21 root@cumin1003: START - Cookbook sre.mysql.pool pool es2036: Maintenance * 11:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2036 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95786 and previous config saved to /var/cache/conftool/dbconfig/20260730-111633-cwilliams.json * 11:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2036.codfw.wmnet with reason: Maintenance * 11:08 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2077.codfw.wmnet with OS trixie * 11:07 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 11:03 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie * 11:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1226: Maintenance * 10:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1226 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95783 and previous config saved to /var/cache/conftool/dbconfig/20260730-104801-cwilliams.json * 10:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1226.eqiad.wmnet with reason: Maintenance * 10:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1214: Maintenance * 10:27 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2035: Maintenance * 10:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2076.codfw.wmnet with OS trixie * 10:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1214: Maintenance * 09:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1214 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95775 and previous config saved to /var/cache/conftool/dbconfig/20260730-095451-cwilliams.json * 09:54 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1214.eqiad.wmnet with reason: Maintenance * 09:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1209: Maintenance * 09:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2076.codfw.wmnet with reason: host reimage * 09:42 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool es2035: Maintenance * 09:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.netbox.update-extras (exit_code=0) rolling restart_daemons on A:netbox * 09:41 ayounsi@cumin1003: START - Cookbook sre.netbox.update-extras rolling restart_daemons on A:netbox * 09:40 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2035: Maintenance * 09:39 ayounsi@cumin1003: END (PASS) - Cookbook sre.netbox.update-extras (exit_code=0) rolling restart_daemons on A:netbox-canary * 09:39 ayounsi@cumin1003: START - Cookbook sre.netbox.update-extras rolling restart_daemons on A:netbox-canary * 09:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2076.codfw.wmnet with reason: host reimage * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 09:34 root@cumin1003: START - Cookbook sre.mysql.pool pool es2035: Maintenance * 09:32 lucaswerkmeister-wmde@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 09:32 lucaswerkmeister-wmde@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 09:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2035 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95771 and previous config saved to /var/cache/conftool/dbconfig/20260730-092910-cwilliams.json * 09:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2035.codfw.wmnet with reason: Maintenance * 09:19 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2076.codfw.wmnet with OS trixie * 09:18 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 09:17 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie * 09:07 root@cumin1003: START - Cookbook sre.mysql.pool pool db1209: Maintenance * 09:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 23 hosts * 09:04 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Remove cable label from interfaces descriptions - ayounsi@cumin1003 * 09:04 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:02 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Remove cable label from interfaces descriptions - ayounsi@cumin1003 * 09:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1209 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95767 and previous config saved to /var/cache/conftool/dbconfig/20260730-090133-cwilliams.json * 09:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1209.eqiad.wmnet with reason: Maintenance * 09:01 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1192: Maintenance * 08:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1252: Maintenance * 08:57 jayme@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 08:56 jayme@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 08:53 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 23 hosts * 08:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:51 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:50 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1263: Maintenance * 08:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:23 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 08:15 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie * 08:14 root@cumin1003: START - Cookbook sre.mysql.pool pool db1192: Maintenance * 08:13 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2075.codfw.wmnet with OS trixie * 08:12 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1252: Maintenance * 08:11 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1252: Maintenance * 08:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1252: Maintenance * 08:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1192 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95754 and previous config saved to /var/cache/conftool/dbconfig/20260730-080611-cwilliams.json * 08:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1192.eqiad.wmnet with reason: Maintenance * 08:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1178: Maintenance * 08:05 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS bullseye * 07:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1263: Maintenance * 07:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1263 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95751 and previous config saved to /var/cache/conftool/dbconfig/20260730-075106-cwilliams.json * 07:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[1260-1262].eqiad.wmnet with reason: Maintenance * 07:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2075.codfw.wmnet with reason: host reimage * 07:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1263.eqiad.wmnet with reason: Maintenance * 07:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2075.codfw.wmnet with reason: host reimage * 07:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance * 07:38 dcausse@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:38 dcausse@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 07:35 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db1252', diff saved to https://phabricator.wikimedia.org/P95748 and previous config saved to /var/cache/conftool/dbconfig/20260730-073510-marostegui.json * 07:26 klausman@dns2004: END - running authdns-update * 07:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 07:25 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2075.codfw.wmnet with OS trixie * 07:24 klausman@dns2004: START - running authdns-update * 07:23 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host an-test-master1003.eqiad.wmnet * 07:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db1178: Maintenance * 07:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1178 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95746 and previous config saved to /var/cache/conftool/dbconfig/20260730-071112-cwilliams.json * 07:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1178.eqiad.wmnet with reason: Maintenance * 07:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1177: Maintenance * 06:24 root@cumin1003: START - Cookbook sre.mysql.pool pool db1177: Maintenance * 06:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1177 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95741 and previous config saved to /var/cache/conftool/dbconfig/20260730-061736-cwilliams.json * 06:17 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1177.eqiad.wmnet with reason: Maintenance * 06:17 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1172: Maintenance * 05:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1218.eqiad.wmnet with reason: crashed * 05:41 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db1217 it crashed', diff saved to https://phabricator.wikimedia.org/P95737 and previous config saved to /var/cache/conftool/dbconfig/20260730-054111-marostegui.json * 05:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95736 and previous config saved to /var/cache/conftool/dbconfig/20260730-053422-cwilliams.json * 05:30 root@cumin1003: START - Cookbook sre.mysql.pool pool db1172: Maintenance * 05:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95734 and previous config saved to /var/cache/conftool/dbconfig/20260730-052414-cwilliams.json * 05:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1172 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95733 and previous config saved to /var/cache/conftool/dbconfig/20260730-052354-cwilliams.json * 05:23 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1172.eqiad.wmnet with reason: Maintenance * 05:23 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1167: Maintenance * 05:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95731 and previous config saved to /var/cache/conftool/dbconfig/20260730-051406-cwilliams.json * 05:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95729 and previous config saved to /var/cache/conftool/dbconfig/20260730-050358-cwilliams.json * 04:47 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 04:47 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 04:47 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 04:35 root@cumin1003: START - Cookbook sre.mysql.pool pool db1167: Maintenance * 04:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1167 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95726 and previous config saved to /var/cache/conftool/dbconfig/20260730-042923-cwilliams.json * 04:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 04:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1167.eqiad.wmnet with reason: Maintenance * 04:22 pt1979@cumin2002: START - Cookbook sre.dns.netbox * 04:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95725 and previous config saved to /var/cache/conftool/dbconfig/20260730-040337-cwilliams.json * 04:03 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:38 brett@cumin2002: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool eqsin [reason: Switch upgrade maintenance window complete, [[phab:T433097|T433097]]] * 01:38 brett@cumin2002: START - Cookbook sre.dns.admin DNS admin: pool eqsin [reason: Switch upgrade maintenance window complete, [[phab:T433097|T433097]]] == 2026-07-29 == * 23:57 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin,mr1-eqsin IPv6,mr1-eqsin.oob,mr1-eqsin.oob IPv6 with reason: connection issue * 22:54 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 22:53 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 22:53 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 22:53 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:25 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2022.codfw.wmnet, repooling source-only afterwards * 22:20 brett@cumin2002: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool eqsin [reason: Switch upgrade maintenance window, [[phab:T433097|T433097]]] * 22:20 brett@cumin2002: START - Cookbook sre.dns.admin DNS admin: depool eqsin [reason: Switch upgrade maintenance window, [[phab:T433097|T433097]]] * 22:01 apine@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 22:00 apine@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 21:59 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 21:58 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 21:58 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 21:58 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 21:32 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1048.eqiad.wmnet with OS trixie * 21:25 pt1979@cumin2002: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be2097.codfw.wmnet with OS bullseye * 21:16 zabe@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki=metawiki 'Mental Health Resource Center' 'Safety Resource Center/Mental Health' Zabe --reason 'per request [[:phab:T433118{{!}}T433118]]' * 21:12 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2022.codfw.wmnet, repooling source-only afterwards * 21:12 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] (duration: 12m 53s) * 21:12 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2015\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 21:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1253: Maintenance * 21:08 aaron@deploy1003: aaron: Continuing with deployment * 21:01 aaron@deploy1003: aaron: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:59 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] * 20:52 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] (duration: 21m 57s) * 20:48 aaron@deploy1003: aaron: Continuing with deployment * 20:32 aaron@deploy1003: aaron: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:30 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] * 20:24 root@cumin1003: START - Cookbook sre.mysql.pool pool db1253: Maintenance * 20:19 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] (duration: 08m 07s) * 20:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1253 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95719 and previous config saved to /var/cache/conftool/dbconfig/20260729-201810-cwilliams.json * 20:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1253.eqiad.wmnet with reason: Maintenance * 20:17 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1231: Maintenance * 20:15 aaron@deploy1003: bpirkle, aaron: Continuing with deployment * 20:13 aaron@deploy1003: bpirkle, aaron: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:12 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie * 20:11 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] * 20:11 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1048.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:09 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1048.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:09 pt1979@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 20:04 pt1979@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 19:47 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 19:43 pt1979@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye * 19:41 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 19:37 zabe: zabe@deploy1003:~$ mwscript-k8s --comment='[[phab:T433529|T433529]]' --follow -- resetAuthenticationThrottle.php --wiki=aawiki --signup --ip=89.36.114.94 * 19:36 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] (duration: 06m 49s) * 19:32 zabe@deploy1003: zabe: Continuing with deployment * 19:31 zabe@deploy1003: zabe: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:31 root@cumin1003: START - Cookbook sre.mysql.pool pool db1231: Maintenance * 19:29 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] * 19:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1231 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95714 and previous config saved to /var/cache/conftool/dbconfig/20260729-192454-cwilliams.json * 19:24 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1231.eqiad.wmnet with reason: Maintenance * 19:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1227: Maintenance * 19:22 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 19:22 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 19:21 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:21 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1048] - vriley@cumin1003" * 19:21 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1048] - vriley@cumin1003" * 19:19 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 19:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1251: Maintenance * 19:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95711 and previous config saved to /var/cache/conftool/dbconfig/20260729-191756-cwilliams.json * 19:16 vriley@cumin1003: START - Cookbook sre.dns.netbox * 19:11 dduvall: rolling back wmf.13 to group0 due to [[phab:T433457|T433457]] (cc [[phab:T430832|T430832]]) * 19:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95709 and previous config saved to /var/cache/conftool/dbconfig/20260729-190748-cwilliams.json * 19:01 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2022.codfw.wmnet with OS bookworm * 19:01 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2021.codfw.wmnet, repooling source-only afterwards * 18:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95707 and previous config saved to /var/cache/conftool/dbconfig/20260729-185740-cwilliams.json * 18:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95704 and previous config saved to /var/cache/conftool/dbconfig/20260729-184732-cwilliams.json * 18:37 root@cumin1003: START - Cookbook sre.mysql.pool pool db1227: Maintenance * 18:34 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2022.codfw.wmnet with reason: host reimage * 18:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1227 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95701 and previous config saved to /var/cache/conftool/dbconfig/20260729-183117-cwilliams.json * 18:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1227.eqiad.wmnet with reason: Maintenance * 18:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1202: Maintenance * 18:30 root@cumin1003: START - Cookbook sre.mysql.pool pool db1251: Maintenance * 18:27 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2022.codfw.wmnet with reason: host reimage * 18:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1251 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95698 and previous config saved to /var/cache/conftool/dbconfig/20260729-182428-cwilliams.json * 18:24 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lvs2014.codfw.wmnet * 18:24 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for lvs2014.codfw.wmnet * 18:24 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1251.eqiad.wmnet with reason: Maintenance * 18:23 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1235: Maintenance * 18:22 brett@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 18:19 brett@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 18:19 brett@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 18:17 brett@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 18:17 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 18:16 mutante: removing jenkins during the train - living on the edge - no, just kidding, jenkins has migrated to dedicated machines, nothing should happen * 18:15 brett@cumin2002: END (ERROR) - Cookbook sre.loadbalancer.restart-pybal (exit_code=97) rolling-restart of pybal on P<nowiki>{</nowiki>lvs2014.codfw.wmnet<nowiki>}</nowiki> and A:lvs ([[phab:T428495|T428495]]) * 18:15 mutante: CI: contint1002/contint2002: apt-get remove --purge jenkins - jenkins be gone - [[phab:T418521|T418521]] * 18:13 brett@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on P<nowiki>{</nowiki>lvs2014.codfw.wmnet<nowiki>}</nowiki> and A:lvs ([[phab:T428495|T428495]]) * 18:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2022 * 18:08 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2022 * 18:03 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T428495|T428495]] * 18:03 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2022 * 18:02 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2022.codfw.wmnet 211.48.192.10.in-addr.arpa 1.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:02 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2022.codfw.wmnet 211.48.192.10.in-addr.arpa 1.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:02 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:02 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2022 - bking@cumin2003" * 18:02 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2022 - bking@cumin2003" * 17:57 bking@cumin2003: START - Cookbook sre.dns.netbox * 17:56 brett@cumin2002: END (FAIL) - Cookbook sre.loadbalancer.restart-pybal (exit_code=1) rolling-restart of pybal on A:lvs-codfw and A:lvs ([[phab:T428495|T428495]]) * 17:55 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo - [[phab:T428495|T428495]] * 17:54 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2022 * 17:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2022.codfw.wmnet with OS bookworm * 17:50 brett@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on A:lvs-codfw and A:lvs ([[phab:T428495|T428495]]) * 17:47 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2021.codfw.wmnet, repooling source-only afterwards * 17:47 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 14s) * 17:47 swfrench-wmf: authdns-update to direct codfw, eqsin, ulsfo etcd clients back to codfw - [[phab:T428495|T428495]] * 17:47 swfrench@dns1004: END - running authdns-update * 17:47 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 17:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95692 and previous config saved to /var/cache/conftool/dbconfig/20260729-174713-cwilliams.json * 17:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance * 17:46 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1249: Maintenance * 17:45 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2015\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 17:45 swfrench@dns1004: START - running authdns-update * 17:44 root@cumin1003: START - Cookbook sre.mysql.pool pool db1202: Maintenance * 17:41 akhatun: Deployed refinery using scap, then deployed onto hdfs * 17:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1202 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95688 and previous config saved to /var/cache/conftool/dbconfig/20260729-173759-cwilliams.json * 17:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1202.eqiad.wmnet with reason: Maintenance * 17:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1194: Maintenance * 17:37 root@cumin1003: START - Cookbook sre.mysql.pool pool db1235: Maintenance * 17:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1230: Maintenance * 17:30 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1235 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95684 and previous config saved to /var/cache/conftool/dbconfig/20260729-173051-cwilliams.json * 17:30 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1235.eqiad.wmnet with reason: Maintenance * 17:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1234: Maintenance * 17:26 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (thin): Regular analytics weekly train THIN [analytics/refinery@56695674] (duration: 02m 02s) * 17:24 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (thin): Regular analytics weekly train THIN [analytics/refinery@56695674] * 17:23 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567]: Regular analytics weekly train [analytics/refinery@56695674] (duration: 06m 20s) * 17:20 dancy@deploy1003: Finished scap sync-world: Testing delay_messageblobstore_purge: true (duration: 06m 29s) * 17:17 akhatun@deploy1003: Started deploy [analytics/refinery@5669567]: Regular analytics weekly train [analytics/refinery@56695674] * 17:17 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] (duration: 00m 22s) * 17:16 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] * 17:13 dancy@deploy1003: Started scap sync-world: Testing delay_messageblobstore_purge: true * 17:05 mutante: CI: contint1002/contint2002 - restarted httpd to be extra sure all is cleaned up - https://integration.wikimedia.org/ci/ is up and running [[phab:T418521|T418521]] * 17:04 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 17:03 mutante: CI: contint1002/contint2002 - rm /etc/apache2/jenkins_proxy - removing legacy jenkins proxy config - jenkins is on new dedicated machines and uses jenkins_proxy_ext config [[phab:T418521|T418521]] * 17:02 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] (duration: 36m 25s) * 17:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1249: Maintenance * 16:59 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 16:54 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2015.codfw.wmnet, repooling source-only afterwards * 16:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1249 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95674 and previous config saved to /var/cache/conftool/dbconfig/20260729-165339-cwilliams.json * 16:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1249.eqiad.wmnet with reason: Maintenance * 16:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1248: Maintenance * 16:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1194: Maintenance * 16:47 swfrench-wmf: silenced EtcdReplicationDown 57b2b421-1cc9-4e38-9276-{{Gerrit|94f223fd231c}} - [[phab:T428495|T428495]] * 16:46 tchin@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/eventstreams-internal: apply * 16:46 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye * 16:46 tchin@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/eventstreams-internal: apply * 16:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1230: Maintenance * 16:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1194 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95669 and previous config saved to /var/cache/conftool/dbconfig/20260729-164422-cwilliams.json * 16:44 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 16:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1194.eqiad.wmnet with reason: Maintenance * 16:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1191: Maintenance * 16:43 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host an-test-master1003.eqiad.wmnet * 16:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db1234: Maintenance * 16:43 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 16:43 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Rolling back deployment * 16:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host an-test-master1004.eqiad.wmnet * 16:41 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 16:40 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 16:40 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 16:39 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 16:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1230 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95667 and previous config saved to /var/cache/conftool/dbconfig/20260729-163932-cwilliams.json * 16:39 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 16:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1230.eqiad.wmnet with reason: Maintenance * 16:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1207: Maintenance * 16:38 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 16:37 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host an-test-master1004.eqiad.wmnet * 16:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1234 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95664 and previous config saved to /var/cache/conftool/dbconfig/20260729-163719-cwilliams.json * 16:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1234.eqiad.wmnet with reason: Maintenance * 16:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1079.eqiad.wmnet with OS trixie * 16:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1232: Maintenance * 16:34 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1259: Maintenance * 16:28 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] (duration: 06m 57s) * 16:28 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:26 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] * 16:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1051 hosts * 16:21 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] * 16:20 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2006.codfw.wmnet with OS bookworm * 16:19 akhatun: Deploying Refinery at {{Gerrit|56695674}} as part of weekly train * 16:18 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1079.eqiad.wmnet with reason: host reimage * 16:16 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] (duration: 15m 36s) * 16:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2021.codfw.wmnet with OS bookworm * 16:14 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1079.eqiad.wmnet with reason: host reimage * 16:12 topranks: hot-swap line card in FPC0 on cr1-eqiad with replacement MPC10E from Juniper [[phab:T426343|T426343]] * 16:10 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Continuing with deployment * 16:07 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db1248: Maintenance * 16:01 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] * 16:00 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 16:00 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 15:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1248 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95651 and previous config saved to /var/cache/conftool/dbconfig/20260729-155956-cwilliams.json * 15:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1248.eqiad.wmnet with reason: Maintenance * 15:59 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2006.codfw.wmnet with reason: host reimage * 15:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1247: Maintenance * 15:59 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2074.codfw.wmnet with OS trixie * 15:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1191: Maintenance * 15:57 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:55 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1079.eqiad.wmnet with OS trixie * 15:55 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2006.codfw.wmnet with reason: host reimage * 15:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1207: Maintenance * 15:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1191 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95646 and previous config saved to /var/cache/conftool/dbconfig/20260729-155104-cwilliams.json * 15:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1191.eqiad.wmnet with reason: Maintenance * 15:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1181: Maintenance * 15:49 root@cumin1003: START - Cookbook sre.mysql.pool pool db1232: Maintenance * 15:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2021.codfw.wmnet with reason: host reimage * 15:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1207 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95643 and previous config saved to /var/cache/conftool/dbconfig/20260729-154735-cwilliams.json * 15:47 root@cumin1003: START - Cookbook sre.mysql.pool pool db1259: Maintenance * 15:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1207.eqiad.wmnet with reason: Maintenance * 15:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1200: Maintenance * 15:46 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] (duration: 31m 59s) * 15:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 15:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1232 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95640 and previous config saved to /var/cache/conftool/dbconfig/20260729-154330-cwilliams.json * 15:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1232.eqiad.wmnet with reason: Maintenance * 15:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1219: Maintenance * 15:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2015.codfw.wmnet, repooling source-only afterwards * 15:41 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 18s) * 15:41 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1259 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95638 and previous config saved to /var/cache/conftool/dbconfig/20260729-154107-cwilliams.json * 15:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1259.eqiad.wmnet with reason: Maintenance * 15:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1254: Maintenance * 15:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2015.codfw.wmnet with OS bookworm * 15:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2021.codfw.wmnet with reason: host reimage * 15:36 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2006.codfw.wmnet with OS bookworm * 15:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 15:35 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Continuing with deployment * 15:33 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2074.codfw.wmnet with OS trixie * 15:33 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 15:32 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:29 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:28 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be2074.codfw.wmnet with OS trixie * 15:28 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2006.codfw.wmnet * 15:26 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1078.eqiad.wmnet with OS trixie * 15:25 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 15:25 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:22 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2006.codfw.wmnet * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2021 * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2021 * 15:19 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2021 * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2021.codfw.wmnet 210.48.192.10.in-addr.arpa 0.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:19 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2021.codfw.wmnet 210.48.192.10.in-addr.arpa 0.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2021 - bking@cumin2003" * 15:19 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2021 - bking@cumin2003" * 15:14 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] * 15:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2015.codfw.wmnet with reason: host reimage * 15:11 root@cumin1003: START - Cookbook sre.mysql.pool pool db1247: Maintenance * 15:11 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on ml-serve2004.codfw.wmnet with reason: [[phab:T433478|T433478]] * 15:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2015.codfw.wmnet with reason: host reimage * 15:10 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on ml-serve2002.codfw.wmnet with reason: [[phab:T433476|T433476]] * 15:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 15:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1247 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95625 and previous config saved to /var/cache/conftool/dbconfig/20260729-150459-cwilliams.json * 15:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1247.eqiad.wmnet with reason: Maintenance * 15:04 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:04 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1244: Maintenance * 15:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1078.eqiad.wmnet with reason: host reimage * 15:03 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1005.wikimedia.org * 15:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db1181: Maintenance * 15:01 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2098.codfw.wmnet with OS bullseye * 15:00 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye * 15:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1200: Maintenance * 14:59 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 14:59 root@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1285.eqiad.wmnet with OS trixie * 14:59 Amir1: mwscript-k8s -- extensions/TimedMediaHandler/maintenance/requeueTranscodes.php --wiki=commonswiki --key '360p.mpeg4.mov' --throttle --video --missing ([[phab:T358266|T358266]]) * 14:58 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1005.wikimedia.org * 14:58 jhancock@cumin2002: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['ms-be2098'] * 14:58 jhancock@cumin2002: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['ms-be2098'] * 14:58 jhancock@cumin2002: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['ms-be2097'] * 14:58 jhancock@cumin2002: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['ms-be2097'] * 14:58 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1078.eqiad.wmnet with reason: host reimage * 14:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1181 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95621 and previous config saved to /var/cache/conftool/dbconfig/20260729-145629-cwilliams.json * 14:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1181.eqiad.wmnet with reason: Maintenance * 14:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1174: Maintenance * 14:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1219: Maintenance * 14:55 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader1006.wikimedia.org on all recursors * 14:55 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader1006.wikimedia.org on all recursors * 14:55 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader1005.wikimedia.org on all recursors * 14:55 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader1005.wikimedia.org on all recursors * 14:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1200 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95618 and previous config saved to /var/cache/conftool/dbconfig/20260729-145336-cwilliams.json * 14:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1254: Maintenance * 14:53 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1200.eqiad.wmnet with reason: Maintenance * 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1185: Maintenance * 14:52 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2021 * 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2015 * 14:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2015 * 14:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1219 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95616 and previous config saved to /var/cache/conftool/dbconfig/20260729-144946-cwilliams.json * 14:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1219.eqiad.wmnet with reason: Maintenance * 14:49 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1218: Maintenance * 14:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2021.codfw.wmnet with OS bookworm * 14:48 dancy@deploy1003: Finished deploy [zuul/deploy@22703a6]: Deploying https://gerrit.wikimedia.org/r/c/integration/zuul/+/1311501 ([[phab:T432491|T432491]]) (duration: 00m 15s) * 14:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2015.codfw.wmnet with OS bookworm * 14:48 dancy@deploy1003: Started deploy [zuul/deploy@22703a6]: Deploying https://gerrit.wikimedia.org/r/c/integration/zuul/+/1311501 ([[phab:T432491|T432491]]) * 14:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1254 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95613 and previous config saved to /var/cache/conftool/dbconfig/20260729-144729-cwilliams.json * 14:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1254.eqiad.wmnet with reason: Maintenance * 14:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1233: Maintenance * 14:46 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2013\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 14:46 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2014\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 14:46 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:45 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:44 root@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1285.eqiad.wmnet with reason: host reimage * 14:43 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:42 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:41 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:40 root@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1285.eqiad.wmnet with reason: host reimage * 14:39 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1078.eqiad.wmnet with OS trixie * 14:39 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2074.codfw.wmnet with OS trixie * 14:32 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:32 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:32 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:31 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2005.codfw.wmnet with OS bookworm * 14:30 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:30 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:29 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:29 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:27 root@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host db1285 * 14:27 root@cumin1003: START - Cookbook sre.hosts.move-vlan for host db1285 * 14:27 root@cumin1003: START - Cookbook sre.hosts.reimage for host db1285.eqiad.wmnet with OS trixie * 14:24 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 14:24 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:24 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:24 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:23 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:22 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:22 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:21 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:17 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:16 root@cumin1003: START - Cookbook sre.mysql.pool pool db1244: Maintenance * 14:15 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:15 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add asw1-604 loopback ipv4 - pt1979@cumin2002" * 14:15 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add asw1-604 loopback ipv4 - pt1979@cumin2002" * 14:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:12 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 14:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95599 and previous config saved to /var/cache/conftool/dbconfig/20260729-141014-cwilliams.json * 14:10 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 14:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1244.eqiad.wmnet with reason: Maintenance * 14:10 pt1979@cumin2002: START - Cookbook sre.dns.netbox * 14:10 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 14:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1243: Maintenance * 14:09 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2005.codfw.wmnet with reason: host reimage * 14:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db1174: Maintenance * 14:08 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad * 14:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db1185: Maintenance * 14:06 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 14:05 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2005.codfw.wmnet with reason: host reimage * 14:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1174 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95595 and previous config saved to /var/cache/conftool/dbconfig/20260729-140309-cwilliams.json * 14:03 sukhe@dns1004: END - running authdns-update * 14:03 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1174.eqiad.wmnet with reason: Maintenance * 14:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1170: Maintenance * 14:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db1218: Maintenance * 14:01 sukhe@dns1004: START - running authdns-update * 14:00 sukhe@dns1004: START - running authdns-update * 13:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db1233: Maintenance * 13:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1185 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95592 and previous config saved to /var/cache/conftool/dbconfig/20260729-135925-cwilliams.json * 13:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1185.eqiad.wmnet with reason: Maintenance * 13:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1161: Maintenance * 13:58 sukhe@puppetserver1001: conftool action : set/pooled=true; selector: dnsdisc=urldownloader * 13:58 root@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1265.eqiad.wmnet with OS trixie * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1218 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95590 and previous config saved to /var/cache/conftool/dbconfig/20260729-135621-cwilliams.json * 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1218.eqiad.wmnet with reason: Maintenance * 13:55 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1206: Maintenance * 13:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2073.codfw.wmnet with OS trixie * 13:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1233 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95587 and previous config saved to /var/cache/conftool/dbconfig/20260729-135335-cwilliams.json * 13:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1233.eqiad.wmnet with reason: Maintenance * 13:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1229: Maintenance * 13:50 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/kartotherian: apply * 13:50 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service * 13:49 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:49 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/kartotherian: apply * 13:48 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 13:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1077.eqiad.wmnet with OS trixie * 13:47 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 13:46 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2005.codfw.wmnet with OS bookworm * 13:44 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 13:44 root@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1265.eqiad.wmnet with reason: host reimage * 13:40 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] (duration: 09m 22s) * 13:39 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:38 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:36 root@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1265.eqiad.wmnet with reason: host reimage * 13:35 stran@deploy1003: stran: Continuing with deployment * 13:33 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 13:32 stran@deploy1003: stran: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified t * 13:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2073.codfw.wmnet with reason: host reimage * 13:30 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ml-build1001.eqiad.wmnet * 13:30 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] * 13:29 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad * 13:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1077.eqiad.wmnet with reason: host reimage * 13:27 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 13:27 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:27 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad * 13:26 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2073.codfw.wmnet with reason: host reimage * 13:26 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] (duration: 07m 56s) * 13:25 klausman@cumin1003: START - Cookbook sre.hosts.reboot-single for host ml-build1001.eqiad.wmnet * 13:24 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 13:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>ml-serve101[2-5].eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 13:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1015.eqiad.wmnet * 13:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1015.eqiad.wmnet * 13:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1077.eqiad.wmnet with reason: host reimage * 13:23 root@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host db1265 * 13:23 root@cumin1003: START - Cookbook sre.hosts.move-vlan for host db1265 * 13:23 root@cumin1003: START - Cookbook sre.hosts.reimage for host db1265.eqiad.wmnet with OS trixie * 13:23 root@cumin1003: START - Cookbook sre.mysql.pool pool db1243: Maintenance * 13:22 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 13:22 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2005.codfw.wmnet * 13:22 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 13:22 samtar@deploy1003: dreamrimmer, samtar: Continuing with deployment * 13:20 samtar@deploy1003: dreamrimmer, samtar: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts an-test-master[1001-1002].eqiad.wmnet * 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-master[1001-1002].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 13:18 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1015.eqiad.wmnet * 13:18 sukhe@cumin1003: END (ERROR) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=97) for role: url_downloader@eqiad * 13:18 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 13:18 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] * 13:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95574 and previous config saved to /var/cache/conftool/dbconfig/20260729-131638-cwilliams.json * 13:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1243.eqiad.wmnet with reason: Maintenance * 13:16 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2005.codfw.wmnet * 13:16 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1242: Maintenance * 13:14 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] (duration: 07m 00s) * 13:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db1170: Maintenance * 13:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1015.eqiad.wmnet * 13:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1014.eqiad.wmnet * 13:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1014.eqiad.wmnet * 13:12 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1223: Maintenance * 13:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db1161: Maintenance * 13:10 samtar@deploy1003: anzx, samtar: Continuing with deployment * 13:09 samtar@deploy1003: anzx, samtar: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db1206: Maintenance * 13:08 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 13:07 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] * 13:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1170 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95566 and previous config saved to /var/cache/conftool/dbconfig/20260729-130730-cwilliams.json * 13:07 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1170.eqiad.wmnet with reason: Maintenance * 13:07 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:07 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt IPs new switches - cmooney@cumin1003" * 13:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1158: Maintenance * 13:06 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1014.eqiad.wmnet * 13:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1161 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95564 and previous config saved to /var/cache/conftool/dbconfig/20260729-130616-cwilliams.json * 13:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 13:06 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1077.eqiad.wmnet with OS trixie * 13:05 root@cumin1003: START - Cookbook sre.mysql.pool pool db1229: Maintenance * 13:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1161.eqiad.wmnet with reason: Maintenance * 13:05 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt IPs new switches - cmooney@cumin1003" * 13:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2073.codfw.wmnet with OS trixie * 13:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Maintenance * 13:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1206 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95562 and previous config saved to /var/cache/conftool/dbconfig/20260729-130258-cwilliams.json * 13:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1206.eqiad.wmnet with reason: Maintenance * 13:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1196: Maintenance * 13:01 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 13:01 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:00 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 13:00 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 12:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1229 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95559 and previous config saved to /var/cache/conftool/dbconfig/20260729-125950-cwilliams.json * 12:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1229.eqiad.wmnet with reason: Maintenance * 12:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1222: Maintenance * 12:57 sukhe: sudo cumin 'A:lvs and (A:eqiad or A:codfw)' 'disable-puppet "adding new service urldownloader"': [[phab:T429175|T429175]] * 12:56 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1014.eqiad.wmnet * 12:56 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1013.eqiad.wmnet * 12:56 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1013.eqiad.wmnet * 12:50 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1013.eqiad.wmnet * 12:50 sukhe: sudo cumin 'O:url_downloader' 'run-puppet-agent --enable "merging CR 1313948"': [[phab:T429175|T429175]] * 12:48 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-master[1001-1002].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 12:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1013.eqiad.wmnet * 12:45 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1012.eqiad.wmnet * 12:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1012.eqiad.wmnet * 12:45 sukhe: sudo cumin 'O:url_downloader' 'disable-puppet "merging CR 1313948"': [[phab:T429175|T429175]] * 12:44 btullis@cumin1003: START - Cookbook sre.dns.netbox * 12:40 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test2001.codfw.wmnet * 12:40 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test2001.codfw.wmnet * 12:38 ayounsi@dns1004: END - running authdns-update * 12:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1012.eqiad.wmnet * 12:35 ayounsi@dns1004: START - running authdns-update * 12:34 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts an-test-master[1001-1002].eqiad.wmnet * 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts an-test-coord1001.eqiad.wmnet * 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-coord1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 12:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1012.eqiad.wmnet * 12:32 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>ml-serve101[2-5].eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 12:29 root@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Maintenance * 12:25 root@cumin1003: START - Cookbook sre.mysql.pool pool db1223: Maintenance * 12:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95544 and previous config saved to /var/cache/conftool/dbconfig/20260729-122254-cwilliams.json * 12:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1242.eqiad.wmnet with reason: Maintenance * 12:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1241: Maintenance * 12:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1051 hosts * 12:20 root@cumin1003: START - Cookbook sre.mysql.pool pool db1158: Maintenance * 12:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1223 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95540 and previous config saved to /var/cache/conftool/dbconfig/20260729-121937-cwilliams.json * 12:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1223.eqiad.wmnet with reason: Maintenance * 12:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1212: Maintenance * 12:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db1159: Maintenance * 12:17 elukey@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: sync * 12:15 elukey@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: sync * 12:15 root@cumin1003: START - Cookbook sre.mysql.pool pool db1196: Maintenance * 12:14 Daimona: Creating new DB tables for the CampaignEvents extension in x1.testwiki, x1.test2wiki, x1.officewiki, and x1.wikishared # [[phab:T429339|T429339]] * 12:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db1222: Maintenance * 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95535 and previous config saved to /var/cache/conftool/dbconfig/20260729-121211-cwilliams.json * 12:12 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 12:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1158.eqiad.wmnet with reason: Maintenance * 12:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1159 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95534 and previous config saved to /var/cache/conftool/dbconfig/20260729-121146-cwilliams.json * 12:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1159.eqiad.wmnet with reason: Maintenance * 12:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1196 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95533 and previous config saved to /var/cache/conftool/dbconfig/20260729-120847-cwilliams.json * 12:08 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 12:08 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1196.eqiad.wmnet with reason: Maintenance * 12:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1195: Maintenance * 12:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1222 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95530 and previous config saved to /var/cache/conftool/dbconfig/20260729-120424-cwilliams.json * 12:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1222.eqiad.wmnet with reason: Maintenance * 12:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1098 hosts * 12:00 marostegui: Rename tables [[phab:T425074|T425074]] * 12:00 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-coord1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 11:58 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1197: Maintenance * 11:55 btullis@cumin1003: START - Cookbook sre.dns.netbox * 11:52 elukey@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: sync * 11:51 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:51 elukey@deploy1003: helmfile [codfw] START helmfile.d/services/proton: sync * 11:51 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:50 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts an-test-coord1001.eqiad.wmnet * 11:50 elukey@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: sync * 11:49 elukey@deploy1003: helmfile [staging] START helmfile.d/services/proton: sync * 11:35 root@cumin1003: START - Cookbook sre.mysql.pool pool db1241: Maintenance * 11:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db1212: Maintenance * 11:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1241 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95520 and previous config saved to /var/cache/conftool/dbconfig/20260729-112918-cwilliams.json * 11:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1241.eqiad.wmnet with reason: Maintenance * 11:29 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1238: Maintenance * 11:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1212 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95517 and previous config saved to /var/cache/conftool/dbconfig/20260729-112727-cwilliams.json * 11:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 11:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1212.eqiad.wmnet with reason: Maintenance * 11:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1198: Maintenance * 11:23 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:22 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:21 root@cumin1003: START - Cookbook sre.mysql.pool pool db1195: Maintenance * 11:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1195 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95514 and previous config saved to /var/cache/conftool/dbconfig/20260729-111450-cwilliams.json * 11:14 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1195.eqiad.wmnet with reason: Maintenance * 11:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1186: Maintenance * 11:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 11:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 11:05 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:54 marostegui: Dropping renamed tables [[phab:T425066|T425066]] * 10:41 root@cumin1003: START - Cookbook sre.mysql.pool pool db1238: Maintenance * 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1198: Maintenance * 10:39 Amir1: ran https://phabricator.wikimedia.org/T432509#12149723 in production ([[phab:T432509|T432509]]) * 10:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db1197: Maintenance * 10:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1238 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95501 and previous config saved to /var/cache/conftool/dbconfig/20260729-103532-cwilliams.json * 10:35 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1238.eqiad.wmnet with reason: Maintenance * 10:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1221: Maintenance * 10:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1198 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95499 and previous config saved to /var/cache/conftool/dbconfig/20260729-103330-cwilliams.json * 10:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1198.eqiad.wmnet with reason: Maintenance * 10:33 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1175: Maintenance * 10:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1197 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95496 and previous config saved to /var/cache/conftool/dbconfig/20260729-103217-cwilliams.json * 10:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1197.eqiad.wmnet with reason: Maintenance * 10:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1188: Maintenance * 10:27 root@cumin1003: START - Cookbook sre.mysql.pool pool db1186: Maintenance * 10:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1186 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95493 and previous config saved to /var/cache/conftool/dbconfig/20260729-102111-cwilliams.json * 10:21 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1186.eqiad.wmnet with reason: Maintenance * 10:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 10:14 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 09:53 XioNoX: reboot cr2-magru - [[phab:T431750|T431750]] * 09:52 XioNoX: drain cr2-magru - [[phab:T431750|T431750]] * 09:48 root@cumin1003: START - Cookbook sre.mysql.pool pool db1221: Maintenance * 09:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zookeeper-test1002.eqiad.wmnet * 09:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1188: Maintenance * 09:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1175: Maintenance * 09:44 btullis@dns1004: END - running authdns-update * 09:42 btullis@dns1004: START - running authdns-update * 09:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1221 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95483 and previous config saved to /var/cache/conftool/dbconfig/20260729-094200-cwilliams.json * 09:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 7 hosts with reason: Maintenance * 09:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1221.eqiad.wmnet with reason: Maintenance * 09:41 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host zookeeper-test1002.eqiad.wmnet * 09:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1199: Maintenance * 09:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1188 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95481 and previous config saved to /var/cache/conftool/dbconfig/20260729-093917-cwilliams.json * 09:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1188.eqiad.wmnet with reason: Maintenance * 09:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1182: Maintenance * 09:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1175 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95479 and previous config saved to /var/cache/conftool/dbconfig/20260729-093842-cwilliams.json * 09:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1175.eqiad.wmnet with reason: Maintenance * 09:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1166: Maintenance * 09:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1169: Maintenance * 09:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1033.eqiad.wmnet,service=s8 * 09:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1033.eqiad.wmnet,service=s5 * 09:33 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1033.eqiad.wmnet,service=s8 * 09:33 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1033.eqiad.wmnet,service=s5 * 09:21 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:21 XioNoX: reboot cr1-magru - [[phab:T431750|T431750]] * 09:17 XioNoX: drain cr1-magru - [[phab:T431750|T431750]] * 09:15 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm1001.wikimedia.org * 09:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply * 09:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply * 09:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 09:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 09:11 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr2-magru,cr2-magru IPv6,cr2-magru.mgmt with reason: router upgrade * 09:11 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm1001.wikimedia.org * 09:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp1005.wikimedia.org * 09:07 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp1005.wikimedia.org * 09:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp2005.wikimedia.org * 09:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp2005.wikimedia.org * 09:00 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr1-magru,cr1-magru IPv6,cr1-magru.mgmt with reason: router upgrade * 09:00 marostegui: Dropping renamed tables [[phab:T426341|T426341]] * 08:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1199: Maintenance * 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1182: Maintenance * 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1166: Maintenance * 08:47 root@cumin1003: START - Cookbook sre.mysql.pool pool db1169: Maintenance * 08:46 ayounsi@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 1:00:00 on cr1-magru,cr1-magru IPv6,cr1-magru.mgmt with reason: router upgrade * 08:45 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2072.codfw.wmnet with OS trixie * 08:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1199 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95464 and previous config saved to /var/cache/conftool/dbconfig/20260729-084534-cwilliams.json * 08:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1199.eqiad.wmnet with reason: Maintenance * 08:45 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 08:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1190: Maintenance * 08:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1182 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95462 and previous config saved to /var/cache/conftool/dbconfig/20260729-084436-cwilliams.json * 08:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1182.eqiad.wmnet with reason: Maintenance * 08:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1156: Maintenance * 08:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1166 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95460 and previous config saved to /var/cache/conftool/dbconfig/20260729-084400-cwilliams.json * 08:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1166.eqiad.wmnet with reason: Maintenance * 08:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1157: Maintenance * 08:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95458 and previous config saved to /var/cache/conftool/dbconfig/20260729-084147-cwilliams.json * 08:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1169.eqiad.wmnet with reason: Maintenance * 08:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1163: Maintenance * 08:30 btullis@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 11 hosts with reason: Replacing the namenodes * 08:23 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2072.codfw.wmnet with reason: host reimage * 08:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1098 hosts * 08:19 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2072.codfw.wmnet with reason: host reimage * 07:58 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2072.codfw.wmnet with OS trixie * 07:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1190: Maintenance * 07:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1156: Maintenance * 07:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1157: Maintenance * 07:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1163: Maintenance * 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1190 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95444 and previous config saved to /var/cache/conftool/dbconfig/20260729-074930-cwilliams.json * 07:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1190.eqiad.wmnet with reason: Maintenance * 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1157 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95443 and previous config saved to /var/cache/conftool/dbconfig/20260729-074914-cwilliams.json * 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1156 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95442 and previous config saved to /var/cache/conftool/dbconfig/20260729-074906-cwilliams.json * 07:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1157.eqiad.wmnet with reason: Maintenance * 07:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 07:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1156.eqiad.wmnet with reason: Maintenance * 07:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1163 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95441 and previous config saved to /var/cache/conftool/dbconfig/20260729-074652-cwilliams.json * 07:46 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1163.eqiad.wmnet with reason: Maintenance * 07:46 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2034.codfw.wmnet * 07:42 ayounsi@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2034.codfw.wmnet * 07:42 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2232.codfw.wmnet with OS trixie * 07:34 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:34 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:33 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:31 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:19 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2232.codfw.wmnet with reason: host reimage * 07:15 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2232.codfw.wmnet with reason: host reimage * 06:58 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db2232.codfw.wmnet with OS trixie * 06:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[2160,2232].codfw.wmnet with reason: Reimage * 06:26 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1164.eqiad.wmnet with OS trixie * 06:05 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1164.eqiad.wmnet with reason: host reimage * 06:01 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1164.eqiad.wmnet with reason: host reimage * 05:47 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1164.eqiad.wmnet with OS trixie * 05:46 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1164.eqiad.wmnet with reason: Reimage == 2026-07-28 == * 22:50 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1138.eqiad.wmnet * 22:50 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1138.eqiad.wmnet * 22:49 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1138.eqiad.wmnet * 22:11 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2014.codfw.wmnet, repooling source-only afterwards * 22:08 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2013.codfw.wmnet, repooling source-only afterwards * 22:03 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 20:58 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] (duration: 08m 19s) * 20:55 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2014.codfw.wmnet, repooling source-only afterwards * 20:55 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2013.codfw.wmnet, repooling source-only afterwards * 20:54 arlolra@deploy1003: arlolra: Continuing with deployment * 20:54 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 14s) * 20:54 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 20:53 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 30s) * 20:53 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 20:52 arlolra@deploy1003: arlolra: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:51 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:50 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] * 20:49 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:43 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 20:34 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] (duration: 06m 54s) * 20:34 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:34 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:31 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:31 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:30 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:30 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:30 arlolra@deploy1003: arlolra: Continuing with deployment * 20:29 arlolra@deploy1003: arlolra: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:27 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] * 20:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2014.codfw.wmnet with OS bookworm * 20:21 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 20:21 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:20 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 20:19 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:19 swfrench-wmf: switched etcd-mirror replication from conf2005 to conf2004 - [[phab:T428495|T428495]] * 20:17 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:17 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:15 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] (duration: 08m 26s) * 20:12 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:11 arlolra@deploy1003: anzx, arlolra: Continuing with deployment * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2013.codfw.wmnet with OS bookworm * 20:09 arlolra@deploy1003: anzx, arlolra: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] * 20:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2216: Maintenance * 19:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2014.codfw.wmnet with reason: host reimage * 19:57 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:54 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:54 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:53 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:52 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2014.codfw.wmnet with reason: host reimage * 19:49 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2013.codfw.wmnet with reason: host reimage * 19:42 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 19:41 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:41 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2013.codfw.wmnet with reason: host reimage * 19:39 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:39 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:39 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-eqiad: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 19:38 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2014 * 19:33 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2014 * 19:29 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2014.codfw.wmnet with OS bookworm * 19:28 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:27 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1006 * 19:26 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2012\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 19:26 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1006 * 19:26 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:26 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 19:25 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2013 * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2013 * 19:21 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2013 * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2013.codfw.wmnet 84.0.192.10.in-addr.arpa 4.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:21 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2013.codfw.wmnet 84.0.192.10.in-addr.arpa 4.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2013 - bking@cumin2003" * 19:21 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2013 - bking@cumin2003" * 19:21 vriley@cumin1003: START - Cookbook sre.dns.netbox * 19:20 root@cumin1003: START - Cookbook sre.mysql.pool pool db2216: Maintenance * 19:13 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2216 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95435 and previous config saved to /var/cache/conftool/dbconfig/20260728-191343-cwilliams.json * 19:13 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2216.codfw.wmnet with reason: Maintenance * 19:13 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2203: Maintenance * 19:06 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1005.eqiad.wmnet with OS trixie * 19:06 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 19:06 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 18:46 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 18:45 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:45 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:43 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:40 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:36 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-eqiad: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 18:35 dancy@deploy1003: Installation of scap version "4.275.0" completed for 3 hosts * 18:33 dancy@deploy1003: Installing scap version "4.275.0" for 3 host(s) * 18:32 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:32 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2097.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:30 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns3003.wikimedia.org [reason: pool for all services after reimaging] * 18:29 sukhe@dns1004: END - running authdns-update * 18:27 sukhe@dns1004: START - running authdns-update * 18:27 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns3003.wikimedia.org,service=authdns-update [reason: pool authdns-update after reimaging] * 18:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db2203: Maintenance * 18:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2203 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95430 and previous config saved to /var/cache/conftool/dbconfig/20260728-181958-cwilliams.json * 18:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2203.codfw.wmnet with reason: Maintenance * 18:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2188: Maintenance * 18:18 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2097.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:17 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-be2098 * 18:17 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host ms-be2098 * 18:17 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-be2097 * 18:16 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host ms-be2097 * 18:15 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:15 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding ms-be2097-8 to codfw - jhancock@cumin2002" * 18:15 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding ms-be2097-8 to codfw - jhancock@cumin2002" * 18:10 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 18:08 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage * 18:05 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns3003.wikimedia.org with OS trixie * 18:03 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage * 17:56 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-codfw: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 17:45 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie * 17:45 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1005.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:41 sukhe@dns1004: END - running authdns-update * 17:39 sukhe@dns1004: START - running authdns-update * 17:36 sukhe@puppetserver1001: conftool action : set/weight=1; selector: cluster=urldownloader,service=squid * 17:36 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1005.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:35 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader,service=squid * 17:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 17:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1005 * 17:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 17:34 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1005 * 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1005] - vriley@cumin1003" * 17:34 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1005] - vriley@cumin1003" * 17:32 root@cumin1003: START - Cookbook sre.mysql.pool pool db2188: Maintenance * 17:29 vriley@cumin1003: START - Cookbook sre.dns.netbox * 17:29 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 17:26 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2188 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95425 and previous config saved to /var/cache/conftool/dbconfig/20260728-172609-cwilliams.json * 17:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2188.codfw.wmnet with reason: Maintenance * 17:25 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2176: Maintenance * 17:19 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1138.eqiad.wmnet with OS trixie * 17:18 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1005 * 17:18 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1005 * 17:18 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:15 vriley@cumin1003: START - Cookbook sre.dns.netbox * 17:13 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns3003.wikimedia.org with reason: host reimage * 17:07 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns3003.wikimedia.org with reason: host reimage * 17:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-eqiad * 17:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1015.eqiad.wmnet * 17:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1015.eqiad.wmnet * 16:59 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1138.eqiad.wmnet with reason: host reimage * 16:55 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-codfw: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 16:54 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1138.eqiad.wmnet with reason: host reimage * 16:53 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1015.eqiad.wmnet * 16:43 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns3003.wikimedia.org with OS trixie * 16:43 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1015.eqiad.wmnet * 16:43 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1014.eqiad.wmnet * 16:43 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1014.eqiad.wmnet * 16:43 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=dns3003.wikimedia.org [reason: depooling for reimage to trixie] * 16:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 290 hosts * 16:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db2176: Maintenance * 16:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1138 * 16:38 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1138 * 16:37 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1138 * 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1138.eqiad.wmnet 193.32.64.10.in-addr.arpa 3.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:37 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1138.eqiad.wmnet 193.32.64.10.in-addr.arpa 3.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1138 - jiji@cumin1003" * 16:37 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1138 - jiji@cumin1003" * 16:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1014.eqiad.wmnet * 16:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2176 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95420 and previous config saved to /var/cache/conftool/dbconfig/20260728-163235-cwilliams.json * 16:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2176.codfw.wmnet with reason: Maintenance * 16:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1014.eqiad.wmnet * 16:32 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1013.eqiad.wmnet * 16:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1013.eqiad.wmnet * 16:32 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2174: Maintenance * 16:28 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2012.codfw.wmnet, repooling source-only afterwards * 16:25 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1013.eqiad.wmnet * 16:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1013.eqiad.wmnet * 16:20 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1012.eqiad.wmnet * 16:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1012.eqiad.wmnet * 16:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1012.eqiad.wmnet * 16:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1012.eqiad.wmnet * 16:03 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1011.eqiad.wmnet * 16:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1011.eqiad.wmnet * 16:00 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2004.codfw.wmnet with OS bookworm * 15:59 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1011.eqiad.wmnet * 15:56 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:55 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 15:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:54 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1011.eqiad.wmnet * 15:54 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1010.eqiad.wmnet * 15:54 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1010.eqiad.wmnet * 15:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:50 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 15:49 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1010.eqiad.wmnet * 15:48 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 15:48 jiji@cumin1003: START - Cookbook sre.dns.netbox * 15:46 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2248: Maintenance * 15:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db2174: Maintenance * 15:44 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1010.eqiad.wmnet * 15:44 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1009.eqiad.wmnet * 15:44 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1009.eqiad.wmnet * 15:42 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1138 * 15:41 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1138.eqiad.wmnet with OS trixie * 15:39 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1009.eqiad.wmnet * 15:39 robh@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on arclamp2001.codfw.wmnet with reason: ram upgrade * 15:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2174 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95413 and previous config saved to /var/cache/conftool/dbconfig/20260728-153844-cwilliams.json * 15:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2174.codfw.wmnet with reason: Maintenance * 15:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2173: Maintenance * 15:37 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 15:35 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 15:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1009.eqiad.wmnet * 15:34 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1008.eqiad.wmnet * 15:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1008.eqiad.wmnet * 15:31 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 15:31 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 15:29 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1008.eqiad.wmnet * 15:27 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 290 hosts * 15:25 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2013 * 15:25 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2195: Maintenance * 15:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1008.eqiad.wmnet * 15:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1007.eqiad.wmnet * 15:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1007.eqiad.wmnet * 15:22 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2004.codfw.wmnet with reason: host reimage * 15:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2013.codfw.wmnet with OS bookworm * 15:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 15:19 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2004.codfw.wmnet with reason: host reimage * 15:19 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 15:17 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1007.eqiad.wmnet * 15:12 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1007.eqiad.wmnet * 15:12 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1006.eqiad.wmnet * 15:12 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1006.eqiad.wmnet * 15:11 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1138.eqiad.wmnet * 15:11 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 15:11 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1138.eqiad.wmnet * 15:11 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1138.eqiad.wmnet * 15:10 brennen@deploy1003: Finished deploy [phabricator/deployment@f8b349f]: deploy phab1004 for [[phab:T433382|T433382]] (duration: 00m 43s) * 15:10 brennen@deploy1003: Started deploy [phabricator/deployment@f8b349f]: deploy phab1004 for [[phab:T433382|T433382]] * 15:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts * 15:09 brennen@deploy1003: Finished deploy [phabricator/deployment@f8b349f]: deploy phab2003 for [[phab:T433382|T433382]] (duration: 00m 55s) * 15:08 brennen@deploy1003: Started deploy [phabricator/deployment@f8b349f]: deploy phab2003 for [[phab:T433382|T433382]] * 15:07 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2012.codfw.wmnet, repooling source-only afterwards * 15:07 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts * 15:06 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1137.eqiad.wmnet * 15:06 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1137.eqiad.wmnet * 15:06 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1137.eqiad.wmnet * 15:05 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1006.eqiad.wmnet * 15:05 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1022\.eqiad\.wmnet,dc=eqiad,cluster=wdqs\-main,service=wdqs\-main * 15:01 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab2003.codfw.wmnet with reason: deployment * 15:01 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1005.eqiad.wmnet with reason: deployment * 15:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1006.eqiad.wmnet * 15:00 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1006.eqiad.wmnet with reason: deployment * 15:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1005.eqiad.wmnet * 15:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1005.eqiad.wmnet * 14:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db2248: Maintenance * 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1004.eqiad.wmnet with reason: deployment * 14:59 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2004.codfw.wmnet with OS bookworm * 14:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1005.eqiad.wmnet * 14:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2248 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95403 and previous config saved to /var/cache/conftool/dbconfig/20260728-145532-cwilliams.json * 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2245-2247].codfw.wmnet with reason: Maintenance * 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2248.codfw.wmnet with reason: Maintenance * 14:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2240: Maintenance * 14:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2173: Maintenance * 14:51 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts * 14:50 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1005.eqiad.wmnet * 14:50 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1004.eqiad.wmnet * 14:50 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1004.eqiad.wmnet * 14:49 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts * 14:45 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2004.codfw.wmnet * 14:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2173 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95399 and previous config saved to /var/cache/conftool/dbconfig/20260728-144453-cwilliams.json * 14:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2173.codfw.wmnet with reason: Maintenance * 14:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2170: Maintenance * 14:44 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1004.eqiad.wmnet * 14:39 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2004.codfw.wmnet * 14:38 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1004.eqiad.wmnet * 14:38 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet * 14:38 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet * 14:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db2195: Maintenance * 14:36 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1022.eqiad.wmnet, repooling source-only afterwards * 14:36 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 14:33 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet * 14:33 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2222: Maintenance * 14:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2195 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95394 and previous config saved to /var/cache/conftool/dbconfig/20260728-143218-cwilliams.json * 14:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2195.codfw.wmnet with reason: Maintenance * 14:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2181: Maintenance * 14:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 14:30 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 14:25 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 14:25 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 14:25 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 14:23 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet * 14:23 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1002.eqiad.wmnet * 14:23 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1002.eqiad.wmnet * 14:23 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 19s) * 14:23 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:18 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1002.eqiad.wmnet * 14:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2012.codfw.wmnet with OS bookworm * 14:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1002.eqiad.wmnet * 14:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1001.eqiad.wmnet * 14:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1001.eqiad.wmnet * 14:11 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 14:11 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 14:08 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1001.eqiad.wmnet * 14:07 elukey@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'. * 14:07 elukey@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'. * 14:06 elukey@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'. * 14:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db2240: Maintenance * 14:06 XioNoX: un-drain cr2-esams - [[phab:T431751|T431751]] * 14:05 elukey@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'. * 14:02 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1001.eqiad.wmnet * 14:02 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-eqiad * 14:01 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 14:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2240 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95384 and previous config saved to /var/cache/conftool/dbconfig/20260728-140011-cwilliams.json * 14:00 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2240.codfw.wmnet with reason: Maintenance * 13:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2237: Maintenance * 13:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db2170: Maintenance * 13:56 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 13:55 XioNoX: reboot cr2-esams - [[phab:T431751|T431751]] * 13:52 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 13:51 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr2-esams,cr2-esams IPv6,cr2-esams.mgmt with reason: router upgrade * 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 13:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2170 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95381 and previous config saved to /var/cache/conftool/dbconfig/20260728-135043-cwilliams.json * 13:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2170.codfw.wmnet with reason: Maintenance * 13:50 XioNoX: drain cr2-esams - [[phab:T431751|T431751]] * 13:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2153: Maintenance * 13:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2012.codfw.wmnet with reason: host reimage * 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2222: Maintenance * 13:45 sukhe: restart pybal on A:lvs-codfw * 13:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2012.codfw.wmnet with reason: host reimage * 13:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db2181: Maintenance * 13:44 btullis@dns1004: END - running authdns-update * 13:42 sukhe: restart pybal on lvs2014 * 13:42 btullis@dns1004: START - running authdns-update * 13:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2222 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95376 and previous config saved to /var/cache/conftool/dbconfig/20260728-133948-cwilliams.json * 13:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2222.codfw.wmnet with reason: Maintenance * 13:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2221: Maintenance * 13:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2181 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95374 and previous config saved to /var/cache/conftool/dbconfig/20260728-133857-cwilliams.json * 13:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2181.codfw.wmnet with reason: Maintenance * 13:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2167: Maintenance * 13:30 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T428495|T428495]] * 13:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 13:29 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 13:29 ayounsi@cumin1003: END (FAIL) - Cookbook sre.dns.admin (exit_code=99) DNS admin: depool esams [reason: router upgrade, [[phab:T431749|T431749]]] * 13:28 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: router upgrade, [[phab:T431749|T431749]]] * 13:27 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1022.eqiad.wmnet, repooling source-only afterwards * 13:27 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo - [[phab:T428495|T428495]] * 13:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2012 * 13:27 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2012 * 13:21 lucaswerkmeister-wmde@deploy1003: mwscript-k8s job started: cleanupTitles bolwiki # [[phab:T429951|T429951]] * 13:21 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] (duration: 07m 19s) * 13:20 swfrench-wmf: authdns-update to direct codfw, eqsin, ulsfo etcd clients to eqiad - [[phab:T428495|T428495]] * 13:18 swfrench@dns1004: END - running authdns-update * 13:17 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, anzx: Continuing with deployment * 13:16 swfrench@dns1004: START - running authdns-update * 13:16 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2012 * 13:16 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2012.codfw.wmnet 57.48.192.10.in-addr.arpa 7.5.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:16 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2012.codfw.wmnet 57.48.192.10.in-addr.arpa 7.5.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:16 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:16 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, anzx: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:14 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 13:14 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] * 13:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db2237: Maintenance * 13:13 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2237: Maintenance * 13:13 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:12 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:12 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback IPV6 for asw1-604 - pt1979@cumin2003" * 13:12 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:12 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback IPV6 for asw1-604 - pt1979@cumin2003" * 13:11 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 20s) * 13:11 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 13:10 esanders@deploy1003: Finished scap sync-world: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] (duration: 08m 11s) * 13:08 pt1979@cumin2003: START - Cookbook sre.dns.netbox * 13:07 root@cumin1003: START - Cookbook sre.mysql.pool pool db2237: Maintenance * 13:06 esanders@deploy1003: esanders: Continuing with deployment * 13:04 esanders@deploy1003: esanders: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2153: Maintenance * 13:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2153: Maintenance * 13:02 esanders@deploy1003: Started scap sync-world: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] * 13:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2237 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95362 and previous config saved to /var/cache/conftool/dbconfig/20260728-130107-cwilliams.json * 13:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2237.codfw.wmnet with reason: Maintenance * 13:00 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2236: Maintenance * 12:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2153: Maintenance * 12:52 root@cumin1003: START - Cookbook sre.mysql.pool pool db2221: Maintenance * 12:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2153 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95358 and previous config saved to /var/cache/conftool/dbconfig/20260728-125214-cwilliams.json * 12:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2153.codfw.wmnet with reason: Maintenance * 12:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2167: Maintenance * 12:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-codfw * 12:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2011.codfw.wmnet * 12:51 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2011.codfw.wmnet * 12:49 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:49 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback for asw1-603 - pt1979@cumin2003" * 12:48 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback for asw1-603 - pt1979@cumin2003" * 12:46 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2011.codfw.wmnet * 12:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2221 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95357 and previous config saved to /var/cache/conftool/dbconfig/20260728-124601-cwilliams.json * 12:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2221.codfw.wmnet with reason: Maintenance * 12:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2218: Maintenance * 12:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2167 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95354 and previous config saved to /var/cache/conftool/dbconfig/20260728-124457-cwilliams.json * 12:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2167.codfw.wmnet with reason: Maintenance * 12:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2166: Maintenance * 12:42 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 12:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2011.codfw.wmnet * 12:41 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2010.codfw.wmnet * 12:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2010.codfw.wmnet * 12:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2010.codfw.wmnet * 12:34 pt1979@cumin2003: START - Cookbook sre.dns.netbox * 12:32 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 12:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2010.codfw.wmnet * 12:31 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2009.codfw.wmnet * 12:31 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2009.codfw.wmnet * 12:27 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2009.codfw.wmnet * 12:22 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2009.codfw.wmnet * 12:22 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2008.codfw.wmnet * 12:21 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2008.codfw.wmnet * 12:16 pt1979@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-604-eqsin * 12:16 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2008.codfw.wmnet * 12:16 pt1979@cumin1003: START - Cookbook sre.network.tls for network device asw1-604-eqsin * 12:14 root@cumin1003: START - Cookbook sre.mysql.pool pool db2236: Maintenance * 12:14 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2236: Maintenance * 12:12 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1137.eqiad.wmnet with OS trixie * 12:11 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2008.codfw.wmnet * 12:11 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2007.codfw.wmnet * 12:11 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2007.codfw.wmnet * 12:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db2236: Maintenance * 12:06 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2007.codfw.wmnet * 12:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2236 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95348 and previous config saved to /var/cache/conftool/dbconfig/20260728-120253-cwilliams.json * 12:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2236.codfw.wmnet with reason: Maintenance * 12:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2007.codfw.wmnet * 12:01 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 12:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 11:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2218: Maintenance * 11:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db2166: Maintenance * 11:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2219: Maintenance * 11:56 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2006.codfw.wmnet * 11:52 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1137.eqiad.wmnet with reason: host reimage * 11:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2218 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95344 and previous config saved to /var/cache/conftool/dbconfig/20260728-115155-cwilliams.json * 11:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2218.codfw.wmnet with reason: Maintenance * 11:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2208: Maintenance * 11:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2166 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95342 and previous config saved to /var/cache/conftool/dbconfig/20260728-115119-cwilliams.json * 11:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2166.codfw.wmnet with reason: Maintenance * 11:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2164: Maintenance * 11:47 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1137.eqiad.wmnet with reason: host reimage * 11:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2006.codfw.wmnet * 11:45 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2005.codfw.wmnet * 11:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2005.codfw.wmnet * 11:40 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2005.codfw.wmnet * 11:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2005.codfw.wmnet * 11:35 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2004.codfw.wmnet * 11:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2004.codfw.wmnet * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1137 * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1137 * 11:30 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1137 * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1137.eqiad.wmnet 192.32.64.10.in-addr.arpa 2.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:30 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1137.eqiad.wmnet 192.32.64.10.in-addr.arpa 2.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1137 - jiji@cumin1003" * 11:25 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2004.codfw.wmnet * 11:19 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2004.codfw.wmnet * 11:19 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2003.codfw.wmnet * 11:19 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2003.codfw.wmnet * 11:14 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2003.codfw.wmnet * 11:11 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2219: Maintenance * 11:10 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2219: Maintenance * 11:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2219: Maintenance * 11:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2003.codfw.wmnet * 11:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2164: Maintenance * 11:03 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2002.codfw.wmnet * 11:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2002.codfw.wmnet * 11:03 root@cumin1003: START - Cookbook sre.mysql.pool pool db2208: Maintenance * 10:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2164 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95332 and previous config saved to /var/cache/conftool/dbconfig/20260728-105749-cwilliams.json * 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2164.codfw.wmnet with reason: Maintenance * 10:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2208 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95331 and previous config saved to /var/cache/conftool/dbconfig/20260728-105711-cwilliams.json * 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2208.codfw.wmnet with reason: Maintenance * 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2219 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95330 and previous config saved to /var/cache/conftool/dbconfig/20260728-105652-cwilliams.json * 10:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2219.codfw.wmnet with reason: Maintenance * 10:53 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1137 - jiji@cumin1003" * 10:52 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2002.codfw.wmnet * 10:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2002.codfw.wmnet * 10:47 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2001.codfw.wmnet * 10:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2001.codfw.wmnet * 10:39 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2001.codfw.wmnet * 10:35 jiji@cumin1003: START - Cookbook sre.dns.netbox * 10:34 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] (duration: 09m 31s) * 10:34 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1137 * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2001.codfw.wmnet * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-codfw * 10:34 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1137.eqiad.wmnet with OS trixie * 10:28 jforrester@deploy1003: jforrester: Continuing with deployment * 10:27 jforrester@deploy1003: jforrester: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:25 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] * 10:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 10:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 10:21 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 10:20 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1137.eqiad.wmnet * 10:20 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 10:20 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1137.eqiad.wmnet * 10:20 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1137.eqiad.wmnet * 10:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2163: Maintenance * 09:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-staging-worker * 09:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2003.codfw.wmnet * 09:37 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2003.codfw.wmnet * 09:32 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 09:31 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2003.codfw.wmnet * 09:30 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 09:30 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2163: Maintenance * 09:30 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 09:30 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 09:30 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:22 klausman@cumin1003: END (ERROR) - Cookbook sre.ganeti.reboot-vm (exit_code=97) for VM ml-serve-ctrl2001.codfw.wmnet * 09:22 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2001.codfw.wmnet * 09:22 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d8-eqiad * 09:22 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d8-eqiad * 09:21 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2003.codfw.wmnet * 09:20 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2002.codfw.wmnet * 09:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2002.codfw.wmnet * 09:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f2-codfw * 09:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f2-codfw * 09:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e4-codfw * 09:18 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2163: Maintenance * 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e4-codfw * 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-codfw * 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-codfw * 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e5-codfw * 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e5-codfw * 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f4-codfw * 09:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2210: Maintenance * 09:16 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f4-codfw * 09:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2182: Maintenance * 09:14 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2002.codfw.wmnet * 09:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db2163: Maintenance * 09:11 XioNoX: rebooting cr2-drmrs - [[phab:T431749|T431749]] * 09:10 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr2-drmrs,cr2-drmrs IPv6,cr2-drmrs.mgmt with reason: router upgrade * 09:06 XioNoX: draining cr2-drmrs - [[phab:T431749|T431749]] * 09:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2163 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95320 and previous config saved to /var/cache/conftool/dbconfig/20260728-090638-cwilliams.json * 09:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2163.codfw.wmnet with reason: Maintenance * 09:06 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2161: Maintenance * 09:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2002.codfw.wmnet * 09:04 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2001.codfw.wmnet * 09:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2001.codfw.wmnet * 08:57 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2001.codfw.wmnet * 08:48 XioNoX: un-drain cr1-drmrs - [[phab:T431749|T431749]] * 08:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2001.codfw.wmnet * 08:47 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-staging-worker * 08:42 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:35 XioNoX: rebooting cr1-drmrs - [[phab:T431749|T431749]] * 08:33 XioNoX: draining cr1-drmrs - [[phab:T431749|T431749]] * 08:31 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2210: Maintenance * 08:29 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2182: Maintenance * 08:21 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2182: Maintenance * 08:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db2161: Maintenance * 08:16 root@cumin1003: START - Cookbook sre.mysql.pool pool db2182: Maintenance * 08:12 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2210: Maintenance * 08:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2161 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95309 and previous config saved to /var/cache/conftool/dbconfig/20260728-081044-cwilliams.json * 08:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2161.codfw.wmnet with reason: Maintenance * 08:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2154: Maintenance * 08:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2182 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95307 and previous config saved to /var/cache/conftool/dbconfig/20260728-080947-cwilliams.json * 08:09 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2182.codfw.wmnet with reason: Maintenance * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2168: Maintenance * 08:06 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr1-drmrs,cr1-drmrs IPv6,cr1-drmrs.mgmt with reason: router upgrade * 08:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db2210: Maintenance * 08:05 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 08:05 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 08:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2210 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95305 and previous config saved to /var/cache/conftool/dbconfig/20260728-080008-cwilliams.json * 08:00 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2210.codfw.wmnet with reason: Maintenance * 07:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2206: Maintenance * 07:50 gkyziridis@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 07:50 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 07:22 root@cumin1003: START - Cookbook sre.mysql.pool pool db2154: Maintenance * 07:22 root@cumin1003: START - Cookbook sre.mysql.pool pool db2168: Maintenance * 07:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2154 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95295 and previous config saved to /var/cache/conftool/dbconfig/20260728-071640-cwilliams.json * 07:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2154.codfw.wmnet with reason: Maintenance * 07:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95294 and previous config saved to /var/cache/conftool/dbconfig/20260728-071604-cwilliams.json * 07:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2168.codfw.wmnet with reason: Maintenance * 07:08 root@cumin1003: START - Cookbook sre.mysql.pool pool db2206: Maintenance * 07:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2206 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95292 and previous config saved to /var/cache/conftool/dbconfig/20260728-070219-cwilliams.json * 07:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2206.codfw.wmnet with reason: Maintenance * 06:44 marostegui: Failover m5 from db1164 to db1228 - [[phab:T432967|T432967]] * 06:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2235].codfw.wmnet,db[1164,1217,1228].eqiad.wmnet with reason: m5 master switch [[phab:T432967|T432967]] * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.10 (duration: 02m 34s) * 03:39 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] (duration: 36m 06s) * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 02:57 dzahn@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1004.eqiad.wmnet with OS trixie * 02:57 dzahn@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - dzahn@cumin1003" * 02:55 dzahn@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - dzahn@cumin1003" * 02:37 dzahn@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1004.eqiad.wmnet with reason: host reimage * 02:31 dzahn@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1004.eqiad.wmnet with reason: host reimage * 02:16 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie * 02:15 dzahn@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host zuul1004.eqiad.wmnet with OS trixie * 01:43 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie * 01:43 dzahn@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1004.eqiad.wmnet with OS trixie * 01:25 pt1979@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-603-eqsin * 01:24 pt1979@cumin1003: START - Cookbook sre.network.tls for network device asw1-603-eqsin * 01:12 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 01:12 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt for new switches in eqsin - pt1979@cumin2003" * 01:12 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt for new switches in eqsin - pt1979@cumin2003" * 01:08 pt1979@cumin2003: START - Cookbook sre.dns.netbox * 00:48 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 00:47 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 00:47 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 00:47 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 00:26 mutante: attempting reimage with trixie on zuul1004 re-purposed physical hardware - dcops reported install issue - host was in busybox shell ([[phab:T427353|T427353]]) * 00:24 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie == 2026-07-27 == * 23:50 Amir1: mass deleting vp8 transcodes * 23:28 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:27 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1004.eqiad.wmnet with OS bullseye * 23:26 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 23:25 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:25 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 22:39 maryum: Deploy security fix for [[phab:T432877|T432877]] * 22:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1022.eqiad.wmnet with OS bookworm * 22:37 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS bullseye * 22:32 sbassett: Deployed security fix for [[phab:T432789|T432789]] * 22:22 sbassett: Deployed security patch for [[phab:T431819|T431819]] * 22:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1022.eqiad.wmnet with reason: host reimage * 22:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1022.eqiad.wmnet with reason: host reimage * 22:01 RScout-WMF: Deployed security fix for [[phab:T431819|T431819]] * 22:00 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2012 * 21:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2012.codfw.wmnet with OS bookworm * 21:55 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2011\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 21:45 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1022 * 21:45 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1022 * 21:44 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1022 * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1022.eqiad.wmnet 239.48.64.10.in-addr.arpa 9.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:44 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1022.eqiad.wmnet 239.48.64.10.in-addr.arpa 9.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:41 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:41 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 21:34 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS bookworm * 21:31 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:22 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:21 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:19 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:17 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1004 * 21:16 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1004 * 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1004] - vriley@cumin1003" * 21:15 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1004] - vriley@cumin1003" * 21:11 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:10 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2011.codfw.wmnet, repooling source-only afterwards * 21:05 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:01 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1022 * 20:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1022.eqiad.wmnet with OS bookworm * 20:53 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1021.eqiad.wmnet, repooling source-only afterwards * 20:51 mutante: zuul1001 - re-enabled puppet - revert "cherry-picked" gerrit:1314120 - [[phab:T431003|T431003]] * 20:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Maintenance * 20:15 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] (duration: 08m 03s) * 20:11 sbisson@deploy1003: sbisson: Continuing with deployment * 20:09 sbisson@deploy1003: sbisson: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] * 19:47 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 46s) * 19:47 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Maintenance * 19:27 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] (duration: 12m 26s) * 19:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2228 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95285 and previous config saved to /var/cache/conftool/dbconfig/20260727-192711-cwilliams.json * 19:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2228.codfw.wmnet with reason: Maintenance * 19:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2223: Maintenance * 19:23 krinkle@deploy1003: krinkle: Continuing with deployment * 19:16 krinkle@deploy1003: krinkle: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:15 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] * 19:12 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2238: Maintenance * 18:58 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 18:57 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 18:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2227: Maintenance * 18:57 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-experimental: apply * 18:55 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-experimental: apply * 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1021.eqiad.wmnet with OS bookworm * 18:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2011.codfw.wmnet with OS bookworm * 18:40 root@cumin1003: START - Cookbook sre.mysql.pool pool db2223: Maintenance * 18:39 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] (duration: 07m 05s) * 18:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2223 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95275 and previous config saved to /var/cache/conftool/dbconfig/20260727-183500-cwilliams.json * 18:34 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2223.codfw.wmnet with reason: Maintenance * 18:34 musikanimal@deploy1003: musikanimal: Continuing with deployment * 18:34 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2213: Maintenance * 18:33 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:32 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] * 18:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db2238: Maintenance * 18:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2011.codfw.wmnet with reason: host reimage * 18:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2238 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95271 and previous config saved to /var/cache/conftool/dbconfig/20260727-181944-cwilliams.json * 18:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2238.codfw.wmnet with reason: Maintenance * 18:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2226: Maintenance * 18:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1021.eqiad.wmnet with reason: host reimage * 18:14 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2011.codfw.wmnet with reason: host reimage * 18:12 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1021.eqiad.wmnet with reason: host reimage * 18:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db2227: Maintenance * 18:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2227 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95265 and previous config saved to /var/cache/conftool/dbconfig/20260727-180256-cwilliams.json * 18:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2227.codfw.wmnet with reason: Maintenance * 18:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2194: Maintenance * 17:57 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2011 * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2011 * 17:56 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2011 * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2011.codfw.wmnet 37.32.192.10.in-addr.arpa 7.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:56 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2011.codfw.wmnet 37.32.192.10.in-addr.arpa 7.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2011 - bking@cumin2003" * 17:56 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2011 - bking@cumin2003" * 17:52 bking@cumin2003: START - Cookbook sre.dns.netbox * 17:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2011 * 17:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1021 * 17:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1021 * 17:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2011.codfw.wmnet with OS bookworm * 17:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1021.eqiad.wmnet with OS bookworm * 17:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Maintenance * 17:38 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2010\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 17:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2213 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95260 and previous config saved to /var/cache/conftool/dbconfig/20260727-173740-cwilliams.json * 17:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2213.codfw.wmnet with reason: Maintenance * 17:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2211: Maintenance * 17:36 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1020\.eqiad\.wmnet,dc=eqiad,cluster=wdqs\-main,service=wdqs\-main * 17:32 root@cumin1003: START - Cookbook sre.mysql.pool pool db2226: Maintenance * 17:31 taavi@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] (duration: 06m 33s) * 17:27 taavi@deploy1003: taavi: Continuing with deployment * 17:27 taavi@deploy1003: taavi: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:26 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2226 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95256 and previous config saved to /var/cache/conftool/dbconfig/20260727-172636-cwilliams.json * 17:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2226.codfw.wmnet with reason: Maintenance * 17:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2225: Maintenance * 17:25 taavi@deploy1003: Started scap sync-world: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] * 17:13 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 17:11 root@cumin1003: START - Cookbook sre.mysql.pool pool db2194: Maintenance * 17:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2194 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95248 and previous config saved to /var/cache/conftool/dbconfig/20260727-170453-cwilliams.json * 17:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2194.codfw.wmnet with reason: Maintenance * 17:04 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2190: Maintenance * 16:52 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 16:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2211: Maintenance * 16:40 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95242 and previous config saved to /var/cache/conftool/dbconfig/20260727-164015-cwilliams.json * 16:40 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2211.codfw.wmnet with reason: Maintenance * 16:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2178: Maintenance * 16:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db2225: Maintenance * 16:39 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2172: Maintenance * 16:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 16:38 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 16:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2225 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95238 and previous config saved to /var/cache/conftool/dbconfig/20260727-163307-cwilliams.json * 16:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2225.codfw.wmnet with reason: Maintenance * 16:32 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2189: Maintenance * 16:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db2190: Maintenance * 16:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2190 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95230 and previous config saved to /var/cache/conftool/dbconfig/20260727-160602-cwilliams.json * 16:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2190.codfw.wmnet with reason: Maintenance * 15:53 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2177: Maintenance * 15:53 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2172: Maintenance * 15:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2178: Maintenance * 15:51 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2172: Maintenance * 15:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2178 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95224 and previous config saved to /var/cache/conftool/dbconfig/20260727-154559-cwilliams.json * 15:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2172: Maintenance * 15:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2178.codfw.wmnet with reason: Maintenance * 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2171: Maintenance * 15:44 root@cumin1003: START - Cookbook sre.mysql.pool pool db2189: Maintenance * 15:43 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:41 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2172 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95222 and previous config saved to /var/cache/conftool/dbconfig/20260727-153927-cwilliams.json * 15:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2172.codfw.wmnet with reason: Maintenance * 15:38 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2189 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95220 and previous config saved to /var/cache/conftool/dbconfig/20260727-153833-cwilliams.json * 15:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2189.codfw.wmnet with reason: Maintenance * 15:34 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:32 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 15:32 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 15:31 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:29 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:26 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 15:22 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] (duration: 07m 00s) * 15:21 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2155: Maintenance * 15:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2175: Maintenance * 15:18 zabe@deploy1003: zabe: Continuing with deployment * 15:17 zabe@deploy1003: zabe: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:15 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 15:15 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:15 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] * 15:15 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 06s) * 15:15 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:12 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2010.codfw.wmnet with OS bookworm * 15:08 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2177: Maintenance * 15:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2177: Maintenance * 14:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2177: Maintenance * 14:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2171: Maintenance * 14:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2171 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95209 and previous config saved to /var/cache/conftool/dbconfig/20260727-145236-cwilliams.json * 14:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2171.codfw.wmnet with reason: Maintenance * 14:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2177 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95208 and previous config saved to /var/cache/conftool/dbconfig/20260727-145206-cwilliams.json * 14:52 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2157: Maintenance * 14:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2177.codfw.wmnet with reason: Maintenance * 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1020.eqiad.wmnet with OS bookworm * 14:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2156: Maintenance * 14:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2010.codfw.wmnet with reason: host reimage * 14:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2010.codfw.wmnet with reason: host reimage * 14:41 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[1031,2024]*: Upgrade Cassandra to 5.0.8 (canary) - eevans@cumin1003 * 14:34 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2155: Maintenance * 14:33 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2175: Maintenance * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2010 * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2010 * 14:24 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2010 * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2010.codfw.wmnet 94.16.192.10.in-addr.arpa 4.9.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2010.codfw.wmnet 94.16.192.10.in-addr.arpa 4.9.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2010 - bking@cumin2003" * 14:24 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2010 - bking@cumin2003" * 14:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1020.eqiad.wmnet with reason: host reimage * 14:23 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[1031,2024]*: Upgrade Cassandra to 5.0.8 (canary) - eevans@cumin1003 * 14:20 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 14:20 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 14:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1020.eqiad.wmnet with reason: host reimage * 14:17 sukhe: sudo gnt-instance reboot urldownloader1005.wikimedia.org * 14:16 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:15 jelto@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:08 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2155: Maintenance * 14:05 root@cumin1003: START - Cookbook sre.mysql.pool pool db2157: Maintenance * 14:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2175: Maintenance * 14:03 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 11 hosts * 14:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db2155: Maintenance * 14:01 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 11 hosts * 14:01 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1136.eqiad.wmnet * 14:01 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1136.eqiad.wmnet * 14:01 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1136.eqiad.wmnet * 14:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db2156: Maintenance * 13:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2157 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95194 and previous config saved to /var/cache/conftool/dbconfig/20260727-135943-cwilliams.json * 13:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2157.codfw.wmnet with reason: Maintenance * 13:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db2175: Maintenance * 13:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 13:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 13:57 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2010 * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2155 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95193 and previous config saved to /var/cache/conftool/dbconfig/20260727-135613-cwilliams.json * 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2155.codfw.wmnet with reason: Maintenance * 13:55 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1020 * 13:55 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1020 * 13:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2156 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95192 and previous config saved to /var/cache/conftool/dbconfig/20260727-135413-cwilliams.json * 13:54 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2156.codfw.wmnet with reason: Maintenance * 13:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2175 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95191 and previous config saved to /var/cache/conftool/dbconfig/20260727-135300-cwilliams.json * 13:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2010.codfw.wmnet with OS bookworm * 13:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2175.codfw.wmnet with reason: Maintenance * 13:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1020.eqiad.wmnet with OS bookworm * 13:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 34 hosts * 13:46 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 34 hosts * 13:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1201: Maintenance * 13:27 Lucas_WMDE: UTC afternoon backport+config window doen * 13:18 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] (duration: 11m 57s) * 13:14 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, sihe: Continuing with deployment * 13:08 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, sihe: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:07 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool ulsfo [reason: router upgrade finished, [[phab:T431752|T431752]]] * 13:07 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool ulsfo [reason: router upgrade finished, [[phab:T431752|T431752]]] * 13:06 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] * 13:03 XioNoX: repool cr4-ulsfo - [[phab:T431752|T431752]] * 12:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db1201: Maintenance * 12:48 gkyziridis@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1201 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95186 and previous config saved to /var/cache/conftool/dbconfig/20260727-124404-cwilliams.json * 12:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1201.eqiad.wmnet with reason: Maintenance * 12:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1187: Maintenance * 12:30 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader.eqiad.wikimedia.org on all recursors * 12:30 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader.eqiad.wikimedia.org on all recursors * 12:30 sukhe@dns1004: END - running authdns-update * 12:30 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] (duration: 09m 32s) * 12:28 sukhe@dns1004: START - running authdns-update * 12:25 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 12:22 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:20 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] * 12:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts * 12:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts * 12:13 XioNoX: rebooting cr4-ulsfo for upgrade - [[phab:T431752|T431752]] * 12:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: es1038 repool * 12:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 38 hosts * 12:08 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 38 hosts * 11:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1187: Maintenance * 11:53 urbanecm@deploy1003: mwscript-k8s job started: foreachwikiindblist growthexperiments GrowthExperiments:cleanMentorList # [[phab:T431804|T431804]] * 11:50 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr4-ulsfo,cr4-ulsfo IPv6,cr4-ulsfo.mgmt with reason: router upgrade * 11:50 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] (duration: 11m 07s) * 11:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1187 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95178 and previous config saved to /var/cache/conftool/dbconfig/20260727-114844-cwilliams.json * 11:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1187.eqiad.wmnet with reason: Maintenance * 11:43 urbanecm@deploy1003: urbanecm: Continuing with deployment * 11:42 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:39 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] * 11:37 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 11:36 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 11:36 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 11:35 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 11:29 XioNoX: start draining cr4-ulsfo - [[phab:T431752|T431752]] * 11:29 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 11:29 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 11:28 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1035: testing * 11:28 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1035: testing * 11:27 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1035: testing * 11:27 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1035: testing * 11:26 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool ulsfo [reason: router upgrade, [[phab:T431752|T431752]]] * 11:26 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1038: es1038 repool * 11:26 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool ulsfo [reason: router upgrade, [[phab:T431752|T431752]]] * 11:26 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1038: testing * 11:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1264: Maintenance * 11:24 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1038: testing * 11:23 marostegui@cumin1003: dbctl commit (dc=all): 'Repool es1050 as master', diff saved to https://phabricator.wikimedia.org/P95170 and previous config saved to /var/cache/conftool/dbconfig/20260727-112326-marostegui.json * 11:23 marostegui@cumin1003: dbctl commit (dc=all): 'Repool es1050', diff saved to https://phabricator.wikimedia.org/P95169 and previous config saved to /var/cache/conftool/dbconfig/20260727-112302-marostegui.json * 11:22 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1050: testing * 11:22 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1050: testing * 11:20 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 11:18 blake@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 11:18 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 11:12 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 11:11 blake@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 11:09 blake@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 11:09 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 11:09 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 11:08 blake@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 11:05 blake@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 11:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 11:02 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 10:50 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 10:43 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:39 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply * 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1264: Maintenance * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply * 10:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply * 10:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 10:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 10:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 10:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1264 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95164 and previous config saved to /var/cache/conftool/dbconfig/20260727-103204-cwilliams.json * 10:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1264.eqiad.wmnet with reason: Maintenance * 10:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 10:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 10:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 10:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 10:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 10:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 10:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 10:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1237: Maintenance * 10:04 elukey: restart burrow main-eqiad on kafkamon2003 to clear some errors on kafka-main1008 * 09:58 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1136.eqiad.wmnet with OS trixie * 09:39 elukey: restart burrow-main-eqiad.service on kafkamon1003 to see if a recurrent kafka error on kafka-main1008 goes away * 09:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1237: Maintenance * 09:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1136.eqiad.wmnet with reason: host reimage * 09:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1237 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95159 and previous config saved to /var/cache/conftool/dbconfig/20260727-093328-cwilliams.json * 09:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1237.eqiad.wmnet with reason: Maintenance * 09:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1136.eqiad.wmnet with reason: host reimage * 09:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1203: Maintenance * 09:17 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1136 * 09:17 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1136 * 09:04 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1136 * 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1136.eqiad.wmnet 191.32.64.10.in-addr.arpa 1.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:04 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1136.eqiad.wmnet 191.32.64.10.in-addr.arpa 1.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1136 - jiji@cumin1003" * 09:04 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1136 - jiji@cumin1003" * 08:52 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 08:52 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 08:52 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 08:51 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 08:50 jiji@cumin1003: START - Cookbook sre.dns.netbox * 08:47 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1136 * 08:46 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1136.eqiad.wmnet with OS trixie * 08:44 marostegui: Rename tables on s3 [[phab:T425066|T425066]] * 08:43 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1136.eqiad.wmnet * 08:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db1203: Maintenance * 08:43 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1136.eqiad.wmnet * 08:43 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1136.eqiad.wmnet * 08:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1203 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95154 and previous config saved to /var/cache/conftool/dbconfig/20260727-083703-cwilliams.json * 08:36 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1203.eqiad.wmnet with reason: Maintenance * 08:16 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1179: Maintenance * 07:44 phuedx: UTC morning backport window done * 07:37 phuedx@deploy1003: Finished scap sync-world: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] (duration: 32m 33s) * 07:28 root@cumin1003: START - Cookbook sre.mysql.pool pool db1179: Maintenance * 07:26 marostegui: Rename tables on s3 [[phab:T426341|T426341]] * 07:25 phuedx@deploy1003: phuedx: Continuing with deployment * 07:22 marostegui: Drop tables in akwiki nawiki pihwiki - growthexperiments_* [[phab:T428885|T428885]] * 07:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95149 and previous config saved to /var/cache/conftool/dbconfig/20260727-072234-cwilliams.json * 07:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1179.eqiad.wmnet with reason: Maintenance * 07:20 phuedx@deploy1003: phuedx: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:16 ryankemper: [[phab:T430880|T430880]] [WDQS] Reimaged `wdqs1018` and `wdqs1019` to Bookworm, restored data using test-cookbook change {{Gerrit|1317128}}, and repooled both; 25/36 hosts complete * 07:04 phuedx@deploy1003: Started scap sync-world: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] * 06:57 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1019.eqiad.wmnet * 06:56 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1018.eqiad.wmnet * 06:40 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1020.eqiad.wmnet with reason: Cloning * 06:35 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db1228.eqiad.wmnet with reason: Rebooting * 06:29 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:29 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:25 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:25 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:25 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1019.eqiad.wmnet, repooling source-only afterwards * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1018.eqiad.wmnet, repooling source-only afterwards * 04:51 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1019.eqiad.wmnet, repooling source-only afterwards * 04:51 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1018.eqiad.wmnet, repooling source-only afterwards * 04:48 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s) * 04:48 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 04:48 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s) * 04:48 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 36s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-26 == * 14:59 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:59 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:59 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:59 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1019.eqiad.wmnet with OS bookworm * 01:05 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1018.eqiad.wmnet with OS bookworm * 00:43 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1019.eqiad.wmnet with reason: host reimage * 00:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1018.eqiad.wmnet with reason: host reimage * 00:34 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1019.eqiad.wmnet with reason: host reimage * 00:33 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1018.eqiad.wmnet with reason: host reimage * 00:16 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 00:16 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 00:15 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 00:15 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1019 * 00:11 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1019 * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1018 * 00:11 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1018 * 00:08 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1019.eqiad.wmnet with OS bookworm * 00:08 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1018.eqiad.wmnet with OS bookworm == 2026-07-25 == * 22:06 ryankemper: [[phab:T430880|T430880]] [WDQS] Repooled `wdqs1017` and `wdqs2024` after reimaging to bookworm, scap deploying, and data xfering * 22:04 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2024.codfw.wmnet * 22:03 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1017.eqiad.wmnet * 21:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1017.eqiad.wmnet, repooling source-only afterwards * 21:06 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2024.codfw.wmnet, repooling source-only afterwards * 20:52 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:52 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:52 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:52 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 20:18 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1017.eqiad.wmnet, repooling source-only afterwards * 20:18 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2024.codfw.wmnet, repooling source-only afterwards * 20:15 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:15 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:15 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:15 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 19:57 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s) * 19:57 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 19:57 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 07s) * 19:57 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 19:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2024.codfw.wmnet with OS bookworm * 19:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1017.eqiad.wmnet with OS bookworm * 19:02 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2024.codfw.wmnet with reason: host reimage * 18:58 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1017.eqiad.wmnet with reason: host reimage * 18:53 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2024.codfw.wmnet with reason: host reimage * 18:52 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1017.eqiad.wmnet with reason: host reimage * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2024 * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2024 * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1017 * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1017 * 18:27 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2024 * 18:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2024.codfw.wmnet 58.16.192.10.in-addr.arpa 8.5.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:26 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2024.codfw.wmnet 58.16.192.10.in-addr.arpa 8.5.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:24 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1017 * 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1017.eqiad.wmnet 238.48.64.10.in-addr.arpa 8.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:24 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1017.eqiad.wmnet 238.48.64.10.in-addr.arpa 8.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1017 - ryankemper@cumin2003" * 18:24 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1017 - ryankemper@cumin2003" * 18:23 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 18:18 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 18:17 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1017 * 18:17 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2024 * 18:14 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1017.eqiad.wmnet with OS bookworm * 18:14 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2024.codfw.wmnet with OS bookworm * 18:05 ryankemper: [WDQS] [[phab:T430880|T430880]] Reimaged `wdqs1016` and `wdqs2023` to Bookworm with `--move-vlan`, restored main and scholarly data, validated postflights, and repooled both hosts. Confirmed PyBal rebuilt both backends with their new addresses * 17:45 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2023.codfw.wmnet * 17:43 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1016.eqiad.wmnet * 06:35 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1016.eqiad.wmnet, repooling source-only afterwards * 06:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2023.codfw.wmnet, repooling source-only afterwards * 05:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2023.codfw.wmnet, repooling source-only afterwards * 05:19 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1016.eqiad.wmnet, repooling source-only afterwards * 05:07 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 07s) * 05:07 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 05:06 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 06s) * 05:06 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 03:27 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2023.codfw.wmnet with OS bookworm * 02:59 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2023.codfw.wmnet with reason: host reimage * 02:56 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2023.codfw.wmnet with reason: host reimage * 02:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2023 * 02:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2023 * 02:30 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2023 * 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2023.codfw.wmnet 35.0.192.10.in-addr.arpa 5.3.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 02:30 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2023.codfw.wmnet 35.0.192.10.in-addr.arpa 5.3.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2023 - ryankemper@cumin2003" * 02:30 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2023 - ryankemper@cumin2003" * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 26s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:15 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1016.eqiad.wmnet with OS bookworm * 00:49 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1016.eqiad.wmnet with reason: host reimage * 00:43 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1016.eqiad.wmnet with reason: host reimage * 00:31 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 00:27 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1016 * 00:27 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1016 * 00:27 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2023 * 00:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1016.eqiad.wmnet with OS bookworm * 00:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2023.codfw.wmnet with OS bookworm * 00:11 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs1014.eqiad.wmnet and wdqs2008.codfw.wmnet after Bookworm reimage, transfer, and postflight; wdqs2008 is serving, while wdqs1014 will remain outside of service until a pybal restart next monday == 2026-07-24 == * 23:54 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1014.eqiad.wmnet * 23:54 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2008.codfw.wmnet * 23:43 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2010.codfw.wmnet with OS trixie * 23:08 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 23:03 jhathaway@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 22:33 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 22:13 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 22:13 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 22:13 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 22:13 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:00 jhathaway@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 21:53 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 21:53 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie * 21:51 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 21:47 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie * 21:43 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 21:39 jhathaway@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 21:38 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 17:21 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1135.eqiad.wmnet * 17:21 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1135.eqiad.wmnet * 17:21 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1135.eqiad.wmnet * 16:34 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 16:34 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 16:34 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 16:34 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 16:33 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 16:33 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 16:28 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:28 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:28 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:28 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2008.codfw.wmnet, repooling source-only afterwards * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1014.eqiad.wmnet, repooling source-only afterwards * 15:56 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1135.eqiad.wmnet with OS trixie * 15:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 40 hosts * 15:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 40 hosts * 15:37 topranks: upgrade SR-Linux OS on lswtest-d8-eqiad * 15:36 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1135.eqiad.wmnet with reason: host reimage * 15:33 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 6 hosts with reason: upgrade lswtest-d8-eqiad * 15:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1135.eqiad.wmnet with reason: host reimage * 15:30 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc-gp2006.codfw.wmnet with OS bookworm * 15:15 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1135 * 15:15 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1135 * 15:13 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc-gp2006.codfw.wmnet with reason: host reimage * 15:08 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc-gp2006.codfw.wmnet with reason: host reimage * 14:49 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm * 14:48 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host mc-gp2006.codfw.wmnet with OS bookworm * 14:34 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] (duration: 41m 12s) * 14:32 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1135 * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1135.eqiad.wmnet 177.32.64.10.in-addr.arpa 7.7.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:32 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1135.eqiad.wmnet 177.32.64.10.in-addr.arpa 7.7.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1135 - jiji@cumin1003" * 14:32 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1135 - jiji@cumin1003" * 14:29 krinkle@deploy1003: krinkle: Continuing with deployment * 14:29 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm * 14:27 jiji@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host mc-gp2006.codfw.wmnet with OS bookworm * 14:26 jiji@cumin1003: START - Cookbook sre.dns.netbox * 14:15 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1135 * 14:14 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1135.eqiad.wmnet with OS trixie * 14:14 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1135.eqiad.wmnet * 14:13 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1135.eqiad.wmnet * 14:13 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1135.eqiad.wmnet * 14:10 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1072.eqiad.wmnet * 14:10 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1072.eqiad.wmnet * 14:10 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1072.eqiad.wmnet * 14:10 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1072.eqiad.wmnet * 14:09 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1071.eqiad.wmnet * 14:09 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1071.eqiad.wmnet * 14:09 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1071.eqiad.wmnet * 14:09 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1071.eqiad.wmnet * 13:58 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 13:58 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:58 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:57 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:55 krinkle@deploy1003: krinkle: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:53 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] * 13:45 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:45 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push new IPs for mc-gp2006 - cmooney@cumin1003" * 13:45 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push new IPs for mc-gp2006 - cmooney@cumin1003" * 13:44 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) mc-gp2006.codfw.wmnet on all recursors * 13:44 cmooney@cumin1003: START - Cookbook sre.dns.wipe-cache mc-gp2006.codfw.wmnet on all recursors * 13:42 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm * 13:41 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:30 papaul: reboot mr1-eqsin for maintenance * 13:24 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb[1029-1031].eqiad.wmnet * 13:10 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb[1029-1031].eqiad.wmnet * 11:33 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 7 hosts * 11:11 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 7 hosts * 10:56 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 7 hosts * 10:47 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 7 hosts * 10:44 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:44 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:41 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:41 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:35 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 8 hosts * 10:34 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:33 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:32 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:32 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:31 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:31 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:30 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 8 hosts * 10:24 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie * 10:19 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:18 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 16 hosts * 10:17 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2001.codfw.wmnet * 10:13 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2001.codfw.wmnet * 10:12 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2001.codfw.wmnet * 10:02 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2001.codfw.wmnet * 10:02 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2002.codfw.wmnet * 09:57 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2002.codfw.wmnet * 09:56 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2002.codfw.wmnet * 09:51 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2002.codfw.wmnet * 09:51 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1002.eqiad.wmnet * 09:47 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1002.eqiad.wmnet * 09:47 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1001.eqiad.wmnet * 09:44 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1001.eqiad.wmnet * 09:34 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2003.codfw.wmnet * 09:32 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2003.codfw.wmnet * 09:32 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2002.codfw.wmnet * 09:30 brouberol@dns1004: END - running authdns-update * 09:29 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2002.codfw.wmnet * 09:29 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2001.codfw.wmnet * 09:27 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 9 hosts * 09:27 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2001.codfw.wmnet * 09:27 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2001.codfw.wmnet * 09:26 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 9 hosts * 09:26 brouberol@dns1004: START - running authdns-update * 09:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 57 hosts * 09:24 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2001.codfw.wmnet * 09:24 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2002.codfw.wmnet * 09:22 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2002.codfw.wmnet * 09:21 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 57 hosts * 09:20 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2003.codfw.wmnet * 09:19 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 16 hosts * 09:16 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2003.codfw.wmnet * 09:16 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1003.eqiad.wmnet * 09:15 urbanecm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 09:15 urbanecm@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 09:13 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1003.eqiad.wmnet * 09:13 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1002.eqiad.wmnet * 09:11 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1002.eqiad.wmnet * 09:11 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1001.eqiad.wmnet * 09:07 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1001.eqiad.wmnet * 08:32 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:24 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:16 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 08:16 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 08:07 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:07 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:07 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 08:02 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 08:01 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:59 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:57 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 07:57 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 06:46 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1025.eqiad.wmnet with reason: Cloning * 06:46 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s4 * 06:45 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s6 * 06:44 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1019.eqiad.wmnet,service=s6 * 06:44 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1019.eqiad.wmnet,service=s4 * 03:40 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:40 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:40 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:40 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 03:37 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:37 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:37 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:36 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:49 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on mr1-eqsin,mr1-eqsin IPv6 with reason: connection issue * 02:38 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on cr[2-3]-eqsin.mgmt,ps1-[603-604]-eqsin with reason: connection issue * 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 27s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-23 == * 23:27 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin.oob,mr1-eqsin.oob IPv6 with reason: switch refresh * 22:21 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Setting storage compatibility to NONE - eevans@cumin1003 * 22:01 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Setting storage compatibility to NONE - eevans@cumin1003 * 21:29 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1014.eqiad.wmnet, repooling source-only afterwards * 21:28 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 46s) * 21:28 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 21:19 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Setting storage compatibility to UPGRADING - eevans@cumin1003 * 21:00 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Setting storage compatibility to UPGRADING - eevans@cumin1003 * 20:17 dani@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] (duration: 11m 57s) * 20:13 dani@deploy1003: dani: Continuing with deployment * 20:07 dani@deploy1003: dani: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:05 dani@deploy1003: Started scap sync-world: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] * 19:24 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:24 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating the rest of the ipv6 dns records. - jhancock@cumin2002" * 19:24 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating the rest of the ipv6 dns records. - jhancock@cumin2002" * 19:14 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 19:05 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wdqs1014.eqiad.wmnet with OS bookworm * 19:04 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.noop (exit_code=99) * 19:04 cwilliams@cumin1003: START - Cookbook sre.mysql.noop * 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1014.eqiad.wmnet with reason: host reimage * 18:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1014.eqiad.wmnet with reason: host reimage * 18:30 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2008.codfw.wmnet, repooling source-only afterwards * 18:28 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 19s) * 18:28 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1014 * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1014 * 18:22 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1014 * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1014.eqiad.wmnet 188.32.64.10.in-addr.arpa 8.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:22 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1014.eqiad.wmnet 188.32.64.10.in-addr.arpa 8.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1014 - bking@cumin2003" * 18:21 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1014 - bking@cumin2003" * 18:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2215: Maintenance * 18:18 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 18:15 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:15 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 18:06 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:06 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 18:05 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 18:04 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2052: codfw rack B8 re-pool after maintenance * 17:54 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 17:54 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:54 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 17:32 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2215: Maintenance * 17:29 cmooney@dns3003: END - running authdns-update * 17:27 cmooney@dns3003: START - running authdns-update * 17:23 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 17:22 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:18 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool es2052: codfw rack B8 re-pool after maintenance * 17:18 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2189: codfw rack B8 re-pool after maintenance * 17:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2215.codfw.wmnet with reason: Maintenance * 17:17 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 17:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2215 [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95126 and previous config saved to /var/cache/conftool/dbconfig/20260723-170903-cwilliams.json * 17:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2191 to x1 primary [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95125 and previous config saved to /var/cache/conftool/dbconfig/20260723-170612-cwilliams.json * 17:05 cezmunsta: Starting x1 codfw failover from db2215 to db2191 - [[phab:T432986|T432986]] * 16:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2191 with weight 0 [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95123 and previous config saved to /var/cache/conftool/dbconfig/20260723-165831-cwilliams.json * 16:58 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 16 hosts with reason: Primary switchover x1 [[phab:T432986|T432986]] * 16:36 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 138128 * 16:35 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 138128 * 16:33 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2189: codfw rack B8 re-pool after maintenance * 16:33 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2164: codfw rack B8 re-pool after maintenance * 16:28 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1072.eqiad.wmnet * 16:27 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1072.eqiad.wmnet with OS trixie * 16:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2249: Maintenance * 16:06 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker1072.eqiad.wmnet with reason: host reimage * 16:06 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1072.eqiad.wmnet with reason: host reimage * 15:50 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1072 * 15:50 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1072 * 15:49 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1072.eqiad.wmnet with OS trixie * 15:48 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2164: codfw rack B8 re-pool after maintenance * 15:48 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] (duration: 06m 37s) * 15:48 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2163: codfw rack B8 re-pool after maintenance * 15:45 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 15:43 musikanimal@deploy1003: musikanimal: Continuing with deployment * 15:43 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:41 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] * 15:36 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1072.eqiad.wmnet * 15:35 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1072.eqiad.wmnet * 15:35 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1072.eqiad.wmnet * 15:34 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:34 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push any outstanding updates - cmooney@cumin1003" * 15:34 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push any outstanding updates - cmooney@cumin1003" * 15:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db2249: Maintenance * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 15:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:26 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:21 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:21 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 15:21 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:21 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 15:20 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 15:19 cmooney@dns2004: END - running authdns-update * 15:17 cmooney@dns2004: START - running authdns-update * 15:14 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns2004.wikimedia.org * 15:12 brouberol@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 15:12 brouberol@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 15:12 klausman@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ml-serve1001.eqiad.wmnet with OS trixie * 15:11 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1071.eqiad.wmnet * 15:11 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1071.eqiad.wmnet with OS trixie * 15:10 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wdqs2008.codfw.wmnet with OS bookworm * 15:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2249.codfw.wmnet with reason: Maintenance * 15:08 brouberol@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 15:08 brouberol@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 15:08 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 15:08 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 15:06 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns1004.wikimedia.org * 15:02 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2002.codfw.wmnet * 15:02 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2002.codfw.wmnet * 15:02 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2163: codfw rack B8 re-pool after maintenance * 15:01 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 15:01 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 14:59 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test2001.codfw.wmnet * 14:57 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test2001.codfw.wmnet * 14:56 ryankemper: [WDQS] [[phab:T430880|T430880]] Reimaged `wdqs2016` to Bookworm, xferred scholarly_articles from `wdqs2024`, validated updater/readiness/federation, and repooled * 14:51 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2016.codfw.wmnet * 14:51 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1001.eqiad.wmnet with reason: host reimage * 14:48 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1071.eqiad.wmnet with reason: host reimage * 14:47 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1001.eqiad.wmnet with reason: host reimage * 14:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2008.codfw.wmnet with reason: host reimage * 14:43 topranks: reboot lsw1-b8-codw to upgrade JunOS [[phab:T430929|T430929]] * 14:41 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1071.eqiad.wmnet with reason: host reimage * 14:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2008.codfw.wmnet with reason: host reimage * 14:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2231: Maintenance * 14:30 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1001.eqiad.wmnet with OS trixie * 14:25 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2002.codfw.wmnet * 14:23 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1071 * 14:23 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1071 * 14:23 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 14:22 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore scholarly data after Bookworm reimage) xfer scholarly_articles from wdqs2024.codfw.wmnet -> wdqs2016.codfw.wmnet, repooling source-only afterwards * 14:22 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2052: codfw rack B8 depool for maintenance * 14:21 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool es2052: codfw rack B8 depool for maintenance * 14:21 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2249: codfw rack B8 depool for maintenance * 14:21 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1071 * 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1071.eqiad.wmnet 166.48.64.10.in-addr.arpa 6.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:21 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1071.eqiad.wmnet 166.48.64.10.in-addr.arpa 6.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1071 - jiji@cumin1003" * 14:21 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1071 - jiji@cumin1003" * 14:21 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2249: codfw rack B8 depool for maintenance * 14:21 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2189: codfw rack B8 depool for maintenance * 14:20 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2002.codfw.wmnet * 14:20 cmooney@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2050.codfw.wmnet * 14:20 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2189: codfw rack B8 depool for maintenance * 14:20 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2164: codfw rack B8 depool for maintenance * 14:20 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2164: codfw rack B8 depool for maintenance * 14:19 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2163: codfw rack B8 depool for maintenance * 14:19 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:19 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1014 * 14:19 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2163: codfw rack B8 depool for maintenance * 14:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1014.eqiad.wmnet with OS bookworm * 14:17 cmooney@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2050.codfw.wmnet * 14:16 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 14:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2008 * 14:14 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2008 * 14:14 cmooney@cumin1003: conftool action : set/pooled=no; selector: name=dns2004.wikimedia.org * 14:14 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2008.codfw.wmnet with OS bookworm * 14:13 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 14:12 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:10 topranks: depool dns2004 before lsw1-b8-codfw switch maintenance [[phab:T430929|T430929]] * 14:10 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b8-codfw,lsw1-b8-codfw IPv6,lsw1-b8-codfw.mgmt,ssw1-a[1,8]-codfw with reason: lsw1-b8-codfw JunOS upgrade * 14:07 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 30 hosts with reason: lsw1-b8-codfw JunOS upgrade * 14:06 elukey: upload python3-docker-report 0.0.19 to apt.wikimedia.org for bookworm and trixie * 13:59 jiji@cumin1003: START - Cookbook sre.dns.netbox * 13:58 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1071 * 13:57 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:57 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:57 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:57 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1071.eqiad.wmnet with OS trixie * 13:55 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1071.eqiad.wmnet * 13:55 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1071.eqiad.wmnet * 13:55 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1071.eqiad.wmnet * 13:53 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:52 logmsgbot: kharlan Deployed security patch for [[phab:T432948|T432948]] * 13:51 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:51 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 13:51 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db2231: Maintenance * 13:50 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:50 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:50 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:50 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:49 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:49 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:49 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2231 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95097 and previous config saved to /var/cache/conftool/dbconfig/20260723-134436-cwilliams.json * 13:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2231.codfw.wmnet with reason: Maintenance * 13:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:39 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:38 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 13:38 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] (duration: 09m 07s) * 13:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1037 hosts * 13:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2196: Maintenance * 13:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:33 kharlan@deploy1003: kharlan, emc-wmf: Continuing with deployment * 13:31 kharlan@deploy1003: kharlan, emc-wmf: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:30 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:28 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] * 13:17 hashar@deploy1003: Finished deploy [integration/docroot@2199146]: build: License GPL2.0+ / updating npm dependencies (duration: 00m 14s) * 13:17 hashar@deploy1003: Started deploy [integration/docroot@2199146]: build: License GPL2.0+ / updating npm dependencies * 13:14 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service * 13:07 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 12:58 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2207: Repooling * 12:49 root@cumin1003: START - Cookbook sre.mysql.pool pool db2196: Maintenance * 12:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2196 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95087 and previous config saved to /var/cache/conftool/dbconfig/20260723-123952-cwilliams.json * 12:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2196.codfw.wmnet with reason: Maintenance * 12:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2191: Maintenance * 12:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:13 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:13 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: Repooling * 12:12 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2207: Repooling * 12:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: Repooling * 11:56 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2235.codfw.wmnet with OS trixie * 11:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db2191: Maintenance * 11:46 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1070.eqiad.wmnet * 11:46 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1070.eqiad.wmnet * 11:46 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1070.eqiad.wmnet * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2191 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95080 and previous config saved to /var/cache/conftool/dbconfig/20260723-114308-cwilliams.json * 11:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2191.codfw.wmnet with reason: Maintenance * 11:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2186: Maintenance * 11:35 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 46375 * 11:34 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 46375 * 11:33 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2235.codfw.wmnet with reason: host reimage * 11:28 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2235.codfw.wmnet with reason: host reimage * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c7-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c7-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c6-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c6-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c5-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c5-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c4-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c4-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c3-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c3-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c2-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c2-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d7-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d7-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d4-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d4-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d3-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d2-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d2-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d8-eqiad * 11:23 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d8-eqiad * 11:23 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d1-eqiad * 11:23 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d1-eqiad * 11:12 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db2235.codfw.wmnet with OS trixie * 11:11 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:11 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[2160,2235].codfw.wmnet with reason: Upgrading * 10:56 root@cumin1003: START - Cookbook sre.mysql.pool pool db2186: Maintenance * 10:54 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1037: testing * 10:53 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1037: testing * 10:53 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1037: testing * 10:53 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1037: testing * 10:52 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: testing * 10:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2186 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95072 and previous config saved to /var/cache/conftool/dbconfig/20260723-104956-cwilliams.json * 10:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2186.codfw.wmnet with reason: Maintenance * 10:43 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1054: testing * 10:41 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1070.eqiad.wmnet with OS trixie * 10:30 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1037 hosts * 10:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 10:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 10:20 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1070.eqiad.wmnet with reason: host reimage * 10:16 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1070.eqiad.wmnet with reason: host reimage * 10:06 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1038: testing * 10:05 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1038: testing * 10:05 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1038: testing * 10:02 marostegui@dns1004: END - running authdns-update * 10:00 marostegui@dns1004: START - running authdns-update * 09:58 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: testing * 09:57 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1054: testing * 09:57 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1070 * 09:57 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1070 * 09:57 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1054: testing * 09:57 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1054: testing * 09:56 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1055: testing * 09:56 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1070 * 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1070.eqiad.wmnet 165.48.64.10.in-addr.arpa 5.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:56 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1070.eqiad.wmnet 165.48.64.10.in-addr.arpa 5.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1070 - jiji@cumin1003" * 09:56 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1070 - jiji@cumin1003" * 09:47 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for 1035 hosts * 09:45 jiji@cumin1003: START - Cookbook sre.dns.netbox * 09:42 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1070 * 09:42 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1070.eqiad.wmnet with OS trixie * 09:42 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1070.eqiad.wmnet * 09:41 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1070.eqiad.wmnet * 09:41 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1070.eqiad.wmnet * 09:27 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es2051: testing * 09:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: testing * 09:12 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es2051: testing * 09:11 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1055: testing * 09:09 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:09 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1055: testing * 09:09 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1055: testing * 08:50 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:50 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:50 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 08:49 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 08:49 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 08:49 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:46 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1069.eqiad.wmnet * 08:46 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1069.eqiad.wmnet * 08:46 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1069.eqiad.wmnet * 08:39 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 08:38 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 08:38 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 08:37 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 08:35 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:10 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1069.eqiad.wmnet with OS trixie * 07:49 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1069.eqiad.wmnet with reason: host reimage * 07:45 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1069.eqiad.wmnet with reason: host reimage * 07:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1035 hosts * 07:33 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 07:32 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 07:29 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1069 * 07:29 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1069 * 07:26 jiji@deploy1003: Finished scap sync-world: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules (duration: 06m 01s) * 07:25 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1069 * 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1069.eqiad.wmnet 164.48.64.10.in-addr.arpa 4.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:25 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1069.eqiad.wmnet 164.48.64.10.in-addr.arpa 4.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1069 - jiji@cumin1003" * 07:25 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1069 - jiji@cumin1003" * 07:25 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 07:25 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 07:24 jiji@deploy1003: jiji: Continuing with deployment * 07:22 jiji@deploy1003: jiji: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:21 jiji@deploy1003: Started scap sync-world: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules * 07:21 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1031.eqiad.wmnet,service=s7 * 07:20 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1031.eqiad.wmnet,service=s2 * 07:20 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1031.eqiad.wmnet,service=s7 * 07:20 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1031.eqiad.wmnet,service=s2 * 07:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts * 07:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts * 07:19 jiji@cumin1003: START - Cookbook sre.dns.netbox * 07:19 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1069 * 07:19 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1069.eqiad.wmnet with OS trixie * 07:19 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1069.eqiad.wmnet * 07:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 45 hosts * 07:17 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1069.eqiad.wmnet * 07:17 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1069.eqiad.wmnet * 07:14 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 45 hosts * 07:13 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 06:16 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs2007 after successful Bookworm reimage, data transfer, and postflight validation; wdqs1013 also passed postflights and is enabled in conftool, but remains out of IPVS pending a rolling pybal restart to clear its stale pre-VLAN-move address * 05:58 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2007.codfw.wmnet * 05:58 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1013.eqiad.wmnet * 05:54 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore scholarly data after Bookworm reimage) xfer scholarly_articles from wdqs2024.codfw.wmnet -> wdqs2016.codfw.wmnet, repooling source-only afterwards == 2026-07-22 == * 23:34 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Apply upgrade to JVM17 - eevans@cumin1003 * 23:14 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Apply upgrade to JVM17 - eevans@cumin1003 * 22:06 ryankemper: [WDQS] Added requestctl per-IP ratelimit `wdqs_heavy_sparql_bots_jul_2026_ratelimit` (chronic heavy-query bot tier driving deadlock-remediation restarts); pruned superseded `wdqs_2026_05_11_worobot` * 21:51 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] (duration: 11m 52s) * 21:44 sbassett@deploy1003: sbassett: Continuing with deployment * 21:43 sbassett@deploy1003: sbassett: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:39 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] * 20:38 dani@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] (duration: 32m 51s) * 20:38 ryankemper: [WDQS] Pruned obsolete requestctl action+pattern `wdqs_20260715_p2003_ring_ja3n` (actor rotated JA3Ns; rule inert) * 20:26 dani@deploy1003: dani, vadymts1: Continuing with deployment * 20:24 dani@deploy1003: dani, vadymts1: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:14 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2244: Testing * 20:06 dani@deploy1003: Started scap sync-world: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] * 19:56 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1013.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2016.codfw.wmnet with OS bookworm * 19:40 mutante: gerrit - one more service restart is needed - restarting * 19:29 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2244: Testing * 19:27 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2244: Testing * 19:27 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2244: Testing * 19:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2016.codfw.wmnet with reason: host reimage * 19:15 dancy@deploy1003: Finished deploy [zuul/deploy@d92e238]: Freshening Zuul installation (duration: 00m 15s) * 19:14 dancy@deploy1003: Started deploy [zuul/deploy@d92e238]: Freshening Zuul installation * 19:11 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2016.codfw.wmnet with reason: host reimage * 18:54 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1013.eqiad.wmnet, repooling source-only afterwards * 18:52 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2016 * 18:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2016 * 18:51 dancy@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 18:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2016.codfw.wmnet with OS bookworm * 18:39 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 11s) * 18:39 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 18:36 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 18:30 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417] (thin): Regular analytics weekly train THIN [analytics/refinery@2a25417d] (duration: 02m 09s) * 18:28 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417] (thin): Regular analytics weekly train THIN [analytics/refinery@2a25417d] * 18:28 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417]: Regular analytics weekly train [analytics/refinery@2a25417d] (duration: 04m 31s) * 18:27 dduvall: deploying https://gerrit.wikimedia.org/r/c/integration/config/+/1314025 (4 jobs updated) * 18:23 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417]: Regular analytics weekly train [analytics/refinery@2a25417d] * 18:22 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@2a25417d] (duration: 01m 59s) * 18:20 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@2a25417d] * 17:56 Raine: deployment server switchover => deploy1003 is primary now * 17:55 kamila@deploy1003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 28m 24s) * 17:54 mutante: restarting gerrit for maintenance * 17:29 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1013.eqiad.wmnet with OS bookworm * 17:27 kamila@deploy1003: Started scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] * 17:20 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] (duration: 22m 50s) * 17:12 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1023.eqiad.wmnet -> wdqs1024.eqiad.wmnet, repooling source-only afterwards * 17:04 Raine: point deployment.eqiad.wmnet to deploy1003 * 17:04 kamila@dns7001: END - running authdns-update * 17:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1013.eqiad.wmnet with reason: host reimage * 17:02 kamila@dns7001: START - running authdns-update * 17:01 kamila@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on releases2003.codfw.wmnet,releases1003.eqiad.wmnet with reason: Deployment server switchover * 17:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1013.eqiad.wmnet with reason: host reimage * 16:58 kamila@deploy2003: Locking from deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] * 16:57 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] (duration: 02m 33s) * 16:55 kamila@deploy2003: Locking from deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] * 16:55 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2003 - [[phab:T240266|T240266]] (duration: 00m 11s) * 16:54 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2003 - [[phab:T240266|T240266]] * 16:40 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1013 * 16:40 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1013 * 16:39 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1013 * 16:39 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1013.eqiad.wmnet 105.32.64.10.in-addr.arpa 5.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:39 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1013.eqiad.wmnet 105.32.64.10.in-addr.arpa 5.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:39 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:39 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1013 - bking@cumin2003" * 16:39 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1013 - bking@cumin2003" * 16:34 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:34 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1013 * 16:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1013.eqiad.wmnet with OS bookworm * 16:28 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1023.eqiad.wmnet -> wdqs1024.eqiad.wmnet, repooling source-only afterwards * 16:27 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-scholarly,name=eqiad * 16:27 eevans@deploy2003: helmfile [eqiad] DONE helmfile.d/services/linked-artifacts: apply * 16:26 eevans@deploy2003: helmfile [eqiad] START helmfile.d/services/linked-artifacts: apply * 16:26 eevans@deploy2003: helmfile [codfw] DONE helmfile.d/services/linked-artifacts: apply * 16:26 eevans@deploy2003: helmfile [codfw] START helmfile.d/services/linked-artifacts: apply * 16:25 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 29s) * 16:25 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 16:24 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 16:21 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 16:18 eevans@deploy2003: helmfile [codfw] DONE helmfile.d/services/linked-artifacts: apply * 16:18 eevans@deploy2003: helmfile [codfw] START helmfile.d/services/linked-artifacts: apply * 16:08 eevans@deploy2003: helmfile [staging] DONE helmfile.d/services/linked-artifacts: apply * 16:07 eevans@deploy2003: helmfile [staging] START helmfile.d/services/linked-artifacts: apply * 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 16:01 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 15:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1024.eqiad.wmnet with OS bookworm * 15:49 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] (duration: 00m 10s) * 15:49 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] * 15:48 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] (duration: 00m 15s) * 15:48 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] * 15:47 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] (duration: 00m 10s) * 15:47 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] * 15:46 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:42 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:42 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:40 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:37 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 15:37 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:36 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:36 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:36 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 15:33 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 15:33 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:31 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:28 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 15:27 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:27 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1024.eqiad.wmnet with reason: host reimage * 15:23 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:23 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:23 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:20 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1068.eqiad.wmnet * 15:20 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1068.eqiad.wmnet * 15:20 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1068.eqiad.wmnet * 15:20 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 15:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1024.eqiad.wmnet with reason: host reimage * 15:11 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:55 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 14:52 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wdqs1024.eqiad.wmnet with OS bookworm * 14:50 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] (duration: 00m 09s) * 14:50 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] * 14:49 jiji@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 14:49 jiji@deploy2003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 14:49 jiji@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 14:48 jiji@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 14:45 ecarg@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:45 ecarg@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:44 ecarg@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:44 ecarg@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:43 ecarg@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:43 ecarg@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:41 sukhe: ipvsadm --delete-service --tcp-service 10.2.1.55:8087: lvs2014 and lvs2013 * 14:39 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:39 sukhe: ipvsadm --delete-service --tcp-service 10.2.2.55:8087: [[phab:T432445|T432445]] * 14:38 ecarg@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:38 ecarg@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:37 ecarg@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:37 ecarg@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:36 ecarg@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:34 ecarg@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts datahubsearch1001.eqiad.wmnet * 14:32 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:32 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 14:31 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 14:31 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 14:28 sukhe: sudo cumin 'A:lvs-low-traffic-codfw' 'systemctl restart pybal': lvs2013 * 14:26 sukhe: sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal': lvs2014 * 14:26 sukhe: sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal' * 14:24 sukhe: restart pybal on lvs1019 * 14:24 sukhe: restart pybal on lvs1020 * 14:19 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for 1036 hosts * 14:17 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:04 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] (duration: 09m 28s) * 13:59 kharlan@deploy2003: dreamyjazz, kharlan: Continuing with deployment * 13:58 bking@cumin2003: START - Cookbook sre.hosts.decommission for hosts datahubsearch1001.eqiad.wmnet * 13:57 kharlan@deploy2003: dreamyjazz, kharlan: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:55 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] * 13:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts datahubsearch[1002-1003].eqiad.wmnet * 13:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:53 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch[1002-1003].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 13:52 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch[1002-1003].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 13:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 13:42 stran@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] (duration: 07m 30s) * 13:42 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:40 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: Test * 13:38 stran@deploy2003: dragoniez, stran: Continuing with deployment * 13:37 stran@deploy2003: dragoniez, stran: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:35 bking@cumin2003: START - Cookbook sre.hosts.decommission for hosts datahubsearch[1002-1003].eqiad.wmnet * 13:35 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024'] * 13:35 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 13:35 stran@deploy2003: Started scap sync-world: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] * 13:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 13:28 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024'] * 13:26 sukhe@dns1004: END - running authdns-update * 13:25 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 13:25 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 13:24 sukhe@dns1004: START - running authdns-update * 13:22 stran@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] (duration: 08m 20s) * 13:21 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 13:20 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 13:19 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 13:19 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 13:19 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 13:18 stran@deploy2003: stran: Continuing with deployment * 13:16 stran@deploy2003: stran: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:14 stran@deploy2003: Started scap sync-world: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] * 13:13 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 13:13 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 13:11 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 13:11 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 13:08 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 12:55 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool es2051: Test * 12:55 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: Test * 12:54 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool es2051: Test * 12:43 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1036 hosts * 12:41 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 12:40 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1048.eqiad.wmnet * 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 12:39 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 12:38 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 12:37 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 12:37 brouberol@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 12:36 brouberol@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 12:36 brouberol@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 12:36 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 12:36 elukey@cumin1003: DONE (PASS) - Cookbook sre.puppet.renew-cert (exit_code=0) for crm2001.codfw.wmnet: Renew puppet certificate - elukey@cumin1003 * 12:35 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:35 brouberol@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 12:34 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 12:31 brouberol@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 12:30 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1048.eqiad.wmnet * 12:30 brouberol@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 12:28 brouberol@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 12:27 brouberol@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 12:20 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1068.eqiad.wmnet with OS trixie * 12:01 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] (duration: 13m 19s) * 11:58 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1068.eqiad.wmnet with reason: host reimage * 11:52 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1068.eqiad.wmnet with reason: host reimage * 11:51 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 11:49 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:47 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] * 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2252: Security updates * 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:43 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 11:42 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2252: Security updates * 11:42 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply * 11:40 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply * 11:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 11:37 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2252.codfw.wmnet with OS trixie * 11:34 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1068 * 11:34 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1068 * 11:26 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1068 * 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1068.eqiad.wmnet 46.48.64.10.in-addr.arpa 6.4.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:26 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1068.eqiad.wmnet 46.48.64.10.in-addr.arpa 6.4.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1068 - jiji@cumin1003" * 11:26 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1068 - jiji@cumin1003" * 11:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2252.codfw.wmnet with reason: host reimage * 11:17 jiji@cumin1003: START - Cookbook sre.dns.netbox * 11:17 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1068 * 11:17 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1068.eqiad.wmnet with OS trixie * 11:17 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2252.codfw.wmnet with reason: host reimage * 11:15 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1068.eqiad.wmnet * 11:15 mvolz@deploy2003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:15 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1068.eqiad.wmnet * 11:15 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1068.eqiad.wmnet * 11:14 mvolz@deploy2003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:13 mvolz@deploy2003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:13 mvolz@deploy2003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:12 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] (duration: 11m 05s) * 11:11 mvolz@deploy2003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:10 mvolz@deploy2003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:07 dreamyjazz@deploy2003: dreamyjazz, kharlan: Continuing with deployment * 11:03 dreamyjazz@deploy2003: dreamyjazz, kharlan: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:03 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2252.codfw.wmnet with OS trixie * 11:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2252: Upgrading db2252.codfw.wmnet * 11:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:02 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 11:02 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2252: Upgrading db2252.codfw.wmnet * 11:02 cwilliams@cumin1003: dbmaint on ms3@codfw [[phab:T432321|T432321]] * 11:01 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 11:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db1153.eqiad.wmnet with reason: Security updates * 11:01 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] * 11:00 fnegri@deploy2003: helmfile [eqiad] DONE helmfile.d/services/toolhub: apply * 10:58 fnegri@deploy2003: helmfile [eqiad] START helmfile.d/services/toolhub: apply * 10:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1151: Security updates * 10:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:57 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1151: Security updates * 10:55 fnegri@deploy2003: helmfile [codfw] DONE helmfile.d/services/toolhub: apply * 10:54 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] (duration: 08m 38s) * 10:53 fnegri@deploy2003: helmfile [codfw] START helmfile.d/services/toolhub: apply * 10:53 fnegri@deploy2003: helmfile [staging] DONE helmfile.d/services/toolhub: apply * 10:52 fnegri@deploy2003: helmfile [staging] START helmfile.d/services/toolhub: apply * 10:50 zabe@deploy2003: zabe: Continuing with deployment * 10:47 zabe@deploy2003: zabe: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:45 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] * 10:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1151: Security updates * 10:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:42 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:42 root@cumin1003: START - Cookbook sre.mysql.depool depool db1151: Security updates * 10:38 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] (duration: 12m 47s) * 10:34 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 10:34 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 10:33 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2253.codfw.wmnet with OS trixie * 10:28 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:26 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] * 10:18 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2253.codfw.wmnet with reason: host reimage * 10:13 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2253.codfw.wmnet with reason: host reimage * 10:00 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2253.codfw.wmnet with OS trixie * 09:58 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db1151.eqiad.wmnet with reason: Security updates * 09:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2253: Upgrading db2253.codfw.wmnet * 09:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:57 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 09:56 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2253: Upgrading db2253.codfw.wmnet * 09:56 cwilliams@cumin1003: dbmaint on ms2@codfw [[phab:T432321|T432321]] * 09:56 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 09:36 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: UI improvement; support url shortener - oblivian@cumin1003" * 09:36 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: UI improvement; support url shortener - oblivian@cumin1003 * 09:35 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: UI improvement; support url shortener - oblivian@cumin1003 * 09:35 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: UI improvement; support url shortener - oblivian@cumin1003" * 09:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1152: Security updates * 09:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:26 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db1152: Security updates * 09:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: Security updates * 09:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:11 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:11 root@cumin1003: START - Cookbook sre.mysql.depool depool db1152: Security updates * 09:10 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1018.eqiad.wmnet with reason: Cloning * 09:09 Dreamy_Jazz: Deployed patch for [[phab:T432453|T432453]] and [[phab:T432454|T432454]] * 09:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 09:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2251.codfw.wmnet with OS trixie * 08:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2251.codfw.wmnet with reason: host reimage * 08:45 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2251.codfw.wmnet with reason: host reimage * 08:40 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] (duration: 12m 26s) * 08:38 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1030.eqiad.wmnet,service=s1 * 08:36 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 08:31 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2251.codfw.wmnet with OS trixie * 08:30 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:28 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] * 08:25 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] (duration: 07m 59s) * 08:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2251: Upgrading db2251.codfw.wmnet * 08:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:22 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 08:22 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2251: Upgrading db2251.codfw.wmnet * 08:20 urbanecm@deploy2003: urbanecm: Continuing with deployment * 08:20 cwilliams@cumin1003: dbmaint on ms1@codfw [[phab:T432321|T432321]] * 08:20 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 08:19 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:17 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] * 08:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade * 08:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade * 08:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db2251.codfw.wmnet,db1152.eqiad.wmnet with reason: OS upgrade * 08:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade * 08:13 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade * 08:11 Dreamy_Jazz: Created cusi_signal, cusi_case, and cusi_user on ukwiki and enwikivoyage in extension1 * 08:11 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade * 08:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade * 08:04 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1030.eqiad.wmnet,service=s1 * 08:04 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1030.eqiad.wmnet,service=s1 * 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply * 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply * 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply * 07:51 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply * 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 07:47 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 07:47 phuedx: End of UTC morning backport window * 07:43 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 07:43 phuedx@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] (duration: 13m 44s) * 07:43 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 07:39 phuedx@deploy2003: phuedx: Continuing with deployment * 07:31 phuedx@deploy2003: phuedx: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:29 phuedx@deploy2003: Started scap sync-world: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] * 07:24 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Turnilo import support - oblivian@cumin1003" * 07:24 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import support - oblivian@cumin1003 * 07:23 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import support - oblivian@cumin1003 * 07:23 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Turnilo import support - oblivian@cumin1003" * 06:42 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs2020 after successful Bookworm reimage, data transfer, and postflight validation * 06:42 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2020.codfw.wmnet * 05:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Managing sanitization for wikis bolwiki in section s5 * 05:25 marostegui@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis bolwiki in section s5 * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 41s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-21 == * 22:50 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2019.codfw.wmnet -> wdqs2020.codfw.wmnet, repooling source-only afterwards * 22:47 cwhite: force reboot arclamp2001 - appears to have run out of memory and gone unresponsive * 22:24 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 01m 26s) * 22:24 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 22:23 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 22:22 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024'] * 22:11 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 21:54 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs1024'] * 21:54 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 21:53 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs1024'] * 21:53 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 21:49 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2019.codfw.wmnet -> wdqs2020.codfw.wmnet, repooling source-only afterwards * 20:57 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] (duration: 09m 10s) * 20:55 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1024.eqiad.wmnet with OS bookworm * 20:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2020.codfw.wmnet with OS bookworm * 20:52 krinkle@deploy2003: krinkle: Continuing with deployment * 20:49 krinkle@deploy2003: krinkle: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:47 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] * 20:45 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] (duration: 05m 42s) * 20:44 krinkle@deploy2003: krinkle: Rolling back deployment * 20:41 krinkle@deploy2003: krinkle: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:39 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] * 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2003.codfw.wmnet * 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1003.eqiad.wmnet * 20:33 dani@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] (duration: 11m 15s) * 20:33 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2003.codfw.wmnet * 20:33 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1003.eqiad.wmnet * 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2020.codfw.wmnet with reason: host reimage * 20:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1002.eqiad.wmnet * 20:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2002.codfw.wmnet * 20:30 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 20:29 dani@deploy2003: dani: Continuing with deployment * 20:29 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2020.codfw.wmnet with reason: host reimage * 20:26 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1002.eqiad.wmnet * 20:26 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2002.codfw.wmnet * 20:24 dani@deploy2003: dani: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2001.codfw.wmnet * 20:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1001.eqiad.wmnet * 20:22 dani@deploy2003: Started scap sync-world: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] * 20:22 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 20:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2001.codfw.wmnet * 20:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1001.eqiad.wmnet * 20:14 sbisson@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] (duration: 09m 01s) * 20:11 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2020.codfw.wmnet with OS bookworm * 20:10 sbisson@deploy2003: sbisson: Continuing with deployment * 20:07 sbisson@deploy2003: sbisson: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:05 sbisson@deploy2003: Started scap sync-world: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] * 20:03 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] (duration: 07m 04s) * 20:01 mutante: Gerrit - tomorrow a new SSH host key will appear - it will be {{Gerrit|ed25519}} and has already been added to wmf-laptop. you can verify it here: https://wikitech.wikimedia.org/wiki/Help:SSH_Fingerprints/gerrit.wikimedia.org:29418 ([[phab:T240266|T240266]]) * 19:59 zabe@deploy2003: zabe: Continuing with deployment * 19:58 zabe@deploy2003: zabe: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:56 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] * 19:52 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] (duration: 07m 25s) * 19:48 zabe@deploy2003: zabe: Continuing with deployment * 19:47 zabe@deploy2003: zabe: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:45 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] * 19:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 19:32 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024'] * 19:27 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 19:26 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024'] * 19:26 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 19:24 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024'] * 19:12 ryankemper: [wdqs] [[phab:T430880|T430880]] Repooled `wdqs-scholarly` discovery in `eqiad` after validating `wdqs1023` end-to-end; `wdqs1024` remains disabled pending reimage recovery * 19:11 ryankemper: [wdqs] [[phab:T430880|T430880]] Repooled wdqs1012.eqiad.wmnet after successful Bookworm reimage, data transfer, service checks, readiness probe, and cross-graph federation query validation * 19:10 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 19:10 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1012.eqiad.wmnet * 19:08 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 18:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for deploy1003.eqiad.wmnet * 18:57 kamila@cumin1003: START - Cookbook sre.hosts.remove-downtime for deploy1003.eqiad.wmnet * 18:37 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:37 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding urldownloader service IPs - sukhe@cumin1003" * 18:37 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding urldownloader service IPs - sukhe@cumin1003" * 18:32 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 18:32 dancy@deploy2003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 18:30 sukhe@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 18:27 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 18:24 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1024.eqiad.wmnet with OS bookworm * 18:20 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host deploy1003.eqiad.wmnet with OS bookworm * 18:09 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deploy1003 reimage (duration: 121m 16s) * 18:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 18:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1180: Security updates * 17:55 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wcqs2003.codfw.wmnet * 17:48 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wcqs2003.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1155.eqiad.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1155.eqiad.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2224.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2224.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2217.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2217.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2193.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2193.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2180.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2180.codfw.wmnet * 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1168.eqiad.wmnet * 17:36 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1168.eqiad.wmnet * 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2169.codfw.wmnet * 17:36 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2169.codfw.wmnet * 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1165.eqiad.wmnet * 17:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1165.eqiad.wmnet * 17:35 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2158.codfw.wmnet * 17:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2158.codfw.wmnet * 17:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wcqs1003.eqiad.wmnet * 17:17 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1180: Security updates * 17:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1180.eqiad.wmnet * 17:16 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1180.eqiad.wmnet * 17:15 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp3073.* * 17:13 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wcqs1003.eqiad.wmnet * 17:11 brett@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp3073.esams.wmnet with OS trixie * 17:11 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 17:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1024 * 17:04 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1024 * 17:03 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 17:00 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 16:59 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2242: codfw rack B7 depool for maintenance * 16:59 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 16:43 brett@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp3073.esams.wmnet with reason: host reimage * 16:42 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 16:39 brett@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cp3073.esams.wmnet with reason: host reimage * 16:32 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on deploy1003.eqiad.wmnet with reason: host reimage * 16:27 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on deploy1003.eqiad.wmnet with reason: host reimage * 16:14 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2242: codfw rack B7 depool for maintenance * 16:14 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: codfw rack B7 depool for maintenance * 16:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1023.eqiad.wmnet with OS bookworm * 16:13 brett@cumin2002: START - Cookbook sre.hosts.reimage for host cp3073.esams.wmnet with OS trixie * 16:08 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host deploy1003.eqiad.wmnet with OS bookworm * 16:08 kamila@deploy2003: Locking from deployment [MediaWiki]: deploy1003 reimage * 16:03 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1012.eqiad.wmnet with OS bookworm * 15:48 inflatador: bking@apt1002 `sudo reprepro copy bookworm-wikimedia bullseye-wikimedia jvmquake` [[phab:T430880|T430880]] * 15:39 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp3073.* * 15:39 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 15:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:34 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1311536{{!}}Set $wgMathInternalRestbaseURL explicitly (take 2) (T349582)]] * 15:29 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:29 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2228: codfw rack B7 depool for maintenance * 15:29 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2229: codfw rack B7 depool for maintenance * 15:27 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1180: Security update * 15:25 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Security update * 15:21 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 15:21 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 15:19 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 24s) * 15:19 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:14 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 15:14 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db1180: Security update * 15:13 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-eqiad * 14:48 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-eqiad * 14:44 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2229: codfw rack B7 depool for maintenance * 14:44 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc2017: codfw rack B7 depool for maintenance * 14:44 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:43 cmooney@cumin2003: START - Cookbook sre.mysql.parsercache * 14:43 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool pc2017: codfw rack B7 depool for maintenance * 14:43 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2003.codfw.wmnet * 14:43 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2003.codfw.wmnet * 14:42 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2009.codfw.wmnet * 14:42 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2009.codfw.wmnet * 14:41 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:41 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:40 cmooney@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 29 hosts * 14:40 cmooney@cumin1003: START - Cookbook sre.hosts.remove-downtime for 29 hosts * 14:35 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 14:34 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 14:32 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] (duration: 07m 56s) * 14:29 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ssw1-a[1,8]-codfw with reason: lsw1-b7-codfw JunOS upgrade * 14:28 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 14:28 elukey@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 14:26 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:24 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] * 14:23 topranks: reboot lsw1-b7-codfw to upgrade JunOS (affects all hosts in rack) [[phab:T430928|T430928]] * 14:18 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2003.codfw.wmnet * 14:14 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2009.codfw.wmnet * 14:14 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Security update * 14:13 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2242: codfw rack B7 depool for maintenance * 14:13 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2242: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2228: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2228: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2229: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2229: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc2017: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.parsercache * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool pc2017: codfw rack B7 depool for maintenance * 14:08 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2003.codfw.wmnet * 14:07 cmooney@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on aux-k8s-etcd2004.codfw.wmnet,ml-etcd2001.codfw.wmnet with reason: lsw1-b7-codfw JunOS upgrade * 14:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95005 and previous config saved to /var/cache/conftool/dbconfig/20260721-140620-cwilliams.json * 14:05 cmooney@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2049.codfw.wmnet * 14:05 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-scholarly,name=eqiad * 14:04 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2009.codfw.wmnet * 14:04 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply * 14:04 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply * 14:03 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:03 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2001.codfw.wmnet * 14:03 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2001.codfw.wmnet * 14:02 cmooney@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2049.codfw.wmnet * 14:00 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:00 Dreamy_Jazz: Created cusi_case, cusi_signal, and cusi_user on svwiki, dewiki, jawiki, eswiki * 13:59 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b7-codfw,lsw1-b7-codfw IPv6,lsw1-b7-codfw.mgmt,ssw1-a[1,8]-codfw.mgmt with reason: lsw1-b7-codfw JunOS upgrade * 13:57 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs1023.eqiad.wmnet, repooling source-only afterwards * 13:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 29 hosts with reason: lsw1-b7-codfw JunOS upgrade * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224', diff saved to https://phabricator.wikimedia.org/P95003 and previous config saved to /var/cache/conftool/dbconfig/20260721-135613-cwilliams.json * 13:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1012.eqiad.wmnet with reason: host reimage * 13:53 cmooney@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 1:00:00 on 30 hosts with reason: lsw1-b7-codfw JunOS upgrade * 13:51 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1012.eqiad.wmnet with reason: host reimage * 13:48 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 13:48 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 13:46 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224', diff saved to https://phabricator.wikimedia.org/P95001 and previous config saved to /var/cache/conftool/dbconfig/20260721-134605-cwilliams.json * 13:46 cmooney@cumin1003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti2033.codfw.wmnet * 13:46 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 13:45 elukey: move the Docker Registry's /v2/wikimedia/machinelearning.* prefix to the ml S3 backend - [[phab:T428022|T428022]] * 13:45 cmooney@cumin1003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti2033.codfw.wmnet * 13:45 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 13:43 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:40 jiji@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 13:40 jiji@deploy2003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 13:39 jiji@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 13:39 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 13:38 cmooney@cumin1003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2032.codfw.wmnet * 13:38 jiji@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 13:38 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 13:37 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2032.codfw.wmnet * 13:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95000 and previous config saved to /var/cache/conftool/dbconfig/20260721-133557-cwilliams.json * 13:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1012 * 13:33 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1012 * 13:33 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1012.eqiad.wmnet with OS bookworm * 13:30 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:30 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:28 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94999 and previous config saved to /var/cache/conftool/dbconfig/20260721-132855-cwilliams.json * 13:28 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2224.codfw.wmnet with reason: Maintenance * 13:28 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94998 and previous config saved to /var/cache/conftool/dbconfig/20260721-132826-cwilliams.json * 13:28 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 13:23 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] (duration: 07m 50s) * 13:20 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 13:18 kharlan@deploy2003: kharlan: Continuing with deployment * 13:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217', diff saved to https://phabricator.wikimedia.org/P94996 and previous config saved to /var/cache/conftool/dbconfig/20260721-131817-cwilliams.json * 13:17 brouberol@dns1004: END - running authdns-update * 13:17 kharlan@deploy2003: kharlan: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:15 brouberol@dns1004: START - running authdns-update * 13:15 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] * 13:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94995 and previous config saved to /var/cache/conftool/dbconfig/20260721-131411-cwilliams.json * 13:13 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs1023.eqiad.wmnet, repooling source-only afterwards * 13:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217', diff saved to https://phabricator.wikimedia.org/P94994 and previous config saved to /var/cache/conftool/dbconfig/20260721-130809-cwilliams.json * 13:07 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 13:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180', diff saved to https://phabricator.wikimedia.org/P94993 and previous config saved to /var/cache/conftool/dbconfig/20260721-130404-cwilliams.json * 13:03 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:03 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:02 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 13:02 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 12:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94992 and previous config saved to /var/cache/conftool/dbconfig/20260721-125801-cwilliams.json * 12:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180', diff saved to https://phabricator.wikimedia.org/P94991 and previous config saved to /var/cache/conftool/dbconfig/20260721-125356-cwilliams.json * 12:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94990 and previous config saved to /var/cache/conftool/dbconfig/20260721-125049-cwilliams.json * 12:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2217.codfw.wmnet with reason: Maintenance * 12:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94989 and previous config saved to /var/cache/conftool/dbconfig/20260721-125017-cwilliams.json * 12:48 elukey: bmc cold reboot for lvs1013 and lvs1015 - [[phab:T426180|T426180]] * 12:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94988 and previous config saved to /var/cache/conftool/dbconfig/20260721-124348-cwilliams.json * 12:40 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193', diff saved to https://phabricator.wikimedia.org/P94987 and previous config saved to /var/cache/conftool/dbconfig/20260721-124009-cwilliams.json * 12:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts * 12:33 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts * 12:33 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts * 12:32 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts * 12:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:30 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193', diff saved to https://phabricator.wikimedia.org/P94986 and previous config saved to /var/cache/conftool/dbconfig/20260721-123001-cwilliams.json * 12:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94985 and previous config saved to /var/cache/conftool/dbconfig/20260721-121953-cwilliams.json * 12:17 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs2007.codfw.wmnet with OS bookworm * 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94983 and previous config saved to /var/cache/conftool/dbconfig/20260721-121257-cwilliams.json * 12:12 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2193.codfw.wmnet with reason: Maintenance * 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94982 and previous config saved to /var/cache/conftool/dbconfig/20260721-121239-cwilliams.json * 12:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180', diff saved to https://phabricator.wikimedia.org/P94980 and previous config saved to /var/cache/conftool/dbconfig/20260721-120231-cwilliams.json * 11:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180', diff saved to https://phabricator.wikimedia.org/P94979 and previous config saved to /var/cache/conftool/dbconfig/20260721-115223-cwilliams.json * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94978 and previous config saved to /var/cache/conftool/dbconfig/20260721-114333-cwilliams.json * 11:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1180.eqiad.wmnet with reason: Maintenance * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94977 and previous config saved to /var/cache/conftool/dbconfig/20260721-114305-cwilliams.json * 11:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94976 and previous config saved to /var/cache/conftool/dbconfig/20260721-114215-cwilliams.json * 11:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94975 and previous config saved to /var/cache/conftool/dbconfig/20260721-113530-cwilliams.json * 11:35 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2180.codfw.wmnet with reason: Maintenance * 11:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94974 and previous config saved to /var/cache/conftool/dbconfig/20260721-113501-cwilliams.json * 11:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168', diff saved to https://phabricator.wikimedia.org/P94973 and previous config saved to /var/cache/conftool/dbconfig/20260721-113258-cwilliams.json * 11:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169', diff saved to https://phabricator.wikimedia.org/P94972 and previous config saved to /var/cache/conftool/dbconfig/20260721-112453-cwilliams.json * 11:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168', diff saved to https://phabricator.wikimedia.org/P94971 and previous config saved to /var/cache/conftool/dbconfig/20260721-112250-cwilliams.json * 11:21 XioNoX: put eqiad-drmrs Arelion link in service * 11:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169', diff saved to https://phabricator.wikimedia.org/P94970 and previous config saved to /var/cache/conftool/dbconfig/20260721-111446-cwilliams.json * 11:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94969 and previous config saved to /var/cache/conftool/dbconfig/20260721-111242-cwilliams.json * 11:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 11:10 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 11:07 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1093 hosts * 11:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94968 and previous config saved to /var/cache/conftool/dbconfig/20260721-110548-cwilliams.json * 11:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1168.eqiad.wmnet with reason: Maintenance * 11:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94967 and previous config saved to /var/cache/conftool/dbconfig/20260721-110520-cwilliams.json * 11:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94966 and previous config saved to /var/cache/conftool/dbconfig/20260721-110439-cwilliams.json * 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94964 and previous config saved to /var/cache/conftool/dbconfig/20260721-105632-cwilliams.json * 10:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2169.codfw.wmnet with reason: Maintenance * 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94963 and previous config saved to /var/cache/conftool/dbconfig/20260721-105603-cwilliams.json * 10:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165', diff saved to https://phabricator.wikimedia.org/P94962 and previous config saved to /var/cache/conftool/dbconfig/20260721-105512-cwilliams.json * 10:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158', diff saved to https://phabricator.wikimedia.org/P94961 and previous config saved to /var/cache/conftool/dbconfig/20260721-104555-cwilliams.json * 10:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165', diff saved to https://phabricator.wikimedia.org/P94960 and previous config saved to /var/cache/conftool/dbconfig/20260721-104504-cwilliams.json * 10:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158', diff saved to https://phabricator.wikimedia.org/P94959 and previous config saved to /var/cache/conftool/dbconfig/20260721-103547-cwilliams.json * 10:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94958 and previous config saved to /var/cache/conftool/dbconfig/20260721-103456-cwilliams.json * 10:29 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2229: Upgraded kernel * 10:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94956 and previous config saved to /var/cache/conftool/dbconfig/20260721-102757-cwilliams.json * 10:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on an-redacteddb1001.eqiad.wmnet,clouddb[1015,1025,1028].eqiad.wmnet,db1155.eqiad.wmnet with reason: Maintenance * 10:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1165.eqiad.wmnet with reason: Maintenance * 10:25 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94955 and previous config saved to /var/cache/conftool/dbconfig/20260721-102539-cwilliams.json * 10:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94954 and previous config saved to /var/cache/conftool/dbconfig/20260721-101848-cwilliams.json * 10:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2158.codfw.wmnet with reason: Maintenance * 09:43 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2229: Upgraded kernel * 09:42 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2229.codfw.wmnet * 09:42 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2229.codfw.wmnet * 09:23 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db2229.codfw.wmnet * 09:23 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2229.codfw.wmnet * 08:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2229 [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94948 and previous config saved to /var/cache/conftool/dbconfig/20260721-085724-cwilliams.json * 08:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2214 to s6 primary [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94947 and previous config saved to /var/cache/conftool/dbconfig/20260721-085442-cwilliams.json * 08:53 cezmunsta: Starting s6 codfw failover from db2229 to db2214 - [[phab:T430964|T430964]] * 08:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2214 with weight 0 [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94946 and previous config saved to /var/cache/conftool/dbconfig/20260721-084613-cwilliams.json * 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 22 hosts with reason: Primary switchover s6 [[phab:T430964|T430964]] * 08:32 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1017.eqiad.wmnet,service=s1 * 08:08 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Add subrated circuit rate to interface descriptions - CR1312476 - ayounsi@cumin1003 * 08:06 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Add subrated circuit rate to interface descriptions - CR1312476 - ayounsi@cumin1003 * 07:58 reedy@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] (duration: 12m 55s) * 07:51 reedy@deploy2003: reedy, neriah: Continuing with deployment * 07:51 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 07:51 reedy@deploy2003: reedy, neriah: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:48 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1093 hosts * 07:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm2001.wikimedia.org * 07:45 reedy@deploy2003: Started scap sync-world: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] * 07:43 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 07:42 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm2001.wikimedia.org * 07:23 elukey: upgrade libtiff6 packages on zuul* trixie hosts for security upgrades * 07:22 elukey: upgrade libtiff6 packages on Wikikube trixie workers for security upgrades * 07:14 elukey@deploy2003: helmfile [codfw] DONE helmfile.d/services/proton: sync * 07:13 elukey@deploy2003: helmfile [codfw] START helmfile.d/services/proton: sync * 07:11 elukey@deploy2003: helmfile [eqiad] DONE helmfile.d/services/proton: sync * 07:10 elukey@deploy2003: helmfile [eqiad] START helmfile.d/services/proton: sync * 07:09 elukey@deploy2003: helmfile [staging] DONE helmfile.d/services/proton: sync * 07:08 elukey@deploy2003: helmfile [staging] START helmfile.d/services/proton: sync * 06:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1023.eqiad.wmnet with reason: host reimage * 06:46 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1023.eqiad.wmnet with reason: host reimage * 06:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 05:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Haproxy-only mode support - oblivian@cumin1003" * 05:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Haproxy-only mode support - oblivian@cumin1003 * 05:42 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Haproxy-only mode support - oblivian@cumin1003 * 05:42 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Haproxy-only mode support - oblivian@cumin1003" * 05:38 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1017.eqiad.wmnet with reason: Cloning * 05:37 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1017.eqiad.wmnet,service=s1 * 05:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1029.eqiad.wmnet,service=s8 * 05:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1029.eqiad.wmnet,service=s5 * 05:32 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:30 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 05:11 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet * 05:04 aokoth@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet * 05:00 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 04:56 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 04:01 mwpresync@deploy2003: Pruned MediaWiki: 1.47.0-wmf.9 (duration: 01m 08s) * 03:41 mwpresync@deploy2003: Finished scap sync-world: testwikis to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] (duration: 36m 30s) * 03:05 mwpresync@deploy2003: Started scap sync-world: testwikis to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 03:01 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:01 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:00 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:00 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:36 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:36 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:36 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:35 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:16 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 47s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 00:56 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm == 2026-07-20 == * 23:38 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 23:07 Amir1: deleting echo notifications from 2015 on group1 wikis * 23:07 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] (duration: 14m 16s) * 23:01 ladsgroup@deploy2003: ladsgroup: Continuing with deployment * 23:00 ladsgroup@deploy2003: ladsgroup: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:53 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] * 22:46 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2007.codfw.wmnet, repooling source-only afterwards * 22:39 maryum: Deployed security fixes for several security bugs * 21:42 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 21:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2007.codfw.wmnet, repooling source-only afterwards * 21:37 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 21:37 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 21:34 sbassett: Deployed security fix for [[phab:T432424|T432424]] * 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs2020.codfw.wmnet * 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1023.eqiad.wmnet * 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1011.eqiad.wmnet * 21:32 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 17s) * 21:32 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 21:27 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 21:13 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2007.codfw.wmnet with reason: host reimage * 21:08 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-internal-main,name=codfw * 21:06 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2007.codfw.wmnet with reason: host reimage * 20:59 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service * 20:58 sukhe: pybal restart for IP changes around wdqs-main hosts * 20:57 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 20:46 ryankemper@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-internal-main,name=codfw * 20:45 ebernhardson@deploy2003: Finished deploy [search/mjolnir/deploy@d4dc3b8]: Update for opensearch 2.x compat (duration: 00m 34s) * 20:44 ebernhardson@deploy2003: Started deploy [search/mjolnir/deploy@d4dc3b8]: Update for opensearch 2.x compat * 20:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2007 * 20:44 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2007 * 20:43 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2007 * 20:43 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2007.codfw.wmnet 156.16.192.10.in-addr.arpa 6.5.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:42 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2007.codfw.wmnet 156.16.192.10.in-addr.arpa 6.5.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:42 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2007 - bking@cumin2003" * 20:41 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2007 - bking@cumin2003" * 20:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94944 and previous config saved to /var/cache/conftool/dbconfig/20260720-203333-cwilliams.json * 20:32 arlolra@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] (duration: 15m 07s) * 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2020.codfw.wmnet * 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1023.eqiad.wmnet * 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1011.eqiad.wmnet * 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs2020.codfw.wmnet * 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1023.eqiad.wmnet * 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1011.eqiad.wmnet * 20:25 arlolra@deploy2003: arlolra, cscott: Continuing with deployment * 20:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257', diff saved to https://phabricator.wikimedia.org/P94943 and previous config saved to /var/cache/conftool/dbconfig/20260720-202325-cwilliams.json * 20:21 arlolra@deploy2003: arlolra, cscott: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:17 arlolra@deploy2003: Started scap sync-world: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] * 20:13 bking@cumin2003: START - Cookbook sre.dns.netbox * 20:13 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257', diff saved to https://phabricator.wikimedia.org/P94942 and previous config saved to /var/cache/conftool/dbconfig/20260720-201318-cwilliams.json * 20:13 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 20:10 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 20:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2007 * 20:04 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2007.codfw.wmnet with OS bookworm * 20:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94941 and previous config saved to /var/cache/conftool/dbconfig/20260720-200310-cwilliams.json * 19:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94940 and previous config saved to /var/cache/conftool/dbconfig/20260720-195633-cwilliams.json * 19:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1257.eqiad.wmnet with reason: Maintenance * 19:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94939 and previous config saved to /var/cache/conftool/dbconfig/20260720-195605-cwilliams.json * 19:51 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 19:50 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 19:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256', diff saved to https://phabricator.wikimedia.org/P94938 and previous config saved to /var/cache/conftool/dbconfig/20260720-194558-cwilliams.json * 19:44 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 19:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 19:41 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wdqs1011.eqiad.wmnet with OS bookworm * 19:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256', diff saved to https://phabricator.wikimedia.org/P94937 and previous config saved to /var/cache/conftool/dbconfig/20260720-193550-cwilliams.json * 19:25 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94936 and previous config saved to /var/cache/conftool/dbconfig/20260720-192542-cwilliams.json * 19:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94935 and previous config saved to /var/cache/conftool/dbconfig/20260720-191856-cwilliams.json * 19:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1256.eqiad.wmnet with reason: Maintenance * 19:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94934 and previous config saved to /var/cache/conftool/dbconfig/20260720-191839-cwilliams.json * 19:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255', diff saved to https://phabricator.wikimedia.org/P94933 and previous config saved to /var/cache/conftool/dbconfig/20260720-190831-cwilliams.json * 18:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255', diff saved to https://phabricator.wikimedia.org/P94932 and previous config saved to /var/cache/conftool/dbconfig/20260720-185824-cwilliams.json * 18:50 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 18:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94931 and previous config saved to /var/cache/conftool/dbconfig/20260720-184816-cwilliams.json * 18:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94930 and previous config saved to /var/cache/conftool/dbconfig/20260720-184224-cwilliams.json * 18:42 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1255.eqiad.wmnet with reason: Maintenance * 18:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94929 and previous config saved to /var/cache/conftool/dbconfig/20260720-184153-cwilliams.json * 18:39 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 18:39 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 16s) * 18:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 18:38 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 59m 26s) * 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 18:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211', diff saved to https://phabricator.wikimedia.org/P94928 and previous config saved to /var/cache/conftool/dbconfig/20260720-183145-cwilliams.json * 18:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211', diff saved to https://phabricator.wikimedia.org/P94927 and previous config saved to /var/cache/conftool/dbconfig/20260720-182137-cwilliams.json * 18:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94926 and previous config saved to /var/cache/conftool/dbconfig/20260720-181129-cwilliams.json * 18:09 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_codfw * 18:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2057.codfw.wmnet * 18:08 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_codfw * 18:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2058.codfw.wmnet * 18:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94925 and previous config saved to /var/cache/conftool/dbconfig/20260720-180452-cwilliams.json * 18:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on clouddb[1016,1020,1022-1023].eqiad.wmnet,db1154.eqiad.wmnet with reason: Maintenance * 18:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1211.eqiad.wmnet with reason: Maintenance * 18:02 sukhe: armed keyholder on acmechief1002.eqiad.wmnet and acmechief2002.codfw.wmnet (active host) * 18:01 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief2002.codfw.wmnet * 17:57 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief2002.codfw.wmnet * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs2020'] * 17:52 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief1002.eqiad.wmnet * 17:50 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 17:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1011.eqiad.wmnet with reason: host reimage * 17:48 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief1002.eqiad.wmnet * 17:47 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test2001.codfw.wmnet * 17:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94924 and previous config saved to /var/cache/conftool/dbconfig/20260720-174717-cwilliams.json * 17:46 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 17:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1011.eqiad.wmnet with reason: host reimage * 17:43 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs2020.codfw.wmnet with OS bookworm * 17:43 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test2001.codfw.wmnet * 17:43 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test1001.eqiad.wmnet * 17:39 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test1001.eqiad.wmnet * 17:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 17:38 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:38 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 17:37 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 17:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244', diff saved to https://phabricator.wikimedia.org/P94923 and previous config saved to /var/cache/conftool/dbconfig/20260720-173709-cwilliams.json * 17:35 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 17:31 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 20m 40s) * 17:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2055.codfw.wmnet * 17:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2056.codfw.wmnet * 17:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1011 * 17:27 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1011 * 17:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1011.eqiad.wmnet with OS bookworm * 17:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244', diff saved to https://phabricator.wikimedia.org/P94922 and previous config saved to /var/cache/conftool/dbconfig/20260720-172701-cwilliams.json * 17:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94921 and previous config saved to /var/cache/conftool/dbconfig/20260720-171653-cwilliams.json * 17:11 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 17:11 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 13m 03s) * 17:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94920 and previous config saved to /var/cache/conftool/dbconfig/20260720-171012-cwilliams.json * 17:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2244.codfw.wmnet with reason: Maintenance * 17:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94919 and previous config saved to /var/cache/conftool/dbconfig/20260720-170941-cwilliams.json * 16:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243', diff saved to https://phabricator.wikimedia.org/P94918 and previous config saved to /var/cache/conftool/dbconfig/20260720-165933-cwilliams.json * 16:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 16:58 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2053.codfw.wmnet * 16:51 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2054.codfw.wmnet * 16:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243', diff saved to https://phabricator.wikimedia.org/P94917 and previous config saved to /var/cache/conftool/dbconfig/20260720-164926-cwilliams.json * 16:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94916 and previous config saved to /var/cache/conftool/dbconfig/20260720-163918-cwilliams.json * 16:35 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 16:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94915 and previous config saved to /var/cache/conftool/dbconfig/20260720-163140-cwilliams.json * 16:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2243.codfw.wmnet with reason: Maintenance * 16:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94914 and previous config saved to /var/cache/conftool/dbconfig/20260720-163111-cwilliams.json * 16:27 btullis@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 16:27 btullis@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 16:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2020 * 16:23 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2020 * 16:21 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2020 * 16:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2020.codfw.wmnet 85.0.192.10.in-addr.arpa 5.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:21 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2020.codfw.wmnet 85.0.192.10.in-addr.arpa 5.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242', diff saved to https://phabricator.wikimedia.org/P94913 and previous config saved to /var/cache/conftool/dbconfig/20260720-162103-cwilliams.json * 16:19 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 16:18 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 16:18 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:18 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 16:17 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:17 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for netbox accounting errors - jhancock@cumin2002" * 16:17 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for netbox accounting errors - jhancock@cumin2002" * 16:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2051.codfw.wmnet * 16:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2052.codfw.wmnet * 16:11 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 16:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242', diff saved to https://phabricator.wikimedia.org/P94912 and previous config saved to /var/cache/conftool/dbconfig/20260720-161055-cwilliams.json * 16:09 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 16:08 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 16:06 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 16:06 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 16:06 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2020 * 16:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2020.codfw.wmnet with OS bookworm * 16:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94911 and previous config saved to /var/cache/conftool/dbconfig/20260720-160047-cwilliams.json * 15:58 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2019.codfw.wmnet, repooling source-only afterwards * 15:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94909 and previous config saved to /var/cache/conftool/dbconfig/20260720-155353-cwilliams.json * 15:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2242.codfw.wmnet with reason: Maintenance * 15:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94908 and previous config saved to /var/cache/conftool/dbconfig/20260720-154433-cwilliams.json * 15:35 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2049.codfw.wmnet * 15:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162', diff saved to https://phabricator.wikimedia.org/P94907 and previous config saved to /var/cache/conftool/dbconfig/20260720-153425-cwilliams.json * 15:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2050.codfw.wmnet * 15:28 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162', diff saved to https://phabricator.wikimedia.org/P94906 and previous config saved to /var/cache/conftool/dbconfig/20260720-152418-cwilliams.json * 15:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1023 * 15:14 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1023 * 15:14 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 15:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94905 and previous config saved to /var/cache/conftool/dbconfig/20260720-151407-cwilliams.json * 15:13 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] (duration: 41m 16s) * 15:08 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 15:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94902 and previous config saved to /var/cache/conftool/dbconfig/20260720-150729-cwilliams.json * 15:07 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2162.codfw.wmnet with reason: Maintenance * 15:05 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2027.codfw.wmnet, repooling source-only afterwards * 15:00 urbanecm@deploy2003: vadymts1, migr, urbanecm: Continuing with deployment * 14:59 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:58 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 07s) * 14:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:58 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 13s) * 14:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:57 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2019.codfw.wmnet, repooling source-only afterwards * 14:57 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2047.codfw.wmnet * 14:55 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2048.codfw.wmnet * 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2019.codfw.wmnet with OS bookworm * 14:49 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 14:47 urbanecm@deploy2003: vadymts1, migr, urbanecm: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:44 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool magru [reason: BGP issues in lvs7003 resolved after liberica restart, no task ID specified] * 14:44 sukhe@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool magru [reason: BGP issues in lvs7003 resolved after liberica restart, no task ID specified] * 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:41 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:39 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:39 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:33 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool magru [reason: no reason specified, no task ID specified] * 14:33 sukhe@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool magru [reason: no reason specified, no task ID specified] * 14:31 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] * 14:24 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:24 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:24 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:24 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2019.codfw.wmnet with reason: host reimage * 14:22 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2027.codfw.wmnet, repooling source-only afterwards * 14:19 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2019.codfw.wmnet with reason: host reimage * 14:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2027.codfw.wmnet with OS bookworm * 14:16 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2046.codfw.wmnet * 14:16 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2045.codfw.wmnet * 14:08 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:08 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:08 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:08 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:07 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:06 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:06 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:06 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:05 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2019 * 14:00 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2019 * 13:56 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1015.eqiad.wmnet * 13:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2027.codfw.wmnet with reason: host reimage * 13:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2071.codfw.wmnet with OS trixie * 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:51 sukhe@cumin1003: END (ERROR) - Cookbook sre.loadbalancer.admin (exit_code=97) rebooting A:liberica and P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica and P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:51 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1015.eqiad.wmnet * 13:50 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1014.eqiad.wmnet * 13:50 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1076.eqiad.wmnet with OS trixie * 13:50 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2027.codfw.wmnet with reason: host reimage * 13:45 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1014.eqiad.wmnet * 13:44 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1013.eqiad.wmnet * 13:39 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 13:38 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1013.eqiad.wmnet * 13:37 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2044.codfw.wmnet * 13:37 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2043.codfw.wmnet * 13:36 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2019 * 13:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2019.codfw.wmnet 156.32.192.10.in-addr.arpa 6.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:36 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2019.codfw.wmnet 156.32.192.10.in-addr.arpa 6.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:36 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2019 - bking@cumin2003" * 13:36 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2019 - bking@cumin2003" * 13:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry2005.codfw.wmnet * 13:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2071.codfw.wmnet with reason: host reimage * 13:31 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:31 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2019 * 13:31 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry2005.codfw.wmnet * 13:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry2004.codfw.wmnet * 13:30 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2019.codfw.wmnet with OS bookworm * 13:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2027 * 13:30 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2027 * 13:30 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2027.codfw.wmnet with OS bookworm * 13:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1076.eqiad.wmnet with reason: host reimage * 13:29 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_codfw * 13:28 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_codfw * 13:26 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry2004.codfw.wmnet * 13:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry1005.eqiad.wmnet * 13:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2071.codfw.wmnet with reason: host reimage * 13:22 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1076.eqiad.wmnet with reason: host reimage * 13:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry1005.eqiad.wmnet * 13:21 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry1004.eqiad.wmnet * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry1004.eqiad.wmnet * 13:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts * 13:13 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts * 13:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts * 13:12 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts * 13:03 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1076.eqiad.wmnet with OS trixie * 13:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2071.codfw.wmnet with OS trixie * 12:55 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:54 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:53 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:46 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 7 hosts * 12:42 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 7 hosts * 12:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts * 12:42 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts * 12:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2070.codfw.wmnet with OS trixie * 12:36 ozge@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:35 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1075.eqiad.wmnet with OS trixie * 12:32 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts * 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts * 12:22 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin1001.eqiad.wmnet * 12:19 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin1001.eqiad.wmnet * 12:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2070.codfw.wmnet with reason: host reimage * 12:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin2001.codfw.wmnet * 12:14 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1075.eqiad.wmnet with reason: host reimage * 12:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2070.codfw.wmnet with reason: host reimage * 12:10 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1075.eqiad.wmnet with reason: host reimage * 12:09 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin2001.codfw.wmnet * 11:17 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1074.eqiad.wmnet with OS trixie * 11:17 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 11:16 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 11:14 ozge@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 6 hosts * 11:09 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 6 hosts * 11:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 324 hosts * 10:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1074.eqiad.wmnet with reason: host reimage * 10:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2069.codfw.wmnet with OS trixie * 10:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1074.eqiad.wmnet with reason: host reimage * 10:30 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2069.codfw.wmnet with reason: host reimage * 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1074.eqiad.wmnet with OS trixie * 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2069.codfw.wmnet with reason: host reimage * 10:06 blake@deploy2003: Stopping before sync operations * 10:06 blake@deploy2003: Started scap sync-world: Non-deployment scap run to populate new release values * 10:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2069.codfw.wmnet with OS trixie * 10:00 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1073.eqiad.wmnet with OS trixie * 09:56 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 324 hosts * 09:39 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 09:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 8 hosts * 09:38 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1073.eqiad.wmnet with reason: host reimage * 09:37 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 8 hosts * 09:34 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1073.eqiad.wmnet with reason: host reimage * 09:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2068.codfw.wmnet with OS trixie * 09:16 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1073.eqiad.wmnet with OS trixie * 09:13 blake@deploy2003: sync-world aborted: Non-deployment scap run to populate new release values (duration: 00m 02s) * 09:13 blake@deploy2003: Started scap sync-world: Non-deployment scap run to populate new release values * 08:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2068.codfw.wmnet with reason: host reimage * 08:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2068.codfw.wmnet with reason: host reimage * 08:50 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 08:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2068.codfw.wmnet with OS trixie * 08:15 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1072.eqiad.wmnet with OS trixie * 07:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2067.codfw.wmnet with OS trixie * 07:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1072.eqiad.wmnet with reason: host reimage * 07:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1072.eqiad.wmnet with reason: host reimage * 07:45 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 07:45 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 07:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2067.codfw.wmnet with reason: host reimage * 07:35 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2067.codfw.wmnet with reason: host reimage * 07:30 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 07:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1072.eqiad.wmnet with OS trixie * 07:30 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 07:17 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 07:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2067.codfw.wmnet with OS trixie * 05:51 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:50 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:25 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:25 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on db2207.codfw.wmnet with reason: Host down * 04:28 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 07m 02s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-18 == * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 29s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 00:11 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2018.codfw.wmnet, repooling source-only afterwards == 2026-07-17 == * 23:53 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2026.codfw.wmnet, repooling source-only afterwards * 23:09 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2018.codfw.wmnet, repooling source-only afterwards * 23:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2026.codfw.wmnet, repooling source-only afterwards * 22:11 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2018.codfw.wmnet with OS bookworm * 22:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2026.codfw.wmnet with OS bookworm * 21:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2018.codfw.wmnet with reason: host reimage * 21:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2018.codfw.wmnet with reason: host reimage * 21:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2026.codfw.wmnet with reason: host reimage * 21:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2026.codfw.wmnet with reason: host reimage * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2018 * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2018 * 21:26 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2018 * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2018.codfw.wmnet 155.32.192.10.in-addr.arpa 5.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:26 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2018.codfw.wmnet 155.32.192.10.in-addr.arpa 5.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2018 - bking@cumin2003" * 21:26 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2018 - bking@cumin2003" * 21:14 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:13 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2018 * 21:13 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2018.codfw.wmnet with OS bookworm * 21:12 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2026 * 21:12 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2026 * 21:12 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2026.codfw.wmnet with OS bookworm * 21:05 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1022.eqiad.wmnet -> wdqs1026.eqiad.wmnet, repooling source-only afterwards * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs2017.codfw.wmnet, repooling source-only afterwards * 20:11 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs2017.codfw.wmnet, repooling source-only afterwards * 20:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2017.codfw.wmnet with OS bookworm * 20:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1022.eqiad.wmnet -> wdqs1026.eqiad.wmnet, repooling source-only afterwards * 20:06 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1026.eqiad.wmnet with OS bookworm * 19:55 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 19:55 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:55 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 09s) * 19:55 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:50 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 08s) * 19:50 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:50 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 10m 03s) * 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2017.codfw.wmnet with reason: host reimage * 19:40 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:40 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 15s) * 19:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1026.eqiad.wmnet with reason: host reimage * 19:37 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 16s) * 19:37 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:34 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2017.codfw.wmnet with reason: host reimage * 19:34 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1026.eqiad.wmnet with reason: host reimage * 19:33 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 19:33 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:16 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1026.eqiad.wmnet with OS bookworm * 19:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2017 * 19:16 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2017 * 19:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2017.codfw.wmnet with OS bookworm * 18:30 bking@dns1004: END - running authdns-update * 18:28 bking@dns1004: START - running authdns-update * 18:16 kamila@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1264.eqiad.wmnet * 18:16 kamila@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1264.eqiad.wmnet * 18:16 kamila@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1264.eqiad.wmnet * 17:49 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 17:46 dzahn@dns1006: END - running authdns-update * 17:44 dzahn@dns1006: START - running authdns-update * 17:44 dzahn@dns1006: END - running authdns-update * 17:42 dzahn@dns1006: START - running authdns-update * 17:28 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 17:21 kamila@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 17:01 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1264 * 17:01 kamila@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1264 * 17:01 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 17:01 kamila@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1264.eqiad.wmnet * 17:01 kamila@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1264.eqiad.wmnet * 17:01 kamila@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1264.eqiad.wmnet * 16:42 reedy@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] (duration: 10m 29s) * 16:34 reedy@deploy2003: reedy, hartman: Continuing with deployment * 16:33 reedy@deploy2003: reedy, hartman: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:31 reedy@deploy2003: Started scap sync-world: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] * 16:26 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 16:10 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in2001.wikimedia.org with reason: [[phab:T431659|T431659]] * 16:07 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in1001.wikimedia.org with reason: [[phab:T431659|T431659]] * 16:05 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 16:01 kamila@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 16:00 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out2001.wikimedia.org with reason: [[phab:T431659|T431659]] * 15:41 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 15:41 kamila@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 15:35 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out1001.wikimedia.org with reason: [[phab:T431659|T431659]] * 15:14 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1339.eqiad.wmnet * 15:13 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1339.eqiad.wmnet * 15:13 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1339.eqiad.wmnet * 14:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1339.eqiad.wmnet with OS trixie * 14:50 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:49 kamila@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:49 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:33 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage * 14:27 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage * 14:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1339 * 14:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1339 * 14:14 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1339 * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1339.eqiad.wmnet 156.32.64.10.in-addr.arpa 6.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:14 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1339.eqiad.wmnet 156.32.64.10.in-addr.arpa 6.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1339 - cgoubert@cumin2003" * 14:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1339 - cgoubert@cumin2003" * 14:09 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 14:06 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1339 * 14:06 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie * 14:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1339.eqiad.wmnet * 14:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1339.eqiad.wmnet * 14:02 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1339.eqiad.wmnet * 13:45 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb1013.eqiad.wmnet * 13:39 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb1013.eqiad.wmnet * 13:27 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:24 blake@dns1004: END - running authdns-update * 13:22 blake@dns1004: START - running authdns-update * 13:20 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 13:11 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2014.codfw.wmnet * 13:06 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb2014.codfw.wmnet * 13:06 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2012.codfw.wmnet * 13:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 13:03 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 13:01 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 15 hosts * 13:01 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb2012.codfw.wmnet * 13:01 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1016.eqiad.wmnet * 13:00 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 15 hosts * 12:55 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb1016.eqiad.wmnet * 12:55 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1014.eqiad.wmnet * 12:49 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb1014.eqiad.wmnet * 12:32 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:32 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:31 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:31 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1338.eqiad.wmnet * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1338.eqiad.wmnet * 12:18 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1338.eqiad.wmnet * 12:17 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:15 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:14 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:13 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1338.eqiad.wmnet with OS trixie * 12:01 klausman@deploy2003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 11:59 klausman@deploy2003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 11:56 klausman@deploy2003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 11:54 klausman@deploy2003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 11:53 klausman@deploy2003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 11:51 klausman@deploy2003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 11:42 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1338.eqiad.wmnet with reason: host reimage * 11:38 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1338.eqiad.wmnet with reason: host reimage * 11:31 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2230.codfw.wmnet * 11:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1338 * 11:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1338 * 11:25 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1338 * 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1338.eqiad.wmnet 155.32.64.10.in-addr.arpa 5.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:25 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1338.eqiad.wmnet 155.32.64.10.in-addr.arpa 5.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1338 - cgoubert@cumin2003" * 11:25 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1338 - cgoubert@cumin2003" * 11:23 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2230.codfw.wmnet * 11:20 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 11:20 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1338 * 11:20 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1338.eqiad.wmnet with OS trixie * 11:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1338.eqiad.wmnet * 11:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1338.eqiad.wmnet * 11:19 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1338.eqiad.wmnet * 11:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1337.eqiad.wmnet * 11:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1337.eqiad.wmnet * 11:17 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1337.eqiad.wmnet * 11:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1337.eqiad.wmnet with OS trixie * 10:51 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[2001-2002].codfw.wmnet * 10:50 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1337.eqiad.wmnet with reason: host reimage * 10:40 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 10:39 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:39 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1337.eqiad.wmnet with reason: host reimage * 10:39 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1001-1003].eqiad.wmnet * 10:34 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:34 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:30 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:28 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1001-1003].eqiad.wmnet * 10:27 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1337 * 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1337 * 10:26 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1337 * 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1337.eqiad.wmnet 154.32.64.10.in-addr.arpa 4.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:26 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1337.eqiad.wmnet 154.32.64.10.in-addr.arpa 4.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1337 - cgoubert@cumin2003" * 10:26 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1337 - cgoubert@cumin2003" * 10:21 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 10:18 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1337 * 10:17 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1337.eqiad.wmnet with OS trixie * 10:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1337.eqiad.wmnet * 10:16 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db1176.eqiad.wmnet * 10:16 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1337.eqiad.wmnet * 10:16 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1337.eqiad.wmnet * 10:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1336.eqiad.wmnet * 10:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1336.eqiad.wmnet * 10:15 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1336.eqiad.wmnet * 10:11 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db1176.eqiad.wmnet * 10:10 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db1176.eqiad.wmnet * 10:09 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db1176.eqiad.wmnet * 10:05 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts (check the cookbook's logs for more details.) * 10:03 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts (check the cookbook's logs for more details.) * 09:58 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1336.eqiad.wmnet with OS trixie * 09:47 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts (check the cookbook's logs for more details.) * 09:47 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts (check the cookbook's logs for more details.) * 09:45 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host acmechief-test2001.codfw.wmnet,acmechief-test1001.eqiad.wmnet,an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet,db-test[2001-2002].codfw.wmnet,db-test[1 * 09:40 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host acmechief-test2001.codfw.wmnet,acmechief-test1001.eqiad.wmnet,an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet,db-test[2001-2002].codfw.wmnet,db-test[1001-1003].eqiad.wmn * 09:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1336.eqiad.wmnet with reason: host reimage * 09:33 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1336.eqiad.wmnet with reason: host reimage * 09:29 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet * 09:29 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet * 09:28 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:26 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:21 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 09:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1336 * 09:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1336 * 09:19 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 09:14 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1336 * 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1336.eqiad.wmnet 152.32.64.10.in-addr.arpa 2.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1336.eqiad.wmnet 152.32.64.10.in-addr.arpa 2.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1336 - cgoubert@cumin2003" * 09:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1336 - cgoubert@cumin2003" * 09:11 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:10 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 09:09 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:09 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1336 * 09:09 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1336.eqiad.wmnet with OS trixie * 09:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1336.eqiad.wmnet * 09:08 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1336.eqiad.wmnet * 09:08 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1336.eqiad.wmnet * 09:06 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1335.eqiad.wmnet * 09:06 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1335.eqiad.wmnet * 09:06 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1335.eqiad.wmnet * 09:04 elukey: uploaded spicerack_13.1.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia * 08:55 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wikikube-worker-exp2001.codfw.wmnet * 08:54 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host testreduce1002.eqiad.wmnet * 08:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1335.eqiad.wmnet with OS trixie * 08:51 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host wikikube-worker-exp2001.codfw.wmnet * 08:51 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wikikube-worker-exp1001.eqiad.wmnet * 08:50 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host testreduce1002.eqiad.wmnet * 08:45 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host wikikube-worker-exp1001.eqiad.wmnet * 08:34 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1335.eqiad.wmnet with reason: host reimage * 08:30 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1335.eqiad.wmnet with reason: host reimage * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1335 * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1335 * 08:18 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1335 * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1335.eqiad.wmnet 150.32.64.10.in-addr.arpa 0.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:18 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1335.eqiad.wmnet 150.32.64.10.in-addr.arpa 0.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1335 - cgoubert@cumin2003" * 08:18 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1335 - cgoubert@cumin2003" * 08:14 elukey@cumin1003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 08:14 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges * 08:13 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 08:10 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1335 * 08:10 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1335.eqiad.wmnet with OS trixie * 08:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1335.eqiad.wmnet * 08:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1335.eqiad.wmnet * 08:09 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1335.eqiad.wmnet * 08:06 elukey@cumin1003: END (FAIL) - Cookbook sre.puppet.disable-merges (exit_code=99) * 08:05 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges * 08:03 elukey@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin1003.eqiad.wmnet * 07:57 elukey@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin1003.eqiad.wmnet * 07:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetdb1003.eqiad.wmnet * 07:46 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetdb1003.eqiad.wmnet * 07:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetdb2003.codfw.wmnet * 07:37 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetdb2003.codfw.wmnet * 07:37 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1001.eqiad.wmnet * 07:28 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver1001.eqiad.wmnet * 07:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet * 07:19 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet * 07:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2002.codfw.wmnet * 07:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver2002.codfw.wmnet * 07:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet * 07:05 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet * 07:04 btullis@cumin1003: END (FAIL) - Cookbook sre.hadoop.reboot-workers (exit_code=99) for Hadoop analytics cluster * 07:04 elukey@cumin1003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 07:04 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges * 06:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox1003.eqiad.wmnet * 06:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox1003.eqiad.wmnet * 02:46 ryankemper@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:46 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:44 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:37 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:37 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-internal-scholarly,name=eqiad * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 49s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 01:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore wdqs1025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling source-only afterwards * 01:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore wdqs1027 after Bookworm reimage) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs1027.eqiad.wmnet, repooling both afterwards * 00:55 urbanecm@deploy2003: helmfile [codfw] DONE helmfile.d/services/linkrecommendation: apply * 00:54 urbanecm@deploy2003: helmfile [eqiad] DONE helmfile.d/services/linkrecommendation: apply * 00:54 urbanecm@deploy2003: helmfile [staging] DONE helmfile.d/services/linkrecommendation: apply * 00:54 urbanecm@deploy2003: helmfile [codfw] START helmfile.d/services/linkrecommendation: apply * 00:53 urbanecm@deploy2003: helmfile [staging] START helmfile.d/services/linkrecommendation: apply * 00:52 urbanecm@deploy2003: helmfile [eqiad] START helmfile.d/services/linkrecommendation: apply * 00:23 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore wdqs1025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling source-only afterwards * 00:23 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore wdqs1027 after Bookworm reimage) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs1027.eqiad.wmnet, repooling both afterwards * 00:14 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1274.eqiad.wmnet * 00:14 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1274.eqiad.wmnet * 00:14 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1274.eqiad.wmnet * 00:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1027.eqiad.wmnet with OS bookworm * 00:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1025.eqiad.wmnet with OS bookworm * 00:04 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1274.eqiad.wmnet with OS trixie == 2026-07-16 == * 23:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], xfer to freshly reimaged/scap-deployed wdqs2025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs2025.codfw.wmnet, repooling source-only afterwards * 23:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1027.eqiad.wmnet with reason: host reimage * 23:47 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1025.eqiad.wmnet with reason: host reimage * 23:43 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1274.eqiad.wmnet with reason: host reimage * 23:41 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1025.eqiad.wmnet with reason: host reimage * 23:39 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1027.eqiad.wmnet with reason: host reimage * 23:38 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1274.eqiad.wmnet with reason: host reimage * 23:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1025 * 23:23 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1025 * 23:22 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1027 * 23:22 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1027 * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1274 * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1274 * 23:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1025.eqiad.wmnet with OS bookworm * 23:19 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1274 * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1274.eqiad.wmnet 145.48.64.10.in-addr.arpa 5.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:19 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1274.eqiad.wmnet 145.48.64.10.in-addr.arpa 5.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1274 - swfrench@cumin1003" * 23:19 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1274 - swfrench@cumin1003" * 23:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1027.eqiad.wmnet with OS bookworm * 23:14 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 23:14 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1274 * 23:13 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1274.eqiad.wmnet with OS trixie * 23:13 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1274.eqiad.wmnet * 23:12 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1274.eqiad.wmnet * 23:12 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1274.eqiad.wmnet * 23:12 ryankemper: [[phab:T430880|T430880]] depooled dnsdisc of wdqs-internal-scholarly-eqiad bc we only have 1 host there * 23:09 ryankemper@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-internal-scholarly,name=eqiad * 23:08 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1272.eqiad.wmnet * 23:08 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1272.eqiad.wmnet * 23:08 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1272.eqiad.wmnet * 23:01 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], xfer to freshly reimaged/scap-deployed wdqs2025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs2025.codfw.wmnet, repooling source-only afterwards * 22:57 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1272.eqiad.wmnet with OS trixie * 22:56 Amir1: deleting echo notifications from 2015 in group0 * 22:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2025.codfw.wmnet with OS bookworm * 22:35 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1272.eqiad.wmnet with reason: host reimage * 22:32 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 27s) * 22:32 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 22:28 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1269.eqiad.wmnet * 22:28 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1269.eqiad.wmnet * 22:28 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1269.eqiad.wmnet * 22:27 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1272.eqiad.wmnet with reason: host reimage * 22:26 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] (duration: 08m 51s) * 22:22 ladsgroup@deploy2003: ladsgroup, urbanecm: Continuing with deployment * 22:19 ladsgroup@deploy2003: ladsgroup, urbanecm: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:17 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] * 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2025.codfw.wmnet with reason: host reimage * 22:06 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1272 * 22:06 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1272 * 22:05 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1272 * 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1272.eqiad.wmnet 127.48.64.10.in-addr.arpa 7.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:05 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1272.eqiad.wmnet 127.48.64.10.in-addr.arpa 7.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1272 - swfrench@cumin1003" * 22:05 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1272 - swfrench@cumin1003" * 22:03 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2025.codfw.wmnet with reason: host reimage * 22:01 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 22:00 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1272 * 22:00 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1272.eqiad.wmnet with OS trixie * 22:00 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1272.eqiad.wmnet * 21:59 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1272.eqiad.wmnet * 21:59 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1272.eqiad.wmnet * 21:55 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1271.eqiad.wmnet * 21:55 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1271.eqiad.wmnet * 21:55 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1271.eqiad.wmnet * 21:46 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1271.eqiad.wmnet with OS trixie * 21:45 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] (duration: 06m 31s) * 21:40 sbassett@deploy2003: sbassett: Continuing with deployment * 21:40 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2025 * 21:40 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2025 * 21:40 sbassett@deploy2003: sbassett: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:38 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] * 21:37 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2025 * 21:37 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2025.codfw.wmnet 220.48.192.10.in-addr.arpa 0.2.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:37 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2025.codfw.wmnet 220.48.192.10.in-addr.arpa 0.2.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:37 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:37 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2025 - bking@cumin2003" * 21:37 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2025 - bking@cumin2003" * 21:30 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] (duration: 08m 19s) * 21:26 sbassett@deploy2003: sbassett: Continuing with deployment * 21:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1269.eqiad.wmnet with OS trixie * 21:24 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1271.eqiad.wmnet with reason: host reimage * 21:23 sbassett@deploy2003: sbassett: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:22 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:22 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] * 21:20 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2025 * 21:19 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2025.codfw.wmnet with OS bookworm * 21:17 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1271.eqiad.wmnet with reason: host reimage * 21:04 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1269.eqiad.wmnet with reason: host reimage * 21:00 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1269.eqiad.wmnet with reason: host reimage * 20:56 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1271 * 20:55 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1271 * 20:54 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1271 * 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1271.eqiad.wmnet 126.48.64.10.in-addr.arpa 6.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:54 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1271.eqiad.wmnet 126.48.64.10.in-addr.arpa 6.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1271 - swfrench@cumin1003" * 20:54 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1271 - swfrench@cumin1003" * 20:51 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:51 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:51 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:50 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 20:49 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 20:49 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1268.eqiad.wmnet * 20:49 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1268.eqiad.wmnet * 20:49 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1268.eqiad.wmnet * 20:48 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1271 * 20:48 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1271.eqiad.wmnet with OS trixie * 20:47 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1271.eqiad.wmnet * 20:46 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1271.eqiad.wmnet * 20:46 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1271.eqiad.wmnet * 20:41 aude@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] (duration: 07m 34s) * 20:39 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1269 * 20:39 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1269 * 20:36 aude@deploy2003: aude: Continuing with deployment * 20:35 aude@deploy2003: aude: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:33 aude@deploy2003: Started scap sync-world: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] * 20:26 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on db2207.codfw.wmnet with reason: Host down * 20:22 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-video: apply * 20:21 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-video: apply * 20:20 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-timeline: apply * 20:20 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-timeline: apply * 20:20 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-syntaxhighlight: apply * 20:19 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-syntaxhighlight: apply * 20:19 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-media: apply * 20:18 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-media: apply * 20:18 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-constraints: apply * 20:17 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-constraints: apply * 20:17 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox: apply * 20:16 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox: apply * 20:13 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1269 * 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1269.eqiad.wmnet 80.32.64.10.in-addr.arpa 0.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:13 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1269.eqiad.wmnet 80.32.64.10.in-addr.arpa 0.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1269 - kamila@cumin1003" * 20:13 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1269 - kamila@cumin1003" * 20:09 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-flink-codfw cluster: Roll restart of jvm daemons. * 20:07 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 20:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 20:03 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-flink-codfw cluster: Roll restart of jvm daemons. * 20:03 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2207 [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94893 and previous config saved to /var/cache/conftool/dbconfig/20260716-200257-marostegui.json * 20:01 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2204 to s2 primary [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94892 and previous config saved to /var/cache/conftool/dbconfig/20260716-200157-marostegui.json * 20:00 marostegui: Starting emergency s2 codfw failover from db2207 to db2204 - [[phab:T432396|T432396]] * 19:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1035.eqiad.wmnet * 19:56 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2204 with weight 0 [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94891 and previous config saved to /var/cache/conftool/dbconfig/20260716-195628-marostegui.json * 19:55 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 26 hosts with reason: Primary switchover s2 [[phab:T432396|T432396]] * 19:54 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1035.eqiad.wmnet * 19:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1034.eqiad.wmnet * 19:48 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1034.eqiad.wmnet * 19:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1033.eqiad.wmnet * 19:43 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-video: apply * 19:43 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1033.eqiad.wmnet * 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1032.eqiad.wmnet * 19:42 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-video: apply * 19:42 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-timeline: apply * 19:41 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-timeline: apply * 19:41 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-syntaxhighlight: apply * 19:41 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-syntaxhighlight: apply * 19:40 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-media: apply * 19:40 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-media: apply * 19:39 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-constraints: apply * 19:36 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-constraints: apply * 19:36 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox: apply * 19:35 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1032.eqiad.wmnet * 19:35 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1031.eqiad.wmnet * 19:35 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox: apply * 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-video: apply * 19:33 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-video: apply * 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-timeline: apply * 19:33 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-timeline: apply * 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-syntaxhighlight: apply * 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-syntaxhighlight: apply * 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-media: apply * 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-media: apply * 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-constraints: apply * 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-constraints: apply * 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox: apply * 19:31 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox: apply * 19:27 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1031.eqiad.wmnet * 19:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1030.eqiad.wmnet * 19:23 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1001.eqiad.wmnet, repooling source-only afterwards * 19:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1030.eqiad.wmnet * 19:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1029.eqiad.wmnet * 19:17 kamila@cumin1003: START - Cookbook sre.dns.netbox * 19:12 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1029.eqiad.wmnet * 19:06 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1269 * 19:05 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1269.eqiad.wmnet with OS trixie * 19:03 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1269.eqiad.wmnet * 19:03 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1269.eqiad.wmnet * 19:03 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1269.eqiad.wmnet * 18:55 dancy@deploy2003: Finished scap sync-world: testing [[phab:T428971|T428971]] (duration: 02m 41s) * 18:53 dancy@deploy2003: Started scap sync-world: testing [[phab:T428971|T428971]] * 18:31 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1268.eqiad.wmnet with OS trixie * 18:18 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 18:16 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1267.eqiad.wmnet * 18:16 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1267.eqiad.wmnet * 18:16 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1267.eqiad.wmnet * 18:09 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1268.eqiad.wmnet with reason: host reimage * 18:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1001.eqiad.wmnet, repooling source-only afterwards * 18:06 swfrench@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] (duration: 07m 34s) * 18:06 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 23s) * 18:06 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 18:06 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1268.eqiad.wmnet with reason: host reimage * 18:03 bd808@deploy2003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 18:02 bd808@deploy2003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 18:02 swfrench@deploy2003: jiji, swfrench: Continuing with deployment * 18:02 bd808@deploy2003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 18:02 bd808@deploy2003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 18:01 bd808@deploy2003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 18:01 swfrench@deploy2003: jiji, swfrench: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:01 bd808@deploy2003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:59 swfrench@deploy2003: Started scap sync-world: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] * 17:45 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1268 * 17:45 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1268 * 17:44 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1267.eqiad.wmnet with OS trixie * 17:43 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter2006.codfw.wmnet * 17:39 swfrench@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter2006.codfw.wmnet * 17:35 swfrench@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] (duration: 07m 27s) * 17:34 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1268 * 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1268.eqiad.wmnet 78.32.64.10.in-addr.arpa 8.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:34 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1268.eqiad.wmnet 78.32.64.10.in-addr.arpa 8.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1268 - kamila@cumin1003" * 17:34 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1268 - kamila@cumin1003" * 17:31 swfrench@deploy2003: jiji, swfrench: Continuing with deployment * 17:29 swfrench@deploy2003: jiji, swfrench: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:28 kamila@cumin1003: START - Cookbook sre.dns.netbox * 17:28 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1268 * 17:28 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1268.eqiad.wmnet with OS trixie * 17:27 swfrench@deploy2003: Started scap sync-world: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] * 17:23 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1267.eqiad.wmnet with reason: host reimage * 17:18 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1267.eqiad.wmnet with reason: host reimage * 17:18 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1268.eqiad.wmnet * 17:17 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1268.eqiad.wmnet * 17:17 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1268.eqiad.wmnet * 17:12 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter2005.codfw.wmnet * 17:11 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1270.eqiad.wmnet * 17:11 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1270.eqiad.wmnet * 17:11 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1270.eqiad.wmnet * 17:09 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter2005.codfw.wmnet * 17:08 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:08 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update reverse dns for moved arelion cct cr2-eqiad - cmooney@cumin1003" * 17:08 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update reverse dns for moved arelion cct cr2-eqiad - cmooney@cumin1003" * 17:08 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] (duration: 07m 34s) * 17:04 jiji@deploy2003: jiji: Continuing with deployment * 17:03 jiji@deploy2003: jiji: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 17:00 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] * 17:00 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2185.codfw.wmnet with OS trixie * 16:59 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:58 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1270.eqiad.wmnet with OS trixie * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1267 * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1267 * 16:57 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1267 * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1267.eqiad.wmnet 77.32.64.10.in-addr.arpa 7.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:57 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1267.eqiad.wmnet 77.32.64.10.in-addr.arpa 7.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1267 - kamila@cumin1003" * 16:56 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1267 - kamila@cumin1003" * 16:56 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_eqsin * 16:56 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5032.eqsin.wmnet * 16:52 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_esams * 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3073.esams.wmnet * 16:50 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_esams * 16:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3081.esams.wmnet * 16:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1266.eqiad.wmnet * 16:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1266.eqiad.wmnet * 16:45 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1266.eqiad.wmnet * 16:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2185.codfw.wmnet with reason: host reimage * 16:41 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_eqiad * 16:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1114.eqiad.wmnet * 16:41 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_eqiad * 16:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1115.eqiad.wmnet * 16:39 kamila@cumin1003: START - Cookbook sre.dns.netbox * 16:39 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1267 * 16:39 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2185.codfw.wmnet with reason: host reimage * 16:38 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1267.eqiad.wmnet with OS trixie * 16:38 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1267.eqiad.wmnet * 16:38 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1270.eqiad.wmnet with reason: host reimage * 16:37 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1267.eqiad.wmnet * 16:37 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1267.eqiad.wmnet * 16:31 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1270.eqiad.wmnet with reason: host reimage * 16:24 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1264.eqiad.wmnet * 16:24 kamila@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 16:24 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 16:23 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 16:21 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2162: switch maintenance completed codfw rack b6 * 16:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2185.codfw.wmnet with OS trixie * 16:19 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 16:16 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter1007.eqiad.wmnet * 16:15 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_eqsin * 16:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5024.eqsin.wmnet * 16:13 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5031.eqsin.wmnet * 16:13 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3072.esams.wmnet * 16:12 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter1007.eqiad.wmnet * 16:11 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] (duration: 09m 47s) * 16:10 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1266.eqiad.wmnet with OS trixie * 16:10 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1270 * 16:10 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1270 * 16:09 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1270 * 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1270.eqiad.wmnet 125.48.64.10.in-addr.arpa 5.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1270.eqiad.wmnet 125.48.64.10.in-addr.arpa 5.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1270 - swfrench@cumin1003" * 16:09 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1270 - swfrench@cumin1003" * 16:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3080.esams.wmnet * 16:07 jiji@deploy2003: jiji: Continuing with deployment * 16:06 jiji@deploy2003: jiji: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:04 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 16:04 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1265.eqiad.wmnet * 16:03 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1265.eqiad.wmnet * 16:03 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1265.eqiad.wmnet * 16:03 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1270 * 16:03 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1270.eqiad.wmnet with OS trixie * 16:02 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1270.eqiad.wmnet * 16:02 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] * 16:01 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1270.eqiad.wmnet * 16:01 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1270.eqiad.wmnet * 16:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1113.eqiad.wmnet * 16:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1112.eqiad.wmnet * 15:49 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1266.eqiad.wmnet with reason: host reimage * 15:47 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter1006.eqiad.wmnet * 15:45 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1265.eqiad.wmnet with OS trixie * 15:44 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1266.eqiad.wmnet with reason: host reimage * 15:43 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter1006.eqiad.wmnet * 15:42 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] (duration: 09m 46s) * 15:37 jiji@deploy2003: jiji: Continuing with deployment * 15:36 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2162: switch maintenance completed codfw rack b6 * 15:36 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2161: switch maintenance completed codfw rack b6 * 15:34 jiji@deploy2003: jiji: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:32 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5023.eqsin.wmnet * 15:32 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] * 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5030.eqsin.wmnet * 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3071.esams.wmnet * 15:27 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3079.esams.wmnet * 15:25 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1265.eqiad.wmnet with reason: host reimage * 15:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1266 * 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1266 * 15:21 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1110.eqiad.wmnet * 15:20 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1111.eqiad.wmnet * 15:16 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1265.eqiad.wmnet with reason: host reimage * 15:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1001.eqiad.wmnet with OS bookworm * 15:15 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1266 * 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1266.eqiad.wmnet 76.32.64.10.in-addr.arpa 6.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:15 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1266.eqiad.wmnet 76.32.64.10.in-addr.arpa 6.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1266 - kamila@cumin1003" * 15:15 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1266 - kamila@cumin1003" * 15:07 kamila@cumin1003: START - Cookbook sre.dns.netbox * 15:04 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1266 * 15:04 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1264 * 15:04 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1264 * 15:04 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1266.eqiad.wmnet with OS trixie * 15:03 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1264 * 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1264.eqiad.wmnet 74.32.64.10.in-addr.arpa 4.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:03 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1264.eqiad.wmnet 74.32.64.10.in-addr.arpa 4.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1264 - kamila@cumin1003" * 15:03 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1264 - kamila@cumin1003" * 15:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-eqiad * 15:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp1001.eqiad.wmnet * 15:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp1001.eqiad.wmnet * 15:01 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp1001.eqiad.wmnet * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp1001.eqiad.wmnet * 15:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1376-1384].eqiad.wmnet * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1376-1384].eqiad.wmnet * 14:59 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1002.eqiad.wmnet * 14:59 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1266.eqiad.wmnet * 14:58 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1266.eqiad.wmnet * 14:58 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1266.eqiad.wmnet * 14:58 kamila@cumin1003: START - Cookbook sre.dns.netbox * 14:57 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1264 * 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1265 * 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1265 * 14:57 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1310596{{!}}Set $wgMathInternalRestbaseURL explicitly (T349582)]] * 14:57 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1265 * 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1265.eqiad.wmnet 75.32.64.10.in-addr.arpa 5.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:56 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1265.eqiad.wmnet 75.32.64.10.in-addr.arpa 5.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:56 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:56 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1265 - kamila@cumin1003" * 14:56 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1265 - kamila@cumin1003" * 14:53 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1002.eqiad.wmnet * 14:53 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1376-1384].eqiad.wmnet * 14:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1001.eqiad.wmnet with reason: host reimage * 14:51 kamila@cumin1003: START - Cookbook sre.dns.netbox * 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 14:50 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:50 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2161: switch maintenance completed codfw rack b6 * 14:50 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5021.eqsin.wmnet * 14:50 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1265 * 14:49 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:49 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:49 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1265.eqiad.wmnet with OS trixie * 14:49 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5029.eqsin.wmnet * 14:49 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1265.eqiad.wmnet * 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3070.esams.wmnet * 14:48 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1264.eqiad.wmnet * 14:48 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1001.eqiad.wmnet with reason: host reimage * 14:48 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1376-1384].eqiad.wmnet * 14:48 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1265.eqiad.wmnet * 14:47 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1265.eqiad.wmnet * 14:47 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1264.eqiad.wmnet * 14:47 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1264.eqiad.wmnet * 14:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:47 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3078.esams.wmnet * 14:44 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1263.eqiad.wmnet * 14:44 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1263.eqiad.wmnet * 14:44 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1263.eqiad.wmnet * 14:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1108.eqiad.wmnet * 14:40 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1109.eqiad.wmnet * 14:40 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:35 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:34 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:34 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2006.codfw.wmnet * 14:34 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-flink-eqiad cluster: Roll restart of jvm daemons. * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf2002.codfw.wmnet * 14:31 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1002.eqiad.wmnet * 14:29 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2006.codfw.wmnet * 14:27 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:27 kamila@deploy2003: Finished scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] (duration: 02m 57s) * 14:27 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-flink-eqiad cluster: Roll restart of jvm daemons. * 14:26 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf2002.codfw.wmnet * 14:26 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf2001.codfw.wmnet * 14:25 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1002.eqiad.wmnet * 14:25 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1001.eqiad.wmnet * 14:25 kamila@deploy2003: Started scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] * 14:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:21 kamila@deploy2003: sync-world aborted: Test deployment to check rsync is working - [[phab:T432108|T432108]] (duration: 00m 36s) * 14:21 topranks: reboot lsw1-b6-codfw to upgrade JunOS [[phab:T430922|T430922]] * 14:21 kamila@deploy2003: Started scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] * 14:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1001.eqiad.wmnet with OS bookworm * 14:20 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b6-codfw,lsw1-b6-codfw IPv6,lsw1-b6-codfw.mgmt,ssw1-a[1,8]-codfw with reason: lsw1-b6-codfw JunOS upgrade * 14:20 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf2001.codfw.wmnet * 14:19 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1001.eqiad.wmnet * 14:19 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 26 hosts with reason: lsw1-b6-codfw JunOS upgrade * 14:14 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 14:13 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc2022: switch maintenance codfw rack b6 * 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:12 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.parsercache * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool pc2022: switch maintenance codfw rack b6 * 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2251: switch maintenance codfw rack b6 * 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.parsercache * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2251: switch maintenance codfw rack b6 * 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2162: switch maintenance codfw rack b6 * 14:12 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1263.eqiad.wmnet with OS trixie * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2162: switch maintenance codfw rack b6 * 14:11 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2161: switch maintenance codfw rack b6 * 14:11 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2161: switch maintenance codfw rack b6 * 14:08 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5020.eqsin.wmnet * 14:07 btullis@cumin1003: START - Cookbook sre.hadoop.reboot-workers for Hadoop analytics cluster * 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3069.esams.wmnet * 14:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5028.eqsin.wmnet * 14:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1338-1347].eqiad.wmnet * 14:06 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1338-1347].eqiad.wmnet * 14:05 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3077.esams.wmnet * 14:02 topranks: beginning depools for lsw1-b6-codfw maintenance [[phab:T430922|T430922]] * 14:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1106.eqiad.wmnet * 14:00 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-misc1002.eqiad.wmnet * 13:59 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1338-1347].eqiad.wmnet * 13:59 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1107.eqiad.wmnet * 13:56 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-codfw * 13:55 sfaci@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply * 13:54 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-misc1002.eqiad.wmnet * 13:54 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-misc1001.eqiad.wmnet * 13:54 sfaci@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply * 13:50 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1263.eqiad.wmnet with reason: host reimage * 13:50 sfaci@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 13:49 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1338-1347].eqiad.wmnet * 13:49 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-misc1001.eqiad.wmnet * 13:49 sfaci@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 13:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:49 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:45 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1263.eqiad.wmnet with reason: host reimage * 13:40 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:40 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-eqiad * 13:35 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:34 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:33 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-reboot (exit_code=0) rolling reboot on A:dnsbox and (A:eqsin or A:drmrs or A:magru) and not (P<nowiki>{</nowiki>dns5003*<nowiki>}</nowiki> or P<nowiki>{</nowiki>dns7002*<nowiki>}</nowiki>) and (A:dnsbox) * 13:33 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns7001.wikimedia.org * 13:27 sukhe@dns1004: END - running authdns-update * 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5019.eqsin.wmnet * 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3076.esams.wmnet * 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3068.esams.wmnet * 13:25 sukhe@dns1004: START - running authdns-update * 13:24 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5027.eqsin.wmnet * 13:24 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1263 * 13:24 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1263 * 13:23 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1263 * 13:23 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:23 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1104.eqiad.wmnet * 13:21 kamila@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:21 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:20 kamila@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:20 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:20 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:20 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1263 - kamila@cumin1003" * 13:20 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1263 - kamila@cumin1003" * 13:19 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1105.eqiad.wmnet * 13:19 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:19 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:18 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:18 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns7001.wikimedia.org * 13:16 cdobbins@cumin2003: conftool action : set/pooled=yes; selector: name=dns7002.* * 13:14 cdobbins@dns1004: END - running authdns-update * 13:13 sbisson@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] (duration: 08m 03s) * 13:13 cdobbins@dns1004: START - running authdns-update * 13:12 kamila@cumin1003: START - Cookbook sre.dns.netbox * 13:12 cdobbins@cumin2003: conftool action : set/pooled=yes; selector: name=dns7002.*,service=authdns-update * 13:12 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1263 * 13:11 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1263.eqiad.wmnet with OS trixie * 13:11 cdobbins@cumin2003: conftool action : set/pooled=no; selector: name=dns7002.* * 13:11 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1263.eqiad.wmnet * 13:10 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1263.eqiad.wmnet * 13:10 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1263.eqiad.wmnet * 13:09 sbisson@deploy2003: sbisson: Continuing with deployment * 13:08 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:07 sbisson@deploy2003: sbisson: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:05 sbisson@deploy2003: Started scap sync-world: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] * 13:03 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns6002.wikimedia.org * 13:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1298-1307].eqiad.wmnet * 13:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1298-1307].eqiad.wmnet * 12:59 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1262.eqiad.wmnet * 12:59 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1262.eqiad.wmnet * 12:59 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1262.eqiad.wmnet * 12:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1298-1307].eqiad.wmnet * 12:49 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns6002.wikimedia.org * 12:46 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1298-1307].eqiad.wmnet * 12:46 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:46 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3075.esams.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3067.esams.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5018.eqsin.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5026.eqsin.wmnet * 12:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1102.eqiad.wmnet * 12:39 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1103.eqiad.wmnet * 12:35 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:34 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns6001.wikimedia.org * 12:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:28 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:18 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns6001.wikimedia.org * 12:14 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1267-1276].eqiad.wmnet * 12:13 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1267-1276].eqiad.wmnet * 12:04 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1267-1276].eqiad.wmnet * 12:03 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns5004.wikimedia.org * 12:02 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1100.eqiad.wmnet * 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3066.esams.wmnet * 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3074.esams.wmnet * 12:01 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:01 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-codfw * 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5017.eqsin.wmnet * 12:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5025.eqsin.wmnet * 12:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1101.eqiad.wmnet * 11:59 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1267-1276].eqiad.wmnet * 11:58 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:58 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:54 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns5004.wikimedia.org * 11:54 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and (A:eqsin or A:drmrs or A:magru) and not (P<nowiki>{</nowiki>dns5003*<nowiki>}</nowiki> or P<nowiki>{</nowiki>dns7002*<nowiki>}</nowiki>) and (A:dnsbox) * 11:54 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:53 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:53 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-eqiad * 11:51 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:51 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:50 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:50 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_eqiad * 11:50 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:50 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_eqiad * 11:49 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_esams * 11:49 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_esams * 11:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_eqsin * 11:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_eqsin * 11:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:44 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-eqiad * 11:43 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-codfw * 11:42 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:41 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:24 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-eqiad * 11:23 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-codfw * 11:22 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:15 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:14 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:09 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2066.codfw.wmnet with OS trixie * 11:05 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1151-1160].eqiad.wmnet * 11:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1151-1160].eqiad.wmnet * 10:59 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1068.eqiad.wmnet with OS trixie * 10:55 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.major-upgrade (exit_code=99) * 10:55 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 10:54 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1151-1160].eqiad.wmnet * 10:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2066.codfw.wmnet with reason: host reimage * 10:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1151-1160].eqiad.wmnet * 10:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:42 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2066.codfw.wmnet with reason: host reimage * 10:39 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:37 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:36 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:23 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:22 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2066.codfw.wmnet with OS trixie * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:07 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2065.codfw.wmnet with OS trixie * 10:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:06 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 10:06 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 10:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 10:03 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:03 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 10:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 09:59 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 09:57 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 09:57 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 09:52 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 09:47 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:46 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:46 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2065.codfw.wmnet with reason: host reimage * 09:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2065.codfw.wmnet with reason: host reimage * 09:40 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:39 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 09:39 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:39 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:39 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:37 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:29 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox2003.codfw.wmnet * 09:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox2003.codfw.wmnet * 09:25 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:25 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:24 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:24 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:24 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:21 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:20 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2065.codfw.wmnet with OS trixie * 09:13 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1068.eqiad.wmnet with OS trixie * 09:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2064.codfw.wmnet with OS trixie * 09:08 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 09:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 09:07 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2162: Repooling after switchover * 09:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1067.eqiad.wmnet with OS trixie * 09:00 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:59 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 08:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 08:57 tappof: bump space for prometheus k8s-dse in eqiad * 08:56 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping2004.codfw.wmnet * 08:52 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host ping2004.codfw.wmnet * 08:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 08:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping1004.eqiad.wmnet * 08:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:51 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:49 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 08:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host ping1004.eqiad.wmnet * 08:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2064.codfw.wmnet with reason: host reimage * 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2064.codfw.wmnet with reason: host reimage * 08:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1067.eqiad.wmnet with reason: host reimage * 08:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:33 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1067.eqiad.wmnet with reason: host reimage * 08:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:21 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2162: Repooling after switchover * 08:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2064.codfw.wmnet with OS trixie * 08:16 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1067.eqiad.wmnet with OS trixie * 08:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:15 cgoubert@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-eqiad * 08:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2062.codfw.wmnet with OS trixie * 08:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1066.eqiad.wmnet with OS trixie * 08:02 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2162: Repooling after switchover * 07:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2162: Repooling after switchover * 07:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2162 [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94870 and previous config saved to /var/cache/conftool/dbconfig/20260716-075530-cwilliams.json * 07:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2241 to x3 primary [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94869 and previous config saved to /var/cache/conftool/dbconfig/20260716-075314-cwilliams.json * 07:52 cezmunsta: Starting x3 codfw failover from db2162 to db2241 - [[phab:T430925|T430925]] * 07:50 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:50 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:47 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 07:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2241 with weight 0 [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94868 and previous config saved to /var/cache/conftool/dbconfig/20260716-074507-cwilliams.json * 07:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 18 hosts with reason: Primary switchover x3 [[phab:T430925|T430925]] * 07:43 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1066.eqiad.wmnet with reason: host reimage * 07:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:dse-k8s-worker-eqiad * 07:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1028.eqiad.wmnet * 07:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1028.eqiad.wmnet * 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 07:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1066.eqiad.wmnet with reason: host reimage * 07:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1028.eqiad.wmnet * 07:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1028.eqiad.wmnet * 07:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1027.eqiad.wmnet * 07:35 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1027.eqiad.wmnet * 07:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1027.eqiad.wmnet * 07:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1027.eqiad.wmnet * 07:28 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1026.eqiad.wmnet * 07:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1026.eqiad.wmnet * 07:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast2003.wikimedia.org * 07:21 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1026.eqiad.wmnet * 07:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1066.eqiad.wmnet with OS trixie * 07:19 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast2003.wikimedia.org * 07:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2062.codfw.wmnet with OS trixie * 06:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1026.eqiad.wmnet * 06:51 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1025.eqiad.wmnet * 06:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1025.eqiad.wmnet * 06:47 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 06:47 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 06:44 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1025.eqiad.wmnet * 06:14 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1025.eqiad.wmnet * 06:14 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1024.eqiad.wmnet * 06:14 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1024.eqiad.wmnet * 06:07 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1024.eqiad.wmnet * 05:37 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1024.eqiad.wmnet * 05:37 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1023.eqiad.wmnet * 05:37 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1023.eqiad.wmnet * 05:26 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1023.eqiad.wmnet * 04:56 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1023.eqiad.wmnet * 04:56 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1022.eqiad.wmnet * 04:56 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1022.eqiad.wmnet * 04:49 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1022.eqiad.wmnet * 04:19 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1022.eqiad.wmnet * 04:19 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1021.eqiad.wmnet * 04:19 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1021.eqiad.wmnet * 04:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1021.eqiad.wmnet * 03:38 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1021.eqiad.wmnet * 03:38 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1020.eqiad.wmnet * 03:38 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1020.eqiad.wmnet * 03:20 btullis@cumin1003: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1020.eqiad.wmnet * 03:18 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1020.eqiad.wmnet * 03:18 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1019.eqiad.wmnet * 03:18 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1019.eqiad.wmnet * 03:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1019.eqiad.wmnet * 02:41 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1019.eqiad.wmnet * 02:41 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1018.eqiad.wmnet * 02:41 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1018.eqiad.wmnet * 02:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling both afterwards * 02:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2003.codfw.wmnet -> wcqs2001.codfw.wmnet, repooling both afterwards * 02:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1018.eqiad.wmnet * 02:30 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1018.eqiad.wmnet * 02:30 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1014.eqiad.wmnet * 02:30 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1014.eqiad.wmnet * 02:24 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1014.eqiad.wmnet * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 01:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1014.eqiad.wmnet * 01:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1013.eqiad.wmnet * 01:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1013.eqiad.wmnet * 01:47 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1013.eqiad.wmnet * 01:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2003.codfw.wmnet -> wcqs2001.codfw.wmnet, repooling both afterwards * 01:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling both afterwards * 01:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1013.eqiad.wmnet * 01:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1012.eqiad.wmnet * 01:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1012.eqiad.wmnet * 01:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1012.eqiad.wmnet * 01:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1012.eqiad.wmnet * 01:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1011.eqiad.wmnet * 01:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1011.eqiad.wmnet * 01:08 ryankemper@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] scap deploy post bookworm reimage (duration: 00m 23s) * 01:08 ryankemper@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] scap deploy post bookworm reimage * 01:08 ryankemper@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): scap deploy post bookworm reimage (duration: 00m 46s) * 01:07 ryankemper@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): scap deploy post bookworm reimage * 01:04 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1011.eqiad.wmnet * 01:04 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1011.eqiad.wmnet * 01:04 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1010.eqiad.wmnet * 01:04 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1010.eqiad.wmnet * 00:57 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1010.eqiad.wmnet * 00:57 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1010.eqiad.wmnet * 00:57 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1009.eqiad.wmnet * 00:57 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1009.eqiad.wmnet * 00:50 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1009.eqiad.wmnet * 00:20 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1009.eqiad.wmnet * 00:20 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1008.eqiad.wmnet * 00:20 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1008.eqiad.wmnet * 00:13 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1008.eqiad.wmnet == 2026-07-15 == * 23:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2001.codfw.wmnet with OS bookworm * 23:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1008.eqiad.wmnet * 23:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1007.eqiad.wmnet * 23:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1007.eqiad.wmnet * 23:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1007.eqiad.wmnet * 23:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1007.eqiad.wmnet * 23:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1006.eqiad.wmnet * 23:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1006.eqiad.wmnet * 23:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1006.eqiad.wmnet * 23:29 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1006.eqiad.wmnet * 23:28 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1005.eqiad.wmnet * 23:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1005.eqiad.wmnet * 23:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1002.eqiad.wmnet with OS bookworm * 23:21 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1005.eqiad.wmnet * 23:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2001.codfw.wmnet with reason: host reimage * 23:15 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host datahubsearch1001.eqiad.wmnet with OS bookworm * 23:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2001.codfw.wmnet with reason: host reimage * 23:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 23:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 22:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 22:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1005.eqiad.wmnet * 22:51 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1004.eqiad.wmnet * 22:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1004.eqiad.wmnet * 22:45 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1004.eqiad.wmnet * 22:44 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 22:44 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS trixie * 22:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host datahubsearch1001.eqiad.wmnet with OS bookworm * 22:34 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host datahubsearch1001.eqiad.wmnet with OS bookworm * 22:16 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on datahubsearch[1002-1003].eqiad.wmnet with reason: Using datahubsearch1001 to test bookworm reimages * 22:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1004.eqiad.wmnet * 22:15 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1003.eqiad.wmnet * 22:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1003.eqiad.wmnet * 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 22:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1003.eqiad.wmnet * 22:08 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1003.eqiad.wmnet * 22:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1002.eqiad.wmnet * 22:08 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1002.eqiad.wmnet * 22:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host datahubsearch1001.eqiad.wmnet with OS bookworm * 22:05 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 22:02 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm * 22:01 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on datahubsearch[1001-1003].eqiad.wmnet with reason: Using datahubsearch1001 to test bookworm reimages * 22:01 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1002.eqiad.wmnet * 22:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1002.eqiad.wmnet * 22:00 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1001.eqiad.wmnet * 22:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1001.eqiad.wmnet * 21:53 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1001.eqiad.wmnet * 21:52 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 21:50 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 21:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS trixie * 21:50 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS bookworm * 21:43 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 21:38 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:30 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wcqs1002'] * 21:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:29 lerickson@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 21:29 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:29 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:29 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS bookworm * 21:28 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 21:28 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm * 21:23 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1001.eqiad.wmnet * 21:23 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:23 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:22 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 21:20 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:18 swfrench-wmf: reprepro include php8.3_8.3.32-1+wmf11u2 into component/php83 for bullseye-wikimedia * 21:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:16 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:15 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-druid-public cluster: Roll restart of jvm daemons. * 21:08 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:05 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1001.eqiad.wmnet * 21:05 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1001.eqiad.wmnet * 21:04 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-druid-public cluster: Roll restart of jvm daemons. * 21:02 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 21:01 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 21:01 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 21:00 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 20:59 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1001.eqiad.wmnet * 20:59 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1001.eqiad.wmnet * 20:59 btullis@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:dse-k8s-worker-eqiad * 20:55 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 20:55 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 20:45 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm * 20:21 jhathaway: puppet is re-enabled, have fun, but not too much fun! * 20:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 20:17 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs2001'] * 20:12 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs2001'] * 20:11 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs2001'] * 20:09 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:08 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 20:05 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:05 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 20:04 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs2001'] * 20:03 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 20:03 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm * 20:02 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:02 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 20:01 jhathaway: disabling puppet fleet wide to roll out kafka patch * 19:55 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host relforge1010.eqiad.wmnet * 19:52 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 19:52 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 19:48 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 19:48 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 19:48 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 19:47 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 19:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 19:45 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 19:45 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 19:44 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1010.eqiad.wmnet * 19:38 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1262.eqiad.wmnet with OS trixie * 19:17 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1262.eqiad.wmnet with reason: host reimage * 19:11 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1262.eqiad.wmnet with reason: host reimage * 18:59 cdobbins@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS trixie * 18:54 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 18:53 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 18:52 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1262 * 18:52 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1262 * 18:51 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1262 * 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1262.eqiad.wmnet 72.32.64.10.in-addr.arpa 2.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:51 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1262.eqiad.wmnet 72.32.64.10.in-addr.arpa 2.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1262 - kamila@cumin1003" * 18:51 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1262 - kamila@cumin1003" * 18:46 kamila@cumin1003: START - Cookbook sre.dns.netbox * 18:46 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1262 * 18:46 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 18:46 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ncmonitor1001.eqiad.wmnet * 18:46 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 18:45 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1262.eqiad.wmnet with OS trixie * 18:45 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 18:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1262.eqiad.wmnet * 18:44 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1262.eqiad.wmnet * 18:44 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1262.eqiad.wmnet * 18:42 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host ncmonitor1001.eqiad.wmnet * 18:29 topranks: pull power on cr1-eqiad to install new switch-control boards [[phab:T426343|T426343]] * 18:29 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs[1018-1020].eqiad.wmnet with reason: line card install in cr1-eqiad * 18:27 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 14 hosts with reason: linecard install in cr1-eqad * 18:22 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_ulsfo * 18:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4052.ulsfo.wmnet * 18:19 cdobbins@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 18:15 cdobbins@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 18:14 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_drmrs * 18:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6016.drmrs.wmnet * 18:12 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_ulsfo * 18:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4044.ulsfo.wmnet * 18:10 sukhe@cumin1003: END (ERROR) - Cookbook sre.cdn.roll-reboot (exit_code=97) rolling reboot on A:cp-upload_drmrs * 18:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2241: Security update * 17:56 topranks: start draining traffic on cr1-eqiad ahead of line card installation [[phab:T426343|T426343]] * 17:47 cdobbins@cumin2003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie * 17:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4051.ulsfo.wmnet * 17:40 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:39 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 17:34 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6007.drmrs.wmnet * 17:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6015.drmrs.wmnet * 17:32 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:31 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 17:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4043.ulsfo.wmnet * 17:27 lerickson@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:25 lerickson@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 17:22 lerickson@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-codfw * 17:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp2001.codfw.wmnet * 17:22 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 17:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp2001.codfw.wmnet * 17:22 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 17:19 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2241: Security update * 17:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2241.codfw.wmnet * 17:17 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2241.codfw.wmnet * 17:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp2001.codfw.wmnet * 17:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp2001.codfw.wmnet * 17:15 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2366-2374].codfw.wmnet * 17:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2366-2374].codfw.wmnet * 17:10 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply * 17:10 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply * 17:08 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2366-2374].codfw.wmnet * 17:06 sukhe: sre.dns.roll-reboot to resume later * 17:06 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-reboot (exit_code=97) rolling reboot on A:dnsbox and not (A:ulsfo or A:magru) and (A:dnsbox) * 17:06 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns5003.wikimedia.org * 17:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2241: Security update * 17:03 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2241: Security update * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply * 17:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2366-2374].codfw.wmnet * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply * 17:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2357-2365].codfw.wmnet * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 17:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2357-2365].codfw.wmnet * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply * 16:55 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2357-2365].codfw.wmnet * 16:55 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 16:53 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6006.drmrs.wmnet * 16:52 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6014.drmrs.wmnet * 16:52 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 16:51 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 16:50 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2357-2365].codfw.wmnet * 16:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4042.ulsfo.wmnet * 16:50 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2347-2356].codfw.wmnet * 16:50 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2347-2356].codfw.wmnet * 16:49 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns5003.wikimedia.org * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply * 16:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4050.ulsfo.wmnet * 16:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2347-2356].codfw.wmnet * 16:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2347-2356].codfw.wmnet * 16:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2337-2346].codfw.wmnet * 16:36 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2337-2346].codfw.wmnet * 16:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:dse-k8s-worker-codfw * 16:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2003.codfw.wmnet * 16:35 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2003.codfw.wmnet * 16:34 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns3004.wikimedia.org * 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply * 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply * 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply * 16:30 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply * 16:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2003.codfw.wmnet * 16:29 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2337-2346].codfw.wmnet * 16:24 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2003.codfw.wmnet * 16:24 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2002.codfw.wmnet * 16:24 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2002.codfw.wmnet * 16:23 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns3004.wikimedia.org * 16:23 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2337-2346].codfw.wmnet * 16:23 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2327-2336].codfw.wmnet * 16:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2327-2336].codfw.wmnet * 16:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2002.codfw.wmnet * 16:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2327-2336].codfw.wmnet * 16:12 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2002.codfw.wmnet * 16:12 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2001.codfw.wmnet * 16:12 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2001.codfw.wmnet * 16:12 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1065.eqiad.wmnet with OS trixie * 16:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6005.drmrs.wmnet * 16:11 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6013.drmrs.wmnet * 16:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4041.ulsfo.wmnet * 16:08 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns3003.wikimedia.org * 16:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2327-2336].codfw.wmnet * 16:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2317-2326].codfw.wmnet * 16:06 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2317-2326].codfw.wmnet * 16:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2001.codfw.wmnet * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply * 16:03 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4049.ulsfo.wmnet * 16:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2001.codfw.wmnet * 16:00 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test2001.codfw.wmnet * 16:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test2001.codfw.wmnet * 16:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2063.codfw.wmnet with OS trixie * 15:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2317-2326].codfw.wmnet * 15:57 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns3003.wikimedia.org * 15:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test2001.codfw.wmnet * 15:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test2001.codfw.wmnet * 15:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2004.codfw.wmnet * 15:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2004.codfw.wmnet * 15:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2317-2326].codfw.wmnet * 15:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2307-2316].codfw.wmnet * 15:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2307-2316].codfw.wmnet * 15:49 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2004.codfw.wmnet * 15:48 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2004.codfw.wmnet * 15:48 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2003.codfw.wmnet * 15:48 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2003.codfw.wmnet * 15:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 15:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2307-2316].codfw.wmnet * 15:42 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2003.codfw.wmnet * 15:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 15:42 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2003.codfw.wmnet * 15:42 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2002.codfw.wmnet * 15:42 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2002.codfw.wmnet * 15:42 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2006.wikimedia.org * 15:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2063.codfw.wmnet with reason: host reimage * 15:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2307-2316].codfw.wmnet * 15:37 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2297-2306].codfw.wmnet * 15:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2297-2306].codfw.wmnet * 15:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2002.codfw.wmnet * 15:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2002.codfw.wmnet * 15:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2001.codfw.wmnet * 15:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2001.codfw.wmnet * 15:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2063.codfw.wmnet with reason: host reimage * 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6004.drmrs.wmnet * 15:31 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2001.codfw.wmnet * 15:31 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2001.codfw.wmnet * 15:31 btullis@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:dse-k8s-worker-codfw * 15:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6012.drmrs.wmnet * 15:28 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2006.wikimedia.org * 15:27 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-analytics cluster: Roll restart of jvm daemons. * 15:27 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2297-2306].codfw.wmnet * 15:27 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4040.ulsfo.wmnet * 15:24 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1065.eqiad.wmnet with OS trixie * 15:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4048.ulsfo.wmnet * 15:21 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-analytics cluster: Roll restart of jvm daemons. * 15:21 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2297-2306].codfw.wmnet * 15:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2287-2296].codfw.wmnet * 15:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2287-2296].codfw.wmnet * 15:20 btullis@cumin1003: END (PASS) - Cookbook sre.druid.reboot-workers (exit_code=0) for Druid public cluster: Reboot Druid nodes * 15:18 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 15:17 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm * 15:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2063.codfw.wmnet with OS trixie * 15:13 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2005.wikimedia.org * 15:11 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2287-2296].codfw.wmnet * 15:11 btullis@cumin1003: END (PASS) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=0) rolling reboot on A:cephosd-eqiad * 15:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1064.eqiad.wmnet with OS trixie * 15:05 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2062.codfw.wmnet with OS trixie * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2287-2296].codfw.wmnet * 15:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2277-2286].codfw.wmnet * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2277-2286].codfw.wmnet * 14:59 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2005.wikimedia.org * 14:57 brouberol@cumin1003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-jumbo-eqiad * 14:52 btullis@cumin1003: END (PASS) - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas (exit_code=0) rolling reboot on A:schema-codfw * 14:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6003.drmrs.wmnet * 14:50 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:50 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host relforge1009.eqiad.wmnet * 14:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2277-2286].codfw.wmnet * 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6011.drmrs.wmnet * 14:47 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 14:46 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>ml-serve1001.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 14:46 klausman@cumin1003: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) pool for host ml-serve1001.eqiad.wmnet * 14:46 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 14:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1001.eqiad.wmnet * 14:45 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4039.ulsfo.wmnet * 14:44 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1009.eqiad.wmnet * 14:44 btullis@cumin1003: START - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas rolling reboot on A:schema-codfw * 14:44 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2004.wikimedia.org * 14:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2277-2286].codfw.wmnet * 14:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2267-2276].codfw.wmnet * 14:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2267-2276].codfw.wmnet * 14:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 14:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4047.ulsfo.wmnet * 14:40 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1001.eqiad.wmnet * 14:38 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 14:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 14:36 topranks: disconnect power on cr2-eqiad to shut down device for switch fabric replacement [[phab:T426343|T426343]] * 14:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2267-2276].codfw.wmnet * 14:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 14:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1001.eqiad.wmnet * 14:35 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>ml-serve1001.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 14:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 14:34 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 14:33 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:33 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:30 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2004.wikimedia.org * 14:29 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2267-2276].codfw.wmnet * 14:29 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2257-2266].codfw.wmnet * 14:29 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2257-2266].codfw.wmnet * 14:24 btullis@cumin1003: END (PASS) - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas (exit_code=0) rolling reboot on A:schema-eqiad * 14:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2257-2266].codfw.wmnet * 14:20 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:20 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:19 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:17 jforrester@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2257-2266].codfw.wmnet * 14:16 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:16 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2062.codfw.wmnet with OS trixie * 14:15 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1064.eqiad.wmnet with OS trixie * 14:15 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1006.wikimedia.org * 14:15 btullis@cumin1003: START - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas rolling reboot on A:schema-eqiad * 14:14 topranks: switch routing-engine on cr2-eqiad resetting all interfaces [[phab:T417873|T417873]] * 14:11 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:11 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:10 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm * 14:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6002.drmrs.wmnet * 14:09 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6010.drmrs.wmnet * 14:06 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1006.wikimedia.org * 14:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:05 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 14:05 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4038.ulsfo.wmnet * 14:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4046.ulsfo.wmnet * 14:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:00 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on cr1-eqiad with reason: switch upgrade and line card install * 14:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:59 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:57 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:57 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:55 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-eqiad * 13:55 btullis@cumin1003: START - Cookbook sre.druid.reboot-workers for Druid public cluster: Reboot Druid nodes * 13:53 topranks: switch routing-engine on cr2-eqiad resetting all interfaces [[phab:T417873|T417873]] * 13:51 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1005.wikimedia.org * 13:50 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:49 brouberol@cumin1003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-test-eqiad * 13:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:44 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2197-2206].codfw.wmnet * 13:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2197-2206].codfw.wmnet * 13:36 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1005.wikimedia.org * 13:35 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2197-2206].codfw.wmnet * 13:30 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2197-2206].codfw.wmnet * 13:28 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6001.drmrs.wmnet * 13:28 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6009.drmrs.wmnet * 13:28 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2187-2196].codfw.wmnet * 13:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2187-2196].codfw.wmnet * 13:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 13:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4037.ulsfo.wmnet * 13:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2001 * 13:22 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2001 * 13:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4045.ulsfo.wmnet * 13:21 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1004.wikimedia.org * 13:19 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on lvs[1018-1020].eqiad.wmnet with reason: switch upgrade and line card install * 13:18 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2009.codfw.wmnet * 13:18 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2001 * 13:18 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2001.codfw.wmnet 26.16.192.10.in-addr.arpa 6.2.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:17 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2001.codfw.wmnet 26.16.192.10.in-addr.arpa 6.2.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:17 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:17 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2001 - bking@cumin2003" * 13:17 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2001 - bking@cumin2003" * 13:17 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2009.codfw.wmnet * 13:17 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_drmrs * 13:17 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2187-2196].codfw.wmnet * 13:17 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_drmrs * 13:17 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on 15 hosts with reason: switch upgrade and line card install * 13:17 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:15 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 13:13 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:13 brouberol@cumin1003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-jumbo-eqiad * 13:13 brouberol@cumin1003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-test-eqiad * 13:13 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1004.wikimedia.org * 13:13 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and not (A:ulsfo or A:magru) and (A:dnsbox) * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:12 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_ulsfo * 13:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:12 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_ulsfo * 13:11 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2187-2196].codfw.wmnet * 13:11 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 13:11 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 13:06 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 13:05 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling source-only afterwards * 13:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2001 * 13:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2009.codfw.wmnet with OS trixie * 13:03 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:03 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling source-only afterwards * 13:01 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 15s) * 13:01 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 13:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 12:57 btullis@cumin1003: END (PASS) - Cookbook sre.druid.reboot-workers (exit_code=0) for Druid analytics cluster: Reboot Druid nodes * 12:54 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 12:54 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2163-2172].codfw.wmnet * 12:54 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2163-2172].codfw.wmnet * 12:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2163-2172].codfw.wmnet * 12:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2009.codfw.wmnet with reason: host reimage * 12:41 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2163-2172].codfw.wmnet * 12:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2153-2162].codfw.wmnet * 12:40 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2153-2162].codfw.wmnet * 12:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2009.codfw.wmnet with reason: host reimage * 12:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2153-2162].codfw.wmnet * 12:29 btullis@cumin1003: END (PASS) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=0) rolling reboot on A:cephosd-codfw * 12:25 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2153-2162].codfw.wmnet * 12:25 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2143-2152].codfw.wmnet * 12:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2143-2152].codfw.wmnet * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2009 * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2009 * 12:22 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2009 * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2009.codfw.wmnet 139.0.192.10.in-addr.arpa 9.3.1.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:22 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2009.codfw.wmnet 139.0.192.10.in-addr.arpa 9.3.1.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2009 - mvernon@cumin2003" * 12:22 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2009 - mvernon@cumin2003" * 12:16 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 12:15 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 12:15 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 12:15 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2009 * 12:15 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 12:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2009.codfw.wmnet with OS trixie * 12:15 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 12:14 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 12:14 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2143-2152].codfw.wmnet * 12:13 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 12:12 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2010.codfw.wmnet * 12:11 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2010.codfw.wmnet * 12:10 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 12:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2143-2152].codfw.wmnet * 12:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2133-2142].codfw.wmnet * 12:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2133-2142].codfw.wmnet * 12:02 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 11:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2133-2142].codfw.wmnet * 11:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2133-2142].codfw.wmnet * 11:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:49 mvolz@deploy2003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:49 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-codfw * 11:48 mvolz@deploy2003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:47 btullis@cumin1003: START - Cookbook sre.druid.reboot-workers for Druid analytics cluster: Reboot Druid nodes * 11:46 mvolz@deploy2003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:46 mvolz@deploy2003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:45 mvolz@deploy2003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:44 mvolz@deploy2003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:40 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] (duration: 11m 38s) * 11:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2010.codfw.wmnet with OS trixie * 11:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2105-2114].codfw.wmnet * 11:36 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2105-2114].codfw.wmnet * 11:36 krinkle@deploy2003: physikerwelt, krinkle: Continuing with deployment * 11:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1018: Security updates * 11:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:36 root@cumin1003: START - Cookbook sre.mysql.parsercache * 11:36 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1018: Security updates * 11:31 krinkle@deploy2003: physikerwelt, krinkle: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:29 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] * 11:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2105-2114].codfw.wmnet * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2105-2114].codfw.wmnet * 11:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2010.codfw.wmnet with reason: host reimage * 11:12 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2010.codfw.wmnet with reason: host reimage * 11:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1018: Security updates * 11:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:10 root@cumin1003: START - Cookbook sre.mysql.parsercache * 11:10 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1018: Security updates * 11:09 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1009.eqiad.wmnet with OS trixie * 11:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow7002.magru.wmnet * 11:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 11:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 11:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=tegola-vector-tiles,name=eqiad * 11:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=kartotherian,name=eqiad * 11:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow7002.magru.wmnet * 10:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2010 * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2010 * 10:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 10:54 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 10:54 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2010 * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2010.codfw.wmnet 76.16.192.10.in-addr.arpa 6.7.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:54 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2010.codfw.wmnet 76.16.192.10.in-addr.arpa 6.7.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2010 - mvernon@cumin2003" * 10:54 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2010 - mvernon@cumin2003" * 10:49 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 10:49 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2010 * 10:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1009.eqiad.wmnet with reason: host reimage * 10:49 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2010.codfw.wmnet with OS trixie * 10:46 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2011.codfw.wmnet * 10:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1011.eqiad.wmnet * 10:44 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2011.codfw.wmnet * 10:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 10:44 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 10:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1009.eqiad.wmnet with reason: host reimage * 10:44 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow6001.drmrs.wmnet * 10:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1017: Security updates * 10:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:39 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1017: Security updates * 10:39 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow6001.drmrs.wmnet * 10:38 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1011.eqiad.wmnet * 10:35 cgoubert@deploy2003: Finished deploy [restbase/deploy@06301bd]: Deploying {{Gerrit|1306088}} {{Gerrit|1308347}} - [[phab:T429944|T429944]] [[phab:T428279|T428279]] (duration: 28m 34s) * 10:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1012.eqiad.wmnet * 10:35 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 10:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow5003.eqsin.wmnet * 10:34 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2011.codfw.wmnet with OS trixie * 10:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1009.eqiad.wmnet with OS trixie * 10:28 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1012.eqiad.wmnet * 10:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1013.eqiad.wmnet * 10:27 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:27 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow5003.eqsin.wmnet * 10:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:26 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow4003.ulsfo.wmnet * 10:25 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 10:25 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 10:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow4003.ulsfo.wmnet * 10:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1013.eqiad.wmnet * 10:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1014.eqiad.wmnet * 10:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:15 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2011.codfw.wmnet with reason: host reimage * 10:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1017: Security updates * 10:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:14 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:14 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1017: Security updates * 10:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow3004.esams.wmnet * 10:11 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2011.codfw.wmnet with reason: host reimage * 10:10 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:10 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1014.eqiad.wmnet * 10:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki2003.codfw.wmnet * 10:09 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 10:09 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 10:09 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow3004.esams.wmnet * 10:08 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2004.codfw.wmnet * 10:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1010.eqiad.wmnet with OS trixie * 10:07 cgoubert@deploy2003: Started deploy [restbase/deploy@06301bd]: Deploying {{Gerrit|1306088}} {{Gerrit|1308347}} - [[phab:T429944|T429944]] [[phab:T428279|T428279]] * 10:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host rpki2003.codfw.wmnet * 10:04 topranks: push out config change to BGP_outfilter on core routers [[phab:T431849|T431849]] * 10:02 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow2004.codfw.wmnet * 09:59 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 09:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2003.codfw.wmnet * 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2011 * 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2011 * 09:53 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 09:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:52 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2011 * 09:52 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2011.codfw.wmnet 36.32.192.10.in-addr.arpa 6.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:52 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2011.codfw.wmnet 36.32.192.10.in-addr.arpa 6.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:51 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:51 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2011 - mvernon@cumin2003" * 09:51 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2011 - mvernon@cumin2003" * 09:51 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow2003.codfw.wmnet * 09:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1003.eqiad.wmnet * 09:49 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 09:49 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 09:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1010.eqiad.wmnet with reason: host reimage * 09:47 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 09:47 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 09:47 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 09:47 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2011 * 09:46 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2011.codfw.wmnet with OS trixie * 09:44 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow1003.eqiad.wmnet * 09:44 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2012.codfw.wmnet * 09:44 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1002.eqiad.wmnet * 09:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1010.eqiad.wmnet with reason: host reimage * 09:43 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2012.codfw.wmnet * 09:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Security updates * 09:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:43 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:43 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Security updates * 09:42 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:40 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow1002.eqiad.wmnet * 09:40 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 09:37 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki1001.eqiad.wmnet * 09:36 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 09:36 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 09:33 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host rpki1001.eqiad.wmnet * 09:32 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:32 cgoubert@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-codfw * 09:31 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=kartotherian,name=eqiad * 09:31 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola-vector-tiles,name=eqiad * 09:31 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 09:31 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2012.codfw.wmnet with OS trixie * 09:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1010.eqiad.wmnet with OS trixie * 09:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1011.eqiad.wmnet with OS trixie * 09:21 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Security updates * 09:21 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:21 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:21 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Security updates * 09:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2012.codfw.wmnet with reason: host reimage * 09:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1011.eqiad.wmnet with reason: host reimage * 09:08 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2012.codfw.wmnet with reason: host reimage * 09:05 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1011.eqiad.wmnet with reason: host reimage * 08:55 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:52 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1011.eqiad.wmnet with OS trixie * 08:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1022: Security updates * 08:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2012 * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2012 * 08:50 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1022: Security updates * 08:50 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2012 * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2012.codfw.wmnet 44.48.192.10.in-addr.arpa 4.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:50 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2012.codfw.wmnet 44.48.192.10.in-addr.arpa 4.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2012 - mvernon@cumin2003" * 08:50 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2012 - mvernon@cumin2003" * 08:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1012.eqiad.wmnet with OS trixie * 08:44 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 08:44 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2012 * 08:43 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2012.codfw.wmnet with OS trixie * 08:42 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2013.codfw.wmnet * 08:41 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2013.codfw.wmnet * 08:35 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 08:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host krb1002.eqiad.wmnet * 08:30 elukey@dns1004: END - running authdns-update * 08:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1012.eqiad.wmnet with reason: host reimage * 08:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Security updates * 08:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:28 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:28 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Security updates * 08:27 elukey@dns1004: START - running authdns-update * 08:26 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 08:26 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host krb1002.eqiad.wmnet * 08:22 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1012.eqiad.wmnet with reason: host reimage * 08:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host krb2002.codfw.wmnet * 08:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast6003.wikimedia.org * 08:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2013.codfw.wmnet with OS trixie * 08:13 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast6003.wikimedia.org * 08:12 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast3007.wikimedia.org * 08:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host krb2002.codfw.wmnet * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Security updates * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:09 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:09 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Security updates * 08:07 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1012.eqiad.wmnet with OS trixie * 08:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast3007.wikimedia.org * 08:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast5005.wikimedia.org * 07:58 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast5005.wikimedia.org * 07:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1013.eqiad.wmnet with OS trixie * 07:53 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2013.codfw.wmnet with reason: host reimage * 07:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1021: Security updates * 07:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:53 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:53 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1021: Security updates * 07:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast1004.wikimedia.org * 07:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2013.codfw.wmnet with reason: host reimage * 07:46 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast1004.wikimedia.org * 07:40 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1013.eqiad.wmnet with reason: host reimage * 07:36 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1013.eqiad.wmnet with reason: host reimage * 07:31 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2013 * 07:31 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2013 * 07:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1021: Security updates * 07:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:30 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:30 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1021: Security updates * 07:24 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2013 * 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2013.codfw.wmnet 87.0.192.10.in-addr.arpa 7.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:24 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2013.codfw.wmnet 87.0.192.10.in-addr.arpa 7.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2013 - mvernon@cumin2003" * 07:24 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2013 - mvernon@cumin2003" * 07:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1013.eqiad.wmnet with OS trixie * 07:19 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 07:19 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2013 * 07:19 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2013.codfw.wmnet with OS trixie * 07:13 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] (duration: 07m 48s) * 07:09 kharlan@deploy2003: kharlan: Continuing with deployment * 07:08 kharlan@deploy2003: kharlan: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:06 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 01:15 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 01:14 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply == 2026-07-14 == * 22:51 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_magru * 22:51 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7016.magru.wmnet * 22:46 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_magru * 22:46 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7008.magru.wmnet * 22:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7015.magru.wmnet * 22:04 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7007.magru.wmnet * 21:29 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7014.magru.wmnet * 21:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7006.magru.wmnet * 21:13 dzahn@cumin2002: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 0:15:00 on gerrit.wikimedia.org with reason: reboot * 21:11 mutante: gerrit2003 (gerrit.wikimedia.org) - reboot for maintenance * 21:11 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on gerrit2003.wikimedia.org with reason: reboot * 20:56 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:56 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:56 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:55 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 20:48 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7013.magru.wmnet * 20:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7005.magru.wmnet * 20:41 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host phab1005.eqiad.wmnet with OS trixie * 20:28 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] (duration: 06m 47s) * 20:24 sbassett@deploy2003: sbassett: Continuing with deployment * 20:23 sbassett@deploy2003: sbassett: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:23 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on phab1005.eqiad.wmnet with reason: host reimage * 20:21 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] * 20:20 aokoth@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on phab1005.eqiad.wmnet with reason: host reimage * 20:12 jhuneidi@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] (duration: 07m 42s) * 20:07 jhuneidi@deploy2003: jhuneidi, priyankar22: Continuing with deployment * 20:06 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7012.magru.wmnet * 20:06 jhuneidi@deploy2003: jhuneidi, priyankar22: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:04 jhuneidi@deploy2003: Started scap sync-world: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] * 20:02 aokoth@cumin1003: START - Cookbook sre.hosts.reimage for host phab1005.eqiad.wmnet with OS trixie * 20:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7004.magru.wmnet * 20:00 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet * 19:57 aokoth@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet * 19:24 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7011.magru.wmnet * 19:19 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7003.magru.wmnet * 19:11 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] (duration: 08m 33s) * 19:07 jforrester@deploy2003: jforrester: Continuing with deployment * 19:04 jforrester@deploy2003: jforrester: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:02 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] * 18:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7010.magru.wmnet * 18:38 mutante: rotating phabricator-gerrit bot token (its-phabricator) * 18:18 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 17:44 swfrench@deploy2003: Finished scap sync-world: Deployment to pick up new production image (duration: 31m 44s) * 17:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7002.magru.wmnet * 17:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7009.magru.wmnet * 17:32 swfrench@deploy2003: swfrench: Continuing with deployment * 17:29 swfrench@deploy2003: swfrench: Deployment to pick up new production image synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:17 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2035: repooling after rack b5 maintenance * 17:16 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool es2035: repooling after rack b5 maintenance * 17:16 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2188: repooling after rack b5 maintenance * 17:12 swfrench@deploy2003: Started scap sync-world: Deployment to pick up new production image * 17:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7001.magru.wmnet * 16:57 swfrench-wmf: reprepro include php8.3_8.3.32-1+wmf12u2 into component/php83 for bookworm-wikimedia * 16:50 sukhe: pool cp2046 * 16:47 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4039.ulsfo.wmnet * 16:44 sukhe: sudo cumin -b31 "A:cp" "run-puppet-agent" * 16:33 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on contint1003.wikimedia.org with reason: reboot * 16:32 mutante: contint1003 - main CI server - rebooting * 16:31 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2188: repooling after rack b5 maintenance * 16:31 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2178: repooling after rack b5 maintenance * 16:29 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 16:28 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 16:28 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 16:28 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 16:18 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2014.codfw.wmnet * 16:18 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 16:17 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2014.codfw.wmnet * 16:10 mvernon@cumin1003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-thanos-proxies (exit_code=0) rolling restart_daemons on A:thanos-fe * 16:09 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 16:07 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp4039.ulsfo.wmnet * 16:07 mvernon@cumin1003: START - Cookbook sre.swift.roll-restart-reboot-swift-thanos-proxies rolling restart_daemons on A:thanos-fe * 16:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2014.codfw.wmnet with OS trixie * 15:56 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1014.eqiad.wmnet with OS trixie * 15:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2014.codfw.wmnet with reason: host reimage * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2014 * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2014 * 15:28 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2014 * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2014.codfw.wmnet 194.16.192.10.in-addr.arpa 4.9.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:28 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2014.codfw.wmnet 194.16.192.10.in-addr.arpa 4.9.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2014 - mvernon@cumin2003" * 15:28 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2014 - mvernon@cumin2003" * 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Apply title-related policies when selecting the name of the entity - kamila@cumin1003" * 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Apply title-related policies when selecting the name of the entity - kamila@cumin1003 * 15:22 kamila@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Apply title-related policies when selecting the name of the entity - kamila@cumin1003 * 15:22 kamila@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Apply title-related policies when selecting the name of the entity - kamila@cumin1003" * 15:20 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 15:20 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2014 * 15:20 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2014.codfw.wmnet with OS trixie * 15:19 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1014.eqiad.wmnet with OS trixie * 15:01 dancy@deploy2003: Installation of scap version "4.274.1" completed for 3 hosts * 15:00 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2177: repooling after rack b5 maintenance * 15:00 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2159: repooling after rack b5 maintenance * 14:59 dancy@deploy2003: Installing scap version "4.274.1" for 3 host(s) * 14:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2015.codfw.wmnet with OS trixie * 14:54 seanleong-wmde: Finished populateSitesTable for isvwiki ([[phab:T429939|T429939]]) * 14:53 javiermonton@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] (duration: 07m 35s) * 14:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1015.eqiad.wmnet with OS trixie * 14:49 javiermonton@deploy2003: javiermonton: Continuing with deployment * 14:48 javiermonton@deploy2003: javiermonton: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:46 javiermonton@deploy2003: Started scap sync-world: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] * 14:42 otto@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 14:41 otto@deploy2003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 14:41 otto@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 14:40 otto@deploy2003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 14:40 otto@deploy2003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 14:39 otto@deploy2003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 14:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2015.codfw.wmnet with reason: host reimage * 14:34 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1015.eqiad.wmnet with reason: host reimage * 14:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2015.codfw.wmnet with reason: host reimage * 14:30 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1015.eqiad.wmnet with reason: host reimage * 14:30 seanleong-wmde@deploy2003: mwscript-k8s job started: foreachwikiindblist wikidataclient extensions/Wikibase/lib/maintenance/populateSitesTable.php --force-protocol https # [[phab:T429939|T429939]] * 14:24 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling reboot on A:durum and not (A:durum-eqiad or A:durum-codfw or A:durum-esams) and A:durum * 14:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2015.codfw.wmnet with OS trixie * 14:15 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2016.codfw.wmnet with OS trixie * 14:14 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2159: repooling after rack b5 maintenance * 14:14 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1015.eqiad.wmnet with OS trixie * 14:12 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1016.eqiad.wmnet with OS trixie * 14:12 cmooney@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=pki,name=codfw * 14:12 sbisson@deploy2003: helmfile [codfw] DONE helmfile.d/services/cxserver: sync * 14:11 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2002.codfw.wmnet * 14:11 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2002.codfw.wmnet * 14:11 sbisson@deploy2003: helmfile [codfw] START helmfile.d/services/cxserver: sync * 14:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1005.wikimedia.org * 14:07 sbisson@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cxserver: sync * 14:07 sbisson@deploy2003: helmfile [eqiad] START helmfile.d/services/cxserver: sync * 14:05 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1005.wikimedia.org * 14:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader2005.wikimedia.org * 14:02 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2003.codfw.wmnet * 14:02 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2003.codfw.wmnet * 14:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=tegola-vector-tiles,name=codfw * 14:00 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=kartotherian,name=codfw * 14:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader2005.wikimedia.org * 13:58 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2016.codfw.wmnet with reason: host reimage * 13:57 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-ncredir (exit_code=0) rolling reboot on A:ncredir and A:ncredir * 13:57 sbisson@deploy2003: helmfile [staging] DONE helmfile.d/services/cxserver: sync * 13:56 sbisson@deploy2003: helmfile [staging] START helmfile.d/services/cxserver: sync * 13:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1016.eqiad.wmnet with reason: host reimage * 13:52 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:52 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:51 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2016.codfw.wmnet with reason: host reimage * 13:50 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1016.eqiad.wmnet with reason: host reimage * 13:49 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy (exit_code=0) rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 13:49 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling reboot on A:wikidough * 13:46 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-tcp-proxy (exit_code=0) rolling reboot on A:tcpproxy and A:tcpproxy * 13:43 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and not (A:durum-eqiad or A:durum-codfw or A:durum-esams) and A:durum * 13:42 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=97) rolling reboot on A:durum and A:durum * 13:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2011.codfw.wmnet * 13:36 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-reboot (exit_code=0) rolling reboot on A:dnsbox and A:ulsfo and (A:dnsbox) * 13:36 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns4004.wikimedia.org * 13:34 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1016.eqiad.wmnet with OS trixie * 13:34 topranks: reboot lsw1-b5-codfw to upgrade JunOS [[phab:T430918|T430918]] * 13:34 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2016.codfw.wmnet with OS trixie * 13:32 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2002.codfw.wmnet * 13:31 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2011.codfw.wmnet * 13:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2012.codfw.wmnet * 13:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2017.codfw.wmnet with OS trixie * 13:25 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1017.eqiad.wmnet with OS trixie * 13:24 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2012.codfw.wmnet * 13:22 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2002.codfw.wmnet * 13:22 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:22 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:22 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns4004.wikimedia.org * 13:21 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2005.codfw.wmnet * 13:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2013.codfw.wmnet * 13:18 elukey@dns1004: END - running authdns-update * 13:17 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2005.codfw.wmnet * 13:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2004.codfw.wmnet * 13:16 elukey@dns1004: START - running authdns-update * 13:16 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1046: es1046 after reimage * 13:14 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1029.eqiad.wmnet,service=s8 * 13:14 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1029.eqiad.wmnet,service=s5 * 13:13 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1029.eqiad.wmnet,service=s5 * 13:13 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1029.eqiad.wmnet,service=s8 * 13:13 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2004.codfw.wmnet * 13:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2013.codfw.wmnet * 13:11 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2014.codfw.wmnet * 13:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2003.codfw.wmnet * 13:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2017.codfw.wmnet with reason: host reimage * 13:07 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:07 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns4003.wikimedia.org * 13:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2003.codfw.wmnet * 13:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1067.eqiad.wmnet * 13:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1067.eqiad.wmnet * 13:06 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1067.eqiad.wmnet * 13:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm-test1001.wikimedia.org * 13:05 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2017.codfw.wmnet with reason: host reimage * 13:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1017.eqiad.wmnet with reason: host reimage * 13:04 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2014.codfw.wmnet * 13:03 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:02 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2188: codfw rack B5 depool for maintenance * 13:02 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_magru * 13:01 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2188: codfw rack B5 depool for maintenance * 13:01 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_magru * 13:01 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2178: codfw rack B5 depool for maintenance * 13:01 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm-test1001.wikimedia.org * 13:01 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2178: codfw rack B5 depool for maintenance * 13:01 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2177: codfw rack B5 depool for maintenance * 13:00 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2177: codfw rack B5 depool for maintenance * 12:59 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1068.eqiad.wmnet * 12:59 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1068.eqiad.wmnet * 12:58 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola-vector-tiles,name=codfw * 12:58 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2159: codfw rack B5 depool for maintenance * 12:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1017.eqiad.wmnet with reason: host reimage * 12:58 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola,name=codfw * 12:57 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=kartotherian,name=codfw * 12:57 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2159: codfw rack B5 depool for maintenance * 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 30 hosts with reason: lsw1-b5-codfw JunOS upgrade * 12:55 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lsw1-b5-codfw,lsw1-b5-codfw IPv6,lsw1-b5-codfw.mgmt,ssw1-a[1,8]-codfw.mgmt with reason: switch upgade lsw1-b5-codfw * 12:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet * 12:49 cmooney@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=pki,name=codfw * 12:49 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1067.eqiad.wmnet with OS trixie * 12:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet * 12:48 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2017.codfw.wmnet with OS trixie * 12:47 topranks: depool codfw pki in dns discovery ahead of lsw1-b5-codfw maintenance [[phab:T430918|T430918]] * 12:47 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns4003.wikimedia.org * 12:47 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and A:ulsfo and (A:dnsbox) * 12:47 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2018.codfw.wmnet * 12:45 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2018.codfw.wmnet * 12:45 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and A:durum * 12:45 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-tcp-proxy rolling reboot on A:tcpproxy and A:tcpproxy * 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host pki-root1002.eqiad.wmnet * 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1009.eqiad.wmnet * 12:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1009.eqiad.wmnet * 12:44 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 12:43 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-ncredir rolling reboot on A:ncredir and A:ncredir * 12:43 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling reboot on A:wikidough * 12:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2018.codfw.wmnet with OS trixie * 12:42 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1017.eqiad.wmnet with OS trixie * 12:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader2006.wikimedia.org * 12:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1018.eqiad.wmnet with OS trixie * 12:39 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1009.eqiad.wmnet * 12:38 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host pki-root1002.eqiad.wmnet * 12:38 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1009.eqiad.wmnet * 12:38 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1008.eqiad.wmnet * 12:38 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1008.eqiad.wmnet * 12:35 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader2006.wikimedia.org * 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1006.wikimedia.org * 12:33 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1008.eqiad.wmnet * 12:30 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1046: es1046 after reimage * 12:29 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host es1046.eqiad.wmnet with OS trixie * 12:29 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1006.wikimedia.org * 12:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test2005.wikimedia.org * 12:28 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1008.eqiad.wmnet * 12:27 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1007.eqiad.wmnet * 12:27 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1007.eqiad.wmnet * 12:27 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1067.eqiad.wmnet with reason: host reimage * 12:25 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1068.eqiad.wmnet with reason: vacuum overlarge container dbs * 12:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2018.codfw.wmnet with reason: host reimage * 12:24 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test2005.wikimedia.org * 12:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test1005.wikimedia.org * 12:23 Amir1: mwscript-k8s --follow --dblist=ores -- extensions/ORES/maintenance/PurgeScoreCache.php --model goodfaith --old ([[phab:T431159|T431159]]) * 12:22 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1007.eqiad.wmnet * 12:22 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test1005.wikimedia.org * 12:22 atsukoito: restarting pybal on lvs2013 `low-traffic` for https://gerrit.wikimedia.org/r/1310535 * 12:22 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1007.eqiad.wmnet * 12:21 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1006.eqiad.wmnet * 12:21 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1006.eqiad.wmnet * 12:20 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1018.eqiad.wmnet with reason: host reimage * 12:19 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2018.codfw.wmnet with reason: host reimage * 12:18 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1067.eqiad.wmnet with reason: host reimage * 12:16 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1006.eqiad.wmnet * 12:15 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1006.eqiad.wmnet * 12:15 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1005.eqiad.wmnet * 12:15 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1005.eqiad.wmnet * 12:15 atsukoito: restarting pybal on lvs2014 for https://gerrit.wikimedia.org/r/1310535 * 12:12 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1018.eqiad.wmnet with reason: host reimage * 12:11 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1005.eqiad.wmnet * 12:11 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1005.eqiad.wmnet * 12:10 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1004.eqiad.wmnet * 12:10 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1004.eqiad.wmnet * 12:09 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on es1046.eqiad.wmnet with reason: host reimage * 12:08 atsukoito: restarting pybal on lvs1019 `low-traffic` for https://gerrit.wikimedia.org/r/1310535 * 12:06 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1004.eqiad.wmnet * 12:06 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1004.eqiad.wmnet * 12:06 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1003.eqiad.wmnet * 12:06 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1003.eqiad.wmnet * 12:05 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on es1046.eqiad.wmnet with reason: host reimage * 12:05 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "set ml-serve1001 back to active state - cmooney@cumin1003" * 12:04 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "set ml-serve1001 back to active state - cmooney@cumin1003" * 12:04 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:02 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1003.eqiad.wmnet * 12:01 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1003.eqiad.wmnet * 12:01 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1002.eqiad.wmnet * 12:01 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1002.eqiad.wmnet * 12:01 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:59 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2018.codfw.wmnet with OS trixie * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1067 * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1067 * 11:59 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1067 * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1067.eqiad.wmnet 17.48.64.10.in-addr.arpa 7.1.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:59 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1067.eqiad.wmnet 17.48.64.10.in-addr.arpa 7.1.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1067 - blake@cumin1003" * 11:58 atsukoito: restarting pybal on lvs1018 `high-traffic2` for https://gerrit.wikimedia.org/r/1310535 * 11:57 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1002.eqiad.wmnet * 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2019.codfw.wmnet with OS trixie * 11:56 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1002.eqiad.wmnet * 11:56 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 11:56 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1018.eqiad.wmnet with OS trixie * 11:54 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 11:54 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:54 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1019.eqiad.wmnet with OS trixie * 11:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:49 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:49 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:49 aikochou@deploy2003: helmfile [codfw] DONE helmfile.d/services/changeprop: sync * 11:48 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host es1046.eqiad.wmnet with OS trixie * 11:48 aikochou@deploy2003: helmfile [codfw] START helmfile.d/services/changeprop: sync * 11:48 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310535 * 11:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1046: Reimage to Trixie * 11:44 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1046: Reimage to Trixie * 11:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5:00:00 on es1046.eqiad.wmnet with reason: Reimage to Trixie * 11:42 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:42 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:42 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 11:42 aikochou@deploy2003: helmfile [eqiad] DONE helmfile.d/services/changeprop: sync * 11:41 aikochou@deploy2003: helmfile [eqiad] START helmfile.d/services/changeprop: sync * 11:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2019.codfw.wmnet with reason: host reimage * 11:36 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] (duration: 09m 41s) * 11:36 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:36 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:35 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:35 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:32 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1019.eqiad.wmnet with reason: host reimage * 11:32 jforrester@deploy2003: jforrester, gengh: Continuing with deployment * 11:29 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2019.codfw.wmnet with reason: host reimage * 11:28 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1019.eqiad.wmnet with reason: host reimage * 11:28 jforrester@deploy2003: jforrester, gengh: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:26 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] * 11:20 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2003.codfw.wmnet * 11:20 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:19 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2003.codfw.wmnet * 11:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:12 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1019.eqiad.wmnet with OS trixie * 11:10 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2019.codfw.wmnet with OS trixie * 11:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1020.eqiad.wmnet with OS trixie * 11:10 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1067 - blake@cumin1003" * 11:09 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] (duration: 12m 12s) * 11:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2020.codfw.wmnet with OS trixie * 11:03 kharlan@deploy2003: kharlan: Continuing with deployment * 11:01 blake@cumin1003: START - Cookbook sre.dns.netbox * 11:01 kharlan@deploy2003: kharlan: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:57 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] * 10:55 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] (duration: 31m 40s) * 10:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1020.eqiad.wmnet with reason: host reimage * 10:52 marostegui@dns1004: START - running authdns-update * 10:49 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2020.codfw.wmnet with reason: host reimage * 10:49 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:48 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1020.eqiad.wmnet with reason: host reimage * 10:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2020.codfw.wmnet with reason: host reimage * 10:43 kharlan@deploy2003: kharlan: Continuing with deployment * 10:42 kharlan@deploy2003: kharlan: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:32 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1020.eqiad.wmnet with OS trixie * 10:29 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2159: Repooling after switchover * 10:29 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1067 * 10:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1021.eqiad.wmnet with OS trixie * 10:27 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1067.eqiad.wmnet with OS trixie * 10:27 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:27 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1067.eqiad.wmnet * 10:27 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:26 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1067.eqiad.wmnet * 10:26 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1067.eqiad.wmnet * 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2020.codfw.wmnet with OS trixie * 10:26 blake@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1055.eqiad.wmnet * 10:26 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1055.eqiad.wmnet * 10:26 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1055.eqiad.wmnet * 10:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2021.codfw.wmnet with OS trixie * 10:24 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] * 10:11 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1055.eqiad.wmnet with OS trixie * 10:09 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1021.eqiad.wmnet with reason: host reimage * 10:05 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2021.codfw.wmnet with reason: host reimage * 10:03 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310129 revert * 10:02 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1021.eqiad.wmnet with reason: host reimage * 10:01 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2021.codfw.wmnet with reason: host reimage * 09:58 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310129 * 09:50 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1055.eqiad.wmnet with reason: host reimage * 09:45 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1055.eqiad.wmnet with reason: host reimage * 09:45 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1021.eqiad.wmnet with OS trixie * 09:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1022.eqiad.wmnet with OS trixie * 09:44 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2159: Repooling after switchover * 09:44 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2021.codfw.wmnet with OS trixie * 09:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2022.codfw.wmnet with OS trixie * 09:31 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2159.codfw.wmnet * 09:28 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1055 * 09:28 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1055 * 09:27 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on ms-fe1022.eqiad.wmnet with reason: host reimage * 09:27 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1022.eqiad.wmnet with reason: host reimage * 09:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2022.codfw.wmnet with reason: host reimage * 09:21 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2022.codfw.wmnet with reason: host reimage * 09:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2159: Rebooting db2159.codfw.wmnet * 09:20 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2159: Rebooting db2159.codfw.wmnet * 09:18 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 09:18 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 09:18 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 09:17 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 09:16 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2159.codfw.wmnet * 09:13 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b] (thin): Regular analytics weekly train THIN [analytics/refinery@ad6e05b8] (duration: 02m 07s) * 09:11 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b] (thin): Regular analytics weekly train THIN [analytics/refinery@ad6e05b8] * 09:10 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1022.eqiad.wmnet with OS trixie * 09:07 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1023.eqiad.wmnet with OS trixie * 09:06 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b]: Regular analytics weekly train [analytics/refinery@ad6e05b8] (duration: 04m 49s) * 09:04 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2022.codfw.wmnet with OS trixie * 09:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2023.codfw.wmnet with OS trixie * 09:01 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b]: Regular analytics weekly train [analytics/refinery@ad6e05b8] * 09:01 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@ad6e05b8] (duration: 02m 01s) * 09:00 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1055 * 09:00 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1055.eqiad.wmnet 50.32.64.10.in-addr.arpa 0.5.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:00 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1055.eqiad.wmnet 50.32.64.10.in-addr.arpa 0.5.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:00 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:00 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1055 - blake@cumin1003" * 09:00 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1055 - blake@cumin1003" * 08:59 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@ad6e05b8] * 08:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2159 [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94811 and previous config saved to /var/cache/conftool/dbconfig/20260714-085624-cwilliams.json * 08:55 blake@cumin1003: START - Cookbook sre.dns.netbox * 08:55 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1055 * 08:54 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1055.eqiad.wmnet with OS trixie * 08:54 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1055.eqiad.wmnet * 08:53 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1055.eqiad.wmnet * 08:53 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1055.eqiad.wmnet * 08:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2220 to s7 primary [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94810 and previous config saved to /var/cache/conftool/dbconfig/20260714-085239-cwilliams.json * 08:51 cezmunsta: Starting s7 codfw failover from db2159 to db2220 - [[phab:T430920|T430920]] * 08:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1023.eqiad.wmnet with reason: host reimage * 08:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2220 with weight 0 [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94809 and previous config saved to /var/cache/conftool/dbconfig/20260714-084553-cwilliams.json * 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s7 [[phab:T430920|T430920]] * 08:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2023.codfw.wmnet with reason: host reimage * 08:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1023.eqiad.wmnet with reason: host reimage * 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2023.codfw.wmnet with reason: host reimage * 08:34 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:34 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:29 marostegui@dns1004: END - running authdns-update * 08:29 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox-dev2003.codfw.wmnet * 08:27 marostegui@dns1004: START - running authdns-update * 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker2*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2009.codfw.wmnet * 08:26 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2009.codfw.wmnet * 08:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1023.eqiad.wmnet with OS trixie * 08:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox-dev2003.codfw.wmnet * 08:24 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:24 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:24 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2023.codfw.wmnet with OS trixie * 08:24 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1029.eqiad.wmnet with reason: reboot * 08:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1027.eqiad.wmnet with reason: reboot * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:21 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2009.codfw.wmnet * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:20 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2009.codfw.wmnet * 08:20 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2008.codfw.wmnet * 08:20 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2008.codfw.wmnet * 08:15 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2008.codfw.wmnet * 08:14 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2008.codfw.wmnet * 08:14 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2007.codfw.wmnet * 08:14 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2007.codfw.wmnet * 08:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1024.eqiad.wmnet with OS trixie * 08:12 elukey@cumin1003: END (PASS) - Cookbook sre.pki.restart-reboot (exit_code=0) rolling reboot on P<nowiki>{</nowiki>pki*<nowiki>}</nowiki> and (A:pki) * 08:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 08:10 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 08:09 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2007.codfw.wmnet * 08:08 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2007.codfw.wmnet * 08:08 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2006.codfw.wmnet * 08:08 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2006.codfw.wmnet * 08:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2024.codfw.wmnet with OS trixie * 08:03 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2006.codfw.wmnet * 08:02 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2006.codfw.wmnet * 08:02 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2005.codfw.wmnet * 08:02 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2005.codfw.wmnet * 07:58 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2005.codfw.wmnet * 07:58 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2005.codfw.wmnet * 07:57 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2004.codfw.wmnet * 07:57 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2004.codfw.wmnet * 07:54 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki.discovery.wmnet. on all recursors * 07:54 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache pki.discovery.wmnet. on all recursors * 07:53 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2004.codfw.wmnet * 07:53 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2004.codfw.wmnet * 07:53 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2003.codfw.wmnet * 07:53 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2003.codfw.wmnet * 07:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1024.eqiad.wmnet with reason: host reimage * 07:49 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki.discovery.wmnet. on all recursors * 07:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2003.codfw.wmnet * 07:49 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache pki.discovery.wmnet. on all recursors * 07:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2024.codfw.wmnet with reason: host reimage * 07:48 elukey@cumin1003: START - Cookbook sre.pki.restart-reboot rolling reboot on P<nowiki>{</nowiki>pki*<nowiki>}</nowiki> and (A:pki) * 07:46 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1024.eqiad.wmnet with reason: host reimage * 07:45 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2024.codfw.wmnet with reason: host reimage * 07:45 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2003.codfw.wmnet * 07:44 elukey@cumin1003: END (PASS) - Cookbook sre.misc-clusters.restart-reboot-config-master (exit_code=0) rolling reboot on P<nowiki>{</nowiki>config-master*<nowiki>}</nowiki> and (A:config-master or A:config-master-eqiad or A:config-master-codfw) * 07:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2002.codfw.wmnet * 07:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2002.codfw.wmnet * 07:39 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2002.codfw.wmnet * 07:39 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) config-master.discovery.wmnet. on all recursors * 07:39 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache config-master.discovery.wmnet. on all recursors * 07:39 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2002.codfw.wmnet * 07:39 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker2*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl200*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl2003.codfw.wmnet * 07:36 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl2003.codfw.wmnet * 07:35 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) config-master.discovery.wmnet. on all recursors * 07:35 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache config-master.discovery.wmnet. on all recursors * 07:34 elukey@cumin1003: START - Cookbook sre.misc-clusters.restart-reboot-config-master rolling reboot on P<nowiki>{</nowiki>config-master*<nowiki>}</nowiki> and (A:config-master or A:config-master-eqiad or A:config-master-codfw) * 07:31 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl2003.codfw.wmnet * 07:31 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl2003.codfw.wmnet * 07:31 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl2002.codfw.wmnet * 07:31 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl2002.codfw.wmnet * 07:29 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1024.eqiad.wmnet with OS trixie * 07:28 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2024.codfw.wmnet with OS trixie * 07:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl2002.codfw.wmnet * 07:26 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl2002.codfw.wmnet * 07:26 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl200*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 07:26 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 07:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 06:50 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lists1004.wikimedia.org * 06:44 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host lists1004.wikimedia.org * 06:25 marostegui@dns1004: END - running authdns-update * 06:23 marostegui@dns1004: START - running authdns-update * 06:22 marostegui@dns1004: END - running authdns-update * 06:20 marostegui@dns1004: START - running authdns-update * 06:17 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1026.eqiad.wmnet with reason: reboot * 06:04 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: sync * 06:04 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: sync * 06:03 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync * 06:03 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync * 06:02 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync * 06:01 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync * 06:01 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync * 06:00 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync * 05:59 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:59 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:40 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:39 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:26 marostegui@dns1004: END - running authdns-update * 05:24 marostegui@dns1004: START - running authdns-update * 05:24 marostegui@dns1004: START - running authdns-update * 05:13 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1004.wikimedia.org * 05:07 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1004.wikimedia.org * 04:01 mwpresync@deploy2003: Pruned MediaWiki: 1.47.0-wmf.8 (duration: 01m 07s) * 03:39 mwpresync@deploy2003: Finished scap sync-world: testwikis to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] (duration: 36m 01s) * 03:03 mwpresync@deploy2003: Started scap sync-world: testwikis to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 29s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-13 == * 23:33 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1064.eqiad.wmnet * 23:33 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1064.eqiad.wmnet * 23:08 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1064.eqiad.wmnet with reason: vacuum overlarge container dbs * 23:06 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1069.eqiad.wmnet * 23:06 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1069.eqiad.wmnet * 22:34 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1069.eqiad.wmnet with reason: vacuum overlarge container dbs * 21:18 maryum: Deployed security fix for [[phab:T321092|T321092]] * 20:28 swfrench-wmf: reprepro include etcd-mirror_0.0.12-1+deb13u1 into main for trixie-wikimedia - [[phab:T424266|T424266]] * 20:26 swfrench-wmf: reprepro include etcd-mirror_0.0.12-1+deb12u1 into main for bookworm-wikimedia - [[phab:T428495|T428495]] * 20:23 dancy@deploy2003: Finished scap sync-world: Testing [[phab:T431635|T431635]] (duration: 03m 36s) * 20:19 dancy@deploy2003: Started scap sync-world: Testing [[phab:T431635|T431635]] * 20:18 dancy@deploy2003: Installation of scap version "4.274.0" completed for 3 hosts * 20:16 dancy@deploy2003: Installing scap version "4.274.0" for 3 host(s) * 20:12 kemayo@deploy2003: Finished scap sync-world: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] (duration: 08m 25s) * 20:07 kemayo@deploy2003: soda, esanders, kemayo: Continuing with deployment * 20:05 kemayo@deploy2003: soda, esanders, kemayo: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there * 20:04 kemayo@deploy2003: Started scap sync-world: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] * 18:22 cwhite: lvextend vg0/srv +500g on centrallog hosts * 18:19 cdobbins@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS trixie * 17:46 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1071.eqiad.wmnet * 17:46 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1071.eqiad.wmnet * 17:13 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1071.eqiad.wmnet with reason: vacuum overlarge container dbs * 17:07 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1065.eqiad.wmnet * 17:07 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1065.eqiad.wmnet * 17:06 dzahn@dns1006: END - running authdns-update * 17:04 dzahn@dns1006: START - running authdns-update * 17:01 dzahn@dns1006: END - running authdns-update * 16:59 dzahn@dns1006: START - running authdns-update * 16:51 dancy@deploy2003: Finished scap sync-world: testing [[phab:T428971|T428971]] (duration: 03m 37s) * 16:47 dancy@deploy2003: Started scap sync-world: testing [[phab:T428971|T428971]] * 16:45 atsukoito: restarting pybal on lvs1019 to flush IP address for `cirrussearch1122.eqiad.wmnet` after moving the vlan [[phab:T431311|T431311]] * 16:42 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:42 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:42 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:42 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:42 Amir1: mwscript-k8s --follow --dblist=ores -- extensions/ORES/maintenance/PurgeScoreCache.php --model damaging --old ([[phab:T431159|T431159]]) * 16:34 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Pool test * 16:34 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 16:34 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 16:34 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Pool test * 16:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Depool test * 16:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 16:33 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 16:33 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Depool test * 16:31 dancy@deploy2003: Installation of scap version "4.273.0" completed for 159 hosts * 16:29 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1065.eqiad.wmnet with reason: vacuum overlarge container dbs * 16:27 dancy@deploy2003: Installing scap version "4.273.0" for 159 host(s) * 16:27 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics-external: sync * 16:27 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics-external: sync * 16:26 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics-external: sync * 16:26 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics-external: sync * 16:22 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync * 16:21 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync * 16:21 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: sync * 16:21 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: sync * 16:19 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync * 16:19 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync * 16:18 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync * 16:17 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync * 15:59 atsukoito: restarting pybal on lvs1018 for https://gerrit.wikimedia.org/r/1310117 * 15:55 aikochou@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 15:50 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310117 * 15:46 aikochou@deploy2003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 15:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host kafka-logging1006.eqiad.wmnet * 15:43 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host kafka-logging1006.eqiad.wmnet * 15:41 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host ganeti-test[2001-2003].codfw.wmnet * 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host ganeti-test[2001-2003].codfw.wmnet * 15:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host netbox1003.eqiad.wmnet * 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host netbox1003.eqiad.wmnet * 15:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host netbox2003.codfw.wmnet * 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host netbox2003.codfw.wmnet * 15:36 sukhe: restart pybal on lvs1020 * 15:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet * 15:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet * 15:08 btullis@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:06 btullis@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 15:01 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:01 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:35 cdobbins@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 14:34 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:33 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:33 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:32 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:29 cdobbins@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 14:28 marostegui@dns1004: END - running authdns-update * 14:28 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:27 marostegui@dns1004: START - running authdns-update * 14:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1023.eqiad.wmnet with reason: reboot * 14:18 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2009.codfw.wmnet with OS trixie * 14:14 swfrench-wmf: start rolling run-puppet-agent on A:cp for ATS config change - [[phab:T428909|T428909]] [[phab:T431838|T431838]] * 14:05 swfrench-wmf: disable-puppet on A:cp for ATS config change - [[phab:T428909|T428909]] [[phab:T431838|T431838]] * 14:05 cdobbins@cumin2002: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie * 14:02 marostegui@dns1004: END - running authdns-update * 14:00 marostegui@dns1004: START - running authdns-update * 14:00 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1070.eqiad.wmnet * 14:00 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1070.eqiad.wmnet * 13:58 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2009.codfw.wmnet with reason: host reimage * 13:57 cdobbins@cumin2002: conftool action : set/pooled=no; selector: name=dns7002.* * 13:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2009.codfw.wmnet with reason: host reimage * 13:48 rscout@deploy2003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply * 13:48 rscout@deploy2003: helmfile [eqiad] START helmfile.d/services/miscweb: apply * 13:48 rscout@deploy2003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply * 13:47 rscout@deploy2003: helmfile [codfw] START helmfile.d/services/miscweb: apply * 13:40 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:33 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2009.codfw.wmnet with OS trixie * 13:30 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1070.eqiad.wmnet with reason: vacuum overlarge container dbs * 13:28 aude@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] (duration: 11m 12s) * 13:23 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:22 aude@deploy2003: aikochou, javiermonton, aude, gkm563: Continuing with deployment * 13:22 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:19 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:19 aude@deploy2003: aikochou, javiermonton, aude, gkm563: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] synced to the testservers * 13:17 aude@deploy2003: Started scap sync-world: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] * 13:01 ladsgroup@deploy2003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 13:01 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:00 ladsgroup@deploy2003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 12:59 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:52 ladsgroup@deploy2003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 12:51 ladsgroup@deploy2003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 12:48 atsuko@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 12:48 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 12:47 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2008.codfw.wmnet with OS trixie * 12:47 atsuko@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 12:47 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply * 12:47 atsuko@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:46 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 12:45 atsuko@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:45 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply * 12:45 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:44 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:43 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] (duration: 07m 02s) * 12:38 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 12:37 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:36 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] * 12:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2008.codfw.wmnet with reason: host reimage * 12:23 Msz2001: Deployed changes to private code for Suggested Investigations * 12:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2008.codfw.wmnet with reason: host reimage * 12:20 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:19 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:17 atsuko@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 12:17 atsuko@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 12:16 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:15 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] (duration: 07m 14s) * 12:10 mszwarc@deploy2003: mszwarc: Continuing with deployment * 12:09 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:07 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] * 12:04 mszwarc@deploy2003: sync-world aborted: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] (duration: 00m 29s) * 12:03 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] * 12:01 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2008.codfw.wmnet with OS trixie * 12:00 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] (duration: 07m 37s) * 11:55 zabe@deploy2003: zabe: Continuing with deployment * 11:54 zabe@deploy2003: zabe: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:52 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] * 11:51 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:43 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:35 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:34 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:33 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:30 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:28 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:27 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:17 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2007.codfw.wmnet with OS trixie * 11:09 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] * 11:06 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=s8 * 11:00 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=x3 * 11:00 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=s5 * 10:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2007.codfw.wmnet with reason: host reimage * 10:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2007.codfw.wmnet with reason: host reimage * 10:51 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host dse-k8s-worker1023 * 10:50 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host dse-k8s-worker1023 * 10:44 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dse-k8s-worker1023 * 10:43 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host dse-k8s-worker1023 * 10:42 marostegui@cumin1003: dbctl commit (dc=all): 'Change x4 masters [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P94804 and previous config saved to /var/cache/conftool/dbconfig/20260713-104248-marostegui.json * 10:37 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:37 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:35 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:35 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:34 atsuko@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 10:34 atsuko@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 10:33 marostegui@cumin1003: dbctl commit (dc=all): 'Push x4 initial dbctl config [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P94803 and previous config saved to /var/cache/conftool/dbconfig/20260713-103259-marostegui.json * 10:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2007.codfw.wmnet with OS trixie * 09:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2006.codfw.wmnet with OS trixie * 09:42 marostegui@dns1004: END - running authdns-update * 09:40 marostegui@dns1004: START - running authdns-update * 09:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2006.codfw.wmnet with reason: host reimage * 09:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2006.codfw.wmnet with reason: host reimage * 09:06 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1024.eqiad.wmnet with reason: reboot * 09:06 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:01 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2006.codfw.wmnet with OS trixie * 08:44 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 08:43 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 08:43 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:42 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:42 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:42 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:41 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 08:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup2004.codfw.wmnet * 08:38 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:33 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host db1208.eqiad.wmnet * 08:30 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=x3 * 08:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2005.codfw.wmnet with OS trixie * 08:28 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup2004.codfw.wmnet * 08:28 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup2003.codfw.wmnet * 08:24 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1039: Repooling after testing * 08:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on clouddb1016.eqiad.wmnet with reason: cloning * 08:23 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s5 * 08:23 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s8 * 08:21 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 08:21 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 08:17 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup2003.codfw.wmnet * 08:17 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1004.eqiad.wmnet * 08:14 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1208.eqiad.wmnet * 08:11 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host phab1005.eqiad.wmnet * 08:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2005.codfw.wmnet with reason: host reimage * 08:07 marostegui@dns1004: END - running authdns-update * 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1004.eqiad.wmnet * 08:07 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1003.eqiad.wmnet * 08:07 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:05 marostegui@dns1004: START - running authdns-update * 08:05 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2005.codfw.wmnet with reason: host reimage * 08:05 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 08:05 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host phab1005.eqiad.wmnet * 08:05 marostegui@dns1004: START - running authdns-update * 08:05 marostegui@dns1004: START - running authdns-update * 08:05 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 08:04 marostegui@dns1004: START - running authdns-update * 08:00 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit1003.wikimedia.org * 07:58 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1003.eqiad.wmnet * 07:58 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1002-dev.eqiad.wmnet * 07:58 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:58 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:54 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1002-dev.eqiad.wmnet * 07:54 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1001-dev.eqiad.wmnet * 07:54 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit1003.wikimedia.org * 07:53 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 07:53 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:52 Msz2001: UTC morning backport+config window done * 07:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2005.codfw.wmnet with OS trixie * {{safesubst:SAL entry|1=07:50 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark (T429943}} * 07:49 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1001-dev.eqiad.wmnet * 07:46 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:46 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:45 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 07:45 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:44 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 07:44 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:43 mszwarc@deploy2003: mszwarc, danielyepezgarces, anzx: Continuing with deployment * {{safesubst:SAL entry|1=07:39 mszwarc@deploy2003: mszwarc, danielyepezgarces, anzx: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark}} * 07:39 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1039: Repooling after testing * {{safesubst:SAL entry|1=07:36 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark (T429943)}} * 07:35 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] (duration: 30m 03s) * 07:25 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit2002.wikimedia.org * 07:22 mszwarc@deploy2003: mszwarc: Continuing with deployment * 07:21 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:19 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit2002.wikimedia.org * 07:15 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aphlict1002.eqiad.wmnet * 07:11 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host aphlict1002.eqiad.wmnet * 07:08 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2003.wikimedia.org * 07:05 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] * 07:02 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2003.wikimedia.org * 07:02 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2002.wikimedia.org * 06:55 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2002.wikimedia.org * 06:55 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1003.wikimedia.org * 06:49 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1003.wikimedia.org * 06:34 marostegui: Drop m5 ipoid database [[phab:T431007|T431007]] * 06:29 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1027.eqiad.wmnet with reason: reboot * 06:24 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1028.eqiad.wmnet with reason: reboot * 06:21 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1025.eqiad.wmnet with reason: reboot * 06:17 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1022.eqiad.wmnet with reason: reboot * 06:03 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on dbproxy[2005-2008].codfw.wmnet with reason: reboot * 05:37 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1217,1228].eqiad.wmnet with reason: cloning * 05:11 marostegui: Drop users_to_rename table [[phab:T431842|T431842]] * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-12 == * 16:01 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2209 [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94792 and previous config saved to /var/cache/conftool/dbconfig/20260712-160124-marostegui.json * 15:58 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2205 to s3 primary [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94791 and previous config saved to /var/cache/conftool/dbconfig/20260712-155853-marostegui.json * 15:58 marostegui: Starting s3 codfw emergency failover from db2209 to db2205 - [[phab:T431950|T431950]] * 15:51 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2205 with weight 0 [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94790 and previous config saved to /var/cache/conftool/dbconfig/20260712-155135-marostegui.json * 15:51 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Primary switchover s3 [[phab:T431950|T431950]] * 02:01 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 01m 17s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-11 == * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 26s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-10 == * 19:12 jhathaway@dns1004: END - running authdns-update * 19:10 jhathaway@dns1004: START - running authdns-update * 18:23 mutante: vrts2002 rebooting (not the active host) * 18:21 mutante: lists2001, phab2003 - rebooting (not the active hosts) * 18:16 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on A:lvs-high-traffic2-codfw * 18:15 swfrench@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on A:lvs-high-traffic2-codfw * 17:15 mutante: [doc1004:~] $ sudo systemctl start rsync-doc-host-data-sync ([[phab:T431856|T431856]]) * 17:09 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1004.eqiad.wmnet * 17:08 jhathaway@dns1004: END - running authdns-update * 17:07 jhathaway@dns1004: START - running authdns-update * 17:06 jhathaway: depooling puppetserver1002, cause of errors is still unknown * 17:03 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1004.eqiad.wmnet * 16:57 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1003.eqiad.wmnet * 16:51 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1003.eqiad.wmnet * 16:48 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 16:48 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2004.codfw.wmnet * 16:42 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2004.codfw.wmnet * 16:41 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2003.codfw.wmnet * 16:35 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2003.codfw.wmnet * 16:33 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2002.codfw.wmnet * 16:27 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2002.codfw.wmnet * 16:25 mutante: gitlab-runners (production) rebooting cluster one by one * 16:17 mutante: etherpad1004/etherpad2002 - (etherpad.wikimedia.org) - rebooting * 16:13 mutante: doc1004/doc2003 (doc.wikimedia.org backends) - rebooting * 16:02 mutante: releases1003/releases2003 (releases.wikimedia.org backends) - rebooting for maintenance * 15:26 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2007-dev.codfw.wmnet * 15:19 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2007-dev.codfw.wmnet * 15:14 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host cloudcephosd2007-dev.codfw.wmnet * 15:14 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2007-dev.codfw.wmnet * 15:14 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host cloudcephosd2006-dev.codfw.wmnet * 15:07 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2006-dev.codfw.wmnet * 15:07 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2005-dev.codfw.wmnet * 14:59 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2005-dev.codfw.wmnet * 14:59 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2004-dev.codfw.wmnet * 14:53 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2004-dev.codfw.wmnet * 14:53 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2007-dev.codfw.wmnet * 14:51 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1054.eqiad.wmnet * 14:51 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1054.eqiad.wmnet * 14:51 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1054.eqiad.wmnet * 14:47 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2007-dev.codfw.wmnet * 14:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2006-dev.codfw.wmnet * 14:41 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2006-dev.codfw.wmnet * 14:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2005-dev.codfw.wmnet * 14:37 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2005-dev.codfw.wmnet * 14:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2005-dev.codfw.wmnet * 14:29 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2005-dev.codfw.wmnet * 14:29 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2006-dev.codfw.wmnet * 14:21 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2006-dev.codfw.wmnet * 14:21 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2010-dev.codfw.wmnet * 14:15 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2010-dev.codfw.wmnet * 14:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudgw2004-dev.codfw.wmnet * 14:10 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1054.eqiad.wmnet with OS trixie * 14:09 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudgw2004-dev.codfw.wmnet * 14:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudgw2003-dev.codfw.wmnet * 14:02 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudgw2003-dev.codfw.wmnet * 14:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2004-dev.codfw.wmnet * 13:53 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2004-dev.codfw.wmnet * 13:53 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2003-dev.codfw.wmnet * 13:48 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:44 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2003-dev.codfw.wmnet * 13:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2002-dev.codfw.wmnet * 13:42 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:41 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:41 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:37 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2002-dev.codfw.wmnet * 13:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudidp2001-dev.codfw.wmnet * 13:33 blake@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker1054.eqiad.wmnet with reason: host reimage * 13:33 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudidp2001-dev.codfw.wmnet * 13:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudnet2006-dev.codfw.wmnet * 13:26 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudnet2006-dev.codfw.wmnet * 13:26 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudnet2005-dev.codfw.wmnet * 13:23 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1054.eqiad.wmnet with reason: host reimage * 13:18 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudnet2005-dev.codfw.wmnet * 13:18 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudservices2005-dev.codfw.wmnet * 13:12 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudservices2005-dev.codfw.wmnet * 13:11 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudservices2004-dev.codfw.wmnet * 13:08 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudservices2004-dev.codfw.wmnet * 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudweb2002-dev.wikimedia.org * 13:05 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 13:05 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1054 * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1054 * 13:04 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1054 * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1054.eqiad.wmnet 49.32.64.10.in-addr.arpa 9.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:04 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1054.eqiad.wmnet 49.32.64.10.in-addr.arpa 9.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1054 - blake@cumin1003" * 13:04 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1054 - blake@cumin1003" * 13:01 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudweb2002-dev.wikimedia.org * 13:00 blake@cumin1003: START - Cookbook sre.dns.netbox * 12:59 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1054 * 12:57 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1054.eqiad.wmnet with OS trixie * 12:57 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1054.eqiad.wmnet * 12:56 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1054.eqiad.wmnet * 12:56 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1054.eqiad.wmnet * 12:47 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS trixie * 12:44 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:39 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 12:39 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 12:38 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:37 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:14 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:10 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:08 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:07 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:00 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:00 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:51 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:49 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:48 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:47 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:44 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:32 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2001.codfw.wmnet * 11:32 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1053.eqiad.wmnet * 11:32 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2001.codfw.wmnet * 11:32 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1053.eqiad.wmnet * 11:32 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1053.eqiad.wmnet * 11:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker2001.codfw.wmnet * 11:31 cgoubert@cumin1003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker2001.codfw.wmnet * 11:31 cgoubert@cumin1003: END (FAIL) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=1) rolling reimage on P<nowiki>{</nowiki>wikikube-worker2001*<nowiki>}</nowiki> and (A:wikikube-master-codfw or A:wikikube-worker-codfw) * 11:30 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:30 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:21 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 18 hosts with reason: reboot & upgrade * 11:20 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker2001.codfw.wmnet with OS trixie * 11:16 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:15 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:14 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:14 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 11 hosts * 11:14 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 11 hosts * 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:08 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 11:02 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1053.eqiad.wmnet with OS trixie * 11:01 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:58 cgoubert@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 10:57 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:38 cgoubert@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker2001.codfw.wmnet with OS trixie * 10:38 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2001.codfw.wmnet * 10:38 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2001.codfw.wmnet * 10:38 cgoubert@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on P<nowiki>{</nowiki>wikikube-worker2001*<nowiki>}</nowiki> and (A:wikikube-master-codfw or A:wikikube-worker-codfw) * 10:35 topranks: adjust IBGP outbound policy on lsw1-e2-codfw [[phab:T423430|T423430]] towards ssw1-e1-codfw * 10:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 cgoubert@cumin1003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:27 cgoubert@cumin1003: END (FAIL) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=1) rolling reimage on A:wikikube-worker-codfw * 10:27 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker2001.codfw.wmnet with OS bookworm * 10:25 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 10:24 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:24 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:15 cgoubert@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 10:11 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 10:11 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 10:08 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:08 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:07 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 10:06 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 10:00 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:55 cgoubert@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker2001.codfw.wmnet with OS bookworm * 09:55 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2005-2006,2011-2012].codfw.wmnet * 09:55 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2005-2006,2011-2012].codfw.wmnet * 09:51 cgoubert@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on A:wikikube-worker-codfw * 09:41 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1053.eqiad.wmnet with reason: host reimage * 09:37 topranks: apply new IBGP outbound policy on lsw1-e2-codfw [[phab:T423430|T423430]] * 09:36 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:36 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1053.eqiad.wmnet with reason: host reimage * 09:16 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1053 * 09:16 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1053 * 09:15 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1053 * 09:15 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1053.eqiad.wmnet 48.32.64.10.in-addr.arpa 8.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:15 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1053.eqiad.wmnet 48.32.64.10.in-addr.arpa 8.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:15 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:15 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1053 - blake@cumin1003" * 09:15 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1053 - blake@cumin1003" * 09:11 blake@cumin1003: START - Cookbook sre.dns.netbox * 09:11 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1053 * 09:08 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1053.eqiad.wmnet with OS trixie * 09:08 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1053.eqiad.wmnet * 09:08 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1053.eqiad.wmnet * 09:08 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1053.eqiad.wmnet * 09:04 brouberol@dns1004: END - running authdns-update * 09:03 brouberol@dns1004: START - running authdns-update * 08:41 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e] (thin): Regular analytics weekly train THIN [analytics/refinery@1abf22ea] (duration: 02m 11s) * 08:38 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e] (thin): Regular analytics weekly train THIN [analytics/refinery@1abf22ea] * 08:38 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e]: Regular analytics weekly train [analytics/refinery@1abf22ea] (duration: 05m 17s) * 08:38 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:34 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:33 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e]: Regular analytics weekly train [analytics/refinery@1abf22ea] * 08:32 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@1abf22ea] (duration: 02m 03s) * 08:30 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@1abf22ea] * 08:30 JavierMonton: Deploying Refinery at {{Gerrit|1abf22ea}} for changes 1308121/T427068 1306491/T430020 and {{Gerrit|1308190}} * 08:29 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:29 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 08:24 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:24 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 08:18 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:18 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 08:00 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db[2183-2184].codfw.wmnet * 08:00 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for db[2183-2184].codfw.wmnet * 07:52 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:52 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 07:49 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 11 hosts with reason: reboot & upgrade * 07:47 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:47 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 07:44 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:44 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 07:23 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 10 hosts * 07:23 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 10 hosts * 06:45 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 10 hosts with reason: reboot & upgrade * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 41s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-09 == * 23:33 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] (duration: 13m 26s) * 23:29 ladsgroup@deploy2003: ladsgroup, jdlrobson: Continuing with deployment * 23:22 ladsgroup@deploy2003: ladsgroup, jdlrobson: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:20 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] * 22:57 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1165.eqiad.wmnet * 22:56 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1165.eqiad.wmnet * 22:56 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1165.eqiad.wmnet * 22:45 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1165.eqiad.wmnet with OS trixie * 22:38 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 22:37 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 22:37 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 22:37 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:37 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 22:25 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1165.eqiad.wmnet with reason: host reimage * 22:17 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1165.eqiad.wmnet with reason: host reimage * 22:13 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 22:12 rzl@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 22:04 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 22:04 rzl@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1165 * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1165 * 22:02 jasmine@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1165 * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1165.eqiad.wmnet 115.48.64.10.in-addr.arpa 5.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:02 jasmine@cumin2002: START - Cookbook sre.dns.wipe-cache wikikube-worker1165.eqiad.wmnet 115.48.64.10.in-addr.arpa 5.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1165 - jasmine@cumin2002" * 22:02 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1165 - jasmine@cumin2002" * 22:02 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 21:57 jasmine@cumin2002: START - Cookbook sre.dns.netbox * 21:55 jasmine@cumin2002: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1165 * 21:54 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-worker1165.eqiad.wmnet with OS trixie * 21:54 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 21:54 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1165.eqiad.wmnet * 21:53 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 21:53 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1165.eqiad.wmnet * 21:53 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1165.eqiad.wmnet * 21:53 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 21:47 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 21:45 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 21:43 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 21:43 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 21:42 maryum: Deploy fix for [[phab:T431684|T431684]] * 21:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 21:27 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] (duration: 34m 14s) * 21:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs1002 * 21:23 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs1002 * 21:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS trixie * 21:22 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 21:20 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 22s) * 21:20 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 21:16 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 21:16 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 21:15 ladsgroup@deploy2003: ladsgroup: Continuing with deployment * 21:13 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2002.codfw.wmnet with OS bookworm * 21:11 ladsgroup@deploy2003: ladsgroup: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:08 ladsgroup@cumin1003: END (PASS) - Cookbook sre.wikireplicas.update-views (exit_code=0) * 21:07 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 6 hosts with reason: reboots * 20:54 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecycle work - bking@cumin2003 * 20:53 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:53 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] * 20:51 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) * 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 20:47 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecycle work - bking@cumin2003 * 20:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 20:41 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:41 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) * 20:40 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host relforge1008.eqiad.wmnet * 20:40 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1009.eqiad.wmnet with OS trixie * 20:33 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:32 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:32 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:31 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:31 ladsgroup@cumin1003: END (PASS) - Cookbook sre.wikireplicas.update-views (exit_code=0) * 20:29 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1008.eqiad.wmnet * 20:24 rzl@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 20:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2002.codfw.wmnet with OS bookworm * 20:23 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host relforge1008.eqiad.wmnet * 20:23 rzl@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 20:23 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1008.eqiad.wmnet * 20:22 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:22 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:21 bking@cumin2003: END (ERROR) - Cookbook sre.elasticsearch.rolling-operation (exit_code=97) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:21 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1009.eqiad.wmnet with reason: host reimage * 20:16 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:15 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1009.eqiad.wmnet with reason: host reimage * 20:12 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) * 20:02 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 19:55 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1009.eqiad.wmnet with OS trixie * 19:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 19:43 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 19:30 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 19:28 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 19:27 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 19:25 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 18:42 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 18:41 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 18:16 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 18:15 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 17:45 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for doh5004.wikimedia.org * 17:45 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for doh5004.wikimedia.org * 17:38 ladsgroup@deploy2003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 17:35 ladsgroup@deploy2003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 17:29 ladsgroup@deploy2003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 17:26 ladsgroup@deploy2003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 17:09 mutante: zuul[12]00[123] - rebooting for maintenance * 17:09 ebernhardson: start full in-place reindex of eqiad cirrussearch cluster * 17:08 dzahn@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-cluster (exit_code=99) * 17:08 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-cluster * 17:03 ebernhardson: start full in-place reindex of codfw cirrussearch cluster * 16:59 mutante: stewards1001/stewards2001 - reboot for maintenance * 16:54 ebernhardson: start full in-place reindex of cloudelastic cluster * 16:53 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 16:52 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply * 16:49 mutante: planet1003/planet2003 - rebooting * 16:47 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on doh5004.wikimedia.org with reason: random high load, investigating * 15:55 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 15:54 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 15:51 jynus: restarting backupmon1001 * 15:49 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 14 hosts * 15:49 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 14 hosts * 15:47 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backupmon1001.eqiad.wmnet with reason: restart * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:06 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 15:06 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 14:59 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply * 14:58 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply * 14:51 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 14 hosts * 14:51 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 14 hosts * 14:49 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 6 hosts with reason: reboot & upgrade * 14:48 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet * 14:48 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet * 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:42 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1052.eqiad.wmnet * 14:42 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1052.eqiad.wmnet * 14:42 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1052.eqiad.wmnet * 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:31 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:31 elukey@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: sync * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:30 elukey@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: sync * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:28 elukey@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: sync * 14:28 elukey@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: sync * 14:26 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:20 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1052.eqiad.wmnet with OS trixie * 14:19 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 6 hosts with reason: reboot & upgrade * 14:18 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:15 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:15 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:13 elukey: update druid indexation job for webrequest_sampled_live - [[phab:T427068|T427068]] * 14:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:09 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:09 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for papaul - jhancock@cumin2002" * 14:09 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for papaul - jhancock@cumin2002" * 14:07 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:07 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:04 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 14:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cuminunpriv1001.eqiad.wmnet * 13:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb1003.eqiad.wmnet * 13:59 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1052.eqiad.wmnet with reason: host reimage * 13:57 moritzm: installing requests security updates * 13:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cuminunpriv1001.eqiad.wmnet * 13:55 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb1003.eqiad.wmnet * 13:53 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1052.eqiad.wmnet with reason: host reimage * 13:50 moritzm: installing python-cryptography security updates * 13:47 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb2003.codfw.wmnet * 13:44 Msz2001: UTC afternoon config+backport window is done * 13:44 Msz2001: Updated `logging` on `metawiki` to fix log performers, [[phab:T431176|T431176]]#12105297 * 13:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb2003.codfw.wmnet * 13:43 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 13:43 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt1002.wikimedia.org * 13:41 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] (duration: 07m 30s) * 13:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt1002.wikimedia.org * 13:37 mszwarc@deploy2003: mszwarc: Continuing with deployment * 13:36 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1052 * 13:36 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1052 * 13:35 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:35 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1052 * 13:35 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1052.eqiad.wmnet 47.32.64.10.in-addr.arpa 7.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:35 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1052.eqiad.wmnet 47.32.64.10.in-addr.arpa 7.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:35 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:35 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1052 - blake@cumin1003" * 13:35 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1052 - blake@cumin1003" * 13:34 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] * 13:31 blake@cumin1003: START - Cookbook sre.dns.netbox * 13:31 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1052 * 13:30 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1052.eqiad.wmnet with OS trixie * 13:30 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1052.eqiad.wmnet * 13:29 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1052.eqiad.wmnet * 13:29 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1052.eqiad.wmnet * 13:17 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] (duration: 11m 26s) * 13:13 jforrester@deploy2003: jforrester: Continuing with deployment * 13:08 jforrester@deploy2003: jforrester: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:06 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] * 12:54 cgoubert@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply * 12:54 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:52 cgoubert@deploy2003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply * 12:45 cgoubert@deploy2003: helmfile [codfw] DONE helmfile.d/services/mobileapps: apply * 12:44 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:44 cgoubert@deploy2003: helmfile [codfw] START helmfile.d/services/mobileapps: apply * 12:43 cgoubert@deploy2003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 12:43 cgoubert@deploy2003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 12:42 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast4006.wikimedia.org * 12:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt2002.wikimedia.org * 12:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast7002.wikimedia.org * 12:18 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast4006.wikimedia.org * 12:18 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host ml-serve1004 * 12:18 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host ml-serve1004 * 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt2002.wikimedia.org * 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast7002.wikimedia.org * 12:10 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backup[2003,2014].codfw.wmnet with reason: reboot & upgrade * 12:10 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt-staging2001.codfw.wmnet * 12:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid1003.eqiad.wmnet * 12:06 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt-staging2001.codfw.wmnet * 12:05 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid1003.eqiad.wmnet * 12:03 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backup[1003,1014].eqiad.wmnet with reason: reboot & upgrade * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid2003.codfw.wmnet * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host irc1003.wikimedia.org * 11:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid2003.codfw.wmnet * 11:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host irc1003.wikimedia.org * 11:55 jmm@dns1004: END - running authdns-update * 11:53 jmm@dns1004: START - running authdns-update * 11:50 jmm@dns1004: END - running authdns-update * 11:48 jmm@dns1004: START - running authdns-update * 11:27 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host irc2003.wikimedia.org * 11:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host irc2003.wikimedia.org * 11:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint2001.codfw.wmnet * 11:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint1001.eqiad.wmnet * 11:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint2001.codfw.wmnet * 11:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint1001.eqiad.wmnet * 11:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-rw2001.wikimedia.org * 11:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-rw1001.wikimedia.org * 11:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-rw2001.wikimedia.org * 11:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-rw1001.wikimedia.org * 11:03 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon1003.wikimedia.org * 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2005.codfw.wmnet * 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2005.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 10:59 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2005.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 10:57 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon1003.wikimedia.org * 10:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon2002.wikimedia.org * 10:55 jmm@cumin2003: START - Cookbook sre.dns.netbox * 10:51 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon2002.wikimedia.org * 10:51 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:50 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2005.codfw.wmnet * 10:41 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:40 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host ml-serve1003 * 10:40 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host ml-serve1003 * 10:39 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2033.codfw.wmnet * 10:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install2005.wikimedia.org * 10:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install1005.wikimedia.org * 10:35 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1004.eqiad.wmnet with OS bookworm * 10:31 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install1005.wikimedia.org * 10:31 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install2005.wikimedia.org * 10:30 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install4004.wikimedia.org * 10:30 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install3004.wikimedia.org * 10:29 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install3004.wikimedia.org * 10:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install4004.wikimedia.org * 10:23 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 10:21 moritzm: failover Ganeti master in codfw/routed to ganeti2034 [[phab:T430928|T430928]] * 10:19 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.addnode (exit_code=0) for new host ganeti2031.codfw.wmnet to cluster codfw and group B * 10:19 moritzm: readded ganeti2031 to the codfw Ganeti cluster [[phab:T430910|T430910]] * 10:18 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1004.eqiad.wmnet with reason: host reimage * 10:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install5004.wikimedia.org * 10:18 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1003 * 10:18 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1003 * 10:17 jmm@cumin2003: START - Cookbook sre.ganeti.addnode for new host ganeti2031.codfw.wmnet to cluster codfw and group B * 10:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install6003.wikimedia.org * 10:16 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install5004.wikimedia.org * 10:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install6003.wikimedia.org * 10:15 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1004.eqiad.wmnet with reason: host reimage * 10:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1001.eqiad.wmnet * 10:14 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 10:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2008.wikimedia.org * 10:00 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ml-serve1004 * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1004 * 09:57 jmm@cumin2003: START - Cookbook sre.dns.netbox * 09:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install7002.wikimedia.org * 09:57 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1004 * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ml-serve1004.eqiad.wmnet 50.48.64.10.in-addr.arpa 0.5.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:57 klausman@cumin1003: START - Cookbook sre.dns.wipe-cache ml-serve1004.eqiad.wmnet 50.48.64.10.in-addr.arpa 0.5.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1004 - klausman@cumin1003" * 09:56 klausman@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1004 - klausman@cumin1003" * 09:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-coord1001.eqiad.wmnet * 09:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 09:55 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow7002.magru.wmnet * 09:52 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-coord1001.eqiad.wmnet * 09:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 09:52 klausman@cumin1003: START - Cookbook sre.dns.netbox * 09:50 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install7002.wikimedia.org * 09:50 klausman@cumin1003: START - Cookbook sre.hosts.move-vlan for host ml-serve1004 * 09:50 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1004.eqiad.wmnet with OS bookworm * 09:50 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1003.eqiad.wmnet with OS bookworm * 09:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1001.eqiad.wmnet * 09:49 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 09:49 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2008.wikimedia.org * 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2007.codfw.wmnet * 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2007.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 09:49 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow7002.magru.wmnet * 09:49 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2007.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 09:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard1003.eqiad.wmnet * 09:39 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard2003.codfw.wmnet * 09:39 jmm@cumin2003: START - Cookbook sre.dns.netbox * 09:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard1003.eqiad.wmnet * 09:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor1003.eqiad.wmnet * 09:35 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard2003.codfw.wmnet * 09:34 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2007.codfw.wmnet * 09:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor1003.eqiad.wmnet * 09:33 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor-dev2001.codfw.wmnet * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor2003.codfw.wmnet * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sretest1006.eqiad.wmnet * 09:27 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 09:25 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor-dev2001.codfw.wmnet * 09:25 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor2003.codfw.wmnet * 09:23 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] (duration: 06m 27s) * 09:23 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2205: codfw rack B4 repool after maintenance * 09:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host sretest1006.eqiad.wmnet * 09:23 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2204: codfw rack B4 repool after maintenance * 09:19 urbanecm@deploy2003: urbanecm: Continuing with deployment * 09:19 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:18 jmm@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 6 hosts with reason: reboot * 09:17 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] * 09:08 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ml-serve1003 * 09:08 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1003 * 09:07 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1003 * 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ml-serve1003.eqiad.wmnet 81.32.64.10.in-addr.arpa 1.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:07 klausman@cumin1003: START - Cookbook sre.dns.wipe-cache ml-serve1003.eqiad.wmnet 81.32.64.10.in-addr.arpa 1.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1003 - klausman@cumin1003" * 09:06 klausman@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1003 - klausman@cumin1003" * 08:58 klausman@cumin1003: START - Cookbook sre.dns.netbox * 08:57 klausman@cumin1003: START - Cookbook sre.hosts.move-vlan for host ml-serve1003 * 08:57 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1003.eqiad.wmnet with OS bookworm * 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=0) rolling reimage on P<nowiki>{</nowiki>ml-serve1003.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet * 08:55 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet * 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1003.eqiad.wmnet with OS bookworm * 08:39 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 08:38 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool db2205: codfw rack B4 repool after maintenance * 08:37 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool db2204: codfw rack B4 repool after maintenance * 08:36 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 08:35 hashar@deploy2003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 08:32 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:32 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:31 hashar@deploy2003: Rolling back deployment * 08:26 moritzm: failover Ganeti master in codfw to ganeti2048 [[phab:T430928|T430928]] * 08:16 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1003.eqiad.wmnet with OS bookworm * 08:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2004.codfw.wmnet * 08:16 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet * 08:16 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet * 08:16 klausman@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on P<nowiki>{</nowiki>ml-serve1003.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 08:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2002.codfw.wmnet * 08:15 XioNoX: lsw1-b4-codfw> request system reboot - [[phab:T430910|T430910]] * 08:15 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b4-codfw,lsw1-b4-codfw IPv6,lsw1-b4-codfw.mgmt with reason: Switch maintenance * 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for codfw rack B4 * 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:10 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2004.codfw.wmnet * 08:10 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2002.codfw.wmnet * 08:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2205: codfw rack B4 depool for maintenance * 08:08 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool db2205: codfw rack B4 depool for maintenance * 08:08 jmm@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin2003.codfw.wmnet * 08:08 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2204: codfw rack B4 depool for maintenance * 08:08 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool db2204: codfw rack B4 depool for maintenance * 08:08 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 27 hosts with reason: codfw rack B4 depool for maintenance * 08:03 jmm@cumin2002: START - Cookbook sre.hosts.reboot-single for host cumin2003.codfw.wmnet * 07:56 ayounsi@cumin1003: START - Cookbook sre.network.depool-rack with action 'depool' for codfw rack B4 * 07:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1008.eqiad.wmnet with OS trixie * 07:49 wmde-fisch@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] (duration: 08m 36s) * 07:44 wmde-fisch@deploy2003: wmde-fisch: Continuing with deployment * 07:43 wmde-fisch@deploy2003: wmde-fisch: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:41 wmde-fisch@deploy2003: Started scap sync-world: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] * 07:35 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1008.eqiad.wmnet with reason: host reimage * 07:31 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1008.eqiad.wmnet with reason: host reimage * 07:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1008.eqiad.wmnet with OS trixie * 07:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 07:00 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 06:59 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1008.eqiad.wmnet with OS trixie * 06:57 Emperor: rebalance thanos swift rings after previous re-image of thanos-fe1004 to trixie * 06:47 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1008.eqiad.wmnet with OS trixie * 04:10 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 14 days, 0:00:00 on cp6008.drmrs.wmnet with reason: Hardware failure - [[phab:T431651|T431651]] * 03:55 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp6008.* * 03:29 ryankemper: [[phab:T431311|T431311]] Repooled eqiad cirrussearch clusters (`chi/omega/psi`) following completion of OpenSearch 2.19 migration * 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad * 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=eqiad * 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 31s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-08 == * 23:52 Amir1: ladsgroup@deploy2003:~$ mwscript-k8s --follow -- extensions/ORES/maintenance/PurgeScoreCache.php --wiki=simplewiki --model damaging --old ([[phab:T431159|T431159]]) * 23:46 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 23:46 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing PTR for 2001:df2:e500:fe08::1 - cmooney@cumin1003" * 23:46 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing PTR for 2001:df2:e500:fe08::1 - cmooney@cumin1003" * 23:40 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 23:16 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 23:15 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 22:42 rzl@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 22:40 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] (duration: 12m 55s) * 22:40 rzl@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 22:37 rzl@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 22:36 rzl@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 22:35 rzl@deploy2003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 22:34 urbanecm@deploy2003: urbanecm: Continuing with deployment * 22:33 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:33 rzl@deploy2003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 22:32 rzl@deploy2003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 22:30 rzl@deploy2003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 22:30 rzl@deploy2003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 22:29 rzl@deploy2003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 22:27 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] * 22:26 rzl@deploy2003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 22:22 rzl@deploy2003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 22:21 rzl@deploy2003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 22:19 rzl@deploy2003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 22:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 22:17 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 22:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 22:13 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 22:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 22:13 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 22:09 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 22:06 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 22:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1094.eqiad.wmnet with OS trixie * 22:01 urbanecm: Make https://test.wikipedia.org/w/index.php?title=MediaWiki:GrowthExperimentsSuggestedEdits.json&diff=prev&oldid=750552 with GrowthExperiments disabled (via mw-experimental), then run `\MediaWiki\MediaWikiServices::getInstance()->get('CommunityConfiguration.ProviderFactory')->newProvider('GrowthSuggestedEdits')->getStore()->invalidate()` ([[phab:T431625|T431625]]) * 21:56 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d2-codfw * 21:55 urbanecm@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 21:55 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d2-codfw * 21:55 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c4-codfw * 21:55 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c4-codfw * 21:55 urbanecm@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2002 * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2002 * 21:54 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2002 * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2002.codfw.wmnet 50.32.192.10.in-addr.arpa 0.5.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:54 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2002.codfw.wmnet 50.32.192.10.in-addr.arpa 0.5.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2002 - bking@cumin2003" * 21:54 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2002 - bking@cumin2003" * 21:49 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:49 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2002 * 21:49 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2002.codfw.wmnet with OS trixie * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1094.eqiad.wmnet with reason: host reimage * 21:42 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 21:39 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 21:37 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1094.eqiad.wmnet with reason: host reimage * 21:36 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 21:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 21:29 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 21:27 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 21:22 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1094.eqiad.wmnet with OS trixie * 21:21 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host restbase2039.codfw.wmnet with OS bullseye * 21:21 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin2002" * 21:21 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin2002" * 21:04 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on restbase2039.codfw.wmnet with reason: host reimage * 21:00 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on restbase2039.codfw.wmnet with reason: host reimage * 20:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1073.eqiad.wmnet with OS trixie * 20:48 mutante: deploy2003 - kill 1102 (stunnel4) ; systemctl start stunnel4 ([[phab:T418262|T418262]]) * 20:42 cjming@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] (duration: 33m 02s) * 20:42 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host restbase2039.codfw.wmnet with OS bullseye * 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1073.eqiad.wmnet with reason: host reimage * 20:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1073.eqiad.wmnet with reason: host reimage * 20:30 cjming@deploy2003: cjming: Continuing with deployment * 20:28 cjming@deploy2003: cjming: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1098.eqiad.wmnet with OS trixie * 20:13 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1073.eqiad.wmnet with OS trixie * 20:09 cjming@deploy2003: Started scap sync-world: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] * 20:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1098.eqiad.wmnet with reason: host reimage * 19:56 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1098.eqiad.wmnet with reason: host reimage * 19:55 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d4-codfw * 19:54 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d4-codfw * 19:54 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c1-codfw * 19:54 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c1-codfw * 19:52 mutante: restarting gerrit on gerrit.wikimedia.org (gerrit2003) * 19:48 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2331.codfw.wmnet * 19:48 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2331.codfw.wmnet * 19:48 mutante: restarting gerrit on gerrit-replica.wikimedia.org (gerrit1003) * 19:46 mutante: restarting gerrit on gerrit-spare.wikimedia.org (gerrit2002) * 19:43 jasmine@cumin2002: conftool action : set/pooled=yes; selector: name=wikikube-worker2331.codfw.wmnet,cluster=kubernetes,service=kubesvc * 19:43 jasmine@cumin2002: conftool action : set/weight=10; selector: name=wikikube-worker2331.codfw.wmnet,cluster=kubernetes,service=kubesvc * 19:40 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1098.eqiad.wmnet with OS trixie * 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d5-codfw * 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d5-codfw * 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c7-codfw * 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c7-codfw * 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c5-codfw * 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c5-codfw * 19:30 jasmine_: ran homer on lsw1-d8-codfw, adding wikikube-worker2331 to cluster * 19:29 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1100.eqiad.wmnet with OS trixie * 19:20 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d8-codfw * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d8-codfw * 19:19 mutante: gerrit - replacing private key for registerEmail verification * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-magru * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device cr2-magru * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d7-codfw * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d7-codfw * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d3-codfw * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d1-codfw * 19:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d1-codfw * 19:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c2-codfw * 19:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c2-codfw * 19:11 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-codfw * 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-magru * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device cr1-magru * 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d8-codfw * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d8-codfw * 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d6-codfw * 19:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1100.eqiad.wmnet with reason: host reimage * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d6-codfw * 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c6-codfw * 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c6-codfw * 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c3-codfw * 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c3-codfw * 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b4-magru * 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device asw1-b4-magru * 19:08 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b3-magru * 19:08 cmooney@cumin1003: START - Cookbook sre.network.tls for network device asw1-b3-magru * 19:05 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1100.eqiad.wmnet with reason: host reimage * 19:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1122.eqiad.wmnet with OS trixie * 19:00 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 18:59 topranks: rolling out update to BGP ACL on Nokia Switches eqiad, codfw & ulsfo [[phab:T425703|T425703]] * 18:58 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 18:57 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 18:55 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 18:53 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 18:52 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 18:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1100.eqiad.wmnet with OS trixie * 18:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1068.eqiad.wmnet with OS trixie * 18:47 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1102.eqiad.wmnet with OS trixie * 18:47 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 18:46 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1122.eqiad.wmnet with reason: host reimage * 18:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1122.eqiad.wmnet with reason: host reimage * 18:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1068.eqiad.wmnet with reason: host reimage * 18:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1122 * 18:26 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1122 * 18:25 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1122 * 18:25 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1122.eqiad.wmnet 31.48.64.10.in-addr.arpa 1.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:25 bking@cumin2003: START - Cookbook sre.dns.wipe-cache cirrussearch1122.eqiad.wmnet 31.48.64.10.in-addr.arpa 1.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:25 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:25 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1122 - bking@cumin2003" * 18:25 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1122 - bking@cumin2003" * 18:21 rzl@deploy2003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 18:21 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1068.eqiad.wmnet with reason: host reimage * 18:21 rzl@deploy2003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 18:21 rzl@deploy2003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 18:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 18:19 rzl@deploy2003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 18:19 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:18 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1122 * 18:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1122.eqiad.wmnet with OS trixie * 18:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 18:15 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 18:13 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 18:13 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 18:10 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 18:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1068.eqiad.wmnet with OS trixie * 18:01 kamila@deploy2003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 18m 29s) * 18:00 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:55 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:42 kamila@deploy2003: Started scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] * 17:42 kamila@deploy2003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 19m 50s) * 17:42 kamila@deploy2003: Rolling back deployment * 17:35 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:31 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet * 17:18 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet * 17:16 kamila@deploy1003: Unlocked for deployment [MediaWiki]: switching deployment server (duration: 22m 07s) * 17:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 17:11 kamila@dns1005: END - running authdns-update * 17:09 kamila@dns1005: START - running authdns-update * 17:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 17:04 jasmine@cumin2002: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1164.eqiad.wmnet * 17:04 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1164.eqiad.wmnet * 17:04 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1164.eqiad.wmnet * 16:56 kamila@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on releases2003.codfw.wmnet,releases1003.eqiad.wmnet with reason: Deployment server switchover * 16:54 kamila@deploy1003: Locking from deployment [MediaWiki]: switching deployment server * 16:53 kamila@deploy1003: Unlocked for deployment [MediaWiki]: switching deployment server (duration: 04m 02s) * 16:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie * 16:49 kamila@deploy1003: Locking from deployment [MediaWiki]: switching deployment server * 16:46 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1095.eqiad.wmnet with OS trixie * 16:45 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1093.eqiad.wmnet with OS trixie * 16:43 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1164.eqiad.wmnet with OS trixie * 16:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1095.eqiad.wmnet with reason: host reimage * 16:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 16:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 16:23 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1164.eqiad.wmnet with reason: host reimage * 16:18 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on cirrussearch1093.eqiad.wmnet with reason: host reimage * 16:16 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1164.eqiad.wmnet with reason: host reimage * 16:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1095.eqiad.wmnet with reason: host reimage * 16:09 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 16:09 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 16:08 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1093.eqiad.wmnet with reason: host reimage * 15:59 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Pool test * 15:59 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:59 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 15:59 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Pool test * 15:58 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Depool test * 15:58 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:58 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 15:58 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Depool test * 15:57 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1164 * 15:57 jasmine@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1164 * 15:57 jasmine@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1164 * 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1164.eqiad.wmnet 114.48.64.10.in-addr.arpa 4.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:56 jasmine@cumin2002: START - Cookbook sre.dns.wipe-cache wikikube-worker1164.eqiad.wmnet 114.48.64.10.in-addr.arpa 4.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1164 - jasmine@cumin2002" * 15:56 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1164 - jasmine@cumin2002" * 15:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1093.eqiad.wmnet with OS trixie * 15:51 jasmine@cumin2002: START - Cookbook sre.dns.netbox * 15:51 jasmine@cumin2002: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1164 * 15:50 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-worker1164.eqiad.wmnet with OS trixie * 15:50 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1164.eqiad.wmnet * 15:50 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1164.eqiad.wmnet * 15:50 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1164.eqiad.wmnet * 15:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1095.eqiad.wmnet with OS trixie * 15:42 jasmine@cumin2002: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1164.eqiad.wmnet * 15:42 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1164.eqiad.wmnet * 15:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:42 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1164.eqiad.wmnet * 15:42 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1164.eqiad.wmnet * 15:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 15:39 elukey@cumin1003: START - Cookbook sre.hosts.provision for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 15:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1007.eqiad.wmnet with OS trixie * 15:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Pool test * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1007.eqiad.wmnet with reason: host reimage * 15:15 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 15:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet * 15:15 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 15:15 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1007.eqiad.wmnet with reason: host reimage * 15:15 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:14 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Pool test * 15:14 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet * 15:14 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2228: Depool test * 15:14 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db2228: Depool test * 15:10 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 15:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:08 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 15:06 blake@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 15:06 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet * 15:06 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 15:06 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 15:05 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet * 15:05 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 15:04 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:04 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 15:04 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:03 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 15:03 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 15:03 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:03 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T430909|T430909]] * 15:03 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:03 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 15:03 swfrench-wmf: restarted eqsin, codfw confds - [[phab:T430909|T430909]] * 15:03 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test1002.eqiad.wmnet * 15:02 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet * 14:59 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:59 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:55 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1007.eqiad.wmnet with OS trixie * 14:52 swfrench-wmf: restarted ulsfo confds, confirmed now connected to codfw backends except those using wikimedia.org SRV record - [[phab:T430909|T430909]] * 14:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:49 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:41 moritzm: uninstalling dhcpcd-base from trixie hosts which still have it installed [[phab:T414341|T414341]] * 14:40 sukhe: sudo cumin -b1 -s120 "P<nowiki>{</nowiki>lvs2011*<nowiki>}</nowiki> or P<nowiki>{</nowiki>lvs2012*<nowiki>}</nowiki>" "systemctl restart pybal.service" * 14:39 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:39 mvernon@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host thanos-be1007.eqiad.wmnet with OS trixie * 14:37 sukhe: restart pybal on lvs2013 to revert back to conf2004 * 14:35 sukhe: restart pybal on lvs2014 to revert back to conf2004 * 14:34 swfrench-wmf: switched codfw, eqsin, ulsfo etcd client SRV records back to codfw - [[phab:T430909|T430909]] * 14:32 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1002.eqiad.wmnet * 14:32 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet * 14:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1007.eqiad.wmnet with OS trixie * 14:31 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:31 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:31 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:30 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:30 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Pool test * 14:30 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:29 swfrench@dns1004: END - running authdns-update * 14:29 moritzm: installing jackson-core security updates * 14:27 swfrench@dns1004: START - running authdns-update * 14:22 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:22 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:22 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1119.eqiad.wmnet with OS trixie * 14:22 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:21 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:20 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:20 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 14:20 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:19 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 14:19 moritzm: installing librabbitmq security updates * 14:19 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1002.eqiad.wmnet * 14:18 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet * 14:16 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:15 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:15 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:14 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Pool test * 14:14 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox) * 14:14 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet * 14:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 14:08 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 14:07 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet * 14:05 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1118.eqiad.wmnet with OS trixie * 14:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test1001.eqiad.wmnet * 14:00 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-worker@eqiad * 14:00 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 13:59 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 13:58 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 13:57 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1119.eqiad.wmnet with reason: host reimage * 13:54 moritzm: installing libcap2 security updates * 13:53 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1119.eqiad.wmnet with reason: host reimage * 13:52 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet * 13:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) * 13:52 fceratto@cumin1003: START - Cookbook sre.mysql.depool * 13:50 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-worker@eqiad * 13:50 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1051.eqiad.wmnet * 13:50 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1051.eqiad.wmnet * 13:50 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1051.eqiad.wmnet * 13:49 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 13:45 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1081.eqiad.wmnet with OS trixie * 13:41 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1119 * 13:41 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1119 * 13:40 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1119 * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1119.eqiad.wmnet 97.32.64.10.in-addr.arpa 7.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1119.eqiad.wmnet 97.32.64.10.in-addr.arpa 7.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1119 - atsuko@cumin1003" * 13:40 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1119 - atsuko@cumin1003" * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1118.eqiad.wmnet with reason: host reimage * 13:39 moritzm: installing krb5 security updates * 13:37 Lucas_WMDE: UTC afternoon backport+config window done * 13:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1006.eqiad.wmnet with OS trixie * 13:36 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1118.eqiad.wmnet with reason: host reimage * 13:36 atsuko@cumin1003: START - Cookbook sre.dns.netbox * 13:35 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] (duration: 07m 46s) * 13:34 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1119 * 13:34 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1119.eqiad.wmnet with OS trixie * 13:30 sbisson@deploy1003: sbisson: Continuing with deployment * 13:30 moritzm: installing openssh security updates * 13:30 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling restart_daemons on A:wikidough * 13:29 sbisson@deploy1003: sbisson: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:27 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] * 13:26 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1051.eqiad.wmnet with OS trixie * 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1081.eqiad.wmnet with reason: host reimage * 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1118 * 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1118 * 13:22 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] (duration: 12m 12s) * 13:21 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1081.eqiad.wmnet with reason: host reimage * 13:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1006.eqiad.wmnet with reason: host reimage * 13:18 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1118 * 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1118.eqiad.wmnet 90.32.64.10.in-addr.arpa 0.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:18 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1118.eqiad.wmnet 90.32.64.10.in-addr.arpa 0.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1118 - atsuko@cumin1003" * 13:18 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1118 - atsuko@cumin1003" * 13:17 stran@deploy1003: stran: Continuing with deployment * 13:16 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough * 13:15 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:13 atsuko@cumin1003: START - Cookbook sre.dns.netbox * 13:12 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1118 * 13:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1006.eqiad.wmnet with reason: host reimage * 13:12 stran@deploy1003: stran: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:12 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1118.eqiad.wmnet with OS trixie * 13:10 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] * 13:05 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-worker@codfw * 13:05 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 13:05 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1081.eqiad.wmnet with OS trixie * 13:05 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1051.eqiad.wmnet with reason: host reimage * 13:04 moritzm: installing jq security updates * 13:04 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 13:01 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1051.eqiad.wmnet with reason: host reimage * 12:58 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-worker@codfw * 12:52 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 12:50 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host thanos-be1006.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1051 * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1051 * 12:43 moritzm: installing Python 3.11 security updates * 12:43 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1051 * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1051.eqiad.wmnet 46.32.64.10.in-addr.arpa 6.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:43 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1051.eqiad.wmnet 46.32.64.10.in-addr.arpa 6.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1051 - blake@cumin1003" * 12:43 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1051 - blake@cumin1003" * 12:38 blake@cumin1003: START - Cookbook sre.dns.netbox * 12:38 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1051 * 12:38 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1051.eqiad.wmnet with OS trixie * 12:37 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1051.eqiad.wmnet * 12:36 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1051.eqiad.wmnet * 12:36 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1051.eqiad.wmnet * 12:34 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1006.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 12:34 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1006.eqiad.wmnet with OS trixie * 12:27 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 12:27 mvernon@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host thanos-be1006.eqiad.wmnet with OS trixie * 12:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:02 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 12:01 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1006.eqiad.wmnet with OS trixie * 11:43 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 11:38 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1076.eqiad.wmnet with OS trixie * 11:26 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1075.eqiad.wmnet with OS trixie * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2047.codfw.wmnet * 11:19 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2047.codfw.wmnet * 11:19 moritzm: temporarily remove ganeti2031 from codfw cluster [[phab:T430910|T430910]] * 11:08 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1076.eqiad.wmnet with reason: host reimage * 11:08 moritzm: installing Linux 6.1.176 on Bookworm servers * 11:03 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1076.eqiad.wmnet with reason: host reimage * 11:00 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1075.eqiad.wmnet with reason: host reimage * 10:56 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1075.eqiad.wmnet with reason: host reimage * 10:47 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1076.eqiad.wmnet with OS trixie * 10:46 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1074.eqiad.wmnet with OS trixie * 10:45 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1005.eqiad.wmnet with OS trixie * 10:40 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1075.eqiad.wmnet with OS trixie * 10:32 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2031.codfw.wmnet * 10:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1005.eqiad.wmnet with reason: host reimage * 10:25 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1005.eqiad.wmnet with reason: host reimage * 10:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1074.eqiad.wmnet with reason: host reimage * 10:17 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1074.eqiad.wmnet with reason: host reimage * 10:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1005.eqiad.wmnet with OS trixie * 10:12 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet * 10:04 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 10:01 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet * 10:01 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1074.eqiad.wmnet with OS trixie * 10:01 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 09:43 cgoubert@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/aux-k8s-services/redioscope: apply * 09:43 cgoubert@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/aux-k8s-services/redioscope: apply * 09:43 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply * 09:35 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply * 09:34 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 09:34 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 09:33 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 41 days, 15:00:00 on db2252.codfw.wmnet with reason: Test * 09:32 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 09:32 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 09:31 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: codfw rack B3 pool after maintenance * 09:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 09:07 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 09:07 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 09:02 ladsgroup@cumin1003: END (PASS) - Cookbook sre.mysql.sanitarium_restart (exit_code=0) * 08:57 topranks: merge patch to shift eqiad <-> esams traffic onto new 40G circuit * 08:54 hashar@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1004.eqiad.wmnet with OS trixie * 08:50 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 08:50 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitarium_restart (exit_code=99) * 08:50 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 08:45 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool es2051: codfw rack B3 pool after maintenance * 08:44 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:44 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:43 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2007.codfw.wmnet * 08:43 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2007.codfw.wmnet * 08:42 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2031.codfw.wmnet * 08:41 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2031.codfw.wmnet * 08:40 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2031.codfw.wmnet * 08:38 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:38 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:35 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.sanitize-wiki (exit_code=97) Managing sanitization for wikis minwikiquote in section s3 * 08:33 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis minwikiquote in section s3 * 08:32 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Checking sanitization for wikis minwikiquote in section s5 * 08:30 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Checking sanitization for wikis minwikiquote in section s5 * 08:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Managing sanitization for wikis minwikiquote in section s5 * 08:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1004.eqiad.wmnet with reason: host reimage * 08:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1004.eqiad.wmnet with reason: host reimage * 08:23 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:22 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis minwikiquote in section s5 * 08:19 XioNoX: lsw1-b3-codfw> request system reboot - [[phab:T430909|T430909]] * 08:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Checking sanitization for wikis minwikiquote in section s5 * 08:17 hashar@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Checking sanitization for wikis minwikiquote in section s5 * 08:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for codfw rack B3 * 08:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2007.codfw.wmnet * 08:15 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lsw1-b3-codfw,lsw1-b3-codfw IPv6,lsw1-b3-codfw.mgmt with reason: Switch maintenance * 08:15 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2007.codfw.wmnet * 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:07 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:06 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: codfw rack B3 depool for maintenance * 08:05 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool es2051: codfw rack B3 depool for maintenance * 08:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1004.eqiad.wmnet with OS trixie * 08:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1005.eqiad.wmnet with OS trixie * 08:03 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 21 hosts with reason: codfw rack B3 depool for maintenance * 07:56 ayounsi@cumin1003: START - Cookbook sre.network.depool-rack with action 'depool' for codfw rack B3 * 07:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1005.eqiad.wmnet with reason: host reimage * 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1005.eqiad.wmnet with reason: host reimage * 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1005.eqiad.wmnet with OS bookworm * 07:29 moritzm: installing gnutls28 security updates * 07:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1005.eqiad.wmnet with OS trixie * 07:13 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1125.eqiad.wmnet with OS trixie * 07:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1005.eqiad.wmnet with reason: host reimage * 07:07 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aux-k8s-etcd1005.eqiad.wmnet with reason: host reimage * 06:56 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1005.eqiad.wmnet with OS bookworm * 06:54 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1125.eqiad.wmnet with reason: host reimage * 06:52 elukey: upgrade all trixie hosts to pywmflib 3.1 - [[phab:T430552|T430552]] * 06:50 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1125.eqiad.wmnet with reason: host reimage * 06:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 06:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 06:38 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1125.eqiad.wmnet with OS trixie * 05:42 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1107.eqiad.wmnet with OS trixie * 05:35 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1124.eqiad.wmnet with OS trixie * 05:31 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1101.eqiad.wmnet with OS trixie * 05:21 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1107.eqiad.wmnet with reason: host reimage * 05:17 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1124.eqiad.wmnet with reason: host reimage * 05:13 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1107.eqiad.wmnet with reason: host reimage * 05:13 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1101.eqiad.wmnet with reason: host reimage * 05:11 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1124.eqiad.wmnet with reason: host reimage * 05:10 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1101.eqiad.wmnet with reason: host reimage * 04:58 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1124.eqiad.wmnet with OS trixie * 04:56 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1107.eqiad.wmnet with OS trixie * 04:55 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1101.eqiad.wmnet with OS trixie * 02:27 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] (duration: 08m 14s) * 02:22 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 02:21 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 02:19 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] * 01:59 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] (duration: 09m 46s) * 01:55 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 01:51 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 01:49 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] * 01:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1099.eqiad.wmnet with OS trixie * 00:57 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1110.eqiad.wmnet with OS trixie * 00:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1099.eqiad.wmnet with reason: host reimage * 00:41 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1099.eqiad.wmnet with reason: host reimage * 00:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1110.eqiad.wmnet with reason: host reimage * 00:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1110.eqiad.wmnet with reason: host reimage * 00:26 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1099.eqiad.wmnet with OS trixie * 00:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1110.eqiad.wmnet with OS trixie == 2026-07-07 == * 22:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1097.eqiad.wmnet with OS trixie * 22:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1097.eqiad.wmnet with reason: host reimage * 22:24 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1097.eqiad.wmnet with reason: host reimage * 22:14 hashar: Restarting Gerrit on gerrit2002 and gerrit1003 (replicas) * 22:09 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1097.eqiad.wmnet with OS trixie * 22:07 hashar: Restarting Gerrit on gerrit2003 * 21:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 21:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 21:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 21:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 21:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1108.eqiad.wmnet with OS trixie * 20:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1091.eqiad.wmnet with OS trixie * 20:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1090.eqiad.wmnet with OS trixie * 20:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1108.eqiad.wmnet with reason: host reimage * 20:36 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1091.eqiad.wmnet with reason: host reimage * 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1090.eqiad.wmnet with reason: host reimage * 20:33 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1091.eqiad.wmnet with reason: host reimage * 20:30 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1108.eqiad.wmnet with reason: host reimage * 20:30 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1006.eqiad.wmnet * 20:30 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1090.eqiad.wmnet with reason: host reimage * 20:30 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1006.eqiad.wmnet * 20:27 jasmine_: "homer lsw1-c2-eqiad* commit "Added new stacked control plane wikikube-ctrl1006"" * 20:22 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] (duration: 07m 29s) * 20:20 jasmine_: "homer "cr*eqiad*" commit "Added new stacked control plane wikikube-ctrl1006"" * 20:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1091.eqiad.wmnet with OS trixie * 20:17 arlolra@deploy1003: arlolra: Continuing with deployment * 20:16 arlolra@deploy1003: arlolra: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:16 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1090.eqiad.wmnet with OS trixie * 20:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1108.eqiad.wmnet with OS trixie * 20:14 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] * 20:09 cwhite: remove 2026-04 swift log archives from centrallog2002 to free some space * 20:01 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=93) for host cirrussearch1108.eqiad.wmnet with OS trixie * 19:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1108.eqiad.wmnet with OS trixie * 19:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1090.eqiad.wmnet with OS trixie * 19:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1109.eqiad.wmnet with OS trixie * 19:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1092.eqiad.wmnet with OS trixie * 19:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1123.eqiad.wmnet with OS trixie * 19:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1109.eqiad.wmnet with reason: host reimage * 19:19 cdobbins@cumin2002: conftool action : set/pooled=yes; selector: name=dns7002.* * 19:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1092.eqiad.wmnet with reason: host reimage * 19:17 jasmine@dns1004: END - running authdns-update * 19:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1109.eqiad.wmnet with reason: host reimage * 19:15 jasmine@dns1004: START - running authdns-update * 19:15 cdobbins@dns1004: END - running authdns-update * 19:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1123.eqiad.wmnet with reason: host reimage * 19:13 cdobbins@dns1004: START - running authdns-update * 19:12 cdobbins@cumin2002: conftool action : set/pooled=yes; selector: name=dns7002.*,service=authdns-update * 19:11 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1092.eqiad.wmnet with reason: host reimage * 19:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1123.eqiad.wmnet with reason: host reimage * 18:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1123.eqiad.wmnet with OS trixie * 18:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1109.eqiad.wmnet with OS trixie * 18:56 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1092.eqiad.wmnet with OS trixie * 18:52 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 18:49 swfrench@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 18:40 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 18:38 swfrench@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 18:11 swfrench-wmf: restarted eqsin, codfw confds - [[phab:T430909|T430909]] * 18:01 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T430909|T430909]] * 17:59 swfrench-wmf: restarted ulsfo confds, confirmed now connected to eqiad backends - [[phab:T430909|T430909]] * 17:52 sukhe: restart pybal on lvs2011 to switch from conf2004 to conf1008: [[phab:T430909|T430909]] * 17:51 sukhe: restart pybal on lvs2012 to switch from conf2004 to conf1008 [puppet re-enabled there]: [[phab:T430909|T430909]] * 17:46 sukhe: restart pybal on lvs2013 to switch from conf2004 to conf1008: [[phab:T430909|T430909]] * 17:44 swfrench-wmf: switched codfw, eqsin, ulsfo etcd client SRV records to eqiad - [[phab:T430909|T430909]] * 17:43 swfrench@dns1004: END - running authdns-update * 17:40 swfrench@dns1004: START - running authdns-update * 17:40 sukhe: restart pybal on lvs2014 to switch from conf2004 to conf1008: [[phab:T430909|T430909]] * 17:21 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1003.eqiad.wmnet * 17:15 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1003.eqiad.wmnet * 17:14 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1002.eqiad.wmnet * 17:06 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1002.eqiad.wmnet * 17:06 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-low-traffic-codfw' 'systemctl restart pybal.service' # lvs2013, [[phab:T416623|T416623]] * 17:04 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1001.eqiad.wmnet * 17:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1111.eqiad.wmnet with OS trixie * 17:00 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal.service' # lvs2014, [[phab:T416623|T416623]] * 16:58 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1001.eqiad.wmnet * 16:58 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS bookworm * 16:55 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-low-traffic-eqiad' 'systemctl restart pybal.service' # lvs1019, [[phab:T416623|T416623]] * 16:53 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-secondary-eqiad' 'systemctl restart pybal.service' # lvs1020, [[phab:T416623|T416623]] * 16:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1111.eqiad.wmnet with reason: host reimage * 16:40 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1111.eqiad.wmnet with reason: host reimage * 16:38 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.peering (exit_code=99) with action 'configure' for AS: 47794 * 16:35 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 47794 * 16:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1111.eqiad.wmnet with OS trixie * 16:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1006.eqiad.wmnet with OS trixie * 16:06 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 16:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1006.eqiad.wmnet with reason: host reimage * 15:58 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1121.eqiad.wmnet with OS trixie * 15:58 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 15:56 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1006.eqiad.wmnet with reason: host reimage * 15:54 mutante: jenkins down in planned maintenance window * 15:42 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1037.eqiad.wmnet * 15:42 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1037.eqiad.wmnet * 15:42 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1037.eqiad.wmnet * 15:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1006.eqiad.wmnet with OS trixie * 15:34 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1121.eqiad.wmnet with reason: host reimage * 15:33 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS bookworm * 15:33 cdobbins@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host dns7002.wikimedia.org with OS trixie * 15:30 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1121.eqiad.wmnet with reason: host reimage * 15:29 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host clouddumps1001.wikimedia.org * 15:20 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1001.wikimedia.org * 15:18 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1121 * 15:18 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1121 * 15:18 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host clouddumps1002.wikimedia.org * 15:17 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1121 * 15:17 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:17 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply * 15:16 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply * 15:16 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:16 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1121 - atsuko@cumin1003" * 15:16 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1121 - atsuko@cumin1003" * 15:14 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1037.eqiad.wmnet with OS trixie * 15:11 atsuko@cumin1003: START - Cookbook sre.dns.netbox * 15:09 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org * 15:09 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1121 * 15:09 andrew@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host clouddumps1002.wikimedia.org * 15:09 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org * 15:09 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1121.eqiad.wmnet with OS trixie * 15:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1007.eqiad.wmnet with OS trixie * 15:08 andrew@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host clouddumps1002.wikimedia.org * 15:08 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org * 15:05 brennen@deploy1003: Finished deploy [phabricator/deployment@7e02037]: deploy phab1004 for [[phab:T431440|T431440]] (duration: 00m 47s) * 15:04 brennen@deploy1003: Started deploy [phabricator/deployment@7e02037]: deploy phab1004 for [[phab:T431440|T431440]] * 15:03 brennen@deploy1003: Finished deploy [phabricator/deployment@7e02037]: deploy phab2003 for [[phab:T431440|T431440]] (duration: 00m 51s) * 15:03 brennen@deploy1003: Started deploy [phabricator/deployment@7e02037]: deploy phab2003 for [[phab:T431440|T431440]] * 15:00 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71] (thin): Regular analytics weekly train THIN [analytics/refinery@7d8dc71f] (duration: 02m 10s) * 14:58 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71] (thin): Regular analytics weekly train THIN [analytics/refinery@7d8dc71f] * 14:58 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71]: Regular analytics weekly train [analytics/refinery@7d8dc71f] (duration: 04m 14s) * 14:54 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1037.eqiad.wmnet with reason: host reimage * 14:53 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71]: Regular analytics weekly train [analytics/refinery@7d8dc71f] * 14:53 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@7d8dc71f] (duration: 02m 00s) * 14:51 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@7d8dc71f] * 14:51 arnaudb@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on phab2003.codfw.wmnet,phab[1004-1006].eqiad.wmnet with reason: maintenance * 14:51 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1037.eqiad.wmnet with reason: host reimage * 14:50 JavierMonton: Deploying Refinery at {{Gerrit|7d8dc71f}} for change {{Gerrit|1308087}} / [[phab:T431318|T431318]] - update filerevision table sqoop and table * 14:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1007.eqiad.wmnet with reason: host reimage * 14:42 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1007.eqiad.wmnet with reason: host reimage * 14:40 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1083.eqiad.wmnet with OS trixie * 14:37 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:36 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:35 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-master@eqiad * 14:35 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 14:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 14:34 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1037 * 14:34 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1037 * 14:34 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:cleanMentorList.php --wiki=frwiki # [[phab:T427386|T427386]] * 14:34 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 14:34 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308112{{!}}Revert^2 "[Growth] frwiki: Deploy automated mentor list cleaner" (T427386)]] (duration: 06m 47s) * 14:34 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 14:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:33 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:32 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1037 * 14:31 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:31 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:29 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:29 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-master@eqiad * 14:29 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:29 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:29 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:28 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:27 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:27 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1308112{{!}}Revert^2 "[Growth] frwiki: Deploy automated mentor list cleaner" (T427386)]] * 14:26 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1007.eqiad.wmnet with OS trixie * 14:26 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:26 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 14:26 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:25 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:25 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:25 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:cleanMentorList.php --wiki=frwiki # [[phab:T427386|T427386]] * 14:24 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:24 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1037 - blake@cumin1003" * 14:24 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1037 - blake@cumin1003" * 14:20 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1083.eqiad.wmnet with reason: host reimage * 14:19 blake@cumin1003: START - Cookbook sre.dns.netbox * 14:19 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-master@codfw * 14:19 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 14:19 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1037 * 14:18 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1037.eqiad.wmnet with OS trixie * 14:18 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1037.eqiad.wmnet * 14:18 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 14:18 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1037.eqiad.wmnet * 14:18 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1037.eqiad.wmnet * 14:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2007.codfw.wmnet with OS trixie * 14:16 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1083.eqiad.wmnet with reason: host reimage * 14:15 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1036.eqiad.wmnet * 14:15 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1036.eqiad.wmnet * 14:14 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1036.eqiad.wmnet * 14:12 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-master@codfw * 14:11 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1120.eqiad.wmnet with OS trixie * 14:05 moritzm: installing distro-info-data updates from trixie/bookworm point releases * 14:04 fabfur: disable puppet on A:cp-text to selectively apply https://gerrit.wikimedia.org/r/c/operations/puppet/+/1308040 * 14:03 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] (duration: 27m 48s) * 14:00 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 14:00 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1083.eqiad.wmnet with OS trixie * 13:58 urbanecm@deploy1003: urbanecm: Continuing with deployment * 13:58 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:58 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1004.eqiad.wmnet with OS bookworm * 13:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2007.codfw.wmnet with reason: host reimage * 13:57 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 13:53 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1120.eqiad.wmnet with reason: host reimage * 13:50 moritzm: installing Linux 5.10.259 on Bullseye hosts * 13:47 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply * 13:47 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply * 13:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2007.codfw.wmnet with reason: host reimage * 13:46 cgoubert@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/aux-k8s-services/redioscope: apply * 13:46 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1120.eqiad.wmnet with reason: host reimage * 13:46 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:46 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:45 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:44 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:44 cgoubert@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/aux-k8s-services/redioscope: apply * 13:40 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 13:39 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:38 moritzm: installing e2fsprogs updates from Trixie point release * 13:35 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] * 13:33 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1120.eqiad.wmnet with OS trixie * 13:33 topranks: reset cr3-eqsin configuration so traffic uses it again after upgrade * 13:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1088.eqiad.wmnet with OS trixie * 13:32 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie * 13:32 cdobbins@cumin1003: conftool action : set/pooled=no; selector: name=dns7002.* * 13:29 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2007.codfw.wmnet with OS trixie * 13:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1004.eqiad.wmnet with reason: host reimage * 13:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2006.codfw.wmnet with OS trixie * 13:18 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1036.eqiad.wmnet with OS trixie * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aux-k8s-etcd1004.eqiad.wmnet with reason: host reimage * 13:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 13:16 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 13:15 jayme: Istio is being upgraded from 1.24.2 to 1.29.4 on wikikube staging eqiad and codfw - [[phab:T427401|T427401]] * 13:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1087.eqiad.wmnet with OS trixie * 13:14 topranks: reboot cr3-eqsin to install new JunOS and set PIC 0/0/0 to 100G * 13:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1088.eqiad.wmnet with reason: host reimage * 13:13 jmm@dns1004: END - running authdns-update * 13:12 jmm@dns1004: START - running authdns-update * 13:09 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1088.eqiad.wmnet with reason: host reimage * 13:07 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1082.eqiad.wmnet with OS trixie * 13:07 jmm@dns1004: END - running authdns-update * 13:06 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1004.eqiad.wmnet with OS bookworm * 13:05 jmm@dns1004: START - running authdns-update * 13:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2006.codfw.wmnet with reason: host reimage * 12:58 topranks: load updated JunOS on cr3-eqsin [[phab:T429386|T429386]] * 12:58 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1036.eqiad.wmnet with reason: host reimage * 12:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2001.codfw.wmnet * 12:57 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2006.codfw.wmnet with reason: host reimage * 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr1-codfw,cr[2-3]-eqsin,cr3-eqsin IPv6,cr3-eqsin.mgmt with reason: upgrade JunOS cr3-eqsin * 12:56 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lvs[5004-5006].eqsin.wmnet with reason: upgrade JunOS cr3-eqsin * 12:55 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 12:55 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 12:53 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1087.eqiad.wmnet with reason: host reimage * 12:53 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 12:52 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1088.eqiad.wmnet with OS trixie * 12:52 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1002.eqiad.wmnet * 12:52 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:51 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2001.codfw.wmnet * 12:49 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1036.eqiad.wmnet with reason: host reimage * 12:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 12:48 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1087.eqiad.wmnet with reason: host reimage * 12:44 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1082.eqiad.wmnet with reason: host reimage * 12:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1002.eqiad.wmnet * 12:42 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 12:42 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 12:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 12:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 12:39 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:39 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: move dumps-nfs IP to the shared one - filippo@cumin1003" * 12:39 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: move dumps-nfs IP to the shared one - filippo@cumin1003" * 12:39 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2006.codfw.wmnet with OS trixie * 12:38 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1082.eqiad.wmnet with reason: host reimage * 12:36 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:33 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:32 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1036 * 12:32 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1036 * 12:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2005.codfw.wmnet with OS trixie * 12:32 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1087.eqiad.wmnet with OS trixie * 12:30 jmm@dns1004: END - running authdns-update * 12:29 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1036 * 12:29 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1036.eqiad.wmnet 21.32.64.10.in-addr.arpa 1.2.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:29 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1036.eqiad.wmnet 21.32.64.10.in-addr.arpa 1.2.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:29 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:29 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1036 - blake@cumin1003" * 12:29 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1036 - blake@cumin1003" * 12:28 jmm@dns1004: START - running authdns-update * 12:26 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:26 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:23 blake@cumin1003: START - Cookbook sre.dns.netbox * 12:23 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1036 * 12:23 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1036.eqiad.wmnet with OS trixie * 12:22 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1036.eqiad.wmnet * 12:22 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1082.eqiad.wmnet with OS trixie * 12:22 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1036.eqiad.wmnet * 12:22 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1036.eqiad.wmnet * 12:21 marostegui: Restart mariadb@s7 on db1155 to pick up new filters - [[phab:T431124|T431124]] * 12:21 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 21 hosts with reason: restarting for replication filter * 12:20 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:19 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:14 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2005.codfw.wmnet with reason: host reimage * 12:14 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:08 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:07 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:07 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2005.codfw.wmnet with reason: host reimage * 12:06 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:06 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:06 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:05 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:05 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:04 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-master-eqiad * 12:04 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl1002.eqiad.wmnet * 12:04 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl1002.eqiad.wmnet * 12:04 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:04 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:03 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:03 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:03 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:03 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 11:59 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl1002.eqiad.wmnet * 11:59 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl1002.eqiad.wmnet * 11:59 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl1001.eqiad.wmnet * 11:59 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl1001.eqiad.wmnet * 11:56 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl1001.eqiad.wmnet * 11:56 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl1001.eqiad.wmnet * 11:56 klausman@cumin2002: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-master-eqiad * 11:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2005.codfw.wmnet with OS trixie * 11:49 blake@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on wikikube-worker1160.eqiad.wmnet with reason: Verifying matchers for silence * 11:42 topranks: cr3-eqsin, begin traffic drain to reset PIC and upgrade JunOS [[phab:T429386|T429386]] * 11:41 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs[5004-5006].eqsin.wmnet with reason: upgrade JunOS cr3-eqsin * 11:39 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr1-codfw,cr[2-3]-eqsin,cr3-eqsin IPv6,cr3-eqsin.mgmt with reason: upgrade JunOS cr3-eqsin * 11:36 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=thanos-fe2004.codfw.wmnet * 11:35 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1086.eqiad.wmnet with OS trixie * 11:35 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=thanos-fe2004.codfw.wmnet * 11:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1085.eqiad.wmnet with OS trixie * 11:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1086.eqiad.wmnet with reason: host reimage * 11:10 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1085.eqiad.wmnet with reason: host reimage * 11:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2004.codfw.wmnet with OS trixie * 11:03 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1086.eqiad.wmnet with reason: host reimage * 11:02 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1085.eqiad.wmnet with reason: host reimage * 10:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2004.codfw.wmnet with reason: host reimage * 10:48 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:46 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1086.eqiad.wmnet with OS trixie * 10:46 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1085.eqiad.wmnet with OS trixie * 10:44 cgoubert@deploy1003: Finished deploy [restbase/deploy@2fc37d4]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] (duration: 16m 44s) * 10:43 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2004.codfw.wmnet with reason: host reimage * 10:35 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:27 cgoubert@deploy1003: Started deploy [restbase/deploy@2fc37d4]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] * 10:27 cgoubert@deploy1003: Finished deploy [restbase/deploy@8a25036]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] (duration: 00m 45s) * 10:26 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1117.eqiad.wmnet with OS trixie * 10:26 cgoubert@deploy1003: Started deploy [restbase/deploy@8a25036]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] * 10:26 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host thanos-fe2004 * 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host thanos-fe2004 * 10:22 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1116.eqiad.wmnet with OS trixie * 10:21 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host thanos-fe2004 * 10:21 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) thanos-fe2004.codfw.wmnet 157.32.192.10.in-addr.arpa 7.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:20 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache thanos-fe2004.codfw.wmnet 157.32.192.10.in-addr.arpa 7.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:20 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:20 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host thanos-fe2004 - mvernon@cumin2003" * 10:20 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host thanos-fe2004 - mvernon@cumin2003" * 10:15 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2252: Repooling after reboot * 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:15 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 10:15 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2252: Repooling after reboot * 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1153.eqiad.wmnet * 10:14 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1153.eqiad.wmnet * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 10:14 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 10:12 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 10:12 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host thanos-fe2004 * 10:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2004.codfw.wmnet with OS trixie * 10:07 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1117.eqiad.wmnet with reason: host reimage * 10:03 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1116.eqiad.wmnet with reason: host reimage * 09:58 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:58 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1117.eqiad.wmnet with reason: host reimage * 09:57 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1116.eqiad.wmnet with reason: host reimage * 09:49 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 41 days, 15:00:00 on db2252.codfw.wmnet with reason: Security updates * 09:45 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1117.eqiad.wmnet with OS trixie * 09:45 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1116.eqiad.wmnet with OS trixie * 09:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1153: Security updates * 09:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:28 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:28 root@cumin1003: START - Cookbook sre.mysql.depool depool db1153: Security updates * 09:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1016: Security updates * 09:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:21 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:21 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1016: Security updates * 09:14 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:14 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 08:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1016: Security updates * 08:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:56 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:56 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1016: Security updates * 08:50 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:50 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:45 filippo@dns1006: END - running authdns-update * 08:43 filippo@dns1006: START - running authdns-update * 08:42 godog: switch dumps-nfs address to be shared with rsync/http - [[phab:T411248|T411248]] * 08:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1016: Security updates * 08:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:40 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:40 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1016: Security updates * 08:29 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host cirrussearch1111.eqiad.wmnet * 08:29 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:27 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:27 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:25 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1015: Security updates * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:09 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:09 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1015: Security updates * 07:42 Msz2001: Deployed private patch for Suggested Ivestigations * 07:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1015: Security updates * 07:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:41 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:41 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1015: Security updates * 07:40 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 07:11 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fingerprint warnings - oblivian@cumin1003" * 07:11 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fingerprint warnings - oblivian@cumin1003 * 07:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1024: Security updates * 07:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:11 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:11 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1024: Security updates * 07:10 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fingerprint warnings - oblivian@cumin1003 * 07:10 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fingerprint warnings - oblivian@cumin1003" * 06:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host cirrussearch1111.eqiad.wmnet * 06:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 06:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1024: Security updates * 06:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 06:48 root@cumin1003: START - Cookbook sre.mysql.parsercache * 06:48 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1024: Security updates * 06:42 moritzm: install nginx security updates * 06:31 root@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool pc1024: Security updates * 06:21 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1024: Security updates * 06:19 moritzm: installing php8.2 security updates * 06:15 moritzm: installing php8.4 security updates * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.7 (duration: 02m 38s) * 03:40 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] (duration: 37m 04s) * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 51s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-06 == * 23:30 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] (duration: 09m 39s) * 23:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1078.eqiad.wmnet with OS trixie * 23:26 jdlrobson@deploy1003: jdlrobson, bwang: Continuing with deployment * 23:22 jdlrobson@deploy1003: jdlrobson, bwang: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug) * 23:21 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] * 23:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1078.eqiad.wmnet with reason: host reimage * 23:06 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1078.eqiad.wmnet with reason: host reimage * 22:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1078.eqiad.wmnet with OS trixie * 22:29 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on cirrussearch1114.eqiad.wmnet with reason: reimage on hold until restore completes * 22:22 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on cirrussearch[1079,1115].eqiad.wmnet with reason: reimage on hold until restore completes * 21:18 maryum: Deployed security fix for [[phab:T428006|T428006]] * 20:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1079.eqiad.wmnet with OS trixie * 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1077.eqiad.wmnet with OS trixie * 20:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1115.eqiad.wmnet with OS trixie * 20:25 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1079.eqiad.wmnet with reason: host reimage * 20:21 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1079.eqiad.wmnet with reason: host reimage * 20:15 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] (duration: 08m 14s) * 20:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1077.eqiad.wmnet with reason: host reimage * 20:10 krinkle@deploy1003: krinkle, pushpaktiwari: Continuing with deployment * 20:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1115.eqiad.wmnet with reason: host reimage * 20:08 krinkle@deploy1003: krinkle, pushpaktiwari: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1077.eqiad.wmnet with reason: host reimage * 20:06 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] * 20:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1079.eqiad.wmnet with OS trixie * 20:04 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1115.eqiad.wmnet with reason: host reimage * 19:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1077.eqiad.wmnet with OS trixie * 19:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1115.eqiad.wmnet with OS trixie * 19:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 19:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 18:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1114.eqiad.wmnet with OS trixie * 18:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1114.eqiad.wmnet with reason: host reimage * 18:35 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1114.eqiad.wmnet with reason: host reimage * 18:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1112.eqiad.wmnet with OS trixie * 18:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1114.eqiad.wmnet with OS trixie * 18:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1072.eqiad.wmnet with OS trixie * 18:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1112.eqiad.wmnet with reason: host reimage * 18:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1112.eqiad.wmnet with reason: host reimage * 17:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1072.eqiad.wmnet with reason: host reimage * 17:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1112.eqiad.wmnet with OS trixie * 17:55 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1072.eqiad.wmnet with reason: host reimage * 17:39 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1072.eqiad.wmnet with OS trixie * 17:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1071.eqiad.wmnet with OS trixie * 17:18 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1070.eqiad.wmnet with OS trixie * 17:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1084.eqiad.wmnet with OS trixie * 16:54 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1071.eqiad.wmnet with reason: host reimage * 16:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1084.eqiad.wmnet with reason: host reimage * 16:51 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1070.eqiad.wmnet with reason: host reimage * 16:49 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1084.eqiad.wmnet with reason: host reimage * 16:38 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1071.eqiad.wmnet with OS trixie * 16:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1096.eqiad.wmnet with OS trixie * 16:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1070.eqiad.wmnet with OS trixie * 16:33 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1084.eqiad.wmnet with OS trixie * 16:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1089.eqiad.wmnet with OS trixie * 16:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1103.eqiad.wmnet with OS trixie * 16:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1096.eqiad.wmnet with reason: host reimage * 16:14 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1096.eqiad.wmnet with reason: host reimage * 16:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1089.eqiad.wmnet with reason: host reimage * 16:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1103.eqiad.wmnet with reason: host reimage * 16:02 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1003.eqiad.wmnet with OS bookworm * 16:01 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1089.eqiad.wmnet with reason: host reimage * 16:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1103.eqiad.wmnet with reason: host reimage * 15:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1096.eqiad.wmnet with OS trixie * 15:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1080.eqiad.wmnet with OS trixie * 15:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1089.eqiad.wmnet with OS trixie * 15:45 dancy@deploy1003: Installation of scap version "4.272.0" completed for 158 hosts * 15:43 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1103.eqiad.wmnet with OS trixie * 15:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1113.eqiad.wmnet with OS trixie * 15:41 dancy@deploy1003: Installing scap version "4.272.0" for 158 host(s) * 15:40 klausman@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 15:39 klausman@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 15:38 klausman@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 15:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1069.eqiad.wmnet with OS trixie * 15:37 klausman@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 15:36 klausman@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 15:34 klausman@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 15:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1080.eqiad.wmnet with reason: host reimage * 15:27 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1080.eqiad.wmnet with reason: host reimage * 15:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1113.eqiad.wmnet with reason: host reimage * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1069.eqiad.wmnet with reason: host reimage * 15:18 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1113.eqiad.wmnet with reason: host reimage * 15:16 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1069.eqiad.wmnet with reason: host reimage * 15:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:11 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1080.eqiad.wmnet with OS trixie * 15:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1113.eqiad.wmnet with OS trixie * 15:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1003.eqiad.wmnet with reason: host reimage * 14:47 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1003.eqiad.wmnet with OS bookworm * 14:33 elukey: rolled out spicerack on all cumin nodes - [[phab:T429699|T429699]] * 14:32 elukey: upgrade all bookworm hosts to pywmflib 3.1 - [[phab:T430552|T430552]] * 14:14 marostegui: Setup x4 eqiad topology [[phab:T404715|T404715]] * 14:13 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 14:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2230.codfw.wmnet * 14:07 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2230.codfw.wmnet * 13:59 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[2001-2002].codfw.wmnet * 13:51 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 13:45 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.major-upgrade (exit_code=97) * 13:45 cwilliams@cumin1003: dbmaint on s4@codfw [[phab:T429893|T429893]] * 13:45 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 13:42 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-master-codfw * 13:42 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl2002.codfw.wmnet * 13:42 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl2002.codfw.wmnet * 13:38 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl2002.codfw.wmnet * 13:38 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl2002.codfw.wmnet * 13:38 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl2001.codfw.wmnet * 13:38 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl2001.codfw.wmnet * 13:35 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl2001.codfw.wmnet * 13:35 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl2001.codfw.wmnet * 13:35 klausman@cumin2002: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-master-codfw * 12:30 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] (duration: 25m 11s) * 12:24 krinkle@deploy1003: krinkle: Continuing with deployment * 12:10 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2048.codfw.wmnet * 12:09 krinkle@deploy1003: krinkle: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:08 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2048.codfw.wmnet * 12:05 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] * 11:57 moritzm: installing curl security updates * 11:49 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:31 moritzm: installing nano security updates * 11:07 moritzm: failover Ganeti master in codfw to ganeti2032 [[phab:T430909|T430909]] * 11:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:04 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest1005.eqiad.wmnet with OS trixie * 11:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:50 jmm@dns1004: END - running authdns-update * 10:47 jmm@dns1004: START - running authdns-update * 10:47 jmm@dns1004: START - running authdns-update * 10:46 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:44 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest1005.eqiad.wmnet with reason: host reimage * 10:38 elukey@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest1005.eqiad.wmnet with reason: host reimage * 10:31 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:31 marostegui: Setup x4 codfw topology [[phab:T404715|T404715]] * 10:31 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 10:24 elukey: spicerack 13.0.0 deployed on cumin2002 * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 10:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 10:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 10:21 elukey@cumin2002: START - Cookbook sre.hosts.reimage for host sretest1005.eqiad.wmnet with OS trixie * 10:20 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:19 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:17 elukey: uploaded spicerack_13.0.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia * 09:54 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:52 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:20 elukey: upgrade all bullseye hosts to pywmflib 3.1 - [[phab:T430552|T430552]] * 09:10 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1015.eqiad.wmnet,service=s4 * 09:10 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1015.eqiad.wmnet,service=s6 * 09:07 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 08:58 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:56 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 08:56 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 08:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 08:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 08:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin2002.codfw.wmnet * 08:06 godog: remove cloudvirt1046, cloudvirt1062, cloudvirt1074, cloudvirt1075 from maintenance aggregate and put them in network-ovs - [[phab:T424802|T424802]] * 08:00 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin2002.codfw.wmnet * 07:58 hashar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] (duration: 32m 53s) * 07:57 fabfur: repooled cp4038 * 07:57 fabfur@cumin1003: conftool action : set/pooled=yes; selector: name=cp4038.* * 07:53 moritzm: installing pyjwt security updates * 07:47 moritzm: installing openjpeg2 security updates * 07:45 hashar@deploy1003: vadymts1, hashar: Continuing with deployment * 07:43 hashar@deploy1003: vadymts1, hashar: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:38 moritzm: installing python-urllib3 security updates * 07:37 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 07:30 fabfur: depooled cp4038 to investigate on possible maxmind failure * 07:30 fabfur@cumin1003: conftool action : set/pooled=no; selector: name=cp4038.* * 07:30 fabfur@cumin1003: conftool action : set/pooled=yes; selector: name=cp4038.* * 07:29 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 07:25 hashar@deploy1003: Started scap sync-world: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] * 06:13 moritzm: installing Linux 6.12.95 on trixie hosts * 05:20 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s6 * 05:20 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s4 * 05:19 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1015.eqiad.wmnet with reason: cloning * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 08s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-05 == * 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 01m 08s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-04 == * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 58s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-03 == * 17:08 topranks: revert protocol preference changes on cr3-ulsfo after upgrade * 16:53 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on cr2-eqord with reason: upgrade JunOS cr3-ulsfo * 16:53 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on cr4-ulsfo with reason: upgrade JunOS cr3-ulsfo * 16:48 topranks: reboot cr3-ulsfo to upgrade JunOS and reset linecard [[phab:T424839|T424839]] * 15:52 topranks: adjust outbound BGP policies on cr3-ulsfo to drain router of traffic [[phab:T424839|T424839]] * 15:45 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on lvs[4008-4010].ulsfo.wmnet with reason: upgrade JunOS cr3-ulsfo * 15:44 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on asw1-[22-23]-ulsfo,cr3-ulsfo,cr3-ulsfo IPv6,cr3-ulsfo.mgmt with reason: upgrade JunOS cr3-ulsfo * 15:36 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 15:35 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 15:35 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 14:40 cmooney@dns3003: END - running authdns-update * 14:26 cmooney@dns3003: START - running authdns-update * 14:26 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:26 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to ulsfo - cmooney@cumin1003" * 14:19 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to ulsfo - cmooney@cumin1003" * 14:16 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:38 sukhe@dns1004: END - running authdns-update * 13:35 sukhe@dns1004: START - running authdns-update * 13:26 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 13:26 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 13:26 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet * 13:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 13:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 13:16 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host sretest1005.eqiad.wmnet * 13:16 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 13:16 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 13:15 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 13:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:14 moritzm: imported samplicator 1.3.8rc1-1+deb13u1 to trixie-wikimedia/main [[phab:T337208|T337208]] * 13:13 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:07 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 13:07 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 13:02 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:02 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:58 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:57 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:57 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:53 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet * 12:50 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 12:47 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:41 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:40 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:39 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:32 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet * 12:26 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet * 12:23 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2005.wikimedia.org * 12:19 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2005.wikimedia.org * 12:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 12:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup[2004-2007].codfw.wmnet * 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[2004-2007].codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin2003" * 12:15 jynus@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[2004-2007].codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin2003" * 12:09 jynus@cumin2003: START - Cookbook sre.dns.netbox * 11:58 jynus@cumin2003: START - Cookbook sre.hosts.decommission for hosts backup[2004-2007].codfw.wmnet * 10:40 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup[1004-1007].eqiad.wmnet * 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[1004-1007].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 10:01 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[1004-1007].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 09:52 jynus@cumin1003: START - Cookbook sre.dns.netbox * 09:39 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:36 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup[1004-1007].eqiad.wmnet * 09:36 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:25 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 09:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 09:16 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 09:05 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:04 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:00 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:59 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:57 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:55 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:50 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 08:50 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 08:49 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 08:49 atsukoito: depooling cirrussearch in codfw because of regression after upgrade [[phab:T431091|T431091]] * 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts mirror1001.wikimedia.org * 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: mirror1001.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 08:29 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: mirror1001.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 08:18 jmm@cumin2003: START - Cookbook sre.dns.netbox * 08:11 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts mirror1001.wikimedia.org * 06:15 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 18s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-02 == * 22:55 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host contint1003.wikimedia.org with OS trixie * 22:29 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on contint1003.wikimedia.org with reason: host reimage * 22:23 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on contint1003.wikimedia.org with reason: host reimage * 22:05 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host contint1003.wikimedia.org with OS trixie * 22:03 mutante: contint1003 (zuul.wikimedia.org) - reimaging because of [[phab:T430510|T430510]]#12067628 [[phab:T418521|T418521]] * 22:03 dzahn@cumin2002: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on zuul.wikimedia.org with reason: reimage * 21:39 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 18s) * 21:39 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 21:20 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1003.eqiad.wmnet, repooling source-only afterwards * 21:19 sbassett: Deployed security fix for [[phab:T428829|T428829]] * 20:58 cmooney@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Release v0.11.2 update for new Aerleon - cmooney@cumin1003 * 20:55 cmooney@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Release v0.11.2 update for new Aerleon - cmooney@cumin1003 * 20:40 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] (duration: 12m 35s) * 20:36 arlolra@deploy1003: cscott, arlolra: Continuing with deployment * 20:35 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 20s) * 20:35 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 20:33 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host contint2003.wikimedia.org with OS trixie * 20:31 arlolra@deploy1003: cscott, arlolra: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Cha * 20:28 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] * 20:17 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] (duration: 08m 13s) * 20:14 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on contint2003.wikimedia.org with reason: host reimage * 20:13 sbassett@deploy1003: sbassett: Continuing with deployment * 20:11 sbassett@deploy1003: sbassett: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:09 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] * 20:08 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 20:08 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 20:08 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on contint2003.wikimedia.org with reason: host reimage * 20:05 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1003.eqiad.wmnet, repooling source-only afterwards * 19:49 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host contint2003.wikimedia.org with OS trixie * 19:48 mutante: contint2003 - reimaging because of [[phab:T430510|T430510]]#12067628 [[phab:T418521|T418521]] * 18:39 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 18:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 18:13 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2002.codfw.wmnet -> wcqs2003.codfw.wmnet, repooling source-only afterwards * 17:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1003.eqiad.wmnet with OS bookworm * 17:52 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1005.eqiad.wmnet * 17:52 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1005.eqiad.wmnet * 17:51 jasmine@cumin2002: conftool action : set/pooled=yes:weight=10; selector: name=wikikube-ctrl1005.eqiad.wmnet * 17:48 jasmine_: homer "cr*eqiad*" commit "Added new stacked control plane wikikube-ctrl1005" * 17:44 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply * 17:44 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply * 17:31 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] (duration: 09m 33s) * 17:26 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 17:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1003.eqiad.wmnet with reason: host reimage * 17:23 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:21 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] * 17:18 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1003.eqiad.wmnet with reason: host reimage * 17:16 rscout@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply * 17:16 rscout@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply * 17:16 rscout@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply * 17:15 rscout@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply * 17:12 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on wcqs[2002-2003].codfw.wmnet,wcqs1002.eqiad.wmnet with reason: reimaging hosts * 17:08 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 17:08 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 17:08 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 17:07 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 17:05 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 17:05 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 17:03 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "running to make sure all updates are synced - cmooney@cumin1003" * 17:03 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "running to make sure all updates are synced - cmooney@cumin1003" * 17:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs1003 * 17:00 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs1003 * 17:00 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 17:00 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1003.eqiad.wmnet with OS bookworm * 16:58 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Re-running - btullis@cumin1003" * 16:58 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Re-running - btullis@cumin1003" * 16:58 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2002.codfw.wmnet -> wcqs2003.codfw.wmnet, repooling source-only afterwards * 16:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-master1004.eqiad.wmnet with OS bookworm * 16:58 btullis@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 16:57 tappof: bump space for prometheus k8s-aux in eqiad * 16:55 cmooney@dns3003: END - running authdns-update * 16:55 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:55 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to eqsin - cmooney@cumin1003" * 16:55 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to eqsin - cmooney@cumin1003" * 16:53 cmooney@dns3003: START - running authdns-update * 16:52 ryankemper: [ml-serve-eqiad] Cleared out 1302 failed (Evicted) pods: `kubectl -n llm delete pods --field-selector=status.phase=Failed`, freeing calico-kube-controllers from OOM crashloop (evictions were caused by disk pressure) * 16:49 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 16:46 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:39 rzl@dns1004: END - running authdns-update * 16:37 rzl@dns1004: START - running authdns-update * 16:36 rzl@dns1004: START - running authdns-update * 16:35 rzl@deploy1003: Finished scap sync-world: [[phab:T416623|T416623]] (duration: 10m 19s) * 16:34 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 16:33 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-master1004.eqiad.wmnet with reason: host reimage * 16:30 rzl@deploy1003: rzl: Continuing with deployment * 16:28 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-master1004.eqiad.wmnet with reason: host reimage * 16:26 rzl@deploy1003: rzl: [[phab:T416623|T416623]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:25 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 16:25 rzl@deploy1003: Started scap sync-world: [[phab:T416623|T416623]] * 16:25 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 16:24 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 16:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: sync * 16:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: sync * 16:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-master1004.eqiad.wmnet with OS bookworm * 16:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-master1003.eqiad.wmnet with OS bookworm * 16:11 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 16:11 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 16:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Security updates * 16:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 16:08 root@cumin1003: START - Cookbook sre.mysql.parsercache * 16:08 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Security updates * 15:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-master1003.eqiad.wmnet with reason: host reimage * 15:54 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:54 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:54 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:54 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-master1003.eqiad.wmnet with reason: host reimage * 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Security updates * 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:45 root@cumin1003: START - Cookbook sre.mysql.parsercache * 15:45 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Security updates * 15:42 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-master1003.eqiad.wmnet with OS bookworm * 15:24 moritzm: installing busybox updates from bookworm point release * 15:20 moritzm: installing busybox updates from trixie point release * 15:15 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1021: Security updates * 15:15 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:15 root@cumin1003: START - Cookbook sre.mysql.parsercache * 15:15 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1021: Security updates * 15:13 moritzm: installing giflib security updates * 15:08 moritzm: installing Tomcat security updates * 14:57 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 14:56 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 14:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:53 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Unblock taavi - oblivian@cumin1003" * 14:53 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Unblock taavi - oblivian@cumin1003 * 14:53 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1021: Security updates * 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:53 root@cumin1003: START - Cookbook sre.mysql.parsercache * 14:53 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1021: Security updates * 14:53 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Unblock taavi - oblivian@cumin1003 * 14:52 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Unblock taavi - oblivian@cumin1003" * 14:46 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94711 and previous config saved to /var/cache/conftool/dbconfig/20260702-144644-fceratto.json * 14:36 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205', diff saved to https://phabricator.wikimedia.org/P94709 and previous config saved to /var/cache/conftool/dbconfig/20260702-143636-fceratto.json * 14:32 moritzm: installing libdbi-perl security updates * 14:26 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205', diff saved to https://phabricator.wikimedia.org/P94708 and previous config saved to /var/cache/conftool/dbconfig/20260702-142628-fceratto.json * 14:16 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94707 and previous config saved to /var/cache/conftool/dbconfig/20260702-141621-fceratto.json * 14:12 moritzm: installing rsync security updates * 14:11 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox) * 14:10 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94706 and previous config saved to /var/cache/conftool/dbconfig/20260702-140959-fceratto.json * 14:09 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2205.codfw.wmnet with reason: Maintenance * 14:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2205: Repooling after switchover * 14:07 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-test-master1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 14:06 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 14:06 Tran: Deployed patch for [[phab:T427287|T427287]] * 14:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:59 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2205: Repooling after switchover * 13:59 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2205: Repooling after switchover * 13:59 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:55 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2205: Repooling after switchover * 13:55 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2205 [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94704 and previous config saved to /var/cache/conftool/dbconfig/20260702-135505-fceratto.json * 13:54 moritzm: installing sed security updates * 13:53 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:52 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2209 to s3 primary [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94703 and previous config saved to /var/cache/conftool/dbconfig/20260702-135235-fceratto.json * 13:52 federico3: Starting s3 codfw failover from db2205 to db2209 - [[phab:T430912|T430912]] * 13:51 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:51 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 13:48 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:47 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2209 with weight 0 [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94702 and previous config saved to /var/cache/conftool/dbconfig/20260702-134719-fceratto.json * 13:47 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Primary switchover s3 [[phab:T430912|T430912]] * 13:44 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:44 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:44 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:40 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 13:38 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 13:37 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 13:36 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 13:36 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:34 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 13:30 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:29 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:29 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:27 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:26 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:25 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling restart_daemons on A:wikidough * 13:23 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 13:22 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 13:17 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 13:17 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns1004.wikimedia.org * 13:12 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:11 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart (exit_code=97) rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough * 13:11 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=97) rolling restart_daemons on A:wikidough * 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough * 13:09 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] (duration: 07m 20s) * 13:05 aude@deploy1003: jdrewniak, aude: Continuing with deployment * 13:04 aude@deploy1003: jdrewniak, aude: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:02 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] * 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts wdqs-categories1001.eqiad.wmnet * 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: wdqs-categories1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 12:10 jmm@dns1004: END - running authdns-update * 12:07 jmm@dns1004: START - running authdns-update * 11:51 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: wdqs-categories1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 11:44 btullis@cumin1003: START - Cookbook sre.dns.netbox * 11:42 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 11:42 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 11:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet * 11:39 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts wdqs-categories1001.eqiad.wmnet * 11:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet * 11:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet * 11:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet * 11:29 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 11:29 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 10:57 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2214: Repooling * 10:49 jmm@dns1004: END - running authdns-update * 10:47 jmm@dns1004: START - running authdns-update * 10:31 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94698 and previous config saved to /var/cache/conftool/dbconfig/20260702-103146-fceratto.json * 10:21 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213', diff saved to https://phabricator.wikimedia.org/P94696 and previous config saved to /var/cache/conftool/dbconfig/20260702-102137-fceratto.json * 10:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:19 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb1017.eqiad.wmnet * 10:18 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 10:18 fceratto@cumin1003: Removing es1033 from zarcillo [[phab:T408772|T408772]] * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts es1033.eqiad.wmnet * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: es1033.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:14 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: es1033.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:13 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb1017.eqiad.wmnet * 10:12 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2214.codfw.wmnet * 10:12 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2214.codfw.wmnet * 10:12 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2214: Repooling * 10:11 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213', diff saved to https://phabricator.wikimedia.org/P94693 and previous config saved to /var/cache/conftool/dbconfig/20260702-101130-fceratto.json * 10:10 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:10 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:03 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts es1033.eqiad.wmnet * 10:03 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 10:01 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94691 and previous config saved to /var/cache/conftool/dbconfig/20260702-100122-fceratto.json * 09:55 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94690 and previous config saved to /var/cache/conftool/dbconfig/20260702-095529-fceratto.json * 09:55 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2213.codfw.wmnet with reason: Maintenance * 09:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 09:53 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2213: Repooling after switchover * 09:51 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover * 09:44 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2213: Repooling after switchover * 09:39 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover * 09:39 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2213 [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94688 and previous config saved to /var/cache/conftool/dbconfig/20260702-093859-fceratto.json * 09:36 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2192 to s5 primary [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94687 and previous config saved to /var/cache/conftool/dbconfig/20260702-093650-fceratto.json * 09:36 federico3: Starting s5 codfw failover from db2213 to db2192 - [[phab:T430923|T430923]] * 09:30 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94686 and previous config saved to /var/cache/conftool/dbconfig/20260702-093004-fceratto.json * 09:24 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2192 with weight 0 [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94685 and previous config saved to /var/cache/conftool/dbconfig/20260702-092455-fceratto.json * 09:24 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 23 hosts with reason: Primary switchover s5 [[phab:T430923|T430923]] * 09:19 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220', diff saved to https://phabricator.wikimedia.org/P94684 and previous config saved to /var/cache/conftool/dbconfig/20260702-091957-fceratto.json * 09:16 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] (duration: 06m 57s) * 09:13 moritzm: installing libgcrypt20 security updates * 09:12 kharlan@deploy1003: kharlan: Continuing with deployment * 09:11 kharlan@deploy1003: kharlan: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:09 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220', diff saved to https://phabricator.wikimedia.org/P94683 and previous config saved to /var/cache/conftool/dbconfig/20260702-090950-fceratto.json * 09:09 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] * 09:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 09:01 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] (duration: 07m 07s) * 08:59 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94682 and previous config saved to /var/cache/conftool/dbconfig/20260702-085942-fceratto.json * 08:57 kharlan@deploy1003: kharlan: Continuing with deployment * 08:56 kharlan@deploy1003: kharlan: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:54 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] * 08:52 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:52 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:52 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94681 and previous config saved to /var/cache/conftool/dbconfig/20260702-085237-fceratto.json * 08:52 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2220.codfw.wmnet with reason: Maintenance * 08:43 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:40 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 08:25 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] (duration: 11m 44s) * 08:21 cscott@deploy1003: cscott: Continuing with deployment * 08:16 cscott@deploy1003: cscott: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:14 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] * 08:08 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 08:08 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1244: Migration of db1244.eqiad.wmnet completed * 08:02 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:02 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:01 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] (duration: 18m 58s) * 08:01 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:59 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 07:59 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:59 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:59 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2006.wikimedia.org * 07:58 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:57 cscott@deploy1003: cscott: Continuing with deployment * 07:56 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:56 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:56 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:55 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:55 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:55 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:54 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2006.wikimedia.org * 07:54 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:54 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 07:44 cscott@deploy1003: cscott: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:44 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2005.wikimedia.org * 07:44 moritzm: installing node-lodash security updates * 07:42 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] * 07:39 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2005.wikimedia.org * 07:30 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] (duration: 07m 28s) * 07:26 cscott@deploy1003: ssastry, cscott: Continuing with deployment * 07:25 cscott@deploy1003: ssastry, cscott: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:23 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1244: Migration of db1244.eqiad.wmnet completed * 07:22 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] * 07:16 wmde-fisch@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] (duration: 06m 55s) * 07:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1244.eqiad.wmnet with OS trixie * 07:11 wmde-fisch@deploy1003: wmde-fisch: Continuing with deployment * 07:11 wmde-fisch@deploy1003: wmde-fisch: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:09 wmde-fisch@deploy1003: Started scap sync-world: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] * 06:54 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1244.eqiad.wmnet with reason: host reimage * 06:50 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1244.eqiad.wmnet with reason: host reimage * 06:38 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1250.eqiad.wmnet with OS trixie * 06:34 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db1244.eqiad.wmnet with OS trixie * 06:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1244: Upgrading db1244.eqiad.wmnet * 06:25 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1244: Upgrading db1244.eqiad.wmnet * 06:25 cwilliams@cumin1003: dbmaint on s4@eqiad [[phab:T429893|T429893]] * 06:25 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 06:15 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1250.eqiad.wmnet with reason: host reimage * 06:14 cwilliams@dns1006: END - running authdns-update * 06:12 cwilliams@dns1006: START - running authdns-update * 06:11 cwilliams@dns1006: END - running authdns-update * 06:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db1244 [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94676 and previous config saved to /var/cache/conftool/dbconfig/20260702-061059-cwilliams.json * 06:09 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1250.eqiad.wmnet with reason: host reimage * 06:09 cwilliams@dns1006: START - running authdns-update * 06:08 aokoth@cumin1003: END (PASS) - Cookbook sre.vrts.upgrade (exit_code=0) on VRTS host vrts1003.eqiad.wmnet * 06:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db1160 to s4 primary and set section read-write [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94675 and previous config saved to /var/cache/conftool/dbconfig/20260702-060746-cwilliams.json * 06:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Set s4 eqiad as read-only for maintenance - [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94674 and previous config saved to /var/cache/conftool/dbconfig/20260702-060704-cwilliams.json * 06:06 cezmunsta: Starting s4 eqiad failover from db1244 to db1160 - [[phab:T430817|T430817]] * 06:04 aokoth@cumin1003: START - Cookbook sre.vrts.upgrade on VRTS host vrts1003.eqiad.wmnet * 05:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db1160 with weight 0 [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94673 and previous config saved to /var/cache/conftool/dbconfig/20260702-055927-cwilliams.json * 05:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 40 hosts with reason: Primary switchover s4 [[phab:T430817|T430817]] * 05:55 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1250.eqiad.wmnet with OS trixie * 05:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on db1250.eqiad.wmnet with reason: m3 master switchover [[phab:T430158|T430158]] * 05:39 marostegui: Failover m3 (phabricator) from db1250 to db1228 - [[phab:T430158|T430158]] * 05:32 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2234].codfw.wmnet,db[1217,1228,1250].eqiad.wmnet with reason: m3 master switchover [[phab:T430158|T430158]] * 04:45 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] (duration: 09m 08s) * 04:41 tstarling@deploy1003: tstarling, reedy: Continuing with deployment * 04:38 tstarling@deploy1003: tstarling, reedy: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 04:36 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 59s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:16 ryankemper: [[phab:T429844|T429844]] [opensearch] completed `cirrussearch2111` reimage; all codfw search clusters are green, all nodes now report `OpenSearch 2.19.5`, and the temporary chi voting exclusion has been removed * 00:57 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2111.codfw.wmnet with OS trixie * 00:29 ryankemper: [[phab:T429844|T429844]] [opensearch] depooled codfw search-omega/search-psi discovery records to match existing codfw search depool during OpenSearch 2.19 migration * 00:29 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2111.codfw.wmnet with reason: host reimage * 00:29 ryankemper@cumin2002: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 00:29 ryankemper@cumin2002: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 00:22 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2111.codfw.wmnet with reason: host reimage * 00:01 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2111.codfw.wmnet with OS trixie * 00:00 ryankemper: [[phab:T429844|T429844]] [opensearch] chi cluster recovered after stopping `opensearch_1@production-search-codfw` on `cirrussearch2111` == 2026-07-01 == * 23:59 ryankemper: [[phab:T429844|T429844]] [opensearch] stopped `opensearch_1@production-search-codfw` on `cirrussearch2111` after chi cluster-manager election churn following `voting_config_exclusions` POST; hoping this triggers a re-election * 23:52 cscott@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 23:51 cscott@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 23:51 cscott@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 23:50 cscott@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2003.codfw.wmnet with OS bookworm * 22:29 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 22:13 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 22:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2084.codfw.wmnet with OS trixie * 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2003.codfw.wmnet with reason: host reimage * 22:03 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 22:01 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2003.codfw.wmnet with reason: host reimage * 21:50 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 21:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2084.codfw.wmnet with reason: host reimage * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2003 * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2003 * 21:42 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2003 * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2003.codfw.wmnet 45.48.192.10.in-addr.arpa 5.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:42 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2003.codfw.wmnet 45.48.192.10.in-addr.arpa 5.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2003 - bking@cumin2003" * 21:42 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2003 - bking@cumin2003" * 21:36 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2084.codfw.wmnet with reason: host reimage * 21:35 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:34 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2003 * 21:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2003.codfw.wmnet with OS bookworm * 21:19 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2084.codfw.wmnet with OS trixie * 21:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2081.codfw.wmnet with OS trixie * 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2108.codfw.wmnet with OS trixie * 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2081.codfw.wmnet with reason: host reimage * 20:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2081.codfw.wmnet with reason: host reimage * 20:28 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2081.codfw.wmnet with OS trixie * 20:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2108.codfw.wmnet with reason: host reimage * 20:19 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2108.codfw.wmnet with reason: host reimage * 19:59 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2108.codfw.wmnet with OS trixie * 19:46 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2093.codfw.wmnet with OS trixie * 19:44 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 19:44 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jasmine@cumin2002" * 19:43 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jasmine@cumin2002" * 19:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2080.codfw.wmnet with OS trixie * 19:28 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 19:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2093.codfw.wmnet with reason: host reimage * 19:18 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 19:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2093.codfw.wmnet with reason: host reimage * 19:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2080.codfw.wmnet with reason: host reimage * 19:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2080.codfw.wmnet with reason: host reimage * 18:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2093.codfw.wmnet with OS trixie * 18:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2080.codfw.wmnet with OS trixie * 18:27 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 18:18 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] (duration: 09m 15s) * 18:13 jgiannelos@deploy1003: jgiannelos, neriah: Continuing with deployment * 18:11 jgiannelos@deploy1003: jgiannelos, neriah: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:09 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] * 17:40 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 16:58 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 30 hosts * 16:57 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for 30 hosts * 16:52 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2202.codfw.wmnet * 16:52 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2202.codfw.wmnet * 16:51 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt * 16:51 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt * 16:51 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lvs2012.codfw.wmnet * 16:51 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for lvs2012.codfw.wmnet * 16:49 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2076.codfw.wmnet with OS trixie * 16:49 brett: Start pybal on lvs2012 - [[phab:T429861|T429861]] * 16:49 pt1979@cumin1003: END (ERROR) - Cookbook sre.hosts.remove-downtime (exit_code=97) for 59 hosts * 16:48 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for 59 hosts * 16:42 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2061.codfw.wmnet with OS trixie * 16:30 dancy@deploy1003: Installation of scap version "4.271.0" completed for 2 hosts * 16:28 dancy@deploy1003: Installing scap version "4.271.0" for 2 host(s) * 16:23 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2076.codfw.wmnet with reason: host reimage * 16:19 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2061.codfw.wmnet with reason: host reimage * 16:18 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2076.codfw.wmnet with reason: host reimage * 16:18 jasmine@dns1004: END - running authdns-update * 16:16 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host restbase2039.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 16:16 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host restbase2039.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 16:16 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2061.codfw.wmnet with reason: host reimage * 16:15 jasmine@dns1004: START - running authdns-update * 16:14 jasmine@dns1004: END - running authdns-update * 16:12 jasmine@dns1004: START - running authdns-update * 16:07 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2202.codfw.wmnet with reason: maintenance * 16:06 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt with reason: Junos upograde * 16:00 papaul: ongoing maintenance on lsw1-b2-codfw * 16:00 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2076.codfw.wmnet with OS trixie * 15:59 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt * 15:59 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt * 15:57 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2061.codfw.wmnet with OS trixie * 15:55 pt1979@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2042,2046].codfw.wmnet * 15:55 pt1979@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2042,2046].codfw.wmnet * 15:51 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 15:51 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2220: Repooling after switchover * 15:50 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 15:50 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 15:48 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2092.codfw.wmnet with OS trixie * 15:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 15:40 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 15:38 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 15:37 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 15:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply * 15:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply * 15:32 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 15:32 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 15:30 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 15:29 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 15:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 15:25 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:22 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs2012.codfw.wmnet with reason: Rack B2 maintenance - [[phab:T429861|T429861]] * 15:21 brett: Stopping pybal on lvs2012 in preparation for codfw rack b2 maintenance - [[phab:T429861|T429861]] * 15:20 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2092.codfw.wmnet with reason: host reimage * 15:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:12 _joe_: restarted manually alertmanager-irc-relay * 15:12 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2092.codfw.wmnet with reason: host reimage * 15:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:12 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt with reason: Junos upograde * 15:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover * 15:07 pt1979@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2042,2046].codfw.wmnet * 15:06 pt1979@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2042,2046].codfw.wmnet * 15:02 papaul: ongoing maintenance on lsw1-a8-codfw * 14:31 topranks: POWERING DOWN CR1-EQIAD for line card installation [[phab:T426343|T426343]] * 14:31 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] (duration: 08m 57s) * 14:29 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:26 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 14:24 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:22 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] * 14:22 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover * 14:16 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:15 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover * 14:14 topranks: re-enable routing-engine graceful-failover on cr1-eqiad [[phab:T417873|T417873]] * 14:13 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:13 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2220: Repooling after switchover * 14:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:12 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:12 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:11 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:08 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] (duration: 10m 01s) * 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:07 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2220 [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94664 and previous config saved to /var/cache/conftool/dbconfig/20260701-140729-fceratto.json * 14:06 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:06 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:06 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:05 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2159 to s7 primary [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94663 and previous config saved to /var/cache/conftool/dbconfig/20260701-140503-fceratto.json * 14:04 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:04 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 14:04 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 14:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:04 dreamyjazz@deploy1003: anzx, dreamyjazz: Continuing with deployment * 14:04 federico3: Starting s7 codfw failover from db2220 to db2159 - [[phab:T430826|T430826]] * 14:03 jmm@dns1004: END - running authdns-update * 14:03 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:03 topranks: flipping cr1-eqiad active routing-enginer back to RE0 [[phab:T417873|T417873]] * 14:03 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudsw1-c8-eqiad,cloudsw1-d5-eqiad with reason: router upgrades eqiad * 14:01 jmm@dns1004: START - running authdns-update * 14:00 dreamyjazz@deploy1003: anzx, dreamyjazz: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:59 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2159 with weight 0 [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94662 and previous config saved to /var/cache/conftool/dbconfig/20260701-135906-fceratto.json * 13:58 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] * 13:57 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s7 [[phab:T430826|T430826]] * 13:56 topranks: reboot routing-enginer RE0 on cr1-eqiad [[phab:T417873|T417873]] * 13:48 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1006.wikimedia.org * 13:44 atsuko@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cirrussearch2092.codfw.wmnet with OS trixie * 13:43 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1006.wikimedia.org * 13:41 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2092.codfw.wmnet with OS trixie * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1005.wikimedia.org * 13:37 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1005.wikimedia.org * 13:37 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on pfw1-eqiad with reason: router upgrades eqiad * 13:35 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on lvs[1017-1020].eqiad.wmnet with reason: router upgrades eqiad * 13:34 caro@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] (duration: 07m 59s) * 13:30 caro@deploy1003: caro: Continuing with deployment * 13:28 caro@deploy1003: caro: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:27 topranks: route-engine failover cr1-eqiad * 13:26 caro@deploy1003: Started scap sync-world: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] * 13:15 topranks: rebooting routing-engine 1 on cr1-eqiad [[phab:T417873|T417873]] * 13:13 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] (duration: 08m 29s) * 13:13 moritzm: installing qemu security updates * 13:11 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 13:11 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 13:09 jgiannelos@deploy1003: jgiannelos: Continuing with deployment * 13:08 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 13:07 jgiannelos@deploy1003: jgiannelos: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:06 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 13:06 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2214.codfw.wmnet with reason: Maintenance * 13:05 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2214: Repooling after switchover * 13:05 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] * 13:04 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2214: Repooling after switchover * 13:04 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2214 [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94660 and previous config saved to /var/cache/conftool/dbconfig/20260701-130413-fceratto.json * 13:01 moritzm: installing python3.13 security updates * 13:00 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2229 to s6 primary [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94659 and previous config saved to /var/cache/conftool/dbconfig/20260701-125959-fceratto.json * 12:59 federico3: Starting s6 codfw failover from db2214 to db2229 - [[phab:T430814|T430814]] * 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on 13 hosts with reason: router upgrade and line card install * 12:51 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2229 with weight 0 [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94658 and previous config saved to /var/cache/conftool/dbconfig/20260701-125149-fceratto.json * 12:51 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 21 hosts with reason: Primary switchover s6 [[phab:T430814|T430814]] * 12:50 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2189.codfw.wmnet * 12:50 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2189.codfw.wmnet * 12:42 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2100.codfw.wmnet with OS trixie * 12:38 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2083.codfw.wmnet with OS trixie * 12:19 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2083.codfw.wmnet with reason: host reimage * 12:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 12:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2240: Migration of db2240.codfw.wmnet completed * 12:14 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2100.codfw.wmnet with reason: host reimage * 12:09 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2083.codfw.wmnet with reason: host reimage * 12:09 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2100.codfw.wmnet with reason: host reimage * 12:00 topranks: drain traffic on cr1-eqiad to allow for line card install and JunOS upgrade [[phab:T426343|T426343]] * 11:52 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2083.codfw.wmnet with OS trixie * 11:50 cmooney@dns2005: END - running authdns-update * 11:49 cmooney@dns2005: START - running authdns-update * 11:48 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2100.codfw.wmnet with OS trixie * 11:40 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/zotero: apply * 11:40 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/zotero: apply * 11:36 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/zotero: apply * 11:36 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/zotero: apply * 11:31 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2240: Migration of db2240.codfw.wmnet completed * 11:30 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply * 11:28 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply * 11:27 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:27 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:27 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:27 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:27 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:23 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2240.codfw.wmnet with OS trixie * 11:20 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:20 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:17 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:16 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:16 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:15 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2086.codfw.wmnet with OS trixie * 11:14 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2106.codfw.wmnet with OS trixie * 11:14 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:13 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:12 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:09 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2115.codfw.wmnet with OS trixie * 11:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2240.codfw.wmnet with reason: host reimage * 11:00 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2240.codfw.wmnet with reason: host reimage * 10:53 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2106.codfw.wmnet with reason: host reimage * 10:49 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2086.codfw.wmnet with reason: host reimage * 10:44 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2115.codfw.wmnet with reason: host reimage * 10:44 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2240.codfw.wmnet with OS trixie * 10:44 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2086.codfw.wmnet with reason: host reimage * 10:42 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2106.codfw.wmnet with reason: host reimage * 10:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2240: Upgrading db2240.codfw.wmnet * 10:41 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2240: Upgrading db2240.codfw.wmnet * 10:41 cwilliams@cumin1003: dbmaint on s4@codfw [[phab:T429893|T429893]] * 10:40 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 10:39 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2115.codfw.wmnet with reason: host reimage * 10:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2240 [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94653 and previous config saved to /var/cache/conftool/dbconfig/20260701-102658-cwilliams.json * 10:26 moritzm: installing nginx security updates * 10:26 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2086.codfw.wmnet with OS trixie * 10:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2179 to s4 primary [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94652 and previous config saved to /var/cache/conftool/dbconfig/20260701-102356-cwilliams.json * 10:23 cezmunsta: Starting s4 codfw failover from db2240 to db2179 - [[phab:T430127|T430127]] * 10:23 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2106.codfw.wmnet with OS trixie * 10:20 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2115.codfw.wmnet with OS trixie * 10:15 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2179 with weight 0 [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94651 and previous config saved to /var/cache/conftool/dbconfig/20260701-101531-cwilliams.json * 10:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 40 hosts with reason: Primary switchover s4 [[phab:T430127|T430127]] * 09:56 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template (take 2) - oblivian@cumin1003" * 09:56 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template (take 2) - oblivian@cumin1003 * 09:55 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template (take 2) - oblivian@cumin1003 * 09:55 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template (take 2) - oblivian@cumin1003" * 09:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:39 mszwarc@deploy1003: Synchronized private/SuggestedInvestigationsSignals/SuggestedInvestigationsSignal4n.php: Update SI signal 4n (duration: 06m 08s) * 09:21 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 09:21 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 09:14 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 09:14 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 09:02 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 09:02 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 08:54 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 08:38 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 08:38 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 08:36 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 08:21 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 08:21 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] (duration: 36m 11s) * 08:15 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 08:09 mszwarc@deploy1003: mszwarc, abi: Continuing with deployment * 08:03 mszwarc@deploy1003: mszwarc, abi: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:55 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 07:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 07:45 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] * 07:30 aqu@deploy1003: Finished deploy [analytics/refinery@410f205]: Regular analytics weekly train 2nd try [analytics/refinery@410f2050] (duration: 00m 22s) * 07:29 aqu@deploy1003: Started deploy [analytics/refinery@410f205]: Regular analytics weekly train 2nd try [analytics/refinery@410f2050] * 07:28 aqu@deploy1003: Finished deploy [analytics/refinery@410f205] (thin): Regular analytics weekly train THIN [analytics/refinery@410f2050] (duration: 01m 59s) * 07:28 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] (duration: 07m 19s) * 07:26 aqu@deploy1003: Started deploy [analytics/refinery@410f205] (thin): Regular analytics weekly train THIN [analytics/refinery@410f2050] * 07:26 aqu@deploy1003: Finished deploy [analytics/refinery@410f205]: Regular analytics weekly train [analytics/refinery@410f2050] (duration: 04m 32s) * 07:24 mszwarc@deploy1003: wmde-fisch, mszwarc: Continuing with deployment * 07:23 mszwarc@deploy1003: wmde-fisch, mszwarc: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:21 aqu@deploy1003: Started deploy [analytics/refinery@410f205]: Regular analytics weekly train [analytics/refinery@410f2050] * 07:21 aqu@deploy1003: Finished deploy [analytics/refinery@410f205] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@410f2050] (duration: 02m 01s) * 07:20 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] * 07:19 aqu@deploy1003: Started deploy [analytics/refinery@410f205] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@410f2050] * 07:13 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] (duration: 09m 13s) * 07:09 mszwarc@deploy1003: mszwarc, chlod, revi: Continuing with deployment * 07:06 mszwarc@deploy1003: mszwarc, chlod, revi: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:04 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] * 06:55 elukey: upgrade all trixie hosts to pywmflib 3.0 - [[phab:T430552|T430552]] * 06:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:43 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:43 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:42 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:42 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:41 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:41 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:35 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:35 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:34 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:34 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:31 jmm@cumin2003: DONE (PASS) - Cookbook sre.idm.logout (exit_code=0) Logging Niharika29 out of all services on: 2453 hosts * 06:30 oblivian@cumin1003: END (FAIL) - Cookbook sre.deploy.hiddenparma (exit_code=99) Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:30 oblivian@cumin1003: END (FAIL) - Cookbook sre.deploy.python-code (exit_code=99) hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:30 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:30 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:01 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2109.codfw.wmnet with OS trixie * 05:45 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on es1039.eqiad.wmnet with reason: issues * 05:41 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1027.eqiad.wmnet * 05:40 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2068.codfw.wmnet with OS trixie * 05:40 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2109.codfw.wmnet with reason: host reimage * 05:40 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1027.eqiad.wmnet,service=s2 * 05:40 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1027.eqiad.wmnet,service=s7 * 05:36 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2109.codfw.wmnet with reason: host reimage * 05:20 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2068.codfw.wmnet with reason: host reimage * 05:16 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2109.codfw.wmnet with OS trixie * 05:15 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2068.codfw.wmnet with reason: host reimage * 05:09 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2067.codfw.wmnet with OS trixie * 04:56 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2068.codfw.wmnet with OS trixie * 04:49 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2067.codfw.wmnet with reason: host reimage * 04:45 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2067.codfw.wmnet with reason: host reimage * 04:27 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2067.codfw.wmnet with OS trixie * 03:47 slyngshede@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1039.eqiad.wmnet with reason: Hardware crash * 03:21 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2107.codfw.wmnet with OS trixie * 02:59 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2107.codfw.wmnet with reason: host reimage * 02:55 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2085.codfw.wmnet with OS trixie * 02:51 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2072.codfw.wmnet with OS trixie * 02:51 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2107.codfw.wmnet with reason: host reimage * 02:35 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2085.codfw.wmnet with reason: host reimage * 02:31 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2107.codfw.wmnet with OS trixie * 02:30 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2072.codfw.wmnet with reason: host reimage * 02:26 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2085.codfw.wmnet with reason: host reimage * 02:22 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2072.codfw.wmnet with reason: host reimage * 02:09 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2085.codfw.wmnet with OS trixie * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 54s) * 02:03 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2072.codfw.wmnet with OS trixie * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es7 eqiad back to read-write - [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94649 and previous config saved to /var/cache/conftool/dbconfig/20260701-010716-ladsgroup.json * 01:05 ladsgroup@dns1004: END - running authdns-update * 01:05 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depool es1039 [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94648 and previous config saved to /var/cache/conftool/dbconfig/20260701-010551-ladsgroup.json * 01:03 ladsgroup@dns1004: START - running authdns-update * 01:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Promote es1035 to es7 primary [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94647 and previous config saved to /var/cache/conftool/dbconfig/20260701-010002-ladsgroup.json * 00:58 Amir1: Starting es7 eqiad failover from es1039 to es1035 - [[phab:T430765|T430765]] * 00:53 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es1035 with weight 0 [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94646 and previous config saved to /var/cache/conftool/dbconfig/20260701-005329-ladsgroup.json * 00:53 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 9 hosts with reason: Primary switchover es7 [[phab:T430765|T430765]] * 00:42 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es7 eqiad as read-only for maintenance - [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94645 and previous config saved to /var/cache/conftool/dbconfig/20260701-004221-ladsgroup.json * 00:20 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2102.codfw.wmnet with OS trixie * 00:15 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2103.codfw.wmnet with OS trixie * 00:05 dr0ptp4kt: DEPLOYED Refinery at {{Gerrit|4e7a2b32}} for changes: pageview allowlist {{Gerrit|1305158}} (+min.wikiquote) {{Gerrit|1305162}} (+bol.wikipedia), {{Gerrit|1305156}} (+isv.wikipedia); {{Gerrit|1305980}} (pv allowlist -api.wikimedia, sqoop +isvwiki); sqoop {{Gerrit|1295064}} (+globalimagelinks) {{Gerrit|1295069}} (+filerevision) using scap, then deployed onto HDFS (manual copyToLocal required additionally) == Other archives == See [[Server Admin Log/Archives]]. <noinclude> [[Category:SAL]] [[Category:Operations]] </noinclude> sxf5nxb1unirr1tgt4j5xb4yn0use4s 2450653 2450652 2026-08-22T16:36:12Z Stashbot 7414 arlolra@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply 2450653 wikitext text/x-wiki == 2026-08-22 == * 16:36 arlolra@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:36 arlolra@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:33 arlolra@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:33 arlolra@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:33 arlolra@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:32 arlolra@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 35s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-21 == * 20:36 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in1001.wikimedia.org with reason: [[phab:T434750|T434750]] * 20:34 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in2001.wikimedia.org with reason: [[phab:T434750|T434750]] * 20:33 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out1001.wikimedia.org with reason: [[phab:T434750|T434750]] * 20:25 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out2001.wikimedia.org with reason: [[phab:T434750|T434750]] * 19:37 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host krb1002.eqiad.wmnet with OS bookworm * 19:00 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 18:59 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 18:51 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 18:51 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 18:35 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:35 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:27 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:27 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:16 bking@cumin2003: START - Cookbook sre.hosts.reimage for host krb1002.eqiad.wmnet with OS bookworm * 17:35 sukhe@dns1004: END - running authdns-update * 17:33 sukhe@dns1004: START - running authdns-update * 17:32 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns5004.wikimedia.org [reason: resolved authdns-update issues] * 17:32 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:32 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: force HEAD to {{Gerrit|be26e30ae101}} - sukhe@cumin1003" * 17:32 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: force HEAD to {{Gerrit|be26e30ae101}} - sukhe@cumin1003" * 17:28 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 17:28 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: service=authdns-update,name=dns5004.wikimedia.org [reason: resolving authdns-update issues] * 17:28 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:28 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: force HEAD to {{Gerrit|be26e30ae101}} - sukhe@cumin1003" * 17:28 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: force HEAD to {{Gerrit|be26e30ae101}} - sukhe@cumin1003" * 17:24 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 17:24 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.netbox (exit_code=97) * 17:23 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 17:16 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 17:12 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 17:10 sukhe@dns1004: END - running authdns-update * 17:08 sukhe@dns1004: START - running authdns-update * 17:08 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=dns5004.wikimedia.org [reason: resolving authdns-update issues] * 17:07 sukhe@dns1004: FAIL - running authdns-update * 17:05 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 17:05 sukhe@dns1004: START - running authdns-update * 17:01 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 16:59 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 16:56 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 16:53 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=dns5004.* [reason: trixie upgrade] * 16:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns5004.wikimedia.org * 16:52 cdobbins@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns5004.wikimedia.org * 16:44 cmooney@dns3003: END - running authdns-update * 16:41 cmooney@dns3003: START - running authdns-update * 16:41 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:41 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on eqsin<->codfw arelion - cmooney@cumin1003" * 16:37 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on eqsin<->codfw arelion - cmooney@cumin1003" * 16:33 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:11 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 16:08 sukhe@dns1004: END - running authdns-update * 16:08 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 16:06 sukhe@dns1004: START - running authdns-update * 16:04 cmooney@dns3003: END - running authdns-update * 16:02 cmooney@dns3003: START - running authdns-update * 16:00 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:00 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on eqord<->codfw arelion - cmooney@cumin1003" * 15:56 cdobbins@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host dns5004.wikimedia.org with OS trixie * 15:55 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on eqord<->codfw arelion - cmooney@cumin1003" * 15:53 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:51 cmooney@cumin1003: END (ERROR) - Cookbook sre.dns.netbox (exit_code=97) * 15:51 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:38 andrew@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudcephosd1042.eqiad.wmnet with OS bookworm * 15:18 andrew@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudcephosd1042.eqiad.wmnet with reason: host reimage * 15:17 dancy@deploy1003: Finished deploy [gerrit/gerrit@2cc11cc]: Deploying https://gerrit.wikimedia.org/r/c/operations/software/gerrit/+/1327669 ([[phab:T434726|T434726]]) (duration: 00m 14s) * 15:17 dancy@deploy1003: Started deploy [gerrit/gerrit@2cc11cc]: Deploying https://gerrit.wikimedia.org/r/c/operations/software/gerrit/+/1327669 ([[phab:T434726|T434726]]) * 15:13 andrew@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudcephosd1042.eqiad.wmnet with reason: host reimage * 15:09 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns5004.wikimedia.org with reason: host reimage * 15:05 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns5004.wikimedia.org with reason: host reimage * 14:53 andrew@cumin2003: START - Cookbook sre.hosts.reimage for host cloudcephosd1042.eqiad.wmnet with OS bookworm * 14:30 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns5004.wikimedia.org with OS trixie * 14:29 cdobbins@cumin1003: conftool action : set/pooled=no; selector: name=dns5004.* [reason: trixie upgrade] * 14:21 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:21 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove entries for cr2-eqord - cmooney@cumin1003" * 14:21 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove entries for cr2-eqord - cmooney@cumin1003" * 14:13 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 14:11 moritzm: imported openjdk 8u504-ga-1~deb12u1 for bookworm-wikimedia (backport of the latest Java 8 security fixes for bookworm) * 13:25 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "sync cr2-eqord router offline - cmooney@cumin1003" * 13:23 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "sync cr2-eqord router offline - cmooney@cumin1003" * 13:14 hashar@deploy1003: Finished deploy [integration/docroot@2d5ff9b]: opensource: add PersonalDashboard docs to MW components - [[phab:T435392|T435392]] (duration: 00m 15s) * 13:14 hashar@deploy1003: Started deploy [integration/docroot@2d5ff9b]: opensource: add PersonalDashboard docs to MW components - [[phab:T435392|T435392]] * 12:16 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:16 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: [[phab:T431682|T431682]] - filippo@cumin1003" * 12:16 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: [[phab:T431682|T431682]] - filippo@cumin1003" * 12:11 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2006.wikimedia.org with OS trixie * 12:00 kevinbazira@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 11:58 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 11:43 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2006.wikimedia.org with reason: host reimage * 11:41 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 11:38 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2006.wikimedia.org with reason: host reimage * 11:20 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2006.wikimedia.org with OS trixie * 11:11 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2005.wikimedia.org with OS trixie * 10:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2005.wikimedia.org with reason: host reimage * 10:53 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2005.wikimedia.org with reason: host reimage * 10:43 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-codfw * 10:43 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2011.codfw.wmnet * 10:43 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2011.codfw.wmnet * 10:40 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2011.codfw.wmnet * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2011.codfw.wmnet * 10:34 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2010.codfw.wmnet * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2010.codfw.wmnet * 10:33 fnegri@cumin1003: END (PASS) - Cookbook sre.wikireplicas.add-wiki (exit_code=0) for database bolwiki ([[phab:T429954|T429954]]) * 10:33 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2005.wikimedia.org with OS trixie * 10:30 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2010.codfw.wmnet * 10:25 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2010.codfw.wmnet * 10:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2009.codfw.wmnet * 10:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2009.codfw.wmnet * 10:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1006.wikimedia.org with OS trixie * 10:18 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2009.codfw.wmnet * 10:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2009.codfw.wmnet * 10:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2008.codfw.wmnet * 10:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2008.codfw.wmnet * 10:06 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2008.codfw.wmnet * 10:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1006.wikimedia.org with reason: host reimage * 10:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2008.codfw.wmnet * 10:01 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2007.codfw.wmnet * 10:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2007.codfw.wmnet * 09:57 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1006.wikimedia.org with reason: host reimage * 09:56 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2007.codfw.wmnet * 09:51 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2007.codfw.wmnet * 09:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 09:51 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 09:46 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2006.codfw.wmnet * 09:46 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1006.wikimedia.org with OS trixie * 09:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1005.wikimedia.org with OS trixie * 09:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2006.codfw.wmnet * 09:41 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2005.codfw.wmnet * 09:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2005.codfw.wmnet * 09:38 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2013.codfw.wmnet * 09:36 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2005.codfw.wmnet * 09:35 fnegri@cumin1003: START - Cookbook sre.wikireplicas.add-wiki for database bolwiki ([[phab:T429954|T429954]]) * 09:35 fnegri@cumin1003: END (PASS) - Cookbook sre.wikireplicas.add-wiki (exit_code=0) for database minwikiquote ([[phab:T429946|T429946]]) * 09:35 fnegri@cumin1003: START - Cookbook sre.wikireplicas.add-wiki for database minwikiquote ([[phab:T429946|T429946]]) * 09:32 blake@cumin1003: START - Cookbook sre.hosts.reboot-single for host rdb2013.codfw.wmnet * 09:30 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2011.codfw.wmnet * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1005.wikimedia.org with reason: host reimage * 09:26 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2005.codfw.wmnet * 09:26 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2004.codfw.wmnet * 09:26 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2004.codfw.wmnet * 09:24 blake@cumin1003: START - Cookbook sre.hosts.reboot-single for host rdb2011.codfw.wmnet * 09:22 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1005.wikimedia.org with reason: host reimage * 09:21 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2004.codfw.wmnet * 09:16 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1015.eqiad.wmnet * 09:16 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2004.codfw.wmnet * 09:16 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2003.codfw.wmnet * 09:16 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2003.codfw.wmnet * 09:13 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on cr[1-2]-eqiad,pfw1-eqiad with reason: upgrade pfw1a-eqiad and pfw1b-eqiad pair * 09:12 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 09:11 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 09:11 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 09:11 blake@cumin1003: START - Cookbook sre.hosts.reboot-single for host rdb1015.eqiad.wmnet * 09:10 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2003.codfw.wmnet * 09:09 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1013.eqiad.wmnet * 09:07 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1005.wikimedia.org with OS trixie * 09:03 blake@cumin1003: START - Cookbook sre.hosts.reboot-single for host rdb1013.eqiad.wmnet * 09:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2003.codfw.wmnet * 09:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2002.codfw.wmnet * 09:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2002.codfw.wmnet * 08:54 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2002.codfw.wmnet * 08:49 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2002.codfw.wmnet * 08:49 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2001.codfw.wmnet * 08:49 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2001.codfw.wmnet * 08:48 jmm@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts netmon2002.wikimedia.org * 08:47 jmm@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts netmon2002.wikimedia.org * 08:44 jmm@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts netmon2002.wikimedia.org * 08:44 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon2002.wikimedia.org * 08:43 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2001.codfw.wmnet * 08:36 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon2002.wikimedia.org * 08:34 jmm@dns1004: END - running authdns-update * 08:33 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2001.codfw.wmnet * 08:33 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-codfw * 08:31 jmm@dns1004: START - running authdns-update * 07:48 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327679{{!}}Block: Disable flaky API test (T435272 T389028)]], [[gerrit:1327678{{!}}API: wfDebugLog for thumberror]] (duration: 15m 34s) * 07:41 krinkle@deploy1003: krinkle: Continuing with deployment * 07:37 krinkle@deploy1003: krinkle: Backport for [[gerrit:1327679{{!}}Block: Disable flaky API test (T435272 T389028)]], [[gerrit:1327678{{!}}API: wfDebugLog for thumberror]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:33 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1327679{{!}}Block: Disable flaky API test (T435272 T389028)]], [[gerrit:1327678{{!}}API: wfDebugLog for thumberror]] * 07:25 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 07:24 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 07:18 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 07:18 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 07:15 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 07:14 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 07:14 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 07:14 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 07:13 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 07:03 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1283: Pool back * 06:42 jmm@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts netmon2002.wikimedia.org * 06:35 moritzm: powercycling netmon2002 * 06:18 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1283: Pool back * 06:17 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1283 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96213 and previous config saved to /var/cache/conftool/dbconfig/20260821-061743-marostegui.json * 04:59 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324963{{!}}Add Produnto to extension-list (T421436)]], [[gerrit:1324964{{!}}Enable Produnto on Beta (T421436)]] (duration: 34m 48s) * 04:45 tstarling@deploy1003: tstarling: Continuing with deployment * 04:44 tstarling@deploy1003: tstarling: Backport for [[gerrit:1324963{{!}}Add Produnto to extension-list (T421436)]], [[gerrit:1324964{{!}}Enable Produnto on Beta (T421436)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 04:24 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1324963{{!}}Add Produnto to extension-list (T421436)]], [[gerrit:1324964{{!}}Enable Produnto on Beta (T421436)]] * 04:21 arlolra@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 04:20 arlolra@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 04:20 arlolra@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 04:20 arlolra@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 41s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-20 == * 23:43 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327652{{!}}RunSingleJob: Add ProfilingContext::init() (T435422)]] (duration: 12m 06s) * 23:38 krinkle@deploy1003: krinkle: Continuing with deployment * 23:33 krinkle@deploy1003: krinkle: Backport for [[gerrit:1327652{{!}}RunSingleJob: Add ProfilingContext::init() (T435422)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:31 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1327652{{!}}RunSingleJob: Add ProfilingContext::init() (T435422)]] * 22:15 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1054.eqiad.wmnet with OS trixie * 22:14 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 22:14 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 21:58 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1054.eqiad.wmnet with reason: host reimage * 21:51 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1054.eqiad.wmnet with reason: host reimage * 21:36 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1054.eqiad.wmnet with OS trixie * 21:36 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:35 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327219{{!}}RunSingleJob: Define MW_ENTRY_POINT for flamegraph sample attribution (T435422)]] (duration: 08m 30s) * 21:31 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:31 krinkle@deploy1003: krinkle: Continuing with deployment * 21:31 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1054 * 21:31 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1054 * 21:29 krinkle@deploy1003: krinkle: Backport for [[gerrit:1327219{{!}}RunSingleJob: Define MW_ENTRY_POINT for flamegraph sample attribution (T435422)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:27 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1327219{{!}}RunSingleJob: Define MW_ENTRY_POINT for flamegraph sample attribution (T435422)]] * 21:17 maryum: Deployed security fix for [[phab:T433020|T433020]] * 20:59 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324752{{!}}InitialiseSettings: Enable 2FA banners on remaining private wikis (T428103)]], [[gerrit:1325920{{!}}Remove sending email to legal team about rejected requests (T374053)]] (duration: 07m 18s) * 20:54 reedy@deploy1003: neriah, reedy: Continuing with deployment * 20:54 reedy@deploy1003: neriah, reedy: Backport for [[gerrit:1324752{{!}}InitialiseSettings: Enable 2FA banners on remaining private wikis (T428103)]], [[gerrit:1325920{{!}}Remove sending email to legal team about rejected requests (T374053)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:51 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324752{{!}}InitialiseSettings: Enable 2FA banners on remaining private wikis (T428103)]], [[gerrit:1325920{{!}}Remove sending email to legal team about rejected requests (T374053)]] * 20:24 reedy@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.15,1.47.0-wmf.16,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/med * 20:23 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324752{{!}}InitialiseSettings: Enable 2FA banners on remaining private wikis (T428103)]], [[gerrit:1325920{{!}}Remove sending email to legal team about rejected requests (T374053)]] * 20:14 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 20:10 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 20:09 cdanis@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "bug fixes & UX fixes - cdanis@cumin1003" * 20:09 cdanis@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: bug fixes & UX fixes - cdanis@cumin1003 * 20:08 cdanis@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: bug fixes & UX fixes - cdanis@cumin1003 * 20:08 cdanis@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "bug fixes & UX fixes - cdanis@cumin1003" * 19:24 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327598{{!}}Make \Omicron non upright (like \Chi) (T434428)]], [[gerrit:1327596{{!}}Render overline of \bar with stretchy=false (T435456)]] (duration: 18m 54s) * 19:20 krinkle@deploy1003: krinkle: Continuing with deployment * 19:07 krinkle@deploy1003: krinkle: Backport for [[gerrit:1327598{{!}}Make \Omicron non upright (like \Chi) (T434428)]], [[gerrit:1327596{{!}}Render overline of \bar with stretchy=false (T435456)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:05 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1327598{{!}}Make \Omicron non upright (like \Chi) (T434428)]], [[gerrit:1327596{{!}}Render overline of \bar with stretchy=false (T435456)]] * 18:51 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327614{{!}}Avoid casting fpxmax to string (T318419)]] (duration: 07m 28s) * 18:50 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:46 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:46 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 18:45 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327614{{!}}Avoid casting fpxmax to string (T318419)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:43 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327614{{!}}Avoid casting fpxmax to string (T318419)]] * 18:37 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:37 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:37 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:36 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:07 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 18:05 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 18:01 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 18:01 sukhe@dns1004: END - running authdns-update * 17:59 sukhe@dns1004: START - running authdns-update * 17:58 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 17:57 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=dns6002.* [reason: depooling for trixie upgrade] * 17:56 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns6002.wikimedia.org * 17:56 cdobbins@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns6002.wikimedia.org * 17:51 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 17:51 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 17:34 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host stat1011.eqiad.wmnet with OS bookworm * 17:31 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns6002.wikimedia.org with OS trixie * 17:30 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 17:30 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 17:30 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 17:30 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 17:29 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 17:29 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 17:29 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 17:29 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:29 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:27 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:24 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327590{{!}}Make sure fpsmax is an int value (T318419)]] (duration: 08m 37s) * 17:20 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 17:17 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327590{{!}}Make sure fpsmax is an int value (T318419)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:16 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 17:15 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327590{{!}}Make sure fpsmax is an int value (T318419)]] * 16:51 swfrench-wmf: disable-puppet on A:cp for ATS Lua change - [[phab:T427666|T427666]] * 16:51 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db2901.codfw.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 16:43 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 16:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on stat1011.eqiad.wmnet with reason: host reimage * 16:39 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns6002.wikimedia.org with reason: host reimage * 16:36 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on stat1011.eqiad.wmnet with reason: host reimage * 16:36 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db2901.codfw.wmnet * 16:34 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns6002.wikimedia.org with reason: host reimage * 16:15 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns6002.wikimedia.org with OS trixie * 16:14 cdobbins@cumin1003: conftool action : set/pooled=no; selector: name=dns6002.* [reason: depooling for trixie upgrade] * 16:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host stat1011 * 16:10 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host stat1011 * 16:09 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host stat1011 * 16:09 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) stat1011.eqiad.wmnet 14.36.64.10.in-addr.arpa 4.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 bking@cumin2003: START - Cookbook sre.dns.wipe-cache stat1011.eqiad.wmnet 14.36.64.10.in-addr.arpa 4.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:09 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host stat1011 - bking@cumin2003" * 16:09 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host stat1011 - bking@cumin2003" * 16:05 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host stat1011 * 16:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host stat1011.eqiad.wmnet with OS bookworm * 16:00 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 15:59 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 15:56 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 15:55 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 15:37 jayme@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:35 jayme@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 15:35 jayme@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:32 jayme@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:32 jayme@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:30 jayme@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 15:30 jayme@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:28 jayme@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:28 jayme@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 15:28 fceratto@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host db1903.eqiad.wmnet * 15:28 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1903.eqiad.wmnet with OS trixie * 15:26 jayme@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 15:26 jayme@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 15:24 jayme@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 15:24 jayme@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 15:21 jayme@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 15:21 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 15:19 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 15:19 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 15:17 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 15:14 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1903.eqiad.wmnet with reason: host reimage * 15:07 fceratto@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1903.eqiad.wmnet with reason: host reimage * 14:54 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db1903.eqiad.wmnet with OS trixie * 14:53 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1903.eqiad.wmnet - fceratto@cumin1003" * 14:53 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1903.eqiad.wmnet - fceratto@cumin1003" * 14:53 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1903.eqiad.wmnet on all recursors * 14:53 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1903.eqiad.wmnet on all recursors * 14:53 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:53 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1903.eqiad.wmnet - fceratto@cumin1003" * 14:53 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1903.eqiad.wmnet - fceratto@cumin1003" * 14:49 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 14:49 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1903.eqiad.wmnet * 14:33 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=cp1100.* * 14:27 topranks: reconfigure eqiad<->codfw bgp settings * 14:22 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:22 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update entries used on new transport backup eqiad codfw - cmooney@cumin1003" * 14:19 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update entries used on new transport backup eqiad codfw - cmooney@cumin1003" * 14:14 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 14:14 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 14:13 moritzm: installing util-linux security updates * 14:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-staging-worker * 14:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2003.codfw.wmnet * 14:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2003.codfw.wmnet * 14:08 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2003.codfw.wmnet * 14:06 moritzm: installing libheif security updates * 13:58 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2003.codfw.wmnet * 13:58 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2002.codfw.wmnet * 13:58 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2002.codfw.wmnet * 13:56 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:56 fnegri@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for clouddb1025.eqiad.wmnet * 13:56 fnegri@cumin1003: START - Cookbook sre.hosts.remove-downtime for clouddb1025.eqiad.wmnet * 13:56 Lucas_WMDE: UTC afternoon backport+config window done * 13:53 moritzm: installing apr-util security updates * 13:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2002.codfw.wmnet * 13:50 fnegri@cumin1003: conftool action : set/weight=100; selector: name=clouddb1025.eqiad.wmnet * 13:49 fnegri@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1025.eqiad.wmnet * 13:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2002.codfw.wmnet * 13:41 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2001.codfw.wmnet * 13:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2001.codfw.wmnet * 13:41 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'. * 13:38 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'. * 13:38 fnegri@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on clouddb1025.eqiad.wmnet with reason: Removing s6 from clouddb1025 * 13:34 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2001.codfw.wmnet * 13:31 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'. * 13:29 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'. * 13:28 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet * 13:26 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host stat1009.eqiad.wmnet with OS bookworm * 13:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2001.codfw.wmnet * 13:24 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-staging-worker * 13:23 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1002.eqiad.wmnet * 13:21 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host stat1010.eqiad.wmnet with OS bookworm * 13:20 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1002.eqiad.wmnet * 13:20 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1001.eqiad.wmnet * 13:17 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1001.eqiad.wmnet * 13:16 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2001.codfw.wmnet * 13:13 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327511{{!}}UIC: Fix page:page instead of page:other in instrumentation]] (duration: 07m 00s) * 13:13 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2001.codfw.wmnet * 13:12 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2002.codfw.wmnet * 13:09 mszwarc@deploy1003: mszwarc: Continuing with deployment * 13:08 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1327511{{!}}UIC: Fix page:page instead of page:other in instrumentation]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2002.codfw.wmnet * 13:07 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2002.codfw.wmnet * 13:06 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1327511{{!}}UIC: Fix page:page instead of page:other in instrumentation]] * 13:04 jmm@dns1004: END - running authdns-update * 13:03 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2002.codfw.wmnet * 13:03 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2001.codfw.wmnet * 13:02 jmm@dns1004: START - running authdns-update * 13:01 cmooney@dns3003: END - running authdns-update * 13:00 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2001.codfw.wmnet * 12:59 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2001.codfw.wmnet * 12:59 cmooney@dns3003: START - running authdns-update * 12:57 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2001.codfw.wmnet * 12:56 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2002.codfw.wmnet * 12:55 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:55 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on drmrs<->eqiad GTT vpls - cmooney@cumin1003" * 12:54 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on drmrs<->eqiad GTT vpls - cmooney@cumin1003" * 12:54 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2002.codfw.wmnet * 12:54 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2003.codfw.wmnet * 12:50 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2003.codfw.wmnet * 12:49 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 12:48 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2001.codfw.wmnet * 12:46 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2001.codfw.wmnet * 12:46 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2002.codfw.wmnet * 12:43 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2002.codfw.wmnet * 12:43 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2003.codfw.wmnet * 12:42 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=1) for new host db1902.eqiad.wmnet * 12:42 fceratto@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host db1902.eqiad.wmnet with OS trixie * 12:41 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2003.codfw.wmnet * 12:40 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1003.eqiad.wmnet * 12:38 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1003.eqiad.wmnet * 12:37 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1002.eqiad.wmnet * 12:35 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1002.eqiad.wmnet * 12:35 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1001.eqiad.wmnet * 12:34 cmooney@dns3003: END - running authdns-update * 12:33 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1001.eqiad.wmnet * 12:32 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on stat1009.eqiad.wmnet with reason: host reimage * 12:31 cmooney@dns3003: START - running authdns-update * 12:31 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:31 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on drmrs<->eqiad cct - cmooney@cumin1003" * 12:28 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on drmrs<->eqiad cct - cmooney@cumin1003" * 12:28 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1902.eqiad.wmnet with reason: host reimage * 12:25 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on stat1009.eqiad.wmnet with reason: host reimage * 12:25 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 12:24 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on stat1010.eqiad.wmnet with reason: host reimage * 12:22 fceratto@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1902.eqiad.wmnet with reason: host reimage * 12:21 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on stat1010.eqiad.wmnet with reason: host reimage * 12:14 elukey: move the Docker Registry's /v2/dev/.* prefix to its dedicated S3 backend - [[phab:T432829|T432829]] * 12:12 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db1902.eqiad.wmnet with OS trixie * 12:09 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1902.eqiad.wmnet - fceratto@cumin1003" * 12:09 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1902.eqiad.wmnet - fceratto@cumin1003" * 12:09 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1902.eqiad.wmnet on all recursors * 12:09 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1902.eqiad.wmnet on all recursors * 12:09 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:08 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1902.eqiad.wmnet - fceratto@cumin1003" * 12:08 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1902.eqiad.wmnet - fceratto@cumin1003" * 12:08 tgr_: [[phab:T413390|T413390]] running CentralAuth:FixRenamedUserGlobalEditCount --wiki=metawiki --since=20250901000000 --until=20260301000000 --fix * 12:04 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1009.eqiad.wmnet with OS bookworm * 12:04 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1010.eqiad.wmnet with OS bookworm * 12:01 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 12:01 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1902.eqiad.wmnet * 12:00 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host stat1010.eqiad.wmnet with OS bookworm * 11:57 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1902.eqiad.wmnet * 11:57 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:57 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1902.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 11:57 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1902.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 11:51 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'. * 11:49 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'. * 11:48 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'. * 11:46 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'. * 11:39 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 11:37 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327518{{!}}Enable thumb.wikimedia.org on cswiki and fawiki (T427465)]] (duration: 10m 40s) * 11:35 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1902.eqiad.wmnet * 11:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1010.eqiad.wmnet with OS bookworm * 11:33 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 11:30 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327518{{!}}Enable thumb.wikimedia.org on cswiki and fawiki (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:26 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327518{{!}}Enable thumb.wikimedia.org on cswiki and fawiki (T427465)]] * 11:24 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host stat1010.eqiad.wmnet with OS bookworm * 11:07 fceratto@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host db1901.eqiad.wmnet * 11:07 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1901.eqiad.wmnet with OS trixie * 10:53 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1901.eqiad.wmnet with reason: host reimage * 10:47 tappof: bump space for prometheus k8s-aux in codfw * 10:47 tappof: bump space for prometheus k8s-dse in eqiad * 10:47 fceratto@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1901.eqiad.wmnet with reason: host reimage * 10:35 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db1901.eqiad.wmnet with OS trixie * 10:32 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:32 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:32 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1901.eqiad.wmnet on all recursors * 10:32 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1901.eqiad.wmnet on all recursors * 10:31 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:31 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:31 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:27 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:27 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1901.eqiad.wmnet * 10:23 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1010.eqiad.wmnet with OS bookworm * 10:20 fceratto@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts db1901.eqiad.wmnet * 10:20 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 10:18 blake@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 10:17 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:16 blake@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 10:13 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1901.eqiad.wmnet * 09:23 jelto@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'. * 09:22 jelto@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'. * 09:22 jelto: update cert-manager to 1.19.6 on wikikube staging-eqiad - [[phab:T427402|T427402]] * 09:20 moritzm: imported squid 7.6-2.1for trixie-wikimedia/main [[phab:T427282|T427282]] * 09:08 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 09:08 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 09:08 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 09:07 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 09:04 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 09:04 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:23 slyngshede@dns1004: END - running authdns-update * 08:21 slyngshede@dns1004: START - running authdns-update * 08:18 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.16 refs [[phab:T430835|T430835]] * 06:27 aokoth@dns1004: END - running authdns-update * 06:25 aokoth@dns1004: START - running authdns-update * 06:22 brennen@deploy1003: Finished deploy [phabricator/deployment@6b9b6ff]: deploy phab1005 for [[phab:T435087|T435087]] (duration: 00m 39s) * 06:21 brennen@deploy1003: Started deploy [phabricator/deployment@6b9b6ff]: deploy phab1005 for [[phab:T435087|T435087]] * 06:20 brennen@deploy1003: Finished deploy [phabricator/deployment@6b9b6ff]: deploy phab1004 for to pick up config values for [[phab:T435087|T435087]] (duration: 01m 46s) * 06:18 brennen@deploy1003: Started deploy [phabricator/deployment@6b9b6ff]: deploy phab1004 for to pick up config values for [[phab:T435087|T435087]] * 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 49s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-19 == * 23:19 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327197{{!}}Enable thumb.wikimedia.org on mediawiki.org (T427465)]] (duration: 10m 50s) * 23:18 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:16 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 23:15 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 23:10 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327197{{!}}Enable thumb.wikimedia.org on mediawiki.org (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:10 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:09 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 23:09 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:09 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 23:08 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327197{{!}}Enable thumb.wikimedia.org on mediawiki.org (T427465)]] * 22:58 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327201{{!}}Enable ReadingLists for all logged in users on test wiki (T435258)]] (duration: 11m 20s) * 22:50 jdlrobson@deploy1003: jdlrobson: Continuing with deployment * 22:49 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1327201{{!}}Enable ReadingLists for all logged in users on test wiki (T435258)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:46 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1327201{{!}}Enable ReadingLists for all logged in users on test wiki (T435258)]] * 22:42 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327162{{!}}Article: Split subjectpageheader by model and disable for wikitext]], [[gerrit:1327169{{!}}Make uppercase greek letters normal (non-italic) font (T434686 T434428)]], [[gerrit:1327176{{!}}Skin: Avoid DB lookup for pagecategorieslink message (T347123)]] (duration: 37m 52s) * 22:29 krinkle@deploy1003: krinkle: Continuing with deployment * 22:25 krinkle@deploy1003: krinkle: Backport for [[gerrit:1327162{{!}}Article: Split subjectpageheader by model and disable for wikitext]], [[gerrit:1327169{{!}}Make uppercase greek letters normal (non-italic) font (T434686 T434428)]], [[gerrit:1327176{{!}}Skin: Avoid DB lookup for pagecategorieslink message (T347123)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:04 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1327162{{!}}Article: Split subjectpageheader by model and disable for wikitext]], [[gerrit:1327169{{!}}Make uppercase greek letters normal (non-italic) font (T434686 T434428)]], [[gerrit:1327176{{!}}Skin: Avoid DB lookup for pagecategorieslink message (T347123)]] * 22:04 krinkle@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: awaiting CI (duration: 03m 06s) * 22:01 krinkle@deploy1003: Locking from deployment [ALL REPOSITORIES]: awaiting CI * 22:00 krinkle@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: awaiting CI (duration: 00m 01s) * 22:00 krinkle@deploy1003: Locking from deployment [ALL REPOSITORIES]: awaiting CI * 21:34 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2001.codfw.wmnet * 21:28 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2001.codfw.wmnet * 21:22 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327178{{!}}AccountRecovery: Notify the email address of the on file of the request (T425799)]] (duration: 47m 02s) * 21:13 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm * 21:09 catrope@deploy1003: catrope: Continuing with deployment * 20:55 catrope@deploy1003: catrope: Backport for [[gerrit:1327178{{!}}AccountRecovery: Notify the email address of the on file of the request (T425799)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:35 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1327178{{!}}AccountRecovery: Notify the email address of the on file of the request (T425799)]] * 20:31 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327128{{!}}Parsoid DataAccess: convert Parsoid fragment markers to/from strip tags (T432547)]] (duration: 07m 30s) * 20:27 catrope@deploy1003: catrope, arlolra: Continuing with deployment * 20:26 catrope@deploy1003: catrope, arlolra: Backport for [[gerrit:1327128{{!}}Parsoid DataAccess: convert Parsoid fragment markers to/from strip tags (T432547)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:24 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1327128{{!}}Parsoid DataAccess: convert Parsoid fragment markers to/from strip tags (T432547)]] * 20:23 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage * 20:17 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage * 20:15 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325878{{!}}[arwiki] Enable restricted user page editing and grant edit permissions (T434878)]] (duration: 08m 46s) * 20:11 catrope@deploy1003: catrope, gergesshamon: Continuing with deployment * 20:08 catrope@deploy1003: catrope, gergesshamon: Backport for [[gerrit:1325878{{!}}[arwiki] Enable restricted user page editing and grant edit permissions (T434878)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:06 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1325878{{!}}[arwiki] Enable restricted user page editing and grant edit permissions (T434878)]] * 19:59 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm * 19:56 eevans@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cassandra-dev2001.codfw.wmnet with OS bookworm * 19:56 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm * 19:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2207.codfw.wmnet with reason: Maintenance * 18:47 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319920{{!}}Allow setting a separate thumbUrl in production (T427465)]], [[gerrit:1327167{{!}}Fix wmgThumbUrl config (T427465)]] (duration: 18m 53s) * 18:43 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 18:30 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1319920{{!}}Allow setting a separate thumbUrl in production (T427465)]], [[gerrit:1327167{{!}}Fix wmgThumbUrl config (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:28 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1319920{{!}}Allow setting a separate thumbUrl in production (T427465)]], [[gerrit:1327167{{!}}Fix wmgThumbUrl config (T427465)]] * 18:26 sukhe@dns1004: END - running authdns-update * 18:24 sukhe@dns1004: START - running authdns-update * 18:09 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1319920{{!}}Allow setting a separate thumbUrl in production (T427465)]] * 18:03 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-eqiad and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 17:56 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-codfw and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 17:50 cmooney@dns3003: END - running authdns-update * 17:42 dancy@deploy1003: Installation of scap version "4.283.0" completed for 3 hosts * 17:41 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling reboot on A:durum and A:durum * 17:40 cmooney@dns3003: START - running authdns-update * 17:40 dancy@deploy1003: Installing scap version "4.283.0" for 3 host(s) * 17:38 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:37 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on GTT VPLS - cmooney@cumin1003" * 17:37 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --verbose --scoreLessThan=0.7 --exceptDatasetChecksums=[[phab:T434319|T434319]]-enwiki-models.txt # [[phab:T434319|T434319]] * 17:32 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on GTT VPLS - cmooney@cumin1003" * 17:29 sbassett: Deployed security fix for [[phab:T435210|T435210]] (wmf.16) * 17:26 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 17:22 sbassett: Deployed security fix for [[phab:T435210|T435210]] (wmf.15) * 17:00 sukhe@dns1004: END - running authdns-update * 16:58 sukhe@dns1004: START - running authdns-update * 16:53 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica-esams and A:liberica * 16:41 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica-esams and A:liberica * 16:41 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-eqiad and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 16:41 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-codfw and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 16:41 cjd91: sudo -i cookbook sre.cdn.roll-upgrade-ats --query 'A:cp-codfw' --task-id [[phab:T434478|T434478]] --reason '9.2.15 upgrade' * 16:41 cjd91: sudo -i cookbook sre.cdn.roll-upgrade-ats --query 'A:cp-eqiad' --task-id [[phab:T434478|T434478]] --reason '9.2.15 upgrade' * 16:40 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and A:durum * 16:28 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326881{{!}}Echo: Start using virtual domains (T380385)]] (duration: 13m 12s) * 16:23 urbanecm@deploy1003: urbanecm: Continuing with deployment * 16:21 urandom: Completed sessionstore Cassandra/JVM upgrade — [[phab:T435154|T435154]] * 16:21 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching sessionstore[2005-2006].codfw.wmnet,sessionstore[1005-1006].eqiad.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 16:19 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1326881{{!}}Echo: Start using virtual domains (T380385)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:15 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:15 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->codfw - cmooney@cumin1003" * 16:14 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1326881{{!}}Echo: Start using virtual domains (T380385)]] * 16:14 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327123{{!}}Revert^2 "Migrate database access to virtual domains" (T435305)]], [[gerrit:1327124{{!}}Pass the mapped domain of virtual-echo-shared to the push NameTableStores (T435305)]] (duration: 07m 42s) * 16:13 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching sessionstore[2005-2006].codfw.wmnet,sessionstore[1005-1006].eqiad.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 16:11 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->codfw - cmooney@cumin1003" * 16:10 urbanecm@deploy1003: urbanecm: Continuing with deployment * 16:08 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1327123{{!}}Revert^2 "Migrate database access to virtual domains" (T435305)]], [[gerrit:1327124{{!}}Pass the mapped domain of virtual-echo-shared to the push NameTableStores (T435305)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:06 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 16:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 16:06 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1327123{{!}}Revert^2 "Migrate database access to virtual domains" (T435305)]], [[gerrit:1327124{{!}}Pass the mapped domain of virtual-echo-shared to the push NameTableStores (T435305)]] * 16:06 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:03 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching sessionstore1004.eqiad.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 16:01 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching sessionstore1004.eqiad.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 15:56 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching sessionstore2004.codfw.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 15:54 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching sessionstore2004.codfw.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 15:53 urandom: beginning sessionstore Cassandra/JVM upgrade — [[phab:T435154|T435154]] * 15:52 urandom: beginning sessionstore Cassandra/JVM upgrade — [[phab:T432944|T432944]] * 15:51 cmooney@dns3003: END - running authdns-update * 15:49 cmooney@dns3003: START - running authdns-update * 15:48 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:48 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->codfw - cmooney@cumin1003" * 15:45 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->codfw - cmooney@cumin1003" * 15:44 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 15:44 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:43 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 15:42 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:42 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:38 cmooney@dns3003: END - running authdns-update * 15:36 cmooney@dns3003: START - running authdns-update * 15:36 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:36 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->eqsin - cmooney@cumin1003" * 15:34 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327138{{!}}Enable redis lock manager everywhere (T366938)]] (duration: 08m 36s) * 15:33 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->eqsin - cmooney@cumin1003" * 15:30 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:29 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 15:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1147.eqiad.wmnet with OS bookworm * 15:28 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327138{{!}}Enable redis lock manager everywhere (T366938)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:25 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327138{{!}}Enable redis lock manager everywhere (T366938)]] * 15:24 jmm@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host krb1002.eqiad.wmnet * 15:19 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2207.codfw.wmnet with reason: Host crashed * 15:17 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327098{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]], [[gerrit:1327101{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]] (duration: 07m 13s) * 15:12 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 15:12 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1327098{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]], [[gerrit:1327101{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:10 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1327098{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]], [[gerrit:1327101{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]] * 15:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2010.codfw.wmnet with OS trixie * 15:05 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 15:05 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1147.eqiad.wmnet with reason: host reimage * 14:59 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 14:59 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:58 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1147.eqiad.wmnet with reason: host reimage * 14:55 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:55 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:49 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 14:48 cmooney@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host durum1001.eqiad.wmnet * 14:46 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:46 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Delete db2902 ipv6 addr - fceratto@cumin1003" * 14:46 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Delete db2902 ipv6 addr - fceratto@cumin1003" * 14:43 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1147.eqiad.wmnet with OS bookworm * 14:42 tgr@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327094{{!}}SpecialMWOAuthListConsumers: Handle newFromMWUser returning null in addNavigationSubtitle (T435167)]] (duration: 19m 25s) * 14:42 cmooney@cumin1003: START - Cookbook sre.hosts.reboot-single for host durum1001.eqiad.wmnet * 14:42 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 14:41 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 14:38 tgr@deploy1003: tgr: Continuing with deployment * 14:36 tgr@deploy1003: tgr: Backport for [[gerrit:1327094{{!}}SpecialMWOAuthListConsumers: Handle newFromMWUser returning null in addNavigationSubtitle (T435167)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:28 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 14:28 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:28 cmooney@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host durum3005.esams.wmnet * 14:25 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 14:23 cmooney@cumin1003: START - Cookbook sre.hosts.reboot-single for host durum3005.esams.wmnet * 14:23 tgr@deploy1003: Started scap sync-world: Backport for [[gerrit:1327094{{!}}SpecialMWOAuthListConsumers: Handle newFromMWUser returning null in addNavigationSubtitle (T435167)]] * 14:18 gengh@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:17 topranks: disable puppet on hosts running BIRD BGP to test merge of patch to systemd healtchcheck service * 14:17 gengh@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:17 gengh@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:16 gengh@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:16 elukey: upgrade spicerack on cumin1003 and cumin2003 * 14:16 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:15 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314962{{!}}static: add new dir bimi/ for BIMI SVG and PEM file (T311685)]] (duration: 10m 00s) * 14:15 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:11 kharlan@deploy1003: kharlan, sukhe: Continuing with deployment * 14:11 gengh@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:09 gengh@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:09 gengh@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:08 kharlan@deploy1003: kharlan, sukhe: Backport for [[gerrit:1314962{{!}}static: add new dir bimi/ for BIMI SVG and PEM file (T311685)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:06 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 14:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:06 gengh@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:05 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1314962{{!}}static: add new dir bimi/ for BIMI SVG and PEM file (T311685)]] * 14:05 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host krb1002.eqiad.wmnet * 14:05 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:04 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:04 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326342{{!}}Revert^2 "wmf-config/ProductionServices: set URL for urldownloader to service record"]] (duration: 07m 40s) * 13:59 kharlan@deploy1003: kharlan, sukhe: Continuing with deployment * 13:58 kharlan@deploy1003: kharlan, sukhe: Backport for [[gerrit:1326342{{!}}Revert^2 "wmf-config/ProductionServices: set URL for urldownloader to service record"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:57 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host stat1008.eqiad.wmnet with OS bookworm * 13:56 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1326342{{!}}Revert^2 "wmf-config/ProductionServices: set URL for urldownloader to service record"]] * 13:56 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 13:54 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325532{{!}}srwiki: Allow bureaucrats to add and remove event-organizer group (T434748)]] (duration: 14m 56s) * 13:54 swfrench@dns1004: END - running authdns-update * 13:52 swfrench@dns1004: START - running authdns-update * 13:51 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:51 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 13:48 kharlan@deploy1003: kharlan, danielyepezgarces: Continuing with deployment * 13:45 swfrench@cumin2003: conftool action : set/pooled=yes; selector: name=wikikube-worker2330.codfw.wmnet * 13:44 swfrench-wmf: finished etcd-main codfw -> eqiad switchover - [[phab:T435103|T435103]] * 13:44 kharlan@deploy1003: kharlan, danielyepezgarces: Backport for [[gerrit:1325532{{!}}srwiki: Allow bureaucrats to add and remove event-organizer group (T434748)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:44 swfrench@cumin2003: conftool action : set/pooled=no; selector: name=wikikube-worker2330.codfw.wmnet * 13:41 swfrench@dns1004: END - running authdns-update * 13:39 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1325532{{!}}srwiki: Allow bureaucrats to add and remove event-organizer group (T434748)]] * 13:39 swfrench@dns1004: START - running authdns-update * 13:37 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327111{{!}}Special:AbuseReview: Add "no further action needed" review action (T435020)]], [[gerrit:1327110{{!}}AbuseReview: Take the review verdict as a REST path parameter (T435020)]] (duration: 31m 43s) * 13:31 swfrench-wmf: starting etcd-main codfw -> eqiad switchover - [[phab:T435103|T435103]] * 13:28 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:28 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 13:25 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=97) rolling reboot on A:durum and A:durum * 13:24 kharlan@deploy1003: kharlan: Continuing with deployment * 13:24 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1147 * 13:24 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1147 * 13:23 kharlan@deploy1003: kharlan: Backport for [[gerrit:1327111{{!}}Special:AbuseReview: Add "no further action needed" review action (T435020)]], [[gerrit:1327110{{!}}AbuseReview: Take the review verdict as a REST path parameter (T435020)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:16 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:16 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 13:12 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:12 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 13:10 cdobbins@cumin1003: END (ERROR) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=97) Rolling upgrade of ATS on A:cp-codfw and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 13:10 cdobbins@cumin1003: END (ERROR) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=97) Rolling upgrade of ATS on A:cp-eqiad and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 13:06 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1327111{{!}}Special:AbuseReview: Add "no further action needed" review action (T435020)]], [[gerrit:1327110{{!}}AbuseReview: Take the review verdict as a REST path parameter (T435020)]] * 13:06 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-eqiad and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 13:05 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-codfw and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 13:05 cjd91: sudo -i cookbook sre.cdn.roll-upgrade-ats --query 'A:cp-eqiad' --task-id [[phab:T434478|T434478]] --reason '9.2.15 upgrade' * 13:03 swfrench@dns1004: END - running authdns-update * 13:01 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:01 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 13:01 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host ncmonitor1001.eqiad.wmnet * 13:00 swfrench@dns1004: START - running authdns-update * 12:59 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 12:59 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 12:59 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 12:59 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 12:57 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and A:durum * 12:53 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on db2902.codfw.wmnet with reason: Cloning * 12:48 cmooney@dns3003: END - running authdns-update * 12:48 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327096{{!}}Switch to redis lock manager on s4 and s8 (T366938)]] (duration: 09m 11s) * 12:46 cmooney@dns3003: START - running authdns-update * 12:45 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:45 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->eqord cct - cmooney@cumin1003" * 12:45 elukey: move the /v2/releng.* prefix on the Docker Registry to its new s3 backend - [[phab:T432829|T432829]] * 12:43 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 12:42 jelto@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 12:42 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->eqord cct - cmooney@cumin1003" * 12:42 jelto@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 12:41 jelto: update cert-manager to 1.19.6 on wikikube staging-codfw - [[phab:T427402|T427402]] * 12:40 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327096{{!}}Switch to redis lock manager on s4 and s8 (T366938)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:38 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327096{{!}}Switch to redis lock manager on s4 and s8 (T366938)]] * 12:38 blake@deploy1003: Finished scap sync-world: non-build deploy for [[phab:T417800|T417800]] (duration: 03m 52s) * 12:36 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 12:35 blake@deploy1003: Started scap sync-world: non-build deploy for [[phab:T417800|T417800]] * 12:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host krb2002.codfw.wmnet * 11:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host krb2002.codfw.wmnet * 11:49 moritzm: installing kerberos security updates * 11:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on stat1008.eqiad.wmnet with reason: host reimage * 11:44 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on stat1008.eqiad.wmnet with reason: host reimage * 11:31 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327084{{!}}Revert "Migrate database access to virtual domains" (T435305)]] (duration: 11m 02s) * 11:29 gkyziridis@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:29 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:29 kart_: Updated MinT to 2026-06-04-131507-production ([[phab:T321316|T321316]]) * 11:28 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/machinetranslation: apply * 11:28 gkyziridis@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:26 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 11:24 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:24 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:23 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:23 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:23 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/machinetranslation: apply * 11:22 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327084{{!}}Revert "Migrate database access to virtual domains" (T435305)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:21 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/machinetranslation: apply * 11:21 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:21 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:20 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327084{{!}}Revert "Migrate database access to virtual domains" (T435305)]] * 11:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1008.eqiad.wmnet with OS bookworm * 11:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet * 11:17 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/machinetranslation: apply * 11:13 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/machinetranslation: apply * 11:12 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:12 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:10 kartik@deploy1003: helmfile [staging] START helmfile.d/services/machinetranslation: apply * 11:08 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-cron: apply * 11:08 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/mw-cron: apply * 11:08 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply * 11:08 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply * 11:07 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet * 11:07 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host stat1008.eqiad.wmnet with OS bookworm * 11:06 moritzm: upgrading the new trixie URL downloaders to Squid 7.6 [[phab:T427282|T427282]] * 11:01 gkyziridis@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin2002.codfw.wmnet * 10:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin2002.codfw.wmnet * 10:45 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1169.eqiad.wmnet with OS bookworm * 10:42 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1185.eqiad.wmnet with OS bookworm * 10:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1169.eqiad.wmnet with reason: host reimage * 10:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1185.eqiad.wmnet with reason: host reimage * 10:14 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1169.eqiad.wmnet with reason: host reimage * 10:14 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1185.eqiad.wmnet with reason: host reimage * 10:11 jmm@cumin2003: END (PASS) - Cookbook sre.netbox.restart-reboot (exit_code=0) rolling reboot on A:netbox * 10:06 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1008.eqiad.wmnet with OS bookworm * 10:04 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 10:04 mpostoronca@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321579{{!}}Register the mediawiki.wikimedia_antiabuse.content_policy_score stream (T432848)]] (duration: 08m 53s) * 10:03 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 10:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 10:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 10:00 mpostoronca@deploy1003: mpostoronca: Continuing with deployment * 09:59 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1185.eqiad.wmnet with OS bookworm * 09:59 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1169.eqiad.wmnet with OS bookworm * 09:59 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.convert-disks (exit_code=0) for host ms-be1065 * 09:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1065.eqiad.wmnet with OS trixie * 09:58 mpostoronca@deploy1003: mpostoronca: Backport for [[gerrit:1321579{{!}}Register the mediawiki.wikimedia_antiabuse.content_policy_score stream (T432848)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:55 jmm@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netbox.discovery.wmnet. on all recursors * 09:55 mpostoronca@deploy1003: Started scap sync-world: Backport for [[gerrit:1321579{{!}}Register the mediawiki.wikimedia_antiabuse.content_policy_score stream (T432848)]] * 09:55 jmm@cumin2003: START - Cookbook sre.dns.wipe-cache netbox.discovery.wmnet. on all recursors * 09:52 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw2001.wikimedia.org with OS trixie * 09:51 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 09:51 jmm@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netbox.discovery.wmnet. on all recursors * 09:51 jmm@cumin2003: START - Cookbook sre.dns.wipe-cache netbox.discovery.wmnet. on all recursors * 09:51 jmm@cumin2003: START - Cookbook sre.netbox.restart-reboot rolling reboot on A:netbox * 09:46 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 09:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1056.eqiad.wmnet with OS trixie * 09:44 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 09:39 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 09:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 09:36 topranks: make HE transport circuits from magru live * 09:36 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 09:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb1003.eqiad.wmnet * 09:33 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage * 09:31 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb1003.eqiad.wmnet * 09:30 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 09:28 arnaudb@dns1006: END - running authdns-update * 09:27 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage * 09:27 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb2003.codfw.wmnet * 09:26 arnaudb@dns1006: START - running authdns-update * 09:24 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1056.eqiad.wmnet with reason: host reimage * 09:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb2003.codfw.wmnet * 09:20 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.convert-disks (exit_code=0) for host ms-be1068 * 09:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1068.eqiad.wmnet with OS trixie * 09:20 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 09:19 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "cloudvirt1057 - filippo@cumin1003" * 09:19 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "cloudvirt1057 - filippo@cumin1003" * 09:18 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1056.eqiad.wmnet with reason: host reimage * 09:18 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1057.eqiad.wmnet with OS trixie * 09:18 filippo@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 09:18 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 09:15 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 09:14 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 09:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host irc1003.wikimedia.org * 09:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:13 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1065.eqiad.wmnet with OS trixie * 09:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 09:09 moritzm: installing Postgresql security updates * 09:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host irc1003.wikimedia.org * 09:07 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 09:07 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:07 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1173.eqiad.wmnet with OS bookworm * 09:06 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw2001.wikimedia.org with OS trixie * 09:03 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.convert-disks (exit_code=0) for host ms-be1064 * 09:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1064.eqiad.wmnet with OS trixie * 09:03 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1056.eqiad.wmnet with OS trixie * 09:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1057.eqiad.wmnet with reason: host reimage * 09:01 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 08:58 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 08:56 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1057.eqiad.wmnet with reason: host reimage * 08:54 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1208.eqiad.wmnet with OS bookworm * 08:53 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2902.codfw.wmnet with OS trixie * 08:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 08:50 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1174.eqiad.wmnet with OS bookworm * 08:46 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1172.eqiad.wmnet with OS bookworm * 08:45 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1173.eqiad.wmnet with reason: host reimage * 08:41 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 08:40 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1057.eqiad.wmnet with OS trixie * 08:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1057.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 08:39 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1207.eqiad.wmnet with OS bookworm * 08:38 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2902.codfw.wmnet with reason: host reimage * 08:37 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 08:34 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1068.eqiad.wmnet with OS trixie * 08:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1208.eqiad.wmnet with reason: host reimage * 08:31 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1057.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 08:29 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 08:28 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1222.eqiad.wmnet onto db1276.eqiad.wmnet * 08:28 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1222: Pool db1222.eqiad.wmnet in after cloning * 08:28 fceratto@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2902.codfw.wmnet with reason: host reimage * 08:27 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1055.eqiad.wmnet with OS trixie * 08:27 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 08:26 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1174.eqiad.wmnet with reason: host reimage * 08:25 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 08:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-misc2002.codfw.wmnet * 08:23 topranks: reboot pfw1-codfw firewall pair to upgrade JunOS [[phab:T434865|T434865]] * 08:22 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1172.eqiad.wmnet with reason: host reimage * 08:20 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1064.eqiad.wmnet with OS trixie * 08:20 mvernon@cumin2003: START - Cookbook sre.swift.convert-disks for host ms-be1065 * 08:18 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1207.eqiad.wmnet with reason: host reimage * 08:17 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1174.eqiad.wmnet with reason: host reimage * 08:17 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1173.eqiad.wmnet with reason: host reimage * 08:17 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1172.eqiad.wmnet with reason: host reimage * 08:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host mc-misc2002.codfw.wmnet * 08:15 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1208.eqiad.wmnet with reason: host reimage * 08:15 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1207.eqiad.wmnet with reason: host reimage * 08:14 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db2902.codfw.wmnet with OS trixie * 08:14 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.16 refs [[phab:T430835|T430835]] * 08:13 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db2902.codfw.wmnet * 08:13 fceratto@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host db2902.codfw.wmnet with OS trixie * 08:10 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr[1-2]-codfw with reason: upgrade pfw1a-codfw and pfw1b-codfw pair * 08:09 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1055.eqiad.wmnet with reason: host reimage * 08:07 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on pfw1-codfw with reason: upgrade pfw1a-codfw and pfw1b-codfw pair * 08:03 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1055.eqiad.wmnet with reason: host reimage * 08:02 arnaudb@dns1006: END - running authdns-update * 08:02 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1208.eqiad.wmnet with OS bookworm * 08:02 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1207.eqiad.wmnet with OS bookworm * 08:01 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1174.eqiad.wmnet with OS bookworm * 08:01 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1173.eqiad.wmnet with OS bookworm * 08:01 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1172.eqiad.wmnet with OS bookworm * 07:59 arnaudb@dns1006: START - running authdns-update * 07:58 arnaudb@dns1006: START - running authdns-update * 07:48 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1055.eqiad.wmnet with OS trixie * 07:45 moritzm: extend the disk of ldap-rw2001 by 80G [[phab:T331699|T331699]] * 07:42 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1222: Pool db1222.eqiad.wmnet in after cloning * 07:36 mvernon@cumin2003: START - Cookbook sre.swift.convert-disks for host ms-be1068 * 07:35 mvernon@cumin2003: START - Cookbook sre.swift.convert-disks for host ms-be1064 * 07:17 moritzm: installing imagemagick security updates * 07:14 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1277: Pool back * 07:14 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon1003.wikimedia.org * 07:07 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon1003.wikimedia.org * 07:03 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1280: Pool back * 07:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon2002.wikimedia.org * 06:55 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon2002.wikimedia.org * 06:54 moritzm: installing php8.2 security updates * 06:51 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1284: Pool back * 06:49 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1222: Depool db1222.eqiad.wmnet to then clone it to db1276.eqiad.wmnet - marostegui@cumin1003 * 06:49 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1222: Depool db1222.eqiad.wmnet to then clone it to db1276.eqiad.wmnet - marostegui@cumin1003 * 06:49 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1222.eqiad.wmnet onto db1276.eqiad.wmnet * 06:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd1005.eqiad.wmnet * 06:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd1005.eqiad.wmnet * 06:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd1004.eqiad.wmnet * 06:36 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2209: db2209 repool * 06:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd1004.eqiad.wmnet * 06:32 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast3007.wikimedia.org * 06:29 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1277: Pool back * 06:28 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1277 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96190 and previous config saved to /var/cache/conftool/dbconfig/20260819-062815-marostegui.json * 06:26 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast3007.wikimedia.org * 06:22 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul1001.eqiad.wmnet * 06:18 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul1001.eqiad.wmnet * 06:18 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul1003.eqiad.wmnet * 06:18 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1280: Pool back * 06:17 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1284 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96186 and previous config saved to /var/cache/conftool/dbconfig/20260819-061743-marostegui.json * 06:14 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul1003.eqiad.wmnet * 06:14 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul1002.eqiad.wmnet * 06:10 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul1002.eqiad.wmnet * 06:10 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2048.codfw.wmnet * 06:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2048.codfw.wmnet * 06:06 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1284: Pool back * 06:06 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1284 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96184 and previous config saved to /var/cache/conftool/dbconfig/20260819-060621-marostegui.json * 06:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2048.codfw.wmnet * 05:59 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2048.codfw.wmnet * 05:51 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2209: db2209 repool * 03:16 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1186.eqiad.wmnet with OS bookworm * 02:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1186.eqiad.wmnet with reason: host reimage * 02:46 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1186.eqiad.wmnet with reason: host reimage * 02:46 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2207 [[phab:T435270|T435270]]', diff saved to https://phabricator.wikimedia.org/P96181 and previous config saved to /var/cache/conftool/dbconfig/20260819-024627-marostegui.json * 02:44 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2204 to s2 primary [[phab:T435270|T435270]]', diff saved to https://phabricator.wikimedia.org/P96180 and previous config saved to /var/cache/conftool/dbconfig/20260819-024403-marostegui.json * 02:43 marostegui: Starting s2 codfw failover from db2207 to db2204 - [[phab:T435270|T435270]] * 02:39 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2204 with weight 0 [[phab:T435270|T435270]]', diff saved to https://phabricator.wikimedia.org/P96179 and previous config saved to /var/cache/conftool/dbconfig/20260819-023951-marostegui.json * 02:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s2 [[phab:T435270|T435270]] * 02:32 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1186.eqiad.wmnet with OS bookworm * 02:29 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-worker1186.eqiad.wmnet with OS bookworm * 02:18 denisse@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2207: Depooling replica * 02:18 denisse@cumin1003: START - Cookbook sre.mysql.depool depool db2207: Depooling replica * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 48s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-18 == * 23:55 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1264845{{!}}Remove unused/redundant wgMFNoindexPages=true setting (T255458)]] (duration: 09m 42s) * 23:51 krinkle@deploy1003: krinkle: Continuing with deployment * 23:48 krinkle@deploy1003: krinkle: Backport for [[gerrit:1264845{{!}}Remove unused/redundant wgMFNoindexPages=true setting (T255458)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:45 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1264845{{!}}Remove unused/redundant wgMFNoindexPages=true setting (T255458)]] * 23:38 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326944{{!}}Retire filebackend lock manager in favour of the default one (T366938)]] (duration: 08m 55s) * 23:34 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 23:31 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326944{{!}}Retire filebackend lock manager in favour of the default one (T366938)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:29 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326944{{!}}Retire filebackend lock manager in favour of the default one (T366938)]] * 23:27 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1170.eqiad.wmnet with OS bookworm * 23:21 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1205.eqiad.wmnet with OS bookworm * 23:20 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1171.eqiad.wmnet with OS bookworm * 23:15 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1206.eqiad.wmnet with OS bookworm * 23:05 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1170.eqiad.wmnet with reason: host reimage * 23:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1205.eqiad.wmnet with reason: host reimage * 22:57 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1171.eqiad.wmnet with reason: host reimage * 22:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1206.eqiad.wmnet with reason: host reimage * 22:53 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1205.eqiad.wmnet with reason: host reimage * 22:51 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1171.eqiad.wmnet with reason: host reimage * 22:51 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1170.eqiad.wmnet with reason: host reimage * 22:50 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1206.eqiad.wmnet with reason: host reimage * 22:36 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1206.eqiad.wmnet with OS bookworm * 22:35 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1205.eqiad.wmnet with OS bookworm * 22:35 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1186.eqiad.wmnet with OS bookworm * 22:35 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1171.eqiad.wmnet with OS bookworm * 22:35 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1170.eqiad.wmnet with OS bookworm * 22:33 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-worker1194.eqiad.wmnet with OS bookworm * 22:22 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326923{{!}}Enable redis lock manager on s6 (T366938)]] (duration: 11m 52s) * 22:18 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 22:13 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326923{{!}}Enable redis lock manager on s6 (T366938)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:10 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326923{{!}}Enable redis lock manager on s6 (T366938)]] * 22:04 sbassett: Deployed security fix for [[phab:T435234|T435234]] (wmf.16) * 21:54 sbassett: Deployed security fix for [[phab:T435234|T435234]] (wmf.15) * 21:38 caro@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326925{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326926{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326929{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]], [[gerrit:1326928{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]] (duration: 0 * 21:34 caro@deploy1003: caro: Continuing with deployment * 21:33 caro@deploy1003: caro: Backport for [[gerrit:1326925{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326926{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326929{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]], [[gerrit:1326928{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]] synced to the testservers (see h * 21:31 caro@deploy1003: Started scap sync-world: Backport for [[gerrit:1326925{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326926{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326929{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]], [[gerrit:1326928{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]] * 21:24 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1204.eqiad.wmnet with reason: 1204 datanode repair [[phab:T434494|T434494]] * 21:02 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326896{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]], [[gerrit:1326897{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]] (duration: 13m 56s) * 20:58 krinkle@deploy1003: krinkle: Continuing with deployment * 20:50 krinkle@deploy1003: krinkle: Backport for [[gerrit:1326896{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]], [[gerrit:1326897{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:49 ryankemper: `an-launcher1003` terminated process group `666809` (`rest_backfill_phase1.sh`) ~20 mins ago with `sudo kill -TERM -- -666809` after its local spark driver (`--driver-memory 64g`) repeatedly exhausted memory on the 32 GB VM and caused SSH to intermittently flap; host recovered to 27 GB available memory * 20:48 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1326896{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]], [[gerrit:1326897{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]] * 20:35 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 20:33 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 20:31 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 20:31 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326870{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]], [[gerrit:1326871{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]] (duration: 07m 35s) * 20:28 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 20:26 kemayo@deploy1003: kemayo: Continuing with deployment * 20:26 ryankemper: `an-launcher1003` confirmed the host is flapping because of memory thrash. chasing down the source of the thrash * 20:25 kemayo@deploy1003: kemayo: Backport for [[gerrit:1326870{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]], [[gerrit:1326871{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:23 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1326870{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]], [[gerrit:1326871{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]] * 20:23 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 20:20 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 20:14 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1168.eqiad.wmnet with OS bookworm * 20:08 zabe: zabe@deploy1003:~$ mwscript extensions/WikimediaMaintenance/maintenance/fixFileRevisionArchiveNameDrift.php enwiki # [[phab:T428406|T428406]] * 20:08 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1204.eqiad.wmnet with OS bookworm * 20:05 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1167.eqiad.wmnet with OS bookworm * 20:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1166.eqiad.wmnet with OS bookworm * 19:54 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1203.eqiad.wmnet with OS bookworm * 19:53 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326907{{!}}Revert "Disable redis lock manager on testwiki"]] (duration: 11m 05s) * 19:50 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1168.eqiad.wmnet with reason: host reimage * 19:47 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1204.eqiad.wmnet with reason: host reimage * 19:46 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 19:44 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326907{{!}}Revert "Disable redis lock manager on testwiki"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:42 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326907{{!}}Revert "Disable redis lock manager on testwiki"]] * 19:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1167.eqiad.wmnet with reason: host reimage * 19:37 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1166.eqiad.wmnet with reason: host reimage * 19:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1203.eqiad.wmnet with reason: host reimage * 19:32 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1167.eqiad.wmnet with reason: host reimage * 19:32 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1168.eqiad.wmnet with reason: host reimage * 19:32 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1166.eqiad.wmnet with reason: host reimage * 19:31 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1204.eqiad.wmnet with reason: host reimage * 19:31 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1203.eqiad.wmnet with reason: host reimage * 19:26 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324751{{!}}InitialiseSettings: Enable 2FA enforcement on more private wikis (T428103)]], [[gerrit:1326875{{!}}Add banner notifying of upcoming 2FA enforcement (T420792)]] (duration: 31m 46s) * 19:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1204.eqiad.wmnet with OS bookworm * 19:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1203.eqiad.wmnet with OS bookworm * 19:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1168.eqiad.wmnet with OS bookworm * 19:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1167.eqiad.wmnet with OS bookworm * 19:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1166.eqiad.wmnet with OS bookworm * 19:15 denisse: rebooting kafkamon2003.codfw.wmnet - [[phab:T435162|T435162]] * 19:14 denisse: rebooting kafkamon1003.eqiad.wmnet [[phab:T435162|T435162]] * 19:13 reedy@deploy1003: reedy: Continuing with deployment * 19:12 reedy@deploy1003: reedy: Backport for [[gerrit:1324751{{!}}InitialiseSettings: Enable 2FA enforcement on more private wikis (T428103)]], [[gerrit:1326875{{!}}Add banner notifying of upcoming 2FA enforcement (T420792)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:54 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324751{{!}}InitialiseSettings: Enable 2FA enforcement on more private wikis (T428103)]], [[gerrit:1326875{{!}}Add banner notifying of upcoming 2FA enforcement (T420792)]] * 18:50 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 18:44 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.16 refs [[phab:T430835|T430835]] * 18:34 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 18:31 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 18:22 aklapper@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326852{{!}}CategoryTree: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]], [[gerrit:1326853{{!}}CategoryViewer: Allow null $html in the CategoryViewerGenerateLink hook (T435161)]], [[gerrit:1326865{{!}}Flow: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]] (duration: 09m 57s) * 18:18 aklapper@deploy1003: jforrester, aklapper: Continuing with deployment * 18:17 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 18:14 aklapper@deploy1003: jforrester, aklapper: Backport for [[gerrit:1326852{{!}}CategoryTree: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]], [[gerrit:1326853{{!}}CategoryViewer: Allow null $html in the CategoryViewerGenerateLink hook (T435161)]], [[gerrit:1326865{{!}}Flow: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki * 18:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1156.eqiad.wmnet with OS bookworm * 18:12 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 18:12 aklapper@deploy1003: Started scap sync-world: Backport for [[gerrit:1326852{{!}}CategoryTree: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]], [[gerrit:1326853{{!}}CategoryViewer: Allow null $html in the CategoryViewerGenerateLink hook (T435161)]], [[gerrit:1326865{{!}}Flow: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]] * 18:11 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 18:08 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1146.eqiad.wmnet with OS bookworm * 18:07 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1177.eqiad.wmnet with OS bookworm * 18:00 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326380{{!}}Introduce main lock manager service (T366938 T427999)]] (duration: 11m 25s) * 17:58 ladsgroup@deploy1003: ladsgroup: Rolling back deployment * 17:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1202.eqiad.wmnet with OS bookworm * 17:55 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1201.eqiad.wmnet with OS bookworm * 17:50 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326380{{!}}Introduce main lock manager service (T366938 T427999)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:48 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326380{{!}}Introduce main lock manager service (T366938 T427999)]] * 17:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1156.eqiad.wmnet with reason: host reimage * 17:46 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1146.eqiad.wmnet with reason: host reimage * 17:45 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-esams and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 17:42 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 17:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1177.eqiad.wmnet with reason: host reimage * 17:38 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1202.eqiad.wmnet with reason: host reimage * 17:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1201.eqiad.wmnet with reason: host reimage * 17:30 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1156.eqiad.wmnet with reason: host reimage * 17:29 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1177.eqiad.wmnet with reason: host reimage * 17:28 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1146.eqiad.wmnet with reason: host reimage * 17:28 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1202.eqiad.wmnet with reason: host reimage * 17:27 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1201.eqiad.wmnet with reason: host reimage * 17:25 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 17:21 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 17:14 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1202.eqiad.wmnet with OS bookworm * 17:14 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1201.eqiad.wmnet with OS bookworm * 17:14 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1177.eqiad.wmnet with OS bookworm * 17:14 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1156.eqiad.wmnet with OS bookworm * 17:14 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1146.eqiad.wmnet with OS bookworm * 17:09 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326885{{!}}w/deployment-info.php: Handle new file format (T434726)]] (duration: 07m 15s) * 17:08 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 17:05 dancy@deploy1003: dancy: Continuing with deployment * 17:04 dancy@deploy1003: dancy: Backport for [[gerrit:1326885{{!}}w/deployment-info.php: Handle new file format (T434726)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:02 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1326885{{!}}w/deployment-info.php: Handle new file format (T434726)]] * 16:50 dancy@deploy1003: Finished scap sync-world: Testing [[phab:T434726|T434726]] (duration: 06m 40s) * 16:43 dancy@deploy1003: Started scap sync-world: Testing [[phab:T434726|T434726]] * 16:43 dancy@deploy1003: Installation of scap version "4.282.0" completed for 3 hosts * 16:41 dancy@deploy1003: Installing scap version "4.282.0" for 3 host(s) * 16:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1145.eqiad.wmnet with OS bookworm * 16:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1200.eqiad.wmnet with OS bookworm * 16:32 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1199.eqiad.wmnet with OS bookworm * 16:18 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1145.eqiad.wmnet with reason: host reimage * 16:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1200.eqiad.wmnet with reason: host reimage * 16:09 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1199.eqiad.wmnet with reason: host reimage * 16:05 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-esams and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 16:05 cjd91: sudo -i cookbook sre.cdn.roll-upgrade-ats --query 'A:cp-esams' --task-id [[phab:T434478|T434478]] --reason '9.2.15 upgrade' * 16:03 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1200.eqiad.wmnet with reason: host reimage * 16:02 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1145.eqiad.wmnet with reason: host reimage * 16:02 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1199.eqiad.wmnet with reason: host reimage * 15:50 moritzm: installing zip security updates * 15:48 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1200.eqiad.wmnet with OS bookworm * 15:47 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1199.eqiad.wmnet with OS bookworm * 15:47 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1145.eqiad.wmnet with OS bookworm * 15:41 topranks: bounce PIC 0/0 on cr1-magru to set port to 40G * 15:37 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1236.eqiad.wmnet with OS bookworm * 15:29 aikochou@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'ores-legacy' for release 'main' . * 15:26 aikochou@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'ores-legacy' for release 'main' . * 15:20 aikochou@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'ores-legacy' for release 'main' . * 15:14 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2002.codfw.wmnet * 15:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1236.eqiad.wmnet with reason: host reimage * 15:12 moritzm: failover ganeti master in codfw to ganeti2047 * 15:09 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1236.eqiad.wmnet with reason: host reimage * 15:09 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2044.codfw.wmnet * 15:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2002.codfw.wmnet * 15:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2044.codfw.wmnet * 15:04 brennen@deploy1003: Finished deploy [phabricator/deployment@6b9b6ff]: deploy phab1004 for [[phab:T435213|T435213]] (duration: 01m 01s) * 15:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2044.codfw.wmnet * 15:03 brennen@deploy1003: Started deploy [phabricator/deployment@6b9b6ff]: deploy phab1004 for [[phab:T435213|T435213]] * 15:03 brennen@deploy1003: Finished deploy [phabricator/deployment@6b9b6ff]: deploy phab2003 for [[phab:T435213|T435213]] (duration: 00m 57s) * 15:02 brennen@deploy1003: Started deploy [phabricator/deployment@6b9b6ff]: deploy phab2003 for [[phab:T435213|T435213]] * 14:57 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply * 14:57 arnaudb@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on phab2003.codfw.wmnet,phab[1004-1006].eqiad.wmnet with reason: maintenance * 14:56 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2044.codfw.wmnet * 14:55 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply * 14:53 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1236.eqiad.wmnet with OS bookworm * 14:51 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2043.codfw.wmnet * 14:51 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2043.codfw.wmnet * 14:45 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2043.codfw.wmnet * 14:33 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2043.codfw.wmnet * 14:22 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2042.codfw.wmnet * 14:22 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2042.codfw.wmnet * 14:21 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db2902.codfw.wmnet with OS trixie * 14:16 elukey: uploaded spicerack_13.2.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia * 14:16 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2042.codfw.wmnet * 14:04 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2042.codfw.wmnet * 14:04 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2041.codfw.wmnet * 14:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2041.codfw.wmnet * 13:59 phuedx@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: apply * 13:59 phuedx@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-main: apply * 13:59 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326838{{!}}Enable Suggested Investigations on hewiki (T435146)]] (duration: 11m 50s) * 13:58 phuedx@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: apply * 13:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2041.codfw.wmnet * 13:57 phuedx@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-main: apply * 13:57 phuedx@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-main: apply * 13:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard1003.eqiad.wmnet * 13:57 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-main: apply * 13:55 phuedx@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-logging-external: apply * 13:55 phuedx@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-logging-external: apply * 13:54 phuedx@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-logging-external: apply * 13:54 stran@deploy1003: stran: Continuing with deployment * 13:54 phuedx@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-logging-external: apply * 13:54 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-logging-external: apply * 13:54 phuedx@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-logging-external: apply * 13:53 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-logging-external: apply * 13:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard1003.eqiad.wmnet * 13:53 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2041.codfw.wmnet * 13:51 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db2902.codfw.wmnet - fceratto@cumin1003" * 13:51 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db2902.codfw.wmnet - fceratto@cumin1003" * 13:51 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2026.codfw.wmnet * 13:51 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard2003.codfw.wmnet * 13:50 phuedx@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: apply * 13:50 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-eqiad * 13:50 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp1001.eqiad.wmnet * 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp1001.eqiad.wmnet * 13:50 phuedx@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: apply * 13:49 phuedx@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: apply * 13:49 stran@deploy1003: stran: Backport for [[gerrit:1326838{{!}}Enable Suggested Investigations on hewiki (T435146)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:48 phuedx@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: apply * 13:48 phuedx@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics: apply * 13:47 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard2003.codfw.wmnet * 13:47 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics: apply * 13:47 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1326838{{!}}Enable Suggested Investigations on hewiki (T435146)]] * 13:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp1001.eqiad.wmnet * 13:44 cdanis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 13:43 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp1001.eqiad.wmnet * 13:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1315-1327].eqiad.wmnet * 13:43 cdanis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 13:43 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1315-1327].eqiad.wmnet * 13:40 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki1001.eqiad.wmnet * 13:35 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1315-1327].eqiad.wmnet * 13:34 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host rpki1001.eqiad.wmnet * 13:27 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1315-1327].eqiad.wmnet * 13:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1302-1314].eqiad.wmnet * 13:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1302-1314].eqiad.wmnet * 13:21 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326775{{!}}SI: Instrument case update on first edit (T435048)]] (duration: 07m 12s) * 13:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1302-1314].eqiad.wmnet * 13:17 stran@deploy1003: stran: Continuing with deployment * 13:16 stran@deploy1003: stran: Backport for [[gerrit:1326775{{!}}SI: Instrument case update on first edit (T435048)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:14 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1326775{{!}}SI: Instrument case update on first edit (T435048)]] * 13:13 moritzm: installing util-linux security updates * 13:11 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1302-1314].eqiad.wmnet * 13:10 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326232{{!}}prv: Enable parsoid rendering for 5 wikis (T435115)]] (duration: 08m 17s) * 13:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1288-1289,1291-1301].eqiad.wmnet * 13:10 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply * 13:10 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1288-1289,1291-1301].eqiad.wmnet * 13:10 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply * 13:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki2003.codfw.wmnet * 13:06 jgiannelos@deploy1003: jgiannelos: Continuing with deployment * 13:05 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host rpki2003.codfw.wmnet * 13:04 jgiannelos@deploy1003: jgiannelos: Backport for [[gerrit:1326232{{!}}prv: Enable parsoid rendering for 5 wikis (T435115)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1288-1289,1291-1301].eqiad.wmnet * 13:02 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1326232{{!}}prv: Enable parsoid rendering for 5 wikis (T435115)]] * 12:53 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1288-1289,1291-1301].eqiad.wmnet * 12:53 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1273,1275-1287].eqiad.wmnet * 12:53 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1273,1275-1287].eqiad.wmnet * 12:52 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt1002.wikimedia.org * 12:51 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db2902.codfw.wmnet on all recursors * 12:51 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db2902.codfw.wmnet on all recursors * 12:51 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:51 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db2902.codfw.wmnet - fceratto@cumin1003" * 12:51 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db2902.codfw.wmnet - fceratto@cumin1003" * 12:46 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt1002.wikimedia.org * 12:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt2002.wikimedia.org * 12:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1273,1275-1287].eqiad.wmnet * 12:42 dhinus: repooled clouddb1032 that was currently <nowiki>{</nowiki>"weight": 0, "pooled": "inactive"<nowiki>}</nowiki> for both s4 and s6 * 12:41 dhinus: also depooled clouddb1017 (forgot it in the previous list) * 12:41 fnegri@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet * 12:40 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 12:40 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db2902.codfw.wmnet * 12:40 fnegri@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032.eqiad.wmnet * 12:40 fnegri@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032 * 12:39 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1017.eqiad.wmnet * 12:39 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt2002.wikimedia.org * 12:38 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1020.eqiad.wmnet * 12:38 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1018.eqiad.wmnet * 12:38 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet * 12:37 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1014.eqiad.wmnet * 12:37 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1013.eqiad.wmnet * 12:37 dhinus: depool again clouddb10[13,14,16,18,20] that were repooled by the cookbook sre.mysql.multiinstance_reboot * 12:36 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1273,1275-1287].eqiad.wmnet * 12:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1248-1261].eqiad.wmnet * 12:36 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1248-1261].eqiad.wmnet * 12:35 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2026.codfw.wmnet * 12:34 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 12:34 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 12:34 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 12:34 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 12:29 lucaswerkmeister-wmde@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 12:28 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2026.codfw.wmnet * 12:28 lucaswerkmeister-wmde@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 12:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1248-1261].eqiad.wmnet * 12:27 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.addnode (exit_code=0) for new host ganeti2046.codfw.wmnet to cluster codfw and group A * 12:26 moritzm: readded ganeti2046 to the codfw cluster following firmware update and reimage [[phab:T434681|T434681]] * 12:23 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply * 12:23 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply * 12:23 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply * 12:22 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply * 12:22 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply * 12:22 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply * 12:21 jmm@cumin2003: START - Cookbook sre.ganeti.addnode for new host ganeti2046.codfw.wmnet to cluster codfw and group A * 12:21 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1169.eqiad.wmnet onto db1283.eqiad.wmnet * 12:21 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1169: Pool db1169.eqiad.wmnet in after cloning * 12:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1248-1261].eqiad.wmnet * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1149-1153,1158,1240-1247].eqiad.wmnet * 12:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1149-1153,1158,1240-1247].eqiad.wmnet * 12:11 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1149-1153,1158,1240-1247].eqiad.wmnet * 12:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1168.eqiad.wmnet onto db1282.eqiad.wmnet * 12:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1168: Pool db1168.eqiad.wmnet in after cloning * 12:06 jmm@cumin2003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-test-eqiad * 12:05 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2026.codfw.wmnet * 12:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1149-1153,1158,1240-1247].eqiad.wmnet * 12:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1128-1134,1142-1148].eqiad.wmnet * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2046.codfw.wmnet * 12:01 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1128-1134,1142-1148].eqiad.wmnet * 11:57 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2025.codfw.wmnet * 11:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2025.codfw.wmnet * 11:54 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2046.codfw.wmnet * 11:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1128-1134,1142-1148].eqiad.wmnet * 11:50 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2025.codfw.wmnet * 11:46 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1128-1134,1142-1148].eqiad.wmnet * 11:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1114-1127].eqiad.wmnet * 11:45 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1114-1127].eqiad.wmnet * 11:39 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db2901.codfw.wmnet * 11:39 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db2901.codfw.wmnet * 11:36 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2025.codfw.wmnet * 11:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1114-1127].eqiad.wmnet * 11:35 fceratto@cumin1003: END (ERROR) - Cookbook sre.ganeti.makevm (exit_code=93) for new host db1901.eqiad.wmnet * 11:35 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1169: Pool db1169.eqiad.wmnet in after cloning * 11:35 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 11:30 jmm@cumin2003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-test-eqiad * 11:27 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1114-1127].eqiad.wmnet * 11:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1076-1081,1084-1087,1093-1095,1113].eqiad.wmnet * 11:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1076-1081,1084-1087,1093-1095,1113].eqiad.wmnet * 11:25 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1168: Pool db1168.eqiad.wmnet in after cloning * 11:24 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1194.eqiad.wmnet with OS bookworm * 11:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1181.eqiad.wmnet with OS bookworm * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2050.codfw.wmnet * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2050.codfw.wmnet * 11:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1076-1081,1084-1087,1093-1095,1113].eqiad.wmnet * 11:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2050.codfw.wmnet * 11:14 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2004.codfw.wmnet * 11:10 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2050.codfw.wmnet * 11:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1076-1081,1084-1087,1093-1095,1113].eqiad.wmnet * 11:09 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1286: Pool back * 11:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1045-1050,1056-1057,1064-1066,1073-1075].eqiad.wmnet * 11:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1045-1050,1056-1057,1064-1066,1073-1075].eqiad.wmnet * 11:08 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2004.codfw.wmnet * 11:07 moritzm: installing PHP 8.4 security updates * 11:06 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on an-worker1194.eqiad.wmnet with reason: host reimage * 11:06 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1194.eqiad.wmnet with reason: host reimage * 11:05 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2049.codfw.wmnet * 11:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2049.codfw.wmnet * 11:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid1003.eqiad.wmnet * 11:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1045-1050,1056-1057,1064-1066,1073-1075].eqiad.wmnet * 10:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1181.eqiad.wmnet with reason: host reimage * 10:59 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2049.codfw.wmnet * 10:59 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid1003.eqiad.wmnet * 10:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid2003.codfw.wmnet * 10:54 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1181.eqiad.wmnet with reason: host reimage * 10:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid2003.codfw.wmnet * 10:51 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1901.eqiad.wmnet on all recursors * 10:51 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1901.eqiad.wmnet on all recursors * 10:51 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:51 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:51 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:50 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2049.codfw.wmnet * 10:50 blake@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:50 blake@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:49 blake@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:48 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1045-1050,1056-1057,1064-1066,1073-1075].eqiad.wmnet * 10:48 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1044].eqiad.wmnet * 10:48 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1044].eqiad.wmnet * 10:47 blake@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:46 blake@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:46 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2047.codfw.wmnet * 10:46 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:46 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2047.codfw.wmnet * 10:46 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1901.eqiad.wmnet * 10:46 blake@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:44 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1168: Depool db1168.eqiad.wmnet to then clone it to db1282.eqiad.wmnet - marostegui@cumin1003 * 10:44 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1168: Depool db1168.eqiad.wmnet to then clone it to db1282.eqiad.wmnet - marostegui@cumin1003 * 10:44 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1168.eqiad.wmnet onto db1282.eqiad.wmnet * 10:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2047.codfw.wmnet * 10:40 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1044].eqiad.wmnet * 10:37 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2047.codfw.wmnet * 10:37 fceratto@cumin1003: END (ERROR) - Cookbook sre.ganeti.makevm (exit_code=93) for new host db1901.eqiad.wmnet * 10:36 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 10:34 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:33 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2032.codfw.wmnet * 10:33 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2032.codfw.wmnet * 10:32 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1044].eqiad.wmnet * 10:32 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-eqiad * 10:27 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2032.codfw.wmnet * 10:24 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1286: Pool back * 10:24 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1286 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96163 and previous config saved to /var/cache/conftool/dbconfig/20260818-102431-marostegui.json * 10:22 moritzm: installing Django security updates * 10:20 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul2002.codfw.wmnet * 10:20 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul2001.codfw.wmnet * 10:17 blake@deploy1003: Finished scap sync-world: no-build deployment for [[phab:T417800|T417800]] (duration: 04m 40s) * 10:16 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul2002.codfw.wmnet * 10:16 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul2001.codfw.wmnet * 10:15 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2032.codfw.wmnet * 10:14 blake@deploy1003: Started scap sync-world: no-build deployment for [[phab:T417800|T417800]] * 10:12 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul2003.codfw.wmnet * 10:12 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy2001.codfw.wmnet * 10:12 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy3001.esams.wmnet * 10:12 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy1001.eqiad.wmnet * 10:08 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2031.codfw.wmnet * 10:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2031.codfw.wmnet * 10:08 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul2003.codfw.wmnet * 10:08 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy2001.codfw.wmnet * 10:08 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy3001.esams.wmnet * 10:08 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy1001.eqiad.wmnet * 10:07 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy1002.eqiad.wmnet * 10:05 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy2002.codfw.wmnet * 10:05 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy3002.esams.wmnet * 10:04 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy1002.eqiad.wmnet * 10:03 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy4003.ulsfo.wmnet * 10:03 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy5003.eqsin.wmnet * 10:02 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2031.codfw.wmnet * 10:02 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1169: Depool db1169.eqiad.wmnet to then clone it to db1283.eqiad.wmnet - marostegui@cumin1003 * 10:01 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy2002.codfw.wmnet * 10:01 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy4004.ulsfo.wmnet * 10:01 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy3002.esams.wmnet * 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1169: Depool db1169.eqiad.wmnet to then clone it to db1283.eqiad.wmnet - marostegui@cumin1003 * 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1169.eqiad.wmnet onto db1283.eqiad.wmnet * 10:01 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy5004.eqsin.wmnet * 09:59 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy4003.ulsfo.wmnet * 09:59 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy4004.ulsfo.wmnet * 09:59 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy5003.eqsin.wmnet * 09:59 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy7001.magru.wmnet * 09:59 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy5004.eqsin.wmnet * 09:58 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy7002.magru.wmnet * 09:57 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2031.codfw.wmnet * 09:57 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy6001.drmrs.wmnet * 09:57 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy6002.drmrs.wmnet * 09:54 filippo@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for 10 hosts * 09:54 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast4006.wikimedia.org * 09:53 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy6001.drmrs.wmnet * 09:53 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy6002.drmrs.wmnet * 09:53 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host releases1003.eqiad.wmnet * 09:53 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy7001.magru.wmnet * 09:53 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host people2004.codfw.wmnet * 09:52 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host people1005.eqiad.wmnet * 09:52 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy7002.magru.wmnet * 09:50 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host releases2003.codfw.wmnet * 09:49 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host releases2003.codfw.wmnet * 09:49 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host releases1003.eqiad.wmnet * 09:49 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host people2004.codfw.wmnet * 09:48 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host people1005.eqiad.wmnet * 09:46 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2040.codfw.wmnet * 09:46 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2040.codfw.wmnet * 09:43 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1901.eqiad.wmnet on all recursors * 09:42 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1901.eqiad.wmnet on all recursors * 09:42 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:42 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:42 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2040.codfw.wmnet * 09:40 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp2005.wikimedia.org * 09:36 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp2005.wikimedia.org * 09:31 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:31 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1901.eqiad.wmnet * 09:31 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1901.eqiad.wmnet * 09:31 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:31 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1901.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 09:31 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1901.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 09:29 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2040.codfw.wmnet * 09:29 slyngshede@dns1004: END - running authdns-update * 09:28 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2039.codfw.wmnet * 09:28 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2039.codfw.wmnet * 09:27 slyngshede@dns1004: START - running authdns-update * 09:26 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test2005.wikimedia.org * 09:22 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2039.codfw.wmnet * 09:22 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test2005.wikimedia.org * 09:22 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp1005.wikimedia.org * 09:21 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:19 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2039.codfw.wmnet * 09:19 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2038.codfw.wmnet * 09:18 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1287: Pool back * 09:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2038.codfw.wmnet * 09:18 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp1005.wikimedia.org * 09:18 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test1005.wikimedia.org * 09:17 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1901.eqiad.wmnet * 09:15 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db1901.eqiad.wmnet * 09:15 fceratto@cumin1003: END (ERROR) - Cookbook sre.dns.netbox (exit_code=97) * 09:14 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test1005.wikimedia.org * 09:13 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2038.codfw.wmnet * 09:13 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:13 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1901.eqiad.wmnet * 09:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast4006.wikimedia.org * 09:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1288: Pool back * 09:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast5005.wikimedia.org * 09:03 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2038.codfw.wmnet * 08:57 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2037.codfw.wmnet * 08:57 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast5005.wikimedia.org * 08:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2037.codfw.wmnet * 08:56 filippo@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 10 hosts * 08:52 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2037.codfw.wmnet * 08:51 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1289: Pool back * 08:35 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudidp2001-dev.codfw.wmnet * 08:34 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1002-dev.eqiad.wmnet * 08:33 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1287: Pool back * 08:33 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1287 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96150 and previous config saved to /var/cache/conftool/dbconfig/20260818-083311-marostegui.json * 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1001-dev.eqiad.wmnet * 08:31 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudidp2001-dev.codfw.wmnet * 08:30 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1002-dev.eqiad.wmnet * 08:30 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2037.codfw.wmnet * 08:29 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1001-dev.eqiad.wmnet * 08:28 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2036.codfw.wmnet * 08:28 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2036.codfw.wmnet * 08:25 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1288: Pool back * 08:23 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 08:23 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1181.eqiad.wmnet with OS bookworm * 08:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2036.codfw.wmnet * 08:22 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1288 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96148 and previous config saved to /var/cache/conftool/dbconfig/20260818-082234-marostegui.json * 08:20 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2036.codfw.wmnet * 08:18 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2035.codfw.wmnet * 08:18 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudvirt1057.eqiad.wmnet with OS trixie * 08:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2035.codfw.wmnet * 08:13 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2035.codfw.wmnet * 08:06 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2035.codfw.wmnet * 08:05 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1289: Pool back * 08:05 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1289 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96145 and previous config saved to /var/cache/conftool/dbconfig/20260818-080531-marostegui.json * 07:54 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 07:51 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326439{{!}}Fix "mathjax_ignore" handling around forcemathmode attribute (T434686)]] (duration: 13m 15s) * 07:48 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 07:47 krinkle@deploy1003: krinkle: Continuing with deployment * 07:44 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ganeti2046.codfw.wmnet with OS bookworm * 07:40 krinkle@deploy1003: krinkle: Backport for [[gerrit:1326439{{!}}Fix "mathjax_ignore" handling around forcemathmode attribute (T434686)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:39 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 07:38 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1326439{{!}}Fix "mathjax_ignore" handling around forcemathmode attribute (T434686)]] * 07:34 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 07:31 samwilson@deploy1003: Finished scap sync-world: Backport for [[gerrit:701016{{!}}InitialiseSettings and -labs: Remove redundant feature flag $wgWikisourceEnableOcr (T285311)]] (duration: 07m 47s) * 07:30 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1048.eqiad.wmnet * 07:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1048.eqiad.wmnet * 07:29 XioNoX: add gnmic 0.47.0 to bookworm and trixie reprepro * 07:28 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ganeti2046.codfw.wmnet with reason: host reimage * 07:27 samwilson@deploy1003: samwilson: Continuing with deployment * 07:25 samwilson@deploy1003: samwilson: Backport for [[gerrit:701016{{!}}InitialiseSettings and -labs: Remove redundant feature flag $wgWikisourceEnableOcr (T285311)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:25 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 07:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1048.eqiad.wmnet * 07:24 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ganeti2046.codfw.wmnet with reason: host reimage * 07:23 samwilson@deploy1003: Started scap sync-world: Backport for [[gerrit:701016{{!}}InitialiseSettings and -labs: Remove redundant feature flag $wgWikisourceEnableOcr (T285311)]] * 07:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1198.eqiad.wmnet with OS bookworm * 07:18 samwilson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326450{{!}}InitialiseSettings.php: Enable Bulk OCR on pawikisource (T434648)]] (duration: 12m 17s) * 07:17 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 07:11 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1048.eqiad.wmnet * 07:11 samwilson@deploy1003: samwilson: Continuing with deployment * 07:11 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ganeti2046.codfw.wmnet with OS bookworm * 07:10 samwilson@deploy1003: samwilson: Backport for [[gerrit:1326450{{!}}InitialiseSettings.php: Enable Bulk OCR on pawikisource (T434648)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1057.eqiad.wmnet with OS trixie * 07:05 samwilson@deploy1003: Started scap sync-world: Backport for [[gerrit:1326450{{!}}InitialiseSettings.php: Enable Bulk OCR on pawikisource (T434648)]] * 07:02 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1198.eqiad.wmnet with reason: host reimage * 07:01 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1056.eqiad.wmnet with OS trixie * 07:01 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1056.eqiad.wmnet with OS trixie * 07:00 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1056.eqiad.wmnet with OS trixie * 07:00 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1056.eqiad.wmnet with OS trixie * 06:59 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudvirt1055.eqiad.wmnet with OS trixie * 06:58 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1198.eqiad.wmnet with reason: host reimage * 06:53 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1055.eqiad.wmnet with OS trixie * 06:52 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudvirt1054.eqiad.wmnet with OS trixie * 06:44 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1194.eqiad.wmnet with OS bookworm * 06:44 XioNoX: upgrade eqsin gnmic to 0.47.0 * 06:43 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1198.eqiad.wmnet with OS bookworm * 06:41 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1054.eqiad.wmnet with OS trixie * 06:09 arnaudb@cumin1003: END (PASS) - Cookbook sre.gerrit.restart-gerrit (exit_code=0) Restarting Gerrit on gerrit2002 * 06:06 arnaudb@cumin1003: START - Cookbook sre.gerrit.restart-gerrit Restarting Gerrit on gerrit2002 * 06:06 arnaudb@cumin1003: END (PASS) - Cookbook sre.gerrit.restart-gerrit (exit_code=0) Restarting Gerrit on gerrit1003 * 06:04 arnaudb@cumin1003: START - Cookbook sre.gerrit.restart-gerrit Restarting Gerrit on gerrit1003 * 06:02 arnaudb@cumin1003: END (PASS) - Cookbook sre.gerrit.restart-gerrit (exit_code=0) Restarting Gerrit on gerrit2003 * 06:00 arnaudb@cumin1003: START - Cookbook sre.gerrit.restart-gerrit Restarting Gerrit on gerrit2003 * 05:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1155.eqiad.wmnet with OS bookworm * 05:38 arnaudb: updating prometheusBearerToken on gerrit * 05:28 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1155.eqiad.wmnet with reason: host reimage * 05:23 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1155.eqiad.wmnet with reason: host reimage * 05:06 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1155.eqiad.wmnet with OS bookworm * 04:57 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1144.eqiad.wmnet with OS bookworm * 04:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1144.eqiad.wmnet with reason: host reimage * 04:29 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1144.eqiad.wmnet with reason: host reimage * 04:14 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1144.eqiad.wmnet with OS bookworm * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.13 (duration: 02m 23s) * 03:45 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1197.eqiad.wmnet with OS bookworm * 03:41 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1165.eqiad.wmnet with OS bookworm * 03:38 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.16 refs [[phab:T430835|T430835]] (duration: 34m 43s) * 03:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1164.eqiad.wmnet with OS bookworm * 03:35 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1196.eqiad.wmnet with OS bookworm * 03:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1163.eqiad.wmnet with OS bookworm * 03:22 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1197.eqiad.wmnet with reason: host reimage * 03:18 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1165.eqiad.wmnet with reason: host reimage * 03:15 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1196.eqiad.wmnet with reason: host reimage * 03:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1164.eqiad.wmnet with reason: host reimage * 03:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1163.eqiad.wmnet with reason: host reimage * 03:05 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1165.eqiad.wmnet with reason: host reimage * 03:04 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1164.eqiad.wmnet with reason: host reimage * 03:04 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1197.eqiad.wmnet with reason: host reimage * 03:04 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1196.eqiad.wmnet with reason: host reimage * 03:04 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1163.eqiad.wmnet with reason: host reimage * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.16 refs [[phab:T430835|T430835]] * 02:50 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1197.eqiad.wmnet with OS bookworm * 02:49 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1196.eqiad.wmnet with OS bookworm * 02:49 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1165.eqiad.wmnet with OS bookworm * 02:49 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1164.eqiad.wmnet with OS bookworm * 02:48 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1163.eqiad.wmnet with OS bookworm * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 46s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:15 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-codfw: Set storage compatability to NONE — [[phab:T433026|T433026]] - eevans@cumin1003 * 00:38 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-codfw: Set storage compatability to NONE — [[phab:T433026|T433026]] - eevans@cumin1003 * 00:11 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326373{{!}}PersonalDashboard: add newly renamed *ReviewChangesMlModel setting (T422148)]] (duration: 07m 07s) * 00:09 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-eqiad: Set storage compatability to NONE — [[phab:T433026|T433026]] - eevans@cumin1003 * 00:07 musikanimal@deploy1003: musikanimal: Continuing with deployment * 00:06 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1326373{{!}}PersonalDashboard: add newly renamed *ReviewChangesMlModel setting (T422148)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:04 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1326373{{!}}PersonalDashboard: add newly renamed *ReviewChangesMlModel setting (T422148)]] == 2026-08-17 == * 23:30 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-eqiad: Set storage compatability to NONE — [[phab:T433026|T433026]] - eevans@cumin1003 * 23:05 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-codfw: Set storage compatability to UPGRADING — [[phab:T433026|T433026]] - eevans@cumin1003 * 22:28 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-codfw: Set storage compatability to UPGRADING — [[phab:T433026|T433026]] - eevans@cumin1003 * 21:46 logmsgbot: jforrester Deployed security patch for [[phab:T435085|T435085]] * 21:39 swfrench@deploy1003: mwscript-k8s job started: purgeList.php # [[phab:T432412|T432412]] * 21:37 maryum: Undeploy security fix for [[phab:T433020|T433020]] * 21:23 maryum: Deployed security fix for [[phab:T433020|T433020]] * 21:21 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 21:21 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 21:14 maryum: Deployed security fix for [[phab:T434967|T434967]] * 20:58 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-eqiad: Set storage compatability to UPGRADING — [[phab:T433026|T433026]] - eevans@cumin1003 * 20:40 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324320{{!}}[itwiki/slwiki/tgwiki] Remove temporary Wikipedia 25 logos permanently (already reverted) (T414265 T414320 T415307)]] (duration: 06m 55s) * 20:40 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 20:39 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 20:36 cjming@deploy1003: cjming, superpes: Continuing with deployment * 20:35 cjming@deploy1003: cjming, superpes: Backport for [[gerrit:1324320{{!}}[itwiki/slwiki/tgwiki] Remove temporary Wikipedia 25 logos permanently (already reverted) (T414265 T414320 T415307)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:33 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1324320{{!}}[itwiki/slwiki/tgwiki] Remove temporary Wikipedia 25 logos permanently (already reverted) (T414265 T414320 T415307)]] * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ttmserver-test: apply * 20:30 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326361{{!}}Remove escaped paths in app site association file (T432412)]] (duration: 13m 28s) * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ttmserver-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-toolhub-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-toolhub-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-toolhub-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-toolhub-test: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-toolhub: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-toolhub: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-toolhub: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-toolhub: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-test: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-test: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 20:27 inflatador: bking@deploy1003 `charlie --services_dir dse-k8s-services -s opensearch-* -e dse-k8s-* apply` [[phab:T435125|T435125]] * 20:27 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-apifeatureusage-test: apply * 20:27 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-apifeatureusage-test: apply * 20:27 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-apifeatureusage-test: apply * 20:27 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-apifeatureusage-test: apply * 20:27 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-apifeatureusage: apply * 20:26 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-apifeatureusage: apply * 20:26 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-apifeatureusage: apply * 20:26 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-apifeatureusage: apply * 20:26 cjming@deploy1003: cjming, tsev: Continuing with deployment * 20:24 inflatador: bking@deploy1003 `charlie --services_dir dse-k8s-services -s opensearch-* -e dse-k8s-*` [[phab:T435125|T435125]] * 20:21 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-eqiad: Set storage compatability to UPGRADING — [[phab:T433026|T433026]] - eevans@cumin1003 * 20:19 cjming@deploy1003: cjming, tsev: Backport for [[gerrit:1326361{{!}}Remove escaped paths in app site association file (T432412)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:17 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1326361{{!}}Remove escaped paths in app site association file (T432412)]] * 20:15 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326077{{!}}InstrumentConstructiveEdits: anchor all runs to the nearest `interval` (T431493)]] (duration: 06m 25s) * 20:11 cjming@deploy1003: cjming: Continuing with deployment * 20:10 cjming@deploy1003: cjming: Backport for [[gerrit:1326077{{!}}InstrumentConstructiveEdits: anchor all runs to the nearest `interval` (T431493)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:08 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1326077{{!}}InstrumentConstructiveEdits: anchor all runs to the nearest `interval` (T431493)]] * 20:06 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1160.eqiad.wmnet with OS bookworm * 20:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1162.eqiad.wmnet with OS bookworm * 19:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1184.eqiad.wmnet with OS bookworm * 19:55 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1161.eqiad.wmnet with OS bookworm * 19:49 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1195.eqiad.wmnet with OS bookworm * 19:49 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-test: apply * 19:49 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-test: apply * 19:44 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1160.eqiad.wmnet with reason: host reimage * 19:41 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-codfw: Upgrade to Java 17 — [[phab:T433026|T433026]] - eevans@cumin1003 * 19:39 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1162.eqiad.wmnet with reason: host reimage * 19:36 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-eqsin and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 19:36 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-test: apply * 19:36 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1184.eqiad.wmnet with reason: host reimage * 19:32 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1161.eqiad.wmnet with reason: host reimage * 19:29 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1195.eqiad.wmnet with reason: host reimage * 19:26 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1161.eqiad.wmnet with reason: host reimage * 19:26 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1160.eqiad.wmnet with reason: host reimage * 19:26 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1184.eqiad.wmnet with reason: host reimage * 19:26 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1162.eqiad.wmnet with reason: host reimage * 19:25 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1195.eqiad.wmnet with reason: host reimage * 19:19 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-test: apply * 19:11 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1195.eqiad.wmnet with OS bookworm * 19:10 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1194.eqiad.wmnet with OS bookworm * 19:10 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1184.eqiad.wmnet with OS bookworm * 19:10 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1162.eqiad.wmnet with OS bookworm * 19:10 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1161.eqiad.wmnet with OS bookworm * 19:10 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1160.eqiad.wmnet with OS bookworm * 19:10 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 19:09 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 19:03 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-codfw: Upgrade to Java 17 — [[phab:T433026|T433026]] - eevans@cumin1003 * 18:58 dancy@deploy1003: Finished scap sync-world: testing [[phab:T375514|T375514]] (duration: 03m 13s) * 18:55 dancy@deploy1003: Started scap sync-world: testing [[phab:T375514|T375514]] * 18:55 dwisehaupt@dns1006: END - running authdns-update * 18:54 dancy@deploy1003: Installation of scap version "4.281.1" completed for 3 hosts * 18:54 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2009.codfw.wmnet * 18:53 dwisehaupt@dns1006: START - running authdns-update * 18:52 dancy@deploy1003: Installing scap version "4.281.1" for 3 host(s) * 18:47 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2009.codfw.wmnet * 18:41 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2008.codfw.wmnet * 18:34 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2008.codfw.wmnet * 18:30 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2007.codfw.wmnet * 18:23 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2007.codfw.wmnet * 18:16 dwisehaupt@dns1005: END - running authdns-update * 18:14 dwisehaupt@dns1005: START - running authdns-update * 18:04 swfrench@deploy1003: Finished scap sync-world: Deploy "Point Test Wiki to new docroot" - [[phab:T432412|T432412]] (duration: 26m 03s) * 18:00 swfrench@deploy1003: swfrench: Continuing with deployment * 17:51 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1057.eqiad.wmnet with OS trixie * 17:47 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1180.eqiad.wmnet with OS bookworm * 17:39 swfrench@deploy1003: swfrench: Deploy "Point Test Wiki to new docroot" - [[phab:T432412|T432412]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:39 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-eqsin and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 17:38 swfrench@deploy1003: Started scap sync-world: Deploy "Point Test Wiki to new docroot" - [[phab:T432412|T432412]] * 17:36 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1159.eqiad.wmnet with OS bookworm * 17:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1158.eqiad.wmnet with OS bookworm * 17:30 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-eqiad: Upgrade to Java 17 — [[phab:T433026|T433026]] - eevans@cumin1003 * 17:30 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-ulsfo or A:cp-drmrs and A:cp - 9.2.15 upgrade ([[phab:T434620|T434620]]) * 17:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1157.eqiad.wmnet with OS bookworm * 17:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1180.eqiad.wmnet with reason: host reimage * 17:22 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 17:21 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1193.eqiad.wmnet with OS bookworm * 17:21 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 17:21 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 17:20 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 17:20 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 17:20 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1183.eqiad.wmnet with OS bookworm * 17:18 swfrench@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 17:18 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 17:17 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1180.eqiad.wmnet with reason: host reimage * 17:17 swfrench@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 17:17 swfrench@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 17:16 swfrench@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 17:16 swfrench@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 17:15 swfrench@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 17:15 swfrench@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 17:15 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1192.eqiad.wmnet with OS bookworm * 17:14 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1159.eqiad.wmnet with reason: host reimage * 17:14 swfrench@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 17:13 swfrench@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 17:12 swfrench@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 17:09 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1158.eqiad.wmnet with reason: host reimage * 17:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1157.eqiad.wmnet with reason: host reimage * 17:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1180 * 17:02 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1180 * 17:01 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica-eqsin and A:liberica * 17:01 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1193.eqiad.wmnet with reason: host reimage * 16:57 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1183.eqiad.wmnet with reason: host reimage * 16:56 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1054.eqiad.wmnet with OS trixie * 16:55 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1056.eqiad.wmnet with OS trixie * 16:54 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1193.eqiad.wmnet with reason: host reimage * 16:54 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1192.eqiad.wmnet with reason: host reimage * 16:51 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-eqiad: Upgrade to Java 17 — [[phab:T433026|T433026]] - eevans@cumin1003 * 16:50 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1180 * 16:50 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1180.eqiad.wmnet 17.36.64.10.in-addr.arpa 7.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:50 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1180.eqiad.wmnet 17.36.64.10.in-addr.arpa 7.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:50 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:50 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1180 - btullis@cumin1003" * 16:50 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1180 - btullis@cumin1003" * 16:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1158.eqiad.wmnet with reason: host reimage * 16:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1157.eqiad.wmnet with reason: host reimage * 16:49 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica-eqsin and A:liberica * 16:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1183.eqiad.wmnet with reason: host reimage * 16:48 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1159.eqiad.wmnet with reason: host reimage * 16:47 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1192.eqiad.wmnet with reason: host reimage * 16:46 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns4004.wikimedia.org * 16:46 sukhe@dns1004: END - running authdns-update * 16:44 sukhe@dns1004: START - running authdns-update * 16:44 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns4004.wikimedia.org,service=authdns-update * 16:43 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns4004.wikimedia.org with OS trixie * 16:39 btullis@cumin1003: START - Cookbook sre.dns.netbox * 16:39 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1180 * 16:39 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1193.eqiad.wmnet with OS bookworm * 16:39 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1180.eqiad.wmnet with OS bookworm * 16:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1192.eqiad.wmnet with OS bookworm * 16:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1183.eqiad.wmnet with OS bookworm * 16:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1159.eqiad.wmnet with OS bookworm * 16:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1158.eqiad.wmnet with OS bookworm * 16:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1157.eqiad.wmnet with OS bookworm * 16:31 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1057.eqiad.wmnet with OS trixie * 16:30 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1057.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 16:29 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica-drmrs and A:liberica * 16:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1154.eqiad.wmnet with OS bookworm * 16:23 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1057.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 16:23 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1055.eqiad.wmnet with OS trixie * 16:22 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1057 * 16:22 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1057 * 16:21 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:21 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1057] - vriley@cumin1003" * 16:21 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1057] - vriley@cumin1003" * 16:19 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica-drmrs and A:liberica * 16:17 vriley@cumin1003: START - Cookbook sre.dns.netbox * 16:16 phuedx@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics-external: apply * 16:15 phuedx@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics-external: apply * 16:13 phuedx@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics-external: apply * 16:12 phuedx@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics-external: apply * 16:11 phuedx@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics-external: apply * 16:09 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics-external: apply * 16:09 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1176.eqiad.wmnet with OS bookworm * 16:07 btullis@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 16:06 btullis@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 16:05 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1191.eqiad.wmnet with OS bookworm * 16:03 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1154.eqiad.wmnet with reason: host reimage * 16:03 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching aqs[2002-2012].codfw.wmnet,aqs[1017-1027].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433026|T433026]] - eevans@cumin1003 * 15:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1190.eqiad.wmnet with OS bookworm * 15:59 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1154.eqiad.wmnet with reason: host reimage * 15:55 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-debug: apply * 15:55 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-debug: apply * 15:55 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-debug: apply * 15:55 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/mw-debug: apply * 15:53 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns4004.wikimedia.org with reason: host reimage * 15:50 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns4004.wikimedia.org with reason: host reimage * 15:46 btullis@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 15:46 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1176.eqiad.wmnet with reason: host reimage * 15:45 btullis@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 15:43 btullis@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 15:42 btullis@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 15:42 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1191.eqiad.wmnet with reason: host reimage * 15:39 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1190.eqiad.wmnet with reason: host reimage * 15:36 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1054.eqiad.wmnet with OS trixie * 15:35 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:35 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1056.eqiad.wmnet with OS trixie * 15:35 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:34 moritzm: failover Ganeti master in eqiad to ganeti1046 * 15:34 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1176.eqiad.wmnet with reason: host reimage * 15:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1191.eqiad.wmnet with reason: host reimage * 15:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1190.eqiad.wmnet with reason: host reimage * 15:31 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns2006.wikimedia.org * 15:31 sukhe@dns1004: END - running authdns-update * 15:29 sukhe@dns1004: START - running authdns-update * 15:29 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns2006.wikimedia.org,service=authdns-update * 15:29 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns2006.wikimedia.org * 15:29 sukhe@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns2006.wikimedia.org * 15:26 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica-ulsfo and A:liberica * 15:26 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns1006.wikimedia.org * 15:26 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:25 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns2006.wikimedia.org with OS trixie * 15:25 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1056 * 15:25 sukhe@dns1004: END - running authdns-update * 15:23 sukhe@dns1004: START - running authdns-update * 15:23 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns1006.wikimedia.org,service=authdns-update * 15:23 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns1006.wikimedia.org * 15:23 sukhe@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns1006.wikimedia.org * 15:20 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1056 * 15:20 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:20 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1056~] - vriley@cumin1003" * 15:19 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1056~] - vriley@cumin1003" * 15:19 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns1006.wikimedia.org with OS trixie * 15:19 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns4004.wikimedia.org with OS trixie * 15:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1190.eqiad.wmnet with OS bookworm * 15:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1191.eqiad.wmnet with OS bookworm * 15:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1176.eqiad.wmnet with OS bookworm * 15:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1154.eqiad.wmnet with OS bookworm * 15:16 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica-ulsfo and A:liberica * 15:13 vriley@cumin1003: START - Cookbook sre.dns.netbox * 15:13 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host dns4004.wikimedia.org with OS trixie * 15:12 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1188.eqiad.wmnet with OS bookworm * 15:09 taavi@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318203{{!}}Undeploy WP25EasterEggs (II) (T418134)]] (duration: 08m 56s) * 15:06 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica-magru and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 15:06 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1051.eqiad.wmnet with OS trixie * 15:06 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 15:05 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 15:05 taavi@deploy1003: taavi: Continuing with deployment * 15:04 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-ncredir (exit_code=0) rolling reboot on A:ncredir and A:ncredir * 15:04 taavi@deploy1003: taavi: Backport for [[gerrit:1318203{{!}}Undeploy WP25EasterEggs (II) (T418134)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:03 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1055.eqiad.wmnet with OS trixie * 15:02 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:00 taavi@deploy1003: Started scap sync-world: Backport for [[gerrit:1318203{{!}}Undeploy WP25EasterEggs (II) (T418134)]] * 14:59 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1047.eqiad.wmnet * 14:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1047.eqiad.wmnet * 14:58 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy (exit_code=0) rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 14:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-codfw * 14:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp2001.codfw.wmnet * 14:57 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp2001.codfw.wmnet * 14:57 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica-magru and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 14:57 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:56 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1055 * 14:56 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1055 * 14:55 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns2006.wikimedia.org with reason: host reimage * 14:55 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:55 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1055] - vriley@cumin1003" * 14:55 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1055] - vriley@cumin1003" * 14:54 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1047.eqiad.wmnet * 14:52 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1188.eqiad.wmnet with reason: host reimage * 14:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp2001.codfw.wmnet * 14:51 vriley@cumin1003: START - Cookbook sre.dns.netbox * 14:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp2001.codfw.wmnet * 14:50 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2016.codfw.wmnet * 14:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2016.codfw.wmnet * 14:50 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1189.eqiad.wmnet with OS bookworm * 14:48 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1188.eqiad.wmnet with reason: host reimage * 14:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1051.eqiad.wmnet with reason: host reimage * 14:46 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1182.eqiad.wmnet with OS bookworm * 14:45 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief2002.codfw.wmnet * 14:44 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns1006.wikimedia.org with reason: host reimage * 14:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2016.codfw.wmnet * 14:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1143.eqiad.wmnet with OS bookworm * 14:43 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2016.codfw.wmnet * 14:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2318-2331].codfw.wmnet * 14:43 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2318-2331].codfw.wmnet * 14:42 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching aqs[2002-2012].codfw.wmnet,aqs[1017-1027].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433026|T433026]] - eevans@cumin1003 * 14:41 cgoubert@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326306{{!}}Add placeholder $wmgRedisLockPassword (T366938 T427999)]] (duration: 06m 56s) * 14:41 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief2002.codfw.wmnet * 14:40 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief1002.eqiad.wmnet * 14:39 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1051.eqiad.wmnet with reason: host reimage * 14:38 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns2006.wikimedia.org with reason: host reimage * 14:37 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns1006.wikimedia.org with reason: host reimage * 14:37 cgoubert@deploy1003: cgoubert: Continuing with deployment * 14:36 cgoubert@deploy1003: cgoubert: Backport for [[gerrit:1326306{{!}}Add placeholder $wmgRedisLockPassword (T366938 T427999)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:36 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief1002.eqiad.wmnet * 14:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2318-2331].codfw.wmnet * 14:34 cgoubert@deploy1003: Started scap sync-world: Backport for [[gerrit:1326306{{!}}Add placeholder $wmgRedisLockPassword (T366938 T427999)]] * 14:29 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2318-2331].codfw.wmnet * 14:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2304-2317].codfw.wmnet * 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2304-2317].codfw.wmnet * 14:28 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1054 * 14:27 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1054 * 14:27 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:27 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1054] - vriley@cumin1003" * 14:27 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1054] - vriley@cumin1003" * 14:26 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1189.eqiad.wmnet with reason: host reimage * 14:26 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test2001.codfw.wmnet * 14:25 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test1001.eqiad.wmnet * 14:25 claime: Deploying wmgRedisLockPassword - [[phab:T366938|T366938]] [[phab:T427999|T427999]] * 14:24 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1051.eqiad.wmnet with OS trixie * 14:23 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 14:22 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1182.eqiad.wmnet with reason: host reimage * 14:22 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1051.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:22 vriley@cumin1003: START - Cookbook sre.dns.netbox * 14:22 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test2001.codfw.wmnet * 14:21 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test1001.eqiad.wmnet * 14:21 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 14:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2304-2317].codfw.wmnet * 14:19 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns4004.wikimedia.org with OS trixie * 14:19 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns2006.wikimedia.org with OS trixie * 14:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1143.eqiad.wmnet with reason: host reimage * 14:19 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns1006.wikimedia.org with OS trixie * 14:17 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1189.eqiad.wmnet with reason: host reimage * 14:15 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1182.eqiad.wmnet with reason: host reimage * 14:14 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1143.eqiad.wmnet with reason: host reimage * 14:13 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1051.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:12 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2304-2317].codfw.wmnet * 14:12 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2290-2303].codfw.wmnet * 14:12 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1049.eqiad.wmnet with OS trixie * 14:12 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2290-2303].codfw.wmnet * 14:12 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1051 * 14:11 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1051 * 14:11 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1170.eqiad.wmnet onto db1284.eqiad.wmnet * 14:11 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1170: Pool db1170.eqiad.wmnet in after cloning * 14:09 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 14:09 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:09 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1051] - vriley@cumin1003" * 14:09 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1051] - vriley@cumin1003" * 14:06 klausman@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:04 vriley@cumin1003: START - Cookbook sre.dns.netbox * 14:04 klausman@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2290-2303].codfw.wmnet * 14:02 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1189.eqiad.wmnet with OS bookworm * 14:02 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1188.eqiad.wmnet with OS bookworm * 14:01 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 14:00 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1182.eqiad.wmnet with OS bookworm * 14:00 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1143.eqiad.wmnet with OS bookworm * 13:59 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 13:57 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply * 13:57 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply * 13:56 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply * 13:56 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-ulsfo or A:cp-drmrs and A:cp - 9.2.15 upgrade ([[phab:T434620|T434620]]) * 13:56 cjd91: sudo -i cookbook sre.cdn.roll-upgrade-ats --query 'A:cp-ulsfo or A:cp-drmrs' --task-id [[phab:T434620|T434620]] --reason '9.2.15 upgrade' * 13:56 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply * 13:55 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply * 13:55 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply * 13:54 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 13:54 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 13:54 phuedx@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:53 phuedx@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics-external: apply * 13:52 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1049.eqiad.wmnet with reason: host reimage * 13:51 phuedx@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2290-2303].codfw.wmnet * 13:50 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1047.eqiad.wmnet * 13:50 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2276-2289].codfw.wmnet * 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2276-2289].codfw.wmnet * 13:49 phuedx@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics-external: apply * 13:49 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1046.eqiad.wmnet * 13:49 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1049.eqiad.wmnet with reason: host reimage * 13:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1046.eqiad.wmnet * 13:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-ncredir rolling reboot on A:ncredir and A:ncredir * 13:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 13:46 phuedx@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:44 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics-external: apply * 13:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1046.eqiad.wmnet * 13:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2276-2289].codfw.wmnet * 13:36 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1046.eqiad.wmnet * 13:36 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1045.eqiad.wmnet * 13:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1045.eqiad.wmnet * 13:34 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2276-2289].codfw.wmnet * 13:34 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1049.eqiad.wmnet with OS trixie * 13:34 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2262-2275].codfw.wmnet * 13:34 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2262-2275].codfw.wmnet * 13:32 Lucas_WMDE: UTC afternoon backport+config window done * 13:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1045.eqiad.wmnet * 13:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2262-2275].codfw.wmnet * 13:26 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1049.eqiad.wmnet with OS trixie * 13:26 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1049.eqiad.wmnet with OS trixie * 13:26 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1170: Pool db1170.eqiad.wmnet in after cloning * 13:23 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1049.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 13:23 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1045.eqiad.wmnet * 13:21 atsukoito: manually done sudo -i docker-registryctl --debug delete-tags 'docker-registry.discovery.wmnet/repos/data-engineering/airflow-dags:airflow-3.3.0-py3.11-2026-08-17-*' to remove incorrect tags * 13:19 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311141{{!}}viwiki: Set `noindex,nofollow` for User and User talk (T432311)]] (duration: 11m 11s) * 13:15 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2262-2275].codfw.wmnet * 13:14 lucaswerkmeister-wmde@deploy1003: ndkdd, lucaswerkmeister-wmde: Continuing with deployment * 13:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2248-2261].codfw.wmnet * 13:14 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2248-2261].codfw.wmnet * 13:13 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1049.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 13:12 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1049 * 13:11 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1049 * 13:10 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:10 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1049] - vriley@cumin1003" * 13:10 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1049] - vriley@cumin1003" * 13:10 lucaswerkmeister-wmde@deploy1003: ndkdd, lucaswerkmeister-wmde: Backport for [[gerrit:1311141{{!}}viwiki: Set `noindex,nofollow` for User and User talk (T432311)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1037.eqiad.wmnet * 13:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1037.eqiad.wmnet * 13:08 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1311141{{!}}viwiki: Set `noindex,nofollow` for User and User talk (T432311)]] * 13:06 vriley@cumin1003: START - Cookbook sre.dns.netbox * 13:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2248-2261].codfw.wmnet * 13:00 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1037.eqiad.wmnet * 12:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2248-2261].codfw.wmnet * 12:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2204-2215,2242-2243].codfw.wmnet * 12:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2204-2215,2242-2243].codfw.wmnet * 12:49 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324832{{!}}Migrate $wgFlaggedRevsTags from flaggedrevs.php to ext-FlaggedRevs.php]] (duration: 14m 02s) * 12:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint2001.codfw.wmnet * 12:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2204-2215,2242-2243].codfw.wmnet * 12:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint2001.codfw.wmnet * 12:41 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1037.eqiad.wmnet * 12:40 ladsgroup@deploy1003: Rolling back deployment * 12:37 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1324832{{!}}Migrate $wgFlaggedRevsTags from flaggedrevs.php to ext-FlaggedRevs.php]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:37 seanleong-wmde: Finished populateSitesTable for [bolwiki] ([[[phab:T429955|T429955]]]) * 12:35 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1324832{{!}}Migrate $wgFlaggedRevsTags from flaggedrevs.php to ext-FlaggedRevs.php]] * 12:35 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2204-2215,2242-2243].codfw.wmnet * 12:34 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2190-2203].codfw.wmnet * 12:34 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2190-2203].codfw.wmnet * 12:32 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1028.eqiad.wmnet * 12:32 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1028.eqiad.wmnet * 12:31 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint1001.eqiad.wmnet * 12:30 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1170: Depool db1170.eqiad.wmnet to then clone it to db1284.eqiad.wmnet - marostegui@cumin1003 * 12:28 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint1001.eqiad.wmnet * 12:28 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1170: Depool db1170.eqiad.wmnet to then clone it to db1284.eqiad.wmnet - marostegui@cumin1003 * 12:27 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1170.eqiad.wmnet onto db1284.eqiad.wmnet * 12:27 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db2901.codfw.wmnet * 12:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2190-2203].codfw.wmnet * 12:27 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 12:26 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1028.eqiad.wmnet * 12:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2190-2203].codfw.wmnet * 12:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2172-2179,2184-2189].codfw.wmnet * 12:18 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2172-2179,2184-2189].codfw.wmnet * 12:13 seanleong-wmde@deploy1003: mwscript-k8s job started: foreachwikiindblist wikidataclient extensions/Wikibase/lib/maintenance/populateSitesTable.php --force-protocol https # [[phab:T429955|T429955]] * 12:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2172-2179,2184-2189].codfw.wmnet * 12:09 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1028.eqiad.wmnet * 12:03 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 12:02 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db2901.codfw.wmnet * 12:02 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db2901.codfw.wmnet * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1027.eqiad.wmnet * 12:02 fceratto@cumin1003: END (ERROR) - Cookbook sre.dns.netbox (exit_code=97) * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1027.eqiad.wmnet * 12:02 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 12:02 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db2901.codfw.wmnet * 12:01 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2172-2179,2184-2189].codfw.wmnet * 12:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2158-2171].codfw.wmnet * 12:01 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2158-2171].codfw.wmnet * 11:58 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db1902.eqiad.wmnet * 11:58 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 11:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1027.eqiad.wmnet * 11:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2158-2171].codfw.wmnet * 11:51 jayme: updated calico to v3.30.7 on wikikube eqiad - [[phab:T427400|T427400]] * 11:50 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1027.eqiad.wmnet * 11:45 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2158-2171].codfw.wmnet * 11:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2144-2157].codfw.wmnet * 11:44 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2144-2157].codfw.wmnet * 11:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1058.eqiad.wmnet * 11:43 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1058.eqiad.wmnet * 11:43 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 11:43 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 11:43 marostegui@cumin1003: Removing db1153 from zarcillo [[phab:T434638|T434638]] * 11:42 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1153.eqiad.wmnet * 11:42 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:42 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1153.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 11:42 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1172.eqiad.wmnet onto db1286.eqiad.wmnet * 11:42 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1172: Pool db1172.eqiad.wmnet in after cloning * 11:42 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1153.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 11:41 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1902.eqiad.wmnet * 11:41 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=97) for new host db2901.codfw.wmnet * 11:41 fceratto@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host db2901.codfw.wmnet with OS trixie * 11:38 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'. * 11:38 marostegui@cumin1003: START - Cookbook sre.dns.netbox * 11:37 marostegui@dns1004: END - running authdns-update * 11:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1058.eqiad.wmnet * 11:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2144-2157].codfw.wmnet * 11:35 marostegui@dns1004: START - running authdns-update * 11:32 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1153.eqiad.wmnet * 11:32 marostegui@cumin1003: START - Cookbook sre.mysql.decommission * 11:28 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326249{{!}}ImagePage: move TOC element below file link (T332644)]] (duration: 09m 56s) * 11:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2144-2157].codfw.wmnet * 11:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2130-2143].codfw.wmnet * 11:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2130-2143].codfw.wmnet * 11:26 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1058.eqiad.wmnet * 11:23 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 11:22 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'. * 11:22 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'. * 11:22 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'. * 11:22 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326249{{!}}ImagePage: move TOC element below file link (T332644)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1057.eqiad.wmnet * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1057.eqiad.wmnet * 11:21 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 11:20 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 11:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2130-2143].codfw.wmnet * 11:19 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326249{{!}}ImagePage: move TOC element below file link (T332644)]] * 11:18 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply * 11:17 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply * 11:17 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply * 11:16 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 11:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1057.eqiad.wmnet * 11:15 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 11:14 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 11:13 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 11:13 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db2901.codfw.wmnet with OS trixie * 11:12 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db2901.codfw.wmnet - fceratto@cumin1003" * 11:12 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db2901.codfw.wmnet - fceratto@cumin1003" * 11:12 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db2901.codfw.wmnet on all recursors * 11:12 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db2901.codfw.wmnet on all recursors * 11:12 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:12 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db2901.codfw.wmnet - fceratto@cumin1003" * 11:12 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2130-2143].codfw.wmnet * 11:11 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2107-2115,2124-2129].codfw.wmnet * 11:11 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2107-2115,2124-2129].codfw.wmnet * 11:11 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 11:11 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 11:11 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply * 11:11 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 11:10 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 11:06 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db2901.codfw.wmnet - fceratto@cumin1003" * 11:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2107-2115,2124-2129].codfw.wmnet * 10:57 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1172: Pool db1172.eqiad.wmnet in after cloning * 10:54 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2107-2115,2124-2129].codfw.wmnet * 10:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2078,2087-2095,2102-2106].codfw.wmnet * 10:54 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2078,2087-2095,2102-2106].codfw.wmnet * 10:47 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply * 10:46 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply * 10:46 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply * 10:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2078,2087-2095,2102-2106].codfw.wmnet * 10:45 blake@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply * 10:38 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1057.eqiad.wmnet * 10:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1056.eqiad.wmnet * 10:38 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1056.eqiad.wmnet * 10:37 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2078,2087-2095,2102-2106].codfw.wmnet * 10:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2061-2062,2064-2065,2067-2077].codfw.wmnet * 10:36 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2061-2062,2064-2065,2067-2077].codfw.wmnet * 10:32 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1056.eqiad.wmnet * 10:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2061-2062,2064-2065,2067-2077].codfw.wmnet * 10:25 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:25 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db2901.codfw.wmnet * 10:25 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db1902.eqiad.wmnet * 10:25 fceratto@cumin1003: END (ERROR) - Cookbook sre.dns.netbox (exit_code=97) * 10:24 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:24 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1902.eqiad.wmnet * 10:24 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=97) for new host db1902.eqiad.wmnet * 10:24 fceratto@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host db1902.eqiad.wmnet with OS trixie * 10:24 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=93) for new host db1903.eqiad.wmnet * 10:24 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 10:20 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1056.eqiad.wmnet * 10:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2061-2062,2064-2065,2067-2077].codfw.wmnet * 10:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2038-2039,2041-2042,2044,2046,2049-2051,2055-2060].codfw.wmnet * 10:18 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2038-2039,2041-2042,2044,2046,2049-2051,2055-2060].codfw.wmnet * 10:15 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1055.eqiad.wmnet * 10:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1055.eqiad.wmnet * 10:13 Amir1: mwscript-k8s --dblist=all -- purgeUserOptions.php --login-age 5 uls-preferences ([[phab:T406724|T406724]]) * 10:11 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:10 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1903.eqiad.wmnet on all recursors * 10:10 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1903.eqiad.wmnet on all recursors * 10:10 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2038-2039,2041-2042,2044,2046,2049-2051,2055-2060].codfw.wmnet * 10:10 moritzm: installing unzip security updates * 10:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1055.eqiad.wmnet * 10:08 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=97) for new host db1901.eqiad.wmnet * 10:08 fceratto@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host db1901.eqiad.wmnet with OS trixie * 10:08 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:08 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 10:08 fceratto@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1903.eqiad.wmnet - fceratto@cumin1003" * 10:07 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db1902.eqiad.wmnet with OS trixie * 10:07 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1902.eqiad.wmnet - fceratto@cumin1003" * 10:07 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1902.eqiad.wmnet - fceratto@cumin1003" * 10:04 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply * 10:04 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326227{{!}}Enable desktop/native lazy loading everywhere (T148047)]] (duration: 07m 13s) * 10:03 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1902.eqiad.wmnet on all recursors * 10:03 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1902.eqiad.wmnet on all recursors * 10:03 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:03 blake@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply * 10:01 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1903.eqiad.wmnet - fceratto@cumin1003" * 10:01 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2038-2039,2041-2042,2044,2046,2049-2051,2055-2060].codfw.wmnet * 10:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2002,2005-2006,2011-2015,2017-2018,2033-2037].codfw.wmnet * 10:01 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2002,2005-2006,2011-2015,2017-2018,2033-2037].codfw.wmnet * 10:00 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:00 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 09:59 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 09:59 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1055.eqiad.wmnet * 09:58 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326227{{!}}Enable desktop/native lazy loading everywhere (T148047)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:56 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326227{{!}}Enable desktop/native lazy loading everywhere (T148047)]] * 09:53 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2002,2005-2006,2011-2015,2017-2018,2033-2037].codfw.wmnet * 09:50 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:48 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1903.eqiad.wmnet * 09:48 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:48 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1902.eqiad.wmnet * 09:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 09:44 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2002,2005-2006,2011-2015,2017-2018,2033-2037].codfw.wmnet * 09:43 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-codfw * 09:43 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1174.eqiad.wmnet onto db1288.eqiad.wmnet * 09:42 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1174: Pool db1174.eqiad.wmnet in after cloning * 09:40 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1172: Depool db1172.eqiad.wmnet to then clone it to db1286.eqiad.wmnet - marostegui@cumin1003 * 09:39 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1172: Depool db1172.eqiad.wmnet to then clone it to db1286.eqiad.wmnet - marostegui@cumin1003 * 09:39 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1172.eqiad.wmnet onto db1286.eqiad.wmnet * 09:33 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1175.eqiad.wmnet onto db1289.eqiad.wmnet * 09:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1175: Pool db1175.eqiad.wmnet in after cloning * 09:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1201.eqiad.wmnet onto db1287.eqiad.wmnet * 09:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1201: Pool db1201.eqiad.wmnet in after cloning * 09:28 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db1901.eqiad.wmnet with OS trixie * 09:27 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:27 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:27 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1901.eqiad.wmnet on all recursors * 09:27 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1901.eqiad.wmnet on all recursors * 09:26 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:26 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:26 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:15 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:15 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1901.eqiad.wmnet * 09:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw1001.wikimedia.org with OS trixie * 08:57 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1174: Pool db1174.eqiad.wmnet in after cloning * 08:54 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1054.eqiad.wmnet * 08:54 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1054.eqiad.wmnet * 08:48 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1054.eqiad.wmnet * 08:47 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1175: Pool db1175.eqiad.wmnet in after cloning * 08:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 08:46 marostegui@cumin1003: Removing db1152 from zarcillo [[phab:T434480|T434480]] * 08:46 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1152.eqiad.wmnet * 08:46 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:46 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1152.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 08:46 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1201: Pool db1201.eqiad.wmnet in after cloning * 08:46 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1152.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 08:46 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1054.eqiad.wmnet * 08:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1053.eqiad.wmnet * 08:43 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1053.eqiad.wmnet * 08:42 marostegui@cumin1003: START - Cookbook sre.dns.netbox * 08:38 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage * 08:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1053.eqiad.wmnet * 08:36 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1152.eqiad.wmnet * 08:36 marostegui@cumin1003: START - Cookbook sre.mysql.decommission * 08:35 phuedx: UTC morning backport window done * 08:35 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1053.eqiad.wmnet * 08:34 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1035.eqiad.wmnet * 08:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1035.eqiad.wmnet * 08:34 phuedx@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324725{{!}}EventStreamConfig: Mark product_metrics.web_base and .web_base_with_ip as Test Kitchen streams (T429898 T430322)]], [[gerrit:1313923{{!}}EventStreamConfig: Remove unused web_ui_scroll* streams (T415370)]], [[gerrit:1325546{{!}}EventStreamConfig: Remove Watchlist click stream (T434790)]] (duration: 12m 42s) * 08:33 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1281: Pool back * 08:32 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage * 08:31 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on db2209.codfw.wmnet with reason: Maintenance * 08:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2209: Maintenance needed * 08:30 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2209: Maintenance needed * 08:26 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1035.eqiad.wmnet * 08:26 phuedx@deploy1003: bearloga, phuedx: Continuing with deployment * 08:23 phuedx@deploy1003: bearloga, phuedx: Backport for [[gerrit:1324725{{!}}EventStreamConfig: Mark product_metrics.web_base and .web_base_with_ip as Test Kitchen streams (T429898 T430322)]], [[gerrit:1313923{{!}}EventStreamConfig: Remove unused web_ui_scroll* streams (T415370)]], [[gerrit:1325546{{!}}EventStreamConfig: Remove Watchlist click stream (T434790)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug * 08:21 phuedx@deploy1003: Started scap sync-world: Backport for [[gerrit:1324725{{!}}EventStreamConfig: Mark product_metrics.web_base and .web_base_with_ip as Test Kitchen streams (T429898 T430322)]], [[gerrit:1313923{{!}}EventStreamConfig: Remove unused web_ui_scroll* streams (T415370)]], [[gerrit:1325546{{!}}EventStreamConfig: Remove Watchlist click stream (T434790)]] * 08:19 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw1001.wikimedia.org with OS trixie * 08:18 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1035.eqiad.wmnet * 08:17 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1032.eqiad.wmnet * 08:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1032.eqiad.wmnet * 08:16 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1279: Pool back * 08:15 phuedx@deploy1003: Finished scap sync-world: Backport for [[gerrit:1216721{{!}}viwikivoyage: enable relatedarticle and pop-up (T405724)]] (duration: 39m 12s) * 08:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1032.eqiad.wmnet * 08:09 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1032.eqiad.wmnet * 08:08 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1031.eqiad.wmnet * 08:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1031.eqiad.wmnet * 08:03 godog: switch production to use dumps-nfs.w.o - [[phab:T432212|T432212]] * 08:02 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1031.eqiad.wmnet * 08:02 phuedx@deploy1003: nvdtn19, phuedx: Continuing with deployment * 08:00 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1201: Depool db1201.eqiad.wmnet to then clone it to db1287.eqiad.wmnet - marostegui@cumin1003 * 08:00 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1201: Depool db1201.eqiad.wmnet to then clone it to db1287.eqiad.wmnet - marostegui@cumin1003 * 08:00 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1201.eqiad.wmnet onto db1287.eqiad.wmnet * 08:00 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1057.eqiad.wmnet * 08:00 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:00 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1057.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:59 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1057.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:59 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1275: Pool back * 07:55 filippo@cumin1003: START - Cookbook sre.dns.netbox * 07:55 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1031.eqiad.wmnet * 07:52 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1030.eqiad.wmnet * 07:52 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1030.eqiad.wmnet * 07:52 phuedx@deploy1003: nvdtn19, phuedx: Backport for [[gerrit:1216721{{!}}viwikivoyage: enable relatedarticle and pop-up (T405724)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:51 tappof: bump space for prometheus k8s-aux in eqiad * 07:50 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1057.eqiad.wmnet * 07:48 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1281: Pool back * 07:47 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1281 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96108 and previous config saved to /var/cache/conftool/dbconfig/20260817-074749-marostegui.json * 07:46 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1030.eqiad.wmnet * 07:42 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1030.eqiad.wmnet * 07:41 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1174: Depool db1174.eqiad.wmnet to then clone it to db1288.eqiad.wmnet - marostegui@cumin1003 * 07:41 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1174: Depool db1174.eqiad.wmnet to then clone it to db1288.eqiad.wmnet - marostegui@cumin1003 * 07:41 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1174.eqiad.wmnet onto db1288.eqiad.wmnet * 07:40 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1029.eqiad.wmnet * 07:40 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm2001.wikimedia.org * 07:40 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1029.eqiad.wmnet * 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1056.eqiad.wmnet * 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1056.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:38 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1056.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:36 phuedx@deploy1003: Started scap sync-world: Backport for [[gerrit:1216721{{!}}viwikivoyage: enable relatedarticle and pop-up (T405724)]] * 07:36 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm2001.wikimedia.org * 07:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1279 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96104 and previous config saved to /var/cache/conftool/dbconfig/20260817-073542-marostegui.json * 07:34 filippo@cumin1003: START - Cookbook sre.dns.netbox * 07:34 slyngshede@dns1004: END - running authdns-update * 07:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1029.eqiad.wmnet * 07:33 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm-test1001.wikimedia.org * 07:32 slyngshede@dns1004: START - running authdns-update * 07:31 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1029.eqiad.wmnet * 07:31 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1279: Pool back * 07:30 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1279 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96102 and previous config saved to /var/cache/conftool/dbconfig/20260817-073038-marostegui.json * 07:29 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm-test1001.wikimedia.org * 07:29 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm1001.wikimedia.org * 07:28 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1056.eqiad.wmnet * 07:28 moritzm: extend the disk of ldap-rw1001 by 80G [[phab:T331699|T331699]] * 07:28 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1055.eqiad.wmnet * 07:28 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:28 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1055.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:27 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1055.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:26 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1044.eqiad.wmnet * 07:26 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1044.eqiad.wmnet * 07:25 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm1001.wikimedia.org * 07:24 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1175: Depool db1175.eqiad.wmnet to then clone it to db1289.eqiad.wmnet - marostegui@cumin1003 * 07:24 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1175: Depool db1175.eqiad.wmnet to then clone it to db1289.eqiad.wmnet - marostegui@cumin1003 * 07:24 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1175.eqiad.wmnet onto db1289.eqiad.wmnet * 07:22 filippo@cumin1003: START - Cookbook sre.dns.netbox * 07:20 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1044.eqiad.wmnet * 07:16 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1055.eqiad.wmnet * 07:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1054.eqiad.wmnet * 07:15 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:15 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1054.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:15 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1044.eqiad.wmnet * 07:15 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1054.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:13 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1275: Pool back * 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1275 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96098 and previous config saved to /var/cache/conftool/dbconfig/20260817-071225-marostegui.json * 07:10 filippo@cumin1003: START - Cookbook sre.dns.netbox * 07:05 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1043.eqiad.wmnet * 07:05 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1054.eqiad.wmnet * 07:05 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1051.eqiad.wmnet * 07:05 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:05 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1051.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:05 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1043.eqiad.wmnet * 07:04 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1051.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 06:59 filippo@cumin1003: START - Cookbook sre.dns.netbox * 06:59 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin1001.eqiad.wmnet * 06:59 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1043.eqiad.wmnet * 06:59 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin2001.codfw.wmnet * 06:55 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin2001.codfw.wmnet * 06:55 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1051.eqiad.wmnet * 06:54 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1049.eqiad.wmnet * 06:54 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:54 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1049.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 06:54 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1049.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 06:54 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin1001.eqiad.wmnet * 06:53 moritzm: installing apr-util security updates * 06:52 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1043.eqiad.wmnet * 06:49 filippo@cumin1003: START - Cookbook sre.dns.netbox * 06:41 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1049.eqiad.wmnet * 06:13 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit2003.wikimedia.org * 06:13 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet * 06:07 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet * 06:06 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit2003.wikimedia.org * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 47s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-16 == * 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 01m 03s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-15 == * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 41s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-14 == * 15:38 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-staging-master-eqiad * 15:38 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster1005.eqiad.wmnet * 15:38 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster1005.eqiad.wmnet * 15:35 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sretest2009.codfw.wmnet * 15:33 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster1005.eqiad.wmnet * 15:33 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster1005.eqiad.wmnet * 15:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster1004.eqiad.wmnet * 15:32 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster1004.eqiad.wmnet * 15:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host sretest2009.codfw.wmnet * 15:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster1004.eqiad.wmnet * 15:27 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster1004.eqiad.wmnet * 15:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster1003.eqiad.wmnet * 15:27 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster1003.eqiad.wmnet * 15:24 dancy@deploy1003: Finished scap sync-world: testing (duration: 03m 23s) * 15:22 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster1003.eqiad.wmnet * 15:22 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster1003.eqiad.wmnet * 15:22 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-staging-master-eqiad * 15:20 dancy@deploy1003: Started scap sync-world: testing * 15:20 dancy@deploy1003: Installation of scap version "4.280.2" completed for 3 hosts * 15:18 dancy@deploy1003: Installing scap version "4.280.2" for 3 host(s) * 15:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sretest2006.codfw.wmnet * 14:54 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host sretest2006.codfw.wmnet * 14:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sretest2003.codfw.wmnet * 14:39 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host sretest2003.codfw.wmnet * 13:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt-staging2001.codfw.wmnet * 13:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt-staging2001.codfw.wmnet * 13:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-staging-master-codfw * 13:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster2005.codfw.wmnet * 13:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster2005.codfw.wmnet * 13:05 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox-dev2003.codfw.wmnet * 13:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster2005.codfw.wmnet * 13:04 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster2005.codfw.wmnet * 13:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster2004.codfw.wmnet * 13:04 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster2004.codfw.wmnet * 13:01 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netbox-dev2003.codfw.wmnet * 12:59 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster2004.codfw.wmnet * 12:59 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster2004.codfw.wmnet * 12:59 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster2003.codfw.wmnet * 12:59 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster2003.codfw.wmnet * 12:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster2003.codfw.wmnet * 12:54 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster2003.codfw.wmnet * 12:54 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-staging-master-codfw * 12:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-staging-worker-eqiad * 12:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage1006.eqiad.wmnet * 12:52 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage1006.eqiad.wmnet * 12:46 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage1006.eqiad.wmnet * 12:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw1001.wikimedia.org with OS trixie * 12:45 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage1006.eqiad.wmnet * 12:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage1005.eqiad.wmnet * 12:45 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage1005.eqiad.wmnet * 12:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage1005.eqiad.wmnet * 12:36 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/ratelimit: apply * 12:35 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/ratelimit: apply * 12:35 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:35 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:33 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage1005.eqiad.wmnet * 12:33 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage1004.eqiad.wmnet * 12:33 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage1004.eqiad.wmnet * 12:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage1004.eqiad.wmnet * 12:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage1004.eqiad.wmnet * 12:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage1003.eqiad.wmnet * 12:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage1003.eqiad.wmnet * 12:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage * 12:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage1003.eqiad.wmnet * 12:17 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage * 12:14 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage1003.eqiad.wmnet * 12:14 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-staging-worker-eqiad * 12:03 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw1001.wikimedia.org with OS trixie * 12:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cuminunpriv1001.eqiad.wmnet * 11:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cuminunpriv1001.eqiad.wmnet * 11:27 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1004.wikimedia.org * 11:24 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1153 from dbctl [[phab:T434638|T434638]]', diff saved to https://phabricator.wikimedia.org/P96097 and previous config saved to /var/cache/conftool/dbconfig/20260814-112449-marostegui.json * 11:21 aokoth@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1004.wikimedia.org * 11:20 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 11:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-staging-worker-codfw * 11:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2004.codfw.wmnet * 11:17 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2004.codfw.wmnet * 11:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2004.codfw.wmnet * 11:10 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2004.codfw.wmnet * 11:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2003.codfw.wmnet * 11:10 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2003.codfw.wmnet * 11:03 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-worker1181.eqiad.wmnet with OS bookworm * 11:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2003.codfw.wmnet * 11:03 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2003.codfw.wmnet * 11:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2002.codfw.wmnet * 11:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2002.codfw.wmnet * 10:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1235.eqiad.wmnet with OS bookworm * 10:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock2003.codfw.wmnet * 10:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock2003.codfw.wmnet with OS trixie * 10:56 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2002.codfw.wmnet * 10:54 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1153.eqiad.wmnet with OS bookworm * 10:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2002.codfw.wmnet * 10:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2001.codfw.wmnet * 10:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2001.codfw.wmnet * 10:50 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 10:49 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1187.eqiad.wmnet with OS bookworm * 10:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2001.codfw.wmnet * 10:44 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2001.codfw.wmnet * 10:44 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-staging-worker-codfw * 10:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock2003.codfw.wmnet with reason: host reimage * 10:38 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock2003.codfw.wmnet with reason: host reimage * 10:35 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1153.eqiad.wmnet with reason: host reimage * 10:32 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1235.eqiad.wmnet with reason: host reimage * 10:29 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1187.eqiad.wmnet with reason: host reimage * 10:24 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1153.eqiad.wmnet with reason: host reimage * 10:23 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1235.eqiad.wmnet with reason: host reimage * 10:21 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1187.eqiad.wmnet with reason: host reimage * 10:16 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock2003.codfw.wmnet with OS trixie * 10:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install1005.wikimedia.org * 10:12 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 10:12 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 10:12 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 10:12 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 10:12 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:12 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 10:12 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 10:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install1005.wikimedia.org * 10:07 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1235.eqiad.wmnet with OS bookworm * 10:07 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1187.eqiad.wmnet with OS bookworm * 10:07 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1181.eqiad.wmnet with OS bookworm * 10:07 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1153.eqiad.wmnet with OS bookworm * 10:07 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install2005.wikimedia.org * 10:03 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 10:03 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2003.codfw.wmnet * 10:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1232.eqiad.wmnet with OS bookworm * 10:00 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install2005.wikimedia.org * 10:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install3004.wikimedia.org * 09:58 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 09:58 fceratto@cumin1003: Removing db1151 from zarcillo [[phab:T434538|T434538]] * 09:56 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1231.eqiad.wmnet with OS bookworm * 09:56 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 09:53 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 09:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install3004.wikimedia.org * 09:50 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1152.eqiad.wmnet with OS bookworm * 09:49 Dreamy_Jazz: `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260808000000" --end-timestamp="20260812120000" --sleep="5" --batch-size="50"` for [[phab:T434688|T434688]] * 09:48 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host centrallog2002.codfw.wmnet * 09:48 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install4004.wikimedia.org * 09:42 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1232.eqiad.wmnet with reason: host reimage * 09:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install4004.wikimedia.org * 09:41 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host centrallog2002.codfw.wmnet * 09:39 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install5004.wikimedia.org * 09:36 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1232.eqiad.wmnet with reason: host reimage * 09:36 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'. * 09:34 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'. * 09:33 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1231.eqiad.wmnet with reason: host reimage * 09:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install5004.wikimedia.org * 09:32 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host centrallog1002.eqiad.wmnet * 09:30 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install6003.wikimedia.org * 09:30 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1231.eqiad.wmnet with reason: host reimage * 09:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1152.eqiad.wmnet with reason: host reimage * 09:25 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host centrallog1002.eqiad.wmnet * 09:25 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1152.eqiad.wmnet with reason: host reimage * 09:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install6003.wikimedia.org * 09:22 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1232 * 09:22 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1232 * 09:22 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1232 * 09:22 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1232.eqiad.wmnet 25.53.64.10.in-addr.arpa 5.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:22 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host titan1001.eqiad.wmnet * 09:22 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1232.eqiad.wmnet 25.53.64.10.in-addr.arpa 5.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:22 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:22 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1232 - btullis@cumin1003" * 09:22 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1232 - btullis@cumin1003" * 09:17 btullis@cumin1003: START - Cookbook sre.dns.netbox * 09:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install7002.wikimedia.org * 09:17 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1232 * 09:16 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1231 * 09:16 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1231 * 09:14 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1231 * 09:14 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1231.eqiad.wmnet 24.53.64.10.in-addr.arpa 4.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host titan1001.eqiad.wmnet * 09:14 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1231.eqiad.wmnet 24.53.64.10.in-addr.arpa 4.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:14 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1231 - btullis@cumin1003" * 09:14 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1231 - btullis@cumin1003" * 09:10 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install7002.wikimedia.org * 09:09 btullis@cumin1003: START - Cookbook sre.dns.netbox * 09:08 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1231 * 09:08 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1152 * 09:08 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1152 * 09:06 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1152 * 09:06 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1152.eqiad.wmnet 16.53.64.10.in-addr.arpa 6.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:06 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1152.eqiad.wmnet 16.53.64.10.in-addr.arpa 6.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:06 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:06 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1152 - btullis@cumin1003" * 09:06 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1152 - btullis@cumin1003" * 09:03 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1151.eqiad.wmnet * 09:03 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:03 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1151.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 08:55 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host titan2001.codfw.wmnet * 08:55 btullis@cumin1003: START - Cookbook sre.dns.netbox * 08:54 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1151.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 08:54 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1152 * 08:53 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1232.eqiad.wmnet with OS bookworm * 08:53 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1231.eqiad.wmnet with OS bookworm * 08:53 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1152.eqiad.wmnet with OS bookworm * 08:51 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2209: Pool back * 08:51 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1201.eqiad.wmnet * 08:50 btullis@cumin1003: START - Cookbook sre.hosts.remove-downtime for an-worker1201.eqiad.wmnet * 08:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1229.eqiad.wmnet with OS bookworm * 08:47 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host titan2001.codfw.wmnet * 08:46 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping2004.codfw.wmnet * 08:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ping2004.codfw.wmnet * 08:40 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host an-worker1230.eqiad.wmnet with OS bookworm * 08:40 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 08:34 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1151.eqiad.wmnet * 08:34 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 08:31 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host titan1002.eqiad.wmnet * 08:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1229.eqiad.wmnet with reason: host reimage * 08:25 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host titan1002.eqiad.wmnet * 08:25 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1229.eqiad.wmnet with reason: host reimage * 08:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping1004.eqiad.wmnet * 08:21 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ping1004.eqiad.wmnet * 08:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1230.eqiad.wmnet with reason: host reimage * 08:12 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1230.eqiad.wmnet with reason: host reimage * 08:11 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1229 * 08:11 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1229 * 08:11 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1229 * 08:11 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1229.eqiad.wmnet 22.53.64.10.in-addr.arpa 2.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:11 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1229.eqiad.wmnet 22.53.64.10.in-addr.arpa 2.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:11 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:11 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1229 - btullis@cumin1003" * 08:11 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1229 - btullis@cumin1003" * 08:10 btullis@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1201.eqiad.wmnet with reason: Fixing a disk * 08:07 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host titan2002.codfw.wmnet * 08:05 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2209: Pool back * 08:05 btullis@cumin1003: START - Cookbook sre.dns.netbox * 08:00 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host titan2002.codfw.wmnet * 07:59 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1229 * 07:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1230 * 07:58 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1230 * 07:55 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1230 * 07:55 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1230.eqiad.wmnet 23.53.64.10.in-addr.arpa 3.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:55 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1230.eqiad.wmnet 23.53.64.10.in-addr.arpa 3.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:55 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:55 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1230 - btullis@cumin1003" * 07:55 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1230 - btullis@cumin1003" * 07:51 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host kubestagemaster2005.codfw.wmnet with OS trixie * 07:48 btullis@cumin1003: START - Cookbook sre.dns.netbox * 07:41 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1230 * 07:41 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1229.eqiad.wmnet with OS bookworm * 07:41 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1230.eqiad.wmnet with OS bookworm * 07:39 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1152 from dbctl [[phab:T434480|T434480]]', diff saved to https://phabricator.wikimedia.org/P96090 and previous config saved to /var/cache/conftool/dbconfig/20260814-073941-marostegui.json * 07:29 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on kubestagemaster2005.codfw.wmnet with reason: host reimage * 07:23 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on kubestagemaster2005.codfw.wmnet with reason: host reimage * 07:04 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host kubestagemaster2005.codfw.wmnet with OS trixie * 06:53 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1228.eqiad.wmnet with OS bookworm * 06:44 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1227.eqiad.wmnet with OS bookworm * 06:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1209.eqiad.wmnet with OS bookworm * 06:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1175.eqiad.wmnet with OS bookworm * 06:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1228.eqiad.wmnet with reason: host reimage * 06:27 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1228.eqiad.wmnet with reason: host reimage * 06:25 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1227.eqiad.wmnet with reason: host reimage * 06:21 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1227.eqiad.wmnet with reason: host reimage * 06:18 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1209.eqiad.wmnet with reason: host reimage * 06:14 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1209.eqiad.wmnet with reason: host reimage * 06:14 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1228 * 06:14 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1228 * 06:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1175.eqiad.wmnet with reason: host reimage * 06:12 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1228 * 06:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1228.eqiad.wmnet 20.53.64.10.in-addr.arpa 0.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:12 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1228.eqiad.wmnet 20.53.64.10.in-addr.arpa 0.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1228 - ryankemper@cumin2003" * 06:12 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1228 - ryankemper@cumin2003" * 06:09 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1175.eqiad.wmnet with reason: host reimage * 06:07 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 06:07 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1228 * 06:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1227 * 06:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1227 * 06:06 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1227 * 06:06 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1227.eqiad.wmnet 19.53.64.10.in-addr.arpa 9.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:06 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1227.eqiad.wmnet 19.53.64.10.in-addr.arpa 9.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:06 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:06 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1227 - ryankemper@cumin2003" * 06:06 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1227 - ryankemper@cumin2003" * 06:00 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 06:00 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1227 * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1209 * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1209 * 06:00 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1209 * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1209.eqiad.wmnet 15.53.64.10.in-addr.arpa 5.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:00 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1209.eqiad.wmnet 15.53.64.10.in-addr.arpa 5.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1209 - ryankemper@cumin2003" * 06:00 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1209 - ryankemper@cumin2003" * 05:54 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 05:54 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1209 * 05:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1175 * 05:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1175 * 05:52 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1175 * 05:52 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1175.eqiad.wmnet 17.53.64.10.in-addr.arpa 7.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 05:52 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1175.eqiad.wmnet 17.53.64.10.in-addr.arpa 7.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 05:52 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 05:52 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1175 - ryankemper@cumin2003" * 05:52 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1175 - ryankemper@cumin2003" * 05:49 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1228.eqiad.wmnet with OS bookworm * 05:49 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1227.eqiad.wmnet with OS bookworm * 05:48 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1209.eqiad.wmnet with OS bookworm * 05:47 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 05:47 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1175 * 05:47 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1175.eqiad.wmnet with OS bookworm * 05:09 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host apifeatureusage1001.eqiad.wmnet with OS bookworm * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 03s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:10 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 01:07 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 01:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 01:02 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 00:59 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1226.eqiad.wmnet with OS bookworm * 00:47 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1225.eqiad.wmnet with OS bookworm * 00:41 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1224.eqiad.wmnet with OS bookworm * 00:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1226.eqiad.wmnet with reason: host reimage * 00:31 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1226.eqiad.wmnet with reason: host reimage * 00:28 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1225.eqiad.wmnet with reason: host reimage * 00:25 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1225.eqiad.wmnet with reason: host reimage * 00:19 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1224.eqiad.wmnet with reason: host reimage * 00:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1226 * 00:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1226 * 00:17 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1226 * 00:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1226.eqiad.wmnet 23.36.64.10.in-addr.arpa 3.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:17 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1226.eqiad.wmnet 23.36.64.10.in-addr.arpa 3.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 00:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1226 - ryankemper@cumin2003" * 00:17 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1226 - ryankemper@cumin2003" * 00:16 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1224.eqiad.wmnet with reason: host reimage * 00:12 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 00:12 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1226 * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1225 * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1225 * 00:10 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1225 * 00:10 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1225.eqiad.wmnet 22.36.64.10.in-addr.arpa 2.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:10 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1225.eqiad.wmnet 22.36.64.10.in-addr.arpa 2.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:10 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 00:10 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1225 - ryankemper@cumin2003" * 00:10 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1225 - ryankemper@cumin2003" * 00:03 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 00:02 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1225 * 00:02 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1224 * 00:02 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1224 * 00:00 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1224 * 00:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1224.eqiad.wmnet 21.36.64.10.in-addr.arpa 1.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:00 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1224.eqiad.wmnet 21.36.64.10.in-addr.arpa 1.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 00:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1224 - ryankemper@cumin2003" * 00:00 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1224 - ryankemper@cumin2003" == 2026-08-13 == * 23:54 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1226.eqiad.wmnet with OS bookworm * 23:53 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1225.eqiad.wmnet with OS bookworm * 23:52 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 23:51 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1224 * 23:51 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1224.eqiad.wmnet with OS bookworm * 23:47 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1223.eqiad.wmnet with OS bookworm * 23:29 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1223.eqiad.wmnet with reason: host reimage * 23:24 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1223.eqiad.wmnet with reason: host reimage * 23:19 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325525{{!}}ve.ui.CodeMirror.less: ensure normal font style]] (duration: 11m 40s) * 23:16 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260807000000" --end-timestamp="20260808000000" --sleep="5" --batch-size="50"` for [[phab:T434688|T434688]] * 23:13 musikanimal@deploy1003: musikanimal: Continuing with deployment * 23:11 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1325525{{!}}ve.ui.CodeMirror.less: ensure normal font style]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1223 * 23:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1223 * 23:08 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1325525{{!}}ve.ui.CodeMirror.less: ensure normal font style]] * 23:07 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1223 * 23:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1223.eqiad.wmnet 20.36.64.10.in-addr.arpa 0.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:07 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1223.eqiad.wmnet 20.36.64.10.in-addr.arpa 0.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 23:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1223 - ryankemper@cumin2003" * 23:03 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1223 - ryankemper@cumin2003" * 23:01 sbassett: Deployed security updates for [[phab:T430596|T430596]], [[phab:T120386|T120386]] * 22:55 bking@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host kubestagemaster2005.codfw.wmnet * 22:55 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host kubestagemaster2005.codfw.wmnet with OS bookworm * 22:54 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 22:53 ryankemper@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 22:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on kubestagemaster2005.codfw.wmnet with reason: host reimage * 22:49 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1222.eqiad.wmnet with OS bookworm * 22:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on kubestagemaster2005.codfw.wmnet with reason: host reimage * 22:29 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1222.eqiad.wmnet with reason: host reimage * 22:26 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1222.eqiad.wmnet with reason: host reimage * 22:23 bking@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host aux-k8s-etcd2003.codfw.wmnet * 22:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd2003.codfw.wmnet with OS bookworm * 22:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host kubestagemaster2005.codfw.wmnet with OS bookworm * 22:23 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM kubestagemaster2005.codfw.wmnet - bking@cumin2003" * 22:23 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM kubestagemaster2005.codfw.wmnet - bking@cumin2003" * 22:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) kubestagemaster2005.codfw.wmnet on all recursors * 22:22 bking@cumin2003: START - Cookbook sre.dns.wipe-cache kubestagemaster2005.codfw.wmnet on all recursors * 22:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:22 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM kubestagemaster2005.codfw.wmnet - bking@cumin2003" * 22:22 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM kubestagemaster2005.codfw.wmnet - bking@cumin2003" * 22:17 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 22:13 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1223 * 22:12 bking@cumin2003: START - Cookbook sre.dns.netbox * 22:12 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host kubestagemaster2005.codfw.wmnet * 22:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1222 * 22:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1222 * 22:11 sbassett: Deployed security updates for [[phab:T429244|T429244]], [[phab:T434039|T434039]] * 22:11 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1222 * 22:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1222.eqiad.wmnet 19.36.64.10.in-addr.arpa 9.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:11 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1222.eqiad.wmnet 19.36.64.10.in-addr.arpa 9.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1222 - ryankemper@cumin2003" * 22:07 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1222 - ryankemper@cumin2003" * 22:05 robh@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-wdqs2001.codfw.wmnet with reason: updating firmware * 22:01 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 22:01 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1223.eqiad.wmnet with OS bookworm * 22:01 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1222 * 22:01 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1222.eqiad.wmnet with OS bookworm * 22:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1212.eqiad.wmnet with OS bookworm * 21:54 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324295{{!}}Revert "Lazily reject pre-fix parser-cache entries for noreferrer/noopener links" (T429090)]] (duration: 06m 42s) * 21:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd2003.codfw.wmnet with reason: host reimage * 21:53 bking@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host dse-k8s-etcd2001.codfw.wmnet * 21:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dse-k8s-etcd2001.codfw.wmnet with OS bookworm * 21:50 sbassett@deploy1003: sbassett, kharlan: Continuing with deployment * 21:49 sbassett@deploy1003: sbassett, kharlan: Backport for [[gerrit:1324295{{!}}Revert "Lazily reject pre-fix parser-cache entries for noreferrer/noopener links" (T429090)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:48 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aux-k8s-etcd2003.codfw.wmnet with reason: host reimage * 21:47 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1324295{{!}}Revert "Lazily reject pre-fix parser-cache entries for noreferrer/noopener links" (T429090)]] * 21:42 maryum: Deployed security patch for [[phab:T434549|T434549]] * 21:39 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1212.eqiad.wmnet with reason: host reimage * 21:34 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1212.eqiad.wmnet with reason: host reimage * 21:33 bking@cumin2003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd2003.codfw.wmnet with OS bookworm * 21:32 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM aux-k8s-etcd2003.codfw.wmnet - bking@cumin2003" * 21:32 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM aux-k8s-etcd2003.codfw.wmnet - bking@cumin2003" * 21:32 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) aux-k8s-etcd2003.codfw.wmnet on all recursors * 21:32 bking@cumin2003: START - Cookbook sre.dns.wipe-cache aux-k8s-etcd2003.codfw.wmnet on all recursors * 21:32 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:32 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM aux-k8s-etcd2003.codfw.wmnet - bking@cumin2003" * 21:31 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM aux-k8s-etcd2003.codfw.wmnet - bking@cumin2003" * 21:28 maryum: Deployed security patch for [[phab:T434619|T434619]] * 21:26 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:26 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host aux-k8s-etcd2003.codfw.wmnet * 21:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-etcd2001.codfw.wmnet with reason: host reimage * 21:20 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1212 * 21:20 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1212 * 21:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1212.eqiad.wmnet with OS bookworm * 21:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on dse-k8s-etcd2001.codfw.wmnet with reason: host reimage * 21:04 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325553{{!}}InstrumentConstructiveEdits: exclude mw-reverted as well (T431493)]] (duration: 06m 25s) * 20:59 kemayo@deploy1003: kemayo: Continuing with deployment * 20:59 kemayo@deploy1003: kemayo: Backport for [[gerrit:1325553{{!}}InstrumentConstructiveEdits: exclude mw-reverted as well (T431493)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host dse-k8s-etcd2001.codfw.wmnet with OS bookworm * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 20:57 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 20:57 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1325553{{!}}InstrumentConstructiveEdits: exclude mw-reverted as well (T431493)]] * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-etcd2001.codfw.wmnet on all recursors * 20:57 bking@cumin2003: START - Cookbook sre.dns.wipe-cache dse-k8s-etcd2001.codfw.wmnet on all recursors * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 20:57 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 20:54 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320163{{!}}Add configurable RestTermsOfServiceUrl (T428147)]] (duration: 21m 39s) * 20:53 bking@cumin2003: START - Cookbook sre.dns.netbox * 20:53 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host dse-k8s-etcd2001.codfw.wmnet * 20:50 samtar@deploy1003: samtar, milazg: Continuing with deployment * 20:35 samtar@deploy1003: samtar, milazg: Backport for [[gerrit:1320163{{!}}Add configurable RestTermsOfServiceUrl (T428147)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:33 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1320163{{!}}Add configurable RestTermsOfServiceUrl (T428147)]] * 20:30 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325549{{!}}Deploy PRV to several LC wikis (T423785)]] (duration: 06m 57s) * 20:26 arlolra@deploy1003: arlolra: Continuing with deployment * 20:25 arlolra@deploy1003: arlolra: Backport for [[gerrit:1325549{{!}}Deploy PRV to several LC wikis (T423785)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:24 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 20:23 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1325549{{!}}Deploy PRV to several LC wikis (T423785)]] * 20:21 ariel@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319906{{!}}Remove boilerplate language from wmf-rest and wmf-math API modules (T433736)]] (duration: 13m 54s) * 20:14 ariel@deploy1003: ariel: Continuing with deployment * 20:11 ariel@deploy1003: ariel: Backport for [[gerrit:1319906{{!}}Remove boilerplate language from wmf-rest and wmf-math API modules (T433736)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 ariel@deploy1003: Started scap sync-world: Backport for [[gerrit:1319906{{!}}Remove boilerplate language from wmf-rest and wmf-math API modules (T433736)]] * 19:58 robh@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-wdqs2001.codfw.wmnet with reason: updating firmware * 19:54 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325551{{!}}Render the focused module view as a full-screen page (T433896)]] (duration: 30m 37s) * 19:52 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching aqs[2001,1016]*: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 19:44 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching aqs[2001,1016]*: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 19:42 musikanimal@deploy1003: musikanimal: Continuing with deployment * 19:41 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1325551{{!}}Render the focused module view as a full-screen page (T433896)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:34 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: quash java safepoint logspam - bking@cumin2003 - [[phab:T434685|T434685]] * 19:34 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 19:34 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 19:24 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1325551{{!}}Render the focused module view as a full-screen page (T433896)]] * 19:20 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 19:20 jhancock@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin1003" * 19:18 jhancock@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin1003" * 19:14 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325548{{!}}Enable image lazy loading on desktop in group1 (T148047)]] (duration: 07m 43s) * 19:10 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 19:10 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1325548{{!}}Enable image lazy loading on desktop in group1 (T148047)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:07 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1325548{{!}}Enable image lazy loading on desktop in group1 (T148047)]] * 19:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1166.eqiad.wmnet onto db1280.eqiad.wmnet * 19:03 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 19:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1166: Pool db1166.eqiad.wmnet in after cloning * 18:59 jhancock@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 18:58 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1234.eqiad.wmnet with OS bookworm * 18:52 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 18:51 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:49 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:46 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 18:46 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 18:44 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=dns3004.* * 18:39 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1221.eqiad.wmnet with OS bookworm * 18:36 inflatador: [bking@ganeti2048] ~$ sudo gnt-instance replace-disks -n ganeti2030.codfw.wmnet aux-k8s-worker2002.codfw.wmnet [[phab:T434681|T434681]] * 18:35 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 18:29 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1220.eqiad.wmnet with OS bookworm * 18:29 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1234.eqiad.wmnet with reason: host reimage * 18:27 bking@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host dse-k8s-etcd2001.codfw.wmnet * 18:27 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-etcd2001.codfw.wmnet on all recursors * 18:27 bking@cumin2003: START - Cookbook sre.dns.wipe-cache dse-k8s-etcd2001.codfw.wmnet on all recursors * 18:27 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:27 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 18:27 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 18:25 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1234.eqiad.wmnet with reason: host reimage * 18:25 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 18:19 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1221.eqiad.wmnet with reason: host reimage * 18:17 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1166: Pool db1166.eqiad.wmnet in after cloning * 18:15 brennen@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 18:15 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1221.eqiad.wmnet with reason: host reimage * 18:14 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:12 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:12 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 18:12 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-etcd2001.codfw.wmnet on all recursors * 18:12 bking@cumin2003: START - Cookbook sre.dns.wipe-cache dse-k8s-etcd2001.codfw.wmnet on all recursors * 18:12 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:12 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 18:12 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 18:10 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: quash java safepoint logspam - bking@cumin2003 - [[phab:T434685|T434685]] * 18:10 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1234 * 18:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1234 * 18:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1220.eqiad.wmnet with reason: host reimage * 18:07 brennen: 1.47.0-wmf.15 train status ([[phab:T430834|T430834]]) - no current blockers, rolling to all wikis * 18:07 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1234 * 18:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1234.eqiad.wmnet 10.36.64.10.in-addr.arpa 0.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:07 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:07 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1234.eqiad.wmnet 10.36.64.10.in-addr.arpa 0.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1234 - ryankemper@cumin2003" * 18:07 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1234 - ryankemper@cumin2003" * 18:05 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1220.eqiad.wmnet with reason: host reimage * 18:04 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host dse-k8s-etcd2001.codfw.wmnet * 18:04 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 18:03 dancy@deploy1003: Installation of scap version "4.280.1" completed for 3 hosts * 18:02 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:02 inflatador: bking@dse-k8s-etcd2002 etcdctl member remove $<nowiki>{</nowiki>UUID of dse-k8s-etcd2001<nowiki>}</nowiki> [[phab:T434681|T434681]] [[phab:T434793|T434793]] * 18:01 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 18:01 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1234 * 18:01 dancy@deploy1003: Installing scap version "4.280.1" for 3 host(s) * 18:01 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1221 * 18:01 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1221 * 18:00 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1221 * 18:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1221.eqiad.wmnet 18.36.64.10.in-addr.arpa 8.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:00 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1221.eqiad.wmnet 18.36.64.10.in-addr.arpa 8.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1221 - ryankemper@cumin2003" * 17:59 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:58 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 17:58 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:57 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 17:57 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:56 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1221 - ryankemper@cumin2003" * 17:53 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 17:52 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 17:52 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 17:51 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1221 * 17:51 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1220 * 17:51 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1220 * 17:51 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:51 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:51 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1220 * 17:51 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1220.eqiad.wmnet 11.36.64.10.in-addr.arpa 1.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:51 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1220.eqiad.wmnet 11.36.64.10.in-addr.arpa 1.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:51 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1220 - ryankemper@cumin2003" * 17:50 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1220 - ryankemper@cumin2003" * 17:47 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:47 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 17:46 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1234.eqiad.wmnet with OS bookworm * 17:46 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 17:45 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1221.eqiad.wmnet with OS bookworm * 17:45 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1220 * 17:45 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1220.eqiad.wmnet with OS bookworm * 17:43 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 17:41 bking@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host dse-k8s-etcd2001.codfw.wmnet * 17:41 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host dse-k8s-etcd2001.codfw.wmnet * 17:40 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:40 inflatador: bking@ganeti2048] `sudo gnt-instance remove --force --ignore-failures --shutdown-timeout=0` on non-DRBD VMs [[phab:T434681|T434681]] * 17:40 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:39 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:38 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:36 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:32 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:32 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:28 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 17:26 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:26 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:24 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:23 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:21 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1218.eqiad.wmnet with OS bookworm * 17:20 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:20 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:19 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1179.eqiad.wmnet with OS bookworm * 17:18 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1150.eqiad.wmnet with OS bookworm * 17:18 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns3004.wikimedia.org with OS trixie * 17:15 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:14 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:13 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:12 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:12 swfrench@deploy1003: Finished scap sync-world: Helmfile-only deployment for mediawiki chart bump - [[phab:T427666|T427666]] (duration: 03m 03s) * 17:09 swfrench@deploy1003: Started scap sync-world: Helmfile-only deployment for mediawiki chart bump - [[phab:T427666|T427666]] * 17:02 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:01 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:01 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1218.eqiad.wmnet with reason: host reimage * 17:01 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:01 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:00 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324810{{!}}deployment-info.php: Report dbname and branch for the requested wiki (T434726)]] (duration: 06m 52s) * 16:58 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 16:57 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1179.eqiad.wmnet with reason: host reimage * 16:56 dancy@deploy1003: dancy: Continuing with deployment * 16:56 dancy@deploy1003: dancy: Backport for [[gerrit:1324810{{!}}deployment-info.php: Report dbname and branch for the requested wiki (T434726)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1150.eqiad.wmnet with reason: host reimage * 16:53 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324810{{!}}deployment-info.php: Report dbname and branch for the requested wiki (T434726)]] * 16:51 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1166: Depool db1166.eqiad.wmnet to then clone it to db1280.eqiad.wmnet - cwilliams@cumin1003 * 16:50 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1166: Depool db1166.eqiad.wmnet to then clone it to db1280.eqiad.wmnet - cwilliams@cumin1003 * 16:50 cwilliams@cumin1003: START - Cookbook sre.mysql.clone of db1166.eqiad.wmnet onto db1280.eqiad.wmnet * 16:49 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1218.eqiad.wmnet with reason: host reimage * 16:48 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1179.eqiad.wmnet with reason: host reimage * 16:47 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1150.eqiad.wmnet with reason: host reimage * 16:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.upgrade (exit_code=0) for 1 hosts * 16:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2220: Upgrade of db2220.codfw.wmnet completed * 16:38 swfrench@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 16:38 swfrench-wmf: kubectl delete node kubestagemaster2005.codfw.wmnet - [[phab:T434681|T434681]] * 16:34 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1218 * 16:34 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1218 * 16:34 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1218.eqiad.wmnet with OS bookworm * 16:34 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1179 * 16:34 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1179 * 16:33 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1179.eqiad.wmnet with OS bookworm * 16:32 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1150 * 16:32 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1150 * 16:31 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1150.eqiad.wmnet with OS bookworm * 16:24 dancy@deploy1003: Installation of scap version "4.280.0" completed for 3 hosts * 16:22 dancy@deploy1003: Installing scap version "4.280.0" for 3 host(s) * 16:20 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: quash java safepoint logspam - bking@cumin2003 - [[phab:T434685|T434685]] * 16:14 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns3004.wikimedia.org with reason: host reimage * 16:08 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 16:07 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 16:07 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns3004.wikimedia.org with reason: host reimage * 16:05 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 16:04 swfrench@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 16:00 swfrench@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 15:59 swfrench@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 15:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Upgrade of db2220.codfw.wmnet completed * 15:48 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2220: Upgrading db2220.codfw.wmnet * 15:48 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2220: Upgrading db2220.codfw.wmnet * 15:48 cwilliams@cumin1003: START - Cookbook sre.mysql.upgrade for 1 hosts * 15:46 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns3004.wikimedia.org with OS trixie * 15:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2220 [[phab:T434802|T434802]]', diff saved to https://phabricator.wikimedia.org/P96079 and previous config saved to /var/cache/conftool/dbconfig/20260813-154624-cwilliams.json * 15:45 cdobbins@cumin1003: conftool action : set/pooled=no; selector: name=dns3004.* * 15:44 cjd91: depooling dns3004 to reimage to trixie * 15:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2159 to s7 primary [[phab:T434802|T434802]]', diff saved to https://phabricator.wikimedia.org/P96078 and previous config saved to /var/cache/conftool/dbconfig/20260813-154405-cwilliams.json * 15:43 cezmunsta: Starting s7 codfw failover from db2220 to db2159 - [[phab:T434802|T434802]] * 15:41 cgoubert@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: Dragonfly supernodes reboot (duration: 09m 42s) * 15:41 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dragonfly-supernode2001.codfw.wmnet * 15:39 inflatador: bking@ganeti2048] ~$ sudo gnt-node failover -f --ignore-consistency ganeti2046.codfw.wmnet [[phab:T434681|T434681]] * 15:39 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 15:39 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 15:39 swfrench@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 15:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2159 with weight 0 [[phab:T434802|T434802]]', diff saved to https://phabricator.wikimedia.org/P96077 and previous config saved to /var/cache/conftool/dbconfig/20260813-153806-cwilliams.json * 15:37 swfrench@dns1004: END - running authdns-update * 15:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 30 hosts with reason: Primary switchover s7 [[phab:T434802|T434802]] * 15:37 cgoubert@cumin2003: START - Cookbook sre.hosts.reboot-single for host dragonfly-supernode2001.codfw.wmnet * 15:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dragonfly-supernode1001.eqiad.wmnet * 15:35 swfrench@dns1004: START - running authdns-update * 15:32 cgoubert@cumin2003: START - Cookbook sre.hosts.reboot-single for host dragonfly-supernode1001.eqiad.wmnet * 15:32 cgoubert@deploy1003: Locking from deployment [ALL REPOSITORIES]: Dragonfly supernodes reboot * 15:30 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-master-codfw * 15:30 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2005.codfw.wmnet * 15:30 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2005.codfw.wmnet * 15:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1169.eqiad.wmnet onto db1277.eqiad.wmnet * 15:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1169: Pool db1169.eqiad.wmnet in after cloning * 15:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2005.codfw.wmnet * 15:23 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2005.codfw.wmnet * 15:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2004.codfw.wmnet * 15:23 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2004.codfw.wmnet * 15:18 inflatador: bking@ganeti2048 sudo gnt-node failover -f ganeti2046.codfw.wmnet [[phab:T434681|T434681]] * 15:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2004.codfw.wmnet * 15:17 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2004.codfw.wmnet * 15:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2003.codfw.wmnet * 15:17 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2003.codfw.wmnet * 15:15 cdobbins@dns1004: END - running authdns-update * 15:13 cdobbins@dns1004: START - running authdns-update * 15:10 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1217.eqiad.wmnet with OS bookworm * 15:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2003.codfw.wmnet * 15:10 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2003.codfw.wmnet * 15:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2002.codfw.wmnet * 15:10 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2002.codfw.wmnet * 15:10 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: quash java safepoint logspam - bking@cumin2003 - [[phab:T434685|T434685]] * 15:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1216.eqiad.wmnet with OS bookworm * 15:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2002.codfw.wmnet * 15:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2002.codfw.wmnet * 15:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2001.codfw.wmnet * 15:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2001.codfw.wmnet * 15:01 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1042.eqiad.wmnet * 15:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1042.eqiad.wmnet * 15:00 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325455{{!}}Api: Use correct query when continue prop=categories (T433922)]] (duration: 09m 47s) * 14:59 bking@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host dse-k8s-etcd2004.codfw.wmnet * 14:58 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-etcd2004.codfw.wmnet on all recursors * 14:58 bking@cumin2003: START - Cookbook sre.dns.wipe-cache dse-k8s-etcd2004.codfw.wmnet on all recursors * 14:58 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:58 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM dse-k8s-etcd2004.codfw.wmnet - bking@cumin2003" * 14:58 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM dse-k8s-etcd2004.codfw.wmnet - bking@cumin2003" * 14:55 zabe@deploy1003: zabe: Continuing with deployment * 14:53 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-etcd2004.codfw.wmnet on all recursors * 14:53 bking@cumin2003: START - Cookbook sre.dns.wipe-cache dse-k8s-etcd2004.codfw.wmnet on all recursors * 14:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:53 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2004.codfw.wmnet - bking@cumin2003" * 14:53 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2004.codfw.wmnet - bking@cumin2003" * 14:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2001.codfw.wmnet * 14:52 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2001.codfw.wmnet * 14:52 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-master-codfw * 14:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2001.codfw.wmnet * 14:52 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2001.codfw.wmnet * 14:52 zabe@deploy1003: zabe: Backport for [[gerrit:1325455{{!}}Api: Use correct query when continue prop=categories (T433922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:50 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1325455{{!}}Api: Use correct query when continue prop=categories (T433922)]] * 14:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1217.eqiad.wmnet with reason: host reimage * 14:48 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:48 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host dse-k8s-etcd2004.codfw.wmnet * 14:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-master-eqiad * 14:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1006.eqiad.wmnet * 14:44 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1006.eqiad.wmnet * 14:44 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1216.eqiad.wmnet with reason: host reimage * 14:40 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1217.eqiad.wmnet with reason: host reimage * 14:39 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1169: Pool db1169.eqiad.wmnet in after cloning * 14:39 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1216.eqiad.wmnet with reason: host reimage * 14:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl1005.eqiad.wmnet * 14:32 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl1005.eqiad.wmnet * 14:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1004.eqiad.wmnet * 14:32 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1004.eqiad.wmnet * 14:31 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus2008.codfw.wmnet * 14:31 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor1003.eqiad.wmnet * 14:29 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1042.eqiad.wmnet * 14:27 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor1003.eqiad.wmnet * 14:26 moritzm: installing Django security updates * 14:26 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling reboot on A:wikidough * 14:25 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1217.eqiad.wmnet with OS bookworm * 14:25 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1216.eqiad.wmnet with OS bookworm * 14:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl1004.eqiad.wmnet * 14:24 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl1004.eqiad.wmnet * 14:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1003.eqiad.wmnet * 14:24 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1003.eqiad.wmnet * 14:23 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus2008.codfw.wmnet * 14:23 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus1008.eqiad.wmnet * 14:23 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor-dev2001.codfw.wmnet * 14:22 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus2006.codfw.wmnet * 14:19 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor-dev2001.codfw.wmnet * 14:18 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1042.eqiad.wmnet * 14:17 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.* * 14:16 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl1003.eqiad.wmnet * 14:16 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl1003.eqiad.wmnet * 14:16 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1002.eqiad.wmnet * 14:16 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1002.eqiad.wmnet * 14:15 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus1008.eqiad.wmnet * 14:14 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus1006.eqiad.wmnet * 14:12 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus2006.codfw.wmnet * 14:12 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor2003.codfw.wmnet * 14:11 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus2007.codfw.wmnet * 14:11 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1041.eqiad.wmnet * 14:11 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1041.eqiad.wmnet * 14:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl1002.eqiad.wmnet * 14:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl1002.eqiad.wmnet * 14:09 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-master-eqiad * 14:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on apifeatureusage1001.eqiad.wmnet with reason: host reimage * 14:08 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1215.eqiad.wmnet with OS bookworm * 14:08 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor2003.codfw.wmnet * 14:08 jayme: updated calico to v3.30.7 on wikikube codfw [[phab:T427400|T427400]] * 14:07 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1149.eqiad.wmnet with OS bookworm * 14:06 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'. * 14:06 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1041.eqiad.wmnet * 14:04 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus1006.eqiad.wmnet * 14:03 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus2007.codfw.wmnet * 14:03 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus2005.codfw.wmnet * 14:03 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus1007.eqiad.wmnet * 14:03 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1002.eqiad.wmnet * 14:02 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on apifeatureusage1001.eqiad.wmnet with reason: host reimage * 14:02 cgoubert@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=helm-charts.*,name=eqiad * 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host chartmuseum1001.eqiad.wmnet * 14:00 moritzm: installing libxml2 security updates * 13:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1214.eqiad.wmnet with OS bookworm * 13:59 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns5003.* * 13:59 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1041.eqiad.wmnet * 13:59 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow1002.eqiad.wmnet * 13:58 cmooney@dns3003: END - running authdns-update * 13:57 cgoubert@cumin2003: START - Cookbook sre.hosts.reboot-single for host chartmuseum1001.eqiad.wmnet * 13:57 cgoubert@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=helm-charts.*,name=eqiad * 13:57 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: cloudelastic cluster restart - bking@cumin2003 * 13:57 cgoubert@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=helm-charts.*,name=codfw * 13:56 cmooney@dns3003: START - running authdns-update * 13:56 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'. * 13:56 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns5003.*,service=authdns-update * 13:55 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host chartmuseum2001.codfw.wmnet * 13:55 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus1007.eqiad.wmnet * 13:55 cmooney@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dns5003.wikimedia.org * 13:55 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus1005.eqiad.wmnet * 13:51 cgoubert@cumin2003: START - Cookbook sre.hosts.reboot-single for host chartmuseum2001.codfw.wmnet * 13:51 cgoubert@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=helm-charts.*,name=codfw * 13:51 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus2005.codfw.wmnet * 13:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host apifeatureusage1001.eqiad.wmnet with OS bookworm * 13:50 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus7002.magru.wmnet * 13:50 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts lvs1015.eqiad.wmnet * 13:50 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:50 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1015.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:49 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1015.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:49 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325480{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] (duration: 06m 39s) * 13:47 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1215.eqiad.wmnet with reason: host reimage * 13:46 cmooney@cumin1003: START - Cookbook sre.hosts.reboot-single for host dns5003.wikimedia.org * 13:46 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1003.eqiad.wmnet * 13:45 cmooney@cumin1003: conftool action : set/pooled=no; selector: name=dns5003.* * 13:45 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1040.eqiad.wmnet * 13:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1040.eqiad.wmnet * 13:45 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 13:45 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus1005.eqiad.wmnet * 13:45 stran@deploy1003: stran: Continuing with deployment * 13:44 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus7002.magru.wmnet * 13:44 stran@deploy1003: stran: Backport for [[gerrit:1325480{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:44 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus6002.drmrs.wmnet * 13:43 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260805000000" --end-timestamp="20260806000000" --sleep="5" --batch-size="10"` for [[phab:T434688|T434688]] * 13:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1149.eqiad.wmnet with reason: host reimage * 13:42 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1325480{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] * 13:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow1003.eqiad.wmnet * 13:40 sukhe@cumin1003: START - Cookbook sre.hosts.decommission for hosts lvs1015.eqiad.wmnet * 13:40 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts lvs1014.eqiad.wmnet * 13:40 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:40 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1014.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:40 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1040.eqiad.wmnet * 13:40 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1014.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:39 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1214.eqiad.wmnet with reason: host reimage * 13:38 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2004.codfw.wmnet * 13:38 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus6002.drmrs.wmnet * 13:38 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus5003.eqsin.wmnet * 13:35 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 13:35 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1149.eqiad.wmnet with reason: host reimage * 13:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1215.eqiad.wmnet with reason: host reimage * 13:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow2004.codfw.wmnet * 13:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1214.eqiad.wmnet with reason: host reimage * 13:33 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1040.eqiad.wmnet * 13:31 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus5003.eqsin.wmnet * 13:31 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: cloudelastic cluster restart - bking@cumin2003 * 13:31 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus4003.ulsfo.wmnet * 13:30 sukhe@cumin1003: START - Cookbook sre.hosts.decommission for hosts lvs1014.eqiad.wmnet * 13:30 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts lvs1013.eqiad.wmnet * 13:30 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:30 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1013.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:30 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1013.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:27 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1039.eqiad.wmnet * 13:27 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1039.eqiad.wmnet * 13:26 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1167.eqiad.wmnet onto db1281.eqiad.wmnet * 13:26 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1167: Pool db1167.eqiad.wmnet in after cloning * 13:25 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus4003.ulsfo.wmnet * 13:25 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260802000000" --end-timestamp="20260803000000" --sleep="5" --batch-size="10"` for [[phab:T434688|T434688]] * 13:24 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 13:24 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus3004.esams.wmnet * 13:24 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1274: New host * 13:24 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325476{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] (duration: 07m 19s) * 13:24 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260804000000" --end-timestamp="20260805000000" --sleep="5" --batch-size="10"` for [[phab:T434688|T434688]] * 13:24 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260803000000" --end-timestamp="20260804000000" --sleep="5" --batch-size="10"` for [[phab:T434688|T434688]] * 13:23 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2003.codfw.wmnet * 13:22 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1039.eqiad.wmnet * 13:20 sukhe@cumin1003: START - Cookbook sre.hosts.decommission for hosts lvs1013.eqiad.wmnet * 13:20 stran@deploy1003: stran: Continuing with deployment * 13:19 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow2003.codfw.wmnet * 13:19 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling reboot on A:wikidough * 13:19 stran@deploy1003: stran: Backport for [[gerrit:1325476{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:18 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus3004.esams.wmnet * 13:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1215.eqiad.wmnet with OS bookworm * 13:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1214.eqiad.wmnet with OS bookworm * 13:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1149.eqiad.wmnet with OS bookworm * 13:17 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1325476{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] * 13:11 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324731{{!}}prv: Enable parsoid rendering for 5 wikisource wikis]] (duration: 07m 26s) * 13:11 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1039.eqiad.wmnet * 13:11 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1036.eqiad.wmnet * 13:11 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1036.eqiad.wmnet * 13:07 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow3004.esams.wmnet * 13:07 jgiannelos@deploy1003: jgiannelos: Continuing with deployment * 13:06 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1233.eqiad.wmnet with OS bookworm * 13:06 jgiannelos@deploy1003: jgiannelos: Backport for [[gerrit:1324731{{!}}prv: Enable parsoid rendering for 5 wikisource wikis]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:04 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1324731{{!}}prv: Enable parsoid rendering for 5 wikisource wikis]] * 13:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow3004.esams.wmnet * 13:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1036.eqiad.wmnet * 13:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1210.eqiad.wmnet with OS bookworm * 13:01 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1036.eqiad.wmnet * 13:00 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1052.eqiad.wmnet * 13:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1052.eqiad.wmnet * 12:58 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow4003.ulsfo.wmnet * 12:55 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1211.eqiad.wmnet with OS bookworm * 12:55 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1052.eqiad.wmnet * 12:52 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow4003.ulsfo.wmnet * 12:51 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1052.eqiad.wmnet * 12:49 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1169: Depool db1169.eqiad.wmnet to then clone it to db1277.eqiad.wmnet - cwilliams@cumin1003 * 12:46 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1169: Depool db1169.eqiad.wmnet to then clone it to db1277.eqiad.wmnet - cwilliams@cumin1003 * 12:46 cwilliams@cumin1003: START - Cookbook sre.mysql.clone of db1169.eqiad.wmnet onto db1277.eqiad.wmnet * 12:45 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:45 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns record for deleted IP reservations eqsin lvs vlan ints - cmooney@cumin1003" * 12:44 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns record for deleted IP reservations eqsin lvs vlan ints - cmooney@cumin1003" * 12:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1051.eqiad.wmnet * 12:43 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1051.eqiad.wmnet * 12:42 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1210.eqiad.wmnet with reason: host reimage * 12:41 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1167: Pool db1167.eqiad.wmnet in after cloning * 12:40 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 12:39 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1274: New host * 12:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Added db1274', diff saved to https://phabricator.wikimedia.org/P96062 and previous config saved to /var/cache/conftool/dbconfig/20260813-123907-cwilliams.json * 12:38 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast7002.wikimedia.org * 12:38 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1233.eqiad.wmnet with reason: host reimage * 12:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1051.eqiad.wmnet * 12:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1211.eqiad.wmnet with reason: host reimage * 12:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1233.eqiad.wmnet with reason: host reimage * 12:32 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1051.eqiad.wmnet * 12:32 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast7002.wikimedia.org * 12:32 marostegui: Drop SecurePoll tables from closed wikis [[phab:T423128|T423128]] * 12:32 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow5003.eqsin.wmnet * 12:31 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1050.eqiad.wmnet * 12:31 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1050.eqiad.wmnet * 12:29 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1210.eqiad.wmnet with reason: host reimage * 12:29 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1211.eqiad.wmnet with reason: host reimage * 12:29 moritzm: installing Wireshark security updates * 12:26 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow5003.eqsin.wmnet * 12:26 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1050.eqiad.wmnet * 12:24 cmooney@dns3003: END - running authdns-update * 12:21 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow6001.drmrs.wmnet * 12:21 cmooney@dns3003: START - running authdns-update * 12:20 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1050.eqiad.wmnet * 12:20 cgoubert@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host rdb-lock2003.codfw.wmnet * 12:20 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:20 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update netbox dns entries for expanded public1-603-eqsin subnet - cmooney@cumin1003" * 12:20 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update netbox dns entries for expanded public1-603-eqsin subnet - cmooney@cumin1003" * 12:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 12:20 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 12:19 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1049.eqiad.wmnet * 12:19 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1049.eqiad.wmnet * 12:19 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 12:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host moss-be1003.eqiad.wmnet * 12:17 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow6001.drmrs.wmnet * 12:17 moritzm: installin curl security updates * 12:15 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 12:15 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1233.eqiad.wmnet with OS bookworm * 12:15 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1210.eqiad.wmnet with OS bookworm * 12:15 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1211.eqiad.wmnet with OS bookworm * 12:13 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1049.eqiad.wmnet * 12:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow7002.magru.wmnet * 12:12 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-worker1178.eqiad.wmnet with OS bookworm * 12:10 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host moss-be1003.eqiad.wmnet * 12:10 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 12:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be1006.eqiad.wmnet * 12:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 12:10 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 12:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 12:10 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 12:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow7002.magru.wmnet * 12:08 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1049.eqiad.wmnet * 12:05 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 12:05 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2003.codfw.wmnet * 12:04 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'. * 12:04 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:03 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be1006.eqiad.wmnet * 12:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be1005.eqiad.wmnet * 12:01 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1038.eqiad.wmnet * 12:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1038.eqiad.wmnet * 12:01 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1213.eqiad.wmnet with OS bookworm * 11:56 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be1005.eqiad.wmnet * 11:55 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be1004.eqiad.wmnet * 11:54 cgoubert@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host rdb-lock2003.codfw.wmnet * 11:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 11:53 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 11:53 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:53 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 11:53 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 11:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1038.eqiad.wmnet * 11:49 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be1004.eqiad.wmnet * 11:44 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:44 moritzm: installing Linux 5.10.262 on Bullseye hosts * 11:41 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1038.eqiad.wmnet * 11:40 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 11:40 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 11:40 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 11:40 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:40 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 11:40 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 11:38 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1213.eqiad.wmnet with reason: host reimage * 11:36 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1167: Depool db1167.eqiad.wmnet to then clone it to db1281.eqiad.wmnet - marostegui@cumin1003 * 11:35 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 11:35 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2003.codfw.wmnet * 11:35 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1167: Depool db1167.eqiad.wmnet to then clone it to db1281.eqiad.wmnet - marostegui@cumin1003 * 11:35 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1167.eqiad.wmnet onto db1281.eqiad.wmnet * 11:34 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 22 hosts with reason: Cloning * 11:34 moritzm: remove ganeti3005 from esams03 cluster, hardware issues [[phab:T434646|T434646]] * 11:32 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1213.eqiad.wmnet with reason: host reimage * 11:28 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1034.eqiad.wmnet * 11:28 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1034.eqiad.wmnet * 11:22 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1034.eqiad.wmnet * 11:19 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1034.eqiad.wmnet * 11:17 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1213.eqiad.wmnet with OS bookworm * 11:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1165.eqiad.wmnet onto db1279.eqiad.wmnet * 11:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1165: Pool db1165.eqiad.wmnet in after cloning * 11:07 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply * 10:57 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply * 10:54 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply * 10:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1274.eqiad.wmnet with reason: Enabling notifications and pooling * 10:45 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1033.eqiad.wmnet * 10:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1033.eqiad.wmnet * 10:44 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply * 10:43 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'. * 10:42 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'. * 10:42 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'. * 10:40 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db[1216,1225,1239-1240].eqiad.wmnet with reason: reboot * 10:39 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1033.eqiad.wmnet * 10:38 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1151 from dbctl [[phab:T434538|T434538]]', diff saved to https://phabricator.wikimedia.org/P96055 and previous config saved to /var/cache/conftool/dbconfig/20260813-103828-marostegui.json * 10:35 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1033.eqiad.wmnet * 10:27 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1165: Pool db1165.eqiad.wmnet in after cloning * 10:24 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1151.eqiad.wmnet with OS bookworm * 10:15 moritzm: installing bind9 security updates (client-side tools/libs only) * 10:07 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-debug: apply * 10:06 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-debug: apply * 10:02 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 7 hosts with reason: reboot * 10:01 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-debug: apply * 10:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.upgrade (exit_code=0) for 1 hosts * 10:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2214: Upgrade of db2214.codfw.wmnet completed * 10:01 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-debug: apply * 10:00 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-debug: apply * 10:00 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-debug: apply * 09:59 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.decommission (exit_code=99) * 09:59 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 09:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1151.eqiad.wmnet with reason: host reimage * 09:59 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 8 hosts * 09:59 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 8 hosts * 09:55 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1151.eqiad.wmnet with reason: host reimage * 09:46 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260801000000" --end-timestamp="20260802000000" --sleep="3" --batch-size="5"` for [[phab:T434688|T434688]] * 09:45 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 8 hosts with reason: reboot * 09:45 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 22 hosts with reason: Cloning * 09:43 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1165: Depool db1165.eqiad.wmnet to then clone it to db1279.eqiad.wmnet - marostegui@cumin1003 * 09:42 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1165: Depool db1165.eqiad.wmnet to then clone it to db1279.eqiad.wmnet - marostegui@cumin1003 * 09:42 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1165.eqiad.wmnet onto db1279.eqiad.wmnet * 09:41 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=testwiki --start-timestamp="20260311000000" --end-timestamp="20260805000000" --sleep="5" --batch-size="2"` for [[phab:T434688|T434688]] * 09:40 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for backupmon1001.eqiad.wmnet * 09:40 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for backupmon1001.eqiad.wmnet * 09:40 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1151 * 09:40 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1151 * 09:37 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325408{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]], [[gerrit:1325407{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]] (duration: 06m 57s) * 09:36 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1151 * 09:36 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1151.eqiad.wmnet 13.36.64.10.in-addr.arpa 3.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:36 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1151.eqiad.wmnet 13.36.64.10.in-addr.arpa 3.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:36 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:36 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1151 - btullis@cumin1003" * 09:36 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on backupmon1001.eqiad.wmnet with reason: reboot * 09:36 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1151 - btullis@cumin1003" * 09:35 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host moss-be2003.codfw.wmnet * 09:33 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 7 hosts * 09:33 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 7 hosts * 09:33 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 09:32 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1325408{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]], [[gerrit:1325407{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1161.eqiad.wmnet onto db1275.eqiad.wmnet * 09:31 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 09:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1161: Pool db1161.eqiad.wmnet in after cloning * 09:30 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1325408{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]], [[gerrit:1325407{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]] * 09:29 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 09:27 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 09:27 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host moss-be2003.codfw.wmnet * 09:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be2006.codfw.wmnet * 09:25 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2004.codfw.wmnet * 09:22 hashar@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 09:21 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be2006.codfw.wmnet * 09:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be2005.codfw.wmnet * 09:19 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2004.codfw.wmnet * 09:19 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2003.codfw.wmnet * 09:18 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 7 hosts with reason: reboot * 09:18 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 7 hosts * 09:18 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 7 hosts * 09:16 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2214: Upgrade of db2214.codfw.wmnet completed * 09:14 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be2005.codfw.wmnet * 09:13 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be2004.codfw.wmnet * 09:12 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2003.codfw.wmnet * 09:12 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2002.codfw.wmnet * 09:09 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2214: Upgrading db2214.codfw.wmnet * 09:09 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2214: Upgrading db2214.codfw.wmnet * 09:09 cwilliams@cumin1003: START - Cookbook sre.mysql.upgrade for 1 hosts * 09:07 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be2004.codfw.wmnet * 09:06 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2002.codfw.wmnet * 09:05 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1004.eqiad.wmnet * 09:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-cluster (exit_code=0) * 09:03 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 7 hosts with reason: reboot * 09:01 btullis@cumin1003: START - Cookbook sre.dns.netbox * 09:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2214 [[phab:T434754|T434754]]', diff saved to https://phabricator.wikimedia.org/P96045 and previous config saved to /var/cache/conftool/dbconfig/20260813-090001-cwilliams.json * 08:59 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1004.eqiad.wmnet * 08:59 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1003.eqiad.wmnet * 08:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2229 to s6 primary [[phab:T434754|T434754]]', diff saved to https://phabricator.wikimedia.org/P96044 and previous config saved to /var/cache/conftool/dbconfig/20260813-085752-cwilliams.json * 08:57 cezmunsta: Starting s6 codfw failover from db2214 to db2229 - [[phab:T434754|T434754]] * 08:54 hashar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325416{{!}}Revert "REST: Enable `GET /lexemes/<nowiki>{</nowiki>lexeme_id<nowiki>}</nowiki>` by default" (T434712)]] (duration: 07m 22s) * 08:53 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1003.eqiad.wmnet * 08:53 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1002.eqiad.wmnet * 08:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2229 with weight 0 [[phab:T434754|T434754]]', diff saved to https://phabricator.wikimedia.org/P96043 and previous config saved to /var/cache/conftool/dbconfig/20260813-085151-cwilliams.json * 08:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 23 hosts with reason: Primary switchover s6 [[phab:T434754|T434754]] * 08:51 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts2002.codfw.wmnet * 08:50 hashar@deploy1003: hashar: Continuing with deployment * 08:49 hashar@deploy1003: hashar: Backport for [[gerrit:1325416{{!}}Revert "REST: Enable `GET /lexemes/<nowiki>{</nowiki>lexeme_id<nowiki>}</nowiki>` by default" (T434712)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:47 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet * 08:47 hashar@deploy1003: Started scap sync-world: Backport for [[gerrit:1325416{{!}}Revert "REST: Enable `GET /lexemes/<nowiki>{</nowiki>lexeme_id<nowiki>}</nowiki>` by default" (T434712)]] * 08:47 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1002.eqiad.wmnet * 08:46 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1161: Pool db1161.eqiad.wmnet in after cloning * 08:46 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host stewards2001.codfw.wmnet * 08:45 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit1003.wikimedia.org * 08:45 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host stewards1001.eqiad.wmnet * 08:45 Emperor: roll-restart apus frontends in codfw * 08:45 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-cluster * 08:44 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts2002.codfw.wmnet * 08:44 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host doc2003.codfw.wmnet * 08:43 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet * 08:43 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host phab2003.codfw.wmnet * 08:42 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host stewards2001.codfw.wmnet * 08:42 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host etherpad2002.codfw.wmnet * 08:41 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host stewards1001.eqiad.wmnet * 08:41 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1273: Pool in s7 * 08:41 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1003.wikimedia.org * 08:40 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host doc2003.codfw.wmnet * 08:40 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit1003.wikimedia.org * 08:39 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host etherpad1004.eqiad.wmnet * 08:39 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host doc1004.eqiad.wmnet * 08:39 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit2002.wikimedia.org * 08:38 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host etherpad2002.codfw.wmnet * 08:37 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host phab2003.codfw.wmnet * 08:36 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lists1004.wikimedia.org * 08:35 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host etherpad1004.eqiad.wmnet * 08:35 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host doc1004.eqiad.wmnet * 08:34 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-cluster (exit_code=0) * 08:34 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1003.wikimedia.org * 08:34 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host planet2003.codfw.wmnet * 08:34 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2003.wikimedia.org * 08:33 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host planet1003.eqiad.wmnet * 08:33 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit2002.wikimedia.org * 08:32 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 22 hosts with reason: Cloning * 08:32 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2002.wikimedia.org * 08:31 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aphlict1002.eqiad.wmnet * 08:30 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host planet2003.codfw.wmnet * 08:29 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host planet1003.eqiad.wmnet * 08:29 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host lists1004.wikimedia.org * 08:28 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lists2001.wikimedia.org * 08:28 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2003.wikimedia.org * 08:27 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host aphlict1002.eqiad.wmnet * 08:27 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aphlict2001.codfw.wmnet * 08:26 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2002.wikimedia.org * 08:23 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host aphlict2001.codfw.wmnet * 08:23 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1151 * 08:22 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1151.eqiad.wmnet with OS bookworm * 08:22 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host lists2001.wikimedia.org * 08:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1161: Depool db1161.eqiad.wmnet to then clone it to db1275.eqiad.wmnet - marostegui@cumin1003 * 08:20 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1161: Depool db1161.eqiad.wmnet to then clone it to db1275.eqiad.wmnet - marostegui@cumin1003 * 08:20 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1161.eqiad.wmnet onto db1275.eqiad.wmnet * 08:15 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-cluster * 08:15 Emperor: roll-restart apus frontends in eqiad * 07:56 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1273: Pool in s7 * 07:56 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1273 to dbctl [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P96035 and previous config saved to /var/cache/conftool/dbconfig/20260813-075611-marostegui.json * 07:38 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db1273.eqiad.wmnet with reason: Reboot * 07:31 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.sanitize-wiki (exit_code=97) Managing sanitization for wikis testwiki in section s3 * 07:24 marostegui@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis testwiki in section s3 * 07:19 jayme@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on kubestagemaster2005.codfw.wmnet with reason: downtime because of hardware failure and no DRBD * 05:42 arnaudb@dns1006: END - running authdns-update * 05:40 arnaudb@dns1006: START - running authdns-update * 05:27 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 04:06 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 02:29 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1151 * 02:29 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1151 * 02:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1219.eqiad.wmnet with OS bookworm * 02:18 ryankemper: [[phab:T434494|T434494]] `ryankemper@deploy1003:~$ echo 'https://stats.wikimedia.org/' {{!}} mwscript-k8s --attach -- purgeList.php` (default page got cached during yesterday's `an-web1001` reimage) * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 46s) * 02:03 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1219.eqiad.wmnet with reason: host reimage * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 02:00 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1219.eqiad.wmnet with reason: host reimage * 01:46 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1219.eqiad.wmnet with OS bookworm * 01:03 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1142.eqiad.wmnet with OS bookworm * 00:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1142.eqiad.wmnet with reason: host reimage * 00:34 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1142.eqiad.wmnet with reason: host reimage * 00:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1142.eqiad.wmnet with OS bookworm * 00:16 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1178 * 00:16 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1178 == 2026-08-12 == * 23:18 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324409{{!}}Remove $wmg = $wg hacks in Collection (T119117)]] (duration: 06m 43s) * 23:14 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 23:13 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1324409{{!}}Remove $wmg = $wg hacks in Collection (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:11 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1324409{{!}}Remove $wmg = $wg hacks in Collection (T119117)]] * 22:59 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 22:49 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260801000000" --end-timestamp="20260802000000" --sleep=2 --batch-size=10` * 22:45 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=testwiki --start-timestamp="20200801010101" --end-timestamp="20260816010101" --sleep=15 --batch-size=5` * 22:40 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=testwiki --start-timestamp="20200101010101" --end-timestamp="20260816010101" --sleep=60` * 22:35 Dreamy_Jazz: Running `mwscript WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260101000000" --end-timestamp="20260102000000" --sleep=10` * 22:21 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324817{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]], [[gerrit:1324818{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]] (duration: 45m 29s) * 22:17 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 21:59 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 21:40 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1324817{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]], [[gerrit:1324818{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:39 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 21:36 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1324817{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]], [[gerrit:1324818{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]] * 21:32 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:31 vriley@cumin1003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:30 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:30 vriley@cumin1003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:17 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:14 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:14 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:11 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:10 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-codfw: Set storage compatability to NONE — [[phab:T433028|T433028]] - eevans@cumin1003 * 21:10 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1006 * 21:09 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1006 * 21:05 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324277{{!}}Improve Math preference labels for SVG/MathJax/MathML (T433891)]] (duration: 31m 42s) * 20:58 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1178.eqiad.wmnet with OS bookworm * 20:54 krinkle@deploy1003: krinkle: Continuing with deployment * 20:51 krinkle@deploy1003: krinkle: Backport for [[gerrit:1324277{{!}}Improve Math preference labels for SVG/MathJax/MathML (T433891)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:39 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 20:38 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:38 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp5022.eqsin.wmnet with OS trixie * 20:38 cdobbins@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - cdobbins@cumin1003" * 20:37 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:36 cdobbins@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - cdobbins@cumin1003" * 20:34 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1178.eqiad.wmnet with reason: host reimage * 20:34 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1324277{{!}}Improve Math preference labels for SVG/MathJax/MathML (T433891)]] * 20:33 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 20:28 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1178.eqiad.wmnet with reason: host reimage * 20:13 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1178.eqiad.wmnet with OS bookworm * 20:11 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-worker1178.eqiad.wmnet with OS bookworm * 20:11 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1178.eqiad.wmnet with OS bookworm * 20:09 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-codfw: Set storage compatability to NONE — [[phab:T433028|T433028]] - eevans@cumin1003 * 20:09 Dreamy_Jazz: Evening UTC backport window done * 20:08 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324794{{!}}WikimediaAntiAbuse: Enable logging channel (T431292)]] (duration: 06m 48s) * 20:08 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage * 20:05 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage * 20:04 dreamyjazz@deploy1003: kharlan, dreamyjazz: Continuing with deployment * 20:04 dreamyjazz@deploy1003: kharlan, dreamyjazz: Backport for [[gerrit:1324794{{!}}WikimediaAntiAbuse: Enable logging channel (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:01 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1324794{{!}}WikimediaAntiAbuse: Enable logging channel (T431292)]] * 19:55 brennen@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 19:47 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 19:35 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 19:35 cdobbins@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cp5022.eqsin.wmnet with OS trixie * 19:32 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:30 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:29 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:26 vriley@cumin1003: START - Cookbook sre.dns.netbox * 19:23 brennen@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 19:19 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-eqiad: Set storage compatability to NONE — [[phab:T433028|T433028]] - eevans@cumin1003 * 19:10 brennen@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 19:09 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 19:09 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 19:00 Amir1: data migrated on wikishared ([[phab:T426102|T426102]]) * 18:57 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321224{{!}}Rename ce_worklist_articles table to ce_invitation_list_articles (T426102)]] (duration: 06m 50s) * 18:53 ladsgroup@deploy1003: ladsgroup, daimona: Continuing with deployment * 18:53 ladsgroup@deploy1003: ladsgroup, daimona: Backport for [[gerrit:1321224{{!}}Rename ce_worklist_articles table to ce_invitation_list_articles (T426102)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:51 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1321224{{!}}Rename ce_worklist_articles table to ce_invitation_list_articles (T426102)]] * 18:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2192: Security update * 18:31 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 18:25 Amir1: ce_invitation_list_articles created as empty on wikishared ([[phab:T426102|T426102]]) * 18:21 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 18:21 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 18:18 Amir1: migrated testwiki entries from ce_worklist_articles to ce_invitation_list_articles ([[phab:T426102|T426102]]) * 18:18 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-eqiad: Set storage compatability to NONE — [[phab:T433028|T433028]] - eevans@cumin1003 * 18:11 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 18:08 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling reboot on A:durum-eqsin and A:durum * 18:07 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Set storage compatability to UPGRADING — [[phab:T433028|T433028]] - eevans@cumin1003 * 18:05 jhancock@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022'] * 17:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2192: Security update * 17:55 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:55 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum-eqsin and A:durum * 17:53 jhancock@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['cp5022'] * 17:47 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:47 jhancock@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['cp5022'] * 17:42 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:41 jhancock@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['cp5022'] * 17:36 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:35 jhancock@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022'] * 17:31 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-magru and not (P<nowiki>{</nowiki>cp7001*<nowiki>}</nowiki> or P<nowiki>{</nowiki>cp7009*<nowiki>}</nowiki>) and A:cp - 9.2.15 upgrade ([[phab:T434620|T434620]]) * 17:28 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:28 jhancock@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['cp5022'] * 17:22 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2192.codfw.wmnet with reason: Maintenance * 17:11 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host apifeatureusage2001.codfw.wmnet with OS bookworm * 17:04 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: apply * 17:03 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-main: apply * 17:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2192 [[phab:T434635|T434635]]', diff saved to https://phabricator.wikimedia.org/P96030 and previous config saved to /var/cache/conftool/dbconfig/20260812-170338-cwilliams.json * 17:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2213 to s5 primary [[phab:T434635|T434635]]', diff saved to https://phabricator.wikimedia.org/P96029 and previous config saved to /var/cache/conftool/dbconfig/20260812-170152-cwilliams.json * 17:01 cezmunsta: Starting s5 codfw failover from db2192 to db2213 - [[phab:T434635|T434635]] * 16:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2213 with weight 0 [[phab:T434635|T434635]]', diff saved to https://phabricator.wikimedia.org/P96028 and previous config saved to /var/cache/conftool/dbconfig/20260812-165544-cwilliams.json * 16:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 27 hosts with reason: Primary switchover s5 [[phab:T434635|T434635]] * 16:53 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: apply * 16:52 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-main: apply * 16:44 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-main: apply * 16:44 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-main: apply * 16:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1272: New host * 16:40 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324765{{!}}WikimediaAntiAbuse: Enable PersonalInfoFlagNotifications (T431292)]] (duration: 07m 02s) * 16:40 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-logging-external: apply * 16:39 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-logging-external: apply * 16:38 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-logging-external: apply * 16:37 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-logging-external: apply * 16:36 kharlan@deploy1003: kharlan: Continuing with deployment * 16:35 kharlan@deploy1003: kharlan: Backport for [[gerrit:1324765{{!}}WikimediaAntiAbuse: Enable PersonalInfoFlagNotifications (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:33 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1324765{{!}}WikimediaAntiAbuse: Enable PersonalInfoFlagNotifications (T431292)]] * 16:27 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-logging-external: apply * 16:27 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-logging-external: apply * 16:18 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 16:18 jhancock@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022'] * 16:17 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 16:16 jhancock@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['cp5022'] * 16:15 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 16:12 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Set storage compatability to UPGRADING — [[phab:T433028|T433028]] - eevans@cumin1003 * 16:10 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: apply * 16:10 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: apply * 16:08 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: apply * 16:08 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: apply * 16:08 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics: apply * 16:07 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics: apply * 16:02 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-magru and not (P<nowiki>{</nowiki>cp7001*<nowiki>}</nowiki> or P<nowiki>{</nowiki>cp7009*<nowiki>}</nowiki>) and A:cp - 9.2.15 upgrade ([[phab:T434620|T434620]]) * 15:57 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1272: New host * 15:52 jmm@cumin2003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti2046.codfw.wmnet * 15:52 jmm@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host ganeti2046.codfw.wmnet * 15:42 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:37 mforns@deploy1003: Finished deploy [analytics/refinery@49c336c] (thin): Regular analytics weekly train THIN [analytics/refinery@49c336cd] (duration: 01m 59s) * 15:37 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324733{{!}}WikimediaAntiAbuse: Enable personal info tag display on enwiki (T431292)]] (duration: 08m 12s) * 15:35 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:35 mforns@deploy1003: Started deploy [analytics/refinery@49c336c] (thin): Regular analytics weekly train THIN [analytics/refinery@49c336cd] * 15:34 mforns@deploy1003: Finished deploy [analytics/refinery@49c336c]: Regular analytics weekly train [analytics/refinery@49c336cd] (duration: 04m 20s) * 15:33 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[2024,1031]*.wmnet: Set storage compatability to UPGRADING — [[phab:T433028|T433028]] - eevans@cumin1003 * 15:33 dreamyjazz@deploy1003: kharlan, dreamyjazz: Continuing with deployment * 15:31 dreamyjazz@deploy1003: kharlan, dreamyjazz: Backport for [[gerrit:1324733{{!}}WikimediaAntiAbuse: Enable personal info tag display on enwiki (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:30 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitize-wiki (exit_code=99) Checking sanitization for wikis testwiki in section s3 * 15:30 mforns@deploy1003: Started deploy [analytics/refinery@49c336c]: Regular analytics weekly train [analytics/refinery@49c336cd] * 15:30 mforns@deploy1003: Finished deploy [analytics/refinery@49c336c] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@49c336cd] (duration: 00m 32s) * 15:29 mforns@deploy1003: Started deploy [analytics/refinery@49c336c] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@49c336cd] * 15:29 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1324733{{!}}WikimediaAntiAbuse: Enable personal info tag display on enwiki (T431292)]] * 15:27 brennen@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324700{{!}}EventDetailsParticipantsModule: populate cache with non-local users (T434597)]] (duration: 06m 38s) * 15:23 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[2024,1031]*.wmnet: Set storage compatability to UPGRADING — [[phab:T433028|T433028]] - eevans@cumin1003 * 15:23 brennen@deploy1003: brennen, daimona: Continuing with deployment * 15:23 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:22 brennen@deploy1003: brennen, daimona: Backport for [[gerrit:1324700{{!}}EventDetailsParticipantsModule: populate cache with non-local users (T434597)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:22 cgoubert@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host rdb-lock2003.codfw.wmnet * 15:21 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 15:21 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 15:21 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:21 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:21 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:20 brennen@deploy1003: Started scap sync-world: Backport for [[gerrit:1324700{{!}}EventDetailsParticipantsModule: populate cache with non-local users (T434597)]] * 15:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1178.eqiad.wmnet with OS bookworm * 15:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock1003.eqiad.wmnet * 15:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock1003.eqiad.wmnet with OS trixie * 15:16 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 15:16 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 15:16 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 15:16 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:16 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:16 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:12 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 15:12 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2003.codfw.wmnet * 15:11 cgoubert@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host rdb-lock2003.codfw.wmnet * 15:11 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 15:11 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 15:11 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:11 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:11 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:07 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324338{{!}}InitialiseSettings: Enable 2FA warnings on more private wikis (T428103)]] (duration: 07m 02s) * 15:04 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 15:04 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock1003.eqiad.wmnet with reason: host reimage * 15:03 reedy@deploy1003: reedy: Continuing with deployment * 15:02 reedy@deploy1003: reedy: Backport for [[gerrit:1324338{{!}}InitialiseSettings: Enable 2FA warnings on more private wikis (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:02 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:00 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324338{{!}}InitialiseSettings: Enable 2FA warnings on more private wikis (T428103)]] * 14:57 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 14:57 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2003.codfw.wmnet * 14:57 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock1003.eqiad.wmnet with reason: host reimage * 14:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock2002.codfw.wmnet * 14:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock2002.codfw.wmnet with OS trixie * 14:56 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 14:56 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 14:56 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 14:55 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 14:55 moritzm: powercycle ganeti2046 * 14:47 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock1003.eqiad.wmnet with OS trixie * 14:46 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1003.eqiad.wmnet - cgoubert@cumin2003" * 14:46 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1003.eqiad.wmnet - cgoubert@cumin2003" * 14:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock1003.eqiad.wmnet on all recursors * 14:45 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock1003.eqiad.wmnet on all recursors * 14:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1003.eqiad.wmnet - cgoubert@cumin2003" * 14:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1272.eqiad.wmnet with reason: Enabling notifications * 14:44 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1003.eqiad.wmnet - cgoubert@cumin2003" * 14:44 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324719{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324720{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324722{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0 (T434187)]] (duration: 11m 02s) * 14:39 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 14:39 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock1003.eqiad.wmnet * 14:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock2002.codfw.wmnet with reason: host reimage * 14:37 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 14:37 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock1002.eqiad.wmnet * 14:37 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock1002.eqiad.wmnet with OS trixie * 14:37 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1324719{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324720{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324722{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0 (T434187)]] synced to the testservers (see https://wikitech. * 14:33 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock2002.codfw.wmnet with reason: host reimage * 14:33 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1324719{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324720{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324722{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0 (T434187)]] * 14:32 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1018.eqiad.wmnet with OS bookworm * 14:32 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2046.codfw.wmnet * 14:31 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1020.eqiad.wmnet with OS bookworm * 14:27 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2046.codfw.wmnet * 14:25 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2045.codfw.wmnet * 14:25 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2045.codfw.wmnet * 14:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock1002.eqiad.wmnet with reason: host reimage * 14:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1019.eqiad.wmnet with OS bookworm * 14:22 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Checking sanitization for wikis testwiki in section s3 * 14:20 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2045.codfw.wmnet * 14:18 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock1002.eqiad.wmnet with reason: host reimage * 14:17 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:16 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324707{{!}}Backport all changes from wmf/1.47.0-wmf.15]] (duration: 40m 51s) * 14:16 moritzm: installing Linux 6.1.180 on Bookworm hosts * 14:15 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock2002.codfw.wmnet with OS trixie * 14:14 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2002.codfw.wmnet - cgoubert@cumin2003" * 14:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2002.codfw.wmnet - cgoubert@cumin2003" * 14:14 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:14 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2045.codfw.wmnet * 14:14 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2002.codfw.wmnet on all recursors * 14:14 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2002.codfw.wmnet on all recursors * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2002.codfw.wmnet - cgoubert@cumin2003" * 14:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2002.codfw.wmnet - cgoubert@cumin2003" * 14:12 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2030.codfw.wmnet * 14:12 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2030.codfw.wmnet * 14:11 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on apifeatureusage2001.codfw.wmnet with reason: host reimage * 14:09 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:08 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:07 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 14:06 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2030.codfw.wmnet * 14:06 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:06 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock1002.eqiad.wmnet with OS trixie * 14:05 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 14:05 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2002.codfw.wmnet * 14:05 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1002.eqiad.wmnet - cgoubert@cumin2003" * 14:05 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1002.eqiad.wmnet - cgoubert@cumin2003" * 14:05 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock1002.eqiad.wmnet on all recursors * 14:05 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock1002.eqiad.wmnet on all recursors * 14:05 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:05 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1002.eqiad.wmnet - cgoubert@cumin2003" * 14:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:04 kharlan@deploy1003: kharlan: Continuing with deployment * 14:04 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock2001.codfw.wmnet * 14:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock2001.codfw.wmnet with OS trixie * 14:02 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1002.eqiad.wmnet - cgoubert@cumin2003" * 14:02 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on apifeatureusage2001.codfw.wmnet with reason: host reimage * 14:01 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2030.codfw.wmnet * 13:59 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2029.codfw.wmnet * 13:58 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2029.codfw.wmnet * 13:58 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 13:58 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock1002.eqiad.wmnet * 13:56 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1019.eqiad.wmnet with reason: host reimage * 13:54 btullis@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on archiva1002.wikimedia.org with reason: Upgrading in-place * 13:53 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock1001.eqiad.wmnet * 13:53 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock1001.eqiad.wmnet with OS trixie * 13:53 kharlan@deploy1003: kharlan: Backport for [[gerrit:1324707{{!}}Backport all changes from wmf/1.47.0-wmf.15]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:52 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2029.codfw.wmnet * 13:52 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1020.eqiad.wmnet with reason: host reimage * 13:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1019.eqiad.wmnet with reason: host reimage * 13:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1020.eqiad.wmnet with reason: host reimage * 13:49 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock2001.codfw.wmnet with reason: host reimage * 13:48 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2029.codfw.wmnet * 13:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Configuring db1272 for s3 pooling', diff saved to https://phabricator.wikimedia.org/P96021 and previous config saved to /var/cache/conftool/dbconfig/20260812-134732-cwilliams.json * 13:44 bking@cumin2003: START - Cookbook sre.hosts.reimage for host apifeatureusage2001.codfw.wmnet with OS bookworm * 13:43 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock2001.codfw.wmnet with reason: host reimage * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2028.codfw.wmnet * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2028.codfw.wmnet * 13:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1018.eqiad.wmnet with reason: host reimage * 13:38 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock1001.eqiad.wmnet with reason: host reimage * 13:36 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1018.eqiad.wmnet with reason: host reimage * 13:35 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1324707{{!}}Backport all changes from wmf/1.47.0-wmf.15]] * 13:35 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2028.codfw.wmnet * 13:32 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1275938{{!}}Enable campaignEvents on bdwikimedia (T424016)]] (duration: 07m 35s) * 13:32 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1020.eqiad.wmnet with OS bookworm * 13:32 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1019.eqiad.wmnet with OS bookworm * 13:32 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock1001.eqiad.wmnet with reason: host reimage * 13:28 kharlan@deploy1003: kharlan, yahya: Continuing with deployment * 13:27 kharlan@deploy1003: kharlan, yahya: Backport for [[gerrit:1275938{{!}}Enable campaignEvents on bdwikimedia (T424016)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:27 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2028.codfw.wmnet * 13:26 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 13:26 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 13:25 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1275938{{!}}Enable campaignEvents on bdwikimedia (T424016)]] * 13:25 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitize-wiki (exit_code=99) Managing sanitization for wikis testwiki in section s3 * 13:24 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock2001.codfw.wmnet with OS trixie * 13:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2001.codfw.wmnet - cgoubert@cumin2003" * 13:24 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2001.codfw.wmnet - cgoubert@cumin2003" * 13:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2001.codfw.wmnet on all recursors * 13:23 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2001.codfw.wmnet on all recursors * 13:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2001.codfw.wmnet - cgoubert@cumin2003" * 13:23 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2001.codfw.wmnet - cgoubert@cumin2003" * 13:23 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324713{{!}}thwiki: reinstate temporary wiki25 logos (T431094)]] (duration: 07m 13s) * 13:20 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:20 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1018.eqiad.wmnet with OS bookworm * 13:20 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2027.codfw.wmnet * 13:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2027.codfw.wmnet * 13:19 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics-external: apply * 13:18 kharlan@deploy1003: anzx, kharlan: Continuing with deployment * 13:18 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:18 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock1001.eqiad.wmnet with OS trixie * 13:17 kharlan@deploy1003: anzx, kharlan: Backport for [[gerrit:1324713{{!}}thwiki: reinstate temporary wiki25 logos (T431094)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1001.eqiad.wmnet - cgoubert@cumin2003" * 13:17 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 13:17 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1001.eqiad.wmnet - cgoubert@cumin2003" * 13:17 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2001.codfw.wmnet * 13:17 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics-external: apply * 13:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock1001.eqiad.wmnet on all recursors * 13:17 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock1001.eqiad.wmnet on all recursors * 13:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1001.eqiad.wmnet - cgoubert@cumin2003" * 13:17 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1001.eqiad.wmnet - cgoubert@cumin2003" * 13:15 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1324713{{!}}thwiki: reinstate temporary wiki25 logos (T431094)]] * 13:15 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:15 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics-external: apply * 13:15 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:14 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics-external: apply * 13:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2027.codfw.wmnet * 13:12 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 13:12 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock1001.eqiad.wmnet * 13:12 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2027.codfw.wmnet * 13:08 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1158.eqiad.wmnet onto db1273.eqiad.wmnet * 13:07 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1158: Pool db1158.eqiad.wmnet in after cloning * 13:02 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti6002.drmrs.wmnet * 13:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti6002.drmrs.wmnet * 12:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1015.eqiad.wmnet with OS bookworm * 12:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti6002.drmrs.wmnet * 12:44 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1017.eqiad.wmnet with OS bookworm * 12:36 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 12:35 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 12:34 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 12:33 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 12:31 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti6002.drmrs.wmnet * 12:24 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1017.eqiad.wmnet with reason: host reimage * 12:22 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1158: Pool db1158.eqiad.wmnet in after cloning * 12:18 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1017.eqiad.wmnet with reason: host reimage * 12:11 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1015.eqiad.wmnet with reason: host reimage * 12:07 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1015.eqiad.wmnet with reason: host reimage * 12:04 moritzm: failover ganeti master in drmrs02 to ganeti6004 * 12:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1009.eqiad.wmnet with OS bookworm * 12:01 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1017.eqiad.wmnet with OS bookworm * 12:00 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti6004.drmrs.wmnet * 12:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti6004.drmrs.wmnet * 11:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti6004.drmrs.wmnet * 11:53 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1015.eqiad.wmnet with OS bookworm * 11:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1159.eqiad.wmnet onto db1274.eqiad.wmnet * 11:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Pool db1159.eqiad.wmnet in after cloning * 11:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1016.eqiad.wmnet with OS bookworm * 11:46 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti6004.drmrs.wmnet * 11:45 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti6001.drmrs.wmnet * 11:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti6001.drmrs.wmnet * 11:43 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324306{{!}}WikimediaAntiAbuse: Enable personal info for enwiki with no display (T431292)]] (duration: 10m 26s) * 11:42 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1015.eqiad.wmnet with OS bookworm * 11:39 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 11:38 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti6001.drmrs.wmnet * 11:34 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1324306{{!}}WikimediaAntiAbuse: Enable personal info for enwiki with no display (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:33 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2212: Security update * 11:33 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti6001.drmrs.wmnet * 11:32 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1324306{{!}}WikimediaAntiAbuse: Enable personal info for enwiki with no display (T431292)]] * 11:22 moritzm: failover ganeti master in drmrs01 to ganeti6003 * 11:20 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:20 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:18 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 22 hosts with reason: Cloning * 11:17 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti6003.drmrs.wmnet * 11:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti6003.drmrs.wmnet * 11:17 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1009.eqiad.wmnet with reason: host reimage * 11:17 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:16 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:14 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1016.eqiad.wmnet with reason: host reimage * 11:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti6003.drmrs.wmnet * 11:11 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1009.eqiad.wmnet with reason: host reimage * 11:10 jmm@cumin2003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti3005.esams.wmnet * 11:10 jmm@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host ganeti3005.esams.wmnet * 11:09 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 11:08 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1158: Depool db1158.eqiad.wmnet to then clone it to db1273.eqiad.wmnet - marostegui@cumin1003 * 11:07 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1016.eqiad.wmnet with reason: host reimage * 11:07 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1158: Depool db1158.eqiad.wmnet to then clone it to db1273.eqiad.wmnet - marostegui@cumin1003 * 11:07 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1158.eqiad.wmnet onto db1273.eqiad.wmnet * 11:06 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:05 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:05 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1159: Pool db1159.eqiad.wmnet in after cloning * 11:04 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 20 hosts with reason: Cloning * 11:02 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti6003.drmrs.wmnet * 11:00 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis testwiki in section s3 * 10:54 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1009.eqiad.wmnet with OS bookworm * 10:51 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1015.eqiad.wmnet with OS bookworm * 10:50 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1016.eqiad.wmnet with OS bookworm * 10:48 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2212: Security update * 10:45 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitize-wiki (exit_code=99) Managing sanitization for wikis testwiki in section s3 * 10:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1013.eqiad.wmnet with OS bookworm * 10:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1014.eqiad.wmnet with OS bookworm * 10:38 cwilliams@cumin1003: START - Cookbook sre.mysql.clone of db1159.eqiad.wmnet onto db1274.eqiad.wmnet * 10:33 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db1274.eqiad.wmnet * 10:33 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db1274.eqiad.wmnet * 10:31 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 396993 * 10:29 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 396993 * 10:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1159: Clone source for db1274 * 10:24 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1159: Clone source for db1274 * 10:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1013.eqiad.wmnet with reason: host reimage * 10:18 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1013.eqiad.wmnet with reason: host reimage * 10:13 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2212.codfw.wmnet with reason: Maintenance * 10:13 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 10:12 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 10:12 blake@deploy1003: Stopping before sync operations * 10:11 blake@deploy1003: Started scap sync-world: Non-deployment scap run to populate new release values for [[phab:T427668|T427668]] * 10:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2212 [[phab:T434644|T434644]]', diff saved to https://phabricator.wikimedia.org/P96003 and previous config saved to /var/cache/conftool/dbconfig/20260812-101053-cwilliams.json * 10:09 moritzm: powercycle ganeti3005 * 10:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2203 to s1 primary [[phab:T434644|T434644]]', diff saved to https://phabricator.wikimedia.org/P96002 and previous config saved to /var/cache/conftool/dbconfig/20260812-100849-cwilliams.json * 10:08 cezmunsta: Starting s1 codfw failover from db2212 to db2203 - [[phab:T434644|T434644]] * 10:03 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1013.eqiad.wmnet with OS bookworm * 10:02 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1013.eqiad.wmnet with OS bookworm * 10:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2203 with weight 0 [[phab:T434644|T434644]]', diff saved to https://phabricator.wikimedia.org/P96001 and previous config saved to /var/cache/conftool/dbconfig/20260812-100134-cwilliams.json * 10:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 32 hosts with reason: Primary switchover s1 [[phab:T434644|T434644]] * 09:53 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1014.eqiad.wmnet with reason: host reimage * 09:50 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 09:50 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti3005.esams.wmnet * 09:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1014.eqiad.wmnet with reason: host reimage * 09:41 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1278: Pool in x1 * 09:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-launcher1003.eqiad.wmnet with OS bookworm * 09:37 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti3005.esams.wmnet * 09:37 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1013.eqiad.wmnet with OS bookworm * 09:34 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-presto1013.eqiad.wmnet with OS bookworm * 09:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1014.eqiad.wmnet with OS bookworm * 09:29 moritzm: failover ganeti master in esams to ganeti3008 * 09:26 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti3006.esams.wmnet * 09:26 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti3006.esams.wmnet * 09:24 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1012.eqiad.wmnet with OS bookworm * 09:23 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1009.eqiad.wmnet with OS bookworm * 09:23 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 09:18 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti3006.esams.wmnet * 09:16 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti3006.esams.wmnet * 09:14 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: fix regexp escaping bug - oblivian@cumin1003" * 09:14 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: fix regexp escaping bug - oblivian@cumin1003 * 09:13 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: fix regexp escaping bug - oblivian@cumin1003 * 09:13 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: fix regexp escaping bug - oblivian@cumin1003" * 09:03 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-launcher1003.eqiad.wmnet with reason: host reimage * 08:58 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-launcher1003.eqiad.wmnet with reason: host reimage * 08:55 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1278: Pool in x1 * 08:55 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1278 to dbctl [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95996 and previous config saved to /var/cache/conftool/dbconfig/20260812-085521-marostegui.json * 08:51 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1012.eqiad.wmnet with reason: host reimage * 08:45 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis testwiki in section s3 * 08:43 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1009.eqiad.wmnet with OS bookworm * 08:42 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1012.eqiad.wmnet with reason: host reimage * 08:41 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-launcher1003.eqiad.wmnet with OS bookworm * 08:40 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1013.eqiad.wmnet with OS bookworm * 08:38 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms3', diff saved to https://phabricator.wikimedia.org/P95995 and previous config saved to /var/cache/conftool/dbconfig/20260812-083816-marostegui.json * 08:38 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-master1003.eqiad.wmnet with OS bookworm * 08:37 marostegui: Failover ms3 [[phab:T434288|T434288]] * 08:37 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1268 to dbctl [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95994 and previous config saved to /var/cache/conftool/dbconfig/20260812-083722-marostegui.json * 08:35 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-web1001.eqiad.wmnet with OS bookworm * 08:32 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db2252.codfw.wmnet,db[1153,1268].eqiad.wmnet with reason: Switching over ms3 * 08:28 marostegui@cumin1003: dbctl commit (dc=all): 'Depool ms3 [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95993 and previous config saved to /var/cache/conftool/dbconfig/20260812-082852-marostegui.json * 08:25 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1012.eqiad.wmnet with OS bookworm * 08:25 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1011.eqiad.wmnet with OS bookworm * 08:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-master1003.eqiad.wmnet with reason: host reimage * 08:07 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-master1003.eqiad.wmnet with reason: host reimage * 08:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-web1001.eqiad.wmnet with reason: host reimage * 07:58 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-web1001.eqiad.wmnet with reason: host reimage * 07:50 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1003.eqiad.wmnet with OS bookworm * 07:38 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1011.eqiad.wmnet with reason: host reimage * 07:38 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-master1003.eqiad.wmnet with OS bookworm * 07:35 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti3007.esams.wmnet * 07:35 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti3007.esams.wmnet * 07:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1011.eqiad.wmnet with reason: host reimage * 07:27 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti3007.esams.wmnet * 07:25 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti3007.esams.wmnet * 07:25 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti3008.esams.wmnet * 07:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti3008.esams.wmnet * 07:22 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-web1001.eqiad.wmnet with OS bookworm * 07:19 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1003.eqiad.wmnet with OS bookworm * 07:18 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-master1003.eqiad.wmnet * 07:18 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host an-master1003.eqiad.wmnet * 07:17 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1011.eqiad.wmnet with OS bookworm * 07:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 07:16 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti3008.esams.wmnet * 07:15 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1009.eqiad.wmnet with OS bookworm * 07:14 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host an-master1003.eqiad.wmnet * 07:13 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-master1003.eqiad.wmnet * 07:13 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-master1003.eqiad.wmnet * 07:12 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-master1003.eqiad.wmnet * 07:11 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti3008.esams.wmnet * 07:07 arnaudb@dns1006: END - running authdns-update * 07:07 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti5007.eqsin.wmnet * 07:07 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti5007.eqsin.wmnet * 07:05 arnaudb@dns1006: START - running authdns-update * 06:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti5007.eqsin.wmnet * 06:54 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti5007.eqsin.wmnet * 06:38 moritzm: failover ganeti master in eqsin to ganeti5004 * 06:36 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti5006.eqsin.wmnet * 06:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti5006.eqsin.wmnet * 06:28 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti5006.eqsin.wmnet * 06:23 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti5006.eqsin.wmnet * 06:20 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti5005.eqsin.wmnet * 06:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti5005.eqsin.wmnet * 06:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti5005.eqsin.wmnet * 06:06 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti5005.eqsin.wmnet * 06:03 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti5004.eqsin.wmnet * 06:03 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti5004.eqsin.wmnet * 05:55 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti5004.eqsin.wmnet * 05:53 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti5004.eqsin.wmnet * 04:40 ryankemper: [[phab:T434494|T434494]] reimaged `an-tool1008.eqiad.wmnet` to bookworm; yarn.wikimedia.org is back up * 04:16 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-tool1008.eqiad.wmnet with OS bookworm * 03:58 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-tool1008.eqiad.wmnet with reason: host reimage * 03:53 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-tool1008.eqiad.wmnet with reason: host reimage * 03:41 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-tool1008.eqiad.wmnet with OS bookworm * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 45s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 00:25 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324427{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]], [[gerrit:1324429{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0]], [[gerrit:1324428{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]] (duration: 07m 55s) * 00:21 kemayo@deploy1003: kemayo: Continuing with deployment * 00:19 kemayo@deploy1003: kemayo: Backport for [[gerrit:1324427{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]], [[gerrit:1324429{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0]], [[gerrit:1324428{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:17 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1324427{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]], [[gerrit:1324429{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0]], [[gerrit:1324428{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]] == 2026-08-11 == * 21:37 sbassett: Deployed security fix for [[phab:T434521|T434521]] (wmf.15) * 21:29 sbassett: Deployed security fix for [[phab:T434521|T434521]] (wmf.14) * 21:19 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324370{{!}}Phase 4 of legal footer deployment (T432796)]], [[gerrit:1319804{{!}}Disable wgMFCustomSiteModules on English Wikipedia (T375538)]] (duration: 15m 26s) * 21:15 jdlrobson@deploy1003: jdlrobson: Continuing with deployment * 21:06 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1324370{{!}}Phase 4 of legal footer deployment (T432796)]], [[gerrit:1319804{{!}}Disable wgMFCustomSiteModules on English Wikipedia (T375538)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:03 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1324370{{!}}Phase 4 of legal footer deployment (T432796)]], [[gerrit:1319804{{!}}Disable wgMFCustomSiteModules on English Wikipedia (T375538)]] * 20:59 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 20:50 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324384{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]], [[gerrit:1324385{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]] (duration: 06m 58s) * 20:46 kemayo@deploy1003: kemayo: Continuing with deployment * 20:45 kemayo@deploy1003: kemayo: Backport for [[gerrit:1324384{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]], [[gerrit:1324385{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:43 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1324384{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]], [[gerrit:1324385{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]] * 20:42 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 20:42 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324386{{!}}build: Updating js-yaml to 3.15.1, 4.3.1]] (duration: 07m 36s) * 20:38 kemayo@deploy1003: kemayo: Continuing with deployment * 20:37 jhancock@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 20:37 kemayo@deploy1003: kemayo: Backport for [[gerrit:1324386{{!}}build: Updating js-yaml to 3.15.1, 4.3.1]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:35 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1324386{{!}}build: Updating js-yaml to 3.15.1, 4.3.1]] * 20:18 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 20:15 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 20:15 jhancock@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin1003" * 20:14 jhancock@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin1003" * 19:59 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 19:54 jhancock@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 19:10 brennen@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] (duration: 06m 41s) * 19:04 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns6001.wikimedia.org * 19:04 sukhe@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns6001.wikimedia.org * 19:04 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns5003.wikimedia.org * 19:04 sukhe@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns5003.wikimedia.org * 19:03 brennen@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 18:59 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns5003.wikimedia.org with OS trixie * 18:55 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns6001.wikimedia.org with OS trixie * 18:19 brett@cumin2002: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on P<nowiki>{</nowiki>cp7009.magru.wmnet<nowiki>}</nowiki> and A:cp - 9.2.15 Upgrade () * 18:14 brett@cumin2002: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on P<nowiki>{</nowiki>cp7009.magru.wmnet<nowiki>}</nowiki> and A:cp - 9.2.15 Upgrade () * 18:13 brennen@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 18:12 brett@cumin2002: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 9.2.15 Upgrade () * 18:09 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns5003.wikimedia.org with reason: host reimage * 18:06 brett@cumin2002: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 9.2.15 Upgrade () * 18:06 brennen: 1.47.0-wmf.15 train status ([[phab:T430834|T430834]]) - no current blockers, rolling to group0 * 18:05 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns5003.wikimedia.org with reason: host reimage * 18:05 brett: import trafficserver-9.2.15~deb13+wmf1 into trixie-wikimedia ([[phab:T434478|T434478]]) * 18:01 ladsgroup@cumin1003: END (PASS) - Cookbook sre.mysql.sanitarium_restart (exit_code=0) * 17:58 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns6001.wikimedia.org with reason: host reimage * 17:53 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324369{{!}}Enable desktop lazy loading on group0 (T148047)]] (duration: 07m 31s) * 17:52 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns6001.wikimedia.org with reason: host reimage * 17:49 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 17:49 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitarium_restart (exit_code=99) * 17:49 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 17:49 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7001.magru.wmnet * 17:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti7001.magru.wmnet * 17:48 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 17:47 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1324369{{!}}Enable desktop lazy loading on group0 (T148047)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:45 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1324369{{!}}Enable desktop lazy loading on group0 (T148047)]] * 17:39 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti7001.magru.wmnet * 17:36 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns5003.wikimedia.org with OS trixie * 17:34 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns6001.wikimedia.org with OS trixie * 17:31 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324361{{!}}Move FR config from IS.php to a dedicated file]], [[gerrit:1324363{{!}}Remove $wmg = $wg hacks in CentralAuth (T119117)]] (duration: 12m 23s) * 17:26 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 17:23 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1324361{{!}}Move FR config from IS.php to a dedicated file]], [[gerrit:1324363{{!}}Remove $wmg = $wg hacks in CentralAuth (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:19 sukhe: sudo cumin "A:cp-magru" "run-puppet-agent --enable 'merging CR 1324355'" * 17:18 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1324361{{!}}Move FR config from IS.php to a dedicated file]], [[gerrit:1324363{{!}}Remove $wmg = $wg hacks in CentralAuth (T119117)]] * 17:11 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-master1004.eqiad.wmnet with OS bookworm * 17:11 sukhe: sukhe@cp7005:~$ sudo puppet agent -tv * 17:02 sukhe: sudo cumin "A:cp-magru" "disable-puppet 'merging CR 1324355'" * 16:54 sukhe@dns1004: END - running authdns-update * 16:53 sukhe@dns1004: START - running authdns-update * 16:53 sukhe@dns1004: FAIL - running authdns-update * 16:51 sukhe@dns1004: START - running authdns-update * 16:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-master1004.eqiad.wmnet with reason: host reimage * 16:44 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-master1004.eqiad.wmnet with reason: host reimage * 16:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1157.eqiad.wmnet onto db1272.eqiad.wmnet * 16:40 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1157: Pool db1157.eqiad.wmnet in after cloning * 16:38 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 16:31 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324356{{!}}InitialiseSettings: Fix wgOATHAuthEnforce2FAForAll]] (duration: 06m 52s) * 16:30 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 16:28 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 16:27 reedy@deploy1003: reedy: Continuing with deployment * 16:26 reedy@deploy1003: reedy: Backport for [[gerrit:1324356{{!}}InitialiseSettings: Fix wgOATHAuthEnforce2FAForAll]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:24 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324356{{!}}InitialiseSettings: Fix wgOATHAuthEnforce2FAForAll]] * 16:13 sukhe: restart ntpsec.serviceon dns7001 * 16:09 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324335{{!}}InitialiseSettings: Enable 2FA enforcement on various private wikis (T428103)]] (duration: 06m 40s) * 16:08 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 16:06 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2204: Security update * 16:04 reedy@deploy1003: reedy: Continuing with deployment * 16:04 reedy@deploy1003: reedy: Backport for [[gerrit:1324335{{!}}InitialiseSettings: Enable 2FA enforcement on various private wikis (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:02 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7001.magru.wmnet * 16:02 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324335{{!}}InitialiseSettings: Enable 2FA enforcement on various private wikis (T428103)]] * 16:01 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 15:55 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1157: Pool db1157.eqiad.wmnet in after cloning * 15:54 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 15:54 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 15:41 dancy@deploy1003: Finished scap sync-world: Testing (duration: 06m 28s) * 15:40 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-master1004.eqiad.wmnet with OS bookworm * 15:40 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 15:35 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti4008.ulsfo.wmnet * 15:35 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti4008.ulsfo.wmnet * 15:34 dancy@deploy1003: Started scap sync-world: Testing * 15:34 dancy@deploy1003: Installation of scap version "4.279.0" completed for 3 hosts * 15:34 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-master1004.eqiad.wmnet with OS bookworm * 15:34 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 15:33 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-master1004.eqiad.wmnet with OS bookworm * 15:32 dancy@deploy1003: Installing scap version "4.279.0" for 3 host(s) * 15:32 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324339{{!}}Add /w/deployment-info.php entrypoint]] (duration: 07m 25s) * 15:30 moritzm: failover ganeti master in magru to ganeti7004 * 15:29 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti4008.ulsfo.wmnet * 15:28 dancy@deploy1003: dancy: Continuing with deployment * 15:28 tappof: remove 2026-05 swift log archives from centrallog to free some space ([[phab:T434502|T434502]]) * 15:27 dancy@deploy1003: dancy: Backport for [[gerrit:1324339{{!}}Add /w/deployment-info.php entrypoint]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:25 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324339{{!}}Add /w/deployment-info.php entrypoint]] * 15:20 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2204: Security update * 15:18 dancy@deploy1003: Installation of scap version "4.278.0" completed for 3 hosts * 15:18 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7004.magru.wmnet * 15:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti7004.magru.wmnet * 15:16 dancy@deploy1003: Installing scap version "4.278.0" for 3 host(s) * 15:14 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2204.codfw.wmnet with reason: Maintenance * 15:12 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 15:11 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-master1004.eqiad.wmnet with OS bookworm * 15:11 hashar: Restarting CI Jenkins on contint1003 due to Java upgrade. * 15:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2204 [[phab:T434565|T434565]]', diff saved to https://phabricator.wikimedia.org/P95984 and previous config saved to /var/cache/conftool/dbconfig/20260811-151126-cwilliams.json * 15:10 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti4008.ulsfo.wmnet * 15:10 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti7004.magru.wmnet * 15:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2207 to s2 primary [[phab:T434565|T434565]]', diff saved to https://phabricator.wikimedia.org/P95983 and previous config saved to /var/cache/conftool/dbconfig/20260811-150905-cwilliams.json * 15:08 cezmunsta: Starting s2 codfw failover from db2204 to db2207 - [[phab:T434565|T434565]] * 15:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2207 with weight 0 [[phab:T434565|T434565]]', diff saved to https://phabricator.wikimedia.org/P95982 and previous config saved to /var/cache/conftool/dbconfig/20260811-150402-cwilliams.json * 15:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s2 [[phab:T434565|T434565]] * 14:55 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1010.eqiad.wmnet with OS bookworm * 14:49 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns4003.wikimedia.org with OS trixie * 14:47 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1010.eqiad.wmnet with OS bookworm * 14:47 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7001.wikimedia.org with OS trixie * 14:44 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-presto1010.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:41 btullis@cumin1003: START - Cookbook sre.hosts.provision for host an-presto1010.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:40 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-presto1010.eqiad.wmnet with OS bookworm * 14:39 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 14:39 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-presto1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:36 btullis@cumin1003: START - Cookbook sre.hosts.provision for host an-presto1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:32 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1009.eqiad.wmnet with OS bookworm * 14:32 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 14:31 cwilliams@cumin1003: START - Cookbook sre.mysql.clone of db1157.eqiad.wmnet onto db1272.eqiad.wmnet * 14:30 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7004.magru.wmnet * 14:28 moritzm: failover ganeti master in ulsfo to ganeti4005 * 14:23 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7003.magru.wmnet * 14:23 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti7003.magru.wmnet * 14:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-master1004.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:22 btullis@cumin1003: START - Cookbook sre.hosts.provision for host an-master1004.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:21 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-master1004.eqiad.wmnet with OS bookworm * 14:19 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti4007.ulsfo.wmnet * 14:19 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti4007.ulsfo.wmnet * 14:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1010.eqiad.wmnet with OS bookworm * 14:17 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1008.eqiad.wmnet with OS bookworm * 14:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti7003.magru.wmnet * 14:12 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7003.magru.wmnet * 14:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti4007.ulsfo.wmnet * 14:11 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7002.magru.wmnet * 14:11 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti7002.magru.wmnet * 14:09 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324318{{!}}Revert "wmf-config/ProductionServices: set URL for urldownloader to service record" (T429175)]] (duration: 06m 46s) * 14:09 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7001.wikimedia.org with reason: host reimage * 14:06 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti4007.ulsfo.wmnet * 14:05 kharlan@deploy1003: kharlan: Continuing with deployment * 14:05 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti4006.ulsfo.wmnet * 14:04 jayme: updated calico to v3.30.7 on staging-eqiad - [[phab:T427400|T427400]] * 14:04 kharlan@deploy1003: kharlan: Backport for [[gerrit:1324318{{!}}Revert "wmf-config/ProductionServices: set URL for urldownloader to service record" (T429175)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti4006.ulsfo.wmnet * 14:03 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns4003.wikimedia.org with reason: host reimage * 14:03 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7001.wikimedia.org with reason: host reimage * 14:02 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti7002.magru.wmnet * 14:02 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1324318{{!}}Revert "wmf-config/ProductionServices: set URL for urldownloader to service record" (T429175)]] * 14:02 btullis@dns1004: FAIL - running authdns-update * 14:00 btullis@dns1004: START - running authdns-update * 13:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1008.eqiad.wmnet with reason: host reimage * 13:59 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'. * 13:59 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1313985{{!}}wmf-config/ProductionServices: set URL for urldownloader to service record (T429175)]] (duration: 25m 06s) * 13:58 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7002.magru.wmnet * 13:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti4006.ulsfo.wmnet * 13:57 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns4003.wikimedia.org with reason: host reimage * 13:56 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'. * 13:56 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7001.magru.wmnet * 13:56 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1008.eqiad.wmnet with reason: host reimage * 13:55 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'. * 13:55 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'. * 13:55 kharlan@deploy1003: kharlan, sukhe: Continuing with deployment * 13:53 marostegui: Failover ms2 [[phab:T434288|T434288]] * 13:52 marostegui: Failover ms1 [[phab:T434288|T434288]] * 13:52 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7001.magru.wmnet * 13:51 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti4006.ulsfo.wmnet * 13:48 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti4005.ulsfo.wmnet * 13:48 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti4005.ulsfo.wmnet * 13:44 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti4005.ulsfo.wmnet * 13:40 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1008.eqiad.wmnet with OS bookworm * 13:39 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns4003.wikimedia.org with OS trixie * 13:38 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns7001.wikimedia.org with OS trixie * 13:38 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1008.eqiad.wmnet with OS bookworm * 13:37 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti4005.ulsfo.wmnet * 13:36 kharlan@deploy1003: kharlan, sukhe: Backport for [[gerrit:1313985{{!}}wmf-config/ProductionServices: set URL for urldownloader to service record (T429175)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:34 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1313985{{!}}wmf-config/ProductionServices: set URL for urldownloader to service record (T429175)]] * 13:29 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2034.codfw.wmnet * 13:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2034.codfw.wmnet * 13:25 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.clone (exit_code=99) of db1157.eqiad.wmnet onto db1272.eqiad.wmnet * 13:25 cwilliams@cumin1003: START - Cookbook sre.mysql.clone of db1157.eqiad.wmnet onto db1272.eqiad.wmnet * 13:21 urbanecm@deploy1003: mwscript-k8s job started: namespaceDupes.php --wiki=frwiktionary --fix # [[phab:T415716|T415716]] * 13:21 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2034.codfw.wmnet * 13:20 urbanecm@deploy1003: mwscript-k8s job started: namespaceDupes.php --wiki=frwiktionary # [[phab:T415716|T415716]] * 13:19 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1323350{{!}}[tgwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T415307)]], [[gerrit:1322961{{!}}[slwiki] Revert temporary logo for Wikipedia 25 (Vector legacy + Vector 2022) (T414265)]], [[gerrit:1323827{{!}}[itwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T414320)]] (duration: 08m 00s) * 13:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-coord1003.eqiad.wmnet with OS bookworm * 13:15 urbanecm@deploy1003: urbanecm, superpes: Continuing with deployment * 13:13 urbanecm@deploy1003: urbanecm, superpes: Backport for [[gerrit:1323350{{!}}[tgwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T415307)]], [[gerrit:1322961{{!}}[slwiki] Revert temporary logo for Wikipedia 25 (Vector legacy + Vector 2022) (T414265)]], [[gerrit:1323827{{!}}[itwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T414320)]] synced to the testservers (see https://wiki * 13:11 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1323350{{!}}[tgwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T415307)]], [[gerrit:1322961{{!}}[slwiki] Revert temporary logo for Wikipedia 25 (Vector legacy + Vector 2022) (T414265)]], [[gerrit:1323827{{!}}[itwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T414320)]] * 13:11 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1323779{{!}}[ukwiki] Remove reviewer usergroup (T434252)]], [[gerrit:1323312{{!}}[frwiktionary] Add new Schème namespace and its talk (T415716)]] (duration: 06m 49s) * 13:10 marostegui@dns1004: END - running authdns-update * 13:08 marostegui@dns1004: START - running authdns-update * 13:07 marostegui@cumin1003: dbctl commit (dc=all): 'Repool ms2 [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95980 and previous config saved to /var/cache/conftool/dbconfig/20260811-130725-marostegui.json * 13:06 urbanecm@deploy1003: urbanecm, superpes: Continuing with deployment * 13:06 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1266 to dbctl [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95979 and previous config saved to /var/cache/conftool/dbconfig/20260811-130627-marostegui.json * 13:06 urbanecm@deploy1003: urbanecm, superpes: Backport for [[gerrit:1323779{{!}}[ukwiki] Remove reviewer usergroup (T434252)]], [[gerrit:1323312{{!}}[frwiktionary] Add new Schème namespace and its talk (T415716)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:04 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1323779{{!}}[ukwiki] Remove reviewer usergroup (T434252)]], [[gerrit:1323312{{!}}[frwiktionary] Add new Schème namespace and its talk (T415716)]] * 12:59 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db2253.codfw.wmnet,db[1151,1266].eqiad.wmnet with reason: Switching over ms2 * 12:54 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1157: Using as clone source * 12:53 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1157: Using as clone source * 12:51 marostegui@cumin1003: dbctl commit (dc=all): 'Depool ms2 [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95977 and previous config saved to /var/cache/conftool/dbconfig/20260811-125129-marostegui.json * 12:47 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 12:46 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 12:45 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 12:44 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'recommendation-api-ng' for release 'main' . * 12:44 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'recommendation-api-ng' for release 'main' . * 12:43 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'recommendation-api-ng' for release 'main' . * 12:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-coord1003.eqiad.wmnet with reason: host reimage * 12:43 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'ores-legacy' for release 'main' . * 12:42 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'ores-legacy' for release 'main' . * 12:42 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2165: Security update * 12:40 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-coord1003.eqiad.wmnet with reason: host reimage * 12:39 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'ores-legacy' for release 'main' . * 12:38 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' . * 12:38 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' . * 12:37 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' . * 12:34 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2034.codfw.wmnet * 12:30 jmm@cumin2002: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti-test2001.codfw.wmnet * 12:30 jmm@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host ganeti-test2001.codfw.wmnet * 12:25 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1179.eqiad.wmnet onto db1278.eqiad.wmnet * 12:25 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1179: Pool db1179.eqiad.wmnet in after cloning * 12:23 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-coord1003.eqiad.wmnet with OS bookworm * 12:22 moritzm: failover ganeti master in codfw/routed to ganeti2033 * 12:22 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2033.codfw.wmnet * 12:22 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2033.codfw.wmnet * 12:19 jmm@cumin2002: START - Cookbook sre.hosts.reboot-single for host ganeti-test2001.codfw.wmnet * 12:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 12:18 jmm@cumin2002: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti-test2001.codfw.wmnet * 12:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1008.eqiad.wmnet with OS bookworm * 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2033.codfw.wmnet * 12:07 moritzm: failover ganeti master in ganeti/test to ganeti-test2003 * 12:04 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 12:03 jmm@cumin2003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti4005.ulsfo.wmnet * 12:03 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti4005.ulsfo.wmnet * 12:00 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324283{{!}}Use maximum compression level in SqlBlobStore and SqlBagOStuff (T428377)]] (duration: 11m 37s) * 11:57 jmm@cumin2002: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti-test2002.codfw.wmnet * 11:57 jmm@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti-test2002.codfw.wmnet * 11:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2165: Security update * 11:54 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 11:52 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1324283{{!}}Use maximum compression level in SqlBlobStore and SqlBagOStuff (T428377)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:51 jmm@cumin2002: START - Cookbook sre.hosts.reboot-single for host ganeti-test2002.codfw.wmnet * 11:50 jmm@cumin2002: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti-test2002.codfw.wmnet * 11:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2165.codfw.wmnet with reason: Maintenance * 11:48 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@050d19e] (releasing): [[phab:T434186|T434186]] (duration: 01m 14s) * 11:48 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1324283{{!}}Use maximum compression level in SqlBlobStore and SqlBagOStuff (T428377)]] * 11:47 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@050d19e] (releasing): [[phab:T434186|T434186]] * 11:44 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@050d19e] (releasing): test jenkins deploy for [[phab:T434186|T434186]] (duration: 01m 08s) * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2165 [[phab:T434514|T434514]]', diff saved to https://phabricator.wikimedia.org/P95969 and previous config saved to /var/cache/conftool/dbconfig/20260811-114352-cwilliams.json * 11:43 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@050d19e] (releasing): test jenkins deploy for [[phab:T434186|T434186]] * 11:42 jmm@cumin2003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti-test2003.codfw.wmnet * 11:42 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti-test2003.codfw.wmnet * 11:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2161 to s8 primary [[phab:T434514|T434514]]', diff saved to https://phabricator.wikimedia.org/P95968 and previous config saved to /var/cache/conftool/dbconfig/20260811-114136-cwilliams.json * 11:40 cezmunsta: Starting s8 codfw failover from db2165 to db2161 - [[phab:T434514|T434514]] * 11:40 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1179: Pool db1179.eqiad.wmnet in after cloning * 11:36 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti-test2003.codfw.wmnet * 11:36 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti-test2003.codfw.wmnet * 11:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2161 with weight 0 [[phab:T434514|T434514]]', diff saved to https://phabricator.wikimedia.org/P95966 and previous config saved to /var/cache/conftool/dbconfig/20260811-113449-cwilliams.json * 11:34 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 25 hosts with reason: Primary switchover s8 [[phab:T434514|T434514]] * 11:29 moritzm: installing Python 3.11 security updates * 11:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-presto1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 11:26 btullis@cumin1003: START - Cookbook sre.hosts.provision for host an-presto1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 11:23 btullis@dns1004: END - running authdns-update * 11:21 btullis@dns1004: START - running authdns-update * 11:20 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-presto1008.eqiad.wmnet with OS bookworm * 11:20 moritzm: installing curl security updates * 11:11 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1007.eqiad.wmnet with OS bookworm * 10:45 tappof: bump space for prometheus k8s-dse in codfw * 10:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1007.eqiad.wmnet with reason: host reimage * 10:38 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1007.eqiad.wmnet with reason: host reimage * 10:37 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-coord1004.eqiad.wmnet with OS bookworm * 10:35 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1008.eqiad.wmnet with OS bookworm * 10:34 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1179: Depool db1179.eqiad.wmnet to then clone it to db1278.eqiad.wmnet - marostegui@cumin1003 * 10:34 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1008.eqiad.wmnet with OS bookworm * 10:25 fceratto@cumin1003: dbctl commit (dc=all): 'Remove db1177 [[phab:T433474|T433474]]', diff saved to https://phabricator.wikimedia.org/P95964 and previous config saved to /var/cache/conftool/dbconfig/20260811-102527-fceratto.json * 10:22 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1008.eqiad.wmnet with OS bookworm * 10:22 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1007.eqiad.wmnet with OS bookworm * 10:21 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1006.eqiad.wmnet with OS bookworm * 10:20 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 10:18 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1179: Depool db1179.eqiad.wmnet to then clone it to db1278.eqiad.wmnet - marostegui@cumin1003 * 10:18 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1179.eqiad.wmnet onto db1278.eqiad.wmnet * 10:17 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 10:17 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 10:14 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 10:09 blake@deploy1003: Stopping before sync operations * 10:09 blake@deploy1003: Started scap sync-world: Non-deployment run to populate release values for [[phab:T427668|T427668]] * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 10:04 fceratto@cumin1003: Removing db1177 from zarcillo [[phab:T433474|T433474]] * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1177.eqiad.wmnet * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1177.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:03 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1177.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:00 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1006.eqiad.wmnet with reason: host reimage * 09:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-coord1004.eqiad.wmnet with reason: host reimage * 09:57 marostegui: Failover m1 from db1164 to db1213 - [[phab:T434493|T434493]] * 09:57 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1006.eqiad.wmnet with reason: host reimage * 09:55 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:54 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2232].codfw.wmnet,db[1164,1213,1217].eqiad.wmnet with reason: Primary switchover m1 [[phab:T434493|T434493]] * 09:52 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-coord1004.eqiad.wmnet with reason: host reimage * 09:49 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1213.eqiad.wmnet with OS trixie * 09:49 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1177.eqiad.wmnet * 09:41 moritzm: installing Linux 6.12.101 on Trixie hosts * 09:40 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1006.eqiad.wmnet with OS bookworm * 09:35 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-coord1004.eqiad.wmnet with OS bookworm * 09:28 moritzm: installing node-tar security updates * 09:27 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1213.eqiad.wmnet with reason: host reimage * 09:22 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1213.eqiad.wmnet with reason: host reimage * 09:09 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1177: Decommission * 09:08 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db1177: Decommission * 09:08 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 09:08 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.decommission (exit_code=99) * 09:06 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1213.eqiad.wmnet with OS trixie * 09:06 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 09:05 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1213.eqiad.wmnet with reason: Reimage * 08:53 marostegui@dns1004: END - running authdns-update * 08:51 marostegui@dns1004: START - running authdns-update * 08:48 marostegui: Switchover ms1 master in eqiad [[phab:T434288|T434288]] * 08:48 marostegui@cumin1003: dbctl commit (dc=all): 'Repool ms1 [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95962 and previous config saved to /var/cache/conftool/dbconfig/20260811-084804-marostegui.json * 08:40 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1267 to dbctl [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95961 and previous config saved to /var/cache/conftool/dbconfig/20260811-084054-marostegui.json * 08:29 marostegui: Failover m1 from db1213 to db1164 - [[phab:T434043|T434043]] * 08:25 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2232].codfw.wmnet,db[1164,1213,1217].eqiad.wmnet with reason: Primary switchover m1 [[phab:T434043|T434043]] * 08:22 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db2251.codfw.wmnet,db[1152,1267].eqiad.wmnet with reason: Switching over ms1 * 08:22 marostegui@cumin1003: dbctl commit (dc=all): 'Depool ms1 [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95960 and previous config saved to /var/cache/conftool/dbconfig/20260811-082201-marostegui.json * 08:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: Switching over ms1 * 08:20 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.parsercache (exit_code=99) * 08:20 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 08:20 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1152: Switching over ms1 * 08:19 slyngshede@dns1004: END - running authdns-update * 08:18 moritzm: installing openjdk-21 security updates * 08:17 slyngshede@dns1004: START - running authdns-update * 08:16 moritzm: imported jenkins 2.568.2 to thirdparty/jenkins for trixie-wikimedia * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.12 (duration: 02m 26s) * 03:36 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] (duration: 33m 33s) * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 35s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-10 == * 14:54 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1323973{{!}}mmv.bootstrap: Fix getUrlParam to account for TIFF lossy/lossless param (T434333)]] (duration: 11m 24s) * 14:50 krinkle@deploy1003: krinkle: Continuing with deployment * 14:45 krinkle@deploy1003: krinkle: Backport for [[gerrit:1323973{{!}}mmv.bootstrap: Fix getUrlParam to account for TIFF lossy/lossless param (T434333)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:43 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1323973{{!}}mmv.bootstrap: Fix getUrlParam to account for TIFF lossy/lossless param (T434333)]] * 14:07 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1323967{{!}}updateIsActiveFlagForMentees: Commit the final partial batch (T432959)]] (duration: 10m 33s) * 13:56 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1323967{{!}}updateIsActiveFlagForMentees: Commit the final partial batch (T432959)]] * 13:45 dani@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply * 13:45 dani@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply * 13:45 dani@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply * 13:45 dani@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply * 13:45 dani@deploy1003: helmfile [staging] DONE helmfile.d/services/miscweb: apply * 13:44 dani@deploy1003: helmfile [staging] START helmfile.d/services/miscweb: apply * 13:38 wmde-fisch@deploy1003: Finished scap sync-world: Backport for [[gerrit:1323939{{!}}Enable sub-references on more group2 wikis (batch3) (T432731)]] (duration: 33m 21s) * 13:25 wmde-fisch@deploy1003: wmde-fisch: Continuing with deployment * 13:22 wmde-fisch@deploy1003: wmde-fisch: Backport for [[gerrit:1323939{{!}}Enable sub-references on more group2 wikis (batch3) (T432731)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:05 wmde-fisch@deploy1003: Started scap sync-world: Backport for [[gerrit:1323939{{!}}Enable sub-references on more group2 wikis (batch3) (T432731)]] * 07:57 hashar@deploy1003: Finished deploy [integration/docroot@7772132]: update build dependencies (duration: 00m 13s) * 07:57 hashar@deploy1003: Started deploy [integration/docroot@7772132]: update build dependencies * 07:35 _joe_: restarting squid on urldownloader1006 * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 48s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-09 == * 16:01 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:01 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:01 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:00 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 36s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-08 == * 05:31 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9] (wcqs): [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) (duration: 02m 36s) * 05:28 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9] (wcqs): [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) * 04:56 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 04:55 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 04:47 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) (duration: 19m 22s) * 04:28 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) * 04:19 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) (duration: 00m 06s) * 04:18 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) * 04:17 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) (duration: 00m 28s) * 04:16 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) * 03:52 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 03:52 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 34s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-07 == * 23:30 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:29 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 22:45 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 22:43 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 22:41 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 22:41 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 20:54 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 20:32 andrewbogott: restarting puppetserver service on puppetserver* for [[phab:T434339|T434339]] * 19:52 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:45 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 19:32 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:25 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:22 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 19:21 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 18:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:41 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 18:35 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 18:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 18:22 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 18:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 18:16 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 18:12 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:09 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:08 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:07 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:04 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:01 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:00 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:00 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 17:59 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 17:25 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 17:14 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 17:13 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 17:13 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 17:13 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:54 maryum: Deployed security fix for [[phab:T434278|T434278]] * 16:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 16:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 16:27 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --verbose --scoreLessThan=0.7 --exceptDatasetChecksums=[[phab:T434319|T434319]]-enwiki-models.txt # [[phab:T434319|T434319]] * 16:06 cdobbins@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp5022.eqsin.wmnet with OS trixie * 15:13 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 14:19 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 14:17 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 13:50 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1156.eqiad.wmnet onto db1271.eqiad.wmnet * 13:50 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1271: Pool db1271.eqiad.wmnet in after cloning * 13:02 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1271: Pool db1271.eqiad.wmnet in after cloning * 12:19 jayme: updated calico to v3.30.7 on staging-codfw - [[phab:T427400|T427400]] * 12:09 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 12:06 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 12:06 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 12:05 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 12:02 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1156: Pool db1156.eqiad.wmnet in after cloning * 11:38 bjensen: sudo -i reprepro -C main include trixie-wikimedia $<nowiki>{</nowiki>HOME<nowiki>}</nowiki>/httpbb/trixie/httpbb_$<nowiki>{</nowiki>VERSION?<nowiki>}</nowiki>-1+deb13u1_amd64.changes #[[phab:T434052|T434052]] * 11:35 bjensen: sudo -i reprepro -C main include bookworm-wikimedia $<nowiki>{</nowiki>HOME<nowiki>}</nowiki>/httpbb/bookworm/httpbb_$<nowiki>{</nowiki>VERSION?<nowiki>}</nowiki>-1_amd64.changes #[[phab:T434052|T434052]] * 11:30 marostegui@cumin1003: dbctl commit (dc=all): 'Adding db1271 to dbctl', diff saved to https://phabricator.wikimedia.org/P95945 and previous config saved to /var/cache/conftool/dbconfig/20260807-113006-marostegui.json * 11:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1156: Pool db1156.eqiad.wmnet in after cloning * 10:23 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 10:22 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 10:22 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 10:21 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 10:20 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 10:20 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 10:19 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 10:18 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 10:06 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on 21 hosts with reason: cloning * 10:01 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1156: Depool db1156.eqiad.wmnet to then clone it to db1271.eqiad.wmnet - marostegui@cumin1003 * 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1156: Depool db1156.eqiad.wmnet to then clone it to db1271.eqiad.wmnet - marostegui@cumin1003 * 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1156.eqiad.wmnet onto db1271.eqiad.wmnet * 09:15 jynus: started stress testing db1245 dbs [[phab:T431115|T431115]] * 08:19 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:18 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:16 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:14 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:13 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:10 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:06 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:05 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:00 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 10 days, 0:00:00 on ml-serve1015.eqiad.wmnet with reason: Downtime to get full picture of current BIOS settings beyond what Redfish shows * 08:00 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 07:54 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 07:54 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 07:53 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:52 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:51 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:50 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:49 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:48 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:47 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:45 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:45 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:41 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:38 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:37 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 06:35 jayme: updated istio to 1.29.4 on wikikube eqiad - [[phab:T427401|T427401]] * 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1178.eqiad.wmnet * 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1178.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 06:06 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1178.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 05:55 marostegui@cumin1003: START - Cookbook sre.dns.netbox * 05:49 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1178.eqiad.wmnet * 05:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 05:46 marostegui@cumin1003: Removing db1178 from zarcillo [[phab:T433471|T433471]] * 05:45 marostegui@cumin1003: START - Cookbook sre.mysql.decommission * 02:42 denisse: Extended volume on prometheus2008 for the disk space alert as per https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_host_running_out_of_space * 02:37 denisse: Extended volume on prometheus2007 tor the disk space alert as per https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_host_running_out_of_space * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 56s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-06 == * 21:39 maryum: Deploy security patch for [[phab:T433070|T433070]] * 21:29 maryum: Deploy security patch for [[phab:T434189|T434189]] * 20:48 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] (duration: 08m 12s) * 20:44 aude@deploy1003: lmora, aude, anzx: Continuing with deployment * 20:41 aude@deploy1003: lmora, aude, anzx: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be * 20:41 ebernhardson@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:41 ebernhardson@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 20:40 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] * 20:37 ebernhardson@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:37 ebernhardson@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 20:32 ebernhardson@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:32 ebernhardson@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 20:31 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] (duration: 06m 41s) * 20:27 cjming@deploy1003: cjming, ebernhardson, chlod: Continuing with deployment * 20:26 cjming@deploy1003: cjming, ebernhardson, chlod: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:24 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] * 20:18 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] (duration: 09m 22s) * 20:14 cjming@deploy1003: cjming, tsev: Continuing with deployment * 20:11 cjming@deploy1003: cjming, tsev: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:09 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] * 19:41 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply * 19:40 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply * 19:31 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 19:31 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 19:00 cdobbins@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cp5022.eqsin.wmnet with OS trixie * 18:25 ladsgroup@deploy1003: Finished scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) (duration: 06m 08s) * 18:19 ladsgroup@deploy1003: Started scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) * 18:18 ladsgroup@deploy1003: Stopping before sync operations * 18:17 ladsgroup@deploy1003: Started scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) * 17:55 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 16:50 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 16:35 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1001.eqiad.wmnet with OS bookworm * 16:19 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1002.eqiad.wmnet with reason: host reimage * 16:16 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1002.eqiad.wmnet with reason: host reimage * 16:05 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1001.eqiad.wmnet with reason: host reimage * 16:00 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1001.eqiad.wmnet with reason: host reimage * 15:57 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 15:43 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm * 15:29 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1001.eqiad.wmnet with OS bookworm * 15:29 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:58 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm * 14:57 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-drmrs ([[phab:T428495|T428495]]) * 14:55 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-drmrs ([[phab:T428495|T428495]]) * 14:55 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-ui1001.eqiad.wmnet with OS bookworm * 14:54 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-presto1001.eqiad.wmnet with OS bookworm * 14:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-magru ([[phab:T428495|T428495]]) * 14:49 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-magru ([[phab:T428495|T428495]]) * 14:48 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1001.eqiad.wmnet with OS bookworm * 14:46 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-esams ([[phab:T428495|T428495]]) * 14:44 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-esams ([[phab:T428495|T428495]]) * 14:43 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:42 brouberol@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:42 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:42 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 14:40 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 14:40 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:38 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-ui1001.eqiad.wmnet with reason: host reimage * 14:34 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-presto1001.eqiad.wmnet with reason: host reimage * 14:28 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-ui1001.eqiad.wmnet with reason: host reimage * 14:27 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-presto1001.eqiad.wmnet with reason: host reimage * 14:23 sukhe: sudo cumin -b2 'A:cp-text' "run-puppet-agent --enable 'merging CR 1290731'": [[phab:T425441|T425441]] * 14:18 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo for hosts in the wikimedia.org domain - [[phab:T428495|T428495]] * 14:16 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-presto1001.eqiad.wmnet with OS bookworm * 14:14 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-ui1001.eqiad.wmnet with OS bookworm * 14:12 sukhe: sudo cumin 'A:cp-text' "disable-puppet 'merging CR 1290731'": [[phab:T425441|T425441]] * 14:11 swfrench-wmf: restarted navtiming on webperf1003 - [[phab:T428495|T428495]] * 14:04 swfrench-wmf: begin rolling restart of confd in drmrs, eqiad, esams, magru - [[phab:T428495|T428495]] * 14:04 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm * 14:04 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:02 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-client1002.eqiad.wmnet with OS bookworm * 13:58 swfrench-wmf: authdns update to direct eqiad-associated etcd clients back to eqiad - [[phab:T428495|T428495]] * 13:58 swfrench@dns1004: END - running authdns-update * 13:56 swfrench@dns1004: START - running authdns-update * 13:49 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:44 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:31 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 13:29 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 13:26 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 13:23 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 13:22 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 13:19 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 13:18 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 13:18 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 13:17 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 13:16 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 13:13 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 13:11 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 13:09 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 13:06 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 13:06 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-client1002.eqiad.wmnet with OS bookworm * 13:05 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revision-models' for release 'main' . * 13:05 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:05 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revision-models' for release 'main' . * 13:04 brouberol@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-test-client1002.eqiad.wmnet with OS bookworm * 13:04 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 13:03 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 13:02 aikochou@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:00 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'readability' for release 'main' . * 12:59 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'readability' for release 'main' . * 12:58 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 12:57 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'logo-detection' for release 'main' . * 12:57 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'logo-detection' for release 'main' . * 12:57 aikochou@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 12:55 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:54 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:53 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 12:53 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:50 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 12:48 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 12:46 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'article-models' for release 'main' . * 12:45 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'article-models' for release 'main' . * 12:41 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . * 12:39 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . * 12:38 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-client1002.eqiad.wmnet with OS bookworm * 12:12 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply * 12:12 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply * 12:09 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:08 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 11:58 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2187: Security update * 11:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:24 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:16 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 11:15 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 11:10 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2187: Security update * 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2187.codfw.wmnet with reason: Maintenance * 10:56 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 10:56 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 10:56 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 10:56 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 10:54 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 10:53 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 10:09 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2187: Security update * 10:07 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2187: Security update * 09:39 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms2', diff saved to https://phabricator.wikimedia.org/P95929 and previous config saved to /var/cache/conftool/dbconfig/20260806-093908-marostegui.json * 09:36 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1178 from dbctl [[phab:T433471|T433471]]', diff saved to https://phabricator.wikimedia.org/P95928 and previous config saved to /var/cache/conftool/dbconfig/20260806-093632-marostegui.json * 09:33 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 09:31 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 09:30 topranks: bounce cr3-eqsin<->cr2-eqiad bgp session to disable no-prepend command * 09:20 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2253.codfw.wmnet,db1151.eqiad.wmnet with reason: cloning * 09:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1151: Cloning * 09:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:19 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 09:19 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1151: Cloning * 09:10 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 09:09 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup2003.codfw.wmnet * 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup2003.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 09:06 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup2003.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 09:03 klausman@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:02 jynus@cumin1003: START - Cookbook sre.dns.netbox * 09:02 klausman@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 08:57 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup2003.codfw.wmnet * 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup1003.eqiad.wmnet * 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:54 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms3', diff saved to https://phabricator.wikimedia.org/P95925 and previous config saved to /var/cache/conftool/dbconfig/20260806-085422-marostegui.json * 08:53 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:46 jynus@cumin1003: START - Cookbook sre.dns.netbox * 08:39 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup1003.eqiad.wmnet * 08:29 XioNoX: push pfw policy - [[phab:T434115|T434115]] * 08:14 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 08:00 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 08:00 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:58 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revision-models' for release 'main' . * 07:56 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 07:54 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'readability' for release 'main' . * 07:53 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'logo-detection' for release 'main' . * 07:51 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'llm' for release 'main' . * 07:48 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . * 07:37 jayme: updated istio to 1.29.4 on wikikube codfw - [[phab:T427401|T427401]] * 07:08 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2252.codfw.wmnet,db1153.eqiad.wmnet with reason: cloning * 07:07 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1153: Cloning * 07:07 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1153: Cloning * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 40s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-05 == * 23:24 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1009.eqiad.wmnet with OS bookworm * 23:03 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1009.eqiad.wmnet with reason: host reimage * 22:59 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1009.eqiad.wmnet with reason: host reimage * 22:43 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1009.eqiad.wmnet with OS bookworm * 22:38 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1009.eqiad.wmnet * 22:34 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1009.eqiad.wmnet * 22:25 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1008.eqiad.wmnet with OS bookworm * 22:04 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1008.eqiad.wmnet with reason: host reimage * 22:00 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1008.eqiad.wmnet with reason: host reimage * 21:48 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:47 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:46 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:44 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1008.eqiad.wmnet with OS bookworm * 21:43 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:41 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1008.eqiad.wmnet * 21:36 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1008.eqiad.wmnet * 21:14 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:12 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad * 21:12 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad * 21:10 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=eqiad * 21:08 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:07 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:07 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=eqiad * 21:04 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:03 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1006 * 21:02 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1006 * 21:00 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:56 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:56 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:55 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 20:55 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 20:55 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1007.eqiad.wmnet with OS bookworm * 20:51 vriley@cumin1003: START - Cookbook sre.dns.netbox * 20:43 ebernhardson: [[phab:T434008|T434008]]: changing cloudelastic:9643 from auto_expand_replicas to number_of_replicas * 20:34 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1007.eqiad.wmnet with reason: host reimage * 20:27 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1007.eqiad.wmnet with reason: host reimage * 20:24 cjming: end of UTC late backport window * 20:23 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] (duration: 06m 26s) * 20:18 cjming@deploy1003: cjming: Continuing with deployment * 20:18 cjming@deploy1003: cjming: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:16 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] * 20:12 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1007.eqiad.wmnet with OS bookworm * 20:12 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] (duration: 08m 41s) * 20:08 swfrench@cumin2002: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host conf1007.eqiad.wmnet with OS bookworm * 20:08 jforrester@deploy1003: jforrester: Continuing with deployment * 20:07 jforrester@deploy1003: jforrester: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:03 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] * 19:51 inflatador: [bking@puppetserver1001] ~$ sudo puppetserver ca sign --certname an-worker1189.eqiad.wmnet [[phab:T434142|T434142]] * 19:47 bking@cumin2003: DONE (FAIL) - Cookbook sre.puppet.renew-cert (exit_code=99) for an-worker1189.eqiad.wmnet: Renew puppet certificate - bking@cumin2003 * 19:46 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:30 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1007.eqiad.wmnet with OS trixie * 19:30 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 19:29 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 19:20 swfrench-wmf: silenced EtcdRelicationDown 0cb709a9-f244-4f1e-971f-{{Gerrit|440ec65e7fd7}} - [[phab:T428495|T428495]] * 19:13 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1007.eqiad.wmnet with OS bookworm * 19:12 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1007.eqiad.wmnet with reason: host reimage * 19:09 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1007.eqiad.wmnet * 19:07 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1007.eqiad.wmnet with reason: host reimage * 19:03 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1007.eqiad.wmnet * 18:52 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1007.eqiad.wmnet with OS trixie * 18:52 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1007.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:35 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 18:34 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 18:34 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 18:30 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1007.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:28 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:28 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1007] - vriley@cumin1003" * 18:27 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1007] - vriley@cumin1003" * 18:23 vriley@cumin1003: START - Cookbook sre.dns.netbox * 18:22 vriley@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 18:22 robh@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:19 vriley@cumin1003: START - Cookbook sre.dns.netbox * 18:13 robh@cumin2002: START - Cookbook sre.hosts.provision for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:31 jasmine@cumin2002: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-main-eqiad * 17:12 mutante: LDAP - added vwalters to group ciadmin - [[phab:T433615|T433615]] * 16:58 aokoth@deploy1003: Finished deploy [phabricator/deployment@e2ebca5]: Deploy Phab (duration: 00m 34s) * 16:57 aokoth@deploy1003: Started deploy [phabricator/deployment@e2ebca5]: Deploy Phab * 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad * 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=eqiad * 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad * 16:53 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:41 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2187.codfw.wmnet * 16:41 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2187.codfw.wmnet * 16:41 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker2187.codfw.wmnet * 16:41 cgoubert@cumin2003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker2187.codfw.wmnet * 16:40 jasmine@cumin2002: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-main-eqiad * 16:40 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:34 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:25 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-magru and A:liberica ([[phab:T428495|T428495]]) * 16:23 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-magru and A:liberica ([[phab:T428495|T428495]]) * 16:20 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-drmrs and A:liberica ([[phab:T428495|T428495]]) * 16:19 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-drmrs and A:liberica ([[phab:T428495|T428495]]) * 16:18 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-esams and A:liberica ([[phab:T428495|T428495]]) * 16:16 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-esams and A:liberica ([[phab:T428495|T428495]]) * 16:06 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1159.eqiad.wmnet * 16:06 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1159.eqiad.wmnet * 16:06 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1159.eqiad.wmnet * 16:05 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] (duration: 09m 11s) * 15:58 reedy@deploy1003: reedy: Continuing with deployment * 15:58 reedy@deploy1003: reedy: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:56 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] * 15:54 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1159.eqiad.wmnet with OS trixie * 15:38 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:33 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1159.eqiad.wmnet with reason: host reimage * 15:32 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:27 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1159.eqiad.wmnet with reason: host reimage * 15:10 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1159 * 15:10 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1159 * 15:00 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo for hosts in the wikimedia.org domain - [[phab:T428495|T428495]] * 14:55 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS trixie * 14:54 swfrench-wmf: restarted navtiming on webperf1003 - [[phab:T428495|T428495]] * 14:52 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1159 * 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1159.eqiad.wmnet 129.48.64.10.in-addr.arpa 9.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:52 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1159.eqiad.wmnet 129.48.64.10.in-addr.arpa 9.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1159 - jayme@cumin1003" * 14:52 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1159 - jayme@cumin1003" * 14:49 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:48 jayme@cumin1003: START - Cookbook sre.dns.netbox * 14:47 swfrench-wmf: begin rolling restart of confd in drmrs, eqiad, esams, magru - [[phab:T428495|T428495]] * 14:47 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1159 * 14:46 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:46 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:46 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1159.eqiad.wmnet with OS trixie * 14:45 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:44 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:44 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1159.eqiad.wmnet * 14:43 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:43 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1159.eqiad.wmnet * 14:43 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:43 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1159.eqiad.wmnet * 14:43 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:43 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1157.eqiad.wmnet * 14:43 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1157.eqiad.wmnet * 14:43 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1157.eqiad.wmnet * 14:42 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:42 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:42 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:41 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:41 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:41 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:41 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1003.eqiad.wmnet with OS bookworm * 14:39 swfrench-wmf: authdns update to direct eqiad-associated etcd clients to codfw - [[phab:T428495|T428495]] * 14:39 swfrench@dns1004: END - running authdns-update * 14:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 14:37 swfrench@dns1004: START - running authdns-update * 14:37 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:37 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:35 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:35 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 14:28 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:27 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1157.eqiad.wmnet with OS trixie * 14:27 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:27 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:27 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 14:26 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 14:26 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:26 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 14:26 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 14:26 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:26 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host search-loader1002.eqiad.wmnet with OS trixie * 14:19 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:19 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:15 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1003.eqiad.wmnet with reason: host reimage * 14:14 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:14 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:13 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 * 14:13 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host mc2046 * 14:13 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS trixie * 14:11 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1003.eqiad.wmnet with reason: host reimage * 14:10 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:09 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:09 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:09 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:08 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:08 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1157.eqiad.wmnet with reason: host reimage * 14:08 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:04 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 14:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on search-loader1002.eqiad.wmnet with reason: host reimage * 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=eqiad * 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=eqiad * 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=eqiad * 14:00 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:59 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:58 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1157.eqiad.wmnet with reason: host reimage * 13:57 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on search-loader1002.eqiad.wmnet with reason: host reimage * 13:54 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1003.eqiad.wmnet with OS bookworm * 13:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host search-loader1002.eqiad.wmnet with OS trixie * 13:43 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1157 * 13:42 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1157 * 13:40 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1157 * 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1157.eqiad.wmnet 183.32.64.10.in-addr.arpa 3.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1157.eqiad.wmnet 183.32.64.10.in-addr.arpa 3.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1157 - jayme@cumin1003" * 13:39 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1157 - jayme@cumin1003" * 13:39 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] (duration: 07m 00s) * 13:35 jayme@cumin1003: START - Cookbook sre.dns.netbox * 13:35 reedy@deploy1003: reedy: Continuing with deployment * 13:34 reedy@deploy1003: reedy: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:32 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] * 13:23 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1157 * 13:22 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1157.eqiad.wmnet with OS trixie * 13:22 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1157.eqiad.wmnet * 13:22 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1157.eqiad.wmnet * 13:21 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1157.eqiad.wmnet * 13:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1156.eqiad.wmnet * 13:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1156.eqiad.wmnet * 13:15 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1156.eqiad.wmnet * 13:01 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1156.eqiad.wmnet with OS trixie * 12:42 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1156.eqiad.wmnet with reason: host reimage * 12:38 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1156.eqiad.wmnet with reason: host reimage * 12:32 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 12:31 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 12:30 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 12:28 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 12:26 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 12:24 topranks: update bgp confed settings in eqsin * 12:22 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1156 * 12:22 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1156 * 12:22 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 12:19 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1156 * 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1156.eqiad.wmnet 110.32.64.10.in-addr.arpa 0.1.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:19 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1156.eqiad.wmnet 110.32.64.10.in-addr.arpa 0.1.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1156 - jayme@cumin1003" * 12:19 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1156 - jayme@cumin1003" * 12:17 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:14 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS trixie * 12:09 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:06 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:04 jayme@cumin1003: START - Cookbook sre.dns.netbox * 12:04 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 12:02 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'article-models' for release 'main' . * 12:01 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1156 * 12:01 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1156.eqiad.wmnet with OS trixie * 11:59 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1156.eqiad.wmnet * 11:59 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1156.eqiad.wmnet * 11:59 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1156.eqiad.wmnet * 11:57 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 11:53 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 11:53 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:52 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:52 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:50 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:50 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:50 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:49 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:48 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:47 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:47 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:45 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:45 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:44 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:44 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:44 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:43 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:42 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:38 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:35 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 * 11:35 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host mc2046 * 11:34 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS trixie * 11:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:27 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:21 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:21 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:18 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:18 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:18 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:18 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:13 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:13 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:09 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:08 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:07 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:06 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:06 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:05 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:05 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:04 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 11:04 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:24 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:24 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:17 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:16 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1155.eqiad.wmnet * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1155.eqiad.wmnet * 10:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1155.eqiad.wmnet * 10:14 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:14 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:11 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:11 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:10 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:09 aikochou@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop: sync * 10:09 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:09 aikochou@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop: sync * 10:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:07 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:05 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:05 aikochou@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop: sync * 10:05 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:05 aikochou@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop: sync * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:04 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:04 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:04 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1155.eqiad.wmnet with OS trixie * 09:52 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms1', diff saved to https://phabricator.wikimedia.org/P95918 and previous config saved to /var/cache/conftool/dbconfig/20260805-095212-marostegui.json * 09:44 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1152: after cloning * 09:44 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.parsercache (exit_code=99) * 09:44 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 09:44 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1152: after cloning * 09:43 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1155.eqiad.wmnet with reason: host reimage * 09:40 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1155.eqiad.wmnet with reason: host reimage * 09:32 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 09:32 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:31 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 09:31 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:27 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1155 * 09:27 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1155 * 09:25 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 09:24 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 09:24 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 09:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:23 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 09:23 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 09:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:22 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 09:22 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 09:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:20 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 09:20 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:17 XioNoX: push pfw policies - [[phab:T434038|T434038]] * 09:14 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1155 * 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1155.eqiad.wmnet 109.32.64.10.in-addr.arpa 9.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1155.eqiad.wmnet 109.32.64.10.in-addr.arpa 9.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1155 - jayme@cumin1003" * 09:14 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1155 - jayme@cumin1003" * 09:10 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 09:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:09 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2251.codfw.wmnet,db1152.eqiad.wmnet with reason: cloning * 09:09 jayme@cumin1003: START - Cookbook sre.dns.netbox * 09:08 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 09:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: Cloning * 09:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:05 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 09:05 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1152: Cloning * 08:38 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1155 * 08:37 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1155.eqiad.wmnet with OS trixie * 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1171.eqiad.wmnet * 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1171.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:29 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95913 and previous config saved to /var/cache/conftool/dbconfig/20260805-082908-ladsgroup.json * 08:27 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1171.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:22 jynus@cumin1003: START - Cookbook sre.dns.netbox * 08:18 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249', diff saved to https://phabricator.wikimedia.org/P95912 and previous config saved to /var/cache/conftool/dbconfig/20260805-081823-ladsgroup.json * 08:17 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1171.eqiad.wmnet * 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1150.eqiad.wmnet * 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1150.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:15 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1150.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:15 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 08:14 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1155.eqiad.wmnet * 08:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1155.eqiad.wmnet * 08:14 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1155.eqiad.wmnet * 08:11 jynus@cumin1003: START - Cookbook sre.dns.netbox * 08:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249', diff saved to https://phabricator.wikimedia.org/P95911 and previous config saved to /var/cache/conftool/dbconfig/20260805-080737-ladsgroup.json * 08:05 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1150.eqiad.wmnet * 08:02 marostegui: Depool clouddb1020 (s5,s8) [[phab:T434048|T434048]] * 08:02 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1020.eqiad.wmnet,service=s8 * 08:02 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1020.eqiad.wmnet,service=s5 * 08:02 marostegui: Depool clouddb1018 (s2,s7) [[phab:T434048|T434048]] * 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1018.eqiad.wmnet,service=s7 * 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1018.eqiad.wmnet,service=s2 * 08:01 marostegui: Depool clouddb1017 (s1) [[phab:T434048|T434048]] * 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1017.eqiad.wmnet,service=s1 * 07:59 marostegui: Depool clouddb1016 (s5,s8) [[phab:T434048|T434048]] * 07:59 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s8 * 07:59 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s5 * 07:57 marostegui: Depool clouddb1015 (s4,s6) [[phab:T434048|T434048]] * 07:57 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s6 * 07:57 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s4 * 07:56 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95910 and previous config saved to /var/cache/conftool/dbconfig/20260805-075650-ladsgroup.json * 07:54 marostegui: Depool clouddb1014 (s2,s7) [[phab:T434048|T434048]] * 07:54 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1014.eqiad.wmnet,service=s7 * 07:54 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1014.eqiad.wmnet,service=s2 * 07:53 marostegui: Depool clouddb1013:s1 [[phab:T434048|T434048]] * 07:53 marostegui: Depool clouddb1013:s1 [[phab:T409557|T409557]] * 07:53 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1013.eqiad.wmnet,service=s1 * 07:25 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95909 and previous config saved to /var/cache/conftool/dbconfig/20260805-072529-ladsgroup.json * 07:24 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2249.codfw.wmnet with reason: Maintenance * 07:24 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95908 and previous config saved to /var/cache/conftool/dbconfig/20260805-072426-ladsgroup.json * 07:21 slyngshede@dns1004: END - running authdns-update * 07:19 slyngshede@dns1004: START - running authdns-update * 07:13 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231', diff saved to https://phabricator.wikimedia.org/P95906 and previous config saved to /var/cache/conftool/dbconfig/20260805-071340-ladsgroup.json * 07:02 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231', diff saved to https://phabricator.wikimedia.org/P95905 and previous config saved to /var/cache/conftool/dbconfig/20260805-070253-ladsgroup.json * 06:52 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95904 and previous config saved to /var/cache/conftool/dbconfig/20260805-065206-ladsgroup.json * 06:45 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 06:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95903 and previous config saved to /var/cache/conftool/dbconfig/20260805-062240-ladsgroup.json * 06:21 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2231.codfw.wmnet with reason: Maintenance * 06:21 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95902 and previous config saved to /var/cache/conftool/dbconfig/20260805-062137-ladsgroup.json * 06:10 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215', diff saved to https://phabricator.wikimedia.org/P95901 and previous config saved to /var/cache/conftool/dbconfig/20260805-061051-ladsgroup.json * 06:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215', diff saved to https://phabricator.wikimedia.org/P95900 and previous config saved to /var/cache/conftool/dbconfig/20260805-060004-ladsgroup.json * 05:49 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95899 and previous config saved to /var/cache/conftool/dbconfig/20260805-054918-ladsgroup.json * 05:19 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95898 and previous config saved to /var/cache/conftool/dbconfig/20260805-051939-ladsgroup.json * 05:18 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2215.codfw.wmnet with reason: Maintenance * 04:30 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2201.codfw.wmnet with reason: Maintenance * 03:40 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2197.codfw.wmnet with reason: Maintenance * 03:40 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95897 and previous config saved to /var/cache/conftool/dbconfig/20260805-034036-ladsgroup.json * 03:29 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196', diff saved to https://phabricator.wikimedia.org/P95896 and previous config saved to /var/cache/conftool/dbconfig/20260805-032948-ladsgroup.json * 03:19 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196', diff saved to https://phabricator.wikimedia.org/P95895 and previous config saved to /var/cache/conftool/dbconfig/20260805-031902-ladsgroup.json * 03:08 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95894 and previous config saved to /var/cache/conftool/dbconfig/20260805-030815-ladsgroup.json * 02:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95893 and previous config saved to /var/cache/conftool/dbconfig/20260805-023413-ladsgroup.json * 02:33 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2196.codfw.wmnet with reason: Maintenance * 02:33 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95892 and previous config saved to /var/cache/conftool/dbconfig/20260805-023310-ladsgroup.json * 02:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186', diff saved to https://phabricator.wikimedia.org/P95891 and previous config saved to /var/cache/conftool/dbconfig/20260805-022223-ladsgroup.json * 02:11 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186', diff saved to https://phabricator.wikimedia.org/P95890 and previous config saved to /var/cache/conftool/dbconfig/20260805-021137-ladsgroup.json * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 02:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95889 and previous config saved to /var/cache/conftool/dbconfig/20260805-020051-ladsgroup.json * 01:30 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95888 and previous config saved to /var/cache/conftool/dbconfig/20260805-013029-ladsgroup.json * 01:29 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2186.codfw.wmnet with reason: Maintenance * 00:34 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on dbstore1009.eqiad.wmnet with reason: Maintenance * 00:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95887 and previous config saved to /var/cache/conftool/dbconfig/20260805-003408-ladsgroup.json * 00:23 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264', diff saved to https://phabricator.wikimedia.org/P95886 and previous config saved to /var/cache/conftool/dbconfig/20260805-002322-ladsgroup.json * 00:12 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264', diff saved to https://phabricator.wikimedia.org/P95885 and previous config saved to /var/cache/conftool/dbconfig/20260805-001235-ladsgroup.json * 00:01 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95884 and previous config saved to /var/cache/conftool/dbconfig/20260805-000148-ladsgroup.json == 2026-08-04 == * 23:45 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95883 and previous config saved to /var/cache/conftool/dbconfig/20260804-234508-ladsgroup.json * 23:44 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1264.eqiad.wmnet with reason: Maintenance * 23:44 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95882 and previous config saved to /var/cache/conftool/dbconfig/20260804-234405-ladsgroup.json * 23:33 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237', diff saved to https://phabricator.wikimedia.org/P95881 and previous config saved to /var/cache/conftool/dbconfig/20260804-233317-ladsgroup.json * 23:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237', diff saved to https://phabricator.wikimedia.org/P95880 and previous config saved to /var/cache/conftool/dbconfig/20260804-232230-ladsgroup.json * 23:11 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95879 and previous config saved to /var/cache/conftool/dbconfig/20260804-231144-ladsgroup.json * 22:23 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95878 and previous config saved to /var/cache/conftool/dbconfig/20260804-222345-ladsgroup.json * 22:23 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1237.eqiad.wmnet with reason: Maintenance * 21:13 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1225.eqiad.wmnet with reason: Maintenance * 20:40 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] (duration: 24m 40s) * 20:33 samtar@deploy1003: samtar, kineticpelagic: Continuing with deployment * 20:28 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS bookworm * 20:21 samtar@deploy1003: samtar, kineticpelagic: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:15 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] * 20:13 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 20:09 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 20:00 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1216.eqiad.wmnet with reason: Maintenance * 20:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95877 and previous config saved to /var/cache/conftool/dbconfig/20260804-195957-ladsgroup.json * 19:51 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 19:50 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) mc2046.codfw.wmnet 120.16.192.10.in-addr.arpa 0.2.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:50 jhancock@cumin2002: START - Cookbook sre.dns.wipe-cache mc2046.codfw.wmnet 120.16.192.10.in-addr.arpa 0.2.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host mc2046 - jhancock@cumin2002" * 19:50 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host mc2046 - jhancock@cumin2002" * 19:49 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203', diff saved to https://phabricator.wikimedia.org/P95876 and previous config saved to /var/cache/conftool/dbconfig/20260804-194911-ladsgroup.json * 19:46 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 19:45 jhancock@cumin2002: START - Cookbook sre.hosts.move-vlan for host mc2046 * 19:45 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS bookworm * 19:38 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203', diff saved to https://phabricator.wikimedia.org/P95875 and previous config saved to /var/cache/conftool/dbconfig/20260804-193825-ladsgroup.json * 19:27 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95874 and previous config saved to /var/cache/conftool/dbconfig/20260804-192738-ladsgroup.json * 19:02 mutante: gerrit ssh -p 29418 gerrit.wikimedia.org gerrit index changes {{Gerrit|1320979}} * 18:20 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 18:18 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 18:14 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 18:14 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 18:13 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 18:10 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 18:08 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 18:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95872 and previous config saved to /var/cache/conftool/dbconfig/20260804-180721-ladsgroup.json * 18:07 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 18:06 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1203.eqiad.wmnet with reason: Maintenance * 18:06 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95871 and previous config saved to /var/cache/conftool/dbconfig/20260804-180618-ladsgroup.json * 17:55 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179', diff saved to https://phabricator.wikimedia.org/P95870 and previous config saved to /var/cache/conftool/dbconfig/20260804-175531-ladsgroup.json * 17:55 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1154.eqiad.wmnet * 17:55 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1154.eqiad.wmnet * 17:55 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1154.eqiad.wmnet * 17:50 swfrench@deploy1003: Finished scap sync-world: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] (duration: 04m 05s) * 17:48 swfrench@deploy1003: swfrench: Continuing with deployment * 17:46 swfrench@deploy1003: swfrench: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:45 swfrench@deploy1003: Started scap sync-world: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] * 17:44 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179', diff saved to https://phabricator.wikimedia.org/P95869 and previous config saved to /var/cache/conftool/dbconfig/20260804-174445-ladsgroup.json * 17:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95868 and previous config saved to /var/cache/conftool/dbconfig/20260804-173359-ladsgroup.json * 17:33 swfrench@deploy1003: Finished scap sync-world: Pick up new PHP production image (duration: 28m 32s) * 17:28 aokoth@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on phab1005.eqiad.wmnet with reason: Puppet Failure * 17:05 swfrench@deploy1003: Started scap sync-world: Pick up new PHP production image * 17:00 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 17:00 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 16:54 cgoubert@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on wikikube-worker2187.codfw.wmnet with reason: Hardware issue * 16:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2187.codfw.wmnet * 16:52 mutante: gerrit2003:/var/log/apache2# ln -s /srv/gerrit/site_path/review_site/logs/ gerrit * 16:52 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2187.codfw.wmnet * 16:48 mutante: gerrit2003 - moving old apache logfiles older than 60 days from /var/log/apache2 to /srv/gerrit/site_path/review_site/logs/old/ * 16:33 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 16:32 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 16:29 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 16:29 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 16:28 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 16:28 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 16:27 dzahn@cumin1003: END (PASS) - Cookbook sre.gerrit.restart-gerrit (exit_code=0) Restarting Gerrit on gerrit2003 * 16:27 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 16:27 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95867 and previous config saved to /var/cache/conftool/dbconfig/20260804-162736-ladsgroup.json * 16:27 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 16:26 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1179.eqiad.wmnet with reason: Maintenance * 16:26 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:25 mutante: restarting gerrit - dropped outdated RSA host key * 16:25 dzahn@cumin1003: START - Cookbook sre.gerrit.restart-gerrit Restarting Gerrit on gerrit2003 * 16:24 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95866 and previous config saved to /var/cache/conftool/dbconfig/20260804-162424-ladsgroup.json * 16:24 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 16:23 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 16:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95865 and previous config saved to /var/cache/conftool/dbconfig/20260804-162236-ladsgroup.json * 16:21 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1179.eqiad.wmnet with reason: Maintenance * 16:17 swfrench-wmf: reprepro include php8.3_8.3.33-1+wmf11u1 into component/php83 for bullseye-wikimedia * 16:17 swfrench-wmf: reprepro include php8.3_8.3.33-1+wmf12u1 into component/php83 for bookworm-wikimedia * 16:11 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply * 16:10 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply * 16:10 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mobileapps: apply * 16:09 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mobileapps: apply * 16:09 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply * 16:08 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply * 16:08 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:08 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:07 aokoth@cumin1003: END (PASS) - Cookbook sre.vrts.upgrade (exit_code=0) on VRTS host vrts1003.eqiad.wmnet * 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:05 aokoth@cumin1003: START - Cookbook sre.vrts.upgrade on VRTS host vrts1003.eqiad.wmnet * 16:04 mutante: gerrit2002/gerrit1003/gerrit2003 - rm /etc/gerrit/ssh_host_rsa_key * 15:59 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:59 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:59 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:59 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:56 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 15:55 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:55 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:55 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:49 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 15:49 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:44 Raine: add php8.5 packages to component/php85 - [[phab:T432983|T432983]] * 15:39 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:33 aaron@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 15:33 aaron@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 15:29 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:19 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:19 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:16 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:16 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1154.eqiad.wmnet with OS trixie * 15:16 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:15 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:15 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:06 brennen@deploy1003: Finished deploy [phabricator/deployment@56f4ffd]: deploy phab1004 for [[phab:T433981|T433981]] (duration: 00m 43s) * 15:05 brennen@deploy1003: Started deploy [phabricator/deployment@56f4ffd]: deploy phab1004 for [[phab:T433981|T433981]] * 15:05 aaron@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 15:04 aaron@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 15:02 brennen@deploy1003: Finished deploy [phabricator/deployment@56f4ffd]: deploy phab2003 for [[phab:T433981|T433981]] (duration: 00m 51s) * 15:01 brennen@deploy1003: Started deploy [phabricator/deployment@56f4ffd]: deploy phab2003 for [[phab:T433981|T433981]] * 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1004.eqiad.wmnet with reason: deployment * 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1005.eqiad.wmnet with reason: deployment * 14:58 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab2003.codfw.wmnet with reason: deployment * 14:55 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1154.eqiad.wmnet with reason: host reimage * 14:51 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1154.eqiad.wmnet with reason: host reimage * 14:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 14:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 14:38 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync * 14:38 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync * 14:38 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync * 14:37 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync * 14:37 ottomata: roll restart eventgate-main to pick up stream config change - [[phab:T433507|T433507]] * 14:37 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-main: sync * 14:36 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-main: sync * 14:36 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1154 * 14:36 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1154 * 14:34 otto@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] (duration: 08m 39s) * 14:34 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1154 * 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1154.eqiad.wmnet 108.32.64.10.in-addr.arpa 8.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:34 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1154.eqiad.wmnet 108.32.64.10.in-addr.arpa 8.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1154 - jayme@cumin1003" * 14:34 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1154 - jayme@cumin1003" * 14:30 otto@deploy1003: otto: Continuing with deployment * 14:30 jayme@cumin1003: START - Cookbook sre.dns.netbox * 14:28 otto@deploy1003: otto: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:26 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1154 * 14:26 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1154.eqiad.wmnet with OS trixie * 14:26 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1154.eqiad.wmnet * 14:26 otto@deploy1003: Started scap sync-world: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] * 14:26 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1154.eqiad.wmnet * 14:26 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1154.eqiad.wmnet * 14:17 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 14:16 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 14:15 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 14:14 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 14:13 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 14:13 swfrench@dns1004: END - running authdns-update * 14:13 Msz2001: Finished deployments for UTC afternoon backport window * 14:13 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 14:13 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] (duration: 07m 58s) * 14:11 swfrench@dns1004: START - running authdns-update * 14:08 mszwarc@deploy1003: javiermonton, mszwarc, mpostoronca: Continuing with deployment * 14:07 mszwarc@deploy1003: javiermonton, mszwarc, mpostoronca: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] synced to the testser * 14:05 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] * 14:03 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 13:49 swfrench@cumin2002: conftool action : set/pooled=yes; selector: name=wikikube-worker2330.codfw.wmnet * 13:49 swfrench@cumin2002: conftool action : set/pooled=no; selector: name=wikikube-worker2330.codfw.wmnet * 13:48 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] (duration: 09m 19s) * 13:45 swfrench@dns1004: END - running authdns-update * 13:44 mszwarc@deploy1003: mszwarc, jforrester: Continuing with deployment * 13:43 swfrench@dns1004: START - running authdns-update * 13:41 mszwarc@deploy1003: mszwarc, jforrester: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:38 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] * 13:33 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 13:33 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1154.eqiad.wmnet * 13:32 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 13:32 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 13:31 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 13:31 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:31 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:29 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1154.eqiad.wmnet * 13:28 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1154.eqiad.wmnet * 13:28 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1154.eqiad.wmnet * 13:28 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1141.eqiad.wmnet * 13:28 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1141.eqiad.wmnet * 13:28 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1141.eqiad.wmnet * 13:22 otto@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply * 13:22 otto@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply * 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1096.eqiad.wmnet with OS trixie * 13:05 swfrench@dns1004: END - running authdns-update * 13:03 swfrench@dns1004: START - running authdns-update * 12:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 12:43 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 1:00:00 on db1171.eqiad.wmnet with reason: decom * 12:42 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 1:00:00 on db1150.eqiad.wmnet with reason: decom * 12:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 12:38 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1164,1217].eqiad.wmnet with reason: cloning * 12:33 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2096.codfw.wmnet with OS trixie * 12:22 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1096.eqiad.wmnet with OS trixie * 12:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2096.codfw.wmnet with reason: host reimage * 12:14 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1141.eqiad.wmnet with OS trixie * 12:10 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2096.codfw.wmnet with reason: host reimage * 12:10 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1289.eqiad.wmnet * 12:05 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1289.eqiad.wmnet * 12:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1288.eqiad.wmnet * 11:59 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1288.eqiad.wmnet * 11:59 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1287.eqiad.wmnet * 11:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1097.eqiad.wmnet with OS trixie * 11:54 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1287.eqiad.wmnet * 11:54 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1286.eqiad.wmnet * 11:53 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1141.eqiad.wmnet with reason: host reimage * 11:51 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2096.codfw.wmnet with OS trixie * 11:49 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1141.eqiad.wmnet with reason: host reimage * 11:48 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1286.eqiad.wmnet * 11:48 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1284.eqiad.wmnet * 11:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2095.codfw.wmnet with OS trixie * 11:43 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1284.eqiad.wmnet * 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1283.eqiad.wmnet * 11:42 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on ml-serve1015.eqiad.wmnet with reason: Downtime to get full picture of current BIOS settings beyond what Redfish shows * 11:39 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad * 11:39 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:37 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1283.eqiad.wmnet * 11:37 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1282.eqiad.wmnet * 11:37 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad * 11:37 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:33 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1141 * 11:33 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1141 * 11:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 11:32 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1141 * 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1141.eqiad.wmnet 156.48.64.10.in-addr.arpa 6.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:32 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1141.eqiad.wmnet 156.48.64.10.in-addr.arpa 6.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1141 - jayme@cumin1003" * 11:32 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1141 - jayme@cumin1003" * 11:32 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1282.eqiad.wmnet * 11:32 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1281.eqiad.wmnet * 11:32 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad * 11:32 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:29 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 11:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1097.eqiad.wmnet with reason: host reimage * 11:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2095.codfw.wmnet with OS trixie * 11:26 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1281.eqiad.wmnet * 11:26 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1280.eqiad.wmnet * 11:25 jayme@cumin1003: START - Cookbook sre.dns.netbox * 11:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1097.eqiad.wmnet with reason: host reimage * 11:22 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1141 * 11:21 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1141.eqiad.wmnet with OS trixie * 11:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1280.eqiad.wmnet * 11:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1279.eqiad.wmnet * 11:20 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin with reason: upgrade new Nokia swtiches in eqsin to SR Linux v26 * 11:17 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1141.eqiad.wmnet * 11:16 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1141.eqiad.wmnet * 11:16 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1141.eqiad.wmnet * 11:16 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be2095.codfw.wmnet with OS trixie * 11:15 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1279.eqiad.wmnet * 11:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1278.eqiad.wmnet * 11:14 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1139.eqiad.wmnet * 11:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1139.eqiad.wmnet * 11:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1070.eqiad.wmnet with OS trixie * 11:13 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1140.eqiad.wmnet * 11:13 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1140.eqiad.wmnet * 11:13 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1140.eqiad.wmnet * 11:09 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1278.eqiad.wmnet * 11:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1071.eqiad.wmnet with OS trixie * 11:05 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1097.eqiad.wmnet with OS trixie * 11:04 mvernon@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be1097.eqiad.wmnet with OS trixie * 11:02 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1140.eqiad.wmnet with OS trixie * 11:02 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1097.eqiad.wmnet with OS trixie * 11:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1096.eqiad.wmnet with OS trixie * 11:00 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1139.eqiad.wmnet * 11:00 jayme@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1139.eqiad.wmnet with OS trixie * 10:56 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1069.eqiad.wmnet with OS trixie * 10:56 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 10:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1070.eqiad.wmnet with reason: host reimage * 10:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 10:45 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1071.eqiad.wmnet with reason: host reimage * 10:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 10:41 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1096.eqiad.wmnet with OS trixie * 10:41 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1140.eqiad.wmnet with reason: host reimage * 10:39 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1071.eqiad.wmnet with reason: host reimage * 10:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1070.eqiad.wmnet with reason: host reimage * 10:38 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1139.eqiad.wmnet with reason: host reimage * 10:37 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1140.eqiad.wmnet with reason: host reimage * 10:35 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1069.eqiad.wmnet with reason: host reimage * 10:33 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1095.eqiad.wmnet with OS trixie * 10:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 10:33 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1139.eqiad.wmnet with reason: host reimage * 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1069.eqiad.wmnet with reason: host reimage * 10:23 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1140 * 10:23 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1140 * 10:23 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 10:22 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1140 * 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1140.eqiad.wmnet 155.48.64.10.in-addr.arpa 5.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:21 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1071.eqiad.wmnet with OS trixie * 10:21 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1140.eqiad.wmnet 155.48.64.10.in-addr.arpa 5.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1140 - jayme@cumin1003" * 10:21 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1140 - jayme@cumin1003" * 10:21 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1071 * 10:21 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1070.eqiad.wmnet with OS trixie * 10:21 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1070 * 10:20 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 10:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1095.eqiad.wmnet with OS trixie * 10:17 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1139 * 10:17 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1139 * 10:17 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be1095.eqiad.wmnet with OS trixie * 10:15 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1139 * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1139.eqiad.wmnet 194.32.64.10.in-addr.arpa 4.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:15 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1139.eqiad.wmnet 194.32.64.10.in-addr.arpa 4.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1139 - jayme@cumin1003" * 10:15 jayme@cumin1003: START - Cookbook sre.dns.netbox * 10:15 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1139 - jayme@cumin1003" * 10:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2095.codfw.wmnet with OS trixie * 10:13 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1069.eqiad.wmnet with OS trixie * 10:12 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1069 * 10:11 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1140 * 10:11 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1140.eqiad.wmnet with OS trixie * 10:11 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1140.eqiad.wmnet * 10:10 jayme@cumin1003: START - Cookbook sre.dns.netbox * 10:10 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1139 * 10:10 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1140.eqiad.wmnet * 10:10 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1140.eqiad.wmnet * 10:10 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1139.eqiad.wmnet with OS trixie * 10:09 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1139.eqiad.wmnet * 10:08 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1139.eqiad.wmnet * 10:08 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1139.eqiad.wmnet * 10:01 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2094.codfw.wmnet with OS trixie * 09:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 09:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 09:53 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 09:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 09:44 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:44 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2094.codfw.wmnet with reason: host reimage * 09:34 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2094.codfw.wmnet with reason: host reimage * 09:34 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1071 * 09:33 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1070 * 09:33 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1095.eqiad.wmnet with OS trixie * 09:27 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1069 * 09:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:23 brouberol@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM archiva1002.wikimedia.org * 09:20 brouberol@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM archiva1002.wikimedia.org * 09:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1277.eqiad.wmnet * 09:13 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2094.codfw.wmnet with OS trixie * 09:13 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 09:12 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 09:12 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:12 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1277.eqiad.wmnet * 09:12 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1276.eqiad.wmnet * 09:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1094.eqiad.wmnet with OS trixie * 09:06 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1276.eqiad.wmnet * 09:06 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1275.eqiad.wmnet * 09:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2093.codfw.wmnet with OS trixie * 09:01 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1275.eqiad.wmnet * 09:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1274.eqiad.wmnet * 08:56 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1274.eqiad.wmnet * 08:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1273.eqiad.wmnet * 08:50 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1273.eqiad.wmnet * 08:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1272.eqiad.wmnet * 08:49 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:49 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1094.eqiad.wmnet with reason: host reimage * 08:45 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1272.eqiad.wmnet * 08:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1094.eqiad.wmnet with reason: host reimage * 08:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2093.codfw.wmnet with reason: host reimage * 08:38 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:38 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2093.codfw.wmnet with reason: host reimage * 08:35 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 08:34 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 08:29 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:28 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:26 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1271.eqiad.wmnet * 08:23 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1094.eqiad.wmnet with OS trixie * 08:21 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 08:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1271.eqiad.wmnet * 08:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1270.eqiad.wmnet * 08:15 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1270.eqiad.wmnet * 08:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1269.eqiad.wmnet * 08:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2093.codfw.wmnet with OS trixie * 08:09 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1269.eqiad.wmnet * 08:09 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1268.eqiad.wmnet * 08:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2092.codfw.wmnet with OS trixie * 08:04 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1268.eqiad.wmnet * 08:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1267.eqiad.wmnet * 07:59 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1267.eqiad.wmnet * 07:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1093.eqiad.wmnet with OS trixie * 07:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1266.eqiad.wmnet * 07:56 jynus: running extra backups to test db1285 [[phab:T433826|T433826]] * 07:51 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1266.eqiad.wmnet * 07:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2092.codfw.wmnet with reason: host reimage * 07:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1093.eqiad.wmnet with reason: host reimage * 07:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2092.codfw.wmnet with reason: host reimage * 07:32 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1093.eqiad.wmnet with reason: host reimage * 07:29 jynus: running extra backups to test db1265 [[phab:T433825|T433825]] * 07:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2092.codfw.wmnet with OS trixie * 07:11 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1093.eqiad.wmnet with OS trixie * 06:50 slyngshede@dns1004: END - running authdns-update * 06:48 slyngshede@dns1004: START - running authdns-update * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.11 (duration: 02m 29s) * 03:38 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] (duration: 32m 57s) * 03:23 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 03:22 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 03:05 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 32s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 00:45 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] (duration: 06m 20s) * 00:41 cjming@deploy1003: cjming: Continuing with deployment * 00:41 cjming@deploy1003: cjming: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:39 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] == 2026-08-03 == * 23:58 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply * 23:57 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply * 23:29 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cp5021.eqsin.wmnet * 23:29 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cp5021.eqsin.wmnet * 23:27 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cp5021.eqsin.wmnet * 23:26 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cp5021.eqsin.wmnet * 23:18 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 23:17 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 22:56 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: sync * 22:56 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: sync * 22:36 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 22:36 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 22:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host search-loader2002.codfw.wmnet with OS trixie * 21:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on search-loader2002.codfw.wmnet with reason: host reimage * 21:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on search-loader2002.codfw.wmnet with reason: host reimage * 21:42 dancy@deploy1003: Stopping before sync operations * 21:41 dancy@deploy1003: Started scap sync-world: testing * 21:39 dancy@deploy1003: Installation of scap version "4.277.0" completed for 3 hosts * 21:37 dancy@deploy1003: Installing scap version "4.277.0" for 3 host(s) * 21:37 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] (duration: 06m 13s) * 21:33 dancy@deploy1003: dancy: Continuing with deployment * 21:32 dancy@deploy1003: dancy: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:31 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] * 21:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host search-loader2002.codfw.wmnet with OS trixie * 21:03 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] (duration: 06m 34s) * 20:59 dancy@deploy1003: dancy: Continuing with deployment * 20:58 dancy@deploy1003: dancy: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:56 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] * 20:52 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] (duration: 06m 23s) * 20:48 cjming@deploy1003: cjming: Continuing with deployment * 20:47 cjming@deploy1003: cjming: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:46 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] * 20:42 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] (duration: 07m 36s) * 20:38 arlolra@deploy1003: arlolra: Continuing with deployment * 20:36 arlolra@deploy1003: arlolra: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:34 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] * 20:16 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] (duration: 08m 26s) * 20:12 krinkle@deploy1003: krinkle: Continuing with deployment * 20:09 krinkle@deploy1003: krinkle: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] * 19:45 jasmine@cumin2002: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-main-codfw * 18:58 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] (duration: 09m 23s) * 18:53 krinkle@deploy1003: krinkle: Continuing with deployment * 18:53 jasmine@cumin2002: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-main-codfw * 18:50 krinkle@deploy1003: krinkle: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:48 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] * 18:37 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] (duration: 10m 13s) * 18:34 dzahn@cumin2002: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host codesearch2001.codfw.wmnet * 18:34 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host codesearch2001.codfw.wmnet with OS trixie * 18:33 krinkle@deploy1003: krinkle: Continuing with deployment * 18:29 krinkle@deploy1003: krinkle: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:27 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] * 18:19 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 18:18 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on codesearch2001.codfw.wmnet with reason: host reimage * 18:14 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 18:14 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:12 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on codesearch2001.codfw.wmnet with reason: host reimage * 18:11 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 18:11 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 18:10 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 18:02 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 18:02 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 18:01 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 18:01 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 17:55 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host codesearch2001.codfw.wmnet with OS trixie * 17:54 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:54 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) codesearch2001.codfw.wmnet on all recursors * 17:53 dzahn@cumin2002: START - Cookbook sre.dns.wipe-cache codesearch2001.codfw.wmnet on all recursors * 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:48 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:41 dzahn@cumin2002: START - Cookbook sre.dns.netbox * 17:41 dzahn@cumin2002: START - Cookbook sre.ganeti.makevm for new host codesearch2001.codfw.wmnet * 17:37 dzahn@cumin2002: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host codesearch1001.eqiad.wmnet * 17:37 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host codesearch1001.eqiad.wmnet with OS trixie * 17:24 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on codesearch1001.eqiad.wmnet with reason: host reimage * 17:17 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on codesearch1001.eqiad.wmnet with reason: host reimage * 17:08 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host codesearch1001.eqiad.wmnet with OS trixie * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 17:06 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:06 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) codesearch1001.eqiad.wmnet on all recursors * 17:06 dzahn@cumin2002: START - Cookbook sre.dns.wipe-cache codesearch1001.eqiad.wmnet on all recursors * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 17:05 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 17:04 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:04 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 16:58 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 16:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2091.codfw.wmnet with OS trixie * 16:54 ebernhardson@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 16:54 ebernhardson@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 16:49 ebernhardson@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 16:49 ebernhardson@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 16:46 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1092.eqiad.wmnet with OS trixie * 16:43 ebernhardson@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 16:43 ebernhardson@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 16:43 dzahn@cumin2002: START - Cookbook sre.dns.netbox * 16:43 dzahn@cumin2002: START - Cookbook sre.ganeti.makevm for new host codesearch1001.eqiad.wmnet * 16:41 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 16:41 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 16:40 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2091.codfw.wmnet with reason: host reimage * 16:37 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 16:35 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2091.codfw.wmnet with reason: host reimage * 16:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1092.eqiad.wmnet with reason: host reimage * 16:24 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 16:24 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 16:23 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1092.eqiad.wmnet with reason: host reimage * 16:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2091.codfw.wmnet with OS trixie * 16:03 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1092.eqiad.wmnet with OS trixie * 16:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2090.codfw.wmnet with OS trixie * 15:51 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 15:51 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 15:51 jiji@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 15:50 jiji@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 15:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2090.codfw.wmnet with reason: host reimage * 15:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2090.codfw.wmnet with reason: host reimage * 15:33 jhathaway@dns1004: END - running authdns-update * 15:31 jhathaway@dns1004: START - running authdns-update * 15:26 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1091.eqiad.wmnet with OS trixie * 15:25 dancy@deploy1003: Installation of scap version "4.276.1" completed for 3 hosts * 15:23 dancy@deploy1003: Installing scap version "4.276.1" for 3 host(s) * 15:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2090.codfw.wmnet with OS trixie * 15:12 marostegui@cumin1003: dbctl commit (dc=all): 'Repool db2245, db2246, db2247 and db2248 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95857 and previous config saved to /var/cache/conftool/dbconfig/20260803-151212-marostegui.json * 15:09 dancy@deploy1003: Started scap sync-world: testing * 15:09 dancy@deploy1003: Installation of scap version "4.277.0" completed for 3 hosts * 15:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1091.eqiad.wmnet with reason: host reimage * 15:07 dancy@deploy1003: Installing scap version "4.277.0" for 3 host(s) * 15:03 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1091.eqiad.wmnet with reason: host reimage * 14:49 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1091.eqiad.wmnet with OS trixie * 14:33 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2089.codfw.wmnet with OS trixie * 14:29 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 14:27 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 14:18 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 14:16 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 14:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2089.codfw.wmnet with reason: host reimage * 14:10 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 14:10 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 14:09 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2089.codfw.wmnet with reason: host reimage * 13:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2089.codfw.wmnet with OS trixie * 13:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2088.codfw.wmnet with OS trixie * 13:40 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1090.eqiad.wmnet with OS trixie * 13:22 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1090.eqiad.wmnet with reason: host reimage * 13:22 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] (duration: 14m 34s) * 13:19 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1090.eqiad.wmnet with reason: host reimage * 13:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2088.codfw.wmnet with reason: host reimage * 13:16 aude@deploy1003: aude, mhorsey: Continuing with deployment * 13:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2088.codfw.wmnet with reason: host reimage * 13:12 aude@deploy1003: aude, mhorsey: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] * 13:05 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1090.eqiad.wmnet with OS trixie * 12:58 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2088.codfw.wmnet with OS trixie * 12:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db[2245-2247].codfw.wmnet * 12:49 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2247: Rebooting db2247.codfw.wmnet * 12:49 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2247: Rebooting db2247.codfw.wmnet * 12:42 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2246: Rebooting db2246.codfw.wmnet * 12:42 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2246: Rebooting db2246.codfw.wmnet * 12:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2087.codfw.wmnet with OS trixie * 12:37 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1089.eqiad.wmnet with OS trixie * 12:34 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2245: Rebooting db2245.codfw.wmnet * 12:34 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2245: Rebooting db2245.codfw.wmnet * 12:34 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db[2245-2247].codfw.wmnet * 12:32 kamila@deploy1003: Finished scap sync-world: rebuild after base image update (duration: 30m 26s) * 12:28 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 12:22 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2087.codfw.wmnet with reason: host reimage * 12:19 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1089.eqiad.wmnet with reason: host reimage * 12:14 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2087.codfw.wmnet with reason: host reimage * 12:14 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1089.eqiad.wmnet with reason: host reimage * 12:03 kamila@deploy1003: Started scap sync-world: rebuild after base image update * 12:00 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1089.eqiad.wmnet with OS trixie * 12:00 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2087.codfw.wmnet with OS trixie * 11:35 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db[2245-2248].codfw.wmnet * 11:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db[2245-2248].codfw.wmnet * 11:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2086.codfw.wmnet with OS trixie * 11:26 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 11:26 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 11:25 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 11:25 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 11:24 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1088.eqiad.wmnet with OS trixie * 11:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db[2245-2248].codfw.wmnet with reason: Checking network * 11:21 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 11:20 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 11:19 marostegui@dns1004: END - running authdns-update * 11:17 marostegui@dns1004: START - running authdns-update * 11:10 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:10 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 11:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2086.codfw.wmnet with reason: host reimage * 11:09 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:08 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 11:08 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:07 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 11:07 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:07 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 11:06 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop: apply * 11:06 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop: apply * 11:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1088.eqiad.wmnet with reason: host reimage * 11:05 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop: apply * 11:04 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop: apply * 11:04 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop: apply * 11:04 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop: apply * 11:02 marostegui@dns1004: END - running authdns-update * 11:02 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2086.codfw.wmnet with reason: host reimage * 11:01 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1088.eqiad.wmnet with reason: host reimage * 11:00 marostegui@dns1004: START - running authdns-update * 10:53 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] (duration: 10m 57s) * 10:51 cmooney@dns3003: END - running authdns-update * 10:49 cmooney@dns3003: START - running authdns-update * 10:47 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1088.eqiad.wmnet with OS trixie * 10:47 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2086.codfw.wmnet with OS trixie * 10:47 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 10:46 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:46 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:46 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new reverse ranges for eqsin CR switch links - cmooney@cumin1003" * 10:46 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new reverse ranges for eqsin CR switch links - cmooney@cumin1003" * 10:42 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] * 10:41 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 10:36 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2245, db2246 and db2247 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95855 and previous config saved to /var/cache/conftool/dbconfig/20260803-103652-marostegui.json * 10:35 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2248 from s4 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95854 and previous config saved to /var/cache/conftool/dbconfig/20260803-103535-marostegui.json * 10:27 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 10:27 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 10:26 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 10:24 kart_: cxserver: Add referencePunctuation config ([[phab:T97231|T97231]]) * 10:24 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 10:23 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:23 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:23 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:22 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:22 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply * 10:21 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply * 10:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2085.codfw.wmnet with OS trixie * 10:20 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply * 10:20 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply * 10:18 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply * 10:18 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply * 10:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1087.eqiad.wmnet with OS trixie * 09:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2085.codfw.wmnet with reason: host reimage * 09:43 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1087.eqiad.wmnet with reason: host reimage * 09:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2085.codfw.wmnet with reason: host reimage * 09:40 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1087.eqiad.wmnet with reason: host reimage * 09:26 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1087.eqiad.wmnet with OS trixie * 09:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2085.codfw.wmnet with OS trixie * 09:13 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2084.codfw.wmnet with OS trixie * 09:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1086.eqiad.wmnet with OS trixie * 08:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2084.codfw.wmnet with reason: host reimage * 08:50 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2084.codfw.wmnet with reason: host reimage * 08:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1086.eqiad.wmnet with reason: host reimage * 08:39 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1086.eqiad.wmnet with reason: host reimage * 08:38 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:38 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:37 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 08:37 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 08:35 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2084.codfw.wmnet with OS trixie * 08:34 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 08:34 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:27 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1086.eqiad.wmnet with OS trixie * 08:09 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1218: Repool after a crash * 08:07 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2083.codfw.wmnet with OS trixie * 08:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1085.eqiad.wmnet with OS trixie * 07:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2083.codfw.wmnet with reason: host reimage * 07:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1085.eqiad.wmnet with reason: host reimage * 07:40 kart_: Updated cxsever to 2026-07-16-140518-production ([[phab:T97231|T97231]]) * 07:39 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply * 07:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2083.codfw.wmnet with reason: host reimage * 07:38 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply * 07:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1085.eqiad.wmnet with reason: host reimage * 07:37 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] (duration: 32m 40s) * 07:33 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply * 07:33 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply * 07:25 jdlrobson@deploy1003: jdlrobson: Continuing with deployment * 07:24 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2083.codfw.wmnet with OS trixie * 07:24 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1085.eqiad.wmnet with OS trixie * 07:23 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1218: Repool after a crash * 07:21 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:09 marostegui: Drop renamed tables [[phab:T425074|T425074]] * 07:04 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] * 06:55 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply * 06:54 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 46s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-02 == * 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 01m 03s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-01 == * 03:30 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:30 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:30 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:30 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 34s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-31 == * 17:41 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 17:41 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 17:40 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 17:40 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 15:33 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2195: Testing * 15:02 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:02 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 15:02 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 14:48 pt1979@cumin2002: START - Cookbook sre.dns.netbox * 14:47 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2195: Testing * 14:22 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 14:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2195: Testing * 14:21 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 14:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2195.codfw.wmnet with reason: Testing * 14:16 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 14:04 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 14:04 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 13:30 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1048.eqiad.wmnet with OS trixie * 13:22 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2195: Testing * 13:22 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 13:19 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2195: Testing * 13:18 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 13:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2195: Testing * 13:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 13:05 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 13:05 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 13:04 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 13:04 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 12:50 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lswtest-d8-eqiad * 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:53 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:42 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 6515 * 11:37 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 6515 * 11:28 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:27 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:07 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2082.codfw.wmnet with OS trixie * 10:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2082.codfw.wmnet with reason: host reimage * 10:42 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2082.codfw.wmnet with reason: host reimage * 10:28 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2082.codfw.wmnet with OS trixie * 10:02 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 09:52 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 09:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts * 09:16 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts * 08:57 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 08:46 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:42 gkyziridis@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 08:37 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 08:37 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 08:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 08:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 08:11 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:11 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:08 filippo@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudvirt1048 * 08:07 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 08:07 filippo@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudvirt1048 * 08:06 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 08:01 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 08:00 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 07:19 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1048.eqiad.wmnet with reason: host reimage * 07:13 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1048.eqiad.wmnet with reason: host reimage * 07:11 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 07:11 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 07:09 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 07:09 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 06:57 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:56 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:48 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:48 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:44 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie * 06:34 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1048.eqiad.wmnet with OS trixie * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 54s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 00:57 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] (duration: 11m 04s) * 00:53 dreamyjazz@deploy1003: dreamyjazz, jforrester: Continuing with deployment * 00:48 dreamyjazz@deploy1003: dreamyjazz, jforrester: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:46 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] == 2026-07-30 == * 21:37 dancy@deploy1003: Installation of scap version "4.276.1" completed for 3 hosts * 21:35 dancy@deploy1003: Installing scap version "4.276.1" for 3 host(s) * 21:24 dancy@deploy1003: Installation of scap version "4.276.0" completed for 3 hosts * 21:22 dancy@deploy1003: Installing scap version "4.276.0" for 3 host(s) * 21:15 maryum: Deployed security fix for [[phab:T430601|T430601]] * 20:13 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] (duration: 09m 20s) * 20:07 arlolra@deploy1003: osleger, arlolra: Continuing with deployment * 20:05 arlolra@deploy1003: osleger, arlolra: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:03 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] * 19:29 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:29 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:25 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service * 19:24 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 19:24 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:24 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:24 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 19:23 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service * 19:20 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1084.eqiad.wmnet with OS trixie * 18:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1084.eqiad.wmnet with reason: host reimage * 18:52 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1084.eqiad.wmnet with reason: host reimage * 18:41 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 18:40 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 18:39 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1084.eqiad.wmnet with OS trixie * 18:25 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 18:15 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 17:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1083.eqiad.wmnet with OS trixie * 17:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1048: Maintenance * 17:36 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new security plugin settings - bking@cumin2003 - [[phab:T350516|T350516]] * 17:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1083.eqiad.wmnet with reason: host reimage * 17:26 inflatador: bking@apt1002 `reprepro --noskipold --component thirdparty/opensearch3 update trixie-wikimedia` [[phab:T433624|T433624]] * 17:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1083.eqiad.wmnet with reason: host reimage * 17:23 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 17:20 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 17:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2081.codfw.wmnet with OS trixie * 17:11 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new security plugin settings - bking@cumin2003 - [[phab:T350516|T350516]] * 17:10 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1083.eqiad.wmnet with OS trixie * 16:55 root@cumin1003: START - Cookbook sre.mysql.pool pool es1048: Maintenance * 16:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2081.codfw.wmnet with reason: host reimage * 16:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1048 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95833 and previous config saved to /var/cache/conftool/dbconfig/20260730-165053-cwilliams.json * 16:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1048.eqiad.wmnet with reason: Maintenance * 16:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1040: Maintenance * 16:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2081.codfw.wmnet with reason: host reimage * 16:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1082.eqiad.wmnet with OS trixie * 16:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2081.codfw.wmnet with OS trixie * 16:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2097.codfw.wmnet with OS trixie * 16:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1082.eqiad.wmnet with reason: host reimage * 16:08 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1082.eqiad.wmnet with reason: host reimage * 16:08 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 16:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1047: Maintenance * 16:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2080.codfw.wmnet with OS trixie * 16:04 root@cumin1003: START - Cookbook sre.mysql.pool pool es1040: Maintenance * 16:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1040: Maintenance * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new logging settings - bking@cumin2003 - [[phab:T324335|T324335]] * 15:58 root@cumin1003: START - Cookbook sre.mysql.pool pool es1040: Maintenance * 15:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1040 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95827 and previous config saved to /var/cache/conftool/dbconfig/20260730-155324-cwilliams.json * 15:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1040.eqiad.wmnet with reason: Maintenance * 15:50 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1082.eqiad.wmnet with OS trixie * 15:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2048: Maintenance * 15:44 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 24s) * 15:43 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2080.codfw.wmnet with reason: host reimage * 15:38 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new logging settings - bking@cumin2003 - [[phab:T324335|T324335]] * 15:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2080.codfw.wmnet with reason: host reimage * 15:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 15:30 mvernon@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be2097.codfw.wmnet with OS trixie * 15:23 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1081.eqiad.wmnet with OS trixie * 15:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2098.codfw.wmnet with OS trixie * 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - mvernon@cumin2003" * 15:18 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be2097.codfw.wmnet with OS trixie * 15:18 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - mvernon@cumin2003" * 15:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 15:17 root@cumin1003: START - Cookbook sre.mysql.pool pool es1047: Maintenance * 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2080.codfw.wmnet with OS trixie * 15:13 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be2097.codfw.wmnet with OS trixie * 15:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1047 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95820 and previous config saved to /var/cache/conftool/dbconfig/20260730-151200-cwilliams.json * 15:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1047.eqiad.wmnet with reason: Maintenance * 15:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: Maintenance * 15:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 15:04 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 15:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1081.eqiad.wmnet with reason: host reimage * 15:00 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1081.eqiad.wmnet with reason: host reimage * 15:00 root@cumin1003: START - Cookbook sre.mysql.pool pool es2048: Maintenance * 15:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 14:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2079.codfw.wmnet with OS trixie * 14:56 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 14:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2048 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95816 and previous config saved to /var/cache/conftool/dbconfig/20260730-145510-cwilliams.json * 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2048.codfw.wmnet with reason: Maintenance * 14:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2040: Maintenance * 14:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 14:51 tchin@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] (duration: 06m 48s) * 14:47 tchin@deploy1003: jforrester, tchin: Continuing with deployment * 14:47 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 14:47 tchin@deploy1003: jforrester, tchin: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:45 tchin@deploy1003: Started scap sync-world: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] * 14:42 sukhe@puppetserver1001: conftool action : set/weight=1; selector: cluster=urldownloader,service=squid * 14:42 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader,service=squid * 14:42 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1081.eqiad.wmnet with OS trixie * 14:39 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 14:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2079.codfw.wmnet with reason: host reimage * 14:36 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2098.codfw.wmnet with OS trixie * 14:32 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2079.codfw.wmnet with reason: host reimage * 14:30 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] (duration: 06m 31s) * 14:27 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 14:26 mszwarc@deploy1003: mszwarc: Continuing with deployment * 14:25 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:25 root@cumin1003: START - Cookbook sre.mysql.pool pool es1038: Maintenance * 14:25 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1038: Maintenance * 14:23 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] * 14:21 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] (duration: 11m 19s) * 14:20 root@cumin1003: START - Cookbook sre.mysql.pool pool es1038: Maintenance * 14:14 stran@deploy1003: stran: Continuing with deployment * 14:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1038 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95810 and previous config saved to /var/cache/conftool/dbconfig/20260730-141439-cwilliams.json * 14:14 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1038.eqiad.wmnet with reason: Maintenance * 14:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1036: Maintenance * 14:13 stran@deploy1003: stran: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2079.codfw.wmnet with OS trixie * 14:09 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] * 14:08 root@cumin1003: START - Cookbook sre.mysql.pool pool es2040: Maintenance * 14:08 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2040: Maintenance * 14:03 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] (duration: 31m 41s) * 14:03 root@cumin1003: START - Cookbook sre.mysql.pool pool es2040: Maintenance * 14:03 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2040 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95806 and previous config saved to /var/cache/conftool/dbconfig/20260730-135643-cwilliams.json * 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2040.codfw.wmnet with reason: Maintenance * 13:56 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:56 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2038: Maintenance * 13:55 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:52 lucaswerkmeister-wmde@deploy1003: migr, lucaswerkmeister-wmde: Continuing with deployment * 13:49 lucaswerkmeister-wmde@deploy1003: migr, lucaswerkmeister-wmde: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:49 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:48 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 13:45 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 13:32 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:32 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] * 13:28 root@cumin1003: START - Cookbook sre.mysql.pool pool es1036: Maintenance * 13:28 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1036: Maintenance * 13:22 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:22 root@cumin1003: START - Cookbook sre.mysql.pool pool es1036: Maintenance * 13:20 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1036 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95800 and previous config saved to /var/cache/conftool/dbconfig/20260730-131727-cwilliams.json * 13:17 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1036.eqiad.wmnet with reason: Maintenance * 13:17 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] (duration: 10m 31s) * 13:16 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2022\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 13:13 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, stran: Continuing with deployment * 13:10 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:10 root@cumin1003: START - Cookbook sre.mysql.pool pool es2038: Maintenance * 13:10 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2038: Maintenance * 13:08 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, stran: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie * 13:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2047: Maintenance * 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048 cloud-private - filippo@cumin1003" * 13:07 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048 cloud-private - filippo@cumin1003" * 13:06 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] * 13:04 root@cumin1003: START - Cookbook sre.mysql.pool pool es2038: Maintenance * 13:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2078.codfw.wmnet with OS trixie * 13:01 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2038 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95797 and previous config saved to /var/cache/conftool/dbconfig/20260730-125919-cwilliams.json * 12:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2038.codfw.wmnet with reason: Maintenance * 12:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2078.codfw.wmnet with reason: host reimage * 12:37 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2078.codfw.wmnet with reason: host reimage * 12:37 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] (duration: 06m 51s) * 12:33 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 12:32 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 12:32 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 12:32 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:30 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] * 12:19 root@cumin1003: START - Cookbook sre.mysql.pool pool es2047: Maintenance * 12:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2078.codfw.wmnet with OS trixie * 12:18 dcausse@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 12:18 dcausse@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 12:15 dcausse@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 12:14 dcausse@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 12:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2047 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95793 and previous config saved to /var/cache/conftool/dbconfig/20260730-121404-cwilliams.json * 12:13 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2047.codfw.wmnet with reason: Maintenance * 12:13 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2036: Maintenance * 12:05 ayounsi@dns1004: END - running authdns-update * 12:02 ayounsi@dns1004: START - running authdns-update * 11:51 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2077.codfw.wmnet with OS trixie * 11:48 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:46 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1080.eqiad.wmnet with OS trixie * 11:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1226: Maintenance * 11:41 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2077.codfw.wmnet with reason: host reimage * 11:28 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2077.codfw.wmnet with reason: host reimage * 11:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1080.eqiad.wmnet with reason: host reimage * 11:27 root@cumin1003: START - Cookbook sre.mysql.pool pool es2036: Maintenance * 11:27 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2036: Maintenance * 11:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1080.eqiad.wmnet with reason: host reimage * 11:21 root@cumin1003: START - Cookbook sre.mysql.pool pool es2036: Maintenance * 11:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2036 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95786 and previous config saved to /var/cache/conftool/dbconfig/20260730-111633-cwilliams.json * 11:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2036.codfw.wmnet with reason: Maintenance * 11:08 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2077.codfw.wmnet with OS trixie * 11:07 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 11:03 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie * 11:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1226: Maintenance * 10:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1226 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95783 and previous config saved to /var/cache/conftool/dbconfig/20260730-104801-cwilliams.json * 10:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1226.eqiad.wmnet with reason: Maintenance * 10:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1214: Maintenance * 10:27 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2035: Maintenance * 10:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2076.codfw.wmnet with OS trixie * 10:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1214: Maintenance * 09:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1214 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95775 and previous config saved to /var/cache/conftool/dbconfig/20260730-095451-cwilliams.json * 09:54 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1214.eqiad.wmnet with reason: Maintenance * 09:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1209: Maintenance * 09:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2076.codfw.wmnet with reason: host reimage * 09:42 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool es2035: Maintenance * 09:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.netbox.update-extras (exit_code=0) rolling restart_daemons on A:netbox * 09:41 ayounsi@cumin1003: START - Cookbook sre.netbox.update-extras rolling restart_daemons on A:netbox * 09:40 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2035: Maintenance * 09:39 ayounsi@cumin1003: END (PASS) - Cookbook sre.netbox.update-extras (exit_code=0) rolling restart_daemons on A:netbox-canary * 09:39 ayounsi@cumin1003: START - Cookbook sre.netbox.update-extras rolling restart_daemons on A:netbox-canary * 09:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2076.codfw.wmnet with reason: host reimage * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 09:34 root@cumin1003: START - Cookbook sre.mysql.pool pool es2035: Maintenance * 09:32 lucaswerkmeister-wmde@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 09:32 lucaswerkmeister-wmde@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 09:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2035 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95771 and previous config saved to /var/cache/conftool/dbconfig/20260730-092910-cwilliams.json * 09:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2035.codfw.wmnet with reason: Maintenance * 09:19 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2076.codfw.wmnet with OS trixie * 09:18 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 09:17 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie * 09:07 root@cumin1003: START - Cookbook sre.mysql.pool pool db1209: Maintenance * 09:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 23 hosts * 09:04 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Remove cable label from interfaces descriptions - ayounsi@cumin1003 * 09:04 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:02 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Remove cable label from interfaces descriptions - ayounsi@cumin1003 * 09:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1209 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95767 and previous config saved to /var/cache/conftool/dbconfig/20260730-090133-cwilliams.json * 09:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1209.eqiad.wmnet with reason: Maintenance * 09:01 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1192: Maintenance * 08:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1252: Maintenance * 08:57 jayme@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 08:56 jayme@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 08:53 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 23 hosts * 08:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:51 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:50 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1263: Maintenance * 08:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:23 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 08:15 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie * 08:14 root@cumin1003: START - Cookbook sre.mysql.pool pool db1192: Maintenance * 08:13 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2075.codfw.wmnet with OS trixie * 08:12 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1252: Maintenance * 08:11 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1252: Maintenance * 08:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1252: Maintenance * 08:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1192 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95754 and previous config saved to /var/cache/conftool/dbconfig/20260730-080611-cwilliams.json * 08:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1192.eqiad.wmnet with reason: Maintenance * 08:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1178: Maintenance * 08:05 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS bullseye * 07:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1263: Maintenance * 07:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1263 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95751 and previous config saved to /var/cache/conftool/dbconfig/20260730-075106-cwilliams.json * 07:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[1260-1262].eqiad.wmnet with reason: Maintenance * 07:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2075.codfw.wmnet with reason: host reimage * 07:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1263.eqiad.wmnet with reason: Maintenance * 07:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2075.codfw.wmnet with reason: host reimage * 07:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance * 07:38 dcausse@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:38 dcausse@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 07:35 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db1252', diff saved to https://phabricator.wikimedia.org/P95748 and previous config saved to /var/cache/conftool/dbconfig/20260730-073510-marostegui.json * 07:26 klausman@dns2004: END - running authdns-update * 07:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 07:25 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2075.codfw.wmnet with OS trixie * 07:24 klausman@dns2004: START - running authdns-update * 07:23 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host an-test-master1003.eqiad.wmnet * 07:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db1178: Maintenance * 07:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1178 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95746 and previous config saved to /var/cache/conftool/dbconfig/20260730-071112-cwilliams.json * 07:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1178.eqiad.wmnet with reason: Maintenance * 07:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1177: Maintenance * 06:24 root@cumin1003: START - Cookbook sre.mysql.pool pool db1177: Maintenance * 06:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1177 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95741 and previous config saved to /var/cache/conftool/dbconfig/20260730-061736-cwilliams.json * 06:17 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1177.eqiad.wmnet with reason: Maintenance * 06:17 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1172: Maintenance * 05:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1218.eqiad.wmnet with reason: crashed * 05:41 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db1217 it crashed', diff saved to https://phabricator.wikimedia.org/P95737 and previous config saved to /var/cache/conftool/dbconfig/20260730-054111-marostegui.json * 05:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95736 and previous config saved to /var/cache/conftool/dbconfig/20260730-053422-cwilliams.json * 05:30 root@cumin1003: START - Cookbook sre.mysql.pool pool db1172: Maintenance * 05:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95734 and previous config saved to /var/cache/conftool/dbconfig/20260730-052414-cwilliams.json * 05:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1172 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95733 and previous config saved to /var/cache/conftool/dbconfig/20260730-052354-cwilliams.json * 05:23 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1172.eqiad.wmnet with reason: Maintenance * 05:23 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1167: Maintenance * 05:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95731 and previous config saved to /var/cache/conftool/dbconfig/20260730-051406-cwilliams.json * 05:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95729 and previous config saved to /var/cache/conftool/dbconfig/20260730-050358-cwilliams.json * 04:47 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 04:47 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 04:47 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 04:35 root@cumin1003: START - Cookbook sre.mysql.pool pool db1167: Maintenance * 04:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1167 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95726 and previous config saved to /var/cache/conftool/dbconfig/20260730-042923-cwilliams.json * 04:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 04:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1167.eqiad.wmnet with reason: Maintenance * 04:22 pt1979@cumin2002: START - Cookbook sre.dns.netbox * 04:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95725 and previous config saved to /var/cache/conftool/dbconfig/20260730-040337-cwilliams.json * 04:03 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:38 brett@cumin2002: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool eqsin [reason: Switch upgrade maintenance window complete, [[phab:T433097|T433097]]] * 01:38 brett@cumin2002: START - Cookbook sre.dns.admin DNS admin: pool eqsin [reason: Switch upgrade maintenance window complete, [[phab:T433097|T433097]]] == 2026-07-29 == * 23:57 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin,mr1-eqsin IPv6,mr1-eqsin.oob,mr1-eqsin.oob IPv6 with reason: connection issue * 22:54 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 22:53 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 22:53 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 22:53 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:25 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2022.codfw.wmnet, repooling source-only afterwards * 22:20 brett@cumin2002: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool eqsin [reason: Switch upgrade maintenance window, [[phab:T433097|T433097]]] * 22:20 brett@cumin2002: START - Cookbook sre.dns.admin DNS admin: depool eqsin [reason: Switch upgrade maintenance window, [[phab:T433097|T433097]]] * 22:01 apine@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 22:00 apine@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 21:59 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 21:58 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 21:58 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 21:58 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 21:32 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1048.eqiad.wmnet with OS trixie * 21:25 pt1979@cumin2002: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be2097.codfw.wmnet with OS bullseye * 21:16 zabe@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki=metawiki 'Mental Health Resource Center' 'Safety Resource Center/Mental Health' Zabe --reason 'per request [[:phab:T433118{{!}}T433118]]' * 21:12 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2022.codfw.wmnet, repooling source-only afterwards * 21:12 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] (duration: 12m 53s) * 21:12 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2015\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 21:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1253: Maintenance * 21:08 aaron@deploy1003: aaron: Continuing with deployment * 21:01 aaron@deploy1003: aaron: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:59 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] * 20:52 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] (duration: 21m 57s) * 20:48 aaron@deploy1003: aaron: Continuing with deployment * 20:32 aaron@deploy1003: aaron: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:30 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] * 20:24 root@cumin1003: START - Cookbook sre.mysql.pool pool db1253: Maintenance * 20:19 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] (duration: 08m 07s) * 20:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1253 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95719 and previous config saved to /var/cache/conftool/dbconfig/20260729-201810-cwilliams.json * 20:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1253.eqiad.wmnet with reason: Maintenance * 20:17 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1231: Maintenance * 20:15 aaron@deploy1003: bpirkle, aaron: Continuing with deployment * 20:13 aaron@deploy1003: bpirkle, aaron: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:12 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie * 20:11 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] * 20:11 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1048.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:09 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1048.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:09 pt1979@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 20:04 pt1979@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 19:47 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 19:43 pt1979@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye * 19:41 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 19:37 zabe: zabe@deploy1003:~$ mwscript-k8s --comment='[[phab:T433529|T433529]]' --follow -- resetAuthenticationThrottle.php --wiki=aawiki --signup --ip=89.36.114.94 * 19:36 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] (duration: 06m 49s) * 19:32 zabe@deploy1003: zabe: Continuing with deployment * 19:31 zabe@deploy1003: zabe: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:31 root@cumin1003: START - Cookbook sre.mysql.pool pool db1231: Maintenance * 19:29 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] * 19:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1231 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95714 and previous config saved to /var/cache/conftool/dbconfig/20260729-192454-cwilliams.json * 19:24 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1231.eqiad.wmnet with reason: Maintenance * 19:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1227: Maintenance * 19:22 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 19:22 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 19:21 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:21 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1048] - vriley@cumin1003" * 19:21 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1048] - vriley@cumin1003" * 19:19 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 19:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1251: Maintenance * 19:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95711 and previous config saved to /var/cache/conftool/dbconfig/20260729-191756-cwilliams.json * 19:16 vriley@cumin1003: START - Cookbook sre.dns.netbox * 19:11 dduvall: rolling back wmf.13 to group0 due to [[phab:T433457|T433457]] (cc [[phab:T430832|T430832]]) * 19:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95709 and previous config saved to /var/cache/conftool/dbconfig/20260729-190748-cwilliams.json * 19:01 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2022.codfw.wmnet with OS bookworm * 19:01 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2021.codfw.wmnet, repooling source-only afterwards * 18:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95707 and previous config saved to /var/cache/conftool/dbconfig/20260729-185740-cwilliams.json * 18:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95704 and previous config saved to /var/cache/conftool/dbconfig/20260729-184732-cwilliams.json * 18:37 root@cumin1003: START - Cookbook sre.mysql.pool pool db1227: Maintenance * 18:34 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2022.codfw.wmnet with reason: host reimage * 18:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1227 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95701 and previous config saved to /var/cache/conftool/dbconfig/20260729-183117-cwilliams.json * 18:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1227.eqiad.wmnet with reason: Maintenance * 18:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1202: Maintenance * 18:30 root@cumin1003: START - Cookbook sre.mysql.pool pool db1251: Maintenance * 18:27 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2022.codfw.wmnet with reason: host reimage * 18:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1251 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95698 and previous config saved to /var/cache/conftool/dbconfig/20260729-182428-cwilliams.json * 18:24 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lvs2014.codfw.wmnet * 18:24 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for lvs2014.codfw.wmnet * 18:24 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1251.eqiad.wmnet with reason: Maintenance * 18:23 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1235: Maintenance * 18:22 brett@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 18:19 brett@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 18:19 brett@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 18:17 brett@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 18:17 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 18:16 mutante: removing jenkins during the train - living on the edge - no, just kidding, jenkins has migrated to dedicated machines, nothing should happen * 18:15 brett@cumin2002: END (ERROR) - Cookbook sre.loadbalancer.restart-pybal (exit_code=97) rolling-restart of pybal on P<nowiki>{</nowiki>lvs2014.codfw.wmnet<nowiki>}</nowiki> and A:lvs ([[phab:T428495|T428495]]) * 18:15 mutante: CI: contint1002/contint2002: apt-get remove --purge jenkins - jenkins be gone - [[phab:T418521|T418521]] * 18:13 brett@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on P<nowiki>{</nowiki>lvs2014.codfw.wmnet<nowiki>}</nowiki> and A:lvs ([[phab:T428495|T428495]]) * 18:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2022 * 18:08 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2022 * 18:03 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T428495|T428495]] * 18:03 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2022 * 18:02 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2022.codfw.wmnet 211.48.192.10.in-addr.arpa 1.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:02 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2022.codfw.wmnet 211.48.192.10.in-addr.arpa 1.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:02 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:02 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2022 - bking@cumin2003" * 18:02 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2022 - bking@cumin2003" * 17:57 bking@cumin2003: START - Cookbook sre.dns.netbox * 17:56 brett@cumin2002: END (FAIL) - Cookbook sre.loadbalancer.restart-pybal (exit_code=1) rolling-restart of pybal on A:lvs-codfw and A:lvs ([[phab:T428495|T428495]]) * 17:55 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo - [[phab:T428495|T428495]] * 17:54 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2022 * 17:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2022.codfw.wmnet with OS bookworm * 17:50 brett@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on A:lvs-codfw and A:lvs ([[phab:T428495|T428495]]) * 17:47 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2021.codfw.wmnet, repooling source-only afterwards * 17:47 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 14s) * 17:47 swfrench-wmf: authdns-update to direct codfw, eqsin, ulsfo etcd clients back to codfw - [[phab:T428495|T428495]] * 17:47 swfrench@dns1004: END - running authdns-update * 17:47 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 17:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95692 and previous config saved to /var/cache/conftool/dbconfig/20260729-174713-cwilliams.json * 17:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance * 17:46 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1249: Maintenance * 17:45 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2015\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 17:45 swfrench@dns1004: START - running authdns-update * 17:44 root@cumin1003: START - Cookbook sre.mysql.pool pool db1202: Maintenance * 17:41 akhatun: Deployed refinery using scap, then deployed onto hdfs * 17:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1202 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95688 and previous config saved to /var/cache/conftool/dbconfig/20260729-173759-cwilliams.json * 17:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1202.eqiad.wmnet with reason: Maintenance * 17:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1194: Maintenance * 17:37 root@cumin1003: START - Cookbook sre.mysql.pool pool db1235: Maintenance * 17:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1230: Maintenance * 17:30 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1235 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95684 and previous config saved to /var/cache/conftool/dbconfig/20260729-173051-cwilliams.json * 17:30 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1235.eqiad.wmnet with reason: Maintenance * 17:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1234: Maintenance * 17:26 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (thin): Regular analytics weekly train THIN [analytics/refinery@56695674] (duration: 02m 02s) * 17:24 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (thin): Regular analytics weekly train THIN [analytics/refinery@56695674] * 17:23 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567]: Regular analytics weekly train [analytics/refinery@56695674] (duration: 06m 20s) * 17:20 dancy@deploy1003: Finished scap sync-world: Testing delay_messageblobstore_purge: true (duration: 06m 29s) * 17:17 akhatun@deploy1003: Started deploy [analytics/refinery@5669567]: Regular analytics weekly train [analytics/refinery@56695674] * 17:17 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] (duration: 00m 22s) * 17:16 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] * 17:13 dancy@deploy1003: Started scap sync-world: Testing delay_messageblobstore_purge: true * 17:05 mutante: CI: contint1002/contint2002 - restarted httpd to be extra sure all is cleaned up - https://integration.wikimedia.org/ci/ is up and running [[phab:T418521|T418521]] * 17:04 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 17:03 mutante: CI: contint1002/contint2002 - rm /etc/apache2/jenkins_proxy - removing legacy jenkins proxy config - jenkins is on new dedicated machines and uses jenkins_proxy_ext config [[phab:T418521|T418521]] * 17:02 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] (duration: 36m 25s) * 17:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1249: Maintenance * 16:59 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 16:54 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2015.codfw.wmnet, repooling source-only afterwards * 16:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1249 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95674 and previous config saved to /var/cache/conftool/dbconfig/20260729-165339-cwilliams.json * 16:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1249.eqiad.wmnet with reason: Maintenance * 16:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1248: Maintenance * 16:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1194: Maintenance * 16:47 swfrench-wmf: silenced EtcdReplicationDown 57b2b421-1cc9-4e38-9276-{{Gerrit|94f223fd231c}} - [[phab:T428495|T428495]] * 16:46 tchin@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/eventstreams-internal: apply * 16:46 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye * 16:46 tchin@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/eventstreams-internal: apply * 16:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1230: Maintenance * 16:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1194 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95669 and previous config saved to /var/cache/conftool/dbconfig/20260729-164422-cwilliams.json * 16:44 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 16:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1194.eqiad.wmnet with reason: Maintenance * 16:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1191: Maintenance * 16:43 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host an-test-master1003.eqiad.wmnet * 16:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db1234: Maintenance * 16:43 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 16:43 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Rolling back deployment * 16:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host an-test-master1004.eqiad.wmnet * 16:41 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 16:40 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 16:40 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 16:39 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 16:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1230 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95667 and previous config saved to /var/cache/conftool/dbconfig/20260729-163932-cwilliams.json * 16:39 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 16:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1230.eqiad.wmnet with reason: Maintenance * 16:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1207: Maintenance * 16:38 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 16:37 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host an-test-master1004.eqiad.wmnet * 16:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1234 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95664 and previous config saved to /var/cache/conftool/dbconfig/20260729-163719-cwilliams.json * 16:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1234.eqiad.wmnet with reason: Maintenance * 16:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1079.eqiad.wmnet with OS trixie * 16:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1232: Maintenance * 16:34 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1259: Maintenance * 16:28 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] (duration: 06m 57s) * 16:28 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:26 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] * 16:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1051 hosts * 16:21 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] * 16:20 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2006.codfw.wmnet with OS bookworm * 16:19 akhatun: Deploying Refinery at {{Gerrit|56695674}} as part of weekly train * 16:18 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1079.eqiad.wmnet with reason: host reimage * 16:16 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] (duration: 15m 36s) * 16:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2021.codfw.wmnet with OS bookworm * 16:14 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1079.eqiad.wmnet with reason: host reimage * 16:12 topranks: hot-swap line card in FPC0 on cr1-eqiad with replacement MPC10E from Juniper [[phab:T426343|T426343]] * 16:10 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Continuing with deployment * 16:07 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db1248: Maintenance * 16:01 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] * 16:00 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 16:00 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 15:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1248 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95651 and previous config saved to /var/cache/conftool/dbconfig/20260729-155956-cwilliams.json * 15:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1248.eqiad.wmnet with reason: Maintenance * 15:59 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2006.codfw.wmnet with reason: host reimage * 15:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1247: Maintenance * 15:59 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2074.codfw.wmnet with OS trixie * 15:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1191: Maintenance * 15:57 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:55 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1079.eqiad.wmnet with OS trixie * 15:55 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2006.codfw.wmnet with reason: host reimage * 15:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1207: Maintenance * 15:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1191 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95646 and previous config saved to /var/cache/conftool/dbconfig/20260729-155104-cwilliams.json * 15:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1191.eqiad.wmnet with reason: Maintenance * 15:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1181: Maintenance * 15:49 root@cumin1003: START - Cookbook sre.mysql.pool pool db1232: Maintenance * 15:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2021.codfw.wmnet with reason: host reimage * 15:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1207 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95643 and previous config saved to /var/cache/conftool/dbconfig/20260729-154735-cwilliams.json * 15:47 root@cumin1003: START - Cookbook sre.mysql.pool pool db1259: Maintenance * 15:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1207.eqiad.wmnet with reason: Maintenance * 15:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1200: Maintenance * 15:46 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] (duration: 31m 59s) * 15:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 15:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1232 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95640 and previous config saved to /var/cache/conftool/dbconfig/20260729-154330-cwilliams.json * 15:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1232.eqiad.wmnet with reason: Maintenance * 15:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1219: Maintenance * 15:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2015.codfw.wmnet, repooling source-only afterwards * 15:41 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 18s) * 15:41 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1259 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95638 and previous config saved to /var/cache/conftool/dbconfig/20260729-154107-cwilliams.json * 15:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1259.eqiad.wmnet with reason: Maintenance * 15:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1254: Maintenance * 15:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2015.codfw.wmnet with OS bookworm * 15:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2021.codfw.wmnet with reason: host reimage * 15:36 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2006.codfw.wmnet with OS bookworm * 15:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 15:35 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Continuing with deployment * 15:33 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2074.codfw.wmnet with OS trixie * 15:33 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 15:32 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:29 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:28 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be2074.codfw.wmnet with OS trixie * 15:28 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2006.codfw.wmnet * 15:26 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1078.eqiad.wmnet with OS trixie * 15:25 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 15:25 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:22 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2006.codfw.wmnet * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2021 * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2021 * 15:19 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2021 * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2021.codfw.wmnet 210.48.192.10.in-addr.arpa 0.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:19 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2021.codfw.wmnet 210.48.192.10.in-addr.arpa 0.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2021 - bking@cumin2003" * 15:19 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2021 - bking@cumin2003" * 15:14 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] * 15:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2015.codfw.wmnet with reason: host reimage * 15:11 root@cumin1003: START - Cookbook sre.mysql.pool pool db1247: Maintenance * 15:11 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on ml-serve2004.codfw.wmnet with reason: [[phab:T433478|T433478]] * 15:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2015.codfw.wmnet with reason: host reimage * 15:10 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on ml-serve2002.codfw.wmnet with reason: [[phab:T433476|T433476]] * 15:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 15:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1247 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95625 and previous config saved to /var/cache/conftool/dbconfig/20260729-150459-cwilliams.json * 15:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1247.eqiad.wmnet with reason: Maintenance * 15:04 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:04 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1244: Maintenance * 15:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1078.eqiad.wmnet with reason: host reimage * 15:03 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1005.wikimedia.org * 15:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db1181: Maintenance * 15:01 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2098.codfw.wmnet with OS bullseye * 15:00 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye * 15:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1200: Maintenance * 14:59 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 14:59 root@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1285.eqiad.wmnet with OS trixie * 14:59 Amir1: mwscript-k8s -- extensions/TimedMediaHandler/maintenance/requeueTranscodes.php --wiki=commonswiki --key '360p.mpeg4.mov' --throttle --video --missing ([[phab:T358266|T358266]]) * 14:58 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1005.wikimedia.org * 14:58 jhancock@cumin2002: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['ms-be2098'] * 14:58 jhancock@cumin2002: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['ms-be2098'] * 14:58 jhancock@cumin2002: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['ms-be2097'] * 14:58 jhancock@cumin2002: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['ms-be2097'] * 14:58 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1078.eqiad.wmnet with reason: host reimage * 14:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1181 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95621 and previous config saved to /var/cache/conftool/dbconfig/20260729-145629-cwilliams.json * 14:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1181.eqiad.wmnet with reason: Maintenance * 14:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1174: Maintenance * 14:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1219: Maintenance * 14:55 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader1006.wikimedia.org on all recursors * 14:55 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader1006.wikimedia.org on all recursors * 14:55 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader1005.wikimedia.org on all recursors * 14:55 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader1005.wikimedia.org on all recursors * 14:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1200 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95618 and previous config saved to /var/cache/conftool/dbconfig/20260729-145336-cwilliams.json * 14:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1254: Maintenance * 14:53 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1200.eqiad.wmnet with reason: Maintenance * 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1185: Maintenance * 14:52 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2021 * 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2015 * 14:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2015 * 14:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1219 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95616 and previous config saved to /var/cache/conftool/dbconfig/20260729-144946-cwilliams.json * 14:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1219.eqiad.wmnet with reason: Maintenance * 14:49 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1218: Maintenance * 14:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2021.codfw.wmnet with OS bookworm * 14:48 dancy@deploy1003: Finished deploy [zuul/deploy@22703a6]: Deploying https://gerrit.wikimedia.org/r/c/integration/zuul/+/1311501 ([[phab:T432491|T432491]]) (duration: 00m 15s) * 14:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2015.codfw.wmnet with OS bookworm * 14:48 dancy@deploy1003: Started deploy [zuul/deploy@22703a6]: Deploying https://gerrit.wikimedia.org/r/c/integration/zuul/+/1311501 ([[phab:T432491|T432491]]) * 14:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1254 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95613 and previous config saved to /var/cache/conftool/dbconfig/20260729-144729-cwilliams.json * 14:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1254.eqiad.wmnet with reason: Maintenance * 14:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1233: Maintenance * 14:46 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2013\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 14:46 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2014\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 14:46 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:45 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:44 root@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1285.eqiad.wmnet with reason: host reimage * 14:43 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:42 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:41 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:40 root@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1285.eqiad.wmnet with reason: host reimage * 14:39 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1078.eqiad.wmnet with OS trixie * 14:39 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2074.codfw.wmnet with OS trixie * 14:32 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:32 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:32 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:31 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2005.codfw.wmnet with OS bookworm * 14:30 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:30 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:29 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:29 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:27 root@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host db1285 * 14:27 root@cumin1003: START - Cookbook sre.hosts.move-vlan for host db1285 * 14:27 root@cumin1003: START - Cookbook sre.hosts.reimage for host db1285.eqiad.wmnet with OS trixie * 14:24 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 14:24 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:24 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:24 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:23 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:22 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:22 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:21 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:17 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:16 root@cumin1003: START - Cookbook sre.mysql.pool pool db1244: Maintenance * 14:15 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:15 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add asw1-604 loopback ipv4 - pt1979@cumin2002" * 14:15 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add asw1-604 loopback ipv4 - pt1979@cumin2002" * 14:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:12 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 14:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95599 and previous config saved to /var/cache/conftool/dbconfig/20260729-141014-cwilliams.json * 14:10 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 14:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1244.eqiad.wmnet with reason: Maintenance * 14:10 pt1979@cumin2002: START - Cookbook sre.dns.netbox * 14:10 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 14:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1243: Maintenance * 14:09 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2005.codfw.wmnet with reason: host reimage * 14:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db1174: Maintenance * 14:08 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad * 14:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db1185: Maintenance * 14:06 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 14:05 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2005.codfw.wmnet with reason: host reimage * 14:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1174 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95595 and previous config saved to /var/cache/conftool/dbconfig/20260729-140309-cwilliams.json * 14:03 sukhe@dns1004: END - running authdns-update * 14:03 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1174.eqiad.wmnet with reason: Maintenance * 14:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1170: Maintenance * 14:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db1218: Maintenance * 14:01 sukhe@dns1004: START - running authdns-update * 14:00 sukhe@dns1004: START - running authdns-update * 13:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db1233: Maintenance * 13:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1185 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95592 and previous config saved to /var/cache/conftool/dbconfig/20260729-135925-cwilliams.json * 13:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1185.eqiad.wmnet with reason: Maintenance * 13:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1161: Maintenance * 13:58 sukhe@puppetserver1001: conftool action : set/pooled=true; selector: dnsdisc=urldownloader * 13:58 root@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1265.eqiad.wmnet with OS trixie * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1218 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95590 and previous config saved to /var/cache/conftool/dbconfig/20260729-135621-cwilliams.json * 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1218.eqiad.wmnet with reason: Maintenance * 13:55 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1206: Maintenance * 13:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2073.codfw.wmnet with OS trixie * 13:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1233 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95587 and previous config saved to /var/cache/conftool/dbconfig/20260729-135335-cwilliams.json * 13:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1233.eqiad.wmnet with reason: Maintenance * 13:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1229: Maintenance * 13:50 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/kartotherian: apply * 13:50 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service * 13:49 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:49 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/kartotherian: apply * 13:48 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 13:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1077.eqiad.wmnet with OS trixie * 13:47 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 13:46 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2005.codfw.wmnet with OS bookworm * 13:44 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 13:44 root@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1265.eqiad.wmnet with reason: host reimage * 13:40 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] (duration: 09m 22s) * 13:39 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:38 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:36 root@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1265.eqiad.wmnet with reason: host reimage * 13:35 stran@deploy1003: stran: Continuing with deployment * 13:33 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 13:32 stran@deploy1003: stran: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified t * 13:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2073.codfw.wmnet with reason: host reimage * 13:30 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ml-build1001.eqiad.wmnet * 13:30 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] * 13:29 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad * 13:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1077.eqiad.wmnet with reason: host reimage * 13:27 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 13:27 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:27 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad * 13:26 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2073.codfw.wmnet with reason: host reimage * 13:26 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] (duration: 07m 56s) * 13:25 klausman@cumin1003: START - Cookbook sre.hosts.reboot-single for host ml-build1001.eqiad.wmnet * 13:24 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 13:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>ml-serve101[2-5].eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 13:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1015.eqiad.wmnet * 13:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1015.eqiad.wmnet * 13:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1077.eqiad.wmnet with reason: host reimage * 13:23 root@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host db1265 * 13:23 root@cumin1003: START - Cookbook sre.hosts.move-vlan for host db1265 * 13:23 root@cumin1003: START - Cookbook sre.hosts.reimage for host db1265.eqiad.wmnet with OS trixie * 13:23 root@cumin1003: START - Cookbook sre.mysql.pool pool db1243: Maintenance * 13:22 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 13:22 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2005.codfw.wmnet * 13:22 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 13:22 samtar@deploy1003: dreamrimmer, samtar: Continuing with deployment * 13:20 samtar@deploy1003: dreamrimmer, samtar: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts an-test-master[1001-1002].eqiad.wmnet * 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-master[1001-1002].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 13:18 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1015.eqiad.wmnet * 13:18 sukhe@cumin1003: END (ERROR) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=97) for role: url_downloader@eqiad * 13:18 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 13:18 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] * 13:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95574 and previous config saved to /var/cache/conftool/dbconfig/20260729-131638-cwilliams.json * 13:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1243.eqiad.wmnet with reason: Maintenance * 13:16 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2005.codfw.wmnet * 13:16 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1242: Maintenance * 13:14 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] (duration: 07m 00s) * 13:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db1170: Maintenance * 13:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1015.eqiad.wmnet * 13:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1014.eqiad.wmnet * 13:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1014.eqiad.wmnet * 13:12 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1223: Maintenance * 13:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db1161: Maintenance * 13:10 samtar@deploy1003: anzx, samtar: Continuing with deployment * 13:09 samtar@deploy1003: anzx, samtar: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db1206: Maintenance * 13:08 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 13:07 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] * 13:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1170 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95566 and previous config saved to /var/cache/conftool/dbconfig/20260729-130730-cwilliams.json * 13:07 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1170.eqiad.wmnet with reason: Maintenance * 13:07 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:07 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt IPs new switches - cmooney@cumin1003" * 13:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1158: Maintenance * 13:06 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1014.eqiad.wmnet * 13:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1161 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95564 and previous config saved to /var/cache/conftool/dbconfig/20260729-130616-cwilliams.json * 13:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 13:06 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1077.eqiad.wmnet with OS trixie * 13:05 root@cumin1003: START - Cookbook sre.mysql.pool pool db1229: Maintenance * 13:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1161.eqiad.wmnet with reason: Maintenance * 13:05 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt IPs new switches - cmooney@cumin1003" * 13:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2073.codfw.wmnet with OS trixie * 13:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Maintenance * 13:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1206 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95562 and previous config saved to /var/cache/conftool/dbconfig/20260729-130258-cwilliams.json * 13:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1206.eqiad.wmnet with reason: Maintenance * 13:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1196: Maintenance * 13:01 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 13:01 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:00 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 13:00 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 12:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1229 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95559 and previous config saved to /var/cache/conftool/dbconfig/20260729-125950-cwilliams.json * 12:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1229.eqiad.wmnet with reason: Maintenance * 12:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1222: Maintenance * 12:57 sukhe: sudo cumin 'A:lvs and (A:eqiad or A:codfw)' 'disable-puppet "adding new service urldownloader"': [[phab:T429175|T429175]] * 12:56 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1014.eqiad.wmnet * 12:56 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1013.eqiad.wmnet * 12:56 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1013.eqiad.wmnet * 12:50 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1013.eqiad.wmnet * 12:50 sukhe: sudo cumin 'O:url_downloader' 'run-puppet-agent --enable "merging CR 1313948"': [[phab:T429175|T429175]] * 12:48 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-master[1001-1002].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 12:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1013.eqiad.wmnet * 12:45 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1012.eqiad.wmnet * 12:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1012.eqiad.wmnet * 12:45 sukhe: sudo cumin 'O:url_downloader' 'disable-puppet "merging CR 1313948"': [[phab:T429175|T429175]] * 12:44 btullis@cumin1003: START - Cookbook sre.dns.netbox * 12:40 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test2001.codfw.wmnet * 12:40 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test2001.codfw.wmnet * 12:38 ayounsi@dns1004: END - running authdns-update * 12:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1012.eqiad.wmnet * 12:35 ayounsi@dns1004: START - running authdns-update * 12:34 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts an-test-master[1001-1002].eqiad.wmnet * 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts an-test-coord1001.eqiad.wmnet * 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-coord1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 12:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1012.eqiad.wmnet * 12:32 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>ml-serve101[2-5].eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 12:29 root@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Maintenance * 12:25 root@cumin1003: START - Cookbook sre.mysql.pool pool db1223: Maintenance * 12:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95544 and previous config saved to /var/cache/conftool/dbconfig/20260729-122254-cwilliams.json * 12:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1242.eqiad.wmnet with reason: Maintenance * 12:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1241: Maintenance * 12:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1051 hosts * 12:20 root@cumin1003: START - Cookbook sre.mysql.pool pool db1158: Maintenance * 12:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1223 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95540 and previous config saved to /var/cache/conftool/dbconfig/20260729-121937-cwilliams.json * 12:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1223.eqiad.wmnet with reason: Maintenance * 12:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1212: Maintenance * 12:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db1159: Maintenance * 12:17 elukey@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: sync * 12:15 elukey@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: sync * 12:15 root@cumin1003: START - Cookbook sre.mysql.pool pool db1196: Maintenance * 12:14 Daimona: Creating new DB tables for the CampaignEvents extension in x1.testwiki, x1.test2wiki, x1.officewiki, and x1.wikishared # [[phab:T429339|T429339]] * 12:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db1222: Maintenance * 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95535 and previous config saved to /var/cache/conftool/dbconfig/20260729-121211-cwilliams.json * 12:12 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 12:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1158.eqiad.wmnet with reason: Maintenance * 12:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1159 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95534 and previous config saved to /var/cache/conftool/dbconfig/20260729-121146-cwilliams.json * 12:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1159.eqiad.wmnet with reason: Maintenance * 12:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1196 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95533 and previous config saved to /var/cache/conftool/dbconfig/20260729-120847-cwilliams.json * 12:08 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 12:08 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1196.eqiad.wmnet with reason: Maintenance * 12:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1195: Maintenance * 12:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1222 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95530 and previous config saved to /var/cache/conftool/dbconfig/20260729-120424-cwilliams.json * 12:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1222.eqiad.wmnet with reason: Maintenance * 12:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1098 hosts * 12:00 marostegui: Rename tables [[phab:T425074|T425074]] * 12:00 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-coord1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 11:58 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1197: Maintenance * 11:55 btullis@cumin1003: START - Cookbook sre.dns.netbox * 11:52 elukey@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: sync * 11:51 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:51 elukey@deploy1003: helmfile [codfw] START helmfile.d/services/proton: sync * 11:51 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:50 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts an-test-coord1001.eqiad.wmnet * 11:50 elukey@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: sync * 11:49 elukey@deploy1003: helmfile [staging] START helmfile.d/services/proton: sync * 11:35 root@cumin1003: START - Cookbook sre.mysql.pool pool db1241: Maintenance * 11:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db1212: Maintenance * 11:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1241 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95520 and previous config saved to /var/cache/conftool/dbconfig/20260729-112918-cwilliams.json * 11:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1241.eqiad.wmnet with reason: Maintenance * 11:29 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1238: Maintenance * 11:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1212 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95517 and previous config saved to /var/cache/conftool/dbconfig/20260729-112727-cwilliams.json * 11:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 11:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1212.eqiad.wmnet with reason: Maintenance * 11:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1198: Maintenance * 11:23 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:22 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:21 root@cumin1003: START - Cookbook sre.mysql.pool pool db1195: Maintenance * 11:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1195 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95514 and previous config saved to /var/cache/conftool/dbconfig/20260729-111450-cwilliams.json * 11:14 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1195.eqiad.wmnet with reason: Maintenance * 11:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1186: Maintenance * 11:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 11:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 11:05 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:54 marostegui: Dropping renamed tables [[phab:T425066|T425066]] * 10:41 root@cumin1003: START - Cookbook sre.mysql.pool pool db1238: Maintenance * 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1198: Maintenance * 10:39 Amir1: ran https://phabricator.wikimedia.org/T432509#12149723 in production ([[phab:T432509|T432509]]) * 10:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db1197: Maintenance * 10:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1238 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95501 and previous config saved to /var/cache/conftool/dbconfig/20260729-103532-cwilliams.json * 10:35 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1238.eqiad.wmnet with reason: Maintenance * 10:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1221: Maintenance * 10:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1198 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95499 and previous config saved to /var/cache/conftool/dbconfig/20260729-103330-cwilliams.json * 10:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1198.eqiad.wmnet with reason: Maintenance * 10:33 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1175: Maintenance * 10:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1197 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95496 and previous config saved to /var/cache/conftool/dbconfig/20260729-103217-cwilliams.json * 10:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1197.eqiad.wmnet with reason: Maintenance * 10:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1188: Maintenance * 10:27 root@cumin1003: START - Cookbook sre.mysql.pool pool db1186: Maintenance * 10:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1186 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95493 and previous config saved to /var/cache/conftool/dbconfig/20260729-102111-cwilliams.json * 10:21 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1186.eqiad.wmnet with reason: Maintenance * 10:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 10:14 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 09:53 XioNoX: reboot cr2-magru - [[phab:T431750|T431750]] * 09:52 XioNoX: drain cr2-magru - [[phab:T431750|T431750]] * 09:48 root@cumin1003: START - Cookbook sre.mysql.pool pool db1221: Maintenance * 09:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zookeeper-test1002.eqiad.wmnet * 09:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1188: Maintenance * 09:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1175: Maintenance * 09:44 btullis@dns1004: END - running authdns-update * 09:42 btullis@dns1004: START - running authdns-update * 09:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1221 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95483 and previous config saved to /var/cache/conftool/dbconfig/20260729-094200-cwilliams.json * 09:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 7 hosts with reason: Maintenance * 09:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1221.eqiad.wmnet with reason: Maintenance * 09:41 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host zookeeper-test1002.eqiad.wmnet * 09:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1199: Maintenance * 09:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1188 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95481 and previous config saved to /var/cache/conftool/dbconfig/20260729-093917-cwilliams.json * 09:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1188.eqiad.wmnet with reason: Maintenance * 09:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1182: Maintenance * 09:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1175 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95479 and previous config saved to /var/cache/conftool/dbconfig/20260729-093842-cwilliams.json * 09:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1175.eqiad.wmnet with reason: Maintenance * 09:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1166: Maintenance * 09:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1169: Maintenance * 09:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1033.eqiad.wmnet,service=s8 * 09:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1033.eqiad.wmnet,service=s5 * 09:33 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1033.eqiad.wmnet,service=s8 * 09:33 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1033.eqiad.wmnet,service=s5 * 09:21 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:21 XioNoX: reboot cr1-magru - [[phab:T431750|T431750]] * 09:17 XioNoX: drain cr1-magru - [[phab:T431750|T431750]] * 09:15 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm1001.wikimedia.org * 09:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply * 09:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply * 09:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 09:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 09:11 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr2-magru,cr2-magru IPv6,cr2-magru.mgmt with reason: router upgrade * 09:11 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm1001.wikimedia.org * 09:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp1005.wikimedia.org * 09:07 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp1005.wikimedia.org * 09:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp2005.wikimedia.org * 09:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp2005.wikimedia.org * 09:00 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr1-magru,cr1-magru IPv6,cr1-magru.mgmt with reason: router upgrade * 09:00 marostegui: Dropping renamed tables [[phab:T426341|T426341]] * 08:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1199: Maintenance * 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1182: Maintenance * 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1166: Maintenance * 08:47 root@cumin1003: START - Cookbook sre.mysql.pool pool db1169: Maintenance * 08:46 ayounsi@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 1:00:00 on cr1-magru,cr1-magru IPv6,cr1-magru.mgmt with reason: router upgrade * 08:45 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2072.codfw.wmnet with OS trixie * 08:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1199 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95464 and previous config saved to /var/cache/conftool/dbconfig/20260729-084534-cwilliams.json * 08:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1199.eqiad.wmnet with reason: Maintenance * 08:45 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 08:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1190: Maintenance * 08:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1182 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95462 and previous config saved to /var/cache/conftool/dbconfig/20260729-084436-cwilliams.json * 08:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1182.eqiad.wmnet with reason: Maintenance * 08:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1156: Maintenance * 08:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1166 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95460 and previous config saved to /var/cache/conftool/dbconfig/20260729-084400-cwilliams.json * 08:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1166.eqiad.wmnet with reason: Maintenance * 08:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1157: Maintenance * 08:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95458 and previous config saved to /var/cache/conftool/dbconfig/20260729-084147-cwilliams.json * 08:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1169.eqiad.wmnet with reason: Maintenance * 08:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1163: Maintenance * 08:30 btullis@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 11 hosts with reason: Replacing the namenodes * 08:23 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2072.codfw.wmnet with reason: host reimage * 08:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1098 hosts * 08:19 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2072.codfw.wmnet with reason: host reimage * 07:58 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2072.codfw.wmnet with OS trixie * 07:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1190: Maintenance * 07:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1156: Maintenance * 07:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1157: Maintenance * 07:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1163: Maintenance * 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1190 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95444 and previous config saved to /var/cache/conftool/dbconfig/20260729-074930-cwilliams.json * 07:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1190.eqiad.wmnet with reason: Maintenance * 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1157 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95443 and previous config saved to /var/cache/conftool/dbconfig/20260729-074914-cwilliams.json * 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1156 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95442 and previous config saved to /var/cache/conftool/dbconfig/20260729-074906-cwilliams.json * 07:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1157.eqiad.wmnet with reason: Maintenance * 07:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 07:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1156.eqiad.wmnet with reason: Maintenance * 07:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1163 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95441 and previous config saved to /var/cache/conftool/dbconfig/20260729-074652-cwilliams.json * 07:46 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1163.eqiad.wmnet with reason: Maintenance * 07:46 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2034.codfw.wmnet * 07:42 ayounsi@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2034.codfw.wmnet * 07:42 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2232.codfw.wmnet with OS trixie * 07:34 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:34 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:33 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:31 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:19 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2232.codfw.wmnet with reason: host reimage * 07:15 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2232.codfw.wmnet with reason: host reimage * 06:58 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db2232.codfw.wmnet with OS trixie * 06:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[2160,2232].codfw.wmnet with reason: Reimage * 06:26 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1164.eqiad.wmnet with OS trixie * 06:05 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1164.eqiad.wmnet with reason: host reimage * 06:01 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1164.eqiad.wmnet with reason: host reimage * 05:47 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1164.eqiad.wmnet with OS trixie * 05:46 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1164.eqiad.wmnet with reason: Reimage == 2026-07-28 == * 22:50 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1138.eqiad.wmnet * 22:50 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1138.eqiad.wmnet * 22:49 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1138.eqiad.wmnet * 22:11 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2014.codfw.wmnet, repooling source-only afterwards * 22:08 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2013.codfw.wmnet, repooling source-only afterwards * 22:03 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 20:58 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] (duration: 08m 19s) * 20:55 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2014.codfw.wmnet, repooling source-only afterwards * 20:55 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2013.codfw.wmnet, repooling source-only afterwards * 20:54 arlolra@deploy1003: arlolra: Continuing with deployment * 20:54 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 14s) * 20:54 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 20:53 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 30s) * 20:53 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 20:52 arlolra@deploy1003: arlolra: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:51 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:50 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] * 20:49 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:43 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 20:34 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] (duration: 06m 54s) * 20:34 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:34 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:31 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:31 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:30 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:30 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:30 arlolra@deploy1003: arlolra: Continuing with deployment * 20:29 arlolra@deploy1003: arlolra: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:27 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] * 20:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2014.codfw.wmnet with OS bookworm * 20:21 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 20:21 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:20 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 20:19 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:19 swfrench-wmf: switched etcd-mirror replication from conf2005 to conf2004 - [[phab:T428495|T428495]] * 20:17 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:17 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:15 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] (duration: 08m 26s) * 20:12 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:11 arlolra@deploy1003: anzx, arlolra: Continuing with deployment * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2013.codfw.wmnet with OS bookworm * 20:09 arlolra@deploy1003: anzx, arlolra: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] * 20:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2216: Maintenance * 19:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2014.codfw.wmnet with reason: host reimage * 19:57 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:54 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:54 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:53 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:52 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2014.codfw.wmnet with reason: host reimage * 19:49 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2013.codfw.wmnet with reason: host reimage * 19:42 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 19:41 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:41 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2013.codfw.wmnet with reason: host reimage * 19:39 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:39 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:39 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-eqiad: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 19:38 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2014 * 19:33 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2014 * 19:29 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2014.codfw.wmnet with OS bookworm * 19:28 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:27 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1006 * 19:26 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2012\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 19:26 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1006 * 19:26 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:26 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 19:25 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2013 * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2013 * 19:21 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2013 * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2013.codfw.wmnet 84.0.192.10.in-addr.arpa 4.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:21 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2013.codfw.wmnet 84.0.192.10.in-addr.arpa 4.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2013 - bking@cumin2003" * 19:21 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2013 - bking@cumin2003" * 19:21 vriley@cumin1003: START - Cookbook sre.dns.netbox * 19:20 root@cumin1003: START - Cookbook sre.mysql.pool pool db2216: Maintenance * 19:13 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2216 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95435 and previous config saved to /var/cache/conftool/dbconfig/20260728-191343-cwilliams.json * 19:13 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2216.codfw.wmnet with reason: Maintenance * 19:13 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2203: Maintenance * 19:06 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1005.eqiad.wmnet with OS trixie * 19:06 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 19:06 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 18:46 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 18:45 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:45 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:43 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:40 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:36 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-eqiad: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 18:35 dancy@deploy1003: Installation of scap version "4.275.0" completed for 3 hosts * 18:33 dancy@deploy1003: Installing scap version "4.275.0" for 3 host(s) * 18:32 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:32 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2097.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:30 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns3003.wikimedia.org [reason: pool for all services after reimaging] * 18:29 sukhe@dns1004: END - running authdns-update * 18:27 sukhe@dns1004: START - running authdns-update * 18:27 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns3003.wikimedia.org,service=authdns-update [reason: pool authdns-update after reimaging] * 18:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db2203: Maintenance * 18:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2203 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95430 and previous config saved to /var/cache/conftool/dbconfig/20260728-181958-cwilliams.json * 18:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2203.codfw.wmnet with reason: Maintenance * 18:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2188: Maintenance * 18:18 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2097.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:17 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-be2098 * 18:17 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host ms-be2098 * 18:17 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-be2097 * 18:16 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host ms-be2097 * 18:15 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:15 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding ms-be2097-8 to codfw - jhancock@cumin2002" * 18:15 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding ms-be2097-8 to codfw - jhancock@cumin2002" * 18:10 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 18:08 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage * 18:05 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns3003.wikimedia.org with OS trixie * 18:03 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage * 17:56 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-codfw: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 17:45 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie * 17:45 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1005.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:41 sukhe@dns1004: END - running authdns-update * 17:39 sukhe@dns1004: START - running authdns-update * 17:36 sukhe@puppetserver1001: conftool action : set/weight=1; selector: cluster=urldownloader,service=squid * 17:36 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1005.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:35 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader,service=squid * 17:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 17:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1005 * 17:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 17:34 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1005 * 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1005] - vriley@cumin1003" * 17:34 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1005] - vriley@cumin1003" * 17:32 root@cumin1003: START - Cookbook sre.mysql.pool pool db2188: Maintenance * 17:29 vriley@cumin1003: START - Cookbook sre.dns.netbox * 17:29 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 17:26 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2188 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95425 and previous config saved to /var/cache/conftool/dbconfig/20260728-172609-cwilliams.json * 17:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2188.codfw.wmnet with reason: Maintenance * 17:25 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2176: Maintenance * 17:19 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1138.eqiad.wmnet with OS trixie * 17:18 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1005 * 17:18 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1005 * 17:18 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:15 vriley@cumin1003: START - Cookbook sre.dns.netbox * 17:13 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns3003.wikimedia.org with reason: host reimage * 17:07 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns3003.wikimedia.org with reason: host reimage * 17:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-eqiad * 17:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1015.eqiad.wmnet * 17:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1015.eqiad.wmnet * 16:59 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1138.eqiad.wmnet with reason: host reimage * 16:55 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-codfw: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 16:54 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1138.eqiad.wmnet with reason: host reimage * 16:53 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1015.eqiad.wmnet * 16:43 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns3003.wikimedia.org with OS trixie * 16:43 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1015.eqiad.wmnet * 16:43 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1014.eqiad.wmnet * 16:43 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1014.eqiad.wmnet * 16:43 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=dns3003.wikimedia.org [reason: depooling for reimage to trixie] * 16:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 290 hosts * 16:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db2176: Maintenance * 16:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1138 * 16:38 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1138 * 16:37 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1138 * 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1138.eqiad.wmnet 193.32.64.10.in-addr.arpa 3.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:37 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1138.eqiad.wmnet 193.32.64.10.in-addr.arpa 3.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1138 - jiji@cumin1003" * 16:37 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1138 - jiji@cumin1003" * 16:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1014.eqiad.wmnet * 16:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2176 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95420 and previous config saved to /var/cache/conftool/dbconfig/20260728-163235-cwilliams.json * 16:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2176.codfw.wmnet with reason: Maintenance * 16:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1014.eqiad.wmnet * 16:32 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1013.eqiad.wmnet * 16:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1013.eqiad.wmnet * 16:32 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2174: Maintenance * 16:28 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2012.codfw.wmnet, repooling source-only afterwards * 16:25 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1013.eqiad.wmnet * 16:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1013.eqiad.wmnet * 16:20 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1012.eqiad.wmnet * 16:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1012.eqiad.wmnet * 16:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1012.eqiad.wmnet * 16:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1012.eqiad.wmnet * 16:03 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1011.eqiad.wmnet * 16:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1011.eqiad.wmnet * 16:00 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2004.codfw.wmnet with OS bookworm * 15:59 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1011.eqiad.wmnet * 15:56 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:55 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 15:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:54 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1011.eqiad.wmnet * 15:54 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1010.eqiad.wmnet * 15:54 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1010.eqiad.wmnet * 15:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:50 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 15:49 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1010.eqiad.wmnet * 15:48 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 15:48 jiji@cumin1003: START - Cookbook sre.dns.netbox * 15:46 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2248: Maintenance * 15:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db2174: Maintenance * 15:44 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1010.eqiad.wmnet * 15:44 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1009.eqiad.wmnet * 15:44 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1009.eqiad.wmnet * 15:42 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1138 * 15:41 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1138.eqiad.wmnet with OS trixie * 15:39 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1009.eqiad.wmnet * 15:39 robh@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on arclamp2001.codfw.wmnet with reason: ram upgrade * 15:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2174 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95413 and previous config saved to /var/cache/conftool/dbconfig/20260728-153844-cwilliams.json * 15:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2174.codfw.wmnet with reason: Maintenance * 15:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2173: Maintenance * 15:37 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 15:35 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 15:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1009.eqiad.wmnet * 15:34 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1008.eqiad.wmnet * 15:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1008.eqiad.wmnet * 15:31 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 15:31 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 15:29 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1008.eqiad.wmnet * 15:27 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 290 hosts * 15:25 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2013 * 15:25 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2195: Maintenance * 15:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1008.eqiad.wmnet * 15:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1007.eqiad.wmnet * 15:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1007.eqiad.wmnet * 15:22 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2004.codfw.wmnet with reason: host reimage * 15:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2013.codfw.wmnet with OS bookworm * 15:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 15:19 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2004.codfw.wmnet with reason: host reimage * 15:19 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 15:17 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1007.eqiad.wmnet * 15:12 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1007.eqiad.wmnet * 15:12 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1006.eqiad.wmnet * 15:12 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1006.eqiad.wmnet * 15:11 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1138.eqiad.wmnet * 15:11 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 15:11 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1138.eqiad.wmnet * 15:11 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1138.eqiad.wmnet * 15:10 brennen@deploy1003: Finished deploy [phabricator/deployment@f8b349f]: deploy phab1004 for [[phab:T433382|T433382]] (duration: 00m 43s) * 15:10 brennen@deploy1003: Started deploy [phabricator/deployment@f8b349f]: deploy phab1004 for [[phab:T433382|T433382]] * 15:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts * 15:09 brennen@deploy1003: Finished deploy [phabricator/deployment@f8b349f]: deploy phab2003 for [[phab:T433382|T433382]] (duration: 00m 55s) * 15:08 brennen@deploy1003: Started deploy [phabricator/deployment@f8b349f]: deploy phab2003 for [[phab:T433382|T433382]] * 15:07 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2012.codfw.wmnet, repooling source-only afterwards * 15:07 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts * 15:06 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1137.eqiad.wmnet * 15:06 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1137.eqiad.wmnet * 15:06 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1137.eqiad.wmnet * 15:05 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1006.eqiad.wmnet * 15:05 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1022\.eqiad\.wmnet,dc=eqiad,cluster=wdqs\-main,service=wdqs\-main * 15:01 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab2003.codfw.wmnet with reason: deployment * 15:01 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1005.eqiad.wmnet with reason: deployment * 15:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1006.eqiad.wmnet * 15:00 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1006.eqiad.wmnet with reason: deployment * 15:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1005.eqiad.wmnet * 15:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1005.eqiad.wmnet * 14:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db2248: Maintenance * 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1004.eqiad.wmnet with reason: deployment * 14:59 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2004.codfw.wmnet with OS bookworm * 14:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1005.eqiad.wmnet * 14:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2248 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95403 and previous config saved to /var/cache/conftool/dbconfig/20260728-145532-cwilliams.json * 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2245-2247].codfw.wmnet with reason: Maintenance * 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2248.codfw.wmnet with reason: Maintenance * 14:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2240: Maintenance * 14:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2173: Maintenance * 14:51 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts * 14:50 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1005.eqiad.wmnet * 14:50 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1004.eqiad.wmnet * 14:50 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1004.eqiad.wmnet * 14:49 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts * 14:45 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2004.codfw.wmnet * 14:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2173 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95399 and previous config saved to /var/cache/conftool/dbconfig/20260728-144453-cwilliams.json * 14:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2173.codfw.wmnet with reason: Maintenance * 14:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2170: Maintenance * 14:44 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1004.eqiad.wmnet * 14:39 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2004.codfw.wmnet * 14:38 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1004.eqiad.wmnet * 14:38 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet * 14:38 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet * 14:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db2195: Maintenance * 14:36 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1022.eqiad.wmnet, repooling source-only afterwards * 14:36 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 14:33 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet * 14:33 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2222: Maintenance * 14:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2195 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95394 and previous config saved to /var/cache/conftool/dbconfig/20260728-143218-cwilliams.json * 14:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2195.codfw.wmnet with reason: Maintenance * 14:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2181: Maintenance * 14:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 14:30 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 14:25 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 14:25 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 14:25 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 14:23 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet * 14:23 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1002.eqiad.wmnet * 14:23 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1002.eqiad.wmnet * 14:23 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 19s) * 14:23 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:18 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1002.eqiad.wmnet * 14:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2012.codfw.wmnet with OS bookworm * 14:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1002.eqiad.wmnet * 14:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1001.eqiad.wmnet * 14:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1001.eqiad.wmnet * 14:11 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 14:11 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 14:08 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1001.eqiad.wmnet * 14:07 elukey@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'. * 14:07 elukey@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'. * 14:06 elukey@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'. * 14:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db2240: Maintenance * 14:06 XioNoX: un-drain cr2-esams - [[phab:T431751|T431751]] * 14:05 elukey@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'. * 14:02 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1001.eqiad.wmnet * 14:02 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-eqiad * 14:01 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 14:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2240 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95384 and previous config saved to /var/cache/conftool/dbconfig/20260728-140011-cwilliams.json * 14:00 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2240.codfw.wmnet with reason: Maintenance * 13:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2237: Maintenance * 13:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db2170: Maintenance * 13:56 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 13:55 XioNoX: reboot cr2-esams - [[phab:T431751|T431751]] * 13:52 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 13:51 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr2-esams,cr2-esams IPv6,cr2-esams.mgmt with reason: router upgrade * 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 13:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2170 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95381 and previous config saved to /var/cache/conftool/dbconfig/20260728-135043-cwilliams.json * 13:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2170.codfw.wmnet with reason: Maintenance * 13:50 XioNoX: drain cr2-esams - [[phab:T431751|T431751]] * 13:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2153: Maintenance * 13:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2012.codfw.wmnet with reason: host reimage * 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2222: Maintenance * 13:45 sukhe: restart pybal on A:lvs-codfw * 13:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2012.codfw.wmnet with reason: host reimage * 13:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db2181: Maintenance * 13:44 btullis@dns1004: END - running authdns-update * 13:42 sukhe: restart pybal on lvs2014 * 13:42 btullis@dns1004: START - running authdns-update * 13:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2222 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95376 and previous config saved to /var/cache/conftool/dbconfig/20260728-133948-cwilliams.json * 13:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2222.codfw.wmnet with reason: Maintenance * 13:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2221: Maintenance * 13:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2181 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95374 and previous config saved to /var/cache/conftool/dbconfig/20260728-133857-cwilliams.json * 13:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2181.codfw.wmnet with reason: Maintenance * 13:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2167: Maintenance * 13:30 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T428495|T428495]] * 13:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 13:29 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 13:29 ayounsi@cumin1003: END (FAIL) - Cookbook sre.dns.admin (exit_code=99) DNS admin: depool esams [reason: router upgrade, [[phab:T431749|T431749]]] * 13:28 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: router upgrade, [[phab:T431749|T431749]]] * 13:27 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1022.eqiad.wmnet, repooling source-only afterwards * 13:27 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo - [[phab:T428495|T428495]] * 13:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2012 * 13:27 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2012 * 13:21 lucaswerkmeister-wmde@deploy1003: mwscript-k8s job started: cleanupTitles bolwiki # [[phab:T429951|T429951]] * 13:21 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] (duration: 07m 19s) * 13:20 swfrench-wmf: authdns-update to direct codfw, eqsin, ulsfo etcd clients to eqiad - [[phab:T428495|T428495]] * 13:18 swfrench@dns1004: END - running authdns-update * 13:17 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, anzx: Continuing with deployment * 13:16 swfrench@dns1004: START - running authdns-update * 13:16 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2012 * 13:16 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2012.codfw.wmnet 57.48.192.10.in-addr.arpa 7.5.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:16 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2012.codfw.wmnet 57.48.192.10.in-addr.arpa 7.5.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:16 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:16 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, anzx: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:14 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 13:14 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] * 13:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db2237: Maintenance * 13:13 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2237: Maintenance * 13:13 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:12 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:12 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback IPV6 for asw1-604 - pt1979@cumin2003" * 13:12 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:12 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback IPV6 for asw1-604 - pt1979@cumin2003" * 13:11 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 20s) * 13:11 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 13:10 esanders@deploy1003: Finished scap sync-world: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] (duration: 08m 11s) * 13:08 pt1979@cumin2003: START - Cookbook sre.dns.netbox * 13:07 root@cumin1003: START - Cookbook sre.mysql.pool pool db2237: Maintenance * 13:06 esanders@deploy1003: esanders: Continuing with deployment * 13:04 esanders@deploy1003: esanders: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2153: Maintenance * 13:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2153: Maintenance * 13:02 esanders@deploy1003: Started scap sync-world: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] * 13:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2237 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95362 and previous config saved to /var/cache/conftool/dbconfig/20260728-130107-cwilliams.json * 13:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2237.codfw.wmnet with reason: Maintenance * 13:00 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2236: Maintenance * 12:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2153: Maintenance * 12:52 root@cumin1003: START - Cookbook sre.mysql.pool pool db2221: Maintenance * 12:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2153 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95358 and previous config saved to /var/cache/conftool/dbconfig/20260728-125214-cwilliams.json * 12:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2153.codfw.wmnet with reason: Maintenance * 12:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2167: Maintenance * 12:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-codfw * 12:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2011.codfw.wmnet * 12:51 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2011.codfw.wmnet * 12:49 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:49 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback for asw1-603 - pt1979@cumin2003" * 12:48 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback for asw1-603 - pt1979@cumin2003" * 12:46 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2011.codfw.wmnet * 12:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2221 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95357 and previous config saved to /var/cache/conftool/dbconfig/20260728-124601-cwilliams.json * 12:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2221.codfw.wmnet with reason: Maintenance * 12:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2218: Maintenance * 12:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2167 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95354 and previous config saved to /var/cache/conftool/dbconfig/20260728-124457-cwilliams.json * 12:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2167.codfw.wmnet with reason: Maintenance * 12:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2166: Maintenance * 12:42 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 12:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2011.codfw.wmnet * 12:41 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2010.codfw.wmnet * 12:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2010.codfw.wmnet * 12:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2010.codfw.wmnet * 12:34 pt1979@cumin2003: START - Cookbook sre.dns.netbox * 12:32 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 12:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2010.codfw.wmnet * 12:31 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2009.codfw.wmnet * 12:31 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2009.codfw.wmnet * 12:27 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2009.codfw.wmnet * 12:22 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2009.codfw.wmnet * 12:22 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2008.codfw.wmnet * 12:21 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2008.codfw.wmnet * 12:16 pt1979@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-604-eqsin * 12:16 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2008.codfw.wmnet * 12:16 pt1979@cumin1003: START - Cookbook sre.network.tls for network device asw1-604-eqsin * 12:14 root@cumin1003: START - Cookbook sre.mysql.pool pool db2236: Maintenance * 12:14 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2236: Maintenance * 12:12 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1137.eqiad.wmnet with OS trixie * 12:11 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2008.codfw.wmnet * 12:11 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2007.codfw.wmnet * 12:11 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2007.codfw.wmnet * 12:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db2236: Maintenance * 12:06 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2007.codfw.wmnet * 12:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2236 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95348 and previous config saved to /var/cache/conftool/dbconfig/20260728-120253-cwilliams.json * 12:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2236.codfw.wmnet with reason: Maintenance * 12:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2007.codfw.wmnet * 12:01 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 12:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 11:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2218: Maintenance * 11:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db2166: Maintenance * 11:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2219: Maintenance * 11:56 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2006.codfw.wmnet * 11:52 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1137.eqiad.wmnet with reason: host reimage * 11:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2218 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95344 and previous config saved to /var/cache/conftool/dbconfig/20260728-115155-cwilliams.json * 11:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2218.codfw.wmnet with reason: Maintenance * 11:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2208: Maintenance * 11:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2166 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95342 and previous config saved to /var/cache/conftool/dbconfig/20260728-115119-cwilliams.json * 11:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2166.codfw.wmnet with reason: Maintenance * 11:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2164: Maintenance * 11:47 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1137.eqiad.wmnet with reason: host reimage * 11:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2006.codfw.wmnet * 11:45 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2005.codfw.wmnet * 11:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2005.codfw.wmnet * 11:40 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2005.codfw.wmnet * 11:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2005.codfw.wmnet * 11:35 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2004.codfw.wmnet * 11:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2004.codfw.wmnet * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1137 * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1137 * 11:30 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1137 * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1137.eqiad.wmnet 192.32.64.10.in-addr.arpa 2.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:30 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1137.eqiad.wmnet 192.32.64.10.in-addr.arpa 2.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1137 - jiji@cumin1003" * 11:25 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2004.codfw.wmnet * 11:19 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2004.codfw.wmnet * 11:19 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2003.codfw.wmnet * 11:19 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2003.codfw.wmnet * 11:14 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2003.codfw.wmnet * 11:11 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2219: Maintenance * 11:10 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2219: Maintenance * 11:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2219: Maintenance * 11:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2003.codfw.wmnet * 11:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2164: Maintenance * 11:03 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2002.codfw.wmnet * 11:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2002.codfw.wmnet * 11:03 root@cumin1003: START - Cookbook sre.mysql.pool pool db2208: Maintenance * 10:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2164 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95332 and previous config saved to /var/cache/conftool/dbconfig/20260728-105749-cwilliams.json * 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2164.codfw.wmnet with reason: Maintenance * 10:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2208 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95331 and previous config saved to /var/cache/conftool/dbconfig/20260728-105711-cwilliams.json * 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2208.codfw.wmnet with reason: Maintenance * 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2219 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95330 and previous config saved to /var/cache/conftool/dbconfig/20260728-105652-cwilliams.json * 10:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2219.codfw.wmnet with reason: Maintenance * 10:53 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1137 - jiji@cumin1003" * 10:52 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2002.codfw.wmnet * 10:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2002.codfw.wmnet * 10:47 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2001.codfw.wmnet * 10:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2001.codfw.wmnet * 10:39 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2001.codfw.wmnet * 10:35 jiji@cumin1003: START - Cookbook sre.dns.netbox * 10:34 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] (duration: 09m 31s) * 10:34 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1137 * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2001.codfw.wmnet * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-codfw * 10:34 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1137.eqiad.wmnet with OS trixie * 10:28 jforrester@deploy1003: jforrester: Continuing with deployment * 10:27 jforrester@deploy1003: jforrester: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:25 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] * 10:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 10:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 10:21 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 10:20 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1137.eqiad.wmnet * 10:20 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 10:20 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1137.eqiad.wmnet * 10:20 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1137.eqiad.wmnet * 10:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2163: Maintenance * 09:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-staging-worker * 09:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2003.codfw.wmnet * 09:37 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2003.codfw.wmnet * 09:32 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 09:31 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2003.codfw.wmnet * 09:30 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 09:30 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2163: Maintenance * 09:30 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 09:30 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 09:30 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:22 klausman@cumin1003: END (ERROR) - Cookbook sre.ganeti.reboot-vm (exit_code=97) for VM ml-serve-ctrl2001.codfw.wmnet * 09:22 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2001.codfw.wmnet * 09:22 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d8-eqiad * 09:22 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d8-eqiad * 09:21 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2003.codfw.wmnet * 09:20 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2002.codfw.wmnet * 09:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2002.codfw.wmnet * 09:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f2-codfw * 09:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f2-codfw * 09:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e4-codfw * 09:18 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2163: Maintenance * 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e4-codfw * 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-codfw * 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-codfw * 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e5-codfw * 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e5-codfw * 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f4-codfw * 09:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2210: Maintenance * 09:16 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f4-codfw * 09:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2182: Maintenance * 09:14 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2002.codfw.wmnet * 09:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db2163: Maintenance * 09:11 XioNoX: rebooting cr2-drmrs - [[phab:T431749|T431749]] * 09:10 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr2-drmrs,cr2-drmrs IPv6,cr2-drmrs.mgmt with reason: router upgrade * 09:06 XioNoX: draining cr2-drmrs - [[phab:T431749|T431749]] * 09:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2163 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95320 and previous config saved to /var/cache/conftool/dbconfig/20260728-090638-cwilliams.json * 09:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2163.codfw.wmnet with reason: Maintenance * 09:06 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2161: Maintenance * 09:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2002.codfw.wmnet * 09:04 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2001.codfw.wmnet * 09:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2001.codfw.wmnet * 08:57 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2001.codfw.wmnet * 08:48 XioNoX: un-drain cr1-drmrs - [[phab:T431749|T431749]] * 08:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2001.codfw.wmnet * 08:47 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-staging-worker * 08:42 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:35 XioNoX: rebooting cr1-drmrs - [[phab:T431749|T431749]] * 08:33 XioNoX: draining cr1-drmrs - [[phab:T431749|T431749]] * 08:31 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2210: Maintenance * 08:29 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2182: Maintenance * 08:21 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2182: Maintenance * 08:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db2161: Maintenance * 08:16 root@cumin1003: START - Cookbook sre.mysql.pool pool db2182: Maintenance * 08:12 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2210: Maintenance * 08:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2161 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95309 and previous config saved to /var/cache/conftool/dbconfig/20260728-081044-cwilliams.json * 08:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2161.codfw.wmnet with reason: Maintenance * 08:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2154: Maintenance * 08:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2182 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95307 and previous config saved to /var/cache/conftool/dbconfig/20260728-080947-cwilliams.json * 08:09 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2182.codfw.wmnet with reason: Maintenance * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2168: Maintenance * 08:06 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr1-drmrs,cr1-drmrs IPv6,cr1-drmrs.mgmt with reason: router upgrade * 08:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db2210: Maintenance * 08:05 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 08:05 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 08:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2210 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95305 and previous config saved to /var/cache/conftool/dbconfig/20260728-080008-cwilliams.json * 08:00 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2210.codfw.wmnet with reason: Maintenance * 07:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2206: Maintenance * 07:50 gkyziridis@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 07:50 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 07:22 root@cumin1003: START - Cookbook sre.mysql.pool pool db2154: Maintenance * 07:22 root@cumin1003: START - Cookbook sre.mysql.pool pool db2168: Maintenance * 07:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2154 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95295 and previous config saved to /var/cache/conftool/dbconfig/20260728-071640-cwilliams.json * 07:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2154.codfw.wmnet with reason: Maintenance * 07:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95294 and previous config saved to /var/cache/conftool/dbconfig/20260728-071604-cwilliams.json * 07:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2168.codfw.wmnet with reason: Maintenance * 07:08 root@cumin1003: START - Cookbook sre.mysql.pool pool db2206: Maintenance * 07:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2206 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95292 and previous config saved to /var/cache/conftool/dbconfig/20260728-070219-cwilliams.json * 07:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2206.codfw.wmnet with reason: Maintenance * 06:44 marostegui: Failover m5 from db1164 to db1228 - [[phab:T432967|T432967]] * 06:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2235].codfw.wmnet,db[1164,1217,1228].eqiad.wmnet with reason: m5 master switch [[phab:T432967|T432967]] * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.10 (duration: 02m 34s) * 03:39 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] (duration: 36m 06s) * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 02:57 dzahn@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1004.eqiad.wmnet with OS trixie * 02:57 dzahn@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - dzahn@cumin1003" * 02:55 dzahn@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - dzahn@cumin1003" * 02:37 dzahn@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1004.eqiad.wmnet with reason: host reimage * 02:31 dzahn@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1004.eqiad.wmnet with reason: host reimage * 02:16 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie * 02:15 dzahn@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host zuul1004.eqiad.wmnet with OS trixie * 01:43 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie * 01:43 dzahn@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1004.eqiad.wmnet with OS trixie * 01:25 pt1979@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-603-eqsin * 01:24 pt1979@cumin1003: START - Cookbook sre.network.tls for network device asw1-603-eqsin * 01:12 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 01:12 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt for new switches in eqsin - pt1979@cumin2003" * 01:12 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt for new switches in eqsin - pt1979@cumin2003" * 01:08 pt1979@cumin2003: START - Cookbook sre.dns.netbox * 00:48 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 00:47 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 00:47 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 00:47 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 00:26 mutante: attempting reimage with trixie on zuul1004 re-purposed physical hardware - dcops reported install issue - host was in busybox shell ([[phab:T427353|T427353]]) * 00:24 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie == 2026-07-27 == * 23:50 Amir1: mass deleting vp8 transcodes * 23:28 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:27 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1004.eqiad.wmnet with OS bullseye * 23:26 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 23:25 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:25 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 22:39 maryum: Deploy security fix for [[phab:T432877|T432877]] * 22:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1022.eqiad.wmnet with OS bookworm * 22:37 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS bullseye * 22:32 sbassett: Deployed security fix for [[phab:T432789|T432789]] * 22:22 sbassett: Deployed security patch for [[phab:T431819|T431819]] * 22:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1022.eqiad.wmnet with reason: host reimage * 22:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1022.eqiad.wmnet with reason: host reimage * 22:01 RScout-WMF: Deployed security fix for [[phab:T431819|T431819]] * 22:00 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2012 * 21:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2012.codfw.wmnet with OS bookworm * 21:55 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2011\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 21:45 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1022 * 21:45 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1022 * 21:44 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1022 * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1022.eqiad.wmnet 239.48.64.10.in-addr.arpa 9.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:44 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1022.eqiad.wmnet 239.48.64.10.in-addr.arpa 9.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:41 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:41 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 21:34 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS bookworm * 21:31 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:22 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:21 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:19 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:17 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1004 * 21:16 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1004 * 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1004] - vriley@cumin1003" * 21:15 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1004] - vriley@cumin1003" * 21:11 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:10 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2011.codfw.wmnet, repooling source-only afterwards * 21:05 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:01 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1022 * 20:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1022.eqiad.wmnet with OS bookworm * 20:53 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1021.eqiad.wmnet, repooling source-only afterwards * 20:51 mutante: zuul1001 - re-enabled puppet - revert "cherry-picked" gerrit:1314120 - [[phab:T431003|T431003]] * 20:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Maintenance * 20:15 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] (duration: 08m 03s) * 20:11 sbisson@deploy1003: sbisson: Continuing with deployment * 20:09 sbisson@deploy1003: sbisson: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] * 19:47 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 46s) * 19:47 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Maintenance * 19:27 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] (duration: 12m 26s) * 19:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2228 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95285 and previous config saved to /var/cache/conftool/dbconfig/20260727-192711-cwilliams.json * 19:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2228.codfw.wmnet with reason: Maintenance * 19:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2223: Maintenance * 19:23 krinkle@deploy1003: krinkle: Continuing with deployment * 19:16 krinkle@deploy1003: krinkle: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:15 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] * 19:12 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2238: Maintenance * 18:58 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 18:57 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 18:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2227: Maintenance * 18:57 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-experimental: apply * 18:55 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-experimental: apply * 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1021.eqiad.wmnet with OS bookworm * 18:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2011.codfw.wmnet with OS bookworm * 18:40 root@cumin1003: START - Cookbook sre.mysql.pool pool db2223: Maintenance * 18:39 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] (duration: 07m 05s) * 18:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2223 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95275 and previous config saved to /var/cache/conftool/dbconfig/20260727-183500-cwilliams.json * 18:34 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2223.codfw.wmnet with reason: Maintenance * 18:34 musikanimal@deploy1003: musikanimal: Continuing with deployment * 18:34 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2213: Maintenance * 18:33 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:32 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] * 18:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db2238: Maintenance * 18:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2011.codfw.wmnet with reason: host reimage * 18:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2238 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95271 and previous config saved to /var/cache/conftool/dbconfig/20260727-181944-cwilliams.json * 18:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2238.codfw.wmnet with reason: Maintenance * 18:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2226: Maintenance * 18:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1021.eqiad.wmnet with reason: host reimage * 18:14 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2011.codfw.wmnet with reason: host reimage * 18:12 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1021.eqiad.wmnet with reason: host reimage * 18:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db2227: Maintenance * 18:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2227 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95265 and previous config saved to /var/cache/conftool/dbconfig/20260727-180256-cwilliams.json * 18:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2227.codfw.wmnet with reason: Maintenance * 18:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2194: Maintenance * 17:57 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2011 * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2011 * 17:56 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2011 * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2011.codfw.wmnet 37.32.192.10.in-addr.arpa 7.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:56 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2011.codfw.wmnet 37.32.192.10.in-addr.arpa 7.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2011 - bking@cumin2003" * 17:56 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2011 - bking@cumin2003" * 17:52 bking@cumin2003: START - Cookbook sre.dns.netbox * 17:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2011 * 17:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1021 * 17:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1021 * 17:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2011.codfw.wmnet with OS bookworm * 17:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1021.eqiad.wmnet with OS bookworm * 17:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Maintenance * 17:38 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2010\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 17:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2213 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95260 and previous config saved to /var/cache/conftool/dbconfig/20260727-173740-cwilliams.json * 17:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2213.codfw.wmnet with reason: Maintenance * 17:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2211: Maintenance * 17:36 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1020\.eqiad\.wmnet,dc=eqiad,cluster=wdqs\-main,service=wdqs\-main * 17:32 root@cumin1003: START - Cookbook sre.mysql.pool pool db2226: Maintenance * 17:31 taavi@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] (duration: 06m 33s) * 17:27 taavi@deploy1003: taavi: Continuing with deployment * 17:27 taavi@deploy1003: taavi: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:26 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2226 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95256 and previous config saved to /var/cache/conftool/dbconfig/20260727-172636-cwilliams.json * 17:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2226.codfw.wmnet with reason: Maintenance * 17:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2225: Maintenance * 17:25 taavi@deploy1003: Started scap sync-world: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] * 17:13 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 17:11 root@cumin1003: START - Cookbook sre.mysql.pool pool db2194: Maintenance * 17:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2194 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95248 and previous config saved to /var/cache/conftool/dbconfig/20260727-170453-cwilliams.json * 17:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2194.codfw.wmnet with reason: Maintenance * 17:04 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2190: Maintenance * 16:52 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 16:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2211: Maintenance * 16:40 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95242 and previous config saved to /var/cache/conftool/dbconfig/20260727-164015-cwilliams.json * 16:40 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2211.codfw.wmnet with reason: Maintenance * 16:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2178: Maintenance * 16:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db2225: Maintenance * 16:39 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2172: Maintenance * 16:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 16:38 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 16:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2225 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95238 and previous config saved to /var/cache/conftool/dbconfig/20260727-163307-cwilliams.json * 16:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2225.codfw.wmnet with reason: Maintenance * 16:32 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2189: Maintenance * 16:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db2190: Maintenance * 16:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2190 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95230 and previous config saved to /var/cache/conftool/dbconfig/20260727-160602-cwilliams.json * 16:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2190.codfw.wmnet with reason: Maintenance * 15:53 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2177: Maintenance * 15:53 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2172: Maintenance * 15:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2178: Maintenance * 15:51 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2172: Maintenance * 15:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2178 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95224 and previous config saved to /var/cache/conftool/dbconfig/20260727-154559-cwilliams.json * 15:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2172: Maintenance * 15:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2178.codfw.wmnet with reason: Maintenance * 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2171: Maintenance * 15:44 root@cumin1003: START - Cookbook sre.mysql.pool pool db2189: Maintenance * 15:43 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:41 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2172 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95222 and previous config saved to /var/cache/conftool/dbconfig/20260727-153927-cwilliams.json * 15:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2172.codfw.wmnet with reason: Maintenance * 15:38 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2189 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95220 and previous config saved to /var/cache/conftool/dbconfig/20260727-153833-cwilliams.json * 15:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2189.codfw.wmnet with reason: Maintenance * 15:34 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:32 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 15:32 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 15:31 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:29 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:26 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 15:22 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] (duration: 07m 00s) * 15:21 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2155: Maintenance * 15:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2175: Maintenance * 15:18 zabe@deploy1003: zabe: Continuing with deployment * 15:17 zabe@deploy1003: zabe: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:15 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 15:15 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:15 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] * 15:15 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 06s) * 15:15 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:12 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2010.codfw.wmnet with OS bookworm * 15:08 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2177: Maintenance * 15:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2177: Maintenance * 14:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2177: Maintenance * 14:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2171: Maintenance * 14:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2171 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95209 and previous config saved to /var/cache/conftool/dbconfig/20260727-145236-cwilliams.json * 14:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2171.codfw.wmnet with reason: Maintenance * 14:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2177 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95208 and previous config saved to /var/cache/conftool/dbconfig/20260727-145206-cwilliams.json * 14:52 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2157: Maintenance * 14:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2177.codfw.wmnet with reason: Maintenance * 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1020.eqiad.wmnet with OS bookworm * 14:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2156: Maintenance * 14:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2010.codfw.wmnet with reason: host reimage * 14:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2010.codfw.wmnet with reason: host reimage * 14:41 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[1031,2024]*: Upgrade Cassandra to 5.0.8 (canary) - eevans@cumin1003 * 14:34 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2155: Maintenance * 14:33 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2175: Maintenance * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2010 * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2010 * 14:24 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2010 * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2010.codfw.wmnet 94.16.192.10.in-addr.arpa 4.9.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2010.codfw.wmnet 94.16.192.10.in-addr.arpa 4.9.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2010 - bking@cumin2003" * 14:24 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2010 - bking@cumin2003" * 14:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1020.eqiad.wmnet with reason: host reimage * 14:23 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[1031,2024]*: Upgrade Cassandra to 5.0.8 (canary) - eevans@cumin1003 * 14:20 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 14:20 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 14:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1020.eqiad.wmnet with reason: host reimage * 14:17 sukhe: sudo gnt-instance reboot urldownloader1005.wikimedia.org * 14:16 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:15 jelto@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:08 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2155: Maintenance * 14:05 root@cumin1003: START - Cookbook sre.mysql.pool pool db2157: Maintenance * 14:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2175: Maintenance * 14:03 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 11 hosts * 14:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db2155: Maintenance * 14:01 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 11 hosts * 14:01 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1136.eqiad.wmnet * 14:01 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1136.eqiad.wmnet * 14:01 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1136.eqiad.wmnet * 14:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db2156: Maintenance * 13:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2157 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95194 and previous config saved to /var/cache/conftool/dbconfig/20260727-135943-cwilliams.json * 13:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2157.codfw.wmnet with reason: Maintenance * 13:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db2175: Maintenance * 13:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 13:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 13:57 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2010 * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2155 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95193 and previous config saved to /var/cache/conftool/dbconfig/20260727-135613-cwilliams.json * 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2155.codfw.wmnet with reason: Maintenance * 13:55 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1020 * 13:55 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1020 * 13:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2156 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95192 and previous config saved to /var/cache/conftool/dbconfig/20260727-135413-cwilliams.json * 13:54 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2156.codfw.wmnet with reason: Maintenance * 13:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2175 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95191 and previous config saved to /var/cache/conftool/dbconfig/20260727-135300-cwilliams.json * 13:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2010.codfw.wmnet with OS bookworm * 13:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2175.codfw.wmnet with reason: Maintenance * 13:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1020.eqiad.wmnet with OS bookworm * 13:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 34 hosts * 13:46 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 34 hosts * 13:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1201: Maintenance * 13:27 Lucas_WMDE: UTC afternoon backport+config window doen * 13:18 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] (duration: 11m 57s) * 13:14 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, sihe: Continuing with deployment * 13:08 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, sihe: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:07 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool ulsfo [reason: router upgrade finished, [[phab:T431752|T431752]]] * 13:07 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool ulsfo [reason: router upgrade finished, [[phab:T431752|T431752]]] * 13:06 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] * 13:03 XioNoX: repool cr4-ulsfo - [[phab:T431752|T431752]] * 12:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db1201: Maintenance * 12:48 gkyziridis@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1201 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95186 and previous config saved to /var/cache/conftool/dbconfig/20260727-124404-cwilliams.json * 12:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1201.eqiad.wmnet with reason: Maintenance * 12:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1187: Maintenance * 12:30 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader.eqiad.wikimedia.org on all recursors * 12:30 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader.eqiad.wikimedia.org on all recursors * 12:30 sukhe@dns1004: END - running authdns-update * 12:30 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] (duration: 09m 32s) * 12:28 sukhe@dns1004: START - running authdns-update * 12:25 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 12:22 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:20 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] * 12:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts * 12:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts * 12:13 XioNoX: rebooting cr4-ulsfo for upgrade - [[phab:T431752|T431752]] * 12:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: es1038 repool * 12:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 38 hosts * 12:08 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 38 hosts * 11:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1187: Maintenance * 11:53 urbanecm@deploy1003: mwscript-k8s job started: foreachwikiindblist growthexperiments GrowthExperiments:cleanMentorList # [[phab:T431804|T431804]] * 11:50 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr4-ulsfo,cr4-ulsfo IPv6,cr4-ulsfo.mgmt with reason: router upgrade * 11:50 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] (duration: 11m 07s) * 11:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1187 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95178 and previous config saved to /var/cache/conftool/dbconfig/20260727-114844-cwilliams.json * 11:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1187.eqiad.wmnet with reason: Maintenance * 11:43 urbanecm@deploy1003: urbanecm: Continuing with deployment * 11:42 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:39 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] * 11:37 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 11:36 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 11:36 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 11:35 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 11:29 XioNoX: start draining cr4-ulsfo - [[phab:T431752|T431752]] * 11:29 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 11:29 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 11:28 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1035: testing * 11:28 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1035: testing * 11:27 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1035: testing * 11:27 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1035: testing * 11:26 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool ulsfo [reason: router upgrade, [[phab:T431752|T431752]]] * 11:26 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1038: es1038 repool * 11:26 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool ulsfo [reason: router upgrade, [[phab:T431752|T431752]]] * 11:26 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1038: testing * 11:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1264: Maintenance * 11:24 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1038: testing * 11:23 marostegui@cumin1003: dbctl commit (dc=all): 'Repool es1050 as master', diff saved to https://phabricator.wikimedia.org/P95170 and previous config saved to /var/cache/conftool/dbconfig/20260727-112326-marostegui.json * 11:23 marostegui@cumin1003: dbctl commit (dc=all): 'Repool es1050', diff saved to https://phabricator.wikimedia.org/P95169 and previous config saved to /var/cache/conftool/dbconfig/20260727-112302-marostegui.json * 11:22 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1050: testing * 11:22 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1050: testing * 11:20 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 11:18 blake@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 11:18 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 11:12 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 11:11 blake@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 11:09 blake@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 11:09 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 11:09 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 11:08 blake@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 11:05 blake@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 11:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 11:02 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 10:50 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 10:43 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:39 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply * 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1264: Maintenance * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply * 10:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply * 10:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 10:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 10:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 10:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1264 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95164 and previous config saved to /var/cache/conftool/dbconfig/20260727-103204-cwilliams.json * 10:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1264.eqiad.wmnet with reason: Maintenance * 10:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 10:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 10:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 10:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 10:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 10:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 10:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 10:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1237: Maintenance * 10:04 elukey: restart burrow main-eqiad on kafkamon2003 to clear some errors on kafka-main1008 * 09:58 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1136.eqiad.wmnet with OS trixie * 09:39 elukey: restart burrow-main-eqiad.service on kafkamon1003 to see if a recurrent kafka error on kafka-main1008 goes away * 09:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1237: Maintenance * 09:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1136.eqiad.wmnet with reason: host reimage * 09:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1237 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95159 and previous config saved to /var/cache/conftool/dbconfig/20260727-093328-cwilliams.json * 09:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1237.eqiad.wmnet with reason: Maintenance * 09:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1136.eqiad.wmnet with reason: host reimage * 09:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1203: Maintenance * 09:17 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1136 * 09:17 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1136 * 09:04 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1136 * 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1136.eqiad.wmnet 191.32.64.10.in-addr.arpa 1.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:04 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1136.eqiad.wmnet 191.32.64.10.in-addr.arpa 1.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1136 - jiji@cumin1003" * 09:04 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1136 - jiji@cumin1003" * 08:52 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 08:52 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 08:52 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 08:51 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 08:50 jiji@cumin1003: START - Cookbook sre.dns.netbox * 08:47 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1136 * 08:46 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1136.eqiad.wmnet with OS trixie * 08:44 marostegui: Rename tables on s3 [[phab:T425066|T425066]] * 08:43 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1136.eqiad.wmnet * 08:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db1203: Maintenance * 08:43 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1136.eqiad.wmnet * 08:43 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1136.eqiad.wmnet * 08:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1203 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95154 and previous config saved to /var/cache/conftool/dbconfig/20260727-083703-cwilliams.json * 08:36 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1203.eqiad.wmnet with reason: Maintenance * 08:16 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1179: Maintenance * 07:44 phuedx: UTC morning backport window done * 07:37 phuedx@deploy1003: Finished scap sync-world: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] (duration: 32m 33s) * 07:28 root@cumin1003: START - Cookbook sre.mysql.pool pool db1179: Maintenance * 07:26 marostegui: Rename tables on s3 [[phab:T426341|T426341]] * 07:25 phuedx@deploy1003: phuedx: Continuing with deployment * 07:22 marostegui: Drop tables in akwiki nawiki pihwiki - growthexperiments_* [[phab:T428885|T428885]] * 07:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95149 and previous config saved to /var/cache/conftool/dbconfig/20260727-072234-cwilliams.json * 07:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1179.eqiad.wmnet with reason: Maintenance * 07:20 phuedx@deploy1003: phuedx: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:16 ryankemper: [[phab:T430880|T430880]] [WDQS] Reimaged `wdqs1018` and `wdqs1019` to Bookworm, restored data using test-cookbook change {{Gerrit|1317128}}, and repooled both; 25/36 hosts complete * 07:04 phuedx@deploy1003: Started scap sync-world: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] * 06:57 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1019.eqiad.wmnet * 06:56 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1018.eqiad.wmnet * 06:40 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1020.eqiad.wmnet with reason: Cloning * 06:35 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db1228.eqiad.wmnet with reason: Rebooting * 06:29 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:29 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:25 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:25 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:25 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1019.eqiad.wmnet, repooling source-only afterwards * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1018.eqiad.wmnet, repooling source-only afterwards * 04:51 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1019.eqiad.wmnet, repooling source-only afterwards * 04:51 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1018.eqiad.wmnet, repooling source-only afterwards * 04:48 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s) * 04:48 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 04:48 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s) * 04:48 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 36s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-26 == * 14:59 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:59 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:59 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:59 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1019.eqiad.wmnet with OS bookworm * 01:05 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1018.eqiad.wmnet with OS bookworm * 00:43 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1019.eqiad.wmnet with reason: host reimage * 00:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1018.eqiad.wmnet with reason: host reimage * 00:34 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1019.eqiad.wmnet with reason: host reimage * 00:33 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1018.eqiad.wmnet with reason: host reimage * 00:16 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 00:16 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 00:15 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 00:15 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1019 * 00:11 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1019 * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1018 * 00:11 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1018 * 00:08 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1019.eqiad.wmnet with OS bookworm * 00:08 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1018.eqiad.wmnet with OS bookworm == 2026-07-25 == * 22:06 ryankemper: [[phab:T430880|T430880]] [WDQS] Repooled `wdqs1017` and `wdqs2024` after reimaging to bookworm, scap deploying, and data xfering * 22:04 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2024.codfw.wmnet * 22:03 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1017.eqiad.wmnet * 21:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1017.eqiad.wmnet, repooling source-only afterwards * 21:06 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2024.codfw.wmnet, repooling source-only afterwards * 20:52 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:52 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:52 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:52 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 20:18 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1017.eqiad.wmnet, repooling source-only afterwards * 20:18 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2024.codfw.wmnet, repooling source-only afterwards * 20:15 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:15 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:15 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:15 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 19:57 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s) * 19:57 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 19:57 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 07s) * 19:57 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 19:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2024.codfw.wmnet with OS bookworm * 19:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1017.eqiad.wmnet with OS bookworm * 19:02 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2024.codfw.wmnet with reason: host reimage * 18:58 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1017.eqiad.wmnet with reason: host reimage * 18:53 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2024.codfw.wmnet with reason: host reimage * 18:52 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1017.eqiad.wmnet with reason: host reimage * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2024 * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2024 * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1017 * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1017 * 18:27 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2024 * 18:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2024.codfw.wmnet 58.16.192.10.in-addr.arpa 8.5.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:26 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2024.codfw.wmnet 58.16.192.10.in-addr.arpa 8.5.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:24 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1017 * 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1017.eqiad.wmnet 238.48.64.10.in-addr.arpa 8.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:24 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1017.eqiad.wmnet 238.48.64.10.in-addr.arpa 8.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1017 - ryankemper@cumin2003" * 18:24 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1017 - ryankemper@cumin2003" * 18:23 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 18:18 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 18:17 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1017 * 18:17 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2024 * 18:14 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1017.eqiad.wmnet with OS bookworm * 18:14 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2024.codfw.wmnet with OS bookworm * 18:05 ryankemper: [WDQS] [[phab:T430880|T430880]] Reimaged `wdqs1016` and `wdqs2023` to Bookworm with `--move-vlan`, restored main and scholarly data, validated postflights, and repooled both hosts. Confirmed PyBal rebuilt both backends with their new addresses * 17:45 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2023.codfw.wmnet * 17:43 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1016.eqiad.wmnet * 06:35 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1016.eqiad.wmnet, repooling source-only afterwards * 06:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2023.codfw.wmnet, repooling source-only afterwards * 05:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2023.codfw.wmnet, repooling source-only afterwards * 05:19 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1016.eqiad.wmnet, repooling source-only afterwards * 05:07 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 07s) * 05:07 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 05:06 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 06s) * 05:06 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 03:27 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2023.codfw.wmnet with OS bookworm * 02:59 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2023.codfw.wmnet with reason: host reimage * 02:56 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2023.codfw.wmnet with reason: host reimage * 02:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2023 * 02:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2023 * 02:30 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2023 * 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2023.codfw.wmnet 35.0.192.10.in-addr.arpa 5.3.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 02:30 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2023.codfw.wmnet 35.0.192.10.in-addr.arpa 5.3.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2023 - ryankemper@cumin2003" * 02:30 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2023 - ryankemper@cumin2003" * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 26s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:15 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1016.eqiad.wmnet with OS bookworm * 00:49 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1016.eqiad.wmnet with reason: host reimage * 00:43 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1016.eqiad.wmnet with reason: host reimage * 00:31 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 00:27 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1016 * 00:27 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1016 * 00:27 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2023 * 00:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1016.eqiad.wmnet with OS bookworm * 00:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2023.codfw.wmnet with OS bookworm * 00:11 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs1014.eqiad.wmnet and wdqs2008.codfw.wmnet after Bookworm reimage, transfer, and postflight; wdqs2008 is serving, while wdqs1014 will remain outside of service until a pybal restart next monday == 2026-07-24 == * 23:54 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1014.eqiad.wmnet * 23:54 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2008.codfw.wmnet * 23:43 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2010.codfw.wmnet with OS trixie * 23:08 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 23:03 jhathaway@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 22:33 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 22:13 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 22:13 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 22:13 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 22:13 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:00 jhathaway@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 21:53 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 21:53 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie * 21:51 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 21:47 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie * 21:43 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 21:39 jhathaway@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 21:38 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 17:21 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1135.eqiad.wmnet * 17:21 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1135.eqiad.wmnet * 17:21 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1135.eqiad.wmnet * 16:34 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 16:34 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 16:34 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 16:34 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 16:33 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 16:33 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 16:28 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:28 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:28 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:28 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2008.codfw.wmnet, repooling source-only afterwards * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1014.eqiad.wmnet, repooling source-only afterwards * 15:56 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1135.eqiad.wmnet with OS trixie * 15:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 40 hosts * 15:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 40 hosts * 15:37 topranks: upgrade SR-Linux OS on lswtest-d8-eqiad * 15:36 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1135.eqiad.wmnet with reason: host reimage * 15:33 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 6 hosts with reason: upgrade lswtest-d8-eqiad * 15:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1135.eqiad.wmnet with reason: host reimage * 15:30 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc-gp2006.codfw.wmnet with OS bookworm * 15:15 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1135 * 15:15 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1135 * 15:13 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc-gp2006.codfw.wmnet with reason: host reimage * 15:08 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc-gp2006.codfw.wmnet with reason: host reimage * 14:49 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm * 14:48 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host mc-gp2006.codfw.wmnet with OS bookworm * 14:34 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] (duration: 41m 12s) * 14:32 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1135 * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1135.eqiad.wmnet 177.32.64.10.in-addr.arpa 7.7.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:32 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1135.eqiad.wmnet 177.32.64.10.in-addr.arpa 7.7.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1135 - jiji@cumin1003" * 14:32 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1135 - jiji@cumin1003" * 14:29 krinkle@deploy1003: krinkle: Continuing with deployment * 14:29 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm * 14:27 jiji@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host mc-gp2006.codfw.wmnet with OS bookworm * 14:26 jiji@cumin1003: START - Cookbook sre.dns.netbox * 14:15 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1135 * 14:14 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1135.eqiad.wmnet with OS trixie * 14:14 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1135.eqiad.wmnet * 14:13 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1135.eqiad.wmnet * 14:13 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1135.eqiad.wmnet * 14:10 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1072.eqiad.wmnet * 14:10 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1072.eqiad.wmnet * 14:10 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1072.eqiad.wmnet * 14:10 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1072.eqiad.wmnet * 14:09 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1071.eqiad.wmnet * 14:09 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1071.eqiad.wmnet * 14:09 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1071.eqiad.wmnet * 14:09 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1071.eqiad.wmnet * 13:58 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 13:58 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:58 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:57 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:55 krinkle@deploy1003: krinkle: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:53 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] * 13:45 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:45 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push new IPs for mc-gp2006 - cmooney@cumin1003" * 13:45 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push new IPs for mc-gp2006 - cmooney@cumin1003" * 13:44 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) mc-gp2006.codfw.wmnet on all recursors * 13:44 cmooney@cumin1003: START - Cookbook sre.dns.wipe-cache mc-gp2006.codfw.wmnet on all recursors * 13:42 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm * 13:41 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:30 papaul: reboot mr1-eqsin for maintenance * 13:24 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb[1029-1031].eqiad.wmnet * 13:10 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb[1029-1031].eqiad.wmnet * 11:33 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 7 hosts * 11:11 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 7 hosts * 10:56 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 7 hosts * 10:47 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 7 hosts * 10:44 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:44 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:41 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:41 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:35 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 8 hosts * 10:34 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:33 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:32 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:32 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:31 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:31 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:30 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 8 hosts * 10:24 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie * 10:19 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:18 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 16 hosts * 10:17 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2001.codfw.wmnet * 10:13 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2001.codfw.wmnet * 10:12 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2001.codfw.wmnet * 10:02 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2001.codfw.wmnet * 10:02 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2002.codfw.wmnet * 09:57 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2002.codfw.wmnet * 09:56 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2002.codfw.wmnet * 09:51 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2002.codfw.wmnet * 09:51 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1002.eqiad.wmnet * 09:47 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1002.eqiad.wmnet * 09:47 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1001.eqiad.wmnet * 09:44 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1001.eqiad.wmnet * 09:34 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2003.codfw.wmnet * 09:32 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2003.codfw.wmnet * 09:32 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2002.codfw.wmnet * 09:30 brouberol@dns1004: END - running authdns-update * 09:29 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2002.codfw.wmnet * 09:29 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2001.codfw.wmnet * 09:27 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 9 hosts * 09:27 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2001.codfw.wmnet * 09:27 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2001.codfw.wmnet * 09:26 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 9 hosts * 09:26 brouberol@dns1004: START - running authdns-update * 09:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 57 hosts * 09:24 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2001.codfw.wmnet * 09:24 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2002.codfw.wmnet * 09:22 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2002.codfw.wmnet * 09:21 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 57 hosts * 09:20 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2003.codfw.wmnet * 09:19 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 16 hosts * 09:16 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2003.codfw.wmnet * 09:16 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1003.eqiad.wmnet * 09:15 urbanecm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 09:15 urbanecm@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 09:13 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1003.eqiad.wmnet * 09:13 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1002.eqiad.wmnet * 09:11 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1002.eqiad.wmnet * 09:11 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1001.eqiad.wmnet * 09:07 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1001.eqiad.wmnet * 08:32 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:24 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:16 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 08:16 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 08:07 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:07 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:07 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 08:02 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 08:01 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:59 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:57 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 07:57 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 06:46 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1025.eqiad.wmnet with reason: Cloning * 06:46 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s4 * 06:45 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s6 * 06:44 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1019.eqiad.wmnet,service=s6 * 06:44 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1019.eqiad.wmnet,service=s4 * 03:40 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:40 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:40 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:40 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 03:37 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:37 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:37 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:36 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:49 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on mr1-eqsin,mr1-eqsin IPv6 with reason: connection issue * 02:38 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on cr[2-3]-eqsin.mgmt,ps1-[603-604]-eqsin with reason: connection issue * 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 27s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-23 == * 23:27 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin.oob,mr1-eqsin.oob IPv6 with reason: switch refresh * 22:21 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Setting storage compatibility to NONE - eevans@cumin1003 * 22:01 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Setting storage compatibility to NONE - eevans@cumin1003 * 21:29 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1014.eqiad.wmnet, repooling source-only afterwards * 21:28 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 46s) * 21:28 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 21:19 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Setting storage compatibility to UPGRADING - eevans@cumin1003 * 21:00 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Setting storage compatibility to UPGRADING - eevans@cumin1003 * 20:17 dani@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] (duration: 11m 57s) * 20:13 dani@deploy1003: dani: Continuing with deployment * 20:07 dani@deploy1003: dani: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:05 dani@deploy1003: Started scap sync-world: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] * 19:24 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:24 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating the rest of the ipv6 dns records. - jhancock@cumin2002" * 19:24 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating the rest of the ipv6 dns records. - jhancock@cumin2002" * 19:14 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 19:05 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wdqs1014.eqiad.wmnet with OS bookworm * 19:04 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.noop (exit_code=99) * 19:04 cwilliams@cumin1003: START - Cookbook sre.mysql.noop * 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1014.eqiad.wmnet with reason: host reimage * 18:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1014.eqiad.wmnet with reason: host reimage * 18:30 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2008.codfw.wmnet, repooling source-only afterwards * 18:28 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 19s) * 18:28 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1014 * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1014 * 18:22 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1014 * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1014.eqiad.wmnet 188.32.64.10.in-addr.arpa 8.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:22 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1014.eqiad.wmnet 188.32.64.10.in-addr.arpa 8.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1014 - bking@cumin2003" * 18:21 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1014 - bking@cumin2003" * 18:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2215: Maintenance * 18:18 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 18:15 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:15 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 18:06 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:06 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 18:05 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 18:04 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2052: codfw rack B8 re-pool after maintenance * 17:54 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 17:54 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:54 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 17:32 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2215: Maintenance * 17:29 cmooney@dns3003: END - running authdns-update * 17:27 cmooney@dns3003: START - running authdns-update * 17:23 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 17:22 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:18 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool es2052: codfw rack B8 re-pool after maintenance * 17:18 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2189: codfw rack B8 re-pool after maintenance * 17:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2215.codfw.wmnet with reason: Maintenance * 17:17 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 17:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2215 [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95126 and previous config saved to /var/cache/conftool/dbconfig/20260723-170903-cwilliams.json * 17:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2191 to x1 primary [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95125 and previous config saved to /var/cache/conftool/dbconfig/20260723-170612-cwilliams.json * 17:05 cezmunsta: Starting x1 codfw failover from db2215 to db2191 - [[phab:T432986|T432986]] * 16:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2191 with weight 0 [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95123 and previous config saved to /var/cache/conftool/dbconfig/20260723-165831-cwilliams.json * 16:58 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 16 hosts with reason: Primary switchover x1 [[phab:T432986|T432986]] * 16:36 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 138128 * 16:35 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 138128 * 16:33 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2189: codfw rack B8 re-pool after maintenance * 16:33 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2164: codfw rack B8 re-pool after maintenance * 16:28 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1072.eqiad.wmnet * 16:27 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1072.eqiad.wmnet with OS trixie * 16:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2249: Maintenance * 16:06 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker1072.eqiad.wmnet with reason: host reimage * 16:06 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1072.eqiad.wmnet with reason: host reimage * 15:50 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1072 * 15:50 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1072 * 15:49 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1072.eqiad.wmnet with OS trixie * 15:48 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2164: codfw rack B8 re-pool after maintenance * 15:48 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] (duration: 06m 37s) * 15:48 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2163: codfw rack B8 re-pool after maintenance * 15:45 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 15:43 musikanimal@deploy1003: musikanimal: Continuing with deployment * 15:43 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:41 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] * 15:36 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1072.eqiad.wmnet * 15:35 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1072.eqiad.wmnet * 15:35 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1072.eqiad.wmnet * 15:34 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:34 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push any outstanding updates - cmooney@cumin1003" * 15:34 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push any outstanding updates - cmooney@cumin1003" * 15:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db2249: Maintenance * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 15:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:26 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:21 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:21 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 15:21 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:21 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 15:20 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 15:19 cmooney@dns2004: END - running authdns-update * 15:17 cmooney@dns2004: START - running authdns-update * 15:14 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns2004.wikimedia.org * 15:12 brouberol@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 15:12 brouberol@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 15:12 klausman@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ml-serve1001.eqiad.wmnet with OS trixie * 15:11 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1071.eqiad.wmnet * 15:11 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1071.eqiad.wmnet with OS trixie * 15:10 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wdqs2008.codfw.wmnet with OS bookworm * 15:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2249.codfw.wmnet with reason: Maintenance * 15:08 brouberol@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 15:08 brouberol@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 15:08 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 15:08 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 15:06 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns1004.wikimedia.org * 15:02 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2002.codfw.wmnet * 15:02 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2002.codfw.wmnet * 15:02 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2163: codfw rack B8 re-pool after maintenance * 15:01 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 15:01 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 14:59 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test2001.codfw.wmnet * 14:57 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test2001.codfw.wmnet * 14:56 ryankemper: [WDQS] [[phab:T430880|T430880]] Reimaged `wdqs2016` to Bookworm, xferred scholarly_articles from `wdqs2024`, validated updater/readiness/federation, and repooled * 14:51 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2016.codfw.wmnet * 14:51 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1001.eqiad.wmnet with reason: host reimage * 14:48 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1071.eqiad.wmnet with reason: host reimage * 14:47 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1001.eqiad.wmnet with reason: host reimage * 14:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2008.codfw.wmnet with reason: host reimage * 14:43 topranks: reboot lsw1-b8-codw to upgrade JunOS [[phab:T430929|T430929]] * 14:41 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1071.eqiad.wmnet with reason: host reimage * 14:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2008.codfw.wmnet with reason: host reimage * 14:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2231: Maintenance * 14:30 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1001.eqiad.wmnet with OS trixie * 14:25 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2002.codfw.wmnet * 14:23 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1071 * 14:23 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1071 * 14:23 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 14:22 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore scholarly data after Bookworm reimage) xfer scholarly_articles from wdqs2024.codfw.wmnet -> wdqs2016.codfw.wmnet, repooling source-only afterwards * 14:22 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2052: codfw rack B8 depool for maintenance * 14:21 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool es2052: codfw rack B8 depool for maintenance * 14:21 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2249: codfw rack B8 depool for maintenance * 14:21 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1071 * 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1071.eqiad.wmnet 166.48.64.10.in-addr.arpa 6.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:21 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1071.eqiad.wmnet 166.48.64.10.in-addr.arpa 6.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1071 - jiji@cumin1003" * 14:21 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1071 - jiji@cumin1003" * 14:21 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2249: codfw rack B8 depool for maintenance * 14:21 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2189: codfw rack B8 depool for maintenance * 14:20 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2002.codfw.wmnet * 14:20 cmooney@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2050.codfw.wmnet * 14:20 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2189: codfw rack B8 depool for maintenance * 14:20 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2164: codfw rack B8 depool for maintenance * 14:20 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2164: codfw rack B8 depool for maintenance * 14:19 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2163: codfw rack B8 depool for maintenance * 14:19 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:19 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1014 * 14:19 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2163: codfw rack B8 depool for maintenance * 14:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1014.eqiad.wmnet with OS bookworm * 14:17 cmooney@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2050.codfw.wmnet * 14:16 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 14:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2008 * 14:14 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2008 * 14:14 cmooney@cumin1003: conftool action : set/pooled=no; selector: name=dns2004.wikimedia.org * 14:14 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2008.codfw.wmnet with OS bookworm * 14:13 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 14:12 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:10 topranks: depool dns2004 before lsw1-b8-codfw switch maintenance [[phab:T430929|T430929]] * 14:10 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b8-codfw,lsw1-b8-codfw IPv6,lsw1-b8-codfw.mgmt,ssw1-a[1,8]-codfw with reason: lsw1-b8-codfw JunOS upgrade * 14:07 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 30 hosts with reason: lsw1-b8-codfw JunOS upgrade * 14:06 elukey: upload python3-docker-report 0.0.19 to apt.wikimedia.org for bookworm and trixie * 13:59 jiji@cumin1003: START - Cookbook sre.dns.netbox * 13:58 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1071 * 13:57 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:57 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:57 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:57 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1071.eqiad.wmnet with OS trixie * 13:55 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1071.eqiad.wmnet * 13:55 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1071.eqiad.wmnet * 13:55 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1071.eqiad.wmnet * 13:53 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:52 logmsgbot: kharlan Deployed security patch for [[phab:T432948|T432948]] * 13:51 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:51 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 13:51 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db2231: Maintenance * 13:50 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:50 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:50 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:50 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:49 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:49 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:49 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2231 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95097 and previous config saved to /var/cache/conftool/dbconfig/20260723-134436-cwilliams.json * 13:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2231.codfw.wmnet with reason: Maintenance * 13:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:39 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:38 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 13:38 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] (duration: 09m 07s) * 13:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1037 hosts * 13:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2196: Maintenance * 13:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:33 kharlan@deploy1003: kharlan, emc-wmf: Continuing with deployment * 13:31 kharlan@deploy1003: kharlan, emc-wmf: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:30 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:28 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] * 13:17 hashar@deploy1003: Finished deploy [integration/docroot@2199146]: build: License GPL2.0+ / updating npm dependencies (duration: 00m 14s) * 13:17 hashar@deploy1003: Started deploy [integration/docroot@2199146]: build: License GPL2.0+ / updating npm dependencies * 13:14 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service * 13:07 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 12:58 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2207: Repooling * 12:49 root@cumin1003: START - Cookbook sre.mysql.pool pool db2196: Maintenance * 12:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2196 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95087 and previous config saved to /var/cache/conftool/dbconfig/20260723-123952-cwilliams.json * 12:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2196.codfw.wmnet with reason: Maintenance * 12:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2191: Maintenance * 12:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:13 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:13 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: Repooling * 12:12 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2207: Repooling * 12:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: Repooling * 11:56 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2235.codfw.wmnet with OS trixie * 11:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db2191: Maintenance * 11:46 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1070.eqiad.wmnet * 11:46 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1070.eqiad.wmnet * 11:46 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1070.eqiad.wmnet * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2191 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95080 and previous config saved to /var/cache/conftool/dbconfig/20260723-114308-cwilliams.json * 11:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2191.codfw.wmnet with reason: Maintenance * 11:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2186: Maintenance * 11:35 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 46375 * 11:34 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 46375 * 11:33 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2235.codfw.wmnet with reason: host reimage * 11:28 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2235.codfw.wmnet with reason: host reimage * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c7-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c7-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c6-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c6-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c5-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c5-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c4-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c4-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c3-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c3-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c2-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c2-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d7-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d7-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d4-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d4-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d3-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d2-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d2-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d8-eqiad * 11:23 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d8-eqiad * 11:23 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d1-eqiad * 11:23 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d1-eqiad * 11:12 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db2235.codfw.wmnet with OS trixie * 11:11 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:11 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[2160,2235].codfw.wmnet with reason: Upgrading * 10:56 root@cumin1003: START - Cookbook sre.mysql.pool pool db2186: Maintenance * 10:54 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1037: testing * 10:53 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1037: testing * 10:53 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1037: testing * 10:53 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1037: testing * 10:52 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: testing * 10:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2186 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95072 and previous config saved to /var/cache/conftool/dbconfig/20260723-104956-cwilliams.json * 10:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2186.codfw.wmnet with reason: Maintenance * 10:43 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1054: testing * 10:41 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1070.eqiad.wmnet with OS trixie * 10:30 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1037 hosts * 10:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 10:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 10:20 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1070.eqiad.wmnet with reason: host reimage * 10:16 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1070.eqiad.wmnet with reason: host reimage * 10:06 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1038: testing * 10:05 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1038: testing * 10:05 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1038: testing * 10:02 marostegui@dns1004: END - running authdns-update * 10:00 marostegui@dns1004: START - running authdns-update * 09:58 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: testing * 09:57 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1054: testing * 09:57 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1070 * 09:57 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1070 * 09:57 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1054: testing * 09:57 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1054: testing * 09:56 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1055: testing * 09:56 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1070 * 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1070.eqiad.wmnet 165.48.64.10.in-addr.arpa 5.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:56 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1070.eqiad.wmnet 165.48.64.10.in-addr.arpa 5.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1070 - jiji@cumin1003" * 09:56 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1070 - jiji@cumin1003" * 09:47 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for 1035 hosts * 09:45 jiji@cumin1003: START - Cookbook sre.dns.netbox * 09:42 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1070 * 09:42 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1070.eqiad.wmnet with OS trixie * 09:42 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1070.eqiad.wmnet * 09:41 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1070.eqiad.wmnet * 09:41 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1070.eqiad.wmnet * 09:27 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es2051: testing * 09:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: testing * 09:12 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es2051: testing * 09:11 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1055: testing * 09:09 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:09 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1055: testing * 09:09 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1055: testing * 08:50 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:50 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:50 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 08:49 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 08:49 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 08:49 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:46 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1069.eqiad.wmnet * 08:46 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1069.eqiad.wmnet * 08:46 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1069.eqiad.wmnet * 08:39 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 08:38 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 08:38 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 08:37 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 08:35 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:10 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1069.eqiad.wmnet with OS trixie * 07:49 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1069.eqiad.wmnet with reason: host reimage * 07:45 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1069.eqiad.wmnet with reason: host reimage * 07:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1035 hosts * 07:33 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 07:32 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 07:29 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1069 * 07:29 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1069 * 07:26 jiji@deploy1003: Finished scap sync-world: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules (duration: 06m 01s) * 07:25 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1069 * 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1069.eqiad.wmnet 164.48.64.10.in-addr.arpa 4.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:25 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1069.eqiad.wmnet 164.48.64.10.in-addr.arpa 4.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1069 - jiji@cumin1003" * 07:25 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1069 - jiji@cumin1003" * 07:25 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 07:25 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 07:24 jiji@deploy1003: jiji: Continuing with deployment * 07:22 jiji@deploy1003: jiji: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:21 jiji@deploy1003: Started scap sync-world: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules * 07:21 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1031.eqiad.wmnet,service=s7 * 07:20 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1031.eqiad.wmnet,service=s2 * 07:20 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1031.eqiad.wmnet,service=s7 * 07:20 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1031.eqiad.wmnet,service=s2 * 07:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts * 07:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts * 07:19 jiji@cumin1003: START - Cookbook sre.dns.netbox * 07:19 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1069 * 07:19 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1069.eqiad.wmnet with OS trixie * 07:19 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1069.eqiad.wmnet * 07:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 45 hosts * 07:17 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1069.eqiad.wmnet * 07:17 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1069.eqiad.wmnet * 07:14 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 45 hosts * 07:13 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 06:16 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs2007 after successful Bookworm reimage, data transfer, and postflight validation; wdqs1013 also passed postflights and is enabled in conftool, but remains out of IPVS pending a rolling pybal restart to clear its stale pre-VLAN-move address * 05:58 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2007.codfw.wmnet * 05:58 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1013.eqiad.wmnet * 05:54 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore scholarly data after Bookworm reimage) xfer scholarly_articles from wdqs2024.codfw.wmnet -> wdqs2016.codfw.wmnet, repooling source-only afterwards == 2026-07-22 == * 23:34 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Apply upgrade to JVM17 - eevans@cumin1003 * 23:14 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Apply upgrade to JVM17 - eevans@cumin1003 * 22:06 ryankemper: [WDQS] Added requestctl per-IP ratelimit `wdqs_heavy_sparql_bots_jul_2026_ratelimit` (chronic heavy-query bot tier driving deadlock-remediation restarts); pruned superseded `wdqs_2026_05_11_worobot` * 21:51 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] (duration: 11m 52s) * 21:44 sbassett@deploy1003: sbassett: Continuing with deployment * 21:43 sbassett@deploy1003: sbassett: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:39 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] * 20:38 dani@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] (duration: 32m 51s) * 20:38 ryankemper: [WDQS] Pruned obsolete requestctl action+pattern `wdqs_20260715_p2003_ring_ja3n` (actor rotated JA3Ns; rule inert) * 20:26 dani@deploy1003: dani, vadymts1: Continuing with deployment * 20:24 dani@deploy1003: dani, vadymts1: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:14 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2244: Testing * 20:06 dani@deploy1003: Started scap sync-world: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] * 19:56 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1013.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2016.codfw.wmnet with OS bookworm * 19:40 mutante: gerrit - one more service restart is needed - restarting * 19:29 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2244: Testing * 19:27 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2244: Testing * 19:27 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2244: Testing * 19:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2016.codfw.wmnet with reason: host reimage * 19:15 dancy@deploy1003: Finished deploy [zuul/deploy@d92e238]: Freshening Zuul installation (duration: 00m 15s) * 19:14 dancy@deploy1003: Started deploy [zuul/deploy@d92e238]: Freshening Zuul installation * 19:11 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2016.codfw.wmnet with reason: host reimage * 18:54 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1013.eqiad.wmnet, repooling source-only afterwards * 18:52 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2016 * 18:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2016 * 18:51 dancy@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 18:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2016.codfw.wmnet with OS bookworm * 18:39 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 11s) * 18:39 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 18:36 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 18:30 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417] (thin): Regular analytics weekly train THIN [analytics/refinery@2a25417d] (duration: 02m 09s) * 18:28 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417] (thin): Regular analytics weekly train THIN [analytics/refinery@2a25417d] * 18:28 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417]: Regular analytics weekly train [analytics/refinery@2a25417d] (duration: 04m 31s) * 18:27 dduvall: deploying https://gerrit.wikimedia.org/r/c/integration/config/+/1314025 (4 jobs updated) * 18:23 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417]: Regular analytics weekly train [analytics/refinery@2a25417d] * 18:22 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@2a25417d] (duration: 01m 59s) * 18:20 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@2a25417d] * 17:56 Raine: deployment server switchover => deploy1003 is primary now * 17:55 kamila@deploy1003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 28m 24s) * 17:54 mutante: restarting gerrit for maintenance * 17:29 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1013.eqiad.wmnet with OS bookworm * 17:27 kamila@deploy1003: Started scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] * 17:20 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] (duration: 22m 50s) * 17:12 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1023.eqiad.wmnet -> wdqs1024.eqiad.wmnet, repooling source-only afterwards * 17:04 Raine: point deployment.eqiad.wmnet to deploy1003 * 17:04 kamila@dns7001: END - running authdns-update * 17:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1013.eqiad.wmnet with reason: host reimage * 17:02 kamila@dns7001: START - running authdns-update * 17:01 kamila@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on releases2003.codfw.wmnet,releases1003.eqiad.wmnet with reason: Deployment server switchover * 17:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1013.eqiad.wmnet with reason: host reimage * 16:58 kamila@deploy2003: Locking from deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] * 16:57 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] (duration: 02m 33s) * 16:55 kamila@deploy2003: Locking from deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] * 16:55 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2003 - [[phab:T240266|T240266]] (duration: 00m 11s) * 16:54 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2003 - [[phab:T240266|T240266]] * 16:40 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1013 * 16:40 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1013 * 16:39 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1013 * 16:39 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1013.eqiad.wmnet 105.32.64.10.in-addr.arpa 5.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:39 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1013.eqiad.wmnet 105.32.64.10.in-addr.arpa 5.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:39 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:39 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1013 - bking@cumin2003" * 16:39 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1013 - bking@cumin2003" * 16:34 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:34 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1013 * 16:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1013.eqiad.wmnet with OS bookworm * 16:28 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1023.eqiad.wmnet -> wdqs1024.eqiad.wmnet, repooling source-only afterwards * 16:27 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-scholarly,name=eqiad * 16:27 eevans@deploy2003: helmfile [eqiad] DONE helmfile.d/services/linked-artifacts: apply * 16:26 eevans@deploy2003: helmfile [eqiad] START helmfile.d/services/linked-artifacts: apply * 16:26 eevans@deploy2003: helmfile [codfw] DONE helmfile.d/services/linked-artifacts: apply * 16:26 eevans@deploy2003: helmfile [codfw] START helmfile.d/services/linked-artifacts: apply * 16:25 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 29s) * 16:25 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 16:24 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 16:21 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 16:18 eevans@deploy2003: helmfile [codfw] DONE helmfile.d/services/linked-artifacts: apply * 16:18 eevans@deploy2003: helmfile [codfw] START helmfile.d/services/linked-artifacts: apply * 16:08 eevans@deploy2003: helmfile [staging] DONE helmfile.d/services/linked-artifacts: apply * 16:07 eevans@deploy2003: helmfile [staging] START helmfile.d/services/linked-artifacts: apply * 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 16:01 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 15:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1024.eqiad.wmnet with OS bookworm * 15:49 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] (duration: 00m 10s) * 15:49 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] * 15:48 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] (duration: 00m 15s) * 15:48 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] * 15:47 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] (duration: 00m 10s) * 15:47 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] * 15:46 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:42 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:42 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:40 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:37 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 15:37 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:36 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:36 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:36 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 15:33 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 15:33 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:31 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:28 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 15:27 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:27 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1024.eqiad.wmnet with reason: host reimage * 15:23 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:23 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:23 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:20 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1068.eqiad.wmnet * 15:20 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1068.eqiad.wmnet * 15:20 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1068.eqiad.wmnet * 15:20 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 15:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1024.eqiad.wmnet with reason: host reimage * 15:11 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:55 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 14:52 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wdqs1024.eqiad.wmnet with OS bookworm * 14:50 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] (duration: 00m 09s) * 14:50 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] * 14:49 jiji@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 14:49 jiji@deploy2003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 14:49 jiji@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 14:48 jiji@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 14:45 ecarg@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:45 ecarg@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:44 ecarg@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:44 ecarg@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:43 ecarg@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:43 ecarg@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:41 sukhe: ipvsadm --delete-service --tcp-service 10.2.1.55:8087: lvs2014 and lvs2013 * 14:39 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:39 sukhe: ipvsadm --delete-service --tcp-service 10.2.2.55:8087: [[phab:T432445|T432445]] * 14:38 ecarg@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:38 ecarg@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:37 ecarg@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:37 ecarg@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:36 ecarg@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:34 ecarg@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts datahubsearch1001.eqiad.wmnet * 14:32 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:32 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 14:31 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 14:31 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 14:28 sukhe: sudo cumin 'A:lvs-low-traffic-codfw' 'systemctl restart pybal': lvs2013 * 14:26 sukhe: sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal': lvs2014 * 14:26 sukhe: sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal' * 14:24 sukhe: restart pybal on lvs1019 * 14:24 sukhe: restart pybal on lvs1020 * 14:19 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for 1036 hosts * 14:17 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:04 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] (duration: 09m 28s) * 13:59 kharlan@deploy2003: dreamyjazz, kharlan: Continuing with deployment * 13:58 bking@cumin2003: START - Cookbook sre.hosts.decommission for hosts datahubsearch1001.eqiad.wmnet * 13:57 kharlan@deploy2003: dreamyjazz, kharlan: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:55 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] * 13:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts datahubsearch[1002-1003].eqiad.wmnet * 13:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:53 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch[1002-1003].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 13:52 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch[1002-1003].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 13:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 13:42 stran@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] (duration: 07m 30s) * 13:42 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:40 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: Test * 13:38 stran@deploy2003: dragoniez, stran: Continuing with deployment * 13:37 stran@deploy2003: dragoniez, stran: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:35 bking@cumin2003: START - Cookbook sre.hosts.decommission for hosts datahubsearch[1002-1003].eqiad.wmnet * 13:35 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024'] * 13:35 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 13:35 stran@deploy2003: Started scap sync-world: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] * 13:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 13:28 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024'] * 13:26 sukhe@dns1004: END - running authdns-update * 13:25 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 13:25 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 13:24 sukhe@dns1004: START - running authdns-update * 13:22 stran@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] (duration: 08m 20s) * 13:21 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 13:20 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 13:19 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 13:19 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 13:19 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 13:18 stran@deploy2003: stran: Continuing with deployment * 13:16 stran@deploy2003: stran: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:14 stran@deploy2003: Started scap sync-world: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] * 13:13 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 13:13 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 13:11 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 13:11 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 13:08 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 12:55 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool es2051: Test * 12:55 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: Test * 12:54 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool es2051: Test * 12:43 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1036 hosts * 12:41 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 12:40 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1048.eqiad.wmnet * 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 12:39 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 12:38 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 12:37 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 12:37 brouberol@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 12:36 brouberol@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 12:36 brouberol@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 12:36 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 12:36 elukey@cumin1003: DONE (PASS) - Cookbook sre.puppet.renew-cert (exit_code=0) for crm2001.codfw.wmnet: Renew puppet certificate - elukey@cumin1003 * 12:35 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:35 brouberol@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 12:34 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 12:31 brouberol@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 12:30 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1048.eqiad.wmnet * 12:30 brouberol@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 12:28 brouberol@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 12:27 brouberol@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 12:20 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1068.eqiad.wmnet with OS trixie * 12:01 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] (duration: 13m 19s) * 11:58 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1068.eqiad.wmnet with reason: host reimage * 11:52 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1068.eqiad.wmnet with reason: host reimage * 11:51 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 11:49 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:47 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] * 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2252: Security updates * 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:43 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 11:42 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2252: Security updates * 11:42 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply * 11:40 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply * 11:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 11:37 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2252.codfw.wmnet with OS trixie * 11:34 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1068 * 11:34 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1068 * 11:26 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1068 * 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1068.eqiad.wmnet 46.48.64.10.in-addr.arpa 6.4.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:26 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1068.eqiad.wmnet 46.48.64.10.in-addr.arpa 6.4.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1068 - jiji@cumin1003" * 11:26 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1068 - jiji@cumin1003" * 11:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2252.codfw.wmnet with reason: host reimage * 11:17 jiji@cumin1003: START - Cookbook sre.dns.netbox * 11:17 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1068 * 11:17 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1068.eqiad.wmnet with OS trixie * 11:17 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2252.codfw.wmnet with reason: host reimage * 11:15 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1068.eqiad.wmnet * 11:15 mvolz@deploy2003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:15 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1068.eqiad.wmnet * 11:15 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1068.eqiad.wmnet * 11:14 mvolz@deploy2003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:13 mvolz@deploy2003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:13 mvolz@deploy2003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:12 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] (duration: 11m 05s) * 11:11 mvolz@deploy2003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:10 mvolz@deploy2003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:07 dreamyjazz@deploy2003: dreamyjazz, kharlan: Continuing with deployment * 11:03 dreamyjazz@deploy2003: dreamyjazz, kharlan: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:03 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2252.codfw.wmnet with OS trixie * 11:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2252: Upgrading db2252.codfw.wmnet * 11:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:02 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 11:02 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2252: Upgrading db2252.codfw.wmnet * 11:02 cwilliams@cumin1003: dbmaint on ms3@codfw [[phab:T432321|T432321]] * 11:01 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 11:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db1153.eqiad.wmnet with reason: Security updates * 11:01 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] * 11:00 fnegri@deploy2003: helmfile [eqiad] DONE helmfile.d/services/toolhub: apply * 10:58 fnegri@deploy2003: helmfile [eqiad] START helmfile.d/services/toolhub: apply * 10:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1151: Security updates * 10:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:57 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1151: Security updates * 10:55 fnegri@deploy2003: helmfile [codfw] DONE helmfile.d/services/toolhub: apply * 10:54 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] (duration: 08m 38s) * 10:53 fnegri@deploy2003: helmfile [codfw] START helmfile.d/services/toolhub: apply * 10:53 fnegri@deploy2003: helmfile [staging] DONE helmfile.d/services/toolhub: apply * 10:52 fnegri@deploy2003: helmfile [staging] START helmfile.d/services/toolhub: apply * 10:50 zabe@deploy2003: zabe: Continuing with deployment * 10:47 zabe@deploy2003: zabe: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:45 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] * 10:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1151: Security updates * 10:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:42 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:42 root@cumin1003: START - Cookbook sre.mysql.depool depool db1151: Security updates * 10:38 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] (duration: 12m 47s) * 10:34 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 10:34 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 10:33 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2253.codfw.wmnet with OS trixie * 10:28 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:26 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] * 10:18 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2253.codfw.wmnet with reason: host reimage * 10:13 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2253.codfw.wmnet with reason: host reimage * 10:00 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2253.codfw.wmnet with OS trixie * 09:58 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db1151.eqiad.wmnet with reason: Security updates * 09:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2253: Upgrading db2253.codfw.wmnet * 09:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:57 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 09:56 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2253: Upgrading db2253.codfw.wmnet * 09:56 cwilliams@cumin1003: dbmaint on ms2@codfw [[phab:T432321|T432321]] * 09:56 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 09:36 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: UI improvement; support url shortener - oblivian@cumin1003" * 09:36 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: UI improvement; support url shortener - oblivian@cumin1003 * 09:35 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: UI improvement; support url shortener - oblivian@cumin1003 * 09:35 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: UI improvement; support url shortener - oblivian@cumin1003" * 09:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1152: Security updates * 09:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:26 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db1152: Security updates * 09:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: Security updates * 09:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:11 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:11 root@cumin1003: START - Cookbook sre.mysql.depool depool db1152: Security updates * 09:10 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1018.eqiad.wmnet with reason: Cloning * 09:09 Dreamy_Jazz: Deployed patch for [[phab:T432453|T432453]] and [[phab:T432454|T432454]] * 09:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 09:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2251.codfw.wmnet with OS trixie * 08:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2251.codfw.wmnet with reason: host reimage * 08:45 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2251.codfw.wmnet with reason: host reimage * 08:40 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] (duration: 12m 26s) * 08:38 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1030.eqiad.wmnet,service=s1 * 08:36 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 08:31 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2251.codfw.wmnet with OS trixie * 08:30 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:28 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] * 08:25 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] (duration: 07m 59s) * 08:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2251: Upgrading db2251.codfw.wmnet * 08:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:22 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 08:22 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2251: Upgrading db2251.codfw.wmnet * 08:20 urbanecm@deploy2003: urbanecm: Continuing with deployment * 08:20 cwilliams@cumin1003: dbmaint on ms1@codfw [[phab:T432321|T432321]] * 08:20 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 08:19 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:17 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] * 08:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade * 08:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade * 08:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db2251.codfw.wmnet,db1152.eqiad.wmnet with reason: OS upgrade * 08:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade * 08:13 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade * 08:11 Dreamy_Jazz: Created cusi_signal, cusi_case, and cusi_user on ukwiki and enwikivoyage in extension1 * 08:11 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade * 08:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade * 08:04 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1030.eqiad.wmnet,service=s1 * 08:04 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1030.eqiad.wmnet,service=s1 * 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply * 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply * 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply * 07:51 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply * 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 07:47 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 07:47 phuedx: End of UTC morning backport window * 07:43 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 07:43 phuedx@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] (duration: 13m 44s) * 07:43 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 07:39 phuedx@deploy2003: phuedx: Continuing with deployment * 07:31 phuedx@deploy2003: phuedx: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:29 phuedx@deploy2003: Started scap sync-world: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] * 07:24 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Turnilo import support - oblivian@cumin1003" * 07:24 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import support - oblivian@cumin1003 * 07:23 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import support - oblivian@cumin1003 * 07:23 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Turnilo import support - oblivian@cumin1003" * 06:42 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs2020 after successful Bookworm reimage, data transfer, and postflight validation * 06:42 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2020.codfw.wmnet * 05:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Managing sanitization for wikis bolwiki in section s5 * 05:25 marostegui@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis bolwiki in section s5 * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 41s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-21 == * 22:50 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2019.codfw.wmnet -> wdqs2020.codfw.wmnet, repooling source-only afterwards * 22:47 cwhite: force reboot arclamp2001 - appears to have run out of memory and gone unresponsive * 22:24 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 01m 26s) * 22:24 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 22:23 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 22:22 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024'] * 22:11 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 21:54 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs1024'] * 21:54 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 21:53 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs1024'] * 21:53 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 21:49 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2019.codfw.wmnet -> wdqs2020.codfw.wmnet, repooling source-only afterwards * 20:57 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] (duration: 09m 10s) * 20:55 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1024.eqiad.wmnet with OS bookworm * 20:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2020.codfw.wmnet with OS bookworm * 20:52 krinkle@deploy2003: krinkle: Continuing with deployment * 20:49 krinkle@deploy2003: krinkle: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:47 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] * 20:45 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] (duration: 05m 42s) * 20:44 krinkle@deploy2003: krinkle: Rolling back deployment * 20:41 krinkle@deploy2003: krinkle: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:39 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] * 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2003.codfw.wmnet * 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1003.eqiad.wmnet * 20:33 dani@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] (duration: 11m 15s) * 20:33 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2003.codfw.wmnet * 20:33 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1003.eqiad.wmnet * 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2020.codfw.wmnet with reason: host reimage * 20:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1002.eqiad.wmnet * 20:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2002.codfw.wmnet * 20:30 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 20:29 dani@deploy2003: dani: Continuing with deployment * 20:29 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2020.codfw.wmnet with reason: host reimage * 20:26 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1002.eqiad.wmnet * 20:26 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2002.codfw.wmnet * 20:24 dani@deploy2003: dani: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2001.codfw.wmnet * 20:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1001.eqiad.wmnet * 20:22 dani@deploy2003: Started scap sync-world: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] * 20:22 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 20:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2001.codfw.wmnet * 20:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1001.eqiad.wmnet * 20:14 sbisson@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] (duration: 09m 01s) * 20:11 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2020.codfw.wmnet with OS bookworm * 20:10 sbisson@deploy2003: sbisson: Continuing with deployment * 20:07 sbisson@deploy2003: sbisson: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:05 sbisson@deploy2003: Started scap sync-world: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] * 20:03 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] (duration: 07m 04s) * 20:01 mutante: Gerrit - tomorrow a new SSH host key will appear - it will be {{Gerrit|ed25519}} and has already been added to wmf-laptop. you can verify it here: https://wikitech.wikimedia.org/wiki/Help:SSH_Fingerprints/gerrit.wikimedia.org:29418 ([[phab:T240266|T240266]]) * 19:59 zabe@deploy2003: zabe: Continuing with deployment * 19:58 zabe@deploy2003: zabe: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:56 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] * 19:52 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] (duration: 07m 25s) * 19:48 zabe@deploy2003: zabe: Continuing with deployment * 19:47 zabe@deploy2003: zabe: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:45 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] * 19:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 19:32 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024'] * 19:27 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 19:26 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024'] * 19:26 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 19:24 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024'] * 19:12 ryankemper: [wdqs] [[phab:T430880|T430880]] Repooled `wdqs-scholarly` discovery in `eqiad` after validating `wdqs1023` end-to-end; `wdqs1024` remains disabled pending reimage recovery * 19:11 ryankemper: [wdqs] [[phab:T430880|T430880]] Repooled wdqs1012.eqiad.wmnet after successful Bookworm reimage, data transfer, service checks, readiness probe, and cross-graph federation query validation * 19:10 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 19:10 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1012.eqiad.wmnet * 19:08 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 18:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for deploy1003.eqiad.wmnet * 18:57 kamila@cumin1003: START - Cookbook sre.hosts.remove-downtime for deploy1003.eqiad.wmnet * 18:37 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:37 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding urldownloader service IPs - sukhe@cumin1003" * 18:37 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding urldownloader service IPs - sukhe@cumin1003" * 18:32 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 18:32 dancy@deploy2003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 18:30 sukhe@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 18:27 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 18:24 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1024.eqiad.wmnet with OS bookworm * 18:20 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host deploy1003.eqiad.wmnet with OS bookworm * 18:09 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deploy1003 reimage (duration: 121m 16s) * 18:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 18:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1180: Security updates * 17:55 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wcqs2003.codfw.wmnet * 17:48 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wcqs2003.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1155.eqiad.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1155.eqiad.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2224.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2224.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2217.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2217.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2193.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2193.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2180.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2180.codfw.wmnet * 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1168.eqiad.wmnet * 17:36 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1168.eqiad.wmnet * 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2169.codfw.wmnet * 17:36 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2169.codfw.wmnet * 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1165.eqiad.wmnet * 17:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1165.eqiad.wmnet * 17:35 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2158.codfw.wmnet * 17:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2158.codfw.wmnet * 17:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wcqs1003.eqiad.wmnet * 17:17 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1180: Security updates * 17:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1180.eqiad.wmnet * 17:16 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1180.eqiad.wmnet * 17:15 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp3073.* * 17:13 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wcqs1003.eqiad.wmnet * 17:11 brett@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp3073.esams.wmnet with OS trixie * 17:11 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 17:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1024 * 17:04 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1024 * 17:03 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 17:00 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 16:59 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2242: codfw rack B7 depool for maintenance * 16:59 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 16:43 brett@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp3073.esams.wmnet with reason: host reimage * 16:42 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 16:39 brett@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cp3073.esams.wmnet with reason: host reimage * 16:32 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on deploy1003.eqiad.wmnet with reason: host reimage * 16:27 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on deploy1003.eqiad.wmnet with reason: host reimage * 16:14 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2242: codfw rack B7 depool for maintenance * 16:14 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: codfw rack B7 depool for maintenance * 16:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1023.eqiad.wmnet with OS bookworm * 16:13 brett@cumin2002: START - Cookbook sre.hosts.reimage for host cp3073.esams.wmnet with OS trixie * 16:08 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host deploy1003.eqiad.wmnet with OS bookworm * 16:08 kamila@deploy2003: Locking from deployment [MediaWiki]: deploy1003 reimage * 16:03 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1012.eqiad.wmnet with OS bookworm * 15:48 inflatador: bking@apt1002 `sudo reprepro copy bookworm-wikimedia bullseye-wikimedia jvmquake` [[phab:T430880|T430880]] * 15:39 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp3073.* * 15:39 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 15:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:34 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1311536{{!}}Set $wgMathInternalRestbaseURL explicitly (take 2) (T349582)]] * 15:29 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:29 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2228: codfw rack B7 depool for maintenance * 15:29 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2229: codfw rack B7 depool for maintenance * 15:27 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1180: Security update * 15:25 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Security update * 15:21 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 15:21 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 15:19 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 24s) * 15:19 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:14 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 15:14 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db1180: Security update * 15:13 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-eqiad * 14:48 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-eqiad * 14:44 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2229: codfw rack B7 depool for maintenance * 14:44 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc2017: codfw rack B7 depool for maintenance * 14:44 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:43 cmooney@cumin2003: START - Cookbook sre.mysql.parsercache * 14:43 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool pc2017: codfw rack B7 depool for maintenance * 14:43 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2003.codfw.wmnet * 14:43 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2003.codfw.wmnet * 14:42 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2009.codfw.wmnet * 14:42 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2009.codfw.wmnet * 14:41 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:41 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:40 cmooney@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 29 hosts * 14:40 cmooney@cumin1003: START - Cookbook sre.hosts.remove-downtime for 29 hosts * 14:35 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 14:34 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 14:32 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] (duration: 07m 56s) * 14:29 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ssw1-a[1,8]-codfw with reason: lsw1-b7-codfw JunOS upgrade * 14:28 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 14:28 elukey@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 14:26 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:24 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] * 14:23 topranks: reboot lsw1-b7-codfw to upgrade JunOS (affects all hosts in rack) [[phab:T430928|T430928]] * 14:18 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2003.codfw.wmnet * 14:14 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2009.codfw.wmnet * 14:14 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Security update * 14:13 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2242: codfw rack B7 depool for maintenance * 14:13 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2242: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2228: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2228: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2229: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2229: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc2017: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.parsercache * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool pc2017: codfw rack B7 depool for maintenance * 14:08 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2003.codfw.wmnet * 14:07 cmooney@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on aux-k8s-etcd2004.codfw.wmnet,ml-etcd2001.codfw.wmnet with reason: lsw1-b7-codfw JunOS upgrade * 14:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95005 and previous config saved to /var/cache/conftool/dbconfig/20260721-140620-cwilliams.json * 14:05 cmooney@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2049.codfw.wmnet * 14:05 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-scholarly,name=eqiad * 14:04 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2009.codfw.wmnet * 14:04 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply * 14:04 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply * 14:03 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:03 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2001.codfw.wmnet * 14:03 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2001.codfw.wmnet * 14:02 cmooney@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2049.codfw.wmnet * 14:00 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:00 Dreamy_Jazz: Created cusi_case, cusi_signal, and cusi_user on svwiki, dewiki, jawiki, eswiki * 13:59 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b7-codfw,lsw1-b7-codfw IPv6,lsw1-b7-codfw.mgmt,ssw1-a[1,8]-codfw.mgmt with reason: lsw1-b7-codfw JunOS upgrade * 13:57 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs1023.eqiad.wmnet, repooling source-only afterwards * 13:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 29 hosts with reason: lsw1-b7-codfw JunOS upgrade * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224', diff saved to https://phabricator.wikimedia.org/P95003 and previous config saved to /var/cache/conftool/dbconfig/20260721-135613-cwilliams.json * 13:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1012.eqiad.wmnet with reason: host reimage * 13:53 cmooney@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 1:00:00 on 30 hosts with reason: lsw1-b7-codfw JunOS upgrade * 13:51 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1012.eqiad.wmnet with reason: host reimage * 13:48 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 13:48 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 13:46 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224', diff saved to https://phabricator.wikimedia.org/P95001 and previous config saved to /var/cache/conftool/dbconfig/20260721-134605-cwilliams.json * 13:46 cmooney@cumin1003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti2033.codfw.wmnet * 13:46 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 13:45 elukey: move the Docker Registry's /v2/wikimedia/machinelearning.* prefix to the ml S3 backend - [[phab:T428022|T428022]] * 13:45 cmooney@cumin1003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti2033.codfw.wmnet * 13:45 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 13:43 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:40 jiji@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 13:40 jiji@deploy2003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 13:39 jiji@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 13:39 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 13:38 cmooney@cumin1003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2032.codfw.wmnet * 13:38 jiji@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 13:38 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 13:37 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2032.codfw.wmnet * 13:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95000 and previous config saved to /var/cache/conftool/dbconfig/20260721-133557-cwilliams.json * 13:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1012 * 13:33 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1012 * 13:33 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1012.eqiad.wmnet with OS bookworm * 13:30 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:30 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:28 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94999 and previous config saved to /var/cache/conftool/dbconfig/20260721-132855-cwilliams.json * 13:28 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2224.codfw.wmnet with reason: Maintenance * 13:28 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94998 and previous config saved to /var/cache/conftool/dbconfig/20260721-132826-cwilliams.json * 13:28 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 13:23 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] (duration: 07m 50s) * 13:20 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 13:18 kharlan@deploy2003: kharlan: Continuing with deployment * 13:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217', diff saved to https://phabricator.wikimedia.org/P94996 and previous config saved to /var/cache/conftool/dbconfig/20260721-131817-cwilliams.json * 13:17 brouberol@dns1004: END - running authdns-update * 13:17 kharlan@deploy2003: kharlan: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:15 brouberol@dns1004: START - running authdns-update * 13:15 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] * 13:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94995 and previous config saved to /var/cache/conftool/dbconfig/20260721-131411-cwilliams.json * 13:13 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs1023.eqiad.wmnet, repooling source-only afterwards * 13:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217', diff saved to https://phabricator.wikimedia.org/P94994 and previous config saved to /var/cache/conftool/dbconfig/20260721-130809-cwilliams.json * 13:07 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 13:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180', diff saved to https://phabricator.wikimedia.org/P94993 and previous config saved to /var/cache/conftool/dbconfig/20260721-130404-cwilliams.json * 13:03 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:03 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:02 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 13:02 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 12:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94992 and previous config saved to /var/cache/conftool/dbconfig/20260721-125801-cwilliams.json * 12:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180', diff saved to https://phabricator.wikimedia.org/P94991 and previous config saved to /var/cache/conftool/dbconfig/20260721-125356-cwilliams.json * 12:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94990 and previous config saved to /var/cache/conftool/dbconfig/20260721-125049-cwilliams.json * 12:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2217.codfw.wmnet with reason: Maintenance * 12:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94989 and previous config saved to /var/cache/conftool/dbconfig/20260721-125017-cwilliams.json * 12:48 elukey: bmc cold reboot for lvs1013 and lvs1015 - [[phab:T426180|T426180]] * 12:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94988 and previous config saved to /var/cache/conftool/dbconfig/20260721-124348-cwilliams.json * 12:40 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193', diff saved to https://phabricator.wikimedia.org/P94987 and previous config saved to /var/cache/conftool/dbconfig/20260721-124009-cwilliams.json * 12:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts * 12:33 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts * 12:33 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts * 12:32 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts * 12:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:30 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193', diff saved to https://phabricator.wikimedia.org/P94986 and previous config saved to /var/cache/conftool/dbconfig/20260721-123001-cwilliams.json * 12:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94985 and previous config saved to /var/cache/conftool/dbconfig/20260721-121953-cwilliams.json * 12:17 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs2007.codfw.wmnet with OS bookworm * 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94983 and previous config saved to /var/cache/conftool/dbconfig/20260721-121257-cwilliams.json * 12:12 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2193.codfw.wmnet with reason: Maintenance * 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94982 and previous config saved to /var/cache/conftool/dbconfig/20260721-121239-cwilliams.json * 12:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180', diff saved to https://phabricator.wikimedia.org/P94980 and previous config saved to /var/cache/conftool/dbconfig/20260721-120231-cwilliams.json * 11:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180', diff saved to https://phabricator.wikimedia.org/P94979 and previous config saved to /var/cache/conftool/dbconfig/20260721-115223-cwilliams.json * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94978 and previous config saved to /var/cache/conftool/dbconfig/20260721-114333-cwilliams.json * 11:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1180.eqiad.wmnet with reason: Maintenance * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94977 and previous config saved to /var/cache/conftool/dbconfig/20260721-114305-cwilliams.json * 11:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94976 and previous config saved to /var/cache/conftool/dbconfig/20260721-114215-cwilliams.json * 11:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94975 and previous config saved to /var/cache/conftool/dbconfig/20260721-113530-cwilliams.json * 11:35 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2180.codfw.wmnet with reason: Maintenance * 11:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94974 and previous config saved to /var/cache/conftool/dbconfig/20260721-113501-cwilliams.json * 11:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168', diff saved to https://phabricator.wikimedia.org/P94973 and previous config saved to /var/cache/conftool/dbconfig/20260721-113258-cwilliams.json * 11:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169', diff saved to https://phabricator.wikimedia.org/P94972 and previous config saved to /var/cache/conftool/dbconfig/20260721-112453-cwilliams.json * 11:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168', diff saved to https://phabricator.wikimedia.org/P94971 and previous config saved to /var/cache/conftool/dbconfig/20260721-112250-cwilliams.json * 11:21 XioNoX: put eqiad-drmrs Arelion link in service * 11:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169', diff saved to https://phabricator.wikimedia.org/P94970 and previous config saved to /var/cache/conftool/dbconfig/20260721-111446-cwilliams.json * 11:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94969 and previous config saved to /var/cache/conftool/dbconfig/20260721-111242-cwilliams.json * 11:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 11:10 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 11:07 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1093 hosts * 11:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94968 and previous config saved to /var/cache/conftool/dbconfig/20260721-110548-cwilliams.json * 11:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1168.eqiad.wmnet with reason: Maintenance * 11:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94967 and previous config saved to /var/cache/conftool/dbconfig/20260721-110520-cwilliams.json * 11:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94966 and previous config saved to /var/cache/conftool/dbconfig/20260721-110439-cwilliams.json * 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94964 and previous config saved to /var/cache/conftool/dbconfig/20260721-105632-cwilliams.json * 10:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2169.codfw.wmnet with reason: Maintenance * 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94963 and previous config saved to /var/cache/conftool/dbconfig/20260721-105603-cwilliams.json * 10:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165', diff saved to https://phabricator.wikimedia.org/P94962 and previous config saved to /var/cache/conftool/dbconfig/20260721-105512-cwilliams.json * 10:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158', diff saved to https://phabricator.wikimedia.org/P94961 and previous config saved to /var/cache/conftool/dbconfig/20260721-104555-cwilliams.json * 10:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165', diff saved to https://phabricator.wikimedia.org/P94960 and previous config saved to /var/cache/conftool/dbconfig/20260721-104504-cwilliams.json * 10:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158', diff saved to https://phabricator.wikimedia.org/P94959 and previous config saved to /var/cache/conftool/dbconfig/20260721-103547-cwilliams.json * 10:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94958 and previous config saved to /var/cache/conftool/dbconfig/20260721-103456-cwilliams.json * 10:29 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2229: Upgraded kernel * 10:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94956 and previous config saved to /var/cache/conftool/dbconfig/20260721-102757-cwilliams.json * 10:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on an-redacteddb1001.eqiad.wmnet,clouddb[1015,1025,1028].eqiad.wmnet,db1155.eqiad.wmnet with reason: Maintenance * 10:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1165.eqiad.wmnet with reason: Maintenance * 10:25 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94955 and previous config saved to /var/cache/conftool/dbconfig/20260721-102539-cwilliams.json * 10:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94954 and previous config saved to /var/cache/conftool/dbconfig/20260721-101848-cwilliams.json * 10:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2158.codfw.wmnet with reason: Maintenance * 09:43 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2229: Upgraded kernel * 09:42 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2229.codfw.wmnet * 09:42 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2229.codfw.wmnet * 09:23 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db2229.codfw.wmnet * 09:23 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2229.codfw.wmnet * 08:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2229 [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94948 and previous config saved to /var/cache/conftool/dbconfig/20260721-085724-cwilliams.json * 08:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2214 to s6 primary [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94947 and previous config saved to /var/cache/conftool/dbconfig/20260721-085442-cwilliams.json * 08:53 cezmunsta: Starting s6 codfw failover from db2229 to db2214 - [[phab:T430964|T430964]] * 08:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2214 with weight 0 [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94946 and previous config saved to /var/cache/conftool/dbconfig/20260721-084613-cwilliams.json * 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 22 hosts with reason: Primary switchover s6 [[phab:T430964|T430964]] * 08:32 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1017.eqiad.wmnet,service=s1 * 08:08 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Add subrated circuit rate to interface descriptions - CR1312476 - ayounsi@cumin1003 * 08:06 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Add subrated circuit rate to interface descriptions - CR1312476 - ayounsi@cumin1003 * 07:58 reedy@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] (duration: 12m 55s) * 07:51 reedy@deploy2003: reedy, neriah: Continuing with deployment * 07:51 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 07:51 reedy@deploy2003: reedy, neriah: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:48 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1093 hosts * 07:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm2001.wikimedia.org * 07:45 reedy@deploy2003: Started scap sync-world: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] * 07:43 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 07:42 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm2001.wikimedia.org * 07:23 elukey: upgrade libtiff6 packages on zuul* trixie hosts for security upgrades * 07:22 elukey: upgrade libtiff6 packages on Wikikube trixie workers for security upgrades * 07:14 elukey@deploy2003: helmfile [codfw] DONE helmfile.d/services/proton: sync * 07:13 elukey@deploy2003: helmfile [codfw] START helmfile.d/services/proton: sync * 07:11 elukey@deploy2003: helmfile [eqiad] DONE helmfile.d/services/proton: sync * 07:10 elukey@deploy2003: helmfile [eqiad] START helmfile.d/services/proton: sync * 07:09 elukey@deploy2003: helmfile [staging] DONE helmfile.d/services/proton: sync * 07:08 elukey@deploy2003: helmfile [staging] START helmfile.d/services/proton: sync * 06:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1023.eqiad.wmnet with reason: host reimage * 06:46 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1023.eqiad.wmnet with reason: host reimage * 06:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 05:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Haproxy-only mode support - oblivian@cumin1003" * 05:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Haproxy-only mode support - oblivian@cumin1003 * 05:42 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Haproxy-only mode support - oblivian@cumin1003 * 05:42 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Haproxy-only mode support - oblivian@cumin1003" * 05:38 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1017.eqiad.wmnet with reason: Cloning * 05:37 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1017.eqiad.wmnet,service=s1 * 05:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1029.eqiad.wmnet,service=s8 * 05:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1029.eqiad.wmnet,service=s5 * 05:32 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:30 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 05:11 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet * 05:04 aokoth@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet * 05:00 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 04:56 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 04:01 mwpresync@deploy2003: Pruned MediaWiki: 1.47.0-wmf.9 (duration: 01m 08s) * 03:41 mwpresync@deploy2003: Finished scap sync-world: testwikis to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] (duration: 36m 30s) * 03:05 mwpresync@deploy2003: Started scap sync-world: testwikis to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 03:01 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:01 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:00 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:00 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:36 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:36 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:36 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:35 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:16 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 47s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 00:56 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm == 2026-07-20 == * 23:38 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 23:07 Amir1: deleting echo notifications from 2015 on group1 wikis * 23:07 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] (duration: 14m 16s) * 23:01 ladsgroup@deploy2003: ladsgroup: Continuing with deployment * 23:00 ladsgroup@deploy2003: ladsgroup: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:53 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] * 22:46 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2007.codfw.wmnet, repooling source-only afterwards * 22:39 maryum: Deployed security fixes for several security bugs * 21:42 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 21:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2007.codfw.wmnet, repooling source-only afterwards * 21:37 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 21:37 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 21:34 sbassett: Deployed security fix for [[phab:T432424|T432424]] * 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs2020.codfw.wmnet * 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1023.eqiad.wmnet * 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1011.eqiad.wmnet * 21:32 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 17s) * 21:32 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 21:27 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 21:13 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2007.codfw.wmnet with reason: host reimage * 21:08 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-internal-main,name=codfw * 21:06 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2007.codfw.wmnet with reason: host reimage * 20:59 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service * 20:58 sukhe: pybal restart for IP changes around wdqs-main hosts * 20:57 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 20:46 ryankemper@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-internal-main,name=codfw * 20:45 ebernhardson@deploy2003: Finished deploy [search/mjolnir/deploy@d4dc3b8]: Update for opensearch 2.x compat (duration: 00m 34s) * 20:44 ebernhardson@deploy2003: Started deploy [search/mjolnir/deploy@d4dc3b8]: Update for opensearch 2.x compat * 20:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2007 * 20:44 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2007 * 20:43 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2007 * 20:43 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2007.codfw.wmnet 156.16.192.10.in-addr.arpa 6.5.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:42 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2007.codfw.wmnet 156.16.192.10.in-addr.arpa 6.5.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:42 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2007 - bking@cumin2003" * 20:41 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2007 - bking@cumin2003" * 20:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94944 and previous config saved to /var/cache/conftool/dbconfig/20260720-203333-cwilliams.json * 20:32 arlolra@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] (duration: 15m 07s) * 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2020.codfw.wmnet * 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1023.eqiad.wmnet * 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1011.eqiad.wmnet * 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs2020.codfw.wmnet * 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1023.eqiad.wmnet * 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1011.eqiad.wmnet * 20:25 arlolra@deploy2003: arlolra, cscott: Continuing with deployment * 20:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257', diff saved to https://phabricator.wikimedia.org/P94943 and previous config saved to /var/cache/conftool/dbconfig/20260720-202325-cwilliams.json * 20:21 arlolra@deploy2003: arlolra, cscott: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:17 arlolra@deploy2003: Started scap sync-world: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] * 20:13 bking@cumin2003: START - Cookbook sre.dns.netbox * 20:13 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257', diff saved to https://phabricator.wikimedia.org/P94942 and previous config saved to /var/cache/conftool/dbconfig/20260720-201318-cwilliams.json * 20:13 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 20:10 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 20:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2007 * 20:04 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2007.codfw.wmnet with OS bookworm * 20:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94941 and previous config saved to /var/cache/conftool/dbconfig/20260720-200310-cwilliams.json * 19:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94940 and previous config saved to /var/cache/conftool/dbconfig/20260720-195633-cwilliams.json * 19:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1257.eqiad.wmnet with reason: Maintenance * 19:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94939 and previous config saved to /var/cache/conftool/dbconfig/20260720-195605-cwilliams.json * 19:51 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 19:50 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 19:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256', diff saved to https://phabricator.wikimedia.org/P94938 and previous config saved to /var/cache/conftool/dbconfig/20260720-194558-cwilliams.json * 19:44 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 19:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 19:41 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wdqs1011.eqiad.wmnet with OS bookworm * 19:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256', diff saved to https://phabricator.wikimedia.org/P94937 and previous config saved to /var/cache/conftool/dbconfig/20260720-193550-cwilliams.json * 19:25 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94936 and previous config saved to /var/cache/conftool/dbconfig/20260720-192542-cwilliams.json * 19:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94935 and previous config saved to /var/cache/conftool/dbconfig/20260720-191856-cwilliams.json * 19:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1256.eqiad.wmnet with reason: Maintenance * 19:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94934 and previous config saved to /var/cache/conftool/dbconfig/20260720-191839-cwilliams.json * 19:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255', diff saved to https://phabricator.wikimedia.org/P94933 and previous config saved to /var/cache/conftool/dbconfig/20260720-190831-cwilliams.json * 18:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255', diff saved to https://phabricator.wikimedia.org/P94932 and previous config saved to /var/cache/conftool/dbconfig/20260720-185824-cwilliams.json * 18:50 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 18:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94931 and previous config saved to /var/cache/conftool/dbconfig/20260720-184816-cwilliams.json * 18:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94930 and previous config saved to /var/cache/conftool/dbconfig/20260720-184224-cwilliams.json * 18:42 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1255.eqiad.wmnet with reason: Maintenance * 18:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94929 and previous config saved to /var/cache/conftool/dbconfig/20260720-184153-cwilliams.json * 18:39 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 18:39 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 16s) * 18:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 18:38 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 59m 26s) * 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 18:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211', diff saved to https://phabricator.wikimedia.org/P94928 and previous config saved to /var/cache/conftool/dbconfig/20260720-183145-cwilliams.json * 18:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211', diff saved to https://phabricator.wikimedia.org/P94927 and previous config saved to /var/cache/conftool/dbconfig/20260720-182137-cwilliams.json * 18:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94926 and previous config saved to /var/cache/conftool/dbconfig/20260720-181129-cwilliams.json * 18:09 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_codfw * 18:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2057.codfw.wmnet * 18:08 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_codfw * 18:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2058.codfw.wmnet * 18:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94925 and previous config saved to /var/cache/conftool/dbconfig/20260720-180452-cwilliams.json * 18:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on clouddb[1016,1020,1022-1023].eqiad.wmnet,db1154.eqiad.wmnet with reason: Maintenance * 18:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1211.eqiad.wmnet with reason: Maintenance * 18:02 sukhe: armed keyholder on acmechief1002.eqiad.wmnet and acmechief2002.codfw.wmnet (active host) * 18:01 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief2002.codfw.wmnet * 17:57 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief2002.codfw.wmnet * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs2020'] * 17:52 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief1002.eqiad.wmnet * 17:50 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 17:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1011.eqiad.wmnet with reason: host reimage * 17:48 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief1002.eqiad.wmnet * 17:47 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test2001.codfw.wmnet * 17:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94924 and previous config saved to /var/cache/conftool/dbconfig/20260720-174717-cwilliams.json * 17:46 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 17:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1011.eqiad.wmnet with reason: host reimage * 17:43 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs2020.codfw.wmnet with OS bookworm * 17:43 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test2001.codfw.wmnet * 17:43 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test1001.eqiad.wmnet * 17:39 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test1001.eqiad.wmnet * 17:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 17:38 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:38 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 17:37 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 17:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244', diff saved to https://phabricator.wikimedia.org/P94923 and previous config saved to /var/cache/conftool/dbconfig/20260720-173709-cwilliams.json * 17:35 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 17:31 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 20m 40s) * 17:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2055.codfw.wmnet * 17:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2056.codfw.wmnet * 17:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1011 * 17:27 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1011 * 17:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1011.eqiad.wmnet with OS bookworm * 17:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244', diff saved to https://phabricator.wikimedia.org/P94922 and previous config saved to /var/cache/conftool/dbconfig/20260720-172701-cwilliams.json * 17:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94921 and previous config saved to /var/cache/conftool/dbconfig/20260720-171653-cwilliams.json * 17:11 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 17:11 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 13m 03s) * 17:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94920 and previous config saved to /var/cache/conftool/dbconfig/20260720-171012-cwilliams.json * 17:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2244.codfw.wmnet with reason: Maintenance * 17:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94919 and previous config saved to /var/cache/conftool/dbconfig/20260720-170941-cwilliams.json * 16:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243', diff saved to https://phabricator.wikimedia.org/P94918 and previous config saved to /var/cache/conftool/dbconfig/20260720-165933-cwilliams.json * 16:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 16:58 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2053.codfw.wmnet * 16:51 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2054.codfw.wmnet * 16:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243', diff saved to https://phabricator.wikimedia.org/P94917 and previous config saved to /var/cache/conftool/dbconfig/20260720-164926-cwilliams.json * 16:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94916 and previous config saved to /var/cache/conftool/dbconfig/20260720-163918-cwilliams.json * 16:35 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 16:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94915 and previous config saved to /var/cache/conftool/dbconfig/20260720-163140-cwilliams.json * 16:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2243.codfw.wmnet with reason: Maintenance * 16:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94914 and previous config saved to /var/cache/conftool/dbconfig/20260720-163111-cwilliams.json * 16:27 btullis@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 16:27 btullis@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 16:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2020 * 16:23 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2020 * 16:21 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2020 * 16:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2020.codfw.wmnet 85.0.192.10.in-addr.arpa 5.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:21 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2020.codfw.wmnet 85.0.192.10.in-addr.arpa 5.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242', diff saved to https://phabricator.wikimedia.org/P94913 and previous config saved to /var/cache/conftool/dbconfig/20260720-162103-cwilliams.json * 16:19 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 16:18 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 16:18 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:18 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 16:17 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:17 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for netbox accounting errors - jhancock@cumin2002" * 16:17 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for netbox accounting errors - jhancock@cumin2002" * 16:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2051.codfw.wmnet * 16:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2052.codfw.wmnet * 16:11 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 16:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242', diff saved to https://phabricator.wikimedia.org/P94912 and previous config saved to /var/cache/conftool/dbconfig/20260720-161055-cwilliams.json * 16:09 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 16:08 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 16:06 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 16:06 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 16:06 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2020 * 16:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2020.codfw.wmnet with OS bookworm * 16:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94911 and previous config saved to /var/cache/conftool/dbconfig/20260720-160047-cwilliams.json * 15:58 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2019.codfw.wmnet, repooling source-only afterwards * 15:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94909 and previous config saved to /var/cache/conftool/dbconfig/20260720-155353-cwilliams.json * 15:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2242.codfw.wmnet with reason: Maintenance * 15:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94908 and previous config saved to /var/cache/conftool/dbconfig/20260720-154433-cwilliams.json * 15:35 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2049.codfw.wmnet * 15:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162', diff saved to https://phabricator.wikimedia.org/P94907 and previous config saved to /var/cache/conftool/dbconfig/20260720-153425-cwilliams.json * 15:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2050.codfw.wmnet * 15:28 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162', diff saved to https://phabricator.wikimedia.org/P94906 and previous config saved to /var/cache/conftool/dbconfig/20260720-152418-cwilliams.json * 15:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1023 * 15:14 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1023 * 15:14 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 15:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94905 and previous config saved to /var/cache/conftool/dbconfig/20260720-151407-cwilliams.json * 15:13 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] (duration: 41m 16s) * 15:08 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 15:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94902 and previous config saved to /var/cache/conftool/dbconfig/20260720-150729-cwilliams.json * 15:07 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2162.codfw.wmnet with reason: Maintenance * 15:05 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2027.codfw.wmnet, repooling source-only afterwards * 15:00 urbanecm@deploy2003: vadymts1, migr, urbanecm: Continuing with deployment * 14:59 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:58 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 07s) * 14:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:58 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 13s) * 14:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:57 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2019.codfw.wmnet, repooling source-only afterwards * 14:57 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2047.codfw.wmnet * 14:55 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2048.codfw.wmnet * 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2019.codfw.wmnet with OS bookworm * 14:49 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 14:47 urbanecm@deploy2003: vadymts1, migr, urbanecm: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:44 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool magru [reason: BGP issues in lvs7003 resolved after liberica restart, no task ID specified] * 14:44 sukhe@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool magru [reason: BGP issues in lvs7003 resolved after liberica restart, no task ID specified] * 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:41 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:39 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:39 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:33 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool magru [reason: no reason specified, no task ID specified] * 14:33 sukhe@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool magru [reason: no reason specified, no task ID specified] * 14:31 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] * 14:24 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:24 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:24 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:24 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2019.codfw.wmnet with reason: host reimage * 14:22 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2027.codfw.wmnet, repooling source-only afterwards * 14:19 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2019.codfw.wmnet with reason: host reimage * 14:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2027.codfw.wmnet with OS bookworm * 14:16 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2046.codfw.wmnet * 14:16 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2045.codfw.wmnet * 14:08 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:08 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:08 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:08 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:07 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:06 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:06 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:06 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:05 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2019 * 14:00 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2019 * 13:56 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1015.eqiad.wmnet * 13:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2027.codfw.wmnet with reason: host reimage * 13:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2071.codfw.wmnet with OS trixie * 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:51 sukhe@cumin1003: END (ERROR) - Cookbook sre.loadbalancer.admin (exit_code=97) rebooting A:liberica and P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica and P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:51 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1015.eqiad.wmnet * 13:50 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1014.eqiad.wmnet * 13:50 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1076.eqiad.wmnet with OS trixie * 13:50 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2027.codfw.wmnet with reason: host reimage * 13:45 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1014.eqiad.wmnet * 13:44 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1013.eqiad.wmnet * 13:39 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 13:38 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1013.eqiad.wmnet * 13:37 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2044.codfw.wmnet * 13:37 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2043.codfw.wmnet * 13:36 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2019 * 13:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2019.codfw.wmnet 156.32.192.10.in-addr.arpa 6.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:36 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2019.codfw.wmnet 156.32.192.10.in-addr.arpa 6.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:36 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2019 - bking@cumin2003" * 13:36 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2019 - bking@cumin2003" * 13:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry2005.codfw.wmnet * 13:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2071.codfw.wmnet with reason: host reimage * 13:31 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:31 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2019 * 13:31 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry2005.codfw.wmnet * 13:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry2004.codfw.wmnet * 13:30 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2019.codfw.wmnet with OS bookworm * 13:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2027 * 13:30 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2027 * 13:30 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2027.codfw.wmnet with OS bookworm * 13:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1076.eqiad.wmnet with reason: host reimage * 13:29 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_codfw * 13:28 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_codfw * 13:26 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry2004.codfw.wmnet * 13:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry1005.eqiad.wmnet * 13:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2071.codfw.wmnet with reason: host reimage * 13:22 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1076.eqiad.wmnet with reason: host reimage * 13:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry1005.eqiad.wmnet * 13:21 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry1004.eqiad.wmnet * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry1004.eqiad.wmnet * 13:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts * 13:13 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts * 13:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts * 13:12 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts * 13:03 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1076.eqiad.wmnet with OS trixie * 13:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2071.codfw.wmnet with OS trixie * 12:55 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:54 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:53 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:46 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 7 hosts * 12:42 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 7 hosts * 12:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts * 12:42 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts * 12:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2070.codfw.wmnet with OS trixie * 12:36 ozge@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:35 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1075.eqiad.wmnet with OS trixie * 12:32 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts * 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts * 12:22 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin1001.eqiad.wmnet * 12:19 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin1001.eqiad.wmnet * 12:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2070.codfw.wmnet with reason: host reimage * 12:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin2001.codfw.wmnet * 12:14 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1075.eqiad.wmnet with reason: host reimage * 12:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2070.codfw.wmnet with reason: host reimage * 12:10 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1075.eqiad.wmnet with reason: host reimage * 12:09 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin2001.codfw.wmnet * 11:17 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1074.eqiad.wmnet with OS trixie * 11:17 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 11:16 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 11:14 ozge@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 6 hosts * 11:09 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 6 hosts * 11:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 324 hosts * 10:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1074.eqiad.wmnet with reason: host reimage * 10:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2069.codfw.wmnet with OS trixie * 10:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1074.eqiad.wmnet with reason: host reimage * 10:30 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2069.codfw.wmnet with reason: host reimage * 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1074.eqiad.wmnet with OS trixie * 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2069.codfw.wmnet with reason: host reimage * 10:06 blake@deploy2003: Stopping before sync operations * 10:06 blake@deploy2003: Started scap sync-world: Non-deployment scap run to populate new release values * 10:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2069.codfw.wmnet with OS trixie * 10:00 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1073.eqiad.wmnet with OS trixie * 09:56 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 324 hosts * 09:39 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 09:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 8 hosts * 09:38 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1073.eqiad.wmnet with reason: host reimage * 09:37 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 8 hosts * 09:34 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1073.eqiad.wmnet with reason: host reimage * 09:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2068.codfw.wmnet with OS trixie * 09:16 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1073.eqiad.wmnet with OS trixie * 09:13 blake@deploy2003: sync-world aborted: Non-deployment scap run to populate new release values (duration: 00m 02s) * 09:13 blake@deploy2003: Started scap sync-world: Non-deployment scap run to populate new release values * 08:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2068.codfw.wmnet with reason: host reimage * 08:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2068.codfw.wmnet with reason: host reimage * 08:50 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 08:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2068.codfw.wmnet with OS trixie * 08:15 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1072.eqiad.wmnet with OS trixie * 07:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2067.codfw.wmnet with OS trixie * 07:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1072.eqiad.wmnet with reason: host reimage * 07:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1072.eqiad.wmnet with reason: host reimage * 07:45 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 07:45 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 07:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2067.codfw.wmnet with reason: host reimage * 07:35 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2067.codfw.wmnet with reason: host reimage * 07:30 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 07:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1072.eqiad.wmnet with OS trixie * 07:30 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 07:17 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 07:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2067.codfw.wmnet with OS trixie * 05:51 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:50 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:25 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:25 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on db2207.codfw.wmnet with reason: Host down * 04:28 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 07m 02s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-18 == * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 29s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 00:11 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2018.codfw.wmnet, repooling source-only afterwards == 2026-07-17 == * 23:53 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2026.codfw.wmnet, repooling source-only afterwards * 23:09 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2018.codfw.wmnet, repooling source-only afterwards * 23:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2026.codfw.wmnet, repooling source-only afterwards * 22:11 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2018.codfw.wmnet with OS bookworm * 22:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2026.codfw.wmnet with OS bookworm * 21:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2018.codfw.wmnet with reason: host reimage * 21:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2018.codfw.wmnet with reason: host reimage * 21:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2026.codfw.wmnet with reason: host reimage * 21:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2026.codfw.wmnet with reason: host reimage * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2018 * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2018 * 21:26 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2018 * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2018.codfw.wmnet 155.32.192.10.in-addr.arpa 5.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:26 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2018.codfw.wmnet 155.32.192.10.in-addr.arpa 5.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2018 - bking@cumin2003" * 21:26 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2018 - bking@cumin2003" * 21:14 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:13 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2018 * 21:13 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2018.codfw.wmnet with OS bookworm * 21:12 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2026 * 21:12 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2026 * 21:12 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2026.codfw.wmnet with OS bookworm * 21:05 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1022.eqiad.wmnet -> wdqs1026.eqiad.wmnet, repooling source-only afterwards * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs2017.codfw.wmnet, repooling source-only afterwards * 20:11 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs2017.codfw.wmnet, repooling source-only afterwards * 20:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2017.codfw.wmnet with OS bookworm * 20:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1022.eqiad.wmnet -> wdqs1026.eqiad.wmnet, repooling source-only afterwards * 20:06 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1026.eqiad.wmnet with OS bookworm * 19:55 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 19:55 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:55 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 09s) * 19:55 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:50 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 08s) * 19:50 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:50 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 10m 03s) * 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2017.codfw.wmnet with reason: host reimage * 19:40 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:40 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 15s) * 19:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1026.eqiad.wmnet with reason: host reimage * 19:37 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 16s) * 19:37 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:34 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2017.codfw.wmnet with reason: host reimage * 19:34 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1026.eqiad.wmnet with reason: host reimage * 19:33 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 19:33 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:16 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1026.eqiad.wmnet with OS bookworm * 19:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2017 * 19:16 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2017 * 19:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2017.codfw.wmnet with OS bookworm * 18:30 bking@dns1004: END - running authdns-update * 18:28 bking@dns1004: START - running authdns-update * 18:16 kamila@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1264.eqiad.wmnet * 18:16 kamila@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1264.eqiad.wmnet * 18:16 kamila@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1264.eqiad.wmnet * 17:49 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 17:46 dzahn@dns1006: END - running authdns-update * 17:44 dzahn@dns1006: START - running authdns-update * 17:44 dzahn@dns1006: END - running authdns-update * 17:42 dzahn@dns1006: START - running authdns-update * 17:28 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 17:21 kamila@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 17:01 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1264 * 17:01 kamila@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1264 * 17:01 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 17:01 kamila@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1264.eqiad.wmnet * 17:01 kamila@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1264.eqiad.wmnet * 17:01 kamila@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1264.eqiad.wmnet * 16:42 reedy@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] (duration: 10m 29s) * 16:34 reedy@deploy2003: reedy, hartman: Continuing with deployment * 16:33 reedy@deploy2003: reedy, hartman: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:31 reedy@deploy2003: Started scap sync-world: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] * 16:26 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 16:10 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in2001.wikimedia.org with reason: [[phab:T431659|T431659]] * 16:07 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in1001.wikimedia.org with reason: [[phab:T431659|T431659]] * 16:05 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 16:01 kamila@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 16:00 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out2001.wikimedia.org with reason: [[phab:T431659|T431659]] * 15:41 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 15:41 kamila@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 15:35 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out1001.wikimedia.org with reason: [[phab:T431659|T431659]] * 15:14 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1339.eqiad.wmnet * 15:13 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1339.eqiad.wmnet * 15:13 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1339.eqiad.wmnet * 14:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1339.eqiad.wmnet with OS trixie * 14:50 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:49 kamila@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:49 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:33 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage * 14:27 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage * 14:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1339 * 14:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1339 * 14:14 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1339 * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1339.eqiad.wmnet 156.32.64.10.in-addr.arpa 6.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:14 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1339.eqiad.wmnet 156.32.64.10.in-addr.arpa 6.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1339 - cgoubert@cumin2003" * 14:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1339 - cgoubert@cumin2003" * 14:09 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 14:06 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1339 * 14:06 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie * 14:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1339.eqiad.wmnet * 14:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1339.eqiad.wmnet * 14:02 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1339.eqiad.wmnet * 13:45 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb1013.eqiad.wmnet * 13:39 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb1013.eqiad.wmnet * 13:27 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:24 blake@dns1004: END - running authdns-update * 13:22 blake@dns1004: START - running authdns-update * 13:20 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 13:11 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2014.codfw.wmnet * 13:06 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb2014.codfw.wmnet * 13:06 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2012.codfw.wmnet * 13:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 13:03 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 13:01 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 15 hosts * 13:01 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb2012.codfw.wmnet * 13:01 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1016.eqiad.wmnet * 13:00 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 15 hosts * 12:55 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb1016.eqiad.wmnet * 12:55 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1014.eqiad.wmnet * 12:49 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb1014.eqiad.wmnet * 12:32 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:32 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:31 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:31 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1338.eqiad.wmnet * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1338.eqiad.wmnet * 12:18 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1338.eqiad.wmnet * 12:17 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:15 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:14 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:13 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1338.eqiad.wmnet with OS trixie * 12:01 klausman@deploy2003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 11:59 klausman@deploy2003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 11:56 klausman@deploy2003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 11:54 klausman@deploy2003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 11:53 klausman@deploy2003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 11:51 klausman@deploy2003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 11:42 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1338.eqiad.wmnet with reason: host reimage * 11:38 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1338.eqiad.wmnet with reason: host reimage * 11:31 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2230.codfw.wmnet * 11:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1338 * 11:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1338 * 11:25 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1338 * 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1338.eqiad.wmnet 155.32.64.10.in-addr.arpa 5.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:25 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1338.eqiad.wmnet 155.32.64.10.in-addr.arpa 5.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1338 - cgoubert@cumin2003" * 11:25 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1338 - cgoubert@cumin2003" * 11:23 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2230.codfw.wmnet * 11:20 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 11:20 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1338 * 11:20 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1338.eqiad.wmnet with OS trixie * 11:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1338.eqiad.wmnet * 11:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1338.eqiad.wmnet * 11:19 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1338.eqiad.wmnet * 11:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1337.eqiad.wmnet * 11:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1337.eqiad.wmnet * 11:17 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1337.eqiad.wmnet * 11:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1337.eqiad.wmnet with OS trixie * 10:51 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[2001-2002].codfw.wmnet * 10:50 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1337.eqiad.wmnet with reason: host reimage * 10:40 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 10:39 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:39 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1337.eqiad.wmnet with reason: host reimage * 10:39 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1001-1003].eqiad.wmnet * 10:34 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:34 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:30 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:28 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1001-1003].eqiad.wmnet * 10:27 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1337 * 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1337 * 10:26 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1337 * 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1337.eqiad.wmnet 154.32.64.10.in-addr.arpa 4.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:26 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1337.eqiad.wmnet 154.32.64.10.in-addr.arpa 4.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1337 - cgoubert@cumin2003" * 10:26 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1337 - cgoubert@cumin2003" * 10:21 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 10:18 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1337 * 10:17 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1337.eqiad.wmnet with OS trixie * 10:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1337.eqiad.wmnet * 10:16 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db1176.eqiad.wmnet * 10:16 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1337.eqiad.wmnet * 10:16 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1337.eqiad.wmnet * 10:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1336.eqiad.wmnet * 10:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1336.eqiad.wmnet * 10:15 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1336.eqiad.wmnet * 10:11 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db1176.eqiad.wmnet * 10:10 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db1176.eqiad.wmnet * 10:09 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db1176.eqiad.wmnet * 10:05 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts (check the cookbook's logs for more details.) * 10:03 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts (check the cookbook's logs for more details.) * 09:58 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1336.eqiad.wmnet with OS trixie * 09:47 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts (check the cookbook's logs for more details.) * 09:47 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts (check the cookbook's logs for more details.) * 09:45 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host acmechief-test2001.codfw.wmnet,acmechief-test1001.eqiad.wmnet,an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet,db-test[2001-2002].codfw.wmnet,db-test[1 * 09:40 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host acmechief-test2001.codfw.wmnet,acmechief-test1001.eqiad.wmnet,an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet,db-test[2001-2002].codfw.wmnet,db-test[1001-1003].eqiad.wmn * 09:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1336.eqiad.wmnet with reason: host reimage * 09:33 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1336.eqiad.wmnet with reason: host reimage * 09:29 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet * 09:29 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet * 09:28 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:26 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:21 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 09:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1336 * 09:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1336 * 09:19 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 09:14 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1336 * 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1336.eqiad.wmnet 152.32.64.10.in-addr.arpa 2.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1336.eqiad.wmnet 152.32.64.10.in-addr.arpa 2.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1336 - cgoubert@cumin2003" * 09:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1336 - cgoubert@cumin2003" * 09:11 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:10 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 09:09 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:09 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1336 * 09:09 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1336.eqiad.wmnet with OS trixie * 09:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1336.eqiad.wmnet * 09:08 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1336.eqiad.wmnet * 09:08 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1336.eqiad.wmnet * 09:06 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1335.eqiad.wmnet * 09:06 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1335.eqiad.wmnet * 09:06 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1335.eqiad.wmnet * 09:04 elukey: uploaded spicerack_13.1.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia * 08:55 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wikikube-worker-exp2001.codfw.wmnet * 08:54 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host testreduce1002.eqiad.wmnet * 08:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1335.eqiad.wmnet with OS trixie * 08:51 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host wikikube-worker-exp2001.codfw.wmnet * 08:51 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wikikube-worker-exp1001.eqiad.wmnet * 08:50 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host testreduce1002.eqiad.wmnet * 08:45 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host wikikube-worker-exp1001.eqiad.wmnet * 08:34 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1335.eqiad.wmnet with reason: host reimage * 08:30 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1335.eqiad.wmnet with reason: host reimage * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1335 * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1335 * 08:18 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1335 * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1335.eqiad.wmnet 150.32.64.10.in-addr.arpa 0.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:18 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1335.eqiad.wmnet 150.32.64.10.in-addr.arpa 0.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1335 - cgoubert@cumin2003" * 08:18 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1335 - cgoubert@cumin2003" * 08:14 elukey@cumin1003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 08:14 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges * 08:13 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 08:10 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1335 * 08:10 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1335.eqiad.wmnet with OS trixie * 08:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1335.eqiad.wmnet * 08:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1335.eqiad.wmnet * 08:09 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1335.eqiad.wmnet * 08:06 elukey@cumin1003: END (FAIL) - Cookbook sre.puppet.disable-merges (exit_code=99) * 08:05 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges * 08:03 elukey@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin1003.eqiad.wmnet * 07:57 elukey@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin1003.eqiad.wmnet * 07:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetdb1003.eqiad.wmnet * 07:46 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetdb1003.eqiad.wmnet * 07:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetdb2003.codfw.wmnet * 07:37 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetdb2003.codfw.wmnet * 07:37 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1001.eqiad.wmnet * 07:28 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver1001.eqiad.wmnet * 07:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet * 07:19 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet * 07:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2002.codfw.wmnet * 07:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver2002.codfw.wmnet * 07:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet * 07:05 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet * 07:04 btullis@cumin1003: END (FAIL) - Cookbook sre.hadoop.reboot-workers (exit_code=99) for Hadoop analytics cluster * 07:04 elukey@cumin1003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 07:04 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges * 06:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox1003.eqiad.wmnet * 06:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox1003.eqiad.wmnet * 02:46 ryankemper@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:46 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:44 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:37 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:37 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-internal-scholarly,name=eqiad * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 49s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 01:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore wdqs1025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling source-only afterwards * 01:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore wdqs1027 after Bookworm reimage) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs1027.eqiad.wmnet, repooling both afterwards * 00:55 urbanecm@deploy2003: helmfile [codfw] DONE helmfile.d/services/linkrecommendation: apply * 00:54 urbanecm@deploy2003: helmfile [eqiad] DONE helmfile.d/services/linkrecommendation: apply * 00:54 urbanecm@deploy2003: helmfile [staging] DONE helmfile.d/services/linkrecommendation: apply * 00:54 urbanecm@deploy2003: helmfile [codfw] START helmfile.d/services/linkrecommendation: apply * 00:53 urbanecm@deploy2003: helmfile [staging] START helmfile.d/services/linkrecommendation: apply * 00:52 urbanecm@deploy2003: helmfile [eqiad] START helmfile.d/services/linkrecommendation: apply * 00:23 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore wdqs1025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling source-only afterwards * 00:23 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore wdqs1027 after Bookworm reimage) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs1027.eqiad.wmnet, repooling both afterwards * 00:14 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1274.eqiad.wmnet * 00:14 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1274.eqiad.wmnet * 00:14 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1274.eqiad.wmnet * 00:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1027.eqiad.wmnet with OS bookworm * 00:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1025.eqiad.wmnet with OS bookworm * 00:04 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1274.eqiad.wmnet with OS trixie == 2026-07-16 == * 23:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], xfer to freshly reimaged/scap-deployed wdqs2025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs2025.codfw.wmnet, repooling source-only afterwards * 23:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1027.eqiad.wmnet with reason: host reimage * 23:47 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1025.eqiad.wmnet with reason: host reimage * 23:43 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1274.eqiad.wmnet with reason: host reimage * 23:41 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1025.eqiad.wmnet with reason: host reimage * 23:39 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1027.eqiad.wmnet with reason: host reimage * 23:38 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1274.eqiad.wmnet with reason: host reimage * 23:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1025 * 23:23 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1025 * 23:22 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1027 * 23:22 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1027 * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1274 * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1274 * 23:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1025.eqiad.wmnet with OS bookworm * 23:19 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1274 * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1274.eqiad.wmnet 145.48.64.10.in-addr.arpa 5.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:19 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1274.eqiad.wmnet 145.48.64.10.in-addr.arpa 5.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1274 - swfrench@cumin1003" * 23:19 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1274 - swfrench@cumin1003" * 23:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1027.eqiad.wmnet with OS bookworm * 23:14 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 23:14 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1274 * 23:13 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1274.eqiad.wmnet with OS trixie * 23:13 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1274.eqiad.wmnet * 23:12 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1274.eqiad.wmnet * 23:12 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1274.eqiad.wmnet * 23:12 ryankemper: [[phab:T430880|T430880]] depooled dnsdisc of wdqs-internal-scholarly-eqiad bc we only have 1 host there * 23:09 ryankemper@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-internal-scholarly,name=eqiad * 23:08 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1272.eqiad.wmnet * 23:08 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1272.eqiad.wmnet * 23:08 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1272.eqiad.wmnet * 23:01 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], xfer to freshly reimaged/scap-deployed wdqs2025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs2025.codfw.wmnet, repooling source-only afterwards * 22:57 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1272.eqiad.wmnet with OS trixie * 22:56 Amir1: deleting echo notifications from 2015 in group0 * 22:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2025.codfw.wmnet with OS bookworm * 22:35 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1272.eqiad.wmnet with reason: host reimage * 22:32 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 27s) * 22:32 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 22:28 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1269.eqiad.wmnet * 22:28 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1269.eqiad.wmnet * 22:28 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1269.eqiad.wmnet * 22:27 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1272.eqiad.wmnet with reason: host reimage * 22:26 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] (duration: 08m 51s) * 22:22 ladsgroup@deploy2003: ladsgroup, urbanecm: Continuing with deployment * 22:19 ladsgroup@deploy2003: ladsgroup, urbanecm: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:17 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] * 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2025.codfw.wmnet with reason: host reimage * 22:06 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1272 * 22:06 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1272 * 22:05 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1272 * 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1272.eqiad.wmnet 127.48.64.10.in-addr.arpa 7.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:05 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1272.eqiad.wmnet 127.48.64.10.in-addr.arpa 7.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1272 - swfrench@cumin1003" * 22:05 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1272 - swfrench@cumin1003" * 22:03 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2025.codfw.wmnet with reason: host reimage * 22:01 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 22:00 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1272 * 22:00 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1272.eqiad.wmnet with OS trixie * 22:00 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1272.eqiad.wmnet * 21:59 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1272.eqiad.wmnet * 21:59 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1272.eqiad.wmnet * 21:55 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1271.eqiad.wmnet * 21:55 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1271.eqiad.wmnet * 21:55 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1271.eqiad.wmnet * 21:46 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1271.eqiad.wmnet with OS trixie * 21:45 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] (duration: 06m 31s) * 21:40 sbassett@deploy2003: sbassett: Continuing with deployment * 21:40 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2025 * 21:40 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2025 * 21:40 sbassett@deploy2003: sbassett: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:38 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] * 21:37 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2025 * 21:37 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2025.codfw.wmnet 220.48.192.10.in-addr.arpa 0.2.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:37 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2025.codfw.wmnet 220.48.192.10.in-addr.arpa 0.2.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:37 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:37 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2025 - bking@cumin2003" * 21:37 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2025 - bking@cumin2003" * 21:30 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] (duration: 08m 19s) * 21:26 sbassett@deploy2003: sbassett: Continuing with deployment * 21:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1269.eqiad.wmnet with OS trixie * 21:24 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1271.eqiad.wmnet with reason: host reimage * 21:23 sbassett@deploy2003: sbassett: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:22 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:22 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] * 21:20 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2025 * 21:19 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2025.codfw.wmnet with OS bookworm * 21:17 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1271.eqiad.wmnet with reason: host reimage * 21:04 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1269.eqiad.wmnet with reason: host reimage * 21:00 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1269.eqiad.wmnet with reason: host reimage * 20:56 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1271 * 20:55 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1271 * 20:54 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1271 * 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1271.eqiad.wmnet 126.48.64.10.in-addr.arpa 6.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:54 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1271.eqiad.wmnet 126.48.64.10.in-addr.arpa 6.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1271 - swfrench@cumin1003" * 20:54 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1271 - swfrench@cumin1003" * 20:51 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:51 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:51 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:50 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 20:49 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 20:49 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1268.eqiad.wmnet * 20:49 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1268.eqiad.wmnet * 20:49 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1268.eqiad.wmnet * 20:48 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1271 * 20:48 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1271.eqiad.wmnet with OS trixie * 20:47 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1271.eqiad.wmnet * 20:46 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1271.eqiad.wmnet * 20:46 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1271.eqiad.wmnet * 20:41 aude@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] (duration: 07m 34s) * 20:39 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1269 * 20:39 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1269 * 20:36 aude@deploy2003: aude: Continuing with deployment * 20:35 aude@deploy2003: aude: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:33 aude@deploy2003: Started scap sync-world: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] * 20:26 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on db2207.codfw.wmnet with reason: Host down * 20:22 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-video: apply * 20:21 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-video: apply * 20:20 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-timeline: apply * 20:20 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-timeline: apply * 20:20 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-syntaxhighlight: apply * 20:19 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-syntaxhighlight: apply * 20:19 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-media: apply * 20:18 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-media: apply * 20:18 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-constraints: apply * 20:17 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-constraints: apply * 20:17 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox: apply * 20:16 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox: apply * 20:13 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1269 * 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1269.eqiad.wmnet 80.32.64.10.in-addr.arpa 0.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:13 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1269.eqiad.wmnet 80.32.64.10.in-addr.arpa 0.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1269 - kamila@cumin1003" * 20:13 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1269 - kamila@cumin1003" * 20:09 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-flink-codfw cluster: Roll restart of jvm daemons. * 20:07 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 20:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 20:03 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-flink-codfw cluster: Roll restart of jvm daemons. * 20:03 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2207 [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94893 and previous config saved to /var/cache/conftool/dbconfig/20260716-200257-marostegui.json * 20:01 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2204 to s2 primary [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94892 and previous config saved to /var/cache/conftool/dbconfig/20260716-200157-marostegui.json * 20:00 marostegui: Starting emergency s2 codfw failover from db2207 to db2204 - [[phab:T432396|T432396]] * 19:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1035.eqiad.wmnet * 19:56 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2204 with weight 0 [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94891 and previous config saved to /var/cache/conftool/dbconfig/20260716-195628-marostegui.json * 19:55 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 26 hosts with reason: Primary switchover s2 [[phab:T432396|T432396]] * 19:54 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1035.eqiad.wmnet * 19:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1034.eqiad.wmnet * 19:48 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1034.eqiad.wmnet * 19:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1033.eqiad.wmnet * 19:43 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-video: apply * 19:43 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1033.eqiad.wmnet * 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1032.eqiad.wmnet * 19:42 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-video: apply * 19:42 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-timeline: apply * 19:41 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-timeline: apply * 19:41 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-syntaxhighlight: apply * 19:41 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-syntaxhighlight: apply * 19:40 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-media: apply * 19:40 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-media: apply * 19:39 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-constraints: apply * 19:36 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-constraints: apply * 19:36 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox: apply * 19:35 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1032.eqiad.wmnet * 19:35 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1031.eqiad.wmnet * 19:35 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox: apply * 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-video: apply * 19:33 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-video: apply * 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-timeline: apply * 19:33 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-timeline: apply * 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-syntaxhighlight: apply * 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-syntaxhighlight: apply * 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-media: apply * 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-media: apply * 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-constraints: apply * 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-constraints: apply * 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox: apply * 19:31 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox: apply * 19:27 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1031.eqiad.wmnet * 19:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1030.eqiad.wmnet * 19:23 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1001.eqiad.wmnet, repooling source-only afterwards * 19:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1030.eqiad.wmnet * 19:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1029.eqiad.wmnet * 19:17 kamila@cumin1003: START - Cookbook sre.dns.netbox * 19:12 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1029.eqiad.wmnet * 19:06 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1269 * 19:05 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1269.eqiad.wmnet with OS trixie * 19:03 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1269.eqiad.wmnet * 19:03 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1269.eqiad.wmnet * 19:03 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1269.eqiad.wmnet * 18:55 dancy@deploy2003: Finished scap sync-world: testing [[phab:T428971|T428971]] (duration: 02m 41s) * 18:53 dancy@deploy2003: Started scap sync-world: testing [[phab:T428971|T428971]] * 18:31 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1268.eqiad.wmnet with OS trixie * 18:18 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 18:16 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1267.eqiad.wmnet * 18:16 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1267.eqiad.wmnet * 18:16 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1267.eqiad.wmnet * 18:09 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1268.eqiad.wmnet with reason: host reimage * 18:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1001.eqiad.wmnet, repooling source-only afterwards * 18:06 swfrench@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] (duration: 07m 34s) * 18:06 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 23s) * 18:06 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 18:06 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1268.eqiad.wmnet with reason: host reimage * 18:03 bd808@deploy2003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 18:02 bd808@deploy2003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 18:02 swfrench@deploy2003: jiji, swfrench: Continuing with deployment * 18:02 bd808@deploy2003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 18:02 bd808@deploy2003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 18:01 bd808@deploy2003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 18:01 swfrench@deploy2003: jiji, swfrench: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:01 bd808@deploy2003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:59 swfrench@deploy2003: Started scap sync-world: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] * 17:45 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1268 * 17:45 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1268 * 17:44 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1267.eqiad.wmnet with OS trixie * 17:43 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter2006.codfw.wmnet * 17:39 swfrench@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter2006.codfw.wmnet * 17:35 swfrench@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] (duration: 07m 27s) * 17:34 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1268 * 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1268.eqiad.wmnet 78.32.64.10.in-addr.arpa 8.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:34 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1268.eqiad.wmnet 78.32.64.10.in-addr.arpa 8.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1268 - kamila@cumin1003" * 17:34 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1268 - kamila@cumin1003" * 17:31 swfrench@deploy2003: jiji, swfrench: Continuing with deployment * 17:29 swfrench@deploy2003: jiji, swfrench: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:28 kamila@cumin1003: START - Cookbook sre.dns.netbox * 17:28 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1268 * 17:28 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1268.eqiad.wmnet with OS trixie * 17:27 swfrench@deploy2003: Started scap sync-world: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] * 17:23 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1267.eqiad.wmnet with reason: host reimage * 17:18 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1267.eqiad.wmnet with reason: host reimage * 17:18 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1268.eqiad.wmnet * 17:17 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1268.eqiad.wmnet * 17:17 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1268.eqiad.wmnet * 17:12 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter2005.codfw.wmnet * 17:11 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1270.eqiad.wmnet * 17:11 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1270.eqiad.wmnet * 17:11 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1270.eqiad.wmnet * 17:09 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter2005.codfw.wmnet * 17:08 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:08 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update reverse dns for moved arelion cct cr2-eqiad - cmooney@cumin1003" * 17:08 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update reverse dns for moved arelion cct cr2-eqiad - cmooney@cumin1003" * 17:08 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] (duration: 07m 34s) * 17:04 jiji@deploy2003: jiji: Continuing with deployment * 17:03 jiji@deploy2003: jiji: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 17:00 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] * 17:00 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2185.codfw.wmnet with OS trixie * 16:59 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:58 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1270.eqiad.wmnet with OS trixie * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1267 * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1267 * 16:57 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1267 * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1267.eqiad.wmnet 77.32.64.10.in-addr.arpa 7.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:57 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1267.eqiad.wmnet 77.32.64.10.in-addr.arpa 7.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1267 - kamila@cumin1003" * 16:56 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1267 - kamila@cumin1003" * 16:56 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_eqsin * 16:56 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5032.eqsin.wmnet * 16:52 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_esams * 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3073.esams.wmnet * 16:50 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_esams * 16:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3081.esams.wmnet * 16:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1266.eqiad.wmnet * 16:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1266.eqiad.wmnet * 16:45 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1266.eqiad.wmnet * 16:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2185.codfw.wmnet with reason: host reimage * 16:41 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_eqiad * 16:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1114.eqiad.wmnet * 16:41 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_eqiad * 16:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1115.eqiad.wmnet * 16:39 kamila@cumin1003: START - Cookbook sre.dns.netbox * 16:39 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1267 * 16:39 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2185.codfw.wmnet with reason: host reimage * 16:38 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1267.eqiad.wmnet with OS trixie * 16:38 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1267.eqiad.wmnet * 16:38 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1270.eqiad.wmnet with reason: host reimage * 16:37 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1267.eqiad.wmnet * 16:37 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1267.eqiad.wmnet * 16:31 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1270.eqiad.wmnet with reason: host reimage * 16:24 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1264.eqiad.wmnet * 16:24 kamila@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 16:24 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 16:23 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 16:21 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2162: switch maintenance completed codfw rack b6 * 16:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2185.codfw.wmnet with OS trixie * 16:19 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 16:16 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter1007.eqiad.wmnet * 16:15 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_eqsin * 16:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5024.eqsin.wmnet * 16:13 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5031.eqsin.wmnet * 16:13 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3072.esams.wmnet * 16:12 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter1007.eqiad.wmnet * 16:11 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] (duration: 09m 47s) * 16:10 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1266.eqiad.wmnet with OS trixie * 16:10 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1270 * 16:10 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1270 * 16:09 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1270 * 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1270.eqiad.wmnet 125.48.64.10.in-addr.arpa 5.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1270.eqiad.wmnet 125.48.64.10.in-addr.arpa 5.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1270 - swfrench@cumin1003" * 16:09 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1270 - swfrench@cumin1003" * 16:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3080.esams.wmnet * 16:07 jiji@deploy2003: jiji: Continuing with deployment * 16:06 jiji@deploy2003: jiji: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:04 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 16:04 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1265.eqiad.wmnet * 16:03 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1265.eqiad.wmnet * 16:03 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1265.eqiad.wmnet * 16:03 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1270 * 16:03 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1270.eqiad.wmnet with OS trixie * 16:02 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1270.eqiad.wmnet * 16:02 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] * 16:01 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1270.eqiad.wmnet * 16:01 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1270.eqiad.wmnet * 16:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1113.eqiad.wmnet * 16:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1112.eqiad.wmnet * 15:49 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1266.eqiad.wmnet with reason: host reimage * 15:47 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter1006.eqiad.wmnet * 15:45 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1265.eqiad.wmnet with OS trixie * 15:44 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1266.eqiad.wmnet with reason: host reimage * 15:43 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter1006.eqiad.wmnet * 15:42 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] (duration: 09m 46s) * 15:37 jiji@deploy2003: jiji: Continuing with deployment * 15:36 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2162: switch maintenance completed codfw rack b6 * 15:36 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2161: switch maintenance completed codfw rack b6 * 15:34 jiji@deploy2003: jiji: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:32 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5023.eqsin.wmnet * 15:32 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] * 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5030.eqsin.wmnet * 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3071.esams.wmnet * 15:27 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3079.esams.wmnet * 15:25 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1265.eqiad.wmnet with reason: host reimage * 15:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1266 * 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1266 * 15:21 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1110.eqiad.wmnet * 15:20 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1111.eqiad.wmnet * 15:16 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1265.eqiad.wmnet with reason: host reimage * 15:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1001.eqiad.wmnet with OS bookworm * 15:15 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1266 * 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1266.eqiad.wmnet 76.32.64.10.in-addr.arpa 6.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:15 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1266.eqiad.wmnet 76.32.64.10.in-addr.arpa 6.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1266 - kamila@cumin1003" * 15:15 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1266 - kamila@cumin1003" * 15:07 kamila@cumin1003: START - Cookbook sre.dns.netbox * 15:04 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1266 * 15:04 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1264 * 15:04 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1264 * 15:04 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1266.eqiad.wmnet with OS trixie * 15:03 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1264 * 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1264.eqiad.wmnet 74.32.64.10.in-addr.arpa 4.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:03 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1264.eqiad.wmnet 74.32.64.10.in-addr.arpa 4.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1264 - kamila@cumin1003" * 15:03 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1264 - kamila@cumin1003" * 15:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-eqiad * 15:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp1001.eqiad.wmnet * 15:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp1001.eqiad.wmnet * 15:01 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp1001.eqiad.wmnet * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp1001.eqiad.wmnet * 15:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1376-1384].eqiad.wmnet * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1376-1384].eqiad.wmnet * 14:59 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1002.eqiad.wmnet * 14:59 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1266.eqiad.wmnet * 14:58 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1266.eqiad.wmnet * 14:58 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1266.eqiad.wmnet * 14:58 kamila@cumin1003: START - Cookbook sre.dns.netbox * 14:57 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1264 * 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1265 * 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1265 * 14:57 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1310596{{!}}Set $wgMathInternalRestbaseURL explicitly (T349582)]] * 14:57 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1265 * 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1265.eqiad.wmnet 75.32.64.10.in-addr.arpa 5.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:56 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1265.eqiad.wmnet 75.32.64.10.in-addr.arpa 5.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:56 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:56 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1265 - kamila@cumin1003" * 14:56 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1265 - kamila@cumin1003" * 14:53 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1002.eqiad.wmnet * 14:53 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1376-1384].eqiad.wmnet * 14:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1001.eqiad.wmnet with reason: host reimage * 14:51 kamila@cumin1003: START - Cookbook sre.dns.netbox * 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 14:50 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:50 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2161: switch maintenance completed codfw rack b6 * 14:50 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5021.eqsin.wmnet * 14:50 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1265 * 14:49 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:49 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:49 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1265.eqiad.wmnet with OS trixie * 14:49 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5029.eqsin.wmnet * 14:49 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1265.eqiad.wmnet * 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3070.esams.wmnet * 14:48 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1264.eqiad.wmnet * 14:48 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1001.eqiad.wmnet with reason: host reimage * 14:48 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1376-1384].eqiad.wmnet * 14:48 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1265.eqiad.wmnet * 14:47 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1265.eqiad.wmnet * 14:47 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1264.eqiad.wmnet * 14:47 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1264.eqiad.wmnet * 14:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:47 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3078.esams.wmnet * 14:44 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1263.eqiad.wmnet * 14:44 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1263.eqiad.wmnet * 14:44 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1263.eqiad.wmnet * 14:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1108.eqiad.wmnet * 14:40 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1109.eqiad.wmnet * 14:40 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:35 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:34 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:34 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2006.codfw.wmnet * 14:34 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-flink-eqiad cluster: Roll restart of jvm daemons. * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf2002.codfw.wmnet * 14:31 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1002.eqiad.wmnet * 14:29 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2006.codfw.wmnet * 14:27 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:27 kamila@deploy2003: Finished scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] (duration: 02m 57s) * 14:27 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-flink-eqiad cluster: Roll restart of jvm daemons. * 14:26 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf2002.codfw.wmnet * 14:26 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf2001.codfw.wmnet * 14:25 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1002.eqiad.wmnet * 14:25 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1001.eqiad.wmnet * 14:25 kamila@deploy2003: Started scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] * 14:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:21 kamila@deploy2003: sync-world aborted: Test deployment to check rsync is working - [[phab:T432108|T432108]] (duration: 00m 36s) * 14:21 topranks: reboot lsw1-b6-codfw to upgrade JunOS [[phab:T430922|T430922]] * 14:21 kamila@deploy2003: Started scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] * 14:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1001.eqiad.wmnet with OS bookworm * 14:20 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b6-codfw,lsw1-b6-codfw IPv6,lsw1-b6-codfw.mgmt,ssw1-a[1,8]-codfw with reason: lsw1-b6-codfw JunOS upgrade * 14:20 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf2001.codfw.wmnet * 14:19 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1001.eqiad.wmnet * 14:19 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 26 hosts with reason: lsw1-b6-codfw JunOS upgrade * 14:14 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 14:13 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc2022: switch maintenance codfw rack b6 * 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:12 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.parsercache * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool pc2022: switch maintenance codfw rack b6 * 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2251: switch maintenance codfw rack b6 * 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.parsercache * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2251: switch maintenance codfw rack b6 * 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2162: switch maintenance codfw rack b6 * 14:12 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1263.eqiad.wmnet with OS trixie * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2162: switch maintenance codfw rack b6 * 14:11 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2161: switch maintenance codfw rack b6 * 14:11 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2161: switch maintenance codfw rack b6 * 14:08 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5020.eqsin.wmnet * 14:07 btullis@cumin1003: START - Cookbook sre.hadoop.reboot-workers for Hadoop analytics cluster * 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3069.esams.wmnet * 14:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5028.eqsin.wmnet * 14:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1338-1347].eqiad.wmnet * 14:06 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1338-1347].eqiad.wmnet * 14:05 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3077.esams.wmnet * 14:02 topranks: beginning depools for lsw1-b6-codfw maintenance [[phab:T430922|T430922]] * 14:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1106.eqiad.wmnet * 14:00 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-misc1002.eqiad.wmnet * 13:59 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1338-1347].eqiad.wmnet * 13:59 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1107.eqiad.wmnet * 13:56 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-codfw * 13:55 sfaci@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply * 13:54 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-misc1002.eqiad.wmnet * 13:54 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-misc1001.eqiad.wmnet * 13:54 sfaci@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply * 13:50 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1263.eqiad.wmnet with reason: host reimage * 13:50 sfaci@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 13:49 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1338-1347].eqiad.wmnet * 13:49 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-misc1001.eqiad.wmnet * 13:49 sfaci@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 13:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:49 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:45 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1263.eqiad.wmnet with reason: host reimage * 13:40 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:40 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-eqiad * 13:35 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:34 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:33 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-reboot (exit_code=0) rolling reboot on A:dnsbox and (A:eqsin or A:drmrs or A:magru) and not (P<nowiki>{</nowiki>dns5003*<nowiki>}</nowiki> or P<nowiki>{</nowiki>dns7002*<nowiki>}</nowiki>) and (A:dnsbox) * 13:33 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns7001.wikimedia.org * 13:27 sukhe@dns1004: END - running authdns-update * 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5019.eqsin.wmnet * 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3076.esams.wmnet * 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3068.esams.wmnet * 13:25 sukhe@dns1004: START - running authdns-update * 13:24 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5027.eqsin.wmnet * 13:24 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1263 * 13:24 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1263 * 13:23 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1263 * 13:23 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:23 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1104.eqiad.wmnet * 13:21 kamila@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:21 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:20 kamila@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:20 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:20 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:20 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1263 - kamila@cumin1003" * 13:20 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1263 - kamila@cumin1003" * 13:19 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1105.eqiad.wmnet * 13:19 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:19 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:18 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:18 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns7001.wikimedia.org * 13:16 cdobbins@cumin2003: conftool action : set/pooled=yes; selector: name=dns7002.* * 13:14 cdobbins@dns1004: END - running authdns-update * 13:13 sbisson@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] (duration: 08m 03s) * 13:13 cdobbins@dns1004: START - running authdns-update * 13:12 kamila@cumin1003: START - Cookbook sre.dns.netbox * 13:12 cdobbins@cumin2003: conftool action : set/pooled=yes; selector: name=dns7002.*,service=authdns-update * 13:12 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1263 * 13:11 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1263.eqiad.wmnet with OS trixie * 13:11 cdobbins@cumin2003: conftool action : set/pooled=no; selector: name=dns7002.* * 13:11 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1263.eqiad.wmnet * 13:10 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1263.eqiad.wmnet * 13:10 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1263.eqiad.wmnet * 13:09 sbisson@deploy2003: sbisson: Continuing with deployment * 13:08 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:07 sbisson@deploy2003: sbisson: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:05 sbisson@deploy2003: Started scap sync-world: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] * 13:03 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns6002.wikimedia.org * 13:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1298-1307].eqiad.wmnet * 13:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1298-1307].eqiad.wmnet * 12:59 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1262.eqiad.wmnet * 12:59 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1262.eqiad.wmnet * 12:59 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1262.eqiad.wmnet * 12:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1298-1307].eqiad.wmnet * 12:49 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns6002.wikimedia.org * 12:46 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1298-1307].eqiad.wmnet * 12:46 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:46 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3075.esams.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3067.esams.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5018.eqsin.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5026.eqsin.wmnet * 12:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1102.eqiad.wmnet * 12:39 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1103.eqiad.wmnet * 12:35 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:34 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns6001.wikimedia.org * 12:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:28 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:18 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns6001.wikimedia.org * 12:14 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1267-1276].eqiad.wmnet * 12:13 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1267-1276].eqiad.wmnet * 12:04 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1267-1276].eqiad.wmnet * 12:03 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns5004.wikimedia.org * 12:02 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1100.eqiad.wmnet * 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3066.esams.wmnet * 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3074.esams.wmnet * 12:01 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:01 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-codfw * 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5017.eqsin.wmnet * 12:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5025.eqsin.wmnet * 12:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1101.eqiad.wmnet * 11:59 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1267-1276].eqiad.wmnet * 11:58 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:58 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:54 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns5004.wikimedia.org * 11:54 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and (A:eqsin or A:drmrs or A:magru) and not (P<nowiki>{</nowiki>dns5003*<nowiki>}</nowiki> or P<nowiki>{</nowiki>dns7002*<nowiki>}</nowiki>) and (A:dnsbox) * 11:54 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:53 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:53 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-eqiad * 11:51 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:51 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:50 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:50 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_eqiad * 11:50 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:50 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_eqiad * 11:49 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_esams * 11:49 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_esams * 11:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_eqsin * 11:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_eqsin * 11:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:44 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-eqiad * 11:43 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-codfw * 11:42 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:41 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:24 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-eqiad * 11:23 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-codfw * 11:22 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:15 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:14 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:09 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2066.codfw.wmnet with OS trixie * 11:05 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1151-1160].eqiad.wmnet * 11:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1151-1160].eqiad.wmnet * 10:59 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1068.eqiad.wmnet with OS trixie * 10:55 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.major-upgrade (exit_code=99) * 10:55 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 10:54 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1151-1160].eqiad.wmnet * 10:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2066.codfw.wmnet with reason: host reimage * 10:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1151-1160].eqiad.wmnet * 10:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:42 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2066.codfw.wmnet with reason: host reimage * 10:39 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:37 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:36 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:23 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:22 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2066.codfw.wmnet with OS trixie * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:07 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2065.codfw.wmnet with OS trixie * 10:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:06 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 10:06 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 10:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 10:03 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:03 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 10:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 09:59 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 09:57 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 09:57 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 09:52 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 09:47 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:46 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:46 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2065.codfw.wmnet with reason: host reimage * 09:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2065.codfw.wmnet with reason: host reimage * 09:40 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:39 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 09:39 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:39 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:39 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:37 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:29 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox2003.codfw.wmnet * 09:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox2003.codfw.wmnet * 09:25 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:25 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:24 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:24 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:24 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:21 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:20 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2065.codfw.wmnet with OS trixie * 09:13 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1068.eqiad.wmnet with OS trixie * 09:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2064.codfw.wmnet with OS trixie * 09:08 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 09:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 09:07 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2162: Repooling after switchover * 09:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1067.eqiad.wmnet with OS trixie * 09:00 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:59 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 08:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 08:57 tappof: bump space for prometheus k8s-dse in eqiad * 08:56 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping2004.codfw.wmnet * 08:52 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host ping2004.codfw.wmnet * 08:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 08:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping1004.eqiad.wmnet * 08:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:51 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:49 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 08:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host ping1004.eqiad.wmnet * 08:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2064.codfw.wmnet with reason: host reimage * 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2064.codfw.wmnet with reason: host reimage * 08:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1067.eqiad.wmnet with reason: host reimage * 08:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:33 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1067.eqiad.wmnet with reason: host reimage * 08:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:21 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2162: Repooling after switchover * 08:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2064.codfw.wmnet with OS trixie * 08:16 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1067.eqiad.wmnet with OS trixie * 08:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:15 cgoubert@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-eqiad * 08:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2062.codfw.wmnet with OS trixie * 08:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1066.eqiad.wmnet with OS trixie * 08:02 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2162: Repooling after switchover * 07:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2162: Repooling after switchover * 07:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2162 [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94870 and previous config saved to /var/cache/conftool/dbconfig/20260716-075530-cwilliams.json * 07:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2241 to x3 primary [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94869 and previous config saved to /var/cache/conftool/dbconfig/20260716-075314-cwilliams.json * 07:52 cezmunsta: Starting x3 codfw failover from db2162 to db2241 - [[phab:T430925|T430925]] * 07:50 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:50 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:47 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 07:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2241 with weight 0 [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94868 and previous config saved to /var/cache/conftool/dbconfig/20260716-074507-cwilliams.json * 07:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 18 hosts with reason: Primary switchover x3 [[phab:T430925|T430925]] * 07:43 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1066.eqiad.wmnet with reason: host reimage * 07:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:dse-k8s-worker-eqiad * 07:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1028.eqiad.wmnet * 07:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1028.eqiad.wmnet * 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 07:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1066.eqiad.wmnet with reason: host reimage * 07:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1028.eqiad.wmnet * 07:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1028.eqiad.wmnet * 07:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1027.eqiad.wmnet * 07:35 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1027.eqiad.wmnet * 07:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1027.eqiad.wmnet * 07:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1027.eqiad.wmnet * 07:28 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1026.eqiad.wmnet * 07:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1026.eqiad.wmnet * 07:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast2003.wikimedia.org * 07:21 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1026.eqiad.wmnet * 07:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1066.eqiad.wmnet with OS trixie * 07:19 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast2003.wikimedia.org * 07:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2062.codfw.wmnet with OS trixie * 06:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1026.eqiad.wmnet * 06:51 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1025.eqiad.wmnet * 06:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1025.eqiad.wmnet * 06:47 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 06:47 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 06:44 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1025.eqiad.wmnet * 06:14 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1025.eqiad.wmnet * 06:14 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1024.eqiad.wmnet * 06:14 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1024.eqiad.wmnet * 06:07 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1024.eqiad.wmnet * 05:37 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1024.eqiad.wmnet * 05:37 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1023.eqiad.wmnet * 05:37 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1023.eqiad.wmnet * 05:26 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1023.eqiad.wmnet * 04:56 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1023.eqiad.wmnet * 04:56 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1022.eqiad.wmnet * 04:56 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1022.eqiad.wmnet * 04:49 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1022.eqiad.wmnet * 04:19 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1022.eqiad.wmnet * 04:19 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1021.eqiad.wmnet * 04:19 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1021.eqiad.wmnet * 04:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1021.eqiad.wmnet * 03:38 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1021.eqiad.wmnet * 03:38 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1020.eqiad.wmnet * 03:38 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1020.eqiad.wmnet * 03:20 btullis@cumin1003: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1020.eqiad.wmnet * 03:18 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1020.eqiad.wmnet * 03:18 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1019.eqiad.wmnet * 03:18 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1019.eqiad.wmnet * 03:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1019.eqiad.wmnet * 02:41 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1019.eqiad.wmnet * 02:41 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1018.eqiad.wmnet * 02:41 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1018.eqiad.wmnet * 02:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling both afterwards * 02:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2003.codfw.wmnet -> wcqs2001.codfw.wmnet, repooling both afterwards * 02:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1018.eqiad.wmnet * 02:30 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1018.eqiad.wmnet * 02:30 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1014.eqiad.wmnet * 02:30 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1014.eqiad.wmnet * 02:24 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1014.eqiad.wmnet * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 01:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1014.eqiad.wmnet * 01:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1013.eqiad.wmnet * 01:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1013.eqiad.wmnet * 01:47 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1013.eqiad.wmnet * 01:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2003.codfw.wmnet -> wcqs2001.codfw.wmnet, repooling both afterwards * 01:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling both afterwards * 01:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1013.eqiad.wmnet * 01:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1012.eqiad.wmnet * 01:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1012.eqiad.wmnet * 01:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1012.eqiad.wmnet * 01:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1012.eqiad.wmnet * 01:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1011.eqiad.wmnet * 01:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1011.eqiad.wmnet * 01:08 ryankemper@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] scap deploy post bookworm reimage (duration: 00m 23s) * 01:08 ryankemper@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] scap deploy post bookworm reimage * 01:08 ryankemper@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): scap deploy post bookworm reimage (duration: 00m 46s) * 01:07 ryankemper@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): scap deploy post bookworm reimage * 01:04 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1011.eqiad.wmnet * 01:04 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1011.eqiad.wmnet * 01:04 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1010.eqiad.wmnet * 01:04 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1010.eqiad.wmnet * 00:57 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1010.eqiad.wmnet * 00:57 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1010.eqiad.wmnet * 00:57 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1009.eqiad.wmnet * 00:57 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1009.eqiad.wmnet * 00:50 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1009.eqiad.wmnet * 00:20 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1009.eqiad.wmnet * 00:20 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1008.eqiad.wmnet * 00:20 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1008.eqiad.wmnet * 00:13 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1008.eqiad.wmnet == 2026-07-15 == * 23:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2001.codfw.wmnet with OS bookworm * 23:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1008.eqiad.wmnet * 23:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1007.eqiad.wmnet * 23:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1007.eqiad.wmnet * 23:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1007.eqiad.wmnet * 23:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1007.eqiad.wmnet * 23:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1006.eqiad.wmnet * 23:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1006.eqiad.wmnet * 23:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1006.eqiad.wmnet * 23:29 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1006.eqiad.wmnet * 23:28 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1005.eqiad.wmnet * 23:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1005.eqiad.wmnet * 23:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1002.eqiad.wmnet with OS bookworm * 23:21 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1005.eqiad.wmnet * 23:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2001.codfw.wmnet with reason: host reimage * 23:15 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host datahubsearch1001.eqiad.wmnet with OS bookworm * 23:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2001.codfw.wmnet with reason: host reimage * 23:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 23:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 22:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 22:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1005.eqiad.wmnet * 22:51 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1004.eqiad.wmnet * 22:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1004.eqiad.wmnet * 22:45 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1004.eqiad.wmnet * 22:44 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 22:44 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS trixie * 22:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host datahubsearch1001.eqiad.wmnet with OS bookworm * 22:34 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host datahubsearch1001.eqiad.wmnet with OS bookworm * 22:16 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on datahubsearch[1002-1003].eqiad.wmnet with reason: Using datahubsearch1001 to test bookworm reimages * 22:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1004.eqiad.wmnet * 22:15 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1003.eqiad.wmnet * 22:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1003.eqiad.wmnet * 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 22:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1003.eqiad.wmnet * 22:08 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1003.eqiad.wmnet * 22:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1002.eqiad.wmnet * 22:08 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1002.eqiad.wmnet * 22:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host datahubsearch1001.eqiad.wmnet with OS bookworm * 22:05 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 22:02 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm * 22:01 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on datahubsearch[1001-1003].eqiad.wmnet with reason: Using datahubsearch1001 to test bookworm reimages * 22:01 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1002.eqiad.wmnet * 22:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1002.eqiad.wmnet * 22:00 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1001.eqiad.wmnet * 22:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1001.eqiad.wmnet * 21:53 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1001.eqiad.wmnet * 21:52 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 21:50 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 21:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS trixie * 21:50 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS bookworm * 21:43 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 21:38 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:30 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wcqs1002'] * 21:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:29 lerickson@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 21:29 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:29 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:29 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS bookworm * 21:28 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 21:28 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm * 21:23 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1001.eqiad.wmnet * 21:23 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:23 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:22 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 21:20 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:18 swfrench-wmf: reprepro include php8.3_8.3.32-1+wmf11u2 into component/php83 for bullseye-wikimedia * 21:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:16 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:15 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-druid-public cluster: Roll restart of jvm daemons. * 21:08 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:05 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1001.eqiad.wmnet * 21:05 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1001.eqiad.wmnet * 21:04 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-druid-public cluster: Roll restart of jvm daemons. * 21:02 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 21:01 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 21:01 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 21:00 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 20:59 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1001.eqiad.wmnet * 20:59 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1001.eqiad.wmnet * 20:59 btullis@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:dse-k8s-worker-eqiad * 20:55 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 20:55 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 20:45 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm * 20:21 jhathaway: puppet is re-enabled, have fun, but not too much fun! * 20:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 20:17 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs2001'] * 20:12 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs2001'] * 20:11 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs2001'] * 20:09 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:08 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 20:05 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:05 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 20:04 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs2001'] * 20:03 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 20:03 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm * 20:02 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:02 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 20:01 jhathaway: disabling puppet fleet wide to roll out kafka patch * 19:55 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host relforge1010.eqiad.wmnet * 19:52 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 19:52 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 19:48 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 19:48 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 19:48 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 19:47 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 19:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 19:45 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 19:45 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 19:44 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1010.eqiad.wmnet * 19:38 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1262.eqiad.wmnet with OS trixie * 19:17 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1262.eqiad.wmnet with reason: host reimage * 19:11 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1262.eqiad.wmnet with reason: host reimage * 18:59 cdobbins@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS trixie * 18:54 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 18:53 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 18:52 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1262 * 18:52 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1262 * 18:51 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1262 * 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1262.eqiad.wmnet 72.32.64.10.in-addr.arpa 2.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:51 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1262.eqiad.wmnet 72.32.64.10.in-addr.arpa 2.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1262 - kamila@cumin1003" * 18:51 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1262 - kamila@cumin1003" * 18:46 kamila@cumin1003: START - Cookbook sre.dns.netbox * 18:46 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1262 * 18:46 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 18:46 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ncmonitor1001.eqiad.wmnet * 18:46 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 18:45 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1262.eqiad.wmnet with OS trixie * 18:45 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 18:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1262.eqiad.wmnet * 18:44 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1262.eqiad.wmnet * 18:44 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1262.eqiad.wmnet * 18:42 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host ncmonitor1001.eqiad.wmnet * 18:29 topranks: pull power on cr1-eqiad to install new switch-control boards [[phab:T426343|T426343]] * 18:29 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs[1018-1020].eqiad.wmnet with reason: line card install in cr1-eqiad * 18:27 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 14 hosts with reason: linecard install in cr1-eqad * 18:22 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_ulsfo * 18:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4052.ulsfo.wmnet * 18:19 cdobbins@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 18:15 cdobbins@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 18:14 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_drmrs * 18:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6016.drmrs.wmnet * 18:12 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_ulsfo * 18:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4044.ulsfo.wmnet * 18:10 sukhe@cumin1003: END (ERROR) - Cookbook sre.cdn.roll-reboot (exit_code=97) rolling reboot on A:cp-upload_drmrs * 18:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2241: Security update * 17:56 topranks: start draining traffic on cr1-eqiad ahead of line card installation [[phab:T426343|T426343]] * 17:47 cdobbins@cumin2003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie * 17:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4051.ulsfo.wmnet * 17:40 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:39 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 17:34 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6007.drmrs.wmnet * 17:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6015.drmrs.wmnet * 17:32 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:31 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 17:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4043.ulsfo.wmnet * 17:27 lerickson@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:25 lerickson@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 17:22 lerickson@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-codfw * 17:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp2001.codfw.wmnet * 17:22 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 17:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp2001.codfw.wmnet * 17:22 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 17:19 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2241: Security update * 17:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2241.codfw.wmnet * 17:17 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2241.codfw.wmnet * 17:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp2001.codfw.wmnet * 17:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp2001.codfw.wmnet * 17:15 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2366-2374].codfw.wmnet * 17:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2366-2374].codfw.wmnet * 17:10 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply * 17:10 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply * 17:08 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2366-2374].codfw.wmnet * 17:06 sukhe: sre.dns.roll-reboot to resume later * 17:06 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-reboot (exit_code=97) rolling reboot on A:dnsbox and not (A:ulsfo or A:magru) and (A:dnsbox) * 17:06 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns5003.wikimedia.org * 17:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2241: Security update * 17:03 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2241: Security update * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply * 17:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2366-2374].codfw.wmnet * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply * 17:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2357-2365].codfw.wmnet * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 17:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2357-2365].codfw.wmnet * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply * 16:55 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2357-2365].codfw.wmnet * 16:55 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 16:53 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6006.drmrs.wmnet * 16:52 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6014.drmrs.wmnet * 16:52 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 16:51 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 16:50 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2357-2365].codfw.wmnet * 16:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4042.ulsfo.wmnet * 16:50 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2347-2356].codfw.wmnet * 16:50 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2347-2356].codfw.wmnet * 16:49 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns5003.wikimedia.org * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply * 16:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4050.ulsfo.wmnet * 16:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2347-2356].codfw.wmnet * 16:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2347-2356].codfw.wmnet * 16:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2337-2346].codfw.wmnet * 16:36 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2337-2346].codfw.wmnet * 16:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:dse-k8s-worker-codfw * 16:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2003.codfw.wmnet * 16:35 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2003.codfw.wmnet * 16:34 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns3004.wikimedia.org * 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply * 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply * 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply * 16:30 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply * 16:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2003.codfw.wmnet * 16:29 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2337-2346].codfw.wmnet * 16:24 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2003.codfw.wmnet * 16:24 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2002.codfw.wmnet * 16:24 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2002.codfw.wmnet * 16:23 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns3004.wikimedia.org * 16:23 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2337-2346].codfw.wmnet * 16:23 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2327-2336].codfw.wmnet * 16:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2327-2336].codfw.wmnet * 16:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2002.codfw.wmnet * 16:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2327-2336].codfw.wmnet * 16:12 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2002.codfw.wmnet * 16:12 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2001.codfw.wmnet * 16:12 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2001.codfw.wmnet * 16:12 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1065.eqiad.wmnet with OS trixie * 16:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6005.drmrs.wmnet * 16:11 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6013.drmrs.wmnet * 16:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4041.ulsfo.wmnet * 16:08 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns3003.wikimedia.org * 16:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2327-2336].codfw.wmnet * 16:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2317-2326].codfw.wmnet * 16:06 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2317-2326].codfw.wmnet * 16:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2001.codfw.wmnet * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply * 16:03 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4049.ulsfo.wmnet * 16:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2001.codfw.wmnet * 16:00 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test2001.codfw.wmnet * 16:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test2001.codfw.wmnet * 16:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2063.codfw.wmnet with OS trixie * 15:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2317-2326].codfw.wmnet * 15:57 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns3003.wikimedia.org * 15:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test2001.codfw.wmnet * 15:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test2001.codfw.wmnet * 15:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2004.codfw.wmnet * 15:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2004.codfw.wmnet * 15:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2317-2326].codfw.wmnet * 15:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2307-2316].codfw.wmnet * 15:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2307-2316].codfw.wmnet * 15:49 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2004.codfw.wmnet * 15:48 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2004.codfw.wmnet * 15:48 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2003.codfw.wmnet * 15:48 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2003.codfw.wmnet * 15:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 15:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2307-2316].codfw.wmnet * 15:42 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2003.codfw.wmnet * 15:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 15:42 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2003.codfw.wmnet * 15:42 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2002.codfw.wmnet * 15:42 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2002.codfw.wmnet * 15:42 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2006.wikimedia.org * 15:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2063.codfw.wmnet with reason: host reimage * 15:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2307-2316].codfw.wmnet * 15:37 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2297-2306].codfw.wmnet * 15:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2297-2306].codfw.wmnet * 15:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2002.codfw.wmnet * 15:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2002.codfw.wmnet * 15:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2001.codfw.wmnet * 15:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2001.codfw.wmnet * 15:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2063.codfw.wmnet with reason: host reimage * 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6004.drmrs.wmnet * 15:31 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2001.codfw.wmnet * 15:31 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2001.codfw.wmnet * 15:31 btullis@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:dse-k8s-worker-codfw * 15:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6012.drmrs.wmnet * 15:28 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2006.wikimedia.org * 15:27 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-analytics cluster: Roll restart of jvm daemons. * 15:27 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2297-2306].codfw.wmnet * 15:27 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4040.ulsfo.wmnet * 15:24 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1065.eqiad.wmnet with OS trixie * 15:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4048.ulsfo.wmnet * 15:21 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-analytics cluster: Roll restart of jvm daemons. * 15:21 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2297-2306].codfw.wmnet * 15:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2287-2296].codfw.wmnet * 15:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2287-2296].codfw.wmnet * 15:20 btullis@cumin1003: END (PASS) - Cookbook sre.druid.reboot-workers (exit_code=0) for Druid public cluster: Reboot Druid nodes * 15:18 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 15:17 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm * 15:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2063.codfw.wmnet with OS trixie * 15:13 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2005.wikimedia.org * 15:11 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2287-2296].codfw.wmnet * 15:11 btullis@cumin1003: END (PASS) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=0) rolling reboot on A:cephosd-eqiad * 15:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1064.eqiad.wmnet with OS trixie * 15:05 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2062.codfw.wmnet with OS trixie * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2287-2296].codfw.wmnet * 15:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2277-2286].codfw.wmnet * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2277-2286].codfw.wmnet * 14:59 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2005.wikimedia.org * 14:57 brouberol@cumin1003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-jumbo-eqiad * 14:52 btullis@cumin1003: END (PASS) - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas (exit_code=0) rolling reboot on A:schema-codfw * 14:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6003.drmrs.wmnet * 14:50 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:50 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host relforge1009.eqiad.wmnet * 14:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2277-2286].codfw.wmnet * 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6011.drmrs.wmnet * 14:47 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 14:46 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>ml-serve1001.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 14:46 klausman@cumin1003: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) pool for host ml-serve1001.eqiad.wmnet * 14:46 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 14:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1001.eqiad.wmnet * 14:45 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4039.ulsfo.wmnet * 14:44 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1009.eqiad.wmnet * 14:44 btullis@cumin1003: START - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas rolling reboot on A:schema-codfw * 14:44 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2004.wikimedia.org * 14:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2277-2286].codfw.wmnet * 14:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2267-2276].codfw.wmnet * 14:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2267-2276].codfw.wmnet * 14:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 14:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4047.ulsfo.wmnet * 14:40 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1001.eqiad.wmnet * 14:38 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 14:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 14:36 topranks: disconnect power on cr2-eqiad to shut down device for switch fabric replacement [[phab:T426343|T426343]] * 14:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2267-2276].codfw.wmnet * 14:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 14:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1001.eqiad.wmnet * 14:35 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>ml-serve1001.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 14:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 14:34 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 14:33 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:33 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:30 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2004.wikimedia.org * 14:29 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2267-2276].codfw.wmnet * 14:29 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2257-2266].codfw.wmnet * 14:29 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2257-2266].codfw.wmnet * 14:24 btullis@cumin1003: END (PASS) - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas (exit_code=0) rolling reboot on A:schema-eqiad * 14:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2257-2266].codfw.wmnet * 14:20 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:20 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:19 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:17 jforrester@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2257-2266].codfw.wmnet * 14:16 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:16 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2062.codfw.wmnet with OS trixie * 14:15 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1064.eqiad.wmnet with OS trixie * 14:15 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1006.wikimedia.org * 14:15 btullis@cumin1003: START - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas rolling reboot on A:schema-eqiad * 14:14 topranks: switch routing-engine on cr2-eqiad resetting all interfaces [[phab:T417873|T417873]] * 14:11 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:11 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:10 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm * 14:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6002.drmrs.wmnet * 14:09 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6010.drmrs.wmnet * 14:06 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1006.wikimedia.org * 14:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:05 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 14:05 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4038.ulsfo.wmnet * 14:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4046.ulsfo.wmnet * 14:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:00 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on cr1-eqiad with reason: switch upgrade and line card install * 14:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:59 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:57 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:57 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:55 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-eqiad * 13:55 btullis@cumin1003: START - Cookbook sre.druid.reboot-workers for Druid public cluster: Reboot Druid nodes * 13:53 topranks: switch routing-engine on cr2-eqiad resetting all interfaces [[phab:T417873|T417873]] * 13:51 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1005.wikimedia.org * 13:50 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:49 brouberol@cumin1003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-test-eqiad * 13:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:44 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2197-2206].codfw.wmnet * 13:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2197-2206].codfw.wmnet * 13:36 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1005.wikimedia.org * 13:35 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2197-2206].codfw.wmnet * 13:30 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2197-2206].codfw.wmnet * 13:28 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6001.drmrs.wmnet * 13:28 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6009.drmrs.wmnet * 13:28 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2187-2196].codfw.wmnet * 13:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2187-2196].codfw.wmnet * 13:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 13:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4037.ulsfo.wmnet * 13:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2001 * 13:22 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2001 * 13:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4045.ulsfo.wmnet * 13:21 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1004.wikimedia.org * 13:19 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on lvs[1018-1020].eqiad.wmnet with reason: switch upgrade and line card install * 13:18 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2009.codfw.wmnet * 13:18 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2001 * 13:18 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2001.codfw.wmnet 26.16.192.10.in-addr.arpa 6.2.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:17 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2001.codfw.wmnet 26.16.192.10.in-addr.arpa 6.2.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:17 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:17 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2001 - bking@cumin2003" * 13:17 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2001 - bking@cumin2003" * 13:17 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2009.codfw.wmnet * 13:17 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_drmrs * 13:17 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2187-2196].codfw.wmnet * 13:17 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_drmrs * 13:17 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on 15 hosts with reason: switch upgrade and line card install * 13:17 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:15 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 13:13 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:13 brouberol@cumin1003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-jumbo-eqiad * 13:13 brouberol@cumin1003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-test-eqiad * 13:13 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1004.wikimedia.org * 13:13 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and not (A:ulsfo or A:magru) and (A:dnsbox) * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:12 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_ulsfo * 13:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:12 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_ulsfo * 13:11 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2187-2196].codfw.wmnet * 13:11 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 13:11 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 13:06 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 13:05 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling source-only afterwards * 13:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2001 * 13:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2009.codfw.wmnet with OS trixie * 13:03 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:03 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling source-only afterwards * 13:01 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 15s) * 13:01 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 13:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 12:57 btullis@cumin1003: END (PASS) - Cookbook sre.druid.reboot-workers (exit_code=0) for Druid analytics cluster: Reboot Druid nodes * 12:54 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 12:54 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2163-2172].codfw.wmnet * 12:54 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2163-2172].codfw.wmnet * 12:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2163-2172].codfw.wmnet * 12:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2009.codfw.wmnet with reason: host reimage * 12:41 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2163-2172].codfw.wmnet * 12:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2153-2162].codfw.wmnet * 12:40 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2153-2162].codfw.wmnet * 12:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2009.codfw.wmnet with reason: host reimage * 12:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2153-2162].codfw.wmnet * 12:29 btullis@cumin1003: END (PASS) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=0) rolling reboot on A:cephosd-codfw * 12:25 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2153-2162].codfw.wmnet * 12:25 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2143-2152].codfw.wmnet * 12:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2143-2152].codfw.wmnet * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2009 * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2009 * 12:22 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2009 * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2009.codfw.wmnet 139.0.192.10.in-addr.arpa 9.3.1.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:22 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2009.codfw.wmnet 139.0.192.10.in-addr.arpa 9.3.1.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2009 - mvernon@cumin2003" * 12:22 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2009 - mvernon@cumin2003" * 12:16 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 12:15 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 12:15 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 12:15 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2009 * 12:15 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 12:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2009.codfw.wmnet with OS trixie * 12:15 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 12:14 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 12:14 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2143-2152].codfw.wmnet * 12:13 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 12:12 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2010.codfw.wmnet * 12:11 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2010.codfw.wmnet * 12:10 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 12:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2143-2152].codfw.wmnet * 12:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2133-2142].codfw.wmnet * 12:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2133-2142].codfw.wmnet * 12:02 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 11:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2133-2142].codfw.wmnet * 11:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2133-2142].codfw.wmnet * 11:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:49 mvolz@deploy2003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:49 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-codfw * 11:48 mvolz@deploy2003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:47 btullis@cumin1003: START - Cookbook sre.druid.reboot-workers for Druid analytics cluster: Reboot Druid nodes * 11:46 mvolz@deploy2003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:46 mvolz@deploy2003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:45 mvolz@deploy2003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:44 mvolz@deploy2003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:40 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] (duration: 11m 38s) * 11:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2010.codfw.wmnet with OS trixie * 11:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2105-2114].codfw.wmnet * 11:36 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2105-2114].codfw.wmnet * 11:36 krinkle@deploy2003: physikerwelt, krinkle: Continuing with deployment * 11:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1018: Security updates * 11:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:36 root@cumin1003: START - Cookbook sre.mysql.parsercache * 11:36 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1018: Security updates * 11:31 krinkle@deploy2003: physikerwelt, krinkle: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:29 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] * 11:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2105-2114].codfw.wmnet * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2105-2114].codfw.wmnet * 11:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2010.codfw.wmnet with reason: host reimage * 11:12 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2010.codfw.wmnet with reason: host reimage * 11:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1018: Security updates * 11:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:10 root@cumin1003: START - Cookbook sre.mysql.parsercache * 11:10 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1018: Security updates * 11:09 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1009.eqiad.wmnet with OS trixie * 11:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow7002.magru.wmnet * 11:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 11:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 11:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=tegola-vector-tiles,name=eqiad * 11:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=kartotherian,name=eqiad * 11:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow7002.magru.wmnet * 10:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2010 * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2010 * 10:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 10:54 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 10:54 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2010 * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2010.codfw.wmnet 76.16.192.10.in-addr.arpa 6.7.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:54 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2010.codfw.wmnet 76.16.192.10.in-addr.arpa 6.7.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2010 - mvernon@cumin2003" * 10:54 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2010 - mvernon@cumin2003" * 10:49 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 10:49 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2010 * 10:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1009.eqiad.wmnet with reason: host reimage * 10:49 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2010.codfw.wmnet with OS trixie * 10:46 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2011.codfw.wmnet * 10:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1011.eqiad.wmnet * 10:44 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2011.codfw.wmnet * 10:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 10:44 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 10:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1009.eqiad.wmnet with reason: host reimage * 10:44 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow6001.drmrs.wmnet * 10:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1017: Security updates * 10:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:39 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1017: Security updates * 10:39 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow6001.drmrs.wmnet * 10:38 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1011.eqiad.wmnet * 10:35 cgoubert@deploy2003: Finished deploy [restbase/deploy@06301bd]: Deploying {{Gerrit|1306088}} {{Gerrit|1308347}} - [[phab:T429944|T429944]] [[phab:T428279|T428279]] (duration: 28m 34s) * 10:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1012.eqiad.wmnet * 10:35 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 10:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow5003.eqsin.wmnet * 10:34 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2011.codfw.wmnet with OS trixie * 10:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1009.eqiad.wmnet with OS trixie * 10:28 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1012.eqiad.wmnet * 10:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1013.eqiad.wmnet * 10:27 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:27 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow5003.eqsin.wmnet * 10:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:26 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow4003.ulsfo.wmnet * 10:25 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 10:25 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 10:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow4003.ulsfo.wmnet * 10:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1013.eqiad.wmnet * 10:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1014.eqiad.wmnet * 10:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:15 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2011.codfw.wmnet with reason: host reimage * 10:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1017: Security updates * 10:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:14 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:14 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1017: Security updates * 10:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow3004.esams.wmnet * 10:11 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2011.codfw.wmnet with reason: host reimage * 10:10 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:10 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1014.eqiad.wmnet * 10:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki2003.codfw.wmnet * 10:09 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 10:09 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 10:09 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow3004.esams.wmnet * 10:08 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2004.codfw.wmnet * 10:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1010.eqiad.wmnet with OS trixie * 10:07 cgoubert@deploy2003: Started deploy [restbase/deploy@06301bd]: Deploying {{Gerrit|1306088}} {{Gerrit|1308347}} - [[phab:T429944|T429944]] [[phab:T428279|T428279]] * 10:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host rpki2003.codfw.wmnet * 10:04 topranks: push out config change to BGP_outfilter on core routers [[phab:T431849|T431849]] * 10:02 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow2004.codfw.wmnet * 09:59 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 09:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2003.codfw.wmnet * 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2011 * 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2011 * 09:53 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 09:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:52 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2011 * 09:52 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2011.codfw.wmnet 36.32.192.10.in-addr.arpa 6.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:52 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2011.codfw.wmnet 36.32.192.10.in-addr.arpa 6.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:51 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:51 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2011 - mvernon@cumin2003" * 09:51 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2011 - mvernon@cumin2003" * 09:51 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow2003.codfw.wmnet * 09:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1003.eqiad.wmnet * 09:49 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 09:49 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 09:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1010.eqiad.wmnet with reason: host reimage * 09:47 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 09:47 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 09:47 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 09:47 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2011 * 09:46 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2011.codfw.wmnet with OS trixie * 09:44 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow1003.eqiad.wmnet * 09:44 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2012.codfw.wmnet * 09:44 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1002.eqiad.wmnet * 09:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1010.eqiad.wmnet with reason: host reimage * 09:43 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2012.codfw.wmnet * 09:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Security updates * 09:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:43 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:43 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Security updates * 09:42 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:40 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow1002.eqiad.wmnet * 09:40 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 09:37 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki1001.eqiad.wmnet * 09:36 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 09:36 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 09:33 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host rpki1001.eqiad.wmnet * 09:32 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:32 cgoubert@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-codfw * 09:31 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=kartotherian,name=eqiad * 09:31 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola-vector-tiles,name=eqiad * 09:31 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 09:31 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2012.codfw.wmnet with OS trixie * 09:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1010.eqiad.wmnet with OS trixie * 09:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1011.eqiad.wmnet with OS trixie * 09:21 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Security updates * 09:21 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:21 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:21 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Security updates * 09:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2012.codfw.wmnet with reason: host reimage * 09:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1011.eqiad.wmnet with reason: host reimage * 09:08 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2012.codfw.wmnet with reason: host reimage * 09:05 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1011.eqiad.wmnet with reason: host reimage * 08:55 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:52 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1011.eqiad.wmnet with OS trixie * 08:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1022: Security updates * 08:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2012 * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2012 * 08:50 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1022: Security updates * 08:50 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2012 * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2012.codfw.wmnet 44.48.192.10.in-addr.arpa 4.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:50 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2012.codfw.wmnet 44.48.192.10.in-addr.arpa 4.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2012 - mvernon@cumin2003" * 08:50 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2012 - mvernon@cumin2003" * 08:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1012.eqiad.wmnet with OS trixie * 08:44 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 08:44 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2012 * 08:43 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2012.codfw.wmnet with OS trixie * 08:42 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2013.codfw.wmnet * 08:41 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2013.codfw.wmnet * 08:35 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 08:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host krb1002.eqiad.wmnet * 08:30 elukey@dns1004: END - running authdns-update * 08:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1012.eqiad.wmnet with reason: host reimage * 08:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Security updates * 08:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:28 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:28 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Security updates * 08:27 elukey@dns1004: START - running authdns-update * 08:26 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 08:26 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host krb1002.eqiad.wmnet * 08:22 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1012.eqiad.wmnet with reason: host reimage * 08:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host krb2002.codfw.wmnet * 08:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast6003.wikimedia.org * 08:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2013.codfw.wmnet with OS trixie * 08:13 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast6003.wikimedia.org * 08:12 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast3007.wikimedia.org * 08:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host krb2002.codfw.wmnet * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Security updates * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:09 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:09 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Security updates * 08:07 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1012.eqiad.wmnet with OS trixie * 08:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast3007.wikimedia.org * 08:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast5005.wikimedia.org * 07:58 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast5005.wikimedia.org * 07:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1013.eqiad.wmnet with OS trixie * 07:53 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2013.codfw.wmnet with reason: host reimage * 07:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1021: Security updates * 07:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:53 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:53 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1021: Security updates * 07:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast1004.wikimedia.org * 07:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2013.codfw.wmnet with reason: host reimage * 07:46 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast1004.wikimedia.org * 07:40 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1013.eqiad.wmnet with reason: host reimage * 07:36 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1013.eqiad.wmnet with reason: host reimage * 07:31 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2013 * 07:31 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2013 * 07:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1021: Security updates * 07:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:30 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:30 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1021: Security updates * 07:24 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2013 * 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2013.codfw.wmnet 87.0.192.10.in-addr.arpa 7.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:24 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2013.codfw.wmnet 87.0.192.10.in-addr.arpa 7.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2013 - mvernon@cumin2003" * 07:24 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2013 - mvernon@cumin2003" * 07:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1013.eqiad.wmnet with OS trixie * 07:19 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 07:19 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2013 * 07:19 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2013.codfw.wmnet with OS trixie * 07:13 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] (duration: 07m 48s) * 07:09 kharlan@deploy2003: kharlan: Continuing with deployment * 07:08 kharlan@deploy2003: kharlan: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:06 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 01:15 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 01:14 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply == 2026-07-14 == * 22:51 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_magru * 22:51 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7016.magru.wmnet * 22:46 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_magru * 22:46 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7008.magru.wmnet * 22:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7015.magru.wmnet * 22:04 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7007.magru.wmnet * 21:29 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7014.magru.wmnet * 21:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7006.magru.wmnet * 21:13 dzahn@cumin2002: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 0:15:00 on gerrit.wikimedia.org with reason: reboot * 21:11 mutante: gerrit2003 (gerrit.wikimedia.org) - reboot for maintenance * 21:11 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on gerrit2003.wikimedia.org with reason: reboot * 20:56 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:56 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:56 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:55 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 20:48 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7013.magru.wmnet * 20:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7005.magru.wmnet * 20:41 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host phab1005.eqiad.wmnet with OS trixie * 20:28 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] (duration: 06m 47s) * 20:24 sbassett@deploy2003: sbassett: Continuing with deployment * 20:23 sbassett@deploy2003: sbassett: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:23 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on phab1005.eqiad.wmnet with reason: host reimage * 20:21 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] * 20:20 aokoth@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on phab1005.eqiad.wmnet with reason: host reimage * 20:12 jhuneidi@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] (duration: 07m 42s) * 20:07 jhuneidi@deploy2003: jhuneidi, priyankar22: Continuing with deployment * 20:06 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7012.magru.wmnet * 20:06 jhuneidi@deploy2003: jhuneidi, priyankar22: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:04 jhuneidi@deploy2003: Started scap sync-world: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] * 20:02 aokoth@cumin1003: START - Cookbook sre.hosts.reimage for host phab1005.eqiad.wmnet with OS trixie * 20:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7004.magru.wmnet * 20:00 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet * 19:57 aokoth@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet * 19:24 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7011.magru.wmnet * 19:19 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7003.magru.wmnet * 19:11 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] (duration: 08m 33s) * 19:07 jforrester@deploy2003: jforrester: Continuing with deployment * 19:04 jforrester@deploy2003: jforrester: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:02 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] * 18:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7010.magru.wmnet * 18:38 mutante: rotating phabricator-gerrit bot token (its-phabricator) * 18:18 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 17:44 swfrench@deploy2003: Finished scap sync-world: Deployment to pick up new production image (duration: 31m 44s) * 17:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7002.magru.wmnet * 17:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7009.magru.wmnet * 17:32 swfrench@deploy2003: swfrench: Continuing with deployment * 17:29 swfrench@deploy2003: swfrench: Deployment to pick up new production image synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:17 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2035: repooling after rack b5 maintenance * 17:16 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool es2035: repooling after rack b5 maintenance * 17:16 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2188: repooling after rack b5 maintenance * 17:12 swfrench@deploy2003: Started scap sync-world: Deployment to pick up new production image * 17:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7001.magru.wmnet * 16:57 swfrench-wmf: reprepro include php8.3_8.3.32-1+wmf12u2 into component/php83 for bookworm-wikimedia * 16:50 sukhe: pool cp2046 * 16:47 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4039.ulsfo.wmnet * 16:44 sukhe: sudo cumin -b31 "A:cp" "run-puppet-agent" * 16:33 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on contint1003.wikimedia.org with reason: reboot * 16:32 mutante: contint1003 - main CI server - rebooting * 16:31 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2188: repooling after rack b5 maintenance * 16:31 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2178: repooling after rack b5 maintenance * 16:29 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 16:28 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 16:28 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 16:28 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 16:18 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2014.codfw.wmnet * 16:18 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 16:17 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2014.codfw.wmnet * 16:10 mvernon@cumin1003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-thanos-proxies (exit_code=0) rolling restart_daemons on A:thanos-fe * 16:09 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 16:07 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp4039.ulsfo.wmnet * 16:07 mvernon@cumin1003: START - Cookbook sre.swift.roll-restart-reboot-swift-thanos-proxies rolling restart_daemons on A:thanos-fe * 16:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2014.codfw.wmnet with OS trixie * 15:56 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1014.eqiad.wmnet with OS trixie * 15:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2014.codfw.wmnet with reason: host reimage * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2014 * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2014 * 15:28 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2014 * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2014.codfw.wmnet 194.16.192.10.in-addr.arpa 4.9.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:28 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2014.codfw.wmnet 194.16.192.10.in-addr.arpa 4.9.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2014 - mvernon@cumin2003" * 15:28 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2014 - mvernon@cumin2003" * 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Apply title-related policies when selecting the name of the entity - kamila@cumin1003" * 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Apply title-related policies when selecting the name of the entity - kamila@cumin1003 * 15:22 kamila@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Apply title-related policies when selecting the name of the entity - kamila@cumin1003 * 15:22 kamila@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Apply title-related policies when selecting the name of the entity - kamila@cumin1003" * 15:20 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 15:20 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2014 * 15:20 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2014.codfw.wmnet with OS trixie * 15:19 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1014.eqiad.wmnet with OS trixie * 15:01 dancy@deploy2003: Installation of scap version "4.274.1" completed for 3 hosts * 15:00 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2177: repooling after rack b5 maintenance * 15:00 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2159: repooling after rack b5 maintenance * 14:59 dancy@deploy2003: Installing scap version "4.274.1" for 3 host(s) * 14:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2015.codfw.wmnet with OS trixie * 14:54 seanleong-wmde: Finished populateSitesTable for isvwiki ([[phab:T429939|T429939]]) * 14:53 javiermonton@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] (duration: 07m 35s) * 14:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1015.eqiad.wmnet with OS trixie * 14:49 javiermonton@deploy2003: javiermonton: Continuing with deployment * 14:48 javiermonton@deploy2003: javiermonton: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:46 javiermonton@deploy2003: Started scap sync-world: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] * 14:42 otto@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 14:41 otto@deploy2003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 14:41 otto@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 14:40 otto@deploy2003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 14:40 otto@deploy2003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 14:39 otto@deploy2003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 14:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2015.codfw.wmnet with reason: host reimage * 14:34 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1015.eqiad.wmnet with reason: host reimage * 14:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2015.codfw.wmnet with reason: host reimage * 14:30 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1015.eqiad.wmnet with reason: host reimage * 14:30 seanleong-wmde@deploy2003: mwscript-k8s job started: foreachwikiindblist wikidataclient extensions/Wikibase/lib/maintenance/populateSitesTable.php --force-protocol https # [[phab:T429939|T429939]] * 14:24 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling reboot on A:durum and not (A:durum-eqiad or A:durum-codfw or A:durum-esams) and A:durum * 14:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2015.codfw.wmnet with OS trixie * 14:15 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2016.codfw.wmnet with OS trixie * 14:14 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2159: repooling after rack b5 maintenance * 14:14 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1015.eqiad.wmnet with OS trixie * 14:12 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1016.eqiad.wmnet with OS trixie * 14:12 cmooney@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=pki,name=codfw * 14:12 sbisson@deploy2003: helmfile [codfw] DONE helmfile.d/services/cxserver: sync * 14:11 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2002.codfw.wmnet * 14:11 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2002.codfw.wmnet * 14:11 sbisson@deploy2003: helmfile [codfw] START helmfile.d/services/cxserver: sync * 14:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1005.wikimedia.org * 14:07 sbisson@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cxserver: sync * 14:07 sbisson@deploy2003: helmfile [eqiad] START helmfile.d/services/cxserver: sync * 14:05 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1005.wikimedia.org * 14:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader2005.wikimedia.org * 14:02 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2003.codfw.wmnet * 14:02 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2003.codfw.wmnet * 14:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=tegola-vector-tiles,name=codfw * 14:00 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=kartotherian,name=codfw * 14:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader2005.wikimedia.org * 13:58 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2016.codfw.wmnet with reason: host reimage * 13:57 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-ncredir (exit_code=0) rolling reboot on A:ncredir and A:ncredir * 13:57 sbisson@deploy2003: helmfile [staging] DONE helmfile.d/services/cxserver: sync * 13:56 sbisson@deploy2003: helmfile [staging] START helmfile.d/services/cxserver: sync * 13:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1016.eqiad.wmnet with reason: host reimage * 13:52 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:52 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:51 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2016.codfw.wmnet with reason: host reimage * 13:50 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1016.eqiad.wmnet with reason: host reimage * 13:49 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy (exit_code=0) rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 13:49 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling reboot on A:wikidough * 13:46 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-tcp-proxy (exit_code=0) rolling reboot on A:tcpproxy and A:tcpproxy * 13:43 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and not (A:durum-eqiad or A:durum-codfw or A:durum-esams) and A:durum * 13:42 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=97) rolling reboot on A:durum and A:durum * 13:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2011.codfw.wmnet * 13:36 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-reboot (exit_code=0) rolling reboot on A:dnsbox and A:ulsfo and (A:dnsbox) * 13:36 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns4004.wikimedia.org * 13:34 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1016.eqiad.wmnet with OS trixie * 13:34 topranks: reboot lsw1-b5-codfw to upgrade JunOS [[phab:T430918|T430918]] * 13:34 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2016.codfw.wmnet with OS trixie * 13:32 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2002.codfw.wmnet * 13:31 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2011.codfw.wmnet * 13:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2012.codfw.wmnet * 13:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2017.codfw.wmnet with OS trixie * 13:25 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1017.eqiad.wmnet with OS trixie * 13:24 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2012.codfw.wmnet * 13:22 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2002.codfw.wmnet * 13:22 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:22 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:22 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns4004.wikimedia.org * 13:21 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2005.codfw.wmnet * 13:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2013.codfw.wmnet * 13:18 elukey@dns1004: END - running authdns-update * 13:17 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2005.codfw.wmnet * 13:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2004.codfw.wmnet * 13:16 elukey@dns1004: START - running authdns-update * 13:16 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1046: es1046 after reimage * 13:14 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1029.eqiad.wmnet,service=s8 * 13:14 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1029.eqiad.wmnet,service=s5 * 13:13 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1029.eqiad.wmnet,service=s5 * 13:13 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1029.eqiad.wmnet,service=s8 * 13:13 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2004.codfw.wmnet * 13:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2013.codfw.wmnet * 13:11 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2014.codfw.wmnet * 13:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2003.codfw.wmnet * 13:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2017.codfw.wmnet with reason: host reimage * 13:07 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:07 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns4003.wikimedia.org * 13:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2003.codfw.wmnet * 13:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1067.eqiad.wmnet * 13:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1067.eqiad.wmnet * 13:06 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1067.eqiad.wmnet * 13:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm-test1001.wikimedia.org * 13:05 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2017.codfw.wmnet with reason: host reimage * 13:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1017.eqiad.wmnet with reason: host reimage * 13:04 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2014.codfw.wmnet * 13:03 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:02 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2188: codfw rack B5 depool for maintenance * 13:02 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_magru * 13:01 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2188: codfw rack B5 depool for maintenance * 13:01 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_magru * 13:01 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2178: codfw rack B5 depool for maintenance * 13:01 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm-test1001.wikimedia.org * 13:01 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2178: codfw rack B5 depool for maintenance * 13:01 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2177: codfw rack B5 depool for maintenance * 13:00 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2177: codfw rack B5 depool for maintenance * 12:59 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1068.eqiad.wmnet * 12:59 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1068.eqiad.wmnet * 12:58 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola-vector-tiles,name=codfw * 12:58 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2159: codfw rack B5 depool for maintenance * 12:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1017.eqiad.wmnet with reason: host reimage * 12:58 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola,name=codfw * 12:57 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=kartotherian,name=codfw * 12:57 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2159: codfw rack B5 depool for maintenance * 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 30 hosts with reason: lsw1-b5-codfw JunOS upgrade * 12:55 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lsw1-b5-codfw,lsw1-b5-codfw IPv6,lsw1-b5-codfw.mgmt,ssw1-a[1,8]-codfw.mgmt with reason: switch upgade lsw1-b5-codfw * 12:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet * 12:49 cmooney@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=pki,name=codfw * 12:49 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1067.eqiad.wmnet with OS trixie * 12:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet * 12:48 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2017.codfw.wmnet with OS trixie * 12:47 topranks: depool codfw pki in dns discovery ahead of lsw1-b5-codfw maintenance [[phab:T430918|T430918]] * 12:47 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns4003.wikimedia.org * 12:47 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and A:ulsfo and (A:dnsbox) * 12:47 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2018.codfw.wmnet * 12:45 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2018.codfw.wmnet * 12:45 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and A:durum * 12:45 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-tcp-proxy rolling reboot on A:tcpproxy and A:tcpproxy * 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host pki-root1002.eqiad.wmnet * 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1009.eqiad.wmnet * 12:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1009.eqiad.wmnet * 12:44 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 12:43 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-ncredir rolling reboot on A:ncredir and A:ncredir * 12:43 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling reboot on A:wikidough * 12:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2018.codfw.wmnet with OS trixie * 12:42 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1017.eqiad.wmnet with OS trixie * 12:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader2006.wikimedia.org * 12:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1018.eqiad.wmnet with OS trixie * 12:39 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1009.eqiad.wmnet * 12:38 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host pki-root1002.eqiad.wmnet * 12:38 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1009.eqiad.wmnet * 12:38 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1008.eqiad.wmnet * 12:38 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1008.eqiad.wmnet * 12:35 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader2006.wikimedia.org * 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1006.wikimedia.org * 12:33 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1008.eqiad.wmnet * 12:30 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1046: es1046 after reimage * 12:29 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host es1046.eqiad.wmnet with OS trixie * 12:29 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1006.wikimedia.org * 12:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test2005.wikimedia.org * 12:28 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1008.eqiad.wmnet * 12:27 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1007.eqiad.wmnet * 12:27 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1007.eqiad.wmnet * 12:27 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1067.eqiad.wmnet with reason: host reimage * 12:25 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1068.eqiad.wmnet with reason: vacuum overlarge container dbs * 12:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2018.codfw.wmnet with reason: host reimage * 12:24 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test2005.wikimedia.org * 12:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test1005.wikimedia.org * 12:23 Amir1: mwscript-k8s --follow --dblist=ores -- extensions/ORES/maintenance/PurgeScoreCache.php --model goodfaith --old ([[phab:T431159|T431159]]) * 12:22 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1007.eqiad.wmnet * 12:22 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test1005.wikimedia.org * 12:22 atsukoito: restarting pybal on lvs2013 `low-traffic` for https://gerrit.wikimedia.org/r/1310535 * 12:22 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1007.eqiad.wmnet * 12:21 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1006.eqiad.wmnet * 12:21 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1006.eqiad.wmnet * 12:20 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1018.eqiad.wmnet with reason: host reimage * 12:19 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2018.codfw.wmnet with reason: host reimage * 12:18 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1067.eqiad.wmnet with reason: host reimage * 12:16 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1006.eqiad.wmnet * 12:15 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1006.eqiad.wmnet * 12:15 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1005.eqiad.wmnet * 12:15 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1005.eqiad.wmnet * 12:15 atsukoito: restarting pybal on lvs2014 for https://gerrit.wikimedia.org/r/1310535 * 12:12 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1018.eqiad.wmnet with reason: host reimage * 12:11 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1005.eqiad.wmnet * 12:11 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1005.eqiad.wmnet * 12:10 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1004.eqiad.wmnet * 12:10 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1004.eqiad.wmnet * 12:09 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on es1046.eqiad.wmnet with reason: host reimage * 12:08 atsukoito: restarting pybal on lvs1019 `low-traffic` for https://gerrit.wikimedia.org/r/1310535 * 12:06 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1004.eqiad.wmnet * 12:06 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1004.eqiad.wmnet * 12:06 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1003.eqiad.wmnet * 12:06 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1003.eqiad.wmnet * 12:05 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on es1046.eqiad.wmnet with reason: host reimage * 12:05 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "set ml-serve1001 back to active state - cmooney@cumin1003" * 12:04 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "set ml-serve1001 back to active state - cmooney@cumin1003" * 12:04 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:02 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1003.eqiad.wmnet * 12:01 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1003.eqiad.wmnet * 12:01 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1002.eqiad.wmnet * 12:01 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1002.eqiad.wmnet * 12:01 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:59 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2018.codfw.wmnet with OS trixie * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1067 * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1067 * 11:59 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1067 * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1067.eqiad.wmnet 17.48.64.10.in-addr.arpa 7.1.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:59 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1067.eqiad.wmnet 17.48.64.10.in-addr.arpa 7.1.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1067 - blake@cumin1003" * 11:58 atsukoito: restarting pybal on lvs1018 `high-traffic2` for https://gerrit.wikimedia.org/r/1310535 * 11:57 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1002.eqiad.wmnet * 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2019.codfw.wmnet with OS trixie * 11:56 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1002.eqiad.wmnet * 11:56 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 11:56 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1018.eqiad.wmnet with OS trixie * 11:54 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 11:54 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:54 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1019.eqiad.wmnet with OS trixie * 11:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:49 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:49 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:49 aikochou@deploy2003: helmfile [codfw] DONE helmfile.d/services/changeprop: sync * 11:48 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host es1046.eqiad.wmnet with OS trixie * 11:48 aikochou@deploy2003: helmfile [codfw] START helmfile.d/services/changeprop: sync * 11:48 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310535 * 11:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1046: Reimage to Trixie * 11:44 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1046: Reimage to Trixie * 11:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5:00:00 on es1046.eqiad.wmnet with reason: Reimage to Trixie * 11:42 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:42 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:42 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 11:42 aikochou@deploy2003: helmfile [eqiad] DONE helmfile.d/services/changeprop: sync * 11:41 aikochou@deploy2003: helmfile [eqiad] START helmfile.d/services/changeprop: sync * 11:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2019.codfw.wmnet with reason: host reimage * 11:36 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] (duration: 09m 41s) * 11:36 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:36 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:35 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:35 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:32 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1019.eqiad.wmnet with reason: host reimage * 11:32 jforrester@deploy2003: jforrester, gengh: Continuing with deployment * 11:29 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2019.codfw.wmnet with reason: host reimage * 11:28 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1019.eqiad.wmnet with reason: host reimage * 11:28 jforrester@deploy2003: jforrester, gengh: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:26 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] * 11:20 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2003.codfw.wmnet * 11:20 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:19 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2003.codfw.wmnet * 11:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:12 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1019.eqiad.wmnet with OS trixie * 11:10 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2019.codfw.wmnet with OS trixie * 11:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1020.eqiad.wmnet with OS trixie * 11:10 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1067 - blake@cumin1003" * 11:09 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] (duration: 12m 12s) * 11:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2020.codfw.wmnet with OS trixie * 11:03 kharlan@deploy2003: kharlan: Continuing with deployment * 11:01 blake@cumin1003: START - Cookbook sre.dns.netbox * 11:01 kharlan@deploy2003: kharlan: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:57 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] * 10:55 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] (duration: 31m 40s) * 10:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1020.eqiad.wmnet with reason: host reimage * 10:52 marostegui@dns1004: START - running authdns-update * 10:49 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2020.codfw.wmnet with reason: host reimage * 10:49 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:48 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1020.eqiad.wmnet with reason: host reimage * 10:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2020.codfw.wmnet with reason: host reimage * 10:43 kharlan@deploy2003: kharlan: Continuing with deployment * 10:42 kharlan@deploy2003: kharlan: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:32 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1020.eqiad.wmnet with OS trixie * 10:29 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2159: Repooling after switchover * 10:29 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1067 * 10:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1021.eqiad.wmnet with OS trixie * 10:27 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1067.eqiad.wmnet with OS trixie * 10:27 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:27 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1067.eqiad.wmnet * 10:27 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:26 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1067.eqiad.wmnet * 10:26 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1067.eqiad.wmnet * 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2020.codfw.wmnet with OS trixie * 10:26 blake@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1055.eqiad.wmnet * 10:26 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1055.eqiad.wmnet * 10:26 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1055.eqiad.wmnet * 10:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2021.codfw.wmnet with OS trixie * 10:24 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] * 10:11 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1055.eqiad.wmnet with OS trixie * 10:09 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1021.eqiad.wmnet with reason: host reimage * 10:05 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2021.codfw.wmnet with reason: host reimage * 10:03 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310129 revert * 10:02 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1021.eqiad.wmnet with reason: host reimage * 10:01 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2021.codfw.wmnet with reason: host reimage * 09:58 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310129 * 09:50 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1055.eqiad.wmnet with reason: host reimage * 09:45 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1055.eqiad.wmnet with reason: host reimage * 09:45 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1021.eqiad.wmnet with OS trixie * 09:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1022.eqiad.wmnet with OS trixie * 09:44 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2159: Repooling after switchover * 09:44 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2021.codfw.wmnet with OS trixie * 09:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2022.codfw.wmnet with OS trixie * 09:31 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2159.codfw.wmnet * 09:28 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1055 * 09:28 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1055 * 09:27 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on ms-fe1022.eqiad.wmnet with reason: host reimage * 09:27 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1022.eqiad.wmnet with reason: host reimage * 09:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2022.codfw.wmnet with reason: host reimage * 09:21 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2022.codfw.wmnet with reason: host reimage * 09:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2159: Rebooting db2159.codfw.wmnet * 09:20 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2159: Rebooting db2159.codfw.wmnet * 09:18 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 09:18 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 09:18 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 09:17 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 09:16 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2159.codfw.wmnet * 09:13 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b] (thin): Regular analytics weekly train THIN [analytics/refinery@ad6e05b8] (duration: 02m 07s) * 09:11 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b] (thin): Regular analytics weekly train THIN [analytics/refinery@ad6e05b8] * 09:10 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1022.eqiad.wmnet with OS trixie * 09:07 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1023.eqiad.wmnet with OS trixie * 09:06 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b]: Regular analytics weekly train [analytics/refinery@ad6e05b8] (duration: 04m 49s) * 09:04 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2022.codfw.wmnet with OS trixie * 09:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2023.codfw.wmnet with OS trixie * 09:01 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b]: Regular analytics weekly train [analytics/refinery@ad6e05b8] * 09:01 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@ad6e05b8] (duration: 02m 01s) * 09:00 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1055 * 09:00 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1055.eqiad.wmnet 50.32.64.10.in-addr.arpa 0.5.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:00 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1055.eqiad.wmnet 50.32.64.10.in-addr.arpa 0.5.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:00 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:00 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1055 - blake@cumin1003" * 09:00 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1055 - blake@cumin1003" * 08:59 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@ad6e05b8] * 08:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2159 [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94811 and previous config saved to /var/cache/conftool/dbconfig/20260714-085624-cwilliams.json * 08:55 blake@cumin1003: START - Cookbook sre.dns.netbox * 08:55 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1055 * 08:54 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1055.eqiad.wmnet with OS trixie * 08:54 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1055.eqiad.wmnet * 08:53 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1055.eqiad.wmnet * 08:53 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1055.eqiad.wmnet * 08:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2220 to s7 primary [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94810 and previous config saved to /var/cache/conftool/dbconfig/20260714-085239-cwilliams.json * 08:51 cezmunsta: Starting s7 codfw failover from db2159 to db2220 - [[phab:T430920|T430920]] * 08:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1023.eqiad.wmnet with reason: host reimage * 08:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2220 with weight 0 [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94809 and previous config saved to /var/cache/conftool/dbconfig/20260714-084553-cwilliams.json * 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s7 [[phab:T430920|T430920]] * 08:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2023.codfw.wmnet with reason: host reimage * 08:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1023.eqiad.wmnet with reason: host reimage * 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2023.codfw.wmnet with reason: host reimage * 08:34 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:34 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:29 marostegui@dns1004: END - running authdns-update * 08:29 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox-dev2003.codfw.wmnet * 08:27 marostegui@dns1004: START - running authdns-update * 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker2*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2009.codfw.wmnet * 08:26 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2009.codfw.wmnet * 08:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1023.eqiad.wmnet with OS trixie * 08:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox-dev2003.codfw.wmnet * 08:24 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:24 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:24 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2023.codfw.wmnet with OS trixie * 08:24 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1029.eqiad.wmnet with reason: reboot * 08:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1027.eqiad.wmnet with reason: reboot * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:21 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2009.codfw.wmnet * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:20 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2009.codfw.wmnet * 08:20 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2008.codfw.wmnet * 08:20 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2008.codfw.wmnet * 08:15 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2008.codfw.wmnet * 08:14 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2008.codfw.wmnet * 08:14 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2007.codfw.wmnet * 08:14 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2007.codfw.wmnet * 08:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1024.eqiad.wmnet with OS trixie * 08:12 elukey@cumin1003: END (PASS) - Cookbook sre.pki.restart-reboot (exit_code=0) rolling reboot on P<nowiki>{</nowiki>pki*<nowiki>}</nowiki> and (A:pki) * 08:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 08:10 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 08:09 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2007.codfw.wmnet * 08:08 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2007.codfw.wmnet * 08:08 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2006.codfw.wmnet * 08:08 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2006.codfw.wmnet * 08:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2024.codfw.wmnet with OS trixie * 08:03 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2006.codfw.wmnet * 08:02 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2006.codfw.wmnet * 08:02 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2005.codfw.wmnet * 08:02 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2005.codfw.wmnet * 07:58 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2005.codfw.wmnet * 07:58 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2005.codfw.wmnet * 07:57 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2004.codfw.wmnet * 07:57 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2004.codfw.wmnet * 07:54 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki.discovery.wmnet. on all recursors * 07:54 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache pki.discovery.wmnet. on all recursors * 07:53 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2004.codfw.wmnet * 07:53 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2004.codfw.wmnet * 07:53 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2003.codfw.wmnet * 07:53 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2003.codfw.wmnet * 07:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1024.eqiad.wmnet with reason: host reimage * 07:49 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki.discovery.wmnet. on all recursors * 07:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2003.codfw.wmnet * 07:49 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache pki.discovery.wmnet. on all recursors * 07:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2024.codfw.wmnet with reason: host reimage * 07:48 elukey@cumin1003: START - Cookbook sre.pki.restart-reboot rolling reboot on P<nowiki>{</nowiki>pki*<nowiki>}</nowiki> and (A:pki) * 07:46 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1024.eqiad.wmnet with reason: host reimage * 07:45 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2024.codfw.wmnet with reason: host reimage * 07:45 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2003.codfw.wmnet * 07:44 elukey@cumin1003: END (PASS) - Cookbook sre.misc-clusters.restart-reboot-config-master (exit_code=0) rolling reboot on P<nowiki>{</nowiki>config-master*<nowiki>}</nowiki> and (A:config-master or A:config-master-eqiad or A:config-master-codfw) * 07:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2002.codfw.wmnet * 07:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2002.codfw.wmnet * 07:39 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2002.codfw.wmnet * 07:39 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) config-master.discovery.wmnet. on all recursors * 07:39 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache config-master.discovery.wmnet. on all recursors * 07:39 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2002.codfw.wmnet * 07:39 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker2*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl200*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl2003.codfw.wmnet * 07:36 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl2003.codfw.wmnet * 07:35 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) config-master.discovery.wmnet. on all recursors * 07:35 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache config-master.discovery.wmnet. on all recursors * 07:34 elukey@cumin1003: START - Cookbook sre.misc-clusters.restart-reboot-config-master rolling reboot on P<nowiki>{</nowiki>config-master*<nowiki>}</nowiki> and (A:config-master or A:config-master-eqiad or A:config-master-codfw) * 07:31 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl2003.codfw.wmnet * 07:31 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl2003.codfw.wmnet * 07:31 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl2002.codfw.wmnet * 07:31 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl2002.codfw.wmnet * 07:29 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1024.eqiad.wmnet with OS trixie * 07:28 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2024.codfw.wmnet with OS trixie * 07:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl2002.codfw.wmnet * 07:26 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl2002.codfw.wmnet * 07:26 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl200*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 07:26 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 07:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 06:50 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lists1004.wikimedia.org * 06:44 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host lists1004.wikimedia.org * 06:25 marostegui@dns1004: END - running authdns-update * 06:23 marostegui@dns1004: START - running authdns-update * 06:22 marostegui@dns1004: END - running authdns-update * 06:20 marostegui@dns1004: START - running authdns-update * 06:17 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1026.eqiad.wmnet with reason: reboot * 06:04 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: sync * 06:04 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: sync * 06:03 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync * 06:03 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync * 06:02 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync * 06:01 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync * 06:01 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync * 06:00 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync * 05:59 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:59 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:40 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:39 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:26 marostegui@dns1004: END - running authdns-update * 05:24 marostegui@dns1004: START - running authdns-update * 05:24 marostegui@dns1004: START - running authdns-update * 05:13 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1004.wikimedia.org * 05:07 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1004.wikimedia.org * 04:01 mwpresync@deploy2003: Pruned MediaWiki: 1.47.0-wmf.8 (duration: 01m 07s) * 03:39 mwpresync@deploy2003: Finished scap sync-world: testwikis to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] (duration: 36m 01s) * 03:03 mwpresync@deploy2003: Started scap sync-world: testwikis to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 29s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-13 == * 23:33 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1064.eqiad.wmnet * 23:33 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1064.eqiad.wmnet * 23:08 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1064.eqiad.wmnet with reason: vacuum overlarge container dbs * 23:06 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1069.eqiad.wmnet * 23:06 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1069.eqiad.wmnet * 22:34 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1069.eqiad.wmnet with reason: vacuum overlarge container dbs * 21:18 maryum: Deployed security fix for [[phab:T321092|T321092]] * 20:28 swfrench-wmf: reprepro include etcd-mirror_0.0.12-1+deb13u1 into main for trixie-wikimedia - [[phab:T424266|T424266]] * 20:26 swfrench-wmf: reprepro include etcd-mirror_0.0.12-1+deb12u1 into main for bookworm-wikimedia - [[phab:T428495|T428495]] * 20:23 dancy@deploy2003: Finished scap sync-world: Testing [[phab:T431635|T431635]] (duration: 03m 36s) * 20:19 dancy@deploy2003: Started scap sync-world: Testing [[phab:T431635|T431635]] * 20:18 dancy@deploy2003: Installation of scap version "4.274.0" completed for 3 hosts * 20:16 dancy@deploy2003: Installing scap version "4.274.0" for 3 host(s) * 20:12 kemayo@deploy2003: Finished scap sync-world: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] (duration: 08m 25s) * 20:07 kemayo@deploy2003: soda, esanders, kemayo: Continuing with deployment * 20:05 kemayo@deploy2003: soda, esanders, kemayo: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there * 20:04 kemayo@deploy2003: Started scap sync-world: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] * 18:22 cwhite: lvextend vg0/srv +500g on centrallog hosts * 18:19 cdobbins@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS trixie * 17:46 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1071.eqiad.wmnet * 17:46 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1071.eqiad.wmnet * 17:13 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1071.eqiad.wmnet with reason: vacuum overlarge container dbs * 17:07 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1065.eqiad.wmnet * 17:07 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1065.eqiad.wmnet * 17:06 dzahn@dns1006: END - running authdns-update * 17:04 dzahn@dns1006: START - running authdns-update * 17:01 dzahn@dns1006: END - running authdns-update * 16:59 dzahn@dns1006: START - running authdns-update * 16:51 dancy@deploy2003: Finished scap sync-world: testing [[phab:T428971|T428971]] (duration: 03m 37s) * 16:47 dancy@deploy2003: Started scap sync-world: testing [[phab:T428971|T428971]] * 16:45 atsukoito: restarting pybal on lvs1019 to flush IP address for `cirrussearch1122.eqiad.wmnet` after moving the vlan [[phab:T431311|T431311]] * 16:42 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:42 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:42 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:42 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:42 Amir1: mwscript-k8s --follow --dblist=ores -- extensions/ORES/maintenance/PurgeScoreCache.php --model damaging --old ([[phab:T431159|T431159]]) * 16:34 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Pool test * 16:34 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 16:34 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 16:34 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Pool test * 16:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Depool test * 16:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 16:33 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 16:33 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Depool test * 16:31 dancy@deploy2003: Installation of scap version "4.273.0" completed for 159 hosts * 16:29 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1065.eqiad.wmnet with reason: vacuum overlarge container dbs * 16:27 dancy@deploy2003: Installing scap version "4.273.0" for 159 host(s) * 16:27 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics-external: sync * 16:27 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics-external: sync * 16:26 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics-external: sync * 16:26 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics-external: sync * 16:22 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync * 16:21 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync * 16:21 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: sync * 16:21 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: sync * 16:19 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync * 16:19 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync * 16:18 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync * 16:17 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync * 15:59 atsukoito: restarting pybal on lvs1018 for https://gerrit.wikimedia.org/r/1310117 * 15:55 aikochou@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 15:50 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310117 * 15:46 aikochou@deploy2003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 15:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host kafka-logging1006.eqiad.wmnet * 15:43 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host kafka-logging1006.eqiad.wmnet * 15:41 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host ganeti-test[2001-2003].codfw.wmnet * 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host ganeti-test[2001-2003].codfw.wmnet * 15:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host netbox1003.eqiad.wmnet * 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host netbox1003.eqiad.wmnet * 15:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host netbox2003.codfw.wmnet * 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host netbox2003.codfw.wmnet * 15:36 sukhe: restart pybal on lvs1020 * 15:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet * 15:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet * 15:08 btullis@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:06 btullis@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 15:01 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:01 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:35 cdobbins@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 14:34 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:33 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:33 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:32 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:29 cdobbins@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 14:28 marostegui@dns1004: END - running authdns-update * 14:28 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:27 marostegui@dns1004: START - running authdns-update * 14:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1023.eqiad.wmnet with reason: reboot * 14:18 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2009.codfw.wmnet with OS trixie * 14:14 swfrench-wmf: start rolling run-puppet-agent on A:cp for ATS config change - [[phab:T428909|T428909]] [[phab:T431838|T431838]] * 14:05 swfrench-wmf: disable-puppet on A:cp for ATS config change - [[phab:T428909|T428909]] [[phab:T431838|T431838]] * 14:05 cdobbins@cumin2002: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie * 14:02 marostegui@dns1004: END - running authdns-update * 14:00 marostegui@dns1004: START - running authdns-update * 14:00 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1070.eqiad.wmnet * 14:00 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1070.eqiad.wmnet * 13:58 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2009.codfw.wmnet with reason: host reimage * 13:57 cdobbins@cumin2002: conftool action : set/pooled=no; selector: name=dns7002.* * 13:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2009.codfw.wmnet with reason: host reimage * 13:48 rscout@deploy2003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply * 13:48 rscout@deploy2003: helmfile [eqiad] START helmfile.d/services/miscweb: apply * 13:48 rscout@deploy2003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply * 13:47 rscout@deploy2003: helmfile [codfw] START helmfile.d/services/miscweb: apply * 13:40 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:33 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2009.codfw.wmnet with OS trixie * 13:30 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1070.eqiad.wmnet with reason: vacuum overlarge container dbs * 13:28 aude@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] (duration: 11m 12s) * 13:23 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:22 aude@deploy2003: aikochou, javiermonton, aude, gkm563: Continuing with deployment * 13:22 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:19 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:19 aude@deploy2003: aikochou, javiermonton, aude, gkm563: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] synced to the testservers * 13:17 aude@deploy2003: Started scap sync-world: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] * 13:01 ladsgroup@deploy2003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 13:01 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:00 ladsgroup@deploy2003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 12:59 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:52 ladsgroup@deploy2003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 12:51 ladsgroup@deploy2003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 12:48 atsuko@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 12:48 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 12:47 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2008.codfw.wmnet with OS trixie * 12:47 atsuko@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 12:47 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply * 12:47 atsuko@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:46 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 12:45 atsuko@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:45 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply * 12:45 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:44 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:43 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] (duration: 07m 02s) * 12:38 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 12:37 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:36 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] * 12:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2008.codfw.wmnet with reason: host reimage * 12:23 Msz2001: Deployed changes to private code for Suggested Investigations * 12:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2008.codfw.wmnet with reason: host reimage * 12:20 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:19 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:17 atsuko@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 12:17 atsuko@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 12:16 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:15 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] (duration: 07m 14s) * 12:10 mszwarc@deploy2003: mszwarc: Continuing with deployment * 12:09 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:07 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] * 12:04 mszwarc@deploy2003: sync-world aborted: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] (duration: 00m 29s) * 12:03 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] * 12:01 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2008.codfw.wmnet with OS trixie * 12:00 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] (duration: 07m 37s) * 11:55 zabe@deploy2003: zabe: Continuing with deployment * 11:54 zabe@deploy2003: zabe: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:52 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] * 11:51 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:43 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:35 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:34 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:33 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:30 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:28 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:27 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:17 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2007.codfw.wmnet with OS trixie * 11:09 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] * 11:06 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=s8 * 11:00 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=x3 * 11:00 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=s5 * 10:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2007.codfw.wmnet with reason: host reimage * 10:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2007.codfw.wmnet with reason: host reimage * 10:51 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host dse-k8s-worker1023 * 10:50 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host dse-k8s-worker1023 * 10:44 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dse-k8s-worker1023 * 10:43 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host dse-k8s-worker1023 * 10:42 marostegui@cumin1003: dbctl commit (dc=all): 'Change x4 masters [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P94804 and previous config saved to /var/cache/conftool/dbconfig/20260713-104248-marostegui.json * 10:37 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:37 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:35 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:35 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:34 atsuko@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 10:34 atsuko@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 10:33 marostegui@cumin1003: dbctl commit (dc=all): 'Push x4 initial dbctl config [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P94803 and previous config saved to /var/cache/conftool/dbconfig/20260713-103259-marostegui.json * 10:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2007.codfw.wmnet with OS trixie * 09:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2006.codfw.wmnet with OS trixie * 09:42 marostegui@dns1004: END - running authdns-update * 09:40 marostegui@dns1004: START - running authdns-update * 09:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2006.codfw.wmnet with reason: host reimage * 09:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2006.codfw.wmnet with reason: host reimage * 09:06 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1024.eqiad.wmnet with reason: reboot * 09:06 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:01 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2006.codfw.wmnet with OS trixie * 08:44 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 08:43 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 08:43 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:42 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:42 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:42 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:41 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 08:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup2004.codfw.wmnet * 08:38 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:33 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host db1208.eqiad.wmnet * 08:30 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=x3 * 08:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2005.codfw.wmnet with OS trixie * 08:28 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup2004.codfw.wmnet * 08:28 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup2003.codfw.wmnet * 08:24 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1039: Repooling after testing * 08:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on clouddb1016.eqiad.wmnet with reason: cloning * 08:23 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s5 * 08:23 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s8 * 08:21 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 08:21 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 08:17 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup2003.codfw.wmnet * 08:17 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1004.eqiad.wmnet * 08:14 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1208.eqiad.wmnet * 08:11 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host phab1005.eqiad.wmnet * 08:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2005.codfw.wmnet with reason: host reimage * 08:07 marostegui@dns1004: END - running authdns-update * 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1004.eqiad.wmnet * 08:07 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1003.eqiad.wmnet * 08:07 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:05 marostegui@dns1004: START - running authdns-update * 08:05 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2005.codfw.wmnet with reason: host reimage * 08:05 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 08:05 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host phab1005.eqiad.wmnet * 08:05 marostegui@dns1004: START - running authdns-update * 08:05 marostegui@dns1004: START - running authdns-update * 08:05 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 08:04 marostegui@dns1004: START - running authdns-update * 08:00 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit1003.wikimedia.org * 07:58 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1003.eqiad.wmnet * 07:58 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1002-dev.eqiad.wmnet * 07:58 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:58 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:54 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1002-dev.eqiad.wmnet * 07:54 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1001-dev.eqiad.wmnet * 07:54 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit1003.wikimedia.org * 07:53 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 07:53 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:52 Msz2001: UTC morning backport+config window done * 07:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2005.codfw.wmnet with OS trixie * {{safesubst:SAL entry|1=07:50 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark (T429943}} * 07:49 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1001-dev.eqiad.wmnet * 07:46 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:46 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:45 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 07:45 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:44 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 07:44 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:43 mszwarc@deploy2003: mszwarc, danielyepezgarces, anzx: Continuing with deployment * {{safesubst:SAL entry|1=07:39 mszwarc@deploy2003: mszwarc, danielyepezgarces, anzx: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark}} * 07:39 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1039: Repooling after testing * {{safesubst:SAL entry|1=07:36 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark (T429943)}} * 07:35 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] (duration: 30m 03s) * 07:25 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit2002.wikimedia.org * 07:22 mszwarc@deploy2003: mszwarc: Continuing with deployment * 07:21 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:19 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit2002.wikimedia.org * 07:15 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aphlict1002.eqiad.wmnet * 07:11 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host aphlict1002.eqiad.wmnet * 07:08 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2003.wikimedia.org * 07:05 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] * 07:02 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2003.wikimedia.org * 07:02 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2002.wikimedia.org * 06:55 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2002.wikimedia.org * 06:55 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1003.wikimedia.org * 06:49 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1003.wikimedia.org * 06:34 marostegui: Drop m5 ipoid database [[phab:T431007|T431007]] * 06:29 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1027.eqiad.wmnet with reason: reboot * 06:24 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1028.eqiad.wmnet with reason: reboot * 06:21 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1025.eqiad.wmnet with reason: reboot * 06:17 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1022.eqiad.wmnet with reason: reboot * 06:03 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on dbproxy[2005-2008].codfw.wmnet with reason: reboot * 05:37 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1217,1228].eqiad.wmnet with reason: cloning * 05:11 marostegui: Drop users_to_rename table [[phab:T431842|T431842]] * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-12 == * 16:01 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2209 [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94792 and previous config saved to /var/cache/conftool/dbconfig/20260712-160124-marostegui.json * 15:58 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2205 to s3 primary [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94791 and previous config saved to /var/cache/conftool/dbconfig/20260712-155853-marostegui.json * 15:58 marostegui: Starting s3 codfw emergency failover from db2209 to db2205 - [[phab:T431950|T431950]] * 15:51 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2205 with weight 0 [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94790 and previous config saved to /var/cache/conftool/dbconfig/20260712-155135-marostegui.json * 15:51 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Primary switchover s3 [[phab:T431950|T431950]] * 02:01 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 01m 17s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-11 == * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 26s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-10 == * 19:12 jhathaway@dns1004: END - running authdns-update * 19:10 jhathaway@dns1004: START - running authdns-update * 18:23 mutante: vrts2002 rebooting (not the active host) * 18:21 mutante: lists2001, phab2003 - rebooting (not the active hosts) * 18:16 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on A:lvs-high-traffic2-codfw * 18:15 swfrench@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on A:lvs-high-traffic2-codfw * 17:15 mutante: [doc1004:~] $ sudo systemctl start rsync-doc-host-data-sync ([[phab:T431856|T431856]]) * 17:09 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1004.eqiad.wmnet * 17:08 jhathaway@dns1004: END - running authdns-update * 17:07 jhathaway@dns1004: START - running authdns-update * 17:06 jhathaway: depooling puppetserver1002, cause of errors is still unknown * 17:03 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1004.eqiad.wmnet * 16:57 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1003.eqiad.wmnet * 16:51 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1003.eqiad.wmnet * 16:48 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 16:48 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2004.codfw.wmnet * 16:42 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2004.codfw.wmnet * 16:41 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2003.codfw.wmnet * 16:35 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2003.codfw.wmnet * 16:33 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2002.codfw.wmnet * 16:27 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2002.codfw.wmnet * 16:25 mutante: gitlab-runners (production) rebooting cluster one by one * 16:17 mutante: etherpad1004/etherpad2002 - (etherpad.wikimedia.org) - rebooting * 16:13 mutante: doc1004/doc2003 (doc.wikimedia.org backends) - rebooting * 16:02 mutante: releases1003/releases2003 (releases.wikimedia.org backends) - rebooting for maintenance * 15:26 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2007-dev.codfw.wmnet * 15:19 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2007-dev.codfw.wmnet * 15:14 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host cloudcephosd2007-dev.codfw.wmnet * 15:14 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2007-dev.codfw.wmnet * 15:14 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host cloudcephosd2006-dev.codfw.wmnet * 15:07 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2006-dev.codfw.wmnet * 15:07 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2005-dev.codfw.wmnet * 14:59 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2005-dev.codfw.wmnet * 14:59 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2004-dev.codfw.wmnet * 14:53 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2004-dev.codfw.wmnet * 14:53 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2007-dev.codfw.wmnet * 14:51 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1054.eqiad.wmnet * 14:51 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1054.eqiad.wmnet * 14:51 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1054.eqiad.wmnet * 14:47 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2007-dev.codfw.wmnet * 14:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2006-dev.codfw.wmnet * 14:41 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2006-dev.codfw.wmnet * 14:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2005-dev.codfw.wmnet * 14:37 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2005-dev.codfw.wmnet * 14:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2005-dev.codfw.wmnet * 14:29 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2005-dev.codfw.wmnet * 14:29 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2006-dev.codfw.wmnet * 14:21 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2006-dev.codfw.wmnet * 14:21 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2010-dev.codfw.wmnet * 14:15 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2010-dev.codfw.wmnet * 14:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudgw2004-dev.codfw.wmnet * 14:10 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1054.eqiad.wmnet with OS trixie * 14:09 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudgw2004-dev.codfw.wmnet * 14:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudgw2003-dev.codfw.wmnet * 14:02 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudgw2003-dev.codfw.wmnet * 14:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2004-dev.codfw.wmnet * 13:53 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2004-dev.codfw.wmnet * 13:53 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2003-dev.codfw.wmnet * 13:48 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:44 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2003-dev.codfw.wmnet * 13:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2002-dev.codfw.wmnet * 13:42 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:41 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:41 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:37 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2002-dev.codfw.wmnet * 13:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudidp2001-dev.codfw.wmnet * 13:33 blake@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker1054.eqiad.wmnet with reason: host reimage * 13:33 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudidp2001-dev.codfw.wmnet * 13:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudnet2006-dev.codfw.wmnet * 13:26 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudnet2006-dev.codfw.wmnet * 13:26 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudnet2005-dev.codfw.wmnet * 13:23 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1054.eqiad.wmnet with reason: host reimage * 13:18 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudnet2005-dev.codfw.wmnet * 13:18 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudservices2005-dev.codfw.wmnet * 13:12 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudservices2005-dev.codfw.wmnet * 13:11 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudservices2004-dev.codfw.wmnet * 13:08 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudservices2004-dev.codfw.wmnet * 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudweb2002-dev.wikimedia.org * 13:05 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 13:05 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1054 * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1054 * 13:04 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1054 * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1054.eqiad.wmnet 49.32.64.10.in-addr.arpa 9.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:04 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1054.eqiad.wmnet 49.32.64.10.in-addr.arpa 9.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1054 - blake@cumin1003" * 13:04 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1054 - blake@cumin1003" * 13:01 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudweb2002-dev.wikimedia.org * 13:00 blake@cumin1003: START - Cookbook sre.dns.netbox * 12:59 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1054 * 12:57 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1054.eqiad.wmnet with OS trixie * 12:57 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1054.eqiad.wmnet * 12:56 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1054.eqiad.wmnet * 12:56 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1054.eqiad.wmnet * 12:47 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS trixie * 12:44 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:39 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 12:39 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 12:38 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:37 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:14 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:10 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:08 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:07 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:00 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:00 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:51 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:49 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:48 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:47 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:44 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:32 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2001.codfw.wmnet * 11:32 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1053.eqiad.wmnet * 11:32 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2001.codfw.wmnet * 11:32 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1053.eqiad.wmnet * 11:32 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1053.eqiad.wmnet * 11:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker2001.codfw.wmnet * 11:31 cgoubert@cumin1003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker2001.codfw.wmnet * 11:31 cgoubert@cumin1003: END (FAIL) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=1) rolling reimage on P<nowiki>{</nowiki>wikikube-worker2001*<nowiki>}</nowiki> and (A:wikikube-master-codfw or A:wikikube-worker-codfw) * 11:30 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:30 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:21 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 18 hosts with reason: reboot & upgrade * 11:20 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker2001.codfw.wmnet with OS trixie * 11:16 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:15 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:14 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:14 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 11 hosts * 11:14 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 11 hosts * 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:08 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 11:02 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1053.eqiad.wmnet with OS trixie * 11:01 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:58 cgoubert@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 10:57 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:38 cgoubert@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker2001.codfw.wmnet with OS trixie * 10:38 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2001.codfw.wmnet * 10:38 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2001.codfw.wmnet * 10:38 cgoubert@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on P<nowiki>{</nowiki>wikikube-worker2001*<nowiki>}</nowiki> and (A:wikikube-master-codfw or A:wikikube-worker-codfw) * 10:35 topranks: adjust IBGP outbound policy on lsw1-e2-codfw [[phab:T423430|T423430]] towards ssw1-e1-codfw * 10:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 cgoubert@cumin1003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:27 cgoubert@cumin1003: END (FAIL) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=1) rolling reimage on A:wikikube-worker-codfw * 10:27 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker2001.codfw.wmnet with OS bookworm * 10:25 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 10:24 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:24 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:15 cgoubert@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 10:11 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 10:11 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 10:08 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:08 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:07 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 10:06 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 10:00 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:55 cgoubert@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker2001.codfw.wmnet with OS bookworm * 09:55 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2005-2006,2011-2012].codfw.wmnet * 09:55 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2005-2006,2011-2012].codfw.wmnet * 09:51 cgoubert@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on A:wikikube-worker-codfw * 09:41 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1053.eqiad.wmnet with reason: host reimage * 09:37 topranks: apply new IBGP outbound policy on lsw1-e2-codfw [[phab:T423430|T423430]] * 09:36 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:36 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1053.eqiad.wmnet with reason: host reimage * 09:16 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1053 * 09:16 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1053 * 09:15 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1053 * 09:15 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1053.eqiad.wmnet 48.32.64.10.in-addr.arpa 8.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:15 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1053.eqiad.wmnet 48.32.64.10.in-addr.arpa 8.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:15 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:15 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1053 - blake@cumin1003" * 09:15 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1053 - blake@cumin1003" * 09:11 blake@cumin1003: START - Cookbook sre.dns.netbox * 09:11 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1053 * 09:08 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1053.eqiad.wmnet with OS trixie * 09:08 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1053.eqiad.wmnet * 09:08 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1053.eqiad.wmnet * 09:08 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1053.eqiad.wmnet * 09:04 brouberol@dns1004: END - running authdns-update * 09:03 brouberol@dns1004: START - running authdns-update * 08:41 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e] (thin): Regular analytics weekly train THIN [analytics/refinery@1abf22ea] (duration: 02m 11s) * 08:38 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e] (thin): Regular analytics weekly train THIN [analytics/refinery@1abf22ea] * 08:38 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e]: Regular analytics weekly train [analytics/refinery@1abf22ea] (duration: 05m 17s) * 08:38 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:34 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:33 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e]: Regular analytics weekly train [analytics/refinery@1abf22ea] * 08:32 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@1abf22ea] (duration: 02m 03s) * 08:30 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@1abf22ea] * 08:30 JavierMonton: Deploying Refinery at {{Gerrit|1abf22ea}} for changes 1308121/T427068 1306491/T430020 and {{Gerrit|1308190}} * 08:29 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:29 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 08:24 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:24 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 08:18 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:18 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 08:00 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db[2183-2184].codfw.wmnet * 08:00 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for db[2183-2184].codfw.wmnet * 07:52 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:52 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 07:49 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 11 hosts with reason: reboot & upgrade * 07:47 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:47 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 07:44 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:44 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 07:23 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 10 hosts * 07:23 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 10 hosts * 06:45 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 10 hosts with reason: reboot & upgrade * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 41s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-09 == * 23:33 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] (duration: 13m 26s) * 23:29 ladsgroup@deploy2003: ladsgroup, jdlrobson: Continuing with deployment * 23:22 ladsgroup@deploy2003: ladsgroup, jdlrobson: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:20 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] * 22:57 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1165.eqiad.wmnet * 22:56 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1165.eqiad.wmnet * 22:56 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1165.eqiad.wmnet * 22:45 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1165.eqiad.wmnet with OS trixie * 22:38 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 22:37 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 22:37 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 22:37 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:37 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 22:25 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1165.eqiad.wmnet with reason: host reimage * 22:17 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1165.eqiad.wmnet with reason: host reimage * 22:13 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 22:12 rzl@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 22:04 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 22:04 rzl@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1165 * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1165 * 22:02 jasmine@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1165 * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1165.eqiad.wmnet 115.48.64.10.in-addr.arpa 5.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:02 jasmine@cumin2002: START - Cookbook sre.dns.wipe-cache wikikube-worker1165.eqiad.wmnet 115.48.64.10.in-addr.arpa 5.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1165 - jasmine@cumin2002" * 22:02 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1165 - jasmine@cumin2002" * 22:02 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 21:57 jasmine@cumin2002: START - Cookbook sre.dns.netbox * 21:55 jasmine@cumin2002: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1165 * 21:54 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-worker1165.eqiad.wmnet with OS trixie * 21:54 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 21:54 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1165.eqiad.wmnet * 21:53 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 21:53 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1165.eqiad.wmnet * 21:53 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1165.eqiad.wmnet * 21:53 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 21:47 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 21:45 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 21:43 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 21:43 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 21:42 maryum: Deploy fix for [[phab:T431684|T431684]] * 21:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 21:27 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] (duration: 34m 14s) * 21:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs1002 * 21:23 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs1002 * 21:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS trixie * 21:22 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 21:20 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 22s) * 21:20 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 21:16 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 21:16 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 21:15 ladsgroup@deploy2003: ladsgroup: Continuing with deployment * 21:13 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2002.codfw.wmnet with OS bookworm * 21:11 ladsgroup@deploy2003: ladsgroup: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:08 ladsgroup@cumin1003: END (PASS) - Cookbook sre.wikireplicas.update-views (exit_code=0) * 21:07 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 6 hosts with reason: reboots * 20:54 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecycle work - bking@cumin2003 * 20:53 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:53 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] * 20:51 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) * 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 20:47 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecycle work - bking@cumin2003 * 20:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 20:41 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:41 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) * 20:40 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host relforge1008.eqiad.wmnet * 20:40 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1009.eqiad.wmnet with OS trixie * 20:33 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:32 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:32 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:31 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:31 ladsgroup@cumin1003: END (PASS) - Cookbook sre.wikireplicas.update-views (exit_code=0) * 20:29 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1008.eqiad.wmnet * 20:24 rzl@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 20:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2002.codfw.wmnet with OS bookworm * 20:23 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host relforge1008.eqiad.wmnet * 20:23 rzl@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 20:23 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1008.eqiad.wmnet * 20:22 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:22 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:21 bking@cumin2003: END (ERROR) - Cookbook sre.elasticsearch.rolling-operation (exit_code=97) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:21 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1009.eqiad.wmnet with reason: host reimage * 20:16 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:15 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1009.eqiad.wmnet with reason: host reimage * 20:12 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) * 20:02 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 19:55 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1009.eqiad.wmnet with OS trixie * 19:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 19:43 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 19:30 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 19:28 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 19:27 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 19:25 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 18:42 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 18:41 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 18:16 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 18:15 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 17:45 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for doh5004.wikimedia.org * 17:45 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for doh5004.wikimedia.org * 17:38 ladsgroup@deploy2003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 17:35 ladsgroup@deploy2003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 17:29 ladsgroup@deploy2003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 17:26 ladsgroup@deploy2003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 17:09 mutante: zuul[12]00[123] - rebooting for maintenance * 17:09 ebernhardson: start full in-place reindex of eqiad cirrussearch cluster * 17:08 dzahn@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-cluster (exit_code=99) * 17:08 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-cluster * 17:03 ebernhardson: start full in-place reindex of codfw cirrussearch cluster * 16:59 mutante: stewards1001/stewards2001 - reboot for maintenance * 16:54 ebernhardson: start full in-place reindex of cloudelastic cluster * 16:53 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 16:52 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply * 16:49 mutante: planet1003/planet2003 - rebooting * 16:47 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on doh5004.wikimedia.org with reason: random high load, investigating * 15:55 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 15:54 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 15:51 jynus: restarting backupmon1001 * 15:49 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 14 hosts * 15:49 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 14 hosts * 15:47 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backupmon1001.eqiad.wmnet with reason: restart * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:06 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 15:06 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 14:59 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply * 14:58 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply * 14:51 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 14 hosts * 14:51 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 14 hosts * 14:49 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 6 hosts with reason: reboot & upgrade * 14:48 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet * 14:48 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet * 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:42 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1052.eqiad.wmnet * 14:42 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1052.eqiad.wmnet * 14:42 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1052.eqiad.wmnet * 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:31 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:31 elukey@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: sync * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:30 elukey@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: sync * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:28 elukey@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: sync * 14:28 elukey@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: sync * 14:26 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:20 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1052.eqiad.wmnet with OS trixie * 14:19 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 6 hosts with reason: reboot & upgrade * 14:18 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:15 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:15 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:13 elukey: update druid indexation job for webrequest_sampled_live - [[phab:T427068|T427068]] * 14:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:09 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:09 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for papaul - jhancock@cumin2002" * 14:09 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for papaul - jhancock@cumin2002" * 14:07 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:07 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:04 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 14:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cuminunpriv1001.eqiad.wmnet * 13:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb1003.eqiad.wmnet * 13:59 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1052.eqiad.wmnet with reason: host reimage * 13:57 moritzm: installing requests security updates * 13:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cuminunpriv1001.eqiad.wmnet * 13:55 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb1003.eqiad.wmnet * 13:53 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1052.eqiad.wmnet with reason: host reimage * 13:50 moritzm: installing python-cryptography security updates * 13:47 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb2003.codfw.wmnet * 13:44 Msz2001: UTC afternoon config+backport window is done * 13:44 Msz2001: Updated `logging` on `metawiki` to fix log performers, [[phab:T431176|T431176]]#12105297 * 13:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb2003.codfw.wmnet * 13:43 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 13:43 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt1002.wikimedia.org * 13:41 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] (duration: 07m 30s) * 13:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt1002.wikimedia.org * 13:37 mszwarc@deploy2003: mszwarc: Continuing with deployment * 13:36 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1052 * 13:36 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1052 * 13:35 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:35 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1052 * 13:35 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1052.eqiad.wmnet 47.32.64.10.in-addr.arpa 7.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:35 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1052.eqiad.wmnet 47.32.64.10.in-addr.arpa 7.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:35 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:35 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1052 - blake@cumin1003" * 13:35 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1052 - blake@cumin1003" * 13:34 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] * 13:31 blake@cumin1003: START - Cookbook sre.dns.netbox * 13:31 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1052 * 13:30 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1052.eqiad.wmnet with OS trixie * 13:30 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1052.eqiad.wmnet * 13:29 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1052.eqiad.wmnet * 13:29 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1052.eqiad.wmnet * 13:17 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] (duration: 11m 26s) * 13:13 jforrester@deploy2003: jforrester: Continuing with deployment * 13:08 jforrester@deploy2003: jforrester: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:06 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] * 12:54 cgoubert@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply * 12:54 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:52 cgoubert@deploy2003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply * 12:45 cgoubert@deploy2003: helmfile [codfw] DONE helmfile.d/services/mobileapps: apply * 12:44 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:44 cgoubert@deploy2003: helmfile [codfw] START helmfile.d/services/mobileapps: apply * 12:43 cgoubert@deploy2003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 12:43 cgoubert@deploy2003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 12:42 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast4006.wikimedia.org * 12:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt2002.wikimedia.org * 12:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast7002.wikimedia.org * 12:18 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast4006.wikimedia.org * 12:18 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host ml-serve1004 * 12:18 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host ml-serve1004 * 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt2002.wikimedia.org * 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast7002.wikimedia.org * 12:10 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backup[2003,2014].codfw.wmnet with reason: reboot & upgrade * 12:10 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt-staging2001.codfw.wmnet * 12:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid1003.eqiad.wmnet * 12:06 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt-staging2001.codfw.wmnet * 12:05 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid1003.eqiad.wmnet * 12:03 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backup[1003,1014].eqiad.wmnet with reason: reboot & upgrade * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid2003.codfw.wmnet * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host irc1003.wikimedia.org * 11:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid2003.codfw.wmnet * 11:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host irc1003.wikimedia.org * 11:55 jmm@dns1004: END - running authdns-update * 11:53 jmm@dns1004: START - running authdns-update * 11:50 jmm@dns1004: END - running authdns-update * 11:48 jmm@dns1004: START - running authdns-update * 11:27 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host irc2003.wikimedia.org * 11:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host irc2003.wikimedia.org * 11:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint2001.codfw.wmnet * 11:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint1001.eqiad.wmnet * 11:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint2001.codfw.wmnet * 11:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint1001.eqiad.wmnet * 11:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-rw2001.wikimedia.org * 11:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-rw1001.wikimedia.org * 11:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-rw2001.wikimedia.org * 11:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-rw1001.wikimedia.org * 11:03 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon1003.wikimedia.org * 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2005.codfw.wmnet * 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2005.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 10:59 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2005.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 10:57 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon1003.wikimedia.org * 10:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon2002.wikimedia.org * 10:55 jmm@cumin2003: START - Cookbook sre.dns.netbox * 10:51 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon2002.wikimedia.org * 10:51 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:50 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2005.codfw.wmnet * 10:41 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:40 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host ml-serve1003 * 10:40 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host ml-serve1003 * 10:39 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2033.codfw.wmnet * 10:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install2005.wikimedia.org * 10:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install1005.wikimedia.org * 10:35 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1004.eqiad.wmnet with OS bookworm * 10:31 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install1005.wikimedia.org * 10:31 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install2005.wikimedia.org * 10:30 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install4004.wikimedia.org * 10:30 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install3004.wikimedia.org * 10:29 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install3004.wikimedia.org * 10:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install4004.wikimedia.org * 10:23 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 10:21 moritzm: failover Ganeti master in codfw/routed to ganeti2034 [[phab:T430928|T430928]] * 10:19 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.addnode (exit_code=0) for new host ganeti2031.codfw.wmnet to cluster codfw and group B * 10:19 moritzm: readded ganeti2031 to the codfw Ganeti cluster [[phab:T430910|T430910]] * 10:18 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1004.eqiad.wmnet with reason: host reimage * 10:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install5004.wikimedia.org * 10:18 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1003 * 10:18 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1003 * 10:17 jmm@cumin2003: START - Cookbook sre.ganeti.addnode for new host ganeti2031.codfw.wmnet to cluster codfw and group B * 10:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install6003.wikimedia.org * 10:16 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install5004.wikimedia.org * 10:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install6003.wikimedia.org * 10:15 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1004.eqiad.wmnet with reason: host reimage * 10:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1001.eqiad.wmnet * 10:14 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 10:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2008.wikimedia.org * 10:00 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ml-serve1004 * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1004 * 09:57 jmm@cumin2003: START - Cookbook sre.dns.netbox * 09:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install7002.wikimedia.org * 09:57 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1004 * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ml-serve1004.eqiad.wmnet 50.48.64.10.in-addr.arpa 0.5.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:57 klausman@cumin1003: START - Cookbook sre.dns.wipe-cache ml-serve1004.eqiad.wmnet 50.48.64.10.in-addr.arpa 0.5.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1004 - klausman@cumin1003" * 09:56 klausman@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1004 - klausman@cumin1003" * 09:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-coord1001.eqiad.wmnet * 09:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 09:55 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow7002.magru.wmnet * 09:52 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-coord1001.eqiad.wmnet * 09:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 09:52 klausman@cumin1003: START - Cookbook sre.dns.netbox * 09:50 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install7002.wikimedia.org * 09:50 klausman@cumin1003: START - Cookbook sre.hosts.move-vlan for host ml-serve1004 * 09:50 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1004.eqiad.wmnet with OS bookworm * 09:50 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1003.eqiad.wmnet with OS bookworm * 09:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1001.eqiad.wmnet * 09:49 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 09:49 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2008.wikimedia.org * 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2007.codfw.wmnet * 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2007.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 09:49 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow7002.magru.wmnet * 09:49 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2007.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 09:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard1003.eqiad.wmnet * 09:39 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard2003.codfw.wmnet * 09:39 jmm@cumin2003: START - Cookbook sre.dns.netbox * 09:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard1003.eqiad.wmnet * 09:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor1003.eqiad.wmnet * 09:35 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard2003.codfw.wmnet * 09:34 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2007.codfw.wmnet * 09:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor1003.eqiad.wmnet * 09:33 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor-dev2001.codfw.wmnet * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor2003.codfw.wmnet * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sretest1006.eqiad.wmnet * 09:27 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 09:25 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor-dev2001.codfw.wmnet * 09:25 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor2003.codfw.wmnet * 09:23 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] (duration: 06m 27s) * 09:23 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2205: codfw rack B4 repool after maintenance * 09:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host sretest1006.eqiad.wmnet * 09:23 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2204: codfw rack B4 repool after maintenance * 09:19 urbanecm@deploy2003: urbanecm: Continuing with deployment * 09:19 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:18 jmm@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 6 hosts with reason: reboot * 09:17 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] * 09:08 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ml-serve1003 * 09:08 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1003 * 09:07 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1003 * 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ml-serve1003.eqiad.wmnet 81.32.64.10.in-addr.arpa 1.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:07 klausman@cumin1003: START - Cookbook sre.dns.wipe-cache ml-serve1003.eqiad.wmnet 81.32.64.10.in-addr.arpa 1.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1003 - klausman@cumin1003" * 09:06 klausman@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1003 - klausman@cumin1003" * 08:58 klausman@cumin1003: START - Cookbook sre.dns.netbox * 08:57 klausman@cumin1003: START - Cookbook sre.hosts.move-vlan for host ml-serve1003 * 08:57 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1003.eqiad.wmnet with OS bookworm * 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=0) rolling reimage on P<nowiki>{</nowiki>ml-serve1003.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet * 08:55 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet * 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1003.eqiad.wmnet with OS bookworm * 08:39 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 08:38 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool db2205: codfw rack B4 repool after maintenance * 08:37 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool db2204: codfw rack B4 repool after maintenance * 08:36 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 08:35 hashar@deploy2003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 08:32 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:32 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:31 hashar@deploy2003: Rolling back deployment * 08:26 moritzm: failover Ganeti master in codfw to ganeti2048 [[phab:T430928|T430928]] * 08:16 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1003.eqiad.wmnet with OS bookworm * 08:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2004.codfw.wmnet * 08:16 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet * 08:16 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet * 08:16 klausman@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on P<nowiki>{</nowiki>ml-serve1003.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 08:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2002.codfw.wmnet * 08:15 XioNoX: lsw1-b4-codfw> request system reboot - [[phab:T430910|T430910]] * 08:15 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b4-codfw,lsw1-b4-codfw IPv6,lsw1-b4-codfw.mgmt with reason: Switch maintenance * 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for codfw rack B4 * 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:10 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2004.codfw.wmnet * 08:10 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2002.codfw.wmnet * 08:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2205: codfw rack B4 depool for maintenance * 08:08 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool db2205: codfw rack B4 depool for maintenance * 08:08 jmm@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin2003.codfw.wmnet * 08:08 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2204: codfw rack B4 depool for maintenance * 08:08 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool db2204: codfw rack B4 depool for maintenance * 08:08 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 27 hosts with reason: codfw rack B4 depool for maintenance * 08:03 jmm@cumin2002: START - Cookbook sre.hosts.reboot-single for host cumin2003.codfw.wmnet * 07:56 ayounsi@cumin1003: START - Cookbook sre.network.depool-rack with action 'depool' for codfw rack B4 * 07:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1008.eqiad.wmnet with OS trixie * 07:49 wmde-fisch@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] (duration: 08m 36s) * 07:44 wmde-fisch@deploy2003: wmde-fisch: Continuing with deployment * 07:43 wmde-fisch@deploy2003: wmde-fisch: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:41 wmde-fisch@deploy2003: Started scap sync-world: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] * 07:35 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1008.eqiad.wmnet with reason: host reimage * 07:31 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1008.eqiad.wmnet with reason: host reimage * 07:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1008.eqiad.wmnet with OS trixie * 07:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 07:00 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 06:59 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1008.eqiad.wmnet with OS trixie * 06:57 Emperor: rebalance thanos swift rings after previous re-image of thanos-fe1004 to trixie * 06:47 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1008.eqiad.wmnet with OS trixie * 04:10 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 14 days, 0:00:00 on cp6008.drmrs.wmnet with reason: Hardware failure - [[phab:T431651|T431651]] * 03:55 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp6008.* * 03:29 ryankemper: [[phab:T431311|T431311]] Repooled eqiad cirrussearch clusters (`chi/omega/psi`) following completion of OpenSearch 2.19 migration * 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad * 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=eqiad * 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 31s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-08 == * 23:52 Amir1: ladsgroup@deploy2003:~$ mwscript-k8s --follow -- extensions/ORES/maintenance/PurgeScoreCache.php --wiki=simplewiki --model damaging --old ([[phab:T431159|T431159]]) * 23:46 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 23:46 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing PTR for 2001:df2:e500:fe08::1 - cmooney@cumin1003" * 23:46 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing PTR for 2001:df2:e500:fe08::1 - cmooney@cumin1003" * 23:40 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 23:16 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 23:15 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 22:42 rzl@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 22:40 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] (duration: 12m 55s) * 22:40 rzl@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 22:37 rzl@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 22:36 rzl@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 22:35 rzl@deploy2003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 22:34 urbanecm@deploy2003: urbanecm: Continuing with deployment * 22:33 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:33 rzl@deploy2003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 22:32 rzl@deploy2003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 22:30 rzl@deploy2003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 22:30 rzl@deploy2003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 22:29 rzl@deploy2003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 22:27 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] * 22:26 rzl@deploy2003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 22:22 rzl@deploy2003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 22:21 rzl@deploy2003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 22:19 rzl@deploy2003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 22:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 22:17 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 22:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 22:13 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 22:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 22:13 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 22:09 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 22:06 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 22:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1094.eqiad.wmnet with OS trixie * 22:01 urbanecm: Make https://test.wikipedia.org/w/index.php?title=MediaWiki:GrowthExperimentsSuggestedEdits.json&diff=prev&oldid=750552 with GrowthExperiments disabled (via mw-experimental), then run `\MediaWiki\MediaWikiServices::getInstance()->get('CommunityConfiguration.ProviderFactory')->newProvider('GrowthSuggestedEdits')->getStore()->invalidate()` ([[phab:T431625|T431625]]) * 21:56 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d2-codfw * 21:55 urbanecm@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 21:55 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d2-codfw * 21:55 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c4-codfw * 21:55 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c4-codfw * 21:55 urbanecm@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2002 * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2002 * 21:54 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2002 * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2002.codfw.wmnet 50.32.192.10.in-addr.arpa 0.5.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:54 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2002.codfw.wmnet 50.32.192.10.in-addr.arpa 0.5.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2002 - bking@cumin2003" * 21:54 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2002 - bking@cumin2003" * 21:49 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:49 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2002 * 21:49 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2002.codfw.wmnet with OS trixie * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1094.eqiad.wmnet with reason: host reimage * 21:42 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 21:39 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 21:37 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1094.eqiad.wmnet with reason: host reimage * 21:36 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 21:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 21:29 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 21:27 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 21:22 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1094.eqiad.wmnet with OS trixie * 21:21 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host restbase2039.codfw.wmnet with OS bullseye * 21:21 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin2002" * 21:21 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin2002" * 21:04 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on restbase2039.codfw.wmnet with reason: host reimage * 21:00 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on restbase2039.codfw.wmnet with reason: host reimage * 20:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1073.eqiad.wmnet with OS trixie * 20:48 mutante: deploy2003 - kill 1102 (stunnel4) ; systemctl start stunnel4 ([[phab:T418262|T418262]]) * 20:42 cjming@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] (duration: 33m 02s) * 20:42 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host restbase2039.codfw.wmnet with OS bullseye * 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1073.eqiad.wmnet with reason: host reimage * 20:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1073.eqiad.wmnet with reason: host reimage * 20:30 cjming@deploy2003: cjming: Continuing with deployment * 20:28 cjming@deploy2003: cjming: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1098.eqiad.wmnet with OS trixie * 20:13 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1073.eqiad.wmnet with OS trixie * 20:09 cjming@deploy2003: Started scap sync-world: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] * 20:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1098.eqiad.wmnet with reason: host reimage * 19:56 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1098.eqiad.wmnet with reason: host reimage * 19:55 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d4-codfw * 19:54 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d4-codfw * 19:54 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c1-codfw * 19:54 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c1-codfw * 19:52 mutante: restarting gerrit on gerrit.wikimedia.org (gerrit2003) * 19:48 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2331.codfw.wmnet * 19:48 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2331.codfw.wmnet * 19:48 mutante: restarting gerrit on gerrit-replica.wikimedia.org (gerrit1003) * 19:46 mutante: restarting gerrit on gerrit-spare.wikimedia.org (gerrit2002) * 19:43 jasmine@cumin2002: conftool action : set/pooled=yes; selector: name=wikikube-worker2331.codfw.wmnet,cluster=kubernetes,service=kubesvc * 19:43 jasmine@cumin2002: conftool action : set/weight=10; selector: name=wikikube-worker2331.codfw.wmnet,cluster=kubernetes,service=kubesvc * 19:40 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1098.eqiad.wmnet with OS trixie * 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d5-codfw * 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d5-codfw * 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c7-codfw * 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c7-codfw * 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c5-codfw * 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c5-codfw * 19:30 jasmine_: ran homer on lsw1-d8-codfw, adding wikikube-worker2331 to cluster * 19:29 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1100.eqiad.wmnet with OS trixie * 19:20 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d8-codfw * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d8-codfw * 19:19 mutante: gerrit - replacing private key for registerEmail verification * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-magru * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device cr2-magru * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d7-codfw * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d7-codfw * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d3-codfw * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d1-codfw * 19:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d1-codfw * 19:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c2-codfw * 19:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c2-codfw * 19:11 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-codfw * 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-magru * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device cr1-magru * 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d8-codfw * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d8-codfw * 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d6-codfw * 19:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1100.eqiad.wmnet with reason: host reimage * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d6-codfw * 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c6-codfw * 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c6-codfw * 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c3-codfw * 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c3-codfw * 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b4-magru * 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device asw1-b4-magru * 19:08 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b3-magru * 19:08 cmooney@cumin1003: START - Cookbook sre.network.tls for network device asw1-b3-magru * 19:05 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1100.eqiad.wmnet with reason: host reimage * 19:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1122.eqiad.wmnet with OS trixie * 19:00 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 18:59 topranks: rolling out update to BGP ACL on Nokia Switches eqiad, codfw & ulsfo [[phab:T425703|T425703]] * 18:58 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 18:57 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 18:55 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 18:53 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 18:52 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 18:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1100.eqiad.wmnet with OS trixie * 18:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1068.eqiad.wmnet with OS trixie * 18:47 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1102.eqiad.wmnet with OS trixie * 18:47 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 18:46 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1122.eqiad.wmnet with reason: host reimage * 18:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1122.eqiad.wmnet with reason: host reimage * 18:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1068.eqiad.wmnet with reason: host reimage * 18:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1122 * 18:26 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1122 * 18:25 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1122 * 18:25 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1122.eqiad.wmnet 31.48.64.10.in-addr.arpa 1.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:25 bking@cumin2003: START - Cookbook sre.dns.wipe-cache cirrussearch1122.eqiad.wmnet 31.48.64.10.in-addr.arpa 1.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:25 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:25 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1122 - bking@cumin2003" * 18:25 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1122 - bking@cumin2003" * 18:21 rzl@deploy2003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 18:21 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1068.eqiad.wmnet with reason: host reimage * 18:21 rzl@deploy2003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 18:21 rzl@deploy2003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 18:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 18:19 rzl@deploy2003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 18:19 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:18 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1122 * 18:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1122.eqiad.wmnet with OS trixie * 18:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 18:15 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 18:13 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 18:13 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 18:10 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 18:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1068.eqiad.wmnet with OS trixie * 18:01 kamila@deploy2003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 18m 29s) * 18:00 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:55 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:42 kamila@deploy2003: Started scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] * 17:42 kamila@deploy2003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 19m 50s) * 17:42 kamila@deploy2003: Rolling back deployment * 17:35 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:31 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet * 17:18 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet * 17:16 kamila@deploy1003: Unlocked for deployment [MediaWiki]: switching deployment server (duration: 22m 07s) * 17:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 17:11 kamila@dns1005: END - running authdns-update * 17:09 kamila@dns1005: START - running authdns-update * 17:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 17:04 jasmine@cumin2002: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1164.eqiad.wmnet * 17:04 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1164.eqiad.wmnet * 17:04 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1164.eqiad.wmnet * 16:56 kamila@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on releases2003.codfw.wmnet,releases1003.eqiad.wmnet with reason: Deployment server switchover * 16:54 kamila@deploy1003: Locking from deployment [MediaWiki]: switching deployment server * 16:53 kamila@deploy1003: Unlocked for deployment [MediaWiki]: switching deployment server (duration: 04m 02s) * 16:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie * 16:49 kamila@deploy1003: Locking from deployment [MediaWiki]: switching deployment server * 16:46 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1095.eqiad.wmnet with OS trixie * 16:45 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1093.eqiad.wmnet with OS trixie * 16:43 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1164.eqiad.wmnet with OS trixie * 16:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1095.eqiad.wmnet with reason: host reimage * 16:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 16:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 16:23 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1164.eqiad.wmnet with reason: host reimage * 16:18 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on cirrussearch1093.eqiad.wmnet with reason: host reimage * 16:16 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1164.eqiad.wmnet with reason: host reimage * 16:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1095.eqiad.wmnet with reason: host reimage * 16:09 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 16:09 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 16:08 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1093.eqiad.wmnet with reason: host reimage * 15:59 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Pool test * 15:59 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:59 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 15:59 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Pool test * 15:58 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Depool test * 15:58 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:58 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 15:58 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Depool test * 15:57 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1164 * 15:57 jasmine@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1164 * 15:57 jasmine@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1164 * 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1164.eqiad.wmnet 114.48.64.10.in-addr.arpa 4.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:56 jasmine@cumin2002: START - Cookbook sre.dns.wipe-cache wikikube-worker1164.eqiad.wmnet 114.48.64.10.in-addr.arpa 4.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1164 - jasmine@cumin2002" * 15:56 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1164 - jasmine@cumin2002" * 15:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1093.eqiad.wmnet with OS trixie * 15:51 jasmine@cumin2002: START - Cookbook sre.dns.netbox * 15:51 jasmine@cumin2002: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1164 * 15:50 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-worker1164.eqiad.wmnet with OS trixie * 15:50 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1164.eqiad.wmnet * 15:50 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1164.eqiad.wmnet * 15:50 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1164.eqiad.wmnet * 15:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1095.eqiad.wmnet with OS trixie * 15:42 jasmine@cumin2002: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1164.eqiad.wmnet * 15:42 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1164.eqiad.wmnet * 15:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:42 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1164.eqiad.wmnet * 15:42 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1164.eqiad.wmnet * 15:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 15:39 elukey@cumin1003: START - Cookbook sre.hosts.provision for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 15:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1007.eqiad.wmnet with OS trixie * 15:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Pool test * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1007.eqiad.wmnet with reason: host reimage * 15:15 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 15:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet * 15:15 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 15:15 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1007.eqiad.wmnet with reason: host reimage * 15:15 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:14 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Pool test * 15:14 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet * 15:14 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2228: Depool test * 15:14 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db2228: Depool test * 15:10 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 15:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:08 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 15:06 blake@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 15:06 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet * 15:06 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 15:06 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 15:05 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet * 15:05 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 15:04 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:04 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 15:04 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:03 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 15:03 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 15:03 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:03 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T430909|T430909]] * 15:03 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:03 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 15:03 swfrench-wmf: restarted eqsin, codfw confds - [[phab:T430909|T430909]] * 15:03 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test1002.eqiad.wmnet * 15:02 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet * 14:59 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:59 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:55 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1007.eqiad.wmnet with OS trixie * 14:52 swfrench-wmf: restarted ulsfo confds, confirmed now connected to codfw backends except those using wikimedia.org SRV record - [[phab:T430909|T430909]] * 14:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:49 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:41 moritzm: uninstalling dhcpcd-base from trixie hosts which still have it installed [[phab:T414341|T414341]] * 14:40 sukhe: sudo cumin -b1 -s120 "P<nowiki>{</nowiki>lvs2011*<nowiki>}</nowiki> or P<nowiki>{</nowiki>lvs2012*<nowiki>}</nowiki>" "systemctl restart pybal.service" * 14:39 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:39 mvernon@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host thanos-be1007.eqiad.wmnet with OS trixie * 14:37 sukhe: restart pybal on lvs2013 to revert back to conf2004 * 14:35 sukhe: restart pybal on lvs2014 to revert back to conf2004 * 14:34 swfrench-wmf: switched codfw, eqsin, ulsfo etcd client SRV records back to codfw - [[phab:T430909|T430909]] * 14:32 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1002.eqiad.wmnet * 14:32 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet * 14:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1007.eqiad.wmnet with OS trixie * 14:31 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:31 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:31 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:30 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:30 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Pool test * 14:30 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:29 swfrench@dns1004: END - running authdns-update * 14:29 moritzm: installing jackson-core security updates * 14:27 swfrench@dns1004: START - running authdns-update * 14:22 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:22 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:22 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1119.eqiad.wmnet with OS trixie * 14:22 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:21 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:20 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:20 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 14:20 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:19 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 14:19 moritzm: installing librabbitmq security updates * 14:19 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1002.eqiad.wmnet * 14:18 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet * 14:16 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:15 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:15 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:14 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Pool test * 14:14 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox) * 14:14 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet * 14:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 14:08 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 14:07 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet * 14:05 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1118.eqiad.wmnet with OS trixie * 14:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test1001.eqiad.wmnet * 14:00 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-worker@eqiad * 14:00 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 13:59 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 13:58 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 13:57 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1119.eqiad.wmnet with reason: host reimage * 13:54 moritzm: installing libcap2 security updates * 13:53 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1119.eqiad.wmnet with reason: host reimage * 13:52 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet * 13:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) * 13:52 fceratto@cumin1003: START - Cookbook sre.mysql.depool * 13:50 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-worker@eqiad * 13:50 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1051.eqiad.wmnet * 13:50 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1051.eqiad.wmnet * 13:50 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1051.eqiad.wmnet * 13:49 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 13:45 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1081.eqiad.wmnet with OS trixie * 13:41 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1119 * 13:41 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1119 * 13:40 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1119 * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1119.eqiad.wmnet 97.32.64.10.in-addr.arpa 7.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1119.eqiad.wmnet 97.32.64.10.in-addr.arpa 7.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1119 - atsuko@cumin1003" * 13:40 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1119 - atsuko@cumin1003" * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1118.eqiad.wmnet with reason: host reimage * 13:39 moritzm: installing krb5 security updates * 13:37 Lucas_WMDE: UTC afternoon backport+config window done * 13:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1006.eqiad.wmnet with OS trixie * 13:36 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1118.eqiad.wmnet with reason: host reimage * 13:36 atsuko@cumin1003: START - Cookbook sre.dns.netbox * 13:35 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] (duration: 07m 46s) * 13:34 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1119 * 13:34 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1119.eqiad.wmnet with OS trixie * 13:30 sbisson@deploy1003: sbisson: Continuing with deployment * 13:30 moritzm: installing openssh security updates * 13:30 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling restart_daemons on A:wikidough * 13:29 sbisson@deploy1003: sbisson: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:27 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] * 13:26 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1051.eqiad.wmnet with OS trixie * 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1081.eqiad.wmnet with reason: host reimage * 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1118 * 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1118 * 13:22 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] (duration: 12m 12s) * 13:21 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1081.eqiad.wmnet with reason: host reimage * 13:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1006.eqiad.wmnet with reason: host reimage * 13:18 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1118 * 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1118.eqiad.wmnet 90.32.64.10.in-addr.arpa 0.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:18 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1118.eqiad.wmnet 90.32.64.10.in-addr.arpa 0.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1118 - atsuko@cumin1003" * 13:18 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1118 - atsuko@cumin1003" * 13:17 stran@deploy1003: stran: Continuing with deployment * 13:16 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough * 13:15 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:13 atsuko@cumin1003: START - Cookbook sre.dns.netbox * 13:12 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1118 * 13:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1006.eqiad.wmnet with reason: host reimage * 13:12 stran@deploy1003: stran: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:12 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1118.eqiad.wmnet with OS trixie * 13:10 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] * 13:05 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-worker@codfw * 13:05 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 13:05 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1081.eqiad.wmnet with OS trixie * 13:05 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1051.eqiad.wmnet with reason: host reimage * 13:04 moritzm: installing jq security updates * 13:04 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 13:01 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1051.eqiad.wmnet with reason: host reimage * 12:58 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-worker@codfw * 12:52 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 12:50 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host thanos-be1006.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1051 * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1051 * 12:43 moritzm: installing Python 3.11 security updates * 12:43 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1051 * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1051.eqiad.wmnet 46.32.64.10.in-addr.arpa 6.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:43 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1051.eqiad.wmnet 46.32.64.10.in-addr.arpa 6.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1051 - blake@cumin1003" * 12:43 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1051 - blake@cumin1003" * 12:38 blake@cumin1003: START - Cookbook sre.dns.netbox * 12:38 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1051 * 12:38 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1051.eqiad.wmnet with OS trixie * 12:37 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1051.eqiad.wmnet * 12:36 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1051.eqiad.wmnet * 12:36 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1051.eqiad.wmnet * 12:34 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1006.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 12:34 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1006.eqiad.wmnet with OS trixie * 12:27 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 12:27 mvernon@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host thanos-be1006.eqiad.wmnet with OS trixie * 12:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:02 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 12:01 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1006.eqiad.wmnet with OS trixie * 11:43 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 11:38 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1076.eqiad.wmnet with OS trixie * 11:26 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1075.eqiad.wmnet with OS trixie * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2047.codfw.wmnet * 11:19 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2047.codfw.wmnet * 11:19 moritzm: temporarily remove ganeti2031 from codfw cluster [[phab:T430910|T430910]] * 11:08 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1076.eqiad.wmnet with reason: host reimage * 11:08 moritzm: installing Linux 6.1.176 on Bookworm servers * 11:03 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1076.eqiad.wmnet with reason: host reimage * 11:00 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1075.eqiad.wmnet with reason: host reimage * 10:56 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1075.eqiad.wmnet with reason: host reimage * 10:47 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1076.eqiad.wmnet with OS trixie * 10:46 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1074.eqiad.wmnet with OS trixie * 10:45 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1005.eqiad.wmnet with OS trixie * 10:40 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1075.eqiad.wmnet with OS trixie * 10:32 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2031.codfw.wmnet * 10:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1005.eqiad.wmnet with reason: host reimage * 10:25 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1005.eqiad.wmnet with reason: host reimage * 10:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1074.eqiad.wmnet with reason: host reimage * 10:17 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1074.eqiad.wmnet with reason: host reimage * 10:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1005.eqiad.wmnet with OS trixie * 10:12 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet * 10:04 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 10:01 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet * 10:01 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1074.eqiad.wmnet with OS trixie * 10:01 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 09:43 cgoubert@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/aux-k8s-services/redioscope: apply * 09:43 cgoubert@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/aux-k8s-services/redioscope: apply * 09:43 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply * 09:35 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply * 09:34 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 09:34 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 09:33 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 41 days, 15:00:00 on db2252.codfw.wmnet with reason: Test * 09:32 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 09:32 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 09:31 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: codfw rack B3 pool after maintenance * 09:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 09:07 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 09:07 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 09:02 ladsgroup@cumin1003: END (PASS) - Cookbook sre.mysql.sanitarium_restart (exit_code=0) * 08:57 topranks: merge patch to shift eqiad <-> esams traffic onto new 40G circuit * 08:54 hashar@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1004.eqiad.wmnet with OS trixie * 08:50 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 08:50 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitarium_restart (exit_code=99) * 08:50 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 08:45 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool es2051: codfw rack B3 pool after maintenance * 08:44 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:44 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:43 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2007.codfw.wmnet * 08:43 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2007.codfw.wmnet * 08:42 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2031.codfw.wmnet * 08:41 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2031.codfw.wmnet * 08:40 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2031.codfw.wmnet * 08:38 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:38 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:35 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.sanitize-wiki (exit_code=97) Managing sanitization for wikis minwikiquote in section s3 * 08:33 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis minwikiquote in section s3 * 08:32 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Checking sanitization for wikis minwikiquote in section s5 * 08:30 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Checking sanitization for wikis minwikiquote in section s5 * 08:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Managing sanitization for wikis minwikiquote in section s5 * 08:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1004.eqiad.wmnet with reason: host reimage * 08:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1004.eqiad.wmnet with reason: host reimage * 08:23 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:22 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis minwikiquote in section s5 * 08:19 XioNoX: lsw1-b3-codfw> request system reboot - [[phab:T430909|T430909]] * 08:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Checking sanitization for wikis minwikiquote in section s5 * 08:17 hashar@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Checking sanitization for wikis minwikiquote in section s5 * 08:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for codfw rack B3 * 08:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2007.codfw.wmnet * 08:15 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lsw1-b3-codfw,lsw1-b3-codfw IPv6,lsw1-b3-codfw.mgmt with reason: Switch maintenance * 08:15 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2007.codfw.wmnet * 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:07 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:06 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: codfw rack B3 depool for maintenance * 08:05 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool es2051: codfw rack B3 depool for maintenance * 08:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1004.eqiad.wmnet with OS trixie * 08:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1005.eqiad.wmnet with OS trixie * 08:03 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 21 hosts with reason: codfw rack B3 depool for maintenance * 07:56 ayounsi@cumin1003: START - Cookbook sre.network.depool-rack with action 'depool' for codfw rack B3 * 07:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1005.eqiad.wmnet with reason: host reimage * 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1005.eqiad.wmnet with reason: host reimage * 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1005.eqiad.wmnet with OS bookworm * 07:29 moritzm: installing gnutls28 security updates * 07:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1005.eqiad.wmnet with OS trixie * 07:13 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1125.eqiad.wmnet with OS trixie * 07:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1005.eqiad.wmnet with reason: host reimage * 07:07 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aux-k8s-etcd1005.eqiad.wmnet with reason: host reimage * 06:56 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1005.eqiad.wmnet with OS bookworm * 06:54 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1125.eqiad.wmnet with reason: host reimage * 06:52 elukey: upgrade all trixie hosts to pywmflib 3.1 - [[phab:T430552|T430552]] * 06:50 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1125.eqiad.wmnet with reason: host reimage * 06:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 06:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 06:38 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1125.eqiad.wmnet with OS trixie * 05:42 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1107.eqiad.wmnet with OS trixie * 05:35 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1124.eqiad.wmnet with OS trixie * 05:31 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1101.eqiad.wmnet with OS trixie * 05:21 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1107.eqiad.wmnet with reason: host reimage * 05:17 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1124.eqiad.wmnet with reason: host reimage * 05:13 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1107.eqiad.wmnet with reason: host reimage * 05:13 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1101.eqiad.wmnet with reason: host reimage * 05:11 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1124.eqiad.wmnet with reason: host reimage * 05:10 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1101.eqiad.wmnet with reason: host reimage * 04:58 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1124.eqiad.wmnet with OS trixie * 04:56 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1107.eqiad.wmnet with OS trixie * 04:55 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1101.eqiad.wmnet with OS trixie * 02:27 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] (duration: 08m 14s) * 02:22 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 02:21 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 02:19 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] * 01:59 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] (duration: 09m 46s) * 01:55 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 01:51 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 01:49 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] * 01:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1099.eqiad.wmnet with OS trixie * 00:57 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1110.eqiad.wmnet with OS trixie * 00:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1099.eqiad.wmnet with reason: host reimage * 00:41 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1099.eqiad.wmnet with reason: host reimage * 00:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1110.eqiad.wmnet with reason: host reimage * 00:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1110.eqiad.wmnet with reason: host reimage * 00:26 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1099.eqiad.wmnet with OS trixie * 00:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1110.eqiad.wmnet with OS trixie == 2026-07-07 == * 22:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1097.eqiad.wmnet with OS trixie * 22:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1097.eqiad.wmnet with reason: host reimage * 22:24 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1097.eqiad.wmnet with reason: host reimage * 22:14 hashar: Restarting Gerrit on gerrit2002 and gerrit1003 (replicas) * 22:09 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1097.eqiad.wmnet with OS trixie * 22:07 hashar: Restarting Gerrit on gerrit2003 * 21:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 21:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 21:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 21:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 21:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1108.eqiad.wmnet with OS trixie * 20:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1091.eqiad.wmnet with OS trixie * 20:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1090.eqiad.wmnet with OS trixie * 20:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1108.eqiad.wmnet with reason: host reimage * 20:36 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1091.eqiad.wmnet with reason: host reimage * 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1090.eqiad.wmnet with reason: host reimage * 20:33 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1091.eqiad.wmnet with reason: host reimage * 20:30 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1108.eqiad.wmnet with reason: host reimage * 20:30 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1006.eqiad.wmnet * 20:30 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1090.eqiad.wmnet with reason: host reimage * 20:30 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1006.eqiad.wmnet * 20:27 jasmine_: "homer lsw1-c2-eqiad* commit "Added new stacked control plane wikikube-ctrl1006"" * 20:22 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] (duration: 07m 29s) * 20:20 jasmine_: "homer "cr*eqiad*" commit "Added new stacked control plane wikikube-ctrl1006"" * 20:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1091.eqiad.wmnet with OS trixie * 20:17 arlolra@deploy1003: arlolra: Continuing with deployment * 20:16 arlolra@deploy1003: arlolra: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:16 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1090.eqiad.wmnet with OS trixie * 20:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1108.eqiad.wmnet with OS trixie * 20:14 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] * 20:09 cwhite: remove 2026-04 swift log archives from centrallog2002 to free some space * 20:01 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=93) for host cirrussearch1108.eqiad.wmnet with OS trixie * 19:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1108.eqiad.wmnet with OS trixie * 19:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1090.eqiad.wmnet with OS trixie * 19:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1109.eqiad.wmnet with OS trixie * 19:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1092.eqiad.wmnet with OS trixie * 19:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1123.eqiad.wmnet with OS trixie * 19:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1109.eqiad.wmnet with reason: host reimage * 19:19 cdobbins@cumin2002: conftool action : set/pooled=yes; selector: name=dns7002.* * 19:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1092.eqiad.wmnet with reason: host reimage * 19:17 jasmine@dns1004: END - running authdns-update * 19:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1109.eqiad.wmnet with reason: host reimage * 19:15 jasmine@dns1004: START - running authdns-update * 19:15 cdobbins@dns1004: END - running authdns-update * 19:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1123.eqiad.wmnet with reason: host reimage * 19:13 cdobbins@dns1004: START - running authdns-update * 19:12 cdobbins@cumin2002: conftool action : set/pooled=yes; selector: name=dns7002.*,service=authdns-update * 19:11 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1092.eqiad.wmnet with reason: host reimage * 19:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1123.eqiad.wmnet with reason: host reimage * 18:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1123.eqiad.wmnet with OS trixie * 18:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1109.eqiad.wmnet with OS trixie * 18:56 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1092.eqiad.wmnet with OS trixie * 18:52 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 18:49 swfrench@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 18:40 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 18:38 swfrench@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 18:11 swfrench-wmf: restarted eqsin, codfw confds - [[phab:T430909|T430909]] * 18:01 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T430909|T430909]] * 17:59 swfrench-wmf: restarted ulsfo confds, confirmed now connected to eqiad backends - [[phab:T430909|T430909]] * 17:52 sukhe: restart pybal on lvs2011 to switch from conf2004 to conf1008: [[phab:T430909|T430909]] * 17:51 sukhe: restart pybal on lvs2012 to switch from conf2004 to conf1008 [puppet re-enabled there]: [[phab:T430909|T430909]] * 17:46 sukhe: restart pybal on lvs2013 to switch from conf2004 to conf1008: [[phab:T430909|T430909]] * 17:44 swfrench-wmf: switched codfw, eqsin, ulsfo etcd client SRV records to eqiad - [[phab:T430909|T430909]] * 17:43 swfrench@dns1004: END - running authdns-update * 17:40 swfrench@dns1004: START - running authdns-update * 17:40 sukhe: restart pybal on lvs2014 to switch from conf2004 to conf1008: [[phab:T430909|T430909]] * 17:21 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1003.eqiad.wmnet * 17:15 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1003.eqiad.wmnet * 17:14 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1002.eqiad.wmnet * 17:06 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1002.eqiad.wmnet * 17:06 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-low-traffic-codfw' 'systemctl restart pybal.service' # lvs2013, [[phab:T416623|T416623]] * 17:04 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1001.eqiad.wmnet * 17:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1111.eqiad.wmnet with OS trixie * 17:00 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal.service' # lvs2014, [[phab:T416623|T416623]] * 16:58 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1001.eqiad.wmnet * 16:58 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS bookworm * 16:55 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-low-traffic-eqiad' 'systemctl restart pybal.service' # lvs1019, [[phab:T416623|T416623]] * 16:53 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-secondary-eqiad' 'systemctl restart pybal.service' # lvs1020, [[phab:T416623|T416623]] * 16:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1111.eqiad.wmnet with reason: host reimage * 16:40 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1111.eqiad.wmnet with reason: host reimage * 16:38 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.peering (exit_code=99) with action 'configure' for AS: 47794 * 16:35 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 47794 * 16:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1111.eqiad.wmnet with OS trixie * 16:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1006.eqiad.wmnet with OS trixie * 16:06 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 16:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1006.eqiad.wmnet with reason: host reimage * 15:58 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1121.eqiad.wmnet with OS trixie * 15:58 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 15:56 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1006.eqiad.wmnet with reason: host reimage * 15:54 mutante: jenkins down in planned maintenance window * 15:42 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1037.eqiad.wmnet * 15:42 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1037.eqiad.wmnet * 15:42 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1037.eqiad.wmnet * 15:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1006.eqiad.wmnet with OS trixie * 15:34 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1121.eqiad.wmnet with reason: host reimage * 15:33 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS bookworm * 15:33 cdobbins@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host dns7002.wikimedia.org with OS trixie * 15:30 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1121.eqiad.wmnet with reason: host reimage * 15:29 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host clouddumps1001.wikimedia.org * 15:20 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1001.wikimedia.org * 15:18 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1121 * 15:18 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1121 * 15:18 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host clouddumps1002.wikimedia.org * 15:17 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1121 * 15:17 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:17 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply * 15:16 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply * 15:16 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:16 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1121 - atsuko@cumin1003" * 15:16 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1121 - atsuko@cumin1003" * 15:14 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1037.eqiad.wmnet with OS trixie * 15:11 atsuko@cumin1003: START - Cookbook sre.dns.netbox * 15:09 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org * 15:09 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1121 * 15:09 andrew@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host clouddumps1002.wikimedia.org * 15:09 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org * 15:09 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1121.eqiad.wmnet with OS trixie * 15:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1007.eqiad.wmnet with OS trixie * 15:08 andrew@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host clouddumps1002.wikimedia.org * 15:08 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org * 15:05 brennen@deploy1003: Finished deploy [phabricator/deployment@7e02037]: deploy phab1004 for [[phab:T431440|T431440]] (duration: 00m 47s) * 15:04 brennen@deploy1003: Started deploy [phabricator/deployment@7e02037]: deploy phab1004 for [[phab:T431440|T431440]] * 15:03 brennen@deploy1003: Finished deploy [phabricator/deployment@7e02037]: deploy phab2003 for [[phab:T431440|T431440]] (duration: 00m 51s) * 15:03 brennen@deploy1003: Started deploy [phabricator/deployment@7e02037]: deploy phab2003 for [[phab:T431440|T431440]] * 15:00 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71] (thin): Regular analytics weekly train THIN [analytics/refinery@7d8dc71f] (duration: 02m 10s) * 14:58 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71] (thin): Regular analytics weekly train THIN [analytics/refinery@7d8dc71f] * 14:58 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71]: Regular analytics weekly train [analytics/refinery@7d8dc71f] (duration: 04m 14s) * 14:54 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1037.eqiad.wmnet with reason: host reimage * 14:53 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71]: Regular analytics weekly train [analytics/refinery@7d8dc71f] * 14:53 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@7d8dc71f] (duration: 02m 00s) * 14:51 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@7d8dc71f] * 14:51 arnaudb@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on phab2003.codfw.wmnet,phab[1004-1006].eqiad.wmnet with reason: maintenance * 14:51 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1037.eqiad.wmnet with reason: host reimage * 14:50 JavierMonton: Deploying Refinery at {{Gerrit|7d8dc71f}} for change {{Gerrit|1308087}} / [[phab:T431318|T431318]] - update filerevision table sqoop and table * 14:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1007.eqiad.wmnet with reason: host reimage * 14:42 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1007.eqiad.wmnet with reason: host reimage * 14:40 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1083.eqiad.wmnet with OS trixie * 14:37 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:36 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:35 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-master@eqiad * 14:35 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 14:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 14:34 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1037 * 14:34 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1037 * 14:34 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:cleanMentorList.php --wiki=frwiki # [[phab:T427386|T427386]] * 14:34 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 14:34 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308112{{!}}Revert^2 "[Growth] frwiki: Deploy automated mentor list cleaner" (T427386)]] (duration: 06m 47s) * 14:34 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 14:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:33 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:32 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1037 * 14:31 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:31 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:29 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:29 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-master@eqiad * 14:29 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:29 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:29 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:28 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:27 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:27 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1308112{{!}}Revert^2 "[Growth] frwiki: Deploy automated mentor list cleaner" (T427386)]] * 14:26 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1007.eqiad.wmnet with OS trixie * 14:26 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:26 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 14:26 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:25 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:25 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:25 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:cleanMentorList.php --wiki=frwiki # [[phab:T427386|T427386]] * 14:24 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:24 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1037 - blake@cumin1003" * 14:24 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1037 - blake@cumin1003" * 14:20 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1083.eqiad.wmnet with reason: host reimage * 14:19 blake@cumin1003: START - Cookbook sre.dns.netbox * 14:19 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-master@codfw * 14:19 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 14:19 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1037 * 14:18 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1037.eqiad.wmnet with OS trixie * 14:18 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1037.eqiad.wmnet * 14:18 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 14:18 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1037.eqiad.wmnet * 14:18 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1037.eqiad.wmnet * 14:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2007.codfw.wmnet with OS trixie * 14:16 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1083.eqiad.wmnet with reason: host reimage * 14:15 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1036.eqiad.wmnet * 14:15 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1036.eqiad.wmnet * 14:14 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1036.eqiad.wmnet * 14:12 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-master@codfw * 14:11 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1120.eqiad.wmnet with OS trixie * 14:05 moritzm: installing distro-info-data updates from trixie/bookworm point releases * 14:04 fabfur: disable puppet on A:cp-text to selectively apply https://gerrit.wikimedia.org/r/c/operations/puppet/+/1308040 * 14:03 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] (duration: 27m 48s) * 14:00 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 14:00 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1083.eqiad.wmnet with OS trixie * 13:58 urbanecm@deploy1003: urbanecm: Continuing with deployment * 13:58 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:58 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1004.eqiad.wmnet with OS bookworm * 13:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2007.codfw.wmnet with reason: host reimage * 13:57 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 13:53 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1120.eqiad.wmnet with reason: host reimage * 13:50 moritzm: installing Linux 5.10.259 on Bullseye hosts * 13:47 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply * 13:47 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply * 13:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2007.codfw.wmnet with reason: host reimage * 13:46 cgoubert@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/aux-k8s-services/redioscope: apply * 13:46 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1120.eqiad.wmnet with reason: host reimage * 13:46 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:46 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:45 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:44 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:44 cgoubert@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/aux-k8s-services/redioscope: apply * 13:40 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 13:39 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:38 moritzm: installing e2fsprogs updates from Trixie point release * 13:35 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] * 13:33 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1120.eqiad.wmnet with OS trixie * 13:33 topranks: reset cr3-eqsin configuration so traffic uses it again after upgrade * 13:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1088.eqiad.wmnet with OS trixie * 13:32 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie * 13:32 cdobbins@cumin1003: conftool action : set/pooled=no; selector: name=dns7002.* * 13:29 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2007.codfw.wmnet with OS trixie * 13:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1004.eqiad.wmnet with reason: host reimage * 13:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2006.codfw.wmnet with OS trixie * 13:18 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1036.eqiad.wmnet with OS trixie * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aux-k8s-etcd1004.eqiad.wmnet with reason: host reimage * 13:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 13:16 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 13:15 jayme: Istio is being upgraded from 1.24.2 to 1.29.4 on wikikube staging eqiad and codfw - [[phab:T427401|T427401]] * 13:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1087.eqiad.wmnet with OS trixie * 13:14 topranks: reboot cr3-eqsin to install new JunOS and set PIC 0/0/0 to 100G * 13:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1088.eqiad.wmnet with reason: host reimage * 13:13 jmm@dns1004: END - running authdns-update * 13:12 jmm@dns1004: START - running authdns-update * 13:09 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1088.eqiad.wmnet with reason: host reimage * 13:07 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1082.eqiad.wmnet with OS trixie * 13:07 jmm@dns1004: END - running authdns-update * 13:06 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1004.eqiad.wmnet with OS bookworm * 13:05 jmm@dns1004: START - running authdns-update * 13:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2006.codfw.wmnet with reason: host reimage * 12:58 topranks: load updated JunOS on cr3-eqsin [[phab:T429386|T429386]] * 12:58 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1036.eqiad.wmnet with reason: host reimage * 12:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2001.codfw.wmnet * 12:57 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2006.codfw.wmnet with reason: host reimage * 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr1-codfw,cr[2-3]-eqsin,cr3-eqsin IPv6,cr3-eqsin.mgmt with reason: upgrade JunOS cr3-eqsin * 12:56 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lvs[5004-5006].eqsin.wmnet with reason: upgrade JunOS cr3-eqsin * 12:55 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 12:55 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 12:53 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1087.eqiad.wmnet with reason: host reimage * 12:53 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 12:52 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1088.eqiad.wmnet with OS trixie * 12:52 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1002.eqiad.wmnet * 12:52 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:51 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2001.codfw.wmnet * 12:49 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1036.eqiad.wmnet with reason: host reimage * 12:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 12:48 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1087.eqiad.wmnet with reason: host reimage * 12:44 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1082.eqiad.wmnet with reason: host reimage * 12:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1002.eqiad.wmnet * 12:42 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 12:42 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 12:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 12:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 12:39 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:39 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: move dumps-nfs IP to the shared one - filippo@cumin1003" * 12:39 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: move dumps-nfs IP to the shared one - filippo@cumin1003" * 12:39 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2006.codfw.wmnet with OS trixie * 12:38 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1082.eqiad.wmnet with reason: host reimage * 12:36 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:33 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:32 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1036 * 12:32 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1036 * 12:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2005.codfw.wmnet with OS trixie * 12:32 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1087.eqiad.wmnet with OS trixie * 12:30 jmm@dns1004: END - running authdns-update * 12:29 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1036 * 12:29 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1036.eqiad.wmnet 21.32.64.10.in-addr.arpa 1.2.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:29 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1036.eqiad.wmnet 21.32.64.10.in-addr.arpa 1.2.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:29 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:29 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1036 - blake@cumin1003" * 12:29 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1036 - blake@cumin1003" * 12:28 jmm@dns1004: START - running authdns-update * 12:26 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:26 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:23 blake@cumin1003: START - Cookbook sre.dns.netbox * 12:23 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1036 * 12:23 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1036.eqiad.wmnet with OS trixie * 12:22 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1036.eqiad.wmnet * 12:22 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1082.eqiad.wmnet with OS trixie * 12:22 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1036.eqiad.wmnet * 12:22 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1036.eqiad.wmnet * 12:21 marostegui: Restart mariadb@s7 on db1155 to pick up new filters - [[phab:T431124|T431124]] * 12:21 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 21 hosts with reason: restarting for replication filter * 12:20 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:19 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:14 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2005.codfw.wmnet with reason: host reimage * 12:14 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:08 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:07 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:07 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2005.codfw.wmnet with reason: host reimage * 12:06 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:06 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:06 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:05 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:05 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:04 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-master-eqiad * 12:04 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl1002.eqiad.wmnet * 12:04 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl1002.eqiad.wmnet * 12:04 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:04 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:03 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:03 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:03 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:03 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 11:59 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl1002.eqiad.wmnet * 11:59 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl1002.eqiad.wmnet * 11:59 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl1001.eqiad.wmnet * 11:59 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl1001.eqiad.wmnet * 11:56 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl1001.eqiad.wmnet * 11:56 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl1001.eqiad.wmnet * 11:56 klausman@cumin2002: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-master-eqiad * 11:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2005.codfw.wmnet with OS trixie * 11:49 blake@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on wikikube-worker1160.eqiad.wmnet with reason: Verifying matchers for silence * 11:42 topranks: cr3-eqsin, begin traffic drain to reset PIC and upgrade JunOS [[phab:T429386|T429386]] * 11:41 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs[5004-5006].eqsin.wmnet with reason: upgrade JunOS cr3-eqsin * 11:39 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr1-codfw,cr[2-3]-eqsin,cr3-eqsin IPv6,cr3-eqsin.mgmt with reason: upgrade JunOS cr3-eqsin * 11:36 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=thanos-fe2004.codfw.wmnet * 11:35 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1086.eqiad.wmnet with OS trixie * 11:35 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=thanos-fe2004.codfw.wmnet * 11:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1085.eqiad.wmnet with OS trixie * 11:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1086.eqiad.wmnet with reason: host reimage * 11:10 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1085.eqiad.wmnet with reason: host reimage * 11:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2004.codfw.wmnet with OS trixie * 11:03 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1086.eqiad.wmnet with reason: host reimage * 11:02 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1085.eqiad.wmnet with reason: host reimage * 10:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2004.codfw.wmnet with reason: host reimage * 10:48 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:46 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1086.eqiad.wmnet with OS trixie * 10:46 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1085.eqiad.wmnet with OS trixie * 10:44 cgoubert@deploy1003: Finished deploy [restbase/deploy@2fc37d4]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] (duration: 16m 44s) * 10:43 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2004.codfw.wmnet with reason: host reimage * 10:35 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:27 cgoubert@deploy1003: Started deploy [restbase/deploy@2fc37d4]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] * 10:27 cgoubert@deploy1003: Finished deploy [restbase/deploy@8a25036]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] (duration: 00m 45s) * 10:26 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1117.eqiad.wmnet with OS trixie * 10:26 cgoubert@deploy1003: Started deploy [restbase/deploy@8a25036]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] * 10:26 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host thanos-fe2004 * 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host thanos-fe2004 * 10:22 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1116.eqiad.wmnet with OS trixie * 10:21 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host thanos-fe2004 * 10:21 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) thanos-fe2004.codfw.wmnet 157.32.192.10.in-addr.arpa 7.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:20 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache thanos-fe2004.codfw.wmnet 157.32.192.10.in-addr.arpa 7.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:20 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:20 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host thanos-fe2004 - mvernon@cumin2003" * 10:20 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host thanos-fe2004 - mvernon@cumin2003" * 10:15 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2252: Repooling after reboot * 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:15 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 10:15 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2252: Repooling after reboot * 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1153.eqiad.wmnet * 10:14 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1153.eqiad.wmnet * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 10:14 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 10:12 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 10:12 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host thanos-fe2004 * 10:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2004.codfw.wmnet with OS trixie * 10:07 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1117.eqiad.wmnet with reason: host reimage * 10:03 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1116.eqiad.wmnet with reason: host reimage * 09:58 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:58 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1117.eqiad.wmnet with reason: host reimage * 09:57 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1116.eqiad.wmnet with reason: host reimage * 09:49 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 41 days, 15:00:00 on db2252.codfw.wmnet with reason: Security updates * 09:45 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1117.eqiad.wmnet with OS trixie * 09:45 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1116.eqiad.wmnet with OS trixie * 09:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1153: Security updates * 09:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:28 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:28 root@cumin1003: START - Cookbook sre.mysql.depool depool db1153: Security updates * 09:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1016: Security updates * 09:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:21 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:21 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1016: Security updates * 09:14 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:14 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 08:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1016: Security updates * 08:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:56 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:56 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1016: Security updates * 08:50 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:50 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:45 filippo@dns1006: END - running authdns-update * 08:43 filippo@dns1006: START - running authdns-update * 08:42 godog: switch dumps-nfs address to be shared with rsync/http - [[phab:T411248|T411248]] * 08:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1016: Security updates * 08:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:40 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:40 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1016: Security updates * 08:29 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host cirrussearch1111.eqiad.wmnet * 08:29 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:27 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:27 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:25 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1015: Security updates * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:09 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:09 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1015: Security updates * 07:42 Msz2001: Deployed private patch for Suggested Ivestigations * 07:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1015: Security updates * 07:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:41 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:41 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1015: Security updates * 07:40 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 07:11 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fingerprint warnings - oblivian@cumin1003" * 07:11 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fingerprint warnings - oblivian@cumin1003 * 07:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1024: Security updates * 07:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:11 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:11 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1024: Security updates * 07:10 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fingerprint warnings - oblivian@cumin1003 * 07:10 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fingerprint warnings - oblivian@cumin1003" * 06:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host cirrussearch1111.eqiad.wmnet * 06:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 06:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1024: Security updates * 06:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 06:48 root@cumin1003: START - Cookbook sre.mysql.parsercache * 06:48 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1024: Security updates * 06:42 moritzm: install nginx security updates * 06:31 root@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool pc1024: Security updates * 06:21 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1024: Security updates * 06:19 moritzm: installing php8.2 security updates * 06:15 moritzm: installing php8.4 security updates * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.7 (duration: 02m 38s) * 03:40 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] (duration: 37m 04s) * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 51s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-06 == * 23:30 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] (duration: 09m 39s) * 23:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1078.eqiad.wmnet with OS trixie * 23:26 jdlrobson@deploy1003: jdlrobson, bwang: Continuing with deployment * 23:22 jdlrobson@deploy1003: jdlrobson, bwang: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug) * 23:21 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] * 23:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1078.eqiad.wmnet with reason: host reimage * 23:06 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1078.eqiad.wmnet with reason: host reimage * 22:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1078.eqiad.wmnet with OS trixie * 22:29 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on cirrussearch1114.eqiad.wmnet with reason: reimage on hold until restore completes * 22:22 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on cirrussearch[1079,1115].eqiad.wmnet with reason: reimage on hold until restore completes * 21:18 maryum: Deployed security fix for [[phab:T428006|T428006]] * 20:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1079.eqiad.wmnet with OS trixie * 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1077.eqiad.wmnet with OS trixie * 20:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1115.eqiad.wmnet with OS trixie * 20:25 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1079.eqiad.wmnet with reason: host reimage * 20:21 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1079.eqiad.wmnet with reason: host reimage * 20:15 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] (duration: 08m 14s) * 20:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1077.eqiad.wmnet with reason: host reimage * 20:10 krinkle@deploy1003: krinkle, pushpaktiwari: Continuing with deployment * 20:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1115.eqiad.wmnet with reason: host reimage * 20:08 krinkle@deploy1003: krinkle, pushpaktiwari: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1077.eqiad.wmnet with reason: host reimage * 20:06 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] * 20:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1079.eqiad.wmnet with OS trixie * 20:04 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1115.eqiad.wmnet with reason: host reimage * 19:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1077.eqiad.wmnet with OS trixie * 19:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1115.eqiad.wmnet with OS trixie * 19:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 19:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 18:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1114.eqiad.wmnet with OS trixie * 18:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1114.eqiad.wmnet with reason: host reimage * 18:35 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1114.eqiad.wmnet with reason: host reimage * 18:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1112.eqiad.wmnet with OS trixie * 18:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1114.eqiad.wmnet with OS trixie * 18:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1072.eqiad.wmnet with OS trixie * 18:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1112.eqiad.wmnet with reason: host reimage * 18:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1112.eqiad.wmnet with reason: host reimage * 17:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1072.eqiad.wmnet with reason: host reimage * 17:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1112.eqiad.wmnet with OS trixie * 17:55 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1072.eqiad.wmnet with reason: host reimage * 17:39 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1072.eqiad.wmnet with OS trixie * 17:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1071.eqiad.wmnet with OS trixie * 17:18 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1070.eqiad.wmnet with OS trixie * 17:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1084.eqiad.wmnet with OS trixie * 16:54 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1071.eqiad.wmnet with reason: host reimage * 16:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1084.eqiad.wmnet with reason: host reimage * 16:51 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1070.eqiad.wmnet with reason: host reimage * 16:49 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1084.eqiad.wmnet with reason: host reimage * 16:38 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1071.eqiad.wmnet with OS trixie * 16:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1096.eqiad.wmnet with OS trixie * 16:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1070.eqiad.wmnet with OS trixie * 16:33 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1084.eqiad.wmnet with OS trixie * 16:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1089.eqiad.wmnet with OS trixie * 16:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1103.eqiad.wmnet with OS trixie * 16:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1096.eqiad.wmnet with reason: host reimage * 16:14 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1096.eqiad.wmnet with reason: host reimage * 16:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1089.eqiad.wmnet with reason: host reimage * 16:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1103.eqiad.wmnet with reason: host reimage * 16:02 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1003.eqiad.wmnet with OS bookworm * 16:01 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1089.eqiad.wmnet with reason: host reimage * 16:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1103.eqiad.wmnet with reason: host reimage * 15:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1096.eqiad.wmnet with OS trixie * 15:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1080.eqiad.wmnet with OS trixie * 15:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1089.eqiad.wmnet with OS trixie * 15:45 dancy@deploy1003: Installation of scap version "4.272.0" completed for 158 hosts * 15:43 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1103.eqiad.wmnet with OS trixie * 15:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1113.eqiad.wmnet with OS trixie * 15:41 dancy@deploy1003: Installing scap version "4.272.0" for 158 host(s) * 15:40 klausman@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 15:39 klausman@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 15:38 klausman@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 15:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1069.eqiad.wmnet with OS trixie * 15:37 klausman@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 15:36 klausman@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 15:34 klausman@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 15:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1080.eqiad.wmnet with reason: host reimage * 15:27 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1080.eqiad.wmnet with reason: host reimage * 15:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1113.eqiad.wmnet with reason: host reimage * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1069.eqiad.wmnet with reason: host reimage * 15:18 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1113.eqiad.wmnet with reason: host reimage * 15:16 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1069.eqiad.wmnet with reason: host reimage * 15:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:11 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1080.eqiad.wmnet with OS trixie * 15:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1113.eqiad.wmnet with OS trixie * 15:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1003.eqiad.wmnet with reason: host reimage * 14:47 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1003.eqiad.wmnet with OS bookworm * 14:33 elukey: rolled out spicerack on all cumin nodes - [[phab:T429699|T429699]] * 14:32 elukey: upgrade all bookworm hosts to pywmflib 3.1 - [[phab:T430552|T430552]] * 14:14 marostegui: Setup x4 eqiad topology [[phab:T404715|T404715]] * 14:13 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 14:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2230.codfw.wmnet * 14:07 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2230.codfw.wmnet * 13:59 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[2001-2002].codfw.wmnet * 13:51 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 13:45 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.major-upgrade (exit_code=97) * 13:45 cwilliams@cumin1003: dbmaint on s4@codfw [[phab:T429893|T429893]] * 13:45 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 13:42 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-master-codfw * 13:42 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl2002.codfw.wmnet * 13:42 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl2002.codfw.wmnet * 13:38 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl2002.codfw.wmnet * 13:38 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl2002.codfw.wmnet * 13:38 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl2001.codfw.wmnet * 13:38 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl2001.codfw.wmnet * 13:35 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl2001.codfw.wmnet * 13:35 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl2001.codfw.wmnet * 13:35 klausman@cumin2002: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-master-codfw * 12:30 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] (duration: 25m 11s) * 12:24 krinkle@deploy1003: krinkle: Continuing with deployment * 12:10 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2048.codfw.wmnet * 12:09 krinkle@deploy1003: krinkle: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:08 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2048.codfw.wmnet * 12:05 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] * 11:57 moritzm: installing curl security updates * 11:49 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:31 moritzm: installing nano security updates * 11:07 moritzm: failover Ganeti master in codfw to ganeti2032 [[phab:T430909|T430909]] * 11:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:04 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest1005.eqiad.wmnet with OS trixie * 11:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:50 jmm@dns1004: END - running authdns-update * 10:47 jmm@dns1004: START - running authdns-update * 10:47 jmm@dns1004: START - running authdns-update * 10:46 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:44 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest1005.eqiad.wmnet with reason: host reimage * 10:38 elukey@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest1005.eqiad.wmnet with reason: host reimage * 10:31 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:31 marostegui: Setup x4 codfw topology [[phab:T404715|T404715]] * 10:31 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 10:24 elukey: spicerack 13.0.0 deployed on cumin2002 * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 10:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 10:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 10:21 elukey@cumin2002: START - Cookbook sre.hosts.reimage for host sretest1005.eqiad.wmnet with OS trixie * 10:20 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:19 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:17 elukey: uploaded spicerack_13.0.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia * 09:54 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:52 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:20 elukey: upgrade all bullseye hosts to pywmflib 3.1 - [[phab:T430552|T430552]] * 09:10 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1015.eqiad.wmnet,service=s4 * 09:10 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1015.eqiad.wmnet,service=s6 * 09:07 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 08:58 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:56 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 08:56 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 08:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 08:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 08:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin2002.codfw.wmnet * 08:06 godog: remove cloudvirt1046, cloudvirt1062, cloudvirt1074, cloudvirt1075 from maintenance aggregate and put them in network-ovs - [[phab:T424802|T424802]] * 08:00 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin2002.codfw.wmnet * 07:58 hashar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] (duration: 32m 53s) * 07:57 fabfur: repooled cp4038 * 07:57 fabfur@cumin1003: conftool action : set/pooled=yes; selector: name=cp4038.* * 07:53 moritzm: installing pyjwt security updates * 07:47 moritzm: installing openjpeg2 security updates * 07:45 hashar@deploy1003: vadymts1, hashar: Continuing with deployment * 07:43 hashar@deploy1003: vadymts1, hashar: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:38 moritzm: installing python-urllib3 security updates * 07:37 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 07:30 fabfur: depooled cp4038 to investigate on possible maxmind failure * 07:30 fabfur@cumin1003: conftool action : set/pooled=no; selector: name=cp4038.* * 07:30 fabfur@cumin1003: conftool action : set/pooled=yes; selector: name=cp4038.* * 07:29 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 07:25 hashar@deploy1003: Started scap sync-world: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] * 06:13 moritzm: installing Linux 6.12.95 on trixie hosts * 05:20 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s6 * 05:20 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s4 * 05:19 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1015.eqiad.wmnet with reason: cloning * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 08s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-05 == * 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 01m 08s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-04 == * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 58s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-03 == * 17:08 topranks: revert protocol preference changes on cr3-ulsfo after upgrade * 16:53 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on cr2-eqord with reason: upgrade JunOS cr3-ulsfo * 16:53 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on cr4-ulsfo with reason: upgrade JunOS cr3-ulsfo * 16:48 topranks: reboot cr3-ulsfo to upgrade JunOS and reset linecard [[phab:T424839|T424839]] * 15:52 topranks: adjust outbound BGP policies on cr3-ulsfo to drain router of traffic [[phab:T424839|T424839]] * 15:45 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on lvs[4008-4010].ulsfo.wmnet with reason: upgrade JunOS cr3-ulsfo * 15:44 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on asw1-[22-23]-ulsfo,cr3-ulsfo,cr3-ulsfo IPv6,cr3-ulsfo.mgmt with reason: upgrade JunOS cr3-ulsfo * 15:36 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 15:35 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 15:35 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 14:40 cmooney@dns3003: END - running authdns-update * 14:26 cmooney@dns3003: START - running authdns-update * 14:26 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:26 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to ulsfo - cmooney@cumin1003" * 14:19 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to ulsfo - cmooney@cumin1003" * 14:16 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:38 sukhe@dns1004: END - running authdns-update * 13:35 sukhe@dns1004: START - running authdns-update * 13:26 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 13:26 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 13:26 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet * 13:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 13:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 13:16 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host sretest1005.eqiad.wmnet * 13:16 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 13:16 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 13:15 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 13:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:14 moritzm: imported samplicator 1.3.8rc1-1+deb13u1 to trixie-wikimedia/main [[phab:T337208|T337208]] * 13:13 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:07 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 13:07 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 13:02 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:02 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:58 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:57 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:57 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:53 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet * 12:50 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 12:47 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:41 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:40 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:39 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:32 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet * 12:26 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet * 12:23 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2005.wikimedia.org * 12:19 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2005.wikimedia.org * 12:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 12:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup[2004-2007].codfw.wmnet * 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[2004-2007].codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin2003" * 12:15 jynus@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[2004-2007].codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin2003" * 12:09 jynus@cumin2003: START - Cookbook sre.dns.netbox * 11:58 jynus@cumin2003: START - Cookbook sre.hosts.decommission for hosts backup[2004-2007].codfw.wmnet * 10:40 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup[1004-1007].eqiad.wmnet * 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[1004-1007].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 10:01 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[1004-1007].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 09:52 jynus@cumin1003: START - Cookbook sre.dns.netbox * 09:39 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:36 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup[1004-1007].eqiad.wmnet * 09:36 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:25 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 09:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 09:16 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 09:05 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:04 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:00 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:59 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:57 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:55 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:50 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 08:50 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 08:49 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 08:49 atsukoito: depooling cirrussearch in codfw because of regression after upgrade [[phab:T431091|T431091]] * 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts mirror1001.wikimedia.org * 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: mirror1001.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 08:29 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: mirror1001.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 08:18 jmm@cumin2003: START - Cookbook sre.dns.netbox * 08:11 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts mirror1001.wikimedia.org * 06:15 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 18s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-02 == * 22:55 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host contint1003.wikimedia.org with OS trixie * 22:29 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on contint1003.wikimedia.org with reason: host reimage * 22:23 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on contint1003.wikimedia.org with reason: host reimage * 22:05 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host contint1003.wikimedia.org with OS trixie * 22:03 mutante: contint1003 (zuul.wikimedia.org) - reimaging because of [[phab:T430510|T430510]]#12067628 [[phab:T418521|T418521]] * 22:03 dzahn@cumin2002: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on zuul.wikimedia.org with reason: reimage * 21:39 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 18s) * 21:39 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 21:20 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1003.eqiad.wmnet, repooling source-only afterwards * 21:19 sbassett: Deployed security fix for [[phab:T428829|T428829]] * 20:58 cmooney@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Release v0.11.2 update for new Aerleon - cmooney@cumin1003 * 20:55 cmooney@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Release v0.11.2 update for new Aerleon - cmooney@cumin1003 * 20:40 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] (duration: 12m 35s) * 20:36 arlolra@deploy1003: cscott, arlolra: Continuing with deployment * 20:35 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 20s) * 20:35 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 20:33 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host contint2003.wikimedia.org with OS trixie * 20:31 arlolra@deploy1003: cscott, arlolra: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Cha * 20:28 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] * 20:17 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] (duration: 08m 13s) * 20:14 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on contint2003.wikimedia.org with reason: host reimage * 20:13 sbassett@deploy1003: sbassett: Continuing with deployment * 20:11 sbassett@deploy1003: sbassett: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:09 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] * 20:08 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 20:08 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 20:08 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on contint2003.wikimedia.org with reason: host reimage * 20:05 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1003.eqiad.wmnet, repooling source-only afterwards * 19:49 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host contint2003.wikimedia.org with OS trixie * 19:48 mutante: contint2003 - reimaging because of [[phab:T430510|T430510]]#12067628 [[phab:T418521|T418521]] * 18:39 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 18:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 18:13 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2002.codfw.wmnet -> wcqs2003.codfw.wmnet, repooling source-only afterwards * 17:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1003.eqiad.wmnet with OS bookworm * 17:52 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1005.eqiad.wmnet * 17:52 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1005.eqiad.wmnet * 17:51 jasmine@cumin2002: conftool action : set/pooled=yes:weight=10; selector: name=wikikube-ctrl1005.eqiad.wmnet * 17:48 jasmine_: homer "cr*eqiad*" commit "Added new stacked control plane wikikube-ctrl1005" * 17:44 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply * 17:44 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply * 17:31 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] (duration: 09m 33s) * 17:26 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 17:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1003.eqiad.wmnet with reason: host reimage * 17:23 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:21 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] * 17:18 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1003.eqiad.wmnet with reason: host reimage * 17:16 rscout@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply * 17:16 rscout@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply * 17:16 rscout@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply * 17:15 rscout@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply * 17:12 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on wcqs[2002-2003].codfw.wmnet,wcqs1002.eqiad.wmnet with reason: reimaging hosts * 17:08 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 17:08 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 17:08 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 17:07 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 17:05 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 17:05 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 17:03 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "running to make sure all updates are synced - cmooney@cumin1003" * 17:03 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "running to make sure all updates are synced - cmooney@cumin1003" * 17:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs1003 * 17:00 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs1003 * 17:00 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 17:00 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1003.eqiad.wmnet with OS bookworm * 16:58 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Re-running - btullis@cumin1003" * 16:58 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Re-running - btullis@cumin1003" * 16:58 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2002.codfw.wmnet -> wcqs2003.codfw.wmnet, repooling source-only afterwards * 16:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-master1004.eqiad.wmnet with OS bookworm * 16:58 btullis@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 16:57 tappof: bump space for prometheus k8s-aux in eqiad * 16:55 cmooney@dns3003: END - running authdns-update * 16:55 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:55 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to eqsin - cmooney@cumin1003" * 16:55 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to eqsin - cmooney@cumin1003" * 16:53 cmooney@dns3003: START - running authdns-update * 16:52 ryankemper: [ml-serve-eqiad] Cleared out 1302 failed (Evicted) pods: `kubectl -n llm delete pods --field-selector=status.phase=Failed`, freeing calico-kube-controllers from OOM crashloop (evictions were caused by disk pressure) * 16:49 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 16:46 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:39 rzl@dns1004: END - running authdns-update * 16:37 rzl@dns1004: START - running authdns-update * 16:36 rzl@dns1004: START - running authdns-update * 16:35 rzl@deploy1003: Finished scap sync-world: [[phab:T416623|T416623]] (duration: 10m 19s) * 16:34 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 16:33 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-master1004.eqiad.wmnet with reason: host reimage * 16:30 rzl@deploy1003: rzl: Continuing with deployment * 16:28 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-master1004.eqiad.wmnet with reason: host reimage * 16:26 rzl@deploy1003: rzl: [[phab:T416623|T416623]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:25 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 16:25 rzl@deploy1003: Started scap sync-world: [[phab:T416623|T416623]] * 16:25 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 16:24 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 16:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: sync * 16:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: sync * 16:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-master1004.eqiad.wmnet with OS bookworm * 16:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-master1003.eqiad.wmnet with OS bookworm * 16:11 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 16:11 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 16:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Security updates * 16:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 16:08 root@cumin1003: START - Cookbook sre.mysql.parsercache * 16:08 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Security updates * 15:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-master1003.eqiad.wmnet with reason: host reimage * 15:54 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:54 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:54 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:54 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-master1003.eqiad.wmnet with reason: host reimage * 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Security updates * 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:45 root@cumin1003: START - Cookbook sre.mysql.parsercache * 15:45 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Security updates * 15:42 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-master1003.eqiad.wmnet with OS bookworm * 15:24 moritzm: installing busybox updates from bookworm point release * 15:20 moritzm: installing busybox updates from trixie point release * 15:15 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1021: Security updates * 15:15 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:15 root@cumin1003: START - Cookbook sre.mysql.parsercache * 15:15 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1021: Security updates * 15:13 moritzm: installing giflib security updates * 15:08 moritzm: installing Tomcat security updates * 14:57 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 14:56 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 14:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:53 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Unblock taavi - oblivian@cumin1003" * 14:53 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Unblock taavi - oblivian@cumin1003 * 14:53 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1021: Security updates * 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:53 root@cumin1003: START - Cookbook sre.mysql.parsercache * 14:53 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1021: Security updates * 14:53 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Unblock taavi - oblivian@cumin1003 * 14:52 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Unblock taavi - oblivian@cumin1003" * 14:46 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94711 and previous config saved to /var/cache/conftool/dbconfig/20260702-144644-fceratto.json * 14:36 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205', diff saved to https://phabricator.wikimedia.org/P94709 and previous config saved to /var/cache/conftool/dbconfig/20260702-143636-fceratto.json * 14:32 moritzm: installing libdbi-perl security updates * 14:26 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205', diff saved to https://phabricator.wikimedia.org/P94708 and previous config saved to /var/cache/conftool/dbconfig/20260702-142628-fceratto.json * 14:16 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94707 and previous config saved to /var/cache/conftool/dbconfig/20260702-141621-fceratto.json * 14:12 moritzm: installing rsync security updates * 14:11 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox) * 14:10 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94706 and previous config saved to /var/cache/conftool/dbconfig/20260702-140959-fceratto.json * 14:09 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2205.codfw.wmnet with reason: Maintenance * 14:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2205: Repooling after switchover * 14:07 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-test-master1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 14:06 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 14:06 Tran: Deployed patch for [[phab:T427287|T427287]] * 14:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:59 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2205: Repooling after switchover * 13:59 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2205: Repooling after switchover * 13:59 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:55 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2205: Repooling after switchover * 13:55 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2205 [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94704 and previous config saved to /var/cache/conftool/dbconfig/20260702-135505-fceratto.json * 13:54 moritzm: installing sed security updates * 13:53 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:52 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2209 to s3 primary [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94703 and previous config saved to /var/cache/conftool/dbconfig/20260702-135235-fceratto.json * 13:52 federico3: Starting s3 codfw failover from db2205 to db2209 - [[phab:T430912|T430912]] * 13:51 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:51 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 13:48 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:47 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2209 with weight 0 [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94702 and previous config saved to /var/cache/conftool/dbconfig/20260702-134719-fceratto.json * 13:47 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Primary switchover s3 [[phab:T430912|T430912]] * 13:44 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:44 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:44 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:40 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 13:38 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 13:37 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 13:36 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 13:36 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:34 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 13:30 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:29 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:29 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:27 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:26 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:25 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling restart_daemons on A:wikidough * 13:23 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 13:22 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 13:17 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 13:17 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns1004.wikimedia.org * 13:12 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:11 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart (exit_code=97) rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough * 13:11 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=97) rolling restart_daemons on A:wikidough * 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough * 13:09 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] (duration: 07m 20s) * 13:05 aude@deploy1003: jdrewniak, aude: Continuing with deployment * 13:04 aude@deploy1003: jdrewniak, aude: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:02 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] * 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts wdqs-categories1001.eqiad.wmnet * 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: wdqs-categories1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 12:10 jmm@dns1004: END - running authdns-update * 12:07 jmm@dns1004: START - running authdns-update * 11:51 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: wdqs-categories1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 11:44 btullis@cumin1003: START - Cookbook sre.dns.netbox * 11:42 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 11:42 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 11:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet * 11:39 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts wdqs-categories1001.eqiad.wmnet * 11:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet * 11:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet * 11:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet * 11:29 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 11:29 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 10:57 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2214: Repooling * 10:49 jmm@dns1004: END - running authdns-update * 10:47 jmm@dns1004: START - running authdns-update * 10:31 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94698 and previous config saved to /var/cache/conftool/dbconfig/20260702-103146-fceratto.json * 10:21 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213', diff saved to https://phabricator.wikimedia.org/P94696 and previous config saved to /var/cache/conftool/dbconfig/20260702-102137-fceratto.json * 10:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:19 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb1017.eqiad.wmnet * 10:18 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 10:18 fceratto@cumin1003: Removing es1033 from zarcillo [[phab:T408772|T408772]] * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts es1033.eqiad.wmnet * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: es1033.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:14 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: es1033.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:13 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb1017.eqiad.wmnet * 10:12 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2214.codfw.wmnet * 10:12 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2214.codfw.wmnet * 10:12 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2214: Repooling * 10:11 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213', diff saved to https://phabricator.wikimedia.org/P94693 and previous config saved to /var/cache/conftool/dbconfig/20260702-101130-fceratto.json * 10:10 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:10 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:03 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts es1033.eqiad.wmnet * 10:03 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 10:01 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94691 and previous config saved to /var/cache/conftool/dbconfig/20260702-100122-fceratto.json * 09:55 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94690 and previous config saved to /var/cache/conftool/dbconfig/20260702-095529-fceratto.json * 09:55 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2213.codfw.wmnet with reason: Maintenance * 09:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 09:53 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2213: Repooling after switchover * 09:51 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover * 09:44 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2213: Repooling after switchover * 09:39 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover * 09:39 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2213 [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94688 and previous config saved to /var/cache/conftool/dbconfig/20260702-093859-fceratto.json * 09:36 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2192 to s5 primary [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94687 and previous config saved to /var/cache/conftool/dbconfig/20260702-093650-fceratto.json * 09:36 federico3: Starting s5 codfw failover from db2213 to db2192 - [[phab:T430923|T430923]] * 09:30 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94686 and previous config saved to /var/cache/conftool/dbconfig/20260702-093004-fceratto.json * 09:24 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2192 with weight 0 [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94685 and previous config saved to /var/cache/conftool/dbconfig/20260702-092455-fceratto.json * 09:24 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 23 hosts with reason: Primary switchover s5 [[phab:T430923|T430923]] * 09:19 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220', diff saved to https://phabricator.wikimedia.org/P94684 and previous config saved to /var/cache/conftool/dbconfig/20260702-091957-fceratto.json * 09:16 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] (duration: 06m 57s) * 09:13 moritzm: installing libgcrypt20 security updates * 09:12 kharlan@deploy1003: kharlan: Continuing with deployment * 09:11 kharlan@deploy1003: kharlan: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:09 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220', diff saved to https://phabricator.wikimedia.org/P94683 and previous config saved to /var/cache/conftool/dbconfig/20260702-090950-fceratto.json * 09:09 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] * 09:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 09:01 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] (duration: 07m 07s) * 08:59 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94682 and previous config saved to /var/cache/conftool/dbconfig/20260702-085942-fceratto.json * 08:57 kharlan@deploy1003: kharlan: Continuing with deployment * 08:56 kharlan@deploy1003: kharlan: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:54 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] * 08:52 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:52 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:52 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94681 and previous config saved to /var/cache/conftool/dbconfig/20260702-085237-fceratto.json * 08:52 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2220.codfw.wmnet with reason: Maintenance * 08:43 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:40 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 08:25 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] (duration: 11m 44s) * 08:21 cscott@deploy1003: cscott: Continuing with deployment * 08:16 cscott@deploy1003: cscott: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:14 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] * 08:08 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 08:08 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1244: Migration of db1244.eqiad.wmnet completed * 08:02 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:02 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:01 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] (duration: 18m 58s) * 08:01 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:59 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 07:59 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:59 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:59 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2006.wikimedia.org * 07:58 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:57 cscott@deploy1003: cscott: Continuing with deployment * 07:56 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:56 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:56 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:55 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:55 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:55 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:54 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2006.wikimedia.org * 07:54 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:54 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 07:44 cscott@deploy1003: cscott: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:44 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2005.wikimedia.org * 07:44 moritzm: installing node-lodash security updates * 07:42 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] * 07:39 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2005.wikimedia.org * 07:30 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] (duration: 07m 28s) * 07:26 cscott@deploy1003: ssastry, cscott: Continuing with deployment * 07:25 cscott@deploy1003: ssastry, cscott: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:23 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1244: Migration of db1244.eqiad.wmnet completed * 07:22 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] * 07:16 wmde-fisch@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] (duration: 06m 55s) * 07:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1244.eqiad.wmnet with OS trixie * 07:11 wmde-fisch@deploy1003: wmde-fisch: Continuing with deployment * 07:11 wmde-fisch@deploy1003: wmde-fisch: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:09 wmde-fisch@deploy1003: Started scap sync-world: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] * 06:54 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1244.eqiad.wmnet with reason: host reimage * 06:50 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1244.eqiad.wmnet with reason: host reimage * 06:38 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1250.eqiad.wmnet with OS trixie * 06:34 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db1244.eqiad.wmnet with OS trixie * 06:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1244: Upgrading db1244.eqiad.wmnet * 06:25 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1244: Upgrading db1244.eqiad.wmnet * 06:25 cwilliams@cumin1003: dbmaint on s4@eqiad [[phab:T429893|T429893]] * 06:25 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 06:15 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1250.eqiad.wmnet with reason: host reimage * 06:14 cwilliams@dns1006: END - running authdns-update * 06:12 cwilliams@dns1006: START - running authdns-update * 06:11 cwilliams@dns1006: END - running authdns-update * 06:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db1244 [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94676 and previous config saved to /var/cache/conftool/dbconfig/20260702-061059-cwilliams.json * 06:09 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1250.eqiad.wmnet with reason: host reimage * 06:09 cwilliams@dns1006: START - running authdns-update * 06:08 aokoth@cumin1003: END (PASS) - Cookbook sre.vrts.upgrade (exit_code=0) on VRTS host vrts1003.eqiad.wmnet * 06:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db1160 to s4 primary and set section read-write [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94675 and previous config saved to /var/cache/conftool/dbconfig/20260702-060746-cwilliams.json * 06:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Set s4 eqiad as read-only for maintenance - [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94674 and previous config saved to /var/cache/conftool/dbconfig/20260702-060704-cwilliams.json * 06:06 cezmunsta: Starting s4 eqiad failover from db1244 to db1160 - [[phab:T430817|T430817]] * 06:04 aokoth@cumin1003: START - Cookbook sre.vrts.upgrade on VRTS host vrts1003.eqiad.wmnet * 05:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db1160 with weight 0 [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94673 and previous config saved to /var/cache/conftool/dbconfig/20260702-055927-cwilliams.json * 05:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 40 hosts with reason: Primary switchover s4 [[phab:T430817|T430817]] * 05:55 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1250.eqiad.wmnet with OS trixie * 05:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on db1250.eqiad.wmnet with reason: m3 master switchover [[phab:T430158|T430158]] * 05:39 marostegui: Failover m3 (phabricator) from db1250 to db1228 - [[phab:T430158|T430158]] * 05:32 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2234].codfw.wmnet,db[1217,1228,1250].eqiad.wmnet with reason: m3 master switchover [[phab:T430158|T430158]] * 04:45 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] (duration: 09m 08s) * 04:41 tstarling@deploy1003: tstarling, reedy: Continuing with deployment * 04:38 tstarling@deploy1003: tstarling, reedy: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 04:36 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 59s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:16 ryankemper: [[phab:T429844|T429844]] [opensearch] completed `cirrussearch2111` reimage; all codfw search clusters are green, all nodes now report `OpenSearch 2.19.5`, and the temporary chi voting exclusion has been removed * 00:57 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2111.codfw.wmnet with OS trixie * 00:29 ryankemper: [[phab:T429844|T429844]] [opensearch] depooled codfw search-omega/search-psi discovery records to match existing codfw search depool during OpenSearch 2.19 migration * 00:29 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2111.codfw.wmnet with reason: host reimage * 00:29 ryankemper@cumin2002: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 00:29 ryankemper@cumin2002: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 00:22 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2111.codfw.wmnet with reason: host reimage * 00:01 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2111.codfw.wmnet with OS trixie * 00:00 ryankemper: [[phab:T429844|T429844]] [opensearch] chi cluster recovered after stopping `opensearch_1@production-search-codfw` on `cirrussearch2111` == 2026-07-01 == * 23:59 ryankemper: [[phab:T429844|T429844]] [opensearch] stopped `opensearch_1@production-search-codfw` on `cirrussearch2111` after chi cluster-manager election churn following `voting_config_exclusions` POST; hoping this triggers a re-election * 23:52 cscott@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 23:51 cscott@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 23:51 cscott@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 23:50 cscott@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2003.codfw.wmnet with OS bookworm * 22:29 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 22:13 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 22:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2084.codfw.wmnet with OS trixie * 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2003.codfw.wmnet with reason: host reimage * 22:03 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 22:01 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2003.codfw.wmnet with reason: host reimage * 21:50 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 21:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2084.codfw.wmnet with reason: host reimage * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2003 * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2003 * 21:42 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2003 * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2003.codfw.wmnet 45.48.192.10.in-addr.arpa 5.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:42 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2003.codfw.wmnet 45.48.192.10.in-addr.arpa 5.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2003 - bking@cumin2003" * 21:42 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2003 - bking@cumin2003" * 21:36 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2084.codfw.wmnet with reason: host reimage * 21:35 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:34 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2003 * 21:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2003.codfw.wmnet with OS bookworm * 21:19 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2084.codfw.wmnet with OS trixie * 21:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2081.codfw.wmnet with OS trixie * 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2108.codfw.wmnet with OS trixie * 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2081.codfw.wmnet with reason: host reimage * 20:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2081.codfw.wmnet with reason: host reimage * 20:28 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2081.codfw.wmnet with OS trixie * 20:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2108.codfw.wmnet with reason: host reimage * 20:19 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2108.codfw.wmnet with reason: host reimage * 19:59 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2108.codfw.wmnet with OS trixie * 19:46 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2093.codfw.wmnet with OS trixie * 19:44 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 19:44 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jasmine@cumin2002" * 19:43 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jasmine@cumin2002" * 19:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2080.codfw.wmnet with OS trixie * 19:28 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 19:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2093.codfw.wmnet with reason: host reimage * 19:18 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 19:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2093.codfw.wmnet with reason: host reimage * 19:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2080.codfw.wmnet with reason: host reimage * 19:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2080.codfw.wmnet with reason: host reimage * 18:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2093.codfw.wmnet with OS trixie * 18:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2080.codfw.wmnet with OS trixie * 18:27 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 18:18 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] (duration: 09m 15s) * 18:13 jgiannelos@deploy1003: jgiannelos, neriah: Continuing with deployment * 18:11 jgiannelos@deploy1003: jgiannelos, neriah: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:09 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] * 17:40 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 16:58 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 30 hosts * 16:57 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for 30 hosts * 16:52 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2202.codfw.wmnet * 16:52 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2202.codfw.wmnet * 16:51 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt * 16:51 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt * 16:51 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lvs2012.codfw.wmnet * 16:51 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for lvs2012.codfw.wmnet * 16:49 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2076.codfw.wmnet with OS trixie * 16:49 brett: Start pybal on lvs2012 - [[phab:T429861|T429861]] * 16:49 pt1979@cumin1003: END (ERROR) - Cookbook sre.hosts.remove-downtime (exit_code=97) for 59 hosts * 16:48 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for 59 hosts * 16:42 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2061.codfw.wmnet with OS trixie * 16:30 dancy@deploy1003: Installation of scap version "4.271.0" completed for 2 hosts * 16:28 dancy@deploy1003: Installing scap version "4.271.0" for 2 host(s) * 16:23 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2076.codfw.wmnet with reason: host reimage * 16:19 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2061.codfw.wmnet with reason: host reimage * 16:18 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2076.codfw.wmnet with reason: host reimage * 16:18 jasmine@dns1004: END - running authdns-update * 16:16 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host restbase2039.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 16:16 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host restbase2039.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 16:16 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2061.codfw.wmnet with reason: host reimage * 16:15 jasmine@dns1004: START - running authdns-update * 16:14 jasmine@dns1004: END - running authdns-update * 16:12 jasmine@dns1004: START - running authdns-update * 16:07 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2202.codfw.wmnet with reason: maintenance * 16:06 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt with reason: Junos upograde * 16:00 papaul: ongoing maintenance on lsw1-b2-codfw * 16:00 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2076.codfw.wmnet with OS trixie * 15:59 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt * 15:59 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt * 15:57 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2061.codfw.wmnet with OS trixie * 15:55 pt1979@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2042,2046].codfw.wmnet * 15:55 pt1979@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2042,2046].codfw.wmnet * 15:51 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 15:51 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2220: Repooling after switchover * 15:50 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 15:50 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 15:48 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2092.codfw.wmnet with OS trixie * 15:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 15:40 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 15:38 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 15:37 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 15:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply * 15:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply * 15:32 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 15:32 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 15:30 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 15:29 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 15:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 15:25 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:22 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs2012.codfw.wmnet with reason: Rack B2 maintenance - [[phab:T429861|T429861]] * 15:21 brett: Stopping pybal on lvs2012 in preparation for codfw rack b2 maintenance - [[phab:T429861|T429861]] * 15:20 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2092.codfw.wmnet with reason: host reimage * 15:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:12 _joe_: restarted manually alertmanager-irc-relay * 15:12 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2092.codfw.wmnet with reason: host reimage * 15:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:12 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt with reason: Junos upograde * 15:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover * 15:07 pt1979@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2042,2046].codfw.wmnet * 15:06 pt1979@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2042,2046].codfw.wmnet * 15:02 papaul: ongoing maintenance on lsw1-a8-codfw * 14:31 topranks: POWERING DOWN CR1-EQIAD for line card installation [[phab:T426343|T426343]] * 14:31 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] (duration: 08m 57s) * 14:29 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:26 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 14:24 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:22 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] * 14:22 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover * 14:16 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:15 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover * 14:14 topranks: re-enable routing-engine graceful-failover on cr1-eqiad [[phab:T417873|T417873]] * 14:13 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:13 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2220: Repooling after switchover * 14:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:12 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:12 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:11 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:08 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] (duration: 10m 01s) * 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:07 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2220 [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94664 and previous config saved to /var/cache/conftool/dbconfig/20260701-140729-fceratto.json * 14:06 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:06 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:06 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:05 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2159 to s7 primary [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94663 and previous config saved to /var/cache/conftool/dbconfig/20260701-140503-fceratto.json * 14:04 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:04 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 14:04 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 14:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:04 dreamyjazz@deploy1003: anzx, dreamyjazz: Continuing with deployment * 14:04 federico3: Starting s7 codfw failover from db2220 to db2159 - [[phab:T430826|T430826]] * 14:03 jmm@dns1004: END - running authdns-update * 14:03 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:03 topranks: flipping cr1-eqiad active routing-enginer back to RE0 [[phab:T417873|T417873]] * 14:03 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudsw1-c8-eqiad,cloudsw1-d5-eqiad with reason: router upgrades eqiad * 14:01 jmm@dns1004: START - running authdns-update * 14:00 dreamyjazz@deploy1003: anzx, dreamyjazz: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:59 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2159 with weight 0 [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94662 and previous config saved to /var/cache/conftool/dbconfig/20260701-135906-fceratto.json * 13:58 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] * 13:57 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s7 [[phab:T430826|T430826]] * 13:56 topranks: reboot routing-enginer RE0 on cr1-eqiad [[phab:T417873|T417873]] * 13:48 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1006.wikimedia.org * 13:44 atsuko@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cirrussearch2092.codfw.wmnet with OS trixie * 13:43 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1006.wikimedia.org * 13:41 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2092.codfw.wmnet with OS trixie * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1005.wikimedia.org * 13:37 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1005.wikimedia.org * 13:37 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on pfw1-eqiad with reason: router upgrades eqiad * 13:35 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on lvs[1017-1020].eqiad.wmnet with reason: router upgrades eqiad * 13:34 caro@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] (duration: 07m 59s) * 13:30 caro@deploy1003: caro: Continuing with deployment * 13:28 caro@deploy1003: caro: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:27 topranks: route-engine failover cr1-eqiad * 13:26 caro@deploy1003: Started scap sync-world: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] * 13:15 topranks: rebooting routing-engine 1 on cr1-eqiad [[phab:T417873|T417873]] * 13:13 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] (duration: 08m 29s) * 13:13 moritzm: installing qemu security updates * 13:11 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 13:11 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 13:09 jgiannelos@deploy1003: jgiannelos: Continuing with deployment * 13:08 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 13:07 jgiannelos@deploy1003: jgiannelos: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:06 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 13:06 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2214.codfw.wmnet with reason: Maintenance * 13:05 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2214: Repooling after switchover * 13:05 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] * 13:04 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2214: Repooling after switchover * 13:04 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2214 [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94660 and previous config saved to /var/cache/conftool/dbconfig/20260701-130413-fceratto.json * 13:01 moritzm: installing python3.13 security updates * 13:00 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2229 to s6 primary [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94659 and previous config saved to /var/cache/conftool/dbconfig/20260701-125959-fceratto.json * 12:59 federico3: Starting s6 codfw failover from db2214 to db2229 - [[phab:T430814|T430814]] * 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on 13 hosts with reason: router upgrade and line card install * 12:51 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2229 with weight 0 [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94658 and previous config saved to /var/cache/conftool/dbconfig/20260701-125149-fceratto.json * 12:51 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 21 hosts with reason: Primary switchover s6 [[phab:T430814|T430814]] * 12:50 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2189.codfw.wmnet * 12:50 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2189.codfw.wmnet * 12:42 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2100.codfw.wmnet with OS trixie * 12:38 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2083.codfw.wmnet with OS trixie * 12:19 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2083.codfw.wmnet with reason: host reimage * 12:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 12:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2240: Migration of db2240.codfw.wmnet completed * 12:14 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2100.codfw.wmnet with reason: host reimage * 12:09 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2083.codfw.wmnet with reason: host reimage * 12:09 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2100.codfw.wmnet with reason: host reimage * 12:00 topranks: drain traffic on cr1-eqiad to allow for line card install and JunOS upgrade [[phab:T426343|T426343]] * 11:52 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2083.codfw.wmnet with OS trixie * 11:50 cmooney@dns2005: END - running authdns-update * 11:49 cmooney@dns2005: START - running authdns-update * 11:48 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2100.codfw.wmnet with OS trixie * 11:40 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/zotero: apply * 11:40 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/zotero: apply * 11:36 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/zotero: apply * 11:36 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/zotero: apply * 11:31 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2240: Migration of db2240.codfw.wmnet completed * 11:30 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply * 11:28 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply * 11:27 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:27 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:27 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:27 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:27 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:23 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2240.codfw.wmnet with OS trixie * 11:20 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:20 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:17 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:16 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:16 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:15 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2086.codfw.wmnet with OS trixie * 11:14 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2106.codfw.wmnet with OS trixie * 11:14 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:13 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:12 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:09 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2115.codfw.wmnet with OS trixie * 11:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2240.codfw.wmnet with reason: host reimage * 11:00 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2240.codfw.wmnet with reason: host reimage * 10:53 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2106.codfw.wmnet with reason: host reimage * 10:49 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2086.codfw.wmnet with reason: host reimage * 10:44 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2115.codfw.wmnet with reason: host reimage * 10:44 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2240.codfw.wmnet with OS trixie * 10:44 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2086.codfw.wmnet with reason: host reimage * 10:42 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2106.codfw.wmnet with reason: host reimage * 10:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2240: Upgrading db2240.codfw.wmnet * 10:41 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2240: Upgrading db2240.codfw.wmnet * 10:41 cwilliams@cumin1003: dbmaint on s4@codfw [[phab:T429893|T429893]] * 10:40 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 10:39 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2115.codfw.wmnet with reason: host reimage * 10:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2240 [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94653 and previous config saved to /var/cache/conftool/dbconfig/20260701-102658-cwilliams.json * 10:26 moritzm: installing nginx security updates * 10:26 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2086.codfw.wmnet with OS trixie * 10:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2179 to s4 primary [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94652 and previous config saved to /var/cache/conftool/dbconfig/20260701-102356-cwilliams.json * 10:23 cezmunsta: Starting s4 codfw failover from db2240 to db2179 - [[phab:T430127|T430127]] * 10:23 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2106.codfw.wmnet with OS trixie * 10:20 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2115.codfw.wmnet with OS trixie * 10:15 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2179 with weight 0 [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94651 and previous config saved to /var/cache/conftool/dbconfig/20260701-101531-cwilliams.json * 10:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 40 hosts with reason: Primary switchover s4 [[phab:T430127|T430127]] * 09:56 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template (take 2) - oblivian@cumin1003" * 09:56 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template (take 2) - oblivian@cumin1003 * 09:55 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template (take 2) - oblivian@cumin1003 * 09:55 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template (take 2) - oblivian@cumin1003" * 09:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:39 mszwarc@deploy1003: Synchronized private/SuggestedInvestigationsSignals/SuggestedInvestigationsSignal4n.php: Update SI signal 4n (duration: 06m 08s) * 09:21 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 09:21 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 09:14 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 09:14 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 09:02 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 09:02 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 08:54 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 08:38 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 08:38 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 08:36 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 08:21 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 08:21 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] (duration: 36m 11s) * 08:15 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 08:09 mszwarc@deploy1003: mszwarc, abi: Continuing with deployment * 08:03 mszwarc@deploy1003: mszwarc, abi: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:55 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 07:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 07:45 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] * 07:30 aqu@deploy1003: Finished deploy [analytics/refinery@410f205]: Regular analytics weekly train 2nd try [analytics/refinery@410f2050] (duration: 00m 22s) * 07:29 aqu@deploy1003: Started deploy [analytics/refinery@410f205]: Regular analytics weekly train 2nd try [analytics/refinery@410f2050] * 07:28 aqu@deploy1003: Finished deploy [analytics/refinery@410f205] (thin): Regular analytics weekly train THIN [analytics/refinery@410f2050] (duration: 01m 59s) * 07:28 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] (duration: 07m 19s) * 07:26 aqu@deploy1003: Started deploy [analytics/refinery@410f205] (thin): Regular analytics weekly train THIN [analytics/refinery@410f2050] * 07:26 aqu@deploy1003: Finished deploy [analytics/refinery@410f205]: Regular analytics weekly train [analytics/refinery@410f2050] (duration: 04m 32s) * 07:24 mszwarc@deploy1003: wmde-fisch, mszwarc: Continuing with deployment * 07:23 mszwarc@deploy1003: wmde-fisch, mszwarc: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:21 aqu@deploy1003: Started deploy [analytics/refinery@410f205]: Regular analytics weekly train [analytics/refinery@410f2050] * 07:21 aqu@deploy1003: Finished deploy [analytics/refinery@410f205] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@410f2050] (duration: 02m 01s) * 07:20 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] * 07:19 aqu@deploy1003: Started deploy [analytics/refinery@410f205] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@410f2050] * 07:13 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] (duration: 09m 13s) * 07:09 mszwarc@deploy1003: mszwarc, chlod, revi: Continuing with deployment * 07:06 mszwarc@deploy1003: mszwarc, chlod, revi: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:04 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] * 06:55 elukey: upgrade all trixie hosts to pywmflib 3.0 - [[phab:T430552|T430552]] * 06:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:43 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:43 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:42 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:42 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:41 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:41 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:35 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:35 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:34 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:34 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:31 jmm@cumin2003: DONE (PASS) - Cookbook sre.idm.logout (exit_code=0) Logging Niharika29 out of all services on: 2453 hosts * 06:30 oblivian@cumin1003: END (FAIL) - Cookbook sre.deploy.hiddenparma (exit_code=99) Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:30 oblivian@cumin1003: END (FAIL) - Cookbook sre.deploy.python-code (exit_code=99) hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:30 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:30 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:01 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2109.codfw.wmnet with OS trixie * 05:45 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on es1039.eqiad.wmnet with reason: issues * 05:41 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1027.eqiad.wmnet * 05:40 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2068.codfw.wmnet with OS trixie * 05:40 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2109.codfw.wmnet with reason: host reimage * 05:40 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1027.eqiad.wmnet,service=s2 * 05:40 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1027.eqiad.wmnet,service=s7 * 05:36 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2109.codfw.wmnet with reason: host reimage * 05:20 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2068.codfw.wmnet with reason: host reimage * 05:16 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2109.codfw.wmnet with OS trixie * 05:15 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2068.codfw.wmnet with reason: host reimage * 05:09 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2067.codfw.wmnet with OS trixie * 04:56 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2068.codfw.wmnet with OS trixie * 04:49 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2067.codfw.wmnet with reason: host reimage * 04:45 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2067.codfw.wmnet with reason: host reimage * 04:27 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2067.codfw.wmnet with OS trixie * 03:47 slyngshede@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1039.eqiad.wmnet with reason: Hardware crash * 03:21 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2107.codfw.wmnet with OS trixie * 02:59 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2107.codfw.wmnet with reason: host reimage * 02:55 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2085.codfw.wmnet with OS trixie * 02:51 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2072.codfw.wmnet with OS trixie * 02:51 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2107.codfw.wmnet with reason: host reimage * 02:35 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2085.codfw.wmnet with reason: host reimage * 02:31 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2107.codfw.wmnet with OS trixie * 02:30 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2072.codfw.wmnet with reason: host reimage * 02:26 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2085.codfw.wmnet with reason: host reimage * 02:22 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2072.codfw.wmnet with reason: host reimage * 02:09 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2085.codfw.wmnet with OS trixie * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 54s) * 02:03 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2072.codfw.wmnet with OS trixie * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es7 eqiad back to read-write - [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94649 and previous config saved to /var/cache/conftool/dbconfig/20260701-010716-ladsgroup.json * 01:05 ladsgroup@dns1004: END - running authdns-update * 01:05 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depool es1039 [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94648 and previous config saved to /var/cache/conftool/dbconfig/20260701-010551-ladsgroup.json * 01:03 ladsgroup@dns1004: START - running authdns-update * 01:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Promote es1035 to es7 primary [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94647 and previous config saved to /var/cache/conftool/dbconfig/20260701-010002-ladsgroup.json * 00:58 Amir1: Starting es7 eqiad failover from es1039 to es1035 - [[phab:T430765|T430765]] * 00:53 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es1035 with weight 0 [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94646 and previous config saved to /var/cache/conftool/dbconfig/20260701-005329-ladsgroup.json * 00:53 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 9 hosts with reason: Primary switchover es7 [[phab:T430765|T430765]] * 00:42 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es7 eqiad as read-only for maintenance - [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94645 and previous config saved to /var/cache/conftool/dbconfig/20260701-004221-ladsgroup.json * 00:20 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2102.codfw.wmnet with OS trixie * 00:15 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2103.codfw.wmnet with OS trixie * 00:05 dr0ptp4kt: DEPLOYED Refinery at {{Gerrit|4e7a2b32}} for changes: pageview allowlist {{Gerrit|1305158}} (+min.wikiquote) {{Gerrit|1305162}} (+bol.wikipedia), {{Gerrit|1305156}} (+isv.wikipedia); {{Gerrit|1305980}} (pv allowlist -api.wikimedia, sqoop +isvwiki); sqoop {{Gerrit|1295064}} (+globalimagelinks) {{Gerrit|1295069}} (+filerevision) using scap, then deployed onto HDFS (manual copyToLocal required additionally) == Other archives == See [[Server Admin Log/Archives]]. <noinclude> [[Category:SAL]] [[Category:Operations]] </noinclude> mqtfthw9n5fw458wqj99up8uq2yzlg8 2450654 2450653 2026-08-22T16:36:14Z Stashbot 7414 arlolra@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply 2450654 wikitext text/x-wiki == 2026-08-22 == * 16:36 arlolra@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:36 arlolra@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:36 arlolra@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:33 arlolra@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:33 arlolra@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:33 arlolra@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:32 arlolra@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 35s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-21 == * 20:36 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in1001.wikimedia.org with reason: [[phab:T434750|T434750]] * 20:34 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in2001.wikimedia.org with reason: [[phab:T434750|T434750]] * 20:33 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out1001.wikimedia.org with reason: [[phab:T434750|T434750]] * 20:25 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out2001.wikimedia.org with reason: [[phab:T434750|T434750]] * 19:37 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host krb1002.eqiad.wmnet with OS bookworm * 19:00 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 18:59 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 18:51 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 18:51 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 18:35 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:35 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:27 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:27 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:16 bking@cumin2003: START - Cookbook sre.hosts.reimage for host krb1002.eqiad.wmnet with OS bookworm * 17:35 sukhe@dns1004: END - running authdns-update * 17:33 sukhe@dns1004: START - running authdns-update * 17:32 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns5004.wikimedia.org [reason: resolved authdns-update issues] * 17:32 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:32 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: force HEAD to {{Gerrit|be26e30ae101}} - sukhe@cumin1003" * 17:32 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: force HEAD to {{Gerrit|be26e30ae101}} - sukhe@cumin1003" * 17:28 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 17:28 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: service=authdns-update,name=dns5004.wikimedia.org [reason: resolving authdns-update issues] * 17:28 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:28 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: force HEAD to {{Gerrit|be26e30ae101}} - sukhe@cumin1003" * 17:28 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: force HEAD to {{Gerrit|be26e30ae101}} - sukhe@cumin1003" * 17:24 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 17:24 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.netbox (exit_code=97) * 17:23 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 17:16 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 17:12 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 17:10 sukhe@dns1004: END - running authdns-update * 17:08 sukhe@dns1004: START - running authdns-update * 17:08 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=dns5004.wikimedia.org [reason: resolving authdns-update issues] * 17:07 sukhe@dns1004: FAIL - running authdns-update * 17:05 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 17:05 sukhe@dns1004: START - running authdns-update * 17:01 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 16:59 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 16:56 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 16:53 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=dns5004.* [reason: trixie upgrade] * 16:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns5004.wikimedia.org * 16:52 cdobbins@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns5004.wikimedia.org * 16:44 cmooney@dns3003: END - running authdns-update * 16:41 cmooney@dns3003: START - running authdns-update * 16:41 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:41 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on eqsin<->codfw arelion - cmooney@cumin1003" * 16:37 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on eqsin<->codfw arelion - cmooney@cumin1003" * 16:33 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:11 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 16:08 sukhe@dns1004: END - running authdns-update * 16:08 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 16:06 sukhe@dns1004: START - running authdns-update * 16:04 cmooney@dns3003: END - running authdns-update * 16:02 cmooney@dns3003: START - running authdns-update * 16:00 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:00 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on eqord<->codfw arelion - cmooney@cumin1003" * 15:56 cdobbins@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host dns5004.wikimedia.org with OS trixie * 15:55 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on eqord<->codfw arelion - cmooney@cumin1003" * 15:53 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:51 cmooney@cumin1003: END (ERROR) - Cookbook sre.dns.netbox (exit_code=97) * 15:51 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:38 andrew@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudcephosd1042.eqiad.wmnet with OS bookworm * 15:18 andrew@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudcephosd1042.eqiad.wmnet with reason: host reimage * 15:17 dancy@deploy1003: Finished deploy [gerrit/gerrit@2cc11cc]: Deploying https://gerrit.wikimedia.org/r/c/operations/software/gerrit/+/1327669 ([[phab:T434726|T434726]]) (duration: 00m 14s) * 15:17 dancy@deploy1003: Started deploy [gerrit/gerrit@2cc11cc]: Deploying https://gerrit.wikimedia.org/r/c/operations/software/gerrit/+/1327669 ([[phab:T434726|T434726]]) * 15:13 andrew@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudcephosd1042.eqiad.wmnet with reason: host reimage * 15:09 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns5004.wikimedia.org with reason: host reimage * 15:05 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns5004.wikimedia.org with reason: host reimage * 14:53 andrew@cumin2003: START - Cookbook sre.hosts.reimage for host cloudcephosd1042.eqiad.wmnet with OS bookworm * 14:30 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns5004.wikimedia.org with OS trixie * 14:29 cdobbins@cumin1003: conftool action : set/pooled=no; selector: name=dns5004.* [reason: trixie upgrade] * 14:21 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:21 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove entries for cr2-eqord - cmooney@cumin1003" * 14:21 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove entries for cr2-eqord - cmooney@cumin1003" * 14:13 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 14:11 moritzm: imported openjdk 8u504-ga-1~deb12u1 for bookworm-wikimedia (backport of the latest Java 8 security fixes for bookworm) * 13:25 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "sync cr2-eqord router offline - cmooney@cumin1003" * 13:23 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "sync cr2-eqord router offline - cmooney@cumin1003" * 13:14 hashar@deploy1003: Finished deploy [integration/docroot@2d5ff9b]: opensource: add PersonalDashboard docs to MW components - [[phab:T435392|T435392]] (duration: 00m 15s) * 13:14 hashar@deploy1003: Started deploy [integration/docroot@2d5ff9b]: opensource: add PersonalDashboard docs to MW components - [[phab:T435392|T435392]] * 12:16 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:16 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: [[phab:T431682|T431682]] - filippo@cumin1003" * 12:16 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: [[phab:T431682|T431682]] - filippo@cumin1003" * 12:11 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2006.wikimedia.org with OS trixie * 12:00 kevinbazira@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 11:58 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 11:43 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2006.wikimedia.org with reason: host reimage * 11:41 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 11:38 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2006.wikimedia.org with reason: host reimage * 11:20 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2006.wikimedia.org with OS trixie * 11:11 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2005.wikimedia.org with OS trixie * 10:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2005.wikimedia.org with reason: host reimage * 10:53 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2005.wikimedia.org with reason: host reimage * 10:43 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-codfw * 10:43 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2011.codfw.wmnet * 10:43 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2011.codfw.wmnet * 10:40 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2011.codfw.wmnet * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2011.codfw.wmnet * 10:34 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2010.codfw.wmnet * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2010.codfw.wmnet * 10:33 fnegri@cumin1003: END (PASS) - Cookbook sre.wikireplicas.add-wiki (exit_code=0) for database bolwiki ([[phab:T429954|T429954]]) * 10:33 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2005.wikimedia.org with OS trixie * 10:30 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2010.codfw.wmnet * 10:25 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2010.codfw.wmnet * 10:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2009.codfw.wmnet * 10:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2009.codfw.wmnet * 10:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1006.wikimedia.org with OS trixie * 10:18 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2009.codfw.wmnet * 10:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2009.codfw.wmnet * 10:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2008.codfw.wmnet * 10:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2008.codfw.wmnet * 10:06 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2008.codfw.wmnet * 10:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1006.wikimedia.org with reason: host reimage * 10:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2008.codfw.wmnet * 10:01 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2007.codfw.wmnet * 10:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2007.codfw.wmnet * 09:57 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1006.wikimedia.org with reason: host reimage * 09:56 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2007.codfw.wmnet * 09:51 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2007.codfw.wmnet * 09:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 09:51 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 09:46 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2006.codfw.wmnet * 09:46 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1006.wikimedia.org with OS trixie * 09:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1005.wikimedia.org with OS trixie * 09:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2006.codfw.wmnet * 09:41 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2005.codfw.wmnet * 09:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2005.codfw.wmnet * 09:38 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2013.codfw.wmnet * 09:36 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2005.codfw.wmnet * 09:35 fnegri@cumin1003: START - Cookbook sre.wikireplicas.add-wiki for database bolwiki ([[phab:T429954|T429954]]) * 09:35 fnegri@cumin1003: END (PASS) - Cookbook sre.wikireplicas.add-wiki (exit_code=0) for database minwikiquote ([[phab:T429946|T429946]]) * 09:35 fnegri@cumin1003: START - Cookbook sre.wikireplicas.add-wiki for database minwikiquote ([[phab:T429946|T429946]]) * 09:32 blake@cumin1003: START - Cookbook sre.hosts.reboot-single for host rdb2013.codfw.wmnet * 09:30 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2011.codfw.wmnet * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1005.wikimedia.org with reason: host reimage * 09:26 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2005.codfw.wmnet * 09:26 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2004.codfw.wmnet * 09:26 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2004.codfw.wmnet * 09:24 blake@cumin1003: START - Cookbook sre.hosts.reboot-single for host rdb2011.codfw.wmnet * 09:22 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1005.wikimedia.org with reason: host reimage * 09:21 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2004.codfw.wmnet * 09:16 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1015.eqiad.wmnet * 09:16 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2004.codfw.wmnet * 09:16 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2003.codfw.wmnet * 09:16 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2003.codfw.wmnet * 09:13 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on cr[1-2]-eqiad,pfw1-eqiad with reason: upgrade pfw1a-eqiad and pfw1b-eqiad pair * 09:12 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 09:11 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 09:11 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 09:11 blake@cumin1003: START - Cookbook sre.hosts.reboot-single for host rdb1015.eqiad.wmnet * 09:10 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2003.codfw.wmnet * 09:09 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1013.eqiad.wmnet * 09:07 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1005.wikimedia.org with OS trixie * 09:03 blake@cumin1003: START - Cookbook sre.hosts.reboot-single for host rdb1013.eqiad.wmnet * 09:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2003.codfw.wmnet * 09:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2002.codfw.wmnet * 09:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2002.codfw.wmnet * 08:54 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2002.codfw.wmnet * 08:49 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2002.codfw.wmnet * 08:49 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2001.codfw.wmnet * 08:49 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2001.codfw.wmnet * 08:48 jmm@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts netmon2002.wikimedia.org * 08:47 jmm@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts netmon2002.wikimedia.org * 08:44 jmm@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts netmon2002.wikimedia.org * 08:44 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon2002.wikimedia.org * 08:43 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2001.codfw.wmnet * 08:36 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon2002.wikimedia.org * 08:34 jmm@dns1004: END - running authdns-update * 08:33 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2001.codfw.wmnet * 08:33 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-codfw * 08:31 jmm@dns1004: START - running authdns-update * 07:48 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327679{{!}}Block: Disable flaky API test (T435272 T389028)]], [[gerrit:1327678{{!}}API: wfDebugLog for thumberror]] (duration: 15m 34s) * 07:41 krinkle@deploy1003: krinkle: Continuing with deployment * 07:37 krinkle@deploy1003: krinkle: Backport for [[gerrit:1327679{{!}}Block: Disable flaky API test (T435272 T389028)]], [[gerrit:1327678{{!}}API: wfDebugLog for thumberror]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:33 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1327679{{!}}Block: Disable flaky API test (T435272 T389028)]], [[gerrit:1327678{{!}}API: wfDebugLog for thumberror]] * 07:25 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 07:24 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 07:18 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 07:18 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 07:15 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 07:14 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 07:14 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 07:14 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 07:13 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 07:03 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1283: Pool back * 06:42 jmm@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts netmon2002.wikimedia.org * 06:35 moritzm: powercycling netmon2002 * 06:18 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1283: Pool back * 06:17 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1283 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96213 and previous config saved to /var/cache/conftool/dbconfig/20260821-061743-marostegui.json * 04:59 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324963{{!}}Add Produnto to extension-list (T421436)]], [[gerrit:1324964{{!}}Enable Produnto on Beta (T421436)]] (duration: 34m 48s) * 04:45 tstarling@deploy1003: tstarling: Continuing with deployment * 04:44 tstarling@deploy1003: tstarling: Backport for [[gerrit:1324963{{!}}Add Produnto to extension-list (T421436)]], [[gerrit:1324964{{!}}Enable Produnto on Beta (T421436)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 04:24 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1324963{{!}}Add Produnto to extension-list (T421436)]], [[gerrit:1324964{{!}}Enable Produnto on Beta (T421436)]] * 04:21 arlolra@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 04:20 arlolra@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 04:20 arlolra@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 04:20 arlolra@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 41s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-20 == * 23:43 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327652{{!}}RunSingleJob: Add ProfilingContext::init() (T435422)]] (duration: 12m 06s) * 23:38 krinkle@deploy1003: krinkle: Continuing with deployment * 23:33 krinkle@deploy1003: krinkle: Backport for [[gerrit:1327652{{!}}RunSingleJob: Add ProfilingContext::init() (T435422)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:31 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1327652{{!}}RunSingleJob: Add ProfilingContext::init() (T435422)]] * 22:15 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1054.eqiad.wmnet with OS trixie * 22:14 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 22:14 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 21:58 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1054.eqiad.wmnet with reason: host reimage * 21:51 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1054.eqiad.wmnet with reason: host reimage * 21:36 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1054.eqiad.wmnet with OS trixie * 21:36 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:35 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327219{{!}}RunSingleJob: Define MW_ENTRY_POINT for flamegraph sample attribution (T435422)]] (duration: 08m 30s) * 21:31 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:31 krinkle@deploy1003: krinkle: Continuing with deployment * 21:31 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1054 * 21:31 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1054 * 21:29 krinkle@deploy1003: krinkle: Backport for [[gerrit:1327219{{!}}RunSingleJob: Define MW_ENTRY_POINT for flamegraph sample attribution (T435422)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:27 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1327219{{!}}RunSingleJob: Define MW_ENTRY_POINT for flamegraph sample attribution (T435422)]] * 21:17 maryum: Deployed security fix for [[phab:T433020|T433020]] * 20:59 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324752{{!}}InitialiseSettings: Enable 2FA banners on remaining private wikis (T428103)]], [[gerrit:1325920{{!}}Remove sending email to legal team about rejected requests (T374053)]] (duration: 07m 18s) * 20:54 reedy@deploy1003: neriah, reedy: Continuing with deployment * 20:54 reedy@deploy1003: neriah, reedy: Backport for [[gerrit:1324752{{!}}InitialiseSettings: Enable 2FA banners on remaining private wikis (T428103)]], [[gerrit:1325920{{!}}Remove sending email to legal team about rejected requests (T374053)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:51 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324752{{!}}InitialiseSettings: Enable 2FA banners on remaining private wikis (T428103)]], [[gerrit:1325920{{!}}Remove sending email to legal team about rejected requests (T374053)]] * 20:24 reedy@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.15,1.47.0-wmf.16,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/med * 20:23 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324752{{!}}InitialiseSettings: Enable 2FA banners on remaining private wikis (T428103)]], [[gerrit:1325920{{!}}Remove sending email to legal team about rejected requests (T374053)]] * 20:14 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 20:10 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 20:09 cdanis@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "bug fixes & UX fixes - cdanis@cumin1003" * 20:09 cdanis@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: bug fixes & UX fixes - cdanis@cumin1003 * 20:08 cdanis@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: bug fixes & UX fixes - cdanis@cumin1003 * 20:08 cdanis@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "bug fixes & UX fixes - cdanis@cumin1003" * 19:24 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327598{{!}}Make \Omicron non upright (like \Chi) (T434428)]], [[gerrit:1327596{{!}}Render overline of \bar with stretchy=false (T435456)]] (duration: 18m 54s) * 19:20 krinkle@deploy1003: krinkle: Continuing with deployment * 19:07 krinkle@deploy1003: krinkle: Backport for [[gerrit:1327598{{!}}Make \Omicron non upright (like \Chi) (T434428)]], [[gerrit:1327596{{!}}Render overline of \bar with stretchy=false (T435456)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:05 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1327598{{!}}Make \Omicron non upright (like \Chi) (T434428)]], [[gerrit:1327596{{!}}Render overline of \bar with stretchy=false (T435456)]] * 18:51 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327614{{!}}Avoid casting fpxmax to string (T318419)]] (duration: 07m 28s) * 18:50 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:46 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:46 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 18:45 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327614{{!}}Avoid casting fpxmax to string (T318419)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:43 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327614{{!}}Avoid casting fpxmax to string (T318419)]] * 18:37 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:37 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:37 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:36 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:07 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 18:05 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 18:01 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 18:01 sukhe@dns1004: END - running authdns-update * 17:59 sukhe@dns1004: START - running authdns-update * 17:58 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 17:57 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=dns6002.* [reason: depooling for trixie upgrade] * 17:56 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns6002.wikimedia.org * 17:56 cdobbins@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns6002.wikimedia.org * 17:51 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 17:51 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 17:34 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host stat1011.eqiad.wmnet with OS bookworm * 17:31 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns6002.wikimedia.org with OS trixie * 17:30 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 17:30 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 17:30 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 17:30 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 17:29 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 17:29 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 17:29 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 17:29 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:29 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:27 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:24 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327590{{!}}Make sure fpsmax is an int value (T318419)]] (duration: 08m 37s) * 17:20 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 17:17 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327590{{!}}Make sure fpsmax is an int value (T318419)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:16 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 17:15 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327590{{!}}Make sure fpsmax is an int value (T318419)]] * 16:51 swfrench-wmf: disable-puppet on A:cp for ATS Lua change - [[phab:T427666|T427666]] * 16:51 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db2901.codfw.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 16:43 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 16:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on stat1011.eqiad.wmnet with reason: host reimage * 16:39 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns6002.wikimedia.org with reason: host reimage * 16:36 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on stat1011.eqiad.wmnet with reason: host reimage * 16:36 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db2901.codfw.wmnet * 16:34 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns6002.wikimedia.org with reason: host reimage * 16:15 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns6002.wikimedia.org with OS trixie * 16:14 cdobbins@cumin1003: conftool action : set/pooled=no; selector: name=dns6002.* [reason: depooling for trixie upgrade] * 16:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host stat1011 * 16:10 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host stat1011 * 16:09 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host stat1011 * 16:09 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) stat1011.eqiad.wmnet 14.36.64.10.in-addr.arpa 4.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 bking@cumin2003: START - Cookbook sre.dns.wipe-cache stat1011.eqiad.wmnet 14.36.64.10.in-addr.arpa 4.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:09 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host stat1011 - bking@cumin2003" * 16:09 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host stat1011 - bking@cumin2003" * 16:05 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host stat1011 * 16:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host stat1011.eqiad.wmnet with OS bookworm * 16:00 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 15:59 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 15:56 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 15:55 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 15:37 jayme@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:35 jayme@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 15:35 jayme@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:32 jayme@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:32 jayme@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:30 jayme@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 15:30 jayme@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:28 jayme@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:28 jayme@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 15:28 fceratto@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host db1903.eqiad.wmnet * 15:28 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1903.eqiad.wmnet with OS trixie * 15:26 jayme@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 15:26 jayme@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 15:24 jayme@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 15:24 jayme@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 15:21 jayme@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 15:21 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 15:19 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 15:19 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 15:17 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 15:14 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1903.eqiad.wmnet with reason: host reimage * 15:07 fceratto@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1903.eqiad.wmnet with reason: host reimage * 14:54 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db1903.eqiad.wmnet with OS trixie * 14:53 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1903.eqiad.wmnet - fceratto@cumin1003" * 14:53 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1903.eqiad.wmnet - fceratto@cumin1003" * 14:53 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1903.eqiad.wmnet on all recursors * 14:53 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1903.eqiad.wmnet on all recursors * 14:53 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:53 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1903.eqiad.wmnet - fceratto@cumin1003" * 14:53 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1903.eqiad.wmnet - fceratto@cumin1003" * 14:49 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 14:49 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1903.eqiad.wmnet * 14:33 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=cp1100.* * 14:27 topranks: reconfigure eqiad<->codfw bgp settings * 14:22 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:22 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update entries used on new transport backup eqiad codfw - cmooney@cumin1003" * 14:19 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update entries used on new transport backup eqiad codfw - cmooney@cumin1003" * 14:14 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 14:14 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 14:13 moritzm: installing util-linux security updates * 14:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-staging-worker * 14:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2003.codfw.wmnet * 14:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2003.codfw.wmnet * 14:08 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2003.codfw.wmnet * 14:06 moritzm: installing libheif security updates * 13:58 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2003.codfw.wmnet * 13:58 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2002.codfw.wmnet * 13:58 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2002.codfw.wmnet * 13:56 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:56 fnegri@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for clouddb1025.eqiad.wmnet * 13:56 fnegri@cumin1003: START - Cookbook sre.hosts.remove-downtime for clouddb1025.eqiad.wmnet * 13:56 Lucas_WMDE: UTC afternoon backport+config window done * 13:53 moritzm: installing apr-util security updates * 13:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2002.codfw.wmnet * 13:50 fnegri@cumin1003: conftool action : set/weight=100; selector: name=clouddb1025.eqiad.wmnet * 13:49 fnegri@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1025.eqiad.wmnet * 13:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2002.codfw.wmnet * 13:41 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2001.codfw.wmnet * 13:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2001.codfw.wmnet * 13:41 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'. * 13:38 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'. * 13:38 fnegri@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on clouddb1025.eqiad.wmnet with reason: Removing s6 from clouddb1025 * 13:34 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2001.codfw.wmnet * 13:31 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'. * 13:29 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'. * 13:28 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet * 13:26 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host stat1009.eqiad.wmnet with OS bookworm * 13:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2001.codfw.wmnet * 13:24 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-staging-worker * 13:23 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1002.eqiad.wmnet * 13:21 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host stat1010.eqiad.wmnet with OS bookworm * 13:20 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1002.eqiad.wmnet * 13:20 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1001.eqiad.wmnet * 13:17 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1001.eqiad.wmnet * 13:16 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2001.codfw.wmnet * 13:13 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327511{{!}}UIC: Fix page:page instead of page:other in instrumentation]] (duration: 07m 00s) * 13:13 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2001.codfw.wmnet * 13:12 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2002.codfw.wmnet * 13:09 mszwarc@deploy1003: mszwarc: Continuing with deployment * 13:08 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1327511{{!}}UIC: Fix page:page instead of page:other in instrumentation]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2002.codfw.wmnet * 13:07 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2002.codfw.wmnet * 13:06 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1327511{{!}}UIC: Fix page:page instead of page:other in instrumentation]] * 13:04 jmm@dns1004: END - running authdns-update * 13:03 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2002.codfw.wmnet * 13:03 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2001.codfw.wmnet * 13:02 jmm@dns1004: START - running authdns-update * 13:01 cmooney@dns3003: END - running authdns-update * 13:00 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2001.codfw.wmnet * 12:59 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2001.codfw.wmnet * 12:59 cmooney@dns3003: START - running authdns-update * 12:57 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2001.codfw.wmnet * 12:56 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2002.codfw.wmnet * 12:55 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:55 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on drmrs<->eqiad GTT vpls - cmooney@cumin1003" * 12:54 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on drmrs<->eqiad GTT vpls - cmooney@cumin1003" * 12:54 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2002.codfw.wmnet * 12:54 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2003.codfw.wmnet * 12:50 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2003.codfw.wmnet * 12:49 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 12:48 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2001.codfw.wmnet * 12:46 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2001.codfw.wmnet * 12:46 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2002.codfw.wmnet * 12:43 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2002.codfw.wmnet * 12:43 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2003.codfw.wmnet * 12:42 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=1) for new host db1902.eqiad.wmnet * 12:42 fceratto@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host db1902.eqiad.wmnet with OS trixie * 12:41 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2003.codfw.wmnet * 12:40 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1003.eqiad.wmnet * 12:38 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1003.eqiad.wmnet * 12:37 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1002.eqiad.wmnet * 12:35 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1002.eqiad.wmnet * 12:35 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1001.eqiad.wmnet * 12:34 cmooney@dns3003: END - running authdns-update * 12:33 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1001.eqiad.wmnet * 12:32 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on stat1009.eqiad.wmnet with reason: host reimage * 12:31 cmooney@dns3003: START - running authdns-update * 12:31 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:31 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on drmrs<->eqiad cct - cmooney@cumin1003" * 12:28 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on drmrs<->eqiad cct - cmooney@cumin1003" * 12:28 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1902.eqiad.wmnet with reason: host reimage * 12:25 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on stat1009.eqiad.wmnet with reason: host reimage * 12:25 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 12:24 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on stat1010.eqiad.wmnet with reason: host reimage * 12:22 fceratto@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1902.eqiad.wmnet with reason: host reimage * 12:21 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on stat1010.eqiad.wmnet with reason: host reimage * 12:14 elukey: move the Docker Registry's /v2/dev/.* prefix to its dedicated S3 backend - [[phab:T432829|T432829]] * 12:12 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db1902.eqiad.wmnet with OS trixie * 12:09 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1902.eqiad.wmnet - fceratto@cumin1003" * 12:09 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1902.eqiad.wmnet - fceratto@cumin1003" * 12:09 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1902.eqiad.wmnet on all recursors * 12:09 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1902.eqiad.wmnet on all recursors * 12:09 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:08 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1902.eqiad.wmnet - fceratto@cumin1003" * 12:08 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1902.eqiad.wmnet - fceratto@cumin1003" * 12:08 tgr_: [[phab:T413390|T413390]] running CentralAuth:FixRenamedUserGlobalEditCount --wiki=metawiki --since=20250901000000 --until=20260301000000 --fix * 12:04 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1009.eqiad.wmnet with OS bookworm * 12:04 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1010.eqiad.wmnet with OS bookworm * 12:01 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 12:01 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1902.eqiad.wmnet * 12:00 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host stat1010.eqiad.wmnet with OS bookworm * 11:57 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1902.eqiad.wmnet * 11:57 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:57 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1902.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 11:57 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1902.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 11:51 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'. * 11:49 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'. * 11:48 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'. * 11:46 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'. * 11:39 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 11:37 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327518{{!}}Enable thumb.wikimedia.org on cswiki and fawiki (T427465)]] (duration: 10m 40s) * 11:35 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1902.eqiad.wmnet * 11:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1010.eqiad.wmnet with OS bookworm * 11:33 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 11:30 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327518{{!}}Enable thumb.wikimedia.org on cswiki and fawiki (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:26 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327518{{!}}Enable thumb.wikimedia.org on cswiki and fawiki (T427465)]] * 11:24 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host stat1010.eqiad.wmnet with OS bookworm * 11:07 fceratto@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host db1901.eqiad.wmnet * 11:07 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1901.eqiad.wmnet with OS trixie * 10:53 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1901.eqiad.wmnet with reason: host reimage * 10:47 tappof: bump space for prometheus k8s-aux in codfw * 10:47 tappof: bump space for prometheus k8s-dse in eqiad * 10:47 fceratto@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1901.eqiad.wmnet with reason: host reimage * 10:35 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db1901.eqiad.wmnet with OS trixie * 10:32 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:32 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:32 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1901.eqiad.wmnet on all recursors * 10:32 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1901.eqiad.wmnet on all recursors * 10:31 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:31 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:31 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:27 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:27 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1901.eqiad.wmnet * 10:23 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1010.eqiad.wmnet with OS bookworm * 10:20 fceratto@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts db1901.eqiad.wmnet * 10:20 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 10:18 blake@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 10:17 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:16 blake@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 10:13 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1901.eqiad.wmnet * 09:23 jelto@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'. * 09:22 jelto@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'. * 09:22 jelto: update cert-manager to 1.19.6 on wikikube staging-eqiad - [[phab:T427402|T427402]] * 09:20 moritzm: imported squid 7.6-2.1for trixie-wikimedia/main [[phab:T427282|T427282]] * 09:08 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 09:08 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 09:08 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 09:07 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 09:04 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 09:04 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:23 slyngshede@dns1004: END - running authdns-update * 08:21 slyngshede@dns1004: START - running authdns-update * 08:18 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.16 refs [[phab:T430835|T430835]] * 06:27 aokoth@dns1004: END - running authdns-update * 06:25 aokoth@dns1004: START - running authdns-update * 06:22 brennen@deploy1003: Finished deploy [phabricator/deployment@6b9b6ff]: deploy phab1005 for [[phab:T435087|T435087]] (duration: 00m 39s) * 06:21 brennen@deploy1003: Started deploy [phabricator/deployment@6b9b6ff]: deploy phab1005 for [[phab:T435087|T435087]] * 06:20 brennen@deploy1003: Finished deploy [phabricator/deployment@6b9b6ff]: deploy phab1004 for to pick up config values for [[phab:T435087|T435087]] (duration: 01m 46s) * 06:18 brennen@deploy1003: Started deploy [phabricator/deployment@6b9b6ff]: deploy phab1004 for to pick up config values for [[phab:T435087|T435087]] * 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 49s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-19 == * 23:19 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327197{{!}}Enable thumb.wikimedia.org on mediawiki.org (T427465)]] (duration: 10m 50s) * 23:18 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:16 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 23:15 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 23:10 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327197{{!}}Enable thumb.wikimedia.org on mediawiki.org (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:10 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:09 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 23:09 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:09 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 23:08 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327197{{!}}Enable thumb.wikimedia.org on mediawiki.org (T427465)]] * 22:58 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327201{{!}}Enable ReadingLists for all logged in users on test wiki (T435258)]] (duration: 11m 20s) * 22:50 jdlrobson@deploy1003: jdlrobson: Continuing with deployment * 22:49 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1327201{{!}}Enable ReadingLists for all logged in users on test wiki (T435258)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:46 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1327201{{!}}Enable ReadingLists for all logged in users on test wiki (T435258)]] * 22:42 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327162{{!}}Article: Split subjectpageheader by model and disable for wikitext]], [[gerrit:1327169{{!}}Make uppercase greek letters normal (non-italic) font (T434686 T434428)]], [[gerrit:1327176{{!}}Skin: Avoid DB lookup for pagecategorieslink message (T347123)]] (duration: 37m 52s) * 22:29 krinkle@deploy1003: krinkle: Continuing with deployment * 22:25 krinkle@deploy1003: krinkle: Backport for [[gerrit:1327162{{!}}Article: Split subjectpageheader by model and disable for wikitext]], [[gerrit:1327169{{!}}Make uppercase greek letters normal (non-italic) font (T434686 T434428)]], [[gerrit:1327176{{!}}Skin: Avoid DB lookup for pagecategorieslink message (T347123)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:04 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1327162{{!}}Article: Split subjectpageheader by model and disable for wikitext]], [[gerrit:1327169{{!}}Make uppercase greek letters normal (non-italic) font (T434686 T434428)]], [[gerrit:1327176{{!}}Skin: Avoid DB lookup for pagecategorieslink message (T347123)]] * 22:04 krinkle@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: awaiting CI (duration: 03m 06s) * 22:01 krinkle@deploy1003: Locking from deployment [ALL REPOSITORIES]: awaiting CI * 22:00 krinkle@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: awaiting CI (duration: 00m 01s) * 22:00 krinkle@deploy1003: Locking from deployment [ALL REPOSITORIES]: awaiting CI * 21:34 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2001.codfw.wmnet * 21:28 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2001.codfw.wmnet * 21:22 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327178{{!}}AccountRecovery: Notify the email address of the on file of the request (T425799)]] (duration: 47m 02s) * 21:13 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm * 21:09 catrope@deploy1003: catrope: Continuing with deployment * 20:55 catrope@deploy1003: catrope: Backport for [[gerrit:1327178{{!}}AccountRecovery: Notify the email address of the on file of the request (T425799)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:35 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1327178{{!}}AccountRecovery: Notify the email address of the on file of the request (T425799)]] * 20:31 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327128{{!}}Parsoid DataAccess: convert Parsoid fragment markers to/from strip tags (T432547)]] (duration: 07m 30s) * 20:27 catrope@deploy1003: catrope, arlolra: Continuing with deployment * 20:26 catrope@deploy1003: catrope, arlolra: Backport for [[gerrit:1327128{{!}}Parsoid DataAccess: convert Parsoid fragment markers to/from strip tags (T432547)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:24 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1327128{{!}}Parsoid DataAccess: convert Parsoid fragment markers to/from strip tags (T432547)]] * 20:23 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage * 20:17 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage * 20:15 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325878{{!}}[arwiki] Enable restricted user page editing and grant edit permissions (T434878)]] (duration: 08m 46s) * 20:11 catrope@deploy1003: catrope, gergesshamon: Continuing with deployment * 20:08 catrope@deploy1003: catrope, gergesshamon: Backport for [[gerrit:1325878{{!}}[arwiki] Enable restricted user page editing and grant edit permissions (T434878)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:06 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1325878{{!}}[arwiki] Enable restricted user page editing and grant edit permissions (T434878)]] * 19:59 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm * 19:56 eevans@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cassandra-dev2001.codfw.wmnet with OS bookworm * 19:56 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm * 19:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2207.codfw.wmnet with reason: Maintenance * 18:47 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319920{{!}}Allow setting a separate thumbUrl in production (T427465)]], [[gerrit:1327167{{!}}Fix wmgThumbUrl config (T427465)]] (duration: 18m 53s) * 18:43 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 18:30 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1319920{{!}}Allow setting a separate thumbUrl in production (T427465)]], [[gerrit:1327167{{!}}Fix wmgThumbUrl config (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:28 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1319920{{!}}Allow setting a separate thumbUrl in production (T427465)]], [[gerrit:1327167{{!}}Fix wmgThumbUrl config (T427465)]] * 18:26 sukhe@dns1004: END - running authdns-update * 18:24 sukhe@dns1004: START - running authdns-update * 18:09 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1319920{{!}}Allow setting a separate thumbUrl in production (T427465)]] * 18:03 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-eqiad and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 17:56 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-codfw and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 17:50 cmooney@dns3003: END - running authdns-update * 17:42 dancy@deploy1003: Installation of scap version "4.283.0" completed for 3 hosts * 17:41 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling reboot on A:durum and A:durum * 17:40 cmooney@dns3003: START - running authdns-update * 17:40 dancy@deploy1003: Installing scap version "4.283.0" for 3 host(s) * 17:38 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:37 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on GTT VPLS - cmooney@cumin1003" * 17:37 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --verbose --scoreLessThan=0.7 --exceptDatasetChecksums=[[phab:T434319|T434319]]-enwiki-models.txt # [[phab:T434319|T434319]] * 17:32 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on GTT VPLS - cmooney@cumin1003" * 17:29 sbassett: Deployed security fix for [[phab:T435210|T435210]] (wmf.16) * 17:26 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 17:22 sbassett: Deployed security fix for [[phab:T435210|T435210]] (wmf.15) * 17:00 sukhe@dns1004: END - running authdns-update * 16:58 sukhe@dns1004: START - running authdns-update * 16:53 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica-esams and A:liberica * 16:41 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica-esams and A:liberica * 16:41 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-eqiad and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 16:41 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-codfw and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 16:41 cjd91: sudo -i cookbook sre.cdn.roll-upgrade-ats --query 'A:cp-codfw' --task-id [[phab:T434478|T434478]] --reason '9.2.15 upgrade' * 16:41 cjd91: sudo -i cookbook sre.cdn.roll-upgrade-ats --query 'A:cp-eqiad' --task-id [[phab:T434478|T434478]] --reason '9.2.15 upgrade' * 16:40 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and A:durum * 16:28 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326881{{!}}Echo: Start using virtual domains (T380385)]] (duration: 13m 12s) * 16:23 urbanecm@deploy1003: urbanecm: Continuing with deployment * 16:21 urandom: Completed sessionstore Cassandra/JVM upgrade — [[phab:T435154|T435154]] * 16:21 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching sessionstore[2005-2006].codfw.wmnet,sessionstore[1005-1006].eqiad.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 16:19 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1326881{{!}}Echo: Start using virtual domains (T380385)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:15 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:15 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->codfw - cmooney@cumin1003" * 16:14 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1326881{{!}}Echo: Start using virtual domains (T380385)]] * 16:14 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327123{{!}}Revert^2 "Migrate database access to virtual domains" (T435305)]], [[gerrit:1327124{{!}}Pass the mapped domain of virtual-echo-shared to the push NameTableStores (T435305)]] (duration: 07m 42s) * 16:13 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching sessionstore[2005-2006].codfw.wmnet,sessionstore[1005-1006].eqiad.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 16:11 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->codfw - cmooney@cumin1003" * 16:10 urbanecm@deploy1003: urbanecm: Continuing with deployment * 16:08 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1327123{{!}}Revert^2 "Migrate database access to virtual domains" (T435305)]], [[gerrit:1327124{{!}}Pass the mapped domain of virtual-echo-shared to the push NameTableStores (T435305)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:06 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 16:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 16:06 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1327123{{!}}Revert^2 "Migrate database access to virtual domains" (T435305)]], [[gerrit:1327124{{!}}Pass the mapped domain of virtual-echo-shared to the push NameTableStores (T435305)]] * 16:06 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:03 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching sessionstore1004.eqiad.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 16:01 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching sessionstore1004.eqiad.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 15:56 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching sessionstore2004.codfw.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 15:54 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching sessionstore2004.codfw.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 15:53 urandom: beginning sessionstore Cassandra/JVM upgrade — [[phab:T435154|T435154]] * 15:52 urandom: beginning sessionstore Cassandra/JVM upgrade — [[phab:T432944|T432944]] * 15:51 cmooney@dns3003: END - running authdns-update * 15:49 cmooney@dns3003: START - running authdns-update * 15:48 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:48 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->codfw - cmooney@cumin1003" * 15:45 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->codfw - cmooney@cumin1003" * 15:44 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 15:44 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:43 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 15:42 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:42 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:38 cmooney@dns3003: END - running authdns-update * 15:36 cmooney@dns3003: START - running authdns-update * 15:36 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:36 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->eqsin - cmooney@cumin1003" * 15:34 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327138{{!}}Enable redis lock manager everywhere (T366938)]] (duration: 08m 36s) * 15:33 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->eqsin - cmooney@cumin1003" * 15:30 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:29 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 15:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1147.eqiad.wmnet with OS bookworm * 15:28 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327138{{!}}Enable redis lock manager everywhere (T366938)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:25 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327138{{!}}Enable redis lock manager everywhere (T366938)]] * 15:24 jmm@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host krb1002.eqiad.wmnet * 15:19 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2207.codfw.wmnet with reason: Host crashed * 15:17 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327098{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]], [[gerrit:1327101{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]] (duration: 07m 13s) * 15:12 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 15:12 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1327098{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]], [[gerrit:1327101{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:10 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1327098{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]], [[gerrit:1327101{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]] * 15:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2010.codfw.wmnet with OS trixie * 15:05 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 15:05 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1147.eqiad.wmnet with reason: host reimage * 14:59 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 14:59 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:58 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1147.eqiad.wmnet with reason: host reimage * 14:55 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:55 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:49 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 14:48 cmooney@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host durum1001.eqiad.wmnet * 14:46 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:46 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Delete db2902 ipv6 addr - fceratto@cumin1003" * 14:46 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Delete db2902 ipv6 addr - fceratto@cumin1003" * 14:43 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1147.eqiad.wmnet with OS bookworm * 14:42 tgr@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327094{{!}}SpecialMWOAuthListConsumers: Handle newFromMWUser returning null in addNavigationSubtitle (T435167)]] (duration: 19m 25s) * 14:42 cmooney@cumin1003: START - Cookbook sre.hosts.reboot-single for host durum1001.eqiad.wmnet * 14:42 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 14:41 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 14:38 tgr@deploy1003: tgr: Continuing with deployment * 14:36 tgr@deploy1003: tgr: Backport for [[gerrit:1327094{{!}}SpecialMWOAuthListConsumers: Handle newFromMWUser returning null in addNavigationSubtitle (T435167)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:28 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 14:28 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:28 cmooney@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host durum3005.esams.wmnet * 14:25 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 14:23 cmooney@cumin1003: START - Cookbook sre.hosts.reboot-single for host durum3005.esams.wmnet * 14:23 tgr@deploy1003: Started scap sync-world: Backport for [[gerrit:1327094{{!}}SpecialMWOAuthListConsumers: Handle newFromMWUser returning null in addNavigationSubtitle (T435167)]] * 14:18 gengh@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:17 topranks: disable puppet on hosts running BIRD BGP to test merge of patch to systemd healtchcheck service * 14:17 gengh@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:17 gengh@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:16 gengh@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:16 elukey: upgrade spicerack on cumin1003 and cumin2003 * 14:16 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:15 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314962{{!}}static: add new dir bimi/ for BIMI SVG and PEM file (T311685)]] (duration: 10m 00s) * 14:15 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:11 kharlan@deploy1003: kharlan, sukhe: Continuing with deployment * 14:11 gengh@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:09 gengh@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:09 gengh@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:08 kharlan@deploy1003: kharlan, sukhe: Backport for [[gerrit:1314962{{!}}static: add new dir bimi/ for BIMI SVG and PEM file (T311685)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:06 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 14:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:06 gengh@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:05 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1314962{{!}}static: add new dir bimi/ for BIMI SVG and PEM file (T311685)]] * 14:05 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host krb1002.eqiad.wmnet * 14:05 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:04 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:04 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326342{{!}}Revert^2 "wmf-config/ProductionServices: set URL for urldownloader to service record"]] (duration: 07m 40s) * 13:59 kharlan@deploy1003: kharlan, sukhe: Continuing with deployment * 13:58 kharlan@deploy1003: kharlan, sukhe: Backport for [[gerrit:1326342{{!}}Revert^2 "wmf-config/ProductionServices: set URL for urldownloader to service record"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:57 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host stat1008.eqiad.wmnet with OS bookworm * 13:56 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1326342{{!}}Revert^2 "wmf-config/ProductionServices: set URL for urldownloader to service record"]] * 13:56 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 13:54 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325532{{!}}srwiki: Allow bureaucrats to add and remove event-organizer group (T434748)]] (duration: 14m 56s) * 13:54 swfrench@dns1004: END - running authdns-update * 13:52 swfrench@dns1004: START - running authdns-update * 13:51 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:51 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 13:48 kharlan@deploy1003: kharlan, danielyepezgarces: Continuing with deployment * 13:45 swfrench@cumin2003: conftool action : set/pooled=yes; selector: name=wikikube-worker2330.codfw.wmnet * 13:44 swfrench-wmf: finished etcd-main codfw -> eqiad switchover - [[phab:T435103|T435103]] * 13:44 kharlan@deploy1003: kharlan, danielyepezgarces: Backport for [[gerrit:1325532{{!}}srwiki: Allow bureaucrats to add and remove event-organizer group (T434748)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:44 swfrench@cumin2003: conftool action : set/pooled=no; selector: name=wikikube-worker2330.codfw.wmnet * 13:41 swfrench@dns1004: END - running authdns-update * 13:39 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1325532{{!}}srwiki: Allow bureaucrats to add and remove event-organizer group (T434748)]] * 13:39 swfrench@dns1004: START - running authdns-update * 13:37 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327111{{!}}Special:AbuseReview: Add "no further action needed" review action (T435020)]], [[gerrit:1327110{{!}}AbuseReview: Take the review verdict as a REST path parameter (T435020)]] (duration: 31m 43s) * 13:31 swfrench-wmf: starting etcd-main codfw -> eqiad switchover - [[phab:T435103|T435103]] * 13:28 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:28 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 13:25 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=97) rolling reboot on A:durum and A:durum * 13:24 kharlan@deploy1003: kharlan: Continuing with deployment * 13:24 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1147 * 13:24 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1147 * 13:23 kharlan@deploy1003: kharlan: Backport for [[gerrit:1327111{{!}}Special:AbuseReview: Add "no further action needed" review action (T435020)]], [[gerrit:1327110{{!}}AbuseReview: Take the review verdict as a REST path parameter (T435020)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:16 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:16 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 13:12 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:12 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 13:10 cdobbins@cumin1003: END (ERROR) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=97) Rolling upgrade of ATS on A:cp-codfw and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 13:10 cdobbins@cumin1003: END (ERROR) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=97) Rolling upgrade of ATS on A:cp-eqiad and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 13:06 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1327111{{!}}Special:AbuseReview: Add "no further action needed" review action (T435020)]], [[gerrit:1327110{{!}}AbuseReview: Take the review verdict as a REST path parameter (T435020)]] * 13:06 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-eqiad and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 13:05 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-codfw and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 13:05 cjd91: sudo -i cookbook sre.cdn.roll-upgrade-ats --query 'A:cp-eqiad' --task-id [[phab:T434478|T434478]] --reason '9.2.15 upgrade' * 13:03 swfrench@dns1004: END - running authdns-update * 13:01 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:01 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 13:01 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host ncmonitor1001.eqiad.wmnet * 13:00 swfrench@dns1004: START - running authdns-update * 12:59 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 12:59 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 12:59 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 12:59 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 12:57 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and A:durum * 12:53 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on db2902.codfw.wmnet with reason: Cloning * 12:48 cmooney@dns3003: END - running authdns-update * 12:48 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327096{{!}}Switch to redis lock manager on s4 and s8 (T366938)]] (duration: 09m 11s) * 12:46 cmooney@dns3003: START - running authdns-update * 12:45 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:45 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->eqord cct - cmooney@cumin1003" * 12:45 elukey: move the /v2/releng.* prefix on the Docker Registry to its new s3 backend - [[phab:T432829|T432829]] * 12:43 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 12:42 jelto@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 12:42 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->eqord cct - cmooney@cumin1003" * 12:42 jelto@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 12:41 jelto: update cert-manager to 1.19.6 on wikikube staging-codfw - [[phab:T427402|T427402]] * 12:40 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327096{{!}}Switch to redis lock manager on s4 and s8 (T366938)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:38 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327096{{!}}Switch to redis lock manager on s4 and s8 (T366938)]] * 12:38 blake@deploy1003: Finished scap sync-world: non-build deploy for [[phab:T417800|T417800]] (duration: 03m 52s) * 12:36 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 12:35 blake@deploy1003: Started scap sync-world: non-build deploy for [[phab:T417800|T417800]] * 12:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host krb2002.codfw.wmnet * 11:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host krb2002.codfw.wmnet * 11:49 moritzm: installing kerberos security updates * 11:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on stat1008.eqiad.wmnet with reason: host reimage * 11:44 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on stat1008.eqiad.wmnet with reason: host reimage * 11:31 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327084{{!}}Revert "Migrate database access to virtual domains" (T435305)]] (duration: 11m 02s) * 11:29 gkyziridis@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:29 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:29 kart_: Updated MinT to 2026-06-04-131507-production ([[phab:T321316|T321316]]) * 11:28 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/machinetranslation: apply * 11:28 gkyziridis@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:26 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 11:24 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:24 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:23 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:23 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:23 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/machinetranslation: apply * 11:22 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327084{{!}}Revert "Migrate database access to virtual domains" (T435305)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:21 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/machinetranslation: apply * 11:21 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:21 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:20 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327084{{!}}Revert "Migrate database access to virtual domains" (T435305)]] * 11:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1008.eqiad.wmnet with OS bookworm * 11:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet * 11:17 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/machinetranslation: apply * 11:13 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/machinetranslation: apply * 11:12 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:12 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:10 kartik@deploy1003: helmfile [staging] START helmfile.d/services/machinetranslation: apply * 11:08 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-cron: apply * 11:08 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/mw-cron: apply * 11:08 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply * 11:08 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply * 11:07 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet * 11:07 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host stat1008.eqiad.wmnet with OS bookworm * 11:06 moritzm: upgrading the new trixie URL downloaders to Squid 7.6 [[phab:T427282|T427282]] * 11:01 gkyziridis@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin2002.codfw.wmnet * 10:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin2002.codfw.wmnet * 10:45 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1169.eqiad.wmnet with OS bookworm * 10:42 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1185.eqiad.wmnet with OS bookworm * 10:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1169.eqiad.wmnet with reason: host reimage * 10:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1185.eqiad.wmnet with reason: host reimage * 10:14 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1169.eqiad.wmnet with reason: host reimage * 10:14 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1185.eqiad.wmnet with reason: host reimage * 10:11 jmm@cumin2003: END (PASS) - Cookbook sre.netbox.restart-reboot (exit_code=0) rolling reboot on A:netbox * 10:06 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1008.eqiad.wmnet with OS bookworm * 10:04 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 10:04 mpostoronca@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321579{{!}}Register the mediawiki.wikimedia_antiabuse.content_policy_score stream (T432848)]] (duration: 08m 53s) * 10:03 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 10:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 10:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 10:00 mpostoronca@deploy1003: mpostoronca: Continuing with deployment * 09:59 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1185.eqiad.wmnet with OS bookworm * 09:59 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1169.eqiad.wmnet with OS bookworm * 09:59 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.convert-disks (exit_code=0) for host ms-be1065 * 09:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1065.eqiad.wmnet with OS trixie * 09:58 mpostoronca@deploy1003: mpostoronca: Backport for [[gerrit:1321579{{!}}Register the mediawiki.wikimedia_antiabuse.content_policy_score stream (T432848)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:55 jmm@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netbox.discovery.wmnet. on all recursors * 09:55 mpostoronca@deploy1003: Started scap sync-world: Backport for [[gerrit:1321579{{!}}Register the mediawiki.wikimedia_antiabuse.content_policy_score stream (T432848)]] * 09:55 jmm@cumin2003: START - Cookbook sre.dns.wipe-cache netbox.discovery.wmnet. on all recursors * 09:52 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw2001.wikimedia.org with OS trixie * 09:51 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 09:51 jmm@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netbox.discovery.wmnet. on all recursors * 09:51 jmm@cumin2003: START - Cookbook sre.dns.wipe-cache netbox.discovery.wmnet. on all recursors * 09:51 jmm@cumin2003: START - Cookbook sre.netbox.restart-reboot rolling reboot on A:netbox * 09:46 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 09:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1056.eqiad.wmnet with OS trixie * 09:44 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 09:39 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 09:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 09:36 topranks: make HE transport circuits from magru live * 09:36 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 09:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb1003.eqiad.wmnet * 09:33 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage * 09:31 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb1003.eqiad.wmnet * 09:30 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 09:28 arnaudb@dns1006: END - running authdns-update * 09:27 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage * 09:27 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb2003.codfw.wmnet * 09:26 arnaudb@dns1006: START - running authdns-update * 09:24 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1056.eqiad.wmnet with reason: host reimage * 09:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb2003.codfw.wmnet * 09:20 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.convert-disks (exit_code=0) for host ms-be1068 * 09:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1068.eqiad.wmnet with OS trixie * 09:20 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 09:19 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "cloudvirt1057 - filippo@cumin1003" * 09:19 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "cloudvirt1057 - filippo@cumin1003" * 09:18 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1056.eqiad.wmnet with reason: host reimage * 09:18 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1057.eqiad.wmnet with OS trixie * 09:18 filippo@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 09:18 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 09:15 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 09:14 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 09:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host irc1003.wikimedia.org * 09:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:13 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1065.eqiad.wmnet with OS trixie * 09:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 09:09 moritzm: installing Postgresql security updates * 09:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host irc1003.wikimedia.org * 09:07 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 09:07 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:07 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1173.eqiad.wmnet with OS bookworm * 09:06 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw2001.wikimedia.org with OS trixie * 09:03 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.convert-disks (exit_code=0) for host ms-be1064 * 09:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1064.eqiad.wmnet with OS trixie * 09:03 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1056.eqiad.wmnet with OS trixie * 09:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1057.eqiad.wmnet with reason: host reimage * 09:01 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 08:58 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 08:56 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1057.eqiad.wmnet with reason: host reimage * 08:54 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1208.eqiad.wmnet with OS bookworm * 08:53 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2902.codfw.wmnet with OS trixie * 08:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 08:50 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1174.eqiad.wmnet with OS bookworm * 08:46 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1172.eqiad.wmnet with OS bookworm * 08:45 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1173.eqiad.wmnet with reason: host reimage * 08:41 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 08:40 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1057.eqiad.wmnet with OS trixie * 08:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1057.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 08:39 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1207.eqiad.wmnet with OS bookworm * 08:38 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2902.codfw.wmnet with reason: host reimage * 08:37 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 08:34 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1068.eqiad.wmnet with OS trixie * 08:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1208.eqiad.wmnet with reason: host reimage * 08:31 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1057.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 08:29 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 08:28 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1222.eqiad.wmnet onto db1276.eqiad.wmnet * 08:28 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1222: Pool db1222.eqiad.wmnet in after cloning * 08:28 fceratto@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2902.codfw.wmnet with reason: host reimage * 08:27 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1055.eqiad.wmnet with OS trixie * 08:27 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 08:26 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1174.eqiad.wmnet with reason: host reimage * 08:25 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 08:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-misc2002.codfw.wmnet * 08:23 topranks: reboot pfw1-codfw firewall pair to upgrade JunOS [[phab:T434865|T434865]] * 08:22 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1172.eqiad.wmnet with reason: host reimage * 08:20 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1064.eqiad.wmnet with OS trixie * 08:20 mvernon@cumin2003: START - Cookbook sre.swift.convert-disks for host ms-be1065 * 08:18 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1207.eqiad.wmnet with reason: host reimage * 08:17 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1174.eqiad.wmnet with reason: host reimage * 08:17 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1173.eqiad.wmnet with reason: host reimage * 08:17 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1172.eqiad.wmnet with reason: host reimage * 08:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host mc-misc2002.codfw.wmnet * 08:15 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1208.eqiad.wmnet with reason: host reimage * 08:15 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1207.eqiad.wmnet with reason: host reimage * 08:14 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db2902.codfw.wmnet with OS trixie * 08:14 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.16 refs [[phab:T430835|T430835]] * 08:13 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db2902.codfw.wmnet * 08:13 fceratto@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host db2902.codfw.wmnet with OS trixie * 08:10 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr[1-2]-codfw with reason: upgrade pfw1a-codfw and pfw1b-codfw pair * 08:09 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1055.eqiad.wmnet with reason: host reimage * 08:07 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on pfw1-codfw with reason: upgrade pfw1a-codfw and pfw1b-codfw pair * 08:03 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1055.eqiad.wmnet with reason: host reimage * 08:02 arnaudb@dns1006: END - running authdns-update * 08:02 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1208.eqiad.wmnet with OS bookworm * 08:02 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1207.eqiad.wmnet with OS bookworm * 08:01 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1174.eqiad.wmnet with OS bookworm * 08:01 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1173.eqiad.wmnet with OS bookworm * 08:01 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1172.eqiad.wmnet with OS bookworm * 07:59 arnaudb@dns1006: START - running authdns-update * 07:58 arnaudb@dns1006: START - running authdns-update * 07:48 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1055.eqiad.wmnet with OS trixie * 07:45 moritzm: extend the disk of ldap-rw2001 by 80G [[phab:T331699|T331699]] * 07:42 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1222: Pool db1222.eqiad.wmnet in after cloning * 07:36 mvernon@cumin2003: START - Cookbook sre.swift.convert-disks for host ms-be1068 * 07:35 mvernon@cumin2003: START - Cookbook sre.swift.convert-disks for host ms-be1064 * 07:17 moritzm: installing imagemagick security updates * 07:14 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1277: Pool back * 07:14 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon1003.wikimedia.org * 07:07 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon1003.wikimedia.org * 07:03 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1280: Pool back * 07:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon2002.wikimedia.org * 06:55 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon2002.wikimedia.org * 06:54 moritzm: installing php8.2 security updates * 06:51 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1284: Pool back * 06:49 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1222: Depool db1222.eqiad.wmnet to then clone it to db1276.eqiad.wmnet - marostegui@cumin1003 * 06:49 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1222: Depool db1222.eqiad.wmnet to then clone it to db1276.eqiad.wmnet - marostegui@cumin1003 * 06:49 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1222.eqiad.wmnet onto db1276.eqiad.wmnet * 06:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd1005.eqiad.wmnet * 06:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd1005.eqiad.wmnet * 06:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd1004.eqiad.wmnet * 06:36 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2209: db2209 repool * 06:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd1004.eqiad.wmnet * 06:32 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast3007.wikimedia.org * 06:29 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1277: Pool back * 06:28 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1277 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96190 and previous config saved to /var/cache/conftool/dbconfig/20260819-062815-marostegui.json * 06:26 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast3007.wikimedia.org * 06:22 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul1001.eqiad.wmnet * 06:18 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul1001.eqiad.wmnet * 06:18 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul1003.eqiad.wmnet * 06:18 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1280: Pool back * 06:17 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1284 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96186 and previous config saved to /var/cache/conftool/dbconfig/20260819-061743-marostegui.json * 06:14 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul1003.eqiad.wmnet * 06:14 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul1002.eqiad.wmnet * 06:10 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul1002.eqiad.wmnet * 06:10 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2048.codfw.wmnet * 06:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2048.codfw.wmnet * 06:06 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1284: Pool back * 06:06 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1284 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96184 and previous config saved to /var/cache/conftool/dbconfig/20260819-060621-marostegui.json * 06:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2048.codfw.wmnet * 05:59 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2048.codfw.wmnet * 05:51 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2209: db2209 repool * 03:16 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1186.eqiad.wmnet with OS bookworm * 02:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1186.eqiad.wmnet with reason: host reimage * 02:46 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1186.eqiad.wmnet with reason: host reimage * 02:46 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2207 [[phab:T435270|T435270]]', diff saved to https://phabricator.wikimedia.org/P96181 and previous config saved to /var/cache/conftool/dbconfig/20260819-024627-marostegui.json * 02:44 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2204 to s2 primary [[phab:T435270|T435270]]', diff saved to https://phabricator.wikimedia.org/P96180 and previous config saved to /var/cache/conftool/dbconfig/20260819-024403-marostegui.json * 02:43 marostegui: Starting s2 codfw failover from db2207 to db2204 - [[phab:T435270|T435270]] * 02:39 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2204 with weight 0 [[phab:T435270|T435270]]', diff saved to https://phabricator.wikimedia.org/P96179 and previous config saved to /var/cache/conftool/dbconfig/20260819-023951-marostegui.json * 02:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s2 [[phab:T435270|T435270]] * 02:32 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1186.eqiad.wmnet with OS bookworm * 02:29 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-worker1186.eqiad.wmnet with OS bookworm * 02:18 denisse@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2207: Depooling replica * 02:18 denisse@cumin1003: START - Cookbook sre.mysql.depool depool db2207: Depooling replica * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 48s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-18 == * 23:55 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1264845{{!}}Remove unused/redundant wgMFNoindexPages=true setting (T255458)]] (duration: 09m 42s) * 23:51 krinkle@deploy1003: krinkle: Continuing with deployment * 23:48 krinkle@deploy1003: krinkle: Backport for [[gerrit:1264845{{!}}Remove unused/redundant wgMFNoindexPages=true setting (T255458)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:45 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1264845{{!}}Remove unused/redundant wgMFNoindexPages=true setting (T255458)]] * 23:38 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326944{{!}}Retire filebackend lock manager in favour of the default one (T366938)]] (duration: 08m 55s) * 23:34 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 23:31 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326944{{!}}Retire filebackend lock manager in favour of the default one (T366938)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:29 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326944{{!}}Retire filebackend lock manager in favour of the default one (T366938)]] * 23:27 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1170.eqiad.wmnet with OS bookworm * 23:21 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1205.eqiad.wmnet with OS bookworm * 23:20 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1171.eqiad.wmnet with OS bookworm * 23:15 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1206.eqiad.wmnet with OS bookworm * 23:05 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1170.eqiad.wmnet with reason: host reimage * 23:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1205.eqiad.wmnet with reason: host reimage * 22:57 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1171.eqiad.wmnet with reason: host reimage * 22:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1206.eqiad.wmnet with reason: host reimage * 22:53 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1205.eqiad.wmnet with reason: host reimage * 22:51 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1171.eqiad.wmnet with reason: host reimage * 22:51 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1170.eqiad.wmnet with reason: host reimage * 22:50 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1206.eqiad.wmnet with reason: host reimage * 22:36 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1206.eqiad.wmnet with OS bookworm * 22:35 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1205.eqiad.wmnet with OS bookworm * 22:35 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1186.eqiad.wmnet with OS bookworm * 22:35 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1171.eqiad.wmnet with OS bookworm * 22:35 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1170.eqiad.wmnet with OS bookworm * 22:33 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-worker1194.eqiad.wmnet with OS bookworm * 22:22 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326923{{!}}Enable redis lock manager on s6 (T366938)]] (duration: 11m 52s) * 22:18 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 22:13 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326923{{!}}Enable redis lock manager on s6 (T366938)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:10 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326923{{!}}Enable redis lock manager on s6 (T366938)]] * 22:04 sbassett: Deployed security fix for [[phab:T435234|T435234]] (wmf.16) * 21:54 sbassett: Deployed security fix for [[phab:T435234|T435234]] (wmf.15) * 21:38 caro@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326925{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326926{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326929{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]], [[gerrit:1326928{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]] (duration: 0 * 21:34 caro@deploy1003: caro: Continuing with deployment * 21:33 caro@deploy1003: caro: Backport for [[gerrit:1326925{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326926{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326929{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]], [[gerrit:1326928{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]] synced to the testservers (see h * 21:31 caro@deploy1003: Started scap sync-world: Backport for [[gerrit:1326925{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326926{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326929{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]], [[gerrit:1326928{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]] * 21:24 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1204.eqiad.wmnet with reason: 1204 datanode repair [[phab:T434494|T434494]] * 21:02 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326896{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]], [[gerrit:1326897{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]] (duration: 13m 56s) * 20:58 krinkle@deploy1003: krinkle: Continuing with deployment * 20:50 krinkle@deploy1003: krinkle: Backport for [[gerrit:1326896{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]], [[gerrit:1326897{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:49 ryankemper: `an-launcher1003` terminated process group `666809` (`rest_backfill_phase1.sh`) ~20 mins ago with `sudo kill -TERM -- -666809` after its local spark driver (`--driver-memory 64g`) repeatedly exhausted memory on the 32 GB VM and caused SSH to intermittently flap; host recovered to 27 GB available memory * 20:48 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1326896{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]], [[gerrit:1326897{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]] * 20:35 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 20:33 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 20:31 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 20:31 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326870{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]], [[gerrit:1326871{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]] (duration: 07m 35s) * 20:28 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 20:26 kemayo@deploy1003: kemayo: Continuing with deployment * 20:26 ryankemper: `an-launcher1003` confirmed the host is flapping because of memory thrash. chasing down the source of the thrash * 20:25 kemayo@deploy1003: kemayo: Backport for [[gerrit:1326870{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]], [[gerrit:1326871{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:23 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1326870{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]], [[gerrit:1326871{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]] * 20:23 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 20:20 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 20:14 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1168.eqiad.wmnet with OS bookworm * 20:08 zabe: zabe@deploy1003:~$ mwscript extensions/WikimediaMaintenance/maintenance/fixFileRevisionArchiveNameDrift.php enwiki # [[phab:T428406|T428406]] * 20:08 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1204.eqiad.wmnet with OS bookworm * 20:05 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1167.eqiad.wmnet with OS bookworm * 20:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1166.eqiad.wmnet with OS bookworm * 19:54 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1203.eqiad.wmnet with OS bookworm * 19:53 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326907{{!}}Revert "Disable redis lock manager on testwiki"]] (duration: 11m 05s) * 19:50 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1168.eqiad.wmnet with reason: host reimage * 19:47 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1204.eqiad.wmnet with reason: host reimage * 19:46 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 19:44 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326907{{!}}Revert "Disable redis lock manager on testwiki"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:42 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326907{{!}}Revert "Disable redis lock manager on testwiki"]] * 19:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1167.eqiad.wmnet with reason: host reimage * 19:37 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1166.eqiad.wmnet with reason: host reimage * 19:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1203.eqiad.wmnet with reason: host reimage * 19:32 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1167.eqiad.wmnet with reason: host reimage * 19:32 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1168.eqiad.wmnet with reason: host reimage * 19:32 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1166.eqiad.wmnet with reason: host reimage * 19:31 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1204.eqiad.wmnet with reason: host reimage * 19:31 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1203.eqiad.wmnet with reason: host reimage * 19:26 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324751{{!}}InitialiseSettings: Enable 2FA enforcement on more private wikis (T428103)]], [[gerrit:1326875{{!}}Add banner notifying of upcoming 2FA enforcement (T420792)]] (duration: 31m 46s) * 19:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1204.eqiad.wmnet with OS bookworm * 19:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1203.eqiad.wmnet with OS bookworm * 19:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1168.eqiad.wmnet with OS bookworm * 19:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1167.eqiad.wmnet with OS bookworm * 19:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1166.eqiad.wmnet with OS bookworm * 19:15 denisse: rebooting kafkamon2003.codfw.wmnet - [[phab:T435162|T435162]] * 19:14 denisse: rebooting kafkamon1003.eqiad.wmnet [[phab:T435162|T435162]] * 19:13 reedy@deploy1003: reedy: Continuing with deployment * 19:12 reedy@deploy1003: reedy: Backport for [[gerrit:1324751{{!}}InitialiseSettings: Enable 2FA enforcement on more private wikis (T428103)]], [[gerrit:1326875{{!}}Add banner notifying of upcoming 2FA enforcement (T420792)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:54 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324751{{!}}InitialiseSettings: Enable 2FA enforcement on more private wikis (T428103)]], [[gerrit:1326875{{!}}Add banner notifying of upcoming 2FA enforcement (T420792)]] * 18:50 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 18:44 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.16 refs [[phab:T430835|T430835]] * 18:34 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 18:31 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 18:22 aklapper@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326852{{!}}CategoryTree: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]], [[gerrit:1326853{{!}}CategoryViewer: Allow null $html in the CategoryViewerGenerateLink hook (T435161)]], [[gerrit:1326865{{!}}Flow: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]] (duration: 09m 57s) * 18:18 aklapper@deploy1003: jforrester, aklapper: Continuing with deployment * 18:17 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 18:14 aklapper@deploy1003: jforrester, aklapper: Backport for [[gerrit:1326852{{!}}CategoryTree: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]], [[gerrit:1326853{{!}}CategoryViewer: Allow null $html in the CategoryViewerGenerateLink hook (T435161)]], [[gerrit:1326865{{!}}Flow: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki * 18:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1156.eqiad.wmnet with OS bookworm * 18:12 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 18:12 aklapper@deploy1003: Started scap sync-world: Backport for [[gerrit:1326852{{!}}CategoryTree: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]], [[gerrit:1326853{{!}}CategoryViewer: Allow null $html in the CategoryViewerGenerateLink hook (T435161)]], [[gerrit:1326865{{!}}Flow: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]] * 18:11 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 18:08 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1146.eqiad.wmnet with OS bookworm * 18:07 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1177.eqiad.wmnet with OS bookworm * 18:00 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326380{{!}}Introduce main lock manager service (T366938 T427999)]] (duration: 11m 25s) * 17:58 ladsgroup@deploy1003: ladsgroup: Rolling back deployment * 17:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1202.eqiad.wmnet with OS bookworm * 17:55 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1201.eqiad.wmnet with OS bookworm * 17:50 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326380{{!}}Introduce main lock manager service (T366938 T427999)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:48 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326380{{!}}Introduce main lock manager service (T366938 T427999)]] * 17:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1156.eqiad.wmnet with reason: host reimage * 17:46 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1146.eqiad.wmnet with reason: host reimage * 17:45 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-esams and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 17:42 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 17:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1177.eqiad.wmnet with reason: host reimage * 17:38 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1202.eqiad.wmnet with reason: host reimage * 17:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1201.eqiad.wmnet with reason: host reimage * 17:30 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1156.eqiad.wmnet with reason: host reimage * 17:29 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1177.eqiad.wmnet with reason: host reimage * 17:28 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1146.eqiad.wmnet with reason: host reimage * 17:28 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1202.eqiad.wmnet with reason: host reimage * 17:27 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1201.eqiad.wmnet with reason: host reimage * 17:25 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 17:21 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 17:14 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1202.eqiad.wmnet with OS bookworm * 17:14 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1201.eqiad.wmnet with OS bookworm * 17:14 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1177.eqiad.wmnet with OS bookworm * 17:14 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1156.eqiad.wmnet with OS bookworm * 17:14 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1146.eqiad.wmnet with OS bookworm * 17:09 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326885{{!}}w/deployment-info.php: Handle new file format (T434726)]] (duration: 07m 15s) * 17:08 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 17:05 dancy@deploy1003: dancy: Continuing with deployment * 17:04 dancy@deploy1003: dancy: Backport for [[gerrit:1326885{{!}}w/deployment-info.php: Handle new file format (T434726)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:02 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1326885{{!}}w/deployment-info.php: Handle new file format (T434726)]] * 16:50 dancy@deploy1003: Finished scap sync-world: Testing [[phab:T434726|T434726]] (duration: 06m 40s) * 16:43 dancy@deploy1003: Started scap sync-world: Testing [[phab:T434726|T434726]] * 16:43 dancy@deploy1003: Installation of scap version "4.282.0" completed for 3 hosts * 16:41 dancy@deploy1003: Installing scap version "4.282.0" for 3 host(s) * 16:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1145.eqiad.wmnet with OS bookworm * 16:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1200.eqiad.wmnet with OS bookworm * 16:32 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1199.eqiad.wmnet with OS bookworm * 16:18 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1145.eqiad.wmnet with reason: host reimage * 16:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1200.eqiad.wmnet with reason: host reimage * 16:09 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1199.eqiad.wmnet with reason: host reimage * 16:05 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-esams and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 16:05 cjd91: sudo -i cookbook sre.cdn.roll-upgrade-ats --query 'A:cp-esams' --task-id [[phab:T434478|T434478]] --reason '9.2.15 upgrade' * 16:03 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1200.eqiad.wmnet with reason: host reimage * 16:02 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1145.eqiad.wmnet with reason: host reimage * 16:02 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1199.eqiad.wmnet with reason: host reimage * 15:50 moritzm: installing zip security updates * 15:48 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1200.eqiad.wmnet with OS bookworm * 15:47 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1199.eqiad.wmnet with OS bookworm * 15:47 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1145.eqiad.wmnet with OS bookworm * 15:41 topranks: bounce PIC 0/0 on cr1-magru to set port to 40G * 15:37 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1236.eqiad.wmnet with OS bookworm * 15:29 aikochou@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'ores-legacy' for release 'main' . * 15:26 aikochou@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'ores-legacy' for release 'main' . * 15:20 aikochou@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'ores-legacy' for release 'main' . * 15:14 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2002.codfw.wmnet * 15:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1236.eqiad.wmnet with reason: host reimage * 15:12 moritzm: failover ganeti master in codfw to ganeti2047 * 15:09 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1236.eqiad.wmnet with reason: host reimage * 15:09 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2044.codfw.wmnet * 15:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2002.codfw.wmnet * 15:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2044.codfw.wmnet * 15:04 brennen@deploy1003: Finished deploy [phabricator/deployment@6b9b6ff]: deploy phab1004 for [[phab:T435213|T435213]] (duration: 01m 01s) * 15:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2044.codfw.wmnet * 15:03 brennen@deploy1003: Started deploy [phabricator/deployment@6b9b6ff]: deploy phab1004 for [[phab:T435213|T435213]] * 15:03 brennen@deploy1003: Finished deploy [phabricator/deployment@6b9b6ff]: deploy phab2003 for [[phab:T435213|T435213]] (duration: 00m 57s) * 15:02 brennen@deploy1003: Started deploy [phabricator/deployment@6b9b6ff]: deploy phab2003 for [[phab:T435213|T435213]] * 14:57 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply * 14:57 arnaudb@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on phab2003.codfw.wmnet,phab[1004-1006].eqiad.wmnet with reason: maintenance * 14:56 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2044.codfw.wmnet * 14:55 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply * 14:53 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1236.eqiad.wmnet with OS bookworm * 14:51 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2043.codfw.wmnet * 14:51 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2043.codfw.wmnet * 14:45 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2043.codfw.wmnet * 14:33 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2043.codfw.wmnet * 14:22 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2042.codfw.wmnet * 14:22 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2042.codfw.wmnet * 14:21 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db2902.codfw.wmnet with OS trixie * 14:16 elukey: uploaded spicerack_13.2.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia * 14:16 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2042.codfw.wmnet * 14:04 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2042.codfw.wmnet * 14:04 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2041.codfw.wmnet * 14:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2041.codfw.wmnet * 13:59 phuedx@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: apply * 13:59 phuedx@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-main: apply * 13:59 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326838{{!}}Enable Suggested Investigations on hewiki (T435146)]] (duration: 11m 50s) * 13:58 phuedx@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: apply * 13:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2041.codfw.wmnet * 13:57 phuedx@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-main: apply * 13:57 phuedx@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-main: apply * 13:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard1003.eqiad.wmnet * 13:57 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-main: apply * 13:55 phuedx@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-logging-external: apply * 13:55 phuedx@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-logging-external: apply * 13:54 phuedx@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-logging-external: apply * 13:54 stran@deploy1003: stran: Continuing with deployment * 13:54 phuedx@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-logging-external: apply * 13:54 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-logging-external: apply * 13:54 phuedx@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-logging-external: apply * 13:53 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-logging-external: apply * 13:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard1003.eqiad.wmnet * 13:53 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2041.codfw.wmnet * 13:51 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db2902.codfw.wmnet - fceratto@cumin1003" * 13:51 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db2902.codfw.wmnet - fceratto@cumin1003" * 13:51 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2026.codfw.wmnet * 13:51 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard2003.codfw.wmnet * 13:50 phuedx@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: apply * 13:50 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-eqiad * 13:50 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp1001.eqiad.wmnet * 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp1001.eqiad.wmnet * 13:50 phuedx@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: apply * 13:49 phuedx@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: apply * 13:49 stran@deploy1003: stran: Backport for [[gerrit:1326838{{!}}Enable Suggested Investigations on hewiki (T435146)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:48 phuedx@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: apply * 13:48 phuedx@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics: apply * 13:47 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard2003.codfw.wmnet * 13:47 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics: apply * 13:47 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1326838{{!}}Enable Suggested Investigations on hewiki (T435146)]] * 13:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp1001.eqiad.wmnet * 13:44 cdanis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 13:43 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp1001.eqiad.wmnet * 13:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1315-1327].eqiad.wmnet * 13:43 cdanis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 13:43 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1315-1327].eqiad.wmnet * 13:40 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki1001.eqiad.wmnet * 13:35 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1315-1327].eqiad.wmnet * 13:34 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host rpki1001.eqiad.wmnet * 13:27 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1315-1327].eqiad.wmnet * 13:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1302-1314].eqiad.wmnet * 13:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1302-1314].eqiad.wmnet * 13:21 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326775{{!}}SI: Instrument case update on first edit (T435048)]] (duration: 07m 12s) * 13:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1302-1314].eqiad.wmnet * 13:17 stran@deploy1003: stran: Continuing with deployment * 13:16 stran@deploy1003: stran: Backport for [[gerrit:1326775{{!}}SI: Instrument case update on first edit (T435048)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:14 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1326775{{!}}SI: Instrument case update on first edit (T435048)]] * 13:13 moritzm: installing util-linux security updates * 13:11 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1302-1314].eqiad.wmnet * 13:10 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326232{{!}}prv: Enable parsoid rendering for 5 wikis (T435115)]] (duration: 08m 17s) * 13:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1288-1289,1291-1301].eqiad.wmnet * 13:10 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply * 13:10 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1288-1289,1291-1301].eqiad.wmnet * 13:10 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply * 13:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki2003.codfw.wmnet * 13:06 jgiannelos@deploy1003: jgiannelos: Continuing with deployment * 13:05 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host rpki2003.codfw.wmnet * 13:04 jgiannelos@deploy1003: jgiannelos: Backport for [[gerrit:1326232{{!}}prv: Enable parsoid rendering for 5 wikis (T435115)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1288-1289,1291-1301].eqiad.wmnet * 13:02 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1326232{{!}}prv: Enable parsoid rendering for 5 wikis (T435115)]] * 12:53 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1288-1289,1291-1301].eqiad.wmnet * 12:53 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1273,1275-1287].eqiad.wmnet * 12:53 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1273,1275-1287].eqiad.wmnet * 12:52 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt1002.wikimedia.org * 12:51 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db2902.codfw.wmnet on all recursors * 12:51 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db2902.codfw.wmnet on all recursors * 12:51 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:51 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db2902.codfw.wmnet - fceratto@cumin1003" * 12:51 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db2902.codfw.wmnet - fceratto@cumin1003" * 12:46 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt1002.wikimedia.org * 12:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt2002.wikimedia.org * 12:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1273,1275-1287].eqiad.wmnet * 12:42 dhinus: repooled clouddb1032 that was currently <nowiki>{</nowiki>"weight": 0, "pooled": "inactive"<nowiki>}</nowiki> for both s4 and s6 * 12:41 dhinus: also depooled clouddb1017 (forgot it in the previous list) * 12:41 fnegri@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet * 12:40 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 12:40 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db2902.codfw.wmnet * 12:40 fnegri@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032.eqiad.wmnet * 12:40 fnegri@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032 * 12:39 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1017.eqiad.wmnet * 12:39 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt2002.wikimedia.org * 12:38 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1020.eqiad.wmnet * 12:38 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1018.eqiad.wmnet * 12:38 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet * 12:37 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1014.eqiad.wmnet * 12:37 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1013.eqiad.wmnet * 12:37 dhinus: depool again clouddb10[13,14,16,18,20] that were repooled by the cookbook sre.mysql.multiinstance_reboot * 12:36 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1273,1275-1287].eqiad.wmnet * 12:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1248-1261].eqiad.wmnet * 12:36 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1248-1261].eqiad.wmnet * 12:35 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2026.codfw.wmnet * 12:34 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 12:34 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 12:34 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 12:34 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 12:29 lucaswerkmeister-wmde@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 12:28 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2026.codfw.wmnet * 12:28 lucaswerkmeister-wmde@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 12:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1248-1261].eqiad.wmnet * 12:27 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.addnode (exit_code=0) for new host ganeti2046.codfw.wmnet to cluster codfw and group A * 12:26 moritzm: readded ganeti2046 to the codfw cluster following firmware update and reimage [[phab:T434681|T434681]] * 12:23 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply * 12:23 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply * 12:23 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply * 12:22 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply * 12:22 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply * 12:22 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply * 12:21 jmm@cumin2003: START - Cookbook sre.ganeti.addnode for new host ganeti2046.codfw.wmnet to cluster codfw and group A * 12:21 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1169.eqiad.wmnet onto db1283.eqiad.wmnet * 12:21 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1169: Pool db1169.eqiad.wmnet in after cloning * 12:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1248-1261].eqiad.wmnet * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1149-1153,1158,1240-1247].eqiad.wmnet * 12:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1149-1153,1158,1240-1247].eqiad.wmnet * 12:11 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1149-1153,1158,1240-1247].eqiad.wmnet * 12:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1168.eqiad.wmnet onto db1282.eqiad.wmnet * 12:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1168: Pool db1168.eqiad.wmnet in after cloning * 12:06 jmm@cumin2003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-test-eqiad * 12:05 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2026.codfw.wmnet * 12:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1149-1153,1158,1240-1247].eqiad.wmnet * 12:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1128-1134,1142-1148].eqiad.wmnet * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2046.codfw.wmnet * 12:01 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1128-1134,1142-1148].eqiad.wmnet * 11:57 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2025.codfw.wmnet * 11:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2025.codfw.wmnet * 11:54 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2046.codfw.wmnet * 11:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1128-1134,1142-1148].eqiad.wmnet * 11:50 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2025.codfw.wmnet * 11:46 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1128-1134,1142-1148].eqiad.wmnet * 11:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1114-1127].eqiad.wmnet * 11:45 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1114-1127].eqiad.wmnet * 11:39 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db2901.codfw.wmnet * 11:39 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db2901.codfw.wmnet * 11:36 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2025.codfw.wmnet * 11:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1114-1127].eqiad.wmnet * 11:35 fceratto@cumin1003: END (ERROR) - Cookbook sre.ganeti.makevm (exit_code=93) for new host db1901.eqiad.wmnet * 11:35 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1169: Pool db1169.eqiad.wmnet in after cloning * 11:35 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 11:30 jmm@cumin2003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-test-eqiad * 11:27 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1114-1127].eqiad.wmnet * 11:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1076-1081,1084-1087,1093-1095,1113].eqiad.wmnet * 11:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1076-1081,1084-1087,1093-1095,1113].eqiad.wmnet * 11:25 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1168: Pool db1168.eqiad.wmnet in after cloning * 11:24 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1194.eqiad.wmnet with OS bookworm * 11:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1181.eqiad.wmnet with OS bookworm * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2050.codfw.wmnet * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2050.codfw.wmnet * 11:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1076-1081,1084-1087,1093-1095,1113].eqiad.wmnet * 11:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2050.codfw.wmnet * 11:14 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2004.codfw.wmnet * 11:10 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2050.codfw.wmnet * 11:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1076-1081,1084-1087,1093-1095,1113].eqiad.wmnet * 11:09 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1286: Pool back * 11:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1045-1050,1056-1057,1064-1066,1073-1075].eqiad.wmnet * 11:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1045-1050,1056-1057,1064-1066,1073-1075].eqiad.wmnet * 11:08 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2004.codfw.wmnet * 11:07 moritzm: installing PHP 8.4 security updates * 11:06 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on an-worker1194.eqiad.wmnet with reason: host reimage * 11:06 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1194.eqiad.wmnet with reason: host reimage * 11:05 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2049.codfw.wmnet * 11:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2049.codfw.wmnet * 11:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid1003.eqiad.wmnet * 11:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1045-1050,1056-1057,1064-1066,1073-1075].eqiad.wmnet * 10:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1181.eqiad.wmnet with reason: host reimage * 10:59 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2049.codfw.wmnet * 10:59 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid1003.eqiad.wmnet * 10:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid2003.codfw.wmnet * 10:54 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1181.eqiad.wmnet with reason: host reimage * 10:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid2003.codfw.wmnet * 10:51 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1901.eqiad.wmnet on all recursors * 10:51 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1901.eqiad.wmnet on all recursors * 10:51 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:51 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:51 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:50 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2049.codfw.wmnet * 10:50 blake@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:50 blake@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:49 blake@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:48 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1045-1050,1056-1057,1064-1066,1073-1075].eqiad.wmnet * 10:48 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1044].eqiad.wmnet * 10:48 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1044].eqiad.wmnet * 10:47 blake@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:46 blake@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:46 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2047.codfw.wmnet * 10:46 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:46 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2047.codfw.wmnet * 10:46 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1901.eqiad.wmnet * 10:46 blake@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:44 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1168: Depool db1168.eqiad.wmnet to then clone it to db1282.eqiad.wmnet - marostegui@cumin1003 * 10:44 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1168: Depool db1168.eqiad.wmnet to then clone it to db1282.eqiad.wmnet - marostegui@cumin1003 * 10:44 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1168.eqiad.wmnet onto db1282.eqiad.wmnet * 10:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2047.codfw.wmnet * 10:40 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1044].eqiad.wmnet * 10:37 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2047.codfw.wmnet * 10:37 fceratto@cumin1003: END (ERROR) - Cookbook sre.ganeti.makevm (exit_code=93) for new host db1901.eqiad.wmnet * 10:36 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 10:34 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:33 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2032.codfw.wmnet * 10:33 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2032.codfw.wmnet * 10:32 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1044].eqiad.wmnet * 10:32 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-eqiad * 10:27 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2032.codfw.wmnet * 10:24 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1286: Pool back * 10:24 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1286 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96163 and previous config saved to /var/cache/conftool/dbconfig/20260818-102431-marostegui.json * 10:22 moritzm: installing Django security updates * 10:20 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul2002.codfw.wmnet * 10:20 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul2001.codfw.wmnet * 10:17 blake@deploy1003: Finished scap sync-world: no-build deployment for [[phab:T417800|T417800]] (duration: 04m 40s) * 10:16 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul2002.codfw.wmnet * 10:16 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul2001.codfw.wmnet * 10:15 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2032.codfw.wmnet * 10:14 blake@deploy1003: Started scap sync-world: no-build deployment for [[phab:T417800|T417800]] * 10:12 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul2003.codfw.wmnet * 10:12 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy2001.codfw.wmnet * 10:12 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy3001.esams.wmnet * 10:12 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy1001.eqiad.wmnet * 10:08 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2031.codfw.wmnet * 10:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2031.codfw.wmnet * 10:08 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul2003.codfw.wmnet * 10:08 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy2001.codfw.wmnet * 10:08 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy3001.esams.wmnet * 10:08 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy1001.eqiad.wmnet * 10:07 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy1002.eqiad.wmnet * 10:05 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy2002.codfw.wmnet * 10:05 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy3002.esams.wmnet * 10:04 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy1002.eqiad.wmnet * 10:03 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy4003.ulsfo.wmnet * 10:03 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy5003.eqsin.wmnet * 10:02 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2031.codfw.wmnet * 10:02 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1169: Depool db1169.eqiad.wmnet to then clone it to db1283.eqiad.wmnet - marostegui@cumin1003 * 10:01 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy2002.codfw.wmnet * 10:01 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy4004.ulsfo.wmnet * 10:01 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy3002.esams.wmnet * 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1169: Depool db1169.eqiad.wmnet to then clone it to db1283.eqiad.wmnet - marostegui@cumin1003 * 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1169.eqiad.wmnet onto db1283.eqiad.wmnet * 10:01 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy5004.eqsin.wmnet * 09:59 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy4003.ulsfo.wmnet * 09:59 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy4004.ulsfo.wmnet * 09:59 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy5003.eqsin.wmnet * 09:59 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy7001.magru.wmnet * 09:59 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy5004.eqsin.wmnet * 09:58 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy7002.magru.wmnet * 09:57 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2031.codfw.wmnet * 09:57 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy6001.drmrs.wmnet * 09:57 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy6002.drmrs.wmnet * 09:54 filippo@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for 10 hosts * 09:54 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast4006.wikimedia.org * 09:53 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy6001.drmrs.wmnet * 09:53 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy6002.drmrs.wmnet * 09:53 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host releases1003.eqiad.wmnet * 09:53 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy7001.magru.wmnet * 09:53 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host people2004.codfw.wmnet * 09:52 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host people1005.eqiad.wmnet * 09:52 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy7002.magru.wmnet * 09:50 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host releases2003.codfw.wmnet * 09:49 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host releases2003.codfw.wmnet * 09:49 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host releases1003.eqiad.wmnet * 09:49 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host people2004.codfw.wmnet * 09:48 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host people1005.eqiad.wmnet * 09:46 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2040.codfw.wmnet * 09:46 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2040.codfw.wmnet * 09:43 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1901.eqiad.wmnet on all recursors * 09:42 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1901.eqiad.wmnet on all recursors * 09:42 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:42 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:42 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2040.codfw.wmnet * 09:40 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp2005.wikimedia.org * 09:36 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp2005.wikimedia.org * 09:31 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:31 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1901.eqiad.wmnet * 09:31 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1901.eqiad.wmnet * 09:31 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:31 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1901.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 09:31 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1901.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 09:29 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2040.codfw.wmnet * 09:29 slyngshede@dns1004: END - running authdns-update * 09:28 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2039.codfw.wmnet * 09:28 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2039.codfw.wmnet * 09:27 slyngshede@dns1004: START - running authdns-update * 09:26 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test2005.wikimedia.org * 09:22 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2039.codfw.wmnet * 09:22 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test2005.wikimedia.org * 09:22 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp1005.wikimedia.org * 09:21 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:19 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2039.codfw.wmnet * 09:19 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2038.codfw.wmnet * 09:18 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1287: Pool back * 09:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2038.codfw.wmnet * 09:18 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp1005.wikimedia.org * 09:18 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test1005.wikimedia.org * 09:17 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1901.eqiad.wmnet * 09:15 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db1901.eqiad.wmnet * 09:15 fceratto@cumin1003: END (ERROR) - Cookbook sre.dns.netbox (exit_code=97) * 09:14 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test1005.wikimedia.org * 09:13 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2038.codfw.wmnet * 09:13 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:13 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1901.eqiad.wmnet * 09:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast4006.wikimedia.org * 09:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1288: Pool back * 09:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast5005.wikimedia.org * 09:03 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2038.codfw.wmnet * 08:57 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2037.codfw.wmnet * 08:57 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast5005.wikimedia.org * 08:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2037.codfw.wmnet * 08:56 filippo@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 10 hosts * 08:52 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2037.codfw.wmnet * 08:51 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1289: Pool back * 08:35 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudidp2001-dev.codfw.wmnet * 08:34 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1002-dev.eqiad.wmnet * 08:33 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1287: Pool back * 08:33 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1287 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96150 and previous config saved to /var/cache/conftool/dbconfig/20260818-083311-marostegui.json * 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1001-dev.eqiad.wmnet * 08:31 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudidp2001-dev.codfw.wmnet * 08:30 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1002-dev.eqiad.wmnet * 08:30 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2037.codfw.wmnet * 08:29 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1001-dev.eqiad.wmnet * 08:28 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2036.codfw.wmnet * 08:28 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2036.codfw.wmnet * 08:25 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1288: Pool back * 08:23 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 08:23 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1181.eqiad.wmnet with OS bookworm * 08:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2036.codfw.wmnet * 08:22 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1288 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96148 and previous config saved to /var/cache/conftool/dbconfig/20260818-082234-marostegui.json * 08:20 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2036.codfw.wmnet * 08:18 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2035.codfw.wmnet * 08:18 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudvirt1057.eqiad.wmnet with OS trixie * 08:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2035.codfw.wmnet * 08:13 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2035.codfw.wmnet * 08:06 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2035.codfw.wmnet * 08:05 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1289: Pool back * 08:05 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1289 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96145 and previous config saved to /var/cache/conftool/dbconfig/20260818-080531-marostegui.json * 07:54 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 07:51 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326439{{!}}Fix "mathjax_ignore" handling around forcemathmode attribute (T434686)]] (duration: 13m 15s) * 07:48 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 07:47 krinkle@deploy1003: krinkle: Continuing with deployment * 07:44 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ganeti2046.codfw.wmnet with OS bookworm * 07:40 krinkle@deploy1003: krinkle: Backport for [[gerrit:1326439{{!}}Fix "mathjax_ignore" handling around forcemathmode attribute (T434686)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:39 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 07:38 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1326439{{!}}Fix "mathjax_ignore" handling around forcemathmode attribute (T434686)]] * 07:34 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 07:31 samwilson@deploy1003: Finished scap sync-world: Backport for [[gerrit:701016{{!}}InitialiseSettings and -labs: Remove redundant feature flag $wgWikisourceEnableOcr (T285311)]] (duration: 07m 47s) * 07:30 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1048.eqiad.wmnet * 07:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1048.eqiad.wmnet * 07:29 XioNoX: add gnmic 0.47.0 to bookworm and trixie reprepro * 07:28 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ganeti2046.codfw.wmnet with reason: host reimage * 07:27 samwilson@deploy1003: samwilson: Continuing with deployment * 07:25 samwilson@deploy1003: samwilson: Backport for [[gerrit:701016{{!}}InitialiseSettings and -labs: Remove redundant feature flag $wgWikisourceEnableOcr (T285311)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:25 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 07:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1048.eqiad.wmnet * 07:24 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ganeti2046.codfw.wmnet with reason: host reimage * 07:23 samwilson@deploy1003: Started scap sync-world: Backport for [[gerrit:701016{{!}}InitialiseSettings and -labs: Remove redundant feature flag $wgWikisourceEnableOcr (T285311)]] * 07:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1198.eqiad.wmnet with OS bookworm * 07:18 samwilson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326450{{!}}InitialiseSettings.php: Enable Bulk OCR on pawikisource (T434648)]] (duration: 12m 17s) * 07:17 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 07:11 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1048.eqiad.wmnet * 07:11 samwilson@deploy1003: samwilson: Continuing with deployment * 07:11 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ganeti2046.codfw.wmnet with OS bookworm * 07:10 samwilson@deploy1003: samwilson: Backport for [[gerrit:1326450{{!}}InitialiseSettings.php: Enable Bulk OCR on pawikisource (T434648)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1057.eqiad.wmnet with OS trixie * 07:05 samwilson@deploy1003: Started scap sync-world: Backport for [[gerrit:1326450{{!}}InitialiseSettings.php: Enable Bulk OCR on pawikisource (T434648)]] * 07:02 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1198.eqiad.wmnet with reason: host reimage * 07:01 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1056.eqiad.wmnet with OS trixie * 07:01 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1056.eqiad.wmnet with OS trixie * 07:00 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1056.eqiad.wmnet with OS trixie * 07:00 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1056.eqiad.wmnet with OS trixie * 06:59 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudvirt1055.eqiad.wmnet with OS trixie * 06:58 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1198.eqiad.wmnet with reason: host reimage * 06:53 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1055.eqiad.wmnet with OS trixie * 06:52 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudvirt1054.eqiad.wmnet with OS trixie * 06:44 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1194.eqiad.wmnet with OS bookworm * 06:44 XioNoX: upgrade eqsin gnmic to 0.47.0 * 06:43 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1198.eqiad.wmnet with OS bookworm * 06:41 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1054.eqiad.wmnet with OS trixie * 06:09 arnaudb@cumin1003: END (PASS) - Cookbook sre.gerrit.restart-gerrit (exit_code=0) Restarting Gerrit on gerrit2002 * 06:06 arnaudb@cumin1003: START - Cookbook sre.gerrit.restart-gerrit Restarting Gerrit on gerrit2002 * 06:06 arnaudb@cumin1003: END (PASS) - Cookbook sre.gerrit.restart-gerrit (exit_code=0) Restarting Gerrit on gerrit1003 * 06:04 arnaudb@cumin1003: START - Cookbook sre.gerrit.restart-gerrit Restarting Gerrit on gerrit1003 * 06:02 arnaudb@cumin1003: END (PASS) - Cookbook sre.gerrit.restart-gerrit (exit_code=0) Restarting Gerrit on gerrit2003 * 06:00 arnaudb@cumin1003: START - Cookbook sre.gerrit.restart-gerrit Restarting Gerrit on gerrit2003 * 05:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1155.eqiad.wmnet with OS bookworm * 05:38 arnaudb: updating prometheusBearerToken on gerrit * 05:28 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1155.eqiad.wmnet with reason: host reimage * 05:23 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1155.eqiad.wmnet with reason: host reimage * 05:06 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1155.eqiad.wmnet with OS bookworm * 04:57 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1144.eqiad.wmnet with OS bookworm * 04:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1144.eqiad.wmnet with reason: host reimage * 04:29 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1144.eqiad.wmnet with reason: host reimage * 04:14 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1144.eqiad.wmnet with OS bookworm * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.13 (duration: 02m 23s) * 03:45 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1197.eqiad.wmnet with OS bookworm * 03:41 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1165.eqiad.wmnet with OS bookworm * 03:38 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.16 refs [[phab:T430835|T430835]] (duration: 34m 43s) * 03:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1164.eqiad.wmnet with OS bookworm * 03:35 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1196.eqiad.wmnet with OS bookworm * 03:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1163.eqiad.wmnet with OS bookworm * 03:22 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1197.eqiad.wmnet with reason: host reimage * 03:18 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1165.eqiad.wmnet with reason: host reimage * 03:15 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1196.eqiad.wmnet with reason: host reimage * 03:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1164.eqiad.wmnet with reason: host reimage * 03:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1163.eqiad.wmnet with reason: host reimage * 03:05 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1165.eqiad.wmnet with reason: host reimage * 03:04 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1164.eqiad.wmnet with reason: host reimage * 03:04 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1197.eqiad.wmnet with reason: host reimage * 03:04 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1196.eqiad.wmnet with reason: host reimage * 03:04 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1163.eqiad.wmnet with reason: host reimage * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.16 refs [[phab:T430835|T430835]] * 02:50 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1197.eqiad.wmnet with OS bookworm * 02:49 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1196.eqiad.wmnet with OS bookworm * 02:49 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1165.eqiad.wmnet with OS bookworm * 02:49 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1164.eqiad.wmnet with OS bookworm * 02:48 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1163.eqiad.wmnet with OS bookworm * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 46s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:15 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-codfw: Set storage compatability to NONE — [[phab:T433026|T433026]] - eevans@cumin1003 * 00:38 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-codfw: Set storage compatability to NONE — [[phab:T433026|T433026]] - eevans@cumin1003 * 00:11 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326373{{!}}PersonalDashboard: add newly renamed *ReviewChangesMlModel setting (T422148)]] (duration: 07m 07s) * 00:09 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-eqiad: Set storage compatability to NONE — [[phab:T433026|T433026]] - eevans@cumin1003 * 00:07 musikanimal@deploy1003: musikanimal: Continuing with deployment * 00:06 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1326373{{!}}PersonalDashboard: add newly renamed *ReviewChangesMlModel setting (T422148)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:04 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1326373{{!}}PersonalDashboard: add newly renamed *ReviewChangesMlModel setting (T422148)]] == 2026-08-17 == * 23:30 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-eqiad: Set storage compatability to NONE — [[phab:T433026|T433026]] - eevans@cumin1003 * 23:05 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-codfw: Set storage compatability to UPGRADING — [[phab:T433026|T433026]] - eevans@cumin1003 * 22:28 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-codfw: Set storage compatability to UPGRADING — [[phab:T433026|T433026]] - eevans@cumin1003 * 21:46 logmsgbot: jforrester Deployed security patch for [[phab:T435085|T435085]] * 21:39 swfrench@deploy1003: mwscript-k8s job started: purgeList.php # [[phab:T432412|T432412]] * 21:37 maryum: Undeploy security fix for [[phab:T433020|T433020]] * 21:23 maryum: Deployed security fix for [[phab:T433020|T433020]] * 21:21 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 21:21 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 21:14 maryum: Deployed security fix for [[phab:T434967|T434967]] * 20:58 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-eqiad: Set storage compatability to UPGRADING — [[phab:T433026|T433026]] - eevans@cumin1003 * 20:40 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324320{{!}}[itwiki/slwiki/tgwiki] Remove temporary Wikipedia 25 logos permanently (already reverted) (T414265 T414320 T415307)]] (duration: 06m 55s) * 20:40 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 20:39 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 20:36 cjming@deploy1003: cjming, superpes: Continuing with deployment * 20:35 cjming@deploy1003: cjming, superpes: Backport for [[gerrit:1324320{{!}}[itwiki/slwiki/tgwiki] Remove temporary Wikipedia 25 logos permanently (already reverted) (T414265 T414320 T415307)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:33 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1324320{{!}}[itwiki/slwiki/tgwiki] Remove temporary Wikipedia 25 logos permanently (already reverted) (T414265 T414320 T415307)]] * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ttmserver-test: apply * 20:30 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326361{{!}}Remove escaped paths in app site association file (T432412)]] (duration: 13m 28s) * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ttmserver-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-toolhub-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-toolhub-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-toolhub-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-toolhub-test: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-toolhub: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-toolhub: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-toolhub: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-toolhub: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-test: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-test: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 20:27 inflatador: bking@deploy1003 `charlie --services_dir dse-k8s-services -s opensearch-* -e dse-k8s-* apply` [[phab:T435125|T435125]] * 20:27 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-apifeatureusage-test: apply * 20:27 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-apifeatureusage-test: apply * 20:27 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-apifeatureusage-test: apply * 20:27 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-apifeatureusage-test: apply * 20:27 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-apifeatureusage: apply * 20:26 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-apifeatureusage: apply * 20:26 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-apifeatureusage: apply * 20:26 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-apifeatureusage: apply * 20:26 cjming@deploy1003: cjming, tsev: Continuing with deployment * 20:24 inflatador: bking@deploy1003 `charlie --services_dir dse-k8s-services -s opensearch-* -e dse-k8s-*` [[phab:T435125|T435125]] * 20:21 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-eqiad: Set storage compatability to UPGRADING — [[phab:T433026|T433026]] - eevans@cumin1003 * 20:19 cjming@deploy1003: cjming, tsev: Backport for [[gerrit:1326361{{!}}Remove escaped paths in app site association file (T432412)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:17 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1326361{{!}}Remove escaped paths in app site association file (T432412)]] * 20:15 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326077{{!}}InstrumentConstructiveEdits: anchor all runs to the nearest `interval` (T431493)]] (duration: 06m 25s) * 20:11 cjming@deploy1003: cjming: Continuing with deployment * 20:10 cjming@deploy1003: cjming: Backport for [[gerrit:1326077{{!}}InstrumentConstructiveEdits: anchor all runs to the nearest `interval` (T431493)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:08 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1326077{{!}}InstrumentConstructiveEdits: anchor all runs to the nearest `interval` (T431493)]] * 20:06 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1160.eqiad.wmnet with OS bookworm * 20:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1162.eqiad.wmnet with OS bookworm * 19:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1184.eqiad.wmnet with OS bookworm * 19:55 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1161.eqiad.wmnet with OS bookworm * 19:49 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1195.eqiad.wmnet with OS bookworm * 19:49 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-test: apply * 19:49 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-test: apply * 19:44 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1160.eqiad.wmnet with reason: host reimage * 19:41 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-codfw: Upgrade to Java 17 — [[phab:T433026|T433026]] - eevans@cumin1003 * 19:39 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1162.eqiad.wmnet with reason: host reimage * 19:36 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-eqsin and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 19:36 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-test: apply * 19:36 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1184.eqiad.wmnet with reason: host reimage * 19:32 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1161.eqiad.wmnet with reason: host reimage * 19:29 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1195.eqiad.wmnet with reason: host reimage * 19:26 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1161.eqiad.wmnet with reason: host reimage * 19:26 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1160.eqiad.wmnet with reason: host reimage * 19:26 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1184.eqiad.wmnet with reason: host reimage * 19:26 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1162.eqiad.wmnet with reason: host reimage * 19:25 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1195.eqiad.wmnet with reason: host reimage * 19:19 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-test: apply * 19:11 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1195.eqiad.wmnet with OS bookworm * 19:10 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1194.eqiad.wmnet with OS bookworm * 19:10 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1184.eqiad.wmnet with OS bookworm * 19:10 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1162.eqiad.wmnet with OS bookworm * 19:10 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1161.eqiad.wmnet with OS bookworm * 19:10 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1160.eqiad.wmnet with OS bookworm * 19:10 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 19:09 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 19:03 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-codfw: Upgrade to Java 17 — [[phab:T433026|T433026]] - eevans@cumin1003 * 18:58 dancy@deploy1003: Finished scap sync-world: testing [[phab:T375514|T375514]] (duration: 03m 13s) * 18:55 dancy@deploy1003: Started scap sync-world: testing [[phab:T375514|T375514]] * 18:55 dwisehaupt@dns1006: END - running authdns-update * 18:54 dancy@deploy1003: Installation of scap version "4.281.1" completed for 3 hosts * 18:54 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2009.codfw.wmnet * 18:53 dwisehaupt@dns1006: START - running authdns-update * 18:52 dancy@deploy1003: Installing scap version "4.281.1" for 3 host(s) * 18:47 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2009.codfw.wmnet * 18:41 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2008.codfw.wmnet * 18:34 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2008.codfw.wmnet * 18:30 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2007.codfw.wmnet * 18:23 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2007.codfw.wmnet * 18:16 dwisehaupt@dns1005: END - running authdns-update * 18:14 dwisehaupt@dns1005: START - running authdns-update * 18:04 swfrench@deploy1003: Finished scap sync-world: Deploy "Point Test Wiki to new docroot" - [[phab:T432412|T432412]] (duration: 26m 03s) * 18:00 swfrench@deploy1003: swfrench: Continuing with deployment * 17:51 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1057.eqiad.wmnet with OS trixie * 17:47 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1180.eqiad.wmnet with OS bookworm * 17:39 swfrench@deploy1003: swfrench: Deploy "Point Test Wiki to new docroot" - [[phab:T432412|T432412]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:39 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-eqsin and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 17:38 swfrench@deploy1003: Started scap sync-world: Deploy "Point Test Wiki to new docroot" - [[phab:T432412|T432412]] * 17:36 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1159.eqiad.wmnet with OS bookworm * 17:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1158.eqiad.wmnet with OS bookworm * 17:30 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-eqiad: Upgrade to Java 17 — [[phab:T433026|T433026]] - eevans@cumin1003 * 17:30 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-ulsfo or A:cp-drmrs and A:cp - 9.2.15 upgrade ([[phab:T434620|T434620]]) * 17:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1157.eqiad.wmnet with OS bookworm * 17:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1180.eqiad.wmnet with reason: host reimage * 17:22 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 17:21 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1193.eqiad.wmnet with OS bookworm * 17:21 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 17:21 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 17:20 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 17:20 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 17:20 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1183.eqiad.wmnet with OS bookworm * 17:18 swfrench@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 17:18 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 17:17 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1180.eqiad.wmnet with reason: host reimage * 17:17 swfrench@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 17:17 swfrench@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 17:16 swfrench@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 17:16 swfrench@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 17:15 swfrench@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 17:15 swfrench@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 17:15 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1192.eqiad.wmnet with OS bookworm * 17:14 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1159.eqiad.wmnet with reason: host reimage * 17:14 swfrench@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 17:13 swfrench@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 17:12 swfrench@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 17:09 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1158.eqiad.wmnet with reason: host reimage * 17:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1157.eqiad.wmnet with reason: host reimage * 17:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1180 * 17:02 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1180 * 17:01 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica-eqsin and A:liberica * 17:01 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1193.eqiad.wmnet with reason: host reimage * 16:57 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1183.eqiad.wmnet with reason: host reimage * 16:56 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1054.eqiad.wmnet with OS trixie * 16:55 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1056.eqiad.wmnet with OS trixie * 16:54 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1193.eqiad.wmnet with reason: host reimage * 16:54 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1192.eqiad.wmnet with reason: host reimage * 16:51 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-eqiad: Upgrade to Java 17 — [[phab:T433026|T433026]] - eevans@cumin1003 * 16:50 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1180 * 16:50 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1180.eqiad.wmnet 17.36.64.10.in-addr.arpa 7.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:50 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1180.eqiad.wmnet 17.36.64.10.in-addr.arpa 7.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:50 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:50 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1180 - btullis@cumin1003" * 16:50 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1180 - btullis@cumin1003" * 16:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1158.eqiad.wmnet with reason: host reimage * 16:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1157.eqiad.wmnet with reason: host reimage * 16:49 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica-eqsin and A:liberica * 16:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1183.eqiad.wmnet with reason: host reimage * 16:48 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1159.eqiad.wmnet with reason: host reimage * 16:47 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1192.eqiad.wmnet with reason: host reimage * 16:46 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns4004.wikimedia.org * 16:46 sukhe@dns1004: END - running authdns-update * 16:44 sukhe@dns1004: START - running authdns-update * 16:44 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns4004.wikimedia.org,service=authdns-update * 16:43 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns4004.wikimedia.org with OS trixie * 16:39 btullis@cumin1003: START - Cookbook sre.dns.netbox * 16:39 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1180 * 16:39 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1193.eqiad.wmnet with OS bookworm * 16:39 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1180.eqiad.wmnet with OS bookworm * 16:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1192.eqiad.wmnet with OS bookworm * 16:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1183.eqiad.wmnet with OS bookworm * 16:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1159.eqiad.wmnet with OS bookworm * 16:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1158.eqiad.wmnet with OS bookworm * 16:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1157.eqiad.wmnet with OS bookworm * 16:31 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1057.eqiad.wmnet with OS trixie * 16:30 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1057.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 16:29 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica-drmrs and A:liberica * 16:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1154.eqiad.wmnet with OS bookworm * 16:23 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1057.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 16:23 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1055.eqiad.wmnet with OS trixie * 16:22 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1057 * 16:22 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1057 * 16:21 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:21 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1057] - vriley@cumin1003" * 16:21 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1057] - vriley@cumin1003" * 16:19 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica-drmrs and A:liberica * 16:17 vriley@cumin1003: START - Cookbook sre.dns.netbox * 16:16 phuedx@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics-external: apply * 16:15 phuedx@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics-external: apply * 16:13 phuedx@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics-external: apply * 16:12 phuedx@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics-external: apply * 16:11 phuedx@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics-external: apply * 16:09 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics-external: apply * 16:09 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1176.eqiad.wmnet with OS bookworm * 16:07 btullis@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 16:06 btullis@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 16:05 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1191.eqiad.wmnet with OS bookworm * 16:03 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1154.eqiad.wmnet with reason: host reimage * 16:03 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching aqs[2002-2012].codfw.wmnet,aqs[1017-1027].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433026|T433026]] - eevans@cumin1003 * 15:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1190.eqiad.wmnet with OS bookworm * 15:59 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1154.eqiad.wmnet with reason: host reimage * 15:55 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-debug: apply * 15:55 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-debug: apply * 15:55 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-debug: apply * 15:55 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/mw-debug: apply * 15:53 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns4004.wikimedia.org with reason: host reimage * 15:50 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns4004.wikimedia.org with reason: host reimage * 15:46 btullis@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 15:46 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1176.eqiad.wmnet with reason: host reimage * 15:45 btullis@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 15:43 btullis@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 15:42 btullis@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 15:42 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1191.eqiad.wmnet with reason: host reimage * 15:39 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1190.eqiad.wmnet with reason: host reimage * 15:36 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1054.eqiad.wmnet with OS trixie * 15:35 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:35 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1056.eqiad.wmnet with OS trixie * 15:35 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:34 moritzm: failover Ganeti master in eqiad to ganeti1046 * 15:34 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1176.eqiad.wmnet with reason: host reimage * 15:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1191.eqiad.wmnet with reason: host reimage * 15:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1190.eqiad.wmnet with reason: host reimage * 15:31 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns2006.wikimedia.org * 15:31 sukhe@dns1004: END - running authdns-update * 15:29 sukhe@dns1004: START - running authdns-update * 15:29 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns2006.wikimedia.org,service=authdns-update * 15:29 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns2006.wikimedia.org * 15:29 sukhe@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns2006.wikimedia.org * 15:26 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica-ulsfo and A:liberica * 15:26 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns1006.wikimedia.org * 15:26 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:25 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns2006.wikimedia.org with OS trixie * 15:25 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1056 * 15:25 sukhe@dns1004: END - running authdns-update * 15:23 sukhe@dns1004: START - running authdns-update * 15:23 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns1006.wikimedia.org,service=authdns-update * 15:23 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns1006.wikimedia.org * 15:23 sukhe@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns1006.wikimedia.org * 15:20 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1056 * 15:20 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:20 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1056~] - vriley@cumin1003" * 15:19 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1056~] - vriley@cumin1003" * 15:19 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns1006.wikimedia.org with OS trixie * 15:19 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns4004.wikimedia.org with OS trixie * 15:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1190.eqiad.wmnet with OS bookworm * 15:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1191.eqiad.wmnet with OS bookworm * 15:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1176.eqiad.wmnet with OS bookworm * 15:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1154.eqiad.wmnet with OS bookworm * 15:16 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica-ulsfo and A:liberica * 15:13 vriley@cumin1003: START - Cookbook sre.dns.netbox * 15:13 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host dns4004.wikimedia.org with OS trixie * 15:12 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1188.eqiad.wmnet with OS bookworm * 15:09 taavi@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318203{{!}}Undeploy WP25EasterEggs (II) (T418134)]] (duration: 08m 56s) * 15:06 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica-magru and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 15:06 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1051.eqiad.wmnet with OS trixie * 15:06 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 15:05 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 15:05 taavi@deploy1003: taavi: Continuing with deployment * 15:04 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-ncredir (exit_code=0) rolling reboot on A:ncredir and A:ncredir * 15:04 taavi@deploy1003: taavi: Backport for [[gerrit:1318203{{!}}Undeploy WP25EasterEggs (II) (T418134)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:03 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1055.eqiad.wmnet with OS trixie * 15:02 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:00 taavi@deploy1003: Started scap sync-world: Backport for [[gerrit:1318203{{!}}Undeploy WP25EasterEggs (II) (T418134)]] * 14:59 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1047.eqiad.wmnet * 14:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1047.eqiad.wmnet * 14:58 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy (exit_code=0) rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 14:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-codfw * 14:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp2001.codfw.wmnet * 14:57 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp2001.codfw.wmnet * 14:57 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica-magru and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 14:57 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:56 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1055 * 14:56 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1055 * 14:55 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns2006.wikimedia.org with reason: host reimage * 14:55 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:55 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1055] - vriley@cumin1003" * 14:55 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1055] - vriley@cumin1003" * 14:54 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1047.eqiad.wmnet * 14:52 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1188.eqiad.wmnet with reason: host reimage * 14:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp2001.codfw.wmnet * 14:51 vriley@cumin1003: START - Cookbook sre.dns.netbox * 14:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp2001.codfw.wmnet * 14:50 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2016.codfw.wmnet * 14:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2016.codfw.wmnet * 14:50 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1189.eqiad.wmnet with OS bookworm * 14:48 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1188.eqiad.wmnet with reason: host reimage * 14:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1051.eqiad.wmnet with reason: host reimage * 14:46 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1182.eqiad.wmnet with OS bookworm * 14:45 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief2002.codfw.wmnet * 14:44 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns1006.wikimedia.org with reason: host reimage * 14:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2016.codfw.wmnet * 14:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1143.eqiad.wmnet with OS bookworm * 14:43 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2016.codfw.wmnet * 14:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2318-2331].codfw.wmnet * 14:43 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2318-2331].codfw.wmnet * 14:42 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching aqs[2002-2012].codfw.wmnet,aqs[1017-1027].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433026|T433026]] - eevans@cumin1003 * 14:41 cgoubert@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326306{{!}}Add placeholder $wmgRedisLockPassword (T366938 T427999)]] (duration: 06m 56s) * 14:41 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief2002.codfw.wmnet * 14:40 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief1002.eqiad.wmnet * 14:39 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1051.eqiad.wmnet with reason: host reimage * 14:38 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns2006.wikimedia.org with reason: host reimage * 14:37 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns1006.wikimedia.org with reason: host reimage * 14:37 cgoubert@deploy1003: cgoubert: Continuing with deployment * 14:36 cgoubert@deploy1003: cgoubert: Backport for [[gerrit:1326306{{!}}Add placeholder $wmgRedisLockPassword (T366938 T427999)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:36 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief1002.eqiad.wmnet * 14:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2318-2331].codfw.wmnet * 14:34 cgoubert@deploy1003: Started scap sync-world: Backport for [[gerrit:1326306{{!}}Add placeholder $wmgRedisLockPassword (T366938 T427999)]] * 14:29 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2318-2331].codfw.wmnet * 14:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2304-2317].codfw.wmnet * 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2304-2317].codfw.wmnet * 14:28 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1054 * 14:27 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1054 * 14:27 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:27 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1054] - vriley@cumin1003" * 14:27 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1054] - vriley@cumin1003" * 14:26 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1189.eqiad.wmnet with reason: host reimage * 14:26 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test2001.codfw.wmnet * 14:25 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test1001.eqiad.wmnet * 14:25 claime: Deploying wmgRedisLockPassword - [[phab:T366938|T366938]] [[phab:T427999|T427999]] * 14:24 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1051.eqiad.wmnet with OS trixie * 14:23 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 14:22 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1182.eqiad.wmnet with reason: host reimage * 14:22 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1051.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:22 vriley@cumin1003: START - Cookbook sre.dns.netbox * 14:22 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test2001.codfw.wmnet * 14:21 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test1001.eqiad.wmnet * 14:21 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 14:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2304-2317].codfw.wmnet * 14:19 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns4004.wikimedia.org with OS trixie * 14:19 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns2006.wikimedia.org with OS trixie * 14:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1143.eqiad.wmnet with reason: host reimage * 14:19 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns1006.wikimedia.org with OS trixie * 14:17 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1189.eqiad.wmnet with reason: host reimage * 14:15 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1182.eqiad.wmnet with reason: host reimage * 14:14 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1143.eqiad.wmnet with reason: host reimage * 14:13 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1051.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:12 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2304-2317].codfw.wmnet * 14:12 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2290-2303].codfw.wmnet * 14:12 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1049.eqiad.wmnet with OS trixie * 14:12 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2290-2303].codfw.wmnet * 14:12 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1051 * 14:11 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1051 * 14:11 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1170.eqiad.wmnet onto db1284.eqiad.wmnet * 14:11 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1170: Pool db1170.eqiad.wmnet in after cloning * 14:09 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 14:09 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:09 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1051] - vriley@cumin1003" * 14:09 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1051] - vriley@cumin1003" * 14:06 klausman@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:04 vriley@cumin1003: START - Cookbook sre.dns.netbox * 14:04 klausman@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2290-2303].codfw.wmnet * 14:02 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1189.eqiad.wmnet with OS bookworm * 14:02 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1188.eqiad.wmnet with OS bookworm * 14:01 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 14:00 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1182.eqiad.wmnet with OS bookworm * 14:00 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1143.eqiad.wmnet with OS bookworm * 13:59 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 13:57 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply * 13:57 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply * 13:56 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply * 13:56 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-ulsfo or A:cp-drmrs and A:cp - 9.2.15 upgrade ([[phab:T434620|T434620]]) * 13:56 cjd91: sudo -i cookbook sre.cdn.roll-upgrade-ats --query 'A:cp-ulsfo or A:cp-drmrs' --task-id [[phab:T434620|T434620]] --reason '9.2.15 upgrade' * 13:56 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply * 13:55 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply * 13:55 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply * 13:54 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 13:54 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 13:54 phuedx@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:53 phuedx@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics-external: apply * 13:52 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1049.eqiad.wmnet with reason: host reimage * 13:51 phuedx@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2290-2303].codfw.wmnet * 13:50 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1047.eqiad.wmnet * 13:50 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2276-2289].codfw.wmnet * 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2276-2289].codfw.wmnet * 13:49 phuedx@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics-external: apply * 13:49 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1046.eqiad.wmnet * 13:49 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1049.eqiad.wmnet with reason: host reimage * 13:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1046.eqiad.wmnet * 13:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-ncredir rolling reboot on A:ncredir and A:ncredir * 13:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 13:46 phuedx@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:44 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics-external: apply * 13:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1046.eqiad.wmnet * 13:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2276-2289].codfw.wmnet * 13:36 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1046.eqiad.wmnet * 13:36 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1045.eqiad.wmnet * 13:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1045.eqiad.wmnet * 13:34 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2276-2289].codfw.wmnet * 13:34 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1049.eqiad.wmnet with OS trixie * 13:34 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2262-2275].codfw.wmnet * 13:34 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2262-2275].codfw.wmnet * 13:32 Lucas_WMDE: UTC afternoon backport+config window done * 13:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1045.eqiad.wmnet * 13:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2262-2275].codfw.wmnet * 13:26 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1049.eqiad.wmnet with OS trixie * 13:26 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1049.eqiad.wmnet with OS trixie * 13:26 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1170: Pool db1170.eqiad.wmnet in after cloning * 13:23 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1049.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 13:23 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1045.eqiad.wmnet * 13:21 atsukoito: manually done sudo -i docker-registryctl --debug delete-tags 'docker-registry.discovery.wmnet/repos/data-engineering/airflow-dags:airflow-3.3.0-py3.11-2026-08-17-*' to remove incorrect tags * 13:19 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311141{{!}}viwiki: Set `noindex,nofollow` for User and User talk (T432311)]] (duration: 11m 11s) * 13:15 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2262-2275].codfw.wmnet * 13:14 lucaswerkmeister-wmde@deploy1003: ndkdd, lucaswerkmeister-wmde: Continuing with deployment * 13:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2248-2261].codfw.wmnet * 13:14 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2248-2261].codfw.wmnet * 13:13 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1049.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 13:12 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1049 * 13:11 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1049 * 13:10 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:10 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1049] - vriley@cumin1003" * 13:10 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1049] - vriley@cumin1003" * 13:10 lucaswerkmeister-wmde@deploy1003: ndkdd, lucaswerkmeister-wmde: Backport for [[gerrit:1311141{{!}}viwiki: Set `noindex,nofollow` for User and User talk (T432311)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1037.eqiad.wmnet * 13:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1037.eqiad.wmnet * 13:08 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1311141{{!}}viwiki: Set `noindex,nofollow` for User and User talk (T432311)]] * 13:06 vriley@cumin1003: START - Cookbook sre.dns.netbox * 13:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2248-2261].codfw.wmnet * 13:00 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1037.eqiad.wmnet * 12:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2248-2261].codfw.wmnet * 12:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2204-2215,2242-2243].codfw.wmnet * 12:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2204-2215,2242-2243].codfw.wmnet * 12:49 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324832{{!}}Migrate $wgFlaggedRevsTags from flaggedrevs.php to ext-FlaggedRevs.php]] (duration: 14m 02s) * 12:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint2001.codfw.wmnet * 12:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2204-2215,2242-2243].codfw.wmnet * 12:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint2001.codfw.wmnet * 12:41 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1037.eqiad.wmnet * 12:40 ladsgroup@deploy1003: Rolling back deployment * 12:37 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1324832{{!}}Migrate $wgFlaggedRevsTags from flaggedrevs.php to ext-FlaggedRevs.php]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:37 seanleong-wmde: Finished populateSitesTable for [bolwiki] ([[[phab:T429955|T429955]]]) * 12:35 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1324832{{!}}Migrate $wgFlaggedRevsTags from flaggedrevs.php to ext-FlaggedRevs.php]] * 12:35 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2204-2215,2242-2243].codfw.wmnet * 12:34 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2190-2203].codfw.wmnet * 12:34 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2190-2203].codfw.wmnet * 12:32 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1028.eqiad.wmnet * 12:32 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1028.eqiad.wmnet * 12:31 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint1001.eqiad.wmnet * 12:30 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1170: Depool db1170.eqiad.wmnet to then clone it to db1284.eqiad.wmnet - marostegui@cumin1003 * 12:28 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint1001.eqiad.wmnet * 12:28 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1170: Depool db1170.eqiad.wmnet to then clone it to db1284.eqiad.wmnet - marostegui@cumin1003 * 12:27 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1170.eqiad.wmnet onto db1284.eqiad.wmnet * 12:27 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db2901.codfw.wmnet * 12:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2190-2203].codfw.wmnet * 12:27 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 12:26 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1028.eqiad.wmnet * 12:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2190-2203].codfw.wmnet * 12:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2172-2179,2184-2189].codfw.wmnet * 12:18 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2172-2179,2184-2189].codfw.wmnet * 12:13 seanleong-wmde@deploy1003: mwscript-k8s job started: foreachwikiindblist wikidataclient extensions/Wikibase/lib/maintenance/populateSitesTable.php --force-protocol https # [[phab:T429955|T429955]] * 12:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2172-2179,2184-2189].codfw.wmnet * 12:09 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1028.eqiad.wmnet * 12:03 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 12:02 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db2901.codfw.wmnet * 12:02 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db2901.codfw.wmnet * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1027.eqiad.wmnet * 12:02 fceratto@cumin1003: END (ERROR) - Cookbook sre.dns.netbox (exit_code=97) * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1027.eqiad.wmnet * 12:02 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 12:02 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db2901.codfw.wmnet * 12:01 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2172-2179,2184-2189].codfw.wmnet * 12:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2158-2171].codfw.wmnet * 12:01 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2158-2171].codfw.wmnet * 11:58 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db1902.eqiad.wmnet * 11:58 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 11:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1027.eqiad.wmnet * 11:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2158-2171].codfw.wmnet * 11:51 jayme: updated calico to v3.30.7 on wikikube eqiad - [[phab:T427400|T427400]] * 11:50 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1027.eqiad.wmnet * 11:45 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2158-2171].codfw.wmnet * 11:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2144-2157].codfw.wmnet * 11:44 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2144-2157].codfw.wmnet * 11:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1058.eqiad.wmnet * 11:43 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1058.eqiad.wmnet * 11:43 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 11:43 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 11:43 marostegui@cumin1003: Removing db1153 from zarcillo [[phab:T434638|T434638]] * 11:42 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1153.eqiad.wmnet * 11:42 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:42 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1153.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 11:42 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1172.eqiad.wmnet onto db1286.eqiad.wmnet * 11:42 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1172: Pool db1172.eqiad.wmnet in after cloning * 11:42 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1153.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 11:41 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1902.eqiad.wmnet * 11:41 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=97) for new host db2901.codfw.wmnet * 11:41 fceratto@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host db2901.codfw.wmnet with OS trixie * 11:38 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'. * 11:38 marostegui@cumin1003: START - Cookbook sre.dns.netbox * 11:37 marostegui@dns1004: END - running authdns-update * 11:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1058.eqiad.wmnet * 11:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2144-2157].codfw.wmnet * 11:35 marostegui@dns1004: START - running authdns-update * 11:32 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1153.eqiad.wmnet * 11:32 marostegui@cumin1003: START - Cookbook sre.mysql.decommission * 11:28 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326249{{!}}ImagePage: move TOC element below file link (T332644)]] (duration: 09m 56s) * 11:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2144-2157].codfw.wmnet * 11:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2130-2143].codfw.wmnet * 11:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2130-2143].codfw.wmnet * 11:26 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1058.eqiad.wmnet * 11:23 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 11:22 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'. * 11:22 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'. * 11:22 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'. * 11:22 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326249{{!}}ImagePage: move TOC element below file link (T332644)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1057.eqiad.wmnet * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1057.eqiad.wmnet * 11:21 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 11:20 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 11:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2130-2143].codfw.wmnet * 11:19 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326249{{!}}ImagePage: move TOC element below file link (T332644)]] * 11:18 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply * 11:17 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply * 11:17 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply * 11:16 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 11:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1057.eqiad.wmnet * 11:15 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 11:14 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 11:13 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 11:13 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db2901.codfw.wmnet with OS trixie * 11:12 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db2901.codfw.wmnet - fceratto@cumin1003" * 11:12 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db2901.codfw.wmnet - fceratto@cumin1003" * 11:12 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db2901.codfw.wmnet on all recursors * 11:12 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db2901.codfw.wmnet on all recursors * 11:12 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:12 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db2901.codfw.wmnet - fceratto@cumin1003" * 11:12 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2130-2143].codfw.wmnet * 11:11 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2107-2115,2124-2129].codfw.wmnet * 11:11 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2107-2115,2124-2129].codfw.wmnet * 11:11 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 11:11 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 11:11 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply * 11:11 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 11:10 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 11:06 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db2901.codfw.wmnet - fceratto@cumin1003" * 11:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2107-2115,2124-2129].codfw.wmnet * 10:57 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1172: Pool db1172.eqiad.wmnet in after cloning * 10:54 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2107-2115,2124-2129].codfw.wmnet * 10:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2078,2087-2095,2102-2106].codfw.wmnet * 10:54 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2078,2087-2095,2102-2106].codfw.wmnet * 10:47 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply * 10:46 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply * 10:46 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply * 10:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2078,2087-2095,2102-2106].codfw.wmnet * 10:45 blake@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply * 10:38 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1057.eqiad.wmnet * 10:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1056.eqiad.wmnet * 10:38 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1056.eqiad.wmnet * 10:37 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2078,2087-2095,2102-2106].codfw.wmnet * 10:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2061-2062,2064-2065,2067-2077].codfw.wmnet * 10:36 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2061-2062,2064-2065,2067-2077].codfw.wmnet * 10:32 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1056.eqiad.wmnet * 10:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2061-2062,2064-2065,2067-2077].codfw.wmnet * 10:25 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:25 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db2901.codfw.wmnet * 10:25 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db1902.eqiad.wmnet * 10:25 fceratto@cumin1003: END (ERROR) - Cookbook sre.dns.netbox (exit_code=97) * 10:24 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:24 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1902.eqiad.wmnet * 10:24 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=97) for new host db1902.eqiad.wmnet * 10:24 fceratto@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host db1902.eqiad.wmnet with OS trixie * 10:24 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=93) for new host db1903.eqiad.wmnet * 10:24 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 10:20 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1056.eqiad.wmnet * 10:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2061-2062,2064-2065,2067-2077].codfw.wmnet * 10:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2038-2039,2041-2042,2044,2046,2049-2051,2055-2060].codfw.wmnet * 10:18 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2038-2039,2041-2042,2044,2046,2049-2051,2055-2060].codfw.wmnet * 10:15 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1055.eqiad.wmnet * 10:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1055.eqiad.wmnet * 10:13 Amir1: mwscript-k8s --dblist=all -- purgeUserOptions.php --login-age 5 uls-preferences ([[phab:T406724|T406724]]) * 10:11 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:10 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1903.eqiad.wmnet on all recursors * 10:10 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1903.eqiad.wmnet on all recursors * 10:10 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2038-2039,2041-2042,2044,2046,2049-2051,2055-2060].codfw.wmnet * 10:10 moritzm: installing unzip security updates * 10:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1055.eqiad.wmnet * 10:08 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=97) for new host db1901.eqiad.wmnet * 10:08 fceratto@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host db1901.eqiad.wmnet with OS trixie * 10:08 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:08 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 10:08 fceratto@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1903.eqiad.wmnet - fceratto@cumin1003" * 10:07 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db1902.eqiad.wmnet with OS trixie * 10:07 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1902.eqiad.wmnet - fceratto@cumin1003" * 10:07 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1902.eqiad.wmnet - fceratto@cumin1003" * 10:04 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply * 10:04 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326227{{!}}Enable desktop/native lazy loading everywhere (T148047)]] (duration: 07m 13s) * 10:03 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1902.eqiad.wmnet on all recursors * 10:03 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1902.eqiad.wmnet on all recursors * 10:03 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:03 blake@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply * 10:01 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1903.eqiad.wmnet - fceratto@cumin1003" * 10:01 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2038-2039,2041-2042,2044,2046,2049-2051,2055-2060].codfw.wmnet * 10:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2002,2005-2006,2011-2015,2017-2018,2033-2037].codfw.wmnet * 10:01 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2002,2005-2006,2011-2015,2017-2018,2033-2037].codfw.wmnet * 10:00 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:00 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 09:59 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 09:59 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1055.eqiad.wmnet * 09:58 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326227{{!}}Enable desktop/native lazy loading everywhere (T148047)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:56 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326227{{!}}Enable desktop/native lazy loading everywhere (T148047)]] * 09:53 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2002,2005-2006,2011-2015,2017-2018,2033-2037].codfw.wmnet * 09:50 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:48 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1903.eqiad.wmnet * 09:48 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:48 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1902.eqiad.wmnet * 09:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 09:44 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2002,2005-2006,2011-2015,2017-2018,2033-2037].codfw.wmnet * 09:43 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-codfw * 09:43 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1174.eqiad.wmnet onto db1288.eqiad.wmnet * 09:42 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1174: Pool db1174.eqiad.wmnet in after cloning * 09:40 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1172: Depool db1172.eqiad.wmnet to then clone it to db1286.eqiad.wmnet - marostegui@cumin1003 * 09:39 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1172: Depool db1172.eqiad.wmnet to then clone it to db1286.eqiad.wmnet - marostegui@cumin1003 * 09:39 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1172.eqiad.wmnet onto db1286.eqiad.wmnet * 09:33 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1175.eqiad.wmnet onto db1289.eqiad.wmnet * 09:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1175: Pool db1175.eqiad.wmnet in after cloning * 09:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1201.eqiad.wmnet onto db1287.eqiad.wmnet * 09:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1201: Pool db1201.eqiad.wmnet in after cloning * 09:28 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db1901.eqiad.wmnet with OS trixie * 09:27 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:27 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:27 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1901.eqiad.wmnet on all recursors * 09:27 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1901.eqiad.wmnet on all recursors * 09:26 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:26 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:26 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:15 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:15 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1901.eqiad.wmnet * 09:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw1001.wikimedia.org with OS trixie * 08:57 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1174: Pool db1174.eqiad.wmnet in after cloning * 08:54 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1054.eqiad.wmnet * 08:54 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1054.eqiad.wmnet * 08:48 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1054.eqiad.wmnet * 08:47 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1175: Pool db1175.eqiad.wmnet in after cloning * 08:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 08:46 marostegui@cumin1003: Removing db1152 from zarcillo [[phab:T434480|T434480]] * 08:46 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1152.eqiad.wmnet * 08:46 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:46 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1152.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 08:46 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1201: Pool db1201.eqiad.wmnet in after cloning * 08:46 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1152.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 08:46 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1054.eqiad.wmnet * 08:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1053.eqiad.wmnet * 08:43 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1053.eqiad.wmnet * 08:42 marostegui@cumin1003: START - Cookbook sre.dns.netbox * 08:38 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage * 08:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1053.eqiad.wmnet * 08:36 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1152.eqiad.wmnet * 08:36 marostegui@cumin1003: START - Cookbook sre.mysql.decommission * 08:35 phuedx: UTC morning backport window done * 08:35 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1053.eqiad.wmnet * 08:34 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1035.eqiad.wmnet * 08:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1035.eqiad.wmnet * 08:34 phuedx@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324725{{!}}EventStreamConfig: Mark product_metrics.web_base and .web_base_with_ip as Test Kitchen streams (T429898 T430322)]], [[gerrit:1313923{{!}}EventStreamConfig: Remove unused web_ui_scroll* streams (T415370)]], [[gerrit:1325546{{!}}EventStreamConfig: Remove Watchlist click stream (T434790)]] (duration: 12m 42s) * 08:33 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1281: Pool back * 08:32 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage * 08:31 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on db2209.codfw.wmnet with reason: Maintenance * 08:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2209: Maintenance needed * 08:30 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2209: Maintenance needed * 08:26 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1035.eqiad.wmnet * 08:26 phuedx@deploy1003: bearloga, phuedx: Continuing with deployment * 08:23 phuedx@deploy1003: bearloga, phuedx: Backport for [[gerrit:1324725{{!}}EventStreamConfig: Mark product_metrics.web_base and .web_base_with_ip as Test Kitchen streams (T429898 T430322)]], [[gerrit:1313923{{!}}EventStreamConfig: Remove unused web_ui_scroll* streams (T415370)]], [[gerrit:1325546{{!}}EventStreamConfig: Remove Watchlist click stream (T434790)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug * 08:21 phuedx@deploy1003: Started scap sync-world: Backport for [[gerrit:1324725{{!}}EventStreamConfig: Mark product_metrics.web_base and .web_base_with_ip as Test Kitchen streams (T429898 T430322)]], [[gerrit:1313923{{!}}EventStreamConfig: Remove unused web_ui_scroll* streams (T415370)]], [[gerrit:1325546{{!}}EventStreamConfig: Remove Watchlist click stream (T434790)]] * 08:19 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw1001.wikimedia.org with OS trixie * 08:18 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1035.eqiad.wmnet * 08:17 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1032.eqiad.wmnet * 08:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1032.eqiad.wmnet * 08:16 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1279: Pool back * 08:15 phuedx@deploy1003: Finished scap sync-world: Backport for [[gerrit:1216721{{!}}viwikivoyage: enable relatedarticle and pop-up (T405724)]] (duration: 39m 12s) * 08:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1032.eqiad.wmnet * 08:09 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1032.eqiad.wmnet * 08:08 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1031.eqiad.wmnet * 08:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1031.eqiad.wmnet * 08:03 godog: switch production to use dumps-nfs.w.o - [[phab:T432212|T432212]] * 08:02 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1031.eqiad.wmnet * 08:02 phuedx@deploy1003: nvdtn19, phuedx: Continuing with deployment * 08:00 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1201: Depool db1201.eqiad.wmnet to then clone it to db1287.eqiad.wmnet - marostegui@cumin1003 * 08:00 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1201: Depool db1201.eqiad.wmnet to then clone it to db1287.eqiad.wmnet - marostegui@cumin1003 * 08:00 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1201.eqiad.wmnet onto db1287.eqiad.wmnet * 08:00 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1057.eqiad.wmnet * 08:00 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:00 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1057.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:59 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1057.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:59 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1275: Pool back * 07:55 filippo@cumin1003: START - Cookbook sre.dns.netbox * 07:55 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1031.eqiad.wmnet * 07:52 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1030.eqiad.wmnet * 07:52 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1030.eqiad.wmnet * 07:52 phuedx@deploy1003: nvdtn19, phuedx: Backport for [[gerrit:1216721{{!}}viwikivoyage: enable relatedarticle and pop-up (T405724)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:51 tappof: bump space for prometheus k8s-aux in eqiad * 07:50 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1057.eqiad.wmnet * 07:48 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1281: Pool back * 07:47 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1281 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96108 and previous config saved to /var/cache/conftool/dbconfig/20260817-074749-marostegui.json * 07:46 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1030.eqiad.wmnet * 07:42 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1030.eqiad.wmnet * 07:41 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1174: Depool db1174.eqiad.wmnet to then clone it to db1288.eqiad.wmnet - marostegui@cumin1003 * 07:41 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1174: Depool db1174.eqiad.wmnet to then clone it to db1288.eqiad.wmnet - marostegui@cumin1003 * 07:41 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1174.eqiad.wmnet onto db1288.eqiad.wmnet * 07:40 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1029.eqiad.wmnet * 07:40 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm2001.wikimedia.org * 07:40 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1029.eqiad.wmnet * 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1056.eqiad.wmnet * 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1056.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:38 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1056.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:36 phuedx@deploy1003: Started scap sync-world: Backport for [[gerrit:1216721{{!}}viwikivoyage: enable relatedarticle and pop-up (T405724)]] * 07:36 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm2001.wikimedia.org * 07:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1279 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96104 and previous config saved to /var/cache/conftool/dbconfig/20260817-073542-marostegui.json * 07:34 filippo@cumin1003: START - Cookbook sre.dns.netbox * 07:34 slyngshede@dns1004: END - running authdns-update * 07:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1029.eqiad.wmnet * 07:33 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm-test1001.wikimedia.org * 07:32 slyngshede@dns1004: START - running authdns-update * 07:31 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1029.eqiad.wmnet * 07:31 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1279: Pool back * 07:30 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1279 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96102 and previous config saved to /var/cache/conftool/dbconfig/20260817-073038-marostegui.json * 07:29 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm-test1001.wikimedia.org * 07:29 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm1001.wikimedia.org * 07:28 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1056.eqiad.wmnet * 07:28 moritzm: extend the disk of ldap-rw1001 by 80G [[phab:T331699|T331699]] * 07:28 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1055.eqiad.wmnet * 07:28 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:28 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1055.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:27 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1055.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:26 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1044.eqiad.wmnet * 07:26 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1044.eqiad.wmnet * 07:25 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm1001.wikimedia.org * 07:24 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1175: Depool db1175.eqiad.wmnet to then clone it to db1289.eqiad.wmnet - marostegui@cumin1003 * 07:24 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1175: Depool db1175.eqiad.wmnet to then clone it to db1289.eqiad.wmnet - marostegui@cumin1003 * 07:24 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1175.eqiad.wmnet onto db1289.eqiad.wmnet * 07:22 filippo@cumin1003: START - Cookbook sre.dns.netbox * 07:20 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1044.eqiad.wmnet * 07:16 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1055.eqiad.wmnet * 07:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1054.eqiad.wmnet * 07:15 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:15 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1054.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:15 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1044.eqiad.wmnet * 07:15 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1054.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:13 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1275: Pool back * 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1275 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96098 and previous config saved to /var/cache/conftool/dbconfig/20260817-071225-marostegui.json * 07:10 filippo@cumin1003: START - Cookbook sre.dns.netbox * 07:05 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1043.eqiad.wmnet * 07:05 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1054.eqiad.wmnet * 07:05 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1051.eqiad.wmnet * 07:05 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:05 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1051.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:05 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1043.eqiad.wmnet * 07:04 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1051.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 06:59 filippo@cumin1003: START - Cookbook sre.dns.netbox * 06:59 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin1001.eqiad.wmnet * 06:59 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1043.eqiad.wmnet * 06:59 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin2001.codfw.wmnet * 06:55 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin2001.codfw.wmnet * 06:55 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1051.eqiad.wmnet * 06:54 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1049.eqiad.wmnet * 06:54 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:54 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1049.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 06:54 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1049.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 06:54 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin1001.eqiad.wmnet * 06:53 moritzm: installing apr-util security updates * 06:52 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1043.eqiad.wmnet * 06:49 filippo@cumin1003: START - Cookbook sre.dns.netbox * 06:41 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1049.eqiad.wmnet * 06:13 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit2003.wikimedia.org * 06:13 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet * 06:07 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet * 06:06 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit2003.wikimedia.org * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 47s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-16 == * 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 01m 03s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-15 == * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 41s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-14 == * 15:38 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-staging-master-eqiad * 15:38 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster1005.eqiad.wmnet * 15:38 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster1005.eqiad.wmnet * 15:35 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sretest2009.codfw.wmnet * 15:33 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster1005.eqiad.wmnet * 15:33 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster1005.eqiad.wmnet * 15:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster1004.eqiad.wmnet * 15:32 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster1004.eqiad.wmnet * 15:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host sretest2009.codfw.wmnet * 15:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster1004.eqiad.wmnet * 15:27 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster1004.eqiad.wmnet * 15:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster1003.eqiad.wmnet * 15:27 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster1003.eqiad.wmnet * 15:24 dancy@deploy1003: Finished scap sync-world: testing (duration: 03m 23s) * 15:22 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster1003.eqiad.wmnet * 15:22 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster1003.eqiad.wmnet * 15:22 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-staging-master-eqiad * 15:20 dancy@deploy1003: Started scap sync-world: testing * 15:20 dancy@deploy1003: Installation of scap version "4.280.2" completed for 3 hosts * 15:18 dancy@deploy1003: Installing scap version "4.280.2" for 3 host(s) * 15:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sretest2006.codfw.wmnet * 14:54 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host sretest2006.codfw.wmnet * 14:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sretest2003.codfw.wmnet * 14:39 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host sretest2003.codfw.wmnet * 13:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt-staging2001.codfw.wmnet * 13:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt-staging2001.codfw.wmnet * 13:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-staging-master-codfw * 13:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster2005.codfw.wmnet * 13:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster2005.codfw.wmnet * 13:05 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox-dev2003.codfw.wmnet * 13:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster2005.codfw.wmnet * 13:04 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster2005.codfw.wmnet * 13:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster2004.codfw.wmnet * 13:04 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster2004.codfw.wmnet * 13:01 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netbox-dev2003.codfw.wmnet * 12:59 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster2004.codfw.wmnet * 12:59 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster2004.codfw.wmnet * 12:59 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster2003.codfw.wmnet * 12:59 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster2003.codfw.wmnet * 12:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster2003.codfw.wmnet * 12:54 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster2003.codfw.wmnet * 12:54 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-staging-master-codfw * 12:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-staging-worker-eqiad * 12:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage1006.eqiad.wmnet * 12:52 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage1006.eqiad.wmnet * 12:46 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage1006.eqiad.wmnet * 12:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw1001.wikimedia.org with OS trixie * 12:45 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage1006.eqiad.wmnet * 12:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage1005.eqiad.wmnet * 12:45 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage1005.eqiad.wmnet * 12:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage1005.eqiad.wmnet * 12:36 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/ratelimit: apply * 12:35 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/ratelimit: apply * 12:35 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:35 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:33 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage1005.eqiad.wmnet * 12:33 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage1004.eqiad.wmnet * 12:33 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage1004.eqiad.wmnet * 12:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage1004.eqiad.wmnet * 12:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage1004.eqiad.wmnet * 12:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage1003.eqiad.wmnet * 12:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage1003.eqiad.wmnet * 12:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage * 12:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage1003.eqiad.wmnet * 12:17 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage * 12:14 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage1003.eqiad.wmnet * 12:14 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-staging-worker-eqiad * 12:03 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw1001.wikimedia.org with OS trixie * 12:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cuminunpriv1001.eqiad.wmnet * 11:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cuminunpriv1001.eqiad.wmnet * 11:27 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1004.wikimedia.org * 11:24 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1153 from dbctl [[phab:T434638|T434638]]', diff saved to https://phabricator.wikimedia.org/P96097 and previous config saved to /var/cache/conftool/dbconfig/20260814-112449-marostegui.json * 11:21 aokoth@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1004.wikimedia.org * 11:20 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 11:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-staging-worker-codfw * 11:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2004.codfw.wmnet * 11:17 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2004.codfw.wmnet * 11:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2004.codfw.wmnet * 11:10 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2004.codfw.wmnet * 11:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2003.codfw.wmnet * 11:10 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2003.codfw.wmnet * 11:03 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-worker1181.eqiad.wmnet with OS bookworm * 11:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2003.codfw.wmnet * 11:03 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2003.codfw.wmnet * 11:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2002.codfw.wmnet * 11:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2002.codfw.wmnet * 10:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1235.eqiad.wmnet with OS bookworm * 10:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock2003.codfw.wmnet * 10:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock2003.codfw.wmnet with OS trixie * 10:56 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2002.codfw.wmnet * 10:54 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1153.eqiad.wmnet with OS bookworm * 10:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2002.codfw.wmnet * 10:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2001.codfw.wmnet * 10:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2001.codfw.wmnet * 10:50 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 10:49 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1187.eqiad.wmnet with OS bookworm * 10:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2001.codfw.wmnet * 10:44 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2001.codfw.wmnet * 10:44 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-staging-worker-codfw * 10:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock2003.codfw.wmnet with reason: host reimage * 10:38 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock2003.codfw.wmnet with reason: host reimage * 10:35 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1153.eqiad.wmnet with reason: host reimage * 10:32 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1235.eqiad.wmnet with reason: host reimage * 10:29 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1187.eqiad.wmnet with reason: host reimage * 10:24 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1153.eqiad.wmnet with reason: host reimage * 10:23 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1235.eqiad.wmnet with reason: host reimage * 10:21 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1187.eqiad.wmnet with reason: host reimage * 10:16 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock2003.codfw.wmnet with OS trixie * 10:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install1005.wikimedia.org * 10:12 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 10:12 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 10:12 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 10:12 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 10:12 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:12 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 10:12 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 10:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install1005.wikimedia.org * 10:07 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1235.eqiad.wmnet with OS bookworm * 10:07 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1187.eqiad.wmnet with OS bookworm * 10:07 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1181.eqiad.wmnet with OS bookworm * 10:07 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1153.eqiad.wmnet with OS bookworm * 10:07 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install2005.wikimedia.org * 10:03 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 10:03 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2003.codfw.wmnet * 10:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1232.eqiad.wmnet with OS bookworm * 10:00 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install2005.wikimedia.org * 10:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install3004.wikimedia.org * 09:58 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 09:58 fceratto@cumin1003: Removing db1151 from zarcillo [[phab:T434538|T434538]] * 09:56 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1231.eqiad.wmnet with OS bookworm * 09:56 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 09:53 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 09:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install3004.wikimedia.org * 09:50 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1152.eqiad.wmnet with OS bookworm * 09:49 Dreamy_Jazz: `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260808000000" --end-timestamp="20260812120000" --sleep="5" --batch-size="50"` for [[phab:T434688|T434688]] * 09:48 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host centrallog2002.codfw.wmnet * 09:48 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install4004.wikimedia.org * 09:42 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1232.eqiad.wmnet with reason: host reimage * 09:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install4004.wikimedia.org * 09:41 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host centrallog2002.codfw.wmnet * 09:39 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install5004.wikimedia.org * 09:36 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1232.eqiad.wmnet with reason: host reimage * 09:36 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'. * 09:34 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'. * 09:33 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1231.eqiad.wmnet with reason: host reimage * 09:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install5004.wikimedia.org * 09:32 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host centrallog1002.eqiad.wmnet * 09:30 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install6003.wikimedia.org * 09:30 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1231.eqiad.wmnet with reason: host reimage * 09:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1152.eqiad.wmnet with reason: host reimage * 09:25 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host centrallog1002.eqiad.wmnet * 09:25 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1152.eqiad.wmnet with reason: host reimage * 09:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install6003.wikimedia.org * 09:22 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1232 * 09:22 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1232 * 09:22 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1232 * 09:22 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1232.eqiad.wmnet 25.53.64.10.in-addr.arpa 5.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:22 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host titan1001.eqiad.wmnet * 09:22 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1232.eqiad.wmnet 25.53.64.10.in-addr.arpa 5.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:22 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:22 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1232 - btullis@cumin1003" * 09:22 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1232 - btullis@cumin1003" * 09:17 btullis@cumin1003: START - Cookbook sre.dns.netbox * 09:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install7002.wikimedia.org * 09:17 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1232 * 09:16 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1231 * 09:16 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1231 * 09:14 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1231 * 09:14 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1231.eqiad.wmnet 24.53.64.10.in-addr.arpa 4.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host titan1001.eqiad.wmnet * 09:14 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1231.eqiad.wmnet 24.53.64.10.in-addr.arpa 4.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:14 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1231 - btullis@cumin1003" * 09:14 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1231 - btullis@cumin1003" * 09:10 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install7002.wikimedia.org * 09:09 btullis@cumin1003: START - Cookbook sre.dns.netbox * 09:08 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1231 * 09:08 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1152 * 09:08 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1152 * 09:06 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1152 * 09:06 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1152.eqiad.wmnet 16.53.64.10.in-addr.arpa 6.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:06 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1152.eqiad.wmnet 16.53.64.10.in-addr.arpa 6.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:06 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:06 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1152 - btullis@cumin1003" * 09:06 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1152 - btullis@cumin1003" * 09:03 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1151.eqiad.wmnet * 09:03 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:03 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1151.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 08:55 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host titan2001.codfw.wmnet * 08:55 btullis@cumin1003: START - Cookbook sre.dns.netbox * 08:54 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1151.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 08:54 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1152 * 08:53 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1232.eqiad.wmnet with OS bookworm * 08:53 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1231.eqiad.wmnet with OS bookworm * 08:53 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1152.eqiad.wmnet with OS bookworm * 08:51 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2209: Pool back * 08:51 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1201.eqiad.wmnet * 08:50 btullis@cumin1003: START - Cookbook sre.hosts.remove-downtime for an-worker1201.eqiad.wmnet * 08:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1229.eqiad.wmnet with OS bookworm * 08:47 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host titan2001.codfw.wmnet * 08:46 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping2004.codfw.wmnet * 08:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ping2004.codfw.wmnet * 08:40 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host an-worker1230.eqiad.wmnet with OS bookworm * 08:40 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 08:34 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1151.eqiad.wmnet * 08:34 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 08:31 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host titan1002.eqiad.wmnet * 08:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1229.eqiad.wmnet with reason: host reimage * 08:25 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host titan1002.eqiad.wmnet * 08:25 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1229.eqiad.wmnet with reason: host reimage * 08:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping1004.eqiad.wmnet * 08:21 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ping1004.eqiad.wmnet * 08:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1230.eqiad.wmnet with reason: host reimage * 08:12 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1230.eqiad.wmnet with reason: host reimage * 08:11 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1229 * 08:11 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1229 * 08:11 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1229 * 08:11 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1229.eqiad.wmnet 22.53.64.10.in-addr.arpa 2.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:11 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1229.eqiad.wmnet 22.53.64.10.in-addr.arpa 2.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:11 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:11 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1229 - btullis@cumin1003" * 08:11 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1229 - btullis@cumin1003" * 08:10 btullis@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1201.eqiad.wmnet with reason: Fixing a disk * 08:07 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host titan2002.codfw.wmnet * 08:05 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2209: Pool back * 08:05 btullis@cumin1003: START - Cookbook sre.dns.netbox * 08:00 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host titan2002.codfw.wmnet * 07:59 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1229 * 07:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1230 * 07:58 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1230 * 07:55 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1230 * 07:55 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1230.eqiad.wmnet 23.53.64.10.in-addr.arpa 3.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:55 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1230.eqiad.wmnet 23.53.64.10.in-addr.arpa 3.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:55 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:55 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1230 - btullis@cumin1003" * 07:55 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1230 - btullis@cumin1003" * 07:51 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host kubestagemaster2005.codfw.wmnet with OS trixie * 07:48 btullis@cumin1003: START - Cookbook sre.dns.netbox * 07:41 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1230 * 07:41 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1229.eqiad.wmnet with OS bookworm * 07:41 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1230.eqiad.wmnet with OS bookworm * 07:39 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1152 from dbctl [[phab:T434480|T434480]]', diff saved to https://phabricator.wikimedia.org/P96090 and previous config saved to /var/cache/conftool/dbconfig/20260814-073941-marostegui.json * 07:29 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on kubestagemaster2005.codfw.wmnet with reason: host reimage * 07:23 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on kubestagemaster2005.codfw.wmnet with reason: host reimage * 07:04 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host kubestagemaster2005.codfw.wmnet with OS trixie * 06:53 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1228.eqiad.wmnet with OS bookworm * 06:44 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1227.eqiad.wmnet with OS bookworm * 06:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1209.eqiad.wmnet with OS bookworm * 06:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1175.eqiad.wmnet with OS bookworm * 06:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1228.eqiad.wmnet with reason: host reimage * 06:27 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1228.eqiad.wmnet with reason: host reimage * 06:25 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1227.eqiad.wmnet with reason: host reimage * 06:21 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1227.eqiad.wmnet with reason: host reimage * 06:18 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1209.eqiad.wmnet with reason: host reimage * 06:14 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1209.eqiad.wmnet with reason: host reimage * 06:14 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1228 * 06:14 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1228 * 06:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1175.eqiad.wmnet with reason: host reimage * 06:12 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1228 * 06:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1228.eqiad.wmnet 20.53.64.10.in-addr.arpa 0.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:12 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1228.eqiad.wmnet 20.53.64.10.in-addr.arpa 0.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1228 - ryankemper@cumin2003" * 06:12 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1228 - ryankemper@cumin2003" * 06:09 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1175.eqiad.wmnet with reason: host reimage * 06:07 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 06:07 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1228 * 06:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1227 * 06:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1227 * 06:06 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1227 * 06:06 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1227.eqiad.wmnet 19.53.64.10.in-addr.arpa 9.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:06 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1227.eqiad.wmnet 19.53.64.10.in-addr.arpa 9.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:06 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:06 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1227 - ryankemper@cumin2003" * 06:06 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1227 - ryankemper@cumin2003" * 06:00 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 06:00 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1227 * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1209 * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1209 * 06:00 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1209 * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1209.eqiad.wmnet 15.53.64.10.in-addr.arpa 5.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:00 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1209.eqiad.wmnet 15.53.64.10.in-addr.arpa 5.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1209 - ryankemper@cumin2003" * 06:00 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1209 - ryankemper@cumin2003" * 05:54 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 05:54 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1209 * 05:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1175 * 05:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1175 * 05:52 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1175 * 05:52 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1175.eqiad.wmnet 17.53.64.10.in-addr.arpa 7.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 05:52 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1175.eqiad.wmnet 17.53.64.10.in-addr.arpa 7.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 05:52 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 05:52 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1175 - ryankemper@cumin2003" * 05:52 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1175 - ryankemper@cumin2003" * 05:49 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1228.eqiad.wmnet with OS bookworm * 05:49 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1227.eqiad.wmnet with OS bookworm * 05:48 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1209.eqiad.wmnet with OS bookworm * 05:47 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 05:47 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1175 * 05:47 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1175.eqiad.wmnet with OS bookworm * 05:09 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host apifeatureusage1001.eqiad.wmnet with OS bookworm * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 03s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:10 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 01:07 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 01:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 01:02 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 00:59 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1226.eqiad.wmnet with OS bookworm * 00:47 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1225.eqiad.wmnet with OS bookworm * 00:41 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1224.eqiad.wmnet with OS bookworm * 00:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1226.eqiad.wmnet with reason: host reimage * 00:31 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1226.eqiad.wmnet with reason: host reimage * 00:28 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1225.eqiad.wmnet with reason: host reimage * 00:25 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1225.eqiad.wmnet with reason: host reimage * 00:19 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1224.eqiad.wmnet with reason: host reimage * 00:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1226 * 00:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1226 * 00:17 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1226 * 00:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1226.eqiad.wmnet 23.36.64.10.in-addr.arpa 3.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:17 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1226.eqiad.wmnet 23.36.64.10.in-addr.arpa 3.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 00:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1226 - ryankemper@cumin2003" * 00:17 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1226 - ryankemper@cumin2003" * 00:16 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1224.eqiad.wmnet with reason: host reimage * 00:12 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 00:12 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1226 * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1225 * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1225 * 00:10 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1225 * 00:10 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1225.eqiad.wmnet 22.36.64.10.in-addr.arpa 2.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:10 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1225.eqiad.wmnet 22.36.64.10.in-addr.arpa 2.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:10 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 00:10 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1225 - ryankemper@cumin2003" * 00:10 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1225 - ryankemper@cumin2003" * 00:03 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 00:02 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1225 * 00:02 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1224 * 00:02 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1224 * 00:00 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1224 * 00:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1224.eqiad.wmnet 21.36.64.10.in-addr.arpa 1.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:00 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1224.eqiad.wmnet 21.36.64.10.in-addr.arpa 1.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 00:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1224 - ryankemper@cumin2003" * 00:00 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1224 - ryankemper@cumin2003" == 2026-08-13 == * 23:54 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1226.eqiad.wmnet with OS bookworm * 23:53 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1225.eqiad.wmnet with OS bookworm * 23:52 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 23:51 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1224 * 23:51 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1224.eqiad.wmnet with OS bookworm * 23:47 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1223.eqiad.wmnet with OS bookworm * 23:29 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1223.eqiad.wmnet with reason: host reimage * 23:24 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1223.eqiad.wmnet with reason: host reimage * 23:19 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325525{{!}}ve.ui.CodeMirror.less: ensure normal font style]] (duration: 11m 40s) * 23:16 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260807000000" --end-timestamp="20260808000000" --sleep="5" --batch-size="50"` for [[phab:T434688|T434688]] * 23:13 musikanimal@deploy1003: musikanimal: Continuing with deployment * 23:11 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1325525{{!}}ve.ui.CodeMirror.less: ensure normal font style]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1223 * 23:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1223 * 23:08 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1325525{{!}}ve.ui.CodeMirror.less: ensure normal font style]] * 23:07 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1223 * 23:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1223.eqiad.wmnet 20.36.64.10.in-addr.arpa 0.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:07 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1223.eqiad.wmnet 20.36.64.10.in-addr.arpa 0.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 23:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1223 - ryankemper@cumin2003" * 23:03 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1223 - ryankemper@cumin2003" * 23:01 sbassett: Deployed security updates for [[phab:T430596|T430596]], [[phab:T120386|T120386]] * 22:55 bking@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host kubestagemaster2005.codfw.wmnet * 22:55 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host kubestagemaster2005.codfw.wmnet with OS bookworm * 22:54 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 22:53 ryankemper@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 22:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on kubestagemaster2005.codfw.wmnet with reason: host reimage * 22:49 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1222.eqiad.wmnet with OS bookworm * 22:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on kubestagemaster2005.codfw.wmnet with reason: host reimage * 22:29 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1222.eqiad.wmnet with reason: host reimage * 22:26 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1222.eqiad.wmnet with reason: host reimage * 22:23 bking@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host aux-k8s-etcd2003.codfw.wmnet * 22:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd2003.codfw.wmnet with OS bookworm * 22:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host kubestagemaster2005.codfw.wmnet with OS bookworm * 22:23 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM kubestagemaster2005.codfw.wmnet - bking@cumin2003" * 22:23 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM kubestagemaster2005.codfw.wmnet - bking@cumin2003" * 22:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) kubestagemaster2005.codfw.wmnet on all recursors * 22:22 bking@cumin2003: START - Cookbook sre.dns.wipe-cache kubestagemaster2005.codfw.wmnet on all recursors * 22:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:22 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM kubestagemaster2005.codfw.wmnet - bking@cumin2003" * 22:22 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM kubestagemaster2005.codfw.wmnet - bking@cumin2003" * 22:17 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 22:13 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1223 * 22:12 bking@cumin2003: START - Cookbook sre.dns.netbox * 22:12 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host kubestagemaster2005.codfw.wmnet * 22:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1222 * 22:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1222 * 22:11 sbassett: Deployed security updates for [[phab:T429244|T429244]], [[phab:T434039|T434039]] * 22:11 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1222 * 22:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1222.eqiad.wmnet 19.36.64.10.in-addr.arpa 9.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:11 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1222.eqiad.wmnet 19.36.64.10.in-addr.arpa 9.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1222 - ryankemper@cumin2003" * 22:07 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1222 - ryankemper@cumin2003" * 22:05 robh@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-wdqs2001.codfw.wmnet with reason: updating firmware * 22:01 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 22:01 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1223.eqiad.wmnet with OS bookworm * 22:01 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1222 * 22:01 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1222.eqiad.wmnet with OS bookworm * 22:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1212.eqiad.wmnet with OS bookworm * 21:54 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324295{{!}}Revert "Lazily reject pre-fix parser-cache entries for noreferrer/noopener links" (T429090)]] (duration: 06m 42s) * 21:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd2003.codfw.wmnet with reason: host reimage * 21:53 bking@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host dse-k8s-etcd2001.codfw.wmnet * 21:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dse-k8s-etcd2001.codfw.wmnet with OS bookworm * 21:50 sbassett@deploy1003: sbassett, kharlan: Continuing with deployment * 21:49 sbassett@deploy1003: sbassett, kharlan: Backport for [[gerrit:1324295{{!}}Revert "Lazily reject pre-fix parser-cache entries for noreferrer/noopener links" (T429090)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:48 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aux-k8s-etcd2003.codfw.wmnet with reason: host reimage * 21:47 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1324295{{!}}Revert "Lazily reject pre-fix parser-cache entries for noreferrer/noopener links" (T429090)]] * 21:42 maryum: Deployed security patch for [[phab:T434549|T434549]] * 21:39 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1212.eqiad.wmnet with reason: host reimage * 21:34 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1212.eqiad.wmnet with reason: host reimage * 21:33 bking@cumin2003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd2003.codfw.wmnet with OS bookworm * 21:32 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM aux-k8s-etcd2003.codfw.wmnet - bking@cumin2003" * 21:32 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM aux-k8s-etcd2003.codfw.wmnet - bking@cumin2003" * 21:32 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) aux-k8s-etcd2003.codfw.wmnet on all recursors * 21:32 bking@cumin2003: START - Cookbook sre.dns.wipe-cache aux-k8s-etcd2003.codfw.wmnet on all recursors * 21:32 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:32 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM aux-k8s-etcd2003.codfw.wmnet - bking@cumin2003" * 21:31 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM aux-k8s-etcd2003.codfw.wmnet - bking@cumin2003" * 21:28 maryum: Deployed security patch for [[phab:T434619|T434619]] * 21:26 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:26 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host aux-k8s-etcd2003.codfw.wmnet * 21:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-etcd2001.codfw.wmnet with reason: host reimage * 21:20 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1212 * 21:20 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1212 * 21:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1212.eqiad.wmnet with OS bookworm * 21:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on dse-k8s-etcd2001.codfw.wmnet with reason: host reimage * 21:04 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325553{{!}}InstrumentConstructiveEdits: exclude mw-reverted as well (T431493)]] (duration: 06m 25s) * 20:59 kemayo@deploy1003: kemayo: Continuing with deployment * 20:59 kemayo@deploy1003: kemayo: Backport for [[gerrit:1325553{{!}}InstrumentConstructiveEdits: exclude mw-reverted as well (T431493)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host dse-k8s-etcd2001.codfw.wmnet with OS bookworm * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 20:57 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 20:57 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1325553{{!}}InstrumentConstructiveEdits: exclude mw-reverted as well (T431493)]] * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-etcd2001.codfw.wmnet on all recursors * 20:57 bking@cumin2003: START - Cookbook sre.dns.wipe-cache dse-k8s-etcd2001.codfw.wmnet on all recursors * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 20:57 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 20:54 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320163{{!}}Add configurable RestTermsOfServiceUrl (T428147)]] (duration: 21m 39s) * 20:53 bking@cumin2003: START - Cookbook sre.dns.netbox * 20:53 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host dse-k8s-etcd2001.codfw.wmnet * 20:50 samtar@deploy1003: samtar, milazg: Continuing with deployment * 20:35 samtar@deploy1003: samtar, milazg: Backport for [[gerrit:1320163{{!}}Add configurable RestTermsOfServiceUrl (T428147)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:33 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1320163{{!}}Add configurable RestTermsOfServiceUrl (T428147)]] * 20:30 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325549{{!}}Deploy PRV to several LC wikis (T423785)]] (duration: 06m 57s) * 20:26 arlolra@deploy1003: arlolra: Continuing with deployment * 20:25 arlolra@deploy1003: arlolra: Backport for [[gerrit:1325549{{!}}Deploy PRV to several LC wikis (T423785)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:24 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 20:23 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1325549{{!}}Deploy PRV to several LC wikis (T423785)]] * 20:21 ariel@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319906{{!}}Remove boilerplate language from wmf-rest and wmf-math API modules (T433736)]] (duration: 13m 54s) * 20:14 ariel@deploy1003: ariel: Continuing with deployment * 20:11 ariel@deploy1003: ariel: Backport for [[gerrit:1319906{{!}}Remove boilerplate language from wmf-rest and wmf-math API modules (T433736)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 ariel@deploy1003: Started scap sync-world: Backport for [[gerrit:1319906{{!}}Remove boilerplate language from wmf-rest and wmf-math API modules (T433736)]] * 19:58 robh@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-wdqs2001.codfw.wmnet with reason: updating firmware * 19:54 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325551{{!}}Render the focused module view as a full-screen page (T433896)]] (duration: 30m 37s) * 19:52 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching aqs[2001,1016]*: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 19:44 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching aqs[2001,1016]*: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 19:42 musikanimal@deploy1003: musikanimal: Continuing with deployment * 19:41 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1325551{{!}}Render the focused module view as a full-screen page (T433896)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:34 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: quash java safepoint logspam - bking@cumin2003 - [[phab:T434685|T434685]] * 19:34 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 19:34 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 19:24 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1325551{{!}}Render the focused module view as a full-screen page (T433896)]] * 19:20 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 19:20 jhancock@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin1003" * 19:18 jhancock@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin1003" * 19:14 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325548{{!}}Enable image lazy loading on desktop in group1 (T148047)]] (duration: 07m 43s) * 19:10 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 19:10 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1325548{{!}}Enable image lazy loading on desktop in group1 (T148047)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:07 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1325548{{!}}Enable image lazy loading on desktop in group1 (T148047)]] * 19:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1166.eqiad.wmnet onto db1280.eqiad.wmnet * 19:03 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 19:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1166: Pool db1166.eqiad.wmnet in after cloning * 18:59 jhancock@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 18:58 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1234.eqiad.wmnet with OS bookworm * 18:52 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 18:51 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:49 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:46 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 18:46 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 18:44 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=dns3004.* * 18:39 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1221.eqiad.wmnet with OS bookworm * 18:36 inflatador: [bking@ganeti2048] ~$ sudo gnt-instance replace-disks -n ganeti2030.codfw.wmnet aux-k8s-worker2002.codfw.wmnet [[phab:T434681|T434681]] * 18:35 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 18:29 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1220.eqiad.wmnet with OS bookworm * 18:29 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1234.eqiad.wmnet with reason: host reimage * 18:27 bking@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host dse-k8s-etcd2001.codfw.wmnet * 18:27 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-etcd2001.codfw.wmnet on all recursors * 18:27 bking@cumin2003: START - Cookbook sre.dns.wipe-cache dse-k8s-etcd2001.codfw.wmnet on all recursors * 18:27 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:27 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 18:27 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 18:25 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1234.eqiad.wmnet with reason: host reimage * 18:25 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 18:19 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1221.eqiad.wmnet with reason: host reimage * 18:17 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1166: Pool db1166.eqiad.wmnet in after cloning * 18:15 brennen@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 18:15 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1221.eqiad.wmnet with reason: host reimage * 18:14 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:12 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:12 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 18:12 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-etcd2001.codfw.wmnet on all recursors * 18:12 bking@cumin2003: START - Cookbook sre.dns.wipe-cache dse-k8s-etcd2001.codfw.wmnet on all recursors * 18:12 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:12 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 18:12 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 18:10 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: quash java safepoint logspam - bking@cumin2003 - [[phab:T434685|T434685]] * 18:10 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1234 * 18:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1234 * 18:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1220.eqiad.wmnet with reason: host reimage * 18:07 brennen: 1.47.0-wmf.15 train status ([[phab:T430834|T430834]]) - no current blockers, rolling to all wikis * 18:07 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1234 * 18:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1234.eqiad.wmnet 10.36.64.10.in-addr.arpa 0.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:07 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:07 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1234.eqiad.wmnet 10.36.64.10.in-addr.arpa 0.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1234 - ryankemper@cumin2003" * 18:07 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1234 - ryankemper@cumin2003" * 18:05 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1220.eqiad.wmnet with reason: host reimage * 18:04 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host dse-k8s-etcd2001.codfw.wmnet * 18:04 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 18:03 dancy@deploy1003: Installation of scap version "4.280.1" completed for 3 hosts * 18:02 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:02 inflatador: bking@dse-k8s-etcd2002 etcdctl member remove $<nowiki>{</nowiki>UUID of dse-k8s-etcd2001<nowiki>}</nowiki> [[phab:T434681|T434681]] [[phab:T434793|T434793]] * 18:01 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 18:01 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1234 * 18:01 dancy@deploy1003: Installing scap version "4.280.1" for 3 host(s) * 18:01 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1221 * 18:01 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1221 * 18:00 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1221 * 18:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1221.eqiad.wmnet 18.36.64.10.in-addr.arpa 8.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:00 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1221.eqiad.wmnet 18.36.64.10.in-addr.arpa 8.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1221 - ryankemper@cumin2003" * 17:59 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:58 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 17:58 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:57 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 17:57 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:56 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1221 - ryankemper@cumin2003" * 17:53 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 17:52 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 17:52 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 17:51 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1221 * 17:51 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1220 * 17:51 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1220 * 17:51 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:51 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:51 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1220 * 17:51 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1220.eqiad.wmnet 11.36.64.10.in-addr.arpa 1.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:51 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1220.eqiad.wmnet 11.36.64.10.in-addr.arpa 1.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:51 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1220 - ryankemper@cumin2003" * 17:50 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1220 - ryankemper@cumin2003" * 17:47 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:47 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 17:46 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1234.eqiad.wmnet with OS bookworm * 17:46 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 17:45 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1221.eqiad.wmnet with OS bookworm * 17:45 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1220 * 17:45 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1220.eqiad.wmnet with OS bookworm * 17:43 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 17:41 bking@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host dse-k8s-etcd2001.codfw.wmnet * 17:41 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host dse-k8s-etcd2001.codfw.wmnet * 17:40 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:40 inflatador: bking@ganeti2048] `sudo gnt-instance remove --force --ignore-failures --shutdown-timeout=0` on non-DRBD VMs [[phab:T434681|T434681]] * 17:40 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:39 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:38 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:36 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:32 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:32 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:28 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 17:26 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:26 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:24 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:23 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:21 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1218.eqiad.wmnet with OS bookworm * 17:20 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:20 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:19 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1179.eqiad.wmnet with OS bookworm * 17:18 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1150.eqiad.wmnet with OS bookworm * 17:18 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns3004.wikimedia.org with OS trixie * 17:15 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:14 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:13 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:12 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:12 swfrench@deploy1003: Finished scap sync-world: Helmfile-only deployment for mediawiki chart bump - [[phab:T427666|T427666]] (duration: 03m 03s) * 17:09 swfrench@deploy1003: Started scap sync-world: Helmfile-only deployment for mediawiki chart bump - [[phab:T427666|T427666]] * 17:02 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:01 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:01 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1218.eqiad.wmnet with reason: host reimage * 17:01 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:01 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:00 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324810{{!}}deployment-info.php: Report dbname and branch for the requested wiki (T434726)]] (duration: 06m 52s) * 16:58 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 16:57 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1179.eqiad.wmnet with reason: host reimage * 16:56 dancy@deploy1003: dancy: Continuing with deployment * 16:56 dancy@deploy1003: dancy: Backport for [[gerrit:1324810{{!}}deployment-info.php: Report dbname and branch for the requested wiki (T434726)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1150.eqiad.wmnet with reason: host reimage * 16:53 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324810{{!}}deployment-info.php: Report dbname and branch for the requested wiki (T434726)]] * 16:51 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1166: Depool db1166.eqiad.wmnet to then clone it to db1280.eqiad.wmnet - cwilliams@cumin1003 * 16:50 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1166: Depool db1166.eqiad.wmnet to then clone it to db1280.eqiad.wmnet - cwilliams@cumin1003 * 16:50 cwilliams@cumin1003: START - Cookbook sre.mysql.clone of db1166.eqiad.wmnet onto db1280.eqiad.wmnet * 16:49 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1218.eqiad.wmnet with reason: host reimage * 16:48 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1179.eqiad.wmnet with reason: host reimage * 16:47 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1150.eqiad.wmnet with reason: host reimage * 16:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.upgrade (exit_code=0) for 1 hosts * 16:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2220: Upgrade of db2220.codfw.wmnet completed * 16:38 swfrench@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 16:38 swfrench-wmf: kubectl delete node kubestagemaster2005.codfw.wmnet - [[phab:T434681|T434681]] * 16:34 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1218 * 16:34 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1218 * 16:34 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1218.eqiad.wmnet with OS bookworm * 16:34 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1179 * 16:34 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1179 * 16:33 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1179.eqiad.wmnet with OS bookworm * 16:32 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1150 * 16:32 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1150 * 16:31 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1150.eqiad.wmnet with OS bookworm * 16:24 dancy@deploy1003: Installation of scap version "4.280.0" completed for 3 hosts * 16:22 dancy@deploy1003: Installing scap version "4.280.0" for 3 host(s) * 16:20 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: quash java safepoint logspam - bking@cumin2003 - [[phab:T434685|T434685]] * 16:14 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns3004.wikimedia.org with reason: host reimage * 16:08 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 16:07 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 16:07 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns3004.wikimedia.org with reason: host reimage * 16:05 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 16:04 swfrench@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 16:00 swfrench@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 15:59 swfrench@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 15:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Upgrade of db2220.codfw.wmnet completed * 15:48 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2220: Upgrading db2220.codfw.wmnet * 15:48 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2220: Upgrading db2220.codfw.wmnet * 15:48 cwilliams@cumin1003: START - Cookbook sre.mysql.upgrade for 1 hosts * 15:46 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns3004.wikimedia.org with OS trixie * 15:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2220 [[phab:T434802|T434802]]', diff saved to https://phabricator.wikimedia.org/P96079 and previous config saved to /var/cache/conftool/dbconfig/20260813-154624-cwilliams.json * 15:45 cdobbins@cumin1003: conftool action : set/pooled=no; selector: name=dns3004.* * 15:44 cjd91: depooling dns3004 to reimage to trixie * 15:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2159 to s7 primary [[phab:T434802|T434802]]', diff saved to https://phabricator.wikimedia.org/P96078 and previous config saved to /var/cache/conftool/dbconfig/20260813-154405-cwilliams.json * 15:43 cezmunsta: Starting s7 codfw failover from db2220 to db2159 - [[phab:T434802|T434802]] * 15:41 cgoubert@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: Dragonfly supernodes reboot (duration: 09m 42s) * 15:41 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dragonfly-supernode2001.codfw.wmnet * 15:39 inflatador: bking@ganeti2048] ~$ sudo gnt-node failover -f --ignore-consistency ganeti2046.codfw.wmnet [[phab:T434681|T434681]] * 15:39 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 15:39 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 15:39 swfrench@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 15:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2159 with weight 0 [[phab:T434802|T434802]]', diff saved to https://phabricator.wikimedia.org/P96077 and previous config saved to /var/cache/conftool/dbconfig/20260813-153806-cwilliams.json * 15:37 swfrench@dns1004: END - running authdns-update * 15:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 30 hosts with reason: Primary switchover s7 [[phab:T434802|T434802]] * 15:37 cgoubert@cumin2003: START - Cookbook sre.hosts.reboot-single for host dragonfly-supernode2001.codfw.wmnet * 15:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dragonfly-supernode1001.eqiad.wmnet * 15:35 swfrench@dns1004: START - running authdns-update * 15:32 cgoubert@cumin2003: START - Cookbook sre.hosts.reboot-single for host dragonfly-supernode1001.eqiad.wmnet * 15:32 cgoubert@deploy1003: Locking from deployment [ALL REPOSITORIES]: Dragonfly supernodes reboot * 15:30 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-master-codfw * 15:30 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2005.codfw.wmnet * 15:30 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2005.codfw.wmnet * 15:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1169.eqiad.wmnet onto db1277.eqiad.wmnet * 15:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1169: Pool db1169.eqiad.wmnet in after cloning * 15:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2005.codfw.wmnet * 15:23 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2005.codfw.wmnet * 15:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2004.codfw.wmnet * 15:23 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2004.codfw.wmnet * 15:18 inflatador: bking@ganeti2048 sudo gnt-node failover -f ganeti2046.codfw.wmnet [[phab:T434681|T434681]] * 15:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2004.codfw.wmnet * 15:17 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2004.codfw.wmnet * 15:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2003.codfw.wmnet * 15:17 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2003.codfw.wmnet * 15:15 cdobbins@dns1004: END - running authdns-update * 15:13 cdobbins@dns1004: START - running authdns-update * 15:10 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1217.eqiad.wmnet with OS bookworm * 15:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2003.codfw.wmnet * 15:10 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2003.codfw.wmnet * 15:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2002.codfw.wmnet * 15:10 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2002.codfw.wmnet * 15:10 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: quash java safepoint logspam - bking@cumin2003 - [[phab:T434685|T434685]] * 15:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1216.eqiad.wmnet with OS bookworm * 15:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2002.codfw.wmnet * 15:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2002.codfw.wmnet * 15:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2001.codfw.wmnet * 15:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2001.codfw.wmnet * 15:01 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1042.eqiad.wmnet * 15:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1042.eqiad.wmnet * 15:00 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325455{{!}}Api: Use correct query when continue prop=categories (T433922)]] (duration: 09m 47s) * 14:59 bking@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host dse-k8s-etcd2004.codfw.wmnet * 14:58 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-etcd2004.codfw.wmnet on all recursors * 14:58 bking@cumin2003: START - Cookbook sre.dns.wipe-cache dse-k8s-etcd2004.codfw.wmnet on all recursors * 14:58 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:58 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM dse-k8s-etcd2004.codfw.wmnet - bking@cumin2003" * 14:58 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM dse-k8s-etcd2004.codfw.wmnet - bking@cumin2003" * 14:55 zabe@deploy1003: zabe: Continuing with deployment * 14:53 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-etcd2004.codfw.wmnet on all recursors * 14:53 bking@cumin2003: START - Cookbook sre.dns.wipe-cache dse-k8s-etcd2004.codfw.wmnet on all recursors * 14:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:53 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2004.codfw.wmnet - bking@cumin2003" * 14:53 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2004.codfw.wmnet - bking@cumin2003" * 14:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2001.codfw.wmnet * 14:52 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2001.codfw.wmnet * 14:52 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-master-codfw * 14:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2001.codfw.wmnet * 14:52 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2001.codfw.wmnet * 14:52 zabe@deploy1003: zabe: Backport for [[gerrit:1325455{{!}}Api: Use correct query when continue prop=categories (T433922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:50 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1325455{{!}}Api: Use correct query when continue prop=categories (T433922)]] * 14:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1217.eqiad.wmnet with reason: host reimage * 14:48 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:48 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host dse-k8s-etcd2004.codfw.wmnet * 14:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-master-eqiad * 14:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1006.eqiad.wmnet * 14:44 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1006.eqiad.wmnet * 14:44 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1216.eqiad.wmnet with reason: host reimage * 14:40 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1217.eqiad.wmnet with reason: host reimage * 14:39 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1169: Pool db1169.eqiad.wmnet in after cloning * 14:39 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1216.eqiad.wmnet with reason: host reimage * 14:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl1005.eqiad.wmnet * 14:32 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl1005.eqiad.wmnet * 14:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1004.eqiad.wmnet * 14:32 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1004.eqiad.wmnet * 14:31 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus2008.codfw.wmnet * 14:31 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor1003.eqiad.wmnet * 14:29 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1042.eqiad.wmnet * 14:27 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor1003.eqiad.wmnet * 14:26 moritzm: installing Django security updates * 14:26 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling reboot on A:wikidough * 14:25 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1217.eqiad.wmnet with OS bookworm * 14:25 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1216.eqiad.wmnet with OS bookworm * 14:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl1004.eqiad.wmnet * 14:24 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl1004.eqiad.wmnet * 14:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1003.eqiad.wmnet * 14:24 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1003.eqiad.wmnet * 14:23 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus2008.codfw.wmnet * 14:23 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus1008.eqiad.wmnet * 14:23 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor-dev2001.codfw.wmnet * 14:22 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus2006.codfw.wmnet * 14:19 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor-dev2001.codfw.wmnet * 14:18 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1042.eqiad.wmnet * 14:17 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.* * 14:16 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl1003.eqiad.wmnet * 14:16 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl1003.eqiad.wmnet * 14:16 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1002.eqiad.wmnet * 14:16 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1002.eqiad.wmnet * 14:15 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus1008.eqiad.wmnet * 14:14 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus1006.eqiad.wmnet * 14:12 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus2006.codfw.wmnet * 14:12 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor2003.codfw.wmnet * 14:11 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus2007.codfw.wmnet * 14:11 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1041.eqiad.wmnet * 14:11 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1041.eqiad.wmnet * 14:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl1002.eqiad.wmnet * 14:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl1002.eqiad.wmnet * 14:09 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-master-eqiad * 14:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on apifeatureusage1001.eqiad.wmnet with reason: host reimage * 14:08 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1215.eqiad.wmnet with OS bookworm * 14:08 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor2003.codfw.wmnet * 14:08 jayme: updated calico to v3.30.7 on wikikube codfw [[phab:T427400|T427400]] * 14:07 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1149.eqiad.wmnet with OS bookworm * 14:06 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'. * 14:06 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1041.eqiad.wmnet * 14:04 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus1006.eqiad.wmnet * 14:03 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus2007.codfw.wmnet * 14:03 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus2005.codfw.wmnet * 14:03 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus1007.eqiad.wmnet * 14:03 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1002.eqiad.wmnet * 14:02 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on apifeatureusage1001.eqiad.wmnet with reason: host reimage * 14:02 cgoubert@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=helm-charts.*,name=eqiad * 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host chartmuseum1001.eqiad.wmnet * 14:00 moritzm: installing libxml2 security updates * 13:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1214.eqiad.wmnet with OS bookworm * 13:59 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns5003.* * 13:59 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1041.eqiad.wmnet * 13:59 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow1002.eqiad.wmnet * 13:58 cmooney@dns3003: END - running authdns-update * 13:57 cgoubert@cumin2003: START - Cookbook sre.hosts.reboot-single for host chartmuseum1001.eqiad.wmnet * 13:57 cgoubert@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=helm-charts.*,name=eqiad * 13:57 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: cloudelastic cluster restart - bking@cumin2003 * 13:57 cgoubert@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=helm-charts.*,name=codfw * 13:56 cmooney@dns3003: START - running authdns-update * 13:56 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'. * 13:56 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns5003.*,service=authdns-update * 13:55 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host chartmuseum2001.codfw.wmnet * 13:55 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus1007.eqiad.wmnet * 13:55 cmooney@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dns5003.wikimedia.org * 13:55 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus1005.eqiad.wmnet * 13:51 cgoubert@cumin2003: START - Cookbook sre.hosts.reboot-single for host chartmuseum2001.codfw.wmnet * 13:51 cgoubert@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=helm-charts.*,name=codfw * 13:51 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus2005.codfw.wmnet * 13:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host apifeatureusage1001.eqiad.wmnet with OS bookworm * 13:50 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus7002.magru.wmnet * 13:50 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts lvs1015.eqiad.wmnet * 13:50 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:50 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1015.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:49 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1015.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:49 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325480{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] (duration: 06m 39s) * 13:47 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1215.eqiad.wmnet with reason: host reimage * 13:46 cmooney@cumin1003: START - Cookbook sre.hosts.reboot-single for host dns5003.wikimedia.org * 13:46 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1003.eqiad.wmnet * 13:45 cmooney@cumin1003: conftool action : set/pooled=no; selector: name=dns5003.* * 13:45 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1040.eqiad.wmnet * 13:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1040.eqiad.wmnet * 13:45 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 13:45 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus1005.eqiad.wmnet * 13:45 stran@deploy1003: stran: Continuing with deployment * 13:44 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus7002.magru.wmnet * 13:44 stran@deploy1003: stran: Backport for [[gerrit:1325480{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:44 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus6002.drmrs.wmnet * 13:43 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260805000000" --end-timestamp="20260806000000" --sleep="5" --batch-size="10"` for [[phab:T434688|T434688]] * 13:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1149.eqiad.wmnet with reason: host reimage * 13:42 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1325480{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] * 13:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow1003.eqiad.wmnet * 13:40 sukhe@cumin1003: START - Cookbook sre.hosts.decommission for hosts lvs1015.eqiad.wmnet * 13:40 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts lvs1014.eqiad.wmnet * 13:40 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:40 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1014.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:40 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1040.eqiad.wmnet * 13:40 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1014.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:39 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1214.eqiad.wmnet with reason: host reimage * 13:38 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2004.codfw.wmnet * 13:38 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus6002.drmrs.wmnet * 13:38 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus5003.eqsin.wmnet * 13:35 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 13:35 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1149.eqiad.wmnet with reason: host reimage * 13:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1215.eqiad.wmnet with reason: host reimage * 13:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow2004.codfw.wmnet * 13:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1214.eqiad.wmnet with reason: host reimage * 13:33 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1040.eqiad.wmnet * 13:31 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus5003.eqsin.wmnet * 13:31 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: cloudelastic cluster restart - bking@cumin2003 * 13:31 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus4003.ulsfo.wmnet * 13:30 sukhe@cumin1003: START - Cookbook sre.hosts.decommission for hosts lvs1014.eqiad.wmnet * 13:30 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts lvs1013.eqiad.wmnet * 13:30 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:30 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1013.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:30 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1013.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:27 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1039.eqiad.wmnet * 13:27 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1039.eqiad.wmnet * 13:26 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1167.eqiad.wmnet onto db1281.eqiad.wmnet * 13:26 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1167: Pool db1167.eqiad.wmnet in after cloning * 13:25 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus4003.ulsfo.wmnet * 13:25 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260802000000" --end-timestamp="20260803000000" --sleep="5" --batch-size="10"` for [[phab:T434688|T434688]] * 13:24 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 13:24 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus3004.esams.wmnet * 13:24 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1274: New host * 13:24 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325476{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] (duration: 07m 19s) * 13:24 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260804000000" --end-timestamp="20260805000000" --sleep="5" --batch-size="10"` for [[phab:T434688|T434688]] * 13:24 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260803000000" --end-timestamp="20260804000000" --sleep="5" --batch-size="10"` for [[phab:T434688|T434688]] * 13:23 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2003.codfw.wmnet * 13:22 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1039.eqiad.wmnet * 13:20 sukhe@cumin1003: START - Cookbook sre.hosts.decommission for hosts lvs1013.eqiad.wmnet * 13:20 stran@deploy1003: stran: Continuing with deployment * 13:19 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow2003.codfw.wmnet * 13:19 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling reboot on A:wikidough * 13:19 stran@deploy1003: stran: Backport for [[gerrit:1325476{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:18 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus3004.esams.wmnet * 13:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1215.eqiad.wmnet with OS bookworm * 13:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1214.eqiad.wmnet with OS bookworm * 13:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1149.eqiad.wmnet with OS bookworm * 13:17 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1325476{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] * 13:11 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324731{{!}}prv: Enable parsoid rendering for 5 wikisource wikis]] (duration: 07m 26s) * 13:11 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1039.eqiad.wmnet * 13:11 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1036.eqiad.wmnet * 13:11 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1036.eqiad.wmnet * 13:07 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow3004.esams.wmnet * 13:07 jgiannelos@deploy1003: jgiannelos: Continuing with deployment * 13:06 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1233.eqiad.wmnet with OS bookworm * 13:06 jgiannelos@deploy1003: jgiannelos: Backport for [[gerrit:1324731{{!}}prv: Enable parsoid rendering for 5 wikisource wikis]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:04 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1324731{{!}}prv: Enable parsoid rendering for 5 wikisource wikis]] * 13:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow3004.esams.wmnet * 13:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1036.eqiad.wmnet * 13:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1210.eqiad.wmnet with OS bookworm * 13:01 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1036.eqiad.wmnet * 13:00 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1052.eqiad.wmnet * 13:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1052.eqiad.wmnet * 12:58 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow4003.ulsfo.wmnet * 12:55 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1211.eqiad.wmnet with OS bookworm * 12:55 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1052.eqiad.wmnet * 12:52 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow4003.ulsfo.wmnet * 12:51 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1052.eqiad.wmnet * 12:49 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1169: Depool db1169.eqiad.wmnet to then clone it to db1277.eqiad.wmnet - cwilliams@cumin1003 * 12:46 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1169: Depool db1169.eqiad.wmnet to then clone it to db1277.eqiad.wmnet - cwilliams@cumin1003 * 12:46 cwilliams@cumin1003: START - Cookbook sre.mysql.clone of db1169.eqiad.wmnet onto db1277.eqiad.wmnet * 12:45 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:45 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns record for deleted IP reservations eqsin lvs vlan ints - cmooney@cumin1003" * 12:44 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns record for deleted IP reservations eqsin lvs vlan ints - cmooney@cumin1003" * 12:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1051.eqiad.wmnet * 12:43 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1051.eqiad.wmnet * 12:42 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1210.eqiad.wmnet with reason: host reimage * 12:41 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1167: Pool db1167.eqiad.wmnet in after cloning * 12:40 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 12:39 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1274: New host * 12:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Added db1274', diff saved to https://phabricator.wikimedia.org/P96062 and previous config saved to /var/cache/conftool/dbconfig/20260813-123907-cwilliams.json * 12:38 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast7002.wikimedia.org * 12:38 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1233.eqiad.wmnet with reason: host reimage * 12:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1051.eqiad.wmnet * 12:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1211.eqiad.wmnet with reason: host reimage * 12:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1233.eqiad.wmnet with reason: host reimage * 12:32 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1051.eqiad.wmnet * 12:32 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast7002.wikimedia.org * 12:32 marostegui: Drop SecurePoll tables from closed wikis [[phab:T423128|T423128]] * 12:32 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow5003.eqsin.wmnet * 12:31 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1050.eqiad.wmnet * 12:31 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1050.eqiad.wmnet * 12:29 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1210.eqiad.wmnet with reason: host reimage * 12:29 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1211.eqiad.wmnet with reason: host reimage * 12:29 moritzm: installing Wireshark security updates * 12:26 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow5003.eqsin.wmnet * 12:26 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1050.eqiad.wmnet * 12:24 cmooney@dns3003: END - running authdns-update * 12:21 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow6001.drmrs.wmnet * 12:21 cmooney@dns3003: START - running authdns-update * 12:20 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1050.eqiad.wmnet * 12:20 cgoubert@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host rdb-lock2003.codfw.wmnet * 12:20 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:20 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update netbox dns entries for expanded public1-603-eqsin subnet - cmooney@cumin1003" * 12:20 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update netbox dns entries for expanded public1-603-eqsin subnet - cmooney@cumin1003" * 12:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 12:20 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 12:19 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1049.eqiad.wmnet * 12:19 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1049.eqiad.wmnet * 12:19 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 12:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host moss-be1003.eqiad.wmnet * 12:17 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow6001.drmrs.wmnet * 12:17 moritzm: installin curl security updates * 12:15 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 12:15 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1233.eqiad.wmnet with OS bookworm * 12:15 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1210.eqiad.wmnet with OS bookworm * 12:15 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1211.eqiad.wmnet with OS bookworm * 12:13 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1049.eqiad.wmnet * 12:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow7002.magru.wmnet * 12:12 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-worker1178.eqiad.wmnet with OS bookworm * 12:10 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host moss-be1003.eqiad.wmnet * 12:10 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 12:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be1006.eqiad.wmnet * 12:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 12:10 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 12:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 12:10 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 12:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow7002.magru.wmnet * 12:08 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1049.eqiad.wmnet * 12:05 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 12:05 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2003.codfw.wmnet * 12:04 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'. * 12:04 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:03 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be1006.eqiad.wmnet * 12:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be1005.eqiad.wmnet * 12:01 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1038.eqiad.wmnet * 12:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1038.eqiad.wmnet * 12:01 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1213.eqiad.wmnet with OS bookworm * 11:56 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be1005.eqiad.wmnet * 11:55 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be1004.eqiad.wmnet * 11:54 cgoubert@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host rdb-lock2003.codfw.wmnet * 11:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 11:53 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 11:53 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:53 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 11:53 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 11:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1038.eqiad.wmnet * 11:49 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be1004.eqiad.wmnet * 11:44 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:44 moritzm: installing Linux 5.10.262 on Bullseye hosts * 11:41 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1038.eqiad.wmnet * 11:40 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 11:40 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 11:40 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 11:40 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:40 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 11:40 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 11:38 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1213.eqiad.wmnet with reason: host reimage * 11:36 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1167: Depool db1167.eqiad.wmnet to then clone it to db1281.eqiad.wmnet - marostegui@cumin1003 * 11:35 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 11:35 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2003.codfw.wmnet * 11:35 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1167: Depool db1167.eqiad.wmnet to then clone it to db1281.eqiad.wmnet - marostegui@cumin1003 * 11:35 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1167.eqiad.wmnet onto db1281.eqiad.wmnet * 11:34 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 22 hosts with reason: Cloning * 11:34 moritzm: remove ganeti3005 from esams03 cluster, hardware issues [[phab:T434646|T434646]] * 11:32 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1213.eqiad.wmnet with reason: host reimage * 11:28 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1034.eqiad.wmnet * 11:28 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1034.eqiad.wmnet * 11:22 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1034.eqiad.wmnet * 11:19 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1034.eqiad.wmnet * 11:17 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1213.eqiad.wmnet with OS bookworm * 11:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1165.eqiad.wmnet onto db1279.eqiad.wmnet * 11:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1165: Pool db1165.eqiad.wmnet in after cloning * 11:07 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply * 10:57 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply * 10:54 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply * 10:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1274.eqiad.wmnet with reason: Enabling notifications and pooling * 10:45 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1033.eqiad.wmnet * 10:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1033.eqiad.wmnet * 10:44 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply * 10:43 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'. * 10:42 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'. * 10:42 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'. * 10:40 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db[1216,1225,1239-1240].eqiad.wmnet with reason: reboot * 10:39 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1033.eqiad.wmnet * 10:38 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1151 from dbctl [[phab:T434538|T434538]]', diff saved to https://phabricator.wikimedia.org/P96055 and previous config saved to /var/cache/conftool/dbconfig/20260813-103828-marostegui.json * 10:35 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1033.eqiad.wmnet * 10:27 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1165: Pool db1165.eqiad.wmnet in after cloning * 10:24 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1151.eqiad.wmnet with OS bookworm * 10:15 moritzm: installing bind9 security updates (client-side tools/libs only) * 10:07 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-debug: apply * 10:06 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-debug: apply * 10:02 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 7 hosts with reason: reboot * 10:01 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-debug: apply * 10:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.upgrade (exit_code=0) for 1 hosts * 10:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2214: Upgrade of db2214.codfw.wmnet completed * 10:01 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-debug: apply * 10:00 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-debug: apply * 10:00 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-debug: apply * 09:59 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.decommission (exit_code=99) * 09:59 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 09:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1151.eqiad.wmnet with reason: host reimage * 09:59 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 8 hosts * 09:59 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 8 hosts * 09:55 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1151.eqiad.wmnet with reason: host reimage * 09:46 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260801000000" --end-timestamp="20260802000000" --sleep="3" --batch-size="5"` for [[phab:T434688|T434688]] * 09:45 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 8 hosts with reason: reboot * 09:45 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 22 hosts with reason: Cloning * 09:43 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1165: Depool db1165.eqiad.wmnet to then clone it to db1279.eqiad.wmnet - marostegui@cumin1003 * 09:42 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1165: Depool db1165.eqiad.wmnet to then clone it to db1279.eqiad.wmnet - marostegui@cumin1003 * 09:42 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1165.eqiad.wmnet onto db1279.eqiad.wmnet * 09:41 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=testwiki --start-timestamp="20260311000000" --end-timestamp="20260805000000" --sleep="5" --batch-size="2"` for [[phab:T434688|T434688]] * 09:40 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for backupmon1001.eqiad.wmnet * 09:40 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for backupmon1001.eqiad.wmnet * 09:40 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1151 * 09:40 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1151 * 09:37 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325408{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]], [[gerrit:1325407{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]] (duration: 06m 57s) * 09:36 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1151 * 09:36 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1151.eqiad.wmnet 13.36.64.10.in-addr.arpa 3.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:36 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1151.eqiad.wmnet 13.36.64.10.in-addr.arpa 3.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:36 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:36 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1151 - btullis@cumin1003" * 09:36 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on backupmon1001.eqiad.wmnet with reason: reboot * 09:36 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1151 - btullis@cumin1003" * 09:35 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host moss-be2003.codfw.wmnet * 09:33 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 7 hosts * 09:33 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 7 hosts * 09:33 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 09:32 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1325408{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]], [[gerrit:1325407{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1161.eqiad.wmnet onto db1275.eqiad.wmnet * 09:31 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 09:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1161: Pool db1161.eqiad.wmnet in after cloning * 09:30 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1325408{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]], [[gerrit:1325407{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]] * 09:29 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 09:27 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 09:27 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host moss-be2003.codfw.wmnet * 09:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be2006.codfw.wmnet * 09:25 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2004.codfw.wmnet * 09:22 hashar@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 09:21 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be2006.codfw.wmnet * 09:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be2005.codfw.wmnet * 09:19 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2004.codfw.wmnet * 09:19 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2003.codfw.wmnet * 09:18 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 7 hosts with reason: reboot * 09:18 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 7 hosts * 09:18 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 7 hosts * 09:16 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2214: Upgrade of db2214.codfw.wmnet completed * 09:14 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be2005.codfw.wmnet * 09:13 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be2004.codfw.wmnet * 09:12 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2003.codfw.wmnet * 09:12 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2002.codfw.wmnet * 09:09 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2214: Upgrading db2214.codfw.wmnet * 09:09 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2214: Upgrading db2214.codfw.wmnet * 09:09 cwilliams@cumin1003: START - Cookbook sre.mysql.upgrade for 1 hosts * 09:07 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be2004.codfw.wmnet * 09:06 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2002.codfw.wmnet * 09:05 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1004.eqiad.wmnet * 09:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-cluster (exit_code=0) * 09:03 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 7 hosts with reason: reboot * 09:01 btullis@cumin1003: START - Cookbook sre.dns.netbox * 09:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2214 [[phab:T434754|T434754]]', diff saved to https://phabricator.wikimedia.org/P96045 and previous config saved to /var/cache/conftool/dbconfig/20260813-090001-cwilliams.json * 08:59 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1004.eqiad.wmnet * 08:59 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1003.eqiad.wmnet * 08:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2229 to s6 primary [[phab:T434754|T434754]]', diff saved to https://phabricator.wikimedia.org/P96044 and previous config saved to /var/cache/conftool/dbconfig/20260813-085752-cwilliams.json * 08:57 cezmunsta: Starting s6 codfw failover from db2214 to db2229 - [[phab:T434754|T434754]] * 08:54 hashar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325416{{!}}Revert "REST: Enable `GET /lexemes/<nowiki>{</nowiki>lexeme_id<nowiki>}</nowiki>` by default" (T434712)]] (duration: 07m 22s) * 08:53 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1003.eqiad.wmnet * 08:53 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1002.eqiad.wmnet * 08:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2229 with weight 0 [[phab:T434754|T434754]]', diff saved to https://phabricator.wikimedia.org/P96043 and previous config saved to /var/cache/conftool/dbconfig/20260813-085151-cwilliams.json * 08:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 23 hosts with reason: Primary switchover s6 [[phab:T434754|T434754]] * 08:51 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts2002.codfw.wmnet * 08:50 hashar@deploy1003: hashar: Continuing with deployment * 08:49 hashar@deploy1003: hashar: Backport for [[gerrit:1325416{{!}}Revert "REST: Enable `GET /lexemes/<nowiki>{</nowiki>lexeme_id<nowiki>}</nowiki>` by default" (T434712)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:47 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet * 08:47 hashar@deploy1003: Started scap sync-world: Backport for [[gerrit:1325416{{!}}Revert "REST: Enable `GET /lexemes/<nowiki>{</nowiki>lexeme_id<nowiki>}</nowiki>` by default" (T434712)]] * 08:47 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1002.eqiad.wmnet * 08:46 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1161: Pool db1161.eqiad.wmnet in after cloning * 08:46 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host stewards2001.codfw.wmnet * 08:45 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit1003.wikimedia.org * 08:45 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host stewards1001.eqiad.wmnet * 08:45 Emperor: roll-restart apus frontends in codfw * 08:45 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-cluster * 08:44 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts2002.codfw.wmnet * 08:44 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host doc2003.codfw.wmnet * 08:43 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet * 08:43 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host phab2003.codfw.wmnet * 08:42 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host stewards2001.codfw.wmnet * 08:42 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host etherpad2002.codfw.wmnet * 08:41 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host stewards1001.eqiad.wmnet * 08:41 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1273: Pool in s7 * 08:41 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1003.wikimedia.org * 08:40 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host doc2003.codfw.wmnet * 08:40 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit1003.wikimedia.org * 08:39 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host etherpad1004.eqiad.wmnet * 08:39 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host doc1004.eqiad.wmnet * 08:39 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit2002.wikimedia.org * 08:38 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host etherpad2002.codfw.wmnet * 08:37 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host phab2003.codfw.wmnet * 08:36 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lists1004.wikimedia.org * 08:35 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host etherpad1004.eqiad.wmnet * 08:35 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host doc1004.eqiad.wmnet * 08:34 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-cluster (exit_code=0) * 08:34 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1003.wikimedia.org * 08:34 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host planet2003.codfw.wmnet * 08:34 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2003.wikimedia.org * 08:33 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host planet1003.eqiad.wmnet * 08:33 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit2002.wikimedia.org * 08:32 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 22 hosts with reason: Cloning * 08:32 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2002.wikimedia.org * 08:31 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aphlict1002.eqiad.wmnet * 08:30 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host planet2003.codfw.wmnet * 08:29 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host planet1003.eqiad.wmnet * 08:29 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host lists1004.wikimedia.org * 08:28 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lists2001.wikimedia.org * 08:28 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2003.wikimedia.org * 08:27 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host aphlict1002.eqiad.wmnet * 08:27 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aphlict2001.codfw.wmnet * 08:26 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2002.wikimedia.org * 08:23 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host aphlict2001.codfw.wmnet * 08:23 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1151 * 08:22 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1151.eqiad.wmnet with OS bookworm * 08:22 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host lists2001.wikimedia.org * 08:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1161: Depool db1161.eqiad.wmnet to then clone it to db1275.eqiad.wmnet - marostegui@cumin1003 * 08:20 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1161: Depool db1161.eqiad.wmnet to then clone it to db1275.eqiad.wmnet - marostegui@cumin1003 * 08:20 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1161.eqiad.wmnet onto db1275.eqiad.wmnet * 08:15 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-cluster * 08:15 Emperor: roll-restart apus frontends in eqiad * 07:56 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1273: Pool in s7 * 07:56 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1273 to dbctl [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P96035 and previous config saved to /var/cache/conftool/dbconfig/20260813-075611-marostegui.json * 07:38 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db1273.eqiad.wmnet with reason: Reboot * 07:31 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.sanitize-wiki (exit_code=97) Managing sanitization for wikis testwiki in section s3 * 07:24 marostegui@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis testwiki in section s3 * 07:19 jayme@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on kubestagemaster2005.codfw.wmnet with reason: downtime because of hardware failure and no DRBD * 05:42 arnaudb@dns1006: END - running authdns-update * 05:40 arnaudb@dns1006: START - running authdns-update * 05:27 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 04:06 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 02:29 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1151 * 02:29 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1151 * 02:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1219.eqiad.wmnet with OS bookworm * 02:18 ryankemper: [[phab:T434494|T434494]] `ryankemper@deploy1003:~$ echo 'https://stats.wikimedia.org/' {{!}} mwscript-k8s --attach -- purgeList.php` (default page got cached during yesterday's `an-web1001` reimage) * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 46s) * 02:03 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1219.eqiad.wmnet with reason: host reimage * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 02:00 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1219.eqiad.wmnet with reason: host reimage * 01:46 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1219.eqiad.wmnet with OS bookworm * 01:03 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1142.eqiad.wmnet with OS bookworm * 00:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1142.eqiad.wmnet with reason: host reimage * 00:34 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1142.eqiad.wmnet with reason: host reimage * 00:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1142.eqiad.wmnet with OS bookworm * 00:16 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1178 * 00:16 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1178 == 2026-08-12 == * 23:18 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324409{{!}}Remove $wmg = $wg hacks in Collection (T119117)]] (duration: 06m 43s) * 23:14 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 23:13 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1324409{{!}}Remove $wmg = $wg hacks in Collection (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:11 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1324409{{!}}Remove $wmg = $wg hacks in Collection (T119117)]] * 22:59 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 22:49 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260801000000" --end-timestamp="20260802000000" --sleep=2 --batch-size=10` * 22:45 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=testwiki --start-timestamp="20200801010101" --end-timestamp="20260816010101" --sleep=15 --batch-size=5` * 22:40 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=testwiki --start-timestamp="20200101010101" --end-timestamp="20260816010101" --sleep=60` * 22:35 Dreamy_Jazz: Running `mwscript WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260101000000" --end-timestamp="20260102000000" --sleep=10` * 22:21 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324817{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]], [[gerrit:1324818{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]] (duration: 45m 29s) * 22:17 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 21:59 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 21:40 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1324817{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]], [[gerrit:1324818{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:39 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 21:36 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1324817{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]], [[gerrit:1324818{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]] * 21:32 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:31 vriley@cumin1003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:30 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:30 vriley@cumin1003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:17 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:14 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:14 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:11 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:10 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-codfw: Set storage compatability to NONE — [[phab:T433028|T433028]] - eevans@cumin1003 * 21:10 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1006 * 21:09 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1006 * 21:05 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324277{{!}}Improve Math preference labels for SVG/MathJax/MathML (T433891)]] (duration: 31m 42s) * 20:58 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1178.eqiad.wmnet with OS bookworm * 20:54 krinkle@deploy1003: krinkle: Continuing with deployment * 20:51 krinkle@deploy1003: krinkle: Backport for [[gerrit:1324277{{!}}Improve Math preference labels for SVG/MathJax/MathML (T433891)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:39 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 20:38 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:38 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp5022.eqsin.wmnet with OS trixie * 20:38 cdobbins@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - cdobbins@cumin1003" * 20:37 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:36 cdobbins@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - cdobbins@cumin1003" * 20:34 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1178.eqiad.wmnet with reason: host reimage * 20:34 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1324277{{!}}Improve Math preference labels for SVG/MathJax/MathML (T433891)]] * 20:33 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 20:28 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1178.eqiad.wmnet with reason: host reimage * 20:13 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1178.eqiad.wmnet with OS bookworm * 20:11 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-worker1178.eqiad.wmnet with OS bookworm * 20:11 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1178.eqiad.wmnet with OS bookworm * 20:09 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-codfw: Set storage compatability to NONE — [[phab:T433028|T433028]] - eevans@cumin1003 * 20:09 Dreamy_Jazz: Evening UTC backport window done * 20:08 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324794{{!}}WikimediaAntiAbuse: Enable logging channel (T431292)]] (duration: 06m 48s) * 20:08 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage * 20:05 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage * 20:04 dreamyjazz@deploy1003: kharlan, dreamyjazz: Continuing with deployment * 20:04 dreamyjazz@deploy1003: kharlan, dreamyjazz: Backport for [[gerrit:1324794{{!}}WikimediaAntiAbuse: Enable logging channel (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:01 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1324794{{!}}WikimediaAntiAbuse: Enable logging channel (T431292)]] * 19:55 brennen@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 19:47 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 19:35 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 19:35 cdobbins@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cp5022.eqsin.wmnet with OS trixie * 19:32 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:30 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:29 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:26 vriley@cumin1003: START - Cookbook sre.dns.netbox * 19:23 brennen@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 19:19 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-eqiad: Set storage compatability to NONE — [[phab:T433028|T433028]] - eevans@cumin1003 * 19:10 brennen@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 19:09 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 19:09 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 19:00 Amir1: data migrated on wikishared ([[phab:T426102|T426102]]) * 18:57 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321224{{!}}Rename ce_worklist_articles table to ce_invitation_list_articles (T426102)]] (duration: 06m 50s) * 18:53 ladsgroup@deploy1003: ladsgroup, daimona: Continuing with deployment * 18:53 ladsgroup@deploy1003: ladsgroup, daimona: Backport for [[gerrit:1321224{{!}}Rename ce_worklist_articles table to ce_invitation_list_articles (T426102)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:51 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1321224{{!}}Rename ce_worklist_articles table to ce_invitation_list_articles (T426102)]] * 18:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2192: Security update * 18:31 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 18:25 Amir1: ce_invitation_list_articles created as empty on wikishared ([[phab:T426102|T426102]]) * 18:21 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 18:21 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 18:18 Amir1: migrated testwiki entries from ce_worklist_articles to ce_invitation_list_articles ([[phab:T426102|T426102]]) * 18:18 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-eqiad: Set storage compatability to NONE — [[phab:T433028|T433028]] - eevans@cumin1003 * 18:11 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 18:08 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling reboot on A:durum-eqsin and A:durum * 18:07 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Set storage compatability to UPGRADING — [[phab:T433028|T433028]] - eevans@cumin1003 * 18:05 jhancock@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022'] * 17:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2192: Security update * 17:55 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:55 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum-eqsin and A:durum * 17:53 jhancock@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['cp5022'] * 17:47 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:47 jhancock@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['cp5022'] * 17:42 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:41 jhancock@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['cp5022'] * 17:36 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:35 jhancock@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022'] * 17:31 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-magru and not (P<nowiki>{</nowiki>cp7001*<nowiki>}</nowiki> or P<nowiki>{</nowiki>cp7009*<nowiki>}</nowiki>) and A:cp - 9.2.15 upgrade ([[phab:T434620|T434620]]) * 17:28 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:28 jhancock@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['cp5022'] * 17:22 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2192.codfw.wmnet with reason: Maintenance * 17:11 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host apifeatureusage2001.codfw.wmnet with OS bookworm * 17:04 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: apply * 17:03 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-main: apply * 17:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2192 [[phab:T434635|T434635]]', diff saved to https://phabricator.wikimedia.org/P96030 and previous config saved to /var/cache/conftool/dbconfig/20260812-170338-cwilliams.json * 17:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2213 to s5 primary [[phab:T434635|T434635]]', diff saved to https://phabricator.wikimedia.org/P96029 and previous config saved to /var/cache/conftool/dbconfig/20260812-170152-cwilliams.json * 17:01 cezmunsta: Starting s5 codfw failover from db2192 to db2213 - [[phab:T434635|T434635]] * 16:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2213 with weight 0 [[phab:T434635|T434635]]', diff saved to https://phabricator.wikimedia.org/P96028 and previous config saved to /var/cache/conftool/dbconfig/20260812-165544-cwilliams.json * 16:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 27 hosts with reason: Primary switchover s5 [[phab:T434635|T434635]] * 16:53 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: apply * 16:52 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-main: apply * 16:44 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-main: apply * 16:44 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-main: apply * 16:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1272: New host * 16:40 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324765{{!}}WikimediaAntiAbuse: Enable PersonalInfoFlagNotifications (T431292)]] (duration: 07m 02s) * 16:40 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-logging-external: apply * 16:39 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-logging-external: apply * 16:38 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-logging-external: apply * 16:37 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-logging-external: apply * 16:36 kharlan@deploy1003: kharlan: Continuing with deployment * 16:35 kharlan@deploy1003: kharlan: Backport for [[gerrit:1324765{{!}}WikimediaAntiAbuse: Enable PersonalInfoFlagNotifications (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:33 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1324765{{!}}WikimediaAntiAbuse: Enable PersonalInfoFlagNotifications (T431292)]] * 16:27 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-logging-external: apply * 16:27 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-logging-external: apply * 16:18 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 16:18 jhancock@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022'] * 16:17 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 16:16 jhancock@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['cp5022'] * 16:15 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 16:12 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Set storage compatability to UPGRADING — [[phab:T433028|T433028]] - eevans@cumin1003 * 16:10 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: apply * 16:10 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: apply * 16:08 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: apply * 16:08 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: apply * 16:08 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics: apply * 16:07 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics: apply * 16:02 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-magru and not (P<nowiki>{</nowiki>cp7001*<nowiki>}</nowiki> or P<nowiki>{</nowiki>cp7009*<nowiki>}</nowiki>) and A:cp - 9.2.15 upgrade ([[phab:T434620|T434620]]) * 15:57 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1272: New host * 15:52 jmm@cumin2003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti2046.codfw.wmnet * 15:52 jmm@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host ganeti2046.codfw.wmnet * 15:42 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:37 mforns@deploy1003: Finished deploy [analytics/refinery@49c336c] (thin): Regular analytics weekly train THIN [analytics/refinery@49c336cd] (duration: 01m 59s) * 15:37 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324733{{!}}WikimediaAntiAbuse: Enable personal info tag display on enwiki (T431292)]] (duration: 08m 12s) * 15:35 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:35 mforns@deploy1003: Started deploy [analytics/refinery@49c336c] (thin): Regular analytics weekly train THIN [analytics/refinery@49c336cd] * 15:34 mforns@deploy1003: Finished deploy [analytics/refinery@49c336c]: Regular analytics weekly train [analytics/refinery@49c336cd] (duration: 04m 20s) * 15:33 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[2024,1031]*.wmnet: Set storage compatability to UPGRADING — [[phab:T433028|T433028]] - eevans@cumin1003 * 15:33 dreamyjazz@deploy1003: kharlan, dreamyjazz: Continuing with deployment * 15:31 dreamyjazz@deploy1003: kharlan, dreamyjazz: Backport for [[gerrit:1324733{{!}}WikimediaAntiAbuse: Enable personal info tag display on enwiki (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:30 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitize-wiki (exit_code=99) Checking sanitization for wikis testwiki in section s3 * 15:30 mforns@deploy1003: Started deploy [analytics/refinery@49c336c]: Regular analytics weekly train [analytics/refinery@49c336cd] * 15:30 mforns@deploy1003: Finished deploy [analytics/refinery@49c336c] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@49c336cd] (duration: 00m 32s) * 15:29 mforns@deploy1003: Started deploy [analytics/refinery@49c336c] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@49c336cd] * 15:29 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1324733{{!}}WikimediaAntiAbuse: Enable personal info tag display on enwiki (T431292)]] * 15:27 brennen@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324700{{!}}EventDetailsParticipantsModule: populate cache with non-local users (T434597)]] (duration: 06m 38s) * 15:23 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[2024,1031]*.wmnet: Set storage compatability to UPGRADING — [[phab:T433028|T433028]] - eevans@cumin1003 * 15:23 brennen@deploy1003: brennen, daimona: Continuing with deployment * 15:23 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:22 brennen@deploy1003: brennen, daimona: Backport for [[gerrit:1324700{{!}}EventDetailsParticipantsModule: populate cache with non-local users (T434597)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:22 cgoubert@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host rdb-lock2003.codfw.wmnet * 15:21 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 15:21 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 15:21 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:21 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:21 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:20 brennen@deploy1003: Started scap sync-world: Backport for [[gerrit:1324700{{!}}EventDetailsParticipantsModule: populate cache with non-local users (T434597)]] * 15:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1178.eqiad.wmnet with OS bookworm * 15:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock1003.eqiad.wmnet * 15:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock1003.eqiad.wmnet with OS trixie * 15:16 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 15:16 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 15:16 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 15:16 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:16 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:16 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:12 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 15:12 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2003.codfw.wmnet * 15:11 cgoubert@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host rdb-lock2003.codfw.wmnet * 15:11 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 15:11 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 15:11 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:11 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:11 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:07 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324338{{!}}InitialiseSettings: Enable 2FA warnings on more private wikis (T428103)]] (duration: 07m 02s) * 15:04 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 15:04 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock1003.eqiad.wmnet with reason: host reimage * 15:03 reedy@deploy1003: reedy: Continuing with deployment * 15:02 reedy@deploy1003: reedy: Backport for [[gerrit:1324338{{!}}InitialiseSettings: Enable 2FA warnings on more private wikis (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:02 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:00 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324338{{!}}InitialiseSettings: Enable 2FA warnings on more private wikis (T428103)]] * 14:57 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 14:57 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2003.codfw.wmnet * 14:57 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock1003.eqiad.wmnet with reason: host reimage * 14:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock2002.codfw.wmnet * 14:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock2002.codfw.wmnet with OS trixie * 14:56 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 14:56 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 14:56 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 14:55 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 14:55 moritzm: powercycle ganeti2046 * 14:47 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock1003.eqiad.wmnet with OS trixie * 14:46 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1003.eqiad.wmnet - cgoubert@cumin2003" * 14:46 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1003.eqiad.wmnet - cgoubert@cumin2003" * 14:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock1003.eqiad.wmnet on all recursors * 14:45 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock1003.eqiad.wmnet on all recursors * 14:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1003.eqiad.wmnet - cgoubert@cumin2003" * 14:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1272.eqiad.wmnet with reason: Enabling notifications * 14:44 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1003.eqiad.wmnet - cgoubert@cumin2003" * 14:44 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324719{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324720{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324722{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0 (T434187)]] (duration: 11m 02s) * 14:39 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 14:39 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock1003.eqiad.wmnet * 14:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock2002.codfw.wmnet with reason: host reimage * 14:37 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 14:37 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock1002.eqiad.wmnet * 14:37 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock1002.eqiad.wmnet with OS trixie * 14:37 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1324719{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324720{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324722{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0 (T434187)]] synced to the testservers (see https://wikitech. * 14:33 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock2002.codfw.wmnet with reason: host reimage * 14:33 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1324719{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324720{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324722{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0 (T434187)]] * 14:32 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1018.eqiad.wmnet with OS bookworm * 14:32 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2046.codfw.wmnet * 14:31 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1020.eqiad.wmnet with OS bookworm * 14:27 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2046.codfw.wmnet * 14:25 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2045.codfw.wmnet * 14:25 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2045.codfw.wmnet * 14:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock1002.eqiad.wmnet with reason: host reimage * 14:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1019.eqiad.wmnet with OS bookworm * 14:22 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Checking sanitization for wikis testwiki in section s3 * 14:20 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2045.codfw.wmnet * 14:18 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock1002.eqiad.wmnet with reason: host reimage * 14:17 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:16 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324707{{!}}Backport all changes from wmf/1.47.0-wmf.15]] (duration: 40m 51s) * 14:16 moritzm: installing Linux 6.1.180 on Bookworm hosts * 14:15 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock2002.codfw.wmnet with OS trixie * 14:14 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2002.codfw.wmnet - cgoubert@cumin2003" * 14:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2002.codfw.wmnet - cgoubert@cumin2003" * 14:14 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:14 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2045.codfw.wmnet * 14:14 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2002.codfw.wmnet on all recursors * 14:14 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2002.codfw.wmnet on all recursors * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2002.codfw.wmnet - cgoubert@cumin2003" * 14:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2002.codfw.wmnet - cgoubert@cumin2003" * 14:12 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2030.codfw.wmnet * 14:12 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2030.codfw.wmnet * 14:11 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on apifeatureusage2001.codfw.wmnet with reason: host reimage * 14:09 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:08 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:07 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 14:06 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2030.codfw.wmnet * 14:06 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:06 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock1002.eqiad.wmnet with OS trixie * 14:05 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 14:05 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2002.codfw.wmnet * 14:05 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1002.eqiad.wmnet - cgoubert@cumin2003" * 14:05 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1002.eqiad.wmnet - cgoubert@cumin2003" * 14:05 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock1002.eqiad.wmnet on all recursors * 14:05 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock1002.eqiad.wmnet on all recursors * 14:05 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:05 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1002.eqiad.wmnet - cgoubert@cumin2003" * 14:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:04 kharlan@deploy1003: kharlan: Continuing with deployment * 14:04 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock2001.codfw.wmnet * 14:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock2001.codfw.wmnet with OS trixie * 14:02 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1002.eqiad.wmnet - cgoubert@cumin2003" * 14:02 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on apifeatureusage2001.codfw.wmnet with reason: host reimage * 14:01 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2030.codfw.wmnet * 13:59 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2029.codfw.wmnet * 13:58 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2029.codfw.wmnet * 13:58 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 13:58 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock1002.eqiad.wmnet * 13:56 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1019.eqiad.wmnet with reason: host reimage * 13:54 btullis@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on archiva1002.wikimedia.org with reason: Upgrading in-place * 13:53 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock1001.eqiad.wmnet * 13:53 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock1001.eqiad.wmnet with OS trixie * 13:53 kharlan@deploy1003: kharlan: Backport for [[gerrit:1324707{{!}}Backport all changes from wmf/1.47.0-wmf.15]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:52 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2029.codfw.wmnet * 13:52 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1020.eqiad.wmnet with reason: host reimage * 13:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1019.eqiad.wmnet with reason: host reimage * 13:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1020.eqiad.wmnet with reason: host reimage * 13:49 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock2001.codfw.wmnet with reason: host reimage * 13:48 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2029.codfw.wmnet * 13:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Configuring db1272 for s3 pooling', diff saved to https://phabricator.wikimedia.org/P96021 and previous config saved to /var/cache/conftool/dbconfig/20260812-134732-cwilliams.json * 13:44 bking@cumin2003: START - Cookbook sre.hosts.reimage for host apifeatureusage2001.codfw.wmnet with OS bookworm * 13:43 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock2001.codfw.wmnet with reason: host reimage * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2028.codfw.wmnet * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2028.codfw.wmnet * 13:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1018.eqiad.wmnet with reason: host reimage * 13:38 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock1001.eqiad.wmnet with reason: host reimage * 13:36 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1018.eqiad.wmnet with reason: host reimage * 13:35 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1324707{{!}}Backport all changes from wmf/1.47.0-wmf.15]] * 13:35 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2028.codfw.wmnet * 13:32 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1275938{{!}}Enable campaignEvents on bdwikimedia (T424016)]] (duration: 07m 35s) * 13:32 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1020.eqiad.wmnet with OS bookworm * 13:32 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1019.eqiad.wmnet with OS bookworm * 13:32 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock1001.eqiad.wmnet with reason: host reimage * 13:28 kharlan@deploy1003: kharlan, yahya: Continuing with deployment * 13:27 kharlan@deploy1003: kharlan, yahya: Backport for [[gerrit:1275938{{!}}Enable campaignEvents on bdwikimedia (T424016)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:27 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2028.codfw.wmnet * 13:26 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 13:26 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 13:25 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1275938{{!}}Enable campaignEvents on bdwikimedia (T424016)]] * 13:25 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitize-wiki (exit_code=99) Managing sanitization for wikis testwiki in section s3 * 13:24 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock2001.codfw.wmnet with OS trixie * 13:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2001.codfw.wmnet - cgoubert@cumin2003" * 13:24 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2001.codfw.wmnet - cgoubert@cumin2003" * 13:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2001.codfw.wmnet on all recursors * 13:23 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2001.codfw.wmnet on all recursors * 13:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2001.codfw.wmnet - cgoubert@cumin2003" * 13:23 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2001.codfw.wmnet - cgoubert@cumin2003" * 13:23 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324713{{!}}thwiki: reinstate temporary wiki25 logos (T431094)]] (duration: 07m 13s) * 13:20 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:20 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1018.eqiad.wmnet with OS bookworm * 13:20 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2027.codfw.wmnet * 13:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2027.codfw.wmnet * 13:19 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics-external: apply * 13:18 kharlan@deploy1003: anzx, kharlan: Continuing with deployment * 13:18 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:18 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock1001.eqiad.wmnet with OS trixie * 13:17 kharlan@deploy1003: anzx, kharlan: Backport for [[gerrit:1324713{{!}}thwiki: reinstate temporary wiki25 logos (T431094)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1001.eqiad.wmnet - cgoubert@cumin2003" * 13:17 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 13:17 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1001.eqiad.wmnet - cgoubert@cumin2003" * 13:17 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2001.codfw.wmnet * 13:17 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics-external: apply * 13:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock1001.eqiad.wmnet on all recursors * 13:17 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock1001.eqiad.wmnet on all recursors * 13:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1001.eqiad.wmnet - cgoubert@cumin2003" * 13:17 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1001.eqiad.wmnet - cgoubert@cumin2003" * 13:15 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1324713{{!}}thwiki: reinstate temporary wiki25 logos (T431094)]] * 13:15 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:15 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics-external: apply * 13:15 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:14 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics-external: apply * 13:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2027.codfw.wmnet * 13:12 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 13:12 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock1001.eqiad.wmnet * 13:12 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2027.codfw.wmnet * 13:08 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1158.eqiad.wmnet onto db1273.eqiad.wmnet * 13:07 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1158: Pool db1158.eqiad.wmnet in after cloning * 13:02 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti6002.drmrs.wmnet * 13:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti6002.drmrs.wmnet * 12:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1015.eqiad.wmnet with OS bookworm * 12:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti6002.drmrs.wmnet * 12:44 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1017.eqiad.wmnet with OS bookworm * 12:36 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 12:35 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 12:34 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 12:33 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 12:31 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti6002.drmrs.wmnet * 12:24 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1017.eqiad.wmnet with reason: host reimage * 12:22 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1158: Pool db1158.eqiad.wmnet in after cloning * 12:18 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1017.eqiad.wmnet with reason: host reimage * 12:11 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1015.eqiad.wmnet with reason: host reimage * 12:07 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1015.eqiad.wmnet with reason: host reimage * 12:04 moritzm: failover ganeti master in drmrs02 to ganeti6004 * 12:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1009.eqiad.wmnet with OS bookworm * 12:01 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1017.eqiad.wmnet with OS bookworm * 12:00 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti6004.drmrs.wmnet * 12:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti6004.drmrs.wmnet * 11:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti6004.drmrs.wmnet * 11:53 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1015.eqiad.wmnet with OS bookworm * 11:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1159.eqiad.wmnet onto db1274.eqiad.wmnet * 11:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Pool db1159.eqiad.wmnet in after cloning * 11:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1016.eqiad.wmnet with OS bookworm * 11:46 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti6004.drmrs.wmnet * 11:45 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti6001.drmrs.wmnet * 11:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti6001.drmrs.wmnet * 11:43 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324306{{!}}WikimediaAntiAbuse: Enable personal info for enwiki with no display (T431292)]] (duration: 10m 26s) * 11:42 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1015.eqiad.wmnet with OS bookworm * 11:39 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 11:38 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti6001.drmrs.wmnet * 11:34 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1324306{{!}}WikimediaAntiAbuse: Enable personal info for enwiki with no display (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:33 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2212: Security update * 11:33 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti6001.drmrs.wmnet * 11:32 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1324306{{!}}WikimediaAntiAbuse: Enable personal info for enwiki with no display (T431292)]] * 11:22 moritzm: failover ganeti master in drmrs01 to ganeti6003 * 11:20 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:20 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:18 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 22 hosts with reason: Cloning * 11:17 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti6003.drmrs.wmnet * 11:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti6003.drmrs.wmnet * 11:17 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1009.eqiad.wmnet with reason: host reimage * 11:17 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:16 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:14 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1016.eqiad.wmnet with reason: host reimage * 11:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti6003.drmrs.wmnet * 11:11 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1009.eqiad.wmnet with reason: host reimage * 11:10 jmm@cumin2003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti3005.esams.wmnet * 11:10 jmm@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host ganeti3005.esams.wmnet * 11:09 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 11:08 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1158: Depool db1158.eqiad.wmnet to then clone it to db1273.eqiad.wmnet - marostegui@cumin1003 * 11:07 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1016.eqiad.wmnet with reason: host reimage * 11:07 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1158: Depool db1158.eqiad.wmnet to then clone it to db1273.eqiad.wmnet - marostegui@cumin1003 * 11:07 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1158.eqiad.wmnet onto db1273.eqiad.wmnet * 11:06 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:05 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:05 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1159: Pool db1159.eqiad.wmnet in after cloning * 11:04 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 20 hosts with reason: Cloning * 11:02 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti6003.drmrs.wmnet * 11:00 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis testwiki in section s3 * 10:54 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1009.eqiad.wmnet with OS bookworm * 10:51 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1015.eqiad.wmnet with OS bookworm * 10:50 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1016.eqiad.wmnet with OS bookworm * 10:48 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2212: Security update * 10:45 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitize-wiki (exit_code=99) Managing sanitization for wikis testwiki in section s3 * 10:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1013.eqiad.wmnet with OS bookworm * 10:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1014.eqiad.wmnet with OS bookworm * 10:38 cwilliams@cumin1003: START - Cookbook sre.mysql.clone of db1159.eqiad.wmnet onto db1274.eqiad.wmnet * 10:33 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db1274.eqiad.wmnet * 10:33 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db1274.eqiad.wmnet * 10:31 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 396993 * 10:29 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 396993 * 10:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1159: Clone source for db1274 * 10:24 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1159: Clone source for db1274 * 10:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1013.eqiad.wmnet with reason: host reimage * 10:18 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1013.eqiad.wmnet with reason: host reimage * 10:13 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2212.codfw.wmnet with reason: Maintenance * 10:13 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 10:12 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 10:12 blake@deploy1003: Stopping before sync operations * 10:11 blake@deploy1003: Started scap sync-world: Non-deployment scap run to populate new release values for [[phab:T427668|T427668]] * 10:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2212 [[phab:T434644|T434644]]', diff saved to https://phabricator.wikimedia.org/P96003 and previous config saved to /var/cache/conftool/dbconfig/20260812-101053-cwilliams.json * 10:09 moritzm: powercycle ganeti3005 * 10:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2203 to s1 primary [[phab:T434644|T434644]]', diff saved to https://phabricator.wikimedia.org/P96002 and previous config saved to /var/cache/conftool/dbconfig/20260812-100849-cwilliams.json * 10:08 cezmunsta: Starting s1 codfw failover from db2212 to db2203 - [[phab:T434644|T434644]] * 10:03 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1013.eqiad.wmnet with OS bookworm * 10:02 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1013.eqiad.wmnet with OS bookworm * 10:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2203 with weight 0 [[phab:T434644|T434644]]', diff saved to https://phabricator.wikimedia.org/P96001 and previous config saved to /var/cache/conftool/dbconfig/20260812-100134-cwilliams.json * 10:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 32 hosts with reason: Primary switchover s1 [[phab:T434644|T434644]] * 09:53 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1014.eqiad.wmnet with reason: host reimage * 09:50 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 09:50 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti3005.esams.wmnet * 09:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1014.eqiad.wmnet with reason: host reimage * 09:41 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1278: Pool in x1 * 09:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-launcher1003.eqiad.wmnet with OS bookworm * 09:37 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti3005.esams.wmnet * 09:37 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1013.eqiad.wmnet with OS bookworm * 09:34 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-presto1013.eqiad.wmnet with OS bookworm * 09:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1014.eqiad.wmnet with OS bookworm * 09:29 moritzm: failover ganeti master in esams to ganeti3008 * 09:26 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti3006.esams.wmnet * 09:26 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti3006.esams.wmnet * 09:24 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1012.eqiad.wmnet with OS bookworm * 09:23 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1009.eqiad.wmnet with OS bookworm * 09:23 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 09:18 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti3006.esams.wmnet * 09:16 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti3006.esams.wmnet * 09:14 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: fix regexp escaping bug - oblivian@cumin1003" * 09:14 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: fix regexp escaping bug - oblivian@cumin1003 * 09:13 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: fix regexp escaping bug - oblivian@cumin1003 * 09:13 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: fix regexp escaping bug - oblivian@cumin1003" * 09:03 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-launcher1003.eqiad.wmnet with reason: host reimage * 08:58 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-launcher1003.eqiad.wmnet with reason: host reimage * 08:55 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1278: Pool in x1 * 08:55 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1278 to dbctl [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95996 and previous config saved to /var/cache/conftool/dbconfig/20260812-085521-marostegui.json * 08:51 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1012.eqiad.wmnet with reason: host reimage * 08:45 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis testwiki in section s3 * 08:43 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1009.eqiad.wmnet with OS bookworm * 08:42 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1012.eqiad.wmnet with reason: host reimage * 08:41 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-launcher1003.eqiad.wmnet with OS bookworm * 08:40 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1013.eqiad.wmnet with OS bookworm * 08:38 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms3', diff saved to https://phabricator.wikimedia.org/P95995 and previous config saved to /var/cache/conftool/dbconfig/20260812-083816-marostegui.json * 08:38 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-master1003.eqiad.wmnet with OS bookworm * 08:37 marostegui: Failover ms3 [[phab:T434288|T434288]] * 08:37 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1268 to dbctl [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95994 and previous config saved to /var/cache/conftool/dbconfig/20260812-083722-marostegui.json * 08:35 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-web1001.eqiad.wmnet with OS bookworm * 08:32 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db2252.codfw.wmnet,db[1153,1268].eqiad.wmnet with reason: Switching over ms3 * 08:28 marostegui@cumin1003: dbctl commit (dc=all): 'Depool ms3 [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95993 and previous config saved to /var/cache/conftool/dbconfig/20260812-082852-marostegui.json * 08:25 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1012.eqiad.wmnet with OS bookworm * 08:25 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1011.eqiad.wmnet with OS bookworm * 08:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-master1003.eqiad.wmnet with reason: host reimage * 08:07 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-master1003.eqiad.wmnet with reason: host reimage * 08:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-web1001.eqiad.wmnet with reason: host reimage * 07:58 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-web1001.eqiad.wmnet with reason: host reimage * 07:50 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1003.eqiad.wmnet with OS bookworm * 07:38 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1011.eqiad.wmnet with reason: host reimage * 07:38 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-master1003.eqiad.wmnet with OS bookworm * 07:35 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti3007.esams.wmnet * 07:35 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti3007.esams.wmnet * 07:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1011.eqiad.wmnet with reason: host reimage * 07:27 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti3007.esams.wmnet * 07:25 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti3007.esams.wmnet * 07:25 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti3008.esams.wmnet * 07:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti3008.esams.wmnet * 07:22 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-web1001.eqiad.wmnet with OS bookworm * 07:19 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1003.eqiad.wmnet with OS bookworm * 07:18 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-master1003.eqiad.wmnet * 07:18 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host an-master1003.eqiad.wmnet * 07:17 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1011.eqiad.wmnet with OS bookworm * 07:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 07:16 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti3008.esams.wmnet * 07:15 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1009.eqiad.wmnet with OS bookworm * 07:14 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host an-master1003.eqiad.wmnet * 07:13 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-master1003.eqiad.wmnet * 07:13 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-master1003.eqiad.wmnet * 07:12 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-master1003.eqiad.wmnet * 07:11 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti3008.esams.wmnet * 07:07 arnaudb@dns1006: END - running authdns-update * 07:07 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti5007.eqsin.wmnet * 07:07 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti5007.eqsin.wmnet * 07:05 arnaudb@dns1006: START - running authdns-update * 06:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti5007.eqsin.wmnet * 06:54 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti5007.eqsin.wmnet * 06:38 moritzm: failover ganeti master in eqsin to ganeti5004 * 06:36 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti5006.eqsin.wmnet * 06:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti5006.eqsin.wmnet * 06:28 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti5006.eqsin.wmnet * 06:23 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti5006.eqsin.wmnet * 06:20 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti5005.eqsin.wmnet * 06:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti5005.eqsin.wmnet * 06:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti5005.eqsin.wmnet * 06:06 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti5005.eqsin.wmnet * 06:03 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti5004.eqsin.wmnet * 06:03 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti5004.eqsin.wmnet * 05:55 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti5004.eqsin.wmnet * 05:53 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti5004.eqsin.wmnet * 04:40 ryankemper: [[phab:T434494|T434494]] reimaged `an-tool1008.eqiad.wmnet` to bookworm; yarn.wikimedia.org is back up * 04:16 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-tool1008.eqiad.wmnet with OS bookworm * 03:58 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-tool1008.eqiad.wmnet with reason: host reimage * 03:53 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-tool1008.eqiad.wmnet with reason: host reimage * 03:41 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-tool1008.eqiad.wmnet with OS bookworm * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 45s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 00:25 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324427{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]], [[gerrit:1324429{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0]], [[gerrit:1324428{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]] (duration: 07m 55s) * 00:21 kemayo@deploy1003: kemayo: Continuing with deployment * 00:19 kemayo@deploy1003: kemayo: Backport for [[gerrit:1324427{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]], [[gerrit:1324429{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0]], [[gerrit:1324428{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:17 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1324427{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]], [[gerrit:1324429{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0]], [[gerrit:1324428{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]] == 2026-08-11 == * 21:37 sbassett: Deployed security fix for [[phab:T434521|T434521]] (wmf.15) * 21:29 sbassett: Deployed security fix for [[phab:T434521|T434521]] (wmf.14) * 21:19 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324370{{!}}Phase 4 of legal footer deployment (T432796)]], [[gerrit:1319804{{!}}Disable wgMFCustomSiteModules on English Wikipedia (T375538)]] (duration: 15m 26s) * 21:15 jdlrobson@deploy1003: jdlrobson: Continuing with deployment * 21:06 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1324370{{!}}Phase 4 of legal footer deployment (T432796)]], [[gerrit:1319804{{!}}Disable wgMFCustomSiteModules on English Wikipedia (T375538)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:03 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1324370{{!}}Phase 4 of legal footer deployment (T432796)]], [[gerrit:1319804{{!}}Disable wgMFCustomSiteModules on English Wikipedia (T375538)]] * 20:59 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 20:50 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324384{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]], [[gerrit:1324385{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]] (duration: 06m 58s) * 20:46 kemayo@deploy1003: kemayo: Continuing with deployment * 20:45 kemayo@deploy1003: kemayo: Backport for [[gerrit:1324384{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]], [[gerrit:1324385{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:43 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1324384{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]], [[gerrit:1324385{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]] * 20:42 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 20:42 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324386{{!}}build: Updating js-yaml to 3.15.1, 4.3.1]] (duration: 07m 36s) * 20:38 kemayo@deploy1003: kemayo: Continuing with deployment * 20:37 jhancock@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 20:37 kemayo@deploy1003: kemayo: Backport for [[gerrit:1324386{{!}}build: Updating js-yaml to 3.15.1, 4.3.1]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:35 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1324386{{!}}build: Updating js-yaml to 3.15.1, 4.3.1]] * 20:18 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 20:15 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 20:15 jhancock@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin1003" * 20:14 jhancock@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin1003" * 19:59 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 19:54 jhancock@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 19:10 brennen@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] (duration: 06m 41s) * 19:04 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns6001.wikimedia.org * 19:04 sukhe@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns6001.wikimedia.org * 19:04 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns5003.wikimedia.org * 19:04 sukhe@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns5003.wikimedia.org * 19:03 brennen@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 18:59 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns5003.wikimedia.org with OS trixie * 18:55 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns6001.wikimedia.org with OS trixie * 18:19 brett@cumin2002: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on P<nowiki>{</nowiki>cp7009.magru.wmnet<nowiki>}</nowiki> and A:cp - 9.2.15 Upgrade () * 18:14 brett@cumin2002: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on P<nowiki>{</nowiki>cp7009.magru.wmnet<nowiki>}</nowiki> and A:cp - 9.2.15 Upgrade () * 18:13 brennen@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 18:12 brett@cumin2002: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 9.2.15 Upgrade () * 18:09 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns5003.wikimedia.org with reason: host reimage * 18:06 brett@cumin2002: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 9.2.15 Upgrade () * 18:06 brennen: 1.47.0-wmf.15 train status ([[phab:T430834|T430834]]) - no current blockers, rolling to group0 * 18:05 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns5003.wikimedia.org with reason: host reimage * 18:05 brett: import trafficserver-9.2.15~deb13+wmf1 into trixie-wikimedia ([[phab:T434478|T434478]]) * 18:01 ladsgroup@cumin1003: END (PASS) - Cookbook sre.mysql.sanitarium_restart (exit_code=0) * 17:58 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns6001.wikimedia.org with reason: host reimage * 17:53 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324369{{!}}Enable desktop lazy loading on group0 (T148047)]] (duration: 07m 31s) * 17:52 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns6001.wikimedia.org with reason: host reimage * 17:49 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 17:49 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitarium_restart (exit_code=99) * 17:49 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 17:49 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7001.magru.wmnet * 17:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti7001.magru.wmnet * 17:48 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 17:47 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1324369{{!}}Enable desktop lazy loading on group0 (T148047)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:45 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1324369{{!}}Enable desktop lazy loading on group0 (T148047)]] * 17:39 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti7001.magru.wmnet * 17:36 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns5003.wikimedia.org with OS trixie * 17:34 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns6001.wikimedia.org with OS trixie * 17:31 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324361{{!}}Move FR config from IS.php to a dedicated file]], [[gerrit:1324363{{!}}Remove $wmg = $wg hacks in CentralAuth (T119117)]] (duration: 12m 23s) * 17:26 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 17:23 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1324361{{!}}Move FR config from IS.php to a dedicated file]], [[gerrit:1324363{{!}}Remove $wmg = $wg hacks in CentralAuth (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:19 sukhe: sudo cumin "A:cp-magru" "run-puppet-agent --enable 'merging CR 1324355'" * 17:18 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1324361{{!}}Move FR config from IS.php to a dedicated file]], [[gerrit:1324363{{!}}Remove $wmg = $wg hacks in CentralAuth (T119117)]] * 17:11 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-master1004.eqiad.wmnet with OS bookworm * 17:11 sukhe: sukhe@cp7005:~$ sudo puppet agent -tv * 17:02 sukhe: sudo cumin "A:cp-magru" "disable-puppet 'merging CR 1324355'" * 16:54 sukhe@dns1004: END - running authdns-update * 16:53 sukhe@dns1004: START - running authdns-update * 16:53 sukhe@dns1004: FAIL - running authdns-update * 16:51 sukhe@dns1004: START - running authdns-update * 16:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-master1004.eqiad.wmnet with reason: host reimage * 16:44 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-master1004.eqiad.wmnet with reason: host reimage * 16:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1157.eqiad.wmnet onto db1272.eqiad.wmnet * 16:40 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1157: Pool db1157.eqiad.wmnet in after cloning * 16:38 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 16:31 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324356{{!}}InitialiseSettings: Fix wgOATHAuthEnforce2FAForAll]] (duration: 06m 52s) * 16:30 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 16:28 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 16:27 reedy@deploy1003: reedy: Continuing with deployment * 16:26 reedy@deploy1003: reedy: Backport for [[gerrit:1324356{{!}}InitialiseSettings: Fix wgOATHAuthEnforce2FAForAll]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:24 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324356{{!}}InitialiseSettings: Fix wgOATHAuthEnforce2FAForAll]] * 16:13 sukhe: restart ntpsec.serviceon dns7001 * 16:09 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324335{{!}}InitialiseSettings: Enable 2FA enforcement on various private wikis (T428103)]] (duration: 06m 40s) * 16:08 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 16:06 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2204: Security update * 16:04 reedy@deploy1003: reedy: Continuing with deployment * 16:04 reedy@deploy1003: reedy: Backport for [[gerrit:1324335{{!}}InitialiseSettings: Enable 2FA enforcement on various private wikis (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:02 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7001.magru.wmnet * 16:02 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324335{{!}}InitialiseSettings: Enable 2FA enforcement on various private wikis (T428103)]] * 16:01 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 15:55 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1157: Pool db1157.eqiad.wmnet in after cloning * 15:54 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 15:54 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 15:41 dancy@deploy1003: Finished scap sync-world: Testing (duration: 06m 28s) * 15:40 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-master1004.eqiad.wmnet with OS bookworm * 15:40 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 15:35 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti4008.ulsfo.wmnet * 15:35 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti4008.ulsfo.wmnet * 15:34 dancy@deploy1003: Started scap sync-world: Testing * 15:34 dancy@deploy1003: Installation of scap version "4.279.0" completed for 3 hosts * 15:34 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-master1004.eqiad.wmnet with OS bookworm * 15:34 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 15:33 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-master1004.eqiad.wmnet with OS bookworm * 15:32 dancy@deploy1003: Installing scap version "4.279.0" for 3 host(s) * 15:32 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324339{{!}}Add /w/deployment-info.php entrypoint]] (duration: 07m 25s) * 15:30 moritzm: failover ganeti master in magru to ganeti7004 * 15:29 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti4008.ulsfo.wmnet * 15:28 dancy@deploy1003: dancy: Continuing with deployment * 15:28 tappof: remove 2026-05 swift log archives from centrallog to free some space ([[phab:T434502|T434502]]) * 15:27 dancy@deploy1003: dancy: Backport for [[gerrit:1324339{{!}}Add /w/deployment-info.php entrypoint]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:25 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324339{{!}}Add /w/deployment-info.php entrypoint]] * 15:20 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2204: Security update * 15:18 dancy@deploy1003: Installation of scap version "4.278.0" completed for 3 hosts * 15:18 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7004.magru.wmnet * 15:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti7004.magru.wmnet * 15:16 dancy@deploy1003: Installing scap version "4.278.0" for 3 host(s) * 15:14 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2204.codfw.wmnet with reason: Maintenance * 15:12 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 15:11 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-master1004.eqiad.wmnet with OS bookworm * 15:11 hashar: Restarting CI Jenkins on contint1003 due to Java upgrade. * 15:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2204 [[phab:T434565|T434565]]', diff saved to https://phabricator.wikimedia.org/P95984 and previous config saved to /var/cache/conftool/dbconfig/20260811-151126-cwilliams.json * 15:10 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti4008.ulsfo.wmnet * 15:10 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti7004.magru.wmnet * 15:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2207 to s2 primary [[phab:T434565|T434565]]', diff saved to https://phabricator.wikimedia.org/P95983 and previous config saved to /var/cache/conftool/dbconfig/20260811-150905-cwilliams.json * 15:08 cezmunsta: Starting s2 codfw failover from db2204 to db2207 - [[phab:T434565|T434565]] * 15:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2207 with weight 0 [[phab:T434565|T434565]]', diff saved to https://phabricator.wikimedia.org/P95982 and previous config saved to /var/cache/conftool/dbconfig/20260811-150402-cwilliams.json * 15:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s2 [[phab:T434565|T434565]] * 14:55 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1010.eqiad.wmnet with OS bookworm * 14:49 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns4003.wikimedia.org with OS trixie * 14:47 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1010.eqiad.wmnet with OS bookworm * 14:47 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7001.wikimedia.org with OS trixie * 14:44 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-presto1010.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:41 btullis@cumin1003: START - Cookbook sre.hosts.provision for host an-presto1010.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:40 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-presto1010.eqiad.wmnet with OS bookworm * 14:39 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 14:39 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-presto1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:36 btullis@cumin1003: START - Cookbook sre.hosts.provision for host an-presto1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:32 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1009.eqiad.wmnet with OS bookworm * 14:32 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 14:31 cwilliams@cumin1003: START - Cookbook sre.mysql.clone of db1157.eqiad.wmnet onto db1272.eqiad.wmnet * 14:30 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7004.magru.wmnet * 14:28 moritzm: failover ganeti master in ulsfo to ganeti4005 * 14:23 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7003.magru.wmnet * 14:23 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti7003.magru.wmnet * 14:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-master1004.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:22 btullis@cumin1003: START - Cookbook sre.hosts.provision for host an-master1004.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:21 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-master1004.eqiad.wmnet with OS bookworm * 14:19 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti4007.ulsfo.wmnet * 14:19 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti4007.ulsfo.wmnet * 14:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1010.eqiad.wmnet with OS bookworm * 14:17 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1008.eqiad.wmnet with OS bookworm * 14:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti7003.magru.wmnet * 14:12 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7003.magru.wmnet * 14:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti4007.ulsfo.wmnet * 14:11 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7002.magru.wmnet * 14:11 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti7002.magru.wmnet * 14:09 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324318{{!}}Revert "wmf-config/ProductionServices: set URL for urldownloader to service record" (T429175)]] (duration: 06m 46s) * 14:09 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7001.wikimedia.org with reason: host reimage * 14:06 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti4007.ulsfo.wmnet * 14:05 kharlan@deploy1003: kharlan: Continuing with deployment * 14:05 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti4006.ulsfo.wmnet * 14:04 jayme: updated calico to v3.30.7 on staging-eqiad - [[phab:T427400|T427400]] * 14:04 kharlan@deploy1003: kharlan: Backport for [[gerrit:1324318{{!}}Revert "wmf-config/ProductionServices: set URL for urldownloader to service record" (T429175)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti4006.ulsfo.wmnet * 14:03 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns4003.wikimedia.org with reason: host reimage * 14:03 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7001.wikimedia.org with reason: host reimage * 14:02 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti7002.magru.wmnet * 14:02 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1324318{{!}}Revert "wmf-config/ProductionServices: set URL for urldownloader to service record" (T429175)]] * 14:02 btullis@dns1004: FAIL - running authdns-update * 14:00 btullis@dns1004: START - running authdns-update * 13:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1008.eqiad.wmnet with reason: host reimage * 13:59 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'. * 13:59 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1313985{{!}}wmf-config/ProductionServices: set URL for urldownloader to service record (T429175)]] (duration: 25m 06s) * 13:58 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7002.magru.wmnet * 13:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti4006.ulsfo.wmnet * 13:57 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns4003.wikimedia.org with reason: host reimage * 13:56 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'. * 13:56 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7001.magru.wmnet * 13:56 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1008.eqiad.wmnet with reason: host reimage * 13:55 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'. * 13:55 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'. * 13:55 kharlan@deploy1003: kharlan, sukhe: Continuing with deployment * 13:53 marostegui: Failover ms2 [[phab:T434288|T434288]] * 13:52 marostegui: Failover ms1 [[phab:T434288|T434288]] * 13:52 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7001.magru.wmnet * 13:51 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti4006.ulsfo.wmnet * 13:48 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti4005.ulsfo.wmnet * 13:48 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti4005.ulsfo.wmnet * 13:44 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti4005.ulsfo.wmnet * 13:40 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1008.eqiad.wmnet with OS bookworm * 13:39 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns4003.wikimedia.org with OS trixie * 13:38 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns7001.wikimedia.org with OS trixie * 13:38 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1008.eqiad.wmnet with OS bookworm * 13:37 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti4005.ulsfo.wmnet * 13:36 kharlan@deploy1003: kharlan, sukhe: Backport for [[gerrit:1313985{{!}}wmf-config/ProductionServices: set URL for urldownloader to service record (T429175)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:34 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1313985{{!}}wmf-config/ProductionServices: set URL for urldownloader to service record (T429175)]] * 13:29 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2034.codfw.wmnet * 13:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2034.codfw.wmnet * 13:25 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.clone (exit_code=99) of db1157.eqiad.wmnet onto db1272.eqiad.wmnet * 13:25 cwilliams@cumin1003: START - Cookbook sre.mysql.clone of db1157.eqiad.wmnet onto db1272.eqiad.wmnet * 13:21 urbanecm@deploy1003: mwscript-k8s job started: namespaceDupes.php --wiki=frwiktionary --fix # [[phab:T415716|T415716]] * 13:21 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2034.codfw.wmnet * 13:20 urbanecm@deploy1003: mwscript-k8s job started: namespaceDupes.php --wiki=frwiktionary # [[phab:T415716|T415716]] * 13:19 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1323350{{!}}[tgwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T415307)]], [[gerrit:1322961{{!}}[slwiki] Revert temporary logo for Wikipedia 25 (Vector legacy + Vector 2022) (T414265)]], [[gerrit:1323827{{!}}[itwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T414320)]] (duration: 08m 00s) * 13:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-coord1003.eqiad.wmnet with OS bookworm * 13:15 urbanecm@deploy1003: urbanecm, superpes: Continuing with deployment * 13:13 urbanecm@deploy1003: urbanecm, superpes: Backport for [[gerrit:1323350{{!}}[tgwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T415307)]], [[gerrit:1322961{{!}}[slwiki] Revert temporary logo for Wikipedia 25 (Vector legacy + Vector 2022) (T414265)]], [[gerrit:1323827{{!}}[itwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T414320)]] synced to the testservers (see https://wiki * 13:11 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1323350{{!}}[tgwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T415307)]], [[gerrit:1322961{{!}}[slwiki] Revert temporary logo for Wikipedia 25 (Vector legacy + Vector 2022) (T414265)]], [[gerrit:1323827{{!}}[itwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T414320)]] * 13:11 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1323779{{!}}[ukwiki] Remove reviewer usergroup (T434252)]], [[gerrit:1323312{{!}}[frwiktionary] Add new Schème namespace and its talk (T415716)]] (duration: 06m 49s) * 13:10 marostegui@dns1004: END - running authdns-update * 13:08 marostegui@dns1004: START - running authdns-update * 13:07 marostegui@cumin1003: dbctl commit (dc=all): 'Repool ms2 [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95980 and previous config saved to /var/cache/conftool/dbconfig/20260811-130725-marostegui.json * 13:06 urbanecm@deploy1003: urbanecm, superpes: Continuing with deployment * 13:06 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1266 to dbctl [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95979 and previous config saved to /var/cache/conftool/dbconfig/20260811-130627-marostegui.json * 13:06 urbanecm@deploy1003: urbanecm, superpes: Backport for [[gerrit:1323779{{!}}[ukwiki] Remove reviewer usergroup (T434252)]], [[gerrit:1323312{{!}}[frwiktionary] Add new Schème namespace and its talk (T415716)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:04 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1323779{{!}}[ukwiki] Remove reviewer usergroup (T434252)]], [[gerrit:1323312{{!}}[frwiktionary] Add new Schème namespace and its talk (T415716)]] * 12:59 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db2253.codfw.wmnet,db[1151,1266].eqiad.wmnet with reason: Switching over ms2 * 12:54 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1157: Using as clone source * 12:53 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1157: Using as clone source * 12:51 marostegui@cumin1003: dbctl commit (dc=all): 'Depool ms2 [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95977 and previous config saved to /var/cache/conftool/dbconfig/20260811-125129-marostegui.json * 12:47 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 12:46 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 12:45 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 12:44 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'recommendation-api-ng' for release 'main' . * 12:44 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'recommendation-api-ng' for release 'main' . * 12:43 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'recommendation-api-ng' for release 'main' . * 12:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-coord1003.eqiad.wmnet with reason: host reimage * 12:43 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'ores-legacy' for release 'main' . * 12:42 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'ores-legacy' for release 'main' . * 12:42 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2165: Security update * 12:40 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-coord1003.eqiad.wmnet with reason: host reimage * 12:39 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'ores-legacy' for release 'main' . * 12:38 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' . * 12:38 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' . * 12:37 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' . * 12:34 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2034.codfw.wmnet * 12:30 jmm@cumin2002: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti-test2001.codfw.wmnet * 12:30 jmm@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host ganeti-test2001.codfw.wmnet * 12:25 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1179.eqiad.wmnet onto db1278.eqiad.wmnet * 12:25 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1179: Pool db1179.eqiad.wmnet in after cloning * 12:23 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-coord1003.eqiad.wmnet with OS bookworm * 12:22 moritzm: failover ganeti master in codfw/routed to ganeti2033 * 12:22 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2033.codfw.wmnet * 12:22 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2033.codfw.wmnet * 12:19 jmm@cumin2002: START - Cookbook sre.hosts.reboot-single for host ganeti-test2001.codfw.wmnet * 12:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 12:18 jmm@cumin2002: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti-test2001.codfw.wmnet * 12:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1008.eqiad.wmnet with OS bookworm * 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2033.codfw.wmnet * 12:07 moritzm: failover ganeti master in ganeti/test to ganeti-test2003 * 12:04 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 12:03 jmm@cumin2003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti4005.ulsfo.wmnet * 12:03 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti4005.ulsfo.wmnet * 12:00 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324283{{!}}Use maximum compression level in SqlBlobStore and SqlBagOStuff (T428377)]] (duration: 11m 37s) * 11:57 jmm@cumin2002: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti-test2002.codfw.wmnet * 11:57 jmm@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti-test2002.codfw.wmnet * 11:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2165: Security update * 11:54 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 11:52 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1324283{{!}}Use maximum compression level in SqlBlobStore and SqlBagOStuff (T428377)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:51 jmm@cumin2002: START - Cookbook sre.hosts.reboot-single for host ganeti-test2002.codfw.wmnet * 11:50 jmm@cumin2002: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti-test2002.codfw.wmnet * 11:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2165.codfw.wmnet with reason: Maintenance * 11:48 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@050d19e] (releasing): [[phab:T434186|T434186]] (duration: 01m 14s) * 11:48 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1324283{{!}}Use maximum compression level in SqlBlobStore and SqlBagOStuff (T428377)]] * 11:47 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@050d19e] (releasing): [[phab:T434186|T434186]] * 11:44 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@050d19e] (releasing): test jenkins deploy for [[phab:T434186|T434186]] (duration: 01m 08s) * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2165 [[phab:T434514|T434514]]', diff saved to https://phabricator.wikimedia.org/P95969 and previous config saved to /var/cache/conftool/dbconfig/20260811-114352-cwilliams.json * 11:43 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@050d19e] (releasing): test jenkins deploy for [[phab:T434186|T434186]] * 11:42 jmm@cumin2003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti-test2003.codfw.wmnet * 11:42 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti-test2003.codfw.wmnet * 11:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2161 to s8 primary [[phab:T434514|T434514]]', diff saved to https://phabricator.wikimedia.org/P95968 and previous config saved to /var/cache/conftool/dbconfig/20260811-114136-cwilliams.json * 11:40 cezmunsta: Starting s8 codfw failover from db2165 to db2161 - [[phab:T434514|T434514]] * 11:40 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1179: Pool db1179.eqiad.wmnet in after cloning * 11:36 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti-test2003.codfw.wmnet * 11:36 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti-test2003.codfw.wmnet * 11:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2161 with weight 0 [[phab:T434514|T434514]]', diff saved to https://phabricator.wikimedia.org/P95966 and previous config saved to /var/cache/conftool/dbconfig/20260811-113449-cwilliams.json * 11:34 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 25 hosts with reason: Primary switchover s8 [[phab:T434514|T434514]] * 11:29 moritzm: installing Python 3.11 security updates * 11:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-presto1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 11:26 btullis@cumin1003: START - Cookbook sre.hosts.provision for host an-presto1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 11:23 btullis@dns1004: END - running authdns-update * 11:21 btullis@dns1004: START - running authdns-update * 11:20 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-presto1008.eqiad.wmnet with OS bookworm * 11:20 moritzm: installing curl security updates * 11:11 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1007.eqiad.wmnet with OS bookworm * 10:45 tappof: bump space for prometheus k8s-dse in codfw * 10:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1007.eqiad.wmnet with reason: host reimage * 10:38 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1007.eqiad.wmnet with reason: host reimage * 10:37 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-coord1004.eqiad.wmnet with OS bookworm * 10:35 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1008.eqiad.wmnet with OS bookworm * 10:34 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1179: Depool db1179.eqiad.wmnet to then clone it to db1278.eqiad.wmnet - marostegui@cumin1003 * 10:34 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1008.eqiad.wmnet with OS bookworm * 10:25 fceratto@cumin1003: dbctl commit (dc=all): 'Remove db1177 [[phab:T433474|T433474]]', diff saved to https://phabricator.wikimedia.org/P95964 and previous config saved to /var/cache/conftool/dbconfig/20260811-102527-fceratto.json * 10:22 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1008.eqiad.wmnet with OS bookworm * 10:22 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1007.eqiad.wmnet with OS bookworm * 10:21 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1006.eqiad.wmnet with OS bookworm * 10:20 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 10:18 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1179: Depool db1179.eqiad.wmnet to then clone it to db1278.eqiad.wmnet - marostegui@cumin1003 * 10:18 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1179.eqiad.wmnet onto db1278.eqiad.wmnet * 10:17 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 10:17 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 10:14 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 10:09 blake@deploy1003: Stopping before sync operations * 10:09 blake@deploy1003: Started scap sync-world: Non-deployment run to populate release values for [[phab:T427668|T427668]] * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 10:04 fceratto@cumin1003: Removing db1177 from zarcillo [[phab:T433474|T433474]] * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1177.eqiad.wmnet * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1177.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:03 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1177.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:00 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1006.eqiad.wmnet with reason: host reimage * 09:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-coord1004.eqiad.wmnet with reason: host reimage * 09:57 marostegui: Failover m1 from db1164 to db1213 - [[phab:T434493|T434493]] * 09:57 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1006.eqiad.wmnet with reason: host reimage * 09:55 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:54 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2232].codfw.wmnet,db[1164,1213,1217].eqiad.wmnet with reason: Primary switchover m1 [[phab:T434493|T434493]] * 09:52 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-coord1004.eqiad.wmnet with reason: host reimage * 09:49 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1213.eqiad.wmnet with OS trixie * 09:49 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1177.eqiad.wmnet * 09:41 moritzm: installing Linux 6.12.101 on Trixie hosts * 09:40 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1006.eqiad.wmnet with OS bookworm * 09:35 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-coord1004.eqiad.wmnet with OS bookworm * 09:28 moritzm: installing node-tar security updates * 09:27 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1213.eqiad.wmnet with reason: host reimage * 09:22 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1213.eqiad.wmnet with reason: host reimage * 09:09 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1177: Decommission * 09:08 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db1177: Decommission * 09:08 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 09:08 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.decommission (exit_code=99) * 09:06 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1213.eqiad.wmnet with OS trixie * 09:06 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 09:05 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1213.eqiad.wmnet with reason: Reimage * 08:53 marostegui@dns1004: END - running authdns-update * 08:51 marostegui@dns1004: START - running authdns-update * 08:48 marostegui: Switchover ms1 master in eqiad [[phab:T434288|T434288]] * 08:48 marostegui@cumin1003: dbctl commit (dc=all): 'Repool ms1 [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95962 and previous config saved to /var/cache/conftool/dbconfig/20260811-084804-marostegui.json * 08:40 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1267 to dbctl [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95961 and previous config saved to /var/cache/conftool/dbconfig/20260811-084054-marostegui.json * 08:29 marostegui: Failover m1 from db1213 to db1164 - [[phab:T434043|T434043]] * 08:25 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2232].codfw.wmnet,db[1164,1213,1217].eqiad.wmnet with reason: Primary switchover m1 [[phab:T434043|T434043]] * 08:22 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db2251.codfw.wmnet,db[1152,1267].eqiad.wmnet with reason: Switching over ms1 * 08:22 marostegui@cumin1003: dbctl commit (dc=all): 'Depool ms1 [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95960 and previous config saved to /var/cache/conftool/dbconfig/20260811-082201-marostegui.json * 08:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: Switching over ms1 * 08:20 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.parsercache (exit_code=99) * 08:20 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 08:20 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1152: Switching over ms1 * 08:19 slyngshede@dns1004: END - running authdns-update * 08:18 moritzm: installing openjdk-21 security updates * 08:17 slyngshede@dns1004: START - running authdns-update * 08:16 moritzm: imported jenkins 2.568.2 to thirdparty/jenkins for trixie-wikimedia * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.12 (duration: 02m 26s) * 03:36 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] (duration: 33m 33s) * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 35s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-10 == * 14:54 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1323973{{!}}mmv.bootstrap: Fix getUrlParam to account for TIFF lossy/lossless param (T434333)]] (duration: 11m 24s) * 14:50 krinkle@deploy1003: krinkle: Continuing with deployment * 14:45 krinkle@deploy1003: krinkle: Backport for [[gerrit:1323973{{!}}mmv.bootstrap: Fix getUrlParam to account for TIFF lossy/lossless param (T434333)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:43 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1323973{{!}}mmv.bootstrap: Fix getUrlParam to account for TIFF lossy/lossless param (T434333)]] * 14:07 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1323967{{!}}updateIsActiveFlagForMentees: Commit the final partial batch (T432959)]] (duration: 10m 33s) * 13:56 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1323967{{!}}updateIsActiveFlagForMentees: Commit the final partial batch (T432959)]] * 13:45 dani@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply * 13:45 dani@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply * 13:45 dani@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply * 13:45 dani@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply * 13:45 dani@deploy1003: helmfile [staging] DONE helmfile.d/services/miscweb: apply * 13:44 dani@deploy1003: helmfile [staging] START helmfile.d/services/miscweb: apply * 13:38 wmde-fisch@deploy1003: Finished scap sync-world: Backport for [[gerrit:1323939{{!}}Enable sub-references on more group2 wikis (batch3) (T432731)]] (duration: 33m 21s) * 13:25 wmde-fisch@deploy1003: wmde-fisch: Continuing with deployment * 13:22 wmde-fisch@deploy1003: wmde-fisch: Backport for [[gerrit:1323939{{!}}Enable sub-references on more group2 wikis (batch3) (T432731)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:05 wmde-fisch@deploy1003: Started scap sync-world: Backport for [[gerrit:1323939{{!}}Enable sub-references on more group2 wikis (batch3) (T432731)]] * 07:57 hashar@deploy1003: Finished deploy [integration/docroot@7772132]: update build dependencies (duration: 00m 13s) * 07:57 hashar@deploy1003: Started deploy [integration/docroot@7772132]: update build dependencies * 07:35 _joe_: restarting squid on urldownloader1006 * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 48s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-09 == * 16:01 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:01 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:01 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:00 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 36s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-08 == * 05:31 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9] (wcqs): [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) (duration: 02m 36s) * 05:28 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9] (wcqs): [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) * 04:56 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 04:55 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 04:47 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) (duration: 19m 22s) * 04:28 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) * 04:19 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) (duration: 00m 06s) * 04:18 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) * 04:17 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) (duration: 00m 28s) * 04:16 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) * 03:52 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 03:52 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 34s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-07 == * 23:30 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:29 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 22:45 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 22:43 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 22:41 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 22:41 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 20:54 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 20:32 andrewbogott: restarting puppetserver service on puppetserver* for [[phab:T434339|T434339]] * 19:52 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:45 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 19:32 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:25 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:22 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 19:21 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 18:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:41 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 18:35 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 18:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 18:22 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 18:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 18:16 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 18:12 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:09 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:08 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:07 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:04 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:01 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:00 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:00 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 17:59 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 17:25 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 17:14 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 17:13 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 17:13 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 17:13 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:54 maryum: Deployed security fix for [[phab:T434278|T434278]] * 16:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 16:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 16:27 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --verbose --scoreLessThan=0.7 --exceptDatasetChecksums=[[phab:T434319|T434319]]-enwiki-models.txt # [[phab:T434319|T434319]] * 16:06 cdobbins@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp5022.eqsin.wmnet with OS trixie * 15:13 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 14:19 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 14:17 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 13:50 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1156.eqiad.wmnet onto db1271.eqiad.wmnet * 13:50 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1271: Pool db1271.eqiad.wmnet in after cloning * 13:02 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1271: Pool db1271.eqiad.wmnet in after cloning * 12:19 jayme: updated calico to v3.30.7 on staging-codfw - [[phab:T427400|T427400]] * 12:09 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 12:06 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 12:06 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 12:05 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 12:02 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1156: Pool db1156.eqiad.wmnet in after cloning * 11:38 bjensen: sudo -i reprepro -C main include trixie-wikimedia $<nowiki>{</nowiki>HOME<nowiki>}</nowiki>/httpbb/trixie/httpbb_$<nowiki>{</nowiki>VERSION?<nowiki>}</nowiki>-1+deb13u1_amd64.changes #[[phab:T434052|T434052]] * 11:35 bjensen: sudo -i reprepro -C main include bookworm-wikimedia $<nowiki>{</nowiki>HOME<nowiki>}</nowiki>/httpbb/bookworm/httpbb_$<nowiki>{</nowiki>VERSION?<nowiki>}</nowiki>-1_amd64.changes #[[phab:T434052|T434052]] * 11:30 marostegui@cumin1003: dbctl commit (dc=all): 'Adding db1271 to dbctl', diff saved to https://phabricator.wikimedia.org/P95945 and previous config saved to /var/cache/conftool/dbconfig/20260807-113006-marostegui.json * 11:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1156: Pool db1156.eqiad.wmnet in after cloning * 10:23 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 10:22 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 10:22 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 10:21 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 10:20 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 10:20 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 10:19 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 10:18 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 10:06 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on 21 hosts with reason: cloning * 10:01 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1156: Depool db1156.eqiad.wmnet to then clone it to db1271.eqiad.wmnet - marostegui@cumin1003 * 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1156: Depool db1156.eqiad.wmnet to then clone it to db1271.eqiad.wmnet - marostegui@cumin1003 * 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1156.eqiad.wmnet onto db1271.eqiad.wmnet * 09:15 jynus: started stress testing db1245 dbs [[phab:T431115|T431115]] * 08:19 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:18 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:16 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:14 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:13 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:10 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:06 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:05 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:00 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 10 days, 0:00:00 on ml-serve1015.eqiad.wmnet with reason: Downtime to get full picture of current BIOS settings beyond what Redfish shows * 08:00 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 07:54 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 07:54 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 07:53 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:52 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:51 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:50 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:49 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:48 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:47 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:45 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:45 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:41 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:38 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:37 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 06:35 jayme: updated istio to 1.29.4 on wikikube eqiad - [[phab:T427401|T427401]] * 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1178.eqiad.wmnet * 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1178.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 06:06 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1178.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 05:55 marostegui@cumin1003: START - Cookbook sre.dns.netbox * 05:49 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1178.eqiad.wmnet * 05:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 05:46 marostegui@cumin1003: Removing db1178 from zarcillo [[phab:T433471|T433471]] * 05:45 marostegui@cumin1003: START - Cookbook sre.mysql.decommission * 02:42 denisse: Extended volume on prometheus2008 for the disk space alert as per https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_host_running_out_of_space * 02:37 denisse: Extended volume on prometheus2007 tor the disk space alert as per https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_host_running_out_of_space * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 56s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-06 == * 21:39 maryum: Deploy security patch for [[phab:T433070|T433070]] * 21:29 maryum: Deploy security patch for [[phab:T434189|T434189]] * 20:48 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] (duration: 08m 12s) * 20:44 aude@deploy1003: lmora, aude, anzx: Continuing with deployment * 20:41 aude@deploy1003: lmora, aude, anzx: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be * 20:41 ebernhardson@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:41 ebernhardson@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 20:40 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] * 20:37 ebernhardson@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:37 ebernhardson@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 20:32 ebernhardson@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:32 ebernhardson@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 20:31 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] (duration: 06m 41s) * 20:27 cjming@deploy1003: cjming, ebernhardson, chlod: Continuing with deployment * 20:26 cjming@deploy1003: cjming, ebernhardson, chlod: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:24 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] * 20:18 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] (duration: 09m 22s) * 20:14 cjming@deploy1003: cjming, tsev: Continuing with deployment * 20:11 cjming@deploy1003: cjming, tsev: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:09 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] * 19:41 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply * 19:40 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply * 19:31 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 19:31 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 19:00 cdobbins@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cp5022.eqsin.wmnet with OS trixie * 18:25 ladsgroup@deploy1003: Finished scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) (duration: 06m 08s) * 18:19 ladsgroup@deploy1003: Started scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) * 18:18 ladsgroup@deploy1003: Stopping before sync operations * 18:17 ladsgroup@deploy1003: Started scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) * 17:55 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 16:50 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 16:35 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1001.eqiad.wmnet with OS bookworm * 16:19 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1002.eqiad.wmnet with reason: host reimage * 16:16 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1002.eqiad.wmnet with reason: host reimage * 16:05 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1001.eqiad.wmnet with reason: host reimage * 16:00 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1001.eqiad.wmnet with reason: host reimage * 15:57 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 15:43 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm * 15:29 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1001.eqiad.wmnet with OS bookworm * 15:29 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:58 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm * 14:57 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-drmrs ([[phab:T428495|T428495]]) * 14:55 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-drmrs ([[phab:T428495|T428495]]) * 14:55 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-ui1001.eqiad.wmnet with OS bookworm * 14:54 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-presto1001.eqiad.wmnet with OS bookworm * 14:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-magru ([[phab:T428495|T428495]]) * 14:49 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-magru ([[phab:T428495|T428495]]) * 14:48 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1001.eqiad.wmnet with OS bookworm * 14:46 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-esams ([[phab:T428495|T428495]]) * 14:44 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-esams ([[phab:T428495|T428495]]) * 14:43 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:42 brouberol@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:42 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:42 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 14:40 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 14:40 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:38 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-ui1001.eqiad.wmnet with reason: host reimage * 14:34 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-presto1001.eqiad.wmnet with reason: host reimage * 14:28 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-ui1001.eqiad.wmnet with reason: host reimage * 14:27 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-presto1001.eqiad.wmnet with reason: host reimage * 14:23 sukhe: sudo cumin -b2 'A:cp-text' "run-puppet-agent --enable 'merging CR 1290731'": [[phab:T425441|T425441]] * 14:18 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo for hosts in the wikimedia.org domain - [[phab:T428495|T428495]] * 14:16 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-presto1001.eqiad.wmnet with OS bookworm * 14:14 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-ui1001.eqiad.wmnet with OS bookworm * 14:12 sukhe: sudo cumin 'A:cp-text' "disable-puppet 'merging CR 1290731'": [[phab:T425441|T425441]] * 14:11 swfrench-wmf: restarted navtiming on webperf1003 - [[phab:T428495|T428495]] * 14:04 swfrench-wmf: begin rolling restart of confd in drmrs, eqiad, esams, magru - [[phab:T428495|T428495]] * 14:04 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm * 14:04 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:02 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-client1002.eqiad.wmnet with OS bookworm * 13:58 swfrench-wmf: authdns update to direct eqiad-associated etcd clients back to eqiad - [[phab:T428495|T428495]] * 13:58 swfrench@dns1004: END - running authdns-update * 13:56 swfrench@dns1004: START - running authdns-update * 13:49 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:44 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:31 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 13:29 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 13:26 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 13:23 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 13:22 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 13:19 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 13:18 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 13:18 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 13:17 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 13:16 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 13:13 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 13:11 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 13:09 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 13:06 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 13:06 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-client1002.eqiad.wmnet with OS bookworm * 13:05 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revision-models' for release 'main' . * 13:05 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:05 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revision-models' for release 'main' . * 13:04 brouberol@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-test-client1002.eqiad.wmnet with OS bookworm * 13:04 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 13:03 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 13:02 aikochou@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:00 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'readability' for release 'main' . * 12:59 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'readability' for release 'main' . * 12:58 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 12:57 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'logo-detection' for release 'main' . * 12:57 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'logo-detection' for release 'main' . * 12:57 aikochou@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 12:55 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:54 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:53 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 12:53 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:50 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 12:48 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 12:46 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'article-models' for release 'main' . * 12:45 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'article-models' for release 'main' . * 12:41 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . * 12:39 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . * 12:38 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-client1002.eqiad.wmnet with OS bookworm * 12:12 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply * 12:12 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply * 12:09 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:08 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 11:58 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2187: Security update * 11:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:24 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:16 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 11:15 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 11:10 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2187: Security update * 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2187.codfw.wmnet with reason: Maintenance * 10:56 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 10:56 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 10:56 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 10:56 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 10:54 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 10:53 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 10:09 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2187: Security update * 10:07 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2187: Security update * 09:39 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms2', diff saved to https://phabricator.wikimedia.org/P95929 and previous config saved to /var/cache/conftool/dbconfig/20260806-093908-marostegui.json * 09:36 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1178 from dbctl [[phab:T433471|T433471]]', diff saved to https://phabricator.wikimedia.org/P95928 and previous config saved to /var/cache/conftool/dbconfig/20260806-093632-marostegui.json * 09:33 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 09:31 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 09:30 topranks: bounce cr3-eqsin<->cr2-eqiad bgp session to disable no-prepend command * 09:20 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2253.codfw.wmnet,db1151.eqiad.wmnet with reason: cloning * 09:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1151: Cloning * 09:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:19 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 09:19 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1151: Cloning * 09:10 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 09:09 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup2003.codfw.wmnet * 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup2003.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 09:06 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup2003.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 09:03 klausman@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:02 jynus@cumin1003: START - Cookbook sre.dns.netbox * 09:02 klausman@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 08:57 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup2003.codfw.wmnet * 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup1003.eqiad.wmnet * 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:54 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms3', diff saved to https://phabricator.wikimedia.org/P95925 and previous config saved to /var/cache/conftool/dbconfig/20260806-085422-marostegui.json * 08:53 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:46 jynus@cumin1003: START - Cookbook sre.dns.netbox * 08:39 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup1003.eqiad.wmnet * 08:29 XioNoX: push pfw policy - [[phab:T434115|T434115]] * 08:14 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 08:00 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 08:00 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:58 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revision-models' for release 'main' . * 07:56 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 07:54 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'readability' for release 'main' . * 07:53 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'logo-detection' for release 'main' . * 07:51 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'llm' for release 'main' . * 07:48 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . * 07:37 jayme: updated istio to 1.29.4 on wikikube codfw - [[phab:T427401|T427401]] * 07:08 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2252.codfw.wmnet,db1153.eqiad.wmnet with reason: cloning * 07:07 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1153: Cloning * 07:07 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1153: Cloning * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 40s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-05 == * 23:24 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1009.eqiad.wmnet with OS bookworm * 23:03 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1009.eqiad.wmnet with reason: host reimage * 22:59 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1009.eqiad.wmnet with reason: host reimage * 22:43 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1009.eqiad.wmnet with OS bookworm * 22:38 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1009.eqiad.wmnet * 22:34 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1009.eqiad.wmnet * 22:25 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1008.eqiad.wmnet with OS bookworm * 22:04 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1008.eqiad.wmnet with reason: host reimage * 22:00 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1008.eqiad.wmnet with reason: host reimage * 21:48 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:47 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:46 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:44 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1008.eqiad.wmnet with OS bookworm * 21:43 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:41 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1008.eqiad.wmnet * 21:36 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1008.eqiad.wmnet * 21:14 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:12 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad * 21:12 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad * 21:10 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=eqiad * 21:08 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:07 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:07 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=eqiad * 21:04 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:03 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1006 * 21:02 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1006 * 21:00 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:56 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:56 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:55 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 20:55 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 20:55 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1007.eqiad.wmnet with OS bookworm * 20:51 vriley@cumin1003: START - Cookbook sre.dns.netbox * 20:43 ebernhardson: [[phab:T434008|T434008]]: changing cloudelastic:9643 from auto_expand_replicas to number_of_replicas * 20:34 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1007.eqiad.wmnet with reason: host reimage * 20:27 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1007.eqiad.wmnet with reason: host reimage * 20:24 cjming: end of UTC late backport window * 20:23 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] (duration: 06m 26s) * 20:18 cjming@deploy1003: cjming: Continuing with deployment * 20:18 cjming@deploy1003: cjming: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:16 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] * 20:12 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1007.eqiad.wmnet with OS bookworm * 20:12 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] (duration: 08m 41s) * 20:08 swfrench@cumin2002: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host conf1007.eqiad.wmnet with OS bookworm * 20:08 jforrester@deploy1003: jforrester: Continuing with deployment * 20:07 jforrester@deploy1003: jforrester: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:03 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] * 19:51 inflatador: [bking@puppetserver1001] ~$ sudo puppetserver ca sign --certname an-worker1189.eqiad.wmnet [[phab:T434142|T434142]] * 19:47 bking@cumin2003: DONE (FAIL) - Cookbook sre.puppet.renew-cert (exit_code=99) for an-worker1189.eqiad.wmnet: Renew puppet certificate - bking@cumin2003 * 19:46 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:30 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1007.eqiad.wmnet with OS trixie * 19:30 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 19:29 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 19:20 swfrench-wmf: silenced EtcdRelicationDown 0cb709a9-f244-4f1e-971f-{{Gerrit|440ec65e7fd7}} - [[phab:T428495|T428495]] * 19:13 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1007.eqiad.wmnet with OS bookworm * 19:12 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1007.eqiad.wmnet with reason: host reimage * 19:09 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1007.eqiad.wmnet * 19:07 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1007.eqiad.wmnet with reason: host reimage * 19:03 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1007.eqiad.wmnet * 18:52 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1007.eqiad.wmnet with OS trixie * 18:52 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1007.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:35 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 18:34 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 18:34 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 18:30 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1007.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:28 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:28 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1007] - vriley@cumin1003" * 18:27 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1007] - vriley@cumin1003" * 18:23 vriley@cumin1003: START - Cookbook sre.dns.netbox * 18:22 vriley@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 18:22 robh@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:19 vriley@cumin1003: START - Cookbook sre.dns.netbox * 18:13 robh@cumin2002: START - Cookbook sre.hosts.provision for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:31 jasmine@cumin2002: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-main-eqiad * 17:12 mutante: LDAP - added vwalters to group ciadmin - [[phab:T433615|T433615]] * 16:58 aokoth@deploy1003: Finished deploy [phabricator/deployment@e2ebca5]: Deploy Phab (duration: 00m 34s) * 16:57 aokoth@deploy1003: Started deploy [phabricator/deployment@e2ebca5]: Deploy Phab * 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad * 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=eqiad * 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad * 16:53 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:41 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2187.codfw.wmnet * 16:41 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2187.codfw.wmnet * 16:41 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker2187.codfw.wmnet * 16:41 cgoubert@cumin2003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker2187.codfw.wmnet * 16:40 jasmine@cumin2002: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-main-eqiad * 16:40 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:34 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:25 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-magru and A:liberica ([[phab:T428495|T428495]]) * 16:23 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-magru and A:liberica ([[phab:T428495|T428495]]) * 16:20 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-drmrs and A:liberica ([[phab:T428495|T428495]]) * 16:19 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-drmrs and A:liberica ([[phab:T428495|T428495]]) * 16:18 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-esams and A:liberica ([[phab:T428495|T428495]]) * 16:16 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-esams and A:liberica ([[phab:T428495|T428495]]) * 16:06 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1159.eqiad.wmnet * 16:06 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1159.eqiad.wmnet * 16:06 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1159.eqiad.wmnet * 16:05 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] (duration: 09m 11s) * 15:58 reedy@deploy1003: reedy: Continuing with deployment * 15:58 reedy@deploy1003: reedy: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:56 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] * 15:54 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1159.eqiad.wmnet with OS trixie * 15:38 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:33 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1159.eqiad.wmnet with reason: host reimage * 15:32 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:27 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1159.eqiad.wmnet with reason: host reimage * 15:10 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1159 * 15:10 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1159 * 15:00 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo for hosts in the wikimedia.org domain - [[phab:T428495|T428495]] * 14:55 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS trixie * 14:54 swfrench-wmf: restarted navtiming on webperf1003 - [[phab:T428495|T428495]] * 14:52 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1159 * 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1159.eqiad.wmnet 129.48.64.10.in-addr.arpa 9.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:52 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1159.eqiad.wmnet 129.48.64.10.in-addr.arpa 9.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1159 - jayme@cumin1003" * 14:52 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1159 - jayme@cumin1003" * 14:49 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:48 jayme@cumin1003: START - Cookbook sre.dns.netbox * 14:47 swfrench-wmf: begin rolling restart of confd in drmrs, eqiad, esams, magru - [[phab:T428495|T428495]] * 14:47 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1159 * 14:46 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:46 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:46 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1159.eqiad.wmnet with OS trixie * 14:45 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:44 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:44 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1159.eqiad.wmnet * 14:43 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:43 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1159.eqiad.wmnet * 14:43 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:43 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1159.eqiad.wmnet * 14:43 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:43 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1157.eqiad.wmnet * 14:43 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1157.eqiad.wmnet * 14:43 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1157.eqiad.wmnet * 14:42 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:42 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:42 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:41 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:41 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:41 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:41 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1003.eqiad.wmnet with OS bookworm * 14:39 swfrench-wmf: authdns update to direct eqiad-associated etcd clients to codfw - [[phab:T428495|T428495]] * 14:39 swfrench@dns1004: END - running authdns-update * 14:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 14:37 swfrench@dns1004: START - running authdns-update * 14:37 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:37 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:35 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:35 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 14:28 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:27 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1157.eqiad.wmnet with OS trixie * 14:27 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:27 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:27 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 14:26 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 14:26 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:26 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 14:26 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 14:26 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:26 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host search-loader1002.eqiad.wmnet with OS trixie * 14:19 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:19 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:15 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1003.eqiad.wmnet with reason: host reimage * 14:14 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:14 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:13 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 * 14:13 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host mc2046 * 14:13 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS trixie * 14:11 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1003.eqiad.wmnet with reason: host reimage * 14:10 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:09 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:09 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:09 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:08 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:08 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1157.eqiad.wmnet with reason: host reimage * 14:08 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:04 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 14:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on search-loader1002.eqiad.wmnet with reason: host reimage * 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=eqiad * 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=eqiad * 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=eqiad * 14:00 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:59 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:58 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1157.eqiad.wmnet with reason: host reimage * 13:57 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on search-loader1002.eqiad.wmnet with reason: host reimage * 13:54 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1003.eqiad.wmnet with OS bookworm * 13:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host search-loader1002.eqiad.wmnet with OS trixie * 13:43 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1157 * 13:42 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1157 * 13:40 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1157 * 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1157.eqiad.wmnet 183.32.64.10.in-addr.arpa 3.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1157.eqiad.wmnet 183.32.64.10.in-addr.arpa 3.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1157 - jayme@cumin1003" * 13:39 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1157 - jayme@cumin1003" * 13:39 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] (duration: 07m 00s) * 13:35 jayme@cumin1003: START - Cookbook sre.dns.netbox * 13:35 reedy@deploy1003: reedy: Continuing with deployment * 13:34 reedy@deploy1003: reedy: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:32 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] * 13:23 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1157 * 13:22 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1157.eqiad.wmnet with OS trixie * 13:22 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1157.eqiad.wmnet * 13:22 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1157.eqiad.wmnet * 13:21 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1157.eqiad.wmnet * 13:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1156.eqiad.wmnet * 13:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1156.eqiad.wmnet * 13:15 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1156.eqiad.wmnet * 13:01 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1156.eqiad.wmnet with OS trixie * 12:42 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1156.eqiad.wmnet with reason: host reimage * 12:38 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1156.eqiad.wmnet with reason: host reimage * 12:32 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 12:31 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 12:30 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 12:28 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 12:26 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 12:24 topranks: update bgp confed settings in eqsin * 12:22 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1156 * 12:22 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1156 * 12:22 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 12:19 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1156 * 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1156.eqiad.wmnet 110.32.64.10.in-addr.arpa 0.1.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:19 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1156.eqiad.wmnet 110.32.64.10.in-addr.arpa 0.1.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1156 - jayme@cumin1003" * 12:19 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1156 - jayme@cumin1003" * 12:17 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:14 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS trixie * 12:09 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:06 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:04 jayme@cumin1003: START - Cookbook sre.dns.netbox * 12:04 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 12:02 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'article-models' for release 'main' . * 12:01 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1156 * 12:01 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1156.eqiad.wmnet with OS trixie * 11:59 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1156.eqiad.wmnet * 11:59 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1156.eqiad.wmnet * 11:59 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1156.eqiad.wmnet * 11:57 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 11:53 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 11:53 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:52 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:52 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:50 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:50 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:50 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:49 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:48 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:47 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:47 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:45 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:45 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:44 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:44 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:44 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:43 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:42 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:38 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:35 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 * 11:35 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host mc2046 * 11:34 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS trixie * 11:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:27 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:21 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:21 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:18 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:18 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:18 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:18 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:13 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:13 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:09 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:08 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:07 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:06 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:06 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:05 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:05 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:04 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 11:04 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:24 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:24 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:17 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:16 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1155.eqiad.wmnet * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1155.eqiad.wmnet * 10:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1155.eqiad.wmnet * 10:14 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:14 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:11 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:11 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:10 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:09 aikochou@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop: sync * 10:09 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:09 aikochou@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop: sync * 10:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:07 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:05 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:05 aikochou@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop: sync * 10:05 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:05 aikochou@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop: sync * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:04 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:04 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:04 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1155.eqiad.wmnet with OS trixie * 09:52 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms1', diff saved to https://phabricator.wikimedia.org/P95918 and previous config saved to /var/cache/conftool/dbconfig/20260805-095212-marostegui.json * 09:44 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1152: after cloning * 09:44 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.parsercache (exit_code=99) * 09:44 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 09:44 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1152: after cloning * 09:43 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1155.eqiad.wmnet with reason: host reimage * 09:40 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1155.eqiad.wmnet with reason: host reimage * 09:32 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 09:32 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:31 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 09:31 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:27 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1155 * 09:27 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1155 * 09:25 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 09:24 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 09:24 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 09:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:23 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 09:23 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 09:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:22 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 09:22 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 09:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:20 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 09:20 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:17 XioNoX: push pfw policies - [[phab:T434038|T434038]] * 09:14 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1155 * 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1155.eqiad.wmnet 109.32.64.10.in-addr.arpa 9.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1155.eqiad.wmnet 109.32.64.10.in-addr.arpa 9.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1155 - jayme@cumin1003" * 09:14 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1155 - jayme@cumin1003" * 09:10 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 09:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:09 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2251.codfw.wmnet,db1152.eqiad.wmnet with reason: cloning * 09:09 jayme@cumin1003: START - Cookbook sre.dns.netbox * 09:08 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 09:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: Cloning * 09:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:05 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 09:05 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1152: Cloning * 08:38 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1155 * 08:37 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1155.eqiad.wmnet with OS trixie * 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1171.eqiad.wmnet * 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1171.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:29 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95913 and previous config saved to /var/cache/conftool/dbconfig/20260805-082908-ladsgroup.json * 08:27 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1171.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:22 jynus@cumin1003: START - Cookbook sre.dns.netbox * 08:18 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249', diff saved to https://phabricator.wikimedia.org/P95912 and previous config saved to /var/cache/conftool/dbconfig/20260805-081823-ladsgroup.json * 08:17 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1171.eqiad.wmnet * 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1150.eqiad.wmnet * 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1150.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:15 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1150.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:15 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 08:14 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1155.eqiad.wmnet * 08:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1155.eqiad.wmnet * 08:14 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1155.eqiad.wmnet * 08:11 jynus@cumin1003: START - Cookbook sre.dns.netbox * 08:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249', diff saved to https://phabricator.wikimedia.org/P95911 and previous config saved to /var/cache/conftool/dbconfig/20260805-080737-ladsgroup.json * 08:05 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1150.eqiad.wmnet * 08:02 marostegui: Depool clouddb1020 (s5,s8) [[phab:T434048|T434048]] * 08:02 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1020.eqiad.wmnet,service=s8 * 08:02 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1020.eqiad.wmnet,service=s5 * 08:02 marostegui: Depool clouddb1018 (s2,s7) [[phab:T434048|T434048]] * 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1018.eqiad.wmnet,service=s7 * 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1018.eqiad.wmnet,service=s2 * 08:01 marostegui: Depool clouddb1017 (s1) [[phab:T434048|T434048]] * 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1017.eqiad.wmnet,service=s1 * 07:59 marostegui: Depool clouddb1016 (s5,s8) [[phab:T434048|T434048]] * 07:59 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s8 * 07:59 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s5 * 07:57 marostegui: Depool clouddb1015 (s4,s6) [[phab:T434048|T434048]] * 07:57 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s6 * 07:57 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s4 * 07:56 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95910 and previous config saved to /var/cache/conftool/dbconfig/20260805-075650-ladsgroup.json * 07:54 marostegui: Depool clouddb1014 (s2,s7) [[phab:T434048|T434048]] * 07:54 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1014.eqiad.wmnet,service=s7 * 07:54 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1014.eqiad.wmnet,service=s2 * 07:53 marostegui: Depool clouddb1013:s1 [[phab:T434048|T434048]] * 07:53 marostegui: Depool clouddb1013:s1 [[phab:T409557|T409557]] * 07:53 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1013.eqiad.wmnet,service=s1 * 07:25 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95909 and previous config saved to /var/cache/conftool/dbconfig/20260805-072529-ladsgroup.json * 07:24 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2249.codfw.wmnet with reason: Maintenance * 07:24 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95908 and previous config saved to /var/cache/conftool/dbconfig/20260805-072426-ladsgroup.json * 07:21 slyngshede@dns1004: END - running authdns-update * 07:19 slyngshede@dns1004: START - running authdns-update * 07:13 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231', diff saved to https://phabricator.wikimedia.org/P95906 and previous config saved to /var/cache/conftool/dbconfig/20260805-071340-ladsgroup.json * 07:02 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231', diff saved to https://phabricator.wikimedia.org/P95905 and previous config saved to /var/cache/conftool/dbconfig/20260805-070253-ladsgroup.json * 06:52 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95904 and previous config saved to /var/cache/conftool/dbconfig/20260805-065206-ladsgroup.json * 06:45 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 06:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95903 and previous config saved to /var/cache/conftool/dbconfig/20260805-062240-ladsgroup.json * 06:21 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2231.codfw.wmnet with reason: Maintenance * 06:21 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95902 and previous config saved to /var/cache/conftool/dbconfig/20260805-062137-ladsgroup.json * 06:10 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215', diff saved to https://phabricator.wikimedia.org/P95901 and previous config saved to /var/cache/conftool/dbconfig/20260805-061051-ladsgroup.json * 06:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215', diff saved to https://phabricator.wikimedia.org/P95900 and previous config saved to /var/cache/conftool/dbconfig/20260805-060004-ladsgroup.json * 05:49 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95899 and previous config saved to /var/cache/conftool/dbconfig/20260805-054918-ladsgroup.json * 05:19 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95898 and previous config saved to /var/cache/conftool/dbconfig/20260805-051939-ladsgroup.json * 05:18 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2215.codfw.wmnet with reason: Maintenance * 04:30 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2201.codfw.wmnet with reason: Maintenance * 03:40 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2197.codfw.wmnet with reason: Maintenance * 03:40 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95897 and previous config saved to /var/cache/conftool/dbconfig/20260805-034036-ladsgroup.json * 03:29 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196', diff saved to https://phabricator.wikimedia.org/P95896 and previous config saved to /var/cache/conftool/dbconfig/20260805-032948-ladsgroup.json * 03:19 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196', diff saved to https://phabricator.wikimedia.org/P95895 and previous config saved to /var/cache/conftool/dbconfig/20260805-031902-ladsgroup.json * 03:08 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95894 and previous config saved to /var/cache/conftool/dbconfig/20260805-030815-ladsgroup.json * 02:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95893 and previous config saved to /var/cache/conftool/dbconfig/20260805-023413-ladsgroup.json * 02:33 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2196.codfw.wmnet with reason: Maintenance * 02:33 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95892 and previous config saved to /var/cache/conftool/dbconfig/20260805-023310-ladsgroup.json * 02:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186', diff saved to https://phabricator.wikimedia.org/P95891 and previous config saved to /var/cache/conftool/dbconfig/20260805-022223-ladsgroup.json * 02:11 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186', diff saved to https://phabricator.wikimedia.org/P95890 and previous config saved to /var/cache/conftool/dbconfig/20260805-021137-ladsgroup.json * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 02:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95889 and previous config saved to /var/cache/conftool/dbconfig/20260805-020051-ladsgroup.json * 01:30 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95888 and previous config saved to /var/cache/conftool/dbconfig/20260805-013029-ladsgroup.json * 01:29 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2186.codfw.wmnet with reason: Maintenance * 00:34 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on dbstore1009.eqiad.wmnet with reason: Maintenance * 00:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95887 and previous config saved to /var/cache/conftool/dbconfig/20260805-003408-ladsgroup.json * 00:23 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264', diff saved to https://phabricator.wikimedia.org/P95886 and previous config saved to /var/cache/conftool/dbconfig/20260805-002322-ladsgroup.json * 00:12 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264', diff saved to https://phabricator.wikimedia.org/P95885 and previous config saved to /var/cache/conftool/dbconfig/20260805-001235-ladsgroup.json * 00:01 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95884 and previous config saved to /var/cache/conftool/dbconfig/20260805-000148-ladsgroup.json == 2026-08-04 == * 23:45 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95883 and previous config saved to /var/cache/conftool/dbconfig/20260804-234508-ladsgroup.json * 23:44 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1264.eqiad.wmnet with reason: Maintenance * 23:44 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95882 and previous config saved to /var/cache/conftool/dbconfig/20260804-234405-ladsgroup.json * 23:33 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237', diff saved to https://phabricator.wikimedia.org/P95881 and previous config saved to /var/cache/conftool/dbconfig/20260804-233317-ladsgroup.json * 23:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237', diff saved to https://phabricator.wikimedia.org/P95880 and previous config saved to /var/cache/conftool/dbconfig/20260804-232230-ladsgroup.json * 23:11 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95879 and previous config saved to /var/cache/conftool/dbconfig/20260804-231144-ladsgroup.json * 22:23 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95878 and previous config saved to /var/cache/conftool/dbconfig/20260804-222345-ladsgroup.json * 22:23 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1237.eqiad.wmnet with reason: Maintenance * 21:13 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1225.eqiad.wmnet with reason: Maintenance * 20:40 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] (duration: 24m 40s) * 20:33 samtar@deploy1003: samtar, kineticpelagic: Continuing with deployment * 20:28 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS bookworm * 20:21 samtar@deploy1003: samtar, kineticpelagic: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:15 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] * 20:13 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 20:09 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 20:00 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1216.eqiad.wmnet with reason: Maintenance * 20:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95877 and previous config saved to /var/cache/conftool/dbconfig/20260804-195957-ladsgroup.json * 19:51 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 19:50 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) mc2046.codfw.wmnet 120.16.192.10.in-addr.arpa 0.2.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:50 jhancock@cumin2002: START - Cookbook sre.dns.wipe-cache mc2046.codfw.wmnet 120.16.192.10.in-addr.arpa 0.2.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host mc2046 - jhancock@cumin2002" * 19:50 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host mc2046 - jhancock@cumin2002" * 19:49 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203', diff saved to https://phabricator.wikimedia.org/P95876 and previous config saved to /var/cache/conftool/dbconfig/20260804-194911-ladsgroup.json * 19:46 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 19:45 jhancock@cumin2002: START - Cookbook sre.hosts.move-vlan for host mc2046 * 19:45 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS bookworm * 19:38 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203', diff saved to https://phabricator.wikimedia.org/P95875 and previous config saved to /var/cache/conftool/dbconfig/20260804-193825-ladsgroup.json * 19:27 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95874 and previous config saved to /var/cache/conftool/dbconfig/20260804-192738-ladsgroup.json * 19:02 mutante: gerrit ssh -p 29418 gerrit.wikimedia.org gerrit index changes {{Gerrit|1320979}} * 18:20 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 18:18 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 18:14 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 18:14 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 18:13 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 18:10 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 18:08 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 18:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95872 and previous config saved to /var/cache/conftool/dbconfig/20260804-180721-ladsgroup.json * 18:07 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 18:06 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1203.eqiad.wmnet with reason: Maintenance * 18:06 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95871 and previous config saved to /var/cache/conftool/dbconfig/20260804-180618-ladsgroup.json * 17:55 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179', diff saved to https://phabricator.wikimedia.org/P95870 and previous config saved to /var/cache/conftool/dbconfig/20260804-175531-ladsgroup.json * 17:55 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1154.eqiad.wmnet * 17:55 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1154.eqiad.wmnet * 17:55 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1154.eqiad.wmnet * 17:50 swfrench@deploy1003: Finished scap sync-world: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] (duration: 04m 05s) * 17:48 swfrench@deploy1003: swfrench: Continuing with deployment * 17:46 swfrench@deploy1003: swfrench: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:45 swfrench@deploy1003: Started scap sync-world: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] * 17:44 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179', diff saved to https://phabricator.wikimedia.org/P95869 and previous config saved to /var/cache/conftool/dbconfig/20260804-174445-ladsgroup.json * 17:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95868 and previous config saved to /var/cache/conftool/dbconfig/20260804-173359-ladsgroup.json * 17:33 swfrench@deploy1003: Finished scap sync-world: Pick up new PHP production image (duration: 28m 32s) * 17:28 aokoth@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on phab1005.eqiad.wmnet with reason: Puppet Failure * 17:05 swfrench@deploy1003: Started scap sync-world: Pick up new PHP production image * 17:00 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 17:00 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 16:54 cgoubert@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on wikikube-worker2187.codfw.wmnet with reason: Hardware issue * 16:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2187.codfw.wmnet * 16:52 mutante: gerrit2003:/var/log/apache2# ln -s /srv/gerrit/site_path/review_site/logs/ gerrit * 16:52 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2187.codfw.wmnet * 16:48 mutante: gerrit2003 - moving old apache logfiles older than 60 days from /var/log/apache2 to /srv/gerrit/site_path/review_site/logs/old/ * 16:33 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 16:32 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 16:29 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 16:29 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 16:28 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 16:28 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 16:27 dzahn@cumin1003: END (PASS) - Cookbook sre.gerrit.restart-gerrit (exit_code=0) Restarting Gerrit on gerrit2003 * 16:27 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 16:27 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95867 and previous config saved to /var/cache/conftool/dbconfig/20260804-162736-ladsgroup.json * 16:27 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 16:26 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1179.eqiad.wmnet with reason: Maintenance * 16:26 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:25 mutante: restarting gerrit - dropped outdated RSA host key * 16:25 dzahn@cumin1003: START - Cookbook sre.gerrit.restart-gerrit Restarting Gerrit on gerrit2003 * 16:24 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95866 and previous config saved to /var/cache/conftool/dbconfig/20260804-162424-ladsgroup.json * 16:24 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 16:23 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 16:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95865 and previous config saved to /var/cache/conftool/dbconfig/20260804-162236-ladsgroup.json * 16:21 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1179.eqiad.wmnet with reason: Maintenance * 16:17 swfrench-wmf: reprepro include php8.3_8.3.33-1+wmf11u1 into component/php83 for bullseye-wikimedia * 16:17 swfrench-wmf: reprepro include php8.3_8.3.33-1+wmf12u1 into component/php83 for bookworm-wikimedia * 16:11 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply * 16:10 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply * 16:10 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mobileapps: apply * 16:09 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mobileapps: apply * 16:09 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply * 16:08 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply * 16:08 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:08 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:07 aokoth@cumin1003: END (PASS) - Cookbook sre.vrts.upgrade (exit_code=0) on VRTS host vrts1003.eqiad.wmnet * 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:05 aokoth@cumin1003: START - Cookbook sre.vrts.upgrade on VRTS host vrts1003.eqiad.wmnet * 16:04 mutante: gerrit2002/gerrit1003/gerrit2003 - rm /etc/gerrit/ssh_host_rsa_key * 15:59 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:59 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:59 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:59 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:56 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 15:55 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:55 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:55 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:49 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 15:49 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:44 Raine: add php8.5 packages to component/php85 - [[phab:T432983|T432983]] * 15:39 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:33 aaron@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 15:33 aaron@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 15:29 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:19 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:19 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:16 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:16 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1154.eqiad.wmnet with OS trixie * 15:16 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:15 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:15 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:06 brennen@deploy1003: Finished deploy [phabricator/deployment@56f4ffd]: deploy phab1004 for [[phab:T433981|T433981]] (duration: 00m 43s) * 15:05 brennen@deploy1003: Started deploy [phabricator/deployment@56f4ffd]: deploy phab1004 for [[phab:T433981|T433981]] * 15:05 aaron@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 15:04 aaron@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 15:02 brennen@deploy1003: Finished deploy [phabricator/deployment@56f4ffd]: deploy phab2003 for [[phab:T433981|T433981]] (duration: 00m 51s) * 15:01 brennen@deploy1003: Started deploy [phabricator/deployment@56f4ffd]: deploy phab2003 for [[phab:T433981|T433981]] * 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1004.eqiad.wmnet with reason: deployment * 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1005.eqiad.wmnet with reason: deployment * 14:58 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab2003.codfw.wmnet with reason: deployment * 14:55 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1154.eqiad.wmnet with reason: host reimage * 14:51 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1154.eqiad.wmnet with reason: host reimage * 14:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 14:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 14:38 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync * 14:38 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync * 14:38 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync * 14:37 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync * 14:37 ottomata: roll restart eventgate-main to pick up stream config change - [[phab:T433507|T433507]] * 14:37 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-main: sync * 14:36 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-main: sync * 14:36 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1154 * 14:36 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1154 * 14:34 otto@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] (duration: 08m 39s) * 14:34 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1154 * 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1154.eqiad.wmnet 108.32.64.10.in-addr.arpa 8.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:34 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1154.eqiad.wmnet 108.32.64.10.in-addr.arpa 8.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1154 - jayme@cumin1003" * 14:34 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1154 - jayme@cumin1003" * 14:30 otto@deploy1003: otto: Continuing with deployment * 14:30 jayme@cumin1003: START - Cookbook sre.dns.netbox * 14:28 otto@deploy1003: otto: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:26 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1154 * 14:26 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1154.eqiad.wmnet with OS trixie * 14:26 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1154.eqiad.wmnet * 14:26 otto@deploy1003: Started scap sync-world: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] * 14:26 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1154.eqiad.wmnet * 14:26 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1154.eqiad.wmnet * 14:17 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 14:16 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 14:15 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 14:14 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 14:13 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 14:13 swfrench@dns1004: END - running authdns-update * 14:13 Msz2001: Finished deployments for UTC afternoon backport window * 14:13 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 14:13 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] (duration: 07m 58s) * 14:11 swfrench@dns1004: START - running authdns-update * 14:08 mszwarc@deploy1003: javiermonton, mszwarc, mpostoronca: Continuing with deployment * 14:07 mszwarc@deploy1003: javiermonton, mszwarc, mpostoronca: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] synced to the testser * 14:05 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] * 14:03 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 13:49 swfrench@cumin2002: conftool action : set/pooled=yes; selector: name=wikikube-worker2330.codfw.wmnet * 13:49 swfrench@cumin2002: conftool action : set/pooled=no; selector: name=wikikube-worker2330.codfw.wmnet * 13:48 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] (duration: 09m 19s) * 13:45 swfrench@dns1004: END - running authdns-update * 13:44 mszwarc@deploy1003: mszwarc, jforrester: Continuing with deployment * 13:43 swfrench@dns1004: START - running authdns-update * 13:41 mszwarc@deploy1003: mszwarc, jforrester: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:38 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] * 13:33 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 13:33 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1154.eqiad.wmnet * 13:32 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 13:32 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 13:31 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 13:31 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:31 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:29 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1154.eqiad.wmnet * 13:28 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1154.eqiad.wmnet * 13:28 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1154.eqiad.wmnet * 13:28 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1141.eqiad.wmnet * 13:28 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1141.eqiad.wmnet * 13:28 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1141.eqiad.wmnet * 13:22 otto@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply * 13:22 otto@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply * 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1096.eqiad.wmnet with OS trixie * 13:05 swfrench@dns1004: END - running authdns-update * 13:03 swfrench@dns1004: START - running authdns-update * 12:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 12:43 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 1:00:00 on db1171.eqiad.wmnet with reason: decom * 12:42 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 1:00:00 on db1150.eqiad.wmnet with reason: decom * 12:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 12:38 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1164,1217].eqiad.wmnet with reason: cloning * 12:33 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2096.codfw.wmnet with OS trixie * 12:22 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1096.eqiad.wmnet with OS trixie * 12:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2096.codfw.wmnet with reason: host reimage * 12:14 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1141.eqiad.wmnet with OS trixie * 12:10 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2096.codfw.wmnet with reason: host reimage * 12:10 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1289.eqiad.wmnet * 12:05 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1289.eqiad.wmnet * 12:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1288.eqiad.wmnet * 11:59 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1288.eqiad.wmnet * 11:59 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1287.eqiad.wmnet * 11:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1097.eqiad.wmnet with OS trixie * 11:54 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1287.eqiad.wmnet * 11:54 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1286.eqiad.wmnet * 11:53 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1141.eqiad.wmnet with reason: host reimage * 11:51 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2096.codfw.wmnet with OS trixie * 11:49 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1141.eqiad.wmnet with reason: host reimage * 11:48 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1286.eqiad.wmnet * 11:48 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1284.eqiad.wmnet * 11:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2095.codfw.wmnet with OS trixie * 11:43 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1284.eqiad.wmnet * 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1283.eqiad.wmnet * 11:42 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on ml-serve1015.eqiad.wmnet with reason: Downtime to get full picture of current BIOS settings beyond what Redfish shows * 11:39 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad * 11:39 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:37 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1283.eqiad.wmnet * 11:37 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1282.eqiad.wmnet * 11:37 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad * 11:37 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:33 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1141 * 11:33 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1141 * 11:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 11:32 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1141 * 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1141.eqiad.wmnet 156.48.64.10.in-addr.arpa 6.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:32 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1141.eqiad.wmnet 156.48.64.10.in-addr.arpa 6.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1141 - jayme@cumin1003" * 11:32 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1141 - jayme@cumin1003" * 11:32 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1282.eqiad.wmnet * 11:32 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1281.eqiad.wmnet * 11:32 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad * 11:32 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:29 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 11:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1097.eqiad.wmnet with reason: host reimage * 11:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2095.codfw.wmnet with OS trixie * 11:26 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1281.eqiad.wmnet * 11:26 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1280.eqiad.wmnet * 11:25 jayme@cumin1003: START - Cookbook sre.dns.netbox * 11:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1097.eqiad.wmnet with reason: host reimage * 11:22 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1141 * 11:21 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1141.eqiad.wmnet with OS trixie * 11:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1280.eqiad.wmnet * 11:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1279.eqiad.wmnet * 11:20 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin with reason: upgrade new Nokia swtiches in eqsin to SR Linux v26 * 11:17 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1141.eqiad.wmnet * 11:16 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1141.eqiad.wmnet * 11:16 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1141.eqiad.wmnet * 11:16 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be2095.codfw.wmnet with OS trixie * 11:15 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1279.eqiad.wmnet * 11:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1278.eqiad.wmnet * 11:14 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1139.eqiad.wmnet * 11:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1139.eqiad.wmnet * 11:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1070.eqiad.wmnet with OS trixie * 11:13 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1140.eqiad.wmnet * 11:13 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1140.eqiad.wmnet * 11:13 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1140.eqiad.wmnet * 11:09 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1278.eqiad.wmnet * 11:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1071.eqiad.wmnet with OS trixie * 11:05 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1097.eqiad.wmnet with OS trixie * 11:04 mvernon@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be1097.eqiad.wmnet with OS trixie * 11:02 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1140.eqiad.wmnet with OS trixie * 11:02 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1097.eqiad.wmnet with OS trixie * 11:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1096.eqiad.wmnet with OS trixie * 11:00 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1139.eqiad.wmnet * 11:00 jayme@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1139.eqiad.wmnet with OS trixie * 10:56 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1069.eqiad.wmnet with OS trixie * 10:56 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 10:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1070.eqiad.wmnet with reason: host reimage * 10:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 10:45 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1071.eqiad.wmnet with reason: host reimage * 10:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 10:41 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1096.eqiad.wmnet with OS trixie * 10:41 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1140.eqiad.wmnet with reason: host reimage * 10:39 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1071.eqiad.wmnet with reason: host reimage * 10:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1070.eqiad.wmnet with reason: host reimage * 10:38 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1139.eqiad.wmnet with reason: host reimage * 10:37 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1140.eqiad.wmnet with reason: host reimage * 10:35 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1069.eqiad.wmnet with reason: host reimage * 10:33 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1095.eqiad.wmnet with OS trixie * 10:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 10:33 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1139.eqiad.wmnet with reason: host reimage * 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1069.eqiad.wmnet with reason: host reimage * 10:23 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1140 * 10:23 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1140 * 10:23 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 10:22 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1140 * 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1140.eqiad.wmnet 155.48.64.10.in-addr.arpa 5.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:21 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1071.eqiad.wmnet with OS trixie * 10:21 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1140.eqiad.wmnet 155.48.64.10.in-addr.arpa 5.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1140 - jayme@cumin1003" * 10:21 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1140 - jayme@cumin1003" * 10:21 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1071 * 10:21 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1070.eqiad.wmnet with OS trixie * 10:21 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1070 * 10:20 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 10:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1095.eqiad.wmnet with OS trixie * 10:17 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1139 * 10:17 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1139 * 10:17 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be1095.eqiad.wmnet with OS trixie * 10:15 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1139 * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1139.eqiad.wmnet 194.32.64.10.in-addr.arpa 4.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:15 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1139.eqiad.wmnet 194.32.64.10.in-addr.arpa 4.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1139 - jayme@cumin1003" * 10:15 jayme@cumin1003: START - Cookbook sre.dns.netbox * 10:15 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1139 - jayme@cumin1003" * 10:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2095.codfw.wmnet with OS trixie * 10:13 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1069.eqiad.wmnet with OS trixie * 10:12 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1069 * 10:11 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1140 * 10:11 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1140.eqiad.wmnet with OS trixie * 10:11 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1140.eqiad.wmnet * 10:10 jayme@cumin1003: START - Cookbook sre.dns.netbox * 10:10 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1139 * 10:10 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1140.eqiad.wmnet * 10:10 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1140.eqiad.wmnet * 10:10 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1139.eqiad.wmnet with OS trixie * 10:09 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1139.eqiad.wmnet * 10:08 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1139.eqiad.wmnet * 10:08 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1139.eqiad.wmnet * 10:01 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2094.codfw.wmnet with OS trixie * 09:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 09:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 09:53 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 09:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 09:44 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:44 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2094.codfw.wmnet with reason: host reimage * 09:34 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2094.codfw.wmnet with reason: host reimage * 09:34 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1071 * 09:33 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1070 * 09:33 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1095.eqiad.wmnet with OS trixie * 09:27 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1069 * 09:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:23 brouberol@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM archiva1002.wikimedia.org * 09:20 brouberol@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM archiva1002.wikimedia.org * 09:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1277.eqiad.wmnet * 09:13 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2094.codfw.wmnet with OS trixie * 09:13 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 09:12 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 09:12 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:12 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1277.eqiad.wmnet * 09:12 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1276.eqiad.wmnet * 09:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1094.eqiad.wmnet with OS trixie * 09:06 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1276.eqiad.wmnet * 09:06 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1275.eqiad.wmnet * 09:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2093.codfw.wmnet with OS trixie * 09:01 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1275.eqiad.wmnet * 09:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1274.eqiad.wmnet * 08:56 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1274.eqiad.wmnet * 08:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1273.eqiad.wmnet * 08:50 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1273.eqiad.wmnet * 08:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1272.eqiad.wmnet * 08:49 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:49 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1094.eqiad.wmnet with reason: host reimage * 08:45 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1272.eqiad.wmnet * 08:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1094.eqiad.wmnet with reason: host reimage * 08:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2093.codfw.wmnet with reason: host reimage * 08:38 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:38 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2093.codfw.wmnet with reason: host reimage * 08:35 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 08:34 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 08:29 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:28 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:26 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1271.eqiad.wmnet * 08:23 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1094.eqiad.wmnet with OS trixie * 08:21 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 08:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1271.eqiad.wmnet * 08:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1270.eqiad.wmnet * 08:15 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1270.eqiad.wmnet * 08:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1269.eqiad.wmnet * 08:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2093.codfw.wmnet with OS trixie * 08:09 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1269.eqiad.wmnet * 08:09 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1268.eqiad.wmnet * 08:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2092.codfw.wmnet with OS trixie * 08:04 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1268.eqiad.wmnet * 08:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1267.eqiad.wmnet * 07:59 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1267.eqiad.wmnet * 07:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1093.eqiad.wmnet with OS trixie * 07:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1266.eqiad.wmnet * 07:56 jynus: running extra backups to test db1285 [[phab:T433826|T433826]] * 07:51 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1266.eqiad.wmnet * 07:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2092.codfw.wmnet with reason: host reimage * 07:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1093.eqiad.wmnet with reason: host reimage * 07:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2092.codfw.wmnet with reason: host reimage * 07:32 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1093.eqiad.wmnet with reason: host reimage * 07:29 jynus: running extra backups to test db1265 [[phab:T433825|T433825]] * 07:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2092.codfw.wmnet with OS trixie * 07:11 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1093.eqiad.wmnet with OS trixie * 06:50 slyngshede@dns1004: END - running authdns-update * 06:48 slyngshede@dns1004: START - running authdns-update * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.11 (duration: 02m 29s) * 03:38 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] (duration: 32m 57s) * 03:23 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 03:22 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 03:05 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 32s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 00:45 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] (duration: 06m 20s) * 00:41 cjming@deploy1003: cjming: Continuing with deployment * 00:41 cjming@deploy1003: cjming: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:39 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] == 2026-08-03 == * 23:58 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply * 23:57 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply * 23:29 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cp5021.eqsin.wmnet * 23:29 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cp5021.eqsin.wmnet * 23:27 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cp5021.eqsin.wmnet * 23:26 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cp5021.eqsin.wmnet * 23:18 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 23:17 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 22:56 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: sync * 22:56 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: sync * 22:36 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 22:36 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 22:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host search-loader2002.codfw.wmnet with OS trixie * 21:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on search-loader2002.codfw.wmnet with reason: host reimage * 21:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on search-loader2002.codfw.wmnet with reason: host reimage * 21:42 dancy@deploy1003: Stopping before sync operations * 21:41 dancy@deploy1003: Started scap sync-world: testing * 21:39 dancy@deploy1003: Installation of scap version "4.277.0" completed for 3 hosts * 21:37 dancy@deploy1003: Installing scap version "4.277.0" for 3 host(s) * 21:37 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] (duration: 06m 13s) * 21:33 dancy@deploy1003: dancy: Continuing with deployment * 21:32 dancy@deploy1003: dancy: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:31 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] * 21:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host search-loader2002.codfw.wmnet with OS trixie * 21:03 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] (duration: 06m 34s) * 20:59 dancy@deploy1003: dancy: Continuing with deployment * 20:58 dancy@deploy1003: dancy: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:56 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] * 20:52 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] (duration: 06m 23s) * 20:48 cjming@deploy1003: cjming: Continuing with deployment * 20:47 cjming@deploy1003: cjming: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:46 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] * 20:42 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] (duration: 07m 36s) * 20:38 arlolra@deploy1003: arlolra: Continuing with deployment * 20:36 arlolra@deploy1003: arlolra: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:34 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] * 20:16 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] (duration: 08m 26s) * 20:12 krinkle@deploy1003: krinkle: Continuing with deployment * 20:09 krinkle@deploy1003: krinkle: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] * 19:45 jasmine@cumin2002: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-main-codfw * 18:58 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] (duration: 09m 23s) * 18:53 krinkle@deploy1003: krinkle: Continuing with deployment * 18:53 jasmine@cumin2002: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-main-codfw * 18:50 krinkle@deploy1003: krinkle: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:48 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] * 18:37 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] (duration: 10m 13s) * 18:34 dzahn@cumin2002: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host codesearch2001.codfw.wmnet * 18:34 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host codesearch2001.codfw.wmnet with OS trixie * 18:33 krinkle@deploy1003: krinkle: Continuing with deployment * 18:29 krinkle@deploy1003: krinkle: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:27 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] * 18:19 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 18:18 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on codesearch2001.codfw.wmnet with reason: host reimage * 18:14 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 18:14 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:12 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on codesearch2001.codfw.wmnet with reason: host reimage * 18:11 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 18:11 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 18:10 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 18:02 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 18:02 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 18:01 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 18:01 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 17:55 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host codesearch2001.codfw.wmnet with OS trixie * 17:54 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:54 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) codesearch2001.codfw.wmnet on all recursors * 17:53 dzahn@cumin2002: START - Cookbook sre.dns.wipe-cache codesearch2001.codfw.wmnet on all recursors * 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:48 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:41 dzahn@cumin2002: START - Cookbook sre.dns.netbox * 17:41 dzahn@cumin2002: START - Cookbook sre.ganeti.makevm for new host codesearch2001.codfw.wmnet * 17:37 dzahn@cumin2002: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host codesearch1001.eqiad.wmnet * 17:37 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host codesearch1001.eqiad.wmnet with OS trixie * 17:24 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on codesearch1001.eqiad.wmnet with reason: host reimage * 17:17 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on codesearch1001.eqiad.wmnet with reason: host reimage * 17:08 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host codesearch1001.eqiad.wmnet with OS trixie * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 17:06 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:06 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) codesearch1001.eqiad.wmnet on all recursors * 17:06 dzahn@cumin2002: START - Cookbook sre.dns.wipe-cache codesearch1001.eqiad.wmnet on all recursors * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 17:05 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 17:04 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:04 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 16:58 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 16:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2091.codfw.wmnet with OS trixie * 16:54 ebernhardson@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 16:54 ebernhardson@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 16:49 ebernhardson@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 16:49 ebernhardson@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 16:46 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1092.eqiad.wmnet with OS trixie * 16:43 ebernhardson@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 16:43 ebernhardson@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 16:43 dzahn@cumin2002: START - Cookbook sre.dns.netbox * 16:43 dzahn@cumin2002: START - Cookbook sre.ganeti.makevm for new host codesearch1001.eqiad.wmnet * 16:41 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 16:41 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 16:40 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2091.codfw.wmnet with reason: host reimage * 16:37 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 16:35 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2091.codfw.wmnet with reason: host reimage * 16:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1092.eqiad.wmnet with reason: host reimage * 16:24 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 16:24 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 16:23 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1092.eqiad.wmnet with reason: host reimage * 16:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2091.codfw.wmnet with OS trixie * 16:03 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1092.eqiad.wmnet with OS trixie * 16:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2090.codfw.wmnet with OS trixie * 15:51 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 15:51 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 15:51 jiji@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 15:50 jiji@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 15:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2090.codfw.wmnet with reason: host reimage * 15:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2090.codfw.wmnet with reason: host reimage * 15:33 jhathaway@dns1004: END - running authdns-update * 15:31 jhathaway@dns1004: START - running authdns-update * 15:26 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1091.eqiad.wmnet with OS trixie * 15:25 dancy@deploy1003: Installation of scap version "4.276.1" completed for 3 hosts * 15:23 dancy@deploy1003: Installing scap version "4.276.1" for 3 host(s) * 15:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2090.codfw.wmnet with OS trixie * 15:12 marostegui@cumin1003: dbctl commit (dc=all): 'Repool db2245, db2246, db2247 and db2248 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95857 and previous config saved to /var/cache/conftool/dbconfig/20260803-151212-marostegui.json * 15:09 dancy@deploy1003: Started scap sync-world: testing * 15:09 dancy@deploy1003: Installation of scap version "4.277.0" completed for 3 hosts * 15:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1091.eqiad.wmnet with reason: host reimage * 15:07 dancy@deploy1003: Installing scap version "4.277.0" for 3 host(s) * 15:03 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1091.eqiad.wmnet with reason: host reimage * 14:49 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1091.eqiad.wmnet with OS trixie * 14:33 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2089.codfw.wmnet with OS trixie * 14:29 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 14:27 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 14:18 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 14:16 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 14:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2089.codfw.wmnet with reason: host reimage * 14:10 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 14:10 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 14:09 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2089.codfw.wmnet with reason: host reimage * 13:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2089.codfw.wmnet with OS trixie * 13:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2088.codfw.wmnet with OS trixie * 13:40 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1090.eqiad.wmnet with OS trixie * 13:22 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1090.eqiad.wmnet with reason: host reimage * 13:22 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] (duration: 14m 34s) * 13:19 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1090.eqiad.wmnet with reason: host reimage * 13:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2088.codfw.wmnet with reason: host reimage * 13:16 aude@deploy1003: aude, mhorsey: Continuing with deployment * 13:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2088.codfw.wmnet with reason: host reimage * 13:12 aude@deploy1003: aude, mhorsey: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] * 13:05 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1090.eqiad.wmnet with OS trixie * 12:58 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2088.codfw.wmnet with OS trixie * 12:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db[2245-2247].codfw.wmnet * 12:49 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2247: Rebooting db2247.codfw.wmnet * 12:49 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2247: Rebooting db2247.codfw.wmnet * 12:42 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2246: Rebooting db2246.codfw.wmnet * 12:42 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2246: Rebooting db2246.codfw.wmnet * 12:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2087.codfw.wmnet with OS trixie * 12:37 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1089.eqiad.wmnet with OS trixie * 12:34 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2245: Rebooting db2245.codfw.wmnet * 12:34 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2245: Rebooting db2245.codfw.wmnet * 12:34 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db[2245-2247].codfw.wmnet * 12:32 kamila@deploy1003: Finished scap sync-world: rebuild after base image update (duration: 30m 26s) * 12:28 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 12:22 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2087.codfw.wmnet with reason: host reimage * 12:19 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1089.eqiad.wmnet with reason: host reimage * 12:14 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2087.codfw.wmnet with reason: host reimage * 12:14 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1089.eqiad.wmnet with reason: host reimage * 12:03 kamila@deploy1003: Started scap sync-world: rebuild after base image update * 12:00 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1089.eqiad.wmnet with OS trixie * 12:00 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2087.codfw.wmnet with OS trixie * 11:35 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db[2245-2248].codfw.wmnet * 11:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db[2245-2248].codfw.wmnet * 11:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2086.codfw.wmnet with OS trixie * 11:26 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 11:26 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 11:25 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 11:25 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 11:24 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1088.eqiad.wmnet with OS trixie * 11:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db[2245-2248].codfw.wmnet with reason: Checking network * 11:21 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 11:20 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 11:19 marostegui@dns1004: END - running authdns-update * 11:17 marostegui@dns1004: START - running authdns-update * 11:10 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:10 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 11:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2086.codfw.wmnet with reason: host reimage * 11:09 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:08 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 11:08 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:07 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 11:07 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:07 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 11:06 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop: apply * 11:06 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop: apply * 11:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1088.eqiad.wmnet with reason: host reimage * 11:05 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop: apply * 11:04 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop: apply * 11:04 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop: apply * 11:04 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop: apply * 11:02 marostegui@dns1004: END - running authdns-update * 11:02 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2086.codfw.wmnet with reason: host reimage * 11:01 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1088.eqiad.wmnet with reason: host reimage * 11:00 marostegui@dns1004: START - running authdns-update * 10:53 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] (duration: 10m 57s) * 10:51 cmooney@dns3003: END - running authdns-update * 10:49 cmooney@dns3003: START - running authdns-update * 10:47 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1088.eqiad.wmnet with OS trixie * 10:47 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2086.codfw.wmnet with OS trixie * 10:47 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 10:46 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:46 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:46 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new reverse ranges for eqsin CR switch links - cmooney@cumin1003" * 10:46 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new reverse ranges for eqsin CR switch links - cmooney@cumin1003" * 10:42 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] * 10:41 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 10:36 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2245, db2246 and db2247 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95855 and previous config saved to /var/cache/conftool/dbconfig/20260803-103652-marostegui.json * 10:35 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2248 from s4 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95854 and previous config saved to /var/cache/conftool/dbconfig/20260803-103535-marostegui.json * 10:27 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 10:27 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 10:26 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 10:24 kart_: cxserver: Add referencePunctuation config ([[phab:T97231|T97231]]) * 10:24 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 10:23 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:23 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:23 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:22 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:22 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply * 10:21 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply * 10:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2085.codfw.wmnet with OS trixie * 10:20 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply * 10:20 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply * 10:18 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply * 10:18 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply * 10:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1087.eqiad.wmnet with OS trixie * 09:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2085.codfw.wmnet with reason: host reimage * 09:43 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1087.eqiad.wmnet with reason: host reimage * 09:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2085.codfw.wmnet with reason: host reimage * 09:40 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1087.eqiad.wmnet with reason: host reimage * 09:26 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1087.eqiad.wmnet with OS trixie * 09:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2085.codfw.wmnet with OS trixie * 09:13 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2084.codfw.wmnet with OS trixie * 09:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1086.eqiad.wmnet with OS trixie * 08:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2084.codfw.wmnet with reason: host reimage * 08:50 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2084.codfw.wmnet with reason: host reimage * 08:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1086.eqiad.wmnet with reason: host reimage * 08:39 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1086.eqiad.wmnet with reason: host reimage * 08:38 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:38 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:37 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 08:37 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 08:35 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2084.codfw.wmnet with OS trixie * 08:34 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 08:34 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:27 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1086.eqiad.wmnet with OS trixie * 08:09 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1218: Repool after a crash * 08:07 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2083.codfw.wmnet with OS trixie * 08:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1085.eqiad.wmnet with OS trixie * 07:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2083.codfw.wmnet with reason: host reimage * 07:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1085.eqiad.wmnet with reason: host reimage * 07:40 kart_: Updated cxsever to 2026-07-16-140518-production ([[phab:T97231|T97231]]) * 07:39 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply * 07:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2083.codfw.wmnet with reason: host reimage * 07:38 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply * 07:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1085.eqiad.wmnet with reason: host reimage * 07:37 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] (duration: 32m 40s) * 07:33 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply * 07:33 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply * 07:25 jdlrobson@deploy1003: jdlrobson: Continuing with deployment * 07:24 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2083.codfw.wmnet with OS trixie * 07:24 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1085.eqiad.wmnet with OS trixie * 07:23 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1218: Repool after a crash * 07:21 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:09 marostegui: Drop renamed tables [[phab:T425074|T425074]] * 07:04 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] * 06:55 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply * 06:54 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 46s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-02 == * 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 01m 03s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-01 == * 03:30 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:30 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:30 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:30 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 34s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-31 == * 17:41 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 17:41 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 17:40 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 17:40 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 15:33 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2195: Testing * 15:02 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:02 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 15:02 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 14:48 pt1979@cumin2002: START - Cookbook sre.dns.netbox * 14:47 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2195: Testing * 14:22 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 14:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2195: Testing * 14:21 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 14:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2195.codfw.wmnet with reason: Testing * 14:16 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 14:04 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 14:04 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 13:30 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1048.eqiad.wmnet with OS trixie * 13:22 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2195: Testing * 13:22 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 13:19 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2195: Testing * 13:18 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 13:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2195: Testing * 13:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 13:05 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 13:05 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 13:04 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 13:04 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 12:50 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lswtest-d8-eqiad * 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:53 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:42 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 6515 * 11:37 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 6515 * 11:28 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:27 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:07 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2082.codfw.wmnet with OS trixie * 10:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2082.codfw.wmnet with reason: host reimage * 10:42 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2082.codfw.wmnet with reason: host reimage * 10:28 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2082.codfw.wmnet with OS trixie * 10:02 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 09:52 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 09:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts * 09:16 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts * 08:57 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 08:46 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:42 gkyziridis@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 08:37 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 08:37 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 08:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 08:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 08:11 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:11 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:08 filippo@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudvirt1048 * 08:07 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 08:07 filippo@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudvirt1048 * 08:06 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 08:01 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 08:00 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 07:19 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1048.eqiad.wmnet with reason: host reimage * 07:13 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1048.eqiad.wmnet with reason: host reimage * 07:11 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 07:11 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 07:09 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 07:09 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 06:57 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:56 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:48 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:48 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:44 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie * 06:34 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1048.eqiad.wmnet with OS trixie * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 54s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 00:57 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] (duration: 11m 04s) * 00:53 dreamyjazz@deploy1003: dreamyjazz, jforrester: Continuing with deployment * 00:48 dreamyjazz@deploy1003: dreamyjazz, jforrester: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:46 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] == 2026-07-30 == * 21:37 dancy@deploy1003: Installation of scap version "4.276.1" completed for 3 hosts * 21:35 dancy@deploy1003: Installing scap version "4.276.1" for 3 host(s) * 21:24 dancy@deploy1003: Installation of scap version "4.276.0" completed for 3 hosts * 21:22 dancy@deploy1003: Installing scap version "4.276.0" for 3 host(s) * 21:15 maryum: Deployed security fix for [[phab:T430601|T430601]] * 20:13 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] (duration: 09m 20s) * 20:07 arlolra@deploy1003: osleger, arlolra: Continuing with deployment * 20:05 arlolra@deploy1003: osleger, arlolra: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:03 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] * 19:29 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:29 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:25 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service * 19:24 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 19:24 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:24 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:24 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 19:23 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service * 19:20 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1084.eqiad.wmnet with OS trixie * 18:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1084.eqiad.wmnet with reason: host reimage * 18:52 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1084.eqiad.wmnet with reason: host reimage * 18:41 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 18:40 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 18:39 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1084.eqiad.wmnet with OS trixie * 18:25 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 18:15 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 17:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1083.eqiad.wmnet with OS trixie * 17:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1048: Maintenance * 17:36 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new security plugin settings - bking@cumin2003 - [[phab:T350516|T350516]] * 17:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1083.eqiad.wmnet with reason: host reimage * 17:26 inflatador: bking@apt1002 `reprepro --noskipold --component thirdparty/opensearch3 update trixie-wikimedia` [[phab:T433624|T433624]] * 17:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1083.eqiad.wmnet with reason: host reimage * 17:23 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 17:20 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 17:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2081.codfw.wmnet with OS trixie * 17:11 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new security plugin settings - bking@cumin2003 - [[phab:T350516|T350516]] * 17:10 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1083.eqiad.wmnet with OS trixie * 16:55 root@cumin1003: START - Cookbook sre.mysql.pool pool es1048: Maintenance * 16:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2081.codfw.wmnet with reason: host reimage * 16:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1048 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95833 and previous config saved to /var/cache/conftool/dbconfig/20260730-165053-cwilliams.json * 16:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1048.eqiad.wmnet with reason: Maintenance * 16:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1040: Maintenance * 16:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2081.codfw.wmnet with reason: host reimage * 16:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1082.eqiad.wmnet with OS trixie * 16:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2081.codfw.wmnet with OS trixie * 16:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2097.codfw.wmnet with OS trixie * 16:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1082.eqiad.wmnet with reason: host reimage * 16:08 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1082.eqiad.wmnet with reason: host reimage * 16:08 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 16:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1047: Maintenance * 16:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2080.codfw.wmnet with OS trixie * 16:04 root@cumin1003: START - Cookbook sre.mysql.pool pool es1040: Maintenance * 16:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1040: Maintenance * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new logging settings - bking@cumin2003 - [[phab:T324335|T324335]] * 15:58 root@cumin1003: START - Cookbook sre.mysql.pool pool es1040: Maintenance * 15:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1040 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95827 and previous config saved to /var/cache/conftool/dbconfig/20260730-155324-cwilliams.json * 15:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1040.eqiad.wmnet with reason: Maintenance * 15:50 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1082.eqiad.wmnet with OS trixie * 15:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2048: Maintenance * 15:44 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 24s) * 15:43 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2080.codfw.wmnet with reason: host reimage * 15:38 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new logging settings - bking@cumin2003 - [[phab:T324335|T324335]] * 15:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2080.codfw.wmnet with reason: host reimage * 15:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 15:30 mvernon@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be2097.codfw.wmnet with OS trixie * 15:23 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1081.eqiad.wmnet with OS trixie * 15:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2098.codfw.wmnet with OS trixie * 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - mvernon@cumin2003" * 15:18 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be2097.codfw.wmnet with OS trixie * 15:18 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - mvernon@cumin2003" * 15:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 15:17 root@cumin1003: START - Cookbook sre.mysql.pool pool es1047: Maintenance * 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2080.codfw.wmnet with OS trixie * 15:13 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be2097.codfw.wmnet with OS trixie * 15:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1047 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95820 and previous config saved to /var/cache/conftool/dbconfig/20260730-151200-cwilliams.json * 15:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1047.eqiad.wmnet with reason: Maintenance * 15:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: Maintenance * 15:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 15:04 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 15:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1081.eqiad.wmnet with reason: host reimage * 15:00 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1081.eqiad.wmnet with reason: host reimage * 15:00 root@cumin1003: START - Cookbook sre.mysql.pool pool es2048: Maintenance * 15:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 14:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2079.codfw.wmnet with OS trixie * 14:56 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 14:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2048 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95816 and previous config saved to /var/cache/conftool/dbconfig/20260730-145510-cwilliams.json * 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2048.codfw.wmnet with reason: Maintenance * 14:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2040: Maintenance * 14:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 14:51 tchin@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] (duration: 06m 48s) * 14:47 tchin@deploy1003: jforrester, tchin: Continuing with deployment * 14:47 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 14:47 tchin@deploy1003: jforrester, tchin: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:45 tchin@deploy1003: Started scap sync-world: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] * 14:42 sukhe@puppetserver1001: conftool action : set/weight=1; selector: cluster=urldownloader,service=squid * 14:42 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader,service=squid * 14:42 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1081.eqiad.wmnet with OS trixie * 14:39 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 14:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2079.codfw.wmnet with reason: host reimage * 14:36 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2098.codfw.wmnet with OS trixie * 14:32 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2079.codfw.wmnet with reason: host reimage * 14:30 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] (duration: 06m 31s) * 14:27 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 14:26 mszwarc@deploy1003: mszwarc: Continuing with deployment * 14:25 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:25 root@cumin1003: START - Cookbook sre.mysql.pool pool es1038: Maintenance * 14:25 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1038: Maintenance * 14:23 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] * 14:21 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] (duration: 11m 19s) * 14:20 root@cumin1003: START - Cookbook sre.mysql.pool pool es1038: Maintenance * 14:14 stran@deploy1003: stran: Continuing with deployment * 14:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1038 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95810 and previous config saved to /var/cache/conftool/dbconfig/20260730-141439-cwilliams.json * 14:14 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1038.eqiad.wmnet with reason: Maintenance * 14:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1036: Maintenance * 14:13 stran@deploy1003: stran: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2079.codfw.wmnet with OS trixie * 14:09 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] * 14:08 root@cumin1003: START - Cookbook sre.mysql.pool pool es2040: Maintenance * 14:08 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2040: Maintenance * 14:03 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] (duration: 31m 41s) * 14:03 root@cumin1003: START - Cookbook sre.mysql.pool pool es2040: Maintenance * 14:03 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2040 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95806 and previous config saved to /var/cache/conftool/dbconfig/20260730-135643-cwilliams.json * 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2040.codfw.wmnet with reason: Maintenance * 13:56 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:56 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2038: Maintenance * 13:55 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:52 lucaswerkmeister-wmde@deploy1003: migr, lucaswerkmeister-wmde: Continuing with deployment * 13:49 lucaswerkmeister-wmde@deploy1003: migr, lucaswerkmeister-wmde: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:49 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:48 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 13:45 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 13:32 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:32 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] * 13:28 root@cumin1003: START - Cookbook sre.mysql.pool pool es1036: Maintenance * 13:28 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1036: Maintenance * 13:22 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:22 root@cumin1003: START - Cookbook sre.mysql.pool pool es1036: Maintenance * 13:20 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1036 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95800 and previous config saved to /var/cache/conftool/dbconfig/20260730-131727-cwilliams.json * 13:17 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1036.eqiad.wmnet with reason: Maintenance * 13:17 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] (duration: 10m 31s) * 13:16 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2022\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 13:13 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, stran: Continuing with deployment * 13:10 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:10 root@cumin1003: START - Cookbook sre.mysql.pool pool es2038: Maintenance * 13:10 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2038: Maintenance * 13:08 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, stran: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie * 13:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2047: Maintenance * 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048 cloud-private - filippo@cumin1003" * 13:07 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048 cloud-private - filippo@cumin1003" * 13:06 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] * 13:04 root@cumin1003: START - Cookbook sre.mysql.pool pool es2038: Maintenance * 13:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2078.codfw.wmnet with OS trixie * 13:01 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2038 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95797 and previous config saved to /var/cache/conftool/dbconfig/20260730-125919-cwilliams.json * 12:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2038.codfw.wmnet with reason: Maintenance * 12:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2078.codfw.wmnet with reason: host reimage * 12:37 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2078.codfw.wmnet with reason: host reimage * 12:37 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] (duration: 06m 51s) * 12:33 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 12:32 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 12:32 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 12:32 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:30 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] * 12:19 root@cumin1003: START - Cookbook sre.mysql.pool pool es2047: Maintenance * 12:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2078.codfw.wmnet with OS trixie * 12:18 dcausse@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 12:18 dcausse@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 12:15 dcausse@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 12:14 dcausse@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 12:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2047 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95793 and previous config saved to /var/cache/conftool/dbconfig/20260730-121404-cwilliams.json * 12:13 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2047.codfw.wmnet with reason: Maintenance * 12:13 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2036: Maintenance * 12:05 ayounsi@dns1004: END - running authdns-update * 12:02 ayounsi@dns1004: START - running authdns-update * 11:51 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2077.codfw.wmnet with OS trixie * 11:48 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:46 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1080.eqiad.wmnet with OS trixie * 11:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1226: Maintenance * 11:41 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2077.codfw.wmnet with reason: host reimage * 11:28 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2077.codfw.wmnet with reason: host reimage * 11:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1080.eqiad.wmnet with reason: host reimage * 11:27 root@cumin1003: START - Cookbook sre.mysql.pool pool es2036: Maintenance * 11:27 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2036: Maintenance * 11:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1080.eqiad.wmnet with reason: host reimage * 11:21 root@cumin1003: START - Cookbook sre.mysql.pool pool es2036: Maintenance * 11:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2036 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95786 and previous config saved to /var/cache/conftool/dbconfig/20260730-111633-cwilliams.json * 11:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2036.codfw.wmnet with reason: Maintenance * 11:08 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2077.codfw.wmnet with OS trixie * 11:07 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 11:03 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie * 11:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1226: Maintenance * 10:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1226 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95783 and previous config saved to /var/cache/conftool/dbconfig/20260730-104801-cwilliams.json * 10:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1226.eqiad.wmnet with reason: Maintenance * 10:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1214: Maintenance * 10:27 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2035: Maintenance * 10:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2076.codfw.wmnet with OS trixie * 10:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1214: Maintenance * 09:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1214 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95775 and previous config saved to /var/cache/conftool/dbconfig/20260730-095451-cwilliams.json * 09:54 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1214.eqiad.wmnet with reason: Maintenance * 09:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1209: Maintenance * 09:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2076.codfw.wmnet with reason: host reimage * 09:42 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool es2035: Maintenance * 09:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.netbox.update-extras (exit_code=0) rolling restart_daemons on A:netbox * 09:41 ayounsi@cumin1003: START - Cookbook sre.netbox.update-extras rolling restart_daemons on A:netbox * 09:40 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2035: Maintenance * 09:39 ayounsi@cumin1003: END (PASS) - Cookbook sre.netbox.update-extras (exit_code=0) rolling restart_daemons on A:netbox-canary * 09:39 ayounsi@cumin1003: START - Cookbook sre.netbox.update-extras rolling restart_daemons on A:netbox-canary * 09:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2076.codfw.wmnet with reason: host reimage * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 09:34 root@cumin1003: START - Cookbook sre.mysql.pool pool es2035: Maintenance * 09:32 lucaswerkmeister-wmde@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 09:32 lucaswerkmeister-wmde@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 09:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2035 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95771 and previous config saved to /var/cache/conftool/dbconfig/20260730-092910-cwilliams.json * 09:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2035.codfw.wmnet with reason: Maintenance * 09:19 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2076.codfw.wmnet with OS trixie * 09:18 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 09:17 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie * 09:07 root@cumin1003: START - Cookbook sre.mysql.pool pool db1209: Maintenance * 09:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 23 hosts * 09:04 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Remove cable label from interfaces descriptions - ayounsi@cumin1003 * 09:04 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:02 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Remove cable label from interfaces descriptions - ayounsi@cumin1003 * 09:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1209 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95767 and previous config saved to /var/cache/conftool/dbconfig/20260730-090133-cwilliams.json * 09:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1209.eqiad.wmnet with reason: Maintenance * 09:01 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1192: Maintenance * 08:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1252: Maintenance * 08:57 jayme@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 08:56 jayme@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 08:53 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 23 hosts * 08:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:51 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:50 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1263: Maintenance * 08:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:23 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 08:15 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie * 08:14 root@cumin1003: START - Cookbook sre.mysql.pool pool db1192: Maintenance * 08:13 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2075.codfw.wmnet with OS trixie * 08:12 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1252: Maintenance * 08:11 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1252: Maintenance * 08:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1252: Maintenance * 08:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1192 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95754 and previous config saved to /var/cache/conftool/dbconfig/20260730-080611-cwilliams.json * 08:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1192.eqiad.wmnet with reason: Maintenance * 08:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1178: Maintenance * 08:05 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS bullseye * 07:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1263: Maintenance * 07:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1263 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95751 and previous config saved to /var/cache/conftool/dbconfig/20260730-075106-cwilliams.json * 07:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[1260-1262].eqiad.wmnet with reason: Maintenance * 07:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2075.codfw.wmnet with reason: host reimage * 07:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1263.eqiad.wmnet with reason: Maintenance * 07:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2075.codfw.wmnet with reason: host reimage * 07:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance * 07:38 dcausse@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:38 dcausse@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 07:35 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db1252', diff saved to https://phabricator.wikimedia.org/P95748 and previous config saved to /var/cache/conftool/dbconfig/20260730-073510-marostegui.json * 07:26 klausman@dns2004: END - running authdns-update * 07:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 07:25 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2075.codfw.wmnet with OS trixie * 07:24 klausman@dns2004: START - running authdns-update * 07:23 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host an-test-master1003.eqiad.wmnet * 07:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db1178: Maintenance * 07:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1178 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95746 and previous config saved to /var/cache/conftool/dbconfig/20260730-071112-cwilliams.json * 07:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1178.eqiad.wmnet with reason: Maintenance * 07:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1177: Maintenance * 06:24 root@cumin1003: START - Cookbook sre.mysql.pool pool db1177: Maintenance * 06:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1177 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95741 and previous config saved to /var/cache/conftool/dbconfig/20260730-061736-cwilliams.json * 06:17 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1177.eqiad.wmnet with reason: Maintenance * 06:17 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1172: Maintenance * 05:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1218.eqiad.wmnet with reason: crashed * 05:41 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db1217 it crashed', diff saved to https://phabricator.wikimedia.org/P95737 and previous config saved to /var/cache/conftool/dbconfig/20260730-054111-marostegui.json * 05:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95736 and previous config saved to /var/cache/conftool/dbconfig/20260730-053422-cwilliams.json * 05:30 root@cumin1003: START - Cookbook sre.mysql.pool pool db1172: Maintenance * 05:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95734 and previous config saved to /var/cache/conftool/dbconfig/20260730-052414-cwilliams.json * 05:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1172 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95733 and previous config saved to /var/cache/conftool/dbconfig/20260730-052354-cwilliams.json * 05:23 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1172.eqiad.wmnet with reason: Maintenance * 05:23 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1167: Maintenance * 05:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95731 and previous config saved to /var/cache/conftool/dbconfig/20260730-051406-cwilliams.json * 05:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95729 and previous config saved to /var/cache/conftool/dbconfig/20260730-050358-cwilliams.json * 04:47 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 04:47 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 04:47 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 04:35 root@cumin1003: START - Cookbook sre.mysql.pool pool db1167: Maintenance * 04:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1167 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95726 and previous config saved to /var/cache/conftool/dbconfig/20260730-042923-cwilliams.json * 04:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 04:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1167.eqiad.wmnet with reason: Maintenance * 04:22 pt1979@cumin2002: START - Cookbook sre.dns.netbox * 04:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95725 and previous config saved to /var/cache/conftool/dbconfig/20260730-040337-cwilliams.json * 04:03 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:38 brett@cumin2002: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool eqsin [reason: Switch upgrade maintenance window complete, [[phab:T433097|T433097]]] * 01:38 brett@cumin2002: START - Cookbook sre.dns.admin DNS admin: pool eqsin [reason: Switch upgrade maintenance window complete, [[phab:T433097|T433097]]] == 2026-07-29 == * 23:57 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin,mr1-eqsin IPv6,mr1-eqsin.oob,mr1-eqsin.oob IPv6 with reason: connection issue * 22:54 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 22:53 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 22:53 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 22:53 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:25 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2022.codfw.wmnet, repooling source-only afterwards * 22:20 brett@cumin2002: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool eqsin [reason: Switch upgrade maintenance window, [[phab:T433097|T433097]]] * 22:20 brett@cumin2002: START - Cookbook sre.dns.admin DNS admin: depool eqsin [reason: Switch upgrade maintenance window, [[phab:T433097|T433097]]] * 22:01 apine@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 22:00 apine@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 21:59 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 21:58 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 21:58 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 21:58 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 21:32 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1048.eqiad.wmnet with OS trixie * 21:25 pt1979@cumin2002: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be2097.codfw.wmnet with OS bullseye * 21:16 zabe@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki=metawiki 'Mental Health Resource Center' 'Safety Resource Center/Mental Health' Zabe --reason 'per request [[:phab:T433118{{!}}T433118]]' * 21:12 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2022.codfw.wmnet, repooling source-only afterwards * 21:12 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] (duration: 12m 53s) * 21:12 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2015\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 21:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1253: Maintenance * 21:08 aaron@deploy1003: aaron: Continuing with deployment * 21:01 aaron@deploy1003: aaron: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:59 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] * 20:52 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] (duration: 21m 57s) * 20:48 aaron@deploy1003: aaron: Continuing with deployment * 20:32 aaron@deploy1003: aaron: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:30 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] * 20:24 root@cumin1003: START - Cookbook sre.mysql.pool pool db1253: Maintenance * 20:19 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] (duration: 08m 07s) * 20:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1253 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95719 and previous config saved to /var/cache/conftool/dbconfig/20260729-201810-cwilliams.json * 20:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1253.eqiad.wmnet with reason: Maintenance * 20:17 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1231: Maintenance * 20:15 aaron@deploy1003: bpirkle, aaron: Continuing with deployment * 20:13 aaron@deploy1003: bpirkle, aaron: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:12 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie * 20:11 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] * 20:11 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1048.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:09 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1048.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:09 pt1979@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 20:04 pt1979@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 19:47 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 19:43 pt1979@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye * 19:41 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 19:37 zabe: zabe@deploy1003:~$ mwscript-k8s --comment='[[phab:T433529|T433529]]' --follow -- resetAuthenticationThrottle.php --wiki=aawiki --signup --ip=89.36.114.94 * 19:36 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] (duration: 06m 49s) * 19:32 zabe@deploy1003: zabe: Continuing with deployment * 19:31 zabe@deploy1003: zabe: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:31 root@cumin1003: START - Cookbook sre.mysql.pool pool db1231: Maintenance * 19:29 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] * 19:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1231 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95714 and previous config saved to /var/cache/conftool/dbconfig/20260729-192454-cwilliams.json * 19:24 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1231.eqiad.wmnet with reason: Maintenance * 19:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1227: Maintenance * 19:22 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 19:22 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 19:21 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:21 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1048] - vriley@cumin1003" * 19:21 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1048] - vriley@cumin1003" * 19:19 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 19:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1251: Maintenance * 19:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95711 and previous config saved to /var/cache/conftool/dbconfig/20260729-191756-cwilliams.json * 19:16 vriley@cumin1003: START - Cookbook sre.dns.netbox * 19:11 dduvall: rolling back wmf.13 to group0 due to [[phab:T433457|T433457]] (cc [[phab:T430832|T430832]]) * 19:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95709 and previous config saved to /var/cache/conftool/dbconfig/20260729-190748-cwilliams.json * 19:01 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2022.codfw.wmnet with OS bookworm * 19:01 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2021.codfw.wmnet, repooling source-only afterwards * 18:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95707 and previous config saved to /var/cache/conftool/dbconfig/20260729-185740-cwilliams.json * 18:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95704 and previous config saved to /var/cache/conftool/dbconfig/20260729-184732-cwilliams.json * 18:37 root@cumin1003: START - Cookbook sre.mysql.pool pool db1227: Maintenance * 18:34 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2022.codfw.wmnet with reason: host reimage * 18:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1227 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95701 and previous config saved to /var/cache/conftool/dbconfig/20260729-183117-cwilliams.json * 18:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1227.eqiad.wmnet with reason: Maintenance * 18:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1202: Maintenance * 18:30 root@cumin1003: START - Cookbook sre.mysql.pool pool db1251: Maintenance * 18:27 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2022.codfw.wmnet with reason: host reimage * 18:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1251 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95698 and previous config saved to /var/cache/conftool/dbconfig/20260729-182428-cwilliams.json * 18:24 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lvs2014.codfw.wmnet * 18:24 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for lvs2014.codfw.wmnet * 18:24 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1251.eqiad.wmnet with reason: Maintenance * 18:23 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1235: Maintenance * 18:22 brett@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 18:19 brett@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 18:19 brett@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 18:17 brett@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 18:17 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 18:16 mutante: removing jenkins during the train - living on the edge - no, just kidding, jenkins has migrated to dedicated machines, nothing should happen * 18:15 brett@cumin2002: END (ERROR) - Cookbook sre.loadbalancer.restart-pybal (exit_code=97) rolling-restart of pybal on P<nowiki>{</nowiki>lvs2014.codfw.wmnet<nowiki>}</nowiki> and A:lvs ([[phab:T428495|T428495]]) * 18:15 mutante: CI: contint1002/contint2002: apt-get remove --purge jenkins - jenkins be gone - [[phab:T418521|T418521]] * 18:13 brett@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on P<nowiki>{</nowiki>lvs2014.codfw.wmnet<nowiki>}</nowiki> and A:lvs ([[phab:T428495|T428495]]) * 18:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2022 * 18:08 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2022 * 18:03 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T428495|T428495]] * 18:03 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2022 * 18:02 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2022.codfw.wmnet 211.48.192.10.in-addr.arpa 1.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:02 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2022.codfw.wmnet 211.48.192.10.in-addr.arpa 1.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:02 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:02 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2022 - bking@cumin2003" * 18:02 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2022 - bking@cumin2003" * 17:57 bking@cumin2003: START - Cookbook sre.dns.netbox * 17:56 brett@cumin2002: END (FAIL) - Cookbook sre.loadbalancer.restart-pybal (exit_code=1) rolling-restart of pybal on A:lvs-codfw and A:lvs ([[phab:T428495|T428495]]) * 17:55 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo - [[phab:T428495|T428495]] * 17:54 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2022 * 17:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2022.codfw.wmnet with OS bookworm * 17:50 brett@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on A:lvs-codfw and A:lvs ([[phab:T428495|T428495]]) * 17:47 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2021.codfw.wmnet, repooling source-only afterwards * 17:47 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 14s) * 17:47 swfrench-wmf: authdns-update to direct codfw, eqsin, ulsfo etcd clients back to codfw - [[phab:T428495|T428495]] * 17:47 swfrench@dns1004: END - running authdns-update * 17:47 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 17:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95692 and previous config saved to /var/cache/conftool/dbconfig/20260729-174713-cwilliams.json * 17:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance * 17:46 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1249: Maintenance * 17:45 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2015\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 17:45 swfrench@dns1004: START - running authdns-update * 17:44 root@cumin1003: START - Cookbook sre.mysql.pool pool db1202: Maintenance * 17:41 akhatun: Deployed refinery using scap, then deployed onto hdfs * 17:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1202 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95688 and previous config saved to /var/cache/conftool/dbconfig/20260729-173759-cwilliams.json * 17:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1202.eqiad.wmnet with reason: Maintenance * 17:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1194: Maintenance * 17:37 root@cumin1003: START - Cookbook sre.mysql.pool pool db1235: Maintenance * 17:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1230: Maintenance * 17:30 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1235 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95684 and previous config saved to /var/cache/conftool/dbconfig/20260729-173051-cwilliams.json * 17:30 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1235.eqiad.wmnet with reason: Maintenance * 17:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1234: Maintenance * 17:26 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (thin): Regular analytics weekly train THIN [analytics/refinery@56695674] (duration: 02m 02s) * 17:24 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (thin): Regular analytics weekly train THIN [analytics/refinery@56695674] * 17:23 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567]: Regular analytics weekly train [analytics/refinery@56695674] (duration: 06m 20s) * 17:20 dancy@deploy1003: Finished scap sync-world: Testing delay_messageblobstore_purge: true (duration: 06m 29s) * 17:17 akhatun@deploy1003: Started deploy [analytics/refinery@5669567]: Regular analytics weekly train [analytics/refinery@56695674] * 17:17 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] (duration: 00m 22s) * 17:16 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] * 17:13 dancy@deploy1003: Started scap sync-world: Testing delay_messageblobstore_purge: true * 17:05 mutante: CI: contint1002/contint2002 - restarted httpd to be extra sure all is cleaned up - https://integration.wikimedia.org/ci/ is up and running [[phab:T418521|T418521]] * 17:04 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 17:03 mutante: CI: contint1002/contint2002 - rm /etc/apache2/jenkins_proxy - removing legacy jenkins proxy config - jenkins is on new dedicated machines and uses jenkins_proxy_ext config [[phab:T418521|T418521]] * 17:02 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] (duration: 36m 25s) * 17:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1249: Maintenance * 16:59 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 16:54 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2015.codfw.wmnet, repooling source-only afterwards * 16:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1249 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95674 and previous config saved to /var/cache/conftool/dbconfig/20260729-165339-cwilliams.json * 16:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1249.eqiad.wmnet with reason: Maintenance * 16:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1248: Maintenance * 16:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1194: Maintenance * 16:47 swfrench-wmf: silenced EtcdReplicationDown 57b2b421-1cc9-4e38-9276-{{Gerrit|94f223fd231c}} - [[phab:T428495|T428495]] * 16:46 tchin@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/eventstreams-internal: apply * 16:46 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye * 16:46 tchin@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/eventstreams-internal: apply * 16:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1230: Maintenance * 16:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1194 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95669 and previous config saved to /var/cache/conftool/dbconfig/20260729-164422-cwilliams.json * 16:44 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 16:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1194.eqiad.wmnet with reason: Maintenance * 16:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1191: Maintenance * 16:43 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host an-test-master1003.eqiad.wmnet * 16:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db1234: Maintenance * 16:43 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 16:43 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Rolling back deployment * 16:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host an-test-master1004.eqiad.wmnet * 16:41 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 16:40 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 16:40 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 16:39 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 16:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1230 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95667 and previous config saved to /var/cache/conftool/dbconfig/20260729-163932-cwilliams.json * 16:39 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 16:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1230.eqiad.wmnet with reason: Maintenance * 16:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1207: Maintenance * 16:38 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 16:37 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host an-test-master1004.eqiad.wmnet * 16:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1234 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95664 and previous config saved to /var/cache/conftool/dbconfig/20260729-163719-cwilliams.json * 16:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1234.eqiad.wmnet with reason: Maintenance * 16:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1079.eqiad.wmnet with OS trixie * 16:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1232: Maintenance * 16:34 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1259: Maintenance * 16:28 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] (duration: 06m 57s) * 16:28 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:26 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] * 16:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1051 hosts * 16:21 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] * 16:20 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2006.codfw.wmnet with OS bookworm * 16:19 akhatun: Deploying Refinery at {{Gerrit|56695674}} as part of weekly train * 16:18 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1079.eqiad.wmnet with reason: host reimage * 16:16 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] (duration: 15m 36s) * 16:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2021.codfw.wmnet with OS bookworm * 16:14 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1079.eqiad.wmnet with reason: host reimage * 16:12 topranks: hot-swap line card in FPC0 on cr1-eqiad with replacement MPC10E from Juniper [[phab:T426343|T426343]] * 16:10 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Continuing with deployment * 16:07 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db1248: Maintenance * 16:01 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] * 16:00 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 16:00 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 15:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1248 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95651 and previous config saved to /var/cache/conftool/dbconfig/20260729-155956-cwilliams.json * 15:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1248.eqiad.wmnet with reason: Maintenance * 15:59 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2006.codfw.wmnet with reason: host reimage * 15:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1247: Maintenance * 15:59 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2074.codfw.wmnet with OS trixie * 15:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1191: Maintenance * 15:57 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:55 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1079.eqiad.wmnet with OS trixie * 15:55 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2006.codfw.wmnet with reason: host reimage * 15:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1207: Maintenance * 15:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1191 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95646 and previous config saved to /var/cache/conftool/dbconfig/20260729-155104-cwilliams.json * 15:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1191.eqiad.wmnet with reason: Maintenance * 15:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1181: Maintenance * 15:49 root@cumin1003: START - Cookbook sre.mysql.pool pool db1232: Maintenance * 15:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2021.codfw.wmnet with reason: host reimage * 15:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1207 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95643 and previous config saved to /var/cache/conftool/dbconfig/20260729-154735-cwilliams.json * 15:47 root@cumin1003: START - Cookbook sre.mysql.pool pool db1259: Maintenance * 15:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1207.eqiad.wmnet with reason: Maintenance * 15:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1200: Maintenance * 15:46 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] (duration: 31m 59s) * 15:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 15:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1232 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95640 and previous config saved to /var/cache/conftool/dbconfig/20260729-154330-cwilliams.json * 15:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1232.eqiad.wmnet with reason: Maintenance * 15:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1219: Maintenance * 15:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2015.codfw.wmnet, repooling source-only afterwards * 15:41 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 18s) * 15:41 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1259 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95638 and previous config saved to /var/cache/conftool/dbconfig/20260729-154107-cwilliams.json * 15:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1259.eqiad.wmnet with reason: Maintenance * 15:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1254: Maintenance * 15:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2015.codfw.wmnet with OS bookworm * 15:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2021.codfw.wmnet with reason: host reimage * 15:36 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2006.codfw.wmnet with OS bookworm * 15:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 15:35 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Continuing with deployment * 15:33 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2074.codfw.wmnet with OS trixie * 15:33 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 15:32 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:29 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:28 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be2074.codfw.wmnet with OS trixie * 15:28 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2006.codfw.wmnet * 15:26 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1078.eqiad.wmnet with OS trixie * 15:25 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 15:25 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:22 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2006.codfw.wmnet * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2021 * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2021 * 15:19 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2021 * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2021.codfw.wmnet 210.48.192.10.in-addr.arpa 0.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:19 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2021.codfw.wmnet 210.48.192.10.in-addr.arpa 0.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2021 - bking@cumin2003" * 15:19 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2021 - bking@cumin2003" * 15:14 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] * 15:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2015.codfw.wmnet with reason: host reimage * 15:11 root@cumin1003: START - Cookbook sre.mysql.pool pool db1247: Maintenance * 15:11 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on ml-serve2004.codfw.wmnet with reason: [[phab:T433478|T433478]] * 15:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2015.codfw.wmnet with reason: host reimage * 15:10 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on ml-serve2002.codfw.wmnet with reason: [[phab:T433476|T433476]] * 15:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 15:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1247 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95625 and previous config saved to /var/cache/conftool/dbconfig/20260729-150459-cwilliams.json * 15:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1247.eqiad.wmnet with reason: Maintenance * 15:04 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:04 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1244: Maintenance * 15:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1078.eqiad.wmnet with reason: host reimage * 15:03 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1005.wikimedia.org * 15:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db1181: Maintenance * 15:01 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2098.codfw.wmnet with OS bullseye * 15:00 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye * 15:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1200: Maintenance * 14:59 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 14:59 root@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1285.eqiad.wmnet with OS trixie * 14:59 Amir1: mwscript-k8s -- extensions/TimedMediaHandler/maintenance/requeueTranscodes.php --wiki=commonswiki --key '360p.mpeg4.mov' --throttle --video --missing ([[phab:T358266|T358266]]) * 14:58 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1005.wikimedia.org * 14:58 jhancock@cumin2002: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['ms-be2098'] * 14:58 jhancock@cumin2002: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['ms-be2098'] * 14:58 jhancock@cumin2002: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['ms-be2097'] * 14:58 jhancock@cumin2002: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['ms-be2097'] * 14:58 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1078.eqiad.wmnet with reason: host reimage * 14:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1181 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95621 and previous config saved to /var/cache/conftool/dbconfig/20260729-145629-cwilliams.json * 14:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1181.eqiad.wmnet with reason: Maintenance * 14:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1174: Maintenance * 14:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1219: Maintenance * 14:55 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader1006.wikimedia.org on all recursors * 14:55 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader1006.wikimedia.org on all recursors * 14:55 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader1005.wikimedia.org on all recursors * 14:55 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader1005.wikimedia.org on all recursors * 14:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1200 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95618 and previous config saved to /var/cache/conftool/dbconfig/20260729-145336-cwilliams.json * 14:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1254: Maintenance * 14:53 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1200.eqiad.wmnet with reason: Maintenance * 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1185: Maintenance * 14:52 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2021 * 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2015 * 14:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2015 * 14:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1219 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95616 and previous config saved to /var/cache/conftool/dbconfig/20260729-144946-cwilliams.json * 14:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1219.eqiad.wmnet with reason: Maintenance * 14:49 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1218: Maintenance * 14:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2021.codfw.wmnet with OS bookworm * 14:48 dancy@deploy1003: Finished deploy [zuul/deploy@22703a6]: Deploying https://gerrit.wikimedia.org/r/c/integration/zuul/+/1311501 ([[phab:T432491|T432491]]) (duration: 00m 15s) * 14:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2015.codfw.wmnet with OS bookworm * 14:48 dancy@deploy1003: Started deploy [zuul/deploy@22703a6]: Deploying https://gerrit.wikimedia.org/r/c/integration/zuul/+/1311501 ([[phab:T432491|T432491]]) * 14:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1254 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95613 and previous config saved to /var/cache/conftool/dbconfig/20260729-144729-cwilliams.json * 14:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1254.eqiad.wmnet with reason: Maintenance * 14:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1233: Maintenance * 14:46 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2013\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 14:46 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2014\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 14:46 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:45 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:44 root@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1285.eqiad.wmnet with reason: host reimage * 14:43 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:42 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:41 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:40 root@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1285.eqiad.wmnet with reason: host reimage * 14:39 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1078.eqiad.wmnet with OS trixie * 14:39 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2074.codfw.wmnet with OS trixie * 14:32 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:32 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:32 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:31 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2005.codfw.wmnet with OS bookworm * 14:30 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:30 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:29 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:29 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:27 root@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host db1285 * 14:27 root@cumin1003: START - Cookbook sre.hosts.move-vlan for host db1285 * 14:27 root@cumin1003: START - Cookbook sre.hosts.reimage for host db1285.eqiad.wmnet with OS trixie * 14:24 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 14:24 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:24 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:24 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:23 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:22 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:22 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:21 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:17 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:16 root@cumin1003: START - Cookbook sre.mysql.pool pool db1244: Maintenance * 14:15 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:15 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add asw1-604 loopback ipv4 - pt1979@cumin2002" * 14:15 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add asw1-604 loopback ipv4 - pt1979@cumin2002" * 14:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:12 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 14:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95599 and previous config saved to /var/cache/conftool/dbconfig/20260729-141014-cwilliams.json * 14:10 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 14:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1244.eqiad.wmnet with reason: Maintenance * 14:10 pt1979@cumin2002: START - Cookbook sre.dns.netbox * 14:10 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 14:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1243: Maintenance * 14:09 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2005.codfw.wmnet with reason: host reimage * 14:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db1174: Maintenance * 14:08 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad * 14:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db1185: Maintenance * 14:06 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 14:05 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2005.codfw.wmnet with reason: host reimage * 14:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1174 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95595 and previous config saved to /var/cache/conftool/dbconfig/20260729-140309-cwilliams.json * 14:03 sukhe@dns1004: END - running authdns-update * 14:03 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1174.eqiad.wmnet with reason: Maintenance * 14:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1170: Maintenance * 14:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db1218: Maintenance * 14:01 sukhe@dns1004: START - running authdns-update * 14:00 sukhe@dns1004: START - running authdns-update * 13:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db1233: Maintenance * 13:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1185 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95592 and previous config saved to /var/cache/conftool/dbconfig/20260729-135925-cwilliams.json * 13:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1185.eqiad.wmnet with reason: Maintenance * 13:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1161: Maintenance * 13:58 sukhe@puppetserver1001: conftool action : set/pooled=true; selector: dnsdisc=urldownloader * 13:58 root@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1265.eqiad.wmnet with OS trixie * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1218 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95590 and previous config saved to /var/cache/conftool/dbconfig/20260729-135621-cwilliams.json * 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1218.eqiad.wmnet with reason: Maintenance * 13:55 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1206: Maintenance * 13:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2073.codfw.wmnet with OS trixie * 13:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1233 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95587 and previous config saved to /var/cache/conftool/dbconfig/20260729-135335-cwilliams.json * 13:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1233.eqiad.wmnet with reason: Maintenance * 13:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1229: Maintenance * 13:50 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/kartotherian: apply * 13:50 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service * 13:49 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:49 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/kartotherian: apply * 13:48 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 13:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1077.eqiad.wmnet with OS trixie * 13:47 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 13:46 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2005.codfw.wmnet with OS bookworm * 13:44 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 13:44 root@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1265.eqiad.wmnet with reason: host reimage * 13:40 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] (duration: 09m 22s) * 13:39 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:38 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:36 root@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1265.eqiad.wmnet with reason: host reimage * 13:35 stran@deploy1003: stran: Continuing with deployment * 13:33 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 13:32 stran@deploy1003: stran: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified t * 13:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2073.codfw.wmnet with reason: host reimage * 13:30 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ml-build1001.eqiad.wmnet * 13:30 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] * 13:29 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad * 13:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1077.eqiad.wmnet with reason: host reimage * 13:27 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 13:27 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:27 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad * 13:26 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2073.codfw.wmnet with reason: host reimage * 13:26 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] (duration: 07m 56s) * 13:25 klausman@cumin1003: START - Cookbook sre.hosts.reboot-single for host ml-build1001.eqiad.wmnet * 13:24 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 13:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>ml-serve101[2-5].eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 13:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1015.eqiad.wmnet * 13:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1015.eqiad.wmnet * 13:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1077.eqiad.wmnet with reason: host reimage * 13:23 root@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host db1265 * 13:23 root@cumin1003: START - Cookbook sre.hosts.move-vlan for host db1265 * 13:23 root@cumin1003: START - Cookbook sre.hosts.reimage for host db1265.eqiad.wmnet with OS trixie * 13:23 root@cumin1003: START - Cookbook sre.mysql.pool pool db1243: Maintenance * 13:22 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 13:22 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2005.codfw.wmnet * 13:22 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 13:22 samtar@deploy1003: dreamrimmer, samtar: Continuing with deployment * 13:20 samtar@deploy1003: dreamrimmer, samtar: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts an-test-master[1001-1002].eqiad.wmnet * 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-master[1001-1002].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 13:18 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1015.eqiad.wmnet * 13:18 sukhe@cumin1003: END (ERROR) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=97) for role: url_downloader@eqiad * 13:18 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 13:18 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] * 13:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95574 and previous config saved to /var/cache/conftool/dbconfig/20260729-131638-cwilliams.json * 13:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1243.eqiad.wmnet with reason: Maintenance * 13:16 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2005.codfw.wmnet * 13:16 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1242: Maintenance * 13:14 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] (duration: 07m 00s) * 13:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db1170: Maintenance * 13:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1015.eqiad.wmnet * 13:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1014.eqiad.wmnet * 13:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1014.eqiad.wmnet * 13:12 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1223: Maintenance * 13:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db1161: Maintenance * 13:10 samtar@deploy1003: anzx, samtar: Continuing with deployment * 13:09 samtar@deploy1003: anzx, samtar: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db1206: Maintenance * 13:08 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 13:07 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] * 13:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1170 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95566 and previous config saved to /var/cache/conftool/dbconfig/20260729-130730-cwilliams.json * 13:07 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1170.eqiad.wmnet with reason: Maintenance * 13:07 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:07 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt IPs new switches - cmooney@cumin1003" * 13:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1158: Maintenance * 13:06 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1014.eqiad.wmnet * 13:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1161 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95564 and previous config saved to /var/cache/conftool/dbconfig/20260729-130616-cwilliams.json * 13:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 13:06 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1077.eqiad.wmnet with OS trixie * 13:05 root@cumin1003: START - Cookbook sre.mysql.pool pool db1229: Maintenance * 13:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1161.eqiad.wmnet with reason: Maintenance * 13:05 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt IPs new switches - cmooney@cumin1003" * 13:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2073.codfw.wmnet with OS trixie * 13:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Maintenance * 13:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1206 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95562 and previous config saved to /var/cache/conftool/dbconfig/20260729-130258-cwilliams.json * 13:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1206.eqiad.wmnet with reason: Maintenance * 13:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1196: Maintenance * 13:01 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 13:01 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:00 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 13:00 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 12:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1229 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95559 and previous config saved to /var/cache/conftool/dbconfig/20260729-125950-cwilliams.json * 12:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1229.eqiad.wmnet with reason: Maintenance * 12:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1222: Maintenance * 12:57 sukhe: sudo cumin 'A:lvs and (A:eqiad or A:codfw)' 'disable-puppet "adding new service urldownloader"': [[phab:T429175|T429175]] * 12:56 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1014.eqiad.wmnet * 12:56 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1013.eqiad.wmnet * 12:56 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1013.eqiad.wmnet * 12:50 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1013.eqiad.wmnet * 12:50 sukhe: sudo cumin 'O:url_downloader' 'run-puppet-agent --enable "merging CR 1313948"': [[phab:T429175|T429175]] * 12:48 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-master[1001-1002].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 12:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1013.eqiad.wmnet * 12:45 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1012.eqiad.wmnet * 12:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1012.eqiad.wmnet * 12:45 sukhe: sudo cumin 'O:url_downloader' 'disable-puppet "merging CR 1313948"': [[phab:T429175|T429175]] * 12:44 btullis@cumin1003: START - Cookbook sre.dns.netbox * 12:40 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test2001.codfw.wmnet * 12:40 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test2001.codfw.wmnet * 12:38 ayounsi@dns1004: END - running authdns-update * 12:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1012.eqiad.wmnet * 12:35 ayounsi@dns1004: START - running authdns-update * 12:34 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts an-test-master[1001-1002].eqiad.wmnet * 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts an-test-coord1001.eqiad.wmnet * 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-coord1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 12:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1012.eqiad.wmnet * 12:32 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>ml-serve101[2-5].eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 12:29 root@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Maintenance * 12:25 root@cumin1003: START - Cookbook sre.mysql.pool pool db1223: Maintenance * 12:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95544 and previous config saved to /var/cache/conftool/dbconfig/20260729-122254-cwilliams.json * 12:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1242.eqiad.wmnet with reason: Maintenance * 12:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1241: Maintenance * 12:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1051 hosts * 12:20 root@cumin1003: START - Cookbook sre.mysql.pool pool db1158: Maintenance * 12:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1223 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95540 and previous config saved to /var/cache/conftool/dbconfig/20260729-121937-cwilliams.json * 12:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1223.eqiad.wmnet with reason: Maintenance * 12:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1212: Maintenance * 12:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db1159: Maintenance * 12:17 elukey@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: sync * 12:15 elukey@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: sync * 12:15 root@cumin1003: START - Cookbook sre.mysql.pool pool db1196: Maintenance * 12:14 Daimona: Creating new DB tables for the CampaignEvents extension in x1.testwiki, x1.test2wiki, x1.officewiki, and x1.wikishared # [[phab:T429339|T429339]] * 12:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db1222: Maintenance * 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95535 and previous config saved to /var/cache/conftool/dbconfig/20260729-121211-cwilliams.json * 12:12 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 12:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1158.eqiad.wmnet with reason: Maintenance * 12:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1159 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95534 and previous config saved to /var/cache/conftool/dbconfig/20260729-121146-cwilliams.json * 12:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1159.eqiad.wmnet with reason: Maintenance * 12:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1196 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95533 and previous config saved to /var/cache/conftool/dbconfig/20260729-120847-cwilliams.json * 12:08 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 12:08 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1196.eqiad.wmnet with reason: Maintenance * 12:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1195: Maintenance * 12:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1222 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95530 and previous config saved to /var/cache/conftool/dbconfig/20260729-120424-cwilliams.json * 12:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1222.eqiad.wmnet with reason: Maintenance * 12:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1098 hosts * 12:00 marostegui: Rename tables [[phab:T425074|T425074]] * 12:00 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-coord1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 11:58 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1197: Maintenance * 11:55 btullis@cumin1003: START - Cookbook sre.dns.netbox * 11:52 elukey@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: sync * 11:51 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:51 elukey@deploy1003: helmfile [codfw] START helmfile.d/services/proton: sync * 11:51 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:50 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts an-test-coord1001.eqiad.wmnet * 11:50 elukey@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: sync * 11:49 elukey@deploy1003: helmfile [staging] START helmfile.d/services/proton: sync * 11:35 root@cumin1003: START - Cookbook sre.mysql.pool pool db1241: Maintenance * 11:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db1212: Maintenance * 11:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1241 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95520 and previous config saved to /var/cache/conftool/dbconfig/20260729-112918-cwilliams.json * 11:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1241.eqiad.wmnet with reason: Maintenance * 11:29 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1238: Maintenance * 11:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1212 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95517 and previous config saved to /var/cache/conftool/dbconfig/20260729-112727-cwilliams.json * 11:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 11:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1212.eqiad.wmnet with reason: Maintenance * 11:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1198: Maintenance * 11:23 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:22 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:21 root@cumin1003: START - Cookbook sre.mysql.pool pool db1195: Maintenance * 11:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1195 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95514 and previous config saved to /var/cache/conftool/dbconfig/20260729-111450-cwilliams.json * 11:14 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1195.eqiad.wmnet with reason: Maintenance * 11:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1186: Maintenance * 11:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 11:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 11:05 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:54 marostegui: Dropping renamed tables [[phab:T425066|T425066]] * 10:41 root@cumin1003: START - Cookbook sre.mysql.pool pool db1238: Maintenance * 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1198: Maintenance * 10:39 Amir1: ran https://phabricator.wikimedia.org/T432509#12149723 in production ([[phab:T432509|T432509]]) * 10:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db1197: Maintenance * 10:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1238 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95501 and previous config saved to /var/cache/conftool/dbconfig/20260729-103532-cwilliams.json * 10:35 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1238.eqiad.wmnet with reason: Maintenance * 10:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1221: Maintenance * 10:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1198 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95499 and previous config saved to /var/cache/conftool/dbconfig/20260729-103330-cwilliams.json * 10:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1198.eqiad.wmnet with reason: Maintenance * 10:33 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1175: Maintenance * 10:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1197 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95496 and previous config saved to /var/cache/conftool/dbconfig/20260729-103217-cwilliams.json * 10:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1197.eqiad.wmnet with reason: Maintenance * 10:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1188: Maintenance * 10:27 root@cumin1003: START - Cookbook sre.mysql.pool pool db1186: Maintenance * 10:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1186 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95493 and previous config saved to /var/cache/conftool/dbconfig/20260729-102111-cwilliams.json * 10:21 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1186.eqiad.wmnet with reason: Maintenance * 10:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 10:14 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 09:53 XioNoX: reboot cr2-magru - [[phab:T431750|T431750]] * 09:52 XioNoX: drain cr2-magru - [[phab:T431750|T431750]] * 09:48 root@cumin1003: START - Cookbook sre.mysql.pool pool db1221: Maintenance * 09:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zookeeper-test1002.eqiad.wmnet * 09:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1188: Maintenance * 09:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1175: Maintenance * 09:44 btullis@dns1004: END - running authdns-update * 09:42 btullis@dns1004: START - running authdns-update * 09:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1221 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95483 and previous config saved to /var/cache/conftool/dbconfig/20260729-094200-cwilliams.json * 09:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 7 hosts with reason: Maintenance * 09:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1221.eqiad.wmnet with reason: Maintenance * 09:41 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host zookeeper-test1002.eqiad.wmnet * 09:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1199: Maintenance * 09:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1188 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95481 and previous config saved to /var/cache/conftool/dbconfig/20260729-093917-cwilliams.json * 09:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1188.eqiad.wmnet with reason: Maintenance * 09:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1182: Maintenance * 09:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1175 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95479 and previous config saved to /var/cache/conftool/dbconfig/20260729-093842-cwilliams.json * 09:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1175.eqiad.wmnet with reason: Maintenance * 09:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1166: Maintenance * 09:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1169: Maintenance * 09:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1033.eqiad.wmnet,service=s8 * 09:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1033.eqiad.wmnet,service=s5 * 09:33 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1033.eqiad.wmnet,service=s8 * 09:33 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1033.eqiad.wmnet,service=s5 * 09:21 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:21 XioNoX: reboot cr1-magru - [[phab:T431750|T431750]] * 09:17 XioNoX: drain cr1-magru - [[phab:T431750|T431750]] * 09:15 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm1001.wikimedia.org * 09:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply * 09:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply * 09:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 09:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 09:11 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr2-magru,cr2-magru IPv6,cr2-magru.mgmt with reason: router upgrade * 09:11 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm1001.wikimedia.org * 09:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp1005.wikimedia.org * 09:07 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp1005.wikimedia.org * 09:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp2005.wikimedia.org * 09:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp2005.wikimedia.org * 09:00 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr1-magru,cr1-magru IPv6,cr1-magru.mgmt with reason: router upgrade * 09:00 marostegui: Dropping renamed tables [[phab:T426341|T426341]] * 08:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1199: Maintenance * 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1182: Maintenance * 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1166: Maintenance * 08:47 root@cumin1003: START - Cookbook sre.mysql.pool pool db1169: Maintenance * 08:46 ayounsi@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 1:00:00 on cr1-magru,cr1-magru IPv6,cr1-magru.mgmt with reason: router upgrade * 08:45 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2072.codfw.wmnet with OS trixie * 08:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1199 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95464 and previous config saved to /var/cache/conftool/dbconfig/20260729-084534-cwilliams.json * 08:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1199.eqiad.wmnet with reason: Maintenance * 08:45 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 08:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1190: Maintenance * 08:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1182 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95462 and previous config saved to /var/cache/conftool/dbconfig/20260729-084436-cwilliams.json * 08:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1182.eqiad.wmnet with reason: Maintenance * 08:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1156: Maintenance * 08:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1166 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95460 and previous config saved to /var/cache/conftool/dbconfig/20260729-084400-cwilliams.json * 08:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1166.eqiad.wmnet with reason: Maintenance * 08:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1157: Maintenance * 08:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95458 and previous config saved to /var/cache/conftool/dbconfig/20260729-084147-cwilliams.json * 08:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1169.eqiad.wmnet with reason: Maintenance * 08:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1163: Maintenance * 08:30 btullis@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 11 hosts with reason: Replacing the namenodes * 08:23 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2072.codfw.wmnet with reason: host reimage * 08:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1098 hosts * 08:19 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2072.codfw.wmnet with reason: host reimage * 07:58 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2072.codfw.wmnet with OS trixie * 07:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1190: Maintenance * 07:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1156: Maintenance * 07:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1157: Maintenance * 07:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1163: Maintenance * 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1190 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95444 and previous config saved to /var/cache/conftool/dbconfig/20260729-074930-cwilliams.json * 07:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1190.eqiad.wmnet with reason: Maintenance * 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1157 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95443 and previous config saved to /var/cache/conftool/dbconfig/20260729-074914-cwilliams.json * 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1156 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95442 and previous config saved to /var/cache/conftool/dbconfig/20260729-074906-cwilliams.json * 07:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1157.eqiad.wmnet with reason: Maintenance * 07:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 07:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1156.eqiad.wmnet with reason: Maintenance * 07:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1163 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95441 and previous config saved to /var/cache/conftool/dbconfig/20260729-074652-cwilliams.json * 07:46 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1163.eqiad.wmnet with reason: Maintenance * 07:46 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2034.codfw.wmnet * 07:42 ayounsi@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2034.codfw.wmnet * 07:42 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2232.codfw.wmnet with OS trixie * 07:34 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:34 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:33 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:31 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:19 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2232.codfw.wmnet with reason: host reimage * 07:15 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2232.codfw.wmnet with reason: host reimage * 06:58 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db2232.codfw.wmnet with OS trixie * 06:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[2160,2232].codfw.wmnet with reason: Reimage * 06:26 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1164.eqiad.wmnet with OS trixie * 06:05 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1164.eqiad.wmnet with reason: host reimage * 06:01 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1164.eqiad.wmnet with reason: host reimage * 05:47 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1164.eqiad.wmnet with OS trixie * 05:46 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1164.eqiad.wmnet with reason: Reimage == 2026-07-28 == * 22:50 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1138.eqiad.wmnet * 22:50 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1138.eqiad.wmnet * 22:49 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1138.eqiad.wmnet * 22:11 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2014.codfw.wmnet, repooling source-only afterwards * 22:08 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2013.codfw.wmnet, repooling source-only afterwards * 22:03 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 20:58 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] (duration: 08m 19s) * 20:55 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2014.codfw.wmnet, repooling source-only afterwards * 20:55 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2013.codfw.wmnet, repooling source-only afterwards * 20:54 arlolra@deploy1003: arlolra: Continuing with deployment * 20:54 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 14s) * 20:54 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 20:53 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 30s) * 20:53 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 20:52 arlolra@deploy1003: arlolra: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:51 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:50 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] * 20:49 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:43 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 20:34 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] (duration: 06m 54s) * 20:34 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:34 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:31 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:31 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:30 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:30 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:30 arlolra@deploy1003: arlolra: Continuing with deployment * 20:29 arlolra@deploy1003: arlolra: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:27 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] * 20:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2014.codfw.wmnet with OS bookworm * 20:21 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 20:21 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:20 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 20:19 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:19 swfrench-wmf: switched etcd-mirror replication from conf2005 to conf2004 - [[phab:T428495|T428495]] * 20:17 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:17 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:15 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] (duration: 08m 26s) * 20:12 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:11 arlolra@deploy1003: anzx, arlolra: Continuing with deployment * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2013.codfw.wmnet with OS bookworm * 20:09 arlolra@deploy1003: anzx, arlolra: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] * 20:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2216: Maintenance * 19:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2014.codfw.wmnet with reason: host reimage * 19:57 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:54 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:54 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:53 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:52 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2014.codfw.wmnet with reason: host reimage * 19:49 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2013.codfw.wmnet with reason: host reimage * 19:42 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 19:41 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:41 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2013.codfw.wmnet with reason: host reimage * 19:39 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:39 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:39 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-eqiad: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 19:38 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2014 * 19:33 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2014 * 19:29 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2014.codfw.wmnet with OS bookworm * 19:28 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:27 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1006 * 19:26 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2012\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 19:26 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1006 * 19:26 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:26 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 19:25 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2013 * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2013 * 19:21 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2013 * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2013.codfw.wmnet 84.0.192.10.in-addr.arpa 4.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:21 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2013.codfw.wmnet 84.0.192.10.in-addr.arpa 4.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2013 - bking@cumin2003" * 19:21 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2013 - bking@cumin2003" * 19:21 vriley@cumin1003: START - Cookbook sre.dns.netbox * 19:20 root@cumin1003: START - Cookbook sre.mysql.pool pool db2216: Maintenance * 19:13 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2216 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95435 and previous config saved to /var/cache/conftool/dbconfig/20260728-191343-cwilliams.json * 19:13 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2216.codfw.wmnet with reason: Maintenance * 19:13 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2203: Maintenance * 19:06 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1005.eqiad.wmnet with OS trixie * 19:06 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 19:06 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 18:46 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 18:45 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:45 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:43 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:40 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:36 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-eqiad: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 18:35 dancy@deploy1003: Installation of scap version "4.275.0" completed for 3 hosts * 18:33 dancy@deploy1003: Installing scap version "4.275.0" for 3 host(s) * 18:32 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:32 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2097.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:30 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns3003.wikimedia.org [reason: pool for all services after reimaging] * 18:29 sukhe@dns1004: END - running authdns-update * 18:27 sukhe@dns1004: START - running authdns-update * 18:27 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns3003.wikimedia.org,service=authdns-update [reason: pool authdns-update after reimaging] * 18:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db2203: Maintenance * 18:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2203 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95430 and previous config saved to /var/cache/conftool/dbconfig/20260728-181958-cwilliams.json * 18:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2203.codfw.wmnet with reason: Maintenance * 18:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2188: Maintenance * 18:18 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2097.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:17 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-be2098 * 18:17 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host ms-be2098 * 18:17 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-be2097 * 18:16 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host ms-be2097 * 18:15 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:15 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding ms-be2097-8 to codfw - jhancock@cumin2002" * 18:15 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding ms-be2097-8 to codfw - jhancock@cumin2002" * 18:10 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 18:08 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage * 18:05 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns3003.wikimedia.org with OS trixie * 18:03 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage * 17:56 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-codfw: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 17:45 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie * 17:45 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1005.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:41 sukhe@dns1004: END - running authdns-update * 17:39 sukhe@dns1004: START - running authdns-update * 17:36 sukhe@puppetserver1001: conftool action : set/weight=1; selector: cluster=urldownloader,service=squid * 17:36 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1005.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:35 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader,service=squid * 17:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 17:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1005 * 17:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 17:34 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1005 * 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1005] - vriley@cumin1003" * 17:34 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1005] - vriley@cumin1003" * 17:32 root@cumin1003: START - Cookbook sre.mysql.pool pool db2188: Maintenance * 17:29 vriley@cumin1003: START - Cookbook sre.dns.netbox * 17:29 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 17:26 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2188 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95425 and previous config saved to /var/cache/conftool/dbconfig/20260728-172609-cwilliams.json * 17:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2188.codfw.wmnet with reason: Maintenance * 17:25 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2176: Maintenance * 17:19 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1138.eqiad.wmnet with OS trixie * 17:18 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1005 * 17:18 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1005 * 17:18 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:15 vriley@cumin1003: START - Cookbook sre.dns.netbox * 17:13 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns3003.wikimedia.org with reason: host reimage * 17:07 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns3003.wikimedia.org with reason: host reimage * 17:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-eqiad * 17:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1015.eqiad.wmnet * 17:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1015.eqiad.wmnet * 16:59 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1138.eqiad.wmnet with reason: host reimage * 16:55 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-codfw: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 16:54 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1138.eqiad.wmnet with reason: host reimage * 16:53 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1015.eqiad.wmnet * 16:43 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns3003.wikimedia.org with OS trixie * 16:43 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1015.eqiad.wmnet * 16:43 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1014.eqiad.wmnet * 16:43 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1014.eqiad.wmnet * 16:43 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=dns3003.wikimedia.org [reason: depooling for reimage to trixie] * 16:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 290 hosts * 16:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db2176: Maintenance * 16:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1138 * 16:38 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1138 * 16:37 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1138 * 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1138.eqiad.wmnet 193.32.64.10.in-addr.arpa 3.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:37 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1138.eqiad.wmnet 193.32.64.10.in-addr.arpa 3.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1138 - jiji@cumin1003" * 16:37 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1138 - jiji@cumin1003" * 16:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1014.eqiad.wmnet * 16:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2176 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95420 and previous config saved to /var/cache/conftool/dbconfig/20260728-163235-cwilliams.json * 16:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2176.codfw.wmnet with reason: Maintenance * 16:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1014.eqiad.wmnet * 16:32 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1013.eqiad.wmnet * 16:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1013.eqiad.wmnet * 16:32 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2174: Maintenance * 16:28 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2012.codfw.wmnet, repooling source-only afterwards * 16:25 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1013.eqiad.wmnet * 16:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1013.eqiad.wmnet * 16:20 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1012.eqiad.wmnet * 16:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1012.eqiad.wmnet * 16:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1012.eqiad.wmnet * 16:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1012.eqiad.wmnet * 16:03 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1011.eqiad.wmnet * 16:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1011.eqiad.wmnet * 16:00 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2004.codfw.wmnet with OS bookworm * 15:59 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1011.eqiad.wmnet * 15:56 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:55 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 15:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:54 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1011.eqiad.wmnet * 15:54 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1010.eqiad.wmnet * 15:54 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1010.eqiad.wmnet * 15:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:50 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 15:49 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1010.eqiad.wmnet * 15:48 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 15:48 jiji@cumin1003: START - Cookbook sre.dns.netbox * 15:46 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2248: Maintenance * 15:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db2174: Maintenance * 15:44 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1010.eqiad.wmnet * 15:44 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1009.eqiad.wmnet * 15:44 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1009.eqiad.wmnet * 15:42 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1138 * 15:41 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1138.eqiad.wmnet with OS trixie * 15:39 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1009.eqiad.wmnet * 15:39 robh@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on arclamp2001.codfw.wmnet with reason: ram upgrade * 15:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2174 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95413 and previous config saved to /var/cache/conftool/dbconfig/20260728-153844-cwilliams.json * 15:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2174.codfw.wmnet with reason: Maintenance * 15:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2173: Maintenance * 15:37 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 15:35 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 15:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1009.eqiad.wmnet * 15:34 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1008.eqiad.wmnet * 15:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1008.eqiad.wmnet * 15:31 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 15:31 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 15:29 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1008.eqiad.wmnet * 15:27 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 290 hosts * 15:25 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2013 * 15:25 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2195: Maintenance * 15:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1008.eqiad.wmnet * 15:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1007.eqiad.wmnet * 15:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1007.eqiad.wmnet * 15:22 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2004.codfw.wmnet with reason: host reimage * 15:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2013.codfw.wmnet with OS bookworm * 15:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 15:19 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2004.codfw.wmnet with reason: host reimage * 15:19 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 15:17 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1007.eqiad.wmnet * 15:12 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1007.eqiad.wmnet * 15:12 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1006.eqiad.wmnet * 15:12 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1006.eqiad.wmnet * 15:11 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1138.eqiad.wmnet * 15:11 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 15:11 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1138.eqiad.wmnet * 15:11 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1138.eqiad.wmnet * 15:10 brennen@deploy1003: Finished deploy [phabricator/deployment@f8b349f]: deploy phab1004 for [[phab:T433382|T433382]] (duration: 00m 43s) * 15:10 brennen@deploy1003: Started deploy [phabricator/deployment@f8b349f]: deploy phab1004 for [[phab:T433382|T433382]] * 15:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts * 15:09 brennen@deploy1003: Finished deploy [phabricator/deployment@f8b349f]: deploy phab2003 for [[phab:T433382|T433382]] (duration: 00m 55s) * 15:08 brennen@deploy1003: Started deploy [phabricator/deployment@f8b349f]: deploy phab2003 for [[phab:T433382|T433382]] * 15:07 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2012.codfw.wmnet, repooling source-only afterwards * 15:07 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts * 15:06 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1137.eqiad.wmnet * 15:06 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1137.eqiad.wmnet * 15:06 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1137.eqiad.wmnet * 15:05 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1006.eqiad.wmnet * 15:05 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1022\.eqiad\.wmnet,dc=eqiad,cluster=wdqs\-main,service=wdqs\-main * 15:01 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab2003.codfw.wmnet with reason: deployment * 15:01 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1005.eqiad.wmnet with reason: deployment * 15:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1006.eqiad.wmnet * 15:00 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1006.eqiad.wmnet with reason: deployment * 15:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1005.eqiad.wmnet * 15:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1005.eqiad.wmnet * 14:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db2248: Maintenance * 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1004.eqiad.wmnet with reason: deployment * 14:59 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2004.codfw.wmnet with OS bookworm * 14:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1005.eqiad.wmnet * 14:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2248 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95403 and previous config saved to /var/cache/conftool/dbconfig/20260728-145532-cwilliams.json * 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2245-2247].codfw.wmnet with reason: Maintenance * 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2248.codfw.wmnet with reason: Maintenance * 14:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2240: Maintenance * 14:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2173: Maintenance * 14:51 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts * 14:50 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1005.eqiad.wmnet * 14:50 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1004.eqiad.wmnet * 14:50 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1004.eqiad.wmnet * 14:49 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts * 14:45 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2004.codfw.wmnet * 14:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2173 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95399 and previous config saved to /var/cache/conftool/dbconfig/20260728-144453-cwilliams.json * 14:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2173.codfw.wmnet with reason: Maintenance * 14:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2170: Maintenance * 14:44 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1004.eqiad.wmnet * 14:39 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2004.codfw.wmnet * 14:38 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1004.eqiad.wmnet * 14:38 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet * 14:38 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet * 14:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db2195: Maintenance * 14:36 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1022.eqiad.wmnet, repooling source-only afterwards * 14:36 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 14:33 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet * 14:33 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2222: Maintenance * 14:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2195 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95394 and previous config saved to /var/cache/conftool/dbconfig/20260728-143218-cwilliams.json * 14:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2195.codfw.wmnet with reason: Maintenance * 14:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2181: Maintenance * 14:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 14:30 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 14:25 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 14:25 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 14:25 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 14:23 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet * 14:23 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1002.eqiad.wmnet * 14:23 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1002.eqiad.wmnet * 14:23 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 19s) * 14:23 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:18 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1002.eqiad.wmnet * 14:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2012.codfw.wmnet with OS bookworm * 14:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1002.eqiad.wmnet * 14:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1001.eqiad.wmnet * 14:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1001.eqiad.wmnet * 14:11 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 14:11 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 14:08 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1001.eqiad.wmnet * 14:07 elukey@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'. * 14:07 elukey@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'. * 14:06 elukey@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'. * 14:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db2240: Maintenance * 14:06 XioNoX: un-drain cr2-esams - [[phab:T431751|T431751]] * 14:05 elukey@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'. * 14:02 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1001.eqiad.wmnet * 14:02 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-eqiad * 14:01 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 14:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2240 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95384 and previous config saved to /var/cache/conftool/dbconfig/20260728-140011-cwilliams.json * 14:00 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2240.codfw.wmnet with reason: Maintenance * 13:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2237: Maintenance * 13:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db2170: Maintenance * 13:56 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 13:55 XioNoX: reboot cr2-esams - [[phab:T431751|T431751]] * 13:52 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 13:51 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr2-esams,cr2-esams IPv6,cr2-esams.mgmt with reason: router upgrade * 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 13:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2170 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95381 and previous config saved to /var/cache/conftool/dbconfig/20260728-135043-cwilliams.json * 13:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2170.codfw.wmnet with reason: Maintenance * 13:50 XioNoX: drain cr2-esams - [[phab:T431751|T431751]] * 13:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2153: Maintenance * 13:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2012.codfw.wmnet with reason: host reimage * 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2222: Maintenance * 13:45 sukhe: restart pybal on A:lvs-codfw * 13:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2012.codfw.wmnet with reason: host reimage * 13:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db2181: Maintenance * 13:44 btullis@dns1004: END - running authdns-update * 13:42 sukhe: restart pybal on lvs2014 * 13:42 btullis@dns1004: START - running authdns-update * 13:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2222 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95376 and previous config saved to /var/cache/conftool/dbconfig/20260728-133948-cwilliams.json * 13:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2222.codfw.wmnet with reason: Maintenance * 13:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2221: Maintenance * 13:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2181 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95374 and previous config saved to /var/cache/conftool/dbconfig/20260728-133857-cwilliams.json * 13:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2181.codfw.wmnet with reason: Maintenance * 13:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2167: Maintenance * 13:30 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T428495|T428495]] * 13:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 13:29 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 13:29 ayounsi@cumin1003: END (FAIL) - Cookbook sre.dns.admin (exit_code=99) DNS admin: depool esams [reason: router upgrade, [[phab:T431749|T431749]]] * 13:28 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: router upgrade, [[phab:T431749|T431749]]] * 13:27 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1022.eqiad.wmnet, repooling source-only afterwards * 13:27 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo - [[phab:T428495|T428495]] * 13:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2012 * 13:27 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2012 * 13:21 lucaswerkmeister-wmde@deploy1003: mwscript-k8s job started: cleanupTitles bolwiki # [[phab:T429951|T429951]] * 13:21 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] (duration: 07m 19s) * 13:20 swfrench-wmf: authdns-update to direct codfw, eqsin, ulsfo etcd clients to eqiad - [[phab:T428495|T428495]] * 13:18 swfrench@dns1004: END - running authdns-update * 13:17 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, anzx: Continuing with deployment * 13:16 swfrench@dns1004: START - running authdns-update * 13:16 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2012 * 13:16 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2012.codfw.wmnet 57.48.192.10.in-addr.arpa 7.5.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:16 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2012.codfw.wmnet 57.48.192.10.in-addr.arpa 7.5.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:16 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:16 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, anzx: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:14 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 13:14 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] * 13:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db2237: Maintenance * 13:13 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2237: Maintenance * 13:13 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:12 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:12 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback IPV6 for asw1-604 - pt1979@cumin2003" * 13:12 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:12 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback IPV6 for asw1-604 - pt1979@cumin2003" * 13:11 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 20s) * 13:11 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 13:10 esanders@deploy1003: Finished scap sync-world: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] (duration: 08m 11s) * 13:08 pt1979@cumin2003: START - Cookbook sre.dns.netbox * 13:07 root@cumin1003: START - Cookbook sre.mysql.pool pool db2237: Maintenance * 13:06 esanders@deploy1003: esanders: Continuing with deployment * 13:04 esanders@deploy1003: esanders: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2153: Maintenance * 13:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2153: Maintenance * 13:02 esanders@deploy1003: Started scap sync-world: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] * 13:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2237 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95362 and previous config saved to /var/cache/conftool/dbconfig/20260728-130107-cwilliams.json * 13:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2237.codfw.wmnet with reason: Maintenance * 13:00 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2236: Maintenance * 12:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2153: Maintenance * 12:52 root@cumin1003: START - Cookbook sre.mysql.pool pool db2221: Maintenance * 12:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2153 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95358 and previous config saved to /var/cache/conftool/dbconfig/20260728-125214-cwilliams.json * 12:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2153.codfw.wmnet with reason: Maintenance * 12:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2167: Maintenance * 12:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-codfw * 12:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2011.codfw.wmnet * 12:51 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2011.codfw.wmnet * 12:49 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:49 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback for asw1-603 - pt1979@cumin2003" * 12:48 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback for asw1-603 - pt1979@cumin2003" * 12:46 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2011.codfw.wmnet * 12:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2221 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95357 and previous config saved to /var/cache/conftool/dbconfig/20260728-124601-cwilliams.json * 12:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2221.codfw.wmnet with reason: Maintenance * 12:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2218: Maintenance * 12:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2167 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95354 and previous config saved to /var/cache/conftool/dbconfig/20260728-124457-cwilliams.json * 12:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2167.codfw.wmnet with reason: Maintenance * 12:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2166: Maintenance * 12:42 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 12:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2011.codfw.wmnet * 12:41 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2010.codfw.wmnet * 12:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2010.codfw.wmnet * 12:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2010.codfw.wmnet * 12:34 pt1979@cumin2003: START - Cookbook sre.dns.netbox * 12:32 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 12:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2010.codfw.wmnet * 12:31 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2009.codfw.wmnet * 12:31 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2009.codfw.wmnet * 12:27 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2009.codfw.wmnet * 12:22 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2009.codfw.wmnet * 12:22 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2008.codfw.wmnet * 12:21 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2008.codfw.wmnet * 12:16 pt1979@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-604-eqsin * 12:16 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2008.codfw.wmnet * 12:16 pt1979@cumin1003: START - Cookbook sre.network.tls for network device asw1-604-eqsin * 12:14 root@cumin1003: START - Cookbook sre.mysql.pool pool db2236: Maintenance * 12:14 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2236: Maintenance * 12:12 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1137.eqiad.wmnet with OS trixie * 12:11 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2008.codfw.wmnet * 12:11 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2007.codfw.wmnet * 12:11 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2007.codfw.wmnet * 12:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db2236: Maintenance * 12:06 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2007.codfw.wmnet * 12:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2236 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95348 and previous config saved to /var/cache/conftool/dbconfig/20260728-120253-cwilliams.json * 12:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2236.codfw.wmnet with reason: Maintenance * 12:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2007.codfw.wmnet * 12:01 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 12:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 11:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2218: Maintenance * 11:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db2166: Maintenance * 11:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2219: Maintenance * 11:56 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2006.codfw.wmnet * 11:52 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1137.eqiad.wmnet with reason: host reimage * 11:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2218 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95344 and previous config saved to /var/cache/conftool/dbconfig/20260728-115155-cwilliams.json * 11:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2218.codfw.wmnet with reason: Maintenance * 11:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2208: Maintenance * 11:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2166 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95342 and previous config saved to /var/cache/conftool/dbconfig/20260728-115119-cwilliams.json * 11:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2166.codfw.wmnet with reason: Maintenance * 11:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2164: Maintenance * 11:47 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1137.eqiad.wmnet with reason: host reimage * 11:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2006.codfw.wmnet * 11:45 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2005.codfw.wmnet * 11:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2005.codfw.wmnet * 11:40 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2005.codfw.wmnet * 11:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2005.codfw.wmnet * 11:35 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2004.codfw.wmnet * 11:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2004.codfw.wmnet * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1137 * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1137 * 11:30 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1137 * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1137.eqiad.wmnet 192.32.64.10.in-addr.arpa 2.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:30 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1137.eqiad.wmnet 192.32.64.10.in-addr.arpa 2.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1137 - jiji@cumin1003" * 11:25 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2004.codfw.wmnet * 11:19 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2004.codfw.wmnet * 11:19 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2003.codfw.wmnet * 11:19 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2003.codfw.wmnet * 11:14 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2003.codfw.wmnet * 11:11 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2219: Maintenance * 11:10 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2219: Maintenance * 11:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2219: Maintenance * 11:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2003.codfw.wmnet * 11:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2164: Maintenance * 11:03 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2002.codfw.wmnet * 11:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2002.codfw.wmnet * 11:03 root@cumin1003: START - Cookbook sre.mysql.pool pool db2208: Maintenance * 10:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2164 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95332 and previous config saved to /var/cache/conftool/dbconfig/20260728-105749-cwilliams.json * 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2164.codfw.wmnet with reason: Maintenance * 10:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2208 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95331 and previous config saved to /var/cache/conftool/dbconfig/20260728-105711-cwilliams.json * 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2208.codfw.wmnet with reason: Maintenance * 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2219 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95330 and previous config saved to /var/cache/conftool/dbconfig/20260728-105652-cwilliams.json * 10:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2219.codfw.wmnet with reason: Maintenance * 10:53 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1137 - jiji@cumin1003" * 10:52 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2002.codfw.wmnet * 10:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2002.codfw.wmnet * 10:47 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2001.codfw.wmnet * 10:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2001.codfw.wmnet * 10:39 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2001.codfw.wmnet * 10:35 jiji@cumin1003: START - Cookbook sre.dns.netbox * 10:34 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] (duration: 09m 31s) * 10:34 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1137 * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2001.codfw.wmnet * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-codfw * 10:34 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1137.eqiad.wmnet with OS trixie * 10:28 jforrester@deploy1003: jforrester: Continuing with deployment * 10:27 jforrester@deploy1003: jforrester: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:25 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] * 10:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 10:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 10:21 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 10:20 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1137.eqiad.wmnet * 10:20 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 10:20 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1137.eqiad.wmnet * 10:20 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1137.eqiad.wmnet * 10:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2163: Maintenance * 09:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-staging-worker * 09:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2003.codfw.wmnet * 09:37 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2003.codfw.wmnet * 09:32 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 09:31 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2003.codfw.wmnet * 09:30 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 09:30 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2163: Maintenance * 09:30 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 09:30 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 09:30 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:22 klausman@cumin1003: END (ERROR) - Cookbook sre.ganeti.reboot-vm (exit_code=97) for VM ml-serve-ctrl2001.codfw.wmnet * 09:22 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2001.codfw.wmnet * 09:22 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d8-eqiad * 09:22 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d8-eqiad * 09:21 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2003.codfw.wmnet * 09:20 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2002.codfw.wmnet * 09:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2002.codfw.wmnet * 09:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f2-codfw * 09:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f2-codfw * 09:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e4-codfw * 09:18 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2163: Maintenance * 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e4-codfw * 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-codfw * 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-codfw * 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e5-codfw * 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e5-codfw * 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f4-codfw * 09:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2210: Maintenance * 09:16 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f4-codfw * 09:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2182: Maintenance * 09:14 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2002.codfw.wmnet * 09:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db2163: Maintenance * 09:11 XioNoX: rebooting cr2-drmrs - [[phab:T431749|T431749]] * 09:10 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr2-drmrs,cr2-drmrs IPv6,cr2-drmrs.mgmt with reason: router upgrade * 09:06 XioNoX: draining cr2-drmrs - [[phab:T431749|T431749]] * 09:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2163 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95320 and previous config saved to /var/cache/conftool/dbconfig/20260728-090638-cwilliams.json * 09:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2163.codfw.wmnet with reason: Maintenance * 09:06 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2161: Maintenance * 09:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2002.codfw.wmnet * 09:04 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2001.codfw.wmnet * 09:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2001.codfw.wmnet * 08:57 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2001.codfw.wmnet * 08:48 XioNoX: un-drain cr1-drmrs - [[phab:T431749|T431749]] * 08:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2001.codfw.wmnet * 08:47 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-staging-worker * 08:42 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:35 XioNoX: rebooting cr1-drmrs - [[phab:T431749|T431749]] * 08:33 XioNoX: draining cr1-drmrs - [[phab:T431749|T431749]] * 08:31 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2210: Maintenance * 08:29 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2182: Maintenance * 08:21 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2182: Maintenance * 08:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db2161: Maintenance * 08:16 root@cumin1003: START - Cookbook sre.mysql.pool pool db2182: Maintenance * 08:12 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2210: Maintenance * 08:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2161 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95309 and previous config saved to /var/cache/conftool/dbconfig/20260728-081044-cwilliams.json * 08:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2161.codfw.wmnet with reason: Maintenance * 08:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2154: Maintenance * 08:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2182 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95307 and previous config saved to /var/cache/conftool/dbconfig/20260728-080947-cwilliams.json * 08:09 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2182.codfw.wmnet with reason: Maintenance * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2168: Maintenance * 08:06 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr1-drmrs,cr1-drmrs IPv6,cr1-drmrs.mgmt with reason: router upgrade * 08:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db2210: Maintenance * 08:05 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 08:05 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 08:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2210 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95305 and previous config saved to /var/cache/conftool/dbconfig/20260728-080008-cwilliams.json * 08:00 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2210.codfw.wmnet with reason: Maintenance * 07:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2206: Maintenance * 07:50 gkyziridis@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 07:50 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 07:22 root@cumin1003: START - Cookbook sre.mysql.pool pool db2154: Maintenance * 07:22 root@cumin1003: START - Cookbook sre.mysql.pool pool db2168: Maintenance * 07:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2154 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95295 and previous config saved to /var/cache/conftool/dbconfig/20260728-071640-cwilliams.json * 07:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2154.codfw.wmnet with reason: Maintenance * 07:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95294 and previous config saved to /var/cache/conftool/dbconfig/20260728-071604-cwilliams.json * 07:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2168.codfw.wmnet with reason: Maintenance * 07:08 root@cumin1003: START - Cookbook sre.mysql.pool pool db2206: Maintenance * 07:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2206 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95292 and previous config saved to /var/cache/conftool/dbconfig/20260728-070219-cwilliams.json * 07:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2206.codfw.wmnet with reason: Maintenance * 06:44 marostegui: Failover m5 from db1164 to db1228 - [[phab:T432967|T432967]] * 06:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2235].codfw.wmnet,db[1164,1217,1228].eqiad.wmnet with reason: m5 master switch [[phab:T432967|T432967]] * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.10 (duration: 02m 34s) * 03:39 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] (duration: 36m 06s) * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 02:57 dzahn@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1004.eqiad.wmnet with OS trixie * 02:57 dzahn@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - dzahn@cumin1003" * 02:55 dzahn@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - dzahn@cumin1003" * 02:37 dzahn@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1004.eqiad.wmnet with reason: host reimage * 02:31 dzahn@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1004.eqiad.wmnet with reason: host reimage * 02:16 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie * 02:15 dzahn@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host zuul1004.eqiad.wmnet with OS trixie * 01:43 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie * 01:43 dzahn@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1004.eqiad.wmnet with OS trixie * 01:25 pt1979@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-603-eqsin * 01:24 pt1979@cumin1003: START - Cookbook sre.network.tls for network device asw1-603-eqsin * 01:12 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 01:12 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt for new switches in eqsin - pt1979@cumin2003" * 01:12 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt for new switches in eqsin - pt1979@cumin2003" * 01:08 pt1979@cumin2003: START - Cookbook sre.dns.netbox * 00:48 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 00:47 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 00:47 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 00:47 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 00:26 mutante: attempting reimage with trixie on zuul1004 re-purposed physical hardware - dcops reported install issue - host was in busybox shell ([[phab:T427353|T427353]]) * 00:24 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie == 2026-07-27 == * 23:50 Amir1: mass deleting vp8 transcodes * 23:28 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:27 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1004.eqiad.wmnet with OS bullseye * 23:26 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 23:25 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:25 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 22:39 maryum: Deploy security fix for [[phab:T432877|T432877]] * 22:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1022.eqiad.wmnet with OS bookworm * 22:37 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS bullseye * 22:32 sbassett: Deployed security fix for [[phab:T432789|T432789]] * 22:22 sbassett: Deployed security patch for [[phab:T431819|T431819]] * 22:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1022.eqiad.wmnet with reason: host reimage * 22:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1022.eqiad.wmnet with reason: host reimage * 22:01 RScout-WMF: Deployed security fix for [[phab:T431819|T431819]] * 22:00 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2012 * 21:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2012.codfw.wmnet with OS bookworm * 21:55 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2011\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 21:45 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1022 * 21:45 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1022 * 21:44 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1022 * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1022.eqiad.wmnet 239.48.64.10.in-addr.arpa 9.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:44 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1022.eqiad.wmnet 239.48.64.10.in-addr.arpa 9.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:41 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:41 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 21:34 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS bookworm * 21:31 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:22 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:21 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:19 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:17 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1004 * 21:16 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1004 * 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1004] - vriley@cumin1003" * 21:15 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1004] - vriley@cumin1003" * 21:11 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:10 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2011.codfw.wmnet, repooling source-only afterwards * 21:05 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:01 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1022 * 20:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1022.eqiad.wmnet with OS bookworm * 20:53 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1021.eqiad.wmnet, repooling source-only afterwards * 20:51 mutante: zuul1001 - re-enabled puppet - revert "cherry-picked" gerrit:1314120 - [[phab:T431003|T431003]] * 20:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Maintenance * 20:15 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] (duration: 08m 03s) * 20:11 sbisson@deploy1003: sbisson: Continuing with deployment * 20:09 sbisson@deploy1003: sbisson: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] * 19:47 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 46s) * 19:47 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Maintenance * 19:27 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] (duration: 12m 26s) * 19:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2228 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95285 and previous config saved to /var/cache/conftool/dbconfig/20260727-192711-cwilliams.json * 19:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2228.codfw.wmnet with reason: Maintenance * 19:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2223: Maintenance * 19:23 krinkle@deploy1003: krinkle: Continuing with deployment * 19:16 krinkle@deploy1003: krinkle: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:15 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] * 19:12 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2238: Maintenance * 18:58 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 18:57 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 18:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2227: Maintenance * 18:57 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-experimental: apply * 18:55 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-experimental: apply * 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1021.eqiad.wmnet with OS bookworm * 18:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2011.codfw.wmnet with OS bookworm * 18:40 root@cumin1003: START - Cookbook sre.mysql.pool pool db2223: Maintenance * 18:39 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] (duration: 07m 05s) * 18:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2223 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95275 and previous config saved to /var/cache/conftool/dbconfig/20260727-183500-cwilliams.json * 18:34 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2223.codfw.wmnet with reason: Maintenance * 18:34 musikanimal@deploy1003: musikanimal: Continuing with deployment * 18:34 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2213: Maintenance * 18:33 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:32 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] * 18:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db2238: Maintenance * 18:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2011.codfw.wmnet with reason: host reimage * 18:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2238 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95271 and previous config saved to /var/cache/conftool/dbconfig/20260727-181944-cwilliams.json * 18:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2238.codfw.wmnet with reason: Maintenance * 18:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2226: Maintenance * 18:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1021.eqiad.wmnet with reason: host reimage * 18:14 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2011.codfw.wmnet with reason: host reimage * 18:12 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1021.eqiad.wmnet with reason: host reimage * 18:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db2227: Maintenance * 18:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2227 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95265 and previous config saved to /var/cache/conftool/dbconfig/20260727-180256-cwilliams.json * 18:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2227.codfw.wmnet with reason: Maintenance * 18:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2194: Maintenance * 17:57 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2011 * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2011 * 17:56 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2011 * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2011.codfw.wmnet 37.32.192.10.in-addr.arpa 7.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:56 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2011.codfw.wmnet 37.32.192.10.in-addr.arpa 7.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2011 - bking@cumin2003" * 17:56 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2011 - bking@cumin2003" * 17:52 bking@cumin2003: START - Cookbook sre.dns.netbox * 17:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2011 * 17:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1021 * 17:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1021 * 17:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2011.codfw.wmnet with OS bookworm * 17:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1021.eqiad.wmnet with OS bookworm * 17:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Maintenance * 17:38 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2010\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 17:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2213 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95260 and previous config saved to /var/cache/conftool/dbconfig/20260727-173740-cwilliams.json * 17:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2213.codfw.wmnet with reason: Maintenance * 17:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2211: Maintenance * 17:36 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1020\.eqiad\.wmnet,dc=eqiad,cluster=wdqs\-main,service=wdqs\-main * 17:32 root@cumin1003: START - Cookbook sre.mysql.pool pool db2226: Maintenance * 17:31 taavi@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] (duration: 06m 33s) * 17:27 taavi@deploy1003: taavi: Continuing with deployment * 17:27 taavi@deploy1003: taavi: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:26 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2226 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95256 and previous config saved to /var/cache/conftool/dbconfig/20260727-172636-cwilliams.json * 17:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2226.codfw.wmnet with reason: Maintenance * 17:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2225: Maintenance * 17:25 taavi@deploy1003: Started scap sync-world: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] * 17:13 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 17:11 root@cumin1003: START - Cookbook sre.mysql.pool pool db2194: Maintenance * 17:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2194 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95248 and previous config saved to /var/cache/conftool/dbconfig/20260727-170453-cwilliams.json * 17:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2194.codfw.wmnet with reason: Maintenance * 17:04 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2190: Maintenance * 16:52 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 16:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2211: Maintenance * 16:40 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95242 and previous config saved to /var/cache/conftool/dbconfig/20260727-164015-cwilliams.json * 16:40 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2211.codfw.wmnet with reason: Maintenance * 16:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2178: Maintenance * 16:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db2225: Maintenance * 16:39 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2172: Maintenance * 16:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 16:38 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 16:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2225 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95238 and previous config saved to /var/cache/conftool/dbconfig/20260727-163307-cwilliams.json * 16:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2225.codfw.wmnet with reason: Maintenance * 16:32 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2189: Maintenance * 16:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db2190: Maintenance * 16:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2190 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95230 and previous config saved to /var/cache/conftool/dbconfig/20260727-160602-cwilliams.json * 16:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2190.codfw.wmnet with reason: Maintenance * 15:53 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2177: Maintenance * 15:53 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2172: Maintenance * 15:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2178: Maintenance * 15:51 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2172: Maintenance * 15:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2178 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95224 and previous config saved to /var/cache/conftool/dbconfig/20260727-154559-cwilliams.json * 15:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2172: Maintenance * 15:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2178.codfw.wmnet with reason: Maintenance * 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2171: Maintenance * 15:44 root@cumin1003: START - Cookbook sre.mysql.pool pool db2189: Maintenance * 15:43 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:41 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2172 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95222 and previous config saved to /var/cache/conftool/dbconfig/20260727-153927-cwilliams.json * 15:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2172.codfw.wmnet with reason: Maintenance * 15:38 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2189 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95220 and previous config saved to /var/cache/conftool/dbconfig/20260727-153833-cwilliams.json * 15:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2189.codfw.wmnet with reason: Maintenance * 15:34 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:32 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 15:32 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 15:31 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:29 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:26 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 15:22 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] (duration: 07m 00s) * 15:21 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2155: Maintenance * 15:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2175: Maintenance * 15:18 zabe@deploy1003: zabe: Continuing with deployment * 15:17 zabe@deploy1003: zabe: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:15 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 15:15 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:15 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] * 15:15 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 06s) * 15:15 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:12 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2010.codfw.wmnet with OS bookworm * 15:08 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2177: Maintenance * 15:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2177: Maintenance * 14:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2177: Maintenance * 14:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2171: Maintenance * 14:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2171 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95209 and previous config saved to /var/cache/conftool/dbconfig/20260727-145236-cwilliams.json * 14:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2171.codfw.wmnet with reason: Maintenance * 14:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2177 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95208 and previous config saved to /var/cache/conftool/dbconfig/20260727-145206-cwilliams.json * 14:52 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2157: Maintenance * 14:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2177.codfw.wmnet with reason: Maintenance * 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1020.eqiad.wmnet with OS bookworm * 14:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2156: Maintenance * 14:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2010.codfw.wmnet with reason: host reimage * 14:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2010.codfw.wmnet with reason: host reimage * 14:41 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[1031,2024]*: Upgrade Cassandra to 5.0.8 (canary) - eevans@cumin1003 * 14:34 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2155: Maintenance * 14:33 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2175: Maintenance * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2010 * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2010 * 14:24 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2010 * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2010.codfw.wmnet 94.16.192.10.in-addr.arpa 4.9.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2010.codfw.wmnet 94.16.192.10.in-addr.arpa 4.9.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2010 - bking@cumin2003" * 14:24 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2010 - bking@cumin2003" * 14:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1020.eqiad.wmnet with reason: host reimage * 14:23 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[1031,2024]*: Upgrade Cassandra to 5.0.8 (canary) - eevans@cumin1003 * 14:20 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 14:20 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 14:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1020.eqiad.wmnet with reason: host reimage * 14:17 sukhe: sudo gnt-instance reboot urldownloader1005.wikimedia.org * 14:16 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:15 jelto@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:08 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2155: Maintenance * 14:05 root@cumin1003: START - Cookbook sre.mysql.pool pool db2157: Maintenance * 14:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2175: Maintenance * 14:03 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 11 hosts * 14:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db2155: Maintenance * 14:01 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 11 hosts * 14:01 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1136.eqiad.wmnet * 14:01 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1136.eqiad.wmnet * 14:01 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1136.eqiad.wmnet * 14:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db2156: Maintenance * 13:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2157 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95194 and previous config saved to /var/cache/conftool/dbconfig/20260727-135943-cwilliams.json * 13:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2157.codfw.wmnet with reason: Maintenance * 13:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db2175: Maintenance * 13:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 13:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 13:57 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2010 * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2155 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95193 and previous config saved to /var/cache/conftool/dbconfig/20260727-135613-cwilliams.json * 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2155.codfw.wmnet with reason: Maintenance * 13:55 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1020 * 13:55 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1020 * 13:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2156 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95192 and previous config saved to /var/cache/conftool/dbconfig/20260727-135413-cwilliams.json * 13:54 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2156.codfw.wmnet with reason: Maintenance * 13:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2175 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95191 and previous config saved to /var/cache/conftool/dbconfig/20260727-135300-cwilliams.json * 13:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2010.codfw.wmnet with OS bookworm * 13:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2175.codfw.wmnet with reason: Maintenance * 13:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1020.eqiad.wmnet with OS bookworm * 13:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 34 hosts * 13:46 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 34 hosts * 13:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1201: Maintenance * 13:27 Lucas_WMDE: UTC afternoon backport+config window doen * 13:18 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] (duration: 11m 57s) * 13:14 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, sihe: Continuing with deployment * 13:08 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, sihe: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:07 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool ulsfo [reason: router upgrade finished, [[phab:T431752|T431752]]] * 13:07 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool ulsfo [reason: router upgrade finished, [[phab:T431752|T431752]]] * 13:06 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] * 13:03 XioNoX: repool cr4-ulsfo - [[phab:T431752|T431752]] * 12:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db1201: Maintenance * 12:48 gkyziridis@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1201 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95186 and previous config saved to /var/cache/conftool/dbconfig/20260727-124404-cwilliams.json * 12:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1201.eqiad.wmnet with reason: Maintenance * 12:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1187: Maintenance * 12:30 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader.eqiad.wikimedia.org on all recursors * 12:30 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader.eqiad.wikimedia.org on all recursors * 12:30 sukhe@dns1004: END - running authdns-update * 12:30 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] (duration: 09m 32s) * 12:28 sukhe@dns1004: START - running authdns-update * 12:25 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 12:22 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:20 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] * 12:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts * 12:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts * 12:13 XioNoX: rebooting cr4-ulsfo for upgrade - [[phab:T431752|T431752]] * 12:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: es1038 repool * 12:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 38 hosts * 12:08 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 38 hosts * 11:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1187: Maintenance * 11:53 urbanecm@deploy1003: mwscript-k8s job started: foreachwikiindblist growthexperiments GrowthExperiments:cleanMentorList # [[phab:T431804|T431804]] * 11:50 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr4-ulsfo,cr4-ulsfo IPv6,cr4-ulsfo.mgmt with reason: router upgrade * 11:50 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] (duration: 11m 07s) * 11:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1187 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95178 and previous config saved to /var/cache/conftool/dbconfig/20260727-114844-cwilliams.json * 11:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1187.eqiad.wmnet with reason: Maintenance * 11:43 urbanecm@deploy1003: urbanecm: Continuing with deployment * 11:42 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:39 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] * 11:37 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 11:36 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 11:36 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 11:35 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 11:29 XioNoX: start draining cr4-ulsfo - [[phab:T431752|T431752]] * 11:29 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 11:29 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 11:28 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1035: testing * 11:28 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1035: testing * 11:27 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1035: testing * 11:27 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1035: testing * 11:26 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool ulsfo [reason: router upgrade, [[phab:T431752|T431752]]] * 11:26 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1038: es1038 repool * 11:26 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool ulsfo [reason: router upgrade, [[phab:T431752|T431752]]] * 11:26 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1038: testing * 11:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1264: Maintenance * 11:24 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1038: testing * 11:23 marostegui@cumin1003: dbctl commit (dc=all): 'Repool es1050 as master', diff saved to https://phabricator.wikimedia.org/P95170 and previous config saved to /var/cache/conftool/dbconfig/20260727-112326-marostegui.json * 11:23 marostegui@cumin1003: dbctl commit (dc=all): 'Repool es1050', diff saved to https://phabricator.wikimedia.org/P95169 and previous config saved to /var/cache/conftool/dbconfig/20260727-112302-marostegui.json * 11:22 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1050: testing * 11:22 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1050: testing * 11:20 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 11:18 blake@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 11:18 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 11:12 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 11:11 blake@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 11:09 blake@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 11:09 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 11:09 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 11:08 blake@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 11:05 blake@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 11:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 11:02 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 10:50 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 10:43 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:39 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply * 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1264: Maintenance * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply * 10:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply * 10:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 10:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 10:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 10:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1264 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95164 and previous config saved to /var/cache/conftool/dbconfig/20260727-103204-cwilliams.json * 10:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1264.eqiad.wmnet with reason: Maintenance * 10:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 10:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 10:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 10:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 10:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 10:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 10:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 10:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1237: Maintenance * 10:04 elukey: restart burrow main-eqiad on kafkamon2003 to clear some errors on kafka-main1008 * 09:58 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1136.eqiad.wmnet with OS trixie * 09:39 elukey: restart burrow-main-eqiad.service on kafkamon1003 to see if a recurrent kafka error on kafka-main1008 goes away * 09:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1237: Maintenance * 09:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1136.eqiad.wmnet with reason: host reimage * 09:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1237 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95159 and previous config saved to /var/cache/conftool/dbconfig/20260727-093328-cwilliams.json * 09:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1237.eqiad.wmnet with reason: Maintenance * 09:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1136.eqiad.wmnet with reason: host reimage * 09:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1203: Maintenance * 09:17 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1136 * 09:17 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1136 * 09:04 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1136 * 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1136.eqiad.wmnet 191.32.64.10.in-addr.arpa 1.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:04 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1136.eqiad.wmnet 191.32.64.10.in-addr.arpa 1.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1136 - jiji@cumin1003" * 09:04 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1136 - jiji@cumin1003" * 08:52 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 08:52 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 08:52 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 08:51 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 08:50 jiji@cumin1003: START - Cookbook sre.dns.netbox * 08:47 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1136 * 08:46 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1136.eqiad.wmnet with OS trixie * 08:44 marostegui: Rename tables on s3 [[phab:T425066|T425066]] * 08:43 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1136.eqiad.wmnet * 08:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db1203: Maintenance * 08:43 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1136.eqiad.wmnet * 08:43 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1136.eqiad.wmnet * 08:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1203 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95154 and previous config saved to /var/cache/conftool/dbconfig/20260727-083703-cwilliams.json * 08:36 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1203.eqiad.wmnet with reason: Maintenance * 08:16 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1179: Maintenance * 07:44 phuedx: UTC morning backport window done * 07:37 phuedx@deploy1003: Finished scap sync-world: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] (duration: 32m 33s) * 07:28 root@cumin1003: START - Cookbook sre.mysql.pool pool db1179: Maintenance * 07:26 marostegui: Rename tables on s3 [[phab:T426341|T426341]] * 07:25 phuedx@deploy1003: phuedx: Continuing with deployment * 07:22 marostegui: Drop tables in akwiki nawiki pihwiki - growthexperiments_* [[phab:T428885|T428885]] * 07:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95149 and previous config saved to /var/cache/conftool/dbconfig/20260727-072234-cwilliams.json * 07:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1179.eqiad.wmnet with reason: Maintenance * 07:20 phuedx@deploy1003: phuedx: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:16 ryankemper: [[phab:T430880|T430880]] [WDQS] Reimaged `wdqs1018` and `wdqs1019` to Bookworm, restored data using test-cookbook change {{Gerrit|1317128}}, and repooled both; 25/36 hosts complete * 07:04 phuedx@deploy1003: Started scap sync-world: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] * 06:57 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1019.eqiad.wmnet * 06:56 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1018.eqiad.wmnet * 06:40 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1020.eqiad.wmnet with reason: Cloning * 06:35 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db1228.eqiad.wmnet with reason: Rebooting * 06:29 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:29 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:25 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:25 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:25 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1019.eqiad.wmnet, repooling source-only afterwards * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1018.eqiad.wmnet, repooling source-only afterwards * 04:51 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1019.eqiad.wmnet, repooling source-only afterwards * 04:51 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1018.eqiad.wmnet, repooling source-only afterwards * 04:48 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s) * 04:48 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 04:48 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s) * 04:48 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 36s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-26 == * 14:59 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:59 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:59 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:59 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1019.eqiad.wmnet with OS bookworm * 01:05 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1018.eqiad.wmnet with OS bookworm * 00:43 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1019.eqiad.wmnet with reason: host reimage * 00:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1018.eqiad.wmnet with reason: host reimage * 00:34 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1019.eqiad.wmnet with reason: host reimage * 00:33 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1018.eqiad.wmnet with reason: host reimage * 00:16 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 00:16 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 00:15 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 00:15 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1019 * 00:11 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1019 * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1018 * 00:11 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1018 * 00:08 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1019.eqiad.wmnet with OS bookworm * 00:08 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1018.eqiad.wmnet with OS bookworm == 2026-07-25 == * 22:06 ryankemper: [[phab:T430880|T430880]] [WDQS] Repooled `wdqs1017` and `wdqs2024` after reimaging to bookworm, scap deploying, and data xfering * 22:04 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2024.codfw.wmnet * 22:03 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1017.eqiad.wmnet * 21:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1017.eqiad.wmnet, repooling source-only afterwards * 21:06 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2024.codfw.wmnet, repooling source-only afterwards * 20:52 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:52 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:52 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:52 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 20:18 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1017.eqiad.wmnet, repooling source-only afterwards * 20:18 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2024.codfw.wmnet, repooling source-only afterwards * 20:15 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:15 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:15 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:15 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 19:57 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s) * 19:57 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 19:57 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 07s) * 19:57 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 19:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2024.codfw.wmnet with OS bookworm * 19:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1017.eqiad.wmnet with OS bookworm * 19:02 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2024.codfw.wmnet with reason: host reimage * 18:58 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1017.eqiad.wmnet with reason: host reimage * 18:53 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2024.codfw.wmnet with reason: host reimage * 18:52 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1017.eqiad.wmnet with reason: host reimage * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2024 * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2024 * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1017 * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1017 * 18:27 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2024 * 18:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2024.codfw.wmnet 58.16.192.10.in-addr.arpa 8.5.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:26 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2024.codfw.wmnet 58.16.192.10.in-addr.arpa 8.5.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:24 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1017 * 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1017.eqiad.wmnet 238.48.64.10.in-addr.arpa 8.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:24 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1017.eqiad.wmnet 238.48.64.10.in-addr.arpa 8.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1017 - ryankemper@cumin2003" * 18:24 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1017 - ryankemper@cumin2003" * 18:23 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 18:18 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 18:17 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1017 * 18:17 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2024 * 18:14 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1017.eqiad.wmnet with OS bookworm * 18:14 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2024.codfw.wmnet with OS bookworm * 18:05 ryankemper: [WDQS] [[phab:T430880|T430880]] Reimaged `wdqs1016` and `wdqs2023` to Bookworm with `--move-vlan`, restored main and scholarly data, validated postflights, and repooled both hosts. Confirmed PyBal rebuilt both backends with their new addresses * 17:45 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2023.codfw.wmnet * 17:43 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1016.eqiad.wmnet * 06:35 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1016.eqiad.wmnet, repooling source-only afterwards * 06:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2023.codfw.wmnet, repooling source-only afterwards * 05:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2023.codfw.wmnet, repooling source-only afterwards * 05:19 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1016.eqiad.wmnet, repooling source-only afterwards * 05:07 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 07s) * 05:07 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 05:06 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 06s) * 05:06 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 03:27 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2023.codfw.wmnet with OS bookworm * 02:59 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2023.codfw.wmnet with reason: host reimage * 02:56 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2023.codfw.wmnet with reason: host reimage * 02:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2023 * 02:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2023 * 02:30 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2023 * 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2023.codfw.wmnet 35.0.192.10.in-addr.arpa 5.3.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 02:30 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2023.codfw.wmnet 35.0.192.10.in-addr.arpa 5.3.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2023 - ryankemper@cumin2003" * 02:30 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2023 - ryankemper@cumin2003" * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 26s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:15 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1016.eqiad.wmnet with OS bookworm * 00:49 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1016.eqiad.wmnet with reason: host reimage * 00:43 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1016.eqiad.wmnet with reason: host reimage * 00:31 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 00:27 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1016 * 00:27 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1016 * 00:27 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2023 * 00:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1016.eqiad.wmnet with OS bookworm * 00:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2023.codfw.wmnet with OS bookworm * 00:11 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs1014.eqiad.wmnet and wdqs2008.codfw.wmnet after Bookworm reimage, transfer, and postflight; wdqs2008 is serving, while wdqs1014 will remain outside of service until a pybal restart next monday == 2026-07-24 == * 23:54 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1014.eqiad.wmnet * 23:54 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2008.codfw.wmnet * 23:43 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2010.codfw.wmnet with OS trixie * 23:08 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 23:03 jhathaway@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 22:33 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 22:13 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 22:13 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 22:13 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 22:13 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:00 jhathaway@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 21:53 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 21:53 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie * 21:51 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 21:47 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie * 21:43 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 21:39 jhathaway@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 21:38 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 17:21 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1135.eqiad.wmnet * 17:21 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1135.eqiad.wmnet * 17:21 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1135.eqiad.wmnet * 16:34 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 16:34 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 16:34 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 16:34 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 16:33 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 16:33 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 16:28 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:28 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:28 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:28 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2008.codfw.wmnet, repooling source-only afterwards * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1014.eqiad.wmnet, repooling source-only afterwards * 15:56 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1135.eqiad.wmnet with OS trixie * 15:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 40 hosts * 15:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 40 hosts * 15:37 topranks: upgrade SR-Linux OS on lswtest-d8-eqiad * 15:36 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1135.eqiad.wmnet with reason: host reimage * 15:33 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 6 hosts with reason: upgrade lswtest-d8-eqiad * 15:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1135.eqiad.wmnet with reason: host reimage * 15:30 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc-gp2006.codfw.wmnet with OS bookworm * 15:15 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1135 * 15:15 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1135 * 15:13 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc-gp2006.codfw.wmnet with reason: host reimage * 15:08 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc-gp2006.codfw.wmnet with reason: host reimage * 14:49 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm * 14:48 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host mc-gp2006.codfw.wmnet with OS bookworm * 14:34 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] (duration: 41m 12s) * 14:32 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1135 * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1135.eqiad.wmnet 177.32.64.10.in-addr.arpa 7.7.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:32 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1135.eqiad.wmnet 177.32.64.10.in-addr.arpa 7.7.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1135 - jiji@cumin1003" * 14:32 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1135 - jiji@cumin1003" * 14:29 krinkle@deploy1003: krinkle: Continuing with deployment * 14:29 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm * 14:27 jiji@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host mc-gp2006.codfw.wmnet with OS bookworm * 14:26 jiji@cumin1003: START - Cookbook sre.dns.netbox * 14:15 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1135 * 14:14 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1135.eqiad.wmnet with OS trixie * 14:14 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1135.eqiad.wmnet * 14:13 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1135.eqiad.wmnet * 14:13 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1135.eqiad.wmnet * 14:10 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1072.eqiad.wmnet * 14:10 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1072.eqiad.wmnet * 14:10 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1072.eqiad.wmnet * 14:10 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1072.eqiad.wmnet * 14:09 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1071.eqiad.wmnet * 14:09 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1071.eqiad.wmnet * 14:09 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1071.eqiad.wmnet * 14:09 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1071.eqiad.wmnet * 13:58 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 13:58 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:58 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:57 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:55 krinkle@deploy1003: krinkle: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:53 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] * 13:45 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:45 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push new IPs for mc-gp2006 - cmooney@cumin1003" * 13:45 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push new IPs for mc-gp2006 - cmooney@cumin1003" * 13:44 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) mc-gp2006.codfw.wmnet on all recursors * 13:44 cmooney@cumin1003: START - Cookbook sre.dns.wipe-cache mc-gp2006.codfw.wmnet on all recursors * 13:42 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm * 13:41 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:30 papaul: reboot mr1-eqsin for maintenance * 13:24 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb[1029-1031].eqiad.wmnet * 13:10 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb[1029-1031].eqiad.wmnet * 11:33 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 7 hosts * 11:11 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 7 hosts * 10:56 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 7 hosts * 10:47 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 7 hosts * 10:44 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:44 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:41 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:41 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:35 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 8 hosts * 10:34 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:33 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:32 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:32 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:31 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:31 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:30 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 8 hosts * 10:24 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie * 10:19 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:18 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 16 hosts * 10:17 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2001.codfw.wmnet * 10:13 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2001.codfw.wmnet * 10:12 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2001.codfw.wmnet * 10:02 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2001.codfw.wmnet * 10:02 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2002.codfw.wmnet * 09:57 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2002.codfw.wmnet * 09:56 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2002.codfw.wmnet * 09:51 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2002.codfw.wmnet * 09:51 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1002.eqiad.wmnet * 09:47 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1002.eqiad.wmnet * 09:47 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1001.eqiad.wmnet * 09:44 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1001.eqiad.wmnet * 09:34 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2003.codfw.wmnet * 09:32 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2003.codfw.wmnet * 09:32 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2002.codfw.wmnet * 09:30 brouberol@dns1004: END - running authdns-update * 09:29 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2002.codfw.wmnet * 09:29 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2001.codfw.wmnet * 09:27 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 9 hosts * 09:27 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2001.codfw.wmnet * 09:27 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2001.codfw.wmnet * 09:26 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 9 hosts * 09:26 brouberol@dns1004: START - running authdns-update * 09:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 57 hosts * 09:24 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2001.codfw.wmnet * 09:24 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2002.codfw.wmnet * 09:22 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2002.codfw.wmnet * 09:21 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 57 hosts * 09:20 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2003.codfw.wmnet * 09:19 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 16 hosts * 09:16 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2003.codfw.wmnet * 09:16 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1003.eqiad.wmnet * 09:15 urbanecm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 09:15 urbanecm@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 09:13 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1003.eqiad.wmnet * 09:13 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1002.eqiad.wmnet * 09:11 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1002.eqiad.wmnet * 09:11 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1001.eqiad.wmnet * 09:07 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1001.eqiad.wmnet * 08:32 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:24 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:16 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 08:16 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 08:07 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:07 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:07 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 08:02 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 08:01 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:59 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:57 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 07:57 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 06:46 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1025.eqiad.wmnet with reason: Cloning * 06:46 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s4 * 06:45 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s6 * 06:44 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1019.eqiad.wmnet,service=s6 * 06:44 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1019.eqiad.wmnet,service=s4 * 03:40 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:40 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:40 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:40 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 03:37 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:37 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:37 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:36 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:49 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on mr1-eqsin,mr1-eqsin IPv6 with reason: connection issue * 02:38 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on cr[2-3]-eqsin.mgmt,ps1-[603-604]-eqsin with reason: connection issue * 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 27s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-23 == * 23:27 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin.oob,mr1-eqsin.oob IPv6 with reason: switch refresh * 22:21 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Setting storage compatibility to NONE - eevans@cumin1003 * 22:01 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Setting storage compatibility to NONE - eevans@cumin1003 * 21:29 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1014.eqiad.wmnet, repooling source-only afterwards * 21:28 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 46s) * 21:28 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 21:19 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Setting storage compatibility to UPGRADING - eevans@cumin1003 * 21:00 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Setting storage compatibility to UPGRADING - eevans@cumin1003 * 20:17 dani@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] (duration: 11m 57s) * 20:13 dani@deploy1003: dani: Continuing with deployment * 20:07 dani@deploy1003: dani: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:05 dani@deploy1003: Started scap sync-world: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] * 19:24 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:24 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating the rest of the ipv6 dns records. - jhancock@cumin2002" * 19:24 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating the rest of the ipv6 dns records. - jhancock@cumin2002" * 19:14 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 19:05 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wdqs1014.eqiad.wmnet with OS bookworm * 19:04 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.noop (exit_code=99) * 19:04 cwilliams@cumin1003: START - Cookbook sre.mysql.noop * 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1014.eqiad.wmnet with reason: host reimage * 18:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1014.eqiad.wmnet with reason: host reimage * 18:30 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2008.codfw.wmnet, repooling source-only afterwards * 18:28 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 19s) * 18:28 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1014 * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1014 * 18:22 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1014 * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1014.eqiad.wmnet 188.32.64.10.in-addr.arpa 8.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:22 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1014.eqiad.wmnet 188.32.64.10.in-addr.arpa 8.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1014 - bking@cumin2003" * 18:21 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1014 - bking@cumin2003" * 18:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2215: Maintenance * 18:18 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 18:15 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:15 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 18:06 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:06 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 18:05 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 18:04 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2052: codfw rack B8 re-pool after maintenance * 17:54 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 17:54 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:54 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 17:32 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2215: Maintenance * 17:29 cmooney@dns3003: END - running authdns-update * 17:27 cmooney@dns3003: START - running authdns-update * 17:23 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 17:22 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:18 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool es2052: codfw rack B8 re-pool after maintenance * 17:18 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2189: codfw rack B8 re-pool after maintenance * 17:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2215.codfw.wmnet with reason: Maintenance * 17:17 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 17:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2215 [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95126 and previous config saved to /var/cache/conftool/dbconfig/20260723-170903-cwilliams.json * 17:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2191 to x1 primary [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95125 and previous config saved to /var/cache/conftool/dbconfig/20260723-170612-cwilliams.json * 17:05 cezmunsta: Starting x1 codfw failover from db2215 to db2191 - [[phab:T432986|T432986]] * 16:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2191 with weight 0 [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95123 and previous config saved to /var/cache/conftool/dbconfig/20260723-165831-cwilliams.json * 16:58 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 16 hosts with reason: Primary switchover x1 [[phab:T432986|T432986]] * 16:36 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 138128 * 16:35 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 138128 * 16:33 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2189: codfw rack B8 re-pool after maintenance * 16:33 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2164: codfw rack B8 re-pool after maintenance * 16:28 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1072.eqiad.wmnet * 16:27 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1072.eqiad.wmnet with OS trixie * 16:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2249: Maintenance * 16:06 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker1072.eqiad.wmnet with reason: host reimage * 16:06 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1072.eqiad.wmnet with reason: host reimage * 15:50 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1072 * 15:50 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1072 * 15:49 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1072.eqiad.wmnet with OS trixie * 15:48 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2164: codfw rack B8 re-pool after maintenance * 15:48 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] (duration: 06m 37s) * 15:48 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2163: codfw rack B8 re-pool after maintenance * 15:45 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 15:43 musikanimal@deploy1003: musikanimal: Continuing with deployment * 15:43 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:41 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] * 15:36 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1072.eqiad.wmnet * 15:35 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1072.eqiad.wmnet * 15:35 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1072.eqiad.wmnet * 15:34 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:34 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push any outstanding updates - cmooney@cumin1003" * 15:34 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push any outstanding updates - cmooney@cumin1003" * 15:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db2249: Maintenance * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 15:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:26 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:21 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:21 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 15:21 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:21 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 15:20 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 15:19 cmooney@dns2004: END - running authdns-update * 15:17 cmooney@dns2004: START - running authdns-update * 15:14 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns2004.wikimedia.org * 15:12 brouberol@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 15:12 brouberol@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 15:12 klausman@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ml-serve1001.eqiad.wmnet with OS trixie * 15:11 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1071.eqiad.wmnet * 15:11 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1071.eqiad.wmnet with OS trixie * 15:10 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wdqs2008.codfw.wmnet with OS bookworm * 15:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2249.codfw.wmnet with reason: Maintenance * 15:08 brouberol@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 15:08 brouberol@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 15:08 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 15:08 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 15:06 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns1004.wikimedia.org * 15:02 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2002.codfw.wmnet * 15:02 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2002.codfw.wmnet * 15:02 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2163: codfw rack B8 re-pool after maintenance * 15:01 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 15:01 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 14:59 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test2001.codfw.wmnet * 14:57 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test2001.codfw.wmnet * 14:56 ryankemper: [WDQS] [[phab:T430880|T430880]] Reimaged `wdqs2016` to Bookworm, xferred scholarly_articles from `wdqs2024`, validated updater/readiness/federation, and repooled * 14:51 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2016.codfw.wmnet * 14:51 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1001.eqiad.wmnet with reason: host reimage * 14:48 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1071.eqiad.wmnet with reason: host reimage * 14:47 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1001.eqiad.wmnet with reason: host reimage * 14:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2008.codfw.wmnet with reason: host reimage * 14:43 topranks: reboot lsw1-b8-codw to upgrade JunOS [[phab:T430929|T430929]] * 14:41 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1071.eqiad.wmnet with reason: host reimage * 14:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2008.codfw.wmnet with reason: host reimage * 14:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2231: Maintenance * 14:30 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1001.eqiad.wmnet with OS trixie * 14:25 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2002.codfw.wmnet * 14:23 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1071 * 14:23 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1071 * 14:23 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 14:22 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore scholarly data after Bookworm reimage) xfer scholarly_articles from wdqs2024.codfw.wmnet -> wdqs2016.codfw.wmnet, repooling source-only afterwards * 14:22 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2052: codfw rack B8 depool for maintenance * 14:21 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool es2052: codfw rack B8 depool for maintenance * 14:21 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2249: codfw rack B8 depool for maintenance * 14:21 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1071 * 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1071.eqiad.wmnet 166.48.64.10.in-addr.arpa 6.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:21 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1071.eqiad.wmnet 166.48.64.10.in-addr.arpa 6.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1071 - jiji@cumin1003" * 14:21 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1071 - jiji@cumin1003" * 14:21 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2249: codfw rack B8 depool for maintenance * 14:21 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2189: codfw rack B8 depool for maintenance * 14:20 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2002.codfw.wmnet * 14:20 cmooney@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2050.codfw.wmnet * 14:20 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2189: codfw rack B8 depool for maintenance * 14:20 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2164: codfw rack B8 depool for maintenance * 14:20 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2164: codfw rack B8 depool for maintenance * 14:19 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2163: codfw rack B8 depool for maintenance * 14:19 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:19 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1014 * 14:19 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2163: codfw rack B8 depool for maintenance * 14:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1014.eqiad.wmnet with OS bookworm * 14:17 cmooney@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2050.codfw.wmnet * 14:16 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 14:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2008 * 14:14 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2008 * 14:14 cmooney@cumin1003: conftool action : set/pooled=no; selector: name=dns2004.wikimedia.org * 14:14 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2008.codfw.wmnet with OS bookworm * 14:13 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 14:12 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:10 topranks: depool dns2004 before lsw1-b8-codfw switch maintenance [[phab:T430929|T430929]] * 14:10 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b8-codfw,lsw1-b8-codfw IPv6,lsw1-b8-codfw.mgmt,ssw1-a[1,8]-codfw with reason: lsw1-b8-codfw JunOS upgrade * 14:07 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 30 hosts with reason: lsw1-b8-codfw JunOS upgrade * 14:06 elukey: upload python3-docker-report 0.0.19 to apt.wikimedia.org for bookworm and trixie * 13:59 jiji@cumin1003: START - Cookbook sre.dns.netbox * 13:58 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1071 * 13:57 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:57 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:57 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:57 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1071.eqiad.wmnet with OS trixie * 13:55 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1071.eqiad.wmnet * 13:55 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1071.eqiad.wmnet * 13:55 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1071.eqiad.wmnet * 13:53 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:52 logmsgbot: kharlan Deployed security patch for [[phab:T432948|T432948]] * 13:51 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:51 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 13:51 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db2231: Maintenance * 13:50 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:50 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:50 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:50 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:49 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:49 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:49 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2231 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95097 and previous config saved to /var/cache/conftool/dbconfig/20260723-134436-cwilliams.json * 13:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2231.codfw.wmnet with reason: Maintenance * 13:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:39 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:38 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 13:38 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] (duration: 09m 07s) * 13:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1037 hosts * 13:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2196: Maintenance * 13:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:33 kharlan@deploy1003: kharlan, emc-wmf: Continuing with deployment * 13:31 kharlan@deploy1003: kharlan, emc-wmf: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:30 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:28 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] * 13:17 hashar@deploy1003: Finished deploy [integration/docroot@2199146]: build: License GPL2.0+ / updating npm dependencies (duration: 00m 14s) * 13:17 hashar@deploy1003: Started deploy [integration/docroot@2199146]: build: License GPL2.0+ / updating npm dependencies * 13:14 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service * 13:07 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 12:58 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2207: Repooling * 12:49 root@cumin1003: START - Cookbook sre.mysql.pool pool db2196: Maintenance * 12:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2196 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95087 and previous config saved to /var/cache/conftool/dbconfig/20260723-123952-cwilliams.json * 12:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2196.codfw.wmnet with reason: Maintenance * 12:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2191: Maintenance * 12:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:13 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:13 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: Repooling * 12:12 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2207: Repooling * 12:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: Repooling * 11:56 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2235.codfw.wmnet with OS trixie * 11:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db2191: Maintenance * 11:46 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1070.eqiad.wmnet * 11:46 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1070.eqiad.wmnet * 11:46 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1070.eqiad.wmnet * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2191 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95080 and previous config saved to /var/cache/conftool/dbconfig/20260723-114308-cwilliams.json * 11:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2191.codfw.wmnet with reason: Maintenance * 11:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2186: Maintenance * 11:35 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 46375 * 11:34 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 46375 * 11:33 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2235.codfw.wmnet with reason: host reimage * 11:28 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2235.codfw.wmnet with reason: host reimage * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c7-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c7-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c6-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c6-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c5-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c5-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c4-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c4-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c3-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c3-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c2-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c2-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d7-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d7-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d4-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d4-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d3-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d2-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d2-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d8-eqiad * 11:23 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d8-eqiad * 11:23 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d1-eqiad * 11:23 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d1-eqiad * 11:12 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db2235.codfw.wmnet with OS trixie * 11:11 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:11 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[2160,2235].codfw.wmnet with reason: Upgrading * 10:56 root@cumin1003: START - Cookbook sre.mysql.pool pool db2186: Maintenance * 10:54 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1037: testing * 10:53 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1037: testing * 10:53 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1037: testing * 10:53 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1037: testing * 10:52 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: testing * 10:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2186 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95072 and previous config saved to /var/cache/conftool/dbconfig/20260723-104956-cwilliams.json * 10:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2186.codfw.wmnet with reason: Maintenance * 10:43 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1054: testing * 10:41 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1070.eqiad.wmnet with OS trixie * 10:30 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1037 hosts * 10:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 10:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 10:20 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1070.eqiad.wmnet with reason: host reimage * 10:16 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1070.eqiad.wmnet with reason: host reimage * 10:06 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1038: testing * 10:05 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1038: testing * 10:05 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1038: testing * 10:02 marostegui@dns1004: END - running authdns-update * 10:00 marostegui@dns1004: START - running authdns-update * 09:58 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: testing * 09:57 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1054: testing * 09:57 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1070 * 09:57 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1070 * 09:57 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1054: testing * 09:57 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1054: testing * 09:56 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1055: testing * 09:56 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1070 * 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1070.eqiad.wmnet 165.48.64.10.in-addr.arpa 5.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:56 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1070.eqiad.wmnet 165.48.64.10.in-addr.arpa 5.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1070 - jiji@cumin1003" * 09:56 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1070 - jiji@cumin1003" * 09:47 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for 1035 hosts * 09:45 jiji@cumin1003: START - Cookbook sre.dns.netbox * 09:42 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1070 * 09:42 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1070.eqiad.wmnet with OS trixie * 09:42 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1070.eqiad.wmnet * 09:41 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1070.eqiad.wmnet * 09:41 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1070.eqiad.wmnet * 09:27 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es2051: testing * 09:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: testing * 09:12 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es2051: testing * 09:11 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1055: testing * 09:09 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:09 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1055: testing * 09:09 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1055: testing * 08:50 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:50 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:50 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 08:49 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 08:49 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 08:49 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:46 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1069.eqiad.wmnet * 08:46 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1069.eqiad.wmnet * 08:46 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1069.eqiad.wmnet * 08:39 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 08:38 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 08:38 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 08:37 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 08:35 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:10 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1069.eqiad.wmnet with OS trixie * 07:49 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1069.eqiad.wmnet with reason: host reimage * 07:45 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1069.eqiad.wmnet with reason: host reimage * 07:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1035 hosts * 07:33 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 07:32 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 07:29 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1069 * 07:29 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1069 * 07:26 jiji@deploy1003: Finished scap sync-world: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules (duration: 06m 01s) * 07:25 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1069 * 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1069.eqiad.wmnet 164.48.64.10.in-addr.arpa 4.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:25 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1069.eqiad.wmnet 164.48.64.10.in-addr.arpa 4.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1069 - jiji@cumin1003" * 07:25 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1069 - jiji@cumin1003" * 07:25 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 07:25 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 07:24 jiji@deploy1003: jiji: Continuing with deployment * 07:22 jiji@deploy1003: jiji: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:21 jiji@deploy1003: Started scap sync-world: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules * 07:21 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1031.eqiad.wmnet,service=s7 * 07:20 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1031.eqiad.wmnet,service=s2 * 07:20 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1031.eqiad.wmnet,service=s7 * 07:20 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1031.eqiad.wmnet,service=s2 * 07:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts * 07:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts * 07:19 jiji@cumin1003: START - Cookbook sre.dns.netbox * 07:19 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1069 * 07:19 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1069.eqiad.wmnet with OS trixie * 07:19 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1069.eqiad.wmnet * 07:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 45 hosts * 07:17 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1069.eqiad.wmnet * 07:17 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1069.eqiad.wmnet * 07:14 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 45 hosts * 07:13 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 06:16 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs2007 after successful Bookworm reimage, data transfer, and postflight validation; wdqs1013 also passed postflights and is enabled in conftool, but remains out of IPVS pending a rolling pybal restart to clear its stale pre-VLAN-move address * 05:58 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2007.codfw.wmnet * 05:58 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1013.eqiad.wmnet * 05:54 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore scholarly data after Bookworm reimage) xfer scholarly_articles from wdqs2024.codfw.wmnet -> wdqs2016.codfw.wmnet, repooling source-only afterwards == 2026-07-22 == * 23:34 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Apply upgrade to JVM17 - eevans@cumin1003 * 23:14 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Apply upgrade to JVM17 - eevans@cumin1003 * 22:06 ryankemper: [WDQS] Added requestctl per-IP ratelimit `wdqs_heavy_sparql_bots_jul_2026_ratelimit` (chronic heavy-query bot tier driving deadlock-remediation restarts); pruned superseded `wdqs_2026_05_11_worobot` * 21:51 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] (duration: 11m 52s) * 21:44 sbassett@deploy1003: sbassett: Continuing with deployment * 21:43 sbassett@deploy1003: sbassett: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:39 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] * 20:38 dani@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] (duration: 32m 51s) * 20:38 ryankemper: [WDQS] Pruned obsolete requestctl action+pattern `wdqs_20260715_p2003_ring_ja3n` (actor rotated JA3Ns; rule inert) * 20:26 dani@deploy1003: dani, vadymts1: Continuing with deployment * 20:24 dani@deploy1003: dani, vadymts1: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:14 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2244: Testing * 20:06 dani@deploy1003: Started scap sync-world: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] * 19:56 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1013.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2016.codfw.wmnet with OS bookworm * 19:40 mutante: gerrit - one more service restart is needed - restarting * 19:29 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2244: Testing * 19:27 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2244: Testing * 19:27 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2244: Testing * 19:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2016.codfw.wmnet with reason: host reimage * 19:15 dancy@deploy1003: Finished deploy [zuul/deploy@d92e238]: Freshening Zuul installation (duration: 00m 15s) * 19:14 dancy@deploy1003: Started deploy [zuul/deploy@d92e238]: Freshening Zuul installation * 19:11 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2016.codfw.wmnet with reason: host reimage * 18:54 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1013.eqiad.wmnet, repooling source-only afterwards * 18:52 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2016 * 18:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2016 * 18:51 dancy@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 18:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2016.codfw.wmnet with OS bookworm * 18:39 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 11s) * 18:39 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 18:36 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 18:30 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417] (thin): Regular analytics weekly train THIN [analytics/refinery@2a25417d] (duration: 02m 09s) * 18:28 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417] (thin): Regular analytics weekly train THIN [analytics/refinery@2a25417d] * 18:28 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417]: Regular analytics weekly train [analytics/refinery@2a25417d] (duration: 04m 31s) * 18:27 dduvall: deploying https://gerrit.wikimedia.org/r/c/integration/config/+/1314025 (4 jobs updated) * 18:23 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417]: Regular analytics weekly train [analytics/refinery@2a25417d] * 18:22 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@2a25417d] (duration: 01m 59s) * 18:20 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@2a25417d] * 17:56 Raine: deployment server switchover => deploy1003 is primary now * 17:55 kamila@deploy1003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 28m 24s) * 17:54 mutante: restarting gerrit for maintenance * 17:29 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1013.eqiad.wmnet with OS bookworm * 17:27 kamila@deploy1003: Started scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] * 17:20 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] (duration: 22m 50s) * 17:12 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1023.eqiad.wmnet -> wdqs1024.eqiad.wmnet, repooling source-only afterwards * 17:04 Raine: point deployment.eqiad.wmnet to deploy1003 * 17:04 kamila@dns7001: END - running authdns-update * 17:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1013.eqiad.wmnet with reason: host reimage * 17:02 kamila@dns7001: START - running authdns-update * 17:01 kamila@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on releases2003.codfw.wmnet,releases1003.eqiad.wmnet with reason: Deployment server switchover * 17:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1013.eqiad.wmnet with reason: host reimage * 16:58 kamila@deploy2003: Locking from deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] * 16:57 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] (duration: 02m 33s) * 16:55 kamila@deploy2003: Locking from deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] * 16:55 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2003 - [[phab:T240266|T240266]] (duration: 00m 11s) * 16:54 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2003 - [[phab:T240266|T240266]] * 16:40 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1013 * 16:40 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1013 * 16:39 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1013 * 16:39 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1013.eqiad.wmnet 105.32.64.10.in-addr.arpa 5.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:39 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1013.eqiad.wmnet 105.32.64.10.in-addr.arpa 5.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:39 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:39 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1013 - bking@cumin2003" * 16:39 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1013 - bking@cumin2003" * 16:34 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:34 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1013 * 16:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1013.eqiad.wmnet with OS bookworm * 16:28 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1023.eqiad.wmnet -> wdqs1024.eqiad.wmnet, repooling source-only afterwards * 16:27 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-scholarly,name=eqiad * 16:27 eevans@deploy2003: helmfile [eqiad] DONE helmfile.d/services/linked-artifacts: apply * 16:26 eevans@deploy2003: helmfile [eqiad] START helmfile.d/services/linked-artifacts: apply * 16:26 eevans@deploy2003: helmfile [codfw] DONE helmfile.d/services/linked-artifacts: apply * 16:26 eevans@deploy2003: helmfile [codfw] START helmfile.d/services/linked-artifacts: apply * 16:25 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 29s) * 16:25 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 16:24 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 16:21 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 16:18 eevans@deploy2003: helmfile [codfw] DONE helmfile.d/services/linked-artifacts: apply * 16:18 eevans@deploy2003: helmfile [codfw] START helmfile.d/services/linked-artifacts: apply * 16:08 eevans@deploy2003: helmfile [staging] DONE helmfile.d/services/linked-artifacts: apply * 16:07 eevans@deploy2003: helmfile [staging] START helmfile.d/services/linked-artifacts: apply * 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 16:01 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 15:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1024.eqiad.wmnet with OS bookworm * 15:49 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] (duration: 00m 10s) * 15:49 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] * 15:48 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] (duration: 00m 15s) * 15:48 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] * 15:47 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] (duration: 00m 10s) * 15:47 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] * 15:46 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:42 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:42 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:40 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:37 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 15:37 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:36 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:36 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:36 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 15:33 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 15:33 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:31 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:28 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 15:27 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:27 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1024.eqiad.wmnet with reason: host reimage * 15:23 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:23 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:23 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:20 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1068.eqiad.wmnet * 15:20 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1068.eqiad.wmnet * 15:20 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1068.eqiad.wmnet * 15:20 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 15:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1024.eqiad.wmnet with reason: host reimage * 15:11 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:55 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 14:52 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wdqs1024.eqiad.wmnet with OS bookworm * 14:50 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] (duration: 00m 09s) * 14:50 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] * 14:49 jiji@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 14:49 jiji@deploy2003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 14:49 jiji@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 14:48 jiji@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 14:45 ecarg@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:45 ecarg@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:44 ecarg@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:44 ecarg@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:43 ecarg@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:43 ecarg@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:41 sukhe: ipvsadm --delete-service --tcp-service 10.2.1.55:8087: lvs2014 and lvs2013 * 14:39 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:39 sukhe: ipvsadm --delete-service --tcp-service 10.2.2.55:8087: [[phab:T432445|T432445]] * 14:38 ecarg@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:38 ecarg@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:37 ecarg@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:37 ecarg@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:36 ecarg@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:34 ecarg@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts datahubsearch1001.eqiad.wmnet * 14:32 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:32 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 14:31 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 14:31 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 14:28 sukhe: sudo cumin 'A:lvs-low-traffic-codfw' 'systemctl restart pybal': lvs2013 * 14:26 sukhe: sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal': lvs2014 * 14:26 sukhe: sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal' * 14:24 sukhe: restart pybal on lvs1019 * 14:24 sukhe: restart pybal on lvs1020 * 14:19 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for 1036 hosts * 14:17 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:04 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] (duration: 09m 28s) * 13:59 kharlan@deploy2003: dreamyjazz, kharlan: Continuing with deployment * 13:58 bking@cumin2003: START - Cookbook sre.hosts.decommission for hosts datahubsearch1001.eqiad.wmnet * 13:57 kharlan@deploy2003: dreamyjazz, kharlan: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:55 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] * 13:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts datahubsearch[1002-1003].eqiad.wmnet * 13:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:53 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch[1002-1003].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 13:52 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch[1002-1003].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 13:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 13:42 stran@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] (duration: 07m 30s) * 13:42 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:40 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: Test * 13:38 stran@deploy2003: dragoniez, stran: Continuing with deployment * 13:37 stran@deploy2003: dragoniez, stran: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:35 bking@cumin2003: START - Cookbook sre.hosts.decommission for hosts datahubsearch[1002-1003].eqiad.wmnet * 13:35 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024'] * 13:35 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 13:35 stran@deploy2003: Started scap sync-world: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] * 13:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 13:28 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024'] * 13:26 sukhe@dns1004: END - running authdns-update * 13:25 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 13:25 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 13:24 sukhe@dns1004: START - running authdns-update * 13:22 stran@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] (duration: 08m 20s) * 13:21 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 13:20 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 13:19 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 13:19 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 13:19 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 13:18 stran@deploy2003: stran: Continuing with deployment * 13:16 stran@deploy2003: stran: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:14 stran@deploy2003: Started scap sync-world: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] * 13:13 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 13:13 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 13:11 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 13:11 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 13:08 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 12:55 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool es2051: Test * 12:55 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: Test * 12:54 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool es2051: Test * 12:43 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1036 hosts * 12:41 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 12:40 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1048.eqiad.wmnet * 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 12:39 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 12:38 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 12:37 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 12:37 brouberol@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 12:36 brouberol@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 12:36 brouberol@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 12:36 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 12:36 elukey@cumin1003: DONE (PASS) - Cookbook sre.puppet.renew-cert (exit_code=0) for crm2001.codfw.wmnet: Renew puppet certificate - elukey@cumin1003 * 12:35 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:35 brouberol@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 12:34 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 12:31 brouberol@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 12:30 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1048.eqiad.wmnet * 12:30 brouberol@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 12:28 brouberol@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 12:27 brouberol@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 12:20 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1068.eqiad.wmnet with OS trixie * 12:01 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] (duration: 13m 19s) * 11:58 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1068.eqiad.wmnet with reason: host reimage * 11:52 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1068.eqiad.wmnet with reason: host reimage * 11:51 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 11:49 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:47 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] * 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2252: Security updates * 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:43 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 11:42 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2252: Security updates * 11:42 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply * 11:40 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply * 11:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 11:37 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2252.codfw.wmnet with OS trixie * 11:34 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1068 * 11:34 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1068 * 11:26 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1068 * 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1068.eqiad.wmnet 46.48.64.10.in-addr.arpa 6.4.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:26 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1068.eqiad.wmnet 46.48.64.10.in-addr.arpa 6.4.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1068 - jiji@cumin1003" * 11:26 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1068 - jiji@cumin1003" * 11:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2252.codfw.wmnet with reason: host reimage * 11:17 jiji@cumin1003: START - Cookbook sre.dns.netbox * 11:17 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1068 * 11:17 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1068.eqiad.wmnet with OS trixie * 11:17 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2252.codfw.wmnet with reason: host reimage * 11:15 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1068.eqiad.wmnet * 11:15 mvolz@deploy2003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:15 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1068.eqiad.wmnet * 11:15 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1068.eqiad.wmnet * 11:14 mvolz@deploy2003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:13 mvolz@deploy2003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:13 mvolz@deploy2003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:12 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] (duration: 11m 05s) * 11:11 mvolz@deploy2003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:10 mvolz@deploy2003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:07 dreamyjazz@deploy2003: dreamyjazz, kharlan: Continuing with deployment * 11:03 dreamyjazz@deploy2003: dreamyjazz, kharlan: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:03 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2252.codfw.wmnet with OS trixie * 11:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2252: Upgrading db2252.codfw.wmnet * 11:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:02 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 11:02 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2252: Upgrading db2252.codfw.wmnet * 11:02 cwilliams@cumin1003: dbmaint on ms3@codfw [[phab:T432321|T432321]] * 11:01 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 11:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db1153.eqiad.wmnet with reason: Security updates * 11:01 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] * 11:00 fnegri@deploy2003: helmfile [eqiad] DONE helmfile.d/services/toolhub: apply * 10:58 fnegri@deploy2003: helmfile [eqiad] START helmfile.d/services/toolhub: apply * 10:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1151: Security updates * 10:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:57 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1151: Security updates * 10:55 fnegri@deploy2003: helmfile [codfw] DONE helmfile.d/services/toolhub: apply * 10:54 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] (duration: 08m 38s) * 10:53 fnegri@deploy2003: helmfile [codfw] START helmfile.d/services/toolhub: apply * 10:53 fnegri@deploy2003: helmfile [staging] DONE helmfile.d/services/toolhub: apply * 10:52 fnegri@deploy2003: helmfile [staging] START helmfile.d/services/toolhub: apply * 10:50 zabe@deploy2003: zabe: Continuing with deployment * 10:47 zabe@deploy2003: zabe: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:45 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] * 10:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1151: Security updates * 10:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:42 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:42 root@cumin1003: START - Cookbook sre.mysql.depool depool db1151: Security updates * 10:38 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] (duration: 12m 47s) * 10:34 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 10:34 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 10:33 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2253.codfw.wmnet with OS trixie * 10:28 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:26 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] * 10:18 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2253.codfw.wmnet with reason: host reimage * 10:13 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2253.codfw.wmnet with reason: host reimage * 10:00 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2253.codfw.wmnet with OS trixie * 09:58 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db1151.eqiad.wmnet with reason: Security updates * 09:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2253: Upgrading db2253.codfw.wmnet * 09:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:57 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 09:56 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2253: Upgrading db2253.codfw.wmnet * 09:56 cwilliams@cumin1003: dbmaint on ms2@codfw [[phab:T432321|T432321]] * 09:56 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 09:36 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: UI improvement; support url shortener - oblivian@cumin1003" * 09:36 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: UI improvement; support url shortener - oblivian@cumin1003 * 09:35 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: UI improvement; support url shortener - oblivian@cumin1003 * 09:35 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: UI improvement; support url shortener - oblivian@cumin1003" * 09:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1152: Security updates * 09:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:26 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db1152: Security updates * 09:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: Security updates * 09:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:11 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:11 root@cumin1003: START - Cookbook sre.mysql.depool depool db1152: Security updates * 09:10 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1018.eqiad.wmnet with reason: Cloning * 09:09 Dreamy_Jazz: Deployed patch for [[phab:T432453|T432453]] and [[phab:T432454|T432454]] * 09:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 09:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2251.codfw.wmnet with OS trixie * 08:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2251.codfw.wmnet with reason: host reimage * 08:45 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2251.codfw.wmnet with reason: host reimage * 08:40 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] (duration: 12m 26s) * 08:38 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1030.eqiad.wmnet,service=s1 * 08:36 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 08:31 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2251.codfw.wmnet with OS trixie * 08:30 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:28 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] * 08:25 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] (duration: 07m 59s) * 08:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2251: Upgrading db2251.codfw.wmnet * 08:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:22 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 08:22 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2251: Upgrading db2251.codfw.wmnet * 08:20 urbanecm@deploy2003: urbanecm: Continuing with deployment * 08:20 cwilliams@cumin1003: dbmaint on ms1@codfw [[phab:T432321|T432321]] * 08:20 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 08:19 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:17 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] * 08:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade * 08:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade * 08:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db2251.codfw.wmnet,db1152.eqiad.wmnet with reason: OS upgrade * 08:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade * 08:13 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade * 08:11 Dreamy_Jazz: Created cusi_signal, cusi_case, and cusi_user on ukwiki and enwikivoyage in extension1 * 08:11 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade * 08:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade * 08:04 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1030.eqiad.wmnet,service=s1 * 08:04 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1030.eqiad.wmnet,service=s1 * 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply * 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply * 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply * 07:51 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply * 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 07:47 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 07:47 phuedx: End of UTC morning backport window * 07:43 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 07:43 phuedx@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] (duration: 13m 44s) * 07:43 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 07:39 phuedx@deploy2003: phuedx: Continuing with deployment * 07:31 phuedx@deploy2003: phuedx: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:29 phuedx@deploy2003: Started scap sync-world: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] * 07:24 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Turnilo import support - oblivian@cumin1003" * 07:24 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import support - oblivian@cumin1003 * 07:23 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import support - oblivian@cumin1003 * 07:23 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Turnilo import support - oblivian@cumin1003" * 06:42 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs2020 after successful Bookworm reimage, data transfer, and postflight validation * 06:42 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2020.codfw.wmnet * 05:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Managing sanitization for wikis bolwiki in section s5 * 05:25 marostegui@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis bolwiki in section s5 * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 41s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-21 == * 22:50 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2019.codfw.wmnet -> wdqs2020.codfw.wmnet, repooling source-only afterwards * 22:47 cwhite: force reboot arclamp2001 - appears to have run out of memory and gone unresponsive * 22:24 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 01m 26s) * 22:24 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 22:23 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 22:22 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024'] * 22:11 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 21:54 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs1024'] * 21:54 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 21:53 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs1024'] * 21:53 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 21:49 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2019.codfw.wmnet -> wdqs2020.codfw.wmnet, repooling source-only afterwards * 20:57 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] (duration: 09m 10s) * 20:55 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1024.eqiad.wmnet with OS bookworm * 20:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2020.codfw.wmnet with OS bookworm * 20:52 krinkle@deploy2003: krinkle: Continuing with deployment * 20:49 krinkle@deploy2003: krinkle: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:47 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] * 20:45 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] (duration: 05m 42s) * 20:44 krinkle@deploy2003: krinkle: Rolling back deployment * 20:41 krinkle@deploy2003: krinkle: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:39 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] * 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2003.codfw.wmnet * 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1003.eqiad.wmnet * 20:33 dani@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] (duration: 11m 15s) * 20:33 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2003.codfw.wmnet * 20:33 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1003.eqiad.wmnet * 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2020.codfw.wmnet with reason: host reimage * 20:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1002.eqiad.wmnet * 20:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2002.codfw.wmnet * 20:30 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 20:29 dani@deploy2003: dani: Continuing with deployment * 20:29 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2020.codfw.wmnet with reason: host reimage * 20:26 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1002.eqiad.wmnet * 20:26 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2002.codfw.wmnet * 20:24 dani@deploy2003: dani: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2001.codfw.wmnet * 20:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1001.eqiad.wmnet * 20:22 dani@deploy2003: Started scap sync-world: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] * 20:22 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 20:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2001.codfw.wmnet * 20:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1001.eqiad.wmnet * 20:14 sbisson@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] (duration: 09m 01s) * 20:11 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2020.codfw.wmnet with OS bookworm * 20:10 sbisson@deploy2003: sbisson: Continuing with deployment * 20:07 sbisson@deploy2003: sbisson: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:05 sbisson@deploy2003: Started scap sync-world: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] * 20:03 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] (duration: 07m 04s) * 20:01 mutante: Gerrit - tomorrow a new SSH host key will appear - it will be {{Gerrit|ed25519}} and has already been added to wmf-laptop. you can verify it here: https://wikitech.wikimedia.org/wiki/Help:SSH_Fingerprints/gerrit.wikimedia.org:29418 ([[phab:T240266|T240266]]) * 19:59 zabe@deploy2003: zabe: Continuing with deployment * 19:58 zabe@deploy2003: zabe: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:56 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] * 19:52 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] (duration: 07m 25s) * 19:48 zabe@deploy2003: zabe: Continuing with deployment * 19:47 zabe@deploy2003: zabe: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:45 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] * 19:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 19:32 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024'] * 19:27 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 19:26 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024'] * 19:26 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 19:24 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024'] * 19:12 ryankemper: [wdqs] [[phab:T430880|T430880]] Repooled `wdqs-scholarly` discovery in `eqiad` after validating `wdqs1023` end-to-end; `wdqs1024` remains disabled pending reimage recovery * 19:11 ryankemper: [wdqs] [[phab:T430880|T430880]] Repooled wdqs1012.eqiad.wmnet after successful Bookworm reimage, data transfer, service checks, readiness probe, and cross-graph federation query validation * 19:10 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 19:10 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1012.eqiad.wmnet * 19:08 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 18:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for deploy1003.eqiad.wmnet * 18:57 kamila@cumin1003: START - Cookbook sre.hosts.remove-downtime for deploy1003.eqiad.wmnet * 18:37 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:37 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding urldownloader service IPs - sukhe@cumin1003" * 18:37 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding urldownloader service IPs - sukhe@cumin1003" * 18:32 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 18:32 dancy@deploy2003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 18:30 sukhe@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 18:27 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 18:24 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1024.eqiad.wmnet with OS bookworm * 18:20 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host deploy1003.eqiad.wmnet with OS bookworm * 18:09 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deploy1003 reimage (duration: 121m 16s) * 18:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 18:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1180: Security updates * 17:55 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wcqs2003.codfw.wmnet * 17:48 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wcqs2003.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1155.eqiad.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1155.eqiad.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2224.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2224.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2217.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2217.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2193.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2193.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2180.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2180.codfw.wmnet * 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1168.eqiad.wmnet * 17:36 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1168.eqiad.wmnet * 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2169.codfw.wmnet * 17:36 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2169.codfw.wmnet * 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1165.eqiad.wmnet * 17:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1165.eqiad.wmnet * 17:35 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2158.codfw.wmnet * 17:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2158.codfw.wmnet * 17:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wcqs1003.eqiad.wmnet * 17:17 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1180: Security updates * 17:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1180.eqiad.wmnet * 17:16 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1180.eqiad.wmnet * 17:15 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp3073.* * 17:13 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wcqs1003.eqiad.wmnet * 17:11 brett@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp3073.esams.wmnet with OS trixie * 17:11 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 17:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1024 * 17:04 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1024 * 17:03 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 17:00 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 16:59 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2242: codfw rack B7 depool for maintenance * 16:59 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 16:43 brett@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp3073.esams.wmnet with reason: host reimage * 16:42 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 16:39 brett@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cp3073.esams.wmnet with reason: host reimage * 16:32 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on deploy1003.eqiad.wmnet with reason: host reimage * 16:27 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on deploy1003.eqiad.wmnet with reason: host reimage * 16:14 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2242: codfw rack B7 depool for maintenance * 16:14 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: codfw rack B7 depool for maintenance * 16:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1023.eqiad.wmnet with OS bookworm * 16:13 brett@cumin2002: START - Cookbook sre.hosts.reimage for host cp3073.esams.wmnet with OS trixie * 16:08 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host deploy1003.eqiad.wmnet with OS bookworm * 16:08 kamila@deploy2003: Locking from deployment [MediaWiki]: deploy1003 reimage * 16:03 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1012.eqiad.wmnet with OS bookworm * 15:48 inflatador: bking@apt1002 `sudo reprepro copy bookworm-wikimedia bullseye-wikimedia jvmquake` [[phab:T430880|T430880]] * 15:39 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp3073.* * 15:39 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 15:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:34 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1311536{{!}}Set $wgMathInternalRestbaseURL explicitly (take 2) (T349582)]] * 15:29 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:29 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2228: codfw rack B7 depool for maintenance * 15:29 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2229: codfw rack B7 depool for maintenance * 15:27 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1180: Security update * 15:25 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Security update * 15:21 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 15:21 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 15:19 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 24s) * 15:19 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:14 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 15:14 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db1180: Security update * 15:13 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-eqiad * 14:48 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-eqiad * 14:44 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2229: codfw rack B7 depool for maintenance * 14:44 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc2017: codfw rack B7 depool for maintenance * 14:44 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:43 cmooney@cumin2003: START - Cookbook sre.mysql.parsercache * 14:43 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool pc2017: codfw rack B7 depool for maintenance * 14:43 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2003.codfw.wmnet * 14:43 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2003.codfw.wmnet * 14:42 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2009.codfw.wmnet * 14:42 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2009.codfw.wmnet * 14:41 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:41 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:40 cmooney@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 29 hosts * 14:40 cmooney@cumin1003: START - Cookbook sre.hosts.remove-downtime for 29 hosts * 14:35 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 14:34 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 14:32 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] (duration: 07m 56s) * 14:29 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ssw1-a[1,8]-codfw with reason: lsw1-b7-codfw JunOS upgrade * 14:28 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 14:28 elukey@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 14:26 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:24 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] * 14:23 topranks: reboot lsw1-b7-codfw to upgrade JunOS (affects all hosts in rack) [[phab:T430928|T430928]] * 14:18 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2003.codfw.wmnet * 14:14 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2009.codfw.wmnet * 14:14 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Security update * 14:13 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2242: codfw rack B7 depool for maintenance * 14:13 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2242: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2228: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2228: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2229: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2229: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc2017: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.parsercache * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool pc2017: codfw rack B7 depool for maintenance * 14:08 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2003.codfw.wmnet * 14:07 cmooney@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on aux-k8s-etcd2004.codfw.wmnet,ml-etcd2001.codfw.wmnet with reason: lsw1-b7-codfw JunOS upgrade * 14:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95005 and previous config saved to /var/cache/conftool/dbconfig/20260721-140620-cwilliams.json * 14:05 cmooney@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2049.codfw.wmnet * 14:05 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-scholarly,name=eqiad * 14:04 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2009.codfw.wmnet * 14:04 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply * 14:04 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply * 14:03 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:03 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2001.codfw.wmnet * 14:03 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2001.codfw.wmnet * 14:02 cmooney@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2049.codfw.wmnet * 14:00 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:00 Dreamy_Jazz: Created cusi_case, cusi_signal, and cusi_user on svwiki, dewiki, jawiki, eswiki * 13:59 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b7-codfw,lsw1-b7-codfw IPv6,lsw1-b7-codfw.mgmt,ssw1-a[1,8]-codfw.mgmt with reason: lsw1-b7-codfw JunOS upgrade * 13:57 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs1023.eqiad.wmnet, repooling source-only afterwards * 13:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 29 hosts with reason: lsw1-b7-codfw JunOS upgrade * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224', diff saved to https://phabricator.wikimedia.org/P95003 and previous config saved to /var/cache/conftool/dbconfig/20260721-135613-cwilliams.json * 13:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1012.eqiad.wmnet with reason: host reimage * 13:53 cmooney@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 1:00:00 on 30 hosts with reason: lsw1-b7-codfw JunOS upgrade * 13:51 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1012.eqiad.wmnet with reason: host reimage * 13:48 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 13:48 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 13:46 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224', diff saved to https://phabricator.wikimedia.org/P95001 and previous config saved to /var/cache/conftool/dbconfig/20260721-134605-cwilliams.json * 13:46 cmooney@cumin1003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti2033.codfw.wmnet * 13:46 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 13:45 elukey: move the Docker Registry's /v2/wikimedia/machinelearning.* prefix to the ml S3 backend - [[phab:T428022|T428022]] * 13:45 cmooney@cumin1003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti2033.codfw.wmnet * 13:45 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 13:43 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:40 jiji@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 13:40 jiji@deploy2003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 13:39 jiji@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 13:39 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 13:38 cmooney@cumin1003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2032.codfw.wmnet * 13:38 jiji@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 13:38 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 13:37 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2032.codfw.wmnet * 13:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95000 and previous config saved to /var/cache/conftool/dbconfig/20260721-133557-cwilliams.json * 13:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1012 * 13:33 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1012 * 13:33 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1012.eqiad.wmnet with OS bookworm * 13:30 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:30 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:28 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94999 and previous config saved to /var/cache/conftool/dbconfig/20260721-132855-cwilliams.json * 13:28 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2224.codfw.wmnet with reason: Maintenance * 13:28 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94998 and previous config saved to /var/cache/conftool/dbconfig/20260721-132826-cwilliams.json * 13:28 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 13:23 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] (duration: 07m 50s) * 13:20 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 13:18 kharlan@deploy2003: kharlan: Continuing with deployment * 13:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217', diff saved to https://phabricator.wikimedia.org/P94996 and previous config saved to /var/cache/conftool/dbconfig/20260721-131817-cwilliams.json * 13:17 brouberol@dns1004: END - running authdns-update * 13:17 kharlan@deploy2003: kharlan: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:15 brouberol@dns1004: START - running authdns-update * 13:15 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] * 13:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94995 and previous config saved to /var/cache/conftool/dbconfig/20260721-131411-cwilliams.json * 13:13 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs1023.eqiad.wmnet, repooling source-only afterwards * 13:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217', diff saved to https://phabricator.wikimedia.org/P94994 and previous config saved to /var/cache/conftool/dbconfig/20260721-130809-cwilliams.json * 13:07 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 13:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180', diff saved to https://phabricator.wikimedia.org/P94993 and previous config saved to /var/cache/conftool/dbconfig/20260721-130404-cwilliams.json * 13:03 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:03 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:02 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 13:02 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 12:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94992 and previous config saved to /var/cache/conftool/dbconfig/20260721-125801-cwilliams.json * 12:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180', diff saved to https://phabricator.wikimedia.org/P94991 and previous config saved to /var/cache/conftool/dbconfig/20260721-125356-cwilliams.json * 12:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94990 and previous config saved to /var/cache/conftool/dbconfig/20260721-125049-cwilliams.json * 12:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2217.codfw.wmnet with reason: Maintenance * 12:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94989 and previous config saved to /var/cache/conftool/dbconfig/20260721-125017-cwilliams.json * 12:48 elukey: bmc cold reboot for lvs1013 and lvs1015 - [[phab:T426180|T426180]] * 12:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94988 and previous config saved to /var/cache/conftool/dbconfig/20260721-124348-cwilliams.json * 12:40 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193', diff saved to https://phabricator.wikimedia.org/P94987 and previous config saved to /var/cache/conftool/dbconfig/20260721-124009-cwilliams.json * 12:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts * 12:33 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts * 12:33 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts * 12:32 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts * 12:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:30 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193', diff saved to https://phabricator.wikimedia.org/P94986 and previous config saved to /var/cache/conftool/dbconfig/20260721-123001-cwilliams.json * 12:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94985 and previous config saved to /var/cache/conftool/dbconfig/20260721-121953-cwilliams.json * 12:17 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs2007.codfw.wmnet with OS bookworm * 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94983 and previous config saved to /var/cache/conftool/dbconfig/20260721-121257-cwilliams.json * 12:12 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2193.codfw.wmnet with reason: Maintenance * 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94982 and previous config saved to /var/cache/conftool/dbconfig/20260721-121239-cwilliams.json * 12:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180', diff saved to https://phabricator.wikimedia.org/P94980 and previous config saved to /var/cache/conftool/dbconfig/20260721-120231-cwilliams.json * 11:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180', diff saved to https://phabricator.wikimedia.org/P94979 and previous config saved to /var/cache/conftool/dbconfig/20260721-115223-cwilliams.json * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94978 and previous config saved to /var/cache/conftool/dbconfig/20260721-114333-cwilliams.json * 11:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1180.eqiad.wmnet with reason: Maintenance * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94977 and previous config saved to /var/cache/conftool/dbconfig/20260721-114305-cwilliams.json * 11:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94976 and previous config saved to /var/cache/conftool/dbconfig/20260721-114215-cwilliams.json * 11:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94975 and previous config saved to /var/cache/conftool/dbconfig/20260721-113530-cwilliams.json * 11:35 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2180.codfw.wmnet with reason: Maintenance * 11:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94974 and previous config saved to /var/cache/conftool/dbconfig/20260721-113501-cwilliams.json * 11:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168', diff saved to https://phabricator.wikimedia.org/P94973 and previous config saved to /var/cache/conftool/dbconfig/20260721-113258-cwilliams.json * 11:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169', diff saved to https://phabricator.wikimedia.org/P94972 and previous config saved to /var/cache/conftool/dbconfig/20260721-112453-cwilliams.json * 11:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168', diff saved to https://phabricator.wikimedia.org/P94971 and previous config saved to /var/cache/conftool/dbconfig/20260721-112250-cwilliams.json * 11:21 XioNoX: put eqiad-drmrs Arelion link in service * 11:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169', diff saved to https://phabricator.wikimedia.org/P94970 and previous config saved to /var/cache/conftool/dbconfig/20260721-111446-cwilliams.json * 11:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94969 and previous config saved to /var/cache/conftool/dbconfig/20260721-111242-cwilliams.json * 11:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 11:10 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 11:07 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1093 hosts * 11:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94968 and previous config saved to /var/cache/conftool/dbconfig/20260721-110548-cwilliams.json * 11:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1168.eqiad.wmnet with reason: Maintenance * 11:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94967 and previous config saved to /var/cache/conftool/dbconfig/20260721-110520-cwilliams.json * 11:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94966 and previous config saved to /var/cache/conftool/dbconfig/20260721-110439-cwilliams.json * 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94964 and previous config saved to /var/cache/conftool/dbconfig/20260721-105632-cwilliams.json * 10:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2169.codfw.wmnet with reason: Maintenance * 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94963 and previous config saved to /var/cache/conftool/dbconfig/20260721-105603-cwilliams.json * 10:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165', diff saved to https://phabricator.wikimedia.org/P94962 and previous config saved to /var/cache/conftool/dbconfig/20260721-105512-cwilliams.json * 10:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158', diff saved to https://phabricator.wikimedia.org/P94961 and previous config saved to /var/cache/conftool/dbconfig/20260721-104555-cwilliams.json * 10:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165', diff saved to https://phabricator.wikimedia.org/P94960 and previous config saved to /var/cache/conftool/dbconfig/20260721-104504-cwilliams.json * 10:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158', diff saved to https://phabricator.wikimedia.org/P94959 and previous config saved to /var/cache/conftool/dbconfig/20260721-103547-cwilliams.json * 10:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94958 and previous config saved to /var/cache/conftool/dbconfig/20260721-103456-cwilliams.json * 10:29 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2229: Upgraded kernel * 10:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94956 and previous config saved to /var/cache/conftool/dbconfig/20260721-102757-cwilliams.json * 10:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on an-redacteddb1001.eqiad.wmnet,clouddb[1015,1025,1028].eqiad.wmnet,db1155.eqiad.wmnet with reason: Maintenance * 10:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1165.eqiad.wmnet with reason: Maintenance * 10:25 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94955 and previous config saved to /var/cache/conftool/dbconfig/20260721-102539-cwilliams.json * 10:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94954 and previous config saved to /var/cache/conftool/dbconfig/20260721-101848-cwilliams.json * 10:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2158.codfw.wmnet with reason: Maintenance * 09:43 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2229: Upgraded kernel * 09:42 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2229.codfw.wmnet * 09:42 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2229.codfw.wmnet * 09:23 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db2229.codfw.wmnet * 09:23 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2229.codfw.wmnet * 08:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2229 [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94948 and previous config saved to /var/cache/conftool/dbconfig/20260721-085724-cwilliams.json * 08:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2214 to s6 primary [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94947 and previous config saved to /var/cache/conftool/dbconfig/20260721-085442-cwilliams.json * 08:53 cezmunsta: Starting s6 codfw failover from db2229 to db2214 - [[phab:T430964|T430964]] * 08:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2214 with weight 0 [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94946 and previous config saved to /var/cache/conftool/dbconfig/20260721-084613-cwilliams.json * 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 22 hosts with reason: Primary switchover s6 [[phab:T430964|T430964]] * 08:32 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1017.eqiad.wmnet,service=s1 * 08:08 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Add subrated circuit rate to interface descriptions - CR1312476 - ayounsi@cumin1003 * 08:06 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Add subrated circuit rate to interface descriptions - CR1312476 - ayounsi@cumin1003 * 07:58 reedy@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] (duration: 12m 55s) * 07:51 reedy@deploy2003: reedy, neriah: Continuing with deployment * 07:51 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 07:51 reedy@deploy2003: reedy, neriah: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:48 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1093 hosts * 07:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm2001.wikimedia.org * 07:45 reedy@deploy2003: Started scap sync-world: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] * 07:43 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 07:42 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm2001.wikimedia.org * 07:23 elukey: upgrade libtiff6 packages on zuul* trixie hosts for security upgrades * 07:22 elukey: upgrade libtiff6 packages on Wikikube trixie workers for security upgrades * 07:14 elukey@deploy2003: helmfile [codfw] DONE helmfile.d/services/proton: sync * 07:13 elukey@deploy2003: helmfile [codfw] START helmfile.d/services/proton: sync * 07:11 elukey@deploy2003: helmfile [eqiad] DONE helmfile.d/services/proton: sync * 07:10 elukey@deploy2003: helmfile [eqiad] START helmfile.d/services/proton: sync * 07:09 elukey@deploy2003: helmfile [staging] DONE helmfile.d/services/proton: sync * 07:08 elukey@deploy2003: helmfile [staging] START helmfile.d/services/proton: sync * 06:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1023.eqiad.wmnet with reason: host reimage * 06:46 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1023.eqiad.wmnet with reason: host reimage * 06:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 05:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Haproxy-only mode support - oblivian@cumin1003" * 05:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Haproxy-only mode support - oblivian@cumin1003 * 05:42 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Haproxy-only mode support - oblivian@cumin1003 * 05:42 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Haproxy-only mode support - oblivian@cumin1003" * 05:38 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1017.eqiad.wmnet with reason: Cloning * 05:37 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1017.eqiad.wmnet,service=s1 * 05:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1029.eqiad.wmnet,service=s8 * 05:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1029.eqiad.wmnet,service=s5 * 05:32 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:30 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 05:11 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet * 05:04 aokoth@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet * 05:00 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 04:56 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 04:01 mwpresync@deploy2003: Pruned MediaWiki: 1.47.0-wmf.9 (duration: 01m 08s) * 03:41 mwpresync@deploy2003: Finished scap sync-world: testwikis to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] (duration: 36m 30s) * 03:05 mwpresync@deploy2003: Started scap sync-world: testwikis to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 03:01 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:01 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:00 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:00 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:36 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:36 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:36 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:35 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:16 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 47s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 00:56 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm == 2026-07-20 == * 23:38 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 23:07 Amir1: deleting echo notifications from 2015 on group1 wikis * 23:07 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] (duration: 14m 16s) * 23:01 ladsgroup@deploy2003: ladsgroup: Continuing with deployment * 23:00 ladsgroup@deploy2003: ladsgroup: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:53 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] * 22:46 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2007.codfw.wmnet, repooling source-only afterwards * 22:39 maryum: Deployed security fixes for several security bugs * 21:42 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 21:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2007.codfw.wmnet, repooling source-only afterwards * 21:37 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 21:37 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 21:34 sbassett: Deployed security fix for [[phab:T432424|T432424]] * 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs2020.codfw.wmnet * 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1023.eqiad.wmnet * 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1011.eqiad.wmnet * 21:32 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 17s) * 21:32 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 21:27 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 21:13 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2007.codfw.wmnet with reason: host reimage * 21:08 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-internal-main,name=codfw * 21:06 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2007.codfw.wmnet with reason: host reimage * 20:59 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service * 20:58 sukhe: pybal restart for IP changes around wdqs-main hosts * 20:57 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 20:46 ryankemper@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-internal-main,name=codfw * 20:45 ebernhardson@deploy2003: Finished deploy [search/mjolnir/deploy@d4dc3b8]: Update for opensearch 2.x compat (duration: 00m 34s) * 20:44 ebernhardson@deploy2003: Started deploy [search/mjolnir/deploy@d4dc3b8]: Update for opensearch 2.x compat * 20:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2007 * 20:44 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2007 * 20:43 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2007 * 20:43 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2007.codfw.wmnet 156.16.192.10.in-addr.arpa 6.5.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:42 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2007.codfw.wmnet 156.16.192.10.in-addr.arpa 6.5.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:42 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2007 - bking@cumin2003" * 20:41 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2007 - bking@cumin2003" * 20:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94944 and previous config saved to /var/cache/conftool/dbconfig/20260720-203333-cwilliams.json * 20:32 arlolra@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] (duration: 15m 07s) * 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2020.codfw.wmnet * 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1023.eqiad.wmnet * 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1011.eqiad.wmnet * 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs2020.codfw.wmnet * 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1023.eqiad.wmnet * 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1011.eqiad.wmnet * 20:25 arlolra@deploy2003: arlolra, cscott: Continuing with deployment * 20:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257', diff saved to https://phabricator.wikimedia.org/P94943 and previous config saved to /var/cache/conftool/dbconfig/20260720-202325-cwilliams.json * 20:21 arlolra@deploy2003: arlolra, cscott: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:17 arlolra@deploy2003: Started scap sync-world: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] * 20:13 bking@cumin2003: START - Cookbook sre.dns.netbox * 20:13 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257', diff saved to https://phabricator.wikimedia.org/P94942 and previous config saved to /var/cache/conftool/dbconfig/20260720-201318-cwilliams.json * 20:13 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 20:10 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 20:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2007 * 20:04 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2007.codfw.wmnet with OS bookworm * 20:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94941 and previous config saved to /var/cache/conftool/dbconfig/20260720-200310-cwilliams.json * 19:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94940 and previous config saved to /var/cache/conftool/dbconfig/20260720-195633-cwilliams.json * 19:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1257.eqiad.wmnet with reason: Maintenance * 19:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94939 and previous config saved to /var/cache/conftool/dbconfig/20260720-195605-cwilliams.json * 19:51 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 19:50 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 19:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256', diff saved to https://phabricator.wikimedia.org/P94938 and previous config saved to /var/cache/conftool/dbconfig/20260720-194558-cwilliams.json * 19:44 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 19:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 19:41 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wdqs1011.eqiad.wmnet with OS bookworm * 19:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256', diff saved to https://phabricator.wikimedia.org/P94937 and previous config saved to /var/cache/conftool/dbconfig/20260720-193550-cwilliams.json * 19:25 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94936 and previous config saved to /var/cache/conftool/dbconfig/20260720-192542-cwilliams.json * 19:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94935 and previous config saved to /var/cache/conftool/dbconfig/20260720-191856-cwilliams.json * 19:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1256.eqiad.wmnet with reason: Maintenance * 19:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94934 and previous config saved to /var/cache/conftool/dbconfig/20260720-191839-cwilliams.json * 19:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255', diff saved to https://phabricator.wikimedia.org/P94933 and previous config saved to /var/cache/conftool/dbconfig/20260720-190831-cwilliams.json * 18:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255', diff saved to https://phabricator.wikimedia.org/P94932 and previous config saved to /var/cache/conftool/dbconfig/20260720-185824-cwilliams.json * 18:50 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 18:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94931 and previous config saved to /var/cache/conftool/dbconfig/20260720-184816-cwilliams.json * 18:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94930 and previous config saved to /var/cache/conftool/dbconfig/20260720-184224-cwilliams.json * 18:42 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1255.eqiad.wmnet with reason: Maintenance * 18:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94929 and previous config saved to /var/cache/conftool/dbconfig/20260720-184153-cwilliams.json * 18:39 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 18:39 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 16s) * 18:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 18:38 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 59m 26s) * 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 18:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211', diff saved to https://phabricator.wikimedia.org/P94928 and previous config saved to /var/cache/conftool/dbconfig/20260720-183145-cwilliams.json * 18:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211', diff saved to https://phabricator.wikimedia.org/P94927 and previous config saved to /var/cache/conftool/dbconfig/20260720-182137-cwilliams.json * 18:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94926 and previous config saved to /var/cache/conftool/dbconfig/20260720-181129-cwilliams.json * 18:09 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_codfw * 18:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2057.codfw.wmnet * 18:08 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_codfw * 18:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2058.codfw.wmnet * 18:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94925 and previous config saved to /var/cache/conftool/dbconfig/20260720-180452-cwilliams.json * 18:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on clouddb[1016,1020,1022-1023].eqiad.wmnet,db1154.eqiad.wmnet with reason: Maintenance * 18:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1211.eqiad.wmnet with reason: Maintenance * 18:02 sukhe: armed keyholder on acmechief1002.eqiad.wmnet and acmechief2002.codfw.wmnet (active host) * 18:01 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief2002.codfw.wmnet * 17:57 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief2002.codfw.wmnet * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs2020'] * 17:52 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief1002.eqiad.wmnet * 17:50 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 17:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1011.eqiad.wmnet with reason: host reimage * 17:48 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief1002.eqiad.wmnet * 17:47 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test2001.codfw.wmnet * 17:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94924 and previous config saved to /var/cache/conftool/dbconfig/20260720-174717-cwilliams.json * 17:46 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 17:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1011.eqiad.wmnet with reason: host reimage * 17:43 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs2020.codfw.wmnet with OS bookworm * 17:43 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test2001.codfw.wmnet * 17:43 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test1001.eqiad.wmnet * 17:39 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test1001.eqiad.wmnet * 17:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 17:38 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:38 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 17:37 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 17:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244', diff saved to https://phabricator.wikimedia.org/P94923 and previous config saved to /var/cache/conftool/dbconfig/20260720-173709-cwilliams.json * 17:35 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 17:31 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 20m 40s) * 17:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2055.codfw.wmnet * 17:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2056.codfw.wmnet * 17:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1011 * 17:27 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1011 * 17:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1011.eqiad.wmnet with OS bookworm * 17:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244', diff saved to https://phabricator.wikimedia.org/P94922 and previous config saved to /var/cache/conftool/dbconfig/20260720-172701-cwilliams.json * 17:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94921 and previous config saved to /var/cache/conftool/dbconfig/20260720-171653-cwilliams.json * 17:11 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 17:11 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 13m 03s) * 17:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94920 and previous config saved to /var/cache/conftool/dbconfig/20260720-171012-cwilliams.json * 17:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2244.codfw.wmnet with reason: Maintenance * 17:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94919 and previous config saved to /var/cache/conftool/dbconfig/20260720-170941-cwilliams.json * 16:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243', diff saved to https://phabricator.wikimedia.org/P94918 and previous config saved to /var/cache/conftool/dbconfig/20260720-165933-cwilliams.json * 16:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 16:58 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2053.codfw.wmnet * 16:51 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2054.codfw.wmnet * 16:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243', diff saved to https://phabricator.wikimedia.org/P94917 and previous config saved to /var/cache/conftool/dbconfig/20260720-164926-cwilliams.json * 16:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94916 and previous config saved to /var/cache/conftool/dbconfig/20260720-163918-cwilliams.json * 16:35 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 16:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94915 and previous config saved to /var/cache/conftool/dbconfig/20260720-163140-cwilliams.json * 16:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2243.codfw.wmnet with reason: Maintenance * 16:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94914 and previous config saved to /var/cache/conftool/dbconfig/20260720-163111-cwilliams.json * 16:27 btullis@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 16:27 btullis@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 16:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2020 * 16:23 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2020 * 16:21 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2020 * 16:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2020.codfw.wmnet 85.0.192.10.in-addr.arpa 5.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:21 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2020.codfw.wmnet 85.0.192.10.in-addr.arpa 5.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242', diff saved to https://phabricator.wikimedia.org/P94913 and previous config saved to /var/cache/conftool/dbconfig/20260720-162103-cwilliams.json * 16:19 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 16:18 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 16:18 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:18 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 16:17 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:17 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for netbox accounting errors - jhancock@cumin2002" * 16:17 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for netbox accounting errors - jhancock@cumin2002" * 16:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2051.codfw.wmnet * 16:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2052.codfw.wmnet * 16:11 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 16:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242', diff saved to https://phabricator.wikimedia.org/P94912 and previous config saved to /var/cache/conftool/dbconfig/20260720-161055-cwilliams.json * 16:09 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 16:08 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 16:06 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 16:06 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 16:06 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2020 * 16:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2020.codfw.wmnet with OS bookworm * 16:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94911 and previous config saved to /var/cache/conftool/dbconfig/20260720-160047-cwilliams.json * 15:58 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2019.codfw.wmnet, repooling source-only afterwards * 15:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94909 and previous config saved to /var/cache/conftool/dbconfig/20260720-155353-cwilliams.json * 15:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2242.codfw.wmnet with reason: Maintenance * 15:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94908 and previous config saved to /var/cache/conftool/dbconfig/20260720-154433-cwilliams.json * 15:35 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2049.codfw.wmnet * 15:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162', diff saved to https://phabricator.wikimedia.org/P94907 and previous config saved to /var/cache/conftool/dbconfig/20260720-153425-cwilliams.json * 15:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2050.codfw.wmnet * 15:28 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162', diff saved to https://phabricator.wikimedia.org/P94906 and previous config saved to /var/cache/conftool/dbconfig/20260720-152418-cwilliams.json * 15:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1023 * 15:14 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1023 * 15:14 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 15:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94905 and previous config saved to /var/cache/conftool/dbconfig/20260720-151407-cwilliams.json * 15:13 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] (duration: 41m 16s) * 15:08 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 15:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94902 and previous config saved to /var/cache/conftool/dbconfig/20260720-150729-cwilliams.json * 15:07 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2162.codfw.wmnet with reason: Maintenance * 15:05 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2027.codfw.wmnet, repooling source-only afterwards * 15:00 urbanecm@deploy2003: vadymts1, migr, urbanecm: Continuing with deployment * 14:59 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:58 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 07s) * 14:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:58 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 13s) * 14:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:57 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2019.codfw.wmnet, repooling source-only afterwards * 14:57 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2047.codfw.wmnet * 14:55 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2048.codfw.wmnet * 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2019.codfw.wmnet with OS bookworm * 14:49 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 14:47 urbanecm@deploy2003: vadymts1, migr, urbanecm: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:44 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool magru [reason: BGP issues in lvs7003 resolved after liberica restart, no task ID specified] * 14:44 sukhe@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool magru [reason: BGP issues in lvs7003 resolved after liberica restart, no task ID specified] * 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:41 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:39 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:39 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:33 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool magru [reason: no reason specified, no task ID specified] * 14:33 sukhe@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool magru [reason: no reason specified, no task ID specified] * 14:31 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] * 14:24 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:24 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:24 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:24 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2019.codfw.wmnet with reason: host reimage * 14:22 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2027.codfw.wmnet, repooling source-only afterwards * 14:19 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2019.codfw.wmnet with reason: host reimage * 14:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2027.codfw.wmnet with OS bookworm * 14:16 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2046.codfw.wmnet * 14:16 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2045.codfw.wmnet * 14:08 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:08 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:08 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:08 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:07 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:06 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:06 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:06 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:05 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2019 * 14:00 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2019 * 13:56 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1015.eqiad.wmnet * 13:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2027.codfw.wmnet with reason: host reimage * 13:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2071.codfw.wmnet with OS trixie * 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:51 sukhe@cumin1003: END (ERROR) - Cookbook sre.loadbalancer.admin (exit_code=97) rebooting A:liberica and P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica and P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:51 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1015.eqiad.wmnet * 13:50 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1014.eqiad.wmnet * 13:50 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1076.eqiad.wmnet with OS trixie * 13:50 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2027.codfw.wmnet with reason: host reimage * 13:45 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1014.eqiad.wmnet * 13:44 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1013.eqiad.wmnet * 13:39 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 13:38 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1013.eqiad.wmnet * 13:37 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2044.codfw.wmnet * 13:37 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2043.codfw.wmnet * 13:36 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2019 * 13:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2019.codfw.wmnet 156.32.192.10.in-addr.arpa 6.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:36 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2019.codfw.wmnet 156.32.192.10.in-addr.arpa 6.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:36 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2019 - bking@cumin2003" * 13:36 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2019 - bking@cumin2003" * 13:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry2005.codfw.wmnet * 13:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2071.codfw.wmnet with reason: host reimage * 13:31 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:31 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2019 * 13:31 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry2005.codfw.wmnet * 13:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry2004.codfw.wmnet * 13:30 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2019.codfw.wmnet with OS bookworm * 13:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2027 * 13:30 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2027 * 13:30 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2027.codfw.wmnet with OS bookworm * 13:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1076.eqiad.wmnet with reason: host reimage * 13:29 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_codfw * 13:28 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_codfw * 13:26 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry2004.codfw.wmnet * 13:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry1005.eqiad.wmnet * 13:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2071.codfw.wmnet with reason: host reimage * 13:22 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1076.eqiad.wmnet with reason: host reimage * 13:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry1005.eqiad.wmnet * 13:21 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry1004.eqiad.wmnet * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry1004.eqiad.wmnet * 13:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts * 13:13 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts * 13:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts * 13:12 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts * 13:03 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1076.eqiad.wmnet with OS trixie * 13:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2071.codfw.wmnet with OS trixie * 12:55 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:54 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:53 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:46 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 7 hosts * 12:42 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 7 hosts * 12:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts * 12:42 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts * 12:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2070.codfw.wmnet with OS trixie * 12:36 ozge@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:35 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1075.eqiad.wmnet with OS trixie * 12:32 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts * 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts * 12:22 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin1001.eqiad.wmnet * 12:19 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin1001.eqiad.wmnet * 12:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2070.codfw.wmnet with reason: host reimage * 12:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin2001.codfw.wmnet * 12:14 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1075.eqiad.wmnet with reason: host reimage * 12:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2070.codfw.wmnet with reason: host reimage * 12:10 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1075.eqiad.wmnet with reason: host reimage * 12:09 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin2001.codfw.wmnet * 11:17 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1074.eqiad.wmnet with OS trixie * 11:17 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 11:16 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 11:14 ozge@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 6 hosts * 11:09 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 6 hosts * 11:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 324 hosts * 10:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1074.eqiad.wmnet with reason: host reimage * 10:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2069.codfw.wmnet with OS trixie * 10:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1074.eqiad.wmnet with reason: host reimage * 10:30 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2069.codfw.wmnet with reason: host reimage * 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1074.eqiad.wmnet with OS trixie * 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2069.codfw.wmnet with reason: host reimage * 10:06 blake@deploy2003: Stopping before sync operations * 10:06 blake@deploy2003: Started scap sync-world: Non-deployment scap run to populate new release values * 10:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2069.codfw.wmnet with OS trixie * 10:00 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1073.eqiad.wmnet with OS trixie * 09:56 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 324 hosts * 09:39 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 09:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 8 hosts * 09:38 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1073.eqiad.wmnet with reason: host reimage * 09:37 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 8 hosts * 09:34 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1073.eqiad.wmnet with reason: host reimage * 09:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2068.codfw.wmnet with OS trixie * 09:16 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1073.eqiad.wmnet with OS trixie * 09:13 blake@deploy2003: sync-world aborted: Non-deployment scap run to populate new release values (duration: 00m 02s) * 09:13 blake@deploy2003: Started scap sync-world: Non-deployment scap run to populate new release values * 08:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2068.codfw.wmnet with reason: host reimage * 08:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2068.codfw.wmnet with reason: host reimage * 08:50 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 08:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2068.codfw.wmnet with OS trixie * 08:15 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1072.eqiad.wmnet with OS trixie * 07:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2067.codfw.wmnet with OS trixie * 07:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1072.eqiad.wmnet with reason: host reimage * 07:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1072.eqiad.wmnet with reason: host reimage * 07:45 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 07:45 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 07:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2067.codfw.wmnet with reason: host reimage * 07:35 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2067.codfw.wmnet with reason: host reimage * 07:30 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 07:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1072.eqiad.wmnet with OS trixie * 07:30 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 07:17 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 07:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2067.codfw.wmnet with OS trixie * 05:51 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:50 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:25 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:25 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on db2207.codfw.wmnet with reason: Host down * 04:28 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 07m 02s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-18 == * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 29s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 00:11 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2018.codfw.wmnet, repooling source-only afterwards == 2026-07-17 == * 23:53 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2026.codfw.wmnet, repooling source-only afterwards * 23:09 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2018.codfw.wmnet, repooling source-only afterwards * 23:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2026.codfw.wmnet, repooling source-only afterwards * 22:11 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2018.codfw.wmnet with OS bookworm * 22:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2026.codfw.wmnet with OS bookworm * 21:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2018.codfw.wmnet with reason: host reimage * 21:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2018.codfw.wmnet with reason: host reimage * 21:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2026.codfw.wmnet with reason: host reimage * 21:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2026.codfw.wmnet with reason: host reimage * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2018 * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2018 * 21:26 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2018 * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2018.codfw.wmnet 155.32.192.10.in-addr.arpa 5.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:26 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2018.codfw.wmnet 155.32.192.10.in-addr.arpa 5.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2018 - bking@cumin2003" * 21:26 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2018 - bking@cumin2003" * 21:14 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:13 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2018 * 21:13 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2018.codfw.wmnet with OS bookworm * 21:12 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2026 * 21:12 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2026 * 21:12 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2026.codfw.wmnet with OS bookworm * 21:05 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1022.eqiad.wmnet -> wdqs1026.eqiad.wmnet, repooling source-only afterwards * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs2017.codfw.wmnet, repooling source-only afterwards * 20:11 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs2017.codfw.wmnet, repooling source-only afterwards * 20:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2017.codfw.wmnet with OS bookworm * 20:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1022.eqiad.wmnet -> wdqs1026.eqiad.wmnet, repooling source-only afterwards * 20:06 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1026.eqiad.wmnet with OS bookworm * 19:55 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 19:55 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:55 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 09s) * 19:55 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:50 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 08s) * 19:50 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:50 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 10m 03s) * 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2017.codfw.wmnet with reason: host reimage * 19:40 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:40 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 15s) * 19:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1026.eqiad.wmnet with reason: host reimage * 19:37 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 16s) * 19:37 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:34 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2017.codfw.wmnet with reason: host reimage * 19:34 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1026.eqiad.wmnet with reason: host reimage * 19:33 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 19:33 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:16 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1026.eqiad.wmnet with OS bookworm * 19:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2017 * 19:16 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2017 * 19:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2017.codfw.wmnet with OS bookworm * 18:30 bking@dns1004: END - running authdns-update * 18:28 bking@dns1004: START - running authdns-update * 18:16 kamila@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1264.eqiad.wmnet * 18:16 kamila@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1264.eqiad.wmnet * 18:16 kamila@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1264.eqiad.wmnet * 17:49 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 17:46 dzahn@dns1006: END - running authdns-update * 17:44 dzahn@dns1006: START - running authdns-update * 17:44 dzahn@dns1006: END - running authdns-update * 17:42 dzahn@dns1006: START - running authdns-update * 17:28 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 17:21 kamila@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 17:01 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1264 * 17:01 kamila@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1264 * 17:01 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 17:01 kamila@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1264.eqiad.wmnet * 17:01 kamila@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1264.eqiad.wmnet * 17:01 kamila@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1264.eqiad.wmnet * 16:42 reedy@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] (duration: 10m 29s) * 16:34 reedy@deploy2003: reedy, hartman: Continuing with deployment * 16:33 reedy@deploy2003: reedy, hartman: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:31 reedy@deploy2003: Started scap sync-world: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] * 16:26 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 16:10 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in2001.wikimedia.org with reason: [[phab:T431659|T431659]] * 16:07 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in1001.wikimedia.org with reason: [[phab:T431659|T431659]] * 16:05 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 16:01 kamila@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 16:00 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out2001.wikimedia.org with reason: [[phab:T431659|T431659]] * 15:41 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 15:41 kamila@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 15:35 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out1001.wikimedia.org with reason: [[phab:T431659|T431659]] * 15:14 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1339.eqiad.wmnet * 15:13 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1339.eqiad.wmnet * 15:13 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1339.eqiad.wmnet * 14:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1339.eqiad.wmnet with OS trixie * 14:50 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:49 kamila@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:49 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:33 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage * 14:27 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage * 14:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1339 * 14:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1339 * 14:14 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1339 * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1339.eqiad.wmnet 156.32.64.10.in-addr.arpa 6.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:14 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1339.eqiad.wmnet 156.32.64.10.in-addr.arpa 6.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1339 - cgoubert@cumin2003" * 14:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1339 - cgoubert@cumin2003" * 14:09 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 14:06 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1339 * 14:06 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie * 14:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1339.eqiad.wmnet * 14:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1339.eqiad.wmnet * 14:02 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1339.eqiad.wmnet * 13:45 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb1013.eqiad.wmnet * 13:39 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb1013.eqiad.wmnet * 13:27 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:24 blake@dns1004: END - running authdns-update * 13:22 blake@dns1004: START - running authdns-update * 13:20 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 13:11 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2014.codfw.wmnet * 13:06 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb2014.codfw.wmnet * 13:06 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2012.codfw.wmnet * 13:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 13:03 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 13:01 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 15 hosts * 13:01 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb2012.codfw.wmnet * 13:01 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1016.eqiad.wmnet * 13:00 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 15 hosts * 12:55 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb1016.eqiad.wmnet * 12:55 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1014.eqiad.wmnet * 12:49 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb1014.eqiad.wmnet * 12:32 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:32 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:31 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:31 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1338.eqiad.wmnet * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1338.eqiad.wmnet * 12:18 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1338.eqiad.wmnet * 12:17 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:15 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:14 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:13 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1338.eqiad.wmnet with OS trixie * 12:01 klausman@deploy2003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 11:59 klausman@deploy2003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 11:56 klausman@deploy2003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 11:54 klausman@deploy2003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 11:53 klausman@deploy2003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 11:51 klausman@deploy2003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 11:42 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1338.eqiad.wmnet with reason: host reimage * 11:38 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1338.eqiad.wmnet with reason: host reimage * 11:31 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2230.codfw.wmnet * 11:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1338 * 11:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1338 * 11:25 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1338 * 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1338.eqiad.wmnet 155.32.64.10.in-addr.arpa 5.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:25 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1338.eqiad.wmnet 155.32.64.10.in-addr.arpa 5.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1338 - cgoubert@cumin2003" * 11:25 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1338 - cgoubert@cumin2003" * 11:23 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2230.codfw.wmnet * 11:20 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 11:20 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1338 * 11:20 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1338.eqiad.wmnet with OS trixie * 11:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1338.eqiad.wmnet * 11:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1338.eqiad.wmnet * 11:19 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1338.eqiad.wmnet * 11:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1337.eqiad.wmnet * 11:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1337.eqiad.wmnet * 11:17 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1337.eqiad.wmnet * 11:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1337.eqiad.wmnet with OS trixie * 10:51 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[2001-2002].codfw.wmnet * 10:50 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1337.eqiad.wmnet with reason: host reimage * 10:40 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 10:39 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:39 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1337.eqiad.wmnet with reason: host reimage * 10:39 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1001-1003].eqiad.wmnet * 10:34 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:34 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:30 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:28 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1001-1003].eqiad.wmnet * 10:27 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1337 * 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1337 * 10:26 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1337 * 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1337.eqiad.wmnet 154.32.64.10.in-addr.arpa 4.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:26 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1337.eqiad.wmnet 154.32.64.10.in-addr.arpa 4.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1337 - cgoubert@cumin2003" * 10:26 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1337 - cgoubert@cumin2003" * 10:21 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 10:18 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1337 * 10:17 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1337.eqiad.wmnet with OS trixie * 10:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1337.eqiad.wmnet * 10:16 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db1176.eqiad.wmnet * 10:16 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1337.eqiad.wmnet * 10:16 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1337.eqiad.wmnet * 10:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1336.eqiad.wmnet * 10:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1336.eqiad.wmnet * 10:15 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1336.eqiad.wmnet * 10:11 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db1176.eqiad.wmnet * 10:10 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db1176.eqiad.wmnet * 10:09 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db1176.eqiad.wmnet * 10:05 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts (check the cookbook's logs for more details.) * 10:03 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts (check the cookbook's logs for more details.) * 09:58 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1336.eqiad.wmnet with OS trixie * 09:47 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts (check the cookbook's logs for more details.) * 09:47 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts (check the cookbook's logs for more details.) * 09:45 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host acmechief-test2001.codfw.wmnet,acmechief-test1001.eqiad.wmnet,an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet,db-test[2001-2002].codfw.wmnet,db-test[1 * 09:40 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host acmechief-test2001.codfw.wmnet,acmechief-test1001.eqiad.wmnet,an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet,db-test[2001-2002].codfw.wmnet,db-test[1001-1003].eqiad.wmn * 09:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1336.eqiad.wmnet with reason: host reimage * 09:33 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1336.eqiad.wmnet with reason: host reimage * 09:29 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet * 09:29 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet * 09:28 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:26 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:21 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 09:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1336 * 09:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1336 * 09:19 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 09:14 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1336 * 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1336.eqiad.wmnet 152.32.64.10.in-addr.arpa 2.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1336.eqiad.wmnet 152.32.64.10.in-addr.arpa 2.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1336 - cgoubert@cumin2003" * 09:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1336 - cgoubert@cumin2003" * 09:11 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:10 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 09:09 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:09 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1336 * 09:09 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1336.eqiad.wmnet with OS trixie * 09:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1336.eqiad.wmnet * 09:08 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1336.eqiad.wmnet * 09:08 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1336.eqiad.wmnet * 09:06 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1335.eqiad.wmnet * 09:06 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1335.eqiad.wmnet * 09:06 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1335.eqiad.wmnet * 09:04 elukey: uploaded spicerack_13.1.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia * 08:55 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wikikube-worker-exp2001.codfw.wmnet * 08:54 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host testreduce1002.eqiad.wmnet * 08:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1335.eqiad.wmnet with OS trixie * 08:51 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host wikikube-worker-exp2001.codfw.wmnet * 08:51 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wikikube-worker-exp1001.eqiad.wmnet * 08:50 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host testreduce1002.eqiad.wmnet * 08:45 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host wikikube-worker-exp1001.eqiad.wmnet * 08:34 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1335.eqiad.wmnet with reason: host reimage * 08:30 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1335.eqiad.wmnet with reason: host reimage * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1335 * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1335 * 08:18 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1335 * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1335.eqiad.wmnet 150.32.64.10.in-addr.arpa 0.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:18 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1335.eqiad.wmnet 150.32.64.10.in-addr.arpa 0.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1335 - cgoubert@cumin2003" * 08:18 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1335 - cgoubert@cumin2003" * 08:14 elukey@cumin1003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 08:14 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges * 08:13 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 08:10 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1335 * 08:10 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1335.eqiad.wmnet with OS trixie * 08:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1335.eqiad.wmnet * 08:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1335.eqiad.wmnet * 08:09 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1335.eqiad.wmnet * 08:06 elukey@cumin1003: END (FAIL) - Cookbook sre.puppet.disable-merges (exit_code=99) * 08:05 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges * 08:03 elukey@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin1003.eqiad.wmnet * 07:57 elukey@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin1003.eqiad.wmnet * 07:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetdb1003.eqiad.wmnet * 07:46 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetdb1003.eqiad.wmnet * 07:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetdb2003.codfw.wmnet * 07:37 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetdb2003.codfw.wmnet * 07:37 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1001.eqiad.wmnet * 07:28 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver1001.eqiad.wmnet * 07:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet * 07:19 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet * 07:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2002.codfw.wmnet * 07:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver2002.codfw.wmnet * 07:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet * 07:05 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet * 07:04 btullis@cumin1003: END (FAIL) - Cookbook sre.hadoop.reboot-workers (exit_code=99) for Hadoop analytics cluster * 07:04 elukey@cumin1003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 07:04 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges * 06:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox1003.eqiad.wmnet * 06:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox1003.eqiad.wmnet * 02:46 ryankemper@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:46 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:44 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:37 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:37 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-internal-scholarly,name=eqiad * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 49s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 01:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore wdqs1025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling source-only afterwards * 01:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore wdqs1027 after Bookworm reimage) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs1027.eqiad.wmnet, repooling both afterwards * 00:55 urbanecm@deploy2003: helmfile [codfw] DONE helmfile.d/services/linkrecommendation: apply * 00:54 urbanecm@deploy2003: helmfile [eqiad] DONE helmfile.d/services/linkrecommendation: apply * 00:54 urbanecm@deploy2003: helmfile [staging] DONE helmfile.d/services/linkrecommendation: apply * 00:54 urbanecm@deploy2003: helmfile [codfw] START helmfile.d/services/linkrecommendation: apply * 00:53 urbanecm@deploy2003: helmfile [staging] START helmfile.d/services/linkrecommendation: apply * 00:52 urbanecm@deploy2003: helmfile [eqiad] START helmfile.d/services/linkrecommendation: apply * 00:23 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore wdqs1025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling source-only afterwards * 00:23 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore wdqs1027 after Bookworm reimage) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs1027.eqiad.wmnet, repooling both afterwards * 00:14 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1274.eqiad.wmnet * 00:14 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1274.eqiad.wmnet * 00:14 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1274.eqiad.wmnet * 00:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1027.eqiad.wmnet with OS bookworm * 00:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1025.eqiad.wmnet with OS bookworm * 00:04 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1274.eqiad.wmnet with OS trixie == 2026-07-16 == * 23:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], xfer to freshly reimaged/scap-deployed wdqs2025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs2025.codfw.wmnet, repooling source-only afterwards * 23:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1027.eqiad.wmnet with reason: host reimage * 23:47 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1025.eqiad.wmnet with reason: host reimage * 23:43 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1274.eqiad.wmnet with reason: host reimage * 23:41 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1025.eqiad.wmnet with reason: host reimage * 23:39 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1027.eqiad.wmnet with reason: host reimage * 23:38 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1274.eqiad.wmnet with reason: host reimage * 23:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1025 * 23:23 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1025 * 23:22 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1027 * 23:22 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1027 * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1274 * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1274 * 23:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1025.eqiad.wmnet with OS bookworm * 23:19 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1274 * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1274.eqiad.wmnet 145.48.64.10.in-addr.arpa 5.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:19 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1274.eqiad.wmnet 145.48.64.10.in-addr.arpa 5.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1274 - swfrench@cumin1003" * 23:19 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1274 - swfrench@cumin1003" * 23:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1027.eqiad.wmnet with OS bookworm * 23:14 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 23:14 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1274 * 23:13 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1274.eqiad.wmnet with OS trixie * 23:13 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1274.eqiad.wmnet * 23:12 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1274.eqiad.wmnet * 23:12 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1274.eqiad.wmnet * 23:12 ryankemper: [[phab:T430880|T430880]] depooled dnsdisc of wdqs-internal-scholarly-eqiad bc we only have 1 host there * 23:09 ryankemper@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-internal-scholarly,name=eqiad * 23:08 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1272.eqiad.wmnet * 23:08 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1272.eqiad.wmnet * 23:08 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1272.eqiad.wmnet * 23:01 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], xfer to freshly reimaged/scap-deployed wdqs2025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs2025.codfw.wmnet, repooling source-only afterwards * 22:57 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1272.eqiad.wmnet with OS trixie * 22:56 Amir1: deleting echo notifications from 2015 in group0 * 22:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2025.codfw.wmnet with OS bookworm * 22:35 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1272.eqiad.wmnet with reason: host reimage * 22:32 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 27s) * 22:32 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 22:28 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1269.eqiad.wmnet * 22:28 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1269.eqiad.wmnet * 22:28 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1269.eqiad.wmnet * 22:27 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1272.eqiad.wmnet with reason: host reimage * 22:26 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] (duration: 08m 51s) * 22:22 ladsgroup@deploy2003: ladsgroup, urbanecm: Continuing with deployment * 22:19 ladsgroup@deploy2003: ladsgroup, urbanecm: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:17 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] * 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2025.codfw.wmnet with reason: host reimage * 22:06 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1272 * 22:06 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1272 * 22:05 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1272 * 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1272.eqiad.wmnet 127.48.64.10.in-addr.arpa 7.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:05 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1272.eqiad.wmnet 127.48.64.10.in-addr.arpa 7.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1272 - swfrench@cumin1003" * 22:05 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1272 - swfrench@cumin1003" * 22:03 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2025.codfw.wmnet with reason: host reimage * 22:01 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 22:00 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1272 * 22:00 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1272.eqiad.wmnet with OS trixie * 22:00 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1272.eqiad.wmnet * 21:59 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1272.eqiad.wmnet * 21:59 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1272.eqiad.wmnet * 21:55 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1271.eqiad.wmnet * 21:55 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1271.eqiad.wmnet * 21:55 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1271.eqiad.wmnet * 21:46 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1271.eqiad.wmnet with OS trixie * 21:45 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] (duration: 06m 31s) * 21:40 sbassett@deploy2003: sbassett: Continuing with deployment * 21:40 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2025 * 21:40 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2025 * 21:40 sbassett@deploy2003: sbassett: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:38 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] * 21:37 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2025 * 21:37 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2025.codfw.wmnet 220.48.192.10.in-addr.arpa 0.2.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:37 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2025.codfw.wmnet 220.48.192.10.in-addr.arpa 0.2.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:37 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:37 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2025 - bking@cumin2003" * 21:37 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2025 - bking@cumin2003" * 21:30 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] (duration: 08m 19s) * 21:26 sbassett@deploy2003: sbassett: Continuing with deployment * 21:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1269.eqiad.wmnet with OS trixie * 21:24 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1271.eqiad.wmnet with reason: host reimage * 21:23 sbassett@deploy2003: sbassett: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:22 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:22 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] * 21:20 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2025 * 21:19 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2025.codfw.wmnet with OS bookworm * 21:17 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1271.eqiad.wmnet with reason: host reimage * 21:04 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1269.eqiad.wmnet with reason: host reimage * 21:00 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1269.eqiad.wmnet with reason: host reimage * 20:56 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1271 * 20:55 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1271 * 20:54 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1271 * 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1271.eqiad.wmnet 126.48.64.10.in-addr.arpa 6.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:54 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1271.eqiad.wmnet 126.48.64.10.in-addr.arpa 6.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1271 - swfrench@cumin1003" * 20:54 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1271 - swfrench@cumin1003" * 20:51 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:51 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:51 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:50 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 20:49 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 20:49 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1268.eqiad.wmnet * 20:49 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1268.eqiad.wmnet * 20:49 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1268.eqiad.wmnet * 20:48 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1271 * 20:48 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1271.eqiad.wmnet with OS trixie * 20:47 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1271.eqiad.wmnet * 20:46 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1271.eqiad.wmnet * 20:46 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1271.eqiad.wmnet * 20:41 aude@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] (duration: 07m 34s) * 20:39 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1269 * 20:39 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1269 * 20:36 aude@deploy2003: aude: Continuing with deployment * 20:35 aude@deploy2003: aude: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:33 aude@deploy2003: Started scap sync-world: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] * 20:26 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on db2207.codfw.wmnet with reason: Host down * 20:22 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-video: apply * 20:21 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-video: apply * 20:20 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-timeline: apply * 20:20 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-timeline: apply * 20:20 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-syntaxhighlight: apply * 20:19 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-syntaxhighlight: apply * 20:19 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-media: apply * 20:18 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-media: apply * 20:18 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-constraints: apply * 20:17 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-constraints: apply * 20:17 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox: apply * 20:16 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox: apply * 20:13 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1269 * 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1269.eqiad.wmnet 80.32.64.10.in-addr.arpa 0.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:13 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1269.eqiad.wmnet 80.32.64.10.in-addr.arpa 0.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1269 - kamila@cumin1003" * 20:13 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1269 - kamila@cumin1003" * 20:09 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-flink-codfw cluster: Roll restart of jvm daemons. * 20:07 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 20:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 20:03 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-flink-codfw cluster: Roll restart of jvm daemons. * 20:03 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2207 [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94893 and previous config saved to /var/cache/conftool/dbconfig/20260716-200257-marostegui.json * 20:01 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2204 to s2 primary [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94892 and previous config saved to /var/cache/conftool/dbconfig/20260716-200157-marostegui.json * 20:00 marostegui: Starting emergency s2 codfw failover from db2207 to db2204 - [[phab:T432396|T432396]] * 19:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1035.eqiad.wmnet * 19:56 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2204 with weight 0 [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94891 and previous config saved to /var/cache/conftool/dbconfig/20260716-195628-marostegui.json * 19:55 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 26 hosts with reason: Primary switchover s2 [[phab:T432396|T432396]] * 19:54 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1035.eqiad.wmnet * 19:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1034.eqiad.wmnet * 19:48 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1034.eqiad.wmnet * 19:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1033.eqiad.wmnet * 19:43 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-video: apply * 19:43 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1033.eqiad.wmnet * 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1032.eqiad.wmnet * 19:42 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-video: apply * 19:42 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-timeline: apply * 19:41 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-timeline: apply * 19:41 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-syntaxhighlight: apply * 19:41 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-syntaxhighlight: apply * 19:40 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-media: apply * 19:40 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-media: apply * 19:39 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-constraints: apply * 19:36 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-constraints: apply * 19:36 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox: apply * 19:35 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1032.eqiad.wmnet * 19:35 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1031.eqiad.wmnet * 19:35 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox: apply * 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-video: apply * 19:33 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-video: apply * 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-timeline: apply * 19:33 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-timeline: apply * 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-syntaxhighlight: apply * 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-syntaxhighlight: apply * 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-media: apply * 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-media: apply * 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-constraints: apply * 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-constraints: apply * 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox: apply * 19:31 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox: apply * 19:27 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1031.eqiad.wmnet * 19:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1030.eqiad.wmnet * 19:23 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1001.eqiad.wmnet, repooling source-only afterwards * 19:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1030.eqiad.wmnet * 19:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1029.eqiad.wmnet * 19:17 kamila@cumin1003: START - Cookbook sre.dns.netbox * 19:12 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1029.eqiad.wmnet * 19:06 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1269 * 19:05 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1269.eqiad.wmnet with OS trixie * 19:03 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1269.eqiad.wmnet * 19:03 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1269.eqiad.wmnet * 19:03 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1269.eqiad.wmnet * 18:55 dancy@deploy2003: Finished scap sync-world: testing [[phab:T428971|T428971]] (duration: 02m 41s) * 18:53 dancy@deploy2003: Started scap sync-world: testing [[phab:T428971|T428971]] * 18:31 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1268.eqiad.wmnet with OS trixie * 18:18 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 18:16 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1267.eqiad.wmnet * 18:16 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1267.eqiad.wmnet * 18:16 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1267.eqiad.wmnet * 18:09 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1268.eqiad.wmnet with reason: host reimage * 18:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1001.eqiad.wmnet, repooling source-only afterwards * 18:06 swfrench@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] (duration: 07m 34s) * 18:06 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 23s) * 18:06 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 18:06 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1268.eqiad.wmnet with reason: host reimage * 18:03 bd808@deploy2003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 18:02 bd808@deploy2003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 18:02 swfrench@deploy2003: jiji, swfrench: Continuing with deployment * 18:02 bd808@deploy2003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 18:02 bd808@deploy2003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 18:01 bd808@deploy2003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 18:01 swfrench@deploy2003: jiji, swfrench: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:01 bd808@deploy2003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:59 swfrench@deploy2003: Started scap sync-world: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] * 17:45 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1268 * 17:45 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1268 * 17:44 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1267.eqiad.wmnet with OS trixie * 17:43 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter2006.codfw.wmnet * 17:39 swfrench@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter2006.codfw.wmnet * 17:35 swfrench@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] (duration: 07m 27s) * 17:34 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1268 * 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1268.eqiad.wmnet 78.32.64.10.in-addr.arpa 8.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:34 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1268.eqiad.wmnet 78.32.64.10.in-addr.arpa 8.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1268 - kamila@cumin1003" * 17:34 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1268 - kamila@cumin1003" * 17:31 swfrench@deploy2003: jiji, swfrench: Continuing with deployment * 17:29 swfrench@deploy2003: jiji, swfrench: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:28 kamila@cumin1003: START - Cookbook sre.dns.netbox * 17:28 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1268 * 17:28 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1268.eqiad.wmnet with OS trixie * 17:27 swfrench@deploy2003: Started scap sync-world: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] * 17:23 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1267.eqiad.wmnet with reason: host reimage * 17:18 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1267.eqiad.wmnet with reason: host reimage * 17:18 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1268.eqiad.wmnet * 17:17 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1268.eqiad.wmnet * 17:17 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1268.eqiad.wmnet * 17:12 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter2005.codfw.wmnet * 17:11 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1270.eqiad.wmnet * 17:11 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1270.eqiad.wmnet * 17:11 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1270.eqiad.wmnet * 17:09 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter2005.codfw.wmnet * 17:08 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:08 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update reverse dns for moved arelion cct cr2-eqiad - cmooney@cumin1003" * 17:08 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update reverse dns for moved arelion cct cr2-eqiad - cmooney@cumin1003" * 17:08 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] (duration: 07m 34s) * 17:04 jiji@deploy2003: jiji: Continuing with deployment * 17:03 jiji@deploy2003: jiji: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 17:00 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] * 17:00 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2185.codfw.wmnet with OS trixie * 16:59 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:58 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1270.eqiad.wmnet with OS trixie * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1267 * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1267 * 16:57 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1267 * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1267.eqiad.wmnet 77.32.64.10.in-addr.arpa 7.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:57 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1267.eqiad.wmnet 77.32.64.10.in-addr.arpa 7.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1267 - kamila@cumin1003" * 16:56 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1267 - kamila@cumin1003" * 16:56 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_eqsin * 16:56 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5032.eqsin.wmnet * 16:52 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_esams * 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3073.esams.wmnet * 16:50 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_esams * 16:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3081.esams.wmnet * 16:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1266.eqiad.wmnet * 16:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1266.eqiad.wmnet * 16:45 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1266.eqiad.wmnet * 16:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2185.codfw.wmnet with reason: host reimage * 16:41 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_eqiad * 16:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1114.eqiad.wmnet * 16:41 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_eqiad * 16:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1115.eqiad.wmnet * 16:39 kamila@cumin1003: START - Cookbook sre.dns.netbox * 16:39 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1267 * 16:39 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2185.codfw.wmnet with reason: host reimage * 16:38 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1267.eqiad.wmnet with OS trixie * 16:38 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1267.eqiad.wmnet * 16:38 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1270.eqiad.wmnet with reason: host reimage * 16:37 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1267.eqiad.wmnet * 16:37 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1267.eqiad.wmnet * 16:31 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1270.eqiad.wmnet with reason: host reimage * 16:24 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1264.eqiad.wmnet * 16:24 kamila@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 16:24 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 16:23 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 16:21 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2162: switch maintenance completed codfw rack b6 * 16:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2185.codfw.wmnet with OS trixie * 16:19 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 16:16 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter1007.eqiad.wmnet * 16:15 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_eqsin * 16:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5024.eqsin.wmnet * 16:13 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5031.eqsin.wmnet * 16:13 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3072.esams.wmnet * 16:12 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter1007.eqiad.wmnet * 16:11 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] (duration: 09m 47s) * 16:10 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1266.eqiad.wmnet with OS trixie * 16:10 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1270 * 16:10 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1270 * 16:09 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1270 * 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1270.eqiad.wmnet 125.48.64.10.in-addr.arpa 5.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1270.eqiad.wmnet 125.48.64.10.in-addr.arpa 5.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1270 - swfrench@cumin1003" * 16:09 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1270 - swfrench@cumin1003" * 16:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3080.esams.wmnet * 16:07 jiji@deploy2003: jiji: Continuing with deployment * 16:06 jiji@deploy2003: jiji: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:04 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 16:04 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1265.eqiad.wmnet * 16:03 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1265.eqiad.wmnet * 16:03 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1265.eqiad.wmnet * 16:03 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1270 * 16:03 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1270.eqiad.wmnet with OS trixie * 16:02 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1270.eqiad.wmnet * 16:02 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] * 16:01 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1270.eqiad.wmnet * 16:01 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1270.eqiad.wmnet * 16:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1113.eqiad.wmnet * 16:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1112.eqiad.wmnet * 15:49 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1266.eqiad.wmnet with reason: host reimage * 15:47 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter1006.eqiad.wmnet * 15:45 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1265.eqiad.wmnet with OS trixie * 15:44 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1266.eqiad.wmnet with reason: host reimage * 15:43 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter1006.eqiad.wmnet * 15:42 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] (duration: 09m 46s) * 15:37 jiji@deploy2003: jiji: Continuing with deployment * 15:36 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2162: switch maintenance completed codfw rack b6 * 15:36 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2161: switch maintenance completed codfw rack b6 * 15:34 jiji@deploy2003: jiji: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:32 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5023.eqsin.wmnet * 15:32 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] * 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5030.eqsin.wmnet * 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3071.esams.wmnet * 15:27 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3079.esams.wmnet * 15:25 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1265.eqiad.wmnet with reason: host reimage * 15:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1266 * 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1266 * 15:21 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1110.eqiad.wmnet * 15:20 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1111.eqiad.wmnet * 15:16 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1265.eqiad.wmnet with reason: host reimage * 15:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1001.eqiad.wmnet with OS bookworm * 15:15 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1266 * 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1266.eqiad.wmnet 76.32.64.10.in-addr.arpa 6.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:15 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1266.eqiad.wmnet 76.32.64.10.in-addr.arpa 6.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1266 - kamila@cumin1003" * 15:15 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1266 - kamila@cumin1003" * 15:07 kamila@cumin1003: START - Cookbook sre.dns.netbox * 15:04 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1266 * 15:04 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1264 * 15:04 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1264 * 15:04 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1266.eqiad.wmnet with OS trixie * 15:03 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1264 * 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1264.eqiad.wmnet 74.32.64.10.in-addr.arpa 4.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:03 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1264.eqiad.wmnet 74.32.64.10.in-addr.arpa 4.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1264 - kamila@cumin1003" * 15:03 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1264 - kamila@cumin1003" * 15:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-eqiad * 15:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp1001.eqiad.wmnet * 15:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp1001.eqiad.wmnet * 15:01 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp1001.eqiad.wmnet * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp1001.eqiad.wmnet * 15:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1376-1384].eqiad.wmnet * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1376-1384].eqiad.wmnet * 14:59 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1002.eqiad.wmnet * 14:59 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1266.eqiad.wmnet * 14:58 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1266.eqiad.wmnet * 14:58 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1266.eqiad.wmnet * 14:58 kamila@cumin1003: START - Cookbook sre.dns.netbox * 14:57 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1264 * 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1265 * 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1265 * 14:57 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1310596{{!}}Set $wgMathInternalRestbaseURL explicitly (T349582)]] * 14:57 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1265 * 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1265.eqiad.wmnet 75.32.64.10.in-addr.arpa 5.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:56 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1265.eqiad.wmnet 75.32.64.10.in-addr.arpa 5.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:56 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:56 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1265 - kamila@cumin1003" * 14:56 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1265 - kamila@cumin1003" * 14:53 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1002.eqiad.wmnet * 14:53 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1376-1384].eqiad.wmnet * 14:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1001.eqiad.wmnet with reason: host reimage * 14:51 kamila@cumin1003: START - Cookbook sre.dns.netbox * 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 14:50 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:50 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2161: switch maintenance completed codfw rack b6 * 14:50 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5021.eqsin.wmnet * 14:50 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1265 * 14:49 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:49 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:49 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1265.eqiad.wmnet with OS trixie * 14:49 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5029.eqsin.wmnet * 14:49 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1265.eqiad.wmnet * 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3070.esams.wmnet * 14:48 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1264.eqiad.wmnet * 14:48 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1001.eqiad.wmnet with reason: host reimage * 14:48 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1376-1384].eqiad.wmnet * 14:48 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1265.eqiad.wmnet * 14:47 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1265.eqiad.wmnet * 14:47 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1264.eqiad.wmnet * 14:47 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1264.eqiad.wmnet * 14:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:47 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3078.esams.wmnet * 14:44 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1263.eqiad.wmnet * 14:44 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1263.eqiad.wmnet * 14:44 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1263.eqiad.wmnet * 14:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1108.eqiad.wmnet * 14:40 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1109.eqiad.wmnet * 14:40 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:35 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:34 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:34 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2006.codfw.wmnet * 14:34 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-flink-eqiad cluster: Roll restart of jvm daemons. * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf2002.codfw.wmnet * 14:31 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1002.eqiad.wmnet * 14:29 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2006.codfw.wmnet * 14:27 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:27 kamila@deploy2003: Finished scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] (duration: 02m 57s) * 14:27 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-flink-eqiad cluster: Roll restart of jvm daemons. * 14:26 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf2002.codfw.wmnet * 14:26 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf2001.codfw.wmnet * 14:25 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1002.eqiad.wmnet * 14:25 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1001.eqiad.wmnet * 14:25 kamila@deploy2003: Started scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] * 14:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:21 kamila@deploy2003: sync-world aborted: Test deployment to check rsync is working - [[phab:T432108|T432108]] (duration: 00m 36s) * 14:21 topranks: reboot lsw1-b6-codfw to upgrade JunOS [[phab:T430922|T430922]] * 14:21 kamila@deploy2003: Started scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] * 14:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1001.eqiad.wmnet with OS bookworm * 14:20 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b6-codfw,lsw1-b6-codfw IPv6,lsw1-b6-codfw.mgmt,ssw1-a[1,8]-codfw with reason: lsw1-b6-codfw JunOS upgrade * 14:20 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf2001.codfw.wmnet * 14:19 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1001.eqiad.wmnet * 14:19 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 26 hosts with reason: lsw1-b6-codfw JunOS upgrade * 14:14 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 14:13 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc2022: switch maintenance codfw rack b6 * 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:12 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.parsercache * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool pc2022: switch maintenance codfw rack b6 * 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2251: switch maintenance codfw rack b6 * 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.parsercache * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2251: switch maintenance codfw rack b6 * 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2162: switch maintenance codfw rack b6 * 14:12 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1263.eqiad.wmnet with OS trixie * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2162: switch maintenance codfw rack b6 * 14:11 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2161: switch maintenance codfw rack b6 * 14:11 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2161: switch maintenance codfw rack b6 * 14:08 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5020.eqsin.wmnet * 14:07 btullis@cumin1003: START - Cookbook sre.hadoop.reboot-workers for Hadoop analytics cluster * 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3069.esams.wmnet * 14:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5028.eqsin.wmnet * 14:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1338-1347].eqiad.wmnet * 14:06 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1338-1347].eqiad.wmnet * 14:05 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3077.esams.wmnet * 14:02 topranks: beginning depools for lsw1-b6-codfw maintenance [[phab:T430922|T430922]] * 14:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1106.eqiad.wmnet * 14:00 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-misc1002.eqiad.wmnet * 13:59 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1338-1347].eqiad.wmnet * 13:59 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1107.eqiad.wmnet * 13:56 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-codfw * 13:55 sfaci@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply * 13:54 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-misc1002.eqiad.wmnet * 13:54 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-misc1001.eqiad.wmnet * 13:54 sfaci@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply * 13:50 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1263.eqiad.wmnet with reason: host reimage * 13:50 sfaci@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 13:49 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1338-1347].eqiad.wmnet * 13:49 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-misc1001.eqiad.wmnet * 13:49 sfaci@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 13:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:49 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:45 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1263.eqiad.wmnet with reason: host reimage * 13:40 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:40 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-eqiad * 13:35 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:34 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:33 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-reboot (exit_code=0) rolling reboot on A:dnsbox and (A:eqsin or A:drmrs or A:magru) and not (P<nowiki>{</nowiki>dns5003*<nowiki>}</nowiki> or P<nowiki>{</nowiki>dns7002*<nowiki>}</nowiki>) and (A:dnsbox) * 13:33 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns7001.wikimedia.org * 13:27 sukhe@dns1004: END - running authdns-update * 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5019.eqsin.wmnet * 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3076.esams.wmnet * 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3068.esams.wmnet * 13:25 sukhe@dns1004: START - running authdns-update * 13:24 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5027.eqsin.wmnet * 13:24 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1263 * 13:24 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1263 * 13:23 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1263 * 13:23 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:23 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1104.eqiad.wmnet * 13:21 kamila@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:21 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:20 kamila@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:20 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:20 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:20 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1263 - kamila@cumin1003" * 13:20 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1263 - kamila@cumin1003" * 13:19 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1105.eqiad.wmnet * 13:19 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:19 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:18 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:18 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns7001.wikimedia.org * 13:16 cdobbins@cumin2003: conftool action : set/pooled=yes; selector: name=dns7002.* * 13:14 cdobbins@dns1004: END - running authdns-update * 13:13 sbisson@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] (duration: 08m 03s) * 13:13 cdobbins@dns1004: START - running authdns-update * 13:12 kamila@cumin1003: START - Cookbook sre.dns.netbox * 13:12 cdobbins@cumin2003: conftool action : set/pooled=yes; selector: name=dns7002.*,service=authdns-update * 13:12 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1263 * 13:11 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1263.eqiad.wmnet with OS trixie * 13:11 cdobbins@cumin2003: conftool action : set/pooled=no; selector: name=dns7002.* * 13:11 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1263.eqiad.wmnet * 13:10 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1263.eqiad.wmnet * 13:10 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1263.eqiad.wmnet * 13:09 sbisson@deploy2003: sbisson: Continuing with deployment * 13:08 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:07 sbisson@deploy2003: sbisson: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:05 sbisson@deploy2003: Started scap sync-world: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] * 13:03 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns6002.wikimedia.org * 13:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1298-1307].eqiad.wmnet * 13:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1298-1307].eqiad.wmnet * 12:59 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1262.eqiad.wmnet * 12:59 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1262.eqiad.wmnet * 12:59 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1262.eqiad.wmnet * 12:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1298-1307].eqiad.wmnet * 12:49 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns6002.wikimedia.org * 12:46 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1298-1307].eqiad.wmnet * 12:46 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:46 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3075.esams.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3067.esams.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5018.eqsin.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5026.eqsin.wmnet * 12:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1102.eqiad.wmnet * 12:39 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1103.eqiad.wmnet * 12:35 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:34 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns6001.wikimedia.org * 12:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:28 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:18 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns6001.wikimedia.org * 12:14 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1267-1276].eqiad.wmnet * 12:13 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1267-1276].eqiad.wmnet * 12:04 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1267-1276].eqiad.wmnet * 12:03 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns5004.wikimedia.org * 12:02 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1100.eqiad.wmnet * 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3066.esams.wmnet * 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3074.esams.wmnet * 12:01 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:01 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-codfw * 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5017.eqsin.wmnet * 12:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5025.eqsin.wmnet * 12:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1101.eqiad.wmnet * 11:59 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1267-1276].eqiad.wmnet * 11:58 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:58 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:54 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns5004.wikimedia.org * 11:54 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and (A:eqsin or A:drmrs or A:magru) and not (P<nowiki>{</nowiki>dns5003*<nowiki>}</nowiki> or P<nowiki>{</nowiki>dns7002*<nowiki>}</nowiki>) and (A:dnsbox) * 11:54 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:53 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:53 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-eqiad * 11:51 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:51 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:50 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:50 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_eqiad * 11:50 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:50 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_eqiad * 11:49 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_esams * 11:49 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_esams * 11:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_eqsin * 11:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_eqsin * 11:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:44 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-eqiad * 11:43 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-codfw * 11:42 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:41 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:24 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-eqiad * 11:23 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-codfw * 11:22 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:15 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:14 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:09 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2066.codfw.wmnet with OS trixie * 11:05 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1151-1160].eqiad.wmnet * 11:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1151-1160].eqiad.wmnet * 10:59 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1068.eqiad.wmnet with OS trixie * 10:55 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.major-upgrade (exit_code=99) * 10:55 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 10:54 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1151-1160].eqiad.wmnet * 10:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2066.codfw.wmnet with reason: host reimage * 10:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1151-1160].eqiad.wmnet * 10:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:42 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2066.codfw.wmnet with reason: host reimage * 10:39 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:37 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:36 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:23 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:22 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2066.codfw.wmnet with OS trixie * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:07 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2065.codfw.wmnet with OS trixie * 10:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:06 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 10:06 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 10:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 10:03 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:03 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 10:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 09:59 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 09:57 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 09:57 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 09:52 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 09:47 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:46 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:46 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2065.codfw.wmnet with reason: host reimage * 09:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2065.codfw.wmnet with reason: host reimage * 09:40 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:39 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 09:39 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:39 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:39 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:37 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:29 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox2003.codfw.wmnet * 09:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox2003.codfw.wmnet * 09:25 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:25 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:24 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:24 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:24 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:21 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:20 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2065.codfw.wmnet with OS trixie * 09:13 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1068.eqiad.wmnet with OS trixie * 09:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2064.codfw.wmnet with OS trixie * 09:08 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 09:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 09:07 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2162: Repooling after switchover * 09:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1067.eqiad.wmnet with OS trixie * 09:00 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:59 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 08:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 08:57 tappof: bump space for prometheus k8s-dse in eqiad * 08:56 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping2004.codfw.wmnet * 08:52 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host ping2004.codfw.wmnet * 08:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 08:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping1004.eqiad.wmnet * 08:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:51 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:49 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 08:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host ping1004.eqiad.wmnet * 08:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2064.codfw.wmnet with reason: host reimage * 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2064.codfw.wmnet with reason: host reimage * 08:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1067.eqiad.wmnet with reason: host reimage * 08:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:33 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1067.eqiad.wmnet with reason: host reimage * 08:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:21 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2162: Repooling after switchover * 08:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2064.codfw.wmnet with OS trixie * 08:16 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1067.eqiad.wmnet with OS trixie * 08:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:15 cgoubert@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-eqiad * 08:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2062.codfw.wmnet with OS trixie * 08:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1066.eqiad.wmnet with OS trixie * 08:02 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2162: Repooling after switchover * 07:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2162: Repooling after switchover * 07:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2162 [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94870 and previous config saved to /var/cache/conftool/dbconfig/20260716-075530-cwilliams.json * 07:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2241 to x3 primary [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94869 and previous config saved to /var/cache/conftool/dbconfig/20260716-075314-cwilliams.json * 07:52 cezmunsta: Starting x3 codfw failover from db2162 to db2241 - [[phab:T430925|T430925]] * 07:50 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:50 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:47 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 07:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2241 with weight 0 [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94868 and previous config saved to /var/cache/conftool/dbconfig/20260716-074507-cwilliams.json * 07:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 18 hosts with reason: Primary switchover x3 [[phab:T430925|T430925]] * 07:43 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1066.eqiad.wmnet with reason: host reimage * 07:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:dse-k8s-worker-eqiad * 07:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1028.eqiad.wmnet * 07:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1028.eqiad.wmnet * 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 07:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1066.eqiad.wmnet with reason: host reimage * 07:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1028.eqiad.wmnet * 07:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1028.eqiad.wmnet * 07:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1027.eqiad.wmnet * 07:35 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1027.eqiad.wmnet * 07:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1027.eqiad.wmnet * 07:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1027.eqiad.wmnet * 07:28 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1026.eqiad.wmnet * 07:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1026.eqiad.wmnet * 07:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast2003.wikimedia.org * 07:21 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1026.eqiad.wmnet * 07:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1066.eqiad.wmnet with OS trixie * 07:19 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast2003.wikimedia.org * 07:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2062.codfw.wmnet with OS trixie * 06:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1026.eqiad.wmnet * 06:51 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1025.eqiad.wmnet * 06:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1025.eqiad.wmnet * 06:47 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 06:47 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 06:44 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1025.eqiad.wmnet * 06:14 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1025.eqiad.wmnet * 06:14 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1024.eqiad.wmnet * 06:14 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1024.eqiad.wmnet * 06:07 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1024.eqiad.wmnet * 05:37 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1024.eqiad.wmnet * 05:37 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1023.eqiad.wmnet * 05:37 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1023.eqiad.wmnet * 05:26 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1023.eqiad.wmnet * 04:56 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1023.eqiad.wmnet * 04:56 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1022.eqiad.wmnet * 04:56 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1022.eqiad.wmnet * 04:49 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1022.eqiad.wmnet * 04:19 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1022.eqiad.wmnet * 04:19 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1021.eqiad.wmnet * 04:19 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1021.eqiad.wmnet * 04:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1021.eqiad.wmnet * 03:38 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1021.eqiad.wmnet * 03:38 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1020.eqiad.wmnet * 03:38 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1020.eqiad.wmnet * 03:20 btullis@cumin1003: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1020.eqiad.wmnet * 03:18 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1020.eqiad.wmnet * 03:18 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1019.eqiad.wmnet * 03:18 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1019.eqiad.wmnet * 03:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1019.eqiad.wmnet * 02:41 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1019.eqiad.wmnet * 02:41 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1018.eqiad.wmnet * 02:41 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1018.eqiad.wmnet * 02:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling both afterwards * 02:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2003.codfw.wmnet -> wcqs2001.codfw.wmnet, repooling both afterwards * 02:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1018.eqiad.wmnet * 02:30 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1018.eqiad.wmnet * 02:30 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1014.eqiad.wmnet * 02:30 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1014.eqiad.wmnet * 02:24 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1014.eqiad.wmnet * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 01:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1014.eqiad.wmnet * 01:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1013.eqiad.wmnet * 01:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1013.eqiad.wmnet * 01:47 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1013.eqiad.wmnet * 01:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2003.codfw.wmnet -> wcqs2001.codfw.wmnet, repooling both afterwards * 01:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling both afterwards * 01:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1013.eqiad.wmnet * 01:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1012.eqiad.wmnet * 01:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1012.eqiad.wmnet * 01:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1012.eqiad.wmnet * 01:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1012.eqiad.wmnet * 01:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1011.eqiad.wmnet * 01:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1011.eqiad.wmnet * 01:08 ryankemper@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] scap deploy post bookworm reimage (duration: 00m 23s) * 01:08 ryankemper@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] scap deploy post bookworm reimage * 01:08 ryankemper@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): scap deploy post bookworm reimage (duration: 00m 46s) * 01:07 ryankemper@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): scap deploy post bookworm reimage * 01:04 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1011.eqiad.wmnet * 01:04 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1011.eqiad.wmnet * 01:04 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1010.eqiad.wmnet * 01:04 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1010.eqiad.wmnet * 00:57 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1010.eqiad.wmnet * 00:57 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1010.eqiad.wmnet * 00:57 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1009.eqiad.wmnet * 00:57 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1009.eqiad.wmnet * 00:50 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1009.eqiad.wmnet * 00:20 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1009.eqiad.wmnet * 00:20 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1008.eqiad.wmnet * 00:20 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1008.eqiad.wmnet * 00:13 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1008.eqiad.wmnet == 2026-07-15 == * 23:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2001.codfw.wmnet with OS bookworm * 23:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1008.eqiad.wmnet * 23:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1007.eqiad.wmnet * 23:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1007.eqiad.wmnet * 23:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1007.eqiad.wmnet * 23:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1007.eqiad.wmnet * 23:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1006.eqiad.wmnet * 23:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1006.eqiad.wmnet * 23:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1006.eqiad.wmnet * 23:29 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1006.eqiad.wmnet * 23:28 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1005.eqiad.wmnet * 23:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1005.eqiad.wmnet * 23:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1002.eqiad.wmnet with OS bookworm * 23:21 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1005.eqiad.wmnet * 23:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2001.codfw.wmnet with reason: host reimage * 23:15 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host datahubsearch1001.eqiad.wmnet with OS bookworm * 23:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2001.codfw.wmnet with reason: host reimage * 23:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 23:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 22:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 22:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1005.eqiad.wmnet * 22:51 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1004.eqiad.wmnet * 22:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1004.eqiad.wmnet * 22:45 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1004.eqiad.wmnet * 22:44 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 22:44 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS trixie * 22:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host datahubsearch1001.eqiad.wmnet with OS bookworm * 22:34 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host datahubsearch1001.eqiad.wmnet with OS bookworm * 22:16 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on datahubsearch[1002-1003].eqiad.wmnet with reason: Using datahubsearch1001 to test bookworm reimages * 22:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1004.eqiad.wmnet * 22:15 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1003.eqiad.wmnet * 22:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1003.eqiad.wmnet * 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 22:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1003.eqiad.wmnet * 22:08 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1003.eqiad.wmnet * 22:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1002.eqiad.wmnet * 22:08 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1002.eqiad.wmnet * 22:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host datahubsearch1001.eqiad.wmnet with OS bookworm * 22:05 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 22:02 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm * 22:01 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on datahubsearch[1001-1003].eqiad.wmnet with reason: Using datahubsearch1001 to test bookworm reimages * 22:01 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1002.eqiad.wmnet * 22:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1002.eqiad.wmnet * 22:00 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1001.eqiad.wmnet * 22:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1001.eqiad.wmnet * 21:53 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1001.eqiad.wmnet * 21:52 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 21:50 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 21:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS trixie * 21:50 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS bookworm * 21:43 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 21:38 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:30 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wcqs1002'] * 21:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:29 lerickson@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 21:29 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:29 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:29 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS bookworm * 21:28 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 21:28 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm * 21:23 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1001.eqiad.wmnet * 21:23 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:23 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:22 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 21:20 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:18 swfrench-wmf: reprepro include php8.3_8.3.32-1+wmf11u2 into component/php83 for bullseye-wikimedia * 21:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:16 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:15 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-druid-public cluster: Roll restart of jvm daemons. * 21:08 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:05 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1001.eqiad.wmnet * 21:05 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1001.eqiad.wmnet * 21:04 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-druid-public cluster: Roll restart of jvm daemons. * 21:02 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 21:01 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 21:01 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 21:00 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 20:59 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1001.eqiad.wmnet * 20:59 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1001.eqiad.wmnet * 20:59 btullis@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:dse-k8s-worker-eqiad * 20:55 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 20:55 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 20:45 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm * 20:21 jhathaway: puppet is re-enabled, have fun, but not too much fun! * 20:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 20:17 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs2001'] * 20:12 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs2001'] * 20:11 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs2001'] * 20:09 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:08 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 20:05 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:05 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 20:04 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs2001'] * 20:03 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 20:03 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm * 20:02 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:02 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 20:01 jhathaway: disabling puppet fleet wide to roll out kafka patch * 19:55 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host relforge1010.eqiad.wmnet * 19:52 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 19:52 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 19:48 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 19:48 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 19:48 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 19:47 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 19:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 19:45 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 19:45 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 19:44 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1010.eqiad.wmnet * 19:38 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1262.eqiad.wmnet with OS trixie * 19:17 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1262.eqiad.wmnet with reason: host reimage * 19:11 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1262.eqiad.wmnet with reason: host reimage * 18:59 cdobbins@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS trixie * 18:54 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 18:53 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 18:52 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1262 * 18:52 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1262 * 18:51 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1262 * 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1262.eqiad.wmnet 72.32.64.10.in-addr.arpa 2.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:51 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1262.eqiad.wmnet 72.32.64.10.in-addr.arpa 2.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1262 - kamila@cumin1003" * 18:51 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1262 - kamila@cumin1003" * 18:46 kamila@cumin1003: START - Cookbook sre.dns.netbox * 18:46 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1262 * 18:46 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 18:46 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ncmonitor1001.eqiad.wmnet * 18:46 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 18:45 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1262.eqiad.wmnet with OS trixie * 18:45 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 18:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1262.eqiad.wmnet * 18:44 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1262.eqiad.wmnet * 18:44 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1262.eqiad.wmnet * 18:42 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host ncmonitor1001.eqiad.wmnet * 18:29 topranks: pull power on cr1-eqiad to install new switch-control boards [[phab:T426343|T426343]] * 18:29 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs[1018-1020].eqiad.wmnet with reason: line card install in cr1-eqiad * 18:27 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 14 hosts with reason: linecard install in cr1-eqad * 18:22 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_ulsfo * 18:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4052.ulsfo.wmnet * 18:19 cdobbins@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 18:15 cdobbins@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 18:14 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_drmrs * 18:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6016.drmrs.wmnet * 18:12 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_ulsfo * 18:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4044.ulsfo.wmnet * 18:10 sukhe@cumin1003: END (ERROR) - Cookbook sre.cdn.roll-reboot (exit_code=97) rolling reboot on A:cp-upload_drmrs * 18:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2241: Security update * 17:56 topranks: start draining traffic on cr1-eqiad ahead of line card installation [[phab:T426343|T426343]] * 17:47 cdobbins@cumin2003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie * 17:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4051.ulsfo.wmnet * 17:40 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:39 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 17:34 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6007.drmrs.wmnet * 17:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6015.drmrs.wmnet * 17:32 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:31 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 17:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4043.ulsfo.wmnet * 17:27 lerickson@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:25 lerickson@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 17:22 lerickson@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-codfw * 17:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp2001.codfw.wmnet * 17:22 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 17:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp2001.codfw.wmnet * 17:22 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 17:19 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2241: Security update * 17:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2241.codfw.wmnet * 17:17 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2241.codfw.wmnet * 17:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp2001.codfw.wmnet * 17:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp2001.codfw.wmnet * 17:15 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2366-2374].codfw.wmnet * 17:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2366-2374].codfw.wmnet * 17:10 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply * 17:10 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply * 17:08 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2366-2374].codfw.wmnet * 17:06 sukhe: sre.dns.roll-reboot to resume later * 17:06 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-reboot (exit_code=97) rolling reboot on A:dnsbox and not (A:ulsfo or A:magru) and (A:dnsbox) * 17:06 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns5003.wikimedia.org * 17:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2241: Security update * 17:03 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2241: Security update * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply * 17:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2366-2374].codfw.wmnet * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply * 17:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2357-2365].codfw.wmnet * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 17:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2357-2365].codfw.wmnet * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply * 16:55 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2357-2365].codfw.wmnet * 16:55 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 16:53 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6006.drmrs.wmnet * 16:52 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6014.drmrs.wmnet * 16:52 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 16:51 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 16:50 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2357-2365].codfw.wmnet * 16:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4042.ulsfo.wmnet * 16:50 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2347-2356].codfw.wmnet * 16:50 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2347-2356].codfw.wmnet * 16:49 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns5003.wikimedia.org * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply * 16:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4050.ulsfo.wmnet * 16:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2347-2356].codfw.wmnet * 16:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2347-2356].codfw.wmnet * 16:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2337-2346].codfw.wmnet * 16:36 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2337-2346].codfw.wmnet * 16:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:dse-k8s-worker-codfw * 16:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2003.codfw.wmnet * 16:35 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2003.codfw.wmnet * 16:34 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns3004.wikimedia.org * 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply * 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply * 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply * 16:30 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply * 16:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2003.codfw.wmnet * 16:29 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2337-2346].codfw.wmnet * 16:24 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2003.codfw.wmnet * 16:24 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2002.codfw.wmnet * 16:24 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2002.codfw.wmnet * 16:23 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns3004.wikimedia.org * 16:23 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2337-2346].codfw.wmnet * 16:23 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2327-2336].codfw.wmnet * 16:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2327-2336].codfw.wmnet * 16:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2002.codfw.wmnet * 16:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2327-2336].codfw.wmnet * 16:12 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2002.codfw.wmnet * 16:12 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2001.codfw.wmnet * 16:12 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2001.codfw.wmnet * 16:12 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1065.eqiad.wmnet with OS trixie * 16:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6005.drmrs.wmnet * 16:11 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6013.drmrs.wmnet * 16:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4041.ulsfo.wmnet * 16:08 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns3003.wikimedia.org * 16:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2327-2336].codfw.wmnet * 16:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2317-2326].codfw.wmnet * 16:06 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2317-2326].codfw.wmnet * 16:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2001.codfw.wmnet * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply * 16:03 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4049.ulsfo.wmnet * 16:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2001.codfw.wmnet * 16:00 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test2001.codfw.wmnet * 16:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test2001.codfw.wmnet * 16:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2063.codfw.wmnet with OS trixie * 15:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2317-2326].codfw.wmnet * 15:57 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns3003.wikimedia.org * 15:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test2001.codfw.wmnet * 15:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test2001.codfw.wmnet * 15:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2004.codfw.wmnet * 15:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2004.codfw.wmnet * 15:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2317-2326].codfw.wmnet * 15:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2307-2316].codfw.wmnet * 15:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2307-2316].codfw.wmnet * 15:49 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2004.codfw.wmnet * 15:48 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2004.codfw.wmnet * 15:48 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2003.codfw.wmnet * 15:48 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2003.codfw.wmnet * 15:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 15:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2307-2316].codfw.wmnet * 15:42 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2003.codfw.wmnet * 15:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 15:42 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2003.codfw.wmnet * 15:42 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2002.codfw.wmnet * 15:42 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2002.codfw.wmnet * 15:42 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2006.wikimedia.org * 15:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2063.codfw.wmnet with reason: host reimage * 15:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2307-2316].codfw.wmnet * 15:37 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2297-2306].codfw.wmnet * 15:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2297-2306].codfw.wmnet * 15:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2002.codfw.wmnet * 15:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2002.codfw.wmnet * 15:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2001.codfw.wmnet * 15:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2001.codfw.wmnet * 15:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2063.codfw.wmnet with reason: host reimage * 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6004.drmrs.wmnet * 15:31 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2001.codfw.wmnet * 15:31 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2001.codfw.wmnet * 15:31 btullis@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:dse-k8s-worker-codfw * 15:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6012.drmrs.wmnet * 15:28 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2006.wikimedia.org * 15:27 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-analytics cluster: Roll restart of jvm daemons. * 15:27 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2297-2306].codfw.wmnet * 15:27 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4040.ulsfo.wmnet * 15:24 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1065.eqiad.wmnet with OS trixie * 15:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4048.ulsfo.wmnet * 15:21 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-analytics cluster: Roll restart of jvm daemons. * 15:21 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2297-2306].codfw.wmnet * 15:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2287-2296].codfw.wmnet * 15:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2287-2296].codfw.wmnet * 15:20 btullis@cumin1003: END (PASS) - Cookbook sre.druid.reboot-workers (exit_code=0) for Druid public cluster: Reboot Druid nodes * 15:18 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 15:17 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm * 15:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2063.codfw.wmnet with OS trixie * 15:13 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2005.wikimedia.org * 15:11 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2287-2296].codfw.wmnet * 15:11 btullis@cumin1003: END (PASS) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=0) rolling reboot on A:cephosd-eqiad * 15:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1064.eqiad.wmnet with OS trixie * 15:05 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2062.codfw.wmnet with OS trixie * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2287-2296].codfw.wmnet * 15:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2277-2286].codfw.wmnet * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2277-2286].codfw.wmnet * 14:59 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2005.wikimedia.org * 14:57 brouberol@cumin1003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-jumbo-eqiad * 14:52 btullis@cumin1003: END (PASS) - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas (exit_code=0) rolling reboot on A:schema-codfw * 14:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6003.drmrs.wmnet * 14:50 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:50 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host relforge1009.eqiad.wmnet * 14:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2277-2286].codfw.wmnet * 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6011.drmrs.wmnet * 14:47 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 14:46 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>ml-serve1001.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 14:46 klausman@cumin1003: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) pool for host ml-serve1001.eqiad.wmnet * 14:46 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 14:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1001.eqiad.wmnet * 14:45 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4039.ulsfo.wmnet * 14:44 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1009.eqiad.wmnet * 14:44 btullis@cumin1003: START - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas rolling reboot on A:schema-codfw * 14:44 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2004.wikimedia.org * 14:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2277-2286].codfw.wmnet * 14:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2267-2276].codfw.wmnet * 14:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2267-2276].codfw.wmnet * 14:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 14:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4047.ulsfo.wmnet * 14:40 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1001.eqiad.wmnet * 14:38 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 14:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 14:36 topranks: disconnect power on cr2-eqiad to shut down device for switch fabric replacement [[phab:T426343|T426343]] * 14:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2267-2276].codfw.wmnet * 14:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 14:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1001.eqiad.wmnet * 14:35 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>ml-serve1001.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 14:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 14:34 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 14:33 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:33 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:30 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2004.wikimedia.org * 14:29 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2267-2276].codfw.wmnet * 14:29 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2257-2266].codfw.wmnet * 14:29 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2257-2266].codfw.wmnet * 14:24 btullis@cumin1003: END (PASS) - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas (exit_code=0) rolling reboot on A:schema-eqiad * 14:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2257-2266].codfw.wmnet * 14:20 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:20 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:19 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:17 jforrester@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2257-2266].codfw.wmnet * 14:16 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:16 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2062.codfw.wmnet with OS trixie * 14:15 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1064.eqiad.wmnet with OS trixie * 14:15 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1006.wikimedia.org * 14:15 btullis@cumin1003: START - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas rolling reboot on A:schema-eqiad * 14:14 topranks: switch routing-engine on cr2-eqiad resetting all interfaces [[phab:T417873|T417873]] * 14:11 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:11 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:10 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm * 14:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6002.drmrs.wmnet * 14:09 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6010.drmrs.wmnet * 14:06 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1006.wikimedia.org * 14:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:05 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 14:05 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4038.ulsfo.wmnet * 14:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4046.ulsfo.wmnet * 14:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:00 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on cr1-eqiad with reason: switch upgrade and line card install * 14:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:59 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:57 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:57 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:55 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-eqiad * 13:55 btullis@cumin1003: START - Cookbook sre.druid.reboot-workers for Druid public cluster: Reboot Druid nodes * 13:53 topranks: switch routing-engine on cr2-eqiad resetting all interfaces [[phab:T417873|T417873]] * 13:51 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1005.wikimedia.org * 13:50 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:49 brouberol@cumin1003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-test-eqiad * 13:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:44 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2197-2206].codfw.wmnet * 13:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2197-2206].codfw.wmnet * 13:36 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1005.wikimedia.org * 13:35 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2197-2206].codfw.wmnet * 13:30 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2197-2206].codfw.wmnet * 13:28 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6001.drmrs.wmnet * 13:28 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6009.drmrs.wmnet * 13:28 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2187-2196].codfw.wmnet * 13:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2187-2196].codfw.wmnet * 13:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 13:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4037.ulsfo.wmnet * 13:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2001 * 13:22 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2001 * 13:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4045.ulsfo.wmnet * 13:21 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1004.wikimedia.org * 13:19 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on lvs[1018-1020].eqiad.wmnet with reason: switch upgrade and line card install * 13:18 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2009.codfw.wmnet * 13:18 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2001 * 13:18 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2001.codfw.wmnet 26.16.192.10.in-addr.arpa 6.2.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:17 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2001.codfw.wmnet 26.16.192.10.in-addr.arpa 6.2.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:17 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:17 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2001 - bking@cumin2003" * 13:17 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2001 - bking@cumin2003" * 13:17 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2009.codfw.wmnet * 13:17 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_drmrs * 13:17 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2187-2196].codfw.wmnet * 13:17 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_drmrs * 13:17 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on 15 hosts with reason: switch upgrade and line card install * 13:17 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:15 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 13:13 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:13 brouberol@cumin1003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-jumbo-eqiad * 13:13 brouberol@cumin1003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-test-eqiad * 13:13 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1004.wikimedia.org * 13:13 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and not (A:ulsfo or A:magru) and (A:dnsbox) * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:12 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_ulsfo * 13:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:12 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_ulsfo * 13:11 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2187-2196].codfw.wmnet * 13:11 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 13:11 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 13:06 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 13:05 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling source-only afterwards * 13:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2001 * 13:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2009.codfw.wmnet with OS trixie * 13:03 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:03 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling source-only afterwards * 13:01 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 15s) * 13:01 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 13:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 12:57 btullis@cumin1003: END (PASS) - Cookbook sre.druid.reboot-workers (exit_code=0) for Druid analytics cluster: Reboot Druid nodes * 12:54 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 12:54 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2163-2172].codfw.wmnet * 12:54 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2163-2172].codfw.wmnet * 12:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2163-2172].codfw.wmnet * 12:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2009.codfw.wmnet with reason: host reimage * 12:41 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2163-2172].codfw.wmnet * 12:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2153-2162].codfw.wmnet * 12:40 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2153-2162].codfw.wmnet * 12:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2009.codfw.wmnet with reason: host reimage * 12:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2153-2162].codfw.wmnet * 12:29 btullis@cumin1003: END (PASS) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=0) rolling reboot on A:cephosd-codfw * 12:25 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2153-2162].codfw.wmnet * 12:25 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2143-2152].codfw.wmnet * 12:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2143-2152].codfw.wmnet * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2009 * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2009 * 12:22 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2009 * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2009.codfw.wmnet 139.0.192.10.in-addr.arpa 9.3.1.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:22 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2009.codfw.wmnet 139.0.192.10.in-addr.arpa 9.3.1.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2009 - mvernon@cumin2003" * 12:22 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2009 - mvernon@cumin2003" * 12:16 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 12:15 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 12:15 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 12:15 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2009 * 12:15 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 12:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2009.codfw.wmnet with OS trixie * 12:15 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 12:14 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 12:14 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2143-2152].codfw.wmnet * 12:13 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 12:12 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2010.codfw.wmnet * 12:11 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2010.codfw.wmnet * 12:10 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 12:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2143-2152].codfw.wmnet * 12:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2133-2142].codfw.wmnet * 12:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2133-2142].codfw.wmnet * 12:02 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 11:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2133-2142].codfw.wmnet * 11:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2133-2142].codfw.wmnet * 11:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:49 mvolz@deploy2003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:49 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-codfw * 11:48 mvolz@deploy2003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:47 btullis@cumin1003: START - Cookbook sre.druid.reboot-workers for Druid analytics cluster: Reboot Druid nodes * 11:46 mvolz@deploy2003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:46 mvolz@deploy2003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:45 mvolz@deploy2003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:44 mvolz@deploy2003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:40 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] (duration: 11m 38s) * 11:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2010.codfw.wmnet with OS trixie * 11:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2105-2114].codfw.wmnet * 11:36 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2105-2114].codfw.wmnet * 11:36 krinkle@deploy2003: physikerwelt, krinkle: Continuing with deployment * 11:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1018: Security updates * 11:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:36 root@cumin1003: START - Cookbook sre.mysql.parsercache * 11:36 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1018: Security updates * 11:31 krinkle@deploy2003: physikerwelt, krinkle: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:29 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] * 11:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2105-2114].codfw.wmnet * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2105-2114].codfw.wmnet * 11:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2010.codfw.wmnet with reason: host reimage * 11:12 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2010.codfw.wmnet with reason: host reimage * 11:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1018: Security updates * 11:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:10 root@cumin1003: START - Cookbook sre.mysql.parsercache * 11:10 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1018: Security updates * 11:09 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1009.eqiad.wmnet with OS trixie * 11:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow7002.magru.wmnet * 11:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 11:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 11:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=tegola-vector-tiles,name=eqiad * 11:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=kartotherian,name=eqiad * 11:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow7002.magru.wmnet * 10:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2010 * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2010 * 10:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 10:54 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 10:54 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2010 * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2010.codfw.wmnet 76.16.192.10.in-addr.arpa 6.7.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:54 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2010.codfw.wmnet 76.16.192.10.in-addr.arpa 6.7.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2010 - mvernon@cumin2003" * 10:54 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2010 - mvernon@cumin2003" * 10:49 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 10:49 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2010 * 10:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1009.eqiad.wmnet with reason: host reimage * 10:49 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2010.codfw.wmnet with OS trixie * 10:46 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2011.codfw.wmnet * 10:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1011.eqiad.wmnet * 10:44 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2011.codfw.wmnet * 10:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 10:44 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 10:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1009.eqiad.wmnet with reason: host reimage * 10:44 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow6001.drmrs.wmnet * 10:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1017: Security updates * 10:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:39 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1017: Security updates * 10:39 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow6001.drmrs.wmnet * 10:38 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1011.eqiad.wmnet * 10:35 cgoubert@deploy2003: Finished deploy [restbase/deploy@06301bd]: Deploying {{Gerrit|1306088}} {{Gerrit|1308347}} - [[phab:T429944|T429944]] [[phab:T428279|T428279]] (duration: 28m 34s) * 10:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1012.eqiad.wmnet * 10:35 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 10:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow5003.eqsin.wmnet * 10:34 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2011.codfw.wmnet with OS trixie * 10:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1009.eqiad.wmnet with OS trixie * 10:28 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1012.eqiad.wmnet * 10:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1013.eqiad.wmnet * 10:27 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:27 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow5003.eqsin.wmnet * 10:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:26 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow4003.ulsfo.wmnet * 10:25 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 10:25 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 10:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow4003.ulsfo.wmnet * 10:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1013.eqiad.wmnet * 10:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1014.eqiad.wmnet * 10:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:15 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2011.codfw.wmnet with reason: host reimage * 10:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1017: Security updates * 10:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:14 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:14 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1017: Security updates * 10:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow3004.esams.wmnet * 10:11 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2011.codfw.wmnet with reason: host reimage * 10:10 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:10 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1014.eqiad.wmnet * 10:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki2003.codfw.wmnet * 10:09 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 10:09 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 10:09 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow3004.esams.wmnet * 10:08 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2004.codfw.wmnet * 10:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1010.eqiad.wmnet with OS trixie * 10:07 cgoubert@deploy2003: Started deploy [restbase/deploy@06301bd]: Deploying {{Gerrit|1306088}} {{Gerrit|1308347}} - [[phab:T429944|T429944]] [[phab:T428279|T428279]] * 10:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host rpki2003.codfw.wmnet * 10:04 topranks: push out config change to BGP_outfilter on core routers [[phab:T431849|T431849]] * 10:02 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow2004.codfw.wmnet * 09:59 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 09:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2003.codfw.wmnet * 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2011 * 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2011 * 09:53 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 09:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:52 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2011 * 09:52 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2011.codfw.wmnet 36.32.192.10.in-addr.arpa 6.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:52 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2011.codfw.wmnet 36.32.192.10.in-addr.arpa 6.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:51 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:51 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2011 - mvernon@cumin2003" * 09:51 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2011 - mvernon@cumin2003" * 09:51 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow2003.codfw.wmnet * 09:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1003.eqiad.wmnet * 09:49 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 09:49 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 09:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1010.eqiad.wmnet with reason: host reimage * 09:47 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 09:47 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 09:47 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 09:47 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2011 * 09:46 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2011.codfw.wmnet with OS trixie * 09:44 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow1003.eqiad.wmnet * 09:44 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2012.codfw.wmnet * 09:44 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1002.eqiad.wmnet * 09:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1010.eqiad.wmnet with reason: host reimage * 09:43 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2012.codfw.wmnet * 09:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Security updates * 09:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:43 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:43 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Security updates * 09:42 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:40 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow1002.eqiad.wmnet * 09:40 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 09:37 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki1001.eqiad.wmnet * 09:36 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 09:36 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 09:33 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host rpki1001.eqiad.wmnet * 09:32 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:32 cgoubert@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-codfw * 09:31 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=kartotherian,name=eqiad * 09:31 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola-vector-tiles,name=eqiad * 09:31 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 09:31 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2012.codfw.wmnet with OS trixie * 09:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1010.eqiad.wmnet with OS trixie * 09:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1011.eqiad.wmnet with OS trixie * 09:21 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Security updates * 09:21 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:21 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:21 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Security updates * 09:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2012.codfw.wmnet with reason: host reimage * 09:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1011.eqiad.wmnet with reason: host reimage * 09:08 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2012.codfw.wmnet with reason: host reimage * 09:05 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1011.eqiad.wmnet with reason: host reimage * 08:55 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:52 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1011.eqiad.wmnet with OS trixie * 08:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1022: Security updates * 08:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2012 * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2012 * 08:50 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1022: Security updates * 08:50 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2012 * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2012.codfw.wmnet 44.48.192.10.in-addr.arpa 4.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:50 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2012.codfw.wmnet 44.48.192.10.in-addr.arpa 4.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2012 - mvernon@cumin2003" * 08:50 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2012 - mvernon@cumin2003" * 08:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1012.eqiad.wmnet with OS trixie * 08:44 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 08:44 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2012 * 08:43 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2012.codfw.wmnet with OS trixie * 08:42 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2013.codfw.wmnet * 08:41 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2013.codfw.wmnet * 08:35 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 08:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host krb1002.eqiad.wmnet * 08:30 elukey@dns1004: END - running authdns-update * 08:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1012.eqiad.wmnet with reason: host reimage * 08:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Security updates * 08:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:28 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:28 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Security updates * 08:27 elukey@dns1004: START - running authdns-update * 08:26 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 08:26 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host krb1002.eqiad.wmnet * 08:22 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1012.eqiad.wmnet with reason: host reimage * 08:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host krb2002.codfw.wmnet * 08:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast6003.wikimedia.org * 08:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2013.codfw.wmnet with OS trixie * 08:13 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast6003.wikimedia.org * 08:12 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast3007.wikimedia.org * 08:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host krb2002.codfw.wmnet * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Security updates * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:09 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:09 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Security updates * 08:07 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1012.eqiad.wmnet with OS trixie * 08:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast3007.wikimedia.org * 08:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast5005.wikimedia.org * 07:58 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast5005.wikimedia.org * 07:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1013.eqiad.wmnet with OS trixie * 07:53 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2013.codfw.wmnet with reason: host reimage * 07:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1021: Security updates * 07:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:53 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:53 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1021: Security updates * 07:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast1004.wikimedia.org * 07:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2013.codfw.wmnet with reason: host reimage * 07:46 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast1004.wikimedia.org * 07:40 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1013.eqiad.wmnet with reason: host reimage * 07:36 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1013.eqiad.wmnet with reason: host reimage * 07:31 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2013 * 07:31 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2013 * 07:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1021: Security updates * 07:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:30 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:30 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1021: Security updates * 07:24 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2013 * 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2013.codfw.wmnet 87.0.192.10.in-addr.arpa 7.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:24 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2013.codfw.wmnet 87.0.192.10.in-addr.arpa 7.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2013 - mvernon@cumin2003" * 07:24 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2013 - mvernon@cumin2003" * 07:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1013.eqiad.wmnet with OS trixie * 07:19 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 07:19 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2013 * 07:19 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2013.codfw.wmnet with OS trixie * 07:13 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] (duration: 07m 48s) * 07:09 kharlan@deploy2003: kharlan: Continuing with deployment * 07:08 kharlan@deploy2003: kharlan: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:06 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 01:15 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 01:14 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply == 2026-07-14 == * 22:51 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_magru * 22:51 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7016.magru.wmnet * 22:46 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_magru * 22:46 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7008.magru.wmnet * 22:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7015.magru.wmnet * 22:04 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7007.magru.wmnet * 21:29 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7014.magru.wmnet * 21:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7006.magru.wmnet * 21:13 dzahn@cumin2002: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 0:15:00 on gerrit.wikimedia.org with reason: reboot * 21:11 mutante: gerrit2003 (gerrit.wikimedia.org) - reboot for maintenance * 21:11 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on gerrit2003.wikimedia.org with reason: reboot * 20:56 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:56 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:56 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:55 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 20:48 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7013.magru.wmnet * 20:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7005.magru.wmnet * 20:41 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host phab1005.eqiad.wmnet with OS trixie * 20:28 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] (duration: 06m 47s) * 20:24 sbassett@deploy2003: sbassett: Continuing with deployment * 20:23 sbassett@deploy2003: sbassett: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:23 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on phab1005.eqiad.wmnet with reason: host reimage * 20:21 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] * 20:20 aokoth@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on phab1005.eqiad.wmnet with reason: host reimage * 20:12 jhuneidi@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] (duration: 07m 42s) * 20:07 jhuneidi@deploy2003: jhuneidi, priyankar22: Continuing with deployment * 20:06 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7012.magru.wmnet * 20:06 jhuneidi@deploy2003: jhuneidi, priyankar22: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:04 jhuneidi@deploy2003: Started scap sync-world: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] * 20:02 aokoth@cumin1003: START - Cookbook sre.hosts.reimage for host phab1005.eqiad.wmnet with OS trixie * 20:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7004.magru.wmnet * 20:00 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet * 19:57 aokoth@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet * 19:24 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7011.magru.wmnet * 19:19 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7003.magru.wmnet * 19:11 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] (duration: 08m 33s) * 19:07 jforrester@deploy2003: jforrester: Continuing with deployment * 19:04 jforrester@deploy2003: jforrester: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:02 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] * 18:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7010.magru.wmnet * 18:38 mutante: rotating phabricator-gerrit bot token (its-phabricator) * 18:18 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 17:44 swfrench@deploy2003: Finished scap sync-world: Deployment to pick up new production image (duration: 31m 44s) * 17:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7002.magru.wmnet * 17:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7009.magru.wmnet * 17:32 swfrench@deploy2003: swfrench: Continuing with deployment * 17:29 swfrench@deploy2003: swfrench: Deployment to pick up new production image synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:17 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2035: repooling after rack b5 maintenance * 17:16 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool es2035: repooling after rack b5 maintenance * 17:16 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2188: repooling after rack b5 maintenance * 17:12 swfrench@deploy2003: Started scap sync-world: Deployment to pick up new production image * 17:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7001.magru.wmnet * 16:57 swfrench-wmf: reprepro include php8.3_8.3.32-1+wmf12u2 into component/php83 for bookworm-wikimedia * 16:50 sukhe: pool cp2046 * 16:47 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4039.ulsfo.wmnet * 16:44 sukhe: sudo cumin -b31 "A:cp" "run-puppet-agent" * 16:33 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on contint1003.wikimedia.org with reason: reboot * 16:32 mutante: contint1003 - main CI server - rebooting * 16:31 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2188: repooling after rack b5 maintenance * 16:31 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2178: repooling after rack b5 maintenance * 16:29 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 16:28 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 16:28 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 16:28 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 16:18 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2014.codfw.wmnet * 16:18 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 16:17 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2014.codfw.wmnet * 16:10 mvernon@cumin1003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-thanos-proxies (exit_code=0) rolling restart_daemons on A:thanos-fe * 16:09 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 16:07 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp4039.ulsfo.wmnet * 16:07 mvernon@cumin1003: START - Cookbook sre.swift.roll-restart-reboot-swift-thanos-proxies rolling restart_daemons on A:thanos-fe * 16:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2014.codfw.wmnet with OS trixie * 15:56 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1014.eqiad.wmnet with OS trixie * 15:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2014.codfw.wmnet with reason: host reimage * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2014 * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2014 * 15:28 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2014 * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2014.codfw.wmnet 194.16.192.10.in-addr.arpa 4.9.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:28 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2014.codfw.wmnet 194.16.192.10.in-addr.arpa 4.9.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2014 - mvernon@cumin2003" * 15:28 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2014 - mvernon@cumin2003" * 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Apply title-related policies when selecting the name of the entity - kamila@cumin1003" * 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Apply title-related policies when selecting the name of the entity - kamila@cumin1003 * 15:22 kamila@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Apply title-related policies when selecting the name of the entity - kamila@cumin1003 * 15:22 kamila@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Apply title-related policies when selecting the name of the entity - kamila@cumin1003" * 15:20 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 15:20 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2014 * 15:20 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2014.codfw.wmnet with OS trixie * 15:19 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1014.eqiad.wmnet with OS trixie * 15:01 dancy@deploy2003: Installation of scap version "4.274.1" completed for 3 hosts * 15:00 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2177: repooling after rack b5 maintenance * 15:00 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2159: repooling after rack b5 maintenance * 14:59 dancy@deploy2003: Installing scap version "4.274.1" for 3 host(s) * 14:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2015.codfw.wmnet with OS trixie * 14:54 seanleong-wmde: Finished populateSitesTable for isvwiki ([[phab:T429939|T429939]]) * 14:53 javiermonton@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] (duration: 07m 35s) * 14:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1015.eqiad.wmnet with OS trixie * 14:49 javiermonton@deploy2003: javiermonton: Continuing with deployment * 14:48 javiermonton@deploy2003: javiermonton: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:46 javiermonton@deploy2003: Started scap sync-world: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] * 14:42 otto@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 14:41 otto@deploy2003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 14:41 otto@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 14:40 otto@deploy2003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 14:40 otto@deploy2003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 14:39 otto@deploy2003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 14:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2015.codfw.wmnet with reason: host reimage * 14:34 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1015.eqiad.wmnet with reason: host reimage * 14:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2015.codfw.wmnet with reason: host reimage * 14:30 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1015.eqiad.wmnet with reason: host reimage * 14:30 seanleong-wmde@deploy2003: mwscript-k8s job started: foreachwikiindblist wikidataclient extensions/Wikibase/lib/maintenance/populateSitesTable.php --force-protocol https # [[phab:T429939|T429939]] * 14:24 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling reboot on A:durum and not (A:durum-eqiad or A:durum-codfw or A:durum-esams) and A:durum * 14:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2015.codfw.wmnet with OS trixie * 14:15 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2016.codfw.wmnet with OS trixie * 14:14 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2159: repooling after rack b5 maintenance * 14:14 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1015.eqiad.wmnet with OS trixie * 14:12 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1016.eqiad.wmnet with OS trixie * 14:12 cmooney@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=pki,name=codfw * 14:12 sbisson@deploy2003: helmfile [codfw] DONE helmfile.d/services/cxserver: sync * 14:11 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2002.codfw.wmnet * 14:11 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2002.codfw.wmnet * 14:11 sbisson@deploy2003: helmfile [codfw] START helmfile.d/services/cxserver: sync * 14:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1005.wikimedia.org * 14:07 sbisson@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cxserver: sync * 14:07 sbisson@deploy2003: helmfile [eqiad] START helmfile.d/services/cxserver: sync * 14:05 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1005.wikimedia.org * 14:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader2005.wikimedia.org * 14:02 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2003.codfw.wmnet * 14:02 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2003.codfw.wmnet * 14:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=tegola-vector-tiles,name=codfw * 14:00 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=kartotherian,name=codfw * 14:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader2005.wikimedia.org * 13:58 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2016.codfw.wmnet with reason: host reimage * 13:57 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-ncredir (exit_code=0) rolling reboot on A:ncredir and A:ncredir * 13:57 sbisson@deploy2003: helmfile [staging] DONE helmfile.d/services/cxserver: sync * 13:56 sbisson@deploy2003: helmfile [staging] START helmfile.d/services/cxserver: sync * 13:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1016.eqiad.wmnet with reason: host reimage * 13:52 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:52 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:51 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2016.codfw.wmnet with reason: host reimage * 13:50 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1016.eqiad.wmnet with reason: host reimage * 13:49 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy (exit_code=0) rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 13:49 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling reboot on A:wikidough * 13:46 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-tcp-proxy (exit_code=0) rolling reboot on A:tcpproxy and A:tcpproxy * 13:43 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and not (A:durum-eqiad or A:durum-codfw or A:durum-esams) and A:durum * 13:42 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=97) rolling reboot on A:durum and A:durum * 13:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2011.codfw.wmnet * 13:36 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-reboot (exit_code=0) rolling reboot on A:dnsbox and A:ulsfo and (A:dnsbox) * 13:36 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns4004.wikimedia.org * 13:34 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1016.eqiad.wmnet with OS trixie * 13:34 topranks: reboot lsw1-b5-codfw to upgrade JunOS [[phab:T430918|T430918]] * 13:34 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2016.codfw.wmnet with OS trixie * 13:32 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2002.codfw.wmnet * 13:31 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2011.codfw.wmnet * 13:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2012.codfw.wmnet * 13:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2017.codfw.wmnet with OS trixie * 13:25 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1017.eqiad.wmnet with OS trixie * 13:24 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2012.codfw.wmnet * 13:22 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2002.codfw.wmnet * 13:22 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:22 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:22 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns4004.wikimedia.org * 13:21 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2005.codfw.wmnet * 13:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2013.codfw.wmnet * 13:18 elukey@dns1004: END - running authdns-update * 13:17 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2005.codfw.wmnet * 13:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2004.codfw.wmnet * 13:16 elukey@dns1004: START - running authdns-update * 13:16 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1046: es1046 after reimage * 13:14 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1029.eqiad.wmnet,service=s8 * 13:14 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1029.eqiad.wmnet,service=s5 * 13:13 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1029.eqiad.wmnet,service=s5 * 13:13 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1029.eqiad.wmnet,service=s8 * 13:13 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2004.codfw.wmnet * 13:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2013.codfw.wmnet * 13:11 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2014.codfw.wmnet * 13:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2003.codfw.wmnet * 13:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2017.codfw.wmnet with reason: host reimage * 13:07 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:07 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns4003.wikimedia.org * 13:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2003.codfw.wmnet * 13:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1067.eqiad.wmnet * 13:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1067.eqiad.wmnet * 13:06 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1067.eqiad.wmnet * 13:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm-test1001.wikimedia.org * 13:05 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2017.codfw.wmnet with reason: host reimage * 13:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1017.eqiad.wmnet with reason: host reimage * 13:04 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2014.codfw.wmnet * 13:03 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:02 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2188: codfw rack B5 depool for maintenance * 13:02 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_magru * 13:01 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2188: codfw rack B5 depool for maintenance * 13:01 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_magru * 13:01 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2178: codfw rack B5 depool for maintenance * 13:01 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm-test1001.wikimedia.org * 13:01 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2178: codfw rack B5 depool for maintenance * 13:01 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2177: codfw rack B5 depool for maintenance * 13:00 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2177: codfw rack B5 depool for maintenance * 12:59 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1068.eqiad.wmnet * 12:59 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1068.eqiad.wmnet * 12:58 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola-vector-tiles,name=codfw * 12:58 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2159: codfw rack B5 depool for maintenance * 12:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1017.eqiad.wmnet with reason: host reimage * 12:58 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola,name=codfw * 12:57 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=kartotherian,name=codfw * 12:57 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2159: codfw rack B5 depool for maintenance * 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 30 hosts with reason: lsw1-b5-codfw JunOS upgrade * 12:55 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lsw1-b5-codfw,lsw1-b5-codfw IPv6,lsw1-b5-codfw.mgmt,ssw1-a[1,8]-codfw.mgmt with reason: switch upgade lsw1-b5-codfw * 12:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet * 12:49 cmooney@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=pki,name=codfw * 12:49 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1067.eqiad.wmnet with OS trixie * 12:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet * 12:48 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2017.codfw.wmnet with OS trixie * 12:47 topranks: depool codfw pki in dns discovery ahead of lsw1-b5-codfw maintenance [[phab:T430918|T430918]] * 12:47 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns4003.wikimedia.org * 12:47 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and A:ulsfo and (A:dnsbox) * 12:47 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2018.codfw.wmnet * 12:45 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2018.codfw.wmnet * 12:45 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and A:durum * 12:45 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-tcp-proxy rolling reboot on A:tcpproxy and A:tcpproxy * 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host pki-root1002.eqiad.wmnet * 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1009.eqiad.wmnet * 12:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1009.eqiad.wmnet * 12:44 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 12:43 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-ncredir rolling reboot on A:ncredir and A:ncredir * 12:43 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling reboot on A:wikidough * 12:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2018.codfw.wmnet with OS trixie * 12:42 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1017.eqiad.wmnet with OS trixie * 12:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader2006.wikimedia.org * 12:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1018.eqiad.wmnet with OS trixie * 12:39 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1009.eqiad.wmnet * 12:38 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host pki-root1002.eqiad.wmnet * 12:38 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1009.eqiad.wmnet * 12:38 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1008.eqiad.wmnet * 12:38 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1008.eqiad.wmnet * 12:35 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader2006.wikimedia.org * 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1006.wikimedia.org * 12:33 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1008.eqiad.wmnet * 12:30 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1046: es1046 after reimage * 12:29 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host es1046.eqiad.wmnet with OS trixie * 12:29 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1006.wikimedia.org * 12:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test2005.wikimedia.org * 12:28 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1008.eqiad.wmnet * 12:27 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1007.eqiad.wmnet * 12:27 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1007.eqiad.wmnet * 12:27 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1067.eqiad.wmnet with reason: host reimage * 12:25 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1068.eqiad.wmnet with reason: vacuum overlarge container dbs * 12:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2018.codfw.wmnet with reason: host reimage * 12:24 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test2005.wikimedia.org * 12:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test1005.wikimedia.org * 12:23 Amir1: mwscript-k8s --follow --dblist=ores -- extensions/ORES/maintenance/PurgeScoreCache.php --model goodfaith --old ([[phab:T431159|T431159]]) * 12:22 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1007.eqiad.wmnet * 12:22 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test1005.wikimedia.org * 12:22 atsukoito: restarting pybal on lvs2013 `low-traffic` for https://gerrit.wikimedia.org/r/1310535 * 12:22 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1007.eqiad.wmnet * 12:21 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1006.eqiad.wmnet * 12:21 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1006.eqiad.wmnet * 12:20 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1018.eqiad.wmnet with reason: host reimage * 12:19 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2018.codfw.wmnet with reason: host reimage * 12:18 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1067.eqiad.wmnet with reason: host reimage * 12:16 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1006.eqiad.wmnet * 12:15 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1006.eqiad.wmnet * 12:15 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1005.eqiad.wmnet * 12:15 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1005.eqiad.wmnet * 12:15 atsukoito: restarting pybal on lvs2014 for https://gerrit.wikimedia.org/r/1310535 * 12:12 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1018.eqiad.wmnet with reason: host reimage * 12:11 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1005.eqiad.wmnet * 12:11 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1005.eqiad.wmnet * 12:10 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1004.eqiad.wmnet * 12:10 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1004.eqiad.wmnet * 12:09 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on es1046.eqiad.wmnet with reason: host reimage * 12:08 atsukoito: restarting pybal on lvs1019 `low-traffic` for https://gerrit.wikimedia.org/r/1310535 * 12:06 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1004.eqiad.wmnet * 12:06 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1004.eqiad.wmnet * 12:06 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1003.eqiad.wmnet * 12:06 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1003.eqiad.wmnet * 12:05 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on es1046.eqiad.wmnet with reason: host reimage * 12:05 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "set ml-serve1001 back to active state - cmooney@cumin1003" * 12:04 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "set ml-serve1001 back to active state - cmooney@cumin1003" * 12:04 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:02 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1003.eqiad.wmnet * 12:01 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1003.eqiad.wmnet * 12:01 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1002.eqiad.wmnet * 12:01 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1002.eqiad.wmnet * 12:01 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:59 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2018.codfw.wmnet with OS trixie * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1067 * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1067 * 11:59 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1067 * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1067.eqiad.wmnet 17.48.64.10.in-addr.arpa 7.1.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:59 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1067.eqiad.wmnet 17.48.64.10.in-addr.arpa 7.1.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1067 - blake@cumin1003" * 11:58 atsukoito: restarting pybal on lvs1018 `high-traffic2` for https://gerrit.wikimedia.org/r/1310535 * 11:57 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1002.eqiad.wmnet * 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2019.codfw.wmnet with OS trixie * 11:56 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1002.eqiad.wmnet * 11:56 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 11:56 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1018.eqiad.wmnet with OS trixie * 11:54 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 11:54 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:54 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1019.eqiad.wmnet with OS trixie * 11:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:49 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:49 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:49 aikochou@deploy2003: helmfile [codfw] DONE helmfile.d/services/changeprop: sync * 11:48 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host es1046.eqiad.wmnet with OS trixie * 11:48 aikochou@deploy2003: helmfile [codfw] START helmfile.d/services/changeprop: sync * 11:48 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310535 * 11:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1046: Reimage to Trixie * 11:44 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1046: Reimage to Trixie * 11:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5:00:00 on es1046.eqiad.wmnet with reason: Reimage to Trixie * 11:42 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:42 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:42 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 11:42 aikochou@deploy2003: helmfile [eqiad] DONE helmfile.d/services/changeprop: sync * 11:41 aikochou@deploy2003: helmfile [eqiad] START helmfile.d/services/changeprop: sync * 11:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2019.codfw.wmnet with reason: host reimage * 11:36 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] (duration: 09m 41s) * 11:36 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:36 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:35 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:35 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:32 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1019.eqiad.wmnet with reason: host reimage * 11:32 jforrester@deploy2003: jforrester, gengh: Continuing with deployment * 11:29 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2019.codfw.wmnet with reason: host reimage * 11:28 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1019.eqiad.wmnet with reason: host reimage * 11:28 jforrester@deploy2003: jforrester, gengh: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:26 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] * 11:20 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2003.codfw.wmnet * 11:20 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:19 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2003.codfw.wmnet * 11:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:12 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1019.eqiad.wmnet with OS trixie * 11:10 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2019.codfw.wmnet with OS trixie * 11:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1020.eqiad.wmnet with OS trixie * 11:10 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1067 - blake@cumin1003" * 11:09 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] (duration: 12m 12s) * 11:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2020.codfw.wmnet with OS trixie * 11:03 kharlan@deploy2003: kharlan: Continuing with deployment * 11:01 blake@cumin1003: START - Cookbook sre.dns.netbox * 11:01 kharlan@deploy2003: kharlan: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:57 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] * 10:55 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] (duration: 31m 40s) * 10:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1020.eqiad.wmnet with reason: host reimage * 10:52 marostegui@dns1004: START - running authdns-update * 10:49 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2020.codfw.wmnet with reason: host reimage * 10:49 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:48 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1020.eqiad.wmnet with reason: host reimage * 10:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2020.codfw.wmnet with reason: host reimage * 10:43 kharlan@deploy2003: kharlan: Continuing with deployment * 10:42 kharlan@deploy2003: kharlan: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:32 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1020.eqiad.wmnet with OS trixie * 10:29 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2159: Repooling after switchover * 10:29 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1067 * 10:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1021.eqiad.wmnet with OS trixie * 10:27 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1067.eqiad.wmnet with OS trixie * 10:27 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:27 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1067.eqiad.wmnet * 10:27 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:26 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1067.eqiad.wmnet * 10:26 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1067.eqiad.wmnet * 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2020.codfw.wmnet with OS trixie * 10:26 blake@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1055.eqiad.wmnet * 10:26 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1055.eqiad.wmnet * 10:26 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1055.eqiad.wmnet * 10:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2021.codfw.wmnet with OS trixie * 10:24 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] * 10:11 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1055.eqiad.wmnet with OS trixie * 10:09 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1021.eqiad.wmnet with reason: host reimage * 10:05 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2021.codfw.wmnet with reason: host reimage * 10:03 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310129 revert * 10:02 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1021.eqiad.wmnet with reason: host reimage * 10:01 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2021.codfw.wmnet with reason: host reimage * 09:58 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310129 * 09:50 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1055.eqiad.wmnet with reason: host reimage * 09:45 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1055.eqiad.wmnet with reason: host reimage * 09:45 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1021.eqiad.wmnet with OS trixie * 09:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1022.eqiad.wmnet with OS trixie * 09:44 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2159: Repooling after switchover * 09:44 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2021.codfw.wmnet with OS trixie * 09:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2022.codfw.wmnet with OS trixie * 09:31 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2159.codfw.wmnet * 09:28 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1055 * 09:28 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1055 * 09:27 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on ms-fe1022.eqiad.wmnet with reason: host reimage * 09:27 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1022.eqiad.wmnet with reason: host reimage * 09:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2022.codfw.wmnet with reason: host reimage * 09:21 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2022.codfw.wmnet with reason: host reimage * 09:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2159: Rebooting db2159.codfw.wmnet * 09:20 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2159: Rebooting db2159.codfw.wmnet * 09:18 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 09:18 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 09:18 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 09:17 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 09:16 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2159.codfw.wmnet * 09:13 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b] (thin): Regular analytics weekly train THIN [analytics/refinery@ad6e05b8] (duration: 02m 07s) * 09:11 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b] (thin): Regular analytics weekly train THIN [analytics/refinery@ad6e05b8] * 09:10 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1022.eqiad.wmnet with OS trixie * 09:07 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1023.eqiad.wmnet with OS trixie * 09:06 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b]: Regular analytics weekly train [analytics/refinery@ad6e05b8] (duration: 04m 49s) * 09:04 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2022.codfw.wmnet with OS trixie * 09:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2023.codfw.wmnet with OS trixie * 09:01 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b]: Regular analytics weekly train [analytics/refinery@ad6e05b8] * 09:01 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@ad6e05b8] (duration: 02m 01s) * 09:00 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1055 * 09:00 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1055.eqiad.wmnet 50.32.64.10.in-addr.arpa 0.5.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:00 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1055.eqiad.wmnet 50.32.64.10.in-addr.arpa 0.5.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:00 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:00 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1055 - blake@cumin1003" * 09:00 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1055 - blake@cumin1003" * 08:59 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@ad6e05b8] * 08:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2159 [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94811 and previous config saved to /var/cache/conftool/dbconfig/20260714-085624-cwilliams.json * 08:55 blake@cumin1003: START - Cookbook sre.dns.netbox * 08:55 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1055 * 08:54 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1055.eqiad.wmnet with OS trixie * 08:54 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1055.eqiad.wmnet * 08:53 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1055.eqiad.wmnet * 08:53 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1055.eqiad.wmnet * 08:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2220 to s7 primary [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94810 and previous config saved to /var/cache/conftool/dbconfig/20260714-085239-cwilliams.json * 08:51 cezmunsta: Starting s7 codfw failover from db2159 to db2220 - [[phab:T430920|T430920]] * 08:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1023.eqiad.wmnet with reason: host reimage * 08:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2220 with weight 0 [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94809 and previous config saved to /var/cache/conftool/dbconfig/20260714-084553-cwilliams.json * 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s7 [[phab:T430920|T430920]] * 08:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2023.codfw.wmnet with reason: host reimage * 08:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1023.eqiad.wmnet with reason: host reimage * 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2023.codfw.wmnet with reason: host reimage * 08:34 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:34 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:29 marostegui@dns1004: END - running authdns-update * 08:29 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox-dev2003.codfw.wmnet * 08:27 marostegui@dns1004: START - running authdns-update * 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker2*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2009.codfw.wmnet * 08:26 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2009.codfw.wmnet * 08:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1023.eqiad.wmnet with OS trixie * 08:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox-dev2003.codfw.wmnet * 08:24 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:24 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:24 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2023.codfw.wmnet with OS trixie * 08:24 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1029.eqiad.wmnet with reason: reboot * 08:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1027.eqiad.wmnet with reason: reboot * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:21 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2009.codfw.wmnet * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:20 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2009.codfw.wmnet * 08:20 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2008.codfw.wmnet * 08:20 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2008.codfw.wmnet * 08:15 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2008.codfw.wmnet * 08:14 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2008.codfw.wmnet * 08:14 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2007.codfw.wmnet * 08:14 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2007.codfw.wmnet * 08:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1024.eqiad.wmnet with OS trixie * 08:12 elukey@cumin1003: END (PASS) - Cookbook sre.pki.restart-reboot (exit_code=0) rolling reboot on P<nowiki>{</nowiki>pki*<nowiki>}</nowiki> and (A:pki) * 08:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 08:10 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 08:09 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2007.codfw.wmnet * 08:08 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2007.codfw.wmnet * 08:08 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2006.codfw.wmnet * 08:08 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2006.codfw.wmnet * 08:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2024.codfw.wmnet with OS trixie * 08:03 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2006.codfw.wmnet * 08:02 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2006.codfw.wmnet * 08:02 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2005.codfw.wmnet * 08:02 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2005.codfw.wmnet * 07:58 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2005.codfw.wmnet * 07:58 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2005.codfw.wmnet * 07:57 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2004.codfw.wmnet * 07:57 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2004.codfw.wmnet * 07:54 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki.discovery.wmnet. on all recursors * 07:54 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache pki.discovery.wmnet. on all recursors * 07:53 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2004.codfw.wmnet * 07:53 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2004.codfw.wmnet * 07:53 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2003.codfw.wmnet * 07:53 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2003.codfw.wmnet * 07:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1024.eqiad.wmnet with reason: host reimage * 07:49 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki.discovery.wmnet. on all recursors * 07:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2003.codfw.wmnet * 07:49 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache pki.discovery.wmnet. on all recursors * 07:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2024.codfw.wmnet with reason: host reimage * 07:48 elukey@cumin1003: START - Cookbook sre.pki.restart-reboot rolling reboot on P<nowiki>{</nowiki>pki*<nowiki>}</nowiki> and (A:pki) * 07:46 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1024.eqiad.wmnet with reason: host reimage * 07:45 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2024.codfw.wmnet with reason: host reimage * 07:45 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2003.codfw.wmnet * 07:44 elukey@cumin1003: END (PASS) - Cookbook sre.misc-clusters.restart-reboot-config-master (exit_code=0) rolling reboot on P<nowiki>{</nowiki>config-master*<nowiki>}</nowiki> and (A:config-master or A:config-master-eqiad or A:config-master-codfw) * 07:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2002.codfw.wmnet * 07:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2002.codfw.wmnet * 07:39 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2002.codfw.wmnet * 07:39 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) config-master.discovery.wmnet. on all recursors * 07:39 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache config-master.discovery.wmnet. on all recursors * 07:39 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2002.codfw.wmnet * 07:39 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker2*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl200*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl2003.codfw.wmnet * 07:36 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl2003.codfw.wmnet * 07:35 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) config-master.discovery.wmnet. on all recursors * 07:35 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache config-master.discovery.wmnet. on all recursors * 07:34 elukey@cumin1003: START - Cookbook sre.misc-clusters.restart-reboot-config-master rolling reboot on P<nowiki>{</nowiki>config-master*<nowiki>}</nowiki> and (A:config-master or A:config-master-eqiad or A:config-master-codfw) * 07:31 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl2003.codfw.wmnet * 07:31 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl2003.codfw.wmnet * 07:31 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl2002.codfw.wmnet * 07:31 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl2002.codfw.wmnet * 07:29 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1024.eqiad.wmnet with OS trixie * 07:28 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2024.codfw.wmnet with OS trixie * 07:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl2002.codfw.wmnet * 07:26 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl2002.codfw.wmnet * 07:26 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl200*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 07:26 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 07:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 06:50 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lists1004.wikimedia.org * 06:44 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host lists1004.wikimedia.org * 06:25 marostegui@dns1004: END - running authdns-update * 06:23 marostegui@dns1004: START - running authdns-update * 06:22 marostegui@dns1004: END - running authdns-update * 06:20 marostegui@dns1004: START - running authdns-update * 06:17 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1026.eqiad.wmnet with reason: reboot * 06:04 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: sync * 06:04 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: sync * 06:03 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync * 06:03 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync * 06:02 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync * 06:01 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync * 06:01 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync * 06:00 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync * 05:59 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:59 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:40 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:39 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:26 marostegui@dns1004: END - running authdns-update * 05:24 marostegui@dns1004: START - running authdns-update * 05:24 marostegui@dns1004: START - running authdns-update * 05:13 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1004.wikimedia.org * 05:07 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1004.wikimedia.org * 04:01 mwpresync@deploy2003: Pruned MediaWiki: 1.47.0-wmf.8 (duration: 01m 07s) * 03:39 mwpresync@deploy2003: Finished scap sync-world: testwikis to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] (duration: 36m 01s) * 03:03 mwpresync@deploy2003: Started scap sync-world: testwikis to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 29s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-13 == * 23:33 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1064.eqiad.wmnet * 23:33 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1064.eqiad.wmnet * 23:08 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1064.eqiad.wmnet with reason: vacuum overlarge container dbs * 23:06 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1069.eqiad.wmnet * 23:06 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1069.eqiad.wmnet * 22:34 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1069.eqiad.wmnet with reason: vacuum overlarge container dbs * 21:18 maryum: Deployed security fix for [[phab:T321092|T321092]] * 20:28 swfrench-wmf: reprepro include etcd-mirror_0.0.12-1+deb13u1 into main for trixie-wikimedia - [[phab:T424266|T424266]] * 20:26 swfrench-wmf: reprepro include etcd-mirror_0.0.12-1+deb12u1 into main for bookworm-wikimedia - [[phab:T428495|T428495]] * 20:23 dancy@deploy2003: Finished scap sync-world: Testing [[phab:T431635|T431635]] (duration: 03m 36s) * 20:19 dancy@deploy2003: Started scap sync-world: Testing [[phab:T431635|T431635]] * 20:18 dancy@deploy2003: Installation of scap version "4.274.0" completed for 3 hosts * 20:16 dancy@deploy2003: Installing scap version "4.274.0" for 3 host(s) * 20:12 kemayo@deploy2003: Finished scap sync-world: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] (duration: 08m 25s) * 20:07 kemayo@deploy2003: soda, esanders, kemayo: Continuing with deployment * 20:05 kemayo@deploy2003: soda, esanders, kemayo: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there * 20:04 kemayo@deploy2003: Started scap sync-world: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] * 18:22 cwhite: lvextend vg0/srv +500g on centrallog hosts * 18:19 cdobbins@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS trixie * 17:46 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1071.eqiad.wmnet * 17:46 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1071.eqiad.wmnet * 17:13 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1071.eqiad.wmnet with reason: vacuum overlarge container dbs * 17:07 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1065.eqiad.wmnet * 17:07 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1065.eqiad.wmnet * 17:06 dzahn@dns1006: END - running authdns-update * 17:04 dzahn@dns1006: START - running authdns-update * 17:01 dzahn@dns1006: END - running authdns-update * 16:59 dzahn@dns1006: START - running authdns-update * 16:51 dancy@deploy2003: Finished scap sync-world: testing [[phab:T428971|T428971]] (duration: 03m 37s) * 16:47 dancy@deploy2003: Started scap sync-world: testing [[phab:T428971|T428971]] * 16:45 atsukoito: restarting pybal on lvs1019 to flush IP address for `cirrussearch1122.eqiad.wmnet` after moving the vlan [[phab:T431311|T431311]] * 16:42 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:42 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:42 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:42 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:42 Amir1: mwscript-k8s --follow --dblist=ores -- extensions/ORES/maintenance/PurgeScoreCache.php --model damaging --old ([[phab:T431159|T431159]]) * 16:34 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Pool test * 16:34 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 16:34 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 16:34 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Pool test * 16:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Depool test * 16:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 16:33 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 16:33 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Depool test * 16:31 dancy@deploy2003: Installation of scap version "4.273.0" completed for 159 hosts * 16:29 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1065.eqiad.wmnet with reason: vacuum overlarge container dbs * 16:27 dancy@deploy2003: Installing scap version "4.273.0" for 159 host(s) * 16:27 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics-external: sync * 16:27 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics-external: sync * 16:26 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics-external: sync * 16:26 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics-external: sync * 16:22 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync * 16:21 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync * 16:21 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: sync * 16:21 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: sync * 16:19 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync * 16:19 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync * 16:18 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync * 16:17 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync * 15:59 atsukoito: restarting pybal on lvs1018 for https://gerrit.wikimedia.org/r/1310117 * 15:55 aikochou@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 15:50 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310117 * 15:46 aikochou@deploy2003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 15:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host kafka-logging1006.eqiad.wmnet * 15:43 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host kafka-logging1006.eqiad.wmnet * 15:41 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host ganeti-test[2001-2003].codfw.wmnet * 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host ganeti-test[2001-2003].codfw.wmnet * 15:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host netbox1003.eqiad.wmnet * 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host netbox1003.eqiad.wmnet * 15:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host netbox2003.codfw.wmnet * 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host netbox2003.codfw.wmnet * 15:36 sukhe: restart pybal on lvs1020 * 15:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet * 15:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet * 15:08 btullis@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:06 btullis@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 15:01 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:01 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:35 cdobbins@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 14:34 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:33 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:33 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:32 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:29 cdobbins@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 14:28 marostegui@dns1004: END - running authdns-update * 14:28 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:27 marostegui@dns1004: START - running authdns-update * 14:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1023.eqiad.wmnet with reason: reboot * 14:18 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2009.codfw.wmnet with OS trixie * 14:14 swfrench-wmf: start rolling run-puppet-agent on A:cp for ATS config change - [[phab:T428909|T428909]] [[phab:T431838|T431838]] * 14:05 swfrench-wmf: disable-puppet on A:cp for ATS config change - [[phab:T428909|T428909]] [[phab:T431838|T431838]] * 14:05 cdobbins@cumin2002: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie * 14:02 marostegui@dns1004: END - running authdns-update * 14:00 marostegui@dns1004: START - running authdns-update * 14:00 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1070.eqiad.wmnet * 14:00 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1070.eqiad.wmnet * 13:58 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2009.codfw.wmnet with reason: host reimage * 13:57 cdobbins@cumin2002: conftool action : set/pooled=no; selector: name=dns7002.* * 13:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2009.codfw.wmnet with reason: host reimage * 13:48 rscout@deploy2003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply * 13:48 rscout@deploy2003: helmfile [eqiad] START helmfile.d/services/miscweb: apply * 13:48 rscout@deploy2003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply * 13:47 rscout@deploy2003: helmfile [codfw] START helmfile.d/services/miscweb: apply * 13:40 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:33 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2009.codfw.wmnet with OS trixie * 13:30 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1070.eqiad.wmnet with reason: vacuum overlarge container dbs * 13:28 aude@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] (duration: 11m 12s) * 13:23 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:22 aude@deploy2003: aikochou, javiermonton, aude, gkm563: Continuing with deployment * 13:22 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:19 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:19 aude@deploy2003: aikochou, javiermonton, aude, gkm563: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] synced to the testservers * 13:17 aude@deploy2003: Started scap sync-world: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] * 13:01 ladsgroup@deploy2003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 13:01 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:00 ladsgroup@deploy2003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 12:59 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:52 ladsgroup@deploy2003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 12:51 ladsgroup@deploy2003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 12:48 atsuko@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 12:48 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 12:47 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2008.codfw.wmnet with OS trixie * 12:47 atsuko@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 12:47 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply * 12:47 atsuko@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:46 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 12:45 atsuko@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:45 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply * 12:45 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:44 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:43 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] (duration: 07m 02s) * 12:38 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 12:37 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:36 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] * 12:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2008.codfw.wmnet with reason: host reimage * 12:23 Msz2001: Deployed changes to private code for Suggested Investigations * 12:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2008.codfw.wmnet with reason: host reimage * 12:20 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:19 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:17 atsuko@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 12:17 atsuko@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 12:16 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:15 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] (duration: 07m 14s) * 12:10 mszwarc@deploy2003: mszwarc: Continuing with deployment * 12:09 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:07 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] * 12:04 mszwarc@deploy2003: sync-world aborted: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] (duration: 00m 29s) * 12:03 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] * 12:01 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2008.codfw.wmnet with OS trixie * 12:00 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] (duration: 07m 37s) * 11:55 zabe@deploy2003: zabe: Continuing with deployment * 11:54 zabe@deploy2003: zabe: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:52 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] * 11:51 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:43 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:35 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:34 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:33 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:30 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:28 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:27 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:17 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2007.codfw.wmnet with OS trixie * 11:09 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] * 11:06 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=s8 * 11:00 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=x3 * 11:00 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=s5 * 10:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2007.codfw.wmnet with reason: host reimage * 10:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2007.codfw.wmnet with reason: host reimage * 10:51 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host dse-k8s-worker1023 * 10:50 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host dse-k8s-worker1023 * 10:44 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dse-k8s-worker1023 * 10:43 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host dse-k8s-worker1023 * 10:42 marostegui@cumin1003: dbctl commit (dc=all): 'Change x4 masters [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P94804 and previous config saved to /var/cache/conftool/dbconfig/20260713-104248-marostegui.json * 10:37 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:37 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:35 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:35 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:34 atsuko@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 10:34 atsuko@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 10:33 marostegui@cumin1003: dbctl commit (dc=all): 'Push x4 initial dbctl config [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P94803 and previous config saved to /var/cache/conftool/dbconfig/20260713-103259-marostegui.json * 10:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2007.codfw.wmnet with OS trixie * 09:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2006.codfw.wmnet with OS trixie * 09:42 marostegui@dns1004: END - running authdns-update * 09:40 marostegui@dns1004: START - running authdns-update * 09:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2006.codfw.wmnet with reason: host reimage * 09:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2006.codfw.wmnet with reason: host reimage * 09:06 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1024.eqiad.wmnet with reason: reboot * 09:06 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:01 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2006.codfw.wmnet with OS trixie * 08:44 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 08:43 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 08:43 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:42 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:42 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:42 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:41 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 08:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup2004.codfw.wmnet * 08:38 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:33 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host db1208.eqiad.wmnet * 08:30 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=x3 * 08:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2005.codfw.wmnet with OS trixie * 08:28 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup2004.codfw.wmnet * 08:28 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup2003.codfw.wmnet * 08:24 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1039: Repooling after testing * 08:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on clouddb1016.eqiad.wmnet with reason: cloning * 08:23 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s5 * 08:23 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s8 * 08:21 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 08:21 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 08:17 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup2003.codfw.wmnet * 08:17 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1004.eqiad.wmnet * 08:14 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1208.eqiad.wmnet * 08:11 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host phab1005.eqiad.wmnet * 08:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2005.codfw.wmnet with reason: host reimage * 08:07 marostegui@dns1004: END - running authdns-update * 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1004.eqiad.wmnet * 08:07 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1003.eqiad.wmnet * 08:07 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:05 marostegui@dns1004: START - running authdns-update * 08:05 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2005.codfw.wmnet with reason: host reimage * 08:05 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 08:05 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host phab1005.eqiad.wmnet * 08:05 marostegui@dns1004: START - running authdns-update * 08:05 marostegui@dns1004: START - running authdns-update * 08:05 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 08:04 marostegui@dns1004: START - running authdns-update * 08:00 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit1003.wikimedia.org * 07:58 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1003.eqiad.wmnet * 07:58 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1002-dev.eqiad.wmnet * 07:58 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:58 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:54 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1002-dev.eqiad.wmnet * 07:54 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1001-dev.eqiad.wmnet * 07:54 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit1003.wikimedia.org * 07:53 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 07:53 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:52 Msz2001: UTC morning backport+config window done * 07:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2005.codfw.wmnet with OS trixie * {{safesubst:SAL entry|1=07:50 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark (T429943}} * 07:49 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1001-dev.eqiad.wmnet * 07:46 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:46 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:45 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 07:45 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:44 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 07:44 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:43 mszwarc@deploy2003: mszwarc, danielyepezgarces, anzx: Continuing with deployment * {{safesubst:SAL entry|1=07:39 mszwarc@deploy2003: mszwarc, danielyepezgarces, anzx: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark}} * 07:39 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1039: Repooling after testing * {{safesubst:SAL entry|1=07:36 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark (T429943)}} * 07:35 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] (duration: 30m 03s) * 07:25 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit2002.wikimedia.org * 07:22 mszwarc@deploy2003: mszwarc: Continuing with deployment * 07:21 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:19 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit2002.wikimedia.org * 07:15 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aphlict1002.eqiad.wmnet * 07:11 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host aphlict1002.eqiad.wmnet * 07:08 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2003.wikimedia.org * 07:05 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] * 07:02 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2003.wikimedia.org * 07:02 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2002.wikimedia.org * 06:55 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2002.wikimedia.org * 06:55 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1003.wikimedia.org * 06:49 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1003.wikimedia.org * 06:34 marostegui: Drop m5 ipoid database [[phab:T431007|T431007]] * 06:29 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1027.eqiad.wmnet with reason: reboot * 06:24 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1028.eqiad.wmnet with reason: reboot * 06:21 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1025.eqiad.wmnet with reason: reboot * 06:17 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1022.eqiad.wmnet with reason: reboot * 06:03 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on dbproxy[2005-2008].codfw.wmnet with reason: reboot * 05:37 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1217,1228].eqiad.wmnet with reason: cloning * 05:11 marostegui: Drop users_to_rename table [[phab:T431842|T431842]] * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-12 == * 16:01 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2209 [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94792 and previous config saved to /var/cache/conftool/dbconfig/20260712-160124-marostegui.json * 15:58 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2205 to s3 primary [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94791 and previous config saved to /var/cache/conftool/dbconfig/20260712-155853-marostegui.json * 15:58 marostegui: Starting s3 codfw emergency failover from db2209 to db2205 - [[phab:T431950|T431950]] * 15:51 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2205 with weight 0 [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94790 and previous config saved to /var/cache/conftool/dbconfig/20260712-155135-marostegui.json * 15:51 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Primary switchover s3 [[phab:T431950|T431950]] * 02:01 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 01m 17s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-11 == * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 26s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-10 == * 19:12 jhathaway@dns1004: END - running authdns-update * 19:10 jhathaway@dns1004: START - running authdns-update * 18:23 mutante: vrts2002 rebooting (not the active host) * 18:21 mutante: lists2001, phab2003 - rebooting (not the active hosts) * 18:16 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on A:lvs-high-traffic2-codfw * 18:15 swfrench@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on A:lvs-high-traffic2-codfw * 17:15 mutante: [doc1004:~] $ sudo systemctl start rsync-doc-host-data-sync ([[phab:T431856|T431856]]) * 17:09 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1004.eqiad.wmnet * 17:08 jhathaway@dns1004: END - running authdns-update * 17:07 jhathaway@dns1004: START - running authdns-update * 17:06 jhathaway: depooling puppetserver1002, cause of errors is still unknown * 17:03 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1004.eqiad.wmnet * 16:57 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1003.eqiad.wmnet * 16:51 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1003.eqiad.wmnet * 16:48 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 16:48 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2004.codfw.wmnet * 16:42 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2004.codfw.wmnet * 16:41 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2003.codfw.wmnet * 16:35 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2003.codfw.wmnet * 16:33 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2002.codfw.wmnet * 16:27 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2002.codfw.wmnet * 16:25 mutante: gitlab-runners (production) rebooting cluster one by one * 16:17 mutante: etherpad1004/etherpad2002 - (etherpad.wikimedia.org) - rebooting * 16:13 mutante: doc1004/doc2003 (doc.wikimedia.org backends) - rebooting * 16:02 mutante: releases1003/releases2003 (releases.wikimedia.org backends) - rebooting for maintenance * 15:26 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2007-dev.codfw.wmnet * 15:19 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2007-dev.codfw.wmnet * 15:14 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host cloudcephosd2007-dev.codfw.wmnet * 15:14 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2007-dev.codfw.wmnet * 15:14 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host cloudcephosd2006-dev.codfw.wmnet * 15:07 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2006-dev.codfw.wmnet * 15:07 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2005-dev.codfw.wmnet * 14:59 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2005-dev.codfw.wmnet * 14:59 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2004-dev.codfw.wmnet * 14:53 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2004-dev.codfw.wmnet * 14:53 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2007-dev.codfw.wmnet * 14:51 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1054.eqiad.wmnet * 14:51 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1054.eqiad.wmnet * 14:51 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1054.eqiad.wmnet * 14:47 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2007-dev.codfw.wmnet * 14:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2006-dev.codfw.wmnet * 14:41 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2006-dev.codfw.wmnet * 14:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2005-dev.codfw.wmnet * 14:37 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2005-dev.codfw.wmnet * 14:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2005-dev.codfw.wmnet * 14:29 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2005-dev.codfw.wmnet * 14:29 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2006-dev.codfw.wmnet * 14:21 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2006-dev.codfw.wmnet * 14:21 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2010-dev.codfw.wmnet * 14:15 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2010-dev.codfw.wmnet * 14:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudgw2004-dev.codfw.wmnet * 14:10 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1054.eqiad.wmnet with OS trixie * 14:09 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudgw2004-dev.codfw.wmnet * 14:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudgw2003-dev.codfw.wmnet * 14:02 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudgw2003-dev.codfw.wmnet * 14:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2004-dev.codfw.wmnet * 13:53 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2004-dev.codfw.wmnet * 13:53 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2003-dev.codfw.wmnet * 13:48 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:44 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2003-dev.codfw.wmnet * 13:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2002-dev.codfw.wmnet * 13:42 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:41 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:41 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:37 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2002-dev.codfw.wmnet * 13:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudidp2001-dev.codfw.wmnet * 13:33 blake@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker1054.eqiad.wmnet with reason: host reimage * 13:33 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudidp2001-dev.codfw.wmnet * 13:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudnet2006-dev.codfw.wmnet * 13:26 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudnet2006-dev.codfw.wmnet * 13:26 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudnet2005-dev.codfw.wmnet * 13:23 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1054.eqiad.wmnet with reason: host reimage * 13:18 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudnet2005-dev.codfw.wmnet * 13:18 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudservices2005-dev.codfw.wmnet * 13:12 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudservices2005-dev.codfw.wmnet * 13:11 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudservices2004-dev.codfw.wmnet * 13:08 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudservices2004-dev.codfw.wmnet * 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudweb2002-dev.wikimedia.org * 13:05 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 13:05 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1054 * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1054 * 13:04 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1054 * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1054.eqiad.wmnet 49.32.64.10.in-addr.arpa 9.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:04 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1054.eqiad.wmnet 49.32.64.10.in-addr.arpa 9.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1054 - blake@cumin1003" * 13:04 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1054 - blake@cumin1003" * 13:01 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudweb2002-dev.wikimedia.org * 13:00 blake@cumin1003: START - Cookbook sre.dns.netbox * 12:59 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1054 * 12:57 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1054.eqiad.wmnet with OS trixie * 12:57 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1054.eqiad.wmnet * 12:56 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1054.eqiad.wmnet * 12:56 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1054.eqiad.wmnet * 12:47 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS trixie * 12:44 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:39 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 12:39 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 12:38 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:37 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:14 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:10 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:08 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:07 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:00 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:00 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:51 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:49 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:48 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:47 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:44 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:32 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2001.codfw.wmnet * 11:32 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1053.eqiad.wmnet * 11:32 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2001.codfw.wmnet * 11:32 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1053.eqiad.wmnet * 11:32 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1053.eqiad.wmnet * 11:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker2001.codfw.wmnet * 11:31 cgoubert@cumin1003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker2001.codfw.wmnet * 11:31 cgoubert@cumin1003: END (FAIL) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=1) rolling reimage on P<nowiki>{</nowiki>wikikube-worker2001*<nowiki>}</nowiki> and (A:wikikube-master-codfw or A:wikikube-worker-codfw) * 11:30 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:30 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:21 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 18 hosts with reason: reboot & upgrade * 11:20 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker2001.codfw.wmnet with OS trixie * 11:16 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:15 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:14 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:14 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 11 hosts * 11:14 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 11 hosts * 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:08 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 11:02 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1053.eqiad.wmnet with OS trixie * 11:01 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:58 cgoubert@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 10:57 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:38 cgoubert@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker2001.codfw.wmnet with OS trixie * 10:38 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2001.codfw.wmnet * 10:38 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2001.codfw.wmnet * 10:38 cgoubert@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on P<nowiki>{</nowiki>wikikube-worker2001*<nowiki>}</nowiki> and (A:wikikube-master-codfw or A:wikikube-worker-codfw) * 10:35 topranks: adjust IBGP outbound policy on lsw1-e2-codfw [[phab:T423430|T423430]] towards ssw1-e1-codfw * 10:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 cgoubert@cumin1003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:27 cgoubert@cumin1003: END (FAIL) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=1) rolling reimage on A:wikikube-worker-codfw * 10:27 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker2001.codfw.wmnet with OS bookworm * 10:25 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 10:24 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:24 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:15 cgoubert@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 10:11 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 10:11 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 10:08 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:08 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:07 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 10:06 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 10:00 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:55 cgoubert@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker2001.codfw.wmnet with OS bookworm * 09:55 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2005-2006,2011-2012].codfw.wmnet * 09:55 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2005-2006,2011-2012].codfw.wmnet * 09:51 cgoubert@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on A:wikikube-worker-codfw * 09:41 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1053.eqiad.wmnet with reason: host reimage * 09:37 topranks: apply new IBGP outbound policy on lsw1-e2-codfw [[phab:T423430|T423430]] * 09:36 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:36 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1053.eqiad.wmnet with reason: host reimage * 09:16 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1053 * 09:16 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1053 * 09:15 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1053 * 09:15 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1053.eqiad.wmnet 48.32.64.10.in-addr.arpa 8.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:15 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1053.eqiad.wmnet 48.32.64.10.in-addr.arpa 8.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:15 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:15 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1053 - blake@cumin1003" * 09:15 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1053 - blake@cumin1003" * 09:11 blake@cumin1003: START - Cookbook sre.dns.netbox * 09:11 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1053 * 09:08 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1053.eqiad.wmnet with OS trixie * 09:08 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1053.eqiad.wmnet * 09:08 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1053.eqiad.wmnet * 09:08 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1053.eqiad.wmnet * 09:04 brouberol@dns1004: END - running authdns-update * 09:03 brouberol@dns1004: START - running authdns-update * 08:41 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e] (thin): Regular analytics weekly train THIN [analytics/refinery@1abf22ea] (duration: 02m 11s) * 08:38 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e] (thin): Regular analytics weekly train THIN [analytics/refinery@1abf22ea] * 08:38 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e]: Regular analytics weekly train [analytics/refinery@1abf22ea] (duration: 05m 17s) * 08:38 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:34 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:33 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e]: Regular analytics weekly train [analytics/refinery@1abf22ea] * 08:32 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@1abf22ea] (duration: 02m 03s) * 08:30 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@1abf22ea] * 08:30 JavierMonton: Deploying Refinery at {{Gerrit|1abf22ea}} for changes 1308121/T427068 1306491/T430020 and {{Gerrit|1308190}} * 08:29 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:29 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 08:24 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:24 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 08:18 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:18 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 08:00 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db[2183-2184].codfw.wmnet * 08:00 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for db[2183-2184].codfw.wmnet * 07:52 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:52 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 07:49 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 11 hosts with reason: reboot & upgrade * 07:47 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:47 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 07:44 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:44 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 07:23 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 10 hosts * 07:23 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 10 hosts * 06:45 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 10 hosts with reason: reboot & upgrade * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 41s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-09 == * 23:33 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] (duration: 13m 26s) * 23:29 ladsgroup@deploy2003: ladsgroup, jdlrobson: Continuing with deployment * 23:22 ladsgroup@deploy2003: ladsgroup, jdlrobson: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:20 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] * 22:57 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1165.eqiad.wmnet * 22:56 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1165.eqiad.wmnet * 22:56 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1165.eqiad.wmnet * 22:45 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1165.eqiad.wmnet with OS trixie * 22:38 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 22:37 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 22:37 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 22:37 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:37 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 22:25 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1165.eqiad.wmnet with reason: host reimage * 22:17 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1165.eqiad.wmnet with reason: host reimage * 22:13 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 22:12 rzl@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 22:04 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 22:04 rzl@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1165 * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1165 * 22:02 jasmine@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1165 * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1165.eqiad.wmnet 115.48.64.10.in-addr.arpa 5.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:02 jasmine@cumin2002: START - Cookbook sre.dns.wipe-cache wikikube-worker1165.eqiad.wmnet 115.48.64.10.in-addr.arpa 5.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1165 - jasmine@cumin2002" * 22:02 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1165 - jasmine@cumin2002" * 22:02 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 21:57 jasmine@cumin2002: START - Cookbook sre.dns.netbox * 21:55 jasmine@cumin2002: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1165 * 21:54 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-worker1165.eqiad.wmnet with OS trixie * 21:54 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 21:54 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1165.eqiad.wmnet * 21:53 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 21:53 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1165.eqiad.wmnet * 21:53 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1165.eqiad.wmnet * 21:53 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 21:47 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 21:45 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 21:43 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 21:43 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 21:42 maryum: Deploy fix for [[phab:T431684|T431684]] * 21:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 21:27 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] (duration: 34m 14s) * 21:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs1002 * 21:23 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs1002 * 21:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS trixie * 21:22 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 21:20 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 22s) * 21:20 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 21:16 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 21:16 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 21:15 ladsgroup@deploy2003: ladsgroup: Continuing with deployment * 21:13 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2002.codfw.wmnet with OS bookworm * 21:11 ladsgroup@deploy2003: ladsgroup: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:08 ladsgroup@cumin1003: END (PASS) - Cookbook sre.wikireplicas.update-views (exit_code=0) * 21:07 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 6 hosts with reason: reboots * 20:54 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecycle work - bking@cumin2003 * 20:53 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:53 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] * 20:51 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) * 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 20:47 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecycle work - bking@cumin2003 * 20:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 20:41 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:41 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) * 20:40 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host relforge1008.eqiad.wmnet * 20:40 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1009.eqiad.wmnet with OS trixie * 20:33 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:32 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:32 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:31 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:31 ladsgroup@cumin1003: END (PASS) - Cookbook sre.wikireplicas.update-views (exit_code=0) * 20:29 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1008.eqiad.wmnet * 20:24 rzl@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 20:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2002.codfw.wmnet with OS bookworm * 20:23 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host relforge1008.eqiad.wmnet * 20:23 rzl@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 20:23 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1008.eqiad.wmnet * 20:22 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:22 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:21 bking@cumin2003: END (ERROR) - Cookbook sre.elasticsearch.rolling-operation (exit_code=97) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:21 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1009.eqiad.wmnet with reason: host reimage * 20:16 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:15 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1009.eqiad.wmnet with reason: host reimage * 20:12 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) * 20:02 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 19:55 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1009.eqiad.wmnet with OS trixie * 19:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 19:43 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 19:30 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 19:28 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 19:27 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 19:25 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 18:42 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 18:41 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 18:16 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 18:15 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 17:45 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for doh5004.wikimedia.org * 17:45 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for doh5004.wikimedia.org * 17:38 ladsgroup@deploy2003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 17:35 ladsgroup@deploy2003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 17:29 ladsgroup@deploy2003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 17:26 ladsgroup@deploy2003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 17:09 mutante: zuul[12]00[123] - rebooting for maintenance * 17:09 ebernhardson: start full in-place reindex of eqiad cirrussearch cluster * 17:08 dzahn@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-cluster (exit_code=99) * 17:08 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-cluster * 17:03 ebernhardson: start full in-place reindex of codfw cirrussearch cluster * 16:59 mutante: stewards1001/stewards2001 - reboot for maintenance * 16:54 ebernhardson: start full in-place reindex of cloudelastic cluster * 16:53 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 16:52 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply * 16:49 mutante: planet1003/planet2003 - rebooting * 16:47 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on doh5004.wikimedia.org with reason: random high load, investigating * 15:55 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 15:54 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 15:51 jynus: restarting backupmon1001 * 15:49 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 14 hosts * 15:49 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 14 hosts * 15:47 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backupmon1001.eqiad.wmnet with reason: restart * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:06 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 15:06 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 14:59 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply * 14:58 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply * 14:51 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 14 hosts * 14:51 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 14 hosts * 14:49 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 6 hosts with reason: reboot & upgrade * 14:48 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet * 14:48 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet * 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:42 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1052.eqiad.wmnet * 14:42 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1052.eqiad.wmnet * 14:42 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1052.eqiad.wmnet * 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:31 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:31 elukey@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: sync * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:30 elukey@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: sync * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:28 elukey@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: sync * 14:28 elukey@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: sync * 14:26 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:20 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1052.eqiad.wmnet with OS trixie * 14:19 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 6 hosts with reason: reboot & upgrade * 14:18 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:15 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:15 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:13 elukey: update druid indexation job for webrequest_sampled_live - [[phab:T427068|T427068]] * 14:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:09 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:09 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for papaul - jhancock@cumin2002" * 14:09 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for papaul - jhancock@cumin2002" * 14:07 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:07 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:04 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 14:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cuminunpriv1001.eqiad.wmnet * 13:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb1003.eqiad.wmnet * 13:59 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1052.eqiad.wmnet with reason: host reimage * 13:57 moritzm: installing requests security updates * 13:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cuminunpriv1001.eqiad.wmnet * 13:55 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb1003.eqiad.wmnet * 13:53 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1052.eqiad.wmnet with reason: host reimage * 13:50 moritzm: installing python-cryptography security updates * 13:47 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb2003.codfw.wmnet * 13:44 Msz2001: UTC afternoon config+backport window is done * 13:44 Msz2001: Updated `logging` on `metawiki` to fix log performers, [[phab:T431176|T431176]]#12105297 * 13:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb2003.codfw.wmnet * 13:43 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 13:43 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt1002.wikimedia.org * 13:41 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] (duration: 07m 30s) * 13:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt1002.wikimedia.org * 13:37 mszwarc@deploy2003: mszwarc: Continuing with deployment * 13:36 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1052 * 13:36 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1052 * 13:35 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:35 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1052 * 13:35 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1052.eqiad.wmnet 47.32.64.10.in-addr.arpa 7.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:35 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1052.eqiad.wmnet 47.32.64.10.in-addr.arpa 7.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:35 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:35 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1052 - blake@cumin1003" * 13:35 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1052 - blake@cumin1003" * 13:34 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] * 13:31 blake@cumin1003: START - Cookbook sre.dns.netbox * 13:31 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1052 * 13:30 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1052.eqiad.wmnet with OS trixie * 13:30 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1052.eqiad.wmnet * 13:29 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1052.eqiad.wmnet * 13:29 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1052.eqiad.wmnet * 13:17 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] (duration: 11m 26s) * 13:13 jforrester@deploy2003: jforrester: Continuing with deployment * 13:08 jforrester@deploy2003: jforrester: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:06 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] * 12:54 cgoubert@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply * 12:54 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:52 cgoubert@deploy2003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply * 12:45 cgoubert@deploy2003: helmfile [codfw] DONE helmfile.d/services/mobileapps: apply * 12:44 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:44 cgoubert@deploy2003: helmfile [codfw] START helmfile.d/services/mobileapps: apply * 12:43 cgoubert@deploy2003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 12:43 cgoubert@deploy2003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 12:42 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast4006.wikimedia.org * 12:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt2002.wikimedia.org * 12:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast7002.wikimedia.org * 12:18 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast4006.wikimedia.org * 12:18 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host ml-serve1004 * 12:18 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host ml-serve1004 * 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt2002.wikimedia.org * 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast7002.wikimedia.org * 12:10 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backup[2003,2014].codfw.wmnet with reason: reboot & upgrade * 12:10 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt-staging2001.codfw.wmnet * 12:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid1003.eqiad.wmnet * 12:06 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt-staging2001.codfw.wmnet * 12:05 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid1003.eqiad.wmnet * 12:03 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backup[1003,1014].eqiad.wmnet with reason: reboot & upgrade * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid2003.codfw.wmnet * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host irc1003.wikimedia.org * 11:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid2003.codfw.wmnet * 11:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host irc1003.wikimedia.org * 11:55 jmm@dns1004: END - running authdns-update * 11:53 jmm@dns1004: START - running authdns-update * 11:50 jmm@dns1004: END - running authdns-update * 11:48 jmm@dns1004: START - running authdns-update * 11:27 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host irc2003.wikimedia.org * 11:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host irc2003.wikimedia.org * 11:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint2001.codfw.wmnet * 11:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint1001.eqiad.wmnet * 11:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint2001.codfw.wmnet * 11:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint1001.eqiad.wmnet * 11:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-rw2001.wikimedia.org * 11:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-rw1001.wikimedia.org * 11:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-rw2001.wikimedia.org * 11:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-rw1001.wikimedia.org * 11:03 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon1003.wikimedia.org * 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2005.codfw.wmnet * 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2005.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 10:59 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2005.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 10:57 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon1003.wikimedia.org * 10:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon2002.wikimedia.org * 10:55 jmm@cumin2003: START - Cookbook sre.dns.netbox * 10:51 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon2002.wikimedia.org * 10:51 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:50 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2005.codfw.wmnet * 10:41 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:40 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host ml-serve1003 * 10:40 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host ml-serve1003 * 10:39 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2033.codfw.wmnet * 10:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install2005.wikimedia.org * 10:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install1005.wikimedia.org * 10:35 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1004.eqiad.wmnet with OS bookworm * 10:31 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install1005.wikimedia.org * 10:31 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install2005.wikimedia.org * 10:30 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install4004.wikimedia.org * 10:30 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install3004.wikimedia.org * 10:29 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install3004.wikimedia.org * 10:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install4004.wikimedia.org * 10:23 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 10:21 moritzm: failover Ganeti master in codfw/routed to ganeti2034 [[phab:T430928|T430928]] * 10:19 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.addnode (exit_code=0) for new host ganeti2031.codfw.wmnet to cluster codfw and group B * 10:19 moritzm: readded ganeti2031 to the codfw Ganeti cluster [[phab:T430910|T430910]] * 10:18 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1004.eqiad.wmnet with reason: host reimage * 10:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install5004.wikimedia.org * 10:18 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1003 * 10:18 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1003 * 10:17 jmm@cumin2003: START - Cookbook sre.ganeti.addnode for new host ganeti2031.codfw.wmnet to cluster codfw and group B * 10:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install6003.wikimedia.org * 10:16 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install5004.wikimedia.org * 10:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install6003.wikimedia.org * 10:15 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1004.eqiad.wmnet with reason: host reimage * 10:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1001.eqiad.wmnet * 10:14 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 10:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2008.wikimedia.org * 10:00 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ml-serve1004 * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1004 * 09:57 jmm@cumin2003: START - Cookbook sre.dns.netbox * 09:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install7002.wikimedia.org * 09:57 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1004 * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ml-serve1004.eqiad.wmnet 50.48.64.10.in-addr.arpa 0.5.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:57 klausman@cumin1003: START - Cookbook sre.dns.wipe-cache ml-serve1004.eqiad.wmnet 50.48.64.10.in-addr.arpa 0.5.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1004 - klausman@cumin1003" * 09:56 klausman@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1004 - klausman@cumin1003" * 09:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-coord1001.eqiad.wmnet * 09:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 09:55 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow7002.magru.wmnet * 09:52 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-coord1001.eqiad.wmnet * 09:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 09:52 klausman@cumin1003: START - Cookbook sre.dns.netbox * 09:50 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install7002.wikimedia.org * 09:50 klausman@cumin1003: START - Cookbook sre.hosts.move-vlan for host ml-serve1004 * 09:50 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1004.eqiad.wmnet with OS bookworm * 09:50 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1003.eqiad.wmnet with OS bookworm * 09:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1001.eqiad.wmnet * 09:49 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 09:49 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2008.wikimedia.org * 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2007.codfw.wmnet * 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2007.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 09:49 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow7002.magru.wmnet * 09:49 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2007.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 09:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard1003.eqiad.wmnet * 09:39 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard2003.codfw.wmnet * 09:39 jmm@cumin2003: START - Cookbook sre.dns.netbox * 09:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard1003.eqiad.wmnet * 09:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor1003.eqiad.wmnet * 09:35 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard2003.codfw.wmnet * 09:34 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2007.codfw.wmnet * 09:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor1003.eqiad.wmnet * 09:33 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor-dev2001.codfw.wmnet * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor2003.codfw.wmnet * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sretest1006.eqiad.wmnet * 09:27 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 09:25 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor-dev2001.codfw.wmnet * 09:25 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor2003.codfw.wmnet * 09:23 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] (duration: 06m 27s) * 09:23 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2205: codfw rack B4 repool after maintenance * 09:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host sretest1006.eqiad.wmnet * 09:23 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2204: codfw rack B4 repool after maintenance * 09:19 urbanecm@deploy2003: urbanecm: Continuing with deployment * 09:19 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:18 jmm@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 6 hosts with reason: reboot * 09:17 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] * 09:08 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ml-serve1003 * 09:08 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1003 * 09:07 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1003 * 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ml-serve1003.eqiad.wmnet 81.32.64.10.in-addr.arpa 1.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:07 klausman@cumin1003: START - Cookbook sre.dns.wipe-cache ml-serve1003.eqiad.wmnet 81.32.64.10.in-addr.arpa 1.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1003 - klausman@cumin1003" * 09:06 klausman@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1003 - klausman@cumin1003" * 08:58 klausman@cumin1003: START - Cookbook sre.dns.netbox * 08:57 klausman@cumin1003: START - Cookbook sre.hosts.move-vlan for host ml-serve1003 * 08:57 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1003.eqiad.wmnet with OS bookworm * 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=0) rolling reimage on P<nowiki>{</nowiki>ml-serve1003.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet * 08:55 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet * 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1003.eqiad.wmnet with OS bookworm * 08:39 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 08:38 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool db2205: codfw rack B4 repool after maintenance * 08:37 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool db2204: codfw rack B4 repool after maintenance * 08:36 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 08:35 hashar@deploy2003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 08:32 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:32 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:31 hashar@deploy2003: Rolling back deployment * 08:26 moritzm: failover Ganeti master in codfw to ganeti2048 [[phab:T430928|T430928]] * 08:16 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1003.eqiad.wmnet with OS bookworm * 08:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2004.codfw.wmnet * 08:16 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet * 08:16 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet * 08:16 klausman@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on P<nowiki>{</nowiki>ml-serve1003.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 08:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2002.codfw.wmnet * 08:15 XioNoX: lsw1-b4-codfw> request system reboot - [[phab:T430910|T430910]] * 08:15 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b4-codfw,lsw1-b4-codfw IPv6,lsw1-b4-codfw.mgmt with reason: Switch maintenance * 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for codfw rack B4 * 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:10 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2004.codfw.wmnet * 08:10 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2002.codfw.wmnet * 08:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2205: codfw rack B4 depool for maintenance * 08:08 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool db2205: codfw rack B4 depool for maintenance * 08:08 jmm@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin2003.codfw.wmnet * 08:08 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2204: codfw rack B4 depool for maintenance * 08:08 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool db2204: codfw rack B4 depool for maintenance * 08:08 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 27 hosts with reason: codfw rack B4 depool for maintenance * 08:03 jmm@cumin2002: START - Cookbook sre.hosts.reboot-single for host cumin2003.codfw.wmnet * 07:56 ayounsi@cumin1003: START - Cookbook sre.network.depool-rack with action 'depool' for codfw rack B4 * 07:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1008.eqiad.wmnet with OS trixie * 07:49 wmde-fisch@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] (duration: 08m 36s) * 07:44 wmde-fisch@deploy2003: wmde-fisch: Continuing with deployment * 07:43 wmde-fisch@deploy2003: wmde-fisch: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:41 wmde-fisch@deploy2003: Started scap sync-world: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] * 07:35 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1008.eqiad.wmnet with reason: host reimage * 07:31 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1008.eqiad.wmnet with reason: host reimage * 07:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1008.eqiad.wmnet with OS trixie * 07:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 07:00 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 06:59 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1008.eqiad.wmnet with OS trixie * 06:57 Emperor: rebalance thanos swift rings after previous re-image of thanos-fe1004 to trixie * 06:47 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1008.eqiad.wmnet with OS trixie * 04:10 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 14 days, 0:00:00 on cp6008.drmrs.wmnet with reason: Hardware failure - [[phab:T431651|T431651]] * 03:55 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp6008.* * 03:29 ryankemper: [[phab:T431311|T431311]] Repooled eqiad cirrussearch clusters (`chi/omega/psi`) following completion of OpenSearch 2.19 migration * 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad * 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=eqiad * 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 31s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-08 == * 23:52 Amir1: ladsgroup@deploy2003:~$ mwscript-k8s --follow -- extensions/ORES/maintenance/PurgeScoreCache.php --wiki=simplewiki --model damaging --old ([[phab:T431159|T431159]]) * 23:46 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 23:46 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing PTR for 2001:df2:e500:fe08::1 - cmooney@cumin1003" * 23:46 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing PTR for 2001:df2:e500:fe08::1 - cmooney@cumin1003" * 23:40 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 23:16 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 23:15 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 22:42 rzl@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 22:40 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] (duration: 12m 55s) * 22:40 rzl@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 22:37 rzl@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 22:36 rzl@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 22:35 rzl@deploy2003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 22:34 urbanecm@deploy2003: urbanecm: Continuing with deployment * 22:33 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:33 rzl@deploy2003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 22:32 rzl@deploy2003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 22:30 rzl@deploy2003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 22:30 rzl@deploy2003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 22:29 rzl@deploy2003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 22:27 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] * 22:26 rzl@deploy2003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 22:22 rzl@deploy2003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 22:21 rzl@deploy2003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 22:19 rzl@deploy2003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 22:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 22:17 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 22:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 22:13 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 22:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 22:13 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 22:09 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 22:06 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 22:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1094.eqiad.wmnet with OS trixie * 22:01 urbanecm: Make https://test.wikipedia.org/w/index.php?title=MediaWiki:GrowthExperimentsSuggestedEdits.json&diff=prev&oldid=750552 with GrowthExperiments disabled (via mw-experimental), then run `\MediaWiki\MediaWikiServices::getInstance()->get('CommunityConfiguration.ProviderFactory')->newProvider('GrowthSuggestedEdits')->getStore()->invalidate()` ([[phab:T431625|T431625]]) * 21:56 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d2-codfw * 21:55 urbanecm@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 21:55 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d2-codfw * 21:55 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c4-codfw * 21:55 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c4-codfw * 21:55 urbanecm@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2002 * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2002 * 21:54 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2002 * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2002.codfw.wmnet 50.32.192.10.in-addr.arpa 0.5.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:54 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2002.codfw.wmnet 50.32.192.10.in-addr.arpa 0.5.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2002 - bking@cumin2003" * 21:54 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2002 - bking@cumin2003" * 21:49 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:49 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2002 * 21:49 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2002.codfw.wmnet with OS trixie * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1094.eqiad.wmnet with reason: host reimage * 21:42 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 21:39 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 21:37 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1094.eqiad.wmnet with reason: host reimage * 21:36 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 21:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 21:29 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 21:27 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 21:22 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1094.eqiad.wmnet with OS trixie * 21:21 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host restbase2039.codfw.wmnet with OS bullseye * 21:21 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin2002" * 21:21 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin2002" * 21:04 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on restbase2039.codfw.wmnet with reason: host reimage * 21:00 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on restbase2039.codfw.wmnet with reason: host reimage * 20:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1073.eqiad.wmnet with OS trixie * 20:48 mutante: deploy2003 - kill 1102 (stunnel4) ; systemctl start stunnel4 ([[phab:T418262|T418262]]) * 20:42 cjming@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] (duration: 33m 02s) * 20:42 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host restbase2039.codfw.wmnet with OS bullseye * 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1073.eqiad.wmnet with reason: host reimage * 20:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1073.eqiad.wmnet with reason: host reimage * 20:30 cjming@deploy2003: cjming: Continuing with deployment * 20:28 cjming@deploy2003: cjming: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1098.eqiad.wmnet with OS trixie * 20:13 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1073.eqiad.wmnet with OS trixie * 20:09 cjming@deploy2003: Started scap sync-world: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] * 20:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1098.eqiad.wmnet with reason: host reimage * 19:56 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1098.eqiad.wmnet with reason: host reimage * 19:55 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d4-codfw * 19:54 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d4-codfw * 19:54 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c1-codfw * 19:54 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c1-codfw * 19:52 mutante: restarting gerrit on gerrit.wikimedia.org (gerrit2003) * 19:48 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2331.codfw.wmnet * 19:48 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2331.codfw.wmnet * 19:48 mutante: restarting gerrit on gerrit-replica.wikimedia.org (gerrit1003) * 19:46 mutante: restarting gerrit on gerrit-spare.wikimedia.org (gerrit2002) * 19:43 jasmine@cumin2002: conftool action : set/pooled=yes; selector: name=wikikube-worker2331.codfw.wmnet,cluster=kubernetes,service=kubesvc * 19:43 jasmine@cumin2002: conftool action : set/weight=10; selector: name=wikikube-worker2331.codfw.wmnet,cluster=kubernetes,service=kubesvc * 19:40 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1098.eqiad.wmnet with OS trixie * 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d5-codfw * 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d5-codfw * 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c7-codfw * 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c7-codfw * 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c5-codfw * 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c5-codfw * 19:30 jasmine_: ran homer on lsw1-d8-codfw, adding wikikube-worker2331 to cluster * 19:29 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1100.eqiad.wmnet with OS trixie * 19:20 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d8-codfw * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d8-codfw * 19:19 mutante: gerrit - replacing private key for registerEmail verification * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-magru * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device cr2-magru * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d7-codfw * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d7-codfw * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d3-codfw * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d1-codfw * 19:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d1-codfw * 19:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c2-codfw * 19:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c2-codfw * 19:11 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-codfw * 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-magru * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device cr1-magru * 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d8-codfw * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d8-codfw * 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d6-codfw * 19:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1100.eqiad.wmnet with reason: host reimage * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d6-codfw * 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c6-codfw * 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c6-codfw * 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c3-codfw * 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c3-codfw * 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b4-magru * 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device asw1-b4-magru * 19:08 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b3-magru * 19:08 cmooney@cumin1003: START - Cookbook sre.network.tls for network device asw1-b3-magru * 19:05 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1100.eqiad.wmnet with reason: host reimage * 19:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1122.eqiad.wmnet with OS trixie * 19:00 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 18:59 topranks: rolling out update to BGP ACL on Nokia Switches eqiad, codfw & ulsfo [[phab:T425703|T425703]] * 18:58 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 18:57 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 18:55 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 18:53 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 18:52 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 18:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1100.eqiad.wmnet with OS trixie * 18:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1068.eqiad.wmnet with OS trixie * 18:47 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1102.eqiad.wmnet with OS trixie * 18:47 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 18:46 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1122.eqiad.wmnet with reason: host reimage * 18:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1122.eqiad.wmnet with reason: host reimage * 18:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1068.eqiad.wmnet with reason: host reimage * 18:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1122 * 18:26 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1122 * 18:25 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1122 * 18:25 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1122.eqiad.wmnet 31.48.64.10.in-addr.arpa 1.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:25 bking@cumin2003: START - Cookbook sre.dns.wipe-cache cirrussearch1122.eqiad.wmnet 31.48.64.10.in-addr.arpa 1.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:25 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:25 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1122 - bking@cumin2003" * 18:25 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1122 - bking@cumin2003" * 18:21 rzl@deploy2003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 18:21 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1068.eqiad.wmnet with reason: host reimage * 18:21 rzl@deploy2003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 18:21 rzl@deploy2003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 18:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 18:19 rzl@deploy2003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 18:19 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:18 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1122 * 18:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1122.eqiad.wmnet with OS trixie * 18:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 18:15 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 18:13 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 18:13 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 18:10 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 18:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1068.eqiad.wmnet with OS trixie * 18:01 kamila@deploy2003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 18m 29s) * 18:00 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:55 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:42 kamila@deploy2003: Started scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] * 17:42 kamila@deploy2003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 19m 50s) * 17:42 kamila@deploy2003: Rolling back deployment * 17:35 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:31 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet * 17:18 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet * 17:16 kamila@deploy1003: Unlocked for deployment [MediaWiki]: switching deployment server (duration: 22m 07s) * 17:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 17:11 kamila@dns1005: END - running authdns-update * 17:09 kamila@dns1005: START - running authdns-update * 17:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 17:04 jasmine@cumin2002: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1164.eqiad.wmnet * 17:04 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1164.eqiad.wmnet * 17:04 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1164.eqiad.wmnet * 16:56 kamila@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on releases2003.codfw.wmnet,releases1003.eqiad.wmnet with reason: Deployment server switchover * 16:54 kamila@deploy1003: Locking from deployment [MediaWiki]: switching deployment server * 16:53 kamila@deploy1003: Unlocked for deployment [MediaWiki]: switching deployment server (duration: 04m 02s) * 16:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie * 16:49 kamila@deploy1003: Locking from deployment [MediaWiki]: switching deployment server * 16:46 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1095.eqiad.wmnet with OS trixie * 16:45 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1093.eqiad.wmnet with OS trixie * 16:43 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1164.eqiad.wmnet with OS trixie * 16:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1095.eqiad.wmnet with reason: host reimage * 16:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 16:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 16:23 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1164.eqiad.wmnet with reason: host reimage * 16:18 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on cirrussearch1093.eqiad.wmnet with reason: host reimage * 16:16 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1164.eqiad.wmnet with reason: host reimage * 16:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1095.eqiad.wmnet with reason: host reimage * 16:09 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 16:09 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 16:08 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1093.eqiad.wmnet with reason: host reimage * 15:59 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Pool test * 15:59 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:59 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 15:59 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Pool test * 15:58 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Depool test * 15:58 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:58 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 15:58 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Depool test * 15:57 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1164 * 15:57 jasmine@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1164 * 15:57 jasmine@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1164 * 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1164.eqiad.wmnet 114.48.64.10.in-addr.arpa 4.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:56 jasmine@cumin2002: START - Cookbook sre.dns.wipe-cache wikikube-worker1164.eqiad.wmnet 114.48.64.10.in-addr.arpa 4.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1164 - jasmine@cumin2002" * 15:56 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1164 - jasmine@cumin2002" * 15:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1093.eqiad.wmnet with OS trixie * 15:51 jasmine@cumin2002: START - Cookbook sre.dns.netbox * 15:51 jasmine@cumin2002: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1164 * 15:50 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-worker1164.eqiad.wmnet with OS trixie * 15:50 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1164.eqiad.wmnet * 15:50 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1164.eqiad.wmnet * 15:50 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1164.eqiad.wmnet * 15:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1095.eqiad.wmnet with OS trixie * 15:42 jasmine@cumin2002: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1164.eqiad.wmnet * 15:42 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1164.eqiad.wmnet * 15:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:42 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1164.eqiad.wmnet * 15:42 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1164.eqiad.wmnet * 15:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 15:39 elukey@cumin1003: START - Cookbook sre.hosts.provision for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 15:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1007.eqiad.wmnet with OS trixie * 15:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Pool test * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1007.eqiad.wmnet with reason: host reimage * 15:15 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 15:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet * 15:15 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 15:15 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1007.eqiad.wmnet with reason: host reimage * 15:15 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:14 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Pool test * 15:14 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet * 15:14 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2228: Depool test * 15:14 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db2228: Depool test * 15:10 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 15:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:08 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 15:06 blake@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 15:06 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet * 15:06 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 15:06 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 15:05 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet * 15:05 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 15:04 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:04 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 15:04 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:03 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 15:03 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 15:03 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:03 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T430909|T430909]] * 15:03 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:03 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 15:03 swfrench-wmf: restarted eqsin, codfw confds - [[phab:T430909|T430909]] * 15:03 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test1002.eqiad.wmnet * 15:02 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet * 14:59 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:59 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:55 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1007.eqiad.wmnet with OS trixie * 14:52 swfrench-wmf: restarted ulsfo confds, confirmed now connected to codfw backends except those using wikimedia.org SRV record - [[phab:T430909|T430909]] * 14:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:49 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:41 moritzm: uninstalling dhcpcd-base from trixie hosts which still have it installed [[phab:T414341|T414341]] * 14:40 sukhe: sudo cumin -b1 -s120 "P<nowiki>{</nowiki>lvs2011*<nowiki>}</nowiki> or P<nowiki>{</nowiki>lvs2012*<nowiki>}</nowiki>" "systemctl restart pybal.service" * 14:39 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:39 mvernon@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host thanos-be1007.eqiad.wmnet with OS trixie * 14:37 sukhe: restart pybal on lvs2013 to revert back to conf2004 * 14:35 sukhe: restart pybal on lvs2014 to revert back to conf2004 * 14:34 swfrench-wmf: switched codfw, eqsin, ulsfo etcd client SRV records back to codfw - [[phab:T430909|T430909]] * 14:32 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1002.eqiad.wmnet * 14:32 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet * 14:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1007.eqiad.wmnet with OS trixie * 14:31 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:31 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:31 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:30 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:30 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Pool test * 14:30 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:29 swfrench@dns1004: END - running authdns-update * 14:29 moritzm: installing jackson-core security updates * 14:27 swfrench@dns1004: START - running authdns-update * 14:22 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:22 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:22 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1119.eqiad.wmnet with OS trixie * 14:22 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:21 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:20 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:20 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 14:20 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:19 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 14:19 moritzm: installing librabbitmq security updates * 14:19 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1002.eqiad.wmnet * 14:18 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet * 14:16 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:15 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:15 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:14 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Pool test * 14:14 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox) * 14:14 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet * 14:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 14:08 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 14:07 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet * 14:05 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1118.eqiad.wmnet with OS trixie * 14:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test1001.eqiad.wmnet * 14:00 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-worker@eqiad * 14:00 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 13:59 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 13:58 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 13:57 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1119.eqiad.wmnet with reason: host reimage * 13:54 moritzm: installing libcap2 security updates * 13:53 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1119.eqiad.wmnet with reason: host reimage * 13:52 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet * 13:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) * 13:52 fceratto@cumin1003: START - Cookbook sre.mysql.depool * 13:50 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-worker@eqiad * 13:50 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1051.eqiad.wmnet * 13:50 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1051.eqiad.wmnet * 13:50 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1051.eqiad.wmnet * 13:49 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 13:45 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1081.eqiad.wmnet with OS trixie * 13:41 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1119 * 13:41 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1119 * 13:40 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1119 * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1119.eqiad.wmnet 97.32.64.10.in-addr.arpa 7.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1119.eqiad.wmnet 97.32.64.10.in-addr.arpa 7.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1119 - atsuko@cumin1003" * 13:40 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1119 - atsuko@cumin1003" * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1118.eqiad.wmnet with reason: host reimage * 13:39 moritzm: installing krb5 security updates * 13:37 Lucas_WMDE: UTC afternoon backport+config window done * 13:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1006.eqiad.wmnet with OS trixie * 13:36 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1118.eqiad.wmnet with reason: host reimage * 13:36 atsuko@cumin1003: START - Cookbook sre.dns.netbox * 13:35 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] (duration: 07m 46s) * 13:34 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1119 * 13:34 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1119.eqiad.wmnet with OS trixie * 13:30 sbisson@deploy1003: sbisson: Continuing with deployment * 13:30 moritzm: installing openssh security updates * 13:30 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling restart_daemons on A:wikidough * 13:29 sbisson@deploy1003: sbisson: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:27 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] * 13:26 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1051.eqiad.wmnet with OS trixie * 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1081.eqiad.wmnet with reason: host reimage * 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1118 * 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1118 * 13:22 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] (duration: 12m 12s) * 13:21 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1081.eqiad.wmnet with reason: host reimage * 13:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1006.eqiad.wmnet with reason: host reimage * 13:18 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1118 * 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1118.eqiad.wmnet 90.32.64.10.in-addr.arpa 0.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:18 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1118.eqiad.wmnet 90.32.64.10.in-addr.arpa 0.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1118 - atsuko@cumin1003" * 13:18 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1118 - atsuko@cumin1003" * 13:17 stran@deploy1003: stran: Continuing with deployment * 13:16 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough * 13:15 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:13 atsuko@cumin1003: START - Cookbook sre.dns.netbox * 13:12 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1118 * 13:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1006.eqiad.wmnet with reason: host reimage * 13:12 stran@deploy1003: stran: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:12 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1118.eqiad.wmnet with OS trixie * 13:10 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] * 13:05 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-worker@codfw * 13:05 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 13:05 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1081.eqiad.wmnet with OS trixie * 13:05 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1051.eqiad.wmnet with reason: host reimage * 13:04 moritzm: installing jq security updates * 13:04 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 13:01 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1051.eqiad.wmnet with reason: host reimage * 12:58 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-worker@codfw * 12:52 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 12:50 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host thanos-be1006.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1051 * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1051 * 12:43 moritzm: installing Python 3.11 security updates * 12:43 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1051 * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1051.eqiad.wmnet 46.32.64.10.in-addr.arpa 6.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:43 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1051.eqiad.wmnet 46.32.64.10.in-addr.arpa 6.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1051 - blake@cumin1003" * 12:43 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1051 - blake@cumin1003" * 12:38 blake@cumin1003: START - Cookbook sre.dns.netbox * 12:38 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1051 * 12:38 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1051.eqiad.wmnet with OS trixie * 12:37 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1051.eqiad.wmnet * 12:36 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1051.eqiad.wmnet * 12:36 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1051.eqiad.wmnet * 12:34 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1006.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 12:34 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1006.eqiad.wmnet with OS trixie * 12:27 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 12:27 mvernon@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host thanos-be1006.eqiad.wmnet with OS trixie * 12:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:02 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 12:01 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1006.eqiad.wmnet with OS trixie * 11:43 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 11:38 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1076.eqiad.wmnet with OS trixie * 11:26 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1075.eqiad.wmnet with OS trixie * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2047.codfw.wmnet * 11:19 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2047.codfw.wmnet * 11:19 moritzm: temporarily remove ganeti2031 from codfw cluster [[phab:T430910|T430910]] * 11:08 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1076.eqiad.wmnet with reason: host reimage * 11:08 moritzm: installing Linux 6.1.176 on Bookworm servers * 11:03 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1076.eqiad.wmnet with reason: host reimage * 11:00 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1075.eqiad.wmnet with reason: host reimage * 10:56 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1075.eqiad.wmnet with reason: host reimage * 10:47 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1076.eqiad.wmnet with OS trixie * 10:46 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1074.eqiad.wmnet with OS trixie * 10:45 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1005.eqiad.wmnet with OS trixie * 10:40 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1075.eqiad.wmnet with OS trixie * 10:32 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2031.codfw.wmnet * 10:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1005.eqiad.wmnet with reason: host reimage * 10:25 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1005.eqiad.wmnet with reason: host reimage * 10:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1074.eqiad.wmnet with reason: host reimage * 10:17 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1074.eqiad.wmnet with reason: host reimage * 10:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1005.eqiad.wmnet with OS trixie * 10:12 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet * 10:04 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 10:01 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet * 10:01 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1074.eqiad.wmnet with OS trixie * 10:01 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 09:43 cgoubert@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/aux-k8s-services/redioscope: apply * 09:43 cgoubert@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/aux-k8s-services/redioscope: apply * 09:43 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply * 09:35 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply * 09:34 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 09:34 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 09:33 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 41 days, 15:00:00 on db2252.codfw.wmnet with reason: Test * 09:32 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 09:32 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 09:31 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: codfw rack B3 pool after maintenance * 09:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 09:07 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 09:07 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 09:02 ladsgroup@cumin1003: END (PASS) - Cookbook sre.mysql.sanitarium_restart (exit_code=0) * 08:57 topranks: merge patch to shift eqiad <-> esams traffic onto new 40G circuit * 08:54 hashar@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1004.eqiad.wmnet with OS trixie * 08:50 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 08:50 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitarium_restart (exit_code=99) * 08:50 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 08:45 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool es2051: codfw rack B3 pool after maintenance * 08:44 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:44 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:43 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2007.codfw.wmnet * 08:43 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2007.codfw.wmnet * 08:42 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2031.codfw.wmnet * 08:41 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2031.codfw.wmnet * 08:40 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2031.codfw.wmnet * 08:38 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:38 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:35 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.sanitize-wiki (exit_code=97) Managing sanitization for wikis minwikiquote in section s3 * 08:33 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis minwikiquote in section s3 * 08:32 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Checking sanitization for wikis minwikiquote in section s5 * 08:30 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Checking sanitization for wikis minwikiquote in section s5 * 08:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Managing sanitization for wikis minwikiquote in section s5 * 08:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1004.eqiad.wmnet with reason: host reimage * 08:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1004.eqiad.wmnet with reason: host reimage * 08:23 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:22 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis minwikiquote in section s5 * 08:19 XioNoX: lsw1-b3-codfw> request system reboot - [[phab:T430909|T430909]] * 08:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Checking sanitization for wikis minwikiquote in section s5 * 08:17 hashar@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Checking sanitization for wikis minwikiquote in section s5 * 08:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for codfw rack B3 * 08:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2007.codfw.wmnet * 08:15 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lsw1-b3-codfw,lsw1-b3-codfw IPv6,lsw1-b3-codfw.mgmt with reason: Switch maintenance * 08:15 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2007.codfw.wmnet * 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:07 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:06 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: codfw rack B3 depool for maintenance * 08:05 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool es2051: codfw rack B3 depool for maintenance * 08:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1004.eqiad.wmnet with OS trixie * 08:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1005.eqiad.wmnet with OS trixie * 08:03 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 21 hosts with reason: codfw rack B3 depool for maintenance * 07:56 ayounsi@cumin1003: START - Cookbook sre.network.depool-rack with action 'depool' for codfw rack B3 * 07:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1005.eqiad.wmnet with reason: host reimage * 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1005.eqiad.wmnet with reason: host reimage * 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1005.eqiad.wmnet with OS bookworm * 07:29 moritzm: installing gnutls28 security updates * 07:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1005.eqiad.wmnet with OS trixie * 07:13 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1125.eqiad.wmnet with OS trixie * 07:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1005.eqiad.wmnet with reason: host reimage * 07:07 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aux-k8s-etcd1005.eqiad.wmnet with reason: host reimage * 06:56 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1005.eqiad.wmnet with OS bookworm * 06:54 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1125.eqiad.wmnet with reason: host reimage * 06:52 elukey: upgrade all trixie hosts to pywmflib 3.1 - [[phab:T430552|T430552]] * 06:50 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1125.eqiad.wmnet with reason: host reimage * 06:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 06:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 06:38 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1125.eqiad.wmnet with OS trixie * 05:42 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1107.eqiad.wmnet with OS trixie * 05:35 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1124.eqiad.wmnet with OS trixie * 05:31 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1101.eqiad.wmnet with OS trixie * 05:21 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1107.eqiad.wmnet with reason: host reimage * 05:17 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1124.eqiad.wmnet with reason: host reimage * 05:13 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1107.eqiad.wmnet with reason: host reimage * 05:13 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1101.eqiad.wmnet with reason: host reimage * 05:11 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1124.eqiad.wmnet with reason: host reimage * 05:10 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1101.eqiad.wmnet with reason: host reimage * 04:58 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1124.eqiad.wmnet with OS trixie * 04:56 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1107.eqiad.wmnet with OS trixie * 04:55 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1101.eqiad.wmnet with OS trixie * 02:27 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] (duration: 08m 14s) * 02:22 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 02:21 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 02:19 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] * 01:59 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] (duration: 09m 46s) * 01:55 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 01:51 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 01:49 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] * 01:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1099.eqiad.wmnet with OS trixie * 00:57 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1110.eqiad.wmnet with OS trixie * 00:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1099.eqiad.wmnet with reason: host reimage * 00:41 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1099.eqiad.wmnet with reason: host reimage * 00:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1110.eqiad.wmnet with reason: host reimage * 00:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1110.eqiad.wmnet with reason: host reimage * 00:26 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1099.eqiad.wmnet with OS trixie * 00:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1110.eqiad.wmnet with OS trixie == 2026-07-07 == * 22:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1097.eqiad.wmnet with OS trixie * 22:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1097.eqiad.wmnet with reason: host reimage * 22:24 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1097.eqiad.wmnet with reason: host reimage * 22:14 hashar: Restarting Gerrit on gerrit2002 and gerrit1003 (replicas) * 22:09 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1097.eqiad.wmnet with OS trixie * 22:07 hashar: Restarting Gerrit on gerrit2003 * 21:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 21:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 21:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 21:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 21:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1108.eqiad.wmnet with OS trixie * 20:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1091.eqiad.wmnet with OS trixie * 20:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1090.eqiad.wmnet with OS trixie * 20:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1108.eqiad.wmnet with reason: host reimage * 20:36 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1091.eqiad.wmnet with reason: host reimage * 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1090.eqiad.wmnet with reason: host reimage * 20:33 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1091.eqiad.wmnet with reason: host reimage * 20:30 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1108.eqiad.wmnet with reason: host reimage * 20:30 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1006.eqiad.wmnet * 20:30 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1090.eqiad.wmnet with reason: host reimage * 20:30 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1006.eqiad.wmnet * 20:27 jasmine_: "homer lsw1-c2-eqiad* commit "Added new stacked control plane wikikube-ctrl1006"" * 20:22 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] (duration: 07m 29s) * 20:20 jasmine_: "homer "cr*eqiad*" commit "Added new stacked control plane wikikube-ctrl1006"" * 20:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1091.eqiad.wmnet with OS trixie * 20:17 arlolra@deploy1003: arlolra: Continuing with deployment * 20:16 arlolra@deploy1003: arlolra: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:16 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1090.eqiad.wmnet with OS trixie * 20:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1108.eqiad.wmnet with OS trixie * 20:14 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] * 20:09 cwhite: remove 2026-04 swift log archives from centrallog2002 to free some space * 20:01 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=93) for host cirrussearch1108.eqiad.wmnet with OS trixie * 19:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1108.eqiad.wmnet with OS trixie * 19:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1090.eqiad.wmnet with OS trixie * 19:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1109.eqiad.wmnet with OS trixie * 19:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1092.eqiad.wmnet with OS trixie * 19:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1123.eqiad.wmnet with OS trixie * 19:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1109.eqiad.wmnet with reason: host reimage * 19:19 cdobbins@cumin2002: conftool action : set/pooled=yes; selector: name=dns7002.* * 19:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1092.eqiad.wmnet with reason: host reimage * 19:17 jasmine@dns1004: END - running authdns-update * 19:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1109.eqiad.wmnet with reason: host reimage * 19:15 jasmine@dns1004: START - running authdns-update * 19:15 cdobbins@dns1004: END - running authdns-update * 19:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1123.eqiad.wmnet with reason: host reimage * 19:13 cdobbins@dns1004: START - running authdns-update * 19:12 cdobbins@cumin2002: conftool action : set/pooled=yes; selector: name=dns7002.*,service=authdns-update * 19:11 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1092.eqiad.wmnet with reason: host reimage * 19:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1123.eqiad.wmnet with reason: host reimage * 18:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1123.eqiad.wmnet with OS trixie * 18:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1109.eqiad.wmnet with OS trixie * 18:56 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1092.eqiad.wmnet with OS trixie * 18:52 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 18:49 swfrench@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 18:40 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 18:38 swfrench@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 18:11 swfrench-wmf: restarted eqsin, codfw confds - [[phab:T430909|T430909]] * 18:01 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T430909|T430909]] * 17:59 swfrench-wmf: restarted ulsfo confds, confirmed now connected to eqiad backends - [[phab:T430909|T430909]] * 17:52 sukhe: restart pybal on lvs2011 to switch from conf2004 to conf1008: [[phab:T430909|T430909]] * 17:51 sukhe: restart pybal on lvs2012 to switch from conf2004 to conf1008 [puppet re-enabled there]: [[phab:T430909|T430909]] * 17:46 sukhe: restart pybal on lvs2013 to switch from conf2004 to conf1008: [[phab:T430909|T430909]] * 17:44 swfrench-wmf: switched codfw, eqsin, ulsfo etcd client SRV records to eqiad - [[phab:T430909|T430909]] * 17:43 swfrench@dns1004: END - running authdns-update * 17:40 swfrench@dns1004: START - running authdns-update * 17:40 sukhe: restart pybal on lvs2014 to switch from conf2004 to conf1008: [[phab:T430909|T430909]] * 17:21 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1003.eqiad.wmnet * 17:15 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1003.eqiad.wmnet * 17:14 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1002.eqiad.wmnet * 17:06 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1002.eqiad.wmnet * 17:06 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-low-traffic-codfw' 'systemctl restart pybal.service' # lvs2013, [[phab:T416623|T416623]] * 17:04 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1001.eqiad.wmnet * 17:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1111.eqiad.wmnet with OS trixie * 17:00 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal.service' # lvs2014, [[phab:T416623|T416623]] * 16:58 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1001.eqiad.wmnet * 16:58 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS bookworm * 16:55 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-low-traffic-eqiad' 'systemctl restart pybal.service' # lvs1019, [[phab:T416623|T416623]] * 16:53 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-secondary-eqiad' 'systemctl restart pybal.service' # lvs1020, [[phab:T416623|T416623]] * 16:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1111.eqiad.wmnet with reason: host reimage * 16:40 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1111.eqiad.wmnet with reason: host reimage * 16:38 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.peering (exit_code=99) with action 'configure' for AS: 47794 * 16:35 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 47794 * 16:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1111.eqiad.wmnet with OS trixie * 16:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1006.eqiad.wmnet with OS trixie * 16:06 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 16:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1006.eqiad.wmnet with reason: host reimage * 15:58 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1121.eqiad.wmnet with OS trixie * 15:58 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 15:56 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1006.eqiad.wmnet with reason: host reimage * 15:54 mutante: jenkins down in planned maintenance window * 15:42 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1037.eqiad.wmnet * 15:42 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1037.eqiad.wmnet * 15:42 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1037.eqiad.wmnet * 15:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1006.eqiad.wmnet with OS trixie * 15:34 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1121.eqiad.wmnet with reason: host reimage * 15:33 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS bookworm * 15:33 cdobbins@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host dns7002.wikimedia.org with OS trixie * 15:30 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1121.eqiad.wmnet with reason: host reimage * 15:29 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host clouddumps1001.wikimedia.org * 15:20 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1001.wikimedia.org * 15:18 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1121 * 15:18 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1121 * 15:18 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host clouddumps1002.wikimedia.org * 15:17 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1121 * 15:17 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:17 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply * 15:16 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply * 15:16 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:16 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1121 - atsuko@cumin1003" * 15:16 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1121 - atsuko@cumin1003" * 15:14 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1037.eqiad.wmnet with OS trixie * 15:11 atsuko@cumin1003: START - Cookbook sre.dns.netbox * 15:09 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org * 15:09 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1121 * 15:09 andrew@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host clouddumps1002.wikimedia.org * 15:09 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org * 15:09 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1121.eqiad.wmnet with OS trixie * 15:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1007.eqiad.wmnet with OS trixie * 15:08 andrew@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host clouddumps1002.wikimedia.org * 15:08 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org * 15:05 brennen@deploy1003: Finished deploy [phabricator/deployment@7e02037]: deploy phab1004 for [[phab:T431440|T431440]] (duration: 00m 47s) * 15:04 brennen@deploy1003: Started deploy [phabricator/deployment@7e02037]: deploy phab1004 for [[phab:T431440|T431440]] * 15:03 brennen@deploy1003: Finished deploy [phabricator/deployment@7e02037]: deploy phab2003 for [[phab:T431440|T431440]] (duration: 00m 51s) * 15:03 brennen@deploy1003: Started deploy [phabricator/deployment@7e02037]: deploy phab2003 for [[phab:T431440|T431440]] * 15:00 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71] (thin): Regular analytics weekly train THIN [analytics/refinery@7d8dc71f] (duration: 02m 10s) * 14:58 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71] (thin): Regular analytics weekly train THIN [analytics/refinery@7d8dc71f] * 14:58 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71]: Regular analytics weekly train [analytics/refinery@7d8dc71f] (duration: 04m 14s) * 14:54 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1037.eqiad.wmnet with reason: host reimage * 14:53 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71]: Regular analytics weekly train [analytics/refinery@7d8dc71f] * 14:53 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@7d8dc71f] (duration: 02m 00s) * 14:51 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@7d8dc71f] * 14:51 arnaudb@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on phab2003.codfw.wmnet,phab[1004-1006].eqiad.wmnet with reason: maintenance * 14:51 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1037.eqiad.wmnet with reason: host reimage * 14:50 JavierMonton: Deploying Refinery at {{Gerrit|7d8dc71f}} for change {{Gerrit|1308087}} / [[phab:T431318|T431318]] - update filerevision table sqoop and table * 14:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1007.eqiad.wmnet with reason: host reimage * 14:42 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1007.eqiad.wmnet with reason: host reimage * 14:40 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1083.eqiad.wmnet with OS trixie * 14:37 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:36 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:35 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-master@eqiad * 14:35 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 14:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 14:34 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1037 * 14:34 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1037 * 14:34 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:cleanMentorList.php --wiki=frwiki # [[phab:T427386|T427386]] * 14:34 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 14:34 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308112{{!}}Revert^2 "[Growth] frwiki: Deploy automated mentor list cleaner" (T427386)]] (duration: 06m 47s) * 14:34 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 14:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:33 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:32 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1037 * 14:31 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:31 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:29 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:29 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-master@eqiad * 14:29 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:29 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:29 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:28 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:27 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:27 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1308112{{!}}Revert^2 "[Growth] frwiki: Deploy automated mentor list cleaner" (T427386)]] * 14:26 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1007.eqiad.wmnet with OS trixie * 14:26 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:26 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 14:26 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:25 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:25 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:25 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:cleanMentorList.php --wiki=frwiki # [[phab:T427386|T427386]] * 14:24 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:24 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1037 - blake@cumin1003" * 14:24 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1037 - blake@cumin1003" * 14:20 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1083.eqiad.wmnet with reason: host reimage * 14:19 blake@cumin1003: START - Cookbook sre.dns.netbox * 14:19 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-master@codfw * 14:19 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 14:19 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1037 * 14:18 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1037.eqiad.wmnet with OS trixie * 14:18 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1037.eqiad.wmnet * 14:18 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 14:18 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1037.eqiad.wmnet * 14:18 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1037.eqiad.wmnet * 14:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2007.codfw.wmnet with OS trixie * 14:16 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1083.eqiad.wmnet with reason: host reimage * 14:15 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1036.eqiad.wmnet * 14:15 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1036.eqiad.wmnet * 14:14 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1036.eqiad.wmnet * 14:12 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-master@codfw * 14:11 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1120.eqiad.wmnet with OS trixie * 14:05 moritzm: installing distro-info-data updates from trixie/bookworm point releases * 14:04 fabfur: disable puppet on A:cp-text to selectively apply https://gerrit.wikimedia.org/r/c/operations/puppet/+/1308040 * 14:03 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] (duration: 27m 48s) * 14:00 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 14:00 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1083.eqiad.wmnet with OS trixie * 13:58 urbanecm@deploy1003: urbanecm: Continuing with deployment * 13:58 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:58 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1004.eqiad.wmnet with OS bookworm * 13:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2007.codfw.wmnet with reason: host reimage * 13:57 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 13:53 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1120.eqiad.wmnet with reason: host reimage * 13:50 moritzm: installing Linux 5.10.259 on Bullseye hosts * 13:47 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply * 13:47 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply * 13:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2007.codfw.wmnet with reason: host reimage * 13:46 cgoubert@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/aux-k8s-services/redioscope: apply * 13:46 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1120.eqiad.wmnet with reason: host reimage * 13:46 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:46 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:45 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:44 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:44 cgoubert@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/aux-k8s-services/redioscope: apply * 13:40 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 13:39 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:38 moritzm: installing e2fsprogs updates from Trixie point release * 13:35 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] * 13:33 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1120.eqiad.wmnet with OS trixie * 13:33 topranks: reset cr3-eqsin configuration so traffic uses it again after upgrade * 13:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1088.eqiad.wmnet with OS trixie * 13:32 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie * 13:32 cdobbins@cumin1003: conftool action : set/pooled=no; selector: name=dns7002.* * 13:29 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2007.codfw.wmnet with OS trixie * 13:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1004.eqiad.wmnet with reason: host reimage * 13:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2006.codfw.wmnet with OS trixie * 13:18 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1036.eqiad.wmnet with OS trixie * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aux-k8s-etcd1004.eqiad.wmnet with reason: host reimage * 13:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 13:16 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 13:15 jayme: Istio is being upgraded from 1.24.2 to 1.29.4 on wikikube staging eqiad and codfw - [[phab:T427401|T427401]] * 13:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1087.eqiad.wmnet with OS trixie * 13:14 topranks: reboot cr3-eqsin to install new JunOS and set PIC 0/0/0 to 100G * 13:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1088.eqiad.wmnet with reason: host reimage * 13:13 jmm@dns1004: END - running authdns-update * 13:12 jmm@dns1004: START - running authdns-update * 13:09 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1088.eqiad.wmnet with reason: host reimage * 13:07 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1082.eqiad.wmnet with OS trixie * 13:07 jmm@dns1004: END - running authdns-update * 13:06 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1004.eqiad.wmnet with OS bookworm * 13:05 jmm@dns1004: START - running authdns-update * 13:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2006.codfw.wmnet with reason: host reimage * 12:58 topranks: load updated JunOS on cr3-eqsin [[phab:T429386|T429386]] * 12:58 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1036.eqiad.wmnet with reason: host reimage * 12:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2001.codfw.wmnet * 12:57 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2006.codfw.wmnet with reason: host reimage * 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr1-codfw,cr[2-3]-eqsin,cr3-eqsin IPv6,cr3-eqsin.mgmt with reason: upgrade JunOS cr3-eqsin * 12:56 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lvs[5004-5006].eqsin.wmnet with reason: upgrade JunOS cr3-eqsin * 12:55 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 12:55 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 12:53 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1087.eqiad.wmnet with reason: host reimage * 12:53 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 12:52 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1088.eqiad.wmnet with OS trixie * 12:52 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1002.eqiad.wmnet * 12:52 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:51 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2001.codfw.wmnet * 12:49 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1036.eqiad.wmnet with reason: host reimage * 12:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 12:48 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1087.eqiad.wmnet with reason: host reimage * 12:44 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1082.eqiad.wmnet with reason: host reimage * 12:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1002.eqiad.wmnet * 12:42 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 12:42 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 12:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 12:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 12:39 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:39 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: move dumps-nfs IP to the shared one - filippo@cumin1003" * 12:39 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: move dumps-nfs IP to the shared one - filippo@cumin1003" * 12:39 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2006.codfw.wmnet with OS trixie * 12:38 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1082.eqiad.wmnet with reason: host reimage * 12:36 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:33 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:32 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1036 * 12:32 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1036 * 12:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2005.codfw.wmnet with OS trixie * 12:32 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1087.eqiad.wmnet with OS trixie * 12:30 jmm@dns1004: END - running authdns-update * 12:29 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1036 * 12:29 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1036.eqiad.wmnet 21.32.64.10.in-addr.arpa 1.2.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:29 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1036.eqiad.wmnet 21.32.64.10.in-addr.arpa 1.2.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:29 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:29 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1036 - blake@cumin1003" * 12:29 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1036 - blake@cumin1003" * 12:28 jmm@dns1004: START - running authdns-update * 12:26 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:26 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:23 blake@cumin1003: START - Cookbook sre.dns.netbox * 12:23 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1036 * 12:23 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1036.eqiad.wmnet with OS trixie * 12:22 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1036.eqiad.wmnet * 12:22 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1082.eqiad.wmnet with OS trixie * 12:22 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1036.eqiad.wmnet * 12:22 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1036.eqiad.wmnet * 12:21 marostegui: Restart mariadb@s7 on db1155 to pick up new filters - [[phab:T431124|T431124]] * 12:21 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 21 hosts with reason: restarting for replication filter * 12:20 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:19 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:14 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2005.codfw.wmnet with reason: host reimage * 12:14 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:08 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:07 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:07 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2005.codfw.wmnet with reason: host reimage * 12:06 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:06 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:06 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:05 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:05 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:04 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-master-eqiad * 12:04 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl1002.eqiad.wmnet * 12:04 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl1002.eqiad.wmnet * 12:04 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:04 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:03 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:03 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:03 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:03 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 11:59 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl1002.eqiad.wmnet * 11:59 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl1002.eqiad.wmnet * 11:59 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl1001.eqiad.wmnet * 11:59 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl1001.eqiad.wmnet * 11:56 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl1001.eqiad.wmnet * 11:56 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl1001.eqiad.wmnet * 11:56 klausman@cumin2002: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-master-eqiad * 11:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2005.codfw.wmnet with OS trixie * 11:49 blake@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on wikikube-worker1160.eqiad.wmnet with reason: Verifying matchers for silence * 11:42 topranks: cr3-eqsin, begin traffic drain to reset PIC and upgrade JunOS [[phab:T429386|T429386]] * 11:41 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs[5004-5006].eqsin.wmnet with reason: upgrade JunOS cr3-eqsin * 11:39 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr1-codfw,cr[2-3]-eqsin,cr3-eqsin IPv6,cr3-eqsin.mgmt with reason: upgrade JunOS cr3-eqsin * 11:36 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=thanos-fe2004.codfw.wmnet * 11:35 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1086.eqiad.wmnet with OS trixie * 11:35 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=thanos-fe2004.codfw.wmnet * 11:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1085.eqiad.wmnet with OS trixie * 11:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1086.eqiad.wmnet with reason: host reimage * 11:10 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1085.eqiad.wmnet with reason: host reimage * 11:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2004.codfw.wmnet with OS trixie * 11:03 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1086.eqiad.wmnet with reason: host reimage * 11:02 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1085.eqiad.wmnet with reason: host reimage * 10:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2004.codfw.wmnet with reason: host reimage * 10:48 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:46 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1086.eqiad.wmnet with OS trixie * 10:46 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1085.eqiad.wmnet with OS trixie * 10:44 cgoubert@deploy1003: Finished deploy [restbase/deploy@2fc37d4]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] (duration: 16m 44s) * 10:43 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2004.codfw.wmnet with reason: host reimage * 10:35 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:27 cgoubert@deploy1003: Started deploy [restbase/deploy@2fc37d4]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] * 10:27 cgoubert@deploy1003: Finished deploy [restbase/deploy@8a25036]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] (duration: 00m 45s) * 10:26 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1117.eqiad.wmnet with OS trixie * 10:26 cgoubert@deploy1003: Started deploy [restbase/deploy@8a25036]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] * 10:26 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host thanos-fe2004 * 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host thanos-fe2004 * 10:22 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1116.eqiad.wmnet with OS trixie * 10:21 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host thanos-fe2004 * 10:21 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) thanos-fe2004.codfw.wmnet 157.32.192.10.in-addr.arpa 7.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:20 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache thanos-fe2004.codfw.wmnet 157.32.192.10.in-addr.arpa 7.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:20 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:20 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host thanos-fe2004 - mvernon@cumin2003" * 10:20 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host thanos-fe2004 - mvernon@cumin2003" * 10:15 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2252: Repooling after reboot * 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:15 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 10:15 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2252: Repooling after reboot * 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1153.eqiad.wmnet * 10:14 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1153.eqiad.wmnet * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 10:14 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 10:12 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 10:12 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host thanos-fe2004 * 10:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2004.codfw.wmnet with OS trixie * 10:07 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1117.eqiad.wmnet with reason: host reimage * 10:03 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1116.eqiad.wmnet with reason: host reimage * 09:58 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:58 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1117.eqiad.wmnet with reason: host reimage * 09:57 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1116.eqiad.wmnet with reason: host reimage * 09:49 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 41 days, 15:00:00 on db2252.codfw.wmnet with reason: Security updates * 09:45 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1117.eqiad.wmnet with OS trixie * 09:45 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1116.eqiad.wmnet with OS trixie * 09:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1153: Security updates * 09:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:28 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:28 root@cumin1003: START - Cookbook sre.mysql.depool depool db1153: Security updates * 09:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1016: Security updates * 09:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:21 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:21 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1016: Security updates * 09:14 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:14 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 08:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1016: Security updates * 08:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:56 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:56 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1016: Security updates * 08:50 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:50 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:45 filippo@dns1006: END - running authdns-update * 08:43 filippo@dns1006: START - running authdns-update * 08:42 godog: switch dumps-nfs address to be shared with rsync/http - [[phab:T411248|T411248]] * 08:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1016: Security updates * 08:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:40 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:40 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1016: Security updates * 08:29 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host cirrussearch1111.eqiad.wmnet * 08:29 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:27 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:27 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:25 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1015: Security updates * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:09 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:09 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1015: Security updates * 07:42 Msz2001: Deployed private patch for Suggested Ivestigations * 07:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1015: Security updates * 07:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:41 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:41 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1015: Security updates * 07:40 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 07:11 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fingerprint warnings - oblivian@cumin1003" * 07:11 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fingerprint warnings - oblivian@cumin1003 * 07:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1024: Security updates * 07:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:11 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:11 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1024: Security updates * 07:10 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fingerprint warnings - oblivian@cumin1003 * 07:10 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fingerprint warnings - oblivian@cumin1003" * 06:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host cirrussearch1111.eqiad.wmnet * 06:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 06:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1024: Security updates * 06:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 06:48 root@cumin1003: START - Cookbook sre.mysql.parsercache * 06:48 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1024: Security updates * 06:42 moritzm: install nginx security updates * 06:31 root@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool pc1024: Security updates * 06:21 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1024: Security updates * 06:19 moritzm: installing php8.2 security updates * 06:15 moritzm: installing php8.4 security updates * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.7 (duration: 02m 38s) * 03:40 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] (duration: 37m 04s) * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 51s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-06 == * 23:30 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] (duration: 09m 39s) * 23:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1078.eqiad.wmnet with OS trixie * 23:26 jdlrobson@deploy1003: jdlrobson, bwang: Continuing with deployment * 23:22 jdlrobson@deploy1003: jdlrobson, bwang: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug) * 23:21 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] * 23:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1078.eqiad.wmnet with reason: host reimage * 23:06 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1078.eqiad.wmnet with reason: host reimage * 22:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1078.eqiad.wmnet with OS trixie * 22:29 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on cirrussearch1114.eqiad.wmnet with reason: reimage on hold until restore completes * 22:22 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on cirrussearch[1079,1115].eqiad.wmnet with reason: reimage on hold until restore completes * 21:18 maryum: Deployed security fix for [[phab:T428006|T428006]] * 20:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1079.eqiad.wmnet with OS trixie * 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1077.eqiad.wmnet with OS trixie * 20:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1115.eqiad.wmnet with OS trixie * 20:25 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1079.eqiad.wmnet with reason: host reimage * 20:21 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1079.eqiad.wmnet with reason: host reimage * 20:15 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] (duration: 08m 14s) * 20:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1077.eqiad.wmnet with reason: host reimage * 20:10 krinkle@deploy1003: krinkle, pushpaktiwari: Continuing with deployment * 20:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1115.eqiad.wmnet with reason: host reimage * 20:08 krinkle@deploy1003: krinkle, pushpaktiwari: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1077.eqiad.wmnet with reason: host reimage * 20:06 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] * 20:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1079.eqiad.wmnet with OS trixie * 20:04 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1115.eqiad.wmnet with reason: host reimage * 19:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1077.eqiad.wmnet with OS trixie * 19:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1115.eqiad.wmnet with OS trixie * 19:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 19:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 18:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1114.eqiad.wmnet with OS trixie * 18:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1114.eqiad.wmnet with reason: host reimage * 18:35 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1114.eqiad.wmnet with reason: host reimage * 18:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1112.eqiad.wmnet with OS trixie * 18:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1114.eqiad.wmnet with OS trixie * 18:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1072.eqiad.wmnet with OS trixie * 18:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1112.eqiad.wmnet with reason: host reimage * 18:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1112.eqiad.wmnet with reason: host reimage * 17:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1072.eqiad.wmnet with reason: host reimage * 17:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1112.eqiad.wmnet with OS trixie * 17:55 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1072.eqiad.wmnet with reason: host reimage * 17:39 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1072.eqiad.wmnet with OS trixie * 17:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1071.eqiad.wmnet with OS trixie * 17:18 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1070.eqiad.wmnet with OS trixie * 17:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1084.eqiad.wmnet with OS trixie * 16:54 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1071.eqiad.wmnet with reason: host reimage * 16:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1084.eqiad.wmnet with reason: host reimage * 16:51 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1070.eqiad.wmnet with reason: host reimage * 16:49 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1084.eqiad.wmnet with reason: host reimage * 16:38 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1071.eqiad.wmnet with OS trixie * 16:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1096.eqiad.wmnet with OS trixie * 16:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1070.eqiad.wmnet with OS trixie * 16:33 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1084.eqiad.wmnet with OS trixie * 16:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1089.eqiad.wmnet with OS trixie * 16:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1103.eqiad.wmnet with OS trixie * 16:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1096.eqiad.wmnet with reason: host reimage * 16:14 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1096.eqiad.wmnet with reason: host reimage * 16:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1089.eqiad.wmnet with reason: host reimage * 16:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1103.eqiad.wmnet with reason: host reimage * 16:02 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1003.eqiad.wmnet with OS bookworm * 16:01 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1089.eqiad.wmnet with reason: host reimage * 16:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1103.eqiad.wmnet with reason: host reimage * 15:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1096.eqiad.wmnet with OS trixie * 15:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1080.eqiad.wmnet with OS trixie * 15:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1089.eqiad.wmnet with OS trixie * 15:45 dancy@deploy1003: Installation of scap version "4.272.0" completed for 158 hosts * 15:43 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1103.eqiad.wmnet with OS trixie * 15:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1113.eqiad.wmnet with OS trixie * 15:41 dancy@deploy1003: Installing scap version "4.272.0" for 158 host(s) * 15:40 klausman@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 15:39 klausman@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 15:38 klausman@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 15:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1069.eqiad.wmnet with OS trixie * 15:37 klausman@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 15:36 klausman@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 15:34 klausman@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 15:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1080.eqiad.wmnet with reason: host reimage * 15:27 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1080.eqiad.wmnet with reason: host reimage * 15:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1113.eqiad.wmnet with reason: host reimage * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1069.eqiad.wmnet with reason: host reimage * 15:18 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1113.eqiad.wmnet with reason: host reimage * 15:16 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1069.eqiad.wmnet with reason: host reimage * 15:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:11 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1080.eqiad.wmnet with OS trixie * 15:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1113.eqiad.wmnet with OS trixie * 15:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1003.eqiad.wmnet with reason: host reimage * 14:47 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1003.eqiad.wmnet with OS bookworm * 14:33 elukey: rolled out spicerack on all cumin nodes - [[phab:T429699|T429699]] * 14:32 elukey: upgrade all bookworm hosts to pywmflib 3.1 - [[phab:T430552|T430552]] * 14:14 marostegui: Setup x4 eqiad topology [[phab:T404715|T404715]] * 14:13 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 14:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2230.codfw.wmnet * 14:07 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2230.codfw.wmnet * 13:59 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[2001-2002].codfw.wmnet * 13:51 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 13:45 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.major-upgrade (exit_code=97) * 13:45 cwilliams@cumin1003: dbmaint on s4@codfw [[phab:T429893|T429893]] * 13:45 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 13:42 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-master-codfw * 13:42 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl2002.codfw.wmnet * 13:42 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl2002.codfw.wmnet * 13:38 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl2002.codfw.wmnet * 13:38 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl2002.codfw.wmnet * 13:38 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl2001.codfw.wmnet * 13:38 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl2001.codfw.wmnet * 13:35 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl2001.codfw.wmnet * 13:35 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl2001.codfw.wmnet * 13:35 klausman@cumin2002: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-master-codfw * 12:30 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] (duration: 25m 11s) * 12:24 krinkle@deploy1003: krinkle: Continuing with deployment * 12:10 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2048.codfw.wmnet * 12:09 krinkle@deploy1003: krinkle: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:08 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2048.codfw.wmnet * 12:05 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] * 11:57 moritzm: installing curl security updates * 11:49 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:31 moritzm: installing nano security updates * 11:07 moritzm: failover Ganeti master in codfw to ganeti2032 [[phab:T430909|T430909]] * 11:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:04 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest1005.eqiad.wmnet with OS trixie * 11:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:50 jmm@dns1004: END - running authdns-update * 10:47 jmm@dns1004: START - running authdns-update * 10:47 jmm@dns1004: START - running authdns-update * 10:46 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:44 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest1005.eqiad.wmnet with reason: host reimage * 10:38 elukey@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest1005.eqiad.wmnet with reason: host reimage * 10:31 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:31 marostegui: Setup x4 codfw topology [[phab:T404715|T404715]] * 10:31 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 10:24 elukey: spicerack 13.0.0 deployed on cumin2002 * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 10:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 10:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 10:21 elukey@cumin2002: START - Cookbook sre.hosts.reimage for host sretest1005.eqiad.wmnet with OS trixie * 10:20 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:19 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:17 elukey: uploaded spicerack_13.0.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia * 09:54 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:52 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:20 elukey: upgrade all bullseye hosts to pywmflib 3.1 - [[phab:T430552|T430552]] * 09:10 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1015.eqiad.wmnet,service=s4 * 09:10 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1015.eqiad.wmnet,service=s6 * 09:07 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 08:58 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:56 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 08:56 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 08:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 08:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 08:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin2002.codfw.wmnet * 08:06 godog: remove cloudvirt1046, cloudvirt1062, cloudvirt1074, cloudvirt1075 from maintenance aggregate and put them in network-ovs - [[phab:T424802|T424802]] * 08:00 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin2002.codfw.wmnet * 07:58 hashar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] (duration: 32m 53s) * 07:57 fabfur: repooled cp4038 * 07:57 fabfur@cumin1003: conftool action : set/pooled=yes; selector: name=cp4038.* * 07:53 moritzm: installing pyjwt security updates * 07:47 moritzm: installing openjpeg2 security updates * 07:45 hashar@deploy1003: vadymts1, hashar: Continuing with deployment * 07:43 hashar@deploy1003: vadymts1, hashar: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:38 moritzm: installing python-urllib3 security updates * 07:37 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 07:30 fabfur: depooled cp4038 to investigate on possible maxmind failure * 07:30 fabfur@cumin1003: conftool action : set/pooled=no; selector: name=cp4038.* * 07:30 fabfur@cumin1003: conftool action : set/pooled=yes; selector: name=cp4038.* * 07:29 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 07:25 hashar@deploy1003: Started scap sync-world: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] * 06:13 moritzm: installing Linux 6.12.95 on trixie hosts * 05:20 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s6 * 05:20 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s4 * 05:19 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1015.eqiad.wmnet with reason: cloning * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 08s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-05 == * 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 01m 08s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-04 == * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 58s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-03 == * 17:08 topranks: revert protocol preference changes on cr3-ulsfo after upgrade * 16:53 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on cr2-eqord with reason: upgrade JunOS cr3-ulsfo * 16:53 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on cr4-ulsfo with reason: upgrade JunOS cr3-ulsfo * 16:48 topranks: reboot cr3-ulsfo to upgrade JunOS and reset linecard [[phab:T424839|T424839]] * 15:52 topranks: adjust outbound BGP policies on cr3-ulsfo to drain router of traffic [[phab:T424839|T424839]] * 15:45 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on lvs[4008-4010].ulsfo.wmnet with reason: upgrade JunOS cr3-ulsfo * 15:44 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on asw1-[22-23]-ulsfo,cr3-ulsfo,cr3-ulsfo IPv6,cr3-ulsfo.mgmt with reason: upgrade JunOS cr3-ulsfo * 15:36 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 15:35 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 15:35 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 14:40 cmooney@dns3003: END - running authdns-update * 14:26 cmooney@dns3003: START - running authdns-update * 14:26 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:26 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to ulsfo - cmooney@cumin1003" * 14:19 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to ulsfo - cmooney@cumin1003" * 14:16 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:38 sukhe@dns1004: END - running authdns-update * 13:35 sukhe@dns1004: START - running authdns-update * 13:26 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 13:26 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 13:26 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet * 13:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 13:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 13:16 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host sretest1005.eqiad.wmnet * 13:16 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 13:16 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 13:15 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 13:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:14 moritzm: imported samplicator 1.3.8rc1-1+deb13u1 to trixie-wikimedia/main [[phab:T337208|T337208]] * 13:13 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:07 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 13:07 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 13:02 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:02 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:58 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:57 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:57 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:53 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet * 12:50 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 12:47 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:41 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:40 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:39 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:32 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet * 12:26 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet * 12:23 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2005.wikimedia.org * 12:19 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2005.wikimedia.org * 12:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 12:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup[2004-2007].codfw.wmnet * 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[2004-2007].codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin2003" * 12:15 jynus@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[2004-2007].codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin2003" * 12:09 jynus@cumin2003: START - Cookbook sre.dns.netbox * 11:58 jynus@cumin2003: START - Cookbook sre.hosts.decommission for hosts backup[2004-2007].codfw.wmnet * 10:40 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup[1004-1007].eqiad.wmnet * 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[1004-1007].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 10:01 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[1004-1007].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 09:52 jynus@cumin1003: START - Cookbook sre.dns.netbox * 09:39 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:36 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup[1004-1007].eqiad.wmnet * 09:36 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:25 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 09:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 09:16 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 09:05 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:04 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:00 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:59 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:57 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:55 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:50 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 08:50 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 08:49 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 08:49 atsukoito: depooling cirrussearch in codfw because of regression after upgrade [[phab:T431091|T431091]] * 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts mirror1001.wikimedia.org * 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: mirror1001.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 08:29 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: mirror1001.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 08:18 jmm@cumin2003: START - Cookbook sre.dns.netbox * 08:11 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts mirror1001.wikimedia.org * 06:15 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 18s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-02 == * 22:55 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host contint1003.wikimedia.org with OS trixie * 22:29 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on contint1003.wikimedia.org with reason: host reimage * 22:23 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on contint1003.wikimedia.org with reason: host reimage * 22:05 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host contint1003.wikimedia.org with OS trixie * 22:03 mutante: contint1003 (zuul.wikimedia.org) - reimaging because of [[phab:T430510|T430510]]#12067628 [[phab:T418521|T418521]] * 22:03 dzahn@cumin2002: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on zuul.wikimedia.org with reason: reimage * 21:39 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 18s) * 21:39 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 21:20 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1003.eqiad.wmnet, repooling source-only afterwards * 21:19 sbassett: Deployed security fix for [[phab:T428829|T428829]] * 20:58 cmooney@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Release v0.11.2 update for new Aerleon - cmooney@cumin1003 * 20:55 cmooney@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Release v0.11.2 update for new Aerleon - cmooney@cumin1003 * 20:40 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] (duration: 12m 35s) * 20:36 arlolra@deploy1003: cscott, arlolra: Continuing with deployment * 20:35 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 20s) * 20:35 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 20:33 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host contint2003.wikimedia.org with OS trixie * 20:31 arlolra@deploy1003: cscott, arlolra: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Cha * 20:28 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] * 20:17 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] (duration: 08m 13s) * 20:14 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on contint2003.wikimedia.org with reason: host reimage * 20:13 sbassett@deploy1003: sbassett: Continuing with deployment * 20:11 sbassett@deploy1003: sbassett: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:09 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] * 20:08 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 20:08 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 20:08 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on contint2003.wikimedia.org with reason: host reimage * 20:05 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1003.eqiad.wmnet, repooling source-only afterwards * 19:49 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host contint2003.wikimedia.org with OS trixie * 19:48 mutante: contint2003 - reimaging because of [[phab:T430510|T430510]]#12067628 [[phab:T418521|T418521]] * 18:39 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 18:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 18:13 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2002.codfw.wmnet -> wcqs2003.codfw.wmnet, repooling source-only afterwards * 17:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1003.eqiad.wmnet with OS bookworm * 17:52 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1005.eqiad.wmnet * 17:52 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1005.eqiad.wmnet * 17:51 jasmine@cumin2002: conftool action : set/pooled=yes:weight=10; selector: name=wikikube-ctrl1005.eqiad.wmnet * 17:48 jasmine_: homer "cr*eqiad*" commit "Added new stacked control plane wikikube-ctrl1005" * 17:44 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply * 17:44 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply * 17:31 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] (duration: 09m 33s) * 17:26 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 17:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1003.eqiad.wmnet with reason: host reimage * 17:23 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:21 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] * 17:18 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1003.eqiad.wmnet with reason: host reimage * 17:16 rscout@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply * 17:16 rscout@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply * 17:16 rscout@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply * 17:15 rscout@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply * 17:12 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on wcqs[2002-2003].codfw.wmnet,wcqs1002.eqiad.wmnet with reason: reimaging hosts * 17:08 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 17:08 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 17:08 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 17:07 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 17:05 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 17:05 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 17:03 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "running to make sure all updates are synced - cmooney@cumin1003" * 17:03 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "running to make sure all updates are synced - cmooney@cumin1003" * 17:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs1003 * 17:00 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs1003 * 17:00 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 17:00 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1003.eqiad.wmnet with OS bookworm * 16:58 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Re-running - btullis@cumin1003" * 16:58 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Re-running - btullis@cumin1003" * 16:58 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2002.codfw.wmnet -> wcqs2003.codfw.wmnet, repooling source-only afterwards * 16:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-master1004.eqiad.wmnet with OS bookworm * 16:58 btullis@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 16:57 tappof: bump space for prometheus k8s-aux in eqiad * 16:55 cmooney@dns3003: END - running authdns-update * 16:55 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:55 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to eqsin - cmooney@cumin1003" * 16:55 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to eqsin - cmooney@cumin1003" * 16:53 cmooney@dns3003: START - running authdns-update * 16:52 ryankemper: [ml-serve-eqiad] Cleared out 1302 failed (Evicted) pods: `kubectl -n llm delete pods --field-selector=status.phase=Failed`, freeing calico-kube-controllers from OOM crashloop (evictions were caused by disk pressure) * 16:49 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 16:46 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:39 rzl@dns1004: END - running authdns-update * 16:37 rzl@dns1004: START - running authdns-update * 16:36 rzl@dns1004: START - running authdns-update * 16:35 rzl@deploy1003: Finished scap sync-world: [[phab:T416623|T416623]] (duration: 10m 19s) * 16:34 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 16:33 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-master1004.eqiad.wmnet with reason: host reimage * 16:30 rzl@deploy1003: rzl: Continuing with deployment * 16:28 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-master1004.eqiad.wmnet with reason: host reimage * 16:26 rzl@deploy1003: rzl: [[phab:T416623|T416623]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:25 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 16:25 rzl@deploy1003: Started scap sync-world: [[phab:T416623|T416623]] * 16:25 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 16:24 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 16:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: sync * 16:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: sync * 16:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-master1004.eqiad.wmnet with OS bookworm * 16:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-master1003.eqiad.wmnet with OS bookworm * 16:11 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 16:11 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 16:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Security updates * 16:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 16:08 root@cumin1003: START - Cookbook sre.mysql.parsercache * 16:08 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Security updates * 15:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-master1003.eqiad.wmnet with reason: host reimage * 15:54 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:54 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:54 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:54 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-master1003.eqiad.wmnet with reason: host reimage * 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Security updates * 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:45 root@cumin1003: START - Cookbook sre.mysql.parsercache * 15:45 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Security updates * 15:42 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-master1003.eqiad.wmnet with OS bookworm * 15:24 moritzm: installing busybox updates from bookworm point release * 15:20 moritzm: installing busybox updates from trixie point release * 15:15 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1021: Security updates * 15:15 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:15 root@cumin1003: START - Cookbook sre.mysql.parsercache * 15:15 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1021: Security updates * 15:13 moritzm: installing giflib security updates * 15:08 moritzm: installing Tomcat security updates * 14:57 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 14:56 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 14:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:53 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Unblock taavi - oblivian@cumin1003" * 14:53 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Unblock taavi - oblivian@cumin1003 * 14:53 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1021: Security updates * 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:53 root@cumin1003: START - Cookbook sre.mysql.parsercache * 14:53 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1021: Security updates * 14:53 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Unblock taavi - oblivian@cumin1003 * 14:52 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Unblock taavi - oblivian@cumin1003" * 14:46 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94711 and previous config saved to /var/cache/conftool/dbconfig/20260702-144644-fceratto.json * 14:36 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205', diff saved to https://phabricator.wikimedia.org/P94709 and previous config saved to /var/cache/conftool/dbconfig/20260702-143636-fceratto.json * 14:32 moritzm: installing libdbi-perl security updates * 14:26 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205', diff saved to https://phabricator.wikimedia.org/P94708 and previous config saved to /var/cache/conftool/dbconfig/20260702-142628-fceratto.json * 14:16 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94707 and previous config saved to /var/cache/conftool/dbconfig/20260702-141621-fceratto.json * 14:12 moritzm: installing rsync security updates * 14:11 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox) * 14:10 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94706 and previous config saved to /var/cache/conftool/dbconfig/20260702-140959-fceratto.json * 14:09 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2205.codfw.wmnet with reason: Maintenance * 14:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2205: Repooling after switchover * 14:07 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-test-master1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 14:06 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 14:06 Tran: Deployed patch for [[phab:T427287|T427287]] * 14:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:59 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2205: Repooling after switchover * 13:59 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2205: Repooling after switchover * 13:59 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:55 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2205: Repooling after switchover * 13:55 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2205 [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94704 and previous config saved to /var/cache/conftool/dbconfig/20260702-135505-fceratto.json * 13:54 moritzm: installing sed security updates * 13:53 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:52 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2209 to s3 primary [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94703 and previous config saved to /var/cache/conftool/dbconfig/20260702-135235-fceratto.json * 13:52 federico3: Starting s3 codfw failover from db2205 to db2209 - [[phab:T430912|T430912]] * 13:51 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:51 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 13:48 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:47 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2209 with weight 0 [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94702 and previous config saved to /var/cache/conftool/dbconfig/20260702-134719-fceratto.json * 13:47 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Primary switchover s3 [[phab:T430912|T430912]] * 13:44 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:44 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:44 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:40 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 13:38 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 13:37 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 13:36 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 13:36 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:34 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 13:30 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:29 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:29 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:27 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:26 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:25 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling restart_daemons on A:wikidough * 13:23 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 13:22 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 13:17 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 13:17 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns1004.wikimedia.org * 13:12 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:11 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart (exit_code=97) rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough * 13:11 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=97) rolling restart_daemons on A:wikidough * 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough * 13:09 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] (duration: 07m 20s) * 13:05 aude@deploy1003: jdrewniak, aude: Continuing with deployment * 13:04 aude@deploy1003: jdrewniak, aude: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:02 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] * 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts wdqs-categories1001.eqiad.wmnet * 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: wdqs-categories1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 12:10 jmm@dns1004: END - running authdns-update * 12:07 jmm@dns1004: START - running authdns-update * 11:51 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: wdqs-categories1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 11:44 btullis@cumin1003: START - Cookbook sre.dns.netbox * 11:42 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 11:42 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 11:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet * 11:39 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts wdqs-categories1001.eqiad.wmnet * 11:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet * 11:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet * 11:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet * 11:29 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 11:29 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 10:57 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2214: Repooling * 10:49 jmm@dns1004: END - running authdns-update * 10:47 jmm@dns1004: START - running authdns-update * 10:31 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94698 and previous config saved to /var/cache/conftool/dbconfig/20260702-103146-fceratto.json * 10:21 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213', diff saved to https://phabricator.wikimedia.org/P94696 and previous config saved to /var/cache/conftool/dbconfig/20260702-102137-fceratto.json * 10:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:19 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb1017.eqiad.wmnet * 10:18 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 10:18 fceratto@cumin1003: Removing es1033 from zarcillo [[phab:T408772|T408772]] * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts es1033.eqiad.wmnet * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: es1033.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:14 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: es1033.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:13 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb1017.eqiad.wmnet * 10:12 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2214.codfw.wmnet * 10:12 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2214.codfw.wmnet * 10:12 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2214: Repooling * 10:11 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213', diff saved to https://phabricator.wikimedia.org/P94693 and previous config saved to /var/cache/conftool/dbconfig/20260702-101130-fceratto.json * 10:10 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:10 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:03 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts es1033.eqiad.wmnet * 10:03 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 10:01 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94691 and previous config saved to /var/cache/conftool/dbconfig/20260702-100122-fceratto.json * 09:55 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94690 and previous config saved to /var/cache/conftool/dbconfig/20260702-095529-fceratto.json * 09:55 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2213.codfw.wmnet with reason: Maintenance * 09:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 09:53 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2213: Repooling after switchover * 09:51 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover * 09:44 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2213: Repooling after switchover * 09:39 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover * 09:39 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2213 [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94688 and previous config saved to /var/cache/conftool/dbconfig/20260702-093859-fceratto.json * 09:36 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2192 to s5 primary [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94687 and previous config saved to /var/cache/conftool/dbconfig/20260702-093650-fceratto.json * 09:36 federico3: Starting s5 codfw failover from db2213 to db2192 - [[phab:T430923|T430923]] * 09:30 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94686 and previous config saved to /var/cache/conftool/dbconfig/20260702-093004-fceratto.json * 09:24 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2192 with weight 0 [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94685 and previous config saved to /var/cache/conftool/dbconfig/20260702-092455-fceratto.json * 09:24 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 23 hosts with reason: Primary switchover s5 [[phab:T430923|T430923]] * 09:19 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220', diff saved to https://phabricator.wikimedia.org/P94684 and previous config saved to /var/cache/conftool/dbconfig/20260702-091957-fceratto.json * 09:16 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] (duration: 06m 57s) * 09:13 moritzm: installing libgcrypt20 security updates * 09:12 kharlan@deploy1003: kharlan: Continuing with deployment * 09:11 kharlan@deploy1003: kharlan: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:09 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220', diff saved to https://phabricator.wikimedia.org/P94683 and previous config saved to /var/cache/conftool/dbconfig/20260702-090950-fceratto.json * 09:09 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] * 09:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 09:01 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] (duration: 07m 07s) * 08:59 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94682 and previous config saved to /var/cache/conftool/dbconfig/20260702-085942-fceratto.json * 08:57 kharlan@deploy1003: kharlan: Continuing with deployment * 08:56 kharlan@deploy1003: kharlan: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:54 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] * 08:52 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:52 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:52 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94681 and previous config saved to /var/cache/conftool/dbconfig/20260702-085237-fceratto.json * 08:52 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2220.codfw.wmnet with reason: Maintenance * 08:43 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:40 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 08:25 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] (duration: 11m 44s) * 08:21 cscott@deploy1003: cscott: Continuing with deployment * 08:16 cscott@deploy1003: cscott: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:14 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] * 08:08 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 08:08 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1244: Migration of db1244.eqiad.wmnet completed * 08:02 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:02 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:01 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] (duration: 18m 58s) * 08:01 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:59 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 07:59 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:59 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:59 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2006.wikimedia.org * 07:58 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:57 cscott@deploy1003: cscott: Continuing with deployment * 07:56 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:56 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:56 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:55 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:55 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:55 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:54 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2006.wikimedia.org * 07:54 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:54 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 07:44 cscott@deploy1003: cscott: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:44 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2005.wikimedia.org * 07:44 moritzm: installing node-lodash security updates * 07:42 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] * 07:39 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2005.wikimedia.org * 07:30 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] (duration: 07m 28s) * 07:26 cscott@deploy1003: ssastry, cscott: Continuing with deployment * 07:25 cscott@deploy1003: ssastry, cscott: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:23 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1244: Migration of db1244.eqiad.wmnet completed * 07:22 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] * 07:16 wmde-fisch@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] (duration: 06m 55s) * 07:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1244.eqiad.wmnet with OS trixie * 07:11 wmde-fisch@deploy1003: wmde-fisch: Continuing with deployment * 07:11 wmde-fisch@deploy1003: wmde-fisch: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:09 wmde-fisch@deploy1003: Started scap sync-world: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] * 06:54 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1244.eqiad.wmnet with reason: host reimage * 06:50 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1244.eqiad.wmnet with reason: host reimage * 06:38 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1250.eqiad.wmnet with OS trixie * 06:34 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db1244.eqiad.wmnet with OS trixie * 06:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1244: Upgrading db1244.eqiad.wmnet * 06:25 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1244: Upgrading db1244.eqiad.wmnet * 06:25 cwilliams@cumin1003: dbmaint on s4@eqiad [[phab:T429893|T429893]] * 06:25 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 06:15 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1250.eqiad.wmnet with reason: host reimage * 06:14 cwilliams@dns1006: END - running authdns-update * 06:12 cwilliams@dns1006: START - running authdns-update * 06:11 cwilliams@dns1006: END - running authdns-update * 06:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db1244 [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94676 and previous config saved to /var/cache/conftool/dbconfig/20260702-061059-cwilliams.json * 06:09 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1250.eqiad.wmnet with reason: host reimage * 06:09 cwilliams@dns1006: START - running authdns-update * 06:08 aokoth@cumin1003: END (PASS) - Cookbook sre.vrts.upgrade (exit_code=0) on VRTS host vrts1003.eqiad.wmnet * 06:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db1160 to s4 primary and set section read-write [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94675 and previous config saved to /var/cache/conftool/dbconfig/20260702-060746-cwilliams.json * 06:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Set s4 eqiad as read-only for maintenance - [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94674 and previous config saved to /var/cache/conftool/dbconfig/20260702-060704-cwilliams.json * 06:06 cezmunsta: Starting s4 eqiad failover from db1244 to db1160 - [[phab:T430817|T430817]] * 06:04 aokoth@cumin1003: START - Cookbook sre.vrts.upgrade on VRTS host vrts1003.eqiad.wmnet * 05:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db1160 with weight 0 [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94673 and previous config saved to /var/cache/conftool/dbconfig/20260702-055927-cwilliams.json * 05:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 40 hosts with reason: Primary switchover s4 [[phab:T430817|T430817]] * 05:55 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1250.eqiad.wmnet with OS trixie * 05:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on db1250.eqiad.wmnet with reason: m3 master switchover [[phab:T430158|T430158]] * 05:39 marostegui: Failover m3 (phabricator) from db1250 to db1228 - [[phab:T430158|T430158]] * 05:32 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2234].codfw.wmnet,db[1217,1228,1250].eqiad.wmnet with reason: m3 master switchover [[phab:T430158|T430158]] * 04:45 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] (duration: 09m 08s) * 04:41 tstarling@deploy1003: tstarling, reedy: Continuing with deployment * 04:38 tstarling@deploy1003: tstarling, reedy: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 04:36 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 59s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:16 ryankemper: [[phab:T429844|T429844]] [opensearch] completed `cirrussearch2111` reimage; all codfw search clusters are green, all nodes now report `OpenSearch 2.19.5`, and the temporary chi voting exclusion has been removed * 00:57 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2111.codfw.wmnet with OS trixie * 00:29 ryankemper: [[phab:T429844|T429844]] [opensearch] depooled codfw search-omega/search-psi discovery records to match existing codfw search depool during OpenSearch 2.19 migration * 00:29 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2111.codfw.wmnet with reason: host reimage * 00:29 ryankemper@cumin2002: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 00:29 ryankemper@cumin2002: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 00:22 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2111.codfw.wmnet with reason: host reimage * 00:01 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2111.codfw.wmnet with OS trixie * 00:00 ryankemper: [[phab:T429844|T429844]] [opensearch] chi cluster recovered after stopping `opensearch_1@production-search-codfw` on `cirrussearch2111` == 2026-07-01 == * 23:59 ryankemper: [[phab:T429844|T429844]] [opensearch] stopped `opensearch_1@production-search-codfw` on `cirrussearch2111` after chi cluster-manager election churn following `voting_config_exclusions` POST; hoping this triggers a re-election * 23:52 cscott@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 23:51 cscott@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 23:51 cscott@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 23:50 cscott@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2003.codfw.wmnet with OS bookworm * 22:29 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 22:13 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 22:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2084.codfw.wmnet with OS trixie * 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2003.codfw.wmnet with reason: host reimage * 22:03 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 22:01 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2003.codfw.wmnet with reason: host reimage * 21:50 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 21:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2084.codfw.wmnet with reason: host reimage * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2003 * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2003 * 21:42 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2003 * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2003.codfw.wmnet 45.48.192.10.in-addr.arpa 5.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:42 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2003.codfw.wmnet 45.48.192.10.in-addr.arpa 5.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2003 - bking@cumin2003" * 21:42 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2003 - bking@cumin2003" * 21:36 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2084.codfw.wmnet with reason: host reimage * 21:35 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:34 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2003 * 21:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2003.codfw.wmnet with OS bookworm * 21:19 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2084.codfw.wmnet with OS trixie * 21:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2081.codfw.wmnet with OS trixie * 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2108.codfw.wmnet with OS trixie * 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2081.codfw.wmnet with reason: host reimage * 20:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2081.codfw.wmnet with reason: host reimage * 20:28 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2081.codfw.wmnet with OS trixie * 20:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2108.codfw.wmnet with reason: host reimage * 20:19 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2108.codfw.wmnet with reason: host reimage * 19:59 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2108.codfw.wmnet with OS trixie * 19:46 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2093.codfw.wmnet with OS trixie * 19:44 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 19:44 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jasmine@cumin2002" * 19:43 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jasmine@cumin2002" * 19:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2080.codfw.wmnet with OS trixie * 19:28 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 19:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2093.codfw.wmnet with reason: host reimage * 19:18 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 19:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2093.codfw.wmnet with reason: host reimage * 19:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2080.codfw.wmnet with reason: host reimage * 19:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2080.codfw.wmnet with reason: host reimage * 18:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2093.codfw.wmnet with OS trixie * 18:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2080.codfw.wmnet with OS trixie * 18:27 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 18:18 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] (duration: 09m 15s) * 18:13 jgiannelos@deploy1003: jgiannelos, neriah: Continuing with deployment * 18:11 jgiannelos@deploy1003: jgiannelos, neriah: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:09 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] * 17:40 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 16:58 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 30 hosts * 16:57 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for 30 hosts * 16:52 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2202.codfw.wmnet * 16:52 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2202.codfw.wmnet * 16:51 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt * 16:51 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt * 16:51 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lvs2012.codfw.wmnet * 16:51 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for lvs2012.codfw.wmnet * 16:49 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2076.codfw.wmnet with OS trixie * 16:49 brett: Start pybal on lvs2012 - [[phab:T429861|T429861]] * 16:49 pt1979@cumin1003: END (ERROR) - Cookbook sre.hosts.remove-downtime (exit_code=97) for 59 hosts * 16:48 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for 59 hosts * 16:42 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2061.codfw.wmnet with OS trixie * 16:30 dancy@deploy1003: Installation of scap version "4.271.0" completed for 2 hosts * 16:28 dancy@deploy1003: Installing scap version "4.271.0" for 2 host(s) * 16:23 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2076.codfw.wmnet with reason: host reimage * 16:19 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2061.codfw.wmnet with reason: host reimage * 16:18 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2076.codfw.wmnet with reason: host reimage * 16:18 jasmine@dns1004: END - running authdns-update * 16:16 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host restbase2039.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 16:16 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host restbase2039.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 16:16 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2061.codfw.wmnet with reason: host reimage * 16:15 jasmine@dns1004: START - running authdns-update * 16:14 jasmine@dns1004: END - running authdns-update * 16:12 jasmine@dns1004: START - running authdns-update * 16:07 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2202.codfw.wmnet with reason: maintenance * 16:06 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt with reason: Junos upograde * 16:00 papaul: ongoing maintenance on lsw1-b2-codfw * 16:00 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2076.codfw.wmnet with OS trixie * 15:59 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt * 15:59 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt * 15:57 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2061.codfw.wmnet with OS trixie * 15:55 pt1979@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2042,2046].codfw.wmnet * 15:55 pt1979@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2042,2046].codfw.wmnet * 15:51 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 15:51 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2220: Repooling after switchover * 15:50 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 15:50 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 15:48 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2092.codfw.wmnet with OS trixie * 15:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 15:40 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 15:38 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 15:37 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 15:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply * 15:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply * 15:32 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 15:32 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 15:30 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 15:29 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 15:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 15:25 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:22 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs2012.codfw.wmnet with reason: Rack B2 maintenance - [[phab:T429861|T429861]] * 15:21 brett: Stopping pybal on lvs2012 in preparation for codfw rack b2 maintenance - [[phab:T429861|T429861]] * 15:20 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2092.codfw.wmnet with reason: host reimage * 15:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:12 _joe_: restarted manually alertmanager-irc-relay * 15:12 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2092.codfw.wmnet with reason: host reimage * 15:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:12 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt with reason: Junos upograde * 15:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover * 15:07 pt1979@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2042,2046].codfw.wmnet * 15:06 pt1979@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2042,2046].codfw.wmnet * 15:02 papaul: ongoing maintenance on lsw1-a8-codfw * 14:31 topranks: POWERING DOWN CR1-EQIAD for line card installation [[phab:T426343|T426343]] * 14:31 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] (duration: 08m 57s) * 14:29 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:26 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 14:24 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:22 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] * 14:22 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover * 14:16 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:15 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover * 14:14 topranks: re-enable routing-engine graceful-failover on cr1-eqiad [[phab:T417873|T417873]] * 14:13 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:13 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2220: Repooling after switchover * 14:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:12 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:12 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:11 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:08 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] (duration: 10m 01s) * 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:07 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2220 [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94664 and previous config saved to /var/cache/conftool/dbconfig/20260701-140729-fceratto.json * 14:06 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:06 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:06 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:05 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2159 to s7 primary [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94663 and previous config saved to /var/cache/conftool/dbconfig/20260701-140503-fceratto.json * 14:04 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:04 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 14:04 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 14:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:04 dreamyjazz@deploy1003: anzx, dreamyjazz: Continuing with deployment * 14:04 federico3: Starting s7 codfw failover from db2220 to db2159 - [[phab:T430826|T430826]] * 14:03 jmm@dns1004: END - running authdns-update * 14:03 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:03 topranks: flipping cr1-eqiad active routing-enginer back to RE0 [[phab:T417873|T417873]] * 14:03 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudsw1-c8-eqiad,cloudsw1-d5-eqiad with reason: router upgrades eqiad * 14:01 jmm@dns1004: START - running authdns-update * 14:00 dreamyjazz@deploy1003: anzx, dreamyjazz: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:59 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2159 with weight 0 [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94662 and previous config saved to /var/cache/conftool/dbconfig/20260701-135906-fceratto.json * 13:58 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] * 13:57 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s7 [[phab:T430826|T430826]] * 13:56 topranks: reboot routing-enginer RE0 on cr1-eqiad [[phab:T417873|T417873]] * 13:48 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1006.wikimedia.org * 13:44 atsuko@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cirrussearch2092.codfw.wmnet with OS trixie * 13:43 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1006.wikimedia.org * 13:41 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2092.codfw.wmnet with OS trixie * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1005.wikimedia.org * 13:37 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1005.wikimedia.org * 13:37 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on pfw1-eqiad with reason: router upgrades eqiad * 13:35 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on lvs[1017-1020].eqiad.wmnet with reason: router upgrades eqiad * 13:34 caro@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] (duration: 07m 59s) * 13:30 caro@deploy1003: caro: Continuing with deployment * 13:28 caro@deploy1003: caro: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:27 topranks: route-engine failover cr1-eqiad * 13:26 caro@deploy1003: Started scap sync-world: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] * 13:15 topranks: rebooting routing-engine 1 on cr1-eqiad [[phab:T417873|T417873]] * 13:13 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] (duration: 08m 29s) * 13:13 moritzm: installing qemu security updates * 13:11 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 13:11 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 13:09 jgiannelos@deploy1003: jgiannelos: Continuing with deployment * 13:08 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 13:07 jgiannelos@deploy1003: jgiannelos: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:06 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 13:06 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2214.codfw.wmnet with reason: Maintenance * 13:05 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2214: Repooling after switchover * 13:05 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] * 13:04 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2214: Repooling after switchover * 13:04 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2214 [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94660 and previous config saved to /var/cache/conftool/dbconfig/20260701-130413-fceratto.json * 13:01 moritzm: installing python3.13 security updates * 13:00 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2229 to s6 primary [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94659 and previous config saved to /var/cache/conftool/dbconfig/20260701-125959-fceratto.json * 12:59 federico3: Starting s6 codfw failover from db2214 to db2229 - [[phab:T430814|T430814]] * 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on 13 hosts with reason: router upgrade and line card install * 12:51 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2229 with weight 0 [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94658 and previous config saved to /var/cache/conftool/dbconfig/20260701-125149-fceratto.json * 12:51 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 21 hosts with reason: Primary switchover s6 [[phab:T430814|T430814]] * 12:50 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2189.codfw.wmnet * 12:50 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2189.codfw.wmnet * 12:42 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2100.codfw.wmnet with OS trixie * 12:38 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2083.codfw.wmnet with OS trixie * 12:19 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2083.codfw.wmnet with reason: host reimage * 12:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 12:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2240: Migration of db2240.codfw.wmnet completed * 12:14 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2100.codfw.wmnet with reason: host reimage * 12:09 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2083.codfw.wmnet with reason: host reimage * 12:09 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2100.codfw.wmnet with reason: host reimage * 12:00 topranks: drain traffic on cr1-eqiad to allow for line card install and JunOS upgrade [[phab:T426343|T426343]] * 11:52 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2083.codfw.wmnet with OS trixie * 11:50 cmooney@dns2005: END - running authdns-update * 11:49 cmooney@dns2005: START - running authdns-update * 11:48 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2100.codfw.wmnet with OS trixie * 11:40 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/zotero: apply * 11:40 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/zotero: apply * 11:36 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/zotero: apply * 11:36 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/zotero: apply * 11:31 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2240: Migration of db2240.codfw.wmnet completed * 11:30 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply * 11:28 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply * 11:27 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:27 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:27 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:27 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:27 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:23 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2240.codfw.wmnet with OS trixie * 11:20 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:20 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:17 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:16 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:16 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:15 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2086.codfw.wmnet with OS trixie * 11:14 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2106.codfw.wmnet with OS trixie * 11:14 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:13 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:12 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:09 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2115.codfw.wmnet with OS trixie * 11:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2240.codfw.wmnet with reason: host reimage * 11:00 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2240.codfw.wmnet with reason: host reimage * 10:53 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2106.codfw.wmnet with reason: host reimage * 10:49 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2086.codfw.wmnet with reason: host reimage * 10:44 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2115.codfw.wmnet with reason: host reimage * 10:44 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2240.codfw.wmnet with OS trixie * 10:44 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2086.codfw.wmnet with reason: host reimage * 10:42 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2106.codfw.wmnet with reason: host reimage * 10:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2240: Upgrading db2240.codfw.wmnet * 10:41 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2240: Upgrading db2240.codfw.wmnet * 10:41 cwilliams@cumin1003: dbmaint on s4@codfw [[phab:T429893|T429893]] * 10:40 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 10:39 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2115.codfw.wmnet with reason: host reimage * 10:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2240 [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94653 and previous config saved to /var/cache/conftool/dbconfig/20260701-102658-cwilliams.json * 10:26 moritzm: installing nginx security updates * 10:26 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2086.codfw.wmnet with OS trixie * 10:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2179 to s4 primary [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94652 and previous config saved to /var/cache/conftool/dbconfig/20260701-102356-cwilliams.json * 10:23 cezmunsta: Starting s4 codfw failover from db2240 to db2179 - [[phab:T430127|T430127]] * 10:23 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2106.codfw.wmnet with OS trixie * 10:20 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2115.codfw.wmnet with OS trixie * 10:15 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2179 with weight 0 [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94651 and previous config saved to /var/cache/conftool/dbconfig/20260701-101531-cwilliams.json * 10:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 40 hosts with reason: Primary switchover s4 [[phab:T430127|T430127]] * 09:56 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template (take 2) - oblivian@cumin1003" * 09:56 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template (take 2) - oblivian@cumin1003 * 09:55 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template (take 2) - oblivian@cumin1003 * 09:55 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template (take 2) - oblivian@cumin1003" * 09:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:39 mszwarc@deploy1003: Synchronized private/SuggestedInvestigationsSignals/SuggestedInvestigationsSignal4n.php: Update SI signal 4n (duration: 06m 08s) * 09:21 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 09:21 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 09:14 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 09:14 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 09:02 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 09:02 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 08:54 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 08:38 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 08:38 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 08:36 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 08:21 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 08:21 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] (duration: 36m 11s) * 08:15 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 08:09 mszwarc@deploy1003: mszwarc, abi: Continuing with deployment * 08:03 mszwarc@deploy1003: mszwarc, abi: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:55 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 07:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 07:45 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] * 07:30 aqu@deploy1003: Finished deploy [analytics/refinery@410f205]: Regular analytics weekly train 2nd try [analytics/refinery@410f2050] (duration: 00m 22s) * 07:29 aqu@deploy1003: Started deploy [analytics/refinery@410f205]: Regular analytics weekly train 2nd try [analytics/refinery@410f2050] * 07:28 aqu@deploy1003: Finished deploy [analytics/refinery@410f205] (thin): Regular analytics weekly train THIN [analytics/refinery@410f2050] (duration: 01m 59s) * 07:28 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] (duration: 07m 19s) * 07:26 aqu@deploy1003: Started deploy [analytics/refinery@410f205] (thin): Regular analytics weekly train THIN [analytics/refinery@410f2050] * 07:26 aqu@deploy1003: Finished deploy [analytics/refinery@410f205]: Regular analytics weekly train [analytics/refinery@410f2050] (duration: 04m 32s) * 07:24 mszwarc@deploy1003: wmde-fisch, mszwarc: Continuing with deployment * 07:23 mszwarc@deploy1003: wmde-fisch, mszwarc: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:21 aqu@deploy1003: Started deploy [analytics/refinery@410f205]: Regular analytics weekly train [analytics/refinery@410f2050] * 07:21 aqu@deploy1003: Finished deploy [analytics/refinery@410f205] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@410f2050] (duration: 02m 01s) * 07:20 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] * 07:19 aqu@deploy1003: Started deploy [analytics/refinery@410f205] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@410f2050] * 07:13 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] (duration: 09m 13s) * 07:09 mszwarc@deploy1003: mszwarc, chlod, revi: Continuing with deployment * 07:06 mszwarc@deploy1003: mszwarc, chlod, revi: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:04 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] * 06:55 elukey: upgrade all trixie hosts to pywmflib 3.0 - [[phab:T430552|T430552]] * 06:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:43 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:43 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:42 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:42 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:41 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:41 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:35 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:35 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:34 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:34 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:31 jmm@cumin2003: DONE (PASS) - Cookbook sre.idm.logout (exit_code=0) Logging Niharika29 out of all services on: 2453 hosts * 06:30 oblivian@cumin1003: END (FAIL) - Cookbook sre.deploy.hiddenparma (exit_code=99) Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:30 oblivian@cumin1003: END (FAIL) - Cookbook sre.deploy.python-code (exit_code=99) hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:30 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:30 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:01 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2109.codfw.wmnet with OS trixie * 05:45 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on es1039.eqiad.wmnet with reason: issues * 05:41 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1027.eqiad.wmnet * 05:40 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2068.codfw.wmnet with OS trixie * 05:40 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2109.codfw.wmnet with reason: host reimage * 05:40 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1027.eqiad.wmnet,service=s2 * 05:40 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1027.eqiad.wmnet,service=s7 * 05:36 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2109.codfw.wmnet with reason: host reimage * 05:20 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2068.codfw.wmnet with reason: host reimage * 05:16 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2109.codfw.wmnet with OS trixie * 05:15 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2068.codfw.wmnet with reason: host reimage * 05:09 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2067.codfw.wmnet with OS trixie * 04:56 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2068.codfw.wmnet with OS trixie * 04:49 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2067.codfw.wmnet with reason: host reimage * 04:45 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2067.codfw.wmnet with reason: host reimage * 04:27 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2067.codfw.wmnet with OS trixie * 03:47 slyngshede@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1039.eqiad.wmnet with reason: Hardware crash * 03:21 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2107.codfw.wmnet with OS trixie * 02:59 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2107.codfw.wmnet with reason: host reimage * 02:55 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2085.codfw.wmnet with OS trixie * 02:51 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2072.codfw.wmnet with OS trixie * 02:51 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2107.codfw.wmnet with reason: host reimage * 02:35 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2085.codfw.wmnet with reason: host reimage * 02:31 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2107.codfw.wmnet with OS trixie * 02:30 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2072.codfw.wmnet with reason: host reimage * 02:26 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2085.codfw.wmnet with reason: host reimage * 02:22 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2072.codfw.wmnet with reason: host reimage * 02:09 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2085.codfw.wmnet with OS trixie * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 54s) * 02:03 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2072.codfw.wmnet with OS trixie * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es7 eqiad back to read-write - [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94649 and previous config saved to /var/cache/conftool/dbconfig/20260701-010716-ladsgroup.json * 01:05 ladsgroup@dns1004: END - running authdns-update * 01:05 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depool es1039 [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94648 and previous config saved to /var/cache/conftool/dbconfig/20260701-010551-ladsgroup.json * 01:03 ladsgroup@dns1004: START - running authdns-update * 01:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Promote es1035 to es7 primary [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94647 and previous config saved to /var/cache/conftool/dbconfig/20260701-010002-ladsgroup.json * 00:58 Amir1: Starting es7 eqiad failover from es1039 to es1035 - [[phab:T430765|T430765]] * 00:53 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es1035 with weight 0 [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94646 and previous config saved to /var/cache/conftool/dbconfig/20260701-005329-ladsgroup.json * 00:53 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 9 hosts with reason: Primary switchover es7 [[phab:T430765|T430765]] * 00:42 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es7 eqiad as read-only for maintenance - [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94645 and previous config saved to /var/cache/conftool/dbconfig/20260701-004221-ladsgroup.json * 00:20 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2102.codfw.wmnet with OS trixie * 00:15 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2103.codfw.wmnet with OS trixie * 00:05 dr0ptp4kt: DEPLOYED Refinery at {{Gerrit|4e7a2b32}} for changes: pageview allowlist {{Gerrit|1305158}} (+min.wikiquote) {{Gerrit|1305162}} (+bol.wikipedia), {{Gerrit|1305156}} (+isv.wikipedia); {{Gerrit|1305980}} (pv allowlist -api.wikimedia, sqoop +isvwiki); sqoop {{Gerrit|1295064}} (+globalimagelinks) {{Gerrit|1295069}} (+filerevision) using scap, then deployed onto HDFS (manual copyToLocal required additionally) == Other archives == See [[Server Admin Log/Archives]]. <noinclude> [[Category:SAL]] [[Category:Operations]] </noinclude> 05s5q5q6edf52q5atg63n6gqadkjfn3 2450655 2450654 2026-08-22T16:36:17Z Stashbot 7414 arlolra@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply 2450655 wikitext text/x-wiki == 2026-08-22 == * 16:36 arlolra@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:36 arlolra@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:36 arlolra@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:36 arlolra@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:33 arlolra@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:33 arlolra@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:33 arlolra@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:32 arlolra@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 35s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-21 == * 20:36 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in1001.wikimedia.org with reason: [[phab:T434750|T434750]] * 20:34 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in2001.wikimedia.org with reason: [[phab:T434750|T434750]] * 20:33 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out1001.wikimedia.org with reason: [[phab:T434750|T434750]] * 20:25 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out2001.wikimedia.org with reason: [[phab:T434750|T434750]] * 19:37 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host krb1002.eqiad.wmnet with OS bookworm * 19:00 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 18:59 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 18:51 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 18:51 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 18:35 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:35 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:27 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:27 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:16 bking@cumin2003: START - Cookbook sre.hosts.reimage for host krb1002.eqiad.wmnet with OS bookworm * 17:35 sukhe@dns1004: END - running authdns-update * 17:33 sukhe@dns1004: START - running authdns-update * 17:32 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns5004.wikimedia.org [reason: resolved authdns-update issues] * 17:32 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:32 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: force HEAD to {{Gerrit|be26e30ae101}} - sukhe@cumin1003" * 17:32 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: force HEAD to {{Gerrit|be26e30ae101}} - sukhe@cumin1003" * 17:28 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 17:28 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: service=authdns-update,name=dns5004.wikimedia.org [reason: resolving authdns-update issues] * 17:28 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:28 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: force HEAD to {{Gerrit|be26e30ae101}} - sukhe@cumin1003" * 17:28 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: force HEAD to {{Gerrit|be26e30ae101}} - sukhe@cumin1003" * 17:24 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 17:24 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.netbox (exit_code=97) * 17:23 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 17:16 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 17:12 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 17:10 sukhe@dns1004: END - running authdns-update * 17:08 sukhe@dns1004: START - running authdns-update * 17:08 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=dns5004.wikimedia.org [reason: resolving authdns-update issues] * 17:07 sukhe@dns1004: FAIL - running authdns-update * 17:05 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 17:05 sukhe@dns1004: START - running authdns-update * 17:01 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 16:59 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 16:56 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 16:53 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=dns5004.* [reason: trixie upgrade] * 16:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns5004.wikimedia.org * 16:52 cdobbins@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns5004.wikimedia.org * 16:44 cmooney@dns3003: END - running authdns-update * 16:41 cmooney@dns3003: START - running authdns-update * 16:41 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:41 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on eqsin<->codfw arelion - cmooney@cumin1003" * 16:37 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on eqsin<->codfw arelion - cmooney@cumin1003" * 16:33 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:11 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 16:08 sukhe@dns1004: END - running authdns-update * 16:08 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 16:06 sukhe@dns1004: START - running authdns-update * 16:04 cmooney@dns3003: END - running authdns-update * 16:02 cmooney@dns3003: START - running authdns-update * 16:00 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:00 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on eqord<->codfw arelion - cmooney@cumin1003" * 15:56 cdobbins@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host dns5004.wikimedia.org with OS trixie * 15:55 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on eqord<->codfw arelion - cmooney@cumin1003" * 15:53 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:51 cmooney@cumin1003: END (ERROR) - Cookbook sre.dns.netbox (exit_code=97) * 15:51 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:38 andrew@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudcephosd1042.eqiad.wmnet with OS bookworm * 15:18 andrew@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudcephosd1042.eqiad.wmnet with reason: host reimage * 15:17 dancy@deploy1003: Finished deploy [gerrit/gerrit@2cc11cc]: Deploying https://gerrit.wikimedia.org/r/c/operations/software/gerrit/+/1327669 ([[phab:T434726|T434726]]) (duration: 00m 14s) * 15:17 dancy@deploy1003: Started deploy [gerrit/gerrit@2cc11cc]: Deploying https://gerrit.wikimedia.org/r/c/operations/software/gerrit/+/1327669 ([[phab:T434726|T434726]]) * 15:13 andrew@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudcephosd1042.eqiad.wmnet with reason: host reimage * 15:09 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns5004.wikimedia.org with reason: host reimage * 15:05 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns5004.wikimedia.org with reason: host reimage * 14:53 andrew@cumin2003: START - Cookbook sre.hosts.reimage for host cloudcephosd1042.eqiad.wmnet with OS bookworm * 14:30 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns5004.wikimedia.org with OS trixie * 14:29 cdobbins@cumin1003: conftool action : set/pooled=no; selector: name=dns5004.* [reason: trixie upgrade] * 14:21 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:21 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove entries for cr2-eqord - cmooney@cumin1003" * 14:21 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove entries for cr2-eqord - cmooney@cumin1003" * 14:13 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 14:11 moritzm: imported openjdk 8u504-ga-1~deb12u1 for bookworm-wikimedia (backport of the latest Java 8 security fixes for bookworm) * 13:25 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "sync cr2-eqord router offline - cmooney@cumin1003" * 13:23 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "sync cr2-eqord router offline - cmooney@cumin1003" * 13:14 hashar@deploy1003: Finished deploy [integration/docroot@2d5ff9b]: opensource: add PersonalDashboard docs to MW components - [[phab:T435392|T435392]] (duration: 00m 15s) * 13:14 hashar@deploy1003: Started deploy [integration/docroot@2d5ff9b]: opensource: add PersonalDashboard docs to MW components - [[phab:T435392|T435392]] * 12:16 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:16 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: [[phab:T431682|T431682]] - filippo@cumin1003" * 12:16 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: [[phab:T431682|T431682]] - filippo@cumin1003" * 12:11 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2006.wikimedia.org with OS trixie * 12:00 kevinbazira@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 11:58 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 11:43 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2006.wikimedia.org with reason: host reimage * 11:41 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 11:38 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2006.wikimedia.org with reason: host reimage * 11:20 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2006.wikimedia.org with OS trixie * 11:11 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2005.wikimedia.org with OS trixie * 10:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2005.wikimedia.org with reason: host reimage * 10:53 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2005.wikimedia.org with reason: host reimage * 10:43 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-codfw * 10:43 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2011.codfw.wmnet * 10:43 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2011.codfw.wmnet * 10:40 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2011.codfw.wmnet * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2011.codfw.wmnet * 10:34 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2010.codfw.wmnet * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2010.codfw.wmnet * 10:33 fnegri@cumin1003: END (PASS) - Cookbook sre.wikireplicas.add-wiki (exit_code=0) for database bolwiki ([[phab:T429954|T429954]]) * 10:33 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2005.wikimedia.org with OS trixie * 10:30 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2010.codfw.wmnet * 10:25 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2010.codfw.wmnet * 10:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2009.codfw.wmnet * 10:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2009.codfw.wmnet * 10:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1006.wikimedia.org with OS trixie * 10:18 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2009.codfw.wmnet * 10:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2009.codfw.wmnet * 10:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2008.codfw.wmnet * 10:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2008.codfw.wmnet * 10:06 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2008.codfw.wmnet * 10:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1006.wikimedia.org with reason: host reimage * 10:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2008.codfw.wmnet * 10:01 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2007.codfw.wmnet * 10:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2007.codfw.wmnet * 09:57 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1006.wikimedia.org with reason: host reimage * 09:56 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2007.codfw.wmnet * 09:51 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2007.codfw.wmnet * 09:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 09:51 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 09:46 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2006.codfw.wmnet * 09:46 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1006.wikimedia.org with OS trixie * 09:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1005.wikimedia.org with OS trixie * 09:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2006.codfw.wmnet * 09:41 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2005.codfw.wmnet * 09:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2005.codfw.wmnet * 09:38 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2013.codfw.wmnet * 09:36 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2005.codfw.wmnet * 09:35 fnegri@cumin1003: START - Cookbook sre.wikireplicas.add-wiki for database bolwiki ([[phab:T429954|T429954]]) * 09:35 fnegri@cumin1003: END (PASS) - Cookbook sre.wikireplicas.add-wiki (exit_code=0) for database minwikiquote ([[phab:T429946|T429946]]) * 09:35 fnegri@cumin1003: START - Cookbook sre.wikireplicas.add-wiki for database minwikiquote ([[phab:T429946|T429946]]) * 09:32 blake@cumin1003: START - Cookbook sre.hosts.reboot-single for host rdb2013.codfw.wmnet * 09:30 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2011.codfw.wmnet * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1005.wikimedia.org with reason: host reimage * 09:26 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2005.codfw.wmnet * 09:26 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2004.codfw.wmnet * 09:26 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2004.codfw.wmnet * 09:24 blake@cumin1003: START - Cookbook sre.hosts.reboot-single for host rdb2011.codfw.wmnet * 09:22 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1005.wikimedia.org with reason: host reimage * 09:21 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2004.codfw.wmnet * 09:16 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1015.eqiad.wmnet * 09:16 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2004.codfw.wmnet * 09:16 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2003.codfw.wmnet * 09:16 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2003.codfw.wmnet * 09:13 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on cr[1-2]-eqiad,pfw1-eqiad with reason: upgrade pfw1a-eqiad and pfw1b-eqiad pair * 09:12 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 09:11 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 09:11 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 09:11 blake@cumin1003: START - Cookbook sre.hosts.reboot-single for host rdb1015.eqiad.wmnet * 09:10 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2003.codfw.wmnet * 09:09 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1013.eqiad.wmnet * 09:07 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1005.wikimedia.org with OS trixie * 09:03 blake@cumin1003: START - Cookbook sre.hosts.reboot-single for host rdb1013.eqiad.wmnet * 09:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2003.codfw.wmnet * 09:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2002.codfw.wmnet * 09:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2002.codfw.wmnet * 08:54 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2002.codfw.wmnet * 08:49 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2002.codfw.wmnet * 08:49 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2001.codfw.wmnet * 08:49 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2001.codfw.wmnet * 08:48 jmm@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts netmon2002.wikimedia.org * 08:47 jmm@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts netmon2002.wikimedia.org * 08:44 jmm@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts netmon2002.wikimedia.org * 08:44 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon2002.wikimedia.org * 08:43 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2001.codfw.wmnet * 08:36 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon2002.wikimedia.org * 08:34 jmm@dns1004: END - running authdns-update * 08:33 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2001.codfw.wmnet * 08:33 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-codfw * 08:31 jmm@dns1004: START - running authdns-update * 07:48 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327679{{!}}Block: Disable flaky API test (T435272 T389028)]], [[gerrit:1327678{{!}}API: wfDebugLog for thumberror]] (duration: 15m 34s) * 07:41 krinkle@deploy1003: krinkle: Continuing with deployment * 07:37 krinkle@deploy1003: krinkle: Backport for [[gerrit:1327679{{!}}Block: Disable flaky API test (T435272 T389028)]], [[gerrit:1327678{{!}}API: wfDebugLog for thumberror]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:33 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1327679{{!}}Block: Disable flaky API test (T435272 T389028)]], [[gerrit:1327678{{!}}API: wfDebugLog for thumberror]] * 07:25 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 07:24 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 07:18 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 07:18 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 07:15 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 07:14 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 07:14 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 07:14 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 07:13 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 07:03 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1283: Pool back * 06:42 jmm@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts netmon2002.wikimedia.org * 06:35 moritzm: powercycling netmon2002 * 06:18 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1283: Pool back * 06:17 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1283 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96213 and previous config saved to /var/cache/conftool/dbconfig/20260821-061743-marostegui.json * 04:59 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324963{{!}}Add Produnto to extension-list (T421436)]], [[gerrit:1324964{{!}}Enable Produnto on Beta (T421436)]] (duration: 34m 48s) * 04:45 tstarling@deploy1003: tstarling: Continuing with deployment * 04:44 tstarling@deploy1003: tstarling: Backport for [[gerrit:1324963{{!}}Add Produnto to extension-list (T421436)]], [[gerrit:1324964{{!}}Enable Produnto on Beta (T421436)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 04:24 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1324963{{!}}Add Produnto to extension-list (T421436)]], [[gerrit:1324964{{!}}Enable Produnto on Beta (T421436)]] * 04:21 arlolra@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 04:20 arlolra@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 04:20 arlolra@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 04:20 arlolra@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 41s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-20 == * 23:43 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327652{{!}}RunSingleJob: Add ProfilingContext::init() (T435422)]] (duration: 12m 06s) * 23:38 krinkle@deploy1003: krinkle: Continuing with deployment * 23:33 krinkle@deploy1003: krinkle: Backport for [[gerrit:1327652{{!}}RunSingleJob: Add ProfilingContext::init() (T435422)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:31 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1327652{{!}}RunSingleJob: Add ProfilingContext::init() (T435422)]] * 22:15 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1054.eqiad.wmnet with OS trixie * 22:14 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 22:14 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 21:58 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1054.eqiad.wmnet with reason: host reimage * 21:51 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1054.eqiad.wmnet with reason: host reimage * 21:36 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1054.eqiad.wmnet with OS trixie * 21:36 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:35 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327219{{!}}RunSingleJob: Define MW_ENTRY_POINT for flamegraph sample attribution (T435422)]] (duration: 08m 30s) * 21:31 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:31 krinkle@deploy1003: krinkle: Continuing with deployment * 21:31 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1054 * 21:31 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1054 * 21:29 krinkle@deploy1003: krinkle: Backport for [[gerrit:1327219{{!}}RunSingleJob: Define MW_ENTRY_POINT for flamegraph sample attribution (T435422)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:27 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1327219{{!}}RunSingleJob: Define MW_ENTRY_POINT for flamegraph sample attribution (T435422)]] * 21:17 maryum: Deployed security fix for [[phab:T433020|T433020]] * 20:59 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324752{{!}}InitialiseSettings: Enable 2FA banners on remaining private wikis (T428103)]], [[gerrit:1325920{{!}}Remove sending email to legal team about rejected requests (T374053)]] (duration: 07m 18s) * 20:54 reedy@deploy1003: neriah, reedy: Continuing with deployment * 20:54 reedy@deploy1003: neriah, reedy: Backport for [[gerrit:1324752{{!}}InitialiseSettings: Enable 2FA banners on remaining private wikis (T428103)]], [[gerrit:1325920{{!}}Remove sending email to legal team about rejected requests (T374053)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:51 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324752{{!}}InitialiseSettings: Enable 2FA banners on remaining private wikis (T428103)]], [[gerrit:1325920{{!}}Remove sending email to legal team about rejected requests (T374053)]] * 20:24 reedy@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.15,1.47.0-wmf.16,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/med * 20:23 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324752{{!}}InitialiseSettings: Enable 2FA banners on remaining private wikis (T428103)]], [[gerrit:1325920{{!}}Remove sending email to legal team about rejected requests (T374053)]] * 20:14 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 20:10 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 20:09 cdanis@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "bug fixes & UX fixes - cdanis@cumin1003" * 20:09 cdanis@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: bug fixes & UX fixes - cdanis@cumin1003 * 20:08 cdanis@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: bug fixes & UX fixes - cdanis@cumin1003 * 20:08 cdanis@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "bug fixes & UX fixes - cdanis@cumin1003" * 19:24 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327598{{!}}Make \Omicron non upright (like \Chi) (T434428)]], [[gerrit:1327596{{!}}Render overline of \bar with stretchy=false (T435456)]] (duration: 18m 54s) * 19:20 krinkle@deploy1003: krinkle: Continuing with deployment * 19:07 krinkle@deploy1003: krinkle: Backport for [[gerrit:1327598{{!}}Make \Omicron non upright (like \Chi) (T434428)]], [[gerrit:1327596{{!}}Render overline of \bar with stretchy=false (T435456)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:05 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1327598{{!}}Make \Omicron non upright (like \Chi) (T434428)]], [[gerrit:1327596{{!}}Render overline of \bar with stretchy=false (T435456)]] * 18:51 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327614{{!}}Avoid casting fpxmax to string (T318419)]] (duration: 07m 28s) * 18:50 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:46 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:46 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 18:45 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327614{{!}}Avoid casting fpxmax to string (T318419)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:43 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327614{{!}}Avoid casting fpxmax to string (T318419)]] * 18:37 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:37 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:37 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:36 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:07 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 18:05 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 18:01 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 18:01 sukhe@dns1004: END - running authdns-update * 17:59 sukhe@dns1004: START - running authdns-update * 17:58 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 17:57 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=dns6002.* [reason: depooling for trixie upgrade] * 17:56 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns6002.wikimedia.org * 17:56 cdobbins@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns6002.wikimedia.org * 17:51 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 17:51 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 17:34 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host stat1011.eqiad.wmnet with OS bookworm * 17:31 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns6002.wikimedia.org with OS trixie * 17:30 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 17:30 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 17:30 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 17:30 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 17:29 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 17:29 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 17:29 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 17:29 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:29 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:27 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:24 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327590{{!}}Make sure fpsmax is an int value (T318419)]] (duration: 08m 37s) * 17:20 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 17:17 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327590{{!}}Make sure fpsmax is an int value (T318419)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:16 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 17:15 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327590{{!}}Make sure fpsmax is an int value (T318419)]] * 16:51 swfrench-wmf: disable-puppet on A:cp for ATS Lua change - [[phab:T427666|T427666]] * 16:51 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db2901.codfw.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 16:43 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 16:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on stat1011.eqiad.wmnet with reason: host reimage * 16:39 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns6002.wikimedia.org with reason: host reimage * 16:36 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on stat1011.eqiad.wmnet with reason: host reimage * 16:36 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db2901.codfw.wmnet * 16:34 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns6002.wikimedia.org with reason: host reimage * 16:15 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns6002.wikimedia.org with OS trixie * 16:14 cdobbins@cumin1003: conftool action : set/pooled=no; selector: name=dns6002.* [reason: depooling for trixie upgrade] * 16:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host stat1011 * 16:10 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host stat1011 * 16:09 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host stat1011 * 16:09 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) stat1011.eqiad.wmnet 14.36.64.10.in-addr.arpa 4.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 bking@cumin2003: START - Cookbook sre.dns.wipe-cache stat1011.eqiad.wmnet 14.36.64.10.in-addr.arpa 4.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:09 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host stat1011 - bking@cumin2003" * 16:09 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host stat1011 - bking@cumin2003" * 16:05 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host stat1011 * 16:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host stat1011.eqiad.wmnet with OS bookworm * 16:00 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 15:59 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 15:56 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 15:55 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 15:37 jayme@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:35 jayme@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 15:35 jayme@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:32 jayme@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:32 jayme@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:30 jayme@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 15:30 jayme@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:28 jayme@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:28 jayme@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 15:28 fceratto@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host db1903.eqiad.wmnet * 15:28 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1903.eqiad.wmnet with OS trixie * 15:26 jayme@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 15:26 jayme@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 15:24 jayme@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 15:24 jayme@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 15:21 jayme@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 15:21 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 15:19 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 15:19 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 15:17 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 15:14 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1903.eqiad.wmnet with reason: host reimage * 15:07 fceratto@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1903.eqiad.wmnet with reason: host reimage * 14:54 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db1903.eqiad.wmnet with OS trixie * 14:53 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1903.eqiad.wmnet - fceratto@cumin1003" * 14:53 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1903.eqiad.wmnet - fceratto@cumin1003" * 14:53 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1903.eqiad.wmnet on all recursors * 14:53 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1903.eqiad.wmnet on all recursors * 14:53 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:53 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1903.eqiad.wmnet - fceratto@cumin1003" * 14:53 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1903.eqiad.wmnet - fceratto@cumin1003" * 14:49 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 14:49 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1903.eqiad.wmnet * 14:33 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=cp1100.* * 14:27 topranks: reconfigure eqiad<->codfw bgp settings * 14:22 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:22 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update entries used on new transport backup eqiad codfw - cmooney@cumin1003" * 14:19 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update entries used on new transport backup eqiad codfw - cmooney@cumin1003" * 14:14 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 14:14 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 14:13 moritzm: installing util-linux security updates * 14:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-staging-worker * 14:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2003.codfw.wmnet * 14:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2003.codfw.wmnet * 14:08 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2003.codfw.wmnet * 14:06 moritzm: installing libheif security updates * 13:58 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2003.codfw.wmnet * 13:58 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2002.codfw.wmnet * 13:58 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2002.codfw.wmnet * 13:56 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:56 fnegri@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for clouddb1025.eqiad.wmnet * 13:56 fnegri@cumin1003: START - Cookbook sre.hosts.remove-downtime for clouddb1025.eqiad.wmnet * 13:56 Lucas_WMDE: UTC afternoon backport+config window done * 13:53 moritzm: installing apr-util security updates * 13:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2002.codfw.wmnet * 13:50 fnegri@cumin1003: conftool action : set/weight=100; selector: name=clouddb1025.eqiad.wmnet * 13:49 fnegri@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1025.eqiad.wmnet * 13:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2002.codfw.wmnet * 13:41 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2001.codfw.wmnet * 13:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2001.codfw.wmnet * 13:41 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'. * 13:38 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'. * 13:38 fnegri@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on clouddb1025.eqiad.wmnet with reason: Removing s6 from clouddb1025 * 13:34 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2001.codfw.wmnet * 13:31 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'. * 13:29 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'. * 13:28 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet * 13:26 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host stat1009.eqiad.wmnet with OS bookworm * 13:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2001.codfw.wmnet * 13:24 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-staging-worker * 13:23 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1002.eqiad.wmnet * 13:21 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host stat1010.eqiad.wmnet with OS bookworm * 13:20 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1002.eqiad.wmnet * 13:20 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1001.eqiad.wmnet * 13:17 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1001.eqiad.wmnet * 13:16 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2001.codfw.wmnet * 13:13 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327511{{!}}UIC: Fix page:page instead of page:other in instrumentation]] (duration: 07m 00s) * 13:13 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2001.codfw.wmnet * 13:12 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2002.codfw.wmnet * 13:09 mszwarc@deploy1003: mszwarc: Continuing with deployment * 13:08 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1327511{{!}}UIC: Fix page:page instead of page:other in instrumentation]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2002.codfw.wmnet * 13:07 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2002.codfw.wmnet * 13:06 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1327511{{!}}UIC: Fix page:page instead of page:other in instrumentation]] * 13:04 jmm@dns1004: END - running authdns-update * 13:03 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2002.codfw.wmnet * 13:03 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2001.codfw.wmnet * 13:02 jmm@dns1004: START - running authdns-update * 13:01 cmooney@dns3003: END - running authdns-update * 13:00 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2001.codfw.wmnet * 12:59 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2001.codfw.wmnet * 12:59 cmooney@dns3003: START - running authdns-update * 12:57 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2001.codfw.wmnet * 12:56 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2002.codfw.wmnet * 12:55 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:55 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on drmrs<->eqiad GTT vpls - cmooney@cumin1003" * 12:54 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on drmrs<->eqiad GTT vpls - cmooney@cumin1003" * 12:54 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2002.codfw.wmnet * 12:54 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2003.codfw.wmnet * 12:50 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2003.codfw.wmnet * 12:49 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 12:48 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2001.codfw.wmnet * 12:46 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2001.codfw.wmnet * 12:46 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2002.codfw.wmnet * 12:43 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2002.codfw.wmnet * 12:43 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2003.codfw.wmnet * 12:42 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=1) for new host db1902.eqiad.wmnet * 12:42 fceratto@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host db1902.eqiad.wmnet with OS trixie * 12:41 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2003.codfw.wmnet * 12:40 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1003.eqiad.wmnet * 12:38 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1003.eqiad.wmnet * 12:37 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1002.eqiad.wmnet * 12:35 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1002.eqiad.wmnet * 12:35 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1001.eqiad.wmnet * 12:34 cmooney@dns3003: END - running authdns-update * 12:33 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1001.eqiad.wmnet * 12:32 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on stat1009.eqiad.wmnet with reason: host reimage * 12:31 cmooney@dns3003: START - running authdns-update * 12:31 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:31 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on drmrs<->eqiad cct - cmooney@cumin1003" * 12:28 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on drmrs<->eqiad cct - cmooney@cumin1003" * 12:28 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1902.eqiad.wmnet with reason: host reimage * 12:25 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on stat1009.eqiad.wmnet with reason: host reimage * 12:25 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 12:24 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on stat1010.eqiad.wmnet with reason: host reimage * 12:22 fceratto@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1902.eqiad.wmnet with reason: host reimage * 12:21 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on stat1010.eqiad.wmnet with reason: host reimage * 12:14 elukey: move the Docker Registry's /v2/dev/.* prefix to its dedicated S3 backend - [[phab:T432829|T432829]] * 12:12 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db1902.eqiad.wmnet with OS trixie * 12:09 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1902.eqiad.wmnet - fceratto@cumin1003" * 12:09 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1902.eqiad.wmnet - fceratto@cumin1003" * 12:09 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1902.eqiad.wmnet on all recursors * 12:09 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1902.eqiad.wmnet on all recursors * 12:09 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:08 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1902.eqiad.wmnet - fceratto@cumin1003" * 12:08 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1902.eqiad.wmnet - fceratto@cumin1003" * 12:08 tgr_: [[phab:T413390|T413390]] running CentralAuth:FixRenamedUserGlobalEditCount --wiki=metawiki --since=20250901000000 --until=20260301000000 --fix * 12:04 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1009.eqiad.wmnet with OS bookworm * 12:04 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1010.eqiad.wmnet with OS bookworm * 12:01 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 12:01 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1902.eqiad.wmnet * 12:00 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host stat1010.eqiad.wmnet with OS bookworm * 11:57 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1902.eqiad.wmnet * 11:57 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:57 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1902.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 11:57 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1902.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 11:51 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'. * 11:49 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'. * 11:48 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'. * 11:46 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'. * 11:39 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 11:37 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327518{{!}}Enable thumb.wikimedia.org on cswiki and fawiki (T427465)]] (duration: 10m 40s) * 11:35 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1902.eqiad.wmnet * 11:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1010.eqiad.wmnet with OS bookworm * 11:33 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 11:30 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327518{{!}}Enable thumb.wikimedia.org on cswiki and fawiki (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:26 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327518{{!}}Enable thumb.wikimedia.org on cswiki and fawiki (T427465)]] * 11:24 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host stat1010.eqiad.wmnet with OS bookworm * 11:07 fceratto@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host db1901.eqiad.wmnet * 11:07 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1901.eqiad.wmnet with OS trixie * 10:53 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1901.eqiad.wmnet with reason: host reimage * 10:47 tappof: bump space for prometheus k8s-aux in codfw * 10:47 tappof: bump space for prometheus k8s-dse in eqiad * 10:47 fceratto@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1901.eqiad.wmnet with reason: host reimage * 10:35 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db1901.eqiad.wmnet with OS trixie * 10:32 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:32 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:32 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1901.eqiad.wmnet on all recursors * 10:32 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1901.eqiad.wmnet on all recursors * 10:31 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:31 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:31 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:27 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:27 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1901.eqiad.wmnet * 10:23 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1010.eqiad.wmnet with OS bookworm * 10:20 fceratto@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts db1901.eqiad.wmnet * 10:20 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 10:18 blake@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 10:17 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:16 blake@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 10:13 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1901.eqiad.wmnet * 09:23 jelto@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'. * 09:22 jelto@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'. * 09:22 jelto: update cert-manager to 1.19.6 on wikikube staging-eqiad - [[phab:T427402|T427402]] * 09:20 moritzm: imported squid 7.6-2.1for trixie-wikimedia/main [[phab:T427282|T427282]] * 09:08 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 09:08 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 09:08 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 09:07 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 09:04 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 09:04 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:23 slyngshede@dns1004: END - running authdns-update * 08:21 slyngshede@dns1004: START - running authdns-update * 08:18 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.16 refs [[phab:T430835|T430835]] * 06:27 aokoth@dns1004: END - running authdns-update * 06:25 aokoth@dns1004: START - running authdns-update * 06:22 brennen@deploy1003: Finished deploy [phabricator/deployment@6b9b6ff]: deploy phab1005 for [[phab:T435087|T435087]] (duration: 00m 39s) * 06:21 brennen@deploy1003: Started deploy [phabricator/deployment@6b9b6ff]: deploy phab1005 for [[phab:T435087|T435087]] * 06:20 brennen@deploy1003: Finished deploy [phabricator/deployment@6b9b6ff]: deploy phab1004 for to pick up config values for [[phab:T435087|T435087]] (duration: 01m 46s) * 06:18 brennen@deploy1003: Started deploy [phabricator/deployment@6b9b6ff]: deploy phab1004 for to pick up config values for [[phab:T435087|T435087]] * 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 49s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-19 == * 23:19 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327197{{!}}Enable thumb.wikimedia.org on mediawiki.org (T427465)]] (duration: 10m 50s) * 23:18 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:16 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 23:15 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 23:10 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327197{{!}}Enable thumb.wikimedia.org on mediawiki.org (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:10 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:09 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 23:09 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:09 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 23:08 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327197{{!}}Enable thumb.wikimedia.org on mediawiki.org (T427465)]] * 22:58 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327201{{!}}Enable ReadingLists for all logged in users on test wiki (T435258)]] (duration: 11m 20s) * 22:50 jdlrobson@deploy1003: jdlrobson: Continuing with deployment * 22:49 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1327201{{!}}Enable ReadingLists for all logged in users on test wiki (T435258)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:46 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1327201{{!}}Enable ReadingLists for all logged in users on test wiki (T435258)]] * 22:42 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327162{{!}}Article: Split subjectpageheader by model and disable for wikitext]], [[gerrit:1327169{{!}}Make uppercase greek letters normal (non-italic) font (T434686 T434428)]], [[gerrit:1327176{{!}}Skin: Avoid DB lookup for pagecategorieslink message (T347123)]] (duration: 37m 52s) * 22:29 krinkle@deploy1003: krinkle: Continuing with deployment * 22:25 krinkle@deploy1003: krinkle: Backport for [[gerrit:1327162{{!}}Article: Split subjectpageheader by model and disable for wikitext]], [[gerrit:1327169{{!}}Make uppercase greek letters normal (non-italic) font (T434686 T434428)]], [[gerrit:1327176{{!}}Skin: Avoid DB lookup for pagecategorieslink message (T347123)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:04 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1327162{{!}}Article: Split subjectpageheader by model and disable for wikitext]], [[gerrit:1327169{{!}}Make uppercase greek letters normal (non-italic) font (T434686 T434428)]], [[gerrit:1327176{{!}}Skin: Avoid DB lookup for pagecategorieslink message (T347123)]] * 22:04 krinkle@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: awaiting CI (duration: 03m 06s) * 22:01 krinkle@deploy1003: Locking from deployment [ALL REPOSITORIES]: awaiting CI * 22:00 krinkle@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: awaiting CI (duration: 00m 01s) * 22:00 krinkle@deploy1003: Locking from deployment [ALL REPOSITORIES]: awaiting CI * 21:34 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2001.codfw.wmnet * 21:28 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2001.codfw.wmnet * 21:22 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327178{{!}}AccountRecovery: Notify the email address of the on file of the request (T425799)]] (duration: 47m 02s) * 21:13 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm * 21:09 catrope@deploy1003: catrope: Continuing with deployment * 20:55 catrope@deploy1003: catrope: Backport for [[gerrit:1327178{{!}}AccountRecovery: Notify the email address of the on file of the request (T425799)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:35 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1327178{{!}}AccountRecovery: Notify the email address of the on file of the request (T425799)]] * 20:31 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327128{{!}}Parsoid DataAccess: convert Parsoid fragment markers to/from strip tags (T432547)]] (duration: 07m 30s) * 20:27 catrope@deploy1003: catrope, arlolra: Continuing with deployment * 20:26 catrope@deploy1003: catrope, arlolra: Backport for [[gerrit:1327128{{!}}Parsoid DataAccess: convert Parsoid fragment markers to/from strip tags (T432547)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:24 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1327128{{!}}Parsoid DataAccess: convert Parsoid fragment markers to/from strip tags (T432547)]] * 20:23 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage * 20:17 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage * 20:15 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325878{{!}}[arwiki] Enable restricted user page editing and grant edit permissions (T434878)]] (duration: 08m 46s) * 20:11 catrope@deploy1003: catrope, gergesshamon: Continuing with deployment * 20:08 catrope@deploy1003: catrope, gergesshamon: Backport for [[gerrit:1325878{{!}}[arwiki] Enable restricted user page editing and grant edit permissions (T434878)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:06 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1325878{{!}}[arwiki] Enable restricted user page editing and grant edit permissions (T434878)]] * 19:59 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm * 19:56 eevans@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cassandra-dev2001.codfw.wmnet with OS bookworm * 19:56 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm * 19:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2207.codfw.wmnet with reason: Maintenance * 18:47 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319920{{!}}Allow setting a separate thumbUrl in production (T427465)]], [[gerrit:1327167{{!}}Fix wmgThumbUrl config (T427465)]] (duration: 18m 53s) * 18:43 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 18:30 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1319920{{!}}Allow setting a separate thumbUrl in production (T427465)]], [[gerrit:1327167{{!}}Fix wmgThumbUrl config (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:28 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1319920{{!}}Allow setting a separate thumbUrl in production (T427465)]], [[gerrit:1327167{{!}}Fix wmgThumbUrl config (T427465)]] * 18:26 sukhe@dns1004: END - running authdns-update * 18:24 sukhe@dns1004: START - running authdns-update * 18:09 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1319920{{!}}Allow setting a separate thumbUrl in production (T427465)]] * 18:03 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-eqiad and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 17:56 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-codfw and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 17:50 cmooney@dns3003: END - running authdns-update * 17:42 dancy@deploy1003: Installation of scap version "4.283.0" completed for 3 hosts * 17:41 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling reboot on A:durum and A:durum * 17:40 cmooney@dns3003: START - running authdns-update * 17:40 dancy@deploy1003: Installing scap version "4.283.0" for 3 host(s) * 17:38 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:37 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on GTT VPLS - cmooney@cumin1003" * 17:37 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --verbose --scoreLessThan=0.7 --exceptDatasetChecksums=[[phab:T434319|T434319]]-enwiki-models.txt # [[phab:T434319|T434319]] * 17:32 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on GTT VPLS - cmooney@cumin1003" * 17:29 sbassett: Deployed security fix for [[phab:T435210|T435210]] (wmf.16) * 17:26 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 17:22 sbassett: Deployed security fix for [[phab:T435210|T435210]] (wmf.15) * 17:00 sukhe@dns1004: END - running authdns-update * 16:58 sukhe@dns1004: START - running authdns-update * 16:53 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica-esams and A:liberica * 16:41 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica-esams and A:liberica * 16:41 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-eqiad and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 16:41 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-codfw and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 16:41 cjd91: sudo -i cookbook sre.cdn.roll-upgrade-ats --query 'A:cp-codfw' --task-id [[phab:T434478|T434478]] --reason '9.2.15 upgrade' * 16:41 cjd91: sudo -i cookbook sre.cdn.roll-upgrade-ats --query 'A:cp-eqiad' --task-id [[phab:T434478|T434478]] --reason '9.2.15 upgrade' * 16:40 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and A:durum * 16:28 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326881{{!}}Echo: Start using virtual domains (T380385)]] (duration: 13m 12s) * 16:23 urbanecm@deploy1003: urbanecm: Continuing with deployment * 16:21 urandom: Completed sessionstore Cassandra/JVM upgrade — [[phab:T435154|T435154]] * 16:21 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching sessionstore[2005-2006].codfw.wmnet,sessionstore[1005-1006].eqiad.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 16:19 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1326881{{!}}Echo: Start using virtual domains (T380385)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:15 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:15 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->codfw - cmooney@cumin1003" * 16:14 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1326881{{!}}Echo: Start using virtual domains (T380385)]] * 16:14 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327123{{!}}Revert^2 "Migrate database access to virtual domains" (T435305)]], [[gerrit:1327124{{!}}Pass the mapped domain of virtual-echo-shared to the push NameTableStores (T435305)]] (duration: 07m 42s) * 16:13 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching sessionstore[2005-2006].codfw.wmnet,sessionstore[1005-1006].eqiad.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 16:11 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->codfw - cmooney@cumin1003" * 16:10 urbanecm@deploy1003: urbanecm: Continuing with deployment * 16:08 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1327123{{!}}Revert^2 "Migrate database access to virtual domains" (T435305)]], [[gerrit:1327124{{!}}Pass the mapped domain of virtual-echo-shared to the push NameTableStores (T435305)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:06 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 16:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 16:06 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1327123{{!}}Revert^2 "Migrate database access to virtual domains" (T435305)]], [[gerrit:1327124{{!}}Pass the mapped domain of virtual-echo-shared to the push NameTableStores (T435305)]] * 16:06 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:03 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching sessionstore1004.eqiad.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 16:01 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching sessionstore1004.eqiad.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 15:56 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching sessionstore2004.codfw.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 15:54 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching sessionstore2004.codfw.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 15:53 urandom: beginning sessionstore Cassandra/JVM upgrade — [[phab:T435154|T435154]] * 15:52 urandom: beginning sessionstore Cassandra/JVM upgrade — [[phab:T432944|T432944]] * 15:51 cmooney@dns3003: END - running authdns-update * 15:49 cmooney@dns3003: START - running authdns-update * 15:48 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:48 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->codfw - cmooney@cumin1003" * 15:45 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->codfw - cmooney@cumin1003" * 15:44 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 15:44 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:43 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 15:42 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:42 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:38 cmooney@dns3003: END - running authdns-update * 15:36 cmooney@dns3003: START - running authdns-update * 15:36 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:36 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->eqsin - cmooney@cumin1003" * 15:34 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327138{{!}}Enable redis lock manager everywhere (T366938)]] (duration: 08m 36s) * 15:33 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->eqsin - cmooney@cumin1003" * 15:30 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:29 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 15:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1147.eqiad.wmnet with OS bookworm * 15:28 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327138{{!}}Enable redis lock manager everywhere (T366938)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:25 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327138{{!}}Enable redis lock manager everywhere (T366938)]] * 15:24 jmm@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host krb1002.eqiad.wmnet * 15:19 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2207.codfw.wmnet with reason: Host crashed * 15:17 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327098{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]], [[gerrit:1327101{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]] (duration: 07m 13s) * 15:12 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 15:12 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1327098{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]], [[gerrit:1327101{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:10 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1327098{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]], [[gerrit:1327101{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]] * 15:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2010.codfw.wmnet with OS trixie * 15:05 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 15:05 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1147.eqiad.wmnet with reason: host reimage * 14:59 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 14:59 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:58 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1147.eqiad.wmnet with reason: host reimage * 14:55 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:55 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:49 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 14:48 cmooney@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host durum1001.eqiad.wmnet * 14:46 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:46 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Delete db2902 ipv6 addr - fceratto@cumin1003" * 14:46 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Delete db2902 ipv6 addr - fceratto@cumin1003" * 14:43 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1147.eqiad.wmnet with OS bookworm * 14:42 tgr@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327094{{!}}SpecialMWOAuthListConsumers: Handle newFromMWUser returning null in addNavigationSubtitle (T435167)]] (duration: 19m 25s) * 14:42 cmooney@cumin1003: START - Cookbook sre.hosts.reboot-single for host durum1001.eqiad.wmnet * 14:42 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 14:41 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 14:38 tgr@deploy1003: tgr: Continuing with deployment * 14:36 tgr@deploy1003: tgr: Backport for [[gerrit:1327094{{!}}SpecialMWOAuthListConsumers: Handle newFromMWUser returning null in addNavigationSubtitle (T435167)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:28 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 14:28 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:28 cmooney@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host durum3005.esams.wmnet * 14:25 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 14:23 cmooney@cumin1003: START - Cookbook sre.hosts.reboot-single for host durum3005.esams.wmnet * 14:23 tgr@deploy1003: Started scap sync-world: Backport for [[gerrit:1327094{{!}}SpecialMWOAuthListConsumers: Handle newFromMWUser returning null in addNavigationSubtitle (T435167)]] * 14:18 gengh@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:17 topranks: disable puppet on hosts running BIRD BGP to test merge of patch to systemd healtchcheck service * 14:17 gengh@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:17 gengh@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:16 gengh@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:16 elukey: upgrade spicerack on cumin1003 and cumin2003 * 14:16 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:15 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314962{{!}}static: add new dir bimi/ for BIMI SVG and PEM file (T311685)]] (duration: 10m 00s) * 14:15 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:11 kharlan@deploy1003: kharlan, sukhe: Continuing with deployment * 14:11 gengh@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:09 gengh@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:09 gengh@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:08 kharlan@deploy1003: kharlan, sukhe: Backport for [[gerrit:1314962{{!}}static: add new dir bimi/ for BIMI SVG and PEM file (T311685)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:06 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 14:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:06 gengh@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:05 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1314962{{!}}static: add new dir bimi/ for BIMI SVG and PEM file (T311685)]] * 14:05 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host krb1002.eqiad.wmnet * 14:05 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:04 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:04 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326342{{!}}Revert^2 "wmf-config/ProductionServices: set URL for urldownloader to service record"]] (duration: 07m 40s) * 13:59 kharlan@deploy1003: kharlan, sukhe: Continuing with deployment * 13:58 kharlan@deploy1003: kharlan, sukhe: Backport for [[gerrit:1326342{{!}}Revert^2 "wmf-config/ProductionServices: set URL for urldownloader to service record"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:57 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host stat1008.eqiad.wmnet with OS bookworm * 13:56 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1326342{{!}}Revert^2 "wmf-config/ProductionServices: set URL for urldownloader to service record"]] * 13:56 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 13:54 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325532{{!}}srwiki: Allow bureaucrats to add and remove event-organizer group (T434748)]] (duration: 14m 56s) * 13:54 swfrench@dns1004: END - running authdns-update * 13:52 swfrench@dns1004: START - running authdns-update * 13:51 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:51 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 13:48 kharlan@deploy1003: kharlan, danielyepezgarces: Continuing with deployment * 13:45 swfrench@cumin2003: conftool action : set/pooled=yes; selector: name=wikikube-worker2330.codfw.wmnet * 13:44 swfrench-wmf: finished etcd-main codfw -> eqiad switchover - [[phab:T435103|T435103]] * 13:44 kharlan@deploy1003: kharlan, danielyepezgarces: Backport for [[gerrit:1325532{{!}}srwiki: Allow bureaucrats to add and remove event-organizer group (T434748)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:44 swfrench@cumin2003: conftool action : set/pooled=no; selector: name=wikikube-worker2330.codfw.wmnet * 13:41 swfrench@dns1004: END - running authdns-update * 13:39 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1325532{{!}}srwiki: Allow bureaucrats to add and remove event-organizer group (T434748)]] * 13:39 swfrench@dns1004: START - running authdns-update * 13:37 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327111{{!}}Special:AbuseReview: Add "no further action needed" review action (T435020)]], [[gerrit:1327110{{!}}AbuseReview: Take the review verdict as a REST path parameter (T435020)]] (duration: 31m 43s) * 13:31 swfrench-wmf: starting etcd-main codfw -> eqiad switchover - [[phab:T435103|T435103]] * 13:28 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:28 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 13:25 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=97) rolling reboot on A:durum and A:durum * 13:24 kharlan@deploy1003: kharlan: Continuing with deployment * 13:24 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1147 * 13:24 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1147 * 13:23 kharlan@deploy1003: kharlan: Backport for [[gerrit:1327111{{!}}Special:AbuseReview: Add "no further action needed" review action (T435020)]], [[gerrit:1327110{{!}}AbuseReview: Take the review verdict as a REST path parameter (T435020)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:16 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:16 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 13:12 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:12 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 13:10 cdobbins@cumin1003: END (ERROR) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=97) Rolling upgrade of ATS on A:cp-codfw and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 13:10 cdobbins@cumin1003: END (ERROR) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=97) Rolling upgrade of ATS on A:cp-eqiad and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 13:06 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1327111{{!}}Special:AbuseReview: Add "no further action needed" review action (T435020)]], [[gerrit:1327110{{!}}AbuseReview: Take the review verdict as a REST path parameter (T435020)]] * 13:06 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-eqiad and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 13:05 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-codfw and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 13:05 cjd91: sudo -i cookbook sre.cdn.roll-upgrade-ats --query 'A:cp-eqiad' --task-id [[phab:T434478|T434478]] --reason '9.2.15 upgrade' * 13:03 swfrench@dns1004: END - running authdns-update * 13:01 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:01 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 13:01 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host ncmonitor1001.eqiad.wmnet * 13:00 swfrench@dns1004: START - running authdns-update * 12:59 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 12:59 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 12:59 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 12:59 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 12:57 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and A:durum * 12:53 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on db2902.codfw.wmnet with reason: Cloning * 12:48 cmooney@dns3003: END - running authdns-update * 12:48 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327096{{!}}Switch to redis lock manager on s4 and s8 (T366938)]] (duration: 09m 11s) * 12:46 cmooney@dns3003: START - running authdns-update * 12:45 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:45 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->eqord cct - cmooney@cumin1003" * 12:45 elukey: move the /v2/releng.* prefix on the Docker Registry to its new s3 backend - [[phab:T432829|T432829]] * 12:43 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 12:42 jelto@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 12:42 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->eqord cct - cmooney@cumin1003" * 12:42 jelto@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 12:41 jelto: update cert-manager to 1.19.6 on wikikube staging-codfw - [[phab:T427402|T427402]] * 12:40 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327096{{!}}Switch to redis lock manager on s4 and s8 (T366938)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:38 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327096{{!}}Switch to redis lock manager on s4 and s8 (T366938)]] * 12:38 blake@deploy1003: Finished scap sync-world: non-build deploy for [[phab:T417800|T417800]] (duration: 03m 52s) * 12:36 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 12:35 blake@deploy1003: Started scap sync-world: non-build deploy for [[phab:T417800|T417800]] * 12:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host krb2002.codfw.wmnet * 11:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host krb2002.codfw.wmnet * 11:49 moritzm: installing kerberos security updates * 11:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on stat1008.eqiad.wmnet with reason: host reimage * 11:44 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on stat1008.eqiad.wmnet with reason: host reimage * 11:31 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327084{{!}}Revert "Migrate database access to virtual domains" (T435305)]] (duration: 11m 02s) * 11:29 gkyziridis@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:29 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:29 kart_: Updated MinT to 2026-06-04-131507-production ([[phab:T321316|T321316]]) * 11:28 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/machinetranslation: apply * 11:28 gkyziridis@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:26 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 11:24 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:24 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:23 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:23 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:23 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/machinetranslation: apply * 11:22 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327084{{!}}Revert "Migrate database access to virtual domains" (T435305)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:21 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/machinetranslation: apply * 11:21 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:21 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:20 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327084{{!}}Revert "Migrate database access to virtual domains" (T435305)]] * 11:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1008.eqiad.wmnet with OS bookworm * 11:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet * 11:17 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/machinetranslation: apply * 11:13 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/machinetranslation: apply * 11:12 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:12 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:10 kartik@deploy1003: helmfile [staging] START helmfile.d/services/machinetranslation: apply * 11:08 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-cron: apply * 11:08 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/mw-cron: apply * 11:08 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply * 11:08 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply * 11:07 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet * 11:07 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host stat1008.eqiad.wmnet with OS bookworm * 11:06 moritzm: upgrading the new trixie URL downloaders to Squid 7.6 [[phab:T427282|T427282]] * 11:01 gkyziridis@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin2002.codfw.wmnet * 10:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin2002.codfw.wmnet * 10:45 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1169.eqiad.wmnet with OS bookworm * 10:42 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1185.eqiad.wmnet with OS bookworm * 10:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1169.eqiad.wmnet with reason: host reimage * 10:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1185.eqiad.wmnet with reason: host reimage * 10:14 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1169.eqiad.wmnet with reason: host reimage * 10:14 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1185.eqiad.wmnet with reason: host reimage * 10:11 jmm@cumin2003: END (PASS) - Cookbook sre.netbox.restart-reboot (exit_code=0) rolling reboot on A:netbox * 10:06 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1008.eqiad.wmnet with OS bookworm * 10:04 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 10:04 mpostoronca@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321579{{!}}Register the mediawiki.wikimedia_antiabuse.content_policy_score stream (T432848)]] (duration: 08m 53s) * 10:03 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 10:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 10:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 10:00 mpostoronca@deploy1003: mpostoronca: Continuing with deployment * 09:59 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1185.eqiad.wmnet with OS bookworm * 09:59 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1169.eqiad.wmnet with OS bookworm * 09:59 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.convert-disks (exit_code=0) for host ms-be1065 * 09:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1065.eqiad.wmnet with OS trixie * 09:58 mpostoronca@deploy1003: mpostoronca: Backport for [[gerrit:1321579{{!}}Register the mediawiki.wikimedia_antiabuse.content_policy_score stream (T432848)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:55 jmm@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netbox.discovery.wmnet. on all recursors * 09:55 mpostoronca@deploy1003: Started scap sync-world: Backport for [[gerrit:1321579{{!}}Register the mediawiki.wikimedia_antiabuse.content_policy_score stream (T432848)]] * 09:55 jmm@cumin2003: START - Cookbook sre.dns.wipe-cache netbox.discovery.wmnet. on all recursors * 09:52 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw2001.wikimedia.org with OS trixie * 09:51 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 09:51 jmm@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netbox.discovery.wmnet. on all recursors * 09:51 jmm@cumin2003: START - Cookbook sre.dns.wipe-cache netbox.discovery.wmnet. on all recursors * 09:51 jmm@cumin2003: START - Cookbook sre.netbox.restart-reboot rolling reboot on A:netbox * 09:46 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 09:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1056.eqiad.wmnet with OS trixie * 09:44 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 09:39 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 09:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 09:36 topranks: make HE transport circuits from magru live * 09:36 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 09:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb1003.eqiad.wmnet * 09:33 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage * 09:31 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb1003.eqiad.wmnet * 09:30 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 09:28 arnaudb@dns1006: END - running authdns-update * 09:27 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage * 09:27 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb2003.codfw.wmnet * 09:26 arnaudb@dns1006: START - running authdns-update * 09:24 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1056.eqiad.wmnet with reason: host reimage * 09:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb2003.codfw.wmnet * 09:20 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.convert-disks (exit_code=0) for host ms-be1068 * 09:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1068.eqiad.wmnet with OS trixie * 09:20 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 09:19 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "cloudvirt1057 - filippo@cumin1003" * 09:19 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "cloudvirt1057 - filippo@cumin1003" * 09:18 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1056.eqiad.wmnet with reason: host reimage * 09:18 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1057.eqiad.wmnet with OS trixie * 09:18 filippo@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 09:18 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 09:15 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 09:14 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 09:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host irc1003.wikimedia.org * 09:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:13 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1065.eqiad.wmnet with OS trixie * 09:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 09:09 moritzm: installing Postgresql security updates * 09:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host irc1003.wikimedia.org * 09:07 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 09:07 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:07 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1173.eqiad.wmnet with OS bookworm * 09:06 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw2001.wikimedia.org with OS trixie * 09:03 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.convert-disks (exit_code=0) for host ms-be1064 * 09:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1064.eqiad.wmnet with OS trixie * 09:03 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1056.eqiad.wmnet with OS trixie * 09:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1057.eqiad.wmnet with reason: host reimage * 09:01 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 08:58 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 08:56 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1057.eqiad.wmnet with reason: host reimage * 08:54 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1208.eqiad.wmnet with OS bookworm * 08:53 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2902.codfw.wmnet with OS trixie * 08:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 08:50 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1174.eqiad.wmnet with OS bookworm * 08:46 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1172.eqiad.wmnet with OS bookworm * 08:45 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1173.eqiad.wmnet with reason: host reimage * 08:41 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 08:40 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1057.eqiad.wmnet with OS trixie * 08:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1057.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 08:39 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1207.eqiad.wmnet with OS bookworm * 08:38 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2902.codfw.wmnet with reason: host reimage * 08:37 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 08:34 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1068.eqiad.wmnet with OS trixie * 08:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1208.eqiad.wmnet with reason: host reimage * 08:31 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1057.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 08:29 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 08:28 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1222.eqiad.wmnet onto db1276.eqiad.wmnet * 08:28 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1222: Pool db1222.eqiad.wmnet in after cloning * 08:28 fceratto@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2902.codfw.wmnet with reason: host reimage * 08:27 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1055.eqiad.wmnet with OS trixie * 08:27 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 08:26 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1174.eqiad.wmnet with reason: host reimage * 08:25 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 08:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-misc2002.codfw.wmnet * 08:23 topranks: reboot pfw1-codfw firewall pair to upgrade JunOS [[phab:T434865|T434865]] * 08:22 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1172.eqiad.wmnet with reason: host reimage * 08:20 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1064.eqiad.wmnet with OS trixie * 08:20 mvernon@cumin2003: START - Cookbook sre.swift.convert-disks for host ms-be1065 * 08:18 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1207.eqiad.wmnet with reason: host reimage * 08:17 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1174.eqiad.wmnet with reason: host reimage * 08:17 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1173.eqiad.wmnet with reason: host reimage * 08:17 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1172.eqiad.wmnet with reason: host reimage * 08:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host mc-misc2002.codfw.wmnet * 08:15 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1208.eqiad.wmnet with reason: host reimage * 08:15 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1207.eqiad.wmnet with reason: host reimage * 08:14 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db2902.codfw.wmnet with OS trixie * 08:14 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.16 refs [[phab:T430835|T430835]] * 08:13 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db2902.codfw.wmnet * 08:13 fceratto@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host db2902.codfw.wmnet with OS trixie * 08:10 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr[1-2]-codfw with reason: upgrade pfw1a-codfw and pfw1b-codfw pair * 08:09 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1055.eqiad.wmnet with reason: host reimage * 08:07 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on pfw1-codfw with reason: upgrade pfw1a-codfw and pfw1b-codfw pair * 08:03 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1055.eqiad.wmnet with reason: host reimage * 08:02 arnaudb@dns1006: END - running authdns-update * 08:02 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1208.eqiad.wmnet with OS bookworm * 08:02 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1207.eqiad.wmnet with OS bookworm * 08:01 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1174.eqiad.wmnet with OS bookworm * 08:01 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1173.eqiad.wmnet with OS bookworm * 08:01 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1172.eqiad.wmnet with OS bookworm * 07:59 arnaudb@dns1006: START - running authdns-update * 07:58 arnaudb@dns1006: START - running authdns-update * 07:48 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1055.eqiad.wmnet with OS trixie * 07:45 moritzm: extend the disk of ldap-rw2001 by 80G [[phab:T331699|T331699]] * 07:42 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1222: Pool db1222.eqiad.wmnet in after cloning * 07:36 mvernon@cumin2003: START - Cookbook sre.swift.convert-disks for host ms-be1068 * 07:35 mvernon@cumin2003: START - Cookbook sre.swift.convert-disks for host ms-be1064 * 07:17 moritzm: installing imagemagick security updates * 07:14 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1277: Pool back * 07:14 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon1003.wikimedia.org * 07:07 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon1003.wikimedia.org * 07:03 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1280: Pool back * 07:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon2002.wikimedia.org * 06:55 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon2002.wikimedia.org * 06:54 moritzm: installing php8.2 security updates * 06:51 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1284: Pool back * 06:49 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1222: Depool db1222.eqiad.wmnet to then clone it to db1276.eqiad.wmnet - marostegui@cumin1003 * 06:49 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1222: Depool db1222.eqiad.wmnet to then clone it to db1276.eqiad.wmnet - marostegui@cumin1003 * 06:49 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1222.eqiad.wmnet onto db1276.eqiad.wmnet * 06:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd1005.eqiad.wmnet * 06:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd1005.eqiad.wmnet * 06:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd1004.eqiad.wmnet * 06:36 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2209: db2209 repool * 06:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd1004.eqiad.wmnet * 06:32 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast3007.wikimedia.org * 06:29 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1277: Pool back * 06:28 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1277 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96190 and previous config saved to /var/cache/conftool/dbconfig/20260819-062815-marostegui.json * 06:26 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast3007.wikimedia.org * 06:22 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul1001.eqiad.wmnet * 06:18 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul1001.eqiad.wmnet * 06:18 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul1003.eqiad.wmnet * 06:18 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1280: Pool back * 06:17 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1284 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96186 and previous config saved to /var/cache/conftool/dbconfig/20260819-061743-marostegui.json * 06:14 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul1003.eqiad.wmnet * 06:14 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul1002.eqiad.wmnet * 06:10 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul1002.eqiad.wmnet * 06:10 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2048.codfw.wmnet * 06:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2048.codfw.wmnet * 06:06 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1284: Pool back * 06:06 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1284 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96184 and previous config saved to /var/cache/conftool/dbconfig/20260819-060621-marostegui.json * 06:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2048.codfw.wmnet * 05:59 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2048.codfw.wmnet * 05:51 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2209: db2209 repool * 03:16 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1186.eqiad.wmnet with OS bookworm * 02:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1186.eqiad.wmnet with reason: host reimage * 02:46 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1186.eqiad.wmnet with reason: host reimage * 02:46 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2207 [[phab:T435270|T435270]]', diff saved to https://phabricator.wikimedia.org/P96181 and previous config saved to /var/cache/conftool/dbconfig/20260819-024627-marostegui.json * 02:44 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2204 to s2 primary [[phab:T435270|T435270]]', diff saved to https://phabricator.wikimedia.org/P96180 and previous config saved to /var/cache/conftool/dbconfig/20260819-024403-marostegui.json * 02:43 marostegui: Starting s2 codfw failover from db2207 to db2204 - [[phab:T435270|T435270]] * 02:39 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2204 with weight 0 [[phab:T435270|T435270]]', diff saved to https://phabricator.wikimedia.org/P96179 and previous config saved to /var/cache/conftool/dbconfig/20260819-023951-marostegui.json * 02:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s2 [[phab:T435270|T435270]] * 02:32 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1186.eqiad.wmnet with OS bookworm * 02:29 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-worker1186.eqiad.wmnet with OS bookworm * 02:18 denisse@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2207: Depooling replica * 02:18 denisse@cumin1003: START - Cookbook sre.mysql.depool depool db2207: Depooling replica * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 48s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-18 == * 23:55 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1264845{{!}}Remove unused/redundant wgMFNoindexPages=true setting (T255458)]] (duration: 09m 42s) * 23:51 krinkle@deploy1003: krinkle: Continuing with deployment * 23:48 krinkle@deploy1003: krinkle: Backport for [[gerrit:1264845{{!}}Remove unused/redundant wgMFNoindexPages=true setting (T255458)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:45 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1264845{{!}}Remove unused/redundant wgMFNoindexPages=true setting (T255458)]] * 23:38 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326944{{!}}Retire filebackend lock manager in favour of the default one (T366938)]] (duration: 08m 55s) * 23:34 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 23:31 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326944{{!}}Retire filebackend lock manager in favour of the default one (T366938)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:29 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326944{{!}}Retire filebackend lock manager in favour of the default one (T366938)]] * 23:27 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1170.eqiad.wmnet with OS bookworm * 23:21 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1205.eqiad.wmnet with OS bookworm * 23:20 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1171.eqiad.wmnet with OS bookworm * 23:15 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1206.eqiad.wmnet with OS bookworm * 23:05 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1170.eqiad.wmnet with reason: host reimage * 23:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1205.eqiad.wmnet with reason: host reimage * 22:57 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1171.eqiad.wmnet with reason: host reimage * 22:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1206.eqiad.wmnet with reason: host reimage * 22:53 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1205.eqiad.wmnet with reason: host reimage * 22:51 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1171.eqiad.wmnet with reason: host reimage * 22:51 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1170.eqiad.wmnet with reason: host reimage * 22:50 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1206.eqiad.wmnet with reason: host reimage * 22:36 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1206.eqiad.wmnet with OS bookworm * 22:35 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1205.eqiad.wmnet with OS bookworm * 22:35 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1186.eqiad.wmnet with OS bookworm * 22:35 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1171.eqiad.wmnet with OS bookworm * 22:35 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1170.eqiad.wmnet with OS bookworm * 22:33 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-worker1194.eqiad.wmnet with OS bookworm * 22:22 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326923{{!}}Enable redis lock manager on s6 (T366938)]] (duration: 11m 52s) * 22:18 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 22:13 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326923{{!}}Enable redis lock manager on s6 (T366938)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:10 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326923{{!}}Enable redis lock manager on s6 (T366938)]] * 22:04 sbassett: Deployed security fix for [[phab:T435234|T435234]] (wmf.16) * 21:54 sbassett: Deployed security fix for [[phab:T435234|T435234]] (wmf.15) * 21:38 caro@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326925{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326926{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326929{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]], [[gerrit:1326928{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]] (duration: 0 * 21:34 caro@deploy1003: caro: Continuing with deployment * 21:33 caro@deploy1003: caro: Backport for [[gerrit:1326925{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326926{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326929{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]], [[gerrit:1326928{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]] synced to the testservers (see h * 21:31 caro@deploy1003: Started scap sync-world: Backport for [[gerrit:1326925{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326926{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326929{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]], [[gerrit:1326928{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]] * 21:24 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1204.eqiad.wmnet with reason: 1204 datanode repair [[phab:T434494|T434494]] * 21:02 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326896{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]], [[gerrit:1326897{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]] (duration: 13m 56s) * 20:58 krinkle@deploy1003: krinkle: Continuing with deployment * 20:50 krinkle@deploy1003: krinkle: Backport for [[gerrit:1326896{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]], [[gerrit:1326897{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:49 ryankemper: `an-launcher1003` terminated process group `666809` (`rest_backfill_phase1.sh`) ~20 mins ago with `sudo kill -TERM -- -666809` after its local spark driver (`--driver-memory 64g`) repeatedly exhausted memory on the 32 GB VM and caused SSH to intermittently flap; host recovered to 27 GB available memory * 20:48 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1326896{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]], [[gerrit:1326897{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]] * 20:35 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 20:33 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 20:31 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 20:31 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326870{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]], [[gerrit:1326871{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]] (duration: 07m 35s) * 20:28 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 20:26 kemayo@deploy1003: kemayo: Continuing with deployment * 20:26 ryankemper: `an-launcher1003` confirmed the host is flapping because of memory thrash. chasing down the source of the thrash * 20:25 kemayo@deploy1003: kemayo: Backport for [[gerrit:1326870{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]], [[gerrit:1326871{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:23 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1326870{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]], [[gerrit:1326871{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]] * 20:23 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 20:20 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 20:14 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1168.eqiad.wmnet with OS bookworm * 20:08 zabe: zabe@deploy1003:~$ mwscript extensions/WikimediaMaintenance/maintenance/fixFileRevisionArchiveNameDrift.php enwiki # [[phab:T428406|T428406]] * 20:08 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1204.eqiad.wmnet with OS bookworm * 20:05 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1167.eqiad.wmnet with OS bookworm * 20:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1166.eqiad.wmnet with OS bookworm * 19:54 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1203.eqiad.wmnet with OS bookworm * 19:53 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326907{{!}}Revert "Disable redis lock manager on testwiki"]] (duration: 11m 05s) * 19:50 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1168.eqiad.wmnet with reason: host reimage * 19:47 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1204.eqiad.wmnet with reason: host reimage * 19:46 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 19:44 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326907{{!}}Revert "Disable redis lock manager on testwiki"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:42 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326907{{!}}Revert "Disable redis lock manager on testwiki"]] * 19:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1167.eqiad.wmnet with reason: host reimage * 19:37 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1166.eqiad.wmnet with reason: host reimage * 19:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1203.eqiad.wmnet with reason: host reimage * 19:32 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1167.eqiad.wmnet with reason: host reimage * 19:32 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1168.eqiad.wmnet with reason: host reimage * 19:32 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1166.eqiad.wmnet with reason: host reimage * 19:31 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1204.eqiad.wmnet with reason: host reimage * 19:31 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1203.eqiad.wmnet with reason: host reimage * 19:26 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324751{{!}}InitialiseSettings: Enable 2FA enforcement on more private wikis (T428103)]], [[gerrit:1326875{{!}}Add banner notifying of upcoming 2FA enforcement (T420792)]] (duration: 31m 46s) * 19:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1204.eqiad.wmnet with OS bookworm * 19:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1203.eqiad.wmnet with OS bookworm * 19:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1168.eqiad.wmnet with OS bookworm * 19:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1167.eqiad.wmnet with OS bookworm * 19:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1166.eqiad.wmnet with OS bookworm * 19:15 denisse: rebooting kafkamon2003.codfw.wmnet - [[phab:T435162|T435162]] * 19:14 denisse: rebooting kafkamon1003.eqiad.wmnet [[phab:T435162|T435162]] * 19:13 reedy@deploy1003: reedy: Continuing with deployment * 19:12 reedy@deploy1003: reedy: Backport for [[gerrit:1324751{{!}}InitialiseSettings: Enable 2FA enforcement on more private wikis (T428103)]], [[gerrit:1326875{{!}}Add banner notifying of upcoming 2FA enforcement (T420792)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:54 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324751{{!}}InitialiseSettings: Enable 2FA enforcement on more private wikis (T428103)]], [[gerrit:1326875{{!}}Add banner notifying of upcoming 2FA enforcement (T420792)]] * 18:50 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 18:44 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.16 refs [[phab:T430835|T430835]] * 18:34 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 18:31 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 18:22 aklapper@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326852{{!}}CategoryTree: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]], [[gerrit:1326853{{!}}CategoryViewer: Allow null $html in the CategoryViewerGenerateLink hook (T435161)]], [[gerrit:1326865{{!}}Flow: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]] (duration: 09m 57s) * 18:18 aklapper@deploy1003: jforrester, aklapper: Continuing with deployment * 18:17 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 18:14 aklapper@deploy1003: jforrester, aklapper: Backport for [[gerrit:1326852{{!}}CategoryTree: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]], [[gerrit:1326853{{!}}CategoryViewer: Allow null $html in the CategoryViewerGenerateLink hook (T435161)]], [[gerrit:1326865{{!}}Flow: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki * 18:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1156.eqiad.wmnet with OS bookworm * 18:12 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 18:12 aklapper@deploy1003: Started scap sync-world: Backport for [[gerrit:1326852{{!}}CategoryTree: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]], [[gerrit:1326853{{!}}CategoryViewer: Allow null $html in the CategoryViewerGenerateLink hook (T435161)]], [[gerrit:1326865{{!}}Flow: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]] * 18:11 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 18:08 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1146.eqiad.wmnet with OS bookworm * 18:07 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1177.eqiad.wmnet with OS bookworm * 18:00 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326380{{!}}Introduce main lock manager service (T366938 T427999)]] (duration: 11m 25s) * 17:58 ladsgroup@deploy1003: ladsgroup: Rolling back deployment * 17:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1202.eqiad.wmnet with OS bookworm * 17:55 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1201.eqiad.wmnet with OS bookworm * 17:50 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326380{{!}}Introduce main lock manager service (T366938 T427999)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:48 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326380{{!}}Introduce main lock manager service (T366938 T427999)]] * 17:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1156.eqiad.wmnet with reason: host reimage * 17:46 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1146.eqiad.wmnet with reason: host reimage * 17:45 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-esams and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 17:42 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 17:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1177.eqiad.wmnet with reason: host reimage * 17:38 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1202.eqiad.wmnet with reason: host reimage * 17:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1201.eqiad.wmnet with reason: host reimage * 17:30 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1156.eqiad.wmnet with reason: host reimage * 17:29 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1177.eqiad.wmnet with reason: host reimage * 17:28 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1146.eqiad.wmnet with reason: host reimage * 17:28 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1202.eqiad.wmnet with reason: host reimage * 17:27 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1201.eqiad.wmnet with reason: host reimage * 17:25 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 17:21 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 17:14 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1202.eqiad.wmnet with OS bookworm * 17:14 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1201.eqiad.wmnet with OS bookworm * 17:14 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1177.eqiad.wmnet with OS bookworm * 17:14 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1156.eqiad.wmnet with OS bookworm * 17:14 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1146.eqiad.wmnet with OS bookworm * 17:09 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326885{{!}}w/deployment-info.php: Handle new file format (T434726)]] (duration: 07m 15s) * 17:08 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 17:05 dancy@deploy1003: dancy: Continuing with deployment * 17:04 dancy@deploy1003: dancy: Backport for [[gerrit:1326885{{!}}w/deployment-info.php: Handle new file format (T434726)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:02 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1326885{{!}}w/deployment-info.php: Handle new file format (T434726)]] * 16:50 dancy@deploy1003: Finished scap sync-world: Testing [[phab:T434726|T434726]] (duration: 06m 40s) * 16:43 dancy@deploy1003: Started scap sync-world: Testing [[phab:T434726|T434726]] * 16:43 dancy@deploy1003: Installation of scap version "4.282.0" completed for 3 hosts * 16:41 dancy@deploy1003: Installing scap version "4.282.0" for 3 host(s) * 16:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1145.eqiad.wmnet with OS bookworm * 16:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1200.eqiad.wmnet with OS bookworm * 16:32 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1199.eqiad.wmnet with OS bookworm * 16:18 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1145.eqiad.wmnet with reason: host reimage * 16:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1200.eqiad.wmnet with reason: host reimage * 16:09 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1199.eqiad.wmnet with reason: host reimage * 16:05 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-esams and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 16:05 cjd91: sudo -i cookbook sre.cdn.roll-upgrade-ats --query 'A:cp-esams' --task-id [[phab:T434478|T434478]] --reason '9.2.15 upgrade' * 16:03 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1200.eqiad.wmnet with reason: host reimage * 16:02 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1145.eqiad.wmnet with reason: host reimage * 16:02 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1199.eqiad.wmnet with reason: host reimage * 15:50 moritzm: installing zip security updates * 15:48 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1200.eqiad.wmnet with OS bookworm * 15:47 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1199.eqiad.wmnet with OS bookworm * 15:47 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1145.eqiad.wmnet with OS bookworm * 15:41 topranks: bounce PIC 0/0 on cr1-magru to set port to 40G * 15:37 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1236.eqiad.wmnet with OS bookworm * 15:29 aikochou@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'ores-legacy' for release 'main' . * 15:26 aikochou@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'ores-legacy' for release 'main' . * 15:20 aikochou@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'ores-legacy' for release 'main' . * 15:14 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2002.codfw.wmnet * 15:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1236.eqiad.wmnet with reason: host reimage * 15:12 moritzm: failover ganeti master in codfw to ganeti2047 * 15:09 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1236.eqiad.wmnet with reason: host reimage * 15:09 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2044.codfw.wmnet * 15:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2002.codfw.wmnet * 15:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2044.codfw.wmnet * 15:04 brennen@deploy1003: Finished deploy [phabricator/deployment@6b9b6ff]: deploy phab1004 for [[phab:T435213|T435213]] (duration: 01m 01s) * 15:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2044.codfw.wmnet * 15:03 brennen@deploy1003: Started deploy [phabricator/deployment@6b9b6ff]: deploy phab1004 for [[phab:T435213|T435213]] * 15:03 brennen@deploy1003: Finished deploy [phabricator/deployment@6b9b6ff]: deploy phab2003 for [[phab:T435213|T435213]] (duration: 00m 57s) * 15:02 brennen@deploy1003: Started deploy [phabricator/deployment@6b9b6ff]: deploy phab2003 for [[phab:T435213|T435213]] * 14:57 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply * 14:57 arnaudb@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on phab2003.codfw.wmnet,phab[1004-1006].eqiad.wmnet with reason: maintenance * 14:56 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2044.codfw.wmnet * 14:55 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply * 14:53 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1236.eqiad.wmnet with OS bookworm * 14:51 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2043.codfw.wmnet * 14:51 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2043.codfw.wmnet * 14:45 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2043.codfw.wmnet * 14:33 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2043.codfw.wmnet * 14:22 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2042.codfw.wmnet * 14:22 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2042.codfw.wmnet * 14:21 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db2902.codfw.wmnet with OS trixie * 14:16 elukey: uploaded spicerack_13.2.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia * 14:16 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2042.codfw.wmnet * 14:04 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2042.codfw.wmnet * 14:04 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2041.codfw.wmnet * 14:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2041.codfw.wmnet * 13:59 phuedx@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: apply * 13:59 phuedx@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-main: apply * 13:59 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326838{{!}}Enable Suggested Investigations on hewiki (T435146)]] (duration: 11m 50s) * 13:58 phuedx@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: apply * 13:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2041.codfw.wmnet * 13:57 phuedx@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-main: apply * 13:57 phuedx@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-main: apply * 13:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard1003.eqiad.wmnet * 13:57 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-main: apply * 13:55 phuedx@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-logging-external: apply * 13:55 phuedx@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-logging-external: apply * 13:54 phuedx@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-logging-external: apply * 13:54 stran@deploy1003: stran: Continuing with deployment * 13:54 phuedx@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-logging-external: apply * 13:54 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-logging-external: apply * 13:54 phuedx@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-logging-external: apply * 13:53 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-logging-external: apply * 13:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard1003.eqiad.wmnet * 13:53 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2041.codfw.wmnet * 13:51 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db2902.codfw.wmnet - fceratto@cumin1003" * 13:51 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db2902.codfw.wmnet - fceratto@cumin1003" * 13:51 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2026.codfw.wmnet * 13:51 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard2003.codfw.wmnet * 13:50 phuedx@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: apply * 13:50 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-eqiad * 13:50 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp1001.eqiad.wmnet * 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp1001.eqiad.wmnet * 13:50 phuedx@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: apply * 13:49 phuedx@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: apply * 13:49 stran@deploy1003: stran: Backport for [[gerrit:1326838{{!}}Enable Suggested Investigations on hewiki (T435146)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:48 phuedx@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: apply * 13:48 phuedx@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics: apply * 13:47 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard2003.codfw.wmnet * 13:47 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics: apply * 13:47 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1326838{{!}}Enable Suggested Investigations on hewiki (T435146)]] * 13:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp1001.eqiad.wmnet * 13:44 cdanis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 13:43 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp1001.eqiad.wmnet * 13:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1315-1327].eqiad.wmnet * 13:43 cdanis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 13:43 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1315-1327].eqiad.wmnet * 13:40 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki1001.eqiad.wmnet * 13:35 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1315-1327].eqiad.wmnet * 13:34 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host rpki1001.eqiad.wmnet * 13:27 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1315-1327].eqiad.wmnet * 13:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1302-1314].eqiad.wmnet * 13:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1302-1314].eqiad.wmnet * 13:21 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326775{{!}}SI: Instrument case update on first edit (T435048)]] (duration: 07m 12s) * 13:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1302-1314].eqiad.wmnet * 13:17 stran@deploy1003: stran: Continuing with deployment * 13:16 stran@deploy1003: stran: Backport for [[gerrit:1326775{{!}}SI: Instrument case update on first edit (T435048)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:14 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1326775{{!}}SI: Instrument case update on first edit (T435048)]] * 13:13 moritzm: installing util-linux security updates * 13:11 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1302-1314].eqiad.wmnet * 13:10 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326232{{!}}prv: Enable parsoid rendering for 5 wikis (T435115)]] (duration: 08m 17s) * 13:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1288-1289,1291-1301].eqiad.wmnet * 13:10 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply * 13:10 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1288-1289,1291-1301].eqiad.wmnet * 13:10 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply * 13:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki2003.codfw.wmnet * 13:06 jgiannelos@deploy1003: jgiannelos: Continuing with deployment * 13:05 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host rpki2003.codfw.wmnet * 13:04 jgiannelos@deploy1003: jgiannelos: Backport for [[gerrit:1326232{{!}}prv: Enable parsoid rendering for 5 wikis (T435115)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1288-1289,1291-1301].eqiad.wmnet * 13:02 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1326232{{!}}prv: Enable parsoid rendering for 5 wikis (T435115)]] * 12:53 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1288-1289,1291-1301].eqiad.wmnet * 12:53 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1273,1275-1287].eqiad.wmnet * 12:53 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1273,1275-1287].eqiad.wmnet * 12:52 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt1002.wikimedia.org * 12:51 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db2902.codfw.wmnet on all recursors * 12:51 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db2902.codfw.wmnet on all recursors * 12:51 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:51 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db2902.codfw.wmnet - fceratto@cumin1003" * 12:51 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db2902.codfw.wmnet - fceratto@cumin1003" * 12:46 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt1002.wikimedia.org * 12:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt2002.wikimedia.org * 12:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1273,1275-1287].eqiad.wmnet * 12:42 dhinus: repooled clouddb1032 that was currently <nowiki>{</nowiki>"weight": 0, "pooled": "inactive"<nowiki>}</nowiki> for both s4 and s6 * 12:41 dhinus: also depooled clouddb1017 (forgot it in the previous list) * 12:41 fnegri@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet * 12:40 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 12:40 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db2902.codfw.wmnet * 12:40 fnegri@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032.eqiad.wmnet * 12:40 fnegri@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032 * 12:39 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1017.eqiad.wmnet * 12:39 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt2002.wikimedia.org * 12:38 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1020.eqiad.wmnet * 12:38 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1018.eqiad.wmnet * 12:38 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet * 12:37 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1014.eqiad.wmnet * 12:37 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1013.eqiad.wmnet * 12:37 dhinus: depool again clouddb10[13,14,16,18,20] that were repooled by the cookbook sre.mysql.multiinstance_reboot * 12:36 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1273,1275-1287].eqiad.wmnet * 12:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1248-1261].eqiad.wmnet * 12:36 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1248-1261].eqiad.wmnet * 12:35 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2026.codfw.wmnet * 12:34 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 12:34 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 12:34 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 12:34 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 12:29 lucaswerkmeister-wmde@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 12:28 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2026.codfw.wmnet * 12:28 lucaswerkmeister-wmde@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 12:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1248-1261].eqiad.wmnet * 12:27 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.addnode (exit_code=0) for new host ganeti2046.codfw.wmnet to cluster codfw and group A * 12:26 moritzm: readded ganeti2046 to the codfw cluster following firmware update and reimage [[phab:T434681|T434681]] * 12:23 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply * 12:23 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply * 12:23 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply * 12:22 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply * 12:22 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply * 12:22 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply * 12:21 jmm@cumin2003: START - Cookbook sre.ganeti.addnode for new host ganeti2046.codfw.wmnet to cluster codfw and group A * 12:21 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1169.eqiad.wmnet onto db1283.eqiad.wmnet * 12:21 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1169: Pool db1169.eqiad.wmnet in after cloning * 12:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1248-1261].eqiad.wmnet * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1149-1153,1158,1240-1247].eqiad.wmnet * 12:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1149-1153,1158,1240-1247].eqiad.wmnet * 12:11 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1149-1153,1158,1240-1247].eqiad.wmnet * 12:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1168.eqiad.wmnet onto db1282.eqiad.wmnet * 12:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1168: Pool db1168.eqiad.wmnet in after cloning * 12:06 jmm@cumin2003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-test-eqiad * 12:05 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2026.codfw.wmnet * 12:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1149-1153,1158,1240-1247].eqiad.wmnet * 12:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1128-1134,1142-1148].eqiad.wmnet * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2046.codfw.wmnet * 12:01 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1128-1134,1142-1148].eqiad.wmnet * 11:57 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2025.codfw.wmnet * 11:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2025.codfw.wmnet * 11:54 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2046.codfw.wmnet * 11:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1128-1134,1142-1148].eqiad.wmnet * 11:50 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2025.codfw.wmnet * 11:46 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1128-1134,1142-1148].eqiad.wmnet * 11:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1114-1127].eqiad.wmnet * 11:45 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1114-1127].eqiad.wmnet * 11:39 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db2901.codfw.wmnet * 11:39 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db2901.codfw.wmnet * 11:36 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2025.codfw.wmnet * 11:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1114-1127].eqiad.wmnet * 11:35 fceratto@cumin1003: END (ERROR) - Cookbook sre.ganeti.makevm (exit_code=93) for new host db1901.eqiad.wmnet * 11:35 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1169: Pool db1169.eqiad.wmnet in after cloning * 11:35 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 11:30 jmm@cumin2003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-test-eqiad * 11:27 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1114-1127].eqiad.wmnet * 11:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1076-1081,1084-1087,1093-1095,1113].eqiad.wmnet * 11:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1076-1081,1084-1087,1093-1095,1113].eqiad.wmnet * 11:25 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1168: Pool db1168.eqiad.wmnet in after cloning * 11:24 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1194.eqiad.wmnet with OS bookworm * 11:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1181.eqiad.wmnet with OS bookworm * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2050.codfw.wmnet * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2050.codfw.wmnet * 11:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1076-1081,1084-1087,1093-1095,1113].eqiad.wmnet * 11:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2050.codfw.wmnet * 11:14 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2004.codfw.wmnet * 11:10 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2050.codfw.wmnet * 11:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1076-1081,1084-1087,1093-1095,1113].eqiad.wmnet * 11:09 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1286: Pool back * 11:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1045-1050,1056-1057,1064-1066,1073-1075].eqiad.wmnet * 11:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1045-1050,1056-1057,1064-1066,1073-1075].eqiad.wmnet * 11:08 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2004.codfw.wmnet * 11:07 moritzm: installing PHP 8.4 security updates * 11:06 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on an-worker1194.eqiad.wmnet with reason: host reimage * 11:06 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1194.eqiad.wmnet with reason: host reimage * 11:05 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2049.codfw.wmnet * 11:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2049.codfw.wmnet * 11:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid1003.eqiad.wmnet * 11:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1045-1050,1056-1057,1064-1066,1073-1075].eqiad.wmnet * 10:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1181.eqiad.wmnet with reason: host reimage * 10:59 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2049.codfw.wmnet * 10:59 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid1003.eqiad.wmnet * 10:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid2003.codfw.wmnet * 10:54 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1181.eqiad.wmnet with reason: host reimage * 10:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid2003.codfw.wmnet * 10:51 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1901.eqiad.wmnet on all recursors * 10:51 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1901.eqiad.wmnet on all recursors * 10:51 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:51 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:51 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:50 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2049.codfw.wmnet * 10:50 blake@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:50 blake@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:49 blake@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:48 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1045-1050,1056-1057,1064-1066,1073-1075].eqiad.wmnet * 10:48 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1044].eqiad.wmnet * 10:48 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1044].eqiad.wmnet * 10:47 blake@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:46 blake@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:46 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2047.codfw.wmnet * 10:46 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:46 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2047.codfw.wmnet * 10:46 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1901.eqiad.wmnet * 10:46 blake@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:44 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1168: Depool db1168.eqiad.wmnet to then clone it to db1282.eqiad.wmnet - marostegui@cumin1003 * 10:44 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1168: Depool db1168.eqiad.wmnet to then clone it to db1282.eqiad.wmnet - marostegui@cumin1003 * 10:44 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1168.eqiad.wmnet onto db1282.eqiad.wmnet * 10:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2047.codfw.wmnet * 10:40 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1044].eqiad.wmnet * 10:37 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2047.codfw.wmnet * 10:37 fceratto@cumin1003: END (ERROR) - Cookbook sre.ganeti.makevm (exit_code=93) for new host db1901.eqiad.wmnet * 10:36 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 10:34 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:33 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2032.codfw.wmnet * 10:33 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2032.codfw.wmnet * 10:32 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1044].eqiad.wmnet * 10:32 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-eqiad * 10:27 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2032.codfw.wmnet * 10:24 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1286: Pool back * 10:24 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1286 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96163 and previous config saved to /var/cache/conftool/dbconfig/20260818-102431-marostegui.json * 10:22 moritzm: installing Django security updates * 10:20 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul2002.codfw.wmnet * 10:20 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul2001.codfw.wmnet * 10:17 blake@deploy1003: Finished scap sync-world: no-build deployment for [[phab:T417800|T417800]] (duration: 04m 40s) * 10:16 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul2002.codfw.wmnet * 10:16 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul2001.codfw.wmnet * 10:15 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2032.codfw.wmnet * 10:14 blake@deploy1003: Started scap sync-world: no-build deployment for [[phab:T417800|T417800]] * 10:12 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul2003.codfw.wmnet * 10:12 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy2001.codfw.wmnet * 10:12 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy3001.esams.wmnet * 10:12 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy1001.eqiad.wmnet * 10:08 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2031.codfw.wmnet * 10:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2031.codfw.wmnet * 10:08 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul2003.codfw.wmnet * 10:08 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy2001.codfw.wmnet * 10:08 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy3001.esams.wmnet * 10:08 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy1001.eqiad.wmnet * 10:07 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy1002.eqiad.wmnet * 10:05 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy2002.codfw.wmnet * 10:05 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy3002.esams.wmnet * 10:04 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy1002.eqiad.wmnet * 10:03 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy4003.ulsfo.wmnet * 10:03 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy5003.eqsin.wmnet * 10:02 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2031.codfw.wmnet * 10:02 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1169: Depool db1169.eqiad.wmnet to then clone it to db1283.eqiad.wmnet - marostegui@cumin1003 * 10:01 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy2002.codfw.wmnet * 10:01 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy4004.ulsfo.wmnet * 10:01 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy3002.esams.wmnet * 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1169: Depool db1169.eqiad.wmnet to then clone it to db1283.eqiad.wmnet - marostegui@cumin1003 * 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1169.eqiad.wmnet onto db1283.eqiad.wmnet * 10:01 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy5004.eqsin.wmnet * 09:59 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy4003.ulsfo.wmnet * 09:59 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy4004.ulsfo.wmnet * 09:59 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy5003.eqsin.wmnet * 09:59 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy7001.magru.wmnet * 09:59 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy5004.eqsin.wmnet * 09:58 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy7002.magru.wmnet * 09:57 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2031.codfw.wmnet * 09:57 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy6001.drmrs.wmnet * 09:57 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy6002.drmrs.wmnet * 09:54 filippo@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for 10 hosts * 09:54 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast4006.wikimedia.org * 09:53 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy6001.drmrs.wmnet * 09:53 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy6002.drmrs.wmnet * 09:53 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host releases1003.eqiad.wmnet * 09:53 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy7001.magru.wmnet * 09:53 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host people2004.codfw.wmnet * 09:52 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host people1005.eqiad.wmnet * 09:52 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy7002.magru.wmnet * 09:50 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host releases2003.codfw.wmnet * 09:49 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host releases2003.codfw.wmnet * 09:49 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host releases1003.eqiad.wmnet * 09:49 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host people2004.codfw.wmnet * 09:48 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host people1005.eqiad.wmnet * 09:46 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2040.codfw.wmnet * 09:46 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2040.codfw.wmnet * 09:43 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1901.eqiad.wmnet on all recursors * 09:42 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1901.eqiad.wmnet on all recursors * 09:42 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:42 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:42 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2040.codfw.wmnet * 09:40 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp2005.wikimedia.org * 09:36 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp2005.wikimedia.org * 09:31 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:31 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1901.eqiad.wmnet * 09:31 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1901.eqiad.wmnet * 09:31 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:31 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1901.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 09:31 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1901.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 09:29 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2040.codfw.wmnet * 09:29 slyngshede@dns1004: END - running authdns-update * 09:28 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2039.codfw.wmnet * 09:28 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2039.codfw.wmnet * 09:27 slyngshede@dns1004: START - running authdns-update * 09:26 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test2005.wikimedia.org * 09:22 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2039.codfw.wmnet * 09:22 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test2005.wikimedia.org * 09:22 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp1005.wikimedia.org * 09:21 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:19 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2039.codfw.wmnet * 09:19 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2038.codfw.wmnet * 09:18 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1287: Pool back * 09:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2038.codfw.wmnet * 09:18 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp1005.wikimedia.org * 09:18 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test1005.wikimedia.org * 09:17 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1901.eqiad.wmnet * 09:15 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db1901.eqiad.wmnet * 09:15 fceratto@cumin1003: END (ERROR) - Cookbook sre.dns.netbox (exit_code=97) * 09:14 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test1005.wikimedia.org * 09:13 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2038.codfw.wmnet * 09:13 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:13 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1901.eqiad.wmnet * 09:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast4006.wikimedia.org * 09:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1288: Pool back * 09:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast5005.wikimedia.org * 09:03 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2038.codfw.wmnet * 08:57 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2037.codfw.wmnet * 08:57 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast5005.wikimedia.org * 08:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2037.codfw.wmnet * 08:56 filippo@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 10 hosts * 08:52 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2037.codfw.wmnet * 08:51 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1289: Pool back * 08:35 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudidp2001-dev.codfw.wmnet * 08:34 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1002-dev.eqiad.wmnet * 08:33 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1287: Pool back * 08:33 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1287 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96150 and previous config saved to /var/cache/conftool/dbconfig/20260818-083311-marostegui.json * 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1001-dev.eqiad.wmnet * 08:31 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudidp2001-dev.codfw.wmnet * 08:30 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1002-dev.eqiad.wmnet * 08:30 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2037.codfw.wmnet * 08:29 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1001-dev.eqiad.wmnet * 08:28 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2036.codfw.wmnet * 08:28 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2036.codfw.wmnet * 08:25 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1288: Pool back * 08:23 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 08:23 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1181.eqiad.wmnet with OS bookworm * 08:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2036.codfw.wmnet * 08:22 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1288 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96148 and previous config saved to /var/cache/conftool/dbconfig/20260818-082234-marostegui.json * 08:20 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2036.codfw.wmnet * 08:18 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2035.codfw.wmnet * 08:18 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudvirt1057.eqiad.wmnet with OS trixie * 08:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2035.codfw.wmnet * 08:13 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2035.codfw.wmnet * 08:06 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2035.codfw.wmnet * 08:05 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1289: Pool back * 08:05 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1289 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96145 and previous config saved to /var/cache/conftool/dbconfig/20260818-080531-marostegui.json * 07:54 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 07:51 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326439{{!}}Fix "mathjax_ignore" handling around forcemathmode attribute (T434686)]] (duration: 13m 15s) * 07:48 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 07:47 krinkle@deploy1003: krinkle: Continuing with deployment * 07:44 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ganeti2046.codfw.wmnet with OS bookworm * 07:40 krinkle@deploy1003: krinkle: Backport for [[gerrit:1326439{{!}}Fix "mathjax_ignore" handling around forcemathmode attribute (T434686)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:39 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 07:38 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1326439{{!}}Fix "mathjax_ignore" handling around forcemathmode attribute (T434686)]] * 07:34 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 07:31 samwilson@deploy1003: Finished scap sync-world: Backport for [[gerrit:701016{{!}}InitialiseSettings and -labs: Remove redundant feature flag $wgWikisourceEnableOcr (T285311)]] (duration: 07m 47s) * 07:30 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1048.eqiad.wmnet * 07:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1048.eqiad.wmnet * 07:29 XioNoX: add gnmic 0.47.0 to bookworm and trixie reprepro * 07:28 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ganeti2046.codfw.wmnet with reason: host reimage * 07:27 samwilson@deploy1003: samwilson: Continuing with deployment * 07:25 samwilson@deploy1003: samwilson: Backport for [[gerrit:701016{{!}}InitialiseSettings and -labs: Remove redundant feature flag $wgWikisourceEnableOcr (T285311)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:25 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 07:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1048.eqiad.wmnet * 07:24 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ganeti2046.codfw.wmnet with reason: host reimage * 07:23 samwilson@deploy1003: Started scap sync-world: Backport for [[gerrit:701016{{!}}InitialiseSettings and -labs: Remove redundant feature flag $wgWikisourceEnableOcr (T285311)]] * 07:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1198.eqiad.wmnet with OS bookworm * 07:18 samwilson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326450{{!}}InitialiseSettings.php: Enable Bulk OCR on pawikisource (T434648)]] (duration: 12m 17s) * 07:17 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 07:11 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1048.eqiad.wmnet * 07:11 samwilson@deploy1003: samwilson: Continuing with deployment * 07:11 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ganeti2046.codfw.wmnet with OS bookworm * 07:10 samwilson@deploy1003: samwilson: Backport for [[gerrit:1326450{{!}}InitialiseSettings.php: Enable Bulk OCR on pawikisource (T434648)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1057.eqiad.wmnet with OS trixie * 07:05 samwilson@deploy1003: Started scap sync-world: Backport for [[gerrit:1326450{{!}}InitialiseSettings.php: Enable Bulk OCR on pawikisource (T434648)]] * 07:02 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1198.eqiad.wmnet with reason: host reimage * 07:01 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1056.eqiad.wmnet with OS trixie * 07:01 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1056.eqiad.wmnet with OS trixie * 07:00 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1056.eqiad.wmnet with OS trixie * 07:00 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1056.eqiad.wmnet with OS trixie * 06:59 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudvirt1055.eqiad.wmnet with OS trixie * 06:58 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1198.eqiad.wmnet with reason: host reimage * 06:53 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1055.eqiad.wmnet with OS trixie * 06:52 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudvirt1054.eqiad.wmnet with OS trixie * 06:44 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1194.eqiad.wmnet with OS bookworm * 06:44 XioNoX: upgrade eqsin gnmic to 0.47.0 * 06:43 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1198.eqiad.wmnet with OS bookworm * 06:41 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1054.eqiad.wmnet with OS trixie * 06:09 arnaudb@cumin1003: END (PASS) - Cookbook sre.gerrit.restart-gerrit (exit_code=0) Restarting Gerrit on gerrit2002 * 06:06 arnaudb@cumin1003: START - Cookbook sre.gerrit.restart-gerrit Restarting Gerrit on gerrit2002 * 06:06 arnaudb@cumin1003: END (PASS) - Cookbook sre.gerrit.restart-gerrit (exit_code=0) Restarting Gerrit on gerrit1003 * 06:04 arnaudb@cumin1003: START - Cookbook sre.gerrit.restart-gerrit Restarting Gerrit on gerrit1003 * 06:02 arnaudb@cumin1003: END (PASS) - Cookbook sre.gerrit.restart-gerrit (exit_code=0) Restarting Gerrit on gerrit2003 * 06:00 arnaudb@cumin1003: START - Cookbook sre.gerrit.restart-gerrit Restarting Gerrit on gerrit2003 * 05:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1155.eqiad.wmnet with OS bookworm * 05:38 arnaudb: updating prometheusBearerToken on gerrit * 05:28 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1155.eqiad.wmnet with reason: host reimage * 05:23 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1155.eqiad.wmnet with reason: host reimage * 05:06 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1155.eqiad.wmnet with OS bookworm * 04:57 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1144.eqiad.wmnet with OS bookworm * 04:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1144.eqiad.wmnet with reason: host reimage * 04:29 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1144.eqiad.wmnet with reason: host reimage * 04:14 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1144.eqiad.wmnet with OS bookworm * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.13 (duration: 02m 23s) * 03:45 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1197.eqiad.wmnet with OS bookworm * 03:41 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1165.eqiad.wmnet with OS bookworm * 03:38 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.16 refs [[phab:T430835|T430835]] (duration: 34m 43s) * 03:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1164.eqiad.wmnet with OS bookworm * 03:35 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1196.eqiad.wmnet with OS bookworm * 03:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1163.eqiad.wmnet with OS bookworm * 03:22 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1197.eqiad.wmnet with reason: host reimage * 03:18 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1165.eqiad.wmnet with reason: host reimage * 03:15 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1196.eqiad.wmnet with reason: host reimage * 03:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1164.eqiad.wmnet with reason: host reimage * 03:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1163.eqiad.wmnet with reason: host reimage * 03:05 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1165.eqiad.wmnet with reason: host reimage * 03:04 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1164.eqiad.wmnet with reason: host reimage * 03:04 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1197.eqiad.wmnet with reason: host reimage * 03:04 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1196.eqiad.wmnet with reason: host reimage * 03:04 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1163.eqiad.wmnet with reason: host reimage * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.16 refs [[phab:T430835|T430835]] * 02:50 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1197.eqiad.wmnet with OS bookworm * 02:49 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1196.eqiad.wmnet with OS bookworm * 02:49 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1165.eqiad.wmnet with OS bookworm * 02:49 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1164.eqiad.wmnet with OS bookworm * 02:48 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1163.eqiad.wmnet with OS bookworm * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 46s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:15 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-codfw: Set storage compatability to NONE — [[phab:T433026|T433026]] - eevans@cumin1003 * 00:38 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-codfw: Set storage compatability to NONE — [[phab:T433026|T433026]] - eevans@cumin1003 * 00:11 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326373{{!}}PersonalDashboard: add newly renamed *ReviewChangesMlModel setting (T422148)]] (duration: 07m 07s) * 00:09 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-eqiad: Set storage compatability to NONE — [[phab:T433026|T433026]] - eevans@cumin1003 * 00:07 musikanimal@deploy1003: musikanimal: Continuing with deployment * 00:06 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1326373{{!}}PersonalDashboard: add newly renamed *ReviewChangesMlModel setting (T422148)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:04 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1326373{{!}}PersonalDashboard: add newly renamed *ReviewChangesMlModel setting (T422148)]] == 2026-08-17 == * 23:30 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-eqiad: Set storage compatability to NONE — [[phab:T433026|T433026]] - eevans@cumin1003 * 23:05 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-codfw: Set storage compatability to UPGRADING — [[phab:T433026|T433026]] - eevans@cumin1003 * 22:28 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-codfw: Set storage compatability to UPGRADING — [[phab:T433026|T433026]] - eevans@cumin1003 * 21:46 logmsgbot: jforrester Deployed security patch for [[phab:T435085|T435085]] * 21:39 swfrench@deploy1003: mwscript-k8s job started: purgeList.php # [[phab:T432412|T432412]] * 21:37 maryum: Undeploy security fix for [[phab:T433020|T433020]] * 21:23 maryum: Deployed security fix for [[phab:T433020|T433020]] * 21:21 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 21:21 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 21:14 maryum: Deployed security fix for [[phab:T434967|T434967]] * 20:58 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-eqiad: Set storage compatability to UPGRADING — [[phab:T433026|T433026]] - eevans@cumin1003 * 20:40 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324320{{!}}[itwiki/slwiki/tgwiki] Remove temporary Wikipedia 25 logos permanently (already reverted) (T414265 T414320 T415307)]] (duration: 06m 55s) * 20:40 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 20:39 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 20:36 cjming@deploy1003: cjming, superpes: Continuing with deployment * 20:35 cjming@deploy1003: cjming, superpes: Backport for [[gerrit:1324320{{!}}[itwiki/slwiki/tgwiki] Remove temporary Wikipedia 25 logos permanently (already reverted) (T414265 T414320 T415307)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:33 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1324320{{!}}[itwiki/slwiki/tgwiki] Remove temporary Wikipedia 25 logos permanently (already reverted) (T414265 T414320 T415307)]] * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ttmserver-test: apply * 20:30 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326361{{!}}Remove escaped paths in app site association file (T432412)]] (duration: 13m 28s) * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ttmserver-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-toolhub-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-toolhub-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-toolhub-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-toolhub-test: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-toolhub: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-toolhub: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-toolhub: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-toolhub: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-test: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-test: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 20:27 inflatador: bking@deploy1003 `charlie --services_dir dse-k8s-services -s opensearch-* -e dse-k8s-* apply` [[phab:T435125|T435125]] * 20:27 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-apifeatureusage-test: apply * 20:27 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-apifeatureusage-test: apply * 20:27 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-apifeatureusage-test: apply * 20:27 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-apifeatureusage-test: apply * 20:27 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-apifeatureusage: apply * 20:26 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-apifeatureusage: apply * 20:26 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-apifeatureusage: apply * 20:26 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-apifeatureusage: apply * 20:26 cjming@deploy1003: cjming, tsev: Continuing with deployment * 20:24 inflatador: bking@deploy1003 `charlie --services_dir dse-k8s-services -s opensearch-* -e dse-k8s-*` [[phab:T435125|T435125]] * 20:21 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-eqiad: Set storage compatability to UPGRADING — [[phab:T433026|T433026]] - eevans@cumin1003 * 20:19 cjming@deploy1003: cjming, tsev: Backport for [[gerrit:1326361{{!}}Remove escaped paths in app site association file (T432412)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:17 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1326361{{!}}Remove escaped paths in app site association file (T432412)]] * 20:15 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326077{{!}}InstrumentConstructiveEdits: anchor all runs to the nearest `interval` (T431493)]] (duration: 06m 25s) * 20:11 cjming@deploy1003: cjming: Continuing with deployment * 20:10 cjming@deploy1003: cjming: Backport for [[gerrit:1326077{{!}}InstrumentConstructiveEdits: anchor all runs to the nearest `interval` (T431493)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:08 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1326077{{!}}InstrumentConstructiveEdits: anchor all runs to the nearest `interval` (T431493)]] * 20:06 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1160.eqiad.wmnet with OS bookworm * 20:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1162.eqiad.wmnet with OS bookworm * 19:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1184.eqiad.wmnet with OS bookworm * 19:55 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1161.eqiad.wmnet with OS bookworm * 19:49 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1195.eqiad.wmnet with OS bookworm * 19:49 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-test: apply * 19:49 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-test: apply * 19:44 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1160.eqiad.wmnet with reason: host reimage * 19:41 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-codfw: Upgrade to Java 17 — [[phab:T433026|T433026]] - eevans@cumin1003 * 19:39 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1162.eqiad.wmnet with reason: host reimage * 19:36 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-eqsin and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 19:36 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-test: apply * 19:36 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1184.eqiad.wmnet with reason: host reimage * 19:32 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1161.eqiad.wmnet with reason: host reimage * 19:29 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1195.eqiad.wmnet with reason: host reimage * 19:26 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1161.eqiad.wmnet with reason: host reimage * 19:26 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1160.eqiad.wmnet with reason: host reimage * 19:26 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1184.eqiad.wmnet with reason: host reimage * 19:26 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1162.eqiad.wmnet with reason: host reimage * 19:25 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1195.eqiad.wmnet with reason: host reimage * 19:19 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-test: apply * 19:11 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1195.eqiad.wmnet with OS bookworm * 19:10 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1194.eqiad.wmnet with OS bookworm * 19:10 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1184.eqiad.wmnet with OS bookworm * 19:10 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1162.eqiad.wmnet with OS bookworm * 19:10 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1161.eqiad.wmnet with OS bookworm * 19:10 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1160.eqiad.wmnet with OS bookworm * 19:10 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 19:09 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 19:03 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-codfw: Upgrade to Java 17 — [[phab:T433026|T433026]] - eevans@cumin1003 * 18:58 dancy@deploy1003: Finished scap sync-world: testing [[phab:T375514|T375514]] (duration: 03m 13s) * 18:55 dancy@deploy1003: Started scap sync-world: testing [[phab:T375514|T375514]] * 18:55 dwisehaupt@dns1006: END - running authdns-update * 18:54 dancy@deploy1003: Installation of scap version "4.281.1" completed for 3 hosts * 18:54 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2009.codfw.wmnet * 18:53 dwisehaupt@dns1006: START - running authdns-update * 18:52 dancy@deploy1003: Installing scap version "4.281.1" for 3 host(s) * 18:47 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2009.codfw.wmnet * 18:41 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2008.codfw.wmnet * 18:34 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2008.codfw.wmnet * 18:30 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2007.codfw.wmnet * 18:23 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2007.codfw.wmnet * 18:16 dwisehaupt@dns1005: END - running authdns-update * 18:14 dwisehaupt@dns1005: START - running authdns-update * 18:04 swfrench@deploy1003: Finished scap sync-world: Deploy "Point Test Wiki to new docroot" - [[phab:T432412|T432412]] (duration: 26m 03s) * 18:00 swfrench@deploy1003: swfrench: Continuing with deployment * 17:51 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1057.eqiad.wmnet with OS trixie * 17:47 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1180.eqiad.wmnet with OS bookworm * 17:39 swfrench@deploy1003: swfrench: Deploy "Point Test Wiki to new docroot" - [[phab:T432412|T432412]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:39 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-eqsin and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 17:38 swfrench@deploy1003: Started scap sync-world: Deploy "Point Test Wiki to new docroot" - [[phab:T432412|T432412]] * 17:36 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1159.eqiad.wmnet with OS bookworm * 17:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1158.eqiad.wmnet with OS bookworm * 17:30 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-eqiad: Upgrade to Java 17 — [[phab:T433026|T433026]] - eevans@cumin1003 * 17:30 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-ulsfo or A:cp-drmrs and A:cp - 9.2.15 upgrade ([[phab:T434620|T434620]]) * 17:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1157.eqiad.wmnet with OS bookworm * 17:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1180.eqiad.wmnet with reason: host reimage * 17:22 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 17:21 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1193.eqiad.wmnet with OS bookworm * 17:21 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 17:21 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 17:20 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 17:20 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 17:20 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1183.eqiad.wmnet with OS bookworm * 17:18 swfrench@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 17:18 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 17:17 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1180.eqiad.wmnet with reason: host reimage * 17:17 swfrench@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 17:17 swfrench@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 17:16 swfrench@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 17:16 swfrench@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 17:15 swfrench@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 17:15 swfrench@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 17:15 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1192.eqiad.wmnet with OS bookworm * 17:14 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1159.eqiad.wmnet with reason: host reimage * 17:14 swfrench@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 17:13 swfrench@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 17:12 swfrench@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 17:09 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1158.eqiad.wmnet with reason: host reimage * 17:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1157.eqiad.wmnet with reason: host reimage * 17:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1180 * 17:02 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1180 * 17:01 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica-eqsin and A:liberica * 17:01 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1193.eqiad.wmnet with reason: host reimage * 16:57 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1183.eqiad.wmnet with reason: host reimage * 16:56 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1054.eqiad.wmnet with OS trixie * 16:55 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1056.eqiad.wmnet with OS trixie * 16:54 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1193.eqiad.wmnet with reason: host reimage * 16:54 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1192.eqiad.wmnet with reason: host reimage * 16:51 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-eqiad: Upgrade to Java 17 — [[phab:T433026|T433026]] - eevans@cumin1003 * 16:50 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1180 * 16:50 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1180.eqiad.wmnet 17.36.64.10.in-addr.arpa 7.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:50 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1180.eqiad.wmnet 17.36.64.10.in-addr.arpa 7.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:50 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:50 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1180 - btullis@cumin1003" * 16:50 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1180 - btullis@cumin1003" * 16:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1158.eqiad.wmnet with reason: host reimage * 16:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1157.eqiad.wmnet with reason: host reimage * 16:49 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica-eqsin and A:liberica * 16:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1183.eqiad.wmnet with reason: host reimage * 16:48 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1159.eqiad.wmnet with reason: host reimage * 16:47 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1192.eqiad.wmnet with reason: host reimage * 16:46 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns4004.wikimedia.org * 16:46 sukhe@dns1004: END - running authdns-update * 16:44 sukhe@dns1004: START - running authdns-update * 16:44 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns4004.wikimedia.org,service=authdns-update * 16:43 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns4004.wikimedia.org with OS trixie * 16:39 btullis@cumin1003: START - Cookbook sre.dns.netbox * 16:39 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1180 * 16:39 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1193.eqiad.wmnet with OS bookworm * 16:39 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1180.eqiad.wmnet with OS bookworm * 16:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1192.eqiad.wmnet with OS bookworm * 16:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1183.eqiad.wmnet with OS bookworm * 16:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1159.eqiad.wmnet with OS bookworm * 16:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1158.eqiad.wmnet with OS bookworm * 16:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1157.eqiad.wmnet with OS bookworm * 16:31 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1057.eqiad.wmnet with OS trixie * 16:30 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1057.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 16:29 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica-drmrs and A:liberica * 16:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1154.eqiad.wmnet with OS bookworm * 16:23 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1057.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 16:23 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1055.eqiad.wmnet with OS trixie * 16:22 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1057 * 16:22 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1057 * 16:21 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:21 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1057] - vriley@cumin1003" * 16:21 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1057] - vriley@cumin1003" * 16:19 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica-drmrs and A:liberica * 16:17 vriley@cumin1003: START - Cookbook sre.dns.netbox * 16:16 phuedx@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics-external: apply * 16:15 phuedx@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics-external: apply * 16:13 phuedx@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics-external: apply * 16:12 phuedx@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics-external: apply * 16:11 phuedx@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics-external: apply * 16:09 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics-external: apply * 16:09 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1176.eqiad.wmnet with OS bookworm * 16:07 btullis@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 16:06 btullis@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 16:05 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1191.eqiad.wmnet with OS bookworm * 16:03 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1154.eqiad.wmnet with reason: host reimage * 16:03 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching aqs[2002-2012].codfw.wmnet,aqs[1017-1027].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433026|T433026]] - eevans@cumin1003 * 15:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1190.eqiad.wmnet with OS bookworm * 15:59 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1154.eqiad.wmnet with reason: host reimage * 15:55 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-debug: apply * 15:55 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-debug: apply * 15:55 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-debug: apply * 15:55 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/mw-debug: apply * 15:53 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns4004.wikimedia.org with reason: host reimage * 15:50 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns4004.wikimedia.org with reason: host reimage * 15:46 btullis@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 15:46 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1176.eqiad.wmnet with reason: host reimage * 15:45 btullis@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 15:43 btullis@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 15:42 btullis@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 15:42 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1191.eqiad.wmnet with reason: host reimage * 15:39 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1190.eqiad.wmnet with reason: host reimage * 15:36 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1054.eqiad.wmnet with OS trixie * 15:35 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:35 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1056.eqiad.wmnet with OS trixie * 15:35 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:34 moritzm: failover Ganeti master in eqiad to ganeti1046 * 15:34 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1176.eqiad.wmnet with reason: host reimage * 15:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1191.eqiad.wmnet with reason: host reimage * 15:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1190.eqiad.wmnet with reason: host reimage * 15:31 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns2006.wikimedia.org * 15:31 sukhe@dns1004: END - running authdns-update * 15:29 sukhe@dns1004: START - running authdns-update * 15:29 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns2006.wikimedia.org,service=authdns-update * 15:29 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns2006.wikimedia.org * 15:29 sukhe@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns2006.wikimedia.org * 15:26 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica-ulsfo and A:liberica * 15:26 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns1006.wikimedia.org * 15:26 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:25 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns2006.wikimedia.org with OS trixie * 15:25 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1056 * 15:25 sukhe@dns1004: END - running authdns-update * 15:23 sukhe@dns1004: START - running authdns-update * 15:23 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns1006.wikimedia.org,service=authdns-update * 15:23 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns1006.wikimedia.org * 15:23 sukhe@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns1006.wikimedia.org * 15:20 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1056 * 15:20 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:20 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1056~] - vriley@cumin1003" * 15:19 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1056~] - vriley@cumin1003" * 15:19 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns1006.wikimedia.org with OS trixie * 15:19 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns4004.wikimedia.org with OS trixie * 15:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1190.eqiad.wmnet with OS bookworm * 15:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1191.eqiad.wmnet with OS bookworm * 15:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1176.eqiad.wmnet with OS bookworm * 15:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1154.eqiad.wmnet with OS bookworm * 15:16 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica-ulsfo and A:liberica * 15:13 vriley@cumin1003: START - Cookbook sre.dns.netbox * 15:13 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host dns4004.wikimedia.org with OS trixie * 15:12 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1188.eqiad.wmnet with OS bookworm * 15:09 taavi@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318203{{!}}Undeploy WP25EasterEggs (II) (T418134)]] (duration: 08m 56s) * 15:06 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica-magru and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 15:06 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1051.eqiad.wmnet with OS trixie * 15:06 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 15:05 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 15:05 taavi@deploy1003: taavi: Continuing with deployment * 15:04 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-ncredir (exit_code=0) rolling reboot on A:ncredir and A:ncredir * 15:04 taavi@deploy1003: taavi: Backport for [[gerrit:1318203{{!}}Undeploy WP25EasterEggs (II) (T418134)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:03 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1055.eqiad.wmnet with OS trixie * 15:02 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:00 taavi@deploy1003: Started scap sync-world: Backport for [[gerrit:1318203{{!}}Undeploy WP25EasterEggs (II) (T418134)]] * 14:59 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1047.eqiad.wmnet * 14:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1047.eqiad.wmnet * 14:58 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy (exit_code=0) rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 14:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-codfw * 14:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp2001.codfw.wmnet * 14:57 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp2001.codfw.wmnet * 14:57 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica-magru and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 14:57 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:56 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1055 * 14:56 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1055 * 14:55 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns2006.wikimedia.org with reason: host reimage * 14:55 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:55 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1055] - vriley@cumin1003" * 14:55 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1055] - vriley@cumin1003" * 14:54 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1047.eqiad.wmnet * 14:52 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1188.eqiad.wmnet with reason: host reimage * 14:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp2001.codfw.wmnet * 14:51 vriley@cumin1003: START - Cookbook sre.dns.netbox * 14:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp2001.codfw.wmnet * 14:50 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2016.codfw.wmnet * 14:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2016.codfw.wmnet * 14:50 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1189.eqiad.wmnet with OS bookworm * 14:48 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1188.eqiad.wmnet with reason: host reimage * 14:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1051.eqiad.wmnet with reason: host reimage * 14:46 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1182.eqiad.wmnet with OS bookworm * 14:45 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief2002.codfw.wmnet * 14:44 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns1006.wikimedia.org with reason: host reimage * 14:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2016.codfw.wmnet * 14:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1143.eqiad.wmnet with OS bookworm * 14:43 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2016.codfw.wmnet * 14:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2318-2331].codfw.wmnet * 14:43 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2318-2331].codfw.wmnet * 14:42 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching aqs[2002-2012].codfw.wmnet,aqs[1017-1027].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433026|T433026]] - eevans@cumin1003 * 14:41 cgoubert@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326306{{!}}Add placeholder $wmgRedisLockPassword (T366938 T427999)]] (duration: 06m 56s) * 14:41 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief2002.codfw.wmnet * 14:40 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief1002.eqiad.wmnet * 14:39 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1051.eqiad.wmnet with reason: host reimage * 14:38 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns2006.wikimedia.org with reason: host reimage * 14:37 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns1006.wikimedia.org with reason: host reimage * 14:37 cgoubert@deploy1003: cgoubert: Continuing with deployment * 14:36 cgoubert@deploy1003: cgoubert: Backport for [[gerrit:1326306{{!}}Add placeholder $wmgRedisLockPassword (T366938 T427999)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:36 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief1002.eqiad.wmnet * 14:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2318-2331].codfw.wmnet * 14:34 cgoubert@deploy1003: Started scap sync-world: Backport for [[gerrit:1326306{{!}}Add placeholder $wmgRedisLockPassword (T366938 T427999)]] * 14:29 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2318-2331].codfw.wmnet * 14:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2304-2317].codfw.wmnet * 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2304-2317].codfw.wmnet * 14:28 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1054 * 14:27 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1054 * 14:27 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:27 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1054] - vriley@cumin1003" * 14:27 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1054] - vriley@cumin1003" * 14:26 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1189.eqiad.wmnet with reason: host reimage * 14:26 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test2001.codfw.wmnet * 14:25 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test1001.eqiad.wmnet * 14:25 claime: Deploying wmgRedisLockPassword - [[phab:T366938|T366938]] [[phab:T427999|T427999]] * 14:24 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1051.eqiad.wmnet with OS trixie * 14:23 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 14:22 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1182.eqiad.wmnet with reason: host reimage * 14:22 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1051.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:22 vriley@cumin1003: START - Cookbook sre.dns.netbox * 14:22 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test2001.codfw.wmnet * 14:21 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test1001.eqiad.wmnet * 14:21 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 14:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2304-2317].codfw.wmnet * 14:19 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns4004.wikimedia.org with OS trixie * 14:19 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns2006.wikimedia.org with OS trixie * 14:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1143.eqiad.wmnet with reason: host reimage * 14:19 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns1006.wikimedia.org with OS trixie * 14:17 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1189.eqiad.wmnet with reason: host reimage * 14:15 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1182.eqiad.wmnet with reason: host reimage * 14:14 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1143.eqiad.wmnet with reason: host reimage * 14:13 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1051.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:12 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2304-2317].codfw.wmnet * 14:12 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2290-2303].codfw.wmnet * 14:12 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1049.eqiad.wmnet with OS trixie * 14:12 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2290-2303].codfw.wmnet * 14:12 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1051 * 14:11 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1051 * 14:11 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1170.eqiad.wmnet onto db1284.eqiad.wmnet * 14:11 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1170: Pool db1170.eqiad.wmnet in after cloning * 14:09 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 14:09 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:09 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1051] - vriley@cumin1003" * 14:09 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1051] - vriley@cumin1003" * 14:06 klausman@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:04 vriley@cumin1003: START - Cookbook sre.dns.netbox * 14:04 klausman@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2290-2303].codfw.wmnet * 14:02 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1189.eqiad.wmnet with OS bookworm * 14:02 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1188.eqiad.wmnet with OS bookworm * 14:01 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 14:00 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1182.eqiad.wmnet with OS bookworm * 14:00 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1143.eqiad.wmnet with OS bookworm * 13:59 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 13:57 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply * 13:57 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply * 13:56 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply * 13:56 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-ulsfo or A:cp-drmrs and A:cp - 9.2.15 upgrade ([[phab:T434620|T434620]]) * 13:56 cjd91: sudo -i cookbook sre.cdn.roll-upgrade-ats --query 'A:cp-ulsfo or A:cp-drmrs' --task-id [[phab:T434620|T434620]] --reason '9.2.15 upgrade' * 13:56 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply * 13:55 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply * 13:55 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply * 13:54 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 13:54 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 13:54 phuedx@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:53 phuedx@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics-external: apply * 13:52 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1049.eqiad.wmnet with reason: host reimage * 13:51 phuedx@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2290-2303].codfw.wmnet * 13:50 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1047.eqiad.wmnet * 13:50 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2276-2289].codfw.wmnet * 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2276-2289].codfw.wmnet * 13:49 phuedx@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics-external: apply * 13:49 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1046.eqiad.wmnet * 13:49 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1049.eqiad.wmnet with reason: host reimage * 13:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1046.eqiad.wmnet * 13:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-ncredir rolling reboot on A:ncredir and A:ncredir * 13:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 13:46 phuedx@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:44 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics-external: apply * 13:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1046.eqiad.wmnet * 13:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2276-2289].codfw.wmnet * 13:36 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1046.eqiad.wmnet * 13:36 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1045.eqiad.wmnet * 13:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1045.eqiad.wmnet * 13:34 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2276-2289].codfw.wmnet * 13:34 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1049.eqiad.wmnet with OS trixie * 13:34 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2262-2275].codfw.wmnet * 13:34 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2262-2275].codfw.wmnet * 13:32 Lucas_WMDE: UTC afternoon backport+config window done * 13:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1045.eqiad.wmnet * 13:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2262-2275].codfw.wmnet * 13:26 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1049.eqiad.wmnet with OS trixie * 13:26 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1049.eqiad.wmnet with OS trixie * 13:26 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1170: Pool db1170.eqiad.wmnet in after cloning * 13:23 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1049.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 13:23 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1045.eqiad.wmnet * 13:21 atsukoito: manually done sudo -i docker-registryctl --debug delete-tags 'docker-registry.discovery.wmnet/repos/data-engineering/airflow-dags:airflow-3.3.0-py3.11-2026-08-17-*' to remove incorrect tags * 13:19 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311141{{!}}viwiki: Set `noindex,nofollow` for User and User talk (T432311)]] (duration: 11m 11s) * 13:15 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2262-2275].codfw.wmnet * 13:14 lucaswerkmeister-wmde@deploy1003: ndkdd, lucaswerkmeister-wmde: Continuing with deployment * 13:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2248-2261].codfw.wmnet * 13:14 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2248-2261].codfw.wmnet * 13:13 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1049.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 13:12 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1049 * 13:11 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1049 * 13:10 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:10 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1049] - vriley@cumin1003" * 13:10 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1049] - vriley@cumin1003" * 13:10 lucaswerkmeister-wmde@deploy1003: ndkdd, lucaswerkmeister-wmde: Backport for [[gerrit:1311141{{!}}viwiki: Set `noindex,nofollow` for User and User talk (T432311)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1037.eqiad.wmnet * 13:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1037.eqiad.wmnet * 13:08 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1311141{{!}}viwiki: Set `noindex,nofollow` for User and User talk (T432311)]] * 13:06 vriley@cumin1003: START - Cookbook sre.dns.netbox * 13:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2248-2261].codfw.wmnet * 13:00 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1037.eqiad.wmnet * 12:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2248-2261].codfw.wmnet * 12:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2204-2215,2242-2243].codfw.wmnet * 12:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2204-2215,2242-2243].codfw.wmnet * 12:49 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324832{{!}}Migrate $wgFlaggedRevsTags from flaggedrevs.php to ext-FlaggedRevs.php]] (duration: 14m 02s) * 12:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint2001.codfw.wmnet * 12:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2204-2215,2242-2243].codfw.wmnet * 12:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint2001.codfw.wmnet * 12:41 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1037.eqiad.wmnet * 12:40 ladsgroup@deploy1003: Rolling back deployment * 12:37 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1324832{{!}}Migrate $wgFlaggedRevsTags from flaggedrevs.php to ext-FlaggedRevs.php]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:37 seanleong-wmde: Finished populateSitesTable for [bolwiki] ([[[phab:T429955|T429955]]]) * 12:35 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1324832{{!}}Migrate $wgFlaggedRevsTags from flaggedrevs.php to ext-FlaggedRevs.php]] * 12:35 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2204-2215,2242-2243].codfw.wmnet * 12:34 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2190-2203].codfw.wmnet * 12:34 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2190-2203].codfw.wmnet * 12:32 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1028.eqiad.wmnet * 12:32 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1028.eqiad.wmnet * 12:31 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint1001.eqiad.wmnet * 12:30 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1170: Depool db1170.eqiad.wmnet to then clone it to db1284.eqiad.wmnet - marostegui@cumin1003 * 12:28 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint1001.eqiad.wmnet * 12:28 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1170: Depool db1170.eqiad.wmnet to then clone it to db1284.eqiad.wmnet - marostegui@cumin1003 * 12:27 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1170.eqiad.wmnet onto db1284.eqiad.wmnet * 12:27 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db2901.codfw.wmnet * 12:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2190-2203].codfw.wmnet * 12:27 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 12:26 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1028.eqiad.wmnet * 12:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2190-2203].codfw.wmnet * 12:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2172-2179,2184-2189].codfw.wmnet * 12:18 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2172-2179,2184-2189].codfw.wmnet * 12:13 seanleong-wmde@deploy1003: mwscript-k8s job started: foreachwikiindblist wikidataclient extensions/Wikibase/lib/maintenance/populateSitesTable.php --force-protocol https # [[phab:T429955|T429955]] * 12:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2172-2179,2184-2189].codfw.wmnet * 12:09 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1028.eqiad.wmnet * 12:03 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 12:02 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db2901.codfw.wmnet * 12:02 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db2901.codfw.wmnet * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1027.eqiad.wmnet * 12:02 fceratto@cumin1003: END (ERROR) - Cookbook sre.dns.netbox (exit_code=97) * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1027.eqiad.wmnet * 12:02 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 12:02 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db2901.codfw.wmnet * 12:01 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2172-2179,2184-2189].codfw.wmnet * 12:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2158-2171].codfw.wmnet * 12:01 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2158-2171].codfw.wmnet * 11:58 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db1902.eqiad.wmnet * 11:58 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 11:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1027.eqiad.wmnet * 11:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2158-2171].codfw.wmnet * 11:51 jayme: updated calico to v3.30.7 on wikikube eqiad - [[phab:T427400|T427400]] * 11:50 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1027.eqiad.wmnet * 11:45 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2158-2171].codfw.wmnet * 11:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2144-2157].codfw.wmnet * 11:44 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2144-2157].codfw.wmnet * 11:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1058.eqiad.wmnet * 11:43 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1058.eqiad.wmnet * 11:43 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 11:43 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 11:43 marostegui@cumin1003: Removing db1153 from zarcillo [[phab:T434638|T434638]] * 11:42 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1153.eqiad.wmnet * 11:42 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:42 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1153.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 11:42 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1172.eqiad.wmnet onto db1286.eqiad.wmnet * 11:42 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1172: Pool db1172.eqiad.wmnet in after cloning * 11:42 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1153.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 11:41 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1902.eqiad.wmnet * 11:41 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=97) for new host db2901.codfw.wmnet * 11:41 fceratto@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host db2901.codfw.wmnet with OS trixie * 11:38 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'. * 11:38 marostegui@cumin1003: START - Cookbook sre.dns.netbox * 11:37 marostegui@dns1004: END - running authdns-update * 11:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1058.eqiad.wmnet * 11:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2144-2157].codfw.wmnet * 11:35 marostegui@dns1004: START - running authdns-update * 11:32 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1153.eqiad.wmnet * 11:32 marostegui@cumin1003: START - Cookbook sre.mysql.decommission * 11:28 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326249{{!}}ImagePage: move TOC element below file link (T332644)]] (duration: 09m 56s) * 11:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2144-2157].codfw.wmnet * 11:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2130-2143].codfw.wmnet * 11:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2130-2143].codfw.wmnet * 11:26 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1058.eqiad.wmnet * 11:23 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 11:22 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'. * 11:22 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'. * 11:22 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'. * 11:22 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326249{{!}}ImagePage: move TOC element below file link (T332644)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1057.eqiad.wmnet * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1057.eqiad.wmnet * 11:21 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 11:20 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 11:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2130-2143].codfw.wmnet * 11:19 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326249{{!}}ImagePage: move TOC element below file link (T332644)]] * 11:18 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply * 11:17 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply * 11:17 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply * 11:16 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 11:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1057.eqiad.wmnet * 11:15 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 11:14 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 11:13 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 11:13 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db2901.codfw.wmnet with OS trixie * 11:12 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db2901.codfw.wmnet - fceratto@cumin1003" * 11:12 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db2901.codfw.wmnet - fceratto@cumin1003" * 11:12 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db2901.codfw.wmnet on all recursors * 11:12 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db2901.codfw.wmnet on all recursors * 11:12 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:12 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db2901.codfw.wmnet - fceratto@cumin1003" * 11:12 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2130-2143].codfw.wmnet * 11:11 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2107-2115,2124-2129].codfw.wmnet * 11:11 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2107-2115,2124-2129].codfw.wmnet * 11:11 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 11:11 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 11:11 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply * 11:11 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 11:10 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 11:06 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db2901.codfw.wmnet - fceratto@cumin1003" * 11:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2107-2115,2124-2129].codfw.wmnet * 10:57 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1172: Pool db1172.eqiad.wmnet in after cloning * 10:54 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2107-2115,2124-2129].codfw.wmnet * 10:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2078,2087-2095,2102-2106].codfw.wmnet * 10:54 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2078,2087-2095,2102-2106].codfw.wmnet * 10:47 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply * 10:46 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply * 10:46 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply * 10:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2078,2087-2095,2102-2106].codfw.wmnet * 10:45 blake@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply * 10:38 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1057.eqiad.wmnet * 10:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1056.eqiad.wmnet * 10:38 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1056.eqiad.wmnet * 10:37 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2078,2087-2095,2102-2106].codfw.wmnet * 10:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2061-2062,2064-2065,2067-2077].codfw.wmnet * 10:36 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2061-2062,2064-2065,2067-2077].codfw.wmnet * 10:32 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1056.eqiad.wmnet * 10:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2061-2062,2064-2065,2067-2077].codfw.wmnet * 10:25 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:25 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db2901.codfw.wmnet * 10:25 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db1902.eqiad.wmnet * 10:25 fceratto@cumin1003: END (ERROR) - Cookbook sre.dns.netbox (exit_code=97) * 10:24 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:24 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1902.eqiad.wmnet * 10:24 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=97) for new host db1902.eqiad.wmnet * 10:24 fceratto@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host db1902.eqiad.wmnet with OS trixie * 10:24 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=93) for new host db1903.eqiad.wmnet * 10:24 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 10:20 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1056.eqiad.wmnet * 10:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2061-2062,2064-2065,2067-2077].codfw.wmnet * 10:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2038-2039,2041-2042,2044,2046,2049-2051,2055-2060].codfw.wmnet * 10:18 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2038-2039,2041-2042,2044,2046,2049-2051,2055-2060].codfw.wmnet * 10:15 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1055.eqiad.wmnet * 10:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1055.eqiad.wmnet * 10:13 Amir1: mwscript-k8s --dblist=all -- purgeUserOptions.php --login-age 5 uls-preferences ([[phab:T406724|T406724]]) * 10:11 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:10 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1903.eqiad.wmnet on all recursors * 10:10 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1903.eqiad.wmnet on all recursors * 10:10 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2038-2039,2041-2042,2044,2046,2049-2051,2055-2060].codfw.wmnet * 10:10 moritzm: installing unzip security updates * 10:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1055.eqiad.wmnet * 10:08 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=97) for new host db1901.eqiad.wmnet * 10:08 fceratto@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host db1901.eqiad.wmnet with OS trixie * 10:08 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:08 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 10:08 fceratto@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1903.eqiad.wmnet - fceratto@cumin1003" * 10:07 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db1902.eqiad.wmnet with OS trixie * 10:07 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1902.eqiad.wmnet - fceratto@cumin1003" * 10:07 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1902.eqiad.wmnet - fceratto@cumin1003" * 10:04 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply * 10:04 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326227{{!}}Enable desktop/native lazy loading everywhere (T148047)]] (duration: 07m 13s) * 10:03 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1902.eqiad.wmnet on all recursors * 10:03 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1902.eqiad.wmnet on all recursors * 10:03 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:03 blake@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply * 10:01 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1903.eqiad.wmnet - fceratto@cumin1003" * 10:01 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2038-2039,2041-2042,2044,2046,2049-2051,2055-2060].codfw.wmnet * 10:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2002,2005-2006,2011-2015,2017-2018,2033-2037].codfw.wmnet * 10:01 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2002,2005-2006,2011-2015,2017-2018,2033-2037].codfw.wmnet * 10:00 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:00 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 09:59 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 09:59 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1055.eqiad.wmnet * 09:58 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326227{{!}}Enable desktop/native lazy loading everywhere (T148047)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:56 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326227{{!}}Enable desktop/native lazy loading everywhere (T148047)]] * 09:53 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2002,2005-2006,2011-2015,2017-2018,2033-2037].codfw.wmnet * 09:50 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:48 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1903.eqiad.wmnet * 09:48 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:48 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1902.eqiad.wmnet * 09:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 09:44 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2002,2005-2006,2011-2015,2017-2018,2033-2037].codfw.wmnet * 09:43 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-codfw * 09:43 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1174.eqiad.wmnet onto db1288.eqiad.wmnet * 09:42 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1174: Pool db1174.eqiad.wmnet in after cloning * 09:40 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1172: Depool db1172.eqiad.wmnet to then clone it to db1286.eqiad.wmnet - marostegui@cumin1003 * 09:39 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1172: Depool db1172.eqiad.wmnet to then clone it to db1286.eqiad.wmnet - marostegui@cumin1003 * 09:39 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1172.eqiad.wmnet onto db1286.eqiad.wmnet * 09:33 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1175.eqiad.wmnet onto db1289.eqiad.wmnet * 09:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1175: Pool db1175.eqiad.wmnet in after cloning * 09:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1201.eqiad.wmnet onto db1287.eqiad.wmnet * 09:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1201: Pool db1201.eqiad.wmnet in after cloning * 09:28 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db1901.eqiad.wmnet with OS trixie * 09:27 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:27 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:27 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1901.eqiad.wmnet on all recursors * 09:27 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1901.eqiad.wmnet on all recursors * 09:26 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:26 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:26 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:15 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:15 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1901.eqiad.wmnet * 09:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw1001.wikimedia.org with OS trixie * 08:57 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1174: Pool db1174.eqiad.wmnet in after cloning * 08:54 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1054.eqiad.wmnet * 08:54 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1054.eqiad.wmnet * 08:48 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1054.eqiad.wmnet * 08:47 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1175: Pool db1175.eqiad.wmnet in after cloning * 08:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 08:46 marostegui@cumin1003: Removing db1152 from zarcillo [[phab:T434480|T434480]] * 08:46 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1152.eqiad.wmnet * 08:46 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:46 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1152.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 08:46 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1201: Pool db1201.eqiad.wmnet in after cloning * 08:46 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1152.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 08:46 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1054.eqiad.wmnet * 08:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1053.eqiad.wmnet * 08:43 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1053.eqiad.wmnet * 08:42 marostegui@cumin1003: START - Cookbook sre.dns.netbox * 08:38 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage * 08:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1053.eqiad.wmnet * 08:36 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1152.eqiad.wmnet * 08:36 marostegui@cumin1003: START - Cookbook sre.mysql.decommission * 08:35 phuedx: UTC morning backport window done * 08:35 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1053.eqiad.wmnet * 08:34 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1035.eqiad.wmnet * 08:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1035.eqiad.wmnet * 08:34 phuedx@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324725{{!}}EventStreamConfig: Mark product_metrics.web_base and .web_base_with_ip as Test Kitchen streams (T429898 T430322)]], [[gerrit:1313923{{!}}EventStreamConfig: Remove unused web_ui_scroll* streams (T415370)]], [[gerrit:1325546{{!}}EventStreamConfig: Remove Watchlist click stream (T434790)]] (duration: 12m 42s) * 08:33 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1281: Pool back * 08:32 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage * 08:31 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on db2209.codfw.wmnet with reason: Maintenance * 08:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2209: Maintenance needed * 08:30 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2209: Maintenance needed * 08:26 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1035.eqiad.wmnet * 08:26 phuedx@deploy1003: bearloga, phuedx: Continuing with deployment * 08:23 phuedx@deploy1003: bearloga, phuedx: Backport for [[gerrit:1324725{{!}}EventStreamConfig: Mark product_metrics.web_base and .web_base_with_ip as Test Kitchen streams (T429898 T430322)]], [[gerrit:1313923{{!}}EventStreamConfig: Remove unused web_ui_scroll* streams (T415370)]], [[gerrit:1325546{{!}}EventStreamConfig: Remove Watchlist click stream (T434790)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug * 08:21 phuedx@deploy1003: Started scap sync-world: Backport for [[gerrit:1324725{{!}}EventStreamConfig: Mark product_metrics.web_base and .web_base_with_ip as Test Kitchen streams (T429898 T430322)]], [[gerrit:1313923{{!}}EventStreamConfig: Remove unused web_ui_scroll* streams (T415370)]], [[gerrit:1325546{{!}}EventStreamConfig: Remove Watchlist click stream (T434790)]] * 08:19 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw1001.wikimedia.org with OS trixie * 08:18 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1035.eqiad.wmnet * 08:17 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1032.eqiad.wmnet * 08:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1032.eqiad.wmnet * 08:16 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1279: Pool back * 08:15 phuedx@deploy1003: Finished scap sync-world: Backport for [[gerrit:1216721{{!}}viwikivoyage: enable relatedarticle and pop-up (T405724)]] (duration: 39m 12s) * 08:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1032.eqiad.wmnet * 08:09 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1032.eqiad.wmnet * 08:08 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1031.eqiad.wmnet * 08:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1031.eqiad.wmnet * 08:03 godog: switch production to use dumps-nfs.w.o - [[phab:T432212|T432212]] * 08:02 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1031.eqiad.wmnet * 08:02 phuedx@deploy1003: nvdtn19, phuedx: Continuing with deployment * 08:00 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1201: Depool db1201.eqiad.wmnet to then clone it to db1287.eqiad.wmnet - marostegui@cumin1003 * 08:00 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1201: Depool db1201.eqiad.wmnet to then clone it to db1287.eqiad.wmnet - marostegui@cumin1003 * 08:00 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1201.eqiad.wmnet onto db1287.eqiad.wmnet * 08:00 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1057.eqiad.wmnet * 08:00 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:00 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1057.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:59 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1057.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:59 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1275: Pool back * 07:55 filippo@cumin1003: START - Cookbook sre.dns.netbox * 07:55 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1031.eqiad.wmnet * 07:52 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1030.eqiad.wmnet * 07:52 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1030.eqiad.wmnet * 07:52 phuedx@deploy1003: nvdtn19, phuedx: Backport for [[gerrit:1216721{{!}}viwikivoyage: enable relatedarticle and pop-up (T405724)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:51 tappof: bump space for prometheus k8s-aux in eqiad * 07:50 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1057.eqiad.wmnet * 07:48 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1281: Pool back * 07:47 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1281 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96108 and previous config saved to /var/cache/conftool/dbconfig/20260817-074749-marostegui.json * 07:46 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1030.eqiad.wmnet * 07:42 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1030.eqiad.wmnet * 07:41 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1174: Depool db1174.eqiad.wmnet to then clone it to db1288.eqiad.wmnet - marostegui@cumin1003 * 07:41 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1174: Depool db1174.eqiad.wmnet to then clone it to db1288.eqiad.wmnet - marostegui@cumin1003 * 07:41 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1174.eqiad.wmnet onto db1288.eqiad.wmnet * 07:40 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1029.eqiad.wmnet * 07:40 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm2001.wikimedia.org * 07:40 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1029.eqiad.wmnet * 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1056.eqiad.wmnet * 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1056.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:38 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1056.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:36 phuedx@deploy1003: Started scap sync-world: Backport for [[gerrit:1216721{{!}}viwikivoyage: enable relatedarticle and pop-up (T405724)]] * 07:36 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm2001.wikimedia.org * 07:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1279 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96104 and previous config saved to /var/cache/conftool/dbconfig/20260817-073542-marostegui.json * 07:34 filippo@cumin1003: START - Cookbook sre.dns.netbox * 07:34 slyngshede@dns1004: END - running authdns-update * 07:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1029.eqiad.wmnet * 07:33 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm-test1001.wikimedia.org * 07:32 slyngshede@dns1004: START - running authdns-update * 07:31 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1029.eqiad.wmnet * 07:31 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1279: Pool back * 07:30 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1279 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96102 and previous config saved to /var/cache/conftool/dbconfig/20260817-073038-marostegui.json * 07:29 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm-test1001.wikimedia.org * 07:29 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm1001.wikimedia.org * 07:28 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1056.eqiad.wmnet * 07:28 moritzm: extend the disk of ldap-rw1001 by 80G [[phab:T331699|T331699]] * 07:28 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1055.eqiad.wmnet * 07:28 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:28 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1055.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:27 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1055.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:26 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1044.eqiad.wmnet * 07:26 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1044.eqiad.wmnet * 07:25 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm1001.wikimedia.org * 07:24 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1175: Depool db1175.eqiad.wmnet to then clone it to db1289.eqiad.wmnet - marostegui@cumin1003 * 07:24 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1175: Depool db1175.eqiad.wmnet to then clone it to db1289.eqiad.wmnet - marostegui@cumin1003 * 07:24 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1175.eqiad.wmnet onto db1289.eqiad.wmnet * 07:22 filippo@cumin1003: START - Cookbook sre.dns.netbox * 07:20 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1044.eqiad.wmnet * 07:16 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1055.eqiad.wmnet * 07:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1054.eqiad.wmnet * 07:15 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:15 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1054.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:15 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1044.eqiad.wmnet * 07:15 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1054.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:13 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1275: Pool back * 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1275 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96098 and previous config saved to /var/cache/conftool/dbconfig/20260817-071225-marostegui.json * 07:10 filippo@cumin1003: START - Cookbook sre.dns.netbox * 07:05 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1043.eqiad.wmnet * 07:05 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1054.eqiad.wmnet * 07:05 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1051.eqiad.wmnet * 07:05 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:05 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1051.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:05 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1043.eqiad.wmnet * 07:04 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1051.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 06:59 filippo@cumin1003: START - Cookbook sre.dns.netbox * 06:59 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin1001.eqiad.wmnet * 06:59 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1043.eqiad.wmnet * 06:59 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin2001.codfw.wmnet * 06:55 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin2001.codfw.wmnet * 06:55 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1051.eqiad.wmnet * 06:54 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1049.eqiad.wmnet * 06:54 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:54 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1049.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 06:54 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1049.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 06:54 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin1001.eqiad.wmnet * 06:53 moritzm: installing apr-util security updates * 06:52 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1043.eqiad.wmnet * 06:49 filippo@cumin1003: START - Cookbook sre.dns.netbox * 06:41 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1049.eqiad.wmnet * 06:13 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit2003.wikimedia.org * 06:13 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet * 06:07 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet * 06:06 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit2003.wikimedia.org * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 47s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-16 == * 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 01m 03s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-15 == * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 41s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-14 == * 15:38 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-staging-master-eqiad * 15:38 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster1005.eqiad.wmnet * 15:38 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster1005.eqiad.wmnet * 15:35 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sretest2009.codfw.wmnet * 15:33 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster1005.eqiad.wmnet * 15:33 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster1005.eqiad.wmnet * 15:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster1004.eqiad.wmnet * 15:32 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster1004.eqiad.wmnet * 15:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host sretest2009.codfw.wmnet * 15:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster1004.eqiad.wmnet * 15:27 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster1004.eqiad.wmnet * 15:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster1003.eqiad.wmnet * 15:27 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster1003.eqiad.wmnet * 15:24 dancy@deploy1003: Finished scap sync-world: testing (duration: 03m 23s) * 15:22 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster1003.eqiad.wmnet * 15:22 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster1003.eqiad.wmnet * 15:22 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-staging-master-eqiad * 15:20 dancy@deploy1003: Started scap sync-world: testing * 15:20 dancy@deploy1003: Installation of scap version "4.280.2" completed for 3 hosts * 15:18 dancy@deploy1003: Installing scap version "4.280.2" for 3 host(s) * 15:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sretest2006.codfw.wmnet * 14:54 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host sretest2006.codfw.wmnet * 14:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sretest2003.codfw.wmnet * 14:39 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host sretest2003.codfw.wmnet * 13:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt-staging2001.codfw.wmnet * 13:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt-staging2001.codfw.wmnet * 13:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-staging-master-codfw * 13:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster2005.codfw.wmnet * 13:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster2005.codfw.wmnet * 13:05 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox-dev2003.codfw.wmnet * 13:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster2005.codfw.wmnet * 13:04 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster2005.codfw.wmnet * 13:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster2004.codfw.wmnet * 13:04 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster2004.codfw.wmnet * 13:01 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netbox-dev2003.codfw.wmnet * 12:59 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster2004.codfw.wmnet * 12:59 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster2004.codfw.wmnet * 12:59 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster2003.codfw.wmnet * 12:59 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster2003.codfw.wmnet * 12:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster2003.codfw.wmnet * 12:54 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster2003.codfw.wmnet * 12:54 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-staging-master-codfw * 12:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-staging-worker-eqiad * 12:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage1006.eqiad.wmnet * 12:52 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage1006.eqiad.wmnet * 12:46 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage1006.eqiad.wmnet * 12:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw1001.wikimedia.org with OS trixie * 12:45 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage1006.eqiad.wmnet * 12:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage1005.eqiad.wmnet * 12:45 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage1005.eqiad.wmnet * 12:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage1005.eqiad.wmnet * 12:36 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/ratelimit: apply * 12:35 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/ratelimit: apply * 12:35 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:35 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:33 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage1005.eqiad.wmnet * 12:33 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage1004.eqiad.wmnet * 12:33 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage1004.eqiad.wmnet * 12:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage1004.eqiad.wmnet * 12:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage1004.eqiad.wmnet * 12:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage1003.eqiad.wmnet * 12:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage1003.eqiad.wmnet * 12:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage * 12:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage1003.eqiad.wmnet * 12:17 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage * 12:14 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage1003.eqiad.wmnet * 12:14 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-staging-worker-eqiad * 12:03 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw1001.wikimedia.org with OS trixie * 12:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cuminunpriv1001.eqiad.wmnet * 11:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cuminunpriv1001.eqiad.wmnet * 11:27 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1004.wikimedia.org * 11:24 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1153 from dbctl [[phab:T434638|T434638]]', diff saved to https://phabricator.wikimedia.org/P96097 and previous config saved to /var/cache/conftool/dbconfig/20260814-112449-marostegui.json * 11:21 aokoth@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1004.wikimedia.org * 11:20 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 11:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-staging-worker-codfw * 11:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2004.codfw.wmnet * 11:17 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2004.codfw.wmnet * 11:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2004.codfw.wmnet * 11:10 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2004.codfw.wmnet * 11:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2003.codfw.wmnet * 11:10 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2003.codfw.wmnet * 11:03 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-worker1181.eqiad.wmnet with OS bookworm * 11:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2003.codfw.wmnet * 11:03 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2003.codfw.wmnet * 11:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2002.codfw.wmnet * 11:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2002.codfw.wmnet * 10:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1235.eqiad.wmnet with OS bookworm * 10:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock2003.codfw.wmnet * 10:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock2003.codfw.wmnet with OS trixie * 10:56 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2002.codfw.wmnet * 10:54 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1153.eqiad.wmnet with OS bookworm * 10:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2002.codfw.wmnet * 10:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2001.codfw.wmnet * 10:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2001.codfw.wmnet * 10:50 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 10:49 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1187.eqiad.wmnet with OS bookworm * 10:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2001.codfw.wmnet * 10:44 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2001.codfw.wmnet * 10:44 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-staging-worker-codfw * 10:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock2003.codfw.wmnet with reason: host reimage * 10:38 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock2003.codfw.wmnet with reason: host reimage * 10:35 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1153.eqiad.wmnet with reason: host reimage * 10:32 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1235.eqiad.wmnet with reason: host reimage * 10:29 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1187.eqiad.wmnet with reason: host reimage * 10:24 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1153.eqiad.wmnet with reason: host reimage * 10:23 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1235.eqiad.wmnet with reason: host reimage * 10:21 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1187.eqiad.wmnet with reason: host reimage * 10:16 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock2003.codfw.wmnet with OS trixie * 10:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install1005.wikimedia.org * 10:12 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 10:12 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 10:12 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 10:12 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 10:12 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:12 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 10:12 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 10:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install1005.wikimedia.org * 10:07 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1235.eqiad.wmnet with OS bookworm * 10:07 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1187.eqiad.wmnet with OS bookworm * 10:07 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1181.eqiad.wmnet with OS bookworm * 10:07 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1153.eqiad.wmnet with OS bookworm * 10:07 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install2005.wikimedia.org * 10:03 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 10:03 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2003.codfw.wmnet * 10:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1232.eqiad.wmnet with OS bookworm * 10:00 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install2005.wikimedia.org * 10:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install3004.wikimedia.org * 09:58 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 09:58 fceratto@cumin1003: Removing db1151 from zarcillo [[phab:T434538|T434538]] * 09:56 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1231.eqiad.wmnet with OS bookworm * 09:56 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 09:53 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 09:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install3004.wikimedia.org * 09:50 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1152.eqiad.wmnet with OS bookworm * 09:49 Dreamy_Jazz: `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260808000000" --end-timestamp="20260812120000" --sleep="5" --batch-size="50"` for [[phab:T434688|T434688]] * 09:48 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host centrallog2002.codfw.wmnet * 09:48 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install4004.wikimedia.org * 09:42 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1232.eqiad.wmnet with reason: host reimage * 09:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install4004.wikimedia.org * 09:41 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host centrallog2002.codfw.wmnet * 09:39 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install5004.wikimedia.org * 09:36 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1232.eqiad.wmnet with reason: host reimage * 09:36 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'. * 09:34 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'. * 09:33 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1231.eqiad.wmnet with reason: host reimage * 09:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install5004.wikimedia.org * 09:32 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host centrallog1002.eqiad.wmnet * 09:30 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install6003.wikimedia.org * 09:30 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1231.eqiad.wmnet with reason: host reimage * 09:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1152.eqiad.wmnet with reason: host reimage * 09:25 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host centrallog1002.eqiad.wmnet * 09:25 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1152.eqiad.wmnet with reason: host reimage * 09:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install6003.wikimedia.org * 09:22 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1232 * 09:22 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1232 * 09:22 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1232 * 09:22 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1232.eqiad.wmnet 25.53.64.10.in-addr.arpa 5.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:22 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host titan1001.eqiad.wmnet * 09:22 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1232.eqiad.wmnet 25.53.64.10.in-addr.arpa 5.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:22 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:22 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1232 - btullis@cumin1003" * 09:22 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1232 - btullis@cumin1003" * 09:17 btullis@cumin1003: START - Cookbook sre.dns.netbox * 09:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install7002.wikimedia.org * 09:17 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1232 * 09:16 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1231 * 09:16 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1231 * 09:14 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1231 * 09:14 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1231.eqiad.wmnet 24.53.64.10.in-addr.arpa 4.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host titan1001.eqiad.wmnet * 09:14 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1231.eqiad.wmnet 24.53.64.10.in-addr.arpa 4.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:14 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1231 - btullis@cumin1003" * 09:14 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1231 - btullis@cumin1003" * 09:10 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install7002.wikimedia.org * 09:09 btullis@cumin1003: START - Cookbook sre.dns.netbox * 09:08 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1231 * 09:08 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1152 * 09:08 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1152 * 09:06 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1152 * 09:06 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1152.eqiad.wmnet 16.53.64.10.in-addr.arpa 6.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:06 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1152.eqiad.wmnet 16.53.64.10.in-addr.arpa 6.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:06 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:06 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1152 - btullis@cumin1003" * 09:06 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1152 - btullis@cumin1003" * 09:03 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1151.eqiad.wmnet * 09:03 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:03 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1151.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 08:55 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host titan2001.codfw.wmnet * 08:55 btullis@cumin1003: START - Cookbook sre.dns.netbox * 08:54 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1151.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 08:54 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1152 * 08:53 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1232.eqiad.wmnet with OS bookworm * 08:53 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1231.eqiad.wmnet with OS bookworm * 08:53 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1152.eqiad.wmnet with OS bookworm * 08:51 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2209: Pool back * 08:51 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1201.eqiad.wmnet * 08:50 btullis@cumin1003: START - Cookbook sre.hosts.remove-downtime for an-worker1201.eqiad.wmnet * 08:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1229.eqiad.wmnet with OS bookworm * 08:47 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host titan2001.codfw.wmnet * 08:46 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping2004.codfw.wmnet * 08:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ping2004.codfw.wmnet * 08:40 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host an-worker1230.eqiad.wmnet with OS bookworm * 08:40 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 08:34 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1151.eqiad.wmnet * 08:34 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 08:31 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host titan1002.eqiad.wmnet * 08:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1229.eqiad.wmnet with reason: host reimage * 08:25 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host titan1002.eqiad.wmnet * 08:25 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1229.eqiad.wmnet with reason: host reimage * 08:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping1004.eqiad.wmnet * 08:21 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ping1004.eqiad.wmnet * 08:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1230.eqiad.wmnet with reason: host reimage * 08:12 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1230.eqiad.wmnet with reason: host reimage * 08:11 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1229 * 08:11 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1229 * 08:11 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1229 * 08:11 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1229.eqiad.wmnet 22.53.64.10.in-addr.arpa 2.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:11 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1229.eqiad.wmnet 22.53.64.10.in-addr.arpa 2.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:11 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:11 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1229 - btullis@cumin1003" * 08:11 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1229 - btullis@cumin1003" * 08:10 btullis@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1201.eqiad.wmnet with reason: Fixing a disk * 08:07 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host titan2002.codfw.wmnet * 08:05 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2209: Pool back * 08:05 btullis@cumin1003: START - Cookbook sre.dns.netbox * 08:00 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host titan2002.codfw.wmnet * 07:59 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1229 * 07:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1230 * 07:58 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1230 * 07:55 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1230 * 07:55 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1230.eqiad.wmnet 23.53.64.10.in-addr.arpa 3.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:55 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1230.eqiad.wmnet 23.53.64.10.in-addr.arpa 3.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:55 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:55 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1230 - btullis@cumin1003" * 07:55 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1230 - btullis@cumin1003" * 07:51 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host kubestagemaster2005.codfw.wmnet with OS trixie * 07:48 btullis@cumin1003: START - Cookbook sre.dns.netbox * 07:41 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1230 * 07:41 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1229.eqiad.wmnet with OS bookworm * 07:41 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1230.eqiad.wmnet with OS bookworm * 07:39 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1152 from dbctl [[phab:T434480|T434480]]', diff saved to https://phabricator.wikimedia.org/P96090 and previous config saved to /var/cache/conftool/dbconfig/20260814-073941-marostegui.json * 07:29 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on kubestagemaster2005.codfw.wmnet with reason: host reimage * 07:23 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on kubestagemaster2005.codfw.wmnet with reason: host reimage * 07:04 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host kubestagemaster2005.codfw.wmnet with OS trixie * 06:53 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1228.eqiad.wmnet with OS bookworm * 06:44 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1227.eqiad.wmnet with OS bookworm * 06:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1209.eqiad.wmnet with OS bookworm * 06:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1175.eqiad.wmnet with OS bookworm * 06:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1228.eqiad.wmnet with reason: host reimage * 06:27 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1228.eqiad.wmnet with reason: host reimage * 06:25 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1227.eqiad.wmnet with reason: host reimage * 06:21 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1227.eqiad.wmnet with reason: host reimage * 06:18 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1209.eqiad.wmnet with reason: host reimage * 06:14 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1209.eqiad.wmnet with reason: host reimage * 06:14 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1228 * 06:14 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1228 * 06:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1175.eqiad.wmnet with reason: host reimage * 06:12 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1228 * 06:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1228.eqiad.wmnet 20.53.64.10.in-addr.arpa 0.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:12 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1228.eqiad.wmnet 20.53.64.10.in-addr.arpa 0.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1228 - ryankemper@cumin2003" * 06:12 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1228 - ryankemper@cumin2003" * 06:09 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1175.eqiad.wmnet with reason: host reimage * 06:07 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 06:07 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1228 * 06:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1227 * 06:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1227 * 06:06 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1227 * 06:06 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1227.eqiad.wmnet 19.53.64.10.in-addr.arpa 9.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:06 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1227.eqiad.wmnet 19.53.64.10.in-addr.arpa 9.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:06 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:06 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1227 - ryankemper@cumin2003" * 06:06 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1227 - ryankemper@cumin2003" * 06:00 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 06:00 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1227 * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1209 * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1209 * 06:00 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1209 * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1209.eqiad.wmnet 15.53.64.10.in-addr.arpa 5.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:00 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1209.eqiad.wmnet 15.53.64.10.in-addr.arpa 5.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1209 - ryankemper@cumin2003" * 06:00 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1209 - ryankemper@cumin2003" * 05:54 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 05:54 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1209 * 05:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1175 * 05:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1175 * 05:52 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1175 * 05:52 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1175.eqiad.wmnet 17.53.64.10.in-addr.arpa 7.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 05:52 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1175.eqiad.wmnet 17.53.64.10.in-addr.arpa 7.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 05:52 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 05:52 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1175 - ryankemper@cumin2003" * 05:52 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1175 - ryankemper@cumin2003" * 05:49 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1228.eqiad.wmnet with OS bookworm * 05:49 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1227.eqiad.wmnet with OS bookworm * 05:48 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1209.eqiad.wmnet with OS bookworm * 05:47 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 05:47 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1175 * 05:47 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1175.eqiad.wmnet with OS bookworm * 05:09 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host apifeatureusage1001.eqiad.wmnet with OS bookworm * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 03s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:10 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 01:07 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 01:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 01:02 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 00:59 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1226.eqiad.wmnet with OS bookworm * 00:47 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1225.eqiad.wmnet with OS bookworm * 00:41 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1224.eqiad.wmnet with OS bookworm * 00:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1226.eqiad.wmnet with reason: host reimage * 00:31 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1226.eqiad.wmnet with reason: host reimage * 00:28 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1225.eqiad.wmnet with reason: host reimage * 00:25 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1225.eqiad.wmnet with reason: host reimage * 00:19 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1224.eqiad.wmnet with reason: host reimage * 00:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1226 * 00:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1226 * 00:17 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1226 * 00:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1226.eqiad.wmnet 23.36.64.10.in-addr.arpa 3.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:17 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1226.eqiad.wmnet 23.36.64.10.in-addr.arpa 3.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 00:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1226 - ryankemper@cumin2003" * 00:17 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1226 - ryankemper@cumin2003" * 00:16 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1224.eqiad.wmnet with reason: host reimage * 00:12 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 00:12 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1226 * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1225 * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1225 * 00:10 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1225 * 00:10 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1225.eqiad.wmnet 22.36.64.10.in-addr.arpa 2.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:10 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1225.eqiad.wmnet 22.36.64.10.in-addr.arpa 2.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:10 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 00:10 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1225 - ryankemper@cumin2003" * 00:10 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1225 - ryankemper@cumin2003" * 00:03 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 00:02 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1225 * 00:02 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1224 * 00:02 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1224 * 00:00 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1224 * 00:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1224.eqiad.wmnet 21.36.64.10.in-addr.arpa 1.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:00 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1224.eqiad.wmnet 21.36.64.10.in-addr.arpa 1.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 00:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1224 - ryankemper@cumin2003" * 00:00 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1224 - ryankemper@cumin2003" == 2026-08-13 == * 23:54 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1226.eqiad.wmnet with OS bookworm * 23:53 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1225.eqiad.wmnet with OS bookworm * 23:52 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 23:51 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1224 * 23:51 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1224.eqiad.wmnet with OS bookworm * 23:47 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1223.eqiad.wmnet with OS bookworm * 23:29 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1223.eqiad.wmnet with reason: host reimage * 23:24 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1223.eqiad.wmnet with reason: host reimage * 23:19 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325525{{!}}ve.ui.CodeMirror.less: ensure normal font style]] (duration: 11m 40s) * 23:16 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260807000000" --end-timestamp="20260808000000" --sleep="5" --batch-size="50"` for [[phab:T434688|T434688]] * 23:13 musikanimal@deploy1003: musikanimal: Continuing with deployment * 23:11 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1325525{{!}}ve.ui.CodeMirror.less: ensure normal font style]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1223 * 23:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1223 * 23:08 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1325525{{!}}ve.ui.CodeMirror.less: ensure normal font style]] * 23:07 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1223 * 23:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1223.eqiad.wmnet 20.36.64.10.in-addr.arpa 0.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:07 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1223.eqiad.wmnet 20.36.64.10.in-addr.arpa 0.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 23:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1223 - ryankemper@cumin2003" * 23:03 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1223 - ryankemper@cumin2003" * 23:01 sbassett: Deployed security updates for [[phab:T430596|T430596]], [[phab:T120386|T120386]] * 22:55 bking@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host kubestagemaster2005.codfw.wmnet * 22:55 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host kubestagemaster2005.codfw.wmnet with OS bookworm * 22:54 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 22:53 ryankemper@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 22:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on kubestagemaster2005.codfw.wmnet with reason: host reimage * 22:49 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1222.eqiad.wmnet with OS bookworm * 22:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on kubestagemaster2005.codfw.wmnet with reason: host reimage * 22:29 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1222.eqiad.wmnet with reason: host reimage * 22:26 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1222.eqiad.wmnet with reason: host reimage * 22:23 bking@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host aux-k8s-etcd2003.codfw.wmnet * 22:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd2003.codfw.wmnet with OS bookworm * 22:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host kubestagemaster2005.codfw.wmnet with OS bookworm * 22:23 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM kubestagemaster2005.codfw.wmnet - bking@cumin2003" * 22:23 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM kubestagemaster2005.codfw.wmnet - bking@cumin2003" * 22:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) kubestagemaster2005.codfw.wmnet on all recursors * 22:22 bking@cumin2003: START - Cookbook sre.dns.wipe-cache kubestagemaster2005.codfw.wmnet on all recursors * 22:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:22 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM kubestagemaster2005.codfw.wmnet - bking@cumin2003" * 22:22 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM kubestagemaster2005.codfw.wmnet - bking@cumin2003" * 22:17 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 22:13 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1223 * 22:12 bking@cumin2003: START - Cookbook sre.dns.netbox * 22:12 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host kubestagemaster2005.codfw.wmnet * 22:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1222 * 22:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1222 * 22:11 sbassett: Deployed security updates for [[phab:T429244|T429244]], [[phab:T434039|T434039]] * 22:11 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1222 * 22:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1222.eqiad.wmnet 19.36.64.10.in-addr.arpa 9.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:11 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1222.eqiad.wmnet 19.36.64.10.in-addr.arpa 9.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1222 - ryankemper@cumin2003" * 22:07 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1222 - ryankemper@cumin2003" * 22:05 robh@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-wdqs2001.codfw.wmnet with reason: updating firmware * 22:01 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 22:01 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1223.eqiad.wmnet with OS bookworm * 22:01 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1222 * 22:01 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1222.eqiad.wmnet with OS bookworm * 22:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1212.eqiad.wmnet with OS bookworm * 21:54 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324295{{!}}Revert "Lazily reject pre-fix parser-cache entries for noreferrer/noopener links" (T429090)]] (duration: 06m 42s) * 21:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd2003.codfw.wmnet with reason: host reimage * 21:53 bking@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host dse-k8s-etcd2001.codfw.wmnet * 21:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dse-k8s-etcd2001.codfw.wmnet with OS bookworm * 21:50 sbassett@deploy1003: sbassett, kharlan: Continuing with deployment * 21:49 sbassett@deploy1003: sbassett, kharlan: Backport for [[gerrit:1324295{{!}}Revert "Lazily reject pre-fix parser-cache entries for noreferrer/noopener links" (T429090)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:48 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aux-k8s-etcd2003.codfw.wmnet with reason: host reimage * 21:47 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1324295{{!}}Revert "Lazily reject pre-fix parser-cache entries for noreferrer/noopener links" (T429090)]] * 21:42 maryum: Deployed security patch for [[phab:T434549|T434549]] * 21:39 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1212.eqiad.wmnet with reason: host reimage * 21:34 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1212.eqiad.wmnet with reason: host reimage * 21:33 bking@cumin2003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd2003.codfw.wmnet with OS bookworm * 21:32 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM aux-k8s-etcd2003.codfw.wmnet - bking@cumin2003" * 21:32 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM aux-k8s-etcd2003.codfw.wmnet - bking@cumin2003" * 21:32 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) aux-k8s-etcd2003.codfw.wmnet on all recursors * 21:32 bking@cumin2003: START - Cookbook sre.dns.wipe-cache aux-k8s-etcd2003.codfw.wmnet on all recursors * 21:32 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:32 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM aux-k8s-etcd2003.codfw.wmnet - bking@cumin2003" * 21:31 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM aux-k8s-etcd2003.codfw.wmnet - bking@cumin2003" * 21:28 maryum: Deployed security patch for [[phab:T434619|T434619]] * 21:26 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:26 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host aux-k8s-etcd2003.codfw.wmnet * 21:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-etcd2001.codfw.wmnet with reason: host reimage * 21:20 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1212 * 21:20 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1212 * 21:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1212.eqiad.wmnet with OS bookworm * 21:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on dse-k8s-etcd2001.codfw.wmnet with reason: host reimage * 21:04 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325553{{!}}InstrumentConstructiveEdits: exclude mw-reverted as well (T431493)]] (duration: 06m 25s) * 20:59 kemayo@deploy1003: kemayo: Continuing with deployment * 20:59 kemayo@deploy1003: kemayo: Backport for [[gerrit:1325553{{!}}InstrumentConstructiveEdits: exclude mw-reverted as well (T431493)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host dse-k8s-etcd2001.codfw.wmnet with OS bookworm * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 20:57 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 20:57 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1325553{{!}}InstrumentConstructiveEdits: exclude mw-reverted as well (T431493)]] * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-etcd2001.codfw.wmnet on all recursors * 20:57 bking@cumin2003: START - Cookbook sre.dns.wipe-cache dse-k8s-etcd2001.codfw.wmnet on all recursors * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 20:57 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 20:54 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320163{{!}}Add configurable RestTermsOfServiceUrl (T428147)]] (duration: 21m 39s) * 20:53 bking@cumin2003: START - Cookbook sre.dns.netbox * 20:53 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host dse-k8s-etcd2001.codfw.wmnet * 20:50 samtar@deploy1003: samtar, milazg: Continuing with deployment * 20:35 samtar@deploy1003: samtar, milazg: Backport for [[gerrit:1320163{{!}}Add configurable RestTermsOfServiceUrl (T428147)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:33 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1320163{{!}}Add configurable RestTermsOfServiceUrl (T428147)]] * 20:30 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325549{{!}}Deploy PRV to several LC wikis (T423785)]] (duration: 06m 57s) * 20:26 arlolra@deploy1003: arlolra: Continuing with deployment * 20:25 arlolra@deploy1003: arlolra: Backport for [[gerrit:1325549{{!}}Deploy PRV to several LC wikis (T423785)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:24 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 20:23 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1325549{{!}}Deploy PRV to several LC wikis (T423785)]] * 20:21 ariel@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319906{{!}}Remove boilerplate language from wmf-rest and wmf-math API modules (T433736)]] (duration: 13m 54s) * 20:14 ariel@deploy1003: ariel: Continuing with deployment * 20:11 ariel@deploy1003: ariel: Backport for [[gerrit:1319906{{!}}Remove boilerplate language from wmf-rest and wmf-math API modules (T433736)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 ariel@deploy1003: Started scap sync-world: Backport for [[gerrit:1319906{{!}}Remove boilerplate language from wmf-rest and wmf-math API modules (T433736)]] * 19:58 robh@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-wdqs2001.codfw.wmnet with reason: updating firmware * 19:54 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325551{{!}}Render the focused module view as a full-screen page (T433896)]] (duration: 30m 37s) * 19:52 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching aqs[2001,1016]*: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 19:44 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching aqs[2001,1016]*: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 19:42 musikanimal@deploy1003: musikanimal: Continuing with deployment * 19:41 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1325551{{!}}Render the focused module view as a full-screen page (T433896)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:34 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: quash java safepoint logspam - bking@cumin2003 - [[phab:T434685|T434685]] * 19:34 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 19:34 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 19:24 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1325551{{!}}Render the focused module view as a full-screen page (T433896)]] * 19:20 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 19:20 jhancock@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin1003" * 19:18 jhancock@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin1003" * 19:14 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325548{{!}}Enable image lazy loading on desktop in group1 (T148047)]] (duration: 07m 43s) * 19:10 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 19:10 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1325548{{!}}Enable image lazy loading on desktop in group1 (T148047)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:07 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1325548{{!}}Enable image lazy loading on desktop in group1 (T148047)]] * 19:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1166.eqiad.wmnet onto db1280.eqiad.wmnet * 19:03 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 19:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1166: Pool db1166.eqiad.wmnet in after cloning * 18:59 jhancock@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 18:58 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1234.eqiad.wmnet with OS bookworm * 18:52 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 18:51 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:49 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:46 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 18:46 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 18:44 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=dns3004.* * 18:39 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1221.eqiad.wmnet with OS bookworm * 18:36 inflatador: [bking@ganeti2048] ~$ sudo gnt-instance replace-disks -n ganeti2030.codfw.wmnet aux-k8s-worker2002.codfw.wmnet [[phab:T434681|T434681]] * 18:35 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 18:29 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1220.eqiad.wmnet with OS bookworm * 18:29 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1234.eqiad.wmnet with reason: host reimage * 18:27 bking@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host dse-k8s-etcd2001.codfw.wmnet * 18:27 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-etcd2001.codfw.wmnet on all recursors * 18:27 bking@cumin2003: START - Cookbook sre.dns.wipe-cache dse-k8s-etcd2001.codfw.wmnet on all recursors * 18:27 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:27 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 18:27 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 18:25 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1234.eqiad.wmnet with reason: host reimage * 18:25 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 18:19 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1221.eqiad.wmnet with reason: host reimage * 18:17 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1166: Pool db1166.eqiad.wmnet in after cloning * 18:15 brennen@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 18:15 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1221.eqiad.wmnet with reason: host reimage * 18:14 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:12 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:12 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 18:12 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-etcd2001.codfw.wmnet on all recursors * 18:12 bking@cumin2003: START - Cookbook sre.dns.wipe-cache dse-k8s-etcd2001.codfw.wmnet on all recursors * 18:12 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:12 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 18:12 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 18:10 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: quash java safepoint logspam - bking@cumin2003 - [[phab:T434685|T434685]] * 18:10 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1234 * 18:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1234 * 18:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1220.eqiad.wmnet with reason: host reimage * 18:07 brennen: 1.47.0-wmf.15 train status ([[phab:T430834|T430834]]) - no current blockers, rolling to all wikis * 18:07 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1234 * 18:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1234.eqiad.wmnet 10.36.64.10.in-addr.arpa 0.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:07 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:07 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1234.eqiad.wmnet 10.36.64.10.in-addr.arpa 0.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1234 - ryankemper@cumin2003" * 18:07 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1234 - ryankemper@cumin2003" * 18:05 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1220.eqiad.wmnet with reason: host reimage * 18:04 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host dse-k8s-etcd2001.codfw.wmnet * 18:04 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 18:03 dancy@deploy1003: Installation of scap version "4.280.1" completed for 3 hosts * 18:02 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:02 inflatador: bking@dse-k8s-etcd2002 etcdctl member remove $<nowiki>{</nowiki>UUID of dse-k8s-etcd2001<nowiki>}</nowiki> [[phab:T434681|T434681]] [[phab:T434793|T434793]] * 18:01 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 18:01 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1234 * 18:01 dancy@deploy1003: Installing scap version "4.280.1" for 3 host(s) * 18:01 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1221 * 18:01 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1221 * 18:00 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1221 * 18:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1221.eqiad.wmnet 18.36.64.10.in-addr.arpa 8.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:00 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1221.eqiad.wmnet 18.36.64.10.in-addr.arpa 8.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1221 - ryankemper@cumin2003" * 17:59 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:58 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 17:58 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:57 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 17:57 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:56 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1221 - ryankemper@cumin2003" * 17:53 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 17:52 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 17:52 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 17:51 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1221 * 17:51 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1220 * 17:51 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1220 * 17:51 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:51 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:51 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1220 * 17:51 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1220.eqiad.wmnet 11.36.64.10.in-addr.arpa 1.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:51 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1220.eqiad.wmnet 11.36.64.10.in-addr.arpa 1.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:51 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1220 - ryankemper@cumin2003" * 17:50 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1220 - ryankemper@cumin2003" * 17:47 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:47 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 17:46 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1234.eqiad.wmnet with OS bookworm * 17:46 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 17:45 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1221.eqiad.wmnet with OS bookworm * 17:45 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1220 * 17:45 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1220.eqiad.wmnet with OS bookworm * 17:43 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 17:41 bking@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host dse-k8s-etcd2001.codfw.wmnet * 17:41 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host dse-k8s-etcd2001.codfw.wmnet * 17:40 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:40 inflatador: bking@ganeti2048] `sudo gnt-instance remove --force --ignore-failures --shutdown-timeout=0` on non-DRBD VMs [[phab:T434681|T434681]] * 17:40 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:39 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:38 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:36 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:32 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:32 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:28 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 17:26 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:26 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:24 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:23 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:21 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1218.eqiad.wmnet with OS bookworm * 17:20 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:20 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:19 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1179.eqiad.wmnet with OS bookworm * 17:18 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1150.eqiad.wmnet with OS bookworm * 17:18 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns3004.wikimedia.org with OS trixie * 17:15 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:14 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:13 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:12 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:12 swfrench@deploy1003: Finished scap sync-world: Helmfile-only deployment for mediawiki chart bump - [[phab:T427666|T427666]] (duration: 03m 03s) * 17:09 swfrench@deploy1003: Started scap sync-world: Helmfile-only deployment for mediawiki chart bump - [[phab:T427666|T427666]] * 17:02 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:01 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:01 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1218.eqiad.wmnet with reason: host reimage * 17:01 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:01 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:00 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324810{{!}}deployment-info.php: Report dbname and branch for the requested wiki (T434726)]] (duration: 06m 52s) * 16:58 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 16:57 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1179.eqiad.wmnet with reason: host reimage * 16:56 dancy@deploy1003: dancy: Continuing with deployment * 16:56 dancy@deploy1003: dancy: Backport for [[gerrit:1324810{{!}}deployment-info.php: Report dbname and branch for the requested wiki (T434726)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1150.eqiad.wmnet with reason: host reimage * 16:53 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324810{{!}}deployment-info.php: Report dbname and branch for the requested wiki (T434726)]] * 16:51 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1166: Depool db1166.eqiad.wmnet to then clone it to db1280.eqiad.wmnet - cwilliams@cumin1003 * 16:50 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1166: Depool db1166.eqiad.wmnet to then clone it to db1280.eqiad.wmnet - cwilliams@cumin1003 * 16:50 cwilliams@cumin1003: START - Cookbook sre.mysql.clone of db1166.eqiad.wmnet onto db1280.eqiad.wmnet * 16:49 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1218.eqiad.wmnet with reason: host reimage * 16:48 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1179.eqiad.wmnet with reason: host reimage * 16:47 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1150.eqiad.wmnet with reason: host reimage * 16:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.upgrade (exit_code=0) for 1 hosts * 16:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2220: Upgrade of db2220.codfw.wmnet completed * 16:38 swfrench@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 16:38 swfrench-wmf: kubectl delete node kubestagemaster2005.codfw.wmnet - [[phab:T434681|T434681]] * 16:34 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1218 * 16:34 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1218 * 16:34 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1218.eqiad.wmnet with OS bookworm * 16:34 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1179 * 16:34 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1179 * 16:33 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1179.eqiad.wmnet with OS bookworm * 16:32 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1150 * 16:32 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1150 * 16:31 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1150.eqiad.wmnet with OS bookworm * 16:24 dancy@deploy1003: Installation of scap version "4.280.0" completed for 3 hosts * 16:22 dancy@deploy1003: Installing scap version "4.280.0" for 3 host(s) * 16:20 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: quash java safepoint logspam - bking@cumin2003 - [[phab:T434685|T434685]] * 16:14 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns3004.wikimedia.org with reason: host reimage * 16:08 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 16:07 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 16:07 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns3004.wikimedia.org with reason: host reimage * 16:05 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 16:04 swfrench@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 16:00 swfrench@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 15:59 swfrench@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 15:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Upgrade of db2220.codfw.wmnet completed * 15:48 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2220: Upgrading db2220.codfw.wmnet * 15:48 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2220: Upgrading db2220.codfw.wmnet * 15:48 cwilliams@cumin1003: START - Cookbook sre.mysql.upgrade for 1 hosts * 15:46 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns3004.wikimedia.org with OS trixie * 15:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2220 [[phab:T434802|T434802]]', diff saved to https://phabricator.wikimedia.org/P96079 and previous config saved to /var/cache/conftool/dbconfig/20260813-154624-cwilliams.json * 15:45 cdobbins@cumin1003: conftool action : set/pooled=no; selector: name=dns3004.* * 15:44 cjd91: depooling dns3004 to reimage to trixie * 15:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2159 to s7 primary [[phab:T434802|T434802]]', diff saved to https://phabricator.wikimedia.org/P96078 and previous config saved to /var/cache/conftool/dbconfig/20260813-154405-cwilliams.json * 15:43 cezmunsta: Starting s7 codfw failover from db2220 to db2159 - [[phab:T434802|T434802]] * 15:41 cgoubert@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: Dragonfly supernodes reboot (duration: 09m 42s) * 15:41 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dragonfly-supernode2001.codfw.wmnet * 15:39 inflatador: bking@ganeti2048] ~$ sudo gnt-node failover -f --ignore-consistency ganeti2046.codfw.wmnet [[phab:T434681|T434681]] * 15:39 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 15:39 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 15:39 swfrench@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 15:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2159 with weight 0 [[phab:T434802|T434802]]', diff saved to https://phabricator.wikimedia.org/P96077 and previous config saved to /var/cache/conftool/dbconfig/20260813-153806-cwilliams.json * 15:37 swfrench@dns1004: END - running authdns-update * 15:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 30 hosts with reason: Primary switchover s7 [[phab:T434802|T434802]] * 15:37 cgoubert@cumin2003: START - Cookbook sre.hosts.reboot-single for host dragonfly-supernode2001.codfw.wmnet * 15:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dragonfly-supernode1001.eqiad.wmnet * 15:35 swfrench@dns1004: START - running authdns-update * 15:32 cgoubert@cumin2003: START - Cookbook sre.hosts.reboot-single for host dragonfly-supernode1001.eqiad.wmnet * 15:32 cgoubert@deploy1003: Locking from deployment [ALL REPOSITORIES]: Dragonfly supernodes reboot * 15:30 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-master-codfw * 15:30 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2005.codfw.wmnet * 15:30 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2005.codfw.wmnet * 15:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1169.eqiad.wmnet onto db1277.eqiad.wmnet * 15:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1169: Pool db1169.eqiad.wmnet in after cloning * 15:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2005.codfw.wmnet * 15:23 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2005.codfw.wmnet * 15:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2004.codfw.wmnet * 15:23 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2004.codfw.wmnet * 15:18 inflatador: bking@ganeti2048 sudo gnt-node failover -f ganeti2046.codfw.wmnet [[phab:T434681|T434681]] * 15:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2004.codfw.wmnet * 15:17 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2004.codfw.wmnet * 15:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2003.codfw.wmnet * 15:17 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2003.codfw.wmnet * 15:15 cdobbins@dns1004: END - running authdns-update * 15:13 cdobbins@dns1004: START - running authdns-update * 15:10 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1217.eqiad.wmnet with OS bookworm * 15:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2003.codfw.wmnet * 15:10 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2003.codfw.wmnet * 15:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2002.codfw.wmnet * 15:10 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2002.codfw.wmnet * 15:10 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: quash java safepoint logspam - bking@cumin2003 - [[phab:T434685|T434685]] * 15:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1216.eqiad.wmnet with OS bookworm * 15:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2002.codfw.wmnet * 15:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2002.codfw.wmnet * 15:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2001.codfw.wmnet * 15:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2001.codfw.wmnet * 15:01 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1042.eqiad.wmnet * 15:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1042.eqiad.wmnet * 15:00 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325455{{!}}Api: Use correct query when continue prop=categories (T433922)]] (duration: 09m 47s) * 14:59 bking@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host dse-k8s-etcd2004.codfw.wmnet * 14:58 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-etcd2004.codfw.wmnet on all recursors * 14:58 bking@cumin2003: START - Cookbook sre.dns.wipe-cache dse-k8s-etcd2004.codfw.wmnet on all recursors * 14:58 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:58 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM dse-k8s-etcd2004.codfw.wmnet - bking@cumin2003" * 14:58 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM dse-k8s-etcd2004.codfw.wmnet - bking@cumin2003" * 14:55 zabe@deploy1003: zabe: Continuing with deployment * 14:53 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-etcd2004.codfw.wmnet on all recursors * 14:53 bking@cumin2003: START - Cookbook sre.dns.wipe-cache dse-k8s-etcd2004.codfw.wmnet on all recursors * 14:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:53 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2004.codfw.wmnet - bking@cumin2003" * 14:53 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2004.codfw.wmnet - bking@cumin2003" * 14:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2001.codfw.wmnet * 14:52 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2001.codfw.wmnet * 14:52 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-master-codfw * 14:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2001.codfw.wmnet * 14:52 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2001.codfw.wmnet * 14:52 zabe@deploy1003: zabe: Backport for [[gerrit:1325455{{!}}Api: Use correct query when continue prop=categories (T433922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:50 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1325455{{!}}Api: Use correct query when continue prop=categories (T433922)]] * 14:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1217.eqiad.wmnet with reason: host reimage * 14:48 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:48 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host dse-k8s-etcd2004.codfw.wmnet * 14:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-master-eqiad * 14:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1006.eqiad.wmnet * 14:44 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1006.eqiad.wmnet * 14:44 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1216.eqiad.wmnet with reason: host reimage * 14:40 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1217.eqiad.wmnet with reason: host reimage * 14:39 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1169: Pool db1169.eqiad.wmnet in after cloning * 14:39 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1216.eqiad.wmnet with reason: host reimage * 14:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl1005.eqiad.wmnet * 14:32 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl1005.eqiad.wmnet * 14:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1004.eqiad.wmnet * 14:32 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1004.eqiad.wmnet * 14:31 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus2008.codfw.wmnet * 14:31 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor1003.eqiad.wmnet * 14:29 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1042.eqiad.wmnet * 14:27 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor1003.eqiad.wmnet * 14:26 moritzm: installing Django security updates * 14:26 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling reboot on A:wikidough * 14:25 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1217.eqiad.wmnet with OS bookworm * 14:25 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1216.eqiad.wmnet with OS bookworm * 14:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl1004.eqiad.wmnet * 14:24 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl1004.eqiad.wmnet * 14:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1003.eqiad.wmnet * 14:24 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1003.eqiad.wmnet * 14:23 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus2008.codfw.wmnet * 14:23 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus1008.eqiad.wmnet * 14:23 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor-dev2001.codfw.wmnet * 14:22 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus2006.codfw.wmnet * 14:19 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor-dev2001.codfw.wmnet * 14:18 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1042.eqiad.wmnet * 14:17 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.* * 14:16 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl1003.eqiad.wmnet * 14:16 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl1003.eqiad.wmnet * 14:16 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1002.eqiad.wmnet * 14:16 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1002.eqiad.wmnet * 14:15 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus1008.eqiad.wmnet * 14:14 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus1006.eqiad.wmnet * 14:12 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus2006.codfw.wmnet * 14:12 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor2003.codfw.wmnet * 14:11 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus2007.codfw.wmnet * 14:11 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1041.eqiad.wmnet * 14:11 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1041.eqiad.wmnet * 14:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl1002.eqiad.wmnet * 14:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl1002.eqiad.wmnet * 14:09 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-master-eqiad * 14:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on apifeatureusage1001.eqiad.wmnet with reason: host reimage * 14:08 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1215.eqiad.wmnet with OS bookworm * 14:08 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor2003.codfw.wmnet * 14:08 jayme: updated calico to v3.30.7 on wikikube codfw [[phab:T427400|T427400]] * 14:07 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1149.eqiad.wmnet with OS bookworm * 14:06 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'. * 14:06 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1041.eqiad.wmnet * 14:04 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus1006.eqiad.wmnet * 14:03 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus2007.codfw.wmnet * 14:03 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus2005.codfw.wmnet * 14:03 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus1007.eqiad.wmnet * 14:03 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1002.eqiad.wmnet * 14:02 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on apifeatureusage1001.eqiad.wmnet with reason: host reimage * 14:02 cgoubert@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=helm-charts.*,name=eqiad * 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host chartmuseum1001.eqiad.wmnet * 14:00 moritzm: installing libxml2 security updates * 13:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1214.eqiad.wmnet with OS bookworm * 13:59 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns5003.* * 13:59 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1041.eqiad.wmnet * 13:59 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow1002.eqiad.wmnet * 13:58 cmooney@dns3003: END - running authdns-update * 13:57 cgoubert@cumin2003: START - Cookbook sre.hosts.reboot-single for host chartmuseum1001.eqiad.wmnet * 13:57 cgoubert@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=helm-charts.*,name=eqiad * 13:57 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: cloudelastic cluster restart - bking@cumin2003 * 13:57 cgoubert@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=helm-charts.*,name=codfw * 13:56 cmooney@dns3003: START - running authdns-update * 13:56 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'. * 13:56 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns5003.*,service=authdns-update * 13:55 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host chartmuseum2001.codfw.wmnet * 13:55 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus1007.eqiad.wmnet * 13:55 cmooney@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dns5003.wikimedia.org * 13:55 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus1005.eqiad.wmnet * 13:51 cgoubert@cumin2003: START - Cookbook sre.hosts.reboot-single for host chartmuseum2001.codfw.wmnet * 13:51 cgoubert@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=helm-charts.*,name=codfw * 13:51 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus2005.codfw.wmnet * 13:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host apifeatureusage1001.eqiad.wmnet with OS bookworm * 13:50 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus7002.magru.wmnet * 13:50 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts lvs1015.eqiad.wmnet * 13:50 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:50 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1015.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:49 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1015.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:49 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325480{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] (duration: 06m 39s) * 13:47 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1215.eqiad.wmnet with reason: host reimage * 13:46 cmooney@cumin1003: START - Cookbook sre.hosts.reboot-single for host dns5003.wikimedia.org * 13:46 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1003.eqiad.wmnet * 13:45 cmooney@cumin1003: conftool action : set/pooled=no; selector: name=dns5003.* * 13:45 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1040.eqiad.wmnet * 13:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1040.eqiad.wmnet * 13:45 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 13:45 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus1005.eqiad.wmnet * 13:45 stran@deploy1003: stran: Continuing with deployment * 13:44 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus7002.magru.wmnet * 13:44 stran@deploy1003: stran: Backport for [[gerrit:1325480{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:44 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus6002.drmrs.wmnet * 13:43 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260805000000" --end-timestamp="20260806000000" --sleep="5" --batch-size="10"` for [[phab:T434688|T434688]] * 13:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1149.eqiad.wmnet with reason: host reimage * 13:42 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1325480{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] * 13:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow1003.eqiad.wmnet * 13:40 sukhe@cumin1003: START - Cookbook sre.hosts.decommission for hosts lvs1015.eqiad.wmnet * 13:40 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts lvs1014.eqiad.wmnet * 13:40 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:40 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1014.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:40 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1040.eqiad.wmnet * 13:40 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1014.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:39 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1214.eqiad.wmnet with reason: host reimage * 13:38 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2004.codfw.wmnet * 13:38 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus6002.drmrs.wmnet * 13:38 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus5003.eqsin.wmnet * 13:35 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 13:35 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1149.eqiad.wmnet with reason: host reimage * 13:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1215.eqiad.wmnet with reason: host reimage * 13:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow2004.codfw.wmnet * 13:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1214.eqiad.wmnet with reason: host reimage * 13:33 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1040.eqiad.wmnet * 13:31 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus5003.eqsin.wmnet * 13:31 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: cloudelastic cluster restart - bking@cumin2003 * 13:31 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus4003.ulsfo.wmnet * 13:30 sukhe@cumin1003: START - Cookbook sre.hosts.decommission for hosts lvs1014.eqiad.wmnet * 13:30 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts lvs1013.eqiad.wmnet * 13:30 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:30 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1013.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:30 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1013.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:27 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1039.eqiad.wmnet * 13:27 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1039.eqiad.wmnet * 13:26 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1167.eqiad.wmnet onto db1281.eqiad.wmnet * 13:26 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1167: Pool db1167.eqiad.wmnet in after cloning * 13:25 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus4003.ulsfo.wmnet * 13:25 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260802000000" --end-timestamp="20260803000000" --sleep="5" --batch-size="10"` for [[phab:T434688|T434688]] * 13:24 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 13:24 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus3004.esams.wmnet * 13:24 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1274: New host * 13:24 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325476{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] (duration: 07m 19s) * 13:24 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260804000000" --end-timestamp="20260805000000" --sleep="5" --batch-size="10"` for [[phab:T434688|T434688]] * 13:24 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260803000000" --end-timestamp="20260804000000" --sleep="5" --batch-size="10"` for [[phab:T434688|T434688]] * 13:23 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2003.codfw.wmnet * 13:22 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1039.eqiad.wmnet * 13:20 sukhe@cumin1003: START - Cookbook sre.hosts.decommission for hosts lvs1013.eqiad.wmnet * 13:20 stran@deploy1003: stran: Continuing with deployment * 13:19 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow2003.codfw.wmnet * 13:19 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling reboot on A:wikidough * 13:19 stran@deploy1003: stran: Backport for [[gerrit:1325476{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:18 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus3004.esams.wmnet * 13:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1215.eqiad.wmnet with OS bookworm * 13:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1214.eqiad.wmnet with OS bookworm * 13:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1149.eqiad.wmnet with OS bookworm * 13:17 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1325476{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] * 13:11 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324731{{!}}prv: Enable parsoid rendering for 5 wikisource wikis]] (duration: 07m 26s) * 13:11 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1039.eqiad.wmnet * 13:11 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1036.eqiad.wmnet * 13:11 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1036.eqiad.wmnet * 13:07 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow3004.esams.wmnet * 13:07 jgiannelos@deploy1003: jgiannelos: Continuing with deployment * 13:06 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1233.eqiad.wmnet with OS bookworm * 13:06 jgiannelos@deploy1003: jgiannelos: Backport for [[gerrit:1324731{{!}}prv: Enable parsoid rendering for 5 wikisource wikis]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:04 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1324731{{!}}prv: Enable parsoid rendering for 5 wikisource wikis]] * 13:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow3004.esams.wmnet * 13:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1036.eqiad.wmnet * 13:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1210.eqiad.wmnet with OS bookworm * 13:01 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1036.eqiad.wmnet * 13:00 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1052.eqiad.wmnet * 13:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1052.eqiad.wmnet * 12:58 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow4003.ulsfo.wmnet * 12:55 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1211.eqiad.wmnet with OS bookworm * 12:55 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1052.eqiad.wmnet * 12:52 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow4003.ulsfo.wmnet * 12:51 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1052.eqiad.wmnet * 12:49 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1169: Depool db1169.eqiad.wmnet to then clone it to db1277.eqiad.wmnet - cwilliams@cumin1003 * 12:46 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1169: Depool db1169.eqiad.wmnet to then clone it to db1277.eqiad.wmnet - cwilliams@cumin1003 * 12:46 cwilliams@cumin1003: START - Cookbook sre.mysql.clone of db1169.eqiad.wmnet onto db1277.eqiad.wmnet * 12:45 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:45 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns record for deleted IP reservations eqsin lvs vlan ints - cmooney@cumin1003" * 12:44 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns record for deleted IP reservations eqsin lvs vlan ints - cmooney@cumin1003" * 12:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1051.eqiad.wmnet * 12:43 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1051.eqiad.wmnet * 12:42 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1210.eqiad.wmnet with reason: host reimage * 12:41 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1167: Pool db1167.eqiad.wmnet in after cloning * 12:40 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 12:39 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1274: New host * 12:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Added db1274', diff saved to https://phabricator.wikimedia.org/P96062 and previous config saved to /var/cache/conftool/dbconfig/20260813-123907-cwilliams.json * 12:38 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast7002.wikimedia.org * 12:38 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1233.eqiad.wmnet with reason: host reimage * 12:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1051.eqiad.wmnet * 12:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1211.eqiad.wmnet with reason: host reimage * 12:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1233.eqiad.wmnet with reason: host reimage * 12:32 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1051.eqiad.wmnet * 12:32 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast7002.wikimedia.org * 12:32 marostegui: Drop SecurePoll tables from closed wikis [[phab:T423128|T423128]] * 12:32 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow5003.eqsin.wmnet * 12:31 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1050.eqiad.wmnet * 12:31 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1050.eqiad.wmnet * 12:29 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1210.eqiad.wmnet with reason: host reimage * 12:29 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1211.eqiad.wmnet with reason: host reimage * 12:29 moritzm: installing Wireshark security updates * 12:26 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow5003.eqsin.wmnet * 12:26 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1050.eqiad.wmnet * 12:24 cmooney@dns3003: END - running authdns-update * 12:21 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow6001.drmrs.wmnet * 12:21 cmooney@dns3003: START - running authdns-update * 12:20 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1050.eqiad.wmnet * 12:20 cgoubert@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host rdb-lock2003.codfw.wmnet * 12:20 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:20 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update netbox dns entries for expanded public1-603-eqsin subnet - cmooney@cumin1003" * 12:20 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update netbox dns entries for expanded public1-603-eqsin subnet - cmooney@cumin1003" * 12:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 12:20 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 12:19 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1049.eqiad.wmnet * 12:19 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1049.eqiad.wmnet * 12:19 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 12:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host moss-be1003.eqiad.wmnet * 12:17 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow6001.drmrs.wmnet * 12:17 moritzm: installin curl security updates * 12:15 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 12:15 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1233.eqiad.wmnet with OS bookworm * 12:15 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1210.eqiad.wmnet with OS bookworm * 12:15 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1211.eqiad.wmnet with OS bookworm * 12:13 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1049.eqiad.wmnet * 12:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow7002.magru.wmnet * 12:12 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-worker1178.eqiad.wmnet with OS bookworm * 12:10 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host moss-be1003.eqiad.wmnet * 12:10 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 12:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be1006.eqiad.wmnet * 12:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 12:10 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 12:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 12:10 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 12:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow7002.magru.wmnet * 12:08 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1049.eqiad.wmnet * 12:05 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 12:05 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2003.codfw.wmnet * 12:04 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'. * 12:04 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:03 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be1006.eqiad.wmnet * 12:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be1005.eqiad.wmnet * 12:01 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1038.eqiad.wmnet * 12:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1038.eqiad.wmnet * 12:01 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1213.eqiad.wmnet with OS bookworm * 11:56 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be1005.eqiad.wmnet * 11:55 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be1004.eqiad.wmnet * 11:54 cgoubert@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host rdb-lock2003.codfw.wmnet * 11:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 11:53 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 11:53 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:53 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 11:53 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 11:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1038.eqiad.wmnet * 11:49 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be1004.eqiad.wmnet * 11:44 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:44 moritzm: installing Linux 5.10.262 on Bullseye hosts * 11:41 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1038.eqiad.wmnet * 11:40 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 11:40 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 11:40 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 11:40 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:40 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 11:40 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 11:38 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1213.eqiad.wmnet with reason: host reimage * 11:36 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1167: Depool db1167.eqiad.wmnet to then clone it to db1281.eqiad.wmnet - marostegui@cumin1003 * 11:35 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 11:35 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2003.codfw.wmnet * 11:35 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1167: Depool db1167.eqiad.wmnet to then clone it to db1281.eqiad.wmnet - marostegui@cumin1003 * 11:35 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1167.eqiad.wmnet onto db1281.eqiad.wmnet * 11:34 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 22 hosts with reason: Cloning * 11:34 moritzm: remove ganeti3005 from esams03 cluster, hardware issues [[phab:T434646|T434646]] * 11:32 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1213.eqiad.wmnet with reason: host reimage * 11:28 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1034.eqiad.wmnet * 11:28 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1034.eqiad.wmnet * 11:22 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1034.eqiad.wmnet * 11:19 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1034.eqiad.wmnet * 11:17 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1213.eqiad.wmnet with OS bookworm * 11:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1165.eqiad.wmnet onto db1279.eqiad.wmnet * 11:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1165: Pool db1165.eqiad.wmnet in after cloning * 11:07 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply * 10:57 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply * 10:54 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply * 10:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1274.eqiad.wmnet with reason: Enabling notifications and pooling * 10:45 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1033.eqiad.wmnet * 10:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1033.eqiad.wmnet * 10:44 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply * 10:43 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'. * 10:42 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'. * 10:42 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'. * 10:40 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db[1216,1225,1239-1240].eqiad.wmnet with reason: reboot * 10:39 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1033.eqiad.wmnet * 10:38 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1151 from dbctl [[phab:T434538|T434538]]', diff saved to https://phabricator.wikimedia.org/P96055 and previous config saved to /var/cache/conftool/dbconfig/20260813-103828-marostegui.json * 10:35 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1033.eqiad.wmnet * 10:27 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1165: Pool db1165.eqiad.wmnet in after cloning * 10:24 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1151.eqiad.wmnet with OS bookworm * 10:15 moritzm: installing bind9 security updates (client-side tools/libs only) * 10:07 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-debug: apply * 10:06 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-debug: apply * 10:02 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 7 hosts with reason: reboot * 10:01 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-debug: apply * 10:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.upgrade (exit_code=0) for 1 hosts * 10:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2214: Upgrade of db2214.codfw.wmnet completed * 10:01 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-debug: apply * 10:00 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-debug: apply * 10:00 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-debug: apply * 09:59 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.decommission (exit_code=99) * 09:59 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 09:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1151.eqiad.wmnet with reason: host reimage * 09:59 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 8 hosts * 09:59 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 8 hosts * 09:55 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1151.eqiad.wmnet with reason: host reimage * 09:46 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260801000000" --end-timestamp="20260802000000" --sleep="3" --batch-size="5"` for [[phab:T434688|T434688]] * 09:45 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 8 hosts with reason: reboot * 09:45 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 22 hosts with reason: Cloning * 09:43 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1165: Depool db1165.eqiad.wmnet to then clone it to db1279.eqiad.wmnet - marostegui@cumin1003 * 09:42 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1165: Depool db1165.eqiad.wmnet to then clone it to db1279.eqiad.wmnet - marostegui@cumin1003 * 09:42 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1165.eqiad.wmnet onto db1279.eqiad.wmnet * 09:41 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=testwiki --start-timestamp="20260311000000" --end-timestamp="20260805000000" --sleep="5" --batch-size="2"` for [[phab:T434688|T434688]] * 09:40 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for backupmon1001.eqiad.wmnet * 09:40 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for backupmon1001.eqiad.wmnet * 09:40 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1151 * 09:40 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1151 * 09:37 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325408{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]], [[gerrit:1325407{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]] (duration: 06m 57s) * 09:36 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1151 * 09:36 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1151.eqiad.wmnet 13.36.64.10.in-addr.arpa 3.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:36 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1151.eqiad.wmnet 13.36.64.10.in-addr.arpa 3.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:36 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:36 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1151 - btullis@cumin1003" * 09:36 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on backupmon1001.eqiad.wmnet with reason: reboot * 09:36 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1151 - btullis@cumin1003" * 09:35 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host moss-be2003.codfw.wmnet * 09:33 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 7 hosts * 09:33 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 7 hosts * 09:33 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 09:32 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1325408{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]], [[gerrit:1325407{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1161.eqiad.wmnet onto db1275.eqiad.wmnet * 09:31 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 09:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1161: Pool db1161.eqiad.wmnet in after cloning * 09:30 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1325408{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]], [[gerrit:1325407{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]] * 09:29 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 09:27 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 09:27 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host moss-be2003.codfw.wmnet * 09:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be2006.codfw.wmnet * 09:25 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2004.codfw.wmnet * 09:22 hashar@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 09:21 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be2006.codfw.wmnet * 09:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be2005.codfw.wmnet * 09:19 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2004.codfw.wmnet * 09:19 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2003.codfw.wmnet * 09:18 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 7 hosts with reason: reboot * 09:18 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 7 hosts * 09:18 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 7 hosts * 09:16 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2214: Upgrade of db2214.codfw.wmnet completed * 09:14 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be2005.codfw.wmnet * 09:13 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be2004.codfw.wmnet * 09:12 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2003.codfw.wmnet * 09:12 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2002.codfw.wmnet * 09:09 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2214: Upgrading db2214.codfw.wmnet * 09:09 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2214: Upgrading db2214.codfw.wmnet * 09:09 cwilliams@cumin1003: START - Cookbook sre.mysql.upgrade for 1 hosts * 09:07 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be2004.codfw.wmnet * 09:06 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2002.codfw.wmnet * 09:05 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1004.eqiad.wmnet * 09:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-cluster (exit_code=0) * 09:03 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 7 hosts with reason: reboot * 09:01 btullis@cumin1003: START - Cookbook sre.dns.netbox * 09:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2214 [[phab:T434754|T434754]]', diff saved to https://phabricator.wikimedia.org/P96045 and previous config saved to /var/cache/conftool/dbconfig/20260813-090001-cwilliams.json * 08:59 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1004.eqiad.wmnet * 08:59 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1003.eqiad.wmnet * 08:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2229 to s6 primary [[phab:T434754|T434754]]', diff saved to https://phabricator.wikimedia.org/P96044 and previous config saved to /var/cache/conftool/dbconfig/20260813-085752-cwilliams.json * 08:57 cezmunsta: Starting s6 codfw failover from db2214 to db2229 - [[phab:T434754|T434754]] * 08:54 hashar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325416{{!}}Revert "REST: Enable `GET /lexemes/<nowiki>{</nowiki>lexeme_id<nowiki>}</nowiki>` by default" (T434712)]] (duration: 07m 22s) * 08:53 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1003.eqiad.wmnet * 08:53 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1002.eqiad.wmnet * 08:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2229 with weight 0 [[phab:T434754|T434754]]', diff saved to https://phabricator.wikimedia.org/P96043 and previous config saved to /var/cache/conftool/dbconfig/20260813-085151-cwilliams.json * 08:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 23 hosts with reason: Primary switchover s6 [[phab:T434754|T434754]] * 08:51 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts2002.codfw.wmnet * 08:50 hashar@deploy1003: hashar: Continuing with deployment * 08:49 hashar@deploy1003: hashar: Backport for [[gerrit:1325416{{!}}Revert "REST: Enable `GET /lexemes/<nowiki>{</nowiki>lexeme_id<nowiki>}</nowiki>` by default" (T434712)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:47 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet * 08:47 hashar@deploy1003: Started scap sync-world: Backport for [[gerrit:1325416{{!}}Revert "REST: Enable `GET /lexemes/<nowiki>{</nowiki>lexeme_id<nowiki>}</nowiki>` by default" (T434712)]] * 08:47 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1002.eqiad.wmnet * 08:46 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1161: Pool db1161.eqiad.wmnet in after cloning * 08:46 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host stewards2001.codfw.wmnet * 08:45 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit1003.wikimedia.org * 08:45 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host stewards1001.eqiad.wmnet * 08:45 Emperor: roll-restart apus frontends in codfw * 08:45 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-cluster * 08:44 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts2002.codfw.wmnet * 08:44 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host doc2003.codfw.wmnet * 08:43 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet * 08:43 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host phab2003.codfw.wmnet * 08:42 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host stewards2001.codfw.wmnet * 08:42 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host etherpad2002.codfw.wmnet * 08:41 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host stewards1001.eqiad.wmnet * 08:41 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1273: Pool in s7 * 08:41 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1003.wikimedia.org * 08:40 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host doc2003.codfw.wmnet * 08:40 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit1003.wikimedia.org * 08:39 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host etherpad1004.eqiad.wmnet * 08:39 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host doc1004.eqiad.wmnet * 08:39 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit2002.wikimedia.org * 08:38 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host etherpad2002.codfw.wmnet * 08:37 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host phab2003.codfw.wmnet * 08:36 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lists1004.wikimedia.org * 08:35 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host etherpad1004.eqiad.wmnet * 08:35 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host doc1004.eqiad.wmnet * 08:34 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-cluster (exit_code=0) * 08:34 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1003.wikimedia.org * 08:34 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host planet2003.codfw.wmnet * 08:34 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2003.wikimedia.org * 08:33 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host planet1003.eqiad.wmnet * 08:33 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit2002.wikimedia.org * 08:32 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 22 hosts with reason: Cloning * 08:32 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2002.wikimedia.org * 08:31 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aphlict1002.eqiad.wmnet * 08:30 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host planet2003.codfw.wmnet * 08:29 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host planet1003.eqiad.wmnet * 08:29 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host lists1004.wikimedia.org * 08:28 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lists2001.wikimedia.org * 08:28 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2003.wikimedia.org * 08:27 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host aphlict1002.eqiad.wmnet * 08:27 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aphlict2001.codfw.wmnet * 08:26 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2002.wikimedia.org * 08:23 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host aphlict2001.codfw.wmnet * 08:23 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1151 * 08:22 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1151.eqiad.wmnet with OS bookworm * 08:22 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host lists2001.wikimedia.org * 08:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1161: Depool db1161.eqiad.wmnet to then clone it to db1275.eqiad.wmnet - marostegui@cumin1003 * 08:20 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1161: Depool db1161.eqiad.wmnet to then clone it to db1275.eqiad.wmnet - marostegui@cumin1003 * 08:20 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1161.eqiad.wmnet onto db1275.eqiad.wmnet * 08:15 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-cluster * 08:15 Emperor: roll-restart apus frontends in eqiad * 07:56 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1273: Pool in s7 * 07:56 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1273 to dbctl [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P96035 and previous config saved to /var/cache/conftool/dbconfig/20260813-075611-marostegui.json * 07:38 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db1273.eqiad.wmnet with reason: Reboot * 07:31 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.sanitize-wiki (exit_code=97) Managing sanitization for wikis testwiki in section s3 * 07:24 marostegui@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis testwiki in section s3 * 07:19 jayme@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on kubestagemaster2005.codfw.wmnet with reason: downtime because of hardware failure and no DRBD * 05:42 arnaudb@dns1006: END - running authdns-update * 05:40 arnaudb@dns1006: START - running authdns-update * 05:27 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 04:06 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 02:29 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1151 * 02:29 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1151 * 02:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1219.eqiad.wmnet with OS bookworm * 02:18 ryankemper: [[phab:T434494|T434494]] `ryankemper@deploy1003:~$ echo 'https://stats.wikimedia.org/' {{!}} mwscript-k8s --attach -- purgeList.php` (default page got cached during yesterday's `an-web1001` reimage) * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 46s) * 02:03 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1219.eqiad.wmnet with reason: host reimage * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 02:00 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1219.eqiad.wmnet with reason: host reimage * 01:46 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1219.eqiad.wmnet with OS bookworm * 01:03 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1142.eqiad.wmnet with OS bookworm * 00:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1142.eqiad.wmnet with reason: host reimage * 00:34 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1142.eqiad.wmnet with reason: host reimage * 00:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1142.eqiad.wmnet with OS bookworm * 00:16 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1178 * 00:16 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1178 == 2026-08-12 == * 23:18 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324409{{!}}Remove $wmg = $wg hacks in Collection (T119117)]] (duration: 06m 43s) * 23:14 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 23:13 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1324409{{!}}Remove $wmg = $wg hacks in Collection (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:11 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1324409{{!}}Remove $wmg = $wg hacks in Collection (T119117)]] * 22:59 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 22:49 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260801000000" --end-timestamp="20260802000000" --sleep=2 --batch-size=10` * 22:45 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=testwiki --start-timestamp="20200801010101" --end-timestamp="20260816010101" --sleep=15 --batch-size=5` * 22:40 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=testwiki --start-timestamp="20200101010101" --end-timestamp="20260816010101" --sleep=60` * 22:35 Dreamy_Jazz: Running `mwscript WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260101000000" --end-timestamp="20260102000000" --sleep=10` * 22:21 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324817{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]], [[gerrit:1324818{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]] (duration: 45m 29s) * 22:17 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 21:59 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 21:40 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1324817{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]], [[gerrit:1324818{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:39 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 21:36 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1324817{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]], [[gerrit:1324818{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]] * 21:32 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:31 vriley@cumin1003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:30 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:30 vriley@cumin1003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:17 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:14 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:14 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:11 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:10 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-codfw: Set storage compatability to NONE — [[phab:T433028|T433028]] - eevans@cumin1003 * 21:10 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1006 * 21:09 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1006 * 21:05 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324277{{!}}Improve Math preference labels for SVG/MathJax/MathML (T433891)]] (duration: 31m 42s) * 20:58 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1178.eqiad.wmnet with OS bookworm * 20:54 krinkle@deploy1003: krinkle: Continuing with deployment * 20:51 krinkle@deploy1003: krinkle: Backport for [[gerrit:1324277{{!}}Improve Math preference labels for SVG/MathJax/MathML (T433891)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:39 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 20:38 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:38 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp5022.eqsin.wmnet with OS trixie * 20:38 cdobbins@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - cdobbins@cumin1003" * 20:37 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:36 cdobbins@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - cdobbins@cumin1003" * 20:34 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1178.eqiad.wmnet with reason: host reimage * 20:34 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1324277{{!}}Improve Math preference labels for SVG/MathJax/MathML (T433891)]] * 20:33 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 20:28 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1178.eqiad.wmnet with reason: host reimage * 20:13 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1178.eqiad.wmnet with OS bookworm * 20:11 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-worker1178.eqiad.wmnet with OS bookworm * 20:11 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1178.eqiad.wmnet with OS bookworm * 20:09 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-codfw: Set storage compatability to NONE — [[phab:T433028|T433028]] - eevans@cumin1003 * 20:09 Dreamy_Jazz: Evening UTC backport window done * 20:08 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324794{{!}}WikimediaAntiAbuse: Enable logging channel (T431292)]] (duration: 06m 48s) * 20:08 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage * 20:05 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage * 20:04 dreamyjazz@deploy1003: kharlan, dreamyjazz: Continuing with deployment * 20:04 dreamyjazz@deploy1003: kharlan, dreamyjazz: Backport for [[gerrit:1324794{{!}}WikimediaAntiAbuse: Enable logging channel (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:01 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1324794{{!}}WikimediaAntiAbuse: Enable logging channel (T431292)]] * 19:55 brennen@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 19:47 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 19:35 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 19:35 cdobbins@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cp5022.eqsin.wmnet with OS trixie * 19:32 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:30 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:29 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:26 vriley@cumin1003: START - Cookbook sre.dns.netbox * 19:23 brennen@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 19:19 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-eqiad: Set storage compatability to NONE — [[phab:T433028|T433028]] - eevans@cumin1003 * 19:10 brennen@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 19:09 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 19:09 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 19:00 Amir1: data migrated on wikishared ([[phab:T426102|T426102]]) * 18:57 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321224{{!}}Rename ce_worklist_articles table to ce_invitation_list_articles (T426102)]] (duration: 06m 50s) * 18:53 ladsgroup@deploy1003: ladsgroup, daimona: Continuing with deployment * 18:53 ladsgroup@deploy1003: ladsgroup, daimona: Backport for [[gerrit:1321224{{!}}Rename ce_worklist_articles table to ce_invitation_list_articles (T426102)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:51 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1321224{{!}}Rename ce_worklist_articles table to ce_invitation_list_articles (T426102)]] * 18:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2192: Security update * 18:31 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 18:25 Amir1: ce_invitation_list_articles created as empty on wikishared ([[phab:T426102|T426102]]) * 18:21 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 18:21 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 18:18 Amir1: migrated testwiki entries from ce_worklist_articles to ce_invitation_list_articles ([[phab:T426102|T426102]]) * 18:18 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-eqiad: Set storage compatability to NONE — [[phab:T433028|T433028]] - eevans@cumin1003 * 18:11 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 18:08 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling reboot on A:durum-eqsin and A:durum * 18:07 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Set storage compatability to UPGRADING — [[phab:T433028|T433028]] - eevans@cumin1003 * 18:05 jhancock@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022'] * 17:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2192: Security update * 17:55 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:55 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum-eqsin and A:durum * 17:53 jhancock@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['cp5022'] * 17:47 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:47 jhancock@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['cp5022'] * 17:42 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:41 jhancock@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['cp5022'] * 17:36 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:35 jhancock@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022'] * 17:31 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-magru and not (P<nowiki>{</nowiki>cp7001*<nowiki>}</nowiki> or P<nowiki>{</nowiki>cp7009*<nowiki>}</nowiki>) and A:cp - 9.2.15 upgrade ([[phab:T434620|T434620]]) * 17:28 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:28 jhancock@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['cp5022'] * 17:22 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2192.codfw.wmnet with reason: Maintenance * 17:11 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host apifeatureusage2001.codfw.wmnet with OS bookworm * 17:04 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: apply * 17:03 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-main: apply * 17:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2192 [[phab:T434635|T434635]]', diff saved to https://phabricator.wikimedia.org/P96030 and previous config saved to /var/cache/conftool/dbconfig/20260812-170338-cwilliams.json * 17:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2213 to s5 primary [[phab:T434635|T434635]]', diff saved to https://phabricator.wikimedia.org/P96029 and previous config saved to /var/cache/conftool/dbconfig/20260812-170152-cwilliams.json * 17:01 cezmunsta: Starting s5 codfw failover from db2192 to db2213 - [[phab:T434635|T434635]] * 16:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2213 with weight 0 [[phab:T434635|T434635]]', diff saved to https://phabricator.wikimedia.org/P96028 and previous config saved to /var/cache/conftool/dbconfig/20260812-165544-cwilliams.json * 16:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 27 hosts with reason: Primary switchover s5 [[phab:T434635|T434635]] * 16:53 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: apply * 16:52 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-main: apply * 16:44 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-main: apply * 16:44 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-main: apply * 16:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1272: New host * 16:40 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324765{{!}}WikimediaAntiAbuse: Enable PersonalInfoFlagNotifications (T431292)]] (duration: 07m 02s) * 16:40 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-logging-external: apply * 16:39 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-logging-external: apply * 16:38 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-logging-external: apply * 16:37 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-logging-external: apply * 16:36 kharlan@deploy1003: kharlan: Continuing with deployment * 16:35 kharlan@deploy1003: kharlan: Backport for [[gerrit:1324765{{!}}WikimediaAntiAbuse: Enable PersonalInfoFlagNotifications (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:33 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1324765{{!}}WikimediaAntiAbuse: Enable PersonalInfoFlagNotifications (T431292)]] * 16:27 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-logging-external: apply * 16:27 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-logging-external: apply * 16:18 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 16:18 jhancock@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022'] * 16:17 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 16:16 jhancock@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['cp5022'] * 16:15 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 16:12 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Set storage compatability to UPGRADING — [[phab:T433028|T433028]] - eevans@cumin1003 * 16:10 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: apply * 16:10 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: apply * 16:08 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: apply * 16:08 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: apply * 16:08 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics: apply * 16:07 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics: apply * 16:02 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-magru and not (P<nowiki>{</nowiki>cp7001*<nowiki>}</nowiki> or P<nowiki>{</nowiki>cp7009*<nowiki>}</nowiki>) and A:cp - 9.2.15 upgrade ([[phab:T434620|T434620]]) * 15:57 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1272: New host * 15:52 jmm@cumin2003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti2046.codfw.wmnet * 15:52 jmm@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host ganeti2046.codfw.wmnet * 15:42 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:37 mforns@deploy1003: Finished deploy [analytics/refinery@49c336c] (thin): Regular analytics weekly train THIN [analytics/refinery@49c336cd] (duration: 01m 59s) * 15:37 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324733{{!}}WikimediaAntiAbuse: Enable personal info tag display on enwiki (T431292)]] (duration: 08m 12s) * 15:35 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:35 mforns@deploy1003: Started deploy [analytics/refinery@49c336c] (thin): Regular analytics weekly train THIN [analytics/refinery@49c336cd] * 15:34 mforns@deploy1003: Finished deploy [analytics/refinery@49c336c]: Regular analytics weekly train [analytics/refinery@49c336cd] (duration: 04m 20s) * 15:33 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[2024,1031]*.wmnet: Set storage compatability to UPGRADING — [[phab:T433028|T433028]] - eevans@cumin1003 * 15:33 dreamyjazz@deploy1003: kharlan, dreamyjazz: Continuing with deployment * 15:31 dreamyjazz@deploy1003: kharlan, dreamyjazz: Backport for [[gerrit:1324733{{!}}WikimediaAntiAbuse: Enable personal info tag display on enwiki (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:30 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitize-wiki (exit_code=99) Checking sanitization for wikis testwiki in section s3 * 15:30 mforns@deploy1003: Started deploy [analytics/refinery@49c336c]: Regular analytics weekly train [analytics/refinery@49c336cd] * 15:30 mforns@deploy1003: Finished deploy [analytics/refinery@49c336c] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@49c336cd] (duration: 00m 32s) * 15:29 mforns@deploy1003: Started deploy [analytics/refinery@49c336c] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@49c336cd] * 15:29 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1324733{{!}}WikimediaAntiAbuse: Enable personal info tag display on enwiki (T431292)]] * 15:27 brennen@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324700{{!}}EventDetailsParticipantsModule: populate cache with non-local users (T434597)]] (duration: 06m 38s) * 15:23 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[2024,1031]*.wmnet: Set storage compatability to UPGRADING — [[phab:T433028|T433028]] - eevans@cumin1003 * 15:23 brennen@deploy1003: brennen, daimona: Continuing with deployment * 15:23 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:22 brennen@deploy1003: brennen, daimona: Backport for [[gerrit:1324700{{!}}EventDetailsParticipantsModule: populate cache with non-local users (T434597)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:22 cgoubert@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host rdb-lock2003.codfw.wmnet * 15:21 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 15:21 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 15:21 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:21 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:21 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:20 brennen@deploy1003: Started scap sync-world: Backport for [[gerrit:1324700{{!}}EventDetailsParticipantsModule: populate cache with non-local users (T434597)]] * 15:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1178.eqiad.wmnet with OS bookworm * 15:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock1003.eqiad.wmnet * 15:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock1003.eqiad.wmnet with OS trixie * 15:16 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 15:16 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 15:16 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 15:16 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:16 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:16 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:12 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 15:12 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2003.codfw.wmnet * 15:11 cgoubert@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host rdb-lock2003.codfw.wmnet * 15:11 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 15:11 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 15:11 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:11 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:11 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:07 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324338{{!}}InitialiseSettings: Enable 2FA warnings on more private wikis (T428103)]] (duration: 07m 02s) * 15:04 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 15:04 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock1003.eqiad.wmnet with reason: host reimage * 15:03 reedy@deploy1003: reedy: Continuing with deployment * 15:02 reedy@deploy1003: reedy: Backport for [[gerrit:1324338{{!}}InitialiseSettings: Enable 2FA warnings on more private wikis (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:02 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:00 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324338{{!}}InitialiseSettings: Enable 2FA warnings on more private wikis (T428103)]] * 14:57 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 14:57 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2003.codfw.wmnet * 14:57 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock1003.eqiad.wmnet with reason: host reimage * 14:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock2002.codfw.wmnet * 14:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock2002.codfw.wmnet with OS trixie * 14:56 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 14:56 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 14:56 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 14:55 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 14:55 moritzm: powercycle ganeti2046 * 14:47 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock1003.eqiad.wmnet with OS trixie * 14:46 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1003.eqiad.wmnet - cgoubert@cumin2003" * 14:46 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1003.eqiad.wmnet - cgoubert@cumin2003" * 14:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock1003.eqiad.wmnet on all recursors * 14:45 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock1003.eqiad.wmnet on all recursors * 14:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1003.eqiad.wmnet - cgoubert@cumin2003" * 14:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1272.eqiad.wmnet with reason: Enabling notifications * 14:44 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1003.eqiad.wmnet - cgoubert@cumin2003" * 14:44 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324719{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324720{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324722{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0 (T434187)]] (duration: 11m 02s) * 14:39 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 14:39 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock1003.eqiad.wmnet * 14:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock2002.codfw.wmnet with reason: host reimage * 14:37 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 14:37 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock1002.eqiad.wmnet * 14:37 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock1002.eqiad.wmnet with OS trixie * 14:37 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1324719{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324720{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324722{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0 (T434187)]] synced to the testservers (see https://wikitech. * 14:33 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock2002.codfw.wmnet with reason: host reimage * 14:33 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1324719{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324720{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324722{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0 (T434187)]] * 14:32 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1018.eqiad.wmnet with OS bookworm * 14:32 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2046.codfw.wmnet * 14:31 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1020.eqiad.wmnet with OS bookworm * 14:27 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2046.codfw.wmnet * 14:25 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2045.codfw.wmnet * 14:25 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2045.codfw.wmnet * 14:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock1002.eqiad.wmnet with reason: host reimage * 14:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1019.eqiad.wmnet with OS bookworm * 14:22 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Checking sanitization for wikis testwiki in section s3 * 14:20 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2045.codfw.wmnet * 14:18 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock1002.eqiad.wmnet with reason: host reimage * 14:17 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:16 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324707{{!}}Backport all changes from wmf/1.47.0-wmf.15]] (duration: 40m 51s) * 14:16 moritzm: installing Linux 6.1.180 on Bookworm hosts * 14:15 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock2002.codfw.wmnet with OS trixie * 14:14 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2002.codfw.wmnet - cgoubert@cumin2003" * 14:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2002.codfw.wmnet - cgoubert@cumin2003" * 14:14 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:14 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2045.codfw.wmnet * 14:14 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2002.codfw.wmnet on all recursors * 14:14 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2002.codfw.wmnet on all recursors * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2002.codfw.wmnet - cgoubert@cumin2003" * 14:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2002.codfw.wmnet - cgoubert@cumin2003" * 14:12 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2030.codfw.wmnet * 14:12 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2030.codfw.wmnet * 14:11 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on apifeatureusage2001.codfw.wmnet with reason: host reimage * 14:09 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:08 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:07 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 14:06 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2030.codfw.wmnet * 14:06 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:06 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock1002.eqiad.wmnet with OS trixie * 14:05 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 14:05 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2002.codfw.wmnet * 14:05 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1002.eqiad.wmnet - cgoubert@cumin2003" * 14:05 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1002.eqiad.wmnet - cgoubert@cumin2003" * 14:05 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock1002.eqiad.wmnet on all recursors * 14:05 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock1002.eqiad.wmnet on all recursors * 14:05 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:05 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1002.eqiad.wmnet - cgoubert@cumin2003" * 14:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:04 kharlan@deploy1003: kharlan: Continuing with deployment * 14:04 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock2001.codfw.wmnet * 14:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock2001.codfw.wmnet with OS trixie * 14:02 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1002.eqiad.wmnet - cgoubert@cumin2003" * 14:02 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on apifeatureusage2001.codfw.wmnet with reason: host reimage * 14:01 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2030.codfw.wmnet * 13:59 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2029.codfw.wmnet * 13:58 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2029.codfw.wmnet * 13:58 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 13:58 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock1002.eqiad.wmnet * 13:56 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1019.eqiad.wmnet with reason: host reimage * 13:54 btullis@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on archiva1002.wikimedia.org with reason: Upgrading in-place * 13:53 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock1001.eqiad.wmnet * 13:53 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock1001.eqiad.wmnet with OS trixie * 13:53 kharlan@deploy1003: kharlan: Backport for [[gerrit:1324707{{!}}Backport all changes from wmf/1.47.0-wmf.15]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:52 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2029.codfw.wmnet * 13:52 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1020.eqiad.wmnet with reason: host reimage * 13:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1019.eqiad.wmnet with reason: host reimage * 13:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1020.eqiad.wmnet with reason: host reimage * 13:49 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock2001.codfw.wmnet with reason: host reimage * 13:48 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2029.codfw.wmnet * 13:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Configuring db1272 for s3 pooling', diff saved to https://phabricator.wikimedia.org/P96021 and previous config saved to /var/cache/conftool/dbconfig/20260812-134732-cwilliams.json * 13:44 bking@cumin2003: START - Cookbook sre.hosts.reimage for host apifeatureusage2001.codfw.wmnet with OS bookworm * 13:43 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock2001.codfw.wmnet with reason: host reimage * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2028.codfw.wmnet * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2028.codfw.wmnet * 13:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1018.eqiad.wmnet with reason: host reimage * 13:38 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock1001.eqiad.wmnet with reason: host reimage * 13:36 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1018.eqiad.wmnet with reason: host reimage * 13:35 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1324707{{!}}Backport all changes from wmf/1.47.0-wmf.15]] * 13:35 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2028.codfw.wmnet * 13:32 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1275938{{!}}Enable campaignEvents on bdwikimedia (T424016)]] (duration: 07m 35s) * 13:32 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1020.eqiad.wmnet with OS bookworm * 13:32 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1019.eqiad.wmnet with OS bookworm * 13:32 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock1001.eqiad.wmnet with reason: host reimage * 13:28 kharlan@deploy1003: kharlan, yahya: Continuing with deployment * 13:27 kharlan@deploy1003: kharlan, yahya: Backport for [[gerrit:1275938{{!}}Enable campaignEvents on bdwikimedia (T424016)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:27 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2028.codfw.wmnet * 13:26 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 13:26 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 13:25 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1275938{{!}}Enable campaignEvents on bdwikimedia (T424016)]] * 13:25 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitize-wiki (exit_code=99) Managing sanitization for wikis testwiki in section s3 * 13:24 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock2001.codfw.wmnet with OS trixie * 13:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2001.codfw.wmnet - cgoubert@cumin2003" * 13:24 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2001.codfw.wmnet - cgoubert@cumin2003" * 13:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2001.codfw.wmnet on all recursors * 13:23 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2001.codfw.wmnet on all recursors * 13:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2001.codfw.wmnet - cgoubert@cumin2003" * 13:23 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2001.codfw.wmnet - cgoubert@cumin2003" * 13:23 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324713{{!}}thwiki: reinstate temporary wiki25 logos (T431094)]] (duration: 07m 13s) * 13:20 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:20 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1018.eqiad.wmnet with OS bookworm * 13:20 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2027.codfw.wmnet * 13:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2027.codfw.wmnet * 13:19 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics-external: apply * 13:18 kharlan@deploy1003: anzx, kharlan: Continuing with deployment * 13:18 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:18 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock1001.eqiad.wmnet with OS trixie * 13:17 kharlan@deploy1003: anzx, kharlan: Backport for [[gerrit:1324713{{!}}thwiki: reinstate temporary wiki25 logos (T431094)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1001.eqiad.wmnet - cgoubert@cumin2003" * 13:17 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 13:17 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1001.eqiad.wmnet - cgoubert@cumin2003" * 13:17 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2001.codfw.wmnet * 13:17 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics-external: apply * 13:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock1001.eqiad.wmnet on all recursors * 13:17 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock1001.eqiad.wmnet on all recursors * 13:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1001.eqiad.wmnet - cgoubert@cumin2003" * 13:17 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1001.eqiad.wmnet - cgoubert@cumin2003" * 13:15 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1324713{{!}}thwiki: reinstate temporary wiki25 logos (T431094)]] * 13:15 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:15 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics-external: apply * 13:15 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:14 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics-external: apply * 13:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2027.codfw.wmnet * 13:12 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 13:12 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock1001.eqiad.wmnet * 13:12 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2027.codfw.wmnet * 13:08 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1158.eqiad.wmnet onto db1273.eqiad.wmnet * 13:07 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1158: Pool db1158.eqiad.wmnet in after cloning * 13:02 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti6002.drmrs.wmnet * 13:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti6002.drmrs.wmnet * 12:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1015.eqiad.wmnet with OS bookworm * 12:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti6002.drmrs.wmnet * 12:44 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1017.eqiad.wmnet with OS bookworm * 12:36 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 12:35 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 12:34 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 12:33 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 12:31 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti6002.drmrs.wmnet * 12:24 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1017.eqiad.wmnet with reason: host reimage * 12:22 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1158: Pool db1158.eqiad.wmnet in after cloning * 12:18 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1017.eqiad.wmnet with reason: host reimage * 12:11 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1015.eqiad.wmnet with reason: host reimage * 12:07 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1015.eqiad.wmnet with reason: host reimage * 12:04 moritzm: failover ganeti master in drmrs02 to ganeti6004 * 12:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1009.eqiad.wmnet with OS bookworm * 12:01 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1017.eqiad.wmnet with OS bookworm * 12:00 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti6004.drmrs.wmnet * 12:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti6004.drmrs.wmnet * 11:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti6004.drmrs.wmnet * 11:53 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1015.eqiad.wmnet with OS bookworm * 11:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1159.eqiad.wmnet onto db1274.eqiad.wmnet * 11:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Pool db1159.eqiad.wmnet in after cloning * 11:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1016.eqiad.wmnet with OS bookworm * 11:46 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti6004.drmrs.wmnet * 11:45 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti6001.drmrs.wmnet * 11:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti6001.drmrs.wmnet * 11:43 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324306{{!}}WikimediaAntiAbuse: Enable personal info for enwiki with no display (T431292)]] (duration: 10m 26s) * 11:42 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1015.eqiad.wmnet with OS bookworm * 11:39 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 11:38 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti6001.drmrs.wmnet * 11:34 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1324306{{!}}WikimediaAntiAbuse: Enable personal info for enwiki with no display (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:33 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2212: Security update * 11:33 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti6001.drmrs.wmnet * 11:32 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1324306{{!}}WikimediaAntiAbuse: Enable personal info for enwiki with no display (T431292)]] * 11:22 moritzm: failover ganeti master in drmrs01 to ganeti6003 * 11:20 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:20 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:18 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 22 hosts with reason: Cloning * 11:17 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti6003.drmrs.wmnet * 11:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti6003.drmrs.wmnet * 11:17 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1009.eqiad.wmnet with reason: host reimage * 11:17 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:16 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:14 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1016.eqiad.wmnet with reason: host reimage * 11:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti6003.drmrs.wmnet * 11:11 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1009.eqiad.wmnet with reason: host reimage * 11:10 jmm@cumin2003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti3005.esams.wmnet * 11:10 jmm@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host ganeti3005.esams.wmnet * 11:09 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 11:08 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1158: Depool db1158.eqiad.wmnet to then clone it to db1273.eqiad.wmnet - marostegui@cumin1003 * 11:07 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1016.eqiad.wmnet with reason: host reimage * 11:07 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1158: Depool db1158.eqiad.wmnet to then clone it to db1273.eqiad.wmnet - marostegui@cumin1003 * 11:07 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1158.eqiad.wmnet onto db1273.eqiad.wmnet * 11:06 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:05 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:05 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1159: Pool db1159.eqiad.wmnet in after cloning * 11:04 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 20 hosts with reason: Cloning * 11:02 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti6003.drmrs.wmnet * 11:00 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis testwiki in section s3 * 10:54 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1009.eqiad.wmnet with OS bookworm * 10:51 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1015.eqiad.wmnet with OS bookworm * 10:50 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1016.eqiad.wmnet with OS bookworm * 10:48 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2212: Security update * 10:45 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitize-wiki (exit_code=99) Managing sanitization for wikis testwiki in section s3 * 10:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1013.eqiad.wmnet with OS bookworm * 10:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1014.eqiad.wmnet with OS bookworm * 10:38 cwilliams@cumin1003: START - Cookbook sre.mysql.clone of db1159.eqiad.wmnet onto db1274.eqiad.wmnet * 10:33 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db1274.eqiad.wmnet * 10:33 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db1274.eqiad.wmnet * 10:31 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 396993 * 10:29 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 396993 * 10:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1159: Clone source for db1274 * 10:24 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1159: Clone source for db1274 * 10:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1013.eqiad.wmnet with reason: host reimage * 10:18 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1013.eqiad.wmnet with reason: host reimage * 10:13 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2212.codfw.wmnet with reason: Maintenance * 10:13 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 10:12 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 10:12 blake@deploy1003: Stopping before sync operations * 10:11 blake@deploy1003: Started scap sync-world: Non-deployment scap run to populate new release values for [[phab:T427668|T427668]] * 10:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2212 [[phab:T434644|T434644]]', diff saved to https://phabricator.wikimedia.org/P96003 and previous config saved to /var/cache/conftool/dbconfig/20260812-101053-cwilliams.json * 10:09 moritzm: powercycle ganeti3005 * 10:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2203 to s1 primary [[phab:T434644|T434644]]', diff saved to https://phabricator.wikimedia.org/P96002 and previous config saved to /var/cache/conftool/dbconfig/20260812-100849-cwilliams.json * 10:08 cezmunsta: Starting s1 codfw failover from db2212 to db2203 - [[phab:T434644|T434644]] * 10:03 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1013.eqiad.wmnet with OS bookworm * 10:02 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1013.eqiad.wmnet with OS bookworm * 10:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2203 with weight 0 [[phab:T434644|T434644]]', diff saved to https://phabricator.wikimedia.org/P96001 and previous config saved to /var/cache/conftool/dbconfig/20260812-100134-cwilliams.json * 10:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 32 hosts with reason: Primary switchover s1 [[phab:T434644|T434644]] * 09:53 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1014.eqiad.wmnet with reason: host reimage * 09:50 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 09:50 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti3005.esams.wmnet * 09:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1014.eqiad.wmnet with reason: host reimage * 09:41 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1278: Pool in x1 * 09:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-launcher1003.eqiad.wmnet with OS bookworm * 09:37 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti3005.esams.wmnet * 09:37 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1013.eqiad.wmnet with OS bookworm * 09:34 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-presto1013.eqiad.wmnet with OS bookworm * 09:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1014.eqiad.wmnet with OS bookworm * 09:29 moritzm: failover ganeti master in esams to ganeti3008 * 09:26 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti3006.esams.wmnet * 09:26 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti3006.esams.wmnet * 09:24 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1012.eqiad.wmnet with OS bookworm * 09:23 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1009.eqiad.wmnet with OS bookworm * 09:23 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 09:18 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti3006.esams.wmnet * 09:16 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti3006.esams.wmnet * 09:14 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: fix regexp escaping bug - oblivian@cumin1003" * 09:14 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: fix regexp escaping bug - oblivian@cumin1003 * 09:13 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: fix regexp escaping bug - oblivian@cumin1003 * 09:13 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: fix regexp escaping bug - oblivian@cumin1003" * 09:03 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-launcher1003.eqiad.wmnet with reason: host reimage * 08:58 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-launcher1003.eqiad.wmnet with reason: host reimage * 08:55 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1278: Pool in x1 * 08:55 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1278 to dbctl [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95996 and previous config saved to /var/cache/conftool/dbconfig/20260812-085521-marostegui.json * 08:51 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1012.eqiad.wmnet with reason: host reimage * 08:45 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis testwiki in section s3 * 08:43 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1009.eqiad.wmnet with OS bookworm * 08:42 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1012.eqiad.wmnet with reason: host reimage * 08:41 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-launcher1003.eqiad.wmnet with OS bookworm * 08:40 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1013.eqiad.wmnet with OS bookworm * 08:38 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms3', diff saved to https://phabricator.wikimedia.org/P95995 and previous config saved to /var/cache/conftool/dbconfig/20260812-083816-marostegui.json * 08:38 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-master1003.eqiad.wmnet with OS bookworm * 08:37 marostegui: Failover ms3 [[phab:T434288|T434288]] * 08:37 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1268 to dbctl [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95994 and previous config saved to /var/cache/conftool/dbconfig/20260812-083722-marostegui.json * 08:35 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-web1001.eqiad.wmnet with OS bookworm * 08:32 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db2252.codfw.wmnet,db[1153,1268].eqiad.wmnet with reason: Switching over ms3 * 08:28 marostegui@cumin1003: dbctl commit (dc=all): 'Depool ms3 [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95993 and previous config saved to /var/cache/conftool/dbconfig/20260812-082852-marostegui.json * 08:25 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1012.eqiad.wmnet with OS bookworm * 08:25 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1011.eqiad.wmnet with OS bookworm * 08:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-master1003.eqiad.wmnet with reason: host reimage * 08:07 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-master1003.eqiad.wmnet with reason: host reimage * 08:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-web1001.eqiad.wmnet with reason: host reimage * 07:58 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-web1001.eqiad.wmnet with reason: host reimage * 07:50 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1003.eqiad.wmnet with OS bookworm * 07:38 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1011.eqiad.wmnet with reason: host reimage * 07:38 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-master1003.eqiad.wmnet with OS bookworm * 07:35 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti3007.esams.wmnet * 07:35 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti3007.esams.wmnet * 07:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1011.eqiad.wmnet with reason: host reimage * 07:27 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti3007.esams.wmnet * 07:25 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti3007.esams.wmnet * 07:25 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti3008.esams.wmnet * 07:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti3008.esams.wmnet * 07:22 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-web1001.eqiad.wmnet with OS bookworm * 07:19 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1003.eqiad.wmnet with OS bookworm * 07:18 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-master1003.eqiad.wmnet * 07:18 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host an-master1003.eqiad.wmnet * 07:17 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1011.eqiad.wmnet with OS bookworm * 07:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 07:16 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti3008.esams.wmnet * 07:15 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1009.eqiad.wmnet with OS bookworm * 07:14 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host an-master1003.eqiad.wmnet * 07:13 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-master1003.eqiad.wmnet * 07:13 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-master1003.eqiad.wmnet * 07:12 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-master1003.eqiad.wmnet * 07:11 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti3008.esams.wmnet * 07:07 arnaudb@dns1006: END - running authdns-update * 07:07 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti5007.eqsin.wmnet * 07:07 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti5007.eqsin.wmnet * 07:05 arnaudb@dns1006: START - running authdns-update * 06:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti5007.eqsin.wmnet * 06:54 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti5007.eqsin.wmnet * 06:38 moritzm: failover ganeti master in eqsin to ganeti5004 * 06:36 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti5006.eqsin.wmnet * 06:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti5006.eqsin.wmnet * 06:28 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti5006.eqsin.wmnet * 06:23 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti5006.eqsin.wmnet * 06:20 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti5005.eqsin.wmnet * 06:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti5005.eqsin.wmnet * 06:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti5005.eqsin.wmnet * 06:06 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti5005.eqsin.wmnet * 06:03 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti5004.eqsin.wmnet * 06:03 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti5004.eqsin.wmnet * 05:55 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti5004.eqsin.wmnet * 05:53 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti5004.eqsin.wmnet * 04:40 ryankemper: [[phab:T434494|T434494]] reimaged `an-tool1008.eqiad.wmnet` to bookworm; yarn.wikimedia.org is back up * 04:16 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-tool1008.eqiad.wmnet with OS bookworm * 03:58 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-tool1008.eqiad.wmnet with reason: host reimage * 03:53 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-tool1008.eqiad.wmnet with reason: host reimage * 03:41 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-tool1008.eqiad.wmnet with OS bookworm * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 45s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 00:25 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324427{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]], [[gerrit:1324429{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0]], [[gerrit:1324428{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]] (duration: 07m 55s) * 00:21 kemayo@deploy1003: kemayo: Continuing with deployment * 00:19 kemayo@deploy1003: kemayo: Backport for [[gerrit:1324427{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]], [[gerrit:1324429{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0]], [[gerrit:1324428{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:17 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1324427{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]], [[gerrit:1324429{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0]], [[gerrit:1324428{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]] == 2026-08-11 == * 21:37 sbassett: Deployed security fix for [[phab:T434521|T434521]] (wmf.15) * 21:29 sbassett: Deployed security fix for [[phab:T434521|T434521]] (wmf.14) * 21:19 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324370{{!}}Phase 4 of legal footer deployment (T432796)]], [[gerrit:1319804{{!}}Disable wgMFCustomSiteModules on English Wikipedia (T375538)]] (duration: 15m 26s) * 21:15 jdlrobson@deploy1003: jdlrobson: Continuing with deployment * 21:06 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1324370{{!}}Phase 4 of legal footer deployment (T432796)]], [[gerrit:1319804{{!}}Disable wgMFCustomSiteModules on English Wikipedia (T375538)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:03 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1324370{{!}}Phase 4 of legal footer deployment (T432796)]], [[gerrit:1319804{{!}}Disable wgMFCustomSiteModules on English Wikipedia (T375538)]] * 20:59 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 20:50 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324384{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]], [[gerrit:1324385{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]] (duration: 06m 58s) * 20:46 kemayo@deploy1003: kemayo: Continuing with deployment * 20:45 kemayo@deploy1003: kemayo: Backport for [[gerrit:1324384{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]], [[gerrit:1324385{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:43 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1324384{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]], [[gerrit:1324385{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]] * 20:42 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 20:42 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324386{{!}}build: Updating js-yaml to 3.15.1, 4.3.1]] (duration: 07m 36s) * 20:38 kemayo@deploy1003: kemayo: Continuing with deployment * 20:37 jhancock@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 20:37 kemayo@deploy1003: kemayo: Backport for [[gerrit:1324386{{!}}build: Updating js-yaml to 3.15.1, 4.3.1]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:35 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1324386{{!}}build: Updating js-yaml to 3.15.1, 4.3.1]] * 20:18 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 20:15 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 20:15 jhancock@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin1003" * 20:14 jhancock@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin1003" * 19:59 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 19:54 jhancock@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 19:10 brennen@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] (duration: 06m 41s) * 19:04 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns6001.wikimedia.org * 19:04 sukhe@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns6001.wikimedia.org * 19:04 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns5003.wikimedia.org * 19:04 sukhe@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns5003.wikimedia.org * 19:03 brennen@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 18:59 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns5003.wikimedia.org with OS trixie * 18:55 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns6001.wikimedia.org with OS trixie * 18:19 brett@cumin2002: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on P<nowiki>{</nowiki>cp7009.magru.wmnet<nowiki>}</nowiki> and A:cp - 9.2.15 Upgrade () * 18:14 brett@cumin2002: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on P<nowiki>{</nowiki>cp7009.magru.wmnet<nowiki>}</nowiki> and A:cp - 9.2.15 Upgrade () * 18:13 brennen@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 18:12 brett@cumin2002: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 9.2.15 Upgrade () * 18:09 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns5003.wikimedia.org with reason: host reimage * 18:06 brett@cumin2002: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 9.2.15 Upgrade () * 18:06 brennen: 1.47.0-wmf.15 train status ([[phab:T430834|T430834]]) - no current blockers, rolling to group0 * 18:05 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns5003.wikimedia.org with reason: host reimage * 18:05 brett: import trafficserver-9.2.15~deb13+wmf1 into trixie-wikimedia ([[phab:T434478|T434478]]) * 18:01 ladsgroup@cumin1003: END (PASS) - Cookbook sre.mysql.sanitarium_restart (exit_code=0) * 17:58 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns6001.wikimedia.org with reason: host reimage * 17:53 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324369{{!}}Enable desktop lazy loading on group0 (T148047)]] (duration: 07m 31s) * 17:52 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns6001.wikimedia.org with reason: host reimage * 17:49 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 17:49 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitarium_restart (exit_code=99) * 17:49 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 17:49 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7001.magru.wmnet * 17:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti7001.magru.wmnet * 17:48 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 17:47 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1324369{{!}}Enable desktop lazy loading on group0 (T148047)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:45 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1324369{{!}}Enable desktop lazy loading on group0 (T148047)]] * 17:39 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti7001.magru.wmnet * 17:36 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns5003.wikimedia.org with OS trixie * 17:34 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns6001.wikimedia.org with OS trixie * 17:31 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324361{{!}}Move FR config from IS.php to a dedicated file]], [[gerrit:1324363{{!}}Remove $wmg = $wg hacks in CentralAuth (T119117)]] (duration: 12m 23s) * 17:26 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 17:23 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1324361{{!}}Move FR config from IS.php to a dedicated file]], [[gerrit:1324363{{!}}Remove $wmg = $wg hacks in CentralAuth (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:19 sukhe: sudo cumin "A:cp-magru" "run-puppet-agent --enable 'merging CR 1324355'" * 17:18 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1324361{{!}}Move FR config from IS.php to a dedicated file]], [[gerrit:1324363{{!}}Remove $wmg = $wg hacks in CentralAuth (T119117)]] * 17:11 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-master1004.eqiad.wmnet with OS bookworm * 17:11 sukhe: sukhe@cp7005:~$ sudo puppet agent -tv * 17:02 sukhe: sudo cumin "A:cp-magru" "disable-puppet 'merging CR 1324355'" * 16:54 sukhe@dns1004: END - running authdns-update * 16:53 sukhe@dns1004: START - running authdns-update * 16:53 sukhe@dns1004: FAIL - running authdns-update * 16:51 sukhe@dns1004: START - running authdns-update * 16:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-master1004.eqiad.wmnet with reason: host reimage * 16:44 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-master1004.eqiad.wmnet with reason: host reimage * 16:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1157.eqiad.wmnet onto db1272.eqiad.wmnet * 16:40 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1157: Pool db1157.eqiad.wmnet in after cloning * 16:38 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 16:31 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324356{{!}}InitialiseSettings: Fix wgOATHAuthEnforce2FAForAll]] (duration: 06m 52s) * 16:30 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 16:28 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 16:27 reedy@deploy1003: reedy: Continuing with deployment * 16:26 reedy@deploy1003: reedy: Backport for [[gerrit:1324356{{!}}InitialiseSettings: Fix wgOATHAuthEnforce2FAForAll]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:24 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324356{{!}}InitialiseSettings: Fix wgOATHAuthEnforce2FAForAll]] * 16:13 sukhe: restart ntpsec.serviceon dns7001 * 16:09 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324335{{!}}InitialiseSettings: Enable 2FA enforcement on various private wikis (T428103)]] (duration: 06m 40s) * 16:08 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 16:06 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2204: Security update * 16:04 reedy@deploy1003: reedy: Continuing with deployment * 16:04 reedy@deploy1003: reedy: Backport for [[gerrit:1324335{{!}}InitialiseSettings: Enable 2FA enforcement on various private wikis (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:02 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7001.magru.wmnet * 16:02 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324335{{!}}InitialiseSettings: Enable 2FA enforcement on various private wikis (T428103)]] * 16:01 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 15:55 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1157: Pool db1157.eqiad.wmnet in after cloning * 15:54 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 15:54 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 15:41 dancy@deploy1003: Finished scap sync-world: Testing (duration: 06m 28s) * 15:40 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-master1004.eqiad.wmnet with OS bookworm * 15:40 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 15:35 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti4008.ulsfo.wmnet * 15:35 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti4008.ulsfo.wmnet * 15:34 dancy@deploy1003: Started scap sync-world: Testing * 15:34 dancy@deploy1003: Installation of scap version "4.279.0" completed for 3 hosts * 15:34 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-master1004.eqiad.wmnet with OS bookworm * 15:34 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 15:33 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-master1004.eqiad.wmnet with OS bookworm * 15:32 dancy@deploy1003: Installing scap version "4.279.0" for 3 host(s) * 15:32 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324339{{!}}Add /w/deployment-info.php entrypoint]] (duration: 07m 25s) * 15:30 moritzm: failover ganeti master in magru to ganeti7004 * 15:29 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti4008.ulsfo.wmnet * 15:28 dancy@deploy1003: dancy: Continuing with deployment * 15:28 tappof: remove 2026-05 swift log archives from centrallog to free some space ([[phab:T434502|T434502]]) * 15:27 dancy@deploy1003: dancy: Backport for [[gerrit:1324339{{!}}Add /w/deployment-info.php entrypoint]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:25 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324339{{!}}Add /w/deployment-info.php entrypoint]] * 15:20 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2204: Security update * 15:18 dancy@deploy1003: Installation of scap version "4.278.0" completed for 3 hosts * 15:18 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7004.magru.wmnet * 15:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti7004.magru.wmnet * 15:16 dancy@deploy1003: Installing scap version "4.278.0" for 3 host(s) * 15:14 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2204.codfw.wmnet with reason: Maintenance * 15:12 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 15:11 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-master1004.eqiad.wmnet with OS bookworm * 15:11 hashar: Restarting CI Jenkins on contint1003 due to Java upgrade. * 15:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2204 [[phab:T434565|T434565]]', diff saved to https://phabricator.wikimedia.org/P95984 and previous config saved to /var/cache/conftool/dbconfig/20260811-151126-cwilliams.json * 15:10 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti4008.ulsfo.wmnet * 15:10 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti7004.magru.wmnet * 15:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2207 to s2 primary [[phab:T434565|T434565]]', diff saved to https://phabricator.wikimedia.org/P95983 and previous config saved to /var/cache/conftool/dbconfig/20260811-150905-cwilliams.json * 15:08 cezmunsta: Starting s2 codfw failover from db2204 to db2207 - [[phab:T434565|T434565]] * 15:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2207 with weight 0 [[phab:T434565|T434565]]', diff saved to https://phabricator.wikimedia.org/P95982 and previous config saved to /var/cache/conftool/dbconfig/20260811-150402-cwilliams.json * 15:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s2 [[phab:T434565|T434565]] * 14:55 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1010.eqiad.wmnet with OS bookworm * 14:49 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns4003.wikimedia.org with OS trixie * 14:47 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1010.eqiad.wmnet with OS bookworm * 14:47 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7001.wikimedia.org with OS trixie * 14:44 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-presto1010.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:41 btullis@cumin1003: START - Cookbook sre.hosts.provision for host an-presto1010.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:40 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-presto1010.eqiad.wmnet with OS bookworm * 14:39 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 14:39 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-presto1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:36 btullis@cumin1003: START - Cookbook sre.hosts.provision for host an-presto1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:32 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1009.eqiad.wmnet with OS bookworm * 14:32 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 14:31 cwilliams@cumin1003: START - Cookbook sre.mysql.clone of db1157.eqiad.wmnet onto db1272.eqiad.wmnet * 14:30 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7004.magru.wmnet * 14:28 moritzm: failover ganeti master in ulsfo to ganeti4005 * 14:23 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7003.magru.wmnet * 14:23 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti7003.magru.wmnet * 14:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-master1004.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:22 btullis@cumin1003: START - Cookbook sre.hosts.provision for host an-master1004.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:21 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-master1004.eqiad.wmnet with OS bookworm * 14:19 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti4007.ulsfo.wmnet * 14:19 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti4007.ulsfo.wmnet * 14:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1010.eqiad.wmnet with OS bookworm * 14:17 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1008.eqiad.wmnet with OS bookworm * 14:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti7003.magru.wmnet * 14:12 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7003.magru.wmnet * 14:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti4007.ulsfo.wmnet * 14:11 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7002.magru.wmnet * 14:11 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti7002.magru.wmnet * 14:09 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324318{{!}}Revert "wmf-config/ProductionServices: set URL for urldownloader to service record" (T429175)]] (duration: 06m 46s) * 14:09 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7001.wikimedia.org with reason: host reimage * 14:06 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti4007.ulsfo.wmnet * 14:05 kharlan@deploy1003: kharlan: Continuing with deployment * 14:05 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti4006.ulsfo.wmnet * 14:04 jayme: updated calico to v3.30.7 on staging-eqiad - [[phab:T427400|T427400]] * 14:04 kharlan@deploy1003: kharlan: Backport for [[gerrit:1324318{{!}}Revert "wmf-config/ProductionServices: set URL for urldownloader to service record" (T429175)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti4006.ulsfo.wmnet * 14:03 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns4003.wikimedia.org with reason: host reimage * 14:03 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7001.wikimedia.org with reason: host reimage * 14:02 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti7002.magru.wmnet * 14:02 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1324318{{!}}Revert "wmf-config/ProductionServices: set URL for urldownloader to service record" (T429175)]] * 14:02 btullis@dns1004: FAIL - running authdns-update * 14:00 btullis@dns1004: START - running authdns-update * 13:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1008.eqiad.wmnet with reason: host reimage * 13:59 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'. * 13:59 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1313985{{!}}wmf-config/ProductionServices: set URL for urldownloader to service record (T429175)]] (duration: 25m 06s) * 13:58 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7002.magru.wmnet * 13:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti4006.ulsfo.wmnet * 13:57 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns4003.wikimedia.org with reason: host reimage * 13:56 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'. * 13:56 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7001.magru.wmnet * 13:56 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1008.eqiad.wmnet with reason: host reimage * 13:55 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'. * 13:55 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'. * 13:55 kharlan@deploy1003: kharlan, sukhe: Continuing with deployment * 13:53 marostegui: Failover ms2 [[phab:T434288|T434288]] * 13:52 marostegui: Failover ms1 [[phab:T434288|T434288]] * 13:52 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7001.magru.wmnet * 13:51 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti4006.ulsfo.wmnet * 13:48 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti4005.ulsfo.wmnet * 13:48 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti4005.ulsfo.wmnet * 13:44 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti4005.ulsfo.wmnet * 13:40 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1008.eqiad.wmnet with OS bookworm * 13:39 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns4003.wikimedia.org with OS trixie * 13:38 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns7001.wikimedia.org with OS trixie * 13:38 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1008.eqiad.wmnet with OS bookworm * 13:37 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti4005.ulsfo.wmnet * 13:36 kharlan@deploy1003: kharlan, sukhe: Backport for [[gerrit:1313985{{!}}wmf-config/ProductionServices: set URL for urldownloader to service record (T429175)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:34 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1313985{{!}}wmf-config/ProductionServices: set URL for urldownloader to service record (T429175)]] * 13:29 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2034.codfw.wmnet * 13:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2034.codfw.wmnet * 13:25 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.clone (exit_code=99) of db1157.eqiad.wmnet onto db1272.eqiad.wmnet * 13:25 cwilliams@cumin1003: START - Cookbook sre.mysql.clone of db1157.eqiad.wmnet onto db1272.eqiad.wmnet * 13:21 urbanecm@deploy1003: mwscript-k8s job started: namespaceDupes.php --wiki=frwiktionary --fix # [[phab:T415716|T415716]] * 13:21 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2034.codfw.wmnet * 13:20 urbanecm@deploy1003: mwscript-k8s job started: namespaceDupes.php --wiki=frwiktionary # [[phab:T415716|T415716]] * 13:19 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1323350{{!}}[tgwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T415307)]], [[gerrit:1322961{{!}}[slwiki] Revert temporary logo for Wikipedia 25 (Vector legacy + Vector 2022) (T414265)]], [[gerrit:1323827{{!}}[itwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T414320)]] (duration: 08m 00s) * 13:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-coord1003.eqiad.wmnet with OS bookworm * 13:15 urbanecm@deploy1003: urbanecm, superpes: Continuing with deployment * 13:13 urbanecm@deploy1003: urbanecm, superpes: Backport for [[gerrit:1323350{{!}}[tgwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T415307)]], [[gerrit:1322961{{!}}[slwiki] Revert temporary logo for Wikipedia 25 (Vector legacy + Vector 2022) (T414265)]], [[gerrit:1323827{{!}}[itwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T414320)]] synced to the testservers (see https://wiki * 13:11 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1323350{{!}}[tgwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T415307)]], [[gerrit:1322961{{!}}[slwiki] Revert temporary logo for Wikipedia 25 (Vector legacy + Vector 2022) (T414265)]], [[gerrit:1323827{{!}}[itwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T414320)]] * 13:11 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1323779{{!}}[ukwiki] Remove reviewer usergroup (T434252)]], [[gerrit:1323312{{!}}[frwiktionary] Add new Schème namespace and its talk (T415716)]] (duration: 06m 49s) * 13:10 marostegui@dns1004: END - running authdns-update * 13:08 marostegui@dns1004: START - running authdns-update * 13:07 marostegui@cumin1003: dbctl commit (dc=all): 'Repool ms2 [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95980 and previous config saved to /var/cache/conftool/dbconfig/20260811-130725-marostegui.json * 13:06 urbanecm@deploy1003: urbanecm, superpes: Continuing with deployment * 13:06 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1266 to dbctl [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95979 and previous config saved to /var/cache/conftool/dbconfig/20260811-130627-marostegui.json * 13:06 urbanecm@deploy1003: urbanecm, superpes: Backport for [[gerrit:1323779{{!}}[ukwiki] Remove reviewer usergroup (T434252)]], [[gerrit:1323312{{!}}[frwiktionary] Add new Schème namespace and its talk (T415716)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:04 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1323779{{!}}[ukwiki] Remove reviewer usergroup (T434252)]], [[gerrit:1323312{{!}}[frwiktionary] Add new Schème namespace and its talk (T415716)]] * 12:59 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db2253.codfw.wmnet,db[1151,1266].eqiad.wmnet with reason: Switching over ms2 * 12:54 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1157: Using as clone source * 12:53 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1157: Using as clone source * 12:51 marostegui@cumin1003: dbctl commit (dc=all): 'Depool ms2 [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95977 and previous config saved to /var/cache/conftool/dbconfig/20260811-125129-marostegui.json * 12:47 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 12:46 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 12:45 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 12:44 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'recommendation-api-ng' for release 'main' . * 12:44 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'recommendation-api-ng' for release 'main' . * 12:43 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'recommendation-api-ng' for release 'main' . * 12:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-coord1003.eqiad.wmnet with reason: host reimage * 12:43 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'ores-legacy' for release 'main' . * 12:42 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'ores-legacy' for release 'main' . * 12:42 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2165: Security update * 12:40 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-coord1003.eqiad.wmnet with reason: host reimage * 12:39 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'ores-legacy' for release 'main' . * 12:38 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' . * 12:38 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' . * 12:37 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' . * 12:34 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2034.codfw.wmnet * 12:30 jmm@cumin2002: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti-test2001.codfw.wmnet * 12:30 jmm@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host ganeti-test2001.codfw.wmnet * 12:25 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1179.eqiad.wmnet onto db1278.eqiad.wmnet * 12:25 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1179: Pool db1179.eqiad.wmnet in after cloning * 12:23 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-coord1003.eqiad.wmnet with OS bookworm * 12:22 moritzm: failover ganeti master in codfw/routed to ganeti2033 * 12:22 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2033.codfw.wmnet * 12:22 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2033.codfw.wmnet * 12:19 jmm@cumin2002: START - Cookbook sre.hosts.reboot-single for host ganeti-test2001.codfw.wmnet * 12:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 12:18 jmm@cumin2002: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti-test2001.codfw.wmnet * 12:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1008.eqiad.wmnet with OS bookworm * 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2033.codfw.wmnet * 12:07 moritzm: failover ganeti master in ganeti/test to ganeti-test2003 * 12:04 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 12:03 jmm@cumin2003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti4005.ulsfo.wmnet * 12:03 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti4005.ulsfo.wmnet * 12:00 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324283{{!}}Use maximum compression level in SqlBlobStore and SqlBagOStuff (T428377)]] (duration: 11m 37s) * 11:57 jmm@cumin2002: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti-test2002.codfw.wmnet * 11:57 jmm@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti-test2002.codfw.wmnet * 11:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2165: Security update * 11:54 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 11:52 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1324283{{!}}Use maximum compression level in SqlBlobStore and SqlBagOStuff (T428377)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:51 jmm@cumin2002: START - Cookbook sre.hosts.reboot-single for host ganeti-test2002.codfw.wmnet * 11:50 jmm@cumin2002: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti-test2002.codfw.wmnet * 11:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2165.codfw.wmnet with reason: Maintenance * 11:48 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@050d19e] (releasing): [[phab:T434186|T434186]] (duration: 01m 14s) * 11:48 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1324283{{!}}Use maximum compression level in SqlBlobStore and SqlBagOStuff (T428377)]] * 11:47 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@050d19e] (releasing): [[phab:T434186|T434186]] * 11:44 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@050d19e] (releasing): test jenkins deploy for [[phab:T434186|T434186]] (duration: 01m 08s) * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2165 [[phab:T434514|T434514]]', diff saved to https://phabricator.wikimedia.org/P95969 and previous config saved to /var/cache/conftool/dbconfig/20260811-114352-cwilliams.json * 11:43 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@050d19e] (releasing): test jenkins deploy for [[phab:T434186|T434186]] * 11:42 jmm@cumin2003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti-test2003.codfw.wmnet * 11:42 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti-test2003.codfw.wmnet * 11:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2161 to s8 primary [[phab:T434514|T434514]]', diff saved to https://phabricator.wikimedia.org/P95968 and previous config saved to /var/cache/conftool/dbconfig/20260811-114136-cwilliams.json * 11:40 cezmunsta: Starting s8 codfw failover from db2165 to db2161 - [[phab:T434514|T434514]] * 11:40 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1179: Pool db1179.eqiad.wmnet in after cloning * 11:36 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti-test2003.codfw.wmnet * 11:36 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti-test2003.codfw.wmnet * 11:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2161 with weight 0 [[phab:T434514|T434514]]', diff saved to https://phabricator.wikimedia.org/P95966 and previous config saved to /var/cache/conftool/dbconfig/20260811-113449-cwilliams.json * 11:34 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 25 hosts with reason: Primary switchover s8 [[phab:T434514|T434514]] * 11:29 moritzm: installing Python 3.11 security updates * 11:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-presto1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 11:26 btullis@cumin1003: START - Cookbook sre.hosts.provision for host an-presto1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 11:23 btullis@dns1004: END - running authdns-update * 11:21 btullis@dns1004: START - running authdns-update * 11:20 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-presto1008.eqiad.wmnet with OS bookworm * 11:20 moritzm: installing curl security updates * 11:11 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1007.eqiad.wmnet with OS bookworm * 10:45 tappof: bump space for prometheus k8s-dse in codfw * 10:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1007.eqiad.wmnet with reason: host reimage * 10:38 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1007.eqiad.wmnet with reason: host reimage * 10:37 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-coord1004.eqiad.wmnet with OS bookworm * 10:35 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1008.eqiad.wmnet with OS bookworm * 10:34 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1179: Depool db1179.eqiad.wmnet to then clone it to db1278.eqiad.wmnet - marostegui@cumin1003 * 10:34 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1008.eqiad.wmnet with OS bookworm * 10:25 fceratto@cumin1003: dbctl commit (dc=all): 'Remove db1177 [[phab:T433474|T433474]]', diff saved to https://phabricator.wikimedia.org/P95964 and previous config saved to /var/cache/conftool/dbconfig/20260811-102527-fceratto.json * 10:22 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1008.eqiad.wmnet with OS bookworm * 10:22 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1007.eqiad.wmnet with OS bookworm * 10:21 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1006.eqiad.wmnet with OS bookworm * 10:20 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 10:18 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1179: Depool db1179.eqiad.wmnet to then clone it to db1278.eqiad.wmnet - marostegui@cumin1003 * 10:18 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1179.eqiad.wmnet onto db1278.eqiad.wmnet * 10:17 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 10:17 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 10:14 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 10:09 blake@deploy1003: Stopping before sync operations * 10:09 blake@deploy1003: Started scap sync-world: Non-deployment run to populate release values for [[phab:T427668|T427668]] * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 10:04 fceratto@cumin1003: Removing db1177 from zarcillo [[phab:T433474|T433474]] * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1177.eqiad.wmnet * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1177.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:03 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1177.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:00 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1006.eqiad.wmnet with reason: host reimage * 09:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-coord1004.eqiad.wmnet with reason: host reimage * 09:57 marostegui: Failover m1 from db1164 to db1213 - [[phab:T434493|T434493]] * 09:57 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1006.eqiad.wmnet with reason: host reimage * 09:55 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:54 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2232].codfw.wmnet,db[1164,1213,1217].eqiad.wmnet with reason: Primary switchover m1 [[phab:T434493|T434493]] * 09:52 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-coord1004.eqiad.wmnet with reason: host reimage * 09:49 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1213.eqiad.wmnet with OS trixie * 09:49 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1177.eqiad.wmnet * 09:41 moritzm: installing Linux 6.12.101 on Trixie hosts * 09:40 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1006.eqiad.wmnet with OS bookworm * 09:35 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-coord1004.eqiad.wmnet with OS bookworm * 09:28 moritzm: installing node-tar security updates * 09:27 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1213.eqiad.wmnet with reason: host reimage * 09:22 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1213.eqiad.wmnet with reason: host reimage * 09:09 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1177: Decommission * 09:08 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db1177: Decommission * 09:08 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 09:08 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.decommission (exit_code=99) * 09:06 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1213.eqiad.wmnet with OS trixie * 09:06 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 09:05 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1213.eqiad.wmnet with reason: Reimage * 08:53 marostegui@dns1004: END - running authdns-update * 08:51 marostegui@dns1004: START - running authdns-update * 08:48 marostegui: Switchover ms1 master in eqiad [[phab:T434288|T434288]] * 08:48 marostegui@cumin1003: dbctl commit (dc=all): 'Repool ms1 [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95962 and previous config saved to /var/cache/conftool/dbconfig/20260811-084804-marostegui.json * 08:40 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1267 to dbctl [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95961 and previous config saved to /var/cache/conftool/dbconfig/20260811-084054-marostegui.json * 08:29 marostegui: Failover m1 from db1213 to db1164 - [[phab:T434043|T434043]] * 08:25 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2232].codfw.wmnet,db[1164,1213,1217].eqiad.wmnet with reason: Primary switchover m1 [[phab:T434043|T434043]] * 08:22 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db2251.codfw.wmnet,db[1152,1267].eqiad.wmnet with reason: Switching over ms1 * 08:22 marostegui@cumin1003: dbctl commit (dc=all): 'Depool ms1 [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95960 and previous config saved to /var/cache/conftool/dbconfig/20260811-082201-marostegui.json * 08:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: Switching over ms1 * 08:20 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.parsercache (exit_code=99) * 08:20 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 08:20 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1152: Switching over ms1 * 08:19 slyngshede@dns1004: END - running authdns-update * 08:18 moritzm: installing openjdk-21 security updates * 08:17 slyngshede@dns1004: START - running authdns-update * 08:16 moritzm: imported jenkins 2.568.2 to thirdparty/jenkins for trixie-wikimedia * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.12 (duration: 02m 26s) * 03:36 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] (duration: 33m 33s) * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 35s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-10 == * 14:54 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1323973{{!}}mmv.bootstrap: Fix getUrlParam to account for TIFF lossy/lossless param (T434333)]] (duration: 11m 24s) * 14:50 krinkle@deploy1003: krinkle: Continuing with deployment * 14:45 krinkle@deploy1003: krinkle: Backport for [[gerrit:1323973{{!}}mmv.bootstrap: Fix getUrlParam to account for TIFF lossy/lossless param (T434333)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:43 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1323973{{!}}mmv.bootstrap: Fix getUrlParam to account for TIFF lossy/lossless param (T434333)]] * 14:07 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1323967{{!}}updateIsActiveFlagForMentees: Commit the final partial batch (T432959)]] (duration: 10m 33s) * 13:56 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1323967{{!}}updateIsActiveFlagForMentees: Commit the final partial batch (T432959)]] * 13:45 dani@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply * 13:45 dani@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply * 13:45 dani@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply * 13:45 dani@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply * 13:45 dani@deploy1003: helmfile [staging] DONE helmfile.d/services/miscweb: apply * 13:44 dani@deploy1003: helmfile [staging] START helmfile.d/services/miscweb: apply * 13:38 wmde-fisch@deploy1003: Finished scap sync-world: Backport for [[gerrit:1323939{{!}}Enable sub-references on more group2 wikis (batch3) (T432731)]] (duration: 33m 21s) * 13:25 wmde-fisch@deploy1003: wmde-fisch: Continuing with deployment * 13:22 wmde-fisch@deploy1003: wmde-fisch: Backport for [[gerrit:1323939{{!}}Enable sub-references on more group2 wikis (batch3) (T432731)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:05 wmde-fisch@deploy1003: Started scap sync-world: Backport for [[gerrit:1323939{{!}}Enable sub-references on more group2 wikis (batch3) (T432731)]] * 07:57 hashar@deploy1003: Finished deploy [integration/docroot@7772132]: update build dependencies (duration: 00m 13s) * 07:57 hashar@deploy1003: Started deploy [integration/docroot@7772132]: update build dependencies * 07:35 _joe_: restarting squid on urldownloader1006 * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 48s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-09 == * 16:01 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:01 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:01 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:00 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 36s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-08 == * 05:31 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9] (wcqs): [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) (duration: 02m 36s) * 05:28 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9] (wcqs): [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) * 04:56 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 04:55 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 04:47 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) (duration: 19m 22s) * 04:28 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) * 04:19 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) (duration: 00m 06s) * 04:18 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) * 04:17 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) (duration: 00m 28s) * 04:16 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) * 03:52 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 03:52 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 34s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-07 == * 23:30 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:29 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 22:45 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 22:43 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 22:41 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 22:41 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 20:54 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 20:32 andrewbogott: restarting puppetserver service on puppetserver* for [[phab:T434339|T434339]] * 19:52 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:45 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 19:32 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:25 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:22 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 19:21 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 18:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:41 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 18:35 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 18:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 18:22 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 18:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 18:16 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 18:12 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:09 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:08 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:07 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:04 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:01 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:00 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:00 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 17:59 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 17:25 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 17:14 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 17:13 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 17:13 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 17:13 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:54 maryum: Deployed security fix for [[phab:T434278|T434278]] * 16:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 16:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 16:27 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --verbose --scoreLessThan=0.7 --exceptDatasetChecksums=[[phab:T434319|T434319]]-enwiki-models.txt # [[phab:T434319|T434319]] * 16:06 cdobbins@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp5022.eqsin.wmnet with OS trixie * 15:13 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 14:19 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 14:17 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 13:50 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1156.eqiad.wmnet onto db1271.eqiad.wmnet * 13:50 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1271: Pool db1271.eqiad.wmnet in after cloning * 13:02 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1271: Pool db1271.eqiad.wmnet in after cloning * 12:19 jayme: updated calico to v3.30.7 on staging-codfw - [[phab:T427400|T427400]] * 12:09 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 12:06 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 12:06 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 12:05 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 12:02 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1156: Pool db1156.eqiad.wmnet in after cloning * 11:38 bjensen: sudo -i reprepro -C main include trixie-wikimedia $<nowiki>{</nowiki>HOME<nowiki>}</nowiki>/httpbb/trixie/httpbb_$<nowiki>{</nowiki>VERSION?<nowiki>}</nowiki>-1+deb13u1_amd64.changes #[[phab:T434052|T434052]] * 11:35 bjensen: sudo -i reprepro -C main include bookworm-wikimedia $<nowiki>{</nowiki>HOME<nowiki>}</nowiki>/httpbb/bookworm/httpbb_$<nowiki>{</nowiki>VERSION?<nowiki>}</nowiki>-1_amd64.changes #[[phab:T434052|T434052]] * 11:30 marostegui@cumin1003: dbctl commit (dc=all): 'Adding db1271 to dbctl', diff saved to https://phabricator.wikimedia.org/P95945 and previous config saved to /var/cache/conftool/dbconfig/20260807-113006-marostegui.json * 11:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1156: Pool db1156.eqiad.wmnet in after cloning * 10:23 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 10:22 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 10:22 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 10:21 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 10:20 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 10:20 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 10:19 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 10:18 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 10:06 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on 21 hosts with reason: cloning * 10:01 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1156: Depool db1156.eqiad.wmnet to then clone it to db1271.eqiad.wmnet - marostegui@cumin1003 * 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1156: Depool db1156.eqiad.wmnet to then clone it to db1271.eqiad.wmnet - marostegui@cumin1003 * 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1156.eqiad.wmnet onto db1271.eqiad.wmnet * 09:15 jynus: started stress testing db1245 dbs [[phab:T431115|T431115]] * 08:19 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:18 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:16 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:14 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:13 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:10 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:06 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:05 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:00 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 10 days, 0:00:00 on ml-serve1015.eqiad.wmnet with reason: Downtime to get full picture of current BIOS settings beyond what Redfish shows * 08:00 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 07:54 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 07:54 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 07:53 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:52 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:51 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:50 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:49 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:48 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:47 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:45 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:45 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:41 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:38 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:37 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 06:35 jayme: updated istio to 1.29.4 on wikikube eqiad - [[phab:T427401|T427401]] * 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1178.eqiad.wmnet * 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1178.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 06:06 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1178.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 05:55 marostegui@cumin1003: START - Cookbook sre.dns.netbox * 05:49 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1178.eqiad.wmnet * 05:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 05:46 marostegui@cumin1003: Removing db1178 from zarcillo [[phab:T433471|T433471]] * 05:45 marostegui@cumin1003: START - Cookbook sre.mysql.decommission * 02:42 denisse: Extended volume on prometheus2008 for the disk space alert as per https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_host_running_out_of_space * 02:37 denisse: Extended volume on prometheus2007 tor the disk space alert as per https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_host_running_out_of_space * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 56s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-06 == * 21:39 maryum: Deploy security patch for [[phab:T433070|T433070]] * 21:29 maryum: Deploy security patch for [[phab:T434189|T434189]] * 20:48 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] (duration: 08m 12s) * 20:44 aude@deploy1003: lmora, aude, anzx: Continuing with deployment * 20:41 aude@deploy1003: lmora, aude, anzx: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be * 20:41 ebernhardson@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:41 ebernhardson@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 20:40 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] * 20:37 ebernhardson@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:37 ebernhardson@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 20:32 ebernhardson@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:32 ebernhardson@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 20:31 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] (duration: 06m 41s) * 20:27 cjming@deploy1003: cjming, ebernhardson, chlod: Continuing with deployment * 20:26 cjming@deploy1003: cjming, ebernhardson, chlod: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:24 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] * 20:18 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] (duration: 09m 22s) * 20:14 cjming@deploy1003: cjming, tsev: Continuing with deployment * 20:11 cjming@deploy1003: cjming, tsev: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:09 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] * 19:41 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply * 19:40 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply * 19:31 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 19:31 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 19:00 cdobbins@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cp5022.eqsin.wmnet with OS trixie * 18:25 ladsgroup@deploy1003: Finished scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) (duration: 06m 08s) * 18:19 ladsgroup@deploy1003: Started scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) * 18:18 ladsgroup@deploy1003: Stopping before sync operations * 18:17 ladsgroup@deploy1003: Started scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) * 17:55 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 16:50 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 16:35 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1001.eqiad.wmnet with OS bookworm * 16:19 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1002.eqiad.wmnet with reason: host reimage * 16:16 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1002.eqiad.wmnet with reason: host reimage * 16:05 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1001.eqiad.wmnet with reason: host reimage * 16:00 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1001.eqiad.wmnet with reason: host reimage * 15:57 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 15:43 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm * 15:29 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1001.eqiad.wmnet with OS bookworm * 15:29 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:58 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm * 14:57 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-drmrs ([[phab:T428495|T428495]]) * 14:55 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-drmrs ([[phab:T428495|T428495]]) * 14:55 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-ui1001.eqiad.wmnet with OS bookworm * 14:54 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-presto1001.eqiad.wmnet with OS bookworm * 14:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-magru ([[phab:T428495|T428495]]) * 14:49 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-magru ([[phab:T428495|T428495]]) * 14:48 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1001.eqiad.wmnet with OS bookworm * 14:46 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-esams ([[phab:T428495|T428495]]) * 14:44 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-esams ([[phab:T428495|T428495]]) * 14:43 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:42 brouberol@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:42 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:42 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 14:40 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 14:40 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:38 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-ui1001.eqiad.wmnet with reason: host reimage * 14:34 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-presto1001.eqiad.wmnet with reason: host reimage * 14:28 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-ui1001.eqiad.wmnet with reason: host reimage * 14:27 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-presto1001.eqiad.wmnet with reason: host reimage * 14:23 sukhe: sudo cumin -b2 'A:cp-text' "run-puppet-agent --enable 'merging CR 1290731'": [[phab:T425441|T425441]] * 14:18 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo for hosts in the wikimedia.org domain - [[phab:T428495|T428495]] * 14:16 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-presto1001.eqiad.wmnet with OS bookworm * 14:14 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-ui1001.eqiad.wmnet with OS bookworm * 14:12 sukhe: sudo cumin 'A:cp-text' "disable-puppet 'merging CR 1290731'": [[phab:T425441|T425441]] * 14:11 swfrench-wmf: restarted navtiming on webperf1003 - [[phab:T428495|T428495]] * 14:04 swfrench-wmf: begin rolling restart of confd in drmrs, eqiad, esams, magru - [[phab:T428495|T428495]] * 14:04 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm * 14:04 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:02 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-client1002.eqiad.wmnet with OS bookworm * 13:58 swfrench-wmf: authdns update to direct eqiad-associated etcd clients back to eqiad - [[phab:T428495|T428495]] * 13:58 swfrench@dns1004: END - running authdns-update * 13:56 swfrench@dns1004: START - running authdns-update * 13:49 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:44 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:31 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 13:29 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 13:26 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 13:23 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 13:22 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 13:19 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 13:18 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 13:18 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 13:17 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 13:16 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 13:13 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 13:11 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 13:09 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 13:06 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 13:06 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-client1002.eqiad.wmnet with OS bookworm * 13:05 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revision-models' for release 'main' . * 13:05 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:05 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revision-models' for release 'main' . * 13:04 brouberol@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-test-client1002.eqiad.wmnet with OS bookworm * 13:04 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 13:03 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 13:02 aikochou@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:00 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'readability' for release 'main' . * 12:59 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'readability' for release 'main' . * 12:58 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 12:57 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'logo-detection' for release 'main' . * 12:57 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'logo-detection' for release 'main' . * 12:57 aikochou@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 12:55 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:54 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:53 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 12:53 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:50 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 12:48 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 12:46 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'article-models' for release 'main' . * 12:45 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'article-models' for release 'main' . * 12:41 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . * 12:39 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . * 12:38 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-client1002.eqiad.wmnet with OS bookworm * 12:12 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply * 12:12 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply * 12:09 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:08 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 11:58 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2187: Security update * 11:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:24 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:16 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 11:15 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 11:10 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2187: Security update * 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2187.codfw.wmnet with reason: Maintenance * 10:56 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 10:56 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 10:56 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 10:56 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 10:54 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 10:53 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 10:09 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2187: Security update * 10:07 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2187: Security update * 09:39 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms2', diff saved to https://phabricator.wikimedia.org/P95929 and previous config saved to /var/cache/conftool/dbconfig/20260806-093908-marostegui.json * 09:36 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1178 from dbctl [[phab:T433471|T433471]]', diff saved to https://phabricator.wikimedia.org/P95928 and previous config saved to /var/cache/conftool/dbconfig/20260806-093632-marostegui.json * 09:33 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 09:31 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 09:30 topranks: bounce cr3-eqsin<->cr2-eqiad bgp session to disable no-prepend command * 09:20 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2253.codfw.wmnet,db1151.eqiad.wmnet with reason: cloning * 09:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1151: Cloning * 09:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:19 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 09:19 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1151: Cloning * 09:10 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 09:09 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup2003.codfw.wmnet * 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup2003.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 09:06 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup2003.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 09:03 klausman@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:02 jynus@cumin1003: START - Cookbook sre.dns.netbox * 09:02 klausman@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 08:57 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup2003.codfw.wmnet * 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup1003.eqiad.wmnet * 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:54 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms3', diff saved to https://phabricator.wikimedia.org/P95925 and previous config saved to /var/cache/conftool/dbconfig/20260806-085422-marostegui.json * 08:53 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:46 jynus@cumin1003: START - Cookbook sre.dns.netbox * 08:39 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup1003.eqiad.wmnet * 08:29 XioNoX: push pfw policy - [[phab:T434115|T434115]] * 08:14 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 08:00 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 08:00 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:58 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revision-models' for release 'main' . * 07:56 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 07:54 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'readability' for release 'main' . * 07:53 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'logo-detection' for release 'main' . * 07:51 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'llm' for release 'main' . * 07:48 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . * 07:37 jayme: updated istio to 1.29.4 on wikikube codfw - [[phab:T427401|T427401]] * 07:08 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2252.codfw.wmnet,db1153.eqiad.wmnet with reason: cloning * 07:07 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1153: Cloning * 07:07 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1153: Cloning * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 40s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-05 == * 23:24 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1009.eqiad.wmnet with OS bookworm * 23:03 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1009.eqiad.wmnet with reason: host reimage * 22:59 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1009.eqiad.wmnet with reason: host reimage * 22:43 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1009.eqiad.wmnet with OS bookworm * 22:38 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1009.eqiad.wmnet * 22:34 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1009.eqiad.wmnet * 22:25 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1008.eqiad.wmnet with OS bookworm * 22:04 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1008.eqiad.wmnet with reason: host reimage * 22:00 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1008.eqiad.wmnet with reason: host reimage * 21:48 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:47 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:46 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:44 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1008.eqiad.wmnet with OS bookworm * 21:43 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:41 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1008.eqiad.wmnet * 21:36 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1008.eqiad.wmnet * 21:14 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:12 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad * 21:12 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad * 21:10 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=eqiad * 21:08 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:07 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:07 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=eqiad * 21:04 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:03 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1006 * 21:02 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1006 * 21:00 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:56 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:56 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:55 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 20:55 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 20:55 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1007.eqiad.wmnet with OS bookworm * 20:51 vriley@cumin1003: START - Cookbook sre.dns.netbox * 20:43 ebernhardson: [[phab:T434008|T434008]]: changing cloudelastic:9643 from auto_expand_replicas to number_of_replicas * 20:34 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1007.eqiad.wmnet with reason: host reimage * 20:27 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1007.eqiad.wmnet with reason: host reimage * 20:24 cjming: end of UTC late backport window * 20:23 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] (duration: 06m 26s) * 20:18 cjming@deploy1003: cjming: Continuing with deployment * 20:18 cjming@deploy1003: cjming: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:16 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] * 20:12 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1007.eqiad.wmnet with OS bookworm * 20:12 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] (duration: 08m 41s) * 20:08 swfrench@cumin2002: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host conf1007.eqiad.wmnet with OS bookworm * 20:08 jforrester@deploy1003: jforrester: Continuing with deployment * 20:07 jforrester@deploy1003: jforrester: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:03 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] * 19:51 inflatador: [bking@puppetserver1001] ~$ sudo puppetserver ca sign --certname an-worker1189.eqiad.wmnet [[phab:T434142|T434142]] * 19:47 bking@cumin2003: DONE (FAIL) - Cookbook sre.puppet.renew-cert (exit_code=99) for an-worker1189.eqiad.wmnet: Renew puppet certificate - bking@cumin2003 * 19:46 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:30 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1007.eqiad.wmnet with OS trixie * 19:30 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 19:29 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 19:20 swfrench-wmf: silenced EtcdRelicationDown 0cb709a9-f244-4f1e-971f-{{Gerrit|440ec65e7fd7}} - [[phab:T428495|T428495]] * 19:13 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1007.eqiad.wmnet with OS bookworm * 19:12 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1007.eqiad.wmnet with reason: host reimage * 19:09 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1007.eqiad.wmnet * 19:07 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1007.eqiad.wmnet with reason: host reimage * 19:03 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1007.eqiad.wmnet * 18:52 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1007.eqiad.wmnet with OS trixie * 18:52 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1007.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:35 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 18:34 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 18:34 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 18:30 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1007.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:28 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:28 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1007] - vriley@cumin1003" * 18:27 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1007] - vriley@cumin1003" * 18:23 vriley@cumin1003: START - Cookbook sre.dns.netbox * 18:22 vriley@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 18:22 robh@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:19 vriley@cumin1003: START - Cookbook sre.dns.netbox * 18:13 robh@cumin2002: START - Cookbook sre.hosts.provision for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:31 jasmine@cumin2002: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-main-eqiad * 17:12 mutante: LDAP - added vwalters to group ciadmin - [[phab:T433615|T433615]] * 16:58 aokoth@deploy1003: Finished deploy [phabricator/deployment@e2ebca5]: Deploy Phab (duration: 00m 34s) * 16:57 aokoth@deploy1003: Started deploy [phabricator/deployment@e2ebca5]: Deploy Phab * 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad * 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=eqiad * 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad * 16:53 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:41 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2187.codfw.wmnet * 16:41 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2187.codfw.wmnet * 16:41 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker2187.codfw.wmnet * 16:41 cgoubert@cumin2003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker2187.codfw.wmnet * 16:40 jasmine@cumin2002: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-main-eqiad * 16:40 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:34 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:25 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-magru and A:liberica ([[phab:T428495|T428495]]) * 16:23 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-magru and A:liberica ([[phab:T428495|T428495]]) * 16:20 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-drmrs and A:liberica ([[phab:T428495|T428495]]) * 16:19 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-drmrs and A:liberica ([[phab:T428495|T428495]]) * 16:18 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-esams and A:liberica ([[phab:T428495|T428495]]) * 16:16 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-esams and A:liberica ([[phab:T428495|T428495]]) * 16:06 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1159.eqiad.wmnet * 16:06 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1159.eqiad.wmnet * 16:06 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1159.eqiad.wmnet * 16:05 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] (duration: 09m 11s) * 15:58 reedy@deploy1003: reedy: Continuing with deployment * 15:58 reedy@deploy1003: reedy: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:56 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] * 15:54 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1159.eqiad.wmnet with OS trixie * 15:38 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:33 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1159.eqiad.wmnet with reason: host reimage * 15:32 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:27 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1159.eqiad.wmnet with reason: host reimage * 15:10 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1159 * 15:10 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1159 * 15:00 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo for hosts in the wikimedia.org domain - [[phab:T428495|T428495]] * 14:55 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS trixie * 14:54 swfrench-wmf: restarted navtiming on webperf1003 - [[phab:T428495|T428495]] * 14:52 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1159 * 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1159.eqiad.wmnet 129.48.64.10.in-addr.arpa 9.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:52 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1159.eqiad.wmnet 129.48.64.10.in-addr.arpa 9.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1159 - jayme@cumin1003" * 14:52 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1159 - jayme@cumin1003" * 14:49 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:48 jayme@cumin1003: START - Cookbook sre.dns.netbox * 14:47 swfrench-wmf: begin rolling restart of confd in drmrs, eqiad, esams, magru - [[phab:T428495|T428495]] * 14:47 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1159 * 14:46 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:46 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:46 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1159.eqiad.wmnet with OS trixie * 14:45 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:44 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:44 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1159.eqiad.wmnet * 14:43 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:43 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1159.eqiad.wmnet * 14:43 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:43 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1159.eqiad.wmnet * 14:43 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:43 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1157.eqiad.wmnet * 14:43 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1157.eqiad.wmnet * 14:43 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1157.eqiad.wmnet * 14:42 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:42 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:42 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:41 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:41 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:41 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:41 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1003.eqiad.wmnet with OS bookworm * 14:39 swfrench-wmf: authdns update to direct eqiad-associated etcd clients to codfw - [[phab:T428495|T428495]] * 14:39 swfrench@dns1004: END - running authdns-update * 14:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 14:37 swfrench@dns1004: START - running authdns-update * 14:37 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:37 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:35 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:35 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 14:28 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:27 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1157.eqiad.wmnet with OS trixie * 14:27 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:27 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:27 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 14:26 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 14:26 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:26 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 14:26 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 14:26 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:26 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host search-loader1002.eqiad.wmnet with OS trixie * 14:19 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:19 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:15 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1003.eqiad.wmnet with reason: host reimage * 14:14 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:14 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:13 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 * 14:13 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host mc2046 * 14:13 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS trixie * 14:11 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1003.eqiad.wmnet with reason: host reimage * 14:10 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:09 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:09 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:09 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:08 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:08 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1157.eqiad.wmnet with reason: host reimage * 14:08 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:04 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 14:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on search-loader1002.eqiad.wmnet with reason: host reimage * 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=eqiad * 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=eqiad * 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=eqiad * 14:00 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:59 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:58 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1157.eqiad.wmnet with reason: host reimage * 13:57 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on search-loader1002.eqiad.wmnet with reason: host reimage * 13:54 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1003.eqiad.wmnet with OS bookworm * 13:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host search-loader1002.eqiad.wmnet with OS trixie * 13:43 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1157 * 13:42 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1157 * 13:40 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1157 * 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1157.eqiad.wmnet 183.32.64.10.in-addr.arpa 3.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1157.eqiad.wmnet 183.32.64.10.in-addr.arpa 3.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1157 - jayme@cumin1003" * 13:39 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1157 - jayme@cumin1003" * 13:39 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] (duration: 07m 00s) * 13:35 jayme@cumin1003: START - Cookbook sre.dns.netbox * 13:35 reedy@deploy1003: reedy: Continuing with deployment * 13:34 reedy@deploy1003: reedy: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:32 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] * 13:23 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1157 * 13:22 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1157.eqiad.wmnet with OS trixie * 13:22 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1157.eqiad.wmnet * 13:22 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1157.eqiad.wmnet * 13:21 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1157.eqiad.wmnet * 13:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1156.eqiad.wmnet * 13:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1156.eqiad.wmnet * 13:15 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1156.eqiad.wmnet * 13:01 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1156.eqiad.wmnet with OS trixie * 12:42 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1156.eqiad.wmnet with reason: host reimage * 12:38 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1156.eqiad.wmnet with reason: host reimage * 12:32 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 12:31 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 12:30 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 12:28 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 12:26 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 12:24 topranks: update bgp confed settings in eqsin * 12:22 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1156 * 12:22 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1156 * 12:22 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 12:19 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1156 * 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1156.eqiad.wmnet 110.32.64.10.in-addr.arpa 0.1.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:19 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1156.eqiad.wmnet 110.32.64.10.in-addr.arpa 0.1.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1156 - jayme@cumin1003" * 12:19 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1156 - jayme@cumin1003" * 12:17 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:14 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS trixie * 12:09 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:06 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:04 jayme@cumin1003: START - Cookbook sre.dns.netbox * 12:04 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 12:02 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'article-models' for release 'main' . * 12:01 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1156 * 12:01 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1156.eqiad.wmnet with OS trixie * 11:59 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1156.eqiad.wmnet * 11:59 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1156.eqiad.wmnet * 11:59 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1156.eqiad.wmnet * 11:57 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 11:53 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 11:53 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:52 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:52 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:50 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:50 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:50 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:49 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:48 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:47 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:47 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:45 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:45 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:44 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:44 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:44 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:43 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:42 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:38 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:35 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 * 11:35 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host mc2046 * 11:34 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS trixie * 11:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:27 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:21 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:21 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:18 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:18 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:18 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:18 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:13 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:13 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:09 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:08 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:07 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:06 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:06 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:05 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:05 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:04 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 11:04 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:24 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:24 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:17 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:16 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1155.eqiad.wmnet * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1155.eqiad.wmnet * 10:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1155.eqiad.wmnet * 10:14 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:14 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:11 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:11 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:10 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:09 aikochou@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop: sync * 10:09 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:09 aikochou@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop: sync * 10:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:07 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:05 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:05 aikochou@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop: sync * 10:05 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:05 aikochou@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop: sync * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:04 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:04 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:04 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1155.eqiad.wmnet with OS trixie * 09:52 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms1', diff saved to https://phabricator.wikimedia.org/P95918 and previous config saved to /var/cache/conftool/dbconfig/20260805-095212-marostegui.json * 09:44 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1152: after cloning * 09:44 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.parsercache (exit_code=99) * 09:44 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 09:44 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1152: after cloning * 09:43 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1155.eqiad.wmnet with reason: host reimage * 09:40 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1155.eqiad.wmnet with reason: host reimage * 09:32 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 09:32 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:31 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 09:31 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:27 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1155 * 09:27 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1155 * 09:25 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 09:24 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 09:24 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 09:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:23 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 09:23 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 09:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:22 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 09:22 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 09:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:20 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 09:20 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:17 XioNoX: push pfw policies - [[phab:T434038|T434038]] * 09:14 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1155 * 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1155.eqiad.wmnet 109.32.64.10.in-addr.arpa 9.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1155.eqiad.wmnet 109.32.64.10.in-addr.arpa 9.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1155 - jayme@cumin1003" * 09:14 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1155 - jayme@cumin1003" * 09:10 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 09:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:09 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2251.codfw.wmnet,db1152.eqiad.wmnet with reason: cloning * 09:09 jayme@cumin1003: START - Cookbook sre.dns.netbox * 09:08 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 09:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: Cloning * 09:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:05 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 09:05 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1152: Cloning * 08:38 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1155 * 08:37 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1155.eqiad.wmnet with OS trixie * 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1171.eqiad.wmnet * 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1171.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:29 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95913 and previous config saved to /var/cache/conftool/dbconfig/20260805-082908-ladsgroup.json * 08:27 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1171.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:22 jynus@cumin1003: START - Cookbook sre.dns.netbox * 08:18 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249', diff saved to https://phabricator.wikimedia.org/P95912 and previous config saved to /var/cache/conftool/dbconfig/20260805-081823-ladsgroup.json * 08:17 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1171.eqiad.wmnet * 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1150.eqiad.wmnet * 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1150.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:15 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1150.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:15 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 08:14 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1155.eqiad.wmnet * 08:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1155.eqiad.wmnet * 08:14 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1155.eqiad.wmnet * 08:11 jynus@cumin1003: START - Cookbook sre.dns.netbox * 08:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249', diff saved to https://phabricator.wikimedia.org/P95911 and previous config saved to /var/cache/conftool/dbconfig/20260805-080737-ladsgroup.json * 08:05 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1150.eqiad.wmnet * 08:02 marostegui: Depool clouddb1020 (s5,s8) [[phab:T434048|T434048]] * 08:02 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1020.eqiad.wmnet,service=s8 * 08:02 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1020.eqiad.wmnet,service=s5 * 08:02 marostegui: Depool clouddb1018 (s2,s7) [[phab:T434048|T434048]] * 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1018.eqiad.wmnet,service=s7 * 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1018.eqiad.wmnet,service=s2 * 08:01 marostegui: Depool clouddb1017 (s1) [[phab:T434048|T434048]] * 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1017.eqiad.wmnet,service=s1 * 07:59 marostegui: Depool clouddb1016 (s5,s8) [[phab:T434048|T434048]] * 07:59 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s8 * 07:59 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s5 * 07:57 marostegui: Depool clouddb1015 (s4,s6) [[phab:T434048|T434048]] * 07:57 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s6 * 07:57 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s4 * 07:56 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95910 and previous config saved to /var/cache/conftool/dbconfig/20260805-075650-ladsgroup.json * 07:54 marostegui: Depool clouddb1014 (s2,s7) [[phab:T434048|T434048]] * 07:54 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1014.eqiad.wmnet,service=s7 * 07:54 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1014.eqiad.wmnet,service=s2 * 07:53 marostegui: Depool clouddb1013:s1 [[phab:T434048|T434048]] * 07:53 marostegui: Depool clouddb1013:s1 [[phab:T409557|T409557]] * 07:53 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1013.eqiad.wmnet,service=s1 * 07:25 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95909 and previous config saved to /var/cache/conftool/dbconfig/20260805-072529-ladsgroup.json * 07:24 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2249.codfw.wmnet with reason: Maintenance * 07:24 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95908 and previous config saved to /var/cache/conftool/dbconfig/20260805-072426-ladsgroup.json * 07:21 slyngshede@dns1004: END - running authdns-update * 07:19 slyngshede@dns1004: START - running authdns-update * 07:13 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231', diff saved to https://phabricator.wikimedia.org/P95906 and previous config saved to /var/cache/conftool/dbconfig/20260805-071340-ladsgroup.json * 07:02 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231', diff saved to https://phabricator.wikimedia.org/P95905 and previous config saved to /var/cache/conftool/dbconfig/20260805-070253-ladsgroup.json * 06:52 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95904 and previous config saved to /var/cache/conftool/dbconfig/20260805-065206-ladsgroup.json * 06:45 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 06:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95903 and previous config saved to /var/cache/conftool/dbconfig/20260805-062240-ladsgroup.json * 06:21 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2231.codfw.wmnet with reason: Maintenance * 06:21 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95902 and previous config saved to /var/cache/conftool/dbconfig/20260805-062137-ladsgroup.json * 06:10 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215', diff saved to https://phabricator.wikimedia.org/P95901 and previous config saved to /var/cache/conftool/dbconfig/20260805-061051-ladsgroup.json * 06:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215', diff saved to https://phabricator.wikimedia.org/P95900 and previous config saved to /var/cache/conftool/dbconfig/20260805-060004-ladsgroup.json * 05:49 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95899 and previous config saved to /var/cache/conftool/dbconfig/20260805-054918-ladsgroup.json * 05:19 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95898 and previous config saved to /var/cache/conftool/dbconfig/20260805-051939-ladsgroup.json * 05:18 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2215.codfw.wmnet with reason: Maintenance * 04:30 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2201.codfw.wmnet with reason: Maintenance * 03:40 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2197.codfw.wmnet with reason: Maintenance * 03:40 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95897 and previous config saved to /var/cache/conftool/dbconfig/20260805-034036-ladsgroup.json * 03:29 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196', diff saved to https://phabricator.wikimedia.org/P95896 and previous config saved to /var/cache/conftool/dbconfig/20260805-032948-ladsgroup.json * 03:19 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196', diff saved to https://phabricator.wikimedia.org/P95895 and previous config saved to /var/cache/conftool/dbconfig/20260805-031902-ladsgroup.json * 03:08 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95894 and previous config saved to /var/cache/conftool/dbconfig/20260805-030815-ladsgroup.json * 02:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95893 and previous config saved to /var/cache/conftool/dbconfig/20260805-023413-ladsgroup.json * 02:33 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2196.codfw.wmnet with reason: Maintenance * 02:33 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95892 and previous config saved to /var/cache/conftool/dbconfig/20260805-023310-ladsgroup.json * 02:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186', diff saved to https://phabricator.wikimedia.org/P95891 and previous config saved to /var/cache/conftool/dbconfig/20260805-022223-ladsgroup.json * 02:11 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186', diff saved to https://phabricator.wikimedia.org/P95890 and previous config saved to /var/cache/conftool/dbconfig/20260805-021137-ladsgroup.json * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 02:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95889 and previous config saved to /var/cache/conftool/dbconfig/20260805-020051-ladsgroup.json * 01:30 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95888 and previous config saved to /var/cache/conftool/dbconfig/20260805-013029-ladsgroup.json * 01:29 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2186.codfw.wmnet with reason: Maintenance * 00:34 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on dbstore1009.eqiad.wmnet with reason: Maintenance * 00:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95887 and previous config saved to /var/cache/conftool/dbconfig/20260805-003408-ladsgroup.json * 00:23 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264', diff saved to https://phabricator.wikimedia.org/P95886 and previous config saved to /var/cache/conftool/dbconfig/20260805-002322-ladsgroup.json * 00:12 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264', diff saved to https://phabricator.wikimedia.org/P95885 and previous config saved to /var/cache/conftool/dbconfig/20260805-001235-ladsgroup.json * 00:01 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95884 and previous config saved to /var/cache/conftool/dbconfig/20260805-000148-ladsgroup.json == 2026-08-04 == * 23:45 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95883 and previous config saved to /var/cache/conftool/dbconfig/20260804-234508-ladsgroup.json * 23:44 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1264.eqiad.wmnet with reason: Maintenance * 23:44 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95882 and previous config saved to /var/cache/conftool/dbconfig/20260804-234405-ladsgroup.json * 23:33 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237', diff saved to https://phabricator.wikimedia.org/P95881 and previous config saved to /var/cache/conftool/dbconfig/20260804-233317-ladsgroup.json * 23:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237', diff saved to https://phabricator.wikimedia.org/P95880 and previous config saved to /var/cache/conftool/dbconfig/20260804-232230-ladsgroup.json * 23:11 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95879 and previous config saved to /var/cache/conftool/dbconfig/20260804-231144-ladsgroup.json * 22:23 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95878 and previous config saved to /var/cache/conftool/dbconfig/20260804-222345-ladsgroup.json * 22:23 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1237.eqiad.wmnet with reason: Maintenance * 21:13 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1225.eqiad.wmnet with reason: Maintenance * 20:40 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] (duration: 24m 40s) * 20:33 samtar@deploy1003: samtar, kineticpelagic: Continuing with deployment * 20:28 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS bookworm * 20:21 samtar@deploy1003: samtar, kineticpelagic: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:15 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] * 20:13 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 20:09 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 20:00 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1216.eqiad.wmnet with reason: Maintenance * 20:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95877 and previous config saved to /var/cache/conftool/dbconfig/20260804-195957-ladsgroup.json * 19:51 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 19:50 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) mc2046.codfw.wmnet 120.16.192.10.in-addr.arpa 0.2.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:50 jhancock@cumin2002: START - Cookbook sre.dns.wipe-cache mc2046.codfw.wmnet 120.16.192.10.in-addr.arpa 0.2.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host mc2046 - jhancock@cumin2002" * 19:50 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host mc2046 - jhancock@cumin2002" * 19:49 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203', diff saved to https://phabricator.wikimedia.org/P95876 and previous config saved to /var/cache/conftool/dbconfig/20260804-194911-ladsgroup.json * 19:46 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 19:45 jhancock@cumin2002: START - Cookbook sre.hosts.move-vlan for host mc2046 * 19:45 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS bookworm * 19:38 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203', diff saved to https://phabricator.wikimedia.org/P95875 and previous config saved to /var/cache/conftool/dbconfig/20260804-193825-ladsgroup.json * 19:27 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95874 and previous config saved to /var/cache/conftool/dbconfig/20260804-192738-ladsgroup.json * 19:02 mutante: gerrit ssh -p 29418 gerrit.wikimedia.org gerrit index changes {{Gerrit|1320979}} * 18:20 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 18:18 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 18:14 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 18:14 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 18:13 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 18:10 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 18:08 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 18:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95872 and previous config saved to /var/cache/conftool/dbconfig/20260804-180721-ladsgroup.json * 18:07 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 18:06 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1203.eqiad.wmnet with reason: Maintenance * 18:06 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95871 and previous config saved to /var/cache/conftool/dbconfig/20260804-180618-ladsgroup.json * 17:55 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179', diff saved to https://phabricator.wikimedia.org/P95870 and previous config saved to /var/cache/conftool/dbconfig/20260804-175531-ladsgroup.json * 17:55 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1154.eqiad.wmnet * 17:55 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1154.eqiad.wmnet * 17:55 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1154.eqiad.wmnet * 17:50 swfrench@deploy1003: Finished scap sync-world: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] (duration: 04m 05s) * 17:48 swfrench@deploy1003: swfrench: Continuing with deployment * 17:46 swfrench@deploy1003: swfrench: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:45 swfrench@deploy1003: Started scap sync-world: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] * 17:44 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179', diff saved to https://phabricator.wikimedia.org/P95869 and previous config saved to /var/cache/conftool/dbconfig/20260804-174445-ladsgroup.json * 17:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95868 and previous config saved to /var/cache/conftool/dbconfig/20260804-173359-ladsgroup.json * 17:33 swfrench@deploy1003: Finished scap sync-world: Pick up new PHP production image (duration: 28m 32s) * 17:28 aokoth@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on phab1005.eqiad.wmnet with reason: Puppet Failure * 17:05 swfrench@deploy1003: Started scap sync-world: Pick up new PHP production image * 17:00 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 17:00 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 16:54 cgoubert@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on wikikube-worker2187.codfw.wmnet with reason: Hardware issue * 16:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2187.codfw.wmnet * 16:52 mutante: gerrit2003:/var/log/apache2# ln -s /srv/gerrit/site_path/review_site/logs/ gerrit * 16:52 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2187.codfw.wmnet * 16:48 mutante: gerrit2003 - moving old apache logfiles older than 60 days from /var/log/apache2 to /srv/gerrit/site_path/review_site/logs/old/ * 16:33 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 16:32 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 16:29 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 16:29 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 16:28 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 16:28 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 16:27 dzahn@cumin1003: END (PASS) - Cookbook sre.gerrit.restart-gerrit (exit_code=0) Restarting Gerrit on gerrit2003 * 16:27 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 16:27 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95867 and previous config saved to /var/cache/conftool/dbconfig/20260804-162736-ladsgroup.json * 16:27 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 16:26 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1179.eqiad.wmnet with reason: Maintenance * 16:26 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:25 mutante: restarting gerrit - dropped outdated RSA host key * 16:25 dzahn@cumin1003: START - Cookbook sre.gerrit.restart-gerrit Restarting Gerrit on gerrit2003 * 16:24 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95866 and previous config saved to /var/cache/conftool/dbconfig/20260804-162424-ladsgroup.json * 16:24 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 16:23 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 16:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95865 and previous config saved to /var/cache/conftool/dbconfig/20260804-162236-ladsgroup.json * 16:21 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1179.eqiad.wmnet with reason: Maintenance * 16:17 swfrench-wmf: reprepro include php8.3_8.3.33-1+wmf11u1 into component/php83 for bullseye-wikimedia * 16:17 swfrench-wmf: reprepro include php8.3_8.3.33-1+wmf12u1 into component/php83 for bookworm-wikimedia * 16:11 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply * 16:10 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply * 16:10 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mobileapps: apply * 16:09 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mobileapps: apply * 16:09 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply * 16:08 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply * 16:08 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:08 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:07 aokoth@cumin1003: END (PASS) - Cookbook sre.vrts.upgrade (exit_code=0) on VRTS host vrts1003.eqiad.wmnet * 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:05 aokoth@cumin1003: START - Cookbook sre.vrts.upgrade on VRTS host vrts1003.eqiad.wmnet * 16:04 mutante: gerrit2002/gerrit1003/gerrit2003 - rm /etc/gerrit/ssh_host_rsa_key * 15:59 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:59 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:59 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:59 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:56 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 15:55 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:55 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:55 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:49 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 15:49 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:44 Raine: add php8.5 packages to component/php85 - [[phab:T432983|T432983]] * 15:39 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:33 aaron@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 15:33 aaron@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 15:29 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:19 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:19 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:16 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:16 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1154.eqiad.wmnet with OS trixie * 15:16 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:15 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:15 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:06 brennen@deploy1003: Finished deploy [phabricator/deployment@56f4ffd]: deploy phab1004 for [[phab:T433981|T433981]] (duration: 00m 43s) * 15:05 brennen@deploy1003: Started deploy [phabricator/deployment@56f4ffd]: deploy phab1004 for [[phab:T433981|T433981]] * 15:05 aaron@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 15:04 aaron@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 15:02 brennen@deploy1003: Finished deploy [phabricator/deployment@56f4ffd]: deploy phab2003 for [[phab:T433981|T433981]] (duration: 00m 51s) * 15:01 brennen@deploy1003: Started deploy [phabricator/deployment@56f4ffd]: deploy phab2003 for [[phab:T433981|T433981]] * 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1004.eqiad.wmnet with reason: deployment * 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1005.eqiad.wmnet with reason: deployment * 14:58 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab2003.codfw.wmnet with reason: deployment * 14:55 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1154.eqiad.wmnet with reason: host reimage * 14:51 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1154.eqiad.wmnet with reason: host reimage * 14:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 14:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 14:38 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync * 14:38 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync * 14:38 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync * 14:37 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync * 14:37 ottomata: roll restart eventgate-main to pick up stream config change - [[phab:T433507|T433507]] * 14:37 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-main: sync * 14:36 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-main: sync * 14:36 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1154 * 14:36 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1154 * 14:34 otto@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] (duration: 08m 39s) * 14:34 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1154 * 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1154.eqiad.wmnet 108.32.64.10.in-addr.arpa 8.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:34 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1154.eqiad.wmnet 108.32.64.10.in-addr.arpa 8.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1154 - jayme@cumin1003" * 14:34 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1154 - jayme@cumin1003" * 14:30 otto@deploy1003: otto: Continuing with deployment * 14:30 jayme@cumin1003: START - Cookbook sre.dns.netbox * 14:28 otto@deploy1003: otto: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:26 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1154 * 14:26 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1154.eqiad.wmnet with OS trixie * 14:26 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1154.eqiad.wmnet * 14:26 otto@deploy1003: Started scap sync-world: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] * 14:26 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1154.eqiad.wmnet * 14:26 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1154.eqiad.wmnet * 14:17 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 14:16 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 14:15 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 14:14 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 14:13 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 14:13 swfrench@dns1004: END - running authdns-update * 14:13 Msz2001: Finished deployments for UTC afternoon backport window * 14:13 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 14:13 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] (duration: 07m 58s) * 14:11 swfrench@dns1004: START - running authdns-update * 14:08 mszwarc@deploy1003: javiermonton, mszwarc, mpostoronca: Continuing with deployment * 14:07 mszwarc@deploy1003: javiermonton, mszwarc, mpostoronca: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] synced to the testser * 14:05 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] * 14:03 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 13:49 swfrench@cumin2002: conftool action : set/pooled=yes; selector: name=wikikube-worker2330.codfw.wmnet * 13:49 swfrench@cumin2002: conftool action : set/pooled=no; selector: name=wikikube-worker2330.codfw.wmnet * 13:48 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] (duration: 09m 19s) * 13:45 swfrench@dns1004: END - running authdns-update * 13:44 mszwarc@deploy1003: mszwarc, jforrester: Continuing with deployment * 13:43 swfrench@dns1004: START - running authdns-update * 13:41 mszwarc@deploy1003: mszwarc, jforrester: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:38 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] * 13:33 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 13:33 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1154.eqiad.wmnet * 13:32 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 13:32 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 13:31 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 13:31 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:31 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:29 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1154.eqiad.wmnet * 13:28 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1154.eqiad.wmnet * 13:28 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1154.eqiad.wmnet * 13:28 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1141.eqiad.wmnet * 13:28 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1141.eqiad.wmnet * 13:28 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1141.eqiad.wmnet * 13:22 otto@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply * 13:22 otto@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply * 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1096.eqiad.wmnet with OS trixie * 13:05 swfrench@dns1004: END - running authdns-update * 13:03 swfrench@dns1004: START - running authdns-update * 12:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 12:43 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 1:00:00 on db1171.eqiad.wmnet with reason: decom * 12:42 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 1:00:00 on db1150.eqiad.wmnet with reason: decom * 12:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 12:38 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1164,1217].eqiad.wmnet with reason: cloning * 12:33 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2096.codfw.wmnet with OS trixie * 12:22 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1096.eqiad.wmnet with OS trixie * 12:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2096.codfw.wmnet with reason: host reimage * 12:14 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1141.eqiad.wmnet with OS trixie * 12:10 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2096.codfw.wmnet with reason: host reimage * 12:10 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1289.eqiad.wmnet * 12:05 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1289.eqiad.wmnet * 12:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1288.eqiad.wmnet * 11:59 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1288.eqiad.wmnet * 11:59 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1287.eqiad.wmnet * 11:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1097.eqiad.wmnet with OS trixie * 11:54 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1287.eqiad.wmnet * 11:54 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1286.eqiad.wmnet * 11:53 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1141.eqiad.wmnet with reason: host reimage * 11:51 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2096.codfw.wmnet with OS trixie * 11:49 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1141.eqiad.wmnet with reason: host reimage * 11:48 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1286.eqiad.wmnet * 11:48 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1284.eqiad.wmnet * 11:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2095.codfw.wmnet with OS trixie * 11:43 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1284.eqiad.wmnet * 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1283.eqiad.wmnet * 11:42 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on ml-serve1015.eqiad.wmnet with reason: Downtime to get full picture of current BIOS settings beyond what Redfish shows * 11:39 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad * 11:39 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:37 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1283.eqiad.wmnet * 11:37 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1282.eqiad.wmnet * 11:37 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad * 11:37 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:33 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1141 * 11:33 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1141 * 11:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 11:32 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1141 * 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1141.eqiad.wmnet 156.48.64.10.in-addr.arpa 6.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:32 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1141.eqiad.wmnet 156.48.64.10.in-addr.arpa 6.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1141 - jayme@cumin1003" * 11:32 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1141 - jayme@cumin1003" * 11:32 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1282.eqiad.wmnet * 11:32 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1281.eqiad.wmnet * 11:32 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad * 11:32 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:29 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 11:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1097.eqiad.wmnet with reason: host reimage * 11:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2095.codfw.wmnet with OS trixie * 11:26 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1281.eqiad.wmnet * 11:26 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1280.eqiad.wmnet * 11:25 jayme@cumin1003: START - Cookbook sre.dns.netbox * 11:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1097.eqiad.wmnet with reason: host reimage * 11:22 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1141 * 11:21 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1141.eqiad.wmnet with OS trixie * 11:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1280.eqiad.wmnet * 11:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1279.eqiad.wmnet * 11:20 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin with reason: upgrade new Nokia swtiches in eqsin to SR Linux v26 * 11:17 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1141.eqiad.wmnet * 11:16 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1141.eqiad.wmnet * 11:16 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1141.eqiad.wmnet * 11:16 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be2095.codfw.wmnet with OS trixie * 11:15 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1279.eqiad.wmnet * 11:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1278.eqiad.wmnet * 11:14 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1139.eqiad.wmnet * 11:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1139.eqiad.wmnet * 11:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1070.eqiad.wmnet with OS trixie * 11:13 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1140.eqiad.wmnet * 11:13 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1140.eqiad.wmnet * 11:13 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1140.eqiad.wmnet * 11:09 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1278.eqiad.wmnet * 11:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1071.eqiad.wmnet with OS trixie * 11:05 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1097.eqiad.wmnet with OS trixie * 11:04 mvernon@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be1097.eqiad.wmnet with OS trixie * 11:02 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1140.eqiad.wmnet with OS trixie * 11:02 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1097.eqiad.wmnet with OS trixie * 11:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1096.eqiad.wmnet with OS trixie * 11:00 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1139.eqiad.wmnet * 11:00 jayme@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1139.eqiad.wmnet with OS trixie * 10:56 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1069.eqiad.wmnet with OS trixie * 10:56 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 10:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1070.eqiad.wmnet with reason: host reimage * 10:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 10:45 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1071.eqiad.wmnet with reason: host reimage * 10:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 10:41 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1096.eqiad.wmnet with OS trixie * 10:41 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1140.eqiad.wmnet with reason: host reimage * 10:39 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1071.eqiad.wmnet with reason: host reimage * 10:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1070.eqiad.wmnet with reason: host reimage * 10:38 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1139.eqiad.wmnet with reason: host reimage * 10:37 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1140.eqiad.wmnet with reason: host reimage * 10:35 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1069.eqiad.wmnet with reason: host reimage * 10:33 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1095.eqiad.wmnet with OS trixie * 10:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 10:33 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1139.eqiad.wmnet with reason: host reimage * 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1069.eqiad.wmnet with reason: host reimage * 10:23 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1140 * 10:23 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1140 * 10:23 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 10:22 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1140 * 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1140.eqiad.wmnet 155.48.64.10.in-addr.arpa 5.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:21 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1071.eqiad.wmnet with OS trixie * 10:21 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1140.eqiad.wmnet 155.48.64.10.in-addr.arpa 5.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1140 - jayme@cumin1003" * 10:21 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1140 - jayme@cumin1003" * 10:21 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1071 * 10:21 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1070.eqiad.wmnet with OS trixie * 10:21 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1070 * 10:20 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 10:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1095.eqiad.wmnet with OS trixie * 10:17 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1139 * 10:17 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1139 * 10:17 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be1095.eqiad.wmnet with OS trixie * 10:15 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1139 * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1139.eqiad.wmnet 194.32.64.10.in-addr.arpa 4.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:15 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1139.eqiad.wmnet 194.32.64.10.in-addr.arpa 4.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1139 - jayme@cumin1003" * 10:15 jayme@cumin1003: START - Cookbook sre.dns.netbox * 10:15 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1139 - jayme@cumin1003" * 10:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2095.codfw.wmnet with OS trixie * 10:13 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1069.eqiad.wmnet with OS trixie * 10:12 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1069 * 10:11 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1140 * 10:11 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1140.eqiad.wmnet with OS trixie * 10:11 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1140.eqiad.wmnet * 10:10 jayme@cumin1003: START - Cookbook sre.dns.netbox * 10:10 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1139 * 10:10 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1140.eqiad.wmnet * 10:10 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1140.eqiad.wmnet * 10:10 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1139.eqiad.wmnet with OS trixie * 10:09 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1139.eqiad.wmnet * 10:08 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1139.eqiad.wmnet * 10:08 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1139.eqiad.wmnet * 10:01 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2094.codfw.wmnet with OS trixie * 09:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 09:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 09:53 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 09:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 09:44 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:44 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2094.codfw.wmnet with reason: host reimage * 09:34 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2094.codfw.wmnet with reason: host reimage * 09:34 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1071 * 09:33 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1070 * 09:33 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1095.eqiad.wmnet with OS trixie * 09:27 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1069 * 09:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:23 brouberol@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM archiva1002.wikimedia.org * 09:20 brouberol@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM archiva1002.wikimedia.org * 09:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1277.eqiad.wmnet * 09:13 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2094.codfw.wmnet with OS trixie * 09:13 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 09:12 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 09:12 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:12 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1277.eqiad.wmnet * 09:12 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1276.eqiad.wmnet * 09:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1094.eqiad.wmnet with OS trixie * 09:06 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1276.eqiad.wmnet * 09:06 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1275.eqiad.wmnet * 09:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2093.codfw.wmnet with OS trixie * 09:01 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1275.eqiad.wmnet * 09:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1274.eqiad.wmnet * 08:56 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1274.eqiad.wmnet * 08:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1273.eqiad.wmnet * 08:50 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1273.eqiad.wmnet * 08:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1272.eqiad.wmnet * 08:49 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:49 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1094.eqiad.wmnet with reason: host reimage * 08:45 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1272.eqiad.wmnet * 08:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1094.eqiad.wmnet with reason: host reimage * 08:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2093.codfw.wmnet with reason: host reimage * 08:38 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:38 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2093.codfw.wmnet with reason: host reimage * 08:35 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 08:34 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 08:29 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:28 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:26 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1271.eqiad.wmnet * 08:23 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1094.eqiad.wmnet with OS trixie * 08:21 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 08:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1271.eqiad.wmnet * 08:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1270.eqiad.wmnet * 08:15 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1270.eqiad.wmnet * 08:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1269.eqiad.wmnet * 08:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2093.codfw.wmnet with OS trixie * 08:09 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1269.eqiad.wmnet * 08:09 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1268.eqiad.wmnet * 08:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2092.codfw.wmnet with OS trixie * 08:04 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1268.eqiad.wmnet * 08:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1267.eqiad.wmnet * 07:59 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1267.eqiad.wmnet * 07:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1093.eqiad.wmnet with OS trixie * 07:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1266.eqiad.wmnet * 07:56 jynus: running extra backups to test db1285 [[phab:T433826|T433826]] * 07:51 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1266.eqiad.wmnet * 07:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2092.codfw.wmnet with reason: host reimage * 07:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1093.eqiad.wmnet with reason: host reimage * 07:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2092.codfw.wmnet with reason: host reimage * 07:32 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1093.eqiad.wmnet with reason: host reimage * 07:29 jynus: running extra backups to test db1265 [[phab:T433825|T433825]] * 07:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2092.codfw.wmnet with OS trixie * 07:11 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1093.eqiad.wmnet with OS trixie * 06:50 slyngshede@dns1004: END - running authdns-update * 06:48 slyngshede@dns1004: START - running authdns-update * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.11 (duration: 02m 29s) * 03:38 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] (duration: 32m 57s) * 03:23 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 03:22 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 03:05 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 32s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 00:45 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] (duration: 06m 20s) * 00:41 cjming@deploy1003: cjming: Continuing with deployment * 00:41 cjming@deploy1003: cjming: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:39 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] == 2026-08-03 == * 23:58 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply * 23:57 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply * 23:29 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cp5021.eqsin.wmnet * 23:29 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cp5021.eqsin.wmnet * 23:27 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cp5021.eqsin.wmnet * 23:26 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cp5021.eqsin.wmnet * 23:18 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 23:17 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 22:56 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: sync * 22:56 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: sync * 22:36 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 22:36 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 22:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host search-loader2002.codfw.wmnet with OS trixie * 21:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on search-loader2002.codfw.wmnet with reason: host reimage * 21:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on search-loader2002.codfw.wmnet with reason: host reimage * 21:42 dancy@deploy1003: Stopping before sync operations * 21:41 dancy@deploy1003: Started scap sync-world: testing * 21:39 dancy@deploy1003: Installation of scap version "4.277.0" completed for 3 hosts * 21:37 dancy@deploy1003: Installing scap version "4.277.0" for 3 host(s) * 21:37 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] (duration: 06m 13s) * 21:33 dancy@deploy1003: dancy: Continuing with deployment * 21:32 dancy@deploy1003: dancy: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:31 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] * 21:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host search-loader2002.codfw.wmnet with OS trixie * 21:03 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] (duration: 06m 34s) * 20:59 dancy@deploy1003: dancy: Continuing with deployment * 20:58 dancy@deploy1003: dancy: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:56 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] * 20:52 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] (duration: 06m 23s) * 20:48 cjming@deploy1003: cjming: Continuing with deployment * 20:47 cjming@deploy1003: cjming: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:46 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] * 20:42 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] (duration: 07m 36s) * 20:38 arlolra@deploy1003: arlolra: Continuing with deployment * 20:36 arlolra@deploy1003: arlolra: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:34 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] * 20:16 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] (duration: 08m 26s) * 20:12 krinkle@deploy1003: krinkle: Continuing with deployment * 20:09 krinkle@deploy1003: krinkle: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] * 19:45 jasmine@cumin2002: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-main-codfw * 18:58 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] (duration: 09m 23s) * 18:53 krinkle@deploy1003: krinkle: Continuing with deployment * 18:53 jasmine@cumin2002: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-main-codfw * 18:50 krinkle@deploy1003: krinkle: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:48 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] * 18:37 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] (duration: 10m 13s) * 18:34 dzahn@cumin2002: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host codesearch2001.codfw.wmnet * 18:34 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host codesearch2001.codfw.wmnet with OS trixie * 18:33 krinkle@deploy1003: krinkle: Continuing with deployment * 18:29 krinkle@deploy1003: krinkle: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:27 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] * 18:19 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 18:18 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on codesearch2001.codfw.wmnet with reason: host reimage * 18:14 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 18:14 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:12 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on codesearch2001.codfw.wmnet with reason: host reimage * 18:11 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 18:11 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 18:10 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 18:02 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 18:02 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 18:01 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 18:01 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 17:55 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host codesearch2001.codfw.wmnet with OS trixie * 17:54 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:54 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) codesearch2001.codfw.wmnet on all recursors * 17:53 dzahn@cumin2002: START - Cookbook sre.dns.wipe-cache codesearch2001.codfw.wmnet on all recursors * 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:48 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:41 dzahn@cumin2002: START - Cookbook sre.dns.netbox * 17:41 dzahn@cumin2002: START - Cookbook sre.ganeti.makevm for new host codesearch2001.codfw.wmnet * 17:37 dzahn@cumin2002: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host codesearch1001.eqiad.wmnet * 17:37 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host codesearch1001.eqiad.wmnet with OS trixie * 17:24 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on codesearch1001.eqiad.wmnet with reason: host reimage * 17:17 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on codesearch1001.eqiad.wmnet with reason: host reimage * 17:08 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host codesearch1001.eqiad.wmnet with OS trixie * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 17:06 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:06 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) codesearch1001.eqiad.wmnet on all recursors * 17:06 dzahn@cumin2002: START - Cookbook sre.dns.wipe-cache codesearch1001.eqiad.wmnet on all recursors * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 17:05 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 17:04 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:04 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 16:58 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 16:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2091.codfw.wmnet with OS trixie * 16:54 ebernhardson@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 16:54 ebernhardson@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 16:49 ebernhardson@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 16:49 ebernhardson@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 16:46 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1092.eqiad.wmnet with OS trixie * 16:43 ebernhardson@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 16:43 ebernhardson@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 16:43 dzahn@cumin2002: START - Cookbook sre.dns.netbox * 16:43 dzahn@cumin2002: START - Cookbook sre.ganeti.makevm for new host codesearch1001.eqiad.wmnet * 16:41 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 16:41 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 16:40 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2091.codfw.wmnet with reason: host reimage * 16:37 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 16:35 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2091.codfw.wmnet with reason: host reimage * 16:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1092.eqiad.wmnet with reason: host reimage * 16:24 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 16:24 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 16:23 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1092.eqiad.wmnet with reason: host reimage * 16:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2091.codfw.wmnet with OS trixie * 16:03 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1092.eqiad.wmnet with OS trixie * 16:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2090.codfw.wmnet with OS trixie * 15:51 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 15:51 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 15:51 jiji@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 15:50 jiji@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 15:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2090.codfw.wmnet with reason: host reimage * 15:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2090.codfw.wmnet with reason: host reimage * 15:33 jhathaway@dns1004: END - running authdns-update * 15:31 jhathaway@dns1004: START - running authdns-update * 15:26 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1091.eqiad.wmnet with OS trixie * 15:25 dancy@deploy1003: Installation of scap version "4.276.1" completed for 3 hosts * 15:23 dancy@deploy1003: Installing scap version "4.276.1" for 3 host(s) * 15:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2090.codfw.wmnet with OS trixie * 15:12 marostegui@cumin1003: dbctl commit (dc=all): 'Repool db2245, db2246, db2247 and db2248 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95857 and previous config saved to /var/cache/conftool/dbconfig/20260803-151212-marostegui.json * 15:09 dancy@deploy1003: Started scap sync-world: testing * 15:09 dancy@deploy1003: Installation of scap version "4.277.0" completed for 3 hosts * 15:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1091.eqiad.wmnet with reason: host reimage * 15:07 dancy@deploy1003: Installing scap version "4.277.0" for 3 host(s) * 15:03 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1091.eqiad.wmnet with reason: host reimage * 14:49 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1091.eqiad.wmnet with OS trixie * 14:33 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2089.codfw.wmnet with OS trixie * 14:29 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 14:27 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 14:18 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 14:16 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 14:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2089.codfw.wmnet with reason: host reimage * 14:10 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 14:10 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 14:09 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2089.codfw.wmnet with reason: host reimage * 13:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2089.codfw.wmnet with OS trixie * 13:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2088.codfw.wmnet with OS trixie * 13:40 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1090.eqiad.wmnet with OS trixie * 13:22 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1090.eqiad.wmnet with reason: host reimage * 13:22 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] (duration: 14m 34s) * 13:19 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1090.eqiad.wmnet with reason: host reimage * 13:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2088.codfw.wmnet with reason: host reimage * 13:16 aude@deploy1003: aude, mhorsey: Continuing with deployment * 13:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2088.codfw.wmnet with reason: host reimage * 13:12 aude@deploy1003: aude, mhorsey: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] * 13:05 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1090.eqiad.wmnet with OS trixie * 12:58 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2088.codfw.wmnet with OS trixie * 12:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db[2245-2247].codfw.wmnet * 12:49 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2247: Rebooting db2247.codfw.wmnet * 12:49 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2247: Rebooting db2247.codfw.wmnet * 12:42 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2246: Rebooting db2246.codfw.wmnet * 12:42 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2246: Rebooting db2246.codfw.wmnet * 12:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2087.codfw.wmnet with OS trixie * 12:37 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1089.eqiad.wmnet with OS trixie * 12:34 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2245: Rebooting db2245.codfw.wmnet * 12:34 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2245: Rebooting db2245.codfw.wmnet * 12:34 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db[2245-2247].codfw.wmnet * 12:32 kamila@deploy1003: Finished scap sync-world: rebuild after base image update (duration: 30m 26s) * 12:28 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 12:22 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2087.codfw.wmnet with reason: host reimage * 12:19 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1089.eqiad.wmnet with reason: host reimage * 12:14 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2087.codfw.wmnet with reason: host reimage * 12:14 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1089.eqiad.wmnet with reason: host reimage * 12:03 kamila@deploy1003: Started scap sync-world: rebuild after base image update * 12:00 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1089.eqiad.wmnet with OS trixie * 12:00 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2087.codfw.wmnet with OS trixie * 11:35 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db[2245-2248].codfw.wmnet * 11:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db[2245-2248].codfw.wmnet * 11:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2086.codfw.wmnet with OS trixie * 11:26 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 11:26 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 11:25 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 11:25 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 11:24 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1088.eqiad.wmnet with OS trixie * 11:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db[2245-2248].codfw.wmnet with reason: Checking network * 11:21 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 11:20 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 11:19 marostegui@dns1004: END - running authdns-update * 11:17 marostegui@dns1004: START - running authdns-update * 11:10 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:10 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 11:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2086.codfw.wmnet with reason: host reimage * 11:09 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:08 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 11:08 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:07 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 11:07 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:07 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 11:06 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop: apply * 11:06 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop: apply * 11:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1088.eqiad.wmnet with reason: host reimage * 11:05 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop: apply * 11:04 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop: apply * 11:04 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop: apply * 11:04 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop: apply * 11:02 marostegui@dns1004: END - running authdns-update * 11:02 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2086.codfw.wmnet with reason: host reimage * 11:01 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1088.eqiad.wmnet with reason: host reimage * 11:00 marostegui@dns1004: START - running authdns-update * 10:53 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] (duration: 10m 57s) * 10:51 cmooney@dns3003: END - running authdns-update * 10:49 cmooney@dns3003: START - running authdns-update * 10:47 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1088.eqiad.wmnet with OS trixie * 10:47 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2086.codfw.wmnet with OS trixie * 10:47 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 10:46 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:46 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:46 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new reverse ranges for eqsin CR switch links - cmooney@cumin1003" * 10:46 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new reverse ranges for eqsin CR switch links - cmooney@cumin1003" * 10:42 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] * 10:41 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 10:36 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2245, db2246 and db2247 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95855 and previous config saved to /var/cache/conftool/dbconfig/20260803-103652-marostegui.json * 10:35 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2248 from s4 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95854 and previous config saved to /var/cache/conftool/dbconfig/20260803-103535-marostegui.json * 10:27 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 10:27 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 10:26 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 10:24 kart_: cxserver: Add referencePunctuation config ([[phab:T97231|T97231]]) * 10:24 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 10:23 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:23 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:23 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:22 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:22 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply * 10:21 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply * 10:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2085.codfw.wmnet with OS trixie * 10:20 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply * 10:20 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply * 10:18 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply * 10:18 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply * 10:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1087.eqiad.wmnet with OS trixie * 09:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2085.codfw.wmnet with reason: host reimage * 09:43 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1087.eqiad.wmnet with reason: host reimage * 09:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2085.codfw.wmnet with reason: host reimage * 09:40 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1087.eqiad.wmnet with reason: host reimage * 09:26 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1087.eqiad.wmnet with OS trixie * 09:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2085.codfw.wmnet with OS trixie * 09:13 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2084.codfw.wmnet with OS trixie * 09:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1086.eqiad.wmnet with OS trixie * 08:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2084.codfw.wmnet with reason: host reimage * 08:50 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2084.codfw.wmnet with reason: host reimage * 08:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1086.eqiad.wmnet with reason: host reimage * 08:39 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1086.eqiad.wmnet with reason: host reimage * 08:38 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:38 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:37 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 08:37 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 08:35 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2084.codfw.wmnet with OS trixie * 08:34 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 08:34 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:27 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1086.eqiad.wmnet with OS trixie * 08:09 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1218: Repool after a crash * 08:07 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2083.codfw.wmnet with OS trixie * 08:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1085.eqiad.wmnet with OS trixie * 07:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2083.codfw.wmnet with reason: host reimage * 07:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1085.eqiad.wmnet with reason: host reimage * 07:40 kart_: Updated cxsever to 2026-07-16-140518-production ([[phab:T97231|T97231]]) * 07:39 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply * 07:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2083.codfw.wmnet with reason: host reimage * 07:38 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply * 07:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1085.eqiad.wmnet with reason: host reimage * 07:37 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] (duration: 32m 40s) * 07:33 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply * 07:33 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply * 07:25 jdlrobson@deploy1003: jdlrobson: Continuing with deployment * 07:24 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2083.codfw.wmnet with OS trixie * 07:24 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1085.eqiad.wmnet with OS trixie * 07:23 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1218: Repool after a crash * 07:21 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:09 marostegui: Drop renamed tables [[phab:T425074|T425074]] * 07:04 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] * 06:55 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply * 06:54 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 46s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-02 == * 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 01m 03s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-01 == * 03:30 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:30 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:30 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:30 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 34s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-31 == * 17:41 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 17:41 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 17:40 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 17:40 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 15:33 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2195: Testing * 15:02 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:02 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 15:02 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 14:48 pt1979@cumin2002: START - Cookbook sre.dns.netbox * 14:47 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2195: Testing * 14:22 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 14:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2195: Testing * 14:21 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 14:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2195.codfw.wmnet with reason: Testing * 14:16 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 14:04 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 14:04 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 13:30 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1048.eqiad.wmnet with OS trixie * 13:22 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2195: Testing * 13:22 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 13:19 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2195: Testing * 13:18 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 13:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2195: Testing * 13:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 13:05 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 13:05 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 13:04 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 13:04 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 12:50 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lswtest-d8-eqiad * 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:53 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:42 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 6515 * 11:37 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 6515 * 11:28 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:27 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:07 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2082.codfw.wmnet with OS trixie * 10:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2082.codfw.wmnet with reason: host reimage * 10:42 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2082.codfw.wmnet with reason: host reimage * 10:28 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2082.codfw.wmnet with OS trixie * 10:02 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 09:52 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 09:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts * 09:16 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts * 08:57 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 08:46 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:42 gkyziridis@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 08:37 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 08:37 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 08:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 08:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 08:11 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:11 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:08 filippo@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudvirt1048 * 08:07 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 08:07 filippo@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudvirt1048 * 08:06 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 08:01 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 08:00 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 07:19 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1048.eqiad.wmnet with reason: host reimage * 07:13 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1048.eqiad.wmnet with reason: host reimage * 07:11 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 07:11 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 07:09 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 07:09 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 06:57 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:56 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:48 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:48 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:44 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie * 06:34 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1048.eqiad.wmnet with OS trixie * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 54s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 00:57 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] (duration: 11m 04s) * 00:53 dreamyjazz@deploy1003: dreamyjazz, jforrester: Continuing with deployment * 00:48 dreamyjazz@deploy1003: dreamyjazz, jforrester: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:46 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] == 2026-07-30 == * 21:37 dancy@deploy1003: Installation of scap version "4.276.1" completed for 3 hosts * 21:35 dancy@deploy1003: Installing scap version "4.276.1" for 3 host(s) * 21:24 dancy@deploy1003: Installation of scap version "4.276.0" completed for 3 hosts * 21:22 dancy@deploy1003: Installing scap version "4.276.0" for 3 host(s) * 21:15 maryum: Deployed security fix for [[phab:T430601|T430601]] * 20:13 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] (duration: 09m 20s) * 20:07 arlolra@deploy1003: osleger, arlolra: Continuing with deployment * 20:05 arlolra@deploy1003: osleger, arlolra: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:03 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] * 19:29 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:29 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:25 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service * 19:24 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 19:24 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:24 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:24 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 19:23 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service * 19:20 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1084.eqiad.wmnet with OS trixie * 18:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1084.eqiad.wmnet with reason: host reimage * 18:52 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1084.eqiad.wmnet with reason: host reimage * 18:41 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 18:40 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 18:39 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1084.eqiad.wmnet with OS trixie * 18:25 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 18:15 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 17:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1083.eqiad.wmnet with OS trixie * 17:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1048: Maintenance * 17:36 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new security plugin settings - bking@cumin2003 - [[phab:T350516|T350516]] * 17:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1083.eqiad.wmnet with reason: host reimage * 17:26 inflatador: bking@apt1002 `reprepro --noskipold --component thirdparty/opensearch3 update trixie-wikimedia` [[phab:T433624|T433624]] * 17:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1083.eqiad.wmnet with reason: host reimage * 17:23 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 17:20 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 17:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2081.codfw.wmnet with OS trixie * 17:11 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new security plugin settings - bking@cumin2003 - [[phab:T350516|T350516]] * 17:10 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1083.eqiad.wmnet with OS trixie * 16:55 root@cumin1003: START - Cookbook sre.mysql.pool pool es1048: Maintenance * 16:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2081.codfw.wmnet with reason: host reimage * 16:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1048 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95833 and previous config saved to /var/cache/conftool/dbconfig/20260730-165053-cwilliams.json * 16:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1048.eqiad.wmnet with reason: Maintenance * 16:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1040: Maintenance * 16:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2081.codfw.wmnet with reason: host reimage * 16:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1082.eqiad.wmnet with OS trixie * 16:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2081.codfw.wmnet with OS trixie * 16:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2097.codfw.wmnet with OS trixie * 16:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1082.eqiad.wmnet with reason: host reimage * 16:08 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1082.eqiad.wmnet with reason: host reimage * 16:08 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 16:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1047: Maintenance * 16:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2080.codfw.wmnet with OS trixie * 16:04 root@cumin1003: START - Cookbook sre.mysql.pool pool es1040: Maintenance * 16:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1040: Maintenance * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new logging settings - bking@cumin2003 - [[phab:T324335|T324335]] * 15:58 root@cumin1003: START - Cookbook sre.mysql.pool pool es1040: Maintenance * 15:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1040 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95827 and previous config saved to /var/cache/conftool/dbconfig/20260730-155324-cwilliams.json * 15:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1040.eqiad.wmnet with reason: Maintenance * 15:50 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1082.eqiad.wmnet with OS trixie * 15:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2048: Maintenance * 15:44 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 24s) * 15:43 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2080.codfw.wmnet with reason: host reimage * 15:38 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new logging settings - bking@cumin2003 - [[phab:T324335|T324335]] * 15:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2080.codfw.wmnet with reason: host reimage * 15:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 15:30 mvernon@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be2097.codfw.wmnet with OS trixie * 15:23 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1081.eqiad.wmnet with OS trixie * 15:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2098.codfw.wmnet with OS trixie * 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - mvernon@cumin2003" * 15:18 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be2097.codfw.wmnet with OS trixie * 15:18 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - mvernon@cumin2003" * 15:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 15:17 root@cumin1003: START - Cookbook sre.mysql.pool pool es1047: Maintenance * 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2080.codfw.wmnet with OS trixie * 15:13 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be2097.codfw.wmnet with OS trixie * 15:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1047 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95820 and previous config saved to /var/cache/conftool/dbconfig/20260730-151200-cwilliams.json * 15:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1047.eqiad.wmnet with reason: Maintenance * 15:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: Maintenance * 15:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 15:04 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 15:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1081.eqiad.wmnet with reason: host reimage * 15:00 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1081.eqiad.wmnet with reason: host reimage * 15:00 root@cumin1003: START - Cookbook sre.mysql.pool pool es2048: Maintenance * 15:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 14:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2079.codfw.wmnet with OS trixie * 14:56 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 14:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2048 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95816 and previous config saved to /var/cache/conftool/dbconfig/20260730-145510-cwilliams.json * 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2048.codfw.wmnet with reason: Maintenance * 14:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2040: Maintenance * 14:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 14:51 tchin@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] (duration: 06m 48s) * 14:47 tchin@deploy1003: jforrester, tchin: Continuing with deployment * 14:47 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 14:47 tchin@deploy1003: jforrester, tchin: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:45 tchin@deploy1003: Started scap sync-world: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] * 14:42 sukhe@puppetserver1001: conftool action : set/weight=1; selector: cluster=urldownloader,service=squid * 14:42 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader,service=squid * 14:42 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1081.eqiad.wmnet with OS trixie * 14:39 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 14:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2079.codfw.wmnet with reason: host reimage * 14:36 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2098.codfw.wmnet with OS trixie * 14:32 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2079.codfw.wmnet with reason: host reimage * 14:30 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] (duration: 06m 31s) * 14:27 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 14:26 mszwarc@deploy1003: mszwarc: Continuing with deployment * 14:25 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:25 root@cumin1003: START - Cookbook sre.mysql.pool pool es1038: Maintenance * 14:25 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1038: Maintenance * 14:23 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] * 14:21 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] (duration: 11m 19s) * 14:20 root@cumin1003: START - Cookbook sre.mysql.pool pool es1038: Maintenance * 14:14 stran@deploy1003: stran: Continuing with deployment * 14:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1038 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95810 and previous config saved to /var/cache/conftool/dbconfig/20260730-141439-cwilliams.json * 14:14 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1038.eqiad.wmnet with reason: Maintenance * 14:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1036: Maintenance * 14:13 stran@deploy1003: stran: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2079.codfw.wmnet with OS trixie * 14:09 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] * 14:08 root@cumin1003: START - Cookbook sre.mysql.pool pool es2040: Maintenance * 14:08 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2040: Maintenance * 14:03 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] (duration: 31m 41s) * 14:03 root@cumin1003: START - Cookbook sre.mysql.pool pool es2040: Maintenance * 14:03 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2040 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95806 and previous config saved to /var/cache/conftool/dbconfig/20260730-135643-cwilliams.json * 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2040.codfw.wmnet with reason: Maintenance * 13:56 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:56 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2038: Maintenance * 13:55 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:52 lucaswerkmeister-wmde@deploy1003: migr, lucaswerkmeister-wmde: Continuing with deployment * 13:49 lucaswerkmeister-wmde@deploy1003: migr, lucaswerkmeister-wmde: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:49 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:48 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 13:45 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 13:32 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:32 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] * 13:28 root@cumin1003: START - Cookbook sre.mysql.pool pool es1036: Maintenance * 13:28 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1036: Maintenance * 13:22 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:22 root@cumin1003: START - Cookbook sre.mysql.pool pool es1036: Maintenance * 13:20 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1036 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95800 and previous config saved to /var/cache/conftool/dbconfig/20260730-131727-cwilliams.json * 13:17 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1036.eqiad.wmnet with reason: Maintenance * 13:17 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] (duration: 10m 31s) * 13:16 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2022\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 13:13 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, stran: Continuing with deployment * 13:10 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:10 root@cumin1003: START - Cookbook sre.mysql.pool pool es2038: Maintenance * 13:10 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2038: Maintenance * 13:08 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, stran: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie * 13:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2047: Maintenance * 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048 cloud-private - filippo@cumin1003" * 13:07 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048 cloud-private - filippo@cumin1003" * 13:06 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] * 13:04 root@cumin1003: START - Cookbook sre.mysql.pool pool es2038: Maintenance * 13:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2078.codfw.wmnet with OS trixie * 13:01 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2038 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95797 and previous config saved to /var/cache/conftool/dbconfig/20260730-125919-cwilliams.json * 12:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2038.codfw.wmnet with reason: Maintenance * 12:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2078.codfw.wmnet with reason: host reimage * 12:37 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2078.codfw.wmnet with reason: host reimage * 12:37 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] (duration: 06m 51s) * 12:33 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 12:32 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 12:32 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 12:32 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:30 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] * 12:19 root@cumin1003: START - Cookbook sre.mysql.pool pool es2047: Maintenance * 12:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2078.codfw.wmnet with OS trixie * 12:18 dcausse@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 12:18 dcausse@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 12:15 dcausse@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 12:14 dcausse@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 12:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2047 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95793 and previous config saved to /var/cache/conftool/dbconfig/20260730-121404-cwilliams.json * 12:13 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2047.codfw.wmnet with reason: Maintenance * 12:13 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2036: Maintenance * 12:05 ayounsi@dns1004: END - running authdns-update * 12:02 ayounsi@dns1004: START - running authdns-update * 11:51 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2077.codfw.wmnet with OS trixie * 11:48 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:46 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1080.eqiad.wmnet with OS trixie * 11:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1226: Maintenance * 11:41 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2077.codfw.wmnet with reason: host reimage * 11:28 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2077.codfw.wmnet with reason: host reimage * 11:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1080.eqiad.wmnet with reason: host reimage * 11:27 root@cumin1003: START - Cookbook sre.mysql.pool pool es2036: Maintenance * 11:27 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2036: Maintenance * 11:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1080.eqiad.wmnet with reason: host reimage * 11:21 root@cumin1003: START - Cookbook sre.mysql.pool pool es2036: Maintenance * 11:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2036 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95786 and previous config saved to /var/cache/conftool/dbconfig/20260730-111633-cwilliams.json * 11:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2036.codfw.wmnet with reason: Maintenance * 11:08 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2077.codfw.wmnet with OS trixie * 11:07 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 11:03 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie * 11:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1226: Maintenance * 10:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1226 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95783 and previous config saved to /var/cache/conftool/dbconfig/20260730-104801-cwilliams.json * 10:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1226.eqiad.wmnet with reason: Maintenance * 10:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1214: Maintenance * 10:27 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2035: Maintenance * 10:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2076.codfw.wmnet with OS trixie * 10:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1214: Maintenance * 09:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1214 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95775 and previous config saved to /var/cache/conftool/dbconfig/20260730-095451-cwilliams.json * 09:54 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1214.eqiad.wmnet with reason: Maintenance * 09:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1209: Maintenance * 09:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2076.codfw.wmnet with reason: host reimage * 09:42 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool es2035: Maintenance * 09:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.netbox.update-extras (exit_code=0) rolling restart_daemons on A:netbox * 09:41 ayounsi@cumin1003: START - Cookbook sre.netbox.update-extras rolling restart_daemons on A:netbox * 09:40 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2035: Maintenance * 09:39 ayounsi@cumin1003: END (PASS) - Cookbook sre.netbox.update-extras (exit_code=0) rolling restart_daemons on A:netbox-canary * 09:39 ayounsi@cumin1003: START - Cookbook sre.netbox.update-extras rolling restart_daemons on A:netbox-canary * 09:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2076.codfw.wmnet with reason: host reimage * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 09:34 root@cumin1003: START - Cookbook sre.mysql.pool pool es2035: Maintenance * 09:32 lucaswerkmeister-wmde@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 09:32 lucaswerkmeister-wmde@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 09:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2035 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95771 and previous config saved to /var/cache/conftool/dbconfig/20260730-092910-cwilliams.json * 09:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2035.codfw.wmnet with reason: Maintenance * 09:19 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2076.codfw.wmnet with OS trixie * 09:18 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 09:17 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie * 09:07 root@cumin1003: START - Cookbook sre.mysql.pool pool db1209: Maintenance * 09:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 23 hosts * 09:04 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Remove cable label from interfaces descriptions - ayounsi@cumin1003 * 09:04 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:02 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Remove cable label from interfaces descriptions - ayounsi@cumin1003 * 09:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1209 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95767 and previous config saved to /var/cache/conftool/dbconfig/20260730-090133-cwilliams.json * 09:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1209.eqiad.wmnet with reason: Maintenance * 09:01 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1192: Maintenance * 08:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1252: Maintenance * 08:57 jayme@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 08:56 jayme@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 08:53 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 23 hosts * 08:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:51 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:50 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1263: Maintenance * 08:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:23 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 08:15 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie * 08:14 root@cumin1003: START - Cookbook sre.mysql.pool pool db1192: Maintenance * 08:13 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2075.codfw.wmnet with OS trixie * 08:12 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1252: Maintenance * 08:11 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1252: Maintenance * 08:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1252: Maintenance * 08:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1192 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95754 and previous config saved to /var/cache/conftool/dbconfig/20260730-080611-cwilliams.json * 08:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1192.eqiad.wmnet with reason: Maintenance * 08:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1178: Maintenance * 08:05 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS bullseye * 07:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1263: Maintenance * 07:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1263 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95751 and previous config saved to /var/cache/conftool/dbconfig/20260730-075106-cwilliams.json * 07:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[1260-1262].eqiad.wmnet with reason: Maintenance * 07:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2075.codfw.wmnet with reason: host reimage * 07:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1263.eqiad.wmnet with reason: Maintenance * 07:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2075.codfw.wmnet with reason: host reimage * 07:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance * 07:38 dcausse@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:38 dcausse@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 07:35 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db1252', diff saved to https://phabricator.wikimedia.org/P95748 and previous config saved to /var/cache/conftool/dbconfig/20260730-073510-marostegui.json * 07:26 klausman@dns2004: END - running authdns-update * 07:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 07:25 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2075.codfw.wmnet with OS trixie * 07:24 klausman@dns2004: START - running authdns-update * 07:23 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host an-test-master1003.eqiad.wmnet * 07:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db1178: Maintenance * 07:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1178 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95746 and previous config saved to /var/cache/conftool/dbconfig/20260730-071112-cwilliams.json * 07:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1178.eqiad.wmnet with reason: Maintenance * 07:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1177: Maintenance * 06:24 root@cumin1003: START - Cookbook sre.mysql.pool pool db1177: Maintenance * 06:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1177 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95741 and previous config saved to /var/cache/conftool/dbconfig/20260730-061736-cwilliams.json * 06:17 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1177.eqiad.wmnet with reason: Maintenance * 06:17 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1172: Maintenance * 05:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1218.eqiad.wmnet with reason: crashed * 05:41 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db1217 it crashed', diff saved to https://phabricator.wikimedia.org/P95737 and previous config saved to /var/cache/conftool/dbconfig/20260730-054111-marostegui.json * 05:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95736 and previous config saved to /var/cache/conftool/dbconfig/20260730-053422-cwilliams.json * 05:30 root@cumin1003: START - Cookbook sre.mysql.pool pool db1172: Maintenance * 05:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95734 and previous config saved to /var/cache/conftool/dbconfig/20260730-052414-cwilliams.json * 05:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1172 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95733 and previous config saved to /var/cache/conftool/dbconfig/20260730-052354-cwilliams.json * 05:23 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1172.eqiad.wmnet with reason: Maintenance * 05:23 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1167: Maintenance * 05:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95731 and previous config saved to /var/cache/conftool/dbconfig/20260730-051406-cwilliams.json * 05:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95729 and previous config saved to /var/cache/conftool/dbconfig/20260730-050358-cwilliams.json * 04:47 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 04:47 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 04:47 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 04:35 root@cumin1003: START - Cookbook sre.mysql.pool pool db1167: Maintenance * 04:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1167 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95726 and previous config saved to /var/cache/conftool/dbconfig/20260730-042923-cwilliams.json * 04:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 04:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1167.eqiad.wmnet with reason: Maintenance * 04:22 pt1979@cumin2002: START - Cookbook sre.dns.netbox * 04:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95725 and previous config saved to /var/cache/conftool/dbconfig/20260730-040337-cwilliams.json * 04:03 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:38 brett@cumin2002: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool eqsin [reason: Switch upgrade maintenance window complete, [[phab:T433097|T433097]]] * 01:38 brett@cumin2002: START - Cookbook sre.dns.admin DNS admin: pool eqsin [reason: Switch upgrade maintenance window complete, [[phab:T433097|T433097]]] == 2026-07-29 == * 23:57 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin,mr1-eqsin IPv6,mr1-eqsin.oob,mr1-eqsin.oob IPv6 with reason: connection issue * 22:54 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 22:53 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 22:53 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 22:53 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:25 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2022.codfw.wmnet, repooling source-only afterwards * 22:20 brett@cumin2002: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool eqsin [reason: Switch upgrade maintenance window, [[phab:T433097|T433097]]] * 22:20 brett@cumin2002: START - Cookbook sre.dns.admin DNS admin: depool eqsin [reason: Switch upgrade maintenance window, [[phab:T433097|T433097]]] * 22:01 apine@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 22:00 apine@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 21:59 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 21:58 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 21:58 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 21:58 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 21:32 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1048.eqiad.wmnet with OS trixie * 21:25 pt1979@cumin2002: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be2097.codfw.wmnet with OS bullseye * 21:16 zabe@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki=metawiki 'Mental Health Resource Center' 'Safety Resource Center/Mental Health' Zabe --reason 'per request [[:phab:T433118{{!}}T433118]]' * 21:12 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2022.codfw.wmnet, repooling source-only afterwards * 21:12 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] (duration: 12m 53s) * 21:12 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2015\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 21:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1253: Maintenance * 21:08 aaron@deploy1003: aaron: Continuing with deployment * 21:01 aaron@deploy1003: aaron: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:59 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] * 20:52 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] (duration: 21m 57s) * 20:48 aaron@deploy1003: aaron: Continuing with deployment * 20:32 aaron@deploy1003: aaron: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:30 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] * 20:24 root@cumin1003: START - Cookbook sre.mysql.pool pool db1253: Maintenance * 20:19 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] (duration: 08m 07s) * 20:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1253 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95719 and previous config saved to /var/cache/conftool/dbconfig/20260729-201810-cwilliams.json * 20:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1253.eqiad.wmnet with reason: Maintenance * 20:17 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1231: Maintenance * 20:15 aaron@deploy1003: bpirkle, aaron: Continuing with deployment * 20:13 aaron@deploy1003: bpirkle, aaron: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:12 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie * 20:11 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] * 20:11 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1048.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:09 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1048.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:09 pt1979@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 20:04 pt1979@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 19:47 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 19:43 pt1979@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye * 19:41 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 19:37 zabe: zabe@deploy1003:~$ mwscript-k8s --comment='[[phab:T433529|T433529]]' --follow -- resetAuthenticationThrottle.php --wiki=aawiki --signup --ip=89.36.114.94 * 19:36 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] (duration: 06m 49s) * 19:32 zabe@deploy1003: zabe: Continuing with deployment * 19:31 zabe@deploy1003: zabe: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:31 root@cumin1003: START - Cookbook sre.mysql.pool pool db1231: Maintenance * 19:29 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] * 19:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1231 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95714 and previous config saved to /var/cache/conftool/dbconfig/20260729-192454-cwilliams.json * 19:24 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1231.eqiad.wmnet with reason: Maintenance * 19:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1227: Maintenance * 19:22 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 19:22 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 19:21 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:21 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1048] - vriley@cumin1003" * 19:21 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1048] - vriley@cumin1003" * 19:19 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 19:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1251: Maintenance * 19:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95711 and previous config saved to /var/cache/conftool/dbconfig/20260729-191756-cwilliams.json * 19:16 vriley@cumin1003: START - Cookbook sre.dns.netbox * 19:11 dduvall: rolling back wmf.13 to group0 due to [[phab:T433457|T433457]] (cc [[phab:T430832|T430832]]) * 19:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95709 and previous config saved to /var/cache/conftool/dbconfig/20260729-190748-cwilliams.json * 19:01 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2022.codfw.wmnet with OS bookworm * 19:01 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2021.codfw.wmnet, repooling source-only afterwards * 18:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95707 and previous config saved to /var/cache/conftool/dbconfig/20260729-185740-cwilliams.json * 18:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95704 and previous config saved to /var/cache/conftool/dbconfig/20260729-184732-cwilliams.json * 18:37 root@cumin1003: START - Cookbook sre.mysql.pool pool db1227: Maintenance * 18:34 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2022.codfw.wmnet with reason: host reimage * 18:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1227 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95701 and previous config saved to /var/cache/conftool/dbconfig/20260729-183117-cwilliams.json * 18:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1227.eqiad.wmnet with reason: Maintenance * 18:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1202: Maintenance * 18:30 root@cumin1003: START - Cookbook sre.mysql.pool pool db1251: Maintenance * 18:27 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2022.codfw.wmnet with reason: host reimage * 18:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1251 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95698 and previous config saved to /var/cache/conftool/dbconfig/20260729-182428-cwilliams.json * 18:24 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lvs2014.codfw.wmnet * 18:24 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for lvs2014.codfw.wmnet * 18:24 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1251.eqiad.wmnet with reason: Maintenance * 18:23 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1235: Maintenance * 18:22 brett@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 18:19 brett@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 18:19 brett@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 18:17 brett@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 18:17 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 18:16 mutante: removing jenkins during the train - living on the edge - no, just kidding, jenkins has migrated to dedicated machines, nothing should happen * 18:15 brett@cumin2002: END (ERROR) - Cookbook sre.loadbalancer.restart-pybal (exit_code=97) rolling-restart of pybal on P<nowiki>{</nowiki>lvs2014.codfw.wmnet<nowiki>}</nowiki> and A:lvs ([[phab:T428495|T428495]]) * 18:15 mutante: CI: contint1002/contint2002: apt-get remove --purge jenkins - jenkins be gone - [[phab:T418521|T418521]] * 18:13 brett@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on P<nowiki>{</nowiki>lvs2014.codfw.wmnet<nowiki>}</nowiki> and A:lvs ([[phab:T428495|T428495]]) * 18:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2022 * 18:08 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2022 * 18:03 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T428495|T428495]] * 18:03 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2022 * 18:02 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2022.codfw.wmnet 211.48.192.10.in-addr.arpa 1.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:02 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2022.codfw.wmnet 211.48.192.10.in-addr.arpa 1.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:02 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:02 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2022 - bking@cumin2003" * 18:02 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2022 - bking@cumin2003" * 17:57 bking@cumin2003: START - Cookbook sre.dns.netbox * 17:56 brett@cumin2002: END (FAIL) - Cookbook sre.loadbalancer.restart-pybal (exit_code=1) rolling-restart of pybal on A:lvs-codfw and A:lvs ([[phab:T428495|T428495]]) * 17:55 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo - [[phab:T428495|T428495]] * 17:54 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2022 * 17:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2022.codfw.wmnet with OS bookworm * 17:50 brett@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on A:lvs-codfw and A:lvs ([[phab:T428495|T428495]]) * 17:47 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2021.codfw.wmnet, repooling source-only afterwards * 17:47 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 14s) * 17:47 swfrench-wmf: authdns-update to direct codfw, eqsin, ulsfo etcd clients back to codfw - [[phab:T428495|T428495]] * 17:47 swfrench@dns1004: END - running authdns-update * 17:47 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 17:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95692 and previous config saved to /var/cache/conftool/dbconfig/20260729-174713-cwilliams.json * 17:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance * 17:46 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1249: Maintenance * 17:45 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2015\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 17:45 swfrench@dns1004: START - running authdns-update * 17:44 root@cumin1003: START - Cookbook sre.mysql.pool pool db1202: Maintenance * 17:41 akhatun: Deployed refinery using scap, then deployed onto hdfs * 17:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1202 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95688 and previous config saved to /var/cache/conftool/dbconfig/20260729-173759-cwilliams.json * 17:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1202.eqiad.wmnet with reason: Maintenance * 17:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1194: Maintenance * 17:37 root@cumin1003: START - Cookbook sre.mysql.pool pool db1235: Maintenance * 17:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1230: Maintenance * 17:30 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1235 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95684 and previous config saved to /var/cache/conftool/dbconfig/20260729-173051-cwilliams.json * 17:30 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1235.eqiad.wmnet with reason: Maintenance * 17:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1234: Maintenance * 17:26 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (thin): Regular analytics weekly train THIN [analytics/refinery@56695674] (duration: 02m 02s) * 17:24 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (thin): Regular analytics weekly train THIN [analytics/refinery@56695674] * 17:23 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567]: Regular analytics weekly train [analytics/refinery@56695674] (duration: 06m 20s) * 17:20 dancy@deploy1003: Finished scap sync-world: Testing delay_messageblobstore_purge: true (duration: 06m 29s) * 17:17 akhatun@deploy1003: Started deploy [analytics/refinery@5669567]: Regular analytics weekly train [analytics/refinery@56695674] * 17:17 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] (duration: 00m 22s) * 17:16 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] * 17:13 dancy@deploy1003: Started scap sync-world: Testing delay_messageblobstore_purge: true * 17:05 mutante: CI: contint1002/contint2002 - restarted httpd to be extra sure all is cleaned up - https://integration.wikimedia.org/ci/ is up and running [[phab:T418521|T418521]] * 17:04 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 17:03 mutante: CI: contint1002/contint2002 - rm /etc/apache2/jenkins_proxy - removing legacy jenkins proxy config - jenkins is on new dedicated machines and uses jenkins_proxy_ext config [[phab:T418521|T418521]] * 17:02 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] (duration: 36m 25s) * 17:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1249: Maintenance * 16:59 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 16:54 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2015.codfw.wmnet, repooling source-only afterwards * 16:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1249 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95674 and previous config saved to /var/cache/conftool/dbconfig/20260729-165339-cwilliams.json * 16:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1249.eqiad.wmnet with reason: Maintenance * 16:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1248: Maintenance * 16:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1194: Maintenance * 16:47 swfrench-wmf: silenced EtcdReplicationDown 57b2b421-1cc9-4e38-9276-{{Gerrit|94f223fd231c}} - [[phab:T428495|T428495]] * 16:46 tchin@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/eventstreams-internal: apply * 16:46 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye * 16:46 tchin@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/eventstreams-internal: apply * 16:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1230: Maintenance * 16:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1194 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95669 and previous config saved to /var/cache/conftool/dbconfig/20260729-164422-cwilliams.json * 16:44 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 16:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1194.eqiad.wmnet with reason: Maintenance * 16:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1191: Maintenance * 16:43 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host an-test-master1003.eqiad.wmnet * 16:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db1234: Maintenance * 16:43 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 16:43 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Rolling back deployment * 16:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host an-test-master1004.eqiad.wmnet * 16:41 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 16:40 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 16:40 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 16:39 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 16:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1230 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95667 and previous config saved to /var/cache/conftool/dbconfig/20260729-163932-cwilliams.json * 16:39 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 16:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1230.eqiad.wmnet with reason: Maintenance * 16:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1207: Maintenance * 16:38 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 16:37 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host an-test-master1004.eqiad.wmnet * 16:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1234 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95664 and previous config saved to /var/cache/conftool/dbconfig/20260729-163719-cwilliams.json * 16:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1234.eqiad.wmnet with reason: Maintenance * 16:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1079.eqiad.wmnet with OS trixie * 16:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1232: Maintenance * 16:34 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1259: Maintenance * 16:28 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] (duration: 06m 57s) * 16:28 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:26 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] * 16:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1051 hosts * 16:21 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] * 16:20 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2006.codfw.wmnet with OS bookworm * 16:19 akhatun: Deploying Refinery at {{Gerrit|56695674}} as part of weekly train * 16:18 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1079.eqiad.wmnet with reason: host reimage * 16:16 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] (duration: 15m 36s) * 16:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2021.codfw.wmnet with OS bookworm * 16:14 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1079.eqiad.wmnet with reason: host reimage * 16:12 topranks: hot-swap line card in FPC0 on cr1-eqiad with replacement MPC10E from Juniper [[phab:T426343|T426343]] * 16:10 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Continuing with deployment * 16:07 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db1248: Maintenance * 16:01 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] * 16:00 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 16:00 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 15:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1248 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95651 and previous config saved to /var/cache/conftool/dbconfig/20260729-155956-cwilliams.json * 15:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1248.eqiad.wmnet with reason: Maintenance * 15:59 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2006.codfw.wmnet with reason: host reimage * 15:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1247: Maintenance * 15:59 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2074.codfw.wmnet with OS trixie * 15:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1191: Maintenance * 15:57 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:55 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1079.eqiad.wmnet with OS trixie * 15:55 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2006.codfw.wmnet with reason: host reimage * 15:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1207: Maintenance * 15:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1191 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95646 and previous config saved to /var/cache/conftool/dbconfig/20260729-155104-cwilliams.json * 15:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1191.eqiad.wmnet with reason: Maintenance * 15:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1181: Maintenance * 15:49 root@cumin1003: START - Cookbook sre.mysql.pool pool db1232: Maintenance * 15:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2021.codfw.wmnet with reason: host reimage * 15:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1207 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95643 and previous config saved to /var/cache/conftool/dbconfig/20260729-154735-cwilliams.json * 15:47 root@cumin1003: START - Cookbook sre.mysql.pool pool db1259: Maintenance * 15:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1207.eqiad.wmnet with reason: Maintenance * 15:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1200: Maintenance * 15:46 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] (duration: 31m 59s) * 15:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 15:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1232 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95640 and previous config saved to /var/cache/conftool/dbconfig/20260729-154330-cwilliams.json * 15:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1232.eqiad.wmnet with reason: Maintenance * 15:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1219: Maintenance * 15:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2015.codfw.wmnet, repooling source-only afterwards * 15:41 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 18s) * 15:41 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1259 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95638 and previous config saved to /var/cache/conftool/dbconfig/20260729-154107-cwilliams.json * 15:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1259.eqiad.wmnet with reason: Maintenance * 15:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1254: Maintenance * 15:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2015.codfw.wmnet with OS bookworm * 15:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2021.codfw.wmnet with reason: host reimage * 15:36 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2006.codfw.wmnet with OS bookworm * 15:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 15:35 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Continuing with deployment * 15:33 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2074.codfw.wmnet with OS trixie * 15:33 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 15:32 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:29 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:28 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be2074.codfw.wmnet with OS trixie * 15:28 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2006.codfw.wmnet * 15:26 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1078.eqiad.wmnet with OS trixie * 15:25 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 15:25 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:22 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2006.codfw.wmnet * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2021 * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2021 * 15:19 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2021 * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2021.codfw.wmnet 210.48.192.10.in-addr.arpa 0.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:19 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2021.codfw.wmnet 210.48.192.10.in-addr.arpa 0.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2021 - bking@cumin2003" * 15:19 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2021 - bking@cumin2003" * 15:14 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] * 15:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2015.codfw.wmnet with reason: host reimage * 15:11 root@cumin1003: START - Cookbook sre.mysql.pool pool db1247: Maintenance * 15:11 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on ml-serve2004.codfw.wmnet with reason: [[phab:T433478|T433478]] * 15:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2015.codfw.wmnet with reason: host reimage * 15:10 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on ml-serve2002.codfw.wmnet with reason: [[phab:T433476|T433476]] * 15:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 15:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1247 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95625 and previous config saved to /var/cache/conftool/dbconfig/20260729-150459-cwilliams.json * 15:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1247.eqiad.wmnet with reason: Maintenance * 15:04 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:04 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1244: Maintenance * 15:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1078.eqiad.wmnet with reason: host reimage * 15:03 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1005.wikimedia.org * 15:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db1181: Maintenance * 15:01 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2098.codfw.wmnet with OS bullseye * 15:00 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye * 15:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1200: Maintenance * 14:59 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 14:59 root@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1285.eqiad.wmnet with OS trixie * 14:59 Amir1: mwscript-k8s -- extensions/TimedMediaHandler/maintenance/requeueTranscodes.php --wiki=commonswiki --key '360p.mpeg4.mov' --throttle --video --missing ([[phab:T358266|T358266]]) * 14:58 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1005.wikimedia.org * 14:58 jhancock@cumin2002: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['ms-be2098'] * 14:58 jhancock@cumin2002: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['ms-be2098'] * 14:58 jhancock@cumin2002: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['ms-be2097'] * 14:58 jhancock@cumin2002: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['ms-be2097'] * 14:58 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1078.eqiad.wmnet with reason: host reimage * 14:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1181 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95621 and previous config saved to /var/cache/conftool/dbconfig/20260729-145629-cwilliams.json * 14:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1181.eqiad.wmnet with reason: Maintenance * 14:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1174: Maintenance * 14:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1219: Maintenance * 14:55 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader1006.wikimedia.org on all recursors * 14:55 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader1006.wikimedia.org on all recursors * 14:55 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader1005.wikimedia.org on all recursors * 14:55 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader1005.wikimedia.org on all recursors * 14:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1200 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95618 and previous config saved to /var/cache/conftool/dbconfig/20260729-145336-cwilliams.json * 14:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1254: Maintenance * 14:53 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1200.eqiad.wmnet with reason: Maintenance * 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1185: Maintenance * 14:52 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2021 * 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2015 * 14:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2015 * 14:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1219 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95616 and previous config saved to /var/cache/conftool/dbconfig/20260729-144946-cwilliams.json * 14:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1219.eqiad.wmnet with reason: Maintenance * 14:49 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1218: Maintenance * 14:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2021.codfw.wmnet with OS bookworm * 14:48 dancy@deploy1003: Finished deploy [zuul/deploy@22703a6]: Deploying https://gerrit.wikimedia.org/r/c/integration/zuul/+/1311501 ([[phab:T432491|T432491]]) (duration: 00m 15s) * 14:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2015.codfw.wmnet with OS bookworm * 14:48 dancy@deploy1003: Started deploy [zuul/deploy@22703a6]: Deploying https://gerrit.wikimedia.org/r/c/integration/zuul/+/1311501 ([[phab:T432491|T432491]]) * 14:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1254 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95613 and previous config saved to /var/cache/conftool/dbconfig/20260729-144729-cwilliams.json * 14:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1254.eqiad.wmnet with reason: Maintenance * 14:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1233: Maintenance * 14:46 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2013\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 14:46 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2014\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 14:46 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:45 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:44 root@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1285.eqiad.wmnet with reason: host reimage * 14:43 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:42 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:41 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:40 root@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1285.eqiad.wmnet with reason: host reimage * 14:39 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1078.eqiad.wmnet with OS trixie * 14:39 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2074.codfw.wmnet with OS trixie * 14:32 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:32 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:32 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:31 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2005.codfw.wmnet with OS bookworm * 14:30 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:30 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:29 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:29 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:27 root@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host db1285 * 14:27 root@cumin1003: START - Cookbook sre.hosts.move-vlan for host db1285 * 14:27 root@cumin1003: START - Cookbook sre.hosts.reimage for host db1285.eqiad.wmnet with OS trixie * 14:24 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 14:24 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:24 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:24 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:23 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:22 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:22 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:21 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:17 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:16 root@cumin1003: START - Cookbook sre.mysql.pool pool db1244: Maintenance * 14:15 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:15 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add asw1-604 loopback ipv4 - pt1979@cumin2002" * 14:15 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add asw1-604 loopback ipv4 - pt1979@cumin2002" * 14:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:12 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 14:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95599 and previous config saved to /var/cache/conftool/dbconfig/20260729-141014-cwilliams.json * 14:10 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 14:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1244.eqiad.wmnet with reason: Maintenance * 14:10 pt1979@cumin2002: START - Cookbook sre.dns.netbox * 14:10 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 14:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1243: Maintenance * 14:09 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2005.codfw.wmnet with reason: host reimage * 14:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db1174: Maintenance * 14:08 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad * 14:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db1185: Maintenance * 14:06 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 14:05 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2005.codfw.wmnet with reason: host reimage * 14:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1174 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95595 and previous config saved to /var/cache/conftool/dbconfig/20260729-140309-cwilliams.json * 14:03 sukhe@dns1004: END - running authdns-update * 14:03 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1174.eqiad.wmnet with reason: Maintenance * 14:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1170: Maintenance * 14:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db1218: Maintenance * 14:01 sukhe@dns1004: START - running authdns-update * 14:00 sukhe@dns1004: START - running authdns-update * 13:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db1233: Maintenance * 13:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1185 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95592 and previous config saved to /var/cache/conftool/dbconfig/20260729-135925-cwilliams.json * 13:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1185.eqiad.wmnet with reason: Maintenance * 13:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1161: Maintenance * 13:58 sukhe@puppetserver1001: conftool action : set/pooled=true; selector: dnsdisc=urldownloader * 13:58 root@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1265.eqiad.wmnet with OS trixie * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1218 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95590 and previous config saved to /var/cache/conftool/dbconfig/20260729-135621-cwilliams.json * 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1218.eqiad.wmnet with reason: Maintenance * 13:55 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1206: Maintenance * 13:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2073.codfw.wmnet with OS trixie * 13:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1233 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95587 and previous config saved to /var/cache/conftool/dbconfig/20260729-135335-cwilliams.json * 13:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1233.eqiad.wmnet with reason: Maintenance * 13:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1229: Maintenance * 13:50 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/kartotherian: apply * 13:50 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service * 13:49 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:49 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/kartotherian: apply * 13:48 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 13:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1077.eqiad.wmnet with OS trixie * 13:47 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 13:46 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2005.codfw.wmnet with OS bookworm * 13:44 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 13:44 root@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1265.eqiad.wmnet with reason: host reimage * 13:40 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] (duration: 09m 22s) * 13:39 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:38 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:36 root@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1265.eqiad.wmnet with reason: host reimage * 13:35 stran@deploy1003: stran: Continuing with deployment * 13:33 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 13:32 stran@deploy1003: stran: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified t * 13:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2073.codfw.wmnet with reason: host reimage * 13:30 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ml-build1001.eqiad.wmnet * 13:30 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] * 13:29 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad * 13:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1077.eqiad.wmnet with reason: host reimage * 13:27 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 13:27 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:27 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad * 13:26 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2073.codfw.wmnet with reason: host reimage * 13:26 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] (duration: 07m 56s) * 13:25 klausman@cumin1003: START - Cookbook sre.hosts.reboot-single for host ml-build1001.eqiad.wmnet * 13:24 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 13:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>ml-serve101[2-5].eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 13:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1015.eqiad.wmnet * 13:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1015.eqiad.wmnet * 13:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1077.eqiad.wmnet with reason: host reimage * 13:23 root@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host db1265 * 13:23 root@cumin1003: START - Cookbook sre.hosts.move-vlan for host db1265 * 13:23 root@cumin1003: START - Cookbook sre.hosts.reimage for host db1265.eqiad.wmnet with OS trixie * 13:23 root@cumin1003: START - Cookbook sre.mysql.pool pool db1243: Maintenance * 13:22 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 13:22 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2005.codfw.wmnet * 13:22 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 13:22 samtar@deploy1003: dreamrimmer, samtar: Continuing with deployment * 13:20 samtar@deploy1003: dreamrimmer, samtar: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts an-test-master[1001-1002].eqiad.wmnet * 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-master[1001-1002].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 13:18 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1015.eqiad.wmnet * 13:18 sukhe@cumin1003: END (ERROR) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=97) for role: url_downloader@eqiad * 13:18 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 13:18 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] * 13:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95574 and previous config saved to /var/cache/conftool/dbconfig/20260729-131638-cwilliams.json * 13:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1243.eqiad.wmnet with reason: Maintenance * 13:16 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2005.codfw.wmnet * 13:16 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1242: Maintenance * 13:14 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] (duration: 07m 00s) * 13:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db1170: Maintenance * 13:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1015.eqiad.wmnet * 13:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1014.eqiad.wmnet * 13:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1014.eqiad.wmnet * 13:12 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1223: Maintenance * 13:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db1161: Maintenance * 13:10 samtar@deploy1003: anzx, samtar: Continuing with deployment * 13:09 samtar@deploy1003: anzx, samtar: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db1206: Maintenance * 13:08 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 13:07 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] * 13:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1170 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95566 and previous config saved to /var/cache/conftool/dbconfig/20260729-130730-cwilliams.json * 13:07 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1170.eqiad.wmnet with reason: Maintenance * 13:07 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:07 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt IPs new switches - cmooney@cumin1003" * 13:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1158: Maintenance * 13:06 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1014.eqiad.wmnet * 13:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1161 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95564 and previous config saved to /var/cache/conftool/dbconfig/20260729-130616-cwilliams.json * 13:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 13:06 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1077.eqiad.wmnet with OS trixie * 13:05 root@cumin1003: START - Cookbook sre.mysql.pool pool db1229: Maintenance * 13:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1161.eqiad.wmnet with reason: Maintenance * 13:05 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt IPs new switches - cmooney@cumin1003" * 13:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2073.codfw.wmnet with OS trixie * 13:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Maintenance * 13:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1206 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95562 and previous config saved to /var/cache/conftool/dbconfig/20260729-130258-cwilliams.json * 13:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1206.eqiad.wmnet with reason: Maintenance * 13:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1196: Maintenance * 13:01 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 13:01 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:00 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 13:00 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 12:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1229 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95559 and previous config saved to /var/cache/conftool/dbconfig/20260729-125950-cwilliams.json * 12:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1229.eqiad.wmnet with reason: Maintenance * 12:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1222: Maintenance * 12:57 sukhe: sudo cumin 'A:lvs and (A:eqiad or A:codfw)' 'disable-puppet "adding new service urldownloader"': [[phab:T429175|T429175]] * 12:56 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1014.eqiad.wmnet * 12:56 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1013.eqiad.wmnet * 12:56 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1013.eqiad.wmnet * 12:50 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1013.eqiad.wmnet * 12:50 sukhe: sudo cumin 'O:url_downloader' 'run-puppet-agent --enable "merging CR 1313948"': [[phab:T429175|T429175]] * 12:48 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-master[1001-1002].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 12:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1013.eqiad.wmnet * 12:45 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1012.eqiad.wmnet * 12:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1012.eqiad.wmnet * 12:45 sukhe: sudo cumin 'O:url_downloader' 'disable-puppet "merging CR 1313948"': [[phab:T429175|T429175]] * 12:44 btullis@cumin1003: START - Cookbook sre.dns.netbox * 12:40 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test2001.codfw.wmnet * 12:40 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test2001.codfw.wmnet * 12:38 ayounsi@dns1004: END - running authdns-update * 12:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1012.eqiad.wmnet * 12:35 ayounsi@dns1004: START - running authdns-update * 12:34 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts an-test-master[1001-1002].eqiad.wmnet * 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts an-test-coord1001.eqiad.wmnet * 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-coord1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 12:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1012.eqiad.wmnet * 12:32 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>ml-serve101[2-5].eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 12:29 root@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Maintenance * 12:25 root@cumin1003: START - Cookbook sre.mysql.pool pool db1223: Maintenance * 12:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95544 and previous config saved to /var/cache/conftool/dbconfig/20260729-122254-cwilliams.json * 12:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1242.eqiad.wmnet with reason: Maintenance * 12:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1241: Maintenance * 12:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1051 hosts * 12:20 root@cumin1003: START - Cookbook sre.mysql.pool pool db1158: Maintenance * 12:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1223 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95540 and previous config saved to /var/cache/conftool/dbconfig/20260729-121937-cwilliams.json * 12:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1223.eqiad.wmnet with reason: Maintenance * 12:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1212: Maintenance * 12:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db1159: Maintenance * 12:17 elukey@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: sync * 12:15 elukey@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: sync * 12:15 root@cumin1003: START - Cookbook sre.mysql.pool pool db1196: Maintenance * 12:14 Daimona: Creating new DB tables for the CampaignEvents extension in x1.testwiki, x1.test2wiki, x1.officewiki, and x1.wikishared # [[phab:T429339|T429339]] * 12:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db1222: Maintenance * 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95535 and previous config saved to /var/cache/conftool/dbconfig/20260729-121211-cwilliams.json * 12:12 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 12:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1158.eqiad.wmnet with reason: Maintenance * 12:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1159 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95534 and previous config saved to /var/cache/conftool/dbconfig/20260729-121146-cwilliams.json * 12:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1159.eqiad.wmnet with reason: Maintenance * 12:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1196 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95533 and previous config saved to /var/cache/conftool/dbconfig/20260729-120847-cwilliams.json * 12:08 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 12:08 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1196.eqiad.wmnet with reason: Maintenance * 12:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1195: Maintenance * 12:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1222 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95530 and previous config saved to /var/cache/conftool/dbconfig/20260729-120424-cwilliams.json * 12:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1222.eqiad.wmnet with reason: Maintenance * 12:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1098 hosts * 12:00 marostegui: Rename tables [[phab:T425074|T425074]] * 12:00 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-coord1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 11:58 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1197: Maintenance * 11:55 btullis@cumin1003: START - Cookbook sre.dns.netbox * 11:52 elukey@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: sync * 11:51 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:51 elukey@deploy1003: helmfile [codfw] START helmfile.d/services/proton: sync * 11:51 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:50 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts an-test-coord1001.eqiad.wmnet * 11:50 elukey@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: sync * 11:49 elukey@deploy1003: helmfile [staging] START helmfile.d/services/proton: sync * 11:35 root@cumin1003: START - Cookbook sre.mysql.pool pool db1241: Maintenance * 11:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db1212: Maintenance * 11:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1241 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95520 and previous config saved to /var/cache/conftool/dbconfig/20260729-112918-cwilliams.json * 11:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1241.eqiad.wmnet with reason: Maintenance * 11:29 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1238: Maintenance * 11:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1212 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95517 and previous config saved to /var/cache/conftool/dbconfig/20260729-112727-cwilliams.json * 11:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 11:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1212.eqiad.wmnet with reason: Maintenance * 11:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1198: Maintenance * 11:23 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:22 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:21 root@cumin1003: START - Cookbook sre.mysql.pool pool db1195: Maintenance * 11:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1195 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95514 and previous config saved to /var/cache/conftool/dbconfig/20260729-111450-cwilliams.json * 11:14 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1195.eqiad.wmnet with reason: Maintenance * 11:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1186: Maintenance * 11:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 11:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 11:05 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:54 marostegui: Dropping renamed tables [[phab:T425066|T425066]] * 10:41 root@cumin1003: START - Cookbook sre.mysql.pool pool db1238: Maintenance * 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1198: Maintenance * 10:39 Amir1: ran https://phabricator.wikimedia.org/T432509#12149723 in production ([[phab:T432509|T432509]]) * 10:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db1197: Maintenance * 10:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1238 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95501 and previous config saved to /var/cache/conftool/dbconfig/20260729-103532-cwilliams.json * 10:35 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1238.eqiad.wmnet with reason: Maintenance * 10:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1221: Maintenance * 10:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1198 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95499 and previous config saved to /var/cache/conftool/dbconfig/20260729-103330-cwilliams.json * 10:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1198.eqiad.wmnet with reason: Maintenance * 10:33 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1175: Maintenance * 10:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1197 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95496 and previous config saved to /var/cache/conftool/dbconfig/20260729-103217-cwilliams.json * 10:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1197.eqiad.wmnet with reason: Maintenance * 10:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1188: Maintenance * 10:27 root@cumin1003: START - Cookbook sre.mysql.pool pool db1186: Maintenance * 10:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1186 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95493 and previous config saved to /var/cache/conftool/dbconfig/20260729-102111-cwilliams.json * 10:21 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1186.eqiad.wmnet with reason: Maintenance * 10:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 10:14 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 09:53 XioNoX: reboot cr2-magru - [[phab:T431750|T431750]] * 09:52 XioNoX: drain cr2-magru - [[phab:T431750|T431750]] * 09:48 root@cumin1003: START - Cookbook sre.mysql.pool pool db1221: Maintenance * 09:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zookeeper-test1002.eqiad.wmnet * 09:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1188: Maintenance * 09:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1175: Maintenance * 09:44 btullis@dns1004: END - running authdns-update * 09:42 btullis@dns1004: START - running authdns-update * 09:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1221 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95483 and previous config saved to /var/cache/conftool/dbconfig/20260729-094200-cwilliams.json * 09:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 7 hosts with reason: Maintenance * 09:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1221.eqiad.wmnet with reason: Maintenance * 09:41 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host zookeeper-test1002.eqiad.wmnet * 09:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1199: Maintenance * 09:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1188 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95481 and previous config saved to /var/cache/conftool/dbconfig/20260729-093917-cwilliams.json * 09:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1188.eqiad.wmnet with reason: Maintenance * 09:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1182: Maintenance * 09:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1175 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95479 and previous config saved to /var/cache/conftool/dbconfig/20260729-093842-cwilliams.json * 09:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1175.eqiad.wmnet with reason: Maintenance * 09:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1166: Maintenance * 09:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1169: Maintenance * 09:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1033.eqiad.wmnet,service=s8 * 09:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1033.eqiad.wmnet,service=s5 * 09:33 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1033.eqiad.wmnet,service=s8 * 09:33 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1033.eqiad.wmnet,service=s5 * 09:21 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:21 XioNoX: reboot cr1-magru - [[phab:T431750|T431750]] * 09:17 XioNoX: drain cr1-magru - [[phab:T431750|T431750]] * 09:15 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm1001.wikimedia.org * 09:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply * 09:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply * 09:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 09:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 09:11 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr2-magru,cr2-magru IPv6,cr2-magru.mgmt with reason: router upgrade * 09:11 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm1001.wikimedia.org * 09:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp1005.wikimedia.org * 09:07 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp1005.wikimedia.org * 09:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp2005.wikimedia.org * 09:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp2005.wikimedia.org * 09:00 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr1-magru,cr1-magru IPv6,cr1-magru.mgmt with reason: router upgrade * 09:00 marostegui: Dropping renamed tables [[phab:T426341|T426341]] * 08:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1199: Maintenance * 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1182: Maintenance * 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1166: Maintenance * 08:47 root@cumin1003: START - Cookbook sre.mysql.pool pool db1169: Maintenance * 08:46 ayounsi@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 1:00:00 on cr1-magru,cr1-magru IPv6,cr1-magru.mgmt with reason: router upgrade * 08:45 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2072.codfw.wmnet with OS trixie * 08:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1199 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95464 and previous config saved to /var/cache/conftool/dbconfig/20260729-084534-cwilliams.json * 08:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1199.eqiad.wmnet with reason: Maintenance * 08:45 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 08:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1190: Maintenance * 08:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1182 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95462 and previous config saved to /var/cache/conftool/dbconfig/20260729-084436-cwilliams.json * 08:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1182.eqiad.wmnet with reason: Maintenance * 08:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1156: Maintenance * 08:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1166 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95460 and previous config saved to /var/cache/conftool/dbconfig/20260729-084400-cwilliams.json * 08:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1166.eqiad.wmnet with reason: Maintenance * 08:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1157: Maintenance * 08:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95458 and previous config saved to /var/cache/conftool/dbconfig/20260729-084147-cwilliams.json * 08:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1169.eqiad.wmnet with reason: Maintenance * 08:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1163: Maintenance * 08:30 btullis@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 11 hosts with reason: Replacing the namenodes * 08:23 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2072.codfw.wmnet with reason: host reimage * 08:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1098 hosts * 08:19 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2072.codfw.wmnet with reason: host reimage * 07:58 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2072.codfw.wmnet with OS trixie * 07:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1190: Maintenance * 07:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1156: Maintenance * 07:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1157: Maintenance * 07:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1163: Maintenance * 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1190 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95444 and previous config saved to /var/cache/conftool/dbconfig/20260729-074930-cwilliams.json * 07:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1190.eqiad.wmnet with reason: Maintenance * 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1157 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95443 and previous config saved to /var/cache/conftool/dbconfig/20260729-074914-cwilliams.json * 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1156 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95442 and previous config saved to /var/cache/conftool/dbconfig/20260729-074906-cwilliams.json * 07:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1157.eqiad.wmnet with reason: Maintenance * 07:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 07:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1156.eqiad.wmnet with reason: Maintenance * 07:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1163 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95441 and previous config saved to /var/cache/conftool/dbconfig/20260729-074652-cwilliams.json * 07:46 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1163.eqiad.wmnet with reason: Maintenance * 07:46 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2034.codfw.wmnet * 07:42 ayounsi@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2034.codfw.wmnet * 07:42 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2232.codfw.wmnet with OS trixie * 07:34 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:34 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:33 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:31 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:19 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2232.codfw.wmnet with reason: host reimage * 07:15 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2232.codfw.wmnet with reason: host reimage * 06:58 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db2232.codfw.wmnet with OS trixie * 06:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[2160,2232].codfw.wmnet with reason: Reimage * 06:26 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1164.eqiad.wmnet with OS trixie * 06:05 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1164.eqiad.wmnet with reason: host reimage * 06:01 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1164.eqiad.wmnet with reason: host reimage * 05:47 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1164.eqiad.wmnet with OS trixie * 05:46 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1164.eqiad.wmnet with reason: Reimage == 2026-07-28 == * 22:50 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1138.eqiad.wmnet * 22:50 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1138.eqiad.wmnet * 22:49 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1138.eqiad.wmnet * 22:11 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2014.codfw.wmnet, repooling source-only afterwards * 22:08 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2013.codfw.wmnet, repooling source-only afterwards * 22:03 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 20:58 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] (duration: 08m 19s) * 20:55 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2014.codfw.wmnet, repooling source-only afterwards * 20:55 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2013.codfw.wmnet, repooling source-only afterwards * 20:54 arlolra@deploy1003: arlolra: Continuing with deployment * 20:54 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 14s) * 20:54 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 20:53 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 30s) * 20:53 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 20:52 arlolra@deploy1003: arlolra: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:51 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:50 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] * 20:49 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:43 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 20:34 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] (duration: 06m 54s) * 20:34 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:34 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:31 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:31 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:30 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:30 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:30 arlolra@deploy1003: arlolra: Continuing with deployment * 20:29 arlolra@deploy1003: arlolra: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:27 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] * 20:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2014.codfw.wmnet with OS bookworm * 20:21 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 20:21 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:20 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 20:19 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:19 swfrench-wmf: switched etcd-mirror replication from conf2005 to conf2004 - [[phab:T428495|T428495]] * 20:17 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:17 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:15 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] (duration: 08m 26s) * 20:12 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:11 arlolra@deploy1003: anzx, arlolra: Continuing with deployment * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2013.codfw.wmnet with OS bookworm * 20:09 arlolra@deploy1003: anzx, arlolra: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] * 20:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2216: Maintenance * 19:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2014.codfw.wmnet with reason: host reimage * 19:57 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:54 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:54 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:53 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:52 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2014.codfw.wmnet with reason: host reimage * 19:49 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2013.codfw.wmnet with reason: host reimage * 19:42 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 19:41 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:41 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2013.codfw.wmnet with reason: host reimage * 19:39 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:39 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:39 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-eqiad: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 19:38 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2014 * 19:33 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2014 * 19:29 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2014.codfw.wmnet with OS bookworm * 19:28 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:27 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1006 * 19:26 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2012\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 19:26 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1006 * 19:26 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:26 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 19:25 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2013 * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2013 * 19:21 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2013 * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2013.codfw.wmnet 84.0.192.10.in-addr.arpa 4.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:21 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2013.codfw.wmnet 84.0.192.10.in-addr.arpa 4.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2013 - bking@cumin2003" * 19:21 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2013 - bking@cumin2003" * 19:21 vriley@cumin1003: START - Cookbook sre.dns.netbox * 19:20 root@cumin1003: START - Cookbook sre.mysql.pool pool db2216: Maintenance * 19:13 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2216 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95435 and previous config saved to /var/cache/conftool/dbconfig/20260728-191343-cwilliams.json * 19:13 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2216.codfw.wmnet with reason: Maintenance * 19:13 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2203: Maintenance * 19:06 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1005.eqiad.wmnet with OS trixie * 19:06 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 19:06 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 18:46 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 18:45 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:45 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:43 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:40 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:36 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-eqiad: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 18:35 dancy@deploy1003: Installation of scap version "4.275.0" completed for 3 hosts * 18:33 dancy@deploy1003: Installing scap version "4.275.0" for 3 host(s) * 18:32 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:32 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2097.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:30 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns3003.wikimedia.org [reason: pool for all services after reimaging] * 18:29 sukhe@dns1004: END - running authdns-update * 18:27 sukhe@dns1004: START - running authdns-update * 18:27 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns3003.wikimedia.org,service=authdns-update [reason: pool authdns-update after reimaging] * 18:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db2203: Maintenance * 18:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2203 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95430 and previous config saved to /var/cache/conftool/dbconfig/20260728-181958-cwilliams.json * 18:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2203.codfw.wmnet with reason: Maintenance * 18:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2188: Maintenance * 18:18 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2097.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:17 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-be2098 * 18:17 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host ms-be2098 * 18:17 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-be2097 * 18:16 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host ms-be2097 * 18:15 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:15 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding ms-be2097-8 to codfw - jhancock@cumin2002" * 18:15 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding ms-be2097-8 to codfw - jhancock@cumin2002" * 18:10 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 18:08 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage * 18:05 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns3003.wikimedia.org with OS trixie * 18:03 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage * 17:56 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-codfw: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 17:45 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie * 17:45 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1005.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:41 sukhe@dns1004: END - running authdns-update * 17:39 sukhe@dns1004: START - running authdns-update * 17:36 sukhe@puppetserver1001: conftool action : set/weight=1; selector: cluster=urldownloader,service=squid * 17:36 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1005.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:35 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader,service=squid * 17:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 17:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1005 * 17:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 17:34 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1005 * 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1005] - vriley@cumin1003" * 17:34 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1005] - vriley@cumin1003" * 17:32 root@cumin1003: START - Cookbook sre.mysql.pool pool db2188: Maintenance * 17:29 vriley@cumin1003: START - Cookbook sre.dns.netbox * 17:29 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 17:26 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2188 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95425 and previous config saved to /var/cache/conftool/dbconfig/20260728-172609-cwilliams.json * 17:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2188.codfw.wmnet with reason: Maintenance * 17:25 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2176: Maintenance * 17:19 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1138.eqiad.wmnet with OS trixie * 17:18 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1005 * 17:18 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1005 * 17:18 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:15 vriley@cumin1003: START - Cookbook sre.dns.netbox * 17:13 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns3003.wikimedia.org with reason: host reimage * 17:07 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns3003.wikimedia.org with reason: host reimage * 17:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-eqiad * 17:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1015.eqiad.wmnet * 17:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1015.eqiad.wmnet * 16:59 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1138.eqiad.wmnet with reason: host reimage * 16:55 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-codfw: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 16:54 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1138.eqiad.wmnet with reason: host reimage * 16:53 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1015.eqiad.wmnet * 16:43 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns3003.wikimedia.org with OS trixie * 16:43 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1015.eqiad.wmnet * 16:43 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1014.eqiad.wmnet * 16:43 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1014.eqiad.wmnet * 16:43 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=dns3003.wikimedia.org [reason: depooling for reimage to trixie] * 16:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 290 hosts * 16:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db2176: Maintenance * 16:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1138 * 16:38 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1138 * 16:37 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1138 * 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1138.eqiad.wmnet 193.32.64.10.in-addr.arpa 3.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:37 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1138.eqiad.wmnet 193.32.64.10.in-addr.arpa 3.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1138 - jiji@cumin1003" * 16:37 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1138 - jiji@cumin1003" * 16:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1014.eqiad.wmnet * 16:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2176 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95420 and previous config saved to /var/cache/conftool/dbconfig/20260728-163235-cwilliams.json * 16:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2176.codfw.wmnet with reason: Maintenance * 16:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1014.eqiad.wmnet * 16:32 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1013.eqiad.wmnet * 16:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1013.eqiad.wmnet * 16:32 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2174: Maintenance * 16:28 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2012.codfw.wmnet, repooling source-only afterwards * 16:25 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1013.eqiad.wmnet * 16:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1013.eqiad.wmnet * 16:20 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1012.eqiad.wmnet * 16:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1012.eqiad.wmnet * 16:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1012.eqiad.wmnet * 16:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1012.eqiad.wmnet * 16:03 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1011.eqiad.wmnet * 16:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1011.eqiad.wmnet * 16:00 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2004.codfw.wmnet with OS bookworm * 15:59 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1011.eqiad.wmnet * 15:56 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:55 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 15:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:54 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1011.eqiad.wmnet * 15:54 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1010.eqiad.wmnet * 15:54 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1010.eqiad.wmnet * 15:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:50 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 15:49 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1010.eqiad.wmnet * 15:48 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 15:48 jiji@cumin1003: START - Cookbook sre.dns.netbox * 15:46 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2248: Maintenance * 15:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db2174: Maintenance * 15:44 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1010.eqiad.wmnet * 15:44 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1009.eqiad.wmnet * 15:44 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1009.eqiad.wmnet * 15:42 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1138 * 15:41 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1138.eqiad.wmnet with OS trixie * 15:39 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1009.eqiad.wmnet * 15:39 robh@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on arclamp2001.codfw.wmnet with reason: ram upgrade * 15:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2174 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95413 and previous config saved to /var/cache/conftool/dbconfig/20260728-153844-cwilliams.json * 15:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2174.codfw.wmnet with reason: Maintenance * 15:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2173: Maintenance * 15:37 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 15:35 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 15:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1009.eqiad.wmnet * 15:34 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1008.eqiad.wmnet * 15:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1008.eqiad.wmnet * 15:31 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 15:31 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 15:29 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1008.eqiad.wmnet * 15:27 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 290 hosts * 15:25 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2013 * 15:25 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2195: Maintenance * 15:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1008.eqiad.wmnet * 15:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1007.eqiad.wmnet * 15:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1007.eqiad.wmnet * 15:22 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2004.codfw.wmnet with reason: host reimage * 15:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2013.codfw.wmnet with OS bookworm * 15:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 15:19 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2004.codfw.wmnet with reason: host reimage * 15:19 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 15:17 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1007.eqiad.wmnet * 15:12 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1007.eqiad.wmnet * 15:12 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1006.eqiad.wmnet * 15:12 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1006.eqiad.wmnet * 15:11 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1138.eqiad.wmnet * 15:11 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 15:11 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1138.eqiad.wmnet * 15:11 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1138.eqiad.wmnet * 15:10 brennen@deploy1003: Finished deploy [phabricator/deployment@f8b349f]: deploy phab1004 for [[phab:T433382|T433382]] (duration: 00m 43s) * 15:10 brennen@deploy1003: Started deploy [phabricator/deployment@f8b349f]: deploy phab1004 for [[phab:T433382|T433382]] * 15:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts * 15:09 brennen@deploy1003: Finished deploy [phabricator/deployment@f8b349f]: deploy phab2003 for [[phab:T433382|T433382]] (duration: 00m 55s) * 15:08 brennen@deploy1003: Started deploy [phabricator/deployment@f8b349f]: deploy phab2003 for [[phab:T433382|T433382]] * 15:07 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2012.codfw.wmnet, repooling source-only afterwards * 15:07 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts * 15:06 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1137.eqiad.wmnet * 15:06 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1137.eqiad.wmnet * 15:06 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1137.eqiad.wmnet * 15:05 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1006.eqiad.wmnet * 15:05 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1022\.eqiad\.wmnet,dc=eqiad,cluster=wdqs\-main,service=wdqs\-main * 15:01 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab2003.codfw.wmnet with reason: deployment * 15:01 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1005.eqiad.wmnet with reason: deployment * 15:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1006.eqiad.wmnet * 15:00 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1006.eqiad.wmnet with reason: deployment * 15:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1005.eqiad.wmnet * 15:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1005.eqiad.wmnet * 14:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db2248: Maintenance * 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1004.eqiad.wmnet with reason: deployment * 14:59 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2004.codfw.wmnet with OS bookworm * 14:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1005.eqiad.wmnet * 14:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2248 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95403 and previous config saved to /var/cache/conftool/dbconfig/20260728-145532-cwilliams.json * 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2245-2247].codfw.wmnet with reason: Maintenance * 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2248.codfw.wmnet with reason: Maintenance * 14:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2240: Maintenance * 14:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2173: Maintenance * 14:51 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts * 14:50 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1005.eqiad.wmnet * 14:50 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1004.eqiad.wmnet * 14:50 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1004.eqiad.wmnet * 14:49 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts * 14:45 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2004.codfw.wmnet * 14:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2173 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95399 and previous config saved to /var/cache/conftool/dbconfig/20260728-144453-cwilliams.json * 14:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2173.codfw.wmnet with reason: Maintenance * 14:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2170: Maintenance * 14:44 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1004.eqiad.wmnet * 14:39 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2004.codfw.wmnet * 14:38 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1004.eqiad.wmnet * 14:38 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet * 14:38 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet * 14:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db2195: Maintenance * 14:36 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1022.eqiad.wmnet, repooling source-only afterwards * 14:36 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 14:33 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet * 14:33 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2222: Maintenance * 14:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2195 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95394 and previous config saved to /var/cache/conftool/dbconfig/20260728-143218-cwilliams.json * 14:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2195.codfw.wmnet with reason: Maintenance * 14:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2181: Maintenance * 14:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 14:30 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 14:25 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 14:25 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 14:25 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 14:23 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet * 14:23 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1002.eqiad.wmnet * 14:23 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1002.eqiad.wmnet * 14:23 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 19s) * 14:23 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:18 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1002.eqiad.wmnet * 14:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2012.codfw.wmnet with OS bookworm * 14:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1002.eqiad.wmnet * 14:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1001.eqiad.wmnet * 14:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1001.eqiad.wmnet * 14:11 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 14:11 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 14:08 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1001.eqiad.wmnet * 14:07 elukey@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'. * 14:07 elukey@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'. * 14:06 elukey@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'. * 14:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db2240: Maintenance * 14:06 XioNoX: un-drain cr2-esams - [[phab:T431751|T431751]] * 14:05 elukey@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'. * 14:02 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1001.eqiad.wmnet * 14:02 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-eqiad * 14:01 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 14:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2240 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95384 and previous config saved to /var/cache/conftool/dbconfig/20260728-140011-cwilliams.json * 14:00 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2240.codfw.wmnet with reason: Maintenance * 13:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2237: Maintenance * 13:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db2170: Maintenance * 13:56 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 13:55 XioNoX: reboot cr2-esams - [[phab:T431751|T431751]] * 13:52 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 13:51 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr2-esams,cr2-esams IPv6,cr2-esams.mgmt with reason: router upgrade * 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 13:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2170 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95381 and previous config saved to /var/cache/conftool/dbconfig/20260728-135043-cwilliams.json * 13:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2170.codfw.wmnet with reason: Maintenance * 13:50 XioNoX: drain cr2-esams - [[phab:T431751|T431751]] * 13:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2153: Maintenance * 13:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2012.codfw.wmnet with reason: host reimage * 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2222: Maintenance * 13:45 sukhe: restart pybal on A:lvs-codfw * 13:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2012.codfw.wmnet with reason: host reimage * 13:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db2181: Maintenance * 13:44 btullis@dns1004: END - running authdns-update * 13:42 sukhe: restart pybal on lvs2014 * 13:42 btullis@dns1004: START - running authdns-update * 13:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2222 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95376 and previous config saved to /var/cache/conftool/dbconfig/20260728-133948-cwilliams.json * 13:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2222.codfw.wmnet with reason: Maintenance * 13:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2221: Maintenance * 13:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2181 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95374 and previous config saved to /var/cache/conftool/dbconfig/20260728-133857-cwilliams.json * 13:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2181.codfw.wmnet with reason: Maintenance * 13:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2167: Maintenance * 13:30 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T428495|T428495]] * 13:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 13:29 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 13:29 ayounsi@cumin1003: END (FAIL) - Cookbook sre.dns.admin (exit_code=99) DNS admin: depool esams [reason: router upgrade, [[phab:T431749|T431749]]] * 13:28 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: router upgrade, [[phab:T431749|T431749]]] * 13:27 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1022.eqiad.wmnet, repooling source-only afterwards * 13:27 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo - [[phab:T428495|T428495]] * 13:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2012 * 13:27 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2012 * 13:21 lucaswerkmeister-wmde@deploy1003: mwscript-k8s job started: cleanupTitles bolwiki # [[phab:T429951|T429951]] * 13:21 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] (duration: 07m 19s) * 13:20 swfrench-wmf: authdns-update to direct codfw, eqsin, ulsfo etcd clients to eqiad - [[phab:T428495|T428495]] * 13:18 swfrench@dns1004: END - running authdns-update * 13:17 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, anzx: Continuing with deployment * 13:16 swfrench@dns1004: START - running authdns-update * 13:16 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2012 * 13:16 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2012.codfw.wmnet 57.48.192.10.in-addr.arpa 7.5.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:16 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2012.codfw.wmnet 57.48.192.10.in-addr.arpa 7.5.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:16 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:16 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, anzx: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:14 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 13:14 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] * 13:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db2237: Maintenance * 13:13 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2237: Maintenance * 13:13 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:12 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:12 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback IPV6 for asw1-604 - pt1979@cumin2003" * 13:12 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:12 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback IPV6 for asw1-604 - pt1979@cumin2003" * 13:11 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 20s) * 13:11 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 13:10 esanders@deploy1003: Finished scap sync-world: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] (duration: 08m 11s) * 13:08 pt1979@cumin2003: START - Cookbook sre.dns.netbox * 13:07 root@cumin1003: START - Cookbook sre.mysql.pool pool db2237: Maintenance * 13:06 esanders@deploy1003: esanders: Continuing with deployment * 13:04 esanders@deploy1003: esanders: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2153: Maintenance * 13:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2153: Maintenance * 13:02 esanders@deploy1003: Started scap sync-world: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] * 13:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2237 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95362 and previous config saved to /var/cache/conftool/dbconfig/20260728-130107-cwilliams.json * 13:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2237.codfw.wmnet with reason: Maintenance * 13:00 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2236: Maintenance * 12:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2153: Maintenance * 12:52 root@cumin1003: START - Cookbook sre.mysql.pool pool db2221: Maintenance * 12:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2153 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95358 and previous config saved to /var/cache/conftool/dbconfig/20260728-125214-cwilliams.json * 12:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2153.codfw.wmnet with reason: Maintenance * 12:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2167: Maintenance * 12:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-codfw * 12:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2011.codfw.wmnet * 12:51 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2011.codfw.wmnet * 12:49 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:49 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback for asw1-603 - pt1979@cumin2003" * 12:48 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback for asw1-603 - pt1979@cumin2003" * 12:46 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2011.codfw.wmnet * 12:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2221 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95357 and previous config saved to /var/cache/conftool/dbconfig/20260728-124601-cwilliams.json * 12:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2221.codfw.wmnet with reason: Maintenance * 12:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2218: Maintenance * 12:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2167 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95354 and previous config saved to /var/cache/conftool/dbconfig/20260728-124457-cwilliams.json * 12:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2167.codfw.wmnet with reason: Maintenance * 12:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2166: Maintenance * 12:42 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 12:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2011.codfw.wmnet * 12:41 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2010.codfw.wmnet * 12:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2010.codfw.wmnet * 12:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2010.codfw.wmnet * 12:34 pt1979@cumin2003: START - Cookbook sre.dns.netbox * 12:32 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 12:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2010.codfw.wmnet * 12:31 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2009.codfw.wmnet * 12:31 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2009.codfw.wmnet * 12:27 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2009.codfw.wmnet * 12:22 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2009.codfw.wmnet * 12:22 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2008.codfw.wmnet * 12:21 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2008.codfw.wmnet * 12:16 pt1979@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-604-eqsin * 12:16 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2008.codfw.wmnet * 12:16 pt1979@cumin1003: START - Cookbook sre.network.tls for network device asw1-604-eqsin * 12:14 root@cumin1003: START - Cookbook sre.mysql.pool pool db2236: Maintenance * 12:14 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2236: Maintenance * 12:12 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1137.eqiad.wmnet with OS trixie * 12:11 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2008.codfw.wmnet * 12:11 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2007.codfw.wmnet * 12:11 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2007.codfw.wmnet * 12:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db2236: Maintenance * 12:06 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2007.codfw.wmnet * 12:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2236 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95348 and previous config saved to /var/cache/conftool/dbconfig/20260728-120253-cwilliams.json * 12:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2236.codfw.wmnet with reason: Maintenance * 12:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2007.codfw.wmnet * 12:01 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 12:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 11:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2218: Maintenance * 11:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db2166: Maintenance * 11:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2219: Maintenance * 11:56 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2006.codfw.wmnet * 11:52 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1137.eqiad.wmnet with reason: host reimage * 11:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2218 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95344 and previous config saved to /var/cache/conftool/dbconfig/20260728-115155-cwilliams.json * 11:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2218.codfw.wmnet with reason: Maintenance * 11:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2208: Maintenance * 11:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2166 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95342 and previous config saved to /var/cache/conftool/dbconfig/20260728-115119-cwilliams.json * 11:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2166.codfw.wmnet with reason: Maintenance * 11:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2164: Maintenance * 11:47 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1137.eqiad.wmnet with reason: host reimage * 11:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2006.codfw.wmnet * 11:45 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2005.codfw.wmnet * 11:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2005.codfw.wmnet * 11:40 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2005.codfw.wmnet * 11:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2005.codfw.wmnet * 11:35 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2004.codfw.wmnet * 11:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2004.codfw.wmnet * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1137 * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1137 * 11:30 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1137 * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1137.eqiad.wmnet 192.32.64.10.in-addr.arpa 2.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:30 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1137.eqiad.wmnet 192.32.64.10.in-addr.arpa 2.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1137 - jiji@cumin1003" * 11:25 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2004.codfw.wmnet * 11:19 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2004.codfw.wmnet * 11:19 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2003.codfw.wmnet * 11:19 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2003.codfw.wmnet * 11:14 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2003.codfw.wmnet * 11:11 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2219: Maintenance * 11:10 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2219: Maintenance * 11:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2219: Maintenance * 11:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2003.codfw.wmnet * 11:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2164: Maintenance * 11:03 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2002.codfw.wmnet * 11:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2002.codfw.wmnet * 11:03 root@cumin1003: START - Cookbook sre.mysql.pool pool db2208: Maintenance * 10:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2164 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95332 and previous config saved to /var/cache/conftool/dbconfig/20260728-105749-cwilliams.json * 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2164.codfw.wmnet with reason: Maintenance * 10:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2208 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95331 and previous config saved to /var/cache/conftool/dbconfig/20260728-105711-cwilliams.json * 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2208.codfw.wmnet with reason: Maintenance * 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2219 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95330 and previous config saved to /var/cache/conftool/dbconfig/20260728-105652-cwilliams.json * 10:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2219.codfw.wmnet with reason: Maintenance * 10:53 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1137 - jiji@cumin1003" * 10:52 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2002.codfw.wmnet * 10:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2002.codfw.wmnet * 10:47 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2001.codfw.wmnet * 10:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2001.codfw.wmnet * 10:39 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2001.codfw.wmnet * 10:35 jiji@cumin1003: START - Cookbook sre.dns.netbox * 10:34 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] (duration: 09m 31s) * 10:34 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1137 * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2001.codfw.wmnet * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-codfw * 10:34 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1137.eqiad.wmnet with OS trixie * 10:28 jforrester@deploy1003: jforrester: Continuing with deployment * 10:27 jforrester@deploy1003: jforrester: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:25 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] * 10:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 10:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 10:21 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 10:20 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1137.eqiad.wmnet * 10:20 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 10:20 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1137.eqiad.wmnet * 10:20 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1137.eqiad.wmnet * 10:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2163: Maintenance * 09:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-staging-worker * 09:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2003.codfw.wmnet * 09:37 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2003.codfw.wmnet * 09:32 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 09:31 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2003.codfw.wmnet * 09:30 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 09:30 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2163: Maintenance * 09:30 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 09:30 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 09:30 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:22 klausman@cumin1003: END (ERROR) - Cookbook sre.ganeti.reboot-vm (exit_code=97) for VM ml-serve-ctrl2001.codfw.wmnet * 09:22 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2001.codfw.wmnet * 09:22 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d8-eqiad * 09:22 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d8-eqiad * 09:21 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2003.codfw.wmnet * 09:20 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2002.codfw.wmnet * 09:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2002.codfw.wmnet * 09:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f2-codfw * 09:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f2-codfw * 09:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e4-codfw * 09:18 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2163: Maintenance * 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e4-codfw * 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-codfw * 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-codfw * 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e5-codfw * 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e5-codfw * 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f4-codfw * 09:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2210: Maintenance * 09:16 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f4-codfw * 09:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2182: Maintenance * 09:14 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2002.codfw.wmnet * 09:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db2163: Maintenance * 09:11 XioNoX: rebooting cr2-drmrs - [[phab:T431749|T431749]] * 09:10 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr2-drmrs,cr2-drmrs IPv6,cr2-drmrs.mgmt with reason: router upgrade * 09:06 XioNoX: draining cr2-drmrs - [[phab:T431749|T431749]] * 09:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2163 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95320 and previous config saved to /var/cache/conftool/dbconfig/20260728-090638-cwilliams.json * 09:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2163.codfw.wmnet with reason: Maintenance * 09:06 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2161: Maintenance * 09:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2002.codfw.wmnet * 09:04 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2001.codfw.wmnet * 09:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2001.codfw.wmnet * 08:57 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2001.codfw.wmnet * 08:48 XioNoX: un-drain cr1-drmrs - [[phab:T431749|T431749]] * 08:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2001.codfw.wmnet * 08:47 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-staging-worker * 08:42 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:35 XioNoX: rebooting cr1-drmrs - [[phab:T431749|T431749]] * 08:33 XioNoX: draining cr1-drmrs - [[phab:T431749|T431749]] * 08:31 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2210: Maintenance * 08:29 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2182: Maintenance * 08:21 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2182: Maintenance * 08:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db2161: Maintenance * 08:16 root@cumin1003: START - Cookbook sre.mysql.pool pool db2182: Maintenance * 08:12 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2210: Maintenance * 08:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2161 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95309 and previous config saved to /var/cache/conftool/dbconfig/20260728-081044-cwilliams.json * 08:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2161.codfw.wmnet with reason: Maintenance * 08:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2154: Maintenance * 08:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2182 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95307 and previous config saved to /var/cache/conftool/dbconfig/20260728-080947-cwilliams.json * 08:09 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2182.codfw.wmnet with reason: Maintenance * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2168: Maintenance * 08:06 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr1-drmrs,cr1-drmrs IPv6,cr1-drmrs.mgmt with reason: router upgrade * 08:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db2210: Maintenance * 08:05 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 08:05 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 08:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2210 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95305 and previous config saved to /var/cache/conftool/dbconfig/20260728-080008-cwilliams.json * 08:00 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2210.codfw.wmnet with reason: Maintenance * 07:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2206: Maintenance * 07:50 gkyziridis@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 07:50 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 07:22 root@cumin1003: START - Cookbook sre.mysql.pool pool db2154: Maintenance * 07:22 root@cumin1003: START - Cookbook sre.mysql.pool pool db2168: Maintenance * 07:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2154 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95295 and previous config saved to /var/cache/conftool/dbconfig/20260728-071640-cwilliams.json * 07:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2154.codfw.wmnet with reason: Maintenance * 07:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95294 and previous config saved to /var/cache/conftool/dbconfig/20260728-071604-cwilliams.json * 07:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2168.codfw.wmnet with reason: Maintenance * 07:08 root@cumin1003: START - Cookbook sre.mysql.pool pool db2206: Maintenance * 07:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2206 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95292 and previous config saved to /var/cache/conftool/dbconfig/20260728-070219-cwilliams.json * 07:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2206.codfw.wmnet with reason: Maintenance * 06:44 marostegui: Failover m5 from db1164 to db1228 - [[phab:T432967|T432967]] * 06:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2235].codfw.wmnet,db[1164,1217,1228].eqiad.wmnet with reason: m5 master switch [[phab:T432967|T432967]] * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.10 (duration: 02m 34s) * 03:39 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] (duration: 36m 06s) * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 02:57 dzahn@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1004.eqiad.wmnet with OS trixie * 02:57 dzahn@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - dzahn@cumin1003" * 02:55 dzahn@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - dzahn@cumin1003" * 02:37 dzahn@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1004.eqiad.wmnet with reason: host reimage * 02:31 dzahn@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1004.eqiad.wmnet with reason: host reimage * 02:16 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie * 02:15 dzahn@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host zuul1004.eqiad.wmnet with OS trixie * 01:43 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie * 01:43 dzahn@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1004.eqiad.wmnet with OS trixie * 01:25 pt1979@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-603-eqsin * 01:24 pt1979@cumin1003: START - Cookbook sre.network.tls for network device asw1-603-eqsin * 01:12 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 01:12 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt for new switches in eqsin - pt1979@cumin2003" * 01:12 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt for new switches in eqsin - pt1979@cumin2003" * 01:08 pt1979@cumin2003: START - Cookbook sre.dns.netbox * 00:48 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 00:47 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 00:47 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 00:47 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 00:26 mutante: attempting reimage with trixie on zuul1004 re-purposed physical hardware - dcops reported install issue - host was in busybox shell ([[phab:T427353|T427353]]) * 00:24 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie == 2026-07-27 == * 23:50 Amir1: mass deleting vp8 transcodes * 23:28 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:27 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1004.eqiad.wmnet with OS bullseye * 23:26 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 23:25 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:25 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 22:39 maryum: Deploy security fix for [[phab:T432877|T432877]] * 22:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1022.eqiad.wmnet with OS bookworm * 22:37 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS bullseye * 22:32 sbassett: Deployed security fix for [[phab:T432789|T432789]] * 22:22 sbassett: Deployed security patch for [[phab:T431819|T431819]] * 22:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1022.eqiad.wmnet with reason: host reimage * 22:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1022.eqiad.wmnet with reason: host reimage * 22:01 RScout-WMF: Deployed security fix for [[phab:T431819|T431819]] * 22:00 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2012 * 21:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2012.codfw.wmnet with OS bookworm * 21:55 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2011\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 21:45 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1022 * 21:45 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1022 * 21:44 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1022 * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1022.eqiad.wmnet 239.48.64.10.in-addr.arpa 9.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:44 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1022.eqiad.wmnet 239.48.64.10.in-addr.arpa 9.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:41 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:41 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 21:34 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS bookworm * 21:31 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:22 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:21 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:19 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:17 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1004 * 21:16 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1004 * 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1004] - vriley@cumin1003" * 21:15 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1004] - vriley@cumin1003" * 21:11 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:10 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2011.codfw.wmnet, repooling source-only afterwards * 21:05 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:01 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1022 * 20:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1022.eqiad.wmnet with OS bookworm * 20:53 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1021.eqiad.wmnet, repooling source-only afterwards * 20:51 mutante: zuul1001 - re-enabled puppet - revert "cherry-picked" gerrit:1314120 - [[phab:T431003|T431003]] * 20:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Maintenance * 20:15 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] (duration: 08m 03s) * 20:11 sbisson@deploy1003: sbisson: Continuing with deployment * 20:09 sbisson@deploy1003: sbisson: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] * 19:47 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 46s) * 19:47 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Maintenance * 19:27 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] (duration: 12m 26s) * 19:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2228 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95285 and previous config saved to /var/cache/conftool/dbconfig/20260727-192711-cwilliams.json * 19:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2228.codfw.wmnet with reason: Maintenance * 19:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2223: Maintenance * 19:23 krinkle@deploy1003: krinkle: Continuing with deployment * 19:16 krinkle@deploy1003: krinkle: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:15 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] * 19:12 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2238: Maintenance * 18:58 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 18:57 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 18:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2227: Maintenance * 18:57 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-experimental: apply * 18:55 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-experimental: apply * 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1021.eqiad.wmnet with OS bookworm * 18:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2011.codfw.wmnet with OS bookworm * 18:40 root@cumin1003: START - Cookbook sre.mysql.pool pool db2223: Maintenance * 18:39 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] (duration: 07m 05s) * 18:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2223 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95275 and previous config saved to /var/cache/conftool/dbconfig/20260727-183500-cwilliams.json * 18:34 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2223.codfw.wmnet with reason: Maintenance * 18:34 musikanimal@deploy1003: musikanimal: Continuing with deployment * 18:34 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2213: Maintenance * 18:33 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:32 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] * 18:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db2238: Maintenance * 18:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2011.codfw.wmnet with reason: host reimage * 18:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2238 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95271 and previous config saved to /var/cache/conftool/dbconfig/20260727-181944-cwilliams.json * 18:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2238.codfw.wmnet with reason: Maintenance * 18:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2226: Maintenance * 18:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1021.eqiad.wmnet with reason: host reimage * 18:14 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2011.codfw.wmnet with reason: host reimage * 18:12 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1021.eqiad.wmnet with reason: host reimage * 18:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db2227: Maintenance * 18:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2227 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95265 and previous config saved to /var/cache/conftool/dbconfig/20260727-180256-cwilliams.json * 18:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2227.codfw.wmnet with reason: Maintenance * 18:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2194: Maintenance * 17:57 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2011 * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2011 * 17:56 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2011 * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2011.codfw.wmnet 37.32.192.10.in-addr.arpa 7.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:56 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2011.codfw.wmnet 37.32.192.10.in-addr.arpa 7.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2011 - bking@cumin2003" * 17:56 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2011 - bking@cumin2003" * 17:52 bking@cumin2003: START - Cookbook sre.dns.netbox * 17:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2011 * 17:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1021 * 17:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1021 * 17:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2011.codfw.wmnet with OS bookworm * 17:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1021.eqiad.wmnet with OS bookworm * 17:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Maintenance * 17:38 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2010\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 17:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2213 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95260 and previous config saved to /var/cache/conftool/dbconfig/20260727-173740-cwilliams.json * 17:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2213.codfw.wmnet with reason: Maintenance * 17:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2211: Maintenance * 17:36 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1020\.eqiad\.wmnet,dc=eqiad,cluster=wdqs\-main,service=wdqs\-main * 17:32 root@cumin1003: START - Cookbook sre.mysql.pool pool db2226: Maintenance * 17:31 taavi@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] (duration: 06m 33s) * 17:27 taavi@deploy1003: taavi: Continuing with deployment * 17:27 taavi@deploy1003: taavi: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:26 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2226 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95256 and previous config saved to /var/cache/conftool/dbconfig/20260727-172636-cwilliams.json * 17:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2226.codfw.wmnet with reason: Maintenance * 17:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2225: Maintenance * 17:25 taavi@deploy1003: Started scap sync-world: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] * 17:13 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 17:11 root@cumin1003: START - Cookbook sre.mysql.pool pool db2194: Maintenance * 17:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2194 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95248 and previous config saved to /var/cache/conftool/dbconfig/20260727-170453-cwilliams.json * 17:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2194.codfw.wmnet with reason: Maintenance * 17:04 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2190: Maintenance * 16:52 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 16:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2211: Maintenance * 16:40 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95242 and previous config saved to /var/cache/conftool/dbconfig/20260727-164015-cwilliams.json * 16:40 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2211.codfw.wmnet with reason: Maintenance * 16:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2178: Maintenance * 16:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db2225: Maintenance * 16:39 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2172: Maintenance * 16:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 16:38 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 16:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2225 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95238 and previous config saved to /var/cache/conftool/dbconfig/20260727-163307-cwilliams.json * 16:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2225.codfw.wmnet with reason: Maintenance * 16:32 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2189: Maintenance * 16:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db2190: Maintenance * 16:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2190 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95230 and previous config saved to /var/cache/conftool/dbconfig/20260727-160602-cwilliams.json * 16:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2190.codfw.wmnet with reason: Maintenance * 15:53 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2177: Maintenance * 15:53 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2172: Maintenance * 15:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2178: Maintenance * 15:51 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2172: Maintenance * 15:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2178 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95224 and previous config saved to /var/cache/conftool/dbconfig/20260727-154559-cwilliams.json * 15:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2172: Maintenance * 15:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2178.codfw.wmnet with reason: Maintenance * 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2171: Maintenance * 15:44 root@cumin1003: START - Cookbook sre.mysql.pool pool db2189: Maintenance * 15:43 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:41 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2172 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95222 and previous config saved to /var/cache/conftool/dbconfig/20260727-153927-cwilliams.json * 15:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2172.codfw.wmnet with reason: Maintenance * 15:38 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2189 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95220 and previous config saved to /var/cache/conftool/dbconfig/20260727-153833-cwilliams.json * 15:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2189.codfw.wmnet with reason: Maintenance * 15:34 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:32 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 15:32 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 15:31 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:29 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:26 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 15:22 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] (duration: 07m 00s) * 15:21 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2155: Maintenance * 15:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2175: Maintenance * 15:18 zabe@deploy1003: zabe: Continuing with deployment * 15:17 zabe@deploy1003: zabe: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:15 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 15:15 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:15 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] * 15:15 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 06s) * 15:15 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:12 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2010.codfw.wmnet with OS bookworm * 15:08 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2177: Maintenance * 15:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2177: Maintenance * 14:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2177: Maintenance * 14:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2171: Maintenance * 14:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2171 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95209 and previous config saved to /var/cache/conftool/dbconfig/20260727-145236-cwilliams.json * 14:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2171.codfw.wmnet with reason: Maintenance * 14:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2177 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95208 and previous config saved to /var/cache/conftool/dbconfig/20260727-145206-cwilliams.json * 14:52 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2157: Maintenance * 14:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2177.codfw.wmnet with reason: Maintenance * 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1020.eqiad.wmnet with OS bookworm * 14:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2156: Maintenance * 14:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2010.codfw.wmnet with reason: host reimage * 14:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2010.codfw.wmnet with reason: host reimage * 14:41 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[1031,2024]*: Upgrade Cassandra to 5.0.8 (canary) - eevans@cumin1003 * 14:34 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2155: Maintenance * 14:33 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2175: Maintenance * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2010 * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2010 * 14:24 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2010 * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2010.codfw.wmnet 94.16.192.10.in-addr.arpa 4.9.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2010.codfw.wmnet 94.16.192.10.in-addr.arpa 4.9.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2010 - bking@cumin2003" * 14:24 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2010 - bking@cumin2003" * 14:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1020.eqiad.wmnet with reason: host reimage * 14:23 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[1031,2024]*: Upgrade Cassandra to 5.0.8 (canary) - eevans@cumin1003 * 14:20 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 14:20 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 14:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1020.eqiad.wmnet with reason: host reimage * 14:17 sukhe: sudo gnt-instance reboot urldownloader1005.wikimedia.org * 14:16 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:15 jelto@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:08 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2155: Maintenance * 14:05 root@cumin1003: START - Cookbook sre.mysql.pool pool db2157: Maintenance * 14:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2175: Maintenance * 14:03 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 11 hosts * 14:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db2155: Maintenance * 14:01 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 11 hosts * 14:01 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1136.eqiad.wmnet * 14:01 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1136.eqiad.wmnet * 14:01 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1136.eqiad.wmnet * 14:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db2156: Maintenance * 13:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2157 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95194 and previous config saved to /var/cache/conftool/dbconfig/20260727-135943-cwilliams.json * 13:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2157.codfw.wmnet with reason: Maintenance * 13:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db2175: Maintenance * 13:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 13:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 13:57 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2010 * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2155 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95193 and previous config saved to /var/cache/conftool/dbconfig/20260727-135613-cwilliams.json * 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2155.codfw.wmnet with reason: Maintenance * 13:55 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1020 * 13:55 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1020 * 13:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2156 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95192 and previous config saved to /var/cache/conftool/dbconfig/20260727-135413-cwilliams.json * 13:54 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2156.codfw.wmnet with reason: Maintenance * 13:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2175 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95191 and previous config saved to /var/cache/conftool/dbconfig/20260727-135300-cwilliams.json * 13:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2010.codfw.wmnet with OS bookworm * 13:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2175.codfw.wmnet with reason: Maintenance * 13:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1020.eqiad.wmnet with OS bookworm * 13:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 34 hosts * 13:46 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 34 hosts * 13:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1201: Maintenance * 13:27 Lucas_WMDE: UTC afternoon backport+config window doen * 13:18 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] (duration: 11m 57s) * 13:14 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, sihe: Continuing with deployment * 13:08 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, sihe: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:07 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool ulsfo [reason: router upgrade finished, [[phab:T431752|T431752]]] * 13:07 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool ulsfo [reason: router upgrade finished, [[phab:T431752|T431752]]] * 13:06 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] * 13:03 XioNoX: repool cr4-ulsfo - [[phab:T431752|T431752]] * 12:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db1201: Maintenance * 12:48 gkyziridis@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1201 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95186 and previous config saved to /var/cache/conftool/dbconfig/20260727-124404-cwilliams.json * 12:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1201.eqiad.wmnet with reason: Maintenance * 12:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1187: Maintenance * 12:30 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader.eqiad.wikimedia.org on all recursors * 12:30 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader.eqiad.wikimedia.org on all recursors * 12:30 sukhe@dns1004: END - running authdns-update * 12:30 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] (duration: 09m 32s) * 12:28 sukhe@dns1004: START - running authdns-update * 12:25 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 12:22 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:20 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] * 12:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts * 12:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts * 12:13 XioNoX: rebooting cr4-ulsfo for upgrade - [[phab:T431752|T431752]] * 12:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: es1038 repool * 12:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 38 hosts * 12:08 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 38 hosts * 11:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1187: Maintenance * 11:53 urbanecm@deploy1003: mwscript-k8s job started: foreachwikiindblist growthexperiments GrowthExperiments:cleanMentorList # [[phab:T431804|T431804]] * 11:50 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr4-ulsfo,cr4-ulsfo IPv6,cr4-ulsfo.mgmt with reason: router upgrade * 11:50 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] (duration: 11m 07s) * 11:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1187 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95178 and previous config saved to /var/cache/conftool/dbconfig/20260727-114844-cwilliams.json * 11:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1187.eqiad.wmnet with reason: Maintenance * 11:43 urbanecm@deploy1003: urbanecm: Continuing with deployment * 11:42 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:39 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] * 11:37 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 11:36 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 11:36 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 11:35 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 11:29 XioNoX: start draining cr4-ulsfo - [[phab:T431752|T431752]] * 11:29 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 11:29 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 11:28 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1035: testing * 11:28 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1035: testing * 11:27 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1035: testing * 11:27 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1035: testing * 11:26 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool ulsfo [reason: router upgrade, [[phab:T431752|T431752]]] * 11:26 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1038: es1038 repool * 11:26 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool ulsfo [reason: router upgrade, [[phab:T431752|T431752]]] * 11:26 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1038: testing * 11:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1264: Maintenance * 11:24 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1038: testing * 11:23 marostegui@cumin1003: dbctl commit (dc=all): 'Repool es1050 as master', diff saved to https://phabricator.wikimedia.org/P95170 and previous config saved to /var/cache/conftool/dbconfig/20260727-112326-marostegui.json * 11:23 marostegui@cumin1003: dbctl commit (dc=all): 'Repool es1050', diff saved to https://phabricator.wikimedia.org/P95169 and previous config saved to /var/cache/conftool/dbconfig/20260727-112302-marostegui.json * 11:22 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1050: testing * 11:22 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1050: testing * 11:20 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 11:18 blake@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 11:18 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 11:12 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 11:11 blake@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 11:09 blake@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 11:09 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 11:09 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 11:08 blake@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 11:05 blake@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 11:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 11:02 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 10:50 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 10:43 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:39 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply * 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1264: Maintenance * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply * 10:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply * 10:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 10:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 10:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 10:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1264 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95164 and previous config saved to /var/cache/conftool/dbconfig/20260727-103204-cwilliams.json * 10:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1264.eqiad.wmnet with reason: Maintenance * 10:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 10:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 10:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 10:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 10:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 10:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 10:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 10:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1237: Maintenance * 10:04 elukey: restart burrow main-eqiad on kafkamon2003 to clear some errors on kafka-main1008 * 09:58 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1136.eqiad.wmnet with OS trixie * 09:39 elukey: restart burrow-main-eqiad.service on kafkamon1003 to see if a recurrent kafka error on kafka-main1008 goes away * 09:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1237: Maintenance * 09:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1136.eqiad.wmnet with reason: host reimage * 09:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1237 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95159 and previous config saved to /var/cache/conftool/dbconfig/20260727-093328-cwilliams.json * 09:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1237.eqiad.wmnet with reason: Maintenance * 09:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1136.eqiad.wmnet with reason: host reimage * 09:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1203: Maintenance * 09:17 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1136 * 09:17 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1136 * 09:04 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1136 * 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1136.eqiad.wmnet 191.32.64.10.in-addr.arpa 1.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:04 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1136.eqiad.wmnet 191.32.64.10.in-addr.arpa 1.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1136 - jiji@cumin1003" * 09:04 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1136 - jiji@cumin1003" * 08:52 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 08:52 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 08:52 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 08:51 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 08:50 jiji@cumin1003: START - Cookbook sre.dns.netbox * 08:47 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1136 * 08:46 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1136.eqiad.wmnet with OS trixie * 08:44 marostegui: Rename tables on s3 [[phab:T425066|T425066]] * 08:43 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1136.eqiad.wmnet * 08:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db1203: Maintenance * 08:43 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1136.eqiad.wmnet * 08:43 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1136.eqiad.wmnet * 08:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1203 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95154 and previous config saved to /var/cache/conftool/dbconfig/20260727-083703-cwilliams.json * 08:36 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1203.eqiad.wmnet with reason: Maintenance * 08:16 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1179: Maintenance * 07:44 phuedx: UTC morning backport window done * 07:37 phuedx@deploy1003: Finished scap sync-world: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] (duration: 32m 33s) * 07:28 root@cumin1003: START - Cookbook sre.mysql.pool pool db1179: Maintenance * 07:26 marostegui: Rename tables on s3 [[phab:T426341|T426341]] * 07:25 phuedx@deploy1003: phuedx: Continuing with deployment * 07:22 marostegui: Drop tables in akwiki nawiki pihwiki - growthexperiments_* [[phab:T428885|T428885]] * 07:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95149 and previous config saved to /var/cache/conftool/dbconfig/20260727-072234-cwilliams.json * 07:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1179.eqiad.wmnet with reason: Maintenance * 07:20 phuedx@deploy1003: phuedx: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:16 ryankemper: [[phab:T430880|T430880]] [WDQS] Reimaged `wdqs1018` and `wdqs1019` to Bookworm, restored data using test-cookbook change {{Gerrit|1317128}}, and repooled both; 25/36 hosts complete * 07:04 phuedx@deploy1003: Started scap sync-world: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] * 06:57 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1019.eqiad.wmnet * 06:56 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1018.eqiad.wmnet * 06:40 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1020.eqiad.wmnet with reason: Cloning * 06:35 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db1228.eqiad.wmnet with reason: Rebooting * 06:29 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:29 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:25 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:25 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:25 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1019.eqiad.wmnet, repooling source-only afterwards * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1018.eqiad.wmnet, repooling source-only afterwards * 04:51 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1019.eqiad.wmnet, repooling source-only afterwards * 04:51 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1018.eqiad.wmnet, repooling source-only afterwards * 04:48 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s) * 04:48 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 04:48 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s) * 04:48 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 36s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-26 == * 14:59 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:59 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:59 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:59 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1019.eqiad.wmnet with OS bookworm * 01:05 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1018.eqiad.wmnet with OS bookworm * 00:43 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1019.eqiad.wmnet with reason: host reimage * 00:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1018.eqiad.wmnet with reason: host reimage * 00:34 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1019.eqiad.wmnet with reason: host reimage * 00:33 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1018.eqiad.wmnet with reason: host reimage * 00:16 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 00:16 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 00:15 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 00:15 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1019 * 00:11 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1019 * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1018 * 00:11 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1018 * 00:08 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1019.eqiad.wmnet with OS bookworm * 00:08 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1018.eqiad.wmnet with OS bookworm == 2026-07-25 == * 22:06 ryankemper: [[phab:T430880|T430880]] [WDQS] Repooled `wdqs1017` and `wdqs2024` after reimaging to bookworm, scap deploying, and data xfering * 22:04 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2024.codfw.wmnet * 22:03 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1017.eqiad.wmnet * 21:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1017.eqiad.wmnet, repooling source-only afterwards * 21:06 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2024.codfw.wmnet, repooling source-only afterwards * 20:52 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:52 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:52 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:52 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 20:18 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1017.eqiad.wmnet, repooling source-only afterwards * 20:18 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2024.codfw.wmnet, repooling source-only afterwards * 20:15 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:15 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:15 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:15 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 19:57 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s) * 19:57 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 19:57 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 07s) * 19:57 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 19:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2024.codfw.wmnet with OS bookworm * 19:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1017.eqiad.wmnet with OS bookworm * 19:02 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2024.codfw.wmnet with reason: host reimage * 18:58 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1017.eqiad.wmnet with reason: host reimage * 18:53 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2024.codfw.wmnet with reason: host reimage * 18:52 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1017.eqiad.wmnet with reason: host reimage * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2024 * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2024 * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1017 * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1017 * 18:27 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2024 * 18:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2024.codfw.wmnet 58.16.192.10.in-addr.arpa 8.5.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:26 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2024.codfw.wmnet 58.16.192.10.in-addr.arpa 8.5.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:24 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1017 * 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1017.eqiad.wmnet 238.48.64.10.in-addr.arpa 8.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:24 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1017.eqiad.wmnet 238.48.64.10.in-addr.arpa 8.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1017 - ryankemper@cumin2003" * 18:24 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1017 - ryankemper@cumin2003" * 18:23 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 18:18 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 18:17 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1017 * 18:17 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2024 * 18:14 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1017.eqiad.wmnet with OS bookworm * 18:14 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2024.codfw.wmnet with OS bookworm * 18:05 ryankemper: [WDQS] [[phab:T430880|T430880]] Reimaged `wdqs1016` and `wdqs2023` to Bookworm with `--move-vlan`, restored main and scholarly data, validated postflights, and repooled both hosts. Confirmed PyBal rebuilt both backends with their new addresses * 17:45 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2023.codfw.wmnet * 17:43 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1016.eqiad.wmnet * 06:35 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1016.eqiad.wmnet, repooling source-only afterwards * 06:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2023.codfw.wmnet, repooling source-only afterwards * 05:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2023.codfw.wmnet, repooling source-only afterwards * 05:19 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1016.eqiad.wmnet, repooling source-only afterwards * 05:07 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 07s) * 05:07 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 05:06 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 06s) * 05:06 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 03:27 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2023.codfw.wmnet with OS bookworm * 02:59 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2023.codfw.wmnet with reason: host reimage * 02:56 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2023.codfw.wmnet with reason: host reimage * 02:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2023 * 02:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2023 * 02:30 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2023 * 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2023.codfw.wmnet 35.0.192.10.in-addr.arpa 5.3.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 02:30 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2023.codfw.wmnet 35.0.192.10.in-addr.arpa 5.3.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2023 - ryankemper@cumin2003" * 02:30 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2023 - ryankemper@cumin2003" * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 26s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:15 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1016.eqiad.wmnet with OS bookworm * 00:49 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1016.eqiad.wmnet with reason: host reimage * 00:43 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1016.eqiad.wmnet with reason: host reimage * 00:31 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 00:27 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1016 * 00:27 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1016 * 00:27 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2023 * 00:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1016.eqiad.wmnet with OS bookworm * 00:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2023.codfw.wmnet with OS bookworm * 00:11 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs1014.eqiad.wmnet and wdqs2008.codfw.wmnet after Bookworm reimage, transfer, and postflight; wdqs2008 is serving, while wdqs1014 will remain outside of service until a pybal restart next monday == 2026-07-24 == * 23:54 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1014.eqiad.wmnet * 23:54 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2008.codfw.wmnet * 23:43 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2010.codfw.wmnet with OS trixie * 23:08 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 23:03 jhathaway@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 22:33 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 22:13 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 22:13 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 22:13 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 22:13 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:00 jhathaway@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 21:53 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 21:53 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie * 21:51 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 21:47 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie * 21:43 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 21:39 jhathaway@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 21:38 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 17:21 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1135.eqiad.wmnet * 17:21 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1135.eqiad.wmnet * 17:21 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1135.eqiad.wmnet * 16:34 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 16:34 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 16:34 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 16:34 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 16:33 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 16:33 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 16:28 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:28 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:28 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:28 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2008.codfw.wmnet, repooling source-only afterwards * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1014.eqiad.wmnet, repooling source-only afterwards * 15:56 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1135.eqiad.wmnet with OS trixie * 15:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 40 hosts * 15:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 40 hosts * 15:37 topranks: upgrade SR-Linux OS on lswtest-d8-eqiad * 15:36 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1135.eqiad.wmnet with reason: host reimage * 15:33 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 6 hosts with reason: upgrade lswtest-d8-eqiad * 15:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1135.eqiad.wmnet with reason: host reimage * 15:30 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc-gp2006.codfw.wmnet with OS bookworm * 15:15 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1135 * 15:15 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1135 * 15:13 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc-gp2006.codfw.wmnet with reason: host reimage * 15:08 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc-gp2006.codfw.wmnet with reason: host reimage * 14:49 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm * 14:48 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host mc-gp2006.codfw.wmnet with OS bookworm * 14:34 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] (duration: 41m 12s) * 14:32 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1135 * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1135.eqiad.wmnet 177.32.64.10.in-addr.arpa 7.7.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:32 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1135.eqiad.wmnet 177.32.64.10.in-addr.arpa 7.7.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1135 - jiji@cumin1003" * 14:32 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1135 - jiji@cumin1003" * 14:29 krinkle@deploy1003: krinkle: Continuing with deployment * 14:29 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm * 14:27 jiji@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host mc-gp2006.codfw.wmnet with OS bookworm * 14:26 jiji@cumin1003: START - Cookbook sre.dns.netbox * 14:15 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1135 * 14:14 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1135.eqiad.wmnet with OS trixie * 14:14 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1135.eqiad.wmnet * 14:13 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1135.eqiad.wmnet * 14:13 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1135.eqiad.wmnet * 14:10 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1072.eqiad.wmnet * 14:10 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1072.eqiad.wmnet * 14:10 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1072.eqiad.wmnet * 14:10 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1072.eqiad.wmnet * 14:09 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1071.eqiad.wmnet * 14:09 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1071.eqiad.wmnet * 14:09 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1071.eqiad.wmnet * 14:09 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1071.eqiad.wmnet * 13:58 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 13:58 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:58 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:57 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:55 krinkle@deploy1003: krinkle: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:53 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] * 13:45 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:45 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push new IPs for mc-gp2006 - cmooney@cumin1003" * 13:45 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push new IPs for mc-gp2006 - cmooney@cumin1003" * 13:44 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) mc-gp2006.codfw.wmnet on all recursors * 13:44 cmooney@cumin1003: START - Cookbook sre.dns.wipe-cache mc-gp2006.codfw.wmnet on all recursors * 13:42 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm * 13:41 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:30 papaul: reboot mr1-eqsin for maintenance * 13:24 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb[1029-1031].eqiad.wmnet * 13:10 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb[1029-1031].eqiad.wmnet * 11:33 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 7 hosts * 11:11 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 7 hosts * 10:56 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 7 hosts * 10:47 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 7 hosts * 10:44 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:44 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:41 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:41 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:35 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 8 hosts * 10:34 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:33 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:32 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:32 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:31 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:31 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:30 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 8 hosts * 10:24 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie * 10:19 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:18 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 16 hosts * 10:17 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2001.codfw.wmnet * 10:13 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2001.codfw.wmnet * 10:12 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2001.codfw.wmnet * 10:02 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2001.codfw.wmnet * 10:02 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2002.codfw.wmnet * 09:57 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2002.codfw.wmnet * 09:56 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2002.codfw.wmnet * 09:51 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2002.codfw.wmnet * 09:51 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1002.eqiad.wmnet * 09:47 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1002.eqiad.wmnet * 09:47 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1001.eqiad.wmnet * 09:44 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1001.eqiad.wmnet * 09:34 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2003.codfw.wmnet * 09:32 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2003.codfw.wmnet * 09:32 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2002.codfw.wmnet * 09:30 brouberol@dns1004: END - running authdns-update * 09:29 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2002.codfw.wmnet * 09:29 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2001.codfw.wmnet * 09:27 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 9 hosts * 09:27 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2001.codfw.wmnet * 09:27 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2001.codfw.wmnet * 09:26 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 9 hosts * 09:26 brouberol@dns1004: START - running authdns-update * 09:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 57 hosts * 09:24 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2001.codfw.wmnet * 09:24 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2002.codfw.wmnet * 09:22 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2002.codfw.wmnet * 09:21 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 57 hosts * 09:20 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2003.codfw.wmnet * 09:19 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 16 hosts * 09:16 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2003.codfw.wmnet * 09:16 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1003.eqiad.wmnet * 09:15 urbanecm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 09:15 urbanecm@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 09:13 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1003.eqiad.wmnet * 09:13 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1002.eqiad.wmnet * 09:11 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1002.eqiad.wmnet * 09:11 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1001.eqiad.wmnet * 09:07 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1001.eqiad.wmnet * 08:32 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:24 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:16 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 08:16 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 08:07 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:07 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:07 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 08:02 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 08:01 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:59 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:57 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 07:57 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 06:46 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1025.eqiad.wmnet with reason: Cloning * 06:46 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s4 * 06:45 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s6 * 06:44 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1019.eqiad.wmnet,service=s6 * 06:44 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1019.eqiad.wmnet,service=s4 * 03:40 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:40 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:40 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:40 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 03:37 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:37 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:37 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:36 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:49 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on mr1-eqsin,mr1-eqsin IPv6 with reason: connection issue * 02:38 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on cr[2-3]-eqsin.mgmt,ps1-[603-604]-eqsin with reason: connection issue * 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 27s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-23 == * 23:27 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin.oob,mr1-eqsin.oob IPv6 with reason: switch refresh * 22:21 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Setting storage compatibility to NONE - eevans@cumin1003 * 22:01 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Setting storage compatibility to NONE - eevans@cumin1003 * 21:29 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1014.eqiad.wmnet, repooling source-only afterwards * 21:28 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 46s) * 21:28 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 21:19 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Setting storage compatibility to UPGRADING - eevans@cumin1003 * 21:00 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Setting storage compatibility to UPGRADING - eevans@cumin1003 * 20:17 dani@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] (duration: 11m 57s) * 20:13 dani@deploy1003: dani: Continuing with deployment * 20:07 dani@deploy1003: dani: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:05 dani@deploy1003: Started scap sync-world: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] * 19:24 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:24 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating the rest of the ipv6 dns records. - jhancock@cumin2002" * 19:24 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating the rest of the ipv6 dns records. - jhancock@cumin2002" * 19:14 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 19:05 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wdqs1014.eqiad.wmnet with OS bookworm * 19:04 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.noop (exit_code=99) * 19:04 cwilliams@cumin1003: START - Cookbook sre.mysql.noop * 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1014.eqiad.wmnet with reason: host reimage * 18:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1014.eqiad.wmnet with reason: host reimage * 18:30 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2008.codfw.wmnet, repooling source-only afterwards * 18:28 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 19s) * 18:28 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1014 * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1014 * 18:22 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1014 * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1014.eqiad.wmnet 188.32.64.10.in-addr.arpa 8.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:22 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1014.eqiad.wmnet 188.32.64.10.in-addr.arpa 8.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1014 - bking@cumin2003" * 18:21 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1014 - bking@cumin2003" * 18:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2215: Maintenance * 18:18 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 18:15 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:15 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 18:06 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:06 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 18:05 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 18:04 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2052: codfw rack B8 re-pool after maintenance * 17:54 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 17:54 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:54 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 17:32 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2215: Maintenance * 17:29 cmooney@dns3003: END - running authdns-update * 17:27 cmooney@dns3003: START - running authdns-update * 17:23 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 17:22 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:18 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool es2052: codfw rack B8 re-pool after maintenance * 17:18 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2189: codfw rack B8 re-pool after maintenance * 17:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2215.codfw.wmnet with reason: Maintenance * 17:17 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 17:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2215 [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95126 and previous config saved to /var/cache/conftool/dbconfig/20260723-170903-cwilliams.json * 17:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2191 to x1 primary [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95125 and previous config saved to /var/cache/conftool/dbconfig/20260723-170612-cwilliams.json * 17:05 cezmunsta: Starting x1 codfw failover from db2215 to db2191 - [[phab:T432986|T432986]] * 16:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2191 with weight 0 [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95123 and previous config saved to /var/cache/conftool/dbconfig/20260723-165831-cwilliams.json * 16:58 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 16 hosts with reason: Primary switchover x1 [[phab:T432986|T432986]] * 16:36 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 138128 * 16:35 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 138128 * 16:33 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2189: codfw rack B8 re-pool after maintenance * 16:33 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2164: codfw rack B8 re-pool after maintenance * 16:28 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1072.eqiad.wmnet * 16:27 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1072.eqiad.wmnet with OS trixie * 16:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2249: Maintenance * 16:06 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker1072.eqiad.wmnet with reason: host reimage * 16:06 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1072.eqiad.wmnet with reason: host reimage * 15:50 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1072 * 15:50 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1072 * 15:49 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1072.eqiad.wmnet with OS trixie * 15:48 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2164: codfw rack B8 re-pool after maintenance * 15:48 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] (duration: 06m 37s) * 15:48 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2163: codfw rack B8 re-pool after maintenance * 15:45 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 15:43 musikanimal@deploy1003: musikanimal: Continuing with deployment * 15:43 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:41 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] * 15:36 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1072.eqiad.wmnet * 15:35 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1072.eqiad.wmnet * 15:35 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1072.eqiad.wmnet * 15:34 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:34 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push any outstanding updates - cmooney@cumin1003" * 15:34 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push any outstanding updates - cmooney@cumin1003" * 15:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db2249: Maintenance * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 15:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:26 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:21 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:21 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 15:21 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:21 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 15:20 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 15:19 cmooney@dns2004: END - running authdns-update * 15:17 cmooney@dns2004: START - running authdns-update * 15:14 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns2004.wikimedia.org * 15:12 brouberol@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 15:12 brouberol@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 15:12 klausman@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ml-serve1001.eqiad.wmnet with OS trixie * 15:11 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1071.eqiad.wmnet * 15:11 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1071.eqiad.wmnet with OS trixie * 15:10 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wdqs2008.codfw.wmnet with OS bookworm * 15:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2249.codfw.wmnet with reason: Maintenance * 15:08 brouberol@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 15:08 brouberol@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 15:08 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 15:08 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 15:06 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns1004.wikimedia.org * 15:02 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2002.codfw.wmnet * 15:02 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2002.codfw.wmnet * 15:02 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2163: codfw rack B8 re-pool after maintenance * 15:01 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 15:01 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 14:59 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test2001.codfw.wmnet * 14:57 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test2001.codfw.wmnet * 14:56 ryankemper: [WDQS] [[phab:T430880|T430880]] Reimaged `wdqs2016` to Bookworm, xferred scholarly_articles from `wdqs2024`, validated updater/readiness/federation, and repooled * 14:51 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2016.codfw.wmnet * 14:51 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1001.eqiad.wmnet with reason: host reimage * 14:48 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1071.eqiad.wmnet with reason: host reimage * 14:47 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1001.eqiad.wmnet with reason: host reimage * 14:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2008.codfw.wmnet with reason: host reimage * 14:43 topranks: reboot lsw1-b8-codw to upgrade JunOS [[phab:T430929|T430929]] * 14:41 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1071.eqiad.wmnet with reason: host reimage * 14:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2008.codfw.wmnet with reason: host reimage * 14:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2231: Maintenance * 14:30 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1001.eqiad.wmnet with OS trixie * 14:25 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2002.codfw.wmnet * 14:23 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1071 * 14:23 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1071 * 14:23 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 14:22 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore scholarly data after Bookworm reimage) xfer scholarly_articles from wdqs2024.codfw.wmnet -> wdqs2016.codfw.wmnet, repooling source-only afterwards * 14:22 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2052: codfw rack B8 depool for maintenance * 14:21 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool es2052: codfw rack B8 depool for maintenance * 14:21 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2249: codfw rack B8 depool for maintenance * 14:21 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1071 * 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1071.eqiad.wmnet 166.48.64.10.in-addr.arpa 6.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:21 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1071.eqiad.wmnet 166.48.64.10.in-addr.arpa 6.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1071 - jiji@cumin1003" * 14:21 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1071 - jiji@cumin1003" * 14:21 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2249: codfw rack B8 depool for maintenance * 14:21 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2189: codfw rack B8 depool for maintenance * 14:20 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2002.codfw.wmnet * 14:20 cmooney@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2050.codfw.wmnet * 14:20 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2189: codfw rack B8 depool for maintenance * 14:20 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2164: codfw rack B8 depool for maintenance * 14:20 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2164: codfw rack B8 depool for maintenance * 14:19 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2163: codfw rack B8 depool for maintenance * 14:19 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:19 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1014 * 14:19 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2163: codfw rack B8 depool for maintenance * 14:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1014.eqiad.wmnet with OS bookworm * 14:17 cmooney@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2050.codfw.wmnet * 14:16 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 14:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2008 * 14:14 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2008 * 14:14 cmooney@cumin1003: conftool action : set/pooled=no; selector: name=dns2004.wikimedia.org * 14:14 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2008.codfw.wmnet with OS bookworm * 14:13 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 14:12 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:10 topranks: depool dns2004 before lsw1-b8-codfw switch maintenance [[phab:T430929|T430929]] * 14:10 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b8-codfw,lsw1-b8-codfw IPv6,lsw1-b8-codfw.mgmt,ssw1-a[1,8]-codfw with reason: lsw1-b8-codfw JunOS upgrade * 14:07 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 30 hosts with reason: lsw1-b8-codfw JunOS upgrade * 14:06 elukey: upload python3-docker-report 0.0.19 to apt.wikimedia.org for bookworm and trixie * 13:59 jiji@cumin1003: START - Cookbook sre.dns.netbox * 13:58 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1071 * 13:57 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:57 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:57 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:57 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1071.eqiad.wmnet with OS trixie * 13:55 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1071.eqiad.wmnet * 13:55 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1071.eqiad.wmnet * 13:55 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1071.eqiad.wmnet * 13:53 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:52 logmsgbot: kharlan Deployed security patch for [[phab:T432948|T432948]] * 13:51 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:51 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 13:51 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db2231: Maintenance * 13:50 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:50 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:50 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:50 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:49 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:49 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:49 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2231 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95097 and previous config saved to /var/cache/conftool/dbconfig/20260723-134436-cwilliams.json * 13:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2231.codfw.wmnet with reason: Maintenance * 13:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:39 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:38 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 13:38 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] (duration: 09m 07s) * 13:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1037 hosts * 13:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2196: Maintenance * 13:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:33 kharlan@deploy1003: kharlan, emc-wmf: Continuing with deployment * 13:31 kharlan@deploy1003: kharlan, emc-wmf: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:30 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:28 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] * 13:17 hashar@deploy1003: Finished deploy [integration/docroot@2199146]: build: License GPL2.0+ / updating npm dependencies (duration: 00m 14s) * 13:17 hashar@deploy1003: Started deploy [integration/docroot@2199146]: build: License GPL2.0+ / updating npm dependencies * 13:14 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service * 13:07 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 12:58 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2207: Repooling * 12:49 root@cumin1003: START - Cookbook sre.mysql.pool pool db2196: Maintenance * 12:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2196 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95087 and previous config saved to /var/cache/conftool/dbconfig/20260723-123952-cwilliams.json * 12:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2196.codfw.wmnet with reason: Maintenance * 12:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2191: Maintenance * 12:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:13 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:13 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: Repooling * 12:12 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2207: Repooling * 12:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: Repooling * 11:56 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2235.codfw.wmnet with OS trixie * 11:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db2191: Maintenance * 11:46 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1070.eqiad.wmnet * 11:46 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1070.eqiad.wmnet * 11:46 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1070.eqiad.wmnet * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2191 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95080 and previous config saved to /var/cache/conftool/dbconfig/20260723-114308-cwilliams.json * 11:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2191.codfw.wmnet with reason: Maintenance * 11:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2186: Maintenance * 11:35 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 46375 * 11:34 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 46375 * 11:33 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2235.codfw.wmnet with reason: host reimage * 11:28 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2235.codfw.wmnet with reason: host reimage * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c7-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c7-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c6-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c6-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c5-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c5-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c4-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c4-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c3-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c3-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c2-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c2-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d7-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d7-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d4-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d4-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d3-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d2-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d2-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d8-eqiad * 11:23 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d8-eqiad * 11:23 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d1-eqiad * 11:23 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d1-eqiad * 11:12 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db2235.codfw.wmnet with OS trixie * 11:11 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:11 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[2160,2235].codfw.wmnet with reason: Upgrading * 10:56 root@cumin1003: START - Cookbook sre.mysql.pool pool db2186: Maintenance * 10:54 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1037: testing * 10:53 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1037: testing * 10:53 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1037: testing * 10:53 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1037: testing * 10:52 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: testing * 10:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2186 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95072 and previous config saved to /var/cache/conftool/dbconfig/20260723-104956-cwilliams.json * 10:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2186.codfw.wmnet with reason: Maintenance * 10:43 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1054: testing * 10:41 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1070.eqiad.wmnet with OS trixie * 10:30 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1037 hosts * 10:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 10:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 10:20 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1070.eqiad.wmnet with reason: host reimage * 10:16 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1070.eqiad.wmnet with reason: host reimage * 10:06 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1038: testing * 10:05 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1038: testing * 10:05 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1038: testing * 10:02 marostegui@dns1004: END - running authdns-update * 10:00 marostegui@dns1004: START - running authdns-update * 09:58 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: testing * 09:57 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1054: testing * 09:57 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1070 * 09:57 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1070 * 09:57 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1054: testing * 09:57 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1054: testing * 09:56 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1055: testing * 09:56 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1070 * 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1070.eqiad.wmnet 165.48.64.10.in-addr.arpa 5.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:56 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1070.eqiad.wmnet 165.48.64.10.in-addr.arpa 5.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1070 - jiji@cumin1003" * 09:56 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1070 - jiji@cumin1003" * 09:47 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for 1035 hosts * 09:45 jiji@cumin1003: START - Cookbook sre.dns.netbox * 09:42 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1070 * 09:42 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1070.eqiad.wmnet with OS trixie * 09:42 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1070.eqiad.wmnet * 09:41 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1070.eqiad.wmnet * 09:41 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1070.eqiad.wmnet * 09:27 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es2051: testing * 09:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: testing * 09:12 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es2051: testing * 09:11 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1055: testing * 09:09 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:09 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1055: testing * 09:09 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1055: testing * 08:50 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:50 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:50 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 08:49 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 08:49 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 08:49 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:46 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1069.eqiad.wmnet * 08:46 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1069.eqiad.wmnet * 08:46 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1069.eqiad.wmnet * 08:39 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 08:38 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 08:38 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 08:37 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 08:35 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:10 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1069.eqiad.wmnet with OS trixie * 07:49 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1069.eqiad.wmnet with reason: host reimage * 07:45 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1069.eqiad.wmnet with reason: host reimage * 07:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1035 hosts * 07:33 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 07:32 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 07:29 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1069 * 07:29 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1069 * 07:26 jiji@deploy1003: Finished scap sync-world: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules (duration: 06m 01s) * 07:25 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1069 * 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1069.eqiad.wmnet 164.48.64.10.in-addr.arpa 4.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:25 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1069.eqiad.wmnet 164.48.64.10.in-addr.arpa 4.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1069 - jiji@cumin1003" * 07:25 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1069 - jiji@cumin1003" * 07:25 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 07:25 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 07:24 jiji@deploy1003: jiji: Continuing with deployment * 07:22 jiji@deploy1003: jiji: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:21 jiji@deploy1003: Started scap sync-world: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules * 07:21 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1031.eqiad.wmnet,service=s7 * 07:20 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1031.eqiad.wmnet,service=s2 * 07:20 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1031.eqiad.wmnet,service=s7 * 07:20 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1031.eqiad.wmnet,service=s2 * 07:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts * 07:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts * 07:19 jiji@cumin1003: START - Cookbook sre.dns.netbox * 07:19 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1069 * 07:19 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1069.eqiad.wmnet with OS trixie * 07:19 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1069.eqiad.wmnet * 07:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 45 hosts * 07:17 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1069.eqiad.wmnet * 07:17 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1069.eqiad.wmnet * 07:14 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 45 hosts * 07:13 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 06:16 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs2007 after successful Bookworm reimage, data transfer, and postflight validation; wdqs1013 also passed postflights and is enabled in conftool, but remains out of IPVS pending a rolling pybal restart to clear its stale pre-VLAN-move address * 05:58 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2007.codfw.wmnet * 05:58 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1013.eqiad.wmnet * 05:54 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore scholarly data after Bookworm reimage) xfer scholarly_articles from wdqs2024.codfw.wmnet -> wdqs2016.codfw.wmnet, repooling source-only afterwards == 2026-07-22 == * 23:34 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Apply upgrade to JVM17 - eevans@cumin1003 * 23:14 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Apply upgrade to JVM17 - eevans@cumin1003 * 22:06 ryankemper: [WDQS] Added requestctl per-IP ratelimit `wdqs_heavy_sparql_bots_jul_2026_ratelimit` (chronic heavy-query bot tier driving deadlock-remediation restarts); pruned superseded `wdqs_2026_05_11_worobot` * 21:51 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] (duration: 11m 52s) * 21:44 sbassett@deploy1003: sbassett: Continuing with deployment * 21:43 sbassett@deploy1003: sbassett: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:39 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] * 20:38 dani@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] (duration: 32m 51s) * 20:38 ryankemper: [WDQS] Pruned obsolete requestctl action+pattern `wdqs_20260715_p2003_ring_ja3n` (actor rotated JA3Ns; rule inert) * 20:26 dani@deploy1003: dani, vadymts1: Continuing with deployment * 20:24 dani@deploy1003: dani, vadymts1: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:14 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2244: Testing * 20:06 dani@deploy1003: Started scap sync-world: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] * 19:56 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1013.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2016.codfw.wmnet with OS bookworm * 19:40 mutante: gerrit - one more service restart is needed - restarting * 19:29 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2244: Testing * 19:27 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2244: Testing * 19:27 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2244: Testing * 19:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2016.codfw.wmnet with reason: host reimage * 19:15 dancy@deploy1003: Finished deploy [zuul/deploy@d92e238]: Freshening Zuul installation (duration: 00m 15s) * 19:14 dancy@deploy1003: Started deploy [zuul/deploy@d92e238]: Freshening Zuul installation * 19:11 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2016.codfw.wmnet with reason: host reimage * 18:54 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1013.eqiad.wmnet, repooling source-only afterwards * 18:52 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2016 * 18:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2016 * 18:51 dancy@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 18:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2016.codfw.wmnet with OS bookworm * 18:39 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 11s) * 18:39 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 18:36 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 18:30 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417] (thin): Regular analytics weekly train THIN [analytics/refinery@2a25417d] (duration: 02m 09s) * 18:28 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417] (thin): Regular analytics weekly train THIN [analytics/refinery@2a25417d] * 18:28 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417]: Regular analytics weekly train [analytics/refinery@2a25417d] (duration: 04m 31s) * 18:27 dduvall: deploying https://gerrit.wikimedia.org/r/c/integration/config/+/1314025 (4 jobs updated) * 18:23 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417]: Regular analytics weekly train [analytics/refinery@2a25417d] * 18:22 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@2a25417d] (duration: 01m 59s) * 18:20 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@2a25417d] * 17:56 Raine: deployment server switchover => deploy1003 is primary now * 17:55 kamila@deploy1003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 28m 24s) * 17:54 mutante: restarting gerrit for maintenance * 17:29 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1013.eqiad.wmnet with OS bookworm * 17:27 kamila@deploy1003: Started scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] * 17:20 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] (duration: 22m 50s) * 17:12 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1023.eqiad.wmnet -> wdqs1024.eqiad.wmnet, repooling source-only afterwards * 17:04 Raine: point deployment.eqiad.wmnet to deploy1003 * 17:04 kamila@dns7001: END - running authdns-update * 17:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1013.eqiad.wmnet with reason: host reimage * 17:02 kamila@dns7001: START - running authdns-update * 17:01 kamila@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on releases2003.codfw.wmnet,releases1003.eqiad.wmnet with reason: Deployment server switchover * 17:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1013.eqiad.wmnet with reason: host reimage * 16:58 kamila@deploy2003: Locking from deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] * 16:57 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] (duration: 02m 33s) * 16:55 kamila@deploy2003: Locking from deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] * 16:55 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2003 - [[phab:T240266|T240266]] (duration: 00m 11s) * 16:54 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2003 - [[phab:T240266|T240266]] * 16:40 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1013 * 16:40 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1013 * 16:39 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1013 * 16:39 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1013.eqiad.wmnet 105.32.64.10.in-addr.arpa 5.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:39 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1013.eqiad.wmnet 105.32.64.10.in-addr.arpa 5.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:39 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:39 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1013 - bking@cumin2003" * 16:39 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1013 - bking@cumin2003" * 16:34 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:34 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1013 * 16:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1013.eqiad.wmnet with OS bookworm * 16:28 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1023.eqiad.wmnet -> wdqs1024.eqiad.wmnet, repooling source-only afterwards * 16:27 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-scholarly,name=eqiad * 16:27 eevans@deploy2003: helmfile [eqiad] DONE helmfile.d/services/linked-artifacts: apply * 16:26 eevans@deploy2003: helmfile [eqiad] START helmfile.d/services/linked-artifacts: apply * 16:26 eevans@deploy2003: helmfile [codfw] DONE helmfile.d/services/linked-artifacts: apply * 16:26 eevans@deploy2003: helmfile [codfw] START helmfile.d/services/linked-artifacts: apply * 16:25 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 29s) * 16:25 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 16:24 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 16:21 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 16:18 eevans@deploy2003: helmfile [codfw] DONE helmfile.d/services/linked-artifacts: apply * 16:18 eevans@deploy2003: helmfile [codfw] START helmfile.d/services/linked-artifacts: apply * 16:08 eevans@deploy2003: helmfile [staging] DONE helmfile.d/services/linked-artifacts: apply * 16:07 eevans@deploy2003: helmfile [staging] START helmfile.d/services/linked-artifacts: apply * 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 16:01 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 15:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1024.eqiad.wmnet with OS bookworm * 15:49 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] (duration: 00m 10s) * 15:49 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] * 15:48 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] (duration: 00m 15s) * 15:48 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] * 15:47 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] (duration: 00m 10s) * 15:47 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] * 15:46 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:42 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:42 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:40 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:37 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 15:37 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:36 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:36 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:36 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 15:33 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 15:33 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:31 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:28 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 15:27 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:27 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1024.eqiad.wmnet with reason: host reimage * 15:23 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:23 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:23 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:20 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1068.eqiad.wmnet * 15:20 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1068.eqiad.wmnet * 15:20 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1068.eqiad.wmnet * 15:20 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 15:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1024.eqiad.wmnet with reason: host reimage * 15:11 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:55 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 14:52 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wdqs1024.eqiad.wmnet with OS bookworm * 14:50 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] (duration: 00m 09s) * 14:50 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] * 14:49 jiji@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 14:49 jiji@deploy2003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 14:49 jiji@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 14:48 jiji@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 14:45 ecarg@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:45 ecarg@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:44 ecarg@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:44 ecarg@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:43 ecarg@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:43 ecarg@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:41 sukhe: ipvsadm --delete-service --tcp-service 10.2.1.55:8087: lvs2014 and lvs2013 * 14:39 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:39 sukhe: ipvsadm --delete-service --tcp-service 10.2.2.55:8087: [[phab:T432445|T432445]] * 14:38 ecarg@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:38 ecarg@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:37 ecarg@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:37 ecarg@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:36 ecarg@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:34 ecarg@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts datahubsearch1001.eqiad.wmnet * 14:32 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:32 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 14:31 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 14:31 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 14:28 sukhe: sudo cumin 'A:lvs-low-traffic-codfw' 'systemctl restart pybal': lvs2013 * 14:26 sukhe: sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal': lvs2014 * 14:26 sukhe: sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal' * 14:24 sukhe: restart pybal on lvs1019 * 14:24 sukhe: restart pybal on lvs1020 * 14:19 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for 1036 hosts * 14:17 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:04 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] (duration: 09m 28s) * 13:59 kharlan@deploy2003: dreamyjazz, kharlan: Continuing with deployment * 13:58 bking@cumin2003: START - Cookbook sre.hosts.decommission for hosts datahubsearch1001.eqiad.wmnet * 13:57 kharlan@deploy2003: dreamyjazz, kharlan: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:55 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] * 13:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts datahubsearch[1002-1003].eqiad.wmnet * 13:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:53 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch[1002-1003].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 13:52 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch[1002-1003].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 13:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 13:42 stran@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] (duration: 07m 30s) * 13:42 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:40 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: Test * 13:38 stran@deploy2003: dragoniez, stran: Continuing with deployment * 13:37 stran@deploy2003: dragoniez, stran: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:35 bking@cumin2003: START - Cookbook sre.hosts.decommission for hosts datahubsearch[1002-1003].eqiad.wmnet * 13:35 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024'] * 13:35 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 13:35 stran@deploy2003: Started scap sync-world: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] * 13:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 13:28 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024'] * 13:26 sukhe@dns1004: END - running authdns-update * 13:25 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 13:25 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 13:24 sukhe@dns1004: START - running authdns-update * 13:22 stran@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] (duration: 08m 20s) * 13:21 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 13:20 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 13:19 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 13:19 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 13:19 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 13:18 stran@deploy2003: stran: Continuing with deployment * 13:16 stran@deploy2003: stran: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:14 stran@deploy2003: Started scap sync-world: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] * 13:13 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 13:13 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 13:11 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 13:11 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 13:08 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 12:55 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool es2051: Test * 12:55 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: Test * 12:54 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool es2051: Test * 12:43 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1036 hosts * 12:41 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 12:40 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1048.eqiad.wmnet * 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 12:39 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 12:38 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 12:37 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 12:37 brouberol@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 12:36 brouberol@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 12:36 brouberol@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 12:36 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 12:36 elukey@cumin1003: DONE (PASS) - Cookbook sre.puppet.renew-cert (exit_code=0) for crm2001.codfw.wmnet: Renew puppet certificate - elukey@cumin1003 * 12:35 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:35 brouberol@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 12:34 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 12:31 brouberol@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 12:30 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1048.eqiad.wmnet * 12:30 brouberol@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 12:28 brouberol@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 12:27 brouberol@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 12:20 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1068.eqiad.wmnet with OS trixie * 12:01 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] (duration: 13m 19s) * 11:58 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1068.eqiad.wmnet with reason: host reimage * 11:52 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1068.eqiad.wmnet with reason: host reimage * 11:51 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 11:49 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:47 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] * 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2252: Security updates * 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:43 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 11:42 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2252: Security updates * 11:42 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply * 11:40 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply * 11:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 11:37 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2252.codfw.wmnet with OS trixie * 11:34 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1068 * 11:34 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1068 * 11:26 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1068 * 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1068.eqiad.wmnet 46.48.64.10.in-addr.arpa 6.4.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:26 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1068.eqiad.wmnet 46.48.64.10.in-addr.arpa 6.4.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1068 - jiji@cumin1003" * 11:26 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1068 - jiji@cumin1003" * 11:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2252.codfw.wmnet with reason: host reimage * 11:17 jiji@cumin1003: START - Cookbook sre.dns.netbox * 11:17 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1068 * 11:17 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1068.eqiad.wmnet with OS trixie * 11:17 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2252.codfw.wmnet with reason: host reimage * 11:15 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1068.eqiad.wmnet * 11:15 mvolz@deploy2003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:15 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1068.eqiad.wmnet * 11:15 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1068.eqiad.wmnet * 11:14 mvolz@deploy2003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:13 mvolz@deploy2003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:13 mvolz@deploy2003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:12 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] (duration: 11m 05s) * 11:11 mvolz@deploy2003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:10 mvolz@deploy2003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:07 dreamyjazz@deploy2003: dreamyjazz, kharlan: Continuing with deployment * 11:03 dreamyjazz@deploy2003: dreamyjazz, kharlan: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:03 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2252.codfw.wmnet with OS trixie * 11:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2252: Upgrading db2252.codfw.wmnet * 11:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:02 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 11:02 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2252: Upgrading db2252.codfw.wmnet * 11:02 cwilliams@cumin1003: dbmaint on ms3@codfw [[phab:T432321|T432321]] * 11:01 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 11:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db1153.eqiad.wmnet with reason: Security updates * 11:01 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] * 11:00 fnegri@deploy2003: helmfile [eqiad] DONE helmfile.d/services/toolhub: apply * 10:58 fnegri@deploy2003: helmfile [eqiad] START helmfile.d/services/toolhub: apply * 10:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1151: Security updates * 10:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:57 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1151: Security updates * 10:55 fnegri@deploy2003: helmfile [codfw] DONE helmfile.d/services/toolhub: apply * 10:54 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] (duration: 08m 38s) * 10:53 fnegri@deploy2003: helmfile [codfw] START helmfile.d/services/toolhub: apply * 10:53 fnegri@deploy2003: helmfile [staging] DONE helmfile.d/services/toolhub: apply * 10:52 fnegri@deploy2003: helmfile [staging] START helmfile.d/services/toolhub: apply * 10:50 zabe@deploy2003: zabe: Continuing with deployment * 10:47 zabe@deploy2003: zabe: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:45 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] * 10:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1151: Security updates * 10:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:42 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:42 root@cumin1003: START - Cookbook sre.mysql.depool depool db1151: Security updates * 10:38 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] (duration: 12m 47s) * 10:34 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 10:34 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 10:33 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2253.codfw.wmnet with OS trixie * 10:28 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:26 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] * 10:18 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2253.codfw.wmnet with reason: host reimage * 10:13 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2253.codfw.wmnet with reason: host reimage * 10:00 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2253.codfw.wmnet with OS trixie * 09:58 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db1151.eqiad.wmnet with reason: Security updates * 09:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2253: Upgrading db2253.codfw.wmnet * 09:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:57 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 09:56 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2253: Upgrading db2253.codfw.wmnet * 09:56 cwilliams@cumin1003: dbmaint on ms2@codfw [[phab:T432321|T432321]] * 09:56 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 09:36 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: UI improvement; support url shortener - oblivian@cumin1003" * 09:36 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: UI improvement; support url shortener - oblivian@cumin1003 * 09:35 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: UI improvement; support url shortener - oblivian@cumin1003 * 09:35 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: UI improvement; support url shortener - oblivian@cumin1003" * 09:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1152: Security updates * 09:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:26 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db1152: Security updates * 09:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: Security updates * 09:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:11 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:11 root@cumin1003: START - Cookbook sre.mysql.depool depool db1152: Security updates * 09:10 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1018.eqiad.wmnet with reason: Cloning * 09:09 Dreamy_Jazz: Deployed patch for [[phab:T432453|T432453]] and [[phab:T432454|T432454]] * 09:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 09:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2251.codfw.wmnet with OS trixie * 08:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2251.codfw.wmnet with reason: host reimage * 08:45 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2251.codfw.wmnet with reason: host reimage * 08:40 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] (duration: 12m 26s) * 08:38 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1030.eqiad.wmnet,service=s1 * 08:36 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 08:31 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2251.codfw.wmnet with OS trixie * 08:30 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:28 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] * 08:25 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] (duration: 07m 59s) * 08:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2251: Upgrading db2251.codfw.wmnet * 08:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:22 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 08:22 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2251: Upgrading db2251.codfw.wmnet * 08:20 urbanecm@deploy2003: urbanecm: Continuing with deployment * 08:20 cwilliams@cumin1003: dbmaint on ms1@codfw [[phab:T432321|T432321]] * 08:20 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 08:19 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:17 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] * 08:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade * 08:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade * 08:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db2251.codfw.wmnet,db1152.eqiad.wmnet with reason: OS upgrade * 08:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade * 08:13 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade * 08:11 Dreamy_Jazz: Created cusi_signal, cusi_case, and cusi_user on ukwiki and enwikivoyage in extension1 * 08:11 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade * 08:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade * 08:04 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1030.eqiad.wmnet,service=s1 * 08:04 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1030.eqiad.wmnet,service=s1 * 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply * 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply * 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply * 07:51 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply * 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 07:47 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 07:47 phuedx: End of UTC morning backport window * 07:43 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 07:43 phuedx@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] (duration: 13m 44s) * 07:43 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 07:39 phuedx@deploy2003: phuedx: Continuing with deployment * 07:31 phuedx@deploy2003: phuedx: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:29 phuedx@deploy2003: Started scap sync-world: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] * 07:24 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Turnilo import support - oblivian@cumin1003" * 07:24 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import support - oblivian@cumin1003 * 07:23 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import support - oblivian@cumin1003 * 07:23 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Turnilo import support - oblivian@cumin1003" * 06:42 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs2020 after successful Bookworm reimage, data transfer, and postflight validation * 06:42 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2020.codfw.wmnet * 05:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Managing sanitization for wikis bolwiki in section s5 * 05:25 marostegui@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis bolwiki in section s5 * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 41s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-21 == * 22:50 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2019.codfw.wmnet -> wdqs2020.codfw.wmnet, repooling source-only afterwards * 22:47 cwhite: force reboot arclamp2001 - appears to have run out of memory and gone unresponsive * 22:24 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 01m 26s) * 22:24 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 22:23 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 22:22 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024'] * 22:11 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 21:54 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs1024'] * 21:54 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 21:53 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs1024'] * 21:53 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 21:49 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2019.codfw.wmnet -> wdqs2020.codfw.wmnet, repooling source-only afterwards * 20:57 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] (duration: 09m 10s) * 20:55 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1024.eqiad.wmnet with OS bookworm * 20:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2020.codfw.wmnet with OS bookworm * 20:52 krinkle@deploy2003: krinkle: Continuing with deployment * 20:49 krinkle@deploy2003: krinkle: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:47 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] * 20:45 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] (duration: 05m 42s) * 20:44 krinkle@deploy2003: krinkle: Rolling back deployment * 20:41 krinkle@deploy2003: krinkle: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:39 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] * 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2003.codfw.wmnet * 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1003.eqiad.wmnet * 20:33 dani@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] (duration: 11m 15s) * 20:33 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2003.codfw.wmnet * 20:33 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1003.eqiad.wmnet * 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2020.codfw.wmnet with reason: host reimage * 20:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1002.eqiad.wmnet * 20:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2002.codfw.wmnet * 20:30 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 20:29 dani@deploy2003: dani: Continuing with deployment * 20:29 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2020.codfw.wmnet with reason: host reimage * 20:26 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1002.eqiad.wmnet * 20:26 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2002.codfw.wmnet * 20:24 dani@deploy2003: dani: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2001.codfw.wmnet * 20:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1001.eqiad.wmnet * 20:22 dani@deploy2003: Started scap sync-world: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] * 20:22 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 20:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2001.codfw.wmnet * 20:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1001.eqiad.wmnet * 20:14 sbisson@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] (duration: 09m 01s) * 20:11 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2020.codfw.wmnet with OS bookworm * 20:10 sbisson@deploy2003: sbisson: Continuing with deployment * 20:07 sbisson@deploy2003: sbisson: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:05 sbisson@deploy2003: Started scap sync-world: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] * 20:03 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] (duration: 07m 04s) * 20:01 mutante: Gerrit - tomorrow a new SSH host key will appear - it will be {{Gerrit|ed25519}} and has already been added to wmf-laptop. you can verify it here: https://wikitech.wikimedia.org/wiki/Help:SSH_Fingerprints/gerrit.wikimedia.org:29418 ([[phab:T240266|T240266]]) * 19:59 zabe@deploy2003: zabe: Continuing with deployment * 19:58 zabe@deploy2003: zabe: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:56 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] * 19:52 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] (duration: 07m 25s) * 19:48 zabe@deploy2003: zabe: Continuing with deployment * 19:47 zabe@deploy2003: zabe: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:45 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] * 19:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 19:32 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024'] * 19:27 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 19:26 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024'] * 19:26 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 19:24 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024'] * 19:12 ryankemper: [wdqs] [[phab:T430880|T430880]] Repooled `wdqs-scholarly` discovery in `eqiad` after validating `wdqs1023` end-to-end; `wdqs1024` remains disabled pending reimage recovery * 19:11 ryankemper: [wdqs] [[phab:T430880|T430880]] Repooled wdqs1012.eqiad.wmnet after successful Bookworm reimage, data transfer, service checks, readiness probe, and cross-graph federation query validation * 19:10 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 19:10 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1012.eqiad.wmnet * 19:08 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 18:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for deploy1003.eqiad.wmnet * 18:57 kamila@cumin1003: START - Cookbook sre.hosts.remove-downtime for deploy1003.eqiad.wmnet * 18:37 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:37 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding urldownloader service IPs - sukhe@cumin1003" * 18:37 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding urldownloader service IPs - sukhe@cumin1003" * 18:32 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 18:32 dancy@deploy2003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 18:30 sukhe@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 18:27 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 18:24 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1024.eqiad.wmnet with OS bookworm * 18:20 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host deploy1003.eqiad.wmnet with OS bookworm * 18:09 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deploy1003 reimage (duration: 121m 16s) * 18:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 18:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1180: Security updates * 17:55 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wcqs2003.codfw.wmnet * 17:48 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wcqs2003.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1155.eqiad.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1155.eqiad.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2224.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2224.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2217.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2217.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2193.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2193.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2180.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2180.codfw.wmnet * 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1168.eqiad.wmnet * 17:36 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1168.eqiad.wmnet * 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2169.codfw.wmnet * 17:36 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2169.codfw.wmnet * 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1165.eqiad.wmnet * 17:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1165.eqiad.wmnet * 17:35 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2158.codfw.wmnet * 17:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2158.codfw.wmnet * 17:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wcqs1003.eqiad.wmnet * 17:17 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1180: Security updates * 17:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1180.eqiad.wmnet * 17:16 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1180.eqiad.wmnet * 17:15 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp3073.* * 17:13 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wcqs1003.eqiad.wmnet * 17:11 brett@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp3073.esams.wmnet with OS trixie * 17:11 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 17:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1024 * 17:04 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1024 * 17:03 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 17:00 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 16:59 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2242: codfw rack B7 depool for maintenance * 16:59 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 16:43 brett@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp3073.esams.wmnet with reason: host reimage * 16:42 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 16:39 brett@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cp3073.esams.wmnet with reason: host reimage * 16:32 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on deploy1003.eqiad.wmnet with reason: host reimage * 16:27 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on deploy1003.eqiad.wmnet with reason: host reimage * 16:14 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2242: codfw rack B7 depool for maintenance * 16:14 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: codfw rack B7 depool for maintenance * 16:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1023.eqiad.wmnet with OS bookworm * 16:13 brett@cumin2002: START - Cookbook sre.hosts.reimage for host cp3073.esams.wmnet with OS trixie * 16:08 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host deploy1003.eqiad.wmnet with OS bookworm * 16:08 kamila@deploy2003: Locking from deployment [MediaWiki]: deploy1003 reimage * 16:03 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1012.eqiad.wmnet with OS bookworm * 15:48 inflatador: bking@apt1002 `sudo reprepro copy bookworm-wikimedia bullseye-wikimedia jvmquake` [[phab:T430880|T430880]] * 15:39 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp3073.* * 15:39 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 15:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:34 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1311536{{!}}Set $wgMathInternalRestbaseURL explicitly (take 2) (T349582)]] * 15:29 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:29 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2228: codfw rack B7 depool for maintenance * 15:29 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2229: codfw rack B7 depool for maintenance * 15:27 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1180: Security update * 15:25 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Security update * 15:21 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 15:21 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 15:19 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 24s) * 15:19 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:14 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 15:14 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db1180: Security update * 15:13 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-eqiad * 14:48 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-eqiad * 14:44 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2229: codfw rack B7 depool for maintenance * 14:44 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc2017: codfw rack B7 depool for maintenance * 14:44 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:43 cmooney@cumin2003: START - Cookbook sre.mysql.parsercache * 14:43 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool pc2017: codfw rack B7 depool for maintenance * 14:43 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2003.codfw.wmnet * 14:43 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2003.codfw.wmnet * 14:42 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2009.codfw.wmnet * 14:42 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2009.codfw.wmnet * 14:41 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:41 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:40 cmooney@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 29 hosts * 14:40 cmooney@cumin1003: START - Cookbook sre.hosts.remove-downtime for 29 hosts * 14:35 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 14:34 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 14:32 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] (duration: 07m 56s) * 14:29 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ssw1-a[1,8]-codfw with reason: lsw1-b7-codfw JunOS upgrade * 14:28 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 14:28 elukey@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 14:26 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:24 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] * 14:23 topranks: reboot lsw1-b7-codfw to upgrade JunOS (affects all hosts in rack) [[phab:T430928|T430928]] * 14:18 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2003.codfw.wmnet * 14:14 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2009.codfw.wmnet * 14:14 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Security update * 14:13 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2242: codfw rack B7 depool for maintenance * 14:13 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2242: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2228: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2228: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2229: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2229: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc2017: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.parsercache * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool pc2017: codfw rack B7 depool for maintenance * 14:08 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2003.codfw.wmnet * 14:07 cmooney@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on aux-k8s-etcd2004.codfw.wmnet,ml-etcd2001.codfw.wmnet with reason: lsw1-b7-codfw JunOS upgrade * 14:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95005 and previous config saved to /var/cache/conftool/dbconfig/20260721-140620-cwilliams.json * 14:05 cmooney@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2049.codfw.wmnet * 14:05 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-scholarly,name=eqiad * 14:04 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2009.codfw.wmnet * 14:04 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply * 14:04 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply * 14:03 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:03 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2001.codfw.wmnet * 14:03 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2001.codfw.wmnet * 14:02 cmooney@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2049.codfw.wmnet * 14:00 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:00 Dreamy_Jazz: Created cusi_case, cusi_signal, and cusi_user on svwiki, dewiki, jawiki, eswiki * 13:59 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b7-codfw,lsw1-b7-codfw IPv6,lsw1-b7-codfw.mgmt,ssw1-a[1,8]-codfw.mgmt with reason: lsw1-b7-codfw JunOS upgrade * 13:57 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs1023.eqiad.wmnet, repooling source-only afterwards * 13:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 29 hosts with reason: lsw1-b7-codfw JunOS upgrade * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224', diff saved to https://phabricator.wikimedia.org/P95003 and previous config saved to /var/cache/conftool/dbconfig/20260721-135613-cwilliams.json * 13:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1012.eqiad.wmnet with reason: host reimage * 13:53 cmooney@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 1:00:00 on 30 hosts with reason: lsw1-b7-codfw JunOS upgrade * 13:51 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1012.eqiad.wmnet with reason: host reimage * 13:48 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 13:48 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 13:46 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224', diff saved to https://phabricator.wikimedia.org/P95001 and previous config saved to /var/cache/conftool/dbconfig/20260721-134605-cwilliams.json * 13:46 cmooney@cumin1003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti2033.codfw.wmnet * 13:46 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 13:45 elukey: move the Docker Registry's /v2/wikimedia/machinelearning.* prefix to the ml S3 backend - [[phab:T428022|T428022]] * 13:45 cmooney@cumin1003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti2033.codfw.wmnet * 13:45 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 13:43 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:40 jiji@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 13:40 jiji@deploy2003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 13:39 jiji@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 13:39 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 13:38 cmooney@cumin1003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2032.codfw.wmnet * 13:38 jiji@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 13:38 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 13:37 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2032.codfw.wmnet * 13:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95000 and previous config saved to /var/cache/conftool/dbconfig/20260721-133557-cwilliams.json * 13:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1012 * 13:33 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1012 * 13:33 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1012.eqiad.wmnet with OS bookworm * 13:30 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:30 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:28 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94999 and previous config saved to /var/cache/conftool/dbconfig/20260721-132855-cwilliams.json * 13:28 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2224.codfw.wmnet with reason: Maintenance * 13:28 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94998 and previous config saved to /var/cache/conftool/dbconfig/20260721-132826-cwilliams.json * 13:28 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 13:23 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] (duration: 07m 50s) * 13:20 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 13:18 kharlan@deploy2003: kharlan: Continuing with deployment * 13:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217', diff saved to https://phabricator.wikimedia.org/P94996 and previous config saved to /var/cache/conftool/dbconfig/20260721-131817-cwilliams.json * 13:17 brouberol@dns1004: END - running authdns-update * 13:17 kharlan@deploy2003: kharlan: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:15 brouberol@dns1004: START - running authdns-update * 13:15 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] * 13:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94995 and previous config saved to /var/cache/conftool/dbconfig/20260721-131411-cwilliams.json * 13:13 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs1023.eqiad.wmnet, repooling source-only afterwards * 13:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217', diff saved to https://phabricator.wikimedia.org/P94994 and previous config saved to /var/cache/conftool/dbconfig/20260721-130809-cwilliams.json * 13:07 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 13:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180', diff saved to https://phabricator.wikimedia.org/P94993 and previous config saved to /var/cache/conftool/dbconfig/20260721-130404-cwilliams.json * 13:03 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:03 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:02 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 13:02 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 12:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94992 and previous config saved to /var/cache/conftool/dbconfig/20260721-125801-cwilliams.json * 12:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180', diff saved to https://phabricator.wikimedia.org/P94991 and previous config saved to /var/cache/conftool/dbconfig/20260721-125356-cwilliams.json * 12:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94990 and previous config saved to /var/cache/conftool/dbconfig/20260721-125049-cwilliams.json * 12:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2217.codfw.wmnet with reason: Maintenance * 12:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94989 and previous config saved to /var/cache/conftool/dbconfig/20260721-125017-cwilliams.json * 12:48 elukey: bmc cold reboot for lvs1013 and lvs1015 - [[phab:T426180|T426180]] * 12:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94988 and previous config saved to /var/cache/conftool/dbconfig/20260721-124348-cwilliams.json * 12:40 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193', diff saved to https://phabricator.wikimedia.org/P94987 and previous config saved to /var/cache/conftool/dbconfig/20260721-124009-cwilliams.json * 12:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts * 12:33 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts * 12:33 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts * 12:32 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts * 12:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:30 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193', diff saved to https://phabricator.wikimedia.org/P94986 and previous config saved to /var/cache/conftool/dbconfig/20260721-123001-cwilliams.json * 12:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94985 and previous config saved to /var/cache/conftool/dbconfig/20260721-121953-cwilliams.json * 12:17 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs2007.codfw.wmnet with OS bookworm * 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94983 and previous config saved to /var/cache/conftool/dbconfig/20260721-121257-cwilliams.json * 12:12 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2193.codfw.wmnet with reason: Maintenance * 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94982 and previous config saved to /var/cache/conftool/dbconfig/20260721-121239-cwilliams.json * 12:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180', diff saved to https://phabricator.wikimedia.org/P94980 and previous config saved to /var/cache/conftool/dbconfig/20260721-120231-cwilliams.json * 11:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180', diff saved to https://phabricator.wikimedia.org/P94979 and previous config saved to /var/cache/conftool/dbconfig/20260721-115223-cwilliams.json * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94978 and previous config saved to /var/cache/conftool/dbconfig/20260721-114333-cwilliams.json * 11:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1180.eqiad.wmnet with reason: Maintenance * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94977 and previous config saved to /var/cache/conftool/dbconfig/20260721-114305-cwilliams.json * 11:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94976 and previous config saved to /var/cache/conftool/dbconfig/20260721-114215-cwilliams.json * 11:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94975 and previous config saved to /var/cache/conftool/dbconfig/20260721-113530-cwilliams.json * 11:35 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2180.codfw.wmnet with reason: Maintenance * 11:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94974 and previous config saved to /var/cache/conftool/dbconfig/20260721-113501-cwilliams.json * 11:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168', diff saved to https://phabricator.wikimedia.org/P94973 and previous config saved to /var/cache/conftool/dbconfig/20260721-113258-cwilliams.json * 11:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169', diff saved to https://phabricator.wikimedia.org/P94972 and previous config saved to /var/cache/conftool/dbconfig/20260721-112453-cwilliams.json * 11:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168', diff saved to https://phabricator.wikimedia.org/P94971 and previous config saved to /var/cache/conftool/dbconfig/20260721-112250-cwilliams.json * 11:21 XioNoX: put eqiad-drmrs Arelion link in service * 11:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169', diff saved to https://phabricator.wikimedia.org/P94970 and previous config saved to /var/cache/conftool/dbconfig/20260721-111446-cwilliams.json * 11:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94969 and previous config saved to /var/cache/conftool/dbconfig/20260721-111242-cwilliams.json * 11:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 11:10 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 11:07 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1093 hosts * 11:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94968 and previous config saved to /var/cache/conftool/dbconfig/20260721-110548-cwilliams.json * 11:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1168.eqiad.wmnet with reason: Maintenance * 11:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94967 and previous config saved to /var/cache/conftool/dbconfig/20260721-110520-cwilliams.json * 11:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94966 and previous config saved to /var/cache/conftool/dbconfig/20260721-110439-cwilliams.json * 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94964 and previous config saved to /var/cache/conftool/dbconfig/20260721-105632-cwilliams.json * 10:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2169.codfw.wmnet with reason: Maintenance * 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94963 and previous config saved to /var/cache/conftool/dbconfig/20260721-105603-cwilliams.json * 10:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165', diff saved to https://phabricator.wikimedia.org/P94962 and previous config saved to /var/cache/conftool/dbconfig/20260721-105512-cwilliams.json * 10:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158', diff saved to https://phabricator.wikimedia.org/P94961 and previous config saved to /var/cache/conftool/dbconfig/20260721-104555-cwilliams.json * 10:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165', diff saved to https://phabricator.wikimedia.org/P94960 and previous config saved to /var/cache/conftool/dbconfig/20260721-104504-cwilliams.json * 10:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158', diff saved to https://phabricator.wikimedia.org/P94959 and previous config saved to /var/cache/conftool/dbconfig/20260721-103547-cwilliams.json * 10:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94958 and previous config saved to /var/cache/conftool/dbconfig/20260721-103456-cwilliams.json * 10:29 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2229: Upgraded kernel * 10:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94956 and previous config saved to /var/cache/conftool/dbconfig/20260721-102757-cwilliams.json * 10:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on an-redacteddb1001.eqiad.wmnet,clouddb[1015,1025,1028].eqiad.wmnet,db1155.eqiad.wmnet with reason: Maintenance * 10:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1165.eqiad.wmnet with reason: Maintenance * 10:25 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94955 and previous config saved to /var/cache/conftool/dbconfig/20260721-102539-cwilliams.json * 10:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94954 and previous config saved to /var/cache/conftool/dbconfig/20260721-101848-cwilliams.json * 10:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2158.codfw.wmnet with reason: Maintenance * 09:43 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2229: Upgraded kernel * 09:42 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2229.codfw.wmnet * 09:42 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2229.codfw.wmnet * 09:23 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db2229.codfw.wmnet * 09:23 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2229.codfw.wmnet * 08:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2229 [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94948 and previous config saved to /var/cache/conftool/dbconfig/20260721-085724-cwilliams.json * 08:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2214 to s6 primary [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94947 and previous config saved to /var/cache/conftool/dbconfig/20260721-085442-cwilliams.json * 08:53 cezmunsta: Starting s6 codfw failover from db2229 to db2214 - [[phab:T430964|T430964]] * 08:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2214 with weight 0 [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94946 and previous config saved to /var/cache/conftool/dbconfig/20260721-084613-cwilliams.json * 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 22 hosts with reason: Primary switchover s6 [[phab:T430964|T430964]] * 08:32 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1017.eqiad.wmnet,service=s1 * 08:08 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Add subrated circuit rate to interface descriptions - CR1312476 - ayounsi@cumin1003 * 08:06 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Add subrated circuit rate to interface descriptions - CR1312476 - ayounsi@cumin1003 * 07:58 reedy@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] (duration: 12m 55s) * 07:51 reedy@deploy2003: reedy, neriah: Continuing with deployment * 07:51 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 07:51 reedy@deploy2003: reedy, neriah: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:48 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1093 hosts * 07:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm2001.wikimedia.org * 07:45 reedy@deploy2003: Started scap sync-world: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] * 07:43 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 07:42 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm2001.wikimedia.org * 07:23 elukey: upgrade libtiff6 packages on zuul* trixie hosts for security upgrades * 07:22 elukey: upgrade libtiff6 packages on Wikikube trixie workers for security upgrades * 07:14 elukey@deploy2003: helmfile [codfw] DONE helmfile.d/services/proton: sync * 07:13 elukey@deploy2003: helmfile [codfw] START helmfile.d/services/proton: sync * 07:11 elukey@deploy2003: helmfile [eqiad] DONE helmfile.d/services/proton: sync * 07:10 elukey@deploy2003: helmfile [eqiad] START helmfile.d/services/proton: sync * 07:09 elukey@deploy2003: helmfile [staging] DONE helmfile.d/services/proton: sync * 07:08 elukey@deploy2003: helmfile [staging] START helmfile.d/services/proton: sync * 06:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1023.eqiad.wmnet with reason: host reimage * 06:46 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1023.eqiad.wmnet with reason: host reimage * 06:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 05:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Haproxy-only mode support - oblivian@cumin1003" * 05:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Haproxy-only mode support - oblivian@cumin1003 * 05:42 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Haproxy-only mode support - oblivian@cumin1003 * 05:42 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Haproxy-only mode support - oblivian@cumin1003" * 05:38 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1017.eqiad.wmnet with reason: Cloning * 05:37 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1017.eqiad.wmnet,service=s1 * 05:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1029.eqiad.wmnet,service=s8 * 05:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1029.eqiad.wmnet,service=s5 * 05:32 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:30 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 05:11 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet * 05:04 aokoth@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet * 05:00 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 04:56 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 04:01 mwpresync@deploy2003: Pruned MediaWiki: 1.47.0-wmf.9 (duration: 01m 08s) * 03:41 mwpresync@deploy2003: Finished scap sync-world: testwikis to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] (duration: 36m 30s) * 03:05 mwpresync@deploy2003: Started scap sync-world: testwikis to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 03:01 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:01 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:00 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:00 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:36 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:36 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:36 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:35 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:16 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 47s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 00:56 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm == 2026-07-20 == * 23:38 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 23:07 Amir1: deleting echo notifications from 2015 on group1 wikis * 23:07 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] (duration: 14m 16s) * 23:01 ladsgroup@deploy2003: ladsgroup: Continuing with deployment * 23:00 ladsgroup@deploy2003: ladsgroup: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:53 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] * 22:46 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2007.codfw.wmnet, repooling source-only afterwards * 22:39 maryum: Deployed security fixes for several security bugs * 21:42 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 21:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2007.codfw.wmnet, repooling source-only afterwards * 21:37 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 21:37 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 21:34 sbassett: Deployed security fix for [[phab:T432424|T432424]] * 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs2020.codfw.wmnet * 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1023.eqiad.wmnet * 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1011.eqiad.wmnet * 21:32 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 17s) * 21:32 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 21:27 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 21:13 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2007.codfw.wmnet with reason: host reimage * 21:08 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-internal-main,name=codfw * 21:06 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2007.codfw.wmnet with reason: host reimage * 20:59 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service * 20:58 sukhe: pybal restart for IP changes around wdqs-main hosts * 20:57 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 20:46 ryankemper@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-internal-main,name=codfw * 20:45 ebernhardson@deploy2003: Finished deploy [search/mjolnir/deploy@d4dc3b8]: Update for opensearch 2.x compat (duration: 00m 34s) * 20:44 ebernhardson@deploy2003: Started deploy [search/mjolnir/deploy@d4dc3b8]: Update for opensearch 2.x compat * 20:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2007 * 20:44 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2007 * 20:43 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2007 * 20:43 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2007.codfw.wmnet 156.16.192.10.in-addr.arpa 6.5.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:42 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2007.codfw.wmnet 156.16.192.10.in-addr.arpa 6.5.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:42 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2007 - bking@cumin2003" * 20:41 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2007 - bking@cumin2003" * 20:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94944 and previous config saved to /var/cache/conftool/dbconfig/20260720-203333-cwilliams.json * 20:32 arlolra@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] (duration: 15m 07s) * 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2020.codfw.wmnet * 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1023.eqiad.wmnet * 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1011.eqiad.wmnet * 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs2020.codfw.wmnet * 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1023.eqiad.wmnet * 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1011.eqiad.wmnet * 20:25 arlolra@deploy2003: arlolra, cscott: Continuing with deployment * 20:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257', diff saved to https://phabricator.wikimedia.org/P94943 and previous config saved to /var/cache/conftool/dbconfig/20260720-202325-cwilliams.json * 20:21 arlolra@deploy2003: arlolra, cscott: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:17 arlolra@deploy2003: Started scap sync-world: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] * 20:13 bking@cumin2003: START - Cookbook sre.dns.netbox * 20:13 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257', diff saved to https://phabricator.wikimedia.org/P94942 and previous config saved to /var/cache/conftool/dbconfig/20260720-201318-cwilliams.json * 20:13 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 20:10 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 20:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2007 * 20:04 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2007.codfw.wmnet with OS bookworm * 20:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94941 and previous config saved to /var/cache/conftool/dbconfig/20260720-200310-cwilliams.json * 19:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94940 and previous config saved to /var/cache/conftool/dbconfig/20260720-195633-cwilliams.json * 19:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1257.eqiad.wmnet with reason: Maintenance * 19:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94939 and previous config saved to /var/cache/conftool/dbconfig/20260720-195605-cwilliams.json * 19:51 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 19:50 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 19:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256', diff saved to https://phabricator.wikimedia.org/P94938 and previous config saved to /var/cache/conftool/dbconfig/20260720-194558-cwilliams.json * 19:44 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 19:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 19:41 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wdqs1011.eqiad.wmnet with OS bookworm * 19:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256', diff saved to https://phabricator.wikimedia.org/P94937 and previous config saved to /var/cache/conftool/dbconfig/20260720-193550-cwilliams.json * 19:25 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94936 and previous config saved to /var/cache/conftool/dbconfig/20260720-192542-cwilliams.json * 19:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94935 and previous config saved to /var/cache/conftool/dbconfig/20260720-191856-cwilliams.json * 19:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1256.eqiad.wmnet with reason: Maintenance * 19:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94934 and previous config saved to /var/cache/conftool/dbconfig/20260720-191839-cwilliams.json * 19:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255', diff saved to https://phabricator.wikimedia.org/P94933 and previous config saved to /var/cache/conftool/dbconfig/20260720-190831-cwilliams.json * 18:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255', diff saved to https://phabricator.wikimedia.org/P94932 and previous config saved to /var/cache/conftool/dbconfig/20260720-185824-cwilliams.json * 18:50 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 18:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94931 and previous config saved to /var/cache/conftool/dbconfig/20260720-184816-cwilliams.json * 18:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94930 and previous config saved to /var/cache/conftool/dbconfig/20260720-184224-cwilliams.json * 18:42 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1255.eqiad.wmnet with reason: Maintenance * 18:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94929 and previous config saved to /var/cache/conftool/dbconfig/20260720-184153-cwilliams.json * 18:39 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 18:39 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 16s) * 18:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 18:38 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 59m 26s) * 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 18:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211', diff saved to https://phabricator.wikimedia.org/P94928 and previous config saved to /var/cache/conftool/dbconfig/20260720-183145-cwilliams.json * 18:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211', diff saved to https://phabricator.wikimedia.org/P94927 and previous config saved to /var/cache/conftool/dbconfig/20260720-182137-cwilliams.json * 18:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94926 and previous config saved to /var/cache/conftool/dbconfig/20260720-181129-cwilliams.json * 18:09 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_codfw * 18:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2057.codfw.wmnet * 18:08 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_codfw * 18:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2058.codfw.wmnet * 18:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94925 and previous config saved to /var/cache/conftool/dbconfig/20260720-180452-cwilliams.json * 18:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on clouddb[1016,1020,1022-1023].eqiad.wmnet,db1154.eqiad.wmnet with reason: Maintenance * 18:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1211.eqiad.wmnet with reason: Maintenance * 18:02 sukhe: armed keyholder on acmechief1002.eqiad.wmnet and acmechief2002.codfw.wmnet (active host) * 18:01 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief2002.codfw.wmnet * 17:57 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief2002.codfw.wmnet * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs2020'] * 17:52 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief1002.eqiad.wmnet * 17:50 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 17:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1011.eqiad.wmnet with reason: host reimage * 17:48 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief1002.eqiad.wmnet * 17:47 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test2001.codfw.wmnet * 17:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94924 and previous config saved to /var/cache/conftool/dbconfig/20260720-174717-cwilliams.json * 17:46 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 17:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1011.eqiad.wmnet with reason: host reimage * 17:43 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs2020.codfw.wmnet with OS bookworm * 17:43 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test2001.codfw.wmnet * 17:43 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test1001.eqiad.wmnet * 17:39 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test1001.eqiad.wmnet * 17:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 17:38 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:38 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 17:37 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 17:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244', diff saved to https://phabricator.wikimedia.org/P94923 and previous config saved to /var/cache/conftool/dbconfig/20260720-173709-cwilliams.json * 17:35 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 17:31 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 20m 40s) * 17:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2055.codfw.wmnet * 17:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2056.codfw.wmnet * 17:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1011 * 17:27 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1011 * 17:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1011.eqiad.wmnet with OS bookworm * 17:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244', diff saved to https://phabricator.wikimedia.org/P94922 and previous config saved to /var/cache/conftool/dbconfig/20260720-172701-cwilliams.json * 17:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94921 and previous config saved to /var/cache/conftool/dbconfig/20260720-171653-cwilliams.json * 17:11 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 17:11 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 13m 03s) * 17:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94920 and previous config saved to /var/cache/conftool/dbconfig/20260720-171012-cwilliams.json * 17:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2244.codfw.wmnet with reason: Maintenance * 17:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94919 and previous config saved to /var/cache/conftool/dbconfig/20260720-170941-cwilliams.json * 16:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243', diff saved to https://phabricator.wikimedia.org/P94918 and previous config saved to /var/cache/conftool/dbconfig/20260720-165933-cwilliams.json * 16:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 16:58 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2053.codfw.wmnet * 16:51 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2054.codfw.wmnet * 16:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243', diff saved to https://phabricator.wikimedia.org/P94917 and previous config saved to /var/cache/conftool/dbconfig/20260720-164926-cwilliams.json * 16:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94916 and previous config saved to /var/cache/conftool/dbconfig/20260720-163918-cwilliams.json * 16:35 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 16:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94915 and previous config saved to /var/cache/conftool/dbconfig/20260720-163140-cwilliams.json * 16:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2243.codfw.wmnet with reason: Maintenance * 16:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94914 and previous config saved to /var/cache/conftool/dbconfig/20260720-163111-cwilliams.json * 16:27 btullis@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 16:27 btullis@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 16:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2020 * 16:23 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2020 * 16:21 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2020 * 16:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2020.codfw.wmnet 85.0.192.10.in-addr.arpa 5.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:21 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2020.codfw.wmnet 85.0.192.10.in-addr.arpa 5.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242', diff saved to https://phabricator.wikimedia.org/P94913 and previous config saved to /var/cache/conftool/dbconfig/20260720-162103-cwilliams.json * 16:19 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 16:18 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 16:18 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:18 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 16:17 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:17 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for netbox accounting errors - jhancock@cumin2002" * 16:17 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for netbox accounting errors - jhancock@cumin2002" * 16:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2051.codfw.wmnet * 16:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2052.codfw.wmnet * 16:11 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 16:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242', diff saved to https://phabricator.wikimedia.org/P94912 and previous config saved to /var/cache/conftool/dbconfig/20260720-161055-cwilliams.json * 16:09 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 16:08 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 16:06 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 16:06 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 16:06 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2020 * 16:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2020.codfw.wmnet with OS bookworm * 16:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94911 and previous config saved to /var/cache/conftool/dbconfig/20260720-160047-cwilliams.json * 15:58 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2019.codfw.wmnet, repooling source-only afterwards * 15:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94909 and previous config saved to /var/cache/conftool/dbconfig/20260720-155353-cwilliams.json * 15:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2242.codfw.wmnet with reason: Maintenance * 15:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94908 and previous config saved to /var/cache/conftool/dbconfig/20260720-154433-cwilliams.json * 15:35 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2049.codfw.wmnet * 15:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162', diff saved to https://phabricator.wikimedia.org/P94907 and previous config saved to /var/cache/conftool/dbconfig/20260720-153425-cwilliams.json * 15:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2050.codfw.wmnet * 15:28 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162', diff saved to https://phabricator.wikimedia.org/P94906 and previous config saved to /var/cache/conftool/dbconfig/20260720-152418-cwilliams.json * 15:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1023 * 15:14 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1023 * 15:14 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 15:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94905 and previous config saved to /var/cache/conftool/dbconfig/20260720-151407-cwilliams.json * 15:13 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] (duration: 41m 16s) * 15:08 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 15:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94902 and previous config saved to /var/cache/conftool/dbconfig/20260720-150729-cwilliams.json * 15:07 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2162.codfw.wmnet with reason: Maintenance * 15:05 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2027.codfw.wmnet, repooling source-only afterwards * 15:00 urbanecm@deploy2003: vadymts1, migr, urbanecm: Continuing with deployment * 14:59 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:58 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 07s) * 14:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:58 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 13s) * 14:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:57 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2019.codfw.wmnet, repooling source-only afterwards * 14:57 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2047.codfw.wmnet * 14:55 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2048.codfw.wmnet * 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2019.codfw.wmnet with OS bookworm * 14:49 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 14:47 urbanecm@deploy2003: vadymts1, migr, urbanecm: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:44 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool magru [reason: BGP issues in lvs7003 resolved after liberica restart, no task ID specified] * 14:44 sukhe@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool magru [reason: BGP issues in lvs7003 resolved after liberica restart, no task ID specified] * 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:41 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:39 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:39 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:33 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool magru [reason: no reason specified, no task ID specified] * 14:33 sukhe@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool magru [reason: no reason specified, no task ID specified] * 14:31 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] * 14:24 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:24 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:24 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:24 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2019.codfw.wmnet with reason: host reimage * 14:22 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2027.codfw.wmnet, repooling source-only afterwards * 14:19 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2019.codfw.wmnet with reason: host reimage * 14:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2027.codfw.wmnet with OS bookworm * 14:16 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2046.codfw.wmnet * 14:16 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2045.codfw.wmnet * 14:08 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:08 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:08 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:08 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:07 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:06 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:06 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:06 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:05 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2019 * 14:00 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2019 * 13:56 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1015.eqiad.wmnet * 13:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2027.codfw.wmnet with reason: host reimage * 13:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2071.codfw.wmnet with OS trixie * 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:51 sukhe@cumin1003: END (ERROR) - Cookbook sre.loadbalancer.admin (exit_code=97) rebooting A:liberica and P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica and P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:51 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1015.eqiad.wmnet * 13:50 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1014.eqiad.wmnet * 13:50 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1076.eqiad.wmnet with OS trixie * 13:50 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2027.codfw.wmnet with reason: host reimage * 13:45 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1014.eqiad.wmnet * 13:44 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1013.eqiad.wmnet * 13:39 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 13:38 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1013.eqiad.wmnet * 13:37 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2044.codfw.wmnet * 13:37 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2043.codfw.wmnet * 13:36 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2019 * 13:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2019.codfw.wmnet 156.32.192.10.in-addr.arpa 6.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:36 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2019.codfw.wmnet 156.32.192.10.in-addr.arpa 6.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:36 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2019 - bking@cumin2003" * 13:36 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2019 - bking@cumin2003" * 13:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry2005.codfw.wmnet * 13:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2071.codfw.wmnet with reason: host reimage * 13:31 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:31 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2019 * 13:31 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry2005.codfw.wmnet * 13:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry2004.codfw.wmnet * 13:30 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2019.codfw.wmnet with OS bookworm * 13:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2027 * 13:30 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2027 * 13:30 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2027.codfw.wmnet with OS bookworm * 13:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1076.eqiad.wmnet with reason: host reimage * 13:29 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_codfw * 13:28 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_codfw * 13:26 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry2004.codfw.wmnet * 13:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry1005.eqiad.wmnet * 13:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2071.codfw.wmnet with reason: host reimage * 13:22 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1076.eqiad.wmnet with reason: host reimage * 13:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry1005.eqiad.wmnet * 13:21 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry1004.eqiad.wmnet * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry1004.eqiad.wmnet * 13:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts * 13:13 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts * 13:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts * 13:12 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts * 13:03 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1076.eqiad.wmnet with OS trixie * 13:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2071.codfw.wmnet with OS trixie * 12:55 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:54 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:53 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:46 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 7 hosts * 12:42 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 7 hosts * 12:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts * 12:42 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts * 12:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2070.codfw.wmnet with OS trixie * 12:36 ozge@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:35 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1075.eqiad.wmnet with OS trixie * 12:32 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts * 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts * 12:22 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin1001.eqiad.wmnet * 12:19 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin1001.eqiad.wmnet * 12:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2070.codfw.wmnet with reason: host reimage * 12:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin2001.codfw.wmnet * 12:14 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1075.eqiad.wmnet with reason: host reimage * 12:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2070.codfw.wmnet with reason: host reimage * 12:10 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1075.eqiad.wmnet with reason: host reimage * 12:09 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin2001.codfw.wmnet * 11:17 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1074.eqiad.wmnet with OS trixie * 11:17 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 11:16 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 11:14 ozge@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 6 hosts * 11:09 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 6 hosts * 11:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 324 hosts * 10:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1074.eqiad.wmnet with reason: host reimage * 10:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2069.codfw.wmnet with OS trixie * 10:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1074.eqiad.wmnet with reason: host reimage * 10:30 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2069.codfw.wmnet with reason: host reimage * 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1074.eqiad.wmnet with OS trixie * 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2069.codfw.wmnet with reason: host reimage * 10:06 blake@deploy2003: Stopping before sync operations * 10:06 blake@deploy2003: Started scap sync-world: Non-deployment scap run to populate new release values * 10:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2069.codfw.wmnet with OS trixie * 10:00 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1073.eqiad.wmnet with OS trixie * 09:56 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 324 hosts * 09:39 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 09:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 8 hosts * 09:38 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1073.eqiad.wmnet with reason: host reimage * 09:37 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 8 hosts * 09:34 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1073.eqiad.wmnet with reason: host reimage * 09:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2068.codfw.wmnet with OS trixie * 09:16 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1073.eqiad.wmnet with OS trixie * 09:13 blake@deploy2003: sync-world aborted: Non-deployment scap run to populate new release values (duration: 00m 02s) * 09:13 blake@deploy2003: Started scap sync-world: Non-deployment scap run to populate new release values * 08:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2068.codfw.wmnet with reason: host reimage * 08:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2068.codfw.wmnet with reason: host reimage * 08:50 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 08:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2068.codfw.wmnet with OS trixie * 08:15 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1072.eqiad.wmnet with OS trixie * 07:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2067.codfw.wmnet with OS trixie * 07:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1072.eqiad.wmnet with reason: host reimage * 07:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1072.eqiad.wmnet with reason: host reimage * 07:45 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 07:45 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 07:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2067.codfw.wmnet with reason: host reimage * 07:35 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2067.codfw.wmnet with reason: host reimage * 07:30 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 07:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1072.eqiad.wmnet with OS trixie * 07:30 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 07:17 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 07:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2067.codfw.wmnet with OS trixie * 05:51 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:50 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:25 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:25 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on db2207.codfw.wmnet with reason: Host down * 04:28 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 07m 02s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-18 == * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 29s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 00:11 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2018.codfw.wmnet, repooling source-only afterwards == 2026-07-17 == * 23:53 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2026.codfw.wmnet, repooling source-only afterwards * 23:09 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2018.codfw.wmnet, repooling source-only afterwards * 23:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2026.codfw.wmnet, repooling source-only afterwards * 22:11 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2018.codfw.wmnet with OS bookworm * 22:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2026.codfw.wmnet with OS bookworm * 21:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2018.codfw.wmnet with reason: host reimage * 21:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2018.codfw.wmnet with reason: host reimage * 21:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2026.codfw.wmnet with reason: host reimage * 21:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2026.codfw.wmnet with reason: host reimage * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2018 * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2018 * 21:26 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2018 * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2018.codfw.wmnet 155.32.192.10.in-addr.arpa 5.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:26 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2018.codfw.wmnet 155.32.192.10.in-addr.arpa 5.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2018 - bking@cumin2003" * 21:26 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2018 - bking@cumin2003" * 21:14 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:13 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2018 * 21:13 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2018.codfw.wmnet with OS bookworm * 21:12 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2026 * 21:12 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2026 * 21:12 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2026.codfw.wmnet with OS bookworm * 21:05 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1022.eqiad.wmnet -> wdqs1026.eqiad.wmnet, repooling source-only afterwards * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs2017.codfw.wmnet, repooling source-only afterwards * 20:11 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs2017.codfw.wmnet, repooling source-only afterwards * 20:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2017.codfw.wmnet with OS bookworm * 20:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1022.eqiad.wmnet -> wdqs1026.eqiad.wmnet, repooling source-only afterwards * 20:06 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1026.eqiad.wmnet with OS bookworm * 19:55 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 19:55 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:55 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 09s) * 19:55 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:50 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 08s) * 19:50 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:50 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 10m 03s) * 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2017.codfw.wmnet with reason: host reimage * 19:40 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:40 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 15s) * 19:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1026.eqiad.wmnet with reason: host reimage * 19:37 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 16s) * 19:37 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:34 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2017.codfw.wmnet with reason: host reimage * 19:34 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1026.eqiad.wmnet with reason: host reimage * 19:33 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 19:33 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:16 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1026.eqiad.wmnet with OS bookworm * 19:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2017 * 19:16 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2017 * 19:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2017.codfw.wmnet with OS bookworm * 18:30 bking@dns1004: END - running authdns-update * 18:28 bking@dns1004: START - running authdns-update * 18:16 kamila@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1264.eqiad.wmnet * 18:16 kamila@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1264.eqiad.wmnet * 18:16 kamila@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1264.eqiad.wmnet * 17:49 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 17:46 dzahn@dns1006: END - running authdns-update * 17:44 dzahn@dns1006: START - running authdns-update * 17:44 dzahn@dns1006: END - running authdns-update * 17:42 dzahn@dns1006: START - running authdns-update * 17:28 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 17:21 kamila@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 17:01 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1264 * 17:01 kamila@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1264 * 17:01 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 17:01 kamila@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1264.eqiad.wmnet * 17:01 kamila@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1264.eqiad.wmnet * 17:01 kamila@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1264.eqiad.wmnet * 16:42 reedy@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] (duration: 10m 29s) * 16:34 reedy@deploy2003: reedy, hartman: Continuing with deployment * 16:33 reedy@deploy2003: reedy, hartman: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:31 reedy@deploy2003: Started scap sync-world: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] * 16:26 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 16:10 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in2001.wikimedia.org with reason: [[phab:T431659|T431659]] * 16:07 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in1001.wikimedia.org with reason: [[phab:T431659|T431659]] * 16:05 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 16:01 kamila@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 16:00 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out2001.wikimedia.org with reason: [[phab:T431659|T431659]] * 15:41 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 15:41 kamila@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 15:35 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out1001.wikimedia.org with reason: [[phab:T431659|T431659]] * 15:14 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1339.eqiad.wmnet * 15:13 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1339.eqiad.wmnet * 15:13 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1339.eqiad.wmnet * 14:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1339.eqiad.wmnet with OS trixie * 14:50 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:49 kamila@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:49 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:33 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage * 14:27 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage * 14:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1339 * 14:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1339 * 14:14 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1339 * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1339.eqiad.wmnet 156.32.64.10.in-addr.arpa 6.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:14 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1339.eqiad.wmnet 156.32.64.10.in-addr.arpa 6.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1339 - cgoubert@cumin2003" * 14:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1339 - cgoubert@cumin2003" * 14:09 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 14:06 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1339 * 14:06 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie * 14:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1339.eqiad.wmnet * 14:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1339.eqiad.wmnet * 14:02 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1339.eqiad.wmnet * 13:45 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb1013.eqiad.wmnet * 13:39 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb1013.eqiad.wmnet * 13:27 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:24 blake@dns1004: END - running authdns-update * 13:22 blake@dns1004: START - running authdns-update * 13:20 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 13:11 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2014.codfw.wmnet * 13:06 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb2014.codfw.wmnet * 13:06 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2012.codfw.wmnet * 13:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 13:03 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 13:01 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 15 hosts * 13:01 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb2012.codfw.wmnet * 13:01 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1016.eqiad.wmnet * 13:00 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 15 hosts * 12:55 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb1016.eqiad.wmnet * 12:55 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1014.eqiad.wmnet * 12:49 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb1014.eqiad.wmnet * 12:32 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:32 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:31 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:31 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1338.eqiad.wmnet * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1338.eqiad.wmnet * 12:18 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1338.eqiad.wmnet * 12:17 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:15 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:14 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:13 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1338.eqiad.wmnet with OS trixie * 12:01 klausman@deploy2003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 11:59 klausman@deploy2003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 11:56 klausman@deploy2003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 11:54 klausman@deploy2003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 11:53 klausman@deploy2003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 11:51 klausman@deploy2003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 11:42 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1338.eqiad.wmnet with reason: host reimage * 11:38 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1338.eqiad.wmnet with reason: host reimage * 11:31 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2230.codfw.wmnet * 11:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1338 * 11:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1338 * 11:25 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1338 * 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1338.eqiad.wmnet 155.32.64.10.in-addr.arpa 5.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:25 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1338.eqiad.wmnet 155.32.64.10.in-addr.arpa 5.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1338 - cgoubert@cumin2003" * 11:25 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1338 - cgoubert@cumin2003" * 11:23 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2230.codfw.wmnet * 11:20 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 11:20 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1338 * 11:20 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1338.eqiad.wmnet with OS trixie * 11:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1338.eqiad.wmnet * 11:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1338.eqiad.wmnet * 11:19 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1338.eqiad.wmnet * 11:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1337.eqiad.wmnet * 11:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1337.eqiad.wmnet * 11:17 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1337.eqiad.wmnet * 11:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1337.eqiad.wmnet with OS trixie * 10:51 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[2001-2002].codfw.wmnet * 10:50 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1337.eqiad.wmnet with reason: host reimage * 10:40 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 10:39 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:39 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1337.eqiad.wmnet with reason: host reimage * 10:39 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1001-1003].eqiad.wmnet * 10:34 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:34 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:30 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:28 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1001-1003].eqiad.wmnet * 10:27 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1337 * 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1337 * 10:26 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1337 * 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1337.eqiad.wmnet 154.32.64.10.in-addr.arpa 4.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:26 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1337.eqiad.wmnet 154.32.64.10.in-addr.arpa 4.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1337 - cgoubert@cumin2003" * 10:26 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1337 - cgoubert@cumin2003" * 10:21 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 10:18 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1337 * 10:17 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1337.eqiad.wmnet with OS trixie * 10:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1337.eqiad.wmnet * 10:16 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db1176.eqiad.wmnet * 10:16 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1337.eqiad.wmnet * 10:16 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1337.eqiad.wmnet * 10:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1336.eqiad.wmnet * 10:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1336.eqiad.wmnet * 10:15 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1336.eqiad.wmnet * 10:11 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db1176.eqiad.wmnet * 10:10 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db1176.eqiad.wmnet * 10:09 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db1176.eqiad.wmnet * 10:05 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts (check the cookbook's logs for more details.) * 10:03 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts (check the cookbook's logs for more details.) * 09:58 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1336.eqiad.wmnet with OS trixie * 09:47 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts (check the cookbook's logs for more details.) * 09:47 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts (check the cookbook's logs for more details.) * 09:45 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host acmechief-test2001.codfw.wmnet,acmechief-test1001.eqiad.wmnet,an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet,db-test[2001-2002].codfw.wmnet,db-test[1 * 09:40 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host acmechief-test2001.codfw.wmnet,acmechief-test1001.eqiad.wmnet,an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet,db-test[2001-2002].codfw.wmnet,db-test[1001-1003].eqiad.wmn * 09:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1336.eqiad.wmnet with reason: host reimage * 09:33 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1336.eqiad.wmnet with reason: host reimage * 09:29 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet * 09:29 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet * 09:28 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:26 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:21 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 09:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1336 * 09:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1336 * 09:19 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 09:14 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1336 * 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1336.eqiad.wmnet 152.32.64.10.in-addr.arpa 2.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1336.eqiad.wmnet 152.32.64.10.in-addr.arpa 2.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1336 - cgoubert@cumin2003" * 09:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1336 - cgoubert@cumin2003" * 09:11 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:10 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 09:09 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:09 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1336 * 09:09 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1336.eqiad.wmnet with OS trixie * 09:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1336.eqiad.wmnet * 09:08 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1336.eqiad.wmnet * 09:08 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1336.eqiad.wmnet * 09:06 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1335.eqiad.wmnet * 09:06 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1335.eqiad.wmnet * 09:06 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1335.eqiad.wmnet * 09:04 elukey: uploaded spicerack_13.1.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia * 08:55 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wikikube-worker-exp2001.codfw.wmnet * 08:54 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host testreduce1002.eqiad.wmnet * 08:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1335.eqiad.wmnet with OS trixie * 08:51 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host wikikube-worker-exp2001.codfw.wmnet * 08:51 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wikikube-worker-exp1001.eqiad.wmnet * 08:50 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host testreduce1002.eqiad.wmnet * 08:45 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host wikikube-worker-exp1001.eqiad.wmnet * 08:34 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1335.eqiad.wmnet with reason: host reimage * 08:30 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1335.eqiad.wmnet with reason: host reimage * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1335 * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1335 * 08:18 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1335 * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1335.eqiad.wmnet 150.32.64.10.in-addr.arpa 0.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:18 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1335.eqiad.wmnet 150.32.64.10.in-addr.arpa 0.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1335 - cgoubert@cumin2003" * 08:18 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1335 - cgoubert@cumin2003" * 08:14 elukey@cumin1003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 08:14 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges * 08:13 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 08:10 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1335 * 08:10 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1335.eqiad.wmnet with OS trixie * 08:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1335.eqiad.wmnet * 08:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1335.eqiad.wmnet * 08:09 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1335.eqiad.wmnet * 08:06 elukey@cumin1003: END (FAIL) - Cookbook sre.puppet.disable-merges (exit_code=99) * 08:05 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges * 08:03 elukey@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin1003.eqiad.wmnet * 07:57 elukey@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin1003.eqiad.wmnet * 07:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetdb1003.eqiad.wmnet * 07:46 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetdb1003.eqiad.wmnet * 07:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetdb2003.codfw.wmnet * 07:37 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetdb2003.codfw.wmnet * 07:37 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1001.eqiad.wmnet * 07:28 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver1001.eqiad.wmnet * 07:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet * 07:19 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet * 07:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2002.codfw.wmnet * 07:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver2002.codfw.wmnet * 07:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet * 07:05 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet * 07:04 btullis@cumin1003: END (FAIL) - Cookbook sre.hadoop.reboot-workers (exit_code=99) for Hadoop analytics cluster * 07:04 elukey@cumin1003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 07:04 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges * 06:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox1003.eqiad.wmnet * 06:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox1003.eqiad.wmnet * 02:46 ryankemper@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:46 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:44 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:37 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:37 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-internal-scholarly,name=eqiad * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 49s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 01:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore wdqs1025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling source-only afterwards * 01:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore wdqs1027 after Bookworm reimage) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs1027.eqiad.wmnet, repooling both afterwards * 00:55 urbanecm@deploy2003: helmfile [codfw] DONE helmfile.d/services/linkrecommendation: apply * 00:54 urbanecm@deploy2003: helmfile [eqiad] DONE helmfile.d/services/linkrecommendation: apply * 00:54 urbanecm@deploy2003: helmfile [staging] DONE helmfile.d/services/linkrecommendation: apply * 00:54 urbanecm@deploy2003: helmfile [codfw] START helmfile.d/services/linkrecommendation: apply * 00:53 urbanecm@deploy2003: helmfile [staging] START helmfile.d/services/linkrecommendation: apply * 00:52 urbanecm@deploy2003: helmfile [eqiad] START helmfile.d/services/linkrecommendation: apply * 00:23 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore wdqs1025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling source-only afterwards * 00:23 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore wdqs1027 after Bookworm reimage) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs1027.eqiad.wmnet, repooling both afterwards * 00:14 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1274.eqiad.wmnet * 00:14 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1274.eqiad.wmnet * 00:14 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1274.eqiad.wmnet * 00:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1027.eqiad.wmnet with OS bookworm * 00:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1025.eqiad.wmnet with OS bookworm * 00:04 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1274.eqiad.wmnet with OS trixie == 2026-07-16 == * 23:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], xfer to freshly reimaged/scap-deployed wdqs2025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs2025.codfw.wmnet, repooling source-only afterwards * 23:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1027.eqiad.wmnet with reason: host reimage * 23:47 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1025.eqiad.wmnet with reason: host reimage * 23:43 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1274.eqiad.wmnet with reason: host reimage * 23:41 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1025.eqiad.wmnet with reason: host reimage * 23:39 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1027.eqiad.wmnet with reason: host reimage * 23:38 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1274.eqiad.wmnet with reason: host reimage * 23:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1025 * 23:23 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1025 * 23:22 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1027 * 23:22 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1027 * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1274 * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1274 * 23:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1025.eqiad.wmnet with OS bookworm * 23:19 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1274 * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1274.eqiad.wmnet 145.48.64.10.in-addr.arpa 5.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:19 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1274.eqiad.wmnet 145.48.64.10.in-addr.arpa 5.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1274 - swfrench@cumin1003" * 23:19 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1274 - swfrench@cumin1003" * 23:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1027.eqiad.wmnet with OS bookworm * 23:14 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 23:14 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1274 * 23:13 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1274.eqiad.wmnet with OS trixie * 23:13 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1274.eqiad.wmnet * 23:12 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1274.eqiad.wmnet * 23:12 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1274.eqiad.wmnet * 23:12 ryankemper: [[phab:T430880|T430880]] depooled dnsdisc of wdqs-internal-scholarly-eqiad bc we only have 1 host there * 23:09 ryankemper@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-internal-scholarly,name=eqiad * 23:08 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1272.eqiad.wmnet * 23:08 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1272.eqiad.wmnet * 23:08 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1272.eqiad.wmnet * 23:01 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], xfer to freshly reimaged/scap-deployed wdqs2025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs2025.codfw.wmnet, repooling source-only afterwards * 22:57 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1272.eqiad.wmnet with OS trixie * 22:56 Amir1: deleting echo notifications from 2015 in group0 * 22:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2025.codfw.wmnet with OS bookworm * 22:35 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1272.eqiad.wmnet with reason: host reimage * 22:32 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 27s) * 22:32 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 22:28 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1269.eqiad.wmnet * 22:28 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1269.eqiad.wmnet * 22:28 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1269.eqiad.wmnet * 22:27 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1272.eqiad.wmnet with reason: host reimage * 22:26 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] (duration: 08m 51s) * 22:22 ladsgroup@deploy2003: ladsgroup, urbanecm: Continuing with deployment * 22:19 ladsgroup@deploy2003: ladsgroup, urbanecm: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:17 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] * 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2025.codfw.wmnet with reason: host reimage * 22:06 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1272 * 22:06 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1272 * 22:05 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1272 * 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1272.eqiad.wmnet 127.48.64.10.in-addr.arpa 7.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:05 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1272.eqiad.wmnet 127.48.64.10.in-addr.arpa 7.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1272 - swfrench@cumin1003" * 22:05 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1272 - swfrench@cumin1003" * 22:03 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2025.codfw.wmnet with reason: host reimage * 22:01 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 22:00 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1272 * 22:00 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1272.eqiad.wmnet with OS trixie * 22:00 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1272.eqiad.wmnet * 21:59 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1272.eqiad.wmnet * 21:59 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1272.eqiad.wmnet * 21:55 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1271.eqiad.wmnet * 21:55 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1271.eqiad.wmnet * 21:55 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1271.eqiad.wmnet * 21:46 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1271.eqiad.wmnet with OS trixie * 21:45 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] (duration: 06m 31s) * 21:40 sbassett@deploy2003: sbassett: Continuing with deployment * 21:40 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2025 * 21:40 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2025 * 21:40 sbassett@deploy2003: sbassett: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:38 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] * 21:37 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2025 * 21:37 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2025.codfw.wmnet 220.48.192.10.in-addr.arpa 0.2.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:37 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2025.codfw.wmnet 220.48.192.10.in-addr.arpa 0.2.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:37 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:37 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2025 - bking@cumin2003" * 21:37 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2025 - bking@cumin2003" * 21:30 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] (duration: 08m 19s) * 21:26 sbassett@deploy2003: sbassett: Continuing with deployment * 21:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1269.eqiad.wmnet with OS trixie * 21:24 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1271.eqiad.wmnet with reason: host reimage * 21:23 sbassett@deploy2003: sbassett: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:22 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:22 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] * 21:20 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2025 * 21:19 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2025.codfw.wmnet with OS bookworm * 21:17 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1271.eqiad.wmnet with reason: host reimage * 21:04 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1269.eqiad.wmnet with reason: host reimage * 21:00 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1269.eqiad.wmnet with reason: host reimage * 20:56 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1271 * 20:55 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1271 * 20:54 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1271 * 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1271.eqiad.wmnet 126.48.64.10.in-addr.arpa 6.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:54 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1271.eqiad.wmnet 126.48.64.10.in-addr.arpa 6.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1271 - swfrench@cumin1003" * 20:54 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1271 - swfrench@cumin1003" * 20:51 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:51 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:51 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:50 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 20:49 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 20:49 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1268.eqiad.wmnet * 20:49 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1268.eqiad.wmnet * 20:49 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1268.eqiad.wmnet * 20:48 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1271 * 20:48 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1271.eqiad.wmnet with OS trixie * 20:47 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1271.eqiad.wmnet * 20:46 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1271.eqiad.wmnet * 20:46 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1271.eqiad.wmnet * 20:41 aude@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] (duration: 07m 34s) * 20:39 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1269 * 20:39 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1269 * 20:36 aude@deploy2003: aude: Continuing with deployment * 20:35 aude@deploy2003: aude: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:33 aude@deploy2003: Started scap sync-world: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] * 20:26 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on db2207.codfw.wmnet with reason: Host down * 20:22 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-video: apply * 20:21 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-video: apply * 20:20 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-timeline: apply * 20:20 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-timeline: apply * 20:20 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-syntaxhighlight: apply * 20:19 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-syntaxhighlight: apply * 20:19 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-media: apply * 20:18 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-media: apply * 20:18 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-constraints: apply * 20:17 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-constraints: apply * 20:17 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox: apply * 20:16 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox: apply * 20:13 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1269 * 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1269.eqiad.wmnet 80.32.64.10.in-addr.arpa 0.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:13 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1269.eqiad.wmnet 80.32.64.10.in-addr.arpa 0.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1269 - kamila@cumin1003" * 20:13 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1269 - kamila@cumin1003" * 20:09 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-flink-codfw cluster: Roll restart of jvm daemons. * 20:07 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 20:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 20:03 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-flink-codfw cluster: Roll restart of jvm daemons. * 20:03 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2207 [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94893 and previous config saved to /var/cache/conftool/dbconfig/20260716-200257-marostegui.json * 20:01 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2204 to s2 primary [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94892 and previous config saved to /var/cache/conftool/dbconfig/20260716-200157-marostegui.json * 20:00 marostegui: Starting emergency s2 codfw failover from db2207 to db2204 - [[phab:T432396|T432396]] * 19:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1035.eqiad.wmnet * 19:56 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2204 with weight 0 [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94891 and previous config saved to /var/cache/conftool/dbconfig/20260716-195628-marostegui.json * 19:55 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 26 hosts with reason: Primary switchover s2 [[phab:T432396|T432396]] * 19:54 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1035.eqiad.wmnet * 19:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1034.eqiad.wmnet * 19:48 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1034.eqiad.wmnet * 19:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1033.eqiad.wmnet * 19:43 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-video: apply * 19:43 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1033.eqiad.wmnet * 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1032.eqiad.wmnet * 19:42 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-video: apply * 19:42 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-timeline: apply * 19:41 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-timeline: apply * 19:41 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-syntaxhighlight: apply * 19:41 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-syntaxhighlight: apply * 19:40 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-media: apply * 19:40 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-media: apply * 19:39 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-constraints: apply * 19:36 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-constraints: apply * 19:36 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox: apply * 19:35 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1032.eqiad.wmnet * 19:35 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1031.eqiad.wmnet * 19:35 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox: apply * 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-video: apply * 19:33 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-video: apply * 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-timeline: apply * 19:33 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-timeline: apply * 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-syntaxhighlight: apply * 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-syntaxhighlight: apply * 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-media: apply * 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-media: apply * 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-constraints: apply * 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-constraints: apply * 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox: apply * 19:31 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox: apply * 19:27 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1031.eqiad.wmnet * 19:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1030.eqiad.wmnet * 19:23 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1001.eqiad.wmnet, repooling source-only afterwards * 19:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1030.eqiad.wmnet * 19:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1029.eqiad.wmnet * 19:17 kamila@cumin1003: START - Cookbook sre.dns.netbox * 19:12 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1029.eqiad.wmnet * 19:06 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1269 * 19:05 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1269.eqiad.wmnet with OS trixie * 19:03 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1269.eqiad.wmnet * 19:03 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1269.eqiad.wmnet * 19:03 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1269.eqiad.wmnet * 18:55 dancy@deploy2003: Finished scap sync-world: testing [[phab:T428971|T428971]] (duration: 02m 41s) * 18:53 dancy@deploy2003: Started scap sync-world: testing [[phab:T428971|T428971]] * 18:31 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1268.eqiad.wmnet with OS trixie * 18:18 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 18:16 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1267.eqiad.wmnet * 18:16 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1267.eqiad.wmnet * 18:16 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1267.eqiad.wmnet * 18:09 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1268.eqiad.wmnet with reason: host reimage * 18:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1001.eqiad.wmnet, repooling source-only afterwards * 18:06 swfrench@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] (duration: 07m 34s) * 18:06 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 23s) * 18:06 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 18:06 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1268.eqiad.wmnet with reason: host reimage * 18:03 bd808@deploy2003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 18:02 bd808@deploy2003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 18:02 swfrench@deploy2003: jiji, swfrench: Continuing with deployment * 18:02 bd808@deploy2003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 18:02 bd808@deploy2003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 18:01 bd808@deploy2003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 18:01 swfrench@deploy2003: jiji, swfrench: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:01 bd808@deploy2003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:59 swfrench@deploy2003: Started scap sync-world: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] * 17:45 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1268 * 17:45 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1268 * 17:44 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1267.eqiad.wmnet with OS trixie * 17:43 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter2006.codfw.wmnet * 17:39 swfrench@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter2006.codfw.wmnet * 17:35 swfrench@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] (duration: 07m 27s) * 17:34 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1268 * 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1268.eqiad.wmnet 78.32.64.10.in-addr.arpa 8.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:34 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1268.eqiad.wmnet 78.32.64.10.in-addr.arpa 8.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1268 - kamila@cumin1003" * 17:34 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1268 - kamila@cumin1003" * 17:31 swfrench@deploy2003: jiji, swfrench: Continuing with deployment * 17:29 swfrench@deploy2003: jiji, swfrench: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:28 kamila@cumin1003: START - Cookbook sre.dns.netbox * 17:28 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1268 * 17:28 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1268.eqiad.wmnet with OS trixie * 17:27 swfrench@deploy2003: Started scap sync-world: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] * 17:23 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1267.eqiad.wmnet with reason: host reimage * 17:18 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1267.eqiad.wmnet with reason: host reimage * 17:18 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1268.eqiad.wmnet * 17:17 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1268.eqiad.wmnet * 17:17 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1268.eqiad.wmnet * 17:12 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter2005.codfw.wmnet * 17:11 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1270.eqiad.wmnet * 17:11 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1270.eqiad.wmnet * 17:11 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1270.eqiad.wmnet * 17:09 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter2005.codfw.wmnet * 17:08 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:08 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update reverse dns for moved arelion cct cr2-eqiad - cmooney@cumin1003" * 17:08 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update reverse dns for moved arelion cct cr2-eqiad - cmooney@cumin1003" * 17:08 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] (duration: 07m 34s) * 17:04 jiji@deploy2003: jiji: Continuing with deployment * 17:03 jiji@deploy2003: jiji: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 17:00 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] * 17:00 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2185.codfw.wmnet with OS trixie * 16:59 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:58 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1270.eqiad.wmnet with OS trixie * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1267 * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1267 * 16:57 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1267 * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1267.eqiad.wmnet 77.32.64.10.in-addr.arpa 7.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:57 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1267.eqiad.wmnet 77.32.64.10.in-addr.arpa 7.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1267 - kamila@cumin1003" * 16:56 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1267 - kamila@cumin1003" * 16:56 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_eqsin * 16:56 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5032.eqsin.wmnet * 16:52 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_esams * 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3073.esams.wmnet * 16:50 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_esams * 16:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3081.esams.wmnet * 16:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1266.eqiad.wmnet * 16:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1266.eqiad.wmnet * 16:45 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1266.eqiad.wmnet * 16:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2185.codfw.wmnet with reason: host reimage * 16:41 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_eqiad * 16:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1114.eqiad.wmnet * 16:41 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_eqiad * 16:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1115.eqiad.wmnet * 16:39 kamila@cumin1003: START - Cookbook sre.dns.netbox * 16:39 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1267 * 16:39 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2185.codfw.wmnet with reason: host reimage * 16:38 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1267.eqiad.wmnet with OS trixie * 16:38 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1267.eqiad.wmnet * 16:38 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1270.eqiad.wmnet with reason: host reimage * 16:37 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1267.eqiad.wmnet * 16:37 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1267.eqiad.wmnet * 16:31 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1270.eqiad.wmnet with reason: host reimage * 16:24 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1264.eqiad.wmnet * 16:24 kamila@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 16:24 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 16:23 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 16:21 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2162: switch maintenance completed codfw rack b6 * 16:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2185.codfw.wmnet with OS trixie * 16:19 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 16:16 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter1007.eqiad.wmnet * 16:15 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_eqsin * 16:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5024.eqsin.wmnet * 16:13 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5031.eqsin.wmnet * 16:13 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3072.esams.wmnet * 16:12 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter1007.eqiad.wmnet * 16:11 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] (duration: 09m 47s) * 16:10 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1266.eqiad.wmnet with OS trixie * 16:10 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1270 * 16:10 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1270 * 16:09 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1270 * 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1270.eqiad.wmnet 125.48.64.10.in-addr.arpa 5.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1270.eqiad.wmnet 125.48.64.10.in-addr.arpa 5.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1270 - swfrench@cumin1003" * 16:09 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1270 - swfrench@cumin1003" * 16:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3080.esams.wmnet * 16:07 jiji@deploy2003: jiji: Continuing with deployment * 16:06 jiji@deploy2003: jiji: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:04 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 16:04 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1265.eqiad.wmnet * 16:03 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1265.eqiad.wmnet * 16:03 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1265.eqiad.wmnet * 16:03 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1270 * 16:03 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1270.eqiad.wmnet with OS trixie * 16:02 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1270.eqiad.wmnet * 16:02 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] * 16:01 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1270.eqiad.wmnet * 16:01 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1270.eqiad.wmnet * 16:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1113.eqiad.wmnet * 16:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1112.eqiad.wmnet * 15:49 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1266.eqiad.wmnet with reason: host reimage * 15:47 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter1006.eqiad.wmnet * 15:45 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1265.eqiad.wmnet with OS trixie * 15:44 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1266.eqiad.wmnet with reason: host reimage * 15:43 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter1006.eqiad.wmnet * 15:42 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] (duration: 09m 46s) * 15:37 jiji@deploy2003: jiji: Continuing with deployment * 15:36 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2162: switch maintenance completed codfw rack b6 * 15:36 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2161: switch maintenance completed codfw rack b6 * 15:34 jiji@deploy2003: jiji: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:32 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5023.eqsin.wmnet * 15:32 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] * 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5030.eqsin.wmnet * 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3071.esams.wmnet * 15:27 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3079.esams.wmnet * 15:25 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1265.eqiad.wmnet with reason: host reimage * 15:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1266 * 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1266 * 15:21 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1110.eqiad.wmnet * 15:20 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1111.eqiad.wmnet * 15:16 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1265.eqiad.wmnet with reason: host reimage * 15:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1001.eqiad.wmnet with OS bookworm * 15:15 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1266 * 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1266.eqiad.wmnet 76.32.64.10.in-addr.arpa 6.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:15 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1266.eqiad.wmnet 76.32.64.10.in-addr.arpa 6.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1266 - kamila@cumin1003" * 15:15 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1266 - kamila@cumin1003" * 15:07 kamila@cumin1003: START - Cookbook sre.dns.netbox * 15:04 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1266 * 15:04 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1264 * 15:04 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1264 * 15:04 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1266.eqiad.wmnet with OS trixie * 15:03 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1264 * 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1264.eqiad.wmnet 74.32.64.10.in-addr.arpa 4.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:03 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1264.eqiad.wmnet 74.32.64.10.in-addr.arpa 4.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1264 - kamila@cumin1003" * 15:03 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1264 - kamila@cumin1003" * 15:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-eqiad * 15:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp1001.eqiad.wmnet * 15:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp1001.eqiad.wmnet * 15:01 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp1001.eqiad.wmnet * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp1001.eqiad.wmnet * 15:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1376-1384].eqiad.wmnet * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1376-1384].eqiad.wmnet * 14:59 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1002.eqiad.wmnet * 14:59 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1266.eqiad.wmnet * 14:58 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1266.eqiad.wmnet * 14:58 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1266.eqiad.wmnet * 14:58 kamila@cumin1003: START - Cookbook sre.dns.netbox * 14:57 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1264 * 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1265 * 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1265 * 14:57 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1310596{{!}}Set $wgMathInternalRestbaseURL explicitly (T349582)]] * 14:57 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1265 * 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1265.eqiad.wmnet 75.32.64.10.in-addr.arpa 5.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:56 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1265.eqiad.wmnet 75.32.64.10.in-addr.arpa 5.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:56 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:56 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1265 - kamila@cumin1003" * 14:56 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1265 - kamila@cumin1003" * 14:53 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1002.eqiad.wmnet * 14:53 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1376-1384].eqiad.wmnet * 14:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1001.eqiad.wmnet with reason: host reimage * 14:51 kamila@cumin1003: START - Cookbook sre.dns.netbox * 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 14:50 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:50 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2161: switch maintenance completed codfw rack b6 * 14:50 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5021.eqsin.wmnet * 14:50 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1265 * 14:49 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:49 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:49 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1265.eqiad.wmnet with OS trixie * 14:49 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5029.eqsin.wmnet * 14:49 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1265.eqiad.wmnet * 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3070.esams.wmnet * 14:48 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1264.eqiad.wmnet * 14:48 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1001.eqiad.wmnet with reason: host reimage * 14:48 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1376-1384].eqiad.wmnet * 14:48 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1265.eqiad.wmnet * 14:47 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1265.eqiad.wmnet * 14:47 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1264.eqiad.wmnet * 14:47 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1264.eqiad.wmnet * 14:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:47 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3078.esams.wmnet * 14:44 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1263.eqiad.wmnet * 14:44 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1263.eqiad.wmnet * 14:44 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1263.eqiad.wmnet * 14:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1108.eqiad.wmnet * 14:40 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1109.eqiad.wmnet * 14:40 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:35 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:34 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:34 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2006.codfw.wmnet * 14:34 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-flink-eqiad cluster: Roll restart of jvm daemons. * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf2002.codfw.wmnet * 14:31 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1002.eqiad.wmnet * 14:29 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2006.codfw.wmnet * 14:27 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:27 kamila@deploy2003: Finished scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] (duration: 02m 57s) * 14:27 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-flink-eqiad cluster: Roll restart of jvm daemons. * 14:26 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf2002.codfw.wmnet * 14:26 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf2001.codfw.wmnet * 14:25 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1002.eqiad.wmnet * 14:25 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1001.eqiad.wmnet * 14:25 kamila@deploy2003: Started scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] * 14:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:21 kamila@deploy2003: sync-world aborted: Test deployment to check rsync is working - [[phab:T432108|T432108]] (duration: 00m 36s) * 14:21 topranks: reboot lsw1-b6-codfw to upgrade JunOS [[phab:T430922|T430922]] * 14:21 kamila@deploy2003: Started scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] * 14:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1001.eqiad.wmnet with OS bookworm * 14:20 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b6-codfw,lsw1-b6-codfw IPv6,lsw1-b6-codfw.mgmt,ssw1-a[1,8]-codfw with reason: lsw1-b6-codfw JunOS upgrade * 14:20 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf2001.codfw.wmnet * 14:19 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1001.eqiad.wmnet * 14:19 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 26 hosts with reason: lsw1-b6-codfw JunOS upgrade * 14:14 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 14:13 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc2022: switch maintenance codfw rack b6 * 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:12 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.parsercache * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool pc2022: switch maintenance codfw rack b6 * 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2251: switch maintenance codfw rack b6 * 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.parsercache * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2251: switch maintenance codfw rack b6 * 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2162: switch maintenance codfw rack b6 * 14:12 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1263.eqiad.wmnet with OS trixie * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2162: switch maintenance codfw rack b6 * 14:11 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2161: switch maintenance codfw rack b6 * 14:11 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2161: switch maintenance codfw rack b6 * 14:08 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5020.eqsin.wmnet * 14:07 btullis@cumin1003: START - Cookbook sre.hadoop.reboot-workers for Hadoop analytics cluster * 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3069.esams.wmnet * 14:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5028.eqsin.wmnet * 14:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1338-1347].eqiad.wmnet * 14:06 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1338-1347].eqiad.wmnet * 14:05 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3077.esams.wmnet * 14:02 topranks: beginning depools for lsw1-b6-codfw maintenance [[phab:T430922|T430922]] * 14:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1106.eqiad.wmnet * 14:00 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-misc1002.eqiad.wmnet * 13:59 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1338-1347].eqiad.wmnet * 13:59 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1107.eqiad.wmnet * 13:56 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-codfw * 13:55 sfaci@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply * 13:54 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-misc1002.eqiad.wmnet * 13:54 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-misc1001.eqiad.wmnet * 13:54 sfaci@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply * 13:50 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1263.eqiad.wmnet with reason: host reimage * 13:50 sfaci@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 13:49 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1338-1347].eqiad.wmnet * 13:49 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-misc1001.eqiad.wmnet * 13:49 sfaci@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 13:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:49 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:45 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1263.eqiad.wmnet with reason: host reimage * 13:40 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:40 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-eqiad * 13:35 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:34 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:33 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-reboot (exit_code=0) rolling reboot on A:dnsbox and (A:eqsin or A:drmrs or A:magru) and not (P<nowiki>{</nowiki>dns5003*<nowiki>}</nowiki> or P<nowiki>{</nowiki>dns7002*<nowiki>}</nowiki>) and (A:dnsbox) * 13:33 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns7001.wikimedia.org * 13:27 sukhe@dns1004: END - running authdns-update * 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5019.eqsin.wmnet * 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3076.esams.wmnet * 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3068.esams.wmnet * 13:25 sukhe@dns1004: START - running authdns-update * 13:24 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5027.eqsin.wmnet * 13:24 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1263 * 13:24 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1263 * 13:23 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1263 * 13:23 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:23 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1104.eqiad.wmnet * 13:21 kamila@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:21 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:20 kamila@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:20 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:20 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:20 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1263 - kamila@cumin1003" * 13:20 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1263 - kamila@cumin1003" * 13:19 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1105.eqiad.wmnet * 13:19 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:19 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:18 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:18 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns7001.wikimedia.org * 13:16 cdobbins@cumin2003: conftool action : set/pooled=yes; selector: name=dns7002.* * 13:14 cdobbins@dns1004: END - running authdns-update * 13:13 sbisson@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] (duration: 08m 03s) * 13:13 cdobbins@dns1004: START - running authdns-update * 13:12 kamila@cumin1003: START - Cookbook sre.dns.netbox * 13:12 cdobbins@cumin2003: conftool action : set/pooled=yes; selector: name=dns7002.*,service=authdns-update * 13:12 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1263 * 13:11 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1263.eqiad.wmnet with OS trixie * 13:11 cdobbins@cumin2003: conftool action : set/pooled=no; selector: name=dns7002.* * 13:11 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1263.eqiad.wmnet * 13:10 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1263.eqiad.wmnet * 13:10 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1263.eqiad.wmnet * 13:09 sbisson@deploy2003: sbisson: Continuing with deployment * 13:08 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:07 sbisson@deploy2003: sbisson: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:05 sbisson@deploy2003: Started scap sync-world: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] * 13:03 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns6002.wikimedia.org * 13:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1298-1307].eqiad.wmnet * 13:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1298-1307].eqiad.wmnet * 12:59 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1262.eqiad.wmnet * 12:59 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1262.eqiad.wmnet * 12:59 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1262.eqiad.wmnet * 12:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1298-1307].eqiad.wmnet * 12:49 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns6002.wikimedia.org * 12:46 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1298-1307].eqiad.wmnet * 12:46 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:46 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3075.esams.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3067.esams.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5018.eqsin.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5026.eqsin.wmnet * 12:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1102.eqiad.wmnet * 12:39 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1103.eqiad.wmnet * 12:35 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:34 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns6001.wikimedia.org * 12:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:28 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:18 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns6001.wikimedia.org * 12:14 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1267-1276].eqiad.wmnet * 12:13 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1267-1276].eqiad.wmnet * 12:04 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1267-1276].eqiad.wmnet * 12:03 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns5004.wikimedia.org * 12:02 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1100.eqiad.wmnet * 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3066.esams.wmnet * 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3074.esams.wmnet * 12:01 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:01 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-codfw * 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5017.eqsin.wmnet * 12:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5025.eqsin.wmnet * 12:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1101.eqiad.wmnet * 11:59 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1267-1276].eqiad.wmnet * 11:58 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:58 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:54 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns5004.wikimedia.org * 11:54 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and (A:eqsin or A:drmrs or A:magru) and not (P<nowiki>{</nowiki>dns5003*<nowiki>}</nowiki> or P<nowiki>{</nowiki>dns7002*<nowiki>}</nowiki>) and (A:dnsbox) * 11:54 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:53 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:53 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-eqiad * 11:51 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:51 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:50 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:50 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_eqiad * 11:50 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:50 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_eqiad * 11:49 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_esams * 11:49 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_esams * 11:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_eqsin * 11:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_eqsin * 11:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:44 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-eqiad * 11:43 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-codfw * 11:42 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:41 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:24 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-eqiad * 11:23 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-codfw * 11:22 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:15 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:14 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:09 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2066.codfw.wmnet with OS trixie * 11:05 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1151-1160].eqiad.wmnet * 11:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1151-1160].eqiad.wmnet * 10:59 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1068.eqiad.wmnet with OS trixie * 10:55 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.major-upgrade (exit_code=99) * 10:55 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 10:54 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1151-1160].eqiad.wmnet * 10:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2066.codfw.wmnet with reason: host reimage * 10:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1151-1160].eqiad.wmnet * 10:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:42 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2066.codfw.wmnet with reason: host reimage * 10:39 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:37 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:36 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:23 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:22 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2066.codfw.wmnet with OS trixie * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:07 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2065.codfw.wmnet with OS trixie * 10:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:06 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 10:06 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 10:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 10:03 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:03 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 10:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 09:59 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 09:57 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 09:57 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 09:52 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 09:47 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:46 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:46 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2065.codfw.wmnet with reason: host reimage * 09:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2065.codfw.wmnet with reason: host reimage * 09:40 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:39 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 09:39 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:39 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:39 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:37 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:29 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox2003.codfw.wmnet * 09:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox2003.codfw.wmnet * 09:25 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:25 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:24 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:24 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:24 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:21 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:20 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2065.codfw.wmnet with OS trixie * 09:13 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1068.eqiad.wmnet with OS trixie * 09:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2064.codfw.wmnet with OS trixie * 09:08 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 09:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 09:07 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2162: Repooling after switchover * 09:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1067.eqiad.wmnet with OS trixie * 09:00 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:59 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 08:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 08:57 tappof: bump space for prometheus k8s-dse in eqiad * 08:56 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping2004.codfw.wmnet * 08:52 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host ping2004.codfw.wmnet * 08:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 08:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping1004.eqiad.wmnet * 08:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:51 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:49 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 08:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host ping1004.eqiad.wmnet * 08:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2064.codfw.wmnet with reason: host reimage * 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2064.codfw.wmnet with reason: host reimage * 08:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1067.eqiad.wmnet with reason: host reimage * 08:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:33 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1067.eqiad.wmnet with reason: host reimage * 08:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:21 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2162: Repooling after switchover * 08:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2064.codfw.wmnet with OS trixie * 08:16 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1067.eqiad.wmnet with OS trixie * 08:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:15 cgoubert@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-eqiad * 08:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2062.codfw.wmnet with OS trixie * 08:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1066.eqiad.wmnet with OS trixie * 08:02 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2162: Repooling after switchover * 07:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2162: Repooling after switchover * 07:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2162 [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94870 and previous config saved to /var/cache/conftool/dbconfig/20260716-075530-cwilliams.json * 07:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2241 to x3 primary [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94869 and previous config saved to /var/cache/conftool/dbconfig/20260716-075314-cwilliams.json * 07:52 cezmunsta: Starting x3 codfw failover from db2162 to db2241 - [[phab:T430925|T430925]] * 07:50 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:50 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:47 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 07:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2241 with weight 0 [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94868 and previous config saved to /var/cache/conftool/dbconfig/20260716-074507-cwilliams.json * 07:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 18 hosts with reason: Primary switchover x3 [[phab:T430925|T430925]] * 07:43 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1066.eqiad.wmnet with reason: host reimage * 07:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:dse-k8s-worker-eqiad * 07:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1028.eqiad.wmnet * 07:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1028.eqiad.wmnet * 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 07:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1066.eqiad.wmnet with reason: host reimage * 07:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1028.eqiad.wmnet * 07:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1028.eqiad.wmnet * 07:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1027.eqiad.wmnet * 07:35 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1027.eqiad.wmnet * 07:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1027.eqiad.wmnet * 07:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1027.eqiad.wmnet * 07:28 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1026.eqiad.wmnet * 07:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1026.eqiad.wmnet * 07:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast2003.wikimedia.org * 07:21 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1026.eqiad.wmnet * 07:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1066.eqiad.wmnet with OS trixie * 07:19 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast2003.wikimedia.org * 07:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2062.codfw.wmnet with OS trixie * 06:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1026.eqiad.wmnet * 06:51 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1025.eqiad.wmnet * 06:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1025.eqiad.wmnet * 06:47 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 06:47 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 06:44 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1025.eqiad.wmnet * 06:14 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1025.eqiad.wmnet * 06:14 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1024.eqiad.wmnet * 06:14 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1024.eqiad.wmnet * 06:07 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1024.eqiad.wmnet * 05:37 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1024.eqiad.wmnet * 05:37 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1023.eqiad.wmnet * 05:37 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1023.eqiad.wmnet * 05:26 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1023.eqiad.wmnet * 04:56 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1023.eqiad.wmnet * 04:56 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1022.eqiad.wmnet * 04:56 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1022.eqiad.wmnet * 04:49 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1022.eqiad.wmnet * 04:19 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1022.eqiad.wmnet * 04:19 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1021.eqiad.wmnet * 04:19 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1021.eqiad.wmnet * 04:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1021.eqiad.wmnet * 03:38 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1021.eqiad.wmnet * 03:38 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1020.eqiad.wmnet * 03:38 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1020.eqiad.wmnet * 03:20 btullis@cumin1003: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1020.eqiad.wmnet * 03:18 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1020.eqiad.wmnet * 03:18 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1019.eqiad.wmnet * 03:18 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1019.eqiad.wmnet * 03:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1019.eqiad.wmnet * 02:41 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1019.eqiad.wmnet * 02:41 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1018.eqiad.wmnet * 02:41 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1018.eqiad.wmnet * 02:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling both afterwards * 02:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2003.codfw.wmnet -> wcqs2001.codfw.wmnet, repooling both afterwards * 02:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1018.eqiad.wmnet * 02:30 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1018.eqiad.wmnet * 02:30 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1014.eqiad.wmnet * 02:30 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1014.eqiad.wmnet * 02:24 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1014.eqiad.wmnet * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 01:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1014.eqiad.wmnet * 01:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1013.eqiad.wmnet * 01:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1013.eqiad.wmnet * 01:47 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1013.eqiad.wmnet * 01:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2003.codfw.wmnet -> wcqs2001.codfw.wmnet, repooling both afterwards * 01:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling both afterwards * 01:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1013.eqiad.wmnet * 01:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1012.eqiad.wmnet * 01:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1012.eqiad.wmnet * 01:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1012.eqiad.wmnet * 01:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1012.eqiad.wmnet * 01:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1011.eqiad.wmnet * 01:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1011.eqiad.wmnet * 01:08 ryankemper@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] scap deploy post bookworm reimage (duration: 00m 23s) * 01:08 ryankemper@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] scap deploy post bookworm reimage * 01:08 ryankemper@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): scap deploy post bookworm reimage (duration: 00m 46s) * 01:07 ryankemper@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): scap deploy post bookworm reimage * 01:04 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1011.eqiad.wmnet * 01:04 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1011.eqiad.wmnet * 01:04 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1010.eqiad.wmnet * 01:04 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1010.eqiad.wmnet * 00:57 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1010.eqiad.wmnet * 00:57 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1010.eqiad.wmnet * 00:57 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1009.eqiad.wmnet * 00:57 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1009.eqiad.wmnet * 00:50 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1009.eqiad.wmnet * 00:20 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1009.eqiad.wmnet * 00:20 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1008.eqiad.wmnet * 00:20 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1008.eqiad.wmnet * 00:13 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1008.eqiad.wmnet == 2026-07-15 == * 23:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2001.codfw.wmnet with OS bookworm * 23:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1008.eqiad.wmnet * 23:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1007.eqiad.wmnet * 23:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1007.eqiad.wmnet * 23:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1007.eqiad.wmnet * 23:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1007.eqiad.wmnet * 23:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1006.eqiad.wmnet * 23:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1006.eqiad.wmnet * 23:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1006.eqiad.wmnet * 23:29 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1006.eqiad.wmnet * 23:28 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1005.eqiad.wmnet * 23:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1005.eqiad.wmnet * 23:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1002.eqiad.wmnet with OS bookworm * 23:21 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1005.eqiad.wmnet * 23:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2001.codfw.wmnet with reason: host reimage * 23:15 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host datahubsearch1001.eqiad.wmnet with OS bookworm * 23:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2001.codfw.wmnet with reason: host reimage * 23:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 23:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 22:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 22:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1005.eqiad.wmnet * 22:51 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1004.eqiad.wmnet * 22:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1004.eqiad.wmnet * 22:45 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1004.eqiad.wmnet * 22:44 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 22:44 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS trixie * 22:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host datahubsearch1001.eqiad.wmnet with OS bookworm * 22:34 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host datahubsearch1001.eqiad.wmnet with OS bookworm * 22:16 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on datahubsearch[1002-1003].eqiad.wmnet with reason: Using datahubsearch1001 to test bookworm reimages * 22:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1004.eqiad.wmnet * 22:15 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1003.eqiad.wmnet * 22:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1003.eqiad.wmnet * 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 22:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1003.eqiad.wmnet * 22:08 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1003.eqiad.wmnet * 22:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1002.eqiad.wmnet * 22:08 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1002.eqiad.wmnet * 22:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host datahubsearch1001.eqiad.wmnet with OS bookworm * 22:05 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 22:02 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm * 22:01 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on datahubsearch[1001-1003].eqiad.wmnet with reason: Using datahubsearch1001 to test bookworm reimages * 22:01 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1002.eqiad.wmnet * 22:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1002.eqiad.wmnet * 22:00 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1001.eqiad.wmnet * 22:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1001.eqiad.wmnet * 21:53 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1001.eqiad.wmnet * 21:52 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 21:50 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 21:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS trixie * 21:50 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS bookworm * 21:43 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 21:38 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:30 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wcqs1002'] * 21:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:29 lerickson@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 21:29 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:29 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:29 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS bookworm * 21:28 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 21:28 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm * 21:23 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1001.eqiad.wmnet * 21:23 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:23 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:22 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 21:20 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:18 swfrench-wmf: reprepro include php8.3_8.3.32-1+wmf11u2 into component/php83 for bullseye-wikimedia * 21:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:16 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:15 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-druid-public cluster: Roll restart of jvm daemons. * 21:08 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:05 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1001.eqiad.wmnet * 21:05 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1001.eqiad.wmnet * 21:04 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-druid-public cluster: Roll restart of jvm daemons. * 21:02 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 21:01 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 21:01 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 21:00 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 20:59 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1001.eqiad.wmnet * 20:59 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1001.eqiad.wmnet * 20:59 btullis@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:dse-k8s-worker-eqiad * 20:55 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 20:55 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 20:45 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm * 20:21 jhathaway: puppet is re-enabled, have fun, but not too much fun! * 20:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 20:17 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs2001'] * 20:12 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs2001'] * 20:11 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs2001'] * 20:09 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:08 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 20:05 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:05 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 20:04 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs2001'] * 20:03 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 20:03 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm * 20:02 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:02 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 20:01 jhathaway: disabling puppet fleet wide to roll out kafka patch * 19:55 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host relforge1010.eqiad.wmnet * 19:52 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 19:52 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 19:48 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 19:48 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 19:48 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 19:47 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 19:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 19:45 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 19:45 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 19:44 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1010.eqiad.wmnet * 19:38 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1262.eqiad.wmnet with OS trixie * 19:17 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1262.eqiad.wmnet with reason: host reimage * 19:11 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1262.eqiad.wmnet with reason: host reimage * 18:59 cdobbins@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS trixie * 18:54 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 18:53 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 18:52 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1262 * 18:52 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1262 * 18:51 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1262 * 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1262.eqiad.wmnet 72.32.64.10.in-addr.arpa 2.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:51 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1262.eqiad.wmnet 72.32.64.10.in-addr.arpa 2.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1262 - kamila@cumin1003" * 18:51 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1262 - kamila@cumin1003" * 18:46 kamila@cumin1003: START - Cookbook sre.dns.netbox * 18:46 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1262 * 18:46 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 18:46 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ncmonitor1001.eqiad.wmnet * 18:46 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 18:45 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1262.eqiad.wmnet with OS trixie * 18:45 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 18:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1262.eqiad.wmnet * 18:44 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1262.eqiad.wmnet * 18:44 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1262.eqiad.wmnet * 18:42 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host ncmonitor1001.eqiad.wmnet * 18:29 topranks: pull power on cr1-eqiad to install new switch-control boards [[phab:T426343|T426343]] * 18:29 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs[1018-1020].eqiad.wmnet with reason: line card install in cr1-eqiad * 18:27 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 14 hosts with reason: linecard install in cr1-eqad * 18:22 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_ulsfo * 18:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4052.ulsfo.wmnet * 18:19 cdobbins@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 18:15 cdobbins@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 18:14 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_drmrs * 18:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6016.drmrs.wmnet * 18:12 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_ulsfo * 18:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4044.ulsfo.wmnet * 18:10 sukhe@cumin1003: END (ERROR) - Cookbook sre.cdn.roll-reboot (exit_code=97) rolling reboot on A:cp-upload_drmrs * 18:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2241: Security update * 17:56 topranks: start draining traffic on cr1-eqiad ahead of line card installation [[phab:T426343|T426343]] * 17:47 cdobbins@cumin2003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie * 17:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4051.ulsfo.wmnet * 17:40 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:39 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 17:34 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6007.drmrs.wmnet * 17:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6015.drmrs.wmnet * 17:32 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:31 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 17:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4043.ulsfo.wmnet * 17:27 lerickson@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:25 lerickson@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 17:22 lerickson@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-codfw * 17:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp2001.codfw.wmnet * 17:22 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 17:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp2001.codfw.wmnet * 17:22 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 17:19 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2241: Security update * 17:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2241.codfw.wmnet * 17:17 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2241.codfw.wmnet * 17:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp2001.codfw.wmnet * 17:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp2001.codfw.wmnet * 17:15 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2366-2374].codfw.wmnet * 17:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2366-2374].codfw.wmnet * 17:10 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply * 17:10 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply * 17:08 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2366-2374].codfw.wmnet * 17:06 sukhe: sre.dns.roll-reboot to resume later * 17:06 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-reboot (exit_code=97) rolling reboot on A:dnsbox and not (A:ulsfo or A:magru) and (A:dnsbox) * 17:06 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns5003.wikimedia.org * 17:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2241: Security update * 17:03 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2241: Security update * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply * 17:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2366-2374].codfw.wmnet * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply * 17:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2357-2365].codfw.wmnet * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 17:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2357-2365].codfw.wmnet * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply * 16:55 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2357-2365].codfw.wmnet * 16:55 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 16:53 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6006.drmrs.wmnet * 16:52 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6014.drmrs.wmnet * 16:52 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 16:51 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 16:50 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2357-2365].codfw.wmnet * 16:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4042.ulsfo.wmnet * 16:50 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2347-2356].codfw.wmnet * 16:50 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2347-2356].codfw.wmnet * 16:49 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns5003.wikimedia.org * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply * 16:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4050.ulsfo.wmnet * 16:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2347-2356].codfw.wmnet * 16:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2347-2356].codfw.wmnet * 16:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2337-2346].codfw.wmnet * 16:36 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2337-2346].codfw.wmnet * 16:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:dse-k8s-worker-codfw * 16:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2003.codfw.wmnet * 16:35 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2003.codfw.wmnet * 16:34 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns3004.wikimedia.org * 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply * 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply * 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply * 16:30 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply * 16:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2003.codfw.wmnet * 16:29 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2337-2346].codfw.wmnet * 16:24 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2003.codfw.wmnet * 16:24 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2002.codfw.wmnet * 16:24 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2002.codfw.wmnet * 16:23 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns3004.wikimedia.org * 16:23 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2337-2346].codfw.wmnet * 16:23 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2327-2336].codfw.wmnet * 16:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2327-2336].codfw.wmnet * 16:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2002.codfw.wmnet * 16:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2327-2336].codfw.wmnet * 16:12 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2002.codfw.wmnet * 16:12 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2001.codfw.wmnet * 16:12 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2001.codfw.wmnet * 16:12 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1065.eqiad.wmnet with OS trixie * 16:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6005.drmrs.wmnet * 16:11 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6013.drmrs.wmnet * 16:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4041.ulsfo.wmnet * 16:08 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns3003.wikimedia.org * 16:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2327-2336].codfw.wmnet * 16:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2317-2326].codfw.wmnet * 16:06 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2317-2326].codfw.wmnet * 16:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2001.codfw.wmnet * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply * 16:03 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4049.ulsfo.wmnet * 16:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2001.codfw.wmnet * 16:00 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test2001.codfw.wmnet * 16:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test2001.codfw.wmnet * 16:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2063.codfw.wmnet with OS trixie * 15:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2317-2326].codfw.wmnet * 15:57 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns3003.wikimedia.org * 15:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test2001.codfw.wmnet * 15:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test2001.codfw.wmnet * 15:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2004.codfw.wmnet * 15:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2004.codfw.wmnet * 15:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2317-2326].codfw.wmnet * 15:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2307-2316].codfw.wmnet * 15:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2307-2316].codfw.wmnet * 15:49 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2004.codfw.wmnet * 15:48 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2004.codfw.wmnet * 15:48 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2003.codfw.wmnet * 15:48 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2003.codfw.wmnet * 15:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 15:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2307-2316].codfw.wmnet * 15:42 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2003.codfw.wmnet * 15:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 15:42 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2003.codfw.wmnet * 15:42 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2002.codfw.wmnet * 15:42 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2002.codfw.wmnet * 15:42 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2006.wikimedia.org * 15:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2063.codfw.wmnet with reason: host reimage * 15:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2307-2316].codfw.wmnet * 15:37 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2297-2306].codfw.wmnet * 15:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2297-2306].codfw.wmnet * 15:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2002.codfw.wmnet * 15:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2002.codfw.wmnet * 15:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2001.codfw.wmnet * 15:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2001.codfw.wmnet * 15:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2063.codfw.wmnet with reason: host reimage * 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6004.drmrs.wmnet * 15:31 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2001.codfw.wmnet * 15:31 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2001.codfw.wmnet * 15:31 btullis@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:dse-k8s-worker-codfw * 15:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6012.drmrs.wmnet * 15:28 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2006.wikimedia.org * 15:27 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-analytics cluster: Roll restart of jvm daemons. * 15:27 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2297-2306].codfw.wmnet * 15:27 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4040.ulsfo.wmnet * 15:24 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1065.eqiad.wmnet with OS trixie * 15:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4048.ulsfo.wmnet * 15:21 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-analytics cluster: Roll restart of jvm daemons. * 15:21 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2297-2306].codfw.wmnet * 15:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2287-2296].codfw.wmnet * 15:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2287-2296].codfw.wmnet * 15:20 btullis@cumin1003: END (PASS) - Cookbook sre.druid.reboot-workers (exit_code=0) for Druid public cluster: Reboot Druid nodes * 15:18 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 15:17 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm * 15:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2063.codfw.wmnet with OS trixie * 15:13 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2005.wikimedia.org * 15:11 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2287-2296].codfw.wmnet * 15:11 btullis@cumin1003: END (PASS) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=0) rolling reboot on A:cephosd-eqiad * 15:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1064.eqiad.wmnet with OS trixie * 15:05 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2062.codfw.wmnet with OS trixie * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2287-2296].codfw.wmnet * 15:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2277-2286].codfw.wmnet * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2277-2286].codfw.wmnet * 14:59 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2005.wikimedia.org * 14:57 brouberol@cumin1003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-jumbo-eqiad * 14:52 btullis@cumin1003: END (PASS) - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas (exit_code=0) rolling reboot on A:schema-codfw * 14:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6003.drmrs.wmnet * 14:50 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:50 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host relforge1009.eqiad.wmnet * 14:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2277-2286].codfw.wmnet * 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6011.drmrs.wmnet * 14:47 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 14:46 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>ml-serve1001.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 14:46 klausman@cumin1003: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) pool for host ml-serve1001.eqiad.wmnet * 14:46 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 14:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1001.eqiad.wmnet * 14:45 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4039.ulsfo.wmnet * 14:44 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1009.eqiad.wmnet * 14:44 btullis@cumin1003: START - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas rolling reboot on A:schema-codfw * 14:44 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2004.wikimedia.org * 14:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2277-2286].codfw.wmnet * 14:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2267-2276].codfw.wmnet * 14:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2267-2276].codfw.wmnet * 14:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 14:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4047.ulsfo.wmnet * 14:40 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1001.eqiad.wmnet * 14:38 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 14:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 14:36 topranks: disconnect power on cr2-eqiad to shut down device for switch fabric replacement [[phab:T426343|T426343]] * 14:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2267-2276].codfw.wmnet * 14:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 14:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1001.eqiad.wmnet * 14:35 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>ml-serve1001.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 14:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 14:34 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 14:33 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:33 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:30 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2004.wikimedia.org * 14:29 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2267-2276].codfw.wmnet * 14:29 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2257-2266].codfw.wmnet * 14:29 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2257-2266].codfw.wmnet * 14:24 btullis@cumin1003: END (PASS) - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas (exit_code=0) rolling reboot on A:schema-eqiad * 14:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2257-2266].codfw.wmnet * 14:20 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:20 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:19 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:17 jforrester@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2257-2266].codfw.wmnet * 14:16 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:16 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2062.codfw.wmnet with OS trixie * 14:15 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1064.eqiad.wmnet with OS trixie * 14:15 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1006.wikimedia.org * 14:15 btullis@cumin1003: START - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas rolling reboot on A:schema-eqiad * 14:14 topranks: switch routing-engine on cr2-eqiad resetting all interfaces [[phab:T417873|T417873]] * 14:11 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:11 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:10 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm * 14:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6002.drmrs.wmnet * 14:09 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6010.drmrs.wmnet * 14:06 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1006.wikimedia.org * 14:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:05 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 14:05 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4038.ulsfo.wmnet * 14:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4046.ulsfo.wmnet * 14:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:00 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on cr1-eqiad with reason: switch upgrade and line card install * 14:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:59 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:57 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:57 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:55 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-eqiad * 13:55 btullis@cumin1003: START - Cookbook sre.druid.reboot-workers for Druid public cluster: Reboot Druid nodes * 13:53 topranks: switch routing-engine on cr2-eqiad resetting all interfaces [[phab:T417873|T417873]] * 13:51 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1005.wikimedia.org * 13:50 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:49 brouberol@cumin1003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-test-eqiad * 13:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:44 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2197-2206].codfw.wmnet * 13:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2197-2206].codfw.wmnet * 13:36 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1005.wikimedia.org * 13:35 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2197-2206].codfw.wmnet * 13:30 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2197-2206].codfw.wmnet * 13:28 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6001.drmrs.wmnet * 13:28 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6009.drmrs.wmnet * 13:28 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2187-2196].codfw.wmnet * 13:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2187-2196].codfw.wmnet * 13:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 13:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4037.ulsfo.wmnet * 13:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2001 * 13:22 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2001 * 13:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4045.ulsfo.wmnet * 13:21 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1004.wikimedia.org * 13:19 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on lvs[1018-1020].eqiad.wmnet with reason: switch upgrade and line card install * 13:18 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2009.codfw.wmnet * 13:18 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2001 * 13:18 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2001.codfw.wmnet 26.16.192.10.in-addr.arpa 6.2.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:17 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2001.codfw.wmnet 26.16.192.10.in-addr.arpa 6.2.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:17 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:17 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2001 - bking@cumin2003" * 13:17 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2001 - bking@cumin2003" * 13:17 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2009.codfw.wmnet * 13:17 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_drmrs * 13:17 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2187-2196].codfw.wmnet * 13:17 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_drmrs * 13:17 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on 15 hosts with reason: switch upgrade and line card install * 13:17 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:15 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 13:13 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:13 brouberol@cumin1003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-jumbo-eqiad * 13:13 brouberol@cumin1003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-test-eqiad * 13:13 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1004.wikimedia.org * 13:13 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and not (A:ulsfo or A:magru) and (A:dnsbox) * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:12 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_ulsfo * 13:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:12 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_ulsfo * 13:11 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2187-2196].codfw.wmnet * 13:11 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 13:11 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 13:06 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 13:05 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling source-only afterwards * 13:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2001 * 13:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2009.codfw.wmnet with OS trixie * 13:03 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:03 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling source-only afterwards * 13:01 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 15s) * 13:01 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 13:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 12:57 btullis@cumin1003: END (PASS) - Cookbook sre.druid.reboot-workers (exit_code=0) for Druid analytics cluster: Reboot Druid nodes * 12:54 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 12:54 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2163-2172].codfw.wmnet * 12:54 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2163-2172].codfw.wmnet * 12:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2163-2172].codfw.wmnet * 12:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2009.codfw.wmnet with reason: host reimage * 12:41 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2163-2172].codfw.wmnet * 12:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2153-2162].codfw.wmnet * 12:40 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2153-2162].codfw.wmnet * 12:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2009.codfw.wmnet with reason: host reimage * 12:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2153-2162].codfw.wmnet * 12:29 btullis@cumin1003: END (PASS) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=0) rolling reboot on A:cephosd-codfw * 12:25 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2153-2162].codfw.wmnet * 12:25 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2143-2152].codfw.wmnet * 12:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2143-2152].codfw.wmnet * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2009 * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2009 * 12:22 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2009 * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2009.codfw.wmnet 139.0.192.10.in-addr.arpa 9.3.1.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:22 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2009.codfw.wmnet 139.0.192.10.in-addr.arpa 9.3.1.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2009 - mvernon@cumin2003" * 12:22 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2009 - mvernon@cumin2003" * 12:16 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 12:15 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 12:15 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 12:15 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2009 * 12:15 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 12:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2009.codfw.wmnet with OS trixie * 12:15 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 12:14 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 12:14 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2143-2152].codfw.wmnet * 12:13 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 12:12 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2010.codfw.wmnet * 12:11 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2010.codfw.wmnet * 12:10 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 12:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2143-2152].codfw.wmnet * 12:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2133-2142].codfw.wmnet * 12:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2133-2142].codfw.wmnet * 12:02 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 11:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2133-2142].codfw.wmnet * 11:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2133-2142].codfw.wmnet * 11:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:49 mvolz@deploy2003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:49 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-codfw * 11:48 mvolz@deploy2003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:47 btullis@cumin1003: START - Cookbook sre.druid.reboot-workers for Druid analytics cluster: Reboot Druid nodes * 11:46 mvolz@deploy2003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:46 mvolz@deploy2003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:45 mvolz@deploy2003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:44 mvolz@deploy2003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:40 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] (duration: 11m 38s) * 11:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2010.codfw.wmnet with OS trixie * 11:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2105-2114].codfw.wmnet * 11:36 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2105-2114].codfw.wmnet * 11:36 krinkle@deploy2003: physikerwelt, krinkle: Continuing with deployment * 11:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1018: Security updates * 11:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:36 root@cumin1003: START - Cookbook sre.mysql.parsercache * 11:36 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1018: Security updates * 11:31 krinkle@deploy2003: physikerwelt, krinkle: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:29 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] * 11:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2105-2114].codfw.wmnet * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2105-2114].codfw.wmnet * 11:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2010.codfw.wmnet with reason: host reimage * 11:12 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2010.codfw.wmnet with reason: host reimage * 11:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1018: Security updates * 11:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:10 root@cumin1003: START - Cookbook sre.mysql.parsercache * 11:10 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1018: Security updates * 11:09 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1009.eqiad.wmnet with OS trixie * 11:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow7002.magru.wmnet * 11:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 11:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 11:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=tegola-vector-tiles,name=eqiad * 11:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=kartotherian,name=eqiad * 11:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow7002.magru.wmnet * 10:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2010 * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2010 * 10:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 10:54 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 10:54 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2010 * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2010.codfw.wmnet 76.16.192.10.in-addr.arpa 6.7.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:54 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2010.codfw.wmnet 76.16.192.10.in-addr.arpa 6.7.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2010 - mvernon@cumin2003" * 10:54 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2010 - mvernon@cumin2003" * 10:49 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 10:49 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2010 * 10:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1009.eqiad.wmnet with reason: host reimage * 10:49 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2010.codfw.wmnet with OS trixie * 10:46 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2011.codfw.wmnet * 10:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1011.eqiad.wmnet * 10:44 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2011.codfw.wmnet * 10:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 10:44 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 10:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1009.eqiad.wmnet with reason: host reimage * 10:44 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow6001.drmrs.wmnet * 10:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1017: Security updates * 10:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:39 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1017: Security updates * 10:39 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow6001.drmrs.wmnet * 10:38 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1011.eqiad.wmnet * 10:35 cgoubert@deploy2003: Finished deploy [restbase/deploy@06301bd]: Deploying {{Gerrit|1306088}} {{Gerrit|1308347}} - [[phab:T429944|T429944]] [[phab:T428279|T428279]] (duration: 28m 34s) * 10:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1012.eqiad.wmnet * 10:35 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 10:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow5003.eqsin.wmnet * 10:34 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2011.codfw.wmnet with OS trixie * 10:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1009.eqiad.wmnet with OS trixie * 10:28 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1012.eqiad.wmnet * 10:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1013.eqiad.wmnet * 10:27 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:27 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow5003.eqsin.wmnet * 10:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:26 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow4003.ulsfo.wmnet * 10:25 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 10:25 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 10:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow4003.ulsfo.wmnet * 10:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1013.eqiad.wmnet * 10:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1014.eqiad.wmnet * 10:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:15 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2011.codfw.wmnet with reason: host reimage * 10:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1017: Security updates * 10:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:14 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:14 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1017: Security updates * 10:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow3004.esams.wmnet * 10:11 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2011.codfw.wmnet with reason: host reimage * 10:10 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:10 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1014.eqiad.wmnet * 10:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki2003.codfw.wmnet * 10:09 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 10:09 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 10:09 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow3004.esams.wmnet * 10:08 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2004.codfw.wmnet * 10:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1010.eqiad.wmnet with OS trixie * 10:07 cgoubert@deploy2003: Started deploy [restbase/deploy@06301bd]: Deploying {{Gerrit|1306088}} {{Gerrit|1308347}} - [[phab:T429944|T429944]] [[phab:T428279|T428279]] * 10:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host rpki2003.codfw.wmnet * 10:04 topranks: push out config change to BGP_outfilter on core routers [[phab:T431849|T431849]] * 10:02 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow2004.codfw.wmnet * 09:59 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 09:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2003.codfw.wmnet * 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2011 * 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2011 * 09:53 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 09:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:52 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2011 * 09:52 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2011.codfw.wmnet 36.32.192.10.in-addr.arpa 6.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:52 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2011.codfw.wmnet 36.32.192.10.in-addr.arpa 6.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:51 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:51 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2011 - mvernon@cumin2003" * 09:51 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2011 - mvernon@cumin2003" * 09:51 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow2003.codfw.wmnet * 09:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1003.eqiad.wmnet * 09:49 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 09:49 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 09:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1010.eqiad.wmnet with reason: host reimage * 09:47 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 09:47 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 09:47 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 09:47 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2011 * 09:46 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2011.codfw.wmnet with OS trixie * 09:44 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow1003.eqiad.wmnet * 09:44 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2012.codfw.wmnet * 09:44 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1002.eqiad.wmnet * 09:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1010.eqiad.wmnet with reason: host reimage * 09:43 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2012.codfw.wmnet * 09:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Security updates * 09:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:43 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:43 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Security updates * 09:42 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:40 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow1002.eqiad.wmnet * 09:40 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 09:37 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki1001.eqiad.wmnet * 09:36 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 09:36 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 09:33 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host rpki1001.eqiad.wmnet * 09:32 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:32 cgoubert@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-codfw * 09:31 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=kartotherian,name=eqiad * 09:31 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola-vector-tiles,name=eqiad * 09:31 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 09:31 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2012.codfw.wmnet with OS trixie * 09:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1010.eqiad.wmnet with OS trixie * 09:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1011.eqiad.wmnet with OS trixie * 09:21 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Security updates * 09:21 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:21 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:21 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Security updates * 09:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2012.codfw.wmnet with reason: host reimage * 09:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1011.eqiad.wmnet with reason: host reimage * 09:08 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2012.codfw.wmnet with reason: host reimage * 09:05 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1011.eqiad.wmnet with reason: host reimage * 08:55 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:52 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1011.eqiad.wmnet with OS trixie * 08:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1022: Security updates * 08:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2012 * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2012 * 08:50 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1022: Security updates * 08:50 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2012 * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2012.codfw.wmnet 44.48.192.10.in-addr.arpa 4.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:50 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2012.codfw.wmnet 44.48.192.10.in-addr.arpa 4.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2012 - mvernon@cumin2003" * 08:50 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2012 - mvernon@cumin2003" * 08:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1012.eqiad.wmnet with OS trixie * 08:44 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 08:44 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2012 * 08:43 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2012.codfw.wmnet with OS trixie * 08:42 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2013.codfw.wmnet * 08:41 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2013.codfw.wmnet * 08:35 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 08:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host krb1002.eqiad.wmnet * 08:30 elukey@dns1004: END - running authdns-update * 08:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1012.eqiad.wmnet with reason: host reimage * 08:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Security updates * 08:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:28 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:28 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Security updates * 08:27 elukey@dns1004: START - running authdns-update * 08:26 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 08:26 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host krb1002.eqiad.wmnet * 08:22 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1012.eqiad.wmnet with reason: host reimage * 08:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host krb2002.codfw.wmnet * 08:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast6003.wikimedia.org * 08:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2013.codfw.wmnet with OS trixie * 08:13 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast6003.wikimedia.org * 08:12 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast3007.wikimedia.org * 08:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host krb2002.codfw.wmnet * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Security updates * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:09 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:09 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Security updates * 08:07 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1012.eqiad.wmnet with OS trixie * 08:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast3007.wikimedia.org * 08:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast5005.wikimedia.org * 07:58 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast5005.wikimedia.org * 07:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1013.eqiad.wmnet with OS trixie * 07:53 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2013.codfw.wmnet with reason: host reimage * 07:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1021: Security updates * 07:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:53 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:53 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1021: Security updates * 07:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast1004.wikimedia.org * 07:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2013.codfw.wmnet with reason: host reimage * 07:46 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast1004.wikimedia.org * 07:40 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1013.eqiad.wmnet with reason: host reimage * 07:36 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1013.eqiad.wmnet with reason: host reimage * 07:31 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2013 * 07:31 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2013 * 07:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1021: Security updates * 07:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:30 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:30 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1021: Security updates * 07:24 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2013 * 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2013.codfw.wmnet 87.0.192.10.in-addr.arpa 7.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:24 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2013.codfw.wmnet 87.0.192.10.in-addr.arpa 7.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2013 - mvernon@cumin2003" * 07:24 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2013 - mvernon@cumin2003" * 07:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1013.eqiad.wmnet with OS trixie * 07:19 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 07:19 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2013 * 07:19 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2013.codfw.wmnet with OS trixie * 07:13 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] (duration: 07m 48s) * 07:09 kharlan@deploy2003: kharlan: Continuing with deployment * 07:08 kharlan@deploy2003: kharlan: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:06 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 01:15 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 01:14 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply == 2026-07-14 == * 22:51 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_magru * 22:51 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7016.magru.wmnet * 22:46 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_magru * 22:46 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7008.magru.wmnet * 22:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7015.magru.wmnet * 22:04 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7007.magru.wmnet * 21:29 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7014.magru.wmnet * 21:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7006.magru.wmnet * 21:13 dzahn@cumin2002: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 0:15:00 on gerrit.wikimedia.org with reason: reboot * 21:11 mutante: gerrit2003 (gerrit.wikimedia.org) - reboot for maintenance * 21:11 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on gerrit2003.wikimedia.org with reason: reboot * 20:56 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:56 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:56 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:55 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 20:48 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7013.magru.wmnet * 20:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7005.magru.wmnet * 20:41 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host phab1005.eqiad.wmnet with OS trixie * 20:28 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] (duration: 06m 47s) * 20:24 sbassett@deploy2003: sbassett: Continuing with deployment * 20:23 sbassett@deploy2003: sbassett: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:23 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on phab1005.eqiad.wmnet with reason: host reimage * 20:21 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] * 20:20 aokoth@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on phab1005.eqiad.wmnet with reason: host reimage * 20:12 jhuneidi@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] (duration: 07m 42s) * 20:07 jhuneidi@deploy2003: jhuneidi, priyankar22: Continuing with deployment * 20:06 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7012.magru.wmnet * 20:06 jhuneidi@deploy2003: jhuneidi, priyankar22: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:04 jhuneidi@deploy2003: Started scap sync-world: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] * 20:02 aokoth@cumin1003: START - Cookbook sre.hosts.reimage for host phab1005.eqiad.wmnet with OS trixie * 20:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7004.magru.wmnet * 20:00 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet * 19:57 aokoth@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet * 19:24 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7011.magru.wmnet * 19:19 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7003.magru.wmnet * 19:11 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] (duration: 08m 33s) * 19:07 jforrester@deploy2003: jforrester: Continuing with deployment * 19:04 jforrester@deploy2003: jforrester: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:02 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] * 18:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7010.magru.wmnet * 18:38 mutante: rotating phabricator-gerrit bot token (its-phabricator) * 18:18 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 17:44 swfrench@deploy2003: Finished scap sync-world: Deployment to pick up new production image (duration: 31m 44s) * 17:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7002.magru.wmnet * 17:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7009.magru.wmnet * 17:32 swfrench@deploy2003: swfrench: Continuing with deployment * 17:29 swfrench@deploy2003: swfrench: Deployment to pick up new production image synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:17 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2035: repooling after rack b5 maintenance * 17:16 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool es2035: repooling after rack b5 maintenance * 17:16 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2188: repooling after rack b5 maintenance * 17:12 swfrench@deploy2003: Started scap sync-world: Deployment to pick up new production image * 17:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7001.magru.wmnet * 16:57 swfrench-wmf: reprepro include php8.3_8.3.32-1+wmf12u2 into component/php83 for bookworm-wikimedia * 16:50 sukhe: pool cp2046 * 16:47 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4039.ulsfo.wmnet * 16:44 sukhe: sudo cumin -b31 "A:cp" "run-puppet-agent" * 16:33 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on contint1003.wikimedia.org with reason: reboot * 16:32 mutante: contint1003 - main CI server - rebooting * 16:31 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2188: repooling after rack b5 maintenance * 16:31 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2178: repooling after rack b5 maintenance * 16:29 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 16:28 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 16:28 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 16:28 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 16:18 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2014.codfw.wmnet * 16:18 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 16:17 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2014.codfw.wmnet * 16:10 mvernon@cumin1003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-thanos-proxies (exit_code=0) rolling restart_daemons on A:thanos-fe * 16:09 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 16:07 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp4039.ulsfo.wmnet * 16:07 mvernon@cumin1003: START - Cookbook sre.swift.roll-restart-reboot-swift-thanos-proxies rolling restart_daemons on A:thanos-fe * 16:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2014.codfw.wmnet with OS trixie * 15:56 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1014.eqiad.wmnet with OS trixie * 15:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2014.codfw.wmnet with reason: host reimage * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2014 * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2014 * 15:28 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2014 * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2014.codfw.wmnet 194.16.192.10.in-addr.arpa 4.9.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:28 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2014.codfw.wmnet 194.16.192.10.in-addr.arpa 4.9.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2014 - mvernon@cumin2003" * 15:28 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2014 - mvernon@cumin2003" * 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Apply title-related policies when selecting the name of the entity - kamila@cumin1003" * 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Apply title-related policies when selecting the name of the entity - kamila@cumin1003 * 15:22 kamila@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Apply title-related policies when selecting the name of the entity - kamila@cumin1003 * 15:22 kamila@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Apply title-related policies when selecting the name of the entity - kamila@cumin1003" * 15:20 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 15:20 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2014 * 15:20 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2014.codfw.wmnet with OS trixie * 15:19 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1014.eqiad.wmnet with OS trixie * 15:01 dancy@deploy2003: Installation of scap version "4.274.1" completed for 3 hosts * 15:00 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2177: repooling after rack b5 maintenance * 15:00 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2159: repooling after rack b5 maintenance * 14:59 dancy@deploy2003: Installing scap version "4.274.1" for 3 host(s) * 14:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2015.codfw.wmnet with OS trixie * 14:54 seanleong-wmde: Finished populateSitesTable for isvwiki ([[phab:T429939|T429939]]) * 14:53 javiermonton@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] (duration: 07m 35s) * 14:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1015.eqiad.wmnet with OS trixie * 14:49 javiermonton@deploy2003: javiermonton: Continuing with deployment * 14:48 javiermonton@deploy2003: javiermonton: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:46 javiermonton@deploy2003: Started scap sync-world: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] * 14:42 otto@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 14:41 otto@deploy2003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 14:41 otto@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 14:40 otto@deploy2003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 14:40 otto@deploy2003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 14:39 otto@deploy2003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 14:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2015.codfw.wmnet with reason: host reimage * 14:34 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1015.eqiad.wmnet with reason: host reimage * 14:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2015.codfw.wmnet with reason: host reimage * 14:30 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1015.eqiad.wmnet with reason: host reimage * 14:30 seanleong-wmde@deploy2003: mwscript-k8s job started: foreachwikiindblist wikidataclient extensions/Wikibase/lib/maintenance/populateSitesTable.php --force-protocol https # [[phab:T429939|T429939]] * 14:24 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling reboot on A:durum and not (A:durum-eqiad or A:durum-codfw or A:durum-esams) and A:durum * 14:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2015.codfw.wmnet with OS trixie * 14:15 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2016.codfw.wmnet with OS trixie * 14:14 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2159: repooling after rack b5 maintenance * 14:14 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1015.eqiad.wmnet with OS trixie * 14:12 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1016.eqiad.wmnet with OS trixie * 14:12 cmooney@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=pki,name=codfw * 14:12 sbisson@deploy2003: helmfile [codfw] DONE helmfile.d/services/cxserver: sync * 14:11 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2002.codfw.wmnet * 14:11 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2002.codfw.wmnet * 14:11 sbisson@deploy2003: helmfile [codfw] START helmfile.d/services/cxserver: sync * 14:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1005.wikimedia.org * 14:07 sbisson@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cxserver: sync * 14:07 sbisson@deploy2003: helmfile [eqiad] START helmfile.d/services/cxserver: sync * 14:05 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1005.wikimedia.org * 14:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader2005.wikimedia.org * 14:02 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2003.codfw.wmnet * 14:02 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2003.codfw.wmnet * 14:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=tegola-vector-tiles,name=codfw * 14:00 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=kartotherian,name=codfw * 14:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader2005.wikimedia.org * 13:58 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2016.codfw.wmnet with reason: host reimage * 13:57 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-ncredir (exit_code=0) rolling reboot on A:ncredir and A:ncredir * 13:57 sbisson@deploy2003: helmfile [staging] DONE helmfile.d/services/cxserver: sync * 13:56 sbisson@deploy2003: helmfile [staging] START helmfile.d/services/cxserver: sync * 13:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1016.eqiad.wmnet with reason: host reimage * 13:52 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:52 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:51 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2016.codfw.wmnet with reason: host reimage * 13:50 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1016.eqiad.wmnet with reason: host reimage * 13:49 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy (exit_code=0) rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 13:49 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling reboot on A:wikidough * 13:46 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-tcp-proxy (exit_code=0) rolling reboot on A:tcpproxy and A:tcpproxy * 13:43 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and not (A:durum-eqiad or A:durum-codfw or A:durum-esams) and A:durum * 13:42 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=97) rolling reboot on A:durum and A:durum * 13:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2011.codfw.wmnet * 13:36 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-reboot (exit_code=0) rolling reboot on A:dnsbox and A:ulsfo and (A:dnsbox) * 13:36 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns4004.wikimedia.org * 13:34 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1016.eqiad.wmnet with OS trixie * 13:34 topranks: reboot lsw1-b5-codfw to upgrade JunOS [[phab:T430918|T430918]] * 13:34 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2016.codfw.wmnet with OS trixie * 13:32 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2002.codfw.wmnet * 13:31 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2011.codfw.wmnet * 13:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2012.codfw.wmnet * 13:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2017.codfw.wmnet with OS trixie * 13:25 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1017.eqiad.wmnet with OS trixie * 13:24 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2012.codfw.wmnet * 13:22 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2002.codfw.wmnet * 13:22 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:22 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:22 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns4004.wikimedia.org * 13:21 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2005.codfw.wmnet * 13:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2013.codfw.wmnet * 13:18 elukey@dns1004: END - running authdns-update * 13:17 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2005.codfw.wmnet * 13:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2004.codfw.wmnet * 13:16 elukey@dns1004: START - running authdns-update * 13:16 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1046: es1046 after reimage * 13:14 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1029.eqiad.wmnet,service=s8 * 13:14 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1029.eqiad.wmnet,service=s5 * 13:13 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1029.eqiad.wmnet,service=s5 * 13:13 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1029.eqiad.wmnet,service=s8 * 13:13 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2004.codfw.wmnet * 13:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2013.codfw.wmnet * 13:11 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2014.codfw.wmnet * 13:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2003.codfw.wmnet * 13:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2017.codfw.wmnet with reason: host reimage * 13:07 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:07 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns4003.wikimedia.org * 13:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2003.codfw.wmnet * 13:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1067.eqiad.wmnet * 13:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1067.eqiad.wmnet * 13:06 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1067.eqiad.wmnet * 13:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm-test1001.wikimedia.org * 13:05 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2017.codfw.wmnet with reason: host reimage * 13:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1017.eqiad.wmnet with reason: host reimage * 13:04 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2014.codfw.wmnet * 13:03 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:02 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2188: codfw rack B5 depool for maintenance * 13:02 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_magru * 13:01 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2188: codfw rack B5 depool for maintenance * 13:01 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_magru * 13:01 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2178: codfw rack B5 depool for maintenance * 13:01 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm-test1001.wikimedia.org * 13:01 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2178: codfw rack B5 depool for maintenance * 13:01 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2177: codfw rack B5 depool for maintenance * 13:00 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2177: codfw rack B5 depool for maintenance * 12:59 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1068.eqiad.wmnet * 12:59 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1068.eqiad.wmnet * 12:58 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola-vector-tiles,name=codfw * 12:58 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2159: codfw rack B5 depool for maintenance * 12:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1017.eqiad.wmnet with reason: host reimage * 12:58 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola,name=codfw * 12:57 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=kartotherian,name=codfw * 12:57 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2159: codfw rack B5 depool for maintenance * 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 30 hosts with reason: lsw1-b5-codfw JunOS upgrade * 12:55 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lsw1-b5-codfw,lsw1-b5-codfw IPv6,lsw1-b5-codfw.mgmt,ssw1-a[1,8]-codfw.mgmt with reason: switch upgade lsw1-b5-codfw * 12:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet * 12:49 cmooney@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=pki,name=codfw * 12:49 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1067.eqiad.wmnet with OS trixie * 12:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet * 12:48 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2017.codfw.wmnet with OS trixie * 12:47 topranks: depool codfw pki in dns discovery ahead of lsw1-b5-codfw maintenance [[phab:T430918|T430918]] * 12:47 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns4003.wikimedia.org * 12:47 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and A:ulsfo and (A:dnsbox) * 12:47 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2018.codfw.wmnet * 12:45 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2018.codfw.wmnet * 12:45 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and A:durum * 12:45 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-tcp-proxy rolling reboot on A:tcpproxy and A:tcpproxy * 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host pki-root1002.eqiad.wmnet * 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1009.eqiad.wmnet * 12:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1009.eqiad.wmnet * 12:44 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 12:43 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-ncredir rolling reboot on A:ncredir and A:ncredir * 12:43 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling reboot on A:wikidough * 12:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2018.codfw.wmnet with OS trixie * 12:42 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1017.eqiad.wmnet with OS trixie * 12:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader2006.wikimedia.org * 12:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1018.eqiad.wmnet with OS trixie * 12:39 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1009.eqiad.wmnet * 12:38 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host pki-root1002.eqiad.wmnet * 12:38 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1009.eqiad.wmnet * 12:38 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1008.eqiad.wmnet * 12:38 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1008.eqiad.wmnet * 12:35 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader2006.wikimedia.org * 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1006.wikimedia.org * 12:33 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1008.eqiad.wmnet * 12:30 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1046: es1046 after reimage * 12:29 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host es1046.eqiad.wmnet with OS trixie * 12:29 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1006.wikimedia.org * 12:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test2005.wikimedia.org * 12:28 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1008.eqiad.wmnet * 12:27 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1007.eqiad.wmnet * 12:27 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1007.eqiad.wmnet * 12:27 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1067.eqiad.wmnet with reason: host reimage * 12:25 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1068.eqiad.wmnet with reason: vacuum overlarge container dbs * 12:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2018.codfw.wmnet with reason: host reimage * 12:24 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test2005.wikimedia.org * 12:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test1005.wikimedia.org * 12:23 Amir1: mwscript-k8s --follow --dblist=ores -- extensions/ORES/maintenance/PurgeScoreCache.php --model goodfaith --old ([[phab:T431159|T431159]]) * 12:22 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1007.eqiad.wmnet * 12:22 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test1005.wikimedia.org * 12:22 atsukoito: restarting pybal on lvs2013 `low-traffic` for https://gerrit.wikimedia.org/r/1310535 * 12:22 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1007.eqiad.wmnet * 12:21 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1006.eqiad.wmnet * 12:21 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1006.eqiad.wmnet * 12:20 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1018.eqiad.wmnet with reason: host reimage * 12:19 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2018.codfw.wmnet with reason: host reimage * 12:18 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1067.eqiad.wmnet with reason: host reimage * 12:16 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1006.eqiad.wmnet * 12:15 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1006.eqiad.wmnet * 12:15 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1005.eqiad.wmnet * 12:15 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1005.eqiad.wmnet * 12:15 atsukoito: restarting pybal on lvs2014 for https://gerrit.wikimedia.org/r/1310535 * 12:12 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1018.eqiad.wmnet with reason: host reimage * 12:11 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1005.eqiad.wmnet * 12:11 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1005.eqiad.wmnet * 12:10 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1004.eqiad.wmnet * 12:10 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1004.eqiad.wmnet * 12:09 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on es1046.eqiad.wmnet with reason: host reimage * 12:08 atsukoito: restarting pybal on lvs1019 `low-traffic` for https://gerrit.wikimedia.org/r/1310535 * 12:06 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1004.eqiad.wmnet * 12:06 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1004.eqiad.wmnet * 12:06 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1003.eqiad.wmnet * 12:06 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1003.eqiad.wmnet * 12:05 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on es1046.eqiad.wmnet with reason: host reimage * 12:05 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "set ml-serve1001 back to active state - cmooney@cumin1003" * 12:04 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "set ml-serve1001 back to active state - cmooney@cumin1003" * 12:04 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:02 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1003.eqiad.wmnet * 12:01 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1003.eqiad.wmnet * 12:01 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1002.eqiad.wmnet * 12:01 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1002.eqiad.wmnet * 12:01 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:59 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2018.codfw.wmnet with OS trixie * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1067 * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1067 * 11:59 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1067 * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1067.eqiad.wmnet 17.48.64.10.in-addr.arpa 7.1.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:59 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1067.eqiad.wmnet 17.48.64.10.in-addr.arpa 7.1.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1067 - blake@cumin1003" * 11:58 atsukoito: restarting pybal on lvs1018 `high-traffic2` for https://gerrit.wikimedia.org/r/1310535 * 11:57 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1002.eqiad.wmnet * 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2019.codfw.wmnet with OS trixie * 11:56 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1002.eqiad.wmnet * 11:56 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 11:56 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1018.eqiad.wmnet with OS trixie * 11:54 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 11:54 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:54 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1019.eqiad.wmnet with OS trixie * 11:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:49 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:49 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:49 aikochou@deploy2003: helmfile [codfw] DONE helmfile.d/services/changeprop: sync * 11:48 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host es1046.eqiad.wmnet with OS trixie * 11:48 aikochou@deploy2003: helmfile [codfw] START helmfile.d/services/changeprop: sync * 11:48 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310535 * 11:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1046: Reimage to Trixie * 11:44 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1046: Reimage to Trixie * 11:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5:00:00 on es1046.eqiad.wmnet with reason: Reimage to Trixie * 11:42 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:42 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:42 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 11:42 aikochou@deploy2003: helmfile [eqiad] DONE helmfile.d/services/changeprop: sync * 11:41 aikochou@deploy2003: helmfile [eqiad] START helmfile.d/services/changeprop: sync * 11:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2019.codfw.wmnet with reason: host reimage * 11:36 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] (duration: 09m 41s) * 11:36 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:36 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:35 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:35 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:32 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1019.eqiad.wmnet with reason: host reimage * 11:32 jforrester@deploy2003: jforrester, gengh: Continuing with deployment * 11:29 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2019.codfw.wmnet with reason: host reimage * 11:28 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1019.eqiad.wmnet with reason: host reimage * 11:28 jforrester@deploy2003: jforrester, gengh: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:26 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] * 11:20 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2003.codfw.wmnet * 11:20 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:19 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2003.codfw.wmnet * 11:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:12 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1019.eqiad.wmnet with OS trixie * 11:10 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2019.codfw.wmnet with OS trixie * 11:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1020.eqiad.wmnet with OS trixie * 11:10 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1067 - blake@cumin1003" * 11:09 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] (duration: 12m 12s) * 11:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2020.codfw.wmnet with OS trixie * 11:03 kharlan@deploy2003: kharlan: Continuing with deployment * 11:01 blake@cumin1003: START - Cookbook sre.dns.netbox * 11:01 kharlan@deploy2003: kharlan: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:57 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] * 10:55 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] (duration: 31m 40s) * 10:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1020.eqiad.wmnet with reason: host reimage * 10:52 marostegui@dns1004: START - running authdns-update * 10:49 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2020.codfw.wmnet with reason: host reimage * 10:49 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:48 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1020.eqiad.wmnet with reason: host reimage * 10:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2020.codfw.wmnet with reason: host reimage * 10:43 kharlan@deploy2003: kharlan: Continuing with deployment * 10:42 kharlan@deploy2003: kharlan: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:32 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1020.eqiad.wmnet with OS trixie * 10:29 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2159: Repooling after switchover * 10:29 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1067 * 10:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1021.eqiad.wmnet with OS trixie * 10:27 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1067.eqiad.wmnet with OS trixie * 10:27 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:27 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1067.eqiad.wmnet * 10:27 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:26 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1067.eqiad.wmnet * 10:26 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1067.eqiad.wmnet * 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2020.codfw.wmnet with OS trixie * 10:26 blake@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1055.eqiad.wmnet * 10:26 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1055.eqiad.wmnet * 10:26 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1055.eqiad.wmnet * 10:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2021.codfw.wmnet with OS trixie * 10:24 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] * 10:11 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1055.eqiad.wmnet with OS trixie * 10:09 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1021.eqiad.wmnet with reason: host reimage * 10:05 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2021.codfw.wmnet with reason: host reimage * 10:03 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310129 revert * 10:02 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1021.eqiad.wmnet with reason: host reimage * 10:01 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2021.codfw.wmnet with reason: host reimage * 09:58 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310129 * 09:50 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1055.eqiad.wmnet with reason: host reimage * 09:45 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1055.eqiad.wmnet with reason: host reimage * 09:45 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1021.eqiad.wmnet with OS trixie * 09:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1022.eqiad.wmnet with OS trixie * 09:44 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2159: Repooling after switchover * 09:44 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2021.codfw.wmnet with OS trixie * 09:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2022.codfw.wmnet with OS trixie * 09:31 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2159.codfw.wmnet * 09:28 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1055 * 09:28 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1055 * 09:27 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on ms-fe1022.eqiad.wmnet with reason: host reimage * 09:27 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1022.eqiad.wmnet with reason: host reimage * 09:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2022.codfw.wmnet with reason: host reimage * 09:21 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2022.codfw.wmnet with reason: host reimage * 09:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2159: Rebooting db2159.codfw.wmnet * 09:20 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2159: Rebooting db2159.codfw.wmnet * 09:18 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 09:18 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 09:18 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 09:17 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 09:16 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2159.codfw.wmnet * 09:13 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b] (thin): Regular analytics weekly train THIN [analytics/refinery@ad6e05b8] (duration: 02m 07s) * 09:11 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b] (thin): Regular analytics weekly train THIN [analytics/refinery@ad6e05b8] * 09:10 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1022.eqiad.wmnet with OS trixie * 09:07 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1023.eqiad.wmnet with OS trixie * 09:06 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b]: Regular analytics weekly train [analytics/refinery@ad6e05b8] (duration: 04m 49s) * 09:04 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2022.codfw.wmnet with OS trixie * 09:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2023.codfw.wmnet with OS trixie * 09:01 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b]: Regular analytics weekly train [analytics/refinery@ad6e05b8] * 09:01 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@ad6e05b8] (duration: 02m 01s) * 09:00 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1055 * 09:00 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1055.eqiad.wmnet 50.32.64.10.in-addr.arpa 0.5.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:00 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1055.eqiad.wmnet 50.32.64.10.in-addr.arpa 0.5.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:00 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:00 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1055 - blake@cumin1003" * 09:00 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1055 - blake@cumin1003" * 08:59 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@ad6e05b8] * 08:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2159 [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94811 and previous config saved to /var/cache/conftool/dbconfig/20260714-085624-cwilliams.json * 08:55 blake@cumin1003: START - Cookbook sre.dns.netbox * 08:55 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1055 * 08:54 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1055.eqiad.wmnet with OS trixie * 08:54 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1055.eqiad.wmnet * 08:53 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1055.eqiad.wmnet * 08:53 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1055.eqiad.wmnet * 08:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2220 to s7 primary [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94810 and previous config saved to /var/cache/conftool/dbconfig/20260714-085239-cwilliams.json * 08:51 cezmunsta: Starting s7 codfw failover from db2159 to db2220 - [[phab:T430920|T430920]] * 08:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1023.eqiad.wmnet with reason: host reimage * 08:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2220 with weight 0 [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94809 and previous config saved to /var/cache/conftool/dbconfig/20260714-084553-cwilliams.json * 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s7 [[phab:T430920|T430920]] * 08:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2023.codfw.wmnet with reason: host reimage * 08:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1023.eqiad.wmnet with reason: host reimage * 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2023.codfw.wmnet with reason: host reimage * 08:34 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:34 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:29 marostegui@dns1004: END - running authdns-update * 08:29 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox-dev2003.codfw.wmnet * 08:27 marostegui@dns1004: START - running authdns-update * 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker2*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2009.codfw.wmnet * 08:26 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2009.codfw.wmnet * 08:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1023.eqiad.wmnet with OS trixie * 08:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox-dev2003.codfw.wmnet * 08:24 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:24 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:24 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2023.codfw.wmnet with OS trixie * 08:24 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1029.eqiad.wmnet with reason: reboot * 08:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1027.eqiad.wmnet with reason: reboot * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:21 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2009.codfw.wmnet * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:20 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2009.codfw.wmnet * 08:20 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2008.codfw.wmnet * 08:20 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2008.codfw.wmnet * 08:15 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2008.codfw.wmnet * 08:14 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2008.codfw.wmnet * 08:14 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2007.codfw.wmnet * 08:14 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2007.codfw.wmnet * 08:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1024.eqiad.wmnet with OS trixie * 08:12 elukey@cumin1003: END (PASS) - Cookbook sre.pki.restart-reboot (exit_code=0) rolling reboot on P<nowiki>{</nowiki>pki*<nowiki>}</nowiki> and (A:pki) * 08:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 08:10 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 08:09 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2007.codfw.wmnet * 08:08 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2007.codfw.wmnet * 08:08 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2006.codfw.wmnet * 08:08 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2006.codfw.wmnet * 08:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2024.codfw.wmnet with OS trixie * 08:03 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2006.codfw.wmnet * 08:02 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2006.codfw.wmnet * 08:02 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2005.codfw.wmnet * 08:02 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2005.codfw.wmnet * 07:58 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2005.codfw.wmnet * 07:58 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2005.codfw.wmnet * 07:57 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2004.codfw.wmnet * 07:57 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2004.codfw.wmnet * 07:54 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki.discovery.wmnet. on all recursors * 07:54 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache pki.discovery.wmnet. on all recursors * 07:53 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2004.codfw.wmnet * 07:53 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2004.codfw.wmnet * 07:53 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2003.codfw.wmnet * 07:53 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2003.codfw.wmnet * 07:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1024.eqiad.wmnet with reason: host reimage * 07:49 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki.discovery.wmnet. on all recursors * 07:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2003.codfw.wmnet * 07:49 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache pki.discovery.wmnet. on all recursors * 07:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2024.codfw.wmnet with reason: host reimage * 07:48 elukey@cumin1003: START - Cookbook sre.pki.restart-reboot rolling reboot on P<nowiki>{</nowiki>pki*<nowiki>}</nowiki> and (A:pki) * 07:46 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1024.eqiad.wmnet with reason: host reimage * 07:45 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2024.codfw.wmnet with reason: host reimage * 07:45 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2003.codfw.wmnet * 07:44 elukey@cumin1003: END (PASS) - Cookbook sre.misc-clusters.restart-reboot-config-master (exit_code=0) rolling reboot on P<nowiki>{</nowiki>config-master*<nowiki>}</nowiki> and (A:config-master or A:config-master-eqiad or A:config-master-codfw) * 07:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2002.codfw.wmnet * 07:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2002.codfw.wmnet * 07:39 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2002.codfw.wmnet * 07:39 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) config-master.discovery.wmnet. on all recursors * 07:39 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache config-master.discovery.wmnet. on all recursors * 07:39 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2002.codfw.wmnet * 07:39 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker2*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl200*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl2003.codfw.wmnet * 07:36 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl2003.codfw.wmnet * 07:35 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) config-master.discovery.wmnet. on all recursors * 07:35 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache config-master.discovery.wmnet. on all recursors * 07:34 elukey@cumin1003: START - Cookbook sre.misc-clusters.restart-reboot-config-master rolling reboot on P<nowiki>{</nowiki>config-master*<nowiki>}</nowiki> and (A:config-master or A:config-master-eqiad or A:config-master-codfw) * 07:31 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl2003.codfw.wmnet * 07:31 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl2003.codfw.wmnet * 07:31 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl2002.codfw.wmnet * 07:31 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl2002.codfw.wmnet * 07:29 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1024.eqiad.wmnet with OS trixie * 07:28 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2024.codfw.wmnet with OS trixie * 07:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl2002.codfw.wmnet * 07:26 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl2002.codfw.wmnet * 07:26 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl200*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 07:26 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 07:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 06:50 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lists1004.wikimedia.org * 06:44 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host lists1004.wikimedia.org * 06:25 marostegui@dns1004: END - running authdns-update * 06:23 marostegui@dns1004: START - running authdns-update * 06:22 marostegui@dns1004: END - running authdns-update * 06:20 marostegui@dns1004: START - running authdns-update * 06:17 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1026.eqiad.wmnet with reason: reboot * 06:04 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: sync * 06:04 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: sync * 06:03 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync * 06:03 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync * 06:02 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync * 06:01 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync * 06:01 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync * 06:00 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync * 05:59 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:59 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:40 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:39 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:26 marostegui@dns1004: END - running authdns-update * 05:24 marostegui@dns1004: START - running authdns-update * 05:24 marostegui@dns1004: START - running authdns-update * 05:13 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1004.wikimedia.org * 05:07 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1004.wikimedia.org * 04:01 mwpresync@deploy2003: Pruned MediaWiki: 1.47.0-wmf.8 (duration: 01m 07s) * 03:39 mwpresync@deploy2003: Finished scap sync-world: testwikis to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] (duration: 36m 01s) * 03:03 mwpresync@deploy2003: Started scap sync-world: testwikis to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 29s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-13 == * 23:33 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1064.eqiad.wmnet * 23:33 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1064.eqiad.wmnet * 23:08 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1064.eqiad.wmnet with reason: vacuum overlarge container dbs * 23:06 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1069.eqiad.wmnet * 23:06 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1069.eqiad.wmnet * 22:34 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1069.eqiad.wmnet with reason: vacuum overlarge container dbs * 21:18 maryum: Deployed security fix for [[phab:T321092|T321092]] * 20:28 swfrench-wmf: reprepro include etcd-mirror_0.0.12-1+deb13u1 into main for trixie-wikimedia - [[phab:T424266|T424266]] * 20:26 swfrench-wmf: reprepro include etcd-mirror_0.0.12-1+deb12u1 into main for bookworm-wikimedia - [[phab:T428495|T428495]] * 20:23 dancy@deploy2003: Finished scap sync-world: Testing [[phab:T431635|T431635]] (duration: 03m 36s) * 20:19 dancy@deploy2003: Started scap sync-world: Testing [[phab:T431635|T431635]] * 20:18 dancy@deploy2003: Installation of scap version "4.274.0" completed for 3 hosts * 20:16 dancy@deploy2003: Installing scap version "4.274.0" for 3 host(s) * 20:12 kemayo@deploy2003: Finished scap sync-world: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] (duration: 08m 25s) * 20:07 kemayo@deploy2003: soda, esanders, kemayo: Continuing with deployment * 20:05 kemayo@deploy2003: soda, esanders, kemayo: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there * 20:04 kemayo@deploy2003: Started scap sync-world: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] * 18:22 cwhite: lvextend vg0/srv +500g on centrallog hosts * 18:19 cdobbins@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS trixie * 17:46 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1071.eqiad.wmnet * 17:46 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1071.eqiad.wmnet * 17:13 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1071.eqiad.wmnet with reason: vacuum overlarge container dbs * 17:07 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1065.eqiad.wmnet * 17:07 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1065.eqiad.wmnet * 17:06 dzahn@dns1006: END - running authdns-update * 17:04 dzahn@dns1006: START - running authdns-update * 17:01 dzahn@dns1006: END - running authdns-update * 16:59 dzahn@dns1006: START - running authdns-update * 16:51 dancy@deploy2003: Finished scap sync-world: testing [[phab:T428971|T428971]] (duration: 03m 37s) * 16:47 dancy@deploy2003: Started scap sync-world: testing [[phab:T428971|T428971]] * 16:45 atsukoito: restarting pybal on lvs1019 to flush IP address for `cirrussearch1122.eqiad.wmnet` after moving the vlan [[phab:T431311|T431311]] * 16:42 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:42 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:42 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:42 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:42 Amir1: mwscript-k8s --follow --dblist=ores -- extensions/ORES/maintenance/PurgeScoreCache.php --model damaging --old ([[phab:T431159|T431159]]) * 16:34 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Pool test * 16:34 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 16:34 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 16:34 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Pool test * 16:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Depool test * 16:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 16:33 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 16:33 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Depool test * 16:31 dancy@deploy2003: Installation of scap version "4.273.0" completed for 159 hosts * 16:29 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1065.eqiad.wmnet with reason: vacuum overlarge container dbs * 16:27 dancy@deploy2003: Installing scap version "4.273.0" for 159 host(s) * 16:27 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics-external: sync * 16:27 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics-external: sync * 16:26 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics-external: sync * 16:26 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics-external: sync * 16:22 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync * 16:21 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync * 16:21 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: sync * 16:21 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: sync * 16:19 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync * 16:19 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync * 16:18 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync * 16:17 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync * 15:59 atsukoito: restarting pybal on lvs1018 for https://gerrit.wikimedia.org/r/1310117 * 15:55 aikochou@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 15:50 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310117 * 15:46 aikochou@deploy2003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 15:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host kafka-logging1006.eqiad.wmnet * 15:43 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host kafka-logging1006.eqiad.wmnet * 15:41 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host ganeti-test[2001-2003].codfw.wmnet * 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host ganeti-test[2001-2003].codfw.wmnet * 15:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host netbox1003.eqiad.wmnet * 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host netbox1003.eqiad.wmnet * 15:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host netbox2003.codfw.wmnet * 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host netbox2003.codfw.wmnet * 15:36 sukhe: restart pybal on lvs1020 * 15:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet * 15:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet * 15:08 btullis@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:06 btullis@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 15:01 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:01 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:35 cdobbins@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 14:34 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:33 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:33 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:32 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:29 cdobbins@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 14:28 marostegui@dns1004: END - running authdns-update * 14:28 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:27 marostegui@dns1004: START - running authdns-update * 14:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1023.eqiad.wmnet with reason: reboot * 14:18 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2009.codfw.wmnet with OS trixie * 14:14 swfrench-wmf: start rolling run-puppet-agent on A:cp for ATS config change - [[phab:T428909|T428909]] [[phab:T431838|T431838]] * 14:05 swfrench-wmf: disable-puppet on A:cp for ATS config change - [[phab:T428909|T428909]] [[phab:T431838|T431838]] * 14:05 cdobbins@cumin2002: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie * 14:02 marostegui@dns1004: END - running authdns-update * 14:00 marostegui@dns1004: START - running authdns-update * 14:00 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1070.eqiad.wmnet * 14:00 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1070.eqiad.wmnet * 13:58 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2009.codfw.wmnet with reason: host reimage * 13:57 cdobbins@cumin2002: conftool action : set/pooled=no; selector: name=dns7002.* * 13:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2009.codfw.wmnet with reason: host reimage * 13:48 rscout@deploy2003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply * 13:48 rscout@deploy2003: helmfile [eqiad] START helmfile.d/services/miscweb: apply * 13:48 rscout@deploy2003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply * 13:47 rscout@deploy2003: helmfile [codfw] START helmfile.d/services/miscweb: apply * 13:40 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:33 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2009.codfw.wmnet with OS trixie * 13:30 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1070.eqiad.wmnet with reason: vacuum overlarge container dbs * 13:28 aude@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] (duration: 11m 12s) * 13:23 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:22 aude@deploy2003: aikochou, javiermonton, aude, gkm563: Continuing with deployment * 13:22 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:19 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:19 aude@deploy2003: aikochou, javiermonton, aude, gkm563: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] synced to the testservers * 13:17 aude@deploy2003: Started scap sync-world: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] * 13:01 ladsgroup@deploy2003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 13:01 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:00 ladsgroup@deploy2003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 12:59 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:52 ladsgroup@deploy2003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 12:51 ladsgroup@deploy2003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 12:48 atsuko@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 12:48 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 12:47 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2008.codfw.wmnet with OS trixie * 12:47 atsuko@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 12:47 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply * 12:47 atsuko@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:46 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 12:45 atsuko@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:45 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply * 12:45 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:44 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:43 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] (duration: 07m 02s) * 12:38 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 12:37 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:36 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] * 12:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2008.codfw.wmnet with reason: host reimage * 12:23 Msz2001: Deployed changes to private code for Suggested Investigations * 12:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2008.codfw.wmnet with reason: host reimage * 12:20 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:19 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:17 atsuko@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 12:17 atsuko@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 12:16 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:15 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] (duration: 07m 14s) * 12:10 mszwarc@deploy2003: mszwarc: Continuing with deployment * 12:09 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:07 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] * 12:04 mszwarc@deploy2003: sync-world aborted: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] (duration: 00m 29s) * 12:03 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] * 12:01 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2008.codfw.wmnet with OS trixie * 12:00 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] (duration: 07m 37s) * 11:55 zabe@deploy2003: zabe: Continuing with deployment * 11:54 zabe@deploy2003: zabe: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:52 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] * 11:51 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:43 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:35 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:34 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:33 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:30 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:28 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:27 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:17 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2007.codfw.wmnet with OS trixie * 11:09 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] * 11:06 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=s8 * 11:00 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=x3 * 11:00 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=s5 * 10:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2007.codfw.wmnet with reason: host reimage * 10:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2007.codfw.wmnet with reason: host reimage * 10:51 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host dse-k8s-worker1023 * 10:50 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host dse-k8s-worker1023 * 10:44 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dse-k8s-worker1023 * 10:43 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host dse-k8s-worker1023 * 10:42 marostegui@cumin1003: dbctl commit (dc=all): 'Change x4 masters [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P94804 and previous config saved to /var/cache/conftool/dbconfig/20260713-104248-marostegui.json * 10:37 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:37 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:35 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:35 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:34 atsuko@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 10:34 atsuko@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 10:33 marostegui@cumin1003: dbctl commit (dc=all): 'Push x4 initial dbctl config [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P94803 and previous config saved to /var/cache/conftool/dbconfig/20260713-103259-marostegui.json * 10:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2007.codfw.wmnet with OS trixie * 09:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2006.codfw.wmnet with OS trixie * 09:42 marostegui@dns1004: END - running authdns-update * 09:40 marostegui@dns1004: START - running authdns-update * 09:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2006.codfw.wmnet with reason: host reimage * 09:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2006.codfw.wmnet with reason: host reimage * 09:06 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1024.eqiad.wmnet with reason: reboot * 09:06 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:01 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2006.codfw.wmnet with OS trixie * 08:44 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 08:43 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 08:43 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:42 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:42 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:42 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:41 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 08:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup2004.codfw.wmnet * 08:38 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:33 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host db1208.eqiad.wmnet * 08:30 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=x3 * 08:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2005.codfw.wmnet with OS trixie * 08:28 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup2004.codfw.wmnet * 08:28 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup2003.codfw.wmnet * 08:24 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1039: Repooling after testing * 08:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on clouddb1016.eqiad.wmnet with reason: cloning * 08:23 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s5 * 08:23 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s8 * 08:21 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 08:21 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 08:17 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup2003.codfw.wmnet * 08:17 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1004.eqiad.wmnet * 08:14 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1208.eqiad.wmnet * 08:11 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host phab1005.eqiad.wmnet * 08:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2005.codfw.wmnet with reason: host reimage * 08:07 marostegui@dns1004: END - running authdns-update * 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1004.eqiad.wmnet * 08:07 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1003.eqiad.wmnet * 08:07 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:05 marostegui@dns1004: START - running authdns-update * 08:05 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2005.codfw.wmnet with reason: host reimage * 08:05 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 08:05 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host phab1005.eqiad.wmnet * 08:05 marostegui@dns1004: START - running authdns-update * 08:05 marostegui@dns1004: START - running authdns-update * 08:05 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 08:04 marostegui@dns1004: START - running authdns-update * 08:00 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit1003.wikimedia.org * 07:58 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1003.eqiad.wmnet * 07:58 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1002-dev.eqiad.wmnet * 07:58 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:58 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:54 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1002-dev.eqiad.wmnet * 07:54 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1001-dev.eqiad.wmnet * 07:54 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit1003.wikimedia.org * 07:53 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 07:53 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:52 Msz2001: UTC morning backport+config window done * 07:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2005.codfw.wmnet with OS trixie * {{safesubst:SAL entry|1=07:50 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark (T429943}} * 07:49 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1001-dev.eqiad.wmnet * 07:46 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:46 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:45 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 07:45 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:44 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 07:44 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:43 mszwarc@deploy2003: mszwarc, danielyepezgarces, anzx: Continuing with deployment * {{safesubst:SAL entry|1=07:39 mszwarc@deploy2003: mszwarc, danielyepezgarces, anzx: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark}} * 07:39 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1039: Repooling after testing * {{safesubst:SAL entry|1=07:36 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark (T429943)}} * 07:35 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] (duration: 30m 03s) * 07:25 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit2002.wikimedia.org * 07:22 mszwarc@deploy2003: mszwarc: Continuing with deployment * 07:21 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:19 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit2002.wikimedia.org * 07:15 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aphlict1002.eqiad.wmnet * 07:11 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host aphlict1002.eqiad.wmnet * 07:08 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2003.wikimedia.org * 07:05 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] * 07:02 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2003.wikimedia.org * 07:02 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2002.wikimedia.org * 06:55 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2002.wikimedia.org * 06:55 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1003.wikimedia.org * 06:49 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1003.wikimedia.org * 06:34 marostegui: Drop m5 ipoid database [[phab:T431007|T431007]] * 06:29 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1027.eqiad.wmnet with reason: reboot * 06:24 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1028.eqiad.wmnet with reason: reboot * 06:21 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1025.eqiad.wmnet with reason: reboot * 06:17 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1022.eqiad.wmnet with reason: reboot * 06:03 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on dbproxy[2005-2008].codfw.wmnet with reason: reboot * 05:37 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1217,1228].eqiad.wmnet with reason: cloning * 05:11 marostegui: Drop users_to_rename table [[phab:T431842|T431842]] * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-12 == * 16:01 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2209 [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94792 and previous config saved to /var/cache/conftool/dbconfig/20260712-160124-marostegui.json * 15:58 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2205 to s3 primary [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94791 and previous config saved to /var/cache/conftool/dbconfig/20260712-155853-marostegui.json * 15:58 marostegui: Starting s3 codfw emergency failover from db2209 to db2205 - [[phab:T431950|T431950]] * 15:51 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2205 with weight 0 [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94790 and previous config saved to /var/cache/conftool/dbconfig/20260712-155135-marostegui.json * 15:51 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Primary switchover s3 [[phab:T431950|T431950]] * 02:01 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 01m 17s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-11 == * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 26s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-10 == * 19:12 jhathaway@dns1004: END - running authdns-update * 19:10 jhathaway@dns1004: START - running authdns-update * 18:23 mutante: vrts2002 rebooting (not the active host) * 18:21 mutante: lists2001, phab2003 - rebooting (not the active hosts) * 18:16 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on A:lvs-high-traffic2-codfw * 18:15 swfrench@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on A:lvs-high-traffic2-codfw * 17:15 mutante: [doc1004:~] $ sudo systemctl start rsync-doc-host-data-sync ([[phab:T431856|T431856]]) * 17:09 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1004.eqiad.wmnet * 17:08 jhathaway@dns1004: END - running authdns-update * 17:07 jhathaway@dns1004: START - running authdns-update * 17:06 jhathaway: depooling puppetserver1002, cause of errors is still unknown * 17:03 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1004.eqiad.wmnet * 16:57 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1003.eqiad.wmnet * 16:51 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1003.eqiad.wmnet * 16:48 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 16:48 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2004.codfw.wmnet * 16:42 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2004.codfw.wmnet * 16:41 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2003.codfw.wmnet * 16:35 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2003.codfw.wmnet * 16:33 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2002.codfw.wmnet * 16:27 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2002.codfw.wmnet * 16:25 mutante: gitlab-runners (production) rebooting cluster one by one * 16:17 mutante: etherpad1004/etherpad2002 - (etherpad.wikimedia.org) - rebooting * 16:13 mutante: doc1004/doc2003 (doc.wikimedia.org backends) - rebooting * 16:02 mutante: releases1003/releases2003 (releases.wikimedia.org backends) - rebooting for maintenance * 15:26 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2007-dev.codfw.wmnet * 15:19 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2007-dev.codfw.wmnet * 15:14 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host cloudcephosd2007-dev.codfw.wmnet * 15:14 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2007-dev.codfw.wmnet * 15:14 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host cloudcephosd2006-dev.codfw.wmnet * 15:07 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2006-dev.codfw.wmnet * 15:07 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2005-dev.codfw.wmnet * 14:59 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2005-dev.codfw.wmnet * 14:59 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2004-dev.codfw.wmnet * 14:53 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2004-dev.codfw.wmnet * 14:53 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2007-dev.codfw.wmnet * 14:51 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1054.eqiad.wmnet * 14:51 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1054.eqiad.wmnet * 14:51 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1054.eqiad.wmnet * 14:47 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2007-dev.codfw.wmnet * 14:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2006-dev.codfw.wmnet * 14:41 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2006-dev.codfw.wmnet * 14:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2005-dev.codfw.wmnet * 14:37 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2005-dev.codfw.wmnet * 14:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2005-dev.codfw.wmnet * 14:29 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2005-dev.codfw.wmnet * 14:29 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2006-dev.codfw.wmnet * 14:21 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2006-dev.codfw.wmnet * 14:21 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2010-dev.codfw.wmnet * 14:15 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2010-dev.codfw.wmnet * 14:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudgw2004-dev.codfw.wmnet * 14:10 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1054.eqiad.wmnet with OS trixie * 14:09 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudgw2004-dev.codfw.wmnet * 14:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudgw2003-dev.codfw.wmnet * 14:02 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudgw2003-dev.codfw.wmnet * 14:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2004-dev.codfw.wmnet * 13:53 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2004-dev.codfw.wmnet * 13:53 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2003-dev.codfw.wmnet * 13:48 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:44 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2003-dev.codfw.wmnet * 13:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2002-dev.codfw.wmnet * 13:42 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:41 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:41 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:37 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2002-dev.codfw.wmnet * 13:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudidp2001-dev.codfw.wmnet * 13:33 blake@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker1054.eqiad.wmnet with reason: host reimage * 13:33 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudidp2001-dev.codfw.wmnet * 13:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudnet2006-dev.codfw.wmnet * 13:26 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudnet2006-dev.codfw.wmnet * 13:26 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudnet2005-dev.codfw.wmnet * 13:23 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1054.eqiad.wmnet with reason: host reimage * 13:18 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudnet2005-dev.codfw.wmnet * 13:18 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudservices2005-dev.codfw.wmnet * 13:12 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudservices2005-dev.codfw.wmnet * 13:11 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudservices2004-dev.codfw.wmnet * 13:08 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudservices2004-dev.codfw.wmnet * 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudweb2002-dev.wikimedia.org * 13:05 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 13:05 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1054 * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1054 * 13:04 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1054 * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1054.eqiad.wmnet 49.32.64.10.in-addr.arpa 9.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:04 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1054.eqiad.wmnet 49.32.64.10.in-addr.arpa 9.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1054 - blake@cumin1003" * 13:04 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1054 - blake@cumin1003" * 13:01 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudweb2002-dev.wikimedia.org * 13:00 blake@cumin1003: START - Cookbook sre.dns.netbox * 12:59 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1054 * 12:57 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1054.eqiad.wmnet with OS trixie * 12:57 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1054.eqiad.wmnet * 12:56 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1054.eqiad.wmnet * 12:56 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1054.eqiad.wmnet * 12:47 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS trixie * 12:44 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:39 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 12:39 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 12:38 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:37 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:14 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:10 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:08 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:07 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:00 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:00 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:51 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:49 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:48 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:47 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:44 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:32 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2001.codfw.wmnet * 11:32 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1053.eqiad.wmnet * 11:32 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2001.codfw.wmnet * 11:32 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1053.eqiad.wmnet * 11:32 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1053.eqiad.wmnet * 11:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker2001.codfw.wmnet * 11:31 cgoubert@cumin1003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker2001.codfw.wmnet * 11:31 cgoubert@cumin1003: END (FAIL) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=1) rolling reimage on P<nowiki>{</nowiki>wikikube-worker2001*<nowiki>}</nowiki> and (A:wikikube-master-codfw or A:wikikube-worker-codfw) * 11:30 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:30 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:21 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 18 hosts with reason: reboot & upgrade * 11:20 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker2001.codfw.wmnet with OS trixie * 11:16 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:15 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:14 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:14 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 11 hosts * 11:14 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 11 hosts * 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:08 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 11:02 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1053.eqiad.wmnet with OS trixie * 11:01 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:58 cgoubert@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 10:57 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:38 cgoubert@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker2001.codfw.wmnet with OS trixie * 10:38 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2001.codfw.wmnet * 10:38 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2001.codfw.wmnet * 10:38 cgoubert@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on P<nowiki>{</nowiki>wikikube-worker2001*<nowiki>}</nowiki> and (A:wikikube-master-codfw or A:wikikube-worker-codfw) * 10:35 topranks: adjust IBGP outbound policy on lsw1-e2-codfw [[phab:T423430|T423430]] towards ssw1-e1-codfw * 10:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 cgoubert@cumin1003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:27 cgoubert@cumin1003: END (FAIL) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=1) rolling reimage on A:wikikube-worker-codfw * 10:27 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker2001.codfw.wmnet with OS bookworm * 10:25 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 10:24 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:24 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:15 cgoubert@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 10:11 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 10:11 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 10:08 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:08 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:07 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 10:06 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 10:00 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:55 cgoubert@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker2001.codfw.wmnet with OS bookworm * 09:55 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2005-2006,2011-2012].codfw.wmnet * 09:55 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2005-2006,2011-2012].codfw.wmnet * 09:51 cgoubert@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on A:wikikube-worker-codfw * 09:41 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1053.eqiad.wmnet with reason: host reimage * 09:37 topranks: apply new IBGP outbound policy on lsw1-e2-codfw [[phab:T423430|T423430]] * 09:36 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:36 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1053.eqiad.wmnet with reason: host reimage * 09:16 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1053 * 09:16 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1053 * 09:15 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1053 * 09:15 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1053.eqiad.wmnet 48.32.64.10.in-addr.arpa 8.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:15 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1053.eqiad.wmnet 48.32.64.10.in-addr.arpa 8.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:15 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:15 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1053 - blake@cumin1003" * 09:15 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1053 - blake@cumin1003" * 09:11 blake@cumin1003: START - Cookbook sre.dns.netbox * 09:11 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1053 * 09:08 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1053.eqiad.wmnet with OS trixie * 09:08 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1053.eqiad.wmnet * 09:08 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1053.eqiad.wmnet * 09:08 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1053.eqiad.wmnet * 09:04 brouberol@dns1004: END - running authdns-update * 09:03 brouberol@dns1004: START - running authdns-update * 08:41 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e] (thin): Regular analytics weekly train THIN [analytics/refinery@1abf22ea] (duration: 02m 11s) * 08:38 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e] (thin): Regular analytics weekly train THIN [analytics/refinery@1abf22ea] * 08:38 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e]: Regular analytics weekly train [analytics/refinery@1abf22ea] (duration: 05m 17s) * 08:38 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:34 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:33 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e]: Regular analytics weekly train [analytics/refinery@1abf22ea] * 08:32 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@1abf22ea] (duration: 02m 03s) * 08:30 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@1abf22ea] * 08:30 JavierMonton: Deploying Refinery at {{Gerrit|1abf22ea}} for changes 1308121/T427068 1306491/T430020 and {{Gerrit|1308190}} * 08:29 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:29 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 08:24 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:24 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 08:18 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:18 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 08:00 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db[2183-2184].codfw.wmnet * 08:00 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for db[2183-2184].codfw.wmnet * 07:52 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:52 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 07:49 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 11 hosts with reason: reboot & upgrade * 07:47 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:47 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 07:44 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:44 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 07:23 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 10 hosts * 07:23 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 10 hosts * 06:45 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 10 hosts with reason: reboot & upgrade * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 41s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-09 == * 23:33 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] (duration: 13m 26s) * 23:29 ladsgroup@deploy2003: ladsgroup, jdlrobson: Continuing with deployment * 23:22 ladsgroup@deploy2003: ladsgroup, jdlrobson: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:20 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] * 22:57 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1165.eqiad.wmnet * 22:56 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1165.eqiad.wmnet * 22:56 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1165.eqiad.wmnet * 22:45 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1165.eqiad.wmnet with OS trixie * 22:38 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 22:37 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 22:37 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 22:37 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:37 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 22:25 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1165.eqiad.wmnet with reason: host reimage * 22:17 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1165.eqiad.wmnet with reason: host reimage * 22:13 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 22:12 rzl@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 22:04 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 22:04 rzl@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1165 * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1165 * 22:02 jasmine@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1165 * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1165.eqiad.wmnet 115.48.64.10.in-addr.arpa 5.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:02 jasmine@cumin2002: START - Cookbook sre.dns.wipe-cache wikikube-worker1165.eqiad.wmnet 115.48.64.10.in-addr.arpa 5.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1165 - jasmine@cumin2002" * 22:02 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1165 - jasmine@cumin2002" * 22:02 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 21:57 jasmine@cumin2002: START - Cookbook sre.dns.netbox * 21:55 jasmine@cumin2002: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1165 * 21:54 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-worker1165.eqiad.wmnet with OS trixie * 21:54 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 21:54 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1165.eqiad.wmnet * 21:53 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 21:53 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1165.eqiad.wmnet * 21:53 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1165.eqiad.wmnet * 21:53 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 21:47 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 21:45 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 21:43 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 21:43 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 21:42 maryum: Deploy fix for [[phab:T431684|T431684]] * 21:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 21:27 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] (duration: 34m 14s) * 21:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs1002 * 21:23 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs1002 * 21:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS trixie * 21:22 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 21:20 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 22s) * 21:20 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 21:16 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 21:16 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 21:15 ladsgroup@deploy2003: ladsgroup: Continuing with deployment * 21:13 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2002.codfw.wmnet with OS bookworm * 21:11 ladsgroup@deploy2003: ladsgroup: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:08 ladsgroup@cumin1003: END (PASS) - Cookbook sre.wikireplicas.update-views (exit_code=0) * 21:07 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 6 hosts with reason: reboots * 20:54 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecycle work - bking@cumin2003 * 20:53 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:53 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] * 20:51 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) * 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 20:47 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecycle work - bking@cumin2003 * 20:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 20:41 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:41 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) * 20:40 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host relforge1008.eqiad.wmnet * 20:40 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1009.eqiad.wmnet with OS trixie * 20:33 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:32 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:32 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:31 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:31 ladsgroup@cumin1003: END (PASS) - Cookbook sre.wikireplicas.update-views (exit_code=0) * 20:29 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1008.eqiad.wmnet * 20:24 rzl@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 20:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2002.codfw.wmnet with OS bookworm * 20:23 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host relforge1008.eqiad.wmnet * 20:23 rzl@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 20:23 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1008.eqiad.wmnet * 20:22 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:22 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:21 bking@cumin2003: END (ERROR) - Cookbook sre.elasticsearch.rolling-operation (exit_code=97) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:21 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1009.eqiad.wmnet with reason: host reimage * 20:16 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:15 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1009.eqiad.wmnet with reason: host reimage * 20:12 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) * 20:02 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 19:55 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1009.eqiad.wmnet with OS trixie * 19:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 19:43 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 19:30 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 19:28 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 19:27 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 19:25 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 18:42 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 18:41 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 18:16 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 18:15 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 17:45 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for doh5004.wikimedia.org * 17:45 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for doh5004.wikimedia.org * 17:38 ladsgroup@deploy2003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 17:35 ladsgroup@deploy2003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 17:29 ladsgroup@deploy2003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 17:26 ladsgroup@deploy2003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 17:09 mutante: zuul[12]00[123] - rebooting for maintenance * 17:09 ebernhardson: start full in-place reindex of eqiad cirrussearch cluster * 17:08 dzahn@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-cluster (exit_code=99) * 17:08 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-cluster * 17:03 ebernhardson: start full in-place reindex of codfw cirrussearch cluster * 16:59 mutante: stewards1001/stewards2001 - reboot for maintenance * 16:54 ebernhardson: start full in-place reindex of cloudelastic cluster * 16:53 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 16:52 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply * 16:49 mutante: planet1003/planet2003 - rebooting * 16:47 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on doh5004.wikimedia.org with reason: random high load, investigating * 15:55 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 15:54 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 15:51 jynus: restarting backupmon1001 * 15:49 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 14 hosts * 15:49 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 14 hosts * 15:47 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backupmon1001.eqiad.wmnet with reason: restart * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:06 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 15:06 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 14:59 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply * 14:58 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply * 14:51 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 14 hosts * 14:51 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 14 hosts * 14:49 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 6 hosts with reason: reboot & upgrade * 14:48 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet * 14:48 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet * 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:42 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1052.eqiad.wmnet * 14:42 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1052.eqiad.wmnet * 14:42 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1052.eqiad.wmnet * 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:31 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:31 elukey@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: sync * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:30 elukey@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: sync * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:28 elukey@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: sync * 14:28 elukey@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: sync * 14:26 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:20 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1052.eqiad.wmnet with OS trixie * 14:19 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 6 hosts with reason: reboot & upgrade * 14:18 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:15 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:15 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:13 elukey: update druid indexation job for webrequest_sampled_live - [[phab:T427068|T427068]] * 14:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:09 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:09 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for papaul - jhancock@cumin2002" * 14:09 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for papaul - jhancock@cumin2002" * 14:07 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:07 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:04 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 14:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cuminunpriv1001.eqiad.wmnet * 13:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb1003.eqiad.wmnet * 13:59 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1052.eqiad.wmnet with reason: host reimage * 13:57 moritzm: installing requests security updates * 13:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cuminunpriv1001.eqiad.wmnet * 13:55 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb1003.eqiad.wmnet * 13:53 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1052.eqiad.wmnet with reason: host reimage * 13:50 moritzm: installing python-cryptography security updates * 13:47 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb2003.codfw.wmnet * 13:44 Msz2001: UTC afternoon config+backport window is done * 13:44 Msz2001: Updated `logging` on `metawiki` to fix log performers, [[phab:T431176|T431176]]#12105297 * 13:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb2003.codfw.wmnet * 13:43 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 13:43 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt1002.wikimedia.org * 13:41 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] (duration: 07m 30s) * 13:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt1002.wikimedia.org * 13:37 mszwarc@deploy2003: mszwarc: Continuing with deployment * 13:36 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1052 * 13:36 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1052 * 13:35 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:35 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1052 * 13:35 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1052.eqiad.wmnet 47.32.64.10.in-addr.arpa 7.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:35 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1052.eqiad.wmnet 47.32.64.10.in-addr.arpa 7.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:35 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:35 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1052 - blake@cumin1003" * 13:35 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1052 - blake@cumin1003" * 13:34 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] * 13:31 blake@cumin1003: START - Cookbook sre.dns.netbox * 13:31 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1052 * 13:30 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1052.eqiad.wmnet with OS trixie * 13:30 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1052.eqiad.wmnet * 13:29 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1052.eqiad.wmnet * 13:29 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1052.eqiad.wmnet * 13:17 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] (duration: 11m 26s) * 13:13 jforrester@deploy2003: jforrester: Continuing with deployment * 13:08 jforrester@deploy2003: jforrester: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:06 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] * 12:54 cgoubert@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply * 12:54 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:52 cgoubert@deploy2003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply * 12:45 cgoubert@deploy2003: helmfile [codfw] DONE helmfile.d/services/mobileapps: apply * 12:44 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:44 cgoubert@deploy2003: helmfile [codfw] START helmfile.d/services/mobileapps: apply * 12:43 cgoubert@deploy2003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 12:43 cgoubert@deploy2003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 12:42 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast4006.wikimedia.org * 12:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt2002.wikimedia.org * 12:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast7002.wikimedia.org * 12:18 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast4006.wikimedia.org * 12:18 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host ml-serve1004 * 12:18 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host ml-serve1004 * 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt2002.wikimedia.org * 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast7002.wikimedia.org * 12:10 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backup[2003,2014].codfw.wmnet with reason: reboot & upgrade * 12:10 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt-staging2001.codfw.wmnet * 12:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid1003.eqiad.wmnet * 12:06 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt-staging2001.codfw.wmnet * 12:05 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid1003.eqiad.wmnet * 12:03 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backup[1003,1014].eqiad.wmnet with reason: reboot & upgrade * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid2003.codfw.wmnet * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host irc1003.wikimedia.org * 11:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid2003.codfw.wmnet * 11:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host irc1003.wikimedia.org * 11:55 jmm@dns1004: END - running authdns-update * 11:53 jmm@dns1004: START - running authdns-update * 11:50 jmm@dns1004: END - running authdns-update * 11:48 jmm@dns1004: START - running authdns-update * 11:27 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host irc2003.wikimedia.org * 11:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host irc2003.wikimedia.org * 11:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint2001.codfw.wmnet * 11:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint1001.eqiad.wmnet * 11:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint2001.codfw.wmnet * 11:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint1001.eqiad.wmnet * 11:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-rw2001.wikimedia.org * 11:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-rw1001.wikimedia.org * 11:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-rw2001.wikimedia.org * 11:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-rw1001.wikimedia.org * 11:03 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon1003.wikimedia.org * 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2005.codfw.wmnet * 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2005.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 10:59 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2005.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 10:57 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon1003.wikimedia.org * 10:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon2002.wikimedia.org * 10:55 jmm@cumin2003: START - Cookbook sre.dns.netbox * 10:51 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon2002.wikimedia.org * 10:51 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:50 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2005.codfw.wmnet * 10:41 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:40 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host ml-serve1003 * 10:40 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host ml-serve1003 * 10:39 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2033.codfw.wmnet * 10:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install2005.wikimedia.org * 10:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install1005.wikimedia.org * 10:35 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1004.eqiad.wmnet with OS bookworm * 10:31 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install1005.wikimedia.org * 10:31 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install2005.wikimedia.org * 10:30 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install4004.wikimedia.org * 10:30 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install3004.wikimedia.org * 10:29 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install3004.wikimedia.org * 10:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install4004.wikimedia.org * 10:23 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 10:21 moritzm: failover Ganeti master in codfw/routed to ganeti2034 [[phab:T430928|T430928]] * 10:19 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.addnode (exit_code=0) for new host ganeti2031.codfw.wmnet to cluster codfw and group B * 10:19 moritzm: readded ganeti2031 to the codfw Ganeti cluster [[phab:T430910|T430910]] * 10:18 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1004.eqiad.wmnet with reason: host reimage * 10:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install5004.wikimedia.org * 10:18 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1003 * 10:18 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1003 * 10:17 jmm@cumin2003: START - Cookbook sre.ganeti.addnode for new host ganeti2031.codfw.wmnet to cluster codfw and group B * 10:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install6003.wikimedia.org * 10:16 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install5004.wikimedia.org * 10:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install6003.wikimedia.org * 10:15 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1004.eqiad.wmnet with reason: host reimage * 10:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1001.eqiad.wmnet * 10:14 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 10:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2008.wikimedia.org * 10:00 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ml-serve1004 * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1004 * 09:57 jmm@cumin2003: START - Cookbook sre.dns.netbox * 09:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install7002.wikimedia.org * 09:57 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1004 * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ml-serve1004.eqiad.wmnet 50.48.64.10.in-addr.arpa 0.5.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:57 klausman@cumin1003: START - Cookbook sre.dns.wipe-cache ml-serve1004.eqiad.wmnet 50.48.64.10.in-addr.arpa 0.5.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1004 - klausman@cumin1003" * 09:56 klausman@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1004 - klausman@cumin1003" * 09:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-coord1001.eqiad.wmnet * 09:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 09:55 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow7002.magru.wmnet * 09:52 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-coord1001.eqiad.wmnet * 09:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 09:52 klausman@cumin1003: START - Cookbook sre.dns.netbox * 09:50 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install7002.wikimedia.org * 09:50 klausman@cumin1003: START - Cookbook sre.hosts.move-vlan for host ml-serve1004 * 09:50 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1004.eqiad.wmnet with OS bookworm * 09:50 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1003.eqiad.wmnet with OS bookworm * 09:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1001.eqiad.wmnet * 09:49 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 09:49 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2008.wikimedia.org * 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2007.codfw.wmnet * 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2007.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 09:49 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow7002.magru.wmnet * 09:49 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2007.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 09:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard1003.eqiad.wmnet * 09:39 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard2003.codfw.wmnet * 09:39 jmm@cumin2003: START - Cookbook sre.dns.netbox * 09:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard1003.eqiad.wmnet * 09:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor1003.eqiad.wmnet * 09:35 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard2003.codfw.wmnet * 09:34 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2007.codfw.wmnet * 09:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor1003.eqiad.wmnet * 09:33 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor-dev2001.codfw.wmnet * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor2003.codfw.wmnet * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sretest1006.eqiad.wmnet * 09:27 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 09:25 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor-dev2001.codfw.wmnet * 09:25 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor2003.codfw.wmnet * 09:23 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] (duration: 06m 27s) * 09:23 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2205: codfw rack B4 repool after maintenance * 09:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host sretest1006.eqiad.wmnet * 09:23 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2204: codfw rack B4 repool after maintenance * 09:19 urbanecm@deploy2003: urbanecm: Continuing with deployment * 09:19 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:18 jmm@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 6 hosts with reason: reboot * 09:17 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] * 09:08 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ml-serve1003 * 09:08 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1003 * 09:07 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1003 * 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ml-serve1003.eqiad.wmnet 81.32.64.10.in-addr.arpa 1.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:07 klausman@cumin1003: START - Cookbook sre.dns.wipe-cache ml-serve1003.eqiad.wmnet 81.32.64.10.in-addr.arpa 1.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1003 - klausman@cumin1003" * 09:06 klausman@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1003 - klausman@cumin1003" * 08:58 klausman@cumin1003: START - Cookbook sre.dns.netbox * 08:57 klausman@cumin1003: START - Cookbook sre.hosts.move-vlan for host ml-serve1003 * 08:57 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1003.eqiad.wmnet with OS bookworm * 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=0) rolling reimage on P<nowiki>{</nowiki>ml-serve1003.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet * 08:55 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet * 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1003.eqiad.wmnet with OS bookworm * 08:39 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 08:38 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool db2205: codfw rack B4 repool after maintenance * 08:37 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool db2204: codfw rack B4 repool after maintenance * 08:36 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 08:35 hashar@deploy2003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 08:32 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:32 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:31 hashar@deploy2003: Rolling back deployment * 08:26 moritzm: failover Ganeti master in codfw to ganeti2048 [[phab:T430928|T430928]] * 08:16 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1003.eqiad.wmnet with OS bookworm * 08:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2004.codfw.wmnet * 08:16 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet * 08:16 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet * 08:16 klausman@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on P<nowiki>{</nowiki>ml-serve1003.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 08:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2002.codfw.wmnet * 08:15 XioNoX: lsw1-b4-codfw> request system reboot - [[phab:T430910|T430910]] * 08:15 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b4-codfw,lsw1-b4-codfw IPv6,lsw1-b4-codfw.mgmt with reason: Switch maintenance * 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for codfw rack B4 * 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:10 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2004.codfw.wmnet * 08:10 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2002.codfw.wmnet * 08:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2205: codfw rack B4 depool for maintenance * 08:08 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool db2205: codfw rack B4 depool for maintenance * 08:08 jmm@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin2003.codfw.wmnet * 08:08 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2204: codfw rack B4 depool for maintenance * 08:08 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool db2204: codfw rack B4 depool for maintenance * 08:08 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 27 hosts with reason: codfw rack B4 depool for maintenance * 08:03 jmm@cumin2002: START - Cookbook sre.hosts.reboot-single for host cumin2003.codfw.wmnet * 07:56 ayounsi@cumin1003: START - Cookbook sre.network.depool-rack with action 'depool' for codfw rack B4 * 07:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1008.eqiad.wmnet with OS trixie * 07:49 wmde-fisch@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] (duration: 08m 36s) * 07:44 wmde-fisch@deploy2003: wmde-fisch: Continuing with deployment * 07:43 wmde-fisch@deploy2003: wmde-fisch: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:41 wmde-fisch@deploy2003: Started scap sync-world: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] * 07:35 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1008.eqiad.wmnet with reason: host reimage * 07:31 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1008.eqiad.wmnet with reason: host reimage * 07:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1008.eqiad.wmnet with OS trixie * 07:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 07:00 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 06:59 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1008.eqiad.wmnet with OS trixie * 06:57 Emperor: rebalance thanos swift rings after previous re-image of thanos-fe1004 to trixie * 06:47 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1008.eqiad.wmnet with OS trixie * 04:10 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 14 days, 0:00:00 on cp6008.drmrs.wmnet with reason: Hardware failure - [[phab:T431651|T431651]] * 03:55 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp6008.* * 03:29 ryankemper: [[phab:T431311|T431311]] Repooled eqiad cirrussearch clusters (`chi/omega/psi`) following completion of OpenSearch 2.19 migration * 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad * 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=eqiad * 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 31s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-08 == * 23:52 Amir1: ladsgroup@deploy2003:~$ mwscript-k8s --follow -- extensions/ORES/maintenance/PurgeScoreCache.php --wiki=simplewiki --model damaging --old ([[phab:T431159|T431159]]) * 23:46 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 23:46 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing PTR for 2001:df2:e500:fe08::1 - cmooney@cumin1003" * 23:46 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing PTR for 2001:df2:e500:fe08::1 - cmooney@cumin1003" * 23:40 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 23:16 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 23:15 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 22:42 rzl@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 22:40 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] (duration: 12m 55s) * 22:40 rzl@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 22:37 rzl@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 22:36 rzl@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 22:35 rzl@deploy2003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 22:34 urbanecm@deploy2003: urbanecm: Continuing with deployment * 22:33 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:33 rzl@deploy2003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 22:32 rzl@deploy2003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 22:30 rzl@deploy2003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 22:30 rzl@deploy2003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 22:29 rzl@deploy2003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 22:27 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] * 22:26 rzl@deploy2003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 22:22 rzl@deploy2003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 22:21 rzl@deploy2003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 22:19 rzl@deploy2003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 22:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 22:17 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 22:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 22:13 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 22:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 22:13 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 22:09 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 22:06 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 22:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1094.eqiad.wmnet with OS trixie * 22:01 urbanecm: Make https://test.wikipedia.org/w/index.php?title=MediaWiki:GrowthExperimentsSuggestedEdits.json&diff=prev&oldid=750552 with GrowthExperiments disabled (via mw-experimental), then run `\MediaWiki\MediaWikiServices::getInstance()->get('CommunityConfiguration.ProviderFactory')->newProvider('GrowthSuggestedEdits')->getStore()->invalidate()` ([[phab:T431625|T431625]]) * 21:56 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d2-codfw * 21:55 urbanecm@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 21:55 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d2-codfw * 21:55 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c4-codfw * 21:55 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c4-codfw * 21:55 urbanecm@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2002 * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2002 * 21:54 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2002 * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2002.codfw.wmnet 50.32.192.10.in-addr.arpa 0.5.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:54 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2002.codfw.wmnet 50.32.192.10.in-addr.arpa 0.5.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2002 - bking@cumin2003" * 21:54 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2002 - bking@cumin2003" * 21:49 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:49 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2002 * 21:49 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2002.codfw.wmnet with OS trixie * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1094.eqiad.wmnet with reason: host reimage * 21:42 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 21:39 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 21:37 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1094.eqiad.wmnet with reason: host reimage * 21:36 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 21:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 21:29 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 21:27 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 21:22 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1094.eqiad.wmnet with OS trixie * 21:21 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host restbase2039.codfw.wmnet with OS bullseye * 21:21 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin2002" * 21:21 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin2002" * 21:04 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on restbase2039.codfw.wmnet with reason: host reimage * 21:00 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on restbase2039.codfw.wmnet with reason: host reimage * 20:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1073.eqiad.wmnet with OS trixie * 20:48 mutante: deploy2003 - kill 1102 (stunnel4) ; systemctl start stunnel4 ([[phab:T418262|T418262]]) * 20:42 cjming@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] (duration: 33m 02s) * 20:42 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host restbase2039.codfw.wmnet with OS bullseye * 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1073.eqiad.wmnet with reason: host reimage * 20:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1073.eqiad.wmnet with reason: host reimage * 20:30 cjming@deploy2003: cjming: Continuing with deployment * 20:28 cjming@deploy2003: cjming: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1098.eqiad.wmnet with OS trixie * 20:13 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1073.eqiad.wmnet with OS trixie * 20:09 cjming@deploy2003: Started scap sync-world: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] * 20:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1098.eqiad.wmnet with reason: host reimage * 19:56 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1098.eqiad.wmnet with reason: host reimage * 19:55 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d4-codfw * 19:54 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d4-codfw * 19:54 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c1-codfw * 19:54 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c1-codfw * 19:52 mutante: restarting gerrit on gerrit.wikimedia.org (gerrit2003) * 19:48 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2331.codfw.wmnet * 19:48 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2331.codfw.wmnet * 19:48 mutante: restarting gerrit on gerrit-replica.wikimedia.org (gerrit1003) * 19:46 mutante: restarting gerrit on gerrit-spare.wikimedia.org (gerrit2002) * 19:43 jasmine@cumin2002: conftool action : set/pooled=yes; selector: name=wikikube-worker2331.codfw.wmnet,cluster=kubernetes,service=kubesvc * 19:43 jasmine@cumin2002: conftool action : set/weight=10; selector: name=wikikube-worker2331.codfw.wmnet,cluster=kubernetes,service=kubesvc * 19:40 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1098.eqiad.wmnet with OS trixie * 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d5-codfw * 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d5-codfw * 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c7-codfw * 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c7-codfw * 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c5-codfw * 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c5-codfw * 19:30 jasmine_: ran homer on lsw1-d8-codfw, adding wikikube-worker2331 to cluster * 19:29 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1100.eqiad.wmnet with OS trixie * 19:20 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d8-codfw * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d8-codfw * 19:19 mutante: gerrit - replacing private key for registerEmail verification * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-magru * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device cr2-magru * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d7-codfw * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d7-codfw * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d3-codfw * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d1-codfw * 19:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d1-codfw * 19:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c2-codfw * 19:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c2-codfw * 19:11 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-codfw * 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-magru * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device cr1-magru * 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d8-codfw * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d8-codfw * 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d6-codfw * 19:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1100.eqiad.wmnet with reason: host reimage * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d6-codfw * 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c6-codfw * 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c6-codfw * 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c3-codfw * 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c3-codfw * 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b4-magru * 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device asw1-b4-magru * 19:08 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b3-magru * 19:08 cmooney@cumin1003: START - Cookbook sre.network.tls for network device asw1-b3-magru * 19:05 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1100.eqiad.wmnet with reason: host reimage * 19:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1122.eqiad.wmnet with OS trixie * 19:00 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 18:59 topranks: rolling out update to BGP ACL on Nokia Switches eqiad, codfw & ulsfo [[phab:T425703|T425703]] * 18:58 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 18:57 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 18:55 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 18:53 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 18:52 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 18:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1100.eqiad.wmnet with OS trixie * 18:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1068.eqiad.wmnet with OS trixie * 18:47 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1102.eqiad.wmnet with OS trixie * 18:47 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 18:46 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1122.eqiad.wmnet with reason: host reimage * 18:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1122.eqiad.wmnet with reason: host reimage * 18:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1068.eqiad.wmnet with reason: host reimage * 18:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1122 * 18:26 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1122 * 18:25 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1122 * 18:25 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1122.eqiad.wmnet 31.48.64.10.in-addr.arpa 1.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:25 bking@cumin2003: START - Cookbook sre.dns.wipe-cache cirrussearch1122.eqiad.wmnet 31.48.64.10.in-addr.arpa 1.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:25 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:25 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1122 - bking@cumin2003" * 18:25 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1122 - bking@cumin2003" * 18:21 rzl@deploy2003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 18:21 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1068.eqiad.wmnet with reason: host reimage * 18:21 rzl@deploy2003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 18:21 rzl@deploy2003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 18:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 18:19 rzl@deploy2003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 18:19 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:18 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1122 * 18:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1122.eqiad.wmnet with OS trixie * 18:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 18:15 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 18:13 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 18:13 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 18:10 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 18:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1068.eqiad.wmnet with OS trixie * 18:01 kamila@deploy2003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 18m 29s) * 18:00 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:55 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:42 kamila@deploy2003: Started scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] * 17:42 kamila@deploy2003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 19m 50s) * 17:42 kamila@deploy2003: Rolling back deployment * 17:35 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:31 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet * 17:18 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet * 17:16 kamila@deploy1003: Unlocked for deployment [MediaWiki]: switching deployment server (duration: 22m 07s) * 17:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 17:11 kamila@dns1005: END - running authdns-update * 17:09 kamila@dns1005: START - running authdns-update * 17:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 17:04 jasmine@cumin2002: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1164.eqiad.wmnet * 17:04 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1164.eqiad.wmnet * 17:04 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1164.eqiad.wmnet * 16:56 kamila@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on releases2003.codfw.wmnet,releases1003.eqiad.wmnet with reason: Deployment server switchover * 16:54 kamila@deploy1003: Locking from deployment [MediaWiki]: switching deployment server * 16:53 kamila@deploy1003: Unlocked for deployment [MediaWiki]: switching deployment server (duration: 04m 02s) * 16:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie * 16:49 kamila@deploy1003: Locking from deployment [MediaWiki]: switching deployment server * 16:46 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1095.eqiad.wmnet with OS trixie * 16:45 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1093.eqiad.wmnet with OS trixie * 16:43 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1164.eqiad.wmnet with OS trixie * 16:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1095.eqiad.wmnet with reason: host reimage * 16:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 16:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 16:23 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1164.eqiad.wmnet with reason: host reimage * 16:18 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on cirrussearch1093.eqiad.wmnet with reason: host reimage * 16:16 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1164.eqiad.wmnet with reason: host reimage * 16:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1095.eqiad.wmnet with reason: host reimage * 16:09 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 16:09 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 16:08 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1093.eqiad.wmnet with reason: host reimage * 15:59 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Pool test * 15:59 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:59 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 15:59 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Pool test * 15:58 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Depool test * 15:58 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:58 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 15:58 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Depool test * 15:57 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1164 * 15:57 jasmine@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1164 * 15:57 jasmine@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1164 * 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1164.eqiad.wmnet 114.48.64.10.in-addr.arpa 4.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:56 jasmine@cumin2002: START - Cookbook sre.dns.wipe-cache wikikube-worker1164.eqiad.wmnet 114.48.64.10.in-addr.arpa 4.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1164 - jasmine@cumin2002" * 15:56 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1164 - jasmine@cumin2002" * 15:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1093.eqiad.wmnet with OS trixie * 15:51 jasmine@cumin2002: START - Cookbook sre.dns.netbox * 15:51 jasmine@cumin2002: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1164 * 15:50 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-worker1164.eqiad.wmnet with OS trixie * 15:50 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1164.eqiad.wmnet * 15:50 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1164.eqiad.wmnet * 15:50 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1164.eqiad.wmnet * 15:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1095.eqiad.wmnet with OS trixie * 15:42 jasmine@cumin2002: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1164.eqiad.wmnet * 15:42 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1164.eqiad.wmnet * 15:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:42 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1164.eqiad.wmnet * 15:42 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1164.eqiad.wmnet * 15:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 15:39 elukey@cumin1003: START - Cookbook sre.hosts.provision for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 15:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1007.eqiad.wmnet with OS trixie * 15:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Pool test * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1007.eqiad.wmnet with reason: host reimage * 15:15 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 15:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet * 15:15 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 15:15 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1007.eqiad.wmnet with reason: host reimage * 15:15 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:14 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Pool test * 15:14 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet * 15:14 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2228: Depool test * 15:14 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db2228: Depool test * 15:10 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 15:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:08 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 15:06 blake@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 15:06 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet * 15:06 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 15:06 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 15:05 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet * 15:05 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 15:04 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:04 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 15:04 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:03 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 15:03 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 15:03 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:03 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T430909|T430909]] * 15:03 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:03 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 15:03 swfrench-wmf: restarted eqsin, codfw confds - [[phab:T430909|T430909]] * 15:03 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test1002.eqiad.wmnet * 15:02 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet * 14:59 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:59 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:55 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1007.eqiad.wmnet with OS trixie * 14:52 swfrench-wmf: restarted ulsfo confds, confirmed now connected to codfw backends except those using wikimedia.org SRV record - [[phab:T430909|T430909]] * 14:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:49 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:41 moritzm: uninstalling dhcpcd-base from trixie hosts which still have it installed [[phab:T414341|T414341]] * 14:40 sukhe: sudo cumin -b1 -s120 "P<nowiki>{</nowiki>lvs2011*<nowiki>}</nowiki> or P<nowiki>{</nowiki>lvs2012*<nowiki>}</nowiki>" "systemctl restart pybal.service" * 14:39 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:39 mvernon@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host thanos-be1007.eqiad.wmnet with OS trixie * 14:37 sukhe: restart pybal on lvs2013 to revert back to conf2004 * 14:35 sukhe: restart pybal on lvs2014 to revert back to conf2004 * 14:34 swfrench-wmf: switched codfw, eqsin, ulsfo etcd client SRV records back to codfw - [[phab:T430909|T430909]] * 14:32 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1002.eqiad.wmnet * 14:32 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet * 14:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1007.eqiad.wmnet with OS trixie * 14:31 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:31 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:31 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:30 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:30 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Pool test * 14:30 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:29 swfrench@dns1004: END - running authdns-update * 14:29 moritzm: installing jackson-core security updates * 14:27 swfrench@dns1004: START - running authdns-update * 14:22 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:22 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:22 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1119.eqiad.wmnet with OS trixie * 14:22 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:21 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:20 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:20 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 14:20 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:19 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 14:19 moritzm: installing librabbitmq security updates * 14:19 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1002.eqiad.wmnet * 14:18 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet * 14:16 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:15 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:15 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:14 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Pool test * 14:14 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox) * 14:14 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet * 14:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 14:08 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 14:07 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet * 14:05 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1118.eqiad.wmnet with OS trixie * 14:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test1001.eqiad.wmnet * 14:00 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-worker@eqiad * 14:00 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 13:59 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 13:58 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 13:57 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1119.eqiad.wmnet with reason: host reimage * 13:54 moritzm: installing libcap2 security updates * 13:53 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1119.eqiad.wmnet with reason: host reimage * 13:52 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet * 13:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) * 13:52 fceratto@cumin1003: START - Cookbook sre.mysql.depool * 13:50 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-worker@eqiad * 13:50 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1051.eqiad.wmnet * 13:50 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1051.eqiad.wmnet * 13:50 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1051.eqiad.wmnet * 13:49 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 13:45 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1081.eqiad.wmnet with OS trixie * 13:41 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1119 * 13:41 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1119 * 13:40 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1119 * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1119.eqiad.wmnet 97.32.64.10.in-addr.arpa 7.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1119.eqiad.wmnet 97.32.64.10.in-addr.arpa 7.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1119 - atsuko@cumin1003" * 13:40 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1119 - atsuko@cumin1003" * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1118.eqiad.wmnet with reason: host reimage * 13:39 moritzm: installing krb5 security updates * 13:37 Lucas_WMDE: UTC afternoon backport+config window done * 13:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1006.eqiad.wmnet with OS trixie * 13:36 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1118.eqiad.wmnet with reason: host reimage * 13:36 atsuko@cumin1003: START - Cookbook sre.dns.netbox * 13:35 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] (duration: 07m 46s) * 13:34 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1119 * 13:34 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1119.eqiad.wmnet with OS trixie * 13:30 sbisson@deploy1003: sbisson: Continuing with deployment * 13:30 moritzm: installing openssh security updates * 13:30 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling restart_daemons on A:wikidough * 13:29 sbisson@deploy1003: sbisson: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:27 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] * 13:26 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1051.eqiad.wmnet with OS trixie * 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1081.eqiad.wmnet with reason: host reimage * 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1118 * 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1118 * 13:22 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] (duration: 12m 12s) * 13:21 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1081.eqiad.wmnet with reason: host reimage * 13:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1006.eqiad.wmnet with reason: host reimage * 13:18 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1118 * 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1118.eqiad.wmnet 90.32.64.10.in-addr.arpa 0.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:18 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1118.eqiad.wmnet 90.32.64.10.in-addr.arpa 0.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1118 - atsuko@cumin1003" * 13:18 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1118 - atsuko@cumin1003" * 13:17 stran@deploy1003: stran: Continuing with deployment * 13:16 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough * 13:15 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:13 atsuko@cumin1003: START - Cookbook sre.dns.netbox * 13:12 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1118 * 13:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1006.eqiad.wmnet with reason: host reimage * 13:12 stran@deploy1003: stran: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:12 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1118.eqiad.wmnet with OS trixie * 13:10 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] * 13:05 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-worker@codfw * 13:05 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 13:05 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1081.eqiad.wmnet with OS trixie * 13:05 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1051.eqiad.wmnet with reason: host reimage * 13:04 moritzm: installing jq security updates * 13:04 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 13:01 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1051.eqiad.wmnet with reason: host reimage * 12:58 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-worker@codfw * 12:52 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 12:50 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host thanos-be1006.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1051 * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1051 * 12:43 moritzm: installing Python 3.11 security updates * 12:43 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1051 * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1051.eqiad.wmnet 46.32.64.10.in-addr.arpa 6.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:43 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1051.eqiad.wmnet 46.32.64.10.in-addr.arpa 6.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1051 - blake@cumin1003" * 12:43 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1051 - blake@cumin1003" * 12:38 blake@cumin1003: START - Cookbook sre.dns.netbox * 12:38 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1051 * 12:38 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1051.eqiad.wmnet with OS trixie * 12:37 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1051.eqiad.wmnet * 12:36 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1051.eqiad.wmnet * 12:36 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1051.eqiad.wmnet * 12:34 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1006.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 12:34 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1006.eqiad.wmnet with OS trixie * 12:27 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 12:27 mvernon@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host thanos-be1006.eqiad.wmnet with OS trixie * 12:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:02 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 12:01 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1006.eqiad.wmnet with OS trixie * 11:43 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 11:38 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1076.eqiad.wmnet with OS trixie * 11:26 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1075.eqiad.wmnet with OS trixie * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2047.codfw.wmnet * 11:19 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2047.codfw.wmnet * 11:19 moritzm: temporarily remove ganeti2031 from codfw cluster [[phab:T430910|T430910]] * 11:08 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1076.eqiad.wmnet with reason: host reimage * 11:08 moritzm: installing Linux 6.1.176 on Bookworm servers * 11:03 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1076.eqiad.wmnet with reason: host reimage * 11:00 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1075.eqiad.wmnet with reason: host reimage * 10:56 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1075.eqiad.wmnet with reason: host reimage * 10:47 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1076.eqiad.wmnet with OS trixie * 10:46 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1074.eqiad.wmnet with OS trixie * 10:45 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1005.eqiad.wmnet with OS trixie * 10:40 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1075.eqiad.wmnet with OS trixie * 10:32 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2031.codfw.wmnet * 10:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1005.eqiad.wmnet with reason: host reimage * 10:25 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1005.eqiad.wmnet with reason: host reimage * 10:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1074.eqiad.wmnet with reason: host reimage * 10:17 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1074.eqiad.wmnet with reason: host reimage * 10:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1005.eqiad.wmnet with OS trixie * 10:12 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet * 10:04 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 10:01 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet * 10:01 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1074.eqiad.wmnet with OS trixie * 10:01 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 09:43 cgoubert@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/aux-k8s-services/redioscope: apply * 09:43 cgoubert@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/aux-k8s-services/redioscope: apply * 09:43 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply * 09:35 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply * 09:34 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 09:34 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 09:33 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 41 days, 15:00:00 on db2252.codfw.wmnet with reason: Test * 09:32 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 09:32 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 09:31 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: codfw rack B3 pool after maintenance * 09:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 09:07 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 09:07 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 09:02 ladsgroup@cumin1003: END (PASS) - Cookbook sre.mysql.sanitarium_restart (exit_code=0) * 08:57 topranks: merge patch to shift eqiad <-> esams traffic onto new 40G circuit * 08:54 hashar@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1004.eqiad.wmnet with OS trixie * 08:50 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 08:50 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitarium_restart (exit_code=99) * 08:50 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 08:45 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool es2051: codfw rack B3 pool after maintenance * 08:44 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:44 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:43 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2007.codfw.wmnet * 08:43 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2007.codfw.wmnet * 08:42 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2031.codfw.wmnet * 08:41 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2031.codfw.wmnet * 08:40 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2031.codfw.wmnet * 08:38 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:38 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:35 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.sanitize-wiki (exit_code=97) Managing sanitization for wikis minwikiquote in section s3 * 08:33 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis minwikiquote in section s3 * 08:32 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Checking sanitization for wikis minwikiquote in section s5 * 08:30 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Checking sanitization for wikis minwikiquote in section s5 * 08:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Managing sanitization for wikis minwikiquote in section s5 * 08:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1004.eqiad.wmnet with reason: host reimage * 08:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1004.eqiad.wmnet with reason: host reimage * 08:23 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:22 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis minwikiquote in section s5 * 08:19 XioNoX: lsw1-b3-codfw> request system reboot - [[phab:T430909|T430909]] * 08:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Checking sanitization for wikis minwikiquote in section s5 * 08:17 hashar@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Checking sanitization for wikis minwikiquote in section s5 * 08:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for codfw rack B3 * 08:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2007.codfw.wmnet * 08:15 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lsw1-b3-codfw,lsw1-b3-codfw IPv6,lsw1-b3-codfw.mgmt with reason: Switch maintenance * 08:15 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2007.codfw.wmnet * 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:07 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:06 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: codfw rack B3 depool for maintenance * 08:05 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool es2051: codfw rack B3 depool for maintenance * 08:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1004.eqiad.wmnet with OS trixie * 08:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1005.eqiad.wmnet with OS trixie * 08:03 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 21 hosts with reason: codfw rack B3 depool for maintenance * 07:56 ayounsi@cumin1003: START - Cookbook sre.network.depool-rack with action 'depool' for codfw rack B3 * 07:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1005.eqiad.wmnet with reason: host reimage * 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1005.eqiad.wmnet with reason: host reimage * 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1005.eqiad.wmnet with OS bookworm * 07:29 moritzm: installing gnutls28 security updates * 07:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1005.eqiad.wmnet with OS trixie * 07:13 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1125.eqiad.wmnet with OS trixie * 07:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1005.eqiad.wmnet with reason: host reimage * 07:07 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aux-k8s-etcd1005.eqiad.wmnet with reason: host reimage * 06:56 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1005.eqiad.wmnet with OS bookworm * 06:54 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1125.eqiad.wmnet with reason: host reimage * 06:52 elukey: upgrade all trixie hosts to pywmflib 3.1 - [[phab:T430552|T430552]] * 06:50 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1125.eqiad.wmnet with reason: host reimage * 06:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 06:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 06:38 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1125.eqiad.wmnet with OS trixie * 05:42 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1107.eqiad.wmnet with OS trixie * 05:35 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1124.eqiad.wmnet with OS trixie * 05:31 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1101.eqiad.wmnet with OS trixie * 05:21 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1107.eqiad.wmnet with reason: host reimage * 05:17 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1124.eqiad.wmnet with reason: host reimage * 05:13 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1107.eqiad.wmnet with reason: host reimage * 05:13 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1101.eqiad.wmnet with reason: host reimage * 05:11 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1124.eqiad.wmnet with reason: host reimage * 05:10 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1101.eqiad.wmnet with reason: host reimage * 04:58 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1124.eqiad.wmnet with OS trixie * 04:56 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1107.eqiad.wmnet with OS trixie * 04:55 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1101.eqiad.wmnet with OS trixie * 02:27 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] (duration: 08m 14s) * 02:22 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 02:21 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 02:19 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] * 01:59 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] (duration: 09m 46s) * 01:55 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 01:51 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 01:49 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] * 01:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1099.eqiad.wmnet with OS trixie * 00:57 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1110.eqiad.wmnet with OS trixie * 00:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1099.eqiad.wmnet with reason: host reimage * 00:41 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1099.eqiad.wmnet with reason: host reimage * 00:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1110.eqiad.wmnet with reason: host reimage * 00:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1110.eqiad.wmnet with reason: host reimage * 00:26 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1099.eqiad.wmnet with OS trixie * 00:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1110.eqiad.wmnet with OS trixie == 2026-07-07 == * 22:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1097.eqiad.wmnet with OS trixie * 22:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1097.eqiad.wmnet with reason: host reimage * 22:24 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1097.eqiad.wmnet with reason: host reimage * 22:14 hashar: Restarting Gerrit on gerrit2002 and gerrit1003 (replicas) * 22:09 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1097.eqiad.wmnet with OS trixie * 22:07 hashar: Restarting Gerrit on gerrit2003 * 21:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 21:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 21:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 21:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 21:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1108.eqiad.wmnet with OS trixie * 20:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1091.eqiad.wmnet with OS trixie * 20:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1090.eqiad.wmnet with OS trixie * 20:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1108.eqiad.wmnet with reason: host reimage * 20:36 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1091.eqiad.wmnet with reason: host reimage * 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1090.eqiad.wmnet with reason: host reimage * 20:33 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1091.eqiad.wmnet with reason: host reimage * 20:30 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1108.eqiad.wmnet with reason: host reimage * 20:30 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1006.eqiad.wmnet * 20:30 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1090.eqiad.wmnet with reason: host reimage * 20:30 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1006.eqiad.wmnet * 20:27 jasmine_: "homer lsw1-c2-eqiad* commit "Added new stacked control plane wikikube-ctrl1006"" * 20:22 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] (duration: 07m 29s) * 20:20 jasmine_: "homer "cr*eqiad*" commit "Added new stacked control plane wikikube-ctrl1006"" * 20:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1091.eqiad.wmnet with OS trixie * 20:17 arlolra@deploy1003: arlolra: Continuing with deployment * 20:16 arlolra@deploy1003: arlolra: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:16 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1090.eqiad.wmnet with OS trixie * 20:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1108.eqiad.wmnet with OS trixie * 20:14 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] * 20:09 cwhite: remove 2026-04 swift log archives from centrallog2002 to free some space * 20:01 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=93) for host cirrussearch1108.eqiad.wmnet with OS trixie * 19:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1108.eqiad.wmnet with OS trixie * 19:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1090.eqiad.wmnet with OS trixie * 19:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1109.eqiad.wmnet with OS trixie * 19:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1092.eqiad.wmnet with OS trixie * 19:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1123.eqiad.wmnet with OS trixie * 19:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1109.eqiad.wmnet with reason: host reimage * 19:19 cdobbins@cumin2002: conftool action : set/pooled=yes; selector: name=dns7002.* * 19:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1092.eqiad.wmnet with reason: host reimage * 19:17 jasmine@dns1004: END - running authdns-update * 19:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1109.eqiad.wmnet with reason: host reimage * 19:15 jasmine@dns1004: START - running authdns-update * 19:15 cdobbins@dns1004: END - running authdns-update * 19:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1123.eqiad.wmnet with reason: host reimage * 19:13 cdobbins@dns1004: START - running authdns-update * 19:12 cdobbins@cumin2002: conftool action : set/pooled=yes; selector: name=dns7002.*,service=authdns-update * 19:11 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1092.eqiad.wmnet with reason: host reimage * 19:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1123.eqiad.wmnet with reason: host reimage * 18:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1123.eqiad.wmnet with OS trixie * 18:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1109.eqiad.wmnet with OS trixie * 18:56 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1092.eqiad.wmnet with OS trixie * 18:52 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 18:49 swfrench@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 18:40 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 18:38 swfrench@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 18:11 swfrench-wmf: restarted eqsin, codfw confds - [[phab:T430909|T430909]] * 18:01 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T430909|T430909]] * 17:59 swfrench-wmf: restarted ulsfo confds, confirmed now connected to eqiad backends - [[phab:T430909|T430909]] * 17:52 sukhe: restart pybal on lvs2011 to switch from conf2004 to conf1008: [[phab:T430909|T430909]] * 17:51 sukhe: restart pybal on lvs2012 to switch from conf2004 to conf1008 [puppet re-enabled there]: [[phab:T430909|T430909]] * 17:46 sukhe: restart pybal on lvs2013 to switch from conf2004 to conf1008: [[phab:T430909|T430909]] * 17:44 swfrench-wmf: switched codfw, eqsin, ulsfo etcd client SRV records to eqiad - [[phab:T430909|T430909]] * 17:43 swfrench@dns1004: END - running authdns-update * 17:40 swfrench@dns1004: START - running authdns-update * 17:40 sukhe: restart pybal on lvs2014 to switch from conf2004 to conf1008: [[phab:T430909|T430909]] * 17:21 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1003.eqiad.wmnet * 17:15 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1003.eqiad.wmnet * 17:14 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1002.eqiad.wmnet * 17:06 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1002.eqiad.wmnet * 17:06 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-low-traffic-codfw' 'systemctl restart pybal.service' # lvs2013, [[phab:T416623|T416623]] * 17:04 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1001.eqiad.wmnet * 17:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1111.eqiad.wmnet with OS trixie * 17:00 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal.service' # lvs2014, [[phab:T416623|T416623]] * 16:58 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1001.eqiad.wmnet * 16:58 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS bookworm * 16:55 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-low-traffic-eqiad' 'systemctl restart pybal.service' # lvs1019, [[phab:T416623|T416623]] * 16:53 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-secondary-eqiad' 'systemctl restart pybal.service' # lvs1020, [[phab:T416623|T416623]] * 16:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1111.eqiad.wmnet with reason: host reimage * 16:40 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1111.eqiad.wmnet with reason: host reimage * 16:38 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.peering (exit_code=99) with action 'configure' for AS: 47794 * 16:35 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 47794 * 16:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1111.eqiad.wmnet with OS trixie * 16:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1006.eqiad.wmnet with OS trixie * 16:06 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 16:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1006.eqiad.wmnet with reason: host reimage * 15:58 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1121.eqiad.wmnet with OS trixie * 15:58 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 15:56 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1006.eqiad.wmnet with reason: host reimage * 15:54 mutante: jenkins down in planned maintenance window * 15:42 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1037.eqiad.wmnet * 15:42 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1037.eqiad.wmnet * 15:42 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1037.eqiad.wmnet * 15:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1006.eqiad.wmnet with OS trixie * 15:34 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1121.eqiad.wmnet with reason: host reimage * 15:33 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS bookworm * 15:33 cdobbins@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host dns7002.wikimedia.org with OS trixie * 15:30 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1121.eqiad.wmnet with reason: host reimage * 15:29 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host clouddumps1001.wikimedia.org * 15:20 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1001.wikimedia.org * 15:18 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1121 * 15:18 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1121 * 15:18 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host clouddumps1002.wikimedia.org * 15:17 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1121 * 15:17 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:17 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply * 15:16 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply * 15:16 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:16 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1121 - atsuko@cumin1003" * 15:16 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1121 - atsuko@cumin1003" * 15:14 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1037.eqiad.wmnet with OS trixie * 15:11 atsuko@cumin1003: START - Cookbook sre.dns.netbox * 15:09 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org * 15:09 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1121 * 15:09 andrew@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host clouddumps1002.wikimedia.org * 15:09 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org * 15:09 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1121.eqiad.wmnet with OS trixie * 15:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1007.eqiad.wmnet with OS trixie * 15:08 andrew@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host clouddumps1002.wikimedia.org * 15:08 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org * 15:05 brennen@deploy1003: Finished deploy [phabricator/deployment@7e02037]: deploy phab1004 for [[phab:T431440|T431440]] (duration: 00m 47s) * 15:04 brennen@deploy1003: Started deploy [phabricator/deployment@7e02037]: deploy phab1004 for [[phab:T431440|T431440]] * 15:03 brennen@deploy1003: Finished deploy [phabricator/deployment@7e02037]: deploy phab2003 for [[phab:T431440|T431440]] (duration: 00m 51s) * 15:03 brennen@deploy1003: Started deploy [phabricator/deployment@7e02037]: deploy phab2003 for [[phab:T431440|T431440]] * 15:00 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71] (thin): Regular analytics weekly train THIN [analytics/refinery@7d8dc71f] (duration: 02m 10s) * 14:58 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71] (thin): Regular analytics weekly train THIN [analytics/refinery@7d8dc71f] * 14:58 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71]: Regular analytics weekly train [analytics/refinery@7d8dc71f] (duration: 04m 14s) * 14:54 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1037.eqiad.wmnet with reason: host reimage * 14:53 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71]: Regular analytics weekly train [analytics/refinery@7d8dc71f] * 14:53 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@7d8dc71f] (duration: 02m 00s) * 14:51 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@7d8dc71f] * 14:51 arnaudb@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on phab2003.codfw.wmnet,phab[1004-1006].eqiad.wmnet with reason: maintenance * 14:51 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1037.eqiad.wmnet with reason: host reimage * 14:50 JavierMonton: Deploying Refinery at {{Gerrit|7d8dc71f}} for change {{Gerrit|1308087}} / [[phab:T431318|T431318]] - update filerevision table sqoop and table * 14:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1007.eqiad.wmnet with reason: host reimage * 14:42 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1007.eqiad.wmnet with reason: host reimage * 14:40 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1083.eqiad.wmnet with OS trixie * 14:37 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:36 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:35 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-master@eqiad * 14:35 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 14:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 14:34 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1037 * 14:34 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1037 * 14:34 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:cleanMentorList.php --wiki=frwiki # [[phab:T427386|T427386]] * 14:34 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 14:34 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308112{{!}}Revert^2 "[Growth] frwiki: Deploy automated mentor list cleaner" (T427386)]] (duration: 06m 47s) * 14:34 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 14:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:33 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:32 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1037 * 14:31 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:31 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:29 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:29 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-master@eqiad * 14:29 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:29 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:29 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:28 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:27 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:27 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1308112{{!}}Revert^2 "[Growth] frwiki: Deploy automated mentor list cleaner" (T427386)]] * 14:26 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1007.eqiad.wmnet with OS trixie * 14:26 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:26 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 14:26 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:25 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:25 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:25 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:cleanMentorList.php --wiki=frwiki # [[phab:T427386|T427386]] * 14:24 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:24 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1037 - blake@cumin1003" * 14:24 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1037 - blake@cumin1003" * 14:20 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1083.eqiad.wmnet with reason: host reimage * 14:19 blake@cumin1003: START - Cookbook sre.dns.netbox * 14:19 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-master@codfw * 14:19 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 14:19 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1037 * 14:18 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1037.eqiad.wmnet with OS trixie * 14:18 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1037.eqiad.wmnet * 14:18 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 14:18 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1037.eqiad.wmnet * 14:18 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1037.eqiad.wmnet * 14:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2007.codfw.wmnet with OS trixie * 14:16 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1083.eqiad.wmnet with reason: host reimage * 14:15 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1036.eqiad.wmnet * 14:15 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1036.eqiad.wmnet * 14:14 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1036.eqiad.wmnet * 14:12 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-master@codfw * 14:11 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1120.eqiad.wmnet with OS trixie * 14:05 moritzm: installing distro-info-data updates from trixie/bookworm point releases * 14:04 fabfur: disable puppet on A:cp-text to selectively apply https://gerrit.wikimedia.org/r/c/operations/puppet/+/1308040 * 14:03 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] (duration: 27m 48s) * 14:00 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 14:00 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1083.eqiad.wmnet with OS trixie * 13:58 urbanecm@deploy1003: urbanecm: Continuing with deployment * 13:58 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:58 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1004.eqiad.wmnet with OS bookworm * 13:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2007.codfw.wmnet with reason: host reimage * 13:57 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 13:53 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1120.eqiad.wmnet with reason: host reimage * 13:50 moritzm: installing Linux 5.10.259 on Bullseye hosts * 13:47 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply * 13:47 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply * 13:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2007.codfw.wmnet with reason: host reimage * 13:46 cgoubert@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/aux-k8s-services/redioscope: apply * 13:46 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1120.eqiad.wmnet with reason: host reimage * 13:46 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:46 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:45 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:44 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:44 cgoubert@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/aux-k8s-services/redioscope: apply * 13:40 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 13:39 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:38 moritzm: installing e2fsprogs updates from Trixie point release * 13:35 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] * 13:33 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1120.eqiad.wmnet with OS trixie * 13:33 topranks: reset cr3-eqsin configuration so traffic uses it again after upgrade * 13:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1088.eqiad.wmnet with OS trixie * 13:32 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie * 13:32 cdobbins@cumin1003: conftool action : set/pooled=no; selector: name=dns7002.* * 13:29 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2007.codfw.wmnet with OS trixie * 13:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1004.eqiad.wmnet with reason: host reimage * 13:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2006.codfw.wmnet with OS trixie * 13:18 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1036.eqiad.wmnet with OS trixie * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aux-k8s-etcd1004.eqiad.wmnet with reason: host reimage * 13:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 13:16 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 13:15 jayme: Istio is being upgraded from 1.24.2 to 1.29.4 on wikikube staging eqiad and codfw - [[phab:T427401|T427401]] * 13:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1087.eqiad.wmnet with OS trixie * 13:14 topranks: reboot cr3-eqsin to install new JunOS and set PIC 0/0/0 to 100G * 13:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1088.eqiad.wmnet with reason: host reimage * 13:13 jmm@dns1004: END - running authdns-update * 13:12 jmm@dns1004: START - running authdns-update * 13:09 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1088.eqiad.wmnet with reason: host reimage * 13:07 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1082.eqiad.wmnet with OS trixie * 13:07 jmm@dns1004: END - running authdns-update * 13:06 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1004.eqiad.wmnet with OS bookworm * 13:05 jmm@dns1004: START - running authdns-update * 13:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2006.codfw.wmnet with reason: host reimage * 12:58 topranks: load updated JunOS on cr3-eqsin [[phab:T429386|T429386]] * 12:58 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1036.eqiad.wmnet with reason: host reimage * 12:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2001.codfw.wmnet * 12:57 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2006.codfw.wmnet with reason: host reimage * 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr1-codfw,cr[2-3]-eqsin,cr3-eqsin IPv6,cr3-eqsin.mgmt with reason: upgrade JunOS cr3-eqsin * 12:56 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lvs[5004-5006].eqsin.wmnet with reason: upgrade JunOS cr3-eqsin * 12:55 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 12:55 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 12:53 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1087.eqiad.wmnet with reason: host reimage * 12:53 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 12:52 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1088.eqiad.wmnet with OS trixie * 12:52 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1002.eqiad.wmnet * 12:52 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:51 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2001.codfw.wmnet * 12:49 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1036.eqiad.wmnet with reason: host reimage * 12:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 12:48 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1087.eqiad.wmnet with reason: host reimage * 12:44 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1082.eqiad.wmnet with reason: host reimage * 12:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1002.eqiad.wmnet * 12:42 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 12:42 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 12:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 12:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 12:39 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:39 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: move dumps-nfs IP to the shared one - filippo@cumin1003" * 12:39 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: move dumps-nfs IP to the shared one - filippo@cumin1003" * 12:39 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2006.codfw.wmnet with OS trixie * 12:38 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1082.eqiad.wmnet with reason: host reimage * 12:36 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:33 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:32 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1036 * 12:32 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1036 * 12:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2005.codfw.wmnet with OS trixie * 12:32 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1087.eqiad.wmnet with OS trixie * 12:30 jmm@dns1004: END - running authdns-update * 12:29 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1036 * 12:29 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1036.eqiad.wmnet 21.32.64.10.in-addr.arpa 1.2.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:29 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1036.eqiad.wmnet 21.32.64.10.in-addr.arpa 1.2.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:29 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:29 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1036 - blake@cumin1003" * 12:29 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1036 - blake@cumin1003" * 12:28 jmm@dns1004: START - running authdns-update * 12:26 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:26 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:23 blake@cumin1003: START - Cookbook sre.dns.netbox * 12:23 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1036 * 12:23 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1036.eqiad.wmnet with OS trixie * 12:22 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1036.eqiad.wmnet * 12:22 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1082.eqiad.wmnet with OS trixie * 12:22 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1036.eqiad.wmnet * 12:22 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1036.eqiad.wmnet * 12:21 marostegui: Restart mariadb@s7 on db1155 to pick up new filters - [[phab:T431124|T431124]] * 12:21 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 21 hosts with reason: restarting for replication filter * 12:20 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:19 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:14 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2005.codfw.wmnet with reason: host reimage * 12:14 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:08 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:07 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:07 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2005.codfw.wmnet with reason: host reimage * 12:06 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:06 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:06 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:05 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:05 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:04 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-master-eqiad * 12:04 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl1002.eqiad.wmnet * 12:04 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl1002.eqiad.wmnet * 12:04 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:04 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:03 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:03 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:03 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:03 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 11:59 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl1002.eqiad.wmnet * 11:59 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl1002.eqiad.wmnet * 11:59 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl1001.eqiad.wmnet * 11:59 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl1001.eqiad.wmnet * 11:56 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl1001.eqiad.wmnet * 11:56 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl1001.eqiad.wmnet * 11:56 klausman@cumin2002: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-master-eqiad * 11:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2005.codfw.wmnet with OS trixie * 11:49 blake@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on wikikube-worker1160.eqiad.wmnet with reason: Verifying matchers for silence * 11:42 topranks: cr3-eqsin, begin traffic drain to reset PIC and upgrade JunOS [[phab:T429386|T429386]] * 11:41 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs[5004-5006].eqsin.wmnet with reason: upgrade JunOS cr3-eqsin * 11:39 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr1-codfw,cr[2-3]-eqsin,cr3-eqsin IPv6,cr3-eqsin.mgmt with reason: upgrade JunOS cr3-eqsin * 11:36 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=thanos-fe2004.codfw.wmnet * 11:35 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1086.eqiad.wmnet with OS trixie * 11:35 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=thanos-fe2004.codfw.wmnet * 11:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1085.eqiad.wmnet with OS trixie * 11:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1086.eqiad.wmnet with reason: host reimage * 11:10 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1085.eqiad.wmnet with reason: host reimage * 11:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2004.codfw.wmnet with OS trixie * 11:03 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1086.eqiad.wmnet with reason: host reimage * 11:02 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1085.eqiad.wmnet with reason: host reimage * 10:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2004.codfw.wmnet with reason: host reimage * 10:48 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:46 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1086.eqiad.wmnet with OS trixie * 10:46 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1085.eqiad.wmnet with OS trixie * 10:44 cgoubert@deploy1003: Finished deploy [restbase/deploy@2fc37d4]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] (duration: 16m 44s) * 10:43 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2004.codfw.wmnet with reason: host reimage * 10:35 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:27 cgoubert@deploy1003: Started deploy [restbase/deploy@2fc37d4]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] * 10:27 cgoubert@deploy1003: Finished deploy [restbase/deploy@8a25036]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] (duration: 00m 45s) * 10:26 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1117.eqiad.wmnet with OS trixie * 10:26 cgoubert@deploy1003: Started deploy [restbase/deploy@8a25036]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] * 10:26 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host thanos-fe2004 * 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host thanos-fe2004 * 10:22 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1116.eqiad.wmnet with OS trixie * 10:21 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host thanos-fe2004 * 10:21 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) thanos-fe2004.codfw.wmnet 157.32.192.10.in-addr.arpa 7.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:20 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache thanos-fe2004.codfw.wmnet 157.32.192.10.in-addr.arpa 7.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:20 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:20 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host thanos-fe2004 - mvernon@cumin2003" * 10:20 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host thanos-fe2004 - mvernon@cumin2003" * 10:15 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2252: Repooling after reboot * 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:15 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 10:15 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2252: Repooling after reboot * 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1153.eqiad.wmnet * 10:14 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1153.eqiad.wmnet * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 10:14 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 10:12 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 10:12 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host thanos-fe2004 * 10:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2004.codfw.wmnet with OS trixie * 10:07 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1117.eqiad.wmnet with reason: host reimage * 10:03 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1116.eqiad.wmnet with reason: host reimage * 09:58 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:58 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1117.eqiad.wmnet with reason: host reimage * 09:57 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1116.eqiad.wmnet with reason: host reimage * 09:49 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 41 days, 15:00:00 on db2252.codfw.wmnet with reason: Security updates * 09:45 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1117.eqiad.wmnet with OS trixie * 09:45 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1116.eqiad.wmnet with OS trixie * 09:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1153: Security updates * 09:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:28 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:28 root@cumin1003: START - Cookbook sre.mysql.depool depool db1153: Security updates * 09:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1016: Security updates * 09:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:21 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:21 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1016: Security updates * 09:14 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:14 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 08:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1016: Security updates * 08:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:56 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:56 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1016: Security updates * 08:50 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:50 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:45 filippo@dns1006: END - running authdns-update * 08:43 filippo@dns1006: START - running authdns-update * 08:42 godog: switch dumps-nfs address to be shared with rsync/http - [[phab:T411248|T411248]] * 08:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1016: Security updates * 08:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:40 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:40 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1016: Security updates * 08:29 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host cirrussearch1111.eqiad.wmnet * 08:29 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:27 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:27 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:25 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1015: Security updates * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:09 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:09 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1015: Security updates * 07:42 Msz2001: Deployed private patch for Suggested Ivestigations * 07:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1015: Security updates * 07:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:41 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:41 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1015: Security updates * 07:40 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 07:11 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fingerprint warnings - oblivian@cumin1003" * 07:11 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fingerprint warnings - oblivian@cumin1003 * 07:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1024: Security updates * 07:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:11 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:11 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1024: Security updates * 07:10 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fingerprint warnings - oblivian@cumin1003 * 07:10 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fingerprint warnings - oblivian@cumin1003" * 06:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host cirrussearch1111.eqiad.wmnet * 06:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 06:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1024: Security updates * 06:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 06:48 root@cumin1003: START - Cookbook sre.mysql.parsercache * 06:48 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1024: Security updates * 06:42 moritzm: install nginx security updates * 06:31 root@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool pc1024: Security updates * 06:21 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1024: Security updates * 06:19 moritzm: installing php8.2 security updates * 06:15 moritzm: installing php8.4 security updates * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.7 (duration: 02m 38s) * 03:40 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] (duration: 37m 04s) * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 51s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-06 == * 23:30 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] (duration: 09m 39s) * 23:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1078.eqiad.wmnet with OS trixie * 23:26 jdlrobson@deploy1003: jdlrobson, bwang: Continuing with deployment * 23:22 jdlrobson@deploy1003: jdlrobson, bwang: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug) * 23:21 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] * 23:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1078.eqiad.wmnet with reason: host reimage * 23:06 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1078.eqiad.wmnet with reason: host reimage * 22:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1078.eqiad.wmnet with OS trixie * 22:29 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on cirrussearch1114.eqiad.wmnet with reason: reimage on hold until restore completes * 22:22 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on cirrussearch[1079,1115].eqiad.wmnet with reason: reimage on hold until restore completes * 21:18 maryum: Deployed security fix for [[phab:T428006|T428006]] * 20:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1079.eqiad.wmnet with OS trixie * 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1077.eqiad.wmnet with OS trixie * 20:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1115.eqiad.wmnet with OS trixie * 20:25 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1079.eqiad.wmnet with reason: host reimage * 20:21 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1079.eqiad.wmnet with reason: host reimage * 20:15 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] (duration: 08m 14s) * 20:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1077.eqiad.wmnet with reason: host reimage * 20:10 krinkle@deploy1003: krinkle, pushpaktiwari: Continuing with deployment * 20:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1115.eqiad.wmnet with reason: host reimage * 20:08 krinkle@deploy1003: krinkle, pushpaktiwari: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1077.eqiad.wmnet with reason: host reimage * 20:06 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] * 20:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1079.eqiad.wmnet with OS trixie * 20:04 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1115.eqiad.wmnet with reason: host reimage * 19:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1077.eqiad.wmnet with OS trixie * 19:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1115.eqiad.wmnet with OS trixie * 19:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 19:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 18:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1114.eqiad.wmnet with OS trixie * 18:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1114.eqiad.wmnet with reason: host reimage * 18:35 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1114.eqiad.wmnet with reason: host reimage * 18:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1112.eqiad.wmnet with OS trixie * 18:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1114.eqiad.wmnet with OS trixie * 18:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1072.eqiad.wmnet with OS trixie * 18:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1112.eqiad.wmnet with reason: host reimage * 18:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1112.eqiad.wmnet with reason: host reimage * 17:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1072.eqiad.wmnet with reason: host reimage * 17:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1112.eqiad.wmnet with OS trixie * 17:55 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1072.eqiad.wmnet with reason: host reimage * 17:39 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1072.eqiad.wmnet with OS trixie * 17:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1071.eqiad.wmnet with OS trixie * 17:18 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1070.eqiad.wmnet with OS trixie * 17:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1084.eqiad.wmnet with OS trixie * 16:54 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1071.eqiad.wmnet with reason: host reimage * 16:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1084.eqiad.wmnet with reason: host reimage * 16:51 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1070.eqiad.wmnet with reason: host reimage * 16:49 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1084.eqiad.wmnet with reason: host reimage * 16:38 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1071.eqiad.wmnet with OS trixie * 16:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1096.eqiad.wmnet with OS trixie * 16:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1070.eqiad.wmnet with OS trixie * 16:33 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1084.eqiad.wmnet with OS trixie * 16:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1089.eqiad.wmnet with OS trixie * 16:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1103.eqiad.wmnet with OS trixie * 16:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1096.eqiad.wmnet with reason: host reimage * 16:14 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1096.eqiad.wmnet with reason: host reimage * 16:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1089.eqiad.wmnet with reason: host reimage * 16:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1103.eqiad.wmnet with reason: host reimage * 16:02 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1003.eqiad.wmnet with OS bookworm * 16:01 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1089.eqiad.wmnet with reason: host reimage * 16:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1103.eqiad.wmnet with reason: host reimage * 15:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1096.eqiad.wmnet with OS trixie * 15:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1080.eqiad.wmnet with OS trixie * 15:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1089.eqiad.wmnet with OS trixie * 15:45 dancy@deploy1003: Installation of scap version "4.272.0" completed for 158 hosts * 15:43 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1103.eqiad.wmnet with OS trixie * 15:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1113.eqiad.wmnet with OS trixie * 15:41 dancy@deploy1003: Installing scap version "4.272.0" for 158 host(s) * 15:40 klausman@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 15:39 klausman@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 15:38 klausman@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 15:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1069.eqiad.wmnet with OS trixie * 15:37 klausman@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 15:36 klausman@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 15:34 klausman@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 15:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1080.eqiad.wmnet with reason: host reimage * 15:27 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1080.eqiad.wmnet with reason: host reimage * 15:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1113.eqiad.wmnet with reason: host reimage * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1069.eqiad.wmnet with reason: host reimage * 15:18 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1113.eqiad.wmnet with reason: host reimage * 15:16 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1069.eqiad.wmnet with reason: host reimage * 15:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:11 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1080.eqiad.wmnet with OS trixie * 15:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1113.eqiad.wmnet with OS trixie * 15:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1003.eqiad.wmnet with reason: host reimage * 14:47 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1003.eqiad.wmnet with OS bookworm * 14:33 elukey: rolled out spicerack on all cumin nodes - [[phab:T429699|T429699]] * 14:32 elukey: upgrade all bookworm hosts to pywmflib 3.1 - [[phab:T430552|T430552]] * 14:14 marostegui: Setup x4 eqiad topology [[phab:T404715|T404715]] * 14:13 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 14:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2230.codfw.wmnet * 14:07 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2230.codfw.wmnet * 13:59 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[2001-2002].codfw.wmnet * 13:51 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 13:45 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.major-upgrade (exit_code=97) * 13:45 cwilliams@cumin1003: dbmaint on s4@codfw [[phab:T429893|T429893]] * 13:45 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 13:42 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-master-codfw * 13:42 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl2002.codfw.wmnet * 13:42 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl2002.codfw.wmnet * 13:38 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl2002.codfw.wmnet * 13:38 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl2002.codfw.wmnet * 13:38 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl2001.codfw.wmnet * 13:38 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl2001.codfw.wmnet * 13:35 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl2001.codfw.wmnet * 13:35 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl2001.codfw.wmnet * 13:35 klausman@cumin2002: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-master-codfw * 12:30 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] (duration: 25m 11s) * 12:24 krinkle@deploy1003: krinkle: Continuing with deployment * 12:10 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2048.codfw.wmnet * 12:09 krinkle@deploy1003: krinkle: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:08 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2048.codfw.wmnet * 12:05 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] * 11:57 moritzm: installing curl security updates * 11:49 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:31 moritzm: installing nano security updates * 11:07 moritzm: failover Ganeti master in codfw to ganeti2032 [[phab:T430909|T430909]] * 11:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:04 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest1005.eqiad.wmnet with OS trixie * 11:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:50 jmm@dns1004: END - running authdns-update * 10:47 jmm@dns1004: START - running authdns-update * 10:47 jmm@dns1004: START - running authdns-update * 10:46 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:44 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest1005.eqiad.wmnet with reason: host reimage * 10:38 elukey@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest1005.eqiad.wmnet with reason: host reimage * 10:31 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:31 marostegui: Setup x4 codfw topology [[phab:T404715|T404715]] * 10:31 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 10:24 elukey: spicerack 13.0.0 deployed on cumin2002 * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 10:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 10:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 10:21 elukey@cumin2002: START - Cookbook sre.hosts.reimage for host sretest1005.eqiad.wmnet with OS trixie * 10:20 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:19 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:17 elukey: uploaded spicerack_13.0.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia * 09:54 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:52 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:20 elukey: upgrade all bullseye hosts to pywmflib 3.1 - [[phab:T430552|T430552]] * 09:10 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1015.eqiad.wmnet,service=s4 * 09:10 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1015.eqiad.wmnet,service=s6 * 09:07 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 08:58 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:56 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 08:56 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 08:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 08:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 08:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin2002.codfw.wmnet * 08:06 godog: remove cloudvirt1046, cloudvirt1062, cloudvirt1074, cloudvirt1075 from maintenance aggregate and put them in network-ovs - [[phab:T424802|T424802]] * 08:00 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin2002.codfw.wmnet * 07:58 hashar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] (duration: 32m 53s) * 07:57 fabfur: repooled cp4038 * 07:57 fabfur@cumin1003: conftool action : set/pooled=yes; selector: name=cp4038.* * 07:53 moritzm: installing pyjwt security updates * 07:47 moritzm: installing openjpeg2 security updates * 07:45 hashar@deploy1003: vadymts1, hashar: Continuing with deployment * 07:43 hashar@deploy1003: vadymts1, hashar: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:38 moritzm: installing python-urllib3 security updates * 07:37 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 07:30 fabfur: depooled cp4038 to investigate on possible maxmind failure * 07:30 fabfur@cumin1003: conftool action : set/pooled=no; selector: name=cp4038.* * 07:30 fabfur@cumin1003: conftool action : set/pooled=yes; selector: name=cp4038.* * 07:29 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 07:25 hashar@deploy1003: Started scap sync-world: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] * 06:13 moritzm: installing Linux 6.12.95 on trixie hosts * 05:20 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s6 * 05:20 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s4 * 05:19 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1015.eqiad.wmnet with reason: cloning * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 08s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-05 == * 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 01m 08s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-04 == * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 58s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-03 == * 17:08 topranks: revert protocol preference changes on cr3-ulsfo after upgrade * 16:53 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on cr2-eqord with reason: upgrade JunOS cr3-ulsfo * 16:53 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on cr4-ulsfo with reason: upgrade JunOS cr3-ulsfo * 16:48 topranks: reboot cr3-ulsfo to upgrade JunOS and reset linecard [[phab:T424839|T424839]] * 15:52 topranks: adjust outbound BGP policies on cr3-ulsfo to drain router of traffic [[phab:T424839|T424839]] * 15:45 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on lvs[4008-4010].ulsfo.wmnet with reason: upgrade JunOS cr3-ulsfo * 15:44 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on asw1-[22-23]-ulsfo,cr3-ulsfo,cr3-ulsfo IPv6,cr3-ulsfo.mgmt with reason: upgrade JunOS cr3-ulsfo * 15:36 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 15:35 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 15:35 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 14:40 cmooney@dns3003: END - running authdns-update * 14:26 cmooney@dns3003: START - running authdns-update * 14:26 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:26 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to ulsfo - cmooney@cumin1003" * 14:19 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to ulsfo - cmooney@cumin1003" * 14:16 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:38 sukhe@dns1004: END - running authdns-update * 13:35 sukhe@dns1004: START - running authdns-update * 13:26 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 13:26 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 13:26 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet * 13:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 13:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 13:16 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host sretest1005.eqiad.wmnet * 13:16 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 13:16 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 13:15 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 13:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:14 moritzm: imported samplicator 1.3.8rc1-1+deb13u1 to trixie-wikimedia/main [[phab:T337208|T337208]] * 13:13 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:07 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 13:07 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 13:02 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:02 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:58 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:57 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:57 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:53 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet * 12:50 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 12:47 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:41 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:40 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:39 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:32 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet * 12:26 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet * 12:23 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2005.wikimedia.org * 12:19 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2005.wikimedia.org * 12:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 12:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup[2004-2007].codfw.wmnet * 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[2004-2007].codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin2003" * 12:15 jynus@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[2004-2007].codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin2003" * 12:09 jynus@cumin2003: START - Cookbook sre.dns.netbox * 11:58 jynus@cumin2003: START - Cookbook sre.hosts.decommission for hosts backup[2004-2007].codfw.wmnet * 10:40 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup[1004-1007].eqiad.wmnet * 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[1004-1007].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 10:01 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[1004-1007].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 09:52 jynus@cumin1003: START - Cookbook sre.dns.netbox * 09:39 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:36 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup[1004-1007].eqiad.wmnet * 09:36 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:25 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 09:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 09:16 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 09:05 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:04 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:00 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:59 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:57 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:55 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:50 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 08:50 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 08:49 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 08:49 atsukoito: depooling cirrussearch in codfw because of regression after upgrade [[phab:T431091|T431091]] * 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts mirror1001.wikimedia.org * 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: mirror1001.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 08:29 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: mirror1001.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 08:18 jmm@cumin2003: START - Cookbook sre.dns.netbox * 08:11 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts mirror1001.wikimedia.org * 06:15 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 18s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-02 == * 22:55 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host contint1003.wikimedia.org with OS trixie * 22:29 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on contint1003.wikimedia.org with reason: host reimage * 22:23 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on contint1003.wikimedia.org with reason: host reimage * 22:05 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host contint1003.wikimedia.org with OS trixie * 22:03 mutante: contint1003 (zuul.wikimedia.org) - reimaging because of [[phab:T430510|T430510]]#12067628 [[phab:T418521|T418521]] * 22:03 dzahn@cumin2002: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on zuul.wikimedia.org with reason: reimage * 21:39 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 18s) * 21:39 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 21:20 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1003.eqiad.wmnet, repooling source-only afterwards * 21:19 sbassett: Deployed security fix for [[phab:T428829|T428829]] * 20:58 cmooney@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Release v0.11.2 update for new Aerleon - cmooney@cumin1003 * 20:55 cmooney@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Release v0.11.2 update for new Aerleon - cmooney@cumin1003 * 20:40 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] (duration: 12m 35s) * 20:36 arlolra@deploy1003: cscott, arlolra: Continuing with deployment * 20:35 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 20s) * 20:35 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 20:33 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host contint2003.wikimedia.org with OS trixie * 20:31 arlolra@deploy1003: cscott, arlolra: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Cha * 20:28 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] * 20:17 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] (duration: 08m 13s) * 20:14 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on contint2003.wikimedia.org with reason: host reimage * 20:13 sbassett@deploy1003: sbassett: Continuing with deployment * 20:11 sbassett@deploy1003: sbassett: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:09 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] * 20:08 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 20:08 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 20:08 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on contint2003.wikimedia.org with reason: host reimage * 20:05 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1003.eqiad.wmnet, repooling source-only afterwards * 19:49 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host contint2003.wikimedia.org with OS trixie * 19:48 mutante: contint2003 - reimaging because of [[phab:T430510|T430510]]#12067628 [[phab:T418521|T418521]] * 18:39 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 18:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 18:13 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2002.codfw.wmnet -> wcqs2003.codfw.wmnet, repooling source-only afterwards * 17:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1003.eqiad.wmnet with OS bookworm * 17:52 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1005.eqiad.wmnet * 17:52 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1005.eqiad.wmnet * 17:51 jasmine@cumin2002: conftool action : set/pooled=yes:weight=10; selector: name=wikikube-ctrl1005.eqiad.wmnet * 17:48 jasmine_: homer "cr*eqiad*" commit "Added new stacked control plane wikikube-ctrl1005" * 17:44 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply * 17:44 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply * 17:31 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] (duration: 09m 33s) * 17:26 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 17:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1003.eqiad.wmnet with reason: host reimage * 17:23 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:21 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] * 17:18 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1003.eqiad.wmnet with reason: host reimage * 17:16 rscout@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply * 17:16 rscout@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply * 17:16 rscout@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply * 17:15 rscout@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply * 17:12 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on wcqs[2002-2003].codfw.wmnet,wcqs1002.eqiad.wmnet with reason: reimaging hosts * 17:08 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 17:08 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 17:08 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 17:07 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 17:05 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 17:05 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 17:03 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "running to make sure all updates are synced - cmooney@cumin1003" * 17:03 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "running to make sure all updates are synced - cmooney@cumin1003" * 17:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs1003 * 17:00 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs1003 * 17:00 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 17:00 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1003.eqiad.wmnet with OS bookworm * 16:58 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Re-running - btullis@cumin1003" * 16:58 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Re-running - btullis@cumin1003" * 16:58 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2002.codfw.wmnet -> wcqs2003.codfw.wmnet, repooling source-only afterwards * 16:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-master1004.eqiad.wmnet with OS bookworm * 16:58 btullis@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 16:57 tappof: bump space for prometheus k8s-aux in eqiad * 16:55 cmooney@dns3003: END - running authdns-update * 16:55 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:55 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to eqsin - cmooney@cumin1003" * 16:55 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to eqsin - cmooney@cumin1003" * 16:53 cmooney@dns3003: START - running authdns-update * 16:52 ryankemper: [ml-serve-eqiad] Cleared out 1302 failed (Evicted) pods: `kubectl -n llm delete pods --field-selector=status.phase=Failed`, freeing calico-kube-controllers from OOM crashloop (evictions were caused by disk pressure) * 16:49 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 16:46 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:39 rzl@dns1004: END - running authdns-update * 16:37 rzl@dns1004: START - running authdns-update * 16:36 rzl@dns1004: START - running authdns-update * 16:35 rzl@deploy1003: Finished scap sync-world: [[phab:T416623|T416623]] (duration: 10m 19s) * 16:34 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 16:33 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-master1004.eqiad.wmnet with reason: host reimage * 16:30 rzl@deploy1003: rzl: Continuing with deployment * 16:28 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-master1004.eqiad.wmnet with reason: host reimage * 16:26 rzl@deploy1003: rzl: [[phab:T416623|T416623]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:25 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 16:25 rzl@deploy1003: Started scap sync-world: [[phab:T416623|T416623]] * 16:25 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 16:24 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 16:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: sync * 16:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: sync * 16:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-master1004.eqiad.wmnet with OS bookworm * 16:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-master1003.eqiad.wmnet with OS bookworm * 16:11 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 16:11 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 16:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Security updates * 16:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 16:08 root@cumin1003: START - Cookbook sre.mysql.parsercache * 16:08 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Security updates * 15:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-master1003.eqiad.wmnet with reason: host reimage * 15:54 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:54 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:54 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:54 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-master1003.eqiad.wmnet with reason: host reimage * 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Security updates * 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:45 root@cumin1003: START - Cookbook sre.mysql.parsercache * 15:45 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Security updates * 15:42 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-master1003.eqiad.wmnet with OS bookworm * 15:24 moritzm: installing busybox updates from bookworm point release * 15:20 moritzm: installing busybox updates from trixie point release * 15:15 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1021: Security updates * 15:15 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:15 root@cumin1003: START - Cookbook sre.mysql.parsercache * 15:15 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1021: Security updates * 15:13 moritzm: installing giflib security updates * 15:08 moritzm: installing Tomcat security updates * 14:57 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 14:56 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 14:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:53 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Unblock taavi - oblivian@cumin1003" * 14:53 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Unblock taavi - oblivian@cumin1003 * 14:53 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1021: Security updates * 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:53 root@cumin1003: START - Cookbook sre.mysql.parsercache * 14:53 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1021: Security updates * 14:53 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Unblock taavi - oblivian@cumin1003 * 14:52 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Unblock taavi - oblivian@cumin1003" * 14:46 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94711 and previous config saved to /var/cache/conftool/dbconfig/20260702-144644-fceratto.json * 14:36 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205', diff saved to https://phabricator.wikimedia.org/P94709 and previous config saved to /var/cache/conftool/dbconfig/20260702-143636-fceratto.json * 14:32 moritzm: installing libdbi-perl security updates * 14:26 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205', diff saved to https://phabricator.wikimedia.org/P94708 and previous config saved to /var/cache/conftool/dbconfig/20260702-142628-fceratto.json * 14:16 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94707 and previous config saved to /var/cache/conftool/dbconfig/20260702-141621-fceratto.json * 14:12 moritzm: installing rsync security updates * 14:11 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox) * 14:10 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94706 and previous config saved to /var/cache/conftool/dbconfig/20260702-140959-fceratto.json * 14:09 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2205.codfw.wmnet with reason: Maintenance * 14:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2205: Repooling after switchover * 14:07 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-test-master1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 14:06 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 14:06 Tran: Deployed patch for [[phab:T427287|T427287]] * 14:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:59 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2205: Repooling after switchover * 13:59 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2205: Repooling after switchover * 13:59 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:55 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2205: Repooling after switchover * 13:55 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2205 [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94704 and previous config saved to /var/cache/conftool/dbconfig/20260702-135505-fceratto.json * 13:54 moritzm: installing sed security updates * 13:53 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:52 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2209 to s3 primary [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94703 and previous config saved to /var/cache/conftool/dbconfig/20260702-135235-fceratto.json * 13:52 federico3: Starting s3 codfw failover from db2205 to db2209 - [[phab:T430912|T430912]] * 13:51 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:51 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 13:48 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:47 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2209 with weight 0 [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94702 and previous config saved to /var/cache/conftool/dbconfig/20260702-134719-fceratto.json * 13:47 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Primary switchover s3 [[phab:T430912|T430912]] * 13:44 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:44 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:44 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:40 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 13:38 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 13:37 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 13:36 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 13:36 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:34 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 13:30 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:29 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:29 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:27 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:26 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:25 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling restart_daemons on A:wikidough * 13:23 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 13:22 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 13:17 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 13:17 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns1004.wikimedia.org * 13:12 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:11 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart (exit_code=97) rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough * 13:11 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=97) rolling restart_daemons on A:wikidough * 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough * 13:09 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] (duration: 07m 20s) * 13:05 aude@deploy1003: jdrewniak, aude: Continuing with deployment * 13:04 aude@deploy1003: jdrewniak, aude: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:02 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] * 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts wdqs-categories1001.eqiad.wmnet * 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: wdqs-categories1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 12:10 jmm@dns1004: END - running authdns-update * 12:07 jmm@dns1004: START - running authdns-update * 11:51 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: wdqs-categories1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 11:44 btullis@cumin1003: START - Cookbook sre.dns.netbox * 11:42 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 11:42 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 11:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet * 11:39 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts wdqs-categories1001.eqiad.wmnet * 11:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet * 11:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet * 11:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet * 11:29 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 11:29 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 10:57 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2214: Repooling * 10:49 jmm@dns1004: END - running authdns-update * 10:47 jmm@dns1004: START - running authdns-update * 10:31 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94698 and previous config saved to /var/cache/conftool/dbconfig/20260702-103146-fceratto.json * 10:21 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213', diff saved to https://phabricator.wikimedia.org/P94696 and previous config saved to /var/cache/conftool/dbconfig/20260702-102137-fceratto.json * 10:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:19 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb1017.eqiad.wmnet * 10:18 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 10:18 fceratto@cumin1003: Removing es1033 from zarcillo [[phab:T408772|T408772]] * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts es1033.eqiad.wmnet * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: es1033.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:14 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: es1033.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:13 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb1017.eqiad.wmnet * 10:12 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2214.codfw.wmnet * 10:12 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2214.codfw.wmnet * 10:12 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2214: Repooling * 10:11 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213', diff saved to https://phabricator.wikimedia.org/P94693 and previous config saved to /var/cache/conftool/dbconfig/20260702-101130-fceratto.json * 10:10 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:10 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:03 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts es1033.eqiad.wmnet * 10:03 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 10:01 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94691 and previous config saved to /var/cache/conftool/dbconfig/20260702-100122-fceratto.json * 09:55 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94690 and previous config saved to /var/cache/conftool/dbconfig/20260702-095529-fceratto.json * 09:55 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2213.codfw.wmnet with reason: Maintenance * 09:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 09:53 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2213: Repooling after switchover * 09:51 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover * 09:44 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2213: Repooling after switchover * 09:39 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover * 09:39 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2213 [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94688 and previous config saved to /var/cache/conftool/dbconfig/20260702-093859-fceratto.json * 09:36 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2192 to s5 primary [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94687 and previous config saved to /var/cache/conftool/dbconfig/20260702-093650-fceratto.json * 09:36 federico3: Starting s5 codfw failover from db2213 to db2192 - [[phab:T430923|T430923]] * 09:30 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94686 and previous config saved to /var/cache/conftool/dbconfig/20260702-093004-fceratto.json * 09:24 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2192 with weight 0 [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94685 and previous config saved to /var/cache/conftool/dbconfig/20260702-092455-fceratto.json * 09:24 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 23 hosts with reason: Primary switchover s5 [[phab:T430923|T430923]] * 09:19 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220', diff saved to https://phabricator.wikimedia.org/P94684 and previous config saved to /var/cache/conftool/dbconfig/20260702-091957-fceratto.json * 09:16 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] (duration: 06m 57s) * 09:13 moritzm: installing libgcrypt20 security updates * 09:12 kharlan@deploy1003: kharlan: Continuing with deployment * 09:11 kharlan@deploy1003: kharlan: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:09 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220', diff saved to https://phabricator.wikimedia.org/P94683 and previous config saved to /var/cache/conftool/dbconfig/20260702-090950-fceratto.json * 09:09 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] * 09:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 09:01 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] (duration: 07m 07s) * 08:59 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94682 and previous config saved to /var/cache/conftool/dbconfig/20260702-085942-fceratto.json * 08:57 kharlan@deploy1003: kharlan: Continuing with deployment * 08:56 kharlan@deploy1003: kharlan: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:54 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] * 08:52 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:52 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:52 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94681 and previous config saved to /var/cache/conftool/dbconfig/20260702-085237-fceratto.json * 08:52 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2220.codfw.wmnet with reason: Maintenance * 08:43 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:40 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 08:25 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] (duration: 11m 44s) * 08:21 cscott@deploy1003: cscott: Continuing with deployment * 08:16 cscott@deploy1003: cscott: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:14 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] * 08:08 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 08:08 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1244: Migration of db1244.eqiad.wmnet completed * 08:02 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:02 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:01 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] (duration: 18m 58s) * 08:01 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:59 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 07:59 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:59 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:59 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2006.wikimedia.org * 07:58 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:57 cscott@deploy1003: cscott: Continuing with deployment * 07:56 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:56 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:56 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:55 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:55 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:55 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:54 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2006.wikimedia.org * 07:54 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:54 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 07:44 cscott@deploy1003: cscott: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:44 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2005.wikimedia.org * 07:44 moritzm: installing node-lodash security updates * 07:42 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] * 07:39 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2005.wikimedia.org * 07:30 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] (duration: 07m 28s) * 07:26 cscott@deploy1003: ssastry, cscott: Continuing with deployment * 07:25 cscott@deploy1003: ssastry, cscott: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:23 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1244: Migration of db1244.eqiad.wmnet completed * 07:22 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] * 07:16 wmde-fisch@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] (duration: 06m 55s) * 07:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1244.eqiad.wmnet with OS trixie * 07:11 wmde-fisch@deploy1003: wmde-fisch: Continuing with deployment * 07:11 wmde-fisch@deploy1003: wmde-fisch: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:09 wmde-fisch@deploy1003: Started scap sync-world: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] * 06:54 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1244.eqiad.wmnet with reason: host reimage * 06:50 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1244.eqiad.wmnet with reason: host reimage * 06:38 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1250.eqiad.wmnet with OS trixie * 06:34 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db1244.eqiad.wmnet with OS trixie * 06:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1244: Upgrading db1244.eqiad.wmnet * 06:25 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1244: Upgrading db1244.eqiad.wmnet * 06:25 cwilliams@cumin1003: dbmaint on s4@eqiad [[phab:T429893|T429893]] * 06:25 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 06:15 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1250.eqiad.wmnet with reason: host reimage * 06:14 cwilliams@dns1006: END - running authdns-update * 06:12 cwilliams@dns1006: START - running authdns-update * 06:11 cwilliams@dns1006: END - running authdns-update * 06:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db1244 [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94676 and previous config saved to /var/cache/conftool/dbconfig/20260702-061059-cwilliams.json * 06:09 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1250.eqiad.wmnet with reason: host reimage * 06:09 cwilliams@dns1006: START - running authdns-update * 06:08 aokoth@cumin1003: END (PASS) - Cookbook sre.vrts.upgrade (exit_code=0) on VRTS host vrts1003.eqiad.wmnet * 06:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db1160 to s4 primary and set section read-write [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94675 and previous config saved to /var/cache/conftool/dbconfig/20260702-060746-cwilliams.json * 06:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Set s4 eqiad as read-only for maintenance - [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94674 and previous config saved to /var/cache/conftool/dbconfig/20260702-060704-cwilliams.json * 06:06 cezmunsta: Starting s4 eqiad failover from db1244 to db1160 - [[phab:T430817|T430817]] * 06:04 aokoth@cumin1003: START - Cookbook sre.vrts.upgrade on VRTS host vrts1003.eqiad.wmnet * 05:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db1160 with weight 0 [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94673 and previous config saved to /var/cache/conftool/dbconfig/20260702-055927-cwilliams.json * 05:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 40 hosts with reason: Primary switchover s4 [[phab:T430817|T430817]] * 05:55 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1250.eqiad.wmnet with OS trixie * 05:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on db1250.eqiad.wmnet with reason: m3 master switchover [[phab:T430158|T430158]] * 05:39 marostegui: Failover m3 (phabricator) from db1250 to db1228 - [[phab:T430158|T430158]] * 05:32 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2234].codfw.wmnet,db[1217,1228,1250].eqiad.wmnet with reason: m3 master switchover [[phab:T430158|T430158]] * 04:45 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] (duration: 09m 08s) * 04:41 tstarling@deploy1003: tstarling, reedy: Continuing with deployment * 04:38 tstarling@deploy1003: tstarling, reedy: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 04:36 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 59s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:16 ryankemper: [[phab:T429844|T429844]] [opensearch] completed `cirrussearch2111` reimage; all codfw search clusters are green, all nodes now report `OpenSearch 2.19.5`, and the temporary chi voting exclusion has been removed * 00:57 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2111.codfw.wmnet with OS trixie * 00:29 ryankemper: [[phab:T429844|T429844]] [opensearch] depooled codfw search-omega/search-psi discovery records to match existing codfw search depool during OpenSearch 2.19 migration * 00:29 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2111.codfw.wmnet with reason: host reimage * 00:29 ryankemper@cumin2002: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 00:29 ryankemper@cumin2002: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 00:22 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2111.codfw.wmnet with reason: host reimage * 00:01 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2111.codfw.wmnet with OS trixie * 00:00 ryankemper: [[phab:T429844|T429844]] [opensearch] chi cluster recovered after stopping `opensearch_1@production-search-codfw` on `cirrussearch2111` == 2026-07-01 == * 23:59 ryankemper: [[phab:T429844|T429844]] [opensearch] stopped `opensearch_1@production-search-codfw` on `cirrussearch2111` after chi cluster-manager election churn following `voting_config_exclusions` POST; hoping this triggers a re-election * 23:52 cscott@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 23:51 cscott@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 23:51 cscott@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 23:50 cscott@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2003.codfw.wmnet with OS bookworm * 22:29 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 22:13 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 22:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2084.codfw.wmnet with OS trixie * 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2003.codfw.wmnet with reason: host reimage * 22:03 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 22:01 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2003.codfw.wmnet with reason: host reimage * 21:50 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 21:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2084.codfw.wmnet with reason: host reimage * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2003 * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2003 * 21:42 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2003 * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2003.codfw.wmnet 45.48.192.10.in-addr.arpa 5.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:42 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2003.codfw.wmnet 45.48.192.10.in-addr.arpa 5.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2003 - bking@cumin2003" * 21:42 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2003 - bking@cumin2003" * 21:36 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2084.codfw.wmnet with reason: host reimage * 21:35 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:34 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2003 * 21:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2003.codfw.wmnet with OS bookworm * 21:19 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2084.codfw.wmnet with OS trixie * 21:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2081.codfw.wmnet with OS trixie * 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2108.codfw.wmnet with OS trixie * 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2081.codfw.wmnet with reason: host reimage * 20:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2081.codfw.wmnet with reason: host reimage * 20:28 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2081.codfw.wmnet with OS trixie * 20:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2108.codfw.wmnet with reason: host reimage * 20:19 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2108.codfw.wmnet with reason: host reimage * 19:59 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2108.codfw.wmnet with OS trixie * 19:46 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2093.codfw.wmnet with OS trixie * 19:44 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 19:44 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jasmine@cumin2002" * 19:43 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jasmine@cumin2002" * 19:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2080.codfw.wmnet with OS trixie * 19:28 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 19:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2093.codfw.wmnet with reason: host reimage * 19:18 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 19:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2093.codfw.wmnet with reason: host reimage * 19:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2080.codfw.wmnet with reason: host reimage * 19:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2080.codfw.wmnet with reason: host reimage * 18:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2093.codfw.wmnet with OS trixie * 18:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2080.codfw.wmnet with OS trixie * 18:27 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 18:18 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] (duration: 09m 15s) * 18:13 jgiannelos@deploy1003: jgiannelos, neriah: Continuing with deployment * 18:11 jgiannelos@deploy1003: jgiannelos, neriah: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:09 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] * 17:40 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 16:58 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 30 hosts * 16:57 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for 30 hosts * 16:52 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2202.codfw.wmnet * 16:52 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2202.codfw.wmnet * 16:51 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt * 16:51 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt * 16:51 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lvs2012.codfw.wmnet * 16:51 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for lvs2012.codfw.wmnet * 16:49 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2076.codfw.wmnet with OS trixie * 16:49 brett: Start pybal on lvs2012 - [[phab:T429861|T429861]] * 16:49 pt1979@cumin1003: END (ERROR) - Cookbook sre.hosts.remove-downtime (exit_code=97) for 59 hosts * 16:48 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for 59 hosts * 16:42 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2061.codfw.wmnet with OS trixie * 16:30 dancy@deploy1003: Installation of scap version "4.271.0" completed for 2 hosts * 16:28 dancy@deploy1003: Installing scap version "4.271.0" for 2 host(s) * 16:23 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2076.codfw.wmnet with reason: host reimage * 16:19 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2061.codfw.wmnet with reason: host reimage * 16:18 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2076.codfw.wmnet with reason: host reimage * 16:18 jasmine@dns1004: END - running authdns-update * 16:16 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host restbase2039.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 16:16 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host restbase2039.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 16:16 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2061.codfw.wmnet with reason: host reimage * 16:15 jasmine@dns1004: START - running authdns-update * 16:14 jasmine@dns1004: END - running authdns-update * 16:12 jasmine@dns1004: START - running authdns-update * 16:07 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2202.codfw.wmnet with reason: maintenance * 16:06 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt with reason: Junos upograde * 16:00 papaul: ongoing maintenance on lsw1-b2-codfw * 16:00 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2076.codfw.wmnet with OS trixie * 15:59 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt * 15:59 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt * 15:57 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2061.codfw.wmnet with OS trixie * 15:55 pt1979@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2042,2046].codfw.wmnet * 15:55 pt1979@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2042,2046].codfw.wmnet * 15:51 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 15:51 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2220: Repooling after switchover * 15:50 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 15:50 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 15:48 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2092.codfw.wmnet with OS trixie * 15:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 15:40 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 15:38 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 15:37 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 15:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply * 15:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply * 15:32 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 15:32 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 15:30 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 15:29 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 15:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 15:25 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:22 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs2012.codfw.wmnet with reason: Rack B2 maintenance - [[phab:T429861|T429861]] * 15:21 brett: Stopping pybal on lvs2012 in preparation for codfw rack b2 maintenance - [[phab:T429861|T429861]] * 15:20 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2092.codfw.wmnet with reason: host reimage * 15:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:12 _joe_: restarted manually alertmanager-irc-relay * 15:12 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2092.codfw.wmnet with reason: host reimage * 15:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:12 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt with reason: Junos upograde * 15:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover * 15:07 pt1979@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2042,2046].codfw.wmnet * 15:06 pt1979@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2042,2046].codfw.wmnet * 15:02 papaul: ongoing maintenance on lsw1-a8-codfw * 14:31 topranks: POWERING DOWN CR1-EQIAD for line card installation [[phab:T426343|T426343]] * 14:31 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] (duration: 08m 57s) * 14:29 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:26 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 14:24 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:22 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] * 14:22 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover * 14:16 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:15 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover * 14:14 topranks: re-enable routing-engine graceful-failover on cr1-eqiad [[phab:T417873|T417873]] * 14:13 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:13 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2220: Repooling after switchover * 14:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:12 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:12 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:11 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:08 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] (duration: 10m 01s) * 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:07 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2220 [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94664 and previous config saved to /var/cache/conftool/dbconfig/20260701-140729-fceratto.json * 14:06 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:06 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:06 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:05 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2159 to s7 primary [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94663 and previous config saved to /var/cache/conftool/dbconfig/20260701-140503-fceratto.json * 14:04 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:04 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 14:04 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 14:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:04 dreamyjazz@deploy1003: anzx, dreamyjazz: Continuing with deployment * 14:04 federico3: Starting s7 codfw failover from db2220 to db2159 - [[phab:T430826|T430826]] * 14:03 jmm@dns1004: END - running authdns-update * 14:03 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:03 topranks: flipping cr1-eqiad active routing-enginer back to RE0 [[phab:T417873|T417873]] * 14:03 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudsw1-c8-eqiad,cloudsw1-d5-eqiad with reason: router upgrades eqiad * 14:01 jmm@dns1004: START - running authdns-update * 14:00 dreamyjazz@deploy1003: anzx, dreamyjazz: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:59 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2159 with weight 0 [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94662 and previous config saved to /var/cache/conftool/dbconfig/20260701-135906-fceratto.json * 13:58 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] * 13:57 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s7 [[phab:T430826|T430826]] * 13:56 topranks: reboot routing-enginer RE0 on cr1-eqiad [[phab:T417873|T417873]] * 13:48 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1006.wikimedia.org * 13:44 atsuko@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cirrussearch2092.codfw.wmnet with OS trixie * 13:43 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1006.wikimedia.org * 13:41 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2092.codfw.wmnet with OS trixie * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1005.wikimedia.org * 13:37 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1005.wikimedia.org * 13:37 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on pfw1-eqiad with reason: router upgrades eqiad * 13:35 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on lvs[1017-1020].eqiad.wmnet with reason: router upgrades eqiad * 13:34 caro@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] (duration: 07m 59s) * 13:30 caro@deploy1003: caro: Continuing with deployment * 13:28 caro@deploy1003: caro: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:27 topranks: route-engine failover cr1-eqiad * 13:26 caro@deploy1003: Started scap sync-world: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] * 13:15 topranks: rebooting routing-engine 1 on cr1-eqiad [[phab:T417873|T417873]] * 13:13 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] (duration: 08m 29s) * 13:13 moritzm: installing qemu security updates * 13:11 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 13:11 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 13:09 jgiannelos@deploy1003: jgiannelos: Continuing with deployment * 13:08 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 13:07 jgiannelos@deploy1003: jgiannelos: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:06 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 13:06 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2214.codfw.wmnet with reason: Maintenance * 13:05 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2214: Repooling after switchover * 13:05 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] * 13:04 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2214: Repooling after switchover * 13:04 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2214 [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94660 and previous config saved to /var/cache/conftool/dbconfig/20260701-130413-fceratto.json * 13:01 moritzm: installing python3.13 security updates * 13:00 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2229 to s6 primary [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94659 and previous config saved to /var/cache/conftool/dbconfig/20260701-125959-fceratto.json * 12:59 federico3: Starting s6 codfw failover from db2214 to db2229 - [[phab:T430814|T430814]] * 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on 13 hosts with reason: router upgrade and line card install * 12:51 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2229 with weight 0 [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94658 and previous config saved to /var/cache/conftool/dbconfig/20260701-125149-fceratto.json * 12:51 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 21 hosts with reason: Primary switchover s6 [[phab:T430814|T430814]] * 12:50 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2189.codfw.wmnet * 12:50 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2189.codfw.wmnet * 12:42 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2100.codfw.wmnet with OS trixie * 12:38 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2083.codfw.wmnet with OS trixie * 12:19 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2083.codfw.wmnet with reason: host reimage * 12:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 12:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2240: Migration of db2240.codfw.wmnet completed * 12:14 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2100.codfw.wmnet with reason: host reimage * 12:09 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2083.codfw.wmnet with reason: host reimage * 12:09 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2100.codfw.wmnet with reason: host reimage * 12:00 topranks: drain traffic on cr1-eqiad to allow for line card install and JunOS upgrade [[phab:T426343|T426343]] * 11:52 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2083.codfw.wmnet with OS trixie * 11:50 cmooney@dns2005: END - running authdns-update * 11:49 cmooney@dns2005: START - running authdns-update * 11:48 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2100.codfw.wmnet with OS trixie * 11:40 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/zotero: apply * 11:40 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/zotero: apply * 11:36 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/zotero: apply * 11:36 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/zotero: apply * 11:31 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2240: Migration of db2240.codfw.wmnet completed * 11:30 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply * 11:28 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply * 11:27 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:27 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:27 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:27 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:27 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:23 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2240.codfw.wmnet with OS trixie * 11:20 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:20 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:17 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:16 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:16 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:15 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2086.codfw.wmnet with OS trixie * 11:14 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2106.codfw.wmnet with OS trixie * 11:14 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:13 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:12 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:09 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2115.codfw.wmnet with OS trixie * 11:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2240.codfw.wmnet with reason: host reimage * 11:00 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2240.codfw.wmnet with reason: host reimage * 10:53 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2106.codfw.wmnet with reason: host reimage * 10:49 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2086.codfw.wmnet with reason: host reimage * 10:44 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2115.codfw.wmnet with reason: host reimage * 10:44 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2240.codfw.wmnet with OS trixie * 10:44 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2086.codfw.wmnet with reason: host reimage * 10:42 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2106.codfw.wmnet with reason: host reimage * 10:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2240: Upgrading db2240.codfw.wmnet * 10:41 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2240: Upgrading db2240.codfw.wmnet * 10:41 cwilliams@cumin1003: dbmaint on s4@codfw [[phab:T429893|T429893]] * 10:40 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 10:39 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2115.codfw.wmnet with reason: host reimage * 10:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2240 [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94653 and previous config saved to /var/cache/conftool/dbconfig/20260701-102658-cwilliams.json * 10:26 moritzm: installing nginx security updates * 10:26 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2086.codfw.wmnet with OS trixie * 10:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2179 to s4 primary [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94652 and previous config saved to /var/cache/conftool/dbconfig/20260701-102356-cwilliams.json * 10:23 cezmunsta: Starting s4 codfw failover from db2240 to db2179 - [[phab:T430127|T430127]] * 10:23 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2106.codfw.wmnet with OS trixie * 10:20 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2115.codfw.wmnet with OS trixie * 10:15 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2179 with weight 0 [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94651 and previous config saved to /var/cache/conftool/dbconfig/20260701-101531-cwilliams.json * 10:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 40 hosts with reason: Primary switchover s4 [[phab:T430127|T430127]] * 09:56 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template (take 2) - oblivian@cumin1003" * 09:56 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template (take 2) - oblivian@cumin1003 * 09:55 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template (take 2) - oblivian@cumin1003 * 09:55 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template (take 2) - oblivian@cumin1003" * 09:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:39 mszwarc@deploy1003: Synchronized private/SuggestedInvestigationsSignals/SuggestedInvestigationsSignal4n.php: Update SI signal 4n (duration: 06m 08s) * 09:21 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 09:21 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 09:14 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 09:14 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 09:02 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 09:02 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 08:54 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 08:38 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 08:38 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 08:36 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 08:21 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 08:21 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] (duration: 36m 11s) * 08:15 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 08:09 mszwarc@deploy1003: mszwarc, abi: Continuing with deployment * 08:03 mszwarc@deploy1003: mszwarc, abi: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:55 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 07:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 07:45 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] * 07:30 aqu@deploy1003: Finished deploy [analytics/refinery@410f205]: Regular analytics weekly train 2nd try [analytics/refinery@410f2050] (duration: 00m 22s) * 07:29 aqu@deploy1003: Started deploy [analytics/refinery@410f205]: Regular analytics weekly train 2nd try [analytics/refinery@410f2050] * 07:28 aqu@deploy1003: Finished deploy [analytics/refinery@410f205] (thin): Regular analytics weekly train THIN [analytics/refinery@410f2050] (duration: 01m 59s) * 07:28 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] (duration: 07m 19s) * 07:26 aqu@deploy1003: Started deploy [analytics/refinery@410f205] (thin): Regular analytics weekly train THIN [analytics/refinery@410f2050] * 07:26 aqu@deploy1003: Finished deploy [analytics/refinery@410f205]: Regular analytics weekly train [analytics/refinery@410f2050] (duration: 04m 32s) * 07:24 mszwarc@deploy1003: wmde-fisch, mszwarc: Continuing with deployment * 07:23 mszwarc@deploy1003: wmde-fisch, mszwarc: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:21 aqu@deploy1003: Started deploy [analytics/refinery@410f205]: Regular analytics weekly train [analytics/refinery@410f2050] * 07:21 aqu@deploy1003: Finished deploy [analytics/refinery@410f205] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@410f2050] (duration: 02m 01s) * 07:20 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] * 07:19 aqu@deploy1003: Started deploy [analytics/refinery@410f205] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@410f2050] * 07:13 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] (duration: 09m 13s) * 07:09 mszwarc@deploy1003: mszwarc, chlod, revi: Continuing with deployment * 07:06 mszwarc@deploy1003: mszwarc, chlod, revi: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:04 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] * 06:55 elukey: upgrade all trixie hosts to pywmflib 3.0 - [[phab:T430552|T430552]] * 06:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:43 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:43 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:42 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:42 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:41 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:41 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:35 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:35 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:34 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:34 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:31 jmm@cumin2003: DONE (PASS) - Cookbook sre.idm.logout (exit_code=0) Logging Niharika29 out of all services on: 2453 hosts * 06:30 oblivian@cumin1003: END (FAIL) - Cookbook sre.deploy.hiddenparma (exit_code=99) Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:30 oblivian@cumin1003: END (FAIL) - Cookbook sre.deploy.python-code (exit_code=99) hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:30 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:30 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:01 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2109.codfw.wmnet with OS trixie * 05:45 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on es1039.eqiad.wmnet with reason: issues * 05:41 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1027.eqiad.wmnet * 05:40 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2068.codfw.wmnet with OS trixie * 05:40 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2109.codfw.wmnet with reason: host reimage * 05:40 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1027.eqiad.wmnet,service=s2 * 05:40 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1027.eqiad.wmnet,service=s7 * 05:36 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2109.codfw.wmnet with reason: host reimage * 05:20 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2068.codfw.wmnet with reason: host reimage * 05:16 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2109.codfw.wmnet with OS trixie * 05:15 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2068.codfw.wmnet with reason: host reimage * 05:09 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2067.codfw.wmnet with OS trixie * 04:56 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2068.codfw.wmnet with OS trixie * 04:49 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2067.codfw.wmnet with reason: host reimage * 04:45 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2067.codfw.wmnet with reason: host reimage * 04:27 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2067.codfw.wmnet with OS trixie * 03:47 slyngshede@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1039.eqiad.wmnet with reason: Hardware crash * 03:21 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2107.codfw.wmnet with OS trixie * 02:59 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2107.codfw.wmnet with reason: host reimage * 02:55 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2085.codfw.wmnet with OS trixie * 02:51 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2072.codfw.wmnet with OS trixie * 02:51 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2107.codfw.wmnet with reason: host reimage * 02:35 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2085.codfw.wmnet with reason: host reimage * 02:31 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2107.codfw.wmnet with OS trixie * 02:30 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2072.codfw.wmnet with reason: host reimage * 02:26 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2085.codfw.wmnet with reason: host reimage * 02:22 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2072.codfw.wmnet with reason: host reimage * 02:09 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2085.codfw.wmnet with OS trixie * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 54s) * 02:03 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2072.codfw.wmnet with OS trixie * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es7 eqiad back to read-write - [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94649 and previous config saved to /var/cache/conftool/dbconfig/20260701-010716-ladsgroup.json * 01:05 ladsgroup@dns1004: END - running authdns-update * 01:05 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depool es1039 [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94648 and previous config saved to /var/cache/conftool/dbconfig/20260701-010551-ladsgroup.json * 01:03 ladsgroup@dns1004: START - running authdns-update * 01:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Promote es1035 to es7 primary [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94647 and previous config saved to /var/cache/conftool/dbconfig/20260701-010002-ladsgroup.json * 00:58 Amir1: Starting es7 eqiad failover from es1039 to es1035 - [[phab:T430765|T430765]] * 00:53 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es1035 with weight 0 [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94646 and previous config saved to /var/cache/conftool/dbconfig/20260701-005329-ladsgroup.json * 00:53 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 9 hosts with reason: Primary switchover es7 [[phab:T430765|T430765]] * 00:42 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es7 eqiad as read-only for maintenance - [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94645 and previous config saved to /var/cache/conftool/dbconfig/20260701-004221-ladsgroup.json * 00:20 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2102.codfw.wmnet with OS trixie * 00:15 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2103.codfw.wmnet with OS trixie * 00:05 dr0ptp4kt: DEPLOYED Refinery at {{Gerrit|4e7a2b32}} for changes: pageview allowlist {{Gerrit|1305158}} (+min.wikiquote) {{Gerrit|1305162}} (+bol.wikipedia), {{Gerrit|1305156}} (+isv.wikipedia); {{Gerrit|1305980}} (pv allowlist -api.wikimedia, sqoop +isvwiki); sqoop {{Gerrit|1295064}} (+globalimagelinks) {{Gerrit|1295069}} (+filerevision) using scap, then deployed onto HDFS (manual copyToLocal required additionally) == Other archives == See [[Server Admin Log/Archives]]. <noinclude> [[Category:SAL]] [[Category:Operations]] </noinclude> n3chy5exiy0fsiapjwn515fj2vb89f7 2450656 2450655 2026-08-22T16:37:21Z Stashbot 7414 arlolra@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply 2450656 wikitext text/x-wiki == 2026-08-22 == * 16:37 arlolra@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:36 arlolra@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:36 arlolra@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:36 arlolra@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:36 arlolra@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:33 arlolra@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:33 arlolra@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:33 arlolra@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:32 arlolra@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 35s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-21 == * 20:36 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in1001.wikimedia.org with reason: [[phab:T434750|T434750]] * 20:34 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in2001.wikimedia.org with reason: [[phab:T434750|T434750]] * 20:33 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out1001.wikimedia.org with reason: [[phab:T434750|T434750]] * 20:25 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out2001.wikimedia.org with reason: [[phab:T434750|T434750]] * 19:37 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host krb1002.eqiad.wmnet with OS bookworm * 19:00 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 18:59 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 18:51 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 18:51 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 18:35 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:35 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:27 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:27 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:16 bking@cumin2003: START - Cookbook sre.hosts.reimage for host krb1002.eqiad.wmnet with OS bookworm * 17:35 sukhe@dns1004: END - running authdns-update * 17:33 sukhe@dns1004: START - running authdns-update * 17:32 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns5004.wikimedia.org [reason: resolved authdns-update issues] * 17:32 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:32 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: force HEAD to {{Gerrit|be26e30ae101}} - sukhe@cumin1003" * 17:32 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: force HEAD to {{Gerrit|be26e30ae101}} - sukhe@cumin1003" * 17:28 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 17:28 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: service=authdns-update,name=dns5004.wikimedia.org [reason: resolving authdns-update issues] * 17:28 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:28 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: force HEAD to {{Gerrit|be26e30ae101}} - sukhe@cumin1003" * 17:28 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: force HEAD to {{Gerrit|be26e30ae101}} - sukhe@cumin1003" * 17:24 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 17:24 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.netbox (exit_code=97) * 17:23 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 17:16 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 17:12 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 17:10 sukhe@dns1004: END - running authdns-update * 17:08 sukhe@dns1004: START - running authdns-update * 17:08 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=dns5004.wikimedia.org [reason: resolving authdns-update issues] * 17:07 sukhe@dns1004: FAIL - running authdns-update * 17:05 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 17:05 sukhe@dns1004: START - running authdns-update * 17:01 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 16:59 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 16:56 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 16:53 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=dns5004.* [reason: trixie upgrade] * 16:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns5004.wikimedia.org * 16:52 cdobbins@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns5004.wikimedia.org * 16:44 cmooney@dns3003: END - running authdns-update * 16:41 cmooney@dns3003: START - running authdns-update * 16:41 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:41 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on eqsin<->codfw arelion - cmooney@cumin1003" * 16:37 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on eqsin<->codfw arelion - cmooney@cumin1003" * 16:33 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:11 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 16:08 sukhe@dns1004: END - running authdns-update * 16:08 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 16:06 sukhe@dns1004: START - running authdns-update * 16:04 cmooney@dns3003: END - running authdns-update * 16:02 cmooney@dns3003: START - running authdns-update * 16:00 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:00 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on eqord<->codfw arelion - cmooney@cumin1003" * 15:56 cdobbins@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host dns5004.wikimedia.org with OS trixie * 15:55 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on eqord<->codfw arelion - cmooney@cumin1003" * 15:53 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:51 cmooney@cumin1003: END (ERROR) - Cookbook sre.dns.netbox (exit_code=97) * 15:51 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:38 andrew@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudcephosd1042.eqiad.wmnet with OS bookworm * 15:18 andrew@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudcephosd1042.eqiad.wmnet with reason: host reimage * 15:17 dancy@deploy1003: Finished deploy [gerrit/gerrit@2cc11cc]: Deploying https://gerrit.wikimedia.org/r/c/operations/software/gerrit/+/1327669 ([[phab:T434726|T434726]]) (duration: 00m 14s) * 15:17 dancy@deploy1003: Started deploy [gerrit/gerrit@2cc11cc]: Deploying https://gerrit.wikimedia.org/r/c/operations/software/gerrit/+/1327669 ([[phab:T434726|T434726]]) * 15:13 andrew@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudcephosd1042.eqiad.wmnet with reason: host reimage * 15:09 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns5004.wikimedia.org with reason: host reimage * 15:05 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns5004.wikimedia.org with reason: host reimage * 14:53 andrew@cumin2003: START - Cookbook sre.hosts.reimage for host cloudcephosd1042.eqiad.wmnet with OS bookworm * 14:30 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns5004.wikimedia.org with OS trixie * 14:29 cdobbins@cumin1003: conftool action : set/pooled=no; selector: name=dns5004.* [reason: trixie upgrade] * 14:21 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:21 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove entries for cr2-eqord - cmooney@cumin1003" * 14:21 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove entries for cr2-eqord - cmooney@cumin1003" * 14:13 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 14:11 moritzm: imported openjdk 8u504-ga-1~deb12u1 for bookworm-wikimedia (backport of the latest Java 8 security fixes for bookworm) * 13:25 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "sync cr2-eqord router offline - cmooney@cumin1003" * 13:23 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "sync cr2-eqord router offline - cmooney@cumin1003" * 13:14 hashar@deploy1003: Finished deploy [integration/docroot@2d5ff9b]: opensource: add PersonalDashboard docs to MW components - [[phab:T435392|T435392]] (duration: 00m 15s) * 13:14 hashar@deploy1003: Started deploy [integration/docroot@2d5ff9b]: opensource: add PersonalDashboard docs to MW components - [[phab:T435392|T435392]] * 12:16 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:16 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: [[phab:T431682|T431682]] - filippo@cumin1003" * 12:16 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: [[phab:T431682|T431682]] - filippo@cumin1003" * 12:11 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2006.wikimedia.org with OS trixie * 12:00 kevinbazira@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 11:58 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 11:43 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2006.wikimedia.org with reason: host reimage * 11:41 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 11:38 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2006.wikimedia.org with reason: host reimage * 11:20 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2006.wikimedia.org with OS trixie * 11:11 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2005.wikimedia.org with OS trixie * 10:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2005.wikimedia.org with reason: host reimage * 10:53 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2005.wikimedia.org with reason: host reimage * 10:43 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-codfw * 10:43 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2011.codfw.wmnet * 10:43 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2011.codfw.wmnet * 10:40 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2011.codfw.wmnet * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2011.codfw.wmnet * 10:34 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2010.codfw.wmnet * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2010.codfw.wmnet * 10:33 fnegri@cumin1003: END (PASS) - Cookbook sre.wikireplicas.add-wiki (exit_code=0) for database bolwiki ([[phab:T429954|T429954]]) * 10:33 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2005.wikimedia.org with OS trixie * 10:30 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2010.codfw.wmnet * 10:25 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2010.codfw.wmnet * 10:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2009.codfw.wmnet * 10:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2009.codfw.wmnet * 10:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1006.wikimedia.org with OS trixie * 10:18 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2009.codfw.wmnet * 10:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2009.codfw.wmnet * 10:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2008.codfw.wmnet * 10:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2008.codfw.wmnet * 10:06 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2008.codfw.wmnet * 10:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1006.wikimedia.org with reason: host reimage * 10:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2008.codfw.wmnet * 10:01 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2007.codfw.wmnet * 10:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2007.codfw.wmnet * 09:57 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1006.wikimedia.org with reason: host reimage * 09:56 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2007.codfw.wmnet * 09:51 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2007.codfw.wmnet * 09:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 09:51 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 09:46 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2006.codfw.wmnet * 09:46 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1006.wikimedia.org with OS trixie * 09:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1005.wikimedia.org with OS trixie * 09:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2006.codfw.wmnet * 09:41 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2005.codfw.wmnet * 09:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2005.codfw.wmnet * 09:38 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2013.codfw.wmnet * 09:36 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2005.codfw.wmnet * 09:35 fnegri@cumin1003: START - Cookbook sre.wikireplicas.add-wiki for database bolwiki ([[phab:T429954|T429954]]) * 09:35 fnegri@cumin1003: END (PASS) - Cookbook sre.wikireplicas.add-wiki (exit_code=0) for database minwikiquote ([[phab:T429946|T429946]]) * 09:35 fnegri@cumin1003: START - Cookbook sre.wikireplicas.add-wiki for database minwikiquote ([[phab:T429946|T429946]]) * 09:32 blake@cumin1003: START - Cookbook sre.hosts.reboot-single for host rdb2013.codfw.wmnet * 09:30 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2011.codfw.wmnet * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1005.wikimedia.org with reason: host reimage * 09:26 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2005.codfw.wmnet * 09:26 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2004.codfw.wmnet * 09:26 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2004.codfw.wmnet * 09:24 blake@cumin1003: START - Cookbook sre.hosts.reboot-single for host rdb2011.codfw.wmnet * 09:22 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1005.wikimedia.org with reason: host reimage * 09:21 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2004.codfw.wmnet * 09:16 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1015.eqiad.wmnet * 09:16 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2004.codfw.wmnet * 09:16 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2003.codfw.wmnet * 09:16 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2003.codfw.wmnet * 09:13 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on cr[1-2]-eqiad,pfw1-eqiad with reason: upgrade pfw1a-eqiad and pfw1b-eqiad pair * 09:12 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 09:11 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 09:11 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 09:11 blake@cumin1003: START - Cookbook sre.hosts.reboot-single for host rdb1015.eqiad.wmnet * 09:10 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2003.codfw.wmnet * 09:09 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1013.eqiad.wmnet * 09:07 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1005.wikimedia.org with OS trixie * 09:03 blake@cumin1003: START - Cookbook sre.hosts.reboot-single for host rdb1013.eqiad.wmnet * 09:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2003.codfw.wmnet * 09:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2002.codfw.wmnet * 09:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2002.codfw.wmnet * 08:54 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2002.codfw.wmnet * 08:49 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2002.codfw.wmnet * 08:49 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2001.codfw.wmnet * 08:49 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2001.codfw.wmnet * 08:48 jmm@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts netmon2002.wikimedia.org * 08:47 jmm@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts netmon2002.wikimedia.org * 08:44 jmm@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts netmon2002.wikimedia.org * 08:44 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon2002.wikimedia.org * 08:43 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2001.codfw.wmnet * 08:36 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon2002.wikimedia.org * 08:34 jmm@dns1004: END - running authdns-update * 08:33 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2001.codfw.wmnet * 08:33 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-codfw * 08:31 jmm@dns1004: START - running authdns-update * 07:48 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327679{{!}}Block: Disable flaky API test (T435272 T389028)]], [[gerrit:1327678{{!}}API: wfDebugLog for thumberror]] (duration: 15m 34s) * 07:41 krinkle@deploy1003: krinkle: Continuing with deployment * 07:37 krinkle@deploy1003: krinkle: Backport for [[gerrit:1327679{{!}}Block: Disable flaky API test (T435272 T389028)]], [[gerrit:1327678{{!}}API: wfDebugLog for thumberror]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:33 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1327679{{!}}Block: Disable flaky API test (T435272 T389028)]], [[gerrit:1327678{{!}}API: wfDebugLog for thumberror]] * 07:25 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 07:24 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 07:18 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 07:18 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 07:15 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 07:14 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 07:14 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 07:14 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 07:13 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 07:03 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1283: Pool back * 06:42 jmm@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts netmon2002.wikimedia.org * 06:35 moritzm: powercycling netmon2002 * 06:18 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1283: Pool back * 06:17 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1283 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96213 and previous config saved to /var/cache/conftool/dbconfig/20260821-061743-marostegui.json * 04:59 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324963{{!}}Add Produnto to extension-list (T421436)]], [[gerrit:1324964{{!}}Enable Produnto on Beta (T421436)]] (duration: 34m 48s) * 04:45 tstarling@deploy1003: tstarling: Continuing with deployment * 04:44 tstarling@deploy1003: tstarling: Backport for [[gerrit:1324963{{!}}Add Produnto to extension-list (T421436)]], [[gerrit:1324964{{!}}Enable Produnto on Beta (T421436)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 04:24 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1324963{{!}}Add Produnto to extension-list (T421436)]], [[gerrit:1324964{{!}}Enable Produnto on Beta (T421436)]] * 04:21 arlolra@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 04:20 arlolra@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 04:20 arlolra@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 04:20 arlolra@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 41s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-20 == * 23:43 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327652{{!}}RunSingleJob: Add ProfilingContext::init() (T435422)]] (duration: 12m 06s) * 23:38 krinkle@deploy1003: krinkle: Continuing with deployment * 23:33 krinkle@deploy1003: krinkle: Backport for [[gerrit:1327652{{!}}RunSingleJob: Add ProfilingContext::init() (T435422)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:31 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1327652{{!}}RunSingleJob: Add ProfilingContext::init() (T435422)]] * 22:15 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1054.eqiad.wmnet with OS trixie * 22:14 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 22:14 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 21:58 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1054.eqiad.wmnet with reason: host reimage * 21:51 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1054.eqiad.wmnet with reason: host reimage * 21:36 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1054.eqiad.wmnet with OS trixie * 21:36 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:35 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327219{{!}}RunSingleJob: Define MW_ENTRY_POINT for flamegraph sample attribution (T435422)]] (duration: 08m 30s) * 21:31 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:31 krinkle@deploy1003: krinkle: Continuing with deployment * 21:31 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1054 * 21:31 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1054 * 21:29 krinkle@deploy1003: krinkle: Backport for [[gerrit:1327219{{!}}RunSingleJob: Define MW_ENTRY_POINT for flamegraph sample attribution (T435422)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:27 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1327219{{!}}RunSingleJob: Define MW_ENTRY_POINT for flamegraph sample attribution (T435422)]] * 21:17 maryum: Deployed security fix for [[phab:T433020|T433020]] * 20:59 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324752{{!}}InitialiseSettings: Enable 2FA banners on remaining private wikis (T428103)]], [[gerrit:1325920{{!}}Remove sending email to legal team about rejected requests (T374053)]] (duration: 07m 18s) * 20:54 reedy@deploy1003: neriah, reedy: Continuing with deployment * 20:54 reedy@deploy1003: neriah, reedy: Backport for [[gerrit:1324752{{!}}InitialiseSettings: Enable 2FA banners on remaining private wikis (T428103)]], [[gerrit:1325920{{!}}Remove sending email to legal team about rejected requests (T374053)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:51 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324752{{!}}InitialiseSettings: Enable 2FA banners on remaining private wikis (T428103)]], [[gerrit:1325920{{!}}Remove sending email to legal team about rejected requests (T374053)]] * 20:24 reedy@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.15,1.47.0-wmf.16,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/med * 20:23 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324752{{!}}InitialiseSettings: Enable 2FA banners on remaining private wikis (T428103)]], [[gerrit:1325920{{!}}Remove sending email to legal team about rejected requests (T374053)]] * 20:14 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 20:10 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 20:09 cdanis@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "bug fixes & UX fixes - cdanis@cumin1003" * 20:09 cdanis@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: bug fixes & UX fixes - cdanis@cumin1003 * 20:08 cdanis@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: bug fixes & UX fixes - cdanis@cumin1003 * 20:08 cdanis@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "bug fixes & UX fixes - cdanis@cumin1003" * 19:24 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327598{{!}}Make \Omicron non upright (like \Chi) (T434428)]], [[gerrit:1327596{{!}}Render overline of \bar with stretchy=false (T435456)]] (duration: 18m 54s) * 19:20 krinkle@deploy1003: krinkle: Continuing with deployment * 19:07 krinkle@deploy1003: krinkle: Backport for [[gerrit:1327598{{!}}Make \Omicron non upright (like \Chi) (T434428)]], [[gerrit:1327596{{!}}Render overline of \bar with stretchy=false (T435456)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:05 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1327598{{!}}Make \Omicron non upright (like \Chi) (T434428)]], [[gerrit:1327596{{!}}Render overline of \bar with stretchy=false (T435456)]] * 18:51 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327614{{!}}Avoid casting fpxmax to string (T318419)]] (duration: 07m 28s) * 18:50 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:46 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:46 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 18:45 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327614{{!}}Avoid casting fpxmax to string (T318419)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:43 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327614{{!}}Avoid casting fpxmax to string (T318419)]] * 18:37 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:37 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:37 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:36 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:07 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 18:05 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 18:01 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 18:01 sukhe@dns1004: END - running authdns-update * 17:59 sukhe@dns1004: START - running authdns-update * 17:58 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 17:57 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=dns6002.* [reason: depooling for trixie upgrade] * 17:56 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns6002.wikimedia.org * 17:56 cdobbins@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns6002.wikimedia.org * 17:51 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 17:51 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 17:34 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host stat1011.eqiad.wmnet with OS bookworm * 17:31 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns6002.wikimedia.org with OS trixie * 17:30 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 17:30 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 17:30 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 17:30 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 17:29 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 17:29 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 17:29 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 17:29 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:29 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:27 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:24 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327590{{!}}Make sure fpsmax is an int value (T318419)]] (duration: 08m 37s) * 17:20 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 17:17 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327590{{!}}Make sure fpsmax is an int value (T318419)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:16 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 17:15 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327590{{!}}Make sure fpsmax is an int value (T318419)]] * 16:51 swfrench-wmf: disable-puppet on A:cp for ATS Lua change - [[phab:T427666|T427666]] * 16:51 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db2901.codfw.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 16:43 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 16:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on stat1011.eqiad.wmnet with reason: host reimage * 16:39 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns6002.wikimedia.org with reason: host reimage * 16:36 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on stat1011.eqiad.wmnet with reason: host reimage * 16:36 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db2901.codfw.wmnet * 16:34 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns6002.wikimedia.org with reason: host reimage * 16:15 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns6002.wikimedia.org with OS trixie * 16:14 cdobbins@cumin1003: conftool action : set/pooled=no; selector: name=dns6002.* [reason: depooling for trixie upgrade] * 16:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host stat1011 * 16:10 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host stat1011 * 16:09 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host stat1011 * 16:09 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) stat1011.eqiad.wmnet 14.36.64.10.in-addr.arpa 4.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 bking@cumin2003: START - Cookbook sre.dns.wipe-cache stat1011.eqiad.wmnet 14.36.64.10.in-addr.arpa 4.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:09 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host stat1011 - bking@cumin2003" * 16:09 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host stat1011 - bking@cumin2003" * 16:05 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host stat1011 * 16:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host stat1011.eqiad.wmnet with OS bookworm * 16:00 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 15:59 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 15:56 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 15:55 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 15:37 jayme@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:35 jayme@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 15:35 jayme@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:32 jayme@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:32 jayme@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:30 jayme@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 15:30 jayme@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:28 jayme@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:28 jayme@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 15:28 fceratto@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host db1903.eqiad.wmnet * 15:28 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1903.eqiad.wmnet with OS trixie * 15:26 jayme@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 15:26 jayme@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 15:24 jayme@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 15:24 jayme@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 15:21 jayme@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 15:21 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 15:19 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 15:19 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 15:17 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 15:14 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1903.eqiad.wmnet with reason: host reimage * 15:07 fceratto@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1903.eqiad.wmnet with reason: host reimage * 14:54 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db1903.eqiad.wmnet with OS trixie * 14:53 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1903.eqiad.wmnet - fceratto@cumin1003" * 14:53 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1903.eqiad.wmnet - fceratto@cumin1003" * 14:53 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1903.eqiad.wmnet on all recursors * 14:53 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1903.eqiad.wmnet on all recursors * 14:53 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:53 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1903.eqiad.wmnet - fceratto@cumin1003" * 14:53 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1903.eqiad.wmnet - fceratto@cumin1003" * 14:49 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 14:49 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1903.eqiad.wmnet * 14:33 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=cp1100.* * 14:27 topranks: reconfigure eqiad<->codfw bgp settings * 14:22 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:22 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update entries used on new transport backup eqiad codfw - cmooney@cumin1003" * 14:19 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update entries used on new transport backup eqiad codfw - cmooney@cumin1003" * 14:14 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 14:14 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 14:13 moritzm: installing util-linux security updates * 14:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-staging-worker * 14:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2003.codfw.wmnet * 14:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2003.codfw.wmnet * 14:08 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2003.codfw.wmnet * 14:06 moritzm: installing libheif security updates * 13:58 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2003.codfw.wmnet * 13:58 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2002.codfw.wmnet * 13:58 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2002.codfw.wmnet * 13:56 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:56 fnegri@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for clouddb1025.eqiad.wmnet * 13:56 fnegri@cumin1003: START - Cookbook sre.hosts.remove-downtime for clouddb1025.eqiad.wmnet * 13:56 Lucas_WMDE: UTC afternoon backport+config window done * 13:53 moritzm: installing apr-util security updates * 13:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2002.codfw.wmnet * 13:50 fnegri@cumin1003: conftool action : set/weight=100; selector: name=clouddb1025.eqiad.wmnet * 13:49 fnegri@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1025.eqiad.wmnet * 13:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2002.codfw.wmnet * 13:41 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2001.codfw.wmnet * 13:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2001.codfw.wmnet * 13:41 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'. * 13:38 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'. * 13:38 fnegri@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on clouddb1025.eqiad.wmnet with reason: Removing s6 from clouddb1025 * 13:34 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2001.codfw.wmnet * 13:31 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'. * 13:29 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'. * 13:28 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet * 13:26 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host stat1009.eqiad.wmnet with OS bookworm * 13:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2001.codfw.wmnet * 13:24 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-staging-worker * 13:23 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1002.eqiad.wmnet * 13:21 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host stat1010.eqiad.wmnet with OS bookworm * 13:20 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1002.eqiad.wmnet * 13:20 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1001.eqiad.wmnet * 13:17 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1001.eqiad.wmnet * 13:16 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2001.codfw.wmnet * 13:13 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327511{{!}}UIC: Fix page:page instead of page:other in instrumentation]] (duration: 07m 00s) * 13:13 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2001.codfw.wmnet * 13:12 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2002.codfw.wmnet * 13:09 mszwarc@deploy1003: mszwarc: Continuing with deployment * 13:08 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1327511{{!}}UIC: Fix page:page instead of page:other in instrumentation]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2002.codfw.wmnet * 13:07 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2002.codfw.wmnet * 13:06 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1327511{{!}}UIC: Fix page:page instead of page:other in instrumentation]] * 13:04 jmm@dns1004: END - running authdns-update * 13:03 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2002.codfw.wmnet * 13:03 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2001.codfw.wmnet * 13:02 jmm@dns1004: START - running authdns-update * 13:01 cmooney@dns3003: END - running authdns-update * 13:00 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2001.codfw.wmnet * 12:59 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2001.codfw.wmnet * 12:59 cmooney@dns3003: START - running authdns-update * 12:57 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2001.codfw.wmnet * 12:56 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2002.codfw.wmnet * 12:55 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:55 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on drmrs<->eqiad GTT vpls - cmooney@cumin1003" * 12:54 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on drmrs<->eqiad GTT vpls - cmooney@cumin1003" * 12:54 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2002.codfw.wmnet * 12:54 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2003.codfw.wmnet * 12:50 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2003.codfw.wmnet * 12:49 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 12:48 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2001.codfw.wmnet * 12:46 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2001.codfw.wmnet * 12:46 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2002.codfw.wmnet * 12:43 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2002.codfw.wmnet * 12:43 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2003.codfw.wmnet * 12:42 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=1) for new host db1902.eqiad.wmnet * 12:42 fceratto@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host db1902.eqiad.wmnet with OS trixie * 12:41 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2003.codfw.wmnet * 12:40 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1003.eqiad.wmnet * 12:38 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1003.eqiad.wmnet * 12:37 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1002.eqiad.wmnet * 12:35 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1002.eqiad.wmnet * 12:35 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1001.eqiad.wmnet * 12:34 cmooney@dns3003: END - running authdns-update * 12:33 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1001.eqiad.wmnet * 12:32 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on stat1009.eqiad.wmnet with reason: host reimage * 12:31 cmooney@dns3003: START - running authdns-update * 12:31 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:31 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on drmrs<->eqiad cct - cmooney@cumin1003" * 12:28 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on drmrs<->eqiad cct - cmooney@cumin1003" * 12:28 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1902.eqiad.wmnet with reason: host reimage * 12:25 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on stat1009.eqiad.wmnet with reason: host reimage * 12:25 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 12:24 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on stat1010.eqiad.wmnet with reason: host reimage * 12:22 fceratto@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1902.eqiad.wmnet with reason: host reimage * 12:21 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on stat1010.eqiad.wmnet with reason: host reimage * 12:14 elukey: move the Docker Registry's /v2/dev/.* prefix to its dedicated S3 backend - [[phab:T432829|T432829]] * 12:12 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db1902.eqiad.wmnet with OS trixie * 12:09 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1902.eqiad.wmnet - fceratto@cumin1003" * 12:09 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1902.eqiad.wmnet - fceratto@cumin1003" * 12:09 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1902.eqiad.wmnet on all recursors * 12:09 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1902.eqiad.wmnet on all recursors * 12:09 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:08 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1902.eqiad.wmnet - fceratto@cumin1003" * 12:08 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1902.eqiad.wmnet - fceratto@cumin1003" * 12:08 tgr_: [[phab:T413390|T413390]] running CentralAuth:FixRenamedUserGlobalEditCount --wiki=metawiki --since=20250901000000 --until=20260301000000 --fix * 12:04 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1009.eqiad.wmnet with OS bookworm * 12:04 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1010.eqiad.wmnet with OS bookworm * 12:01 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 12:01 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1902.eqiad.wmnet * 12:00 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host stat1010.eqiad.wmnet with OS bookworm * 11:57 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1902.eqiad.wmnet * 11:57 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:57 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1902.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 11:57 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1902.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 11:51 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'. * 11:49 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'. * 11:48 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'. * 11:46 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'. * 11:39 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 11:37 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327518{{!}}Enable thumb.wikimedia.org on cswiki and fawiki (T427465)]] (duration: 10m 40s) * 11:35 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1902.eqiad.wmnet * 11:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1010.eqiad.wmnet with OS bookworm * 11:33 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 11:30 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327518{{!}}Enable thumb.wikimedia.org on cswiki and fawiki (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:26 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327518{{!}}Enable thumb.wikimedia.org on cswiki and fawiki (T427465)]] * 11:24 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host stat1010.eqiad.wmnet with OS bookworm * 11:07 fceratto@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host db1901.eqiad.wmnet * 11:07 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1901.eqiad.wmnet with OS trixie * 10:53 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1901.eqiad.wmnet with reason: host reimage * 10:47 tappof: bump space for prometheus k8s-aux in codfw * 10:47 tappof: bump space for prometheus k8s-dse in eqiad * 10:47 fceratto@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1901.eqiad.wmnet with reason: host reimage * 10:35 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db1901.eqiad.wmnet with OS trixie * 10:32 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:32 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:32 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1901.eqiad.wmnet on all recursors * 10:32 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1901.eqiad.wmnet on all recursors * 10:31 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:31 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:31 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:27 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:27 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1901.eqiad.wmnet * 10:23 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1010.eqiad.wmnet with OS bookworm * 10:20 fceratto@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts db1901.eqiad.wmnet * 10:20 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 10:18 blake@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 10:17 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:16 blake@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 10:13 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1901.eqiad.wmnet * 09:23 jelto@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'. * 09:22 jelto@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'. * 09:22 jelto: update cert-manager to 1.19.6 on wikikube staging-eqiad - [[phab:T427402|T427402]] * 09:20 moritzm: imported squid 7.6-2.1for trixie-wikimedia/main [[phab:T427282|T427282]] * 09:08 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 09:08 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 09:08 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 09:07 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 09:04 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 09:04 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:23 slyngshede@dns1004: END - running authdns-update * 08:21 slyngshede@dns1004: START - running authdns-update * 08:18 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.16 refs [[phab:T430835|T430835]] * 06:27 aokoth@dns1004: END - running authdns-update * 06:25 aokoth@dns1004: START - running authdns-update * 06:22 brennen@deploy1003: Finished deploy [phabricator/deployment@6b9b6ff]: deploy phab1005 for [[phab:T435087|T435087]] (duration: 00m 39s) * 06:21 brennen@deploy1003: Started deploy [phabricator/deployment@6b9b6ff]: deploy phab1005 for [[phab:T435087|T435087]] * 06:20 brennen@deploy1003: Finished deploy [phabricator/deployment@6b9b6ff]: deploy phab1004 for to pick up config values for [[phab:T435087|T435087]] (duration: 01m 46s) * 06:18 brennen@deploy1003: Started deploy [phabricator/deployment@6b9b6ff]: deploy phab1004 for to pick up config values for [[phab:T435087|T435087]] * 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 49s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-19 == * 23:19 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327197{{!}}Enable thumb.wikimedia.org on mediawiki.org (T427465)]] (duration: 10m 50s) * 23:18 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:16 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 23:15 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 23:10 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327197{{!}}Enable thumb.wikimedia.org on mediawiki.org (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:10 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:09 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 23:09 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:09 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 23:08 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327197{{!}}Enable thumb.wikimedia.org on mediawiki.org (T427465)]] * 22:58 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327201{{!}}Enable ReadingLists for all logged in users on test wiki (T435258)]] (duration: 11m 20s) * 22:50 jdlrobson@deploy1003: jdlrobson: Continuing with deployment * 22:49 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1327201{{!}}Enable ReadingLists for all logged in users on test wiki (T435258)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:46 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1327201{{!}}Enable ReadingLists for all logged in users on test wiki (T435258)]] * 22:42 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327162{{!}}Article: Split subjectpageheader by model and disable for wikitext]], [[gerrit:1327169{{!}}Make uppercase greek letters normal (non-italic) font (T434686 T434428)]], [[gerrit:1327176{{!}}Skin: Avoid DB lookup for pagecategorieslink message (T347123)]] (duration: 37m 52s) * 22:29 krinkle@deploy1003: krinkle: Continuing with deployment * 22:25 krinkle@deploy1003: krinkle: Backport for [[gerrit:1327162{{!}}Article: Split subjectpageheader by model and disable for wikitext]], [[gerrit:1327169{{!}}Make uppercase greek letters normal (non-italic) font (T434686 T434428)]], [[gerrit:1327176{{!}}Skin: Avoid DB lookup for pagecategorieslink message (T347123)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:04 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1327162{{!}}Article: Split subjectpageheader by model and disable for wikitext]], [[gerrit:1327169{{!}}Make uppercase greek letters normal (non-italic) font (T434686 T434428)]], [[gerrit:1327176{{!}}Skin: Avoid DB lookup for pagecategorieslink message (T347123)]] * 22:04 krinkle@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: awaiting CI (duration: 03m 06s) * 22:01 krinkle@deploy1003: Locking from deployment [ALL REPOSITORIES]: awaiting CI * 22:00 krinkle@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: awaiting CI (duration: 00m 01s) * 22:00 krinkle@deploy1003: Locking from deployment [ALL REPOSITORIES]: awaiting CI * 21:34 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2001.codfw.wmnet * 21:28 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2001.codfw.wmnet * 21:22 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327178{{!}}AccountRecovery: Notify the email address of the on file of the request (T425799)]] (duration: 47m 02s) * 21:13 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm * 21:09 catrope@deploy1003: catrope: Continuing with deployment * 20:55 catrope@deploy1003: catrope: Backport for [[gerrit:1327178{{!}}AccountRecovery: Notify the email address of the on file of the request (T425799)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:35 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1327178{{!}}AccountRecovery: Notify the email address of the on file of the request (T425799)]] * 20:31 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327128{{!}}Parsoid DataAccess: convert Parsoid fragment markers to/from strip tags (T432547)]] (duration: 07m 30s) * 20:27 catrope@deploy1003: catrope, arlolra: Continuing with deployment * 20:26 catrope@deploy1003: catrope, arlolra: Backport for [[gerrit:1327128{{!}}Parsoid DataAccess: convert Parsoid fragment markers to/from strip tags (T432547)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:24 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1327128{{!}}Parsoid DataAccess: convert Parsoid fragment markers to/from strip tags (T432547)]] * 20:23 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage * 20:17 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage * 20:15 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325878{{!}}[arwiki] Enable restricted user page editing and grant edit permissions (T434878)]] (duration: 08m 46s) * 20:11 catrope@deploy1003: catrope, gergesshamon: Continuing with deployment * 20:08 catrope@deploy1003: catrope, gergesshamon: Backport for [[gerrit:1325878{{!}}[arwiki] Enable restricted user page editing and grant edit permissions (T434878)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:06 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1325878{{!}}[arwiki] Enable restricted user page editing and grant edit permissions (T434878)]] * 19:59 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm * 19:56 eevans@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cassandra-dev2001.codfw.wmnet with OS bookworm * 19:56 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm * 19:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2207.codfw.wmnet with reason: Maintenance * 18:47 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319920{{!}}Allow setting a separate thumbUrl in production (T427465)]], [[gerrit:1327167{{!}}Fix wmgThumbUrl config (T427465)]] (duration: 18m 53s) * 18:43 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 18:30 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1319920{{!}}Allow setting a separate thumbUrl in production (T427465)]], [[gerrit:1327167{{!}}Fix wmgThumbUrl config (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:28 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1319920{{!}}Allow setting a separate thumbUrl in production (T427465)]], [[gerrit:1327167{{!}}Fix wmgThumbUrl config (T427465)]] * 18:26 sukhe@dns1004: END - running authdns-update * 18:24 sukhe@dns1004: START - running authdns-update * 18:09 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1319920{{!}}Allow setting a separate thumbUrl in production (T427465)]] * 18:03 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-eqiad and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 17:56 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-codfw and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 17:50 cmooney@dns3003: END - running authdns-update * 17:42 dancy@deploy1003: Installation of scap version "4.283.0" completed for 3 hosts * 17:41 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling reboot on A:durum and A:durum * 17:40 cmooney@dns3003: START - running authdns-update * 17:40 dancy@deploy1003: Installing scap version "4.283.0" for 3 host(s) * 17:38 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:37 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on GTT VPLS - cmooney@cumin1003" * 17:37 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --verbose --scoreLessThan=0.7 --exceptDatasetChecksums=[[phab:T434319|T434319]]-enwiki-models.txt # [[phab:T434319|T434319]] * 17:32 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on GTT VPLS - cmooney@cumin1003" * 17:29 sbassett: Deployed security fix for [[phab:T435210|T435210]] (wmf.16) * 17:26 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 17:22 sbassett: Deployed security fix for [[phab:T435210|T435210]] (wmf.15) * 17:00 sukhe@dns1004: END - running authdns-update * 16:58 sukhe@dns1004: START - running authdns-update * 16:53 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica-esams and A:liberica * 16:41 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica-esams and A:liberica * 16:41 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-eqiad and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 16:41 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-codfw and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 16:41 cjd91: sudo -i cookbook sre.cdn.roll-upgrade-ats --query 'A:cp-codfw' --task-id [[phab:T434478|T434478]] --reason '9.2.15 upgrade' * 16:41 cjd91: sudo -i cookbook sre.cdn.roll-upgrade-ats --query 'A:cp-eqiad' --task-id [[phab:T434478|T434478]] --reason '9.2.15 upgrade' * 16:40 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and A:durum * 16:28 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326881{{!}}Echo: Start using virtual domains (T380385)]] (duration: 13m 12s) * 16:23 urbanecm@deploy1003: urbanecm: Continuing with deployment * 16:21 urandom: Completed sessionstore Cassandra/JVM upgrade — [[phab:T435154|T435154]] * 16:21 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching sessionstore[2005-2006].codfw.wmnet,sessionstore[1005-1006].eqiad.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 16:19 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1326881{{!}}Echo: Start using virtual domains (T380385)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:15 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:15 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->codfw - cmooney@cumin1003" * 16:14 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1326881{{!}}Echo: Start using virtual domains (T380385)]] * 16:14 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327123{{!}}Revert^2 "Migrate database access to virtual domains" (T435305)]], [[gerrit:1327124{{!}}Pass the mapped domain of virtual-echo-shared to the push NameTableStores (T435305)]] (duration: 07m 42s) * 16:13 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching sessionstore[2005-2006].codfw.wmnet,sessionstore[1005-1006].eqiad.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 16:11 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->codfw - cmooney@cumin1003" * 16:10 urbanecm@deploy1003: urbanecm: Continuing with deployment * 16:08 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1327123{{!}}Revert^2 "Migrate database access to virtual domains" (T435305)]], [[gerrit:1327124{{!}}Pass the mapped domain of virtual-echo-shared to the push NameTableStores (T435305)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:06 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 16:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 16:06 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1327123{{!}}Revert^2 "Migrate database access to virtual domains" (T435305)]], [[gerrit:1327124{{!}}Pass the mapped domain of virtual-echo-shared to the push NameTableStores (T435305)]] * 16:06 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:03 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching sessionstore1004.eqiad.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 16:01 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching sessionstore1004.eqiad.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 15:56 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching sessionstore2004.codfw.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 15:54 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching sessionstore2004.codfw.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 15:53 urandom: beginning sessionstore Cassandra/JVM upgrade — [[phab:T435154|T435154]] * 15:52 urandom: beginning sessionstore Cassandra/JVM upgrade — [[phab:T432944|T432944]] * 15:51 cmooney@dns3003: END - running authdns-update * 15:49 cmooney@dns3003: START - running authdns-update * 15:48 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:48 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->codfw - cmooney@cumin1003" * 15:45 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->codfw - cmooney@cumin1003" * 15:44 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 15:44 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:43 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 15:42 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:42 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:38 cmooney@dns3003: END - running authdns-update * 15:36 cmooney@dns3003: START - running authdns-update * 15:36 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:36 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->eqsin - cmooney@cumin1003" * 15:34 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327138{{!}}Enable redis lock manager everywhere (T366938)]] (duration: 08m 36s) * 15:33 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->eqsin - cmooney@cumin1003" * 15:30 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:29 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 15:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1147.eqiad.wmnet with OS bookworm * 15:28 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327138{{!}}Enable redis lock manager everywhere (T366938)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:25 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327138{{!}}Enable redis lock manager everywhere (T366938)]] * 15:24 jmm@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host krb1002.eqiad.wmnet * 15:19 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2207.codfw.wmnet with reason: Host crashed * 15:17 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327098{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]], [[gerrit:1327101{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]] (duration: 07m 13s) * 15:12 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 15:12 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1327098{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]], [[gerrit:1327101{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:10 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1327098{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]], [[gerrit:1327101{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]] * 15:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2010.codfw.wmnet with OS trixie * 15:05 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 15:05 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1147.eqiad.wmnet with reason: host reimage * 14:59 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 14:59 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:58 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1147.eqiad.wmnet with reason: host reimage * 14:55 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:55 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:49 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 14:48 cmooney@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host durum1001.eqiad.wmnet * 14:46 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:46 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Delete db2902 ipv6 addr - fceratto@cumin1003" * 14:46 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Delete db2902 ipv6 addr - fceratto@cumin1003" * 14:43 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1147.eqiad.wmnet with OS bookworm * 14:42 tgr@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327094{{!}}SpecialMWOAuthListConsumers: Handle newFromMWUser returning null in addNavigationSubtitle (T435167)]] (duration: 19m 25s) * 14:42 cmooney@cumin1003: START - Cookbook sre.hosts.reboot-single for host durum1001.eqiad.wmnet * 14:42 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 14:41 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 14:38 tgr@deploy1003: tgr: Continuing with deployment * 14:36 tgr@deploy1003: tgr: Backport for [[gerrit:1327094{{!}}SpecialMWOAuthListConsumers: Handle newFromMWUser returning null in addNavigationSubtitle (T435167)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:28 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 14:28 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:28 cmooney@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host durum3005.esams.wmnet * 14:25 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 14:23 cmooney@cumin1003: START - Cookbook sre.hosts.reboot-single for host durum3005.esams.wmnet * 14:23 tgr@deploy1003: Started scap sync-world: Backport for [[gerrit:1327094{{!}}SpecialMWOAuthListConsumers: Handle newFromMWUser returning null in addNavigationSubtitle (T435167)]] * 14:18 gengh@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:17 topranks: disable puppet on hosts running BIRD BGP to test merge of patch to systemd healtchcheck service * 14:17 gengh@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:17 gengh@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:16 gengh@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:16 elukey: upgrade spicerack on cumin1003 and cumin2003 * 14:16 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:15 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314962{{!}}static: add new dir bimi/ for BIMI SVG and PEM file (T311685)]] (duration: 10m 00s) * 14:15 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:11 kharlan@deploy1003: kharlan, sukhe: Continuing with deployment * 14:11 gengh@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:09 gengh@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:09 gengh@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:08 kharlan@deploy1003: kharlan, sukhe: Backport for [[gerrit:1314962{{!}}static: add new dir bimi/ for BIMI SVG and PEM file (T311685)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:06 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 14:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:06 gengh@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:05 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1314962{{!}}static: add new dir bimi/ for BIMI SVG and PEM file (T311685)]] * 14:05 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host krb1002.eqiad.wmnet * 14:05 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:04 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:04 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326342{{!}}Revert^2 "wmf-config/ProductionServices: set URL for urldownloader to service record"]] (duration: 07m 40s) * 13:59 kharlan@deploy1003: kharlan, sukhe: Continuing with deployment * 13:58 kharlan@deploy1003: kharlan, sukhe: Backport for [[gerrit:1326342{{!}}Revert^2 "wmf-config/ProductionServices: set URL for urldownloader to service record"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:57 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host stat1008.eqiad.wmnet with OS bookworm * 13:56 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1326342{{!}}Revert^2 "wmf-config/ProductionServices: set URL for urldownloader to service record"]] * 13:56 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 13:54 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325532{{!}}srwiki: Allow bureaucrats to add and remove event-organizer group (T434748)]] (duration: 14m 56s) * 13:54 swfrench@dns1004: END - running authdns-update * 13:52 swfrench@dns1004: START - running authdns-update * 13:51 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:51 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 13:48 kharlan@deploy1003: kharlan, danielyepezgarces: Continuing with deployment * 13:45 swfrench@cumin2003: conftool action : set/pooled=yes; selector: name=wikikube-worker2330.codfw.wmnet * 13:44 swfrench-wmf: finished etcd-main codfw -> eqiad switchover - [[phab:T435103|T435103]] * 13:44 kharlan@deploy1003: kharlan, danielyepezgarces: Backport for [[gerrit:1325532{{!}}srwiki: Allow bureaucrats to add and remove event-organizer group (T434748)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:44 swfrench@cumin2003: conftool action : set/pooled=no; selector: name=wikikube-worker2330.codfw.wmnet * 13:41 swfrench@dns1004: END - running authdns-update * 13:39 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1325532{{!}}srwiki: Allow bureaucrats to add and remove event-organizer group (T434748)]] * 13:39 swfrench@dns1004: START - running authdns-update * 13:37 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327111{{!}}Special:AbuseReview: Add "no further action needed" review action (T435020)]], [[gerrit:1327110{{!}}AbuseReview: Take the review verdict as a REST path parameter (T435020)]] (duration: 31m 43s) * 13:31 swfrench-wmf: starting etcd-main codfw -> eqiad switchover - [[phab:T435103|T435103]] * 13:28 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:28 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 13:25 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=97) rolling reboot on A:durum and A:durum * 13:24 kharlan@deploy1003: kharlan: Continuing with deployment * 13:24 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1147 * 13:24 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1147 * 13:23 kharlan@deploy1003: kharlan: Backport for [[gerrit:1327111{{!}}Special:AbuseReview: Add "no further action needed" review action (T435020)]], [[gerrit:1327110{{!}}AbuseReview: Take the review verdict as a REST path parameter (T435020)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:16 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:16 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 13:12 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:12 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 13:10 cdobbins@cumin1003: END (ERROR) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=97) Rolling upgrade of ATS on A:cp-codfw and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 13:10 cdobbins@cumin1003: END (ERROR) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=97) Rolling upgrade of ATS on A:cp-eqiad and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 13:06 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1327111{{!}}Special:AbuseReview: Add "no further action needed" review action (T435020)]], [[gerrit:1327110{{!}}AbuseReview: Take the review verdict as a REST path parameter (T435020)]] * 13:06 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-eqiad and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 13:05 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-codfw and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 13:05 cjd91: sudo -i cookbook sre.cdn.roll-upgrade-ats --query 'A:cp-eqiad' --task-id [[phab:T434478|T434478]] --reason '9.2.15 upgrade' * 13:03 swfrench@dns1004: END - running authdns-update * 13:01 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:01 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 13:01 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host ncmonitor1001.eqiad.wmnet * 13:00 swfrench@dns1004: START - running authdns-update * 12:59 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 12:59 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 12:59 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 12:59 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 12:57 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and A:durum * 12:53 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on db2902.codfw.wmnet with reason: Cloning * 12:48 cmooney@dns3003: END - running authdns-update * 12:48 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327096{{!}}Switch to redis lock manager on s4 and s8 (T366938)]] (duration: 09m 11s) * 12:46 cmooney@dns3003: START - running authdns-update * 12:45 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:45 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->eqord cct - cmooney@cumin1003" * 12:45 elukey: move the /v2/releng.* prefix on the Docker Registry to its new s3 backend - [[phab:T432829|T432829]] * 12:43 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 12:42 jelto@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 12:42 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->eqord cct - cmooney@cumin1003" * 12:42 jelto@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 12:41 jelto: update cert-manager to 1.19.6 on wikikube staging-codfw - [[phab:T427402|T427402]] * 12:40 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327096{{!}}Switch to redis lock manager on s4 and s8 (T366938)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:38 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327096{{!}}Switch to redis lock manager on s4 and s8 (T366938)]] * 12:38 blake@deploy1003: Finished scap sync-world: non-build deploy for [[phab:T417800|T417800]] (duration: 03m 52s) * 12:36 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 12:35 blake@deploy1003: Started scap sync-world: non-build deploy for [[phab:T417800|T417800]] * 12:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host krb2002.codfw.wmnet * 11:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host krb2002.codfw.wmnet * 11:49 moritzm: installing kerberos security updates * 11:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on stat1008.eqiad.wmnet with reason: host reimage * 11:44 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on stat1008.eqiad.wmnet with reason: host reimage * 11:31 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327084{{!}}Revert "Migrate database access to virtual domains" (T435305)]] (duration: 11m 02s) * 11:29 gkyziridis@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:29 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:29 kart_: Updated MinT to 2026-06-04-131507-production ([[phab:T321316|T321316]]) * 11:28 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/machinetranslation: apply * 11:28 gkyziridis@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:26 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 11:24 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:24 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:23 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:23 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:23 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/machinetranslation: apply * 11:22 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327084{{!}}Revert "Migrate database access to virtual domains" (T435305)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:21 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/machinetranslation: apply * 11:21 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:21 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:20 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327084{{!}}Revert "Migrate database access to virtual domains" (T435305)]] * 11:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1008.eqiad.wmnet with OS bookworm * 11:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet * 11:17 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/machinetranslation: apply * 11:13 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/machinetranslation: apply * 11:12 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:12 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:10 kartik@deploy1003: helmfile [staging] START helmfile.d/services/machinetranslation: apply * 11:08 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-cron: apply * 11:08 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/mw-cron: apply * 11:08 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply * 11:08 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply * 11:07 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet * 11:07 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host stat1008.eqiad.wmnet with OS bookworm * 11:06 moritzm: upgrading the new trixie URL downloaders to Squid 7.6 [[phab:T427282|T427282]] * 11:01 gkyziridis@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin2002.codfw.wmnet * 10:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin2002.codfw.wmnet * 10:45 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1169.eqiad.wmnet with OS bookworm * 10:42 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1185.eqiad.wmnet with OS bookworm * 10:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1169.eqiad.wmnet with reason: host reimage * 10:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1185.eqiad.wmnet with reason: host reimage * 10:14 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1169.eqiad.wmnet with reason: host reimage * 10:14 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1185.eqiad.wmnet with reason: host reimage * 10:11 jmm@cumin2003: END (PASS) - Cookbook sre.netbox.restart-reboot (exit_code=0) rolling reboot on A:netbox * 10:06 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1008.eqiad.wmnet with OS bookworm * 10:04 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 10:04 mpostoronca@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321579{{!}}Register the mediawiki.wikimedia_antiabuse.content_policy_score stream (T432848)]] (duration: 08m 53s) * 10:03 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 10:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 10:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 10:00 mpostoronca@deploy1003: mpostoronca: Continuing with deployment * 09:59 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1185.eqiad.wmnet with OS bookworm * 09:59 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1169.eqiad.wmnet with OS bookworm * 09:59 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.convert-disks (exit_code=0) for host ms-be1065 * 09:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1065.eqiad.wmnet with OS trixie * 09:58 mpostoronca@deploy1003: mpostoronca: Backport for [[gerrit:1321579{{!}}Register the mediawiki.wikimedia_antiabuse.content_policy_score stream (T432848)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:55 jmm@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netbox.discovery.wmnet. on all recursors * 09:55 mpostoronca@deploy1003: Started scap sync-world: Backport for [[gerrit:1321579{{!}}Register the mediawiki.wikimedia_antiabuse.content_policy_score stream (T432848)]] * 09:55 jmm@cumin2003: START - Cookbook sre.dns.wipe-cache netbox.discovery.wmnet. on all recursors * 09:52 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw2001.wikimedia.org with OS trixie * 09:51 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 09:51 jmm@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netbox.discovery.wmnet. on all recursors * 09:51 jmm@cumin2003: START - Cookbook sre.dns.wipe-cache netbox.discovery.wmnet. on all recursors * 09:51 jmm@cumin2003: START - Cookbook sre.netbox.restart-reboot rolling reboot on A:netbox * 09:46 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 09:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1056.eqiad.wmnet with OS trixie * 09:44 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 09:39 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 09:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 09:36 topranks: make HE transport circuits from magru live * 09:36 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 09:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb1003.eqiad.wmnet * 09:33 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage * 09:31 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb1003.eqiad.wmnet * 09:30 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 09:28 arnaudb@dns1006: END - running authdns-update * 09:27 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage * 09:27 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb2003.codfw.wmnet * 09:26 arnaudb@dns1006: START - running authdns-update * 09:24 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1056.eqiad.wmnet with reason: host reimage * 09:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb2003.codfw.wmnet * 09:20 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.convert-disks (exit_code=0) for host ms-be1068 * 09:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1068.eqiad.wmnet with OS trixie * 09:20 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 09:19 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "cloudvirt1057 - filippo@cumin1003" * 09:19 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "cloudvirt1057 - filippo@cumin1003" * 09:18 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1056.eqiad.wmnet with reason: host reimage * 09:18 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1057.eqiad.wmnet with OS trixie * 09:18 filippo@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 09:18 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 09:15 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 09:14 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 09:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host irc1003.wikimedia.org * 09:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:13 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1065.eqiad.wmnet with OS trixie * 09:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 09:09 moritzm: installing Postgresql security updates * 09:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host irc1003.wikimedia.org * 09:07 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 09:07 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:07 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1173.eqiad.wmnet with OS bookworm * 09:06 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw2001.wikimedia.org with OS trixie * 09:03 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.convert-disks (exit_code=0) for host ms-be1064 * 09:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1064.eqiad.wmnet with OS trixie * 09:03 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1056.eqiad.wmnet with OS trixie * 09:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1057.eqiad.wmnet with reason: host reimage * 09:01 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 08:58 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 08:56 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1057.eqiad.wmnet with reason: host reimage * 08:54 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1208.eqiad.wmnet with OS bookworm * 08:53 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2902.codfw.wmnet with OS trixie * 08:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 08:50 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1174.eqiad.wmnet with OS bookworm * 08:46 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1172.eqiad.wmnet with OS bookworm * 08:45 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1173.eqiad.wmnet with reason: host reimage * 08:41 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 08:40 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1057.eqiad.wmnet with OS trixie * 08:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1057.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 08:39 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1207.eqiad.wmnet with OS bookworm * 08:38 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2902.codfw.wmnet with reason: host reimage * 08:37 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 08:34 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1068.eqiad.wmnet with OS trixie * 08:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1208.eqiad.wmnet with reason: host reimage * 08:31 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1057.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 08:29 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 08:28 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1222.eqiad.wmnet onto db1276.eqiad.wmnet * 08:28 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1222: Pool db1222.eqiad.wmnet in after cloning * 08:28 fceratto@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2902.codfw.wmnet with reason: host reimage * 08:27 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1055.eqiad.wmnet with OS trixie * 08:27 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 08:26 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1174.eqiad.wmnet with reason: host reimage * 08:25 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 08:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-misc2002.codfw.wmnet * 08:23 topranks: reboot pfw1-codfw firewall pair to upgrade JunOS [[phab:T434865|T434865]] * 08:22 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1172.eqiad.wmnet with reason: host reimage * 08:20 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1064.eqiad.wmnet with OS trixie * 08:20 mvernon@cumin2003: START - Cookbook sre.swift.convert-disks for host ms-be1065 * 08:18 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1207.eqiad.wmnet with reason: host reimage * 08:17 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1174.eqiad.wmnet with reason: host reimage * 08:17 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1173.eqiad.wmnet with reason: host reimage * 08:17 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1172.eqiad.wmnet with reason: host reimage * 08:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host mc-misc2002.codfw.wmnet * 08:15 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1208.eqiad.wmnet with reason: host reimage * 08:15 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1207.eqiad.wmnet with reason: host reimage * 08:14 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db2902.codfw.wmnet with OS trixie * 08:14 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.16 refs [[phab:T430835|T430835]] * 08:13 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db2902.codfw.wmnet * 08:13 fceratto@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host db2902.codfw.wmnet with OS trixie * 08:10 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr[1-2]-codfw with reason: upgrade pfw1a-codfw and pfw1b-codfw pair * 08:09 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1055.eqiad.wmnet with reason: host reimage * 08:07 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on pfw1-codfw with reason: upgrade pfw1a-codfw and pfw1b-codfw pair * 08:03 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1055.eqiad.wmnet with reason: host reimage * 08:02 arnaudb@dns1006: END - running authdns-update * 08:02 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1208.eqiad.wmnet with OS bookworm * 08:02 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1207.eqiad.wmnet with OS bookworm * 08:01 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1174.eqiad.wmnet with OS bookworm * 08:01 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1173.eqiad.wmnet with OS bookworm * 08:01 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1172.eqiad.wmnet with OS bookworm * 07:59 arnaudb@dns1006: START - running authdns-update * 07:58 arnaudb@dns1006: START - running authdns-update * 07:48 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1055.eqiad.wmnet with OS trixie * 07:45 moritzm: extend the disk of ldap-rw2001 by 80G [[phab:T331699|T331699]] * 07:42 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1222: Pool db1222.eqiad.wmnet in after cloning * 07:36 mvernon@cumin2003: START - Cookbook sre.swift.convert-disks for host ms-be1068 * 07:35 mvernon@cumin2003: START - Cookbook sre.swift.convert-disks for host ms-be1064 * 07:17 moritzm: installing imagemagick security updates * 07:14 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1277: Pool back * 07:14 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon1003.wikimedia.org * 07:07 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon1003.wikimedia.org * 07:03 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1280: Pool back * 07:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon2002.wikimedia.org * 06:55 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon2002.wikimedia.org * 06:54 moritzm: installing php8.2 security updates * 06:51 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1284: Pool back * 06:49 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1222: Depool db1222.eqiad.wmnet to then clone it to db1276.eqiad.wmnet - marostegui@cumin1003 * 06:49 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1222: Depool db1222.eqiad.wmnet to then clone it to db1276.eqiad.wmnet - marostegui@cumin1003 * 06:49 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1222.eqiad.wmnet onto db1276.eqiad.wmnet * 06:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd1005.eqiad.wmnet * 06:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd1005.eqiad.wmnet * 06:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd1004.eqiad.wmnet * 06:36 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2209: db2209 repool * 06:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd1004.eqiad.wmnet * 06:32 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast3007.wikimedia.org * 06:29 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1277: Pool back * 06:28 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1277 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96190 and previous config saved to /var/cache/conftool/dbconfig/20260819-062815-marostegui.json * 06:26 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast3007.wikimedia.org * 06:22 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul1001.eqiad.wmnet * 06:18 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul1001.eqiad.wmnet * 06:18 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul1003.eqiad.wmnet * 06:18 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1280: Pool back * 06:17 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1284 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96186 and previous config saved to /var/cache/conftool/dbconfig/20260819-061743-marostegui.json * 06:14 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul1003.eqiad.wmnet * 06:14 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul1002.eqiad.wmnet * 06:10 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul1002.eqiad.wmnet * 06:10 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2048.codfw.wmnet * 06:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2048.codfw.wmnet * 06:06 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1284: Pool back * 06:06 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1284 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96184 and previous config saved to /var/cache/conftool/dbconfig/20260819-060621-marostegui.json * 06:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2048.codfw.wmnet * 05:59 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2048.codfw.wmnet * 05:51 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2209: db2209 repool * 03:16 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1186.eqiad.wmnet with OS bookworm * 02:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1186.eqiad.wmnet with reason: host reimage * 02:46 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1186.eqiad.wmnet with reason: host reimage * 02:46 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2207 [[phab:T435270|T435270]]', diff saved to https://phabricator.wikimedia.org/P96181 and previous config saved to /var/cache/conftool/dbconfig/20260819-024627-marostegui.json * 02:44 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2204 to s2 primary [[phab:T435270|T435270]]', diff saved to https://phabricator.wikimedia.org/P96180 and previous config saved to /var/cache/conftool/dbconfig/20260819-024403-marostegui.json * 02:43 marostegui: Starting s2 codfw failover from db2207 to db2204 - [[phab:T435270|T435270]] * 02:39 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2204 with weight 0 [[phab:T435270|T435270]]', diff saved to https://phabricator.wikimedia.org/P96179 and previous config saved to /var/cache/conftool/dbconfig/20260819-023951-marostegui.json * 02:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s2 [[phab:T435270|T435270]] * 02:32 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1186.eqiad.wmnet with OS bookworm * 02:29 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-worker1186.eqiad.wmnet with OS bookworm * 02:18 denisse@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2207: Depooling replica * 02:18 denisse@cumin1003: START - Cookbook sre.mysql.depool depool db2207: Depooling replica * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 48s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-18 == * 23:55 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1264845{{!}}Remove unused/redundant wgMFNoindexPages=true setting (T255458)]] (duration: 09m 42s) * 23:51 krinkle@deploy1003: krinkle: Continuing with deployment * 23:48 krinkle@deploy1003: krinkle: Backport for [[gerrit:1264845{{!}}Remove unused/redundant wgMFNoindexPages=true setting (T255458)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:45 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1264845{{!}}Remove unused/redundant wgMFNoindexPages=true setting (T255458)]] * 23:38 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326944{{!}}Retire filebackend lock manager in favour of the default one (T366938)]] (duration: 08m 55s) * 23:34 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 23:31 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326944{{!}}Retire filebackend lock manager in favour of the default one (T366938)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:29 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326944{{!}}Retire filebackend lock manager in favour of the default one (T366938)]] * 23:27 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1170.eqiad.wmnet with OS bookworm * 23:21 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1205.eqiad.wmnet with OS bookworm * 23:20 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1171.eqiad.wmnet with OS bookworm * 23:15 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1206.eqiad.wmnet with OS bookworm * 23:05 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1170.eqiad.wmnet with reason: host reimage * 23:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1205.eqiad.wmnet with reason: host reimage * 22:57 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1171.eqiad.wmnet with reason: host reimage * 22:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1206.eqiad.wmnet with reason: host reimage * 22:53 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1205.eqiad.wmnet with reason: host reimage * 22:51 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1171.eqiad.wmnet with reason: host reimage * 22:51 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1170.eqiad.wmnet with reason: host reimage * 22:50 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1206.eqiad.wmnet with reason: host reimage * 22:36 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1206.eqiad.wmnet with OS bookworm * 22:35 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1205.eqiad.wmnet with OS bookworm * 22:35 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1186.eqiad.wmnet with OS bookworm * 22:35 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1171.eqiad.wmnet with OS bookworm * 22:35 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1170.eqiad.wmnet with OS bookworm * 22:33 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-worker1194.eqiad.wmnet with OS bookworm * 22:22 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326923{{!}}Enable redis lock manager on s6 (T366938)]] (duration: 11m 52s) * 22:18 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 22:13 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326923{{!}}Enable redis lock manager on s6 (T366938)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:10 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326923{{!}}Enable redis lock manager on s6 (T366938)]] * 22:04 sbassett: Deployed security fix for [[phab:T435234|T435234]] (wmf.16) * 21:54 sbassett: Deployed security fix for [[phab:T435234|T435234]] (wmf.15) * 21:38 caro@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326925{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326926{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326929{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]], [[gerrit:1326928{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]] (duration: 0 * 21:34 caro@deploy1003: caro: Continuing with deployment * 21:33 caro@deploy1003: caro: Backport for [[gerrit:1326925{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326926{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326929{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]], [[gerrit:1326928{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]] synced to the testservers (see h * 21:31 caro@deploy1003: Started scap sync-world: Backport for [[gerrit:1326925{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326926{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326929{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]], [[gerrit:1326928{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]] * 21:24 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1204.eqiad.wmnet with reason: 1204 datanode repair [[phab:T434494|T434494]] * 21:02 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326896{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]], [[gerrit:1326897{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]] (duration: 13m 56s) * 20:58 krinkle@deploy1003: krinkle: Continuing with deployment * 20:50 krinkle@deploy1003: krinkle: Backport for [[gerrit:1326896{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]], [[gerrit:1326897{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:49 ryankemper: `an-launcher1003` terminated process group `666809` (`rest_backfill_phase1.sh`) ~20 mins ago with `sudo kill -TERM -- -666809` after its local spark driver (`--driver-memory 64g`) repeatedly exhausted memory on the 32 GB VM and caused SSH to intermittently flap; host recovered to 27 GB available memory * 20:48 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1326896{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]], [[gerrit:1326897{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]] * 20:35 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 20:33 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 20:31 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 20:31 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326870{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]], [[gerrit:1326871{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]] (duration: 07m 35s) * 20:28 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 20:26 kemayo@deploy1003: kemayo: Continuing with deployment * 20:26 ryankemper: `an-launcher1003` confirmed the host is flapping because of memory thrash. chasing down the source of the thrash * 20:25 kemayo@deploy1003: kemayo: Backport for [[gerrit:1326870{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]], [[gerrit:1326871{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:23 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1326870{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]], [[gerrit:1326871{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]] * 20:23 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 20:20 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 20:14 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1168.eqiad.wmnet with OS bookworm * 20:08 zabe: zabe@deploy1003:~$ mwscript extensions/WikimediaMaintenance/maintenance/fixFileRevisionArchiveNameDrift.php enwiki # [[phab:T428406|T428406]] * 20:08 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1204.eqiad.wmnet with OS bookworm * 20:05 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1167.eqiad.wmnet with OS bookworm * 20:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1166.eqiad.wmnet with OS bookworm * 19:54 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1203.eqiad.wmnet with OS bookworm * 19:53 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326907{{!}}Revert "Disable redis lock manager on testwiki"]] (duration: 11m 05s) * 19:50 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1168.eqiad.wmnet with reason: host reimage * 19:47 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1204.eqiad.wmnet with reason: host reimage * 19:46 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 19:44 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326907{{!}}Revert "Disable redis lock manager on testwiki"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:42 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326907{{!}}Revert "Disable redis lock manager on testwiki"]] * 19:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1167.eqiad.wmnet with reason: host reimage * 19:37 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1166.eqiad.wmnet with reason: host reimage * 19:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1203.eqiad.wmnet with reason: host reimage * 19:32 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1167.eqiad.wmnet with reason: host reimage * 19:32 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1168.eqiad.wmnet with reason: host reimage * 19:32 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1166.eqiad.wmnet with reason: host reimage * 19:31 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1204.eqiad.wmnet with reason: host reimage * 19:31 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1203.eqiad.wmnet with reason: host reimage * 19:26 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324751{{!}}InitialiseSettings: Enable 2FA enforcement on more private wikis (T428103)]], [[gerrit:1326875{{!}}Add banner notifying of upcoming 2FA enforcement (T420792)]] (duration: 31m 46s) * 19:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1204.eqiad.wmnet with OS bookworm * 19:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1203.eqiad.wmnet with OS bookworm * 19:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1168.eqiad.wmnet with OS bookworm * 19:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1167.eqiad.wmnet with OS bookworm * 19:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1166.eqiad.wmnet with OS bookworm * 19:15 denisse: rebooting kafkamon2003.codfw.wmnet - [[phab:T435162|T435162]] * 19:14 denisse: rebooting kafkamon1003.eqiad.wmnet [[phab:T435162|T435162]] * 19:13 reedy@deploy1003: reedy: Continuing with deployment * 19:12 reedy@deploy1003: reedy: Backport for [[gerrit:1324751{{!}}InitialiseSettings: Enable 2FA enforcement on more private wikis (T428103)]], [[gerrit:1326875{{!}}Add banner notifying of upcoming 2FA enforcement (T420792)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:54 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324751{{!}}InitialiseSettings: Enable 2FA enforcement on more private wikis (T428103)]], [[gerrit:1326875{{!}}Add banner notifying of upcoming 2FA enforcement (T420792)]] * 18:50 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 18:44 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.16 refs [[phab:T430835|T430835]] * 18:34 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 18:31 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 18:22 aklapper@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326852{{!}}CategoryTree: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]], [[gerrit:1326853{{!}}CategoryViewer: Allow null $html in the CategoryViewerGenerateLink hook (T435161)]], [[gerrit:1326865{{!}}Flow: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]] (duration: 09m 57s) * 18:18 aklapper@deploy1003: jforrester, aklapper: Continuing with deployment * 18:17 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 18:14 aklapper@deploy1003: jforrester, aklapper: Backport for [[gerrit:1326852{{!}}CategoryTree: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]], [[gerrit:1326853{{!}}CategoryViewer: Allow null $html in the CategoryViewerGenerateLink hook (T435161)]], [[gerrit:1326865{{!}}Flow: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki * 18:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1156.eqiad.wmnet with OS bookworm * 18:12 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 18:12 aklapper@deploy1003: Started scap sync-world: Backport for [[gerrit:1326852{{!}}CategoryTree: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]], [[gerrit:1326853{{!}}CategoryViewer: Allow null $html in the CategoryViewerGenerateLink hook (T435161)]], [[gerrit:1326865{{!}}Flow: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]] * 18:11 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 18:08 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1146.eqiad.wmnet with OS bookworm * 18:07 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1177.eqiad.wmnet with OS bookworm * 18:00 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326380{{!}}Introduce main lock manager service (T366938 T427999)]] (duration: 11m 25s) * 17:58 ladsgroup@deploy1003: ladsgroup: Rolling back deployment * 17:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1202.eqiad.wmnet with OS bookworm * 17:55 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1201.eqiad.wmnet with OS bookworm * 17:50 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326380{{!}}Introduce main lock manager service (T366938 T427999)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:48 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326380{{!}}Introduce main lock manager service (T366938 T427999)]] * 17:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1156.eqiad.wmnet with reason: host reimage * 17:46 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1146.eqiad.wmnet with reason: host reimage * 17:45 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-esams and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 17:42 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 17:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1177.eqiad.wmnet with reason: host reimage * 17:38 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1202.eqiad.wmnet with reason: host reimage * 17:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1201.eqiad.wmnet with reason: host reimage * 17:30 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1156.eqiad.wmnet with reason: host reimage * 17:29 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1177.eqiad.wmnet with reason: host reimage * 17:28 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1146.eqiad.wmnet with reason: host reimage * 17:28 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1202.eqiad.wmnet with reason: host reimage * 17:27 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1201.eqiad.wmnet with reason: host reimage * 17:25 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 17:21 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 17:14 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1202.eqiad.wmnet with OS bookworm * 17:14 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1201.eqiad.wmnet with OS bookworm * 17:14 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1177.eqiad.wmnet with OS bookworm * 17:14 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1156.eqiad.wmnet with OS bookworm * 17:14 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1146.eqiad.wmnet with OS bookworm * 17:09 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326885{{!}}w/deployment-info.php: Handle new file format (T434726)]] (duration: 07m 15s) * 17:08 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 17:05 dancy@deploy1003: dancy: Continuing with deployment * 17:04 dancy@deploy1003: dancy: Backport for [[gerrit:1326885{{!}}w/deployment-info.php: Handle new file format (T434726)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:02 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1326885{{!}}w/deployment-info.php: Handle new file format (T434726)]] * 16:50 dancy@deploy1003: Finished scap sync-world: Testing [[phab:T434726|T434726]] (duration: 06m 40s) * 16:43 dancy@deploy1003: Started scap sync-world: Testing [[phab:T434726|T434726]] * 16:43 dancy@deploy1003: Installation of scap version "4.282.0" completed for 3 hosts * 16:41 dancy@deploy1003: Installing scap version "4.282.0" for 3 host(s) * 16:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1145.eqiad.wmnet with OS bookworm * 16:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1200.eqiad.wmnet with OS bookworm * 16:32 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1199.eqiad.wmnet with OS bookworm * 16:18 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1145.eqiad.wmnet with reason: host reimage * 16:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1200.eqiad.wmnet with reason: host reimage * 16:09 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1199.eqiad.wmnet with reason: host reimage * 16:05 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-esams and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 16:05 cjd91: sudo -i cookbook sre.cdn.roll-upgrade-ats --query 'A:cp-esams' --task-id [[phab:T434478|T434478]] --reason '9.2.15 upgrade' * 16:03 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1200.eqiad.wmnet with reason: host reimage * 16:02 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1145.eqiad.wmnet with reason: host reimage * 16:02 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1199.eqiad.wmnet with reason: host reimage * 15:50 moritzm: installing zip security updates * 15:48 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1200.eqiad.wmnet with OS bookworm * 15:47 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1199.eqiad.wmnet with OS bookworm * 15:47 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1145.eqiad.wmnet with OS bookworm * 15:41 topranks: bounce PIC 0/0 on cr1-magru to set port to 40G * 15:37 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1236.eqiad.wmnet with OS bookworm * 15:29 aikochou@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'ores-legacy' for release 'main' . * 15:26 aikochou@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'ores-legacy' for release 'main' . * 15:20 aikochou@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'ores-legacy' for release 'main' . * 15:14 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2002.codfw.wmnet * 15:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1236.eqiad.wmnet with reason: host reimage * 15:12 moritzm: failover ganeti master in codfw to ganeti2047 * 15:09 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1236.eqiad.wmnet with reason: host reimage * 15:09 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2044.codfw.wmnet * 15:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2002.codfw.wmnet * 15:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2044.codfw.wmnet * 15:04 brennen@deploy1003: Finished deploy [phabricator/deployment@6b9b6ff]: deploy phab1004 for [[phab:T435213|T435213]] (duration: 01m 01s) * 15:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2044.codfw.wmnet * 15:03 brennen@deploy1003: Started deploy [phabricator/deployment@6b9b6ff]: deploy phab1004 for [[phab:T435213|T435213]] * 15:03 brennen@deploy1003: Finished deploy [phabricator/deployment@6b9b6ff]: deploy phab2003 for [[phab:T435213|T435213]] (duration: 00m 57s) * 15:02 brennen@deploy1003: Started deploy [phabricator/deployment@6b9b6ff]: deploy phab2003 for [[phab:T435213|T435213]] * 14:57 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply * 14:57 arnaudb@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on phab2003.codfw.wmnet,phab[1004-1006].eqiad.wmnet with reason: maintenance * 14:56 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2044.codfw.wmnet * 14:55 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply * 14:53 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1236.eqiad.wmnet with OS bookworm * 14:51 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2043.codfw.wmnet * 14:51 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2043.codfw.wmnet * 14:45 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2043.codfw.wmnet * 14:33 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2043.codfw.wmnet * 14:22 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2042.codfw.wmnet * 14:22 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2042.codfw.wmnet * 14:21 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db2902.codfw.wmnet with OS trixie * 14:16 elukey: uploaded spicerack_13.2.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia * 14:16 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2042.codfw.wmnet * 14:04 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2042.codfw.wmnet * 14:04 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2041.codfw.wmnet * 14:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2041.codfw.wmnet * 13:59 phuedx@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: apply * 13:59 phuedx@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-main: apply * 13:59 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326838{{!}}Enable Suggested Investigations on hewiki (T435146)]] (duration: 11m 50s) * 13:58 phuedx@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: apply * 13:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2041.codfw.wmnet * 13:57 phuedx@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-main: apply * 13:57 phuedx@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-main: apply * 13:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard1003.eqiad.wmnet * 13:57 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-main: apply * 13:55 phuedx@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-logging-external: apply * 13:55 phuedx@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-logging-external: apply * 13:54 phuedx@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-logging-external: apply * 13:54 stran@deploy1003: stran: Continuing with deployment * 13:54 phuedx@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-logging-external: apply * 13:54 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-logging-external: apply * 13:54 phuedx@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-logging-external: apply * 13:53 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-logging-external: apply * 13:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard1003.eqiad.wmnet * 13:53 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2041.codfw.wmnet * 13:51 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db2902.codfw.wmnet - fceratto@cumin1003" * 13:51 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db2902.codfw.wmnet - fceratto@cumin1003" * 13:51 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2026.codfw.wmnet * 13:51 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard2003.codfw.wmnet * 13:50 phuedx@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: apply * 13:50 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-eqiad * 13:50 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp1001.eqiad.wmnet * 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp1001.eqiad.wmnet * 13:50 phuedx@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: apply * 13:49 phuedx@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: apply * 13:49 stran@deploy1003: stran: Backport for [[gerrit:1326838{{!}}Enable Suggested Investigations on hewiki (T435146)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:48 phuedx@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: apply * 13:48 phuedx@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics: apply * 13:47 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard2003.codfw.wmnet * 13:47 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics: apply * 13:47 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1326838{{!}}Enable Suggested Investigations on hewiki (T435146)]] * 13:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp1001.eqiad.wmnet * 13:44 cdanis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 13:43 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp1001.eqiad.wmnet * 13:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1315-1327].eqiad.wmnet * 13:43 cdanis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 13:43 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1315-1327].eqiad.wmnet * 13:40 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki1001.eqiad.wmnet * 13:35 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1315-1327].eqiad.wmnet * 13:34 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host rpki1001.eqiad.wmnet * 13:27 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1315-1327].eqiad.wmnet * 13:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1302-1314].eqiad.wmnet * 13:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1302-1314].eqiad.wmnet * 13:21 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326775{{!}}SI: Instrument case update on first edit (T435048)]] (duration: 07m 12s) * 13:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1302-1314].eqiad.wmnet * 13:17 stran@deploy1003: stran: Continuing with deployment * 13:16 stran@deploy1003: stran: Backport for [[gerrit:1326775{{!}}SI: Instrument case update on first edit (T435048)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:14 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1326775{{!}}SI: Instrument case update on first edit (T435048)]] * 13:13 moritzm: installing util-linux security updates * 13:11 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1302-1314].eqiad.wmnet * 13:10 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326232{{!}}prv: Enable parsoid rendering for 5 wikis (T435115)]] (duration: 08m 17s) * 13:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1288-1289,1291-1301].eqiad.wmnet * 13:10 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply * 13:10 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1288-1289,1291-1301].eqiad.wmnet * 13:10 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply * 13:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki2003.codfw.wmnet * 13:06 jgiannelos@deploy1003: jgiannelos: Continuing with deployment * 13:05 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host rpki2003.codfw.wmnet * 13:04 jgiannelos@deploy1003: jgiannelos: Backport for [[gerrit:1326232{{!}}prv: Enable parsoid rendering for 5 wikis (T435115)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1288-1289,1291-1301].eqiad.wmnet * 13:02 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1326232{{!}}prv: Enable parsoid rendering for 5 wikis (T435115)]] * 12:53 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1288-1289,1291-1301].eqiad.wmnet * 12:53 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1273,1275-1287].eqiad.wmnet * 12:53 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1273,1275-1287].eqiad.wmnet * 12:52 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt1002.wikimedia.org * 12:51 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db2902.codfw.wmnet on all recursors * 12:51 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db2902.codfw.wmnet on all recursors * 12:51 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:51 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db2902.codfw.wmnet - fceratto@cumin1003" * 12:51 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db2902.codfw.wmnet - fceratto@cumin1003" * 12:46 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt1002.wikimedia.org * 12:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt2002.wikimedia.org * 12:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1273,1275-1287].eqiad.wmnet * 12:42 dhinus: repooled clouddb1032 that was currently <nowiki>{</nowiki>"weight": 0, "pooled": "inactive"<nowiki>}</nowiki> for both s4 and s6 * 12:41 dhinus: also depooled clouddb1017 (forgot it in the previous list) * 12:41 fnegri@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet * 12:40 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 12:40 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db2902.codfw.wmnet * 12:40 fnegri@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032.eqiad.wmnet * 12:40 fnegri@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032 * 12:39 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1017.eqiad.wmnet * 12:39 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt2002.wikimedia.org * 12:38 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1020.eqiad.wmnet * 12:38 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1018.eqiad.wmnet * 12:38 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet * 12:37 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1014.eqiad.wmnet * 12:37 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1013.eqiad.wmnet * 12:37 dhinus: depool again clouddb10[13,14,16,18,20] that were repooled by the cookbook sre.mysql.multiinstance_reboot * 12:36 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1273,1275-1287].eqiad.wmnet * 12:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1248-1261].eqiad.wmnet * 12:36 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1248-1261].eqiad.wmnet * 12:35 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2026.codfw.wmnet * 12:34 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 12:34 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 12:34 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 12:34 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 12:29 lucaswerkmeister-wmde@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 12:28 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2026.codfw.wmnet * 12:28 lucaswerkmeister-wmde@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 12:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1248-1261].eqiad.wmnet * 12:27 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.addnode (exit_code=0) for new host ganeti2046.codfw.wmnet to cluster codfw and group A * 12:26 moritzm: readded ganeti2046 to the codfw cluster following firmware update and reimage [[phab:T434681|T434681]] * 12:23 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply * 12:23 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply * 12:23 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply * 12:22 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply * 12:22 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply * 12:22 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply * 12:21 jmm@cumin2003: START - Cookbook sre.ganeti.addnode for new host ganeti2046.codfw.wmnet to cluster codfw and group A * 12:21 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1169.eqiad.wmnet onto db1283.eqiad.wmnet * 12:21 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1169: Pool db1169.eqiad.wmnet in after cloning * 12:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1248-1261].eqiad.wmnet * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1149-1153,1158,1240-1247].eqiad.wmnet * 12:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1149-1153,1158,1240-1247].eqiad.wmnet * 12:11 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1149-1153,1158,1240-1247].eqiad.wmnet * 12:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1168.eqiad.wmnet onto db1282.eqiad.wmnet * 12:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1168: Pool db1168.eqiad.wmnet in after cloning * 12:06 jmm@cumin2003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-test-eqiad * 12:05 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2026.codfw.wmnet * 12:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1149-1153,1158,1240-1247].eqiad.wmnet * 12:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1128-1134,1142-1148].eqiad.wmnet * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2046.codfw.wmnet * 12:01 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1128-1134,1142-1148].eqiad.wmnet * 11:57 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2025.codfw.wmnet * 11:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2025.codfw.wmnet * 11:54 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2046.codfw.wmnet * 11:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1128-1134,1142-1148].eqiad.wmnet * 11:50 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2025.codfw.wmnet * 11:46 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1128-1134,1142-1148].eqiad.wmnet * 11:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1114-1127].eqiad.wmnet * 11:45 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1114-1127].eqiad.wmnet * 11:39 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db2901.codfw.wmnet * 11:39 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db2901.codfw.wmnet * 11:36 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2025.codfw.wmnet * 11:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1114-1127].eqiad.wmnet * 11:35 fceratto@cumin1003: END (ERROR) - Cookbook sre.ganeti.makevm (exit_code=93) for new host db1901.eqiad.wmnet * 11:35 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1169: Pool db1169.eqiad.wmnet in after cloning * 11:35 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 11:30 jmm@cumin2003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-test-eqiad * 11:27 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1114-1127].eqiad.wmnet * 11:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1076-1081,1084-1087,1093-1095,1113].eqiad.wmnet * 11:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1076-1081,1084-1087,1093-1095,1113].eqiad.wmnet * 11:25 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1168: Pool db1168.eqiad.wmnet in after cloning * 11:24 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1194.eqiad.wmnet with OS bookworm * 11:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1181.eqiad.wmnet with OS bookworm * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2050.codfw.wmnet * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2050.codfw.wmnet * 11:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1076-1081,1084-1087,1093-1095,1113].eqiad.wmnet * 11:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2050.codfw.wmnet * 11:14 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2004.codfw.wmnet * 11:10 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2050.codfw.wmnet * 11:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1076-1081,1084-1087,1093-1095,1113].eqiad.wmnet * 11:09 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1286: Pool back * 11:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1045-1050,1056-1057,1064-1066,1073-1075].eqiad.wmnet * 11:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1045-1050,1056-1057,1064-1066,1073-1075].eqiad.wmnet * 11:08 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2004.codfw.wmnet * 11:07 moritzm: installing PHP 8.4 security updates * 11:06 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on an-worker1194.eqiad.wmnet with reason: host reimage * 11:06 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1194.eqiad.wmnet with reason: host reimage * 11:05 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2049.codfw.wmnet * 11:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2049.codfw.wmnet * 11:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid1003.eqiad.wmnet * 11:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1045-1050,1056-1057,1064-1066,1073-1075].eqiad.wmnet * 10:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1181.eqiad.wmnet with reason: host reimage * 10:59 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2049.codfw.wmnet * 10:59 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid1003.eqiad.wmnet * 10:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid2003.codfw.wmnet * 10:54 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1181.eqiad.wmnet with reason: host reimage * 10:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid2003.codfw.wmnet * 10:51 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1901.eqiad.wmnet on all recursors * 10:51 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1901.eqiad.wmnet on all recursors * 10:51 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:51 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:51 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:50 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2049.codfw.wmnet * 10:50 blake@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:50 blake@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:49 blake@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:48 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1045-1050,1056-1057,1064-1066,1073-1075].eqiad.wmnet * 10:48 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1044].eqiad.wmnet * 10:48 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1044].eqiad.wmnet * 10:47 blake@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:46 blake@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:46 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2047.codfw.wmnet * 10:46 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:46 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2047.codfw.wmnet * 10:46 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1901.eqiad.wmnet * 10:46 blake@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:44 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1168: Depool db1168.eqiad.wmnet to then clone it to db1282.eqiad.wmnet - marostegui@cumin1003 * 10:44 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1168: Depool db1168.eqiad.wmnet to then clone it to db1282.eqiad.wmnet - marostegui@cumin1003 * 10:44 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1168.eqiad.wmnet onto db1282.eqiad.wmnet * 10:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2047.codfw.wmnet * 10:40 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1044].eqiad.wmnet * 10:37 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2047.codfw.wmnet * 10:37 fceratto@cumin1003: END (ERROR) - Cookbook sre.ganeti.makevm (exit_code=93) for new host db1901.eqiad.wmnet * 10:36 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 10:34 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:33 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2032.codfw.wmnet * 10:33 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2032.codfw.wmnet * 10:32 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1044].eqiad.wmnet * 10:32 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-eqiad * 10:27 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2032.codfw.wmnet * 10:24 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1286: Pool back * 10:24 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1286 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96163 and previous config saved to /var/cache/conftool/dbconfig/20260818-102431-marostegui.json * 10:22 moritzm: installing Django security updates * 10:20 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul2002.codfw.wmnet * 10:20 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul2001.codfw.wmnet * 10:17 blake@deploy1003: Finished scap sync-world: no-build deployment for [[phab:T417800|T417800]] (duration: 04m 40s) * 10:16 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul2002.codfw.wmnet * 10:16 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul2001.codfw.wmnet * 10:15 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2032.codfw.wmnet * 10:14 blake@deploy1003: Started scap sync-world: no-build deployment for [[phab:T417800|T417800]] * 10:12 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul2003.codfw.wmnet * 10:12 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy2001.codfw.wmnet * 10:12 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy3001.esams.wmnet * 10:12 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy1001.eqiad.wmnet * 10:08 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2031.codfw.wmnet * 10:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2031.codfw.wmnet * 10:08 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul2003.codfw.wmnet * 10:08 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy2001.codfw.wmnet * 10:08 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy3001.esams.wmnet * 10:08 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy1001.eqiad.wmnet * 10:07 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy1002.eqiad.wmnet * 10:05 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy2002.codfw.wmnet * 10:05 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy3002.esams.wmnet * 10:04 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy1002.eqiad.wmnet * 10:03 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy4003.ulsfo.wmnet * 10:03 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy5003.eqsin.wmnet * 10:02 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2031.codfw.wmnet * 10:02 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1169: Depool db1169.eqiad.wmnet to then clone it to db1283.eqiad.wmnet - marostegui@cumin1003 * 10:01 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy2002.codfw.wmnet * 10:01 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy4004.ulsfo.wmnet * 10:01 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy3002.esams.wmnet * 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1169: Depool db1169.eqiad.wmnet to then clone it to db1283.eqiad.wmnet - marostegui@cumin1003 * 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1169.eqiad.wmnet onto db1283.eqiad.wmnet * 10:01 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy5004.eqsin.wmnet * 09:59 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy4003.ulsfo.wmnet * 09:59 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy4004.ulsfo.wmnet * 09:59 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy5003.eqsin.wmnet * 09:59 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy7001.magru.wmnet * 09:59 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy5004.eqsin.wmnet * 09:58 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy7002.magru.wmnet * 09:57 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2031.codfw.wmnet * 09:57 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy6001.drmrs.wmnet * 09:57 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy6002.drmrs.wmnet * 09:54 filippo@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for 10 hosts * 09:54 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast4006.wikimedia.org * 09:53 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy6001.drmrs.wmnet * 09:53 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy6002.drmrs.wmnet * 09:53 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host releases1003.eqiad.wmnet * 09:53 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy7001.magru.wmnet * 09:53 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host people2004.codfw.wmnet * 09:52 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host people1005.eqiad.wmnet * 09:52 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy7002.magru.wmnet * 09:50 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host releases2003.codfw.wmnet * 09:49 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host releases2003.codfw.wmnet * 09:49 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host releases1003.eqiad.wmnet * 09:49 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host people2004.codfw.wmnet * 09:48 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host people1005.eqiad.wmnet * 09:46 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2040.codfw.wmnet * 09:46 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2040.codfw.wmnet * 09:43 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1901.eqiad.wmnet on all recursors * 09:42 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1901.eqiad.wmnet on all recursors * 09:42 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:42 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:42 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2040.codfw.wmnet * 09:40 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp2005.wikimedia.org * 09:36 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp2005.wikimedia.org * 09:31 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:31 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1901.eqiad.wmnet * 09:31 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1901.eqiad.wmnet * 09:31 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:31 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1901.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 09:31 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1901.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 09:29 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2040.codfw.wmnet * 09:29 slyngshede@dns1004: END - running authdns-update * 09:28 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2039.codfw.wmnet * 09:28 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2039.codfw.wmnet * 09:27 slyngshede@dns1004: START - running authdns-update * 09:26 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test2005.wikimedia.org * 09:22 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2039.codfw.wmnet * 09:22 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test2005.wikimedia.org * 09:22 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp1005.wikimedia.org * 09:21 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:19 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2039.codfw.wmnet * 09:19 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2038.codfw.wmnet * 09:18 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1287: Pool back * 09:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2038.codfw.wmnet * 09:18 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp1005.wikimedia.org * 09:18 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test1005.wikimedia.org * 09:17 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1901.eqiad.wmnet * 09:15 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db1901.eqiad.wmnet * 09:15 fceratto@cumin1003: END (ERROR) - Cookbook sre.dns.netbox (exit_code=97) * 09:14 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test1005.wikimedia.org * 09:13 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2038.codfw.wmnet * 09:13 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:13 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1901.eqiad.wmnet * 09:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast4006.wikimedia.org * 09:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1288: Pool back * 09:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast5005.wikimedia.org * 09:03 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2038.codfw.wmnet * 08:57 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2037.codfw.wmnet * 08:57 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast5005.wikimedia.org * 08:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2037.codfw.wmnet * 08:56 filippo@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 10 hosts * 08:52 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2037.codfw.wmnet * 08:51 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1289: Pool back * 08:35 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudidp2001-dev.codfw.wmnet * 08:34 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1002-dev.eqiad.wmnet * 08:33 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1287: Pool back * 08:33 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1287 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96150 and previous config saved to /var/cache/conftool/dbconfig/20260818-083311-marostegui.json * 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1001-dev.eqiad.wmnet * 08:31 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudidp2001-dev.codfw.wmnet * 08:30 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1002-dev.eqiad.wmnet * 08:30 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2037.codfw.wmnet * 08:29 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1001-dev.eqiad.wmnet * 08:28 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2036.codfw.wmnet * 08:28 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2036.codfw.wmnet * 08:25 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1288: Pool back * 08:23 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 08:23 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1181.eqiad.wmnet with OS bookworm * 08:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2036.codfw.wmnet * 08:22 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1288 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96148 and previous config saved to /var/cache/conftool/dbconfig/20260818-082234-marostegui.json * 08:20 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2036.codfw.wmnet * 08:18 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2035.codfw.wmnet * 08:18 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudvirt1057.eqiad.wmnet with OS trixie * 08:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2035.codfw.wmnet * 08:13 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2035.codfw.wmnet * 08:06 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2035.codfw.wmnet * 08:05 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1289: Pool back * 08:05 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1289 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96145 and previous config saved to /var/cache/conftool/dbconfig/20260818-080531-marostegui.json * 07:54 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 07:51 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326439{{!}}Fix "mathjax_ignore" handling around forcemathmode attribute (T434686)]] (duration: 13m 15s) * 07:48 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 07:47 krinkle@deploy1003: krinkle: Continuing with deployment * 07:44 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ganeti2046.codfw.wmnet with OS bookworm * 07:40 krinkle@deploy1003: krinkle: Backport for [[gerrit:1326439{{!}}Fix "mathjax_ignore" handling around forcemathmode attribute (T434686)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:39 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 07:38 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1326439{{!}}Fix "mathjax_ignore" handling around forcemathmode attribute (T434686)]] * 07:34 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 07:31 samwilson@deploy1003: Finished scap sync-world: Backport for [[gerrit:701016{{!}}InitialiseSettings and -labs: Remove redundant feature flag $wgWikisourceEnableOcr (T285311)]] (duration: 07m 47s) * 07:30 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1048.eqiad.wmnet * 07:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1048.eqiad.wmnet * 07:29 XioNoX: add gnmic 0.47.0 to bookworm and trixie reprepro * 07:28 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ganeti2046.codfw.wmnet with reason: host reimage * 07:27 samwilson@deploy1003: samwilson: Continuing with deployment * 07:25 samwilson@deploy1003: samwilson: Backport for [[gerrit:701016{{!}}InitialiseSettings and -labs: Remove redundant feature flag $wgWikisourceEnableOcr (T285311)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:25 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 07:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1048.eqiad.wmnet * 07:24 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ganeti2046.codfw.wmnet with reason: host reimage * 07:23 samwilson@deploy1003: Started scap sync-world: Backport for [[gerrit:701016{{!}}InitialiseSettings and -labs: Remove redundant feature flag $wgWikisourceEnableOcr (T285311)]] * 07:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1198.eqiad.wmnet with OS bookworm * 07:18 samwilson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326450{{!}}InitialiseSettings.php: Enable Bulk OCR on pawikisource (T434648)]] (duration: 12m 17s) * 07:17 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 07:11 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1048.eqiad.wmnet * 07:11 samwilson@deploy1003: samwilson: Continuing with deployment * 07:11 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ganeti2046.codfw.wmnet with OS bookworm * 07:10 samwilson@deploy1003: samwilson: Backport for [[gerrit:1326450{{!}}InitialiseSettings.php: Enable Bulk OCR on pawikisource (T434648)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1057.eqiad.wmnet with OS trixie * 07:05 samwilson@deploy1003: Started scap sync-world: Backport for [[gerrit:1326450{{!}}InitialiseSettings.php: Enable Bulk OCR on pawikisource (T434648)]] * 07:02 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1198.eqiad.wmnet with reason: host reimage * 07:01 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1056.eqiad.wmnet with OS trixie * 07:01 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1056.eqiad.wmnet with OS trixie * 07:00 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1056.eqiad.wmnet with OS trixie * 07:00 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1056.eqiad.wmnet with OS trixie * 06:59 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudvirt1055.eqiad.wmnet with OS trixie * 06:58 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1198.eqiad.wmnet with reason: host reimage * 06:53 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1055.eqiad.wmnet with OS trixie * 06:52 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudvirt1054.eqiad.wmnet with OS trixie * 06:44 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1194.eqiad.wmnet with OS bookworm * 06:44 XioNoX: upgrade eqsin gnmic to 0.47.0 * 06:43 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1198.eqiad.wmnet with OS bookworm * 06:41 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1054.eqiad.wmnet with OS trixie * 06:09 arnaudb@cumin1003: END (PASS) - Cookbook sre.gerrit.restart-gerrit (exit_code=0) Restarting Gerrit on gerrit2002 * 06:06 arnaudb@cumin1003: START - Cookbook sre.gerrit.restart-gerrit Restarting Gerrit on gerrit2002 * 06:06 arnaudb@cumin1003: END (PASS) - Cookbook sre.gerrit.restart-gerrit (exit_code=0) Restarting Gerrit on gerrit1003 * 06:04 arnaudb@cumin1003: START - Cookbook sre.gerrit.restart-gerrit Restarting Gerrit on gerrit1003 * 06:02 arnaudb@cumin1003: END (PASS) - Cookbook sre.gerrit.restart-gerrit (exit_code=0) Restarting Gerrit on gerrit2003 * 06:00 arnaudb@cumin1003: START - Cookbook sre.gerrit.restart-gerrit Restarting Gerrit on gerrit2003 * 05:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1155.eqiad.wmnet with OS bookworm * 05:38 arnaudb: updating prometheusBearerToken on gerrit * 05:28 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1155.eqiad.wmnet with reason: host reimage * 05:23 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1155.eqiad.wmnet with reason: host reimage * 05:06 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1155.eqiad.wmnet with OS bookworm * 04:57 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1144.eqiad.wmnet with OS bookworm * 04:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1144.eqiad.wmnet with reason: host reimage * 04:29 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1144.eqiad.wmnet with reason: host reimage * 04:14 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1144.eqiad.wmnet with OS bookworm * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.13 (duration: 02m 23s) * 03:45 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1197.eqiad.wmnet with OS bookworm * 03:41 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1165.eqiad.wmnet with OS bookworm * 03:38 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.16 refs [[phab:T430835|T430835]] (duration: 34m 43s) * 03:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1164.eqiad.wmnet with OS bookworm * 03:35 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1196.eqiad.wmnet with OS bookworm * 03:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1163.eqiad.wmnet with OS bookworm * 03:22 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1197.eqiad.wmnet with reason: host reimage * 03:18 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1165.eqiad.wmnet with reason: host reimage * 03:15 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1196.eqiad.wmnet with reason: host reimage * 03:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1164.eqiad.wmnet with reason: host reimage * 03:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1163.eqiad.wmnet with reason: host reimage * 03:05 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1165.eqiad.wmnet with reason: host reimage * 03:04 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1164.eqiad.wmnet with reason: host reimage * 03:04 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1197.eqiad.wmnet with reason: host reimage * 03:04 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1196.eqiad.wmnet with reason: host reimage * 03:04 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1163.eqiad.wmnet with reason: host reimage * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.16 refs [[phab:T430835|T430835]] * 02:50 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1197.eqiad.wmnet with OS bookworm * 02:49 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1196.eqiad.wmnet with OS bookworm * 02:49 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1165.eqiad.wmnet with OS bookworm * 02:49 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1164.eqiad.wmnet with OS bookworm * 02:48 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1163.eqiad.wmnet with OS bookworm * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 46s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:15 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-codfw: Set storage compatability to NONE — [[phab:T433026|T433026]] - eevans@cumin1003 * 00:38 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-codfw: Set storage compatability to NONE — [[phab:T433026|T433026]] - eevans@cumin1003 * 00:11 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326373{{!}}PersonalDashboard: add newly renamed *ReviewChangesMlModel setting (T422148)]] (duration: 07m 07s) * 00:09 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-eqiad: Set storage compatability to NONE — [[phab:T433026|T433026]] - eevans@cumin1003 * 00:07 musikanimal@deploy1003: musikanimal: Continuing with deployment * 00:06 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1326373{{!}}PersonalDashboard: add newly renamed *ReviewChangesMlModel setting (T422148)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:04 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1326373{{!}}PersonalDashboard: add newly renamed *ReviewChangesMlModel setting (T422148)]] == 2026-08-17 == * 23:30 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-eqiad: Set storage compatability to NONE — [[phab:T433026|T433026]] - eevans@cumin1003 * 23:05 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-codfw: Set storage compatability to UPGRADING — [[phab:T433026|T433026]] - eevans@cumin1003 * 22:28 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-codfw: Set storage compatability to UPGRADING — [[phab:T433026|T433026]] - eevans@cumin1003 * 21:46 logmsgbot: jforrester Deployed security patch for [[phab:T435085|T435085]] * 21:39 swfrench@deploy1003: mwscript-k8s job started: purgeList.php # [[phab:T432412|T432412]] * 21:37 maryum: Undeploy security fix for [[phab:T433020|T433020]] * 21:23 maryum: Deployed security fix for [[phab:T433020|T433020]] * 21:21 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 21:21 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 21:14 maryum: Deployed security fix for [[phab:T434967|T434967]] * 20:58 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-eqiad: Set storage compatability to UPGRADING — [[phab:T433026|T433026]] - eevans@cumin1003 * 20:40 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324320{{!}}[itwiki/slwiki/tgwiki] Remove temporary Wikipedia 25 logos permanently (already reverted) (T414265 T414320 T415307)]] (duration: 06m 55s) * 20:40 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 20:39 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 20:36 cjming@deploy1003: cjming, superpes: Continuing with deployment * 20:35 cjming@deploy1003: cjming, superpes: Backport for [[gerrit:1324320{{!}}[itwiki/slwiki/tgwiki] Remove temporary Wikipedia 25 logos permanently (already reverted) (T414265 T414320 T415307)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:33 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1324320{{!}}[itwiki/slwiki/tgwiki] Remove temporary Wikipedia 25 logos permanently (already reverted) (T414265 T414320 T415307)]] * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ttmserver-test: apply * 20:30 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326361{{!}}Remove escaped paths in app site association file (T432412)]] (duration: 13m 28s) * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ttmserver-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-toolhub-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-toolhub-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-toolhub-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-toolhub-test: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-toolhub: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-toolhub: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-toolhub: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-toolhub: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-test: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-test: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 20:27 inflatador: bking@deploy1003 `charlie --services_dir dse-k8s-services -s opensearch-* -e dse-k8s-* apply` [[phab:T435125|T435125]] * 20:27 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-apifeatureusage-test: apply * 20:27 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-apifeatureusage-test: apply * 20:27 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-apifeatureusage-test: apply * 20:27 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-apifeatureusage-test: apply * 20:27 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-apifeatureusage: apply * 20:26 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-apifeatureusage: apply * 20:26 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-apifeatureusage: apply * 20:26 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-apifeatureusage: apply * 20:26 cjming@deploy1003: cjming, tsev: Continuing with deployment * 20:24 inflatador: bking@deploy1003 `charlie --services_dir dse-k8s-services -s opensearch-* -e dse-k8s-*` [[phab:T435125|T435125]] * 20:21 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-eqiad: Set storage compatability to UPGRADING — [[phab:T433026|T433026]] - eevans@cumin1003 * 20:19 cjming@deploy1003: cjming, tsev: Backport for [[gerrit:1326361{{!}}Remove escaped paths in app site association file (T432412)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:17 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1326361{{!}}Remove escaped paths in app site association file (T432412)]] * 20:15 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326077{{!}}InstrumentConstructiveEdits: anchor all runs to the nearest `interval` (T431493)]] (duration: 06m 25s) * 20:11 cjming@deploy1003: cjming: Continuing with deployment * 20:10 cjming@deploy1003: cjming: Backport for [[gerrit:1326077{{!}}InstrumentConstructiveEdits: anchor all runs to the nearest `interval` (T431493)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:08 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1326077{{!}}InstrumentConstructiveEdits: anchor all runs to the nearest `interval` (T431493)]] * 20:06 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1160.eqiad.wmnet with OS bookworm * 20:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1162.eqiad.wmnet with OS bookworm * 19:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1184.eqiad.wmnet with OS bookworm * 19:55 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1161.eqiad.wmnet with OS bookworm * 19:49 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1195.eqiad.wmnet with OS bookworm * 19:49 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-test: apply * 19:49 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-test: apply * 19:44 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1160.eqiad.wmnet with reason: host reimage * 19:41 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-codfw: Upgrade to Java 17 — [[phab:T433026|T433026]] - eevans@cumin1003 * 19:39 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1162.eqiad.wmnet with reason: host reimage * 19:36 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-eqsin and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 19:36 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-test: apply * 19:36 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1184.eqiad.wmnet with reason: host reimage * 19:32 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1161.eqiad.wmnet with reason: host reimage * 19:29 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1195.eqiad.wmnet with reason: host reimage * 19:26 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1161.eqiad.wmnet with reason: host reimage * 19:26 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1160.eqiad.wmnet with reason: host reimage * 19:26 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1184.eqiad.wmnet with reason: host reimage * 19:26 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1162.eqiad.wmnet with reason: host reimage * 19:25 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1195.eqiad.wmnet with reason: host reimage * 19:19 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-test: apply * 19:11 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1195.eqiad.wmnet with OS bookworm * 19:10 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1194.eqiad.wmnet with OS bookworm * 19:10 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1184.eqiad.wmnet with OS bookworm * 19:10 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1162.eqiad.wmnet with OS bookworm * 19:10 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1161.eqiad.wmnet with OS bookworm * 19:10 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1160.eqiad.wmnet with OS bookworm * 19:10 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 19:09 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 19:03 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-codfw: Upgrade to Java 17 — [[phab:T433026|T433026]] - eevans@cumin1003 * 18:58 dancy@deploy1003: Finished scap sync-world: testing [[phab:T375514|T375514]] (duration: 03m 13s) * 18:55 dancy@deploy1003: Started scap sync-world: testing [[phab:T375514|T375514]] * 18:55 dwisehaupt@dns1006: END - running authdns-update * 18:54 dancy@deploy1003: Installation of scap version "4.281.1" completed for 3 hosts * 18:54 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2009.codfw.wmnet * 18:53 dwisehaupt@dns1006: START - running authdns-update * 18:52 dancy@deploy1003: Installing scap version "4.281.1" for 3 host(s) * 18:47 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2009.codfw.wmnet * 18:41 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2008.codfw.wmnet * 18:34 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2008.codfw.wmnet * 18:30 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2007.codfw.wmnet * 18:23 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2007.codfw.wmnet * 18:16 dwisehaupt@dns1005: END - running authdns-update * 18:14 dwisehaupt@dns1005: START - running authdns-update * 18:04 swfrench@deploy1003: Finished scap sync-world: Deploy "Point Test Wiki to new docroot" - [[phab:T432412|T432412]] (duration: 26m 03s) * 18:00 swfrench@deploy1003: swfrench: Continuing with deployment * 17:51 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1057.eqiad.wmnet with OS trixie * 17:47 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1180.eqiad.wmnet with OS bookworm * 17:39 swfrench@deploy1003: swfrench: Deploy "Point Test Wiki to new docroot" - [[phab:T432412|T432412]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:39 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-eqsin and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 17:38 swfrench@deploy1003: Started scap sync-world: Deploy "Point Test Wiki to new docroot" - [[phab:T432412|T432412]] * 17:36 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1159.eqiad.wmnet with OS bookworm * 17:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1158.eqiad.wmnet with OS bookworm * 17:30 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-eqiad: Upgrade to Java 17 — [[phab:T433026|T433026]] - eevans@cumin1003 * 17:30 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-ulsfo or A:cp-drmrs and A:cp - 9.2.15 upgrade ([[phab:T434620|T434620]]) * 17:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1157.eqiad.wmnet with OS bookworm * 17:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1180.eqiad.wmnet with reason: host reimage * 17:22 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 17:21 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1193.eqiad.wmnet with OS bookworm * 17:21 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 17:21 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 17:20 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 17:20 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 17:20 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1183.eqiad.wmnet with OS bookworm * 17:18 swfrench@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 17:18 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 17:17 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1180.eqiad.wmnet with reason: host reimage * 17:17 swfrench@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 17:17 swfrench@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 17:16 swfrench@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 17:16 swfrench@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 17:15 swfrench@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 17:15 swfrench@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 17:15 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1192.eqiad.wmnet with OS bookworm * 17:14 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1159.eqiad.wmnet with reason: host reimage * 17:14 swfrench@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 17:13 swfrench@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 17:12 swfrench@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 17:09 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1158.eqiad.wmnet with reason: host reimage * 17:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1157.eqiad.wmnet with reason: host reimage * 17:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1180 * 17:02 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1180 * 17:01 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica-eqsin and A:liberica * 17:01 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1193.eqiad.wmnet with reason: host reimage * 16:57 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1183.eqiad.wmnet with reason: host reimage * 16:56 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1054.eqiad.wmnet with OS trixie * 16:55 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1056.eqiad.wmnet with OS trixie * 16:54 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1193.eqiad.wmnet with reason: host reimage * 16:54 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1192.eqiad.wmnet with reason: host reimage * 16:51 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-eqiad: Upgrade to Java 17 — [[phab:T433026|T433026]] - eevans@cumin1003 * 16:50 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1180 * 16:50 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1180.eqiad.wmnet 17.36.64.10.in-addr.arpa 7.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:50 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1180.eqiad.wmnet 17.36.64.10.in-addr.arpa 7.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:50 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:50 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1180 - btullis@cumin1003" * 16:50 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1180 - btullis@cumin1003" * 16:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1158.eqiad.wmnet with reason: host reimage * 16:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1157.eqiad.wmnet with reason: host reimage * 16:49 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica-eqsin and A:liberica * 16:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1183.eqiad.wmnet with reason: host reimage * 16:48 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1159.eqiad.wmnet with reason: host reimage * 16:47 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1192.eqiad.wmnet with reason: host reimage * 16:46 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns4004.wikimedia.org * 16:46 sukhe@dns1004: END - running authdns-update * 16:44 sukhe@dns1004: START - running authdns-update * 16:44 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns4004.wikimedia.org,service=authdns-update * 16:43 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns4004.wikimedia.org with OS trixie * 16:39 btullis@cumin1003: START - Cookbook sre.dns.netbox * 16:39 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1180 * 16:39 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1193.eqiad.wmnet with OS bookworm * 16:39 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1180.eqiad.wmnet with OS bookworm * 16:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1192.eqiad.wmnet with OS bookworm * 16:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1183.eqiad.wmnet with OS bookworm * 16:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1159.eqiad.wmnet with OS bookworm * 16:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1158.eqiad.wmnet with OS bookworm * 16:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1157.eqiad.wmnet with OS bookworm * 16:31 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1057.eqiad.wmnet with OS trixie * 16:30 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1057.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 16:29 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica-drmrs and A:liberica * 16:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1154.eqiad.wmnet with OS bookworm * 16:23 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1057.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 16:23 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1055.eqiad.wmnet with OS trixie * 16:22 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1057 * 16:22 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1057 * 16:21 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:21 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1057] - vriley@cumin1003" * 16:21 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1057] - vriley@cumin1003" * 16:19 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica-drmrs and A:liberica * 16:17 vriley@cumin1003: START - Cookbook sre.dns.netbox * 16:16 phuedx@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics-external: apply * 16:15 phuedx@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics-external: apply * 16:13 phuedx@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics-external: apply * 16:12 phuedx@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics-external: apply * 16:11 phuedx@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics-external: apply * 16:09 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics-external: apply * 16:09 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1176.eqiad.wmnet with OS bookworm * 16:07 btullis@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 16:06 btullis@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 16:05 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1191.eqiad.wmnet with OS bookworm * 16:03 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1154.eqiad.wmnet with reason: host reimage * 16:03 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching aqs[2002-2012].codfw.wmnet,aqs[1017-1027].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433026|T433026]] - eevans@cumin1003 * 15:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1190.eqiad.wmnet with OS bookworm * 15:59 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1154.eqiad.wmnet with reason: host reimage * 15:55 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-debug: apply * 15:55 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-debug: apply * 15:55 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-debug: apply * 15:55 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/mw-debug: apply * 15:53 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns4004.wikimedia.org with reason: host reimage * 15:50 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns4004.wikimedia.org with reason: host reimage * 15:46 btullis@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 15:46 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1176.eqiad.wmnet with reason: host reimage * 15:45 btullis@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 15:43 btullis@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 15:42 btullis@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 15:42 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1191.eqiad.wmnet with reason: host reimage * 15:39 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1190.eqiad.wmnet with reason: host reimage * 15:36 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1054.eqiad.wmnet with OS trixie * 15:35 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:35 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1056.eqiad.wmnet with OS trixie * 15:35 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:34 moritzm: failover Ganeti master in eqiad to ganeti1046 * 15:34 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1176.eqiad.wmnet with reason: host reimage * 15:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1191.eqiad.wmnet with reason: host reimage * 15:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1190.eqiad.wmnet with reason: host reimage * 15:31 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns2006.wikimedia.org * 15:31 sukhe@dns1004: END - running authdns-update * 15:29 sukhe@dns1004: START - running authdns-update * 15:29 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns2006.wikimedia.org,service=authdns-update * 15:29 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns2006.wikimedia.org * 15:29 sukhe@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns2006.wikimedia.org * 15:26 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica-ulsfo and A:liberica * 15:26 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns1006.wikimedia.org * 15:26 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:25 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns2006.wikimedia.org with OS trixie * 15:25 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1056 * 15:25 sukhe@dns1004: END - running authdns-update * 15:23 sukhe@dns1004: START - running authdns-update * 15:23 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns1006.wikimedia.org,service=authdns-update * 15:23 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns1006.wikimedia.org * 15:23 sukhe@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns1006.wikimedia.org * 15:20 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1056 * 15:20 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:20 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1056~] - vriley@cumin1003" * 15:19 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1056~] - vriley@cumin1003" * 15:19 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns1006.wikimedia.org with OS trixie * 15:19 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns4004.wikimedia.org with OS trixie * 15:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1190.eqiad.wmnet with OS bookworm * 15:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1191.eqiad.wmnet with OS bookworm * 15:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1176.eqiad.wmnet with OS bookworm * 15:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1154.eqiad.wmnet with OS bookworm * 15:16 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica-ulsfo and A:liberica * 15:13 vriley@cumin1003: START - Cookbook sre.dns.netbox * 15:13 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host dns4004.wikimedia.org with OS trixie * 15:12 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1188.eqiad.wmnet with OS bookworm * 15:09 taavi@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318203{{!}}Undeploy WP25EasterEggs (II) (T418134)]] (duration: 08m 56s) * 15:06 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica-magru and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 15:06 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1051.eqiad.wmnet with OS trixie * 15:06 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 15:05 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 15:05 taavi@deploy1003: taavi: Continuing with deployment * 15:04 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-ncredir (exit_code=0) rolling reboot on A:ncredir and A:ncredir * 15:04 taavi@deploy1003: taavi: Backport for [[gerrit:1318203{{!}}Undeploy WP25EasterEggs (II) (T418134)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:03 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1055.eqiad.wmnet with OS trixie * 15:02 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:00 taavi@deploy1003: Started scap sync-world: Backport for [[gerrit:1318203{{!}}Undeploy WP25EasterEggs (II) (T418134)]] * 14:59 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1047.eqiad.wmnet * 14:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1047.eqiad.wmnet * 14:58 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy (exit_code=0) rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 14:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-codfw * 14:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp2001.codfw.wmnet * 14:57 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp2001.codfw.wmnet * 14:57 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica-magru and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 14:57 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:56 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1055 * 14:56 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1055 * 14:55 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns2006.wikimedia.org with reason: host reimage * 14:55 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:55 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1055] - vriley@cumin1003" * 14:55 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1055] - vriley@cumin1003" * 14:54 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1047.eqiad.wmnet * 14:52 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1188.eqiad.wmnet with reason: host reimage * 14:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp2001.codfw.wmnet * 14:51 vriley@cumin1003: START - Cookbook sre.dns.netbox * 14:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp2001.codfw.wmnet * 14:50 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2016.codfw.wmnet * 14:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2016.codfw.wmnet * 14:50 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1189.eqiad.wmnet with OS bookworm * 14:48 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1188.eqiad.wmnet with reason: host reimage * 14:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1051.eqiad.wmnet with reason: host reimage * 14:46 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1182.eqiad.wmnet with OS bookworm * 14:45 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief2002.codfw.wmnet * 14:44 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns1006.wikimedia.org with reason: host reimage * 14:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2016.codfw.wmnet * 14:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1143.eqiad.wmnet with OS bookworm * 14:43 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2016.codfw.wmnet * 14:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2318-2331].codfw.wmnet * 14:43 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2318-2331].codfw.wmnet * 14:42 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching aqs[2002-2012].codfw.wmnet,aqs[1017-1027].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433026|T433026]] - eevans@cumin1003 * 14:41 cgoubert@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326306{{!}}Add placeholder $wmgRedisLockPassword (T366938 T427999)]] (duration: 06m 56s) * 14:41 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief2002.codfw.wmnet * 14:40 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief1002.eqiad.wmnet * 14:39 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1051.eqiad.wmnet with reason: host reimage * 14:38 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns2006.wikimedia.org with reason: host reimage * 14:37 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns1006.wikimedia.org with reason: host reimage * 14:37 cgoubert@deploy1003: cgoubert: Continuing with deployment * 14:36 cgoubert@deploy1003: cgoubert: Backport for [[gerrit:1326306{{!}}Add placeholder $wmgRedisLockPassword (T366938 T427999)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:36 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief1002.eqiad.wmnet * 14:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2318-2331].codfw.wmnet * 14:34 cgoubert@deploy1003: Started scap sync-world: Backport for [[gerrit:1326306{{!}}Add placeholder $wmgRedisLockPassword (T366938 T427999)]] * 14:29 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2318-2331].codfw.wmnet * 14:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2304-2317].codfw.wmnet * 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2304-2317].codfw.wmnet * 14:28 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1054 * 14:27 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1054 * 14:27 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:27 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1054] - vriley@cumin1003" * 14:27 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1054] - vriley@cumin1003" * 14:26 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1189.eqiad.wmnet with reason: host reimage * 14:26 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test2001.codfw.wmnet * 14:25 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test1001.eqiad.wmnet * 14:25 claime: Deploying wmgRedisLockPassword - [[phab:T366938|T366938]] [[phab:T427999|T427999]] * 14:24 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1051.eqiad.wmnet with OS trixie * 14:23 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 14:22 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1182.eqiad.wmnet with reason: host reimage * 14:22 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1051.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:22 vriley@cumin1003: START - Cookbook sre.dns.netbox * 14:22 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test2001.codfw.wmnet * 14:21 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test1001.eqiad.wmnet * 14:21 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 14:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2304-2317].codfw.wmnet * 14:19 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns4004.wikimedia.org with OS trixie * 14:19 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns2006.wikimedia.org with OS trixie * 14:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1143.eqiad.wmnet with reason: host reimage * 14:19 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns1006.wikimedia.org with OS trixie * 14:17 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1189.eqiad.wmnet with reason: host reimage * 14:15 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1182.eqiad.wmnet with reason: host reimage * 14:14 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1143.eqiad.wmnet with reason: host reimage * 14:13 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1051.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:12 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2304-2317].codfw.wmnet * 14:12 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2290-2303].codfw.wmnet * 14:12 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1049.eqiad.wmnet with OS trixie * 14:12 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2290-2303].codfw.wmnet * 14:12 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1051 * 14:11 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1051 * 14:11 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1170.eqiad.wmnet onto db1284.eqiad.wmnet * 14:11 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1170: Pool db1170.eqiad.wmnet in after cloning * 14:09 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 14:09 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:09 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1051] - vriley@cumin1003" * 14:09 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1051] - vriley@cumin1003" * 14:06 klausman@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:04 vriley@cumin1003: START - Cookbook sre.dns.netbox * 14:04 klausman@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2290-2303].codfw.wmnet * 14:02 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1189.eqiad.wmnet with OS bookworm * 14:02 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1188.eqiad.wmnet with OS bookworm * 14:01 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 14:00 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1182.eqiad.wmnet with OS bookworm * 14:00 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1143.eqiad.wmnet with OS bookworm * 13:59 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 13:57 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply * 13:57 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply * 13:56 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply * 13:56 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-ulsfo or A:cp-drmrs and A:cp - 9.2.15 upgrade ([[phab:T434620|T434620]]) * 13:56 cjd91: sudo -i cookbook sre.cdn.roll-upgrade-ats --query 'A:cp-ulsfo or A:cp-drmrs' --task-id [[phab:T434620|T434620]] --reason '9.2.15 upgrade' * 13:56 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply * 13:55 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply * 13:55 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply * 13:54 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 13:54 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 13:54 phuedx@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:53 phuedx@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics-external: apply * 13:52 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1049.eqiad.wmnet with reason: host reimage * 13:51 phuedx@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2290-2303].codfw.wmnet * 13:50 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1047.eqiad.wmnet * 13:50 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2276-2289].codfw.wmnet * 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2276-2289].codfw.wmnet * 13:49 phuedx@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics-external: apply * 13:49 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1046.eqiad.wmnet * 13:49 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1049.eqiad.wmnet with reason: host reimage * 13:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1046.eqiad.wmnet * 13:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-ncredir rolling reboot on A:ncredir and A:ncredir * 13:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 13:46 phuedx@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:44 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics-external: apply * 13:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1046.eqiad.wmnet * 13:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2276-2289].codfw.wmnet * 13:36 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1046.eqiad.wmnet * 13:36 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1045.eqiad.wmnet * 13:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1045.eqiad.wmnet * 13:34 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2276-2289].codfw.wmnet * 13:34 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1049.eqiad.wmnet with OS trixie * 13:34 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2262-2275].codfw.wmnet * 13:34 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2262-2275].codfw.wmnet * 13:32 Lucas_WMDE: UTC afternoon backport+config window done * 13:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1045.eqiad.wmnet * 13:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2262-2275].codfw.wmnet * 13:26 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1049.eqiad.wmnet with OS trixie * 13:26 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1049.eqiad.wmnet with OS trixie * 13:26 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1170: Pool db1170.eqiad.wmnet in after cloning * 13:23 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1049.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 13:23 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1045.eqiad.wmnet * 13:21 atsukoito: manually done sudo -i docker-registryctl --debug delete-tags 'docker-registry.discovery.wmnet/repos/data-engineering/airflow-dags:airflow-3.3.0-py3.11-2026-08-17-*' to remove incorrect tags * 13:19 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311141{{!}}viwiki: Set `noindex,nofollow` for User and User talk (T432311)]] (duration: 11m 11s) * 13:15 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2262-2275].codfw.wmnet * 13:14 lucaswerkmeister-wmde@deploy1003: ndkdd, lucaswerkmeister-wmde: Continuing with deployment * 13:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2248-2261].codfw.wmnet * 13:14 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2248-2261].codfw.wmnet * 13:13 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1049.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 13:12 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1049 * 13:11 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1049 * 13:10 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:10 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1049] - vriley@cumin1003" * 13:10 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1049] - vriley@cumin1003" * 13:10 lucaswerkmeister-wmde@deploy1003: ndkdd, lucaswerkmeister-wmde: Backport for [[gerrit:1311141{{!}}viwiki: Set `noindex,nofollow` for User and User talk (T432311)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1037.eqiad.wmnet * 13:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1037.eqiad.wmnet * 13:08 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1311141{{!}}viwiki: Set `noindex,nofollow` for User and User talk (T432311)]] * 13:06 vriley@cumin1003: START - Cookbook sre.dns.netbox * 13:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2248-2261].codfw.wmnet * 13:00 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1037.eqiad.wmnet * 12:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2248-2261].codfw.wmnet * 12:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2204-2215,2242-2243].codfw.wmnet * 12:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2204-2215,2242-2243].codfw.wmnet * 12:49 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324832{{!}}Migrate $wgFlaggedRevsTags from flaggedrevs.php to ext-FlaggedRevs.php]] (duration: 14m 02s) * 12:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint2001.codfw.wmnet * 12:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2204-2215,2242-2243].codfw.wmnet * 12:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint2001.codfw.wmnet * 12:41 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1037.eqiad.wmnet * 12:40 ladsgroup@deploy1003: Rolling back deployment * 12:37 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1324832{{!}}Migrate $wgFlaggedRevsTags from flaggedrevs.php to ext-FlaggedRevs.php]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:37 seanleong-wmde: Finished populateSitesTable for [bolwiki] ([[[phab:T429955|T429955]]]) * 12:35 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1324832{{!}}Migrate $wgFlaggedRevsTags from flaggedrevs.php to ext-FlaggedRevs.php]] * 12:35 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2204-2215,2242-2243].codfw.wmnet * 12:34 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2190-2203].codfw.wmnet * 12:34 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2190-2203].codfw.wmnet * 12:32 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1028.eqiad.wmnet * 12:32 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1028.eqiad.wmnet * 12:31 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint1001.eqiad.wmnet * 12:30 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1170: Depool db1170.eqiad.wmnet to then clone it to db1284.eqiad.wmnet - marostegui@cumin1003 * 12:28 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint1001.eqiad.wmnet * 12:28 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1170: Depool db1170.eqiad.wmnet to then clone it to db1284.eqiad.wmnet - marostegui@cumin1003 * 12:27 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1170.eqiad.wmnet onto db1284.eqiad.wmnet * 12:27 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db2901.codfw.wmnet * 12:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2190-2203].codfw.wmnet * 12:27 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 12:26 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1028.eqiad.wmnet * 12:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2190-2203].codfw.wmnet * 12:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2172-2179,2184-2189].codfw.wmnet * 12:18 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2172-2179,2184-2189].codfw.wmnet * 12:13 seanleong-wmde@deploy1003: mwscript-k8s job started: foreachwikiindblist wikidataclient extensions/Wikibase/lib/maintenance/populateSitesTable.php --force-protocol https # [[phab:T429955|T429955]] * 12:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2172-2179,2184-2189].codfw.wmnet * 12:09 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1028.eqiad.wmnet * 12:03 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 12:02 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db2901.codfw.wmnet * 12:02 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db2901.codfw.wmnet * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1027.eqiad.wmnet * 12:02 fceratto@cumin1003: END (ERROR) - Cookbook sre.dns.netbox (exit_code=97) * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1027.eqiad.wmnet * 12:02 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 12:02 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db2901.codfw.wmnet * 12:01 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2172-2179,2184-2189].codfw.wmnet * 12:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2158-2171].codfw.wmnet * 12:01 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2158-2171].codfw.wmnet * 11:58 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db1902.eqiad.wmnet * 11:58 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 11:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1027.eqiad.wmnet * 11:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2158-2171].codfw.wmnet * 11:51 jayme: updated calico to v3.30.7 on wikikube eqiad - [[phab:T427400|T427400]] * 11:50 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1027.eqiad.wmnet * 11:45 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2158-2171].codfw.wmnet * 11:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2144-2157].codfw.wmnet * 11:44 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2144-2157].codfw.wmnet * 11:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1058.eqiad.wmnet * 11:43 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1058.eqiad.wmnet * 11:43 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 11:43 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 11:43 marostegui@cumin1003: Removing db1153 from zarcillo [[phab:T434638|T434638]] * 11:42 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1153.eqiad.wmnet * 11:42 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:42 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1153.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 11:42 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1172.eqiad.wmnet onto db1286.eqiad.wmnet * 11:42 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1172: Pool db1172.eqiad.wmnet in after cloning * 11:42 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1153.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 11:41 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1902.eqiad.wmnet * 11:41 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=97) for new host db2901.codfw.wmnet * 11:41 fceratto@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host db2901.codfw.wmnet with OS trixie * 11:38 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'. * 11:38 marostegui@cumin1003: START - Cookbook sre.dns.netbox * 11:37 marostegui@dns1004: END - running authdns-update * 11:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1058.eqiad.wmnet * 11:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2144-2157].codfw.wmnet * 11:35 marostegui@dns1004: START - running authdns-update * 11:32 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1153.eqiad.wmnet * 11:32 marostegui@cumin1003: START - Cookbook sre.mysql.decommission * 11:28 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326249{{!}}ImagePage: move TOC element below file link (T332644)]] (duration: 09m 56s) * 11:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2144-2157].codfw.wmnet * 11:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2130-2143].codfw.wmnet * 11:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2130-2143].codfw.wmnet * 11:26 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1058.eqiad.wmnet * 11:23 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 11:22 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'. * 11:22 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'. * 11:22 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'. * 11:22 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326249{{!}}ImagePage: move TOC element below file link (T332644)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1057.eqiad.wmnet * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1057.eqiad.wmnet * 11:21 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 11:20 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 11:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2130-2143].codfw.wmnet * 11:19 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326249{{!}}ImagePage: move TOC element below file link (T332644)]] * 11:18 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply * 11:17 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply * 11:17 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply * 11:16 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 11:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1057.eqiad.wmnet * 11:15 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 11:14 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 11:13 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 11:13 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db2901.codfw.wmnet with OS trixie * 11:12 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db2901.codfw.wmnet - fceratto@cumin1003" * 11:12 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db2901.codfw.wmnet - fceratto@cumin1003" * 11:12 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db2901.codfw.wmnet on all recursors * 11:12 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db2901.codfw.wmnet on all recursors * 11:12 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:12 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db2901.codfw.wmnet - fceratto@cumin1003" * 11:12 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2130-2143].codfw.wmnet * 11:11 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2107-2115,2124-2129].codfw.wmnet * 11:11 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2107-2115,2124-2129].codfw.wmnet * 11:11 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 11:11 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 11:11 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply * 11:11 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 11:10 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 11:06 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db2901.codfw.wmnet - fceratto@cumin1003" * 11:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2107-2115,2124-2129].codfw.wmnet * 10:57 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1172: Pool db1172.eqiad.wmnet in after cloning * 10:54 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2107-2115,2124-2129].codfw.wmnet * 10:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2078,2087-2095,2102-2106].codfw.wmnet * 10:54 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2078,2087-2095,2102-2106].codfw.wmnet * 10:47 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply * 10:46 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply * 10:46 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply * 10:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2078,2087-2095,2102-2106].codfw.wmnet * 10:45 blake@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply * 10:38 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1057.eqiad.wmnet * 10:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1056.eqiad.wmnet * 10:38 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1056.eqiad.wmnet * 10:37 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2078,2087-2095,2102-2106].codfw.wmnet * 10:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2061-2062,2064-2065,2067-2077].codfw.wmnet * 10:36 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2061-2062,2064-2065,2067-2077].codfw.wmnet * 10:32 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1056.eqiad.wmnet * 10:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2061-2062,2064-2065,2067-2077].codfw.wmnet * 10:25 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:25 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db2901.codfw.wmnet * 10:25 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db1902.eqiad.wmnet * 10:25 fceratto@cumin1003: END (ERROR) - Cookbook sre.dns.netbox (exit_code=97) * 10:24 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:24 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1902.eqiad.wmnet * 10:24 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=97) for new host db1902.eqiad.wmnet * 10:24 fceratto@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host db1902.eqiad.wmnet with OS trixie * 10:24 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=93) for new host db1903.eqiad.wmnet * 10:24 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 10:20 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1056.eqiad.wmnet * 10:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2061-2062,2064-2065,2067-2077].codfw.wmnet * 10:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2038-2039,2041-2042,2044,2046,2049-2051,2055-2060].codfw.wmnet * 10:18 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2038-2039,2041-2042,2044,2046,2049-2051,2055-2060].codfw.wmnet * 10:15 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1055.eqiad.wmnet * 10:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1055.eqiad.wmnet * 10:13 Amir1: mwscript-k8s --dblist=all -- purgeUserOptions.php --login-age 5 uls-preferences ([[phab:T406724|T406724]]) * 10:11 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:10 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1903.eqiad.wmnet on all recursors * 10:10 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1903.eqiad.wmnet on all recursors * 10:10 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2038-2039,2041-2042,2044,2046,2049-2051,2055-2060].codfw.wmnet * 10:10 moritzm: installing unzip security updates * 10:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1055.eqiad.wmnet * 10:08 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=97) for new host db1901.eqiad.wmnet * 10:08 fceratto@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host db1901.eqiad.wmnet with OS trixie * 10:08 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:08 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 10:08 fceratto@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1903.eqiad.wmnet - fceratto@cumin1003" * 10:07 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db1902.eqiad.wmnet with OS trixie * 10:07 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1902.eqiad.wmnet - fceratto@cumin1003" * 10:07 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1902.eqiad.wmnet - fceratto@cumin1003" * 10:04 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply * 10:04 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326227{{!}}Enable desktop/native lazy loading everywhere (T148047)]] (duration: 07m 13s) * 10:03 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1902.eqiad.wmnet on all recursors * 10:03 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1902.eqiad.wmnet on all recursors * 10:03 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:03 blake@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply * 10:01 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1903.eqiad.wmnet - fceratto@cumin1003" * 10:01 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2038-2039,2041-2042,2044,2046,2049-2051,2055-2060].codfw.wmnet * 10:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2002,2005-2006,2011-2015,2017-2018,2033-2037].codfw.wmnet * 10:01 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2002,2005-2006,2011-2015,2017-2018,2033-2037].codfw.wmnet * 10:00 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:00 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 09:59 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 09:59 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1055.eqiad.wmnet * 09:58 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326227{{!}}Enable desktop/native lazy loading everywhere (T148047)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:56 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326227{{!}}Enable desktop/native lazy loading everywhere (T148047)]] * 09:53 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2002,2005-2006,2011-2015,2017-2018,2033-2037].codfw.wmnet * 09:50 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:48 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1903.eqiad.wmnet * 09:48 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:48 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1902.eqiad.wmnet * 09:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 09:44 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2002,2005-2006,2011-2015,2017-2018,2033-2037].codfw.wmnet * 09:43 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-codfw * 09:43 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1174.eqiad.wmnet onto db1288.eqiad.wmnet * 09:42 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1174: Pool db1174.eqiad.wmnet in after cloning * 09:40 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1172: Depool db1172.eqiad.wmnet to then clone it to db1286.eqiad.wmnet - marostegui@cumin1003 * 09:39 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1172: Depool db1172.eqiad.wmnet to then clone it to db1286.eqiad.wmnet - marostegui@cumin1003 * 09:39 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1172.eqiad.wmnet onto db1286.eqiad.wmnet * 09:33 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1175.eqiad.wmnet onto db1289.eqiad.wmnet * 09:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1175: Pool db1175.eqiad.wmnet in after cloning * 09:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1201.eqiad.wmnet onto db1287.eqiad.wmnet * 09:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1201: Pool db1201.eqiad.wmnet in after cloning * 09:28 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db1901.eqiad.wmnet with OS trixie * 09:27 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:27 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:27 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1901.eqiad.wmnet on all recursors * 09:27 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1901.eqiad.wmnet on all recursors * 09:26 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:26 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:26 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:15 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:15 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1901.eqiad.wmnet * 09:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw1001.wikimedia.org with OS trixie * 08:57 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1174: Pool db1174.eqiad.wmnet in after cloning * 08:54 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1054.eqiad.wmnet * 08:54 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1054.eqiad.wmnet * 08:48 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1054.eqiad.wmnet * 08:47 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1175: Pool db1175.eqiad.wmnet in after cloning * 08:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 08:46 marostegui@cumin1003: Removing db1152 from zarcillo [[phab:T434480|T434480]] * 08:46 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1152.eqiad.wmnet * 08:46 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:46 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1152.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 08:46 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1201: Pool db1201.eqiad.wmnet in after cloning * 08:46 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1152.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 08:46 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1054.eqiad.wmnet * 08:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1053.eqiad.wmnet * 08:43 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1053.eqiad.wmnet * 08:42 marostegui@cumin1003: START - Cookbook sre.dns.netbox * 08:38 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage * 08:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1053.eqiad.wmnet * 08:36 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1152.eqiad.wmnet * 08:36 marostegui@cumin1003: START - Cookbook sre.mysql.decommission * 08:35 phuedx: UTC morning backport window done * 08:35 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1053.eqiad.wmnet * 08:34 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1035.eqiad.wmnet * 08:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1035.eqiad.wmnet * 08:34 phuedx@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324725{{!}}EventStreamConfig: Mark product_metrics.web_base and .web_base_with_ip as Test Kitchen streams (T429898 T430322)]], [[gerrit:1313923{{!}}EventStreamConfig: Remove unused web_ui_scroll* streams (T415370)]], [[gerrit:1325546{{!}}EventStreamConfig: Remove Watchlist click stream (T434790)]] (duration: 12m 42s) * 08:33 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1281: Pool back * 08:32 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage * 08:31 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on db2209.codfw.wmnet with reason: Maintenance * 08:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2209: Maintenance needed * 08:30 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2209: Maintenance needed * 08:26 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1035.eqiad.wmnet * 08:26 phuedx@deploy1003: bearloga, phuedx: Continuing with deployment * 08:23 phuedx@deploy1003: bearloga, phuedx: Backport for [[gerrit:1324725{{!}}EventStreamConfig: Mark product_metrics.web_base and .web_base_with_ip as Test Kitchen streams (T429898 T430322)]], [[gerrit:1313923{{!}}EventStreamConfig: Remove unused web_ui_scroll* streams (T415370)]], [[gerrit:1325546{{!}}EventStreamConfig: Remove Watchlist click stream (T434790)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug * 08:21 phuedx@deploy1003: Started scap sync-world: Backport for [[gerrit:1324725{{!}}EventStreamConfig: Mark product_metrics.web_base and .web_base_with_ip as Test Kitchen streams (T429898 T430322)]], [[gerrit:1313923{{!}}EventStreamConfig: Remove unused web_ui_scroll* streams (T415370)]], [[gerrit:1325546{{!}}EventStreamConfig: Remove Watchlist click stream (T434790)]] * 08:19 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw1001.wikimedia.org with OS trixie * 08:18 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1035.eqiad.wmnet * 08:17 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1032.eqiad.wmnet * 08:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1032.eqiad.wmnet * 08:16 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1279: Pool back * 08:15 phuedx@deploy1003: Finished scap sync-world: Backport for [[gerrit:1216721{{!}}viwikivoyage: enable relatedarticle and pop-up (T405724)]] (duration: 39m 12s) * 08:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1032.eqiad.wmnet * 08:09 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1032.eqiad.wmnet * 08:08 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1031.eqiad.wmnet * 08:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1031.eqiad.wmnet * 08:03 godog: switch production to use dumps-nfs.w.o - [[phab:T432212|T432212]] * 08:02 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1031.eqiad.wmnet * 08:02 phuedx@deploy1003: nvdtn19, phuedx: Continuing with deployment * 08:00 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1201: Depool db1201.eqiad.wmnet to then clone it to db1287.eqiad.wmnet - marostegui@cumin1003 * 08:00 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1201: Depool db1201.eqiad.wmnet to then clone it to db1287.eqiad.wmnet - marostegui@cumin1003 * 08:00 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1201.eqiad.wmnet onto db1287.eqiad.wmnet * 08:00 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1057.eqiad.wmnet * 08:00 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:00 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1057.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:59 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1057.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:59 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1275: Pool back * 07:55 filippo@cumin1003: START - Cookbook sre.dns.netbox * 07:55 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1031.eqiad.wmnet * 07:52 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1030.eqiad.wmnet * 07:52 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1030.eqiad.wmnet * 07:52 phuedx@deploy1003: nvdtn19, phuedx: Backport for [[gerrit:1216721{{!}}viwikivoyage: enable relatedarticle and pop-up (T405724)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:51 tappof: bump space for prometheus k8s-aux in eqiad * 07:50 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1057.eqiad.wmnet * 07:48 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1281: Pool back * 07:47 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1281 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96108 and previous config saved to /var/cache/conftool/dbconfig/20260817-074749-marostegui.json * 07:46 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1030.eqiad.wmnet * 07:42 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1030.eqiad.wmnet * 07:41 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1174: Depool db1174.eqiad.wmnet to then clone it to db1288.eqiad.wmnet - marostegui@cumin1003 * 07:41 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1174: Depool db1174.eqiad.wmnet to then clone it to db1288.eqiad.wmnet - marostegui@cumin1003 * 07:41 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1174.eqiad.wmnet onto db1288.eqiad.wmnet * 07:40 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1029.eqiad.wmnet * 07:40 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm2001.wikimedia.org * 07:40 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1029.eqiad.wmnet * 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1056.eqiad.wmnet * 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1056.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:38 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1056.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:36 phuedx@deploy1003: Started scap sync-world: Backport for [[gerrit:1216721{{!}}viwikivoyage: enable relatedarticle and pop-up (T405724)]] * 07:36 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm2001.wikimedia.org * 07:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1279 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96104 and previous config saved to /var/cache/conftool/dbconfig/20260817-073542-marostegui.json * 07:34 filippo@cumin1003: START - Cookbook sre.dns.netbox * 07:34 slyngshede@dns1004: END - running authdns-update * 07:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1029.eqiad.wmnet * 07:33 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm-test1001.wikimedia.org * 07:32 slyngshede@dns1004: START - running authdns-update * 07:31 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1029.eqiad.wmnet * 07:31 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1279: Pool back * 07:30 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1279 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96102 and previous config saved to /var/cache/conftool/dbconfig/20260817-073038-marostegui.json * 07:29 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm-test1001.wikimedia.org * 07:29 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm1001.wikimedia.org * 07:28 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1056.eqiad.wmnet * 07:28 moritzm: extend the disk of ldap-rw1001 by 80G [[phab:T331699|T331699]] * 07:28 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1055.eqiad.wmnet * 07:28 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:28 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1055.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:27 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1055.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:26 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1044.eqiad.wmnet * 07:26 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1044.eqiad.wmnet * 07:25 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm1001.wikimedia.org * 07:24 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1175: Depool db1175.eqiad.wmnet to then clone it to db1289.eqiad.wmnet - marostegui@cumin1003 * 07:24 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1175: Depool db1175.eqiad.wmnet to then clone it to db1289.eqiad.wmnet - marostegui@cumin1003 * 07:24 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1175.eqiad.wmnet onto db1289.eqiad.wmnet * 07:22 filippo@cumin1003: START - Cookbook sre.dns.netbox * 07:20 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1044.eqiad.wmnet * 07:16 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1055.eqiad.wmnet * 07:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1054.eqiad.wmnet * 07:15 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:15 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1054.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:15 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1044.eqiad.wmnet * 07:15 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1054.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:13 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1275: Pool back * 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1275 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96098 and previous config saved to /var/cache/conftool/dbconfig/20260817-071225-marostegui.json * 07:10 filippo@cumin1003: START - Cookbook sre.dns.netbox * 07:05 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1043.eqiad.wmnet * 07:05 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1054.eqiad.wmnet * 07:05 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1051.eqiad.wmnet * 07:05 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:05 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1051.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:05 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1043.eqiad.wmnet * 07:04 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1051.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 06:59 filippo@cumin1003: START - Cookbook sre.dns.netbox * 06:59 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin1001.eqiad.wmnet * 06:59 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1043.eqiad.wmnet * 06:59 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin2001.codfw.wmnet * 06:55 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin2001.codfw.wmnet * 06:55 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1051.eqiad.wmnet * 06:54 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1049.eqiad.wmnet * 06:54 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:54 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1049.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 06:54 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1049.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 06:54 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin1001.eqiad.wmnet * 06:53 moritzm: installing apr-util security updates * 06:52 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1043.eqiad.wmnet * 06:49 filippo@cumin1003: START - Cookbook sre.dns.netbox * 06:41 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1049.eqiad.wmnet * 06:13 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit2003.wikimedia.org * 06:13 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet * 06:07 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet * 06:06 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit2003.wikimedia.org * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 47s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-16 == * 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 01m 03s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-15 == * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 41s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-14 == * 15:38 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-staging-master-eqiad * 15:38 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster1005.eqiad.wmnet * 15:38 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster1005.eqiad.wmnet * 15:35 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sretest2009.codfw.wmnet * 15:33 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster1005.eqiad.wmnet * 15:33 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster1005.eqiad.wmnet * 15:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster1004.eqiad.wmnet * 15:32 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster1004.eqiad.wmnet * 15:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host sretest2009.codfw.wmnet * 15:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster1004.eqiad.wmnet * 15:27 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster1004.eqiad.wmnet * 15:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster1003.eqiad.wmnet * 15:27 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster1003.eqiad.wmnet * 15:24 dancy@deploy1003: Finished scap sync-world: testing (duration: 03m 23s) * 15:22 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster1003.eqiad.wmnet * 15:22 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster1003.eqiad.wmnet * 15:22 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-staging-master-eqiad * 15:20 dancy@deploy1003: Started scap sync-world: testing * 15:20 dancy@deploy1003: Installation of scap version "4.280.2" completed for 3 hosts * 15:18 dancy@deploy1003: Installing scap version "4.280.2" for 3 host(s) * 15:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sretest2006.codfw.wmnet * 14:54 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host sretest2006.codfw.wmnet * 14:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sretest2003.codfw.wmnet * 14:39 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host sretest2003.codfw.wmnet * 13:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt-staging2001.codfw.wmnet * 13:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt-staging2001.codfw.wmnet * 13:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-staging-master-codfw * 13:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster2005.codfw.wmnet * 13:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster2005.codfw.wmnet * 13:05 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox-dev2003.codfw.wmnet * 13:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster2005.codfw.wmnet * 13:04 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster2005.codfw.wmnet * 13:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster2004.codfw.wmnet * 13:04 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster2004.codfw.wmnet * 13:01 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netbox-dev2003.codfw.wmnet * 12:59 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster2004.codfw.wmnet * 12:59 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster2004.codfw.wmnet * 12:59 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster2003.codfw.wmnet * 12:59 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster2003.codfw.wmnet * 12:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster2003.codfw.wmnet * 12:54 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster2003.codfw.wmnet * 12:54 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-staging-master-codfw * 12:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-staging-worker-eqiad * 12:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage1006.eqiad.wmnet * 12:52 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage1006.eqiad.wmnet * 12:46 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage1006.eqiad.wmnet * 12:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw1001.wikimedia.org with OS trixie * 12:45 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage1006.eqiad.wmnet * 12:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage1005.eqiad.wmnet * 12:45 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage1005.eqiad.wmnet * 12:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage1005.eqiad.wmnet * 12:36 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/ratelimit: apply * 12:35 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/ratelimit: apply * 12:35 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:35 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:33 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage1005.eqiad.wmnet * 12:33 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage1004.eqiad.wmnet * 12:33 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage1004.eqiad.wmnet * 12:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage1004.eqiad.wmnet * 12:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage1004.eqiad.wmnet * 12:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage1003.eqiad.wmnet * 12:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage1003.eqiad.wmnet * 12:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage * 12:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage1003.eqiad.wmnet * 12:17 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage * 12:14 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage1003.eqiad.wmnet * 12:14 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-staging-worker-eqiad * 12:03 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw1001.wikimedia.org with OS trixie * 12:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cuminunpriv1001.eqiad.wmnet * 11:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cuminunpriv1001.eqiad.wmnet * 11:27 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1004.wikimedia.org * 11:24 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1153 from dbctl [[phab:T434638|T434638]]', diff saved to https://phabricator.wikimedia.org/P96097 and previous config saved to /var/cache/conftool/dbconfig/20260814-112449-marostegui.json * 11:21 aokoth@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1004.wikimedia.org * 11:20 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 11:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-staging-worker-codfw * 11:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2004.codfw.wmnet * 11:17 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2004.codfw.wmnet * 11:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2004.codfw.wmnet * 11:10 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2004.codfw.wmnet * 11:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2003.codfw.wmnet * 11:10 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2003.codfw.wmnet * 11:03 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-worker1181.eqiad.wmnet with OS bookworm * 11:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2003.codfw.wmnet * 11:03 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2003.codfw.wmnet * 11:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2002.codfw.wmnet * 11:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2002.codfw.wmnet * 10:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1235.eqiad.wmnet with OS bookworm * 10:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock2003.codfw.wmnet * 10:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock2003.codfw.wmnet with OS trixie * 10:56 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2002.codfw.wmnet * 10:54 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1153.eqiad.wmnet with OS bookworm * 10:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2002.codfw.wmnet * 10:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2001.codfw.wmnet * 10:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2001.codfw.wmnet * 10:50 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 10:49 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1187.eqiad.wmnet with OS bookworm * 10:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2001.codfw.wmnet * 10:44 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2001.codfw.wmnet * 10:44 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-staging-worker-codfw * 10:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock2003.codfw.wmnet with reason: host reimage * 10:38 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock2003.codfw.wmnet with reason: host reimage * 10:35 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1153.eqiad.wmnet with reason: host reimage * 10:32 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1235.eqiad.wmnet with reason: host reimage * 10:29 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1187.eqiad.wmnet with reason: host reimage * 10:24 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1153.eqiad.wmnet with reason: host reimage * 10:23 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1235.eqiad.wmnet with reason: host reimage * 10:21 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1187.eqiad.wmnet with reason: host reimage * 10:16 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock2003.codfw.wmnet with OS trixie * 10:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install1005.wikimedia.org * 10:12 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 10:12 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 10:12 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 10:12 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 10:12 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:12 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 10:12 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 10:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install1005.wikimedia.org * 10:07 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1235.eqiad.wmnet with OS bookworm * 10:07 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1187.eqiad.wmnet with OS bookworm * 10:07 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1181.eqiad.wmnet with OS bookworm * 10:07 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1153.eqiad.wmnet with OS bookworm * 10:07 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install2005.wikimedia.org * 10:03 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 10:03 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2003.codfw.wmnet * 10:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1232.eqiad.wmnet with OS bookworm * 10:00 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install2005.wikimedia.org * 10:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install3004.wikimedia.org * 09:58 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 09:58 fceratto@cumin1003: Removing db1151 from zarcillo [[phab:T434538|T434538]] * 09:56 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1231.eqiad.wmnet with OS bookworm * 09:56 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 09:53 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 09:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install3004.wikimedia.org * 09:50 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1152.eqiad.wmnet with OS bookworm * 09:49 Dreamy_Jazz: `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260808000000" --end-timestamp="20260812120000" --sleep="5" --batch-size="50"` for [[phab:T434688|T434688]] * 09:48 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host centrallog2002.codfw.wmnet * 09:48 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install4004.wikimedia.org * 09:42 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1232.eqiad.wmnet with reason: host reimage * 09:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install4004.wikimedia.org * 09:41 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host centrallog2002.codfw.wmnet * 09:39 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install5004.wikimedia.org * 09:36 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1232.eqiad.wmnet with reason: host reimage * 09:36 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'. * 09:34 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'. * 09:33 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1231.eqiad.wmnet with reason: host reimage * 09:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install5004.wikimedia.org * 09:32 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host centrallog1002.eqiad.wmnet * 09:30 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install6003.wikimedia.org * 09:30 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1231.eqiad.wmnet with reason: host reimage * 09:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1152.eqiad.wmnet with reason: host reimage * 09:25 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host centrallog1002.eqiad.wmnet * 09:25 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1152.eqiad.wmnet with reason: host reimage * 09:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install6003.wikimedia.org * 09:22 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1232 * 09:22 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1232 * 09:22 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1232 * 09:22 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1232.eqiad.wmnet 25.53.64.10.in-addr.arpa 5.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:22 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host titan1001.eqiad.wmnet * 09:22 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1232.eqiad.wmnet 25.53.64.10.in-addr.arpa 5.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:22 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:22 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1232 - btullis@cumin1003" * 09:22 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1232 - btullis@cumin1003" * 09:17 btullis@cumin1003: START - Cookbook sre.dns.netbox * 09:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install7002.wikimedia.org * 09:17 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1232 * 09:16 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1231 * 09:16 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1231 * 09:14 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1231 * 09:14 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1231.eqiad.wmnet 24.53.64.10.in-addr.arpa 4.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host titan1001.eqiad.wmnet * 09:14 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1231.eqiad.wmnet 24.53.64.10.in-addr.arpa 4.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:14 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1231 - btullis@cumin1003" * 09:14 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1231 - btullis@cumin1003" * 09:10 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install7002.wikimedia.org * 09:09 btullis@cumin1003: START - Cookbook sre.dns.netbox * 09:08 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1231 * 09:08 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1152 * 09:08 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1152 * 09:06 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1152 * 09:06 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1152.eqiad.wmnet 16.53.64.10.in-addr.arpa 6.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:06 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1152.eqiad.wmnet 16.53.64.10.in-addr.arpa 6.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:06 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:06 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1152 - btullis@cumin1003" * 09:06 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1152 - btullis@cumin1003" * 09:03 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1151.eqiad.wmnet * 09:03 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:03 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1151.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 08:55 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host titan2001.codfw.wmnet * 08:55 btullis@cumin1003: START - Cookbook sre.dns.netbox * 08:54 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1151.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 08:54 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1152 * 08:53 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1232.eqiad.wmnet with OS bookworm * 08:53 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1231.eqiad.wmnet with OS bookworm * 08:53 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1152.eqiad.wmnet with OS bookworm * 08:51 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2209: Pool back * 08:51 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1201.eqiad.wmnet * 08:50 btullis@cumin1003: START - Cookbook sre.hosts.remove-downtime for an-worker1201.eqiad.wmnet * 08:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1229.eqiad.wmnet with OS bookworm * 08:47 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host titan2001.codfw.wmnet * 08:46 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping2004.codfw.wmnet * 08:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ping2004.codfw.wmnet * 08:40 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host an-worker1230.eqiad.wmnet with OS bookworm * 08:40 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 08:34 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1151.eqiad.wmnet * 08:34 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 08:31 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host titan1002.eqiad.wmnet * 08:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1229.eqiad.wmnet with reason: host reimage * 08:25 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host titan1002.eqiad.wmnet * 08:25 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1229.eqiad.wmnet with reason: host reimage * 08:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping1004.eqiad.wmnet * 08:21 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ping1004.eqiad.wmnet * 08:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1230.eqiad.wmnet with reason: host reimage * 08:12 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1230.eqiad.wmnet with reason: host reimage * 08:11 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1229 * 08:11 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1229 * 08:11 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1229 * 08:11 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1229.eqiad.wmnet 22.53.64.10.in-addr.arpa 2.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:11 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1229.eqiad.wmnet 22.53.64.10.in-addr.arpa 2.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:11 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:11 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1229 - btullis@cumin1003" * 08:11 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1229 - btullis@cumin1003" * 08:10 btullis@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1201.eqiad.wmnet with reason: Fixing a disk * 08:07 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host titan2002.codfw.wmnet * 08:05 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2209: Pool back * 08:05 btullis@cumin1003: START - Cookbook sre.dns.netbox * 08:00 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host titan2002.codfw.wmnet * 07:59 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1229 * 07:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1230 * 07:58 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1230 * 07:55 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1230 * 07:55 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1230.eqiad.wmnet 23.53.64.10.in-addr.arpa 3.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:55 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1230.eqiad.wmnet 23.53.64.10.in-addr.arpa 3.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:55 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:55 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1230 - btullis@cumin1003" * 07:55 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1230 - btullis@cumin1003" * 07:51 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host kubestagemaster2005.codfw.wmnet with OS trixie * 07:48 btullis@cumin1003: START - Cookbook sre.dns.netbox * 07:41 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1230 * 07:41 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1229.eqiad.wmnet with OS bookworm * 07:41 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1230.eqiad.wmnet with OS bookworm * 07:39 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1152 from dbctl [[phab:T434480|T434480]]', diff saved to https://phabricator.wikimedia.org/P96090 and previous config saved to /var/cache/conftool/dbconfig/20260814-073941-marostegui.json * 07:29 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on kubestagemaster2005.codfw.wmnet with reason: host reimage * 07:23 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on kubestagemaster2005.codfw.wmnet with reason: host reimage * 07:04 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host kubestagemaster2005.codfw.wmnet with OS trixie * 06:53 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1228.eqiad.wmnet with OS bookworm * 06:44 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1227.eqiad.wmnet with OS bookworm * 06:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1209.eqiad.wmnet with OS bookworm * 06:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1175.eqiad.wmnet with OS bookworm * 06:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1228.eqiad.wmnet with reason: host reimage * 06:27 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1228.eqiad.wmnet with reason: host reimage * 06:25 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1227.eqiad.wmnet with reason: host reimage * 06:21 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1227.eqiad.wmnet with reason: host reimage * 06:18 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1209.eqiad.wmnet with reason: host reimage * 06:14 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1209.eqiad.wmnet with reason: host reimage * 06:14 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1228 * 06:14 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1228 * 06:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1175.eqiad.wmnet with reason: host reimage * 06:12 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1228 * 06:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1228.eqiad.wmnet 20.53.64.10.in-addr.arpa 0.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:12 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1228.eqiad.wmnet 20.53.64.10.in-addr.arpa 0.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1228 - ryankemper@cumin2003" * 06:12 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1228 - ryankemper@cumin2003" * 06:09 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1175.eqiad.wmnet with reason: host reimage * 06:07 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 06:07 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1228 * 06:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1227 * 06:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1227 * 06:06 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1227 * 06:06 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1227.eqiad.wmnet 19.53.64.10.in-addr.arpa 9.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:06 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1227.eqiad.wmnet 19.53.64.10.in-addr.arpa 9.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:06 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:06 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1227 - ryankemper@cumin2003" * 06:06 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1227 - ryankemper@cumin2003" * 06:00 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 06:00 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1227 * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1209 * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1209 * 06:00 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1209 * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1209.eqiad.wmnet 15.53.64.10.in-addr.arpa 5.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:00 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1209.eqiad.wmnet 15.53.64.10.in-addr.arpa 5.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1209 - ryankemper@cumin2003" * 06:00 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1209 - ryankemper@cumin2003" * 05:54 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 05:54 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1209 * 05:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1175 * 05:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1175 * 05:52 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1175 * 05:52 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1175.eqiad.wmnet 17.53.64.10.in-addr.arpa 7.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 05:52 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1175.eqiad.wmnet 17.53.64.10.in-addr.arpa 7.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 05:52 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 05:52 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1175 - ryankemper@cumin2003" * 05:52 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1175 - ryankemper@cumin2003" * 05:49 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1228.eqiad.wmnet with OS bookworm * 05:49 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1227.eqiad.wmnet with OS bookworm * 05:48 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1209.eqiad.wmnet with OS bookworm * 05:47 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 05:47 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1175 * 05:47 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1175.eqiad.wmnet with OS bookworm * 05:09 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host apifeatureusage1001.eqiad.wmnet with OS bookworm * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 03s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:10 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 01:07 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 01:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 01:02 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 00:59 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1226.eqiad.wmnet with OS bookworm * 00:47 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1225.eqiad.wmnet with OS bookworm * 00:41 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1224.eqiad.wmnet with OS bookworm * 00:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1226.eqiad.wmnet with reason: host reimage * 00:31 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1226.eqiad.wmnet with reason: host reimage * 00:28 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1225.eqiad.wmnet with reason: host reimage * 00:25 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1225.eqiad.wmnet with reason: host reimage * 00:19 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1224.eqiad.wmnet with reason: host reimage * 00:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1226 * 00:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1226 * 00:17 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1226 * 00:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1226.eqiad.wmnet 23.36.64.10.in-addr.arpa 3.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:17 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1226.eqiad.wmnet 23.36.64.10.in-addr.arpa 3.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 00:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1226 - ryankemper@cumin2003" * 00:17 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1226 - ryankemper@cumin2003" * 00:16 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1224.eqiad.wmnet with reason: host reimage * 00:12 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 00:12 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1226 * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1225 * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1225 * 00:10 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1225 * 00:10 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1225.eqiad.wmnet 22.36.64.10.in-addr.arpa 2.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:10 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1225.eqiad.wmnet 22.36.64.10.in-addr.arpa 2.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:10 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 00:10 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1225 - ryankemper@cumin2003" * 00:10 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1225 - ryankemper@cumin2003" * 00:03 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 00:02 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1225 * 00:02 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1224 * 00:02 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1224 * 00:00 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1224 * 00:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1224.eqiad.wmnet 21.36.64.10.in-addr.arpa 1.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:00 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1224.eqiad.wmnet 21.36.64.10.in-addr.arpa 1.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 00:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1224 - ryankemper@cumin2003" * 00:00 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1224 - ryankemper@cumin2003" == 2026-08-13 == * 23:54 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1226.eqiad.wmnet with OS bookworm * 23:53 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1225.eqiad.wmnet with OS bookworm * 23:52 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 23:51 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1224 * 23:51 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1224.eqiad.wmnet with OS bookworm * 23:47 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1223.eqiad.wmnet with OS bookworm * 23:29 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1223.eqiad.wmnet with reason: host reimage * 23:24 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1223.eqiad.wmnet with reason: host reimage * 23:19 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325525{{!}}ve.ui.CodeMirror.less: ensure normal font style]] (duration: 11m 40s) * 23:16 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260807000000" --end-timestamp="20260808000000" --sleep="5" --batch-size="50"` for [[phab:T434688|T434688]] * 23:13 musikanimal@deploy1003: musikanimal: Continuing with deployment * 23:11 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1325525{{!}}ve.ui.CodeMirror.less: ensure normal font style]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1223 * 23:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1223 * 23:08 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1325525{{!}}ve.ui.CodeMirror.less: ensure normal font style]] * 23:07 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1223 * 23:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1223.eqiad.wmnet 20.36.64.10.in-addr.arpa 0.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:07 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1223.eqiad.wmnet 20.36.64.10.in-addr.arpa 0.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 23:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1223 - ryankemper@cumin2003" * 23:03 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1223 - ryankemper@cumin2003" * 23:01 sbassett: Deployed security updates for [[phab:T430596|T430596]], [[phab:T120386|T120386]] * 22:55 bking@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host kubestagemaster2005.codfw.wmnet * 22:55 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host kubestagemaster2005.codfw.wmnet with OS bookworm * 22:54 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 22:53 ryankemper@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 22:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on kubestagemaster2005.codfw.wmnet with reason: host reimage * 22:49 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1222.eqiad.wmnet with OS bookworm * 22:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on kubestagemaster2005.codfw.wmnet with reason: host reimage * 22:29 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1222.eqiad.wmnet with reason: host reimage * 22:26 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1222.eqiad.wmnet with reason: host reimage * 22:23 bking@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host aux-k8s-etcd2003.codfw.wmnet * 22:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd2003.codfw.wmnet with OS bookworm * 22:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host kubestagemaster2005.codfw.wmnet with OS bookworm * 22:23 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM kubestagemaster2005.codfw.wmnet - bking@cumin2003" * 22:23 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM kubestagemaster2005.codfw.wmnet - bking@cumin2003" * 22:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) kubestagemaster2005.codfw.wmnet on all recursors * 22:22 bking@cumin2003: START - Cookbook sre.dns.wipe-cache kubestagemaster2005.codfw.wmnet on all recursors * 22:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:22 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM kubestagemaster2005.codfw.wmnet - bking@cumin2003" * 22:22 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM kubestagemaster2005.codfw.wmnet - bking@cumin2003" * 22:17 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 22:13 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1223 * 22:12 bking@cumin2003: START - Cookbook sre.dns.netbox * 22:12 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host kubestagemaster2005.codfw.wmnet * 22:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1222 * 22:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1222 * 22:11 sbassett: Deployed security updates for [[phab:T429244|T429244]], [[phab:T434039|T434039]] * 22:11 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1222 * 22:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1222.eqiad.wmnet 19.36.64.10.in-addr.arpa 9.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:11 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1222.eqiad.wmnet 19.36.64.10.in-addr.arpa 9.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1222 - ryankemper@cumin2003" * 22:07 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1222 - ryankemper@cumin2003" * 22:05 robh@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-wdqs2001.codfw.wmnet with reason: updating firmware * 22:01 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 22:01 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1223.eqiad.wmnet with OS bookworm * 22:01 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1222 * 22:01 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1222.eqiad.wmnet with OS bookworm * 22:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1212.eqiad.wmnet with OS bookworm * 21:54 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324295{{!}}Revert "Lazily reject pre-fix parser-cache entries for noreferrer/noopener links" (T429090)]] (duration: 06m 42s) * 21:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd2003.codfw.wmnet with reason: host reimage * 21:53 bking@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host dse-k8s-etcd2001.codfw.wmnet * 21:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dse-k8s-etcd2001.codfw.wmnet with OS bookworm * 21:50 sbassett@deploy1003: sbassett, kharlan: Continuing with deployment * 21:49 sbassett@deploy1003: sbassett, kharlan: Backport for [[gerrit:1324295{{!}}Revert "Lazily reject pre-fix parser-cache entries for noreferrer/noopener links" (T429090)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:48 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aux-k8s-etcd2003.codfw.wmnet with reason: host reimage * 21:47 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1324295{{!}}Revert "Lazily reject pre-fix parser-cache entries for noreferrer/noopener links" (T429090)]] * 21:42 maryum: Deployed security patch for [[phab:T434549|T434549]] * 21:39 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1212.eqiad.wmnet with reason: host reimage * 21:34 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1212.eqiad.wmnet with reason: host reimage * 21:33 bking@cumin2003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd2003.codfw.wmnet with OS bookworm * 21:32 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM aux-k8s-etcd2003.codfw.wmnet - bking@cumin2003" * 21:32 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM aux-k8s-etcd2003.codfw.wmnet - bking@cumin2003" * 21:32 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) aux-k8s-etcd2003.codfw.wmnet on all recursors * 21:32 bking@cumin2003: START - Cookbook sre.dns.wipe-cache aux-k8s-etcd2003.codfw.wmnet on all recursors * 21:32 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:32 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM aux-k8s-etcd2003.codfw.wmnet - bking@cumin2003" * 21:31 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM aux-k8s-etcd2003.codfw.wmnet - bking@cumin2003" * 21:28 maryum: Deployed security patch for [[phab:T434619|T434619]] * 21:26 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:26 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host aux-k8s-etcd2003.codfw.wmnet * 21:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-etcd2001.codfw.wmnet with reason: host reimage * 21:20 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1212 * 21:20 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1212 * 21:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1212.eqiad.wmnet with OS bookworm * 21:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on dse-k8s-etcd2001.codfw.wmnet with reason: host reimage * 21:04 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325553{{!}}InstrumentConstructiveEdits: exclude mw-reverted as well (T431493)]] (duration: 06m 25s) * 20:59 kemayo@deploy1003: kemayo: Continuing with deployment * 20:59 kemayo@deploy1003: kemayo: Backport for [[gerrit:1325553{{!}}InstrumentConstructiveEdits: exclude mw-reverted as well (T431493)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host dse-k8s-etcd2001.codfw.wmnet with OS bookworm * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 20:57 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 20:57 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1325553{{!}}InstrumentConstructiveEdits: exclude mw-reverted as well (T431493)]] * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-etcd2001.codfw.wmnet on all recursors * 20:57 bking@cumin2003: START - Cookbook sre.dns.wipe-cache dse-k8s-etcd2001.codfw.wmnet on all recursors * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 20:57 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 20:54 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320163{{!}}Add configurable RestTermsOfServiceUrl (T428147)]] (duration: 21m 39s) * 20:53 bking@cumin2003: START - Cookbook sre.dns.netbox * 20:53 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host dse-k8s-etcd2001.codfw.wmnet * 20:50 samtar@deploy1003: samtar, milazg: Continuing with deployment * 20:35 samtar@deploy1003: samtar, milazg: Backport for [[gerrit:1320163{{!}}Add configurable RestTermsOfServiceUrl (T428147)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:33 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1320163{{!}}Add configurable RestTermsOfServiceUrl (T428147)]] * 20:30 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325549{{!}}Deploy PRV to several LC wikis (T423785)]] (duration: 06m 57s) * 20:26 arlolra@deploy1003: arlolra: Continuing with deployment * 20:25 arlolra@deploy1003: arlolra: Backport for [[gerrit:1325549{{!}}Deploy PRV to several LC wikis (T423785)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:24 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 20:23 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1325549{{!}}Deploy PRV to several LC wikis (T423785)]] * 20:21 ariel@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319906{{!}}Remove boilerplate language from wmf-rest and wmf-math API modules (T433736)]] (duration: 13m 54s) * 20:14 ariel@deploy1003: ariel: Continuing with deployment * 20:11 ariel@deploy1003: ariel: Backport for [[gerrit:1319906{{!}}Remove boilerplate language from wmf-rest and wmf-math API modules (T433736)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 ariel@deploy1003: Started scap sync-world: Backport for [[gerrit:1319906{{!}}Remove boilerplate language from wmf-rest and wmf-math API modules (T433736)]] * 19:58 robh@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-wdqs2001.codfw.wmnet with reason: updating firmware * 19:54 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325551{{!}}Render the focused module view as a full-screen page (T433896)]] (duration: 30m 37s) * 19:52 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching aqs[2001,1016]*: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 19:44 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching aqs[2001,1016]*: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 19:42 musikanimal@deploy1003: musikanimal: Continuing with deployment * 19:41 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1325551{{!}}Render the focused module view as a full-screen page (T433896)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:34 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: quash java safepoint logspam - bking@cumin2003 - [[phab:T434685|T434685]] * 19:34 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 19:34 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 19:24 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1325551{{!}}Render the focused module view as a full-screen page (T433896)]] * 19:20 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 19:20 jhancock@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin1003" * 19:18 jhancock@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin1003" * 19:14 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325548{{!}}Enable image lazy loading on desktop in group1 (T148047)]] (duration: 07m 43s) * 19:10 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 19:10 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1325548{{!}}Enable image lazy loading on desktop in group1 (T148047)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:07 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1325548{{!}}Enable image lazy loading on desktop in group1 (T148047)]] * 19:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1166.eqiad.wmnet onto db1280.eqiad.wmnet * 19:03 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 19:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1166: Pool db1166.eqiad.wmnet in after cloning * 18:59 jhancock@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 18:58 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1234.eqiad.wmnet with OS bookworm * 18:52 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 18:51 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:49 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:46 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 18:46 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 18:44 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=dns3004.* * 18:39 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1221.eqiad.wmnet with OS bookworm * 18:36 inflatador: [bking@ganeti2048] ~$ sudo gnt-instance replace-disks -n ganeti2030.codfw.wmnet aux-k8s-worker2002.codfw.wmnet [[phab:T434681|T434681]] * 18:35 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 18:29 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1220.eqiad.wmnet with OS bookworm * 18:29 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1234.eqiad.wmnet with reason: host reimage * 18:27 bking@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host dse-k8s-etcd2001.codfw.wmnet * 18:27 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-etcd2001.codfw.wmnet on all recursors * 18:27 bking@cumin2003: START - Cookbook sre.dns.wipe-cache dse-k8s-etcd2001.codfw.wmnet on all recursors * 18:27 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:27 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 18:27 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 18:25 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1234.eqiad.wmnet with reason: host reimage * 18:25 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 18:19 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1221.eqiad.wmnet with reason: host reimage * 18:17 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1166: Pool db1166.eqiad.wmnet in after cloning * 18:15 brennen@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 18:15 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1221.eqiad.wmnet with reason: host reimage * 18:14 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:12 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:12 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 18:12 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-etcd2001.codfw.wmnet on all recursors * 18:12 bking@cumin2003: START - Cookbook sre.dns.wipe-cache dse-k8s-etcd2001.codfw.wmnet on all recursors * 18:12 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:12 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 18:12 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 18:10 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: quash java safepoint logspam - bking@cumin2003 - [[phab:T434685|T434685]] * 18:10 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1234 * 18:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1234 * 18:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1220.eqiad.wmnet with reason: host reimage * 18:07 brennen: 1.47.0-wmf.15 train status ([[phab:T430834|T430834]]) - no current blockers, rolling to all wikis * 18:07 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1234 * 18:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1234.eqiad.wmnet 10.36.64.10.in-addr.arpa 0.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:07 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:07 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1234.eqiad.wmnet 10.36.64.10.in-addr.arpa 0.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1234 - ryankemper@cumin2003" * 18:07 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1234 - ryankemper@cumin2003" * 18:05 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1220.eqiad.wmnet with reason: host reimage * 18:04 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host dse-k8s-etcd2001.codfw.wmnet * 18:04 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 18:03 dancy@deploy1003: Installation of scap version "4.280.1" completed for 3 hosts * 18:02 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:02 inflatador: bking@dse-k8s-etcd2002 etcdctl member remove $<nowiki>{</nowiki>UUID of dse-k8s-etcd2001<nowiki>}</nowiki> [[phab:T434681|T434681]] [[phab:T434793|T434793]] * 18:01 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 18:01 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1234 * 18:01 dancy@deploy1003: Installing scap version "4.280.1" for 3 host(s) * 18:01 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1221 * 18:01 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1221 * 18:00 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1221 * 18:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1221.eqiad.wmnet 18.36.64.10.in-addr.arpa 8.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:00 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1221.eqiad.wmnet 18.36.64.10.in-addr.arpa 8.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1221 - ryankemper@cumin2003" * 17:59 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:58 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 17:58 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:57 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 17:57 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:56 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1221 - ryankemper@cumin2003" * 17:53 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 17:52 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 17:52 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 17:51 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1221 * 17:51 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1220 * 17:51 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1220 * 17:51 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:51 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:51 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1220 * 17:51 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1220.eqiad.wmnet 11.36.64.10.in-addr.arpa 1.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:51 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1220.eqiad.wmnet 11.36.64.10.in-addr.arpa 1.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:51 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1220 - ryankemper@cumin2003" * 17:50 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1220 - ryankemper@cumin2003" * 17:47 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:47 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 17:46 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1234.eqiad.wmnet with OS bookworm * 17:46 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 17:45 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1221.eqiad.wmnet with OS bookworm * 17:45 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1220 * 17:45 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1220.eqiad.wmnet with OS bookworm * 17:43 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 17:41 bking@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host dse-k8s-etcd2001.codfw.wmnet * 17:41 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host dse-k8s-etcd2001.codfw.wmnet * 17:40 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:40 inflatador: bking@ganeti2048] `sudo gnt-instance remove --force --ignore-failures --shutdown-timeout=0` on non-DRBD VMs [[phab:T434681|T434681]] * 17:40 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:39 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:38 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:36 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:32 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:32 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:28 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 17:26 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:26 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:24 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:23 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:21 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1218.eqiad.wmnet with OS bookworm * 17:20 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:20 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:19 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1179.eqiad.wmnet with OS bookworm * 17:18 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1150.eqiad.wmnet with OS bookworm * 17:18 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns3004.wikimedia.org with OS trixie * 17:15 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:14 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:13 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:12 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:12 swfrench@deploy1003: Finished scap sync-world: Helmfile-only deployment for mediawiki chart bump - [[phab:T427666|T427666]] (duration: 03m 03s) * 17:09 swfrench@deploy1003: Started scap sync-world: Helmfile-only deployment for mediawiki chart bump - [[phab:T427666|T427666]] * 17:02 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:01 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:01 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1218.eqiad.wmnet with reason: host reimage * 17:01 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:01 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:00 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324810{{!}}deployment-info.php: Report dbname and branch for the requested wiki (T434726)]] (duration: 06m 52s) * 16:58 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 16:57 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1179.eqiad.wmnet with reason: host reimage * 16:56 dancy@deploy1003: dancy: Continuing with deployment * 16:56 dancy@deploy1003: dancy: Backport for [[gerrit:1324810{{!}}deployment-info.php: Report dbname and branch for the requested wiki (T434726)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1150.eqiad.wmnet with reason: host reimage * 16:53 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324810{{!}}deployment-info.php: Report dbname and branch for the requested wiki (T434726)]] * 16:51 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1166: Depool db1166.eqiad.wmnet to then clone it to db1280.eqiad.wmnet - cwilliams@cumin1003 * 16:50 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1166: Depool db1166.eqiad.wmnet to then clone it to db1280.eqiad.wmnet - cwilliams@cumin1003 * 16:50 cwilliams@cumin1003: START - Cookbook sre.mysql.clone of db1166.eqiad.wmnet onto db1280.eqiad.wmnet * 16:49 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1218.eqiad.wmnet with reason: host reimage * 16:48 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1179.eqiad.wmnet with reason: host reimage * 16:47 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1150.eqiad.wmnet with reason: host reimage * 16:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.upgrade (exit_code=0) for 1 hosts * 16:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2220: Upgrade of db2220.codfw.wmnet completed * 16:38 swfrench@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 16:38 swfrench-wmf: kubectl delete node kubestagemaster2005.codfw.wmnet - [[phab:T434681|T434681]] * 16:34 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1218 * 16:34 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1218 * 16:34 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1218.eqiad.wmnet with OS bookworm * 16:34 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1179 * 16:34 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1179 * 16:33 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1179.eqiad.wmnet with OS bookworm * 16:32 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1150 * 16:32 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1150 * 16:31 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1150.eqiad.wmnet with OS bookworm * 16:24 dancy@deploy1003: Installation of scap version "4.280.0" completed for 3 hosts * 16:22 dancy@deploy1003: Installing scap version "4.280.0" for 3 host(s) * 16:20 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: quash java safepoint logspam - bking@cumin2003 - [[phab:T434685|T434685]] * 16:14 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns3004.wikimedia.org with reason: host reimage * 16:08 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 16:07 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 16:07 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns3004.wikimedia.org with reason: host reimage * 16:05 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 16:04 swfrench@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 16:00 swfrench@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 15:59 swfrench@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 15:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Upgrade of db2220.codfw.wmnet completed * 15:48 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2220: Upgrading db2220.codfw.wmnet * 15:48 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2220: Upgrading db2220.codfw.wmnet * 15:48 cwilliams@cumin1003: START - Cookbook sre.mysql.upgrade for 1 hosts * 15:46 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns3004.wikimedia.org with OS trixie * 15:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2220 [[phab:T434802|T434802]]', diff saved to https://phabricator.wikimedia.org/P96079 and previous config saved to /var/cache/conftool/dbconfig/20260813-154624-cwilliams.json * 15:45 cdobbins@cumin1003: conftool action : set/pooled=no; selector: name=dns3004.* * 15:44 cjd91: depooling dns3004 to reimage to trixie * 15:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2159 to s7 primary [[phab:T434802|T434802]]', diff saved to https://phabricator.wikimedia.org/P96078 and previous config saved to /var/cache/conftool/dbconfig/20260813-154405-cwilliams.json * 15:43 cezmunsta: Starting s7 codfw failover from db2220 to db2159 - [[phab:T434802|T434802]] * 15:41 cgoubert@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: Dragonfly supernodes reboot (duration: 09m 42s) * 15:41 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dragonfly-supernode2001.codfw.wmnet * 15:39 inflatador: bking@ganeti2048] ~$ sudo gnt-node failover -f --ignore-consistency ganeti2046.codfw.wmnet [[phab:T434681|T434681]] * 15:39 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 15:39 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 15:39 swfrench@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 15:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2159 with weight 0 [[phab:T434802|T434802]]', diff saved to https://phabricator.wikimedia.org/P96077 and previous config saved to /var/cache/conftool/dbconfig/20260813-153806-cwilliams.json * 15:37 swfrench@dns1004: END - running authdns-update * 15:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 30 hosts with reason: Primary switchover s7 [[phab:T434802|T434802]] * 15:37 cgoubert@cumin2003: START - Cookbook sre.hosts.reboot-single for host dragonfly-supernode2001.codfw.wmnet * 15:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dragonfly-supernode1001.eqiad.wmnet * 15:35 swfrench@dns1004: START - running authdns-update * 15:32 cgoubert@cumin2003: START - Cookbook sre.hosts.reboot-single for host dragonfly-supernode1001.eqiad.wmnet * 15:32 cgoubert@deploy1003: Locking from deployment [ALL REPOSITORIES]: Dragonfly supernodes reboot * 15:30 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-master-codfw * 15:30 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2005.codfw.wmnet * 15:30 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2005.codfw.wmnet * 15:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1169.eqiad.wmnet onto db1277.eqiad.wmnet * 15:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1169: Pool db1169.eqiad.wmnet in after cloning * 15:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2005.codfw.wmnet * 15:23 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2005.codfw.wmnet * 15:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2004.codfw.wmnet * 15:23 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2004.codfw.wmnet * 15:18 inflatador: bking@ganeti2048 sudo gnt-node failover -f ganeti2046.codfw.wmnet [[phab:T434681|T434681]] * 15:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2004.codfw.wmnet * 15:17 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2004.codfw.wmnet * 15:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2003.codfw.wmnet * 15:17 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2003.codfw.wmnet * 15:15 cdobbins@dns1004: END - running authdns-update * 15:13 cdobbins@dns1004: START - running authdns-update * 15:10 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1217.eqiad.wmnet with OS bookworm * 15:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2003.codfw.wmnet * 15:10 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2003.codfw.wmnet * 15:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2002.codfw.wmnet * 15:10 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2002.codfw.wmnet * 15:10 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: quash java safepoint logspam - bking@cumin2003 - [[phab:T434685|T434685]] * 15:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1216.eqiad.wmnet with OS bookworm * 15:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2002.codfw.wmnet * 15:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2002.codfw.wmnet * 15:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2001.codfw.wmnet * 15:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2001.codfw.wmnet * 15:01 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1042.eqiad.wmnet * 15:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1042.eqiad.wmnet * 15:00 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325455{{!}}Api: Use correct query when continue prop=categories (T433922)]] (duration: 09m 47s) * 14:59 bking@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host dse-k8s-etcd2004.codfw.wmnet * 14:58 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-etcd2004.codfw.wmnet on all recursors * 14:58 bking@cumin2003: START - Cookbook sre.dns.wipe-cache dse-k8s-etcd2004.codfw.wmnet on all recursors * 14:58 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:58 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM dse-k8s-etcd2004.codfw.wmnet - bking@cumin2003" * 14:58 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM dse-k8s-etcd2004.codfw.wmnet - bking@cumin2003" * 14:55 zabe@deploy1003: zabe: Continuing with deployment * 14:53 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-etcd2004.codfw.wmnet on all recursors * 14:53 bking@cumin2003: START - Cookbook sre.dns.wipe-cache dse-k8s-etcd2004.codfw.wmnet on all recursors * 14:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:53 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2004.codfw.wmnet - bking@cumin2003" * 14:53 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2004.codfw.wmnet - bking@cumin2003" * 14:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2001.codfw.wmnet * 14:52 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2001.codfw.wmnet * 14:52 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-master-codfw * 14:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2001.codfw.wmnet * 14:52 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2001.codfw.wmnet * 14:52 zabe@deploy1003: zabe: Backport for [[gerrit:1325455{{!}}Api: Use correct query when continue prop=categories (T433922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:50 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1325455{{!}}Api: Use correct query when continue prop=categories (T433922)]] * 14:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1217.eqiad.wmnet with reason: host reimage * 14:48 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:48 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host dse-k8s-etcd2004.codfw.wmnet * 14:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-master-eqiad * 14:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1006.eqiad.wmnet * 14:44 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1006.eqiad.wmnet * 14:44 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1216.eqiad.wmnet with reason: host reimage * 14:40 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1217.eqiad.wmnet with reason: host reimage * 14:39 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1169: Pool db1169.eqiad.wmnet in after cloning * 14:39 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1216.eqiad.wmnet with reason: host reimage * 14:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl1005.eqiad.wmnet * 14:32 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl1005.eqiad.wmnet * 14:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1004.eqiad.wmnet * 14:32 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1004.eqiad.wmnet * 14:31 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus2008.codfw.wmnet * 14:31 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor1003.eqiad.wmnet * 14:29 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1042.eqiad.wmnet * 14:27 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor1003.eqiad.wmnet * 14:26 moritzm: installing Django security updates * 14:26 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling reboot on A:wikidough * 14:25 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1217.eqiad.wmnet with OS bookworm * 14:25 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1216.eqiad.wmnet with OS bookworm * 14:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl1004.eqiad.wmnet * 14:24 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl1004.eqiad.wmnet * 14:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1003.eqiad.wmnet * 14:24 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1003.eqiad.wmnet * 14:23 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus2008.codfw.wmnet * 14:23 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus1008.eqiad.wmnet * 14:23 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor-dev2001.codfw.wmnet * 14:22 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus2006.codfw.wmnet * 14:19 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor-dev2001.codfw.wmnet * 14:18 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1042.eqiad.wmnet * 14:17 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.* * 14:16 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl1003.eqiad.wmnet * 14:16 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl1003.eqiad.wmnet * 14:16 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1002.eqiad.wmnet * 14:16 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1002.eqiad.wmnet * 14:15 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus1008.eqiad.wmnet * 14:14 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus1006.eqiad.wmnet * 14:12 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus2006.codfw.wmnet * 14:12 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor2003.codfw.wmnet * 14:11 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus2007.codfw.wmnet * 14:11 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1041.eqiad.wmnet * 14:11 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1041.eqiad.wmnet * 14:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl1002.eqiad.wmnet * 14:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl1002.eqiad.wmnet * 14:09 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-master-eqiad * 14:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on apifeatureusage1001.eqiad.wmnet with reason: host reimage * 14:08 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1215.eqiad.wmnet with OS bookworm * 14:08 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor2003.codfw.wmnet * 14:08 jayme: updated calico to v3.30.7 on wikikube codfw [[phab:T427400|T427400]] * 14:07 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1149.eqiad.wmnet with OS bookworm * 14:06 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'. * 14:06 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1041.eqiad.wmnet * 14:04 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus1006.eqiad.wmnet * 14:03 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus2007.codfw.wmnet * 14:03 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus2005.codfw.wmnet * 14:03 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus1007.eqiad.wmnet * 14:03 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1002.eqiad.wmnet * 14:02 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on apifeatureusage1001.eqiad.wmnet with reason: host reimage * 14:02 cgoubert@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=helm-charts.*,name=eqiad * 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host chartmuseum1001.eqiad.wmnet * 14:00 moritzm: installing libxml2 security updates * 13:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1214.eqiad.wmnet with OS bookworm * 13:59 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns5003.* * 13:59 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1041.eqiad.wmnet * 13:59 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow1002.eqiad.wmnet * 13:58 cmooney@dns3003: END - running authdns-update * 13:57 cgoubert@cumin2003: START - Cookbook sre.hosts.reboot-single for host chartmuseum1001.eqiad.wmnet * 13:57 cgoubert@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=helm-charts.*,name=eqiad * 13:57 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: cloudelastic cluster restart - bking@cumin2003 * 13:57 cgoubert@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=helm-charts.*,name=codfw * 13:56 cmooney@dns3003: START - running authdns-update * 13:56 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'. * 13:56 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns5003.*,service=authdns-update * 13:55 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host chartmuseum2001.codfw.wmnet * 13:55 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus1007.eqiad.wmnet * 13:55 cmooney@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dns5003.wikimedia.org * 13:55 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus1005.eqiad.wmnet * 13:51 cgoubert@cumin2003: START - Cookbook sre.hosts.reboot-single for host chartmuseum2001.codfw.wmnet * 13:51 cgoubert@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=helm-charts.*,name=codfw * 13:51 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus2005.codfw.wmnet * 13:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host apifeatureusage1001.eqiad.wmnet with OS bookworm * 13:50 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus7002.magru.wmnet * 13:50 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts lvs1015.eqiad.wmnet * 13:50 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:50 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1015.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:49 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1015.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:49 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325480{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] (duration: 06m 39s) * 13:47 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1215.eqiad.wmnet with reason: host reimage * 13:46 cmooney@cumin1003: START - Cookbook sre.hosts.reboot-single for host dns5003.wikimedia.org * 13:46 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1003.eqiad.wmnet * 13:45 cmooney@cumin1003: conftool action : set/pooled=no; selector: name=dns5003.* * 13:45 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1040.eqiad.wmnet * 13:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1040.eqiad.wmnet * 13:45 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 13:45 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus1005.eqiad.wmnet * 13:45 stran@deploy1003: stran: Continuing with deployment * 13:44 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus7002.magru.wmnet * 13:44 stran@deploy1003: stran: Backport for [[gerrit:1325480{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:44 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus6002.drmrs.wmnet * 13:43 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260805000000" --end-timestamp="20260806000000" --sleep="5" --batch-size="10"` for [[phab:T434688|T434688]] * 13:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1149.eqiad.wmnet with reason: host reimage * 13:42 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1325480{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] * 13:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow1003.eqiad.wmnet * 13:40 sukhe@cumin1003: START - Cookbook sre.hosts.decommission for hosts lvs1015.eqiad.wmnet * 13:40 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts lvs1014.eqiad.wmnet * 13:40 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:40 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1014.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:40 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1040.eqiad.wmnet * 13:40 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1014.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:39 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1214.eqiad.wmnet with reason: host reimage * 13:38 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2004.codfw.wmnet * 13:38 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus6002.drmrs.wmnet * 13:38 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus5003.eqsin.wmnet * 13:35 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 13:35 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1149.eqiad.wmnet with reason: host reimage * 13:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1215.eqiad.wmnet with reason: host reimage * 13:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow2004.codfw.wmnet * 13:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1214.eqiad.wmnet with reason: host reimage * 13:33 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1040.eqiad.wmnet * 13:31 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus5003.eqsin.wmnet * 13:31 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: cloudelastic cluster restart - bking@cumin2003 * 13:31 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus4003.ulsfo.wmnet * 13:30 sukhe@cumin1003: START - Cookbook sre.hosts.decommission for hosts lvs1014.eqiad.wmnet * 13:30 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts lvs1013.eqiad.wmnet * 13:30 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:30 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1013.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:30 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1013.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:27 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1039.eqiad.wmnet * 13:27 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1039.eqiad.wmnet * 13:26 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1167.eqiad.wmnet onto db1281.eqiad.wmnet * 13:26 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1167: Pool db1167.eqiad.wmnet in after cloning * 13:25 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus4003.ulsfo.wmnet * 13:25 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260802000000" --end-timestamp="20260803000000" --sleep="5" --batch-size="10"` for [[phab:T434688|T434688]] * 13:24 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 13:24 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus3004.esams.wmnet * 13:24 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1274: New host * 13:24 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325476{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] (duration: 07m 19s) * 13:24 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260804000000" --end-timestamp="20260805000000" --sleep="5" --batch-size="10"` for [[phab:T434688|T434688]] * 13:24 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260803000000" --end-timestamp="20260804000000" --sleep="5" --batch-size="10"` for [[phab:T434688|T434688]] * 13:23 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2003.codfw.wmnet * 13:22 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1039.eqiad.wmnet * 13:20 sukhe@cumin1003: START - Cookbook sre.hosts.decommission for hosts lvs1013.eqiad.wmnet * 13:20 stran@deploy1003: stran: Continuing with deployment * 13:19 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow2003.codfw.wmnet * 13:19 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling reboot on A:wikidough * 13:19 stran@deploy1003: stran: Backport for [[gerrit:1325476{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:18 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus3004.esams.wmnet * 13:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1215.eqiad.wmnet with OS bookworm * 13:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1214.eqiad.wmnet with OS bookworm * 13:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1149.eqiad.wmnet with OS bookworm * 13:17 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1325476{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] * 13:11 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324731{{!}}prv: Enable parsoid rendering for 5 wikisource wikis]] (duration: 07m 26s) * 13:11 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1039.eqiad.wmnet * 13:11 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1036.eqiad.wmnet * 13:11 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1036.eqiad.wmnet * 13:07 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow3004.esams.wmnet * 13:07 jgiannelos@deploy1003: jgiannelos: Continuing with deployment * 13:06 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1233.eqiad.wmnet with OS bookworm * 13:06 jgiannelos@deploy1003: jgiannelos: Backport for [[gerrit:1324731{{!}}prv: Enable parsoid rendering for 5 wikisource wikis]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:04 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1324731{{!}}prv: Enable parsoid rendering for 5 wikisource wikis]] * 13:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow3004.esams.wmnet * 13:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1036.eqiad.wmnet * 13:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1210.eqiad.wmnet with OS bookworm * 13:01 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1036.eqiad.wmnet * 13:00 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1052.eqiad.wmnet * 13:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1052.eqiad.wmnet * 12:58 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow4003.ulsfo.wmnet * 12:55 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1211.eqiad.wmnet with OS bookworm * 12:55 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1052.eqiad.wmnet * 12:52 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow4003.ulsfo.wmnet * 12:51 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1052.eqiad.wmnet * 12:49 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1169: Depool db1169.eqiad.wmnet to then clone it to db1277.eqiad.wmnet - cwilliams@cumin1003 * 12:46 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1169: Depool db1169.eqiad.wmnet to then clone it to db1277.eqiad.wmnet - cwilliams@cumin1003 * 12:46 cwilliams@cumin1003: START - Cookbook sre.mysql.clone of db1169.eqiad.wmnet onto db1277.eqiad.wmnet * 12:45 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:45 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns record for deleted IP reservations eqsin lvs vlan ints - cmooney@cumin1003" * 12:44 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns record for deleted IP reservations eqsin lvs vlan ints - cmooney@cumin1003" * 12:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1051.eqiad.wmnet * 12:43 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1051.eqiad.wmnet * 12:42 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1210.eqiad.wmnet with reason: host reimage * 12:41 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1167: Pool db1167.eqiad.wmnet in after cloning * 12:40 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 12:39 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1274: New host * 12:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Added db1274', diff saved to https://phabricator.wikimedia.org/P96062 and previous config saved to /var/cache/conftool/dbconfig/20260813-123907-cwilliams.json * 12:38 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast7002.wikimedia.org * 12:38 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1233.eqiad.wmnet with reason: host reimage * 12:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1051.eqiad.wmnet * 12:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1211.eqiad.wmnet with reason: host reimage * 12:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1233.eqiad.wmnet with reason: host reimage * 12:32 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1051.eqiad.wmnet * 12:32 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast7002.wikimedia.org * 12:32 marostegui: Drop SecurePoll tables from closed wikis [[phab:T423128|T423128]] * 12:32 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow5003.eqsin.wmnet * 12:31 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1050.eqiad.wmnet * 12:31 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1050.eqiad.wmnet * 12:29 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1210.eqiad.wmnet with reason: host reimage * 12:29 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1211.eqiad.wmnet with reason: host reimage * 12:29 moritzm: installing Wireshark security updates * 12:26 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow5003.eqsin.wmnet * 12:26 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1050.eqiad.wmnet * 12:24 cmooney@dns3003: END - running authdns-update * 12:21 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow6001.drmrs.wmnet * 12:21 cmooney@dns3003: START - running authdns-update * 12:20 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1050.eqiad.wmnet * 12:20 cgoubert@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host rdb-lock2003.codfw.wmnet * 12:20 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:20 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update netbox dns entries for expanded public1-603-eqsin subnet - cmooney@cumin1003" * 12:20 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update netbox dns entries for expanded public1-603-eqsin subnet - cmooney@cumin1003" * 12:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 12:20 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 12:19 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1049.eqiad.wmnet * 12:19 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1049.eqiad.wmnet * 12:19 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 12:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host moss-be1003.eqiad.wmnet * 12:17 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow6001.drmrs.wmnet * 12:17 moritzm: installin curl security updates * 12:15 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 12:15 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1233.eqiad.wmnet with OS bookworm * 12:15 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1210.eqiad.wmnet with OS bookworm * 12:15 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1211.eqiad.wmnet with OS bookworm * 12:13 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1049.eqiad.wmnet * 12:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow7002.magru.wmnet * 12:12 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-worker1178.eqiad.wmnet with OS bookworm * 12:10 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host moss-be1003.eqiad.wmnet * 12:10 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 12:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be1006.eqiad.wmnet * 12:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 12:10 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 12:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 12:10 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 12:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow7002.magru.wmnet * 12:08 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1049.eqiad.wmnet * 12:05 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 12:05 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2003.codfw.wmnet * 12:04 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'. * 12:04 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:03 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be1006.eqiad.wmnet * 12:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be1005.eqiad.wmnet * 12:01 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1038.eqiad.wmnet * 12:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1038.eqiad.wmnet * 12:01 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1213.eqiad.wmnet with OS bookworm * 11:56 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be1005.eqiad.wmnet * 11:55 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be1004.eqiad.wmnet * 11:54 cgoubert@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host rdb-lock2003.codfw.wmnet * 11:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 11:53 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 11:53 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:53 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 11:53 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 11:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1038.eqiad.wmnet * 11:49 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be1004.eqiad.wmnet * 11:44 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:44 moritzm: installing Linux 5.10.262 on Bullseye hosts * 11:41 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1038.eqiad.wmnet * 11:40 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 11:40 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 11:40 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 11:40 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:40 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 11:40 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 11:38 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1213.eqiad.wmnet with reason: host reimage * 11:36 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1167: Depool db1167.eqiad.wmnet to then clone it to db1281.eqiad.wmnet - marostegui@cumin1003 * 11:35 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 11:35 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2003.codfw.wmnet * 11:35 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1167: Depool db1167.eqiad.wmnet to then clone it to db1281.eqiad.wmnet - marostegui@cumin1003 * 11:35 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1167.eqiad.wmnet onto db1281.eqiad.wmnet * 11:34 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 22 hosts with reason: Cloning * 11:34 moritzm: remove ganeti3005 from esams03 cluster, hardware issues [[phab:T434646|T434646]] * 11:32 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1213.eqiad.wmnet with reason: host reimage * 11:28 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1034.eqiad.wmnet * 11:28 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1034.eqiad.wmnet * 11:22 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1034.eqiad.wmnet * 11:19 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1034.eqiad.wmnet * 11:17 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1213.eqiad.wmnet with OS bookworm * 11:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1165.eqiad.wmnet onto db1279.eqiad.wmnet * 11:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1165: Pool db1165.eqiad.wmnet in after cloning * 11:07 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply * 10:57 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply * 10:54 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply * 10:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1274.eqiad.wmnet with reason: Enabling notifications and pooling * 10:45 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1033.eqiad.wmnet * 10:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1033.eqiad.wmnet * 10:44 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply * 10:43 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'. * 10:42 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'. * 10:42 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'. * 10:40 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db[1216,1225,1239-1240].eqiad.wmnet with reason: reboot * 10:39 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1033.eqiad.wmnet * 10:38 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1151 from dbctl [[phab:T434538|T434538]]', diff saved to https://phabricator.wikimedia.org/P96055 and previous config saved to /var/cache/conftool/dbconfig/20260813-103828-marostegui.json * 10:35 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1033.eqiad.wmnet * 10:27 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1165: Pool db1165.eqiad.wmnet in after cloning * 10:24 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1151.eqiad.wmnet with OS bookworm * 10:15 moritzm: installing bind9 security updates (client-side tools/libs only) * 10:07 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-debug: apply * 10:06 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-debug: apply * 10:02 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 7 hosts with reason: reboot * 10:01 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-debug: apply * 10:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.upgrade (exit_code=0) for 1 hosts * 10:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2214: Upgrade of db2214.codfw.wmnet completed * 10:01 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-debug: apply * 10:00 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-debug: apply * 10:00 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-debug: apply * 09:59 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.decommission (exit_code=99) * 09:59 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 09:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1151.eqiad.wmnet with reason: host reimage * 09:59 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 8 hosts * 09:59 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 8 hosts * 09:55 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1151.eqiad.wmnet with reason: host reimage * 09:46 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260801000000" --end-timestamp="20260802000000" --sleep="3" --batch-size="5"` for [[phab:T434688|T434688]] * 09:45 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 8 hosts with reason: reboot * 09:45 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 22 hosts with reason: Cloning * 09:43 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1165: Depool db1165.eqiad.wmnet to then clone it to db1279.eqiad.wmnet - marostegui@cumin1003 * 09:42 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1165: Depool db1165.eqiad.wmnet to then clone it to db1279.eqiad.wmnet - marostegui@cumin1003 * 09:42 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1165.eqiad.wmnet onto db1279.eqiad.wmnet * 09:41 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=testwiki --start-timestamp="20260311000000" --end-timestamp="20260805000000" --sleep="5" --batch-size="2"` for [[phab:T434688|T434688]] * 09:40 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for backupmon1001.eqiad.wmnet * 09:40 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for backupmon1001.eqiad.wmnet * 09:40 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1151 * 09:40 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1151 * 09:37 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325408{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]], [[gerrit:1325407{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]] (duration: 06m 57s) * 09:36 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1151 * 09:36 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1151.eqiad.wmnet 13.36.64.10.in-addr.arpa 3.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:36 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1151.eqiad.wmnet 13.36.64.10.in-addr.arpa 3.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:36 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:36 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1151 - btullis@cumin1003" * 09:36 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on backupmon1001.eqiad.wmnet with reason: reboot * 09:36 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1151 - btullis@cumin1003" * 09:35 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host moss-be2003.codfw.wmnet * 09:33 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 7 hosts * 09:33 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 7 hosts * 09:33 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 09:32 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1325408{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]], [[gerrit:1325407{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1161.eqiad.wmnet onto db1275.eqiad.wmnet * 09:31 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 09:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1161: Pool db1161.eqiad.wmnet in after cloning * 09:30 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1325408{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]], [[gerrit:1325407{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]] * 09:29 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 09:27 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 09:27 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host moss-be2003.codfw.wmnet * 09:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be2006.codfw.wmnet * 09:25 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2004.codfw.wmnet * 09:22 hashar@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 09:21 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be2006.codfw.wmnet * 09:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be2005.codfw.wmnet * 09:19 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2004.codfw.wmnet * 09:19 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2003.codfw.wmnet * 09:18 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 7 hosts with reason: reboot * 09:18 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 7 hosts * 09:18 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 7 hosts * 09:16 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2214: Upgrade of db2214.codfw.wmnet completed * 09:14 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be2005.codfw.wmnet * 09:13 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be2004.codfw.wmnet * 09:12 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2003.codfw.wmnet * 09:12 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2002.codfw.wmnet * 09:09 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2214: Upgrading db2214.codfw.wmnet * 09:09 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2214: Upgrading db2214.codfw.wmnet * 09:09 cwilliams@cumin1003: START - Cookbook sre.mysql.upgrade for 1 hosts * 09:07 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be2004.codfw.wmnet * 09:06 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2002.codfw.wmnet * 09:05 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1004.eqiad.wmnet * 09:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-cluster (exit_code=0) * 09:03 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 7 hosts with reason: reboot * 09:01 btullis@cumin1003: START - Cookbook sre.dns.netbox * 09:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2214 [[phab:T434754|T434754]]', diff saved to https://phabricator.wikimedia.org/P96045 and previous config saved to /var/cache/conftool/dbconfig/20260813-090001-cwilliams.json * 08:59 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1004.eqiad.wmnet * 08:59 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1003.eqiad.wmnet * 08:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2229 to s6 primary [[phab:T434754|T434754]]', diff saved to https://phabricator.wikimedia.org/P96044 and previous config saved to /var/cache/conftool/dbconfig/20260813-085752-cwilliams.json * 08:57 cezmunsta: Starting s6 codfw failover from db2214 to db2229 - [[phab:T434754|T434754]] * 08:54 hashar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325416{{!}}Revert "REST: Enable `GET /lexemes/<nowiki>{</nowiki>lexeme_id<nowiki>}</nowiki>` by default" (T434712)]] (duration: 07m 22s) * 08:53 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1003.eqiad.wmnet * 08:53 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1002.eqiad.wmnet * 08:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2229 with weight 0 [[phab:T434754|T434754]]', diff saved to https://phabricator.wikimedia.org/P96043 and previous config saved to /var/cache/conftool/dbconfig/20260813-085151-cwilliams.json * 08:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 23 hosts with reason: Primary switchover s6 [[phab:T434754|T434754]] * 08:51 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts2002.codfw.wmnet * 08:50 hashar@deploy1003: hashar: Continuing with deployment * 08:49 hashar@deploy1003: hashar: Backport for [[gerrit:1325416{{!}}Revert "REST: Enable `GET /lexemes/<nowiki>{</nowiki>lexeme_id<nowiki>}</nowiki>` by default" (T434712)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:47 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet * 08:47 hashar@deploy1003: Started scap sync-world: Backport for [[gerrit:1325416{{!}}Revert "REST: Enable `GET /lexemes/<nowiki>{</nowiki>lexeme_id<nowiki>}</nowiki>` by default" (T434712)]] * 08:47 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1002.eqiad.wmnet * 08:46 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1161: Pool db1161.eqiad.wmnet in after cloning * 08:46 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host stewards2001.codfw.wmnet * 08:45 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit1003.wikimedia.org * 08:45 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host stewards1001.eqiad.wmnet * 08:45 Emperor: roll-restart apus frontends in codfw * 08:45 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-cluster * 08:44 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts2002.codfw.wmnet * 08:44 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host doc2003.codfw.wmnet * 08:43 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet * 08:43 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host phab2003.codfw.wmnet * 08:42 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host stewards2001.codfw.wmnet * 08:42 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host etherpad2002.codfw.wmnet * 08:41 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host stewards1001.eqiad.wmnet * 08:41 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1273: Pool in s7 * 08:41 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1003.wikimedia.org * 08:40 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host doc2003.codfw.wmnet * 08:40 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit1003.wikimedia.org * 08:39 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host etherpad1004.eqiad.wmnet * 08:39 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host doc1004.eqiad.wmnet * 08:39 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit2002.wikimedia.org * 08:38 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host etherpad2002.codfw.wmnet * 08:37 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host phab2003.codfw.wmnet * 08:36 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lists1004.wikimedia.org * 08:35 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host etherpad1004.eqiad.wmnet * 08:35 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host doc1004.eqiad.wmnet * 08:34 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-cluster (exit_code=0) * 08:34 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1003.wikimedia.org * 08:34 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host planet2003.codfw.wmnet * 08:34 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2003.wikimedia.org * 08:33 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host planet1003.eqiad.wmnet * 08:33 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit2002.wikimedia.org * 08:32 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 22 hosts with reason: Cloning * 08:32 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2002.wikimedia.org * 08:31 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aphlict1002.eqiad.wmnet * 08:30 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host planet2003.codfw.wmnet * 08:29 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host planet1003.eqiad.wmnet * 08:29 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host lists1004.wikimedia.org * 08:28 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lists2001.wikimedia.org * 08:28 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2003.wikimedia.org * 08:27 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host aphlict1002.eqiad.wmnet * 08:27 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aphlict2001.codfw.wmnet * 08:26 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2002.wikimedia.org * 08:23 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host aphlict2001.codfw.wmnet * 08:23 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1151 * 08:22 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1151.eqiad.wmnet with OS bookworm * 08:22 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host lists2001.wikimedia.org * 08:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1161: Depool db1161.eqiad.wmnet to then clone it to db1275.eqiad.wmnet - marostegui@cumin1003 * 08:20 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1161: Depool db1161.eqiad.wmnet to then clone it to db1275.eqiad.wmnet - marostegui@cumin1003 * 08:20 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1161.eqiad.wmnet onto db1275.eqiad.wmnet * 08:15 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-cluster * 08:15 Emperor: roll-restart apus frontends in eqiad * 07:56 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1273: Pool in s7 * 07:56 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1273 to dbctl [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P96035 and previous config saved to /var/cache/conftool/dbconfig/20260813-075611-marostegui.json * 07:38 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db1273.eqiad.wmnet with reason: Reboot * 07:31 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.sanitize-wiki (exit_code=97) Managing sanitization for wikis testwiki in section s3 * 07:24 marostegui@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis testwiki in section s3 * 07:19 jayme@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on kubestagemaster2005.codfw.wmnet with reason: downtime because of hardware failure and no DRBD * 05:42 arnaudb@dns1006: END - running authdns-update * 05:40 arnaudb@dns1006: START - running authdns-update * 05:27 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 04:06 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 02:29 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1151 * 02:29 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1151 * 02:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1219.eqiad.wmnet with OS bookworm * 02:18 ryankemper: [[phab:T434494|T434494]] `ryankemper@deploy1003:~$ echo 'https://stats.wikimedia.org/' {{!}} mwscript-k8s --attach -- purgeList.php` (default page got cached during yesterday's `an-web1001` reimage) * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 46s) * 02:03 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1219.eqiad.wmnet with reason: host reimage * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 02:00 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1219.eqiad.wmnet with reason: host reimage * 01:46 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1219.eqiad.wmnet with OS bookworm * 01:03 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1142.eqiad.wmnet with OS bookworm * 00:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1142.eqiad.wmnet with reason: host reimage * 00:34 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1142.eqiad.wmnet with reason: host reimage * 00:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1142.eqiad.wmnet with OS bookworm * 00:16 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1178 * 00:16 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1178 == 2026-08-12 == * 23:18 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324409{{!}}Remove $wmg = $wg hacks in Collection (T119117)]] (duration: 06m 43s) * 23:14 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 23:13 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1324409{{!}}Remove $wmg = $wg hacks in Collection (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:11 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1324409{{!}}Remove $wmg = $wg hacks in Collection (T119117)]] * 22:59 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 22:49 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260801000000" --end-timestamp="20260802000000" --sleep=2 --batch-size=10` * 22:45 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=testwiki --start-timestamp="20200801010101" --end-timestamp="20260816010101" --sleep=15 --batch-size=5` * 22:40 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=testwiki --start-timestamp="20200101010101" --end-timestamp="20260816010101" --sleep=60` * 22:35 Dreamy_Jazz: Running `mwscript WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260101000000" --end-timestamp="20260102000000" --sleep=10` * 22:21 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324817{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]], [[gerrit:1324818{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]] (duration: 45m 29s) * 22:17 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 21:59 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 21:40 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1324817{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]], [[gerrit:1324818{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:39 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 21:36 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1324817{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]], [[gerrit:1324818{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]] * 21:32 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:31 vriley@cumin1003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:30 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:30 vriley@cumin1003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:17 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:14 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:14 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:11 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:10 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-codfw: Set storage compatability to NONE — [[phab:T433028|T433028]] - eevans@cumin1003 * 21:10 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1006 * 21:09 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1006 * 21:05 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324277{{!}}Improve Math preference labels for SVG/MathJax/MathML (T433891)]] (duration: 31m 42s) * 20:58 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1178.eqiad.wmnet with OS bookworm * 20:54 krinkle@deploy1003: krinkle: Continuing with deployment * 20:51 krinkle@deploy1003: krinkle: Backport for [[gerrit:1324277{{!}}Improve Math preference labels for SVG/MathJax/MathML (T433891)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:39 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 20:38 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:38 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp5022.eqsin.wmnet with OS trixie * 20:38 cdobbins@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - cdobbins@cumin1003" * 20:37 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:36 cdobbins@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - cdobbins@cumin1003" * 20:34 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1178.eqiad.wmnet with reason: host reimage * 20:34 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1324277{{!}}Improve Math preference labels for SVG/MathJax/MathML (T433891)]] * 20:33 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 20:28 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1178.eqiad.wmnet with reason: host reimage * 20:13 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1178.eqiad.wmnet with OS bookworm * 20:11 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-worker1178.eqiad.wmnet with OS bookworm * 20:11 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1178.eqiad.wmnet with OS bookworm * 20:09 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-codfw: Set storage compatability to NONE — [[phab:T433028|T433028]] - eevans@cumin1003 * 20:09 Dreamy_Jazz: Evening UTC backport window done * 20:08 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324794{{!}}WikimediaAntiAbuse: Enable logging channel (T431292)]] (duration: 06m 48s) * 20:08 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage * 20:05 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage * 20:04 dreamyjazz@deploy1003: kharlan, dreamyjazz: Continuing with deployment * 20:04 dreamyjazz@deploy1003: kharlan, dreamyjazz: Backport for [[gerrit:1324794{{!}}WikimediaAntiAbuse: Enable logging channel (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:01 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1324794{{!}}WikimediaAntiAbuse: Enable logging channel (T431292)]] * 19:55 brennen@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 19:47 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 19:35 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 19:35 cdobbins@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cp5022.eqsin.wmnet with OS trixie * 19:32 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:30 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:29 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:26 vriley@cumin1003: START - Cookbook sre.dns.netbox * 19:23 brennen@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 19:19 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-eqiad: Set storage compatability to NONE — [[phab:T433028|T433028]] - eevans@cumin1003 * 19:10 brennen@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 19:09 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 19:09 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 19:00 Amir1: data migrated on wikishared ([[phab:T426102|T426102]]) * 18:57 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321224{{!}}Rename ce_worklist_articles table to ce_invitation_list_articles (T426102)]] (duration: 06m 50s) * 18:53 ladsgroup@deploy1003: ladsgroup, daimona: Continuing with deployment * 18:53 ladsgroup@deploy1003: ladsgroup, daimona: Backport for [[gerrit:1321224{{!}}Rename ce_worklist_articles table to ce_invitation_list_articles (T426102)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:51 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1321224{{!}}Rename ce_worklist_articles table to ce_invitation_list_articles (T426102)]] * 18:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2192: Security update * 18:31 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 18:25 Amir1: ce_invitation_list_articles created as empty on wikishared ([[phab:T426102|T426102]]) * 18:21 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 18:21 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 18:18 Amir1: migrated testwiki entries from ce_worklist_articles to ce_invitation_list_articles ([[phab:T426102|T426102]]) * 18:18 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-eqiad: Set storage compatability to NONE — [[phab:T433028|T433028]] - eevans@cumin1003 * 18:11 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 18:08 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling reboot on A:durum-eqsin and A:durum * 18:07 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Set storage compatability to UPGRADING — [[phab:T433028|T433028]] - eevans@cumin1003 * 18:05 jhancock@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022'] * 17:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2192: Security update * 17:55 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:55 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum-eqsin and A:durum * 17:53 jhancock@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['cp5022'] * 17:47 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:47 jhancock@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['cp5022'] * 17:42 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:41 jhancock@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['cp5022'] * 17:36 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:35 jhancock@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022'] * 17:31 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-magru and not (P<nowiki>{</nowiki>cp7001*<nowiki>}</nowiki> or P<nowiki>{</nowiki>cp7009*<nowiki>}</nowiki>) and A:cp - 9.2.15 upgrade ([[phab:T434620|T434620]]) * 17:28 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:28 jhancock@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['cp5022'] * 17:22 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2192.codfw.wmnet with reason: Maintenance * 17:11 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host apifeatureusage2001.codfw.wmnet with OS bookworm * 17:04 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: apply * 17:03 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-main: apply * 17:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2192 [[phab:T434635|T434635]]', diff saved to https://phabricator.wikimedia.org/P96030 and previous config saved to /var/cache/conftool/dbconfig/20260812-170338-cwilliams.json * 17:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2213 to s5 primary [[phab:T434635|T434635]]', diff saved to https://phabricator.wikimedia.org/P96029 and previous config saved to /var/cache/conftool/dbconfig/20260812-170152-cwilliams.json * 17:01 cezmunsta: Starting s5 codfw failover from db2192 to db2213 - [[phab:T434635|T434635]] * 16:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2213 with weight 0 [[phab:T434635|T434635]]', diff saved to https://phabricator.wikimedia.org/P96028 and previous config saved to /var/cache/conftool/dbconfig/20260812-165544-cwilliams.json * 16:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 27 hosts with reason: Primary switchover s5 [[phab:T434635|T434635]] * 16:53 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: apply * 16:52 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-main: apply * 16:44 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-main: apply * 16:44 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-main: apply * 16:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1272: New host * 16:40 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324765{{!}}WikimediaAntiAbuse: Enable PersonalInfoFlagNotifications (T431292)]] (duration: 07m 02s) * 16:40 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-logging-external: apply * 16:39 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-logging-external: apply * 16:38 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-logging-external: apply * 16:37 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-logging-external: apply * 16:36 kharlan@deploy1003: kharlan: Continuing with deployment * 16:35 kharlan@deploy1003: kharlan: Backport for [[gerrit:1324765{{!}}WikimediaAntiAbuse: Enable PersonalInfoFlagNotifications (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:33 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1324765{{!}}WikimediaAntiAbuse: Enable PersonalInfoFlagNotifications (T431292)]] * 16:27 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-logging-external: apply * 16:27 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-logging-external: apply * 16:18 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 16:18 jhancock@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022'] * 16:17 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 16:16 jhancock@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['cp5022'] * 16:15 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 16:12 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Set storage compatability to UPGRADING — [[phab:T433028|T433028]] - eevans@cumin1003 * 16:10 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: apply * 16:10 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: apply * 16:08 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: apply * 16:08 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: apply * 16:08 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics: apply * 16:07 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics: apply * 16:02 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-magru and not (P<nowiki>{</nowiki>cp7001*<nowiki>}</nowiki> or P<nowiki>{</nowiki>cp7009*<nowiki>}</nowiki>) and A:cp - 9.2.15 upgrade ([[phab:T434620|T434620]]) * 15:57 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1272: New host * 15:52 jmm@cumin2003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti2046.codfw.wmnet * 15:52 jmm@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host ganeti2046.codfw.wmnet * 15:42 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:37 mforns@deploy1003: Finished deploy [analytics/refinery@49c336c] (thin): Regular analytics weekly train THIN [analytics/refinery@49c336cd] (duration: 01m 59s) * 15:37 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324733{{!}}WikimediaAntiAbuse: Enable personal info tag display on enwiki (T431292)]] (duration: 08m 12s) * 15:35 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:35 mforns@deploy1003: Started deploy [analytics/refinery@49c336c] (thin): Regular analytics weekly train THIN [analytics/refinery@49c336cd] * 15:34 mforns@deploy1003: Finished deploy [analytics/refinery@49c336c]: Regular analytics weekly train [analytics/refinery@49c336cd] (duration: 04m 20s) * 15:33 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[2024,1031]*.wmnet: Set storage compatability to UPGRADING — [[phab:T433028|T433028]] - eevans@cumin1003 * 15:33 dreamyjazz@deploy1003: kharlan, dreamyjazz: Continuing with deployment * 15:31 dreamyjazz@deploy1003: kharlan, dreamyjazz: Backport for [[gerrit:1324733{{!}}WikimediaAntiAbuse: Enable personal info tag display on enwiki (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:30 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitize-wiki (exit_code=99) Checking sanitization for wikis testwiki in section s3 * 15:30 mforns@deploy1003: Started deploy [analytics/refinery@49c336c]: Regular analytics weekly train [analytics/refinery@49c336cd] * 15:30 mforns@deploy1003: Finished deploy [analytics/refinery@49c336c] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@49c336cd] (duration: 00m 32s) * 15:29 mforns@deploy1003: Started deploy [analytics/refinery@49c336c] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@49c336cd] * 15:29 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1324733{{!}}WikimediaAntiAbuse: Enable personal info tag display on enwiki (T431292)]] * 15:27 brennen@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324700{{!}}EventDetailsParticipantsModule: populate cache with non-local users (T434597)]] (duration: 06m 38s) * 15:23 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[2024,1031]*.wmnet: Set storage compatability to UPGRADING — [[phab:T433028|T433028]] - eevans@cumin1003 * 15:23 brennen@deploy1003: brennen, daimona: Continuing with deployment * 15:23 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:22 brennen@deploy1003: brennen, daimona: Backport for [[gerrit:1324700{{!}}EventDetailsParticipantsModule: populate cache with non-local users (T434597)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:22 cgoubert@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host rdb-lock2003.codfw.wmnet * 15:21 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 15:21 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 15:21 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:21 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:21 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:20 brennen@deploy1003: Started scap sync-world: Backport for [[gerrit:1324700{{!}}EventDetailsParticipantsModule: populate cache with non-local users (T434597)]] * 15:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1178.eqiad.wmnet with OS bookworm * 15:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock1003.eqiad.wmnet * 15:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock1003.eqiad.wmnet with OS trixie * 15:16 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 15:16 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 15:16 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 15:16 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:16 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:16 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:12 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 15:12 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2003.codfw.wmnet * 15:11 cgoubert@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host rdb-lock2003.codfw.wmnet * 15:11 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 15:11 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 15:11 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:11 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:11 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:07 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324338{{!}}InitialiseSettings: Enable 2FA warnings on more private wikis (T428103)]] (duration: 07m 02s) * 15:04 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 15:04 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock1003.eqiad.wmnet with reason: host reimage * 15:03 reedy@deploy1003: reedy: Continuing with deployment * 15:02 reedy@deploy1003: reedy: Backport for [[gerrit:1324338{{!}}InitialiseSettings: Enable 2FA warnings on more private wikis (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:02 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:00 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324338{{!}}InitialiseSettings: Enable 2FA warnings on more private wikis (T428103)]] * 14:57 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 14:57 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2003.codfw.wmnet * 14:57 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock1003.eqiad.wmnet with reason: host reimage * 14:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock2002.codfw.wmnet * 14:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock2002.codfw.wmnet with OS trixie * 14:56 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 14:56 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 14:56 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 14:55 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 14:55 moritzm: powercycle ganeti2046 * 14:47 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock1003.eqiad.wmnet with OS trixie * 14:46 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1003.eqiad.wmnet - cgoubert@cumin2003" * 14:46 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1003.eqiad.wmnet - cgoubert@cumin2003" * 14:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock1003.eqiad.wmnet on all recursors * 14:45 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock1003.eqiad.wmnet on all recursors * 14:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1003.eqiad.wmnet - cgoubert@cumin2003" * 14:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1272.eqiad.wmnet with reason: Enabling notifications * 14:44 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1003.eqiad.wmnet - cgoubert@cumin2003" * 14:44 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324719{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324720{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324722{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0 (T434187)]] (duration: 11m 02s) * 14:39 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 14:39 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock1003.eqiad.wmnet * 14:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock2002.codfw.wmnet with reason: host reimage * 14:37 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 14:37 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock1002.eqiad.wmnet * 14:37 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock1002.eqiad.wmnet with OS trixie * 14:37 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1324719{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324720{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324722{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0 (T434187)]] synced to the testservers (see https://wikitech. * 14:33 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock2002.codfw.wmnet with reason: host reimage * 14:33 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1324719{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324720{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324722{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0 (T434187)]] * 14:32 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1018.eqiad.wmnet with OS bookworm * 14:32 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2046.codfw.wmnet * 14:31 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1020.eqiad.wmnet with OS bookworm * 14:27 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2046.codfw.wmnet * 14:25 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2045.codfw.wmnet * 14:25 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2045.codfw.wmnet * 14:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock1002.eqiad.wmnet with reason: host reimage * 14:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1019.eqiad.wmnet with OS bookworm * 14:22 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Checking sanitization for wikis testwiki in section s3 * 14:20 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2045.codfw.wmnet * 14:18 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock1002.eqiad.wmnet with reason: host reimage * 14:17 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:16 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324707{{!}}Backport all changes from wmf/1.47.0-wmf.15]] (duration: 40m 51s) * 14:16 moritzm: installing Linux 6.1.180 on Bookworm hosts * 14:15 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock2002.codfw.wmnet with OS trixie * 14:14 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2002.codfw.wmnet - cgoubert@cumin2003" * 14:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2002.codfw.wmnet - cgoubert@cumin2003" * 14:14 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:14 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2045.codfw.wmnet * 14:14 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2002.codfw.wmnet on all recursors * 14:14 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2002.codfw.wmnet on all recursors * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2002.codfw.wmnet - cgoubert@cumin2003" * 14:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2002.codfw.wmnet - cgoubert@cumin2003" * 14:12 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2030.codfw.wmnet * 14:12 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2030.codfw.wmnet * 14:11 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on apifeatureusage2001.codfw.wmnet with reason: host reimage * 14:09 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:08 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:07 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 14:06 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2030.codfw.wmnet * 14:06 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:06 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock1002.eqiad.wmnet with OS trixie * 14:05 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 14:05 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2002.codfw.wmnet * 14:05 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1002.eqiad.wmnet - cgoubert@cumin2003" * 14:05 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1002.eqiad.wmnet - cgoubert@cumin2003" * 14:05 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock1002.eqiad.wmnet on all recursors * 14:05 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock1002.eqiad.wmnet on all recursors * 14:05 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:05 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1002.eqiad.wmnet - cgoubert@cumin2003" * 14:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:04 kharlan@deploy1003: kharlan: Continuing with deployment * 14:04 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock2001.codfw.wmnet * 14:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock2001.codfw.wmnet with OS trixie * 14:02 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1002.eqiad.wmnet - cgoubert@cumin2003" * 14:02 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on apifeatureusage2001.codfw.wmnet with reason: host reimage * 14:01 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2030.codfw.wmnet * 13:59 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2029.codfw.wmnet * 13:58 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2029.codfw.wmnet * 13:58 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 13:58 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock1002.eqiad.wmnet * 13:56 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1019.eqiad.wmnet with reason: host reimage * 13:54 btullis@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on archiva1002.wikimedia.org with reason: Upgrading in-place * 13:53 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock1001.eqiad.wmnet * 13:53 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock1001.eqiad.wmnet with OS trixie * 13:53 kharlan@deploy1003: kharlan: Backport for [[gerrit:1324707{{!}}Backport all changes from wmf/1.47.0-wmf.15]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:52 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2029.codfw.wmnet * 13:52 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1020.eqiad.wmnet with reason: host reimage * 13:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1019.eqiad.wmnet with reason: host reimage * 13:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1020.eqiad.wmnet with reason: host reimage * 13:49 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock2001.codfw.wmnet with reason: host reimage * 13:48 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2029.codfw.wmnet * 13:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Configuring db1272 for s3 pooling', diff saved to https://phabricator.wikimedia.org/P96021 and previous config saved to /var/cache/conftool/dbconfig/20260812-134732-cwilliams.json * 13:44 bking@cumin2003: START - Cookbook sre.hosts.reimage for host apifeatureusage2001.codfw.wmnet with OS bookworm * 13:43 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock2001.codfw.wmnet with reason: host reimage * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2028.codfw.wmnet * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2028.codfw.wmnet * 13:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1018.eqiad.wmnet with reason: host reimage * 13:38 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock1001.eqiad.wmnet with reason: host reimage * 13:36 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1018.eqiad.wmnet with reason: host reimage * 13:35 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1324707{{!}}Backport all changes from wmf/1.47.0-wmf.15]] * 13:35 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2028.codfw.wmnet * 13:32 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1275938{{!}}Enable campaignEvents on bdwikimedia (T424016)]] (duration: 07m 35s) * 13:32 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1020.eqiad.wmnet with OS bookworm * 13:32 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1019.eqiad.wmnet with OS bookworm * 13:32 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock1001.eqiad.wmnet with reason: host reimage * 13:28 kharlan@deploy1003: kharlan, yahya: Continuing with deployment * 13:27 kharlan@deploy1003: kharlan, yahya: Backport for [[gerrit:1275938{{!}}Enable campaignEvents on bdwikimedia (T424016)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:27 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2028.codfw.wmnet * 13:26 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 13:26 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 13:25 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1275938{{!}}Enable campaignEvents on bdwikimedia (T424016)]] * 13:25 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitize-wiki (exit_code=99) Managing sanitization for wikis testwiki in section s3 * 13:24 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock2001.codfw.wmnet with OS trixie * 13:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2001.codfw.wmnet - cgoubert@cumin2003" * 13:24 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2001.codfw.wmnet - cgoubert@cumin2003" * 13:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2001.codfw.wmnet on all recursors * 13:23 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2001.codfw.wmnet on all recursors * 13:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2001.codfw.wmnet - cgoubert@cumin2003" * 13:23 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2001.codfw.wmnet - cgoubert@cumin2003" * 13:23 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324713{{!}}thwiki: reinstate temporary wiki25 logos (T431094)]] (duration: 07m 13s) * 13:20 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:20 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1018.eqiad.wmnet with OS bookworm * 13:20 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2027.codfw.wmnet * 13:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2027.codfw.wmnet * 13:19 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics-external: apply * 13:18 kharlan@deploy1003: anzx, kharlan: Continuing with deployment * 13:18 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:18 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock1001.eqiad.wmnet with OS trixie * 13:17 kharlan@deploy1003: anzx, kharlan: Backport for [[gerrit:1324713{{!}}thwiki: reinstate temporary wiki25 logos (T431094)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1001.eqiad.wmnet - cgoubert@cumin2003" * 13:17 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 13:17 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1001.eqiad.wmnet - cgoubert@cumin2003" * 13:17 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2001.codfw.wmnet * 13:17 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics-external: apply * 13:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock1001.eqiad.wmnet on all recursors * 13:17 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock1001.eqiad.wmnet on all recursors * 13:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1001.eqiad.wmnet - cgoubert@cumin2003" * 13:17 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1001.eqiad.wmnet - cgoubert@cumin2003" * 13:15 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1324713{{!}}thwiki: reinstate temporary wiki25 logos (T431094)]] * 13:15 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:15 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics-external: apply * 13:15 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:14 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics-external: apply * 13:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2027.codfw.wmnet * 13:12 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 13:12 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock1001.eqiad.wmnet * 13:12 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2027.codfw.wmnet * 13:08 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1158.eqiad.wmnet onto db1273.eqiad.wmnet * 13:07 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1158: Pool db1158.eqiad.wmnet in after cloning * 13:02 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti6002.drmrs.wmnet * 13:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti6002.drmrs.wmnet * 12:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1015.eqiad.wmnet with OS bookworm * 12:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti6002.drmrs.wmnet * 12:44 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1017.eqiad.wmnet with OS bookworm * 12:36 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 12:35 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 12:34 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 12:33 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 12:31 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti6002.drmrs.wmnet * 12:24 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1017.eqiad.wmnet with reason: host reimage * 12:22 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1158: Pool db1158.eqiad.wmnet in after cloning * 12:18 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1017.eqiad.wmnet with reason: host reimage * 12:11 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1015.eqiad.wmnet with reason: host reimage * 12:07 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1015.eqiad.wmnet with reason: host reimage * 12:04 moritzm: failover ganeti master in drmrs02 to ganeti6004 * 12:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1009.eqiad.wmnet with OS bookworm * 12:01 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1017.eqiad.wmnet with OS bookworm * 12:00 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti6004.drmrs.wmnet * 12:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti6004.drmrs.wmnet * 11:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti6004.drmrs.wmnet * 11:53 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1015.eqiad.wmnet with OS bookworm * 11:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1159.eqiad.wmnet onto db1274.eqiad.wmnet * 11:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Pool db1159.eqiad.wmnet in after cloning * 11:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1016.eqiad.wmnet with OS bookworm * 11:46 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti6004.drmrs.wmnet * 11:45 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti6001.drmrs.wmnet * 11:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti6001.drmrs.wmnet * 11:43 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324306{{!}}WikimediaAntiAbuse: Enable personal info for enwiki with no display (T431292)]] (duration: 10m 26s) * 11:42 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1015.eqiad.wmnet with OS bookworm * 11:39 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 11:38 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti6001.drmrs.wmnet * 11:34 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1324306{{!}}WikimediaAntiAbuse: Enable personal info for enwiki with no display (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:33 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2212: Security update * 11:33 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti6001.drmrs.wmnet * 11:32 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1324306{{!}}WikimediaAntiAbuse: Enable personal info for enwiki with no display (T431292)]] * 11:22 moritzm: failover ganeti master in drmrs01 to ganeti6003 * 11:20 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:20 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:18 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 22 hosts with reason: Cloning * 11:17 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti6003.drmrs.wmnet * 11:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti6003.drmrs.wmnet * 11:17 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1009.eqiad.wmnet with reason: host reimage * 11:17 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:16 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:14 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1016.eqiad.wmnet with reason: host reimage * 11:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti6003.drmrs.wmnet * 11:11 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1009.eqiad.wmnet with reason: host reimage * 11:10 jmm@cumin2003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti3005.esams.wmnet * 11:10 jmm@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host ganeti3005.esams.wmnet * 11:09 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 11:08 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1158: Depool db1158.eqiad.wmnet to then clone it to db1273.eqiad.wmnet - marostegui@cumin1003 * 11:07 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1016.eqiad.wmnet with reason: host reimage * 11:07 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1158: Depool db1158.eqiad.wmnet to then clone it to db1273.eqiad.wmnet - marostegui@cumin1003 * 11:07 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1158.eqiad.wmnet onto db1273.eqiad.wmnet * 11:06 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:05 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:05 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1159: Pool db1159.eqiad.wmnet in after cloning * 11:04 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 20 hosts with reason: Cloning * 11:02 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti6003.drmrs.wmnet * 11:00 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis testwiki in section s3 * 10:54 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1009.eqiad.wmnet with OS bookworm * 10:51 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1015.eqiad.wmnet with OS bookworm * 10:50 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1016.eqiad.wmnet with OS bookworm * 10:48 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2212: Security update * 10:45 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitize-wiki (exit_code=99) Managing sanitization for wikis testwiki in section s3 * 10:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1013.eqiad.wmnet with OS bookworm * 10:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1014.eqiad.wmnet with OS bookworm * 10:38 cwilliams@cumin1003: START - Cookbook sre.mysql.clone of db1159.eqiad.wmnet onto db1274.eqiad.wmnet * 10:33 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db1274.eqiad.wmnet * 10:33 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db1274.eqiad.wmnet * 10:31 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 396993 * 10:29 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 396993 * 10:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1159: Clone source for db1274 * 10:24 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1159: Clone source for db1274 * 10:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1013.eqiad.wmnet with reason: host reimage * 10:18 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1013.eqiad.wmnet with reason: host reimage * 10:13 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2212.codfw.wmnet with reason: Maintenance * 10:13 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 10:12 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 10:12 blake@deploy1003: Stopping before sync operations * 10:11 blake@deploy1003: Started scap sync-world: Non-deployment scap run to populate new release values for [[phab:T427668|T427668]] * 10:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2212 [[phab:T434644|T434644]]', diff saved to https://phabricator.wikimedia.org/P96003 and previous config saved to /var/cache/conftool/dbconfig/20260812-101053-cwilliams.json * 10:09 moritzm: powercycle ganeti3005 * 10:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2203 to s1 primary [[phab:T434644|T434644]]', diff saved to https://phabricator.wikimedia.org/P96002 and previous config saved to /var/cache/conftool/dbconfig/20260812-100849-cwilliams.json * 10:08 cezmunsta: Starting s1 codfw failover from db2212 to db2203 - [[phab:T434644|T434644]] * 10:03 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1013.eqiad.wmnet with OS bookworm * 10:02 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1013.eqiad.wmnet with OS bookworm * 10:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2203 with weight 0 [[phab:T434644|T434644]]', diff saved to https://phabricator.wikimedia.org/P96001 and previous config saved to /var/cache/conftool/dbconfig/20260812-100134-cwilliams.json * 10:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 32 hosts with reason: Primary switchover s1 [[phab:T434644|T434644]] * 09:53 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1014.eqiad.wmnet with reason: host reimage * 09:50 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 09:50 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti3005.esams.wmnet * 09:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1014.eqiad.wmnet with reason: host reimage * 09:41 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1278: Pool in x1 * 09:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-launcher1003.eqiad.wmnet with OS bookworm * 09:37 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti3005.esams.wmnet * 09:37 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1013.eqiad.wmnet with OS bookworm * 09:34 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-presto1013.eqiad.wmnet with OS bookworm * 09:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1014.eqiad.wmnet with OS bookworm * 09:29 moritzm: failover ganeti master in esams to ganeti3008 * 09:26 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti3006.esams.wmnet * 09:26 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti3006.esams.wmnet * 09:24 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1012.eqiad.wmnet with OS bookworm * 09:23 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1009.eqiad.wmnet with OS bookworm * 09:23 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 09:18 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti3006.esams.wmnet * 09:16 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti3006.esams.wmnet * 09:14 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: fix regexp escaping bug - oblivian@cumin1003" * 09:14 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: fix regexp escaping bug - oblivian@cumin1003 * 09:13 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: fix regexp escaping bug - oblivian@cumin1003 * 09:13 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: fix regexp escaping bug - oblivian@cumin1003" * 09:03 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-launcher1003.eqiad.wmnet with reason: host reimage * 08:58 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-launcher1003.eqiad.wmnet with reason: host reimage * 08:55 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1278: Pool in x1 * 08:55 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1278 to dbctl [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95996 and previous config saved to /var/cache/conftool/dbconfig/20260812-085521-marostegui.json * 08:51 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1012.eqiad.wmnet with reason: host reimage * 08:45 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis testwiki in section s3 * 08:43 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1009.eqiad.wmnet with OS bookworm * 08:42 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1012.eqiad.wmnet with reason: host reimage * 08:41 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-launcher1003.eqiad.wmnet with OS bookworm * 08:40 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1013.eqiad.wmnet with OS bookworm * 08:38 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms3', diff saved to https://phabricator.wikimedia.org/P95995 and previous config saved to /var/cache/conftool/dbconfig/20260812-083816-marostegui.json * 08:38 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-master1003.eqiad.wmnet with OS bookworm * 08:37 marostegui: Failover ms3 [[phab:T434288|T434288]] * 08:37 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1268 to dbctl [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95994 and previous config saved to /var/cache/conftool/dbconfig/20260812-083722-marostegui.json * 08:35 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-web1001.eqiad.wmnet with OS bookworm * 08:32 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db2252.codfw.wmnet,db[1153,1268].eqiad.wmnet with reason: Switching over ms3 * 08:28 marostegui@cumin1003: dbctl commit (dc=all): 'Depool ms3 [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95993 and previous config saved to /var/cache/conftool/dbconfig/20260812-082852-marostegui.json * 08:25 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1012.eqiad.wmnet with OS bookworm * 08:25 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1011.eqiad.wmnet with OS bookworm * 08:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-master1003.eqiad.wmnet with reason: host reimage * 08:07 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-master1003.eqiad.wmnet with reason: host reimage * 08:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-web1001.eqiad.wmnet with reason: host reimage * 07:58 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-web1001.eqiad.wmnet with reason: host reimage * 07:50 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1003.eqiad.wmnet with OS bookworm * 07:38 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1011.eqiad.wmnet with reason: host reimage * 07:38 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-master1003.eqiad.wmnet with OS bookworm * 07:35 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti3007.esams.wmnet * 07:35 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti3007.esams.wmnet * 07:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1011.eqiad.wmnet with reason: host reimage * 07:27 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti3007.esams.wmnet * 07:25 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti3007.esams.wmnet * 07:25 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti3008.esams.wmnet * 07:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti3008.esams.wmnet * 07:22 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-web1001.eqiad.wmnet with OS bookworm * 07:19 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1003.eqiad.wmnet with OS bookworm * 07:18 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-master1003.eqiad.wmnet * 07:18 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host an-master1003.eqiad.wmnet * 07:17 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1011.eqiad.wmnet with OS bookworm * 07:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 07:16 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti3008.esams.wmnet * 07:15 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1009.eqiad.wmnet with OS bookworm * 07:14 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host an-master1003.eqiad.wmnet * 07:13 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-master1003.eqiad.wmnet * 07:13 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-master1003.eqiad.wmnet * 07:12 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-master1003.eqiad.wmnet * 07:11 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti3008.esams.wmnet * 07:07 arnaudb@dns1006: END - running authdns-update * 07:07 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti5007.eqsin.wmnet * 07:07 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti5007.eqsin.wmnet * 07:05 arnaudb@dns1006: START - running authdns-update * 06:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti5007.eqsin.wmnet * 06:54 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti5007.eqsin.wmnet * 06:38 moritzm: failover ganeti master in eqsin to ganeti5004 * 06:36 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti5006.eqsin.wmnet * 06:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti5006.eqsin.wmnet * 06:28 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti5006.eqsin.wmnet * 06:23 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti5006.eqsin.wmnet * 06:20 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti5005.eqsin.wmnet * 06:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti5005.eqsin.wmnet * 06:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti5005.eqsin.wmnet * 06:06 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti5005.eqsin.wmnet * 06:03 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti5004.eqsin.wmnet * 06:03 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti5004.eqsin.wmnet * 05:55 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti5004.eqsin.wmnet * 05:53 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti5004.eqsin.wmnet * 04:40 ryankemper: [[phab:T434494|T434494]] reimaged `an-tool1008.eqiad.wmnet` to bookworm; yarn.wikimedia.org is back up * 04:16 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-tool1008.eqiad.wmnet with OS bookworm * 03:58 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-tool1008.eqiad.wmnet with reason: host reimage * 03:53 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-tool1008.eqiad.wmnet with reason: host reimage * 03:41 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-tool1008.eqiad.wmnet with OS bookworm * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 45s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 00:25 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324427{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]], [[gerrit:1324429{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0]], [[gerrit:1324428{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]] (duration: 07m 55s) * 00:21 kemayo@deploy1003: kemayo: Continuing with deployment * 00:19 kemayo@deploy1003: kemayo: Backport for [[gerrit:1324427{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]], [[gerrit:1324429{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0]], [[gerrit:1324428{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:17 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1324427{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]], [[gerrit:1324429{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0]], [[gerrit:1324428{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]] == 2026-08-11 == * 21:37 sbassett: Deployed security fix for [[phab:T434521|T434521]] (wmf.15) * 21:29 sbassett: Deployed security fix for [[phab:T434521|T434521]] (wmf.14) * 21:19 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324370{{!}}Phase 4 of legal footer deployment (T432796)]], [[gerrit:1319804{{!}}Disable wgMFCustomSiteModules on English Wikipedia (T375538)]] (duration: 15m 26s) * 21:15 jdlrobson@deploy1003: jdlrobson: Continuing with deployment * 21:06 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1324370{{!}}Phase 4 of legal footer deployment (T432796)]], [[gerrit:1319804{{!}}Disable wgMFCustomSiteModules on English Wikipedia (T375538)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:03 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1324370{{!}}Phase 4 of legal footer deployment (T432796)]], [[gerrit:1319804{{!}}Disable wgMFCustomSiteModules on English Wikipedia (T375538)]] * 20:59 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 20:50 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324384{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]], [[gerrit:1324385{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]] (duration: 06m 58s) * 20:46 kemayo@deploy1003: kemayo: Continuing with deployment * 20:45 kemayo@deploy1003: kemayo: Backport for [[gerrit:1324384{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]], [[gerrit:1324385{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:43 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1324384{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]], [[gerrit:1324385{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]] * 20:42 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 20:42 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324386{{!}}build: Updating js-yaml to 3.15.1, 4.3.1]] (duration: 07m 36s) * 20:38 kemayo@deploy1003: kemayo: Continuing with deployment * 20:37 jhancock@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 20:37 kemayo@deploy1003: kemayo: Backport for [[gerrit:1324386{{!}}build: Updating js-yaml to 3.15.1, 4.3.1]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:35 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1324386{{!}}build: Updating js-yaml to 3.15.1, 4.3.1]] * 20:18 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 20:15 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 20:15 jhancock@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin1003" * 20:14 jhancock@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin1003" * 19:59 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 19:54 jhancock@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 19:10 brennen@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] (duration: 06m 41s) * 19:04 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns6001.wikimedia.org * 19:04 sukhe@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns6001.wikimedia.org * 19:04 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns5003.wikimedia.org * 19:04 sukhe@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns5003.wikimedia.org * 19:03 brennen@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 18:59 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns5003.wikimedia.org with OS trixie * 18:55 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns6001.wikimedia.org with OS trixie * 18:19 brett@cumin2002: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on P<nowiki>{</nowiki>cp7009.magru.wmnet<nowiki>}</nowiki> and A:cp - 9.2.15 Upgrade () * 18:14 brett@cumin2002: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on P<nowiki>{</nowiki>cp7009.magru.wmnet<nowiki>}</nowiki> and A:cp - 9.2.15 Upgrade () * 18:13 brennen@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 18:12 brett@cumin2002: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 9.2.15 Upgrade () * 18:09 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns5003.wikimedia.org with reason: host reimage * 18:06 brett@cumin2002: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 9.2.15 Upgrade () * 18:06 brennen: 1.47.0-wmf.15 train status ([[phab:T430834|T430834]]) - no current blockers, rolling to group0 * 18:05 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns5003.wikimedia.org with reason: host reimage * 18:05 brett: import trafficserver-9.2.15~deb13+wmf1 into trixie-wikimedia ([[phab:T434478|T434478]]) * 18:01 ladsgroup@cumin1003: END (PASS) - Cookbook sre.mysql.sanitarium_restart (exit_code=0) * 17:58 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns6001.wikimedia.org with reason: host reimage * 17:53 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324369{{!}}Enable desktop lazy loading on group0 (T148047)]] (duration: 07m 31s) * 17:52 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns6001.wikimedia.org with reason: host reimage * 17:49 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 17:49 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitarium_restart (exit_code=99) * 17:49 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 17:49 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7001.magru.wmnet * 17:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti7001.magru.wmnet * 17:48 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 17:47 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1324369{{!}}Enable desktop lazy loading on group0 (T148047)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:45 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1324369{{!}}Enable desktop lazy loading on group0 (T148047)]] * 17:39 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti7001.magru.wmnet * 17:36 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns5003.wikimedia.org with OS trixie * 17:34 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns6001.wikimedia.org with OS trixie * 17:31 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324361{{!}}Move FR config from IS.php to a dedicated file]], [[gerrit:1324363{{!}}Remove $wmg = $wg hacks in CentralAuth (T119117)]] (duration: 12m 23s) * 17:26 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 17:23 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1324361{{!}}Move FR config from IS.php to a dedicated file]], [[gerrit:1324363{{!}}Remove $wmg = $wg hacks in CentralAuth (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:19 sukhe: sudo cumin "A:cp-magru" "run-puppet-agent --enable 'merging CR 1324355'" * 17:18 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1324361{{!}}Move FR config from IS.php to a dedicated file]], [[gerrit:1324363{{!}}Remove $wmg = $wg hacks in CentralAuth (T119117)]] * 17:11 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-master1004.eqiad.wmnet with OS bookworm * 17:11 sukhe: sukhe@cp7005:~$ sudo puppet agent -tv * 17:02 sukhe: sudo cumin "A:cp-magru" "disable-puppet 'merging CR 1324355'" * 16:54 sukhe@dns1004: END - running authdns-update * 16:53 sukhe@dns1004: START - running authdns-update * 16:53 sukhe@dns1004: FAIL - running authdns-update * 16:51 sukhe@dns1004: START - running authdns-update * 16:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-master1004.eqiad.wmnet with reason: host reimage * 16:44 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-master1004.eqiad.wmnet with reason: host reimage * 16:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1157.eqiad.wmnet onto db1272.eqiad.wmnet * 16:40 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1157: Pool db1157.eqiad.wmnet in after cloning * 16:38 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 16:31 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324356{{!}}InitialiseSettings: Fix wgOATHAuthEnforce2FAForAll]] (duration: 06m 52s) * 16:30 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 16:28 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 16:27 reedy@deploy1003: reedy: Continuing with deployment * 16:26 reedy@deploy1003: reedy: Backport for [[gerrit:1324356{{!}}InitialiseSettings: Fix wgOATHAuthEnforce2FAForAll]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:24 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324356{{!}}InitialiseSettings: Fix wgOATHAuthEnforce2FAForAll]] * 16:13 sukhe: restart ntpsec.serviceon dns7001 * 16:09 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324335{{!}}InitialiseSettings: Enable 2FA enforcement on various private wikis (T428103)]] (duration: 06m 40s) * 16:08 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 16:06 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2204: Security update * 16:04 reedy@deploy1003: reedy: Continuing with deployment * 16:04 reedy@deploy1003: reedy: Backport for [[gerrit:1324335{{!}}InitialiseSettings: Enable 2FA enforcement on various private wikis (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:02 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7001.magru.wmnet * 16:02 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324335{{!}}InitialiseSettings: Enable 2FA enforcement on various private wikis (T428103)]] * 16:01 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 15:55 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1157: Pool db1157.eqiad.wmnet in after cloning * 15:54 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 15:54 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 15:41 dancy@deploy1003: Finished scap sync-world: Testing (duration: 06m 28s) * 15:40 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-master1004.eqiad.wmnet with OS bookworm * 15:40 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 15:35 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti4008.ulsfo.wmnet * 15:35 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti4008.ulsfo.wmnet * 15:34 dancy@deploy1003: Started scap sync-world: Testing * 15:34 dancy@deploy1003: Installation of scap version "4.279.0" completed for 3 hosts * 15:34 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-master1004.eqiad.wmnet with OS bookworm * 15:34 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 15:33 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-master1004.eqiad.wmnet with OS bookworm * 15:32 dancy@deploy1003: Installing scap version "4.279.0" for 3 host(s) * 15:32 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324339{{!}}Add /w/deployment-info.php entrypoint]] (duration: 07m 25s) * 15:30 moritzm: failover ganeti master in magru to ganeti7004 * 15:29 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti4008.ulsfo.wmnet * 15:28 dancy@deploy1003: dancy: Continuing with deployment * 15:28 tappof: remove 2026-05 swift log archives from centrallog to free some space ([[phab:T434502|T434502]]) * 15:27 dancy@deploy1003: dancy: Backport for [[gerrit:1324339{{!}}Add /w/deployment-info.php entrypoint]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:25 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324339{{!}}Add /w/deployment-info.php entrypoint]] * 15:20 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2204: Security update * 15:18 dancy@deploy1003: Installation of scap version "4.278.0" completed for 3 hosts * 15:18 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7004.magru.wmnet * 15:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti7004.magru.wmnet * 15:16 dancy@deploy1003: Installing scap version "4.278.0" for 3 host(s) * 15:14 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2204.codfw.wmnet with reason: Maintenance * 15:12 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 15:11 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-master1004.eqiad.wmnet with OS bookworm * 15:11 hashar: Restarting CI Jenkins on contint1003 due to Java upgrade. * 15:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2204 [[phab:T434565|T434565]]', diff saved to https://phabricator.wikimedia.org/P95984 and previous config saved to /var/cache/conftool/dbconfig/20260811-151126-cwilliams.json * 15:10 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti4008.ulsfo.wmnet * 15:10 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti7004.magru.wmnet * 15:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2207 to s2 primary [[phab:T434565|T434565]]', diff saved to https://phabricator.wikimedia.org/P95983 and previous config saved to /var/cache/conftool/dbconfig/20260811-150905-cwilliams.json * 15:08 cezmunsta: Starting s2 codfw failover from db2204 to db2207 - [[phab:T434565|T434565]] * 15:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2207 with weight 0 [[phab:T434565|T434565]]', diff saved to https://phabricator.wikimedia.org/P95982 and previous config saved to /var/cache/conftool/dbconfig/20260811-150402-cwilliams.json * 15:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s2 [[phab:T434565|T434565]] * 14:55 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1010.eqiad.wmnet with OS bookworm * 14:49 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns4003.wikimedia.org with OS trixie * 14:47 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1010.eqiad.wmnet with OS bookworm * 14:47 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7001.wikimedia.org with OS trixie * 14:44 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-presto1010.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:41 btullis@cumin1003: START - Cookbook sre.hosts.provision for host an-presto1010.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:40 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-presto1010.eqiad.wmnet with OS bookworm * 14:39 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 14:39 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-presto1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:36 btullis@cumin1003: START - Cookbook sre.hosts.provision for host an-presto1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:32 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1009.eqiad.wmnet with OS bookworm * 14:32 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 14:31 cwilliams@cumin1003: START - Cookbook sre.mysql.clone of db1157.eqiad.wmnet onto db1272.eqiad.wmnet * 14:30 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7004.magru.wmnet * 14:28 moritzm: failover ganeti master in ulsfo to ganeti4005 * 14:23 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7003.magru.wmnet * 14:23 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti7003.magru.wmnet * 14:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-master1004.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:22 btullis@cumin1003: START - Cookbook sre.hosts.provision for host an-master1004.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:21 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-master1004.eqiad.wmnet with OS bookworm * 14:19 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti4007.ulsfo.wmnet * 14:19 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti4007.ulsfo.wmnet * 14:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1010.eqiad.wmnet with OS bookworm * 14:17 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1008.eqiad.wmnet with OS bookworm * 14:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti7003.magru.wmnet * 14:12 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7003.magru.wmnet * 14:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti4007.ulsfo.wmnet * 14:11 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7002.magru.wmnet * 14:11 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti7002.magru.wmnet * 14:09 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324318{{!}}Revert "wmf-config/ProductionServices: set URL for urldownloader to service record" (T429175)]] (duration: 06m 46s) * 14:09 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7001.wikimedia.org with reason: host reimage * 14:06 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti4007.ulsfo.wmnet * 14:05 kharlan@deploy1003: kharlan: Continuing with deployment * 14:05 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti4006.ulsfo.wmnet * 14:04 jayme: updated calico to v3.30.7 on staging-eqiad - [[phab:T427400|T427400]] * 14:04 kharlan@deploy1003: kharlan: Backport for [[gerrit:1324318{{!}}Revert "wmf-config/ProductionServices: set URL for urldownloader to service record" (T429175)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti4006.ulsfo.wmnet * 14:03 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns4003.wikimedia.org with reason: host reimage * 14:03 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7001.wikimedia.org with reason: host reimage * 14:02 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti7002.magru.wmnet * 14:02 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1324318{{!}}Revert "wmf-config/ProductionServices: set URL for urldownloader to service record" (T429175)]] * 14:02 btullis@dns1004: FAIL - running authdns-update * 14:00 btullis@dns1004: START - running authdns-update * 13:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1008.eqiad.wmnet with reason: host reimage * 13:59 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'. * 13:59 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1313985{{!}}wmf-config/ProductionServices: set URL for urldownloader to service record (T429175)]] (duration: 25m 06s) * 13:58 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7002.magru.wmnet * 13:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti4006.ulsfo.wmnet * 13:57 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns4003.wikimedia.org with reason: host reimage * 13:56 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'. * 13:56 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7001.magru.wmnet * 13:56 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1008.eqiad.wmnet with reason: host reimage * 13:55 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'. * 13:55 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'. * 13:55 kharlan@deploy1003: kharlan, sukhe: Continuing with deployment * 13:53 marostegui: Failover ms2 [[phab:T434288|T434288]] * 13:52 marostegui: Failover ms1 [[phab:T434288|T434288]] * 13:52 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7001.magru.wmnet * 13:51 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti4006.ulsfo.wmnet * 13:48 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti4005.ulsfo.wmnet * 13:48 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti4005.ulsfo.wmnet * 13:44 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti4005.ulsfo.wmnet * 13:40 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1008.eqiad.wmnet with OS bookworm * 13:39 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns4003.wikimedia.org with OS trixie * 13:38 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns7001.wikimedia.org with OS trixie * 13:38 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1008.eqiad.wmnet with OS bookworm * 13:37 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti4005.ulsfo.wmnet * 13:36 kharlan@deploy1003: kharlan, sukhe: Backport for [[gerrit:1313985{{!}}wmf-config/ProductionServices: set URL for urldownloader to service record (T429175)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:34 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1313985{{!}}wmf-config/ProductionServices: set URL for urldownloader to service record (T429175)]] * 13:29 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2034.codfw.wmnet * 13:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2034.codfw.wmnet * 13:25 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.clone (exit_code=99) of db1157.eqiad.wmnet onto db1272.eqiad.wmnet * 13:25 cwilliams@cumin1003: START - Cookbook sre.mysql.clone of db1157.eqiad.wmnet onto db1272.eqiad.wmnet * 13:21 urbanecm@deploy1003: mwscript-k8s job started: namespaceDupes.php --wiki=frwiktionary --fix # [[phab:T415716|T415716]] * 13:21 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2034.codfw.wmnet * 13:20 urbanecm@deploy1003: mwscript-k8s job started: namespaceDupes.php --wiki=frwiktionary # [[phab:T415716|T415716]] * 13:19 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1323350{{!}}[tgwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T415307)]], [[gerrit:1322961{{!}}[slwiki] Revert temporary logo for Wikipedia 25 (Vector legacy + Vector 2022) (T414265)]], [[gerrit:1323827{{!}}[itwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T414320)]] (duration: 08m 00s) * 13:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-coord1003.eqiad.wmnet with OS bookworm * 13:15 urbanecm@deploy1003: urbanecm, superpes: Continuing with deployment * 13:13 urbanecm@deploy1003: urbanecm, superpes: Backport for [[gerrit:1323350{{!}}[tgwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T415307)]], [[gerrit:1322961{{!}}[slwiki] Revert temporary logo for Wikipedia 25 (Vector legacy + Vector 2022) (T414265)]], [[gerrit:1323827{{!}}[itwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T414320)]] synced to the testservers (see https://wiki * 13:11 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1323350{{!}}[tgwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T415307)]], [[gerrit:1322961{{!}}[slwiki] Revert temporary logo for Wikipedia 25 (Vector legacy + Vector 2022) (T414265)]], [[gerrit:1323827{{!}}[itwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T414320)]] * 13:11 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1323779{{!}}[ukwiki] Remove reviewer usergroup (T434252)]], [[gerrit:1323312{{!}}[frwiktionary] Add new Schème namespace and its talk (T415716)]] (duration: 06m 49s) * 13:10 marostegui@dns1004: END - running authdns-update * 13:08 marostegui@dns1004: START - running authdns-update * 13:07 marostegui@cumin1003: dbctl commit (dc=all): 'Repool ms2 [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95980 and previous config saved to /var/cache/conftool/dbconfig/20260811-130725-marostegui.json * 13:06 urbanecm@deploy1003: urbanecm, superpes: Continuing with deployment * 13:06 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1266 to dbctl [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95979 and previous config saved to /var/cache/conftool/dbconfig/20260811-130627-marostegui.json * 13:06 urbanecm@deploy1003: urbanecm, superpes: Backport for [[gerrit:1323779{{!}}[ukwiki] Remove reviewer usergroup (T434252)]], [[gerrit:1323312{{!}}[frwiktionary] Add new Schème namespace and its talk (T415716)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:04 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1323779{{!}}[ukwiki] Remove reviewer usergroup (T434252)]], [[gerrit:1323312{{!}}[frwiktionary] Add new Schème namespace and its talk (T415716)]] * 12:59 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db2253.codfw.wmnet,db[1151,1266].eqiad.wmnet with reason: Switching over ms2 * 12:54 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1157: Using as clone source * 12:53 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1157: Using as clone source * 12:51 marostegui@cumin1003: dbctl commit (dc=all): 'Depool ms2 [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95977 and previous config saved to /var/cache/conftool/dbconfig/20260811-125129-marostegui.json * 12:47 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 12:46 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 12:45 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 12:44 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'recommendation-api-ng' for release 'main' . * 12:44 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'recommendation-api-ng' for release 'main' . * 12:43 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'recommendation-api-ng' for release 'main' . * 12:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-coord1003.eqiad.wmnet with reason: host reimage * 12:43 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'ores-legacy' for release 'main' . * 12:42 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'ores-legacy' for release 'main' . * 12:42 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2165: Security update * 12:40 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-coord1003.eqiad.wmnet with reason: host reimage * 12:39 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'ores-legacy' for release 'main' . * 12:38 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' . * 12:38 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' . * 12:37 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' . * 12:34 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2034.codfw.wmnet * 12:30 jmm@cumin2002: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti-test2001.codfw.wmnet * 12:30 jmm@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host ganeti-test2001.codfw.wmnet * 12:25 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1179.eqiad.wmnet onto db1278.eqiad.wmnet * 12:25 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1179: Pool db1179.eqiad.wmnet in after cloning * 12:23 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-coord1003.eqiad.wmnet with OS bookworm * 12:22 moritzm: failover ganeti master in codfw/routed to ganeti2033 * 12:22 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2033.codfw.wmnet * 12:22 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2033.codfw.wmnet * 12:19 jmm@cumin2002: START - Cookbook sre.hosts.reboot-single for host ganeti-test2001.codfw.wmnet * 12:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 12:18 jmm@cumin2002: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti-test2001.codfw.wmnet * 12:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1008.eqiad.wmnet with OS bookworm * 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2033.codfw.wmnet * 12:07 moritzm: failover ganeti master in ganeti/test to ganeti-test2003 * 12:04 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 12:03 jmm@cumin2003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti4005.ulsfo.wmnet * 12:03 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti4005.ulsfo.wmnet * 12:00 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324283{{!}}Use maximum compression level in SqlBlobStore and SqlBagOStuff (T428377)]] (duration: 11m 37s) * 11:57 jmm@cumin2002: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti-test2002.codfw.wmnet * 11:57 jmm@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti-test2002.codfw.wmnet * 11:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2165: Security update * 11:54 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 11:52 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1324283{{!}}Use maximum compression level in SqlBlobStore and SqlBagOStuff (T428377)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:51 jmm@cumin2002: START - Cookbook sre.hosts.reboot-single for host ganeti-test2002.codfw.wmnet * 11:50 jmm@cumin2002: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti-test2002.codfw.wmnet * 11:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2165.codfw.wmnet with reason: Maintenance * 11:48 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@050d19e] (releasing): [[phab:T434186|T434186]] (duration: 01m 14s) * 11:48 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1324283{{!}}Use maximum compression level in SqlBlobStore and SqlBagOStuff (T428377)]] * 11:47 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@050d19e] (releasing): [[phab:T434186|T434186]] * 11:44 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@050d19e] (releasing): test jenkins deploy for [[phab:T434186|T434186]] (duration: 01m 08s) * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2165 [[phab:T434514|T434514]]', diff saved to https://phabricator.wikimedia.org/P95969 and previous config saved to /var/cache/conftool/dbconfig/20260811-114352-cwilliams.json * 11:43 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@050d19e] (releasing): test jenkins deploy for [[phab:T434186|T434186]] * 11:42 jmm@cumin2003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti-test2003.codfw.wmnet * 11:42 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti-test2003.codfw.wmnet * 11:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2161 to s8 primary [[phab:T434514|T434514]]', diff saved to https://phabricator.wikimedia.org/P95968 and previous config saved to /var/cache/conftool/dbconfig/20260811-114136-cwilliams.json * 11:40 cezmunsta: Starting s8 codfw failover from db2165 to db2161 - [[phab:T434514|T434514]] * 11:40 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1179: Pool db1179.eqiad.wmnet in after cloning * 11:36 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti-test2003.codfw.wmnet * 11:36 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti-test2003.codfw.wmnet * 11:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2161 with weight 0 [[phab:T434514|T434514]]', diff saved to https://phabricator.wikimedia.org/P95966 and previous config saved to /var/cache/conftool/dbconfig/20260811-113449-cwilliams.json * 11:34 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 25 hosts with reason: Primary switchover s8 [[phab:T434514|T434514]] * 11:29 moritzm: installing Python 3.11 security updates * 11:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-presto1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 11:26 btullis@cumin1003: START - Cookbook sre.hosts.provision for host an-presto1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 11:23 btullis@dns1004: END - running authdns-update * 11:21 btullis@dns1004: START - running authdns-update * 11:20 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-presto1008.eqiad.wmnet with OS bookworm * 11:20 moritzm: installing curl security updates * 11:11 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1007.eqiad.wmnet with OS bookworm * 10:45 tappof: bump space for prometheus k8s-dse in codfw * 10:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1007.eqiad.wmnet with reason: host reimage * 10:38 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1007.eqiad.wmnet with reason: host reimage * 10:37 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-coord1004.eqiad.wmnet with OS bookworm * 10:35 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1008.eqiad.wmnet with OS bookworm * 10:34 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1179: Depool db1179.eqiad.wmnet to then clone it to db1278.eqiad.wmnet - marostegui@cumin1003 * 10:34 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1008.eqiad.wmnet with OS bookworm * 10:25 fceratto@cumin1003: dbctl commit (dc=all): 'Remove db1177 [[phab:T433474|T433474]]', diff saved to https://phabricator.wikimedia.org/P95964 and previous config saved to /var/cache/conftool/dbconfig/20260811-102527-fceratto.json * 10:22 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1008.eqiad.wmnet with OS bookworm * 10:22 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1007.eqiad.wmnet with OS bookworm * 10:21 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1006.eqiad.wmnet with OS bookworm * 10:20 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 10:18 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1179: Depool db1179.eqiad.wmnet to then clone it to db1278.eqiad.wmnet - marostegui@cumin1003 * 10:18 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1179.eqiad.wmnet onto db1278.eqiad.wmnet * 10:17 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 10:17 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 10:14 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 10:09 blake@deploy1003: Stopping before sync operations * 10:09 blake@deploy1003: Started scap sync-world: Non-deployment run to populate release values for [[phab:T427668|T427668]] * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 10:04 fceratto@cumin1003: Removing db1177 from zarcillo [[phab:T433474|T433474]] * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1177.eqiad.wmnet * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1177.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:03 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1177.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:00 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1006.eqiad.wmnet with reason: host reimage * 09:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-coord1004.eqiad.wmnet with reason: host reimage * 09:57 marostegui: Failover m1 from db1164 to db1213 - [[phab:T434493|T434493]] * 09:57 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1006.eqiad.wmnet with reason: host reimage * 09:55 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:54 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2232].codfw.wmnet,db[1164,1213,1217].eqiad.wmnet with reason: Primary switchover m1 [[phab:T434493|T434493]] * 09:52 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-coord1004.eqiad.wmnet with reason: host reimage * 09:49 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1213.eqiad.wmnet with OS trixie * 09:49 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1177.eqiad.wmnet * 09:41 moritzm: installing Linux 6.12.101 on Trixie hosts * 09:40 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1006.eqiad.wmnet with OS bookworm * 09:35 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-coord1004.eqiad.wmnet with OS bookworm * 09:28 moritzm: installing node-tar security updates * 09:27 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1213.eqiad.wmnet with reason: host reimage * 09:22 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1213.eqiad.wmnet with reason: host reimage * 09:09 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1177: Decommission * 09:08 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db1177: Decommission * 09:08 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 09:08 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.decommission (exit_code=99) * 09:06 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1213.eqiad.wmnet with OS trixie * 09:06 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 09:05 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1213.eqiad.wmnet with reason: Reimage * 08:53 marostegui@dns1004: END - running authdns-update * 08:51 marostegui@dns1004: START - running authdns-update * 08:48 marostegui: Switchover ms1 master in eqiad [[phab:T434288|T434288]] * 08:48 marostegui@cumin1003: dbctl commit (dc=all): 'Repool ms1 [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95962 and previous config saved to /var/cache/conftool/dbconfig/20260811-084804-marostegui.json * 08:40 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1267 to dbctl [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95961 and previous config saved to /var/cache/conftool/dbconfig/20260811-084054-marostegui.json * 08:29 marostegui: Failover m1 from db1213 to db1164 - [[phab:T434043|T434043]] * 08:25 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2232].codfw.wmnet,db[1164,1213,1217].eqiad.wmnet with reason: Primary switchover m1 [[phab:T434043|T434043]] * 08:22 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db2251.codfw.wmnet,db[1152,1267].eqiad.wmnet with reason: Switching over ms1 * 08:22 marostegui@cumin1003: dbctl commit (dc=all): 'Depool ms1 [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95960 and previous config saved to /var/cache/conftool/dbconfig/20260811-082201-marostegui.json * 08:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: Switching over ms1 * 08:20 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.parsercache (exit_code=99) * 08:20 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 08:20 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1152: Switching over ms1 * 08:19 slyngshede@dns1004: END - running authdns-update * 08:18 moritzm: installing openjdk-21 security updates * 08:17 slyngshede@dns1004: START - running authdns-update * 08:16 moritzm: imported jenkins 2.568.2 to thirdparty/jenkins for trixie-wikimedia * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.12 (duration: 02m 26s) * 03:36 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] (duration: 33m 33s) * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 35s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-10 == * 14:54 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1323973{{!}}mmv.bootstrap: Fix getUrlParam to account for TIFF lossy/lossless param (T434333)]] (duration: 11m 24s) * 14:50 krinkle@deploy1003: krinkle: Continuing with deployment * 14:45 krinkle@deploy1003: krinkle: Backport for [[gerrit:1323973{{!}}mmv.bootstrap: Fix getUrlParam to account for TIFF lossy/lossless param (T434333)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:43 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1323973{{!}}mmv.bootstrap: Fix getUrlParam to account for TIFF lossy/lossless param (T434333)]] * 14:07 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1323967{{!}}updateIsActiveFlagForMentees: Commit the final partial batch (T432959)]] (duration: 10m 33s) * 13:56 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1323967{{!}}updateIsActiveFlagForMentees: Commit the final partial batch (T432959)]] * 13:45 dani@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply * 13:45 dani@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply * 13:45 dani@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply * 13:45 dani@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply * 13:45 dani@deploy1003: helmfile [staging] DONE helmfile.d/services/miscweb: apply * 13:44 dani@deploy1003: helmfile [staging] START helmfile.d/services/miscweb: apply * 13:38 wmde-fisch@deploy1003: Finished scap sync-world: Backport for [[gerrit:1323939{{!}}Enable sub-references on more group2 wikis (batch3) (T432731)]] (duration: 33m 21s) * 13:25 wmde-fisch@deploy1003: wmde-fisch: Continuing with deployment * 13:22 wmde-fisch@deploy1003: wmde-fisch: Backport for [[gerrit:1323939{{!}}Enable sub-references on more group2 wikis (batch3) (T432731)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:05 wmde-fisch@deploy1003: Started scap sync-world: Backport for [[gerrit:1323939{{!}}Enable sub-references on more group2 wikis (batch3) (T432731)]] * 07:57 hashar@deploy1003: Finished deploy [integration/docroot@7772132]: update build dependencies (duration: 00m 13s) * 07:57 hashar@deploy1003: Started deploy [integration/docroot@7772132]: update build dependencies * 07:35 _joe_: restarting squid on urldownloader1006 * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 48s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-09 == * 16:01 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:01 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:01 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:00 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 36s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-08 == * 05:31 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9] (wcqs): [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) (duration: 02m 36s) * 05:28 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9] (wcqs): [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) * 04:56 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 04:55 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 04:47 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) (duration: 19m 22s) * 04:28 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) * 04:19 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) (duration: 00m 06s) * 04:18 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) * 04:17 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) (duration: 00m 28s) * 04:16 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) * 03:52 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 03:52 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 34s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-07 == * 23:30 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:29 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 22:45 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 22:43 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 22:41 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 22:41 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 20:54 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 20:32 andrewbogott: restarting puppetserver service on puppetserver* for [[phab:T434339|T434339]] * 19:52 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:45 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 19:32 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:25 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:22 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 19:21 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 18:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:41 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 18:35 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 18:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 18:22 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 18:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 18:16 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 18:12 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:09 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:08 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:07 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:04 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:01 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:00 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:00 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 17:59 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 17:25 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 17:14 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 17:13 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 17:13 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 17:13 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:54 maryum: Deployed security fix for [[phab:T434278|T434278]] * 16:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 16:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 16:27 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --verbose --scoreLessThan=0.7 --exceptDatasetChecksums=[[phab:T434319|T434319]]-enwiki-models.txt # [[phab:T434319|T434319]] * 16:06 cdobbins@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp5022.eqsin.wmnet with OS trixie * 15:13 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 14:19 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 14:17 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 13:50 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1156.eqiad.wmnet onto db1271.eqiad.wmnet * 13:50 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1271: Pool db1271.eqiad.wmnet in after cloning * 13:02 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1271: Pool db1271.eqiad.wmnet in after cloning * 12:19 jayme: updated calico to v3.30.7 on staging-codfw - [[phab:T427400|T427400]] * 12:09 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 12:06 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 12:06 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 12:05 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 12:02 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1156: Pool db1156.eqiad.wmnet in after cloning * 11:38 bjensen: sudo -i reprepro -C main include trixie-wikimedia $<nowiki>{</nowiki>HOME<nowiki>}</nowiki>/httpbb/trixie/httpbb_$<nowiki>{</nowiki>VERSION?<nowiki>}</nowiki>-1+deb13u1_amd64.changes #[[phab:T434052|T434052]] * 11:35 bjensen: sudo -i reprepro -C main include bookworm-wikimedia $<nowiki>{</nowiki>HOME<nowiki>}</nowiki>/httpbb/bookworm/httpbb_$<nowiki>{</nowiki>VERSION?<nowiki>}</nowiki>-1_amd64.changes #[[phab:T434052|T434052]] * 11:30 marostegui@cumin1003: dbctl commit (dc=all): 'Adding db1271 to dbctl', diff saved to https://phabricator.wikimedia.org/P95945 and previous config saved to /var/cache/conftool/dbconfig/20260807-113006-marostegui.json * 11:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1156: Pool db1156.eqiad.wmnet in after cloning * 10:23 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 10:22 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 10:22 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 10:21 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 10:20 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 10:20 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 10:19 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 10:18 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 10:06 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on 21 hosts with reason: cloning * 10:01 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1156: Depool db1156.eqiad.wmnet to then clone it to db1271.eqiad.wmnet - marostegui@cumin1003 * 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1156: Depool db1156.eqiad.wmnet to then clone it to db1271.eqiad.wmnet - marostegui@cumin1003 * 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1156.eqiad.wmnet onto db1271.eqiad.wmnet * 09:15 jynus: started stress testing db1245 dbs [[phab:T431115|T431115]] * 08:19 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:18 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:16 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:14 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:13 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:10 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:06 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:05 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:00 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 10 days, 0:00:00 on ml-serve1015.eqiad.wmnet with reason: Downtime to get full picture of current BIOS settings beyond what Redfish shows * 08:00 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 07:54 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 07:54 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 07:53 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:52 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:51 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:50 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:49 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:48 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:47 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:45 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:45 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:41 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:38 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:37 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 06:35 jayme: updated istio to 1.29.4 on wikikube eqiad - [[phab:T427401|T427401]] * 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1178.eqiad.wmnet * 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1178.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 06:06 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1178.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 05:55 marostegui@cumin1003: START - Cookbook sre.dns.netbox * 05:49 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1178.eqiad.wmnet * 05:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 05:46 marostegui@cumin1003: Removing db1178 from zarcillo [[phab:T433471|T433471]] * 05:45 marostegui@cumin1003: START - Cookbook sre.mysql.decommission * 02:42 denisse: Extended volume on prometheus2008 for the disk space alert as per https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_host_running_out_of_space * 02:37 denisse: Extended volume on prometheus2007 tor the disk space alert as per https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_host_running_out_of_space * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 56s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-06 == * 21:39 maryum: Deploy security patch for [[phab:T433070|T433070]] * 21:29 maryum: Deploy security patch for [[phab:T434189|T434189]] * 20:48 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] (duration: 08m 12s) * 20:44 aude@deploy1003: lmora, aude, anzx: Continuing with deployment * 20:41 aude@deploy1003: lmora, aude, anzx: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be * 20:41 ebernhardson@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:41 ebernhardson@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 20:40 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] * 20:37 ebernhardson@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:37 ebernhardson@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 20:32 ebernhardson@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:32 ebernhardson@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 20:31 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] (duration: 06m 41s) * 20:27 cjming@deploy1003: cjming, ebernhardson, chlod: Continuing with deployment * 20:26 cjming@deploy1003: cjming, ebernhardson, chlod: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:24 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] * 20:18 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] (duration: 09m 22s) * 20:14 cjming@deploy1003: cjming, tsev: Continuing with deployment * 20:11 cjming@deploy1003: cjming, tsev: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:09 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] * 19:41 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply * 19:40 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply * 19:31 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 19:31 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 19:00 cdobbins@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cp5022.eqsin.wmnet with OS trixie * 18:25 ladsgroup@deploy1003: Finished scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) (duration: 06m 08s) * 18:19 ladsgroup@deploy1003: Started scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) * 18:18 ladsgroup@deploy1003: Stopping before sync operations * 18:17 ladsgroup@deploy1003: Started scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) * 17:55 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 16:50 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 16:35 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1001.eqiad.wmnet with OS bookworm * 16:19 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1002.eqiad.wmnet with reason: host reimage * 16:16 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1002.eqiad.wmnet with reason: host reimage * 16:05 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1001.eqiad.wmnet with reason: host reimage * 16:00 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1001.eqiad.wmnet with reason: host reimage * 15:57 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 15:43 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm * 15:29 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1001.eqiad.wmnet with OS bookworm * 15:29 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:58 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm * 14:57 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-drmrs ([[phab:T428495|T428495]]) * 14:55 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-drmrs ([[phab:T428495|T428495]]) * 14:55 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-ui1001.eqiad.wmnet with OS bookworm * 14:54 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-presto1001.eqiad.wmnet with OS bookworm * 14:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-magru ([[phab:T428495|T428495]]) * 14:49 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-magru ([[phab:T428495|T428495]]) * 14:48 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1001.eqiad.wmnet with OS bookworm * 14:46 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-esams ([[phab:T428495|T428495]]) * 14:44 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-esams ([[phab:T428495|T428495]]) * 14:43 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:42 brouberol@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:42 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:42 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 14:40 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 14:40 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:38 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-ui1001.eqiad.wmnet with reason: host reimage * 14:34 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-presto1001.eqiad.wmnet with reason: host reimage * 14:28 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-ui1001.eqiad.wmnet with reason: host reimage * 14:27 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-presto1001.eqiad.wmnet with reason: host reimage * 14:23 sukhe: sudo cumin -b2 'A:cp-text' "run-puppet-agent --enable 'merging CR 1290731'": [[phab:T425441|T425441]] * 14:18 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo for hosts in the wikimedia.org domain - [[phab:T428495|T428495]] * 14:16 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-presto1001.eqiad.wmnet with OS bookworm * 14:14 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-ui1001.eqiad.wmnet with OS bookworm * 14:12 sukhe: sudo cumin 'A:cp-text' "disable-puppet 'merging CR 1290731'": [[phab:T425441|T425441]] * 14:11 swfrench-wmf: restarted navtiming on webperf1003 - [[phab:T428495|T428495]] * 14:04 swfrench-wmf: begin rolling restart of confd in drmrs, eqiad, esams, magru - [[phab:T428495|T428495]] * 14:04 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm * 14:04 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:02 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-client1002.eqiad.wmnet with OS bookworm * 13:58 swfrench-wmf: authdns update to direct eqiad-associated etcd clients back to eqiad - [[phab:T428495|T428495]] * 13:58 swfrench@dns1004: END - running authdns-update * 13:56 swfrench@dns1004: START - running authdns-update * 13:49 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:44 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:31 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 13:29 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 13:26 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 13:23 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 13:22 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 13:19 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 13:18 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 13:18 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 13:17 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 13:16 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 13:13 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 13:11 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 13:09 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 13:06 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 13:06 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-client1002.eqiad.wmnet with OS bookworm * 13:05 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revision-models' for release 'main' . * 13:05 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:05 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revision-models' for release 'main' . * 13:04 brouberol@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-test-client1002.eqiad.wmnet with OS bookworm * 13:04 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 13:03 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 13:02 aikochou@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:00 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'readability' for release 'main' . * 12:59 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'readability' for release 'main' . * 12:58 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 12:57 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'logo-detection' for release 'main' . * 12:57 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'logo-detection' for release 'main' . * 12:57 aikochou@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 12:55 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:54 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:53 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 12:53 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:50 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 12:48 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 12:46 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'article-models' for release 'main' . * 12:45 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'article-models' for release 'main' . * 12:41 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . * 12:39 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . * 12:38 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-client1002.eqiad.wmnet with OS bookworm * 12:12 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply * 12:12 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply * 12:09 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:08 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 11:58 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2187: Security update * 11:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:24 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:16 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 11:15 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 11:10 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2187: Security update * 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2187.codfw.wmnet with reason: Maintenance * 10:56 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 10:56 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 10:56 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 10:56 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 10:54 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 10:53 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 10:09 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2187: Security update * 10:07 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2187: Security update * 09:39 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms2', diff saved to https://phabricator.wikimedia.org/P95929 and previous config saved to /var/cache/conftool/dbconfig/20260806-093908-marostegui.json * 09:36 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1178 from dbctl [[phab:T433471|T433471]]', diff saved to https://phabricator.wikimedia.org/P95928 and previous config saved to /var/cache/conftool/dbconfig/20260806-093632-marostegui.json * 09:33 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 09:31 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 09:30 topranks: bounce cr3-eqsin<->cr2-eqiad bgp session to disable no-prepend command * 09:20 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2253.codfw.wmnet,db1151.eqiad.wmnet with reason: cloning * 09:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1151: Cloning * 09:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:19 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 09:19 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1151: Cloning * 09:10 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 09:09 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup2003.codfw.wmnet * 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup2003.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 09:06 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup2003.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 09:03 klausman@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:02 jynus@cumin1003: START - Cookbook sre.dns.netbox * 09:02 klausman@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 08:57 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup2003.codfw.wmnet * 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup1003.eqiad.wmnet * 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:54 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms3', diff saved to https://phabricator.wikimedia.org/P95925 and previous config saved to /var/cache/conftool/dbconfig/20260806-085422-marostegui.json * 08:53 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:46 jynus@cumin1003: START - Cookbook sre.dns.netbox * 08:39 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup1003.eqiad.wmnet * 08:29 XioNoX: push pfw policy - [[phab:T434115|T434115]] * 08:14 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 08:00 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 08:00 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:58 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revision-models' for release 'main' . * 07:56 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 07:54 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'readability' for release 'main' . * 07:53 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'logo-detection' for release 'main' . * 07:51 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'llm' for release 'main' . * 07:48 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . * 07:37 jayme: updated istio to 1.29.4 on wikikube codfw - [[phab:T427401|T427401]] * 07:08 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2252.codfw.wmnet,db1153.eqiad.wmnet with reason: cloning * 07:07 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1153: Cloning * 07:07 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1153: Cloning * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 40s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-05 == * 23:24 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1009.eqiad.wmnet with OS bookworm * 23:03 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1009.eqiad.wmnet with reason: host reimage * 22:59 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1009.eqiad.wmnet with reason: host reimage * 22:43 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1009.eqiad.wmnet with OS bookworm * 22:38 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1009.eqiad.wmnet * 22:34 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1009.eqiad.wmnet * 22:25 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1008.eqiad.wmnet with OS bookworm * 22:04 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1008.eqiad.wmnet with reason: host reimage * 22:00 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1008.eqiad.wmnet with reason: host reimage * 21:48 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:47 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:46 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:44 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1008.eqiad.wmnet with OS bookworm * 21:43 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:41 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1008.eqiad.wmnet * 21:36 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1008.eqiad.wmnet * 21:14 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:12 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad * 21:12 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad * 21:10 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=eqiad * 21:08 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:07 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:07 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=eqiad * 21:04 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:03 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1006 * 21:02 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1006 * 21:00 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:56 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:56 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:55 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 20:55 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 20:55 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1007.eqiad.wmnet with OS bookworm * 20:51 vriley@cumin1003: START - Cookbook sre.dns.netbox * 20:43 ebernhardson: [[phab:T434008|T434008]]: changing cloudelastic:9643 from auto_expand_replicas to number_of_replicas * 20:34 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1007.eqiad.wmnet with reason: host reimage * 20:27 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1007.eqiad.wmnet with reason: host reimage * 20:24 cjming: end of UTC late backport window * 20:23 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] (duration: 06m 26s) * 20:18 cjming@deploy1003: cjming: Continuing with deployment * 20:18 cjming@deploy1003: cjming: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:16 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] * 20:12 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1007.eqiad.wmnet with OS bookworm * 20:12 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] (duration: 08m 41s) * 20:08 swfrench@cumin2002: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host conf1007.eqiad.wmnet with OS bookworm * 20:08 jforrester@deploy1003: jforrester: Continuing with deployment * 20:07 jforrester@deploy1003: jforrester: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:03 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] * 19:51 inflatador: [bking@puppetserver1001] ~$ sudo puppetserver ca sign --certname an-worker1189.eqiad.wmnet [[phab:T434142|T434142]] * 19:47 bking@cumin2003: DONE (FAIL) - Cookbook sre.puppet.renew-cert (exit_code=99) for an-worker1189.eqiad.wmnet: Renew puppet certificate - bking@cumin2003 * 19:46 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:30 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1007.eqiad.wmnet with OS trixie * 19:30 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 19:29 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 19:20 swfrench-wmf: silenced EtcdRelicationDown 0cb709a9-f244-4f1e-971f-{{Gerrit|440ec65e7fd7}} - [[phab:T428495|T428495]] * 19:13 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1007.eqiad.wmnet with OS bookworm * 19:12 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1007.eqiad.wmnet with reason: host reimage * 19:09 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1007.eqiad.wmnet * 19:07 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1007.eqiad.wmnet with reason: host reimage * 19:03 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1007.eqiad.wmnet * 18:52 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1007.eqiad.wmnet with OS trixie * 18:52 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1007.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:35 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 18:34 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 18:34 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 18:30 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1007.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:28 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:28 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1007] - vriley@cumin1003" * 18:27 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1007] - vriley@cumin1003" * 18:23 vriley@cumin1003: START - Cookbook sre.dns.netbox * 18:22 vriley@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 18:22 robh@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:19 vriley@cumin1003: START - Cookbook sre.dns.netbox * 18:13 robh@cumin2002: START - Cookbook sre.hosts.provision for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:31 jasmine@cumin2002: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-main-eqiad * 17:12 mutante: LDAP - added vwalters to group ciadmin - [[phab:T433615|T433615]] * 16:58 aokoth@deploy1003: Finished deploy [phabricator/deployment@e2ebca5]: Deploy Phab (duration: 00m 34s) * 16:57 aokoth@deploy1003: Started deploy [phabricator/deployment@e2ebca5]: Deploy Phab * 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad * 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=eqiad * 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad * 16:53 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:41 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2187.codfw.wmnet * 16:41 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2187.codfw.wmnet * 16:41 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker2187.codfw.wmnet * 16:41 cgoubert@cumin2003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker2187.codfw.wmnet * 16:40 jasmine@cumin2002: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-main-eqiad * 16:40 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:34 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:25 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-magru and A:liberica ([[phab:T428495|T428495]]) * 16:23 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-magru and A:liberica ([[phab:T428495|T428495]]) * 16:20 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-drmrs and A:liberica ([[phab:T428495|T428495]]) * 16:19 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-drmrs and A:liberica ([[phab:T428495|T428495]]) * 16:18 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-esams and A:liberica ([[phab:T428495|T428495]]) * 16:16 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-esams and A:liberica ([[phab:T428495|T428495]]) * 16:06 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1159.eqiad.wmnet * 16:06 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1159.eqiad.wmnet * 16:06 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1159.eqiad.wmnet * 16:05 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] (duration: 09m 11s) * 15:58 reedy@deploy1003: reedy: Continuing with deployment * 15:58 reedy@deploy1003: reedy: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:56 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] * 15:54 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1159.eqiad.wmnet with OS trixie * 15:38 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:33 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1159.eqiad.wmnet with reason: host reimage * 15:32 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:27 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1159.eqiad.wmnet with reason: host reimage * 15:10 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1159 * 15:10 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1159 * 15:00 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo for hosts in the wikimedia.org domain - [[phab:T428495|T428495]] * 14:55 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS trixie * 14:54 swfrench-wmf: restarted navtiming on webperf1003 - [[phab:T428495|T428495]] * 14:52 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1159 * 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1159.eqiad.wmnet 129.48.64.10.in-addr.arpa 9.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:52 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1159.eqiad.wmnet 129.48.64.10.in-addr.arpa 9.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1159 - jayme@cumin1003" * 14:52 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1159 - jayme@cumin1003" * 14:49 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:48 jayme@cumin1003: START - Cookbook sre.dns.netbox * 14:47 swfrench-wmf: begin rolling restart of confd in drmrs, eqiad, esams, magru - [[phab:T428495|T428495]] * 14:47 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1159 * 14:46 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:46 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:46 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1159.eqiad.wmnet with OS trixie * 14:45 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:44 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:44 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1159.eqiad.wmnet * 14:43 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:43 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1159.eqiad.wmnet * 14:43 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:43 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1159.eqiad.wmnet * 14:43 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:43 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1157.eqiad.wmnet * 14:43 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1157.eqiad.wmnet * 14:43 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1157.eqiad.wmnet * 14:42 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:42 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:42 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:41 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:41 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:41 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:41 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1003.eqiad.wmnet with OS bookworm * 14:39 swfrench-wmf: authdns update to direct eqiad-associated etcd clients to codfw - [[phab:T428495|T428495]] * 14:39 swfrench@dns1004: END - running authdns-update * 14:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 14:37 swfrench@dns1004: START - running authdns-update * 14:37 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:37 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:35 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:35 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 14:28 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:27 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1157.eqiad.wmnet with OS trixie * 14:27 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:27 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:27 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 14:26 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 14:26 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:26 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 14:26 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 14:26 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:26 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host search-loader1002.eqiad.wmnet with OS trixie * 14:19 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:19 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:15 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1003.eqiad.wmnet with reason: host reimage * 14:14 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:14 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:13 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 * 14:13 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host mc2046 * 14:13 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS trixie * 14:11 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1003.eqiad.wmnet with reason: host reimage * 14:10 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:09 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:09 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:09 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:08 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:08 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1157.eqiad.wmnet with reason: host reimage * 14:08 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:04 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 14:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on search-loader1002.eqiad.wmnet with reason: host reimage * 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=eqiad * 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=eqiad * 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=eqiad * 14:00 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:59 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:58 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1157.eqiad.wmnet with reason: host reimage * 13:57 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on search-loader1002.eqiad.wmnet with reason: host reimage * 13:54 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1003.eqiad.wmnet with OS bookworm * 13:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host search-loader1002.eqiad.wmnet with OS trixie * 13:43 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1157 * 13:42 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1157 * 13:40 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1157 * 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1157.eqiad.wmnet 183.32.64.10.in-addr.arpa 3.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1157.eqiad.wmnet 183.32.64.10.in-addr.arpa 3.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1157 - jayme@cumin1003" * 13:39 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1157 - jayme@cumin1003" * 13:39 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] (duration: 07m 00s) * 13:35 jayme@cumin1003: START - Cookbook sre.dns.netbox * 13:35 reedy@deploy1003: reedy: Continuing with deployment * 13:34 reedy@deploy1003: reedy: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:32 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] * 13:23 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1157 * 13:22 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1157.eqiad.wmnet with OS trixie * 13:22 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1157.eqiad.wmnet * 13:22 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1157.eqiad.wmnet * 13:21 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1157.eqiad.wmnet * 13:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1156.eqiad.wmnet * 13:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1156.eqiad.wmnet * 13:15 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1156.eqiad.wmnet * 13:01 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1156.eqiad.wmnet with OS trixie * 12:42 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1156.eqiad.wmnet with reason: host reimage * 12:38 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1156.eqiad.wmnet with reason: host reimage * 12:32 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 12:31 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 12:30 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 12:28 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 12:26 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 12:24 topranks: update bgp confed settings in eqsin * 12:22 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1156 * 12:22 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1156 * 12:22 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 12:19 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1156 * 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1156.eqiad.wmnet 110.32.64.10.in-addr.arpa 0.1.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:19 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1156.eqiad.wmnet 110.32.64.10.in-addr.arpa 0.1.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1156 - jayme@cumin1003" * 12:19 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1156 - jayme@cumin1003" * 12:17 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:14 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS trixie * 12:09 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:06 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:04 jayme@cumin1003: START - Cookbook sre.dns.netbox * 12:04 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 12:02 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'article-models' for release 'main' . * 12:01 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1156 * 12:01 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1156.eqiad.wmnet with OS trixie * 11:59 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1156.eqiad.wmnet * 11:59 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1156.eqiad.wmnet * 11:59 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1156.eqiad.wmnet * 11:57 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 11:53 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 11:53 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:52 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:52 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:50 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:50 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:50 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:49 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:48 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:47 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:47 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:45 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:45 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:44 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:44 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:44 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:43 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:42 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:38 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:35 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 * 11:35 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host mc2046 * 11:34 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS trixie * 11:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:27 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:21 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:21 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:18 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:18 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:18 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:18 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:13 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:13 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:09 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:08 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:07 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:06 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:06 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:05 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:05 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:04 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 11:04 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:24 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:24 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:17 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:16 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1155.eqiad.wmnet * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1155.eqiad.wmnet * 10:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1155.eqiad.wmnet * 10:14 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:14 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:11 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:11 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:10 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:09 aikochou@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop: sync * 10:09 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:09 aikochou@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop: sync * 10:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:07 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:05 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:05 aikochou@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop: sync * 10:05 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:05 aikochou@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop: sync * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:04 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:04 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:04 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1155.eqiad.wmnet with OS trixie * 09:52 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms1', diff saved to https://phabricator.wikimedia.org/P95918 and previous config saved to /var/cache/conftool/dbconfig/20260805-095212-marostegui.json * 09:44 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1152: after cloning * 09:44 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.parsercache (exit_code=99) * 09:44 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 09:44 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1152: after cloning * 09:43 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1155.eqiad.wmnet with reason: host reimage * 09:40 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1155.eqiad.wmnet with reason: host reimage * 09:32 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 09:32 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:31 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 09:31 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:27 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1155 * 09:27 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1155 * 09:25 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 09:24 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 09:24 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 09:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:23 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 09:23 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 09:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:22 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 09:22 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 09:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:20 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 09:20 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:17 XioNoX: push pfw policies - [[phab:T434038|T434038]] * 09:14 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1155 * 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1155.eqiad.wmnet 109.32.64.10.in-addr.arpa 9.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1155.eqiad.wmnet 109.32.64.10.in-addr.arpa 9.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1155 - jayme@cumin1003" * 09:14 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1155 - jayme@cumin1003" * 09:10 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 09:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:09 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2251.codfw.wmnet,db1152.eqiad.wmnet with reason: cloning * 09:09 jayme@cumin1003: START - Cookbook sre.dns.netbox * 09:08 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 09:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: Cloning * 09:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:05 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 09:05 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1152: Cloning * 08:38 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1155 * 08:37 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1155.eqiad.wmnet with OS trixie * 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1171.eqiad.wmnet * 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1171.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:29 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95913 and previous config saved to /var/cache/conftool/dbconfig/20260805-082908-ladsgroup.json * 08:27 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1171.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:22 jynus@cumin1003: START - Cookbook sre.dns.netbox * 08:18 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249', diff saved to https://phabricator.wikimedia.org/P95912 and previous config saved to /var/cache/conftool/dbconfig/20260805-081823-ladsgroup.json * 08:17 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1171.eqiad.wmnet * 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1150.eqiad.wmnet * 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1150.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:15 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1150.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:15 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 08:14 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1155.eqiad.wmnet * 08:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1155.eqiad.wmnet * 08:14 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1155.eqiad.wmnet * 08:11 jynus@cumin1003: START - Cookbook sre.dns.netbox * 08:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249', diff saved to https://phabricator.wikimedia.org/P95911 and previous config saved to /var/cache/conftool/dbconfig/20260805-080737-ladsgroup.json * 08:05 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1150.eqiad.wmnet * 08:02 marostegui: Depool clouddb1020 (s5,s8) [[phab:T434048|T434048]] * 08:02 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1020.eqiad.wmnet,service=s8 * 08:02 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1020.eqiad.wmnet,service=s5 * 08:02 marostegui: Depool clouddb1018 (s2,s7) [[phab:T434048|T434048]] * 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1018.eqiad.wmnet,service=s7 * 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1018.eqiad.wmnet,service=s2 * 08:01 marostegui: Depool clouddb1017 (s1) [[phab:T434048|T434048]] * 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1017.eqiad.wmnet,service=s1 * 07:59 marostegui: Depool clouddb1016 (s5,s8) [[phab:T434048|T434048]] * 07:59 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s8 * 07:59 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s5 * 07:57 marostegui: Depool clouddb1015 (s4,s6) [[phab:T434048|T434048]] * 07:57 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s6 * 07:57 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s4 * 07:56 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95910 and previous config saved to /var/cache/conftool/dbconfig/20260805-075650-ladsgroup.json * 07:54 marostegui: Depool clouddb1014 (s2,s7) [[phab:T434048|T434048]] * 07:54 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1014.eqiad.wmnet,service=s7 * 07:54 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1014.eqiad.wmnet,service=s2 * 07:53 marostegui: Depool clouddb1013:s1 [[phab:T434048|T434048]] * 07:53 marostegui: Depool clouddb1013:s1 [[phab:T409557|T409557]] * 07:53 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1013.eqiad.wmnet,service=s1 * 07:25 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95909 and previous config saved to /var/cache/conftool/dbconfig/20260805-072529-ladsgroup.json * 07:24 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2249.codfw.wmnet with reason: Maintenance * 07:24 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95908 and previous config saved to /var/cache/conftool/dbconfig/20260805-072426-ladsgroup.json * 07:21 slyngshede@dns1004: END - running authdns-update * 07:19 slyngshede@dns1004: START - running authdns-update * 07:13 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231', diff saved to https://phabricator.wikimedia.org/P95906 and previous config saved to /var/cache/conftool/dbconfig/20260805-071340-ladsgroup.json * 07:02 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231', diff saved to https://phabricator.wikimedia.org/P95905 and previous config saved to /var/cache/conftool/dbconfig/20260805-070253-ladsgroup.json * 06:52 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95904 and previous config saved to /var/cache/conftool/dbconfig/20260805-065206-ladsgroup.json * 06:45 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 06:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95903 and previous config saved to /var/cache/conftool/dbconfig/20260805-062240-ladsgroup.json * 06:21 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2231.codfw.wmnet with reason: Maintenance * 06:21 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95902 and previous config saved to /var/cache/conftool/dbconfig/20260805-062137-ladsgroup.json * 06:10 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215', diff saved to https://phabricator.wikimedia.org/P95901 and previous config saved to /var/cache/conftool/dbconfig/20260805-061051-ladsgroup.json * 06:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215', diff saved to https://phabricator.wikimedia.org/P95900 and previous config saved to /var/cache/conftool/dbconfig/20260805-060004-ladsgroup.json * 05:49 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95899 and previous config saved to /var/cache/conftool/dbconfig/20260805-054918-ladsgroup.json * 05:19 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95898 and previous config saved to /var/cache/conftool/dbconfig/20260805-051939-ladsgroup.json * 05:18 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2215.codfw.wmnet with reason: Maintenance * 04:30 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2201.codfw.wmnet with reason: Maintenance * 03:40 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2197.codfw.wmnet with reason: Maintenance * 03:40 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95897 and previous config saved to /var/cache/conftool/dbconfig/20260805-034036-ladsgroup.json * 03:29 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196', diff saved to https://phabricator.wikimedia.org/P95896 and previous config saved to /var/cache/conftool/dbconfig/20260805-032948-ladsgroup.json * 03:19 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196', diff saved to https://phabricator.wikimedia.org/P95895 and previous config saved to /var/cache/conftool/dbconfig/20260805-031902-ladsgroup.json * 03:08 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95894 and previous config saved to /var/cache/conftool/dbconfig/20260805-030815-ladsgroup.json * 02:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95893 and previous config saved to /var/cache/conftool/dbconfig/20260805-023413-ladsgroup.json * 02:33 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2196.codfw.wmnet with reason: Maintenance * 02:33 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95892 and previous config saved to /var/cache/conftool/dbconfig/20260805-023310-ladsgroup.json * 02:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186', diff saved to https://phabricator.wikimedia.org/P95891 and previous config saved to /var/cache/conftool/dbconfig/20260805-022223-ladsgroup.json * 02:11 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186', diff saved to https://phabricator.wikimedia.org/P95890 and previous config saved to /var/cache/conftool/dbconfig/20260805-021137-ladsgroup.json * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 02:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95889 and previous config saved to /var/cache/conftool/dbconfig/20260805-020051-ladsgroup.json * 01:30 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95888 and previous config saved to /var/cache/conftool/dbconfig/20260805-013029-ladsgroup.json * 01:29 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2186.codfw.wmnet with reason: Maintenance * 00:34 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on dbstore1009.eqiad.wmnet with reason: Maintenance * 00:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95887 and previous config saved to /var/cache/conftool/dbconfig/20260805-003408-ladsgroup.json * 00:23 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264', diff saved to https://phabricator.wikimedia.org/P95886 and previous config saved to /var/cache/conftool/dbconfig/20260805-002322-ladsgroup.json * 00:12 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264', diff saved to https://phabricator.wikimedia.org/P95885 and previous config saved to /var/cache/conftool/dbconfig/20260805-001235-ladsgroup.json * 00:01 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95884 and previous config saved to /var/cache/conftool/dbconfig/20260805-000148-ladsgroup.json == 2026-08-04 == * 23:45 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95883 and previous config saved to /var/cache/conftool/dbconfig/20260804-234508-ladsgroup.json * 23:44 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1264.eqiad.wmnet with reason: Maintenance * 23:44 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95882 and previous config saved to /var/cache/conftool/dbconfig/20260804-234405-ladsgroup.json * 23:33 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237', diff saved to https://phabricator.wikimedia.org/P95881 and previous config saved to /var/cache/conftool/dbconfig/20260804-233317-ladsgroup.json * 23:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237', diff saved to https://phabricator.wikimedia.org/P95880 and previous config saved to /var/cache/conftool/dbconfig/20260804-232230-ladsgroup.json * 23:11 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95879 and previous config saved to /var/cache/conftool/dbconfig/20260804-231144-ladsgroup.json * 22:23 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95878 and previous config saved to /var/cache/conftool/dbconfig/20260804-222345-ladsgroup.json * 22:23 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1237.eqiad.wmnet with reason: Maintenance * 21:13 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1225.eqiad.wmnet with reason: Maintenance * 20:40 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] (duration: 24m 40s) * 20:33 samtar@deploy1003: samtar, kineticpelagic: Continuing with deployment * 20:28 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS bookworm * 20:21 samtar@deploy1003: samtar, kineticpelagic: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:15 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] * 20:13 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 20:09 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 20:00 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1216.eqiad.wmnet with reason: Maintenance * 20:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95877 and previous config saved to /var/cache/conftool/dbconfig/20260804-195957-ladsgroup.json * 19:51 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 19:50 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) mc2046.codfw.wmnet 120.16.192.10.in-addr.arpa 0.2.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:50 jhancock@cumin2002: START - Cookbook sre.dns.wipe-cache mc2046.codfw.wmnet 120.16.192.10.in-addr.arpa 0.2.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host mc2046 - jhancock@cumin2002" * 19:50 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host mc2046 - jhancock@cumin2002" * 19:49 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203', diff saved to https://phabricator.wikimedia.org/P95876 and previous config saved to /var/cache/conftool/dbconfig/20260804-194911-ladsgroup.json * 19:46 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 19:45 jhancock@cumin2002: START - Cookbook sre.hosts.move-vlan for host mc2046 * 19:45 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS bookworm * 19:38 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203', diff saved to https://phabricator.wikimedia.org/P95875 and previous config saved to /var/cache/conftool/dbconfig/20260804-193825-ladsgroup.json * 19:27 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95874 and previous config saved to /var/cache/conftool/dbconfig/20260804-192738-ladsgroup.json * 19:02 mutante: gerrit ssh -p 29418 gerrit.wikimedia.org gerrit index changes {{Gerrit|1320979}} * 18:20 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 18:18 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 18:14 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 18:14 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 18:13 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 18:10 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 18:08 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 18:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95872 and previous config saved to /var/cache/conftool/dbconfig/20260804-180721-ladsgroup.json * 18:07 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 18:06 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1203.eqiad.wmnet with reason: Maintenance * 18:06 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95871 and previous config saved to /var/cache/conftool/dbconfig/20260804-180618-ladsgroup.json * 17:55 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179', diff saved to https://phabricator.wikimedia.org/P95870 and previous config saved to /var/cache/conftool/dbconfig/20260804-175531-ladsgroup.json * 17:55 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1154.eqiad.wmnet * 17:55 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1154.eqiad.wmnet * 17:55 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1154.eqiad.wmnet * 17:50 swfrench@deploy1003: Finished scap sync-world: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] (duration: 04m 05s) * 17:48 swfrench@deploy1003: swfrench: Continuing with deployment * 17:46 swfrench@deploy1003: swfrench: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:45 swfrench@deploy1003: Started scap sync-world: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] * 17:44 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179', diff saved to https://phabricator.wikimedia.org/P95869 and previous config saved to /var/cache/conftool/dbconfig/20260804-174445-ladsgroup.json * 17:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95868 and previous config saved to /var/cache/conftool/dbconfig/20260804-173359-ladsgroup.json * 17:33 swfrench@deploy1003: Finished scap sync-world: Pick up new PHP production image (duration: 28m 32s) * 17:28 aokoth@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on phab1005.eqiad.wmnet with reason: Puppet Failure * 17:05 swfrench@deploy1003: Started scap sync-world: Pick up new PHP production image * 17:00 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 17:00 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 16:54 cgoubert@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on wikikube-worker2187.codfw.wmnet with reason: Hardware issue * 16:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2187.codfw.wmnet * 16:52 mutante: gerrit2003:/var/log/apache2# ln -s /srv/gerrit/site_path/review_site/logs/ gerrit * 16:52 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2187.codfw.wmnet * 16:48 mutante: gerrit2003 - moving old apache logfiles older than 60 days from /var/log/apache2 to /srv/gerrit/site_path/review_site/logs/old/ * 16:33 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 16:32 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 16:29 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 16:29 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 16:28 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 16:28 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 16:27 dzahn@cumin1003: END (PASS) - Cookbook sre.gerrit.restart-gerrit (exit_code=0) Restarting Gerrit on gerrit2003 * 16:27 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 16:27 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95867 and previous config saved to /var/cache/conftool/dbconfig/20260804-162736-ladsgroup.json * 16:27 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 16:26 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1179.eqiad.wmnet with reason: Maintenance * 16:26 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:25 mutante: restarting gerrit - dropped outdated RSA host key * 16:25 dzahn@cumin1003: START - Cookbook sre.gerrit.restart-gerrit Restarting Gerrit on gerrit2003 * 16:24 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95866 and previous config saved to /var/cache/conftool/dbconfig/20260804-162424-ladsgroup.json * 16:24 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 16:23 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 16:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95865 and previous config saved to /var/cache/conftool/dbconfig/20260804-162236-ladsgroup.json * 16:21 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1179.eqiad.wmnet with reason: Maintenance * 16:17 swfrench-wmf: reprepro include php8.3_8.3.33-1+wmf11u1 into component/php83 for bullseye-wikimedia * 16:17 swfrench-wmf: reprepro include php8.3_8.3.33-1+wmf12u1 into component/php83 for bookworm-wikimedia * 16:11 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply * 16:10 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply * 16:10 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mobileapps: apply * 16:09 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mobileapps: apply * 16:09 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply * 16:08 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply * 16:08 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:08 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:07 aokoth@cumin1003: END (PASS) - Cookbook sre.vrts.upgrade (exit_code=0) on VRTS host vrts1003.eqiad.wmnet * 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:05 aokoth@cumin1003: START - Cookbook sre.vrts.upgrade on VRTS host vrts1003.eqiad.wmnet * 16:04 mutante: gerrit2002/gerrit1003/gerrit2003 - rm /etc/gerrit/ssh_host_rsa_key * 15:59 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:59 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:59 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:59 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:56 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 15:55 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:55 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:55 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:49 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 15:49 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:44 Raine: add php8.5 packages to component/php85 - [[phab:T432983|T432983]] * 15:39 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:33 aaron@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 15:33 aaron@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 15:29 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:19 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:19 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:16 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:16 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1154.eqiad.wmnet with OS trixie * 15:16 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:15 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:15 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:06 brennen@deploy1003: Finished deploy [phabricator/deployment@56f4ffd]: deploy phab1004 for [[phab:T433981|T433981]] (duration: 00m 43s) * 15:05 brennen@deploy1003: Started deploy [phabricator/deployment@56f4ffd]: deploy phab1004 for [[phab:T433981|T433981]] * 15:05 aaron@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 15:04 aaron@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 15:02 brennen@deploy1003: Finished deploy [phabricator/deployment@56f4ffd]: deploy phab2003 for [[phab:T433981|T433981]] (duration: 00m 51s) * 15:01 brennen@deploy1003: Started deploy [phabricator/deployment@56f4ffd]: deploy phab2003 for [[phab:T433981|T433981]] * 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1004.eqiad.wmnet with reason: deployment * 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1005.eqiad.wmnet with reason: deployment * 14:58 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab2003.codfw.wmnet with reason: deployment * 14:55 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1154.eqiad.wmnet with reason: host reimage * 14:51 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1154.eqiad.wmnet with reason: host reimage * 14:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 14:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 14:38 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync * 14:38 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync * 14:38 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync * 14:37 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync * 14:37 ottomata: roll restart eventgate-main to pick up stream config change - [[phab:T433507|T433507]] * 14:37 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-main: sync * 14:36 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-main: sync * 14:36 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1154 * 14:36 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1154 * 14:34 otto@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] (duration: 08m 39s) * 14:34 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1154 * 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1154.eqiad.wmnet 108.32.64.10.in-addr.arpa 8.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:34 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1154.eqiad.wmnet 108.32.64.10.in-addr.arpa 8.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1154 - jayme@cumin1003" * 14:34 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1154 - jayme@cumin1003" * 14:30 otto@deploy1003: otto: Continuing with deployment * 14:30 jayme@cumin1003: START - Cookbook sre.dns.netbox * 14:28 otto@deploy1003: otto: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:26 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1154 * 14:26 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1154.eqiad.wmnet with OS trixie * 14:26 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1154.eqiad.wmnet * 14:26 otto@deploy1003: Started scap sync-world: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] * 14:26 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1154.eqiad.wmnet * 14:26 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1154.eqiad.wmnet * 14:17 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 14:16 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 14:15 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 14:14 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 14:13 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 14:13 swfrench@dns1004: END - running authdns-update * 14:13 Msz2001: Finished deployments for UTC afternoon backport window * 14:13 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 14:13 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] (duration: 07m 58s) * 14:11 swfrench@dns1004: START - running authdns-update * 14:08 mszwarc@deploy1003: javiermonton, mszwarc, mpostoronca: Continuing with deployment * 14:07 mszwarc@deploy1003: javiermonton, mszwarc, mpostoronca: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] synced to the testser * 14:05 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] * 14:03 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 13:49 swfrench@cumin2002: conftool action : set/pooled=yes; selector: name=wikikube-worker2330.codfw.wmnet * 13:49 swfrench@cumin2002: conftool action : set/pooled=no; selector: name=wikikube-worker2330.codfw.wmnet * 13:48 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] (duration: 09m 19s) * 13:45 swfrench@dns1004: END - running authdns-update * 13:44 mszwarc@deploy1003: mszwarc, jforrester: Continuing with deployment * 13:43 swfrench@dns1004: START - running authdns-update * 13:41 mszwarc@deploy1003: mszwarc, jforrester: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:38 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] * 13:33 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 13:33 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1154.eqiad.wmnet * 13:32 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 13:32 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 13:31 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 13:31 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:31 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:29 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1154.eqiad.wmnet * 13:28 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1154.eqiad.wmnet * 13:28 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1154.eqiad.wmnet * 13:28 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1141.eqiad.wmnet * 13:28 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1141.eqiad.wmnet * 13:28 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1141.eqiad.wmnet * 13:22 otto@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply * 13:22 otto@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply * 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1096.eqiad.wmnet with OS trixie * 13:05 swfrench@dns1004: END - running authdns-update * 13:03 swfrench@dns1004: START - running authdns-update * 12:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 12:43 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 1:00:00 on db1171.eqiad.wmnet with reason: decom * 12:42 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 1:00:00 on db1150.eqiad.wmnet with reason: decom * 12:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 12:38 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1164,1217].eqiad.wmnet with reason: cloning * 12:33 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2096.codfw.wmnet with OS trixie * 12:22 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1096.eqiad.wmnet with OS trixie * 12:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2096.codfw.wmnet with reason: host reimage * 12:14 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1141.eqiad.wmnet with OS trixie * 12:10 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2096.codfw.wmnet with reason: host reimage * 12:10 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1289.eqiad.wmnet * 12:05 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1289.eqiad.wmnet * 12:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1288.eqiad.wmnet * 11:59 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1288.eqiad.wmnet * 11:59 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1287.eqiad.wmnet * 11:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1097.eqiad.wmnet with OS trixie * 11:54 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1287.eqiad.wmnet * 11:54 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1286.eqiad.wmnet * 11:53 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1141.eqiad.wmnet with reason: host reimage * 11:51 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2096.codfw.wmnet with OS trixie * 11:49 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1141.eqiad.wmnet with reason: host reimage * 11:48 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1286.eqiad.wmnet * 11:48 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1284.eqiad.wmnet * 11:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2095.codfw.wmnet with OS trixie * 11:43 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1284.eqiad.wmnet * 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1283.eqiad.wmnet * 11:42 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on ml-serve1015.eqiad.wmnet with reason: Downtime to get full picture of current BIOS settings beyond what Redfish shows * 11:39 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad * 11:39 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:37 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1283.eqiad.wmnet * 11:37 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1282.eqiad.wmnet * 11:37 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad * 11:37 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:33 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1141 * 11:33 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1141 * 11:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 11:32 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1141 * 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1141.eqiad.wmnet 156.48.64.10.in-addr.arpa 6.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:32 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1141.eqiad.wmnet 156.48.64.10.in-addr.arpa 6.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1141 - jayme@cumin1003" * 11:32 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1141 - jayme@cumin1003" * 11:32 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1282.eqiad.wmnet * 11:32 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1281.eqiad.wmnet * 11:32 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad * 11:32 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:29 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 11:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1097.eqiad.wmnet with reason: host reimage * 11:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2095.codfw.wmnet with OS trixie * 11:26 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1281.eqiad.wmnet * 11:26 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1280.eqiad.wmnet * 11:25 jayme@cumin1003: START - Cookbook sre.dns.netbox * 11:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1097.eqiad.wmnet with reason: host reimage * 11:22 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1141 * 11:21 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1141.eqiad.wmnet with OS trixie * 11:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1280.eqiad.wmnet * 11:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1279.eqiad.wmnet * 11:20 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin with reason: upgrade new Nokia swtiches in eqsin to SR Linux v26 * 11:17 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1141.eqiad.wmnet * 11:16 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1141.eqiad.wmnet * 11:16 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1141.eqiad.wmnet * 11:16 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be2095.codfw.wmnet with OS trixie * 11:15 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1279.eqiad.wmnet * 11:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1278.eqiad.wmnet * 11:14 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1139.eqiad.wmnet * 11:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1139.eqiad.wmnet * 11:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1070.eqiad.wmnet with OS trixie * 11:13 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1140.eqiad.wmnet * 11:13 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1140.eqiad.wmnet * 11:13 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1140.eqiad.wmnet * 11:09 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1278.eqiad.wmnet * 11:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1071.eqiad.wmnet with OS trixie * 11:05 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1097.eqiad.wmnet with OS trixie * 11:04 mvernon@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be1097.eqiad.wmnet with OS trixie * 11:02 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1140.eqiad.wmnet with OS trixie * 11:02 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1097.eqiad.wmnet with OS trixie * 11:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1096.eqiad.wmnet with OS trixie * 11:00 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1139.eqiad.wmnet * 11:00 jayme@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1139.eqiad.wmnet with OS trixie * 10:56 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1069.eqiad.wmnet with OS trixie * 10:56 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 10:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1070.eqiad.wmnet with reason: host reimage * 10:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 10:45 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1071.eqiad.wmnet with reason: host reimage * 10:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 10:41 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1096.eqiad.wmnet with OS trixie * 10:41 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1140.eqiad.wmnet with reason: host reimage * 10:39 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1071.eqiad.wmnet with reason: host reimage * 10:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1070.eqiad.wmnet with reason: host reimage * 10:38 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1139.eqiad.wmnet with reason: host reimage * 10:37 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1140.eqiad.wmnet with reason: host reimage * 10:35 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1069.eqiad.wmnet with reason: host reimage * 10:33 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1095.eqiad.wmnet with OS trixie * 10:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 10:33 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1139.eqiad.wmnet with reason: host reimage * 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1069.eqiad.wmnet with reason: host reimage * 10:23 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1140 * 10:23 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1140 * 10:23 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 10:22 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1140 * 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1140.eqiad.wmnet 155.48.64.10.in-addr.arpa 5.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:21 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1071.eqiad.wmnet with OS trixie * 10:21 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1140.eqiad.wmnet 155.48.64.10.in-addr.arpa 5.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1140 - jayme@cumin1003" * 10:21 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1140 - jayme@cumin1003" * 10:21 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1071 * 10:21 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1070.eqiad.wmnet with OS trixie * 10:21 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1070 * 10:20 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 10:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1095.eqiad.wmnet with OS trixie * 10:17 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1139 * 10:17 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1139 * 10:17 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be1095.eqiad.wmnet with OS trixie * 10:15 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1139 * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1139.eqiad.wmnet 194.32.64.10.in-addr.arpa 4.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:15 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1139.eqiad.wmnet 194.32.64.10.in-addr.arpa 4.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1139 - jayme@cumin1003" * 10:15 jayme@cumin1003: START - Cookbook sre.dns.netbox * 10:15 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1139 - jayme@cumin1003" * 10:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2095.codfw.wmnet with OS trixie * 10:13 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1069.eqiad.wmnet with OS trixie * 10:12 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1069 * 10:11 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1140 * 10:11 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1140.eqiad.wmnet with OS trixie * 10:11 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1140.eqiad.wmnet * 10:10 jayme@cumin1003: START - Cookbook sre.dns.netbox * 10:10 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1139 * 10:10 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1140.eqiad.wmnet * 10:10 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1140.eqiad.wmnet * 10:10 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1139.eqiad.wmnet with OS trixie * 10:09 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1139.eqiad.wmnet * 10:08 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1139.eqiad.wmnet * 10:08 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1139.eqiad.wmnet * 10:01 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2094.codfw.wmnet with OS trixie * 09:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 09:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 09:53 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 09:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 09:44 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:44 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2094.codfw.wmnet with reason: host reimage * 09:34 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2094.codfw.wmnet with reason: host reimage * 09:34 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1071 * 09:33 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1070 * 09:33 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1095.eqiad.wmnet with OS trixie * 09:27 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1069 * 09:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:23 brouberol@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM archiva1002.wikimedia.org * 09:20 brouberol@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM archiva1002.wikimedia.org * 09:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1277.eqiad.wmnet * 09:13 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2094.codfw.wmnet with OS trixie * 09:13 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 09:12 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 09:12 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:12 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1277.eqiad.wmnet * 09:12 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1276.eqiad.wmnet * 09:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1094.eqiad.wmnet with OS trixie * 09:06 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1276.eqiad.wmnet * 09:06 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1275.eqiad.wmnet * 09:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2093.codfw.wmnet with OS trixie * 09:01 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1275.eqiad.wmnet * 09:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1274.eqiad.wmnet * 08:56 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1274.eqiad.wmnet * 08:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1273.eqiad.wmnet * 08:50 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1273.eqiad.wmnet * 08:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1272.eqiad.wmnet * 08:49 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:49 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1094.eqiad.wmnet with reason: host reimage * 08:45 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1272.eqiad.wmnet * 08:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1094.eqiad.wmnet with reason: host reimage * 08:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2093.codfw.wmnet with reason: host reimage * 08:38 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:38 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2093.codfw.wmnet with reason: host reimage * 08:35 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 08:34 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 08:29 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:28 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:26 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1271.eqiad.wmnet * 08:23 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1094.eqiad.wmnet with OS trixie * 08:21 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 08:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1271.eqiad.wmnet * 08:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1270.eqiad.wmnet * 08:15 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1270.eqiad.wmnet * 08:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1269.eqiad.wmnet * 08:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2093.codfw.wmnet with OS trixie * 08:09 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1269.eqiad.wmnet * 08:09 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1268.eqiad.wmnet * 08:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2092.codfw.wmnet with OS trixie * 08:04 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1268.eqiad.wmnet * 08:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1267.eqiad.wmnet * 07:59 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1267.eqiad.wmnet * 07:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1093.eqiad.wmnet with OS trixie * 07:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1266.eqiad.wmnet * 07:56 jynus: running extra backups to test db1285 [[phab:T433826|T433826]] * 07:51 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1266.eqiad.wmnet * 07:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2092.codfw.wmnet with reason: host reimage * 07:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1093.eqiad.wmnet with reason: host reimage * 07:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2092.codfw.wmnet with reason: host reimage * 07:32 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1093.eqiad.wmnet with reason: host reimage * 07:29 jynus: running extra backups to test db1265 [[phab:T433825|T433825]] * 07:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2092.codfw.wmnet with OS trixie * 07:11 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1093.eqiad.wmnet with OS trixie * 06:50 slyngshede@dns1004: END - running authdns-update * 06:48 slyngshede@dns1004: START - running authdns-update * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.11 (duration: 02m 29s) * 03:38 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] (duration: 32m 57s) * 03:23 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 03:22 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 03:05 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 32s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 00:45 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] (duration: 06m 20s) * 00:41 cjming@deploy1003: cjming: Continuing with deployment * 00:41 cjming@deploy1003: cjming: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:39 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] == 2026-08-03 == * 23:58 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply * 23:57 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply * 23:29 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cp5021.eqsin.wmnet * 23:29 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cp5021.eqsin.wmnet * 23:27 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cp5021.eqsin.wmnet * 23:26 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cp5021.eqsin.wmnet * 23:18 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 23:17 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 22:56 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: sync * 22:56 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: sync * 22:36 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 22:36 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 22:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host search-loader2002.codfw.wmnet with OS trixie * 21:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on search-loader2002.codfw.wmnet with reason: host reimage * 21:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on search-loader2002.codfw.wmnet with reason: host reimage * 21:42 dancy@deploy1003: Stopping before sync operations * 21:41 dancy@deploy1003: Started scap sync-world: testing * 21:39 dancy@deploy1003: Installation of scap version "4.277.0" completed for 3 hosts * 21:37 dancy@deploy1003: Installing scap version "4.277.0" for 3 host(s) * 21:37 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] (duration: 06m 13s) * 21:33 dancy@deploy1003: dancy: Continuing with deployment * 21:32 dancy@deploy1003: dancy: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:31 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] * 21:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host search-loader2002.codfw.wmnet with OS trixie * 21:03 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] (duration: 06m 34s) * 20:59 dancy@deploy1003: dancy: Continuing with deployment * 20:58 dancy@deploy1003: dancy: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:56 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] * 20:52 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] (duration: 06m 23s) * 20:48 cjming@deploy1003: cjming: Continuing with deployment * 20:47 cjming@deploy1003: cjming: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:46 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] * 20:42 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] (duration: 07m 36s) * 20:38 arlolra@deploy1003: arlolra: Continuing with deployment * 20:36 arlolra@deploy1003: arlolra: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:34 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] * 20:16 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] (duration: 08m 26s) * 20:12 krinkle@deploy1003: krinkle: Continuing with deployment * 20:09 krinkle@deploy1003: krinkle: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] * 19:45 jasmine@cumin2002: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-main-codfw * 18:58 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] (duration: 09m 23s) * 18:53 krinkle@deploy1003: krinkle: Continuing with deployment * 18:53 jasmine@cumin2002: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-main-codfw * 18:50 krinkle@deploy1003: krinkle: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:48 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] * 18:37 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] (duration: 10m 13s) * 18:34 dzahn@cumin2002: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host codesearch2001.codfw.wmnet * 18:34 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host codesearch2001.codfw.wmnet with OS trixie * 18:33 krinkle@deploy1003: krinkle: Continuing with deployment * 18:29 krinkle@deploy1003: krinkle: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:27 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] * 18:19 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 18:18 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on codesearch2001.codfw.wmnet with reason: host reimage * 18:14 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 18:14 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:12 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on codesearch2001.codfw.wmnet with reason: host reimage * 18:11 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 18:11 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 18:10 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 18:02 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 18:02 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 18:01 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 18:01 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 17:55 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host codesearch2001.codfw.wmnet with OS trixie * 17:54 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:54 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) codesearch2001.codfw.wmnet on all recursors * 17:53 dzahn@cumin2002: START - Cookbook sre.dns.wipe-cache codesearch2001.codfw.wmnet on all recursors * 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:48 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:41 dzahn@cumin2002: START - Cookbook sre.dns.netbox * 17:41 dzahn@cumin2002: START - Cookbook sre.ganeti.makevm for new host codesearch2001.codfw.wmnet * 17:37 dzahn@cumin2002: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host codesearch1001.eqiad.wmnet * 17:37 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host codesearch1001.eqiad.wmnet with OS trixie * 17:24 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on codesearch1001.eqiad.wmnet with reason: host reimage * 17:17 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on codesearch1001.eqiad.wmnet with reason: host reimage * 17:08 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host codesearch1001.eqiad.wmnet with OS trixie * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 17:06 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:06 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) codesearch1001.eqiad.wmnet on all recursors * 17:06 dzahn@cumin2002: START - Cookbook sre.dns.wipe-cache codesearch1001.eqiad.wmnet on all recursors * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 17:05 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 17:04 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:04 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 16:58 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 16:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2091.codfw.wmnet with OS trixie * 16:54 ebernhardson@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 16:54 ebernhardson@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 16:49 ebernhardson@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 16:49 ebernhardson@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 16:46 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1092.eqiad.wmnet with OS trixie * 16:43 ebernhardson@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 16:43 ebernhardson@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 16:43 dzahn@cumin2002: START - Cookbook sre.dns.netbox * 16:43 dzahn@cumin2002: START - Cookbook sre.ganeti.makevm for new host codesearch1001.eqiad.wmnet * 16:41 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 16:41 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 16:40 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2091.codfw.wmnet with reason: host reimage * 16:37 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 16:35 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2091.codfw.wmnet with reason: host reimage * 16:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1092.eqiad.wmnet with reason: host reimage * 16:24 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 16:24 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 16:23 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1092.eqiad.wmnet with reason: host reimage * 16:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2091.codfw.wmnet with OS trixie * 16:03 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1092.eqiad.wmnet with OS trixie * 16:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2090.codfw.wmnet with OS trixie * 15:51 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 15:51 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 15:51 jiji@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 15:50 jiji@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 15:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2090.codfw.wmnet with reason: host reimage * 15:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2090.codfw.wmnet with reason: host reimage * 15:33 jhathaway@dns1004: END - running authdns-update * 15:31 jhathaway@dns1004: START - running authdns-update * 15:26 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1091.eqiad.wmnet with OS trixie * 15:25 dancy@deploy1003: Installation of scap version "4.276.1" completed for 3 hosts * 15:23 dancy@deploy1003: Installing scap version "4.276.1" for 3 host(s) * 15:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2090.codfw.wmnet with OS trixie * 15:12 marostegui@cumin1003: dbctl commit (dc=all): 'Repool db2245, db2246, db2247 and db2248 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95857 and previous config saved to /var/cache/conftool/dbconfig/20260803-151212-marostegui.json * 15:09 dancy@deploy1003: Started scap sync-world: testing * 15:09 dancy@deploy1003: Installation of scap version "4.277.0" completed for 3 hosts * 15:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1091.eqiad.wmnet with reason: host reimage * 15:07 dancy@deploy1003: Installing scap version "4.277.0" for 3 host(s) * 15:03 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1091.eqiad.wmnet with reason: host reimage * 14:49 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1091.eqiad.wmnet with OS trixie * 14:33 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2089.codfw.wmnet with OS trixie * 14:29 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 14:27 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 14:18 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 14:16 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 14:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2089.codfw.wmnet with reason: host reimage * 14:10 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 14:10 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 14:09 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2089.codfw.wmnet with reason: host reimage * 13:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2089.codfw.wmnet with OS trixie * 13:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2088.codfw.wmnet with OS trixie * 13:40 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1090.eqiad.wmnet with OS trixie * 13:22 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1090.eqiad.wmnet with reason: host reimage * 13:22 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] (duration: 14m 34s) * 13:19 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1090.eqiad.wmnet with reason: host reimage * 13:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2088.codfw.wmnet with reason: host reimage * 13:16 aude@deploy1003: aude, mhorsey: Continuing with deployment * 13:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2088.codfw.wmnet with reason: host reimage * 13:12 aude@deploy1003: aude, mhorsey: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] * 13:05 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1090.eqiad.wmnet with OS trixie * 12:58 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2088.codfw.wmnet with OS trixie * 12:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db[2245-2247].codfw.wmnet * 12:49 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2247: Rebooting db2247.codfw.wmnet * 12:49 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2247: Rebooting db2247.codfw.wmnet * 12:42 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2246: Rebooting db2246.codfw.wmnet * 12:42 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2246: Rebooting db2246.codfw.wmnet * 12:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2087.codfw.wmnet with OS trixie * 12:37 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1089.eqiad.wmnet with OS trixie * 12:34 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2245: Rebooting db2245.codfw.wmnet * 12:34 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2245: Rebooting db2245.codfw.wmnet * 12:34 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db[2245-2247].codfw.wmnet * 12:32 kamila@deploy1003: Finished scap sync-world: rebuild after base image update (duration: 30m 26s) * 12:28 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 12:22 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2087.codfw.wmnet with reason: host reimage * 12:19 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1089.eqiad.wmnet with reason: host reimage * 12:14 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2087.codfw.wmnet with reason: host reimage * 12:14 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1089.eqiad.wmnet with reason: host reimage * 12:03 kamila@deploy1003: Started scap sync-world: rebuild after base image update * 12:00 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1089.eqiad.wmnet with OS trixie * 12:00 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2087.codfw.wmnet with OS trixie * 11:35 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db[2245-2248].codfw.wmnet * 11:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db[2245-2248].codfw.wmnet * 11:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2086.codfw.wmnet with OS trixie * 11:26 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 11:26 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 11:25 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 11:25 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 11:24 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1088.eqiad.wmnet with OS trixie * 11:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db[2245-2248].codfw.wmnet with reason: Checking network * 11:21 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 11:20 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 11:19 marostegui@dns1004: END - running authdns-update * 11:17 marostegui@dns1004: START - running authdns-update * 11:10 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:10 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 11:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2086.codfw.wmnet with reason: host reimage * 11:09 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:08 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 11:08 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:07 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 11:07 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:07 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 11:06 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop: apply * 11:06 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop: apply * 11:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1088.eqiad.wmnet with reason: host reimage * 11:05 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop: apply * 11:04 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop: apply * 11:04 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop: apply * 11:04 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop: apply * 11:02 marostegui@dns1004: END - running authdns-update * 11:02 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2086.codfw.wmnet with reason: host reimage * 11:01 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1088.eqiad.wmnet with reason: host reimage * 11:00 marostegui@dns1004: START - running authdns-update * 10:53 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] (duration: 10m 57s) * 10:51 cmooney@dns3003: END - running authdns-update * 10:49 cmooney@dns3003: START - running authdns-update * 10:47 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1088.eqiad.wmnet with OS trixie * 10:47 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2086.codfw.wmnet with OS trixie * 10:47 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 10:46 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:46 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:46 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new reverse ranges for eqsin CR switch links - cmooney@cumin1003" * 10:46 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new reverse ranges for eqsin CR switch links - cmooney@cumin1003" * 10:42 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] * 10:41 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 10:36 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2245, db2246 and db2247 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95855 and previous config saved to /var/cache/conftool/dbconfig/20260803-103652-marostegui.json * 10:35 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2248 from s4 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95854 and previous config saved to /var/cache/conftool/dbconfig/20260803-103535-marostegui.json * 10:27 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 10:27 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 10:26 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 10:24 kart_: cxserver: Add referencePunctuation config ([[phab:T97231|T97231]]) * 10:24 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 10:23 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:23 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:23 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:22 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:22 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply * 10:21 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply * 10:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2085.codfw.wmnet with OS trixie * 10:20 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply * 10:20 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply * 10:18 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply * 10:18 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply * 10:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1087.eqiad.wmnet with OS trixie * 09:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2085.codfw.wmnet with reason: host reimage * 09:43 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1087.eqiad.wmnet with reason: host reimage * 09:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2085.codfw.wmnet with reason: host reimage * 09:40 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1087.eqiad.wmnet with reason: host reimage * 09:26 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1087.eqiad.wmnet with OS trixie * 09:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2085.codfw.wmnet with OS trixie * 09:13 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2084.codfw.wmnet with OS trixie * 09:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1086.eqiad.wmnet with OS trixie * 08:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2084.codfw.wmnet with reason: host reimage * 08:50 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2084.codfw.wmnet with reason: host reimage * 08:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1086.eqiad.wmnet with reason: host reimage * 08:39 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1086.eqiad.wmnet with reason: host reimage * 08:38 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:38 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:37 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 08:37 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 08:35 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2084.codfw.wmnet with OS trixie * 08:34 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 08:34 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:27 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1086.eqiad.wmnet with OS trixie * 08:09 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1218: Repool after a crash * 08:07 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2083.codfw.wmnet with OS trixie * 08:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1085.eqiad.wmnet with OS trixie * 07:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2083.codfw.wmnet with reason: host reimage * 07:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1085.eqiad.wmnet with reason: host reimage * 07:40 kart_: Updated cxsever to 2026-07-16-140518-production ([[phab:T97231|T97231]]) * 07:39 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply * 07:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2083.codfw.wmnet with reason: host reimage * 07:38 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply * 07:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1085.eqiad.wmnet with reason: host reimage * 07:37 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] (duration: 32m 40s) * 07:33 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply * 07:33 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply * 07:25 jdlrobson@deploy1003: jdlrobson: Continuing with deployment * 07:24 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2083.codfw.wmnet with OS trixie * 07:24 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1085.eqiad.wmnet with OS trixie * 07:23 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1218: Repool after a crash * 07:21 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:09 marostegui: Drop renamed tables [[phab:T425074|T425074]] * 07:04 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] * 06:55 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply * 06:54 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 46s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-02 == * 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 01m 03s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-01 == * 03:30 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:30 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:30 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:30 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 34s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-31 == * 17:41 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 17:41 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 17:40 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 17:40 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 15:33 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2195: Testing * 15:02 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:02 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 15:02 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 14:48 pt1979@cumin2002: START - Cookbook sre.dns.netbox * 14:47 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2195: Testing * 14:22 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 14:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2195: Testing * 14:21 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 14:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2195.codfw.wmnet with reason: Testing * 14:16 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 14:04 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 14:04 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 13:30 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1048.eqiad.wmnet with OS trixie * 13:22 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2195: Testing * 13:22 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 13:19 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2195: Testing * 13:18 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 13:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2195: Testing * 13:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 13:05 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 13:05 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 13:04 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 13:04 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 12:50 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lswtest-d8-eqiad * 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:53 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:42 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 6515 * 11:37 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 6515 * 11:28 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:27 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:07 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2082.codfw.wmnet with OS trixie * 10:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2082.codfw.wmnet with reason: host reimage * 10:42 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2082.codfw.wmnet with reason: host reimage * 10:28 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2082.codfw.wmnet with OS trixie * 10:02 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 09:52 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 09:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts * 09:16 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts * 08:57 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 08:46 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:42 gkyziridis@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 08:37 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 08:37 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 08:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 08:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 08:11 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:11 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:08 filippo@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudvirt1048 * 08:07 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 08:07 filippo@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudvirt1048 * 08:06 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 08:01 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 08:00 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 07:19 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1048.eqiad.wmnet with reason: host reimage * 07:13 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1048.eqiad.wmnet with reason: host reimage * 07:11 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 07:11 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 07:09 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 07:09 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 06:57 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:56 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:48 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:48 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:44 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie * 06:34 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1048.eqiad.wmnet with OS trixie * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 54s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 00:57 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] (duration: 11m 04s) * 00:53 dreamyjazz@deploy1003: dreamyjazz, jforrester: Continuing with deployment * 00:48 dreamyjazz@deploy1003: dreamyjazz, jforrester: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:46 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] == 2026-07-30 == * 21:37 dancy@deploy1003: Installation of scap version "4.276.1" completed for 3 hosts * 21:35 dancy@deploy1003: Installing scap version "4.276.1" for 3 host(s) * 21:24 dancy@deploy1003: Installation of scap version "4.276.0" completed for 3 hosts * 21:22 dancy@deploy1003: Installing scap version "4.276.0" for 3 host(s) * 21:15 maryum: Deployed security fix for [[phab:T430601|T430601]] * 20:13 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] (duration: 09m 20s) * 20:07 arlolra@deploy1003: osleger, arlolra: Continuing with deployment * 20:05 arlolra@deploy1003: osleger, arlolra: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:03 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] * 19:29 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:29 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:25 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service * 19:24 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 19:24 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:24 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:24 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 19:23 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service * 19:20 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1084.eqiad.wmnet with OS trixie * 18:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1084.eqiad.wmnet with reason: host reimage * 18:52 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1084.eqiad.wmnet with reason: host reimage * 18:41 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 18:40 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 18:39 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1084.eqiad.wmnet with OS trixie * 18:25 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 18:15 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 17:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1083.eqiad.wmnet with OS trixie * 17:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1048: Maintenance * 17:36 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new security plugin settings - bking@cumin2003 - [[phab:T350516|T350516]] * 17:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1083.eqiad.wmnet with reason: host reimage * 17:26 inflatador: bking@apt1002 `reprepro --noskipold --component thirdparty/opensearch3 update trixie-wikimedia` [[phab:T433624|T433624]] * 17:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1083.eqiad.wmnet with reason: host reimage * 17:23 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 17:20 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 17:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2081.codfw.wmnet with OS trixie * 17:11 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new security plugin settings - bking@cumin2003 - [[phab:T350516|T350516]] * 17:10 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1083.eqiad.wmnet with OS trixie * 16:55 root@cumin1003: START - Cookbook sre.mysql.pool pool es1048: Maintenance * 16:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2081.codfw.wmnet with reason: host reimage * 16:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1048 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95833 and previous config saved to /var/cache/conftool/dbconfig/20260730-165053-cwilliams.json * 16:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1048.eqiad.wmnet with reason: Maintenance * 16:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1040: Maintenance * 16:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2081.codfw.wmnet with reason: host reimage * 16:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1082.eqiad.wmnet with OS trixie * 16:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2081.codfw.wmnet with OS trixie * 16:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2097.codfw.wmnet with OS trixie * 16:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1082.eqiad.wmnet with reason: host reimage * 16:08 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1082.eqiad.wmnet with reason: host reimage * 16:08 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 16:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1047: Maintenance * 16:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2080.codfw.wmnet with OS trixie * 16:04 root@cumin1003: START - Cookbook sre.mysql.pool pool es1040: Maintenance * 16:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1040: Maintenance * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new logging settings - bking@cumin2003 - [[phab:T324335|T324335]] * 15:58 root@cumin1003: START - Cookbook sre.mysql.pool pool es1040: Maintenance * 15:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1040 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95827 and previous config saved to /var/cache/conftool/dbconfig/20260730-155324-cwilliams.json * 15:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1040.eqiad.wmnet with reason: Maintenance * 15:50 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1082.eqiad.wmnet with OS trixie * 15:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2048: Maintenance * 15:44 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 24s) * 15:43 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2080.codfw.wmnet with reason: host reimage * 15:38 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new logging settings - bking@cumin2003 - [[phab:T324335|T324335]] * 15:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2080.codfw.wmnet with reason: host reimage * 15:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 15:30 mvernon@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be2097.codfw.wmnet with OS trixie * 15:23 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1081.eqiad.wmnet with OS trixie * 15:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2098.codfw.wmnet with OS trixie * 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - mvernon@cumin2003" * 15:18 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be2097.codfw.wmnet with OS trixie * 15:18 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - mvernon@cumin2003" * 15:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 15:17 root@cumin1003: START - Cookbook sre.mysql.pool pool es1047: Maintenance * 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2080.codfw.wmnet with OS trixie * 15:13 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be2097.codfw.wmnet with OS trixie * 15:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1047 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95820 and previous config saved to /var/cache/conftool/dbconfig/20260730-151200-cwilliams.json * 15:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1047.eqiad.wmnet with reason: Maintenance * 15:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: Maintenance * 15:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 15:04 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 15:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1081.eqiad.wmnet with reason: host reimage * 15:00 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1081.eqiad.wmnet with reason: host reimage * 15:00 root@cumin1003: START - Cookbook sre.mysql.pool pool es2048: Maintenance * 15:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 14:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2079.codfw.wmnet with OS trixie * 14:56 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 14:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2048 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95816 and previous config saved to /var/cache/conftool/dbconfig/20260730-145510-cwilliams.json * 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2048.codfw.wmnet with reason: Maintenance * 14:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2040: Maintenance * 14:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 14:51 tchin@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] (duration: 06m 48s) * 14:47 tchin@deploy1003: jforrester, tchin: Continuing with deployment * 14:47 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 14:47 tchin@deploy1003: jforrester, tchin: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:45 tchin@deploy1003: Started scap sync-world: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] * 14:42 sukhe@puppetserver1001: conftool action : set/weight=1; selector: cluster=urldownloader,service=squid * 14:42 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader,service=squid * 14:42 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1081.eqiad.wmnet with OS trixie * 14:39 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 14:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2079.codfw.wmnet with reason: host reimage * 14:36 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2098.codfw.wmnet with OS trixie * 14:32 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2079.codfw.wmnet with reason: host reimage * 14:30 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] (duration: 06m 31s) * 14:27 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 14:26 mszwarc@deploy1003: mszwarc: Continuing with deployment * 14:25 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:25 root@cumin1003: START - Cookbook sre.mysql.pool pool es1038: Maintenance * 14:25 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1038: Maintenance * 14:23 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] * 14:21 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] (duration: 11m 19s) * 14:20 root@cumin1003: START - Cookbook sre.mysql.pool pool es1038: Maintenance * 14:14 stran@deploy1003: stran: Continuing with deployment * 14:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1038 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95810 and previous config saved to /var/cache/conftool/dbconfig/20260730-141439-cwilliams.json * 14:14 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1038.eqiad.wmnet with reason: Maintenance * 14:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1036: Maintenance * 14:13 stran@deploy1003: stran: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2079.codfw.wmnet with OS trixie * 14:09 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] * 14:08 root@cumin1003: START - Cookbook sre.mysql.pool pool es2040: Maintenance * 14:08 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2040: Maintenance * 14:03 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] (duration: 31m 41s) * 14:03 root@cumin1003: START - Cookbook sre.mysql.pool pool es2040: Maintenance * 14:03 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2040 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95806 and previous config saved to /var/cache/conftool/dbconfig/20260730-135643-cwilliams.json * 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2040.codfw.wmnet with reason: Maintenance * 13:56 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:56 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2038: Maintenance * 13:55 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:52 lucaswerkmeister-wmde@deploy1003: migr, lucaswerkmeister-wmde: Continuing with deployment * 13:49 lucaswerkmeister-wmde@deploy1003: migr, lucaswerkmeister-wmde: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:49 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:48 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 13:45 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 13:32 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:32 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] * 13:28 root@cumin1003: START - Cookbook sre.mysql.pool pool es1036: Maintenance * 13:28 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1036: Maintenance * 13:22 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:22 root@cumin1003: START - Cookbook sre.mysql.pool pool es1036: Maintenance * 13:20 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1036 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95800 and previous config saved to /var/cache/conftool/dbconfig/20260730-131727-cwilliams.json * 13:17 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1036.eqiad.wmnet with reason: Maintenance * 13:17 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] (duration: 10m 31s) * 13:16 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2022\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 13:13 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, stran: Continuing with deployment * 13:10 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:10 root@cumin1003: START - Cookbook sre.mysql.pool pool es2038: Maintenance * 13:10 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2038: Maintenance * 13:08 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, stran: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie * 13:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2047: Maintenance * 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048 cloud-private - filippo@cumin1003" * 13:07 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048 cloud-private - filippo@cumin1003" * 13:06 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] * 13:04 root@cumin1003: START - Cookbook sre.mysql.pool pool es2038: Maintenance * 13:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2078.codfw.wmnet with OS trixie * 13:01 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2038 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95797 and previous config saved to /var/cache/conftool/dbconfig/20260730-125919-cwilliams.json * 12:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2038.codfw.wmnet with reason: Maintenance * 12:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2078.codfw.wmnet with reason: host reimage * 12:37 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2078.codfw.wmnet with reason: host reimage * 12:37 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] (duration: 06m 51s) * 12:33 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 12:32 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 12:32 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 12:32 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:30 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] * 12:19 root@cumin1003: START - Cookbook sre.mysql.pool pool es2047: Maintenance * 12:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2078.codfw.wmnet with OS trixie * 12:18 dcausse@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 12:18 dcausse@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 12:15 dcausse@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 12:14 dcausse@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 12:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2047 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95793 and previous config saved to /var/cache/conftool/dbconfig/20260730-121404-cwilliams.json * 12:13 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2047.codfw.wmnet with reason: Maintenance * 12:13 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2036: Maintenance * 12:05 ayounsi@dns1004: END - running authdns-update * 12:02 ayounsi@dns1004: START - running authdns-update * 11:51 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2077.codfw.wmnet with OS trixie * 11:48 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:46 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1080.eqiad.wmnet with OS trixie * 11:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1226: Maintenance * 11:41 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2077.codfw.wmnet with reason: host reimage * 11:28 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2077.codfw.wmnet with reason: host reimage * 11:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1080.eqiad.wmnet with reason: host reimage * 11:27 root@cumin1003: START - Cookbook sre.mysql.pool pool es2036: Maintenance * 11:27 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2036: Maintenance * 11:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1080.eqiad.wmnet with reason: host reimage * 11:21 root@cumin1003: START - Cookbook sre.mysql.pool pool es2036: Maintenance * 11:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2036 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95786 and previous config saved to /var/cache/conftool/dbconfig/20260730-111633-cwilliams.json * 11:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2036.codfw.wmnet with reason: Maintenance * 11:08 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2077.codfw.wmnet with OS trixie * 11:07 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 11:03 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie * 11:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1226: Maintenance * 10:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1226 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95783 and previous config saved to /var/cache/conftool/dbconfig/20260730-104801-cwilliams.json * 10:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1226.eqiad.wmnet with reason: Maintenance * 10:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1214: Maintenance * 10:27 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2035: Maintenance * 10:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2076.codfw.wmnet with OS trixie * 10:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1214: Maintenance * 09:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1214 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95775 and previous config saved to /var/cache/conftool/dbconfig/20260730-095451-cwilliams.json * 09:54 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1214.eqiad.wmnet with reason: Maintenance * 09:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1209: Maintenance * 09:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2076.codfw.wmnet with reason: host reimage * 09:42 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool es2035: Maintenance * 09:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.netbox.update-extras (exit_code=0) rolling restart_daemons on A:netbox * 09:41 ayounsi@cumin1003: START - Cookbook sre.netbox.update-extras rolling restart_daemons on A:netbox * 09:40 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2035: Maintenance * 09:39 ayounsi@cumin1003: END (PASS) - Cookbook sre.netbox.update-extras (exit_code=0) rolling restart_daemons on A:netbox-canary * 09:39 ayounsi@cumin1003: START - Cookbook sre.netbox.update-extras rolling restart_daemons on A:netbox-canary * 09:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2076.codfw.wmnet with reason: host reimage * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 09:34 root@cumin1003: START - Cookbook sre.mysql.pool pool es2035: Maintenance * 09:32 lucaswerkmeister-wmde@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 09:32 lucaswerkmeister-wmde@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 09:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2035 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95771 and previous config saved to /var/cache/conftool/dbconfig/20260730-092910-cwilliams.json * 09:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2035.codfw.wmnet with reason: Maintenance * 09:19 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2076.codfw.wmnet with OS trixie * 09:18 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 09:17 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie * 09:07 root@cumin1003: START - Cookbook sre.mysql.pool pool db1209: Maintenance * 09:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 23 hosts * 09:04 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Remove cable label from interfaces descriptions - ayounsi@cumin1003 * 09:04 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:02 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Remove cable label from interfaces descriptions - ayounsi@cumin1003 * 09:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1209 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95767 and previous config saved to /var/cache/conftool/dbconfig/20260730-090133-cwilliams.json * 09:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1209.eqiad.wmnet with reason: Maintenance * 09:01 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1192: Maintenance * 08:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1252: Maintenance * 08:57 jayme@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 08:56 jayme@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 08:53 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 23 hosts * 08:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:51 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:50 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1263: Maintenance * 08:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:23 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 08:15 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie * 08:14 root@cumin1003: START - Cookbook sre.mysql.pool pool db1192: Maintenance * 08:13 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2075.codfw.wmnet with OS trixie * 08:12 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1252: Maintenance * 08:11 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1252: Maintenance * 08:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1252: Maintenance * 08:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1192 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95754 and previous config saved to /var/cache/conftool/dbconfig/20260730-080611-cwilliams.json * 08:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1192.eqiad.wmnet with reason: Maintenance * 08:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1178: Maintenance * 08:05 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS bullseye * 07:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1263: Maintenance * 07:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1263 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95751 and previous config saved to /var/cache/conftool/dbconfig/20260730-075106-cwilliams.json * 07:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[1260-1262].eqiad.wmnet with reason: Maintenance * 07:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2075.codfw.wmnet with reason: host reimage * 07:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1263.eqiad.wmnet with reason: Maintenance * 07:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2075.codfw.wmnet with reason: host reimage * 07:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance * 07:38 dcausse@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:38 dcausse@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 07:35 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db1252', diff saved to https://phabricator.wikimedia.org/P95748 and previous config saved to /var/cache/conftool/dbconfig/20260730-073510-marostegui.json * 07:26 klausman@dns2004: END - running authdns-update * 07:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 07:25 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2075.codfw.wmnet with OS trixie * 07:24 klausman@dns2004: START - running authdns-update * 07:23 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host an-test-master1003.eqiad.wmnet * 07:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db1178: Maintenance * 07:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1178 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95746 and previous config saved to /var/cache/conftool/dbconfig/20260730-071112-cwilliams.json * 07:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1178.eqiad.wmnet with reason: Maintenance * 07:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1177: Maintenance * 06:24 root@cumin1003: START - Cookbook sre.mysql.pool pool db1177: Maintenance * 06:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1177 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95741 and previous config saved to /var/cache/conftool/dbconfig/20260730-061736-cwilliams.json * 06:17 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1177.eqiad.wmnet with reason: Maintenance * 06:17 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1172: Maintenance * 05:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1218.eqiad.wmnet with reason: crashed * 05:41 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db1217 it crashed', diff saved to https://phabricator.wikimedia.org/P95737 and previous config saved to /var/cache/conftool/dbconfig/20260730-054111-marostegui.json * 05:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95736 and previous config saved to /var/cache/conftool/dbconfig/20260730-053422-cwilliams.json * 05:30 root@cumin1003: START - Cookbook sre.mysql.pool pool db1172: Maintenance * 05:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95734 and previous config saved to /var/cache/conftool/dbconfig/20260730-052414-cwilliams.json * 05:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1172 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95733 and previous config saved to /var/cache/conftool/dbconfig/20260730-052354-cwilliams.json * 05:23 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1172.eqiad.wmnet with reason: Maintenance * 05:23 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1167: Maintenance * 05:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95731 and previous config saved to /var/cache/conftool/dbconfig/20260730-051406-cwilliams.json * 05:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95729 and previous config saved to /var/cache/conftool/dbconfig/20260730-050358-cwilliams.json * 04:47 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 04:47 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 04:47 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 04:35 root@cumin1003: START - Cookbook sre.mysql.pool pool db1167: Maintenance * 04:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1167 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95726 and previous config saved to /var/cache/conftool/dbconfig/20260730-042923-cwilliams.json * 04:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 04:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1167.eqiad.wmnet with reason: Maintenance * 04:22 pt1979@cumin2002: START - Cookbook sre.dns.netbox * 04:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95725 and previous config saved to /var/cache/conftool/dbconfig/20260730-040337-cwilliams.json * 04:03 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:38 brett@cumin2002: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool eqsin [reason: Switch upgrade maintenance window complete, [[phab:T433097|T433097]]] * 01:38 brett@cumin2002: START - Cookbook sre.dns.admin DNS admin: pool eqsin [reason: Switch upgrade maintenance window complete, [[phab:T433097|T433097]]] == 2026-07-29 == * 23:57 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin,mr1-eqsin IPv6,mr1-eqsin.oob,mr1-eqsin.oob IPv6 with reason: connection issue * 22:54 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 22:53 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 22:53 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 22:53 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:25 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2022.codfw.wmnet, repooling source-only afterwards * 22:20 brett@cumin2002: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool eqsin [reason: Switch upgrade maintenance window, [[phab:T433097|T433097]]] * 22:20 brett@cumin2002: START - Cookbook sre.dns.admin DNS admin: depool eqsin [reason: Switch upgrade maintenance window, [[phab:T433097|T433097]]] * 22:01 apine@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 22:00 apine@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 21:59 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 21:58 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 21:58 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 21:58 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 21:32 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1048.eqiad.wmnet with OS trixie * 21:25 pt1979@cumin2002: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be2097.codfw.wmnet with OS bullseye * 21:16 zabe@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki=metawiki 'Mental Health Resource Center' 'Safety Resource Center/Mental Health' Zabe --reason 'per request [[:phab:T433118{{!}}T433118]]' * 21:12 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2022.codfw.wmnet, repooling source-only afterwards * 21:12 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] (duration: 12m 53s) * 21:12 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2015\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 21:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1253: Maintenance * 21:08 aaron@deploy1003: aaron: Continuing with deployment * 21:01 aaron@deploy1003: aaron: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:59 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] * 20:52 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] (duration: 21m 57s) * 20:48 aaron@deploy1003: aaron: Continuing with deployment * 20:32 aaron@deploy1003: aaron: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:30 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] * 20:24 root@cumin1003: START - Cookbook sre.mysql.pool pool db1253: Maintenance * 20:19 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] (duration: 08m 07s) * 20:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1253 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95719 and previous config saved to /var/cache/conftool/dbconfig/20260729-201810-cwilliams.json * 20:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1253.eqiad.wmnet with reason: Maintenance * 20:17 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1231: Maintenance * 20:15 aaron@deploy1003: bpirkle, aaron: Continuing with deployment * 20:13 aaron@deploy1003: bpirkle, aaron: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:12 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie * 20:11 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] * 20:11 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1048.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:09 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1048.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:09 pt1979@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 20:04 pt1979@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 19:47 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 19:43 pt1979@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye * 19:41 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 19:37 zabe: zabe@deploy1003:~$ mwscript-k8s --comment='[[phab:T433529|T433529]]' --follow -- resetAuthenticationThrottle.php --wiki=aawiki --signup --ip=89.36.114.94 * 19:36 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] (duration: 06m 49s) * 19:32 zabe@deploy1003: zabe: Continuing with deployment * 19:31 zabe@deploy1003: zabe: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:31 root@cumin1003: START - Cookbook sre.mysql.pool pool db1231: Maintenance * 19:29 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] * 19:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1231 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95714 and previous config saved to /var/cache/conftool/dbconfig/20260729-192454-cwilliams.json * 19:24 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1231.eqiad.wmnet with reason: Maintenance * 19:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1227: Maintenance * 19:22 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 19:22 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 19:21 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:21 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1048] - vriley@cumin1003" * 19:21 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1048] - vriley@cumin1003" * 19:19 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 19:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1251: Maintenance * 19:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95711 and previous config saved to /var/cache/conftool/dbconfig/20260729-191756-cwilliams.json * 19:16 vriley@cumin1003: START - Cookbook sre.dns.netbox * 19:11 dduvall: rolling back wmf.13 to group0 due to [[phab:T433457|T433457]] (cc [[phab:T430832|T430832]]) * 19:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95709 and previous config saved to /var/cache/conftool/dbconfig/20260729-190748-cwilliams.json * 19:01 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2022.codfw.wmnet with OS bookworm * 19:01 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2021.codfw.wmnet, repooling source-only afterwards * 18:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95707 and previous config saved to /var/cache/conftool/dbconfig/20260729-185740-cwilliams.json * 18:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95704 and previous config saved to /var/cache/conftool/dbconfig/20260729-184732-cwilliams.json * 18:37 root@cumin1003: START - Cookbook sre.mysql.pool pool db1227: Maintenance * 18:34 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2022.codfw.wmnet with reason: host reimage * 18:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1227 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95701 and previous config saved to /var/cache/conftool/dbconfig/20260729-183117-cwilliams.json * 18:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1227.eqiad.wmnet with reason: Maintenance * 18:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1202: Maintenance * 18:30 root@cumin1003: START - Cookbook sre.mysql.pool pool db1251: Maintenance * 18:27 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2022.codfw.wmnet with reason: host reimage * 18:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1251 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95698 and previous config saved to /var/cache/conftool/dbconfig/20260729-182428-cwilliams.json * 18:24 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lvs2014.codfw.wmnet * 18:24 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for lvs2014.codfw.wmnet * 18:24 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1251.eqiad.wmnet with reason: Maintenance * 18:23 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1235: Maintenance * 18:22 brett@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 18:19 brett@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 18:19 brett@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 18:17 brett@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 18:17 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 18:16 mutante: removing jenkins during the train - living on the edge - no, just kidding, jenkins has migrated to dedicated machines, nothing should happen * 18:15 brett@cumin2002: END (ERROR) - Cookbook sre.loadbalancer.restart-pybal (exit_code=97) rolling-restart of pybal on P<nowiki>{</nowiki>lvs2014.codfw.wmnet<nowiki>}</nowiki> and A:lvs ([[phab:T428495|T428495]]) * 18:15 mutante: CI: contint1002/contint2002: apt-get remove --purge jenkins - jenkins be gone - [[phab:T418521|T418521]] * 18:13 brett@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on P<nowiki>{</nowiki>lvs2014.codfw.wmnet<nowiki>}</nowiki> and A:lvs ([[phab:T428495|T428495]]) * 18:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2022 * 18:08 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2022 * 18:03 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T428495|T428495]] * 18:03 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2022 * 18:02 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2022.codfw.wmnet 211.48.192.10.in-addr.arpa 1.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:02 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2022.codfw.wmnet 211.48.192.10.in-addr.arpa 1.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:02 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:02 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2022 - bking@cumin2003" * 18:02 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2022 - bking@cumin2003" * 17:57 bking@cumin2003: START - Cookbook sre.dns.netbox * 17:56 brett@cumin2002: END (FAIL) - Cookbook sre.loadbalancer.restart-pybal (exit_code=1) rolling-restart of pybal on A:lvs-codfw and A:lvs ([[phab:T428495|T428495]]) * 17:55 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo - [[phab:T428495|T428495]] * 17:54 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2022 * 17:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2022.codfw.wmnet with OS bookworm * 17:50 brett@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on A:lvs-codfw and A:lvs ([[phab:T428495|T428495]]) * 17:47 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2021.codfw.wmnet, repooling source-only afterwards * 17:47 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 14s) * 17:47 swfrench-wmf: authdns-update to direct codfw, eqsin, ulsfo etcd clients back to codfw - [[phab:T428495|T428495]] * 17:47 swfrench@dns1004: END - running authdns-update * 17:47 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 17:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95692 and previous config saved to /var/cache/conftool/dbconfig/20260729-174713-cwilliams.json * 17:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance * 17:46 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1249: Maintenance * 17:45 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2015\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 17:45 swfrench@dns1004: START - running authdns-update * 17:44 root@cumin1003: START - Cookbook sre.mysql.pool pool db1202: Maintenance * 17:41 akhatun: Deployed refinery using scap, then deployed onto hdfs * 17:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1202 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95688 and previous config saved to /var/cache/conftool/dbconfig/20260729-173759-cwilliams.json * 17:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1202.eqiad.wmnet with reason: Maintenance * 17:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1194: Maintenance * 17:37 root@cumin1003: START - Cookbook sre.mysql.pool pool db1235: Maintenance * 17:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1230: Maintenance * 17:30 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1235 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95684 and previous config saved to /var/cache/conftool/dbconfig/20260729-173051-cwilliams.json * 17:30 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1235.eqiad.wmnet with reason: Maintenance * 17:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1234: Maintenance * 17:26 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (thin): Regular analytics weekly train THIN [analytics/refinery@56695674] (duration: 02m 02s) * 17:24 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (thin): Regular analytics weekly train THIN [analytics/refinery@56695674] * 17:23 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567]: Regular analytics weekly train [analytics/refinery@56695674] (duration: 06m 20s) * 17:20 dancy@deploy1003: Finished scap sync-world: Testing delay_messageblobstore_purge: true (duration: 06m 29s) * 17:17 akhatun@deploy1003: Started deploy [analytics/refinery@5669567]: Regular analytics weekly train [analytics/refinery@56695674] * 17:17 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] (duration: 00m 22s) * 17:16 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] * 17:13 dancy@deploy1003: Started scap sync-world: Testing delay_messageblobstore_purge: true * 17:05 mutante: CI: contint1002/contint2002 - restarted httpd to be extra sure all is cleaned up - https://integration.wikimedia.org/ci/ is up and running [[phab:T418521|T418521]] * 17:04 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 17:03 mutante: CI: contint1002/contint2002 - rm /etc/apache2/jenkins_proxy - removing legacy jenkins proxy config - jenkins is on new dedicated machines and uses jenkins_proxy_ext config [[phab:T418521|T418521]] * 17:02 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] (duration: 36m 25s) * 17:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1249: Maintenance * 16:59 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 16:54 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2015.codfw.wmnet, repooling source-only afterwards * 16:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1249 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95674 and previous config saved to /var/cache/conftool/dbconfig/20260729-165339-cwilliams.json * 16:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1249.eqiad.wmnet with reason: Maintenance * 16:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1248: Maintenance * 16:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1194: Maintenance * 16:47 swfrench-wmf: silenced EtcdReplicationDown 57b2b421-1cc9-4e38-9276-{{Gerrit|94f223fd231c}} - [[phab:T428495|T428495]] * 16:46 tchin@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/eventstreams-internal: apply * 16:46 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye * 16:46 tchin@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/eventstreams-internal: apply * 16:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1230: Maintenance * 16:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1194 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95669 and previous config saved to /var/cache/conftool/dbconfig/20260729-164422-cwilliams.json * 16:44 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 16:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1194.eqiad.wmnet with reason: Maintenance * 16:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1191: Maintenance * 16:43 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host an-test-master1003.eqiad.wmnet * 16:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db1234: Maintenance * 16:43 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 16:43 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Rolling back deployment * 16:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host an-test-master1004.eqiad.wmnet * 16:41 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 16:40 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 16:40 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 16:39 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 16:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1230 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95667 and previous config saved to /var/cache/conftool/dbconfig/20260729-163932-cwilliams.json * 16:39 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 16:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1230.eqiad.wmnet with reason: Maintenance * 16:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1207: Maintenance * 16:38 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 16:37 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host an-test-master1004.eqiad.wmnet * 16:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1234 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95664 and previous config saved to /var/cache/conftool/dbconfig/20260729-163719-cwilliams.json * 16:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1234.eqiad.wmnet with reason: Maintenance * 16:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1079.eqiad.wmnet with OS trixie * 16:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1232: Maintenance * 16:34 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1259: Maintenance * 16:28 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] (duration: 06m 57s) * 16:28 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:26 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] * 16:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1051 hosts * 16:21 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] * 16:20 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2006.codfw.wmnet with OS bookworm * 16:19 akhatun: Deploying Refinery at {{Gerrit|56695674}} as part of weekly train * 16:18 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1079.eqiad.wmnet with reason: host reimage * 16:16 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] (duration: 15m 36s) * 16:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2021.codfw.wmnet with OS bookworm * 16:14 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1079.eqiad.wmnet with reason: host reimage * 16:12 topranks: hot-swap line card in FPC0 on cr1-eqiad with replacement MPC10E from Juniper [[phab:T426343|T426343]] * 16:10 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Continuing with deployment * 16:07 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db1248: Maintenance * 16:01 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] * 16:00 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 16:00 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 15:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1248 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95651 and previous config saved to /var/cache/conftool/dbconfig/20260729-155956-cwilliams.json * 15:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1248.eqiad.wmnet with reason: Maintenance * 15:59 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2006.codfw.wmnet with reason: host reimage * 15:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1247: Maintenance * 15:59 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2074.codfw.wmnet with OS trixie * 15:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1191: Maintenance * 15:57 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:55 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1079.eqiad.wmnet with OS trixie * 15:55 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2006.codfw.wmnet with reason: host reimage * 15:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1207: Maintenance * 15:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1191 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95646 and previous config saved to /var/cache/conftool/dbconfig/20260729-155104-cwilliams.json * 15:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1191.eqiad.wmnet with reason: Maintenance * 15:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1181: Maintenance * 15:49 root@cumin1003: START - Cookbook sre.mysql.pool pool db1232: Maintenance * 15:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2021.codfw.wmnet with reason: host reimage * 15:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1207 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95643 and previous config saved to /var/cache/conftool/dbconfig/20260729-154735-cwilliams.json * 15:47 root@cumin1003: START - Cookbook sre.mysql.pool pool db1259: Maintenance * 15:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1207.eqiad.wmnet with reason: Maintenance * 15:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1200: Maintenance * 15:46 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] (duration: 31m 59s) * 15:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 15:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1232 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95640 and previous config saved to /var/cache/conftool/dbconfig/20260729-154330-cwilliams.json * 15:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1232.eqiad.wmnet with reason: Maintenance * 15:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1219: Maintenance * 15:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2015.codfw.wmnet, repooling source-only afterwards * 15:41 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 18s) * 15:41 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1259 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95638 and previous config saved to /var/cache/conftool/dbconfig/20260729-154107-cwilliams.json * 15:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1259.eqiad.wmnet with reason: Maintenance * 15:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1254: Maintenance * 15:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2015.codfw.wmnet with OS bookworm * 15:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2021.codfw.wmnet with reason: host reimage * 15:36 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2006.codfw.wmnet with OS bookworm * 15:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 15:35 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Continuing with deployment * 15:33 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2074.codfw.wmnet with OS trixie * 15:33 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 15:32 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:29 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:28 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be2074.codfw.wmnet with OS trixie * 15:28 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2006.codfw.wmnet * 15:26 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1078.eqiad.wmnet with OS trixie * 15:25 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 15:25 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:22 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2006.codfw.wmnet * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2021 * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2021 * 15:19 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2021 * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2021.codfw.wmnet 210.48.192.10.in-addr.arpa 0.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:19 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2021.codfw.wmnet 210.48.192.10.in-addr.arpa 0.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2021 - bking@cumin2003" * 15:19 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2021 - bking@cumin2003" * 15:14 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] * 15:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2015.codfw.wmnet with reason: host reimage * 15:11 root@cumin1003: START - Cookbook sre.mysql.pool pool db1247: Maintenance * 15:11 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on ml-serve2004.codfw.wmnet with reason: [[phab:T433478|T433478]] * 15:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2015.codfw.wmnet with reason: host reimage * 15:10 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on ml-serve2002.codfw.wmnet with reason: [[phab:T433476|T433476]] * 15:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 15:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1247 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95625 and previous config saved to /var/cache/conftool/dbconfig/20260729-150459-cwilliams.json * 15:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1247.eqiad.wmnet with reason: Maintenance * 15:04 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:04 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1244: Maintenance * 15:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1078.eqiad.wmnet with reason: host reimage * 15:03 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1005.wikimedia.org * 15:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db1181: Maintenance * 15:01 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2098.codfw.wmnet with OS bullseye * 15:00 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye * 15:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1200: Maintenance * 14:59 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 14:59 root@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1285.eqiad.wmnet with OS trixie * 14:59 Amir1: mwscript-k8s -- extensions/TimedMediaHandler/maintenance/requeueTranscodes.php --wiki=commonswiki --key '360p.mpeg4.mov' --throttle --video --missing ([[phab:T358266|T358266]]) * 14:58 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1005.wikimedia.org * 14:58 jhancock@cumin2002: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['ms-be2098'] * 14:58 jhancock@cumin2002: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['ms-be2098'] * 14:58 jhancock@cumin2002: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['ms-be2097'] * 14:58 jhancock@cumin2002: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['ms-be2097'] * 14:58 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1078.eqiad.wmnet with reason: host reimage * 14:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1181 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95621 and previous config saved to /var/cache/conftool/dbconfig/20260729-145629-cwilliams.json * 14:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1181.eqiad.wmnet with reason: Maintenance * 14:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1174: Maintenance * 14:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1219: Maintenance * 14:55 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader1006.wikimedia.org on all recursors * 14:55 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader1006.wikimedia.org on all recursors * 14:55 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader1005.wikimedia.org on all recursors * 14:55 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader1005.wikimedia.org on all recursors * 14:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1200 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95618 and previous config saved to /var/cache/conftool/dbconfig/20260729-145336-cwilliams.json * 14:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1254: Maintenance * 14:53 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1200.eqiad.wmnet with reason: Maintenance * 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1185: Maintenance * 14:52 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2021 * 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2015 * 14:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2015 * 14:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1219 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95616 and previous config saved to /var/cache/conftool/dbconfig/20260729-144946-cwilliams.json * 14:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1219.eqiad.wmnet with reason: Maintenance * 14:49 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1218: Maintenance * 14:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2021.codfw.wmnet with OS bookworm * 14:48 dancy@deploy1003: Finished deploy [zuul/deploy@22703a6]: Deploying https://gerrit.wikimedia.org/r/c/integration/zuul/+/1311501 ([[phab:T432491|T432491]]) (duration: 00m 15s) * 14:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2015.codfw.wmnet with OS bookworm * 14:48 dancy@deploy1003: Started deploy [zuul/deploy@22703a6]: Deploying https://gerrit.wikimedia.org/r/c/integration/zuul/+/1311501 ([[phab:T432491|T432491]]) * 14:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1254 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95613 and previous config saved to /var/cache/conftool/dbconfig/20260729-144729-cwilliams.json * 14:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1254.eqiad.wmnet with reason: Maintenance * 14:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1233: Maintenance * 14:46 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2013\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 14:46 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2014\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 14:46 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:45 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:44 root@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1285.eqiad.wmnet with reason: host reimage * 14:43 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:42 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:41 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:40 root@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1285.eqiad.wmnet with reason: host reimage * 14:39 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1078.eqiad.wmnet with OS trixie * 14:39 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2074.codfw.wmnet with OS trixie * 14:32 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:32 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:32 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:31 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2005.codfw.wmnet with OS bookworm * 14:30 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:30 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:29 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:29 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:27 root@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host db1285 * 14:27 root@cumin1003: START - Cookbook sre.hosts.move-vlan for host db1285 * 14:27 root@cumin1003: START - Cookbook sre.hosts.reimage for host db1285.eqiad.wmnet with OS trixie * 14:24 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 14:24 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:24 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:24 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:23 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:22 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:22 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:21 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:17 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:16 root@cumin1003: START - Cookbook sre.mysql.pool pool db1244: Maintenance * 14:15 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:15 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add asw1-604 loopback ipv4 - pt1979@cumin2002" * 14:15 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add asw1-604 loopback ipv4 - pt1979@cumin2002" * 14:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:12 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 14:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95599 and previous config saved to /var/cache/conftool/dbconfig/20260729-141014-cwilliams.json * 14:10 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 14:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1244.eqiad.wmnet with reason: Maintenance * 14:10 pt1979@cumin2002: START - Cookbook sre.dns.netbox * 14:10 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 14:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1243: Maintenance * 14:09 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2005.codfw.wmnet with reason: host reimage * 14:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db1174: Maintenance * 14:08 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad * 14:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db1185: Maintenance * 14:06 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 14:05 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2005.codfw.wmnet with reason: host reimage * 14:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1174 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95595 and previous config saved to /var/cache/conftool/dbconfig/20260729-140309-cwilliams.json * 14:03 sukhe@dns1004: END - running authdns-update * 14:03 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1174.eqiad.wmnet with reason: Maintenance * 14:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1170: Maintenance * 14:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db1218: Maintenance * 14:01 sukhe@dns1004: START - running authdns-update * 14:00 sukhe@dns1004: START - running authdns-update * 13:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db1233: Maintenance * 13:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1185 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95592 and previous config saved to /var/cache/conftool/dbconfig/20260729-135925-cwilliams.json * 13:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1185.eqiad.wmnet with reason: Maintenance * 13:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1161: Maintenance * 13:58 sukhe@puppetserver1001: conftool action : set/pooled=true; selector: dnsdisc=urldownloader * 13:58 root@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1265.eqiad.wmnet with OS trixie * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1218 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95590 and previous config saved to /var/cache/conftool/dbconfig/20260729-135621-cwilliams.json * 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1218.eqiad.wmnet with reason: Maintenance * 13:55 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1206: Maintenance * 13:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2073.codfw.wmnet with OS trixie * 13:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1233 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95587 and previous config saved to /var/cache/conftool/dbconfig/20260729-135335-cwilliams.json * 13:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1233.eqiad.wmnet with reason: Maintenance * 13:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1229: Maintenance * 13:50 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/kartotherian: apply * 13:50 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service * 13:49 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:49 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/kartotherian: apply * 13:48 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 13:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1077.eqiad.wmnet with OS trixie * 13:47 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 13:46 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2005.codfw.wmnet with OS bookworm * 13:44 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 13:44 root@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1265.eqiad.wmnet with reason: host reimage * 13:40 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] (duration: 09m 22s) * 13:39 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:38 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:36 root@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1265.eqiad.wmnet with reason: host reimage * 13:35 stran@deploy1003: stran: Continuing with deployment * 13:33 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 13:32 stran@deploy1003: stran: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified t * 13:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2073.codfw.wmnet with reason: host reimage * 13:30 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ml-build1001.eqiad.wmnet * 13:30 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] * 13:29 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad * 13:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1077.eqiad.wmnet with reason: host reimage * 13:27 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 13:27 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:27 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad * 13:26 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2073.codfw.wmnet with reason: host reimage * 13:26 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] (duration: 07m 56s) * 13:25 klausman@cumin1003: START - Cookbook sre.hosts.reboot-single for host ml-build1001.eqiad.wmnet * 13:24 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 13:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>ml-serve101[2-5].eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 13:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1015.eqiad.wmnet * 13:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1015.eqiad.wmnet * 13:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1077.eqiad.wmnet with reason: host reimage * 13:23 root@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host db1265 * 13:23 root@cumin1003: START - Cookbook sre.hosts.move-vlan for host db1265 * 13:23 root@cumin1003: START - Cookbook sre.hosts.reimage for host db1265.eqiad.wmnet with OS trixie * 13:23 root@cumin1003: START - Cookbook sre.mysql.pool pool db1243: Maintenance * 13:22 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 13:22 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2005.codfw.wmnet * 13:22 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 13:22 samtar@deploy1003: dreamrimmer, samtar: Continuing with deployment * 13:20 samtar@deploy1003: dreamrimmer, samtar: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts an-test-master[1001-1002].eqiad.wmnet * 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-master[1001-1002].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 13:18 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1015.eqiad.wmnet * 13:18 sukhe@cumin1003: END (ERROR) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=97) for role: url_downloader@eqiad * 13:18 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 13:18 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] * 13:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95574 and previous config saved to /var/cache/conftool/dbconfig/20260729-131638-cwilliams.json * 13:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1243.eqiad.wmnet with reason: Maintenance * 13:16 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2005.codfw.wmnet * 13:16 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1242: Maintenance * 13:14 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] (duration: 07m 00s) * 13:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db1170: Maintenance * 13:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1015.eqiad.wmnet * 13:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1014.eqiad.wmnet * 13:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1014.eqiad.wmnet * 13:12 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1223: Maintenance * 13:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db1161: Maintenance * 13:10 samtar@deploy1003: anzx, samtar: Continuing with deployment * 13:09 samtar@deploy1003: anzx, samtar: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db1206: Maintenance * 13:08 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 13:07 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] * 13:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1170 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95566 and previous config saved to /var/cache/conftool/dbconfig/20260729-130730-cwilliams.json * 13:07 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1170.eqiad.wmnet with reason: Maintenance * 13:07 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:07 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt IPs new switches - cmooney@cumin1003" * 13:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1158: Maintenance * 13:06 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1014.eqiad.wmnet * 13:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1161 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95564 and previous config saved to /var/cache/conftool/dbconfig/20260729-130616-cwilliams.json * 13:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 13:06 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1077.eqiad.wmnet with OS trixie * 13:05 root@cumin1003: START - Cookbook sre.mysql.pool pool db1229: Maintenance * 13:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1161.eqiad.wmnet with reason: Maintenance * 13:05 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt IPs new switches - cmooney@cumin1003" * 13:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2073.codfw.wmnet with OS trixie * 13:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Maintenance * 13:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1206 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95562 and previous config saved to /var/cache/conftool/dbconfig/20260729-130258-cwilliams.json * 13:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1206.eqiad.wmnet with reason: Maintenance * 13:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1196: Maintenance * 13:01 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 13:01 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:00 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 13:00 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 12:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1229 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95559 and previous config saved to /var/cache/conftool/dbconfig/20260729-125950-cwilliams.json * 12:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1229.eqiad.wmnet with reason: Maintenance * 12:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1222: Maintenance * 12:57 sukhe: sudo cumin 'A:lvs and (A:eqiad or A:codfw)' 'disable-puppet "adding new service urldownloader"': [[phab:T429175|T429175]] * 12:56 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1014.eqiad.wmnet * 12:56 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1013.eqiad.wmnet * 12:56 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1013.eqiad.wmnet * 12:50 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1013.eqiad.wmnet * 12:50 sukhe: sudo cumin 'O:url_downloader' 'run-puppet-agent --enable "merging CR 1313948"': [[phab:T429175|T429175]] * 12:48 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-master[1001-1002].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 12:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1013.eqiad.wmnet * 12:45 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1012.eqiad.wmnet * 12:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1012.eqiad.wmnet * 12:45 sukhe: sudo cumin 'O:url_downloader' 'disable-puppet "merging CR 1313948"': [[phab:T429175|T429175]] * 12:44 btullis@cumin1003: START - Cookbook sre.dns.netbox * 12:40 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test2001.codfw.wmnet * 12:40 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test2001.codfw.wmnet * 12:38 ayounsi@dns1004: END - running authdns-update * 12:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1012.eqiad.wmnet * 12:35 ayounsi@dns1004: START - running authdns-update * 12:34 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts an-test-master[1001-1002].eqiad.wmnet * 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts an-test-coord1001.eqiad.wmnet * 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-coord1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 12:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1012.eqiad.wmnet * 12:32 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>ml-serve101[2-5].eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 12:29 root@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Maintenance * 12:25 root@cumin1003: START - Cookbook sre.mysql.pool pool db1223: Maintenance * 12:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95544 and previous config saved to /var/cache/conftool/dbconfig/20260729-122254-cwilliams.json * 12:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1242.eqiad.wmnet with reason: Maintenance * 12:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1241: Maintenance * 12:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1051 hosts * 12:20 root@cumin1003: START - Cookbook sre.mysql.pool pool db1158: Maintenance * 12:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1223 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95540 and previous config saved to /var/cache/conftool/dbconfig/20260729-121937-cwilliams.json * 12:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1223.eqiad.wmnet with reason: Maintenance * 12:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1212: Maintenance * 12:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db1159: Maintenance * 12:17 elukey@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: sync * 12:15 elukey@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: sync * 12:15 root@cumin1003: START - Cookbook sre.mysql.pool pool db1196: Maintenance * 12:14 Daimona: Creating new DB tables for the CampaignEvents extension in x1.testwiki, x1.test2wiki, x1.officewiki, and x1.wikishared # [[phab:T429339|T429339]] * 12:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db1222: Maintenance * 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95535 and previous config saved to /var/cache/conftool/dbconfig/20260729-121211-cwilliams.json * 12:12 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 12:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1158.eqiad.wmnet with reason: Maintenance * 12:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1159 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95534 and previous config saved to /var/cache/conftool/dbconfig/20260729-121146-cwilliams.json * 12:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1159.eqiad.wmnet with reason: Maintenance * 12:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1196 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95533 and previous config saved to /var/cache/conftool/dbconfig/20260729-120847-cwilliams.json * 12:08 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 12:08 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1196.eqiad.wmnet with reason: Maintenance * 12:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1195: Maintenance * 12:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1222 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95530 and previous config saved to /var/cache/conftool/dbconfig/20260729-120424-cwilliams.json * 12:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1222.eqiad.wmnet with reason: Maintenance * 12:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1098 hosts * 12:00 marostegui: Rename tables [[phab:T425074|T425074]] * 12:00 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-coord1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 11:58 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1197: Maintenance * 11:55 btullis@cumin1003: START - Cookbook sre.dns.netbox * 11:52 elukey@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: sync * 11:51 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:51 elukey@deploy1003: helmfile [codfw] START helmfile.d/services/proton: sync * 11:51 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:50 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts an-test-coord1001.eqiad.wmnet * 11:50 elukey@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: sync * 11:49 elukey@deploy1003: helmfile [staging] START helmfile.d/services/proton: sync * 11:35 root@cumin1003: START - Cookbook sre.mysql.pool pool db1241: Maintenance * 11:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db1212: Maintenance * 11:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1241 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95520 and previous config saved to /var/cache/conftool/dbconfig/20260729-112918-cwilliams.json * 11:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1241.eqiad.wmnet with reason: Maintenance * 11:29 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1238: Maintenance * 11:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1212 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95517 and previous config saved to /var/cache/conftool/dbconfig/20260729-112727-cwilliams.json * 11:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 11:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1212.eqiad.wmnet with reason: Maintenance * 11:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1198: Maintenance * 11:23 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:22 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:21 root@cumin1003: START - Cookbook sre.mysql.pool pool db1195: Maintenance * 11:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1195 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95514 and previous config saved to /var/cache/conftool/dbconfig/20260729-111450-cwilliams.json * 11:14 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1195.eqiad.wmnet with reason: Maintenance * 11:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1186: Maintenance * 11:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 11:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 11:05 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:54 marostegui: Dropping renamed tables [[phab:T425066|T425066]] * 10:41 root@cumin1003: START - Cookbook sre.mysql.pool pool db1238: Maintenance * 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1198: Maintenance * 10:39 Amir1: ran https://phabricator.wikimedia.org/T432509#12149723 in production ([[phab:T432509|T432509]]) * 10:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db1197: Maintenance * 10:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1238 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95501 and previous config saved to /var/cache/conftool/dbconfig/20260729-103532-cwilliams.json * 10:35 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1238.eqiad.wmnet with reason: Maintenance * 10:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1221: Maintenance * 10:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1198 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95499 and previous config saved to /var/cache/conftool/dbconfig/20260729-103330-cwilliams.json * 10:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1198.eqiad.wmnet with reason: Maintenance * 10:33 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1175: Maintenance * 10:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1197 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95496 and previous config saved to /var/cache/conftool/dbconfig/20260729-103217-cwilliams.json * 10:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1197.eqiad.wmnet with reason: Maintenance * 10:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1188: Maintenance * 10:27 root@cumin1003: START - Cookbook sre.mysql.pool pool db1186: Maintenance * 10:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1186 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95493 and previous config saved to /var/cache/conftool/dbconfig/20260729-102111-cwilliams.json * 10:21 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1186.eqiad.wmnet with reason: Maintenance * 10:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 10:14 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 09:53 XioNoX: reboot cr2-magru - [[phab:T431750|T431750]] * 09:52 XioNoX: drain cr2-magru - [[phab:T431750|T431750]] * 09:48 root@cumin1003: START - Cookbook sre.mysql.pool pool db1221: Maintenance * 09:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zookeeper-test1002.eqiad.wmnet * 09:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1188: Maintenance * 09:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1175: Maintenance * 09:44 btullis@dns1004: END - running authdns-update * 09:42 btullis@dns1004: START - running authdns-update * 09:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1221 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95483 and previous config saved to /var/cache/conftool/dbconfig/20260729-094200-cwilliams.json * 09:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 7 hosts with reason: Maintenance * 09:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1221.eqiad.wmnet with reason: Maintenance * 09:41 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host zookeeper-test1002.eqiad.wmnet * 09:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1199: Maintenance * 09:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1188 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95481 and previous config saved to /var/cache/conftool/dbconfig/20260729-093917-cwilliams.json * 09:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1188.eqiad.wmnet with reason: Maintenance * 09:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1182: Maintenance * 09:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1175 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95479 and previous config saved to /var/cache/conftool/dbconfig/20260729-093842-cwilliams.json * 09:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1175.eqiad.wmnet with reason: Maintenance * 09:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1166: Maintenance * 09:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1169: Maintenance * 09:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1033.eqiad.wmnet,service=s8 * 09:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1033.eqiad.wmnet,service=s5 * 09:33 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1033.eqiad.wmnet,service=s8 * 09:33 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1033.eqiad.wmnet,service=s5 * 09:21 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:21 XioNoX: reboot cr1-magru - [[phab:T431750|T431750]] * 09:17 XioNoX: drain cr1-magru - [[phab:T431750|T431750]] * 09:15 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm1001.wikimedia.org * 09:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply * 09:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply * 09:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 09:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 09:11 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr2-magru,cr2-magru IPv6,cr2-magru.mgmt with reason: router upgrade * 09:11 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm1001.wikimedia.org * 09:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp1005.wikimedia.org * 09:07 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp1005.wikimedia.org * 09:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp2005.wikimedia.org * 09:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp2005.wikimedia.org * 09:00 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr1-magru,cr1-magru IPv6,cr1-magru.mgmt with reason: router upgrade * 09:00 marostegui: Dropping renamed tables [[phab:T426341|T426341]] * 08:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1199: Maintenance * 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1182: Maintenance * 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1166: Maintenance * 08:47 root@cumin1003: START - Cookbook sre.mysql.pool pool db1169: Maintenance * 08:46 ayounsi@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 1:00:00 on cr1-magru,cr1-magru IPv6,cr1-magru.mgmt with reason: router upgrade * 08:45 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2072.codfw.wmnet with OS trixie * 08:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1199 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95464 and previous config saved to /var/cache/conftool/dbconfig/20260729-084534-cwilliams.json * 08:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1199.eqiad.wmnet with reason: Maintenance * 08:45 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 08:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1190: Maintenance * 08:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1182 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95462 and previous config saved to /var/cache/conftool/dbconfig/20260729-084436-cwilliams.json * 08:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1182.eqiad.wmnet with reason: Maintenance * 08:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1156: Maintenance * 08:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1166 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95460 and previous config saved to /var/cache/conftool/dbconfig/20260729-084400-cwilliams.json * 08:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1166.eqiad.wmnet with reason: Maintenance * 08:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1157: Maintenance * 08:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95458 and previous config saved to /var/cache/conftool/dbconfig/20260729-084147-cwilliams.json * 08:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1169.eqiad.wmnet with reason: Maintenance * 08:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1163: Maintenance * 08:30 btullis@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 11 hosts with reason: Replacing the namenodes * 08:23 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2072.codfw.wmnet with reason: host reimage * 08:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1098 hosts * 08:19 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2072.codfw.wmnet with reason: host reimage * 07:58 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2072.codfw.wmnet with OS trixie * 07:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1190: Maintenance * 07:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1156: Maintenance * 07:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1157: Maintenance * 07:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1163: Maintenance * 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1190 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95444 and previous config saved to /var/cache/conftool/dbconfig/20260729-074930-cwilliams.json * 07:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1190.eqiad.wmnet with reason: Maintenance * 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1157 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95443 and previous config saved to /var/cache/conftool/dbconfig/20260729-074914-cwilliams.json * 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1156 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95442 and previous config saved to /var/cache/conftool/dbconfig/20260729-074906-cwilliams.json * 07:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1157.eqiad.wmnet with reason: Maintenance * 07:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 07:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1156.eqiad.wmnet with reason: Maintenance * 07:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1163 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95441 and previous config saved to /var/cache/conftool/dbconfig/20260729-074652-cwilliams.json * 07:46 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1163.eqiad.wmnet with reason: Maintenance * 07:46 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2034.codfw.wmnet * 07:42 ayounsi@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2034.codfw.wmnet * 07:42 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2232.codfw.wmnet with OS trixie * 07:34 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:34 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:33 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:31 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:19 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2232.codfw.wmnet with reason: host reimage * 07:15 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2232.codfw.wmnet with reason: host reimage * 06:58 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db2232.codfw.wmnet with OS trixie * 06:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[2160,2232].codfw.wmnet with reason: Reimage * 06:26 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1164.eqiad.wmnet with OS trixie * 06:05 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1164.eqiad.wmnet with reason: host reimage * 06:01 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1164.eqiad.wmnet with reason: host reimage * 05:47 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1164.eqiad.wmnet with OS trixie * 05:46 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1164.eqiad.wmnet with reason: Reimage == 2026-07-28 == * 22:50 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1138.eqiad.wmnet * 22:50 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1138.eqiad.wmnet * 22:49 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1138.eqiad.wmnet * 22:11 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2014.codfw.wmnet, repooling source-only afterwards * 22:08 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2013.codfw.wmnet, repooling source-only afterwards * 22:03 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 20:58 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] (duration: 08m 19s) * 20:55 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2014.codfw.wmnet, repooling source-only afterwards * 20:55 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2013.codfw.wmnet, repooling source-only afterwards * 20:54 arlolra@deploy1003: arlolra: Continuing with deployment * 20:54 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 14s) * 20:54 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 20:53 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 30s) * 20:53 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 20:52 arlolra@deploy1003: arlolra: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:51 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:50 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] * 20:49 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:43 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 20:34 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] (duration: 06m 54s) * 20:34 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:34 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:31 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:31 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:30 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:30 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:30 arlolra@deploy1003: arlolra: Continuing with deployment * 20:29 arlolra@deploy1003: arlolra: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:27 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] * 20:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2014.codfw.wmnet with OS bookworm * 20:21 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 20:21 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:20 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 20:19 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:19 swfrench-wmf: switched etcd-mirror replication from conf2005 to conf2004 - [[phab:T428495|T428495]] * 20:17 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:17 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:15 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] (duration: 08m 26s) * 20:12 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:11 arlolra@deploy1003: anzx, arlolra: Continuing with deployment * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2013.codfw.wmnet with OS bookworm * 20:09 arlolra@deploy1003: anzx, arlolra: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] * 20:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2216: Maintenance * 19:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2014.codfw.wmnet with reason: host reimage * 19:57 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:54 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:54 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:53 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:52 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2014.codfw.wmnet with reason: host reimage * 19:49 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2013.codfw.wmnet with reason: host reimage * 19:42 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 19:41 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:41 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2013.codfw.wmnet with reason: host reimage * 19:39 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:39 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:39 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-eqiad: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 19:38 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2014 * 19:33 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2014 * 19:29 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2014.codfw.wmnet with OS bookworm * 19:28 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:27 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1006 * 19:26 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2012\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 19:26 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1006 * 19:26 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:26 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 19:25 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2013 * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2013 * 19:21 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2013 * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2013.codfw.wmnet 84.0.192.10.in-addr.arpa 4.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:21 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2013.codfw.wmnet 84.0.192.10.in-addr.arpa 4.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2013 - bking@cumin2003" * 19:21 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2013 - bking@cumin2003" * 19:21 vriley@cumin1003: START - Cookbook sre.dns.netbox * 19:20 root@cumin1003: START - Cookbook sre.mysql.pool pool db2216: Maintenance * 19:13 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2216 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95435 and previous config saved to /var/cache/conftool/dbconfig/20260728-191343-cwilliams.json * 19:13 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2216.codfw.wmnet with reason: Maintenance * 19:13 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2203: Maintenance * 19:06 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1005.eqiad.wmnet with OS trixie * 19:06 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 19:06 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 18:46 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 18:45 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:45 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:43 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:40 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:36 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-eqiad: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 18:35 dancy@deploy1003: Installation of scap version "4.275.0" completed for 3 hosts * 18:33 dancy@deploy1003: Installing scap version "4.275.0" for 3 host(s) * 18:32 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:32 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2097.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:30 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns3003.wikimedia.org [reason: pool for all services after reimaging] * 18:29 sukhe@dns1004: END - running authdns-update * 18:27 sukhe@dns1004: START - running authdns-update * 18:27 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns3003.wikimedia.org,service=authdns-update [reason: pool authdns-update after reimaging] * 18:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db2203: Maintenance * 18:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2203 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95430 and previous config saved to /var/cache/conftool/dbconfig/20260728-181958-cwilliams.json * 18:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2203.codfw.wmnet with reason: Maintenance * 18:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2188: Maintenance * 18:18 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2097.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:17 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-be2098 * 18:17 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host ms-be2098 * 18:17 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-be2097 * 18:16 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host ms-be2097 * 18:15 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:15 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding ms-be2097-8 to codfw - jhancock@cumin2002" * 18:15 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding ms-be2097-8 to codfw - jhancock@cumin2002" * 18:10 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 18:08 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage * 18:05 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns3003.wikimedia.org with OS trixie * 18:03 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage * 17:56 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-codfw: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 17:45 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie * 17:45 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1005.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:41 sukhe@dns1004: END - running authdns-update * 17:39 sukhe@dns1004: START - running authdns-update * 17:36 sukhe@puppetserver1001: conftool action : set/weight=1; selector: cluster=urldownloader,service=squid * 17:36 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1005.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:35 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader,service=squid * 17:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 17:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1005 * 17:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 17:34 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1005 * 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1005] - vriley@cumin1003" * 17:34 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1005] - vriley@cumin1003" * 17:32 root@cumin1003: START - Cookbook sre.mysql.pool pool db2188: Maintenance * 17:29 vriley@cumin1003: START - Cookbook sre.dns.netbox * 17:29 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 17:26 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2188 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95425 and previous config saved to /var/cache/conftool/dbconfig/20260728-172609-cwilliams.json * 17:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2188.codfw.wmnet with reason: Maintenance * 17:25 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2176: Maintenance * 17:19 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1138.eqiad.wmnet with OS trixie * 17:18 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1005 * 17:18 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1005 * 17:18 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:15 vriley@cumin1003: START - Cookbook sre.dns.netbox * 17:13 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns3003.wikimedia.org with reason: host reimage * 17:07 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns3003.wikimedia.org with reason: host reimage * 17:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-eqiad * 17:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1015.eqiad.wmnet * 17:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1015.eqiad.wmnet * 16:59 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1138.eqiad.wmnet with reason: host reimage * 16:55 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-codfw: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 16:54 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1138.eqiad.wmnet with reason: host reimage * 16:53 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1015.eqiad.wmnet * 16:43 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns3003.wikimedia.org with OS trixie * 16:43 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1015.eqiad.wmnet * 16:43 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1014.eqiad.wmnet * 16:43 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1014.eqiad.wmnet * 16:43 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=dns3003.wikimedia.org [reason: depooling for reimage to trixie] * 16:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 290 hosts * 16:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db2176: Maintenance * 16:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1138 * 16:38 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1138 * 16:37 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1138 * 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1138.eqiad.wmnet 193.32.64.10.in-addr.arpa 3.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:37 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1138.eqiad.wmnet 193.32.64.10.in-addr.arpa 3.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1138 - jiji@cumin1003" * 16:37 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1138 - jiji@cumin1003" * 16:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1014.eqiad.wmnet * 16:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2176 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95420 and previous config saved to /var/cache/conftool/dbconfig/20260728-163235-cwilliams.json * 16:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2176.codfw.wmnet with reason: Maintenance * 16:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1014.eqiad.wmnet * 16:32 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1013.eqiad.wmnet * 16:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1013.eqiad.wmnet * 16:32 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2174: Maintenance * 16:28 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2012.codfw.wmnet, repooling source-only afterwards * 16:25 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1013.eqiad.wmnet * 16:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1013.eqiad.wmnet * 16:20 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1012.eqiad.wmnet * 16:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1012.eqiad.wmnet * 16:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1012.eqiad.wmnet * 16:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1012.eqiad.wmnet * 16:03 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1011.eqiad.wmnet * 16:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1011.eqiad.wmnet * 16:00 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2004.codfw.wmnet with OS bookworm * 15:59 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1011.eqiad.wmnet * 15:56 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:55 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 15:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:54 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1011.eqiad.wmnet * 15:54 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1010.eqiad.wmnet * 15:54 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1010.eqiad.wmnet * 15:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:50 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 15:49 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1010.eqiad.wmnet * 15:48 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 15:48 jiji@cumin1003: START - Cookbook sre.dns.netbox * 15:46 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2248: Maintenance * 15:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db2174: Maintenance * 15:44 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1010.eqiad.wmnet * 15:44 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1009.eqiad.wmnet * 15:44 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1009.eqiad.wmnet * 15:42 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1138 * 15:41 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1138.eqiad.wmnet with OS trixie * 15:39 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1009.eqiad.wmnet * 15:39 robh@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on arclamp2001.codfw.wmnet with reason: ram upgrade * 15:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2174 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95413 and previous config saved to /var/cache/conftool/dbconfig/20260728-153844-cwilliams.json * 15:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2174.codfw.wmnet with reason: Maintenance * 15:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2173: Maintenance * 15:37 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 15:35 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 15:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1009.eqiad.wmnet * 15:34 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1008.eqiad.wmnet * 15:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1008.eqiad.wmnet * 15:31 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 15:31 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 15:29 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1008.eqiad.wmnet * 15:27 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 290 hosts * 15:25 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2013 * 15:25 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2195: Maintenance * 15:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1008.eqiad.wmnet * 15:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1007.eqiad.wmnet * 15:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1007.eqiad.wmnet * 15:22 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2004.codfw.wmnet with reason: host reimage * 15:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2013.codfw.wmnet with OS bookworm * 15:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 15:19 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2004.codfw.wmnet with reason: host reimage * 15:19 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 15:17 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1007.eqiad.wmnet * 15:12 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1007.eqiad.wmnet * 15:12 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1006.eqiad.wmnet * 15:12 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1006.eqiad.wmnet * 15:11 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1138.eqiad.wmnet * 15:11 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 15:11 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1138.eqiad.wmnet * 15:11 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1138.eqiad.wmnet * 15:10 brennen@deploy1003: Finished deploy [phabricator/deployment@f8b349f]: deploy phab1004 for [[phab:T433382|T433382]] (duration: 00m 43s) * 15:10 brennen@deploy1003: Started deploy [phabricator/deployment@f8b349f]: deploy phab1004 for [[phab:T433382|T433382]] * 15:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts * 15:09 brennen@deploy1003: Finished deploy [phabricator/deployment@f8b349f]: deploy phab2003 for [[phab:T433382|T433382]] (duration: 00m 55s) * 15:08 brennen@deploy1003: Started deploy [phabricator/deployment@f8b349f]: deploy phab2003 for [[phab:T433382|T433382]] * 15:07 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2012.codfw.wmnet, repooling source-only afterwards * 15:07 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts * 15:06 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1137.eqiad.wmnet * 15:06 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1137.eqiad.wmnet * 15:06 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1137.eqiad.wmnet * 15:05 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1006.eqiad.wmnet * 15:05 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1022\.eqiad\.wmnet,dc=eqiad,cluster=wdqs\-main,service=wdqs\-main * 15:01 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab2003.codfw.wmnet with reason: deployment * 15:01 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1005.eqiad.wmnet with reason: deployment * 15:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1006.eqiad.wmnet * 15:00 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1006.eqiad.wmnet with reason: deployment * 15:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1005.eqiad.wmnet * 15:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1005.eqiad.wmnet * 14:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db2248: Maintenance * 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1004.eqiad.wmnet with reason: deployment * 14:59 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2004.codfw.wmnet with OS bookworm * 14:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1005.eqiad.wmnet * 14:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2248 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95403 and previous config saved to /var/cache/conftool/dbconfig/20260728-145532-cwilliams.json * 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2245-2247].codfw.wmnet with reason: Maintenance * 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2248.codfw.wmnet with reason: Maintenance * 14:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2240: Maintenance * 14:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2173: Maintenance * 14:51 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts * 14:50 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1005.eqiad.wmnet * 14:50 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1004.eqiad.wmnet * 14:50 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1004.eqiad.wmnet * 14:49 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts * 14:45 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2004.codfw.wmnet * 14:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2173 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95399 and previous config saved to /var/cache/conftool/dbconfig/20260728-144453-cwilliams.json * 14:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2173.codfw.wmnet with reason: Maintenance * 14:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2170: Maintenance * 14:44 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1004.eqiad.wmnet * 14:39 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2004.codfw.wmnet * 14:38 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1004.eqiad.wmnet * 14:38 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet * 14:38 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet * 14:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db2195: Maintenance * 14:36 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1022.eqiad.wmnet, repooling source-only afterwards * 14:36 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 14:33 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet * 14:33 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2222: Maintenance * 14:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2195 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95394 and previous config saved to /var/cache/conftool/dbconfig/20260728-143218-cwilliams.json * 14:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2195.codfw.wmnet with reason: Maintenance * 14:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2181: Maintenance * 14:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 14:30 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 14:25 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 14:25 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 14:25 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 14:23 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet * 14:23 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1002.eqiad.wmnet * 14:23 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1002.eqiad.wmnet * 14:23 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 19s) * 14:23 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:18 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1002.eqiad.wmnet * 14:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2012.codfw.wmnet with OS bookworm * 14:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1002.eqiad.wmnet * 14:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1001.eqiad.wmnet * 14:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1001.eqiad.wmnet * 14:11 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 14:11 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 14:08 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1001.eqiad.wmnet * 14:07 elukey@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'. * 14:07 elukey@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'. * 14:06 elukey@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'. * 14:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db2240: Maintenance * 14:06 XioNoX: un-drain cr2-esams - [[phab:T431751|T431751]] * 14:05 elukey@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'. * 14:02 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1001.eqiad.wmnet * 14:02 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-eqiad * 14:01 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 14:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2240 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95384 and previous config saved to /var/cache/conftool/dbconfig/20260728-140011-cwilliams.json * 14:00 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2240.codfw.wmnet with reason: Maintenance * 13:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2237: Maintenance * 13:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db2170: Maintenance * 13:56 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 13:55 XioNoX: reboot cr2-esams - [[phab:T431751|T431751]] * 13:52 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 13:51 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr2-esams,cr2-esams IPv6,cr2-esams.mgmt with reason: router upgrade * 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 13:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2170 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95381 and previous config saved to /var/cache/conftool/dbconfig/20260728-135043-cwilliams.json * 13:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2170.codfw.wmnet with reason: Maintenance * 13:50 XioNoX: drain cr2-esams - [[phab:T431751|T431751]] * 13:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2153: Maintenance * 13:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2012.codfw.wmnet with reason: host reimage * 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2222: Maintenance * 13:45 sukhe: restart pybal on A:lvs-codfw * 13:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2012.codfw.wmnet with reason: host reimage * 13:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db2181: Maintenance * 13:44 btullis@dns1004: END - running authdns-update * 13:42 sukhe: restart pybal on lvs2014 * 13:42 btullis@dns1004: START - running authdns-update * 13:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2222 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95376 and previous config saved to /var/cache/conftool/dbconfig/20260728-133948-cwilliams.json * 13:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2222.codfw.wmnet with reason: Maintenance * 13:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2221: Maintenance * 13:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2181 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95374 and previous config saved to /var/cache/conftool/dbconfig/20260728-133857-cwilliams.json * 13:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2181.codfw.wmnet with reason: Maintenance * 13:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2167: Maintenance * 13:30 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T428495|T428495]] * 13:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 13:29 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 13:29 ayounsi@cumin1003: END (FAIL) - Cookbook sre.dns.admin (exit_code=99) DNS admin: depool esams [reason: router upgrade, [[phab:T431749|T431749]]] * 13:28 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: router upgrade, [[phab:T431749|T431749]]] * 13:27 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1022.eqiad.wmnet, repooling source-only afterwards * 13:27 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo - [[phab:T428495|T428495]] * 13:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2012 * 13:27 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2012 * 13:21 lucaswerkmeister-wmde@deploy1003: mwscript-k8s job started: cleanupTitles bolwiki # [[phab:T429951|T429951]] * 13:21 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] (duration: 07m 19s) * 13:20 swfrench-wmf: authdns-update to direct codfw, eqsin, ulsfo etcd clients to eqiad - [[phab:T428495|T428495]] * 13:18 swfrench@dns1004: END - running authdns-update * 13:17 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, anzx: Continuing with deployment * 13:16 swfrench@dns1004: START - running authdns-update * 13:16 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2012 * 13:16 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2012.codfw.wmnet 57.48.192.10.in-addr.arpa 7.5.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:16 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2012.codfw.wmnet 57.48.192.10.in-addr.arpa 7.5.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:16 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:16 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, anzx: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:14 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 13:14 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] * 13:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db2237: Maintenance * 13:13 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2237: Maintenance * 13:13 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:12 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:12 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback IPV6 for asw1-604 - pt1979@cumin2003" * 13:12 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:12 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback IPV6 for asw1-604 - pt1979@cumin2003" * 13:11 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 20s) * 13:11 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 13:10 esanders@deploy1003: Finished scap sync-world: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] (duration: 08m 11s) * 13:08 pt1979@cumin2003: START - Cookbook sre.dns.netbox * 13:07 root@cumin1003: START - Cookbook sre.mysql.pool pool db2237: Maintenance * 13:06 esanders@deploy1003: esanders: Continuing with deployment * 13:04 esanders@deploy1003: esanders: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2153: Maintenance * 13:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2153: Maintenance * 13:02 esanders@deploy1003: Started scap sync-world: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] * 13:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2237 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95362 and previous config saved to /var/cache/conftool/dbconfig/20260728-130107-cwilliams.json * 13:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2237.codfw.wmnet with reason: Maintenance * 13:00 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2236: Maintenance * 12:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2153: Maintenance * 12:52 root@cumin1003: START - Cookbook sre.mysql.pool pool db2221: Maintenance * 12:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2153 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95358 and previous config saved to /var/cache/conftool/dbconfig/20260728-125214-cwilliams.json * 12:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2153.codfw.wmnet with reason: Maintenance * 12:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2167: Maintenance * 12:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-codfw * 12:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2011.codfw.wmnet * 12:51 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2011.codfw.wmnet * 12:49 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:49 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback for asw1-603 - pt1979@cumin2003" * 12:48 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback for asw1-603 - pt1979@cumin2003" * 12:46 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2011.codfw.wmnet * 12:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2221 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95357 and previous config saved to /var/cache/conftool/dbconfig/20260728-124601-cwilliams.json * 12:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2221.codfw.wmnet with reason: Maintenance * 12:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2218: Maintenance * 12:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2167 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95354 and previous config saved to /var/cache/conftool/dbconfig/20260728-124457-cwilliams.json * 12:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2167.codfw.wmnet with reason: Maintenance * 12:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2166: Maintenance * 12:42 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 12:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2011.codfw.wmnet * 12:41 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2010.codfw.wmnet * 12:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2010.codfw.wmnet * 12:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2010.codfw.wmnet * 12:34 pt1979@cumin2003: START - Cookbook sre.dns.netbox * 12:32 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 12:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2010.codfw.wmnet * 12:31 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2009.codfw.wmnet * 12:31 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2009.codfw.wmnet * 12:27 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2009.codfw.wmnet * 12:22 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2009.codfw.wmnet * 12:22 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2008.codfw.wmnet * 12:21 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2008.codfw.wmnet * 12:16 pt1979@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-604-eqsin * 12:16 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2008.codfw.wmnet * 12:16 pt1979@cumin1003: START - Cookbook sre.network.tls for network device asw1-604-eqsin * 12:14 root@cumin1003: START - Cookbook sre.mysql.pool pool db2236: Maintenance * 12:14 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2236: Maintenance * 12:12 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1137.eqiad.wmnet with OS trixie * 12:11 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2008.codfw.wmnet * 12:11 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2007.codfw.wmnet * 12:11 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2007.codfw.wmnet * 12:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db2236: Maintenance * 12:06 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2007.codfw.wmnet * 12:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2236 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95348 and previous config saved to /var/cache/conftool/dbconfig/20260728-120253-cwilliams.json * 12:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2236.codfw.wmnet with reason: Maintenance * 12:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2007.codfw.wmnet * 12:01 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 12:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 11:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2218: Maintenance * 11:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db2166: Maintenance * 11:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2219: Maintenance * 11:56 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2006.codfw.wmnet * 11:52 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1137.eqiad.wmnet with reason: host reimage * 11:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2218 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95344 and previous config saved to /var/cache/conftool/dbconfig/20260728-115155-cwilliams.json * 11:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2218.codfw.wmnet with reason: Maintenance * 11:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2208: Maintenance * 11:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2166 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95342 and previous config saved to /var/cache/conftool/dbconfig/20260728-115119-cwilliams.json * 11:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2166.codfw.wmnet with reason: Maintenance * 11:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2164: Maintenance * 11:47 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1137.eqiad.wmnet with reason: host reimage * 11:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2006.codfw.wmnet * 11:45 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2005.codfw.wmnet * 11:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2005.codfw.wmnet * 11:40 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2005.codfw.wmnet * 11:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2005.codfw.wmnet * 11:35 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2004.codfw.wmnet * 11:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2004.codfw.wmnet * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1137 * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1137 * 11:30 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1137 * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1137.eqiad.wmnet 192.32.64.10.in-addr.arpa 2.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:30 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1137.eqiad.wmnet 192.32.64.10.in-addr.arpa 2.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1137 - jiji@cumin1003" * 11:25 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2004.codfw.wmnet * 11:19 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2004.codfw.wmnet * 11:19 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2003.codfw.wmnet * 11:19 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2003.codfw.wmnet * 11:14 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2003.codfw.wmnet * 11:11 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2219: Maintenance * 11:10 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2219: Maintenance * 11:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2219: Maintenance * 11:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2003.codfw.wmnet * 11:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2164: Maintenance * 11:03 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2002.codfw.wmnet * 11:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2002.codfw.wmnet * 11:03 root@cumin1003: START - Cookbook sre.mysql.pool pool db2208: Maintenance * 10:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2164 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95332 and previous config saved to /var/cache/conftool/dbconfig/20260728-105749-cwilliams.json * 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2164.codfw.wmnet with reason: Maintenance * 10:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2208 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95331 and previous config saved to /var/cache/conftool/dbconfig/20260728-105711-cwilliams.json * 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2208.codfw.wmnet with reason: Maintenance * 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2219 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95330 and previous config saved to /var/cache/conftool/dbconfig/20260728-105652-cwilliams.json * 10:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2219.codfw.wmnet with reason: Maintenance * 10:53 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1137 - jiji@cumin1003" * 10:52 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2002.codfw.wmnet * 10:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2002.codfw.wmnet * 10:47 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2001.codfw.wmnet * 10:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2001.codfw.wmnet * 10:39 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2001.codfw.wmnet * 10:35 jiji@cumin1003: START - Cookbook sre.dns.netbox * 10:34 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] (duration: 09m 31s) * 10:34 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1137 * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2001.codfw.wmnet * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-codfw * 10:34 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1137.eqiad.wmnet with OS trixie * 10:28 jforrester@deploy1003: jforrester: Continuing with deployment * 10:27 jforrester@deploy1003: jforrester: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:25 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] * 10:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 10:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 10:21 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 10:20 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1137.eqiad.wmnet * 10:20 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 10:20 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1137.eqiad.wmnet * 10:20 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1137.eqiad.wmnet * 10:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2163: Maintenance * 09:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-staging-worker * 09:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2003.codfw.wmnet * 09:37 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2003.codfw.wmnet * 09:32 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 09:31 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2003.codfw.wmnet * 09:30 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 09:30 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2163: Maintenance * 09:30 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 09:30 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 09:30 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:22 klausman@cumin1003: END (ERROR) - Cookbook sre.ganeti.reboot-vm (exit_code=97) for VM ml-serve-ctrl2001.codfw.wmnet * 09:22 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2001.codfw.wmnet * 09:22 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d8-eqiad * 09:22 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d8-eqiad * 09:21 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2003.codfw.wmnet * 09:20 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2002.codfw.wmnet * 09:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2002.codfw.wmnet * 09:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f2-codfw * 09:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f2-codfw * 09:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e4-codfw * 09:18 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2163: Maintenance * 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e4-codfw * 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-codfw * 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-codfw * 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e5-codfw * 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e5-codfw * 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f4-codfw * 09:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2210: Maintenance * 09:16 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f4-codfw * 09:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2182: Maintenance * 09:14 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2002.codfw.wmnet * 09:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db2163: Maintenance * 09:11 XioNoX: rebooting cr2-drmrs - [[phab:T431749|T431749]] * 09:10 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr2-drmrs,cr2-drmrs IPv6,cr2-drmrs.mgmt with reason: router upgrade * 09:06 XioNoX: draining cr2-drmrs - [[phab:T431749|T431749]] * 09:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2163 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95320 and previous config saved to /var/cache/conftool/dbconfig/20260728-090638-cwilliams.json * 09:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2163.codfw.wmnet with reason: Maintenance * 09:06 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2161: Maintenance * 09:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2002.codfw.wmnet * 09:04 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2001.codfw.wmnet * 09:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2001.codfw.wmnet * 08:57 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2001.codfw.wmnet * 08:48 XioNoX: un-drain cr1-drmrs - [[phab:T431749|T431749]] * 08:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2001.codfw.wmnet * 08:47 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-staging-worker * 08:42 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:35 XioNoX: rebooting cr1-drmrs - [[phab:T431749|T431749]] * 08:33 XioNoX: draining cr1-drmrs - [[phab:T431749|T431749]] * 08:31 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2210: Maintenance * 08:29 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2182: Maintenance * 08:21 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2182: Maintenance * 08:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db2161: Maintenance * 08:16 root@cumin1003: START - Cookbook sre.mysql.pool pool db2182: Maintenance * 08:12 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2210: Maintenance * 08:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2161 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95309 and previous config saved to /var/cache/conftool/dbconfig/20260728-081044-cwilliams.json * 08:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2161.codfw.wmnet with reason: Maintenance * 08:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2154: Maintenance * 08:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2182 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95307 and previous config saved to /var/cache/conftool/dbconfig/20260728-080947-cwilliams.json * 08:09 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2182.codfw.wmnet with reason: Maintenance * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2168: Maintenance * 08:06 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr1-drmrs,cr1-drmrs IPv6,cr1-drmrs.mgmt with reason: router upgrade * 08:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db2210: Maintenance * 08:05 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 08:05 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 08:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2210 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95305 and previous config saved to /var/cache/conftool/dbconfig/20260728-080008-cwilliams.json * 08:00 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2210.codfw.wmnet with reason: Maintenance * 07:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2206: Maintenance * 07:50 gkyziridis@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 07:50 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 07:22 root@cumin1003: START - Cookbook sre.mysql.pool pool db2154: Maintenance * 07:22 root@cumin1003: START - Cookbook sre.mysql.pool pool db2168: Maintenance * 07:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2154 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95295 and previous config saved to /var/cache/conftool/dbconfig/20260728-071640-cwilliams.json * 07:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2154.codfw.wmnet with reason: Maintenance * 07:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95294 and previous config saved to /var/cache/conftool/dbconfig/20260728-071604-cwilliams.json * 07:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2168.codfw.wmnet with reason: Maintenance * 07:08 root@cumin1003: START - Cookbook sre.mysql.pool pool db2206: Maintenance * 07:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2206 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95292 and previous config saved to /var/cache/conftool/dbconfig/20260728-070219-cwilliams.json * 07:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2206.codfw.wmnet with reason: Maintenance * 06:44 marostegui: Failover m5 from db1164 to db1228 - [[phab:T432967|T432967]] * 06:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2235].codfw.wmnet,db[1164,1217,1228].eqiad.wmnet with reason: m5 master switch [[phab:T432967|T432967]] * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.10 (duration: 02m 34s) * 03:39 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] (duration: 36m 06s) * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 02:57 dzahn@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1004.eqiad.wmnet with OS trixie * 02:57 dzahn@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - dzahn@cumin1003" * 02:55 dzahn@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - dzahn@cumin1003" * 02:37 dzahn@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1004.eqiad.wmnet with reason: host reimage * 02:31 dzahn@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1004.eqiad.wmnet with reason: host reimage * 02:16 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie * 02:15 dzahn@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host zuul1004.eqiad.wmnet with OS trixie * 01:43 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie * 01:43 dzahn@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1004.eqiad.wmnet with OS trixie * 01:25 pt1979@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-603-eqsin * 01:24 pt1979@cumin1003: START - Cookbook sre.network.tls for network device asw1-603-eqsin * 01:12 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 01:12 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt for new switches in eqsin - pt1979@cumin2003" * 01:12 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt for new switches in eqsin - pt1979@cumin2003" * 01:08 pt1979@cumin2003: START - Cookbook sre.dns.netbox * 00:48 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 00:47 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 00:47 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 00:47 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 00:26 mutante: attempting reimage with trixie on zuul1004 re-purposed physical hardware - dcops reported install issue - host was in busybox shell ([[phab:T427353|T427353]]) * 00:24 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie == 2026-07-27 == * 23:50 Amir1: mass deleting vp8 transcodes * 23:28 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:27 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1004.eqiad.wmnet with OS bullseye * 23:26 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 23:25 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:25 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 22:39 maryum: Deploy security fix for [[phab:T432877|T432877]] * 22:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1022.eqiad.wmnet with OS bookworm * 22:37 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS bullseye * 22:32 sbassett: Deployed security fix for [[phab:T432789|T432789]] * 22:22 sbassett: Deployed security patch for [[phab:T431819|T431819]] * 22:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1022.eqiad.wmnet with reason: host reimage * 22:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1022.eqiad.wmnet with reason: host reimage * 22:01 RScout-WMF: Deployed security fix for [[phab:T431819|T431819]] * 22:00 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2012 * 21:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2012.codfw.wmnet with OS bookworm * 21:55 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2011\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 21:45 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1022 * 21:45 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1022 * 21:44 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1022 * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1022.eqiad.wmnet 239.48.64.10.in-addr.arpa 9.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:44 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1022.eqiad.wmnet 239.48.64.10.in-addr.arpa 9.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:41 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:41 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 21:34 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS bookworm * 21:31 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:22 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:21 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:19 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:17 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1004 * 21:16 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1004 * 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1004] - vriley@cumin1003" * 21:15 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1004] - vriley@cumin1003" * 21:11 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:10 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2011.codfw.wmnet, repooling source-only afterwards * 21:05 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:01 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1022 * 20:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1022.eqiad.wmnet with OS bookworm * 20:53 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1021.eqiad.wmnet, repooling source-only afterwards * 20:51 mutante: zuul1001 - re-enabled puppet - revert "cherry-picked" gerrit:1314120 - [[phab:T431003|T431003]] * 20:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Maintenance * 20:15 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] (duration: 08m 03s) * 20:11 sbisson@deploy1003: sbisson: Continuing with deployment * 20:09 sbisson@deploy1003: sbisson: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] * 19:47 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 46s) * 19:47 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Maintenance * 19:27 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] (duration: 12m 26s) * 19:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2228 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95285 and previous config saved to /var/cache/conftool/dbconfig/20260727-192711-cwilliams.json * 19:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2228.codfw.wmnet with reason: Maintenance * 19:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2223: Maintenance * 19:23 krinkle@deploy1003: krinkle: Continuing with deployment * 19:16 krinkle@deploy1003: krinkle: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:15 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] * 19:12 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2238: Maintenance * 18:58 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 18:57 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 18:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2227: Maintenance * 18:57 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-experimental: apply * 18:55 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-experimental: apply * 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1021.eqiad.wmnet with OS bookworm * 18:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2011.codfw.wmnet with OS bookworm * 18:40 root@cumin1003: START - Cookbook sre.mysql.pool pool db2223: Maintenance * 18:39 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] (duration: 07m 05s) * 18:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2223 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95275 and previous config saved to /var/cache/conftool/dbconfig/20260727-183500-cwilliams.json * 18:34 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2223.codfw.wmnet with reason: Maintenance * 18:34 musikanimal@deploy1003: musikanimal: Continuing with deployment * 18:34 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2213: Maintenance * 18:33 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:32 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] * 18:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db2238: Maintenance * 18:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2011.codfw.wmnet with reason: host reimage * 18:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2238 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95271 and previous config saved to /var/cache/conftool/dbconfig/20260727-181944-cwilliams.json * 18:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2238.codfw.wmnet with reason: Maintenance * 18:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2226: Maintenance * 18:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1021.eqiad.wmnet with reason: host reimage * 18:14 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2011.codfw.wmnet with reason: host reimage * 18:12 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1021.eqiad.wmnet with reason: host reimage * 18:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db2227: Maintenance * 18:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2227 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95265 and previous config saved to /var/cache/conftool/dbconfig/20260727-180256-cwilliams.json * 18:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2227.codfw.wmnet with reason: Maintenance * 18:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2194: Maintenance * 17:57 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2011 * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2011 * 17:56 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2011 * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2011.codfw.wmnet 37.32.192.10.in-addr.arpa 7.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:56 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2011.codfw.wmnet 37.32.192.10.in-addr.arpa 7.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2011 - bking@cumin2003" * 17:56 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2011 - bking@cumin2003" * 17:52 bking@cumin2003: START - Cookbook sre.dns.netbox * 17:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2011 * 17:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1021 * 17:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1021 * 17:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2011.codfw.wmnet with OS bookworm * 17:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1021.eqiad.wmnet with OS bookworm * 17:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Maintenance * 17:38 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2010\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 17:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2213 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95260 and previous config saved to /var/cache/conftool/dbconfig/20260727-173740-cwilliams.json * 17:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2213.codfw.wmnet with reason: Maintenance * 17:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2211: Maintenance * 17:36 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1020\.eqiad\.wmnet,dc=eqiad,cluster=wdqs\-main,service=wdqs\-main * 17:32 root@cumin1003: START - Cookbook sre.mysql.pool pool db2226: Maintenance * 17:31 taavi@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] (duration: 06m 33s) * 17:27 taavi@deploy1003: taavi: Continuing with deployment * 17:27 taavi@deploy1003: taavi: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:26 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2226 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95256 and previous config saved to /var/cache/conftool/dbconfig/20260727-172636-cwilliams.json * 17:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2226.codfw.wmnet with reason: Maintenance * 17:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2225: Maintenance * 17:25 taavi@deploy1003: Started scap sync-world: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] * 17:13 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 17:11 root@cumin1003: START - Cookbook sre.mysql.pool pool db2194: Maintenance * 17:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2194 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95248 and previous config saved to /var/cache/conftool/dbconfig/20260727-170453-cwilliams.json * 17:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2194.codfw.wmnet with reason: Maintenance * 17:04 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2190: Maintenance * 16:52 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 16:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2211: Maintenance * 16:40 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95242 and previous config saved to /var/cache/conftool/dbconfig/20260727-164015-cwilliams.json * 16:40 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2211.codfw.wmnet with reason: Maintenance * 16:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2178: Maintenance * 16:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db2225: Maintenance * 16:39 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2172: Maintenance * 16:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 16:38 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 16:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2225 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95238 and previous config saved to /var/cache/conftool/dbconfig/20260727-163307-cwilliams.json * 16:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2225.codfw.wmnet with reason: Maintenance * 16:32 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2189: Maintenance * 16:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db2190: Maintenance * 16:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2190 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95230 and previous config saved to /var/cache/conftool/dbconfig/20260727-160602-cwilliams.json * 16:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2190.codfw.wmnet with reason: Maintenance * 15:53 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2177: Maintenance * 15:53 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2172: Maintenance * 15:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2178: Maintenance * 15:51 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2172: Maintenance * 15:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2178 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95224 and previous config saved to /var/cache/conftool/dbconfig/20260727-154559-cwilliams.json * 15:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2172: Maintenance * 15:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2178.codfw.wmnet with reason: Maintenance * 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2171: Maintenance * 15:44 root@cumin1003: START - Cookbook sre.mysql.pool pool db2189: Maintenance * 15:43 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:41 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2172 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95222 and previous config saved to /var/cache/conftool/dbconfig/20260727-153927-cwilliams.json * 15:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2172.codfw.wmnet with reason: Maintenance * 15:38 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2189 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95220 and previous config saved to /var/cache/conftool/dbconfig/20260727-153833-cwilliams.json * 15:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2189.codfw.wmnet with reason: Maintenance * 15:34 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:32 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 15:32 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 15:31 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:29 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:26 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 15:22 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] (duration: 07m 00s) * 15:21 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2155: Maintenance * 15:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2175: Maintenance * 15:18 zabe@deploy1003: zabe: Continuing with deployment * 15:17 zabe@deploy1003: zabe: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:15 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 15:15 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:15 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] * 15:15 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 06s) * 15:15 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:12 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2010.codfw.wmnet with OS bookworm * 15:08 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2177: Maintenance * 15:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2177: Maintenance * 14:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2177: Maintenance * 14:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2171: Maintenance * 14:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2171 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95209 and previous config saved to /var/cache/conftool/dbconfig/20260727-145236-cwilliams.json * 14:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2171.codfw.wmnet with reason: Maintenance * 14:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2177 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95208 and previous config saved to /var/cache/conftool/dbconfig/20260727-145206-cwilliams.json * 14:52 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2157: Maintenance * 14:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2177.codfw.wmnet with reason: Maintenance * 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1020.eqiad.wmnet with OS bookworm * 14:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2156: Maintenance * 14:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2010.codfw.wmnet with reason: host reimage * 14:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2010.codfw.wmnet with reason: host reimage * 14:41 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[1031,2024]*: Upgrade Cassandra to 5.0.8 (canary) - eevans@cumin1003 * 14:34 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2155: Maintenance * 14:33 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2175: Maintenance * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2010 * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2010 * 14:24 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2010 * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2010.codfw.wmnet 94.16.192.10.in-addr.arpa 4.9.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2010.codfw.wmnet 94.16.192.10.in-addr.arpa 4.9.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2010 - bking@cumin2003" * 14:24 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2010 - bking@cumin2003" * 14:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1020.eqiad.wmnet with reason: host reimage * 14:23 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[1031,2024]*: Upgrade Cassandra to 5.0.8 (canary) - eevans@cumin1003 * 14:20 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 14:20 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 14:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1020.eqiad.wmnet with reason: host reimage * 14:17 sukhe: sudo gnt-instance reboot urldownloader1005.wikimedia.org * 14:16 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:15 jelto@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:08 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2155: Maintenance * 14:05 root@cumin1003: START - Cookbook sre.mysql.pool pool db2157: Maintenance * 14:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2175: Maintenance * 14:03 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 11 hosts * 14:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db2155: Maintenance * 14:01 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 11 hosts * 14:01 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1136.eqiad.wmnet * 14:01 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1136.eqiad.wmnet * 14:01 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1136.eqiad.wmnet * 14:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db2156: Maintenance * 13:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2157 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95194 and previous config saved to /var/cache/conftool/dbconfig/20260727-135943-cwilliams.json * 13:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2157.codfw.wmnet with reason: Maintenance * 13:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db2175: Maintenance * 13:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 13:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 13:57 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2010 * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2155 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95193 and previous config saved to /var/cache/conftool/dbconfig/20260727-135613-cwilliams.json * 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2155.codfw.wmnet with reason: Maintenance * 13:55 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1020 * 13:55 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1020 * 13:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2156 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95192 and previous config saved to /var/cache/conftool/dbconfig/20260727-135413-cwilliams.json * 13:54 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2156.codfw.wmnet with reason: Maintenance * 13:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2175 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95191 and previous config saved to /var/cache/conftool/dbconfig/20260727-135300-cwilliams.json * 13:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2010.codfw.wmnet with OS bookworm * 13:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2175.codfw.wmnet with reason: Maintenance * 13:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1020.eqiad.wmnet with OS bookworm * 13:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 34 hosts * 13:46 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 34 hosts * 13:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1201: Maintenance * 13:27 Lucas_WMDE: UTC afternoon backport+config window doen * 13:18 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] (duration: 11m 57s) * 13:14 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, sihe: Continuing with deployment * 13:08 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, sihe: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:07 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool ulsfo [reason: router upgrade finished, [[phab:T431752|T431752]]] * 13:07 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool ulsfo [reason: router upgrade finished, [[phab:T431752|T431752]]] * 13:06 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] * 13:03 XioNoX: repool cr4-ulsfo - [[phab:T431752|T431752]] * 12:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db1201: Maintenance * 12:48 gkyziridis@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1201 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95186 and previous config saved to /var/cache/conftool/dbconfig/20260727-124404-cwilliams.json * 12:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1201.eqiad.wmnet with reason: Maintenance * 12:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1187: Maintenance * 12:30 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader.eqiad.wikimedia.org on all recursors * 12:30 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader.eqiad.wikimedia.org on all recursors * 12:30 sukhe@dns1004: END - running authdns-update * 12:30 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] (duration: 09m 32s) * 12:28 sukhe@dns1004: START - running authdns-update * 12:25 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 12:22 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:20 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] * 12:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts * 12:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts * 12:13 XioNoX: rebooting cr4-ulsfo for upgrade - [[phab:T431752|T431752]] * 12:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: es1038 repool * 12:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 38 hosts * 12:08 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 38 hosts * 11:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1187: Maintenance * 11:53 urbanecm@deploy1003: mwscript-k8s job started: foreachwikiindblist growthexperiments GrowthExperiments:cleanMentorList # [[phab:T431804|T431804]] * 11:50 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr4-ulsfo,cr4-ulsfo IPv6,cr4-ulsfo.mgmt with reason: router upgrade * 11:50 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] (duration: 11m 07s) * 11:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1187 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95178 and previous config saved to /var/cache/conftool/dbconfig/20260727-114844-cwilliams.json * 11:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1187.eqiad.wmnet with reason: Maintenance * 11:43 urbanecm@deploy1003: urbanecm: Continuing with deployment * 11:42 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:39 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] * 11:37 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 11:36 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 11:36 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 11:35 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 11:29 XioNoX: start draining cr4-ulsfo - [[phab:T431752|T431752]] * 11:29 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 11:29 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 11:28 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1035: testing * 11:28 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1035: testing * 11:27 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1035: testing * 11:27 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1035: testing * 11:26 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool ulsfo [reason: router upgrade, [[phab:T431752|T431752]]] * 11:26 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1038: es1038 repool * 11:26 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool ulsfo [reason: router upgrade, [[phab:T431752|T431752]]] * 11:26 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1038: testing * 11:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1264: Maintenance * 11:24 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1038: testing * 11:23 marostegui@cumin1003: dbctl commit (dc=all): 'Repool es1050 as master', diff saved to https://phabricator.wikimedia.org/P95170 and previous config saved to /var/cache/conftool/dbconfig/20260727-112326-marostegui.json * 11:23 marostegui@cumin1003: dbctl commit (dc=all): 'Repool es1050', diff saved to https://phabricator.wikimedia.org/P95169 and previous config saved to /var/cache/conftool/dbconfig/20260727-112302-marostegui.json * 11:22 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1050: testing * 11:22 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1050: testing * 11:20 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 11:18 blake@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 11:18 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 11:12 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 11:11 blake@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 11:09 blake@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 11:09 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 11:09 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 11:08 blake@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 11:05 blake@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 11:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 11:02 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 10:50 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 10:43 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:39 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply * 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1264: Maintenance * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply * 10:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply * 10:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 10:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 10:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 10:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1264 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95164 and previous config saved to /var/cache/conftool/dbconfig/20260727-103204-cwilliams.json * 10:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1264.eqiad.wmnet with reason: Maintenance * 10:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 10:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 10:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 10:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 10:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 10:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 10:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 10:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1237: Maintenance * 10:04 elukey: restart burrow main-eqiad on kafkamon2003 to clear some errors on kafka-main1008 * 09:58 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1136.eqiad.wmnet with OS trixie * 09:39 elukey: restart burrow-main-eqiad.service on kafkamon1003 to see if a recurrent kafka error on kafka-main1008 goes away * 09:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1237: Maintenance * 09:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1136.eqiad.wmnet with reason: host reimage * 09:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1237 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95159 and previous config saved to /var/cache/conftool/dbconfig/20260727-093328-cwilliams.json * 09:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1237.eqiad.wmnet with reason: Maintenance * 09:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1136.eqiad.wmnet with reason: host reimage * 09:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1203: Maintenance * 09:17 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1136 * 09:17 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1136 * 09:04 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1136 * 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1136.eqiad.wmnet 191.32.64.10.in-addr.arpa 1.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:04 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1136.eqiad.wmnet 191.32.64.10.in-addr.arpa 1.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1136 - jiji@cumin1003" * 09:04 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1136 - jiji@cumin1003" * 08:52 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 08:52 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 08:52 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 08:51 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 08:50 jiji@cumin1003: START - Cookbook sre.dns.netbox * 08:47 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1136 * 08:46 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1136.eqiad.wmnet with OS trixie * 08:44 marostegui: Rename tables on s3 [[phab:T425066|T425066]] * 08:43 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1136.eqiad.wmnet * 08:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db1203: Maintenance * 08:43 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1136.eqiad.wmnet * 08:43 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1136.eqiad.wmnet * 08:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1203 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95154 and previous config saved to /var/cache/conftool/dbconfig/20260727-083703-cwilliams.json * 08:36 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1203.eqiad.wmnet with reason: Maintenance * 08:16 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1179: Maintenance * 07:44 phuedx: UTC morning backport window done * 07:37 phuedx@deploy1003: Finished scap sync-world: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] (duration: 32m 33s) * 07:28 root@cumin1003: START - Cookbook sre.mysql.pool pool db1179: Maintenance * 07:26 marostegui: Rename tables on s3 [[phab:T426341|T426341]] * 07:25 phuedx@deploy1003: phuedx: Continuing with deployment * 07:22 marostegui: Drop tables in akwiki nawiki pihwiki - growthexperiments_* [[phab:T428885|T428885]] * 07:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95149 and previous config saved to /var/cache/conftool/dbconfig/20260727-072234-cwilliams.json * 07:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1179.eqiad.wmnet with reason: Maintenance * 07:20 phuedx@deploy1003: phuedx: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:16 ryankemper: [[phab:T430880|T430880]] [WDQS] Reimaged `wdqs1018` and `wdqs1019` to Bookworm, restored data using test-cookbook change {{Gerrit|1317128}}, and repooled both; 25/36 hosts complete * 07:04 phuedx@deploy1003: Started scap sync-world: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] * 06:57 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1019.eqiad.wmnet * 06:56 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1018.eqiad.wmnet * 06:40 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1020.eqiad.wmnet with reason: Cloning * 06:35 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db1228.eqiad.wmnet with reason: Rebooting * 06:29 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:29 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:25 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:25 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:25 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1019.eqiad.wmnet, repooling source-only afterwards * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1018.eqiad.wmnet, repooling source-only afterwards * 04:51 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1019.eqiad.wmnet, repooling source-only afterwards * 04:51 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1018.eqiad.wmnet, repooling source-only afterwards * 04:48 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s) * 04:48 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 04:48 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s) * 04:48 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 36s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-26 == * 14:59 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:59 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:59 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:59 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1019.eqiad.wmnet with OS bookworm * 01:05 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1018.eqiad.wmnet with OS bookworm * 00:43 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1019.eqiad.wmnet with reason: host reimage * 00:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1018.eqiad.wmnet with reason: host reimage * 00:34 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1019.eqiad.wmnet with reason: host reimage * 00:33 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1018.eqiad.wmnet with reason: host reimage * 00:16 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 00:16 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 00:15 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 00:15 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1019 * 00:11 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1019 * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1018 * 00:11 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1018 * 00:08 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1019.eqiad.wmnet with OS bookworm * 00:08 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1018.eqiad.wmnet with OS bookworm == 2026-07-25 == * 22:06 ryankemper: [[phab:T430880|T430880]] [WDQS] Repooled `wdqs1017` and `wdqs2024` after reimaging to bookworm, scap deploying, and data xfering * 22:04 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2024.codfw.wmnet * 22:03 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1017.eqiad.wmnet * 21:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1017.eqiad.wmnet, repooling source-only afterwards * 21:06 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2024.codfw.wmnet, repooling source-only afterwards * 20:52 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:52 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:52 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:52 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 20:18 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1017.eqiad.wmnet, repooling source-only afterwards * 20:18 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2024.codfw.wmnet, repooling source-only afterwards * 20:15 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:15 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:15 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:15 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 19:57 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s) * 19:57 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 19:57 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 07s) * 19:57 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 19:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2024.codfw.wmnet with OS bookworm * 19:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1017.eqiad.wmnet with OS bookworm * 19:02 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2024.codfw.wmnet with reason: host reimage * 18:58 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1017.eqiad.wmnet with reason: host reimage * 18:53 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2024.codfw.wmnet with reason: host reimage * 18:52 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1017.eqiad.wmnet with reason: host reimage * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2024 * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2024 * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1017 * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1017 * 18:27 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2024 * 18:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2024.codfw.wmnet 58.16.192.10.in-addr.arpa 8.5.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:26 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2024.codfw.wmnet 58.16.192.10.in-addr.arpa 8.5.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:24 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1017 * 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1017.eqiad.wmnet 238.48.64.10.in-addr.arpa 8.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:24 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1017.eqiad.wmnet 238.48.64.10.in-addr.arpa 8.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1017 - ryankemper@cumin2003" * 18:24 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1017 - ryankemper@cumin2003" * 18:23 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 18:18 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 18:17 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1017 * 18:17 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2024 * 18:14 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1017.eqiad.wmnet with OS bookworm * 18:14 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2024.codfw.wmnet with OS bookworm * 18:05 ryankemper: [WDQS] [[phab:T430880|T430880]] Reimaged `wdqs1016` and `wdqs2023` to Bookworm with `--move-vlan`, restored main and scholarly data, validated postflights, and repooled both hosts. Confirmed PyBal rebuilt both backends with their new addresses * 17:45 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2023.codfw.wmnet * 17:43 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1016.eqiad.wmnet * 06:35 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1016.eqiad.wmnet, repooling source-only afterwards * 06:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2023.codfw.wmnet, repooling source-only afterwards * 05:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2023.codfw.wmnet, repooling source-only afterwards * 05:19 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1016.eqiad.wmnet, repooling source-only afterwards * 05:07 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 07s) * 05:07 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 05:06 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 06s) * 05:06 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 03:27 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2023.codfw.wmnet with OS bookworm * 02:59 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2023.codfw.wmnet with reason: host reimage * 02:56 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2023.codfw.wmnet with reason: host reimage * 02:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2023 * 02:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2023 * 02:30 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2023 * 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2023.codfw.wmnet 35.0.192.10.in-addr.arpa 5.3.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 02:30 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2023.codfw.wmnet 35.0.192.10.in-addr.arpa 5.3.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2023 - ryankemper@cumin2003" * 02:30 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2023 - ryankemper@cumin2003" * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 26s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:15 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1016.eqiad.wmnet with OS bookworm * 00:49 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1016.eqiad.wmnet with reason: host reimage * 00:43 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1016.eqiad.wmnet with reason: host reimage * 00:31 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 00:27 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1016 * 00:27 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1016 * 00:27 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2023 * 00:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1016.eqiad.wmnet with OS bookworm * 00:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2023.codfw.wmnet with OS bookworm * 00:11 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs1014.eqiad.wmnet and wdqs2008.codfw.wmnet after Bookworm reimage, transfer, and postflight; wdqs2008 is serving, while wdqs1014 will remain outside of service until a pybal restart next monday == 2026-07-24 == * 23:54 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1014.eqiad.wmnet * 23:54 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2008.codfw.wmnet * 23:43 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2010.codfw.wmnet with OS trixie * 23:08 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 23:03 jhathaway@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 22:33 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 22:13 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 22:13 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 22:13 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 22:13 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:00 jhathaway@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 21:53 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 21:53 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie * 21:51 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 21:47 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie * 21:43 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 21:39 jhathaway@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 21:38 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 17:21 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1135.eqiad.wmnet * 17:21 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1135.eqiad.wmnet * 17:21 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1135.eqiad.wmnet * 16:34 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 16:34 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 16:34 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 16:34 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 16:33 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 16:33 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 16:28 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:28 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:28 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:28 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2008.codfw.wmnet, repooling source-only afterwards * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1014.eqiad.wmnet, repooling source-only afterwards * 15:56 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1135.eqiad.wmnet with OS trixie * 15:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 40 hosts * 15:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 40 hosts * 15:37 topranks: upgrade SR-Linux OS on lswtest-d8-eqiad * 15:36 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1135.eqiad.wmnet with reason: host reimage * 15:33 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 6 hosts with reason: upgrade lswtest-d8-eqiad * 15:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1135.eqiad.wmnet with reason: host reimage * 15:30 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc-gp2006.codfw.wmnet with OS bookworm * 15:15 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1135 * 15:15 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1135 * 15:13 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc-gp2006.codfw.wmnet with reason: host reimage * 15:08 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc-gp2006.codfw.wmnet with reason: host reimage * 14:49 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm * 14:48 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host mc-gp2006.codfw.wmnet with OS bookworm * 14:34 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] (duration: 41m 12s) * 14:32 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1135 * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1135.eqiad.wmnet 177.32.64.10.in-addr.arpa 7.7.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:32 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1135.eqiad.wmnet 177.32.64.10.in-addr.arpa 7.7.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1135 - jiji@cumin1003" * 14:32 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1135 - jiji@cumin1003" * 14:29 krinkle@deploy1003: krinkle: Continuing with deployment * 14:29 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm * 14:27 jiji@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host mc-gp2006.codfw.wmnet with OS bookworm * 14:26 jiji@cumin1003: START - Cookbook sre.dns.netbox * 14:15 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1135 * 14:14 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1135.eqiad.wmnet with OS trixie * 14:14 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1135.eqiad.wmnet * 14:13 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1135.eqiad.wmnet * 14:13 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1135.eqiad.wmnet * 14:10 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1072.eqiad.wmnet * 14:10 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1072.eqiad.wmnet * 14:10 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1072.eqiad.wmnet * 14:10 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1072.eqiad.wmnet * 14:09 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1071.eqiad.wmnet * 14:09 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1071.eqiad.wmnet * 14:09 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1071.eqiad.wmnet * 14:09 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1071.eqiad.wmnet * 13:58 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 13:58 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:58 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:57 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:55 krinkle@deploy1003: krinkle: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:53 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] * 13:45 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:45 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push new IPs for mc-gp2006 - cmooney@cumin1003" * 13:45 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push new IPs for mc-gp2006 - cmooney@cumin1003" * 13:44 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) mc-gp2006.codfw.wmnet on all recursors * 13:44 cmooney@cumin1003: START - Cookbook sre.dns.wipe-cache mc-gp2006.codfw.wmnet on all recursors * 13:42 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm * 13:41 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:30 papaul: reboot mr1-eqsin for maintenance * 13:24 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb[1029-1031].eqiad.wmnet * 13:10 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb[1029-1031].eqiad.wmnet * 11:33 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 7 hosts * 11:11 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 7 hosts * 10:56 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 7 hosts * 10:47 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 7 hosts * 10:44 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:44 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:41 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:41 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:35 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 8 hosts * 10:34 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:33 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:32 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:32 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:31 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:31 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:30 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 8 hosts * 10:24 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie * 10:19 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:18 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 16 hosts * 10:17 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2001.codfw.wmnet * 10:13 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2001.codfw.wmnet * 10:12 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2001.codfw.wmnet * 10:02 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2001.codfw.wmnet * 10:02 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2002.codfw.wmnet * 09:57 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2002.codfw.wmnet * 09:56 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2002.codfw.wmnet * 09:51 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2002.codfw.wmnet * 09:51 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1002.eqiad.wmnet * 09:47 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1002.eqiad.wmnet * 09:47 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1001.eqiad.wmnet * 09:44 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1001.eqiad.wmnet * 09:34 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2003.codfw.wmnet * 09:32 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2003.codfw.wmnet * 09:32 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2002.codfw.wmnet * 09:30 brouberol@dns1004: END - running authdns-update * 09:29 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2002.codfw.wmnet * 09:29 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2001.codfw.wmnet * 09:27 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 9 hosts * 09:27 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2001.codfw.wmnet * 09:27 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2001.codfw.wmnet * 09:26 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 9 hosts * 09:26 brouberol@dns1004: START - running authdns-update * 09:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 57 hosts * 09:24 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2001.codfw.wmnet * 09:24 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2002.codfw.wmnet * 09:22 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2002.codfw.wmnet * 09:21 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 57 hosts * 09:20 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2003.codfw.wmnet * 09:19 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 16 hosts * 09:16 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2003.codfw.wmnet * 09:16 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1003.eqiad.wmnet * 09:15 urbanecm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 09:15 urbanecm@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 09:13 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1003.eqiad.wmnet * 09:13 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1002.eqiad.wmnet * 09:11 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1002.eqiad.wmnet * 09:11 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1001.eqiad.wmnet * 09:07 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1001.eqiad.wmnet * 08:32 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:24 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:16 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 08:16 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 08:07 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:07 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:07 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 08:02 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 08:01 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:59 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:57 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 07:57 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 06:46 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1025.eqiad.wmnet with reason: Cloning * 06:46 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s4 * 06:45 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s6 * 06:44 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1019.eqiad.wmnet,service=s6 * 06:44 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1019.eqiad.wmnet,service=s4 * 03:40 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:40 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:40 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:40 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 03:37 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:37 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:37 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:36 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:49 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on mr1-eqsin,mr1-eqsin IPv6 with reason: connection issue * 02:38 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on cr[2-3]-eqsin.mgmt,ps1-[603-604]-eqsin with reason: connection issue * 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 27s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-23 == * 23:27 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin.oob,mr1-eqsin.oob IPv6 with reason: switch refresh * 22:21 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Setting storage compatibility to NONE - eevans@cumin1003 * 22:01 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Setting storage compatibility to NONE - eevans@cumin1003 * 21:29 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1014.eqiad.wmnet, repooling source-only afterwards * 21:28 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 46s) * 21:28 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 21:19 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Setting storage compatibility to UPGRADING - eevans@cumin1003 * 21:00 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Setting storage compatibility to UPGRADING - eevans@cumin1003 * 20:17 dani@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] (duration: 11m 57s) * 20:13 dani@deploy1003: dani: Continuing with deployment * 20:07 dani@deploy1003: dani: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:05 dani@deploy1003: Started scap sync-world: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] * 19:24 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:24 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating the rest of the ipv6 dns records. - jhancock@cumin2002" * 19:24 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating the rest of the ipv6 dns records. - jhancock@cumin2002" * 19:14 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 19:05 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wdqs1014.eqiad.wmnet with OS bookworm * 19:04 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.noop (exit_code=99) * 19:04 cwilliams@cumin1003: START - Cookbook sre.mysql.noop * 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1014.eqiad.wmnet with reason: host reimage * 18:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1014.eqiad.wmnet with reason: host reimage * 18:30 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2008.codfw.wmnet, repooling source-only afterwards * 18:28 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 19s) * 18:28 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1014 * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1014 * 18:22 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1014 * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1014.eqiad.wmnet 188.32.64.10.in-addr.arpa 8.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:22 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1014.eqiad.wmnet 188.32.64.10.in-addr.arpa 8.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1014 - bking@cumin2003" * 18:21 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1014 - bking@cumin2003" * 18:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2215: Maintenance * 18:18 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 18:15 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:15 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 18:06 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:06 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 18:05 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 18:04 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2052: codfw rack B8 re-pool after maintenance * 17:54 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 17:54 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:54 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 17:32 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2215: Maintenance * 17:29 cmooney@dns3003: END - running authdns-update * 17:27 cmooney@dns3003: START - running authdns-update * 17:23 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 17:22 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:18 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool es2052: codfw rack B8 re-pool after maintenance * 17:18 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2189: codfw rack B8 re-pool after maintenance * 17:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2215.codfw.wmnet with reason: Maintenance * 17:17 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 17:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2215 [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95126 and previous config saved to /var/cache/conftool/dbconfig/20260723-170903-cwilliams.json * 17:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2191 to x1 primary [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95125 and previous config saved to /var/cache/conftool/dbconfig/20260723-170612-cwilliams.json * 17:05 cezmunsta: Starting x1 codfw failover from db2215 to db2191 - [[phab:T432986|T432986]] * 16:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2191 with weight 0 [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95123 and previous config saved to /var/cache/conftool/dbconfig/20260723-165831-cwilliams.json * 16:58 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 16 hosts with reason: Primary switchover x1 [[phab:T432986|T432986]] * 16:36 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 138128 * 16:35 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 138128 * 16:33 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2189: codfw rack B8 re-pool after maintenance * 16:33 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2164: codfw rack B8 re-pool after maintenance * 16:28 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1072.eqiad.wmnet * 16:27 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1072.eqiad.wmnet with OS trixie * 16:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2249: Maintenance * 16:06 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker1072.eqiad.wmnet with reason: host reimage * 16:06 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1072.eqiad.wmnet with reason: host reimage * 15:50 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1072 * 15:50 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1072 * 15:49 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1072.eqiad.wmnet with OS trixie * 15:48 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2164: codfw rack B8 re-pool after maintenance * 15:48 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] (duration: 06m 37s) * 15:48 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2163: codfw rack B8 re-pool after maintenance * 15:45 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 15:43 musikanimal@deploy1003: musikanimal: Continuing with deployment * 15:43 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:41 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] * 15:36 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1072.eqiad.wmnet * 15:35 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1072.eqiad.wmnet * 15:35 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1072.eqiad.wmnet * 15:34 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:34 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push any outstanding updates - cmooney@cumin1003" * 15:34 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push any outstanding updates - cmooney@cumin1003" * 15:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db2249: Maintenance * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 15:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:26 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:21 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:21 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 15:21 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:21 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 15:20 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 15:19 cmooney@dns2004: END - running authdns-update * 15:17 cmooney@dns2004: START - running authdns-update * 15:14 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns2004.wikimedia.org * 15:12 brouberol@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 15:12 brouberol@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 15:12 klausman@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ml-serve1001.eqiad.wmnet with OS trixie * 15:11 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1071.eqiad.wmnet * 15:11 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1071.eqiad.wmnet with OS trixie * 15:10 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wdqs2008.codfw.wmnet with OS bookworm * 15:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2249.codfw.wmnet with reason: Maintenance * 15:08 brouberol@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 15:08 brouberol@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 15:08 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 15:08 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 15:06 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns1004.wikimedia.org * 15:02 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2002.codfw.wmnet * 15:02 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2002.codfw.wmnet * 15:02 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2163: codfw rack B8 re-pool after maintenance * 15:01 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 15:01 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 14:59 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test2001.codfw.wmnet * 14:57 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test2001.codfw.wmnet * 14:56 ryankemper: [WDQS] [[phab:T430880|T430880]] Reimaged `wdqs2016` to Bookworm, xferred scholarly_articles from `wdqs2024`, validated updater/readiness/federation, and repooled * 14:51 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2016.codfw.wmnet * 14:51 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1001.eqiad.wmnet with reason: host reimage * 14:48 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1071.eqiad.wmnet with reason: host reimage * 14:47 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1001.eqiad.wmnet with reason: host reimage * 14:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2008.codfw.wmnet with reason: host reimage * 14:43 topranks: reboot lsw1-b8-codw to upgrade JunOS [[phab:T430929|T430929]] * 14:41 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1071.eqiad.wmnet with reason: host reimage * 14:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2008.codfw.wmnet with reason: host reimage * 14:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2231: Maintenance * 14:30 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1001.eqiad.wmnet with OS trixie * 14:25 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2002.codfw.wmnet * 14:23 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1071 * 14:23 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1071 * 14:23 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 14:22 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore scholarly data after Bookworm reimage) xfer scholarly_articles from wdqs2024.codfw.wmnet -> wdqs2016.codfw.wmnet, repooling source-only afterwards * 14:22 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2052: codfw rack B8 depool for maintenance * 14:21 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool es2052: codfw rack B8 depool for maintenance * 14:21 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2249: codfw rack B8 depool for maintenance * 14:21 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1071 * 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1071.eqiad.wmnet 166.48.64.10.in-addr.arpa 6.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:21 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1071.eqiad.wmnet 166.48.64.10.in-addr.arpa 6.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1071 - jiji@cumin1003" * 14:21 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1071 - jiji@cumin1003" * 14:21 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2249: codfw rack B8 depool for maintenance * 14:21 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2189: codfw rack B8 depool for maintenance * 14:20 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2002.codfw.wmnet * 14:20 cmooney@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2050.codfw.wmnet * 14:20 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2189: codfw rack B8 depool for maintenance * 14:20 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2164: codfw rack B8 depool for maintenance * 14:20 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2164: codfw rack B8 depool for maintenance * 14:19 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2163: codfw rack B8 depool for maintenance * 14:19 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:19 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1014 * 14:19 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2163: codfw rack B8 depool for maintenance * 14:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1014.eqiad.wmnet with OS bookworm * 14:17 cmooney@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2050.codfw.wmnet * 14:16 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 14:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2008 * 14:14 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2008 * 14:14 cmooney@cumin1003: conftool action : set/pooled=no; selector: name=dns2004.wikimedia.org * 14:14 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2008.codfw.wmnet with OS bookworm * 14:13 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 14:12 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:10 topranks: depool dns2004 before lsw1-b8-codfw switch maintenance [[phab:T430929|T430929]] * 14:10 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b8-codfw,lsw1-b8-codfw IPv6,lsw1-b8-codfw.mgmt,ssw1-a[1,8]-codfw with reason: lsw1-b8-codfw JunOS upgrade * 14:07 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 30 hosts with reason: lsw1-b8-codfw JunOS upgrade * 14:06 elukey: upload python3-docker-report 0.0.19 to apt.wikimedia.org for bookworm and trixie * 13:59 jiji@cumin1003: START - Cookbook sre.dns.netbox * 13:58 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1071 * 13:57 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:57 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:57 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:57 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1071.eqiad.wmnet with OS trixie * 13:55 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1071.eqiad.wmnet * 13:55 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1071.eqiad.wmnet * 13:55 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1071.eqiad.wmnet * 13:53 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:52 logmsgbot: kharlan Deployed security patch for [[phab:T432948|T432948]] * 13:51 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:51 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 13:51 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db2231: Maintenance * 13:50 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:50 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:50 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:50 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:49 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:49 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:49 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2231 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95097 and previous config saved to /var/cache/conftool/dbconfig/20260723-134436-cwilliams.json * 13:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2231.codfw.wmnet with reason: Maintenance * 13:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:39 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:38 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 13:38 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] (duration: 09m 07s) * 13:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1037 hosts * 13:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2196: Maintenance * 13:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:33 kharlan@deploy1003: kharlan, emc-wmf: Continuing with deployment * 13:31 kharlan@deploy1003: kharlan, emc-wmf: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:30 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:28 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] * 13:17 hashar@deploy1003: Finished deploy [integration/docroot@2199146]: build: License GPL2.0+ / updating npm dependencies (duration: 00m 14s) * 13:17 hashar@deploy1003: Started deploy [integration/docroot@2199146]: build: License GPL2.0+ / updating npm dependencies * 13:14 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service * 13:07 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 12:58 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2207: Repooling * 12:49 root@cumin1003: START - Cookbook sre.mysql.pool pool db2196: Maintenance * 12:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2196 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95087 and previous config saved to /var/cache/conftool/dbconfig/20260723-123952-cwilliams.json * 12:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2196.codfw.wmnet with reason: Maintenance * 12:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2191: Maintenance * 12:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:13 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:13 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: Repooling * 12:12 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2207: Repooling * 12:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: Repooling * 11:56 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2235.codfw.wmnet with OS trixie * 11:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db2191: Maintenance * 11:46 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1070.eqiad.wmnet * 11:46 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1070.eqiad.wmnet * 11:46 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1070.eqiad.wmnet * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2191 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95080 and previous config saved to /var/cache/conftool/dbconfig/20260723-114308-cwilliams.json * 11:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2191.codfw.wmnet with reason: Maintenance * 11:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2186: Maintenance * 11:35 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 46375 * 11:34 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 46375 * 11:33 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2235.codfw.wmnet with reason: host reimage * 11:28 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2235.codfw.wmnet with reason: host reimage * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c7-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c7-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c6-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c6-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c5-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c5-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c4-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c4-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c3-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c3-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c2-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c2-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d7-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d7-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d4-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d4-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d3-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d2-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d2-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d8-eqiad * 11:23 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d8-eqiad * 11:23 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d1-eqiad * 11:23 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d1-eqiad * 11:12 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db2235.codfw.wmnet with OS trixie * 11:11 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:11 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[2160,2235].codfw.wmnet with reason: Upgrading * 10:56 root@cumin1003: START - Cookbook sre.mysql.pool pool db2186: Maintenance * 10:54 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1037: testing * 10:53 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1037: testing * 10:53 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1037: testing * 10:53 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1037: testing * 10:52 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: testing * 10:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2186 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95072 and previous config saved to /var/cache/conftool/dbconfig/20260723-104956-cwilliams.json * 10:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2186.codfw.wmnet with reason: Maintenance * 10:43 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1054: testing * 10:41 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1070.eqiad.wmnet with OS trixie * 10:30 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1037 hosts * 10:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 10:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 10:20 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1070.eqiad.wmnet with reason: host reimage * 10:16 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1070.eqiad.wmnet with reason: host reimage * 10:06 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1038: testing * 10:05 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1038: testing * 10:05 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1038: testing * 10:02 marostegui@dns1004: END - running authdns-update * 10:00 marostegui@dns1004: START - running authdns-update * 09:58 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: testing * 09:57 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1054: testing * 09:57 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1070 * 09:57 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1070 * 09:57 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1054: testing * 09:57 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1054: testing * 09:56 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1055: testing * 09:56 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1070 * 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1070.eqiad.wmnet 165.48.64.10.in-addr.arpa 5.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:56 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1070.eqiad.wmnet 165.48.64.10.in-addr.arpa 5.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1070 - jiji@cumin1003" * 09:56 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1070 - jiji@cumin1003" * 09:47 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for 1035 hosts * 09:45 jiji@cumin1003: START - Cookbook sre.dns.netbox * 09:42 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1070 * 09:42 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1070.eqiad.wmnet with OS trixie * 09:42 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1070.eqiad.wmnet * 09:41 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1070.eqiad.wmnet * 09:41 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1070.eqiad.wmnet * 09:27 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es2051: testing * 09:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: testing * 09:12 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es2051: testing * 09:11 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1055: testing * 09:09 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:09 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1055: testing * 09:09 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1055: testing * 08:50 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:50 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:50 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 08:49 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 08:49 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 08:49 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:46 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1069.eqiad.wmnet * 08:46 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1069.eqiad.wmnet * 08:46 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1069.eqiad.wmnet * 08:39 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 08:38 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 08:38 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 08:37 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 08:35 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:10 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1069.eqiad.wmnet with OS trixie * 07:49 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1069.eqiad.wmnet with reason: host reimage * 07:45 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1069.eqiad.wmnet with reason: host reimage * 07:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1035 hosts * 07:33 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 07:32 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 07:29 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1069 * 07:29 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1069 * 07:26 jiji@deploy1003: Finished scap sync-world: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules (duration: 06m 01s) * 07:25 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1069 * 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1069.eqiad.wmnet 164.48.64.10.in-addr.arpa 4.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:25 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1069.eqiad.wmnet 164.48.64.10.in-addr.arpa 4.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1069 - jiji@cumin1003" * 07:25 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1069 - jiji@cumin1003" * 07:25 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 07:25 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 07:24 jiji@deploy1003: jiji: Continuing with deployment * 07:22 jiji@deploy1003: jiji: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:21 jiji@deploy1003: Started scap sync-world: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules * 07:21 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1031.eqiad.wmnet,service=s7 * 07:20 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1031.eqiad.wmnet,service=s2 * 07:20 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1031.eqiad.wmnet,service=s7 * 07:20 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1031.eqiad.wmnet,service=s2 * 07:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts * 07:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts * 07:19 jiji@cumin1003: START - Cookbook sre.dns.netbox * 07:19 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1069 * 07:19 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1069.eqiad.wmnet with OS trixie * 07:19 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1069.eqiad.wmnet * 07:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 45 hosts * 07:17 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1069.eqiad.wmnet * 07:17 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1069.eqiad.wmnet * 07:14 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 45 hosts * 07:13 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 06:16 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs2007 after successful Bookworm reimage, data transfer, and postflight validation; wdqs1013 also passed postflights and is enabled in conftool, but remains out of IPVS pending a rolling pybal restart to clear its stale pre-VLAN-move address * 05:58 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2007.codfw.wmnet * 05:58 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1013.eqiad.wmnet * 05:54 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore scholarly data after Bookworm reimage) xfer scholarly_articles from wdqs2024.codfw.wmnet -> wdqs2016.codfw.wmnet, repooling source-only afterwards == 2026-07-22 == * 23:34 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Apply upgrade to JVM17 - eevans@cumin1003 * 23:14 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Apply upgrade to JVM17 - eevans@cumin1003 * 22:06 ryankemper: [WDQS] Added requestctl per-IP ratelimit `wdqs_heavy_sparql_bots_jul_2026_ratelimit` (chronic heavy-query bot tier driving deadlock-remediation restarts); pruned superseded `wdqs_2026_05_11_worobot` * 21:51 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] (duration: 11m 52s) * 21:44 sbassett@deploy1003: sbassett: Continuing with deployment * 21:43 sbassett@deploy1003: sbassett: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:39 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] * 20:38 dani@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] (duration: 32m 51s) * 20:38 ryankemper: [WDQS] Pruned obsolete requestctl action+pattern `wdqs_20260715_p2003_ring_ja3n` (actor rotated JA3Ns; rule inert) * 20:26 dani@deploy1003: dani, vadymts1: Continuing with deployment * 20:24 dani@deploy1003: dani, vadymts1: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:14 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2244: Testing * 20:06 dani@deploy1003: Started scap sync-world: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] * 19:56 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1013.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2016.codfw.wmnet with OS bookworm * 19:40 mutante: gerrit - one more service restart is needed - restarting * 19:29 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2244: Testing * 19:27 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2244: Testing * 19:27 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2244: Testing * 19:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2016.codfw.wmnet with reason: host reimage * 19:15 dancy@deploy1003: Finished deploy [zuul/deploy@d92e238]: Freshening Zuul installation (duration: 00m 15s) * 19:14 dancy@deploy1003: Started deploy [zuul/deploy@d92e238]: Freshening Zuul installation * 19:11 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2016.codfw.wmnet with reason: host reimage * 18:54 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1013.eqiad.wmnet, repooling source-only afterwards * 18:52 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2016 * 18:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2016 * 18:51 dancy@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 18:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2016.codfw.wmnet with OS bookworm * 18:39 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 11s) * 18:39 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 18:36 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 18:30 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417] (thin): Regular analytics weekly train THIN [analytics/refinery@2a25417d] (duration: 02m 09s) * 18:28 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417] (thin): Regular analytics weekly train THIN [analytics/refinery@2a25417d] * 18:28 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417]: Regular analytics weekly train [analytics/refinery@2a25417d] (duration: 04m 31s) * 18:27 dduvall: deploying https://gerrit.wikimedia.org/r/c/integration/config/+/1314025 (4 jobs updated) * 18:23 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417]: Regular analytics weekly train [analytics/refinery@2a25417d] * 18:22 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@2a25417d] (duration: 01m 59s) * 18:20 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@2a25417d] * 17:56 Raine: deployment server switchover => deploy1003 is primary now * 17:55 kamila@deploy1003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 28m 24s) * 17:54 mutante: restarting gerrit for maintenance * 17:29 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1013.eqiad.wmnet with OS bookworm * 17:27 kamila@deploy1003: Started scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] * 17:20 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] (duration: 22m 50s) * 17:12 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1023.eqiad.wmnet -> wdqs1024.eqiad.wmnet, repooling source-only afterwards * 17:04 Raine: point deployment.eqiad.wmnet to deploy1003 * 17:04 kamila@dns7001: END - running authdns-update * 17:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1013.eqiad.wmnet with reason: host reimage * 17:02 kamila@dns7001: START - running authdns-update * 17:01 kamila@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on releases2003.codfw.wmnet,releases1003.eqiad.wmnet with reason: Deployment server switchover * 17:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1013.eqiad.wmnet with reason: host reimage * 16:58 kamila@deploy2003: Locking from deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] * 16:57 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] (duration: 02m 33s) * 16:55 kamila@deploy2003: Locking from deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] * 16:55 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2003 - [[phab:T240266|T240266]] (duration: 00m 11s) * 16:54 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2003 - [[phab:T240266|T240266]] * 16:40 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1013 * 16:40 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1013 * 16:39 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1013 * 16:39 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1013.eqiad.wmnet 105.32.64.10.in-addr.arpa 5.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:39 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1013.eqiad.wmnet 105.32.64.10.in-addr.arpa 5.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:39 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:39 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1013 - bking@cumin2003" * 16:39 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1013 - bking@cumin2003" * 16:34 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:34 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1013 * 16:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1013.eqiad.wmnet with OS bookworm * 16:28 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1023.eqiad.wmnet -> wdqs1024.eqiad.wmnet, repooling source-only afterwards * 16:27 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-scholarly,name=eqiad * 16:27 eevans@deploy2003: helmfile [eqiad] DONE helmfile.d/services/linked-artifacts: apply * 16:26 eevans@deploy2003: helmfile [eqiad] START helmfile.d/services/linked-artifacts: apply * 16:26 eevans@deploy2003: helmfile [codfw] DONE helmfile.d/services/linked-artifacts: apply * 16:26 eevans@deploy2003: helmfile [codfw] START helmfile.d/services/linked-artifacts: apply * 16:25 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 29s) * 16:25 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 16:24 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 16:21 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 16:18 eevans@deploy2003: helmfile [codfw] DONE helmfile.d/services/linked-artifacts: apply * 16:18 eevans@deploy2003: helmfile [codfw] START helmfile.d/services/linked-artifacts: apply * 16:08 eevans@deploy2003: helmfile [staging] DONE helmfile.d/services/linked-artifacts: apply * 16:07 eevans@deploy2003: helmfile [staging] START helmfile.d/services/linked-artifacts: apply * 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 16:01 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 15:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1024.eqiad.wmnet with OS bookworm * 15:49 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] (duration: 00m 10s) * 15:49 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] * 15:48 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] (duration: 00m 15s) * 15:48 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] * 15:47 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] (duration: 00m 10s) * 15:47 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] * 15:46 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:42 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:42 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:40 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:37 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 15:37 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:36 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:36 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:36 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 15:33 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 15:33 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:31 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:28 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 15:27 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:27 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1024.eqiad.wmnet with reason: host reimage * 15:23 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:23 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:23 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:20 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1068.eqiad.wmnet * 15:20 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1068.eqiad.wmnet * 15:20 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1068.eqiad.wmnet * 15:20 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 15:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1024.eqiad.wmnet with reason: host reimage * 15:11 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:55 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 14:52 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wdqs1024.eqiad.wmnet with OS bookworm * 14:50 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] (duration: 00m 09s) * 14:50 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] * 14:49 jiji@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 14:49 jiji@deploy2003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 14:49 jiji@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 14:48 jiji@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 14:45 ecarg@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:45 ecarg@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:44 ecarg@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:44 ecarg@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:43 ecarg@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:43 ecarg@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:41 sukhe: ipvsadm --delete-service --tcp-service 10.2.1.55:8087: lvs2014 and lvs2013 * 14:39 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:39 sukhe: ipvsadm --delete-service --tcp-service 10.2.2.55:8087: [[phab:T432445|T432445]] * 14:38 ecarg@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:38 ecarg@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:37 ecarg@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:37 ecarg@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:36 ecarg@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:34 ecarg@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts datahubsearch1001.eqiad.wmnet * 14:32 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:32 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 14:31 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 14:31 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 14:28 sukhe: sudo cumin 'A:lvs-low-traffic-codfw' 'systemctl restart pybal': lvs2013 * 14:26 sukhe: sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal': lvs2014 * 14:26 sukhe: sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal' * 14:24 sukhe: restart pybal on lvs1019 * 14:24 sukhe: restart pybal on lvs1020 * 14:19 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for 1036 hosts * 14:17 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:04 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] (duration: 09m 28s) * 13:59 kharlan@deploy2003: dreamyjazz, kharlan: Continuing with deployment * 13:58 bking@cumin2003: START - Cookbook sre.hosts.decommission for hosts datahubsearch1001.eqiad.wmnet * 13:57 kharlan@deploy2003: dreamyjazz, kharlan: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:55 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] * 13:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts datahubsearch[1002-1003].eqiad.wmnet * 13:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:53 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch[1002-1003].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 13:52 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch[1002-1003].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 13:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 13:42 stran@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] (duration: 07m 30s) * 13:42 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:40 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: Test * 13:38 stran@deploy2003: dragoniez, stran: Continuing with deployment * 13:37 stran@deploy2003: dragoniez, stran: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:35 bking@cumin2003: START - Cookbook sre.hosts.decommission for hosts datahubsearch[1002-1003].eqiad.wmnet * 13:35 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024'] * 13:35 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 13:35 stran@deploy2003: Started scap sync-world: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] * 13:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 13:28 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024'] * 13:26 sukhe@dns1004: END - running authdns-update * 13:25 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 13:25 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 13:24 sukhe@dns1004: START - running authdns-update * 13:22 stran@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] (duration: 08m 20s) * 13:21 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 13:20 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 13:19 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 13:19 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 13:19 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 13:18 stran@deploy2003: stran: Continuing with deployment * 13:16 stran@deploy2003: stran: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:14 stran@deploy2003: Started scap sync-world: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] * 13:13 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 13:13 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 13:11 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 13:11 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 13:08 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 12:55 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool es2051: Test * 12:55 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: Test * 12:54 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool es2051: Test * 12:43 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1036 hosts * 12:41 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 12:40 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1048.eqiad.wmnet * 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 12:39 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 12:38 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 12:37 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 12:37 brouberol@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 12:36 brouberol@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 12:36 brouberol@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 12:36 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 12:36 elukey@cumin1003: DONE (PASS) - Cookbook sre.puppet.renew-cert (exit_code=0) for crm2001.codfw.wmnet: Renew puppet certificate - elukey@cumin1003 * 12:35 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:35 brouberol@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 12:34 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 12:31 brouberol@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 12:30 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1048.eqiad.wmnet * 12:30 brouberol@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 12:28 brouberol@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 12:27 brouberol@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 12:20 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1068.eqiad.wmnet with OS trixie * 12:01 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] (duration: 13m 19s) * 11:58 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1068.eqiad.wmnet with reason: host reimage * 11:52 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1068.eqiad.wmnet with reason: host reimage * 11:51 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 11:49 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:47 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] * 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2252: Security updates * 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:43 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 11:42 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2252: Security updates * 11:42 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply * 11:40 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply * 11:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 11:37 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2252.codfw.wmnet with OS trixie * 11:34 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1068 * 11:34 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1068 * 11:26 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1068 * 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1068.eqiad.wmnet 46.48.64.10.in-addr.arpa 6.4.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:26 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1068.eqiad.wmnet 46.48.64.10.in-addr.arpa 6.4.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1068 - jiji@cumin1003" * 11:26 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1068 - jiji@cumin1003" * 11:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2252.codfw.wmnet with reason: host reimage * 11:17 jiji@cumin1003: START - Cookbook sre.dns.netbox * 11:17 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1068 * 11:17 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1068.eqiad.wmnet with OS trixie * 11:17 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2252.codfw.wmnet with reason: host reimage * 11:15 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1068.eqiad.wmnet * 11:15 mvolz@deploy2003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:15 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1068.eqiad.wmnet * 11:15 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1068.eqiad.wmnet * 11:14 mvolz@deploy2003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:13 mvolz@deploy2003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:13 mvolz@deploy2003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:12 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] (duration: 11m 05s) * 11:11 mvolz@deploy2003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:10 mvolz@deploy2003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:07 dreamyjazz@deploy2003: dreamyjazz, kharlan: Continuing with deployment * 11:03 dreamyjazz@deploy2003: dreamyjazz, kharlan: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:03 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2252.codfw.wmnet with OS trixie * 11:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2252: Upgrading db2252.codfw.wmnet * 11:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:02 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 11:02 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2252: Upgrading db2252.codfw.wmnet * 11:02 cwilliams@cumin1003: dbmaint on ms3@codfw [[phab:T432321|T432321]] * 11:01 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 11:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db1153.eqiad.wmnet with reason: Security updates * 11:01 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] * 11:00 fnegri@deploy2003: helmfile [eqiad] DONE helmfile.d/services/toolhub: apply * 10:58 fnegri@deploy2003: helmfile [eqiad] START helmfile.d/services/toolhub: apply * 10:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1151: Security updates * 10:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:57 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1151: Security updates * 10:55 fnegri@deploy2003: helmfile [codfw] DONE helmfile.d/services/toolhub: apply * 10:54 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] (duration: 08m 38s) * 10:53 fnegri@deploy2003: helmfile [codfw] START helmfile.d/services/toolhub: apply * 10:53 fnegri@deploy2003: helmfile [staging] DONE helmfile.d/services/toolhub: apply * 10:52 fnegri@deploy2003: helmfile [staging] START helmfile.d/services/toolhub: apply * 10:50 zabe@deploy2003: zabe: Continuing with deployment * 10:47 zabe@deploy2003: zabe: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:45 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] * 10:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1151: Security updates * 10:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:42 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:42 root@cumin1003: START - Cookbook sre.mysql.depool depool db1151: Security updates * 10:38 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] (duration: 12m 47s) * 10:34 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 10:34 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 10:33 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2253.codfw.wmnet with OS trixie * 10:28 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:26 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] * 10:18 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2253.codfw.wmnet with reason: host reimage * 10:13 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2253.codfw.wmnet with reason: host reimage * 10:00 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2253.codfw.wmnet with OS trixie * 09:58 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db1151.eqiad.wmnet with reason: Security updates * 09:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2253: Upgrading db2253.codfw.wmnet * 09:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:57 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 09:56 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2253: Upgrading db2253.codfw.wmnet * 09:56 cwilliams@cumin1003: dbmaint on ms2@codfw [[phab:T432321|T432321]] * 09:56 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 09:36 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: UI improvement; support url shortener - oblivian@cumin1003" * 09:36 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: UI improvement; support url shortener - oblivian@cumin1003 * 09:35 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: UI improvement; support url shortener - oblivian@cumin1003 * 09:35 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: UI improvement; support url shortener - oblivian@cumin1003" * 09:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1152: Security updates * 09:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:26 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db1152: Security updates * 09:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: Security updates * 09:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:11 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:11 root@cumin1003: START - Cookbook sre.mysql.depool depool db1152: Security updates * 09:10 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1018.eqiad.wmnet with reason: Cloning * 09:09 Dreamy_Jazz: Deployed patch for [[phab:T432453|T432453]] and [[phab:T432454|T432454]] * 09:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 09:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2251.codfw.wmnet with OS trixie * 08:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2251.codfw.wmnet with reason: host reimage * 08:45 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2251.codfw.wmnet with reason: host reimage * 08:40 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] (duration: 12m 26s) * 08:38 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1030.eqiad.wmnet,service=s1 * 08:36 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 08:31 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2251.codfw.wmnet with OS trixie * 08:30 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:28 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] * 08:25 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] (duration: 07m 59s) * 08:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2251: Upgrading db2251.codfw.wmnet * 08:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:22 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 08:22 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2251: Upgrading db2251.codfw.wmnet * 08:20 urbanecm@deploy2003: urbanecm: Continuing with deployment * 08:20 cwilliams@cumin1003: dbmaint on ms1@codfw [[phab:T432321|T432321]] * 08:20 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 08:19 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:17 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] * 08:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade * 08:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade * 08:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db2251.codfw.wmnet,db1152.eqiad.wmnet with reason: OS upgrade * 08:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade * 08:13 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade * 08:11 Dreamy_Jazz: Created cusi_signal, cusi_case, and cusi_user on ukwiki and enwikivoyage in extension1 * 08:11 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade * 08:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade * 08:04 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1030.eqiad.wmnet,service=s1 * 08:04 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1030.eqiad.wmnet,service=s1 * 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply * 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply * 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply * 07:51 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply * 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 07:47 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 07:47 phuedx: End of UTC morning backport window * 07:43 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 07:43 phuedx@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] (duration: 13m 44s) * 07:43 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 07:39 phuedx@deploy2003: phuedx: Continuing with deployment * 07:31 phuedx@deploy2003: phuedx: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:29 phuedx@deploy2003: Started scap sync-world: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] * 07:24 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Turnilo import support - oblivian@cumin1003" * 07:24 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import support - oblivian@cumin1003 * 07:23 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import support - oblivian@cumin1003 * 07:23 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Turnilo import support - oblivian@cumin1003" * 06:42 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs2020 after successful Bookworm reimage, data transfer, and postflight validation * 06:42 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2020.codfw.wmnet * 05:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Managing sanitization for wikis bolwiki in section s5 * 05:25 marostegui@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis bolwiki in section s5 * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 41s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-21 == * 22:50 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2019.codfw.wmnet -> wdqs2020.codfw.wmnet, repooling source-only afterwards * 22:47 cwhite: force reboot arclamp2001 - appears to have run out of memory and gone unresponsive * 22:24 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 01m 26s) * 22:24 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 22:23 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 22:22 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024'] * 22:11 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 21:54 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs1024'] * 21:54 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 21:53 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs1024'] * 21:53 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 21:49 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2019.codfw.wmnet -> wdqs2020.codfw.wmnet, repooling source-only afterwards * 20:57 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] (duration: 09m 10s) * 20:55 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1024.eqiad.wmnet with OS bookworm * 20:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2020.codfw.wmnet with OS bookworm * 20:52 krinkle@deploy2003: krinkle: Continuing with deployment * 20:49 krinkle@deploy2003: krinkle: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:47 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] * 20:45 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] (duration: 05m 42s) * 20:44 krinkle@deploy2003: krinkle: Rolling back deployment * 20:41 krinkle@deploy2003: krinkle: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:39 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] * 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2003.codfw.wmnet * 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1003.eqiad.wmnet * 20:33 dani@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] (duration: 11m 15s) * 20:33 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2003.codfw.wmnet * 20:33 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1003.eqiad.wmnet * 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2020.codfw.wmnet with reason: host reimage * 20:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1002.eqiad.wmnet * 20:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2002.codfw.wmnet * 20:30 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 20:29 dani@deploy2003: dani: Continuing with deployment * 20:29 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2020.codfw.wmnet with reason: host reimage * 20:26 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1002.eqiad.wmnet * 20:26 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2002.codfw.wmnet * 20:24 dani@deploy2003: dani: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2001.codfw.wmnet * 20:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1001.eqiad.wmnet * 20:22 dani@deploy2003: Started scap sync-world: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] * 20:22 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 20:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2001.codfw.wmnet * 20:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1001.eqiad.wmnet * 20:14 sbisson@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] (duration: 09m 01s) * 20:11 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2020.codfw.wmnet with OS bookworm * 20:10 sbisson@deploy2003: sbisson: Continuing with deployment * 20:07 sbisson@deploy2003: sbisson: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:05 sbisson@deploy2003: Started scap sync-world: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] * 20:03 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] (duration: 07m 04s) * 20:01 mutante: Gerrit - tomorrow a new SSH host key will appear - it will be {{Gerrit|ed25519}} and has already been added to wmf-laptop. you can verify it here: https://wikitech.wikimedia.org/wiki/Help:SSH_Fingerprints/gerrit.wikimedia.org:29418 ([[phab:T240266|T240266]]) * 19:59 zabe@deploy2003: zabe: Continuing with deployment * 19:58 zabe@deploy2003: zabe: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:56 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] * 19:52 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] (duration: 07m 25s) * 19:48 zabe@deploy2003: zabe: Continuing with deployment * 19:47 zabe@deploy2003: zabe: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:45 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] * 19:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 19:32 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024'] * 19:27 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 19:26 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024'] * 19:26 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 19:24 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024'] * 19:12 ryankemper: [wdqs] [[phab:T430880|T430880]] Repooled `wdqs-scholarly` discovery in `eqiad` after validating `wdqs1023` end-to-end; `wdqs1024` remains disabled pending reimage recovery * 19:11 ryankemper: [wdqs] [[phab:T430880|T430880]] Repooled wdqs1012.eqiad.wmnet after successful Bookworm reimage, data transfer, service checks, readiness probe, and cross-graph federation query validation * 19:10 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 19:10 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1012.eqiad.wmnet * 19:08 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 18:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for deploy1003.eqiad.wmnet * 18:57 kamila@cumin1003: START - Cookbook sre.hosts.remove-downtime for deploy1003.eqiad.wmnet * 18:37 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:37 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding urldownloader service IPs - sukhe@cumin1003" * 18:37 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding urldownloader service IPs - sukhe@cumin1003" * 18:32 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 18:32 dancy@deploy2003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 18:30 sukhe@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 18:27 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 18:24 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1024.eqiad.wmnet with OS bookworm * 18:20 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host deploy1003.eqiad.wmnet with OS bookworm * 18:09 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deploy1003 reimage (duration: 121m 16s) * 18:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 18:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1180: Security updates * 17:55 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wcqs2003.codfw.wmnet * 17:48 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wcqs2003.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1155.eqiad.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1155.eqiad.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2224.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2224.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2217.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2217.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2193.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2193.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2180.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2180.codfw.wmnet * 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1168.eqiad.wmnet * 17:36 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1168.eqiad.wmnet * 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2169.codfw.wmnet * 17:36 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2169.codfw.wmnet * 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1165.eqiad.wmnet * 17:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1165.eqiad.wmnet * 17:35 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2158.codfw.wmnet * 17:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2158.codfw.wmnet * 17:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wcqs1003.eqiad.wmnet * 17:17 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1180: Security updates * 17:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1180.eqiad.wmnet * 17:16 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1180.eqiad.wmnet * 17:15 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp3073.* * 17:13 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wcqs1003.eqiad.wmnet * 17:11 brett@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp3073.esams.wmnet with OS trixie * 17:11 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 17:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1024 * 17:04 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1024 * 17:03 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 17:00 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 16:59 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2242: codfw rack B7 depool for maintenance * 16:59 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 16:43 brett@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp3073.esams.wmnet with reason: host reimage * 16:42 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 16:39 brett@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cp3073.esams.wmnet with reason: host reimage * 16:32 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on deploy1003.eqiad.wmnet with reason: host reimage * 16:27 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on deploy1003.eqiad.wmnet with reason: host reimage * 16:14 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2242: codfw rack B7 depool for maintenance * 16:14 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: codfw rack B7 depool for maintenance * 16:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1023.eqiad.wmnet with OS bookworm * 16:13 brett@cumin2002: START - Cookbook sre.hosts.reimage for host cp3073.esams.wmnet with OS trixie * 16:08 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host deploy1003.eqiad.wmnet with OS bookworm * 16:08 kamila@deploy2003: Locking from deployment [MediaWiki]: deploy1003 reimage * 16:03 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1012.eqiad.wmnet with OS bookworm * 15:48 inflatador: bking@apt1002 `sudo reprepro copy bookworm-wikimedia bullseye-wikimedia jvmquake` [[phab:T430880|T430880]] * 15:39 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp3073.* * 15:39 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 15:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:34 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1311536{{!}}Set $wgMathInternalRestbaseURL explicitly (take 2) (T349582)]] * 15:29 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:29 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2228: codfw rack B7 depool for maintenance * 15:29 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2229: codfw rack B7 depool for maintenance * 15:27 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1180: Security update * 15:25 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Security update * 15:21 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 15:21 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 15:19 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 24s) * 15:19 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:14 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 15:14 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db1180: Security update * 15:13 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-eqiad * 14:48 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-eqiad * 14:44 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2229: codfw rack B7 depool for maintenance * 14:44 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc2017: codfw rack B7 depool for maintenance * 14:44 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:43 cmooney@cumin2003: START - Cookbook sre.mysql.parsercache * 14:43 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool pc2017: codfw rack B7 depool for maintenance * 14:43 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2003.codfw.wmnet * 14:43 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2003.codfw.wmnet * 14:42 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2009.codfw.wmnet * 14:42 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2009.codfw.wmnet * 14:41 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:41 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:40 cmooney@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 29 hosts * 14:40 cmooney@cumin1003: START - Cookbook sre.hosts.remove-downtime for 29 hosts * 14:35 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 14:34 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 14:32 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] (duration: 07m 56s) * 14:29 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ssw1-a[1,8]-codfw with reason: lsw1-b7-codfw JunOS upgrade * 14:28 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 14:28 elukey@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 14:26 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:24 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] * 14:23 topranks: reboot lsw1-b7-codfw to upgrade JunOS (affects all hosts in rack) [[phab:T430928|T430928]] * 14:18 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2003.codfw.wmnet * 14:14 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2009.codfw.wmnet * 14:14 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Security update * 14:13 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2242: codfw rack B7 depool for maintenance * 14:13 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2242: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2228: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2228: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2229: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2229: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc2017: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.parsercache * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool pc2017: codfw rack B7 depool for maintenance * 14:08 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2003.codfw.wmnet * 14:07 cmooney@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on aux-k8s-etcd2004.codfw.wmnet,ml-etcd2001.codfw.wmnet with reason: lsw1-b7-codfw JunOS upgrade * 14:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95005 and previous config saved to /var/cache/conftool/dbconfig/20260721-140620-cwilliams.json * 14:05 cmooney@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2049.codfw.wmnet * 14:05 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-scholarly,name=eqiad * 14:04 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2009.codfw.wmnet * 14:04 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply * 14:04 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply * 14:03 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:03 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2001.codfw.wmnet * 14:03 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2001.codfw.wmnet * 14:02 cmooney@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2049.codfw.wmnet * 14:00 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:00 Dreamy_Jazz: Created cusi_case, cusi_signal, and cusi_user on svwiki, dewiki, jawiki, eswiki * 13:59 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b7-codfw,lsw1-b7-codfw IPv6,lsw1-b7-codfw.mgmt,ssw1-a[1,8]-codfw.mgmt with reason: lsw1-b7-codfw JunOS upgrade * 13:57 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs1023.eqiad.wmnet, repooling source-only afterwards * 13:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 29 hosts with reason: lsw1-b7-codfw JunOS upgrade * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224', diff saved to https://phabricator.wikimedia.org/P95003 and previous config saved to /var/cache/conftool/dbconfig/20260721-135613-cwilliams.json * 13:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1012.eqiad.wmnet with reason: host reimage * 13:53 cmooney@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 1:00:00 on 30 hosts with reason: lsw1-b7-codfw JunOS upgrade * 13:51 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1012.eqiad.wmnet with reason: host reimage * 13:48 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 13:48 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 13:46 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224', diff saved to https://phabricator.wikimedia.org/P95001 and previous config saved to /var/cache/conftool/dbconfig/20260721-134605-cwilliams.json * 13:46 cmooney@cumin1003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti2033.codfw.wmnet * 13:46 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 13:45 elukey: move the Docker Registry's /v2/wikimedia/machinelearning.* prefix to the ml S3 backend - [[phab:T428022|T428022]] * 13:45 cmooney@cumin1003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti2033.codfw.wmnet * 13:45 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 13:43 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:40 jiji@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 13:40 jiji@deploy2003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 13:39 jiji@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 13:39 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 13:38 cmooney@cumin1003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2032.codfw.wmnet * 13:38 jiji@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 13:38 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 13:37 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2032.codfw.wmnet * 13:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95000 and previous config saved to /var/cache/conftool/dbconfig/20260721-133557-cwilliams.json * 13:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1012 * 13:33 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1012 * 13:33 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1012.eqiad.wmnet with OS bookworm * 13:30 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:30 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:28 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94999 and previous config saved to /var/cache/conftool/dbconfig/20260721-132855-cwilliams.json * 13:28 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2224.codfw.wmnet with reason: Maintenance * 13:28 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94998 and previous config saved to /var/cache/conftool/dbconfig/20260721-132826-cwilliams.json * 13:28 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 13:23 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] (duration: 07m 50s) * 13:20 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 13:18 kharlan@deploy2003: kharlan: Continuing with deployment * 13:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217', diff saved to https://phabricator.wikimedia.org/P94996 and previous config saved to /var/cache/conftool/dbconfig/20260721-131817-cwilliams.json * 13:17 brouberol@dns1004: END - running authdns-update * 13:17 kharlan@deploy2003: kharlan: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:15 brouberol@dns1004: START - running authdns-update * 13:15 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] * 13:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94995 and previous config saved to /var/cache/conftool/dbconfig/20260721-131411-cwilliams.json * 13:13 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs1023.eqiad.wmnet, repooling source-only afterwards * 13:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217', diff saved to https://phabricator.wikimedia.org/P94994 and previous config saved to /var/cache/conftool/dbconfig/20260721-130809-cwilliams.json * 13:07 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 13:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180', diff saved to https://phabricator.wikimedia.org/P94993 and previous config saved to /var/cache/conftool/dbconfig/20260721-130404-cwilliams.json * 13:03 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:03 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:02 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 13:02 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 12:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94992 and previous config saved to /var/cache/conftool/dbconfig/20260721-125801-cwilliams.json * 12:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180', diff saved to https://phabricator.wikimedia.org/P94991 and previous config saved to /var/cache/conftool/dbconfig/20260721-125356-cwilliams.json * 12:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94990 and previous config saved to /var/cache/conftool/dbconfig/20260721-125049-cwilliams.json * 12:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2217.codfw.wmnet with reason: Maintenance * 12:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94989 and previous config saved to /var/cache/conftool/dbconfig/20260721-125017-cwilliams.json * 12:48 elukey: bmc cold reboot for lvs1013 and lvs1015 - [[phab:T426180|T426180]] * 12:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94988 and previous config saved to /var/cache/conftool/dbconfig/20260721-124348-cwilliams.json * 12:40 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193', diff saved to https://phabricator.wikimedia.org/P94987 and previous config saved to /var/cache/conftool/dbconfig/20260721-124009-cwilliams.json * 12:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts * 12:33 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts * 12:33 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts * 12:32 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts * 12:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:30 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193', diff saved to https://phabricator.wikimedia.org/P94986 and previous config saved to /var/cache/conftool/dbconfig/20260721-123001-cwilliams.json * 12:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94985 and previous config saved to /var/cache/conftool/dbconfig/20260721-121953-cwilliams.json * 12:17 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs2007.codfw.wmnet with OS bookworm * 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94983 and previous config saved to /var/cache/conftool/dbconfig/20260721-121257-cwilliams.json * 12:12 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2193.codfw.wmnet with reason: Maintenance * 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94982 and previous config saved to /var/cache/conftool/dbconfig/20260721-121239-cwilliams.json * 12:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180', diff saved to https://phabricator.wikimedia.org/P94980 and previous config saved to /var/cache/conftool/dbconfig/20260721-120231-cwilliams.json * 11:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180', diff saved to https://phabricator.wikimedia.org/P94979 and previous config saved to /var/cache/conftool/dbconfig/20260721-115223-cwilliams.json * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94978 and previous config saved to /var/cache/conftool/dbconfig/20260721-114333-cwilliams.json * 11:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1180.eqiad.wmnet with reason: Maintenance * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94977 and previous config saved to /var/cache/conftool/dbconfig/20260721-114305-cwilliams.json * 11:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94976 and previous config saved to /var/cache/conftool/dbconfig/20260721-114215-cwilliams.json * 11:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94975 and previous config saved to /var/cache/conftool/dbconfig/20260721-113530-cwilliams.json * 11:35 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2180.codfw.wmnet with reason: Maintenance * 11:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94974 and previous config saved to /var/cache/conftool/dbconfig/20260721-113501-cwilliams.json * 11:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168', diff saved to https://phabricator.wikimedia.org/P94973 and previous config saved to /var/cache/conftool/dbconfig/20260721-113258-cwilliams.json * 11:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169', diff saved to https://phabricator.wikimedia.org/P94972 and previous config saved to /var/cache/conftool/dbconfig/20260721-112453-cwilliams.json * 11:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168', diff saved to https://phabricator.wikimedia.org/P94971 and previous config saved to /var/cache/conftool/dbconfig/20260721-112250-cwilliams.json * 11:21 XioNoX: put eqiad-drmrs Arelion link in service * 11:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169', diff saved to https://phabricator.wikimedia.org/P94970 and previous config saved to /var/cache/conftool/dbconfig/20260721-111446-cwilliams.json * 11:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94969 and previous config saved to /var/cache/conftool/dbconfig/20260721-111242-cwilliams.json * 11:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 11:10 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 11:07 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1093 hosts * 11:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94968 and previous config saved to /var/cache/conftool/dbconfig/20260721-110548-cwilliams.json * 11:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1168.eqiad.wmnet with reason: Maintenance * 11:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94967 and previous config saved to /var/cache/conftool/dbconfig/20260721-110520-cwilliams.json * 11:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94966 and previous config saved to /var/cache/conftool/dbconfig/20260721-110439-cwilliams.json * 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94964 and previous config saved to /var/cache/conftool/dbconfig/20260721-105632-cwilliams.json * 10:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2169.codfw.wmnet with reason: Maintenance * 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94963 and previous config saved to /var/cache/conftool/dbconfig/20260721-105603-cwilliams.json * 10:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165', diff saved to https://phabricator.wikimedia.org/P94962 and previous config saved to /var/cache/conftool/dbconfig/20260721-105512-cwilliams.json * 10:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158', diff saved to https://phabricator.wikimedia.org/P94961 and previous config saved to /var/cache/conftool/dbconfig/20260721-104555-cwilliams.json * 10:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165', diff saved to https://phabricator.wikimedia.org/P94960 and previous config saved to /var/cache/conftool/dbconfig/20260721-104504-cwilliams.json * 10:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158', diff saved to https://phabricator.wikimedia.org/P94959 and previous config saved to /var/cache/conftool/dbconfig/20260721-103547-cwilliams.json * 10:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94958 and previous config saved to /var/cache/conftool/dbconfig/20260721-103456-cwilliams.json * 10:29 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2229: Upgraded kernel * 10:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94956 and previous config saved to /var/cache/conftool/dbconfig/20260721-102757-cwilliams.json * 10:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on an-redacteddb1001.eqiad.wmnet,clouddb[1015,1025,1028].eqiad.wmnet,db1155.eqiad.wmnet with reason: Maintenance * 10:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1165.eqiad.wmnet with reason: Maintenance * 10:25 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94955 and previous config saved to /var/cache/conftool/dbconfig/20260721-102539-cwilliams.json * 10:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94954 and previous config saved to /var/cache/conftool/dbconfig/20260721-101848-cwilliams.json * 10:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2158.codfw.wmnet with reason: Maintenance * 09:43 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2229: Upgraded kernel * 09:42 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2229.codfw.wmnet * 09:42 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2229.codfw.wmnet * 09:23 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db2229.codfw.wmnet * 09:23 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2229.codfw.wmnet * 08:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2229 [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94948 and previous config saved to /var/cache/conftool/dbconfig/20260721-085724-cwilliams.json * 08:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2214 to s6 primary [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94947 and previous config saved to /var/cache/conftool/dbconfig/20260721-085442-cwilliams.json * 08:53 cezmunsta: Starting s6 codfw failover from db2229 to db2214 - [[phab:T430964|T430964]] * 08:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2214 with weight 0 [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94946 and previous config saved to /var/cache/conftool/dbconfig/20260721-084613-cwilliams.json * 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 22 hosts with reason: Primary switchover s6 [[phab:T430964|T430964]] * 08:32 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1017.eqiad.wmnet,service=s1 * 08:08 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Add subrated circuit rate to interface descriptions - CR1312476 - ayounsi@cumin1003 * 08:06 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Add subrated circuit rate to interface descriptions - CR1312476 - ayounsi@cumin1003 * 07:58 reedy@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] (duration: 12m 55s) * 07:51 reedy@deploy2003: reedy, neriah: Continuing with deployment * 07:51 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 07:51 reedy@deploy2003: reedy, neriah: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:48 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1093 hosts * 07:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm2001.wikimedia.org * 07:45 reedy@deploy2003: Started scap sync-world: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] * 07:43 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 07:42 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm2001.wikimedia.org * 07:23 elukey: upgrade libtiff6 packages on zuul* trixie hosts for security upgrades * 07:22 elukey: upgrade libtiff6 packages on Wikikube trixie workers for security upgrades * 07:14 elukey@deploy2003: helmfile [codfw] DONE helmfile.d/services/proton: sync * 07:13 elukey@deploy2003: helmfile [codfw] START helmfile.d/services/proton: sync * 07:11 elukey@deploy2003: helmfile [eqiad] DONE helmfile.d/services/proton: sync * 07:10 elukey@deploy2003: helmfile [eqiad] START helmfile.d/services/proton: sync * 07:09 elukey@deploy2003: helmfile [staging] DONE helmfile.d/services/proton: sync * 07:08 elukey@deploy2003: helmfile [staging] START helmfile.d/services/proton: sync * 06:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1023.eqiad.wmnet with reason: host reimage * 06:46 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1023.eqiad.wmnet with reason: host reimage * 06:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 05:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Haproxy-only mode support - oblivian@cumin1003" * 05:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Haproxy-only mode support - oblivian@cumin1003 * 05:42 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Haproxy-only mode support - oblivian@cumin1003 * 05:42 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Haproxy-only mode support - oblivian@cumin1003" * 05:38 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1017.eqiad.wmnet with reason: Cloning * 05:37 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1017.eqiad.wmnet,service=s1 * 05:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1029.eqiad.wmnet,service=s8 * 05:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1029.eqiad.wmnet,service=s5 * 05:32 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:30 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 05:11 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet * 05:04 aokoth@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet * 05:00 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 04:56 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 04:01 mwpresync@deploy2003: Pruned MediaWiki: 1.47.0-wmf.9 (duration: 01m 08s) * 03:41 mwpresync@deploy2003: Finished scap sync-world: testwikis to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] (duration: 36m 30s) * 03:05 mwpresync@deploy2003: Started scap sync-world: testwikis to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 03:01 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:01 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:00 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:00 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:36 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:36 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:36 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:35 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:16 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 47s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 00:56 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm == 2026-07-20 == * 23:38 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 23:07 Amir1: deleting echo notifications from 2015 on group1 wikis * 23:07 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] (duration: 14m 16s) * 23:01 ladsgroup@deploy2003: ladsgroup: Continuing with deployment * 23:00 ladsgroup@deploy2003: ladsgroup: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:53 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] * 22:46 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2007.codfw.wmnet, repooling source-only afterwards * 22:39 maryum: Deployed security fixes for several security bugs * 21:42 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 21:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2007.codfw.wmnet, repooling source-only afterwards * 21:37 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 21:37 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 21:34 sbassett: Deployed security fix for [[phab:T432424|T432424]] * 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs2020.codfw.wmnet * 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1023.eqiad.wmnet * 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1011.eqiad.wmnet * 21:32 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 17s) * 21:32 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 21:27 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 21:13 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2007.codfw.wmnet with reason: host reimage * 21:08 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-internal-main,name=codfw * 21:06 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2007.codfw.wmnet with reason: host reimage * 20:59 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service * 20:58 sukhe: pybal restart for IP changes around wdqs-main hosts * 20:57 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 20:46 ryankemper@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-internal-main,name=codfw * 20:45 ebernhardson@deploy2003: Finished deploy [search/mjolnir/deploy@d4dc3b8]: Update for opensearch 2.x compat (duration: 00m 34s) * 20:44 ebernhardson@deploy2003: Started deploy [search/mjolnir/deploy@d4dc3b8]: Update for opensearch 2.x compat * 20:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2007 * 20:44 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2007 * 20:43 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2007 * 20:43 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2007.codfw.wmnet 156.16.192.10.in-addr.arpa 6.5.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:42 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2007.codfw.wmnet 156.16.192.10.in-addr.arpa 6.5.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:42 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2007 - bking@cumin2003" * 20:41 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2007 - bking@cumin2003" * 20:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94944 and previous config saved to /var/cache/conftool/dbconfig/20260720-203333-cwilliams.json * 20:32 arlolra@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] (duration: 15m 07s) * 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2020.codfw.wmnet * 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1023.eqiad.wmnet * 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1011.eqiad.wmnet * 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs2020.codfw.wmnet * 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1023.eqiad.wmnet * 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1011.eqiad.wmnet * 20:25 arlolra@deploy2003: arlolra, cscott: Continuing with deployment * 20:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257', diff saved to https://phabricator.wikimedia.org/P94943 and previous config saved to /var/cache/conftool/dbconfig/20260720-202325-cwilliams.json * 20:21 arlolra@deploy2003: arlolra, cscott: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:17 arlolra@deploy2003: Started scap sync-world: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] * 20:13 bking@cumin2003: START - Cookbook sre.dns.netbox * 20:13 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257', diff saved to https://phabricator.wikimedia.org/P94942 and previous config saved to /var/cache/conftool/dbconfig/20260720-201318-cwilliams.json * 20:13 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 20:10 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 20:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2007 * 20:04 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2007.codfw.wmnet with OS bookworm * 20:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94941 and previous config saved to /var/cache/conftool/dbconfig/20260720-200310-cwilliams.json * 19:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94940 and previous config saved to /var/cache/conftool/dbconfig/20260720-195633-cwilliams.json * 19:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1257.eqiad.wmnet with reason: Maintenance * 19:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94939 and previous config saved to /var/cache/conftool/dbconfig/20260720-195605-cwilliams.json * 19:51 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 19:50 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 19:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256', diff saved to https://phabricator.wikimedia.org/P94938 and previous config saved to /var/cache/conftool/dbconfig/20260720-194558-cwilliams.json * 19:44 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 19:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 19:41 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wdqs1011.eqiad.wmnet with OS bookworm * 19:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256', diff saved to https://phabricator.wikimedia.org/P94937 and previous config saved to /var/cache/conftool/dbconfig/20260720-193550-cwilliams.json * 19:25 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94936 and previous config saved to /var/cache/conftool/dbconfig/20260720-192542-cwilliams.json * 19:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94935 and previous config saved to /var/cache/conftool/dbconfig/20260720-191856-cwilliams.json * 19:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1256.eqiad.wmnet with reason: Maintenance * 19:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94934 and previous config saved to /var/cache/conftool/dbconfig/20260720-191839-cwilliams.json * 19:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255', diff saved to https://phabricator.wikimedia.org/P94933 and previous config saved to /var/cache/conftool/dbconfig/20260720-190831-cwilliams.json * 18:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255', diff saved to https://phabricator.wikimedia.org/P94932 and previous config saved to /var/cache/conftool/dbconfig/20260720-185824-cwilliams.json * 18:50 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 18:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94931 and previous config saved to /var/cache/conftool/dbconfig/20260720-184816-cwilliams.json * 18:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94930 and previous config saved to /var/cache/conftool/dbconfig/20260720-184224-cwilliams.json * 18:42 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1255.eqiad.wmnet with reason: Maintenance * 18:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94929 and previous config saved to /var/cache/conftool/dbconfig/20260720-184153-cwilliams.json * 18:39 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 18:39 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 16s) * 18:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 18:38 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 59m 26s) * 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 18:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211', diff saved to https://phabricator.wikimedia.org/P94928 and previous config saved to /var/cache/conftool/dbconfig/20260720-183145-cwilliams.json * 18:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211', diff saved to https://phabricator.wikimedia.org/P94927 and previous config saved to /var/cache/conftool/dbconfig/20260720-182137-cwilliams.json * 18:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94926 and previous config saved to /var/cache/conftool/dbconfig/20260720-181129-cwilliams.json * 18:09 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_codfw * 18:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2057.codfw.wmnet * 18:08 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_codfw * 18:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2058.codfw.wmnet * 18:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94925 and previous config saved to /var/cache/conftool/dbconfig/20260720-180452-cwilliams.json * 18:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on clouddb[1016,1020,1022-1023].eqiad.wmnet,db1154.eqiad.wmnet with reason: Maintenance * 18:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1211.eqiad.wmnet with reason: Maintenance * 18:02 sukhe: armed keyholder on acmechief1002.eqiad.wmnet and acmechief2002.codfw.wmnet (active host) * 18:01 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief2002.codfw.wmnet * 17:57 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief2002.codfw.wmnet * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs2020'] * 17:52 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief1002.eqiad.wmnet * 17:50 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 17:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1011.eqiad.wmnet with reason: host reimage * 17:48 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief1002.eqiad.wmnet * 17:47 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test2001.codfw.wmnet * 17:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94924 and previous config saved to /var/cache/conftool/dbconfig/20260720-174717-cwilliams.json * 17:46 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 17:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1011.eqiad.wmnet with reason: host reimage * 17:43 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs2020.codfw.wmnet with OS bookworm * 17:43 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test2001.codfw.wmnet * 17:43 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test1001.eqiad.wmnet * 17:39 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test1001.eqiad.wmnet * 17:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 17:38 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:38 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 17:37 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 17:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244', diff saved to https://phabricator.wikimedia.org/P94923 and previous config saved to /var/cache/conftool/dbconfig/20260720-173709-cwilliams.json * 17:35 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 17:31 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 20m 40s) * 17:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2055.codfw.wmnet * 17:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2056.codfw.wmnet * 17:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1011 * 17:27 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1011 * 17:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1011.eqiad.wmnet with OS bookworm * 17:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244', diff saved to https://phabricator.wikimedia.org/P94922 and previous config saved to /var/cache/conftool/dbconfig/20260720-172701-cwilliams.json * 17:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94921 and previous config saved to /var/cache/conftool/dbconfig/20260720-171653-cwilliams.json * 17:11 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 17:11 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 13m 03s) * 17:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94920 and previous config saved to /var/cache/conftool/dbconfig/20260720-171012-cwilliams.json * 17:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2244.codfw.wmnet with reason: Maintenance * 17:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94919 and previous config saved to /var/cache/conftool/dbconfig/20260720-170941-cwilliams.json * 16:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243', diff saved to https://phabricator.wikimedia.org/P94918 and previous config saved to /var/cache/conftool/dbconfig/20260720-165933-cwilliams.json * 16:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 16:58 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2053.codfw.wmnet * 16:51 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2054.codfw.wmnet * 16:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243', diff saved to https://phabricator.wikimedia.org/P94917 and previous config saved to /var/cache/conftool/dbconfig/20260720-164926-cwilliams.json * 16:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94916 and previous config saved to /var/cache/conftool/dbconfig/20260720-163918-cwilliams.json * 16:35 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 16:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94915 and previous config saved to /var/cache/conftool/dbconfig/20260720-163140-cwilliams.json * 16:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2243.codfw.wmnet with reason: Maintenance * 16:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94914 and previous config saved to /var/cache/conftool/dbconfig/20260720-163111-cwilliams.json * 16:27 btullis@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 16:27 btullis@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 16:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2020 * 16:23 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2020 * 16:21 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2020 * 16:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2020.codfw.wmnet 85.0.192.10.in-addr.arpa 5.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:21 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2020.codfw.wmnet 85.0.192.10.in-addr.arpa 5.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242', diff saved to https://phabricator.wikimedia.org/P94913 and previous config saved to /var/cache/conftool/dbconfig/20260720-162103-cwilliams.json * 16:19 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 16:18 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 16:18 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:18 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 16:17 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:17 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for netbox accounting errors - jhancock@cumin2002" * 16:17 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for netbox accounting errors - jhancock@cumin2002" * 16:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2051.codfw.wmnet * 16:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2052.codfw.wmnet * 16:11 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 16:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242', diff saved to https://phabricator.wikimedia.org/P94912 and previous config saved to /var/cache/conftool/dbconfig/20260720-161055-cwilliams.json * 16:09 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 16:08 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 16:06 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 16:06 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 16:06 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2020 * 16:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2020.codfw.wmnet with OS bookworm * 16:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94911 and previous config saved to /var/cache/conftool/dbconfig/20260720-160047-cwilliams.json * 15:58 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2019.codfw.wmnet, repooling source-only afterwards * 15:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94909 and previous config saved to /var/cache/conftool/dbconfig/20260720-155353-cwilliams.json * 15:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2242.codfw.wmnet with reason: Maintenance * 15:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94908 and previous config saved to /var/cache/conftool/dbconfig/20260720-154433-cwilliams.json * 15:35 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2049.codfw.wmnet * 15:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162', diff saved to https://phabricator.wikimedia.org/P94907 and previous config saved to /var/cache/conftool/dbconfig/20260720-153425-cwilliams.json * 15:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2050.codfw.wmnet * 15:28 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162', diff saved to https://phabricator.wikimedia.org/P94906 and previous config saved to /var/cache/conftool/dbconfig/20260720-152418-cwilliams.json * 15:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1023 * 15:14 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1023 * 15:14 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 15:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94905 and previous config saved to /var/cache/conftool/dbconfig/20260720-151407-cwilliams.json * 15:13 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] (duration: 41m 16s) * 15:08 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 15:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94902 and previous config saved to /var/cache/conftool/dbconfig/20260720-150729-cwilliams.json * 15:07 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2162.codfw.wmnet with reason: Maintenance * 15:05 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2027.codfw.wmnet, repooling source-only afterwards * 15:00 urbanecm@deploy2003: vadymts1, migr, urbanecm: Continuing with deployment * 14:59 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:58 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 07s) * 14:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:58 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 13s) * 14:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:57 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2019.codfw.wmnet, repooling source-only afterwards * 14:57 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2047.codfw.wmnet * 14:55 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2048.codfw.wmnet * 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2019.codfw.wmnet with OS bookworm * 14:49 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 14:47 urbanecm@deploy2003: vadymts1, migr, urbanecm: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:44 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool magru [reason: BGP issues in lvs7003 resolved after liberica restart, no task ID specified] * 14:44 sukhe@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool magru [reason: BGP issues in lvs7003 resolved after liberica restart, no task ID specified] * 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:41 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:39 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:39 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:33 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool magru [reason: no reason specified, no task ID specified] * 14:33 sukhe@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool magru [reason: no reason specified, no task ID specified] * 14:31 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] * 14:24 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:24 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:24 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:24 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2019.codfw.wmnet with reason: host reimage * 14:22 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2027.codfw.wmnet, repooling source-only afterwards * 14:19 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2019.codfw.wmnet with reason: host reimage * 14:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2027.codfw.wmnet with OS bookworm * 14:16 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2046.codfw.wmnet * 14:16 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2045.codfw.wmnet * 14:08 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:08 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:08 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:08 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:07 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:06 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:06 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:06 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:05 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2019 * 14:00 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2019 * 13:56 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1015.eqiad.wmnet * 13:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2027.codfw.wmnet with reason: host reimage * 13:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2071.codfw.wmnet with OS trixie * 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:51 sukhe@cumin1003: END (ERROR) - Cookbook sre.loadbalancer.admin (exit_code=97) rebooting A:liberica and P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica and P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:51 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1015.eqiad.wmnet * 13:50 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1014.eqiad.wmnet * 13:50 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1076.eqiad.wmnet with OS trixie * 13:50 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2027.codfw.wmnet with reason: host reimage * 13:45 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1014.eqiad.wmnet * 13:44 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1013.eqiad.wmnet * 13:39 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 13:38 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1013.eqiad.wmnet * 13:37 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2044.codfw.wmnet * 13:37 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2043.codfw.wmnet * 13:36 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2019 * 13:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2019.codfw.wmnet 156.32.192.10.in-addr.arpa 6.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:36 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2019.codfw.wmnet 156.32.192.10.in-addr.arpa 6.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:36 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2019 - bking@cumin2003" * 13:36 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2019 - bking@cumin2003" * 13:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry2005.codfw.wmnet * 13:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2071.codfw.wmnet with reason: host reimage * 13:31 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:31 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2019 * 13:31 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry2005.codfw.wmnet * 13:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry2004.codfw.wmnet * 13:30 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2019.codfw.wmnet with OS bookworm * 13:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2027 * 13:30 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2027 * 13:30 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2027.codfw.wmnet with OS bookworm * 13:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1076.eqiad.wmnet with reason: host reimage * 13:29 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_codfw * 13:28 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_codfw * 13:26 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry2004.codfw.wmnet * 13:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry1005.eqiad.wmnet * 13:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2071.codfw.wmnet with reason: host reimage * 13:22 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1076.eqiad.wmnet with reason: host reimage * 13:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry1005.eqiad.wmnet * 13:21 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry1004.eqiad.wmnet * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry1004.eqiad.wmnet * 13:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts * 13:13 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts * 13:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts * 13:12 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts * 13:03 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1076.eqiad.wmnet with OS trixie * 13:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2071.codfw.wmnet with OS trixie * 12:55 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:54 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:53 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:46 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 7 hosts * 12:42 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 7 hosts * 12:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts * 12:42 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts * 12:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2070.codfw.wmnet with OS trixie * 12:36 ozge@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:35 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1075.eqiad.wmnet with OS trixie * 12:32 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts * 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts * 12:22 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin1001.eqiad.wmnet * 12:19 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin1001.eqiad.wmnet * 12:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2070.codfw.wmnet with reason: host reimage * 12:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin2001.codfw.wmnet * 12:14 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1075.eqiad.wmnet with reason: host reimage * 12:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2070.codfw.wmnet with reason: host reimage * 12:10 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1075.eqiad.wmnet with reason: host reimage * 12:09 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin2001.codfw.wmnet * 11:17 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1074.eqiad.wmnet with OS trixie * 11:17 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 11:16 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 11:14 ozge@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 6 hosts * 11:09 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 6 hosts * 11:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 324 hosts * 10:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1074.eqiad.wmnet with reason: host reimage * 10:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2069.codfw.wmnet with OS trixie * 10:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1074.eqiad.wmnet with reason: host reimage * 10:30 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2069.codfw.wmnet with reason: host reimage * 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1074.eqiad.wmnet with OS trixie * 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2069.codfw.wmnet with reason: host reimage * 10:06 blake@deploy2003: Stopping before sync operations * 10:06 blake@deploy2003: Started scap sync-world: Non-deployment scap run to populate new release values * 10:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2069.codfw.wmnet with OS trixie * 10:00 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1073.eqiad.wmnet with OS trixie * 09:56 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 324 hosts * 09:39 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 09:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 8 hosts * 09:38 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1073.eqiad.wmnet with reason: host reimage * 09:37 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 8 hosts * 09:34 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1073.eqiad.wmnet with reason: host reimage * 09:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2068.codfw.wmnet with OS trixie * 09:16 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1073.eqiad.wmnet with OS trixie * 09:13 blake@deploy2003: sync-world aborted: Non-deployment scap run to populate new release values (duration: 00m 02s) * 09:13 blake@deploy2003: Started scap sync-world: Non-deployment scap run to populate new release values * 08:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2068.codfw.wmnet with reason: host reimage * 08:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2068.codfw.wmnet with reason: host reimage * 08:50 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 08:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2068.codfw.wmnet with OS trixie * 08:15 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1072.eqiad.wmnet with OS trixie * 07:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2067.codfw.wmnet with OS trixie * 07:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1072.eqiad.wmnet with reason: host reimage * 07:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1072.eqiad.wmnet with reason: host reimage * 07:45 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 07:45 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 07:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2067.codfw.wmnet with reason: host reimage * 07:35 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2067.codfw.wmnet with reason: host reimage * 07:30 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 07:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1072.eqiad.wmnet with OS trixie * 07:30 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 07:17 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 07:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2067.codfw.wmnet with OS trixie * 05:51 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:50 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:25 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:25 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on db2207.codfw.wmnet with reason: Host down * 04:28 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 07m 02s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-18 == * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 29s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 00:11 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2018.codfw.wmnet, repooling source-only afterwards == 2026-07-17 == * 23:53 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2026.codfw.wmnet, repooling source-only afterwards * 23:09 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2018.codfw.wmnet, repooling source-only afterwards * 23:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2026.codfw.wmnet, repooling source-only afterwards * 22:11 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2018.codfw.wmnet with OS bookworm * 22:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2026.codfw.wmnet with OS bookworm * 21:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2018.codfw.wmnet with reason: host reimage * 21:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2018.codfw.wmnet with reason: host reimage * 21:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2026.codfw.wmnet with reason: host reimage * 21:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2026.codfw.wmnet with reason: host reimage * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2018 * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2018 * 21:26 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2018 * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2018.codfw.wmnet 155.32.192.10.in-addr.arpa 5.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:26 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2018.codfw.wmnet 155.32.192.10.in-addr.arpa 5.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2018 - bking@cumin2003" * 21:26 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2018 - bking@cumin2003" * 21:14 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:13 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2018 * 21:13 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2018.codfw.wmnet with OS bookworm * 21:12 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2026 * 21:12 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2026 * 21:12 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2026.codfw.wmnet with OS bookworm * 21:05 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1022.eqiad.wmnet -> wdqs1026.eqiad.wmnet, repooling source-only afterwards * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs2017.codfw.wmnet, repooling source-only afterwards * 20:11 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs2017.codfw.wmnet, repooling source-only afterwards * 20:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2017.codfw.wmnet with OS bookworm * 20:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1022.eqiad.wmnet -> wdqs1026.eqiad.wmnet, repooling source-only afterwards * 20:06 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1026.eqiad.wmnet with OS bookworm * 19:55 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 19:55 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:55 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 09s) * 19:55 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:50 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 08s) * 19:50 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:50 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 10m 03s) * 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2017.codfw.wmnet with reason: host reimage * 19:40 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:40 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 15s) * 19:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1026.eqiad.wmnet with reason: host reimage * 19:37 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 16s) * 19:37 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:34 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2017.codfw.wmnet with reason: host reimage * 19:34 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1026.eqiad.wmnet with reason: host reimage * 19:33 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 19:33 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:16 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1026.eqiad.wmnet with OS bookworm * 19:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2017 * 19:16 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2017 * 19:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2017.codfw.wmnet with OS bookworm * 18:30 bking@dns1004: END - running authdns-update * 18:28 bking@dns1004: START - running authdns-update * 18:16 kamila@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1264.eqiad.wmnet * 18:16 kamila@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1264.eqiad.wmnet * 18:16 kamila@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1264.eqiad.wmnet * 17:49 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 17:46 dzahn@dns1006: END - running authdns-update * 17:44 dzahn@dns1006: START - running authdns-update * 17:44 dzahn@dns1006: END - running authdns-update * 17:42 dzahn@dns1006: START - running authdns-update * 17:28 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 17:21 kamila@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 17:01 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1264 * 17:01 kamila@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1264 * 17:01 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 17:01 kamila@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1264.eqiad.wmnet * 17:01 kamila@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1264.eqiad.wmnet * 17:01 kamila@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1264.eqiad.wmnet * 16:42 reedy@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] (duration: 10m 29s) * 16:34 reedy@deploy2003: reedy, hartman: Continuing with deployment * 16:33 reedy@deploy2003: reedy, hartman: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:31 reedy@deploy2003: Started scap sync-world: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] * 16:26 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 16:10 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in2001.wikimedia.org with reason: [[phab:T431659|T431659]] * 16:07 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in1001.wikimedia.org with reason: [[phab:T431659|T431659]] * 16:05 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 16:01 kamila@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 16:00 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out2001.wikimedia.org with reason: [[phab:T431659|T431659]] * 15:41 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 15:41 kamila@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 15:35 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out1001.wikimedia.org with reason: [[phab:T431659|T431659]] * 15:14 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1339.eqiad.wmnet * 15:13 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1339.eqiad.wmnet * 15:13 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1339.eqiad.wmnet * 14:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1339.eqiad.wmnet with OS trixie * 14:50 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:49 kamila@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:49 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:33 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage * 14:27 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage * 14:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1339 * 14:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1339 * 14:14 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1339 * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1339.eqiad.wmnet 156.32.64.10.in-addr.arpa 6.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:14 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1339.eqiad.wmnet 156.32.64.10.in-addr.arpa 6.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1339 - cgoubert@cumin2003" * 14:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1339 - cgoubert@cumin2003" * 14:09 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 14:06 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1339 * 14:06 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie * 14:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1339.eqiad.wmnet * 14:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1339.eqiad.wmnet * 14:02 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1339.eqiad.wmnet * 13:45 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb1013.eqiad.wmnet * 13:39 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb1013.eqiad.wmnet * 13:27 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:24 blake@dns1004: END - running authdns-update * 13:22 blake@dns1004: START - running authdns-update * 13:20 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 13:11 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2014.codfw.wmnet * 13:06 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb2014.codfw.wmnet * 13:06 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2012.codfw.wmnet * 13:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 13:03 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 13:01 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 15 hosts * 13:01 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb2012.codfw.wmnet * 13:01 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1016.eqiad.wmnet * 13:00 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 15 hosts * 12:55 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb1016.eqiad.wmnet * 12:55 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1014.eqiad.wmnet * 12:49 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb1014.eqiad.wmnet * 12:32 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:32 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:31 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:31 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1338.eqiad.wmnet * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1338.eqiad.wmnet * 12:18 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1338.eqiad.wmnet * 12:17 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:15 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:14 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:13 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1338.eqiad.wmnet with OS trixie * 12:01 klausman@deploy2003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 11:59 klausman@deploy2003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 11:56 klausman@deploy2003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 11:54 klausman@deploy2003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 11:53 klausman@deploy2003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 11:51 klausman@deploy2003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 11:42 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1338.eqiad.wmnet with reason: host reimage * 11:38 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1338.eqiad.wmnet with reason: host reimage * 11:31 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2230.codfw.wmnet * 11:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1338 * 11:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1338 * 11:25 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1338 * 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1338.eqiad.wmnet 155.32.64.10.in-addr.arpa 5.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:25 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1338.eqiad.wmnet 155.32.64.10.in-addr.arpa 5.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1338 - cgoubert@cumin2003" * 11:25 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1338 - cgoubert@cumin2003" * 11:23 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2230.codfw.wmnet * 11:20 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 11:20 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1338 * 11:20 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1338.eqiad.wmnet with OS trixie * 11:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1338.eqiad.wmnet * 11:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1338.eqiad.wmnet * 11:19 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1338.eqiad.wmnet * 11:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1337.eqiad.wmnet * 11:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1337.eqiad.wmnet * 11:17 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1337.eqiad.wmnet * 11:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1337.eqiad.wmnet with OS trixie * 10:51 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[2001-2002].codfw.wmnet * 10:50 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1337.eqiad.wmnet with reason: host reimage * 10:40 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 10:39 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:39 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1337.eqiad.wmnet with reason: host reimage * 10:39 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1001-1003].eqiad.wmnet * 10:34 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:34 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:30 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:28 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1001-1003].eqiad.wmnet * 10:27 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1337 * 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1337 * 10:26 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1337 * 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1337.eqiad.wmnet 154.32.64.10.in-addr.arpa 4.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:26 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1337.eqiad.wmnet 154.32.64.10.in-addr.arpa 4.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1337 - cgoubert@cumin2003" * 10:26 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1337 - cgoubert@cumin2003" * 10:21 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 10:18 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1337 * 10:17 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1337.eqiad.wmnet with OS trixie * 10:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1337.eqiad.wmnet * 10:16 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db1176.eqiad.wmnet * 10:16 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1337.eqiad.wmnet * 10:16 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1337.eqiad.wmnet * 10:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1336.eqiad.wmnet * 10:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1336.eqiad.wmnet * 10:15 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1336.eqiad.wmnet * 10:11 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db1176.eqiad.wmnet * 10:10 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db1176.eqiad.wmnet * 10:09 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db1176.eqiad.wmnet * 10:05 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts (check the cookbook's logs for more details.) * 10:03 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts (check the cookbook's logs for more details.) * 09:58 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1336.eqiad.wmnet with OS trixie * 09:47 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts (check the cookbook's logs for more details.) * 09:47 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts (check the cookbook's logs for more details.) * 09:45 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host acmechief-test2001.codfw.wmnet,acmechief-test1001.eqiad.wmnet,an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet,db-test[2001-2002].codfw.wmnet,db-test[1 * 09:40 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host acmechief-test2001.codfw.wmnet,acmechief-test1001.eqiad.wmnet,an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet,db-test[2001-2002].codfw.wmnet,db-test[1001-1003].eqiad.wmn * 09:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1336.eqiad.wmnet with reason: host reimage * 09:33 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1336.eqiad.wmnet with reason: host reimage * 09:29 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet * 09:29 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet * 09:28 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:26 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:21 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 09:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1336 * 09:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1336 * 09:19 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 09:14 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1336 * 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1336.eqiad.wmnet 152.32.64.10.in-addr.arpa 2.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1336.eqiad.wmnet 152.32.64.10.in-addr.arpa 2.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1336 - cgoubert@cumin2003" * 09:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1336 - cgoubert@cumin2003" * 09:11 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:10 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 09:09 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:09 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1336 * 09:09 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1336.eqiad.wmnet with OS trixie * 09:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1336.eqiad.wmnet * 09:08 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1336.eqiad.wmnet * 09:08 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1336.eqiad.wmnet * 09:06 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1335.eqiad.wmnet * 09:06 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1335.eqiad.wmnet * 09:06 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1335.eqiad.wmnet * 09:04 elukey: uploaded spicerack_13.1.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia * 08:55 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wikikube-worker-exp2001.codfw.wmnet * 08:54 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host testreduce1002.eqiad.wmnet * 08:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1335.eqiad.wmnet with OS trixie * 08:51 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host wikikube-worker-exp2001.codfw.wmnet * 08:51 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wikikube-worker-exp1001.eqiad.wmnet * 08:50 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host testreduce1002.eqiad.wmnet * 08:45 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host wikikube-worker-exp1001.eqiad.wmnet * 08:34 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1335.eqiad.wmnet with reason: host reimage * 08:30 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1335.eqiad.wmnet with reason: host reimage * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1335 * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1335 * 08:18 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1335 * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1335.eqiad.wmnet 150.32.64.10.in-addr.arpa 0.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:18 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1335.eqiad.wmnet 150.32.64.10.in-addr.arpa 0.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1335 - cgoubert@cumin2003" * 08:18 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1335 - cgoubert@cumin2003" * 08:14 elukey@cumin1003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 08:14 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges * 08:13 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 08:10 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1335 * 08:10 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1335.eqiad.wmnet with OS trixie * 08:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1335.eqiad.wmnet * 08:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1335.eqiad.wmnet * 08:09 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1335.eqiad.wmnet * 08:06 elukey@cumin1003: END (FAIL) - Cookbook sre.puppet.disable-merges (exit_code=99) * 08:05 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges * 08:03 elukey@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin1003.eqiad.wmnet * 07:57 elukey@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin1003.eqiad.wmnet * 07:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetdb1003.eqiad.wmnet * 07:46 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetdb1003.eqiad.wmnet * 07:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetdb2003.codfw.wmnet * 07:37 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetdb2003.codfw.wmnet * 07:37 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1001.eqiad.wmnet * 07:28 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver1001.eqiad.wmnet * 07:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet * 07:19 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet * 07:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2002.codfw.wmnet * 07:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver2002.codfw.wmnet * 07:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet * 07:05 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet * 07:04 btullis@cumin1003: END (FAIL) - Cookbook sre.hadoop.reboot-workers (exit_code=99) for Hadoop analytics cluster * 07:04 elukey@cumin1003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 07:04 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges * 06:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox1003.eqiad.wmnet * 06:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox1003.eqiad.wmnet * 02:46 ryankemper@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:46 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:44 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:37 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:37 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-internal-scholarly,name=eqiad * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 49s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 01:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore wdqs1025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling source-only afterwards * 01:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore wdqs1027 after Bookworm reimage) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs1027.eqiad.wmnet, repooling both afterwards * 00:55 urbanecm@deploy2003: helmfile [codfw] DONE helmfile.d/services/linkrecommendation: apply * 00:54 urbanecm@deploy2003: helmfile [eqiad] DONE helmfile.d/services/linkrecommendation: apply * 00:54 urbanecm@deploy2003: helmfile [staging] DONE helmfile.d/services/linkrecommendation: apply * 00:54 urbanecm@deploy2003: helmfile [codfw] START helmfile.d/services/linkrecommendation: apply * 00:53 urbanecm@deploy2003: helmfile [staging] START helmfile.d/services/linkrecommendation: apply * 00:52 urbanecm@deploy2003: helmfile [eqiad] START helmfile.d/services/linkrecommendation: apply * 00:23 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore wdqs1025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling source-only afterwards * 00:23 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore wdqs1027 after Bookworm reimage) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs1027.eqiad.wmnet, repooling both afterwards * 00:14 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1274.eqiad.wmnet * 00:14 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1274.eqiad.wmnet * 00:14 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1274.eqiad.wmnet * 00:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1027.eqiad.wmnet with OS bookworm * 00:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1025.eqiad.wmnet with OS bookworm * 00:04 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1274.eqiad.wmnet with OS trixie == 2026-07-16 == * 23:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], xfer to freshly reimaged/scap-deployed wdqs2025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs2025.codfw.wmnet, repooling source-only afterwards * 23:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1027.eqiad.wmnet with reason: host reimage * 23:47 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1025.eqiad.wmnet with reason: host reimage * 23:43 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1274.eqiad.wmnet with reason: host reimage * 23:41 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1025.eqiad.wmnet with reason: host reimage * 23:39 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1027.eqiad.wmnet with reason: host reimage * 23:38 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1274.eqiad.wmnet with reason: host reimage * 23:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1025 * 23:23 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1025 * 23:22 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1027 * 23:22 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1027 * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1274 * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1274 * 23:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1025.eqiad.wmnet with OS bookworm * 23:19 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1274 * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1274.eqiad.wmnet 145.48.64.10.in-addr.arpa 5.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:19 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1274.eqiad.wmnet 145.48.64.10.in-addr.arpa 5.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1274 - swfrench@cumin1003" * 23:19 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1274 - swfrench@cumin1003" * 23:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1027.eqiad.wmnet with OS bookworm * 23:14 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 23:14 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1274 * 23:13 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1274.eqiad.wmnet with OS trixie * 23:13 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1274.eqiad.wmnet * 23:12 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1274.eqiad.wmnet * 23:12 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1274.eqiad.wmnet * 23:12 ryankemper: [[phab:T430880|T430880]] depooled dnsdisc of wdqs-internal-scholarly-eqiad bc we only have 1 host there * 23:09 ryankemper@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-internal-scholarly,name=eqiad * 23:08 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1272.eqiad.wmnet * 23:08 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1272.eqiad.wmnet * 23:08 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1272.eqiad.wmnet * 23:01 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], xfer to freshly reimaged/scap-deployed wdqs2025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs2025.codfw.wmnet, repooling source-only afterwards * 22:57 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1272.eqiad.wmnet with OS trixie * 22:56 Amir1: deleting echo notifications from 2015 in group0 * 22:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2025.codfw.wmnet with OS bookworm * 22:35 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1272.eqiad.wmnet with reason: host reimage * 22:32 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 27s) * 22:32 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 22:28 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1269.eqiad.wmnet * 22:28 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1269.eqiad.wmnet * 22:28 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1269.eqiad.wmnet * 22:27 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1272.eqiad.wmnet with reason: host reimage * 22:26 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] (duration: 08m 51s) * 22:22 ladsgroup@deploy2003: ladsgroup, urbanecm: Continuing with deployment * 22:19 ladsgroup@deploy2003: ladsgroup, urbanecm: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:17 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] * 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2025.codfw.wmnet with reason: host reimage * 22:06 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1272 * 22:06 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1272 * 22:05 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1272 * 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1272.eqiad.wmnet 127.48.64.10.in-addr.arpa 7.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:05 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1272.eqiad.wmnet 127.48.64.10.in-addr.arpa 7.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1272 - swfrench@cumin1003" * 22:05 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1272 - swfrench@cumin1003" * 22:03 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2025.codfw.wmnet with reason: host reimage * 22:01 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 22:00 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1272 * 22:00 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1272.eqiad.wmnet with OS trixie * 22:00 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1272.eqiad.wmnet * 21:59 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1272.eqiad.wmnet * 21:59 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1272.eqiad.wmnet * 21:55 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1271.eqiad.wmnet * 21:55 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1271.eqiad.wmnet * 21:55 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1271.eqiad.wmnet * 21:46 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1271.eqiad.wmnet with OS trixie * 21:45 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] (duration: 06m 31s) * 21:40 sbassett@deploy2003: sbassett: Continuing with deployment * 21:40 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2025 * 21:40 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2025 * 21:40 sbassett@deploy2003: sbassett: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:38 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] * 21:37 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2025 * 21:37 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2025.codfw.wmnet 220.48.192.10.in-addr.arpa 0.2.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:37 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2025.codfw.wmnet 220.48.192.10.in-addr.arpa 0.2.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:37 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:37 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2025 - bking@cumin2003" * 21:37 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2025 - bking@cumin2003" * 21:30 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] (duration: 08m 19s) * 21:26 sbassett@deploy2003: sbassett: Continuing with deployment * 21:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1269.eqiad.wmnet with OS trixie * 21:24 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1271.eqiad.wmnet with reason: host reimage * 21:23 sbassett@deploy2003: sbassett: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:22 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:22 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] * 21:20 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2025 * 21:19 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2025.codfw.wmnet with OS bookworm * 21:17 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1271.eqiad.wmnet with reason: host reimage * 21:04 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1269.eqiad.wmnet with reason: host reimage * 21:00 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1269.eqiad.wmnet with reason: host reimage * 20:56 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1271 * 20:55 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1271 * 20:54 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1271 * 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1271.eqiad.wmnet 126.48.64.10.in-addr.arpa 6.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:54 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1271.eqiad.wmnet 126.48.64.10.in-addr.arpa 6.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1271 - swfrench@cumin1003" * 20:54 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1271 - swfrench@cumin1003" * 20:51 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:51 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:51 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:50 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 20:49 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 20:49 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1268.eqiad.wmnet * 20:49 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1268.eqiad.wmnet * 20:49 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1268.eqiad.wmnet * 20:48 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1271 * 20:48 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1271.eqiad.wmnet with OS trixie * 20:47 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1271.eqiad.wmnet * 20:46 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1271.eqiad.wmnet * 20:46 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1271.eqiad.wmnet * 20:41 aude@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] (duration: 07m 34s) * 20:39 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1269 * 20:39 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1269 * 20:36 aude@deploy2003: aude: Continuing with deployment * 20:35 aude@deploy2003: aude: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:33 aude@deploy2003: Started scap sync-world: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] * 20:26 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on db2207.codfw.wmnet with reason: Host down * 20:22 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-video: apply * 20:21 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-video: apply * 20:20 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-timeline: apply * 20:20 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-timeline: apply * 20:20 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-syntaxhighlight: apply * 20:19 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-syntaxhighlight: apply * 20:19 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-media: apply * 20:18 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-media: apply * 20:18 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-constraints: apply * 20:17 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-constraints: apply * 20:17 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox: apply * 20:16 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox: apply * 20:13 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1269 * 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1269.eqiad.wmnet 80.32.64.10.in-addr.arpa 0.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:13 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1269.eqiad.wmnet 80.32.64.10.in-addr.arpa 0.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1269 - kamila@cumin1003" * 20:13 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1269 - kamila@cumin1003" * 20:09 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-flink-codfw cluster: Roll restart of jvm daemons. * 20:07 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 20:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 20:03 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-flink-codfw cluster: Roll restart of jvm daemons. * 20:03 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2207 [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94893 and previous config saved to /var/cache/conftool/dbconfig/20260716-200257-marostegui.json * 20:01 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2204 to s2 primary [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94892 and previous config saved to /var/cache/conftool/dbconfig/20260716-200157-marostegui.json * 20:00 marostegui: Starting emergency s2 codfw failover from db2207 to db2204 - [[phab:T432396|T432396]] * 19:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1035.eqiad.wmnet * 19:56 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2204 with weight 0 [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94891 and previous config saved to /var/cache/conftool/dbconfig/20260716-195628-marostegui.json * 19:55 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 26 hosts with reason: Primary switchover s2 [[phab:T432396|T432396]] * 19:54 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1035.eqiad.wmnet * 19:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1034.eqiad.wmnet * 19:48 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1034.eqiad.wmnet * 19:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1033.eqiad.wmnet * 19:43 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-video: apply * 19:43 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1033.eqiad.wmnet * 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1032.eqiad.wmnet * 19:42 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-video: apply * 19:42 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-timeline: apply * 19:41 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-timeline: apply * 19:41 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-syntaxhighlight: apply * 19:41 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-syntaxhighlight: apply * 19:40 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-media: apply * 19:40 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-media: apply * 19:39 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-constraints: apply * 19:36 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-constraints: apply * 19:36 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox: apply * 19:35 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1032.eqiad.wmnet * 19:35 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1031.eqiad.wmnet * 19:35 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox: apply * 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-video: apply * 19:33 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-video: apply * 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-timeline: apply * 19:33 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-timeline: apply * 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-syntaxhighlight: apply * 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-syntaxhighlight: apply * 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-media: apply * 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-media: apply * 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-constraints: apply * 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-constraints: apply * 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox: apply * 19:31 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox: apply * 19:27 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1031.eqiad.wmnet * 19:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1030.eqiad.wmnet * 19:23 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1001.eqiad.wmnet, repooling source-only afterwards * 19:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1030.eqiad.wmnet * 19:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1029.eqiad.wmnet * 19:17 kamila@cumin1003: START - Cookbook sre.dns.netbox * 19:12 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1029.eqiad.wmnet * 19:06 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1269 * 19:05 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1269.eqiad.wmnet with OS trixie * 19:03 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1269.eqiad.wmnet * 19:03 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1269.eqiad.wmnet * 19:03 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1269.eqiad.wmnet * 18:55 dancy@deploy2003: Finished scap sync-world: testing [[phab:T428971|T428971]] (duration: 02m 41s) * 18:53 dancy@deploy2003: Started scap sync-world: testing [[phab:T428971|T428971]] * 18:31 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1268.eqiad.wmnet with OS trixie * 18:18 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 18:16 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1267.eqiad.wmnet * 18:16 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1267.eqiad.wmnet * 18:16 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1267.eqiad.wmnet * 18:09 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1268.eqiad.wmnet with reason: host reimage * 18:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1001.eqiad.wmnet, repooling source-only afterwards * 18:06 swfrench@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] (duration: 07m 34s) * 18:06 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 23s) * 18:06 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 18:06 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1268.eqiad.wmnet with reason: host reimage * 18:03 bd808@deploy2003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 18:02 bd808@deploy2003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 18:02 swfrench@deploy2003: jiji, swfrench: Continuing with deployment * 18:02 bd808@deploy2003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 18:02 bd808@deploy2003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 18:01 bd808@deploy2003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 18:01 swfrench@deploy2003: jiji, swfrench: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:01 bd808@deploy2003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:59 swfrench@deploy2003: Started scap sync-world: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] * 17:45 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1268 * 17:45 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1268 * 17:44 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1267.eqiad.wmnet with OS trixie * 17:43 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter2006.codfw.wmnet * 17:39 swfrench@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter2006.codfw.wmnet * 17:35 swfrench@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] (duration: 07m 27s) * 17:34 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1268 * 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1268.eqiad.wmnet 78.32.64.10.in-addr.arpa 8.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:34 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1268.eqiad.wmnet 78.32.64.10.in-addr.arpa 8.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1268 - kamila@cumin1003" * 17:34 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1268 - kamila@cumin1003" * 17:31 swfrench@deploy2003: jiji, swfrench: Continuing with deployment * 17:29 swfrench@deploy2003: jiji, swfrench: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:28 kamila@cumin1003: START - Cookbook sre.dns.netbox * 17:28 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1268 * 17:28 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1268.eqiad.wmnet with OS trixie * 17:27 swfrench@deploy2003: Started scap sync-world: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] * 17:23 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1267.eqiad.wmnet with reason: host reimage * 17:18 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1267.eqiad.wmnet with reason: host reimage * 17:18 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1268.eqiad.wmnet * 17:17 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1268.eqiad.wmnet * 17:17 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1268.eqiad.wmnet * 17:12 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter2005.codfw.wmnet * 17:11 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1270.eqiad.wmnet * 17:11 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1270.eqiad.wmnet * 17:11 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1270.eqiad.wmnet * 17:09 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter2005.codfw.wmnet * 17:08 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:08 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update reverse dns for moved arelion cct cr2-eqiad - cmooney@cumin1003" * 17:08 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update reverse dns for moved arelion cct cr2-eqiad - cmooney@cumin1003" * 17:08 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] (duration: 07m 34s) * 17:04 jiji@deploy2003: jiji: Continuing with deployment * 17:03 jiji@deploy2003: jiji: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 17:00 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] * 17:00 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2185.codfw.wmnet with OS trixie * 16:59 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:58 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1270.eqiad.wmnet with OS trixie * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1267 * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1267 * 16:57 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1267 * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1267.eqiad.wmnet 77.32.64.10.in-addr.arpa 7.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:57 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1267.eqiad.wmnet 77.32.64.10.in-addr.arpa 7.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1267 - kamila@cumin1003" * 16:56 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1267 - kamila@cumin1003" * 16:56 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_eqsin * 16:56 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5032.eqsin.wmnet * 16:52 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_esams * 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3073.esams.wmnet * 16:50 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_esams * 16:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3081.esams.wmnet * 16:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1266.eqiad.wmnet * 16:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1266.eqiad.wmnet * 16:45 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1266.eqiad.wmnet * 16:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2185.codfw.wmnet with reason: host reimage * 16:41 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_eqiad * 16:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1114.eqiad.wmnet * 16:41 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_eqiad * 16:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1115.eqiad.wmnet * 16:39 kamila@cumin1003: START - Cookbook sre.dns.netbox * 16:39 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1267 * 16:39 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2185.codfw.wmnet with reason: host reimage * 16:38 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1267.eqiad.wmnet with OS trixie * 16:38 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1267.eqiad.wmnet * 16:38 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1270.eqiad.wmnet with reason: host reimage * 16:37 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1267.eqiad.wmnet * 16:37 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1267.eqiad.wmnet * 16:31 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1270.eqiad.wmnet with reason: host reimage * 16:24 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1264.eqiad.wmnet * 16:24 kamila@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 16:24 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 16:23 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 16:21 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2162: switch maintenance completed codfw rack b6 * 16:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2185.codfw.wmnet with OS trixie * 16:19 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 16:16 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter1007.eqiad.wmnet * 16:15 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_eqsin * 16:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5024.eqsin.wmnet * 16:13 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5031.eqsin.wmnet * 16:13 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3072.esams.wmnet * 16:12 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter1007.eqiad.wmnet * 16:11 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] (duration: 09m 47s) * 16:10 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1266.eqiad.wmnet with OS trixie * 16:10 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1270 * 16:10 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1270 * 16:09 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1270 * 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1270.eqiad.wmnet 125.48.64.10.in-addr.arpa 5.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1270.eqiad.wmnet 125.48.64.10.in-addr.arpa 5.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1270 - swfrench@cumin1003" * 16:09 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1270 - swfrench@cumin1003" * 16:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3080.esams.wmnet * 16:07 jiji@deploy2003: jiji: Continuing with deployment * 16:06 jiji@deploy2003: jiji: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:04 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 16:04 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1265.eqiad.wmnet * 16:03 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1265.eqiad.wmnet * 16:03 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1265.eqiad.wmnet * 16:03 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1270 * 16:03 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1270.eqiad.wmnet with OS trixie * 16:02 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1270.eqiad.wmnet * 16:02 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] * 16:01 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1270.eqiad.wmnet * 16:01 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1270.eqiad.wmnet * 16:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1113.eqiad.wmnet * 16:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1112.eqiad.wmnet * 15:49 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1266.eqiad.wmnet with reason: host reimage * 15:47 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter1006.eqiad.wmnet * 15:45 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1265.eqiad.wmnet with OS trixie * 15:44 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1266.eqiad.wmnet with reason: host reimage * 15:43 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter1006.eqiad.wmnet * 15:42 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] (duration: 09m 46s) * 15:37 jiji@deploy2003: jiji: Continuing with deployment * 15:36 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2162: switch maintenance completed codfw rack b6 * 15:36 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2161: switch maintenance completed codfw rack b6 * 15:34 jiji@deploy2003: jiji: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:32 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5023.eqsin.wmnet * 15:32 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] * 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5030.eqsin.wmnet * 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3071.esams.wmnet * 15:27 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3079.esams.wmnet * 15:25 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1265.eqiad.wmnet with reason: host reimage * 15:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1266 * 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1266 * 15:21 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1110.eqiad.wmnet * 15:20 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1111.eqiad.wmnet * 15:16 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1265.eqiad.wmnet with reason: host reimage * 15:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1001.eqiad.wmnet with OS bookworm * 15:15 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1266 * 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1266.eqiad.wmnet 76.32.64.10.in-addr.arpa 6.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:15 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1266.eqiad.wmnet 76.32.64.10.in-addr.arpa 6.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1266 - kamila@cumin1003" * 15:15 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1266 - kamila@cumin1003" * 15:07 kamila@cumin1003: START - Cookbook sre.dns.netbox * 15:04 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1266 * 15:04 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1264 * 15:04 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1264 * 15:04 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1266.eqiad.wmnet with OS trixie * 15:03 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1264 * 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1264.eqiad.wmnet 74.32.64.10.in-addr.arpa 4.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:03 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1264.eqiad.wmnet 74.32.64.10.in-addr.arpa 4.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1264 - kamila@cumin1003" * 15:03 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1264 - kamila@cumin1003" * 15:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-eqiad * 15:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp1001.eqiad.wmnet * 15:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp1001.eqiad.wmnet * 15:01 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp1001.eqiad.wmnet * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp1001.eqiad.wmnet * 15:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1376-1384].eqiad.wmnet * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1376-1384].eqiad.wmnet * 14:59 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1002.eqiad.wmnet * 14:59 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1266.eqiad.wmnet * 14:58 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1266.eqiad.wmnet * 14:58 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1266.eqiad.wmnet * 14:58 kamila@cumin1003: START - Cookbook sre.dns.netbox * 14:57 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1264 * 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1265 * 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1265 * 14:57 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1310596{{!}}Set $wgMathInternalRestbaseURL explicitly (T349582)]] * 14:57 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1265 * 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1265.eqiad.wmnet 75.32.64.10.in-addr.arpa 5.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:56 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1265.eqiad.wmnet 75.32.64.10.in-addr.arpa 5.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:56 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:56 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1265 - kamila@cumin1003" * 14:56 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1265 - kamila@cumin1003" * 14:53 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1002.eqiad.wmnet * 14:53 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1376-1384].eqiad.wmnet * 14:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1001.eqiad.wmnet with reason: host reimage * 14:51 kamila@cumin1003: START - Cookbook sre.dns.netbox * 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 14:50 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:50 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2161: switch maintenance completed codfw rack b6 * 14:50 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5021.eqsin.wmnet * 14:50 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1265 * 14:49 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:49 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:49 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1265.eqiad.wmnet with OS trixie * 14:49 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5029.eqsin.wmnet * 14:49 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1265.eqiad.wmnet * 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3070.esams.wmnet * 14:48 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1264.eqiad.wmnet * 14:48 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1001.eqiad.wmnet with reason: host reimage * 14:48 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1376-1384].eqiad.wmnet * 14:48 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1265.eqiad.wmnet * 14:47 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1265.eqiad.wmnet * 14:47 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1264.eqiad.wmnet * 14:47 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1264.eqiad.wmnet * 14:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:47 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3078.esams.wmnet * 14:44 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1263.eqiad.wmnet * 14:44 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1263.eqiad.wmnet * 14:44 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1263.eqiad.wmnet * 14:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1108.eqiad.wmnet * 14:40 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1109.eqiad.wmnet * 14:40 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:35 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:34 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:34 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2006.codfw.wmnet * 14:34 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-flink-eqiad cluster: Roll restart of jvm daemons. * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf2002.codfw.wmnet * 14:31 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1002.eqiad.wmnet * 14:29 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2006.codfw.wmnet * 14:27 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:27 kamila@deploy2003: Finished scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] (duration: 02m 57s) * 14:27 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-flink-eqiad cluster: Roll restart of jvm daemons. * 14:26 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf2002.codfw.wmnet * 14:26 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf2001.codfw.wmnet * 14:25 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1002.eqiad.wmnet * 14:25 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1001.eqiad.wmnet * 14:25 kamila@deploy2003: Started scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] * 14:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:21 kamila@deploy2003: sync-world aborted: Test deployment to check rsync is working - [[phab:T432108|T432108]] (duration: 00m 36s) * 14:21 topranks: reboot lsw1-b6-codfw to upgrade JunOS [[phab:T430922|T430922]] * 14:21 kamila@deploy2003: Started scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] * 14:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1001.eqiad.wmnet with OS bookworm * 14:20 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b6-codfw,lsw1-b6-codfw IPv6,lsw1-b6-codfw.mgmt,ssw1-a[1,8]-codfw with reason: lsw1-b6-codfw JunOS upgrade * 14:20 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf2001.codfw.wmnet * 14:19 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1001.eqiad.wmnet * 14:19 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 26 hosts with reason: lsw1-b6-codfw JunOS upgrade * 14:14 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 14:13 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc2022: switch maintenance codfw rack b6 * 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:12 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.parsercache * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool pc2022: switch maintenance codfw rack b6 * 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2251: switch maintenance codfw rack b6 * 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.parsercache * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2251: switch maintenance codfw rack b6 * 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2162: switch maintenance codfw rack b6 * 14:12 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1263.eqiad.wmnet with OS trixie * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2162: switch maintenance codfw rack b6 * 14:11 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2161: switch maintenance codfw rack b6 * 14:11 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2161: switch maintenance codfw rack b6 * 14:08 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5020.eqsin.wmnet * 14:07 btullis@cumin1003: START - Cookbook sre.hadoop.reboot-workers for Hadoop analytics cluster * 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3069.esams.wmnet * 14:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5028.eqsin.wmnet * 14:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1338-1347].eqiad.wmnet * 14:06 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1338-1347].eqiad.wmnet * 14:05 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3077.esams.wmnet * 14:02 topranks: beginning depools for lsw1-b6-codfw maintenance [[phab:T430922|T430922]] * 14:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1106.eqiad.wmnet * 14:00 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-misc1002.eqiad.wmnet * 13:59 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1338-1347].eqiad.wmnet * 13:59 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1107.eqiad.wmnet * 13:56 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-codfw * 13:55 sfaci@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply * 13:54 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-misc1002.eqiad.wmnet * 13:54 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-misc1001.eqiad.wmnet * 13:54 sfaci@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply * 13:50 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1263.eqiad.wmnet with reason: host reimage * 13:50 sfaci@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 13:49 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1338-1347].eqiad.wmnet * 13:49 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-misc1001.eqiad.wmnet * 13:49 sfaci@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 13:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:49 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:45 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1263.eqiad.wmnet with reason: host reimage * 13:40 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:40 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-eqiad * 13:35 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:34 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:33 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-reboot (exit_code=0) rolling reboot on A:dnsbox and (A:eqsin or A:drmrs or A:magru) and not (P<nowiki>{</nowiki>dns5003*<nowiki>}</nowiki> or P<nowiki>{</nowiki>dns7002*<nowiki>}</nowiki>) and (A:dnsbox) * 13:33 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns7001.wikimedia.org * 13:27 sukhe@dns1004: END - running authdns-update * 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5019.eqsin.wmnet * 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3076.esams.wmnet * 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3068.esams.wmnet * 13:25 sukhe@dns1004: START - running authdns-update * 13:24 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5027.eqsin.wmnet * 13:24 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1263 * 13:24 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1263 * 13:23 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1263 * 13:23 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:23 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1104.eqiad.wmnet * 13:21 kamila@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:21 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:20 kamila@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:20 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:20 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:20 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1263 - kamila@cumin1003" * 13:20 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1263 - kamila@cumin1003" * 13:19 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1105.eqiad.wmnet * 13:19 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:19 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:18 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:18 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns7001.wikimedia.org * 13:16 cdobbins@cumin2003: conftool action : set/pooled=yes; selector: name=dns7002.* * 13:14 cdobbins@dns1004: END - running authdns-update * 13:13 sbisson@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] (duration: 08m 03s) * 13:13 cdobbins@dns1004: START - running authdns-update * 13:12 kamila@cumin1003: START - Cookbook sre.dns.netbox * 13:12 cdobbins@cumin2003: conftool action : set/pooled=yes; selector: name=dns7002.*,service=authdns-update * 13:12 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1263 * 13:11 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1263.eqiad.wmnet with OS trixie * 13:11 cdobbins@cumin2003: conftool action : set/pooled=no; selector: name=dns7002.* * 13:11 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1263.eqiad.wmnet * 13:10 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1263.eqiad.wmnet * 13:10 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1263.eqiad.wmnet * 13:09 sbisson@deploy2003: sbisson: Continuing with deployment * 13:08 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:07 sbisson@deploy2003: sbisson: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:05 sbisson@deploy2003: Started scap sync-world: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] * 13:03 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns6002.wikimedia.org * 13:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1298-1307].eqiad.wmnet * 13:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1298-1307].eqiad.wmnet * 12:59 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1262.eqiad.wmnet * 12:59 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1262.eqiad.wmnet * 12:59 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1262.eqiad.wmnet * 12:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1298-1307].eqiad.wmnet * 12:49 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns6002.wikimedia.org * 12:46 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1298-1307].eqiad.wmnet * 12:46 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:46 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3075.esams.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3067.esams.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5018.eqsin.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5026.eqsin.wmnet * 12:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1102.eqiad.wmnet * 12:39 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1103.eqiad.wmnet * 12:35 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:34 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns6001.wikimedia.org * 12:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:28 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:18 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns6001.wikimedia.org * 12:14 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1267-1276].eqiad.wmnet * 12:13 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1267-1276].eqiad.wmnet * 12:04 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1267-1276].eqiad.wmnet * 12:03 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns5004.wikimedia.org * 12:02 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1100.eqiad.wmnet * 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3066.esams.wmnet * 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3074.esams.wmnet * 12:01 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:01 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-codfw * 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5017.eqsin.wmnet * 12:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5025.eqsin.wmnet * 12:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1101.eqiad.wmnet * 11:59 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1267-1276].eqiad.wmnet * 11:58 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:58 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:54 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns5004.wikimedia.org * 11:54 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and (A:eqsin or A:drmrs or A:magru) and not (P<nowiki>{</nowiki>dns5003*<nowiki>}</nowiki> or P<nowiki>{</nowiki>dns7002*<nowiki>}</nowiki>) and (A:dnsbox) * 11:54 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:53 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:53 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-eqiad * 11:51 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:51 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:50 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:50 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_eqiad * 11:50 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:50 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_eqiad * 11:49 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_esams * 11:49 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_esams * 11:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_eqsin * 11:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_eqsin * 11:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:44 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-eqiad * 11:43 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-codfw * 11:42 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:41 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:24 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-eqiad * 11:23 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-codfw * 11:22 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:15 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:14 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:09 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2066.codfw.wmnet with OS trixie * 11:05 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1151-1160].eqiad.wmnet * 11:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1151-1160].eqiad.wmnet * 10:59 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1068.eqiad.wmnet with OS trixie * 10:55 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.major-upgrade (exit_code=99) * 10:55 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 10:54 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1151-1160].eqiad.wmnet * 10:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2066.codfw.wmnet with reason: host reimage * 10:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1151-1160].eqiad.wmnet * 10:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:42 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2066.codfw.wmnet with reason: host reimage * 10:39 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:37 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:36 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:23 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:22 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2066.codfw.wmnet with OS trixie * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:07 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2065.codfw.wmnet with OS trixie * 10:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:06 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 10:06 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 10:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 10:03 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:03 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 10:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 09:59 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 09:57 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 09:57 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 09:52 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 09:47 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:46 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:46 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2065.codfw.wmnet with reason: host reimage * 09:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2065.codfw.wmnet with reason: host reimage * 09:40 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:39 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 09:39 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:39 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:39 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:37 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:29 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox2003.codfw.wmnet * 09:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox2003.codfw.wmnet * 09:25 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:25 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:24 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:24 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:24 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:21 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:20 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2065.codfw.wmnet with OS trixie * 09:13 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1068.eqiad.wmnet with OS trixie * 09:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2064.codfw.wmnet with OS trixie * 09:08 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 09:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 09:07 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2162: Repooling after switchover * 09:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1067.eqiad.wmnet with OS trixie * 09:00 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:59 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 08:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 08:57 tappof: bump space for prometheus k8s-dse in eqiad * 08:56 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping2004.codfw.wmnet * 08:52 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host ping2004.codfw.wmnet * 08:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 08:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping1004.eqiad.wmnet * 08:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:51 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:49 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 08:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host ping1004.eqiad.wmnet * 08:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2064.codfw.wmnet with reason: host reimage * 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2064.codfw.wmnet with reason: host reimage * 08:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1067.eqiad.wmnet with reason: host reimage * 08:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:33 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1067.eqiad.wmnet with reason: host reimage * 08:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:21 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2162: Repooling after switchover * 08:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2064.codfw.wmnet with OS trixie * 08:16 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1067.eqiad.wmnet with OS trixie * 08:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:15 cgoubert@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-eqiad * 08:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2062.codfw.wmnet with OS trixie * 08:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1066.eqiad.wmnet with OS trixie * 08:02 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2162: Repooling after switchover * 07:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2162: Repooling after switchover * 07:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2162 [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94870 and previous config saved to /var/cache/conftool/dbconfig/20260716-075530-cwilliams.json * 07:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2241 to x3 primary [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94869 and previous config saved to /var/cache/conftool/dbconfig/20260716-075314-cwilliams.json * 07:52 cezmunsta: Starting x3 codfw failover from db2162 to db2241 - [[phab:T430925|T430925]] * 07:50 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:50 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:47 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 07:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2241 with weight 0 [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94868 and previous config saved to /var/cache/conftool/dbconfig/20260716-074507-cwilliams.json * 07:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 18 hosts with reason: Primary switchover x3 [[phab:T430925|T430925]] * 07:43 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1066.eqiad.wmnet with reason: host reimage * 07:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:dse-k8s-worker-eqiad * 07:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1028.eqiad.wmnet * 07:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1028.eqiad.wmnet * 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 07:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1066.eqiad.wmnet with reason: host reimage * 07:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1028.eqiad.wmnet * 07:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1028.eqiad.wmnet * 07:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1027.eqiad.wmnet * 07:35 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1027.eqiad.wmnet * 07:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1027.eqiad.wmnet * 07:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1027.eqiad.wmnet * 07:28 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1026.eqiad.wmnet * 07:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1026.eqiad.wmnet * 07:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast2003.wikimedia.org * 07:21 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1026.eqiad.wmnet * 07:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1066.eqiad.wmnet with OS trixie * 07:19 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast2003.wikimedia.org * 07:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2062.codfw.wmnet with OS trixie * 06:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1026.eqiad.wmnet * 06:51 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1025.eqiad.wmnet * 06:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1025.eqiad.wmnet * 06:47 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 06:47 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 06:44 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1025.eqiad.wmnet * 06:14 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1025.eqiad.wmnet * 06:14 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1024.eqiad.wmnet * 06:14 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1024.eqiad.wmnet * 06:07 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1024.eqiad.wmnet * 05:37 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1024.eqiad.wmnet * 05:37 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1023.eqiad.wmnet * 05:37 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1023.eqiad.wmnet * 05:26 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1023.eqiad.wmnet * 04:56 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1023.eqiad.wmnet * 04:56 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1022.eqiad.wmnet * 04:56 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1022.eqiad.wmnet * 04:49 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1022.eqiad.wmnet * 04:19 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1022.eqiad.wmnet * 04:19 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1021.eqiad.wmnet * 04:19 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1021.eqiad.wmnet * 04:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1021.eqiad.wmnet * 03:38 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1021.eqiad.wmnet * 03:38 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1020.eqiad.wmnet * 03:38 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1020.eqiad.wmnet * 03:20 btullis@cumin1003: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1020.eqiad.wmnet * 03:18 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1020.eqiad.wmnet * 03:18 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1019.eqiad.wmnet * 03:18 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1019.eqiad.wmnet * 03:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1019.eqiad.wmnet * 02:41 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1019.eqiad.wmnet * 02:41 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1018.eqiad.wmnet * 02:41 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1018.eqiad.wmnet * 02:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling both afterwards * 02:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2003.codfw.wmnet -> wcqs2001.codfw.wmnet, repooling both afterwards * 02:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1018.eqiad.wmnet * 02:30 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1018.eqiad.wmnet * 02:30 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1014.eqiad.wmnet * 02:30 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1014.eqiad.wmnet * 02:24 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1014.eqiad.wmnet * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 01:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1014.eqiad.wmnet * 01:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1013.eqiad.wmnet * 01:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1013.eqiad.wmnet * 01:47 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1013.eqiad.wmnet * 01:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2003.codfw.wmnet -> wcqs2001.codfw.wmnet, repooling both afterwards * 01:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling both afterwards * 01:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1013.eqiad.wmnet * 01:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1012.eqiad.wmnet * 01:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1012.eqiad.wmnet * 01:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1012.eqiad.wmnet * 01:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1012.eqiad.wmnet * 01:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1011.eqiad.wmnet * 01:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1011.eqiad.wmnet * 01:08 ryankemper@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] scap deploy post bookworm reimage (duration: 00m 23s) * 01:08 ryankemper@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] scap deploy post bookworm reimage * 01:08 ryankemper@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): scap deploy post bookworm reimage (duration: 00m 46s) * 01:07 ryankemper@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): scap deploy post bookworm reimage * 01:04 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1011.eqiad.wmnet * 01:04 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1011.eqiad.wmnet * 01:04 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1010.eqiad.wmnet * 01:04 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1010.eqiad.wmnet * 00:57 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1010.eqiad.wmnet * 00:57 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1010.eqiad.wmnet * 00:57 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1009.eqiad.wmnet * 00:57 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1009.eqiad.wmnet * 00:50 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1009.eqiad.wmnet * 00:20 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1009.eqiad.wmnet * 00:20 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1008.eqiad.wmnet * 00:20 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1008.eqiad.wmnet * 00:13 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1008.eqiad.wmnet == 2026-07-15 == * 23:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2001.codfw.wmnet with OS bookworm * 23:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1008.eqiad.wmnet * 23:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1007.eqiad.wmnet * 23:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1007.eqiad.wmnet * 23:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1007.eqiad.wmnet * 23:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1007.eqiad.wmnet * 23:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1006.eqiad.wmnet * 23:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1006.eqiad.wmnet * 23:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1006.eqiad.wmnet * 23:29 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1006.eqiad.wmnet * 23:28 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1005.eqiad.wmnet * 23:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1005.eqiad.wmnet * 23:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1002.eqiad.wmnet with OS bookworm * 23:21 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1005.eqiad.wmnet * 23:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2001.codfw.wmnet with reason: host reimage * 23:15 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host datahubsearch1001.eqiad.wmnet with OS bookworm * 23:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2001.codfw.wmnet with reason: host reimage * 23:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 23:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 22:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 22:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1005.eqiad.wmnet * 22:51 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1004.eqiad.wmnet * 22:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1004.eqiad.wmnet * 22:45 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1004.eqiad.wmnet * 22:44 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 22:44 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS trixie * 22:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host datahubsearch1001.eqiad.wmnet with OS bookworm * 22:34 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host datahubsearch1001.eqiad.wmnet with OS bookworm * 22:16 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on datahubsearch[1002-1003].eqiad.wmnet with reason: Using datahubsearch1001 to test bookworm reimages * 22:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1004.eqiad.wmnet * 22:15 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1003.eqiad.wmnet * 22:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1003.eqiad.wmnet * 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 22:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1003.eqiad.wmnet * 22:08 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1003.eqiad.wmnet * 22:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1002.eqiad.wmnet * 22:08 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1002.eqiad.wmnet * 22:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host datahubsearch1001.eqiad.wmnet with OS bookworm * 22:05 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 22:02 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm * 22:01 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on datahubsearch[1001-1003].eqiad.wmnet with reason: Using datahubsearch1001 to test bookworm reimages * 22:01 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1002.eqiad.wmnet * 22:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1002.eqiad.wmnet * 22:00 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1001.eqiad.wmnet * 22:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1001.eqiad.wmnet * 21:53 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1001.eqiad.wmnet * 21:52 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 21:50 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 21:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS trixie * 21:50 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS bookworm * 21:43 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 21:38 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:30 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wcqs1002'] * 21:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:29 lerickson@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 21:29 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:29 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:29 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS bookworm * 21:28 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 21:28 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm * 21:23 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1001.eqiad.wmnet * 21:23 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:23 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:22 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 21:20 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:18 swfrench-wmf: reprepro include php8.3_8.3.32-1+wmf11u2 into component/php83 for bullseye-wikimedia * 21:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:16 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:15 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-druid-public cluster: Roll restart of jvm daemons. * 21:08 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:05 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1001.eqiad.wmnet * 21:05 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1001.eqiad.wmnet * 21:04 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-druid-public cluster: Roll restart of jvm daemons. * 21:02 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 21:01 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 21:01 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 21:00 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 20:59 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1001.eqiad.wmnet * 20:59 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1001.eqiad.wmnet * 20:59 btullis@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:dse-k8s-worker-eqiad * 20:55 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 20:55 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 20:45 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm * 20:21 jhathaway: puppet is re-enabled, have fun, but not too much fun! * 20:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 20:17 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs2001'] * 20:12 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs2001'] * 20:11 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs2001'] * 20:09 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:08 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 20:05 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:05 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 20:04 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs2001'] * 20:03 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 20:03 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm * 20:02 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:02 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 20:01 jhathaway: disabling puppet fleet wide to roll out kafka patch * 19:55 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host relforge1010.eqiad.wmnet * 19:52 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 19:52 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 19:48 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 19:48 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 19:48 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 19:47 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 19:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 19:45 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 19:45 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 19:44 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1010.eqiad.wmnet * 19:38 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1262.eqiad.wmnet with OS trixie * 19:17 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1262.eqiad.wmnet with reason: host reimage * 19:11 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1262.eqiad.wmnet with reason: host reimage * 18:59 cdobbins@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS trixie * 18:54 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 18:53 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 18:52 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1262 * 18:52 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1262 * 18:51 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1262 * 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1262.eqiad.wmnet 72.32.64.10.in-addr.arpa 2.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:51 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1262.eqiad.wmnet 72.32.64.10.in-addr.arpa 2.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1262 - kamila@cumin1003" * 18:51 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1262 - kamila@cumin1003" * 18:46 kamila@cumin1003: START - Cookbook sre.dns.netbox * 18:46 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1262 * 18:46 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 18:46 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ncmonitor1001.eqiad.wmnet * 18:46 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 18:45 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1262.eqiad.wmnet with OS trixie * 18:45 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 18:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1262.eqiad.wmnet * 18:44 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1262.eqiad.wmnet * 18:44 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1262.eqiad.wmnet * 18:42 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host ncmonitor1001.eqiad.wmnet * 18:29 topranks: pull power on cr1-eqiad to install new switch-control boards [[phab:T426343|T426343]] * 18:29 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs[1018-1020].eqiad.wmnet with reason: line card install in cr1-eqiad * 18:27 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 14 hosts with reason: linecard install in cr1-eqad * 18:22 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_ulsfo * 18:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4052.ulsfo.wmnet * 18:19 cdobbins@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 18:15 cdobbins@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 18:14 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_drmrs * 18:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6016.drmrs.wmnet * 18:12 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_ulsfo * 18:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4044.ulsfo.wmnet * 18:10 sukhe@cumin1003: END (ERROR) - Cookbook sre.cdn.roll-reboot (exit_code=97) rolling reboot on A:cp-upload_drmrs * 18:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2241: Security update * 17:56 topranks: start draining traffic on cr1-eqiad ahead of line card installation [[phab:T426343|T426343]] * 17:47 cdobbins@cumin2003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie * 17:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4051.ulsfo.wmnet * 17:40 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:39 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 17:34 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6007.drmrs.wmnet * 17:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6015.drmrs.wmnet * 17:32 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:31 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 17:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4043.ulsfo.wmnet * 17:27 lerickson@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:25 lerickson@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 17:22 lerickson@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-codfw * 17:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp2001.codfw.wmnet * 17:22 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 17:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp2001.codfw.wmnet * 17:22 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 17:19 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2241: Security update * 17:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2241.codfw.wmnet * 17:17 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2241.codfw.wmnet * 17:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp2001.codfw.wmnet * 17:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp2001.codfw.wmnet * 17:15 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2366-2374].codfw.wmnet * 17:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2366-2374].codfw.wmnet * 17:10 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply * 17:10 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply * 17:08 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2366-2374].codfw.wmnet * 17:06 sukhe: sre.dns.roll-reboot to resume later * 17:06 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-reboot (exit_code=97) rolling reboot on A:dnsbox and not (A:ulsfo or A:magru) and (A:dnsbox) * 17:06 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns5003.wikimedia.org * 17:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2241: Security update * 17:03 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2241: Security update * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply * 17:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2366-2374].codfw.wmnet * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply * 17:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2357-2365].codfw.wmnet * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 17:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2357-2365].codfw.wmnet * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply * 16:55 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2357-2365].codfw.wmnet * 16:55 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 16:53 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6006.drmrs.wmnet * 16:52 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6014.drmrs.wmnet * 16:52 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 16:51 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 16:50 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2357-2365].codfw.wmnet * 16:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4042.ulsfo.wmnet * 16:50 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2347-2356].codfw.wmnet * 16:50 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2347-2356].codfw.wmnet * 16:49 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns5003.wikimedia.org * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply * 16:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4050.ulsfo.wmnet * 16:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2347-2356].codfw.wmnet * 16:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2347-2356].codfw.wmnet * 16:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2337-2346].codfw.wmnet * 16:36 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2337-2346].codfw.wmnet * 16:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:dse-k8s-worker-codfw * 16:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2003.codfw.wmnet * 16:35 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2003.codfw.wmnet * 16:34 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns3004.wikimedia.org * 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply * 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply * 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply * 16:30 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply * 16:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2003.codfw.wmnet * 16:29 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2337-2346].codfw.wmnet * 16:24 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2003.codfw.wmnet * 16:24 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2002.codfw.wmnet * 16:24 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2002.codfw.wmnet * 16:23 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns3004.wikimedia.org * 16:23 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2337-2346].codfw.wmnet * 16:23 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2327-2336].codfw.wmnet * 16:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2327-2336].codfw.wmnet * 16:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2002.codfw.wmnet * 16:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2327-2336].codfw.wmnet * 16:12 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2002.codfw.wmnet * 16:12 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2001.codfw.wmnet * 16:12 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2001.codfw.wmnet * 16:12 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1065.eqiad.wmnet with OS trixie * 16:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6005.drmrs.wmnet * 16:11 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6013.drmrs.wmnet * 16:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4041.ulsfo.wmnet * 16:08 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns3003.wikimedia.org * 16:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2327-2336].codfw.wmnet * 16:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2317-2326].codfw.wmnet * 16:06 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2317-2326].codfw.wmnet * 16:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2001.codfw.wmnet * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply * 16:03 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4049.ulsfo.wmnet * 16:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2001.codfw.wmnet * 16:00 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test2001.codfw.wmnet * 16:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test2001.codfw.wmnet * 16:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2063.codfw.wmnet with OS trixie * 15:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2317-2326].codfw.wmnet * 15:57 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns3003.wikimedia.org * 15:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test2001.codfw.wmnet * 15:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test2001.codfw.wmnet * 15:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2004.codfw.wmnet * 15:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2004.codfw.wmnet * 15:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2317-2326].codfw.wmnet * 15:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2307-2316].codfw.wmnet * 15:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2307-2316].codfw.wmnet * 15:49 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2004.codfw.wmnet * 15:48 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2004.codfw.wmnet * 15:48 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2003.codfw.wmnet * 15:48 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2003.codfw.wmnet * 15:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 15:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2307-2316].codfw.wmnet * 15:42 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2003.codfw.wmnet * 15:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 15:42 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2003.codfw.wmnet * 15:42 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2002.codfw.wmnet * 15:42 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2002.codfw.wmnet * 15:42 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2006.wikimedia.org * 15:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2063.codfw.wmnet with reason: host reimage * 15:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2307-2316].codfw.wmnet * 15:37 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2297-2306].codfw.wmnet * 15:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2297-2306].codfw.wmnet * 15:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2002.codfw.wmnet * 15:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2002.codfw.wmnet * 15:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2001.codfw.wmnet * 15:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2001.codfw.wmnet * 15:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2063.codfw.wmnet with reason: host reimage * 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6004.drmrs.wmnet * 15:31 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2001.codfw.wmnet * 15:31 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2001.codfw.wmnet * 15:31 btullis@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:dse-k8s-worker-codfw * 15:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6012.drmrs.wmnet * 15:28 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2006.wikimedia.org * 15:27 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-analytics cluster: Roll restart of jvm daemons. * 15:27 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2297-2306].codfw.wmnet * 15:27 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4040.ulsfo.wmnet * 15:24 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1065.eqiad.wmnet with OS trixie * 15:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4048.ulsfo.wmnet * 15:21 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-analytics cluster: Roll restart of jvm daemons. * 15:21 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2297-2306].codfw.wmnet * 15:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2287-2296].codfw.wmnet * 15:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2287-2296].codfw.wmnet * 15:20 btullis@cumin1003: END (PASS) - Cookbook sre.druid.reboot-workers (exit_code=0) for Druid public cluster: Reboot Druid nodes * 15:18 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 15:17 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm * 15:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2063.codfw.wmnet with OS trixie * 15:13 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2005.wikimedia.org * 15:11 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2287-2296].codfw.wmnet * 15:11 btullis@cumin1003: END (PASS) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=0) rolling reboot on A:cephosd-eqiad * 15:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1064.eqiad.wmnet with OS trixie * 15:05 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2062.codfw.wmnet with OS trixie * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2287-2296].codfw.wmnet * 15:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2277-2286].codfw.wmnet * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2277-2286].codfw.wmnet * 14:59 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2005.wikimedia.org * 14:57 brouberol@cumin1003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-jumbo-eqiad * 14:52 btullis@cumin1003: END (PASS) - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas (exit_code=0) rolling reboot on A:schema-codfw * 14:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6003.drmrs.wmnet * 14:50 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:50 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host relforge1009.eqiad.wmnet * 14:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2277-2286].codfw.wmnet * 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6011.drmrs.wmnet * 14:47 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 14:46 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>ml-serve1001.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 14:46 klausman@cumin1003: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) pool for host ml-serve1001.eqiad.wmnet * 14:46 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 14:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1001.eqiad.wmnet * 14:45 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4039.ulsfo.wmnet * 14:44 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1009.eqiad.wmnet * 14:44 btullis@cumin1003: START - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas rolling reboot on A:schema-codfw * 14:44 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2004.wikimedia.org * 14:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2277-2286].codfw.wmnet * 14:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2267-2276].codfw.wmnet * 14:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2267-2276].codfw.wmnet * 14:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 14:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4047.ulsfo.wmnet * 14:40 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1001.eqiad.wmnet * 14:38 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 14:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 14:36 topranks: disconnect power on cr2-eqiad to shut down device for switch fabric replacement [[phab:T426343|T426343]] * 14:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2267-2276].codfw.wmnet * 14:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 14:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1001.eqiad.wmnet * 14:35 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>ml-serve1001.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 14:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 14:34 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 14:33 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:33 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:30 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2004.wikimedia.org * 14:29 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2267-2276].codfw.wmnet * 14:29 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2257-2266].codfw.wmnet * 14:29 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2257-2266].codfw.wmnet * 14:24 btullis@cumin1003: END (PASS) - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas (exit_code=0) rolling reboot on A:schema-eqiad * 14:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2257-2266].codfw.wmnet * 14:20 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:20 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:19 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:17 jforrester@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2257-2266].codfw.wmnet * 14:16 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:16 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2062.codfw.wmnet with OS trixie * 14:15 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1064.eqiad.wmnet with OS trixie * 14:15 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1006.wikimedia.org * 14:15 btullis@cumin1003: START - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas rolling reboot on A:schema-eqiad * 14:14 topranks: switch routing-engine on cr2-eqiad resetting all interfaces [[phab:T417873|T417873]] * 14:11 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:11 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:10 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm * 14:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6002.drmrs.wmnet * 14:09 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6010.drmrs.wmnet * 14:06 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1006.wikimedia.org * 14:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:05 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 14:05 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4038.ulsfo.wmnet * 14:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4046.ulsfo.wmnet * 14:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:00 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on cr1-eqiad with reason: switch upgrade and line card install * 14:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:59 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:57 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:57 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:55 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-eqiad * 13:55 btullis@cumin1003: START - Cookbook sre.druid.reboot-workers for Druid public cluster: Reboot Druid nodes * 13:53 topranks: switch routing-engine on cr2-eqiad resetting all interfaces [[phab:T417873|T417873]] * 13:51 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1005.wikimedia.org * 13:50 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:49 brouberol@cumin1003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-test-eqiad * 13:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:44 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2197-2206].codfw.wmnet * 13:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2197-2206].codfw.wmnet * 13:36 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1005.wikimedia.org * 13:35 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2197-2206].codfw.wmnet * 13:30 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2197-2206].codfw.wmnet * 13:28 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6001.drmrs.wmnet * 13:28 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6009.drmrs.wmnet * 13:28 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2187-2196].codfw.wmnet * 13:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2187-2196].codfw.wmnet * 13:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 13:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4037.ulsfo.wmnet * 13:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2001 * 13:22 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2001 * 13:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4045.ulsfo.wmnet * 13:21 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1004.wikimedia.org * 13:19 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on lvs[1018-1020].eqiad.wmnet with reason: switch upgrade and line card install * 13:18 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2009.codfw.wmnet * 13:18 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2001 * 13:18 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2001.codfw.wmnet 26.16.192.10.in-addr.arpa 6.2.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:17 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2001.codfw.wmnet 26.16.192.10.in-addr.arpa 6.2.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:17 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:17 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2001 - bking@cumin2003" * 13:17 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2001 - bking@cumin2003" * 13:17 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2009.codfw.wmnet * 13:17 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_drmrs * 13:17 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2187-2196].codfw.wmnet * 13:17 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_drmrs * 13:17 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on 15 hosts with reason: switch upgrade and line card install * 13:17 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:15 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 13:13 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:13 brouberol@cumin1003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-jumbo-eqiad * 13:13 brouberol@cumin1003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-test-eqiad * 13:13 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1004.wikimedia.org * 13:13 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and not (A:ulsfo or A:magru) and (A:dnsbox) * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:12 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_ulsfo * 13:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:12 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_ulsfo * 13:11 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2187-2196].codfw.wmnet * 13:11 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 13:11 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 13:06 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 13:05 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling source-only afterwards * 13:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2001 * 13:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2009.codfw.wmnet with OS trixie * 13:03 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:03 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling source-only afterwards * 13:01 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 15s) * 13:01 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 13:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 12:57 btullis@cumin1003: END (PASS) - Cookbook sre.druid.reboot-workers (exit_code=0) for Druid analytics cluster: Reboot Druid nodes * 12:54 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 12:54 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2163-2172].codfw.wmnet * 12:54 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2163-2172].codfw.wmnet * 12:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2163-2172].codfw.wmnet * 12:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2009.codfw.wmnet with reason: host reimage * 12:41 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2163-2172].codfw.wmnet * 12:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2153-2162].codfw.wmnet * 12:40 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2153-2162].codfw.wmnet * 12:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2009.codfw.wmnet with reason: host reimage * 12:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2153-2162].codfw.wmnet * 12:29 btullis@cumin1003: END (PASS) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=0) rolling reboot on A:cephosd-codfw * 12:25 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2153-2162].codfw.wmnet * 12:25 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2143-2152].codfw.wmnet * 12:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2143-2152].codfw.wmnet * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2009 * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2009 * 12:22 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2009 * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2009.codfw.wmnet 139.0.192.10.in-addr.arpa 9.3.1.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:22 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2009.codfw.wmnet 139.0.192.10.in-addr.arpa 9.3.1.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2009 - mvernon@cumin2003" * 12:22 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2009 - mvernon@cumin2003" * 12:16 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 12:15 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 12:15 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 12:15 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2009 * 12:15 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 12:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2009.codfw.wmnet with OS trixie * 12:15 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 12:14 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 12:14 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2143-2152].codfw.wmnet * 12:13 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 12:12 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2010.codfw.wmnet * 12:11 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2010.codfw.wmnet * 12:10 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 12:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2143-2152].codfw.wmnet * 12:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2133-2142].codfw.wmnet * 12:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2133-2142].codfw.wmnet * 12:02 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 11:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2133-2142].codfw.wmnet * 11:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2133-2142].codfw.wmnet * 11:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:49 mvolz@deploy2003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:49 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-codfw * 11:48 mvolz@deploy2003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:47 btullis@cumin1003: START - Cookbook sre.druid.reboot-workers for Druid analytics cluster: Reboot Druid nodes * 11:46 mvolz@deploy2003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:46 mvolz@deploy2003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:45 mvolz@deploy2003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:44 mvolz@deploy2003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:40 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] (duration: 11m 38s) * 11:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2010.codfw.wmnet with OS trixie * 11:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2105-2114].codfw.wmnet * 11:36 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2105-2114].codfw.wmnet * 11:36 krinkle@deploy2003: physikerwelt, krinkle: Continuing with deployment * 11:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1018: Security updates * 11:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:36 root@cumin1003: START - Cookbook sre.mysql.parsercache * 11:36 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1018: Security updates * 11:31 krinkle@deploy2003: physikerwelt, krinkle: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:29 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] * 11:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2105-2114].codfw.wmnet * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2105-2114].codfw.wmnet * 11:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2010.codfw.wmnet with reason: host reimage * 11:12 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2010.codfw.wmnet with reason: host reimage * 11:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1018: Security updates * 11:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:10 root@cumin1003: START - Cookbook sre.mysql.parsercache * 11:10 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1018: Security updates * 11:09 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1009.eqiad.wmnet with OS trixie * 11:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow7002.magru.wmnet * 11:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 11:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 11:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=tegola-vector-tiles,name=eqiad * 11:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=kartotherian,name=eqiad * 11:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow7002.magru.wmnet * 10:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2010 * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2010 * 10:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 10:54 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 10:54 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2010 * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2010.codfw.wmnet 76.16.192.10.in-addr.arpa 6.7.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:54 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2010.codfw.wmnet 76.16.192.10.in-addr.arpa 6.7.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2010 - mvernon@cumin2003" * 10:54 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2010 - mvernon@cumin2003" * 10:49 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 10:49 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2010 * 10:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1009.eqiad.wmnet with reason: host reimage * 10:49 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2010.codfw.wmnet with OS trixie * 10:46 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2011.codfw.wmnet * 10:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1011.eqiad.wmnet * 10:44 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2011.codfw.wmnet * 10:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 10:44 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 10:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1009.eqiad.wmnet with reason: host reimage * 10:44 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow6001.drmrs.wmnet * 10:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1017: Security updates * 10:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:39 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1017: Security updates * 10:39 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow6001.drmrs.wmnet * 10:38 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1011.eqiad.wmnet * 10:35 cgoubert@deploy2003: Finished deploy [restbase/deploy@06301bd]: Deploying {{Gerrit|1306088}} {{Gerrit|1308347}} - [[phab:T429944|T429944]] [[phab:T428279|T428279]] (duration: 28m 34s) * 10:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1012.eqiad.wmnet * 10:35 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 10:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow5003.eqsin.wmnet * 10:34 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2011.codfw.wmnet with OS trixie * 10:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1009.eqiad.wmnet with OS trixie * 10:28 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1012.eqiad.wmnet * 10:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1013.eqiad.wmnet * 10:27 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:27 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow5003.eqsin.wmnet * 10:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:26 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow4003.ulsfo.wmnet * 10:25 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 10:25 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 10:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow4003.ulsfo.wmnet * 10:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1013.eqiad.wmnet * 10:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1014.eqiad.wmnet * 10:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:15 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2011.codfw.wmnet with reason: host reimage * 10:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1017: Security updates * 10:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:14 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:14 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1017: Security updates * 10:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow3004.esams.wmnet * 10:11 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2011.codfw.wmnet with reason: host reimage * 10:10 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:10 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1014.eqiad.wmnet * 10:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki2003.codfw.wmnet * 10:09 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 10:09 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 10:09 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow3004.esams.wmnet * 10:08 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2004.codfw.wmnet * 10:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1010.eqiad.wmnet with OS trixie * 10:07 cgoubert@deploy2003: Started deploy [restbase/deploy@06301bd]: Deploying {{Gerrit|1306088}} {{Gerrit|1308347}} - [[phab:T429944|T429944]] [[phab:T428279|T428279]] * 10:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host rpki2003.codfw.wmnet * 10:04 topranks: push out config change to BGP_outfilter on core routers [[phab:T431849|T431849]] * 10:02 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow2004.codfw.wmnet * 09:59 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 09:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2003.codfw.wmnet * 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2011 * 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2011 * 09:53 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 09:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:52 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2011 * 09:52 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2011.codfw.wmnet 36.32.192.10.in-addr.arpa 6.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:52 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2011.codfw.wmnet 36.32.192.10.in-addr.arpa 6.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:51 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:51 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2011 - mvernon@cumin2003" * 09:51 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2011 - mvernon@cumin2003" * 09:51 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow2003.codfw.wmnet * 09:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1003.eqiad.wmnet * 09:49 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 09:49 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 09:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1010.eqiad.wmnet with reason: host reimage * 09:47 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 09:47 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 09:47 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 09:47 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2011 * 09:46 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2011.codfw.wmnet with OS trixie * 09:44 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow1003.eqiad.wmnet * 09:44 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2012.codfw.wmnet * 09:44 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1002.eqiad.wmnet * 09:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1010.eqiad.wmnet with reason: host reimage * 09:43 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2012.codfw.wmnet * 09:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Security updates * 09:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:43 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:43 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Security updates * 09:42 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:40 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow1002.eqiad.wmnet * 09:40 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 09:37 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki1001.eqiad.wmnet * 09:36 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 09:36 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 09:33 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host rpki1001.eqiad.wmnet * 09:32 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:32 cgoubert@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-codfw * 09:31 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=kartotherian,name=eqiad * 09:31 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola-vector-tiles,name=eqiad * 09:31 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 09:31 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2012.codfw.wmnet with OS trixie * 09:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1010.eqiad.wmnet with OS trixie * 09:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1011.eqiad.wmnet with OS trixie * 09:21 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Security updates * 09:21 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:21 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:21 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Security updates * 09:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2012.codfw.wmnet with reason: host reimage * 09:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1011.eqiad.wmnet with reason: host reimage * 09:08 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2012.codfw.wmnet with reason: host reimage * 09:05 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1011.eqiad.wmnet with reason: host reimage * 08:55 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:52 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1011.eqiad.wmnet with OS trixie * 08:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1022: Security updates * 08:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2012 * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2012 * 08:50 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1022: Security updates * 08:50 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2012 * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2012.codfw.wmnet 44.48.192.10.in-addr.arpa 4.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:50 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2012.codfw.wmnet 44.48.192.10.in-addr.arpa 4.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2012 - mvernon@cumin2003" * 08:50 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2012 - mvernon@cumin2003" * 08:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1012.eqiad.wmnet with OS trixie * 08:44 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 08:44 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2012 * 08:43 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2012.codfw.wmnet with OS trixie * 08:42 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2013.codfw.wmnet * 08:41 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2013.codfw.wmnet * 08:35 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 08:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host krb1002.eqiad.wmnet * 08:30 elukey@dns1004: END - running authdns-update * 08:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1012.eqiad.wmnet with reason: host reimage * 08:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Security updates * 08:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:28 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:28 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Security updates * 08:27 elukey@dns1004: START - running authdns-update * 08:26 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 08:26 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host krb1002.eqiad.wmnet * 08:22 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1012.eqiad.wmnet with reason: host reimage * 08:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host krb2002.codfw.wmnet * 08:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast6003.wikimedia.org * 08:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2013.codfw.wmnet with OS trixie * 08:13 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast6003.wikimedia.org * 08:12 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast3007.wikimedia.org * 08:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host krb2002.codfw.wmnet * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Security updates * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:09 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:09 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Security updates * 08:07 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1012.eqiad.wmnet with OS trixie * 08:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast3007.wikimedia.org * 08:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast5005.wikimedia.org * 07:58 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast5005.wikimedia.org * 07:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1013.eqiad.wmnet with OS trixie * 07:53 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2013.codfw.wmnet with reason: host reimage * 07:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1021: Security updates * 07:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:53 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:53 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1021: Security updates * 07:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast1004.wikimedia.org * 07:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2013.codfw.wmnet with reason: host reimage * 07:46 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast1004.wikimedia.org * 07:40 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1013.eqiad.wmnet with reason: host reimage * 07:36 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1013.eqiad.wmnet with reason: host reimage * 07:31 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2013 * 07:31 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2013 * 07:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1021: Security updates * 07:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:30 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:30 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1021: Security updates * 07:24 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2013 * 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2013.codfw.wmnet 87.0.192.10.in-addr.arpa 7.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:24 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2013.codfw.wmnet 87.0.192.10.in-addr.arpa 7.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2013 - mvernon@cumin2003" * 07:24 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2013 - mvernon@cumin2003" * 07:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1013.eqiad.wmnet with OS trixie * 07:19 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 07:19 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2013 * 07:19 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2013.codfw.wmnet with OS trixie * 07:13 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] (duration: 07m 48s) * 07:09 kharlan@deploy2003: kharlan: Continuing with deployment * 07:08 kharlan@deploy2003: kharlan: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:06 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 01:15 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 01:14 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply == 2026-07-14 == * 22:51 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_magru * 22:51 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7016.magru.wmnet * 22:46 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_magru * 22:46 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7008.magru.wmnet * 22:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7015.magru.wmnet * 22:04 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7007.magru.wmnet * 21:29 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7014.magru.wmnet * 21:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7006.magru.wmnet * 21:13 dzahn@cumin2002: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 0:15:00 on gerrit.wikimedia.org with reason: reboot * 21:11 mutante: gerrit2003 (gerrit.wikimedia.org) - reboot for maintenance * 21:11 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on gerrit2003.wikimedia.org with reason: reboot * 20:56 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:56 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:56 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:55 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 20:48 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7013.magru.wmnet * 20:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7005.magru.wmnet * 20:41 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host phab1005.eqiad.wmnet with OS trixie * 20:28 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] (duration: 06m 47s) * 20:24 sbassett@deploy2003: sbassett: Continuing with deployment * 20:23 sbassett@deploy2003: sbassett: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:23 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on phab1005.eqiad.wmnet with reason: host reimage * 20:21 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] * 20:20 aokoth@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on phab1005.eqiad.wmnet with reason: host reimage * 20:12 jhuneidi@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] (duration: 07m 42s) * 20:07 jhuneidi@deploy2003: jhuneidi, priyankar22: Continuing with deployment * 20:06 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7012.magru.wmnet * 20:06 jhuneidi@deploy2003: jhuneidi, priyankar22: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:04 jhuneidi@deploy2003: Started scap sync-world: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] * 20:02 aokoth@cumin1003: START - Cookbook sre.hosts.reimage for host phab1005.eqiad.wmnet with OS trixie * 20:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7004.magru.wmnet * 20:00 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet * 19:57 aokoth@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet * 19:24 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7011.magru.wmnet * 19:19 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7003.magru.wmnet * 19:11 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] (duration: 08m 33s) * 19:07 jforrester@deploy2003: jforrester: Continuing with deployment * 19:04 jforrester@deploy2003: jforrester: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:02 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] * 18:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7010.magru.wmnet * 18:38 mutante: rotating phabricator-gerrit bot token (its-phabricator) * 18:18 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 17:44 swfrench@deploy2003: Finished scap sync-world: Deployment to pick up new production image (duration: 31m 44s) * 17:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7002.magru.wmnet * 17:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7009.magru.wmnet * 17:32 swfrench@deploy2003: swfrench: Continuing with deployment * 17:29 swfrench@deploy2003: swfrench: Deployment to pick up new production image synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:17 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2035: repooling after rack b5 maintenance * 17:16 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool es2035: repooling after rack b5 maintenance * 17:16 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2188: repooling after rack b5 maintenance * 17:12 swfrench@deploy2003: Started scap sync-world: Deployment to pick up new production image * 17:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7001.magru.wmnet * 16:57 swfrench-wmf: reprepro include php8.3_8.3.32-1+wmf12u2 into component/php83 for bookworm-wikimedia * 16:50 sukhe: pool cp2046 * 16:47 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4039.ulsfo.wmnet * 16:44 sukhe: sudo cumin -b31 "A:cp" "run-puppet-agent" * 16:33 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on contint1003.wikimedia.org with reason: reboot * 16:32 mutante: contint1003 - main CI server - rebooting * 16:31 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2188: repooling after rack b5 maintenance * 16:31 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2178: repooling after rack b5 maintenance * 16:29 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 16:28 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 16:28 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 16:28 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 16:18 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2014.codfw.wmnet * 16:18 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 16:17 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2014.codfw.wmnet * 16:10 mvernon@cumin1003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-thanos-proxies (exit_code=0) rolling restart_daemons on A:thanos-fe * 16:09 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 16:07 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp4039.ulsfo.wmnet * 16:07 mvernon@cumin1003: START - Cookbook sre.swift.roll-restart-reboot-swift-thanos-proxies rolling restart_daemons on A:thanos-fe * 16:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2014.codfw.wmnet with OS trixie * 15:56 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1014.eqiad.wmnet with OS trixie * 15:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2014.codfw.wmnet with reason: host reimage * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2014 * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2014 * 15:28 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2014 * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2014.codfw.wmnet 194.16.192.10.in-addr.arpa 4.9.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:28 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2014.codfw.wmnet 194.16.192.10.in-addr.arpa 4.9.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2014 - mvernon@cumin2003" * 15:28 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2014 - mvernon@cumin2003" * 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Apply title-related policies when selecting the name of the entity - kamila@cumin1003" * 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Apply title-related policies when selecting the name of the entity - kamila@cumin1003 * 15:22 kamila@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Apply title-related policies when selecting the name of the entity - kamila@cumin1003 * 15:22 kamila@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Apply title-related policies when selecting the name of the entity - kamila@cumin1003" * 15:20 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 15:20 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2014 * 15:20 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2014.codfw.wmnet with OS trixie * 15:19 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1014.eqiad.wmnet with OS trixie * 15:01 dancy@deploy2003: Installation of scap version "4.274.1" completed for 3 hosts * 15:00 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2177: repooling after rack b5 maintenance * 15:00 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2159: repooling after rack b5 maintenance * 14:59 dancy@deploy2003: Installing scap version "4.274.1" for 3 host(s) * 14:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2015.codfw.wmnet with OS trixie * 14:54 seanleong-wmde: Finished populateSitesTable for isvwiki ([[phab:T429939|T429939]]) * 14:53 javiermonton@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] (duration: 07m 35s) * 14:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1015.eqiad.wmnet with OS trixie * 14:49 javiermonton@deploy2003: javiermonton: Continuing with deployment * 14:48 javiermonton@deploy2003: javiermonton: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:46 javiermonton@deploy2003: Started scap sync-world: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] * 14:42 otto@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 14:41 otto@deploy2003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 14:41 otto@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 14:40 otto@deploy2003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 14:40 otto@deploy2003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 14:39 otto@deploy2003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 14:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2015.codfw.wmnet with reason: host reimage * 14:34 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1015.eqiad.wmnet with reason: host reimage * 14:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2015.codfw.wmnet with reason: host reimage * 14:30 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1015.eqiad.wmnet with reason: host reimage * 14:30 seanleong-wmde@deploy2003: mwscript-k8s job started: foreachwikiindblist wikidataclient extensions/Wikibase/lib/maintenance/populateSitesTable.php --force-protocol https # [[phab:T429939|T429939]] * 14:24 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling reboot on A:durum and not (A:durum-eqiad or A:durum-codfw or A:durum-esams) and A:durum * 14:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2015.codfw.wmnet with OS trixie * 14:15 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2016.codfw.wmnet with OS trixie * 14:14 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2159: repooling after rack b5 maintenance * 14:14 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1015.eqiad.wmnet with OS trixie * 14:12 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1016.eqiad.wmnet with OS trixie * 14:12 cmooney@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=pki,name=codfw * 14:12 sbisson@deploy2003: helmfile [codfw] DONE helmfile.d/services/cxserver: sync * 14:11 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2002.codfw.wmnet * 14:11 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2002.codfw.wmnet * 14:11 sbisson@deploy2003: helmfile [codfw] START helmfile.d/services/cxserver: sync * 14:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1005.wikimedia.org * 14:07 sbisson@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cxserver: sync * 14:07 sbisson@deploy2003: helmfile [eqiad] START helmfile.d/services/cxserver: sync * 14:05 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1005.wikimedia.org * 14:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader2005.wikimedia.org * 14:02 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2003.codfw.wmnet * 14:02 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2003.codfw.wmnet * 14:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=tegola-vector-tiles,name=codfw * 14:00 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=kartotherian,name=codfw * 14:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader2005.wikimedia.org * 13:58 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2016.codfw.wmnet with reason: host reimage * 13:57 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-ncredir (exit_code=0) rolling reboot on A:ncredir and A:ncredir * 13:57 sbisson@deploy2003: helmfile [staging] DONE helmfile.d/services/cxserver: sync * 13:56 sbisson@deploy2003: helmfile [staging] START helmfile.d/services/cxserver: sync * 13:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1016.eqiad.wmnet with reason: host reimage * 13:52 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:52 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:51 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2016.codfw.wmnet with reason: host reimage * 13:50 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1016.eqiad.wmnet with reason: host reimage * 13:49 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy (exit_code=0) rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 13:49 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling reboot on A:wikidough * 13:46 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-tcp-proxy (exit_code=0) rolling reboot on A:tcpproxy and A:tcpproxy * 13:43 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and not (A:durum-eqiad or A:durum-codfw or A:durum-esams) and A:durum * 13:42 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=97) rolling reboot on A:durum and A:durum * 13:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2011.codfw.wmnet * 13:36 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-reboot (exit_code=0) rolling reboot on A:dnsbox and A:ulsfo and (A:dnsbox) * 13:36 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns4004.wikimedia.org * 13:34 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1016.eqiad.wmnet with OS trixie * 13:34 topranks: reboot lsw1-b5-codfw to upgrade JunOS [[phab:T430918|T430918]] * 13:34 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2016.codfw.wmnet with OS trixie * 13:32 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2002.codfw.wmnet * 13:31 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2011.codfw.wmnet * 13:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2012.codfw.wmnet * 13:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2017.codfw.wmnet with OS trixie * 13:25 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1017.eqiad.wmnet with OS trixie * 13:24 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2012.codfw.wmnet * 13:22 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2002.codfw.wmnet * 13:22 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:22 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:22 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns4004.wikimedia.org * 13:21 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2005.codfw.wmnet * 13:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2013.codfw.wmnet * 13:18 elukey@dns1004: END - running authdns-update * 13:17 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2005.codfw.wmnet * 13:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2004.codfw.wmnet * 13:16 elukey@dns1004: START - running authdns-update * 13:16 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1046: es1046 after reimage * 13:14 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1029.eqiad.wmnet,service=s8 * 13:14 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1029.eqiad.wmnet,service=s5 * 13:13 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1029.eqiad.wmnet,service=s5 * 13:13 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1029.eqiad.wmnet,service=s8 * 13:13 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2004.codfw.wmnet * 13:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2013.codfw.wmnet * 13:11 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2014.codfw.wmnet * 13:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2003.codfw.wmnet * 13:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2017.codfw.wmnet with reason: host reimage * 13:07 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:07 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns4003.wikimedia.org * 13:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2003.codfw.wmnet * 13:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1067.eqiad.wmnet * 13:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1067.eqiad.wmnet * 13:06 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1067.eqiad.wmnet * 13:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm-test1001.wikimedia.org * 13:05 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2017.codfw.wmnet with reason: host reimage * 13:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1017.eqiad.wmnet with reason: host reimage * 13:04 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2014.codfw.wmnet * 13:03 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:02 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2188: codfw rack B5 depool for maintenance * 13:02 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_magru * 13:01 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2188: codfw rack B5 depool for maintenance * 13:01 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_magru * 13:01 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2178: codfw rack B5 depool for maintenance * 13:01 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm-test1001.wikimedia.org * 13:01 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2178: codfw rack B5 depool for maintenance * 13:01 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2177: codfw rack B5 depool for maintenance * 13:00 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2177: codfw rack B5 depool for maintenance * 12:59 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1068.eqiad.wmnet * 12:59 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1068.eqiad.wmnet * 12:58 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola-vector-tiles,name=codfw * 12:58 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2159: codfw rack B5 depool for maintenance * 12:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1017.eqiad.wmnet with reason: host reimage * 12:58 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola,name=codfw * 12:57 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=kartotherian,name=codfw * 12:57 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2159: codfw rack B5 depool for maintenance * 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 30 hosts with reason: lsw1-b5-codfw JunOS upgrade * 12:55 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lsw1-b5-codfw,lsw1-b5-codfw IPv6,lsw1-b5-codfw.mgmt,ssw1-a[1,8]-codfw.mgmt with reason: switch upgade lsw1-b5-codfw * 12:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet * 12:49 cmooney@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=pki,name=codfw * 12:49 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1067.eqiad.wmnet with OS trixie * 12:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet * 12:48 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2017.codfw.wmnet with OS trixie * 12:47 topranks: depool codfw pki in dns discovery ahead of lsw1-b5-codfw maintenance [[phab:T430918|T430918]] * 12:47 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns4003.wikimedia.org * 12:47 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and A:ulsfo and (A:dnsbox) * 12:47 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2018.codfw.wmnet * 12:45 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2018.codfw.wmnet * 12:45 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and A:durum * 12:45 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-tcp-proxy rolling reboot on A:tcpproxy and A:tcpproxy * 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host pki-root1002.eqiad.wmnet * 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1009.eqiad.wmnet * 12:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1009.eqiad.wmnet * 12:44 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 12:43 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-ncredir rolling reboot on A:ncredir and A:ncredir * 12:43 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling reboot on A:wikidough * 12:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2018.codfw.wmnet with OS trixie * 12:42 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1017.eqiad.wmnet with OS trixie * 12:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader2006.wikimedia.org * 12:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1018.eqiad.wmnet with OS trixie * 12:39 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1009.eqiad.wmnet * 12:38 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host pki-root1002.eqiad.wmnet * 12:38 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1009.eqiad.wmnet * 12:38 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1008.eqiad.wmnet * 12:38 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1008.eqiad.wmnet * 12:35 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader2006.wikimedia.org * 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1006.wikimedia.org * 12:33 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1008.eqiad.wmnet * 12:30 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1046: es1046 after reimage * 12:29 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host es1046.eqiad.wmnet with OS trixie * 12:29 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1006.wikimedia.org * 12:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test2005.wikimedia.org * 12:28 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1008.eqiad.wmnet * 12:27 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1007.eqiad.wmnet * 12:27 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1007.eqiad.wmnet * 12:27 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1067.eqiad.wmnet with reason: host reimage * 12:25 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1068.eqiad.wmnet with reason: vacuum overlarge container dbs * 12:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2018.codfw.wmnet with reason: host reimage * 12:24 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test2005.wikimedia.org * 12:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test1005.wikimedia.org * 12:23 Amir1: mwscript-k8s --follow --dblist=ores -- extensions/ORES/maintenance/PurgeScoreCache.php --model goodfaith --old ([[phab:T431159|T431159]]) * 12:22 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1007.eqiad.wmnet * 12:22 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test1005.wikimedia.org * 12:22 atsukoito: restarting pybal on lvs2013 `low-traffic` for https://gerrit.wikimedia.org/r/1310535 * 12:22 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1007.eqiad.wmnet * 12:21 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1006.eqiad.wmnet * 12:21 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1006.eqiad.wmnet * 12:20 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1018.eqiad.wmnet with reason: host reimage * 12:19 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2018.codfw.wmnet with reason: host reimage * 12:18 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1067.eqiad.wmnet with reason: host reimage * 12:16 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1006.eqiad.wmnet * 12:15 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1006.eqiad.wmnet * 12:15 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1005.eqiad.wmnet * 12:15 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1005.eqiad.wmnet * 12:15 atsukoito: restarting pybal on lvs2014 for https://gerrit.wikimedia.org/r/1310535 * 12:12 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1018.eqiad.wmnet with reason: host reimage * 12:11 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1005.eqiad.wmnet * 12:11 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1005.eqiad.wmnet * 12:10 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1004.eqiad.wmnet * 12:10 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1004.eqiad.wmnet * 12:09 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on es1046.eqiad.wmnet with reason: host reimage * 12:08 atsukoito: restarting pybal on lvs1019 `low-traffic` for https://gerrit.wikimedia.org/r/1310535 * 12:06 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1004.eqiad.wmnet * 12:06 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1004.eqiad.wmnet * 12:06 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1003.eqiad.wmnet * 12:06 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1003.eqiad.wmnet * 12:05 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on es1046.eqiad.wmnet with reason: host reimage * 12:05 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "set ml-serve1001 back to active state - cmooney@cumin1003" * 12:04 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "set ml-serve1001 back to active state - cmooney@cumin1003" * 12:04 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:02 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1003.eqiad.wmnet * 12:01 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1003.eqiad.wmnet * 12:01 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1002.eqiad.wmnet * 12:01 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1002.eqiad.wmnet * 12:01 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:59 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2018.codfw.wmnet with OS trixie * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1067 * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1067 * 11:59 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1067 * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1067.eqiad.wmnet 17.48.64.10.in-addr.arpa 7.1.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:59 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1067.eqiad.wmnet 17.48.64.10.in-addr.arpa 7.1.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1067 - blake@cumin1003" * 11:58 atsukoito: restarting pybal on lvs1018 `high-traffic2` for https://gerrit.wikimedia.org/r/1310535 * 11:57 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1002.eqiad.wmnet * 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2019.codfw.wmnet with OS trixie * 11:56 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1002.eqiad.wmnet * 11:56 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 11:56 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1018.eqiad.wmnet with OS trixie * 11:54 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 11:54 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:54 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1019.eqiad.wmnet with OS trixie * 11:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:49 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:49 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:49 aikochou@deploy2003: helmfile [codfw] DONE helmfile.d/services/changeprop: sync * 11:48 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host es1046.eqiad.wmnet with OS trixie * 11:48 aikochou@deploy2003: helmfile [codfw] START helmfile.d/services/changeprop: sync * 11:48 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310535 * 11:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1046: Reimage to Trixie * 11:44 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1046: Reimage to Trixie * 11:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5:00:00 on es1046.eqiad.wmnet with reason: Reimage to Trixie * 11:42 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:42 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:42 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 11:42 aikochou@deploy2003: helmfile [eqiad] DONE helmfile.d/services/changeprop: sync * 11:41 aikochou@deploy2003: helmfile [eqiad] START helmfile.d/services/changeprop: sync * 11:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2019.codfw.wmnet with reason: host reimage * 11:36 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] (duration: 09m 41s) * 11:36 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:36 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:35 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:35 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:32 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1019.eqiad.wmnet with reason: host reimage * 11:32 jforrester@deploy2003: jforrester, gengh: Continuing with deployment * 11:29 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2019.codfw.wmnet with reason: host reimage * 11:28 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1019.eqiad.wmnet with reason: host reimage * 11:28 jforrester@deploy2003: jforrester, gengh: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:26 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] * 11:20 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2003.codfw.wmnet * 11:20 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:19 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2003.codfw.wmnet * 11:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:12 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1019.eqiad.wmnet with OS trixie * 11:10 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2019.codfw.wmnet with OS trixie * 11:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1020.eqiad.wmnet with OS trixie * 11:10 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1067 - blake@cumin1003" * 11:09 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] (duration: 12m 12s) * 11:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2020.codfw.wmnet with OS trixie * 11:03 kharlan@deploy2003: kharlan: Continuing with deployment * 11:01 blake@cumin1003: START - Cookbook sre.dns.netbox * 11:01 kharlan@deploy2003: kharlan: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:57 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] * 10:55 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] (duration: 31m 40s) * 10:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1020.eqiad.wmnet with reason: host reimage * 10:52 marostegui@dns1004: START - running authdns-update * 10:49 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2020.codfw.wmnet with reason: host reimage * 10:49 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:48 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1020.eqiad.wmnet with reason: host reimage * 10:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2020.codfw.wmnet with reason: host reimage * 10:43 kharlan@deploy2003: kharlan: Continuing with deployment * 10:42 kharlan@deploy2003: kharlan: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:32 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1020.eqiad.wmnet with OS trixie * 10:29 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2159: Repooling after switchover * 10:29 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1067 * 10:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1021.eqiad.wmnet with OS trixie * 10:27 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1067.eqiad.wmnet with OS trixie * 10:27 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:27 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1067.eqiad.wmnet * 10:27 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:26 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1067.eqiad.wmnet * 10:26 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1067.eqiad.wmnet * 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2020.codfw.wmnet with OS trixie * 10:26 blake@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1055.eqiad.wmnet * 10:26 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1055.eqiad.wmnet * 10:26 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1055.eqiad.wmnet * 10:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2021.codfw.wmnet with OS trixie * 10:24 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] * 10:11 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1055.eqiad.wmnet with OS trixie * 10:09 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1021.eqiad.wmnet with reason: host reimage * 10:05 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2021.codfw.wmnet with reason: host reimage * 10:03 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310129 revert * 10:02 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1021.eqiad.wmnet with reason: host reimage * 10:01 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2021.codfw.wmnet with reason: host reimage * 09:58 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310129 * 09:50 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1055.eqiad.wmnet with reason: host reimage * 09:45 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1055.eqiad.wmnet with reason: host reimage * 09:45 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1021.eqiad.wmnet with OS trixie * 09:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1022.eqiad.wmnet with OS trixie * 09:44 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2159: Repooling after switchover * 09:44 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2021.codfw.wmnet with OS trixie * 09:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2022.codfw.wmnet with OS trixie * 09:31 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2159.codfw.wmnet * 09:28 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1055 * 09:28 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1055 * 09:27 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on ms-fe1022.eqiad.wmnet with reason: host reimage * 09:27 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1022.eqiad.wmnet with reason: host reimage * 09:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2022.codfw.wmnet with reason: host reimage * 09:21 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2022.codfw.wmnet with reason: host reimage * 09:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2159: Rebooting db2159.codfw.wmnet * 09:20 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2159: Rebooting db2159.codfw.wmnet * 09:18 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 09:18 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 09:18 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 09:17 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 09:16 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2159.codfw.wmnet * 09:13 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b] (thin): Regular analytics weekly train THIN [analytics/refinery@ad6e05b8] (duration: 02m 07s) * 09:11 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b] (thin): Regular analytics weekly train THIN [analytics/refinery@ad6e05b8] * 09:10 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1022.eqiad.wmnet with OS trixie * 09:07 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1023.eqiad.wmnet with OS trixie * 09:06 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b]: Regular analytics weekly train [analytics/refinery@ad6e05b8] (duration: 04m 49s) * 09:04 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2022.codfw.wmnet with OS trixie * 09:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2023.codfw.wmnet with OS trixie * 09:01 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b]: Regular analytics weekly train [analytics/refinery@ad6e05b8] * 09:01 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@ad6e05b8] (duration: 02m 01s) * 09:00 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1055 * 09:00 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1055.eqiad.wmnet 50.32.64.10.in-addr.arpa 0.5.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:00 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1055.eqiad.wmnet 50.32.64.10.in-addr.arpa 0.5.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:00 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:00 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1055 - blake@cumin1003" * 09:00 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1055 - blake@cumin1003" * 08:59 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@ad6e05b8] * 08:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2159 [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94811 and previous config saved to /var/cache/conftool/dbconfig/20260714-085624-cwilliams.json * 08:55 blake@cumin1003: START - Cookbook sre.dns.netbox * 08:55 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1055 * 08:54 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1055.eqiad.wmnet with OS trixie * 08:54 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1055.eqiad.wmnet * 08:53 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1055.eqiad.wmnet * 08:53 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1055.eqiad.wmnet * 08:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2220 to s7 primary [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94810 and previous config saved to /var/cache/conftool/dbconfig/20260714-085239-cwilliams.json * 08:51 cezmunsta: Starting s7 codfw failover from db2159 to db2220 - [[phab:T430920|T430920]] * 08:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1023.eqiad.wmnet with reason: host reimage * 08:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2220 with weight 0 [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94809 and previous config saved to /var/cache/conftool/dbconfig/20260714-084553-cwilliams.json * 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s7 [[phab:T430920|T430920]] * 08:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2023.codfw.wmnet with reason: host reimage * 08:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1023.eqiad.wmnet with reason: host reimage * 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2023.codfw.wmnet with reason: host reimage * 08:34 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:34 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:29 marostegui@dns1004: END - running authdns-update * 08:29 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox-dev2003.codfw.wmnet * 08:27 marostegui@dns1004: START - running authdns-update * 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker2*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2009.codfw.wmnet * 08:26 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2009.codfw.wmnet * 08:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1023.eqiad.wmnet with OS trixie * 08:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox-dev2003.codfw.wmnet * 08:24 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:24 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:24 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2023.codfw.wmnet with OS trixie * 08:24 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1029.eqiad.wmnet with reason: reboot * 08:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1027.eqiad.wmnet with reason: reboot * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:21 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2009.codfw.wmnet * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:20 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2009.codfw.wmnet * 08:20 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2008.codfw.wmnet * 08:20 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2008.codfw.wmnet * 08:15 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2008.codfw.wmnet * 08:14 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2008.codfw.wmnet * 08:14 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2007.codfw.wmnet * 08:14 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2007.codfw.wmnet * 08:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1024.eqiad.wmnet with OS trixie * 08:12 elukey@cumin1003: END (PASS) - Cookbook sre.pki.restart-reboot (exit_code=0) rolling reboot on P<nowiki>{</nowiki>pki*<nowiki>}</nowiki> and (A:pki) * 08:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 08:10 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 08:09 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2007.codfw.wmnet * 08:08 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2007.codfw.wmnet * 08:08 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2006.codfw.wmnet * 08:08 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2006.codfw.wmnet * 08:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2024.codfw.wmnet with OS trixie * 08:03 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2006.codfw.wmnet * 08:02 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2006.codfw.wmnet * 08:02 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2005.codfw.wmnet * 08:02 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2005.codfw.wmnet * 07:58 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2005.codfw.wmnet * 07:58 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2005.codfw.wmnet * 07:57 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2004.codfw.wmnet * 07:57 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2004.codfw.wmnet * 07:54 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki.discovery.wmnet. on all recursors * 07:54 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache pki.discovery.wmnet. on all recursors * 07:53 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2004.codfw.wmnet * 07:53 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2004.codfw.wmnet * 07:53 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2003.codfw.wmnet * 07:53 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2003.codfw.wmnet * 07:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1024.eqiad.wmnet with reason: host reimage * 07:49 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki.discovery.wmnet. on all recursors * 07:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2003.codfw.wmnet * 07:49 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache pki.discovery.wmnet. on all recursors * 07:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2024.codfw.wmnet with reason: host reimage * 07:48 elukey@cumin1003: START - Cookbook sre.pki.restart-reboot rolling reboot on P<nowiki>{</nowiki>pki*<nowiki>}</nowiki> and (A:pki) * 07:46 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1024.eqiad.wmnet with reason: host reimage * 07:45 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2024.codfw.wmnet with reason: host reimage * 07:45 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2003.codfw.wmnet * 07:44 elukey@cumin1003: END (PASS) - Cookbook sre.misc-clusters.restart-reboot-config-master (exit_code=0) rolling reboot on P<nowiki>{</nowiki>config-master*<nowiki>}</nowiki> and (A:config-master or A:config-master-eqiad or A:config-master-codfw) * 07:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2002.codfw.wmnet * 07:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2002.codfw.wmnet * 07:39 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2002.codfw.wmnet * 07:39 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) config-master.discovery.wmnet. on all recursors * 07:39 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache config-master.discovery.wmnet. on all recursors * 07:39 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2002.codfw.wmnet * 07:39 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker2*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl200*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl2003.codfw.wmnet * 07:36 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl2003.codfw.wmnet * 07:35 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) config-master.discovery.wmnet. on all recursors * 07:35 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache config-master.discovery.wmnet. on all recursors * 07:34 elukey@cumin1003: START - Cookbook sre.misc-clusters.restart-reboot-config-master rolling reboot on P<nowiki>{</nowiki>config-master*<nowiki>}</nowiki> and (A:config-master or A:config-master-eqiad or A:config-master-codfw) * 07:31 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl2003.codfw.wmnet * 07:31 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl2003.codfw.wmnet * 07:31 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl2002.codfw.wmnet * 07:31 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl2002.codfw.wmnet * 07:29 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1024.eqiad.wmnet with OS trixie * 07:28 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2024.codfw.wmnet with OS trixie * 07:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl2002.codfw.wmnet * 07:26 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl2002.codfw.wmnet * 07:26 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl200*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 07:26 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 07:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 06:50 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lists1004.wikimedia.org * 06:44 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host lists1004.wikimedia.org * 06:25 marostegui@dns1004: END - running authdns-update * 06:23 marostegui@dns1004: START - running authdns-update * 06:22 marostegui@dns1004: END - running authdns-update * 06:20 marostegui@dns1004: START - running authdns-update * 06:17 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1026.eqiad.wmnet with reason: reboot * 06:04 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: sync * 06:04 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: sync * 06:03 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync * 06:03 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync * 06:02 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync * 06:01 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync * 06:01 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync * 06:00 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync * 05:59 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:59 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:40 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:39 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:26 marostegui@dns1004: END - running authdns-update * 05:24 marostegui@dns1004: START - running authdns-update * 05:24 marostegui@dns1004: START - running authdns-update * 05:13 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1004.wikimedia.org * 05:07 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1004.wikimedia.org * 04:01 mwpresync@deploy2003: Pruned MediaWiki: 1.47.0-wmf.8 (duration: 01m 07s) * 03:39 mwpresync@deploy2003: Finished scap sync-world: testwikis to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] (duration: 36m 01s) * 03:03 mwpresync@deploy2003: Started scap sync-world: testwikis to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 29s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-13 == * 23:33 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1064.eqiad.wmnet * 23:33 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1064.eqiad.wmnet * 23:08 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1064.eqiad.wmnet with reason: vacuum overlarge container dbs * 23:06 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1069.eqiad.wmnet * 23:06 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1069.eqiad.wmnet * 22:34 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1069.eqiad.wmnet with reason: vacuum overlarge container dbs * 21:18 maryum: Deployed security fix for [[phab:T321092|T321092]] * 20:28 swfrench-wmf: reprepro include etcd-mirror_0.0.12-1+deb13u1 into main for trixie-wikimedia - [[phab:T424266|T424266]] * 20:26 swfrench-wmf: reprepro include etcd-mirror_0.0.12-1+deb12u1 into main for bookworm-wikimedia - [[phab:T428495|T428495]] * 20:23 dancy@deploy2003: Finished scap sync-world: Testing [[phab:T431635|T431635]] (duration: 03m 36s) * 20:19 dancy@deploy2003: Started scap sync-world: Testing [[phab:T431635|T431635]] * 20:18 dancy@deploy2003: Installation of scap version "4.274.0" completed for 3 hosts * 20:16 dancy@deploy2003: Installing scap version "4.274.0" for 3 host(s) * 20:12 kemayo@deploy2003: Finished scap sync-world: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] (duration: 08m 25s) * 20:07 kemayo@deploy2003: soda, esanders, kemayo: Continuing with deployment * 20:05 kemayo@deploy2003: soda, esanders, kemayo: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there * 20:04 kemayo@deploy2003: Started scap sync-world: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] * 18:22 cwhite: lvextend vg0/srv +500g on centrallog hosts * 18:19 cdobbins@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS trixie * 17:46 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1071.eqiad.wmnet * 17:46 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1071.eqiad.wmnet * 17:13 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1071.eqiad.wmnet with reason: vacuum overlarge container dbs * 17:07 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1065.eqiad.wmnet * 17:07 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1065.eqiad.wmnet * 17:06 dzahn@dns1006: END - running authdns-update * 17:04 dzahn@dns1006: START - running authdns-update * 17:01 dzahn@dns1006: END - running authdns-update * 16:59 dzahn@dns1006: START - running authdns-update * 16:51 dancy@deploy2003: Finished scap sync-world: testing [[phab:T428971|T428971]] (duration: 03m 37s) * 16:47 dancy@deploy2003: Started scap sync-world: testing [[phab:T428971|T428971]] * 16:45 atsukoito: restarting pybal on lvs1019 to flush IP address for `cirrussearch1122.eqiad.wmnet` after moving the vlan [[phab:T431311|T431311]] * 16:42 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:42 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:42 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:42 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:42 Amir1: mwscript-k8s --follow --dblist=ores -- extensions/ORES/maintenance/PurgeScoreCache.php --model damaging --old ([[phab:T431159|T431159]]) * 16:34 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Pool test * 16:34 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 16:34 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 16:34 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Pool test * 16:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Depool test * 16:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 16:33 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 16:33 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Depool test * 16:31 dancy@deploy2003: Installation of scap version "4.273.0" completed for 159 hosts * 16:29 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1065.eqiad.wmnet with reason: vacuum overlarge container dbs * 16:27 dancy@deploy2003: Installing scap version "4.273.0" for 159 host(s) * 16:27 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics-external: sync * 16:27 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics-external: sync * 16:26 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics-external: sync * 16:26 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics-external: sync * 16:22 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync * 16:21 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync * 16:21 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: sync * 16:21 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: sync * 16:19 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync * 16:19 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync * 16:18 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync * 16:17 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync * 15:59 atsukoito: restarting pybal on lvs1018 for https://gerrit.wikimedia.org/r/1310117 * 15:55 aikochou@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 15:50 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310117 * 15:46 aikochou@deploy2003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 15:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host kafka-logging1006.eqiad.wmnet * 15:43 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host kafka-logging1006.eqiad.wmnet * 15:41 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host ganeti-test[2001-2003].codfw.wmnet * 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host ganeti-test[2001-2003].codfw.wmnet * 15:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host netbox1003.eqiad.wmnet * 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host netbox1003.eqiad.wmnet * 15:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host netbox2003.codfw.wmnet * 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host netbox2003.codfw.wmnet * 15:36 sukhe: restart pybal on lvs1020 * 15:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet * 15:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet * 15:08 btullis@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:06 btullis@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 15:01 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:01 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:35 cdobbins@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 14:34 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:33 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:33 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:32 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:29 cdobbins@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 14:28 marostegui@dns1004: END - running authdns-update * 14:28 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:27 marostegui@dns1004: START - running authdns-update * 14:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1023.eqiad.wmnet with reason: reboot * 14:18 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2009.codfw.wmnet with OS trixie * 14:14 swfrench-wmf: start rolling run-puppet-agent on A:cp for ATS config change - [[phab:T428909|T428909]] [[phab:T431838|T431838]] * 14:05 swfrench-wmf: disable-puppet on A:cp for ATS config change - [[phab:T428909|T428909]] [[phab:T431838|T431838]] * 14:05 cdobbins@cumin2002: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie * 14:02 marostegui@dns1004: END - running authdns-update * 14:00 marostegui@dns1004: START - running authdns-update * 14:00 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1070.eqiad.wmnet * 14:00 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1070.eqiad.wmnet * 13:58 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2009.codfw.wmnet with reason: host reimage * 13:57 cdobbins@cumin2002: conftool action : set/pooled=no; selector: name=dns7002.* * 13:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2009.codfw.wmnet with reason: host reimage * 13:48 rscout@deploy2003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply * 13:48 rscout@deploy2003: helmfile [eqiad] START helmfile.d/services/miscweb: apply * 13:48 rscout@deploy2003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply * 13:47 rscout@deploy2003: helmfile [codfw] START helmfile.d/services/miscweb: apply * 13:40 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:33 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2009.codfw.wmnet with OS trixie * 13:30 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1070.eqiad.wmnet with reason: vacuum overlarge container dbs * 13:28 aude@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] (duration: 11m 12s) * 13:23 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:22 aude@deploy2003: aikochou, javiermonton, aude, gkm563: Continuing with deployment * 13:22 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:19 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:19 aude@deploy2003: aikochou, javiermonton, aude, gkm563: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] synced to the testservers * 13:17 aude@deploy2003: Started scap sync-world: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] * 13:01 ladsgroup@deploy2003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 13:01 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:00 ladsgroup@deploy2003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 12:59 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:52 ladsgroup@deploy2003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 12:51 ladsgroup@deploy2003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 12:48 atsuko@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 12:48 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 12:47 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2008.codfw.wmnet with OS trixie * 12:47 atsuko@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 12:47 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply * 12:47 atsuko@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:46 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 12:45 atsuko@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:45 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply * 12:45 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:44 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:43 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] (duration: 07m 02s) * 12:38 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 12:37 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:36 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] * 12:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2008.codfw.wmnet with reason: host reimage * 12:23 Msz2001: Deployed changes to private code for Suggested Investigations * 12:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2008.codfw.wmnet with reason: host reimage * 12:20 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:19 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:17 atsuko@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 12:17 atsuko@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 12:16 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:15 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] (duration: 07m 14s) * 12:10 mszwarc@deploy2003: mszwarc: Continuing with deployment * 12:09 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:07 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] * 12:04 mszwarc@deploy2003: sync-world aborted: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] (duration: 00m 29s) * 12:03 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] * 12:01 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2008.codfw.wmnet with OS trixie * 12:00 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] (duration: 07m 37s) * 11:55 zabe@deploy2003: zabe: Continuing with deployment * 11:54 zabe@deploy2003: zabe: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:52 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] * 11:51 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:43 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:35 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:34 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:33 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:30 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:28 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:27 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:17 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2007.codfw.wmnet with OS trixie * 11:09 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] * 11:06 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=s8 * 11:00 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=x3 * 11:00 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=s5 * 10:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2007.codfw.wmnet with reason: host reimage * 10:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2007.codfw.wmnet with reason: host reimage * 10:51 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host dse-k8s-worker1023 * 10:50 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host dse-k8s-worker1023 * 10:44 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dse-k8s-worker1023 * 10:43 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host dse-k8s-worker1023 * 10:42 marostegui@cumin1003: dbctl commit (dc=all): 'Change x4 masters [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P94804 and previous config saved to /var/cache/conftool/dbconfig/20260713-104248-marostegui.json * 10:37 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:37 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:35 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:35 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:34 atsuko@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 10:34 atsuko@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 10:33 marostegui@cumin1003: dbctl commit (dc=all): 'Push x4 initial dbctl config [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P94803 and previous config saved to /var/cache/conftool/dbconfig/20260713-103259-marostegui.json * 10:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2007.codfw.wmnet with OS trixie * 09:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2006.codfw.wmnet with OS trixie * 09:42 marostegui@dns1004: END - running authdns-update * 09:40 marostegui@dns1004: START - running authdns-update * 09:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2006.codfw.wmnet with reason: host reimage * 09:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2006.codfw.wmnet with reason: host reimage * 09:06 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1024.eqiad.wmnet with reason: reboot * 09:06 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:01 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2006.codfw.wmnet with OS trixie * 08:44 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 08:43 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 08:43 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:42 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:42 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:42 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:41 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 08:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup2004.codfw.wmnet * 08:38 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:33 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host db1208.eqiad.wmnet * 08:30 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=x3 * 08:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2005.codfw.wmnet with OS trixie * 08:28 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup2004.codfw.wmnet * 08:28 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup2003.codfw.wmnet * 08:24 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1039: Repooling after testing * 08:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on clouddb1016.eqiad.wmnet with reason: cloning * 08:23 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s5 * 08:23 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s8 * 08:21 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 08:21 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 08:17 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup2003.codfw.wmnet * 08:17 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1004.eqiad.wmnet * 08:14 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1208.eqiad.wmnet * 08:11 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host phab1005.eqiad.wmnet * 08:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2005.codfw.wmnet with reason: host reimage * 08:07 marostegui@dns1004: END - running authdns-update * 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1004.eqiad.wmnet * 08:07 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1003.eqiad.wmnet * 08:07 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:05 marostegui@dns1004: START - running authdns-update * 08:05 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2005.codfw.wmnet with reason: host reimage * 08:05 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 08:05 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host phab1005.eqiad.wmnet * 08:05 marostegui@dns1004: START - running authdns-update * 08:05 marostegui@dns1004: START - running authdns-update * 08:05 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 08:04 marostegui@dns1004: START - running authdns-update * 08:00 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit1003.wikimedia.org * 07:58 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1003.eqiad.wmnet * 07:58 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1002-dev.eqiad.wmnet * 07:58 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:58 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:54 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1002-dev.eqiad.wmnet * 07:54 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1001-dev.eqiad.wmnet * 07:54 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit1003.wikimedia.org * 07:53 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 07:53 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:52 Msz2001: UTC morning backport+config window done * 07:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2005.codfw.wmnet with OS trixie * {{safesubst:SAL entry|1=07:50 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark (T429943}} * 07:49 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1001-dev.eqiad.wmnet * 07:46 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:46 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:45 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 07:45 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:44 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 07:44 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:43 mszwarc@deploy2003: mszwarc, danielyepezgarces, anzx: Continuing with deployment * {{safesubst:SAL entry|1=07:39 mszwarc@deploy2003: mszwarc, danielyepezgarces, anzx: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark}} * 07:39 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1039: Repooling after testing * {{safesubst:SAL entry|1=07:36 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark (T429943)}} * 07:35 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] (duration: 30m 03s) * 07:25 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit2002.wikimedia.org * 07:22 mszwarc@deploy2003: mszwarc: Continuing with deployment * 07:21 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:19 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit2002.wikimedia.org * 07:15 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aphlict1002.eqiad.wmnet * 07:11 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host aphlict1002.eqiad.wmnet * 07:08 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2003.wikimedia.org * 07:05 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] * 07:02 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2003.wikimedia.org * 07:02 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2002.wikimedia.org * 06:55 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2002.wikimedia.org * 06:55 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1003.wikimedia.org * 06:49 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1003.wikimedia.org * 06:34 marostegui: Drop m5 ipoid database [[phab:T431007|T431007]] * 06:29 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1027.eqiad.wmnet with reason: reboot * 06:24 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1028.eqiad.wmnet with reason: reboot * 06:21 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1025.eqiad.wmnet with reason: reboot * 06:17 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1022.eqiad.wmnet with reason: reboot * 06:03 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on dbproxy[2005-2008].codfw.wmnet with reason: reboot * 05:37 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1217,1228].eqiad.wmnet with reason: cloning * 05:11 marostegui: Drop users_to_rename table [[phab:T431842|T431842]] * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-12 == * 16:01 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2209 [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94792 and previous config saved to /var/cache/conftool/dbconfig/20260712-160124-marostegui.json * 15:58 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2205 to s3 primary [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94791 and previous config saved to /var/cache/conftool/dbconfig/20260712-155853-marostegui.json * 15:58 marostegui: Starting s3 codfw emergency failover from db2209 to db2205 - [[phab:T431950|T431950]] * 15:51 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2205 with weight 0 [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94790 and previous config saved to /var/cache/conftool/dbconfig/20260712-155135-marostegui.json * 15:51 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Primary switchover s3 [[phab:T431950|T431950]] * 02:01 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 01m 17s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-11 == * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 26s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-10 == * 19:12 jhathaway@dns1004: END - running authdns-update * 19:10 jhathaway@dns1004: START - running authdns-update * 18:23 mutante: vrts2002 rebooting (not the active host) * 18:21 mutante: lists2001, phab2003 - rebooting (not the active hosts) * 18:16 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on A:lvs-high-traffic2-codfw * 18:15 swfrench@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on A:lvs-high-traffic2-codfw * 17:15 mutante: [doc1004:~] $ sudo systemctl start rsync-doc-host-data-sync ([[phab:T431856|T431856]]) * 17:09 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1004.eqiad.wmnet * 17:08 jhathaway@dns1004: END - running authdns-update * 17:07 jhathaway@dns1004: START - running authdns-update * 17:06 jhathaway: depooling puppetserver1002, cause of errors is still unknown * 17:03 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1004.eqiad.wmnet * 16:57 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1003.eqiad.wmnet * 16:51 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1003.eqiad.wmnet * 16:48 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 16:48 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2004.codfw.wmnet * 16:42 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2004.codfw.wmnet * 16:41 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2003.codfw.wmnet * 16:35 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2003.codfw.wmnet * 16:33 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2002.codfw.wmnet * 16:27 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2002.codfw.wmnet * 16:25 mutante: gitlab-runners (production) rebooting cluster one by one * 16:17 mutante: etherpad1004/etherpad2002 - (etherpad.wikimedia.org) - rebooting * 16:13 mutante: doc1004/doc2003 (doc.wikimedia.org backends) - rebooting * 16:02 mutante: releases1003/releases2003 (releases.wikimedia.org backends) - rebooting for maintenance * 15:26 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2007-dev.codfw.wmnet * 15:19 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2007-dev.codfw.wmnet * 15:14 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host cloudcephosd2007-dev.codfw.wmnet * 15:14 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2007-dev.codfw.wmnet * 15:14 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host cloudcephosd2006-dev.codfw.wmnet * 15:07 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2006-dev.codfw.wmnet * 15:07 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2005-dev.codfw.wmnet * 14:59 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2005-dev.codfw.wmnet * 14:59 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2004-dev.codfw.wmnet * 14:53 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2004-dev.codfw.wmnet * 14:53 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2007-dev.codfw.wmnet * 14:51 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1054.eqiad.wmnet * 14:51 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1054.eqiad.wmnet * 14:51 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1054.eqiad.wmnet * 14:47 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2007-dev.codfw.wmnet * 14:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2006-dev.codfw.wmnet * 14:41 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2006-dev.codfw.wmnet * 14:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2005-dev.codfw.wmnet * 14:37 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2005-dev.codfw.wmnet * 14:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2005-dev.codfw.wmnet * 14:29 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2005-dev.codfw.wmnet * 14:29 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2006-dev.codfw.wmnet * 14:21 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2006-dev.codfw.wmnet * 14:21 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2010-dev.codfw.wmnet * 14:15 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2010-dev.codfw.wmnet * 14:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudgw2004-dev.codfw.wmnet * 14:10 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1054.eqiad.wmnet with OS trixie * 14:09 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudgw2004-dev.codfw.wmnet * 14:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudgw2003-dev.codfw.wmnet * 14:02 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudgw2003-dev.codfw.wmnet * 14:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2004-dev.codfw.wmnet * 13:53 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2004-dev.codfw.wmnet * 13:53 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2003-dev.codfw.wmnet * 13:48 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:44 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2003-dev.codfw.wmnet * 13:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2002-dev.codfw.wmnet * 13:42 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:41 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:41 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:37 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2002-dev.codfw.wmnet * 13:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudidp2001-dev.codfw.wmnet * 13:33 blake@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker1054.eqiad.wmnet with reason: host reimage * 13:33 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudidp2001-dev.codfw.wmnet * 13:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudnet2006-dev.codfw.wmnet * 13:26 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudnet2006-dev.codfw.wmnet * 13:26 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudnet2005-dev.codfw.wmnet * 13:23 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1054.eqiad.wmnet with reason: host reimage * 13:18 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudnet2005-dev.codfw.wmnet * 13:18 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudservices2005-dev.codfw.wmnet * 13:12 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudservices2005-dev.codfw.wmnet * 13:11 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudservices2004-dev.codfw.wmnet * 13:08 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudservices2004-dev.codfw.wmnet * 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudweb2002-dev.wikimedia.org * 13:05 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 13:05 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1054 * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1054 * 13:04 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1054 * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1054.eqiad.wmnet 49.32.64.10.in-addr.arpa 9.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:04 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1054.eqiad.wmnet 49.32.64.10.in-addr.arpa 9.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1054 - blake@cumin1003" * 13:04 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1054 - blake@cumin1003" * 13:01 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudweb2002-dev.wikimedia.org * 13:00 blake@cumin1003: START - Cookbook sre.dns.netbox * 12:59 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1054 * 12:57 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1054.eqiad.wmnet with OS trixie * 12:57 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1054.eqiad.wmnet * 12:56 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1054.eqiad.wmnet * 12:56 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1054.eqiad.wmnet * 12:47 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS trixie * 12:44 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:39 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 12:39 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 12:38 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:37 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:14 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:10 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:08 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:07 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:00 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:00 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:51 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:49 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:48 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:47 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:44 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:32 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2001.codfw.wmnet * 11:32 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1053.eqiad.wmnet * 11:32 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2001.codfw.wmnet * 11:32 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1053.eqiad.wmnet * 11:32 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1053.eqiad.wmnet * 11:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker2001.codfw.wmnet * 11:31 cgoubert@cumin1003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker2001.codfw.wmnet * 11:31 cgoubert@cumin1003: END (FAIL) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=1) rolling reimage on P<nowiki>{</nowiki>wikikube-worker2001*<nowiki>}</nowiki> and (A:wikikube-master-codfw or A:wikikube-worker-codfw) * 11:30 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:30 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:21 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 18 hosts with reason: reboot & upgrade * 11:20 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker2001.codfw.wmnet with OS trixie * 11:16 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:15 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:14 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:14 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 11 hosts * 11:14 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 11 hosts * 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:08 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 11:02 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1053.eqiad.wmnet with OS trixie * 11:01 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:58 cgoubert@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 10:57 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:38 cgoubert@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker2001.codfw.wmnet with OS trixie * 10:38 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2001.codfw.wmnet * 10:38 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2001.codfw.wmnet * 10:38 cgoubert@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on P<nowiki>{</nowiki>wikikube-worker2001*<nowiki>}</nowiki> and (A:wikikube-master-codfw or A:wikikube-worker-codfw) * 10:35 topranks: adjust IBGP outbound policy on lsw1-e2-codfw [[phab:T423430|T423430]] towards ssw1-e1-codfw * 10:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 cgoubert@cumin1003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:27 cgoubert@cumin1003: END (FAIL) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=1) rolling reimage on A:wikikube-worker-codfw * 10:27 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker2001.codfw.wmnet with OS bookworm * 10:25 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 10:24 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:24 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:15 cgoubert@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 10:11 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 10:11 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 10:08 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:08 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:07 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 10:06 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 10:00 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:55 cgoubert@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker2001.codfw.wmnet with OS bookworm * 09:55 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2005-2006,2011-2012].codfw.wmnet * 09:55 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2005-2006,2011-2012].codfw.wmnet * 09:51 cgoubert@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on A:wikikube-worker-codfw * 09:41 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1053.eqiad.wmnet with reason: host reimage * 09:37 topranks: apply new IBGP outbound policy on lsw1-e2-codfw [[phab:T423430|T423430]] * 09:36 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:36 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1053.eqiad.wmnet with reason: host reimage * 09:16 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1053 * 09:16 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1053 * 09:15 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1053 * 09:15 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1053.eqiad.wmnet 48.32.64.10.in-addr.arpa 8.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:15 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1053.eqiad.wmnet 48.32.64.10.in-addr.arpa 8.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:15 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:15 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1053 - blake@cumin1003" * 09:15 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1053 - blake@cumin1003" * 09:11 blake@cumin1003: START - Cookbook sre.dns.netbox * 09:11 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1053 * 09:08 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1053.eqiad.wmnet with OS trixie * 09:08 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1053.eqiad.wmnet * 09:08 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1053.eqiad.wmnet * 09:08 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1053.eqiad.wmnet * 09:04 brouberol@dns1004: END - running authdns-update * 09:03 brouberol@dns1004: START - running authdns-update * 08:41 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e] (thin): Regular analytics weekly train THIN [analytics/refinery@1abf22ea] (duration: 02m 11s) * 08:38 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e] (thin): Regular analytics weekly train THIN [analytics/refinery@1abf22ea] * 08:38 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e]: Regular analytics weekly train [analytics/refinery@1abf22ea] (duration: 05m 17s) * 08:38 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:34 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:33 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e]: Regular analytics weekly train [analytics/refinery@1abf22ea] * 08:32 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@1abf22ea] (duration: 02m 03s) * 08:30 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@1abf22ea] * 08:30 JavierMonton: Deploying Refinery at {{Gerrit|1abf22ea}} for changes 1308121/T427068 1306491/T430020 and {{Gerrit|1308190}} * 08:29 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:29 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 08:24 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:24 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 08:18 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:18 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 08:00 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db[2183-2184].codfw.wmnet * 08:00 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for db[2183-2184].codfw.wmnet * 07:52 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:52 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 07:49 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 11 hosts with reason: reboot & upgrade * 07:47 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:47 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 07:44 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:44 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 07:23 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 10 hosts * 07:23 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 10 hosts * 06:45 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 10 hosts with reason: reboot & upgrade * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 41s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-09 == * 23:33 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] (duration: 13m 26s) * 23:29 ladsgroup@deploy2003: ladsgroup, jdlrobson: Continuing with deployment * 23:22 ladsgroup@deploy2003: ladsgroup, jdlrobson: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:20 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] * 22:57 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1165.eqiad.wmnet * 22:56 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1165.eqiad.wmnet * 22:56 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1165.eqiad.wmnet * 22:45 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1165.eqiad.wmnet with OS trixie * 22:38 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 22:37 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 22:37 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 22:37 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:37 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 22:25 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1165.eqiad.wmnet with reason: host reimage * 22:17 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1165.eqiad.wmnet with reason: host reimage * 22:13 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 22:12 rzl@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 22:04 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 22:04 rzl@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1165 * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1165 * 22:02 jasmine@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1165 * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1165.eqiad.wmnet 115.48.64.10.in-addr.arpa 5.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:02 jasmine@cumin2002: START - Cookbook sre.dns.wipe-cache wikikube-worker1165.eqiad.wmnet 115.48.64.10.in-addr.arpa 5.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1165 - jasmine@cumin2002" * 22:02 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1165 - jasmine@cumin2002" * 22:02 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 21:57 jasmine@cumin2002: START - Cookbook sre.dns.netbox * 21:55 jasmine@cumin2002: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1165 * 21:54 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-worker1165.eqiad.wmnet with OS trixie * 21:54 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 21:54 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1165.eqiad.wmnet * 21:53 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 21:53 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1165.eqiad.wmnet * 21:53 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1165.eqiad.wmnet * 21:53 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 21:47 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 21:45 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 21:43 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 21:43 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 21:42 maryum: Deploy fix for [[phab:T431684|T431684]] * 21:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 21:27 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] (duration: 34m 14s) * 21:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs1002 * 21:23 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs1002 * 21:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS trixie * 21:22 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 21:20 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 22s) * 21:20 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 21:16 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 21:16 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 21:15 ladsgroup@deploy2003: ladsgroup: Continuing with deployment * 21:13 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2002.codfw.wmnet with OS bookworm * 21:11 ladsgroup@deploy2003: ladsgroup: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:08 ladsgroup@cumin1003: END (PASS) - Cookbook sre.wikireplicas.update-views (exit_code=0) * 21:07 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 6 hosts with reason: reboots * 20:54 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecycle work - bking@cumin2003 * 20:53 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:53 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] * 20:51 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) * 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 20:47 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecycle work - bking@cumin2003 * 20:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 20:41 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:41 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) * 20:40 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host relforge1008.eqiad.wmnet * 20:40 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1009.eqiad.wmnet with OS trixie * 20:33 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:32 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:32 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:31 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:31 ladsgroup@cumin1003: END (PASS) - Cookbook sre.wikireplicas.update-views (exit_code=0) * 20:29 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1008.eqiad.wmnet * 20:24 rzl@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 20:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2002.codfw.wmnet with OS bookworm * 20:23 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host relforge1008.eqiad.wmnet * 20:23 rzl@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 20:23 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1008.eqiad.wmnet * 20:22 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:22 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:21 bking@cumin2003: END (ERROR) - Cookbook sre.elasticsearch.rolling-operation (exit_code=97) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:21 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1009.eqiad.wmnet with reason: host reimage * 20:16 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:15 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1009.eqiad.wmnet with reason: host reimage * 20:12 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) * 20:02 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 19:55 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1009.eqiad.wmnet with OS trixie * 19:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 19:43 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 19:30 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 19:28 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 19:27 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 19:25 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 18:42 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 18:41 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 18:16 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 18:15 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 17:45 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for doh5004.wikimedia.org * 17:45 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for doh5004.wikimedia.org * 17:38 ladsgroup@deploy2003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 17:35 ladsgroup@deploy2003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 17:29 ladsgroup@deploy2003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 17:26 ladsgroup@deploy2003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 17:09 mutante: zuul[12]00[123] - rebooting for maintenance * 17:09 ebernhardson: start full in-place reindex of eqiad cirrussearch cluster * 17:08 dzahn@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-cluster (exit_code=99) * 17:08 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-cluster * 17:03 ebernhardson: start full in-place reindex of codfw cirrussearch cluster * 16:59 mutante: stewards1001/stewards2001 - reboot for maintenance * 16:54 ebernhardson: start full in-place reindex of cloudelastic cluster * 16:53 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 16:52 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply * 16:49 mutante: planet1003/planet2003 - rebooting * 16:47 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on doh5004.wikimedia.org with reason: random high load, investigating * 15:55 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 15:54 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 15:51 jynus: restarting backupmon1001 * 15:49 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 14 hosts * 15:49 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 14 hosts * 15:47 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backupmon1001.eqiad.wmnet with reason: restart * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:06 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 15:06 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 14:59 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply * 14:58 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply * 14:51 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 14 hosts * 14:51 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 14 hosts * 14:49 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 6 hosts with reason: reboot & upgrade * 14:48 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet * 14:48 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet * 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:42 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1052.eqiad.wmnet * 14:42 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1052.eqiad.wmnet * 14:42 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1052.eqiad.wmnet * 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:31 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:31 elukey@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: sync * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:30 elukey@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: sync * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:28 elukey@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: sync * 14:28 elukey@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: sync * 14:26 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:20 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1052.eqiad.wmnet with OS trixie * 14:19 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 6 hosts with reason: reboot & upgrade * 14:18 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:15 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:15 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:13 elukey: update druid indexation job for webrequest_sampled_live - [[phab:T427068|T427068]] * 14:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:09 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:09 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for papaul - jhancock@cumin2002" * 14:09 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for papaul - jhancock@cumin2002" * 14:07 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:07 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:04 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 14:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cuminunpriv1001.eqiad.wmnet * 13:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb1003.eqiad.wmnet * 13:59 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1052.eqiad.wmnet with reason: host reimage * 13:57 moritzm: installing requests security updates * 13:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cuminunpriv1001.eqiad.wmnet * 13:55 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb1003.eqiad.wmnet * 13:53 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1052.eqiad.wmnet with reason: host reimage * 13:50 moritzm: installing python-cryptography security updates * 13:47 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb2003.codfw.wmnet * 13:44 Msz2001: UTC afternoon config+backport window is done * 13:44 Msz2001: Updated `logging` on `metawiki` to fix log performers, [[phab:T431176|T431176]]#12105297 * 13:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb2003.codfw.wmnet * 13:43 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 13:43 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt1002.wikimedia.org * 13:41 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] (duration: 07m 30s) * 13:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt1002.wikimedia.org * 13:37 mszwarc@deploy2003: mszwarc: Continuing with deployment * 13:36 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1052 * 13:36 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1052 * 13:35 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:35 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1052 * 13:35 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1052.eqiad.wmnet 47.32.64.10.in-addr.arpa 7.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:35 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1052.eqiad.wmnet 47.32.64.10.in-addr.arpa 7.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:35 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:35 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1052 - blake@cumin1003" * 13:35 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1052 - blake@cumin1003" * 13:34 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] * 13:31 blake@cumin1003: START - Cookbook sre.dns.netbox * 13:31 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1052 * 13:30 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1052.eqiad.wmnet with OS trixie * 13:30 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1052.eqiad.wmnet * 13:29 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1052.eqiad.wmnet * 13:29 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1052.eqiad.wmnet * 13:17 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] (duration: 11m 26s) * 13:13 jforrester@deploy2003: jforrester: Continuing with deployment * 13:08 jforrester@deploy2003: jforrester: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:06 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] * 12:54 cgoubert@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply * 12:54 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:52 cgoubert@deploy2003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply * 12:45 cgoubert@deploy2003: helmfile [codfw] DONE helmfile.d/services/mobileapps: apply * 12:44 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:44 cgoubert@deploy2003: helmfile [codfw] START helmfile.d/services/mobileapps: apply * 12:43 cgoubert@deploy2003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 12:43 cgoubert@deploy2003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 12:42 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast4006.wikimedia.org * 12:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt2002.wikimedia.org * 12:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast7002.wikimedia.org * 12:18 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast4006.wikimedia.org * 12:18 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host ml-serve1004 * 12:18 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host ml-serve1004 * 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt2002.wikimedia.org * 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast7002.wikimedia.org * 12:10 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backup[2003,2014].codfw.wmnet with reason: reboot & upgrade * 12:10 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt-staging2001.codfw.wmnet * 12:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid1003.eqiad.wmnet * 12:06 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt-staging2001.codfw.wmnet * 12:05 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid1003.eqiad.wmnet * 12:03 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backup[1003,1014].eqiad.wmnet with reason: reboot & upgrade * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid2003.codfw.wmnet * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host irc1003.wikimedia.org * 11:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid2003.codfw.wmnet * 11:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host irc1003.wikimedia.org * 11:55 jmm@dns1004: END - running authdns-update * 11:53 jmm@dns1004: START - running authdns-update * 11:50 jmm@dns1004: END - running authdns-update * 11:48 jmm@dns1004: START - running authdns-update * 11:27 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host irc2003.wikimedia.org * 11:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host irc2003.wikimedia.org * 11:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint2001.codfw.wmnet * 11:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint1001.eqiad.wmnet * 11:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint2001.codfw.wmnet * 11:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint1001.eqiad.wmnet * 11:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-rw2001.wikimedia.org * 11:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-rw1001.wikimedia.org * 11:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-rw2001.wikimedia.org * 11:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-rw1001.wikimedia.org * 11:03 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon1003.wikimedia.org * 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2005.codfw.wmnet * 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2005.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 10:59 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2005.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 10:57 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon1003.wikimedia.org * 10:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon2002.wikimedia.org * 10:55 jmm@cumin2003: START - Cookbook sre.dns.netbox * 10:51 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon2002.wikimedia.org * 10:51 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:50 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2005.codfw.wmnet * 10:41 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:40 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host ml-serve1003 * 10:40 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host ml-serve1003 * 10:39 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2033.codfw.wmnet * 10:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install2005.wikimedia.org * 10:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install1005.wikimedia.org * 10:35 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1004.eqiad.wmnet with OS bookworm * 10:31 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install1005.wikimedia.org * 10:31 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install2005.wikimedia.org * 10:30 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install4004.wikimedia.org * 10:30 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install3004.wikimedia.org * 10:29 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install3004.wikimedia.org * 10:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install4004.wikimedia.org * 10:23 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 10:21 moritzm: failover Ganeti master in codfw/routed to ganeti2034 [[phab:T430928|T430928]] * 10:19 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.addnode (exit_code=0) for new host ganeti2031.codfw.wmnet to cluster codfw and group B * 10:19 moritzm: readded ganeti2031 to the codfw Ganeti cluster [[phab:T430910|T430910]] * 10:18 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1004.eqiad.wmnet with reason: host reimage * 10:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install5004.wikimedia.org * 10:18 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1003 * 10:18 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1003 * 10:17 jmm@cumin2003: START - Cookbook sre.ganeti.addnode for new host ganeti2031.codfw.wmnet to cluster codfw and group B * 10:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install6003.wikimedia.org * 10:16 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install5004.wikimedia.org * 10:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install6003.wikimedia.org * 10:15 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1004.eqiad.wmnet with reason: host reimage * 10:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1001.eqiad.wmnet * 10:14 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 10:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2008.wikimedia.org * 10:00 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ml-serve1004 * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1004 * 09:57 jmm@cumin2003: START - Cookbook sre.dns.netbox * 09:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install7002.wikimedia.org * 09:57 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1004 * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ml-serve1004.eqiad.wmnet 50.48.64.10.in-addr.arpa 0.5.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:57 klausman@cumin1003: START - Cookbook sre.dns.wipe-cache ml-serve1004.eqiad.wmnet 50.48.64.10.in-addr.arpa 0.5.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1004 - klausman@cumin1003" * 09:56 klausman@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1004 - klausman@cumin1003" * 09:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-coord1001.eqiad.wmnet * 09:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 09:55 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow7002.magru.wmnet * 09:52 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-coord1001.eqiad.wmnet * 09:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 09:52 klausman@cumin1003: START - Cookbook sre.dns.netbox * 09:50 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install7002.wikimedia.org * 09:50 klausman@cumin1003: START - Cookbook sre.hosts.move-vlan for host ml-serve1004 * 09:50 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1004.eqiad.wmnet with OS bookworm * 09:50 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1003.eqiad.wmnet with OS bookworm * 09:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1001.eqiad.wmnet * 09:49 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 09:49 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2008.wikimedia.org * 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2007.codfw.wmnet * 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2007.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 09:49 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow7002.magru.wmnet * 09:49 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2007.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 09:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard1003.eqiad.wmnet * 09:39 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard2003.codfw.wmnet * 09:39 jmm@cumin2003: START - Cookbook sre.dns.netbox * 09:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard1003.eqiad.wmnet * 09:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor1003.eqiad.wmnet * 09:35 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard2003.codfw.wmnet * 09:34 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2007.codfw.wmnet * 09:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor1003.eqiad.wmnet * 09:33 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor-dev2001.codfw.wmnet * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor2003.codfw.wmnet * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sretest1006.eqiad.wmnet * 09:27 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 09:25 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor-dev2001.codfw.wmnet * 09:25 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor2003.codfw.wmnet * 09:23 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] (duration: 06m 27s) * 09:23 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2205: codfw rack B4 repool after maintenance * 09:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host sretest1006.eqiad.wmnet * 09:23 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2204: codfw rack B4 repool after maintenance * 09:19 urbanecm@deploy2003: urbanecm: Continuing with deployment * 09:19 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:18 jmm@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 6 hosts with reason: reboot * 09:17 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] * 09:08 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ml-serve1003 * 09:08 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1003 * 09:07 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1003 * 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ml-serve1003.eqiad.wmnet 81.32.64.10.in-addr.arpa 1.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:07 klausman@cumin1003: START - Cookbook sre.dns.wipe-cache ml-serve1003.eqiad.wmnet 81.32.64.10.in-addr.arpa 1.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1003 - klausman@cumin1003" * 09:06 klausman@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1003 - klausman@cumin1003" * 08:58 klausman@cumin1003: START - Cookbook sre.dns.netbox * 08:57 klausman@cumin1003: START - Cookbook sre.hosts.move-vlan for host ml-serve1003 * 08:57 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1003.eqiad.wmnet with OS bookworm * 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=0) rolling reimage on P<nowiki>{</nowiki>ml-serve1003.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet * 08:55 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet * 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1003.eqiad.wmnet with OS bookworm * 08:39 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 08:38 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool db2205: codfw rack B4 repool after maintenance * 08:37 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool db2204: codfw rack B4 repool after maintenance * 08:36 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 08:35 hashar@deploy2003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 08:32 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:32 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:31 hashar@deploy2003: Rolling back deployment * 08:26 moritzm: failover Ganeti master in codfw to ganeti2048 [[phab:T430928|T430928]] * 08:16 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1003.eqiad.wmnet with OS bookworm * 08:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2004.codfw.wmnet * 08:16 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet * 08:16 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet * 08:16 klausman@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on P<nowiki>{</nowiki>ml-serve1003.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 08:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2002.codfw.wmnet * 08:15 XioNoX: lsw1-b4-codfw> request system reboot - [[phab:T430910|T430910]] * 08:15 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b4-codfw,lsw1-b4-codfw IPv6,lsw1-b4-codfw.mgmt with reason: Switch maintenance * 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for codfw rack B4 * 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:10 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2004.codfw.wmnet * 08:10 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2002.codfw.wmnet * 08:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2205: codfw rack B4 depool for maintenance * 08:08 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool db2205: codfw rack B4 depool for maintenance * 08:08 jmm@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin2003.codfw.wmnet * 08:08 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2204: codfw rack B4 depool for maintenance * 08:08 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool db2204: codfw rack B4 depool for maintenance * 08:08 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 27 hosts with reason: codfw rack B4 depool for maintenance * 08:03 jmm@cumin2002: START - Cookbook sre.hosts.reboot-single for host cumin2003.codfw.wmnet * 07:56 ayounsi@cumin1003: START - Cookbook sre.network.depool-rack with action 'depool' for codfw rack B4 * 07:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1008.eqiad.wmnet with OS trixie * 07:49 wmde-fisch@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] (duration: 08m 36s) * 07:44 wmde-fisch@deploy2003: wmde-fisch: Continuing with deployment * 07:43 wmde-fisch@deploy2003: wmde-fisch: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:41 wmde-fisch@deploy2003: Started scap sync-world: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] * 07:35 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1008.eqiad.wmnet with reason: host reimage * 07:31 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1008.eqiad.wmnet with reason: host reimage * 07:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1008.eqiad.wmnet with OS trixie * 07:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 07:00 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 06:59 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1008.eqiad.wmnet with OS trixie * 06:57 Emperor: rebalance thanos swift rings after previous re-image of thanos-fe1004 to trixie * 06:47 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1008.eqiad.wmnet with OS trixie * 04:10 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 14 days, 0:00:00 on cp6008.drmrs.wmnet with reason: Hardware failure - [[phab:T431651|T431651]] * 03:55 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp6008.* * 03:29 ryankemper: [[phab:T431311|T431311]] Repooled eqiad cirrussearch clusters (`chi/omega/psi`) following completion of OpenSearch 2.19 migration * 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad * 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=eqiad * 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 31s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-08 == * 23:52 Amir1: ladsgroup@deploy2003:~$ mwscript-k8s --follow -- extensions/ORES/maintenance/PurgeScoreCache.php --wiki=simplewiki --model damaging --old ([[phab:T431159|T431159]]) * 23:46 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 23:46 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing PTR for 2001:df2:e500:fe08::1 - cmooney@cumin1003" * 23:46 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing PTR for 2001:df2:e500:fe08::1 - cmooney@cumin1003" * 23:40 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 23:16 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 23:15 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 22:42 rzl@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 22:40 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] (duration: 12m 55s) * 22:40 rzl@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 22:37 rzl@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 22:36 rzl@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 22:35 rzl@deploy2003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 22:34 urbanecm@deploy2003: urbanecm: Continuing with deployment * 22:33 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:33 rzl@deploy2003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 22:32 rzl@deploy2003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 22:30 rzl@deploy2003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 22:30 rzl@deploy2003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 22:29 rzl@deploy2003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 22:27 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] * 22:26 rzl@deploy2003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 22:22 rzl@deploy2003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 22:21 rzl@deploy2003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 22:19 rzl@deploy2003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 22:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 22:17 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 22:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 22:13 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 22:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 22:13 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 22:09 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 22:06 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 22:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1094.eqiad.wmnet with OS trixie * 22:01 urbanecm: Make https://test.wikipedia.org/w/index.php?title=MediaWiki:GrowthExperimentsSuggestedEdits.json&diff=prev&oldid=750552 with GrowthExperiments disabled (via mw-experimental), then run `\MediaWiki\MediaWikiServices::getInstance()->get('CommunityConfiguration.ProviderFactory')->newProvider('GrowthSuggestedEdits')->getStore()->invalidate()` ([[phab:T431625|T431625]]) * 21:56 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d2-codfw * 21:55 urbanecm@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 21:55 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d2-codfw * 21:55 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c4-codfw * 21:55 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c4-codfw * 21:55 urbanecm@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2002 * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2002 * 21:54 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2002 * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2002.codfw.wmnet 50.32.192.10.in-addr.arpa 0.5.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:54 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2002.codfw.wmnet 50.32.192.10.in-addr.arpa 0.5.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2002 - bking@cumin2003" * 21:54 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2002 - bking@cumin2003" * 21:49 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:49 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2002 * 21:49 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2002.codfw.wmnet with OS trixie * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1094.eqiad.wmnet with reason: host reimage * 21:42 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 21:39 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 21:37 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1094.eqiad.wmnet with reason: host reimage * 21:36 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 21:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 21:29 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 21:27 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 21:22 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1094.eqiad.wmnet with OS trixie * 21:21 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host restbase2039.codfw.wmnet with OS bullseye * 21:21 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin2002" * 21:21 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin2002" * 21:04 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on restbase2039.codfw.wmnet with reason: host reimage * 21:00 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on restbase2039.codfw.wmnet with reason: host reimage * 20:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1073.eqiad.wmnet with OS trixie * 20:48 mutante: deploy2003 - kill 1102 (stunnel4) ; systemctl start stunnel4 ([[phab:T418262|T418262]]) * 20:42 cjming@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] (duration: 33m 02s) * 20:42 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host restbase2039.codfw.wmnet with OS bullseye * 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1073.eqiad.wmnet with reason: host reimage * 20:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1073.eqiad.wmnet with reason: host reimage * 20:30 cjming@deploy2003: cjming: Continuing with deployment * 20:28 cjming@deploy2003: cjming: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1098.eqiad.wmnet with OS trixie * 20:13 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1073.eqiad.wmnet with OS trixie * 20:09 cjming@deploy2003: Started scap sync-world: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] * 20:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1098.eqiad.wmnet with reason: host reimage * 19:56 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1098.eqiad.wmnet with reason: host reimage * 19:55 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d4-codfw * 19:54 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d4-codfw * 19:54 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c1-codfw * 19:54 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c1-codfw * 19:52 mutante: restarting gerrit on gerrit.wikimedia.org (gerrit2003) * 19:48 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2331.codfw.wmnet * 19:48 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2331.codfw.wmnet * 19:48 mutante: restarting gerrit on gerrit-replica.wikimedia.org (gerrit1003) * 19:46 mutante: restarting gerrit on gerrit-spare.wikimedia.org (gerrit2002) * 19:43 jasmine@cumin2002: conftool action : set/pooled=yes; selector: name=wikikube-worker2331.codfw.wmnet,cluster=kubernetes,service=kubesvc * 19:43 jasmine@cumin2002: conftool action : set/weight=10; selector: name=wikikube-worker2331.codfw.wmnet,cluster=kubernetes,service=kubesvc * 19:40 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1098.eqiad.wmnet with OS trixie * 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d5-codfw * 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d5-codfw * 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c7-codfw * 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c7-codfw * 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c5-codfw * 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c5-codfw * 19:30 jasmine_: ran homer on lsw1-d8-codfw, adding wikikube-worker2331 to cluster * 19:29 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1100.eqiad.wmnet with OS trixie * 19:20 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d8-codfw * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d8-codfw * 19:19 mutante: gerrit - replacing private key for registerEmail verification * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-magru * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device cr2-magru * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d7-codfw * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d7-codfw * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d3-codfw * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d1-codfw * 19:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d1-codfw * 19:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c2-codfw * 19:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c2-codfw * 19:11 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-codfw * 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-magru * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device cr1-magru * 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d8-codfw * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d8-codfw * 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d6-codfw * 19:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1100.eqiad.wmnet with reason: host reimage * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d6-codfw * 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c6-codfw * 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c6-codfw * 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c3-codfw * 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c3-codfw * 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b4-magru * 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device asw1-b4-magru * 19:08 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b3-magru * 19:08 cmooney@cumin1003: START - Cookbook sre.network.tls for network device asw1-b3-magru * 19:05 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1100.eqiad.wmnet with reason: host reimage * 19:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1122.eqiad.wmnet with OS trixie * 19:00 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 18:59 topranks: rolling out update to BGP ACL on Nokia Switches eqiad, codfw & ulsfo [[phab:T425703|T425703]] * 18:58 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 18:57 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 18:55 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 18:53 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 18:52 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 18:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1100.eqiad.wmnet with OS trixie * 18:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1068.eqiad.wmnet with OS trixie * 18:47 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1102.eqiad.wmnet with OS trixie * 18:47 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 18:46 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1122.eqiad.wmnet with reason: host reimage * 18:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1122.eqiad.wmnet with reason: host reimage * 18:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1068.eqiad.wmnet with reason: host reimage * 18:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1122 * 18:26 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1122 * 18:25 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1122 * 18:25 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1122.eqiad.wmnet 31.48.64.10.in-addr.arpa 1.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:25 bking@cumin2003: START - Cookbook sre.dns.wipe-cache cirrussearch1122.eqiad.wmnet 31.48.64.10.in-addr.arpa 1.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:25 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:25 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1122 - bking@cumin2003" * 18:25 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1122 - bking@cumin2003" * 18:21 rzl@deploy2003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 18:21 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1068.eqiad.wmnet with reason: host reimage * 18:21 rzl@deploy2003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 18:21 rzl@deploy2003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 18:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 18:19 rzl@deploy2003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 18:19 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:18 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1122 * 18:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1122.eqiad.wmnet with OS trixie * 18:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 18:15 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 18:13 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 18:13 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 18:10 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 18:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1068.eqiad.wmnet with OS trixie * 18:01 kamila@deploy2003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 18m 29s) * 18:00 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:55 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:42 kamila@deploy2003: Started scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] * 17:42 kamila@deploy2003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 19m 50s) * 17:42 kamila@deploy2003: Rolling back deployment * 17:35 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:31 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet * 17:18 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet * 17:16 kamila@deploy1003: Unlocked for deployment [MediaWiki]: switching deployment server (duration: 22m 07s) * 17:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 17:11 kamila@dns1005: END - running authdns-update * 17:09 kamila@dns1005: START - running authdns-update * 17:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 17:04 jasmine@cumin2002: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1164.eqiad.wmnet * 17:04 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1164.eqiad.wmnet * 17:04 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1164.eqiad.wmnet * 16:56 kamila@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on releases2003.codfw.wmnet,releases1003.eqiad.wmnet with reason: Deployment server switchover * 16:54 kamila@deploy1003: Locking from deployment [MediaWiki]: switching deployment server * 16:53 kamila@deploy1003: Unlocked for deployment [MediaWiki]: switching deployment server (duration: 04m 02s) * 16:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie * 16:49 kamila@deploy1003: Locking from deployment [MediaWiki]: switching deployment server * 16:46 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1095.eqiad.wmnet with OS trixie * 16:45 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1093.eqiad.wmnet with OS trixie * 16:43 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1164.eqiad.wmnet with OS trixie * 16:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1095.eqiad.wmnet with reason: host reimage * 16:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 16:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 16:23 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1164.eqiad.wmnet with reason: host reimage * 16:18 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on cirrussearch1093.eqiad.wmnet with reason: host reimage * 16:16 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1164.eqiad.wmnet with reason: host reimage * 16:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1095.eqiad.wmnet with reason: host reimage * 16:09 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 16:09 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 16:08 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1093.eqiad.wmnet with reason: host reimage * 15:59 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Pool test * 15:59 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:59 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 15:59 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Pool test * 15:58 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Depool test * 15:58 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:58 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 15:58 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Depool test * 15:57 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1164 * 15:57 jasmine@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1164 * 15:57 jasmine@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1164 * 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1164.eqiad.wmnet 114.48.64.10.in-addr.arpa 4.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:56 jasmine@cumin2002: START - Cookbook sre.dns.wipe-cache wikikube-worker1164.eqiad.wmnet 114.48.64.10.in-addr.arpa 4.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1164 - jasmine@cumin2002" * 15:56 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1164 - jasmine@cumin2002" * 15:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1093.eqiad.wmnet with OS trixie * 15:51 jasmine@cumin2002: START - Cookbook sre.dns.netbox * 15:51 jasmine@cumin2002: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1164 * 15:50 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-worker1164.eqiad.wmnet with OS trixie * 15:50 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1164.eqiad.wmnet * 15:50 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1164.eqiad.wmnet * 15:50 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1164.eqiad.wmnet * 15:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1095.eqiad.wmnet with OS trixie * 15:42 jasmine@cumin2002: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1164.eqiad.wmnet * 15:42 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1164.eqiad.wmnet * 15:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:42 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1164.eqiad.wmnet * 15:42 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1164.eqiad.wmnet * 15:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 15:39 elukey@cumin1003: START - Cookbook sre.hosts.provision for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 15:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1007.eqiad.wmnet with OS trixie * 15:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Pool test * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1007.eqiad.wmnet with reason: host reimage * 15:15 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 15:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet * 15:15 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 15:15 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1007.eqiad.wmnet with reason: host reimage * 15:15 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:14 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Pool test * 15:14 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet * 15:14 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2228: Depool test * 15:14 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db2228: Depool test * 15:10 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 15:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:08 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 15:06 blake@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 15:06 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet * 15:06 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 15:06 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 15:05 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet * 15:05 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 15:04 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:04 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 15:04 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:03 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 15:03 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 15:03 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:03 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T430909|T430909]] * 15:03 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:03 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 15:03 swfrench-wmf: restarted eqsin, codfw confds - [[phab:T430909|T430909]] * 15:03 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test1002.eqiad.wmnet * 15:02 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet * 14:59 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:59 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:55 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1007.eqiad.wmnet with OS trixie * 14:52 swfrench-wmf: restarted ulsfo confds, confirmed now connected to codfw backends except those using wikimedia.org SRV record - [[phab:T430909|T430909]] * 14:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:49 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:41 moritzm: uninstalling dhcpcd-base from trixie hosts which still have it installed [[phab:T414341|T414341]] * 14:40 sukhe: sudo cumin -b1 -s120 "P<nowiki>{</nowiki>lvs2011*<nowiki>}</nowiki> or P<nowiki>{</nowiki>lvs2012*<nowiki>}</nowiki>" "systemctl restart pybal.service" * 14:39 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:39 mvernon@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host thanos-be1007.eqiad.wmnet with OS trixie * 14:37 sukhe: restart pybal on lvs2013 to revert back to conf2004 * 14:35 sukhe: restart pybal on lvs2014 to revert back to conf2004 * 14:34 swfrench-wmf: switched codfw, eqsin, ulsfo etcd client SRV records back to codfw - [[phab:T430909|T430909]] * 14:32 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1002.eqiad.wmnet * 14:32 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet * 14:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1007.eqiad.wmnet with OS trixie * 14:31 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:31 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:31 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:30 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:30 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Pool test * 14:30 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:29 swfrench@dns1004: END - running authdns-update * 14:29 moritzm: installing jackson-core security updates * 14:27 swfrench@dns1004: START - running authdns-update * 14:22 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:22 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:22 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1119.eqiad.wmnet with OS trixie * 14:22 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:21 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:20 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:20 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 14:20 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:19 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 14:19 moritzm: installing librabbitmq security updates * 14:19 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1002.eqiad.wmnet * 14:18 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet * 14:16 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:15 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:15 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:14 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Pool test * 14:14 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox) * 14:14 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet * 14:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 14:08 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 14:07 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet * 14:05 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1118.eqiad.wmnet with OS trixie * 14:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test1001.eqiad.wmnet * 14:00 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-worker@eqiad * 14:00 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 13:59 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 13:58 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 13:57 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1119.eqiad.wmnet with reason: host reimage * 13:54 moritzm: installing libcap2 security updates * 13:53 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1119.eqiad.wmnet with reason: host reimage * 13:52 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet * 13:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) * 13:52 fceratto@cumin1003: START - Cookbook sre.mysql.depool * 13:50 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-worker@eqiad * 13:50 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1051.eqiad.wmnet * 13:50 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1051.eqiad.wmnet * 13:50 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1051.eqiad.wmnet * 13:49 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 13:45 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1081.eqiad.wmnet with OS trixie * 13:41 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1119 * 13:41 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1119 * 13:40 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1119 * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1119.eqiad.wmnet 97.32.64.10.in-addr.arpa 7.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1119.eqiad.wmnet 97.32.64.10.in-addr.arpa 7.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1119 - atsuko@cumin1003" * 13:40 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1119 - atsuko@cumin1003" * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1118.eqiad.wmnet with reason: host reimage * 13:39 moritzm: installing krb5 security updates * 13:37 Lucas_WMDE: UTC afternoon backport+config window done * 13:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1006.eqiad.wmnet with OS trixie * 13:36 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1118.eqiad.wmnet with reason: host reimage * 13:36 atsuko@cumin1003: START - Cookbook sre.dns.netbox * 13:35 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] (duration: 07m 46s) * 13:34 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1119 * 13:34 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1119.eqiad.wmnet with OS trixie * 13:30 sbisson@deploy1003: sbisson: Continuing with deployment * 13:30 moritzm: installing openssh security updates * 13:30 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling restart_daemons on A:wikidough * 13:29 sbisson@deploy1003: sbisson: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:27 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] * 13:26 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1051.eqiad.wmnet with OS trixie * 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1081.eqiad.wmnet with reason: host reimage * 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1118 * 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1118 * 13:22 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] (duration: 12m 12s) * 13:21 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1081.eqiad.wmnet with reason: host reimage * 13:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1006.eqiad.wmnet with reason: host reimage * 13:18 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1118 * 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1118.eqiad.wmnet 90.32.64.10.in-addr.arpa 0.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:18 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1118.eqiad.wmnet 90.32.64.10.in-addr.arpa 0.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1118 - atsuko@cumin1003" * 13:18 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1118 - atsuko@cumin1003" * 13:17 stran@deploy1003: stran: Continuing with deployment * 13:16 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough * 13:15 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:13 atsuko@cumin1003: START - Cookbook sre.dns.netbox * 13:12 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1118 * 13:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1006.eqiad.wmnet with reason: host reimage * 13:12 stran@deploy1003: stran: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:12 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1118.eqiad.wmnet with OS trixie * 13:10 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] * 13:05 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-worker@codfw * 13:05 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 13:05 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1081.eqiad.wmnet with OS trixie * 13:05 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1051.eqiad.wmnet with reason: host reimage * 13:04 moritzm: installing jq security updates * 13:04 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 13:01 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1051.eqiad.wmnet with reason: host reimage * 12:58 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-worker@codfw * 12:52 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 12:50 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host thanos-be1006.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1051 * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1051 * 12:43 moritzm: installing Python 3.11 security updates * 12:43 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1051 * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1051.eqiad.wmnet 46.32.64.10.in-addr.arpa 6.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:43 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1051.eqiad.wmnet 46.32.64.10.in-addr.arpa 6.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1051 - blake@cumin1003" * 12:43 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1051 - blake@cumin1003" * 12:38 blake@cumin1003: START - Cookbook sre.dns.netbox * 12:38 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1051 * 12:38 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1051.eqiad.wmnet with OS trixie * 12:37 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1051.eqiad.wmnet * 12:36 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1051.eqiad.wmnet * 12:36 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1051.eqiad.wmnet * 12:34 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1006.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 12:34 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1006.eqiad.wmnet with OS trixie * 12:27 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 12:27 mvernon@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host thanos-be1006.eqiad.wmnet with OS trixie * 12:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:02 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 12:01 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1006.eqiad.wmnet with OS trixie * 11:43 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 11:38 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1076.eqiad.wmnet with OS trixie * 11:26 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1075.eqiad.wmnet with OS trixie * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2047.codfw.wmnet * 11:19 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2047.codfw.wmnet * 11:19 moritzm: temporarily remove ganeti2031 from codfw cluster [[phab:T430910|T430910]] * 11:08 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1076.eqiad.wmnet with reason: host reimage * 11:08 moritzm: installing Linux 6.1.176 on Bookworm servers * 11:03 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1076.eqiad.wmnet with reason: host reimage * 11:00 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1075.eqiad.wmnet with reason: host reimage * 10:56 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1075.eqiad.wmnet with reason: host reimage * 10:47 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1076.eqiad.wmnet with OS trixie * 10:46 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1074.eqiad.wmnet with OS trixie * 10:45 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1005.eqiad.wmnet with OS trixie * 10:40 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1075.eqiad.wmnet with OS trixie * 10:32 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2031.codfw.wmnet * 10:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1005.eqiad.wmnet with reason: host reimage * 10:25 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1005.eqiad.wmnet with reason: host reimage * 10:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1074.eqiad.wmnet with reason: host reimage * 10:17 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1074.eqiad.wmnet with reason: host reimage * 10:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1005.eqiad.wmnet with OS trixie * 10:12 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet * 10:04 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 10:01 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet * 10:01 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1074.eqiad.wmnet with OS trixie * 10:01 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 09:43 cgoubert@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/aux-k8s-services/redioscope: apply * 09:43 cgoubert@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/aux-k8s-services/redioscope: apply * 09:43 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply * 09:35 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply * 09:34 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 09:34 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 09:33 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 41 days, 15:00:00 on db2252.codfw.wmnet with reason: Test * 09:32 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 09:32 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 09:31 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: codfw rack B3 pool after maintenance * 09:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 09:07 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 09:07 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 09:02 ladsgroup@cumin1003: END (PASS) - Cookbook sre.mysql.sanitarium_restart (exit_code=0) * 08:57 topranks: merge patch to shift eqiad <-> esams traffic onto new 40G circuit * 08:54 hashar@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1004.eqiad.wmnet with OS trixie * 08:50 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 08:50 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitarium_restart (exit_code=99) * 08:50 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 08:45 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool es2051: codfw rack B3 pool after maintenance * 08:44 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:44 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:43 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2007.codfw.wmnet * 08:43 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2007.codfw.wmnet * 08:42 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2031.codfw.wmnet * 08:41 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2031.codfw.wmnet * 08:40 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2031.codfw.wmnet * 08:38 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:38 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:35 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.sanitize-wiki (exit_code=97) Managing sanitization for wikis minwikiquote in section s3 * 08:33 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis minwikiquote in section s3 * 08:32 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Checking sanitization for wikis minwikiquote in section s5 * 08:30 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Checking sanitization for wikis minwikiquote in section s5 * 08:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Managing sanitization for wikis minwikiquote in section s5 * 08:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1004.eqiad.wmnet with reason: host reimage * 08:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1004.eqiad.wmnet with reason: host reimage * 08:23 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:22 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis minwikiquote in section s5 * 08:19 XioNoX: lsw1-b3-codfw> request system reboot - [[phab:T430909|T430909]] * 08:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Checking sanitization for wikis minwikiquote in section s5 * 08:17 hashar@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Checking sanitization for wikis minwikiquote in section s5 * 08:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for codfw rack B3 * 08:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2007.codfw.wmnet * 08:15 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lsw1-b3-codfw,lsw1-b3-codfw IPv6,lsw1-b3-codfw.mgmt with reason: Switch maintenance * 08:15 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2007.codfw.wmnet * 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:07 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:06 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: codfw rack B3 depool for maintenance * 08:05 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool es2051: codfw rack B3 depool for maintenance * 08:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1004.eqiad.wmnet with OS trixie * 08:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1005.eqiad.wmnet with OS trixie * 08:03 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 21 hosts with reason: codfw rack B3 depool for maintenance * 07:56 ayounsi@cumin1003: START - Cookbook sre.network.depool-rack with action 'depool' for codfw rack B3 * 07:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1005.eqiad.wmnet with reason: host reimage * 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1005.eqiad.wmnet with reason: host reimage * 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1005.eqiad.wmnet with OS bookworm * 07:29 moritzm: installing gnutls28 security updates * 07:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1005.eqiad.wmnet with OS trixie * 07:13 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1125.eqiad.wmnet with OS trixie * 07:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1005.eqiad.wmnet with reason: host reimage * 07:07 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aux-k8s-etcd1005.eqiad.wmnet with reason: host reimage * 06:56 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1005.eqiad.wmnet with OS bookworm * 06:54 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1125.eqiad.wmnet with reason: host reimage * 06:52 elukey: upgrade all trixie hosts to pywmflib 3.1 - [[phab:T430552|T430552]] * 06:50 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1125.eqiad.wmnet with reason: host reimage * 06:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 06:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 06:38 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1125.eqiad.wmnet with OS trixie * 05:42 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1107.eqiad.wmnet with OS trixie * 05:35 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1124.eqiad.wmnet with OS trixie * 05:31 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1101.eqiad.wmnet with OS trixie * 05:21 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1107.eqiad.wmnet with reason: host reimage * 05:17 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1124.eqiad.wmnet with reason: host reimage * 05:13 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1107.eqiad.wmnet with reason: host reimage * 05:13 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1101.eqiad.wmnet with reason: host reimage * 05:11 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1124.eqiad.wmnet with reason: host reimage * 05:10 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1101.eqiad.wmnet with reason: host reimage * 04:58 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1124.eqiad.wmnet with OS trixie * 04:56 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1107.eqiad.wmnet with OS trixie * 04:55 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1101.eqiad.wmnet with OS trixie * 02:27 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] (duration: 08m 14s) * 02:22 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 02:21 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 02:19 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] * 01:59 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] (duration: 09m 46s) * 01:55 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 01:51 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 01:49 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] * 01:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1099.eqiad.wmnet with OS trixie * 00:57 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1110.eqiad.wmnet with OS trixie * 00:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1099.eqiad.wmnet with reason: host reimage * 00:41 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1099.eqiad.wmnet with reason: host reimage * 00:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1110.eqiad.wmnet with reason: host reimage * 00:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1110.eqiad.wmnet with reason: host reimage * 00:26 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1099.eqiad.wmnet with OS trixie * 00:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1110.eqiad.wmnet with OS trixie == 2026-07-07 == * 22:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1097.eqiad.wmnet with OS trixie * 22:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1097.eqiad.wmnet with reason: host reimage * 22:24 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1097.eqiad.wmnet with reason: host reimage * 22:14 hashar: Restarting Gerrit on gerrit2002 and gerrit1003 (replicas) * 22:09 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1097.eqiad.wmnet with OS trixie * 22:07 hashar: Restarting Gerrit on gerrit2003 * 21:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 21:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 21:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 21:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 21:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1108.eqiad.wmnet with OS trixie * 20:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1091.eqiad.wmnet with OS trixie * 20:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1090.eqiad.wmnet with OS trixie * 20:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1108.eqiad.wmnet with reason: host reimage * 20:36 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1091.eqiad.wmnet with reason: host reimage * 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1090.eqiad.wmnet with reason: host reimage * 20:33 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1091.eqiad.wmnet with reason: host reimage * 20:30 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1108.eqiad.wmnet with reason: host reimage * 20:30 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1006.eqiad.wmnet * 20:30 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1090.eqiad.wmnet with reason: host reimage * 20:30 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1006.eqiad.wmnet * 20:27 jasmine_: "homer lsw1-c2-eqiad* commit "Added new stacked control plane wikikube-ctrl1006"" * 20:22 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] (duration: 07m 29s) * 20:20 jasmine_: "homer "cr*eqiad*" commit "Added new stacked control plane wikikube-ctrl1006"" * 20:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1091.eqiad.wmnet with OS trixie * 20:17 arlolra@deploy1003: arlolra: Continuing with deployment * 20:16 arlolra@deploy1003: arlolra: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:16 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1090.eqiad.wmnet with OS trixie * 20:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1108.eqiad.wmnet with OS trixie * 20:14 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] * 20:09 cwhite: remove 2026-04 swift log archives from centrallog2002 to free some space * 20:01 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=93) for host cirrussearch1108.eqiad.wmnet with OS trixie * 19:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1108.eqiad.wmnet with OS trixie * 19:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1090.eqiad.wmnet with OS trixie * 19:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1109.eqiad.wmnet with OS trixie * 19:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1092.eqiad.wmnet with OS trixie * 19:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1123.eqiad.wmnet with OS trixie * 19:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1109.eqiad.wmnet with reason: host reimage * 19:19 cdobbins@cumin2002: conftool action : set/pooled=yes; selector: name=dns7002.* * 19:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1092.eqiad.wmnet with reason: host reimage * 19:17 jasmine@dns1004: END - running authdns-update * 19:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1109.eqiad.wmnet with reason: host reimage * 19:15 jasmine@dns1004: START - running authdns-update * 19:15 cdobbins@dns1004: END - running authdns-update * 19:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1123.eqiad.wmnet with reason: host reimage * 19:13 cdobbins@dns1004: START - running authdns-update * 19:12 cdobbins@cumin2002: conftool action : set/pooled=yes; selector: name=dns7002.*,service=authdns-update * 19:11 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1092.eqiad.wmnet with reason: host reimage * 19:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1123.eqiad.wmnet with reason: host reimage * 18:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1123.eqiad.wmnet with OS trixie * 18:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1109.eqiad.wmnet with OS trixie * 18:56 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1092.eqiad.wmnet with OS trixie * 18:52 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 18:49 swfrench@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 18:40 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 18:38 swfrench@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 18:11 swfrench-wmf: restarted eqsin, codfw confds - [[phab:T430909|T430909]] * 18:01 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T430909|T430909]] * 17:59 swfrench-wmf: restarted ulsfo confds, confirmed now connected to eqiad backends - [[phab:T430909|T430909]] * 17:52 sukhe: restart pybal on lvs2011 to switch from conf2004 to conf1008: [[phab:T430909|T430909]] * 17:51 sukhe: restart pybal on lvs2012 to switch from conf2004 to conf1008 [puppet re-enabled there]: [[phab:T430909|T430909]] * 17:46 sukhe: restart pybal on lvs2013 to switch from conf2004 to conf1008: [[phab:T430909|T430909]] * 17:44 swfrench-wmf: switched codfw, eqsin, ulsfo etcd client SRV records to eqiad - [[phab:T430909|T430909]] * 17:43 swfrench@dns1004: END - running authdns-update * 17:40 swfrench@dns1004: START - running authdns-update * 17:40 sukhe: restart pybal on lvs2014 to switch from conf2004 to conf1008: [[phab:T430909|T430909]] * 17:21 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1003.eqiad.wmnet * 17:15 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1003.eqiad.wmnet * 17:14 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1002.eqiad.wmnet * 17:06 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1002.eqiad.wmnet * 17:06 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-low-traffic-codfw' 'systemctl restart pybal.service' # lvs2013, [[phab:T416623|T416623]] * 17:04 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1001.eqiad.wmnet * 17:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1111.eqiad.wmnet with OS trixie * 17:00 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal.service' # lvs2014, [[phab:T416623|T416623]] * 16:58 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1001.eqiad.wmnet * 16:58 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS bookworm * 16:55 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-low-traffic-eqiad' 'systemctl restart pybal.service' # lvs1019, [[phab:T416623|T416623]] * 16:53 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-secondary-eqiad' 'systemctl restart pybal.service' # lvs1020, [[phab:T416623|T416623]] * 16:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1111.eqiad.wmnet with reason: host reimage * 16:40 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1111.eqiad.wmnet with reason: host reimage * 16:38 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.peering (exit_code=99) with action 'configure' for AS: 47794 * 16:35 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 47794 * 16:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1111.eqiad.wmnet with OS trixie * 16:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1006.eqiad.wmnet with OS trixie * 16:06 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 16:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1006.eqiad.wmnet with reason: host reimage * 15:58 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1121.eqiad.wmnet with OS trixie * 15:58 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 15:56 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1006.eqiad.wmnet with reason: host reimage * 15:54 mutante: jenkins down in planned maintenance window * 15:42 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1037.eqiad.wmnet * 15:42 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1037.eqiad.wmnet * 15:42 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1037.eqiad.wmnet * 15:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1006.eqiad.wmnet with OS trixie * 15:34 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1121.eqiad.wmnet with reason: host reimage * 15:33 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS bookworm * 15:33 cdobbins@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host dns7002.wikimedia.org with OS trixie * 15:30 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1121.eqiad.wmnet with reason: host reimage * 15:29 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host clouddumps1001.wikimedia.org * 15:20 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1001.wikimedia.org * 15:18 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1121 * 15:18 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1121 * 15:18 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host clouddumps1002.wikimedia.org * 15:17 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1121 * 15:17 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:17 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply * 15:16 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply * 15:16 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:16 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1121 - atsuko@cumin1003" * 15:16 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1121 - atsuko@cumin1003" * 15:14 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1037.eqiad.wmnet with OS trixie * 15:11 atsuko@cumin1003: START - Cookbook sre.dns.netbox * 15:09 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org * 15:09 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1121 * 15:09 andrew@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host clouddumps1002.wikimedia.org * 15:09 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org * 15:09 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1121.eqiad.wmnet with OS trixie * 15:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1007.eqiad.wmnet with OS trixie * 15:08 andrew@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host clouddumps1002.wikimedia.org * 15:08 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org * 15:05 brennen@deploy1003: Finished deploy [phabricator/deployment@7e02037]: deploy phab1004 for [[phab:T431440|T431440]] (duration: 00m 47s) * 15:04 brennen@deploy1003: Started deploy [phabricator/deployment@7e02037]: deploy phab1004 for [[phab:T431440|T431440]] * 15:03 brennen@deploy1003: Finished deploy [phabricator/deployment@7e02037]: deploy phab2003 for [[phab:T431440|T431440]] (duration: 00m 51s) * 15:03 brennen@deploy1003: Started deploy [phabricator/deployment@7e02037]: deploy phab2003 for [[phab:T431440|T431440]] * 15:00 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71] (thin): Regular analytics weekly train THIN [analytics/refinery@7d8dc71f] (duration: 02m 10s) * 14:58 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71] (thin): Regular analytics weekly train THIN [analytics/refinery@7d8dc71f] * 14:58 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71]: Regular analytics weekly train [analytics/refinery@7d8dc71f] (duration: 04m 14s) * 14:54 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1037.eqiad.wmnet with reason: host reimage * 14:53 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71]: Regular analytics weekly train [analytics/refinery@7d8dc71f] * 14:53 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@7d8dc71f] (duration: 02m 00s) * 14:51 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@7d8dc71f] * 14:51 arnaudb@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on phab2003.codfw.wmnet,phab[1004-1006].eqiad.wmnet with reason: maintenance * 14:51 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1037.eqiad.wmnet with reason: host reimage * 14:50 JavierMonton: Deploying Refinery at {{Gerrit|7d8dc71f}} for change {{Gerrit|1308087}} / [[phab:T431318|T431318]] - update filerevision table sqoop and table * 14:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1007.eqiad.wmnet with reason: host reimage * 14:42 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1007.eqiad.wmnet with reason: host reimage * 14:40 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1083.eqiad.wmnet with OS trixie * 14:37 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:36 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:35 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-master@eqiad * 14:35 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 14:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 14:34 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1037 * 14:34 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1037 * 14:34 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:cleanMentorList.php --wiki=frwiki # [[phab:T427386|T427386]] * 14:34 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 14:34 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308112{{!}}Revert^2 "[Growth] frwiki: Deploy automated mentor list cleaner" (T427386)]] (duration: 06m 47s) * 14:34 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 14:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:33 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:32 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1037 * 14:31 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:31 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:29 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:29 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-master@eqiad * 14:29 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:29 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:29 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:28 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:27 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:27 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1308112{{!}}Revert^2 "[Growth] frwiki: Deploy automated mentor list cleaner" (T427386)]] * 14:26 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1007.eqiad.wmnet with OS trixie * 14:26 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:26 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 14:26 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:25 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:25 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:25 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:cleanMentorList.php --wiki=frwiki # [[phab:T427386|T427386]] * 14:24 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:24 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1037 - blake@cumin1003" * 14:24 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1037 - blake@cumin1003" * 14:20 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1083.eqiad.wmnet with reason: host reimage * 14:19 blake@cumin1003: START - Cookbook sre.dns.netbox * 14:19 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-master@codfw * 14:19 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 14:19 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1037 * 14:18 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1037.eqiad.wmnet with OS trixie * 14:18 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1037.eqiad.wmnet * 14:18 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 14:18 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1037.eqiad.wmnet * 14:18 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1037.eqiad.wmnet * 14:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2007.codfw.wmnet with OS trixie * 14:16 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1083.eqiad.wmnet with reason: host reimage * 14:15 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1036.eqiad.wmnet * 14:15 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1036.eqiad.wmnet * 14:14 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1036.eqiad.wmnet * 14:12 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-master@codfw * 14:11 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1120.eqiad.wmnet with OS trixie * 14:05 moritzm: installing distro-info-data updates from trixie/bookworm point releases * 14:04 fabfur: disable puppet on A:cp-text to selectively apply https://gerrit.wikimedia.org/r/c/operations/puppet/+/1308040 * 14:03 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] (duration: 27m 48s) * 14:00 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 14:00 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1083.eqiad.wmnet with OS trixie * 13:58 urbanecm@deploy1003: urbanecm: Continuing with deployment * 13:58 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:58 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1004.eqiad.wmnet with OS bookworm * 13:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2007.codfw.wmnet with reason: host reimage * 13:57 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 13:53 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1120.eqiad.wmnet with reason: host reimage * 13:50 moritzm: installing Linux 5.10.259 on Bullseye hosts * 13:47 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply * 13:47 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply * 13:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2007.codfw.wmnet with reason: host reimage * 13:46 cgoubert@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/aux-k8s-services/redioscope: apply * 13:46 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1120.eqiad.wmnet with reason: host reimage * 13:46 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:46 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:45 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:44 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:44 cgoubert@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/aux-k8s-services/redioscope: apply * 13:40 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 13:39 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:38 moritzm: installing e2fsprogs updates from Trixie point release * 13:35 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] * 13:33 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1120.eqiad.wmnet with OS trixie * 13:33 topranks: reset cr3-eqsin configuration so traffic uses it again after upgrade * 13:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1088.eqiad.wmnet with OS trixie * 13:32 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie * 13:32 cdobbins@cumin1003: conftool action : set/pooled=no; selector: name=dns7002.* * 13:29 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2007.codfw.wmnet with OS trixie * 13:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1004.eqiad.wmnet with reason: host reimage * 13:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2006.codfw.wmnet with OS trixie * 13:18 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1036.eqiad.wmnet with OS trixie * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aux-k8s-etcd1004.eqiad.wmnet with reason: host reimage * 13:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 13:16 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 13:15 jayme: Istio is being upgraded from 1.24.2 to 1.29.4 on wikikube staging eqiad and codfw - [[phab:T427401|T427401]] * 13:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1087.eqiad.wmnet with OS trixie * 13:14 topranks: reboot cr3-eqsin to install new JunOS and set PIC 0/0/0 to 100G * 13:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1088.eqiad.wmnet with reason: host reimage * 13:13 jmm@dns1004: END - running authdns-update * 13:12 jmm@dns1004: START - running authdns-update * 13:09 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1088.eqiad.wmnet with reason: host reimage * 13:07 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1082.eqiad.wmnet with OS trixie * 13:07 jmm@dns1004: END - running authdns-update * 13:06 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1004.eqiad.wmnet with OS bookworm * 13:05 jmm@dns1004: START - running authdns-update * 13:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2006.codfw.wmnet with reason: host reimage * 12:58 topranks: load updated JunOS on cr3-eqsin [[phab:T429386|T429386]] * 12:58 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1036.eqiad.wmnet with reason: host reimage * 12:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2001.codfw.wmnet * 12:57 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2006.codfw.wmnet with reason: host reimage * 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr1-codfw,cr[2-3]-eqsin,cr3-eqsin IPv6,cr3-eqsin.mgmt with reason: upgrade JunOS cr3-eqsin * 12:56 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lvs[5004-5006].eqsin.wmnet with reason: upgrade JunOS cr3-eqsin * 12:55 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 12:55 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 12:53 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1087.eqiad.wmnet with reason: host reimage * 12:53 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 12:52 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1088.eqiad.wmnet with OS trixie * 12:52 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1002.eqiad.wmnet * 12:52 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:51 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2001.codfw.wmnet * 12:49 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1036.eqiad.wmnet with reason: host reimage * 12:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 12:48 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1087.eqiad.wmnet with reason: host reimage * 12:44 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1082.eqiad.wmnet with reason: host reimage * 12:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1002.eqiad.wmnet * 12:42 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 12:42 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 12:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 12:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 12:39 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:39 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: move dumps-nfs IP to the shared one - filippo@cumin1003" * 12:39 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: move dumps-nfs IP to the shared one - filippo@cumin1003" * 12:39 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2006.codfw.wmnet with OS trixie * 12:38 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1082.eqiad.wmnet with reason: host reimage * 12:36 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:33 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:32 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1036 * 12:32 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1036 * 12:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2005.codfw.wmnet with OS trixie * 12:32 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1087.eqiad.wmnet with OS trixie * 12:30 jmm@dns1004: END - running authdns-update * 12:29 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1036 * 12:29 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1036.eqiad.wmnet 21.32.64.10.in-addr.arpa 1.2.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:29 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1036.eqiad.wmnet 21.32.64.10.in-addr.arpa 1.2.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:29 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:29 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1036 - blake@cumin1003" * 12:29 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1036 - blake@cumin1003" * 12:28 jmm@dns1004: START - running authdns-update * 12:26 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:26 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:23 blake@cumin1003: START - Cookbook sre.dns.netbox * 12:23 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1036 * 12:23 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1036.eqiad.wmnet with OS trixie * 12:22 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1036.eqiad.wmnet * 12:22 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1082.eqiad.wmnet with OS trixie * 12:22 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1036.eqiad.wmnet * 12:22 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1036.eqiad.wmnet * 12:21 marostegui: Restart mariadb@s7 on db1155 to pick up new filters - [[phab:T431124|T431124]] * 12:21 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 21 hosts with reason: restarting for replication filter * 12:20 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:19 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:14 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2005.codfw.wmnet with reason: host reimage * 12:14 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:08 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:07 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:07 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2005.codfw.wmnet with reason: host reimage * 12:06 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:06 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:06 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:05 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:05 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:04 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-master-eqiad * 12:04 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl1002.eqiad.wmnet * 12:04 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl1002.eqiad.wmnet * 12:04 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:04 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:03 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:03 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:03 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:03 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 11:59 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl1002.eqiad.wmnet * 11:59 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl1002.eqiad.wmnet * 11:59 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl1001.eqiad.wmnet * 11:59 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl1001.eqiad.wmnet * 11:56 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl1001.eqiad.wmnet * 11:56 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl1001.eqiad.wmnet * 11:56 klausman@cumin2002: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-master-eqiad * 11:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2005.codfw.wmnet with OS trixie * 11:49 blake@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on wikikube-worker1160.eqiad.wmnet with reason: Verifying matchers for silence * 11:42 topranks: cr3-eqsin, begin traffic drain to reset PIC and upgrade JunOS [[phab:T429386|T429386]] * 11:41 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs[5004-5006].eqsin.wmnet with reason: upgrade JunOS cr3-eqsin * 11:39 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr1-codfw,cr[2-3]-eqsin,cr3-eqsin IPv6,cr3-eqsin.mgmt with reason: upgrade JunOS cr3-eqsin * 11:36 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=thanos-fe2004.codfw.wmnet * 11:35 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1086.eqiad.wmnet with OS trixie * 11:35 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=thanos-fe2004.codfw.wmnet * 11:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1085.eqiad.wmnet with OS trixie * 11:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1086.eqiad.wmnet with reason: host reimage * 11:10 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1085.eqiad.wmnet with reason: host reimage * 11:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2004.codfw.wmnet with OS trixie * 11:03 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1086.eqiad.wmnet with reason: host reimage * 11:02 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1085.eqiad.wmnet with reason: host reimage * 10:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2004.codfw.wmnet with reason: host reimage * 10:48 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:46 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1086.eqiad.wmnet with OS trixie * 10:46 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1085.eqiad.wmnet with OS trixie * 10:44 cgoubert@deploy1003: Finished deploy [restbase/deploy@2fc37d4]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] (duration: 16m 44s) * 10:43 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2004.codfw.wmnet with reason: host reimage * 10:35 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:27 cgoubert@deploy1003: Started deploy [restbase/deploy@2fc37d4]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] * 10:27 cgoubert@deploy1003: Finished deploy [restbase/deploy@8a25036]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] (duration: 00m 45s) * 10:26 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1117.eqiad.wmnet with OS trixie * 10:26 cgoubert@deploy1003: Started deploy [restbase/deploy@8a25036]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] * 10:26 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host thanos-fe2004 * 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host thanos-fe2004 * 10:22 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1116.eqiad.wmnet with OS trixie * 10:21 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host thanos-fe2004 * 10:21 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) thanos-fe2004.codfw.wmnet 157.32.192.10.in-addr.arpa 7.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:20 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache thanos-fe2004.codfw.wmnet 157.32.192.10.in-addr.arpa 7.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:20 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:20 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host thanos-fe2004 - mvernon@cumin2003" * 10:20 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host thanos-fe2004 - mvernon@cumin2003" * 10:15 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2252: Repooling after reboot * 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:15 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 10:15 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2252: Repooling after reboot * 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1153.eqiad.wmnet * 10:14 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1153.eqiad.wmnet * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 10:14 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 10:12 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 10:12 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host thanos-fe2004 * 10:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2004.codfw.wmnet with OS trixie * 10:07 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1117.eqiad.wmnet with reason: host reimage * 10:03 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1116.eqiad.wmnet with reason: host reimage * 09:58 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:58 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1117.eqiad.wmnet with reason: host reimage * 09:57 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1116.eqiad.wmnet with reason: host reimage * 09:49 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 41 days, 15:00:00 on db2252.codfw.wmnet with reason: Security updates * 09:45 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1117.eqiad.wmnet with OS trixie * 09:45 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1116.eqiad.wmnet with OS trixie * 09:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1153: Security updates * 09:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:28 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:28 root@cumin1003: START - Cookbook sre.mysql.depool depool db1153: Security updates * 09:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1016: Security updates * 09:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:21 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:21 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1016: Security updates * 09:14 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:14 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 08:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1016: Security updates * 08:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:56 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:56 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1016: Security updates * 08:50 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:50 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:45 filippo@dns1006: END - running authdns-update * 08:43 filippo@dns1006: START - running authdns-update * 08:42 godog: switch dumps-nfs address to be shared with rsync/http - [[phab:T411248|T411248]] * 08:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1016: Security updates * 08:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:40 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:40 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1016: Security updates * 08:29 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host cirrussearch1111.eqiad.wmnet * 08:29 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:27 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:27 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:25 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1015: Security updates * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:09 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:09 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1015: Security updates * 07:42 Msz2001: Deployed private patch for Suggested Ivestigations * 07:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1015: Security updates * 07:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:41 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:41 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1015: Security updates * 07:40 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 07:11 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fingerprint warnings - oblivian@cumin1003" * 07:11 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fingerprint warnings - oblivian@cumin1003 * 07:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1024: Security updates * 07:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:11 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:11 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1024: Security updates * 07:10 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fingerprint warnings - oblivian@cumin1003 * 07:10 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fingerprint warnings - oblivian@cumin1003" * 06:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host cirrussearch1111.eqiad.wmnet * 06:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 06:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1024: Security updates * 06:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 06:48 root@cumin1003: START - Cookbook sre.mysql.parsercache * 06:48 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1024: Security updates * 06:42 moritzm: install nginx security updates * 06:31 root@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool pc1024: Security updates * 06:21 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1024: Security updates * 06:19 moritzm: installing php8.2 security updates * 06:15 moritzm: installing php8.4 security updates * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.7 (duration: 02m 38s) * 03:40 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] (duration: 37m 04s) * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 51s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-06 == * 23:30 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] (duration: 09m 39s) * 23:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1078.eqiad.wmnet with OS trixie * 23:26 jdlrobson@deploy1003: jdlrobson, bwang: Continuing with deployment * 23:22 jdlrobson@deploy1003: jdlrobson, bwang: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug) * 23:21 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] * 23:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1078.eqiad.wmnet with reason: host reimage * 23:06 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1078.eqiad.wmnet with reason: host reimage * 22:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1078.eqiad.wmnet with OS trixie * 22:29 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on cirrussearch1114.eqiad.wmnet with reason: reimage on hold until restore completes * 22:22 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on cirrussearch[1079,1115].eqiad.wmnet with reason: reimage on hold until restore completes * 21:18 maryum: Deployed security fix for [[phab:T428006|T428006]] * 20:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1079.eqiad.wmnet with OS trixie * 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1077.eqiad.wmnet with OS trixie * 20:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1115.eqiad.wmnet with OS trixie * 20:25 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1079.eqiad.wmnet with reason: host reimage * 20:21 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1079.eqiad.wmnet with reason: host reimage * 20:15 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] (duration: 08m 14s) * 20:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1077.eqiad.wmnet with reason: host reimage * 20:10 krinkle@deploy1003: krinkle, pushpaktiwari: Continuing with deployment * 20:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1115.eqiad.wmnet with reason: host reimage * 20:08 krinkle@deploy1003: krinkle, pushpaktiwari: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1077.eqiad.wmnet with reason: host reimage * 20:06 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] * 20:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1079.eqiad.wmnet with OS trixie * 20:04 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1115.eqiad.wmnet with reason: host reimage * 19:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1077.eqiad.wmnet with OS trixie * 19:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1115.eqiad.wmnet with OS trixie * 19:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 19:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 18:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1114.eqiad.wmnet with OS trixie * 18:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1114.eqiad.wmnet with reason: host reimage * 18:35 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1114.eqiad.wmnet with reason: host reimage * 18:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1112.eqiad.wmnet with OS trixie * 18:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1114.eqiad.wmnet with OS trixie * 18:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1072.eqiad.wmnet with OS trixie * 18:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1112.eqiad.wmnet with reason: host reimage * 18:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1112.eqiad.wmnet with reason: host reimage * 17:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1072.eqiad.wmnet with reason: host reimage * 17:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1112.eqiad.wmnet with OS trixie * 17:55 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1072.eqiad.wmnet with reason: host reimage * 17:39 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1072.eqiad.wmnet with OS trixie * 17:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1071.eqiad.wmnet with OS trixie * 17:18 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1070.eqiad.wmnet with OS trixie * 17:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1084.eqiad.wmnet with OS trixie * 16:54 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1071.eqiad.wmnet with reason: host reimage * 16:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1084.eqiad.wmnet with reason: host reimage * 16:51 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1070.eqiad.wmnet with reason: host reimage * 16:49 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1084.eqiad.wmnet with reason: host reimage * 16:38 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1071.eqiad.wmnet with OS trixie * 16:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1096.eqiad.wmnet with OS trixie * 16:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1070.eqiad.wmnet with OS trixie * 16:33 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1084.eqiad.wmnet with OS trixie * 16:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1089.eqiad.wmnet with OS trixie * 16:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1103.eqiad.wmnet with OS trixie * 16:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1096.eqiad.wmnet with reason: host reimage * 16:14 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1096.eqiad.wmnet with reason: host reimage * 16:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1089.eqiad.wmnet with reason: host reimage * 16:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1103.eqiad.wmnet with reason: host reimage * 16:02 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1003.eqiad.wmnet with OS bookworm * 16:01 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1089.eqiad.wmnet with reason: host reimage * 16:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1103.eqiad.wmnet with reason: host reimage * 15:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1096.eqiad.wmnet with OS trixie * 15:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1080.eqiad.wmnet with OS trixie * 15:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1089.eqiad.wmnet with OS trixie * 15:45 dancy@deploy1003: Installation of scap version "4.272.0" completed for 158 hosts * 15:43 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1103.eqiad.wmnet with OS trixie * 15:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1113.eqiad.wmnet with OS trixie * 15:41 dancy@deploy1003: Installing scap version "4.272.0" for 158 host(s) * 15:40 klausman@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 15:39 klausman@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 15:38 klausman@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 15:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1069.eqiad.wmnet with OS trixie * 15:37 klausman@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 15:36 klausman@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 15:34 klausman@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 15:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1080.eqiad.wmnet with reason: host reimage * 15:27 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1080.eqiad.wmnet with reason: host reimage * 15:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1113.eqiad.wmnet with reason: host reimage * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1069.eqiad.wmnet with reason: host reimage * 15:18 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1113.eqiad.wmnet with reason: host reimage * 15:16 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1069.eqiad.wmnet with reason: host reimage * 15:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:11 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1080.eqiad.wmnet with OS trixie * 15:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1113.eqiad.wmnet with OS trixie * 15:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1003.eqiad.wmnet with reason: host reimage * 14:47 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1003.eqiad.wmnet with OS bookworm * 14:33 elukey: rolled out spicerack on all cumin nodes - [[phab:T429699|T429699]] * 14:32 elukey: upgrade all bookworm hosts to pywmflib 3.1 - [[phab:T430552|T430552]] * 14:14 marostegui: Setup x4 eqiad topology [[phab:T404715|T404715]] * 14:13 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 14:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2230.codfw.wmnet * 14:07 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2230.codfw.wmnet * 13:59 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[2001-2002].codfw.wmnet * 13:51 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 13:45 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.major-upgrade (exit_code=97) * 13:45 cwilliams@cumin1003: dbmaint on s4@codfw [[phab:T429893|T429893]] * 13:45 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 13:42 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-master-codfw * 13:42 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl2002.codfw.wmnet * 13:42 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl2002.codfw.wmnet * 13:38 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl2002.codfw.wmnet * 13:38 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl2002.codfw.wmnet * 13:38 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl2001.codfw.wmnet * 13:38 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl2001.codfw.wmnet * 13:35 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl2001.codfw.wmnet * 13:35 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl2001.codfw.wmnet * 13:35 klausman@cumin2002: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-master-codfw * 12:30 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] (duration: 25m 11s) * 12:24 krinkle@deploy1003: krinkle: Continuing with deployment * 12:10 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2048.codfw.wmnet * 12:09 krinkle@deploy1003: krinkle: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:08 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2048.codfw.wmnet * 12:05 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] * 11:57 moritzm: installing curl security updates * 11:49 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:31 moritzm: installing nano security updates * 11:07 moritzm: failover Ganeti master in codfw to ganeti2032 [[phab:T430909|T430909]] * 11:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:04 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest1005.eqiad.wmnet with OS trixie * 11:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:50 jmm@dns1004: END - running authdns-update * 10:47 jmm@dns1004: START - running authdns-update * 10:47 jmm@dns1004: START - running authdns-update * 10:46 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:44 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest1005.eqiad.wmnet with reason: host reimage * 10:38 elukey@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest1005.eqiad.wmnet with reason: host reimage * 10:31 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:31 marostegui: Setup x4 codfw topology [[phab:T404715|T404715]] * 10:31 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 10:24 elukey: spicerack 13.0.0 deployed on cumin2002 * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 10:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 10:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 10:21 elukey@cumin2002: START - Cookbook sre.hosts.reimage for host sretest1005.eqiad.wmnet with OS trixie * 10:20 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:19 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:17 elukey: uploaded spicerack_13.0.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia * 09:54 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:52 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:20 elukey: upgrade all bullseye hosts to pywmflib 3.1 - [[phab:T430552|T430552]] * 09:10 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1015.eqiad.wmnet,service=s4 * 09:10 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1015.eqiad.wmnet,service=s6 * 09:07 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 08:58 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:56 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 08:56 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 08:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 08:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 08:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin2002.codfw.wmnet * 08:06 godog: remove cloudvirt1046, cloudvirt1062, cloudvirt1074, cloudvirt1075 from maintenance aggregate and put them in network-ovs - [[phab:T424802|T424802]] * 08:00 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin2002.codfw.wmnet * 07:58 hashar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] (duration: 32m 53s) * 07:57 fabfur: repooled cp4038 * 07:57 fabfur@cumin1003: conftool action : set/pooled=yes; selector: name=cp4038.* * 07:53 moritzm: installing pyjwt security updates * 07:47 moritzm: installing openjpeg2 security updates * 07:45 hashar@deploy1003: vadymts1, hashar: Continuing with deployment * 07:43 hashar@deploy1003: vadymts1, hashar: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:38 moritzm: installing python-urllib3 security updates * 07:37 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 07:30 fabfur: depooled cp4038 to investigate on possible maxmind failure * 07:30 fabfur@cumin1003: conftool action : set/pooled=no; selector: name=cp4038.* * 07:30 fabfur@cumin1003: conftool action : set/pooled=yes; selector: name=cp4038.* * 07:29 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 07:25 hashar@deploy1003: Started scap sync-world: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] * 06:13 moritzm: installing Linux 6.12.95 on trixie hosts * 05:20 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s6 * 05:20 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s4 * 05:19 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1015.eqiad.wmnet with reason: cloning * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 08s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-05 == * 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 01m 08s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-04 == * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 58s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-03 == * 17:08 topranks: revert protocol preference changes on cr3-ulsfo after upgrade * 16:53 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on cr2-eqord with reason: upgrade JunOS cr3-ulsfo * 16:53 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on cr4-ulsfo with reason: upgrade JunOS cr3-ulsfo * 16:48 topranks: reboot cr3-ulsfo to upgrade JunOS and reset linecard [[phab:T424839|T424839]] * 15:52 topranks: adjust outbound BGP policies on cr3-ulsfo to drain router of traffic [[phab:T424839|T424839]] * 15:45 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on lvs[4008-4010].ulsfo.wmnet with reason: upgrade JunOS cr3-ulsfo * 15:44 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on asw1-[22-23]-ulsfo,cr3-ulsfo,cr3-ulsfo IPv6,cr3-ulsfo.mgmt with reason: upgrade JunOS cr3-ulsfo * 15:36 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 15:35 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 15:35 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 14:40 cmooney@dns3003: END - running authdns-update * 14:26 cmooney@dns3003: START - running authdns-update * 14:26 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:26 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to ulsfo - cmooney@cumin1003" * 14:19 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to ulsfo - cmooney@cumin1003" * 14:16 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:38 sukhe@dns1004: END - running authdns-update * 13:35 sukhe@dns1004: START - running authdns-update * 13:26 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 13:26 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 13:26 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet * 13:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 13:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 13:16 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host sretest1005.eqiad.wmnet * 13:16 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 13:16 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 13:15 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 13:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:14 moritzm: imported samplicator 1.3.8rc1-1+deb13u1 to trixie-wikimedia/main [[phab:T337208|T337208]] * 13:13 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:07 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 13:07 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 13:02 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:02 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:58 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:57 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:57 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:53 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet * 12:50 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 12:47 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:41 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:40 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:39 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:32 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet * 12:26 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet * 12:23 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2005.wikimedia.org * 12:19 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2005.wikimedia.org * 12:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 12:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup[2004-2007].codfw.wmnet * 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[2004-2007].codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin2003" * 12:15 jynus@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[2004-2007].codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin2003" * 12:09 jynus@cumin2003: START - Cookbook sre.dns.netbox * 11:58 jynus@cumin2003: START - Cookbook sre.hosts.decommission for hosts backup[2004-2007].codfw.wmnet * 10:40 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup[1004-1007].eqiad.wmnet * 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[1004-1007].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 10:01 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[1004-1007].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 09:52 jynus@cumin1003: START - Cookbook sre.dns.netbox * 09:39 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:36 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup[1004-1007].eqiad.wmnet * 09:36 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:25 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 09:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 09:16 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 09:05 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:04 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:00 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:59 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:57 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:55 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:50 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 08:50 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 08:49 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 08:49 atsukoito: depooling cirrussearch in codfw because of regression after upgrade [[phab:T431091|T431091]] * 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts mirror1001.wikimedia.org * 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: mirror1001.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 08:29 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: mirror1001.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 08:18 jmm@cumin2003: START - Cookbook sre.dns.netbox * 08:11 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts mirror1001.wikimedia.org * 06:15 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 18s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-02 == * 22:55 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host contint1003.wikimedia.org with OS trixie * 22:29 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on contint1003.wikimedia.org with reason: host reimage * 22:23 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on contint1003.wikimedia.org with reason: host reimage * 22:05 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host contint1003.wikimedia.org with OS trixie * 22:03 mutante: contint1003 (zuul.wikimedia.org) - reimaging because of [[phab:T430510|T430510]]#12067628 [[phab:T418521|T418521]] * 22:03 dzahn@cumin2002: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on zuul.wikimedia.org with reason: reimage * 21:39 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 18s) * 21:39 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 21:20 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1003.eqiad.wmnet, repooling source-only afterwards * 21:19 sbassett: Deployed security fix for [[phab:T428829|T428829]] * 20:58 cmooney@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Release v0.11.2 update for new Aerleon - cmooney@cumin1003 * 20:55 cmooney@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Release v0.11.2 update for new Aerleon - cmooney@cumin1003 * 20:40 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] (duration: 12m 35s) * 20:36 arlolra@deploy1003: cscott, arlolra: Continuing with deployment * 20:35 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 20s) * 20:35 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 20:33 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host contint2003.wikimedia.org with OS trixie * 20:31 arlolra@deploy1003: cscott, arlolra: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Cha * 20:28 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] * 20:17 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] (duration: 08m 13s) * 20:14 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on contint2003.wikimedia.org with reason: host reimage * 20:13 sbassett@deploy1003: sbassett: Continuing with deployment * 20:11 sbassett@deploy1003: sbassett: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:09 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] * 20:08 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 20:08 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 20:08 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on contint2003.wikimedia.org with reason: host reimage * 20:05 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1003.eqiad.wmnet, repooling source-only afterwards * 19:49 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host contint2003.wikimedia.org with OS trixie * 19:48 mutante: contint2003 - reimaging because of [[phab:T430510|T430510]]#12067628 [[phab:T418521|T418521]] * 18:39 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 18:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 18:13 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2002.codfw.wmnet -> wcqs2003.codfw.wmnet, repooling source-only afterwards * 17:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1003.eqiad.wmnet with OS bookworm * 17:52 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1005.eqiad.wmnet * 17:52 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1005.eqiad.wmnet * 17:51 jasmine@cumin2002: conftool action : set/pooled=yes:weight=10; selector: name=wikikube-ctrl1005.eqiad.wmnet * 17:48 jasmine_: homer "cr*eqiad*" commit "Added new stacked control plane wikikube-ctrl1005" * 17:44 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply * 17:44 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply * 17:31 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] (duration: 09m 33s) * 17:26 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 17:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1003.eqiad.wmnet with reason: host reimage * 17:23 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:21 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] * 17:18 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1003.eqiad.wmnet with reason: host reimage * 17:16 rscout@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply * 17:16 rscout@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply * 17:16 rscout@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply * 17:15 rscout@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply * 17:12 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on wcqs[2002-2003].codfw.wmnet,wcqs1002.eqiad.wmnet with reason: reimaging hosts * 17:08 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 17:08 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 17:08 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 17:07 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 17:05 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 17:05 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 17:03 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "running to make sure all updates are synced - cmooney@cumin1003" * 17:03 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "running to make sure all updates are synced - cmooney@cumin1003" * 17:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs1003 * 17:00 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs1003 * 17:00 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 17:00 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1003.eqiad.wmnet with OS bookworm * 16:58 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Re-running - btullis@cumin1003" * 16:58 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Re-running - btullis@cumin1003" * 16:58 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2002.codfw.wmnet -> wcqs2003.codfw.wmnet, repooling source-only afterwards * 16:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-master1004.eqiad.wmnet with OS bookworm * 16:58 btullis@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 16:57 tappof: bump space for prometheus k8s-aux in eqiad * 16:55 cmooney@dns3003: END - running authdns-update * 16:55 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:55 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to eqsin - cmooney@cumin1003" * 16:55 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to eqsin - cmooney@cumin1003" * 16:53 cmooney@dns3003: START - running authdns-update * 16:52 ryankemper: [ml-serve-eqiad] Cleared out 1302 failed (Evicted) pods: `kubectl -n llm delete pods --field-selector=status.phase=Failed`, freeing calico-kube-controllers from OOM crashloop (evictions were caused by disk pressure) * 16:49 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 16:46 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:39 rzl@dns1004: END - running authdns-update * 16:37 rzl@dns1004: START - running authdns-update * 16:36 rzl@dns1004: START - running authdns-update * 16:35 rzl@deploy1003: Finished scap sync-world: [[phab:T416623|T416623]] (duration: 10m 19s) * 16:34 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 16:33 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-master1004.eqiad.wmnet with reason: host reimage * 16:30 rzl@deploy1003: rzl: Continuing with deployment * 16:28 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-master1004.eqiad.wmnet with reason: host reimage * 16:26 rzl@deploy1003: rzl: [[phab:T416623|T416623]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:25 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 16:25 rzl@deploy1003: Started scap sync-world: [[phab:T416623|T416623]] * 16:25 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 16:24 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 16:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: sync * 16:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: sync * 16:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-master1004.eqiad.wmnet with OS bookworm * 16:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-master1003.eqiad.wmnet with OS bookworm * 16:11 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 16:11 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 16:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Security updates * 16:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 16:08 root@cumin1003: START - Cookbook sre.mysql.parsercache * 16:08 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Security updates * 15:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-master1003.eqiad.wmnet with reason: host reimage * 15:54 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:54 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:54 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:54 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-master1003.eqiad.wmnet with reason: host reimage * 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Security updates * 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:45 root@cumin1003: START - Cookbook sre.mysql.parsercache * 15:45 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Security updates * 15:42 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-master1003.eqiad.wmnet with OS bookworm * 15:24 moritzm: installing busybox updates from bookworm point release * 15:20 moritzm: installing busybox updates from trixie point release * 15:15 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1021: Security updates * 15:15 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:15 root@cumin1003: START - Cookbook sre.mysql.parsercache * 15:15 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1021: Security updates * 15:13 moritzm: installing giflib security updates * 15:08 moritzm: installing Tomcat security updates * 14:57 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 14:56 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 14:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:53 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Unblock taavi - oblivian@cumin1003" * 14:53 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Unblock taavi - oblivian@cumin1003 * 14:53 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1021: Security updates * 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:53 root@cumin1003: START - Cookbook sre.mysql.parsercache * 14:53 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1021: Security updates * 14:53 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Unblock taavi - oblivian@cumin1003 * 14:52 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Unblock taavi - oblivian@cumin1003" * 14:46 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94711 and previous config saved to /var/cache/conftool/dbconfig/20260702-144644-fceratto.json * 14:36 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205', diff saved to https://phabricator.wikimedia.org/P94709 and previous config saved to /var/cache/conftool/dbconfig/20260702-143636-fceratto.json * 14:32 moritzm: installing libdbi-perl security updates * 14:26 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205', diff saved to https://phabricator.wikimedia.org/P94708 and previous config saved to /var/cache/conftool/dbconfig/20260702-142628-fceratto.json * 14:16 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94707 and previous config saved to /var/cache/conftool/dbconfig/20260702-141621-fceratto.json * 14:12 moritzm: installing rsync security updates * 14:11 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox) * 14:10 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94706 and previous config saved to /var/cache/conftool/dbconfig/20260702-140959-fceratto.json * 14:09 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2205.codfw.wmnet with reason: Maintenance * 14:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2205: Repooling after switchover * 14:07 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-test-master1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 14:06 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 14:06 Tran: Deployed patch for [[phab:T427287|T427287]] * 14:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:59 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2205: Repooling after switchover * 13:59 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2205: Repooling after switchover * 13:59 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:55 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2205: Repooling after switchover * 13:55 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2205 [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94704 and previous config saved to /var/cache/conftool/dbconfig/20260702-135505-fceratto.json * 13:54 moritzm: installing sed security updates * 13:53 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:52 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2209 to s3 primary [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94703 and previous config saved to /var/cache/conftool/dbconfig/20260702-135235-fceratto.json * 13:52 federico3: Starting s3 codfw failover from db2205 to db2209 - [[phab:T430912|T430912]] * 13:51 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:51 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 13:48 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:47 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2209 with weight 0 [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94702 and previous config saved to /var/cache/conftool/dbconfig/20260702-134719-fceratto.json * 13:47 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Primary switchover s3 [[phab:T430912|T430912]] * 13:44 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:44 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:44 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:40 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 13:38 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 13:37 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 13:36 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 13:36 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:34 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 13:30 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:29 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:29 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:27 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:26 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:25 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling restart_daemons on A:wikidough * 13:23 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 13:22 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 13:17 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 13:17 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns1004.wikimedia.org * 13:12 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:11 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart (exit_code=97) rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough * 13:11 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=97) rolling restart_daemons on A:wikidough * 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough * 13:09 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] (duration: 07m 20s) * 13:05 aude@deploy1003: jdrewniak, aude: Continuing with deployment * 13:04 aude@deploy1003: jdrewniak, aude: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:02 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] * 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts wdqs-categories1001.eqiad.wmnet * 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: wdqs-categories1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 12:10 jmm@dns1004: END - running authdns-update * 12:07 jmm@dns1004: START - running authdns-update * 11:51 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: wdqs-categories1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 11:44 btullis@cumin1003: START - Cookbook sre.dns.netbox * 11:42 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 11:42 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 11:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet * 11:39 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts wdqs-categories1001.eqiad.wmnet * 11:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet * 11:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet * 11:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet * 11:29 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 11:29 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 10:57 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2214: Repooling * 10:49 jmm@dns1004: END - running authdns-update * 10:47 jmm@dns1004: START - running authdns-update * 10:31 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94698 and previous config saved to /var/cache/conftool/dbconfig/20260702-103146-fceratto.json * 10:21 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213', diff saved to https://phabricator.wikimedia.org/P94696 and previous config saved to /var/cache/conftool/dbconfig/20260702-102137-fceratto.json * 10:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:19 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb1017.eqiad.wmnet * 10:18 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 10:18 fceratto@cumin1003: Removing es1033 from zarcillo [[phab:T408772|T408772]] * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts es1033.eqiad.wmnet * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: es1033.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:14 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: es1033.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:13 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb1017.eqiad.wmnet * 10:12 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2214.codfw.wmnet * 10:12 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2214.codfw.wmnet * 10:12 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2214: Repooling * 10:11 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213', diff saved to https://phabricator.wikimedia.org/P94693 and previous config saved to /var/cache/conftool/dbconfig/20260702-101130-fceratto.json * 10:10 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:10 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:03 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts es1033.eqiad.wmnet * 10:03 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 10:01 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94691 and previous config saved to /var/cache/conftool/dbconfig/20260702-100122-fceratto.json * 09:55 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94690 and previous config saved to /var/cache/conftool/dbconfig/20260702-095529-fceratto.json * 09:55 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2213.codfw.wmnet with reason: Maintenance * 09:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 09:53 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2213: Repooling after switchover * 09:51 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover * 09:44 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2213: Repooling after switchover * 09:39 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover * 09:39 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2213 [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94688 and previous config saved to /var/cache/conftool/dbconfig/20260702-093859-fceratto.json * 09:36 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2192 to s5 primary [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94687 and previous config saved to /var/cache/conftool/dbconfig/20260702-093650-fceratto.json * 09:36 federico3: Starting s5 codfw failover from db2213 to db2192 - [[phab:T430923|T430923]] * 09:30 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94686 and previous config saved to /var/cache/conftool/dbconfig/20260702-093004-fceratto.json * 09:24 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2192 with weight 0 [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94685 and previous config saved to /var/cache/conftool/dbconfig/20260702-092455-fceratto.json * 09:24 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 23 hosts with reason: Primary switchover s5 [[phab:T430923|T430923]] * 09:19 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220', diff saved to https://phabricator.wikimedia.org/P94684 and previous config saved to /var/cache/conftool/dbconfig/20260702-091957-fceratto.json * 09:16 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] (duration: 06m 57s) * 09:13 moritzm: installing libgcrypt20 security updates * 09:12 kharlan@deploy1003: kharlan: Continuing with deployment * 09:11 kharlan@deploy1003: kharlan: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:09 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220', diff saved to https://phabricator.wikimedia.org/P94683 and previous config saved to /var/cache/conftool/dbconfig/20260702-090950-fceratto.json * 09:09 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] * 09:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 09:01 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] (duration: 07m 07s) * 08:59 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94682 and previous config saved to /var/cache/conftool/dbconfig/20260702-085942-fceratto.json * 08:57 kharlan@deploy1003: kharlan: Continuing with deployment * 08:56 kharlan@deploy1003: kharlan: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:54 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] * 08:52 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:52 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:52 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94681 and previous config saved to /var/cache/conftool/dbconfig/20260702-085237-fceratto.json * 08:52 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2220.codfw.wmnet with reason: Maintenance * 08:43 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:40 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 08:25 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] (duration: 11m 44s) * 08:21 cscott@deploy1003: cscott: Continuing with deployment * 08:16 cscott@deploy1003: cscott: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:14 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] * 08:08 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 08:08 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1244: Migration of db1244.eqiad.wmnet completed * 08:02 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:02 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:01 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] (duration: 18m 58s) * 08:01 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:59 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 07:59 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:59 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:59 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2006.wikimedia.org * 07:58 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:57 cscott@deploy1003: cscott: Continuing with deployment * 07:56 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:56 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:56 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:55 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:55 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:55 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:54 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2006.wikimedia.org * 07:54 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:54 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 07:44 cscott@deploy1003: cscott: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:44 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2005.wikimedia.org * 07:44 moritzm: installing node-lodash security updates * 07:42 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] * 07:39 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2005.wikimedia.org * 07:30 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] (duration: 07m 28s) * 07:26 cscott@deploy1003: ssastry, cscott: Continuing with deployment * 07:25 cscott@deploy1003: ssastry, cscott: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:23 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1244: Migration of db1244.eqiad.wmnet completed * 07:22 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] * 07:16 wmde-fisch@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] (duration: 06m 55s) * 07:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1244.eqiad.wmnet with OS trixie * 07:11 wmde-fisch@deploy1003: wmde-fisch: Continuing with deployment * 07:11 wmde-fisch@deploy1003: wmde-fisch: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:09 wmde-fisch@deploy1003: Started scap sync-world: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] * 06:54 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1244.eqiad.wmnet with reason: host reimage * 06:50 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1244.eqiad.wmnet with reason: host reimage * 06:38 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1250.eqiad.wmnet with OS trixie * 06:34 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db1244.eqiad.wmnet with OS trixie * 06:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1244: Upgrading db1244.eqiad.wmnet * 06:25 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1244: Upgrading db1244.eqiad.wmnet * 06:25 cwilliams@cumin1003: dbmaint on s4@eqiad [[phab:T429893|T429893]] * 06:25 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 06:15 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1250.eqiad.wmnet with reason: host reimage * 06:14 cwilliams@dns1006: END - running authdns-update * 06:12 cwilliams@dns1006: START - running authdns-update * 06:11 cwilliams@dns1006: END - running authdns-update * 06:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db1244 [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94676 and previous config saved to /var/cache/conftool/dbconfig/20260702-061059-cwilliams.json * 06:09 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1250.eqiad.wmnet with reason: host reimage * 06:09 cwilliams@dns1006: START - running authdns-update * 06:08 aokoth@cumin1003: END (PASS) - Cookbook sre.vrts.upgrade (exit_code=0) on VRTS host vrts1003.eqiad.wmnet * 06:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db1160 to s4 primary and set section read-write [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94675 and previous config saved to /var/cache/conftool/dbconfig/20260702-060746-cwilliams.json * 06:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Set s4 eqiad as read-only for maintenance - [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94674 and previous config saved to /var/cache/conftool/dbconfig/20260702-060704-cwilliams.json * 06:06 cezmunsta: Starting s4 eqiad failover from db1244 to db1160 - [[phab:T430817|T430817]] * 06:04 aokoth@cumin1003: START - Cookbook sre.vrts.upgrade on VRTS host vrts1003.eqiad.wmnet * 05:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db1160 with weight 0 [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94673 and previous config saved to /var/cache/conftool/dbconfig/20260702-055927-cwilliams.json * 05:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 40 hosts with reason: Primary switchover s4 [[phab:T430817|T430817]] * 05:55 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1250.eqiad.wmnet with OS trixie * 05:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on db1250.eqiad.wmnet with reason: m3 master switchover [[phab:T430158|T430158]] * 05:39 marostegui: Failover m3 (phabricator) from db1250 to db1228 - [[phab:T430158|T430158]] * 05:32 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2234].codfw.wmnet,db[1217,1228,1250].eqiad.wmnet with reason: m3 master switchover [[phab:T430158|T430158]] * 04:45 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] (duration: 09m 08s) * 04:41 tstarling@deploy1003: tstarling, reedy: Continuing with deployment * 04:38 tstarling@deploy1003: tstarling, reedy: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 04:36 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 59s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:16 ryankemper: [[phab:T429844|T429844]] [opensearch] completed `cirrussearch2111` reimage; all codfw search clusters are green, all nodes now report `OpenSearch 2.19.5`, and the temporary chi voting exclusion has been removed * 00:57 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2111.codfw.wmnet with OS trixie * 00:29 ryankemper: [[phab:T429844|T429844]] [opensearch] depooled codfw search-omega/search-psi discovery records to match existing codfw search depool during OpenSearch 2.19 migration * 00:29 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2111.codfw.wmnet with reason: host reimage * 00:29 ryankemper@cumin2002: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 00:29 ryankemper@cumin2002: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 00:22 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2111.codfw.wmnet with reason: host reimage * 00:01 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2111.codfw.wmnet with OS trixie * 00:00 ryankemper: [[phab:T429844|T429844]] [opensearch] chi cluster recovered after stopping `opensearch_1@production-search-codfw` on `cirrussearch2111` == 2026-07-01 == * 23:59 ryankemper: [[phab:T429844|T429844]] [opensearch] stopped `opensearch_1@production-search-codfw` on `cirrussearch2111` after chi cluster-manager election churn following `voting_config_exclusions` POST; hoping this triggers a re-election * 23:52 cscott@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 23:51 cscott@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 23:51 cscott@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 23:50 cscott@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2003.codfw.wmnet with OS bookworm * 22:29 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 22:13 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 22:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2084.codfw.wmnet with OS trixie * 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2003.codfw.wmnet with reason: host reimage * 22:03 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 22:01 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2003.codfw.wmnet with reason: host reimage * 21:50 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 21:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2084.codfw.wmnet with reason: host reimage * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2003 * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2003 * 21:42 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2003 * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2003.codfw.wmnet 45.48.192.10.in-addr.arpa 5.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:42 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2003.codfw.wmnet 45.48.192.10.in-addr.arpa 5.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2003 - bking@cumin2003" * 21:42 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2003 - bking@cumin2003" * 21:36 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2084.codfw.wmnet with reason: host reimage * 21:35 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:34 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2003 * 21:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2003.codfw.wmnet with OS bookworm * 21:19 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2084.codfw.wmnet with OS trixie * 21:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2081.codfw.wmnet with OS trixie * 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2108.codfw.wmnet with OS trixie * 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2081.codfw.wmnet with reason: host reimage * 20:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2081.codfw.wmnet with reason: host reimage * 20:28 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2081.codfw.wmnet with OS trixie * 20:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2108.codfw.wmnet with reason: host reimage * 20:19 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2108.codfw.wmnet with reason: host reimage * 19:59 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2108.codfw.wmnet with OS trixie * 19:46 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2093.codfw.wmnet with OS trixie * 19:44 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 19:44 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jasmine@cumin2002" * 19:43 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jasmine@cumin2002" * 19:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2080.codfw.wmnet with OS trixie * 19:28 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 19:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2093.codfw.wmnet with reason: host reimage * 19:18 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 19:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2093.codfw.wmnet with reason: host reimage * 19:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2080.codfw.wmnet with reason: host reimage * 19:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2080.codfw.wmnet with reason: host reimage * 18:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2093.codfw.wmnet with OS trixie * 18:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2080.codfw.wmnet with OS trixie * 18:27 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 18:18 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] (duration: 09m 15s) * 18:13 jgiannelos@deploy1003: jgiannelos, neriah: Continuing with deployment * 18:11 jgiannelos@deploy1003: jgiannelos, neriah: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:09 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] * 17:40 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 16:58 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 30 hosts * 16:57 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for 30 hosts * 16:52 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2202.codfw.wmnet * 16:52 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2202.codfw.wmnet * 16:51 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt * 16:51 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt * 16:51 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lvs2012.codfw.wmnet * 16:51 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for lvs2012.codfw.wmnet * 16:49 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2076.codfw.wmnet with OS trixie * 16:49 brett: Start pybal on lvs2012 - [[phab:T429861|T429861]] * 16:49 pt1979@cumin1003: END (ERROR) - Cookbook sre.hosts.remove-downtime (exit_code=97) for 59 hosts * 16:48 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for 59 hosts * 16:42 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2061.codfw.wmnet with OS trixie * 16:30 dancy@deploy1003: Installation of scap version "4.271.0" completed for 2 hosts * 16:28 dancy@deploy1003: Installing scap version "4.271.0" for 2 host(s) * 16:23 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2076.codfw.wmnet with reason: host reimage * 16:19 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2061.codfw.wmnet with reason: host reimage * 16:18 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2076.codfw.wmnet with reason: host reimage * 16:18 jasmine@dns1004: END - running authdns-update * 16:16 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host restbase2039.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 16:16 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host restbase2039.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 16:16 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2061.codfw.wmnet with reason: host reimage * 16:15 jasmine@dns1004: START - running authdns-update * 16:14 jasmine@dns1004: END - running authdns-update * 16:12 jasmine@dns1004: START - running authdns-update * 16:07 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2202.codfw.wmnet with reason: maintenance * 16:06 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt with reason: Junos upograde * 16:00 papaul: ongoing maintenance on lsw1-b2-codfw * 16:00 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2076.codfw.wmnet with OS trixie * 15:59 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt * 15:59 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt * 15:57 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2061.codfw.wmnet with OS trixie * 15:55 pt1979@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2042,2046].codfw.wmnet * 15:55 pt1979@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2042,2046].codfw.wmnet * 15:51 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 15:51 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2220: Repooling after switchover * 15:50 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 15:50 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 15:48 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2092.codfw.wmnet with OS trixie * 15:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 15:40 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 15:38 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 15:37 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 15:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply * 15:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply * 15:32 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 15:32 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 15:30 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 15:29 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 15:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 15:25 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:22 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs2012.codfw.wmnet with reason: Rack B2 maintenance - [[phab:T429861|T429861]] * 15:21 brett: Stopping pybal on lvs2012 in preparation for codfw rack b2 maintenance - [[phab:T429861|T429861]] * 15:20 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2092.codfw.wmnet with reason: host reimage * 15:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:12 _joe_: restarted manually alertmanager-irc-relay * 15:12 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2092.codfw.wmnet with reason: host reimage * 15:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:12 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt with reason: Junos upograde * 15:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover * 15:07 pt1979@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2042,2046].codfw.wmnet * 15:06 pt1979@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2042,2046].codfw.wmnet * 15:02 papaul: ongoing maintenance on lsw1-a8-codfw * 14:31 topranks: POWERING DOWN CR1-EQIAD for line card installation [[phab:T426343|T426343]] * 14:31 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] (duration: 08m 57s) * 14:29 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:26 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 14:24 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:22 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] * 14:22 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover * 14:16 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:15 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover * 14:14 topranks: re-enable routing-engine graceful-failover on cr1-eqiad [[phab:T417873|T417873]] * 14:13 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:13 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2220: Repooling after switchover * 14:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:12 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:12 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:11 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:08 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] (duration: 10m 01s) * 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:07 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2220 [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94664 and previous config saved to /var/cache/conftool/dbconfig/20260701-140729-fceratto.json * 14:06 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:06 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:06 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:05 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2159 to s7 primary [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94663 and previous config saved to /var/cache/conftool/dbconfig/20260701-140503-fceratto.json * 14:04 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:04 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 14:04 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 14:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:04 dreamyjazz@deploy1003: anzx, dreamyjazz: Continuing with deployment * 14:04 federico3: Starting s7 codfw failover from db2220 to db2159 - [[phab:T430826|T430826]] * 14:03 jmm@dns1004: END - running authdns-update * 14:03 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:03 topranks: flipping cr1-eqiad active routing-enginer back to RE0 [[phab:T417873|T417873]] * 14:03 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudsw1-c8-eqiad,cloudsw1-d5-eqiad with reason: router upgrades eqiad * 14:01 jmm@dns1004: START - running authdns-update * 14:00 dreamyjazz@deploy1003: anzx, dreamyjazz: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:59 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2159 with weight 0 [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94662 and previous config saved to /var/cache/conftool/dbconfig/20260701-135906-fceratto.json * 13:58 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] * 13:57 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s7 [[phab:T430826|T430826]] * 13:56 topranks: reboot routing-enginer RE0 on cr1-eqiad [[phab:T417873|T417873]] * 13:48 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1006.wikimedia.org * 13:44 atsuko@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cirrussearch2092.codfw.wmnet with OS trixie * 13:43 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1006.wikimedia.org * 13:41 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2092.codfw.wmnet with OS trixie * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1005.wikimedia.org * 13:37 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1005.wikimedia.org * 13:37 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on pfw1-eqiad with reason: router upgrades eqiad * 13:35 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on lvs[1017-1020].eqiad.wmnet with reason: router upgrades eqiad * 13:34 caro@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] (duration: 07m 59s) * 13:30 caro@deploy1003: caro: Continuing with deployment * 13:28 caro@deploy1003: caro: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:27 topranks: route-engine failover cr1-eqiad * 13:26 caro@deploy1003: Started scap sync-world: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] * 13:15 topranks: rebooting routing-engine 1 on cr1-eqiad [[phab:T417873|T417873]] * 13:13 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] (duration: 08m 29s) * 13:13 moritzm: installing qemu security updates * 13:11 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 13:11 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 13:09 jgiannelos@deploy1003: jgiannelos: Continuing with deployment * 13:08 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 13:07 jgiannelos@deploy1003: jgiannelos: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:06 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 13:06 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2214.codfw.wmnet with reason: Maintenance * 13:05 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2214: Repooling after switchover * 13:05 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] * 13:04 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2214: Repooling after switchover * 13:04 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2214 [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94660 and previous config saved to /var/cache/conftool/dbconfig/20260701-130413-fceratto.json * 13:01 moritzm: installing python3.13 security updates * 13:00 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2229 to s6 primary [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94659 and previous config saved to /var/cache/conftool/dbconfig/20260701-125959-fceratto.json * 12:59 federico3: Starting s6 codfw failover from db2214 to db2229 - [[phab:T430814|T430814]] * 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on 13 hosts with reason: router upgrade and line card install * 12:51 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2229 with weight 0 [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94658 and previous config saved to /var/cache/conftool/dbconfig/20260701-125149-fceratto.json * 12:51 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 21 hosts with reason: Primary switchover s6 [[phab:T430814|T430814]] * 12:50 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2189.codfw.wmnet * 12:50 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2189.codfw.wmnet * 12:42 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2100.codfw.wmnet with OS trixie * 12:38 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2083.codfw.wmnet with OS trixie * 12:19 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2083.codfw.wmnet with reason: host reimage * 12:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 12:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2240: Migration of db2240.codfw.wmnet completed * 12:14 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2100.codfw.wmnet with reason: host reimage * 12:09 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2083.codfw.wmnet with reason: host reimage * 12:09 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2100.codfw.wmnet with reason: host reimage * 12:00 topranks: drain traffic on cr1-eqiad to allow for line card install and JunOS upgrade [[phab:T426343|T426343]] * 11:52 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2083.codfw.wmnet with OS trixie * 11:50 cmooney@dns2005: END - running authdns-update * 11:49 cmooney@dns2005: START - running authdns-update * 11:48 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2100.codfw.wmnet with OS trixie * 11:40 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/zotero: apply * 11:40 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/zotero: apply * 11:36 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/zotero: apply * 11:36 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/zotero: apply * 11:31 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2240: Migration of db2240.codfw.wmnet completed * 11:30 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply * 11:28 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply * 11:27 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:27 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:27 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:27 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:27 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:23 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2240.codfw.wmnet with OS trixie * 11:20 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:20 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:17 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:16 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:16 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:15 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2086.codfw.wmnet with OS trixie * 11:14 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2106.codfw.wmnet with OS trixie * 11:14 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:13 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:12 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:09 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2115.codfw.wmnet with OS trixie * 11:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2240.codfw.wmnet with reason: host reimage * 11:00 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2240.codfw.wmnet with reason: host reimage * 10:53 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2106.codfw.wmnet with reason: host reimage * 10:49 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2086.codfw.wmnet with reason: host reimage * 10:44 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2115.codfw.wmnet with reason: host reimage * 10:44 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2240.codfw.wmnet with OS trixie * 10:44 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2086.codfw.wmnet with reason: host reimage * 10:42 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2106.codfw.wmnet with reason: host reimage * 10:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2240: Upgrading db2240.codfw.wmnet * 10:41 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2240: Upgrading db2240.codfw.wmnet * 10:41 cwilliams@cumin1003: dbmaint on s4@codfw [[phab:T429893|T429893]] * 10:40 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 10:39 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2115.codfw.wmnet with reason: host reimage * 10:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2240 [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94653 and previous config saved to /var/cache/conftool/dbconfig/20260701-102658-cwilliams.json * 10:26 moritzm: installing nginx security updates * 10:26 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2086.codfw.wmnet with OS trixie * 10:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2179 to s4 primary [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94652 and previous config saved to /var/cache/conftool/dbconfig/20260701-102356-cwilliams.json * 10:23 cezmunsta: Starting s4 codfw failover from db2240 to db2179 - [[phab:T430127|T430127]] * 10:23 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2106.codfw.wmnet with OS trixie * 10:20 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2115.codfw.wmnet with OS trixie * 10:15 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2179 with weight 0 [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94651 and previous config saved to /var/cache/conftool/dbconfig/20260701-101531-cwilliams.json * 10:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 40 hosts with reason: Primary switchover s4 [[phab:T430127|T430127]] * 09:56 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template (take 2) - oblivian@cumin1003" * 09:56 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template (take 2) - oblivian@cumin1003 * 09:55 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template (take 2) - oblivian@cumin1003 * 09:55 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template (take 2) - oblivian@cumin1003" * 09:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:39 mszwarc@deploy1003: Synchronized private/SuggestedInvestigationsSignals/SuggestedInvestigationsSignal4n.php: Update SI signal 4n (duration: 06m 08s) * 09:21 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 09:21 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 09:14 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 09:14 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 09:02 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 09:02 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 08:54 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 08:38 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 08:38 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 08:36 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 08:21 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 08:21 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] (duration: 36m 11s) * 08:15 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 08:09 mszwarc@deploy1003: mszwarc, abi: Continuing with deployment * 08:03 mszwarc@deploy1003: mszwarc, abi: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:55 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 07:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 07:45 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] * 07:30 aqu@deploy1003: Finished deploy [analytics/refinery@410f205]: Regular analytics weekly train 2nd try [analytics/refinery@410f2050] (duration: 00m 22s) * 07:29 aqu@deploy1003: Started deploy [analytics/refinery@410f205]: Regular analytics weekly train 2nd try [analytics/refinery@410f2050] * 07:28 aqu@deploy1003: Finished deploy [analytics/refinery@410f205] (thin): Regular analytics weekly train THIN [analytics/refinery@410f2050] (duration: 01m 59s) * 07:28 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] (duration: 07m 19s) * 07:26 aqu@deploy1003: Started deploy [analytics/refinery@410f205] (thin): Regular analytics weekly train THIN [analytics/refinery@410f2050] * 07:26 aqu@deploy1003: Finished deploy [analytics/refinery@410f205]: Regular analytics weekly train [analytics/refinery@410f2050] (duration: 04m 32s) * 07:24 mszwarc@deploy1003: wmde-fisch, mszwarc: Continuing with deployment * 07:23 mszwarc@deploy1003: wmde-fisch, mszwarc: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:21 aqu@deploy1003: Started deploy [analytics/refinery@410f205]: Regular analytics weekly train [analytics/refinery@410f2050] * 07:21 aqu@deploy1003: Finished deploy [analytics/refinery@410f205] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@410f2050] (duration: 02m 01s) * 07:20 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] * 07:19 aqu@deploy1003: Started deploy [analytics/refinery@410f205] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@410f2050] * 07:13 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] (duration: 09m 13s) * 07:09 mszwarc@deploy1003: mszwarc, chlod, revi: Continuing with deployment * 07:06 mszwarc@deploy1003: mszwarc, chlod, revi: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:04 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] * 06:55 elukey: upgrade all trixie hosts to pywmflib 3.0 - [[phab:T430552|T430552]] * 06:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:43 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:43 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:42 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:42 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:41 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:41 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:35 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:35 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:34 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:34 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:31 jmm@cumin2003: DONE (PASS) - Cookbook sre.idm.logout (exit_code=0) Logging Niharika29 out of all services on: 2453 hosts * 06:30 oblivian@cumin1003: END (FAIL) - Cookbook sre.deploy.hiddenparma (exit_code=99) Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:30 oblivian@cumin1003: END (FAIL) - Cookbook sre.deploy.python-code (exit_code=99) hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:30 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:30 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:01 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2109.codfw.wmnet with OS trixie * 05:45 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on es1039.eqiad.wmnet with reason: issues * 05:41 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1027.eqiad.wmnet * 05:40 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2068.codfw.wmnet with OS trixie * 05:40 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2109.codfw.wmnet with reason: host reimage * 05:40 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1027.eqiad.wmnet,service=s2 * 05:40 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1027.eqiad.wmnet,service=s7 * 05:36 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2109.codfw.wmnet with reason: host reimage * 05:20 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2068.codfw.wmnet with reason: host reimage * 05:16 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2109.codfw.wmnet with OS trixie * 05:15 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2068.codfw.wmnet with reason: host reimage * 05:09 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2067.codfw.wmnet with OS trixie * 04:56 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2068.codfw.wmnet with OS trixie * 04:49 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2067.codfw.wmnet with reason: host reimage * 04:45 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2067.codfw.wmnet with reason: host reimage * 04:27 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2067.codfw.wmnet with OS trixie * 03:47 slyngshede@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1039.eqiad.wmnet with reason: Hardware crash * 03:21 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2107.codfw.wmnet with OS trixie * 02:59 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2107.codfw.wmnet with reason: host reimage * 02:55 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2085.codfw.wmnet with OS trixie * 02:51 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2072.codfw.wmnet with OS trixie * 02:51 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2107.codfw.wmnet with reason: host reimage * 02:35 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2085.codfw.wmnet with reason: host reimage * 02:31 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2107.codfw.wmnet with OS trixie * 02:30 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2072.codfw.wmnet with reason: host reimage * 02:26 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2085.codfw.wmnet with reason: host reimage * 02:22 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2072.codfw.wmnet with reason: host reimage * 02:09 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2085.codfw.wmnet with OS trixie * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 54s) * 02:03 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2072.codfw.wmnet with OS trixie * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es7 eqiad back to read-write - [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94649 and previous config saved to /var/cache/conftool/dbconfig/20260701-010716-ladsgroup.json * 01:05 ladsgroup@dns1004: END - running authdns-update * 01:05 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depool es1039 [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94648 and previous config saved to /var/cache/conftool/dbconfig/20260701-010551-ladsgroup.json * 01:03 ladsgroup@dns1004: START - running authdns-update * 01:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Promote es1035 to es7 primary [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94647 and previous config saved to /var/cache/conftool/dbconfig/20260701-010002-ladsgroup.json * 00:58 Amir1: Starting es7 eqiad failover from es1039 to es1035 - [[phab:T430765|T430765]] * 00:53 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es1035 with weight 0 [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94646 and previous config saved to /var/cache/conftool/dbconfig/20260701-005329-ladsgroup.json * 00:53 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 9 hosts with reason: Primary switchover es7 [[phab:T430765|T430765]] * 00:42 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es7 eqiad as read-only for maintenance - [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94645 and previous config saved to /var/cache/conftool/dbconfig/20260701-004221-ladsgroup.json * 00:20 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2102.codfw.wmnet with OS trixie * 00:15 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2103.codfw.wmnet with OS trixie * 00:05 dr0ptp4kt: DEPLOYED Refinery at {{Gerrit|4e7a2b32}} for changes: pageview allowlist {{Gerrit|1305158}} (+min.wikiquote) {{Gerrit|1305162}} (+bol.wikipedia), {{Gerrit|1305156}} (+isv.wikipedia); {{Gerrit|1305980}} (pv allowlist -api.wikimedia, sqoop +isvwiki); sqoop {{Gerrit|1295064}} (+globalimagelinks) {{Gerrit|1295069}} (+filerevision) using scap, then deployed onto HDFS (manual copyToLocal required additionally) == Other archives == See [[Server Admin Log/Archives]]. <noinclude> [[Category:SAL]] [[Category:Operations]] </noinclude> akunfqdy8e5tsg2ixr0ixpa5w0y1tbh 2450657 2450656 2026-08-22T16:37:24Z Stashbot 7414 arlolra@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply 2450657 wikitext text/x-wiki == 2026-08-22 == * 16:37 arlolra@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:37 arlolra@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:36 arlolra@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:36 arlolra@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:36 arlolra@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:36 arlolra@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:33 arlolra@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:33 arlolra@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:33 arlolra@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:32 arlolra@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 35s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-21 == * 20:36 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in1001.wikimedia.org with reason: [[phab:T434750|T434750]] * 20:34 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in2001.wikimedia.org with reason: [[phab:T434750|T434750]] * 20:33 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out1001.wikimedia.org with reason: [[phab:T434750|T434750]] * 20:25 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out2001.wikimedia.org with reason: [[phab:T434750|T434750]] * 19:37 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host krb1002.eqiad.wmnet with OS bookworm * 19:00 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 18:59 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 18:51 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 18:51 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 18:35 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:35 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:27 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:27 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:16 bking@cumin2003: START - Cookbook sre.hosts.reimage for host krb1002.eqiad.wmnet with OS bookworm * 17:35 sukhe@dns1004: END - running authdns-update * 17:33 sukhe@dns1004: START - running authdns-update * 17:32 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns5004.wikimedia.org [reason: resolved authdns-update issues] * 17:32 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:32 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: force HEAD to {{Gerrit|be26e30ae101}} - sukhe@cumin1003" * 17:32 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: force HEAD to {{Gerrit|be26e30ae101}} - sukhe@cumin1003" * 17:28 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 17:28 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: service=authdns-update,name=dns5004.wikimedia.org [reason: resolving authdns-update issues] * 17:28 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:28 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: force HEAD to {{Gerrit|be26e30ae101}} - sukhe@cumin1003" * 17:28 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: force HEAD to {{Gerrit|be26e30ae101}} - sukhe@cumin1003" * 17:24 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 17:24 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.netbox (exit_code=97) * 17:23 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 17:16 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 17:12 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 17:10 sukhe@dns1004: END - running authdns-update * 17:08 sukhe@dns1004: START - running authdns-update * 17:08 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=dns5004.wikimedia.org [reason: resolving authdns-update issues] * 17:07 sukhe@dns1004: FAIL - running authdns-update * 17:05 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 17:05 sukhe@dns1004: START - running authdns-update * 17:01 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 16:59 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 16:56 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 16:53 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=dns5004.* [reason: trixie upgrade] * 16:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns5004.wikimedia.org * 16:52 cdobbins@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns5004.wikimedia.org * 16:44 cmooney@dns3003: END - running authdns-update * 16:41 cmooney@dns3003: START - running authdns-update * 16:41 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:41 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on eqsin<->codfw arelion - cmooney@cumin1003" * 16:37 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on eqsin<->codfw arelion - cmooney@cumin1003" * 16:33 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:11 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 16:08 sukhe@dns1004: END - running authdns-update * 16:08 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 16:06 sukhe@dns1004: START - running authdns-update * 16:04 cmooney@dns3003: END - running authdns-update * 16:02 cmooney@dns3003: START - running authdns-update * 16:00 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:00 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on eqord<->codfw arelion - cmooney@cumin1003" * 15:56 cdobbins@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host dns5004.wikimedia.org with OS trixie * 15:55 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on eqord<->codfw arelion - cmooney@cumin1003" * 15:53 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:51 cmooney@cumin1003: END (ERROR) - Cookbook sre.dns.netbox (exit_code=97) * 15:51 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:38 andrew@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudcephosd1042.eqiad.wmnet with OS bookworm * 15:18 andrew@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudcephosd1042.eqiad.wmnet with reason: host reimage * 15:17 dancy@deploy1003: Finished deploy [gerrit/gerrit@2cc11cc]: Deploying https://gerrit.wikimedia.org/r/c/operations/software/gerrit/+/1327669 ([[phab:T434726|T434726]]) (duration: 00m 14s) * 15:17 dancy@deploy1003: Started deploy [gerrit/gerrit@2cc11cc]: Deploying https://gerrit.wikimedia.org/r/c/operations/software/gerrit/+/1327669 ([[phab:T434726|T434726]]) * 15:13 andrew@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudcephosd1042.eqiad.wmnet with reason: host reimage * 15:09 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns5004.wikimedia.org with reason: host reimage * 15:05 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns5004.wikimedia.org with reason: host reimage * 14:53 andrew@cumin2003: START - Cookbook sre.hosts.reimage for host cloudcephosd1042.eqiad.wmnet with OS bookworm * 14:30 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns5004.wikimedia.org with OS trixie * 14:29 cdobbins@cumin1003: conftool action : set/pooled=no; selector: name=dns5004.* [reason: trixie upgrade] * 14:21 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:21 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove entries for cr2-eqord - cmooney@cumin1003" * 14:21 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove entries for cr2-eqord - cmooney@cumin1003" * 14:13 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 14:11 moritzm: imported openjdk 8u504-ga-1~deb12u1 for bookworm-wikimedia (backport of the latest Java 8 security fixes for bookworm) * 13:25 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "sync cr2-eqord router offline - cmooney@cumin1003" * 13:23 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "sync cr2-eqord router offline - cmooney@cumin1003" * 13:14 hashar@deploy1003: Finished deploy [integration/docroot@2d5ff9b]: opensource: add PersonalDashboard docs to MW components - [[phab:T435392|T435392]] (duration: 00m 15s) * 13:14 hashar@deploy1003: Started deploy [integration/docroot@2d5ff9b]: opensource: add PersonalDashboard docs to MW components - [[phab:T435392|T435392]] * 12:16 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:16 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: [[phab:T431682|T431682]] - filippo@cumin1003" * 12:16 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: [[phab:T431682|T431682]] - filippo@cumin1003" * 12:11 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2006.wikimedia.org with OS trixie * 12:00 kevinbazira@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 11:58 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 11:43 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2006.wikimedia.org with reason: host reimage * 11:41 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 11:38 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2006.wikimedia.org with reason: host reimage * 11:20 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2006.wikimedia.org with OS trixie * 11:11 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2005.wikimedia.org with OS trixie * 10:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2005.wikimedia.org with reason: host reimage * 10:53 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2005.wikimedia.org with reason: host reimage * 10:43 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-codfw * 10:43 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2011.codfw.wmnet * 10:43 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2011.codfw.wmnet * 10:40 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2011.codfw.wmnet * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2011.codfw.wmnet * 10:34 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2010.codfw.wmnet * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2010.codfw.wmnet * 10:33 fnegri@cumin1003: END (PASS) - Cookbook sre.wikireplicas.add-wiki (exit_code=0) for database bolwiki ([[phab:T429954|T429954]]) * 10:33 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2005.wikimedia.org with OS trixie * 10:30 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2010.codfw.wmnet * 10:25 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2010.codfw.wmnet * 10:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2009.codfw.wmnet * 10:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2009.codfw.wmnet * 10:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1006.wikimedia.org with OS trixie * 10:18 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2009.codfw.wmnet * 10:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2009.codfw.wmnet * 10:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2008.codfw.wmnet * 10:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2008.codfw.wmnet * 10:06 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2008.codfw.wmnet * 10:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1006.wikimedia.org with reason: host reimage * 10:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2008.codfw.wmnet * 10:01 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2007.codfw.wmnet * 10:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2007.codfw.wmnet * 09:57 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1006.wikimedia.org with reason: host reimage * 09:56 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2007.codfw.wmnet * 09:51 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2007.codfw.wmnet * 09:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 09:51 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 09:46 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2006.codfw.wmnet * 09:46 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1006.wikimedia.org with OS trixie * 09:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1005.wikimedia.org with OS trixie * 09:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2006.codfw.wmnet * 09:41 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2005.codfw.wmnet * 09:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2005.codfw.wmnet * 09:38 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2013.codfw.wmnet * 09:36 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2005.codfw.wmnet * 09:35 fnegri@cumin1003: START - Cookbook sre.wikireplicas.add-wiki for database bolwiki ([[phab:T429954|T429954]]) * 09:35 fnegri@cumin1003: END (PASS) - Cookbook sre.wikireplicas.add-wiki (exit_code=0) for database minwikiquote ([[phab:T429946|T429946]]) * 09:35 fnegri@cumin1003: START - Cookbook sre.wikireplicas.add-wiki for database minwikiquote ([[phab:T429946|T429946]]) * 09:32 blake@cumin1003: START - Cookbook sre.hosts.reboot-single for host rdb2013.codfw.wmnet * 09:30 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2011.codfw.wmnet * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1005.wikimedia.org with reason: host reimage * 09:26 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2005.codfw.wmnet * 09:26 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2004.codfw.wmnet * 09:26 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2004.codfw.wmnet * 09:24 blake@cumin1003: START - Cookbook sre.hosts.reboot-single for host rdb2011.codfw.wmnet * 09:22 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1005.wikimedia.org with reason: host reimage * 09:21 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2004.codfw.wmnet * 09:16 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1015.eqiad.wmnet * 09:16 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2004.codfw.wmnet * 09:16 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2003.codfw.wmnet * 09:16 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2003.codfw.wmnet * 09:13 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on cr[1-2]-eqiad,pfw1-eqiad with reason: upgrade pfw1a-eqiad and pfw1b-eqiad pair * 09:12 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 09:11 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 09:11 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 09:11 blake@cumin1003: START - Cookbook sre.hosts.reboot-single for host rdb1015.eqiad.wmnet * 09:10 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2003.codfw.wmnet * 09:09 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1013.eqiad.wmnet * 09:07 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1005.wikimedia.org with OS trixie * 09:03 blake@cumin1003: START - Cookbook sre.hosts.reboot-single for host rdb1013.eqiad.wmnet * 09:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2003.codfw.wmnet * 09:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2002.codfw.wmnet * 09:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2002.codfw.wmnet * 08:54 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2002.codfw.wmnet * 08:49 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2002.codfw.wmnet * 08:49 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2001.codfw.wmnet * 08:49 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2001.codfw.wmnet * 08:48 jmm@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts netmon2002.wikimedia.org * 08:47 jmm@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts netmon2002.wikimedia.org * 08:44 jmm@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts netmon2002.wikimedia.org * 08:44 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon2002.wikimedia.org * 08:43 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2001.codfw.wmnet * 08:36 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon2002.wikimedia.org * 08:34 jmm@dns1004: END - running authdns-update * 08:33 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2001.codfw.wmnet * 08:33 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-codfw * 08:31 jmm@dns1004: START - running authdns-update * 07:48 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327679{{!}}Block: Disable flaky API test (T435272 T389028)]], [[gerrit:1327678{{!}}API: wfDebugLog for thumberror]] (duration: 15m 34s) * 07:41 krinkle@deploy1003: krinkle: Continuing with deployment * 07:37 krinkle@deploy1003: krinkle: Backport for [[gerrit:1327679{{!}}Block: Disable flaky API test (T435272 T389028)]], [[gerrit:1327678{{!}}API: wfDebugLog for thumberror]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:33 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1327679{{!}}Block: Disable flaky API test (T435272 T389028)]], [[gerrit:1327678{{!}}API: wfDebugLog for thumberror]] * 07:25 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 07:24 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 07:18 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 07:18 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 07:15 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 07:14 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 07:14 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 07:14 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 07:13 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 07:03 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1283: Pool back * 06:42 jmm@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts netmon2002.wikimedia.org * 06:35 moritzm: powercycling netmon2002 * 06:18 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1283: Pool back * 06:17 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1283 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96213 and previous config saved to /var/cache/conftool/dbconfig/20260821-061743-marostegui.json * 04:59 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324963{{!}}Add Produnto to extension-list (T421436)]], [[gerrit:1324964{{!}}Enable Produnto on Beta (T421436)]] (duration: 34m 48s) * 04:45 tstarling@deploy1003: tstarling: Continuing with deployment * 04:44 tstarling@deploy1003: tstarling: Backport for [[gerrit:1324963{{!}}Add Produnto to extension-list (T421436)]], [[gerrit:1324964{{!}}Enable Produnto on Beta (T421436)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 04:24 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1324963{{!}}Add Produnto to extension-list (T421436)]], [[gerrit:1324964{{!}}Enable Produnto on Beta (T421436)]] * 04:21 arlolra@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 04:20 arlolra@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 04:20 arlolra@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 04:20 arlolra@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 41s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-20 == * 23:43 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327652{{!}}RunSingleJob: Add ProfilingContext::init() (T435422)]] (duration: 12m 06s) * 23:38 krinkle@deploy1003: krinkle: Continuing with deployment * 23:33 krinkle@deploy1003: krinkle: Backport for [[gerrit:1327652{{!}}RunSingleJob: Add ProfilingContext::init() (T435422)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:31 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1327652{{!}}RunSingleJob: Add ProfilingContext::init() (T435422)]] * 22:15 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1054.eqiad.wmnet with OS trixie * 22:14 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 22:14 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 21:58 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1054.eqiad.wmnet with reason: host reimage * 21:51 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1054.eqiad.wmnet with reason: host reimage * 21:36 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1054.eqiad.wmnet with OS trixie * 21:36 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:35 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327219{{!}}RunSingleJob: Define MW_ENTRY_POINT for flamegraph sample attribution (T435422)]] (duration: 08m 30s) * 21:31 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:31 krinkle@deploy1003: krinkle: Continuing with deployment * 21:31 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1054 * 21:31 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1054 * 21:29 krinkle@deploy1003: krinkle: Backport for [[gerrit:1327219{{!}}RunSingleJob: Define MW_ENTRY_POINT for flamegraph sample attribution (T435422)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:27 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1327219{{!}}RunSingleJob: Define MW_ENTRY_POINT for flamegraph sample attribution (T435422)]] * 21:17 maryum: Deployed security fix for [[phab:T433020|T433020]] * 20:59 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324752{{!}}InitialiseSettings: Enable 2FA banners on remaining private wikis (T428103)]], [[gerrit:1325920{{!}}Remove sending email to legal team about rejected requests (T374053)]] (duration: 07m 18s) * 20:54 reedy@deploy1003: neriah, reedy: Continuing with deployment * 20:54 reedy@deploy1003: neriah, reedy: Backport for [[gerrit:1324752{{!}}InitialiseSettings: Enable 2FA banners on remaining private wikis (T428103)]], [[gerrit:1325920{{!}}Remove sending email to legal team about rejected requests (T374053)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:51 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324752{{!}}InitialiseSettings: Enable 2FA banners on remaining private wikis (T428103)]], [[gerrit:1325920{{!}}Remove sending email to legal team about rejected requests (T374053)]] * 20:24 reedy@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.15,1.47.0-wmf.16,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/med * 20:23 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324752{{!}}InitialiseSettings: Enable 2FA banners on remaining private wikis (T428103)]], [[gerrit:1325920{{!}}Remove sending email to legal team about rejected requests (T374053)]] * 20:14 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 20:10 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 20:09 cdanis@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "bug fixes & UX fixes - cdanis@cumin1003" * 20:09 cdanis@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: bug fixes & UX fixes - cdanis@cumin1003 * 20:08 cdanis@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: bug fixes & UX fixes - cdanis@cumin1003 * 20:08 cdanis@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "bug fixes & UX fixes - cdanis@cumin1003" * 19:24 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327598{{!}}Make \Omicron non upright (like \Chi) (T434428)]], [[gerrit:1327596{{!}}Render overline of \bar with stretchy=false (T435456)]] (duration: 18m 54s) * 19:20 krinkle@deploy1003: krinkle: Continuing with deployment * 19:07 krinkle@deploy1003: krinkle: Backport for [[gerrit:1327598{{!}}Make \Omicron non upright (like \Chi) (T434428)]], [[gerrit:1327596{{!}}Render overline of \bar with stretchy=false (T435456)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:05 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1327598{{!}}Make \Omicron non upright (like \Chi) (T434428)]], [[gerrit:1327596{{!}}Render overline of \bar with stretchy=false (T435456)]] * 18:51 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327614{{!}}Avoid casting fpxmax to string (T318419)]] (duration: 07m 28s) * 18:50 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:46 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:46 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 18:45 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327614{{!}}Avoid casting fpxmax to string (T318419)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:43 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327614{{!}}Avoid casting fpxmax to string (T318419)]] * 18:37 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:37 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:37 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:36 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:07 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 18:05 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 18:01 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 18:01 sukhe@dns1004: END - running authdns-update * 17:59 sukhe@dns1004: START - running authdns-update * 17:58 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 17:57 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=dns6002.* [reason: depooling for trixie upgrade] * 17:56 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns6002.wikimedia.org * 17:56 cdobbins@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns6002.wikimedia.org * 17:51 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 17:51 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 17:34 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host stat1011.eqiad.wmnet with OS bookworm * 17:31 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns6002.wikimedia.org with OS trixie * 17:30 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 17:30 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 17:30 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 17:30 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 17:29 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 17:29 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 17:29 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 17:29 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:29 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:27 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:24 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327590{{!}}Make sure fpsmax is an int value (T318419)]] (duration: 08m 37s) * 17:20 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 17:17 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327590{{!}}Make sure fpsmax is an int value (T318419)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:16 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 17:15 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327590{{!}}Make sure fpsmax is an int value (T318419)]] * 16:51 swfrench-wmf: disable-puppet on A:cp for ATS Lua change - [[phab:T427666|T427666]] * 16:51 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db2901.codfw.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 16:43 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 16:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on stat1011.eqiad.wmnet with reason: host reimage * 16:39 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns6002.wikimedia.org with reason: host reimage * 16:36 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on stat1011.eqiad.wmnet with reason: host reimage * 16:36 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db2901.codfw.wmnet * 16:34 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns6002.wikimedia.org with reason: host reimage * 16:15 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns6002.wikimedia.org with OS trixie * 16:14 cdobbins@cumin1003: conftool action : set/pooled=no; selector: name=dns6002.* [reason: depooling for trixie upgrade] * 16:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host stat1011 * 16:10 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host stat1011 * 16:09 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host stat1011 * 16:09 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) stat1011.eqiad.wmnet 14.36.64.10.in-addr.arpa 4.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 bking@cumin2003: START - Cookbook sre.dns.wipe-cache stat1011.eqiad.wmnet 14.36.64.10.in-addr.arpa 4.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:09 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host stat1011 - bking@cumin2003" * 16:09 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host stat1011 - bking@cumin2003" * 16:05 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host stat1011 * 16:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host stat1011.eqiad.wmnet with OS bookworm * 16:00 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 15:59 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 15:56 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 15:55 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 15:37 jayme@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:35 jayme@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 15:35 jayme@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:32 jayme@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:32 jayme@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:30 jayme@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 15:30 jayme@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:28 jayme@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:28 jayme@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 15:28 fceratto@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host db1903.eqiad.wmnet * 15:28 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1903.eqiad.wmnet with OS trixie * 15:26 jayme@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 15:26 jayme@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 15:24 jayme@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 15:24 jayme@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 15:21 jayme@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 15:21 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 15:19 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 15:19 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 15:17 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 15:14 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1903.eqiad.wmnet with reason: host reimage * 15:07 fceratto@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1903.eqiad.wmnet with reason: host reimage * 14:54 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db1903.eqiad.wmnet with OS trixie * 14:53 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1903.eqiad.wmnet - fceratto@cumin1003" * 14:53 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1903.eqiad.wmnet - fceratto@cumin1003" * 14:53 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1903.eqiad.wmnet on all recursors * 14:53 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1903.eqiad.wmnet on all recursors * 14:53 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:53 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1903.eqiad.wmnet - fceratto@cumin1003" * 14:53 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1903.eqiad.wmnet - fceratto@cumin1003" * 14:49 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 14:49 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1903.eqiad.wmnet * 14:33 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=cp1100.* * 14:27 topranks: reconfigure eqiad<->codfw bgp settings * 14:22 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:22 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update entries used on new transport backup eqiad codfw - cmooney@cumin1003" * 14:19 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update entries used on new transport backup eqiad codfw - cmooney@cumin1003" * 14:14 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 14:14 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 14:13 moritzm: installing util-linux security updates * 14:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-staging-worker * 14:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2003.codfw.wmnet * 14:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2003.codfw.wmnet * 14:08 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2003.codfw.wmnet * 14:06 moritzm: installing libheif security updates * 13:58 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2003.codfw.wmnet * 13:58 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2002.codfw.wmnet * 13:58 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2002.codfw.wmnet * 13:56 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:56 fnegri@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for clouddb1025.eqiad.wmnet * 13:56 fnegri@cumin1003: START - Cookbook sre.hosts.remove-downtime for clouddb1025.eqiad.wmnet * 13:56 Lucas_WMDE: UTC afternoon backport+config window done * 13:53 moritzm: installing apr-util security updates * 13:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2002.codfw.wmnet * 13:50 fnegri@cumin1003: conftool action : set/weight=100; selector: name=clouddb1025.eqiad.wmnet * 13:49 fnegri@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1025.eqiad.wmnet * 13:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2002.codfw.wmnet * 13:41 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2001.codfw.wmnet * 13:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2001.codfw.wmnet * 13:41 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'. * 13:38 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'. * 13:38 fnegri@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on clouddb1025.eqiad.wmnet with reason: Removing s6 from clouddb1025 * 13:34 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2001.codfw.wmnet * 13:31 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'. * 13:29 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'. * 13:28 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet * 13:26 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host stat1009.eqiad.wmnet with OS bookworm * 13:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2001.codfw.wmnet * 13:24 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-staging-worker * 13:23 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1002.eqiad.wmnet * 13:21 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host stat1010.eqiad.wmnet with OS bookworm * 13:20 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1002.eqiad.wmnet * 13:20 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1001.eqiad.wmnet * 13:17 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1001.eqiad.wmnet * 13:16 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2001.codfw.wmnet * 13:13 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327511{{!}}UIC: Fix page:page instead of page:other in instrumentation]] (duration: 07m 00s) * 13:13 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2001.codfw.wmnet * 13:12 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2002.codfw.wmnet * 13:09 mszwarc@deploy1003: mszwarc: Continuing with deployment * 13:08 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1327511{{!}}UIC: Fix page:page instead of page:other in instrumentation]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2002.codfw.wmnet * 13:07 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2002.codfw.wmnet * 13:06 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1327511{{!}}UIC: Fix page:page instead of page:other in instrumentation]] * 13:04 jmm@dns1004: END - running authdns-update * 13:03 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2002.codfw.wmnet * 13:03 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2001.codfw.wmnet * 13:02 jmm@dns1004: START - running authdns-update * 13:01 cmooney@dns3003: END - running authdns-update * 13:00 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2001.codfw.wmnet * 12:59 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2001.codfw.wmnet * 12:59 cmooney@dns3003: START - running authdns-update * 12:57 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2001.codfw.wmnet * 12:56 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2002.codfw.wmnet * 12:55 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:55 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on drmrs<->eqiad GTT vpls - cmooney@cumin1003" * 12:54 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on drmrs<->eqiad GTT vpls - cmooney@cumin1003" * 12:54 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2002.codfw.wmnet * 12:54 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2003.codfw.wmnet * 12:50 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2003.codfw.wmnet * 12:49 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 12:48 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2001.codfw.wmnet * 12:46 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2001.codfw.wmnet * 12:46 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2002.codfw.wmnet * 12:43 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2002.codfw.wmnet * 12:43 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2003.codfw.wmnet * 12:42 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=1) for new host db1902.eqiad.wmnet * 12:42 fceratto@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host db1902.eqiad.wmnet with OS trixie * 12:41 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2003.codfw.wmnet * 12:40 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1003.eqiad.wmnet * 12:38 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1003.eqiad.wmnet * 12:37 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1002.eqiad.wmnet * 12:35 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1002.eqiad.wmnet * 12:35 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1001.eqiad.wmnet * 12:34 cmooney@dns3003: END - running authdns-update * 12:33 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1001.eqiad.wmnet * 12:32 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on stat1009.eqiad.wmnet with reason: host reimage * 12:31 cmooney@dns3003: START - running authdns-update * 12:31 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:31 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on drmrs<->eqiad cct - cmooney@cumin1003" * 12:28 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on drmrs<->eqiad cct - cmooney@cumin1003" * 12:28 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1902.eqiad.wmnet with reason: host reimage * 12:25 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on stat1009.eqiad.wmnet with reason: host reimage * 12:25 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 12:24 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on stat1010.eqiad.wmnet with reason: host reimage * 12:22 fceratto@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1902.eqiad.wmnet with reason: host reimage * 12:21 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on stat1010.eqiad.wmnet with reason: host reimage * 12:14 elukey: move the Docker Registry's /v2/dev/.* prefix to its dedicated S3 backend - [[phab:T432829|T432829]] * 12:12 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db1902.eqiad.wmnet with OS trixie * 12:09 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1902.eqiad.wmnet - fceratto@cumin1003" * 12:09 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1902.eqiad.wmnet - fceratto@cumin1003" * 12:09 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1902.eqiad.wmnet on all recursors * 12:09 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1902.eqiad.wmnet on all recursors * 12:09 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:08 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1902.eqiad.wmnet - fceratto@cumin1003" * 12:08 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1902.eqiad.wmnet - fceratto@cumin1003" * 12:08 tgr_: [[phab:T413390|T413390]] running CentralAuth:FixRenamedUserGlobalEditCount --wiki=metawiki --since=20250901000000 --until=20260301000000 --fix * 12:04 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1009.eqiad.wmnet with OS bookworm * 12:04 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1010.eqiad.wmnet with OS bookworm * 12:01 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 12:01 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1902.eqiad.wmnet * 12:00 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host stat1010.eqiad.wmnet with OS bookworm * 11:57 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1902.eqiad.wmnet * 11:57 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:57 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1902.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 11:57 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1902.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 11:51 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'. * 11:49 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'. * 11:48 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'. * 11:46 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'. * 11:39 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 11:37 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327518{{!}}Enable thumb.wikimedia.org on cswiki and fawiki (T427465)]] (duration: 10m 40s) * 11:35 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1902.eqiad.wmnet * 11:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1010.eqiad.wmnet with OS bookworm * 11:33 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 11:30 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327518{{!}}Enable thumb.wikimedia.org on cswiki and fawiki (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:26 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327518{{!}}Enable thumb.wikimedia.org on cswiki and fawiki (T427465)]] * 11:24 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host stat1010.eqiad.wmnet with OS bookworm * 11:07 fceratto@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host db1901.eqiad.wmnet * 11:07 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1901.eqiad.wmnet with OS trixie * 10:53 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1901.eqiad.wmnet with reason: host reimage * 10:47 tappof: bump space for prometheus k8s-aux in codfw * 10:47 tappof: bump space for prometheus k8s-dse in eqiad * 10:47 fceratto@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1901.eqiad.wmnet with reason: host reimage * 10:35 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db1901.eqiad.wmnet with OS trixie * 10:32 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:32 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:32 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1901.eqiad.wmnet on all recursors * 10:32 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1901.eqiad.wmnet on all recursors * 10:31 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:31 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:31 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:27 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:27 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1901.eqiad.wmnet * 10:23 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1010.eqiad.wmnet with OS bookworm * 10:20 fceratto@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts db1901.eqiad.wmnet * 10:20 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 10:18 blake@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 10:17 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:16 blake@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 10:13 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1901.eqiad.wmnet * 09:23 jelto@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'. * 09:22 jelto@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'. * 09:22 jelto: update cert-manager to 1.19.6 on wikikube staging-eqiad - [[phab:T427402|T427402]] * 09:20 moritzm: imported squid 7.6-2.1for trixie-wikimedia/main [[phab:T427282|T427282]] * 09:08 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 09:08 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 09:08 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 09:07 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 09:04 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 09:04 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:23 slyngshede@dns1004: END - running authdns-update * 08:21 slyngshede@dns1004: START - running authdns-update * 08:18 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.16 refs [[phab:T430835|T430835]] * 06:27 aokoth@dns1004: END - running authdns-update * 06:25 aokoth@dns1004: START - running authdns-update * 06:22 brennen@deploy1003: Finished deploy [phabricator/deployment@6b9b6ff]: deploy phab1005 for [[phab:T435087|T435087]] (duration: 00m 39s) * 06:21 brennen@deploy1003: Started deploy [phabricator/deployment@6b9b6ff]: deploy phab1005 for [[phab:T435087|T435087]] * 06:20 brennen@deploy1003: Finished deploy [phabricator/deployment@6b9b6ff]: deploy phab1004 for to pick up config values for [[phab:T435087|T435087]] (duration: 01m 46s) * 06:18 brennen@deploy1003: Started deploy [phabricator/deployment@6b9b6ff]: deploy phab1004 for to pick up config values for [[phab:T435087|T435087]] * 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 49s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-19 == * 23:19 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327197{{!}}Enable thumb.wikimedia.org on mediawiki.org (T427465)]] (duration: 10m 50s) * 23:18 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:16 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 23:15 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 23:10 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327197{{!}}Enable thumb.wikimedia.org on mediawiki.org (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:10 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:09 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 23:09 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:09 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 23:08 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327197{{!}}Enable thumb.wikimedia.org on mediawiki.org (T427465)]] * 22:58 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327201{{!}}Enable ReadingLists for all logged in users on test wiki (T435258)]] (duration: 11m 20s) * 22:50 jdlrobson@deploy1003: jdlrobson: Continuing with deployment * 22:49 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1327201{{!}}Enable ReadingLists for all logged in users on test wiki (T435258)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:46 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1327201{{!}}Enable ReadingLists for all logged in users on test wiki (T435258)]] * 22:42 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327162{{!}}Article: Split subjectpageheader by model and disable for wikitext]], [[gerrit:1327169{{!}}Make uppercase greek letters normal (non-italic) font (T434686 T434428)]], [[gerrit:1327176{{!}}Skin: Avoid DB lookup for pagecategorieslink message (T347123)]] (duration: 37m 52s) * 22:29 krinkle@deploy1003: krinkle: Continuing with deployment * 22:25 krinkle@deploy1003: krinkle: Backport for [[gerrit:1327162{{!}}Article: Split subjectpageheader by model and disable for wikitext]], [[gerrit:1327169{{!}}Make uppercase greek letters normal (non-italic) font (T434686 T434428)]], [[gerrit:1327176{{!}}Skin: Avoid DB lookup for pagecategorieslink message (T347123)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:04 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1327162{{!}}Article: Split subjectpageheader by model and disable for wikitext]], [[gerrit:1327169{{!}}Make uppercase greek letters normal (non-italic) font (T434686 T434428)]], [[gerrit:1327176{{!}}Skin: Avoid DB lookup for pagecategorieslink message (T347123)]] * 22:04 krinkle@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: awaiting CI (duration: 03m 06s) * 22:01 krinkle@deploy1003: Locking from deployment [ALL REPOSITORIES]: awaiting CI * 22:00 krinkle@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: awaiting CI (duration: 00m 01s) * 22:00 krinkle@deploy1003: Locking from deployment [ALL REPOSITORIES]: awaiting CI * 21:34 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2001.codfw.wmnet * 21:28 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2001.codfw.wmnet * 21:22 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327178{{!}}AccountRecovery: Notify the email address of the on file of the request (T425799)]] (duration: 47m 02s) * 21:13 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm * 21:09 catrope@deploy1003: catrope: Continuing with deployment * 20:55 catrope@deploy1003: catrope: Backport for [[gerrit:1327178{{!}}AccountRecovery: Notify the email address of the on file of the request (T425799)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:35 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1327178{{!}}AccountRecovery: Notify the email address of the on file of the request (T425799)]] * 20:31 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327128{{!}}Parsoid DataAccess: convert Parsoid fragment markers to/from strip tags (T432547)]] (duration: 07m 30s) * 20:27 catrope@deploy1003: catrope, arlolra: Continuing with deployment * 20:26 catrope@deploy1003: catrope, arlolra: Backport for [[gerrit:1327128{{!}}Parsoid DataAccess: convert Parsoid fragment markers to/from strip tags (T432547)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:24 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1327128{{!}}Parsoid DataAccess: convert Parsoid fragment markers to/from strip tags (T432547)]] * 20:23 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage * 20:17 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage * 20:15 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325878{{!}}[arwiki] Enable restricted user page editing and grant edit permissions (T434878)]] (duration: 08m 46s) * 20:11 catrope@deploy1003: catrope, gergesshamon: Continuing with deployment * 20:08 catrope@deploy1003: catrope, gergesshamon: Backport for [[gerrit:1325878{{!}}[arwiki] Enable restricted user page editing and grant edit permissions (T434878)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:06 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1325878{{!}}[arwiki] Enable restricted user page editing and grant edit permissions (T434878)]] * 19:59 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm * 19:56 eevans@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cassandra-dev2001.codfw.wmnet with OS bookworm * 19:56 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm * 19:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2207.codfw.wmnet with reason: Maintenance * 18:47 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319920{{!}}Allow setting a separate thumbUrl in production (T427465)]], [[gerrit:1327167{{!}}Fix wmgThumbUrl config (T427465)]] (duration: 18m 53s) * 18:43 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 18:30 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1319920{{!}}Allow setting a separate thumbUrl in production (T427465)]], [[gerrit:1327167{{!}}Fix wmgThumbUrl config (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:28 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1319920{{!}}Allow setting a separate thumbUrl in production (T427465)]], [[gerrit:1327167{{!}}Fix wmgThumbUrl config (T427465)]] * 18:26 sukhe@dns1004: END - running authdns-update * 18:24 sukhe@dns1004: START - running authdns-update * 18:09 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1319920{{!}}Allow setting a separate thumbUrl in production (T427465)]] * 18:03 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-eqiad and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 17:56 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-codfw and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 17:50 cmooney@dns3003: END - running authdns-update * 17:42 dancy@deploy1003: Installation of scap version "4.283.0" completed for 3 hosts * 17:41 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling reboot on A:durum and A:durum * 17:40 cmooney@dns3003: START - running authdns-update * 17:40 dancy@deploy1003: Installing scap version "4.283.0" for 3 host(s) * 17:38 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:37 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on GTT VPLS - cmooney@cumin1003" * 17:37 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --verbose --scoreLessThan=0.7 --exceptDatasetChecksums=[[phab:T434319|T434319]]-enwiki-models.txt # [[phab:T434319|T434319]] * 17:32 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on GTT VPLS - cmooney@cumin1003" * 17:29 sbassett: Deployed security fix for [[phab:T435210|T435210]] (wmf.16) * 17:26 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 17:22 sbassett: Deployed security fix for [[phab:T435210|T435210]] (wmf.15) * 17:00 sukhe@dns1004: END - running authdns-update * 16:58 sukhe@dns1004: START - running authdns-update * 16:53 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica-esams and A:liberica * 16:41 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica-esams and A:liberica * 16:41 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-eqiad and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 16:41 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-codfw and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 16:41 cjd91: sudo -i cookbook sre.cdn.roll-upgrade-ats --query 'A:cp-codfw' --task-id [[phab:T434478|T434478]] --reason '9.2.15 upgrade' * 16:41 cjd91: sudo -i cookbook sre.cdn.roll-upgrade-ats --query 'A:cp-eqiad' --task-id [[phab:T434478|T434478]] --reason '9.2.15 upgrade' * 16:40 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and A:durum * 16:28 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326881{{!}}Echo: Start using virtual domains (T380385)]] (duration: 13m 12s) * 16:23 urbanecm@deploy1003: urbanecm: Continuing with deployment * 16:21 urandom: Completed sessionstore Cassandra/JVM upgrade — [[phab:T435154|T435154]] * 16:21 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching sessionstore[2005-2006].codfw.wmnet,sessionstore[1005-1006].eqiad.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 16:19 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1326881{{!}}Echo: Start using virtual domains (T380385)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:15 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:15 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->codfw - cmooney@cumin1003" * 16:14 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1326881{{!}}Echo: Start using virtual domains (T380385)]] * 16:14 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327123{{!}}Revert^2 "Migrate database access to virtual domains" (T435305)]], [[gerrit:1327124{{!}}Pass the mapped domain of virtual-echo-shared to the push NameTableStores (T435305)]] (duration: 07m 42s) * 16:13 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching sessionstore[2005-2006].codfw.wmnet,sessionstore[1005-1006].eqiad.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 16:11 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->codfw - cmooney@cumin1003" * 16:10 urbanecm@deploy1003: urbanecm: Continuing with deployment * 16:08 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1327123{{!}}Revert^2 "Migrate database access to virtual domains" (T435305)]], [[gerrit:1327124{{!}}Pass the mapped domain of virtual-echo-shared to the push NameTableStores (T435305)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:06 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 16:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 16:06 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1327123{{!}}Revert^2 "Migrate database access to virtual domains" (T435305)]], [[gerrit:1327124{{!}}Pass the mapped domain of virtual-echo-shared to the push NameTableStores (T435305)]] * 16:06 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:03 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching sessionstore1004.eqiad.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 16:01 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching sessionstore1004.eqiad.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 15:56 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching sessionstore2004.codfw.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 15:54 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching sessionstore2004.codfw.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 15:53 urandom: beginning sessionstore Cassandra/JVM upgrade — [[phab:T435154|T435154]] * 15:52 urandom: beginning sessionstore Cassandra/JVM upgrade — [[phab:T432944|T432944]] * 15:51 cmooney@dns3003: END - running authdns-update * 15:49 cmooney@dns3003: START - running authdns-update * 15:48 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:48 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->codfw - cmooney@cumin1003" * 15:45 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->codfw - cmooney@cumin1003" * 15:44 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 15:44 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:43 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 15:42 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:42 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:38 cmooney@dns3003: END - running authdns-update * 15:36 cmooney@dns3003: START - running authdns-update * 15:36 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:36 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->eqsin - cmooney@cumin1003" * 15:34 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327138{{!}}Enable redis lock manager everywhere (T366938)]] (duration: 08m 36s) * 15:33 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->eqsin - cmooney@cumin1003" * 15:30 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:29 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 15:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1147.eqiad.wmnet with OS bookworm * 15:28 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327138{{!}}Enable redis lock manager everywhere (T366938)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:25 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327138{{!}}Enable redis lock manager everywhere (T366938)]] * 15:24 jmm@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host krb1002.eqiad.wmnet * 15:19 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2207.codfw.wmnet with reason: Host crashed * 15:17 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327098{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]], [[gerrit:1327101{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]] (duration: 07m 13s) * 15:12 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 15:12 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1327098{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]], [[gerrit:1327101{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:10 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1327098{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]], [[gerrit:1327101{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]] * 15:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2010.codfw.wmnet with OS trixie * 15:05 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 15:05 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1147.eqiad.wmnet with reason: host reimage * 14:59 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 14:59 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:58 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1147.eqiad.wmnet with reason: host reimage * 14:55 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:55 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:49 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 14:48 cmooney@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host durum1001.eqiad.wmnet * 14:46 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:46 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Delete db2902 ipv6 addr - fceratto@cumin1003" * 14:46 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Delete db2902 ipv6 addr - fceratto@cumin1003" * 14:43 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1147.eqiad.wmnet with OS bookworm * 14:42 tgr@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327094{{!}}SpecialMWOAuthListConsumers: Handle newFromMWUser returning null in addNavigationSubtitle (T435167)]] (duration: 19m 25s) * 14:42 cmooney@cumin1003: START - Cookbook sre.hosts.reboot-single for host durum1001.eqiad.wmnet * 14:42 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 14:41 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 14:38 tgr@deploy1003: tgr: Continuing with deployment * 14:36 tgr@deploy1003: tgr: Backport for [[gerrit:1327094{{!}}SpecialMWOAuthListConsumers: Handle newFromMWUser returning null in addNavigationSubtitle (T435167)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:28 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 14:28 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:28 cmooney@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host durum3005.esams.wmnet * 14:25 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 14:23 cmooney@cumin1003: START - Cookbook sre.hosts.reboot-single for host durum3005.esams.wmnet * 14:23 tgr@deploy1003: Started scap sync-world: Backport for [[gerrit:1327094{{!}}SpecialMWOAuthListConsumers: Handle newFromMWUser returning null in addNavigationSubtitle (T435167)]] * 14:18 gengh@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:17 topranks: disable puppet on hosts running BIRD BGP to test merge of patch to systemd healtchcheck service * 14:17 gengh@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:17 gengh@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:16 gengh@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:16 elukey: upgrade spicerack on cumin1003 and cumin2003 * 14:16 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:15 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314962{{!}}static: add new dir bimi/ for BIMI SVG and PEM file (T311685)]] (duration: 10m 00s) * 14:15 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:11 kharlan@deploy1003: kharlan, sukhe: Continuing with deployment * 14:11 gengh@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:09 gengh@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:09 gengh@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:08 kharlan@deploy1003: kharlan, sukhe: Backport for [[gerrit:1314962{{!}}static: add new dir bimi/ for BIMI SVG and PEM file (T311685)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:06 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 14:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:06 gengh@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:05 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1314962{{!}}static: add new dir bimi/ for BIMI SVG and PEM file (T311685)]] * 14:05 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host krb1002.eqiad.wmnet * 14:05 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:04 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:04 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326342{{!}}Revert^2 "wmf-config/ProductionServices: set URL for urldownloader to service record"]] (duration: 07m 40s) * 13:59 kharlan@deploy1003: kharlan, sukhe: Continuing with deployment * 13:58 kharlan@deploy1003: kharlan, sukhe: Backport for [[gerrit:1326342{{!}}Revert^2 "wmf-config/ProductionServices: set URL for urldownloader to service record"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:57 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host stat1008.eqiad.wmnet with OS bookworm * 13:56 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1326342{{!}}Revert^2 "wmf-config/ProductionServices: set URL for urldownloader to service record"]] * 13:56 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 13:54 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325532{{!}}srwiki: Allow bureaucrats to add and remove event-organizer group (T434748)]] (duration: 14m 56s) * 13:54 swfrench@dns1004: END - running authdns-update * 13:52 swfrench@dns1004: START - running authdns-update * 13:51 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:51 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 13:48 kharlan@deploy1003: kharlan, danielyepezgarces: Continuing with deployment * 13:45 swfrench@cumin2003: conftool action : set/pooled=yes; selector: name=wikikube-worker2330.codfw.wmnet * 13:44 swfrench-wmf: finished etcd-main codfw -> eqiad switchover - [[phab:T435103|T435103]] * 13:44 kharlan@deploy1003: kharlan, danielyepezgarces: Backport for [[gerrit:1325532{{!}}srwiki: Allow bureaucrats to add and remove event-organizer group (T434748)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:44 swfrench@cumin2003: conftool action : set/pooled=no; selector: name=wikikube-worker2330.codfw.wmnet * 13:41 swfrench@dns1004: END - running authdns-update * 13:39 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1325532{{!}}srwiki: Allow bureaucrats to add and remove event-organizer group (T434748)]] * 13:39 swfrench@dns1004: START - running authdns-update * 13:37 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327111{{!}}Special:AbuseReview: Add "no further action needed" review action (T435020)]], [[gerrit:1327110{{!}}AbuseReview: Take the review verdict as a REST path parameter (T435020)]] (duration: 31m 43s) * 13:31 swfrench-wmf: starting etcd-main codfw -> eqiad switchover - [[phab:T435103|T435103]] * 13:28 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:28 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 13:25 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=97) rolling reboot on A:durum and A:durum * 13:24 kharlan@deploy1003: kharlan: Continuing with deployment * 13:24 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1147 * 13:24 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1147 * 13:23 kharlan@deploy1003: kharlan: Backport for [[gerrit:1327111{{!}}Special:AbuseReview: Add "no further action needed" review action (T435020)]], [[gerrit:1327110{{!}}AbuseReview: Take the review verdict as a REST path parameter (T435020)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:16 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:16 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 13:12 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:12 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 13:10 cdobbins@cumin1003: END (ERROR) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=97) Rolling upgrade of ATS on A:cp-codfw and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 13:10 cdobbins@cumin1003: END (ERROR) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=97) Rolling upgrade of ATS on A:cp-eqiad and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 13:06 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1327111{{!}}Special:AbuseReview: Add "no further action needed" review action (T435020)]], [[gerrit:1327110{{!}}AbuseReview: Take the review verdict as a REST path parameter (T435020)]] * 13:06 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-eqiad and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 13:05 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-codfw and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 13:05 cjd91: sudo -i cookbook sre.cdn.roll-upgrade-ats --query 'A:cp-eqiad' --task-id [[phab:T434478|T434478]] --reason '9.2.15 upgrade' * 13:03 swfrench@dns1004: END - running authdns-update * 13:01 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:01 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 13:01 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host ncmonitor1001.eqiad.wmnet * 13:00 swfrench@dns1004: START - running authdns-update * 12:59 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 12:59 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 12:59 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 12:59 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 12:57 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and A:durum * 12:53 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on db2902.codfw.wmnet with reason: Cloning * 12:48 cmooney@dns3003: END - running authdns-update * 12:48 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327096{{!}}Switch to redis lock manager on s4 and s8 (T366938)]] (duration: 09m 11s) * 12:46 cmooney@dns3003: START - running authdns-update * 12:45 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:45 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->eqord cct - cmooney@cumin1003" * 12:45 elukey: move the /v2/releng.* prefix on the Docker Registry to its new s3 backend - [[phab:T432829|T432829]] * 12:43 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 12:42 jelto@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 12:42 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->eqord cct - cmooney@cumin1003" * 12:42 jelto@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 12:41 jelto: update cert-manager to 1.19.6 on wikikube staging-codfw - [[phab:T427402|T427402]] * 12:40 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327096{{!}}Switch to redis lock manager on s4 and s8 (T366938)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:38 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327096{{!}}Switch to redis lock manager on s4 and s8 (T366938)]] * 12:38 blake@deploy1003: Finished scap sync-world: non-build deploy for [[phab:T417800|T417800]] (duration: 03m 52s) * 12:36 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 12:35 blake@deploy1003: Started scap sync-world: non-build deploy for [[phab:T417800|T417800]] * 12:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host krb2002.codfw.wmnet * 11:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host krb2002.codfw.wmnet * 11:49 moritzm: installing kerberos security updates * 11:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on stat1008.eqiad.wmnet with reason: host reimage * 11:44 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on stat1008.eqiad.wmnet with reason: host reimage * 11:31 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327084{{!}}Revert "Migrate database access to virtual domains" (T435305)]] (duration: 11m 02s) * 11:29 gkyziridis@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:29 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:29 kart_: Updated MinT to 2026-06-04-131507-production ([[phab:T321316|T321316]]) * 11:28 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/machinetranslation: apply * 11:28 gkyziridis@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:26 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 11:24 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:24 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:23 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:23 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:23 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/machinetranslation: apply * 11:22 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327084{{!}}Revert "Migrate database access to virtual domains" (T435305)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:21 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/machinetranslation: apply * 11:21 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:21 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:20 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327084{{!}}Revert "Migrate database access to virtual domains" (T435305)]] * 11:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1008.eqiad.wmnet with OS bookworm * 11:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet * 11:17 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/machinetranslation: apply * 11:13 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/machinetranslation: apply * 11:12 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:12 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:10 kartik@deploy1003: helmfile [staging] START helmfile.d/services/machinetranslation: apply * 11:08 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-cron: apply * 11:08 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/mw-cron: apply * 11:08 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply * 11:08 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply * 11:07 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet * 11:07 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host stat1008.eqiad.wmnet with OS bookworm * 11:06 moritzm: upgrading the new trixie URL downloaders to Squid 7.6 [[phab:T427282|T427282]] * 11:01 gkyziridis@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin2002.codfw.wmnet * 10:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin2002.codfw.wmnet * 10:45 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1169.eqiad.wmnet with OS bookworm * 10:42 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1185.eqiad.wmnet with OS bookworm * 10:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1169.eqiad.wmnet with reason: host reimage * 10:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1185.eqiad.wmnet with reason: host reimage * 10:14 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1169.eqiad.wmnet with reason: host reimage * 10:14 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1185.eqiad.wmnet with reason: host reimage * 10:11 jmm@cumin2003: END (PASS) - Cookbook sre.netbox.restart-reboot (exit_code=0) rolling reboot on A:netbox * 10:06 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1008.eqiad.wmnet with OS bookworm * 10:04 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 10:04 mpostoronca@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321579{{!}}Register the mediawiki.wikimedia_antiabuse.content_policy_score stream (T432848)]] (duration: 08m 53s) * 10:03 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 10:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 10:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 10:00 mpostoronca@deploy1003: mpostoronca: Continuing with deployment * 09:59 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1185.eqiad.wmnet with OS bookworm * 09:59 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1169.eqiad.wmnet with OS bookworm * 09:59 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.convert-disks (exit_code=0) for host ms-be1065 * 09:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1065.eqiad.wmnet with OS trixie * 09:58 mpostoronca@deploy1003: mpostoronca: Backport for [[gerrit:1321579{{!}}Register the mediawiki.wikimedia_antiabuse.content_policy_score stream (T432848)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:55 jmm@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netbox.discovery.wmnet. on all recursors * 09:55 mpostoronca@deploy1003: Started scap sync-world: Backport for [[gerrit:1321579{{!}}Register the mediawiki.wikimedia_antiabuse.content_policy_score stream (T432848)]] * 09:55 jmm@cumin2003: START - Cookbook sre.dns.wipe-cache netbox.discovery.wmnet. on all recursors * 09:52 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw2001.wikimedia.org with OS trixie * 09:51 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 09:51 jmm@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netbox.discovery.wmnet. on all recursors * 09:51 jmm@cumin2003: START - Cookbook sre.dns.wipe-cache netbox.discovery.wmnet. on all recursors * 09:51 jmm@cumin2003: START - Cookbook sre.netbox.restart-reboot rolling reboot on A:netbox * 09:46 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 09:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1056.eqiad.wmnet with OS trixie * 09:44 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 09:39 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 09:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 09:36 topranks: make HE transport circuits from magru live * 09:36 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 09:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb1003.eqiad.wmnet * 09:33 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage * 09:31 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb1003.eqiad.wmnet * 09:30 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 09:28 arnaudb@dns1006: END - running authdns-update * 09:27 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage * 09:27 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb2003.codfw.wmnet * 09:26 arnaudb@dns1006: START - running authdns-update * 09:24 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1056.eqiad.wmnet with reason: host reimage * 09:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb2003.codfw.wmnet * 09:20 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.convert-disks (exit_code=0) for host ms-be1068 * 09:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1068.eqiad.wmnet with OS trixie * 09:20 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 09:19 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "cloudvirt1057 - filippo@cumin1003" * 09:19 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "cloudvirt1057 - filippo@cumin1003" * 09:18 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1056.eqiad.wmnet with reason: host reimage * 09:18 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1057.eqiad.wmnet with OS trixie * 09:18 filippo@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 09:18 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 09:15 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 09:14 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 09:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host irc1003.wikimedia.org * 09:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:13 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1065.eqiad.wmnet with OS trixie * 09:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 09:09 moritzm: installing Postgresql security updates * 09:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host irc1003.wikimedia.org * 09:07 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 09:07 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:07 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1173.eqiad.wmnet with OS bookworm * 09:06 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw2001.wikimedia.org with OS trixie * 09:03 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.convert-disks (exit_code=0) for host ms-be1064 * 09:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1064.eqiad.wmnet with OS trixie * 09:03 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1056.eqiad.wmnet with OS trixie * 09:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1057.eqiad.wmnet with reason: host reimage * 09:01 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 08:58 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 08:56 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1057.eqiad.wmnet with reason: host reimage * 08:54 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1208.eqiad.wmnet with OS bookworm * 08:53 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2902.codfw.wmnet with OS trixie * 08:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 08:50 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1174.eqiad.wmnet with OS bookworm * 08:46 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1172.eqiad.wmnet with OS bookworm * 08:45 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1173.eqiad.wmnet with reason: host reimage * 08:41 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 08:40 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1057.eqiad.wmnet with OS trixie * 08:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1057.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 08:39 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1207.eqiad.wmnet with OS bookworm * 08:38 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2902.codfw.wmnet with reason: host reimage * 08:37 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 08:34 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1068.eqiad.wmnet with OS trixie * 08:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1208.eqiad.wmnet with reason: host reimage * 08:31 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1057.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 08:29 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 08:28 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1222.eqiad.wmnet onto db1276.eqiad.wmnet * 08:28 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1222: Pool db1222.eqiad.wmnet in after cloning * 08:28 fceratto@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2902.codfw.wmnet with reason: host reimage * 08:27 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1055.eqiad.wmnet with OS trixie * 08:27 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 08:26 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1174.eqiad.wmnet with reason: host reimage * 08:25 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 08:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-misc2002.codfw.wmnet * 08:23 topranks: reboot pfw1-codfw firewall pair to upgrade JunOS [[phab:T434865|T434865]] * 08:22 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1172.eqiad.wmnet with reason: host reimage * 08:20 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1064.eqiad.wmnet with OS trixie * 08:20 mvernon@cumin2003: START - Cookbook sre.swift.convert-disks for host ms-be1065 * 08:18 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1207.eqiad.wmnet with reason: host reimage * 08:17 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1174.eqiad.wmnet with reason: host reimage * 08:17 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1173.eqiad.wmnet with reason: host reimage * 08:17 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1172.eqiad.wmnet with reason: host reimage * 08:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host mc-misc2002.codfw.wmnet * 08:15 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1208.eqiad.wmnet with reason: host reimage * 08:15 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1207.eqiad.wmnet with reason: host reimage * 08:14 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db2902.codfw.wmnet with OS trixie * 08:14 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.16 refs [[phab:T430835|T430835]] * 08:13 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db2902.codfw.wmnet * 08:13 fceratto@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host db2902.codfw.wmnet with OS trixie * 08:10 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr[1-2]-codfw with reason: upgrade pfw1a-codfw and pfw1b-codfw pair * 08:09 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1055.eqiad.wmnet with reason: host reimage * 08:07 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on pfw1-codfw with reason: upgrade pfw1a-codfw and pfw1b-codfw pair * 08:03 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1055.eqiad.wmnet with reason: host reimage * 08:02 arnaudb@dns1006: END - running authdns-update * 08:02 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1208.eqiad.wmnet with OS bookworm * 08:02 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1207.eqiad.wmnet with OS bookworm * 08:01 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1174.eqiad.wmnet with OS bookworm * 08:01 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1173.eqiad.wmnet with OS bookworm * 08:01 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1172.eqiad.wmnet with OS bookworm * 07:59 arnaudb@dns1006: START - running authdns-update * 07:58 arnaudb@dns1006: START - running authdns-update * 07:48 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1055.eqiad.wmnet with OS trixie * 07:45 moritzm: extend the disk of ldap-rw2001 by 80G [[phab:T331699|T331699]] * 07:42 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1222: Pool db1222.eqiad.wmnet in after cloning * 07:36 mvernon@cumin2003: START - Cookbook sre.swift.convert-disks for host ms-be1068 * 07:35 mvernon@cumin2003: START - Cookbook sre.swift.convert-disks for host ms-be1064 * 07:17 moritzm: installing imagemagick security updates * 07:14 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1277: Pool back * 07:14 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon1003.wikimedia.org * 07:07 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon1003.wikimedia.org * 07:03 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1280: Pool back * 07:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon2002.wikimedia.org * 06:55 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon2002.wikimedia.org * 06:54 moritzm: installing php8.2 security updates * 06:51 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1284: Pool back * 06:49 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1222: Depool db1222.eqiad.wmnet to then clone it to db1276.eqiad.wmnet - marostegui@cumin1003 * 06:49 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1222: Depool db1222.eqiad.wmnet to then clone it to db1276.eqiad.wmnet - marostegui@cumin1003 * 06:49 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1222.eqiad.wmnet onto db1276.eqiad.wmnet * 06:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd1005.eqiad.wmnet * 06:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd1005.eqiad.wmnet * 06:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd1004.eqiad.wmnet * 06:36 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2209: db2209 repool * 06:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd1004.eqiad.wmnet * 06:32 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast3007.wikimedia.org * 06:29 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1277: Pool back * 06:28 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1277 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96190 and previous config saved to /var/cache/conftool/dbconfig/20260819-062815-marostegui.json * 06:26 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast3007.wikimedia.org * 06:22 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul1001.eqiad.wmnet * 06:18 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul1001.eqiad.wmnet * 06:18 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul1003.eqiad.wmnet * 06:18 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1280: Pool back * 06:17 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1284 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96186 and previous config saved to /var/cache/conftool/dbconfig/20260819-061743-marostegui.json * 06:14 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul1003.eqiad.wmnet * 06:14 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul1002.eqiad.wmnet * 06:10 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul1002.eqiad.wmnet * 06:10 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2048.codfw.wmnet * 06:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2048.codfw.wmnet * 06:06 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1284: Pool back * 06:06 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1284 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96184 and previous config saved to /var/cache/conftool/dbconfig/20260819-060621-marostegui.json * 06:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2048.codfw.wmnet * 05:59 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2048.codfw.wmnet * 05:51 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2209: db2209 repool * 03:16 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1186.eqiad.wmnet with OS bookworm * 02:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1186.eqiad.wmnet with reason: host reimage * 02:46 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1186.eqiad.wmnet with reason: host reimage * 02:46 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2207 [[phab:T435270|T435270]]', diff saved to https://phabricator.wikimedia.org/P96181 and previous config saved to /var/cache/conftool/dbconfig/20260819-024627-marostegui.json * 02:44 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2204 to s2 primary [[phab:T435270|T435270]]', diff saved to https://phabricator.wikimedia.org/P96180 and previous config saved to /var/cache/conftool/dbconfig/20260819-024403-marostegui.json * 02:43 marostegui: Starting s2 codfw failover from db2207 to db2204 - [[phab:T435270|T435270]] * 02:39 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2204 with weight 0 [[phab:T435270|T435270]]', diff saved to https://phabricator.wikimedia.org/P96179 and previous config saved to /var/cache/conftool/dbconfig/20260819-023951-marostegui.json * 02:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s2 [[phab:T435270|T435270]] * 02:32 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1186.eqiad.wmnet with OS bookworm * 02:29 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-worker1186.eqiad.wmnet with OS bookworm * 02:18 denisse@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2207: Depooling replica * 02:18 denisse@cumin1003: START - Cookbook sre.mysql.depool depool db2207: Depooling replica * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 48s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-18 == * 23:55 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1264845{{!}}Remove unused/redundant wgMFNoindexPages=true setting (T255458)]] (duration: 09m 42s) * 23:51 krinkle@deploy1003: krinkle: Continuing with deployment * 23:48 krinkle@deploy1003: krinkle: Backport for [[gerrit:1264845{{!}}Remove unused/redundant wgMFNoindexPages=true setting (T255458)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:45 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1264845{{!}}Remove unused/redundant wgMFNoindexPages=true setting (T255458)]] * 23:38 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326944{{!}}Retire filebackend lock manager in favour of the default one (T366938)]] (duration: 08m 55s) * 23:34 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 23:31 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326944{{!}}Retire filebackend lock manager in favour of the default one (T366938)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:29 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326944{{!}}Retire filebackend lock manager in favour of the default one (T366938)]] * 23:27 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1170.eqiad.wmnet with OS bookworm * 23:21 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1205.eqiad.wmnet with OS bookworm * 23:20 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1171.eqiad.wmnet with OS bookworm * 23:15 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1206.eqiad.wmnet with OS bookworm * 23:05 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1170.eqiad.wmnet with reason: host reimage * 23:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1205.eqiad.wmnet with reason: host reimage * 22:57 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1171.eqiad.wmnet with reason: host reimage * 22:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1206.eqiad.wmnet with reason: host reimage * 22:53 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1205.eqiad.wmnet with reason: host reimage * 22:51 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1171.eqiad.wmnet with reason: host reimage * 22:51 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1170.eqiad.wmnet with reason: host reimage * 22:50 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1206.eqiad.wmnet with reason: host reimage * 22:36 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1206.eqiad.wmnet with OS bookworm * 22:35 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1205.eqiad.wmnet with OS bookworm * 22:35 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1186.eqiad.wmnet with OS bookworm * 22:35 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1171.eqiad.wmnet with OS bookworm * 22:35 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1170.eqiad.wmnet with OS bookworm * 22:33 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-worker1194.eqiad.wmnet with OS bookworm * 22:22 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326923{{!}}Enable redis lock manager on s6 (T366938)]] (duration: 11m 52s) * 22:18 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 22:13 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326923{{!}}Enable redis lock manager on s6 (T366938)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:10 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326923{{!}}Enable redis lock manager on s6 (T366938)]] * 22:04 sbassett: Deployed security fix for [[phab:T435234|T435234]] (wmf.16) * 21:54 sbassett: Deployed security fix for [[phab:T435234|T435234]] (wmf.15) * 21:38 caro@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326925{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326926{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326929{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]], [[gerrit:1326928{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]] (duration: 0 * 21:34 caro@deploy1003: caro: Continuing with deployment * 21:33 caro@deploy1003: caro: Backport for [[gerrit:1326925{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326926{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326929{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]], [[gerrit:1326928{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]] synced to the testservers (see h * 21:31 caro@deploy1003: Started scap sync-world: Backport for [[gerrit:1326925{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326926{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326929{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]], [[gerrit:1326928{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]] * 21:24 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1204.eqiad.wmnet with reason: 1204 datanode repair [[phab:T434494|T434494]] * 21:02 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326896{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]], [[gerrit:1326897{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]] (duration: 13m 56s) * 20:58 krinkle@deploy1003: krinkle: Continuing with deployment * 20:50 krinkle@deploy1003: krinkle: Backport for [[gerrit:1326896{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]], [[gerrit:1326897{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:49 ryankemper: `an-launcher1003` terminated process group `666809` (`rest_backfill_phase1.sh`) ~20 mins ago with `sudo kill -TERM -- -666809` after its local spark driver (`--driver-memory 64g`) repeatedly exhausted memory on the 32 GB VM and caused SSH to intermittently flap; host recovered to 27 GB available memory * 20:48 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1326896{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]], [[gerrit:1326897{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]] * 20:35 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 20:33 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 20:31 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 20:31 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326870{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]], [[gerrit:1326871{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]] (duration: 07m 35s) * 20:28 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 20:26 kemayo@deploy1003: kemayo: Continuing with deployment * 20:26 ryankemper: `an-launcher1003` confirmed the host is flapping because of memory thrash. chasing down the source of the thrash * 20:25 kemayo@deploy1003: kemayo: Backport for [[gerrit:1326870{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]], [[gerrit:1326871{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:23 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1326870{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]], [[gerrit:1326871{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]] * 20:23 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 20:20 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 20:14 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1168.eqiad.wmnet with OS bookworm * 20:08 zabe: zabe@deploy1003:~$ mwscript extensions/WikimediaMaintenance/maintenance/fixFileRevisionArchiveNameDrift.php enwiki # [[phab:T428406|T428406]] * 20:08 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1204.eqiad.wmnet with OS bookworm * 20:05 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1167.eqiad.wmnet with OS bookworm * 20:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1166.eqiad.wmnet with OS bookworm * 19:54 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1203.eqiad.wmnet with OS bookworm * 19:53 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326907{{!}}Revert "Disable redis lock manager on testwiki"]] (duration: 11m 05s) * 19:50 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1168.eqiad.wmnet with reason: host reimage * 19:47 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1204.eqiad.wmnet with reason: host reimage * 19:46 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 19:44 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326907{{!}}Revert "Disable redis lock manager on testwiki"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:42 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326907{{!}}Revert "Disable redis lock manager on testwiki"]] * 19:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1167.eqiad.wmnet with reason: host reimage * 19:37 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1166.eqiad.wmnet with reason: host reimage * 19:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1203.eqiad.wmnet with reason: host reimage * 19:32 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1167.eqiad.wmnet with reason: host reimage * 19:32 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1168.eqiad.wmnet with reason: host reimage * 19:32 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1166.eqiad.wmnet with reason: host reimage * 19:31 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1204.eqiad.wmnet with reason: host reimage * 19:31 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1203.eqiad.wmnet with reason: host reimage * 19:26 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324751{{!}}InitialiseSettings: Enable 2FA enforcement on more private wikis (T428103)]], [[gerrit:1326875{{!}}Add banner notifying of upcoming 2FA enforcement (T420792)]] (duration: 31m 46s) * 19:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1204.eqiad.wmnet with OS bookworm * 19:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1203.eqiad.wmnet with OS bookworm * 19:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1168.eqiad.wmnet with OS bookworm * 19:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1167.eqiad.wmnet with OS bookworm * 19:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1166.eqiad.wmnet with OS bookworm * 19:15 denisse: rebooting kafkamon2003.codfw.wmnet - [[phab:T435162|T435162]] * 19:14 denisse: rebooting kafkamon1003.eqiad.wmnet [[phab:T435162|T435162]] * 19:13 reedy@deploy1003: reedy: Continuing with deployment * 19:12 reedy@deploy1003: reedy: Backport for [[gerrit:1324751{{!}}InitialiseSettings: Enable 2FA enforcement on more private wikis (T428103)]], [[gerrit:1326875{{!}}Add banner notifying of upcoming 2FA enforcement (T420792)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:54 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324751{{!}}InitialiseSettings: Enable 2FA enforcement on more private wikis (T428103)]], [[gerrit:1326875{{!}}Add banner notifying of upcoming 2FA enforcement (T420792)]] * 18:50 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 18:44 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.16 refs [[phab:T430835|T430835]] * 18:34 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 18:31 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 18:22 aklapper@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326852{{!}}CategoryTree: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]], [[gerrit:1326853{{!}}CategoryViewer: Allow null $html in the CategoryViewerGenerateLink hook (T435161)]], [[gerrit:1326865{{!}}Flow: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]] (duration: 09m 57s) * 18:18 aklapper@deploy1003: jforrester, aklapper: Continuing with deployment * 18:17 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 18:14 aklapper@deploy1003: jforrester, aklapper: Backport for [[gerrit:1326852{{!}}CategoryTree: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]], [[gerrit:1326853{{!}}CategoryViewer: Allow null $html in the CategoryViewerGenerateLink hook (T435161)]], [[gerrit:1326865{{!}}Flow: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki * 18:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1156.eqiad.wmnet with OS bookworm * 18:12 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 18:12 aklapper@deploy1003: Started scap sync-world: Backport for [[gerrit:1326852{{!}}CategoryTree: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]], [[gerrit:1326853{{!}}CategoryViewer: Allow null $html in the CategoryViewerGenerateLink hook (T435161)]], [[gerrit:1326865{{!}}Flow: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]] * 18:11 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 18:08 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1146.eqiad.wmnet with OS bookworm * 18:07 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1177.eqiad.wmnet with OS bookworm * 18:00 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326380{{!}}Introduce main lock manager service (T366938 T427999)]] (duration: 11m 25s) * 17:58 ladsgroup@deploy1003: ladsgroup: Rolling back deployment * 17:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1202.eqiad.wmnet with OS bookworm * 17:55 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1201.eqiad.wmnet with OS bookworm * 17:50 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326380{{!}}Introduce main lock manager service (T366938 T427999)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:48 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326380{{!}}Introduce main lock manager service (T366938 T427999)]] * 17:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1156.eqiad.wmnet with reason: host reimage * 17:46 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1146.eqiad.wmnet with reason: host reimage * 17:45 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-esams and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 17:42 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 17:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1177.eqiad.wmnet with reason: host reimage * 17:38 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1202.eqiad.wmnet with reason: host reimage * 17:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1201.eqiad.wmnet with reason: host reimage * 17:30 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1156.eqiad.wmnet with reason: host reimage * 17:29 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1177.eqiad.wmnet with reason: host reimage * 17:28 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1146.eqiad.wmnet with reason: host reimage * 17:28 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1202.eqiad.wmnet with reason: host reimage * 17:27 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1201.eqiad.wmnet with reason: host reimage * 17:25 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 17:21 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 17:14 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1202.eqiad.wmnet with OS bookworm * 17:14 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1201.eqiad.wmnet with OS bookworm * 17:14 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1177.eqiad.wmnet with OS bookworm * 17:14 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1156.eqiad.wmnet with OS bookworm * 17:14 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1146.eqiad.wmnet with OS bookworm * 17:09 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326885{{!}}w/deployment-info.php: Handle new file format (T434726)]] (duration: 07m 15s) * 17:08 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 17:05 dancy@deploy1003: dancy: Continuing with deployment * 17:04 dancy@deploy1003: dancy: Backport for [[gerrit:1326885{{!}}w/deployment-info.php: Handle new file format (T434726)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:02 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1326885{{!}}w/deployment-info.php: Handle new file format (T434726)]] * 16:50 dancy@deploy1003: Finished scap sync-world: Testing [[phab:T434726|T434726]] (duration: 06m 40s) * 16:43 dancy@deploy1003: Started scap sync-world: Testing [[phab:T434726|T434726]] * 16:43 dancy@deploy1003: Installation of scap version "4.282.0" completed for 3 hosts * 16:41 dancy@deploy1003: Installing scap version "4.282.0" for 3 host(s) * 16:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1145.eqiad.wmnet with OS bookworm * 16:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1200.eqiad.wmnet with OS bookworm * 16:32 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1199.eqiad.wmnet with OS bookworm * 16:18 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1145.eqiad.wmnet with reason: host reimage * 16:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1200.eqiad.wmnet with reason: host reimage * 16:09 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1199.eqiad.wmnet with reason: host reimage * 16:05 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-esams and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 16:05 cjd91: sudo -i cookbook sre.cdn.roll-upgrade-ats --query 'A:cp-esams' --task-id [[phab:T434478|T434478]] --reason '9.2.15 upgrade' * 16:03 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1200.eqiad.wmnet with reason: host reimage * 16:02 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1145.eqiad.wmnet with reason: host reimage * 16:02 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1199.eqiad.wmnet with reason: host reimage * 15:50 moritzm: installing zip security updates * 15:48 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1200.eqiad.wmnet with OS bookworm * 15:47 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1199.eqiad.wmnet with OS bookworm * 15:47 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1145.eqiad.wmnet with OS bookworm * 15:41 topranks: bounce PIC 0/0 on cr1-magru to set port to 40G * 15:37 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1236.eqiad.wmnet with OS bookworm * 15:29 aikochou@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'ores-legacy' for release 'main' . * 15:26 aikochou@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'ores-legacy' for release 'main' . * 15:20 aikochou@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'ores-legacy' for release 'main' . * 15:14 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2002.codfw.wmnet * 15:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1236.eqiad.wmnet with reason: host reimage * 15:12 moritzm: failover ganeti master in codfw to ganeti2047 * 15:09 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1236.eqiad.wmnet with reason: host reimage * 15:09 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2044.codfw.wmnet * 15:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2002.codfw.wmnet * 15:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2044.codfw.wmnet * 15:04 brennen@deploy1003: Finished deploy [phabricator/deployment@6b9b6ff]: deploy phab1004 for [[phab:T435213|T435213]] (duration: 01m 01s) * 15:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2044.codfw.wmnet * 15:03 brennen@deploy1003: Started deploy [phabricator/deployment@6b9b6ff]: deploy phab1004 for [[phab:T435213|T435213]] * 15:03 brennen@deploy1003: Finished deploy [phabricator/deployment@6b9b6ff]: deploy phab2003 for [[phab:T435213|T435213]] (duration: 00m 57s) * 15:02 brennen@deploy1003: Started deploy [phabricator/deployment@6b9b6ff]: deploy phab2003 for [[phab:T435213|T435213]] * 14:57 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply * 14:57 arnaudb@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on phab2003.codfw.wmnet,phab[1004-1006].eqiad.wmnet with reason: maintenance * 14:56 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2044.codfw.wmnet * 14:55 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply * 14:53 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1236.eqiad.wmnet with OS bookworm * 14:51 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2043.codfw.wmnet * 14:51 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2043.codfw.wmnet * 14:45 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2043.codfw.wmnet * 14:33 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2043.codfw.wmnet * 14:22 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2042.codfw.wmnet * 14:22 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2042.codfw.wmnet * 14:21 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db2902.codfw.wmnet with OS trixie * 14:16 elukey: uploaded spicerack_13.2.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia * 14:16 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2042.codfw.wmnet * 14:04 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2042.codfw.wmnet * 14:04 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2041.codfw.wmnet * 14:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2041.codfw.wmnet * 13:59 phuedx@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: apply * 13:59 phuedx@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-main: apply * 13:59 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326838{{!}}Enable Suggested Investigations on hewiki (T435146)]] (duration: 11m 50s) * 13:58 phuedx@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: apply * 13:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2041.codfw.wmnet * 13:57 phuedx@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-main: apply * 13:57 phuedx@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-main: apply * 13:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard1003.eqiad.wmnet * 13:57 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-main: apply * 13:55 phuedx@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-logging-external: apply * 13:55 phuedx@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-logging-external: apply * 13:54 phuedx@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-logging-external: apply * 13:54 stran@deploy1003: stran: Continuing with deployment * 13:54 phuedx@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-logging-external: apply * 13:54 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-logging-external: apply * 13:54 phuedx@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-logging-external: apply * 13:53 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-logging-external: apply * 13:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard1003.eqiad.wmnet * 13:53 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2041.codfw.wmnet * 13:51 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db2902.codfw.wmnet - fceratto@cumin1003" * 13:51 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db2902.codfw.wmnet - fceratto@cumin1003" * 13:51 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2026.codfw.wmnet * 13:51 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard2003.codfw.wmnet * 13:50 phuedx@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: apply * 13:50 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-eqiad * 13:50 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp1001.eqiad.wmnet * 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp1001.eqiad.wmnet * 13:50 phuedx@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: apply * 13:49 phuedx@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: apply * 13:49 stran@deploy1003: stran: Backport for [[gerrit:1326838{{!}}Enable Suggested Investigations on hewiki (T435146)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:48 phuedx@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: apply * 13:48 phuedx@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics: apply * 13:47 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard2003.codfw.wmnet * 13:47 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics: apply * 13:47 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1326838{{!}}Enable Suggested Investigations on hewiki (T435146)]] * 13:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp1001.eqiad.wmnet * 13:44 cdanis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 13:43 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp1001.eqiad.wmnet * 13:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1315-1327].eqiad.wmnet * 13:43 cdanis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 13:43 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1315-1327].eqiad.wmnet * 13:40 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki1001.eqiad.wmnet * 13:35 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1315-1327].eqiad.wmnet * 13:34 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host rpki1001.eqiad.wmnet * 13:27 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1315-1327].eqiad.wmnet * 13:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1302-1314].eqiad.wmnet * 13:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1302-1314].eqiad.wmnet * 13:21 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326775{{!}}SI: Instrument case update on first edit (T435048)]] (duration: 07m 12s) * 13:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1302-1314].eqiad.wmnet * 13:17 stran@deploy1003: stran: Continuing with deployment * 13:16 stran@deploy1003: stran: Backport for [[gerrit:1326775{{!}}SI: Instrument case update on first edit (T435048)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:14 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1326775{{!}}SI: Instrument case update on first edit (T435048)]] * 13:13 moritzm: installing util-linux security updates * 13:11 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1302-1314].eqiad.wmnet * 13:10 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326232{{!}}prv: Enable parsoid rendering for 5 wikis (T435115)]] (duration: 08m 17s) * 13:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1288-1289,1291-1301].eqiad.wmnet * 13:10 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply * 13:10 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1288-1289,1291-1301].eqiad.wmnet * 13:10 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply * 13:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki2003.codfw.wmnet * 13:06 jgiannelos@deploy1003: jgiannelos: Continuing with deployment * 13:05 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host rpki2003.codfw.wmnet * 13:04 jgiannelos@deploy1003: jgiannelos: Backport for [[gerrit:1326232{{!}}prv: Enable parsoid rendering for 5 wikis (T435115)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1288-1289,1291-1301].eqiad.wmnet * 13:02 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1326232{{!}}prv: Enable parsoid rendering for 5 wikis (T435115)]] * 12:53 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1288-1289,1291-1301].eqiad.wmnet * 12:53 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1273,1275-1287].eqiad.wmnet * 12:53 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1273,1275-1287].eqiad.wmnet * 12:52 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt1002.wikimedia.org * 12:51 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db2902.codfw.wmnet on all recursors * 12:51 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db2902.codfw.wmnet on all recursors * 12:51 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:51 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db2902.codfw.wmnet - fceratto@cumin1003" * 12:51 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db2902.codfw.wmnet - fceratto@cumin1003" * 12:46 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt1002.wikimedia.org * 12:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt2002.wikimedia.org * 12:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1273,1275-1287].eqiad.wmnet * 12:42 dhinus: repooled clouddb1032 that was currently <nowiki>{</nowiki>"weight": 0, "pooled": "inactive"<nowiki>}</nowiki> for both s4 and s6 * 12:41 dhinus: also depooled clouddb1017 (forgot it in the previous list) * 12:41 fnegri@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet * 12:40 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 12:40 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db2902.codfw.wmnet * 12:40 fnegri@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032.eqiad.wmnet * 12:40 fnegri@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032 * 12:39 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1017.eqiad.wmnet * 12:39 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt2002.wikimedia.org * 12:38 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1020.eqiad.wmnet * 12:38 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1018.eqiad.wmnet * 12:38 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet * 12:37 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1014.eqiad.wmnet * 12:37 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1013.eqiad.wmnet * 12:37 dhinus: depool again clouddb10[13,14,16,18,20] that were repooled by the cookbook sre.mysql.multiinstance_reboot * 12:36 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1273,1275-1287].eqiad.wmnet * 12:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1248-1261].eqiad.wmnet * 12:36 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1248-1261].eqiad.wmnet * 12:35 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2026.codfw.wmnet * 12:34 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 12:34 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 12:34 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 12:34 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 12:29 lucaswerkmeister-wmde@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 12:28 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2026.codfw.wmnet * 12:28 lucaswerkmeister-wmde@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 12:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1248-1261].eqiad.wmnet * 12:27 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.addnode (exit_code=0) for new host ganeti2046.codfw.wmnet to cluster codfw and group A * 12:26 moritzm: readded ganeti2046 to the codfw cluster following firmware update and reimage [[phab:T434681|T434681]] * 12:23 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply * 12:23 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply * 12:23 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply * 12:22 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply * 12:22 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply * 12:22 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply * 12:21 jmm@cumin2003: START - Cookbook sre.ganeti.addnode for new host ganeti2046.codfw.wmnet to cluster codfw and group A * 12:21 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1169.eqiad.wmnet onto db1283.eqiad.wmnet * 12:21 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1169: Pool db1169.eqiad.wmnet in after cloning * 12:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1248-1261].eqiad.wmnet * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1149-1153,1158,1240-1247].eqiad.wmnet * 12:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1149-1153,1158,1240-1247].eqiad.wmnet * 12:11 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1149-1153,1158,1240-1247].eqiad.wmnet * 12:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1168.eqiad.wmnet onto db1282.eqiad.wmnet * 12:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1168: Pool db1168.eqiad.wmnet in after cloning * 12:06 jmm@cumin2003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-test-eqiad * 12:05 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2026.codfw.wmnet * 12:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1149-1153,1158,1240-1247].eqiad.wmnet * 12:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1128-1134,1142-1148].eqiad.wmnet * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2046.codfw.wmnet * 12:01 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1128-1134,1142-1148].eqiad.wmnet * 11:57 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2025.codfw.wmnet * 11:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2025.codfw.wmnet * 11:54 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2046.codfw.wmnet * 11:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1128-1134,1142-1148].eqiad.wmnet * 11:50 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2025.codfw.wmnet * 11:46 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1128-1134,1142-1148].eqiad.wmnet * 11:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1114-1127].eqiad.wmnet * 11:45 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1114-1127].eqiad.wmnet * 11:39 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db2901.codfw.wmnet * 11:39 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db2901.codfw.wmnet * 11:36 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2025.codfw.wmnet * 11:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1114-1127].eqiad.wmnet * 11:35 fceratto@cumin1003: END (ERROR) - Cookbook sre.ganeti.makevm (exit_code=93) for new host db1901.eqiad.wmnet * 11:35 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1169: Pool db1169.eqiad.wmnet in after cloning * 11:35 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 11:30 jmm@cumin2003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-test-eqiad * 11:27 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1114-1127].eqiad.wmnet * 11:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1076-1081,1084-1087,1093-1095,1113].eqiad.wmnet * 11:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1076-1081,1084-1087,1093-1095,1113].eqiad.wmnet * 11:25 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1168: Pool db1168.eqiad.wmnet in after cloning * 11:24 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1194.eqiad.wmnet with OS bookworm * 11:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1181.eqiad.wmnet with OS bookworm * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2050.codfw.wmnet * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2050.codfw.wmnet * 11:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1076-1081,1084-1087,1093-1095,1113].eqiad.wmnet * 11:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2050.codfw.wmnet * 11:14 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2004.codfw.wmnet * 11:10 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2050.codfw.wmnet * 11:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1076-1081,1084-1087,1093-1095,1113].eqiad.wmnet * 11:09 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1286: Pool back * 11:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1045-1050,1056-1057,1064-1066,1073-1075].eqiad.wmnet * 11:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1045-1050,1056-1057,1064-1066,1073-1075].eqiad.wmnet * 11:08 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2004.codfw.wmnet * 11:07 moritzm: installing PHP 8.4 security updates * 11:06 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on an-worker1194.eqiad.wmnet with reason: host reimage * 11:06 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1194.eqiad.wmnet with reason: host reimage * 11:05 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2049.codfw.wmnet * 11:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2049.codfw.wmnet * 11:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid1003.eqiad.wmnet * 11:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1045-1050,1056-1057,1064-1066,1073-1075].eqiad.wmnet * 10:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1181.eqiad.wmnet with reason: host reimage * 10:59 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2049.codfw.wmnet * 10:59 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid1003.eqiad.wmnet * 10:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid2003.codfw.wmnet * 10:54 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1181.eqiad.wmnet with reason: host reimage * 10:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid2003.codfw.wmnet * 10:51 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1901.eqiad.wmnet on all recursors * 10:51 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1901.eqiad.wmnet on all recursors * 10:51 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:51 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:51 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:50 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2049.codfw.wmnet * 10:50 blake@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:50 blake@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:49 blake@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:48 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1045-1050,1056-1057,1064-1066,1073-1075].eqiad.wmnet * 10:48 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1044].eqiad.wmnet * 10:48 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1044].eqiad.wmnet * 10:47 blake@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:46 blake@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:46 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2047.codfw.wmnet * 10:46 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:46 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2047.codfw.wmnet * 10:46 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1901.eqiad.wmnet * 10:46 blake@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:44 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1168: Depool db1168.eqiad.wmnet to then clone it to db1282.eqiad.wmnet - marostegui@cumin1003 * 10:44 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1168: Depool db1168.eqiad.wmnet to then clone it to db1282.eqiad.wmnet - marostegui@cumin1003 * 10:44 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1168.eqiad.wmnet onto db1282.eqiad.wmnet * 10:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2047.codfw.wmnet * 10:40 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1044].eqiad.wmnet * 10:37 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2047.codfw.wmnet * 10:37 fceratto@cumin1003: END (ERROR) - Cookbook sre.ganeti.makevm (exit_code=93) for new host db1901.eqiad.wmnet * 10:36 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 10:34 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:33 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2032.codfw.wmnet * 10:33 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2032.codfw.wmnet * 10:32 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1044].eqiad.wmnet * 10:32 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-eqiad * 10:27 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2032.codfw.wmnet * 10:24 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1286: Pool back * 10:24 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1286 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96163 and previous config saved to /var/cache/conftool/dbconfig/20260818-102431-marostegui.json * 10:22 moritzm: installing Django security updates * 10:20 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul2002.codfw.wmnet * 10:20 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul2001.codfw.wmnet * 10:17 blake@deploy1003: Finished scap sync-world: no-build deployment for [[phab:T417800|T417800]] (duration: 04m 40s) * 10:16 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul2002.codfw.wmnet * 10:16 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul2001.codfw.wmnet * 10:15 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2032.codfw.wmnet * 10:14 blake@deploy1003: Started scap sync-world: no-build deployment for [[phab:T417800|T417800]] * 10:12 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul2003.codfw.wmnet * 10:12 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy2001.codfw.wmnet * 10:12 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy3001.esams.wmnet * 10:12 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy1001.eqiad.wmnet * 10:08 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2031.codfw.wmnet * 10:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2031.codfw.wmnet * 10:08 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul2003.codfw.wmnet * 10:08 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy2001.codfw.wmnet * 10:08 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy3001.esams.wmnet * 10:08 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy1001.eqiad.wmnet * 10:07 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy1002.eqiad.wmnet * 10:05 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy2002.codfw.wmnet * 10:05 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy3002.esams.wmnet * 10:04 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy1002.eqiad.wmnet * 10:03 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy4003.ulsfo.wmnet * 10:03 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy5003.eqsin.wmnet * 10:02 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2031.codfw.wmnet * 10:02 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1169: Depool db1169.eqiad.wmnet to then clone it to db1283.eqiad.wmnet - marostegui@cumin1003 * 10:01 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy2002.codfw.wmnet * 10:01 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy4004.ulsfo.wmnet * 10:01 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy3002.esams.wmnet * 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1169: Depool db1169.eqiad.wmnet to then clone it to db1283.eqiad.wmnet - marostegui@cumin1003 * 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1169.eqiad.wmnet onto db1283.eqiad.wmnet * 10:01 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy5004.eqsin.wmnet * 09:59 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy4003.ulsfo.wmnet * 09:59 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy4004.ulsfo.wmnet * 09:59 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy5003.eqsin.wmnet * 09:59 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy7001.magru.wmnet * 09:59 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy5004.eqsin.wmnet * 09:58 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy7002.magru.wmnet * 09:57 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2031.codfw.wmnet * 09:57 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy6001.drmrs.wmnet * 09:57 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy6002.drmrs.wmnet * 09:54 filippo@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for 10 hosts * 09:54 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast4006.wikimedia.org * 09:53 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy6001.drmrs.wmnet * 09:53 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy6002.drmrs.wmnet * 09:53 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host releases1003.eqiad.wmnet * 09:53 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy7001.magru.wmnet * 09:53 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host people2004.codfw.wmnet * 09:52 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host people1005.eqiad.wmnet * 09:52 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy7002.magru.wmnet * 09:50 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host releases2003.codfw.wmnet * 09:49 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host releases2003.codfw.wmnet * 09:49 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host releases1003.eqiad.wmnet * 09:49 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host people2004.codfw.wmnet * 09:48 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host people1005.eqiad.wmnet * 09:46 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2040.codfw.wmnet * 09:46 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2040.codfw.wmnet * 09:43 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1901.eqiad.wmnet on all recursors * 09:42 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1901.eqiad.wmnet on all recursors * 09:42 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:42 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:42 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2040.codfw.wmnet * 09:40 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp2005.wikimedia.org * 09:36 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp2005.wikimedia.org * 09:31 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:31 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1901.eqiad.wmnet * 09:31 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1901.eqiad.wmnet * 09:31 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:31 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1901.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 09:31 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1901.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 09:29 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2040.codfw.wmnet * 09:29 slyngshede@dns1004: END - running authdns-update * 09:28 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2039.codfw.wmnet * 09:28 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2039.codfw.wmnet * 09:27 slyngshede@dns1004: START - running authdns-update * 09:26 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test2005.wikimedia.org * 09:22 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2039.codfw.wmnet * 09:22 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test2005.wikimedia.org * 09:22 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp1005.wikimedia.org * 09:21 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:19 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2039.codfw.wmnet * 09:19 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2038.codfw.wmnet * 09:18 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1287: Pool back * 09:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2038.codfw.wmnet * 09:18 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp1005.wikimedia.org * 09:18 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test1005.wikimedia.org * 09:17 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1901.eqiad.wmnet * 09:15 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db1901.eqiad.wmnet * 09:15 fceratto@cumin1003: END (ERROR) - Cookbook sre.dns.netbox (exit_code=97) * 09:14 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test1005.wikimedia.org * 09:13 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2038.codfw.wmnet * 09:13 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:13 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1901.eqiad.wmnet * 09:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast4006.wikimedia.org * 09:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1288: Pool back * 09:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast5005.wikimedia.org * 09:03 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2038.codfw.wmnet * 08:57 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2037.codfw.wmnet * 08:57 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast5005.wikimedia.org * 08:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2037.codfw.wmnet * 08:56 filippo@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 10 hosts * 08:52 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2037.codfw.wmnet * 08:51 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1289: Pool back * 08:35 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudidp2001-dev.codfw.wmnet * 08:34 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1002-dev.eqiad.wmnet * 08:33 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1287: Pool back * 08:33 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1287 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96150 and previous config saved to /var/cache/conftool/dbconfig/20260818-083311-marostegui.json * 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1001-dev.eqiad.wmnet * 08:31 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudidp2001-dev.codfw.wmnet * 08:30 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1002-dev.eqiad.wmnet * 08:30 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2037.codfw.wmnet * 08:29 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1001-dev.eqiad.wmnet * 08:28 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2036.codfw.wmnet * 08:28 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2036.codfw.wmnet * 08:25 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1288: Pool back * 08:23 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 08:23 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1181.eqiad.wmnet with OS bookworm * 08:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2036.codfw.wmnet * 08:22 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1288 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96148 and previous config saved to /var/cache/conftool/dbconfig/20260818-082234-marostegui.json * 08:20 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2036.codfw.wmnet * 08:18 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2035.codfw.wmnet * 08:18 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudvirt1057.eqiad.wmnet with OS trixie * 08:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2035.codfw.wmnet * 08:13 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2035.codfw.wmnet * 08:06 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2035.codfw.wmnet * 08:05 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1289: Pool back * 08:05 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1289 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96145 and previous config saved to /var/cache/conftool/dbconfig/20260818-080531-marostegui.json * 07:54 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 07:51 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326439{{!}}Fix "mathjax_ignore" handling around forcemathmode attribute (T434686)]] (duration: 13m 15s) * 07:48 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 07:47 krinkle@deploy1003: krinkle: Continuing with deployment * 07:44 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ganeti2046.codfw.wmnet with OS bookworm * 07:40 krinkle@deploy1003: krinkle: Backport for [[gerrit:1326439{{!}}Fix "mathjax_ignore" handling around forcemathmode attribute (T434686)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:39 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 07:38 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1326439{{!}}Fix "mathjax_ignore" handling around forcemathmode attribute (T434686)]] * 07:34 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 07:31 samwilson@deploy1003: Finished scap sync-world: Backport for [[gerrit:701016{{!}}InitialiseSettings and -labs: Remove redundant feature flag $wgWikisourceEnableOcr (T285311)]] (duration: 07m 47s) * 07:30 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1048.eqiad.wmnet * 07:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1048.eqiad.wmnet * 07:29 XioNoX: add gnmic 0.47.0 to bookworm and trixie reprepro * 07:28 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ganeti2046.codfw.wmnet with reason: host reimage * 07:27 samwilson@deploy1003: samwilson: Continuing with deployment * 07:25 samwilson@deploy1003: samwilson: Backport for [[gerrit:701016{{!}}InitialiseSettings and -labs: Remove redundant feature flag $wgWikisourceEnableOcr (T285311)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:25 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 07:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1048.eqiad.wmnet * 07:24 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ganeti2046.codfw.wmnet with reason: host reimage * 07:23 samwilson@deploy1003: Started scap sync-world: Backport for [[gerrit:701016{{!}}InitialiseSettings and -labs: Remove redundant feature flag $wgWikisourceEnableOcr (T285311)]] * 07:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1198.eqiad.wmnet with OS bookworm * 07:18 samwilson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326450{{!}}InitialiseSettings.php: Enable Bulk OCR on pawikisource (T434648)]] (duration: 12m 17s) * 07:17 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 07:11 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1048.eqiad.wmnet * 07:11 samwilson@deploy1003: samwilson: Continuing with deployment * 07:11 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ganeti2046.codfw.wmnet with OS bookworm * 07:10 samwilson@deploy1003: samwilson: Backport for [[gerrit:1326450{{!}}InitialiseSettings.php: Enable Bulk OCR on pawikisource (T434648)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1057.eqiad.wmnet with OS trixie * 07:05 samwilson@deploy1003: Started scap sync-world: Backport for [[gerrit:1326450{{!}}InitialiseSettings.php: Enable Bulk OCR on pawikisource (T434648)]] * 07:02 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1198.eqiad.wmnet with reason: host reimage * 07:01 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1056.eqiad.wmnet with OS trixie * 07:01 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1056.eqiad.wmnet with OS trixie * 07:00 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1056.eqiad.wmnet with OS trixie * 07:00 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1056.eqiad.wmnet with OS trixie * 06:59 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudvirt1055.eqiad.wmnet with OS trixie * 06:58 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1198.eqiad.wmnet with reason: host reimage * 06:53 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1055.eqiad.wmnet with OS trixie * 06:52 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudvirt1054.eqiad.wmnet with OS trixie * 06:44 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1194.eqiad.wmnet with OS bookworm * 06:44 XioNoX: upgrade eqsin gnmic to 0.47.0 * 06:43 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1198.eqiad.wmnet with OS bookworm * 06:41 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1054.eqiad.wmnet with OS trixie * 06:09 arnaudb@cumin1003: END (PASS) - Cookbook sre.gerrit.restart-gerrit (exit_code=0) Restarting Gerrit on gerrit2002 * 06:06 arnaudb@cumin1003: START - Cookbook sre.gerrit.restart-gerrit Restarting Gerrit on gerrit2002 * 06:06 arnaudb@cumin1003: END (PASS) - Cookbook sre.gerrit.restart-gerrit (exit_code=0) Restarting Gerrit on gerrit1003 * 06:04 arnaudb@cumin1003: START - Cookbook sre.gerrit.restart-gerrit Restarting Gerrit on gerrit1003 * 06:02 arnaudb@cumin1003: END (PASS) - Cookbook sre.gerrit.restart-gerrit (exit_code=0) Restarting Gerrit on gerrit2003 * 06:00 arnaudb@cumin1003: START - Cookbook sre.gerrit.restart-gerrit Restarting Gerrit on gerrit2003 * 05:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1155.eqiad.wmnet with OS bookworm * 05:38 arnaudb: updating prometheusBearerToken on gerrit * 05:28 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1155.eqiad.wmnet with reason: host reimage * 05:23 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1155.eqiad.wmnet with reason: host reimage * 05:06 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1155.eqiad.wmnet with OS bookworm * 04:57 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1144.eqiad.wmnet with OS bookworm * 04:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1144.eqiad.wmnet with reason: host reimage * 04:29 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1144.eqiad.wmnet with reason: host reimage * 04:14 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1144.eqiad.wmnet with OS bookworm * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.13 (duration: 02m 23s) * 03:45 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1197.eqiad.wmnet with OS bookworm * 03:41 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1165.eqiad.wmnet with OS bookworm * 03:38 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.16 refs [[phab:T430835|T430835]] (duration: 34m 43s) * 03:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1164.eqiad.wmnet with OS bookworm * 03:35 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1196.eqiad.wmnet with OS bookworm * 03:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1163.eqiad.wmnet with OS bookworm * 03:22 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1197.eqiad.wmnet with reason: host reimage * 03:18 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1165.eqiad.wmnet with reason: host reimage * 03:15 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1196.eqiad.wmnet with reason: host reimage * 03:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1164.eqiad.wmnet with reason: host reimage * 03:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1163.eqiad.wmnet with reason: host reimage * 03:05 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1165.eqiad.wmnet with reason: host reimage * 03:04 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1164.eqiad.wmnet with reason: host reimage * 03:04 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1197.eqiad.wmnet with reason: host reimage * 03:04 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1196.eqiad.wmnet with reason: host reimage * 03:04 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1163.eqiad.wmnet with reason: host reimage * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.16 refs [[phab:T430835|T430835]] * 02:50 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1197.eqiad.wmnet with OS bookworm * 02:49 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1196.eqiad.wmnet with OS bookworm * 02:49 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1165.eqiad.wmnet with OS bookworm * 02:49 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1164.eqiad.wmnet with OS bookworm * 02:48 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1163.eqiad.wmnet with OS bookworm * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 46s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:15 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-codfw: Set storage compatability to NONE — [[phab:T433026|T433026]] - eevans@cumin1003 * 00:38 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-codfw: Set storage compatability to NONE — [[phab:T433026|T433026]] - eevans@cumin1003 * 00:11 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326373{{!}}PersonalDashboard: add newly renamed *ReviewChangesMlModel setting (T422148)]] (duration: 07m 07s) * 00:09 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-eqiad: Set storage compatability to NONE — [[phab:T433026|T433026]] - eevans@cumin1003 * 00:07 musikanimal@deploy1003: musikanimal: Continuing with deployment * 00:06 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1326373{{!}}PersonalDashboard: add newly renamed *ReviewChangesMlModel setting (T422148)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:04 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1326373{{!}}PersonalDashboard: add newly renamed *ReviewChangesMlModel setting (T422148)]] == 2026-08-17 == * 23:30 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-eqiad: Set storage compatability to NONE — [[phab:T433026|T433026]] - eevans@cumin1003 * 23:05 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-codfw: Set storage compatability to UPGRADING — [[phab:T433026|T433026]] - eevans@cumin1003 * 22:28 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-codfw: Set storage compatability to UPGRADING — [[phab:T433026|T433026]] - eevans@cumin1003 * 21:46 logmsgbot: jforrester Deployed security patch for [[phab:T435085|T435085]] * 21:39 swfrench@deploy1003: mwscript-k8s job started: purgeList.php # [[phab:T432412|T432412]] * 21:37 maryum: Undeploy security fix for [[phab:T433020|T433020]] * 21:23 maryum: Deployed security fix for [[phab:T433020|T433020]] * 21:21 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 21:21 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 21:14 maryum: Deployed security fix for [[phab:T434967|T434967]] * 20:58 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-eqiad: Set storage compatability to UPGRADING — [[phab:T433026|T433026]] - eevans@cumin1003 * 20:40 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324320{{!}}[itwiki/slwiki/tgwiki] Remove temporary Wikipedia 25 logos permanently (already reverted) (T414265 T414320 T415307)]] (duration: 06m 55s) * 20:40 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 20:39 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 20:36 cjming@deploy1003: cjming, superpes: Continuing with deployment * 20:35 cjming@deploy1003: cjming, superpes: Backport for [[gerrit:1324320{{!}}[itwiki/slwiki/tgwiki] Remove temporary Wikipedia 25 logos permanently (already reverted) (T414265 T414320 T415307)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:33 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1324320{{!}}[itwiki/slwiki/tgwiki] Remove temporary Wikipedia 25 logos permanently (already reverted) (T414265 T414320 T415307)]] * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ttmserver-test: apply * 20:30 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326361{{!}}Remove escaped paths in app site association file (T432412)]] (duration: 13m 28s) * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ttmserver-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-toolhub-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-toolhub-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-toolhub-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-toolhub-test: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-toolhub: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-toolhub: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-toolhub: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-toolhub: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-test: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-test: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 20:27 inflatador: bking@deploy1003 `charlie --services_dir dse-k8s-services -s opensearch-* -e dse-k8s-* apply` [[phab:T435125|T435125]] * 20:27 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-apifeatureusage-test: apply * 20:27 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-apifeatureusage-test: apply * 20:27 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-apifeatureusage-test: apply * 20:27 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-apifeatureusage-test: apply * 20:27 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-apifeatureusage: apply * 20:26 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-apifeatureusage: apply * 20:26 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-apifeatureusage: apply * 20:26 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-apifeatureusage: apply * 20:26 cjming@deploy1003: cjming, tsev: Continuing with deployment * 20:24 inflatador: bking@deploy1003 `charlie --services_dir dse-k8s-services -s opensearch-* -e dse-k8s-*` [[phab:T435125|T435125]] * 20:21 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-eqiad: Set storage compatability to UPGRADING — [[phab:T433026|T433026]] - eevans@cumin1003 * 20:19 cjming@deploy1003: cjming, tsev: Backport for [[gerrit:1326361{{!}}Remove escaped paths in app site association file (T432412)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:17 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1326361{{!}}Remove escaped paths in app site association file (T432412)]] * 20:15 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326077{{!}}InstrumentConstructiveEdits: anchor all runs to the nearest `interval` (T431493)]] (duration: 06m 25s) * 20:11 cjming@deploy1003: cjming: Continuing with deployment * 20:10 cjming@deploy1003: cjming: Backport for [[gerrit:1326077{{!}}InstrumentConstructiveEdits: anchor all runs to the nearest `interval` (T431493)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:08 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1326077{{!}}InstrumentConstructiveEdits: anchor all runs to the nearest `interval` (T431493)]] * 20:06 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1160.eqiad.wmnet with OS bookworm * 20:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1162.eqiad.wmnet with OS bookworm * 19:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1184.eqiad.wmnet with OS bookworm * 19:55 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1161.eqiad.wmnet with OS bookworm * 19:49 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1195.eqiad.wmnet with OS bookworm * 19:49 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-test: apply * 19:49 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-test: apply * 19:44 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1160.eqiad.wmnet with reason: host reimage * 19:41 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-codfw: Upgrade to Java 17 — [[phab:T433026|T433026]] - eevans@cumin1003 * 19:39 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1162.eqiad.wmnet with reason: host reimage * 19:36 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-eqsin and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 19:36 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-test: apply * 19:36 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1184.eqiad.wmnet with reason: host reimage * 19:32 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1161.eqiad.wmnet with reason: host reimage * 19:29 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1195.eqiad.wmnet with reason: host reimage * 19:26 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1161.eqiad.wmnet with reason: host reimage * 19:26 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1160.eqiad.wmnet with reason: host reimage * 19:26 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1184.eqiad.wmnet with reason: host reimage * 19:26 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1162.eqiad.wmnet with reason: host reimage * 19:25 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1195.eqiad.wmnet with reason: host reimage * 19:19 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-test: apply * 19:11 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1195.eqiad.wmnet with OS bookworm * 19:10 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1194.eqiad.wmnet with OS bookworm * 19:10 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1184.eqiad.wmnet with OS bookworm * 19:10 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1162.eqiad.wmnet with OS bookworm * 19:10 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1161.eqiad.wmnet with OS bookworm * 19:10 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1160.eqiad.wmnet with OS bookworm * 19:10 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 19:09 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 19:03 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-codfw: Upgrade to Java 17 — [[phab:T433026|T433026]] - eevans@cumin1003 * 18:58 dancy@deploy1003: Finished scap sync-world: testing [[phab:T375514|T375514]] (duration: 03m 13s) * 18:55 dancy@deploy1003: Started scap sync-world: testing [[phab:T375514|T375514]] * 18:55 dwisehaupt@dns1006: END - running authdns-update * 18:54 dancy@deploy1003: Installation of scap version "4.281.1" completed for 3 hosts * 18:54 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2009.codfw.wmnet * 18:53 dwisehaupt@dns1006: START - running authdns-update * 18:52 dancy@deploy1003: Installing scap version "4.281.1" for 3 host(s) * 18:47 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2009.codfw.wmnet * 18:41 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2008.codfw.wmnet * 18:34 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2008.codfw.wmnet * 18:30 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2007.codfw.wmnet * 18:23 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2007.codfw.wmnet * 18:16 dwisehaupt@dns1005: END - running authdns-update * 18:14 dwisehaupt@dns1005: START - running authdns-update * 18:04 swfrench@deploy1003: Finished scap sync-world: Deploy "Point Test Wiki to new docroot" - [[phab:T432412|T432412]] (duration: 26m 03s) * 18:00 swfrench@deploy1003: swfrench: Continuing with deployment * 17:51 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1057.eqiad.wmnet with OS trixie * 17:47 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1180.eqiad.wmnet with OS bookworm * 17:39 swfrench@deploy1003: swfrench: Deploy "Point Test Wiki to new docroot" - [[phab:T432412|T432412]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:39 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-eqsin and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 17:38 swfrench@deploy1003: Started scap sync-world: Deploy "Point Test Wiki to new docroot" - [[phab:T432412|T432412]] * 17:36 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1159.eqiad.wmnet with OS bookworm * 17:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1158.eqiad.wmnet with OS bookworm * 17:30 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-eqiad: Upgrade to Java 17 — [[phab:T433026|T433026]] - eevans@cumin1003 * 17:30 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-ulsfo or A:cp-drmrs and A:cp - 9.2.15 upgrade ([[phab:T434620|T434620]]) * 17:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1157.eqiad.wmnet with OS bookworm * 17:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1180.eqiad.wmnet with reason: host reimage * 17:22 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 17:21 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1193.eqiad.wmnet with OS bookworm * 17:21 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 17:21 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 17:20 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 17:20 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 17:20 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1183.eqiad.wmnet with OS bookworm * 17:18 swfrench@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 17:18 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 17:17 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1180.eqiad.wmnet with reason: host reimage * 17:17 swfrench@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 17:17 swfrench@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 17:16 swfrench@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 17:16 swfrench@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 17:15 swfrench@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 17:15 swfrench@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 17:15 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1192.eqiad.wmnet with OS bookworm * 17:14 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1159.eqiad.wmnet with reason: host reimage * 17:14 swfrench@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 17:13 swfrench@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 17:12 swfrench@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 17:09 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1158.eqiad.wmnet with reason: host reimage * 17:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1157.eqiad.wmnet with reason: host reimage * 17:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1180 * 17:02 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1180 * 17:01 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica-eqsin and A:liberica * 17:01 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1193.eqiad.wmnet with reason: host reimage * 16:57 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1183.eqiad.wmnet with reason: host reimage * 16:56 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1054.eqiad.wmnet with OS trixie * 16:55 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1056.eqiad.wmnet with OS trixie * 16:54 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1193.eqiad.wmnet with reason: host reimage * 16:54 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1192.eqiad.wmnet with reason: host reimage * 16:51 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-eqiad: Upgrade to Java 17 — [[phab:T433026|T433026]] - eevans@cumin1003 * 16:50 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1180 * 16:50 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1180.eqiad.wmnet 17.36.64.10.in-addr.arpa 7.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:50 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1180.eqiad.wmnet 17.36.64.10.in-addr.arpa 7.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:50 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:50 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1180 - btullis@cumin1003" * 16:50 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1180 - btullis@cumin1003" * 16:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1158.eqiad.wmnet with reason: host reimage * 16:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1157.eqiad.wmnet with reason: host reimage * 16:49 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica-eqsin and A:liberica * 16:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1183.eqiad.wmnet with reason: host reimage * 16:48 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1159.eqiad.wmnet with reason: host reimage * 16:47 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1192.eqiad.wmnet with reason: host reimage * 16:46 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns4004.wikimedia.org * 16:46 sukhe@dns1004: END - running authdns-update * 16:44 sukhe@dns1004: START - running authdns-update * 16:44 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns4004.wikimedia.org,service=authdns-update * 16:43 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns4004.wikimedia.org with OS trixie * 16:39 btullis@cumin1003: START - Cookbook sre.dns.netbox * 16:39 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1180 * 16:39 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1193.eqiad.wmnet with OS bookworm * 16:39 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1180.eqiad.wmnet with OS bookworm * 16:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1192.eqiad.wmnet with OS bookworm * 16:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1183.eqiad.wmnet with OS bookworm * 16:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1159.eqiad.wmnet with OS bookworm * 16:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1158.eqiad.wmnet with OS bookworm * 16:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1157.eqiad.wmnet with OS bookworm * 16:31 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1057.eqiad.wmnet with OS trixie * 16:30 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1057.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 16:29 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica-drmrs and A:liberica * 16:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1154.eqiad.wmnet with OS bookworm * 16:23 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1057.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 16:23 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1055.eqiad.wmnet with OS trixie * 16:22 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1057 * 16:22 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1057 * 16:21 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:21 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1057] - vriley@cumin1003" * 16:21 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1057] - vriley@cumin1003" * 16:19 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica-drmrs and A:liberica * 16:17 vriley@cumin1003: START - Cookbook sre.dns.netbox * 16:16 phuedx@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics-external: apply * 16:15 phuedx@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics-external: apply * 16:13 phuedx@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics-external: apply * 16:12 phuedx@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics-external: apply * 16:11 phuedx@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics-external: apply * 16:09 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics-external: apply * 16:09 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1176.eqiad.wmnet with OS bookworm * 16:07 btullis@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 16:06 btullis@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 16:05 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1191.eqiad.wmnet with OS bookworm * 16:03 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1154.eqiad.wmnet with reason: host reimage * 16:03 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching aqs[2002-2012].codfw.wmnet,aqs[1017-1027].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433026|T433026]] - eevans@cumin1003 * 15:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1190.eqiad.wmnet with OS bookworm * 15:59 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1154.eqiad.wmnet with reason: host reimage * 15:55 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-debug: apply * 15:55 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-debug: apply * 15:55 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-debug: apply * 15:55 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/mw-debug: apply * 15:53 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns4004.wikimedia.org with reason: host reimage * 15:50 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns4004.wikimedia.org with reason: host reimage * 15:46 btullis@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 15:46 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1176.eqiad.wmnet with reason: host reimage * 15:45 btullis@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 15:43 btullis@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 15:42 btullis@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 15:42 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1191.eqiad.wmnet with reason: host reimage * 15:39 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1190.eqiad.wmnet with reason: host reimage * 15:36 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1054.eqiad.wmnet with OS trixie * 15:35 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:35 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1056.eqiad.wmnet with OS trixie * 15:35 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:34 moritzm: failover Ganeti master in eqiad to ganeti1046 * 15:34 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1176.eqiad.wmnet with reason: host reimage * 15:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1191.eqiad.wmnet with reason: host reimage * 15:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1190.eqiad.wmnet with reason: host reimage * 15:31 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns2006.wikimedia.org * 15:31 sukhe@dns1004: END - running authdns-update * 15:29 sukhe@dns1004: START - running authdns-update * 15:29 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns2006.wikimedia.org,service=authdns-update * 15:29 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns2006.wikimedia.org * 15:29 sukhe@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns2006.wikimedia.org * 15:26 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica-ulsfo and A:liberica * 15:26 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns1006.wikimedia.org * 15:26 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:25 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns2006.wikimedia.org with OS trixie * 15:25 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1056 * 15:25 sukhe@dns1004: END - running authdns-update * 15:23 sukhe@dns1004: START - running authdns-update * 15:23 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns1006.wikimedia.org,service=authdns-update * 15:23 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns1006.wikimedia.org * 15:23 sukhe@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns1006.wikimedia.org * 15:20 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1056 * 15:20 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:20 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1056~] - vriley@cumin1003" * 15:19 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1056~] - vriley@cumin1003" * 15:19 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns1006.wikimedia.org with OS trixie * 15:19 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns4004.wikimedia.org with OS trixie * 15:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1190.eqiad.wmnet with OS bookworm * 15:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1191.eqiad.wmnet with OS bookworm * 15:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1176.eqiad.wmnet with OS bookworm * 15:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1154.eqiad.wmnet with OS bookworm * 15:16 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica-ulsfo and A:liberica * 15:13 vriley@cumin1003: START - Cookbook sre.dns.netbox * 15:13 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host dns4004.wikimedia.org with OS trixie * 15:12 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1188.eqiad.wmnet with OS bookworm * 15:09 taavi@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318203{{!}}Undeploy WP25EasterEggs (II) (T418134)]] (duration: 08m 56s) * 15:06 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica-magru and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 15:06 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1051.eqiad.wmnet with OS trixie * 15:06 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 15:05 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 15:05 taavi@deploy1003: taavi: Continuing with deployment * 15:04 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-ncredir (exit_code=0) rolling reboot on A:ncredir and A:ncredir * 15:04 taavi@deploy1003: taavi: Backport for [[gerrit:1318203{{!}}Undeploy WP25EasterEggs (II) (T418134)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:03 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1055.eqiad.wmnet with OS trixie * 15:02 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:00 taavi@deploy1003: Started scap sync-world: Backport for [[gerrit:1318203{{!}}Undeploy WP25EasterEggs (II) (T418134)]] * 14:59 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1047.eqiad.wmnet * 14:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1047.eqiad.wmnet * 14:58 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy (exit_code=0) rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 14:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-codfw * 14:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp2001.codfw.wmnet * 14:57 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp2001.codfw.wmnet * 14:57 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica-magru and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 14:57 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:56 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1055 * 14:56 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1055 * 14:55 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns2006.wikimedia.org with reason: host reimage * 14:55 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:55 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1055] - vriley@cumin1003" * 14:55 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1055] - vriley@cumin1003" * 14:54 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1047.eqiad.wmnet * 14:52 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1188.eqiad.wmnet with reason: host reimage * 14:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp2001.codfw.wmnet * 14:51 vriley@cumin1003: START - Cookbook sre.dns.netbox * 14:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp2001.codfw.wmnet * 14:50 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2016.codfw.wmnet * 14:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2016.codfw.wmnet * 14:50 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1189.eqiad.wmnet with OS bookworm * 14:48 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1188.eqiad.wmnet with reason: host reimage * 14:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1051.eqiad.wmnet with reason: host reimage * 14:46 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1182.eqiad.wmnet with OS bookworm * 14:45 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief2002.codfw.wmnet * 14:44 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns1006.wikimedia.org with reason: host reimage * 14:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2016.codfw.wmnet * 14:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1143.eqiad.wmnet with OS bookworm * 14:43 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2016.codfw.wmnet * 14:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2318-2331].codfw.wmnet * 14:43 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2318-2331].codfw.wmnet * 14:42 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching aqs[2002-2012].codfw.wmnet,aqs[1017-1027].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433026|T433026]] - eevans@cumin1003 * 14:41 cgoubert@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326306{{!}}Add placeholder $wmgRedisLockPassword (T366938 T427999)]] (duration: 06m 56s) * 14:41 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief2002.codfw.wmnet * 14:40 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief1002.eqiad.wmnet * 14:39 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1051.eqiad.wmnet with reason: host reimage * 14:38 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns2006.wikimedia.org with reason: host reimage * 14:37 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns1006.wikimedia.org with reason: host reimage * 14:37 cgoubert@deploy1003: cgoubert: Continuing with deployment * 14:36 cgoubert@deploy1003: cgoubert: Backport for [[gerrit:1326306{{!}}Add placeholder $wmgRedisLockPassword (T366938 T427999)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:36 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief1002.eqiad.wmnet * 14:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2318-2331].codfw.wmnet * 14:34 cgoubert@deploy1003: Started scap sync-world: Backport for [[gerrit:1326306{{!}}Add placeholder $wmgRedisLockPassword (T366938 T427999)]] * 14:29 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2318-2331].codfw.wmnet * 14:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2304-2317].codfw.wmnet * 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2304-2317].codfw.wmnet * 14:28 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1054 * 14:27 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1054 * 14:27 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:27 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1054] - vriley@cumin1003" * 14:27 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1054] - vriley@cumin1003" * 14:26 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1189.eqiad.wmnet with reason: host reimage * 14:26 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test2001.codfw.wmnet * 14:25 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test1001.eqiad.wmnet * 14:25 claime: Deploying wmgRedisLockPassword - [[phab:T366938|T366938]] [[phab:T427999|T427999]] * 14:24 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1051.eqiad.wmnet with OS trixie * 14:23 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 14:22 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1182.eqiad.wmnet with reason: host reimage * 14:22 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1051.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:22 vriley@cumin1003: START - Cookbook sre.dns.netbox * 14:22 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test2001.codfw.wmnet * 14:21 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test1001.eqiad.wmnet * 14:21 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 14:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2304-2317].codfw.wmnet * 14:19 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns4004.wikimedia.org with OS trixie * 14:19 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns2006.wikimedia.org with OS trixie * 14:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1143.eqiad.wmnet with reason: host reimage * 14:19 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns1006.wikimedia.org with OS trixie * 14:17 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1189.eqiad.wmnet with reason: host reimage * 14:15 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1182.eqiad.wmnet with reason: host reimage * 14:14 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1143.eqiad.wmnet with reason: host reimage * 14:13 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1051.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:12 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2304-2317].codfw.wmnet * 14:12 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2290-2303].codfw.wmnet * 14:12 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1049.eqiad.wmnet with OS trixie * 14:12 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2290-2303].codfw.wmnet * 14:12 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1051 * 14:11 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1051 * 14:11 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1170.eqiad.wmnet onto db1284.eqiad.wmnet * 14:11 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1170: Pool db1170.eqiad.wmnet in after cloning * 14:09 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 14:09 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:09 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1051] - vriley@cumin1003" * 14:09 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1051] - vriley@cumin1003" * 14:06 klausman@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:04 vriley@cumin1003: START - Cookbook sre.dns.netbox * 14:04 klausman@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2290-2303].codfw.wmnet * 14:02 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1189.eqiad.wmnet with OS bookworm * 14:02 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1188.eqiad.wmnet with OS bookworm * 14:01 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 14:00 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1182.eqiad.wmnet with OS bookworm * 14:00 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1143.eqiad.wmnet with OS bookworm * 13:59 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 13:57 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply * 13:57 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply * 13:56 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply * 13:56 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-ulsfo or A:cp-drmrs and A:cp - 9.2.15 upgrade ([[phab:T434620|T434620]]) * 13:56 cjd91: sudo -i cookbook sre.cdn.roll-upgrade-ats --query 'A:cp-ulsfo or A:cp-drmrs' --task-id [[phab:T434620|T434620]] --reason '9.2.15 upgrade' * 13:56 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply * 13:55 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply * 13:55 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply * 13:54 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 13:54 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 13:54 phuedx@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:53 phuedx@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics-external: apply * 13:52 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1049.eqiad.wmnet with reason: host reimage * 13:51 phuedx@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2290-2303].codfw.wmnet * 13:50 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1047.eqiad.wmnet * 13:50 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2276-2289].codfw.wmnet * 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2276-2289].codfw.wmnet * 13:49 phuedx@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics-external: apply * 13:49 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1046.eqiad.wmnet * 13:49 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1049.eqiad.wmnet with reason: host reimage * 13:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1046.eqiad.wmnet * 13:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-ncredir rolling reboot on A:ncredir and A:ncredir * 13:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 13:46 phuedx@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:44 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics-external: apply * 13:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1046.eqiad.wmnet * 13:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2276-2289].codfw.wmnet * 13:36 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1046.eqiad.wmnet * 13:36 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1045.eqiad.wmnet * 13:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1045.eqiad.wmnet * 13:34 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2276-2289].codfw.wmnet * 13:34 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1049.eqiad.wmnet with OS trixie * 13:34 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2262-2275].codfw.wmnet * 13:34 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2262-2275].codfw.wmnet * 13:32 Lucas_WMDE: UTC afternoon backport+config window done * 13:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1045.eqiad.wmnet * 13:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2262-2275].codfw.wmnet * 13:26 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1049.eqiad.wmnet with OS trixie * 13:26 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1049.eqiad.wmnet with OS trixie * 13:26 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1170: Pool db1170.eqiad.wmnet in after cloning * 13:23 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1049.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 13:23 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1045.eqiad.wmnet * 13:21 atsukoito: manually done sudo -i docker-registryctl --debug delete-tags 'docker-registry.discovery.wmnet/repos/data-engineering/airflow-dags:airflow-3.3.0-py3.11-2026-08-17-*' to remove incorrect tags * 13:19 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311141{{!}}viwiki: Set `noindex,nofollow` for User and User talk (T432311)]] (duration: 11m 11s) * 13:15 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2262-2275].codfw.wmnet * 13:14 lucaswerkmeister-wmde@deploy1003: ndkdd, lucaswerkmeister-wmde: Continuing with deployment * 13:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2248-2261].codfw.wmnet * 13:14 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2248-2261].codfw.wmnet * 13:13 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1049.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 13:12 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1049 * 13:11 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1049 * 13:10 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:10 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1049] - vriley@cumin1003" * 13:10 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1049] - vriley@cumin1003" * 13:10 lucaswerkmeister-wmde@deploy1003: ndkdd, lucaswerkmeister-wmde: Backport for [[gerrit:1311141{{!}}viwiki: Set `noindex,nofollow` for User and User talk (T432311)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1037.eqiad.wmnet * 13:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1037.eqiad.wmnet * 13:08 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1311141{{!}}viwiki: Set `noindex,nofollow` for User and User talk (T432311)]] * 13:06 vriley@cumin1003: START - Cookbook sre.dns.netbox * 13:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2248-2261].codfw.wmnet * 13:00 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1037.eqiad.wmnet * 12:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2248-2261].codfw.wmnet * 12:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2204-2215,2242-2243].codfw.wmnet * 12:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2204-2215,2242-2243].codfw.wmnet * 12:49 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324832{{!}}Migrate $wgFlaggedRevsTags from flaggedrevs.php to ext-FlaggedRevs.php]] (duration: 14m 02s) * 12:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint2001.codfw.wmnet * 12:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2204-2215,2242-2243].codfw.wmnet * 12:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint2001.codfw.wmnet * 12:41 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1037.eqiad.wmnet * 12:40 ladsgroup@deploy1003: Rolling back deployment * 12:37 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1324832{{!}}Migrate $wgFlaggedRevsTags from flaggedrevs.php to ext-FlaggedRevs.php]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:37 seanleong-wmde: Finished populateSitesTable for [bolwiki] ([[[phab:T429955|T429955]]]) * 12:35 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1324832{{!}}Migrate $wgFlaggedRevsTags from flaggedrevs.php to ext-FlaggedRevs.php]] * 12:35 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2204-2215,2242-2243].codfw.wmnet * 12:34 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2190-2203].codfw.wmnet * 12:34 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2190-2203].codfw.wmnet * 12:32 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1028.eqiad.wmnet * 12:32 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1028.eqiad.wmnet * 12:31 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint1001.eqiad.wmnet * 12:30 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1170: Depool db1170.eqiad.wmnet to then clone it to db1284.eqiad.wmnet - marostegui@cumin1003 * 12:28 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint1001.eqiad.wmnet * 12:28 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1170: Depool db1170.eqiad.wmnet to then clone it to db1284.eqiad.wmnet - marostegui@cumin1003 * 12:27 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1170.eqiad.wmnet onto db1284.eqiad.wmnet * 12:27 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db2901.codfw.wmnet * 12:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2190-2203].codfw.wmnet * 12:27 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 12:26 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1028.eqiad.wmnet * 12:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2190-2203].codfw.wmnet * 12:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2172-2179,2184-2189].codfw.wmnet * 12:18 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2172-2179,2184-2189].codfw.wmnet * 12:13 seanleong-wmde@deploy1003: mwscript-k8s job started: foreachwikiindblist wikidataclient extensions/Wikibase/lib/maintenance/populateSitesTable.php --force-protocol https # [[phab:T429955|T429955]] * 12:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2172-2179,2184-2189].codfw.wmnet * 12:09 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1028.eqiad.wmnet * 12:03 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 12:02 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db2901.codfw.wmnet * 12:02 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db2901.codfw.wmnet * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1027.eqiad.wmnet * 12:02 fceratto@cumin1003: END (ERROR) - Cookbook sre.dns.netbox (exit_code=97) * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1027.eqiad.wmnet * 12:02 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 12:02 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db2901.codfw.wmnet * 12:01 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2172-2179,2184-2189].codfw.wmnet * 12:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2158-2171].codfw.wmnet * 12:01 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2158-2171].codfw.wmnet * 11:58 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db1902.eqiad.wmnet * 11:58 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 11:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1027.eqiad.wmnet * 11:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2158-2171].codfw.wmnet * 11:51 jayme: updated calico to v3.30.7 on wikikube eqiad - [[phab:T427400|T427400]] * 11:50 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1027.eqiad.wmnet * 11:45 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2158-2171].codfw.wmnet * 11:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2144-2157].codfw.wmnet * 11:44 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2144-2157].codfw.wmnet * 11:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1058.eqiad.wmnet * 11:43 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1058.eqiad.wmnet * 11:43 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 11:43 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 11:43 marostegui@cumin1003: Removing db1153 from zarcillo [[phab:T434638|T434638]] * 11:42 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1153.eqiad.wmnet * 11:42 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:42 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1153.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 11:42 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1172.eqiad.wmnet onto db1286.eqiad.wmnet * 11:42 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1172: Pool db1172.eqiad.wmnet in after cloning * 11:42 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1153.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 11:41 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1902.eqiad.wmnet * 11:41 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=97) for new host db2901.codfw.wmnet * 11:41 fceratto@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host db2901.codfw.wmnet with OS trixie * 11:38 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'. * 11:38 marostegui@cumin1003: START - Cookbook sre.dns.netbox * 11:37 marostegui@dns1004: END - running authdns-update * 11:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1058.eqiad.wmnet * 11:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2144-2157].codfw.wmnet * 11:35 marostegui@dns1004: START - running authdns-update * 11:32 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1153.eqiad.wmnet * 11:32 marostegui@cumin1003: START - Cookbook sre.mysql.decommission * 11:28 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326249{{!}}ImagePage: move TOC element below file link (T332644)]] (duration: 09m 56s) * 11:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2144-2157].codfw.wmnet * 11:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2130-2143].codfw.wmnet * 11:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2130-2143].codfw.wmnet * 11:26 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1058.eqiad.wmnet * 11:23 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 11:22 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'. * 11:22 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'. * 11:22 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'. * 11:22 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326249{{!}}ImagePage: move TOC element below file link (T332644)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1057.eqiad.wmnet * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1057.eqiad.wmnet * 11:21 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 11:20 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 11:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2130-2143].codfw.wmnet * 11:19 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326249{{!}}ImagePage: move TOC element below file link (T332644)]] * 11:18 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply * 11:17 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply * 11:17 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply * 11:16 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 11:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1057.eqiad.wmnet * 11:15 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 11:14 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 11:13 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 11:13 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db2901.codfw.wmnet with OS trixie * 11:12 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db2901.codfw.wmnet - fceratto@cumin1003" * 11:12 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db2901.codfw.wmnet - fceratto@cumin1003" * 11:12 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db2901.codfw.wmnet on all recursors * 11:12 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db2901.codfw.wmnet on all recursors * 11:12 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:12 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db2901.codfw.wmnet - fceratto@cumin1003" * 11:12 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2130-2143].codfw.wmnet * 11:11 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2107-2115,2124-2129].codfw.wmnet * 11:11 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2107-2115,2124-2129].codfw.wmnet * 11:11 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 11:11 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 11:11 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply * 11:11 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 11:10 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 11:06 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db2901.codfw.wmnet - fceratto@cumin1003" * 11:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2107-2115,2124-2129].codfw.wmnet * 10:57 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1172: Pool db1172.eqiad.wmnet in after cloning * 10:54 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2107-2115,2124-2129].codfw.wmnet * 10:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2078,2087-2095,2102-2106].codfw.wmnet * 10:54 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2078,2087-2095,2102-2106].codfw.wmnet * 10:47 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply * 10:46 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply * 10:46 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply * 10:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2078,2087-2095,2102-2106].codfw.wmnet * 10:45 blake@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply * 10:38 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1057.eqiad.wmnet * 10:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1056.eqiad.wmnet * 10:38 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1056.eqiad.wmnet * 10:37 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2078,2087-2095,2102-2106].codfw.wmnet * 10:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2061-2062,2064-2065,2067-2077].codfw.wmnet * 10:36 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2061-2062,2064-2065,2067-2077].codfw.wmnet * 10:32 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1056.eqiad.wmnet * 10:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2061-2062,2064-2065,2067-2077].codfw.wmnet * 10:25 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:25 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db2901.codfw.wmnet * 10:25 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db1902.eqiad.wmnet * 10:25 fceratto@cumin1003: END (ERROR) - Cookbook sre.dns.netbox (exit_code=97) * 10:24 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:24 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1902.eqiad.wmnet * 10:24 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=97) for new host db1902.eqiad.wmnet * 10:24 fceratto@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host db1902.eqiad.wmnet with OS trixie * 10:24 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=93) for new host db1903.eqiad.wmnet * 10:24 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 10:20 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1056.eqiad.wmnet * 10:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2061-2062,2064-2065,2067-2077].codfw.wmnet * 10:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2038-2039,2041-2042,2044,2046,2049-2051,2055-2060].codfw.wmnet * 10:18 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2038-2039,2041-2042,2044,2046,2049-2051,2055-2060].codfw.wmnet * 10:15 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1055.eqiad.wmnet * 10:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1055.eqiad.wmnet * 10:13 Amir1: mwscript-k8s --dblist=all -- purgeUserOptions.php --login-age 5 uls-preferences ([[phab:T406724|T406724]]) * 10:11 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:10 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1903.eqiad.wmnet on all recursors * 10:10 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1903.eqiad.wmnet on all recursors * 10:10 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2038-2039,2041-2042,2044,2046,2049-2051,2055-2060].codfw.wmnet * 10:10 moritzm: installing unzip security updates * 10:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1055.eqiad.wmnet * 10:08 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=97) for new host db1901.eqiad.wmnet * 10:08 fceratto@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host db1901.eqiad.wmnet with OS trixie * 10:08 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:08 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 10:08 fceratto@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1903.eqiad.wmnet - fceratto@cumin1003" * 10:07 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db1902.eqiad.wmnet with OS trixie * 10:07 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1902.eqiad.wmnet - fceratto@cumin1003" * 10:07 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1902.eqiad.wmnet - fceratto@cumin1003" * 10:04 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply * 10:04 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326227{{!}}Enable desktop/native lazy loading everywhere (T148047)]] (duration: 07m 13s) * 10:03 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1902.eqiad.wmnet on all recursors * 10:03 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1902.eqiad.wmnet on all recursors * 10:03 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:03 blake@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply * 10:01 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1903.eqiad.wmnet - fceratto@cumin1003" * 10:01 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2038-2039,2041-2042,2044,2046,2049-2051,2055-2060].codfw.wmnet * 10:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2002,2005-2006,2011-2015,2017-2018,2033-2037].codfw.wmnet * 10:01 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2002,2005-2006,2011-2015,2017-2018,2033-2037].codfw.wmnet * 10:00 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:00 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 09:59 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 09:59 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1055.eqiad.wmnet * 09:58 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326227{{!}}Enable desktop/native lazy loading everywhere (T148047)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:56 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326227{{!}}Enable desktop/native lazy loading everywhere (T148047)]] * 09:53 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2002,2005-2006,2011-2015,2017-2018,2033-2037].codfw.wmnet * 09:50 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:48 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1903.eqiad.wmnet * 09:48 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:48 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1902.eqiad.wmnet * 09:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 09:44 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2002,2005-2006,2011-2015,2017-2018,2033-2037].codfw.wmnet * 09:43 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-codfw * 09:43 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1174.eqiad.wmnet onto db1288.eqiad.wmnet * 09:42 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1174: Pool db1174.eqiad.wmnet in after cloning * 09:40 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1172: Depool db1172.eqiad.wmnet to then clone it to db1286.eqiad.wmnet - marostegui@cumin1003 * 09:39 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1172: Depool db1172.eqiad.wmnet to then clone it to db1286.eqiad.wmnet - marostegui@cumin1003 * 09:39 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1172.eqiad.wmnet onto db1286.eqiad.wmnet * 09:33 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1175.eqiad.wmnet onto db1289.eqiad.wmnet * 09:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1175: Pool db1175.eqiad.wmnet in after cloning * 09:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1201.eqiad.wmnet onto db1287.eqiad.wmnet * 09:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1201: Pool db1201.eqiad.wmnet in after cloning * 09:28 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db1901.eqiad.wmnet with OS trixie * 09:27 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:27 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:27 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1901.eqiad.wmnet on all recursors * 09:27 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1901.eqiad.wmnet on all recursors * 09:26 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:26 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:26 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:15 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:15 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1901.eqiad.wmnet * 09:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw1001.wikimedia.org with OS trixie * 08:57 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1174: Pool db1174.eqiad.wmnet in after cloning * 08:54 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1054.eqiad.wmnet * 08:54 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1054.eqiad.wmnet * 08:48 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1054.eqiad.wmnet * 08:47 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1175: Pool db1175.eqiad.wmnet in after cloning * 08:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 08:46 marostegui@cumin1003: Removing db1152 from zarcillo [[phab:T434480|T434480]] * 08:46 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1152.eqiad.wmnet * 08:46 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:46 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1152.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 08:46 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1201: Pool db1201.eqiad.wmnet in after cloning * 08:46 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1152.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 08:46 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1054.eqiad.wmnet * 08:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1053.eqiad.wmnet * 08:43 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1053.eqiad.wmnet * 08:42 marostegui@cumin1003: START - Cookbook sre.dns.netbox * 08:38 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage * 08:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1053.eqiad.wmnet * 08:36 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1152.eqiad.wmnet * 08:36 marostegui@cumin1003: START - Cookbook sre.mysql.decommission * 08:35 phuedx: UTC morning backport window done * 08:35 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1053.eqiad.wmnet * 08:34 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1035.eqiad.wmnet * 08:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1035.eqiad.wmnet * 08:34 phuedx@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324725{{!}}EventStreamConfig: Mark product_metrics.web_base and .web_base_with_ip as Test Kitchen streams (T429898 T430322)]], [[gerrit:1313923{{!}}EventStreamConfig: Remove unused web_ui_scroll* streams (T415370)]], [[gerrit:1325546{{!}}EventStreamConfig: Remove Watchlist click stream (T434790)]] (duration: 12m 42s) * 08:33 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1281: Pool back * 08:32 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage * 08:31 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on db2209.codfw.wmnet with reason: Maintenance * 08:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2209: Maintenance needed * 08:30 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2209: Maintenance needed * 08:26 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1035.eqiad.wmnet * 08:26 phuedx@deploy1003: bearloga, phuedx: Continuing with deployment * 08:23 phuedx@deploy1003: bearloga, phuedx: Backport for [[gerrit:1324725{{!}}EventStreamConfig: Mark product_metrics.web_base and .web_base_with_ip as Test Kitchen streams (T429898 T430322)]], [[gerrit:1313923{{!}}EventStreamConfig: Remove unused web_ui_scroll* streams (T415370)]], [[gerrit:1325546{{!}}EventStreamConfig: Remove Watchlist click stream (T434790)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug * 08:21 phuedx@deploy1003: Started scap sync-world: Backport for [[gerrit:1324725{{!}}EventStreamConfig: Mark product_metrics.web_base and .web_base_with_ip as Test Kitchen streams (T429898 T430322)]], [[gerrit:1313923{{!}}EventStreamConfig: Remove unused web_ui_scroll* streams (T415370)]], [[gerrit:1325546{{!}}EventStreamConfig: Remove Watchlist click stream (T434790)]] * 08:19 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw1001.wikimedia.org with OS trixie * 08:18 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1035.eqiad.wmnet * 08:17 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1032.eqiad.wmnet * 08:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1032.eqiad.wmnet * 08:16 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1279: Pool back * 08:15 phuedx@deploy1003: Finished scap sync-world: Backport for [[gerrit:1216721{{!}}viwikivoyage: enable relatedarticle and pop-up (T405724)]] (duration: 39m 12s) * 08:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1032.eqiad.wmnet * 08:09 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1032.eqiad.wmnet * 08:08 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1031.eqiad.wmnet * 08:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1031.eqiad.wmnet * 08:03 godog: switch production to use dumps-nfs.w.o - [[phab:T432212|T432212]] * 08:02 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1031.eqiad.wmnet * 08:02 phuedx@deploy1003: nvdtn19, phuedx: Continuing with deployment * 08:00 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1201: Depool db1201.eqiad.wmnet to then clone it to db1287.eqiad.wmnet - marostegui@cumin1003 * 08:00 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1201: Depool db1201.eqiad.wmnet to then clone it to db1287.eqiad.wmnet - marostegui@cumin1003 * 08:00 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1201.eqiad.wmnet onto db1287.eqiad.wmnet * 08:00 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1057.eqiad.wmnet * 08:00 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:00 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1057.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:59 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1057.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:59 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1275: Pool back * 07:55 filippo@cumin1003: START - Cookbook sre.dns.netbox * 07:55 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1031.eqiad.wmnet * 07:52 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1030.eqiad.wmnet * 07:52 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1030.eqiad.wmnet * 07:52 phuedx@deploy1003: nvdtn19, phuedx: Backport for [[gerrit:1216721{{!}}viwikivoyage: enable relatedarticle and pop-up (T405724)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:51 tappof: bump space for prometheus k8s-aux in eqiad * 07:50 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1057.eqiad.wmnet * 07:48 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1281: Pool back * 07:47 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1281 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96108 and previous config saved to /var/cache/conftool/dbconfig/20260817-074749-marostegui.json * 07:46 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1030.eqiad.wmnet * 07:42 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1030.eqiad.wmnet * 07:41 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1174: Depool db1174.eqiad.wmnet to then clone it to db1288.eqiad.wmnet - marostegui@cumin1003 * 07:41 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1174: Depool db1174.eqiad.wmnet to then clone it to db1288.eqiad.wmnet - marostegui@cumin1003 * 07:41 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1174.eqiad.wmnet onto db1288.eqiad.wmnet * 07:40 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1029.eqiad.wmnet * 07:40 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm2001.wikimedia.org * 07:40 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1029.eqiad.wmnet * 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1056.eqiad.wmnet * 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1056.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:38 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1056.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:36 phuedx@deploy1003: Started scap sync-world: Backport for [[gerrit:1216721{{!}}viwikivoyage: enable relatedarticle and pop-up (T405724)]] * 07:36 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm2001.wikimedia.org * 07:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1279 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96104 and previous config saved to /var/cache/conftool/dbconfig/20260817-073542-marostegui.json * 07:34 filippo@cumin1003: START - Cookbook sre.dns.netbox * 07:34 slyngshede@dns1004: END - running authdns-update * 07:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1029.eqiad.wmnet * 07:33 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm-test1001.wikimedia.org * 07:32 slyngshede@dns1004: START - running authdns-update * 07:31 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1029.eqiad.wmnet * 07:31 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1279: Pool back * 07:30 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1279 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96102 and previous config saved to /var/cache/conftool/dbconfig/20260817-073038-marostegui.json * 07:29 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm-test1001.wikimedia.org * 07:29 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm1001.wikimedia.org * 07:28 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1056.eqiad.wmnet * 07:28 moritzm: extend the disk of ldap-rw1001 by 80G [[phab:T331699|T331699]] * 07:28 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1055.eqiad.wmnet * 07:28 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:28 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1055.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:27 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1055.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:26 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1044.eqiad.wmnet * 07:26 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1044.eqiad.wmnet * 07:25 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm1001.wikimedia.org * 07:24 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1175: Depool db1175.eqiad.wmnet to then clone it to db1289.eqiad.wmnet - marostegui@cumin1003 * 07:24 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1175: Depool db1175.eqiad.wmnet to then clone it to db1289.eqiad.wmnet - marostegui@cumin1003 * 07:24 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1175.eqiad.wmnet onto db1289.eqiad.wmnet * 07:22 filippo@cumin1003: START - Cookbook sre.dns.netbox * 07:20 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1044.eqiad.wmnet * 07:16 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1055.eqiad.wmnet * 07:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1054.eqiad.wmnet * 07:15 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:15 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1054.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:15 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1044.eqiad.wmnet * 07:15 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1054.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:13 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1275: Pool back * 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1275 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96098 and previous config saved to /var/cache/conftool/dbconfig/20260817-071225-marostegui.json * 07:10 filippo@cumin1003: START - Cookbook sre.dns.netbox * 07:05 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1043.eqiad.wmnet * 07:05 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1054.eqiad.wmnet * 07:05 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1051.eqiad.wmnet * 07:05 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:05 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1051.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:05 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1043.eqiad.wmnet * 07:04 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1051.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 06:59 filippo@cumin1003: START - Cookbook sre.dns.netbox * 06:59 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin1001.eqiad.wmnet * 06:59 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1043.eqiad.wmnet * 06:59 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin2001.codfw.wmnet * 06:55 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin2001.codfw.wmnet * 06:55 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1051.eqiad.wmnet * 06:54 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1049.eqiad.wmnet * 06:54 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:54 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1049.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 06:54 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1049.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 06:54 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin1001.eqiad.wmnet * 06:53 moritzm: installing apr-util security updates * 06:52 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1043.eqiad.wmnet * 06:49 filippo@cumin1003: START - Cookbook sre.dns.netbox * 06:41 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1049.eqiad.wmnet * 06:13 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit2003.wikimedia.org * 06:13 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet * 06:07 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet * 06:06 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit2003.wikimedia.org * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 47s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-16 == * 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 01m 03s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-15 == * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 41s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-14 == * 15:38 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-staging-master-eqiad * 15:38 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster1005.eqiad.wmnet * 15:38 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster1005.eqiad.wmnet * 15:35 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sretest2009.codfw.wmnet * 15:33 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster1005.eqiad.wmnet * 15:33 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster1005.eqiad.wmnet * 15:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster1004.eqiad.wmnet * 15:32 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster1004.eqiad.wmnet * 15:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host sretest2009.codfw.wmnet * 15:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster1004.eqiad.wmnet * 15:27 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster1004.eqiad.wmnet * 15:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster1003.eqiad.wmnet * 15:27 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster1003.eqiad.wmnet * 15:24 dancy@deploy1003: Finished scap sync-world: testing (duration: 03m 23s) * 15:22 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster1003.eqiad.wmnet * 15:22 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster1003.eqiad.wmnet * 15:22 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-staging-master-eqiad * 15:20 dancy@deploy1003: Started scap sync-world: testing * 15:20 dancy@deploy1003: Installation of scap version "4.280.2" completed for 3 hosts * 15:18 dancy@deploy1003: Installing scap version "4.280.2" for 3 host(s) * 15:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sretest2006.codfw.wmnet * 14:54 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host sretest2006.codfw.wmnet * 14:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sretest2003.codfw.wmnet * 14:39 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host sretest2003.codfw.wmnet * 13:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt-staging2001.codfw.wmnet * 13:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt-staging2001.codfw.wmnet * 13:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-staging-master-codfw * 13:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster2005.codfw.wmnet * 13:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster2005.codfw.wmnet * 13:05 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox-dev2003.codfw.wmnet * 13:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster2005.codfw.wmnet * 13:04 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster2005.codfw.wmnet * 13:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster2004.codfw.wmnet * 13:04 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster2004.codfw.wmnet * 13:01 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netbox-dev2003.codfw.wmnet * 12:59 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster2004.codfw.wmnet * 12:59 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster2004.codfw.wmnet * 12:59 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster2003.codfw.wmnet * 12:59 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster2003.codfw.wmnet * 12:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster2003.codfw.wmnet * 12:54 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster2003.codfw.wmnet * 12:54 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-staging-master-codfw * 12:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-staging-worker-eqiad * 12:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage1006.eqiad.wmnet * 12:52 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage1006.eqiad.wmnet * 12:46 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage1006.eqiad.wmnet * 12:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw1001.wikimedia.org with OS trixie * 12:45 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage1006.eqiad.wmnet * 12:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage1005.eqiad.wmnet * 12:45 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage1005.eqiad.wmnet * 12:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage1005.eqiad.wmnet * 12:36 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/ratelimit: apply * 12:35 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/ratelimit: apply * 12:35 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:35 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:33 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage1005.eqiad.wmnet * 12:33 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage1004.eqiad.wmnet * 12:33 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage1004.eqiad.wmnet * 12:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage1004.eqiad.wmnet * 12:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage1004.eqiad.wmnet * 12:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage1003.eqiad.wmnet * 12:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage1003.eqiad.wmnet * 12:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage * 12:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage1003.eqiad.wmnet * 12:17 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage * 12:14 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage1003.eqiad.wmnet * 12:14 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-staging-worker-eqiad * 12:03 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw1001.wikimedia.org with OS trixie * 12:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cuminunpriv1001.eqiad.wmnet * 11:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cuminunpriv1001.eqiad.wmnet * 11:27 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1004.wikimedia.org * 11:24 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1153 from dbctl [[phab:T434638|T434638]]', diff saved to https://phabricator.wikimedia.org/P96097 and previous config saved to /var/cache/conftool/dbconfig/20260814-112449-marostegui.json * 11:21 aokoth@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1004.wikimedia.org * 11:20 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 11:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-staging-worker-codfw * 11:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2004.codfw.wmnet * 11:17 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2004.codfw.wmnet * 11:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2004.codfw.wmnet * 11:10 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2004.codfw.wmnet * 11:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2003.codfw.wmnet * 11:10 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2003.codfw.wmnet * 11:03 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-worker1181.eqiad.wmnet with OS bookworm * 11:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2003.codfw.wmnet * 11:03 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2003.codfw.wmnet * 11:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2002.codfw.wmnet * 11:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2002.codfw.wmnet * 10:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1235.eqiad.wmnet with OS bookworm * 10:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock2003.codfw.wmnet * 10:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock2003.codfw.wmnet with OS trixie * 10:56 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2002.codfw.wmnet * 10:54 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1153.eqiad.wmnet with OS bookworm * 10:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2002.codfw.wmnet * 10:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2001.codfw.wmnet * 10:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2001.codfw.wmnet * 10:50 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 10:49 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1187.eqiad.wmnet with OS bookworm * 10:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2001.codfw.wmnet * 10:44 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2001.codfw.wmnet * 10:44 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-staging-worker-codfw * 10:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock2003.codfw.wmnet with reason: host reimage * 10:38 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock2003.codfw.wmnet with reason: host reimage * 10:35 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1153.eqiad.wmnet with reason: host reimage * 10:32 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1235.eqiad.wmnet with reason: host reimage * 10:29 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1187.eqiad.wmnet with reason: host reimage * 10:24 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1153.eqiad.wmnet with reason: host reimage * 10:23 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1235.eqiad.wmnet with reason: host reimage * 10:21 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1187.eqiad.wmnet with reason: host reimage * 10:16 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock2003.codfw.wmnet with OS trixie * 10:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install1005.wikimedia.org * 10:12 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 10:12 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 10:12 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 10:12 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 10:12 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:12 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 10:12 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 10:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install1005.wikimedia.org * 10:07 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1235.eqiad.wmnet with OS bookworm * 10:07 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1187.eqiad.wmnet with OS bookworm * 10:07 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1181.eqiad.wmnet with OS bookworm * 10:07 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1153.eqiad.wmnet with OS bookworm * 10:07 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install2005.wikimedia.org * 10:03 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 10:03 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2003.codfw.wmnet * 10:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1232.eqiad.wmnet with OS bookworm * 10:00 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install2005.wikimedia.org * 10:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install3004.wikimedia.org * 09:58 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 09:58 fceratto@cumin1003: Removing db1151 from zarcillo [[phab:T434538|T434538]] * 09:56 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1231.eqiad.wmnet with OS bookworm * 09:56 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 09:53 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 09:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install3004.wikimedia.org * 09:50 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1152.eqiad.wmnet with OS bookworm * 09:49 Dreamy_Jazz: `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260808000000" --end-timestamp="20260812120000" --sleep="5" --batch-size="50"` for [[phab:T434688|T434688]] * 09:48 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host centrallog2002.codfw.wmnet * 09:48 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install4004.wikimedia.org * 09:42 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1232.eqiad.wmnet with reason: host reimage * 09:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install4004.wikimedia.org * 09:41 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host centrallog2002.codfw.wmnet * 09:39 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install5004.wikimedia.org * 09:36 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1232.eqiad.wmnet with reason: host reimage * 09:36 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'. * 09:34 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'. * 09:33 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1231.eqiad.wmnet with reason: host reimage * 09:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install5004.wikimedia.org * 09:32 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host centrallog1002.eqiad.wmnet * 09:30 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install6003.wikimedia.org * 09:30 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1231.eqiad.wmnet with reason: host reimage * 09:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1152.eqiad.wmnet with reason: host reimage * 09:25 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host centrallog1002.eqiad.wmnet * 09:25 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1152.eqiad.wmnet with reason: host reimage * 09:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install6003.wikimedia.org * 09:22 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1232 * 09:22 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1232 * 09:22 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1232 * 09:22 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1232.eqiad.wmnet 25.53.64.10.in-addr.arpa 5.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:22 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host titan1001.eqiad.wmnet * 09:22 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1232.eqiad.wmnet 25.53.64.10.in-addr.arpa 5.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:22 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:22 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1232 - btullis@cumin1003" * 09:22 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1232 - btullis@cumin1003" * 09:17 btullis@cumin1003: START - Cookbook sre.dns.netbox * 09:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install7002.wikimedia.org * 09:17 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1232 * 09:16 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1231 * 09:16 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1231 * 09:14 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1231 * 09:14 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1231.eqiad.wmnet 24.53.64.10.in-addr.arpa 4.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host titan1001.eqiad.wmnet * 09:14 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1231.eqiad.wmnet 24.53.64.10.in-addr.arpa 4.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:14 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1231 - btullis@cumin1003" * 09:14 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1231 - btullis@cumin1003" * 09:10 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install7002.wikimedia.org * 09:09 btullis@cumin1003: START - Cookbook sre.dns.netbox * 09:08 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1231 * 09:08 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1152 * 09:08 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1152 * 09:06 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1152 * 09:06 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1152.eqiad.wmnet 16.53.64.10.in-addr.arpa 6.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:06 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1152.eqiad.wmnet 16.53.64.10.in-addr.arpa 6.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:06 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:06 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1152 - btullis@cumin1003" * 09:06 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1152 - btullis@cumin1003" * 09:03 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1151.eqiad.wmnet * 09:03 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:03 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1151.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 08:55 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host titan2001.codfw.wmnet * 08:55 btullis@cumin1003: START - Cookbook sre.dns.netbox * 08:54 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1151.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 08:54 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1152 * 08:53 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1232.eqiad.wmnet with OS bookworm * 08:53 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1231.eqiad.wmnet with OS bookworm * 08:53 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1152.eqiad.wmnet with OS bookworm * 08:51 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2209: Pool back * 08:51 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1201.eqiad.wmnet * 08:50 btullis@cumin1003: START - Cookbook sre.hosts.remove-downtime for an-worker1201.eqiad.wmnet * 08:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1229.eqiad.wmnet with OS bookworm * 08:47 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host titan2001.codfw.wmnet * 08:46 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping2004.codfw.wmnet * 08:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ping2004.codfw.wmnet * 08:40 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host an-worker1230.eqiad.wmnet with OS bookworm * 08:40 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 08:34 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1151.eqiad.wmnet * 08:34 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 08:31 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host titan1002.eqiad.wmnet * 08:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1229.eqiad.wmnet with reason: host reimage * 08:25 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host titan1002.eqiad.wmnet * 08:25 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1229.eqiad.wmnet with reason: host reimage * 08:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping1004.eqiad.wmnet * 08:21 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ping1004.eqiad.wmnet * 08:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1230.eqiad.wmnet with reason: host reimage * 08:12 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1230.eqiad.wmnet with reason: host reimage * 08:11 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1229 * 08:11 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1229 * 08:11 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1229 * 08:11 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1229.eqiad.wmnet 22.53.64.10.in-addr.arpa 2.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:11 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1229.eqiad.wmnet 22.53.64.10.in-addr.arpa 2.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:11 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:11 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1229 - btullis@cumin1003" * 08:11 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1229 - btullis@cumin1003" * 08:10 btullis@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1201.eqiad.wmnet with reason: Fixing a disk * 08:07 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host titan2002.codfw.wmnet * 08:05 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2209: Pool back * 08:05 btullis@cumin1003: START - Cookbook sre.dns.netbox * 08:00 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host titan2002.codfw.wmnet * 07:59 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1229 * 07:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1230 * 07:58 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1230 * 07:55 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1230 * 07:55 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1230.eqiad.wmnet 23.53.64.10.in-addr.arpa 3.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:55 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1230.eqiad.wmnet 23.53.64.10.in-addr.arpa 3.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:55 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:55 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1230 - btullis@cumin1003" * 07:55 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1230 - btullis@cumin1003" * 07:51 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host kubestagemaster2005.codfw.wmnet with OS trixie * 07:48 btullis@cumin1003: START - Cookbook sre.dns.netbox * 07:41 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1230 * 07:41 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1229.eqiad.wmnet with OS bookworm * 07:41 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1230.eqiad.wmnet with OS bookworm * 07:39 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1152 from dbctl [[phab:T434480|T434480]]', diff saved to https://phabricator.wikimedia.org/P96090 and previous config saved to /var/cache/conftool/dbconfig/20260814-073941-marostegui.json * 07:29 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on kubestagemaster2005.codfw.wmnet with reason: host reimage * 07:23 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on kubestagemaster2005.codfw.wmnet with reason: host reimage * 07:04 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host kubestagemaster2005.codfw.wmnet with OS trixie * 06:53 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1228.eqiad.wmnet with OS bookworm * 06:44 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1227.eqiad.wmnet with OS bookworm * 06:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1209.eqiad.wmnet with OS bookworm * 06:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1175.eqiad.wmnet with OS bookworm * 06:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1228.eqiad.wmnet with reason: host reimage * 06:27 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1228.eqiad.wmnet with reason: host reimage * 06:25 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1227.eqiad.wmnet with reason: host reimage * 06:21 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1227.eqiad.wmnet with reason: host reimage * 06:18 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1209.eqiad.wmnet with reason: host reimage * 06:14 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1209.eqiad.wmnet with reason: host reimage * 06:14 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1228 * 06:14 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1228 * 06:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1175.eqiad.wmnet with reason: host reimage * 06:12 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1228 * 06:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1228.eqiad.wmnet 20.53.64.10.in-addr.arpa 0.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:12 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1228.eqiad.wmnet 20.53.64.10.in-addr.arpa 0.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1228 - ryankemper@cumin2003" * 06:12 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1228 - ryankemper@cumin2003" * 06:09 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1175.eqiad.wmnet with reason: host reimage * 06:07 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 06:07 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1228 * 06:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1227 * 06:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1227 * 06:06 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1227 * 06:06 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1227.eqiad.wmnet 19.53.64.10.in-addr.arpa 9.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:06 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1227.eqiad.wmnet 19.53.64.10.in-addr.arpa 9.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:06 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:06 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1227 - ryankemper@cumin2003" * 06:06 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1227 - ryankemper@cumin2003" * 06:00 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 06:00 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1227 * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1209 * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1209 * 06:00 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1209 * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1209.eqiad.wmnet 15.53.64.10.in-addr.arpa 5.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:00 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1209.eqiad.wmnet 15.53.64.10.in-addr.arpa 5.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1209 - ryankemper@cumin2003" * 06:00 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1209 - ryankemper@cumin2003" * 05:54 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 05:54 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1209 * 05:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1175 * 05:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1175 * 05:52 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1175 * 05:52 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1175.eqiad.wmnet 17.53.64.10.in-addr.arpa 7.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 05:52 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1175.eqiad.wmnet 17.53.64.10.in-addr.arpa 7.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 05:52 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 05:52 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1175 - ryankemper@cumin2003" * 05:52 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1175 - ryankemper@cumin2003" * 05:49 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1228.eqiad.wmnet with OS bookworm * 05:49 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1227.eqiad.wmnet with OS bookworm * 05:48 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1209.eqiad.wmnet with OS bookworm * 05:47 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 05:47 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1175 * 05:47 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1175.eqiad.wmnet with OS bookworm * 05:09 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host apifeatureusage1001.eqiad.wmnet with OS bookworm * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 03s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:10 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 01:07 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 01:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 01:02 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 00:59 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1226.eqiad.wmnet with OS bookworm * 00:47 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1225.eqiad.wmnet with OS bookworm * 00:41 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1224.eqiad.wmnet with OS bookworm * 00:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1226.eqiad.wmnet with reason: host reimage * 00:31 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1226.eqiad.wmnet with reason: host reimage * 00:28 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1225.eqiad.wmnet with reason: host reimage * 00:25 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1225.eqiad.wmnet with reason: host reimage * 00:19 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1224.eqiad.wmnet with reason: host reimage * 00:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1226 * 00:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1226 * 00:17 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1226 * 00:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1226.eqiad.wmnet 23.36.64.10.in-addr.arpa 3.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:17 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1226.eqiad.wmnet 23.36.64.10.in-addr.arpa 3.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 00:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1226 - ryankemper@cumin2003" * 00:17 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1226 - ryankemper@cumin2003" * 00:16 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1224.eqiad.wmnet with reason: host reimage * 00:12 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 00:12 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1226 * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1225 * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1225 * 00:10 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1225 * 00:10 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1225.eqiad.wmnet 22.36.64.10.in-addr.arpa 2.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:10 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1225.eqiad.wmnet 22.36.64.10.in-addr.arpa 2.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:10 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 00:10 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1225 - ryankemper@cumin2003" * 00:10 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1225 - ryankemper@cumin2003" * 00:03 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 00:02 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1225 * 00:02 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1224 * 00:02 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1224 * 00:00 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1224 * 00:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1224.eqiad.wmnet 21.36.64.10.in-addr.arpa 1.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:00 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1224.eqiad.wmnet 21.36.64.10.in-addr.arpa 1.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 00:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1224 - ryankemper@cumin2003" * 00:00 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1224 - ryankemper@cumin2003" == 2026-08-13 == * 23:54 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1226.eqiad.wmnet with OS bookworm * 23:53 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1225.eqiad.wmnet with OS bookworm * 23:52 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 23:51 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1224 * 23:51 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1224.eqiad.wmnet with OS bookworm * 23:47 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1223.eqiad.wmnet with OS bookworm * 23:29 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1223.eqiad.wmnet with reason: host reimage * 23:24 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1223.eqiad.wmnet with reason: host reimage * 23:19 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325525{{!}}ve.ui.CodeMirror.less: ensure normal font style]] (duration: 11m 40s) * 23:16 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260807000000" --end-timestamp="20260808000000" --sleep="5" --batch-size="50"` for [[phab:T434688|T434688]] * 23:13 musikanimal@deploy1003: musikanimal: Continuing with deployment * 23:11 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1325525{{!}}ve.ui.CodeMirror.less: ensure normal font style]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1223 * 23:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1223 * 23:08 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1325525{{!}}ve.ui.CodeMirror.less: ensure normal font style]] * 23:07 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1223 * 23:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1223.eqiad.wmnet 20.36.64.10.in-addr.arpa 0.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:07 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1223.eqiad.wmnet 20.36.64.10.in-addr.arpa 0.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 23:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1223 - ryankemper@cumin2003" * 23:03 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1223 - ryankemper@cumin2003" * 23:01 sbassett: Deployed security updates for [[phab:T430596|T430596]], [[phab:T120386|T120386]] * 22:55 bking@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host kubestagemaster2005.codfw.wmnet * 22:55 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host kubestagemaster2005.codfw.wmnet with OS bookworm * 22:54 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 22:53 ryankemper@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 22:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on kubestagemaster2005.codfw.wmnet with reason: host reimage * 22:49 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1222.eqiad.wmnet with OS bookworm * 22:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on kubestagemaster2005.codfw.wmnet with reason: host reimage * 22:29 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1222.eqiad.wmnet with reason: host reimage * 22:26 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1222.eqiad.wmnet with reason: host reimage * 22:23 bking@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host aux-k8s-etcd2003.codfw.wmnet * 22:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd2003.codfw.wmnet with OS bookworm * 22:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host kubestagemaster2005.codfw.wmnet with OS bookworm * 22:23 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM kubestagemaster2005.codfw.wmnet - bking@cumin2003" * 22:23 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM kubestagemaster2005.codfw.wmnet - bking@cumin2003" * 22:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) kubestagemaster2005.codfw.wmnet on all recursors * 22:22 bking@cumin2003: START - Cookbook sre.dns.wipe-cache kubestagemaster2005.codfw.wmnet on all recursors * 22:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:22 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM kubestagemaster2005.codfw.wmnet - bking@cumin2003" * 22:22 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM kubestagemaster2005.codfw.wmnet - bking@cumin2003" * 22:17 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 22:13 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1223 * 22:12 bking@cumin2003: START - Cookbook sre.dns.netbox * 22:12 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host kubestagemaster2005.codfw.wmnet * 22:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1222 * 22:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1222 * 22:11 sbassett: Deployed security updates for [[phab:T429244|T429244]], [[phab:T434039|T434039]] * 22:11 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1222 * 22:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1222.eqiad.wmnet 19.36.64.10.in-addr.arpa 9.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:11 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1222.eqiad.wmnet 19.36.64.10.in-addr.arpa 9.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1222 - ryankemper@cumin2003" * 22:07 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1222 - ryankemper@cumin2003" * 22:05 robh@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-wdqs2001.codfw.wmnet with reason: updating firmware * 22:01 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 22:01 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1223.eqiad.wmnet with OS bookworm * 22:01 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1222 * 22:01 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1222.eqiad.wmnet with OS bookworm * 22:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1212.eqiad.wmnet with OS bookworm * 21:54 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324295{{!}}Revert "Lazily reject pre-fix parser-cache entries for noreferrer/noopener links" (T429090)]] (duration: 06m 42s) * 21:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd2003.codfw.wmnet with reason: host reimage * 21:53 bking@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host dse-k8s-etcd2001.codfw.wmnet * 21:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dse-k8s-etcd2001.codfw.wmnet with OS bookworm * 21:50 sbassett@deploy1003: sbassett, kharlan: Continuing with deployment * 21:49 sbassett@deploy1003: sbassett, kharlan: Backport for [[gerrit:1324295{{!}}Revert "Lazily reject pre-fix parser-cache entries for noreferrer/noopener links" (T429090)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:48 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aux-k8s-etcd2003.codfw.wmnet with reason: host reimage * 21:47 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1324295{{!}}Revert "Lazily reject pre-fix parser-cache entries for noreferrer/noopener links" (T429090)]] * 21:42 maryum: Deployed security patch for [[phab:T434549|T434549]] * 21:39 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1212.eqiad.wmnet with reason: host reimage * 21:34 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1212.eqiad.wmnet with reason: host reimage * 21:33 bking@cumin2003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd2003.codfw.wmnet with OS bookworm * 21:32 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM aux-k8s-etcd2003.codfw.wmnet - bking@cumin2003" * 21:32 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM aux-k8s-etcd2003.codfw.wmnet - bking@cumin2003" * 21:32 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) aux-k8s-etcd2003.codfw.wmnet on all recursors * 21:32 bking@cumin2003: START - Cookbook sre.dns.wipe-cache aux-k8s-etcd2003.codfw.wmnet on all recursors * 21:32 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:32 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM aux-k8s-etcd2003.codfw.wmnet - bking@cumin2003" * 21:31 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM aux-k8s-etcd2003.codfw.wmnet - bking@cumin2003" * 21:28 maryum: Deployed security patch for [[phab:T434619|T434619]] * 21:26 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:26 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host aux-k8s-etcd2003.codfw.wmnet * 21:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-etcd2001.codfw.wmnet with reason: host reimage * 21:20 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1212 * 21:20 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1212 * 21:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1212.eqiad.wmnet with OS bookworm * 21:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on dse-k8s-etcd2001.codfw.wmnet with reason: host reimage * 21:04 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325553{{!}}InstrumentConstructiveEdits: exclude mw-reverted as well (T431493)]] (duration: 06m 25s) * 20:59 kemayo@deploy1003: kemayo: Continuing with deployment * 20:59 kemayo@deploy1003: kemayo: Backport for [[gerrit:1325553{{!}}InstrumentConstructiveEdits: exclude mw-reverted as well (T431493)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host dse-k8s-etcd2001.codfw.wmnet with OS bookworm * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 20:57 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 20:57 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1325553{{!}}InstrumentConstructiveEdits: exclude mw-reverted as well (T431493)]] * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-etcd2001.codfw.wmnet on all recursors * 20:57 bking@cumin2003: START - Cookbook sre.dns.wipe-cache dse-k8s-etcd2001.codfw.wmnet on all recursors * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 20:57 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 20:54 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320163{{!}}Add configurable RestTermsOfServiceUrl (T428147)]] (duration: 21m 39s) * 20:53 bking@cumin2003: START - Cookbook sre.dns.netbox * 20:53 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host dse-k8s-etcd2001.codfw.wmnet * 20:50 samtar@deploy1003: samtar, milazg: Continuing with deployment * 20:35 samtar@deploy1003: samtar, milazg: Backport for [[gerrit:1320163{{!}}Add configurable RestTermsOfServiceUrl (T428147)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:33 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1320163{{!}}Add configurable RestTermsOfServiceUrl (T428147)]] * 20:30 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325549{{!}}Deploy PRV to several LC wikis (T423785)]] (duration: 06m 57s) * 20:26 arlolra@deploy1003: arlolra: Continuing with deployment * 20:25 arlolra@deploy1003: arlolra: Backport for [[gerrit:1325549{{!}}Deploy PRV to several LC wikis (T423785)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:24 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 20:23 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1325549{{!}}Deploy PRV to several LC wikis (T423785)]] * 20:21 ariel@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319906{{!}}Remove boilerplate language from wmf-rest and wmf-math API modules (T433736)]] (duration: 13m 54s) * 20:14 ariel@deploy1003: ariel: Continuing with deployment * 20:11 ariel@deploy1003: ariel: Backport for [[gerrit:1319906{{!}}Remove boilerplate language from wmf-rest and wmf-math API modules (T433736)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 ariel@deploy1003: Started scap sync-world: Backport for [[gerrit:1319906{{!}}Remove boilerplate language from wmf-rest and wmf-math API modules (T433736)]] * 19:58 robh@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-wdqs2001.codfw.wmnet with reason: updating firmware * 19:54 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325551{{!}}Render the focused module view as a full-screen page (T433896)]] (duration: 30m 37s) * 19:52 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching aqs[2001,1016]*: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 19:44 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching aqs[2001,1016]*: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 19:42 musikanimal@deploy1003: musikanimal: Continuing with deployment * 19:41 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1325551{{!}}Render the focused module view as a full-screen page (T433896)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:34 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: quash java safepoint logspam - bking@cumin2003 - [[phab:T434685|T434685]] * 19:34 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 19:34 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 19:24 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1325551{{!}}Render the focused module view as a full-screen page (T433896)]] * 19:20 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 19:20 jhancock@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin1003" * 19:18 jhancock@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin1003" * 19:14 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325548{{!}}Enable image lazy loading on desktop in group1 (T148047)]] (duration: 07m 43s) * 19:10 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 19:10 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1325548{{!}}Enable image lazy loading on desktop in group1 (T148047)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:07 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1325548{{!}}Enable image lazy loading on desktop in group1 (T148047)]] * 19:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1166.eqiad.wmnet onto db1280.eqiad.wmnet * 19:03 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 19:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1166: Pool db1166.eqiad.wmnet in after cloning * 18:59 jhancock@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 18:58 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1234.eqiad.wmnet with OS bookworm * 18:52 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 18:51 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:49 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:46 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 18:46 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 18:44 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=dns3004.* * 18:39 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1221.eqiad.wmnet with OS bookworm * 18:36 inflatador: [bking@ganeti2048] ~$ sudo gnt-instance replace-disks -n ganeti2030.codfw.wmnet aux-k8s-worker2002.codfw.wmnet [[phab:T434681|T434681]] * 18:35 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 18:29 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1220.eqiad.wmnet with OS bookworm * 18:29 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1234.eqiad.wmnet with reason: host reimage * 18:27 bking@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host dse-k8s-etcd2001.codfw.wmnet * 18:27 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-etcd2001.codfw.wmnet on all recursors * 18:27 bking@cumin2003: START - Cookbook sre.dns.wipe-cache dse-k8s-etcd2001.codfw.wmnet on all recursors * 18:27 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:27 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 18:27 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 18:25 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1234.eqiad.wmnet with reason: host reimage * 18:25 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 18:19 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1221.eqiad.wmnet with reason: host reimage * 18:17 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1166: Pool db1166.eqiad.wmnet in after cloning * 18:15 brennen@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 18:15 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1221.eqiad.wmnet with reason: host reimage * 18:14 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:12 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:12 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 18:12 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-etcd2001.codfw.wmnet on all recursors * 18:12 bking@cumin2003: START - Cookbook sre.dns.wipe-cache dse-k8s-etcd2001.codfw.wmnet on all recursors * 18:12 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:12 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 18:12 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 18:10 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: quash java safepoint logspam - bking@cumin2003 - [[phab:T434685|T434685]] * 18:10 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1234 * 18:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1234 * 18:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1220.eqiad.wmnet with reason: host reimage * 18:07 brennen: 1.47.0-wmf.15 train status ([[phab:T430834|T430834]]) - no current blockers, rolling to all wikis * 18:07 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1234 * 18:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1234.eqiad.wmnet 10.36.64.10.in-addr.arpa 0.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:07 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:07 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1234.eqiad.wmnet 10.36.64.10.in-addr.arpa 0.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1234 - ryankemper@cumin2003" * 18:07 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1234 - ryankemper@cumin2003" * 18:05 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1220.eqiad.wmnet with reason: host reimage * 18:04 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host dse-k8s-etcd2001.codfw.wmnet * 18:04 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 18:03 dancy@deploy1003: Installation of scap version "4.280.1" completed for 3 hosts * 18:02 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:02 inflatador: bking@dse-k8s-etcd2002 etcdctl member remove $<nowiki>{</nowiki>UUID of dse-k8s-etcd2001<nowiki>}</nowiki> [[phab:T434681|T434681]] [[phab:T434793|T434793]] * 18:01 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 18:01 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1234 * 18:01 dancy@deploy1003: Installing scap version "4.280.1" for 3 host(s) * 18:01 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1221 * 18:01 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1221 * 18:00 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1221 * 18:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1221.eqiad.wmnet 18.36.64.10.in-addr.arpa 8.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:00 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1221.eqiad.wmnet 18.36.64.10.in-addr.arpa 8.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1221 - ryankemper@cumin2003" * 17:59 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:58 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 17:58 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:57 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 17:57 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:56 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1221 - ryankemper@cumin2003" * 17:53 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 17:52 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 17:52 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 17:51 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1221 * 17:51 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1220 * 17:51 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1220 * 17:51 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:51 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:51 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1220 * 17:51 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1220.eqiad.wmnet 11.36.64.10.in-addr.arpa 1.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:51 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1220.eqiad.wmnet 11.36.64.10.in-addr.arpa 1.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:51 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1220 - ryankemper@cumin2003" * 17:50 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1220 - ryankemper@cumin2003" * 17:47 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:47 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 17:46 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1234.eqiad.wmnet with OS bookworm * 17:46 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 17:45 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1221.eqiad.wmnet with OS bookworm * 17:45 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1220 * 17:45 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1220.eqiad.wmnet with OS bookworm * 17:43 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 17:41 bking@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host dse-k8s-etcd2001.codfw.wmnet * 17:41 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host dse-k8s-etcd2001.codfw.wmnet * 17:40 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:40 inflatador: bking@ganeti2048] `sudo gnt-instance remove --force --ignore-failures --shutdown-timeout=0` on non-DRBD VMs [[phab:T434681|T434681]] * 17:40 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:39 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:38 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:36 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:32 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:32 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:28 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 17:26 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:26 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:24 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:23 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:21 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1218.eqiad.wmnet with OS bookworm * 17:20 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:20 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:19 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1179.eqiad.wmnet with OS bookworm * 17:18 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1150.eqiad.wmnet with OS bookworm * 17:18 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns3004.wikimedia.org with OS trixie * 17:15 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:14 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:13 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:12 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:12 swfrench@deploy1003: Finished scap sync-world: Helmfile-only deployment for mediawiki chart bump - [[phab:T427666|T427666]] (duration: 03m 03s) * 17:09 swfrench@deploy1003: Started scap sync-world: Helmfile-only deployment for mediawiki chart bump - [[phab:T427666|T427666]] * 17:02 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:01 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:01 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1218.eqiad.wmnet with reason: host reimage * 17:01 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:01 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:00 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324810{{!}}deployment-info.php: Report dbname and branch for the requested wiki (T434726)]] (duration: 06m 52s) * 16:58 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 16:57 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1179.eqiad.wmnet with reason: host reimage * 16:56 dancy@deploy1003: dancy: Continuing with deployment * 16:56 dancy@deploy1003: dancy: Backport for [[gerrit:1324810{{!}}deployment-info.php: Report dbname and branch for the requested wiki (T434726)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1150.eqiad.wmnet with reason: host reimage * 16:53 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324810{{!}}deployment-info.php: Report dbname and branch for the requested wiki (T434726)]] * 16:51 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1166: Depool db1166.eqiad.wmnet to then clone it to db1280.eqiad.wmnet - cwilliams@cumin1003 * 16:50 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1166: Depool db1166.eqiad.wmnet to then clone it to db1280.eqiad.wmnet - cwilliams@cumin1003 * 16:50 cwilliams@cumin1003: START - Cookbook sre.mysql.clone of db1166.eqiad.wmnet onto db1280.eqiad.wmnet * 16:49 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1218.eqiad.wmnet with reason: host reimage * 16:48 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1179.eqiad.wmnet with reason: host reimage * 16:47 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1150.eqiad.wmnet with reason: host reimage * 16:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.upgrade (exit_code=0) for 1 hosts * 16:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2220: Upgrade of db2220.codfw.wmnet completed * 16:38 swfrench@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 16:38 swfrench-wmf: kubectl delete node kubestagemaster2005.codfw.wmnet - [[phab:T434681|T434681]] * 16:34 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1218 * 16:34 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1218 * 16:34 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1218.eqiad.wmnet with OS bookworm * 16:34 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1179 * 16:34 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1179 * 16:33 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1179.eqiad.wmnet with OS bookworm * 16:32 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1150 * 16:32 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1150 * 16:31 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1150.eqiad.wmnet with OS bookworm * 16:24 dancy@deploy1003: Installation of scap version "4.280.0" completed for 3 hosts * 16:22 dancy@deploy1003: Installing scap version "4.280.0" for 3 host(s) * 16:20 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: quash java safepoint logspam - bking@cumin2003 - [[phab:T434685|T434685]] * 16:14 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns3004.wikimedia.org with reason: host reimage * 16:08 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 16:07 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 16:07 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns3004.wikimedia.org with reason: host reimage * 16:05 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 16:04 swfrench@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 16:00 swfrench@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 15:59 swfrench@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 15:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Upgrade of db2220.codfw.wmnet completed * 15:48 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2220: Upgrading db2220.codfw.wmnet * 15:48 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2220: Upgrading db2220.codfw.wmnet * 15:48 cwilliams@cumin1003: START - Cookbook sre.mysql.upgrade for 1 hosts * 15:46 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns3004.wikimedia.org with OS trixie * 15:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2220 [[phab:T434802|T434802]]', diff saved to https://phabricator.wikimedia.org/P96079 and previous config saved to /var/cache/conftool/dbconfig/20260813-154624-cwilliams.json * 15:45 cdobbins@cumin1003: conftool action : set/pooled=no; selector: name=dns3004.* * 15:44 cjd91: depooling dns3004 to reimage to trixie * 15:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2159 to s7 primary [[phab:T434802|T434802]]', diff saved to https://phabricator.wikimedia.org/P96078 and previous config saved to /var/cache/conftool/dbconfig/20260813-154405-cwilliams.json * 15:43 cezmunsta: Starting s7 codfw failover from db2220 to db2159 - [[phab:T434802|T434802]] * 15:41 cgoubert@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: Dragonfly supernodes reboot (duration: 09m 42s) * 15:41 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dragonfly-supernode2001.codfw.wmnet * 15:39 inflatador: bking@ganeti2048] ~$ sudo gnt-node failover -f --ignore-consistency ganeti2046.codfw.wmnet [[phab:T434681|T434681]] * 15:39 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 15:39 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 15:39 swfrench@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 15:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2159 with weight 0 [[phab:T434802|T434802]]', diff saved to https://phabricator.wikimedia.org/P96077 and previous config saved to /var/cache/conftool/dbconfig/20260813-153806-cwilliams.json * 15:37 swfrench@dns1004: END - running authdns-update * 15:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 30 hosts with reason: Primary switchover s7 [[phab:T434802|T434802]] * 15:37 cgoubert@cumin2003: START - Cookbook sre.hosts.reboot-single for host dragonfly-supernode2001.codfw.wmnet * 15:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dragonfly-supernode1001.eqiad.wmnet * 15:35 swfrench@dns1004: START - running authdns-update * 15:32 cgoubert@cumin2003: START - Cookbook sre.hosts.reboot-single for host dragonfly-supernode1001.eqiad.wmnet * 15:32 cgoubert@deploy1003: Locking from deployment [ALL REPOSITORIES]: Dragonfly supernodes reboot * 15:30 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-master-codfw * 15:30 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2005.codfw.wmnet * 15:30 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2005.codfw.wmnet * 15:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1169.eqiad.wmnet onto db1277.eqiad.wmnet * 15:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1169: Pool db1169.eqiad.wmnet in after cloning * 15:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2005.codfw.wmnet * 15:23 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2005.codfw.wmnet * 15:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2004.codfw.wmnet * 15:23 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2004.codfw.wmnet * 15:18 inflatador: bking@ganeti2048 sudo gnt-node failover -f ganeti2046.codfw.wmnet [[phab:T434681|T434681]] * 15:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2004.codfw.wmnet * 15:17 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2004.codfw.wmnet * 15:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2003.codfw.wmnet * 15:17 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2003.codfw.wmnet * 15:15 cdobbins@dns1004: END - running authdns-update * 15:13 cdobbins@dns1004: START - running authdns-update * 15:10 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1217.eqiad.wmnet with OS bookworm * 15:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2003.codfw.wmnet * 15:10 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2003.codfw.wmnet * 15:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2002.codfw.wmnet * 15:10 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2002.codfw.wmnet * 15:10 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: quash java safepoint logspam - bking@cumin2003 - [[phab:T434685|T434685]] * 15:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1216.eqiad.wmnet with OS bookworm * 15:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2002.codfw.wmnet * 15:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2002.codfw.wmnet * 15:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2001.codfw.wmnet * 15:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2001.codfw.wmnet * 15:01 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1042.eqiad.wmnet * 15:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1042.eqiad.wmnet * 15:00 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325455{{!}}Api: Use correct query when continue prop=categories (T433922)]] (duration: 09m 47s) * 14:59 bking@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host dse-k8s-etcd2004.codfw.wmnet * 14:58 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-etcd2004.codfw.wmnet on all recursors * 14:58 bking@cumin2003: START - Cookbook sre.dns.wipe-cache dse-k8s-etcd2004.codfw.wmnet on all recursors * 14:58 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:58 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM dse-k8s-etcd2004.codfw.wmnet - bking@cumin2003" * 14:58 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM dse-k8s-etcd2004.codfw.wmnet - bking@cumin2003" * 14:55 zabe@deploy1003: zabe: Continuing with deployment * 14:53 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-etcd2004.codfw.wmnet on all recursors * 14:53 bking@cumin2003: START - Cookbook sre.dns.wipe-cache dse-k8s-etcd2004.codfw.wmnet on all recursors * 14:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:53 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2004.codfw.wmnet - bking@cumin2003" * 14:53 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2004.codfw.wmnet - bking@cumin2003" * 14:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2001.codfw.wmnet * 14:52 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2001.codfw.wmnet * 14:52 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-master-codfw * 14:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2001.codfw.wmnet * 14:52 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2001.codfw.wmnet * 14:52 zabe@deploy1003: zabe: Backport for [[gerrit:1325455{{!}}Api: Use correct query when continue prop=categories (T433922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:50 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1325455{{!}}Api: Use correct query when continue prop=categories (T433922)]] * 14:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1217.eqiad.wmnet with reason: host reimage * 14:48 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:48 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host dse-k8s-etcd2004.codfw.wmnet * 14:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-master-eqiad * 14:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1006.eqiad.wmnet * 14:44 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1006.eqiad.wmnet * 14:44 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1216.eqiad.wmnet with reason: host reimage * 14:40 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1217.eqiad.wmnet with reason: host reimage * 14:39 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1169: Pool db1169.eqiad.wmnet in after cloning * 14:39 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1216.eqiad.wmnet with reason: host reimage * 14:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl1005.eqiad.wmnet * 14:32 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl1005.eqiad.wmnet * 14:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1004.eqiad.wmnet * 14:32 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1004.eqiad.wmnet * 14:31 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus2008.codfw.wmnet * 14:31 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor1003.eqiad.wmnet * 14:29 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1042.eqiad.wmnet * 14:27 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor1003.eqiad.wmnet * 14:26 moritzm: installing Django security updates * 14:26 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling reboot on A:wikidough * 14:25 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1217.eqiad.wmnet with OS bookworm * 14:25 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1216.eqiad.wmnet with OS bookworm * 14:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl1004.eqiad.wmnet * 14:24 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl1004.eqiad.wmnet * 14:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1003.eqiad.wmnet * 14:24 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1003.eqiad.wmnet * 14:23 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus2008.codfw.wmnet * 14:23 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus1008.eqiad.wmnet * 14:23 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor-dev2001.codfw.wmnet * 14:22 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus2006.codfw.wmnet * 14:19 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor-dev2001.codfw.wmnet * 14:18 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1042.eqiad.wmnet * 14:17 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.* * 14:16 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl1003.eqiad.wmnet * 14:16 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl1003.eqiad.wmnet * 14:16 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1002.eqiad.wmnet * 14:16 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1002.eqiad.wmnet * 14:15 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus1008.eqiad.wmnet * 14:14 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus1006.eqiad.wmnet * 14:12 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus2006.codfw.wmnet * 14:12 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor2003.codfw.wmnet * 14:11 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus2007.codfw.wmnet * 14:11 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1041.eqiad.wmnet * 14:11 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1041.eqiad.wmnet * 14:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl1002.eqiad.wmnet * 14:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl1002.eqiad.wmnet * 14:09 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-master-eqiad * 14:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on apifeatureusage1001.eqiad.wmnet with reason: host reimage * 14:08 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1215.eqiad.wmnet with OS bookworm * 14:08 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor2003.codfw.wmnet * 14:08 jayme: updated calico to v3.30.7 on wikikube codfw [[phab:T427400|T427400]] * 14:07 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1149.eqiad.wmnet with OS bookworm * 14:06 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'. * 14:06 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1041.eqiad.wmnet * 14:04 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus1006.eqiad.wmnet * 14:03 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus2007.codfw.wmnet * 14:03 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus2005.codfw.wmnet * 14:03 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus1007.eqiad.wmnet * 14:03 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1002.eqiad.wmnet * 14:02 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on apifeatureusage1001.eqiad.wmnet with reason: host reimage * 14:02 cgoubert@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=helm-charts.*,name=eqiad * 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host chartmuseum1001.eqiad.wmnet * 14:00 moritzm: installing libxml2 security updates * 13:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1214.eqiad.wmnet with OS bookworm * 13:59 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns5003.* * 13:59 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1041.eqiad.wmnet * 13:59 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow1002.eqiad.wmnet * 13:58 cmooney@dns3003: END - running authdns-update * 13:57 cgoubert@cumin2003: START - Cookbook sre.hosts.reboot-single for host chartmuseum1001.eqiad.wmnet * 13:57 cgoubert@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=helm-charts.*,name=eqiad * 13:57 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: cloudelastic cluster restart - bking@cumin2003 * 13:57 cgoubert@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=helm-charts.*,name=codfw * 13:56 cmooney@dns3003: START - running authdns-update * 13:56 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'. * 13:56 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns5003.*,service=authdns-update * 13:55 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host chartmuseum2001.codfw.wmnet * 13:55 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus1007.eqiad.wmnet * 13:55 cmooney@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dns5003.wikimedia.org * 13:55 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus1005.eqiad.wmnet * 13:51 cgoubert@cumin2003: START - Cookbook sre.hosts.reboot-single for host chartmuseum2001.codfw.wmnet * 13:51 cgoubert@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=helm-charts.*,name=codfw * 13:51 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus2005.codfw.wmnet * 13:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host apifeatureusage1001.eqiad.wmnet with OS bookworm * 13:50 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus7002.magru.wmnet * 13:50 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts lvs1015.eqiad.wmnet * 13:50 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:50 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1015.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:49 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1015.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:49 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325480{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] (duration: 06m 39s) * 13:47 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1215.eqiad.wmnet with reason: host reimage * 13:46 cmooney@cumin1003: START - Cookbook sre.hosts.reboot-single for host dns5003.wikimedia.org * 13:46 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1003.eqiad.wmnet * 13:45 cmooney@cumin1003: conftool action : set/pooled=no; selector: name=dns5003.* * 13:45 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1040.eqiad.wmnet * 13:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1040.eqiad.wmnet * 13:45 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 13:45 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus1005.eqiad.wmnet * 13:45 stran@deploy1003: stran: Continuing with deployment * 13:44 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus7002.magru.wmnet * 13:44 stran@deploy1003: stran: Backport for [[gerrit:1325480{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:44 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus6002.drmrs.wmnet * 13:43 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260805000000" --end-timestamp="20260806000000" --sleep="5" --batch-size="10"` for [[phab:T434688|T434688]] * 13:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1149.eqiad.wmnet with reason: host reimage * 13:42 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1325480{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] * 13:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow1003.eqiad.wmnet * 13:40 sukhe@cumin1003: START - Cookbook sre.hosts.decommission for hosts lvs1015.eqiad.wmnet * 13:40 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts lvs1014.eqiad.wmnet * 13:40 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:40 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1014.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:40 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1040.eqiad.wmnet * 13:40 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1014.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:39 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1214.eqiad.wmnet with reason: host reimage * 13:38 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2004.codfw.wmnet * 13:38 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus6002.drmrs.wmnet * 13:38 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus5003.eqsin.wmnet * 13:35 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 13:35 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1149.eqiad.wmnet with reason: host reimage * 13:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1215.eqiad.wmnet with reason: host reimage * 13:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow2004.codfw.wmnet * 13:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1214.eqiad.wmnet with reason: host reimage * 13:33 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1040.eqiad.wmnet * 13:31 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus5003.eqsin.wmnet * 13:31 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: cloudelastic cluster restart - bking@cumin2003 * 13:31 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus4003.ulsfo.wmnet * 13:30 sukhe@cumin1003: START - Cookbook sre.hosts.decommission for hosts lvs1014.eqiad.wmnet * 13:30 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts lvs1013.eqiad.wmnet * 13:30 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:30 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1013.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:30 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1013.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:27 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1039.eqiad.wmnet * 13:27 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1039.eqiad.wmnet * 13:26 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1167.eqiad.wmnet onto db1281.eqiad.wmnet * 13:26 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1167: Pool db1167.eqiad.wmnet in after cloning * 13:25 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus4003.ulsfo.wmnet * 13:25 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260802000000" --end-timestamp="20260803000000" --sleep="5" --batch-size="10"` for [[phab:T434688|T434688]] * 13:24 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 13:24 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus3004.esams.wmnet * 13:24 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1274: New host * 13:24 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325476{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] (duration: 07m 19s) * 13:24 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260804000000" --end-timestamp="20260805000000" --sleep="5" --batch-size="10"` for [[phab:T434688|T434688]] * 13:24 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260803000000" --end-timestamp="20260804000000" --sleep="5" --batch-size="10"` for [[phab:T434688|T434688]] * 13:23 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2003.codfw.wmnet * 13:22 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1039.eqiad.wmnet * 13:20 sukhe@cumin1003: START - Cookbook sre.hosts.decommission for hosts lvs1013.eqiad.wmnet * 13:20 stran@deploy1003: stran: Continuing with deployment * 13:19 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow2003.codfw.wmnet * 13:19 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling reboot on A:wikidough * 13:19 stran@deploy1003: stran: Backport for [[gerrit:1325476{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:18 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus3004.esams.wmnet * 13:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1215.eqiad.wmnet with OS bookworm * 13:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1214.eqiad.wmnet with OS bookworm * 13:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1149.eqiad.wmnet with OS bookworm * 13:17 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1325476{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] * 13:11 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324731{{!}}prv: Enable parsoid rendering for 5 wikisource wikis]] (duration: 07m 26s) * 13:11 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1039.eqiad.wmnet * 13:11 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1036.eqiad.wmnet * 13:11 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1036.eqiad.wmnet * 13:07 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow3004.esams.wmnet * 13:07 jgiannelos@deploy1003: jgiannelos: Continuing with deployment * 13:06 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1233.eqiad.wmnet with OS bookworm * 13:06 jgiannelos@deploy1003: jgiannelos: Backport for [[gerrit:1324731{{!}}prv: Enable parsoid rendering for 5 wikisource wikis]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:04 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1324731{{!}}prv: Enable parsoid rendering for 5 wikisource wikis]] * 13:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow3004.esams.wmnet * 13:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1036.eqiad.wmnet * 13:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1210.eqiad.wmnet with OS bookworm * 13:01 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1036.eqiad.wmnet * 13:00 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1052.eqiad.wmnet * 13:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1052.eqiad.wmnet * 12:58 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow4003.ulsfo.wmnet * 12:55 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1211.eqiad.wmnet with OS bookworm * 12:55 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1052.eqiad.wmnet * 12:52 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow4003.ulsfo.wmnet * 12:51 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1052.eqiad.wmnet * 12:49 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1169: Depool db1169.eqiad.wmnet to then clone it to db1277.eqiad.wmnet - cwilliams@cumin1003 * 12:46 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1169: Depool db1169.eqiad.wmnet to then clone it to db1277.eqiad.wmnet - cwilliams@cumin1003 * 12:46 cwilliams@cumin1003: START - Cookbook sre.mysql.clone of db1169.eqiad.wmnet onto db1277.eqiad.wmnet * 12:45 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:45 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns record for deleted IP reservations eqsin lvs vlan ints - cmooney@cumin1003" * 12:44 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns record for deleted IP reservations eqsin lvs vlan ints - cmooney@cumin1003" * 12:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1051.eqiad.wmnet * 12:43 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1051.eqiad.wmnet * 12:42 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1210.eqiad.wmnet with reason: host reimage * 12:41 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1167: Pool db1167.eqiad.wmnet in after cloning * 12:40 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 12:39 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1274: New host * 12:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Added db1274', diff saved to https://phabricator.wikimedia.org/P96062 and previous config saved to /var/cache/conftool/dbconfig/20260813-123907-cwilliams.json * 12:38 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast7002.wikimedia.org * 12:38 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1233.eqiad.wmnet with reason: host reimage * 12:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1051.eqiad.wmnet * 12:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1211.eqiad.wmnet with reason: host reimage * 12:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1233.eqiad.wmnet with reason: host reimage * 12:32 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1051.eqiad.wmnet * 12:32 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast7002.wikimedia.org * 12:32 marostegui: Drop SecurePoll tables from closed wikis [[phab:T423128|T423128]] * 12:32 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow5003.eqsin.wmnet * 12:31 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1050.eqiad.wmnet * 12:31 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1050.eqiad.wmnet * 12:29 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1210.eqiad.wmnet with reason: host reimage * 12:29 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1211.eqiad.wmnet with reason: host reimage * 12:29 moritzm: installing Wireshark security updates * 12:26 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow5003.eqsin.wmnet * 12:26 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1050.eqiad.wmnet * 12:24 cmooney@dns3003: END - running authdns-update * 12:21 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow6001.drmrs.wmnet * 12:21 cmooney@dns3003: START - running authdns-update * 12:20 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1050.eqiad.wmnet * 12:20 cgoubert@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host rdb-lock2003.codfw.wmnet * 12:20 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:20 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update netbox dns entries for expanded public1-603-eqsin subnet - cmooney@cumin1003" * 12:20 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update netbox dns entries for expanded public1-603-eqsin subnet - cmooney@cumin1003" * 12:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 12:20 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 12:19 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1049.eqiad.wmnet * 12:19 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1049.eqiad.wmnet * 12:19 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 12:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host moss-be1003.eqiad.wmnet * 12:17 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow6001.drmrs.wmnet * 12:17 moritzm: installin curl security updates * 12:15 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 12:15 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1233.eqiad.wmnet with OS bookworm * 12:15 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1210.eqiad.wmnet with OS bookworm * 12:15 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1211.eqiad.wmnet with OS bookworm * 12:13 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1049.eqiad.wmnet * 12:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow7002.magru.wmnet * 12:12 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-worker1178.eqiad.wmnet with OS bookworm * 12:10 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host moss-be1003.eqiad.wmnet * 12:10 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 12:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be1006.eqiad.wmnet * 12:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 12:10 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 12:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 12:10 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 12:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow7002.magru.wmnet * 12:08 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1049.eqiad.wmnet * 12:05 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 12:05 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2003.codfw.wmnet * 12:04 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'. * 12:04 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:03 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be1006.eqiad.wmnet * 12:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be1005.eqiad.wmnet * 12:01 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1038.eqiad.wmnet * 12:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1038.eqiad.wmnet * 12:01 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1213.eqiad.wmnet with OS bookworm * 11:56 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be1005.eqiad.wmnet * 11:55 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be1004.eqiad.wmnet * 11:54 cgoubert@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host rdb-lock2003.codfw.wmnet * 11:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 11:53 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 11:53 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:53 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 11:53 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 11:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1038.eqiad.wmnet * 11:49 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be1004.eqiad.wmnet * 11:44 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:44 moritzm: installing Linux 5.10.262 on Bullseye hosts * 11:41 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1038.eqiad.wmnet * 11:40 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 11:40 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 11:40 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 11:40 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:40 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 11:40 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 11:38 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1213.eqiad.wmnet with reason: host reimage * 11:36 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1167: Depool db1167.eqiad.wmnet to then clone it to db1281.eqiad.wmnet - marostegui@cumin1003 * 11:35 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 11:35 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2003.codfw.wmnet * 11:35 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1167: Depool db1167.eqiad.wmnet to then clone it to db1281.eqiad.wmnet - marostegui@cumin1003 * 11:35 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1167.eqiad.wmnet onto db1281.eqiad.wmnet * 11:34 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 22 hosts with reason: Cloning * 11:34 moritzm: remove ganeti3005 from esams03 cluster, hardware issues [[phab:T434646|T434646]] * 11:32 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1213.eqiad.wmnet with reason: host reimage * 11:28 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1034.eqiad.wmnet * 11:28 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1034.eqiad.wmnet * 11:22 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1034.eqiad.wmnet * 11:19 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1034.eqiad.wmnet * 11:17 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1213.eqiad.wmnet with OS bookworm * 11:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1165.eqiad.wmnet onto db1279.eqiad.wmnet * 11:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1165: Pool db1165.eqiad.wmnet in after cloning * 11:07 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply * 10:57 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply * 10:54 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply * 10:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1274.eqiad.wmnet with reason: Enabling notifications and pooling * 10:45 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1033.eqiad.wmnet * 10:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1033.eqiad.wmnet * 10:44 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply * 10:43 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'. * 10:42 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'. * 10:42 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'. * 10:40 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db[1216,1225,1239-1240].eqiad.wmnet with reason: reboot * 10:39 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1033.eqiad.wmnet * 10:38 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1151 from dbctl [[phab:T434538|T434538]]', diff saved to https://phabricator.wikimedia.org/P96055 and previous config saved to /var/cache/conftool/dbconfig/20260813-103828-marostegui.json * 10:35 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1033.eqiad.wmnet * 10:27 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1165: Pool db1165.eqiad.wmnet in after cloning * 10:24 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1151.eqiad.wmnet with OS bookworm * 10:15 moritzm: installing bind9 security updates (client-side tools/libs only) * 10:07 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-debug: apply * 10:06 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-debug: apply * 10:02 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 7 hosts with reason: reboot * 10:01 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-debug: apply * 10:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.upgrade (exit_code=0) for 1 hosts * 10:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2214: Upgrade of db2214.codfw.wmnet completed * 10:01 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-debug: apply * 10:00 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-debug: apply * 10:00 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-debug: apply * 09:59 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.decommission (exit_code=99) * 09:59 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 09:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1151.eqiad.wmnet with reason: host reimage * 09:59 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 8 hosts * 09:59 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 8 hosts * 09:55 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1151.eqiad.wmnet with reason: host reimage * 09:46 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260801000000" --end-timestamp="20260802000000" --sleep="3" --batch-size="5"` for [[phab:T434688|T434688]] * 09:45 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 8 hosts with reason: reboot * 09:45 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 22 hosts with reason: Cloning * 09:43 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1165: Depool db1165.eqiad.wmnet to then clone it to db1279.eqiad.wmnet - marostegui@cumin1003 * 09:42 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1165: Depool db1165.eqiad.wmnet to then clone it to db1279.eqiad.wmnet - marostegui@cumin1003 * 09:42 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1165.eqiad.wmnet onto db1279.eqiad.wmnet * 09:41 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=testwiki --start-timestamp="20260311000000" --end-timestamp="20260805000000" --sleep="5" --batch-size="2"` for [[phab:T434688|T434688]] * 09:40 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for backupmon1001.eqiad.wmnet * 09:40 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for backupmon1001.eqiad.wmnet * 09:40 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1151 * 09:40 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1151 * 09:37 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325408{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]], [[gerrit:1325407{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]] (duration: 06m 57s) * 09:36 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1151 * 09:36 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1151.eqiad.wmnet 13.36.64.10.in-addr.arpa 3.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:36 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1151.eqiad.wmnet 13.36.64.10.in-addr.arpa 3.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:36 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:36 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1151 - btullis@cumin1003" * 09:36 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on backupmon1001.eqiad.wmnet with reason: reboot * 09:36 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1151 - btullis@cumin1003" * 09:35 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host moss-be2003.codfw.wmnet * 09:33 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 7 hosts * 09:33 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 7 hosts * 09:33 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 09:32 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1325408{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]], [[gerrit:1325407{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1161.eqiad.wmnet onto db1275.eqiad.wmnet * 09:31 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 09:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1161: Pool db1161.eqiad.wmnet in after cloning * 09:30 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1325408{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]], [[gerrit:1325407{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]] * 09:29 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 09:27 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 09:27 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host moss-be2003.codfw.wmnet * 09:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be2006.codfw.wmnet * 09:25 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2004.codfw.wmnet * 09:22 hashar@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 09:21 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be2006.codfw.wmnet * 09:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be2005.codfw.wmnet * 09:19 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2004.codfw.wmnet * 09:19 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2003.codfw.wmnet * 09:18 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 7 hosts with reason: reboot * 09:18 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 7 hosts * 09:18 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 7 hosts * 09:16 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2214: Upgrade of db2214.codfw.wmnet completed * 09:14 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be2005.codfw.wmnet * 09:13 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be2004.codfw.wmnet * 09:12 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2003.codfw.wmnet * 09:12 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2002.codfw.wmnet * 09:09 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2214: Upgrading db2214.codfw.wmnet * 09:09 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2214: Upgrading db2214.codfw.wmnet * 09:09 cwilliams@cumin1003: START - Cookbook sre.mysql.upgrade for 1 hosts * 09:07 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be2004.codfw.wmnet * 09:06 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2002.codfw.wmnet * 09:05 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1004.eqiad.wmnet * 09:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-cluster (exit_code=0) * 09:03 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 7 hosts with reason: reboot * 09:01 btullis@cumin1003: START - Cookbook sre.dns.netbox * 09:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2214 [[phab:T434754|T434754]]', diff saved to https://phabricator.wikimedia.org/P96045 and previous config saved to /var/cache/conftool/dbconfig/20260813-090001-cwilliams.json * 08:59 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1004.eqiad.wmnet * 08:59 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1003.eqiad.wmnet * 08:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2229 to s6 primary [[phab:T434754|T434754]]', diff saved to https://phabricator.wikimedia.org/P96044 and previous config saved to /var/cache/conftool/dbconfig/20260813-085752-cwilliams.json * 08:57 cezmunsta: Starting s6 codfw failover from db2214 to db2229 - [[phab:T434754|T434754]] * 08:54 hashar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325416{{!}}Revert "REST: Enable `GET /lexemes/<nowiki>{</nowiki>lexeme_id<nowiki>}</nowiki>` by default" (T434712)]] (duration: 07m 22s) * 08:53 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1003.eqiad.wmnet * 08:53 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1002.eqiad.wmnet * 08:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2229 with weight 0 [[phab:T434754|T434754]]', diff saved to https://phabricator.wikimedia.org/P96043 and previous config saved to /var/cache/conftool/dbconfig/20260813-085151-cwilliams.json * 08:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 23 hosts with reason: Primary switchover s6 [[phab:T434754|T434754]] * 08:51 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts2002.codfw.wmnet * 08:50 hashar@deploy1003: hashar: Continuing with deployment * 08:49 hashar@deploy1003: hashar: Backport for [[gerrit:1325416{{!}}Revert "REST: Enable `GET /lexemes/<nowiki>{</nowiki>lexeme_id<nowiki>}</nowiki>` by default" (T434712)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:47 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet * 08:47 hashar@deploy1003: Started scap sync-world: Backport for [[gerrit:1325416{{!}}Revert "REST: Enable `GET /lexemes/<nowiki>{</nowiki>lexeme_id<nowiki>}</nowiki>` by default" (T434712)]] * 08:47 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1002.eqiad.wmnet * 08:46 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1161: Pool db1161.eqiad.wmnet in after cloning * 08:46 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host stewards2001.codfw.wmnet * 08:45 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit1003.wikimedia.org * 08:45 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host stewards1001.eqiad.wmnet * 08:45 Emperor: roll-restart apus frontends in codfw * 08:45 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-cluster * 08:44 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts2002.codfw.wmnet * 08:44 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host doc2003.codfw.wmnet * 08:43 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet * 08:43 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host phab2003.codfw.wmnet * 08:42 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host stewards2001.codfw.wmnet * 08:42 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host etherpad2002.codfw.wmnet * 08:41 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host stewards1001.eqiad.wmnet * 08:41 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1273: Pool in s7 * 08:41 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1003.wikimedia.org * 08:40 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host doc2003.codfw.wmnet * 08:40 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit1003.wikimedia.org * 08:39 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host etherpad1004.eqiad.wmnet * 08:39 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host doc1004.eqiad.wmnet * 08:39 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit2002.wikimedia.org * 08:38 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host etherpad2002.codfw.wmnet * 08:37 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host phab2003.codfw.wmnet * 08:36 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lists1004.wikimedia.org * 08:35 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host etherpad1004.eqiad.wmnet * 08:35 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host doc1004.eqiad.wmnet * 08:34 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-cluster (exit_code=0) * 08:34 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1003.wikimedia.org * 08:34 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host planet2003.codfw.wmnet * 08:34 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2003.wikimedia.org * 08:33 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host planet1003.eqiad.wmnet * 08:33 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit2002.wikimedia.org * 08:32 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 22 hosts with reason: Cloning * 08:32 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2002.wikimedia.org * 08:31 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aphlict1002.eqiad.wmnet * 08:30 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host planet2003.codfw.wmnet * 08:29 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host planet1003.eqiad.wmnet * 08:29 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host lists1004.wikimedia.org * 08:28 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lists2001.wikimedia.org * 08:28 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2003.wikimedia.org * 08:27 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host aphlict1002.eqiad.wmnet * 08:27 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aphlict2001.codfw.wmnet * 08:26 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2002.wikimedia.org * 08:23 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host aphlict2001.codfw.wmnet * 08:23 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1151 * 08:22 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1151.eqiad.wmnet with OS bookworm * 08:22 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host lists2001.wikimedia.org * 08:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1161: Depool db1161.eqiad.wmnet to then clone it to db1275.eqiad.wmnet - marostegui@cumin1003 * 08:20 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1161: Depool db1161.eqiad.wmnet to then clone it to db1275.eqiad.wmnet - marostegui@cumin1003 * 08:20 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1161.eqiad.wmnet onto db1275.eqiad.wmnet * 08:15 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-cluster * 08:15 Emperor: roll-restart apus frontends in eqiad * 07:56 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1273: Pool in s7 * 07:56 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1273 to dbctl [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P96035 and previous config saved to /var/cache/conftool/dbconfig/20260813-075611-marostegui.json * 07:38 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db1273.eqiad.wmnet with reason: Reboot * 07:31 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.sanitize-wiki (exit_code=97) Managing sanitization for wikis testwiki in section s3 * 07:24 marostegui@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis testwiki in section s3 * 07:19 jayme@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on kubestagemaster2005.codfw.wmnet with reason: downtime because of hardware failure and no DRBD * 05:42 arnaudb@dns1006: END - running authdns-update * 05:40 arnaudb@dns1006: START - running authdns-update * 05:27 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 04:06 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 02:29 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1151 * 02:29 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1151 * 02:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1219.eqiad.wmnet with OS bookworm * 02:18 ryankemper: [[phab:T434494|T434494]] `ryankemper@deploy1003:~$ echo 'https://stats.wikimedia.org/' {{!}} mwscript-k8s --attach -- purgeList.php` (default page got cached during yesterday's `an-web1001` reimage) * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 46s) * 02:03 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1219.eqiad.wmnet with reason: host reimage * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 02:00 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1219.eqiad.wmnet with reason: host reimage * 01:46 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1219.eqiad.wmnet with OS bookworm * 01:03 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1142.eqiad.wmnet with OS bookworm * 00:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1142.eqiad.wmnet with reason: host reimage * 00:34 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1142.eqiad.wmnet with reason: host reimage * 00:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1142.eqiad.wmnet with OS bookworm * 00:16 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1178 * 00:16 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1178 == 2026-08-12 == * 23:18 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324409{{!}}Remove $wmg = $wg hacks in Collection (T119117)]] (duration: 06m 43s) * 23:14 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 23:13 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1324409{{!}}Remove $wmg = $wg hacks in Collection (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:11 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1324409{{!}}Remove $wmg = $wg hacks in Collection (T119117)]] * 22:59 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 22:49 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260801000000" --end-timestamp="20260802000000" --sleep=2 --batch-size=10` * 22:45 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=testwiki --start-timestamp="20200801010101" --end-timestamp="20260816010101" --sleep=15 --batch-size=5` * 22:40 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=testwiki --start-timestamp="20200101010101" --end-timestamp="20260816010101" --sleep=60` * 22:35 Dreamy_Jazz: Running `mwscript WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260101000000" --end-timestamp="20260102000000" --sleep=10` * 22:21 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324817{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]], [[gerrit:1324818{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]] (duration: 45m 29s) * 22:17 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 21:59 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 21:40 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1324817{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]], [[gerrit:1324818{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:39 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 21:36 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1324817{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]], [[gerrit:1324818{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]] * 21:32 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:31 vriley@cumin1003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:30 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:30 vriley@cumin1003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:17 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:14 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:14 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:11 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:10 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-codfw: Set storage compatability to NONE — [[phab:T433028|T433028]] - eevans@cumin1003 * 21:10 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1006 * 21:09 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1006 * 21:05 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324277{{!}}Improve Math preference labels for SVG/MathJax/MathML (T433891)]] (duration: 31m 42s) * 20:58 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1178.eqiad.wmnet with OS bookworm * 20:54 krinkle@deploy1003: krinkle: Continuing with deployment * 20:51 krinkle@deploy1003: krinkle: Backport for [[gerrit:1324277{{!}}Improve Math preference labels for SVG/MathJax/MathML (T433891)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:39 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 20:38 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:38 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp5022.eqsin.wmnet with OS trixie * 20:38 cdobbins@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - cdobbins@cumin1003" * 20:37 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:36 cdobbins@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - cdobbins@cumin1003" * 20:34 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1178.eqiad.wmnet with reason: host reimage * 20:34 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1324277{{!}}Improve Math preference labels for SVG/MathJax/MathML (T433891)]] * 20:33 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 20:28 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1178.eqiad.wmnet with reason: host reimage * 20:13 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1178.eqiad.wmnet with OS bookworm * 20:11 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-worker1178.eqiad.wmnet with OS bookworm * 20:11 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1178.eqiad.wmnet with OS bookworm * 20:09 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-codfw: Set storage compatability to NONE — [[phab:T433028|T433028]] - eevans@cumin1003 * 20:09 Dreamy_Jazz: Evening UTC backport window done * 20:08 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324794{{!}}WikimediaAntiAbuse: Enable logging channel (T431292)]] (duration: 06m 48s) * 20:08 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage * 20:05 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage * 20:04 dreamyjazz@deploy1003: kharlan, dreamyjazz: Continuing with deployment * 20:04 dreamyjazz@deploy1003: kharlan, dreamyjazz: Backport for [[gerrit:1324794{{!}}WikimediaAntiAbuse: Enable logging channel (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:01 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1324794{{!}}WikimediaAntiAbuse: Enable logging channel (T431292)]] * 19:55 brennen@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 19:47 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 19:35 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 19:35 cdobbins@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cp5022.eqsin.wmnet with OS trixie * 19:32 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:30 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:29 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:26 vriley@cumin1003: START - Cookbook sre.dns.netbox * 19:23 brennen@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 19:19 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-eqiad: Set storage compatability to NONE — [[phab:T433028|T433028]] - eevans@cumin1003 * 19:10 brennen@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 19:09 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 19:09 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 19:00 Amir1: data migrated on wikishared ([[phab:T426102|T426102]]) * 18:57 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321224{{!}}Rename ce_worklist_articles table to ce_invitation_list_articles (T426102)]] (duration: 06m 50s) * 18:53 ladsgroup@deploy1003: ladsgroup, daimona: Continuing with deployment * 18:53 ladsgroup@deploy1003: ladsgroup, daimona: Backport for [[gerrit:1321224{{!}}Rename ce_worklist_articles table to ce_invitation_list_articles (T426102)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:51 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1321224{{!}}Rename ce_worklist_articles table to ce_invitation_list_articles (T426102)]] * 18:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2192: Security update * 18:31 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 18:25 Amir1: ce_invitation_list_articles created as empty on wikishared ([[phab:T426102|T426102]]) * 18:21 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 18:21 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 18:18 Amir1: migrated testwiki entries from ce_worklist_articles to ce_invitation_list_articles ([[phab:T426102|T426102]]) * 18:18 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-eqiad: Set storage compatability to NONE — [[phab:T433028|T433028]] - eevans@cumin1003 * 18:11 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 18:08 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling reboot on A:durum-eqsin and A:durum * 18:07 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Set storage compatability to UPGRADING — [[phab:T433028|T433028]] - eevans@cumin1003 * 18:05 jhancock@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022'] * 17:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2192: Security update * 17:55 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:55 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum-eqsin and A:durum * 17:53 jhancock@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['cp5022'] * 17:47 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:47 jhancock@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['cp5022'] * 17:42 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:41 jhancock@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['cp5022'] * 17:36 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:35 jhancock@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022'] * 17:31 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-magru and not (P<nowiki>{</nowiki>cp7001*<nowiki>}</nowiki> or P<nowiki>{</nowiki>cp7009*<nowiki>}</nowiki>) and A:cp - 9.2.15 upgrade ([[phab:T434620|T434620]]) * 17:28 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:28 jhancock@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['cp5022'] * 17:22 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2192.codfw.wmnet with reason: Maintenance * 17:11 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host apifeatureusage2001.codfw.wmnet with OS bookworm * 17:04 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: apply * 17:03 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-main: apply * 17:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2192 [[phab:T434635|T434635]]', diff saved to https://phabricator.wikimedia.org/P96030 and previous config saved to /var/cache/conftool/dbconfig/20260812-170338-cwilliams.json * 17:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2213 to s5 primary [[phab:T434635|T434635]]', diff saved to https://phabricator.wikimedia.org/P96029 and previous config saved to /var/cache/conftool/dbconfig/20260812-170152-cwilliams.json * 17:01 cezmunsta: Starting s5 codfw failover from db2192 to db2213 - [[phab:T434635|T434635]] * 16:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2213 with weight 0 [[phab:T434635|T434635]]', diff saved to https://phabricator.wikimedia.org/P96028 and previous config saved to /var/cache/conftool/dbconfig/20260812-165544-cwilliams.json * 16:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 27 hosts with reason: Primary switchover s5 [[phab:T434635|T434635]] * 16:53 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: apply * 16:52 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-main: apply * 16:44 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-main: apply * 16:44 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-main: apply * 16:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1272: New host * 16:40 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324765{{!}}WikimediaAntiAbuse: Enable PersonalInfoFlagNotifications (T431292)]] (duration: 07m 02s) * 16:40 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-logging-external: apply * 16:39 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-logging-external: apply * 16:38 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-logging-external: apply * 16:37 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-logging-external: apply * 16:36 kharlan@deploy1003: kharlan: Continuing with deployment * 16:35 kharlan@deploy1003: kharlan: Backport for [[gerrit:1324765{{!}}WikimediaAntiAbuse: Enable PersonalInfoFlagNotifications (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:33 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1324765{{!}}WikimediaAntiAbuse: Enable PersonalInfoFlagNotifications (T431292)]] * 16:27 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-logging-external: apply * 16:27 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-logging-external: apply * 16:18 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 16:18 jhancock@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022'] * 16:17 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 16:16 jhancock@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['cp5022'] * 16:15 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 16:12 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Set storage compatability to UPGRADING — [[phab:T433028|T433028]] - eevans@cumin1003 * 16:10 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: apply * 16:10 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: apply * 16:08 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: apply * 16:08 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: apply * 16:08 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics: apply * 16:07 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics: apply * 16:02 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-magru and not (P<nowiki>{</nowiki>cp7001*<nowiki>}</nowiki> or P<nowiki>{</nowiki>cp7009*<nowiki>}</nowiki>) and A:cp - 9.2.15 upgrade ([[phab:T434620|T434620]]) * 15:57 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1272: New host * 15:52 jmm@cumin2003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti2046.codfw.wmnet * 15:52 jmm@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host ganeti2046.codfw.wmnet * 15:42 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:37 mforns@deploy1003: Finished deploy [analytics/refinery@49c336c] (thin): Regular analytics weekly train THIN [analytics/refinery@49c336cd] (duration: 01m 59s) * 15:37 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324733{{!}}WikimediaAntiAbuse: Enable personal info tag display on enwiki (T431292)]] (duration: 08m 12s) * 15:35 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:35 mforns@deploy1003: Started deploy [analytics/refinery@49c336c] (thin): Regular analytics weekly train THIN [analytics/refinery@49c336cd] * 15:34 mforns@deploy1003: Finished deploy [analytics/refinery@49c336c]: Regular analytics weekly train [analytics/refinery@49c336cd] (duration: 04m 20s) * 15:33 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[2024,1031]*.wmnet: Set storage compatability to UPGRADING — [[phab:T433028|T433028]] - eevans@cumin1003 * 15:33 dreamyjazz@deploy1003: kharlan, dreamyjazz: Continuing with deployment * 15:31 dreamyjazz@deploy1003: kharlan, dreamyjazz: Backport for [[gerrit:1324733{{!}}WikimediaAntiAbuse: Enable personal info tag display on enwiki (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:30 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitize-wiki (exit_code=99) Checking sanitization for wikis testwiki in section s3 * 15:30 mforns@deploy1003: Started deploy [analytics/refinery@49c336c]: Regular analytics weekly train [analytics/refinery@49c336cd] * 15:30 mforns@deploy1003: Finished deploy [analytics/refinery@49c336c] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@49c336cd] (duration: 00m 32s) * 15:29 mforns@deploy1003: Started deploy [analytics/refinery@49c336c] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@49c336cd] * 15:29 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1324733{{!}}WikimediaAntiAbuse: Enable personal info tag display on enwiki (T431292)]] * 15:27 brennen@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324700{{!}}EventDetailsParticipantsModule: populate cache with non-local users (T434597)]] (duration: 06m 38s) * 15:23 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[2024,1031]*.wmnet: Set storage compatability to UPGRADING — [[phab:T433028|T433028]] - eevans@cumin1003 * 15:23 brennen@deploy1003: brennen, daimona: Continuing with deployment * 15:23 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:22 brennen@deploy1003: brennen, daimona: Backport for [[gerrit:1324700{{!}}EventDetailsParticipantsModule: populate cache with non-local users (T434597)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:22 cgoubert@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host rdb-lock2003.codfw.wmnet * 15:21 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 15:21 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 15:21 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:21 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:21 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:20 brennen@deploy1003: Started scap sync-world: Backport for [[gerrit:1324700{{!}}EventDetailsParticipantsModule: populate cache with non-local users (T434597)]] * 15:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1178.eqiad.wmnet with OS bookworm * 15:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock1003.eqiad.wmnet * 15:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock1003.eqiad.wmnet with OS trixie * 15:16 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 15:16 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 15:16 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 15:16 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:16 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:16 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:12 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 15:12 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2003.codfw.wmnet * 15:11 cgoubert@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host rdb-lock2003.codfw.wmnet * 15:11 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 15:11 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 15:11 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:11 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:11 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:07 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324338{{!}}InitialiseSettings: Enable 2FA warnings on more private wikis (T428103)]] (duration: 07m 02s) * 15:04 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 15:04 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock1003.eqiad.wmnet with reason: host reimage * 15:03 reedy@deploy1003: reedy: Continuing with deployment * 15:02 reedy@deploy1003: reedy: Backport for [[gerrit:1324338{{!}}InitialiseSettings: Enable 2FA warnings on more private wikis (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:02 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:00 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324338{{!}}InitialiseSettings: Enable 2FA warnings on more private wikis (T428103)]] * 14:57 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 14:57 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2003.codfw.wmnet * 14:57 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock1003.eqiad.wmnet with reason: host reimage * 14:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock2002.codfw.wmnet * 14:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock2002.codfw.wmnet with OS trixie * 14:56 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 14:56 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 14:56 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 14:55 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 14:55 moritzm: powercycle ganeti2046 * 14:47 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock1003.eqiad.wmnet with OS trixie * 14:46 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1003.eqiad.wmnet - cgoubert@cumin2003" * 14:46 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1003.eqiad.wmnet - cgoubert@cumin2003" * 14:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock1003.eqiad.wmnet on all recursors * 14:45 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock1003.eqiad.wmnet on all recursors * 14:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1003.eqiad.wmnet - cgoubert@cumin2003" * 14:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1272.eqiad.wmnet with reason: Enabling notifications * 14:44 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1003.eqiad.wmnet - cgoubert@cumin2003" * 14:44 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324719{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324720{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324722{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0 (T434187)]] (duration: 11m 02s) * 14:39 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 14:39 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock1003.eqiad.wmnet * 14:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock2002.codfw.wmnet with reason: host reimage * 14:37 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 14:37 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock1002.eqiad.wmnet * 14:37 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock1002.eqiad.wmnet with OS trixie * 14:37 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1324719{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324720{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324722{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0 (T434187)]] synced to the testservers (see https://wikitech. * 14:33 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock2002.codfw.wmnet with reason: host reimage * 14:33 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1324719{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324720{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324722{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0 (T434187)]] * 14:32 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1018.eqiad.wmnet with OS bookworm * 14:32 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2046.codfw.wmnet * 14:31 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1020.eqiad.wmnet with OS bookworm * 14:27 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2046.codfw.wmnet * 14:25 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2045.codfw.wmnet * 14:25 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2045.codfw.wmnet * 14:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock1002.eqiad.wmnet with reason: host reimage * 14:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1019.eqiad.wmnet with OS bookworm * 14:22 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Checking sanitization for wikis testwiki in section s3 * 14:20 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2045.codfw.wmnet * 14:18 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock1002.eqiad.wmnet with reason: host reimage * 14:17 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:16 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324707{{!}}Backport all changes from wmf/1.47.0-wmf.15]] (duration: 40m 51s) * 14:16 moritzm: installing Linux 6.1.180 on Bookworm hosts * 14:15 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock2002.codfw.wmnet with OS trixie * 14:14 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2002.codfw.wmnet - cgoubert@cumin2003" * 14:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2002.codfw.wmnet - cgoubert@cumin2003" * 14:14 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:14 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2045.codfw.wmnet * 14:14 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2002.codfw.wmnet on all recursors * 14:14 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2002.codfw.wmnet on all recursors * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2002.codfw.wmnet - cgoubert@cumin2003" * 14:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2002.codfw.wmnet - cgoubert@cumin2003" * 14:12 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2030.codfw.wmnet * 14:12 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2030.codfw.wmnet * 14:11 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on apifeatureusage2001.codfw.wmnet with reason: host reimage * 14:09 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:08 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:07 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 14:06 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2030.codfw.wmnet * 14:06 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:06 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock1002.eqiad.wmnet with OS trixie * 14:05 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 14:05 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2002.codfw.wmnet * 14:05 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1002.eqiad.wmnet - cgoubert@cumin2003" * 14:05 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1002.eqiad.wmnet - cgoubert@cumin2003" * 14:05 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock1002.eqiad.wmnet on all recursors * 14:05 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock1002.eqiad.wmnet on all recursors * 14:05 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:05 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1002.eqiad.wmnet - cgoubert@cumin2003" * 14:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:04 kharlan@deploy1003: kharlan: Continuing with deployment * 14:04 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock2001.codfw.wmnet * 14:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock2001.codfw.wmnet with OS trixie * 14:02 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1002.eqiad.wmnet - cgoubert@cumin2003" * 14:02 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on apifeatureusage2001.codfw.wmnet with reason: host reimage * 14:01 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2030.codfw.wmnet * 13:59 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2029.codfw.wmnet * 13:58 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2029.codfw.wmnet * 13:58 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 13:58 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock1002.eqiad.wmnet * 13:56 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1019.eqiad.wmnet with reason: host reimage * 13:54 btullis@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on archiva1002.wikimedia.org with reason: Upgrading in-place * 13:53 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock1001.eqiad.wmnet * 13:53 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock1001.eqiad.wmnet with OS trixie * 13:53 kharlan@deploy1003: kharlan: Backport for [[gerrit:1324707{{!}}Backport all changes from wmf/1.47.0-wmf.15]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:52 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2029.codfw.wmnet * 13:52 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1020.eqiad.wmnet with reason: host reimage * 13:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1019.eqiad.wmnet with reason: host reimage * 13:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1020.eqiad.wmnet with reason: host reimage * 13:49 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock2001.codfw.wmnet with reason: host reimage * 13:48 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2029.codfw.wmnet * 13:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Configuring db1272 for s3 pooling', diff saved to https://phabricator.wikimedia.org/P96021 and previous config saved to /var/cache/conftool/dbconfig/20260812-134732-cwilliams.json * 13:44 bking@cumin2003: START - Cookbook sre.hosts.reimage for host apifeatureusage2001.codfw.wmnet with OS bookworm * 13:43 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock2001.codfw.wmnet with reason: host reimage * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2028.codfw.wmnet * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2028.codfw.wmnet * 13:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1018.eqiad.wmnet with reason: host reimage * 13:38 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock1001.eqiad.wmnet with reason: host reimage * 13:36 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1018.eqiad.wmnet with reason: host reimage * 13:35 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1324707{{!}}Backport all changes from wmf/1.47.0-wmf.15]] * 13:35 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2028.codfw.wmnet * 13:32 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1275938{{!}}Enable campaignEvents on bdwikimedia (T424016)]] (duration: 07m 35s) * 13:32 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1020.eqiad.wmnet with OS bookworm * 13:32 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1019.eqiad.wmnet with OS bookworm * 13:32 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock1001.eqiad.wmnet with reason: host reimage * 13:28 kharlan@deploy1003: kharlan, yahya: Continuing with deployment * 13:27 kharlan@deploy1003: kharlan, yahya: Backport for [[gerrit:1275938{{!}}Enable campaignEvents on bdwikimedia (T424016)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:27 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2028.codfw.wmnet * 13:26 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 13:26 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 13:25 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1275938{{!}}Enable campaignEvents on bdwikimedia (T424016)]] * 13:25 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitize-wiki (exit_code=99) Managing sanitization for wikis testwiki in section s3 * 13:24 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock2001.codfw.wmnet with OS trixie * 13:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2001.codfw.wmnet - cgoubert@cumin2003" * 13:24 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2001.codfw.wmnet - cgoubert@cumin2003" * 13:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2001.codfw.wmnet on all recursors * 13:23 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2001.codfw.wmnet on all recursors * 13:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2001.codfw.wmnet - cgoubert@cumin2003" * 13:23 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2001.codfw.wmnet - cgoubert@cumin2003" * 13:23 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324713{{!}}thwiki: reinstate temporary wiki25 logos (T431094)]] (duration: 07m 13s) * 13:20 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:20 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1018.eqiad.wmnet with OS bookworm * 13:20 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2027.codfw.wmnet * 13:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2027.codfw.wmnet * 13:19 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics-external: apply * 13:18 kharlan@deploy1003: anzx, kharlan: Continuing with deployment * 13:18 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:18 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock1001.eqiad.wmnet with OS trixie * 13:17 kharlan@deploy1003: anzx, kharlan: Backport for [[gerrit:1324713{{!}}thwiki: reinstate temporary wiki25 logos (T431094)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1001.eqiad.wmnet - cgoubert@cumin2003" * 13:17 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 13:17 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1001.eqiad.wmnet - cgoubert@cumin2003" * 13:17 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2001.codfw.wmnet * 13:17 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics-external: apply * 13:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock1001.eqiad.wmnet on all recursors * 13:17 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock1001.eqiad.wmnet on all recursors * 13:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1001.eqiad.wmnet - cgoubert@cumin2003" * 13:17 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1001.eqiad.wmnet - cgoubert@cumin2003" * 13:15 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1324713{{!}}thwiki: reinstate temporary wiki25 logos (T431094)]] * 13:15 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:15 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics-external: apply * 13:15 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:14 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics-external: apply * 13:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2027.codfw.wmnet * 13:12 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 13:12 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock1001.eqiad.wmnet * 13:12 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2027.codfw.wmnet * 13:08 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1158.eqiad.wmnet onto db1273.eqiad.wmnet * 13:07 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1158: Pool db1158.eqiad.wmnet in after cloning * 13:02 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti6002.drmrs.wmnet * 13:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti6002.drmrs.wmnet * 12:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1015.eqiad.wmnet with OS bookworm * 12:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti6002.drmrs.wmnet * 12:44 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1017.eqiad.wmnet with OS bookworm * 12:36 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 12:35 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 12:34 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 12:33 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 12:31 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti6002.drmrs.wmnet * 12:24 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1017.eqiad.wmnet with reason: host reimage * 12:22 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1158: Pool db1158.eqiad.wmnet in after cloning * 12:18 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1017.eqiad.wmnet with reason: host reimage * 12:11 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1015.eqiad.wmnet with reason: host reimage * 12:07 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1015.eqiad.wmnet with reason: host reimage * 12:04 moritzm: failover ganeti master in drmrs02 to ganeti6004 * 12:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1009.eqiad.wmnet with OS bookworm * 12:01 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1017.eqiad.wmnet with OS bookworm * 12:00 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti6004.drmrs.wmnet * 12:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti6004.drmrs.wmnet * 11:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti6004.drmrs.wmnet * 11:53 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1015.eqiad.wmnet with OS bookworm * 11:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1159.eqiad.wmnet onto db1274.eqiad.wmnet * 11:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Pool db1159.eqiad.wmnet in after cloning * 11:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1016.eqiad.wmnet with OS bookworm * 11:46 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti6004.drmrs.wmnet * 11:45 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti6001.drmrs.wmnet * 11:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti6001.drmrs.wmnet * 11:43 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324306{{!}}WikimediaAntiAbuse: Enable personal info for enwiki with no display (T431292)]] (duration: 10m 26s) * 11:42 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1015.eqiad.wmnet with OS bookworm * 11:39 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 11:38 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti6001.drmrs.wmnet * 11:34 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1324306{{!}}WikimediaAntiAbuse: Enable personal info for enwiki with no display (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:33 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2212: Security update * 11:33 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti6001.drmrs.wmnet * 11:32 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1324306{{!}}WikimediaAntiAbuse: Enable personal info for enwiki with no display (T431292)]] * 11:22 moritzm: failover ganeti master in drmrs01 to ganeti6003 * 11:20 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:20 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:18 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 22 hosts with reason: Cloning * 11:17 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti6003.drmrs.wmnet * 11:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti6003.drmrs.wmnet * 11:17 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1009.eqiad.wmnet with reason: host reimage * 11:17 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:16 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:14 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1016.eqiad.wmnet with reason: host reimage * 11:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti6003.drmrs.wmnet * 11:11 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1009.eqiad.wmnet with reason: host reimage * 11:10 jmm@cumin2003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti3005.esams.wmnet * 11:10 jmm@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host ganeti3005.esams.wmnet * 11:09 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 11:08 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1158: Depool db1158.eqiad.wmnet to then clone it to db1273.eqiad.wmnet - marostegui@cumin1003 * 11:07 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1016.eqiad.wmnet with reason: host reimage * 11:07 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1158: Depool db1158.eqiad.wmnet to then clone it to db1273.eqiad.wmnet - marostegui@cumin1003 * 11:07 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1158.eqiad.wmnet onto db1273.eqiad.wmnet * 11:06 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:05 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:05 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1159: Pool db1159.eqiad.wmnet in after cloning * 11:04 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 20 hosts with reason: Cloning * 11:02 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti6003.drmrs.wmnet * 11:00 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis testwiki in section s3 * 10:54 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1009.eqiad.wmnet with OS bookworm * 10:51 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1015.eqiad.wmnet with OS bookworm * 10:50 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1016.eqiad.wmnet with OS bookworm * 10:48 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2212: Security update * 10:45 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitize-wiki (exit_code=99) Managing sanitization for wikis testwiki in section s3 * 10:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1013.eqiad.wmnet with OS bookworm * 10:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1014.eqiad.wmnet with OS bookworm * 10:38 cwilliams@cumin1003: START - Cookbook sre.mysql.clone of db1159.eqiad.wmnet onto db1274.eqiad.wmnet * 10:33 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db1274.eqiad.wmnet * 10:33 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db1274.eqiad.wmnet * 10:31 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 396993 * 10:29 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 396993 * 10:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1159: Clone source for db1274 * 10:24 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1159: Clone source for db1274 * 10:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1013.eqiad.wmnet with reason: host reimage * 10:18 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1013.eqiad.wmnet with reason: host reimage * 10:13 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2212.codfw.wmnet with reason: Maintenance * 10:13 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 10:12 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 10:12 blake@deploy1003: Stopping before sync operations * 10:11 blake@deploy1003: Started scap sync-world: Non-deployment scap run to populate new release values for [[phab:T427668|T427668]] * 10:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2212 [[phab:T434644|T434644]]', diff saved to https://phabricator.wikimedia.org/P96003 and previous config saved to /var/cache/conftool/dbconfig/20260812-101053-cwilliams.json * 10:09 moritzm: powercycle ganeti3005 * 10:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2203 to s1 primary [[phab:T434644|T434644]]', diff saved to https://phabricator.wikimedia.org/P96002 and previous config saved to /var/cache/conftool/dbconfig/20260812-100849-cwilliams.json * 10:08 cezmunsta: Starting s1 codfw failover from db2212 to db2203 - [[phab:T434644|T434644]] * 10:03 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1013.eqiad.wmnet with OS bookworm * 10:02 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1013.eqiad.wmnet with OS bookworm * 10:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2203 with weight 0 [[phab:T434644|T434644]]', diff saved to https://phabricator.wikimedia.org/P96001 and previous config saved to /var/cache/conftool/dbconfig/20260812-100134-cwilliams.json * 10:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 32 hosts with reason: Primary switchover s1 [[phab:T434644|T434644]] * 09:53 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1014.eqiad.wmnet with reason: host reimage * 09:50 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 09:50 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti3005.esams.wmnet * 09:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1014.eqiad.wmnet with reason: host reimage * 09:41 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1278: Pool in x1 * 09:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-launcher1003.eqiad.wmnet with OS bookworm * 09:37 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti3005.esams.wmnet * 09:37 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1013.eqiad.wmnet with OS bookworm * 09:34 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-presto1013.eqiad.wmnet with OS bookworm * 09:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1014.eqiad.wmnet with OS bookworm * 09:29 moritzm: failover ganeti master in esams to ganeti3008 * 09:26 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti3006.esams.wmnet * 09:26 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti3006.esams.wmnet * 09:24 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1012.eqiad.wmnet with OS bookworm * 09:23 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1009.eqiad.wmnet with OS bookworm * 09:23 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 09:18 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti3006.esams.wmnet * 09:16 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti3006.esams.wmnet * 09:14 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: fix regexp escaping bug - oblivian@cumin1003" * 09:14 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: fix regexp escaping bug - oblivian@cumin1003 * 09:13 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: fix regexp escaping bug - oblivian@cumin1003 * 09:13 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: fix regexp escaping bug - oblivian@cumin1003" * 09:03 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-launcher1003.eqiad.wmnet with reason: host reimage * 08:58 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-launcher1003.eqiad.wmnet with reason: host reimage * 08:55 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1278: Pool in x1 * 08:55 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1278 to dbctl [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95996 and previous config saved to /var/cache/conftool/dbconfig/20260812-085521-marostegui.json * 08:51 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1012.eqiad.wmnet with reason: host reimage * 08:45 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis testwiki in section s3 * 08:43 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1009.eqiad.wmnet with OS bookworm * 08:42 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1012.eqiad.wmnet with reason: host reimage * 08:41 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-launcher1003.eqiad.wmnet with OS bookworm * 08:40 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1013.eqiad.wmnet with OS bookworm * 08:38 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms3', diff saved to https://phabricator.wikimedia.org/P95995 and previous config saved to /var/cache/conftool/dbconfig/20260812-083816-marostegui.json * 08:38 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-master1003.eqiad.wmnet with OS bookworm * 08:37 marostegui: Failover ms3 [[phab:T434288|T434288]] * 08:37 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1268 to dbctl [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95994 and previous config saved to /var/cache/conftool/dbconfig/20260812-083722-marostegui.json * 08:35 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-web1001.eqiad.wmnet with OS bookworm * 08:32 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db2252.codfw.wmnet,db[1153,1268].eqiad.wmnet with reason: Switching over ms3 * 08:28 marostegui@cumin1003: dbctl commit (dc=all): 'Depool ms3 [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95993 and previous config saved to /var/cache/conftool/dbconfig/20260812-082852-marostegui.json * 08:25 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1012.eqiad.wmnet with OS bookworm * 08:25 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1011.eqiad.wmnet with OS bookworm * 08:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-master1003.eqiad.wmnet with reason: host reimage * 08:07 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-master1003.eqiad.wmnet with reason: host reimage * 08:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-web1001.eqiad.wmnet with reason: host reimage * 07:58 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-web1001.eqiad.wmnet with reason: host reimage * 07:50 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1003.eqiad.wmnet with OS bookworm * 07:38 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1011.eqiad.wmnet with reason: host reimage * 07:38 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-master1003.eqiad.wmnet with OS bookworm * 07:35 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti3007.esams.wmnet * 07:35 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti3007.esams.wmnet * 07:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1011.eqiad.wmnet with reason: host reimage * 07:27 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti3007.esams.wmnet * 07:25 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti3007.esams.wmnet * 07:25 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti3008.esams.wmnet * 07:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti3008.esams.wmnet * 07:22 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-web1001.eqiad.wmnet with OS bookworm * 07:19 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1003.eqiad.wmnet with OS bookworm * 07:18 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-master1003.eqiad.wmnet * 07:18 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host an-master1003.eqiad.wmnet * 07:17 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1011.eqiad.wmnet with OS bookworm * 07:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 07:16 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti3008.esams.wmnet * 07:15 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1009.eqiad.wmnet with OS bookworm * 07:14 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host an-master1003.eqiad.wmnet * 07:13 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-master1003.eqiad.wmnet * 07:13 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-master1003.eqiad.wmnet * 07:12 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-master1003.eqiad.wmnet * 07:11 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti3008.esams.wmnet * 07:07 arnaudb@dns1006: END - running authdns-update * 07:07 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti5007.eqsin.wmnet * 07:07 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti5007.eqsin.wmnet * 07:05 arnaudb@dns1006: START - running authdns-update * 06:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti5007.eqsin.wmnet * 06:54 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti5007.eqsin.wmnet * 06:38 moritzm: failover ganeti master in eqsin to ganeti5004 * 06:36 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti5006.eqsin.wmnet * 06:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti5006.eqsin.wmnet * 06:28 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti5006.eqsin.wmnet * 06:23 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti5006.eqsin.wmnet * 06:20 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti5005.eqsin.wmnet * 06:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti5005.eqsin.wmnet * 06:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti5005.eqsin.wmnet * 06:06 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti5005.eqsin.wmnet * 06:03 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti5004.eqsin.wmnet * 06:03 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti5004.eqsin.wmnet * 05:55 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti5004.eqsin.wmnet * 05:53 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti5004.eqsin.wmnet * 04:40 ryankemper: [[phab:T434494|T434494]] reimaged `an-tool1008.eqiad.wmnet` to bookworm; yarn.wikimedia.org is back up * 04:16 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-tool1008.eqiad.wmnet with OS bookworm * 03:58 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-tool1008.eqiad.wmnet with reason: host reimage * 03:53 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-tool1008.eqiad.wmnet with reason: host reimage * 03:41 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-tool1008.eqiad.wmnet with OS bookworm * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 45s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 00:25 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324427{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]], [[gerrit:1324429{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0]], [[gerrit:1324428{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]] (duration: 07m 55s) * 00:21 kemayo@deploy1003: kemayo: Continuing with deployment * 00:19 kemayo@deploy1003: kemayo: Backport for [[gerrit:1324427{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]], [[gerrit:1324429{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0]], [[gerrit:1324428{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:17 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1324427{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]], [[gerrit:1324429{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0]], [[gerrit:1324428{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]] == 2026-08-11 == * 21:37 sbassett: Deployed security fix for [[phab:T434521|T434521]] (wmf.15) * 21:29 sbassett: Deployed security fix for [[phab:T434521|T434521]] (wmf.14) * 21:19 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324370{{!}}Phase 4 of legal footer deployment (T432796)]], [[gerrit:1319804{{!}}Disable wgMFCustomSiteModules on English Wikipedia (T375538)]] (duration: 15m 26s) * 21:15 jdlrobson@deploy1003: jdlrobson: Continuing with deployment * 21:06 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1324370{{!}}Phase 4 of legal footer deployment (T432796)]], [[gerrit:1319804{{!}}Disable wgMFCustomSiteModules on English Wikipedia (T375538)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:03 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1324370{{!}}Phase 4 of legal footer deployment (T432796)]], [[gerrit:1319804{{!}}Disable wgMFCustomSiteModules on English Wikipedia (T375538)]] * 20:59 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 20:50 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324384{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]], [[gerrit:1324385{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]] (duration: 06m 58s) * 20:46 kemayo@deploy1003: kemayo: Continuing with deployment * 20:45 kemayo@deploy1003: kemayo: Backport for [[gerrit:1324384{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]], [[gerrit:1324385{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:43 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1324384{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]], [[gerrit:1324385{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]] * 20:42 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 20:42 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324386{{!}}build: Updating js-yaml to 3.15.1, 4.3.1]] (duration: 07m 36s) * 20:38 kemayo@deploy1003: kemayo: Continuing with deployment * 20:37 jhancock@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 20:37 kemayo@deploy1003: kemayo: Backport for [[gerrit:1324386{{!}}build: Updating js-yaml to 3.15.1, 4.3.1]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:35 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1324386{{!}}build: Updating js-yaml to 3.15.1, 4.3.1]] * 20:18 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 20:15 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 20:15 jhancock@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin1003" * 20:14 jhancock@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin1003" * 19:59 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 19:54 jhancock@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 19:10 brennen@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] (duration: 06m 41s) * 19:04 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns6001.wikimedia.org * 19:04 sukhe@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns6001.wikimedia.org * 19:04 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns5003.wikimedia.org * 19:04 sukhe@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns5003.wikimedia.org * 19:03 brennen@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 18:59 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns5003.wikimedia.org with OS trixie * 18:55 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns6001.wikimedia.org with OS trixie * 18:19 brett@cumin2002: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on P<nowiki>{</nowiki>cp7009.magru.wmnet<nowiki>}</nowiki> and A:cp - 9.2.15 Upgrade () * 18:14 brett@cumin2002: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on P<nowiki>{</nowiki>cp7009.magru.wmnet<nowiki>}</nowiki> and A:cp - 9.2.15 Upgrade () * 18:13 brennen@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 18:12 brett@cumin2002: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 9.2.15 Upgrade () * 18:09 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns5003.wikimedia.org with reason: host reimage * 18:06 brett@cumin2002: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 9.2.15 Upgrade () * 18:06 brennen: 1.47.0-wmf.15 train status ([[phab:T430834|T430834]]) - no current blockers, rolling to group0 * 18:05 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns5003.wikimedia.org with reason: host reimage * 18:05 brett: import trafficserver-9.2.15~deb13+wmf1 into trixie-wikimedia ([[phab:T434478|T434478]]) * 18:01 ladsgroup@cumin1003: END (PASS) - Cookbook sre.mysql.sanitarium_restart (exit_code=0) * 17:58 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns6001.wikimedia.org with reason: host reimage * 17:53 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324369{{!}}Enable desktop lazy loading on group0 (T148047)]] (duration: 07m 31s) * 17:52 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns6001.wikimedia.org with reason: host reimage * 17:49 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 17:49 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitarium_restart (exit_code=99) * 17:49 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 17:49 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7001.magru.wmnet * 17:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti7001.magru.wmnet * 17:48 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 17:47 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1324369{{!}}Enable desktop lazy loading on group0 (T148047)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:45 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1324369{{!}}Enable desktop lazy loading on group0 (T148047)]] * 17:39 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti7001.magru.wmnet * 17:36 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns5003.wikimedia.org with OS trixie * 17:34 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns6001.wikimedia.org with OS trixie * 17:31 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324361{{!}}Move FR config from IS.php to a dedicated file]], [[gerrit:1324363{{!}}Remove $wmg = $wg hacks in CentralAuth (T119117)]] (duration: 12m 23s) * 17:26 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 17:23 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1324361{{!}}Move FR config from IS.php to a dedicated file]], [[gerrit:1324363{{!}}Remove $wmg = $wg hacks in CentralAuth (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:19 sukhe: sudo cumin "A:cp-magru" "run-puppet-agent --enable 'merging CR 1324355'" * 17:18 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1324361{{!}}Move FR config from IS.php to a dedicated file]], [[gerrit:1324363{{!}}Remove $wmg = $wg hacks in CentralAuth (T119117)]] * 17:11 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-master1004.eqiad.wmnet with OS bookworm * 17:11 sukhe: sukhe@cp7005:~$ sudo puppet agent -tv * 17:02 sukhe: sudo cumin "A:cp-magru" "disable-puppet 'merging CR 1324355'" * 16:54 sukhe@dns1004: END - running authdns-update * 16:53 sukhe@dns1004: START - running authdns-update * 16:53 sukhe@dns1004: FAIL - running authdns-update * 16:51 sukhe@dns1004: START - running authdns-update * 16:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-master1004.eqiad.wmnet with reason: host reimage * 16:44 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-master1004.eqiad.wmnet with reason: host reimage * 16:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1157.eqiad.wmnet onto db1272.eqiad.wmnet * 16:40 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1157: Pool db1157.eqiad.wmnet in after cloning * 16:38 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 16:31 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324356{{!}}InitialiseSettings: Fix wgOATHAuthEnforce2FAForAll]] (duration: 06m 52s) * 16:30 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 16:28 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 16:27 reedy@deploy1003: reedy: Continuing with deployment * 16:26 reedy@deploy1003: reedy: Backport for [[gerrit:1324356{{!}}InitialiseSettings: Fix wgOATHAuthEnforce2FAForAll]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:24 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324356{{!}}InitialiseSettings: Fix wgOATHAuthEnforce2FAForAll]] * 16:13 sukhe: restart ntpsec.serviceon dns7001 * 16:09 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324335{{!}}InitialiseSettings: Enable 2FA enforcement on various private wikis (T428103)]] (duration: 06m 40s) * 16:08 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 16:06 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2204: Security update * 16:04 reedy@deploy1003: reedy: Continuing with deployment * 16:04 reedy@deploy1003: reedy: Backport for [[gerrit:1324335{{!}}InitialiseSettings: Enable 2FA enforcement on various private wikis (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:02 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7001.magru.wmnet * 16:02 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324335{{!}}InitialiseSettings: Enable 2FA enforcement on various private wikis (T428103)]] * 16:01 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 15:55 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1157: Pool db1157.eqiad.wmnet in after cloning * 15:54 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 15:54 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 15:41 dancy@deploy1003: Finished scap sync-world: Testing (duration: 06m 28s) * 15:40 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-master1004.eqiad.wmnet with OS bookworm * 15:40 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 15:35 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti4008.ulsfo.wmnet * 15:35 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti4008.ulsfo.wmnet * 15:34 dancy@deploy1003: Started scap sync-world: Testing * 15:34 dancy@deploy1003: Installation of scap version "4.279.0" completed for 3 hosts * 15:34 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-master1004.eqiad.wmnet with OS bookworm * 15:34 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 15:33 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-master1004.eqiad.wmnet with OS bookworm * 15:32 dancy@deploy1003: Installing scap version "4.279.0" for 3 host(s) * 15:32 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324339{{!}}Add /w/deployment-info.php entrypoint]] (duration: 07m 25s) * 15:30 moritzm: failover ganeti master in magru to ganeti7004 * 15:29 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti4008.ulsfo.wmnet * 15:28 dancy@deploy1003: dancy: Continuing with deployment * 15:28 tappof: remove 2026-05 swift log archives from centrallog to free some space ([[phab:T434502|T434502]]) * 15:27 dancy@deploy1003: dancy: Backport for [[gerrit:1324339{{!}}Add /w/deployment-info.php entrypoint]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:25 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324339{{!}}Add /w/deployment-info.php entrypoint]] * 15:20 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2204: Security update * 15:18 dancy@deploy1003: Installation of scap version "4.278.0" completed for 3 hosts * 15:18 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7004.magru.wmnet * 15:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti7004.magru.wmnet * 15:16 dancy@deploy1003: Installing scap version "4.278.0" for 3 host(s) * 15:14 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2204.codfw.wmnet with reason: Maintenance * 15:12 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 15:11 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-master1004.eqiad.wmnet with OS bookworm * 15:11 hashar: Restarting CI Jenkins on contint1003 due to Java upgrade. * 15:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2204 [[phab:T434565|T434565]]', diff saved to https://phabricator.wikimedia.org/P95984 and previous config saved to /var/cache/conftool/dbconfig/20260811-151126-cwilliams.json * 15:10 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti4008.ulsfo.wmnet * 15:10 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti7004.magru.wmnet * 15:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2207 to s2 primary [[phab:T434565|T434565]]', diff saved to https://phabricator.wikimedia.org/P95983 and previous config saved to /var/cache/conftool/dbconfig/20260811-150905-cwilliams.json * 15:08 cezmunsta: Starting s2 codfw failover from db2204 to db2207 - [[phab:T434565|T434565]] * 15:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2207 with weight 0 [[phab:T434565|T434565]]', diff saved to https://phabricator.wikimedia.org/P95982 and previous config saved to /var/cache/conftool/dbconfig/20260811-150402-cwilliams.json * 15:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s2 [[phab:T434565|T434565]] * 14:55 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1010.eqiad.wmnet with OS bookworm * 14:49 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns4003.wikimedia.org with OS trixie * 14:47 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1010.eqiad.wmnet with OS bookworm * 14:47 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7001.wikimedia.org with OS trixie * 14:44 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-presto1010.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:41 btullis@cumin1003: START - Cookbook sre.hosts.provision for host an-presto1010.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:40 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-presto1010.eqiad.wmnet with OS bookworm * 14:39 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 14:39 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-presto1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:36 btullis@cumin1003: START - Cookbook sre.hosts.provision for host an-presto1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:32 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1009.eqiad.wmnet with OS bookworm * 14:32 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 14:31 cwilliams@cumin1003: START - Cookbook sre.mysql.clone of db1157.eqiad.wmnet onto db1272.eqiad.wmnet * 14:30 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7004.magru.wmnet * 14:28 moritzm: failover ganeti master in ulsfo to ganeti4005 * 14:23 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7003.magru.wmnet * 14:23 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti7003.magru.wmnet * 14:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-master1004.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:22 btullis@cumin1003: START - Cookbook sre.hosts.provision for host an-master1004.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:21 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-master1004.eqiad.wmnet with OS bookworm * 14:19 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti4007.ulsfo.wmnet * 14:19 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti4007.ulsfo.wmnet * 14:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1010.eqiad.wmnet with OS bookworm * 14:17 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1008.eqiad.wmnet with OS bookworm * 14:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti7003.magru.wmnet * 14:12 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7003.magru.wmnet * 14:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti4007.ulsfo.wmnet * 14:11 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7002.magru.wmnet * 14:11 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti7002.magru.wmnet * 14:09 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324318{{!}}Revert "wmf-config/ProductionServices: set URL for urldownloader to service record" (T429175)]] (duration: 06m 46s) * 14:09 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7001.wikimedia.org with reason: host reimage * 14:06 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti4007.ulsfo.wmnet * 14:05 kharlan@deploy1003: kharlan: Continuing with deployment * 14:05 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti4006.ulsfo.wmnet * 14:04 jayme: updated calico to v3.30.7 on staging-eqiad - [[phab:T427400|T427400]] * 14:04 kharlan@deploy1003: kharlan: Backport for [[gerrit:1324318{{!}}Revert "wmf-config/ProductionServices: set URL for urldownloader to service record" (T429175)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti4006.ulsfo.wmnet * 14:03 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns4003.wikimedia.org with reason: host reimage * 14:03 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7001.wikimedia.org with reason: host reimage * 14:02 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti7002.magru.wmnet * 14:02 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1324318{{!}}Revert "wmf-config/ProductionServices: set URL for urldownloader to service record" (T429175)]] * 14:02 btullis@dns1004: FAIL - running authdns-update * 14:00 btullis@dns1004: START - running authdns-update * 13:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1008.eqiad.wmnet with reason: host reimage * 13:59 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'. * 13:59 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1313985{{!}}wmf-config/ProductionServices: set URL for urldownloader to service record (T429175)]] (duration: 25m 06s) * 13:58 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7002.magru.wmnet * 13:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti4006.ulsfo.wmnet * 13:57 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns4003.wikimedia.org with reason: host reimage * 13:56 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'. * 13:56 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7001.magru.wmnet * 13:56 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1008.eqiad.wmnet with reason: host reimage * 13:55 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'. * 13:55 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'. * 13:55 kharlan@deploy1003: kharlan, sukhe: Continuing with deployment * 13:53 marostegui: Failover ms2 [[phab:T434288|T434288]] * 13:52 marostegui: Failover ms1 [[phab:T434288|T434288]] * 13:52 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7001.magru.wmnet * 13:51 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti4006.ulsfo.wmnet * 13:48 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti4005.ulsfo.wmnet * 13:48 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti4005.ulsfo.wmnet * 13:44 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti4005.ulsfo.wmnet * 13:40 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1008.eqiad.wmnet with OS bookworm * 13:39 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns4003.wikimedia.org with OS trixie * 13:38 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns7001.wikimedia.org with OS trixie * 13:38 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1008.eqiad.wmnet with OS bookworm * 13:37 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti4005.ulsfo.wmnet * 13:36 kharlan@deploy1003: kharlan, sukhe: Backport for [[gerrit:1313985{{!}}wmf-config/ProductionServices: set URL for urldownloader to service record (T429175)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:34 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1313985{{!}}wmf-config/ProductionServices: set URL for urldownloader to service record (T429175)]] * 13:29 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2034.codfw.wmnet * 13:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2034.codfw.wmnet * 13:25 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.clone (exit_code=99) of db1157.eqiad.wmnet onto db1272.eqiad.wmnet * 13:25 cwilliams@cumin1003: START - Cookbook sre.mysql.clone of db1157.eqiad.wmnet onto db1272.eqiad.wmnet * 13:21 urbanecm@deploy1003: mwscript-k8s job started: namespaceDupes.php --wiki=frwiktionary --fix # [[phab:T415716|T415716]] * 13:21 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2034.codfw.wmnet * 13:20 urbanecm@deploy1003: mwscript-k8s job started: namespaceDupes.php --wiki=frwiktionary # [[phab:T415716|T415716]] * 13:19 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1323350{{!}}[tgwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T415307)]], [[gerrit:1322961{{!}}[slwiki] Revert temporary logo for Wikipedia 25 (Vector legacy + Vector 2022) (T414265)]], [[gerrit:1323827{{!}}[itwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T414320)]] (duration: 08m 00s) * 13:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-coord1003.eqiad.wmnet with OS bookworm * 13:15 urbanecm@deploy1003: urbanecm, superpes: Continuing with deployment * 13:13 urbanecm@deploy1003: urbanecm, superpes: Backport for [[gerrit:1323350{{!}}[tgwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T415307)]], [[gerrit:1322961{{!}}[slwiki] Revert temporary logo for Wikipedia 25 (Vector legacy + Vector 2022) (T414265)]], [[gerrit:1323827{{!}}[itwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T414320)]] synced to the testservers (see https://wiki * 13:11 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1323350{{!}}[tgwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T415307)]], [[gerrit:1322961{{!}}[slwiki] Revert temporary logo for Wikipedia 25 (Vector legacy + Vector 2022) (T414265)]], [[gerrit:1323827{{!}}[itwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T414320)]] * 13:11 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1323779{{!}}[ukwiki] Remove reviewer usergroup (T434252)]], [[gerrit:1323312{{!}}[frwiktionary] Add new Schème namespace and its talk (T415716)]] (duration: 06m 49s) * 13:10 marostegui@dns1004: END - running authdns-update * 13:08 marostegui@dns1004: START - running authdns-update * 13:07 marostegui@cumin1003: dbctl commit (dc=all): 'Repool ms2 [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95980 and previous config saved to /var/cache/conftool/dbconfig/20260811-130725-marostegui.json * 13:06 urbanecm@deploy1003: urbanecm, superpes: Continuing with deployment * 13:06 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1266 to dbctl [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95979 and previous config saved to /var/cache/conftool/dbconfig/20260811-130627-marostegui.json * 13:06 urbanecm@deploy1003: urbanecm, superpes: Backport for [[gerrit:1323779{{!}}[ukwiki] Remove reviewer usergroup (T434252)]], [[gerrit:1323312{{!}}[frwiktionary] Add new Schème namespace and its talk (T415716)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:04 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1323779{{!}}[ukwiki] Remove reviewer usergroup (T434252)]], [[gerrit:1323312{{!}}[frwiktionary] Add new Schème namespace and its talk (T415716)]] * 12:59 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db2253.codfw.wmnet,db[1151,1266].eqiad.wmnet with reason: Switching over ms2 * 12:54 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1157: Using as clone source * 12:53 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1157: Using as clone source * 12:51 marostegui@cumin1003: dbctl commit (dc=all): 'Depool ms2 [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95977 and previous config saved to /var/cache/conftool/dbconfig/20260811-125129-marostegui.json * 12:47 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 12:46 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 12:45 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 12:44 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'recommendation-api-ng' for release 'main' . * 12:44 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'recommendation-api-ng' for release 'main' . * 12:43 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'recommendation-api-ng' for release 'main' . * 12:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-coord1003.eqiad.wmnet with reason: host reimage * 12:43 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'ores-legacy' for release 'main' . * 12:42 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'ores-legacy' for release 'main' . * 12:42 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2165: Security update * 12:40 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-coord1003.eqiad.wmnet with reason: host reimage * 12:39 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'ores-legacy' for release 'main' . * 12:38 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' . * 12:38 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' . * 12:37 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' . * 12:34 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2034.codfw.wmnet * 12:30 jmm@cumin2002: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti-test2001.codfw.wmnet * 12:30 jmm@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host ganeti-test2001.codfw.wmnet * 12:25 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1179.eqiad.wmnet onto db1278.eqiad.wmnet * 12:25 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1179: Pool db1179.eqiad.wmnet in after cloning * 12:23 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-coord1003.eqiad.wmnet with OS bookworm * 12:22 moritzm: failover ganeti master in codfw/routed to ganeti2033 * 12:22 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2033.codfw.wmnet * 12:22 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2033.codfw.wmnet * 12:19 jmm@cumin2002: START - Cookbook sre.hosts.reboot-single for host ganeti-test2001.codfw.wmnet * 12:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 12:18 jmm@cumin2002: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti-test2001.codfw.wmnet * 12:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1008.eqiad.wmnet with OS bookworm * 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2033.codfw.wmnet * 12:07 moritzm: failover ganeti master in ganeti/test to ganeti-test2003 * 12:04 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 12:03 jmm@cumin2003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti4005.ulsfo.wmnet * 12:03 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti4005.ulsfo.wmnet * 12:00 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324283{{!}}Use maximum compression level in SqlBlobStore and SqlBagOStuff (T428377)]] (duration: 11m 37s) * 11:57 jmm@cumin2002: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti-test2002.codfw.wmnet * 11:57 jmm@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti-test2002.codfw.wmnet * 11:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2165: Security update * 11:54 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 11:52 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1324283{{!}}Use maximum compression level in SqlBlobStore and SqlBagOStuff (T428377)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:51 jmm@cumin2002: START - Cookbook sre.hosts.reboot-single for host ganeti-test2002.codfw.wmnet * 11:50 jmm@cumin2002: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti-test2002.codfw.wmnet * 11:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2165.codfw.wmnet with reason: Maintenance * 11:48 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@050d19e] (releasing): [[phab:T434186|T434186]] (duration: 01m 14s) * 11:48 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1324283{{!}}Use maximum compression level in SqlBlobStore and SqlBagOStuff (T428377)]] * 11:47 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@050d19e] (releasing): [[phab:T434186|T434186]] * 11:44 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@050d19e] (releasing): test jenkins deploy for [[phab:T434186|T434186]] (duration: 01m 08s) * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2165 [[phab:T434514|T434514]]', diff saved to https://phabricator.wikimedia.org/P95969 and previous config saved to /var/cache/conftool/dbconfig/20260811-114352-cwilliams.json * 11:43 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@050d19e] (releasing): test jenkins deploy for [[phab:T434186|T434186]] * 11:42 jmm@cumin2003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti-test2003.codfw.wmnet * 11:42 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti-test2003.codfw.wmnet * 11:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2161 to s8 primary [[phab:T434514|T434514]]', diff saved to https://phabricator.wikimedia.org/P95968 and previous config saved to /var/cache/conftool/dbconfig/20260811-114136-cwilliams.json * 11:40 cezmunsta: Starting s8 codfw failover from db2165 to db2161 - [[phab:T434514|T434514]] * 11:40 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1179: Pool db1179.eqiad.wmnet in after cloning * 11:36 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti-test2003.codfw.wmnet * 11:36 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti-test2003.codfw.wmnet * 11:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2161 with weight 0 [[phab:T434514|T434514]]', diff saved to https://phabricator.wikimedia.org/P95966 and previous config saved to /var/cache/conftool/dbconfig/20260811-113449-cwilliams.json * 11:34 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 25 hosts with reason: Primary switchover s8 [[phab:T434514|T434514]] * 11:29 moritzm: installing Python 3.11 security updates * 11:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-presto1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 11:26 btullis@cumin1003: START - Cookbook sre.hosts.provision for host an-presto1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 11:23 btullis@dns1004: END - running authdns-update * 11:21 btullis@dns1004: START - running authdns-update * 11:20 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-presto1008.eqiad.wmnet with OS bookworm * 11:20 moritzm: installing curl security updates * 11:11 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1007.eqiad.wmnet with OS bookworm * 10:45 tappof: bump space for prometheus k8s-dse in codfw * 10:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1007.eqiad.wmnet with reason: host reimage * 10:38 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1007.eqiad.wmnet with reason: host reimage * 10:37 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-coord1004.eqiad.wmnet with OS bookworm * 10:35 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1008.eqiad.wmnet with OS bookworm * 10:34 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1179: Depool db1179.eqiad.wmnet to then clone it to db1278.eqiad.wmnet - marostegui@cumin1003 * 10:34 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1008.eqiad.wmnet with OS bookworm * 10:25 fceratto@cumin1003: dbctl commit (dc=all): 'Remove db1177 [[phab:T433474|T433474]]', diff saved to https://phabricator.wikimedia.org/P95964 and previous config saved to /var/cache/conftool/dbconfig/20260811-102527-fceratto.json * 10:22 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1008.eqiad.wmnet with OS bookworm * 10:22 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1007.eqiad.wmnet with OS bookworm * 10:21 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1006.eqiad.wmnet with OS bookworm * 10:20 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 10:18 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1179: Depool db1179.eqiad.wmnet to then clone it to db1278.eqiad.wmnet - marostegui@cumin1003 * 10:18 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1179.eqiad.wmnet onto db1278.eqiad.wmnet * 10:17 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 10:17 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 10:14 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 10:09 blake@deploy1003: Stopping before sync operations * 10:09 blake@deploy1003: Started scap sync-world: Non-deployment run to populate release values for [[phab:T427668|T427668]] * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 10:04 fceratto@cumin1003: Removing db1177 from zarcillo [[phab:T433474|T433474]] * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1177.eqiad.wmnet * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1177.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:03 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1177.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:00 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1006.eqiad.wmnet with reason: host reimage * 09:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-coord1004.eqiad.wmnet with reason: host reimage * 09:57 marostegui: Failover m1 from db1164 to db1213 - [[phab:T434493|T434493]] * 09:57 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1006.eqiad.wmnet with reason: host reimage * 09:55 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:54 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2232].codfw.wmnet,db[1164,1213,1217].eqiad.wmnet with reason: Primary switchover m1 [[phab:T434493|T434493]] * 09:52 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-coord1004.eqiad.wmnet with reason: host reimage * 09:49 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1213.eqiad.wmnet with OS trixie * 09:49 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1177.eqiad.wmnet * 09:41 moritzm: installing Linux 6.12.101 on Trixie hosts * 09:40 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1006.eqiad.wmnet with OS bookworm * 09:35 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-coord1004.eqiad.wmnet with OS bookworm * 09:28 moritzm: installing node-tar security updates * 09:27 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1213.eqiad.wmnet with reason: host reimage * 09:22 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1213.eqiad.wmnet with reason: host reimage * 09:09 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1177: Decommission * 09:08 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db1177: Decommission * 09:08 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 09:08 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.decommission (exit_code=99) * 09:06 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1213.eqiad.wmnet with OS trixie * 09:06 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 09:05 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1213.eqiad.wmnet with reason: Reimage * 08:53 marostegui@dns1004: END - running authdns-update * 08:51 marostegui@dns1004: START - running authdns-update * 08:48 marostegui: Switchover ms1 master in eqiad [[phab:T434288|T434288]] * 08:48 marostegui@cumin1003: dbctl commit (dc=all): 'Repool ms1 [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95962 and previous config saved to /var/cache/conftool/dbconfig/20260811-084804-marostegui.json * 08:40 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1267 to dbctl [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95961 and previous config saved to /var/cache/conftool/dbconfig/20260811-084054-marostegui.json * 08:29 marostegui: Failover m1 from db1213 to db1164 - [[phab:T434043|T434043]] * 08:25 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2232].codfw.wmnet,db[1164,1213,1217].eqiad.wmnet with reason: Primary switchover m1 [[phab:T434043|T434043]] * 08:22 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db2251.codfw.wmnet,db[1152,1267].eqiad.wmnet with reason: Switching over ms1 * 08:22 marostegui@cumin1003: dbctl commit (dc=all): 'Depool ms1 [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95960 and previous config saved to /var/cache/conftool/dbconfig/20260811-082201-marostegui.json * 08:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: Switching over ms1 * 08:20 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.parsercache (exit_code=99) * 08:20 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 08:20 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1152: Switching over ms1 * 08:19 slyngshede@dns1004: END - running authdns-update * 08:18 moritzm: installing openjdk-21 security updates * 08:17 slyngshede@dns1004: START - running authdns-update * 08:16 moritzm: imported jenkins 2.568.2 to thirdparty/jenkins for trixie-wikimedia * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.12 (duration: 02m 26s) * 03:36 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] (duration: 33m 33s) * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 35s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-10 == * 14:54 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1323973{{!}}mmv.bootstrap: Fix getUrlParam to account for TIFF lossy/lossless param (T434333)]] (duration: 11m 24s) * 14:50 krinkle@deploy1003: krinkle: Continuing with deployment * 14:45 krinkle@deploy1003: krinkle: Backport for [[gerrit:1323973{{!}}mmv.bootstrap: Fix getUrlParam to account for TIFF lossy/lossless param (T434333)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:43 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1323973{{!}}mmv.bootstrap: Fix getUrlParam to account for TIFF lossy/lossless param (T434333)]] * 14:07 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1323967{{!}}updateIsActiveFlagForMentees: Commit the final partial batch (T432959)]] (duration: 10m 33s) * 13:56 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1323967{{!}}updateIsActiveFlagForMentees: Commit the final partial batch (T432959)]] * 13:45 dani@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply * 13:45 dani@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply * 13:45 dani@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply * 13:45 dani@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply * 13:45 dani@deploy1003: helmfile [staging] DONE helmfile.d/services/miscweb: apply * 13:44 dani@deploy1003: helmfile [staging] START helmfile.d/services/miscweb: apply * 13:38 wmde-fisch@deploy1003: Finished scap sync-world: Backport for [[gerrit:1323939{{!}}Enable sub-references on more group2 wikis (batch3) (T432731)]] (duration: 33m 21s) * 13:25 wmde-fisch@deploy1003: wmde-fisch: Continuing with deployment * 13:22 wmde-fisch@deploy1003: wmde-fisch: Backport for [[gerrit:1323939{{!}}Enable sub-references on more group2 wikis (batch3) (T432731)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:05 wmde-fisch@deploy1003: Started scap sync-world: Backport for [[gerrit:1323939{{!}}Enable sub-references on more group2 wikis (batch3) (T432731)]] * 07:57 hashar@deploy1003: Finished deploy [integration/docroot@7772132]: update build dependencies (duration: 00m 13s) * 07:57 hashar@deploy1003: Started deploy [integration/docroot@7772132]: update build dependencies * 07:35 _joe_: restarting squid on urldownloader1006 * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 48s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-09 == * 16:01 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:01 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:01 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:00 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 36s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-08 == * 05:31 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9] (wcqs): [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) (duration: 02m 36s) * 05:28 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9] (wcqs): [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) * 04:56 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 04:55 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 04:47 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) (duration: 19m 22s) * 04:28 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) * 04:19 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) (duration: 00m 06s) * 04:18 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) * 04:17 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) (duration: 00m 28s) * 04:16 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) * 03:52 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 03:52 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 34s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-07 == * 23:30 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:29 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 22:45 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 22:43 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 22:41 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 22:41 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 20:54 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 20:32 andrewbogott: restarting puppetserver service on puppetserver* for [[phab:T434339|T434339]] * 19:52 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:45 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 19:32 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:25 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:22 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 19:21 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 18:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:41 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 18:35 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 18:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 18:22 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 18:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 18:16 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 18:12 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:09 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:08 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:07 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:04 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:01 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:00 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:00 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 17:59 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 17:25 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 17:14 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 17:13 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 17:13 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 17:13 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:54 maryum: Deployed security fix for [[phab:T434278|T434278]] * 16:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 16:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 16:27 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --verbose --scoreLessThan=0.7 --exceptDatasetChecksums=[[phab:T434319|T434319]]-enwiki-models.txt # [[phab:T434319|T434319]] * 16:06 cdobbins@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp5022.eqsin.wmnet with OS trixie * 15:13 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 14:19 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 14:17 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 13:50 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1156.eqiad.wmnet onto db1271.eqiad.wmnet * 13:50 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1271: Pool db1271.eqiad.wmnet in after cloning * 13:02 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1271: Pool db1271.eqiad.wmnet in after cloning * 12:19 jayme: updated calico to v3.30.7 on staging-codfw - [[phab:T427400|T427400]] * 12:09 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 12:06 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 12:06 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 12:05 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 12:02 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1156: Pool db1156.eqiad.wmnet in after cloning * 11:38 bjensen: sudo -i reprepro -C main include trixie-wikimedia $<nowiki>{</nowiki>HOME<nowiki>}</nowiki>/httpbb/trixie/httpbb_$<nowiki>{</nowiki>VERSION?<nowiki>}</nowiki>-1+deb13u1_amd64.changes #[[phab:T434052|T434052]] * 11:35 bjensen: sudo -i reprepro -C main include bookworm-wikimedia $<nowiki>{</nowiki>HOME<nowiki>}</nowiki>/httpbb/bookworm/httpbb_$<nowiki>{</nowiki>VERSION?<nowiki>}</nowiki>-1_amd64.changes #[[phab:T434052|T434052]] * 11:30 marostegui@cumin1003: dbctl commit (dc=all): 'Adding db1271 to dbctl', diff saved to https://phabricator.wikimedia.org/P95945 and previous config saved to /var/cache/conftool/dbconfig/20260807-113006-marostegui.json * 11:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1156: Pool db1156.eqiad.wmnet in after cloning * 10:23 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 10:22 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 10:22 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 10:21 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 10:20 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 10:20 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 10:19 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 10:18 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 10:06 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on 21 hosts with reason: cloning * 10:01 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1156: Depool db1156.eqiad.wmnet to then clone it to db1271.eqiad.wmnet - marostegui@cumin1003 * 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1156: Depool db1156.eqiad.wmnet to then clone it to db1271.eqiad.wmnet - marostegui@cumin1003 * 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1156.eqiad.wmnet onto db1271.eqiad.wmnet * 09:15 jynus: started stress testing db1245 dbs [[phab:T431115|T431115]] * 08:19 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:18 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:16 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:14 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:13 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:10 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:06 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:05 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:00 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 10 days, 0:00:00 on ml-serve1015.eqiad.wmnet with reason: Downtime to get full picture of current BIOS settings beyond what Redfish shows * 08:00 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 07:54 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 07:54 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 07:53 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:52 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:51 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:50 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:49 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:48 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:47 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:45 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:45 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:41 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:38 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:37 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 06:35 jayme: updated istio to 1.29.4 on wikikube eqiad - [[phab:T427401|T427401]] * 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1178.eqiad.wmnet * 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1178.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 06:06 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1178.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 05:55 marostegui@cumin1003: START - Cookbook sre.dns.netbox * 05:49 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1178.eqiad.wmnet * 05:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 05:46 marostegui@cumin1003: Removing db1178 from zarcillo [[phab:T433471|T433471]] * 05:45 marostegui@cumin1003: START - Cookbook sre.mysql.decommission * 02:42 denisse: Extended volume on prometheus2008 for the disk space alert as per https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_host_running_out_of_space * 02:37 denisse: Extended volume on prometheus2007 tor the disk space alert as per https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_host_running_out_of_space * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 56s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-06 == * 21:39 maryum: Deploy security patch for [[phab:T433070|T433070]] * 21:29 maryum: Deploy security patch for [[phab:T434189|T434189]] * 20:48 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] (duration: 08m 12s) * 20:44 aude@deploy1003: lmora, aude, anzx: Continuing with deployment * 20:41 aude@deploy1003: lmora, aude, anzx: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be * 20:41 ebernhardson@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:41 ebernhardson@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 20:40 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] * 20:37 ebernhardson@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:37 ebernhardson@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 20:32 ebernhardson@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:32 ebernhardson@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 20:31 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] (duration: 06m 41s) * 20:27 cjming@deploy1003: cjming, ebernhardson, chlod: Continuing with deployment * 20:26 cjming@deploy1003: cjming, ebernhardson, chlod: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:24 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] * 20:18 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] (duration: 09m 22s) * 20:14 cjming@deploy1003: cjming, tsev: Continuing with deployment * 20:11 cjming@deploy1003: cjming, tsev: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:09 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] * 19:41 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply * 19:40 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply * 19:31 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 19:31 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 19:00 cdobbins@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cp5022.eqsin.wmnet with OS trixie * 18:25 ladsgroup@deploy1003: Finished scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) (duration: 06m 08s) * 18:19 ladsgroup@deploy1003: Started scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) * 18:18 ladsgroup@deploy1003: Stopping before sync operations * 18:17 ladsgroup@deploy1003: Started scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) * 17:55 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 16:50 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 16:35 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1001.eqiad.wmnet with OS bookworm * 16:19 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1002.eqiad.wmnet with reason: host reimage * 16:16 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1002.eqiad.wmnet with reason: host reimage * 16:05 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1001.eqiad.wmnet with reason: host reimage * 16:00 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1001.eqiad.wmnet with reason: host reimage * 15:57 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 15:43 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm * 15:29 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1001.eqiad.wmnet with OS bookworm * 15:29 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:58 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm * 14:57 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-drmrs ([[phab:T428495|T428495]]) * 14:55 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-drmrs ([[phab:T428495|T428495]]) * 14:55 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-ui1001.eqiad.wmnet with OS bookworm * 14:54 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-presto1001.eqiad.wmnet with OS bookworm * 14:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-magru ([[phab:T428495|T428495]]) * 14:49 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-magru ([[phab:T428495|T428495]]) * 14:48 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1001.eqiad.wmnet with OS bookworm * 14:46 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-esams ([[phab:T428495|T428495]]) * 14:44 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-esams ([[phab:T428495|T428495]]) * 14:43 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:42 brouberol@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:42 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:42 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 14:40 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 14:40 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:38 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-ui1001.eqiad.wmnet with reason: host reimage * 14:34 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-presto1001.eqiad.wmnet with reason: host reimage * 14:28 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-ui1001.eqiad.wmnet with reason: host reimage * 14:27 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-presto1001.eqiad.wmnet with reason: host reimage * 14:23 sukhe: sudo cumin -b2 'A:cp-text' "run-puppet-agent --enable 'merging CR 1290731'": [[phab:T425441|T425441]] * 14:18 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo for hosts in the wikimedia.org domain - [[phab:T428495|T428495]] * 14:16 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-presto1001.eqiad.wmnet with OS bookworm * 14:14 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-ui1001.eqiad.wmnet with OS bookworm * 14:12 sukhe: sudo cumin 'A:cp-text' "disable-puppet 'merging CR 1290731'": [[phab:T425441|T425441]] * 14:11 swfrench-wmf: restarted navtiming on webperf1003 - [[phab:T428495|T428495]] * 14:04 swfrench-wmf: begin rolling restart of confd in drmrs, eqiad, esams, magru - [[phab:T428495|T428495]] * 14:04 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm * 14:04 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:02 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-client1002.eqiad.wmnet with OS bookworm * 13:58 swfrench-wmf: authdns update to direct eqiad-associated etcd clients back to eqiad - [[phab:T428495|T428495]] * 13:58 swfrench@dns1004: END - running authdns-update * 13:56 swfrench@dns1004: START - running authdns-update * 13:49 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:44 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:31 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 13:29 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 13:26 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 13:23 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 13:22 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 13:19 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 13:18 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 13:18 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 13:17 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 13:16 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 13:13 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 13:11 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 13:09 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 13:06 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 13:06 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-client1002.eqiad.wmnet with OS bookworm * 13:05 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revision-models' for release 'main' . * 13:05 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:05 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revision-models' for release 'main' . * 13:04 brouberol@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-test-client1002.eqiad.wmnet with OS bookworm * 13:04 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 13:03 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 13:02 aikochou@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:00 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'readability' for release 'main' . * 12:59 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'readability' for release 'main' . * 12:58 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 12:57 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'logo-detection' for release 'main' . * 12:57 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'logo-detection' for release 'main' . * 12:57 aikochou@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 12:55 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:54 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:53 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 12:53 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:50 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 12:48 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 12:46 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'article-models' for release 'main' . * 12:45 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'article-models' for release 'main' . * 12:41 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . * 12:39 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . * 12:38 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-client1002.eqiad.wmnet with OS bookworm * 12:12 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply * 12:12 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply * 12:09 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:08 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 11:58 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2187: Security update * 11:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:24 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:16 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 11:15 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 11:10 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2187: Security update * 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2187.codfw.wmnet with reason: Maintenance * 10:56 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 10:56 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 10:56 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 10:56 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 10:54 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 10:53 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 10:09 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2187: Security update * 10:07 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2187: Security update * 09:39 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms2', diff saved to https://phabricator.wikimedia.org/P95929 and previous config saved to /var/cache/conftool/dbconfig/20260806-093908-marostegui.json * 09:36 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1178 from dbctl [[phab:T433471|T433471]]', diff saved to https://phabricator.wikimedia.org/P95928 and previous config saved to /var/cache/conftool/dbconfig/20260806-093632-marostegui.json * 09:33 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 09:31 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 09:30 topranks: bounce cr3-eqsin<->cr2-eqiad bgp session to disable no-prepend command * 09:20 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2253.codfw.wmnet,db1151.eqiad.wmnet with reason: cloning * 09:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1151: Cloning * 09:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:19 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 09:19 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1151: Cloning * 09:10 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 09:09 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup2003.codfw.wmnet * 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup2003.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 09:06 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup2003.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 09:03 klausman@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:02 jynus@cumin1003: START - Cookbook sre.dns.netbox * 09:02 klausman@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 08:57 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup2003.codfw.wmnet * 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup1003.eqiad.wmnet * 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:54 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms3', diff saved to https://phabricator.wikimedia.org/P95925 and previous config saved to /var/cache/conftool/dbconfig/20260806-085422-marostegui.json * 08:53 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:46 jynus@cumin1003: START - Cookbook sre.dns.netbox * 08:39 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup1003.eqiad.wmnet * 08:29 XioNoX: push pfw policy - [[phab:T434115|T434115]] * 08:14 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 08:00 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 08:00 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:58 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revision-models' for release 'main' . * 07:56 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 07:54 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'readability' for release 'main' . * 07:53 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'logo-detection' for release 'main' . * 07:51 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'llm' for release 'main' . * 07:48 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . * 07:37 jayme: updated istio to 1.29.4 on wikikube codfw - [[phab:T427401|T427401]] * 07:08 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2252.codfw.wmnet,db1153.eqiad.wmnet with reason: cloning * 07:07 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1153: Cloning * 07:07 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1153: Cloning * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 40s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-05 == * 23:24 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1009.eqiad.wmnet with OS bookworm * 23:03 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1009.eqiad.wmnet with reason: host reimage * 22:59 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1009.eqiad.wmnet with reason: host reimage * 22:43 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1009.eqiad.wmnet with OS bookworm * 22:38 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1009.eqiad.wmnet * 22:34 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1009.eqiad.wmnet * 22:25 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1008.eqiad.wmnet with OS bookworm * 22:04 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1008.eqiad.wmnet with reason: host reimage * 22:00 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1008.eqiad.wmnet with reason: host reimage * 21:48 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:47 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:46 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:44 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1008.eqiad.wmnet with OS bookworm * 21:43 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:41 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1008.eqiad.wmnet * 21:36 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1008.eqiad.wmnet * 21:14 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:12 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad * 21:12 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad * 21:10 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=eqiad * 21:08 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:07 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:07 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=eqiad * 21:04 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:03 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1006 * 21:02 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1006 * 21:00 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:56 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:56 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:55 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 20:55 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 20:55 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1007.eqiad.wmnet with OS bookworm * 20:51 vriley@cumin1003: START - Cookbook sre.dns.netbox * 20:43 ebernhardson: [[phab:T434008|T434008]]: changing cloudelastic:9643 from auto_expand_replicas to number_of_replicas * 20:34 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1007.eqiad.wmnet with reason: host reimage * 20:27 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1007.eqiad.wmnet with reason: host reimage * 20:24 cjming: end of UTC late backport window * 20:23 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] (duration: 06m 26s) * 20:18 cjming@deploy1003: cjming: Continuing with deployment * 20:18 cjming@deploy1003: cjming: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:16 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] * 20:12 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1007.eqiad.wmnet with OS bookworm * 20:12 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] (duration: 08m 41s) * 20:08 swfrench@cumin2002: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host conf1007.eqiad.wmnet with OS bookworm * 20:08 jforrester@deploy1003: jforrester: Continuing with deployment * 20:07 jforrester@deploy1003: jforrester: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:03 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] * 19:51 inflatador: [bking@puppetserver1001] ~$ sudo puppetserver ca sign --certname an-worker1189.eqiad.wmnet [[phab:T434142|T434142]] * 19:47 bking@cumin2003: DONE (FAIL) - Cookbook sre.puppet.renew-cert (exit_code=99) for an-worker1189.eqiad.wmnet: Renew puppet certificate - bking@cumin2003 * 19:46 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:30 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1007.eqiad.wmnet with OS trixie * 19:30 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 19:29 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 19:20 swfrench-wmf: silenced EtcdRelicationDown 0cb709a9-f244-4f1e-971f-{{Gerrit|440ec65e7fd7}} - [[phab:T428495|T428495]] * 19:13 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1007.eqiad.wmnet with OS bookworm * 19:12 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1007.eqiad.wmnet with reason: host reimage * 19:09 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1007.eqiad.wmnet * 19:07 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1007.eqiad.wmnet with reason: host reimage * 19:03 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1007.eqiad.wmnet * 18:52 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1007.eqiad.wmnet with OS trixie * 18:52 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1007.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:35 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 18:34 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 18:34 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 18:30 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1007.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:28 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:28 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1007] - vriley@cumin1003" * 18:27 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1007] - vriley@cumin1003" * 18:23 vriley@cumin1003: START - Cookbook sre.dns.netbox * 18:22 vriley@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 18:22 robh@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:19 vriley@cumin1003: START - Cookbook sre.dns.netbox * 18:13 robh@cumin2002: START - Cookbook sre.hosts.provision for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:31 jasmine@cumin2002: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-main-eqiad * 17:12 mutante: LDAP - added vwalters to group ciadmin - [[phab:T433615|T433615]] * 16:58 aokoth@deploy1003: Finished deploy [phabricator/deployment@e2ebca5]: Deploy Phab (duration: 00m 34s) * 16:57 aokoth@deploy1003: Started deploy [phabricator/deployment@e2ebca5]: Deploy Phab * 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad * 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=eqiad * 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad * 16:53 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:41 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2187.codfw.wmnet * 16:41 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2187.codfw.wmnet * 16:41 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker2187.codfw.wmnet * 16:41 cgoubert@cumin2003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker2187.codfw.wmnet * 16:40 jasmine@cumin2002: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-main-eqiad * 16:40 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:34 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:25 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-magru and A:liberica ([[phab:T428495|T428495]]) * 16:23 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-magru and A:liberica ([[phab:T428495|T428495]]) * 16:20 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-drmrs and A:liberica ([[phab:T428495|T428495]]) * 16:19 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-drmrs and A:liberica ([[phab:T428495|T428495]]) * 16:18 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-esams and A:liberica ([[phab:T428495|T428495]]) * 16:16 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-esams and A:liberica ([[phab:T428495|T428495]]) * 16:06 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1159.eqiad.wmnet * 16:06 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1159.eqiad.wmnet * 16:06 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1159.eqiad.wmnet * 16:05 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] (duration: 09m 11s) * 15:58 reedy@deploy1003: reedy: Continuing with deployment * 15:58 reedy@deploy1003: reedy: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:56 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] * 15:54 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1159.eqiad.wmnet with OS trixie * 15:38 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:33 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1159.eqiad.wmnet with reason: host reimage * 15:32 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:27 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1159.eqiad.wmnet with reason: host reimage * 15:10 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1159 * 15:10 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1159 * 15:00 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo for hosts in the wikimedia.org domain - [[phab:T428495|T428495]] * 14:55 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS trixie * 14:54 swfrench-wmf: restarted navtiming on webperf1003 - [[phab:T428495|T428495]] * 14:52 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1159 * 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1159.eqiad.wmnet 129.48.64.10.in-addr.arpa 9.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:52 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1159.eqiad.wmnet 129.48.64.10.in-addr.arpa 9.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1159 - jayme@cumin1003" * 14:52 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1159 - jayme@cumin1003" * 14:49 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:48 jayme@cumin1003: START - Cookbook sre.dns.netbox * 14:47 swfrench-wmf: begin rolling restart of confd in drmrs, eqiad, esams, magru - [[phab:T428495|T428495]] * 14:47 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1159 * 14:46 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:46 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:46 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1159.eqiad.wmnet with OS trixie * 14:45 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:44 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:44 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1159.eqiad.wmnet * 14:43 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:43 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1159.eqiad.wmnet * 14:43 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:43 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1159.eqiad.wmnet * 14:43 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:43 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1157.eqiad.wmnet * 14:43 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1157.eqiad.wmnet * 14:43 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1157.eqiad.wmnet * 14:42 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:42 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:42 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:41 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:41 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:41 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:41 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1003.eqiad.wmnet with OS bookworm * 14:39 swfrench-wmf: authdns update to direct eqiad-associated etcd clients to codfw - [[phab:T428495|T428495]] * 14:39 swfrench@dns1004: END - running authdns-update * 14:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 14:37 swfrench@dns1004: START - running authdns-update * 14:37 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:37 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:35 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:35 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 14:28 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:27 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1157.eqiad.wmnet with OS trixie * 14:27 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:27 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:27 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 14:26 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 14:26 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:26 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 14:26 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 14:26 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:26 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host search-loader1002.eqiad.wmnet with OS trixie * 14:19 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:19 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:15 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1003.eqiad.wmnet with reason: host reimage * 14:14 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:14 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:13 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 * 14:13 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host mc2046 * 14:13 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS trixie * 14:11 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1003.eqiad.wmnet with reason: host reimage * 14:10 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:09 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:09 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:09 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:08 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:08 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1157.eqiad.wmnet with reason: host reimage * 14:08 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:04 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 14:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on search-loader1002.eqiad.wmnet with reason: host reimage * 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=eqiad * 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=eqiad * 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=eqiad * 14:00 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:59 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:58 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1157.eqiad.wmnet with reason: host reimage * 13:57 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on search-loader1002.eqiad.wmnet with reason: host reimage * 13:54 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1003.eqiad.wmnet with OS bookworm * 13:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host search-loader1002.eqiad.wmnet with OS trixie * 13:43 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1157 * 13:42 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1157 * 13:40 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1157 * 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1157.eqiad.wmnet 183.32.64.10.in-addr.arpa 3.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1157.eqiad.wmnet 183.32.64.10.in-addr.arpa 3.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1157 - jayme@cumin1003" * 13:39 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1157 - jayme@cumin1003" * 13:39 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] (duration: 07m 00s) * 13:35 jayme@cumin1003: START - Cookbook sre.dns.netbox * 13:35 reedy@deploy1003: reedy: Continuing with deployment * 13:34 reedy@deploy1003: reedy: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:32 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] * 13:23 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1157 * 13:22 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1157.eqiad.wmnet with OS trixie * 13:22 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1157.eqiad.wmnet * 13:22 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1157.eqiad.wmnet * 13:21 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1157.eqiad.wmnet * 13:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1156.eqiad.wmnet * 13:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1156.eqiad.wmnet * 13:15 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1156.eqiad.wmnet * 13:01 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1156.eqiad.wmnet with OS trixie * 12:42 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1156.eqiad.wmnet with reason: host reimage * 12:38 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1156.eqiad.wmnet with reason: host reimage * 12:32 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 12:31 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 12:30 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 12:28 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 12:26 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 12:24 topranks: update bgp confed settings in eqsin * 12:22 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1156 * 12:22 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1156 * 12:22 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 12:19 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1156 * 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1156.eqiad.wmnet 110.32.64.10.in-addr.arpa 0.1.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:19 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1156.eqiad.wmnet 110.32.64.10.in-addr.arpa 0.1.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1156 - jayme@cumin1003" * 12:19 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1156 - jayme@cumin1003" * 12:17 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:14 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS trixie * 12:09 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:06 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:04 jayme@cumin1003: START - Cookbook sre.dns.netbox * 12:04 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 12:02 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'article-models' for release 'main' . * 12:01 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1156 * 12:01 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1156.eqiad.wmnet with OS trixie * 11:59 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1156.eqiad.wmnet * 11:59 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1156.eqiad.wmnet * 11:59 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1156.eqiad.wmnet * 11:57 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 11:53 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 11:53 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:52 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:52 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:50 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:50 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:50 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:49 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:48 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:47 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:47 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:45 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:45 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:44 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:44 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:44 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:43 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:42 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:38 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:35 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 * 11:35 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host mc2046 * 11:34 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS trixie * 11:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:27 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:21 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:21 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:18 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:18 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:18 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:18 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:13 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:13 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:09 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:08 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:07 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:06 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:06 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:05 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:05 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:04 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 11:04 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:24 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:24 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:17 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:16 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1155.eqiad.wmnet * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1155.eqiad.wmnet * 10:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1155.eqiad.wmnet * 10:14 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:14 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:11 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:11 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:10 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:09 aikochou@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop: sync * 10:09 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:09 aikochou@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop: sync * 10:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:07 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:05 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:05 aikochou@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop: sync * 10:05 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:05 aikochou@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop: sync * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:04 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:04 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:04 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1155.eqiad.wmnet with OS trixie * 09:52 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms1', diff saved to https://phabricator.wikimedia.org/P95918 and previous config saved to /var/cache/conftool/dbconfig/20260805-095212-marostegui.json * 09:44 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1152: after cloning * 09:44 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.parsercache (exit_code=99) * 09:44 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 09:44 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1152: after cloning * 09:43 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1155.eqiad.wmnet with reason: host reimage * 09:40 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1155.eqiad.wmnet with reason: host reimage * 09:32 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 09:32 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:31 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 09:31 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:27 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1155 * 09:27 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1155 * 09:25 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 09:24 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 09:24 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 09:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:23 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 09:23 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 09:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:22 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 09:22 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 09:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:20 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 09:20 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:17 XioNoX: push pfw policies - [[phab:T434038|T434038]] * 09:14 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1155 * 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1155.eqiad.wmnet 109.32.64.10.in-addr.arpa 9.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1155.eqiad.wmnet 109.32.64.10.in-addr.arpa 9.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1155 - jayme@cumin1003" * 09:14 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1155 - jayme@cumin1003" * 09:10 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 09:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:09 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2251.codfw.wmnet,db1152.eqiad.wmnet with reason: cloning * 09:09 jayme@cumin1003: START - Cookbook sre.dns.netbox * 09:08 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 09:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: Cloning * 09:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:05 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 09:05 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1152: Cloning * 08:38 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1155 * 08:37 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1155.eqiad.wmnet with OS trixie * 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1171.eqiad.wmnet * 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1171.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:29 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95913 and previous config saved to /var/cache/conftool/dbconfig/20260805-082908-ladsgroup.json * 08:27 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1171.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:22 jynus@cumin1003: START - Cookbook sre.dns.netbox * 08:18 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249', diff saved to https://phabricator.wikimedia.org/P95912 and previous config saved to /var/cache/conftool/dbconfig/20260805-081823-ladsgroup.json * 08:17 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1171.eqiad.wmnet * 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1150.eqiad.wmnet * 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1150.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:15 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1150.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:15 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 08:14 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1155.eqiad.wmnet * 08:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1155.eqiad.wmnet * 08:14 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1155.eqiad.wmnet * 08:11 jynus@cumin1003: START - Cookbook sre.dns.netbox * 08:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249', diff saved to https://phabricator.wikimedia.org/P95911 and previous config saved to /var/cache/conftool/dbconfig/20260805-080737-ladsgroup.json * 08:05 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1150.eqiad.wmnet * 08:02 marostegui: Depool clouddb1020 (s5,s8) [[phab:T434048|T434048]] * 08:02 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1020.eqiad.wmnet,service=s8 * 08:02 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1020.eqiad.wmnet,service=s5 * 08:02 marostegui: Depool clouddb1018 (s2,s7) [[phab:T434048|T434048]] * 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1018.eqiad.wmnet,service=s7 * 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1018.eqiad.wmnet,service=s2 * 08:01 marostegui: Depool clouddb1017 (s1) [[phab:T434048|T434048]] * 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1017.eqiad.wmnet,service=s1 * 07:59 marostegui: Depool clouddb1016 (s5,s8) [[phab:T434048|T434048]] * 07:59 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s8 * 07:59 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s5 * 07:57 marostegui: Depool clouddb1015 (s4,s6) [[phab:T434048|T434048]] * 07:57 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s6 * 07:57 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s4 * 07:56 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95910 and previous config saved to /var/cache/conftool/dbconfig/20260805-075650-ladsgroup.json * 07:54 marostegui: Depool clouddb1014 (s2,s7) [[phab:T434048|T434048]] * 07:54 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1014.eqiad.wmnet,service=s7 * 07:54 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1014.eqiad.wmnet,service=s2 * 07:53 marostegui: Depool clouddb1013:s1 [[phab:T434048|T434048]] * 07:53 marostegui: Depool clouddb1013:s1 [[phab:T409557|T409557]] * 07:53 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1013.eqiad.wmnet,service=s1 * 07:25 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95909 and previous config saved to /var/cache/conftool/dbconfig/20260805-072529-ladsgroup.json * 07:24 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2249.codfw.wmnet with reason: Maintenance * 07:24 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95908 and previous config saved to /var/cache/conftool/dbconfig/20260805-072426-ladsgroup.json * 07:21 slyngshede@dns1004: END - running authdns-update * 07:19 slyngshede@dns1004: START - running authdns-update * 07:13 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231', diff saved to https://phabricator.wikimedia.org/P95906 and previous config saved to /var/cache/conftool/dbconfig/20260805-071340-ladsgroup.json * 07:02 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231', diff saved to https://phabricator.wikimedia.org/P95905 and previous config saved to /var/cache/conftool/dbconfig/20260805-070253-ladsgroup.json * 06:52 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95904 and previous config saved to /var/cache/conftool/dbconfig/20260805-065206-ladsgroup.json * 06:45 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 06:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95903 and previous config saved to /var/cache/conftool/dbconfig/20260805-062240-ladsgroup.json * 06:21 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2231.codfw.wmnet with reason: Maintenance * 06:21 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95902 and previous config saved to /var/cache/conftool/dbconfig/20260805-062137-ladsgroup.json * 06:10 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215', diff saved to https://phabricator.wikimedia.org/P95901 and previous config saved to /var/cache/conftool/dbconfig/20260805-061051-ladsgroup.json * 06:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215', diff saved to https://phabricator.wikimedia.org/P95900 and previous config saved to /var/cache/conftool/dbconfig/20260805-060004-ladsgroup.json * 05:49 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95899 and previous config saved to /var/cache/conftool/dbconfig/20260805-054918-ladsgroup.json * 05:19 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95898 and previous config saved to /var/cache/conftool/dbconfig/20260805-051939-ladsgroup.json * 05:18 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2215.codfw.wmnet with reason: Maintenance * 04:30 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2201.codfw.wmnet with reason: Maintenance * 03:40 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2197.codfw.wmnet with reason: Maintenance * 03:40 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95897 and previous config saved to /var/cache/conftool/dbconfig/20260805-034036-ladsgroup.json * 03:29 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196', diff saved to https://phabricator.wikimedia.org/P95896 and previous config saved to /var/cache/conftool/dbconfig/20260805-032948-ladsgroup.json * 03:19 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196', diff saved to https://phabricator.wikimedia.org/P95895 and previous config saved to /var/cache/conftool/dbconfig/20260805-031902-ladsgroup.json * 03:08 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95894 and previous config saved to /var/cache/conftool/dbconfig/20260805-030815-ladsgroup.json * 02:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95893 and previous config saved to /var/cache/conftool/dbconfig/20260805-023413-ladsgroup.json * 02:33 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2196.codfw.wmnet with reason: Maintenance * 02:33 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95892 and previous config saved to /var/cache/conftool/dbconfig/20260805-023310-ladsgroup.json * 02:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186', diff saved to https://phabricator.wikimedia.org/P95891 and previous config saved to /var/cache/conftool/dbconfig/20260805-022223-ladsgroup.json * 02:11 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186', diff saved to https://phabricator.wikimedia.org/P95890 and previous config saved to /var/cache/conftool/dbconfig/20260805-021137-ladsgroup.json * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 02:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95889 and previous config saved to /var/cache/conftool/dbconfig/20260805-020051-ladsgroup.json * 01:30 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95888 and previous config saved to /var/cache/conftool/dbconfig/20260805-013029-ladsgroup.json * 01:29 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2186.codfw.wmnet with reason: Maintenance * 00:34 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on dbstore1009.eqiad.wmnet with reason: Maintenance * 00:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95887 and previous config saved to /var/cache/conftool/dbconfig/20260805-003408-ladsgroup.json * 00:23 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264', diff saved to https://phabricator.wikimedia.org/P95886 and previous config saved to /var/cache/conftool/dbconfig/20260805-002322-ladsgroup.json * 00:12 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264', diff saved to https://phabricator.wikimedia.org/P95885 and previous config saved to /var/cache/conftool/dbconfig/20260805-001235-ladsgroup.json * 00:01 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95884 and previous config saved to /var/cache/conftool/dbconfig/20260805-000148-ladsgroup.json == 2026-08-04 == * 23:45 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95883 and previous config saved to /var/cache/conftool/dbconfig/20260804-234508-ladsgroup.json * 23:44 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1264.eqiad.wmnet with reason: Maintenance * 23:44 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95882 and previous config saved to /var/cache/conftool/dbconfig/20260804-234405-ladsgroup.json * 23:33 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237', diff saved to https://phabricator.wikimedia.org/P95881 and previous config saved to /var/cache/conftool/dbconfig/20260804-233317-ladsgroup.json * 23:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237', diff saved to https://phabricator.wikimedia.org/P95880 and previous config saved to /var/cache/conftool/dbconfig/20260804-232230-ladsgroup.json * 23:11 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95879 and previous config saved to /var/cache/conftool/dbconfig/20260804-231144-ladsgroup.json * 22:23 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95878 and previous config saved to /var/cache/conftool/dbconfig/20260804-222345-ladsgroup.json * 22:23 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1237.eqiad.wmnet with reason: Maintenance * 21:13 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1225.eqiad.wmnet with reason: Maintenance * 20:40 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] (duration: 24m 40s) * 20:33 samtar@deploy1003: samtar, kineticpelagic: Continuing with deployment * 20:28 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS bookworm * 20:21 samtar@deploy1003: samtar, kineticpelagic: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:15 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] * 20:13 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 20:09 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 20:00 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1216.eqiad.wmnet with reason: Maintenance * 20:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95877 and previous config saved to /var/cache/conftool/dbconfig/20260804-195957-ladsgroup.json * 19:51 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 19:50 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) mc2046.codfw.wmnet 120.16.192.10.in-addr.arpa 0.2.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:50 jhancock@cumin2002: START - Cookbook sre.dns.wipe-cache mc2046.codfw.wmnet 120.16.192.10.in-addr.arpa 0.2.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host mc2046 - jhancock@cumin2002" * 19:50 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host mc2046 - jhancock@cumin2002" * 19:49 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203', diff saved to https://phabricator.wikimedia.org/P95876 and previous config saved to /var/cache/conftool/dbconfig/20260804-194911-ladsgroup.json * 19:46 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 19:45 jhancock@cumin2002: START - Cookbook sre.hosts.move-vlan for host mc2046 * 19:45 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS bookworm * 19:38 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203', diff saved to https://phabricator.wikimedia.org/P95875 and previous config saved to /var/cache/conftool/dbconfig/20260804-193825-ladsgroup.json * 19:27 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95874 and previous config saved to /var/cache/conftool/dbconfig/20260804-192738-ladsgroup.json * 19:02 mutante: gerrit ssh -p 29418 gerrit.wikimedia.org gerrit index changes {{Gerrit|1320979}} * 18:20 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 18:18 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 18:14 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 18:14 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 18:13 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 18:10 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 18:08 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 18:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95872 and previous config saved to /var/cache/conftool/dbconfig/20260804-180721-ladsgroup.json * 18:07 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 18:06 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1203.eqiad.wmnet with reason: Maintenance * 18:06 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95871 and previous config saved to /var/cache/conftool/dbconfig/20260804-180618-ladsgroup.json * 17:55 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179', diff saved to https://phabricator.wikimedia.org/P95870 and previous config saved to /var/cache/conftool/dbconfig/20260804-175531-ladsgroup.json * 17:55 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1154.eqiad.wmnet * 17:55 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1154.eqiad.wmnet * 17:55 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1154.eqiad.wmnet * 17:50 swfrench@deploy1003: Finished scap sync-world: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] (duration: 04m 05s) * 17:48 swfrench@deploy1003: swfrench: Continuing with deployment * 17:46 swfrench@deploy1003: swfrench: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:45 swfrench@deploy1003: Started scap sync-world: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] * 17:44 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179', diff saved to https://phabricator.wikimedia.org/P95869 and previous config saved to /var/cache/conftool/dbconfig/20260804-174445-ladsgroup.json * 17:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95868 and previous config saved to /var/cache/conftool/dbconfig/20260804-173359-ladsgroup.json * 17:33 swfrench@deploy1003: Finished scap sync-world: Pick up new PHP production image (duration: 28m 32s) * 17:28 aokoth@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on phab1005.eqiad.wmnet with reason: Puppet Failure * 17:05 swfrench@deploy1003: Started scap sync-world: Pick up new PHP production image * 17:00 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 17:00 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 16:54 cgoubert@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on wikikube-worker2187.codfw.wmnet with reason: Hardware issue * 16:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2187.codfw.wmnet * 16:52 mutante: gerrit2003:/var/log/apache2# ln -s /srv/gerrit/site_path/review_site/logs/ gerrit * 16:52 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2187.codfw.wmnet * 16:48 mutante: gerrit2003 - moving old apache logfiles older than 60 days from /var/log/apache2 to /srv/gerrit/site_path/review_site/logs/old/ * 16:33 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 16:32 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 16:29 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 16:29 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 16:28 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 16:28 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 16:27 dzahn@cumin1003: END (PASS) - Cookbook sre.gerrit.restart-gerrit (exit_code=0) Restarting Gerrit on gerrit2003 * 16:27 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 16:27 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95867 and previous config saved to /var/cache/conftool/dbconfig/20260804-162736-ladsgroup.json * 16:27 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 16:26 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1179.eqiad.wmnet with reason: Maintenance * 16:26 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:25 mutante: restarting gerrit - dropped outdated RSA host key * 16:25 dzahn@cumin1003: START - Cookbook sre.gerrit.restart-gerrit Restarting Gerrit on gerrit2003 * 16:24 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95866 and previous config saved to /var/cache/conftool/dbconfig/20260804-162424-ladsgroup.json * 16:24 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 16:23 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 16:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95865 and previous config saved to /var/cache/conftool/dbconfig/20260804-162236-ladsgroup.json * 16:21 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1179.eqiad.wmnet with reason: Maintenance * 16:17 swfrench-wmf: reprepro include php8.3_8.3.33-1+wmf11u1 into component/php83 for bullseye-wikimedia * 16:17 swfrench-wmf: reprepro include php8.3_8.3.33-1+wmf12u1 into component/php83 for bookworm-wikimedia * 16:11 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply * 16:10 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply * 16:10 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mobileapps: apply * 16:09 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mobileapps: apply * 16:09 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply * 16:08 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply * 16:08 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:08 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:07 aokoth@cumin1003: END (PASS) - Cookbook sre.vrts.upgrade (exit_code=0) on VRTS host vrts1003.eqiad.wmnet * 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:05 aokoth@cumin1003: START - Cookbook sre.vrts.upgrade on VRTS host vrts1003.eqiad.wmnet * 16:04 mutante: gerrit2002/gerrit1003/gerrit2003 - rm /etc/gerrit/ssh_host_rsa_key * 15:59 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:59 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:59 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:59 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:56 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 15:55 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:55 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:55 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:49 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 15:49 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:44 Raine: add php8.5 packages to component/php85 - [[phab:T432983|T432983]] * 15:39 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:33 aaron@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 15:33 aaron@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 15:29 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:19 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:19 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:16 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:16 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1154.eqiad.wmnet with OS trixie * 15:16 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:15 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:15 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:06 brennen@deploy1003: Finished deploy [phabricator/deployment@56f4ffd]: deploy phab1004 for [[phab:T433981|T433981]] (duration: 00m 43s) * 15:05 brennen@deploy1003: Started deploy [phabricator/deployment@56f4ffd]: deploy phab1004 for [[phab:T433981|T433981]] * 15:05 aaron@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 15:04 aaron@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 15:02 brennen@deploy1003: Finished deploy [phabricator/deployment@56f4ffd]: deploy phab2003 for [[phab:T433981|T433981]] (duration: 00m 51s) * 15:01 brennen@deploy1003: Started deploy [phabricator/deployment@56f4ffd]: deploy phab2003 for [[phab:T433981|T433981]] * 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1004.eqiad.wmnet with reason: deployment * 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1005.eqiad.wmnet with reason: deployment * 14:58 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab2003.codfw.wmnet with reason: deployment * 14:55 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1154.eqiad.wmnet with reason: host reimage * 14:51 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1154.eqiad.wmnet with reason: host reimage * 14:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 14:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 14:38 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync * 14:38 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync * 14:38 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync * 14:37 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync * 14:37 ottomata: roll restart eventgate-main to pick up stream config change - [[phab:T433507|T433507]] * 14:37 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-main: sync * 14:36 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-main: sync * 14:36 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1154 * 14:36 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1154 * 14:34 otto@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] (duration: 08m 39s) * 14:34 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1154 * 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1154.eqiad.wmnet 108.32.64.10.in-addr.arpa 8.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:34 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1154.eqiad.wmnet 108.32.64.10.in-addr.arpa 8.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1154 - jayme@cumin1003" * 14:34 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1154 - jayme@cumin1003" * 14:30 otto@deploy1003: otto: Continuing with deployment * 14:30 jayme@cumin1003: START - Cookbook sre.dns.netbox * 14:28 otto@deploy1003: otto: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:26 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1154 * 14:26 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1154.eqiad.wmnet with OS trixie * 14:26 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1154.eqiad.wmnet * 14:26 otto@deploy1003: Started scap sync-world: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] * 14:26 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1154.eqiad.wmnet * 14:26 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1154.eqiad.wmnet * 14:17 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 14:16 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 14:15 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 14:14 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 14:13 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 14:13 swfrench@dns1004: END - running authdns-update * 14:13 Msz2001: Finished deployments for UTC afternoon backport window * 14:13 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 14:13 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] (duration: 07m 58s) * 14:11 swfrench@dns1004: START - running authdns-update * 14:08 mszwarc@deploy1003: javiermonton, mszwarc, mpostoronca: Continuing with deployment * 14:07 mszwarc@deploy1003: javiermonton, mszwarc, mpostoronca: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] synced to the testser * 14:05 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] * 14:03 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 13:49 swfrench@cumin2002: conftool action : set/pooled=yes; selector: name=wikikube-worker2330.codfw.wmnet * 13:49 swfrench@cumin2002: conftool action : set/pooled=no; selector: name=wikikube-worker2330.codfw.wmnet * 13:48 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] (duration: 09m 19s) * 13:45 swfrench@dns1004: END - running authdns-update * 13:44 mszwarc@deploy1003: mszwarc, jforrester: Continuing with deployment * 13:43 swfrench@dns1004: START - running authdns-update * 13:41 mszwarc@deploy1003: mszwarc, jforrester: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:38 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] * 13:33 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 13:33 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1154.eqiad.wmnet * 13:32 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 13:32 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 13:31 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 13:31 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:31 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:29 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1154.eqiad.wmnet * 13:28 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1154.eqiad.wmnet * 13:28 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1154.eqiad.wmnet * 13:28 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1141.eqiad.wmnet * 13:28 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1141.eqiad.wmnet * 13:28 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1141.eqiad.wmnet * 13:22 otto@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply * 13:22 otto@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply * 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1096.eqiad.wmnet with OS trixie * 13:05 swfrench@dns1004: END - running authdns-update * 13:03 swfrench@dns1004: START - running authdns-update * 12:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 12:43 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 1:00:00 on db1171.eqiad.wmnet with reason: decom * 12:42 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 1:00:00 on db1150.eqiad.wmnet with reason: decom * 12:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 12:38 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1164,1217].eqiad.wmnet with reason: cloning * 12:33 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2096.codfw.wmnet with OS trixie * 12:22 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1096.eqiad.wmnet with OS trixie * 12:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2096.codfw.wmnet with reason: host reimage * 12:14 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1141.eqiad.wmnet with OS trixie * 12:10 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2096.codfw.wmnet with reason: host reimage * 12:10 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1289.eqiad.wmnet * 12:05 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1289.eqiad.wmnet * 12:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1288.eqiad.wmnet * 11:59 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1288.eqiad.wmnet * 11:59 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1287.eqiad.wmnet * 11:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1097.eqiad.wmnet with OS trixie * 11:54 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1287.eqiad.wmnet * 11:54 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1286.eqiad.wmnet * 11:53 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1141.eqiad.wmnet with reason: host reimage * 11:51 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2096.codfw.wmnet with OS trixie * 11:49 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1141.eqiad.wmnet with reason: host reimage * 11:48 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1286.eqiad.wmnet * 11:48 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1284.eqiad.wmnet * 11:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2095.codfw.wmnet with OS trixie * 11:43 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1284.eqiad.wmnet * 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1283.eqiad.wmnet * 11:42 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on ml-serve1015.eqiad.wmnet with reason: Downtime to get full picture of current BIOS settings beyond what Redfish shows * 11:39 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad * 11:39 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:37 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1283.eqiad.wmnet * 11:37 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1282.eqiad.wmnet * 11:37 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad * 11:37 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:33 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1141 * 11:33 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1141 * 11:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 11:32 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1141 * 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1141.eqiad.wmnet 156.48.64.10.in-addr.arpa 6.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:32 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1141.eqiad.wmnet 156.48.64.10.in-addr.arpa 6.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1141 - jayme@cumin1003" * 11:32 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1141 - jayme@cumin1003" * 11:32 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1282.eqiad.wmnet * 11:32 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1281.eqiad.wmnet * 11:32 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad * 11:32 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:29 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 11:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1097.eqiad.wmnet with reason: host reimage * 11:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2095.codfw.wmnet with OS trixie * 11:26 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1281.eqiad.wmnet * 11:26 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1280.eqiad.wmnet * 11:25 jayme@cumin1003: START - Cookbook sre.dns.netbox * 11:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1097.eqiad.wmnet with reason: host reimage * 11:22 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1141 * 11:21 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1141.eqiad.wmnet with OS trixie * 11:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1280.eqiad.wmnet * 11:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1279.eqiad.wmnet * 11:20 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin with reason: upgrade new Nokia swtiches in eqsin to SR Linux v26 * 11:17 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1141.eqiad.wmnet * 11:16 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1141.eqiad.wmnet * 11:16 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1141.eqiad.wmnet * 11:16 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be2095.codfw.wmnet with OS trixie * 11:15 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1279.eqiad.wmnet * 11:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1278.eqiad.wmnet * 11:14 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1139.eqiad.wmnet * 11:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1139.eqiad.wmnet * 11:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1070.eqiad.wmnet with OS trixie * 11:13 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1140.eqiad.wmnet * 11:13 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1140.eqiad.wmnet * 11:13 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1140.eqiad.wmnet * 11:09 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1278.eqiad.wmnet * 11:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1071.eqiad.wmnet with OS trixie * 11:05 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1097.eqiad.wmnet with OS trixie * 11:04 mvernon@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be1097.eqiad.wmnet with OS trixie * 11:02 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1140.eqiad.wmnet with OS trixie * 11:02 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1097.eqiad.wmnet with OS trixie * 11:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1096.eqiad.wmnet with OS trixie * 11:00 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1139.eqiad.wmnet * 11:00 jayme@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1139.eqiad.wmnet with OS trixie * 10:56 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1069.eqiad.wmnet with OS trixie * 10:56 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 10:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1070.eqiad.wmnet with reason: host reimage * 10:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 10:45 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1071.eqiad.wmnet with reason: host reimage * 10:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 10:41 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1096.eqiad.wmnet with OS trixie * 10:41 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1140.eqiad.wmnet with reason: host reimage * 10:39 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1071.eqiad.wmnet with reason: host reimage * 10:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1070.eqiad.wmnet with reason: host reimage * 10:38 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1139.eqiad.wmnet with reason: host reimage * 10:37 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1140.eqiad.wmnet with reason: host reimage * 10:35 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1069.eqiad.wmnet with reason: host reimage * 10:33 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1095.eqiad.wmnet with OS trixie * 10:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 10:33 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1139.eqiad.wmnet with reason: host reimage * 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1069.eqiad.wmnet with reason: host reimage * 10:23 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1140 * 10:23 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1140 * 10:23 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 10:22 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1140 * 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1140.eqiad.wmnet 155.48.64.10.in-addr.arpa 5.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:21 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1071.eqiad.wmnet with OS trixie * 10:21 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1140.eqiad.wmnet 155.48.64.10.in-addr.arpa 5.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1140 - jayme@cumin1003" * 10:21 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1140 - jayme@cumin1003" * 10:21 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1071 * 10:21 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1070.eqiad.wmnet with OS trixie * 10:21 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1070 * 10:20 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 10:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1095.eqiad.wmnet with OS trixie * 10:17 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1139 * 10:17 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1139 * 10:17 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be1095.eqiad.wmnet with OS trixie * 10:15 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1139 * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1139.eqiad.wmnet 194.32.64.10.in-addr.arpa 4.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:15 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1139.eqiad.wmnet 194.32.64.10.in-addr.arpa 4.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1139 - jayme@cumin1003" * 10:15 jayme@cumin1003: START - Cookbook sre.dns.netbox * 10:15 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1139 - jayme@cumin1003" * 10:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2095.codfw.wmnet with OS trixie * 10:13 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1069.eqiad.wmnet with OS trixie * 10:12 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1069 * 10:11 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1140 * 10:11 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1140.eqiad.wmnet with OS trixie * 10:11 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1140.eqiad.wmnet * 10:10 jayme@cumin1003: START - Cookbook sre.dns.netbox * 10:10 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1139 * 10:10 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1140.eqiad.wmnet * 10:10 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1140.eqiad.wmnet * 10:10 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1139.eqiad.wmnet with OS trixie * 10:09 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1139.eqiad.wmnet * 10:08 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1139.eqiad.wmnet * 10:08 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1139.eqiad.wmnet * 10:01 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2094.codfw.wmnet with OS trixie * 09:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 09:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 09:53 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 09:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 09:44 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:44 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2094.codfw.wmnet with reason: host reimage * 09:34 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2094.codfw.wmnet with reason: host reimage * 09:34 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1071 * 09:33 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1070 * 09:33 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1095.eqiad.wmnet with OS trixie * 09:27 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1069 * 09:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:23 brouberol@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM archiva1002.wikimedia.org * 09:20 brouberol@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM archiva1002.wikimedia.org * 09:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1277.eqiad.wmnet * 09:13 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2094.codfw.wmnet with OS trixie * 09:13 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 09:12 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 09:12 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:12 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1277.eqiad.wmnet * 09:12 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1276.eqiad.wmnet * 09:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1094.eqiad.wmnet with OS trixie * 09:06 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1276.eqiad.wmnet * 09:06 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1275.eqiad.wmnet * 09:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2093.codfw.wmnet with OS trixie * 09:01 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1275.eqiad.wmnet * 09:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1274.eqiad.wmnet * 08:56 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1274.eqiad.wmnet * 08:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1273.eqiad.wmnet * 08:50 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1273.eqiad.wmnet * 08:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1272.eqiad.wmnet * 08:49 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:49 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1094.eqiad.wmnet with reason: host reimage * 08:45 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1272.eqiad.wmnet * 08:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1094.eqiad.wmnet with reason: host reimage * 08:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2093.codfw.wmnet with reason: host reimage * 08:38 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:38 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2093.codfw.wmnet with reason: host reimage * 08:35 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 08:34 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 08:29 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:28 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:26 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1271.eqiad.wmnet * 08:23 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1094.eqiad.wmnet with OS trixie * 08:21 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 08:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1271.eqiad.wmnet * 08:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1270.eqiad.wmnet * 08:15 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1270.eqiad.wmnet * 08:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1269.eqiad.wmnet * 08:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2093.codfw.wmnet with OS trixie * 08:09 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1269.eqiad.wmnet * 08:09 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1268.eqiad.wmnet * 08:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2092.codfw.wmnet with OS trixie * 08:04 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1268.eqiad.wmnet * 08:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1267.eqiad.wmnet * 07:59 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1267.eqiad.wmnet * 07:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1093.eqiad.wmnet with OS trixie * 07:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1266.eqiad.wmnet * 07:56 jynus: running extra backups to test db1285 [[phab:T433826|T433826]] * 07:51 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1266.eqiad.wmnet * 07:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2092.codfw.wmnet with reason: host reimage * 07:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1093.eqiad.wmnet with reason: host reimage * 07:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2092.codfw.wmnet with reason: host reimage * 07:32 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1093.eqiad.wmnet with reason: host reimage * 07:29 jynus: running extra backups to test db1265 [[phab:T433825|T433825]] * 07:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2092.codfw.wmnet with OS trixie * 07:11 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1093.eqiad.wmnet with OS trixie * 06:50 slyngshede@dns1004: END - running authdns-update * 06:48 slyngshede@dns1004: START - running authdns-update * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.11 (duration: 02m 29s) * 03:38 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] (duration: 32m 57s) * 03:23 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 03:22 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 03:05 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 32s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 00:45 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] (duration: 06m 20s) * 00:41 cjming@deploy1003: cjming: Continuing with deployment * 00:41 cjming@deploy1003: cjming: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:39 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] == 2026-08-03 == * 23:58 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply * 23:57 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply * 23:29 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cp5021.eqsin.wmnet * 23:29 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cp5021.eqsin.wmnet * 23:27 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cp5021.eqsin.wmnet * 23:26 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cp5021.eqsin.wmnet * 23:18 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 23:17 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 22:56 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: sync * 22:56 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: sync * 22:36 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 22:36 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 22:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host search-loader2002.codfw.wmnet with OS trixie * 21:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on search-loader2002.codfw.wmnet with reason: host reimage * 21:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on search-loader2002.codfw.wmnet with reason: host reimage * 21:42 dancy@deploy1003: Stopping before sync operations * 21:41 dancy@deploy1003: Started scap sync-world: testing * 21:39 dancy@deploy1003: Installation of scap version "4.277.0" completed for 3 hosts * 21:37 dancy@deploy1003: Installing scap version "4.277.0" for 3 host(s) * 21:37 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] (duration: 06m 13s) * 21:33 dancy@deploy1003: dancy: Continuing with deployment * 21:32 dancy@deploy1003: dancy: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:31 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] * 21:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host search-loader2002.codfw.wmnet with OS trixie * 21:03 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] (duration: 06m 34s) * 20:59 dancy@deploy1003: dancy: Continuing with deployment * 20:58 dancy@deploy1003: dancy: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:56 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] * 20:52 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] (duration: 06m 23s) * 20:48 cjming@deploy1003: cjming: Continuing with deployment * 20:47 cjming@deploy1003: cjming: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:46 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] * 20:42 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] (duration: 07m 36s) * 20:38 arlolra@deploy1003: arlolra: Continuing with deployment * 20:36 arlolra@deploy1003: arlolra: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:34 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] * 20:16 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] (duration: 08m 26s) * 20:12 krinkle@deploy1003: krinkle: Continuing with deployment * 20:09 krinkle@deploy1003: krinkle: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] * 19:45 jasmine@cumin2002: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-main-codfw * 18:58 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] (duration: 09m 23s) * 18:53 krinkle@deploy1003: krinkle: Continuing with deployment * 18:53 jasmine@cumin2002: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-main-codfw * 18:50 krinkle@deploy1003: krinkle: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:48 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] * 18:37 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] (duration: 10m 13s) * 18:34 dzahn@cumin2002: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host codesearch2001.codfw.wmnet * 18:34 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host codesearch2001.codfw.wmnet with OS trixie * 18:33 krinkle@deploy1003: krinkle: Continuing with deployment * 18:29 krinkle@deploy1003: krinkle: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:27 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] * 18:19 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 18:18 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on codesearch2001.codfw.wmnet with reason: host reimage * 18:14 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 18:14 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:12 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on codesearch2001.codfw.wmnet with reason: host reimage * 18:11 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 18:11 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 18:10 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 18:02 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 18:02 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 18:01 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 18:01 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 17:55 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host codesearch2001.codfw.wmnet with OS trixie * 17:54 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:54 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) codesearch2001.codfw.wmnet on all recursors * 17:53 dzahn@cumin2002: START - Cookbook sre.dns.wipe-cache codesearch2001.codfw.wmnet on all recursors * 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:48 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:41 dzahn@cumin2002: START - Cookbook sre.dns.netbox * 17:41 dzahn@cumin2002: START - Cookbook sre.ganeti.makevm for new host codesearch2001.codfw.wmnet * 17:37 dzahn@cumin2002: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host codesearch1001.eqiad.wmnet * 17:37 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host codesearch1001.eqiad.wmnet with OS trixie * 17:24 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on codesearch1001.eqiad.wmnet with reason: host reimage * 17:17 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on codesearch1001.eqiad.wmnet with reason: host reimage * 17:08 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host codesearch1001.eqiad.wmnet with OS trixie * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 17:06 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:06 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) codesearch1001.eqiad.wmnet on all recursors * 17:06 dzahn@cumin2002: START - Cookbook sre.dns.wipe-cache codesearch1001.eqiad.wmnet on all recursors * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 17:05 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 17:04 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:04 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 16:58 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 16:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2091.codfw.wmnet with OS trixie * 16:54 ebernhardson@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 16:54 ebernhardson@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 16:49 ebernhardson@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 16:49 ebernhardson@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 16:46 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1092.eqiad.wmnet with OS trixie * 16:43 ebernhardson@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 16:43 ebernhardson@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 16:43 dzahn@cumin2002: START - Cookbook sre.dns.netbox * 16:43 dzahn@cumin2002: START - Cookbook sre.ganeti.makevm for new host codesearch1001.eqiad.wmnet * 16:41 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 16:41 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 16:40 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2091.codfw.wmnet with reason: host reimage * 16:37 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 16:35 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2091.codfw.wmnet with reason: host reimage * 16:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1092.eqiad.wmnet with reason: host reimage * 16:24 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 16:24 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 16:23 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1092.eqiad.wmnet with reason: host reimage * 16:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2091.codfw.wmnet with OS trixie * 16:03 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1092.eqiad.wmnet with OS trixie * 16:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2090.codfw.wmnet with OS trixie * 15:51 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 15:51 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 15:51 jiji@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 15:50 jiji@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 15:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2090.codfw.wmnet with reason: host reimage * 15:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2090.codfw.wmnet with reason: host reimage * 15:33 jhathaway@dns1004: END - running authdns-update * 15:31 jhathaway@dns1004: START - running authdns-update * 15:26 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1091.eqiad.wmnet with OS trixie * 15:25 dancy@deploy1003: Installation of scap version "4.276.1" completed for 3 hosts * 15:23 dancy@deploy1003: Installing scap version "4.276.1" for 3 host(s) * 15:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2090.codfw.wmnet with OS trixie * 15:12 marostegui@cumin1003: dbctl commit (dc=all): 'Repool db2245, db2246, db2247 and db2248 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95857 and previous config saved to /var/cache/conftool/dbconfig/20260803-151212-marostegui.json * 15:09 dancy@deploy1003: Started scap sync-world: testing * 15:09 dancy@deploy1003: Installation of scap version "4.277.0" completed for 3 hosts * 15:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1091.eqiad.wmnet with reason: host reimage * 15:07 dancy@deploy1003: Installing scap version "4.277.0" for 3 host(s) * 15:03 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1091.eqiad.wmnet with reason: host reimage * 14:49 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1091.eqiad.wmnet with OS trixie * 14:33 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2089.codfw.wmnet with OS trixie * 14:29 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 14:27 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 14:18 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 14:16 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 14:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2089.codfw.wmnet with reason: host reimage * 14:10 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 14:10 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 14:09 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2089.codfw.wmnet with reason: host reimage * 13:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2089.codfw.wmnet with OS trixie * 13:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2088.codfw.wmnet with OS trixie * 13:40 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1090.eqiad.wmnet with OS trixie * 13:22 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1090.eqiad.wmnet with reason: host reimage * 13:22 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] (duration: 14m 34s) * 13:19 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1090.eqiad.wmnet with reason: host reimage * 13:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2088.codfw.wmnet with reason: host reimage * 13:16 aude@deploy1003: aude, mhorsey: Continuing with deployment * 13:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2088.codfw.wmnet with reason: host reimage * 13:12 aude@deploy1003: aude, mhorsey: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] * 13:05 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1090.eqiad.wmnet with OS trixie * 12:58 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2088.codfw.wmnet with OS trixie * 12:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db[2245-2247].codfw.wmnet * 12:49 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2247: Rebooting db2247.codfw.wmnet * 12:49 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2247: Rebooting db2247.codfw.wmnet * 12:42 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2246: Rebooting db2246.codfw.wmnet * 12:42 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2246: Rebooting db2246.codfw.wmnet * 12:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2087.codfw.wmnet with OS trixie * 12:37 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1089.eqiad.wmnet with OS trixie * 12:34 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2245: Rebooting db2245.codfw.wmnet * 12:34 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2245: Rebooting db2245.codfw.wmnet * 12:34 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db[2245-2247].codfw.wmnet * 12:32 kamila@deploy1003: Finished scap sync-world: rebuild after base image update (duration: 30m 26s) * 12:28 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 12:22 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2087.codfw.wmnet with reason: host reimage * 12:19 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1089.eqiad.wmnet with reason: host reimage * 12:14 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2087.codfw.wmnet with reason: host reimage * 12:14 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1089.eqiad.wmnet with reason: host reimage * 12:03 kamila@deploy1003: Started scap sync-world: rebuild after base image update * 12:00 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1089.eqiad.wmnet with OS trixie * 12:00 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2087.codfw.wmnet with OS trixie * 11:35 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db[2245-2248].codfw.wmnet * 11:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db[2245-2248].codfw.wmnet * 11:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2086.codfw.wmnet with OS trixie * 11:26 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 11:26 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 11:25 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 11:25 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 11:24 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1088.eqiad.wmnet with OS trixie * 11:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db[2245-2248].codfw.wmnet with reason: Checking network * 11:21 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 11:20 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 11:19 marostegui@dns1004: END - running authdns-update * 11:17 marostegui@dns1004: START - running authdns-update * 11:10 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:10 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 11:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2086.codfw.wmnet with reason: host reimage * 11:09 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:08 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 11:08 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:07 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 11:07 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:07 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 11:06 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop: apply * 11:06 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop: apply * 11:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1088.eqiad.wmnet with reason: host reimage * 11:05 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop: apply * 11:04 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop: apply * 11:04 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop: apply * 11:04 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop: apply * 11:02 marostegui@dns1004: END - running authdns-update * 11:02 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2086.codfw.wmnet with reason: host reimage * 11:01 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1088.eqiad.wmnet with reason: host reimage * 11:00 marostegui@dns1004: START - running authdns-update * 10:53 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] (duration: 10m 57s) * 10:51 cmooney@dns3003: END - running authdns-update * 10:49 cmooney@dns3003: START - running authdns-update * 10:47 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1088.eqiad.wmnet with OS trixie * 10:47 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2086.codfw.wmnet with OS trixie * 10:47 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 10:46 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:46 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:46 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new reverse ranges for eqsin CR switch links - cmooney@cumin1003" * 10:46 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new reverse ranges for eqsin CR switch links - cmooney@cumin1003" * 10:42 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] * 10:41 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 10:36 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2245, db2246 and db2247 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95855 and previous config saved to /var/cache/conftool/dbconfig/20260803-103652-marostegui.json * 10:35 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2248 from s4 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95854 and previous config saved to /var/cache/conftool/dbconfig/20260803-103535-marostegui.json * 10:27 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 10:27 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 10:26 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 10:24 kart_: cxserver: Add referencePunctuation config ([[phab:T97231|T97231]]) * 10:24 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 10:23 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:23 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:23 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:22 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:22 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply * 10:21 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply * 10:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2085.codfw.wmnet with OS trixie * 10:20 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply * 10:20 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply * 10:18 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply * 10:18 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply * 10:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1087.eqiad.wmnet with OS trixie * 09:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2085.codfw.wmnet with reason: host reimage * 09:43 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1087.eqiad.wmnet with reason: host reimage * 09:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2085.codfw.wmnet with reason: host reimage * 09:40 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1087.eqiad.wmnet with reason: host reimage * 09:26 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1087.eqiad.wmnet with OS trixie * 09:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2085.codfw.wmnet with OS trixie * 09:13 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2084.codfw.wmnet with OS trixie * 09:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1086.eqiad.wmnet with OS trixie * 08:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2084.codfw.wmnet with reason: host reimage * 08:50 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2084.codfw.wmnet with reason: host reimage * 08:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1086.eqiad.wmnet with reason: host reimage * 08:39 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1086.eqiad.wmnet with reason: host reimage * 08:38 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:38 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:37 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 08:37 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 08:35 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2084.codfw.wmnet with OS trixie * 08:34 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 08:34 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:27 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1086.eqiad.wmnet with OS trixie * 08:09 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1218: Repool after a crash * 08:07 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2083.codfw.wmnet with OS trixie * 08:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1085.eqiad.wmnet with OS trixie * 07:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2083.codfw.wmnet with reason: host reimage * 07:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1085.eqiad.wmnet with reason: host reimage * 07:40 kart_: Updated cxsever to 2026-07-16-140518-production ([[phab:T97231|T97231]]) * 07:39 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply * 07:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2083.codfw.wmnet with reason: host reimage * 07:38 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply * 07:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1085.eqiad.wmnet with reason: host reimage * 07:37 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] (duration: 32m 40s) * 07:33 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply * 07:33 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply * 07:25 jdlrobson@deploy1003: jdlrobson: Continuing with deployment * 07:24 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2083.codfw.wmnet with OS trixie * 07:24 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1085.eqiad.wmnet with OS trixie * 07:23 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1218: Repool after a crash * 07:21 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:09 marostegui: Drop renamed tables [[phab:T425074|T425074]] * 07:04 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] * 06:55 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply * 06:54 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 46s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-02 == * 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 01m 03s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-01 == * 03:30 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:30 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:30 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:30 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 34s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-31 == * 17:41 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 17:41 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 17:40 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 17:40 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 15:33 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2195: Testing * 15:02 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:02 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 15:02 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 14:48 pt1979@cumin2002: START - Cookbook sre.dns.netbox * 14:47 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2195: Testing * 14:22 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 14:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2195: Testing * 14:21 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 14:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2195.codfw.wmnet with reason: Testing * 14:16 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 14:04 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 14:04 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 13:30 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1048.eqiad.wmnet with OS trixie * 13:22 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2195: Testing * 13:22 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 13:19 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2195: Testing * 13:18 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 13:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2195: Testing * 13:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 13:05 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 13:05 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 13:04 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 13:04 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 12:50 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lswtest-d8-eqiad * 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:53 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:42 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 6515 * 11:37 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 6515 * 11:28 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:27 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:07 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2082.codfw.wmnet with OS trixie * 10:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2082.codfw.wmnet with reason: host reimage * 10:42 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2082.codfw.wmnet with reason: host reimage * 10:28 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2082.codfw.wmnet with OS trixie * 10:02 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 09:52 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 09:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts * 09:16 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts * 08:57 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 08:46 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:42 gkyziridis@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 08:37 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 08:37 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 08:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 08:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 08:11 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:11 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:08 filippo@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudvirt1048 * 08:07 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 08:07 filippo@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudvirt1048 * 08:06 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 08:01 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 08:00 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 07:19 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1048.eqiad.wmnet with reason: host reimage * 07:13 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1048.eqiad.wmnet with reason: host reimage * 07:11 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 07:11 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 07:09 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 07:09 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 06:57 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:56 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:48 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:48 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:44 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie * 06:34 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1048.eqiad.wmnet with OS trixie * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 54s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 00:57 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] (duration: 11m 04s) * 00:53 dreamyjazz@deploy1003: dreamyjazz, jforrester: Continuing with deployment * 00:48 dreamyjazz@deploy1003: dreamyjazz, jforrester: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:46 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] == 2026-07-30 == * 21:37 dancy@deploy1003: Installation of scap version "4.276.1" completed for 3 hosts * 21:35 dancy@deploy1003: Installing scap version "4.276.1" for 3 host(s) * 21:24 dancy@deploy1003: Installation of scap version "4.276.0" completed for 3 hosts * 21:22 dancy@deploy1003: Installing scap version "4.276.0" for 3 host(s) * 21:15 maryum: Deployed security fix for [[phab:T430601|T430601]] * 20:13 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] (duration: 09m 20s) * 20:07 arlolra@deploy1003: osleger, arlolra: Continuing with deployment * 20:05 arlolra@deploy1003: osleger, arlolra: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:03 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] * 19:29 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:29 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:25 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service * 19:24 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 19:24 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:24 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:24 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 19:23 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service * 19:20 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1084.eqiad.wmnet with OS trixie * 18:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1084.eqiad.wmnet with reason: host reimage * 18:52 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1084.eqiad.wmnet with reason: host reimage * 18:41 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 18:40 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 18:39 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1084.eqiad.wmnet with OS trixie * 18:25 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 18:15 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 17:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1083.eqiad.wmnet with OS trixie * 17:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1048: Maintenance * 17:36 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new security plugin settings - bking@cumin2003 - [[phab:T350516|T350516]] * 17:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1083.eqiad.wmnet with reason: host reimage * 17:26 inflatador: bking@apt1002 `reprepro --noskipold --component thirdparty/opensearch3 update trixie-wikimedia` [[phab:T433624|T433624]] * 17:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1083.eqiad.wmnet with reason: host reimage * 17:23 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 17:20 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 17:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2081.codfw.wmnet with OS trixie * 17:11 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new security plugin settings - bking@cumin2003 - [[phab:T350516|T350516]] * 17:10 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1083.eqiad.wmnet with OS trixie * 16:55 root@cumin1003: START - Cookbook sre.mysql.pool pool es1048: Maintenance * 16:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2081.codfw.wmnet with reason: host reimage * 16:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1048 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95833 and previous config saved to /var/cache/conftool/dbconfig/20260730-165053-cwilliams.json * 16:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1048.eqiad.wmnet with reason: Maintenance * 16:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1040: Maintenance * 16:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2081.codfw.wmnet with reason: host reimage * 16:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1082.eqiad.wmnet with OS trixie * 16:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2081.codfw.wmnet with OS trixie * 16:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2097.codfw.wmnet with OS trixie * 16:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1082.eqiad.wmnet with reason: host reimage * 16:08 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1082.eqiad.wmnet with reason: host reimage * 16:08 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 16:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1047: Maintenance * 16:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2080.codfw.wmnet with OS trixie * 16:04 root@cumin1003: START - Cookbook sre.mysql.pool pool es1040: Maintenance * 16:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1040: Maintenance * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new logging settings - bking@cumin2003 - [[phab:T324335|T324335]] * 15:58 root@cumin1003: START - Cookbook sre.mysql.pool pool es1040: Maintenance * 15:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1040 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95827 and previous config saved to /var/cache/conftool/dbconfig/20260730-155324-cwilliams.json * 15:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1040.eqiad.wmnet with reason: Maintenance * 15:50 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1082.eqiad.wmnet with OS trixie * 15:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2048: Maintenance * 15:44 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 24s) * 15:43 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2080.codfw.wmnet with reason: host reimage * 15:38 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new logging settings - bking@cumin2003 - [[phab:T324335|T324335]] * 15:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2080.codfw.wmnet with reason: host reimage * 15:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 15:30 mvernon@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be2097.codfw.wmnet with OS trixie * 15:23 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1081.eqiad.wmnet with OS trixie * 15:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2098.codfw.wmnet with OS trixie * 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - mvernon@cumin2003" * 15:18 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be2097.codfw.wmnet with OS trixie * 15:18 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - mvernon@cumin2003" * 15:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 15:17 root@cumin1003: START - Cookbook sre.mysql.pool pool es1047: Maintenance * 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2080.codfw.wmnet with OS trixie * 15:13 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be2097.codfw.wmnet with OS trixie * 15:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1047 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95820 and previous config saved to /var/cache/conftool/dbconfig/20260730-151200-cwilliams.json * 15:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1047.eqiad.wmnet with reason: Maintenance * 15:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: Maintenance * 15:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 15:04 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 15:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1081.eqiad.wmnet with reason: host reimage * 15:00 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1081.eqiad.wmnet with reason: host reimage * 15:00 root@cumin1003: START - Cookbook sre.mysql.pool pool es2048: Maintenance * 15:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 14:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2079.codfw.wmnet with OS trixie * 14:56 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 14:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2048 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95816 and previous config saved to /var/cache/conftool/dbconfig/20260730-145510-cwilliams.json * 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2048.codfw.wmnet with reason: Maintenance * 14:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2040: Maintenance * 14:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 14:51 tchin@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] (duration: 06m 48s) * 14:47 tchin@deploy1003: jforrester, tchin: Continuing with deployment * 14:47 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 14:47 tchin@deploy1003: jforrester, tchin: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:45 tchin@deploy1003: Started scap sync-world: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] * 14:42 sukhe@puppetserver1001: conftool action : set/weight=1; selector: cluster=urldownloader,service=squid * 14:42 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader,service=squid * 14:42 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1081.eqiad.wmnet with OS trixie * 14:39 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 14:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2079.codfw.wmnet with reason: host reimage * 14:36 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2098.codfw.wmnet with OS trixie * 14:32 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2079.codfw.wmnet with reason: host reimage * 14:30 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] (duration: 06m 31s) * 14:27 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 14:26 mszwarc@deploy1003: mszwarc: Continuing with deployment * 14:25 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:25 root@cumin1003: START - Cookbook sre.mysql.pool pool es1038: Maintenance * 14:25 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1038: Maintenance * 14:23 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] * 14:21 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] (duration: 11m 19s) * 14:20 root@cumin1003: START - Cookbook sre.mysql.pool pool es1038: Maintenance * 14:14 stran@deploy1003: stran: Continuing with deployment * 14:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1038 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95810 and previous config saved to /var/cache/conftool/dbconfig/20260730-141439-cwilliams.json * 14:14 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1038.eqiad.wmnet with reason: Maintenance * 14:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1036: Maintenance * 14:13 stran@deploy1003: stran: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2079.codfw.wmnet with OS trixie * 14:09 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] * 14:08 root@cumin1003: START - Cookbook sre.mysql.pool pool es2040: Maintenance * 14:08 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2040: Maintenance * 14:03 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] (duration: 31m 41s) * 14:03 root@cumin1003: START - Cookbook sre.mysql.pool pool es2040: Maintenance * 14:03 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2040 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95806 and previous config saved to /var/cache/conftool/dbconfig/20260730-135643-cwilliams.json * 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2040.codfw.wmnet with reason: Maintenance * 13:56 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:56 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2038: Maintenance * 13:55 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:52 lucaswerkmeister-wmde@deploy1003: migr, lucaswerkmeister-wmde: Continuing with deployment * 13:49 lucaswerkmeister-wmde@deploy1003: migr, lucaswerkmeister-wmde: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:49 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:48 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 13:45 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 13:32 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:32 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] * 13:28 root@cumin1003: START - Cookbook sre.mysql.pool pool es1036: Maintenance * 13:28 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1036: Maintenance * 13:22 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:22 root@cumin1003: START - Cookbook sre.mysql.pool pool es1036: Maintenance * 13:20 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1036 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95800 and previous config saved to /var/cache/conftool/dbconfig/20260730-131727-cwilliams.json * 13:17 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1036.eqiad.wmnet with reason: Maintenance * 13:17 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] (duration: 10m 31s) * 13:16 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2022\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 13:13 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, stran: Continuing with deployment * 13:10 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:10 root@cumin1003: START - Cookbook sre.mysql.pool pool es2038: Maintenance * 13:10 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2038: Maintenance * 13:08 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, stran: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie * 13:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2047: Maintenance * 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048 cloud-private - filippo@cumin1003" * 13:07 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048 cloud-private - filippo@cumin1003" * 13:06 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] * 13:04 root@cumin1003: START - Cookbook sre.mysql.pool pool es2038: Maintenance * 13:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2078.codfw.wmnet with OS trixie * 13:01 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2038 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95797 and previous config saved to /var/cache/conftool/dbconfig/20260730-125919-cwilliams.json * 12:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2038.codfw.wmnet with reason: Maintenance * 12:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2078.codfw.wmnet with reason: host reimage * 12:37 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2078.codfw.wmnet with reason: host reimage * 12:37 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] (duration: 06m 51s) * 12:33 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 12:32 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 12:32 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 12:32 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:30 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] * 12:19 root@cumin1003: START - Cookbook sre.mysql.pool pool es2047: Maintenance * 12:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2078.codfw.wmnet with OS trixie * 12:18 dcausse@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 12:18 dcausse@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 12:15 dcausse@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 12:14 dcausse@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 12:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2047 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95793 and previous config saved to /var/cache/conftool/dbconfig/20260730-121404-cwilliams.json * 12:13 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2047.codfw.wmnet with reason: Maintenance * 12:13 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2036: Maintenance * 12:05 ayounsi@dns1004: END - running authdns-update * 12:02 ayounsi@dns1004: START - running authdns-update * 11:51 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2077.codfw.wmnet with OS trixie * 11:48 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:46 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1080.eqiad.wmnet with OS trixie * 11:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1226: Maintenance * 11:41 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2077.codfw.wmnet with reason: host reimage * 11:28 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2077.codfw.wmnet with reason: host reimage * 11:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1080.eqiad.wmnet with reason: host reimage * 11:27 root@cumin1003: START - Cookbook sre.mysql.pool pool es2036: Maintenance * 11:27 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2036: Maintenance * 11:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1080.eqiad.wmnet with reason: host reimage * 11:21 root@cumin1003: START - Cookbook sre.mysql.pool pool es2036: Maintenance * 11:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2036 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95786 and previous config saved to /var/cache/conftool/dbconfig/20260730-111633-cwilliams.json * 11:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2036.codfw.wmnet with reason: Maintenance * 11:08 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2077.codfw.wmnet with OS trixie * 11:07 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 11:03 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie * 11:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1226: Maintenance * 10:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1226 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95783 and previous config saved to /var/cache/conftool/dbconfig/20260730-104801-cwilliams.json * 10:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1226.eqiad.wmnet with reason: Maintenance * 10:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1214: Maintenance * 10:27 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2035: Maintenance * 10:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2076.codfw.wmnet with OS trixie * 10:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1214: Maintenance * 09:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1214 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95775 and previous config saved to /var/cache/conftool/dbconfig/20260730-095451-cwilliams.json * 09:54 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1214.eqiad.wmnet with reason: Maintenance * 09:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1209: Maintenance * 09:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2076.codfw.wmnet with reason: host reimage * 09:42 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool es2035: Maintenance * 09:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.netbox.update-extras (exit_code=0) rolling restart_daemons on A:netbox * 09:41 ayounsi@cumin1003: START - Cookbook sre.netbox.update-extras rolling restart_daemons on A:netbox * 09:40 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2035: Maintenance * 09:39 ayounsi@cumin1003: END (PASS) - Cookbook sre.netbox.update-extras (exit_code=0) rolling restart_daemons on A:netbox-canary * 09:39 ayounsi@cumin1003: START - Cookbook sre.netbox.update-extras rolling restart_daemons on A:netbox-canary * 09:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2076.codfw.wmnet with reason: host reimage * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 09:34 root@cumin1003: START - Cookbook sre.mysql.pool pool es2035: Maintenance * 09:32 lucaswerkmeister-wmde@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 09:32 lucaswerkmeister-wmde@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 09:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2035 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95771 and previous config saved to /var/cache/conftool/dbconfig/20260730-092910-cwilliams.json * 09:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2035.codfw.wmnet with reason: Maintenance * 09:19 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2076.codfw.wmnet with OS trixie * 09:18 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 09:17 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie * 09:07 root@cumin1003: START - Cookbook sre.mysql.pool pool db1209: Maintenance * 09:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 23 hosts * 09:04 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Remove cable label from interfaces descriptions - ayounsi@cumin1003 * 09:04 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:02 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Remove cable label from interfaces descriptions - ayounsi@cumin1003 * 09:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1209 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95767 and previous config saved to /var/cache/conftool/dbconfig/20260730-090133-cwilliams.json * 09:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1209.eqiad.wmnet with reason: Maintenance * 09:01 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1192: Maintenance * 08:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1252: Maintenance * 08:57 jayme@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 08:56 jayme@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 08:53 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 23 hosts * 08:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:51 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:50 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1263: Maintenance * 08:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:23 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 08:15 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie * 08:14 root@cumin1003: START - Cookbook sre.mysql.pool pool db1192: Maintenance * 08:13 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2075.codfw.wmnet with OS trixie * 08:12 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1252: Maintenance * 08:11 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1252: Maintenance * 08:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1252: Maintenance * 08:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1192 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95754 and previous config saved to /var/cache/conftool/dbconfig/20260730-080611-cwilliams.json * 08:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1192.eqiad.wmnet with reason: Maintenance * 08:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1178: Maintenance * 08:05 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS bullseye * 07:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1263: Maintenance * 07:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1263 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95751 and previous config saved to /var/cache/conftool/dbconfig/20260730-075106-cwilliams.json * 07:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[1260-1262].eqiad.wmnet with reason: Maintenance * 07:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2075.codfw.wmnet with reason: host reimage * 07:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1263.eqiad.wmnet with reason: Maintenance * 07:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2075.codfw.wmnet with reason: host reimage * 07:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance * 07:38 dcausse@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:38 dcausse@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 07:35 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db1252', diff saved to https://phabricator.wikimedia.org/P95748 and previous config saved to /var/cache/conftool/dbconfig/20260730-073510-marostegui.json * 07:26 klausman@dns2004: END - running authdns-update * 07:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 07:25 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2075.codfw.wmnet with OS trixie * 07:24 klausman@dns2004: START - running authdns-update * 07:23 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host an-test-master1003.eqiad.wmnet * 07:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db1178: Maintenance * 07:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1178 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95746 and previous config saved to /var/cache/conftool/dbconfig/20260730-071112-cwilliams.json * 07:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1178.eqiad.wmnet with reason: Maintenance * 07:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1177: Maintenance * 06:24 root@cumin1003: START - Cookbook sre.mysql.pool pool db1177: Maintenance * 06:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1177 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95741 and previous config saved to /var/cache/conftool/dbconfig/20260730-061736-cwilliams.json * 06:17 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1177.eqiad.wmnet with reason: Maintenance * 06:17 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1172: Maintenance * 05:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1218.eqiad.wmnet with reason: crashed * 05:41 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db1217 it crashed', diff saved to https://phabricator.wikimedia.org/P95737 and previous config saved to /var/cache/conftool/dbconfig/20260730-054111-marostegui.json * 05:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95736 and previous config saved to /var/cache/conftool/dbconfig/20260730-053422-cwilliams.json * 05:30 root@cumin1003: START - Cookbook sre.mysql.pool pool db1172: Maintenance * 05:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95734 and previous config saved to /var/cache/conftool/dbconfig/20260730-052414-cwilliams.json * 05:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1172 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95733 and previous config saved to /var/cache/conftool/dbconfig/20260730-052354-cwilliams.json * 05:23 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1172.eqiad.wmnet with reason: Maintenance * 05:23 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1167: Maintenance * 05:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95731 and previous config saved to /var/cache/conftool/dbconfig/20260730-051406-cwilliams.json * 05:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95729 and previous config saved to /var/cache/conftool/dbconfig/20260730-050358-cwilliams.json * 04:47 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 04:47 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 04:47 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 04:35 root@cumin1003: START - Cookbook sre.mysql.pool pool db1167: Maintenance * 04:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1167 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95726 and previous config saved to /var/cache/conftool/dbconfig/20260730-042923-cwilliams.json * 04:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 04:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1167.eqiad.wmnet with reason: Maintenance * 04:22 pt1979@cumin2002: START - Cookbook sre.dns.netbox * 04:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95725 and previous config saved to /var/cache/conftool/dbconfig/20260730-040337-cwilliams.json * 04:03 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:38 brett@cumin2002: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool eqsin [reason: Switch upgrade maintenance window complete, [[phab:T433097|T433097]]] * 01:38 brett@cumin2002: START - Cookbook sre.dns.admin DNS admin: pool eqsin [reason: Switch upgrade maintenance window complete, [[phab:T433097|T433097]]] == 2026-07-29 == * 23:57 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin,mr1-eqsin IPv6,mr1-eqsin.oob,mr1-eqsin.oob IPv6 with reason: connection issue * 22:54 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 22:53 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 22:53 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 22:53 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:25 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2022.codfw.wmnet, repooling source-only afterwards * 22:20 brett@cumin2002: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool eqsin [reason: Switch upgrade maintenance window, [[phab:T433097|T433097]]] * 22:20 brett@cumin2002: START - Cookbook sre.dns.admin DNS admin: depool eqsin [reason: Switch upgrade maintenance window, [[phab:T433097|T433097]]] * 22:01 apine@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 22:00 apine@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 21:59 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 21:58 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 21:58 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 21:58 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 21:32 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1048.eqiad.wmnet with OS trixie * 21:25 pt1979@cumin2002: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be2097.codfw.wmnet with OS bullseye * 21:16 zabe@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki=metawiki 'Mental Health Resource Center' 'Safety Resource Center/Mental Health' Zabe --reason 'per request [[:phab:T433118{{!}}T433118]]' * 21:12 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2022.codfw.wmnet, repooling source-only afterwards * 21:12 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] (duration: 12m 53s) * 21:12 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2015\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 21:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1253: Maintenance * 21:08 aaron@deploy1003: aaron: Continuing with deployment * 21:01 aaron@deploy1003: aaron: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:59 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] * 20:52 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] (duration: 21m 57s) * 20:48 aaron@deploy1003: aaron: Continuing with deployment * 20:32 aaron@deploy1003: aaron: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:30 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] * 20:24 root@cumin1003: START - Cookbook sre.mysql.pool pool db1253: Maintenance * 20:19 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] (duration: 08m 07s) * 20:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1253 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95719 and previous config saved to /var/cache/conftool/dbconfig/20260729-201810-cwilliams.json * 20:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1253.eqiad.wmnet with reason: Maintenance * 20:17 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1231: Maintenance * 20:15 aaron@deploy1003: bpirkle, aaron: Continuing with deployment * 20:13 aaron@deploy1003: bpirkle, aaron: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:12 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie * 20:11 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] * 20:11 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1048.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:09 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1048.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:09 pt1979@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 20:04 pt1979@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 19:47 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 19:43 pt1979@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye * 19:41 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 19:37 zabe: zabe@deploy1003:~$ mwscript-k8s --comment='[[phab:T433529|T433529]]' --follow -- resetAuthenticationThrottle.php --wiki=aawiki --signup --ip=89.36.114.94 * 19:36 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] (duration: 06m 49s) * 19:32 zabe@deploy1003: zabe: Continuing with deployment * 19:31 zabe@deploy1003: zabe: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:31 root@cumin1003: START - Cookbook sre.mysql.pool pool db1231: Maintenance * 19:29 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] * 19:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1231 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95714 and previous config saved to /var/cache/conftool/dbconfig/20260729-192454-cwilliams.json * 19:24 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1231.eqiad.wmnet with reason: Maintenance * 19:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1227: Maintenance * 19:22 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 19:22 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 19:21 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:21 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1048] - vriley@cumin1003" * 19:21 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1048] - vriley@cumin1003" * 19:19 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 19:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1251: Maintenance * 19:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95711 and previous config saved to /var/cache/conftool/dbconfig/20260729-191756-cwilliams.json * 19:16 vriley@cumin1003: START - Cookbook sre.dns.netbox * 19:11 dduvall: rolling back wmf.13 to group0 due to [[phab:T433457|T433457]] (cc [[phab:T430832|T430832]]) * 19:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95709 and previous config saved to /var/cache/conftool/dbconfig/20260729-190748-cwilliams.json * 19:01 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2022.codfw.wmnet with OS bookworm * 19:01 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2021.codfw.wmnet, repooling source-only afterwards * 18:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95707 and previous config saved to /var/cache/conftool/dbconfig/20260729-185740-cwilliams.json * 18:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95704 and previous config saved to /var/cache/conftool/dbconfig/20260729-184732-cwilliams.json * 18:37 root@cumin1003: START - Cookbook sre.mysql.pool pool db1227: Maintenance * 18:34 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2022.codfw.wmnet with reason: host reimage * 18:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1227 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95701 and previous config saved to /var/cache/conftool/dbconfig/20260729-183117-cwilliams.json * 18:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1227.eqiad.wmnet with reason: Maintenance * 18:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1202: Maintenance * 18:30 root@cumin1003: START - Cookbook sre.mysql.pool pool db1251: Maintenance * 18:27 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2022.codfw.wmnet with reason: host reimage * 18:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1251 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95698 and previous config saved to /var/cache/conftool/dbconfig/20260729-182428-cwilliams.json * 18:24 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lvs2014.codfw.wmnet * 18:24 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for lvs2014.codfw.wmnet * 18:24 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1251.eqiad.wmnet with reason: Maintenance * 18:23 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1235: Maintenance * 18:22 brett@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 18:19 brett@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 18:19 brett@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 18:17 brett@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 18:17 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 18:16 mutante: removing jenkins during the train - living on the edge - no, just kidding, jenkins has migrated to dedicated machines, nothing should happen * 18:15 brett@cumin2002: END (ERROR) - Cookbook sre.loadbalancer.restart-pybal (exit_code=97) rolling-restart of pybal on P<nowiki>{</nowiki>lvs2014.codfw.wmnet<nowiki>}</nowiki> and A:lvs ([[phab:T428495|T428495]]) * 18:15 mutante: CI: contint1002/contint2002: apt-get remove --purge jenkins - jenkins be gone - [[phab:T418521|T418521]] * 18:13 brett@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on P<nowiki>{</nowiki>lvs2014.codfw.wmnet<nowiki>}</nowiki> and A:lvs ([[phab:T428495|T428495]]) * 18:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2022 * 18:08 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2022 * 18:03 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T428495|T428495]] * 18:03 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2022 * 18:02 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2022.codfw.wmnet 211.48.192.10.in-addr.arpa 1.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:02 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2022.codfw.wmnet 211.48.192.10.in-addr.arpa 1.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:02 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:02 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2022 - bking@cumin2003" * 18:02 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2022 - bking@cumin2003" * 17:57 bking@cumin2003: START - Cookbook sre.dns.netbox * 17:56 brett@cumin2002: END (FAIL) - Cookbook sre.loadbalancer.restart-pybal (exit_code=1) rolling-restart of pybal on A:lvs-codfw and A:lvs ([[phab:T428495|T428495]]) * 17:55 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo - [[phab:T428495|T428495]] * 17:54 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2022 * 17:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2022.codfw.wmnet with OS bookworm * 17:50 brett@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on A:lvs-codfw and A:lvs ([[phab:T428495|T428495]]) * 17:47 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2021.codfw.wmnet, repooling source-only afterwards * 17:47 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 14s) * 17:47 swfrench-wmf: authdns-update to direct codfw, eqsin, ulsfo etcd clients back to codfw - [[phab:T428495|T428495]] * 17:47 swfrench@dns1004: END - running authdns-update * 17:47 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 17:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95692 and previous config saved to /var/cache/conftool/dbconfig/20260729-174713-cwilliams.json * 17:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance * 17:46 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1249: Maintenance * 17:45 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2015\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 17:45 swfrench@dns1004: START - running authdns-update * 17:44 root@cumin1003: START - Cookbook sre.mysql.pool pool db1202: Maintenance * 17:41 akhatun: Deployed refinery using scap, then deployed onto hdfs * 17:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1202 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95688 and previous config saved to /var/cache/conftool/dbconfig/20260729-173759-cwilliams.json * 17:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1202.eqiad.wmnet with reason: Maintenance * 17:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1194: Maintenance * 17:37 root@cumin1003: START - Cookbook sre.mysql.pool pool db1235: Maintenance * 17:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1230: Maintenance * 17:30 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1235 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95684 and previous config saved to /var/cache/conftool/dbconfig/20260729-173051-cwilliams.json * 17:30 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1235.eqiad.wmnet with reason: Maintenance * 17:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1234: Maintenance * 17:26 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (thin): Regular analytics weekly train THIN [analytics/refinery@56695674] (duration: 02m 02s) * 17:24 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (thin): Regular analytics weekly train THIN [analytics/refinery@56695674] * 17:23 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567]: Regular analytics weekly train [analytics/refinery@56695674] (duration: 06m 20s) * 17:20 dancy@deploy1003: Finished scap sync-world: Testing delay_messageblobstore_purge: true (duration: 06m 29s) * 17:17 akhatun@deploy1003: Started deploy [analytics/refinery@5669567]: Regular analytics weekly train [analytics/refinery@56695674] * 17:17 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] (duration: 00m 22s) * 17:16 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] * 17:13 dancy@deploy1003: Started scap sync-world: Testing delay_messageblobstore_purge: true * 17:05 mutante: CI: contint1002/contint2002 - restarted httpd to be extra sure all is cleaned up - https://integration.wikimedia.org/ci/ is up and running [[phab:T418521|T418521]] * 17:04 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 17:03 mutante: CI: contint1002/contint2002 - rm /etc/apache2/jenkins_proxy - removing legacy jenkins proxy config - jenkins is on new dedicated machines and uses jenkins_proxy_ext config [[phab:T418521|T418521]] * 17:02 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] (duration: 36m 25s) * 17:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1249: Maintenance * 16:59 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 16:54 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2015.codfw.wmnet, repooling source-only afterwards * 16:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1249 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95674 and previous config saved to /var/cache/conftool/dbconfig/20260729-165339-cwilliams.json * 16:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1249.eqiad.wmnet with reason: Maintenance * 16:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1248: Maintenance * 16:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1194: Maintenance * 16:47 swfrench-wmf: silenced EtcdReplicationDown 57b2b421-1cc9-4e38-9276-{{Gerrit|94f223fd231c}} - [[phab:T428495|T428495]] * 16:46 tchin@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/eventstreams-internal: apply * 16:46 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye * 16:46 tchin@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/eventstreams-internal: apply * 16:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1230: Maintenance * 16:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1194 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95669 and previous config saved to /var/cache/conftool/dbconfig/20260729-164422-cwilliams.json * 16:44 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 16:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1194.eqiad.wmnet with reason: Maintenance * 16:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1191: Maintenance * 16:43 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host an-test-master1003.eqiad.wmnet * 16:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db1234: Maintenance * 16:43 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 16:43 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Rolling back deployment * 16:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host an-test-master1004.eqiad.wmnet * 16:41 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 16:40 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 16:40 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 16:39 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 16:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1230 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95667 and previous config saved to /var/cache/conftool/dbconfig/20260729-163932-cwilliams.json * 16:39 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 16:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1230.eqiad.wmnet with reason: Maintenance * 16:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1207: Maintenance * 16:38 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 16:37 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host an-test-master1004.eqiad.wmnet * 16:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1234 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95664 and previous config saved to /var/cache/conftool/dbconfig/20260729-163719-cwilliams.json * 16:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1234.eqiad.wmnet with reason: Maintenance * 16:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1079.eqiad.wmnet with OS trixie * 16:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1232: Maintenance * 16:34 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1259: Maintenance * 16:28 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] (duration: 06m 57s) * 16:28 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:26 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] * 16:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1051 hosts * 16:21 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] * 16:20 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2006.codfw.wmnet with OS bookworm * 16:19 akhatun: Deploying Refinery at {{Gerrit|56695674}} as part of weekly train * 16:18 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1079.eqiad.wmnet with reason: host reimage * 16:16 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] (duration: 15m 36s) * 16:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2021.codfw.wmnet with OS bookworm * 16:14 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1079.eqiad.wmnet with reason: host reimage * 16:12 topranks: hot-swap line card in FPC0 on cr1-eqiad with replacement MPC10E from Juniper [[phab:T426343|T426343]] * 16:10 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Continuing with deployment * 16:07 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db1248: Maintenance * 16:01 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] * 16:00 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 16:00 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 15:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1248 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95651 and previous config saved to /var/cache/conftool/dbconfig/20260729-155956-cwilliams.json * 15:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1248.eqiad.wmnet with reason: Maintenance * 15:59 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2006.codfw.wmnet with reason: host reimage * 15:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1247: Maintenance * 15:59 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2074.codfw.wmnet with OS trixie * 15:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1191: Maintenance * 15:57 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:55 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1079.eqiad.wmnet with OS trixie * 15:55 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2006.codfw.wmnet with reason: host reimage * 15:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1207: Maintenance * 15:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1191 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95646 and previous config saved to /var/cache/conftool/dbconfig/20260729-155104-cwilliams.json * 15:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1191.eqiad.wmnet with reason: Maintenance * 15:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1181: Maintenance * 15:49 root@cumin1003: START - Cookbook sre.mysql.pool pool db1232: Maintenance * 15:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2021.codfw.wmnet with reason: host reimage * 15:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1207 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95643 and previous config saved to /var/cache/conftool/dbconfig/20260729-154735-cwilliams.json * 15:47 root@cumin1003: START - Cookbook sre.mysql.pool pool db1259: Maintenance * 15:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1207.eqiad.wmnet with reason: Maintenance * 15:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1200: Maintenance * 15:46 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] (duration: 31m 59s) * 15:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 15:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1232 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95640 and previous config saved to /var/cache/conftool/dbconfig/20260729-154330-cwilliams.json * 15:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1232.eqiad.wmnet with reason: Maintenance * 15:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1219: Maintenance * 15:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2015.codfw.wmnet, repooling source-only afterwards * 15:41 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 18s) * 15:41 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1259 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95638 and previous config saved to /var/cache/conftool/dbconfig/20260729-154107-cwilliams.json * 15:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1259.eqiad.wmnet with reason: Maintenance * 15:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1254: Maintenance * 15:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2015.codfw.wmnet with OS bookworm * 15:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2021.codfw.wmnet with reason: host reimage * 15:36 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2006.codfw.wmnet with OS bookworm * 15:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 15:35 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Continuing with deployment * 15:33 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2074.codfw.wmnet with OS trixie * 15:33 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 15:32 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:29 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:28 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be2074.codfw.wmnet with OS trixie * 15:28 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2006.codfw.wmnet * 15:26 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1078.eqiad.wmnet with OS trixie * 15:25 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 15:25 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:22 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2006.codfw.wmnet * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2021 * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2021 * 15:19 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2021 * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2021.codfw.wmnet 210.48.192.10.in-addr.arpa 0.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:19 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2021.codfw.wmnet 210.48.192.10.in-addr.arpa 0.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2021 - bking@cumin2003" * 15:19 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2021 - bking@cumin2003" * 15:14 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] * 15:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2015.codfw.wmnet with reason: host reimage * 15:11 root@cumin1003: START - Cookbook sre.mysql.pool pool db1247: Maintenance * 15:11 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on ml-serve2004.codfw.wmnet with reason: [[phab:T433478|T433478]] * 15:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2015.codfw.wmnet with reason: host reimage * 15:10 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on ml-serve2002.codfw.wmnet with reason: [[phab:T433476|T433476]] * 15:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 15:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1247 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95625 and previous config saved to /var/cache/conftool/dbconfig/20260729-150459-cwilliams.json * 15:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1247.eqiad.wmnet with reason: Maintenance * 15:04 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:04 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1244: Maintenance * 15:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1078.eqiad.wmnet with reason: host reimage * 15:03 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1005.wikimedia.org * 15:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db1181: Maintenance * 15:01 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2098.codfw.wmnet with OS bullseye * 15:00 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye * 15:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1200: Maintenance * 14:59 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 14:59 root@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1285.eqiad.wmnet with OS trixie * 14:59 Amir1: mwscript-k8s -- extensions/TimedMediaHandler/maintenance/requeueTranscodes.php --wiki=commonswiki --key '360p.mpeg4.mov' --throttle --video --missing ([[phab:T358266|T358266]]) * 14:58 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1005.wikimedia.org * 14:58 jhancock@cumin2002: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['ms-be2098'] * 14:58 jhancock@cumin2002: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['ms-be2098'] * 14:58 jhancock@cumin2002: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['ms-be2097'] * 14:58 jhancock@cumin2002: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['ms-be2097'] * 14:58 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1078.eqiad.wmnet with reason: host reimage * 14:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1181 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95621 and previous config saved to /var/cache/conftool/dbconfig/20260729-145629-cwilliams.json * 14:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1181.eqiad.wmnet with reason: Maintenance * 14:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1174: Maintenance * 14:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1219: Maintenance * 14:55 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader1006.wikimedia.org on all recursors * 14:55 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader1006.wikimedia.org on all recursors * 14:55 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader1005.wikimedia.org on all recursors * 14:55 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader1005.wikimedia.org on all recursors * 14:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1200 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95618 and previous config saved to /var/cache/conftool/dbconfig/20260729-145336-cwilliams.json * 14:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1254: Maintenance * 14:53 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1200.eqiad.wmnet with reason: Maintenance * 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1185: Maintenance * 14:52 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2021 * 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2015 * 14:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2015 * 14:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1219 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95616 and previous config saved to /var/cache/conftool/dbconfig/20260729-144946-cwilliams.json * 14:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1219.eqiad.wmnet with reason: Maintenance * 14:49 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1218: Maintenance * 14:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2021.codfw.wmnet with OS bookworm * 14:48 dancy@deploy1003: Finished deploy [zuul/deploy@22703a6]: Deploying https://gerrit.wikimedia.org/r/c/integration/zuul/+/1311501 ([[phab:T432491|T432491]]) (duration: 00m 15s) * 14:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2015.codfw.wmnet with OS bookworm * 14:48 dancy@deploy1003: Started deploy [zuul/deploy@22703a6]: Deploying https://gerrit.wikimedia.org/r/c/integration/zuul/+/1311501 ([[phab:T432491|T432491]]) * 14:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1254 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95613 and previous config saved to /var/cache/conftool/dbconfig/20260729-144729-cwilliams.json * 14:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1254.eqiad.wmnet with reason: Maintenance * 14:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1233: Maintenance * 14:46 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2013\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 14:46 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2014\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 14:46 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:45 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:44 root@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1285.eqiad.wmnet with reason: host reimage * 14:43 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:42 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:41 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:40 root@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1285.eqiad.wmnet with reason: host reimage * 14:39 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1078.eqiad.wmnet with OS trixie * 14:39 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2074.codfw.wmnet with OS trixie * 14:32 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:32 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:32 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:31 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2005.codfw.wmnet with OS bookworm * 14:30 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:30 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:29 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:29 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:27 root@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host db1285 * 14:27 root@cumin1003: START - Cookbook sre.hosts.move-vlan for host db1285 * 14:27 root@cumin1003: START - Cookbook sre.hosts.reimage for host db1285.eqiad.wmnet with OS trixie * 14:24 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 14:24 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:24 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:24 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:23 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:22 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:22 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:21 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:17 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:16 root@cumin1003: START - Cookbook sre.mysql.pool pool db1244: Maintenance * 14:15 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:15 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add asw1-604 loopback ipv4 - pt1979@cumin2002" * 14:15 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add asw1-604 loopback ipv4 - pt1979@cumin2002" * 14:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:12 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 14:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95599 and previous config saved to /var/cache/conftool/dbconfig/20260729-141014-cwilliams.json * 14:10 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 14:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1244.eqiad.wmnet with reason: Maintenance * 14:10 pt1979@cumin2002: START - Cookbook sre.dns.netbox * 14:10 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 14:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1243: Maintenance * 14:09 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2005.codfw.wmnet with reason: host reimage * 14:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db1174: Maintenance * 14:08 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad * 14:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db1185: Maintenance * 14:06 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 14:05 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2005.codfw.wmnet with reason: host reimage * 14:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1174 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95595 and previous config saved to /var/cache/conftool/dbconfig/20260729-140309-cwilliams.json * 14:03 sukhe@dns1004: END - running authdns-update * 14:03 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1174.eqiad.wmnet with reason: Maintenance * 14:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1170: Maintenance * 14:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db1218: Maintenance * 14:01 sukhe@dns1004: START - running authdns-update * 14:00 sukhe@dns1004: START - running authdns-update * 13:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db1233: Maintenance * 13:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1185 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95592 and previous config saved to /var/cache/conftool/dbconfig/20260729-135925-cwilliams.json * 13:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1185.eqiad.wmnet with reason: Maintenance * 13:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1161: Maintenance * 13:58 sukhe@puppetserver1001: conftool action : set/pooled=true; selector: dnsdisc=urldownloader * 13:58 root@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1265.eqiad.wmnet with OS trixie * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1218 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95590 and previous config saved to /var/cache/conftool/dbconfig/20260729-135621-cwilliams.json * 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1218.eqiad.wmnet with reason: Maintenance * 13:55 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1206: Maintenance * 13:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2073.codfw.wmnet with OS trixie * 13:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1233 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95587 and previous config saved to /var/cache/conftool/dbconfig/20260729-135335-cwilliams.json * 13:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1233.eqiad.wmnet with reason: Maintenance * 13:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1229: Maintenance * 13:50 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/kartotherian: apply * 13:50 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service * 13:49 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:49 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/kartotherian: apply * 13:48 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 13:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1077.eqiad.wmnet with OS trixie * 13:47 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 13:46 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2005.codfw.wmnet with OS bookworm * 13:44 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 13:44 root@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1265.eqiad.wmnet with reason: host reimage * 13:40 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] (duration: 09m 22s) * 13:39 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:38 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:36 root@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1265.eqiad.wmnet with reason: host reimage * 13:35 stran@deploy1003: stran: Continuing with deployment * 13:33 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 13:32 stran@deploy1003: stran: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified t * 13:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2073.codfw.wmnet with reason: host reimage * 13:30 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ml-build1001.eqiad.wmnet * 13:30 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] * 13:29 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad * 13:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1077.eqiad.wmnet with reason: host reimage * 13:27 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 13:27 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:27 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad * 13:26 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2073.codfw.wmnet with reason: host reimage * 13:26 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] (duration: 07m 56s) * 13:25 klausman@cumin1003: START - Cookbook sre.hosts.reboot-single for host ml-build1001.eqiad.wmnet * 13:24 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 13:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>ml-serve101[2-5].eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 13:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1015.eqiad.wmnet * 13:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1015.eqiad.wmnet * 13:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1077.eqiad.wmnet with reason: host reimage * 13:23 root@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host db1265 * 13:23 root@cumin1003: START - Cookbook sre.hosts.move-vlan for host db1265 * 13:23 root@cumin1003: START - Cookbook sre.hosts.reimage for host db1265.eqiad.wmnet with OS trixie * 13:23 root@cumin1003: START - Cookbook sre.mysql.pool pool db1243: Maintenance * 13:22 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 13:22 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2005.codfw.wmnet * 13:22 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 13:22 samtar@deploy1003: dreamrimmer, samtar: Continuing with deployment * 13:20 samtar@deploy1003: dreamrimmer, samtar: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts an-test-master[1001-1002].eqiad.wmnet * 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-master[1001-1002].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 13:18 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1015.eqiad.wmnet * 13:18 sukhe@cumin1003: END (ERROR) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=97) for role: url_downloader@eqiad * 13:18 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 13:18 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] * 13:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95574 and previous config saved to /var/cache/conftool/dbconfig/20260729-131638-cwilliams.json * 13:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1243.eqiad.wmnet with reason: Maintenance * 13:16 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2005.codfw.wmnet * 13:16 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1242: Maintenance * 13:14 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] (duration: 07m 00s) * 13:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db1170: Maintenance * 13:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1015.eqiad.wmnet * 13:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1014.eqiad.wmnet * 13:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1014.eqiad.wmnet * 13:12 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1223: Maintenance * 13:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db1161: Maintenance * 13:10 samtar@deploy1003: anzx, samtar: Continuing with deployment * 13:09 samtar@deploy1003: anzx, samtar: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db1206: Maintenance * 13:08 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 13:07 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] * 13:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1170 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95566 and previous config saved to /var/cache/conftool/dbconfig/20260729-130730-cwilliams.json * 13:07 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1170.eqiad.wmnet with reason: Maintenance * 13:07 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:07 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt IPs new switches - cmooney@cumin1003" * 13:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1158: Maintenance * 13:06 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1014.eqiad.wmnet * 13:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1161 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95564 and previous config saved to /var/cache/conftool/dbconfig/20260729-130616-cwilliams.json * 13:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 13:06 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1077.eqiad.wmnet with OS trixie * 13:05 root@cumin1003: START - Cookbook sre.mysql.pool pool db1229: Maintenance * 13:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1161.eqiad.wmnet with reason: Maintenance * 13:05 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt IPs new switches - cmooney@cumin1003" * 13:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2073.codfw.wmnet with OS trixie * 13:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Maintenance * 13:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1206 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95562 and previous config saved to /var/cache/conftool/dbconfig/20260729-130258-cwilliams.json * 13:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1206.eqiad.wmnet with reason: Maintenance * 13:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1196: Maintenance * 13:01 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 13:01 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:00 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 13:00 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 12:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1229 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95559 and previous config saved to /var/cache/conftool/dbconfig/20260729-125950-cwilliams.json * 12:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1229.eqiad.wmnet with reason: Maintenance * 12:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1222: Maintenance * 12:57 sukhe: sudo cumin 'A:lvs and (A:eqiad or A:codfw)' 'disable-puppet "adding new service urldownloader"': [[phab:T429175|T429175]] * 12:56 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1014.eqiad.wmnet * 12:56 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1013.eqiad.wmnet * 12:56 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1013.eqiad.wmnet * 12:50 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1013.eqiad.wmnet * 12:50 sukhe: sudo cumin 'O:url_downloader' 'run-puppet-agent --enable "merging CR 1313948"': [[phab:T429175|T429175]] * 12:48 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-master[1001-1002].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 12:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1013.eqiad.wmnet * 12:45 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1012.eqiad.wmnet * 12:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1012.eqiad.wmnet * 12:45 sukhe: sudo cumin 'O:url_downloader' 'disable-puppet "merging CR 1313948"': [[phab:T429175|T429175]] * 12:44 btullis@cumin1003: START - Cookbook sre.dns.netbox * 12:40 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test2001.codfw.wmnet * 12:40 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test2001.codfw.wmnet * 12:38 ayounsi@dns1004: END - running authdns-update * 12:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1012.eqiad.wmnet * 12:35 ayounsi@dns1004: START - running authdns-update * 12:34 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts an-test-master[1001-1002].eqiad.wmnet * 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts an-test-coord1001.eqiad.wmnet * 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-coord1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 12:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1012.eqiad.wmnet * 12:32 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>ml-serve101[2-5].eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 12:29 root@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Maintenance * 12:25 root@cumin1003: START - Cookbook sre.mysql.pool pool db1223: Maintenance * 12:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95544 and previous config saved to /var/cache/conftool/dbconfig/20260729-122254-cwilliams.json * 12:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1242.eqiad.wmnet with reason: Maintenance * 12:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1241: Maintenance * 12:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1051 hosts * 12:20 root@cumin1003: START - Cookbook sre.mysql.pool pool db1158: Maintenance * 12:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1223 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95540 and previous config saved to /var/cache/conftool/dbconfig/20260729-121937-cwilliams.json * 12:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1223.eqiad.wmnet with reason: Maintenance * 12:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1212: Maintenance * 12:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db1159: Maintenance * 12:17 elukey@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: sync * 12:15 elukey@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: sync * 12:15 root@cumin1003: START - Cookbook sre.mysql.pool pool db1196: Maintenance * 12:14 Daimona: Creating new DB tables for the CampaignEvents extension in x1.testwiki, x1.test2wiki, x1.officewiki, and x1.wikishared # [[phab:T429339|T429339]] * 12:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db1222: Maintenance * 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95535 and previous config saved to /var/cache/conftool/dbconfig/20260729-121211-cwilliams.json * 12:12 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 12:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1158.eqiad.wmnet with reason: Maintenance * 12:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1159 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95534 and previous config saved to /var/cache/conftool/dbconfig/20260729-121146-cwilliams.json * 12:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1159.eqiad.wmnet with reason: Maintenance * 12:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1196 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95533 and previous config saved to /var/cache/conftool/dbconfig/20260729-120847-cwilliams.json * 12:08 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 12:08 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1196.eqiad.wmnet with reason: Maintenance * 12:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1195: Maintenance * 12:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1222 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95530 and previous config saved to /var/cache/conftool/dbconfig/20260729-120424-cwilliams.json * 12:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1222.eqiad.wmnet with reason: Maintenance * 12:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1098 hosts * 12:00 marostegui: Rename tables [[phab:T425074|T425074]] * 12:00 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-coord1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 11:58 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1197: Maintenance * 11:55 btullis@cumin1003: START - Cookbook sre.dns.netbox * 11:52 elukey@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: sync * 11:51 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:51 elukey@deploy1003: helmfile [codfw] START helmfile.d/services/proton: sync * 11:51 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:50 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts an-test-coord1001.eqiad.wmnet * 11:50 elukey@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: sync * 11:49 elukey@deploy1003: helmfile [staging] START helmfile.d/services/proton: sync * 11:35 root@cumin1003: START - Cookbook sre.mysql.pool pool db1241: Maintenance * 11:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db1212: Maintenance * 11:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1241 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95520 and previous config saved to /var/cache/conftool/dbconfig/20260729-112918-cwilliams.json * 11:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1241.eqiad.wmnet with reason: Maintenance * 11:29 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1238: Maintenance * 11:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1212 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95517 and previous config saved to /var/cache/conftool/dbconfig/20260729-112727-cwilliams.json * 11:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 11:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1212.eqiad.wmnet with reason: Maintenance * 11:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1198: Maintenance * 11:23 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:22 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:21 root@cumin1003: START - Cookbook sre.mysql.pool pool db1195: Maintenance * 11:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1195 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95514 and previous config saved to /var/cache/conftool/dbconfig/20260729-111450-cwilliams.json * 11:14 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1195.eqiad.wmnet with reason: Maintenance * 11:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1186: Maintenance * 11:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 11:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 11:05 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:54 marostegui: Dropping renamed tables [[phab:T425066|T425066]] * 10:41 root@cumin1003: START - Cookbook sre.mysql.pool pool db1238: Maintenance * 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1198: Maintenance * 10:39 Amir1: ran https://phabricator.wikimedia.org/T432509#12149723 in production ([[phab:T432509|T432509]]) * 10:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db1197: Maintenance * 10:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1238 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95501 and previous config saved to /var/cache/conftool/dbconfig/20260729-103532-cwilliams.json * 10:35 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1238.eqiad.wmnet with reason: Maintenance * 10:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1221: Maintenance * 10:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1198 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95499 and previous config saved to /var/cache/conftool/dbconfig/20260729-103330-cwilliams.json * 10:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1198.eqiad.wmnet with reason: Maintenance * 10:33 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1175: Maintenance * 10:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1197 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95496 and previous config saved to /var/cache/conftool/dbconfig/20260729-103217-cwilliams.json * 10:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1197.eqiad.wmnet with reason: Maintenance * 10:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1188: Maintenance * 10:27 root@cumin1003: START - Cookbook sre.mysql.pool pool db1186: Maintenance * 10:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1186 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95493 and previous config saved to /var/cache/conftool/dbconfig/20260729-102111-cwilliams.json * 10:21 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1186.eqiad.wmnet with reason: Maintenance * 10:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 10:14 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 09:53 XioNoX: reboot cr2-magru - [[phab:T431750|T431750]] * 09:52 XioNoX: drain cr2-magru - [[phab:T431750|T431750]] * 09:48 root@cumin1003: START - Cookbook sre.mysql.pool pool db1221: Maintenance * 09:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zookeeper-test1002.eqiad.wmnet * 09:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1188: Maintenance * 09:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1175: Maintenance * 09:44 btullis@dns1004: END - running authdns-update * 09:42 btullis@dns1004: START - running authdns-update * 09:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1221 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95483 and previous config saved to /var/cache/conftool/dbconfig/20260729-094200-cwilliams.json * 09:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 7 hosts with reason: Maintenance * 09:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1221.eqiad.wmnet with reason: Maintenance * 09:41 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host zookeeper-test1002.eqiad.wmnet * 09:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1199: Maintenance * 09:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1188 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95481 and previous config saved to /var/cache/conftool/dbconfig/20260729-093917-cwilliams.json * 09:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1188.eqiad.wmnet with reason: Maintenance * 09:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1182: Maintenance * 09:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1175 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95479 and previous config saved to /var/cache/conftool/dbconfig/20260729-093842-cwilliams.json * 09:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1175.eqiad.wmnet with reason: Maintenance * 09:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1166: Maintenance * 09:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1169: Maintenance * 09:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1033.eqiad.wmnet,service=s8 * 09:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1033.eqiad.wmnet,service=s5 * 09:33 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1033.eqiad.wmnet,service=s8 * 09:33 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1033.eqiad.wmnet,service=s5 * 09:21 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:21 XioNoX: reboot cr1-magru - [[phab:T431750|T431750]] * 09:17 XioNoX: drain cr1-magru - [[phab:T431750|T431750]] * 09:15 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm1001.wikimedia.org * 09:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply * 09:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply * 09:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 09:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 09:11 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr2-magru,cr2-magru IPv6,cr2-magru.mgmt with reason: router upgrade * 09:11 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm1001.wikimedia.org * 09:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp1005.wikimedia.org * 09:07 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp1005.wikimedia.org * 09:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp2005.wikimedia.org * 09:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp2005.wikimedia.org * 09:00 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr1-magru,cr1-magru IPv6,cr1-magru.mgmt with reason: router upgrade * 09:00 marostegui: Dropping renamed tables [[phab:T426341|T426341]] * 08:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1199: Maintenance * 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1182: Maintenance * 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1166: Maintenance * 08:47 root@cumin1003: START - Cookbook sre.mysql.pool pool db1169: Maintenance * 08:46 ayounsi@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 1:00:00 on cr1-magru,cr1-magru IPv6,cr1-magru.mgmt with reason: router upgrade * 08:45 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2072.codfw.wmnet with OS trixie * 08:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1199 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95464 and previous config saved to /var/cache/conftool/dbconfig/20260729-084534-cwilliams.json * 08:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1199.eqiad.wmnet with reason: Maintenance * 08:45 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 08:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1190: Maintenance * 08:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1182 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95462 and previous config saved to /var/cache/conftool/dbconfig/20260729-084436-cwilliams.json * 08:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1182.eqiad.wmnet with reason: Maintenance * 08:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1156: Maintenance * 08:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1166 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95460 and previous config saved to /var/cache/conftool/dbconfig/20260729-084400-cwilliams.json * 08:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1166.eqiad.wmnet with reason: Maintenance * 08:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1157: Maintenance * 08:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95458 and previous config saved to /var/cache/conftool/dbconfig/20260729-084147-cwilliams.json * 08:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1169.eqiad.wmnet with reason: Maintenance * 08:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1163: Maintenance * 08:30 btullis@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 11 hosts with reason: Replacing the namenodes * 08:23 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2072.codfw.wmnet with reason: host reimage * 08:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1098 hosts * 08:19 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2072.codfw.wmnet with reason: host reimage * 07:58 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2072.codfw.wmnet with OS trixie * 07:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1190: Maintenance * 07:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1156: Maintenance * 07:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1157: Maintenance * 07:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1163: Maintenance * 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1190 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95444 and previous config saved to /var/cache/conftool/dbconfig/20260729-074930-cwilliams.json * 07:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1190.eqiad.wmnet with reason: Maintenance * 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1157 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95443 and previous config saved to /var/cache/conftool/dbconfig/20260729-074914-cwilliams.json * 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1156 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95442 and previous config saved to /var/cache/conftool/dbconfig/20260729-074906-cwilliams.json * 07:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1157.eqiad.wmnet with reason: Maintenance * 07:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 07:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1156.eqiad.wmnet with reason: Maintenance * 07:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1163 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95441 and previous config saved to /var/cache/conftool/dbconfig/20260729-074652-cwilliams.json * 07:46 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1163.eqiad.wmnet with reason: Maintenance * 07:46 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2034.codfw.wmnet * 07:42 ayounsi@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2034.codfw.wmnet * 07:42 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2232.codfw.wmnet with OS trixie * 07:34 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:34 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:33 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:31 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:19 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2232.codfw.wmnet with reason: host reimage * 07:15 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2232.codfw.wmnet with reason: host reimage * 06:58 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db2232.codfw.wmnet with OS trixie * 06:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[2160,2232].codfw.wmnet with reason: Reimage * 06:26 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1164.eqiad.wmnet with OS trixie * 06:05 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1164.eqiad.wmnet with reason: host reimage * 06:01 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1164.eqiad.wmnet with reason: host reimage * 05:47 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1164.eqiad.wmnet with OS trixie * 05:46 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1164.eqiad.wmnet with reason: Reimage == 2026-07-28 == * 22:50 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1138.eqiad.wmnet * 22:50 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1138.eqiad.wmnet * 22:49 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1138.eqiad.wmnet * 22:11 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2014.codfw.wmnet, repooling source-only afterwards * 22:08 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2013.codfw.wmnet, repooling source-only afterwards * 22:03 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 20:58 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] (duration: 08m 19s) * 20:55 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2014.codfw.wmnet, repooling source-only afterwards * 20:55 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2013.codfw.wmnet, repooling source-only afterwards * 20:54 arlolra@deploy1003: arlolra: Continuing with deployment * 20:54 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 14s) * 20:54 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 20:53 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 30s) * 20:53 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 20:52 arlolra@deploy1003: arlolra: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:51 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:50 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] * 20:49 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:43 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 20:34 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] (duration: 06m 54s) * 20:34 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:34 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:31 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:31 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:30 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:30 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:30 arlolra@deploy1003: arlolra: Continuing with deployment * 20:29 arlolra@deploy1003: arlolra: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:27 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] * 20:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2014.codfw.wmnet with OS bookworm * 20:21 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 20:21 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:20 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 20:19 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:19 swfrench-wmf: switched etcd-mirror replication from conf2005 to conf2004 - [[phab:T428495|T428495]] * 20:17 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:17 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:15 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] (duration: 08m 26s) * 20:12 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:11 arlolra@deploy1003: anzx, arlolra: Continuing with deployment * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2013.codfw.wmnet with OS bookworm * 20:09 arlolra@deploy1003: anzx, arlolra: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] * 20:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2216: Maintenance * 19:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2014.codfw.wmnet with reason: host reimage * 19:57 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:54 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:54 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:53 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:52 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2014.codfw.wmnet with reason: host reimage * 19:49 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2013.codfw.wmnet with reason: host reimage * 19:42 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 19:41 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:41 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2013.codfw.wmnet with reason: host reimage * 19:39 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:39 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:39 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-eqiad: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 19:38 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2014 * 19:33 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2014 * 19:29 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2014.codfw.wmnet with OS bookworm * 19:28 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:27 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1006 * 19:26 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2012\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 19:26 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1006 * 19:26 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:26 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 19:25 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2013 * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2013 * 19:21 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2013 * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2013.codfw.wmnet 84.0.192.10.in-addr.arpa 4.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:21 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2013.codfw.wmnet 84.0.192.10.in-addr.arpa 4.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2013 - bking@cumin2003" * 19:21 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2013 - bking@cumin2003" * 19:21 vriley@cumin1003: START - Cookbook sre.dns.netbox * 19:20 root@cumin1003: START - Cookbook sre.mysql.pool pool db2216: Maintenance * 19:13 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2216 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95435 and previous config saved to /var/cache/conftool/dbconfig/20260728-191343-cwilliams.json * 19:13 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2216.codfw.wmnet with reason: Maintenance * 19:13 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2203: Maintenance * 19:06 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1005.eqiad.wmnet with OS trixie * 19:06 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 19:06 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 18:46 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 18:45 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:45 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:43 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:40 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:36 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-eqiad: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 18:35 dancy@deploy1003: Installation of scap version "4.275.0" completed for 3 hosts * 18:33 dancy@deploy1003: Installing scap version "4.275.0" for 3 host(s) * 18:32 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:32 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2097.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:30 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns3003.wikimedia.org [reason: pool for all services after reimaging] * 18:29 sukhe@dns1004: END - running authdns-update * 18:27 sukhe@dns1004: START - running authdns-update * 18:27 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns3003.wikimedia.org,service=authdns-update [reason: pool authdns-update after reimaging] * 18:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db2203: Maintenance * 18:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2203 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95430 and previous config saved to /var/cache/conftool/dbconfig/20260728-181958-cwilliams.json * 18:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2203.codfw.wmnet with reason: Maintenance * 18:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2188: Maintenance * 18:18 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2097.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:17 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-be2098 * 18:17 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host ms-be2098 * 18:17 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-be2097 * 18:16 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host ms-be2097 * 18:15 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:15 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding ms-be2097-8 to codfw - jhancock@cumin2002" * 18:15 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding ms-be2097-8 to codfw - jhancock@cumin2002" * 18:10 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 18:08 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage * 18:05 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns3003.wikimedia.org with OS trixie * 18:03 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage * 17:56 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-codfw: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 17:45 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie * 17:45 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1005.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:41 sukhe@dns1004: END - running authdns-update * 17:39 sukhe@dns1004: START - running authdns-update * 17:36 sukhe@puppetserver1001: conftool action : set/weight=1; selector: cluster=urldownloader,service=squid * 17:36 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1005.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:35 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader,service=squid * 17:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 17:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1005 * 17:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 17:34 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1005 * 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1005] - vriley@cumin1003" * 17:34 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1005] - vriley@cumin1003" * 17:32 root@cumin1003: START - Cookbook sre.mysql.pool pool db2188: Maintenance * 17:29 vriley@cumin1003: START - Cookbook sre.dns.netbox * 17:29 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 17:26 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2188 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95425 and previous config saved to /var/cache/conftool/dbconfig/20260728-172609-cwilliams.json * 17:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2188.codfw.wmnet with reason: Maintenance * 17:25 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2176: Maintenance * 17:19 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1138.eqiad.wmnet with OS trixie * 17:18 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1005 * 17:18 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1005 * 17:18 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:15 vriley@cumin1003: START - Cookbook sre.dns.netbox * 17:13 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns3003.wikimedia.org with reason: host reimage * 17:07 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns3003.wikimedia.org with reason: host reimage * 17:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-eqiad * 17:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1015.eqiad.wmnet * 17:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1015.eqiad.wmnet * 16:59 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1138.eqiad.wmnet with reason: host reimage * 16:55 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-codfw: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 16:54 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1138.eqiad.wmnet with reason: host reimage * 16:53 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1015.eqiad.wmnet * 16:43 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns3003.wikimedia.org with OS trixie * 16:43 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1015.eqiad.wmnet * 16:43 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1014.eqiad.wmnet * 16:43 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1014.eqiad.wmnet * 16:43 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=dns3003.wikimedia.org [reason: depooling for reimage to trixie] * 16:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 290 hosts * 16:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db2176: Maintenance * 16:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1138 * 16:38 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1138 * 16:37 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1138 * 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1138.eqiad.wmnet 193.32.64.10.in-addr.arpa 3.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:37 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1138.eqiad.wmnet 193.32.64.10.in-addr.arpa 3.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1138 - jiji@cumin1003" * 16:37 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1138 - jiji@cumin1003" * 16:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1014.eqiad.wmnet * 16:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2176 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95420 and previous config saved to /var/cache/conftool/dbconfig/20260728-163235-cwilliams.json * 16:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2176.codfw.wmnet with reason: Maintenance * 16:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1014.eqiad.wmnet * 16:32 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1013.eqiad.wmnet * 16:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1013.eqiad.wmnet * 16:32 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2174: Maintenance * 16:28 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2012.codfw.wmnet, repooling source-only afterwards * 16:25 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1013.eqiad.wmnet * 16:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1013.eqiad.wmnet * 16:20 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1012.eqiad.wmnet * 16:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1012.eqiad.wmnet * 16:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1012.eqiad.wmnet * 16:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1012.eqiad.wmnet * 16:03 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1011.eqiad.wmnet * 16:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1011.eqiad.wmnet * 16:00 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2004.codfw.wmnet with OS bookworm * 15:59 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1011.eqiad.wmnet * 15:56 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:55 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 15:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:54 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1011.eqiad.wmnet * 15:54 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1010.eqiad.wmnet * 15:54 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1010.eqiad.wmnet * 15:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:50 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 15:49 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1010.eqiad.wmnet * 15:48 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 15:48 jiji@cumin1003: START - Cookbook sre.dns.netbox * 15:46 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2248: Maintenance * 15:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db2174: Maintenance * 15:44 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1010.eqiad.wmnet * 15:44 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1009.eqiad.wmnet * 15:44 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1009.eqiad.wmnet * 15:42 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1138 * 15:41 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1138.eqiad.wmnet with OS trixie * 15:39 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1009.eqiad.wmnet * 15:39 robh@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on arclamp2001.codfw.wmnet with reason: ram upgrade * 15:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2174 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95413 and previous config saved to /var/cache/conftool/dbconfig/20260728-153844-cwilliams.json * 15:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2174.codfw.wmnet with reason: Maintenance * 15:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2173: Maintenance * 15:37 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 15:35 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 15:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1009.eqiad.wmnet * 15:34 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1008.eqiad.wmnet * 15:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1008.eqiad.wmnet * 15:31 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 15:31 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 15:29 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1008.eqiad.wmnet * 15:27 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 290 hosts * 15:25 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2013 * 15:25 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2195: Maintenance * 15:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1008.eqiad.wmnet * 15:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1007.eqiad.wmnet * 15:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1007.eqiad.wmnet * 15:22 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2004.codfw.wmnet with reason: host reimage * 15:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2013.codfw.wmnet with OS bookworm * 15:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 15:19 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2004.codfw.wmnet with reason: host reimage * 15:19 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 15:17 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1007.eqiad.wmnet * 15:12 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1007.eqiad.wmnet * 15:12 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1006.eqiad.wmnet * 15:12 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1006.eqiad.wmnet * 15:11 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1138.eqiad.wmnet * 15:11 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 15:11 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1138.eqiad.wmnet * 15:11 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1138.eqiad.wmnet * 15:10 brennen@deploy1003: Finished deploy [phabricator/deployment@f8b349f]: deploy phab1004 for [[phab:T433382|T433382]] (duration: 00m 43s) * 15:10 brennen@deploy1003: Started deploy [phabricator/deployment@f8b349f]: deploy phab1004 for [[phab:T433382|T433382]] * 15:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts * 15:09 brennen@deploy1003: Finished deploy [phabricator/deployment@f8b349f]: deploy phab2003 for [[phab:T433382|T433382]] (duration: 00m 55s) * 15:08 brennen@deploy1003: Started deploy [phabricator/deployment@f8b349f]: deploy phab2003 for [[phab:T433382|T433382]] * 15:07 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2012.codfw.wmnet, repooling source-only afterwards * 15:07 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts * 15:06 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1137.eqiad.wmnet * 15:06 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1137.eqiad.wmnet * 15:06 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1137.eqiad.wmnet * 15:05 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1006.eqiad.wmnet * 15:05 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1022\.eqiad\.wmnet,dc=eqiad,cluster=wdqs\-main,service=wdqs\-main * 15:01 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab2003.codfw.wmnet with reason: deployment * 15:01 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1005.eqiad.wmnet with reason: deployment * 15:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1006.eqiad.wmnet * 15:00 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1006.eqiad.wmnet with reason: deployment * 15:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1005.eqiad.wmnet * 15:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1005.eqiad.wmnet * 14:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db2248: Maintenance * 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1004.eqiad.wmnet with reason: deployment * 14:59 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2004.codfw.wmnet with OS bookworm * 14:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1005.eqiad.wmnet * 14:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2248 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95403 and previous config saved to /var/cache/conftool/dbconfig/20260728-145532-cwilliams.json * 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2245-2247].codfw.wmnet with reason: Maintenance * 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2248.codfw.wmnet with reason: Maintenance * 14:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2240: Maintenance * 14:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2173: Maintenance * 14:51 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts * 14:50 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1005.eqiad.wmnet * 14:50 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1004.eqiad.wmnet * 14:50 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1004.eqiad.wmnet * 14:49 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts * 14:45 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2004.codfw.wmnet * 14:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2173 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95399 and previous config saved to /var/cache/conftool/dbconfig/20260728-144453-cwilliams.json * 14:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2173.codfw.wmnet with reason: Maintenance * 14:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2170: Maintenance * 14:44 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1004.eqiad.wmnet * 14:39 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2004.codfw.wmnet * 14:38 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1004.eqiad.wmnet * 14:38 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet * 14:38 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet * 14:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db2195: Maintenance * 14:36 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1022.eqiad.wmnet, repooling source-only afterwards * 14:36 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 14:33 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet * 14:33 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2222: Maintenance * 14:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2195 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95394 and previous config saved to /var/cache/conftool/dbconfig/20260728-143218-cwilliams.json * 14:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2195.codfw.wmnet with reason: Maintenance * 14:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2181: Maintenance * 14:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 14:30 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 14:25 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 14:25 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 14:25 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 14:23 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet * 14:23 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1002.eqiad.wmnet * 14:23 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1002.eqiad.wmnet * 14:23 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 19s) * 14:23 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:18 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1002.eqiad.wmnet * 14:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2012.codfw.wmnet with OS bookworm * 14:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1002.eqiad.wmnet * 14:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1001.eqiad.wmnet * 14:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1001.eqiad.wmnet * 14:11 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 14:11 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 14:08 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1001.eqiad.wmnet * 14:07 elukey@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'. * 14:07 elukey@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'. * 14:06 elukey@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'. * 14:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db2240: Maintenance * 14:06 XioNoX: un-drain cr2-esams - [[phab:T431751|T431751]] * 14:05 elukey@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'. * 14:02 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1001.eqiad.wmnet * 14:02 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-eqiad * 14:01 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 14:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2240 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95384 and previous config saved to /var/cache/conftool/dbconfig/20260728-140011-cwilliams.json * 14:00 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2240.codfw.wmnet with reason: Maintenance * 13:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2237: Maintenance * 13:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db2170: Maintenance * 13:56 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 13:55 XioNoX: reboot cr2-esams - [[phab:T431751|T431751]] * 13:52 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 13:51 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr2-esams,cr2-esams IPv6,cr2-esams.mgmt with reason: router upgrade * 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 13:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2170 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95381 and previous config saved to /var/cache/conftool/dbconfig/20260728-135043-cwilliams.json * 13:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2170.codfw.wmnet with reason: Maintenance * 13:50 XioNoX: drain cr2-esams - [[phab:T431751|T431751]] * 13:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2153: Maintenance * 13:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2012.codfw.wmnet with reason: host reimage * 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2222: Maintenance * 13:45 sukhe: restart pybal on A:lvs-codfw * 13:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2012.codfw.wmnet with reason: host reimage * 13:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db2181: Maintenance * 13:44 btullis@dns1004: END - running authdns-update * 13:42 sukhe: restart pybal on lvs2014 * 13:42 btullis@dns1004: START - running authdns-update * 13:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2222 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95376 and previous config saved to /var/cache/conftool/dbconfig/20260728-133948-cwilliams.json * 13:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2222.codfw.wmnet with reason: Maintenance * 13:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2221: Maintenance * 13:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2181 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95374 and previous config saved to /var/cache/conftool/dbconfig/20260728-133857-cwilliams.json * 13:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2181.codfw.wmnet with reason: Maintenance * 13:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2167: Maintenance * 13:30 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T428495|T428495]] * 13:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 13:29 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 13:29 ayounsi@cumin1003: END (FAIL) - Cookbook sre.dns.admin (exit_code=99) DNS admin: depool esams [reason: router upgrade, [[phab:T431749|T431749]]] * 13:28 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: router upgrade, [[phab:T431749|T431749]]] * 13:27 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1022.eqiad.wmnet, repooling source-only afterwards * 13:27 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo - [[phab:T428495|T428495]] * 13:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2012 * 13:27 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2012 * 13:21 lucaswerkmeister-wmde@deploy1003: mwscript-k8s job started: cleanupTitles bolwiki # [[phab:T429951|T429951]] * 13:21 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] (duration: 07m 19s) * 13:20 swfrench-wmf: authdns-update to direct codfw, eqsin, ulsfo etcd clients to eqiad - [[phab:T428495|T428495]] * 13:18 swfrench@dns1004: END - running authdns-update * 13:17 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, anzx: Continuing with deployment * 13:16 swfrench@dns1004: START - running authdns-update * 13:16 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2012 * 13:16 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2012.codfw.wmnet 57.48.192.10.in-addr.arpa 7.5.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:16 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2012.codfw.wmnet 57.48.192.10.in-addr.arpa 7.5.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:16 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:16 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, anzx: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:14 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 13:14 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] * 13:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db2237: Maintenance * 13:13 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2237: Maintenance * 13:13 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:12 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:12 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback IPV6 for asw1-604 - pt1979@cumin2003" * 13:12 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:12 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback IPV6 for asw1-604 - pt1979@cumin2003" * 13:11 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 20s) * 13:11 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 13:10 esanders@deploy1003: Finished scap sync-world: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] (duration: 08m 11s) * 13:08 pt1979@cumin2003: START - Cookbook sre.dns.netbox * 13:07 root@cumin1003: START - Cookbook sre.mysql.pool pool db2237: Maintenance * 13:06 esanders@deploy1003: esanders: Continuing with deployment * 13:04 esanders@deploy1003: esanders: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2153: Maintenance * 13:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2153: Maintenance * 13:02 esanders@deploy1003: Started scap sync-world: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] * 13:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2237 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95362 and previous config saved to /var/cache/conftool/dbconfig/20260728-130107-cwilliams.json * 13:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2237.codfw.wmnet with reason: Maintenance * 13:00 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2236: Maintenance * 12:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2153: Maintenance * 12:52 root@cumin1003: START - Cookbook sre.mysql.pool pool db2221: Maintenance * 12:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2153 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95358 and previous config saved to /var/cache/conftool/dbconfig/20260728-125214-cwilliams.json * 12:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2153.codfw.wmnet with reason: Maintenance * 12:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2167: Maintenance * 12:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-codfw * 12:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2011.codfw.wmnet * 12:51 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2011.codfw.wmnet * 12:49 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:49 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback for asw1-603 - pt1979@cumin2003" * 12:48 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback for asw1-603 - pt1979@cumin2003" * 12:46 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2011.codfw.wmnet * 12:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2221 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95357 and previous config saved to /var/cache/conftool/dbconfig/20260728-124601-cwilliams.json * 12:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2221.codfw.wmnet with reason: Maintenance * 12:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2218: Maintenance * 12:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2167 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95354 and previous config saved to /var/cache/conftool/dbconfig/20260728-124457-cwilliams.json * 12:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2167.codfw.wmnet with reason: Maintenance * 12:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2166: Maintenance * 12:42 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 12:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2011.codfw.wmnet * 12:41 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2010.codfw.wmnet * 12:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2010.codfw.wmnet * 12:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2010.codfw.wmnet * 12:34 pt1979@cumin2003: START - Cookbook sre.dns.netbox * 12:32 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 12:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2010.codfw.wmnet * 12:31 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2009.codfw.wmnet * 12:31 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2009.codfw.wmnet * 12:27 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2009.codfw.wmnet * 12:22 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2009.codfw.wmnet * 12:22 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2008.codfw.wmnet * 12:21 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2008.codfw.wmnet * 12:16 pt1979@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-604-eqsin * 12:16 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2008.codfw.wmnet * 12:16 pt1979@cumin1003: START - Cookbook sre.network.tls for network device asw1-604-eqsin * 12:14 root@cumin1003: START - Cookbook sre.mysql.pool pool db2236: Maintenance * 12:14 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2236: Maintenance * 12:12 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1137.eqiad.wmnet with OS trixie * 12:11 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2008.codfw.wmnet * 12:11 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2007.codfw.wmnet * 12:11 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2007.codfw.wmnet * 12:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db2236: Maintenance * 12:06 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2007.codfw.wmnet * 12:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2236 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95348 and previous config saved to /var/cache/conftool/dbconfig/20260728-120253-cwilliams.json * 12:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2236.codfw.wmnet with reason: Maintenance * 12:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2007.codfw.wmnet * 12:01 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 12:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 11:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2218: Maintenance * 11:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db2166: Maintenance * 11:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2219: Maintenance * 11:56 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2006.codfw.wmnet * 11:52 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1137.eqiad.wmnet with reason: host reimage * 11:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2218 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95344 and previous config saved to /var/cache/conftool/dbconfig/20260728-115155-cwilliams.json * 11:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2218.codfw.wmnet with reason: Maintenance * 11:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2208: Maintenance * 11:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2166 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95342 and previous config saved to /var/cache/conftool/dbconfig/20260728-115119-cwilliams.json * 11:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2166.codfw.wmnet with reason: Maintenance * 11:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2164: Maintenance * 11:47 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1137.eqiad.wmnet with reason: host reimage * 11:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2006.codfw.wmnet * 11:45 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2005.codfw.wmnet * 11:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2005.codfw.wmnet * 11:40 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2005.codfw.wmnet * 11:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2005.codfw.wmnet * 11:35 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2004.codfw.wmnet * 11:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2004.codfw.wmnet * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1137 * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1137 * 11:30 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1137 * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1137.eqiad.wmnet 192.32.64.10.in-addr.arpa 2.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:30 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1137.eqiad.wmnet 192.32.64.10.in-addr.arpa 2.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1137 - jiji@cumin1003" * 11:25 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2004.codfw.wmnet * 11:19 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2004.codfw.wmnet * 11:19 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2003.codfw.wmnet * 11:19 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2003.codfw.wmnet * 11:14 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2003.codfw.wmnet * 11:11 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2219: Maintenance * 11:10 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2219: Maintenance * 11:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2219: Maintenance * 11:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2003.codfw.wmnet * 11:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2164: Maintenance * 11:03 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2002.codfw.wmnet * 11:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2002.codfw.wmnet * 11:03 root@cumin1003: START - Cookbook sre.mysql.pool pool db2208: Maintenance * 10:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2164 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95332 and previous config saved to /var/cache/conftool/dbconfig/20260728-105749-cwilliams.json * 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2164.codfw.wmnet with reason: Maintenance * 10:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2208 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95331 and previous config saved to /var/cache/conftool/dbconfig/20260728-105711-cwilliams.json * 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2208.codfw.wmnet with reason: Maintenance * 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2219 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95330 and previous config saved to /var/cache/conftool/dbconfig/20260728-105652-cwilliams.json * 10:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2219.codfw.wmnet with reason: Maintenance * 10:53 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1137 - jiji@cumin1003" * 10:52 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2002.codfw.wmnet * 10:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2002.codfw.wmnet * 10:47 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2001.codfw.wmnet * 10:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2001.codfw.wmnet * 10:39 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2001.codfw.wmnet * 10:35 jiji@cumin1003: START - Cookbook sre.dns.netbox * 10:34 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] (duration: 09m 31s) * 10:34 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1137 * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2001.codfw.wmnet * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-codfw * 10:34 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1137.eqiad.wmnet with OS trixie * 10:28 jforrester@deploy1003: jforrester: Continuing with deployment * 10:27 jforrester@deploy1003: jforrester: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:25 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] * 10:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 10:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 10:21 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 10:20 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1137.eqiad.wmnet * 10:20 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 10:20 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1137.eqiad.wmnet * 10:20 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1137.eqiad.wmnet * 10:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2163: Maintenance * 09:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-staging-worker * 09:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2003.codfw.wmnet * 09:37 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2003.codfw.wmnet * 09:32 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 09:31 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2003.codfw.wmnet * 09:30 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 09:30 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2163: Maintenance * 09:30 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 09:30 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 09:30 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:22 klausman@cumin1003: END (ERROR) - Cookbook sre.ganeti.reboot-vm (exit_code=97) for VM ml-serve-ctrl2001.codfw.wmnet * 09:22 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2001.codfw.wmnet * 09:22 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d8-eqiad * 09:22 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d8-eqiad * 09:21 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2003.codfw.wmnet * 09:20 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2002.codfw.wmnet * 09:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2002.codfw.wmnet * 09:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f2-codfw * 09:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f2-codfw * 09:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e4-codfw * 09:18 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2163: Maintenance * 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e4-codfw * 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-codfw * 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-codfw * 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e5-codfw * 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e5-codfw * 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f4-codfw * 09:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2210: Maintenance * 09:16 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f4-codfw * 09:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2182: Maintenance * 09:14 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2002.codfw.wmnet * 09:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db2163: Maintenance * 09:11 XioNoX: rebooting cr2-drmrs - [[phab:T431749|T431749]] * 09:10 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr2-drmrs,cr2-drmrs IPv6,cr2-drmrs.mgmt with reason: router upgrade * 09:06 XioNoX: draining cr2-drmrs - [[phab:T431749|T431749]] * 09:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2163 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95320 and previous config saved to /var/cache/conftool/dbconfig/20260728-090638-cwilliams.json * 09:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2163.codfw.wmnet with reason: Maintenance * 09:06 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2161: Maintenance * 09:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2002.codfw.wmnet * 09:04 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2001.codfw.wmnet * 09:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2001.codfw.wmnet * 08:57 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2001.codfw.wmnet * 08:48 XioNoX: un-drain cr1-drmrs - [[phab:T431749|T431749]] * 08:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2001.codfw.wmnet * 08:47 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-staging-worker * 08:42 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:35 XioNoX: rebooting cr1-drmrs - [[phab:T431749|T431749]] * 08:33 XioNoX: draining cr1-drmrs - [[phab:T431749|T431749]] * 08:31 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2210: Maintenance * 08:29 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2182: Maintenance * 08:21 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2182: Maintenance * 08:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db2161: Maintenance * 08:16 root@cumin1003: START - Cookbook sre.mysql.pool pool db2182: Maintenance * 08:12 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2210: Maintenance * 08:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2161 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95309 and previous config saved to /var/cache/conftool/dbconfig/20260728-081044-cwilliams.json * 08:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2161.codfw.wmnet with reason: Maintenance * 08:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2154: Maintenance * 08:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2182 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95307 and previous config saved to /var/cache/conftool/dbconfig/20260728-080947-cwilliams.json * 08:09 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2182.codfw.wmnet with reason: Maintenance * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2168: Maintenance * 08:06 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr1-drmrs,cr1-drmrs IPv6,cr1-drmrs.mgmt with reason: router upgrade * 08:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db2210: Maintenance * 08:05 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 08:05 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 08:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2210 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95305 and previous config saved to /var/cache/conftool/dbconfig/20260728-080008-cwilliams.json * 08:00 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2210.codfw.wmnet with reason: Maintenance * 07:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2206: Maintenance * 07:50 gkyziridis@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 07:50 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 07:22 root@cumin1003: START - Cookbook sre.mysql.pool pool db2154: Maintenance * 07:22 root@cumin1003: START - Cookbook sre.mysql.pool pool db2168: Maintenance * 07:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2154 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95295 and previous config saved to /var/cache/conftool/dbconfig/20260728-071640-cwilliams.json * 07:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2154.codfw.wmnet with reason: Maintenance * 07:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95294 and previous config saved to /var/cache/conftool/dbconfig/20260728-071604-cwilliams.json * 07:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2168.codfw.wmnet with reason: Maintenance * 07:08 root@cumin1003: START - Cookbook sre.mysql.pool pool db2206: Maintenance * 07:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2206 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95292 and previous config saved to /var/cache/conftool/dbconfig/20260728-070219-cwilliams.json * 07:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2206.codfw.wmnet with reason: Maintenance * 06:44 marostegui: Failover m5 from db1164 to db1228 - [[phab:T432967|T432967]] * 06:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2235].codfw.wmnet,db[1164,1217,1228].eqiad.wmnet with reason: m5 master switch [[phab:T432967|T432967]] * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.10 (duration: 02m 34s) * 03:39 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] (duration: 36m 06s) * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 02:57 dzahn@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1004.eqiad.wmnet with OS trixie * 02:57 dzahn@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - dzahn@cumin1003" * 02:55 dzahn@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - dzahn@cumin1003" * 02:37 dzahn@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1004.eqiad.wmnet with reason: host reimage * 02:31 dzahn@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1004.eqiad.wmnet with reason: host reimage * 02:16 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie * 02:15 dzahn@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host zuul1004.eqiad.wmnet with OS trixie * 01:43 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie * 01:43 dzahn@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1004.eqiad.wmnet with OS trixie * 01:25 pt1979@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-603-eqsin * 01:24 pt1979@cumin1003: START - Cookbook sre.network.tls for network device asw1-603-eqsin * 01:12 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 01:12 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt for new switches in eqsin - pt1979@cumin2003" * 01:12 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt for new switches in eqsin - pt1979@cumin2003" * 01:08 pt1979@cumin2003: START - Cookbook sre.dns.netbox * 00:48 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 00:47 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 00:47 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 00:47 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 00:26 mutante: attempting reimage with trixie on zuul1004 re-purposed physical hardware - dcops reported install issue - host was in busybox shell ([[phab:T427353|T427353]]) * 00:24 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie == 2026-07-27 == * 23:50 Amir1: mass deleting vp8 transcodes * 23:28 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:27 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1004.eqiad.wmnet with OS bullseye * 23:26 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 23:25 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:25 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 22:39 maryum: Deploy security fix for [[phab:T432877|T432877]] * 22:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1022.eqiad.wmnet with OS bookworm * 22:37 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS bullseye * 22:32 sbassett: Deployed security fix for [[phab:T432789|T432789]] * 22:22 sbassett: Deployed security patch for [[phab:T431819|T431819]] * 22:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1022.eqiad.wmnet with reason: host reimage * 22:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1022.eqiad.wmnet with reason: host reimage * 22:01 RScout-WMF: Deployed security fix for [[phab:T431819|T431819]] * 22:00 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2012 * 21:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2012.codfw.wmnet with OS bookworm * 21:55 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2011\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 21:45 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1022 * 21:45 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1022 * 21:44 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1022 * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1022.eqiad.wmnet 239.48.64.10.in-addr.arpa 9.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:44 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1022.eqiad.wmnet 239.48.64.10.in-addr.arpa 9.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:41 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:41 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 21:34 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS bookworm * 21:31 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:22 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:21 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:19 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:17 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1004 * 21:16 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1004 * 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1004] - vriley@cumin1003" * 21:15 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1004] - vriley@cumin1003" * 21:11 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:10 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2011.codfw.wmnet, repooling source-only afterwards * 21:05 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:01 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1022 * 20:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1022.eqiad.wmnet with OS bookworm * 20:53 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1021.eqiad.wmnet, repooling source-only afterwards * 20:51 mutante: zuul1001 - re-enabled puppet - revert "cherry-picked" gerrit:1314120 - [[phab:T431003|T431003]] * 20:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Maintenance * 20:15 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] (duration: 08m 03s) * 20:11 sbisson@deploy1003: sbisson: Continuing with deployment * 20:09 sbisson@deploy1003: sbisson: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] * 19:47 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 46s) * 19:47 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Maintenance * 19:27 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] (duration: 12m 26s) * 19:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2228 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95285 and previous config saved to /var/cache/conftool/dbconfig/20260727-192711-cwilliams.json * 19:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2228.codfw.wmnet with reason: Maintenance * 19:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2223: Maintenance * 19:23 krinkle@deploy1003: krinkle: Continuing with deployment * 19:16 krinkle@deploy1003: krinkle: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:15 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] * 19:12 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2238: Maintenance * 18:58 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 18:57 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 18:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2227: Maintenance * 18:57 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-experimental: apply * 18:55 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-experimental: apply * 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1021.eqiad.wmnet with OS bookworm * 18:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2011.codfw.wmnet with OS bookworm * 18:40 root@cumin1003: START - Cookbook sre.mysql.pool pool db2223: Maintenance * 18:39 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] (duration: 07m 05s) * 18:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2223 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95275 and previous config saved to /var/cache/conftool/dbconfig/20260727-183500-cwilliams.json * 18:34 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2223.codfw.wmnet with reason: Maintenance * 18:34 musikanimal@deploy1003: musikanimal: Continuing with deployment * 18:34 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2213: Maintenance * 18:33 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:32 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] * 18:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db2238: Maintenance * 18:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2011.codfw.wmnet with reason: host reimage * 18:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2238 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95271 and previous config saved to /var/cache/conftool/dbconfig/20260727-181944-cwilliams.json * 18:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2238.codfw.wmnet with reason: Maintenance * 18:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2226: Maintenance * 18:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1021.eqiad.wmnet with reason: host reimage * 18:14 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2011.codfw.wmnet with reason: host reimage * 18:12 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1021.eqiad.wmnet with reason: host reimage * 18:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db2227: Maintenance * 18:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2227 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95265 and previous config saved to /var/cache/conftool/dbconfig/20260727-180256-cwilliams.json * 18:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2227.codfw.wmnet with reason: Maintenance * 18:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2194: Maintenance * 17:57 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2011 * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2011 * 17:56 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2011 * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2011.codfw.wmnet 37.32.192.10.in-addr.arpa 7.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:56 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2011.codfw.wmnet 37.32.192.10.in-addr.arpa 7.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2011 - bking@cumin2003" * 17:56 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2011 - bking@cumin2003" * 17:52 bking@cumin2003: START - Cookbook sre.dns.netbox * 17:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2011 * 17:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1021 * 17:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1021 * 17:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2011.codfw.wmnet with OS bookworm * 17:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1021.eqiad.wmnet with OS bookworm * 17:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Maintenance * 17:38 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2010\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 17:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2213 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95260 and previous config saved to /var/cache/conftool/dbconfig/20260727-173740-cwilliams.json * 17:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2213.codfw.wmnet with reason: Maintenance * 17:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2211: Maintenance * 17:36 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1020\.eqiad\.wmnet,dc=eqiad,cluster=wdqs\-main,service=wdqs\-main * 17:32 root@cumin1003: START - Cookbook sre.mysql.pool pool db2226: Maintenance * 17:31 taavi@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] (duration: 06m 33s) * 17:27 taavi@deploy1003: taavi: Continuing with deployment * 17:27 taavi@deploy1003: taavi: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:26 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2226 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95256 and previous config saved to /var/cache/conftool/dbconfig/20260727-172636-cwilliams.json * 17:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2226.codfw.wmnet with reason: Maintenance * 17:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2225: Maintenance * 17:25 taavi@deploy1003: Started scap sync-world: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] * 17:13 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 17:11 root@cumin1003: START - Cookbook sre.mysql.pool pool db2194: Maintenance * 17:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2194 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95248 and previous config saved to /var/cache/conftool/dbconfig/20260727-170453-cwilliams.json * 17:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2194.codfw.wmnet with reason: Maintenance * 17:04 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2190: Maintenance * 16:52 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 16:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2211: Maintenance * 16:40 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95242 and previous config saved to /var/cache/conftool/dbconfig/20260727-164015-cwilliams.json * 16:40 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2211.codfw.wmnet with reason: Maintenance * 16:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2178: Maintenance * 16:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db2225: Maintenance * 16:39 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2172: Maintenance * 16:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 16:38 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 16:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2225 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95238 and previous config saved to /var/cache/conftool/dbconfig/20260727-163307-cwilliams.json * 16:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2225.codfw.wmnet with reason: Maintenance * 16:32 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2189: Maintenance * 16:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db2190: Maintenance * 16:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2190 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95230 and previous config saved to /var/cache/conftool/dbconfig/20260727-160602-cwilliams.json * 16:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2190.codfw.wmnet with reason: Maintenance * 15:53 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2177: Maintenance * 15:53 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2172: Maintenance * 15:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2178: Maintenance * 15:51 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2172: Maintenance * 15:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2178 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95224 and previous config saved to /var/cache/conftool/dbconfig/20260727-154559-cwilliams.json * 15:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2172: Maintenance * 15:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2178.codfw.wmnet with reason: Maintenance * 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2171: Maintenance * 15:44 root@cumin1003: START - Cookbook sre.mysql.pool pool db2189: Maintenance * 15:43 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:41 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2172 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95222 and previous config saved to /var/cache/conftool/dbconfig/20260727-153927-cwilliams.json * 15:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2172.codfw.wmnet with reason: Maintenance * 15:38 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2189 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95220 and previous config saved to /var/cache/conftool/dbconfig/20260727-153833-cwilliams.json * 15:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2189.codfw.wmnet with reason: Maintenance * 15:34 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:32 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 15:32 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 15:31 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:29 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:26 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 15:22 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] (duration: 07m 00s) * 15:21 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2155: Maintenance * 15:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2175: Maintenance * 15:18 zabe@deploy1003: zabe: Continuing with deployment * 15:17 zabe@deploy1003: zabe: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:15 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 15:15 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:15 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] * 15:15 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 06s) * 15:15 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:12 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2010.codfw.wmnet with OS bookworm * 15:08 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2177: Maintenance * 15:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2177: Maintenance * 14:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2177: Maintenance * 14:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2171: Maintenance * 14:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2171 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95209 and previous config saved to /var/cache/conftool/dbconfig/20260727-145236-cwilliams.json * 14:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2171.codfw.wmnet with reason: Maintenance * 14:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2177 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95208 and previous config saved to /var/cache/conftool/dbconfig/20260727-145206-cwilliams.json * 14:52 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2157: Maintenance * 14:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2177.codfw.wmnet with reason: Maintenance * 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1020.eqiad.wmnet with OS bookworm * 14:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2156: Maintenance * 14:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2010.codfw.wmnet with reason: host reimage * 14:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2010.codfw.wmnet with reason: host reimage * 14:41 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[1031,2024]*: Upgrade Cassandra to 5.0.8 (canary) - eevans@cumin1003 * 14:34 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2155: Maintenance * 14:33 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2175: Maintenance * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2010 * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2010 * 14:24 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2010 * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2010.codfw.wmnet 94.16.192.10.in-addr.arpa 4.9.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2010.codfw.wmnet 94.16.192.10.in-addr.arpa 4.9.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2010 - bking@cumin2003" * 14:24 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2010 - bking@cumin2003" * 14:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1020.eqiad.wmnet with reason: host reimage * 14:23 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[1031,2024]*: Upgrade Cassandra to 5.0.8 (canary) - eevans@cumin1003 * 14:20 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 14:20 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 14:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1020.eqiad.wmnet with reason: host reimage * 14:17 sukhe: sudo gnt-instance reboot urldownloader1005.wikimedia.org * 14:16 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:15 jelto@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:08 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2155: Maintenance * 14:05 root@cumin1003: START - Cookbook sre.mysql.pool pool db2157: Maintenance * 14:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2175: Maintenance * 14:03 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 11 hosts * 14:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db2155: Maintenance * 14:01 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 11 hosts * 14:01 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1136.eqiad.wmnet * 14:01 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1136.eqiad.wmnet * 14:01 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1136.eqiad.wmnet * 14:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db2156: Maintenance * 13:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2157 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95194 and previous config saved to /var/cache/conftool/dbconfig/20260727-135943-cwilliams.json * 13:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2157.codfw.wmnet with reason: Maintenance * 13:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db2175: Maintenance * 13:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 13:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 13:57 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2010 * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2155 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95193 and previous config saved to /var/cache/conftool/dbconfig/20260727-135613-cwilliams.json * 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2155.codfw.wmnet with reason: Maintenance * 13:55 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1020 * 13:55 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1020 * 13:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2156 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95192 and previous config saved to /var/cache/conftool/dbconfig/20260727-135413-cwilliams.json * 13:54 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2156.codfw.wmnet with reason: Maintenance * 13:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2175 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95191 and previous config saved to /var/cache/conftool/dbconfig/20260727-135300-cwilliams.json * 13:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2010.codfw.wmnet with OS bookworm * 13:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2175.codfw.wmnet with reason: Maintenance * 13:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1020.eqiad.wmnet with OS bookworm * 13:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 34 hosts * 13:46 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 34 hosts * 13:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1201: Maintenance * 13:27 Lucas_WMDE: UTC afternoon backport+config window doen * 13:18 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] (duration: 11m 57s) * 13:14 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, sihe: Continuing with deployment * 13:08 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, sihe: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:07 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool ulsfo [reason: router upgrade finished, [[phab:T431752|T431752]]] * 13:07 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool ulsfo [reason: router upgrade finished, [[phab:T431752|T431752]]] * 13:06 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] * 13:03 XioNoX: repool cr4-ulsfo - [[phab:T431752|T431752]] * 12:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db1201: Maintenance * 12:48 gkyziridis@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1201 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95186 and previous config saved to /var/cache/conftool/dbconfig/20260727-124404-cwilliams.json * 12:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1201.eqiad.wmnet with reason: Maintenance * 12:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1187: Maintenance * 12:30 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader.eqiad.wikimedia.org on all recursors * 12:30 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader.eqiad.wikimedia.org on all recursors * 12:30 sukhe@dns1004: END - running authdns-update * 12:30 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] (duration: 09m 32s) * 12:28 sukhe@dns1004: START - running authdns-update * 12:25 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 12:22 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:20 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] * 12:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts * 12:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts * 12:13 XioNoX: rebooting cr4-ulsfo for upgrade - [[phab:T431752|T431752]] * 12:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: es1038 repool * 12:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 38 hosts * 12:08 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 38 hosts * 11:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1187: Maintenance * 11:53 urbanecm@deploy1003: mwscript-k8s job started: foreachwikiindblist growthexperiments GrowthExperiments:cleanMentorList # [[phab:T431804|T431804]] * 11:50 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr4-ulsfo,cr4-ulsfo IPv6,cr4-ulsfo.mgmt with reason: router upgrade * 11:50 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] (duration: 11m 07s) * 11:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1187 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95178 and previous config saved to /var/cache/conftool/dbconfig/20260727-114844-cwilliams.json * 11:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1187.eqiad.wmnet with reason: Maintenance * 11:43 urbanecm@deploy1003: urbanecm: Continuing with deployment * 11:42 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:39 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] * 11:37 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 11:36 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 11:36 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 11:35 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 11:29 XioNoX: start draining cr4-ulsfo - [[phab:T431752|T431752]] * 11:29 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 11:29 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 11:28 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1035: testing * 11:28 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1035: testing * 11:27 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1035: testing * 11:27 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1035: testing * 11:26 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool ulsfo [reason: router upgrade, [[phab:T431752|T431752]]] * 11:26 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1038: es1038 repool * 11:26 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool ulsfo [reason: router upgrade, [[phab:T431752|T431752]]] * 11:26 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1038: testing * 11:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1264: Maintenance * 11:24 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1038: testing * 11:23 marostegui@cumin1003: dbctl commit (dc=all): 'Repool es1050 as master', diff saved to https://phabricator.wikimedia.org/P95170 and previous config saved to /var/cache/conftool/dbconfig/20260727-112326-marostegui.json * 11:23 marostegui@cumin1003: dbctl commit (dc=all): 'Repool es1050', diff saved to https://phabricator.wikimedia.org/P95169 and previous config saved to /var/cache/conftool/dbconfig/20260727-112302-marostegui.json * 11:22 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1050: testing * 11:22 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1050: testing * 11:20 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 11:18 blake@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 11:18 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 11:12 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 11:11 blake@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 11:09 blake@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 11:09 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 11:09 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 11:08 blake@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 11:05 blake@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 11:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 11:02 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 10:50 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 10:43 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:39 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply * 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1264: Maintenance * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply * 10:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply * 10:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 10:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 10:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 10:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1264 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95164 and previous config saved to /var/cache/conftool/dbconfig/20260727-103204-cwilliams.json * 10:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1264.eqiad.wmnet with reason: Maintenance * 10:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 10:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 10:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 10:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 10:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 10:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 10:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 10:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1237: Maintenance * 10:04 elukey: restart burrow main-eqiad on kafkamon2003 to clear some errors on kafka-main1008 * 09:58 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1136.eqiad.wmnet with OS trixie * 09:39 elukey: restart burrow-main-eqiad.service on kafkamon1003 to see if a recurrent kafka error on kafka-main1008 goes away * 09:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1237: Maintenance * 09:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1136.eqiad.wmnet with reason: host reimage * 09:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1237 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95159 and previous config saved to /var/cache/conftool/dbconfig/20260727-093328-cwilliams.json * 09:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1237.eqiad.wmnet with reason: Maintenance * 09:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1136.eqiad.wmnet with reason: host reimage * 09:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1203: Maintenance * 09:17 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1136 * 09:17 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1136 * 09:04 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1136 * 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1136.eqiad.wmnet 191.32.64.10.in-addr.arpa 1.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:04 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1136.eqiad.wmnet 191.32.64.10.in-addr.arpa 1.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1136 - jiji@cumin1003" * 09:04 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1136 - jiji@cumin1003" * 08:52 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 08:52 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 08:52 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 08:51 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 08:50 jiji@cumin1003: START - Cookbook sre.dns.netbox * 08:47 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1136 * 08:46 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1136.eqiad.wmnet with OS trixie * 08:44 marostegui: Rename tables on s3 [[phab:T425066|T425066]] * 08:43 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1136.eqiad.wmnet * 08:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db1203: Maintenance * 08:43 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1136.eqiad.wmnet * 08:43 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1136.eqiad.wmnet * 08:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1203 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95154 and previous config saved to /var/cache/conftool/dbconfig/20260727-083703-cwilliams.json * 08:36 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1203.eqiad.wmnet with reason: Maintenance * 08:16 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1179: Maintenance * 07:44 phuedx: UTC morning backport window done * 07:37 phuedx@deploy1003: Finished scap sync-world: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] (duration: 32m 33s) * 07:28 root@cumin1003: START - Cookbook sre.mysql.pool pool db1179: Maintenance * 07:26 marostegui: Rename tables on s3 [[phab:T426341|T426341]] * 07:25 phuedx@deploy1003: phuedx: Continuing with deployment * 07:22 marostegui: Drop tables in akwiki nawiki pihwiki - growthexperiments_* [[phab:T428885|T428885]] * 07:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95149 and previous config saved to /var/cache/conftool/dbconfig/20260727-072234-cwilliams.json * 07:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1179.eqiad.wmnet with reason: Maintenance * 07:20 phuedx@deploy1003: phuedx: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:16 ryankemper: [[phab:T430880|T430880]] [WDQS] Reimaged `wdqs1018` and `wdqs1019` to Bookworm, restored data using test-cookbook change {{Gerrit|1317128}}, and repooled both; 25/36 hosts complete * 07:04 phuedx@deploy1003: Started scap sync-world: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] * 06:57 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1019.eqiad.wmnet * 06:56 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1018.eqiad.wmnet * 06:40 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1020.eqiad.wmnet with reason: Cloning * 06:35 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db1228.eqiad.wmnet with reason: Rebooting * 06:29 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:29 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:25 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:25 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:25 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1019.eqiad.wmnet, repooling source-only afterwards * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1018.eqiad.wmnet, repooling source-only afterwards * 04:51 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1019.eqiad.wmnet, repooling source-only afterwards * 04:51 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1018.eqiad.wmnet, repooling source-only afterwards * 04:48 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s) * 04:48 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 04:48 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s) * 04:48 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 36s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-26 == * 14:59 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:59 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:59 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:59 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1019.eqiad.wmnet with OS bookworm * 01:05 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1018.eqiad.wmnet with OS bookworm * 00:43 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1019.eqiad.wmnet with reason: host reimage * 00:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1018.eqiad.wmnet with reason: host reimage * 00:34 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1019.eqiad.wmnet with reason: host reimage * 00:33 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1018.eqiad.wmnet with reason: host reimage * 00:16 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 00:16 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 00:15 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 00:15 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1019 * 00:11 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1019 * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1018 * 00:11 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1018 * 00:08 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1019.eqiad.wmnet with OS bookworm * 00:08 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1018.eqiad.wmnet with OS bookworm == 2026-07-25 == * 22:06 ryankemper: [[phab:T430880|T430880]] [WDQS] Repooled `wdqs1017` and `wdqs2024` after reimaging to bookworm, scap deploying, and data xfering * 22:04 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2024.codfw.wmnet * 22:03 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1017.eqiad.wmnet * 21:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1017.eqiad.wmnet, repooling source-only afterwards * 21:06 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2024.codfw.wmnet, repooling source-only afterwards * 20:52 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:52 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:52 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:52 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 20:18 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1017.eqiad.wmnet, repooling source-only afterwards * 20:18 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2024.codfw.wmnet, repooling source-only afterwards * 20:15 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:15 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:15 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:15 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 19:57 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s) * 19:57 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 19:57 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 07s) * 19:57 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 19:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2024.codfw.wmnet with OS bookworm * 19:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1017.eqiad.wmnet with OS bookworm * 19:02 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2024.codfw.wmnet with reason: host reimage * 18:58 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1017.eqiad.wmnet with reason: host reimage * 18:53 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2024.codfw.wmnet with reason: host reimage * 18:52 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1017.eqiad.wmnet with reason: host reimage * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2024 * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2024 * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1017 * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1017 * 18:27 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2024 * 18:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2024.codfw.wmnet 58.16.192.10.in-addr.arpa 8.5.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:26 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2024.codfw.wmnet 58.16.192.10.in-addr.arpa 8.5.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:24 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1017 * 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1017.eqiad.wmnet 238.48.64.10.in-addr.arpa 8.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:24 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1017.eqiad.wmnet 238.48.64.10.in-addr.arpa 8.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1017 - ryankemper@cumin2003" * 18:24 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1017 - ryankemper@cumin2003" * 18:23 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 18:18 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 18:17 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1017 * 18:17 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2024 * 18:14 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1017.eqiad.wmnet with OS bookworm * 18:14 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2024.codfw.wmnet with OS bookworm * 18:05 ryankemper: [WDQS] [[phab:T430880|T430880]] Reimaged `wdqs1016` and `wdqs2023` to Bookworm with `--move-vlan`, restored main and scholarly data, validated postflights, and repooled both hosts. Confirmed PyBal rebuilt both backends with their new addresses * 17:45 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2023.codfw.wmnet * 17:43 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1016.eqiad.wmnet * 06:35 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1016.eqiad.wmnet, repooling source-only afterwards * 06:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2023.codfw.wmnet, repooling source-only afterwards * 05:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2023.codfw.wmnet, repooling source-only afterwards * 05:19 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1016.eqiad.wmnet, repooling source-only afterwards * 05:07 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 07s) * 05:07 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 05:06 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 06s) * 05:06 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 03:27 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2023.codfw.wmnet with OS bookworm * 02:59 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2023.codfw.wmnet with reason: host reimage * 02:56 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2023.codfw.wmnet with reason: host reimage * 02:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2023 * 02:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2023 * 02:30 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2023 * 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2023.codfw.wmnet 35.0.192.10.in-addr.arpa 5.3.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 02:30 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2023.codfw.wmnet 35.0.192.10.in-addr.arpa 5.3.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2023 - ryankemper@cumin2003" * 02:30 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2023 - ryankemper@cumin2003" * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 26s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:15 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1016.eqiad.wmnet with OS bookworm * 00:49 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1016.eqiad.wmnet with reason: host reimage * 00:43 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1016.eqiad.wmnet with reason: host reimage * 00:31 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 00:27 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1016 * 00:27 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1016 * 00:27 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2023 * 00:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1016.eqiad.wmnet with OS bookworm * 00:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2023.codfw.wmnet with OS bookworm * 00:11 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs1014.eqiad.wmnet and wdqs2008.codfw.wmnet after Bookworm reimage, transfer, and postflight; wdqs2008 is serving, while wdqs1014 will remain outside of service until a pybal restart next monday == 2026-07-24 == * 23:54 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1014.eqiad.wmnet * 23:54 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2008.codfw.wmnet * 23:43 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2010.codfw.wmnet with OS trixie * 23:08 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 23:03 jhathaway@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 22:33 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 22:13 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 22:13 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 22:13 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 22:13 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:00 jhathaway@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 21:53 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 21:53 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie * 21:51 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 21:47 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie * 21:43 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 21:39 jhathaway@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 21:38 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 17:21 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1135.eqiad.wmnet * 17:21 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1135.eqiad.wmnet * 17:21 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1135.eqiad.wmnet * 16:34 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 16:34 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 16:34 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 16:34 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 16:33 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 16:33 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 16:28 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:28 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:28 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:28 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2008.codfw.wmnet, repooling source-only afterwards * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1014.eqiad.wmnet, repooling source-only afterwards * 15:56 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1135.eqiad.wmnet with OS trixie * 15:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 40 hosts * 15:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 40 hosts * 15:37 topranks: upgrade SR-Linux OS on lswtest-d8-eqiad * 15:36 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1135.eqiad.wmnet with reason: host reimage * 15:33 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 6 hosts with reason: upgrade lswtest-d8-eqiad * 15:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1135.eqiad.wmnet with reason: host reimage * 15:30 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc-gp2006.codfw.wmnet with OS bookworm * 15:15 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1135 * 15:15 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1135 * 15:13 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc-gp2006.codfw.wmnet with reason: host reimage * 15:08 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc-gp2006.codfw.wmnet with reason: host reimage * 14:49 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm * 14:48 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host mc-gp2006.codfw.wmnet with OS bookworm * 14:34 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] (duration: 41m 12s) * 14:32 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1135 * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1135.eqiad.wmnet 177.32.64.10.in-addr.arpa 7.7.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:32 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1135.eqiad.wmnet 177.32.64.10.in-addr.arpa 7.7.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1135 - jiji@cumin1003" * 14:32 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1135 - jiji@cumin1003" * 14:29 krinkle@deploy1003: krinkle: Continuing with deployment * 14:29 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm * 14:27 jiji@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host mc-gp2006.codfw.wmnet with OS bookworm * 14:26 jiji@cumin1003: START - Cookbook sre.dns.netbox * 14:15 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1135 * 14:14 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1135.eqiad.wmnet with OS trixie * 14:14 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1135.eqiad.wmnet * 14:13 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1135.eqiad.wmnet * 14:13 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1135.eqiad.wmnet * 14:10 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1072.eqiad.wmnet * 14:10 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1072.eqiad.wmnet * 14:10 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1072.eqiad.wmnet * 14:10 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1072.eqiad.wmnet * 14:09 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1071.eqiad.wmnet * 14:09 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1071.eqiad.wmnet * 14:09 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1071.eqiad.wmnet * 14:09 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1071.eqiad.wmnet * 13:58 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 13:58 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:58 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:57 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:55 krinkle@deploy1003: krinkle: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:53 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] * 13:45 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:45 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push new IPs for mc-gp2006 - cmooney@cumin1003" * 13:45 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push new IPs for mc-gp2006 - cmooney@cumin1003" * 13:44 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) mc-gp2006.codfw.wmnet on all recursors * 13:44 cmooney@cumin1003: START - Cookbook sre.dns.wipe-cache mc-gp2006.codfw.wmnet on all recursors * 13:42 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm * 13:41 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:30 papaul: reboot mr1-eqsin for maintenance * 13:24 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb[1029-1031].eqiad.wmnet * 13:10 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb[1029-1031].eqiad.wmnet * 11:33 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 7 hosts * 11:11 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 7 hosts * 10:56 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 7 hosts * 10:47 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 7 hosts * 10:44 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:44 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:41 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:41 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:35 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 8 hosts * 10:34 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:33 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:32 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:32 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:31 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:31 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:30 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 8 hosts * 10:24 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie * 10:19 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:18 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 16 hosts * 10:17 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2001.codfw.wmnet * 10:13 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2001.codfw.wmnet * 10:12 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2001.codfw.wmnet * 10:02 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2001.codfw.wmnet * 10:02 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2002.codfw.wmnet * 09:57 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2002.codfw.wmnet * 09:56 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2002.codfw.wmnet * 09:51 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2002.codfw.wmnet * 09:51 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1002.eqiad.wmnet * 09:47 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1002.eqiad.wmnet * 09:47 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1001.eqiad.wmnet * 09:44 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1001.eqiad.wmnet * 09:34 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2003.codfw.wmnet * 09:32 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2003.codfw.wmnet * 09:32 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2002.codfw.wmnet * 09:30 brouberol@dns1004: END - running authdns-update * 09:29 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2002.codfw.wmnet * 09:29 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2001.codfw.wmnet * 09:27 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 9 hosts * 09:27 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2001.codfw.wmnet * 09:27 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2001.codfw.wmnet * 09:26 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 9 hosts * 09:26 brouberol@dns1004: START - running authdns-update * 09:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 57 hosts * 09:24 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2001.codfw.wmnet * 09:24 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2002.codfw.wmnet * 09:22 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2002.codfw.wmnet * 09:21 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 57 hosts * 09:20 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2003.codfw.wmnet * 09:19 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 16 hosts * 09:16 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2003.codfw.wmnet * 09:16 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1003.eqiad.wmnet * 09:15 urbanecm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 09:15 urbanecm@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 09:13 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1003.eqiad.wmnet * 09:13 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1002.eqiad.wmnet * 09:11 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1002.eqiad.wmnet * 09:11 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1001.eqiad.wmnet * 09:07 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1001.eqiad.wmnet * 08:32 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:24 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:16 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 08:16 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 08:07 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:07 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:07 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 08:02 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 08:01 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:59 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:57 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 07:57 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 06:46 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1025.eqiad.wmnet with reason: Cloning * 06:46 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s4 * 06:45 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s6 * 06:44 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1019.eqiad.wmnet,service=s6 * 06:44 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1019.eqiad.wmnet,service=s4 * 03:40 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:40 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:40 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:40 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 03:37 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:37 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:37 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:36 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:49 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on mr1-eqsin,mr1-eqsin IPv6 with reason: connection issue * 02:38 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on cr[2-3]-eqsin.mgmt,ps1-[603-604]-eqsin with reason: connection issue * 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 27s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-23 == * 23:27 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin.oob,mr1-eqsin.oob IPv6 with reason: switch refresh * 22:21 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Setting storage compatibility to NONE - eevans@cumin1003 * 22:01 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Setting storage compatibility to NONE - eevans@cumin1003 * 21:29 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1014.eqiad.wmnet, repooling source-only afterwards * 21:28 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 46s) * 21:28 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 21:19 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Setting storage compatibility to UPGRADING - eevans@cumin1003 * 21:00 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Setting storage compatibility to UPGRADING - eevans@cumin1003 * 20:17 dani@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] (duration: 11m 57s) * 20:13 dani@deploy1003: dani: Continuing with deployment * 20:07 dani@deploy1003: dani: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:05 dani@deploy1003: Started scap sync-world: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] * 19:24 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:24 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating the rest of the ipv6 dns records. - jhancock@cumin2002" * 19:24 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating the rest of the ipv6 dns records. - jhancock@cumin2002" * 19:14 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 19:05 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wdqs1014.eqiad.wmnet with OS bookworm * 19:04 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.noop (exit_code=99) * 19:04 cwilliams@cumin1003: START - Cookbook sre.mysql.noop * 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1014.eqiad.wmnet with reason: host reimage * 18:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1014.eqiad.wmnet with reason: host reimage * 18:30 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2008.codfw.wmnet, repooling source-only afterwards * 18:28 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 19s) * 18:28 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1014 * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1014 * 18:22 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1014 * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1014.eqiad.wmnet 188.32.64.10.in-addr.arpa 8.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:22 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1014.eqiad.wmnet 188.32.64.10.in-addr.arpa 8.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1014 - bking@cumin2003" * 18:21 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1014 - bking@cumin2003" * 18:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2215: Maintenance * 18:18 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 18:15 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:15 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 18:06 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:06 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 18:05 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 18:04 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2052: codfw rack B8 re-pool after maintenance * 17:54 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 17:54 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:54 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 17:32 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2215: Maintenance * 17:29 cmooney@dns3003: END - running authdns-update * 17:27 cmooney@dns3003: START - running authdns-update * 17:23 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 17:22 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:18 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool es2052: codfw rack B8 re-pool after maintenance * 17:18 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2189: codfw rack B8 re-pool after maintenance * 17:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2215.codfw.wmnet with reason: Maintenance * 17:17 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 17:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2215 [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95126 and previous config saved to /var/cache/conftool/dbconfig/20260723-170903-cwilliams.json * 17:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2191 to x1 primary [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95125 and previous config saved to /var/cache/conftool/dbconfig/20260723-170612-cwilliams.json * 17:05 cezmunsta: Starting x1 codfw failover from db2215 to db2191 - [[phab:T432986|T432986]] * 16:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2191 with weight 0 [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95123 and previous config saved to /var/cache/conftool/dbconfig/20260723-165831-cwilliams.json * 16:58 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 16 hosts with reason: Primary switchover x1 [[phab:T432986|T432986]] * 16:36 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 138128 * 16:35 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 138128 * 16:33 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2189: codfw rack B8 re-pool after maintenance * 16:33 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2164: codfw rack B8 re-pool after maintenance * 16:28 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1072.eqiad.wmnet * 16:27 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1072.eqiad.wmnet with OS trixie * 16:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2249: Maintenance * 16:06 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker1072.eqiad.wmnet with reason: host reimage * 16:06 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1072.eqiad.wmnet with reason: host reimage * 15:50 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1072 * 15:50 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1072 * 15:49 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1072.eqiad.wmnet with OS trixie * 15:48 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2164: codfw rack B8 re-pool after maintenance * 15:48 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] (duration: 06m 37s) * 15:48 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2163: codfw rack B8 re-pool after maintenance * 15:45 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 15:43 musikanimal@deploy1003: musikanimal: Continuing with deployment * 15:43 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:41 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] * 15:36 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1072.eqiad.wmnet * 15:35 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1072.eqiad.wmnet * 15:35 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1072.eqiad.wmnet * 15:34 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:34 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push any outstanding updates - cmooney@cumin1003" * 15:34 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push any outstanding updates - cmooney@cumin1003" * 15:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db2249: Maintenance * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 15:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:26 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:21 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:21 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 15:21 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:21 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 15:20 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 15:19 cmooney@dns2004: END - running authdns-update * 15:17 cmooney@dns2004: START - running authdns-update * 15:14 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns2004.wikimedia.org * 15:12 brouberol@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 15:12 brouberol@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 15:12 klausman@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ml-serve1001.eqiad.wmnet with OS trixie * 15:11 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1071.eqiad.wmnet * 15:11 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1071.eqiad.wmnet with OS trixie * 15:10 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wdqs2008.codfw.wmnet with OS bookworm * 15:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2249.codfw.wmnet with reason: Maintenance * 15:08 brouberol@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 15:08 brouberol@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 15:08 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 15:08 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 15:06 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns1004.wikimedia.org * 15:02 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2002.codfw.wmnet * 15:02 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2002.codfw.wmnet * 15:02 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2163: codfw rack B8 re-pool after maintenance * 15:01 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 15:01 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 14:59 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test2001.codfw.wmnet * 14:57 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test2001.codfw.wmnet * 14:56 ryankemper: [WDQS] [[phab:T430880|T430880]] Reimaged `wdqs2016` to Bookworm, xferred scholarly_articles from `wdqs2024`, validated updater/readiness/federation, and repooled * 14:51 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2016.codfw.wmnet * 14:51 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1001.eqiad.wmnet with reason: host reimage * 14:48 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1071.eqiad.wmnet with reason: host reimage * 14:47 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1001.eqiad.wmnet with reason: host reimage * 14:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2008.codfw.wmnet with reason: host reimage * 14:43 topranks: reboot lsw1-b8-codw to upgrade JunOS [[phab:T430929|T430929]] * 14:41 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1071.eqiad.wmnet with reason: host reimage * 14:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2008.codfw.wmnet with reason: host reimage * 14:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2231: Maintenance * 14:30 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1001.eqiad.wmnet with OS trixie * 14:25 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2002.codfw.wmnet * 14:23 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1071 * 14:23 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1071 * 14:23 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 14:22 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore scholarly data after Bookworm reimage) xfer scholarly_articles from wdqs2024.codfw.wmnet -> wdqs2016.codfw.wmnet, repooling source-only afterwards * 14:22 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2052: codfw rack B8 depool for maintenance * 14:21 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool es2052: codfw rack B8 depool for maintenance * 14:21 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2249: codfw rack B8 depool for maintenance * 14:21 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1071 * 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1071.eqiad.wmnet 166.48.64.10.in-addr.arpa 6.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:21 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1071.eqiad.wmnet 166.48.64.10.in-addr.arpa 6.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1071 - jiji@cumin1003" * 14:21 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1071 - jiji@cumin1003" * 14:21 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2249: codfw rack B8 depool for maintenance * 14:21 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2189: codfw rack B8 depool for maintenance * 14:20 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2002.codfw.wmnet * 14:20 cmooney@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2050.codfw.wmnet * 14:20 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2189: codfw rack B8 depool for maintenance * 14:20 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2164: codfw rack B8 depool for maintenance * 14:20 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2164: codfw rack B8 depool for maintenance * 14:19 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2163: codfw rack B8 depool for maintenance * 14:19 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:19 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1014 * 14:19 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2163: codfw rack B8 depool for maintenance * 14:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1014.eqiad.wmnet with OS bookworm * 14:17 cmooney@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2050.codfw.wmnet * 14:16 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 14:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2008 * 14:14 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2008 * 14:14 cmooney@cumin1003: conftool action : set/pooled=no; selector: name=dns2004.wikimedia.org * 14:14 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2008.codfw.wmnet with OS bookworm * 14:13 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 14:12 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:10 topranks: depool dns2004 before lsw1-b8-codfw switch maintenance [[phab:T430929|T430929]] * 14:10 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b8-codfw,lsw1-b8-codfw IPv6,lsw1-b8-codfw.mgmt,ssw1-a[1,8]-codfw with reason: lsw1-b8-codfw JunOS upgrade * 14:07 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 30 hosts with reason: lsw1-b8-codfw JunOS upgrade * 14:06 elukey: upload python3-docker-report 0.0.19 to apt.wikimedia.org for bookworm and trixie * 13:59 jiji@cumin1003: START - Cookbook sre.dns.netbox * 13:58 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1071 * 13:57 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:57 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:57 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:57 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1071.eqiad.wmnet with OS trixie * 13:55 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1071.eqiad.wmnet * 13:55 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1071.eqiad.wmnet * 13:55 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1071.eqiad.wmnet * 13:53 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:52 logmsgbot: kharlan Deployed security patch for [[phab:T432948|T432948]] * 13:51 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:51 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 13:51 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db2231: Maintenance * 13:50 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:50 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:50 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:50 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:49 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:49 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:49 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2231 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95097 and previous config saved to /var/cache/conftool/dbconfig/20260723-134436-cwilliams.json * 13:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2231.codfw.wmnet with reason: Maintenance * 13:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:39 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:38 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 13:38 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] (duration: 09m 07s) * 13:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1037 hosts * 13:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2196: Maintenance * 13:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:33 kharlan@deploy1003: kharlan, emc-wmf: Continuing with deployment * 13:31 kharlan@deploy1003: kharlan, emc-wmf: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:30 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:28 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] * 13:17 hashar@deploy1003: Finished deploy [integration/docroot@2199146]: build: License GPL2.0+ / updating npm dependencies (duration: 00m 14s) * 13:17 hashar@deploy1003: Started deploy [integration/docroot@2199146]: build: License GPL2.0+ / updating npm dependencies * 13:14 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service * 13:07 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 12:58 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2207: Repooling * 12:49 root@cumin1003: START - Cookbook sre.mysql.pool pool db2196: Maintenance * 12:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2196 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95087 and previous config saved to /var/cache/conftool/dbconfig/20260723-123952-cwilliams.json * 12:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2196.codfw.wmnet with reason: Maintenance * 12:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2191: Maintenance * 12:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:13 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:13 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: Repooling * 12:12 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2207: Repooling * 12:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: Repooling * 11:56 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2235.codfw.wmnet with OS trixie * 11:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db2191: Maintenance * 11:46 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1070.eqiad.wmnet * 11:46 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1070.eqiad.wmnet * 11:46 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1070.eqiad.wmnet * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2191 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95080 and previous config saved to /var/cache/conftool/dbconfig/20260723-114308-cwilliams.json * 11:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2191.codfw.wmnet with reason: Maintenance * 11:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2186: Maintenance * 11:35 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 46375 * 11:34 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 46375 * 11:33 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2235.codfw.wmnet with reason: host reimage * 11:28 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2235.codfw.wmnet with reason: host reimage * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c7-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c7-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c6-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c6-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c5-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c5-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c4-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c4-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c3-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c3-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c2-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c2-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d7-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d7-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d4-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d4-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d3-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d2-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d2-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d8-eqiad * 11:23 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d8-eqiad * 11:23 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d1-eqiad * 11:23 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d1-eqiad * 11:12 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db2235.codfw.wmnet with OS trixie * 11:11 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:11 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[2160,2235].codfw.wmnet with reason: Upgrading * 10:56 root@cumin1003: START - Cookbook sre.mysql.pool pool db2186: Maintenance * 10:54 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1037: testing * 10:53 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1037: testing * 10:53 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1037: testing * 10:53 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1037: testing * 10:52 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: testing * 10:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2186 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95072 and previous config saved to /var/cache/conftool/dbconfig/20260723-104956-cwilliams.json * 10:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2186.codfw.wmnet with reason: Maintenance * 10:43 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1054: testing * 10:41 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1070.eqiad.wmnet with OS trixie * 10:30 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1037 hosts * 10:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 10:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 10:20 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1070.eqiad.wmnet with reason: host reimage * 10:16 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1070.eqiad.wmnet with reason: host reimage * 10:06 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1038: testing * 10:05 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1038: testing * 10:05 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1038: testing * 10:02 marostegui@dns1004: END - running authdns-update * 10:00 marostegui@dns1004: START - running authdns-update * 09:58 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: testing * 09:57 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1054: testing * 09:57 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1070 * 09:57 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1070 * 09:57 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1054: testing * 09:57 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1054: testing * 09:56 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1055: testing * 09:56 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1070 * 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1070.eqiad.wmnet 165.48.64.10.in-addr.arpa 5.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:56 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1070.eqiad.wmnet 165.48.64.10.in-addr.arpa 5.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1070 - jiji@cumin1003" * 09:56 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1070 - jiji@cumin1003" * 09:47 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for 1035 hosts * 09:45 jiji@cumin1003: START - Cookbook sre.dns.netbox * 09:42 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1070 * 09:42 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1070.eqiad.wmnet with OS trixie * 09:42 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1070.eqiad.wmnet * 09:41 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1070.eqiad.wmnet * 09:41 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1070.eqiad.wmnet * 09:27 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es2051: testing * 09:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: testing * 09:12 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es2051: testing * 09:11 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1055: testing * 09:09 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:09 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1055: testing * 09:09 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1055: testing * 08:50 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:50 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:50 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 08:49 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 08:49 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 08:49 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:46 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1069.eqiad.wmnet * 08:46 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1069.eqiad.wmnet * 08:46 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1069.eqiad.wmnet * 08:39 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 08:38 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 08:38 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 08:37 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 08:35 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:10 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1069.eqiad.wmnet with OS trixie * 07:49 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1069.eqiad.wmnet with reason: host reimage * 07:45 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1069.eqiad.wmnet with reason: host reimage * 07:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1035 hosts * 07:33 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 07:32 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 07:29 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1069 * 07:29 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1069 * 07:26 jiji@deploy1003: Finished scap sync-world: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules (duration: 06m 01s) * 07:25 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1069 * 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1069.eqiad.wmnet 164.48.64.10.in-addr.arpa 4.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:25 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1069.eqiad.wmnet 164.48.64.10.in-addr.arpa 4.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1069 - jiji@cumin1003" * 07:25 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1069 - jiji@cumin1003" * 07:25 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 07:25 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 07:24 jiji@deploy1003: jiji: Continuing with deployment * 07:22 jiji@deploy1003: jiji: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:21 jiji@deploy1003: Started scap sync-world: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules * 07:21 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1031.eqiad.wmnet,service=s7 * 07:20 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1031.eqiad.wmnet,service=s2 * 07:20 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1031.eqiad.wmnet,service=s7 * 07:20 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1031.eqiad.wmnet,service=s2 * 07:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts * 07:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts * 07:19 jiji@cumin1003: START - Cookbook sre.dns.netbox * 07:19 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1069 * 07:19 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1069.eqiad.wmnet with OS trixie * 07:19 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1069.eqiad.wmnet * 07:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 45 hosts * 07:17 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1069.eqiad.wmnet * 07:17 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1069.eqiad.wmnet * 07:14 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 45 hosts * 07:13 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 06:16 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs2007 after successful Bookworm reimage, data transfer, and postflight validation; wdqs1013 also passed postflights and is enabled in conftool, but remains out of IPVS pending a rolling pybal restart to clear its stale pre-VLAN-move address * 05:58 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2007.codfw.wmnet * 05:58 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1013.eqiad.wmnet * 05:54 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore scholarly data after Bookworm reimage) xfer scholarly_articles from wdqs2024.codfw.wmnet -> wdqs2016.codfw.wmnet, repooling source-only afterwards == 2026-07-22 == * 23:34 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Apply upgrade to JVM17 - eevans@cumin1003 * 23:14 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Apply upgrade to JVM17 - eevans@cumin1003 * 22:06 ryankemper: [WDQS] Added requestctl per-IP ratelimit `wdqs_heavy_sparql_bots_jul_2026_ratelimit` (chronic heavy-query bot tier driving deadlock-remediation restarts); pruned superseded `wdqs_2026_05_11_worobot` * 21:51 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] (duration: 11m 52s) * 21:44 sbassett@deploy1003: sbassett: Continuing with deployment * 21:43 sbassett@deploy1003: sbassett: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:39 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] * 20:38 dani@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] (duration: 32m 51s) * 20:38 ryankemper: [WDQS] Pruned obsolete requestctl action+pattern `wdqs_20260715_p2003_ring_ja3n` (actor rotated JA3Ns; rule inert) * 20:26 dani@deploy1003: dani, vadymts1: Continuing with deployment * 20:24 dani@deploy1003: dani, vadymts1: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:14 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2244: Testing * 20:06 dani@deploy1003: Started scap sync-world: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] * 19:56 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1013.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2016.codfw.wmnet with OS bookworm * 19:40 mutante: gerrit - one more service restart is needed - restarting * 19:29 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2244: Testing * 19:27 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2244: Testing * 19:27 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2244: Testing * 19:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2016.codfw.wmnet with reason: host reimage * 19:15 dancy@deploy1003: Finished deploy [zuul/deploy@d92e238]: Freshening Zuul installation (duration: 00m 15s) * 19:14 dancy@deploy1003: Started deploy [zuul/deploy@d92e238]: Freshening Zuul installation * 19:11 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2016.codfw.wmnet with reason: host reimage * 18:54 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1013.eqiad.wmnet, repooling source-only afterwards * 18:52 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2016 * 18:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2016 * 18:51 dancy@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 18:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2016.codfw.wmnet with OS bookworm * 18:39 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 11s) * 18:39 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 18:36 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 18:30 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417] (thin): Regular analytics weekly train THIN [analytics/refinery@2a25417d] (duration: 02m 09s) * 18:28 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417] (thin): Regular analytics weekly train THIN [analytics/refinery@2a25417d] * 18:28 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417]: Regular analytics weekly train [analytics/refinery@2a25417d] (duration: 04m 31s) * 18:27 dduvall: deploying https://gerrit.wikimedia.org/r/c/integration/config/+/1314025 (4 jobs updated) * 18:23 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417]: Regular analytics weekly train [analytics/refinery@2a25417d] * 18:22 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@2a25417d] (duration: 01m 59s) * 18:20 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@2a25417d] * 17:56 Raine: deployment server switchover => deploy1003 is primary now * 17:55 kamila@deploy1003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 28m 24s) * 17:54 mutante: restarting gerrit for maintenance * 17:29 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1013.eqiad.wmnet with OS bookworm * 17:27 kamila@deploy1003: Started scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] * 17:20 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] (duration: 22m 50s) * 17:12 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1023.eqiad.wmnet -> wdqs1024.eqiad.wmnet, repooling source-only afterwards * 17:04 Raine: point deployment.eqiad.wmnet to deploy1003 * 17:04 kamila@dns7001: END - running authdns-update * 17:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1013.eqiad.wmnet with reason: host reimage * 17:02 kamila@dns7001: START - running authdns-update * 17:01 kamila@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on releases2003.codfw.wmnet,releases1003.eqiad.wmnet with reason: Deployment server switchover * 17:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1013.eqiad.wmnet with reason: host reimage * 16:58 kamila@deploy2003: Locking from deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] * 16:57 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] (duration: 02m 33s) * 16:55 kamila@deploy2003: Locking from deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] * 16:55 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2003 - [[phab:T240266|T240266]] (duration: 00m 11s) * 16:54 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2003 - [[phab:T240266|T240266]] * 16:40 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1013 * 16:40 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1013 * 16:39 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1013 * 16:39 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1013.eqiad.wmnet 105.32.64.10.in-addr.arpa 5.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:39 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1013.eqiad.wmnet 105.32.64.10.in-addr.arpa 5.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:39 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:39 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1013 - bking@cumin2003" * 16:39 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1013 - bking@cumin2003" * 16:34 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:34 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1013 * 16:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1013.eqiad.wmnet with OS bookworm * 16:28 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1023.eqiad.wmnet -> wdqs1024.eqiad.wmnet, repooling source-only afterwards * 16:27 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-scholarly,name=eqiad * 16:27 eevans@deploy2003: helmfile [eqiad] DONE helmfile.d/services/linked-artifacts: apply * 16:26 eevans@deploy2003: helmfile [eqiad] START helmfile.d/services/linked-artifacts: apply * 16:26 eevans@deploy2003: helmfile [codfw] DONE helmfile.d/services/linked-artifacts: apply * 16:26 eevans@deploy2003: helmfile [codfw] START helmfile.d/services/linked-artifacts: apply * 16:25 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 29s) * 16:25 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 16:24 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 16:21 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 16:18 eevans@deploy2003: helmfile [codfw] DONE helmfile.d/services/linked-artifacts: apply * 16:18 eevans@deploy2003: helmfile [codfw] START helmfile.d/services/linked-artifacts: apply * 16:08 eevans@deploy2003: helmfile [staging] DONE helmfile.d/services/linked-artifacts: apply * 16:07 eevans@deploy2003: helmfile [staging] START helmfile.d/services/linked-artifacts: apply * 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 16:01 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 15:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1024.eqiad.wmnet with OS bookworm * 15:49 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] (duration: 00m 10s) * 15:49 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] * 15:48 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] (duration: 00m 15s) * 15:48 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] * 15:47 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] (duration: 00m 10s) * 15:47 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] * 15:46 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:42 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:42 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:40 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:37 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 15:37 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:36 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:36 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:36 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 15:33 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 15:33 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:31 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:28 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 15:27 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:27 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1024.eqiad.wmnet with reason: host reimage * 15:23 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:23 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:23 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:20 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1068.eqiad.wmnet * 15:20 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1068.eqiad.wmnet * 15:20 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1068.eqiad.wmnet * 15:20 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 15:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1024.eqiad.wmnet with reason: host reimage * 15:11 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:55 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 14:52 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wdqs1024.eqiad.wmnet with OS bookworm * 14:50 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] (duration: 00m 09s) * 14:50 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] * 14:49 jiji@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 14:49 jiji@deploy2003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 14:49 jiji@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 14:48 jiji@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 14:45 ecarg@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:45 ecarg@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:44 ecarg@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:44 ecarg@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:43 ecarg@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:43 ecarg@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:41 sukhe: ipvsadm --delete-service --tcp-service 10.2.1.55:8087: lvs2014 and lvs2013 * 14:39 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:39 sukhe: ipvsadm --delete-service --tcp-service 10.2.2.55:8087: [[phab:T432445|T432445]] * 14:38 ecarg@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:38 ecarg@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:37 ecarg@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:37 ecarg@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:36 ecarg@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:34 ecarg@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts datahubsearch1001.eqiad.wmnet * 14:32 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:32 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 14:31 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 14:31 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 14:28 sukhe: sudo cumin 'A:lvs-low-traffic-codfw' 'systemctl restart pybal': lvs2013 * 14:26 sukhe: sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal': lvs2014 * 14:26 sukhe: sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal' * 14:24 sukhe: restart pybal on lvs1019 * 14:24 sukhe: restart pybal on lvs1020 * 14:19 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for 1036 hosts * 14:17 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:04 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] (duration: 09m 28s) * 13:59 kharlan@deploy2003: dreamyjazz, kharlan: Continuing with deployment * 13:58 bking@cumin2003: START - Cookbook sre.hosts.decommission for hosts datahubsearch1001.eqiad.wmnet * 13:57 kharlan@deploy2003: dreamyjazz, kharlan: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:55 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] * 13:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts datahubsearch[1002-1003].eqiad.wmnet * 13:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:53 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch[1002-1003].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 13:52 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch[1002-1003].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 13:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 13:42 stran@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] (duration: 07m 30s) * 13:42 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:40 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: Test * 13:38 stran@deploy2003: dragoniez, stran: Continuing with deployment * 13:37 stran@deploy2003: dragoniez, stran: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:35 bking@cumin2003: START - Cookbook sre.hosts.decommission for hosts datahubsearch[1002-1003].eqiad.wmnet * 13:35 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024'] * 13:35 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 13:35 stran@deploy2003: Started scap sync-world: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] * 13:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 13:28 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024'] * 13:26 sukhe@dns1004: END - running authdns-update * 13:25 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 13:25 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 13:24 sukhe@dns1004: START - running authdns-update * 13:22 stran@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] (duration: 08m 20s) * 13:21 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 13:20 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 13:19 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 13:19 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 13:19 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 13:18 stran@deploy2003: stran: Continuing with deployment * 13:16 stran@deploy2003: stran: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:14 stran@deploy2003: Started scap sync-world: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] * 13:13 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 13:13 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 13:11 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 13:11 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 13:08 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 12:55 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool es2051: Test * 12:55 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: Test * 12:54 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool es2051: Test * 12:43 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1036 hosts * 12:41 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 12:40 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1048.eqiad.wmnet * 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 12:39 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 12:38 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 12:37 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 12:37 brouberol@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 12:36 brouberol@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 12:36 brouberol@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 12:36 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 12:36 elukey@cumin1003: DONE (PASS) - Cookbook sre.puppet.renew-cert (exit_code=0) for crm2001.codfw.wmnet: Renew puppet certificate - elukey@cumin1003 * 12:35 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:35 brouberol@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 12:34 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 12:31 brouberol@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 12:30 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1048.eqiad.wmnet * 12:30 brouberol@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 12:28 brouberol@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 12:27 brouberol@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 12:20 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1068.eqiad.wmnet with OS trixie * 12:01 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] (duration: 13m 19s) * 11:58 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1068.eqiad.wmnet with reason: host reimage * 11:52 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1068.eqiad.wmnet with reason: host reimage * 11:51 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 11:49 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:47 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] * 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2252: Security updates * 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:43 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 11:42 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2252: Security updates * 11:42 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply * 11:40 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply * 11:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 11:37 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2252.codfw.wmnet with OS trixie * 11:34 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1068 * 11:34 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1068 * 11:26 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1068 * 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1068.eqiad.wmnet 46.48.64.10.in-addr.arpa 6.4.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:26 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1068.eqiad.wmnet 46.48.64.10.in-addr.arpa 6.4.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1068 - jiji@cumin1003" * 11:26 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1068 - jiji@cumin1003" * 11:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2252.codfw.wmnet with reason: host reimage * 11:17 jiji@cumin1003: START - Cookbook sre.dns.netbox * 11:17 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1068 * 11:17 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1068.eqiad.wmnet with OS trixie * 11:17 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2252.codfw.wmnet with reason: host reimage * 11:15 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1068.eqiad.wmnet * 11:15 mvolz@deploy2003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:15 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1068.eqiad.wmnet * 11:15 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1068.eqiad.wmnet * 11:14 mvolz@deploy2003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:13 mvolz@deploy2003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:13 mvolz@deploy2003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:12 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] (duration: 11m 05s) * 11:11 mvolz@deploy2003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:10 mvolz@deploy2003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:07 dreamyjazz@deploy2003: dreamyjazz, kharlan: Continuing with deployment * 11:03 dreamyjazz@deploy2003: dreamyjazz, kharlan: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:03 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2252.codfw.wmnet with OS trixie * 11:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2252: Upgrading db2252.codfw.wmnet * 11:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:02 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 11:02 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2252: Upgrading db2252.codfw.wmnet * 11:02 cwilliams@cumin1003: dbmaint on ms3@codfw [[phab:T432321|T432321]] * 11:01 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 11:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db1153.eqiad.wmnet with reason: Security updates * 11:01 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] * 11:00 fnegri@deploy2003: helmfile [eqiad] DONE helmfile.d/services/toolhub: apply * 10:58 fnegri@deploy2003: helmfile [eqiad] START helmfile.d/services/toolhub: apply * 10:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1151: Security updates * 10:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:57 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1151: Security updates * 10:55 fnegri@deploy2003: helmfile [codfw] DONE helmfile.d/services/toolhub: apply * 10:54 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] (duration: 08m 38s) * 10:53 fnegri@deploy2003: helmfile [codfw] START helmfile.d/services/toolhub: apply * 10:53 fnegri@deploy2003: helmfile [staging] DONE helmfile.d/services/toolhub: apply * 10:52 fnegri@deploy2003: helmfile [staging] START helmfile.d/services/toolhub: apply * 10:50 zabe@deploy2003: zabe: Continuing with deployment * 10:47 zabe@deploy2003: zabe: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:45 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] * 10:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1151: Security updates * 10:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:42 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:42 root@cumin1003: START - Cookbook sre.mysql.depool depool db1151: Security updates * 10:38 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] (duration: 12m 47s) * 10:34 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 10:34 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 10:33 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2253.codfw.wmnet with OS trixie * 10:28 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:26 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] * 10:18 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2253.codfw.wmnet with reason: host reimage * 10:13 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2253.codfw.wmnet with reason: host reimage * 10:00 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2253.codfw.wmnet with OS trixie * 09:58 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db1151.eqiad.wmnet with reason: Security updates * 09:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2253: Upgrading db2253.codfw.wmnet * 09:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:57 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 09:56 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2253: Upgrading db2253.codfw.wmnet * 09:56 cwilliams@cumin1003: dbmaint on ms2@codfw [[phab:T432321|T432321]] * 09:56 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 09:36 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: UI improvement; support url shortener - oblivian@cumin1003" * 09:36 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: UI improvement; support url shortener - oblivian@cumin1003 * 09:35 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: UI improvement; support url shortener - oblivian@cumin1003 * 09:35 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: UI improvement; support url shortener - oblivian@cumin1003" * 09:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1152: Security updates * 09:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:26 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db1152: Security updates * 09:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: Security updates * 09:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:11 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:11 root@cumin1003: START - Cookbook sre.mysql.depool depool db1152: Security updates * 09:10 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1018.eqiad.wmnet with reason: Cloning * 09:09 Dreamy_Jazz: Deployed patch for [[phab:T432453|T432453]] and [[phab:T432454|T432454]] * 09:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 09:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2251.codfw.wmnet with OS trixie * 08:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2251.codfw.wmnet with reason: host reimage * 08:45 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2251.codfw.wmnet with reason: host reimage * 08:40 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] (duration: 12m 26s) * 08:38 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1030.eqiad.wmnet,service=s1 * 08:36 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 08:31 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2251.codfw.wmnet with OS trixie * 08:30 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:28 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] * 08:25 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] (duration: 07m 59s) * 08:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2251: Upgrading db2251.codfw.wmnet * 08:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:22 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 08:22 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2251: Upgrading db2251.codfw.wmnet * 08:20 urbanecm@deploy2003: urbanecm: Continuing with deployment * 08:20 cwilliams@cumin1003: dbmaint on ms1@codfw [[phab:T432321|T432321]] * 08:20 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 08:19 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:17 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] * 08:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade * 08:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade * 08:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db2251.codfw.wmnet,db1152.eqiad.wmnet with reason: OS upgrade * 08:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade * 08:13 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade * 08:11 Dreamy_Jazz: Created cusi_signal, cusi_case, and cusi_user on ukwiki and enwikivoyage in extension1 * 08:11 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade * 08:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade * 08:04 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1030.eqiad.wmnet,service=s1 * 08:04 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1030.eqiad.wmnet,service=s1 * 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply * 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply * 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply * 07:51 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply * 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 07:47 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 07:47 phuedx: End of UTC morning backport window * 07:43 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 07:43 phuedx@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] (duration: 13m 44s) * 07:43 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 07:39 phuedx@deploy2003: phuedx: Continuing with deployment * 07:31 phuedx@deploy2003: phuedx: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:29 phuedx@deploy2003: Started scap sync-world: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] * 07:24 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Turnilo import support - oblivian@cumin1003" * 07:24 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import support - oblivian@cumin1003 * 07:23 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import support - oblivian@cumin1003 * 07:23 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Turnilo import support - oblivian@cumin1003" * 06:42 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs2020 after successful Bookworm reimage, data transfer, and postflight validation * 06:42 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2020.codfw.wmnet * 05:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Managing sanitization for wikis bolwiki in section s5 * 05:25 marostegui@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis bolwiki in section s5 * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 41s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-21 == * 22:50 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2019.codfw.wmnet -> wdqs2020.codfw.wmnet, repooling source-only afterwards * 22:47 cwhite: force reboot arclamp2001 - appears to have run out of memory and gone unresponsive * 22:24 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 01m 26s) * 22:24 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 22:23 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 22:22 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024'] * 22:11 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 21:54 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs1024'] * 21:54 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 21:53 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs1024'] * 21:53 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 21:49 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2019.codfw.wmnet -> wdqs2020.codfw.wmnet, repooling source-only afterwards * 20:57 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] (duration: 09m 10s) * 20:55 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1024.eqiad.wmnet with OS bookworm * 20:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2020.codfw.wmnet with OS bookworm * 20:52 krinkle@deploy2003: krinkle: Continuing with deployment * 20:49 krinkle@deploy2003: krinkle: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:47 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] * 20:45 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] (duration: 05m 42s) * 20:44 krinkle@deploy2003: krinkle: Rolling back deployment * 20:41 krinkle@deploy2003: krinkle: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:39 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] * 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2003.codfw.wmnet * 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1003.eqiad.wmnet * 20:33 dani@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] (duration: 11m 15s) * 20:33 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2003.codfw.wmnet * 20:33 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1003.eqiad.wmnet * 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2020.codfw.wmnet with reason: host reimage * 20:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1002.eqiad.wmnet * 20:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2002.codfw.wmnet * 20:30 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 20:29 dani@deploy2003: dani: Continuing with deployment * 20:29 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2020.codfw.wmnet with reason: host reimage * 20:26 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1002.eqiad.wmnet * 20:26 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2002.codfw.wmnet * 20:24 dani@deploy2003: dani: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2001.codfw.wmnet * 20:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1001.eqiad.wmnet * 20:22 dani@deploy2003: Started scap sync-world: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] * 20:22 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 20:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2001.codfw.wmnet * 20:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1001.eqiad.wmnet * 20:14 sbisson@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] (duration: 09m 01s) * 20:11 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2020.codfw.wmnet with OS bookworm * 20:10 sbisson@deploy2003: sbisson: Continuing with deployment * 20:07 sbisson@deploy2003: sbisson: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:05 sbisson@deploy2003: Started scap sync-world: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] * 20:03 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] (duration: 07m 04s) * 20:01 mutante: Gerrit - tomorrow a new SSH host key will appear - it will be {{Gerrit|ed25519}} and has already been added to wmf-laptop. you can verify it here: https://wikitech.wikimedia.org/wiki/Help:SSH_Fingerprints/gerrit.wikimedia.org:29418 ([[phab:T240266|T240266]]) * 19:59 zabe@deploy2003: zabe: Continuing with deployment * 19:58 zabe@deploy2003: zabe: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:56 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] * 19:52 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] (duration: 07m 25s) * 19:48 zabe@deploy2003: zabe: Continuing with deployment * 19:47 zabe@deploy2003: zabe: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:45 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] * 19:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 19:32 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024'] * 19:27 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 19:26 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024'] * 19:26 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 19:24 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024'] * 19:12 ryankemper: [wdqs] [[phab:T430880|T430880]] Repooled `wdqs-scholarly` discovery in `eqiad` after validating `wdqs1023` end-to-end; `wdqs1024` remains disabled pending reimage recovery * 19:11 ryankemper: [wdqs] [[phab:T430880|T430880]] Repooled wdqs1012.eqiad.wmnet after successful Bookworm reimage, data transfer, service checks, readiness probe, and cross-graph federation query validation * 19:10 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 19:10 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1012.eqiad.wmnet * 19:08 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 18:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for deploy1003.eqiad.wmnet * 18:57 kamila@cumin1003: START - Cookbook sre.hosts.remove-downtime for deploy1003.eqiad.wmnet * 18:37 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:37 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding urldownloader service IPs - sukhe@cumin1003" * 18:37 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding urldownloader service IPs - sukhe@cumin1003" * 18:32 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 18:32 dancy@deploy2003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 18:30 sukhe@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 18:27 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 18:24 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1024.eqiad.wmnet with OS bookworm * 18:20 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host deploy1003.eqiad.wmnet with OS bookworm * 18:09 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deploy1003 reimage (duration: 121m 16s) * 18:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 18:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1180: Security updates * 17:55 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wcqs2003.codfw.wmnet * 17:48 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wcqs2003.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1155.eqiad.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1155.eqiad.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2224.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2224.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2217.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2217.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2193.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2193.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2180.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2180.codfw.wmnet * 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1168.eqiad.wmnet * 17:36 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1168.eqiad.wmnet * 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2169.codfw.wmnet * 17:36 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2169.codfw.wmnet * 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1165.eqiad.wmnet * 17:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1165.eqiad.wmnet * 17:35 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2158.codfw.wmnet * 17:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2158.codfw.wmnet * 17:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wcqs1003.eqiad.wmnet * 17:17 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1180: Security updates * 17:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1180.eqiad.wmnet * 17:16 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1180.eqiad.wmnet * 17:15 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp3073.* * 17:13 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wcqs1003.eqiad.wmnet * 17:11 brett@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp3073.esams.wmnet with OS trixie * 17:11 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 17:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1024 * 17:04 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1024 * 17:03 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 17:00 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 16:59 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2242: codfw rack B7 depool for maintenance * 16:59 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 16:43 brett@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp3073.esams.wmnet with reason: host reimage * 16:42 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 16:39 brett@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cp3073.esams.wmnet with reason: host reimage * 16:32 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on deploy1003.eqiad.wmnet with reason: host reimage * 16:27 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on deploy1003.eqiad.wmnet with reason: host reimage * 16:14 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2242: codfw rack B7 depool for maintenance * 16:14 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: codfw rack B7 depool for maintenance * 16:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1023.eqiad.wmnet with OS bookworm * 16:13 brett@cumin2002: START - Cookbook sre.hosts.reimage for host cp3073.esams.wmnet with OS trixie * 16:08 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host deploy1003.eqiad.wmnet with OS bookworm * 16:08 kamila@deploy2003: Locking from deployment [MediaWiki]: deploy1003 reimage * 16:03 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1012.eqiad.wmnet with OS bookworm * 15:48 inflatador: bking@apt1002 `sudo reprepro copy bookworm-wikimedia bullseye-wikimedia jvmquake` [[phab:T430880|T430880]] * 15:39 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp3073.* * 15:39 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 15:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:34 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1311536{{!}}Set $wgMathInternalRestbaseURL explicitly (take 2) (T349582)]] * 15:29 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:29 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2228: codfw rack B7 depool for maintenance * 15:29 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2229: codfw rack B7 depool for maintenance * 15:27 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1180: Security update * 15:25 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Security update * 15:21 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 15:21 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 15:19 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 24s) * 15:19 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:14 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 15:14 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db1180: Security update * 15:13 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-eqiad * 14:48 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-eqiad * 14:44 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2229: codfw rack B7 depool for maintenance * 14:44 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc2017: codfw rack B7 depool for maintenance * 14:44 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:43 cmooney@cumin2003: START - Cookbook sre.mysql.parsercache * 14:43 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool pc2017: codfw rack B7 depool for maintenance * 14:43 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2003.codfw.wmnet * 14:43 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2003.codfw.wmnet * 14:42 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2009.codfw.wmnet * 14:42 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2009.codfw.wmnet * 14:41 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:41 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:40 cmooney@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 29 hosts * 14:40 cmooney@cumin1003: START - Cookbook sre.hosts.remove-downtime for 29 hosts * 14:35 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 14:34 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 14:32 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] (duration: 07m 56s) * 14:29 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ssw1-a[1,8]-codfw with reason: lsw1-b7-codfw JunOS upgrade * 14:28 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 14:28 elukey@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 14:26 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:24 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] * 14:23 topranks: reboot lsw1-b7-codfw to upgrade JunOS (affects all hosts in rack) [[phab:T430928|T430928]] * 14:18 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2003.codfw.wmnet * 14:14 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2009.codfw.wmnet * 14:14 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Security update * 14:13 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2242: codfw rack B7 depool for maintenance * 14:13 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2242: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2228: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2228: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2229: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2229: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc2017: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.parsercache * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool pc2017: codfw rack B7 depool for maintenance * 14:08 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2003.codfw.wmnet * 14:07 cmooney@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on aux-k8s-etcd2004.codfw.wmnet,ml-etcd2001.codfw.wmnet with reason: lsw1-b7-codfw JunOS upgrade * 14:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95005 and previous config saved to /var/cache/conftool/dbconfig/20260721-140620-cwilliams.json * 14:05 cmooney@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2049.codfw.wmnet * 14:05 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-scholarly,name=eqiad * 14:04 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2009.codfw.wmnet * 14:04 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply * 14:04 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply * 14:03 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:03 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2001.codfw.wmnet * 14:03 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2001.codfw.wmnet * 14:02 cmooney@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2049.codfw.wmnet * 14:00 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:00 Dreamy_Jazz: Created cusi_case, cusi_signal, and cusi_user on svwiki, dewiki, jawiki, eswiki * 13:59 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b7-codfw,lsw1-b7-codfw IPv6,lsw1-b7-codfw.mgmt,ssw1-a[1,8]-codfw.mgmt with reason: lsw1-b7-codfw JunOS upgrade * 13:57 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs1023.eqiad.wmnet, repooling source-only afterwards * 13:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 29 hosts with reason: lsw1-b7-codfw JunOS upgrade * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224', diff saved to https://phabricator.wikimedia.org/P95003 and previous config saved to /var/cache/conftool/dbconfig/20260721-135613-cwilliams.json * 13:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1012.eqiad.wmnet with reason: host reimage * 13:53 cmooney@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 1:00:00 on 30 hosts with reason: lsw1-b7-codfw JunOS upgrade * 13:51 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1012.eqiad.wmnet with reason: host reimage * 13:48 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 13:48 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 13:46 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224', diff saved to https://phabricator.wikimedia.org/P95001 and previous config saved to /var/cache/conftool/dbconfig/20260721-134605-cwilliams.json * 13:46 cmooney@cumin1003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti2033.codfw.wmnet * 13:46 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 13:45 elukey: move the Docker Registry's /v2/wikimedia/machinelearning.* prefix to the ml S3 backend - [[phab:T428022|T428022]] * 13:45 cmooney@cumin1003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti2033.codfw.wmnet * 13:45 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 13:43 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:40 jiji@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 13:40 jiji@deploy2003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 13:39 jiji@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 13:39 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 13:38 cmooney@cumin1003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2032.codfw.wmnet * 13:38 jiji@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 13:38 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 13:37 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2032.codfw.wmnet * 13:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95000 and previous config saved to /var/cache/conftool/dbconfig/20260721-133557-cwilliams.json * 13:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1012 * 13:33 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1012 * 13:33 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1012.eqiad.wmnet with OS bookworm * 13:30 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:30 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:28 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94999 and previous config saved to /var/cache/conftool/dbconfig/20260721-132855-cwilliams.json * 13:28 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2224.codfw.wmnet with reason: Maintenance * 13:28 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94998 and previous config saved to /var/cache/conftool/dbconfig/20260721-132826-cwilliams.json * 13:28 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 13:23 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] (duration: 07m 50s) * 13:20 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 13:18 kharlan@deploy2003: kharlan: Continuing with deployment * 13:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217', diff saved to https://phabricator.wikimedia.org/P94996 and previous config saved to /var/cache/conftool/dbconfig/20260721-131817-cwilliams.json * 13:17 brouberol@dns1004: END - running authdns-update * 13:17 kharlan@deploy2003: kharlan: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:15 brouberol@dns1004: START - running authdns-update * 13:15 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] * 13:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94995 and previous config saved to /var/cache/conftool/dbconfig/20260721-131411-cwilliams.json * 13:13 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs1023.eqiad.wmnet, repooling source-only afterwards * 13:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217', diff saved to https://phabricator.wikimedia.org/P94994 and previous config saved to /var/cache/conftool/dbconfig/20260721-130809-cwilliams.json * 13:07 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 13:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180', diff saved to https://phabricator.wikimedia.org/P94993 and previous config saved to /var/cache/conftool/dbconfig/20260721-130404-cwilliams.json * 13:03 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:03 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:02 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 13:02 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 12:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94992 and previous config saved to /var/cache/conftool/dbconfig/20260721-125801-cwilliams.json * 12:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180', diff saved to https://phabricator.wikimedia.org/P94991 and previous config saved to /var/cache/conftool/dbconfig/20260721-125356-cwilliams.json * 12:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94990 and previous config saved to /var/cache/conftool/dbconfig/20260721-125049-cwilliams.json * 12:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2217.codfw.wmnet with reason: Maintenance * 12:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94989 and previous config saved to /var/cache/conftool/dbconfig/20260721-125017-cwilliams.json * 12:48 elukey: bmc cold reboot for lvs1013 and lvs1015 - [[phab:T426180|T426180]] * 12:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94988 and previous config saved to /var/cache/conftool/dbconfig/20260721-124348-cwilliams.json * 12:40 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193', diff saved to https://phabricator.wikimedia.org/P94987 and previous config saved to /var/cache/conftool/dbconfig/20260721-124009-cwilliams.json * 12:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts * 12:33 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts * 12:33 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts * 12:32 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts * 12:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:30 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193', diff saved to https://phabricator.wikimedia.org/P94986 and previous config saved to /var/cache/conftool/dbconfig/20260721-123001-cwilliams.json * 12:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94985 and previous config saved to /var/cache/conftool/dbconfig/20260721-121953-cwilliams.json * 12:17 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs2007.codfw.wmnet with OS bookworm * 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94983 and previous config saved to /var/cache/conftool/dbconfig/20260721-121257-cwilliams.json * 12:12 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2193.codfw.wmnet with reason: Maintenance * 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94982 and previous config saved to /var/cache/conftool/dbconfig/20260721-121239-cwilliams.json * 12:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180', diff saved to https://phabricator.wikimedia.org/P94980 and previous config saved to /var/cache/conftool/dbconfig/20260721-120231-cwilliams.json * 11:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180', diff saved to https://phabricator.wikimedia.org/P94979 and previous config saved to /var/cache/conftool/dbconfig/20260721-115223-cwilliams.json * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94978 and previous config saved to /var/cache/conftool/dbconfig/20260721-114333-cwilliams.json * 11:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1180.eqiad.wmnet with reason: Maintenance * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94977 and previous config saved to /var/cache/conftool/dbconfig/20260721-114305-cwilliams.json * 11:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94976 and previous config saved to /var/cache/conftool/dbconfig/20260721-114215-cwilliams.json * 11:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94975 and previous config saved to /var/cache/conftool/dbconfig/20260721-113530-cwilliams.json * 11:35 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2180.codfw.wmnet with reason: Maintenance * 11:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94974 and previous config saved to /var/cache/conftool/dbconfig/20260721-113501-cwilliams.json * 11:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168', diff saved to https://phabricator.wikimedia.org/P94973 and previous config saved to /var/cache/conftool/dbconfig/20260721-113258-cwilliams.json * 11:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169', diff saved to https://phabricator.wikimedia.org/P94972 and previous config saved to /var/cache/conftool/dbconfig/20260721-112453-cwilliams.json * 11:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168', diff saved to https://phabricator.wikimedia.org/P94971 and previous config saved to /var/cache/conftool/dbconfig/20260721-112250-cwilliams.json * 11:21 XioNoX: put eqiad-drmrs Arelion link in service * 11:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169', diff saved to https://phabricator.wikimedia.org/P94970 and previous config saved to /var/cache/conftool/dbconfig/20260721-111446-cwilliams.json * 11:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94969 and previous config saved to /var/cache/conftool/dbconfig/20260721-111242-cwilliams.json * 11:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 11:10 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 11:07 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1093 hosts * 11:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94968 and previous config saved to /var/cache/conftool/dbconfig/20260721-110548-cwilliams.json * 11:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1168.eqiad.wmnet with reason: Maintenance * 11:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94967 and previous config saved to /var/cache/conftool/dbconfig/20260721-110520-cwilliams.json * 11:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94966 and previous config saved to /var/cache/conftool/dbconfig/20260721-110439-cwilliams.json * 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94964 and previous config saved to /var/cache/conftool/dbconfig/20260721-105632-cwilliams.json * 10:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2169.codfw.wmnet with reason: Maintenance * 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94963 and previous config saved to /var/cache/conftool/dbconfig/20260721-105603-cwilliams.json * 10:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165', diff saved to https://phabricator.wikimedia.org/P94962 and previous config saved to /var/cache/conftool/dbconfig/20260721-105512-cwilliams.json * 10:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158', diff saved to https://phabricator.wikimedia.org/P94961 and previous config saved to /var/cache/conftool/dbconfig/20260721-104555-cwilliams.json * 10:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165', diff saved to https://phabricator.wikimedia.org/P94960 and previous config saved to /var/cache/conftool/dbconfig/20260721-104504-cwilliams.json * 10:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158', diff saved to https://phabricator.wikimedia.org/P94959 and previous config saved to /var/cache/conftool/dbconfig/20260721-103547-cwilliams.json * 10:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94958 and previous config saved to /var/cache/conftool/dbconfig/20260721-103456-cwilliams.json * 10:29 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2229: Upgraded kernel * 10:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94956 and previous config saved to /var/cache/conftool/dbconfig/20260721-102757-cwilliams.json * 10:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on an-redacteddb1001.eqiad.wmnet,clouddb[1015,1025,1028].eqiad.wmnet,db1155.eqiad.wmnet with reason: Maintenance * 10:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1165.eqiad.wmnet with reason: Maintenance * 10:25 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94955 and previous config saved to /var/cache/conftool/dbconfig/20260721-102539-cwilliams.json * 10:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94954 and previous config saved to /var/cache/conftool/dbconfig/20260721-101848-cwilliams.json * 10:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2158.codfw.wmnet with reason: Maintenance * 09:43 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2229: Upgraded kernel * 09:42 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2229.codfw.wmnet * 09:42 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2229.codfw.wmnet * 09:23 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db2229.codfw.wmnet * 09:23 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2229.codfw.wmnet * 08:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2229 [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94948 and previous config saved to /var/cache/conftool/dbconfig/20260721-085724-cwilliams.json * 08:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2214 to s6 primary [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94947 and previous config saved to /var/cache/conftool/dbconfig/20260721-085442-cwilliams.json * 08:53 cezmunsta: Starting s6 codfw failover from db2229 to db2214 - [[phab:T430964|T430964]] * 08:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2214 with weight 0 [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94946 and previous config saved to /var/cache/conftool/dbconfig/20260721-084613-cwilliams.json * 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 22 hosts with reason: Primary switchover s6 [[phab:T430964|T430964]] * 08:32 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1017.eqiad.wmnet,service=s1 * 08:08 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Add subrated circuit rate to interface descriptions - CR1312476 - ayounsi@cumin1003 * 08:06 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Add subrated circuit rate to interface descriptions - CR1312476 - ayounsi@cumin1003 * 07:58 reedy@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] (duration: 12m 55s) * 07:51 reedy@deploy2003: reedy, neriah: Continuing with deployment * 07:51 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 07:51 reedy@deploy2003: reedy, neriah: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:48 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1093 hosts * 07:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm2001.wikimedia.org * 07:45 reedy@deploy2003: Started scap sync-world: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] * 07:43 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 07:42 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm2001.wikimedia.org * 07:23 elukey: upgrade libtiff6 packages on zuul* trixie hosts for security upgrades * 07:22 elukey: upgrade libtiff6 packages on Wikikube trixie workers for security upgrades * 07:14 elukey@deploy2003: helmfile [codfw] DONE helmfile.d/services/proton: sync * 07:13 elukey@deploy2003: helmfile [codfw] START helmfile.d/services/proton: sync * 07:11 elukey@deploy2003: helmfile [eqiad] DONE helmfile.d/services/proton: sync * 07:10 elukey@deploy2003: helmfile [eqiad] START helmfile.d/services/proton: sync * 07:09 elukey@deploy2003: helmfile [staging] DONE helmfile.d/services/proton: sync * 07:08 elukey@deploy2003: helmfile [staging] START helmfile.d/services/proton: sync * 06:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1023.eqiad.wmnet with reason: host reimage * 06:46 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1023.eqiad.wmnet with reason: host reimage * 06:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 05:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Haproxy-only mode support - oblivian@cumin1003" * 05:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Haproxy-only mode support - oblivian@cumin1003 * 05:42 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Haproxy-only mode support - oblivian@cumin1003 * 05:42 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Haproxy-only mode support - oblivian@cumin1003" * 05:38 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1017.eqiad.wmnet with reason: Cloning * 05:37 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1017.eqiad.wmnet,service=s1 * 05:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1029.eqiad.wmnet,service=s8 * 05:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1029.eqiad.wmnet,service=s5 * 05:32 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:30 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 05:11 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet * 05:04 aokoth@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet * 05:00 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 04:56 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 04:01 mwpresync@deploy2003: Pruned MediaWiki: 1.47.0-wmf.9 (duration: 01m 08s) * 03:41 mwpresync@deploy2003: Finished scap sync-world: testwikis to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] (duration: 36m 30s) * 03:05 mwpresync@deploy2003: Started scap sync-world: testwikis to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 03:01 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:01 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:00 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:00 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:36 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:36 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:36 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:35 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:16 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 47s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 00:56 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm == 2026-07-20 == * 23:38 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 23:07 Amir1: deleting echo notifications from 2015 on group1 wikis * 23:07 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] (duration: 14m 16s) * 23:01 ladsgroup@deploy2003: ladsgroup: Continuing with deployment * 23:00 ladsgroup@deploy2003: ladsgroup: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:53 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] * 22:46 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2007.codfw.wmnet, repooling source-only afterwards * 22:39 maryum: Deployed security fixes for several security bugs * 21:42 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 21:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2007.codfw.wmnet, repooling source-only afterwards * 21:37 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 21:37 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 21:34 sbassett: Deployed security fix for [[phab:T432424|T432424]] * 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs2020.codfw.wmnet * 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1023.eqiad.wmnet * 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1011.eqiad.wmnet * 21:32 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 17s) * 21:32 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 21:27 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 21:13 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2007.codfw.wmnet with reason: host reimage * 21:08 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-internal-main,name=codfw * 21:06 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2007.codfw.wmnet with reason: host reimage * 20:59 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service * 20:58 sukhe: pybal restart for IP changes around wdqs-main hosts * 20:57 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 20:46 ryankemper@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-internal-main,name=codfw * 20:45 ebernhardson@deploy2003: Finished deploy [search/mjolnir/deploy@d4dc3b8]: Update for opensearch 2.x compat (duration: 00m 34s) * 20:44 ebernhardson@deploy2003: Started deploy [search/mjolnir/deploy@d4dc3b8]: Update for opensearch 2.x compat * 20:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2007 * 20:44 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2007 * 20:43 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2007 * 20:43 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2007.codfw.wmnet 156.16.192.10.in-addr.arpa 6.5.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:42 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2007.codfw.wmnet 156.16.192.10.in-addr.arpa 6.5.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:42 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2007 - bking@cumin2003" * 20:41 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2007 - bking@cumin2003" * 20:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94944 and previous config saved to /var/cache/conftool/dbconfig/20260720-203333-cwilliams.json * 20:32 arlolra@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] (duration: 15m 07s) * 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2020.codfw.wmnet * 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1023.eqiad.wmnet * 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1011.eqiad.wmnet * 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs2020.codfw.wmnet * 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1023.eqiad.wmnet * 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1011.eqiad.wmnet * 20:25 arlolra@deploy2003: arlolra, cscott: Continuing with deployment * 20:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257', diff saved to https://phabricator.wikimedia.org/P94943 and previous config saved to /var/cache/conftool/dbconfig/20260720-202325-cwilliams.json * 20:21 arlolra@deploy2003: arlolra, cscott: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:17 arlolra@deploy2003: Started scap sync-world: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] * 20:13 bking@cumin2003: START - Cookbook sre.dns.netbox * 20:13 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257', diff saved to https://phabricator.wikimedia.org/P94942 and previous config saved to /var/cache/conftool/dbconfig/20260720-201318-cwilliams.json * 20:13 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 20:10 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 20:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2007 * 20:04 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2007.codfw.wmnet with OS bookworm * 20:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94941 and previous config saved to /var/cache/conftool/dbconfig/20260720-200310-cwilliams.json * 19:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94940 and previous config saved to /var/cache/conftool/dbconfig/20260720-195633-cwilliams.json * 19:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1257.eqiad.wmnet with reason: Maintenance * 19:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94939 and previous config saved to /var/cache/conftool/dbconfig/20260720-195605-cwilliams.json * 19:51 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 19:50 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 19:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256', diff saved to https://phabricator.wikimedia.org/P94938 and previous config saved to /var/cache/conftool/dbconfig/20260720-194558-cwilliams.json * 19:44 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 19:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 19:41 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wdqs1011.eqiad.wmnet with OS bookworm * 19:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256', diff saved to https://phabricator.wikimedia.org/P94937 and previous config saved to /var/cache/conftool/dbconfig/20260720-193550-cwilliams.json * 19:25 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94936 and previous config saved to /var/cache/conftool/dbconfig/20260720-192542-cwilliams.json * 19:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94935 and previous config saved to /var/cache/conftool/dbconfig/20260720-191856-cwilliams.json * 19:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1256.eqiad.wmnet with reason: Maintenance * 19:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94934 and previous config saved to /var/cache/conftool/dbconfig/20260720-191839-cwilliams.json * 19:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255', diff saved to https://phabricator.wikimedia.org/P94933 and previous config saved to /var/cache/conftool/dbconfig/20260720-190831-cwilliams.json * 18:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255', diff saved to https://phabricator.wikimedia.org/P94932 and previous config saved to /var/cache/conftool/dbconfig/20260720-185824-cwilliams.json * 18:50 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 18:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94931 and previous config saved to /var/cache/conftool/dbconfig/20260720-184816-cwilliams.json * 18:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94930 and previous config saved to /var/cache/conftool/dbconfig/20260720-184224-cwilliams.json * 18:42 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1255.eqiad.wmnet with reason: Maintenance * 18:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94929 and previous config saved to /var/cache/conftool/dbconfig/20260720-184153-cwilliams.json * 18:39 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 18:39 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 16s) * 18:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 18:38 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 59m 26s) * 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 18:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211', diff saved to https://phabricator.wikimedia.org/P94928 and previous config saved to /var/cache/conftool/dbconfig/20260720-183145-cwilliams.json * 18:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211', diff saved to https://phabricator.wikimedia.org/P94927 and previous config saved to /var/cache/conftool/dbconfig/20260720-182137-cwilliams.json * 18:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94926 and previous config saved to /var/cache/conftool/dbconfig/20260720-181129-cwilliams.json * 18:09 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_codfw * 18:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2057.codfw.wmnet * 18:08 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_codfw * 18:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2058.codfw.wmnet * 18:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94925 and previous config saved to /var/cache/conftool/dbconfig/20260720-180452-cwilliams.json * 18:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on clouddb[1016,1020,1022-1023].eqiad.wmnet,db1154.eqiad.wmnet with reason: Maintenance * 18:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1211.eqiad.wmnet with reason: Maintenance * 18:02 sukhe: armed keyholder on acmechief1002.eqiad.wmnet and acmechief2002.codfw.wmnet (active host) * 18:01 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief2002.codfw.wmnet * 17:57 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief2002.codfw.wmnet * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs2020'] * 17:52 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief1002.eqiad.wmnet * 17:50 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 17:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1011.eqiad.wmnet with reason: host reimage * 17:48 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief1002.eqiad.wmnet * 17:47 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test2001.codfw.wmnet * 17:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94924 and previous config saved to /var/cache/conftool/dbconfig/20260720-174717-cwilliams.json * 17:46 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 17:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1011.eqiad.wmnet with reason: host reimage * 17:43 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs2020.codfw.wmnet with OS bookworm * 17:43 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test2001.codfw.wmnet * 17:43 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test1001.eqiad.wmnet * 17:39 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test1001.eqiad.wmnet * 17:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 17:38 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:38 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 17:37 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 17:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244', diff saved to https://phabricator.wikimedia.org/P94923 and previous config saved to /var/cache/conftool/dbconfig/20260720-173709-cwilliams.json * 17:35 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 17:31 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 20m 40s) * 17:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2055.codfw.wmnet * 17:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2056.codfw.wmnet * 17:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1011 * 17:27 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1011 * 17:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1011.eqiad.wmnet with OS bookworm * 17:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244', diff saved to https://phabricator.wikimedia.org/P94922 and previous config saved to /var/cache/conftool/dbconfig/20260720-172701-cwilliams.json * 17:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94921 and previous config saved to /var/cache/conftool/dbconfig/20260720-171653-cwilliams.json * 17:11 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 17:11 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 13m 03s) * 17:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94920 and previous config saved to /var/cache/conftool/dbconfig/20260720-171012-cwilliams.json * 17:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2244.codfw.wmnet with reason: Maintenance * 17:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94919 and previous config saved to /var/cache/conftool/dbconfig/20260720-170941-cwilliams.json * 16:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243', diff saved to https://phabricator.wikimedia.org/P94918 and previous config saved to /var/cache/conftool/dbconfig/20260720-165933-cwilliams.json * 16:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 16:58 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2053.codfw.wmnet * 16:51 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2054.codfw.wmnet * 16:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243', diff saved to https://phabricator.wikimedia.org/P94917 and previous config saved to /var/cache/conftool/dbconfig/20260720-164926-cwilliams.json * 16:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94916 and previous config saved to /var/cache/conftool/dbconfig/20260720-163918-cwilliams.json * 16:35 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 16:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94915 and previous config saved to /var/cache/conftool/dbconfig/20260720-163140-cwilliams.json * 16:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2243.codfw.wmnet with reason: Maintenance * 16:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94914 and previous config saved to /var/cache/conftool/dbconfig/20260720-163111-cwilliams.json * 16:27 btullis@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 16:27 btullis@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 16:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2020 * 16:23 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2020 * 16:21 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2020 * 16:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2020.codfw.wmnet 85.0.192.10.in-addr.arpa 5.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:21 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2020.codfw.wmnet 85.0.192.10.in-addr.arpa 5.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242', diff saved to https://phabricator.wikimedia.org/P94913 and previous config saved to /var/cache/conftool/dbconfig/20260720-162103-cwilliams.json * 16:19 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 16:18 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 16:18 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:18 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 16:17 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:17 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for netbox accounting errors - jhancock@cumin2002" * 16:17 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for netbox accounting errors - jhancock@cumin2002" * 16:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2051.codfw.wmnet * 16:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2052.codfw.wmnet * 16:11 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 16:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242', diff saved to https://phabricator.wikimedia.org/P94912 and previous config saved to /var/cache/conftool/dbconfig/20260720-161055-cwilliams.json * 16:09 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 16:08 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 16:06 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 16:06 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 16:06 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2020 * 16:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2020.codfw.wmnet with OS bookworm * 16:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94911 and previous config saved to /var/cache/conftool/dbconfig/20260720-160047-cwilliams.json * 15:58 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2019.codfw.wmnet, repooling source-only afterwards * 15:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94909 and previous config saved to /var/cache/conftool/dbconfig/20260720-155353-cwilliams.json * 15:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2242.codfw.wmnet with reason: Maintenance * 15:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94908 and previous config saved to /var/cache/conftool/dbconfig/20260720-154433-cwilliams.json * 15:35 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2049.codfw.wmnet * 15:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162', diff saved to https://phabricator.wikimedia.org/P94907 and previous config saved to /var/cache/conftool/dbconfig/20260720-153425-cwilliams.json * 15:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2050.codfw.wmnet * 15:28 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162', diff saved to https://phabricator.wikimedia.org/P94906 and previous config saved to /var/cache/conftool/dbconfig/20260720-152418-cwilliams.json * 15:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1023 * 15:14 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1023 * 15:14 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 15:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94905 and previous config saved to /var/cache/conftool/dbconfig/20260720-151407-cwilliams.json * 15:13 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] (duration: 41m 16s) * 15:08 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 15:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94902 and previous config saved to /var/cache/conftool/dbconfig/20260720-150729-cwilliams.json * 15:07 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2162.codfw.wmnet with reason: Maintenance * 15:05 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2027.codfw.wmnet, repooling source-only afterwards * 15:00 urbanecm@deploy2003: vadymts1, migr, urbanecm: Continuing with deployment * 14:59 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:58 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 07s) * 14:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:58 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 13s) * 14:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:57 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2019.codfw.wmnet, repooling source-only afterwards * 14:57 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2047.codfw.wmnet * 14:55 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2048.codfw.wmnet * 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2019.codfw.wmnet with OS bookworm * 14:49 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 14:47 urbanecm@deploy2003: vadymts1, migr, urbanecm: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:44 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool magru [reason: BGP issues in lvs7003 resolved after liberica restart, no task ID specified] * 14:44 sukhe@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool magru [reason: BGP issues in lvs7003 resolved after liberica restart, no task ID specified] * 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:41 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:39 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:39 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:33 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool magru [reason: no reason specified, no task ID specified] * 14:33 sukhe@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool magru [reason: no reason specified, no task ID specified] * 14:31 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] * 14:24 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:24 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:24 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:24 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2019.codfw.wmnet with reason: host reimage * 14:22 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2027.codfw.wmnet, repooling source-only afterwards * 14:19 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2019.codfw.wmnet with reason: host reimage * 14:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2027.codfw.wmnet with OS bookworm * 14:16 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2046.codfw.wmnet * 14:16 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2045.codfw.wmnet * 14:08 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:08 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:08 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:08 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:07 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:06 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:06 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:06 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:05 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2019 * 14:00 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2019 * 13:56 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1015.eqiad.wmnet * 13:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2027.codfw.wmnet with reason: host reimage * 13:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2071.codfw.wmnet with OS trixie * 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:51 sukhe@cumin1003: END (ERROR) - Cookbook sre.loadbalancer.admin (exit_code=97) rebooting A:liberica and P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica and P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:51 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1015.eqiad.wmnet * 13:50 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1014.eqiad.wmnet * 13:50 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1076.eqiad.wmnet with OS trixie * 13:50 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2027.codfw.wmnet with reason: host reimage * 13:45 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1014.eqiad.wmnet * 13:44 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1013.eqiad.wmnet * 13:39 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 13:38 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1013.eqiad.wmnet * 13:37 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2044.codfw.wmnet * 13:37 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2043.codfw.wmnet * 13:36 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2019 * 13:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2019.codfw.wmnet 156.32.192.10.in-addr.arpa 6.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:36 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2019.codfw.wmnet 156.32.192.10.in-addr.arpa 6.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:36 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2019 - bking@cumin2003" * 13:36 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2019 - bking@cumin2003" * 13:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry2005.codfw.wmnet * 13:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2071.codfw.wmnet with reason: host reimage * 13:31 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:31 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2019 * 13:31 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry2005.codfw.wmnet * 13:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry2004.codfw.wmnet * 13:30 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2019.codfw.wmnet with OS bookworm * 13:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2027 * 13:30 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2027 * 13:30 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2027.codfw.wmnet with OS bookworm * 13:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1076.eqiad.wmnet with reason: host reimage * 13:29 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_codfw * 13:28 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_codfw * 13:26 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry2004.codfw.wmnet * 13:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry1005.eqiad.wmnet * 13:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2071.codfw.wmnet with reason: host reimage * 13:22 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1076.eqiad.wmnet with reason: host reimage * 13:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry1005.eqiad.wmnet * 13:21 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry1004.eqiad.wmnet * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry1004.eqiad.wmnet * 13:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts * 13:13 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts * 13:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts * 13:12 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts * 13:03 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1076.eqiad.wmnet with OS trixie * 13:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2071.codfw.wmnet with OS trixie * 12:55 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:54 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:53 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:46 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 7 hosts * 12:42 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 7 hosts * 12:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts * 12:42 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts * 12:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2070.codfw.wmnet with OS trixie * 12:36 ozge@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:35 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1075.eqiad.wmnet with OS trixie * 12:32 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts * 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts * 12:22 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin1001.eqiad.wmnet * 12:19 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin1001.eqiad.wmnet * 12:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2070.codfw.wmnet with reason: host reimage * 12:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin2001.codfw.wmnet * 12:14 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1075.eqiad.wmnet with reason: host reimage * 12:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2070.codfw.wmnet with reason: host reimage * 12:10 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1075.eqiad.wmnet with reason: host reimage * 12:09 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin2001.codfw.wmnet * 11:17 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1074.eqiad.wmnet with OS trixie * 11:17 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 11:16 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 11:14 ozge@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 6 hosts * 11:09 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 6 hosts * 11:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 324 hosts * 10:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1074.eqiad.wmnet with reason: host reimage * 10:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2069.codfw.wmnet with OS trixie * 10:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1074.eqiad.wmnet with reason: host reimage * 10:30 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2069.codfw.wmnet with reason: host reimage * 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1074.eqiad.wmnet with OS trixie * 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2069.codfw.wmnet with reason: host reimage * 10:06 blake@deploy2003: Stopping before sync operations * 10:06 blake@deploy2003: Started scap sync-world: Non-deployment scap run to populate new release values * 10:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2069.codfw.wmnet with OS trixie * 10:00 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1073.eqiad.wmnet with OS trixie * 09:56 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 324 hosts * 09:39 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 09:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 8 hosts * 09:38 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1073.eqiad.wmnet with reason: host reimage * 09:37 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 8 hosts * 09:34 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1073.eqiad.wmnet with reason: host reimage * 09:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2068.codfw.wmnet with OS trixie * 09:16 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1073.eqiad.wmnet with OS trixie * 09:13 blake@deploy2003: sync-world aborted: Non-deployment scap run to populate new release values (duration: 00m 02s) * 09:13 blake@deploy2003: Started scap sync-world: Non-deployment scap run to populate new release values * 08:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2068.codfw.wmnet with reason: host reimage * 08:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2068.codfw.wmnet with reason: host reimage * 08:50 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 08:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2068.codfw.wmnet with OS trixie * 08:15 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1072.eqiad.wmnet with OS trixie * 07:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2067.codfw.wmnet with OS trixie * 07:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1072.eqiad.wmnet with reason: host reimage * 07:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1072.eqiad.wmnet with reason: host reimage * 07:45 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 07:45 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 07:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2067.codfw.wmnet with reason: host reimage * 07:35 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2067.codfw.wmnet with reason: host reimage * 07:30 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 07:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1072.eqiad.wmnet with OS trixie * 07:30 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 07:17 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 07:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2067.codfw.wmnet with OS trixie * 05:51 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:50 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:25 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:25 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on db2207.codfw.wmnet with reason: Host down * 04:28 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 07m 02s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-18 == * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 29s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 00:11 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2018.codfw.wmnet, repooling source-only afterwards == 2026-07-17 == * 23:53 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2026.codfw.wmnet, repooling source-only afterwards * 23:09 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2018.codfw.wmnet, repooling source-only afterwards * 23:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2026.codfw.wmnet, repooling source-only afterwards * 22:11 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2018.codfw.wmnet with OS bookworm * 22:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2026.codfw.wmnet with OS bookworm * 21:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2018.codfw.wmnet with reason: host reimage * 21:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2018.codfw.wmnet with reason: host reimage * 21:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2026.codfw.wmnet with reason: host reimage * 21:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2026.codfw.wmnet with reason: host reimage * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2018 * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2018 * 21:26 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2018 * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2018.codfw.wmnet 155.32.192.10.in-addr.arpa 5.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:26 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2018.codfw.wmnet 155.32.192.10.in-addr.arpa 5.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2018 - bking@cumin2003" * 21:26 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2018 - bking@cumin2003" * 21:14 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:13 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2018 * 21:13 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2018.codfw.wmnet with OS bookworm * 21:12 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2026 * 21:12 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2026 * 21:12 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2026.codfw.wmnet with OS bookworm * 21:05 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1022.eqiad.wmnet -> wdqs1026.eqiad.wmnet, repooling source-only afterwards * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs2017.codfw.wmnet, repooling source-only afterwards * 20:11 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs2017.codfw.wmnet, repooling source-only afterwards * 20:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2017.codfw.wmnet with OS bookworm * 20:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1022.eqiad.wmnet -> wdqs1026.eqiad.wmnet, repooling source-only afterwards * 20:06 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1026.eqiad.wmnet with OS bookworm * 19:55 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 19:55 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:55 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 09s) * 19:55 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:50 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 08s) * 19:50 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:50 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 10m 03s) * 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2017.codfw.wmnet with reason: host reimage * 19:40 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:40 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 15s) * 19:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1026.eqiad.wmnet with reason: host reimage * 19:37 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 16s) * 19:37 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:34 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2017.codfw.wmnet with reason: host reimage * 19:34 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1026.eqiad.wmnet with reason: host reimage * 19:33 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 19:33 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:16 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1026.eqiad.wmnet with OS bookworm * 19:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2017 * 19:16 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2017 * 19:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2017.codfw.wmnet with OS bookworm * 18:30 bking@dns1004: END - running authdns-update * 18:28 bking@dns1004: START - running authdns-update * 18:16 kamila@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1264.eqiad.wmnet * 18:16 kamila@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1264.eqiad.wmnet * 18:16 kamila@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1264.eqiad.wmnet * 17:49 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 17:46 dzahn@dns1006: END - running authdns-update * 17:44 dzahn@dns1006: START - running authdns-update * 17:44 dzahn@dns1006: END - running authdns-update * 17:42 dzahn@dns1006: START - running authdns-update * 17:28 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 17:21 kamila@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 17:01 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1264 * 17:01 kamila@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1264 * 17:01 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 17:01 kamila@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1264.eqiad.wmnet * 17:01 kamila@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1264.eqiad.wmnet * 17:01 kamila@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1264.eqiad.wmnet * 16:42 reedy@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] (duration: 10m 29s) * 16:34 reedy@deploy2003: reedy, hartman: Continuing with deployment * 16:33 reedy@deploy2003: reedy, hartman: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:31 reedy@deploy2003: Started scap sync-world: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] * 16:26 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 16:10 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in2001.wikimedia.org with reason: [[phab:T431659|T431659]] * 16:07 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in1001.wikimedia.org with reason: [[phab:T431659|T431659]] * 16:05 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 16:01 kamila@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 16:00 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out2001.wikimedia.org with reason: [[phab:T431659|T431659]] * 15:41 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 15:41 kamila@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 15:35 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out1001.wikimedia.org with reason: [[phab:T431659|T431659]] * 15:14 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1339.eqiad.wmnet * 15:13 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1339.eqiad.wmnet * 15:13 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1339.eqiad.wmnet * 14:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1339.eqiad.wmnet with OS trixie * 14:50 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:49 kamila@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:49 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:33 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage * 14:27 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage * 14:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1339 * 14:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1339 * 14:14 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1339 * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1339.eqiad.wmnet 156.32.64.10.in-addr.arpa 6.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:14 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1339.eqiad.wmnet 156.32.64.10.in-addr.arpa 6.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1339 - cgoubert@cumin2003" * 14:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1339 - cgoubert@cumin2003" * 14:09 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 14:06 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1339 * 14:06 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie * 14:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1339.eqiad.wmnet * 14:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1339.eqiad.wmnet * 14:02 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1339.eqiad.wmnet * 13:45 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb1013.eqiad.wmnet * 13:39 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb1013.eqiad.wmnet * 13:27 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:24 blake@dns1004: END - running authdns-update * 13:22 blake@dns1004: START - running authdns-update * 13:20 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 13:11 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2014.codfw.wmnet * 13:06 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb2014.codfw.wmnet * 13:06 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2012.codfw.wmnet * 13:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 13:03 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 13:01 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 15 hosts * 13:01 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb2012.codfw.wmnet * 13:01 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1016.eqiad.wmnet * 13:00 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 15 hosts * 12:55 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb1016.eqiad.wmnet * 12:55 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1014.eqiad.wmnet * 12:49 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb1014.eqiad.wmnet * 12:32 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:32 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:31 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:31 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1338.eqiad.wmnet * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1338.eqiad.wmnet * 12:18 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1338.eqiad.wmnet * 12:17 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:15 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:14 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:13 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1338.eqiad.wmnet with OS trixie * 12:01 klausman@deploy2003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 11:59 klausman@deploy2003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 11:56 klausman@deploy2003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 11:54 klausman@deploy2003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 11:53 klausman@deploy2003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 11:51 klausman@deploy2003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 11:42 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1338.eqiad.wmnet with reason: host reimage * 11:38 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1338.eqiad.wmnet with reason: host reimage * 11:31 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2230.codfw.wmnet * 11:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1338 * 11:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1338 * 11:25 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1338 * 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1338.eqiad.wmnet 155.32.64.10.in-addr.arpa 5.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:25 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1338.eqiad.wmnet 155.32.64.10.in-addr.arpa 5.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1338 - cgoubert@cumin2003" * 11:25 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1338 - cgoubert@cumin2003" * 11:23 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2230.codfw.wmnet * 11:20 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 11:20 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1338 * 11:20 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1338.eqiad.wmnet with OS trixie * 11:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1338.eqiad.wmnet * 11:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1338.eqiad.wmnet * 11:19 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1338.eqiad.wmnet * 11:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1337.eqiad.wmnet * 11:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1337.eqiad.wmnet * 11:17 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1337.eqiad.wmnet * 11:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1337.eqiad.wmnet with OS trixie * 10:51 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[2001-2002].codfw.wmnet * 10:50 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1337.eqiad.wmnet with reason: host reimage * 10:40 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 10:39 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:39 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1337.eqiad.wmnet with reason: host reimage * 10:39 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1001-1003].eqiad.wmnet * 10:34 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:34 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:30 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:28 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1001-1003].eqiad.wmnet * 10:27 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1337 * 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1337 * 10:26 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1337 * 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1337.eqiad.wmnet 154.32.64.10.in-addr.arpa 4.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:26 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1337.eqiad.wmnet 154.32.64.10.in-addr.arpa 4.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1337 - cgoubert@cumin2003" * 10:26 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1337 - cgoubert@cumin2003" * 10:21 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 10:18 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1337 * 10:17 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1337.eqiad.wmnet with OS trixie * 10:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1337.eqiad.wmnet * 10:16 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db1176.eqiad.wmnet * 10:16 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1337.eqiad.wmnet * 10:16 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1337.eqiad.wmnet * 10:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1336.eqiad.wmnet * 10:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1336.eqiad.wmnet * 10:15 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1336.eqiad.wmnet * 10:11 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db1176.eqiad.wmnet * 10:10 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db1176.eqiad.wmnet * 10:09 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db1176.eqiad.wmnet * 10:05 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts (check the cookbook's logs for more details.) * 10:03 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts (check the cookbook's logs for more details.) * 09:58 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1336.eqiad.wmnet with OS trixie * 09:47 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts (check the cookbook's logs for more details.) * 09:47 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts (check the cookbook's logs for more details.) * 09:45 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host acmechief-test2001.codfw.wmnet,acmechief-test1001.eqiad.wmnet,an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet,db-test[2001-2002].codfw.wmnet,db-test[1 * 09:40 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host acmechief-test2001.codfw.wmnet,acmechief-test1001.eqiad.wmnet,an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet,db-test[2001-2002].codfw.wmnet,db-test[1001-1003].eqiad.wmn * 09:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1336.eqiad.wmnet with reason: host reimage * 09:33 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1336.eqiad.wmnet with reason: host reimage * 09:29 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet * 09:29 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet * 09:28 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:26 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:21 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 09:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1336 * 09:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1336 * 09:19 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 09:14 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1336 * 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1336.eqiad.wmnet 152.32.64.10.in-addr.arpa 2.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1336.eqiad.wmnet 152.32.64.10.in-addr.arpa 2.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1336 - cgoubert@cumin2003" * 09:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1336 - cgoubert@cumin2003" * 09:11 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:10 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 09:09 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:09 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1336 * 09:09 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1336.eqiad.wmnet with OS trixie * 09:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1336.eqiad.wmnet * 09:08 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1336.eqiad.wmnet * 09:08 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1336.eqiad.wmnet * 09:06 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1335.eqiad.wmnet * 09:06 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1335.eqiad.wmnet * 09:06 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1335.eqiad.wmnet * 09:04 elukey: uploaded spicerack_13.1.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia * 08:55 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wikikube-worker-exp2001.codfw.wmnet * 08:54 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host testreduce1002.eqiad.wmnet * 08:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1335.eqiad.wmnet with OS trixie * 08:51 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host wikikube-worker-exp2001.codfw.wmnet * 08:51 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wikikube-worker-exp1001.eqiad.wmnet * 08:50 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host testreduce1002.eqiad.wmnet * 08:45 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host wikikube-worker-exp1001.eqiad.wmnet * 08:34 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1335.eqiad.wmnet with reason: host reimage * 08:30 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1335.eqiad.wmnet with reason: host reimage * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1335 * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1335 * 08:18 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1335 * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1335.eqiad.wmnet 150.32.64.10.in-addr.arpa 0.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:18 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1335.eqiad.wmnet 150.32.64.10.in-addr.arpa 0.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1335 - cgoubert@cumin2003" * 08:18 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1335 - cgoubert@cumin2003" * 08:14 elukey@cumin1003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 08:14 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges * 08:13 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 08:10 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1335 * 08:10 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1335.eqiad.wmnet with OS trixie * 08:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1335.eqiad.wmnet * 08:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1335.eqiad.wmnet * 08:09 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1335.eqiad.wmnet * 08:06 elukey@cumin1003: END (FAIL) - Cookbook sre.puppet.disable-merges (exit_code=99) * 08:05 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges * 08:03 elukey@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin1003.eqiad.wmnet * 07:57 elukey@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin1003.eqiad.wmnet * 07:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetdb1003.eqiad.wmnet * 07:46 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetdb1003.eqiad.wmnet * 07:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetdb2003.codfw.wmnet * 07:37 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetdb2003.codfw.wmnet * 07:37 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1001.eqiad.wmnet * 07:28 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver1001.eqiad.wmnet * 07:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet * 07:19 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet * 07:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2002.codfw.wmnet * 07:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver2002.codfw.wmnet * 07:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet * 07:05 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet * 07:04 btullis@cumin1003: END (FAIL) - Cookbook sre.hadoop.reboot-workers (exit_code=99) for Hadoop analytics cluster * 07:04 elukey@cumin1003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 07:04 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges * 06:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox1003.eqiad.wmnet * 06:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox1003.eqiad.wmnet * 02:46 ryankemper@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:46 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:44 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:37 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:37 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-internal-scholarly,name=eqiad * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 49s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 01:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore wdqs1025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling source-only afterwards * 01:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore wdqs1027 after Bookworm reimage) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs1027.eqiad.wmnet, repooling both afterwards * 00:55 urbanecm@deploy2003: helmfile [codfw] DONE helmfile.d/services/linkrecommendation: apply * 00:54 urbanecm@deploy2003: helmfile [eqiad] DONE helmfile.d/services/linkrecommendation: apply * 00:54 urbanecm@deploy2003: helmfile [staging] DONE helmfile.d/services/linkrecommendation: apply * 00:54 urbanecm@deploy2003: helmfile [codfw] START helmfile.d/services/linkrecommendation: apply * 00:53 urbanecm@deploy2003: helmfile [staging] START helmfile.d/services/linkrecommendation: apply * 00:52 urbanecm@deploy2003: helmfile [eqiad] START helmfile.d/services/linkrecommendation: apply * 00:23 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore wdqs1025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling source-only afterwards * 00:23 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore wdqs1027 after Bookworm reimage) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs1027.eqiad.wmnet, repooling both afterwards * 00:14 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1274.eqiad.wmnet * 00:14 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1274.eqiad.wmnet * 00:14 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1274.eqiad.wmnet * 00:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1027.eqiad.wmnet with OS bookworm * 00:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1025.eqiad.wmnet with OS bookworm * 00:04 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1274.eqiad.wmnet with OS trixie == 2026-07-16 == * 23:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], xfer to freshly reimaged/scap-deployed wdqs2025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs2025.codfw.wmnet, repooling source-only afterwards * 23:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1027.eqiad.wmnet with reason: host reimage * 23:47 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1025.eqiad.wmnet with reason: host reimage * 23:43 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1274.eqiad.wmnet with reason: host reimage * 23:41 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1025.eqiad.wmnet with reason: host reimage * 23:39 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1027.eqiad.wmnet with reason: host reimage * 23:38 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1274.eqiad.wmnet with reason: host reimage * 23:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1025 * 23:23 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1025 * 23:22 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1027 * 23:22 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1027 * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1274 * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1274 * 23:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1025.eqiad.wmnet with OS bookworm * 23:19 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1274 * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1274.eqiad.wmnet 145.48.64.10.in-addr.arpa 5.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:19 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1274.eqiad.wmnet 145.48.64.10.in-addr.arpa 5.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1274 - swfrench@cumin1003" * 23:19 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1274 - swfrench@cumin1003" * 23:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1027.eqiad.wmnet with OS bookworm * 23:14 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 23:14 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1274 * 23:13 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1274.eqiad.wmnet with OS trixie * 23:13 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1274.eqiad.wmnet * 23:12 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1274.eqiad.wmnet * 23:12 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1274.eqiad.wmnet * 23:12 ryankemper: [[phab:T430880|T430880]] depooled dnsdisc of wdqs-internal-scholarly-eqiad bc we only have 1 host there * 23:09 ryankemper@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-internal-scholarly,name=eqiad * 23:08 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1272.eqiad.wmnet * 23:08 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1272.eqiad.wmnet * 23:08 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1272.eqiad.wmnet * 23:01 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], xfer to freshly reimaged/scap-deployed wdqs2025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs2025.codfw.wmnet, repooling source-only afterwards * 22:57 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1272.eqiad.wmnet with OS trixie * 22:56 Amir1: deleting echo notifications from 2015 in group0 * 22:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2025.codfw.wmnet with OS bookworm * 22:35 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1272.eqiad.wmnet with reason: host reimage * 22:32 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 27s) * 22:32 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 22:28 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1269.eqiad.wmnet * 22:28 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1269.eqiad.wmnet * 22:28 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1269.eqiad.wmnet * 22:27 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1272.eqiad.wmnet with reason: host reimage * 22:26 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] (duration: 08m 51s) * 22:22 ladsgroup@deploy2003: ladsgroup, urbanecm: Continuing with deployment * 22:19 ladsgroup@deploy2003: ladsgroup, urbanecm: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:17 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] * 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2025.codfw.wmnet with reason: host reimage * 22:06 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1272 * 22:06 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1272 * 22:05 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1272 * 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1272.eqiad.wmnet 127.48.64.10.in-addr.arpa 7.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:05 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1272.eqiad.wmnet 127.48.64.10.in-addr.arpa 7.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1272 - swfrench@cumin1003" * 22:05 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1272 - swfrench@cumin1003" * 22:03 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2025.codfw.wmnet with reason: host reimage * 22:01 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 22:00 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1272 * 22:00 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1272.eqiad.wmnet with OS trixie * 22:00 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1272.eqiad.wmnet * 21:59 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1272.eqiad.wmnet * 21:59 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1272.eqiad.wmnet * 21:55 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1271.eqiad.wmnet * 21:55 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1271.eqiad.wmnet * 21:55 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1271.eqiad.wmnet * 21:46 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1271.eqiad.wmnet with OS trixie * 21:45 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] (duration: 06m 31s) * 21:40 sbassett@deploy2003: sbassett: Continuing with deployment * 21:40 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2025 * 21:40 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2025 * 21:40 sbassett@deploy2003: sbassett: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:38 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] * 21:37 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2025 * 21:37 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2025.codfw.wmnet 220.48.192.10.in-addr.arpa 0.2.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:37 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2025.codfw.wmnet 220.48.192.10.in-addr.arpa 0.2.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:37 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:37 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2025 - bking@cumin2003" * 21:37 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2025 - bking@cumin2003" * 21:30 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] (duration: 08m 19s) * 21:26 sbassett@deploy2003: sbassett: Continuing with deployment * 21:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1269.eqiad.wmnet with OS trixie * 21:24 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1271.eqiad.wmnet with reason: host reimage * 21:23 sbassett@deploy2003: sbassett: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:22 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:22 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] * 21:20 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2025 * 21:19 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2025.codfw.wmnet with OS bookworm * 21:17 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1271.eqiad.wmnet with reason: host reimage * 21:04 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1269.eqiad.wmnet with reason: host reimage * 21:00 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1269.eqiad.wmnet with reason: host reimage * 20:56 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1271 * 20:55 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1271 * 20:54 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1271 * 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1271.eqiad.wmnet 126.48.64.10.in-addr.arpa 6.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:54 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1271.eqiad.wmnet 126.48.64.10.in-addr.arpa 6.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1271 - swfrench@cumin1003" * 20:54 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1271 - swfrench@cumin1003" * 20:51 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:51 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:51 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:50 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 20:49 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 20:49 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1268.eqiad.wmnet * 20:49 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1268.eqiad.wmnet * 20:49 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1268.eqiad.wmnet * 20:48 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1271 * 20:48 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1271.eqiad.wmnet with OS trixie * 20:47 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1271.eqiad.wmnet * 20:46 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1271.eqiad.wmnet * 20:46 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1271.eqiad.wmnet * 20:41 aude@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] (duration: 07m 34s) * 20:39 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1269 * 20:39 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1269 * 20:36 aude@deploy2003: aude: Continuing with deployment * 20:35 aude@deploy2003: aude: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:33 aude@deploy2003: Started scap sync-world: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] * 20:26 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on db2207.codfw.wmnet with reason: Host down * 20:22 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-video: apply * 20:21 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-video: apply * 20:20 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-timeline: apply * 20:20 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-timeline: apply * 20:20 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-syntaxhighlight: apply * 20:19 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-syntaxhighlight: apply * 20:19 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-media: apply * 20:18 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-media: apply * 20:18 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-constraints: apply * 20:17 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-constraints: apply * 20:17 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox: apply * 20:16 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox: apply * 20:13 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1269 * 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1269.eqiad.wmnet 80.32.64.10.in-addr.arpa 0.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:13 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1269.eqiad.wmnet 80.32.64.10.in-addr.arpa 0.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1269 - kamila@cumin1003" * 20:13 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1269 - kamila@cumin1003" * 20:09 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-flink-codfw cluster: Roll restart of jvm daemons. * 20:07 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 20:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 20:03 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-flink-codfw cluster: Roll restart of jvm daemons. * 20:03 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2207 [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94893 and previous config saved to /var/cache/conftool/dbconfig/20260716-200257-marostegui.json * 20:01 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2204 to s2 primary [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94892 and previous config saved to /var/cache/conftool/dbconfig/20260716-200157-marostegui.json * 20:00 marostegui: Starting emergency s2 codfw failover from db2207 to db2204 - [[phab:T432396|T432396]] * 19:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1035.eqiad.wmnet * 19:56 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2204 with weight 0 [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94891 and previous config saved to /var/cache/conftool/dbconfig/20260716-195628-marostegui.json * 19:55 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 26 hosts with reason: Primary switchover s2 [[phab:T432396|T432396]] * 19:54 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1035.eqiad.wmnet * 19:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1034.eqiad.wmnet * 19:48 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1034.eqiad.wmnet * 19:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1033.eqiad.wmnet * 19:43 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-video: apply * 19:43 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1033.eqiad.wmnet * 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1032.eqiad.wmnet * 19:42 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-video: apply * 19:42 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-timeline: apply * 19:41 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-timeline: apply * 19:41 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-syntaxhighlight: apply * 19:41 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-syntaxhighlight: apply * 19:40 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-media: apply * 19:40 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-media: apply * 19:39 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-constraints: apply * 19:36 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-constraints: apply * 19:36 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox: apply * 19:35 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1032.eqiad.wmnet * 19:35 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1031.eqiad.wmnet * 19:35 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox: apply * 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-video: apply * 19:33 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-video: apply * 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-timeline: apply * 19:33 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-timeline: apply * 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-syntaxhighlight: apply * 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-syntaxhighlight: apply * 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-media: apply * 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-media: apply * 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-constraints: apply * 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-constraints: apply * 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox: apply * 19:31 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox: apply * 19:27 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1031.eqiad.wmnet * 19:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1030.eqiad.wmnet * 19:23 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1001.eqiad.wmnet, repooling source-only afterwards * 19:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1030.eqiad.wmnet * 19:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1029.eqiad.wmnet * 19:17 kamila@cumin1003: START - Cookbook sre.dns.netbox * 19:12 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1029.eqiad.wmnet * 19:06 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1269 * 19:05 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1269.eqiad.wmnet with OS trixie * 19:03 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1269.eqiad.wmnet * 19:03 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1269.eqiad.wmnet * 19:03 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1269.eqiad.wmnet * 18:55 dancy@deploy2003: Finished scap sync-world: testing [[phab:T428971|T428971]] (duration: 02m 41s) * 18:53 dancy@deploy2003: Started scap sync-world: testing [[phab:T428971|T428971]] * 18:31 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1268.eqiad.wmnet with OS trixie * 18:18 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 18:16 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1267.eqiad.wmnet * 18:16 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1267.eqiad.wmnet * 18:16 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1267.eqiad.wmnet * 18:09 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1268.eqiad.wmnet with reason: host reimage * 18:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1001.eqiad.wmnet, repooling source-only afterwards * 18:06 swfrench@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] (duration: 07m 34s) * 18:06 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 23s) * 18:06 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 18:06 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1268.eqiad.wmnet with reason: host reimage * 18:03 bd808@deploy2003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 18:02 bd808@deploy2003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 18:02 swfrench@deploy2003: jiji, swfrench: Continuing with deployment * 18:02 bd808@deploy2003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 18:02 bd808@deploy2003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 18:01 bd808@deploy2003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 18:01 swfrench@deploy2003: jiji, swfrench: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:01 bd808@deploy2003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:59 swfrench@deploy2003: Started scap sync-world: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] * 17:45 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1268 * 17:45 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1268 * 17:44 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1267.eqiad.wmnet with OS trixie * 17:43 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter2006.codfw.wmnet * 17:39 swfrench@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter2006.codfw.wmnet * 17:35 swfrench@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] (duration: 07m 27s) * 17:34 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1268 * 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1268.eqiad.wmnet 78.32.64.10.in-addr.arpa 8.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:34 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1268.eqiad.wmnet 78.32.64.10.in-addr.arpa 8.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1268 - kamila@cumin1003" * 17:34 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1268 - kamila@cumin1003" * 17:31 swfrench@deploy2003: jiji, swfrench: Continuing with deployment * 17:29 swfrench@deploy2003: jiji, swfrench: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:28 kamila@cumin1003: START - Cookbook sre.dns.netbox * 17:28 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1268 * 17:28 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1268.eqiad.wmnet with OS trixie * 17:27 swfrench@deploy2003: Started scap sync-world: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] * 17:23 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1267.eqiad.wmnet with reason: host reimage * 17:18 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1267.eqiad.wmnet with reason: host reimage * 17:18 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1268.eqiad.wmnet * 17:17 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1268.eqiad.wmnet * 17:17 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1268.eqiad.wmnet * 17:12 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter2005.codfw.wmnet * 17:11 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1270.eqiad.wmnet * 17:11 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1270.eqiad.wmnet * 17:11 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1270.eqiad.wmnet * 17:09 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter2005.codfw.wmnet * 17:08 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:08 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update reverse dns for moved arelion cct cr2-eqiad - cmooney@cumin1003" * 17:08 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update reverse dns for moved arelion cct cr2-eqiad - cmooney@cumin1003" * 17:08 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] (duration: 07m 34s) * 17:04 jiji@deploy2003: jiji: Continuing with deployment * 17:03 jiji@deploy2003: jiji: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 17:00 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] * 17:00 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2185.codfw.wmnet with OS trixie * 16:59 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:58 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1270.eqiad.wmnet with OS trixie * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1267 * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1267 * 16:57 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1267 * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1267.eqiad.wmnet 77.32.64.10.in-addr.arpa 7.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:57 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1267.eqiad.wmnet 77.32.64.10.in-addr.arpa 7.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1267 - kamila@cumin1003" * 16:56 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1267 - kamila@cumin1003" * 16:56 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_eqsin * 16:56 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5032.eqsin.wmnet * 16:52 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_esams * 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3073.esams.wmnet * 16:50 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_esams * 16:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3081.esams.wmnet * 16:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1266.eqiad.wmnet * 16:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1266.eqiad.wmnet * 16:45 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1266.eqiad.wmnet * 16:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2185.codfw.wmnet with reason: host reimage * 16:41 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_eqiad * 16:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1114.eqiad.wmnet * 16:41 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_eqiad * 16:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1115.eqiad.wmnet * 16:39 kamila@cumin1003: START - Cookbook sre.dns.netbox * 16:39 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1267 * 16:39 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2185.codfw.wmnet with reason: host reimage * 16:38 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1267.eqiad.wmnet with OS trixie * 16:38 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1267.eqiad.wmnet * 16:38 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1270.eqiad.wmnet with reason: host reimage * 16:37 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1267.eqiad.wmnet * 16:37 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1267.eqiad.wmnet * 16:31 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1270.eqiad.wmnet with reason: host reimage * 16:24 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1264.eqiad.wmnet * 16:24 kamila@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 16:24 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 16:23 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 16:21 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2162: switch maintenance completed codfw rack b6 * 16:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2185.codfw.wmnet with OS trixie * 16:19 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 16:16 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter1007.eqiad.wmnet * 16:15 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_eqsin * 16:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5024.eqsin.wmnet * 16:13 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5031.eqsin.wmnet * 16:13 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3072.esams.wmnet * 16:12 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter1007.eqiad.wmnet * 16:11 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] (duration: 09m 47s) * 16:10 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1266.eqiad.wmnet with OS trixie * 16:10 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1270 * 16:10 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1270 * 16:09 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1270 * 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1270.eqiad.wmnet 125.48.64.10.in-addr.arpa 5.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1270.eqiad.wmnet 125.48.64.10.in-addr.arpa 5.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1270 - swfrench@cumin1003" * 16:09 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1270 - swfrench@cumin1003" * 16:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3080.esams.wmnet * 16:07 jiji@deploy2003: jiji: Continuing with deployment * 16:06 jiji@deploy2003: jiji: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:04 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 16:04 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1265.eqiad.wmnet * 16:03 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1265.eqiad.wmnet * 16:03 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1265.eqiad.wmnet * 16:03 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1270 * 16:03 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1270.eqiad.wmnet with OS trixie * 16:02 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1270.eqiad.wmnet * 16:02 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] * 16:01 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1270.eqiad.wmnet * 16:01 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1270.eqiad.wmnet * 16:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1113.eqiad.wmnet * 16:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1112.eqiad.wmnet * 15:49 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1266.eqiad.wmnet with reason: host reimage * 15:47 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter1006.eqiad.wmnet * 15:45 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1265.eqiad.wmnet with OS trixie * 15:44 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1266.eqiad.wmnet with reason: host reimage * 15:43 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter1006.eqiad.wmnet * 15:42 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] (duration: 09m 46s) * 15:37 jiji@deploy2003: jiji: Continuing with deployment * 15:36 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2162: switch maintenance completed codfw rack b6 * 15:36 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2161: switch maintenance completed codfw rack b6 * 15:34 jiji@deploy2003: jiji: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:32 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5023.eqsin.wmnet * 15:32 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] * 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5030.eqsin.wmnet * 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3071.esams.wmnet * 15:27 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3079.esams.wmnet * 15:25 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1265.eqiad.wmnet with reason: host reimage * 15:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1266 * 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1266 * 15:21 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1110.eqiad.wmnet * 15:20 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1111.eqiad.wmnet * 15:16 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1265.eqiad.wmnet with reason: host reimage * 15:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1001.eqiad.wmnet with OS bookworm * 15:15 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1266 * 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1266.eqiad.wmnet 76.32.64.10.in-addr.arpa 6.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:15 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1266.eqiad.wmnet 76.32.64.10.in-addr.arpa 6.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1266 - kamila@cumin1003" * 15:15 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1266 - kamila@cumin1003" * 15:07 kamila@cumin1003: START - Cookbook sre.dns.netbox * 15:04 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1266 * 15:04 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1264 * 15:04 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1264 * 15:04 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1266.eqiad.wmnet with OS trixie * 15:03 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1264 * 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1264.eqiad.wmnet 74.32.64.10.in-addr.arpa 4.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:03 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1264.eqiad.wmnet 74.32.64.10.in-addr.arpa 4.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1264 - kamila@cumin1003" * 15:03 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1264 - kamila@cumin1003" * 15:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-eqiad * 15:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp1001.eqiad.wmnet * 15:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp1001.eqiad.wmnet * 15:01 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp1001.eqiad.wmnet * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp1001.eqiad.wmnet * 15:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1376-1384].eqiad.wmnet * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1376-1384].eqiad.wmnet * 14:59 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1002.eqiad.wmnet * 14:59 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1266.eqiad.wmnet * 14:58 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1266.eqiad.wmnet * 14:58 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1266.eqiad.wmnet * 14:58 kamila@cumin1003: START - Cookbook sre.dns.netbox * 14:57 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1264 * 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1265 * 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1265 * 14:57 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1310596{{!}}Set $wgMathInternalRestbaseURL explicitly (T349582)]] * 14:57 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1265 * 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1265.eqiad.wmnet 75.32.64.10.in-addr.arpa 5.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:56 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1265.eqiad.wmnet 75.32.64.10.in-addr.arpa 5.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:56 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:56 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1265 - kamila@cumin1003" * 14:56 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1265 - kamila@cumin1003" * 14:53 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1002.eqiad.wmnet * 14:53 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1376-1384].eqiad.wmnet * 14:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1001.eqiad.wmnet with reason: host reimage * 14:51 kamila@cumin1003: START - Cookbook sre.dns.netbox * 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 14:50 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:50 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2161: switch maintenance completed codfw rack b6 * 14:50 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5021.eqsin.wmnet * 14:50 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1265 * 14:49 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:49 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:49 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1265.eqiad.wmnet with OS trixie * 14:49 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5029.eqsin.wmnet * 14:49 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1265.eqiad.wmnet * 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3070.esams.wmnet * 14:48 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1264.eqiad.wmnet * 14:48 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1001.eqiad.wmnet with reason: host reimage * 14:48 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1376-1384].eqiad.wmnet * 14:48 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1265.eqiad.wmnet * 14:47 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1265.eqiad.wmnet * 14:47 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1264.eqiad.wmnet * 14:47 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1264.eqiad.wmnet * 14:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:47 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3078.esams.wmnet * 14:44 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1263.eqiad.wmnet * 14:44 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1263.eqiad.wmnet * 14:44 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1263.eqiad.wmnet * 14:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1108.eqiad.wmnet * 14:40 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1109.eqiad.wmnet * 14:40 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:35 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:34 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:34 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2006.codfw.wmnet * 14:34 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-flink-eqiad cluster: Roll restart of jvm daemons. * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf2002.codfw.wmnet * 14:31 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1002.eqiad.wmnet * 14:29 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2006.codfw.wmnet * 14:27 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:27 kamila@deploy2003: Finished scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] (duration: 02m 57s) * 14:27 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-flink-eqiad cluster: Roll restart of jvm daemons. * 14:26 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf2002.codfw.wmnet * 14:26 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf2001.codfw.wmnet * 14:25 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1002.eqiad.wmnet * 14:25 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1001.eqiad.wmnet * 14:25 kamila@deploy2003: Started scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] * 14:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:21 kamila@deploy2003: sync-world aborted: Test deployment to check rsync is working - [[phab:T432108|T432108]] (duration: 00m 36s) * 14:21 topranks: reboot lsw1-b6-codfw to upgrade JunOS [[phab:T430922|T430922]] * 14:21 kamila@deploy2003: Started scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] * 14:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1001.eqiad.wmnet with OS bookworm * 14:20 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b6-codfw,lsw1-b6-codfw IPv6,lsw1-b6-codfw.mgmt,ssw1-a[1,8]-codfw with reason: lsw1-b6-codfw JunOS upgrade * 14:20 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf2001.codfw.wmnet * 14:19 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1001.eqiad.wmnet * 14:19 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 26 hosts with reason: lsw1-b6-codfw JunOS upgrade * 14:14 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 14:13 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc2022: switch maintenance codfw rack b6 * 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:12 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.parsercache * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool pc2022: switch maintenance codfw rack b6 * 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2251: switch maintenance codfw rack b6 * 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.parsercache * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2251: switch maintenance codfw rack b6 * 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2162: switch maintenance codfw rack b6 * 14:12 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1263.eqiad.wmnet with OS trixie * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2162: switch maintenance codfw rack b6 * 14:11 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2161: switch maintenance codfw rack b6 * 14:11 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2161: switch maintenance codfw rack b6 * 14:08 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5020.eqsin.wmnet * 14:07 btullis@cumin1003: START - Cookbook sre.hadoop.reboot-workers for Hadoop analytics cluster * 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3069.esams.wmnet * 14:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5028.eqsin.wmnet * 14:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1338-1347].eqiad.wmnet * 14:06 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1338-1347].eqiad.wmnet * 14:05 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3077.esams.wmnet * 14:02 topranks: beginning depools for lsw1-b6-codfw maintenance [[phab:T430922|T430922]] * 14:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1106.eqiad.wmnet * 14:00 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-misc1002.eqiad.wmnet * 13:59 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1338-1347].eqiad.wmnet * 13:59 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1107.eqiad.wmnet * 13:56 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-codfw * 13:55 sfaci@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply * 13:54 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-misc1002.eqiad.wmnet * 13:54 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-misc1001.eqiad.wmnet * 13:54 sfaci@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply * 13:50 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1263.eqiad.wmnet with reason: host reimage * 13:50 sfaci@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 13:49 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1338-1347].eqiad.wmnet * 13:49 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-misc1001.eqiad.wmnet * 13:49 sfaci@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 13:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:49 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:45 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1263.eqiad.wmnet with reason: host reimage * 13:40 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:40 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-eqiad * 13:35 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:34 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:33 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-reboot (exit_code=0) rolling reboot on A:dnsbox and (A:eqsin or A:drmrs or A:magru) and not (P<nowiki>{</nowiki>dns5003*<nowiki>}</nowiki> or P<nowiki>{</nowiki>dns7002*<nowiki>}</nowiki>) and (A:dnsbox) * 13:33 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns7001.wikimedia.org * 13:27 sukhe@dns1004: END - running authdns-update * 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5019.eqsin.wmnet * 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3076.esams.wmnet * 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3068.esams.wmnet * 13:25 sukhe@dns1004: START - running authdns-update * 13:24 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5027.eqsin.wmnet * 13:24 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1263 * 13:24 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1263 * 13:23 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1263 * 13:23 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:23 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1104.eqiad.wmnet * 13:21 kamila@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:21 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:20 kamila@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:20 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:20 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:20 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1263 - kamila@cumin1003" * 13:20 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1263 - kamila@cumin1003" * 13:19 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1105.eqiad.wmnet * 13:19 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:19 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:18 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:18 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns7001.wikimedia.org * 13:16 cdobbins@cumin2003: conftool action : set/pooled=yes; selector: name=dns7002.* * 13:14 cdobbins@dns1004: END - running authdns-update * 13:13 sbisson@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] (duration: 08m 03s) * 13:13 cdobbins@dns1004: START - running authdns-update * 13:12 kamila@cumin1003: START - Cookbook sre.dns.netbox * 13:12 cdobbins@cumin2003: conftool action : set/pooled=yes; selector: name=dns7002.*,service=authdns-update * 13:12 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1263 * 13:11 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1263.eqiad.wmnet with OS trixie * 13:11 cdobbins@cumin2003: conftool action : set/pooled=no; selector: name=dns7002.* * 13:11 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1263.eqiad.wmnet * 13:10 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1263.eqiad.wmnet * 13:10 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1263.eqiad.wmnet * 13:09 sbisson@deploy2003: sbisson: Continuing with deployment * 13:08 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:07 sbisson@deploy2003: sbisson: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:05 sbisson@deploy2003: Started scap sync-world: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] * 13:03 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns6002.wikimedia.org * 13:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1298-1307].eqiad.wmnet * 13:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1298-1307].eqiad.wmnet * 12:59 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1262.eqiad.wmnet * 12:59 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1262.eqiad.wmnet * 12:59 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1262.eqiad.wmnet * 12:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1298-1307].eqiad.wmnet * 12:49 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns6002.wikimedia.org * 12:46 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1298-1307].eqiad.wmnet * 12:46 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:46 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3075.esams.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3067.esams.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5018.eqsin.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5026.eqsin.wmnet * 12:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1102.eqiad.wmnet * 12:39 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1103.eqiad.wmnet * 12:35 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:34 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns6001.wikimedia.org * 12:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:28 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:18 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns6001.wikimedia.org * 12:14 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1267-1276].eqiad.wmnet * 12:13 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1267-1276].eqiad.wmnet * 12:04 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1267-1276].eqiad.wmnet * 12:03 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns5004.wikimedia.org * 12:02 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1100.eqiad.wmnet * 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3066.esams.wmnet * 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3074.esams.wmnet * 12:01 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:01 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-codfw * 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5017.eqsin.wmnet * 12:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5025.eqsin.wmnet * 12:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1101.eqiad.wmnet * 11:59 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1267-1276].eqiad.wmnet * 11:58 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:58 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:54 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns5004.wikimedia.org * 11:54 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and (A:eqsin or A:drmrs or A:magru) and not (P<nowiki>{</nowiki>dns5003*<nowiki>}</nowiki> or P<nowiki>{</nowiki>dns7002*<nowiki>}</nowiki>) and (A:dnsbox) * 11:54 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:53 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:53 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-eqiad * 11:51 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:51 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:50 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:50 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_eqiad * 11:50 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:50 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_eqiad * 11:49 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_esams * 11:49 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_esams * 11:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_eqsin * 11:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_eqsin * 11:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:44 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-eqiad * 11:43 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-codfw * 11:42 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:41 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:24 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-eqiad * 11:23 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-codfw * 11:22 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:15 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:14 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:09 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2066.codfw.wmnet with OS trixie * 11:05 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1151-1160].eqiad.wmnet * 11:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1151-1160].eqiad.wmnet * 10:59 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1068.eqiad.wmnet with OS trixie * 10:55 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.major-upgrade (exit_code=99) * 10:55 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 10:54 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1151-1160].eqiad.wmnet * 10:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2066.codfw.wmnet with reason: host reimage * 10:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1151-1160].eqiad.wmnet * 10:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:42 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2066.codfw.wmnet with reason: host reimage * 10:39 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:37 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:36 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:23 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:22 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2066.codfw.wmnet with OS trixie * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:07 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2065.codfw.wmnet with OS trixie * 10:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:06 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 10:06 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 10:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 10:03 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:03 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 10:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 09:59 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 09:57 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 09:57 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 09:52 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 09:47 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:46 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:46 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2065.codfw.wmnet with reason: host reimage * 09:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2065.codfw.wmnet with reason: host reimage * 09:40 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:39 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 09:39 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:39 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:39 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:37 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:29 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox2003.codfw.wmnet * 09:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox2003.codfw.wmnet * 09:25 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:25 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:24 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:24 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:24 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:21 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:20 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2065.codfw.wmnet with OS trixie * 09:13 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1068.eqiad.wmnet with OS trixie * 09:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2064.codfw.wmnet with OS trixie * 09:08 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 09:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 09:07 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2162: Repooling after switchover * 09:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1067.eqiad.wmnet with OS trixie * 09:00 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:59 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 08:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 08:57 tappof: bump space for prometheus k8s-dse in eqiad * 08:56 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping2004.codfw.wmnet * 08:52 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host ping2004.codfw.wmnet * 08:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 08:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping1004.eqiad.wmnet * 08:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:51 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:49 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 08:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host ping1004.eqiad.wmnet * 08:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2064.codfw.wmnet with reason: host reimage * 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2064.codfw.wmnet with reason: host reimage * 08:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1067.eqiad.wmnet with reason: host reimage * 08:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:33 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1067.eqiad.wmnet with reason: host reimage * 08:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:21 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2162: Repooling after switchover * 08:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2064.codfw.wmnet with OS trixie * 08:16 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1067.eqiad.wmnet with OS trixie * 08:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:15 cgoubert@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-eqiad * 08:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2062.codfw.wmnet with OS trixie * 08:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1066.eqiad.wmnet with OS trixie * 08:02 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2162: Repooling after switchover * 07:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2162: Repooling after switchover * 07:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2162 [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94870 and previous config saved to /var/cache/conftool/dbconfig/20260716-075530-cwilliams.json * 07:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2241 to x3 primary [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94869 and previous config saved to /var/cache/conftool/dbconfig/20260716-075314-cwilliams.json * 07:52 cezmunsta: Starting x3 codfw failover from db2162 to db2241 - [[phab:T430925|T430925]] * 07:50 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:50 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:47 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 07:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2241 with weight 0 [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94868 and previous config saved to /var/cache/conftool/dbconfig/20260716-074507-cwilliams.json * 07:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 18 hosts with reason: Primary switchover x3 [[phab:T430925|T430925]] * 07:43 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1066.eqiad.wmnet with reason: host reimage * 07:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:dse-k8s-worker-eqiad * 07:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1028.eqiad.wmnet * 07:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1028.eqiad.wmnet * 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 07:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1066.eqiad.wmnet with reason: host reimage * 07:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1028.eqiad.wmnet * 07:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1028.eqiad.wmnet * 07:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1027.eqiad.wmnet * 07:35 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1027.eqiad.wmnet * 07:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1027.eqiad.wmnet * 07:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1027.eqiad.wmnet * 07:28 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1026.eqiad.wmnet * 07:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1026.eqiad.wmnet * 07:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast2003.wikimedia.org * 07:21 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1026.eqiad.wmnet * 07:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1066.eqiad.wmnet with OS trixie * 07:19 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast2003.wikimedia.org * 07:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2062.codfw.wmnet with OS trixie * 06:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1026.eqiad.wmnet * 06:51 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1025.eqiad.wmnet * 06:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1025.eqiad.wmnet * 06:47 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 06:47 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 06:44 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1025.eqiad.wmnet * 06:14 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1025.eqiad.wmnet * 06:14 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1024.eqiad.wmnet * 06:14 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1024.eqiad.wmnet * 06:07 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1024.eqiad.wmnet * 05:37 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1024.eqiad.wmnet * 05:37 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1023.eqiad.wmnet * 05:37 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1023.eqiad.wmnet * 05:26 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1023.eqiad.wmnet * 04:56 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1023.eqiad.wmnet * 04:56 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1022.eqiad.wmnet * 04:56 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1022.eqiad.wmnet * 04:49 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1022.eqiad.wmnet * 04:19 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1022.eqiad.wmnet * 04:19 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1021.eqiad.wmnet * 04:19 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1021.eqiad.wmnet * 04:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1021.eqiad.wmnet * 03:38 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1021.eqiad.wmnet * 03:38 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1020.eqiad.wmnet * 03:38 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1020.eqiad.wmnet * 03:20 btullis@cumin1003: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1020.eqiad.wmnet * 03:18 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1020.eqiad.wmnet * 03:18 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1019.eqiad.wmnet * 03:18 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1019.eqiad.wmnet * 03:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1019.eqiad.wmnet * 02:41 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1019.eqiad.wmnet * 02:41 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1018.eqiad.wmnet * 02:41 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1018.eqiad.wmnet * 02:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling both afterwards * 02:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2003.codfw.wmnet -> wcqs2001.codfw.wmnet, repooling both afterwards * 02:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1018.eqiad.wmnet * 02:30 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1018.eqiad.wmnet * 02:30 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1014.eqiad.wmnet * 02:30 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1014.eqiad.wmnet * 02:24 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1014.eqiad.wmnet * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 01:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1014.eqiad.wmnet * 01:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1013.eqiad.wmnet * 01:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1013.eqiad.wmnet * 01:47 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1013.eqiad.wmnet * 01:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2003.codfw.wmnet -> wcqs2001.codfw.wmnet, repooling both afterwards * 01:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling both afterwards * 01:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1013.eqiad.wmnet * 01:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1012.eqiad.wmnet * 01:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1012.eqiad.wmnet * 01:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1012.eqiad.wmnet * 01:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1012.eqiad.wmnet * 01:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1011.eqiad.wmnet * 01:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1011.eqiad.wmnet * 01:08 ryankemper@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] scap deploy post bookworm reimage (duration: 00m 23s) * 01:08 ryankemper@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] scap deploy post bookworm reimage * 01:08 ryankemper@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): scap deploy post bookworm reimage (duration: 00m 46s) * 01:07 ryankemper@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): scap deploy post bookworm reimage * 01:04 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1011.eqiad.wmnet * 01:04 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1011.eqiad.wmnet * 01:04 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1010.eqiad.wmnet * 01:04 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1010.eqiad.wmnet * 00:57 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1010.eqiad.wmnet * 00:57 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1010.eqiad.wmnet * 00:57 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1009.eqiad.wmnet * 00:57 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1009.eqiad.wmnet * 00:50 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1009.eqiad.wmnet * 00:20 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1009.eqiad.wmnet * 00:20 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1008.eqiad.wmnet * 00:20 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1008.eqiad.wmnet * 00:13 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1008.eqiad.wmnet == 2026-07-15 == * 23:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2001.codfw.wmnet with OS bookworm * 23:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1008.eqiad.wmnet * 23:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1007.eqiad.wmnet * 23:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1007.eqiad.wmnet * 23:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1007.eqiad.wmnet * 23:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1007.eqiad.wmnet * 23:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1006.eqiad.wmnet * 23:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1006.eqiad.wmnet * 23:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1006.eqiad.wmnet * 23:29 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1006.eqiad.wmnet * 23:28 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1005.eqiad.wmnet * 23:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1005.eqiad.wmnet * 23:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1002.eqiad.wmnet with OS bookworm * 23:21 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1005.eqiad.wmnet * 23:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2001.codfw.wmnet with reason: host reimage * 23:15 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host datahubsearch1001.eqiad.wmnet with OS bookworm * 23:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2001.codfw.wmnet with reason: host reimage * 23:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 23:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 22:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 22:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1005.eqiad.wmnet * 22:51 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1004.eqiad.wmnet * 22:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1004.eqiad.wmnet * 22:45 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1004.eqiad.wmnet * 22:44 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 22:44 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS trixie * 22:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host datahubsearch1001.eqiad.wmnet with OS bookworm * 22:34 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host datahubsearch1001.eqiad.wmnet with OS bookworm * 22:16 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on datahubsearch[1002-1003].eqiad.wmnet with reason: Using datahubsearch1001 to test bookworm reimages * 22:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1004.eqiad.wmnet * 22:15 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1003.eqiad.wmnet * 22:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1003.eqiad.wmnet * 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 22:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1003.eqiad.wmnet * 22:08 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1003.eqiad.wmnet * 22:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1002.eqiad.wmnet * 22:08 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1002.eqiad.wmnet * 22:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host datahubsearch1001.eqiad.wmnet with OS bookworm * 22:05 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 22:02 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm * 22:01 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on datahubsearch[1001-1003].eqiad.wmnet with reason: Using datahubsearch1001 to test bookworm reimages * 22:01 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1002.eqiad.wmnet * 22:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1002.eqiad.wmnet * 22:00 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1001.eqiad.wmnet * 22:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1001.eqiad.wmnet * 21:53 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1001.eqiad.wmnet * 21:52 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 21:50 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 21:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS trixie * 21:50 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS bookworm * 21:43 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 21:38 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:30 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wcqs1002'] * 21:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:29 lerickson@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 21:29 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:29 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:29 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS bookworm * 21:28 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 21:28 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm * 21:23 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1001.eqiad.wmnet * 21:23 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:23 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:22 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 21:20 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:18 swfrench-wmf: reprepro include php8.3_8.3.32-1+wmf11u2 into component/php83 for bullseye-wikimedia * 21:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:16 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:15 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-druid-public cluster: Roll restart of jvm daemons. * 21:08 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:05 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1001.eqiad.wmnet * 21:05 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1001.eqiad.wmnet * 21:04 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-druid-public cluster: Roll restart of jvm daemons. * 21:02 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 21:01 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 21:01 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 21:00 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 20:59 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1001.eqiad.wmnet * 20:59 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1001.eqiad.wmnet * 20:59 btullis@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:dse-k8s-worker-eqiad * 20:55 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 20:55 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 20:45 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm * 20:21 jhathaway: puppet is re-enabled, have fun, but not too much fun! * 20:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 20:17 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs2001'] * 20:12 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs2001'] * 20:11 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs2001'] * 20:09 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:08 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 20:05 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:05 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 20:04 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs2001'] * 20:03 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 20:03 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm * 20:02 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:02 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 20:01 jhathaway: disabling puppet fleet wide to roll out kafka patch * 19:55 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host relforge1010.eqiad.wmnet * 19:52 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 19:52 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 19:48 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 19:48 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 19:48 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 19:47 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 19:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 19:45 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 19:45 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 19:44 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1010.eqiad.wmnet * 19:38 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1262.eqiad.wmnet with OS trixie * 19:17 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1262.eqiad.wmnet with reason: host reimage * 19:11 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1262.eqiad.wmnet with reason: host reimage * 18:59 cdobbins@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS trixie * 18:54 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 18:53 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 18:52 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1262 * 18:52 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1262 * 18:51 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1262 * 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1262.eqiad.wmnet 72.32.64.10.in-addr.arpa 2.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:51 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1262.eqiad.wmnet 72.32.64.10.in-addr.arpa 2.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1262 - kamila@cumin1003" * 18:51 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1262 - kamila@cumin1003" * 18:46 kamila@cumin1003: START - Cookbook sre.dns.netbox * 18:46 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1262 * 18:46 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 18:46 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ncmonitor1001.eqiad.wmnet * 18:46 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 18:45 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1262.eqiad.wmnet with OS trixie * 18:45 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 18:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1262.eqiad.wmnet * 18:44 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1262.eqiad.wmnet * 18:44 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1262.eqiad.wmnet * 18:42 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host ncmonitor1001.eqiad.wmnet * 18:29 topranks: pull power on cr1-eqiad to install new switch-control boards [[phab:T426343|T426343]] * 18:29 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs[1018-1020].eqiad.wmnet with reason: line card install in cr1-eqiad * 18:27 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 14 hosts with reason: linecard install in cr1-eqad * 18:22 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_ulsfo * 18:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4052.ulsfo.wmnet * 18:19 cdobbins@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 18:15 cdobbins@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 18:14 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_drmrs * 18:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6016.drmrs.wmnet * 18:12 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_ulsfo * 18:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4044.ulsfo.wmnet * 18:10 sukhe@cumin1003: END (ERROR) - Cookbook sre.cdn.roll-reboot (exit_code=97) rolling reboot on A:cp-upload_drmrs * 18:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2241: Security update * 17:56 topranks: start draining traffic on cr1-eqiad ahead of line card installation [[phab:T426343|T426343]] * 17:47 cdobbins@cumin2003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie * 17:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4051.ulsfo.wmnet * 17:40 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:39 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 17:34 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6007.drmrs.wmnet * 17:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6015.drmrs.wmnet * 17:32 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:31 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 17:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4043.ulsfo.wmnet * 17:27 lerickson@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:25 lerickson@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 17:22 lerickson@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-codfw * 17:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp2001.codfw.wmnet * 17:22 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 17:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp2001.codfw.wmnet * 17:22 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 17:19 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2241: Security update * 17:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2241.codfw.wmnet * 17:17 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2241.codfw.wmnet * 17:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp2001.codfw.wmnet * 17:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp2001.codfw.wmnet * 17:15 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2366-2374].codfw.wmnet * 17:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2366-2374].codfw.wmnet * 17:10 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply * 17:10 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply * 17:08 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2366-2374].codfw.wmnet * 17:06 sukhe: sre.dns.roll-reboot to resume later * 17:06 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-reboot (exit_code=97) rolling reboot on A:dnsbox and not (A:ulsfo or A:magru) and (A:dnsbox) * 17:06 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns5003.wikimedia.org * 17:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2241: Security update * 17:03 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2241: Security update * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply * 17:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2366-2374].codfw.wmnet * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply * 17:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2357-2365].codfw.wmnet * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 17:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2357-2365].codfw.wmnet * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply * 16:55 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2357-2365].codfw.wmnet * 16:55 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 16:53 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6006.drmrs.wmnet * 16:52 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6014.drmrs.wmnet * 16:52 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 16:51 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 16:50 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2357-2365].codfw.wmnet * 16:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4042.ulsfo.wmnet * 16:50 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2347-2356].codfw.wmnet * 16:50 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2347-2356].codfw.wmnet * 16:49 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns5003.wikimedia.org * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply * 16:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4050.ulsfo.wmnet * 16:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2347-2356].codfw.wmnet * 16:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2347-2356].codfw.wmnet * 16:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2337-2346].codfw.wmnet * 16:36 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2337-2346].codfw.wmnet * 16:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:dse-k8s-worker-codfw * 16:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2003.codfw.wmnet * 16:35 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2003.codfw.wmnet * 16:34 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns3004.wikimedia.org * 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply * 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply * 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply * 16:30 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply * 16:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2003.codfw.wmnet * 16:29 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2337-2346].codfw.wmnet * 16:24 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2003.codfw.wmnet * 16:24 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2002.codfw.wmnet * 16:24 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2002.codfw.wmnet * 16:23 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns3004.wikimedia.org * 16:23 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2337-2346].codfw.wmnet * 16:23 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2327-2336].codfw.wmnet * 16:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2327-2336].codfw.wmnet * 16:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2002.codfw.wmnet * 16:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2327-2336].codfw.wmnet * 16:12 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2002.codfw.wmnet * 16:12 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2001.codfw.wmnet * 16:12 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2001.codfw.wmnet * 16:12 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1065.eqiad.wmnet with OS trixie * 16:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6005.drmrs.wmnet * 16:11 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6013.drmrs.wmnet * 16:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4041.ulsfo.wmnet * 16:08 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns3003.wikimedia.org * 16:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2327-2336].codfw.wmnet * 16:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2317-2326].codfw.wmnet * 16:06 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2317-2326].codfw.wmnet * 16:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2001.codfw.wmnet * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply * 16:03 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4049.ulsfo.wmnet * 16:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2001.codfw.wmnet * 16:00 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test2001.codfw.wmnet * 16:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test2001.codfw.wmnet * 16:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2063.codfw.wmnet with OS trixie * 15:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2317-2326].codfw.wmnet * 15:57 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns3003.wikimedia.org * 15:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test2001.codfw.wmnet * 15:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test2001.codfw.wmnet * 15:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2004.codfw.wmnet * 15:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2004.codfw.wmnet * 15:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2317-2326].codfw.wmnet * 15:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2307-2316].codfw.wmnet * 15:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2307-2316].codfw.wmnet * 15:49 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2004.codfw.wmnet * 15:48 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2004.codfw.wmnet * 15:48 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2003.codfw.wmnet * 15:48 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2003.codfw.wmnet * 15:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 15:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2307-2316].codfw.wmnet * 15:42 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2003.codfw.wmnet * 15:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 15:42 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2003.codfw.wmnet * 15:42 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2002.codfw.wmnet * 15:42 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2002.codfw.wmnet * 15:42 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2006.wikimedia.org * 15:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2063.codfw.wmnet with reason: host reimage * 15:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2307-2316].codfw.wmnet * 15:37 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2297-2306].codfw.wmnet * 15:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2297-2306].codfw.wmnet * 15:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2002.codfw.wmnet * 15:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2002.codfw.wmnet * 15:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2001.codfw.wmnet * 15:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2001.codfw.wmnet * 15:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2063.codfw.wmnet with reason: host reimage * 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6004.drmrs.wmnet * 15:31 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2001.codfw.wmnet * 15:31 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2001.codfw.wmnet * 15:31 btullis@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:dse-k8s-worker-codfw * 15:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6012.drmrs.wmnet * 15:28 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2006.wikimedia.org * 15:27 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-analytics cluster: Roll restart of jvm daemons. * 15:27 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2297-2306].codfw.wmnet * 15:27 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4040.ulsfo.wmnet * 15:24 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1065.eqiad.wmnet with OS trixie * 15:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4048.ulsfo.wmnet * 15:21 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-analytics cluster: Roll restart of jvm daemons. * 15:21 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2297-2306].codfw.wmnet * 15:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2287-2296].codfw.wmnet * 15:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2287-2296].codfw.wmnet * 15:20 btullis@cumin1003: END (PASS) - Cookbook sre.druid.reboot-workers (exit_code=0) for Druid public cluster: Reboot Druid nodes * 15:18 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 15:17 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm * 15:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2063.codfw.wmnet with OS trixie * 15:13 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2005.wikimedia.org * 15:11 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2287-2296].codfw.wmnet * 15:11 btullis@cumin1003: END (PASS) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=0) rolling reboot on A:cephosd-eqiad * 15:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1064.eqiad.wmnet with OS trixie * 15:05 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2062.codfw.wmnet with OS trixie * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2287-2296].codfw.wmnet * 15:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2277-2286].codfw.wmnet * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2277-2286].codfw.wmnet * 14:59 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2005.wikimedia.org * 14:57 brouberol@cumin1003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-jumbo-eqiad * 14:52 btullis@cumin1003: END (PASS) - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas (exit_code=0) rolling reboot on A:schema-codfw * 14:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6003.drmrs.wmnet * 14:50 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:50 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host relforge1009.eqiad.wmnet * 14:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2277-2286].codfw.wmnet * 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6011.drmrs.wmnet * 14:47 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 14:46 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>ml-serve1001.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 14:46 klausman@cumin1003: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) pool for host ml-serve1001.eqiad.wmnet * 14:46 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 14:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1001.eqiad.wmnet * 14:45 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4039.ulsfo.wmnet * 14:44 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1009.eqiad.wmnet * 14:44 btullis@cumin1003: START - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas rolling reboot on A:schema-codfw * 14:44 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2004.wikimedia.org * 14:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2277-2286].codfw.wmnet * 14:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2267-2276].codfw.wmnet * 14:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2267-2276].codfw.wmnet * 14:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 14:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4047.ulsfo.wmnet * 14:40 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1001.eqiad.wmnet * 14:38 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 14:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 14:36 topranks: disconnect power on cr2-eqiad to shut down device for switch fabric replacement [[phab:T426343|T426343]] * 14:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2267-2276].codfw.wmnet * 14:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 14:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1001.eqiad.wmnet * 14:35 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>ml-serve1001.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 14:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 14:34 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 14:33 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:33 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:30 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2004.wikimedia.org * 14:29 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2267-2276].codfw.wmnet * 14:29 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2257-2266].codfw.wmnet * 14:29 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2257-2266].codfw.wmnet * 14:24 btullis@cumin1003: END (PASS) - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas (exit_code=0) rolling reboot on A:schema-eqiad * 14:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2257-2266].codfw.wmnet * 14:20 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:20 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:19 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:17 jforrester@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2257-2266].codfw.wmnet * 14:16 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:16 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2062.codfw.wmnet with OS trixie * 14:15 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1064.eqiad.wmnet with OS trixie * 14:15 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1006.wikimedia.org * 14:15 btullis@cumin1003: START - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas rolling reboot on A:schema-eqiad * 14:14 topranks: switch routing-engine on cr2-eqiad resetting all interfaces [[phab:T417873|T417873]] * 14:11 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:11 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:10 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm * 14:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6002.drmrs.wmnet * 14:09 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6010.drmrs.wmnet * 14:06 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1006.wikimedia.org * 14:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:05 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 14:05 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4038.ulsfo.wmnet * 14:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4046.ulsfo.wmnet * 14:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:00 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on cr1-eqiad with reason: switch upgrade and line card install * 14:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:59 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:57 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:57 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:55 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-eqiad * 13:55 btullis@cumin1003: START - Cookbook sre.druid.reboot-workers for Druid public cluster: Reboot Druid nodes * 13:53 topranks: switch routing-engine on cr2-eqiad resetting all interfaces [[phab:T417873|T417873]] * 13:51 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1005.wikimedia.org * 13:50 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:49 brouberol@cumin1003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-test-eqiad * 13:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:44 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2197-2206].codfw.wmnet * 13:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2197-2206].codfw.wmnet * 13:36 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1005.wikimedia.org * 13:35 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2197-2206].codfw.wmnet * 13:30 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2197-2206].codfw.wmnet * 13:28 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6001.drmrs.wmnet * 13:28 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6009.drmrs.wmnet * 13:28 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2187-2196].codfw.wmnet * 13:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2187-2196].codfw.wmnet * 13:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 13:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4037.ulsfo.wmnet * 13:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2001 * 13:22 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2001 * 13:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4045.ulsfo.wmnet * 13:21 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1004.wikimedia.org * 13:19 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on lvs[1018-1020].eqiad.wmnet with reason: switch upgrade and line card install * 13:18 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2009.codfw.wmnet * 13:18 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2001 * 13:18 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2001.codfw.wmnet 26.16.192.10.in-addr.arpa 6.2.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:17 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2001.codfw.wmnet 26.16.192.10.in-addr.arpa 6.2.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:17 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:17 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2001 - bking@cumin2003" * 13:17 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2001 - bking@cumin2003" * 13:17 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2009.codfw.wmnet * 13:17 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_drmrs * 13:17 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2187-2196].codfw.wmnet * 13:17 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_drmrs * 13:17 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on 15 hosts with reason: switch upgrade and line card install * 13:17 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:15 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 13:13 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:13 brouberol@cumin1003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-jumbo-eqiad * 13:13 brouberol@cumin1003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-test-eqiad * 13:13 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1004.wikimedia.org * 13:13 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and not (A:ulsfo or A:magru) and (A:dnsbox) * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:12 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_ulsfo * 13:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:12 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_ulsfo * 13:11 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2187-2196].codfw.wmnet * 13:11 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 13:11 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 13:06 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 13:05 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling source-only afterwards * 13:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2001 * 13:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2009.codfw.wmnet with OS trixie * 13:03 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:03 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling source-only afterwards * 13:01 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 15s) * 13:01 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 13:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 12:57 btullis@cumin1003: END (PASS) - Cookbook sre.druid.reboot-workers (exit_code=0) for Druid analytics cluster: Reboot Druid nodes * 12:54 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 12:54 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2163-2172].codfw.wmnet * 12:54 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2163-2172].codfw.wmnet * 12:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2163-2172].codfw.wmnet * 12:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2009.codfw.wmnet with reason: host reimage * 12:41 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2163-2172].codfw.wmnet * 12:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2153-2162].codfw.wmnet * 12:40 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2153-2162].codfw.wmnet * 12:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2009.codfw.wmnet with reason: host reimage * 12:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2153-2162].codfw.wmnet * 12:29 btullis@cumin1003: END (PASS) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=0) rolling reboot on A:cephosd-codfw * 12:25 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2153-2162].codfw.wmnet * 12:25 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2143-2152].codfw.wmnet * 12:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2143-2152].codfw.wmnet * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2009 * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2009 * 12:22 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2009 * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2009.codfw.wmnet 139.0.192.10.in-addr.arpa 9.3.1.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:22 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2009.codfw.wmnet 139.0.192.10.in-addr.arpa 9.3.1.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2009 - mvernon@cumin2003" * 12:22 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2009 - mvernon@cumin2003" * 12:16 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 12:15 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 12:15 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 12:15 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2009 * 12:15 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 12:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2009.codfw.wmnet with OS trixie * 12:15 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 12:14 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 12:14 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2143-2152].codfw.wmnet * 12:13 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 12:12 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2010.codfw.wmnet * 12:11 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2010.codfw.wmnet * 12:10 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 12:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2143-2152].codfw.wmnet * 12:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2133-2142].codfw.wmnet * 12:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2133-2142].codfw.wmnet * 12:02 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 11:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2133-2142].codfw.wmnet * 11:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2133-2142].codfw.wmnet * 11:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:49 mvolz@deploy2003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:49 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-codfw * 11:48 mvolz@deploy2003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:47 btullis@cumin1003: START - Cookbook sre.druid.reboot-workers for Druid analytics cluster: Reboot Druid nodes * 11:46 mvolz@deploy2003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:46 mvolz@deploy2003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:45 mvolz@deploy2003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:44 mvolz@deploy2003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:40 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] (duration: 11m 38s) * 11:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2010.codfw.wmnet with OS trixie * 11:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2105-2114].codfw.wmnet * 11:36 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2105-2114].codfw.wmnet * 11:36 krinkle@deploy2003: physikerwelt, krinkle: Continuing with deployment * 11:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1018: Security updates * 11:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:36 root@cumin1003: START - Cookbook sre.mysql.parsercache * 11:36 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1018: Security updates * 11:31 krinkle@deploy2003: physikerwelt, krinkle: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:29 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] * 11:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2105-2114].codfw.wmnet * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2105-2114].codfw.wmnet * 11:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2010.codfw.wmnet with reason: host reimage * 11:12 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2010.codfw.wmnet with reason: host reimage * 11:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1018: Security updates * 11:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:10 root@cumin1003: START - Cookbook sre.mysql.parsercache * 11:10 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1018: Security updates * 11:09 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1009.eqiad.wmnet with OS trixie * 11:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow7002.magru.wmnet * 11:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 11:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 11:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=tegola-vector-tiles,name=eqiad * 11:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=kartotherian,name=eqiad * 11:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow7002.magru.wmnet * 10:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2010 * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2010 * 10:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 10:54 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 10:54 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2010 * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2010.codfw.wmnet 76.16.192.10.in-addr.arpa 6.7.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:54 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2010.codfw.wmnet 76.16.192.10.in-addr.arpa 6.7.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2010 - mvernon@cumin2003" * 10:54 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2010 - mvernon@cumin2003" * 10:49 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 10:49 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2010 * 10:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1009.eqiad.wmnet with reason: host reimage * 10:49 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2010.codfw.wmnet with OS trixie * 10:46 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2011.codfw.wmnet * 10:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1011.eqiad.wmnet * 10:44 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2011.codfw.wmnet * 10:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 10:44 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 10:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1009.eqiad.wmnet with reason: host reimage * 10:44 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow6001.drmrs.wmnet * 10:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1017: Security updates * 10:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:39 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1017: Security updates * 10:39 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow6001.drmrs.wmnet * 10:38 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1011.eqiad.wmnet * 10:35 cgoubert@deploy2003: Finished deploy [restbase/deploy@06301bd]: Deploying {{Gerrit|1306088}} {{Gerrit|1308347}} - [[phab:T429944|T429944]] [[phab:T428279|T428279]] (duration: 28m 34s) * 10:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1012.eqiad.wmnet * 10:35 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 10:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow5003.eqsin.wmnet * 10:34 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2011.codfw.wmnet with OS trixie * 10:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1009.eqiad.wmnet with OS trixie * 10:28 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1012.eqiad.wmnet * 10:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1013.eqiad.wmnet * 10:27 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:27 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow5003.eqsin.wmnet * 10:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:26 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow4003.ulsfo.wmnet * 10:25 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 10:25 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 10:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow4003.ulsfo.wmnet * 10:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1013.eqiad.wmnet * 10:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1014.eqiad.wmnet * 10:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:15 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2011.codfw.wmnet with reason: host reimage * 10:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1017: Security updates * 10:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:14 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:14 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1017: Security updates * 10:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow3004.esams.wmnet * 10:11 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2011.codfw.wmnet with reason: host reimage * 10:10 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:10 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1014.eqiad.wmnet * 10:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki2003.codfw.wmnet * 10:09 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 10:09 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 10:09 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow3004.esams.wmnet * 10:08 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2004.codfw.wmnet * 10:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1010.eqiad.wmnet with OS trixie * 10:07 cgoubert@deploy2003: Started deploy [restbase/deploy@06301bd]: Deploying {{Gerrit|1306088}} {{Gerrit|1308347}} - [[phab:T429944|T429944]] [[phab:T428279|T428279]] * 10:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host rpki2003.codfw.wmnet * 10:04 topranks: push out config change to BGP_outfilter on core routers [[phab:T431849|T431849]] * 10:02 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow2004.codfw.wmnet * 09:59 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 09:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2003.codfw.wmnet * 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2011 * 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2011 * 09:53 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 09:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:52 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2011 * 09:52 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2011.codfw.wmnet 36.32.192.10.in-addr.arpa 6.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:52 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2011.codfw.wmnet 36.32.192.10.in-addr.arpa 6.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:51 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:51 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2011 - mvernon@cumin2003" * 09:51 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2011 - mvernon@cumin2003" * 09:51 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow2003.codfw.wmnet * 09:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1003.eqiad.wmnet * 09:49 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 09:49 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 09:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1010.eqiad.wmnet with reason: host reimage * 09:47 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 09:47 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 09:47 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 09:47 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2011 * 09:46 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2011.codfw.wmnet with OS trixie * 09:44 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow1003.eqiad.wmnet * 09:44 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2012.codfw.wmnet * 09:44 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1002.eqiad.wmnet * 09:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1010.eqiad.wmnet with reason: host reimage * 09:43 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2012.codfw.wmnet * 09:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Security updates * 09:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:43 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:43 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Security updates * 09:42 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:40 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow1002.eqiad.wmnet * 09:40 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 09:37 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki1001.eqiad.wmnet * 09:36 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 09:36 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 09:33 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host rpki1001.eqiad.wmnet * 09:32 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:32 cgoubert@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-codfw * 09:31 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=kartotherian,name=eqiad * 09:31 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola-vector-tiles,name=eqiad * 09:31 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 09:31 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2012.codfw.wmnet with OS trixie * 09:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1010.eqiad.wmnet with OS trixie * 09:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1011.eqiad.wmnet with OS trixie * 09:21 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Security updates * 09:21 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:21 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:21 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Security updates * 09:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2012.codfw.wmnet with reason: host reimage * 09:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1011.eqiad.wmnet with reason: host reimage * 09:08 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2012.codfw.wmnet with reason: host reimage * 09:05 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1011.eqiad.wmnet with reason: host reimage * 08:55 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:52 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1011.eqiad.wmnet with OS trixie * 08:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1022: Security updates * 08:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2012 * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2012 * 08:50 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1022: Security updates * 08:50 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2012 * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2012.codfw.wmnet 44.48.192.10.in-addr.arpa 4.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:50 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2012.codfw.wmnet 44.48.192.10.in-addr.arpa 4.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2012 - mvernon@cumin2003" * 08:50 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2012 - mvernon@cumin2003" * 08:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1012.eqiad.wmnet with OS trixie * 08:44 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 08:44 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2012 * 08:43 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2012.codfw.wmnet with OS trixie * 08:42 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2013.codfw.wmnet * 08:41 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2013.codfw.wmnet * 08:35 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 08:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host krb1002.eqiad.wmnet * 08:30 elukey@dns1004: END - running authdns-update * 08:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1012.eqiad.wmnet with reason: host reimage * 08:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Security updates * 08:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:28 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:28 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Security updates * 08:27 elukey@dns1004: START - running authdns-update * 08:26 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 08:26 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host krb1002.eqiad.wmnet * 08:22 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1012.eqiad.wmnet with reason: host reimage * 08:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host krb2002.codfw.wmnet * 08:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast6003.wikimedia.org * 08:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2013.codfw.wmnet with OS trixie * 08:13 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast6003.wikimedia.org * 08:12 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast3007.wikimedia.org * 08:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host krb2002.codfw.wmnet * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Security updates * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:09 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:09 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Security updates * 08:07 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1012.eqiad.wmnet with OS trixie * 08:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast3007.wikimedia.org * 08:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast5005.wikimedia.org * 07:58 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast5005.wikimedia.org * 07:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1013.eqiad.wmnet with OS trixie * 07:53 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2013.codfw.wmnet with reason: host reimage * 07:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1021: Security updates * 07:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:53 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:53 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1021: Security updates * 07:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast1004.wikimedia.org * 07:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2013.codfw.wmnet with reason: host reimage * 07:46 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast1004.wikimedia.org * 07:40 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1013.eqiad.wmnet with reason: host reimage * 07:36 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1013.eqiad.wmnet with reason: host reimage * 07:31 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2013 * 07:31 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2013 * 07:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1021: Security updates * 07:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:30 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:30 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1021: Security updates * 07:24 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2013 * 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2013.codfw.wmnet 87.0.192.10.in-addr.arpa 7.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:24 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2013.codfw.wmnet 87.0.192.10.in-addr.arpa 7.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2013 - mvernon@cumin2003" * 07:24 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2013 - mvernon@cumin2003" * 07:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1013.eqiad.wmnet with OS trixie * 07:19 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 07:19 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2013 * 07:19 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2013.codfw.wmnet with OS trixie * 07:13 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] (duration: 07m 48s) * 07:09 kharlan@deploy2003: kharlan: Continuing with deployment * 07:08 kharlan@deploy2003: kharlan: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:06 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 01:15 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 01:14 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply == 2026-07-14 == * 22:51 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_magru * 22:51 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7016.magru.wmnet * 22:46 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_magru * 22:46 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7008.magru.wmnet * 22:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7015.magru.wmnet * 22:04 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7007.magru.wmnet * 21:29 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7014.magru.wmnet * 21:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7006.magru.wmnet * 21:13 dzahn@cumin2002: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 0:15:00 on gerrit.wikimedia.org with reason: reboot * 21:11 mutante: gerrit2003 (gerrit.wikimedia.org) - reboot for maintenance * 21:11 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on gerrit2003.wikimedia.org with reason: reboot * 20:56 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:56 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:56 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:55 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 20:48 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7013.magru.wmnet * 20:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7005.magru.wmnet * 20:41 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host phab1005.eqiad.wmnet with OS trixie * 20:28 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] (duration: 06m 47s) * 20:24 sbassett@deploy2003: sbassett: Continuing with deployment * 20:23 sbassett@deploy2003: sbassett: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:23 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on phab1005.eqiad.wmnet with reason: host reimage * 20:21 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] * 20:20 aokoth@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on phab1005.eqiad.wmnet with reason: host reimage * 20:12 jhuneidi@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] (duration: 07m 42s) * 20:07 jhuneidi@deploy2003: jhuneidi, priyankar22: Continuing with deployment * 20:06 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7012.magru.wmnet * 20:06 jhuneidi@deploy2003: jhuneidi, priyankar22: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:04 jhuneidi@deploy2003: Started scap sync-world: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] * 20:02 aokoth@cumin1003: START - Cookbook sre.hosts.reimage for host phab1005.eqiad.wmnet with OS trixie * 20:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7004.magru.wmnet * 20:00 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet * 19:57 aokoth@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet * 19:24 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7011.magru.wmnet * 19:19 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7003.magru.wmnet * 19:11 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] (duration: 08m 33s) * 19:07 jforrester@deploy2003: jforrester: Continuing with deployment * 19:04 jforrester@deploy2003: jforrester: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:02 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] * 18:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7010.magru.wmnet * 18:38 mutante: rotating phabricator-gerrit bot token (its-phabricator) * 18:18 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 17:44 swfrench@deploy2003: Finished scap sync-world: Deployment to pick up new production image (duration: 31m 44s) * 17:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7002.magru.wmnet * 17:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7009.magru.wmnet * 17:32 swfrench@deploy2003: swfrench: Continuing with deployment * 17:29 swfrench@deploy2003: swfrench: Deployment to pick up new production image synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:17 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2035: repooling after rack b5 maintenance * 17:16 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool es2035: repooling after rack b5 maintenance * 17:16 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2188: repooling after rack b5 maintenance * 17:12 swfrench@deploy2003: Started scap sync-world: Deployment to pick up new production image * 17:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7001.magru.wmnet * 16:57 swfrench-wmf: reprepro include php8.3_8.3.32-1+wmf12u2 into component/php83 for bookworm-wikimedia * 16:50 sukhe: pool cp2046 * 16:47 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4039.ulsfo.wmnet * 16:44 sukhe: sudo cumin -b31 "A:cp" "run-puppet-agent" * 16:33 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on contint1003.wikimedia.org with reason: reboot * 16:32 mutante: contint1003 - main CI server - rebooting * 16:31 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2188: repooling after rack b5 maintenance * 16:31 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2178: repooling after rack b5 maintenance * 16:29 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 16:28 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 16:28 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 16:28 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 16:18 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2014.codfw.wmnet * 16:18 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 16:17 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2014.codfw.wmnet * 16:10 mvernon@cumin1003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-thanos-proxies (exit_code=0) rolling restart_daemons on A:thanos-fe * 16:09 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 16:07 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp4039.ulsfo.wmnet * 16:07 mvernon@cumin1003: START - Cookbook sre.swift.roll-restart-reboot-swift-thanos-proxies rolling restart_daemons on A:thanos-fe * 16:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2014.codfw.wmnet with OS trixie * 15:56 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1014.eqiad.wmnet with OS trixie * 15:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2014.codfw.wmnet with reason: host reimage * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2014 * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2014 * 15:28 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2014 * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2014.codfw.wmnet 194.16.192.10.in-addr.arpa 4.9.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:28 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2014.codfw.wmnet 194.16.192.10.in-addr.arpa 4.9.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2014 - mvernon@cumin2003" * 15:28 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2014 - mvernon@cumin2003" * 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Apply title-related policies when selecting the name of the entity - kamila@cumin1003" * 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Apply title-related policies when selecting the name of the entity - kamila@cumin1003 * 15:22 kamila@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Apply title-related policies when selecting the name of the entity - kamila@cumin1003 * 15:22 kamila@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Apply title-related policies when selecting the name of the entity - kamila@cumin1003" * 15:20 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 15:20 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2014 * 15:20 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2014.codfw.wmnet with OS trixie * 15:19 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1014.eqiad.wmnet with OS trixie * 15:01 dancy@deploy2003: Installation of scap version "4.274.1" completed for 3 hosts * 15:00 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2177: repooling after rack b5 maintenance * 15:00 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2159: repooling after rack b5 maintenance * 14:59 dancy@deploy2003: Installing scap version "4.274.1" for 3 host(s) * 14:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2015.codfw.wmnet with OS trixie * 14:54 seanleong-wmde: Finished populateSitesTable for isvwiki ([[phab:T429939|T429939]]) * 14:53 javiermonton@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] (duration: 07m 35s) * 14:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1015.eqiad.wmnet with OS trixie * 14:49 javiermonton@deploy2003: javiermonton: Continuing with deployment * 14:48 javiermonton@deploy2003: javiermonton: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:46 javiermonton@deploy2003: Started scap sync-world: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] * 14:42 otto@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 14:41 otto@deploy2003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 14:41 otto@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 14:40 otto@deploy2003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 14:40 otto@deploy2003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 14:39 otto@deploy2003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 14:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2015.codfw.wmnet with reason: host reimage * 14:34 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1015.eqiad.wmnet with reason: host reimage * 14:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2015.codfw.wmnet with reason: host reimage * 14:30 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1015.eqiad.wmnet with reason: host reimage * 14:30 seanleong-wmde@deploy2003: mwscript-k8s job started: foreachwikiindblist wikidataclient extensions/Wikibase/lib/maintenance/populateSitesTable.php --force-protocol https # [[phab:T429939|T429939]] * 14:24 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling reboot on A:durum and not (A:durum-eqiad or A:durum-codfw or A:durum-esams) and A:durum * 14:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2015.codfw.wmnet with OS trixie * 14:15 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2016.codfw.wmnet with OS trixie * 14:14 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2159: repooling after rack b5 maintenance * 14:14 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1015.eqiad.wmnet with OS trixie * 14:12 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1016.eqiad.wmnet with OS trixie * 14:12 cmooney@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=pki,name=codfw * 14:12 sbisson@deploy2003: helmfile [codfw] DONE helmfile.d/services/cxserver: sync * 14:11 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2002.codfw.wmnet * 14:11 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2002.codfw.wmnet * 14:11 sbisson@deploy2003: helmfile [codfw] START helmfile.d/services/cxserver: sync * 14:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1005.wikimedia.org * 14:07 sbisson@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cxserver: sync * 14:07 sbisson@deploy2003: helmfile [eqiad] START helmfile.d/services/cxserver: sync * 14:05 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1005.wikimedia.org * 14:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader2005.wikimedia.org * 14:02 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2003.codfw.wmnet * 14:02 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2003.codfw.wmnet * 14:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=tegola-vector-tiles,name=codfw * 14:00 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=kartotherian,name=codfw * 14:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader2005.wikimedia.org * 13:58 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2016.codfw.wmnet with reason: host reimage * 13:57 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-ncredir (exit_code=0) rolling reboot on A:ncredir and A:ncredir * 13:57 sbisson@deploy2003: helmfile [staging] DONE helmfile.d/services/cxserver: sync * 13:56 sbisson@deploy2003: helmfile [staging] START helmfile.d/services/cxserver: sync * 13:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1016.eqiad.wmnet with reason: host reimage * 13:52 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:52 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:51 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2016.codfw.wmnet with reason: host reimage * 13:50 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1016.eqiad.wmnet with reason: host reimage * 13:49 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy (exit_code=0) rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 13:49 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling reboot on A:wikidough * 13:46 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-tcp-proxy (exit_code=0) rolling reboot on A:tcpproxy and A:tcpproxy * 13:43 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and not (A:durum-eqiad or A:durum-codfw or A:durum-esams) and A:durum * 13:42 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=97) rolling reboot on A:durum and A:durum * 13:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2011.codfw.wmnet * 13:36 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-reboot (exit_code=0) rolling reboot on A:dnsbox and A:ulsfo and (A:dnsbox) * 13:36 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns4004.wikimedia.org * 13:34 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1016.eqiad.wmnet with OS trixie * 13:34 topranks: reboot lsw1-b5-codfw to upgrade JunOS [[phab:T430918|T430918]] * 13:34 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2016.codfw.wmnet with OS trixie * 13:32 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2002.codfw.wmnet * 13:31 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2011.codfw.wmnet * 13:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2012.codfw.wmnet * 13:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2017.codfw.wmnet with OS trixie * 13:25 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1017.eqiad.wmnet with OS trixie * 13:24 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2012.codfw.wmnet * 13:22 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2002.codfw.wmnet * 13:22 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:22 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:22 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns4004.wikimedia.org * 13:21 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2005.codfw.wmnet * 13:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2013.codfw.wmnet * 13:18 elukey@dns1004: END - running authdns-update * 13:17 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2005.codfw.wmnet * 13:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2004.codfw.wmnet * 13:16 elukey@dns1004: START - running authdns-update * 13:16 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1046: es1046 after reimage * 13:14 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1029.eqiad.wmnet,service=s8 * 13:14 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1029.eqiad.wmnet,service=s5 * 13:13 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1029.eqiad.wmnet,service=s5 * 13:13 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1029.eqiad.wmnet,service=s8 * 13:13 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2004.codfw.wmnet * 13:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2013.codfw.wmnet * 13:11 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2014.codfw.wmnet * 13:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2003.codfw.wmnet * 13:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2017.codfw.wmnet with reason: host reimage * 13:07 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:07 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns4003.wikimedia.org * 13:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2003.codfw.wmnet * 13:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1067.eqiad.wmnet * 13:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1067.eqiad.wmnet * 13:06 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1067.eqiad.wmnet * 13:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm-test1001.wikimedia.org * 13:05 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2017.codfw.wmnet with reason: host reimage * 13:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1017.eqiad.wmnet with reason: host reimage * 13:04 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2014.codfw.wmnet * 13:03 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:02 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2188: codfw rack B5 depool for maintenance * 13:02 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_magru * 13:01 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2188: codfw rack B5 depool for maintenance * 13:01 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_magru * 13:01 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2178: codfw rack B5 depool for maintenance * 13:01 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm-test1001.wikimedia.org * 13:01 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2178: codfw rack B5 depool for maintenance * 13:01 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2177: codfw rack B5 depool for maintenance * 13:00 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2177: codfw rack B5 depool for maintenance * 12:59 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1068.eqiad.wmnet * 12:59 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1068.eqiad.wmnet * 12:58 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola-vector-tiles,name=codfw * 12:58 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2159: codfw rack B5 depool for maintenance * 12:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1017.eqiad.wmnet with reason: host reimage * 12:58 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola,name=codfw * 12:57 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=kartotherian,name=codfw * 12:57 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2159: codfw rack B5 depool for maintenance * 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 30 hosts with reason: lsw1-b5-codfw JunOS upgrade * 12:55 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lsw1-b5-codfw,lsw1-b5-codfw IPv6,lsw1-b5-codfw.mgmt,ssw1-a[1,8]-codfw.mgmt with reason: switch upgade lsw1-b5-codfw * 12:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet * 12:49 cmooney@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=pki,name=codfw * 12:49 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1067.eqiad.wmnet with OS trixie * 12:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet * 12:48 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2017.codfw.wmnet with OS trixie * 12:47 topranks: depool codfw pki in dns discovery ahead of lsw1-b5-codfw maintenance [[phab:T430918|T430918]] * 12:47 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns4003.wikimedia.org * 12:47 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and A:ulsfo and (A:dnsbox) * 12:47 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2018.codfw.wmnet * 12:45 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2018.codfw.wmnet * 12:45 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and A:durum * 12:45 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-tcp-proxy rolling reboot on A:tcpproxy and A:tcpproxy * 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host pki-root1002.eqiad.wmnet * 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1009.eqiad.wmnet * 12:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1009.eqiad.wmnet * 12:44 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 12:43 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-ncredir rolling reboot on A:ncredir and A:ncredir * 12:43 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling reboot on A:wikidough * 12:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2018.codfw.wmnet with OS trixie * 12:42 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1017.eqiad.wmnet with OS trixie * 12:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader2006.wikimedia.org * 12:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1018.eqiad.wmnet with OS trixie * 12:39 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1009.eqiad.wmnet * 12:38 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host pki-root1002.eqiad.wmnet * 12:38 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1009.eqiad.wmnet * 12:38 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1008.eqiad.wmnet * 12:38 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1008.eqiad.wmnet * 12:35 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader2006.wikimedia.org * 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1006.wikimedia.org * 12:33 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1008.eqiad.wmnet * 12:30 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1046: es1046 after reimage * 12:29 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host es1046.eqiad.wmnet with OS trixie * 12:29 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1006.wikimedia.org * 12:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test2005.wikimedia.org * 12:28 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1008.eqiad.wmnet * 12:27 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1007.eqiad.wmnet * 12:27 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1007.eqiad.wmnet * 12:27 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1067.eqiad.wmnet with reason: host reimage * 12:25 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1068.eqiad.wmnet with reason: vacuum overlarge container dbs * 12:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2018.codfw.wmnet with reason: host reimage * 12:24 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test2005.wikimedia.org * 12:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test1005.wikimedia.org * 12:23 Amir1: mwscript-k8s --follow --dblist=ores -- extensions/ORES/maintenance/PurgeScoreCache.php --model goodfaith --old ([[phab:T431159|T431159]]) * 12:22 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1007.eqiad.wmnet * 12:22 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test1005.wikimedia.org * 12:22 atsukoito: restarting pybal on lvs2013 `low-traffic` for https://gerrit.wikimedia.org/r/1310535 * 12:22 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1007.eqiad.wmnet * 12:21 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1006.eqiad.wmnet * 12:21 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1006.eqiad.wmnet * 12:20 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1018.eqiad.wmnet with reason: host reimage * 12:19 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2018.codfw.wmnet with reason: host reimage * 12:18 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1067.eqiad.wmnet with reason: host reimage * 12:16 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1006.eqiad.wmnet * 12:15 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1006.eqiad.wmnet * 12:15 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1005.eqiad.wmnet * 12:15 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1005.eqiad.wmnet * 12:15 atsukoito: restarting pybal on lvs2014 for https://gerrit.wikimedia.org/r/1310535 * 12:12 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1018.eqiad.wmnet with reason: host reimage * 12:11 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1005.eqiad.wmnet * 12:11 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1005.eqiad.wmnet * 12:10 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1004.eqiad.wmnet * 12:10 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1004.eqiad.wmnet * 12:09 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on es1046.eqiad.wmnet with reason: host reimage * 12:08 atsukoito: restarting pybal on lvs1019 `low-traffic` for https://gerrit.wikimedia.org/r/1310535 * 12:06 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1004.eqiad.wmnet * 12:06 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1004.eqiad.wmnet * 12:06 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1003.eqiad.wmnet * 12:06 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1003.eqiad.wmnet * 12:05 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on es1046.eqiad.wmnet with reason: host reimage * 12:05 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "set ml-serve1001 back to active state - cmooney@cumin1003" * 12:04 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "set ml-serve1001 back to active state - cmooney@cumin1003" * 12:04 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:02 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1003.eqiad.wmnet * 12:01 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1003.eqiad.wmnet * 12:01 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1002.eqiad.wmnet * 12:01 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1002.eqiad.wmnet * 12:01 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:59 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2018.codfw.wmnet with OS trixie * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1067 * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1067 * 11:59 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1067 * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1067.eqiad.wmnet 17.48.64.10.in-addr.arpa 7.1.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:59 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1067.eqiad.wmnet 17.48.64.10.in-addr.arpa 7.1.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1067 - blake@cumin1003" * 11:58 atsukoito: restarting pybal on lvs1018 `high-traffic2` for https://gerrit.wikimedia.org/r/1310535 * 11:57 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1002.eqiad.wmnet * 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2019.codfw.wmnet with OS trixie * 11:56 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1002.eqiad.wmnet * 11:56 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 11:56 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1018.eqiad.wmnet with OS trixie * 11:54 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 11:54 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:54 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1019.eqiad.wmnet with OS trixie * 11:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:49 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:49 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:49 aikochou@deploy2003: helmfile [codfw] DONE helmfile.d/services/changeprop: sync * 11:48 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host es1046.eqiad.wmnet with OS trixie * 11:48 aikochou@deploy2003: helmfile [codfw] START helmfile.d/services/changeprop: sync * 11:48 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310535 * 11:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1046: Reimage to Trixie * 11:44 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1046: Reimage to Trixie * 11:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5:00:00 on es1046.eqiad.wmnet with reason: Reimage to Trixie * 11:42 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:42 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:42 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 11:42 aikochou@deploy2003: helmfile [eqiad] DONE helmfile.d/services/changeprop: sync * 11:41 aikochou@deploy2003: helmfile [eqiad] START helmfile.d/services/changeprop: sync * 11:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2019.codfw.wmnet with reason: host reimage * 11:36 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] (duration: 09m 41s) * 11:36 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:36 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:35 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:35 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:32 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1019.eqiad.wmnet with reason: host reimage * 11:32 jforrester@deploy2003: jforrester, gengh: Continuing with deployment * 11:29 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2019.codfw.wmnet with reason: host reimage * 11:28 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1019.eqiad.wmnet with reason: host reimage * 11:28 jforrester@deploy2003: jforrester, gengh: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:26 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] * 11:20 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2003.codfw.wmnet * 11:20 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:19 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2003.codfw.wmnet * 11:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:12 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1019.eqiad.wmnet with OS trixie * 11:10 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2019.codfw.wmnet with OS trixie * 11:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1020.eqiad.wmnet with OS trixie * 11:10 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1067 - blake@cumin1003" * 11:09 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] (duration: 12m 12s) * 11:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2020.codfw.wmnet with OS trixie * 11:03 kharlan@deploy2003: kharlan: Continuing with deployment * 11:01 blake@cumin1003: START - Cookbook sre.dns.netbox * 11:01 kharlan@deploy2003: kharlan: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:57 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] * 10:55 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] (duration: 31m 40s) * 10:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1020.eqiad.wmnet with reason: host reimage * 10:52 marostegui@dns1004: START - running authdns-update * 10:49 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2020.codfw.wmnet with reason: host reimage * 10:49 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:48 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1020.eqiad.wmnet with reason: host reimage * 10:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2020.codfw.wmnet with reason: host reimage * 10:43 kharlan@deploy2003: kharlan: Continuing with deployment * 10:42 kharlan@deploy2003: kharlan: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:32 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1020.eqiad.wmnet with OS trixie * 10:29 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2159: Repooling after switchover * 10:29 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1067 * 10:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1021.eqiad.wmnet with OS trixie * 10:27 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1067.eqiad.wmnet with OS trixie * 10:27 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:27 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1067.eqiad.wmnet * 10:27 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:26 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1067.eqiad.wmnet * 10:26 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1067.eqiad.wmnet * 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2020.codfw.wmnet with OS trixie * 10:26 blake@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1055.eqiad.wmnet * 10:26 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1055.eqiad.wmnet * 10:26 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1055.eqiad.wmnet * 10:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2021.codfw.wmnet with OS trixie * 10:24 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] * 10:11 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1055.eqiad.wmnet with OS trixie * 10:09 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1021.eqiad.wmnet with reason: host reimage * 10:05 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2021.codfw.wmnet with reason: host reimage * 10:03 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310129 revert * 10:02 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1021.eqiad.wmnet with reason: host reimage * 10:01 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2021.codfw.wmnet with reason: host reimage * 09:58 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310129 * 09:50 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1055.eqiad.wmnet with reason: host reimage * 09:45 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1055.eqiad.wmnet with reason: host reimage * 09:45 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1021.eqiad.wmnet with OS trixie * 09:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1022.eqiad.wmnet with OS trixie * 09:44 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2159: Repooling after switchover * 09:44 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2021.codfw.wmnet with OS trixie * 09:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2022.codfw.wmnet with OS trixie * 09:31 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2159.codfw.wmnet * 09:28 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1055 * 09:28 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1055 * 09:27 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on ms-fe1022.eqiad.wmnet with reason: host reimage * 09:27 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1022.eqiad.wmnet with reason: host reimage * 09:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2022.codfw.wmnet with reason: host reimage * 09:21 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2022.codfw.wmnet with reason: host reimage * 09:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2159: Rebooting db2159.codfw.wmnet * 09:20 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2159: Rebooting db2159.codfw.wmnet * 09:18 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 09:18 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 09:18 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 09:17 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 09:16 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2159.codfw.wmnet * 09:13 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b] (thin): Regular analytics weekly train THIN [analytics/refinery@ad6e05b8] (duration: 02m 07s) * 09:11 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b] (thin): Regular analytics weekly train THIN [analytics/refinery@ad6e05b8] * 09:10 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1022.eqiad.wmnet with OS trixie * 09:07 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1023.eqiad.wmnet with OS trixie * 09:06 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b]: Regular analytics weekly train [analytics/refinery@ad6e05b8] (duration: 04m 49s) * 09:04 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2022.codfw.wmnet with OS trixie * 09:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2023.codfw.wmnet with OS trixie * 09:01 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b]: Regular analytics weekly train [analytics/refinery@ad6e05b8] * 09:01 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@ad6e05b8] (duration: 02m 01s) * 09:00 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1055 * 09:00 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1055.eqiad.wmnet 50.32.64.10.in-addr.arpa 0.5.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:00 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1055.eqiad.wmnet 50.32.64.10.in-addr.arpa 0.5.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:00 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:00 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1055 - blake@cumin1003" * 09:00 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1055 - blake@cumin1003" * 08:59 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@ad6e05b8] * 08:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2159 [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94811 and previous config saved to /var/cache/conftool/dbconfig/20260714-085624-cwilliams.json * 08:55 blake@cumin1003: START - Cookbook sre.dns.netbox * 08:55 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1055 * 08:54 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1055.eqiad.wmnet with OS trixie * 08:54 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1055.eqiad.wmnet * 08:53 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1055.eqiad.wmnet * 08:53 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1055.eqiad.wmnet * 08:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2220 to s7 primary [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94810 and previous config saved to /var/cache/conftool/dbconfig/20260714-085239-cwilliams.json * 08:51 cezmunsta: Starting s7 codfw failover from db2159 to db2220 - [[phab:T430920|T430920]] * 08:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1023.eqiad.wmnet with reason: host reimage * 08:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2220 with weight 0 [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94809 and previous config saved to /var/cache/conftool/dbconfig/20260714-084553-cwilliams.json * 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s7 [[phab:T430920|T430920]] * 08:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2023.codfw.wmnet with reason: host reimage * 08:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1023.eqiad.wmnet with reason: host reimage * 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2023.codfw.wmnet with reason: host reimage * 08:34 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:34 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:29 marostegui@dns1004: END - running authdns-update * 08:29 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox-dev2003.codfw.wmnet * 08:27 marostegui@dns1004: START - running authdns-update * 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker2*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2009.codfw.wmnet * 08:26 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2009.codfw.wmnet * 08:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1023.eqiad.wmnet with OS trixie * 08:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox-dev2003.codfw.wmnet * 08:24 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:24 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:24 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2023.codfw.wmnet with OS trixie * 08:24 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1029.eqiad.wmnet with reason: reboot * 08:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1027.eqiad.wmnet with reason: reboot * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:21 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2009.codfw.wmnet * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:20 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2009.codfw.wmnet * 08:20 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2008.codfw.wmnet * 08:20 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2008.codfw.wmnet * 08:15 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2008.codfw.wmnet * 08:14 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2008.codfw.wmnet * 08:14 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2007.codfw.wmnet * 08:14 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2007.codfw.wmnet * 08:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1024.eqiad.wmnet with OS trixie * 08:12 elukey@cumin1003: END (PASS) - Cookbook sre.pki.restart-reboot (exit_code=0) rolling reboot on P<nowiki>{</nowiki>pki*<nowiki>}</nowiki> and (A:pki) * 08:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 08:10 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 08:09 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2007.codfw.wmnet * 08:08 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2007.codfw.wmnet * 08:08 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2006.codfw.wmnet * 08:08 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2006.codfw.wmnet * 08:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2024.codfw.wmnet with OS trixie * 08:03 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2006.codfw.wmnet * 08:02 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2006.codfw.wmnet * 08:02 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2005.codfw.wmnet * 08:02 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2005.codfw.wmnet * 07:58 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2005.codfw.wmnet * 07:58 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2005.codfw.wmnet * 07:57 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2004.codfw.wmnet * 07:57 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2004.codfw.wmnet * 07:54 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki.discovery.wmnet. on all recursors * 07:54 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache pki.discovery.wmnet. on all recursors * 07:53 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2004.codfw.wmnet * 07:53 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2004.codfw.wmnet * 07:53 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2003.codfw.wmnet * 07:53 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2003.codfw.wmnet * 07:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1024.eqiad.wmnet with reason: host reimage * 07:49 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki.discovery.wmnet. on all recursors * 07:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2003.codfw.wmnet * 07:49 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache pki.discovery.wmnet. on all recursors * 07:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2024.codfw.wmnet with reason: host reimage * 07:48 elukey@cumin1003: START - Cookbook sre.pki.restart-reboot rolling reboot on P<nowiki>{</nowiki>pki*<nowiki>}</nowiki> and (A:pki) * 07:46 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1024.eqiad.wmnet with reason: host reimage * 07:45 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2024.codfw.wmnet with reason: host reimage * 07:45 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2003.codfw.wmnet * 07:44 elukey@cumin1003: END (PASS) - Cookbook sre.misc-clusters.restart-reboot-config-master (exit_code=0) rolling reboot on P<nowiki>{</nowiki>config-master*<nowiki>}</nowiki> and (A:config-master or A:config-master-eqiad or A:config-master-codfw) * 07:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2002.codfw.wmnet * 07:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2002.codfw.wmnet * 07:39 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2002.codfw.wmnet * 07:39 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) config-master.discovery.wmnet. on all recursors * 07:39 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache config-master.discovery.wmnet. on all recursors * 07:39 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2002.codfw.wmnet * 07:39 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker2*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl200*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl2003.codfw.wmnet * 07:36 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl2003.codfw.wmnet * 07:35 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) config-master.discovery.wmnet. on all recursors * 07:35 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache config-master.discovery.wmnet. on all recursors * 07:34 elukey@cumin1003: START - Cookbook sre.misc-clusters.restart-reboot-config-master rolling reboot on P<nowiki>{</nowiki>config-master*<nowiki>}</nowiki> and (A:config-master or A:config-master-eqiad or A:config-master-codfw) * 07:31 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl2003.codfw.wmnet * 07:31 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl2003.codfw.wmnet * 07:31 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl2002.codfw.wmnet * 07:31 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl2002.codfw.wmnet * 07:29 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1024.eqiad.wmnet with OS trixie * 07:28 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2024.codfw.wmnet with OS trixie * 07:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl2002.codfw.wmnet * 07:26 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl2002.codfw.wmnet * 07:26 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl200*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 07:26 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 07:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 06:50 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lists1004.wikimedia.org * 06:44 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host lists1004.wikimedia.org * 06:25 marostegui@dns1004: END - running authdns-update * 06:23 marostegui@dns1004: START - running authdns-update * 06:22 marostegui@dns1004: END - running authdns-update * 06:20 marostegui@dns1004: START - running authdns-update * 06:17 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1026.eqiad.wmnet with reason: reboot * 06:04 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: sync * 06:04 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: sync * 06:03 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync * 06:03 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync * 06:02 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync * 06:01 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync * 06:01 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync * 06:00 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync * 05:59 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:59 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:40 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:39 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:26 marostegui@dns1004: END - running authdns-update * 05:24 marostegui@dns1004: START - running authdns-update * 05:24 marostegui@dns1004: START - running authdns-update * 05:13 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1004.wikimedia.org * 05:07 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1004.wikimedia.org * 04:01 mwpresync@deploy2003: Pruned MediaWiki: 1.47.0-wmf.8 (duration: 01m 07s) * 03:39 mwpresync@deploy2003: Finished scap sync-world: testwikis to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] (duration: 36m 01s) * 03:03 mwpresync@deploy2003: Started scap sync-world: testwikis to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 29s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-13 == * 23:33 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1064.eqiad.wmnet * 23:33 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1064.eqiad.wmnet * 23:08 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1064.eqiad.wmnet with reason: vacuum overlarge container dbs * 23:06 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1069.eqiad.wmnet * 23:06 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1069.eqiad.wmnet * 22:34 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1069.eqiad.wmnet with reason: vacuum overlarge container dbs * 21:18 maryum: Deployed security fix for [[phab:T321092|T321092]] * 20:28 swfrench-wmf: reprepro include etcd-mirror_0.0.12-1+deb13u1 into main for trixie-wikimedia - [[phab:T424266|T424266]] * 20:26 swfrench-wmf: reprepro include etcd-mirror_0.0.12-1+deb12u1 into main for bookworm-wikimedia - [[phab:T428495|T428495]] * 20:23 dancy@deploy2003: Finished scap sync-world: Testing [[phab:T431635|T431635]] (duration: 03m 36s) * 20:19 dancy@deploy2003: Started scap sync-world: Testing [[phab:T431635|T431635]] * 20:18 dancy@deploy2003: Installation of scap version "4.274.0" completed for 3 hosts * 20:16 dancy@deploy2003: Installing scap version "4.274.0" for 3 host(s) * 20:12 kemayo@deploy2003: Finished scap sync-world: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] (duration: 08m 25s) * 20:07 kemayo@deploy2003: soda, esanders, kemayo: Continuing with deployment * 20:05 kemayo@deploy2003: soda, esanders, kemayo: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there * 20:04 kemayo@deploy2003: Started scap sync-world: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] * 18:22 cwhite: lvextend vg0/srv +500g on centrallog hosts * 18:19 cdobbins@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS trixie * 17:46 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1071.eqiad.wmnet * 17:46 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1071.eqiad.wmnet * 17:13 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1071.eqiad.wmnet with reason: vacuum overlarge container dbs * 17:07 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1065.eqiad.wmnet * 17:07 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1065.eqiad.wmnet * 17:06 dzahn@dns1006: END - running authdns-update * 17:04 dzahn@dns1006: START - running authdns-update * 17:01 dzahn@dns1006: END - running authdns-update * 16:59 dzahn@dns1006: START - running authdns-update * 16:51 dancy@deploy2003: Finished scap sync-world: testing [[phab:T428971|T428971]] (duration: 03m 37s) * 16:47 dancy@deploy2003: Started scap sync-world: testing [[phab:T428971|T428971]] * 16:45 atsukoito: restarting pybal on lvs1019 to flush IP address for `cirrussearch1122.eqiad.wmnet` after moving the vlan [[phab:T431311|T431311]] * 16:42 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:42 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:42 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:42 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:42 Amir1: mwscript-k8s --follow --dblist=ores -- extensions/ORES/maintenance/PurgeScoreCache.php --model damaging --old ([[phab:T431159|T431159]]) * 16:34 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Pool test * 16:34 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 16:34 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 16:34 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Pool test * 16:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Depool test * 16:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 16:33 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 16:33 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Depool test * 16:31 dancy@deploy2003: Installation of scap version "4.273.0" completed for 159 hosts * 16:29 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1065.eqiad.wmnet with reason: vacuum overlarge container dbs * 16:27 dancy@deploy2003: Installing scap version "4.273.0" for 159 host(s) * 16:27 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics-external: sync * 16:27 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics-external: sync * 16:26 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics-external: sync * 16:26 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics-external: sync * 16:22 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync * 16:21 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync * 16:21 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: sync * 16:21 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: sync * 16:19 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync * 16:19 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync * 16:18 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync * 16:17 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync * 15:59 atsukoito: restarting pybal on lvs1018 for https://gerrit.wikimedia.org/r/1310117 * 15:55 aikochou@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 15:50 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310117 * 15:46 aikochou@deploy2003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 15:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host kafka-logging1006.eqiad.wmnet * 15:43 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host kafka-logging1006.eqiad.wmnet * 15:41 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host ganeti-test[2001-2003].codfw.wmnet * 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host ganeti-test[2001-2003].codfw.wmnet * 15:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host netbox1003.eqiad.wmnet * 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host netbox1003.eqiad.wmnet * 15:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host netbox2003.codfw.wmnet * 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host netbox2003.codfw.wmnet * 15:36 sukhe: restart pybal on lvs1020 * 15:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet * 15:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet * 15:08 btullis@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:06 btullis@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 15:01 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:01 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:35 cdobbins@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 14:34 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:33 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:33 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:32 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:29 cdobbins@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 14:28 marostegui@dns1004: END - running authdns-update * 14:28 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:27 marostegui@dns1004: START - running authdns-update * 14:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1023.eqiad.wmnet with reason: reboot * 14:18 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2009.codfw.wmnet with OS trixie * 14:14 swfrench-wmf: start rolling run-puppet-agent on A:cp for ATS config change - [[phab:T428909|T428909]] [[phab:T431838|T431838]] * 14:05 swfrench-wmf: disable-puppet on A:cp for ATS config change - [[phab:T428909|T428909]] [[phab:T431838|T431838]] * 14:05 cdobbins@cumin2002: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie * 14:02 marostegui@dns1004: END - running authdns-update * 14:00 marostegui@dns1004: START - running authdns-update * 14:00 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1070.eqiad.wmnet * 14:00 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1070.eqiad.wmnet * 13:58 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2009.codfw.wmnet with reason: host reimage * 13:57 cdobbins@cumin2002: conftool action : set/pooled=no; selector: name=dns7002.* * 13:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2009.codfw.wmnet with reason: host reimage * 13:48 rscout@deploy2003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply * 13:48 rscout@deploy2003: helmfile [eqiad] START helmfile.d/services/miscweb: apply * 13:48 rscout@deploy2003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply * 13:47 rscout@deploy2003: helmfile [codfw] START helmfile.d/services/miscweb: apply * 13:40 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:33 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2009.codfw.wmnet with OS trixie * 13:30 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1070.eqiad.wmnet with reason: vacuum overlarge container dbs * 13:28 aude@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] (duration: 11m 12s) * 13:23 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:22 aude@deploy2003: aikochou, javiermonton, aude, gkm563: Continuing with deployment * 13:22 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:19 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:19 aude@deploy2003: aikochou, javiermonton, aude, gkm563: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] synced to the testservers * 13:17 aude@deploy2003: Started scap sync-world: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] * 13:01 ladsgroup@deploy2003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 13:01 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:00 ladsgroup@deploy2003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 12:59 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:52 ladsgroup@deploy2003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 12:51 ladsgroup@deploy2003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 12:48 atsuko@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 12:48 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 12:47 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2008.codfw.wmnet with OS trixie * 12:47 atsuko@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 12:47 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply * 12:47 atsuko@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:46 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 12:45 atsuko@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:45 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply * 12:45 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:44 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:43 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] (duration: 07m 02s) * 12:38 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 12:37 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:36 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] * 12:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2008.codfw.wmnet with reason: host reimage * 12:23 Msz2001: Deployed changes to private code for Suggested Investigations * 12:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2008.codfw.wmnet with reason: host reimage * 12:20 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:19 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:17 atsuko@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 12:17 atsuko@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 12:16 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:15 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] (duration: 07m 14s) * 12:10 mszwarc@deploy2003: mszwarc: Continuing with deployment * 12:09 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:07 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] * 12:04 mszwarc@deploy2003: sync-world aborted: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] (duration: 00m 29s) * 12:03 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] * 12:01 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2008.codfw.wmnet with OS trixie * 12:00 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] (duration: 07m 37s) * 11:55 zabe@deploy2003: zabe: Continuing with deployment * 11:54 zabe@deploy2003: zabe: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:52 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] * 11:51 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:43 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:35 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:34 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:33 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:30 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:28 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:27 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:17 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2007.codfw.wmnet with OS trixie * 11:09 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] * 11:06 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=s8 * 11:00 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=x3 * 11:00 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=s5 * 10:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2007.codfw.wmnet with reason: host reimage * 10:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2007.codfw.wmnet with reason: host reimage * 10:51 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host dse-k8s-worker1023 * 10:50 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host dse-k8s-worker1023 * 10:44 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dse-k8s-worker1023 * 10:43 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host dse-k8s-worker1023 * 10:42 marostegui@cumin1003: dbctl commit (dc=all): 'Change x4 masters [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P94804 and previous config saved to /var/cache/conftool/dbconfig/20260713-104248-marostegui.json * 10:37 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:37 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:35 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:35 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:34 atsuko@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 10:34 atsuko@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 10:33 marostegui@cumin1003: dbctl commit (dc=all): 'Push x4 initial dbctl config [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P94803 and previous config saved to /var/cache/conftool/dbconfig/20260713-103259-marostegui.json * 10:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2007.codfw.wmnet with OS trixie * 09:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2006.codfw.wmnet with OS trixie * 09:42 marostegui@dns1004: END - running authdns-update * 09:40 marostegui@dns1004: START - running authdns-update * 09:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2006.codfw.wmnet with reason: host reimage * 09:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2006.codfw.wmnet with reason: host reimage * 09:06 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1024.eqiad.wmnet with reason: reboot * 09:06 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:01 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2006.codfw.wmnet with OS trixie * 08:44 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 08:43 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 08:43 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:42 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:42 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:42 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:41 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 08:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup2004.codfw.wmnet * 08:38 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:33 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host db1208.eqiad.wmnet * 08:30 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=x3 * 08:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2005.codfw.wmnet with OS trixie * 08:28 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup2004.codfw.wmnet * 08:28 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup2003.codfw.wmnet * 08:24 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1039: Repooling after testing * 08:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on clouddb1016.eqiad.wmnet with reason: cloning * 08:23 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s5 * 08:23 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s8 * 08:21 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 08:21 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 08:17 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup2003.codfw.wmnet * 08:17 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1004.eqiad.wmnet * 08:14 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1208.eqiad.wmnet * 08:11 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host phab1005.eqiad.wmnet * 08:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2005.codfw.wmnet with reason: host reimage * 08:07 marostegui@dns1004: END - running authdns-update * 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1004.eqiad.wmnet * 08:07 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1003.eqiad.wmnet * 08:07 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:05 marostegui@dns1004: START - running authdns-update * 08:05 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2005.codfw.wmnet with reason: host reimage * 08:05 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 08:05 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host phab1005.eqiad.wmnet * 08:05 marostegui@dns1004: START - running authdns-update * 08:05 marostegui@dns1004: START - running authdns-update * 08:05 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 08:04 marostegui@dns1004: START - running authdns-update * 08:00 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit1003.wikimedia.org * 07:58 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1003.eqiad.wmnet * 07:58 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1002-dev.eqiad.wmnet * 07:58 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:58 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:54 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1002-dev.eqiad.wmnet * 07:54 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1001-dev.eqiad.wmnet * 07:54 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit1003.wikimedia.org * 07:53 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 07:53 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:52 Msz2001: UTC morning backport+config window done * 07:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2005.codfw.wmnet with OS trixie * {{safesubst:SAL entry|1=07:50 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark (T429943}} * 07:49 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1001-dev.eqiad.wmnet * 07:46 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:46 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:45 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 07:45 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:44 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 07:44 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:43 mszwarc@deploy2003: mszwarc, danielyepezgarces, anzx: Continuing with deployment * {{safesubst:SAL entry|1=07:39 mszwarc@deploy2003: mszwarc, danielyepezgarces, anzx: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark}} * 07:39 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1039: Repooling after testing * {{safesubst:SAL entry|1=07:36 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark (T429943)}} * 07:35 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] (duration: 30m 03s) * 07:25 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit2002.wikimedia.org * 07:22 mszwarc@deploy2003: mszwarc: Continuing with deployment * 07:21 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:19 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit2002.wikimedia.org * 07:15 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aphlict1002.eqiad.wmnet * 07:11 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host aphlict1002.eqiad.wmnet * 07:08 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2003.wikimedia.org * 07:05 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] * 07:02 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2003.wikimedia.org * 07:02 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2002.wikimedia.org * 06:55 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2002.wikimedia.org * 06:55 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1003.wikimedia.org * 06:49 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1003.wikimedia.org * 06:34 marostegui: Drop m5 ipoid database [[phab:T431007|T431007]] * 06:29 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1027.eqiad.wmnet with reason: reboot * 06:24 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1028.eqiad.wmnet with reason: reboot * 06:21 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1025.eqiad.wmnet with reason: reboot * 06:17 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1022.eqiad.wmnet with reason: reboot * 06:03 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on dbproxy[2005-2008].codfw.wmnet with reason: reboot * 05:37 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1217,1228].eqiad.wmnet with reason: cloning * 05:11 marostegui: Drop users_to_rename table [[phab:T431842|T431842]] * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-12 == * 16:01 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2209 [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94792 and previous config saved to /var/cache/conftool/dbconfig/20260712-160124-marostegui.json * 15:58 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2205 to s3 primary [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94791 and previous config saved to /var/cache/conftool/dbconfig/20260712-155853-marostegui.json * 15:58 marostegui: Starting s3 codfw emergency failover from db2209 to db2205 - [[phab:T431950|T431950]] * 15:51 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2205 with weight 0 [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94790 and previous config saved to /var/cache/conftool/dbconfig/20260712-155135-marostegui.json * 15:51 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Primary switchover s3 [[phab:T431950|T431950]] * 02:01 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 01m 17s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-11 == * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 26s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-10 == * 19:12 jhathaway@dns1004: END - running authdns-update * 19:10 jhathaway@dns1004: START - running authdns-update * 18:23 mutante: vrts2002 rebooting (not the active host) * 18:21 mutante: lists2001, phab2003 - rebooting (not the active hosts) * 18:16 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on A:lvs-high-traffic2-codfw * 18:15 swfrench@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on A:lvs-high-traffic2-codfw * 17:15 mutante: [doc1004:~] $ sudo systemctl start rsync-doc-host-data-sync ([[phab:T431856|T431856]]) * 17:09 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1004.eqiad.wmnet * 17:08 jhathaway@dns1004: END - running authdns-update * 17:07 jhathaway@dns1004: START - running authdns-update * 17:06 jhathaway: depooling puppetserver1002, cause of errors is still unknown * 17:03 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1004.eqiad.wmnet * 16:57 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1003.eqiad.wmnet * 16:51 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1003.eqiad.wmnet * 16:48 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 16:48 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2004.codfw.wmnet * 16:42 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2004.codfw.wmnet * 16:41 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2003.codfw.wmnet * 16:35 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2003.codfw.wmnet * 16:33 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2002.codfw.wmnet * 16:27 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2002.codfw.wmnet * 16:25 mutante: gitlab-runners (production) rebooting cluster one by one * 16:17 mutante: etherpad1004/etherpad2002 - (etherpad.wikimedia.org) - rebooting * 16:13 mutante: doc1004/doc2003 (doc.wikimedia.org backends) - rebooting * 16:02 mutante: releases1003/releases2003 (releases.wikimedia.org backends) - rebooting for maintenance * 15:26 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2007-dev.codfw.wmnet * 15:19 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2007-dev.codfw.wmnet * 15:14 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host cloudcephosd2007-dev.codfw.wmnet * 15:14 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2007-dev.codfw.wmnet * 15:14 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host cloudcephosd2006-dev.codfw.wmnet * 15:07 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2006-dev.codfw.wmnet * 15:07 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2005-dev.codfw.wmnet * 14:59 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2005-dev.codfw.wmnet * 14:59 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2004-dev.codfw.wmnet * 14:53 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2004-dev.codfw.wmnet * 14:53 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2007-dev.codfw.wmnet * 14:51 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1054.eqiad.wmnet * 14:51 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1054.eqiad.wmnet * 14:51 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1054.eqiad.wmnet * 14:47 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2007-dev.codfw.wmnet * 14:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2006-dev.codfw.wmnet * 14:41 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2006-dev.codfw.wmnet * 14:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2005-dev.codfw.wmnet * 14:37 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2005-dev.codfw.wmnet * 14:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2005-dev.codfw.wmnet * 14:29 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2005-dev.codfw.wmnet * 14:29 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2006-dev.codfw.wmnet * 14:21 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2006-dev.codfw.wmnet * 14:21 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2010-dev.codfw.wmnet * 14:15 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2010-dev.codfw.wmnet * 14:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudgw2004-dev.codfw.wmnet * 14:10 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1054.eqiad.wmnet with OS trixie * 14:09 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudgw2004-dev.codfw.wmnet * 14:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudgw2003-dev.codfw.wmnet * 14:02 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudgw2003-dev.codfw.wmnet * 14:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2004-dev.codfw.wmnet * 13:53 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2004-dev.codfw.wmnet * 13:53 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2003-dev.codfw.wmnet * 13:48 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:44 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2003-dev.codfw.wmnet * 13:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2002-dev.codfw.wmnet * 13:42 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:41 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:41 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:37 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2002-dev.codfw.wmnet * 13:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudidp2001-dev.codfw.wmnet * 13:33 blake@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker1054.eqiad.wmnet with reason: host reimage * 13:33 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudidp2001-dev.codfw.wmnet * 13:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudnet2006-dev.codfw.wmnet * 13:26 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudnet2006-dev.codfw.wmnet * 13:26 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudnet2005-dev.codfw.wmnet * 13:23 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1054.eqiad.wmnet with reason: host reimage * 13:18 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudnet2005-dev.codfw.wmnet * 13:18 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudservices2005-dev.codfw.wmnet * 13:12 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudservices2005-dev.codfw.wmnet * 13:11 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudservices2004-dev.codfw.wmnet * 13:08 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudservices2004-dev.codfw.wmnet * 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudweb2002-dev.wikimedia.org * 13:05 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 13:05 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1054 * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1054 * 13:04 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1054 * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1054.eqiad.wmnet 49.32.64.10.in-addr.arpa 9.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:04 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1054.eqiad.wmnet 49.32.64.10.in-addr.arpa 9.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1054 - blake@cumin1003" * 13:04 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1054 - blake@cumin1003" * 13:01 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudweb2002-dev.wikimedia.org * 13:00 blake@cumin1003: START - Cookbook sre.dns.netbox * 12:59 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1054 * 12:57 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1054.eqiad.wmnet with OS trixie * 12:57 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1054.eqiad.wmnet * 12:56 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1054.eqiad.wmnet * 12:56 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1054.eqiad.wmnet * 12:47 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS trixie * 12:44 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:39 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 12:39 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 12:38 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:37 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:14 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:10 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:08 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:07 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:00 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:00 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:51 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:49 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:48 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:47 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:44 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:32 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2001.codfw.wmnet * 11:32 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1053.eqiad.wmnet * 11:32 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2001.codfw.wmnet * 11:32 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1053.eqiad.wmnet * 11:32 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1053.eqiad.wmnet * 11:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker2001.codfw.wmnet * 11:31 cgoubert@cumin1003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker2001.codfw.wmnet * 11:31 cgoubert@cumin1003: END (FAIL) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=1) rolling reimage on P<nowiki>{</nowiki>wikikube-worker2001*<nowiki>}</nowiki> and (A:wikikube-master-codfw or A:wikikube-worker-codfw) * 11:30 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:30 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:21 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 18 hosts with reason: reboot & upgrade * 11:20 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker2001.codfw.wmnet with OS trixie * 11:16 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:15 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:14 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:14 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 11 hosts * 11:14 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 11 hosts * 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:08 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 11:02 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1053.eqiad.wmnet with OS trixie * 11:01 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:58 cgoubert@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 10:57 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:38 cgoubert@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker2001.codfw.wmnet with OS trixie * 10:38 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2001.codfw.wmnet * 10:38 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2001.codfw.wmnet * 10:38 cgoubert@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on P<nowiki>{</nowiki>wikikube-worker2001*<nowiki>}</nowiki> and (A:wikikube-master-codfw or A:wikikube-worker-codfw) * 10:35 topranks: adjust IBGP outbound policy on lsw1-e2-codfw [[phab:T423430|T423430]] towards ssw1-e1-codfw * 10:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 cgoubert@cumin1003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:27 cgoubert@cumin1003: END (FAIL) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=1) rolling reimage on A:wikikube-worker-codfw * 10:27 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker2001.codfw.wmnet with OS bookworm * 10:25 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 10:24 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:24 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:15 cgoubert@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 10:11 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 10:11 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 10:08 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:08 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:07 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 10:06 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 10:00 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:55 cgoubert@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker2001.codfw.wmnet with OS bookworm * 09:55 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2005-2006,2011-2012].codfw.wmnet * 09:55 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2005-2006,2011-2012].codfw.wmnet * 09:51 cgoubert@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on A:wikikube-worker-codfw * 09:41 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1053.eqiad.wmnet with reason: host reimage * 09:37 topranks: apply new IBGP outbound policy on lsw1-e2-codfw [[phab:T423430|T423430]] * 09:36 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:36 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1053.eqiad.wmnet with reason: host reimage * 09:16 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1053 * 09:16 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1053 * 09:15 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1053 * 09:15 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1053.eqiad.wmnet 48.32.64.10.in-addr.arpa 8.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:15 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1053.eqiad.wmnet 48.32.64.10.in-addr.arpa 8.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:15 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:15 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1053 - blake@cumin1003" * 09:15 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1053 - blake@cumin1003" * 09:11 blake@cumin1003: START - Cookbook sre.dns.netbox * 09:11 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1053 * 09:08 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1053.eqiad.wmnet with OS trixie * 09:08 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1053.eqiad.wmnet * 09:08 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1053.eqiad.wmnet * 09:08 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1053.eqiad.wmnet * 09:04 brouberol@dns1004: END - running authdns-update * 09:03 brouberol@dns1004: START - running authdns-update * 08:41 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e] (thin): Regular analytics weekly train THIN [analytics/refinery@1abf22ea] (duration: 02m 11s) * 08:38 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e] (thin): Regular analytics weekly train THIN [analytics/refinery@1abf22ea] * 08:38 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e]: Regular analytics weekly train [analytics/refinery@1abf22ea] (duration: 05m 17s) * 08:38 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:34 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:33 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e]: Regular analytics weekly train [analytics/refinery@1abf22ea] * 08:32 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@1abf22ea] (duration: 02m 03s) * 08:30 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@1abf22ea] * 08:30 JavierMonton: Deploying Refinery at {{Gerrit|1abf22ea}} for changes 1308121/T427068 1306491/T430020 and {{Gerrit|1308190}} * 08:29 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:29 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 08:24 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:24 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 08:18 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:18 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 08:00 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db[2183-2184].codfw.wmnet * 08:00 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for db[2183-2184].codfw.wmnet * 07:52 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:52 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 07:49 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 11 hosts with reason: reboot & upgrade * 07:47 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:47 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 07:44 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:44 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 07:23 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 10 hosts * 07:23 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 10 hosts * 06:45 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 10 hosts with reason: reboot & upgrade * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 41s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-09 == * 23:33 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] (duration: 13m 26s) * 23:29 ladsgroup@deploy2003: ladsgroup, jdlrobson: Continuing with deployment * 23:22 ladsgroup@deploy2003: ladsgroup, jdlrobson: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:20 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] * 22:57 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1165.eqiad.wmnet * 22:56 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1165.eqiad.wmnet * 22:56 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1165.eqiad.wmnet * 22:45 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1165.eqiad.wmnet with OS trixie * 22:38 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 22:37 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 22:37 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 22:37 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:37 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 22:25 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1165.eqiad.wmnet with reason: host reimage * 22:17 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1165.eqiad.wmnet with reason: host reimage * 22:13 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 22:12 rzl@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 22:04 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 22:04 rzl@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1165 * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1165 * 22:02 jasmine@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1165 * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1165.eqiad.wmnet 115.48.64.10.in-addr.arpa 5.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:02 jasmine@cumin2002: START - Cookbook sre.dns.wipe-cache wikikube-worker1165.eqiad.wmnet 115.48.64.10.in-addr.arpa 5.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1165 - jasmine@cumin2002" * 22:02 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1165 - jasmine@cumin2002" * 22:02 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 21:57 jasmine@cumin2002: START - Cookbook sre.dns.netbox * 21:55 jasmine@cumin2002: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1165 * 21:54 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-worker1165.eqiad.wmnet with OS trixie * 21:54 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 21:54 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1165.eqiad.wmnet * 21:53 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 21:53 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1165.eqiad.wmnet * 21:53 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1165.eqiad.wmnet * 21:53 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 21:47 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 21:45 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 21:43 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 21:43 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 21:42 maryum: Deploy fix for [[phab:T431684|T431684]] * 21:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 21:27 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] (duration: 34m 14s) * 21:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs1002 * 21:23 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs1002 * 21:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS trixie * 21:22 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 21:20 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 22s) * 21:20 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 21:16 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 21:16 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 21:15 ladsgroup@deploy2003: ladsgroup: Continuing with deployment * 21:13 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2002.codfw.wmnet with OS bookworm * 21:11 ladsgroup@deploy2003: ladsgroup: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:08 ladsgroup@cumin1003: END (PASS) - Cookbook sre.wikireplicas.update-views (exit_code=0) * 21:07 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 6 hosts with reason: reboots * 20:54 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecycle work - bking@cumin2003 * 20:53 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:53 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] * 20:51 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) * 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 20:47 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecycle work - bking@cumin2003 * 20:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 20:41 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:41 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) * 20:40 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host relforge1008.eqiad.wmnet * 20:40 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1009.eqiad.wmnet with OS trixie * 20:33 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:32 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:32 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:31 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:31 ladsgroup@cumin1003: END (PASS) - Cookbook sre.wikireplicas.update-views (exit_code=0) * 20:29 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1008.eqiad.wmnet * 20:24 rzl@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 20:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2002.codfw.wmnet with OS bookworm * 20:23 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host relforge1008.eqiad.wmnet * 20:23 rzl@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 20:23 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1008.eqiad.wmnet * 20:22 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:22 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:21 bking@cumin2003: END (ERROR) - Cookbook sre.elasticsearch.rolling-operation (exit_code=97) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:21 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1009.eqiad.wmnet with reason: host reimage * 20:16 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:15 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1009.eqiad.wmnet with reason: host reimage * 20:12 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) * 20:02 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 19:55 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1009.eqiad.wmnet with OS trixie * 19:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 19:43 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 19:30 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 19:28 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 19:27 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 19:25 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 18:42 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 18:41 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 18:16 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 18:15 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 17:45 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for doh5004.wikimedia.org * 17:45 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for doh5004.wikimedia.org * 17:38 ladsgroup@deploy2003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 17:35 ladsgroup@deploy2003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 17:29 ladsgroup@deploy2003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 17:26 ladsgroup@deploy2003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 17:09 mutante: zuul[12]00[123] - rebooting for maintenance * 17:09 ebernhardson: start full in-place reindex of eqiad cirrussearch cluster * 17:08 dzahn@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-cluster (exit_code=99) * 17:08 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-cluster * 17:03 ebernhardson: start full in-place reindex of codfw cirrussearch cluster * 16:59 mutante: stewards1001/stewards2001 - reboot for maintenance * 16:54 ebernhardson: start full in-place reindex of cloudelastic cluster * 16:53 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 16:52 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply * 16:49 mutante: planet1003/planet2003 - rebooting * 16:47 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on doh5004.wikimedia.org with reason: random high load, investigating * 15:55 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 15:54 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 15:51 jynus: restarting backupmon1001 * 15:49 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 14 hosts * 15:49 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 14 hosts * 15:47 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backupmon1001.eqiad.wmnet with reason: restart * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:06 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 15:06 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 14:59 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply * 14:58 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply * 14:51 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 14 hosts * 14:51 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 14 hosts * 14:49 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 6 hosts with reason: reboot & upgrade * 14:48 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet * 14:48 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet * 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:42 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1052.eqiad.wmnet * 14:42 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1052.eqiad.wmnet * 14:42 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1052.eqiad.wmnet * 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:31 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:31 elukey@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: sync * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:30 elukey@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: sync * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:28 elukey@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: sync * 14:28 elukey@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: sync * 14:26 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:20 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1052.eqiad.wmnet with OS trixie * 14:19 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 6 hosts with reason: reboot & upgrade * 14:18 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:15 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:15 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:13 elukey: update druid indexation job for webrequest_sampled_live - [[phab:T427068|T427068]] * 14:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:09 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:09 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for papaul - jhancock@cumin2002" * 14:09 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for papaul - jhancock@cumin2002" * 14:07 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:07 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:04 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 14:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cuminunpriv1001.eqiad.wmnet * 13:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb1003.eqiad.wmnet * 13:59 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1052.eqiad.wmnet with reason: host reimage * 13:57 moritzm: installing requests security updates * 13:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cuminunpriv1001.eqiad.wmnet * 13:55 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb1003.eqiad.wmnet * 13:53 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1052.eqiad.wmnet with reason: host reimage * 13:50 moritzm: installing python-cryptography security updates * 13:47 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb2003.codfw.wmnet * 13:44 Msz2001: UTC afternoon config+backport window is done * 13:44 Msz2001: Updated `logging` on `metawiki` to fix log performers, [[phab:T431176|T431176]]#12105297 * 13:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb2003.codfw.wmnet * 13:43 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 13:43 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt1002.wikimedia.org * 13:41 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] (duration: 07m 30s) * 13:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt1002.wikimedia.org * 13:37 mszwarc@deploy2003: mszwarc: Continuing with deployment * 13:36 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1052 * 13:36 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1052 * 13:35 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:35 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1052 * 13:35 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1052.eqiad.wmnet 47.32.64.10.in-addr.arpa 7.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:35 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1052.eqiad.wmnet 47.32.64.10.in-addr.arpa 7.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:35 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:35 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1052 - blake@cumin1003" * 13:35 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1052 - blake@cumin1003" * 13:34 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] * 13:31 blake@cumin1003: START - Cookbook sre.dns.netbox * 13:31 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1052 * 13:30 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1052.eqiad.wmnet with OS trixie * 13:30 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1052.eqiad.wmnet * 13:29 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1052.eqiad.wmnet * 13:29 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1052.eqiad.wmnet * 13:17 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] (duration: 11m 26s) * 13:13 jforrester@deploy2003: jforrester: Continuing with deployment * 13:08 jforrester@deploy2003: jforrester: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:06 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] * 12:54 cgoubert@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply * 12:54 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:52 cgoubert@deploy2003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply * 12:45 cgoubert@deploy2003: helmfile [codfw] DONE helmfile.d/services/mobileapps: apply * 12:44 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:44 cgoubert@deploy2003: helmfile [codfw] START helmfile.d/services/mobileapps: apply * 12:43 cgoubert@deploy2003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 12:43 cgoubert@deploy2003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 12:42 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast4006.wikimedia.org * 12:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt2002.wikimedia.org * 12:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast7002.wikimedia.org * 12:18 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast4006.wikimedia.org * 12:18 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host ml-serve1004 * 12:18 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host ml-serve1004 * 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt2002.wikimedia.org * 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast7002.wikimedia.org * 12:10 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backup[2003,2014].codfw.wmnet with reason: reboot & upgrade * 12:10 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt-staging2001.codfw.wmnet * 12:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid1003.eqiad.wmnet * 12:06 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt-staging2001.codfw.wmnet * 12:05 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid1003.eqiad.wmnet * 12:03 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backup[1003,1014].eqiad.wmnet with reason: reboot & upgrade * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid2003.codfw.wmnet * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host irc1003.wikimedia.org * 11:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid2003.codfw.wmnet * 11:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host irc1003.wikimedia.org * 11:55 jmm@dns1004: END - running authdns-update * 11:53 jmm@dns1004: START - running authdns-update * 11:50 jmm@dns1004: END - running authdns-update * 11:48 jmm@dns1004: START - running authdns-update * 11:27 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host irc2003.wikimedia.org * 11:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host irc2003.wikimedia.org * 11:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint2001.codfw.wmnet * 11:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint1001.eqiad.wmnet * 11:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint2001.codfw.wmnet * 11:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint1001.eqiad.wmnet * 11:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-rw2001.wikimedia.org * 11:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-rw1001.wikimedia.org * 11:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-rw2001.wikimedia.org * 11:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-rw1001.wikimedia.org * 11:03 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon1003.wikimedia.org * 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2005.codfw.wmnet * 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2005.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 10:59 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2005.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 10:57 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon1003.wikimedia.org * 10:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon2002.wikimedia.org * 10:55 jmm@cumin2003: START - Cookbook sre.dns.netbox * 10:51 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon2002.wikimedia.org * 10:51 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:50 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2005.codfw.wmnet * 10:41 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:40 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host ml-serve1003 * 10:40 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host ml-serve1003 * 10:39 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2033.codfw.wmnet * 10:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install2005.wikimedia.org * 10:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install1005.wikimedia.org * 10:35 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1004.eqiad.wmnet with OS bookworm * 10:31 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install1005.wikimedia.org * 10:31 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install2005.wikimedia.org * 10:30 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install4004.wikimedia.org * 10:30 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install3004.wikimedia.org * 10:29 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install3004.wikimedia.org * 10:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install4004.wikimedia.org * 10:23 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 10:21 moritzm: failover Ganeti master in codfw/routed to ganeti2034 [[phab:T430928|T430928]] * 10:19 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.addnode (exit_code=0) for new host ganeti2031.codfw.wmnet to cluster codfw and group B * 10:19 moritzm: readded ganeti2031 to the codfw Ganeti cluster [[phab:T430910|T430910]] * 10:18 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1004.eqiad.wmnet with reason: host reimage * 10:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install5004.wikimedia.org * 10:18 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1003 * 10:18 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1003 * 10:17 jmm@cumin2003: START - Cookbook sre.ganeti.addnode for new host ganeti2031.codfw.wmnet to cluster codfw and group B * 10:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install6003.wikimedia.org * 10:16 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install5004.wikimedia.org * 10:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install6003.wikimedia.org * 10:15 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1004.eqiad.wmnet with reason: host reimage * 10:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1001.eqiad.wmnet * 10:14 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 10:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2008.wikimedia.org * 10:00 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ml-serve1004 * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1004 * 09:57 jmm@cumin2003: START - Cookbook sre.dns.netbox * 09:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install7002.wikimedia.org * 09:57 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1004 * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ml-serve1004.eqiad.wmnet 50.48.64.10.in-addr.arpa 0.5.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:57 klausman@cumin1003: START - Cookbook sre.dns.wipe-cache ml-serve1004.eqiad.wmnet 50.48.64.10.in-addr.arpa 0.5.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1004 - klausman@cumin1003" * 09:56 klausman@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1004 - klausman@cumin1003" * 09:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-coord1001.eqiad.wmnet * 09:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 09:55 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow7002.magru.wmnet * 09:52 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-coord1001.eqiad.wmnet * 09:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 09:52 klausman@cumin1003: START - Cookbook sre.dns.netbox * 09:50 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install7002.wikimedia.org * 09:50 klausman@cumin1003: START - Cookbook sre.hosts.move-vlan for host ml-serve1004 * 09:50 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1004.eqiad.wmnet with OS bookworm * 09:50 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1003.eqiad.wmnet with OS bookworm * 09:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1001.eqiad.wmnet * 09:49 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 09:49 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2008.wikimedia.org * 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2007.codfw.wmnet * 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2007.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 09:49 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow7002.magru.wmnet * 09:49 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2007.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 09:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard1003.eqiad.wmnet * 09:39 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard2003.codfw.wmnet * 09:39 jmm@cumin2003: START - Cookbook sre.dns.netbox * 09:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard1003.eqiad.wmnet * 09:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor1003.eqiad.wmnet * 09:35 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard2003.codfw.wmnet * 09:34 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2007.codfw.wmnet * 09:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor1003.eqiad.wmnet * 09:33 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor-dev2001.codfw.wmnet * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor2003.codfw.wmnet * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sretest1006.eqiad.wmnet * 09:27 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 09:25 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor-dev2001.codfw.wmnet * 09:25 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor2003.codfw.wmnet * 09:23 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] (duration: 06m 27s) * 09:23 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2205: codfw rack B4 repool after maintenance * 09:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host sretest1006.eqiad.wmnet * 09:23 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2204: codfw rack B4 repool after maintenance * 09:19 urbanecm@deploy2003: urbanecm: Continuing with deployment * 09:19 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:18 jmm@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 6 hosts with reason: reboot * 09:17 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] * 09:08 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ml-serve1003 * 09:08 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1003 * 09:07 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1003 * 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ml-serve1003.eqiad.wmnet 81.32.64.10.in-addr.arpa 1.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:07 klausman@cumin1003: START - Cookbook sre.dns.wipe-cache ml-serve1003.eqiad.wmnet 81.32.64.10.in-addr.arpa 1.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1003 - klausman@cumin1003" * 09:06 klausman@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1003 - klausman@cumin1003" * 08:58 klausman@cumin1003: START - Cookbook sre.dns.netbox * 08:57 klausman@cumin1003: START - Cookbook sre.hosts.move-vlan for host ml-serve1003 * 08:57 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1003.eqiad.wmnet with OS bookworm * 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=0) rolling reimage on P<nowiki>{</nowiki>ml-serve1003.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet * 08:55 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet * 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1003.eqiad.wmnet with OS bookworm * 08:39 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 08:38 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool db2205: codfw rack B4 repool after maintenance * 08:37 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool db2204: codfw rack B4 repool after maintenance * 08:36 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 08:35 hashar@deploy2003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 08:32 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:32 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:31 hashar@deploy2003: Rolling back deployment * 08:26 moritzm: failover Ganeti master in codfw to ganeti2048 [[phab:T430928|T430928]] * 08:16 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1003.eqiad.wmnet with OS bookworm * 08:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2004.codfw.wmnet * 08:16 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet * 08:16 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet * 08:16 klausman@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on P<nowiki>{</nowiki>ml-serve1003.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 08:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2002.codfw.wmnet * 08:15 XioNoX: lsw1-b4-codfw> request system reboot - [[phab:T430910|T430910]] * 08:15 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b4-codfw,lsw1-b4-codfw IPv6,lsw1-b4-codfw.mgmt with reason: Switch maintenance * 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for codfw rack B4 * 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:10 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2004.codfw.wmnet * 08:10 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2002.codfw.wmnet * 08:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2205: codfw rack B4 depool for maintenance * 08:08 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool db2205: codfw rack B4 depool for maintenance * 08:08 jmm@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin2003.codfw.wmnet * 08:08 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2204: codfw rack B4 depool for maintenance * 08:08 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool db2204: codfw rack B4 depool for maintenance * 08:08 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 27 hosts with reason: codfw rack B4 depool for maintenance * 08:03 jmm@cumin2002: START - Cookbook sre.hosts.reboot-single for host cumin2003.codfw.wmnet * 07:56 ayounsi@cumin1003: START - Cookbook sre.network.depool-rack with action 'depool' for codfw rack B4 * 07:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1008.eqiad.wmnet with OS trixie * 07:49 wmde-fisch@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] (duration: 08m 36s) * 07:44 wmde-fisch@deploy2003: wmde-fisch: Continuing with deployment * 07:43 wmde-fisch@deploy2003: wmde-fisch: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:41 wmde-fisch@deploy2003: Started scap sync-world: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] * 07:35 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1008.eqiad.wmnet with reason: host reimage * 07:31 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1008.eqiad.wmnet with reason: host reimage * 07:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1008.eqiad.wmnet with OS trixie * 07:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 07:00 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 06:59 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1008.eqiad.wmnet with OS trixie * 06:57 Emperor: rebalance thanos swift rings after previous re-image of thanos-fe1004 to trixie * 06:47 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1008.eqiad.wmnet with OS trixie * 04:10 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 14 days, 0:00:00 on cp6008.drmrs.wmnet with reason: Hardware failure - [[phab:T431651|T431651]] * 03:55 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp6008.* * 03:29 ryankemper: [[phab:T431311|T431311]] Repooled eqiad cirrussearch clusters (`chi/omega/psi`) following completion of OpenSearch 2.19 migration * 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad * 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=eqiad * 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 31s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-08 == * 23:52 Amir1: ladsgroup@deploy2003:~$ mwscript-k8s --follow -- extensions/ORES/maintenance/PurgeScoreCache.php --wiki=simplewiki --model damaging --old ([[phab:T431159|T431159]]) * 23:46 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 23:46 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing PTR for 2001:df2:e500:fe08::1 - cmooney@cumin1003" * 23:46 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing PTR for 2001:df2:e500:fe08::1 - cmooney@cumin1003" * 23:40 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 23:16 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 23:15 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 22:42 rzl@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 22:40 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] (duration: 12m 55s) * 22:40 rzl@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 22:37 rzl@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 22:36 rzl@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 22:35 rzl@deploy2003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 22:34 urbanecm@deploy2003: urbanecm: Continuing with deployment * 22:33 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:33 rzl@deploy2003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 22:32 rzl@deploy2003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 22:30 rzl@deploy2003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 22:30 rzl@deploy2003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 22:29 rzl@deploy2003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 22:27 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] * 22:26 rzl@deploy2003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 22:22 rzl@deploy2003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 22:21 rzl@deploy2003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 22:19 rzl@deploy2003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 22:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 22:17 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 22:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 22:13 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 22:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 22:13 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 22:09 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 22:06 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 22:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1094.eqiad.wmnet with OS trixie * 22:01 urbanecm: Make https://test.wikipedia.org/w/index.php?title=MediaWiki:GrowthExperimentsSuggestedEdits.json&diff=prev&oldid=750552 with GrowthExperiments disabled (via mw-experimental), then run `\MediaWiki\MediaWikiServices::getInstance()->get('CommunityConfiguration.ProviderFactory')->newProvider('GrowthSuggestedEdits')->getStore()->invalidate()` ([[phab:T431625|T431625]]) * 21:56 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d2-codfw * 21:55 urbanecm@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 21:55 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d2-codfw * 21:55 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c4-codfw * 21:55 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c4-codfw * 21:55 urbanecm@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2002 * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2002 * 21:54 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2002 * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2002.codfw.wmnet 50.32.192.10.in-addr.arpa 0.5.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:54 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2002.codfw.wmnet 50.32.192.10.in-addr.arpa 0.5.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2002 - bking@cumin2003" * 21:54 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2002 - bking@cumin2003" * 21:49 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:49 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2002 * 21:49 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2002.codfw.wmnet with OS trixie * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1094.eqiad.wmnet with reason: host reimage * 21:42 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 21:39 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 21:37 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1094.eqiad.wmnet with reason: host reimage * 21:36 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 21:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 21:29 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 21:27 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 21:22 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1094.eqiad.wmnet with OS trixie * 21:21 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host restbase2039.codfw.wmnet with OS bullseye * 21:21 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin2002" * 21:21 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin2002" * 21:04 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on restbase2039.codfw.wmnet with reason: host reimage * 21:00 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on restbase2039.codfw.wmnet with reason: host reimage * 20:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1073.eqiad.wmnet with OS trixie * 20:48 mutante: deploy2003 - kill 1102 (stunnel4) ; systemctl start stunnel4 ([[phab:T418262|T418262]]) * 20:42 cjming@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] (duration: 33m 02s) * 20:42 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host restbase2039.codfw.wmnet with OS bullseye * 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1073.eqiad.wmnet with reason: host reimage * 20:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1073.eqiad.wmnet with reason: host reimage * 20:30 cjming@deploy2003: cjming: Continuing with deployment * 20:28 cjming@deploy2003: cjming: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1098.eqiad.wmnet with OS trixie * 20:13 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1073.eqiad.wmnet with OS trixie * 20:09 cjming@deploy2003: Started scap sync-world: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] * 20:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1098.eqiad.wmnet with reason: host reimage * 19:56 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1098.eqiad.wmnet with reason: host reimage * 19:55 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d4-codfw * 19:54 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d4-codfw * 19:54 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c1-codfw * 19:54 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c1-codfw * 19:52 mutante: restarting gerrit on gerrit.wikimedia.org (gerrit2003) * 19:48 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2331.codfw.wmnet * 19:48 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2331.codfw.wmnet * 19:48 mutante: restarting gerrit on gerrit-replica.wikimedia.org (gerrit1003) * 19:46 mutante: restarting gerrit on gerrit-spare.wikimedia.org (gerrit2002) * 19:43 jasmine@cumin2002: conftool action : set/pooled=yes; selector: name=wikikube-worker2331.codfw.wmnet,cluster=kubernetes,service=kubesvc * 19:43 jasmine@cumin2002: conftool action : set/weight=10; selector: name=wikikube-worker2331.codfw.wmnet,cluster=kubernetes,service=kubesvc * 19:40 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1098.eqiad.wmnet with OS trixie * 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d5-codfw * 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d5-codfw * 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c7-codfw * 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c7-codfw * 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c5-codfw * 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c5-codfw * 19:30 jasmine_: ran homer on lsw1-d8-codfw, adding wikikube-worker2331 to cluster * 19:29 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1100.eqiad.wmnet with OS trixie * 19:20 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d8-codfw * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d8-codfw * 19:19 mutante: gerrit - replacing private key for registerEmail verification * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-magru * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device cr2-magru * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d7-codfw * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d7-codfw * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d3-codfw * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d1-codfw * 19:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d1-codfw * 19:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c2-codfw * 19:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c2-codfw * 19:11 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-codfw * 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-magru * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device cr1-magru * 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d8-codfw * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d8-codfw * 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d6-codfw * 19:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1100.eqiad.wmnet with reason: host reimage * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d6-codfw * 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c6-codfw * 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c6-codfw * 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c3-codfw * 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c3-codfw * 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b4-magru * 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device asw1-b4-magru * 19:08 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b3-magru * 19:08 cmooney@cumin1003: START - Cookbook sre.network.tls for network device asw1-b3-magru * 19:05 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1100.eqiad.wmnet with reason: host reimage * 19:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1122.eqiad.wmnet with OS trixie * 19:00 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 18:59 topranks: rolling out update to BGP ACL on Nokia Switches eqiad, codfw & ulsfo [[phab:T425703|T425703]] * 18:58 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 18:57 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 18:55 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 18:53 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 18:52 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 18:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1100.eqiad.wmnet with OS trixie * 18:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1068.eqiad.wmnet with OS trixie * 18:47 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1102.eqiad.wmnet with OS trixie * 18:47 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 18:46 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1122.eqiad.wmnet with reason: host reimage * 18:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1122.eqiad.wmnet with reason: host reimage * 18:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1068.eqiad.wmnet with reason: host reimage * 18:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1122 * 18:26 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1122 * 18:25 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1122 * 18:25 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1122.eqiad.wmnet 31.48.64.10.in-addr.arpa 1.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:25 bking@cumin2003: START - Cookbook sre.dns.wipe-cache cirrussearch1122.eqiad.wmnet 31.48.64.10.in-addr.arpa 1.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:25 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:25 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1122 - bking@cumin2003" * 18:25 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1122 - bking@cumin2003" * 18:21 rzl@deploy2003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 18:21 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1068.eqiad.wmnet with reason: host reimage * 18:21 rzl@deploy2003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 18:21 rzl@deploy2003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 18:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 18:19 rzl@deploy2003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 18:19 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:18 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1122 * 18:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1122.eqiad.wmnet with OS trixie * 18:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 18:15 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 18:13 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 18:13 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 18:10 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 18:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1068.eqiad.wmnet with OS trixie * 18:01 kamila@deploy2003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 18m 29s) * 18:00 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:55 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:42 kamila@deploy2003: Started scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] * 17:42 kamila@deploy2003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 19m 50s) * 17:42 kamila@deploy2003: Rolling back deployment * 17:35 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:31 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet * 17:18 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet * 17:16 kamila@deploy1003: Unlocked for deployment [MediaWiki]: switching deployment server (duration: 22m 07s) * 17:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 17:11 kamila@dns1005: END - running authdns-update * 17:09 kamila@dns1005: START - running authdns-update * 17:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 17:04 jasmine@cumin2002: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1164.eqiad.wmnet * 17:04 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1164.eqiad.wmnet * 17:04 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1164.eqiad.wmnet * 16:56 kamila@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on releases2003.codfw.wmnet,releases1003.eqiad.wmnet with reason: Deployment server switchover * 16:54 kamila@deploy1003: Locking from deployment [MediaWiki]: switching deployment server * 16:53 kamila@deploy1003: Unlocked for deployment [MediaWiki]: switching deployment server (duration: 04m 02s) * 16:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie * 16:49 kamila@deploy1003: Locking from deployment [MediaWiki]: switching deployment server * 16:46 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1095.eqiad.wmnet with OS trixie * 16:45 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1093.eqiad.wmnet with OS trixie * 16:43 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1164.eqiad.wmnet with OS trixie * 16:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1095.eqiad.wmnet with reason: host reimage * 16:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 16:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 16:23 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1164.eqiad.wmnet with reason: host reimage * 16:18 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on cirrussearch1093.eqiad.wmnet with reason: host reimage * 16:16 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1164.eqiad.wmnet with reason: host reimage * 16:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1095.eqiad.wmnet with reason: host reimage * 16:09 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 16:09 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 16:08 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1093.eqiad.wmnet with reason: host reimage * 15:59 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Pool test * 15:59 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:59 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 15:59 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Pool test * 15:58 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Depool test * 15:58 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:58 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 15:58 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Depool test * 15:57 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1164 * 15:57 jasmine@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1164 * 15:57 jasmine@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1164 * 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1164.eqiad.wmnet 114.48.64.10.in-addr.arpa 4.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:56 jasmine@cumin2002: START - Cookbook sre.dns.wipe-cache wikikube-worker1164.eqiad.wmnet 114.48.64.10.in-addr.arpa 4.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1164 - jasmine@cumin2002" * 15:56 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1164 - jasmine@cumin2002" * 15:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1093.eqiad.wmnet with OS trixie * 15:51 jasmine@cumin2002: START - Cookbook sre.dns.netbox * 15:51 jasmine@cumin2002: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1164 * 15:50 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-worker1164.eqiad.wmnet with OS trixie * 15:50 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1164.eqiad.wmnet * 15:50 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1164.eqiad.wmnet * 15:50 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1164.eqiad.wmnet * 15:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1095.eqiad.wmnet with OS trixie * 15:42 jasmine@cumin2002: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1164.eqiad.wmnet * 15:42 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1164.eqiad.wmnet * 15:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:42 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1164.eqiad.wmnet * 15:42 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1164.eqiad.wmnet * 15:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 15:39 elukey@cumin1003: START - Cookbook sre.hosts.provision for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 15:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1007.eqiad.wmnet with OS trixie * 15:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Pool test * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1007.eqiad.wmnet with reason: host reimage * 15:15 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 15:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet * 15:15 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 15:15 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1007.eqiad.wmnet with reason: host reimage * 15:15 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:14 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Pool test * 15:14 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet * 15:14 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2228: Depool test * 15:14 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db2228: Depool test * 15:10 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 15:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:08 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 15:06 blake@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 15:06 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet * 15:06 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 15:06 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 15:05 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet * 15:05 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 15:04 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:04 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 15:04 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:03 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 15:03 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 15:03 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:03 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T430909|T430909]] * 15:03 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:03 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 15:03 swfrench-wmf: restarted eqsin, codfw confds - [[phab:T430909|T430909]] * 15:03 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test1002.eqiad.wmnet * 15:02 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet * 14:59 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:59 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:55 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1007.eqiad.wmnet with OS trixie * 14:52 swfrench-wmf: restarted ulsfo confds, confirmed now connected to codfw backends except those using wikimedia.org SRV record - [[phab:T430909|T430909]] * 14:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:49 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:41 moritzm: uninstalling dhcpcd-base from trixie hosts which still have it installed [[phab:T414341|T414341]] * 14:40 sukhe: sudo cumin -b1 -s120 "P<nowiki>{</nowiki>lvs2011*<nowiki>}</nowiki> or P<nowiki>{</nowiki>lvs2012*<nowiki>}</nowiki>" "systemctl restart pybal.service" * 14:39 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:39 mvernon@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host thanos-be1007.eqiad.wmnet with OS trixie * 14:37 sukhe: restart pybal on lvs2013 to revert back to conf2004 * 14:35 sukhe: restart pybal on lvs2014 to revert back to conf2004 * 14:34 swfrench-wmf: switched codfw, eqsin, ulsfo etcd client SRV records back to codfw - [[phab:T430909|T430909]] * 14:32 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1002.eqiad.wmnet * 14:32 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet * 14:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1007.eqiad.wmnet with OS trixie * 14:31 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:31 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:31 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:30 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:30 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Pool test * 14:30 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:29 swfrench@dns1004: END - running authdns-update * 14:29 moritzm: installing jackson-core security updates * 14:27 swfrench@dns1004: START - running authdns-update * 14:22 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:22 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:22 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1119.eqiad.wmnet with OS trixie * 14:22 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:21 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:20 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:20 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 14:20 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:19 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 14:19 moritzm: installing librabbitmq security updates * 14:19 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1002.eqiad.wmnet * 14:18 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet * 14:16 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:15 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:15 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:14 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Pool test * 14:14 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox) * 14:14 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet * 14:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 14:08 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 14:07 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet * 14:05 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1118.eqiad.wmnet with OS trixie * 14:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test1001.eqiad.wmnet * 14:00 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-worker@eqiad * 14:00 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 13:59 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 13:58 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 13:57 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1119.eqiad.wmnet with reason: host reimage * 13:54 moritzm: installing libcap2 security updates * 13:53 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1119.eqiad.wmnet with reason: host reimage * 13:52 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet * 13:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) * 13:52 fceratto@cumin1003: START - Cookbook sre.mysql.depool * 13:50 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-worker@eqiad * 13:50 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1051.eqiad.wmnet * 13:50 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1051.eqiad.wmnet * 13:50 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1051.eqiad.wmnet * 13:49 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 13:45 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1081.eqiad.wmnet with OS trixie * 13:41 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1119 * 13:41 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1119 * 13:40 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1119 * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1119.eqiad.wmnet 97.32.64.10.in-addr.arpa 7.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1119.eqiad.wmnet 97.32.64.10.in-addr.arpa 7.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1119 - atsuko@cumin1003" * 13:40 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1119 - atsuko@cumin1003" * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1118.eqiad.wmnet with reason: host reimage * 13:39 moritzm: installing krb5 security updates * 13:37 Lucas_WMDE: UTC afternoon backport+config window done * 13:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1006.eqiad.wmnet with OS trixie * 13:36 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1118.eqiad.wmnet with reason: host reimage * 13:36 atsuko@cumin1003: START - Cookbook sre.dns.netbox * 13:35 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] (duration: 07m 46s) * 13:34 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1119 * 13:34 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1119.eqiad.wmnet with OS trixie * 13:30 sbisson@deploy1003: sbisson: Continuing with deployment * 13:30 moritzm: installing openssh security updates * 13:30 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling restart_daemons on A:wikidough * 13:29 sbisson@deploy1003: sbisson: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:27 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] * 13:26 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1051.eqiad.wmnet with OS trixie * 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1081.eqiad.wmnet with reason: host reimage * 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1118 * 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1118 * 13:22 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] (duration: 12m 12s) * 13:21 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1081.eqiad.wmnet with reason: host reimage * 13:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1006.eqiad.wmnet with reason: host reimage * 13:18 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1118 * 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1118.eqiad.wmnet 90.32.64.10.in-addr.arpa 0.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:18 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1118.eqiad.wmnet 90.32.64.10.in-addr.arpa 0.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1118 - atsuko@cumin1003" * 13:18 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1118 - atsuko@cumin1003" * 13:17 stran@deploy1003: stran: Continuing with deployment * 13:16 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough * 13:15 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:13 atsuko@cumin1003: START - Cookbook sre.dns.netbox * 13:12 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1118 * 13:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1006.eqiad.wmnet with reason: host reimage * 13:12 stran@deploy1003: stran: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:12 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1118.eqiad.wmnet with OS trixie * 13:10 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] * 13:05 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-worker@codfw * 13:05 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 13:05 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1081.eqiad.wmnet with OS trixie * 13:05 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1051.eqiad.wmnet with reason: host reimage * 13:04 moritzm: installing jq security updates * 13:04 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 13:01 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1051.eqiad.wmnet with reason: host reimage * 12:58 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-worker@codfw * 12:52 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 12:50 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host thanos-be1006.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1051 * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1051 * 12:43 moritzm: installing Python 3.11 security updates * 12:43 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1051 * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1051.eqiad.wmnet 46.32.64.10.in-addr.arpa 6.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:43 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1051.eqiad.wmnet 46.32.64.10.in-addr.arpa 6.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1051 - blake@cumin1003" * 12:43 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1051 - blake@cumin1003" * 12:38 blake@cumin1003: START - Cookbook sre.dns.netbox * 12:38 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1051 * 12:38 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1051.eqiad.wmnet with OS trixie * 12:37 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1051.eqiad.wmnet * 12:36 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1051.eqiad.wmnet * 12:36 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1051.eqiad.wmnet * 12:34 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1006.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 12:34 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1006.eqiad.wmnet with OS trixie * 12:27 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 12:27 mvernon@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host thanos-be1006.eqiad.wmnet with OS trixie * 12:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:02 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 12:01 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1006.eqiad.wmnet with OS trixie * 11:43 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 11:38 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1076.eqiad.wmnet with OS trixie * 11:26 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1075.eqiad.wmnet with OS trixie * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2047.codfw.wmnet * 11:19 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2047.codfw.wmnet * 11:19 moritzm: temporarily remove ganeti2031 from codfw cluster [[phab:T430910|T430910]] * 11:08 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1076.eqiad.wmnet with reason: host reimage * 11:08 moritzm: installing Linux 6.1.176 on Bookworm servers * 11:03 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1076.eqiad.wmnet with reason: host reimage * 11:00 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1075.eqiad.wmnet with reason: host reimage * 10:56 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1075.eqiad.wmnet with reason: host reimage * 10:47 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1076.eqiad.wmnet with OS trixie * 10:46 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1074.eqiad.wmnet with OS trixie * 10:45 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1005.eqiad.wmnet with OS trixie * 10:40 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1075.eqiad.wmnet with OS trixie * 10:32 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2031.codfw.wmnet * 10:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1005.eqiad.wmnet with reason: host reimage * 10:25 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1005.eqiad.wmnet with reason: host reimage * 10:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1074.eqiad.wmnet with reason: host reimage * 10:17 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1074.eqiad.wmnet with reason: host reimage * 10:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1005.eqiad.wmnet with OS trixie * 10:12 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet * 10:04 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 10:01 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet * 10:01 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1074.eqiad.wmnet with OS trixie * 10:01 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 09:43 cgoubert@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/aux-k8s-services/redioscope: apply * 09:43 cgoubert@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/aux-k8s-services/redioscope: apply * 09:43 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply * 09:35 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply * 09:34 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 09:34 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 09:33 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 41 days, 15:00:00 on db2252.codfw.wmnet with reason: Test * 09:32 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 09:32 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 09:31 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: codfw rack B3 pool after maintenance * 09:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 09:07 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 09:07 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 09:02 ladsgroup@cumin1003: END (PASS) - Cookbook sre.mysql.sanitarium_restart (exit_code=0) * 08:57 topranks: merge patch to shift eqiad <-> esams traffic onto new 40G circuit * 08:54 hashar@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1004.eqiad.wmnet with OS trixie * 08:50 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 08:50 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitarium_restart (exit_code=99) * 08:50 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 08:45 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool es2051: codfw rack B3 pool after maintenance * 08:44 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:44 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:43 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2007.codfw.wmnet * 08:43 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2007.codfw.wmnet * 08:42 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2031.codfw.wmnet * 08:41 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2031.codfw.wmnet * 08:40 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2031.codfw.wmnet * 08:38 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:38 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:35 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.sanitize-wiki (exit_code=97) Managing sanitization for wikis minwikiquote in section s3 * 08:33 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis minwikiquote in section s3 * 08:32 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Checking sanitization for wikis minwikiquote in section s5 * 08:30 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Checking sanitization for wikis minwikiquote in section s5 * 08:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Managing sanitization for wikis minwikiquote in section s5 * 08:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1004.eqiad.wmnet with reason: host reimage * 08:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1004.eqiad.wmnet with reason: host reimage * 08:23 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:22 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis minwikiquote in section s5 * 08:19 XioNoX: lsw1-b3-codfw> request system reboot - [[phab:T430909|T430909]] * 08:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Checking sanitization for wikis minwikiquote in section s5 * 08:17 hashar@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Checking sanitization for wikis minwikiquote in section s5 * 08:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for codfw rack B3 * 08:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2007.codfw.wmnet * 08:15 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lsw1-b3-codfw,lsw1-b3-codfw IPv6,lsw1-b3-codfw.mgmt with reason: Switch maintenance * 08:15 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2007.codfw.wmnet * 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:07 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:06 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: codfw rack B3 depool for maintenance * 08:05 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool es2051: codfw rack B3 depool for maintenance * 08:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1004.eqiad.wmnet with OS trixie * 08:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1005.eqiad.wmnet with OS trixie * 08:03 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 21 hosts with reason: codfw rack B3 depool for maintenance * 07:56 ayounsi@cumin1003: START - Cookbook sre.network.depool-rack with action 'depool' for codfw rack B3 * 07:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1005.eqiad.wmnet with reason: host reimage * 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1005.eqiad.wmnet with reason: host reimage * 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1005.eqiad.wmnet with OS bookworm * 07:29 moritzm: installing gnutls28 security updates * 07:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1005.eqiad.wmnet with OS trixie * 07:13 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1125.eqiad.wmnet with OS trixie * 07:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1005.eqiad.wmnet with reason: host reimage * 07:07 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aux-k8s-etcd1005.eqiad.wmnet with reason: host reimage * 06:56 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1005.eqiad.wmnet with OS bookworm * 06:54 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1125.eqiad.wmnet with reason: host reimage * 06:52 elukey: upgrade all trixie hosts to pywmflib 3.1 - [[phab:T430552|T430552]] * 06:50 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1125.eqiad.wmnet with reason: host reimage * 06:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 06:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 06:38 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1125.eqiad.wmnet with OS trixie * 05:42 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1107.eqiad.wmnet with OS trixie * 05:35 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1124.eqiad.wmnet with OS trixie * 05:31 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1101.eqiad.wmnet with OS trixie * 05:21 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1107.eqiad.wmnet with reason: host reimage * 05:17 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1124.eqiad.wmnet with reason: host reimage * 05:13 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1107.eqiad.wmnet with reason: host reimage * 05:13 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1101.eqiad.wmnet with reason: host reimage * 05:11 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1124.eqiad.wmnet with reason: host reimage * 05:10 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1101.eqiad.wmnet with reason: host reimage * 04:58 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1124.eqiad.wmnet with OS trixie * 04:56 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1107.eqiad.wmnet with OS trixie * 04:55 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1101.eqiad.wmnet with OS trixie * 02:27 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] (duration: 08m 14s) * 02:22 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 02:21 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 02:19 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] * 01:59 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] (duration: 09m 46s) * 01:55 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 01:51 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 01:49 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] * 01:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1099.eqiad.wmnet with OS trixie * 00:57 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1110.eqiad.wmnet with OS trixie * 00:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1099.eqiad.wmnet with reason: host reimage * 00:41 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1099.eqiad.wmnet with reason: host reimage * 00:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1110.eqiad.wmnet with reason: host reimage * 00:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1110.eqiad.wmnet with reason: host reimage * 00:26 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1099.eqiad.wmnet with OS trixie * 00:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1110.eqiad.wmnet with OS trixie == 2026-07-07 == * 22:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1097.eqiad.wmnet with OS trixie * 22:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1097.eqiad.wmnet with reason: host reimage * 22:24 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1097.eqiad.wmnet with reason: host reimage * 22:14 hashar: Restarting Gerrit on gerrit2002 and gerrit1003 (replicas) * 22:09 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1097.eqiad.wmnet with OS trixie * 22:07 hashar: Restarting Gerrit on gerrit2003 * 21:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 21:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 21:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 21:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 21:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1108.eqiad.wmnet with OS trixie * 20:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1091.eqiad.wmnet with OS trixie * 20:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1090.eqiad.wmnet with OS trixie * 20:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1108.eqiad.wmnet with reason: host reimage * 20:36 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1091.eqiad.wmnet with reason: host reimage * 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1090.eqiad.wmnet with reason: host reimage * 20:33 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1091.eqiad.wmnet with reason: host reimage * 20:30 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1108.eqiad.wmnet with reason: host reimage * 20:30 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1006.eqiad.wmnet * 20:30 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1090.eqiad.wmnet with reason: host reimage * 20:30 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1006.eqiad.wmnet * 20:27 jasmine_: "homer lsw1-c2-eqiad* commit "Added new stacked control plane wikikube-ctrl1006"" * 20:22 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] (duration: 07m 29s) * 20:20 jasmine_: "homer "cr*eqiad*" commit "Added new stacked control plane wikikube-ctrl1006"" * 20:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1091.eqiad.wmnet with OS trixie * 20:17 arlolra@deploy1003: arlolra: Continuing with deployment * 20:16 arlolra@deploy1003: arlolra: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:16 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1090.eqiad.wmnet with OS trixie * 20:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1108.eqiad.wmnet with OS trixie * 20:14 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] * 20:09 cwhite: remove 2026-04 swift log archives from centrallog2002 to free some space * 20:01 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=93) for host cirrussearch1108.eqiad.wmnet with OS trixie * 19:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1108.eqiad.wmnet with OS trixie * 19:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1090.eqiad.wmnet with OS trixie * 19:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1109.eqiad.wmnet with OS trixie * 19:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1092.eqiad.wmnet with OS trixie * 19:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1123.eqiad.wmnet with OS trixie * 19:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1109.eqiad.wmnet with reason: host reimage * 19:19 cdobbins@cumin2002: conftool action : set/pooled=yes; selector: name=dns7002.* * 19:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1092.eqiad.wmnet with reason: host reimage * 19:17 jasmine@dns1004: END - running authdns-update * 19:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1109.eqiad.wmnet with reason: host reimage * 19:15 jasmine@dns1004: START - running authdns-update * 19:15 cdobbins@dns1004: END - running authdns-update * 19:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1123.eqiad.wmnet with reason: host reimage * 19:13 cdobbins@dns1004: START - running authdns-update * 19:12 cdobbins@cumin2002: conftool action : set/pooled=yes; selector: name=dns7002.*,service=authdns-update * 19:11 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1092.eqiad.wmnet with reason: host reimage * 19:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1123.eqiad.wmnet with reason: host reimage * 18:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1123.eqiad.wmnet with OS trixie * 18:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1109.eqiad.wmnet with OS trixie * 18:56 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1092.eqiad.wmnet with OS trixie * 18:52 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 18:49 swfrench@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 18:40 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 18:38 swfrench@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 18:11 swfrench-wmf: restarted eqsin, codfw confds - [[phab:T430909|T430909]] * 18:01 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T430909|T430909]] * 17:59 swfrench-wmf: restarted ulsfo confds, confirmed now connected to eqiad backends - [[phab:T430909|T430909]] * 17:52 sukhe: restart pybal on lvs2011 to switch from conf2004 to conf1008: [[phab:T430909|T430909]] * 17:51 sukhe: restart pybal on lvs2012 to switch from conf2004 to conf1008 [puppet re-enabled there]: [[phab:T430909|T430909]] * 17:46 sukhe: restart pybal on lvs2013 to switch from conf2004 to conf1008: [[phab:T430909|T430909]] * 17:44 swfrench-wmf: switched codfw, eqsin, ulsfo etcd client SRV records to eqiad - [[phab:T430909|T430909]] * 17:43 swfrench@dns1004: END - running authdns-update * 17:40 swfrench@dns1004: START - running authdns-update * 17:40 sukhe: restart pybal on lvs2014 to switch from conf2004 to conf1008: [[phab:T430909|T430909]] * 17:21 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1003.eqiad.wmnet * 17:15 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1003.eqiad.wmnet * 17:14 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1002.eqiad.wmnet * 17:06 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1002.eqiad.wmnet * 17:06 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-low-traffic-codfw' 'systemctl restart pybal.service' # lvs2013, [[phab:T416623|T416623]] * 17:04 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1001.eqiad.wmnet * 17:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1111.eqiad.wmnet with OS trixie * 17:00 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal.service' # lvs2014, [[phab:T416623|T416623]] * 16:58 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1001.eqiad.wmnet * 16:58 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS bookworm * 16:55 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-low-traffic-eqiad' 'systemctl restart pybal.service' # lvs1019, [[phab:T416623|T416623]] * 16:53 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-secondary-eqiad' 'systemctl restart pybal.service' # lvs1020, [[phab:T416623|T416623]] * 16:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1111.eqiad.wmnet with reason: host reimage * 16:40 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1111.eqiad.wmnet with reason: host reimage * 16:38 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.peering (exit_code=99) with action 'configure' for AS: 47794 * 16:35 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 47794 * 16:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1111.eqiad.wmnet with OS trixie * 16:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1006.eqiad.wmnet with OS trixie * 16:06 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 16:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1006.eqiad.wmnet with reason: host reimage * 15:58 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1121.eqiad.wmnet with OS trixie * 15:58 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 15:56 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1006.eqiad.wmnet with reason: host reimage * 15:54 mutante: jenkins down in planned maintenance window * 15:42 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1037.eqiad.wmnet * 15:42 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1037.eqiad.wmnet * 15:42 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1037.eqiad.wmnet * 15:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1006.eqiad.wmnet with OS trixie * 15:34 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1121.eqiad.wmnet with reason: host reimage * 15:33 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS bookworm * 15:33 cdobbins@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host dns7002.wikimedia.org with OS trixie * 15:30 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1121.eqiad.wmnet with reason: host reimage * 15:29 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host clouddumps1001.wikimedia.org * 15:20 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1001.wikimedia.org * 15:18 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1121 * 15:18 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1121 * 15:18 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host clouddumps1002.wikimedia.org * 15:17 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1121 * 15:17 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:17 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply * 15:16 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply * 15:16 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:16 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1121 - atsuko@cumin1003" * 15:16 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1121 - atsuko@cumin1003" * 15:14 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1037.eqiad.wmnet with OS trixie * 15:11 atsuko@cumin1003: START - Cookbook sre.dns.netbox * 15:09 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org * 15:09 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1121 * 15:09 andrew@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host clouddumps1002.wikimedia.org * 15:09 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org * 15:09 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1121.eqiad.wmnet with OS trixie * 15:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1007.eqiad.wmnet with OS trixie * 15:08 andrew@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host clouddumps1002.wikimedia.org * 15:08 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org * 15:05 brennen@deploy1003: Finished deploy [phabricator/deployment@7e02037]: deploy phab1004 for [[phab:T431440|T431440]] (duration: 00m 47s) * 15:04 brennen@deploy1003: Started deploy [phabricator/deployment@7e02037]: deploy phab1004 for [[phab:T431440|T431440]] * 15:03 brennen@deploy1003: Finished deploy [phabricator/deployment@7e02037]: deploy phab2003 for [[phab:T431440|T431440]] (duration: 00m 51s) * 15:03 brennen@deploy1003: Started deploy [phabricator/deployment@7e02037]: deploy phab2003 for [[phab:T431440|T431440]] * 15:00 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71] (thin): Regular analytics weekly train THIN [analytics/refinery@7d8dc71f] (duration: 02m 10s) * 14:58 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71] (thin): Regular analytics weekly train THIN [analytics/refinery@7d8dc71f] * 14:58 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71]: Regular analytics weekly train [analytics/refinery@7d8dc71f] (duration: 04m 14s) * 14:54 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1037.eqiad.wmnet with reason: host reimage * 14:53 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71]: Regular analytics weekly train [analytics/refinery@7d8dc71f] * 14:53 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@7d8dc71f] (duration: 02m 00s) * 14:51 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@7d8dc71f] * 14:51 arnaudb@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on phab2003.codfw.wmnet,phab[1004-1006].eqiad.wmnet with reason: maintenance * 14:51 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1037.eqiad.wmnet with reason: host reimage * 14:50 JavierMonton: Deploying Refinery at {{Gerrit|7d8dc71f}} for change {{Gerrit|1308087}} / [[phab:T431318|T431318]] - update filerevision table sqoop and table * 14:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1007.eqiad.wmnet with reason: host reimage * 14:42 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1007.eqiad.wmnet with reason: host reimage * 14:40 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1083.eqiad.wmnet with OS trixie * 14:37 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:36 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:35 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-master@eqiad * 14:35 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 14:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 14:34 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1037 * 14:34 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1037 * 14:34 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:cleanMentorList.php --wiki=frwiki # [[phab:T427386|T427386]] * 14:34 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 14:34 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308112{{!}}Revert^2 "[Growth] frwiki: Deploy automated mentor list cleaner" (T427386)]] (duration: 06m 47s) * 14:34 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 14:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:33 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:32 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1037 * 14:31 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:31 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:29 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:29 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-master@eqiad * 14:29 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:29 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:29 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:28 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:27 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:27 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1308112{{!}}Revert^2 "[Growth] frwiki: Deploy automated mentor list cleaner" (T427386)]] * 14:26 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1007.eqiad.wmnet with OS trixie * 14:26 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:26 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 14:26 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:25 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:25 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:25 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:cleanMentorList.php --wiki=frwiki # [[phab:T427386|T427386]] * 14:24 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:24 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1037 - blake@cumin1003" * 14:24 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1037 - blake@cumin1003" * 14:20 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1083.eqiad.wmnet with reason: host reimage * 14:19 blake@cumin1003: START - Cookbook sre.dns.netbox * 14:19 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-master@codfw * 14:19 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 14:19 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1037 * 14:18 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1037.eqiad.wmnet with OS trixie * 14:18 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1037.eqiad.wmnet * 14:18 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 14:18 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1037.eqiad.wmnet * 14:18 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1037.eqiad.wmnet * 14:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2007.codfw.wmnet with OS trixie * 14:16 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1083.eqiad.wmnet with reason: host reimage * 14:15 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1036.eqiad.wmnet * 14:15 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1036.eqiad.wmnet * 14:14 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1036.eqiad.wmnet * 14:12 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-master@codfw * 14:11 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1120.eqiad.wmnet with OS trixie * 14:05 moritzm: installing distro-info-data updates from trixie/bookworm point releases * 14:04 fabfur: disable puppet on A:cp-text to selectively apply https://gerrit.wikimedia.org/r/c/operations/puppet/+/1308040 * 14:03 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] (duration: 27m 48s) * 14:00 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 14:00 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1083.eqiad.wmnet with OS trixie * 13:58 urbanecm@deploy1003: urbanecm: Continuing with deployment * 13:58 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:58 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1004.eqiad.wmnet with OS bookworm * 13:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2007.codfw.wmnet with reason: host reimage * 13:57 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 13:53 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1120.eqiad.wmnet with reason: host reimage * 13:50 moritzm: installing Linux 5.10.259 on Bullseye hosts * 13:47 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply * 13:47 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply * 13:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2007.codfw.wmnet with reason: host reimage * 13:46 cgoubert@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/aux-k8s-services/redioscope: apply * 13:46 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1120.eqiad.wmnet with reason: host reimage * 13:46 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:46 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:45 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:44 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:44 cgoubert@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/aux-k8s-services/redioscope: apply * 13:40 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 13:39 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:38 moritzm: installing e2fsprogs updates from Trixie point release * 13:35 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] * 13:33 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1120.eqiad.wmnet with OS trixie * 13:33 topranks: reset cr3-eqsin configuration so traffic uses it again after upgrade * 13:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1088.eqiad.wmnet with OS trixie * 13:32 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie * 13:32 cdobbins@cumin1003: conftool action : set/pooled=no; selector: name=dns7002.* * 13:29 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2007.codfw.wmnet with OS trixie * 13:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1004.eqiad.wmnet with reason: host reimage * 13:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2006.codfw.wmnet with OS trixie * 13:18 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1036.eqiad.wmnet with OS trixie * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aux-k8s-etcd1004.eqiad.wmnet with reason: host reimage * 13:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 13:16 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 13:15 jayme: Istio is being upgraded from 1.24.2 to 1.29.4 on wikikube staging eqiad and codfw - [[phab:T427401|T427401]] * 13:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1087.eqiad.wmnet with OS trixie * 13:14 topranks: reboot cr3-eqsin to install new JunOS and set PIC 0/0/0 to 100G * 13:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1088.eqiad.wmnet with reason: host reimage * 13:13 jmm@dns1004: END - running authdns-update * 13:12 jmm@dns1004: START - running authdns-update * 13:09 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1088.eqiad.wmnet with reason: host reimage * 13:07 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1082.eqiad.wmnet with OS trixie * 13:07 jmm@dns1004: END - running authdns-update * 13:06 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1004.eqiad.wmnet with OS bookworm * 13:05 jmm@dns1004: START - running authdns-update * 13:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2006.codfw.wmnet with reason: host reimage * 12:58 topranks: load updated JunOS on cr3-eqsin [[phab:T429386|T429386]] * 12:58 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1036.eqiad.wmnet with reason: host reimage * 12:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2001.codfw.wmnet * 12:57 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2006.codfw.wmnet with reason: host reimage * 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr1-codfw,cr[2-3]-eqsin,cr3-eqsin IPv6,cr3-eqsin.mgmt with reason: upgrade JunOS cr3-eqsin * 12:56 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lvs[5004-5006].eqsin.wmnet with reason: upgrade JunOS cr3-eqsin * 12:55 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 12:55 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 12:53 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1087.eqiad.wmnet with reason: host reimage * 12:53 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 12:52 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1088.eqiad.wmnet with OS trixie * 12:52 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1002.eqiad.wmnet * 12:52 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:51 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2001.codfw.wmnet * 12:49 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1036.eqiad.wmnet with reason: host reimage * 12:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 12:48 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1087.eqiad.wmnet with reason: host reimage * 12:44 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1082.eqiad.wmnet with reason: host reimage * 12:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1002.eqiad.wmnet * 12:42 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 12:42 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 12:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 12:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 12:39 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:39 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: move dumps-nfs IP to the shared one - filippo@cumin1003" * 12:39 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: move dumps-nfs IP to the shared one - filippo@cumin1003" * 12:39 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2006.codfw.wmnet with OS trixie * 12:38 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1082.eqiad.wmnet with reason: host reimage * 12:36 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:33 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:32 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1036 * 12:32 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1036 * 12:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2005.codfw.wmnet with OS trixie * 12:32 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1087.eqiad.wmnet with OS trixie * 12:30 jmm@dns1004: END - running authdns-update * 12:29 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1036 * 12:29 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1036.eqiad.wmnet 21.32.64.10.in-addr.arpa 1.2.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:29 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1036.eqiad.wmnet 21.32.64.10.in-addr.arpa 1.2.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:29 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:29 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1036 - blake@cumin1003" * 12:29 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1036 - blake@cumin1003" * 12:28 jmm@dns1004: START - running authdns-update * 12:26 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:26 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:23 blake@cumin1003: START - Cookbook sre.dns.netbox * 12:23 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1036 * 12:23 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1036.eqiad.wmnet with OS trixie * 12:22 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1036.eqiad.wmnet * 12:22 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1082.eqiad.wmnet with OS trixie * 12:22 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1036.eqiad.wmnet * 12:22 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1036.eqiad.wmnet * 12:21 marostegui: Restart mariadb@s7 on db1155 to pick up new filters - [[phab:T431124|T431124]] * 12:21 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 21 hosts with reason: restarting for replication filter * 12:20 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:19 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:14 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2005.codfw.wmnet with reason: host reimage * 12:14 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:08 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:07 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:07 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2005.codfw.wmnet with reason: host reimage * 12:06 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:06 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:06 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:05 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:05 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:04 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-master-eqiad * 12:04 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl1002.eqiad.wmnet * 12:04 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl1002.eqiad.wmnet * 12:04 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:04 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:03 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:03 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:03 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:03 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 11:59 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl1002.eqiad.wmnet * 11:59 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl1002.eqiad.wmnet * 11:59 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl1001.eqiad.wmnet * 11:59 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl1001.eqiad.wmnet * 11:56 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl1001.eqiad.wmnet * 11:56 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl1001.eqiad.wmnet * 11:56 klausman@cumin2002: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-master-eqiad * 11:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2005.codfw.wmnet with OS trixie * 11:49 blake@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on wikikube-worker1160.eqiad.wmnet with reason: Verifying matchers for silence * 11:42 topranks: cr3-eqsin, begin traffic drain to reset PIC and upgrade JunOS [[phab:T429386|T429386]] * 11:41 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs[5004-5006].eqsin.wmnet with reason: upgrade JunOS cr3-eqsin * 11:39 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr1-codfw,cr[2-3]-eqsin,cr3-eqsin IPv6,cr3-eqsin.mgmt with reason: upgrade JunOS cr3-eqsin * 11:36 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=thanos-fe2004.codfw.wmnet * 11:35 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1086.eqiad.wmnet with OS trixie * 11:35 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=thanos-fe2004.codfw.wmnet * 11:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1085.eqiad.wmnet with OS trixie * 11:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1086.eqiad.wmnet with reason: host reimage * 11:10 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1085.eqiad.wmnet with reason: host reimage * 11:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2004.codfw.wmnet with OS trixie * 11:03 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1086.eqiad.wmnet with reason: host reimage * 11:02 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1085.eqiad.wmnet with reason: host reimage * 10:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2004.codfw.wmnet with reason: host reimage * 10:48 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:46 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1086.eqiad.wmnet with OS trixie * 10:46 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1085.eqiad.wmnet with OS trixie * 10:44 cgoubert@deploy1003: Finished deploy [restbase/deploy@2fc37d4]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] (duration: 16m 44s) * 10:43 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2004.codfw.wmnet with reason: host reimage * 10:35 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:27 cgoubert@deploy1003: Started deploy [restbase/deploy@2fc37d4]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] * 10:27 cgoubert@deploy1003: Finished deploy [restbase/deploy@8a25036]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] (duration: 00m 45s) * 10:26 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1117.eqiad.wmnet with OS trixie * 10:26 cgoubert@deploy1003: Started deploy [restbase/deploy@8a25036]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] * 10:26 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host thanos-fe2004 * 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host thanos-fe2004 * 10:22 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1116.eqiad.wmnet with OS trixie * 10:21 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host thanos-fe2004 * 10:21 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) thanos-fe2004.codfw.wmnet 157.32.192.10.in-addr.arpa 7.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:20 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache thanos-fe2004.codfw.wmnet 157.32.192.10.in-addr.arpa 7.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:20 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:20 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host thanos-fe2004 - mvernon@cumin2003" * 10:20 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host thanos-fe2004 - mvernon@cumin2003" * 10:15 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2252: Repooling after reboot * 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:15 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 10:15 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2252: Repooling after reboot * 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1153.eqiad.wmnet * 10:14 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1153.eqiad.wmnet * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 10:14 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 10:12 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 10:12 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host thanos-fe2004 * 10:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2004.codfw.wmnet with OS trixie * 10:07 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1117.eqiad.wmnet with reason: host reimage * 10:03 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1116.eqiad.wmnet with reason: host reimage * 09:58 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:58 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1117.eqiad.wmnet with reason: host reimage * 09:57 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1116.eqiad.wmnet with reason: host reimage * 09:49 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 41 days, 15:00:00 on db2252.codfw.wmnet with reason: Security updates * 09:45 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1117.eqiad.wmnet with OS trixie * 09:45 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1116.eqiad.wmnet with OS trixie * 09:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1153: Security updates * 09:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:28 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:28 root@cumin1003: START - Cookbook sre.mysql.depool depool db1153: Security updates * 09:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1016: Security updates * 09:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:21 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:21 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1016: Security updates * 09:14 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:14 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 08:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1016: Security updates * 08:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:56 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:56 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1016: Security updates * 08:50 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:50 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:45 filippo@dns1006: END - running authdns-update * 08:43 filippo@dns1006: START - running authdns-update * 08:42 godog: switch dumps-nfs address to be shared with rsync/http - [[phab:T411248|T411248]] * 08:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1016: Security updates * 08:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:40 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:40 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1016: Security updates * 08:29 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host cirrussearch1111.eqiad.wmnet * 08:29 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:27 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:27 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:25 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1015: Security updates * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:09 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:09 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1015: Security updates * 07:42 Msz2001: Deployed private patch for Suggested Ivestigations * 07:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1015: Security updates * 07:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:41 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:41 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1015: Security updates * 07:40 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 07:11 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fingerprint warnings - oblivian@cumin1003" * 07:11 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fingerprint warnings - oblivian@cumin1003 * 07:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1024: Security updates * 07:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:11 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:11 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1024: Security updates * 07:10 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fingerprint warnings - oblivian@cumin1003 * 07:10 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fingerprint warnings - oblivian@cumin1003" * 06:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host cirrussearch1111.eqiad.wmnet * 06:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 06:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1024: Security updates * 06:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 06:48 root@cumin1003: START - Cookbook sre.mysql.parsercache * 06:48 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1024: Security updates * 06:42 moritzm: install nginx security updates * 06:31 root@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool pc1024: Security updates * 06:21 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1024: Security updates * 06:19 moritzm: installing php8.2 security updates * 06:15 moritzm: installing php8.4 security updates * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.7 (duration: 02m 38s) * 03:40 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] (duration: 37m 04s) * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 51s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-06 == * 23:30 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] (duration: 09m 39s) * 23:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1078.eqiad.wmnet with OS trixie * 23:26 jdlrobson@deploy1003: jdlrobson, bwang: Continuing with deployment * 23:22 jdlrobson@deploy1003: jdlrobson, bwang: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug) * 23:21 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] * 23:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1078.eqiad.wmnet with reason: host reimage * 23:06 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1078.eqiad.wmnet with reason: host reimage * 22:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1078.eqiad.wmnet with OS trixie * 22:29 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on cirrussearch1114.eqiad.wmnet with reason: reimage on hold until restore completes * 22:22 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on cirrussearch[1079,1115].eqiad.wmnet with reason: reimage on hold until restore completes * 21:18 maryum: Deployed security fix for [[phab:T428006|T428006]] * 20:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1079.eqiad.wmnet with OS trixie * 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1077.eqiad.wmnet with OS trixie * 20:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1115.eqiad.wmnet with OS trixie * 20:25 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1079.eqiad.wmnet with reason: host reimage * 20:21 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1079.eqiad.wmnet with reason: host reimage * 20:15 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] (duration: 08m 14s) * 20:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1077.eqiad.wmnet with reason: host reimage * 20:10 krinkle@deploy1003: krinkle, pushpaktiwari: Continuing with deployment * 20:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1115.eqiad.wmnet with reason: host reimage * 20:08 krinkle@deploy1003: krinkle, pushpaktiwari: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1077.eqiad.wmnet with reason: host reimage * 20:06 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] * 20:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1079.eqiad.wmnet with OS trixie * 20:04 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1115.eqiad.wmnet with reason: host reimage * 19:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1077.eqiad.wmnet with OS trixie * 19:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1115.eqiad.wmnet with OS trixie * 19:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 19:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 18:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1114.eqiad.wmnet with OS trixie * 18:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1114.eqiad.wmnet with reason: host reimage * 18:35 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1114.eqiad.wmnet with reason: host reimage * 18:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1112.eqiad.wmnet with OS trixie * 18:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1114.eqiad.wmnet with OS trixie * 18:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1072.eqiad.wmnet with OS trixie * 18:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1112.eqiad.wmnet with reason: host reimage * 18:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1112.eqiad.wmnet with reason: host reimage * 17:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1072.eqiad.wmnet with reason: host reimage * 17:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1112.eqiad.wmnet with OS trixie * 17:55 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1072.eqiad.wmnet with reason: host reimage * 17:39 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1072.eqiad.wmnet with OS trixie * 17:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1071.eqiad.wmnet with OS trixie * 17:18 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1070.eqiad.wmnet with OS trixie * 17:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1084.eqiad.wmnet with OS trixie * 16:54 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1071.eqiad.wmnet with reason: host reimage * 16:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1084.eqiad.wmnet with reason: host reimage * 16:51 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1070.eqiad.wmnet with reason: host reimage * 16:49 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1084.eqiad.wmnet with reason: host reimage * 16:38 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1071.eqiad.wmnet with OS trixie * 16:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1096.eqiad.wmnet with OS trixie * 16:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1070.eqiad.wmnet with OS trixie * 16:33 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1084.eqiad.wmnet with OS trixie * 16:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1089.eqiad.wmnet with OS trixie * 16:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1103.eqiad.wmnet with OS trixie * 16:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1096.eqiad.wmnet with reason: host reimage * 16:14 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1096.eqiad.wmnet with reason: host reimage * 16:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1089.eqiad.wmnet with reason: host reimage * 16:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1103.eqiad.wmnet with reason: host reimage * 16:02 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1003.eqiad.wmnet with OS bookworm * 16:01 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1089.eqiad.wmnet with reason: host reimage * 16:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1103.eqiad.wmnet with reason: host reimage * 15:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1096.eqiad.wmnet with OS trixie * 15:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1080.eqiad.wmnet with OS trixie * 15:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1089.eqiad.wmnet with OS trixie * 15:45 dancy@deploy1003: Installation of scap version "4.272.0" completed for 158 hosts * 15:43 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1103.eqiad.wmnet with OS trixie * 15:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1113.eqiad.wmnet with OS trixie * 15:41 dancy@deploy1003: Installing scap version "4.272.0" for 158 host(s) * 15:40 klausman@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 15:39 klausman@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 15:38 klausman@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 15:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1069.eqiad.wmnet with OS trixie * 15:37 klausman@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 15:36 klausman@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 15:34 klausman@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 15:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1080.eqiad.wmnet with reason: host reimage * 15:27 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1080.eqiad.wmnet with reason: host reimage * 15:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1113.eqiad.wmnet with reason: host reimage * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1069.eqiad.wmnet with reason: host reimage * 15:18 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1113.eqiad.wmnet with reason: host reimage * 15:16 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1069.eqiad.wmnet with reason: host reimage * 15:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:11 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1080.eqiad.wmnet with OS trixie * 15:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1113.eqiad.wmnet with OS trixie * 15:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1003.eqiad.wmnet with reason: host reimage * 14:47 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1003.eqiad.wmnet with OS bookworm * 14:33 elukey: rolled out spicerack on all cumin nodes - [[phab:T429699|T429699]] * 14:32 elukey: upgrade all bookworm hosts to pywmflib 3.1 - [[phab:T430552|T430552]] * 14:14 marostegui: Setup x4 eqiad topology [[phab:T404715|T404715]] * 14:13 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 14:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2230.codfw.wmnet * 14:07 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2230.codfw.wmnet * 13:59 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[2001-2002].codfw.wmnet * 13:51 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 13:45 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.major-upgrade (exit_code=97) * 13:45 cwilliams@cumin1003: dbmaint on s4@codfw [[phab:T429893|T429893]] * 13:45 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 13:42 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-master-codfw * 13:42 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl2002.codfw.wmnet * 13:42 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl2002.codfw.wmnet * 13:38 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl2002.codfw.wmnet * 13:38 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl2002.codfw.wmnet * 13:38 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl2001.codfw.wmnet * 13:38 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl2001.codfw.wmnet * 13:35 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl2001.codfw.wmnet * 13:35 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl2001.codfw.wmnet * 13:35 klausman@cumin2002: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-master-codfw * 12:30 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] (duration: 25m 11s) * 12:24 krinkle@deploy1003: krinkle: Continuing with deployment * 12:10 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2048.codfw.wmnet * 12:09 krinkle@deploy1003: krinkle: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:08 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2048.codfw.wmnet * 12:05 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] * 11:57 moritzm: installing curl security updates * 11:49 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:31 moritzm: installing nano security updates * 11:07 moritzm: failover Ganeti master in codfw to ganeti2032 [[phab:T430909|T430909]] * 11:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:04 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest1005.eqiad.wmnet with OS trixie * 11:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:50 jmm@dns1004: END - running authdns-update * 10:47 jmm@dns1004: START - running authdns-update * 10:47 jmm@dns1004: START - running authdns-update * 10:46 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:44 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest1005.eqiad.wmnet with reason: host reimage * 10:38 elukey@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest1005.eqiad.wmnet with reason: host reimage * 10:31 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:31 marostegui: Setup x4 codfw topology [[phab:T404715|T404715]] * 10:31 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 10:24 elukey: spicerack 13.0.0 deployed on cumin2002 * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 10:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 10:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 10:21 elukey@cumin2002: START - Cookbook sre.hosts.reimage for host sretest1005.eqiad.wmnet with OS trixie * 10:20 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:19 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:17 elukey: uploaded spicerack_13.0.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia * 09:54 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:52 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:20 elukey: upgrade all bullseye hosts to pywmflib 3.1 - [[phab:T430552|T430552]] * 09:10 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1015.eqiad.wmnet,service=s4 * 09:10 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1015.eqiad.wmnet,service=s6 * 09:07 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 08:58 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:56 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 08:56 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 08:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 08:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 08:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin2002.codfw.wmnet * 08:06 godog: remove cloudvirt1046, cloudvirt1062, cloudvirt1074, cloudvirt1075 from maintenance aggregate and put them in network-ovs - [[phab:T424802|T424802]] * 08:00 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin2002.codfw.wmnet * 07:58 hashar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] (duration: 32m 53s) * 07:57 fabfur: repooled cp4038 * 07:57 fabfur@cumin1003: conftool action : set/pooled=yes; selector: name=cp4038.* * 07:53 moritzm: installing pyjwt security updates * 07:47 moritzm: installing openjpeg2 security updates * 07:45 hashar@deploy1003: vadymts1, hashar: Continuing with deployment * 07:43 hashar@deploy1003: vadymts1, hashar: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:38 moritzm: installing python-urllib3 security updates * 07:37 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 07:30 fabfur: depooled cp4038 to investigate on possible maxmind failure * 07:30 fabfur@cumin1003: conftool action : set/pooled=no; selector: name=cp4038.* * 07:30 fabfur@cumin1003: conftool action : set/pooled=yes; selector: name=cp4038.* * 07:29 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 07:25 hashar@deploy1003: Started scap sync-world: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] * 06:13 moritzm: installing Linux 6.12.95 on trixie hosts * 05:20 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s6 * 05:20 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s4 * 05:19 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1015.eqiad.wmnet with reason: cloning * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 08s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-05 == * 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 01m 08s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-04 == * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 58s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-03 == * 17:08 topranks: revert protocol preference changes on cr3-ulsfo after upgrade * 16:53 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on cr2-eqord with reason: upgrade JunOS cr3-ulsfo * 16:53 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on cr4-ulsfo with reason: upgrade JunOS cr3-ulsfo * 16:48 topranks: reboot cr3-ulsfo to upgrade JunOS and reset linecard [[phab:T424839|T424839]] * 15:52 topranks: adjust outbound BGP policies on cr3-ulsfo to drain router of traffic [[phab:T424839|T424839]] * 15:45 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on lvs[4008-4010].ulsfo.wmnet with reason: upgrade JunOS cr3-ulsfo * 15:44 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on asw1-[22-23]-ulsfo,cr3-ulsfo,cr3-ulsfo IPv6,cr3-ulsfo.mgmt with reason: upgrade JunOS cr3-ulsfo * 15:36 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 15:35 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 15:35 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 14:40 cmooney@dns3003: END - running authdns-update * 14:26 cmooney@dns3003: START - running authdns-update * 14:26 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:26 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to ulsfo - cmooney@cumin1003" * 14:19 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to ulsfo - cmooney@cumin1003" * 14:16 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:38 sukhe@dns1004: END - running authdns-update * 13:35 sukhe@dns1004: START - running authdns-update * 13:26 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 13:26 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 13:26 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet * 13:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 13:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 13:16 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host sretest1005.eqiad.wmnet * 13:16 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 13:16 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 13:15 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 13:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:14 moritzm: imported samplicator 1.3.8rc1-1+deb13u1 to trixie-wikimedia/main [[phab:T337208|T337208]] * 13:13 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:07 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 13:07 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 13:02 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:02 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:58 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:57 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:57 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:53 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet * 12:50 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 12:47 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:41 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:40 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:39 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:32 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet * 12:26 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet * 12:23 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2005.wikimedia.org * 12:19 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2005.wikimedia.org * 12:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 12:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup[2004-2007].codfw.wmnet * 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[2004-2007].codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin2003" * 12:15 jynus@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[2004-2007].codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin2003" * 12:09 jynus@cumin2003: START - Cookbook sre.dns.netbox * 11:58 jynus@cumin2003: START - Cookbook sre.hosts.decommission for hosts backup[2004-2007].codfw.wmnet * 10:40 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup[1004-1007].eqiad.wmnet * 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[1004-1007].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 10:01 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[1004-1007].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 09:52 jynus@cumin1003: START - Cookbook sre.dns.netbox * 09:39 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:36 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup[1004-1007].eqiad.wmnet * 09:36 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:25 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 09:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 09:16 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 09:05 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:04 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:00 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:59 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:57 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:55 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:50 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 08:50 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 08:49 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 08:49 atsukoito: depooling cirrussearch in codfw because of regression after upgrade [[phab:T431091|T431091]] * 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts mirror1001.wikimedia.org * 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: mirror1001.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 08:29 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: mirror1001.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 08:18 jmm@cumin2003: START - Cookbook sre.dns.netbox * 08:11 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts mirror1001.wikimedia.org * 06:15 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 18s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-02 == * 22:55 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host contint1003.wikimedia.org with OS trixie * 22:29 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on contint1003.wikimedia.org with reason: host reimage * 22:23 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on contint1003.wikimedia.org with reason: host reimage * 22:05 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host contint1003.wikimedia.org with OS trixie * 22:03 mutante: contint1003 (zuul.wikimedia.org) - reimaging because of [[phab:T430510|T430510]]#12067628 [[phab:T418521|T418521]] * 22:03 dzahn@cumin2002: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on zuul.wikimedia.org with reason: reimage * 21:39 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 18s) * 21:39 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 21:20 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1003.eqiad.wmnet, repooling source-only afterwards * 21:19 sbassett: Deployed security fix for [[phab:T428829|T428829]] * 20:58 cmooney@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Release v0.11.2 update for new Aerleon - cmooney@cumin1003 * 20:55 cmooney@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Release v0.11.2 update for new Aerleon - cmooney@cumin1003 * 20:40 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] (duration: 12m 35s) * 20:36 arlolra@deploy1003: cscott, arlolra: Continuing with deployment * 20:35 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 20s) * 20:35 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 20:33 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host contint2003.wikimedia.org with OS trixie * 20:31 arlolra@deploy1003: cscott, arlolra: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Cha * 20:28 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] * 20:17 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] (duration: 08m 13s) * 20:14 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on contint2003.wikimedia.org with reason: host reimage * 20:13 sbassett@deploy1003: sbassett: Continuing with deployment * 20:11 sbassett@deploy1003: sbassett: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:09 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] * 20:08 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 20:08 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 20:08 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on contint2003.wikimedia.org with reason: host reimage * 20:05 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1003.eqiad.wmnet, repooling source-only afterwards * 19:49 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host contint2003.wikimedia.org with OS trixie * 19:48 mutante: contint2003 - reimaging because of [[phab:T430510|T430510]]#12067628 [[phab:T418521|T418521]] * 18:39 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 18:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 18:13 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2002.codfw.wmnet -> wcqs2003.codfw.wmnet, repooling source-only afterwards * 17:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1003.eqiad.wmnet with OS bookworm * 17:52 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1005.eqiad.wmnet * 17:52 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1005.eqiad.wmnet * 17:51 jasmine@cumin2002: conftool action : set/pooled=yes:weight=10; selector: name=wikikube-ctrl1005.eqiad.wmnet * 17:48 jasmine_: homer "cr*eqiad*" commit "Added new stacked control plane wikikube-ctrl1005" * 17:44 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply * 17:44 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply * 17:31 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] (duration: 09m 33s) * 17:26 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 17:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1003.eqiad.wmnet with reason: host reimage * 17:23 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:21 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] * 17:18 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1003.eqiad.wmnet with reason: host reimage * 17:16 rscout@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply * 17:16 rscout@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply * 17:16 rscout@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply * 17:15 rscout@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply * 17:12 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on wcqs[2002-2003].codfw.wmnet,wcqs1002.eqiad.wmnet with reason: reimaging hosts * 17:08 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 17:08 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 17:08 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 17:07 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 17:05 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 17:05 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 17:03 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "running to make sure all updates are synced - cmooney@cumin1003" * 17:03 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "running to make sure all updates are synced - cmooney@cumin1003" * 17:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs1003 * 17:00 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs1003 * 17:00 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 17:00 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1003.eqiad.wmnet with OS bookworm * 16:58 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Re-running - btullis@cumin1003" * 16:58 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Re-running - btullis@cumin1003" * 16:58 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2002.codfw.wmnet -> wcqs2003.codfw.wmnet, repooling source-only afterwards * 16:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-master1004.eqiad.wmnet with OS bookworm * 16:58 btullis@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 16:57 tappof: bump space for prometheus k8s-aux in eqiad * 16:55 cmooney@dns3003: END - running authdns-update * 16:55 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:55 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to eqsin - cmooney@cumin1003" * 16:55 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to eqsin - cmooney@cumin1003" * 16:53 cmooney@dns3003: START - running authdns-update * 16:52 ryankemper: [ml-serve-eqiad] Cleared out 1302 failed (Evicted) pods: `kubectl -n llm delete pods --field-selector=status.phase=Failed`, freeing calico-kube-controllers from OOM crashloop (evictions were caused by disk pressure) * 16:49 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 16:46 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:39 rzl@dns1004: END - running authdns-update * 16:37 rzl@dns1004: START - running authdns-update * 16:36 rzl@dns1004: START - running authdns-update * 16:35 rzl@deploy1003: Finished scap sync-world: [[phab:T416623|T416623]] (duration: 10m 19s) * 16:34 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 16:33 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-master1004.eqiad.wmnet with reason: host reimage * 16:30 rzl@deploy1003: rzl: Continuing with deployment * 16:28 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-master1004.eqiad.wmnet with reason: host reimage * 16:26 rzl@deploy1003: rzl: [[phab:T416623|T416623]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:25 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 16:25 rzl@deploy1003: Started scap sync-world: [[phab:T416623|T416623]] * 16:25 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 16:24 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 16:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: sync * 16:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: sync * 16:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-master1004.eqiad.wmnet with OS bookworm * 16:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-master1003.eqiad.wmnet with OS bookworm * 16:11 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 16:11 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 16:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Security updates * 16:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 16:08 root@cumin1003: START - Cookbook sre.mysql.parsercache * 16:08 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Security updates * 15:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-master1003.eqiad.wmnet with reason: host reimage * 15:54 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:54 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:54 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:54 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-master1003.eqiad.wmnet with reason: host reimage * 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Security updates * 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:45 root@cumin1003: START - Cookbook sre.mysql.parsercache * 15:45 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Security updates * 15:42 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-master1003.eqiad.wmnet with OS bookworm * 15:24 moritzm: installing busybox updates from bookworm point release * 15:20 moritzm: installing busybox updates from trixie point release * 15:15 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1021: Security updates * 15:15 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:15 root@cumin1003: START - Cookbook sre.mysql.parsercache * 15:15 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1021: Security updates * 15:13 moritzm: installing giflib security updates * 15:08 moritzm: installing Tomcat security updates * 14:57 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 14:56 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 14:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:53 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Unblock taavi - oblivian@cumin1003" * 14:53 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Unblock taavi - oblivian@cumin1003 * 14:53 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1021: Security updates * 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:53 root@cumin1003: START - Cookbook sre.mysql.parsercache * 14:53 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1021: Security updates * 14:53 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Unblock taavi - oblivian@cumin1003 * 14:52 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Unblock taavi - oblivian@cumin1003" * 14:46 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94711 and previous config saved to /var/cache/conftool/dbconfig/20260702-144644-fceratto.json * 14:36 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205', diff saved to https://phabricator.wikimedia.org/P94709 and previous config saved to /var/cache/conftool/dbconfig/20260702-143636-fceratto.json * 14:32 moritzm: installing libdbi-perl security updates * 14:26 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205', diff saved to https://phabricator.wikimedia.org/P94708 and previous config saved to /var/cache/conftool/dbconfig/20260702-142628-fceratto.json * 14:16 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94707 and previous config saved to /var/cache/conftool/dbconfig/20260702-141621-fceratto.json * 14:12 moritzm: installing rsync security updates * 14:11 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox) * 14:10 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94706 and previous config saved to /var/cache/conftool/dbconfig/20260702-140959-fceratto.json * 14:09 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2205.codfw.wmnet with reason: Maintenance * 14:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2205: Repooling after switchover * 14:07 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-test-master1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 14:06 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 14:06 Tran: Deployed patch for [[phab:T427287|T427287]] * 14:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:59 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2205: Repooling after switchover * 13:59 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2205: Repooling after switchover * 13:59 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:55 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2205: Repooling after switchover * 13:55 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2205 [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94704 and previous config saved to /var/cache/conftool/dbconfig/20260702-135505-fceratto.json * 13:54 moritzm: installing sed security updates * 13:53 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:52 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2209 to s3 primary [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94703 and previous config saved to /var/cache/conftool/dbconfig/20260702-135235-fceratto.json * 13:52 federico3: Starting s3 codfw failover from db2205 to db2209 - [[phab:T430912|T430912]] * 13:51 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:51 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 13:48 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:47 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2209 with weight 0 [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94702 and previous config saved to /var/cache/conftool/dbconfig/20260702-134719-fceratto.json * 13:47 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Primary switchover s3 [[phab:T430912|T430912]] * 13:44 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:44 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:44 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:40 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 13:38 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 13:37 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 13:36 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 13:36 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:34 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 13:30 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:29 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:29 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:27 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:26 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:25 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling restart_daemons on A:wikidough * 13:23 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 13:22 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 13:17 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 13:17 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns1004.wikimedia.org * 13:12 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:11 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart (exit_code=97) rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough * 13:11 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=97) rolling restart_daemons on A:wikidough * 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough * 13:09 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] (duration: 07m 20s) * 13:05 aude@deploy1003: jdrewniak, aude: Continuing with deployment * 13:04 aude@deploy1003: jdrewniak, aude: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:02 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] * 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts wdqs-categories1001.eqiad.wmnet * 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: wdqs-categories1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 12:10 jmm@dns1004: END - running authdns-update * 12:07 jmm@dns1004: START - running authdns-update * 11:51 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: wdqs-categories1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 11:44 btullis@cumin1003: START - Cookbook sre.dns.netbox * 11:42 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 11:42 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 11:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet * 11:39 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts wdqs-categories1001.eqiad.wmnet * 11:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet * 11:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet * 11:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet * 11:29 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 11:29 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 10:57 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2214: Repooling * 10:49 jmm@dns1004: END - running authdns-update * 10:47 jmm@dns1004: START - running authdns-update * 10:31 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94698 and previous config saved to /var/cache/conftool/dbconfig/20260702-103146-fceratto.json * 10:21 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213', diff saved to https://phabricator.wikimedia.org/P94696 and previous config saved to /var/cache/conftool/dbconfig/20260702-102137-fceratto.json * 10:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:19 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb1017.eqiad.wmnet * 10:18 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 10:18 fceratto@cumin1003: Removing es1033 from zarcillo [[phab:T408772|T408772]] * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts es1033.eqiad.wmnet * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: es1033.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:14 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: es1033.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:13 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb1017.eqiad.wmnet * 10:12 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2214.codfw.wmnet * 10:12 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2214.codfw.wmnet * 10:12 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2214: Repooling * 10:11 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213', diff saved to https://phabricator.wikimedia.org/P94693 and previous config saved to /var/cache/conftool/dbconfig/20260702-101130-fceratto.json * 10:10 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:10 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:03 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts es1033.eqiad.wmnet * 10:03 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 10:01 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94691 and previous config saved to /var/cache/conftool/dbconfig/20260702-100122-fceratto.json * 09:55 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94690 and previous config saved to /var/cache/conftool/dbconfig/20260702-095529-fceratto.json * 09:55 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2213.codfw.wmnet with reason: Maintenance * 09:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 09:53 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2213: Repooling after switchover * 09:51 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover * 09:44 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2213: Repooling after switchover * 09:39 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover * 09:39 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2213 [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94688 and previous config saved to /var/cache/conftool/dbconfig/20260702-093859-fceratto.json * 09:36 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2192 to s5 primary [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94687 and previous config saved to /var/cache/conftool/dbconfig/20260702-093650-fceratto.json * 09:36 federico3: Starting s5 codfw failover from db2213 to db2192 - [[phab:T430923|T430923]] * 09:30 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94686 and previous config saved to /var/cache/conftool/dbconfig/20260702-093004-fceratto.json * 09:24 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2192 with weight 0 [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94685 and previous config saved to /var/cache/conftool/dbconfig/20260702-092455-fceratto.json * 09:24 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 23 hosts with reason: Primary switchover s5 [[phab:T430923|T430923]] * 09:19 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220', diff saved to https://phabricator.wikimedia.org/P94684 and previous config saved to /var/cache/conftool/dbconfig/20260702-091957-fceratto.json * 09:16 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] (duration: 06m 57s) * 09:13 moritzm: installing libgcrypt20 security updates * 09:12 kharlan@deploy1003: kharlan: Continuing with deployment * 09:11 kharlan@deploy1003: kharlan: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:09 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220', diff saved to https://phabricator.wikimedia.org/P94683 and previous config saved to /var/cache/conftool/dbconfig/20260702-090950-fceratto.json * 09:09 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] * 09:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 09:01 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] (duration: 07m 07s) * 08:59 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94682 and previous config saved to /var/cache/conftool/dbconfig/20260702-085942-fceratto.json * 08:57 kharlan@deploy1003: kharlan: Continuing with deployment * 08:56 kharlan@deploy1003: kharlan: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:54 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] * 08:52 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:52 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:52 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94681 and previous config saved to /var/cache/conftool/dbconfig/20260702-085237-fceratto.json * 08:52 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2220.codfw.wmnet with reason: Maintenance * 08:43 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:40 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 08:25 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] (duration: 11m 44s) * 08:21 cscott@deploy1003: cscott: Continuing with deployment * 08:16 cscott@deploy1003: cscott: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:14 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] * 08:08 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 08:08 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1244: Migration of db1244.eqiad.wmnet completed * 08:02 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:02 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:01 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] (duration: 18m 58s) * 08:01 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:59 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 07:59 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:59 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:59 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2006.wikimedia.org * 07:58 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:57 cscott@deploy1003: cscott: Continuing with deployment * 07:56 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:56 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:56 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:55 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:55 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:55 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:54 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2006.wikimedia.org * 07:54 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:54 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 07:44 cscott@deploy1003: cscott: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:44 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2005.wikimedia.org * 07:44 moritzm: installing node-lodash security updates * 07:42 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] * 07:39 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2005.wikimedia.org * 07:30 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] (duration: 07m 28s) * 07:26 cscott@deploy1003: ssastry, cscott: Continuing with deployment * 07:25 cscott@deploy1003: ssastry, cscott: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:23 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1244: Migration of db1244.eqiad.wmnet completed * 07:22 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] * 07:16 wmde-fisch@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] (duration: 06m 55s) * 07:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1244.eqiad.wmnet with OS trixie * 07:11 wmde-fisch@deploy1003: wmde-fisch: Continuing with deployment * 07:11 wmde-fisch@deploy1003: wmde-fisch: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:09 wmde-fisch@deploy1003: Started scap sync-world: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] * 06:54 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1244.eqiad.wmnet with reason: host reimage * 06:50 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1244.eqiad.wmnet with reason: host reimage * 06:38 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1250.eqiad.wmnet with OS trixie * 06:34 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db1244.eqiad.wmnet with OS trixie * 06:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1244: Upgrading db1244.eqiad.wmnet * 06:25 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1244: Upgrading db1244.eqiad.wmnet * 06:25 cwilliams@cumin1003: dbmaint on s4@eqiad [[phab:T429893|T429893]] * 06:25 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 06:15 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1250.eqiad.wmnet with reason: host reimage * 06:14 cwilliams@dns1006: END - running authdns-update * 06:12 cwilliams@dns1006: START - running authdns-update * 06:11 cwilliams@dns1006: END - running authdns-update * 06:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db1244 [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94676 and previous config saved to /var/cache/conftool/dbconfig/20260702-061059-cwilliams.json * 06:09 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1250.eqiad.wmnet with reason: host reimage * 06:09 cwilliams@dns1006: START - running authdns-update * 06:08 aokoth@cumin1003: END (PASS) - Cookbook sre.vrts.upgrade (exit_code=0) on VRTS host vrts1003.eqiad.wmnet * 06:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db1160 to s4 primary and set section read-write [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94675 and previous config saved to /var/cache/conftool/dbconfig/20260702-060746-cwilliams.json * 06:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Set s4 eqiad as read-only for maintenance - [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94674 and previous config saved to /var/cache/conftool/dbconfig/20260702-060704-cwilliams.json * 06:06 cezmunsta: Starting s4 eqiad failover from db1244 to db1160 - [[phab:T430817|T430817]] * 06:04 aokoth@cumin1003: START - Cookbook sre.vrts.upgrade on VRTS host vrts1003.eqiad.wmnet * 05:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db1160 with weight 0 [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94673 and previous config saved to /var/cache/conftool/dbconfig/20260702-055927-cwilliams.json * 05:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 40 hosts with reason: Primary switchover s4 [[phab:T430817|T430817]] * 05:55 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1250.eqiad.wmnet with OS trixie * 05:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on db1250.eqiad.wmnet with reason: m3 master switchover [[phab:T430158|T430158]] * 05:39 marostegui: Failover m3 (phabricator) from db1250 to db1228 - [[phab:T430158|T430158]] * 05:32 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2234].codfw.wmnet,db[1217,1228,1250].eqiad.wmnet with reason: m3 master switchover [[phab:T430158|T430158]] * 04:45 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] (duration: 09m 08s) * 04:41 tstarling@deploy1003: tstarling, reedy: Continuing with deployment * 04:38 tstarling@deploy1003: tstarling, reedy: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 04:36 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 59s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:16 ryankemper: [[phab:T429844|T429844]] [opensearch] completed `cirrussearch2111` reimage; all codfw search clusters are green, all nodes now report `OpenSearch 2.19.5`, and the temporary chi voting exclusion has been removed * 00:57 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2111.codfw.wmnet with OS trixie * 00:29 ryankemper: [[phab:T429844|T429844]] [opensearch] depooled codfw search-omega/search-psi discovery records to match existing codfw search depool during OpenSearch 2.19 migration * 00:29 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2111.codfw.wmnet with reason: host reimage * 00:29 ryankemper@cumin2002: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 00:29 ryankemper@cumin2002: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 00:22 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2111.codfw.wmnet with reason: host reimage * 00:01 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2111.codfw.wmnet with OS trixie * 00:00 ryankemper: [[phab:T429844|T429844]] [opensearch] chi cluster recovered after stopping `opensearch_1@production-search-codfw` on `cirrussearch2111` == 2026-07-01 == * 23:59 ryankemper: [[phab:T429844|T429844]] [opensearch] stopped `opensearch_1@production-search-codfw` on `cirrussearch2111` after chi cluster-manager election churn following `voting_config_exclusions` POST; hoping this triggers a re-election * 23:52 cscott@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 23:51 cscott@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 23:51 cscott@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 23:50 cscott@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2003.codfw.wmnet with OS bookworm * 22:29 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 22:13 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 22:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2084.codfw.wmnet with OS trixie * 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2003.codfw.wmnet with reason: host reimage * 22:03 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 22:01 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2003.codfw.wmnet with reason: host reimage * 21:50 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 21:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2084.codfw.wmnet with reason: host reimage * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2003 * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2003 * 21:42 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2003 * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2003.codfw.wmnet 45.48.192.10.in-addr.arpa 5.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:42 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2003.codfw.wmnet 45.48.192.10.in-addr.arpa 5.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2003 - bking@cumin2003" * 21:42 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2003 - bking@cumin2003" * 21:36 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2084.codfw.wmnet with reason: host reimage * 21:35 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:34 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2003 * 21:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2003.codfw.wmnet with OS bookworm * 21:19 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2084.codfw.wmnet with OS trixie * 21:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2081.codfw.wmnet with OS trixie * 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2108.codfw.wmnet with OS trixie * 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2081.codfw.wmnet with reason: host reimage * 20:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2081.codfw.wmnet with reason: host reimage * 20:28 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2081.codfw.wmnet with OS trixie * 20:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2108.codfw.wmnet with reason: host reimage * 20:19 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2108.codfw.wmnet with reason: host reimage * 19:59 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2108.codfw.wmnet with OS trixie * 19:46 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2093.codfw.wmnet with OS trixie * 19:44 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 19:44 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jasmine@cumin2002" * 19:43 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jasmine@cumin2002" * 19:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2080.codfw.wmnet with OS trixie * 19:28 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 19:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2093.codfw.wmnet with reason: host reimage * 19:18 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 19:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2093.codfw.wmnet with reason: host reimage * 19:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2080.codfw.wmnet with reason: host reimage * 19:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2080.codfw.wmnet with reason: host reimage * 18:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2093.codfw.wmnet with OS trixie * 18:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2080.codfw.wmnet with OS trixie * 18:27 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 18:18 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] (duration: 09m 15s) * 18:13 jgiannelos@deploy1003: jgiannelos, neriah: Continuing with deployment * 18:11 jgiannelos@deploy1003: jgiannelos, neriah: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:09 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] * 17:40 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 16:58 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 30 hosts * 16:57 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for 30 hosts * 16:52 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2202.codfw.wmnet * 16:52 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2202.codfw.wmnet * 16:51 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt * 16:51 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt * 16:51 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lvs2012.codfw.wmnet * 16:51 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for lvs2012.codfw.wmnet * 16:49 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2076.codfw.wmnet with OS trixie * 16:49 brett: Start pybal on lvs2012 - [[phab:T429861|T429861]] * 16:49 pt1979@cumin1003: END (ERROR) - Cookbook sre.hosts.remove-downtime (exit_code=97) for 59 hosts * 16:48 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for 59 hosts * 16:42 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2061.codfw.wmnet with OS trixie * 16:30 dancy@deploy1003: Installation of scap version "4.271.0" completed for 2 hosts * 16:28 dancy@deploy1003: Installing scap version "4.271.0" for 2 host(s) * 16:23 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2076.codfw.wmnet with reason: host reimage * 16:19 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2061.codfw.wmnet with reason: host reimage * 16:18 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2076.codfw.wmnet with reason: host reimage * 16:18 jasmine@dns1004: END - running authdns-update * 16:16 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host restbase2039.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 16:16 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host restbase2039.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 16:16 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2061.codfw.wmnet with reason: host reimage * 16:15 jasmine@dns1004: START - running authdns-update * 16:14 jasmine@dns1004: END - running authdns-update * 16:12 jasmine@dns1004: START - running authdns-update * 16:07 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2202.codfw.wmnet with reason: maintenance * 16:06 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt with reason: Junos upograde * 16:00 papaul: ongoing maintenance on lsw1-b2-codfw * 16:00 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2076.codfw.wmnet with OS trixie * 15:59 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt * 15:59 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt * 15:57 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2061.codfw.wmnet with OS trixie * 15:55 pt1979@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2042,2046].codfw.wmnet * 15:55 pt1979@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2042,2046].codfw.wmnet * 15:51 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 15:51 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2220: Repooling after switchover * 15:50 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 15:50 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 15:48 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2092.codfw.wmnet with OS trixie * 15:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 15:40 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 15:38 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 15:37 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 15:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply * 15:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply * 15:32 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 15:32 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 15:30 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 15:29 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 15:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 15:25 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:22 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs2012.codfw.wmnet with reason: Rack B2 maintenance - [[phab:T429861|T429861]] * 15:21 brett: Stopping pybal on lvs2012 in preparation for codfw rack b2 maintenance - [[phab:T429861|T429861]] * 15:20 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2092.codfw.wmnet with reason: host reimage * 15:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:12 _joe_: restarted manually alertmanager-irc-relay * 15:12 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2092.codfw.wmnet with reason: host reimage * 15:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:12 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt with reason: Junos upograde * 15:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover * 15:07 pt1979@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2042,2046].codfw.wmnet * 15:06 pt1979@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2042,2046].codfw.wmnet * 15:02 papaul: ongoing maintenance on lsw1-a8-codfw * 14:31 topranks: POWERING DOWN CR1-EQIAD for line card installation [[phab:T426343|T426343]] * 14:31 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] (duration: 08m 57s) * 14:29 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:26 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 14:24 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:22 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] * 14:22 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover * 14:16 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:15 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover * 14:14 topranks: re-enable routing-engine graceful-failover on cr1-eqiad [[phab:T417873|T417873]] * 14:13 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:13 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2220: Repooling after switchover * 14:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:12 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:12 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:11 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:08 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] (duration: 10m 01s) * 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:07 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2220 [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94664 and previous config saved to /var/cache/conftool/dbconfig/20260701-140729-fceratto.json * 14:06 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:06 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:06 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:05 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2159 to s7 primary [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94663 and previous config saved to /var/cache/conftool/dbconfig/20260701-140503-fceratto.json * 14:04 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:04 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 14:04 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 14:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:04 dreamyjazz@deploy1003: anzx, dreamyjazz: Continuing with deployment * 14:04 federico3: Starting s7 codfw failover from db2220 to db2159 - [[phab:T430826|T430826]] * 14:03 jmm@dns1004: END - running authdns-update * 14:03 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:03 topranks: flipping cr1-eqiad active routing-enginer back to RE0 [[phab:T417873|T417873]] * 14:03 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudsw1-c8-eqiad,cloudsw1-d5-eqiad with reason: router upgrades eqiad * 14:01 jmm@dns1004: START - running authdns-update * 14:00 dreamyjazz@deploy1003: anzx, dreamyjazz: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:59 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2159 with weight 0 [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94662 and previous config saved to /var/cache/conftool/dbconfig/20260701-135906-fceratto.json * 13:58 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] * 13:57 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s7 [[phab:T430826|T430826]] * 13:56 topranks: reboot routing-enginer RE0 on cr1-eqiad [[phab:T417873|T417873]] * 13:48 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1006.wikimedia.org * 13:44 atsuko@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cirrussearch2092.codfw.wmnet with OS trixie * 13:43 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1006.wikimedia.org * 13:41 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2092.codfw.wmnet with OS trixie * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1005.wikimedia.org * 13:37 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1005.wikimedia.org * 13:37 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on pfw1-eqiad with reason: router upgrades eqiad * 13:35 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on lvs[1017-1020].eqiad.wmnet with reason: router upgrades eqiad * 13:34 caro@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] (duration: 07m 59s) * 13:30 caro@deploy1003: caro: Continuing with deployment * 13:28 caro@deploy1003: caro: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:27 topranks: route-engine failover cr1-eqiad * 13:26 caro@deploy1003: Started scap sync-world: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] * 13:15 topranks: rebooting routing-engine 1 on cr1-eqiad [[phab:T417873|T417873]] * 13:13 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] (duration: 08m 29s) * 13:13 moritzm: installing qemu security updates * 13:11 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 13:11 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 13:09 jgiannelos@deploy1003: jgiannelos: Continuing with deployment * 13:08 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 13:07 jgiannelos@deploy1003: jgiannelos: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:06 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 13:06 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2214.codfw.wmnet with reason: Maintenance * 13:05 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2214: Repooling after switchover * 13:05 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] * 13:04 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2214: Repooling after switchover * 13:04 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2214 [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94660 and previous config saved to /var/cache/conftool/dbconfig/20260701-130413-fceratto.json * 13:01 moritzm: installing python3.13 security updates * 13:00 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2229 to s6 primary [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94659 and previous config saved to /var/cache/conftool/dbconfig/20260701-125959-fceratto.json * 12:59 federico3: Starting s6 codfw failover from db2214 to db2229 - [[phab:T430814|T430814]] * 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on 13 hosts with reason: router upgrade and line card install * 12:51 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2229 with weight 0 [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94658 and previous config saved to /var/cache/conftool/dbconfig/20260701-125149-fceratto.json * 12:51 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 21 hosts with reason: Primary switchover s6 [[phab:T430814|T430814]] * 12:50 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2189.codfw.wmnet * 12:50 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2189.codfw.wmnet * 12:42 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2100.codfw.wmnet with OS trixie * 12:38 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2083.codfw.wmnet with OS trixie * 12:19 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2083.codfw.wmnet with reason: host reimage * 12:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 12:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2240: Migration of db2240.codfw.wmnet completed * 12:14 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2100.codfw.wmnet with reason: host reimage * 12:09 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2083.codfw.wmnet with reason: host reimage * 12:09 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2100.codfw.wmnet with reason: host reimage * 12:00 topranks: drain traffic on cr1-eqiad to allow for line card install and JunOS upgrade [[phab:T426343|T426343]] * 11:52 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2083.codfw.wmnet with OS trixie * 11:50 cmooney@dns2005: END - running authdns-update * 11:49 cmooney@dns2005: START - running authdns-update * 11:48 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2100.codfw.wmnet with OS trixie * 11:40 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/zotero: apply * 11:40 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/zotero: apply * 11:36 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/zotero: apply * 11:36 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/zotero: apply * 11:31 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2240: Migration of db2240.codfw.wmnet completed * 11:30 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply * 11:28 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply * 11:27 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:27 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:27 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:27 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:27 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:23 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2240.codfw.wmnet with OS trixie * 11:20 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:20 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:17 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:16 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:16 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:15 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2086.codfw.wmnet with OS trixie * 11:14 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2106.codfw.wmnet with OS trixie * 11:14 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:13 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:12 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:09 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2115.codfw.wmnet with OS trixie * 11:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2240.codfw.wmnet with reason: host reimage * 11:00 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2240.codfw.wmnet with reason: host reimage * 10:53 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2106.codfw.wmnet with reason: host reimage * 10:49 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2086.codfw.wmnet with reason: host reimage * 10:44 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2115.codfw.wmnet with reason: host reimage * 10:44 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2240.codfw.wmnet with OS trixie * 10:44 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2086.codfw.wmnet with reason: host reimage * 10:42 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2106.codfw.wmnet with reason: host reimage * 10:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2240: Upgrading db2240.codfw.wmnet * 10:41 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2240: Upgrading db2240.codfw.wmnet * 10:41 cwilliams@cumin1003: dbmaint on s4@codfw [[phab:T429893|T429893]] * 10:40 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 10:39 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2115.codfw.wmnet with reason: host reimage * 10:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2240 [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94653 and previous config saved to /var/cache/conftool/dbconfig/20260701-102658-cwilliams.json * 10:26 moritzm: installing nginx security updates * 10:26 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2086.codfw.wmnet with OS trixie * 10:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2179 to s4 primary [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94652 and previous config saved to /var/cache/conftool/dbconfig/20260701-102356-cwilliams.json * 10:23 cezmunsta: Starting s4 codfw failover from db2240 to db2179 - [[phab:T430127|T430127]] * 10:23 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2106.codfw.wmnet with OS trixie * 10:20 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2115.codfw.wmnet with OS trixie * 10:15 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2179 with weight 0 [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94651 and previous config saved to /var/cache/conftool/dbconfig/20260701-101531-cwilliams.json * 10:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 40 hosts with reason: Primary switchover s4 [[phab:T430127|T430127]] * 09:56 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template (take 2) - oblivian@cumin1003" * 09:56 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template (take 2) - oblivian@cumin1003 * 09:55 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template (take 2) - oblivian@cumin1003 * 09:55 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template (take 2) - oblivian@cumin1003" * 09:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:39 mszwarc@deploy1003: Synchronized private/SuggestedInvestigationsSignals/SuggestedInvestigationsSignal4n.php: Update SI signal 4n (duration: 06m 08s) * 09:21 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 09:21 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 09:14 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 09:14 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 09:02 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 09:02 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 08:54 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 08:38 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 08:38 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 08:36 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 08:21 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 08:21 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] (duration: 36m 11s) * 08:15 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 08:09 mszwarc@deploy1003: mszwarc, abi: Continuing with deployment * 08:03 mszwarc@deploy1003: mszwarc, abi: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:55 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 07:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 07:45 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] * 07:30 aqu@deploy1003: Finished deploy [analytics/refinery@410f205]: Regular analytics weekly train 2nd try [analytics/refinery@410f2050] (duration: 00m 22s) * 07:29 aqu@deploy1003: Started deploy [analytics/refinery@410f205]: Regular analytics weekly train 2nd try [analytics/refinery@410f2050] * 07:28 aqu@deploy1003: Finished deploy [analytics/refinery@410f205] (thin): Regular analytics weekly train THIN [analytics/refinery@410f2050] (duration: 01m 59s) * 07:28 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] (duration: 07m 19s) * 07:26 aqu@deploy1003: Started deploy [analytics/refinery@410f205] (thin): Regular analytics weekly train THIN [analytics/refinery@410f2050] * 07:26 aqu@deploy1003: Finished deploy [analytics/refinery@410f205]: Regular analytics weekly train [analytics/refinery@410f2050] (duration: 04m 32s) * 07:24 mszwarc@deploy1003: wmde-fisch, mszwarc: Continuing with deployment * 07:23 mszwarc@deploy1003: wmde-fisch, mszwarc: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:21 aqu@deploy1003: Started deploy [analytics/refinery@410f205]: Regular analytics weekly train [analytics/refinery@410f2050] * 07:21 aqu@deploy1003: Finished deploy [analytics/refinery@410f205] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@410f2050] (duration: 02m 01s) * 07:20 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] * 07:19 aqu@deploy1003: Started deploy [analytics/refinery@410f205] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@410f2050] * 07:13 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] (duration: 09m 13s) * 07:09 mszwarc@deploy1003: mszwarc, chlod, revi: Continuing with deployment * 07:06 mszwarc@deploy1003: mszwarc, chlod, revi: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:04 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] * 06:55 elukey: upgrade all trixie hosts to pywmflib 3.0 - [[phab:T430552|T430552]] * 06:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:43 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:43 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:42 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:42 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:41 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:41 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:35 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:35 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:34 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:34 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:31 jmm@cumin2003: DONE (PASS) - Cookbook sre.idm.logout (exit_code=0) Logging Niharika29 out of all services on: 2453 hosts * 06:30 oblivian@cumin1003: END (FAIL) - Cookbook sre.deploy.hiddenparma (exit_code=99) Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:30 oblivian@cumin1003: END (FAIL) - Cookbook sre.deploy.python-code (exit_code=99) hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:30 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:30 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:01 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2109.codfw.wmnet with OS trixie * 05:45 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on es1039.eqiad.wmnet with reason: issues * 05:41 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1027.eqiad.wmnet * 05:40 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2068.codfw.wmnet with OS trixie * 05:40 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2109.codfw.wmnet with reason: host reimage * 05:40 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1027.eqiad.wmnet,service=s2 * 05:40 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1027.eqiad.wmnet,service=s7 * 05:36 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2109.codfw.wmnet with reason: host reimage * 05:20 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2068.codfw.wmnet with reason: host reimage * 05:16 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2109.codfw.wmnet with OS trixie * 05:15 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2068.codfw.wmnet with reason: host reimage * 05:09 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2067.codfw.wmnet with OS trixie * 04:56 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2068.codfw.wmnet with OS trixie * 04:49 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2067.codfw.wmnet with reason: host reimage * 04:45 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2067.codfw.wmnet with reason: host reimage * 04:27 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2067.codfw.wmnet with OS trixie * 03:47 slyngshede@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1039.eqiad.wmnet with reason: Hardware crash * 03:21 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2107.codfw.wmnet with OS trixie * 02:59 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2107.codfw.wmnet with reason: host reimage * 02:55 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2085.codfw.wmnet with OS trixie * 02:51 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2072.codfw.wmnet with OS trixie * 02:51 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2107.codfw.wmnet with reason: host reimage * 02:35 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2085.codfw.wmnet with reason: host reimage * 02:31 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2107.codfw.wmnet with OS trixie * 02:30 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2072.codfw.wmnet with reason: host reimage * 02:26 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2085.codfw.wmnet with reason: host reimage * 02:22 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2072.codfw.wmnet with reason: host reimage * 02:09 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2085.codfw.wmnet with OS trixie * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 54s) * 02:03 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2072.codfw.wmnet with OS trixie * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es7 eqiad back to read-write - [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94649 and previous config saved to /var/cache/conftool/dbconfig/20260701-010716-ladsgroup.json * 01:05 ladsgroup@dns1004: END - running authdns-update * 01:05 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depool es1039 [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94648 and previous config saved to /var/cache/conftool/dbconfig/20260701-010551-ladsgroup.json * 01:03 ladsgroup@dns1004: START - running authdns-update * 01:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Promote es1035 to es7 primary [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94647 and previous config saved to /var/cache/conftool/dbconfig/20260701-010002-ladsgroup.json * 00:58 Amir1: Starting es7 eqiad failover from es1039 to es1035 - [[phab:T430765|T430765]] * 00:53 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es1035 with weight 0 [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94646 and previous config saved to /var/cache/conftool/dbconfig/20260701-005329-ladsgroup.json * 00:53 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 9 hosts with reason: Primary switchover es7 [[phab:T430765|T430765]] * 00:42 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es7 eqiad as read-only for maintenance - [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94645 and previous config saved to /var/cache/conftool/dbconfig/20260701-004221-ladsgroup.json * 00:20 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2102.codfw.wmnet with OS trixie * 00:15 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2103.codfw.wmnet with OS trixie * 00:05 dr0ptp4kt: DEPLOYED Refinery at {{Gerrit|4e7a2b32}} for changes: pageview allowlist {{Gerrit|1305158}} (+min.wikiquote) {{Gerrit|1305162}} (+bol.wikipedia), {{Gerrit|1305156}} (+isv.wikipedia); {{Gerrit|1305980}} (pv allowlist -api.wikimedia, sqoop +isvwiki); sqoop {{Gerrit|1295064}} (+globalimagelinks) {{Gerrit|1295069}} (+filerevision) using scap, then deployed onto HDFS (manual copyToLocal required additionally) == Other archives == See [[Server Admin Log/Archives]]. <noinclude> [[Category:SAL]] [[Category:Operations]] </noinclude> ry5ciq6nlsjq71jz53r0n2cvtdpu8xn 2450658 2450657 2026-08-22T16:37:27Z Stashbot 7414 arlolra@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply 2450658 wikitext text/x-wiki == 2026-08-22 == * 16:37 arlolra@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:37 arlolra@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:37 arlolra@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:36 arlolra@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:36 arlolra@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:36 arlolra@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:36 arlolra@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:33 arlolra@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:33 arlolra@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:33 arlolra@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:32 arlolra@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 35s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-21 == * 20:36 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in1001.wikimedia.org with reason: [[phab:T434750|T434750]] * 20:34 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in2001.wikimedia.org with reason: [[phab:T434750|T434750]] * 20:33 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out1001.wikimedia.org with reason: [[phab:T434750|T434750]] * 20:25 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out2001.wikimedia.org with reason: [[phab:T434750|T434750]] * 19:37 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host krb1002.eqiad.wmnet with OS bookworm * 19:00 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 18:59 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 18:51 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 18:51 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 18:35 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:35 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:27 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:27 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:16 bking@cumin2003: START - Cookbook sre.hosts.reimage for host krb1002.eqiad.wmnet with OS bookworm * 17:35 sukhe@dns1004: END - running authdns-update * 17:33 sukhe@dns1004: START - running authdns-update * 17:32 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns5004.wikimedia.org [reason: resolved authdns-update issues] * 17:32 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:32 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: force HEAD to {{Gerrit|be26e30ae101}} - sukhe@cumin1003" * 17:32 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: force HEAD to {{Gerrit|be26e30ae101}} - sukhe@cumin1003" * 17:28 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 17:28 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: service=authdns-update,name=dns5004.wikimedia.org [reason: resolving authdns-update issues] * 17:28 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:28 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: force HEAD to {{Gerrit|be26e30ae101}} - sukhe@cumin1003" * 17:28 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: force HEAD to {{Gerrit|be26e30ae101}} - sukhe@cumin1003" * 17:24 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 17:24 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.netbox (exit_code=97) * 17:23 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 17:16 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 17:12 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 17:10 sukhe@dns1004: END - running authdns-update * 17:08 sukhe@dns1004: START - running authdns-update * 17:08 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=dns5004.wikimedia.org [reason: resolving authdns-update issues] * 17:07 sukhe@dns1004: FAIL - running authdns-update * 17:05 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 17:05 sukhe@dns1004: START - running authdns-update * 17:01 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 16:59 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 16:56 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 16:53 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=dns5004.* [reason: trixie upgrade] * 16:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns5004.wikimedia.org * 16:52 cdobbins@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns5004.wikimedia.org * 16:44 cmooney@dns3003: END - running authdns-update * 16:41 cmooney@dns3003: START - running authdns-update * 16:41 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:41 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on eqsin<->codfw arelion - cmooney@cumin1003" * 16:37 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on eqsin<->codfw arelion - cmooney@cumin1003" * 16:33 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:11 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 16:08 sukhe@dns1004: END - running authdns-update * 16:08 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 16:06 sukhe@dns1004: START - running authdns-update * 16:04 cmooney@dns3003: END - running authdns-update * 16:02 cmooney@dns3003: START - running authdns-update * 16:00 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:00 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on eqord<->codfw arelion - cmooney@cumin1003" * 15:56 cdobbins@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host dns5004.wikimedia.org with OS trixie * 15:55 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on eqord<->codfw arelion - cmooney@cumin1003" * 15:53 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:51 cmooney@cumin1003: END (ERROR) - Cookbook sre.dns.netbox (exit_code=97) * 15:51 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:38 andrew@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudcephosd1042.eqiad.wmnet with OS bookworm * 15:18 andrew@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudcephosd1042.eqiad.wmnet with reason: host reimage * 15:17 dancy@deploy1003: Finished deploy [gerrit/gerrit@2cc11cc]: Deploying https://gerrit.wikimedia.org/r/c/operations/software/gerrit/+/1327669 ([[phab:T434726|T434726]]) (duration: 00m 14s) * 15:17 dancy@deploy1003: Started deploy [gerrit/gerrit@2cc11cc]: Deploying https://gerrit.wikimedia.org/r/c/operations/software/gerrit/+/1327669 ([[phab:T434726|T434726]]) * 15:13 andrew@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudcephosd1042.eqiad.wmnet with reason: host reimage * 15:09 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns5004.wikimedia.org with reason: host reimage * 15:05 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns5004.wikimedia.org with reason: host reimage * 14:53 andrew@cumin2003: START - Cookbook sre.hosts.reimage for host cloudcephosd1042.eqiad.wmnet with OS bookworm * 14:30 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns5004.wikimedia.org with OS trixie * 14:29 cdobbins@cumin1003: conftool action : set/pooled=no; selector: name=dns5004.* [reason: trixie upgrade] * 14:21 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:21 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove entries for cr2-eqord - cmooney@cumin1003" * 14:21 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove entries for cr2-eqord - cmooney@cumin1003" * 14:13 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 14:11 moritzm: imported openjdk 8u504-ga-1~deb12u1 for bookworm-wikimedia (backport of the latest Java 8 security fixes for bookworm) * 13:25 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "sync cr2-eqord router offline - cmooney@cumin1003" * 13:23 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "sync cr2-eqord router offline - cmooney@cumin1003" * 13:14 hashar@deploy1003: Finished deploy [integration/docroot@2d5ff9b]: opensource: add PersonalDashboard docs to MW components - [[phab:T435392|T435392]] (duration: 00m 15s) * 13:14 hashar@deploy1003: Started deploy [integration/docroot@2d5ff9b]: opensource: add PersonalDashboard docs to MW components - [[phab:T435392|T435392]] * 12:16 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:16 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: [[phab:T431682|T431682]] - filippo@cumin1003" * 12:16 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: [[phab:T431682|T431682]] - filippo@cumin1003" * 12:11 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2006.wikimedia.org with OS trixie * 12:00 kevinbazira@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 11:58 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 11:43 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2006.wikimedia.org with reason: host reimage * 11:41 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 11:38 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2006.wikimedia.org with reason: host reimage * 11:20 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2006.wikimedia.org with OS trixie * 11:11 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2005.wikimedia.org with OS trixie * 10:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2005.wikimedia.org with reason: host reimage * 10:53 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2005.wikimedia.org with reason: host reimage * 10:43 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-codfw * 10:43 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2011.codfw.wmnet * 10:43 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2011.codfw.wmnet * 10:40 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2011.codfw.wmnet * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2011.codfw.wmnet * 10:34 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2010.codfw.wmnet * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2010.codfw.wmnet * 10:33 fnegri@cumin1003: END (PASS) - Cookbook sre.wikireplicas.add-wiki (exit_code=0) for database bolwiki ([[phab:T429954|T429954]]) * 10:33 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2005.wikimedia.org with OS trixie * 10:30 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2010.codfw.wmnet * 10:25 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2010.codfw.wmnet * 10:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2009.codfw.wmnet * 10:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2009.codfw.wmnet * 10:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1006.wikimedia.org with OS trixie * 10:18 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2009.codfw.wmnet * 10:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2009.codfw.wmnet * 10:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2008.codfw.wmnet * 10:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2008.codfw.wmnet * 10:06 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2008.codfw.wmnet * 10:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1006.wikimedia.org with reason: host reimage * 10:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2008.codfw.wmnet * 10:01 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2007.codfw.wmnet * 10:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2007.codfw.wmnet * 09:57 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1006.wikimedia.org with reason: host reimage * 09:56 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2007.codfw.wmnet * 09:51 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2007.codfw.wmnet * 09:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 09:51 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 09:46 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2006.codfw.wmnet * 09:46 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1006.wikimedia.org with OS trixie * 09:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1005.wikimedia.org with OS trixie * 09:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2006.codfw.wmnet * 09:41 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2005.codfw.wmnet * 09:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2005.codfw.wmnet * 09:38 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2013.codfw.wmnet * 09:36 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2005.codfw.wmnet * 09:35 fnegri@cumin1003: START - Cookbook sre.wikireplicas.add-wiki for database bolwiki ([[phab:T429954|T429954]]) * 09:35 fnegri@cumin1003: END (PASS) - Cookbook sre.wikireplicas.add-wiki (exit_code=0) for database minwikiquote ([[phab:T429946|T429946]]) * 09:35 fnegri@cumin1003: START - Cookbook sre.wikireplicas.add-wiki for database minwikiquote ([[phab:T429946|T429946]]) * 09:32 blake@cumin1003: START - Cookbook sre.hosts.reboot-single for host rdb2013.codfw.wmnet * 09:30 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2011.codfw.wmnet * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1005.wikimedia.org with reason: host reimage * 09:26 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2005.codfw.wmnet * 09:26 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2004.codfw.wmnet * 09:26 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2004.codfw.wmnet * 09:24 blake@cumin1003: START - Cookbook sre.hosts.reboot-single for host rdb2011.codfw.wmnet * 09:22 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1005.wikimedia.org with reason: host reimage * 09:21 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2004.codfw.wmnet * 09:16 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1015.eqiad.wmnet * 09:16 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2004.codfw.wmnet * 09:16 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2003.codfw.wmnet * 09:16 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2003.codfw.wmnet * 09:13 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on cr[1-2]-eqiad,pfw1-eqiad with reason: upgrade pfw1a-eqiad and pfw1b-eqiad pair * 09:12 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 09:11 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 09:11 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 09:11 blake@cumin1003: START - Cookbook sre.hosts.reboot-single for host rdb1015.eqiad.wmnet * 09:10 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2003.codfw.wmnet * 09:09 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1013.eqiad.wmnet * 09:07 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1005.wikimedia.org with OS trixie * 09:03 blake@cumin1003: START - Cookbook sre.hosts.reboot-single for host rdb1013.eqiad.wmnet * 09:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2003.codfw.wmnet * 09:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2002.codfw.wmnet * 09:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2002.codfw.wmnet * 08:54 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2002.codfw.wmnet * 08:49 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2002.codfw.wmnet * 08:49 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2001.codfw.wmnet * 08:49 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2001.codfw.wmnet * 08:48 jmm@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts netmon2002.wikimedia.org * 08:47 jmm@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts netmon2002.wikimedia.org * 08:44 jmm@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts netmon2002.wikimedia.org * 08:44 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon2002.wikimedia.org * 08:43 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2001.codfw.wmnet * 08:36 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon2002.wikimedia.org * 08:34 jmm@dns1004: END - running authdns-update * 08:33 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2001.codfw.wmnet * 08:33 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-codfw * 08:31 jmm@dns1004: START - running authdns-update * 07:48 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327679{{!}}Block: Disable flaky API test (T435272 T389028)]], [[gerrit:1327678{{!}}API: wfDebugLog for thumberror]] (duration: 15m 34s) * 07:41 krinkle@deploy1003: krinkle: Continuing with deployment * 07:37 krinkle@deploy1003: krinkle: Backport for [[gerrit:1327679{{!}}Block: Disable flaky API test (T435272 T389028)]], [[gerrit:1327678{{!}}API: wfDebugLog for thumberror]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:33 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1327679{{!}}Block: Disable flaky API test (T435272 T389028)]], [[gerrit:1327678{{!}}API: wfDebugLog for thumberror]] * 07:25 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 07:24 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 07:18 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 07:18 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 07:15 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 07:14 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 07:14 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 07:14 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 07:13 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 07:03 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1283: Pool back * 06:42 jmm@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts netmon2002.wikimedia.org * 06:35 moritzm: powercycling netmon2002 * 06:18 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1283: Pool back * 06:17 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1283 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96213 and previous config saved to /var/cache/conftool/dbconfig/20260821-061743-marostegui.json * 04:59 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324963{{!}}Add Produnto to extension-list (T421436)]], [[gerrit:1324964{{!}}Enable Produnto on Beta (T421436)]] (duration: 34m 48s) * 04:45 tstarling@deploy1003: tstarling: Continuing with deployment * 04:44 tstarling@deploy1003: tstarling: Backport for [[gerrit:1324963{{!}}Add Produnto to extension-list (T421436)]], [[gerrit:1324964{{!}}Enable Produnto on Beta (T421436)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 04:24 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1324963{{!}}Add Produnto to extension-list (T421436)]], [[gerrit:1324964{{!}}Enable Produnto on Beta (T421436)]] * 04:21 arlolra@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 04:20 arlolra@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 04:20 arlolra@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 04:20 arlolra@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 41s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-20 == * 23:43 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327652{{!}}RunSingleJob: Add ProfilingContext::init() (T435422)]] (duration: 12m 06s) * 23:38 krinkle@deploy1003: krinkle: Continuing with deployment * 23:33 krinkle@deploy1003: krinkle: Backport for [[gerrit:1327652{{!}}RunSingleJob: Add ProfilingContext::init() (T435422)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:31 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1327652{{!}}RunSingleJob: Add ProfilingContext::init() (T435422)]] * 22:15 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1054.eqiad.wmnet with OS trixie * 22:14 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 22:14 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 21:58 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1054.eqiad.wmnet with reason: host reimage * 21:51 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1054.eqiad.wmnet with reason: host reimage * 21:36 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1054.eqiad.wmnet with OS trixie * 21:36 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:35 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327219{{!}}RunSingleJob: Define MW_ENTRY_POINT for flamegraph sample attribution (T435422)]] (duration: 08m 30s) * 21:31 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:31 krinkle@deploy1003: krinkle: Continuing with deployment * 21:31 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1054 * 21:31 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1054 * 21:29 krinkle@deploy1003: krinkle: Backport for [[gerrit:1327219{{!}}RunSingleJob: Define MW_ENTRY_POINT for flamegraph sample attribution (T435422)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:27 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1327219{{!}}RunSingleJob: Define MW_ENTRY_POINT for flamegraph sample attribution (T435422)]] * 21:17 maryum: Deployed security fix for [[phab:T433020|T433020]] * 20:59 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324752{{!}}InitialiseSettings: Enable 2FA banners on remaining private wikis (T428103)]], [[gerrit:1325920{{!}}Remove sending email to legal team about rejected requests (T374053)]] (duration: 07m 18s) * 20:54 reedy@deploy1003: neriah, reedy: Continuing with deployment * 20:54 reedy@deploy1003: neriah, reedy: Backport for [[gerrit:1324752{{!}}InitialiseSettings: Enable 2FA banners on remaining private wikis (T428103)]], [[gerrit:1325920{{!}}Remove sending email to legal team about rejected requests (T374053)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:51 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324752{{!}}InitialiseSettings: Enable 2FA banners on remaining private wikis (T428103)]], [[gerrit:1325920{{!}}Remove sending email to legal team about rejected requests (T374053)]] * 20:24 reedy@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.15,1.47.0-wmf.16,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/med * 20:23 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324752{{!}}InitialiseSettings: Enable 2FA banners on remaining private wikis (T428103)]], [[gerrit:1325920{{!}}Remove sending email to legal team about rejected requests (T374053)]] * 20:14 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 20:10 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 20:09 cdanis@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "bug fixes & UX fixes - cdanis@cumin1003" * 20:09 cdanis@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: bug fixes & UX fixes - cdanis@cumin1003 * 20:08 cdanis@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: bug fixes & UX fixes - cdanis@cumin1003 * 20:08 cdanis@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "bug fixes & UX fixes - cdanis@cumin1003" * 19:24 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327598{{!}}Make \Omicron non upright (like \Chi) (T434428)]], [[gerrit:1327596{{!}}Render overline of \bar with stretchy=false (T435456)]] (duration: 18m 54s) * 19:20 krinkle@deploy1003: krinkle: Continuing with deployment * 19:07 krinkle@deploy1003: krinkle: Backport for [[gerrit:1327598{{!}}Make \Omicron non upright (like \Chi) (T434428)]], [[gerrit:1327596{{!}}Render overline of \bar with stretchy=false (T435456)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:05 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1327598{{!}}Make \Omicron non upright (like \Chi) (T434428)]], [[gerrit:1327596{{!}}Render overline of \bar with stretchy=false (T435456)]] * 18:51 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327614{{!}}Avoid casting fpxmax to string (T318419)]] (duration: 07m 28s) * 18:50 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:46 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:46 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 18:45 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327614{{!}}Avoid casting fpxmax to string (T318419)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:43 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327614{{!}}Avoid casting fpxmax to string (T318419)]] * 18:37 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:37 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:37 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:36 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:07 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 18:05 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 18:01 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 18:01 sukhe@dns1004: END - running authdns-update * 17:59 sukhe@dns1004: START - running authdns-update * 17:58 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 17:57 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=dns6002.* [reason: depooling for trixie upgrade] * 17:56 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns6002.wikimedia.org * 17:56 cdobbins@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns6002.wikimedia.org * 17:51 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 17:51 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 17:34 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host stat1011.eqiad.wmnet with OS bookworm * 17:31 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns6002.wikimedia.org with OS trixie * 17:30 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 17:30 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 17:30 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 17:30 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 17:29 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 17:29 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 17:29 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 17:29 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:29 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:27 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:24 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327590{{!}}Make sure fpsmax is an int value (T318419)]] (duration: 08m 37s) * 17:20 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 17:17 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327590{{!}}Make sure fpsmax is an int value (T318419)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:16 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 17:15 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327590{{!}}Make sure fpsmax is an int value (T318419)]] * 16:51 swfrench-wmf: disable-puppet on A:cp for ATS Lua change - [[phab:T427666|T427666]] * 16:51 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db2901.codfw.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 16:43 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 16:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on stat1011.eqiad.wmnet with reason: host reimage * 16:39 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns6002.wikimedia.org with reason: host reimage * 16:36 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on stat1011.eqiad.wmnet with reason: host reimage * 16:36 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db2901.codfw.wmnet * 16:34 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns6002.wikimedia.org with reason: host reimage * 16:15 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns6002.wikimedia.org with OS trixie * 16:14 cdobbins@cumin1003: conftool action : set/pooled=no; selector: name=dns6002.* [reason: depooling for trixie upgrade] * 16:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host stat1011 * 16:10 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host stat1011 * 16:09 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host stat1011 * 16:09 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) stat1011.eqiad.wmnet 14.36.64.10.in-addr.arpa 4.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 bking@cumin2003: START - Cookbook sre.dns.wipe-cache stat1011.eqiad.wmnet 14.36.64.10.in-addr.arpa 4.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:09 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host stat1011 - bking@cumin2003" * 16:09 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host stat1011 - bking@cumin2003" * 16:05 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host stat1011 * 16:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host stat1011.eqiad.wmnet with OS bookworm * 16:00 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 15:59 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 15:56 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 15:55 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 15:37 jayme@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:35 jayme@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 15:35 jayme@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:32 jayme@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:32 jayme@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:30 jayme@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 15:30 jayme@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:28 jayme@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:28 jayme@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 15:28 fceratto@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host db1903.eqiad.wmnet * 15:28 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1903.eqiad.wmnet with OS trixie * 15:26 jayme@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 15:26 jayme@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 15:24 jayme@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 15:24 jayme@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 15:21 jayme@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 15:21 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 15:19 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 15:19 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 15:17 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 15:14 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1903.eqiad.wmnet with reason: host reimage * 15:07 fceratto@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1903.eqiad.wmnet with reason: host reimage * 14:54 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db1903.eqiad.wmnet with OS trixie * 14:53 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1903.eqiad.wmnet - fceratto@cumin1003" * 14:53 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1903.eqiad.wmnet - fceratto@cumin1003" * 14:53 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1903.eqiad.wmnet on all recursors * 14:53 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1903.eqiad.wmnet on all recursors * 14:53 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:53 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1903.eqiad.wmnet - fceratto@cumin1003" * 14:53 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1903.eqiad.wmnet - fceratto@cumin1003" * 14:49 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 14:49 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1903.eqiad.wmnet * 14:33 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=cp1100.* * 14:27 topranks: reconfigure eqiad<->codfw bgp settings * 14:22 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:22 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update entries used on new transport backup eqiad codfw - cmooney@cumin1003" * 14:19 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update entries used on new transport backup eqiad codfw - cmooney@cumin1003" * 14:14 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 14:14 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 14:13 moritzm: installing util-linux security updates * 14:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-staging-worker * 14:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2003.codfw.wmnet * 14:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2003.codfw.wmnet * 14:08 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2003.codfw.wmnet * 14:06 moritzm: installing libheif security updates * 13:58 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2003.codfw.wmnet * 13:58 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2002.codfw.wmnet * 13:58 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2002.codfw.wmnet * 13:56 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:56 fnegri@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for clouddb1025.eqiad.wmnet * 13:56 fnegri@cumin1003: START - Cookbook sre.hosts.remove-downtime for clouddb1025.eqiad.wmnet * 13:56 Lucas_WMDE: UTC afternoon backport+config window done * 13:53 moritzm: installing apr-util security updates * 13:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2002.codfw.wmnet * 13:50 fnegri@cumin1003: conftool action : set/weight=100; selector: name=clouddb1025.eqiad.wmnet * 13:49 fnegri@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1025.eqiad.wmnet * 13:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2002.codfw.wmnet * 13:41 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2001.codfw.wmnet * 13:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2001.codfw.wmnet * 13:41 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'. * 13:38 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'. * 13:38 fnegri@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on clouddb1025.eqiad.wmnet with reason: Removing s6 from clouddb1025 * 13:34 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2001.codfw.wmnet * 13:31 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'. * 13:29 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'. * 13:28 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet * 13:26 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host stat1009.eqiad.wmnet with OS bookworm * 13:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2001.codfw.wmnet * 13:24 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-staging-worker * 13:23 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1002.eqiad.wmnet * 13:21 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host stat1010.eqiad.wmnet with OS bookworm * 13:20 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1002.eqiad.wmnet * 13:20 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1001.eqiad.wmnet * 13:17 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1001.eqiad.wmnet * 13:16 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2001.codfw.wmnet * 13:13 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327511{{!}}UIC: Fix page:page instead of page:other in instrumentation]] (duration: 07m 00s) * 13:13 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2001.codfw.wmnet * 13:12 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2002.codfw.wmnet * 13:09 mszwarc@deploy1003: mszwarc: Continuing with deployment * 13:08 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1327511{{!}}UIC: Fix page:page instead of page:other in instrumentation]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2002.codfw.wmnet * 13:07 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2002.codfw.wmnet * 13:06 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1327511{{!}}UIC: Fix page:page instead of page:other in instrumentation]] * 13:04 jmm@dns1004: END - running authdns-update * 13:03 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2002.codfw.wmnet * 13:03 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2001.codfw.wmnet * 13:02 jmm@dns1004: START - running authdns-update * 13:01 cmooney@dns3003: END - running authdns-update * 13:00 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2001.codfw.wmnet * 12:59 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2001.codfw.wmnet * 12:59 cmooney@dns3003: START - running authdns-update * 12:57 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2001.codfw.wmnet * 12:56 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2002.codfw.wmnet * 12:55 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:55 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on drmrs<->eqiad GTT vpls - cmooney@cumin1003" * 12:54 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on drmrs<->eqiad GTT vpls - cmooney@cumin1003" * 12:54 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2002.codfw.wmnet * 12:54 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2003.codfw.wmnet * 12:50 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2003.codfw.wmnet * 12:49 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 12:48 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2001.codfw.wmnet * 12:46 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2001.codfw.wmnet * 12:46 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2002.codfw.wmnet * 12:43 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2002.codfw.wmnet * 12:43 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2003.codfw.wmnet * 12:42 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=1) for new host db1902.eqiad.wmnet * 12:42 fceratto@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host db1902.eqiad.wmnet with OS trixie * 12:41 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2003.codfw.wmnet * 12:40 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1003.eqiad.wmnet * 12:38 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1003.eqiad.wmnet * 12:37 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1002.eqiad.wmnet * 12:35 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1002.eqiad.wmnet * 12:35 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1001.eqiad.wmnet * 12:34 cmooney@dns3003: END - running authdns-update * 12:33 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1001.eqiad.wmnet * 12:32 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on stat1009.eqiad.wmnet with reason: host reimage * 12:31 cmooney@dns3003: START - running authdns-update * 12:31 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:31 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on drmrs<->eqiad cct - cmooney@cumin1003" * 12:28 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on drmrs<->eqiad cct - cmooney@cumin1003" * 12:28 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1902.eqiad.wmnet with reason: host reimage * 12:25 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on stat1009.eqiad.wmnet with reason: host reimage * 12:25 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 12:24 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on stat1010.eqiad.wmnet with reason: host reimage * 12:22 fceratto@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1902.eqiad.wmnet with reason: host reimage * 12:21 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on stat1010.eqiad.wmnet with reason: host reimage * 12:14 elukey: move the Docker Registry's /v2/dev/.* prefix to its dedicated S3 backend - [[phab:T432829|T432829]] * 12:12 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db1902.eqiad.wmnet with OS trixie * 12:09 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1902.eqiad.wmnet - fceratto@cumin1003" * 12:09 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1902.eqiad.wmnet - fceratto@cumin1003" * 12:09 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1902.eqiad.wmnet on all recursors * 12:09 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1902.eqiad.wmnet on all recursors * 12:09 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:08 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1902.eqiad.wmnet - fceratto@cumin1003" * 12:08 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1902.eqiad.wmnet - fceratto@cumin1003" * 12:08 tgr_: [[phab:T413390|T413390]] running CentralAuth:FixRenamedUserGlobalEditCount --wiki=metawiki --since=20250901000000 --until=20260301000000 --fix * 12:04 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1009.eqiad.wmnet with OS bookworm * 12:04 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1010.eqiad.wmnet with OS bookworm * 12:01 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 12:01 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1902.eqiad.wmnet * 12:00 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host stat1010.eqiad.wmnet with OS bookworm * 11:57 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1902.eqiad.wmnet * 11:57 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:57 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1902.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 11:57 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1902.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 11:51 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'. * 11:49 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'. * 11:48 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'. * 11:46 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'. * 11:39 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 11:37 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327518{{!}}Enable thumb.wikimedia.org on cswiki and fawiki (T427465)]] (duration: 10m 40s) * 11:35 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1902.eqiad.wmnet * 11:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1010.eqiad.wmnet with OS bookworm * 11:33 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 11:30 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327518{{!}}Enable thumb.wikimedia.org on cswiki and fawiki (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:26 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327518{{!}}Enable thumb.wikimedia.org on cswiki and fawiki (T427465)]] * 11:24 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host stat1010.eqiad.wmnet with OS bookworm * 11:07 fceratto@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host db1901.eqiad.wmnet * 11:07 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1901.eqiad.wmnet with OS trixie * 10:53 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1901.eqiad.wmnet with reason: host reimage * 10:47 tappof: bump space for prometheus k8s-aux in codfw * 10:47 tappof: bump space for prometheus k8s-dse in eqiad * 10:47 fceratto@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1901.eqiad.wmnet with reason: host reimage * 10:35 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db1901.eqiad.wmnet with OS trixie * 10:32 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:32 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:32 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1901.eqiad.wmnet on all recursors * 10:32 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1901.eqiad.wmnet on all recursors * 10:31 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:31 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:31 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:27 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:27 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1901.eqiad.wmnet * 10:23 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1010.eqiad.wmnet with OS bookworm * 10:20 fceratto@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts db1901.eqiad.wmnet * 10:20 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 10:18 blake@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 10:17 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:16 blake@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 10:13 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1901.eqiad.wmnet * 09:23 jelto@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'. * 09:22 jelto@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'. * 09:22 jelto: update cert-manager to 1.19.6 on wikikube staging-eqiad - [[phab:T427402|T427402]] * 09:20 moritzm: imported squid 7.6-2.1for trixie-wikimedia/main [[phab:T427282|T427282]] * 09:08 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 09:08 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 09:08 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 09:07 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 09:04 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 09:04 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:23 slyngshede@dns1004: END - running authdns-update * 08:21 slyngshede@dns1004: START - running authdns-update * 08:18 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.16 refs [[phab:T430835|T430835]] * 06:27 aokoth@dns1004: END - running authdns-update * 06:25 aokoth@dns1004: START - running authdns-update * 06:22 brennen@deploy1003: Finished deploy [phabricator/deployment@6b9b6ff]: deploy phab1005 for [[phab:T435087|T435087]] (duration: 00m 39s) * 06:21 brennen@deploy1003: Started deploy [phabricator/deployment@6b9b6ff]: deploy phab1005 for [[phab:T435087|T435087]] * 06:20 brennen@deploy1003: Finished deploy [phabricator/deployment@6b9b6ff]: deploy phab1004 for to pick up config values for [[phab:T435087|T435087]] (duration: 01m 46s) * 06:18 brennen@deploy1003: Started deploy [phabricator/deployment@6b9b6ff]: deploy phab1004 for to pick up config values for [[phab:T435087|T435087]] * 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 49s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-19 == * 23:19 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327197{{!}}Enable thumb.wikimedia.org on mediawiki.org (T427465)]] (duration: 10m 50s) * 23:18 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:16 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 23:15 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 23:10 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327197{{!}}Enable thumb.wikimedia.org on mediawiki.org (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:10 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:09 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 23:09 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:09 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 23:08 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327197{{!}}Enable thumb.wikimedia.org on mediawiki.org (T427465)]] * 22:58 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327201{{!}}Enable ReadingLists for all logged in users on test wiki (T435258)]] (duration: 11m 20s) * 22:50 jdlrobson@deploy1003: jdlrobson: Continuing with deployment * 22:49 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1327201{{!}}Enable ReadingLists for all logged in users on test wiki (T435258)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:46 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1327201{{!}}Enable ReadingLists for all logged in users on test wiki (T435258)]] * 22:42 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327162{{!}}Article: Split subjectpageheader by model and disable for wikitext]], [[gerrit:1327169{{!}}Make uppercase greek letters normal (non-italic) font (T434686 T434428)]], [[gerrit:1327176{{!}}Skin: Avoid DB lookup for pagecategorieslink message (T347123)]] (duration: 37m 52s) * 22:29 krinkle@deploy1003: krinkle: Continuing with deployment * 22:25 krinkle@deploy1003: krinkle: Backport for [[gerrit:1327162{{!}}Article: Split subjectpageheader by model and disable for wikitext]], [[gerrit:1327169{{!}}Make uppercase greek letters normal (non-italic) font (T434686 T434428)]], [[gerrit:1327176{{!}}Skin: Avoid DB lookup for pagecategorieslink message (T347123)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:04 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1327162{{!}}Article: Split subjectpageheader by model and disable for wikitext]], [[gerrit:1327169{{!}}Make uppercase greek letters normal (non-italic) font (T434686 T434428)]], [[gerrit:1327176{{!}}Skin: Avoid DB lookup for pagecategorieslink message (T347123)]] * 22:04 krinkle@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: awaiting CI (duration: 03m 06s) * 22:01 krinkle@deploy1003: Locking from deployment [ALL REPOSITORIES]: awaiting CI * 22:00 krinkle@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: awaiting CI (duration: 00m 01s) * 22:00 krinkle@deploy1003: Locking from deployment [ALL REPOSITORIES]: awaiting CI * 21:34 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2001.codfw.wmnet * 21:28 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2001.codfw.wmnet * 21:22 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327178{{!}}AccountRecovery: Notify the email address of the on file of the request (T425799)]] (duration: 47m 02s) * 21:13 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm * 21:09 catrope@deploy1003: catrope: Continuing with deployment * 20:55 catrope@deploy1003: catrope: Backport for [[gerrit:1327178{{!}}AccountRecovery: Notify the email address of the on file of the request (T425799)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:35 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1327178{{!}}AccountRecovery: Notify the email address of the on file of the request (T425799)]] * 20:31 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327128{{!}}Parsoid DataAccess: convert Parsoid fragment markers to/from strip tags (T432547)]] (duration: 07m 30s) * 20:27 catrope@deploy1003: catrope, arlolra: Continuing with deployment * 20:26 catrope@deploy1003: catrope, arlolra: Backport for [[gerrit:1327128{{!}}Parsoid DataAccess: convert Parsoid fragment markers to/from strip tags (T432547)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:24 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1327128{{!}}Parsoid DataAccess: convert Parsoid fragment markers to/from strip tags (T432547)]] * 20:23 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage * 20:17 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage * 20:15 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325878{{!}}[arwiki] Enable restricted user page editing and grant edit permissions (T434878)]] (duration: 08m 46s) * 20:11 catrope@deploy1003: catrope, gergesshamon: Continuing with deployment * 20:08 catrope@deploy1003: catrope, gergesshamon: Backport for [[gerrit:1325878{{!}}[arwiki] Enable restricted user page editing and grant edit permissions (T434878)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:06 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1325878{{!}}[arwiki] Enable restricted user page editing and grant edit permissions (T434878)]] * 19:59 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm * 19:56 eevans@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cassandra-dev2001.codfw.wmnet with OS bookworm * 19:56 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm * 19:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2207.codfw.wmnet with reason: Maintenance * 18:47 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319920{{!}}Allow setting a separate thumbUrl in production (T427465)]], [[gerrit:1327167{{!}}Fix wmgThumbUrl config (T427465)]] (duration: 18m 53s) * 18:43 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 18:30 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1319920{{!}}Allow setting a separate thumbUrl in production (T427465)]], [[gerrit:1327167{{!}}Fix wmgThumbUrl config (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:28 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1319920{{!}}Allow setting a separate thumbUrl in production (T427465)]], [[gerrit:1327167{{!}}Fix wmgThumbUrl config (T427465)]] * 18:26 sukhe@dns1004: END - running authdns-update * 18:24 sukhe@dns1004: START - running authdns-update * 18:09 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1319920{{!}}Allow setting a separate thumbUrl in production (T427465)]] * 18:03 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-eqiad and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 17:56 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-codfw and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 17:50 cmooney@dns3003: END - running authdns-update * 17:42 dancy@deploy1003: Installation of scap version "4.283.0" completed for 3 hosts * 17:41 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling reboot on A:durum and A:durum * 17:40 cmooney@dns3003: START - running authdns-update * 17:40 dancy@deploy1003: Installing scap version "4.283.0" for 3 host(s) * 17:38 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:37 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on GTT VPLS - cmooney@cumin1003" * 17:37 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --verbose --scoreLessThan=0.7 --exceptDatasetChecksums=[[phab:T434319|T434319]]-enwiki-models.txt # [[phab:T434319|T434319]] * 17:32 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on GTT VPLS - cmooney@cumin1003" * 17:29 sbassett: Deployed security fix for [[phab:T435210|T435210]] (wmf.16) * 17:26 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 17:22 sbassett: Deployed security fix for [[phab:T435210|T435210]] (wmf.15) * 17:00 sukhe@dns1004: END - running authdns-update * 16:58 sukhe@dns1004: START - running authdns-update * 16:53 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica-esams and A:liberica * 16:41 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica-esams and A:liberica * 16:41 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-eqiad and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 16:41 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-codfw and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 16:41 cjd91: sudo -i cookbook sre.cdn.roll-upgrade-ats --query 'A:cp-codfw' --task-id [[phab:T434478|T434478]] --reason '9.2.15 upgrade' * 16:41 cjd91: sudo -i cookbook sre.cdn.roll-upgrade-ats --query 'A:cp-eqiad' --task-id [[phab:T434478|T434478]] --reason '9.2.15 upgrade' * 16:40 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and A:durum * 16:28 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326881{{!}}Echo: Start using virtual domains (T380385)]] (duration: 13m 12s) * 16:23 urbanecm@deploy1003: urbanecm: Continuing with deployment * 16:21 urandom: Completed sessionstore Cassandra/JVM upgrade — [[phab:T435154|T435154]] * 16:21 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching sessionstore[2005-2006].codfw.wmnet,sessionstore[1005-1006].eqiad.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 16:19 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1326881{{!}}Echo: Start using virtual domains (T380385)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:15 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:15 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->codfw - cmooney@cumin1003" * 16:14 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1326881{{!}}Echo: Start using virtual domains (T380385)]] * 16:14 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327123{{!}}Revert^2 "Migrate database access to virtual domains" (T435305)]], [[gerrit:1327124{{!}}Pass the mapped domain of virtual-echo-shared to the push NameTableStores (T435305)]] (duration: 07m 42s) * 16:13 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching sessionstore[2005-2006].codfw.wmnet,sessionstore[1005-1006].eqiad.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 16:11 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->codfw - cmooney@cumin1003" * 16:10 urbanecm@deploy1003: urbanecm: Continuing with deployment * 16:08 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1327123{{!}}Revert^2 "Migrate database access to virtual domains" (T435305)]], [[gerrit:1327124{{!}}Pass the mapped domain of virtual-echo-shared to the push NameTableStores (T435305)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:06 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 16:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 16:06 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1327123{{!}}Revert^2 "Migrate database access to virtual domains" (T435305)]], [[gerrit:1327124{{!}}Pass the mapped domain of virtual-echo-shared to the push NameTableStores (T435305)]] * 16:06 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:03 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching sessionstore1004.eqiad.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 16:01 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching sessionstore1004.eqiad.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 15:56 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching sessionstore2004.codfw.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 15:54 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching sessionstore2004.codfw.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 15:53 urandom: beginning sessionstore Cassandra/JVM upgrade — [[phab:T435154|T435154]] * 15:52 urandom: beginning sessionstore Cassandra/JVM upgrade — [[phab:T432944|T432944]] * 15:51 cmooney@dns3003: END - running authdns-update * 15:49 cmooney@dns3003: START - running authdns-update * 15:48 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:48 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->codfw - cmooney@cumin1003" * 15:45 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->codfw - cmooney@cumin1003" * 15:44 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 15:44 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:43 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 15:42 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:42 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:38 cmooney@dns3003: END - running authdns-update * 15:36 cmooney@dns3003: START - running authdns-update * 15:36 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:36 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->eqsin - cmooney@cumin1003" * 15:34 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327138{{!}}Enable redis lock manager everywhere (T366938)]] (duration: 08m 36s) * 15:33 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->eqsin - cmooney@cumin1003" * 15:30 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:29 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 15:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1147.eqiad.wmnet with OS bookworm * 15:28 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327138{{!}}Enable redis lock manager everywhere (T366938)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:25 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327138{{!}}Enable redis lock manager everywhere (T366938)]] * 15:24 jmm@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host krb1002.eqiad.wmnet * 15:19 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2207.codfw.wmnet with reason: Host crashed * 15:17 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327098{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]], [[gerrit:1327101{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]] (duration: 07m 13s) * 15:12 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 15:12 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1327098{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]], [[gerrit:1327101{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:10 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1327098{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]], [[gerrit:1327101{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]] * 15:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2010.codfw.wmnet with OS trixie * 15:05 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 15:05 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1147.eqiad.wmnet with reason: host reimage * 14:59 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 14:59 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:58 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1147.eqiad.wmnet with reason: host reimage * 14:55 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:55 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:49 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 14:48 cmooney@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host durum1001.eqiad.wmnet * 14:46 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:46 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Delete db2902 ipv6 addr - fceratto@cumin1003" * 14:46 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Delete db2902 ipv6 addr - fceratto@cumin1003" * 14:43 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1147.eqiad.wmnet with OS bookworm * 14:42 tgr@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327094{{!}}SpecialMWOAuthListConsumers: Handle newFromMWUser returning null in addNavigationSubtitle (T435167)]] (duration: 19m 25s) * 14:42 cmooney@cumin1003: START - Cookbook sre.hosts.reboot-single for host durum1001.eqiad.wmnet * 14:42 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 14:41 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 14:38 tgr@deploy1003: tgr: Continuing with deployment * 14:36 tgr@deploy1003: tgr: Backport for [[gerrit:1327094{{!}}SpecialMWOAuthListConsumers: Handle newFromMWUser returning null in addNavigationSubtitle (T435167)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:28 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 14:28 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:28 cmooney@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host durum3005.esams.wmnet * 14:25 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 14:23 cmooney@cumin1003: START - Cookbook sre.hosts.reboot-single for host durum3005.esams.wmnet * 14:23 tgr@deploy1003: Started scap sync-world: Backport for [[gerrit:1327094{{!}}SpecialMWOAuthListConsumers: Handle newFromMWUser returning null in addNavigationSubtitle (T435167)]] * 14:18 gengh@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:17 topranks: disable puppet on hosts running BIRD BGP to test merge of patch to systemd healtchcheck service * 14:17 gengh@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:17 gengh@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:16 gengh@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:16 elukey: upgrade spicerack on cumin1003 and cumin2003 * 14:16 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:15 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314962{{!}}static: add new dir bimi/ for BIMI SVG and PEM file (T311685)]] (duration: 10m 00s) * 14:15 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:11 kharlan@deploy1003: kharlan, sukhe: Continuing with deployment * 14:11 gengh@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:09 gengh@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:09 gengh@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:08 kharlan@deploy1003: kharlan, sukhe: Backport for [[gerrit:1314962{{!}}static: add new dir bimi/ for BIMI SVG and PEM file (T311685)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:06 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 14:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:06 gengh@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:05 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1314962{{!}}static: add new dir bimi/ for BIMI SVG and PEM file (T311685)]] * 14:05 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host krb1002.eqiad.wmnet * 14:05 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:04 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:04 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326342{{!}}Revert^2 "wmf-config/ProductionServices: set URL for urldownloader to service record"]] (duration: 07m 40s) * 13:59 kharlan@deploy1003: kharlan, sukhe: Continuing with deployment * 13:58 kharlan@deploy1003: kharlan, sukhe: Backport for [[gerrit:1326342{{!}}Revert^2 "wmf-config/ProductionServices: set URL for urldownloader to service record"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:57 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host stat1008.eqiad.wmnet with OS bookworm * 13:56 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1326342{{!}}Revert^2 "wmf-config/ProductionServices: set URL for urldownloader to service record"]] * 13:56 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 13:54 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325532{{!}}srwiki: Allow bureaucrats to add and remove event-organizer group (T434748)]] (duration: 14m 56s) * 13:54 swfrench@dns1004: END - running authdns-update * 13:52 swfrench@dns1004: START - running authdns-update * 13:51 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:51 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 13:48 kharlan@deploy1003: kharlan, danielyepezgarces: Continuing with deployment * 13:45 swfrench@cumin2003: conftool action : set/pooled=yes; selector: name=wikikube-worker2330.codfw.wmnet * 13:44 swfrench-wmf: finished etcd-main codfw -> eqiad switchover - [[phab:T435103|T435103]] * 13:44 kharlan@deploy1003: kharlan, danielyepezgarces: Backport for [[gerrit:1325532{{!}}srwiki: Allow bureaucrats to add and remove event-organizer group (T434748)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:44 swfrench@cumin2003: conftool action : set/pooled=no; selector: name=wikikube-worker2330.codfw.wmnet * 13:41 swfrench@dns1004: END - running authdns-update * 13:39 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1325532{{!}}srwiki: Allow bureaucrats to add and remove event-organizer group (T434748)]] * 13:39 swfrench@dns1004: START - running authdns-update * 13:37 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327111{{!}}Special:AbuseReview: Add "no further action needed" review action (T435020)]], [[gerrit:1327110{{!}}AbuseReview: Take the review verdict as a REST path parameter (T435020)]] (duration: 31m 43s) * 13:31 swfrench-wmf: starting etcd-main codfw -> eqiad switchover - [[phab:T435103|T435103]] * 13:28 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:28 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 13:25 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=97) rolling reboot on A:durum and A:durum * 13:24 kharlan@deploy1003: kharlan: Continuing with deployment * 13:24 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1147 * 13:24 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1147 * 13:23 kharlan@deploy1003: kharlan: Backport for [[gerrit:1327111{{!}}Special:AbuseReview: Add "no further action needed" review action (T435020)]], [[gerrit:1327110{{!}}AbuseReview: Take the review verdict as a REST path parameter (T435020)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:16 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:16 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 13:12 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:12 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 13:10 cdobbins@cumin1003: END (ERROR) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=97) Rolling upgrade of ATS on A:cp-codfw and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 13:10 cdobbins@cumin1003: END (ERROR) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=97) Rolling upgrade of ATS on A:cp-eqiad and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 13:06 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1327111{{!}}Special:AbuseReview: Add "no further action needed" review action (T435020)]], [[gerrit:1327110{{!}}AbuseReview: Take the review verdict as a REST path parameter (T435020)]] * 13:06 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-eqiad and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 13:05 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-codfw and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 13:05 cjd91: sudo -i cookbook sre.cdn.roll-upgrade-ats --query 'A:cp-eqiad' --task-id [[phab:T434478|T434478]] --reason '9.2.15 upgrade' * 13:03 swfrench@dns1004: END - running authdns-update * 13:01 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:01 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 13:01 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host ncmonitor1001.eqiad.wmnet * 13:00 swfrench@dns1004: START - running authdns-update * 12:59 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 12:59 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 12:59 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 12:59 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 12:57 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and A:durum * 12:53 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on db2902.codfw.wmnet with reason: Cloning * 12:48 cmooney@dns3003: END - running authdns-update * 12:48 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327096{{!}}Switch to redis lock manager on s4 and s8 (T366938)]] (duration: 09m 11s) * 12:46 cmooney@dns3003: START - running authdns-update * 12:45 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:45 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->eqord cct - cmooney@cumin1003" * 12:45 elukey: move the /v2/releng.* prefix on the Docker Registry to its new s3 backend - [[phab:T432829|T432829]] * 12:43 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 12:42 jelto@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 12:42 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->eqord cct - cmooney@cumin1003" * 12:42 jelto@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 12:41 jelto: update cert-manager to 1.19.6 on wikikube staging-codfw - [[phab:T427402|T427402]] * 12:40 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327096{{!}}Switch to redis lock manager on s4 and s8 (T366938)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:38 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327096{{!}}Switch to redis lock manager on s4 and s8 (T366938)]] * 12:38 blake@deploy1003: Finished scap sync-world: non-build deploy for [[phab:T417800|T417800]] (duration: 03m 52s) * 12:36 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 12:35 blake@deploy1003: Started scap sync-world: non-build deploy for [[phab:T417800|T417800]] * 12:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host krb2002.codfw.wmnet * 11:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host krb2002.codfw.wmnet * 11:49 moritzm: installing kerberos security updates * 11:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on stat1008.eqiad.wmnet with reason: host reimage * 11:44 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on stat1008.eqiad.wmnet with reason: host reimage * 11:31 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327084{{!}}Revert "Migrate database access to virtual domains" (T435305)]] (duration: 11m 02s) * 11:29 gkyziridis@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:29 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:29 kart_: Updated MinT to 2026-06-04-131507-production ([[phab:T321316|T321316]]) * 11:28 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/machinetranslation: apply * 11:28 gkyziridis@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:26 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 11:24 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:24 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:23 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:23 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:23 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/machinetranslation: apply * 11:22 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327084{{!}}Revert "Migrate database access to virtual domains" (T435305)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:21 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/machinetranslation: apply * 11:21 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:21 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:20 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327084{{!}}Revert "Migrate database access to virtual domains" (T435305)]] * 11:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1008.eqiad.wmnet with OS bookworm * 11:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet * 11:17 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/machinetranslation: apply * 11:13 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/machinetranslation: apply * 11:12 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:12 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:10 kartik@deploy1003: helmfile [staging] START helmfile.d/services/machinetranslation: apply * 11:08 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-cron: apply * 11:08 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/mw-cron: apply * 11:08 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply * 11:08 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply * 11:07 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet * 11:07 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host stat1008.eqiad.wmnet with OS bookworm * 11:06 moritzm: upgrading the new trixie URL downloaders to Squid 7.6 [[phab:T427282|T427282]] * 11:01 gkyziridis@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin2002.codfw.wmnet * 10:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin2002.codfw.wmnet * 10:45 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1169.eqiad.wmnet with OS bookworm * 10:42 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1185.eqiad.wmnet with OS bookworm * 10:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1169.eqiad.wmnet with reason: host reimage * 10:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1185.eqiad.wmnet with reason: host reimage * 10:14 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1169.eqiad.wmnet with reason: host reimage * 10:14 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1185.eqiad.wmnet with reason: host reimage * 10:11 jmm@cumin2003: END (PASS) - Cookbook sre.netbox.restart-reboot (exit_code=0) rolling reboot on A:netbox * 10:06 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1008.eqiad.wmnet with OS bookworm * 10:04 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 10:04 mpostoronca@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321579{{!}}Register the mediawiki.wikimedia_antiabuse.content_policy_score stream (T432848)]] (duration: 08m 53s) * 10:03 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 10:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 10:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 10:00 mpostoronca@deploy1003: mpostoronca: Continuing with deployment * 09:59 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1185.eqiad.wmnet with OS bookworm * 09:59 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1169.eqiad.wmnet with OS bookworm * 09:59 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.convert-disks (exit_code=0) for host ms-be1065 * 09:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1065.eqiad.wmnet with OS trixie * 09:58 mpostoronca@deploy1003: mpostoronca: Backport for [[gerrit:1321579{{!}}Register the mediawiki.wikimedia_antiabuse.content_policy_score stream (T432848)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:55 jmm@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netbox.discovery.wmnet. on all recursors * 09:55 mpostoronca@deploy1003: Started scap sync-world: Backport for [[gerrit:1321579{{!}}Register the mediawiki.wikimedia_antiabuse.content_policy_score stream (T432848)]] * 09:55 jmm@cumin2003: START - Cookbook sre.dns.wipe-cache netbox.discovery.wmnet. on all recursors * 09:52 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw2001.wikimedia.org with OS trixie * 09:51 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 09:51 jmm@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netbox.discovery.wmnet. on all recursors * 09:51 jmm@cumin2003: START - Cookbook sre.dns.wipe-cache netbox.discovery.wmnet. on all recursors * 09:51 jmm@cumin2003: START - Cookbook sre.netbox.restart-reboot rolling reboot on A:netbox * 09:46 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 09:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1056.eqiad.wmnet with OS trixie * 09:44 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 09:39 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 09:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 09:36 topranks: make HE transport circuits from magru live * 09:36 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 09:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb1003.eqiad.wmnet * 09:33 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage * 09:31 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb1003.eqiad.wmnet * 09:30 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 09:28 arnaudb@dns1006: END - running authdns-update * 09:27 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage * 09:27 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb2003.codfw.wmnet * 09:26 arnaudb@dns1006: START - running authdns-update * 09:24 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1056.eqiad.wmnet with reason: host reimage * 09:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb2003.codfw.wmnet * 09:20 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.convert-disks (exit_code=0) for host ms-be1068 * 09:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1068.eqiad.wmnet with OS trixie * 09:20 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 09:19 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "cloudvirt1057 - filippo@cumin1003" * 09:19 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "cloudvirt1057 - filippo@cumin1003" * 09:18 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1056.eqiad.wmnet with reason: host reimage * 09:18 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1057.eqiad.wmnet with OS trixie * 09:18 filippo@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 09:18 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 09:15 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 09:14 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 09:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host irc1003.wikimedia.org * 09:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:13 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1065.eqiad.wmnet with OS trixie * 09:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 09:09 moritzm: installing Postgresql security updates * 09:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host irc1003.wikimedia.org * 09:07 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 09:07 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:07 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1173.eqiad.wmnet with OS bookworm * 09:06 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw2001.wikimedia.org with OS trixie * 09:03 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.convert-disks (exit_code=0) for host ms-be1064 * 09:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1064.eqiad.wmnet with OS trixie * 09:03 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1056.eqiad.wmnet with OS trixie * 09:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1057.eqiad.wmnet with reason: host reimage * 09:01 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 08:58 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 08:56 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1057.eqiad.wmnet with reason: host reimage * 08:54 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1208.eqiad.wmnet with OS bookworm * 08:53 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2902.codfw.wmnet with OS trixie * 08:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 08:50 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1174.eqiad.wmnet with OS bookworm * 08:46 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1172.eqiad.wmnet with OS bookworm * 08:45 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1173.eqiad.wmnet with reason: host reimage * 08:41 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 08:40 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1057.eqiad.wmnet with OS trixie * 08:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1057.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 08:39 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1207.eqiad.wmnet with OS bookworm * 08:38 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2902.codfw.wmnet with reason: host reimage * 08:37 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 08:34 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1068.eqiad.wmnet with OS trixie * 08:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1208.eqiad.wmnet with reason: host reimage * 08:31 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1057.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 08:29 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 08:28 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1222.eqiad.wmnet onto db1276.eqiad.wmnet * 08:28 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1222: Pool db1222.eqiad.wmnet in after cloning * 08:28 fceratto@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2902.codfw.wmnet with reason: host reimage * 08:27 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1055.eqiad.wmnet with OS trixie * 08:27 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 08:26 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1174.eqiad.wmnet with reason: host reimage * 08:25 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 08:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-misc2002.codfw.wmnet * 08:23 topranks: reboot pfw1-codfw firewall pair to upgrade JunOS [[phab:T434865|T434865]] * 08:22 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1172.eqiad.wmnet with reason: host reimage * 08:20 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1064.eqiad.wmnet with OS trixie * 08:20 mvernon@cumin2003: START - Cookbook sre.swift.convert-disks for host ms-be1065 * 08:18 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1207.eqiad.wmnet with reason: host reimage * 08:17 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1174.eqiad.wmnet with reason: host reimage * 08:17 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1173.eqiad.wmnet with reason: host reimage * 08:17 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1172.eqiad.wmnet with reason: host reimage * 08:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host mc-misc2002.codfw.wmnet * 08:15 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1208.eqiad.wmnet with reason: host reimage * 08:15 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1207.eqiad.wmnet with reason: host reimage * 08:14 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db2902.codfw.wmnet with OS trixie * 08:14 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.16 refs [[phab:T430835|T430835]] * 08:13 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db2902.codfw.wmnet * 08:13 fceratto@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host db2902.codfw.wmnet with OS trixie * 08:10 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr[1-2]-codfw with reason: upgrade pfw1a-codfw and pfw1b-codfw pair * 08:09 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1055.eqiad.wmnet with reason: host reimage * 08:07 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on pfw1-codfw with reason: upgrade pfw1a-codfw and pfw1b-codfw pair * 08:03 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1055.eqiad.wmnet with reason: host reimage * 08:02 arnaudb@dns1006: END - running authdns-update * 08:02 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1208.eqiad.wmnet with OS bookworm * 08:02 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1207.eqiad.wmnet with OS bookworm * 08:01 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1174.eqiad.wmnet with OS bookworm * 08:01 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1173.eqiad.wmnet with OS bookworm * 08:01 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1172.eqiad.wmnet with OS bookworm * 07:59 arnaudb@dns1006: START - running authdns-update * 07:58 arnaudb@dns1006: START - running authdns-update * 07:48 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1055.eqiad.wmnet with OS trixie * 07:45 moritzm: extend the disk of ldap-rw2001 by 80G [[phab:T331699|T331699]] * 07:42 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1222: Pool db1222.eqiad.wmnet in after cloning * 07:36 mvernon@cumin2003: START - Cookbook sre.swift.convert-disks for host ms-be1068 * 07:35 mvernon@cumin2003: START - Cookbook sre.swift.convert-disks for host ms-be1064 * 07:17 moritzm: installing imagemagick security updates * 07:14 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1277: Pool back * 07:14 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon1003.wikimedia.org * 07:07 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon1003.wikimedia.org * 07:03 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1280: Pool back * 07:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon2002.wikimedia.org * 06:55 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon2002.wikimedia.org * 06:54 moritzm: installing php8.2 security updates * 06:51 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1284: Pool back * 06:49 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1222: Depool db1222.eqiad.wmnet to then clone it to db1276.eqiad.wmnet - marostegui@cumin1003 * 06:49 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1222: Depool db1222.eqiad.wmnet to then clone it to db1276.eqiad.wmnet - marostegui@cumin1003 * 06:49 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1222.eqiad.wmnet onto db1276.eqiad.wmnet * 06:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd1005.eqiad.wmnet * 06:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd1005.eqiad.wmnet * 06:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd1004.eqiad.wmnet * 06:36 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2209: db2209 repool * 06:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd1004.eqiad.wmnet * 06:32 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast3007.wikimedia.org * 06:29 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1277: Pool back * 06:28 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1277 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96190 and previous config saved to /var/cache/conftool/dbconfig/20260819-062815-marostegui.json * 06:26 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast3007.wikimedia.org * 06:22 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul1001.eqiad.wmnet * 06:18 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul1001.eqiad.wmnet * 06:18 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul1003.eqiad.wmnet * 06:18 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1280: Pool back * 06:17 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1284 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96186 and previous config saved to /var/cache/conftool/dbconfig/20260819-061743-marostegui.json * 06:14 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul1003.eqiad.wmnet * 06:14 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul1002.eqiad.wmnet * 06:10 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul1002.eqiad.wmnet * 06:10 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2048.codfw.wmnet * 06:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2048.codfw.wmnet * 06:06 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1284: Pool back * 06:06 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1284 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96184 and previous config saved to /var/cache/conftool/dbconfig/20260819-060621-marostegui.json * 06:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2048.codfw.wmnet * 05:59 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2048.codfw.wmnet * 05:51 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2209: db2209 repool * 03:16 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1186.eqiad.wmnet with OS bookworm * 02:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1186.eqiad.wmnet with reason: host reimage * 02:46 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1186.eqiad.wmnet with reason: host reimage * 02:46 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2207 [[phab:T435270|T435270]]', diff saved to https://phabricator.wikimedia.org/P96181 and previous config saved to /var/cache/conftool/dbconfig/20260819-024627-marostegui.json * 02:44 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2204 to s2 primary [[phab:T435270|T435270]]', diff saved to https://phabricator.wikimedia.org/P96180 and previous config saved to /var/cache/conftool/dbconfig/20260819-024403-marostegui.json * 02:43 marostegui: Starting s2 codfw failover from db2207 to db2204 - [[phab:T435270|T435270]] * 02:39 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2204 with weight 0 [[phab:T435270|T435270]]', diff saved to https://phabricator.wikimedia.org/P96179 and previous config saved to /var/cache/conftool/dbconfig/20260819-023951-marostegui.json * 02:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s2 [[phab:T435270|T435270]] * 02:32 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1186.eqiad.wmnet with OS bookworm * 02:29 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-worker1186.eqiad.wmnet with OS bookworm * 02:18 denisse@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2207: Depooling replica * 02:18 denisse@cumin1003: START - Cookbook sre.mysql.depool depool db2207: Depooling replica * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 48s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-18 == * 23:55 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1264845{{!}}Remove unused/redundant wgMFNoindexPages=true setting (T255458)]] (duration: 09m 42s) * 23:51 krinkle@deploy1003: krinkle: Continuing with deployment * 23:48 krinkle@deploy1003: krinkle: Backport for [[gerrit:1264845{{!}}Remove unused/redundant wgMFNoindexPages=true setting (T255458)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:45 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1264845{{!}}Remove unused/redundant wgMFNoindexPages=true setting (T255458)]] * 23:38 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326944{{!}}Retire filebackend lock manager in favour of the default one (T366938)]] (duration: 08m 55s) * 23:34 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 23:31 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326944{{!}}Retire filebackend lock manager in favour of the default one (T366938)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:29 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326944{{!}}Retire filebackend lock manager in favour of the default one (T366938)]] * 23:27 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1170.eqiad.wmnet with OS bookworm * 23:21 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1205.eqiad.wmnet with OS bookworm * 23:20 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1171.eqiad.wmnet with OS bookworm * 23:15 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1206.eqiad.wmnet with OS bookworm * 23:05 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1170.eqiad.wmnet with reason: host reimage * 23:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1205.eqiad.wmnet with reason: host reimage * 22:57 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1171.eqiad.wmnet with reason: host reimage * 22:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1206.eqiad.wmnet with reason: host reimage * 22:53 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1205.eqiad.wmnet with reason: host reimage * 22:51 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1171.eqiad.wmnet with reason: host reimage * 22:51 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1170.eqiad.wmnet with reason: host reimage * 22:50 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1206.eqiad.wmnet with reason: host reimage * 22:36 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1206.eqiad.wmnet with OS bookworm * 22:35 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1205.eqiad.wmnet with OS bookworm * 22:35 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1186.eqiad.wmnet with OS bookworm * 22:35 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1171.eqiad.wmnet with OS bookworm * 22:35 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1170.eqiad.wmnet with OS bookworm * 22:33 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-worker1194.eqiad.wmnet with OS bookworm * 22:22 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326923{{!}}Enable redis lock manager on s6 (T366938)]] (duration: 11m 52s) * 22:18 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 22:13 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326923{{!}}Enable redis lock manager on s6 (T366938)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:10 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326923{{!}}Enable redis lock manager on s6 (T366938)]] * 22:04 sbassett: Deployed security fix for [[phab:T435234|T435234]] (wmf.16) * 21:54 sbassett: Deployed security fix for [[phab:T435234|T435234]] (wmf.15) * 21:38 caro@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326925{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326926{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326929{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]], [[gerrit:1326928{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]] (duration: 0 * 21:34 caro@deploy1003: caro: Continuing with deployment * 21:33 caro@deploy1003: caro: Backport for [[gerrit:1326925{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326926{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326929{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]], [[gerrit:1326928{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]] synced to the testservers (see h * 21:31 caro@deploy1003: Started scap sync-world: Backport for [[gerrit:1326925{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326926{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326929{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]], [[gerrit:1326928{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]] * 21:24 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1204.eqiad.wmnet with reason: 1204 datanode repair [[phab:T434494|T434494]] * 21:02 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326896{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]], [[gerrit:1326897{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]] (duration: 13m 56s) * 20:58 krinkle@deploy1003: krinkle: Continuing with deployment * 20:50 krinkle@deploy1003: krinkle: Backport for [[gerrit:1326896{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]], [[gerrit:1326897{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:49 ryankemper: `an-launcher1003` terminated process group `666809` (`rest_backfill_phase1.sh`) ~20 mins ago with `sudo kill -TERM -- -666809` after its local spark driver (`--driver-memory 64g`) repeatedly exhausted memory on the 32 GB VM and caused SSH to intermittently flap; host recovered to 27 GB available memory * 20:48 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1326896{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]], [[gerrit:1326897{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]] * 20:35 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 20:33 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 20:31 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 20:31 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326870{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]], [[gerrit:1326871{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]] (duration: 07m 35s) * 20:28 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 20:26 kemayo@deploy1003: kemayo: Continuing with deployment * 20:26 ryankemper: `an-launcher1003` confirmed the host is flapping because of memory thrash. chasing down the source of the thrash * 20:25 kemayo@deploy1003: kemayo: Backport for [[gerrit:1326870{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]], [[gerrit:1326871{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:23 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1326870{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]], [[gerrit:1326871{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]] * 20:23 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 20:20 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 20:14 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1168.eqiad.wmnet with OS bookworm * 20:08 zabe: zabe@deploy1003:~$ mwscript extensions/WikimediaMaintenance/maintenance/fixFileRevisionArchiveNameDrift.php enwiki # [[phab:T428406|T428406]] * 20:08 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1204.eqiad.wmnet with OS bookworm * 20:05 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1167.eqiad.wmnet with OS bookworm * 20:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1166.eqiad.wmnet with OS bookworm * 19:54 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1203.eqiad.wmnet with OS bookworm * 19:53 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326907{{!}}Revert "Disable redis lock manager on testwiki"]] (duration: 11m 05s) * 19:50 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1168.eqiad.wmnet with reason: host reimage * 19:47 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1204.eqiad.wmnet with reason: host reimage * 19:46 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 19:44 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326907{{!}}Revert "Disable redis lock manager on testwiki"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:42 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326907{{!}}Revert "Disable redis lock manager on testwiki"]] * 19:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1167.eqiad.wmnet with reason: host reimage * 19:37 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1166.eqiad.wmnet with reason: host reimage * 19:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1203.eqiad.wmnet with reason: host reimage * 19:32 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1167.eqiad.wmnet with reason: host reimage * 19:32 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1168.eqiad.wmnet with reason: host reimage * 19:32 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1166.eqiad.wmnet with reason: host reimage * 19:31 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1204.eqiad.wmnet with reason: host reimage * 19:31 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1203.eqiad.wmnet with reason: host reimage * 19:26 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324751{{!}}InitialiseSettings: Enable 2FA enforcement on more private wikis (T428103)]], [[gerrit:1326875{{!}}Add banner notifying of upcoming 2FA enforcement (T420792)]] (duration: 31m 46s) * 19:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1204.eqiad.wmnet with OS bookworm * 19:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1203.eqiad.wmnet with OS bookworm * 19:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1168.eqiad.wmnet with OS bookworm * 19:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1167.eqiad.wmnet with OS bookworm * 19:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1166.eqiad.wmnet with OS bookworm * 19:15 denisse: rebooting kafkamon2003.codfw.wmnet - [[phab:T435162|T435162]] * 19:14 denisse: rebooting kafkamon1003.eqiad.wmnet [[phab:T435162|T435162]] * 19:13 reedy@deploy1003: reedy: Continuing with deployment * 19:12 reedy@deploy1003: reedy: Backport for [[gerrit:1324751{{!}}InitialiseSettings: Enable 2FA enforcement on more private wikis (T428103)]], [[gerrit:1326875{{!}}Add banner notifying of upcoming 2FA enforcement (T420792)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:54 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324751{{!}}InitialiseSettings: Enable 2FA enforcement on more private wikis (T428103)]], [[gerrit:1326875{{!}}Add banner notifying of upcoming 2FA enforcement (T420792)]] * 18:50 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 18:44 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.16 refs [[phab:T430835|T430835]] * 18:34 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 18:31 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 18:22 aklapper@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326852{{!}}CategoryTree: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]], [[gerrit:1326853{{!}}CategoryViewer: Allow null $html in the CategoryViewerGenerateLink hook (T435161)]], [[gerrit:1326865{{!}}Flow: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]] (duration: 09m 57s) * 18:18 aklapper@deploy1003: jforrester, aklapper: Continuing with deployment * 18:17 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 18:14 aklapper@deploy1003: jforrester, aklapper: Backport for [[gerrit:1326852{{!}}CategoryTree: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]], [[gerrit:1326853{{!}}CategoryViewer: Allow null $html in the CategoryViewerGenerateLink hook (T435161)]], [[gerrit:1326865{{!}}Flow: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki * 18:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1156.eqiad.wmnet with OS bookworm * 18:12 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 18:12 aklapper@deploy1003: Started scap sync-world: Backport for [[gerrit:1326852{{!}}CategoryTree: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]], [[gerrit:1326853{{!}}CategoryViewer: Allow null $html in the CategoryViewerGenerateLink hook (T435161)]], [[gerrit:1326865{{!}}Flow: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]] * 18:11 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 18:08 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1146.eqiad.wmnet with OS bookworm * 18:07 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1177.eqiad.wmnet with OS bookworm * 18:00 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326380{{!}}Introduce main lock manager service (T366938 T427999)]] (duration: 11m 25s) * 17:58 ladsgroup@deploy1003: ladsgroup: Rolling back deployment * 17:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1202.eqiad.wmnet with OS bookworm * 17:55 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1201.eqiad.wmnet with OS bookworm * 17:50 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326380{{!}}Introduce main lock manager service (T366938 T427999)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:48 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326380{{!}}Introduce main lock manager service (T366938 T427999)]] * 17:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1156.eqiad.wmnet with reason: host reimage * 17:46 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1146.eqiad.wmnet with reason: host reimage * 17:45 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-esams and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 17:42 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 17:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1177.eqiad.wmnet with reason: host reimage * 17:38 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1202.eqiad.wmnet with reason: host reimage * 17:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1201.eqiad.wmnet with reason: host reimage * 17:30 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1156.eqiad.wmnet with reason: host reimage * 17:29 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1177.eqiad.wmnet with reason: host reimage * 17:28 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1146.eqiad.wmnet with reason: host reimage * 17:28 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1202.eqiad.wmnet with reason: host reimage * 17:27 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1201.eqiad.wmnet with reason: host reimage * 17:25 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 17:21 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 17:14 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1202.eqiad.wmnet with OS bookworm * 17:14 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1201.eqiad.wmnet with OS bookworm * 17:14 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1177.eqiad.wmnet with OS bookworm * 17:14 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1156.eqiad.wmnet with OS bookworm * 17:14 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1146.eqiad.wmnet with OS bookworm * 17:09 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326885{{!}}w/deployment-info.php: Handle new file format (T434726)]] (duration: 07m 15s) * 17:08 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 17:05 dancy@deploy1003: dancy: Continuing with deployment * 17:04 dancy@deploy1003: dancy: Backport for [[gerrit:1326885{{!}}w/deployment-info.php: Handle new file format (T434726)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:02 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1326885{{!}}w/deployment-info.php: Handle new file format (T434726)]] * 16:50 dancy@deploy1003: Finished scap sync-world: Testing [[phab:T434726|T434726]] (duration: 06m 40s) * 16:43 dancy@deploy1003: Started scap sync-world: Testing [[phab:T434726|T434726]] * 16:43 dancy@deploy1003: Installation of scap version "4.282.0" completed for 3 hosts * 16:41 dancy@deploy1003: Installing scap version "4.282.0" for 3 host(s) * 16:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1145.eqiad.wmnet with OS bookworm * 16:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1200.eqiad.wmnet with OS bookworm * 16:32 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1199.eqiad.wmnet with OS bookworm * 16:18 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1145.eqiad.wmnet with reason: host reimage * 16:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1200.eqiad.wmnet with reason: host reimage * 16:09 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1199.eqiad.wmnet with reason: host reimage * 16:05 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-esams and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 16:05 cjd91: sudo -i cookbook sre.cdn.roll-upgrade-ats --query 'A:cp-esams' --task-id [[phab:T434478|T434478]] --reason '9.2.15 upgrade' * 16:03 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1200.eqiad.wmnet with reason: host reimage * 16:02 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1145.eqiad.wmnet with reason: host reimage * 16:02 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1199.eqiad.wmnet with reason: host reimage * 15:50 moritzm: installing zip security updates * 15:48 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1200.eqiad.wmnet with OS bookworm * 15:47 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1199.eqiad.wmnet with OS bookworm * 15:47 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1145.eqiad.wmnet with OS bookworm * 15:41 topranks: bounce PIC 0/0 on cr1-magru to set port to 40G * 15:37 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1236.eqiad.wmnet with OS bookworm * 15:29 aikochou@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'ores-legacy' for release 'main' . * 15:26 aikochou@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'ores-legacy' for release 'main' . * 15:20 aikochou@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'ores-legacy' for release 'main' . * 15:14 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2002.codfw.wmnet * 15:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1236.eqiad.wmnet with reason: host reimage * 15:12 moritzm: failover ganeti master in codfw to ganeti2047 * 15:09 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1236.eqiad.wmnet with reason: host reimage * 15:09 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2044.codfw.wmnet * 15:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2002.codfw.wmnet * 15:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2044.codfw.wmnet * 15:04 brennen@deploy1003: Finished deploy [phabricator/deployment@6b9b6ff]: deploy phab1004 for [[phab:T435213|T435213]] (duration: 01m 01s) * 15:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2044.codfw.wmnet * 15:03 brennen@deploy1003: Started deploy [phabricator/deployment@6b9b6ff]: deploy phab1004 for [[phab:T435213|T435213]] * 15:03 brennen@deploy1003: Finished deploy [phabricator/deployment@6b9b6ff]: deploy phab2003 for [[phab:T435213|T435213]] (duration: 00m 57s) * 15:02 brennen@deploy1003: Started deploy [phabricator/deployment@6b9b6ff]: deploy phab2003 for [[phab:T435213|T435213]] * 14:57 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply * 14:57 arnaudb@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on phab2003.codfw.wmnet,phab[1004-1006].eqiad.wmnet with reason: maintenance * 14:56 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2044.codfw.wmnet * 14:55 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply * 14:53 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1236.eqiad.wmnet with OS bookworm * 14:51 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2043.codfw.wmnet * 14:51 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2043.codfw.wmnet * 14:45 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2043.codfw.wmnet * 14:33 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2043.codfw.wmnet * 14:22 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2042.codfw.wmnet * 14:22 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2042.codfw.wmnet * 14:21 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db2902.codfw.wmnet with OS trixie * 14:16 elukey: uploaded spicerack_13.2.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia * 14:16 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2042.codfw.wmnet * 14:04 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2042.codfw.wmnet * 14:04 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2041.codfw.wmnet * 14:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2041.codfw.wmnet * 13:59 phuedx@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: apply * 13:59 phuedx@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-main: apply * 13:59 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326838{{!}}Enable Suggested Investigations on hewiki (T435146)]] (duration: 11m 50s) * 13:58 phuedx@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: apply * 13:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2041.codfw.wmnet * 13:57 phuedx@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-main: apply * 13:57 phuedx@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-main: apply * 13:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard1003.eqiad.wmnet * 13:57 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-main: apply * 13:55 phuedx@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-logging-external: apply * 13:55 phuedx@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-logging-external: apply * 13:54 phuedx@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-logging-external: apply * 13:54 stran@deploy1003: stran: Continuing with deployment * 13:54 phuedx@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-logging-external: apply * 13:54 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-logging-external: apply * 13:54 phuedx@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-logging-external: apply * 13:53 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-logging-external: apply * 13:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard1003.eqiad.wmnet * 13:53 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2041.codfw.wmnet * 13:51 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db2902.codfw.wmnet - fceratto@cumin1003" * 13:51 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db2902.codfw.wmnet - fceratto@cumin1003" * 13:51 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2026.codfw.wmnet * 13:51 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard2003.codfw.wmnet * 13:50 phuedx@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: apply * 13:50 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-eqiad * 13:50 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp1001.eqiad.wmnet * 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp1001.eqiad.wmnet * 13:50 phuedx@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: apply * 13:49 phuedx@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: apply * 13:49 stran@deploy1003: stran: Backport for [[gerrit:1326838{{!}}Enable Suggested Investigations on hewiki (T435146)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:48 phuedx@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: apply * 13:48 phuedx@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics: apply * 13:47 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard2003.codfw.wmnet * 13:47 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics: apply * 13:47 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1326838{{!}}Enable Suggested Investigations on hewiki (T435146)]] * 13:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp1001.eqiad.wmnet * 13:44 cdanis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 13:43 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp1001.eqiad.wmnet * 13:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1315-1327].eqiad.wmnet * 13:43 cdanis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 13:43 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1315-1327].eqiad.wmnet * 13:40 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki1001.eqiad.wmnet * 13:35 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1315-1327].eqiad.wmnet * 13:34 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host rpki1001.eqiad.wmnet * 13:27 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1315-1327].eqiad.wmnet * 13:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1302-1314].eqiad.wmnet * 13:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1302-1314].eqiad.wmnet * 13:21 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326775{{!}}SI: Instrument case update on first edit (T435048)]] (duration: 07m 12s) * 13:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1302-1314].eqiad.wmnet * 13:17 stran@deploy1003: stran: Continuing with deployment * 13:16 stran@deploy1003: stran: Backport for [[gerrit:1326775{{!}}SI: Instrument case update on first edit (T435048)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:14 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1326775{{!}}SI: Instrument case update on first edit (T435048)]] * 13:13 moritzm: installing util-linux security updates * 13:11 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1302-1314].eqiad.wmnet * 13:10 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326232{{!}}prv: Enable parsoid rendering for 5 wikis (T435115)]] (duration: 08m 17s) * 13:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1288-1289,1291-1301].eqiad.wmnet * 13:10 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply * 13:10 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1288-1289,1291-1301].eqiad.wmnet * 13:10 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply * 13:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki2003.codfw.wmnet * 13:06 jgiannelos@deploy1003: jgiannelos: Continuing with deployment * 13:05 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host rpki2003.codfw.wmnet * 13:04 jgiannelos@deploy1003: jgiannelos: Backport for [[gerrit:1326232{{!}}prv: Enable parsoid rendering for 5 wikis (T435115)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1288-1289,1291-1301].eqiad.wmnet * 13:02 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1326232{{!}}prv: Enable parsoid rendering for 5 wikis (T435115)]] * 12:53 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1288-1289,1291-1301].eqiad.wmnet * 12:53 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1273,1275-1287].eqiad.wmnet * 12:53 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1273,1275-1287].eqiad.wmnet * 12:52 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt1002.wikimedia.org * 12:51 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db2902.codfw.wmnet on all recursors * 12:51 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db2902.codfw.wmnet on all recursors * 12:51 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:51 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db2902.codfw.wmnet - fceratto@cumin1003" * 12:51 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db2902.codfw.wmnet - fceratto@cumin1003" * 12:46 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt1002.wikimedia.org * 12:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt2002.wikimedia.org * 12:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1273,1275-1287].eqiad.wmnet * 12:42 dhinus: repooled clouddb1032 that was currently <nowiki>{</nowiki>"weight": 0, "pooled": "inactive"<nowiki>}</nowiki> for both s4 and s6 * 12:41 dhinus: also depooled clouddb1017 (forgot it in the previous list) * 12:41 fnegri@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet * 12:40 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 12:40 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db2902.codfw.wmnet * 12:40 fnegri@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032.eqiad.wmnet * 12:40 fnegri@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032 * 12:39 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1017.eqiad.wmnet * 12:39 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt2002.wikimedia.org * 12:38 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1020.eqiad.wmnet * 12:38 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1018.eqiad.wmnet * 12:38 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet * 12:37 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1014.eqiad.wmnet * 12:37 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1013.eqiad.wmnet * 12:37 dhinus: depool again clouddb10[13,14,16,18,20] that were repooled by the cookbook sre.mysql.multiinstance_reboot * 12:36 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1273,1275-1287].eqiad.wmnet * 12:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1248-1261].eqiad.wmnet * 12:36 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1248-1261].eqiad.wmnet * 12:35 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2026.codfw.wmnet * 12:34 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 12:34 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 12:34 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 12:34 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 12:29 lucaswerkmeister-wmde@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 12:28 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2026.codfw.wmnet * 12:28 lucaswerkmeister-wmde@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 12:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1248-1261].eqiad.wmnet * 12:27 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.addnode (exit_code=0) for new host ganeti2046.codfw.wmnet to cluster codfw and group A * 12:26 moritzm: readded ganeti2046 to the codfw cluster following firmware update and reimage [[phab:T434681|T434681]] * 12:23 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply * 12:23 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply * 12:23 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply * 12:22 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply * 12:22 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply * 12:22 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply * 12:21 jmm@cumin2003: START - Cookbook sre.ganeti.addnode for new host ganeti2046.codfw.wmnet to cluster codfw and group A * 12:21 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1169.eqiad.wmnet onto db1283.eqiad.wmnet * 12:21 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1169: Pool db1169.eqiad.wmnet in after cloning * 12:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1248-1261].eqiad.wmnet * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1149-1153,1158,1240-1247].eqiad.wmnet * 12:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1149-1153,1158,1240-1247].eqiad.wmnet * 12:11 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1149-1153,1158,1240-1247].eqiad.wmnet * 12:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1168.eqiad.wmnet onto db1282.eqiad.wmnet * 12:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1168: Pool db1168.eqiad.wmnet in after cloning * 12:06 jmm@cumin2003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-test-eqiad * 12:05 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2026.codfw.wmnet * 12:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1149-1153,1158,1240-1247].eqiad.wmnet * 12:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1128-1134,1142-1148].eqiad.wmnet * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2046.codfw.wmnet * 12:01 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1128-1134,1142-1148].eqiad.wmnet * 11:57 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2025.codfw.wmnet * 11:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2025.codfw.wmnet * 11:54 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2046.codfw.wmnet * 11:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1128-1134,1142-1148].eqiad.wmnet * 11:50 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2025.codfw.wmnet * 11:46 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1128-1134,1142-1148].eqiad.wmnet * 11:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1114-1127].eqiad.wmnet * 11:45 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1114-1127].eqiad.wmnet * 11:39 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db2901.codfw.wmnet * 11:39 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db2901.codfw.wmnet * 11:36 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2025.codfw.wmnet * 11:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1114-1127].eqiad.wmnet * 11:35 fceratto@cumin1003: END (ERROR) - Cookbook sre.ganeti.makevm (exit_code=93) for new host db1901.eqiad.wmnet * 11:35 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1169: Pool db1169.eqiad.wmnet in after cloning * 11:35 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 11:30 jmm@cumin2003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-test-eqiad * 11:27 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1114-1127].eqiad.wmnet * 11:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1076-1081,1084-1087,1093-1095,1113].eqiad.wmnet * 11:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1076-1081,1084-1087,1093-1095,1113].eqiad.wmnet * 11:25 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1168: Pool db1168.eqiad.wmnet in after cloning * 11:24 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1194.eqiad.wmnet with OS bookworm * 11:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1181.eqiad.wmnet with OS bookworm * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2050.codfw.wmnet * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2050.codfw.wmnet * 11:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1076-1081,1084-1087,1093-1095,1113].eqiad.wmnet * 11:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2050.codfw.wmnet * 11:14 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2004.codfw.wmnet * 11:10 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2050.codfw.wmnet * 11:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1076-1081,1084-1087,1093-1095,1113].eqiad.wmnet * 11:09 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1286: Pool back * 11:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1045-1050,1056-1057,1064-1066,1073-1075].eqiad.wmnet * 11:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1045-1050,1056-1057,1064-1066,1073-1075].eqiad.wmnet * 11:08 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2004.codfw.wmnet * 11:07 moritzm: installing PHP 8.4 security updates * 11:06 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on an-worker1194.eqiad.wmnet with reason: host reimage * 11:06 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1194.eqiad.wmnet with reason: host reimage * 11:05 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2049.codfw.wmnet * 11:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2049.codfw.wmnet * 11:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid1003.eqiad.wmnet * 11:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1045-1050,1056-1057,1064-1066,1073-1075].eqiad.wmnet * 10:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1181.eqiad.wmnet with reason: host reimage * 10:59 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2049.codfw.wmnet * 10:59 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid1003.eqiad.wmnet * 10:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid2003.codfw.wmnet * 10:54 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1181.eqiad.wmnet with reason: host reimage * 10:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid2003.codfw.wmnet * 10:51 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1901.eqiad.wmnet on all recursors * 10:51 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1901.eqiad.wmnet on all recursors * 10:51 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:51 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:51 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:50 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2049.codfw.wmnet * 10:50 blake@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:50 blake@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:49 blake@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:48 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1045-1050,1056-1057,1064-1066,1073-1075].eqiad.wmnet * 10:48 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1044].eqiad.wmnet * 10:48 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1044].eqiad.wmnet * 10:47 blake@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:46 blake@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:46 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2047.codfw.wmnet * 10:46 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:46 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2047.codfw.wmnet * 10:46 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1901.eqiad.wmnet * 10:46 blake@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:44 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1168: Depool db1168.eqiad.wmnet to then clone it to db1282.eqiad.wmnet - marostegui@cumin1003 * 10:44 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1168: Depool db1168.eqiad.wmnet to then clone it to db1282.eqiad.wmnet - marostegui@cumin1003 * 10:44 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1168.eqiad.wmnet onto db1282.eqiad.wmnet * 10:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2047.codfw.wmnet * 10:40 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1044].eqiad.wmnet * 10:37 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2047.codfw.wmnet * 10:37 fceratto@cumin1003: END (ERROR) - Cookbook sre.ganeti.makevm (exit_code=93) for new host db1901.eqiad.wmnet * 10:36 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 10:34 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:33 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2032.codfw.wmnet * 10:33 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2032.codfw.wmnet * 10:32 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1044].eqiad.wmnet * 10:32 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-eqiad * 10:27 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2032.codfw.wmnet * 10:24 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1286: Pool back * 10:24 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1286 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96163 and previous config saved to /var/cache/conftool/dbconfig/20260818-102431-marostegui.json * 10:22 moritzm: installing Django security updates * 10:20 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul2002.codfw.wmnet * 10:20 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul2001.codfw.wmnet * 10:17 blake@deploy1003: Finished scap sync-world: no-build deployment for [[phab:T417800|T417800]] (duration: 04m 40s) * 10:16 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul2002.codfw.wmnet * 10:16 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul2001.codfw.wmnet * 10:15 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2032.codfw.wmnet * 10:14 blake@deploy1003: Started scap sync-world: no-build deployment for [[phab:T417800|T417800]] * 10:12 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul2003.codfw.wmnet * 10:12 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy2001.codfw.wmnet * 10:12 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy3001.esams.wmnet * 10:12 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy1001.eqiad.wmnet * 10:08 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2031.codfw.wmnet * 10:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2031.codfw.wmnet * 10:08 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul2003.codfw.wmnet * 10:08 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy2001.codfw.wmnet * 10:08 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy3001.esams.wmnet * 10:08 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy1001.eqiad.wmnet * 10:07 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy1002.eqiad.wmnet * 10:05 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy2002.codfw.wmnet * 10:05 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy3002.esams.wmnet * 10:04 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy1002.eqiad.wmnet * 10:03 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy4003.ulsfo.wmnet * 10:03 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy5003.eqsin.wmnet * 10:02 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2031.codfw.wmnet * 10:02 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1169: Depool db1169.eqiad.wmnet to then clone it to db1283.eqiad.wmnet - marostegui@cumin1003 * 10:01 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy2002.codfw.wmnet * 10:01 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy4004.ulsfo.wmnet * 10:01 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy3002.esams.wmnet * 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1169: Depool db1169.eqiad.wmnet to then clone it to db1283.eqiad.wmnet - marostegui@cumin1003 * 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1169.eqiad.wmnet onto db1283.eqiad.wmnet * 10:01 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy5004.eqsin.wmnet * 09:59 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy4003.ulsfo.wmnet * 09:59 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy4004.ulsfo.wmnet * 09:59 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy5003.eqsin.wmnet * 09:59 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy7001.magru.wmnet * 09:59 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy5004.eqsin.wmnet * 09:58 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy7002.magru.wmnet * 09:57 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2031.codfw.wmnet * 09:57 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy6001.drmrs.wmnet * 09:57 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy6002.drmrs.wmnet * 09:54 filippo@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for 10 hosts * 09:54 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast4006.wikimedia.org * 09:53 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy6001.drmrs.wmnet * 09:53 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy6002.drmrs.wmnet * 09:53 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host releases1003.eqiad.wmnet * 09:53 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy7001.magru.wmnet * 09:53 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host people2004.codfw.wmnet * 09:52 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host people1005.eqiad.wmnet * 09:52 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy7002.magru.wmnet * 09:50 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host releases2003.codfw.wmnet * 09:49 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host releases2003.codfw.wmnet * 09:49 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host releases1003.eqiad.wmnet * 09:49 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host people2004.codfw.wmnet * 09:48 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host people1005.eqiad.wmnet * 09:46 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2040.codfw.wmnet * 09:46 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2040.codfw.wmnet * 09:43 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1901.eqiad.wmnet on all recursors * 09:42 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1901.eqiad.wmnet on all recursors * 09:42 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:42 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:42 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2040.codfw.wmnet * 09:40 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp2005.wikimedia.org * 09:36 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp2005.wikimedia.org * 09:31 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:31 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1901.eqiad.wmnet * 09:31 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1901.eqiad.wmnet * 09:31 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:31 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1901.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 09:31 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1901.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 09:29 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2040.codfw.wmnet * 09:29 slyngshede@dns1004: END - running authdns-update * 09:28 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2039.codfw.wmnet * 09:28 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2039.codfw.wmnet * 09:27 slyngshede@dns1004: START - running authdns-update * 09:26 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test2005.wikimedia.org * 09:22 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2039.codfw.wmnet * 09:22 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test2005.wikimedia.org * 09:22 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp1005.wikimedia.org * 09:21 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:19 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2039.codfw.wmnet * 09:19 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2038.codfw.wmnet * 09:18 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1287: Pool back * 09:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2038.codfw.wmnet * 09:18 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp1005.wikimedia.org * 09:18 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test1005.wikimedia.org * 09:17 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1901.eqiad.wmnet * 09:15 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db1901.eqiad.wmnet * 09:15 fceratto@cumin1003: END (ERROR) - Cookbook sre.dns.netbox (exit_code=97) * 09:14 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test1005.wikimedia.org * 09:13 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2038.codfw.wmnet * 09:13 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:13 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1901.eqiad.wmnet * 09:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast4006.wikimedia.org * 09:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1288: Pool back * 09:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast5005.wikimedia.org * 09:03 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2038.codfw.wmnet * 08:57 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2037.codfw.wmnet * 08:57 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast5005.wikimedia.org * 08:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2037.codfw.wmnet * 08:56 filippo@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 10 hosts * 08:52 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2037.codfw.wmnet * 08:51 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1289: Pool back * 08:35 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudidp2001-dev.codfw.wmnet * 08:34 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1002-dev.eqiad.wmnet * 08:33 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1287: Pool back * 08:33 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1287 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96150 and previous config saved to /var/cache/conftool/dbconfig/20260818-083311-marostegui.json * 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1001-dev.eqiad.wmnet * 08:31 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudidp2001-dev.codfw.wmnet * 08:30 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1002-dev.eqiad.wmnet * 08:30 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2037.codfw.wmnet * 08:29 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1001-dev.eqiad.wmnet * 08:28 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2036.codfw.wmnet * 08:28 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2036.codfw.wmnet * 08:25 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1288: Pool back * 08:23 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 08:23 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1181.eqiad.wmnet with OS bookworm * 08:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2036.codfw.wmnet * 08:22 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1288 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96148 and previous config saved to /var/cache/conftool/dbconfig/20260818-082234-marostegui.json * 08:20 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2036.codfw.wmnet * 08:18 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2035.codfw.wmnet * 08:18 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudvirt1057.eqiad.wmnet with OS trixie * 08:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2035.codfw.wmnet * 08:13 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2035.codfw.wmnet * 08:06 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2035.codfw.wmnet * 08:05 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1289: Pool back * 08:05 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1289 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96145 and previous config saved to /var/cache/conftool/dbconfig/20260818-080531-marostegui.json * 07:54 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 07:51 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326439{{!}}Fix "mathjax_ignore" handling around forcemathmode attribute (T434686)]] (duration: 13m 15s) * 07:48 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 07:47 krinkle@deploy1003: krinkle: Continuing with deployment * 07:44 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ganeti2046.codfw.wmnet with OS bookworm * 07:40 krinkle@deploy1003: krinkle: Backport for [[gerrit:1326439{{!}}Fix "mathjax_ignore" handling around forcemathmode attribute (T434686)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:39 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 07:38 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1326439{{!}}Fix "mathjax_ignore" handling around forcemathmode attribute (T434686)]] * 07:34 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 07:31 samwilson@deploy1003: Finished scap sync-world: Backport for [[gerrit:701016{{!}}InitialiseSettings and -labs: Remove redundant feature flag $wgWikisourceEnableOcr (T285311)]] (duration: 07m 47s) * 07:30 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1048.eqiad.wmnet * 07:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1048.eqiad.wmnet * 07:29 XioNoX: add gnmic 0.47.0 to bookworm and trixie reprepro * 07:28 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ganeti2046.codfw.wmnet with reason: host reimage * 07:27 samwilson@deploy1003: samwilson: Continuing with deployment * 07:25 samwilson@deploy1003: samwilson: Backport for [[gerrit:701016{{!}}InitialiseSettings and -labs: Remove redundant feature flag $wgWikisourceEnableOcr (T285311)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:25 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 07:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1048.eqiad.wmnet * 07:24 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ganeti2046.codfw.wmnet with reason: host reimage * 07:23 samwilson@deploy1003: Started scap sync-world: Backport for [[gerrit:701016{{!}}InitialiseSettings and -labs: Remove redundant feature flag $wgWikisourceEnableOcr (T285311)]] * 07:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1198.eqiad.wmnet with OS bookworm * 07:18 samwilson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326450{{!}}InitialiseSettings.php: Enable Bulk OCR on pawikisource (T434648)]] (duration: 12m 17s) * 07:17 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 07:11 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1048.eqiad.wmnet * 07:11 samwilson@deploy1003: samwilson: Continuing with deployment * 07:11 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ganeti2046.codfw.wmnet with OS bookworm * 07:10 samwilson@deploy1003: samwilson: Backport for [[gerrit:1326450{{!}}InitialiseSettings.php: Enable Bulk OCR on pawikisource (T434648)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1057.eqiad.wmnet with OS trixie * 07:05 samwilson@deploy1003: Started scap sync-world: Backport for [[gerrit:1326450{{!}}InitialiseSettings.php: Enable Bulk OCR on pawikisource (T434648)]] * 07:02 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1198.eqiad.wmnet with reason: host reimage * 07:01 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1056.eqiad.wmnet with OS trixie * 07:01 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1056.eqiad.wmnet with OS trixie * 07:00 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1056.eqiad.wmnet with OS trixie * 07:00 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1056.eqiad.wmnet with OS trixie * 06:59 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudvirt1055.eqiad.wmnet with OS trixie * 06:58 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1198.eqiad.wmnet with reason: host reimage * 06:53 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1055.eqiad.wmnet with OS trixie * 06:52 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudvirt1054.eqiad.wmnet with OS trixie * 06:44 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1194.eqiad.wmnet with OS bookworm * 06:44 XioNoX: upgrade eqsin gnmic to 0.47.0 * 06:43 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1198.eqiad.wmnet with OS bookworm * 06:41 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1054.eqiad.wmnet with OS trixie * 06:09 arnaudb@cumin1003: END (PASS) - Cookbook sre.gerrit.restart-gerrit (exit_code=0) Restarting Gerrit on gerrit2002 * 06:06 arnaudb@cumin1003: START - Cookbook sre.gerrit.restart-gerrit Restarting Gerrit on gerrit2002 * 06:06 arnaudb@cumin1003: END (PASS) - Cookbook sre.gerrit.restart-gerrit (exit_code=0) Restarting Gerrit on gerrit1003 * 06:04 arnaudb@cumin1003: START - Cookbook sre.gerrit.restart-gerrit Restarting Gerrit on gerrit1003 * 06:02 arnaudb@cumin1003: END (PASS) - Cookbook sre.gerrit.restart-gerrit (exit_code=0) Restarting Gerrit on gerrit2003 * 06:00 arnaudb@cumin1003: START - Cookbook sre.gerrit.restart-gerrit Restarting Gerrit on gerrit2003 * 05:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1155.eqiad.wmnet with OS bookworm * 05:38 arnaudb: updating prometheusBearerToken on gerrit * 05:28 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1155.eqiad.wmnet with reason: host reimage * 05:23 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1155.eqiad.wmnet with reason: host reimage * 05:06 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1155.eqiad.wmnet with OS bookworm * 04:57 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1144.eqiad.wmnet with OS bookworm * 04:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1144.eqiad.wmnet with reason: host reimage * 04:29 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1144.eqiad.wmnet with reason: host reimage * 04:14 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1144.eqiad.wmnet with OS bookworm * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.13 (duration: 02m 23s) * 03:45 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1197.eqiad.wmnet with OS bookworm * 03:41 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1165.eqiad.wmnet with OS bookworm * 03:38 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.16 refs [[phab:T430835|T430835]] (duration: 34m 43s) * 03:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1164.eqiad.wmnet with OS bookworm * 03:35 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1196.eqiad.wmnet with OS bookworm * 03:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1163.eqiad.wmnet with OS bookworm * 03:22 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1197.eqiad.wmnet with reason: host reimage * 03:18 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1165.eqiad.wmnet with reason: host reimage * 03:15 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1196.eqiad.wmnet with reason: host reimage * 03:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1164.eqiad.wmnet with reason: host reimage * 03:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1163.eqiad.wmnet with reason: host reimage * 03:05 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1165.eqiad.wmnet with reason: host reimage * 03:04 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1164.eqiad.wmnet with reason: host reimage * 03:04 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1197.eqiad.wmnet with reason: host reimage * 03:04 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1196.eqiad.wmnet with reason: host reimage * 03:04 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1163.eqiad.wmnet with reason: host reimage * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.16 refs [[phab:T430835|T430835]] * 02:50 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1197.eqiad.wmnet with OS bookworm * 02:49 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1196.eqiad.wmnet with OS bookworm * 02:49 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1165.eqiad.wmnet with OS bookworm * 02:49 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1164.eqiad.wmnet with OS bookworm * 02:48 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1163.eqiad.wmnet with OS bookworm * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 46s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:15 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-codfw: Set storage compatability to NONE — [[phab:T433026|T433026]] - eevans@cumin1003 * 00:38 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-codfw: Set storage compatability to NONE — [[phab:T433026|T433026]] - eevans@cumin1003 * 00:11 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326373{{!}}PersonalDashboard: add newly renamed *ReviewChangesMlModel setting (T422148)]] (duration: 07m 07s) * 00:09 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-eqiad: Set storage compatability to NONE — [[phab:T433026|T433026]] - eevans@cumin1003 * 00:07 musikanimal@deploy1003: musikanimal: Continuing with deployment * 00:06 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1326373{{!}}PersonalDashboard: add newly renamed *ReviewChangesMlModel setting (T422148)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:04 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1326373{{!}}PersonalDashboard: add newly renamed *ReviewChangesMlModel setting (T422148)]] == 2026-08-17 == * 23:30 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-eqiad: Set storage compatability to NONE — [[phab:T433026|T433026]] - eevans@cumin1003 * 23:05 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-codfw: Set storage compatability to UPGRADING — [[phab:T433026|T433026]] - eevans@cumin1003 * 22:28 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-codfw: Set storage compatability to UPGRADING — [[phab:T433026|T433026]] - eevans@cumin1003 * 21:46 logmsgbot: jforrester Deployed security patch for [[phab:T435085|T435085]] * 21:39 swfrench@deploy1003: mwscript-k8s job started: purgeList.php # [[phab:T432412|T432412]] * 21:37 maryum: Undeploy security fix for [[phab:T433020|T433020]] * 21:23 maryum: Deployed security fix for [[phab:T433020|T433020]] * 21:21 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 21:21 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 21:14 maryum: Deployed security fix for [[phab:T434967|T434967]] * 20:58 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-eqiad: Set storage compatability to UPGRADING — [[phab:T433026|T433026]] - eevans@cumin1003 * 20:40 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324320{{!}}[itwiki/slwiki/tgwiki] Remove temporary Wikipedia 25 logos permanently (already reverted) (T414265 T414320 T415307)]] (duration: 06m 55s) * 20:40 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 20:39 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 20:36 cjming@deploy1003: cjming, superpes: Continuing with deployment * 20:35 cjming@deploy1003: cjming, superpes: Backport for [[gerrit:1324320{{!}}[itwiki/slwiki/tgwiki] Remove temporary Wikipedia 25 logos permanently (already reverted) (T414265 T414320 T415307)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:33 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1324320{{!}}[itwiki/slwiki/tgwiki] Remove temporary Wikipedia 25 logos permanently (already reverted) (T414265 T414320 T415307)]] * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ttmserver-test: apply * 20:30 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326361{{!}}Remove escaped paths in app site association file (T432412)]] (duration: 13m 28s) * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ttmserver-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-toolhub-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-toolhub-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-toolhub-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-toolhub-test: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-toolhub: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-toolhub: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-toolhub: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-toolhub: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-test: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-test: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 20:27 inflatador: bking@deploy1003 `charlie --services_dir dse-k8s-services -s opensearch-* -e dse-k8s-* apply` [[phab:T435125|T435125]] * 20:27 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-apifeatureusage-test: apply * 20:27 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-apifeatureusage-test: apply * 20:27 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-apifeatureusage-test: apply * 20:27 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-apifeatureusage-test: apply * 20:27 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-apifeatureusage: apply * 20:26 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-apifeatureusage: apply * 20:26 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-apifeatureusage: apply * 20:26 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-apifeatureusage: apply * 20:26 cjming@deploy1003: cjming, tsev: Continuing with deployment * 20:24 inflatador: bking@deploy1003 `charlie --services_dir dse-k8s-services -s opensearch-* -e dse-k8s-*` [[phab:T435125|T435125]] * 20:21 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-eqiad: Set storage compatability to UPGRADING — [[phab:T433026|T433026]] - eevans@cumin1003 * 20:19 cjming@deploy1003: cjming, tsev: Backport for [[gerrit:1326361{{!}}Remove escaped paths in app site association file (T432412)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:17 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1326361{{!}}Remove escaped paths in app site association file (T432412)]] * 20:15 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326077{{!}}InstrumentConstructiveEdits: anchor all runs to the nearest `interval` (T431493)]] (duration: 06m 25s) * 20:11 cjming@deploy1003: cjming: Continuing with deployment * 20:10 cjming@deploy1003: cjming: Backport for [[gerrit:1326077{{!}}InstrumentConstructiveEdits: anchor all runs to the nearest `interval` (T431493)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:08 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1326077{{!}}InstrumentConstructiveEdits: anchor all runs to the nearest `interval` (T431493)]] * 20:06 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1160.eqiad.wmnet with OS bookworm * 20:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1162.eqiad.wmnet with OS bookworm * 19:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1184.eqiad.wmnet with OS bookworm * 19:55 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1161.eqiad.wmnet with OS bookworm * 19:49 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1195.eqiad.wmnet with OS bookworm * 19:49 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-test: apply * 19:49 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-test: apply * 19:44 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1160.eqiad.wmnet with reason: host reimage * 19:41 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-codfw: Upgrade to Java 17 — [[phab:T433026|T433026]] - eevans@cumin1003 * 19:39 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1162.eqiad.wmnet with reason: host reimage * 19:36 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-eqsin and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 19:36 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-test: apply * 19:36 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1184.eqiad.wmnet with reason: host reimage * 19:32 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1161.eqiad.wmnet with reason: host reimage * 19:29 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1195.eqiad.wmnet with reason: host reimage * 19:26 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1161.eqiad.wmnet with reason: host reimage * 19:26 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1160.eqiad.wmnet with reason: host reimage * 19:26 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1184.eqiad.wmnet with reason: host reimage * 19:26 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1162.eqiad.wmnet with reason: host reimage * 19:25 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1195.eqiad.wmnet with reason: host reimage * 19:19 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-test: apply * 19:11 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1195.eqiad.wmnet with OS bookworm * 19:10 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1194.eqiad.wmnet with OS bookworm * 19:10 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1184.eqiad.wmnet with OS bookworm * 19:10 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1162.eqiad.wmnet with OS bookworm * 19:10 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1161.eqiad.wmnet with OS bookworm * 19:10 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1160.eqiad.wmnet with OS bookworm * 19:10 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 19:09 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 19:03 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-codfw: Upgrade to Java 17 — [[phab:T433026|T433026]] - eevans@cumin1003 * 18:58 dancy@deploy1003: Finished scap sync-world: testing [[phab:T375514|T375514]] (duration: 03m 13s) * 18:55 dancy@deploy1003: Started scap sync-world: testing [[phab:T375514|T375514]] * 18:55 dwisehaupt@dns1006: END - running authdns-update * 18:54 dancy@deploy1003: Installation of scap version "4.281.1" completed for 3 hosts * 18:54 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2009.codfw.wmnet * 18:53 dwisehaupt@dns1006: START - running authdns-update * 18:52 dancy@deploy1003: Installing scap version "4.281.1" for 3 host(s) * 18:47 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2009.codfw.wmnet * 18:41 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2008.codfw.wmnet * 18:34 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2008.codfw.wmnet * 18:30 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2007.codfw.wmnet * 18:23 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2007.codfw.wmnet * 18:16 dwisehaupt@dns1005: END - running authdns-update * 18:14 dwisehaupt@dns1005: START - running authdns-update * 18:04 swfrench@deploy1003: Finished scap sync-world: Deploy "Point Test Wiki to new docroot" - [[phab:T432412|T432412]] (duration: 26m 03s) * 18:00 swfrench@deploy1003: swfrench: Continuing with deployment * 17:51 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1057.eqiad.wmnet with OS trixie * 17:47 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1180.eqiad.wmnet with OS bookworm * 17:39 swfrench@deploy1003: swfrench: Deploy "Point Test Wiki to new docroot" - [[phab:T432412|T432412]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:39 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-eqsin and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 17:38 swfrench@deploy1003: Started scap sync-world: Deploy "Point Test Wiki to new docroot" - [[phab:T432412|T432412]] * 17:36 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1159.eqiad.wmnet with OS bookworm * 17:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1158.eqiad.wmnet with OS bookworm * 17:30 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-eqiad: Upgrade to Java 17 — [[phab:T433026|T433026]] - eevans@cumin1003 * 17:30 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-ulsfo or A:cp-drmrs and A:cp - 9.2.15 upgrade ([[phab:T434620|T434620]]) * 17:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1157.eqiad.wmnet with OS bookworm * 17:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1180.eqiad.wmnet with reason: host reimage * 17:22 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 17:21 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1193.eqiad.wmnet with OS bookworm * 17:21 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 17:21 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 17:20 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 17:20 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 17:20 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1183.eqiad.wmnet with OS bookworm * 17:18 swfrench@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 17:18 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 17:17 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1180.eqiad.wmnet with reason: host reimage * 17:17 swfrench@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 17:17 swfrench@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 17:16 swfrench@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 17:16 swfrench@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 17:15 swfrench@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 17:15 swfrench@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 17:15 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1192.eqiad.wmnet with OS bookworm * 17:14 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1159.eqiad.wmnet with reason: host reimage * 17:14 swfrench@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 17:13 swfrench@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 17:12 swfrench@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 17:09 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1158.eqiad.wmnet with reason: host reimage * 17:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1157.eqiad.wmnet with reason: host reimage * 17:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1180 * 17:02 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1180 * 17:01 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica-eqsin and A:liberica * 17:01 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1193.eqiad.wmnet with reason: host reimage * 16:57 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1183.eqiad.wmnet with reason: host reimage * 16:56 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1054.eqiad.wmnet with OS trixie * 16:55 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1056.eqiad.wmnet with OS trixie * 16:54 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1193.eqiad.wmnet with reason: host reimage * 16:54 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1192.eqiad.wmnet with reason: host reimage * 16:51 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-eqiad: Upgrade to Java 17 — [[phab:T433026|T433026]] - eevans@cumin1003 * 16:50 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1180 * 16:50 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1180.eqiad.wmnet 17.36.64.10.in-addr.arpa 7.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:50 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1180.eqiad.wmnet 17.36.64.10.in-addr.arpa 7.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:50 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:50 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1180 - btullis@cumin1003" * 16:50 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1180 - btullis@cumin1003" * 16:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1158.eqiad.wmnet with reason: host reimage * 16:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1157.eqiad.wmnet with reason: host reimage * 16:49 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica-eqsin and A:liberica * 16:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1183.eqiad.wmnet with reason: host reimage * 16:48 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1159.eqiad.wmnet with reason: host reimage * 16:47 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1192.eqiad.wmnet with reason: host reimage * 16:46 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns4004.wikimedia.org * 16:46 sukhe@dns1004: END - running authdns-update * 16:44 sukhe@dns1004: START - running authdns-update * 16:44 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns4004.wikimedia.org,service=authdns-update * 16:43 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns4004.wikimedia.org with OS trixie * 16:39 btullis@cumin1003: START - Cookbook sre.dns.netbox * 16:39 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1180 * 16:39 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1193.eqiad.wmnet with OS bookworm * 16:39 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1180.eqiad.wmnet with OS bookworm * 16:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1192.eqiad.wmnet with OS bookworm * 16:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1183.eqiad.wmnet with OS bookworm * 16:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1159.eqiad.wmnet with OS bookworm * 16:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1158.eqiad.wmnet with OS bookworm * 16:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1157.eqiad.wmnet with OS bookworm * 16:31 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1057.eqiad.wmnet with OS trixie * 16:30 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1057.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 16:29 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica-drmrs and A:liberica * 16:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1154.eqiad.wmnet with OS bookworm * 16:23 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1057.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 16:23 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1055.eqiad.wmnet with OS trixie * 16:22 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1057 * 16:22 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1057 * 16:21 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:21 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1057] - vriley@cumin1003" * 16:21 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1057] - vriley@cumin1003" * 16:19 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica-drmrs and A:liberica * 16:17 vriley@cumin1003: START - Cookbook sre.dns.netbox * 16:16 phuedx@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics-external: apply * 16:15 phuedx@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics-external: apply * 16:13 phuedx@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics-external: apply * 16:12 phuedx@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics-external: apply * 16:11 phuedx@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics-external: apply * 16:09 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics-external: apply * 16:09 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1176.eqiad.wmnet with OS bookworm * 16:07 btullis@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 16:06 btullis@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 16:05 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1191.eqiad.wmnet with OS bookworm * 16:03 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1154.eqiad.wmnet with reason: host reimage * 16:03 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching aqs[2002-2012].codfw.wmnet,aqs[1017-1027].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433026|T433026]] - eevans@cumin1003 * 15:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1190.eqiad.wmnet with OS bookworm * 15:59 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1154.eqiad.wmnet with reason: host reimage * 15:55 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-debug: apply * 15:55 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-debug: apply * 15:55 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-debug: apply * 15:55 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/mw-debug: apply * 15:53 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns4004.wikimedia.org with reason: host reimage * 15:50 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns4004.wikimedia.org with reason: host reimage * 15:46 btullis@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 15:46 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1176.eqiad.wmnet with reason: host reimage * 15:45 btullis@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 15:43 btullis@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 15:42 btullis@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 15:42 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1191.eqiad.wmnet with reason: host reimage * 15:39 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1190.eqiad.wmnet with reason: host reimage * 15:36 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1054.eqiad.wmnet with OS trixie * 15:35 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:35 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1056.eqiad.wmnet with OS trixie * 15:35 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:34 moritzm: failover Ganeti master in eqiad to ganeti1046 * 15:34 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1176.eqiad.wmnet with reason: host reimage * 15:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1191.eqiad.wmnet with reason: host reimage * 15:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1190.eqiad.wmnet with reason: host reimage * 15:31 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns2006.wikimedia.org * 15:31 sukhe@dns1004: END - running authdns-update * 15:29 sukhe@dns1004: START - running authdns-update * 15:29 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns2006.wikimedia.org,service=authdns-update * 15:29 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns2006.wikimedia.org * 15:29 sukhe@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns2006.wikimedia.org * 15:26 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica-ulsfo and A:liberica * 15:26 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns1006.wikimedia.org * 15:26 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:25 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns2006.wikimedia.org with OS trixie * 15:25 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1056 * 15:25 sukhe@dns1004: END - running authdns-update * 15:23 sukhe@dns1004: START - running authdns-update * 15:23 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns1006.wikimedia.org,service=authdns-update * 15:23 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns1006.wikimedia.org * 15:23 sukhe@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns1006.wikimedia.org * 15:20 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1056 * 15:20 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:20 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1056~] - vriley@cumin1003" * 15:19 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1056~] - vriley@cumin1003" * 15:19 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns1006.wikimedia.org with OS trixie * 15:19 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns4004.wikimedia.org with OS trixie * 15:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1190.eqiad.wmnet with OS bookworm * 15:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1191.eqiad.wmnet with OS bookworm * 15:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1176.eqiad.wmnet with OS bookworm * 15:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1154.eqiad.wmnet with OS bookworm * 15:16 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica-ulsfo and A:liberica * 15:13 vriley@cumin1003: START - Cookbook sre.dns.netbox * 15:13 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host dns4004.wikimedia.org with OS trixie * 15:12 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1188.eqiad.wmnet with OS bookworm * 15:09 taavi@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318203{{!}}Undeploy WP25EasterEggs (II) (T418134)]] (duration: 08m 56s) * 15:06 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica-magru and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 15:06 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1051.eqiad.wmnet with OS trixie * 15:06 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 15:05 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 15:05 taavi@deploy1003: taavi: Continuing with deployment * 15:04 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-ncredir (exit_code=0) rolling reboot on A:ncredir and A:ncredir * 15:04 taavi@deploy1003: taavi: Backport for [[gerrit:1318203{{!}}Undeploy WP25EasterEggs (II) (T418134)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:03 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1055.eqiad.wmnet with OS trixie * 15:02 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:00 taavi@deploy1003: Started scap sync-world: Backport for [[gerrit:1318203{{!}}Undeploy WP25EasterEggs (II) (T418134)]] * 14:59 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1047.eqiad.wmnet * 14:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1047.eqiad.wmnet * 14:58 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy (exit_code=0) rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 14:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-codfw * 14:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp2001.codfw.wmnet * 14:57 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp2001.codfw.wmnet * 14:57 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica-magru and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 14:57 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:56 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1055 * 14:56 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1055 * 14:55 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns2006.wikimedia.org with reason: host reimage * 14:55 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:55 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1055] - vriley@cumin1003" * 14:55 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1055] - vriley@cumin1003" * 14:54 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1047.eqiad.wmnet * 14:52 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1188.eqiad.wmnet with reason: host reimage * 14:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp2001.codfw.wmnet * 14:51 vriley@cumin1003: START - Cookbook sre.dns.netbox * 14:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp2001.codfw.wmnet * 14:50 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2016.codfw.wmnet * 14:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2016.codfw.wmnet * 14:50 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1189.eqiad.wmnet with OS bookworm * 14:48 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1188.eqiad.wmnet with reason: host reimage * 14:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1051.eqiad.wmnet with reason: host reimage * 14:46 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1182.eqiad.wmnet with OS bookworm * 14:45 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief2002.codfw.wmnet * 14:44 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns1006.wikimedia.org with reason: host reimage * 14:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2016.codfw.wmnet * 14:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1143.eqiad.wmnet with OS bookworm * 14:43 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2016.codfw.wmnet * 14:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2318-2331].codfw.wmnet * 14:43 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2318-2331].codfw.wmnet * 14:42 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching aqs[2002-2012].codfw.wmnet,aqs[1017-1027].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433026|T433026]] - eevans@cumin1003 * 14:41 cgoubert@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326306{{!}}Add placeholder $wmgRedisLockPassword (T366938 T427999)]] (duration: 06m 56s) * 14:41 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief2002.codfw.wmnet * 14:40 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief1002.eqiad.wmnet * 14:39 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1051.eqiad.wmnet with reason: host reimage * 14:38 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns2006.wikimedia.org with reason: host reimage * 14:37 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns1006.wikimedia.org with reason: host reimage * 14:37 cgoubert@deploy1003: cgoubert: Continuing with deployment * 14:36 cgoubert@deploy1003: cgoubert: Backport for [[gerrit:1326306{{!}}Add placeholder $wmgRedisLockPassword (T366938 T427999)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:36 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief1002.eqiad.wmnet * 14:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2318-2331].codfw.wmnet * 14:34 cgoubert@deploy1003: Started scap sync-world: Backport for [[gerrit:1326306{{!}}Add placeholder $wmgRedisLockPassword (T366938 T427999)]] * 14:29 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2318-2331].codfw.wmnet * 14:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2304-2317].codfw.wmnet * 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2304-2317].codfw.wmnet * 14:28 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1054 * 14:27 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1054 * 14:27 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:27 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1054] - vriley@cumin1003" * 14:27 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1054] - vriley@cumin1003" * 14:26 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1189.eqiad.wmnet with reason: host reimage * 14:26 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test2001.codfw.wmnet * 14:25 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test1001.eqiad.wmnet * 14:25 claime: Deploying wmgRedisLockPassword - [[phab:T366938|T366938]] [[phab:T427999|T427999]] * 14:24 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1051.eqiad.wmnet with OS trixie * 14:23 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 14:22 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1182.eqiad.wmnet with reason: host reimage * 14:22 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1051.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:22 vriley@cumin1003: START - Cookbook sre.dns.netbox * 14:22 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test2001.codfw.wmnet * 14:21 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test1001.eqiad.wmnet * 14:21 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 14:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2304-2317].codfw.wmnet * 14:19 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns4004.wikimedia.org with OS trixie * 14:19 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns2006.wikimedia.org with OS trixie * 14:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1143.eqiad.wmnet with reason: host reimage * 14:19 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns1006.wikimedia.org with OS trixie * 14:17 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1189.eqiad.wmnet with reason: host reimage * 14:15 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1182.eqiad.wmnet with reason: host reimage * 14:14 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1143.eqiad.wmnet with reason: host reimage * 14:13 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1051.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:12 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2304-2317].codfw.wmnet * 14:12 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2290-2303].codfw.wmnet * 14:12 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1049.eqiad.wmnet with OS trixie * 14:12 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2290-2303].codfw.wmnet * 14:12 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1051 * 14:11 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1051 * 14:11 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1170.eqiad.wmnet onto db1284.eqiad.wmnet * 14:11 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1170: Pool db1170.eqiad.wmnet in after cloning * 14:09 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 14:09 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:09 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1051] - vriley@cumin1003" * 14:09 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1051] - vriley@cumin1003" * 14:06 klausman@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:04 vriley@cumin1003: START - Cookbook sre.dns.netbox * 14:04 klausman@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2290-2303].codfw.wmnet * 14:02 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1189.eqiad.wmnet with OS bookworm * 14:02 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1188.eqiad.wmnet with OS bookworm * 14:01 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 14:00 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1182.eqiad.wmnet with OS bookworm * 14:00 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1143.eqiad.wmnet with OS bookworm * 13:59 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 13:57 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply * 13:57 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply * 13:56 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply * 13:56 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-ulsfo or A:cp-drmrs and A:cp - 9.2.15 upgrade ([[phab:T434620|T434620]]) * 13:56 cjd91: sudo -i cookbook sre.cdn.roll-upgrade-ats --query 'A:cp-ulsfo or A:cp-drmrs' --task-id [[phab:T434620|T434620]] --reason '9.2.15 upgrade' * 13:56 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply * 13:55 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply * 13:55 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply * 13:54 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 13:54 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 13:54 phuedx@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:53 phuedx@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics-external: apply * 13:52 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1049.eqiad.wmnet with reason: host reimage * 13:51 phuedx@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2290-2303].codfw.wmnet * 13:50 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1047.eqiad.wmnet * 13:50 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2276-2289].codfw.wmnet * 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2276-2289].codfw.wmnet * 13:49 phuedx@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics-external: apply * 13:49 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1046.eqiad.wmnet * 13:49 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1049.eqiad.wmnet with reason: host reimage * 13:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1046.eqiad.wmnet * 13:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-ncredir rolling reboot on A:ncredir and A:ncredir * 13:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 13:46 phuedx@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:44 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics-external: apply * 13:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1046.eqiad.wmnet * 13:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2276-2289].codfw.wmnet * 13:36 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1046.eqiad.wmnet * 13:36 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1045.eqiad.wmnet * 13:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1045.eqiad.wmnet * 13:34 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2276-2289].codfw.wmnet * 13:34 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1049.eqiad.wmnet with OS trixie * 13:34 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2262-2275].codfw.wmnet * 13:34 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2262-2275].codfw.wmnet * 13:32 Lucas_WMDE: UTC afternoon backport+config window done * 13:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1045.eqiad.wmnet * 13:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2262-2275].codfw.wmnet * 13:26 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1049.eqiad.wmnet with OS trixie * 13:26 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1049.eqiad.wmnet with OS trixie * 13:26 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1170: Pool db1170.eqiad.wmnet in after cloning * 13:23 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1049.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 13:23 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1045.eqiad.wmnet * 13:21 atsukoito: manually done sudo -i docker-registryctl --debug delete-tags 'docker-registry.discovery.wmnet/repos/data-engineering/airflow-dags:airflow-3.3.0-py3.11-2026-08-17-*' to remove incorrect tags * 13:19 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311141{{!}}viwiki: Set `noindex,nofollow` for User and User talk (T432311)]] (duration: 11m 11s) * 13:15 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2262-2275].codfw.wmnet * 13:14 lucaswerkmeister-wmde@deploy1003: ndkdd, lucaswerkmeister-wmde: Continuing with deployment * 13:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2248-2261].codfw.wmnet * 13:14 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2248-2261].codfw.wmnet * 13:13 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1049.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 13:12 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1049 * 13:11 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1049 * 13:10 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:10 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1049] - vriley@cumin1003" * 13:10 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1049] - vriley@cumin1003" * 13:10 lucaswerkmeister-wmde@deploy1003: ndkdd, lucaswerkmeister-wmde: Backport for [[gerrit:1311141{{!}}viwiki: Set `noindex,nofollow` for User and User talk (T432311)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1037.eqiad.wmnet * 13:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1037.eqiad.wmnet * 13:08 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1311141{{!}}viwiki: Set `noindex,nofollow` for User and User talk (T432311)]] * 13:06 vriley@cumin1003: START - Cookbook sre.dns.netbox * 13:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2248-2261].codfw.wmnet * 13:00 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1037.eqiad.wmnet * 12:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2248-2261].codfw.wmnet * 12:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2204-2215,2242-2243].codfw.wmnet * 12:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2204-2215,2242-2243].codfw.wmnet * 12:49 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324832{{!}}Migrate $wgFlaggedRevsTags from flaggedrevs.php to ext-FlaggedRevs.php]] (duration: 14m 02s) * 12:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint2001.codfw.wmnet * 12:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2204-2215,2242-2243].codfw.wmnet * 12:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint2001.codfw.wmnet * 12:41 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1037.eqiad.wmnet * 12:40 ladsgroup@deploy1003: Rolling back deployment * 12:37 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1324832{{!}}Migrate $wgFlaggedRevsTags from flaggedrevs.php to ext-FlaggedRevs.php]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:37 seanleong-wmde: Finished populateSitesTable for [bolwiki] ([[[phab:T429955|T429955]]]) * 12:35 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1324832{{!}}Migrate $wgFlaggedRevsTags from flaggedrevs.php to ext-FlaggedRevs.php]] * 12:35 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2204-2215,2242-2243].codfw.wmnet * 12:34 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2190-2203].codfw.wmnet * 12:34 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2190-2203].codfw.wmnet * 12:32 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1028.eqiad.wmnet * 12:32 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1028.eqiad.wmnet * 12:31 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint1001.eqiad.wmnet * 12:30 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1170: Depool db1170.eqiad.wmnet to then clone it to db1284.eqiad.wmnet - marostegui@cumin1003 * 12:28 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint1001.eqiad.wmnet * 12:28 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1170: Depool db1170.eqiad.wmnet to then clone it to db1284.eqiad.wmnet - marostegui@cumin1003 * 12:27 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1170.eqiad.wmnet onto db1284.eqiad.wmnet * 12:27 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db2901.codfw.wmnet * 12:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2190-2203].codfw.wmnet * 12:27 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 12:26 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1028.eqiad.wmnet * 12:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2190-2203].codfw.wmnet * 12:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2172-2179,2184-2189].codfw.wmnet * 12:18 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2172-2179,2184-2189].codfw.wmnet * 12:13 seanleong-wmde@deploy1003: mwscript-k8s job started: foreachwikiindblist wikidataclient extensions/Wikibase/lib/maintenance/populateSitesTable.php --force-protocol https # [[phab:T429955|T429955]] * 12:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2172-2179,2184-2189].codfw.wmnet * 12:09 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1028.eqiad.wmnet * 12:03 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 12:02 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db2901.codfw.wmnet * 12:02 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db2901.codfw.wmnet * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1027.eqiad.wmnet * 12:02 fceratto@cumin1003: END (ERROR) - Cookbook sre.dns.netbox (exit_code=97) * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1027.eqiad.wmnet * 12:02 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 12:02 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db2901.codfw.wmnet * 12:01 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2172-2179,2184-2189].codfw.wmnet * 12:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2158-2171].codfw.wmnet * 12:01 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2158-2171].codfw.wmnet * 11:58 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db1902.eqiad.wmnet * 11:58 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 11:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1027.eqiad.wmnet * 11:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2158-2171].codfw.wmnet * 11:51 jayme: updated calico to v3.30.7 on wikikube eqiad - [[phab:T427400|T427400]] * 11:50 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1027.eqiad.wmnet * 11:45 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2158-2171].codfw.wmnet * 11:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2144-2157].codfw.wmnet * 11:44 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2144-2157].codfw.wmnet * 11:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1058.eqiad.wmnet * 11:43 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1058.eqiad.wmnet * 11:43 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 11:43 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 11:43 marostegui@cumin1003: Removing db1153 from zarcillo [[phab:T434638|T434638]] * 11:42 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1153.eqiad.wmnet * 11:42 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:42 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1153.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 11:42 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1172.eqiad.wmnet onto db1286.eqiad.wmnet * 11:42 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1172: Pool db1172.eqiad.wmnet in after cloning * 11:42 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1153.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 11:41 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1902.eqiad.wmnet * 11:41 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=97) for new host db2901.codfw.wmnet * 11:41 fceratto@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host db2901.codfw.wmnet with OS trixie * 11:38 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'. * 11:38 marostegui@cumin1003: START - Cookbook sre.dns.netbox * 11:37 marostegui@dns1004: END - running authdns-update * 11:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1058.eqiad.wmnet * 11:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2144-2157].codfw.wmnet * 11:35 marostegui@dns1004: START - running authdns-update * 11:32 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1153.eqiad.wmnet * 11:32 marostegui@cumin1003: START - Cookbook sre.mysql.decommission * 11:28 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326249{{!}}ImagePage: move TOC element below file link (T332644)]] (duration: 09m 56s) * 11:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2144-2157].codfw.wmnet * 11:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2130-2143].codfw.wmnet * 11:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2130-2143].codfw.wmnet * 11:26 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1058.eqiad.wmnet * 11:23 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 11:22 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'. * 11:22 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'. * 11:22 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'. * 11:22 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326249{{!}}ImagePage: move TOC element below file link (T332644)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1057.eqiad.wmnet * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1057.eqiad.wmnet * 11:21 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 11:20 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 11:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2130-2143].codfw.wmnet * 11:19 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326249{{!}}ImagePage: move TOC element below file link (T332644)]] * 11:18 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply * 11:17 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply * 11:17 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply * 11:16 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 11:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1057.eqiad.wmnet * 11:15 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 11:14 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 11:13 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 11:13 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db2901.codfw.wmnet with OS trixie * 11:12 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db2901.codfw.wmnet - fceratto@cumin1003" * 11:12 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db2901.codfw.wmnet - fceratto@cumin1003" * 11:12 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db2901.codfw.wmnet on all recursors * 11:12 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db2901.codfw.wmnet on all recursors * 11:12 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:12 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db2901.codfw.wmnet - fceratto@cumin1003" * 11:12 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2130-2143].codfw.wmnet * 11:11 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2107-2115,2124-2129].codfw.wmnet * 11:11 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2107-2115,2124-2129].codfw.wmnet * 11:11 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 11:11 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 11:11 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply * 11:11 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 11:10 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 11:06 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db2901.codfw.wmnet - fceratto@cumin1003" * 11:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2107-2115,2124-2129].codfw.wmnet * 10:57 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1172: Pool db1172.eqiad.wmnet in after cloning * 10:54 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2107-2115,2124-2129].codfw.wmnet * 10:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2078,2087-2095,2102-2106].codfw.wmnet * 10:54 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2078,2087-2095,2102-2106].codfw.wmnet * 10:47 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply * 10:46 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply * 10:46 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply * 10:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2078,2087-2095,2102-2106].codfw.wmnet * 10:45 blake@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply * 10:38 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1057.eqiad.wmnet * 10:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1056.eqiad.wmnet * 10:38 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1056.eqiad.wmnet * 10:37 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2078,2087-2095,2102-2106].codfw.wmnet * 10:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2061-2062,2064-2065,2067-2077].codfw.wmnet * 10:36 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2061-2062,2064-2065,2067-2077].codfw.wmnet * 10:32 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1056.eqiad.wmnet * 10:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2061-2062,2064-2065,2067-2077].codfw.wmnet * 10:25 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:25 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db2901.codfw.wmnet * 10:25 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db1902.eqiad.wmnet * 10:25 fceratto@cumin1003: END (ERROR) - Cookbook sre.dns.netbox (exit_code=97) * 10:24 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:24 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1902.eqiad.wmnet * 10:24 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=97) for new host db1902.eqiad.wmnet * 10:24 fceratto@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host db1902.eqiad.wmnet with OS trixie * 10:24 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=93) for new host db1903.eqiad.wmnet * 10:24 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 10:20 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1056.eqiad.wmnet * 10:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2061-2062,2064-2065,2067-2077].codfw.wmnet * 10:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2038-2039,2041-2042,2044,2046,2049-2051,2055-2060].codfw.wmnet * 10:18 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2038-2039,2041-2042,2044,2046,2049-2051,2055-2060].codfw.wmnet * 10:15 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1055.eqiad.wmnet * 10:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1055.eqiad.wmnet * 10:13 Amir1: mwscript-k8s --dblist=all -- purgeUserOptions.php --login-age 5 uls-preferences ([[phab:T406724|T406724]]) * 10:11 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:10 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1903.eqiad.wmnet on all recursors * 10:10 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1903.eqiad.wmnet on all recursors * 10:10 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2038-2039,2041-2042,2044,2046,2049-2051,2055-2060].codfw.wmnet * 10:10 moritzm: installing unzip security updates * 10:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1055.eqiad.wmnet * 10:08 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=97) for new host db1901.eqiad.wmnet * 10:08 fceratto@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host db1901.eqiad.wmnet with OS trixie * 10:08 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:08 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 10:08 fceratto@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1903.eqiad.wmnet - fceratto@cumin1003" * 10:07 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db1902.eqiad.wmnet with OS trixie * 10:07 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1902.eqiad.wmnet - fceratto@cumin1003" * 10:07 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1902.eqiad.wmnet - fceratto@cumin1003" * 10:04 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply * 10:04 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326227{{!}}Enable desktop/native lazy loading everywhere (T148047)]] (duration: 07m 13s) * 10:03 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1902.eqiad.wmnet on all recursors * 10:03 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1902.eqiad.wmnet on all recursors * 10:03 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:03 blake@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply * 10:01 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1903.eqiad.wmnet - fceratto@cumin1003" * 10:01 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2038-2039,2041-2042,2044,2046,2049-2051,2055-2060].codfw.wmnet * 10:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2002,2005-2006,2011-2015,2017-2018,2033-2037].codfw.wmnet * 10:01 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2002,2005-2006,2011-2015,2017-2018,2033-2037].codfw.wmnet * 10:00 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:00 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 09:59 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 09:59 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1055.eqiad.wmnet * 09:58 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326227{{!}}Enable desktop/native lazy loading everywhere (T148047)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:56 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326227{{!}}Enable desktop/native lazy loading everywhere (T148047)]] * 09:53 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2002,2005-2006,2011-2015,2017-2018,2033-2037].codfw.wmnet * 09:50 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:48 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1903.eqiad.wmnet * 09:48 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:48 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1902.eqiad.wmnet * 09:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 09:44 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2002,2005-2006,2011-2015,2017-2018,2033-2037].codfw.wmnet * 09:43 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-codfw * 09:43 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1174.eqiad.wmnet onto db1288.eqiad.wmnet * 09:42 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1174: Pool db1174.eqiad.wmnet in after cloning * 09:40 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1172: Depool db1172.eqiad.wmnet to then clone it to db1286.eqiad.wmnet - marostegui@cumin1003 * 09:39 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1172: Depool db1172.eqiad.wmnet to then clone it to db1286.eqiad.wmnet - marostegui@cumin1003 * 09:39 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1172.eqiad.wmnet onto db1286.eqiad.wmnet * 09:33 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1175.eqiad.wmnet onto db1289.eqiad.wmnet * 09:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1175: Pool db1175.eqiad.wmnet in after cloning * 09:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1201.eqiad.wmnet onto db1287.eqiad.wmnet * 09:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1201: Pool db1201.eqiad.wmnet in after cloning * 09:28 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db1901.eqiad.wmnet with OS trixie * 09:27 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:27 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:27 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1901.eqiad.wmnet on all recursors * 09:27 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1901.eqiad.wmnet on all recursors * 09:26 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:26 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:26 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:15 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:15 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1901.eqiad.wmnet * 09:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw1001.wikimedia.org with OS trixie * 08:57 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1174: Pool db1174.eqiad.wmnet in after cloning * 08:54 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1054.eqiad.wmnet * 08:54 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1054.eqiad.wmnet * 08:48 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1054.eqiad.wmnet * 08:47 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1175: Pool db1175.eqiad.wmnet in after cloning * 08:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 08:46 marostegui@cumin1003: Removing db1152 from zarcillo [[phab:T434480|T434480]] * 08:46 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1152.eqiad.wmnet * 08:46 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:46 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1152.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 08:46 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1201: Pool db1201.eqiad.wmnet in after cloning * 08:46 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1152.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 08:46 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1054.eqiad.wmnet * 08:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1053.eqiad.wmnet * 08:43 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1053.eqiad.wmnet * 08:42 marostegui@cumin1003: START - Cookbook sre.dns.netbox * 08:38 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage * 08:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1053.eqiad.wmnet * 08:36 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1152.eqiad.wmnet * 08:36 marostegui@cumin1003: START - Cookbook sre.mysql.decommission * 08:35 phuedx: UTC morning backport window done * 08:35 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1053.eqiad.wmnet * 08:34 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1035.eqiad.wmnet * 08:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1035.eqiad.wmnet * 08:34 phuedx@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324725{{!}}EventStreamConfig: Mark product_metrics.web_base and .web_base_with_ip as Test Kitchen streams (T429898 T430322)]], [[gerrit:1313923{{!}}EventStreamConfig: Remove unused web_ui_scroll* streams (T415370)]], [[gerrit:1325546{{!}}EventStreamConfig: Remove Watchlist click stream (T434790)]] (duration: 12m 42s) * 08:33 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1281: Pool back * 08:32 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage * 08:31 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on db2209.codfw.wmnet with reason: Maintenance * 08:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2209: Maintenance needed * 08:30 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2209: Maintenance needed * 08:26 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1035.eqiad.wmnet * 08:26 phuedx@deploy1003: bearloga, phuedx: Continuing with deployment * 08:23 phuedx@deploy1003: bearloga, phuedx: Backport for [[gerrit:1324725{{!}}EventStreamConfig: Mark product_metrics.web_base and .web_base_with_ip as Test Kitchen streams (T429898 T430322)]], [[gerrit:1313923{{!}}EventStreamConfig: Remove unused web_ui_scroll* streams (T415370)]], [[gerrit:1325546{{!}}EventStreamConfig: Remove Watchlist click stream (T434790)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug * 08:21 phuedx@deploy1003: Started scap sync-world: Backport for [[gerrit:1324725{{!}}EventStreamConfig: Mark product_metrics.web_base and .web_base_with_ip as Test Kitchen streams (T429898 T430322)]], [[gerrit:1313923{{!}}EventStreamConfig: Remove unused web_ui_scroll* streams (T415370)]], [[gerrit:1325546{{!}}EventStreamConfig: Remove Watchlist click stream (T434790)]] * 08:19 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw1001.wikimedia.org with OS trixie * 08:18 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1035.eqiad.wmnet * 08:17 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1032.eqiad.wmnet * 08:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1032.eqiad.wmnet * 08:16 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1279: Pool back * 08:15 phuedx@deploy1003: Finished scap sync-world: Backport for [[gerrit:1216721{{!}}viwikivoyage: enable relatedarticle and pop-up (T405724)]] (duration: 39m 12s) * 08:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1032.eqiad.wmnet * 08:09 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1032.eqiad.wmnet * 08:08 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1031.eqiad.wmnet * 08:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1031.eqiad.wmnet * 08:03 godog: switch production to use dumps-nfs.w.o - [[phab:T432212|T432212]] * 08:02 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1031.eqiad.wmnet * 08:02 phuedx@deploy1003: nvdtn19, phuedx: Continuing with deployment * 08:00 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1201: Depool db1201.eqiad.wmnet to then clone it to db1287.eqiad.wmnet - marostegui@cumin1003 * 08:00 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1201: Depool db1201.eqiad.wmnet to then clone it to db1287.eqiad.wmnet - marostegui@cumin1003 * 08:00 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1201.eqiad.wmnet onto db1287.eqiad.wmnet * 08:00 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1057.eqiad.wmnet * 08:00 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:00 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1057.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:59 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1057.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:59 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1275: Pool back * 07:55 filippo@cumin1003: START - Cookbook sre.dns.netbox * 07:55 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1031.eqiad.wmnet * 07:52 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1030.eqiad.wmnet * 07:52 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1030.eqiad.wmnet * 07:52 phuedx@deploy1003: nvdtn19, phuedx: Backport for [[gerrit:1216721{{!}}viwikivoyage: enable relatedarticle and pop-up (T405724)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:51 tappof: bump space for prometheus k8s-aux in eqiad * 07:50 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1057.eqiad.wmnet * 07:48 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1281: Pool back * 07:47 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1281 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96108 and previous config saved to /var/cache/conftool/dbconfig/20260817-074749-marostegui.json * 07:46 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1030.eqiad.wmnet * 07:42 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1030.eqiad.wmnet * 07:41 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1174: Depool db1174.eqiad.wmnet to then clone it to db1288.eqiad.wmnet - marostegui@cumin1003 * 07:41 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1174: Depool db1174.eqiad.wmnet to then clone it to db1288.eqiad.wmnet - marostegui@cumin1003 * 07:41 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1174.eqiad.wmnet onto db1288.eqiad.wmnet * 07:40 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1029.eqiad.wmnet * 07:40 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm2001.wikimedia.org * 07:40 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1029.eqiad.wmnet * 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1056.eqiad.wmnet * 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1056.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:38 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1056.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:36 phuedx@deploy1003: Started scap sync-world: Backport for [[gerrit:1216721{{!}}viwikivoyage: enable relatedarticle and pop-up (T405724)]] * 07:36 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm2001.wikimedia.org * 07:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1279 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96104 and previous config saved to /var/cache/conftool/dbconfig/20260817-073542-marostegui.json * 07:34 filippo@cumin1003: START - Cookbook sre.dns.netbox * 07:34 slyngshede@dns1004: END - running authdns-update * 07:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1029.eqiad.wmnet * 07:33 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm-test1001.wikimedia.org * 07:32 slyngshede@dns1004: START - running authdns-update * 07:31 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1029.eqiad.wmnet * 07:31 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1279: Pool back * 07:30 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1279 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96102 and previous config saved to /var/cache/conftool/dbconfig/20260817-073038-marostegui.json * 07:29 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm-test1001.wikimedia.org * 07:29 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm1001.wikimedia.org * 07:28 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1056.eqiad.wmnet * 07:28 moritzm: extend the disk of ldap-rw1001 by 80G [[phab:T331699|T331699]] * 07:28 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1055.eqiad.wmnet * 07:28 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:28 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1055.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:27 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1055.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:26 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1044.eqiad.wmnet * 07:26 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1044.eqiad.wmnet * 07:25 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm1001.wikimedia.org * 07:24 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1175: Depool db1175.eqiad.wmnet to then clone it to db1289.eqiad.wmnet - marostegui@cumin1003 * 07:24 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1175: Depool db1175.eqiad.wmnet to then clone it to db1289.eqiad.wmnet - marostegui@cumin1003 * 07:24 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1175.eqiad.wmnet onto db1289.eqiad.wmnet * 07:22 filippo@cumin1003: START - Cookbook sre.dns.netbox * 07:20 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1044.eqiad.wmnet * 07:16 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1055.eqiad.wmnet * 07:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1054.eqiad.wmnet * 07:15 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:15 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1054.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:15 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1044.eqiad.wmnet * 07:15 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1054.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:13 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1275: Pool back * 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1275 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96098 and previous config saved to /var/cache/conftool/dbconfig/20260817-071225-marostegui.json * 07:10 filippo@cumin1003: START - Cookbook sre.dns.netbox * 07:05 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1043.eqiad.wmnet * 07:05 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1054.eqiad.wmnet * 07:05 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1051.eqiad.wmnet * 07:05 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:05 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1051.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:05 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1043.eqiad.wmnet * 07:04 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1051.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 06:59 filippo@cumin1003: START - Cookbook sre.dns.netbox * 06:59 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin1001.eqiad.wmnet * 06:59 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1043.eqiad.wmnet * 06:59 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin2001.codfw.wmnet * 06:55 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin2001.codfw.wmnet * 06:55 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1051.eqiad.wmnet * 06:54 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1049.eqiad.wmnet * 06:54 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:54 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1049.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 06:54 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1049.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 06:54 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin1001.eqiad.wmnet * 06:53 moritzm: installing apr-util security updates * 06:52 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1043.eqiad.wmnet * 06:49 filippo@cumin1003: START - Cookbook sre.dns.netbox * 06:41 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1049.eqiad.wmnet * 06:13 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit2003.wikimedia.org * 06:13 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet * 06:07 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet * 06:06 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit2003.wikimedia.org * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 47s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-16 == * 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 01m 03s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-15 == * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 41s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-14 == * 15:38 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-staging-master-eqiad * 15:38 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster1005.eqiad.wmnet * 15:38 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster1005.eqiad.wmnet * 15:35 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sretest2009.codfw.wmnet * 15:33 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster1005.eqiad.wmnet * 15:33 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster1005.eqiad.wmnet * 15:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster1004.eqiad.wmnet * 15:32 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster1004.eqiad.wmnet * 15:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host sretest2009.codfw.wmnet * 15:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster1004.eqiad.wmnet * 15:27 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster1004.eqiad.wmnet * 15:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster1003.eqiad.wmnet * 15:27 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster1003.eqiad.wmnet * 15:24 dancy@deploy1003: Finished scap sync-world: testing (duration: 03m 23s) * 15:22 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster1003.eqiad.wmnet * 15:22 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster1003.eqiad.wmnet * 15:22 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-staging-master-eqiad * 15:20 dancy@deploy1003: Started scap sync-world: testing * 15:20 dancy@deploy1003: Installation of scap version "4.280.2" completed for 3 hosts * 15:18 dancy@deploy1003: Installing scap version "4.280.2" for 3 host(s) * 15:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sretest2006.codfw.wmnet * 14:54 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host sretest2006.codfw.wmnet * 14:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sretest2003.codfw.wmnet * 14:39 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host sretest2003.codfw.wmnet * 13:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt-staging2001.codfw.wmnet * 13:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt-staging2001.codfw.wmnet * 13:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-staging-master-codfw * 13:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster2005.codfw.wmnet * 13:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster2005.codfw.wmnet * 13:05 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox-dev2003.codfw.wmnet * 13:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster2005.codfw.wmnet * 13:04 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster2005.codfw.wmnet * 13:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster2004.codfw.wmnet * 13:04 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster2004.codfw.wmnet * 13:01 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netbox-dev2003.codfw.wmnet * 12:59 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster2004.codfw.wmnet * 12:59 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster2004.codfw.wmnet * 12:59 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster2003.codfw.wmnet * 12:59 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster2003.codfw.wmnet * 12:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster2003.codfw.wmnet * 12:54 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster2003.codfw.wmnet * 12:54 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-staging-master-codfw * 12:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-staging-worker-eqiad * 12:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage1006.eqiad.wmnet * 12:52 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage1006.eqiad.wmnet * 12:46 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage1006.eqiad.wmnet * 12:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw1001.wikimedia.org with OS trixie * 12:45 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage1006.eqiad.wmnet * 12:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage1005.eqiad.wmnet * 12:45 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage1005.eqiad.wmnet * 12:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage1005.eqiad.wmnet * 12:36 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/ratelimit: apply * 12:35 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/ratelimit: apply * 12:35 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:35 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:33 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage1005.eqiad.wmnet * 12:33 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage1004.eqiad.wmnet * 12:33 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage1004.eqiad.wmnet * 12:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage1004.eqiad.wmnet * 12:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage1004.eqiad.wmnet * 12:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage1003.eqiad.wmnet * 12:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage1003.eqiad.wmnet * 12:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage * 12:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage1003.eqiad.wmnet * 12:17 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage * 12:14 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage1003.eqiad.wmnet * 12:14 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-staging-worker-eqiad * 12:03 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw1001.wikimedia.org with OS trixie * 12:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cuminunpriv1001.eqiad.wmnet * 11:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cuminunpriv1001.eqiad.wmnet * 11:27 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1004.wikimedia.org * 11:24 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1153 from dbctl [[phab:T434638|T434638]]', diff saved to https://phabricator.wikimedia.org/P96097 and previous config saved to /var/cache/conftool/dbconfig/20260814-112449-marostegui.json * 11:21 aokoth@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1004.wikimedia.org * 11:20 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 11:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-staging-worker-codfw * 11:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2004.codfw.wmnet * 11:17 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2004.codfw.wmnet * 11:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2004.codfw.wmnet * 11:10 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2004.codfw.wmnet * 11:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2003.codfw.wmnet * 11:10 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2003.codfw.wmnet * 11:03 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-worker1181.eqiad.wmnet with OS bookworm * 11:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2003.codfw.wmnet * 11:03 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2003.codfw.wmnet * 11:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2002.codfw.wmnet * 11:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2002.codfw.wmnet * 10:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1235.eqiad.wmnet with OS bookworm * 10:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock2003.codfw.wmnet * 10:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock2003.codfw.wmnet with OS trixie * 10:56 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2002.codfw.wmnet * 10:54 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1153.eqiad.wmnet with OS bookworm * 10:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2002.codfw.wmnet * 10:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2001.codfw.wmnet * 10:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2001.codfw.wmnet * 10:50 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 10:49 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1187.eqiad.wmnet with OS bookworm * 10:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2001.codfw.wmnet * 10:44 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2001.codfw.wmnet * 10:44 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-staging-worker-codfw * 10:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock2003.codfw.wmnet with reason: host reimage * 10:38 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock2003.codfw.wmnet with reason: host reimage * 10:35 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1153.eqiad.wmnet with reason: host reimage * 10:32 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1235.eqiad.wmnet with reason: host reimage * 10:29 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1187.eqiad.wmnet with reason: host reimage * 10:24 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1153.eqiad.wmnet with reason: host reimage * 10:23 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1235.eqiad.wmnet with reason: host reimage * 10:21 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1187.eqiad.wmnet with reason: host reimage * 10:16 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock2003.codfw.wmnet with OS trixie * 10:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install1005.wikimedia.org * 10:12 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 10:12 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 10:12 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 10:12 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 10:12 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:12 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 10:12 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 10:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install1005.wikimedia.org * 10:07 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1235.eqiad.wmnet with OS bookworm * 10:07 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1187.eqiad.wmnet with OS bookworm * 10:07 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1181.eqiad.wmnet with OS bookworm * 10:07 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1153.eqiad.wmnet with OS bookworm * 10:07 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install2005.wikimedia.org * 10:03 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 10:03 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2003.codfw.wmnet * 10:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1232.eqiad.wmnet with OS bookworm * 10:00 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install2005.wikimedia.org * 10:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install3004.wikimedia.org * 09:58 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 09:58 fceratto@cumin1003: Removing db1151 from zarcillo [[phab:T434538|T434538]] * 09:56 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1231.eqiad.wmnet with OS bookworm * 09:56 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 09:53 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 09:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install3004.wikimedia.org * 09:50 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1152.eqiad.wmnet with OS bookworm * 09:49 Dreamy_Jazz: `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260808000000" --end-timestamp="20260812120000" --sleep="5" --batch-size="50"` for [[phab:T434688|T434688]] * 09:48 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host centrallog2002.codfw.wmnet * 09:48 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install4004.wikimedia.org * 09:42 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1232.eqiad.wmnet with reason: host reimage * 09:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install4004.wikimedia.org * 09:41 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host centrallog2002.codfw.wmnet * 09:39 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install5004.wikimedia.org * 09:36 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1232.eqiad.wmnet with reason: host reimage * 09:36 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'. * 09:34 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'. * 09:33 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1231.eqiad.wmnet with reason: host reimage * 09:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install5004.wikimedia.org * 09:32 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host centrallog1002.eqiad.wmnet * 09:30 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install6003.wikimedia.org * 09:30 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1231.eqiad.wmnet with reason: host reimage * 09:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1152.eqiad.wmnet with reason: host reimage * 09:25 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host centrallog1002.eqiad.wmnet * 09:25 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1152.eqiad.wmnet with reason: host reimage * 09:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install6003.wikimedia.org * 09:22 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1232 * 09:22 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1232 * 09:22 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1232 * 09:22 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1232.eqiad.wmnet 25.53.64.10.in-addr.arpa 5.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:22 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host titan1001.eqiad.wmnet * 09:22 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1232.eqiad.wmnet 25.53.64.10.in-addr.arpa 5.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:22 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:22 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1232 - btullis@cumin1003" * 09:22 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1232 - btullis@cumin1003" * 09:17 btullis@cumin1003: START - Cookbook sre.dns.netbox * 09:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install7002.wikimedia.org * 09:17 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1232 * 09:16 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1231 * 09:16 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1231 * 09:14 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1231 * 09:14 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1231.eqiad.wmnet 24.53.64.10.in-addr.arpa 4.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host titan1001.eqiad.wmnet * 09:14 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1231.eqiad.wmnet 24.53.64.10.in-addr.arpa 4.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:14 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1231 - btullis@cumin1003" * 09:14 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1231 - btullis@cumin1003" * 09:10 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install7002.wikimedia.org * 09:09 btullis@cumin1003: START - Cookbook sre.dns.netbox * 09:08 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1231 * 09:08 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1152 * 09:08 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1152 * 09:06 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1152 * 09:06 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1152.eqiad.wmnet 16.53.64.10.in-addr.arpa 6.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:06 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1152.eqiad.wmnet 16.53.64.10.in-addr.arpa 6.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:06 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:06 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1152 - btullis@cumin1003" * 09:06 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1152 - btullis@cumin1003" * 09:03 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1151.eqiad.wmnet * 09:03 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:03 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1151.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 08:55 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host titan2001.codfw.wmnet * 08:55 btullis@cumin1003: START - Cookbook sre.dns.netbox * 08:54 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1151.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 08:54 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1152 * 08:53 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1232.eqiad.wmnet with OS bookworm * 08:53 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1231.eqiad.wmnet with OS bookworm * 08:53 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1152.eqiad.wmnet with OS bookworm * 08:51 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2209: Pool back * 08:51 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1201.eqiad.wmnet * 08:50 btullis@cumin1003: START - Cookbook sre.hosts.remove-downtime for an-worker1201.eqiad.wmnet * 08:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1229.eqiad.wmnet with OS bookworm * 08:47 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host titan2001.codfw.wmnet * 08:46 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping2004.codfw.wmnet * 08:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ping2004.codfw.wmnet * 08:40 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host an-worker1230.eqiad.wmnet with OS bookworm * 08:40 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 08:34 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1151.eqiad.wmnet * 08:34 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 08:31 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host titan1002.eqiad.wmnet * 08:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1229.eqiad.wmnet with reason: host reimage * 08:25 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host titan1002.eqiad.wmnet * 08:25 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1229.eqiad.wmnet with reason: host reimage * 08:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping1004.eqiad.wmnet * 08:21 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ping1004.eqiad.wmnet * 08:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1230.eqiad.wmnet with reason: host reimage * 08:12 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1230.eqiad.wmnet with reason: host reimage * 08:11 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1229 * 08:11 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1229 * 08:11 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1229 * 08:11 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1229.eqiad.wmnet 22.53.64.10.in-addr.arpa 2.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:11 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1229.eqiad.wmnet 22.53.64.10.in-addr.arpa 2.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:11 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:11 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1229 - btullis@cumin1003" * 08:11 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1229 - btullis@cumin1003" * 08:10 btullis@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1201.eqiad.wmnet with reason: Fixing a disk * 08:07 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host titan2002.codfw.wmnet * 08:05 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2209: Pool back * 08:05 btullis@cumin1003: START - Cookbook sre.dns.netbox * 08:00 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host titan2002.codfw.wmnet * 07:59 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1229 * 07:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1230 * 07:58 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1230 * 07:55 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1230 * 07:55 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1230.eqiad.wmnet 23.53.64.10.in-addr.arpa 3.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:55 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1230.eqiad.wmnet 23.53.64.10.in-addr.arpa 3.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:55 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:55 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1230 - btullis@cumin1003" * 07:55 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1230 - btullis@cumin1003" * 07:51 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host kubestagemaster2005.codfw.wmnet with OS trixie * 07:48 btullis@cumin1003: START - Cookbook sre.dns.netbox * 07:41 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1230 * 07:41 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1229.eqiad.wmnet with OS bookworm * 07:41 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1230.eqiad.wmnet with OS bookworm * 07:39 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1152 from dbctl [[phab:T434480|T434480]]', diff saved to https://phabricator.wikimedia.org/P96090 and previous config saved to /var/cache/conftool/dbconfig/20260814-073941-marostegui.json * 07:29 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on kubestagemaster2005.codfw.wmnet with reason: host reimage * 07:23 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on kubestagemaster2005.codfw.wmnet with reason: host reimage * 07:04 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host kubestagemaster2005.codfw.wmnet with OS trixie * 06:53 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1228.eqiad.wmnet with OS bookworm * 06:44 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1227.eqiad.wmnet with OS bookworm * 06:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1209.eqiad.wmnet with OS bookworm * 06:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1175.eqiad.wmnet with OS bookworm * 06:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1228.eqiad.wmnet with reason: host reimage * 06:27 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1228.eqiad.wmnet with reason: host reimage * 06:25 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1227.eqiad.wmnet with reason: host reimage * 06:21 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1227.eqiad.wmnet with reason: host reimage * 06:18 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1209.eqiad.wmnet with reason: host reimage * 06:14 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1209.eqiad.wmnet with reason: host reimage * 06:14 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1228 * 06:14 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1228 * 06:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1175.eqiad.wmnet with reason: host reimage * 06:12 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1228 * 06:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1228.eqiad.wmnet 20.53.64.10.in-addr.arpa 0.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:12 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1228.eqiad.wmnet 20.53.64.10.in-addr.arpa 0.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1228 - ryankemper@cumin2003" * 06:12 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1228 - ryankemper@cumin2003" * 06:09 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1175.eqiad.wmnet with reason: host reimage * 06:07 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 06:07 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1228 * 06:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1227 * 06:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1227 * 06:06 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1227 * 06:06 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1227.eqiad.wmnet 19.53.64.10.in-addr.arpa 9.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:06 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1227.eqiad.wmnet 19.53.64.10.in-addr.arpa 9.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:06 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:06 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1227 - ryankemper@cumin2003" * 06:06 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1227 - ryankemper@cumin2003" * 06:00 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 06:00 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1227 * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1209 * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1209 * 06:00 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1209 * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1209.eqiad.wmnet 15.53.64.10.in-addr.arpa 5.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:00 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1209.eqiad.wmnet 15.53.64.10.in-addr.arpa 5.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1209 - ryankemper@cumin2003" * 06:00 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1209 - ryankemper@cumin2003" * 05:54 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 05:54 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1209 * 05:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1175 * 05:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1175 * 05:52 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1175 * 05:52 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1175.eqiad.wmnet 17.53.64.10.in-addr.arpa 7.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 05:52 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1175.eqiad.wmnet 17.53.64.10.in-addr.arpa 7.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 05:52 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 05:52 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1175 - ryankemper@cumin2003" * 05:52 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1175 - ryankemper@cumin2003" * 05:49 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1228.eqiad.wmnet with OS bookworm * 05:49 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1227.eqiad.wmnet with OS bookworm * 05:48 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1209.eqiad.wmnet with OS bookworm * 05:47 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 05:47 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1175 * 05:47 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1175.eqiad.wmnet with OS bookworm * 05:09 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host apifeatureusage1001.eqiad.wmnet with OS bookworm * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 03s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:10 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 01:07 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 01:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 01:02 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 00:59 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1226.eqiad.wmnet with OS bookworm * 00:47 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1225.eqiad.wmnet with OS bookworm * 00:41 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1224.eqiad.wmnet with OS bookworm * 00:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1226.eqiad.wmnet with reason: host reimage * 00:31 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1226.eqiad.wmnet with reason: host reimage * 00:28 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1225.eqiad.wmnet with reason: host reimage * 00:25 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1225.eqiad.wmnet with reason: host reimage * 00:19 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1224.eqiad.wmnet with reason: host reimage * 00:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1226 * 00:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1226 * 00:17 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1226 * 00:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1226.eqiad.wmnet 23.36.64.10.in-addr.arpa 3.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:17 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1226.eqiad.wmnet 23.36.64.10.in-addr.arpa 3.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 00:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1226 - ryankemper@cumin2003" * 00:17 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1226 - ryankemper@cumin2003" * 00:16 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1224.eqiad.wmnet with reason: host reimage * 00:12 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 00:12 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1226 * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1225 * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1225 * 00:10 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1225 * 00:10 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1225.eqiad.wmnet 22.36.64.10.in-addr.arpa 2.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:10 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1225.eqiad.wmnet 22.36.64.10.in-addr.arpa 2.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:10 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 00:10 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1225 - ryankemper@cumin2003" * 00:10 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1225 - ryankemper@cumin2003" * 00:03 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 00:02 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1225 * 00:02 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1224 * 00:02 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1224 * 00:00 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1224 * 00:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1224.eqiad.wmnet 21.36.64.10.in-addr.arpa 1.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:00 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1224.eqiad.wmnet 21.36.64.10.in-addr.arpa 1.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 00:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1224 - ryankemper@cumin2003" * 00:00 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1224 - ryankemper@cumin2003" == 2026-08-13 == * 23:54 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1226.eqiad.wmnet with OS bookworm * 23:53 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1225.eqiad.wmnet with OS bookworm * 23:52 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 23:51 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1224 * 23:51 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1224.eqiad.wmnet with OS bookworm * 23:47 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1223.eqiad.wmnet with OS bookworm * 23:29 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1223.eqiad.wmnet with reason: host reimage * 23:24 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1223.eqiad.wmnet with reason: host reimage * 23:19 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325525{{!}}ve.ui.CodeMirror.less: ensure normal font style]] (duration: 11m 40s) * 23:16 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260807000000" --end-timestamp="20260808000000" --sleep="5" --batch-size="50"` for [[phab:T434688|T434688]] * 23:13 musikanimal@deploy1003: musikanimal: Continuing with deployment * 23:11 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1325525{{!}}ve.ui.CodeMirror.less: ensure normal font style]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1223 * 23:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1223 * 23:08 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1325525{{!}}ve.ui.CodeMirror.less: ensure normal font style]] * 23:07 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1223 * 23:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1223.eqiad.wmnet 20.36.64.10.in-addr.arpa 0.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:07 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1223.eqiad.wmnet 20.36.64.10.in-addr.arpa 0.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 23:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1223 - ryankemper@cumin2003" * 23:03 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1223 - ryankemper@cumin2003" * 23:01 sbassett: Deployed security updates for [[phab:T430596|T430596]], [[phab:T120386|T120386]] * 22:55 bking@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host kubestagemaster2005.codfw.wmnet * 22:55 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host kubestagemaster2005.codfw.wmnet with OS bookworm * 22:54 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 22:53 ryankemper@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 22:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on kubestagemaster2005.codfw.wmnet with reason: host reimage * 22:49 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1222.eqiad.wmnet with OS bookworm * 22:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on kubestagemaster2005.codfw.wmnet with reason: host reimage * 22:29 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1222.eqiad.wmnet with reason: host reimage * 22:26 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1222.eqiad.wmnet with reason: host reimage * 22:23 bking@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host aux-k8s-etcd2003.codfw.wmnet * 22:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd2003.codfw.wmnet with OS bookworm * 22:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host kubestagemaster2005.codfw.wmnet with OS bookworm * 22:23 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM kubestagemaster2005.codfw.wmnet - bking@cumin2003" * 22:23 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM kubestagemaster2005.codfw.wmnet - bking@cumin2003" * 22:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) kubestagemaster2005.codfw.wmnet on all recursors * 22:22 bking@cumin2003: START - Cookbook sre.dns.wipe-cache kubestagemaster2005.codfw.wmnet on all recursors * 22:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:22 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM kubestagemaster2005.codfw.wmnet - bking@cumin2003" * 22:22 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM kubestagemaster2005.codfw.wmnet - bking@cumin2003" * 22:17 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 22:13 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1223 * 22:12 bking@cumin2003: START - Cookbook sre.dns.netbox * 22:12 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host kubestagemaster2005.codfw.wmnet * 22:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1222 * 22:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1222 * 22:11 sbassett: Deployed security updates for [[phab:T429244|T429244]], [[phab:T434039|T434039]] * 22:11 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1222 * 22:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1222.eqiad.wmnet 19.36.64.10.in-addr.arpa 9.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:11 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1222.eqiad.wmnet 19.36.64.10.in-addr.arpa 9.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1222 - ryankemper@cumin2003" * 22:07 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1222 - ryankemper@cumin2003" * 22:05 robh@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-wdqs2001.codfw.wmnet with reason: updating firmware * 22:01 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 22:01 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1223.eqiad.wmnet with OS bookworm * 22:01 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1222 * 22:01 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1222.eqiad.wmnet with OS bookworm * 22:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1212.eqiad.wmnet with OS bookworm * 21:54 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324295{{!}}Revert "Lazily reject pre-fix parser-cache entries for noreferrer/noopener links" (T429090)]] (duration: 06m 42s) * 21:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd2003.codfw.wmnet with reason: host reimage * 21:53 bking@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host dse-k8s-etcd2001.codfw.wmnet * 21:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dse-k8s-etcd2001.codfw.wmnet with OS bookworm * 21:50 sbassett@deploy1003: sbassett, kharlan: Continuing with deployment * 21:49 sbassett@deploy1003: sbassett, kharlan: Backport for [[gerrit:1324295{{!}}Revert "Lazily reject pre-fix parser-cache entries for noreferrer/noopener links" (T429090)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:48 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aux-k8s-etcd2003.codfw.wmnet with reason: host reimage * 21:47 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1324295{{!}}Revert "Lazily reject pre-fix parser-cache entries for noreferrer/noopener links" (T429090)]] * 21:42 maryum: Deployed security patch for [[phab:T434549|T434549]] * 21:39 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1212.eqiad.wmnet with reason: host reimage * 21:34 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1212.eqiad.wmnet with reason: host reimage * 21:33 bking@cumin2003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd2003.codfw.wmnet with OS bookworm * 21:32 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM aux-k8s-etcd2003.codfw.wmnet - bking@cumin2003" * 21:32 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM aux-k8s-etcd2003.codfw.wmnet - bking@cumin2003" * 21:32 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) aux-k8s-etcd2003.codfw.wmnet on all recursors * 21:32 bking@cumin2003: START - Cookbook sre.dns.wipe-cache aux-k8s-etcd2003.codfw.wmnet on all recursors * 21:32 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:32 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM aux-k8s-etcd2003.codfw.wmnet - bking@cumin2003" * 21:31 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM aux-k8s-etcd2003.codfw.wmnet - bking@cumin2003" * 21:28 maryum: Deployed security patch for [[phab:T434619|T434619]] * 21:26 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:26 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host aux-k8s-etcd2003.codfw.wmnet * 21:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-etcd2001.codfw.wmnet with reason: host reimage * 21:20 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1212 * 21:20 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1212 * 21:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1212.eqiad.wmnet with OS bookworm * 21:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on dse-k8s-etcd2001.codfw.wmnet with reason: host reimage * 21:04 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325553{{!}}InstrumentConstructiveEdits: exclude mw-reverted as well (T431493)]] (duration: 06m 25s) * 20:59 kemayo@deploy1003: kemayo: Continuing with deployment * 20:59 kemayo@deploy1003: kemayo: Backport for [[gerrit:1325553{{!}}InstrumentConstructiveEdits: exclude mw-reverted as well (T431493)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host dse-k8s-etcd2001.codfw.wmnet with OS bookworm * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 20:57 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 20:57 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1325553{{!}}InstrumentConstructiveEdits: exclude mw-reverted as well (T431493)]] * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-etcd2001.codfw.wmnet on all recursors * 20:57 bking@cumin2003: START - Cookbook sre.dns.wipe-cache dse-k8s-etcd2001.codfw.wmnet on all recursors * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 20:57 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 20:54 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320163{{!}}Add configurable RestTermsOfServiceUrl (T428147)]] (duration: 21m 39s) * 20:53 bking@cumin2003: START - Cookbook sre.dns.netbox * 20:53 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host dse-k8s-etcd2001.codfw.wmnet * 20:50 samtar@deploy1003: samtar, milazg: Continuing with deployment * 20:35 samtar@deploy1003: samtar, milazg: Backport for [[gerrit:1320163{{!}}Add configurable RestTermsOfServiceUrl (T428147)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:33 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1320163{{!}}Add configurable RestTermsOfServiceUrl (T428147)]] * 20:30 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325549{{!}}Deploy PRV to several LC wikis (T423785)]] (duration: 06m 57s) * 20:26 arlolra@deploy1003: arlolra: Continuing with deployment * 20:25 arlolra@deploy1003: arlolra: Backport for [[gerrit:1325549{{!}}Deploy PRV to several LC wikis (T423785)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:24 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 20:23 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1325549{{!}}Deploy PRV to several LC wikis (T423785)]] * 20:21 ariel@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319906{{!}}Remove boilerplate language from wmf-rest and wmf-math API modules (T433736)]] (duration: 13m 54s) * 20:14 ariel@deploy1003: ariel: Continuing with deployment * 20:11 ariel@deploy1003: ariel: Backport for [[gerrit:1319906{{!}}Remove boilerplate language from wmf-rest and wmf-math API modules (T433736)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 ariel@deploy1003: Started scap sync-world: Backport for [[gerrit:1319906{{!}}Remove boilerplate language from wmf-rest and wmf-math API modules (T433736)]] * 19:58 robh@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-wdqs2001.codfw.wmnet with reason: updating firmware * 19:54 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325551{{!}}Render the focused module view as a full-screen page (T433896)]] (duration: 30m 37s) * 19:52 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching aqs[2001,1016]*: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 19:44 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching aqs[2001,1016]*: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 19:42 musikanimal@deploy1003: musikanimal: Continuing with deployment * 19:41 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1325551{{!}}Render the focused module view as a full-screen page (T433896)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:34 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: quash java safepoint logspam - bking@cumin2003 - [[phab:T434685|T434685]] * 19:34 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 19:34 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 19:24 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1325551{{!}}Render the focused module view as a full-screen page (T433896)]] * 19:20 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 19:20 jhancock@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin1003" * 19:18 jhancock@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin1003" * 19:14 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325548{{!}}Enable image lazy loading on desktop in group1 (T148047)]] (duration: 07m 43s) * 19:10 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 19:10 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1325548{{!}}Enable image lazy loading on desktop in group1 (T148047)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:07 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1325548{{!}}Enable image lazy loading on desktop in group1 (T148047)]] * 19:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1166.eqiad.wmnet onto db1280.eqiad.wmnet * 19:03 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 19:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1166: Pool db1166.eqiad.wmnet in after cloning * 18:59 jhancock@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 18:58 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1234.eqiad.wmnet with OS bookworm * 18:52 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 18:51 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:49 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:46 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 18:46 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 18:44 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=dns3004.* * 18:39 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1221.eqiad.wmnet with OS bookworm * 18:36 inflatador: [bking@ganeti2048] ~$ sudo gnt-instance replace-disks -n ganeti2030.codfw.wmnet aux-k8s-worker2002.codfw.wmnet [[phab:T434681|T434681]] * 18:35 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 18:29 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1220.eqiad.wmnet with OS bookworm * 18:29 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1234.eqiad.wmnet with reason: host reimage * 18:27 bking@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host dse-k8s-etcd2001.codfw.wmnet * 18:27 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-etcd2001.codfw.wmnet on all recursors * 18:27 bking@cumin2003: START - Cookbook sre.dns.wipe-cache dse-k8s-etcd2001.codfw.wmnet on all recursors * 18:27 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:27 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 18:27 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 18:25 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1234.eqiad.wmnet with reason: host reimage * 18:25 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 18:19 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1221.eqiad.wmnet with reason: host reimage * 18:17 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1166: Pool db1166.eqiad.wmnet in after cloning * 18:15 brennen@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 18:15 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1221.eqiad.wmnet with reason: host reimage * 18:14 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:12 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:12 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 18:12 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-etcd2001.codfw.wmnet on all recursors * 18:12 bking@cumin2003: START - Cookbook sre.dns.wipe-cache dse-k8s-etcd2001.codfw.wmnet on all recursors * 18:12 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:12 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 18:12 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 18:10 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: quash java safepoint logspam - bking@cumin2003 - [[phab:T434685|T434685]] * 18:10 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1234 * 18:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1234 * 18:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1220.eqiad.wmnet with reason: host reimage * 18:07 brennen: 1.47.0-wmf.15 train status ([[phab:T430834|T430834]]) - no current blockers, rolling to all wikis * 18:07 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1234 * 18:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1234.eqiad.wmnet 10.36.64.10.in-addr.arpa 0.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:07 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:07 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1234.eqiad.wmnet 10.36.64.10.in-addr.arpa 0.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1234 - ryankemper@cumin2003" * 18:07 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1234 - ryankemper@cumin2003" * 18:05 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1220.eqiad.wmnet with reason: host reimage * 18:04 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host dse-k8s-etcd2001.codfw.wmnet * 18:04 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 18:03 dancy@deploy1003: Installation of scap version "4.280.1" completed for 3 hosts * 18:02 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:02 inflatador: bking@dse-k8s-etcd2002 etcdctl member remove $<nowiki>{</nowiki>UUID of dse-k8s-etcd2001<nowiki>}</nowiki> [[phab:T434681|T434681]] [[phab:T434793|T434793]] * 18:01 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 18:01 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1234 * 18:01 dancy@deploy1003: Installing scap version "4.280.1" for 3 host(s) * 18:01 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1221 * 18:01 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1221 * 18:00 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1221 * 18:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1221.eqiad.wmnet 18.36.64.10.in-addr.arpa 8.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:00 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1221.eqiad.wmnet 18.36.64.10.in-addr.arpa 8.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1221 - ryankemper@cumin2003" * 17:59 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:58 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 17:58 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:57 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 17:57 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:56 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1221 - ryankemper@cumin2003" * 17:53 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 17:52 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 17:52 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 17:51 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1221 * 17:51 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1220 * 17:51 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1220 * 17:51 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:51 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:51 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1220 * 17:51 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1220.eqiad.wmnet 11.36.64.10.in-addr.arpa 1.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:51 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1220.eqiad.wmnet 11.36.64.10.in-addr.arpa 1.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:51 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1220 - ryankemper@cumin2003" * 17:50 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1220 - ryankemper@cumin2003" * 17:47 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:47 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 17:46 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1234.eqiad.wmnet with OS bookworm * 17:46 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 17:45 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1221.eqiad.wmnet with OS bookworm * 17:45 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1220 * 17:45 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1220.eqiad.wmnet with OS bookworm * 17:43 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 17:41 bking@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host dse-k8s-etcd2001.codfw.wmnet * 17:41 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host dse-k8s-etcd2001.codfw.wmnet * 17:40 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:40 inflatador: bking@ganeti2048] `sudo gnt-instance remove --force --ignore-failures --shutdown-timeout=0` on non-DRBD VMs [[phab:T434681|T434681]] * 17:40 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:39 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:38 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:36 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:32 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:32 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:28 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 17:26 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:26 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:24 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:23 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:21 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1218.eqiad.wmnet with OS bookworm * 17:20 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:20 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:19 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1179.eqiad.wmnet with OS bookworm * 17:18 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1150.eqiad.wmnet with OS bookworm * 17:18 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns3004.wikimedia.org with OS trixie * 17:15 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:14 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:13 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:12 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:12 swfrench@deploy1003: Finished scap sync-world: Helmfile-only deployment for mediawiki chart bump - [[phab:T427666|T427666]] (duration: 03m 03s) * 17:09 swfrench@deploy1003: Started scap sync-world: Helmfile-only deployment for mediawiki chart bump - [[phab:T427666|T427666]] * 17:02 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:01 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:01 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1218.eqiad.wmnet with reason: host reimage * 17:01 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:01 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:00 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324810{{!}}deployment-info.php: Report dbname and branch for the requested wiki (T434726)]] (duration: 06m 52s) * 16:58 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 16:57 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1179.eqiad.wmnet with reason: host reimage * 16:56 dancy@deploy1003: dancy: Continuing with deployment * 16:56 dancy@deploy1003: dancy: Backport for [[gerrit:1324810{{!}}deployment-info.php: Report dbname and branch for the requested wiki (T434726)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1150.eqiad.wmnet with reason: host reimage * 16:53 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324810{{!}}deployment-info.php: Report dbname and branch for the requested wiki (T434726)]] * 16:51 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1166: Depool db1166.eqiad.wmnet to then clone it to db1280.eqiad.wmnet - cwilliams@cumin1003 * 16:50 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1166: Depool db1166.eqiad.wmnet to then clone it to db1280.eqiad.wmnet - cwilliams@cumin1003 * 16:50 cwilliams@cumin1003: START - Cookbook sre.mysql.clone of db1166.eqiad.wmnet onto db1280.eqiad.wmnet * 16:49 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1218.eqiad.wmnet with reason: host reimage * 16:48 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1179.eqiad.wmnet with reason: host reimage * 16:47 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1150.eqiad.wmnet with reason: host reimage * 16:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.upgrade (exit_code=0) for 1 hosts * 16:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2220: Upgrade of db2220.codfw.wmnet completed * 16:38 swfrench@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 16:38 swfrench-wmf: kubectl delete node kubestagemaster2005.codfw.wmnet - [[phab:T434681|T434681]] * 16:34 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1218 * 16:34 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1218 * 16:34 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1218.eqiad.wmnet with OS bookworm * 16:34 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1179 * 16:34 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1179 * 16:33 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1179.eqiad.wmnet with OS bookworm * 16:32 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1150 * 16:32 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1150 * 16:31 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1150.eqiad.wmnet with OS bookworm * 16:24 dancy@deploy1003: Installation of scap version "4.280.0" completed for 3 hosts * 16:22 dancy@deploy1003: Installing scap version "4.280.0" for 3 host(s) * 16:20 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: quash java safepoint logspam - bking@cumin2003 - [[phab:T434685|T434685]] * 16:14 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns3004.wikimedia.org with reason: host reimage * 16:08 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 16:07 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 16:07 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns3004.wikimedia.org with reason: host reimage * 16:05 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 16:04 swfrench@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 16:00 swfrench@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 15:59 swfrench@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 15:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Upgrade of db2220.codfw.wmnet completed * 15:48 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2220: Upgrading db2220.codfw.wmnet * 15:48 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2220: Upgrading db2220.codfw.wmnet * 15:48 cwilliams@cumin1003: START - Cookbook sre.mysql.upgrade for 1 hosts * 15:46 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns3004.wikimedia.org with OS trixie * 15:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2220 [[phab:T434802|T434802]]', diff saved to https://phabricator.wikimedia.org/P96079 and previous config saved to /var/cache/conftool/dbconfig/20260813-154624-cwilliams.json * 15:45 cdobbins@cumin1003: conftool action : set/pooled=no; selector: name=dns3004.* * 15:44 cjd91: depooling dns3004 to reimage to trixie * 15:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2159 to s7 primary [[phab:T434802|T434802]]', diff saved to https://phabricator.wikimedia.org/P96078 and previous config saved to /var/cache/conftool/dbconfig/20260813-154405-cwilliams.json * 15:43 cezmunsta: Starting s7 codfw failover from db2220 to db2159 - [[phab:T434802|T434802]] * 15:41 cgoubert@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: Dragonfly supernodes reboot (duration: 09m 42s) * 15:41 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dragonfly-supernode2001.codfw.wmnet * 15:39 inflatador: bking@ganeti2048] ~$ sudo gnt-node failover -f --ignore-consistency ganeti2046.codfw.wmnet [[phab:T434681|T434681]] * 15:39 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 15:39 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 15:39 swfrench@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 15:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2159 with weight 0 [[phab:T434802|T434802]]', diff saved to https://phabricator.wikimedia.org/P96077 and previous config saved to /var/cache/conftool/dbconfig/20260813-153806-cwilliams.json * 15:37 swfrench@dns1004: END - running authdns-update * 15:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 30 hosts with reason: Primary switchover s7 [[phab:T434802|T434802]] * 15:37 cgoubert@cumin2003: START - Cookbook sre.hosts.reboot-single for host dragonfly-supernode2001.codfw.wmnet * 15:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dragonfly-supernode1001.eqiad.wmnet * 15:35 swfrench@dns1004: START - running authdns-update * 15:32 cgoubert@cumin2003: START - Cookbook sre.hosts.reboot-single for host dragonfly-supernode1001.eqiad.wmnet * 15:32 cgoubert@deploy1003: Locking from deployment [ALL REPOSITORIES]: Dragonfly supernodes reboot * 15:30 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-master-codfw * 15:30 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2005.codfw.wmnet * 15:30 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2005.codfw.wmnet * 15:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1169.eqiad.wmnet onto db1277.eqiad.wmnet * 15:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1169: Pool db1169.eqiad.wmnet in after cloning * 15:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2005.codfw.wmnet * 15:23 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2005.codfw.wmnet * 15:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2004.codfw.wmnet * 15:23 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2004.codfw.wmnet * 15:18 inflatador: bking@ganeti2048 sudo gnt-node failover -f ganeti2046.codfw.wmnet [[phab:T434681|T434681]] * 15:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2004.codfw.wmnet * 15:17 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2004.codfw.wmnet * 15:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2003.codfw.wmnet * 15:17 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2003.codfw.wmnet * 15:15 cdobbins@dns1004: END - running authdns-update * 15:13 cdobbins@dns1004: START - running authdns-update * 15:10 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1217.eqiad.wmnet with OS bookworm * 15:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2003.codfw.wmnet * 15:10 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2003.codfw.wmnet * 15:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2002.codfw.wmnet * 15:10 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2002.codfw.wmnet * 15:10 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: quash java safepoint logspam - bking@cumin2003 - [[phab:T434685|T434685]] * 15:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1216.eqiad.wmnet with OS bookworm * 15:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2002.codfw.wmnet * 15:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2002.codfw.wmnet * 15:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2001.codfw.wmnet * 15:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2001.codfw.wmnet * 15:01 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1042.eqiad.wmnet * 15:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1042.eqiad.wmnet * 15:00 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325455{{!}}Api: Use correct query when continue prop=categories (T433922)]] (duration: 09m 47s) * 14:59 bking@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host dse-k8s-etcd2004.codfw.wmnet * 14:58 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-etcd2004.codfw.wmnet on all recursors * 14:58 bking@cumin2003: START - Cookbook sre.dns.wipe-cache dse-k8s-etcd2004.codfw.wmnet on all recursors * 14:58 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:58 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM dse-k8s-etcd2004.codfw.wmnet - bking@cumin2003" * 14:58 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM dse-k8s-etcd2004.codfw.wmnet - bking@cumin2003" * 14:55 zabe@deploy1003: zabe: Continuing with deployment * 14:53 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-etcd2004.codfw.wmnet on all recursors * 14:53 bking@cumin2003: START - Cookbook sre.dns.wipe-cache dse-k8s-etcd2004.codfw.wmnet on all recursors * 14:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:53 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2004.codfw.wmnet - bking@cumin2003" * 14:53 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2004.codfw.wmnet - bking@cumin2003" * 14:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2001.codfw.wmnet * 14:52 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2001.codfw.wmnet * 14:52 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-master-codfw * 14:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2001.codfw.wmnet * 14:52 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2001.codfw.wmnet * 14:52 zabe@deploy1003: zabe: Backport for [[gerrit:1325455{{!}}Api: Use correct query when continue prop=categories (T433922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:50 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1325455{{!}}Api: Use correct query when continue prop=categories (T433922)]] * 14:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1217.eqiad.wmnet with reason: host reimage * 14:48 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:48 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host dse-k8s-etcd2004.codfw.wmnet * 14:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-master-eqiad * 14:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1006.eqiad.wmnet * 14:44 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1006.eqiad.wmnet * 14:44 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1216.eqiad.wmnet with reason: host reimage * 14:40 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1217.eqiad.wmnet with reason: host reimage * 14:39 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1169: Pool db1169.eqiad.wmnet in after cloning * 14:39 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1216.eqiad.wmnet with reason: host reimage * 14:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl1005.eqiad.wmnet * 14:32 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl1005.eqiad.wmnet * 14:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1004.eqiad.wmnet * 14:32 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1004.eqiad.wmnet * 14:31 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus2008.codfw.wmnet * 14:31 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor1003.eqiad.wmnet * 14:29 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1042.eqiad.wmnet * 14:27 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor1003.eqiad.wmnet * 14:26 moritzm: installing Django security updates * 14:26 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling reboot on A:wikidough * 14:25 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1217.eqiad.wmnet with OS bookworm * 14:25 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1216.eqiad.wmnet with OS bookworm * 14:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl1004.eqiad.wmnet * 14:24 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl1004.eqiad.wmnet * 14:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1003.eqiad.wmnet * 14:24 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1003.eqiad.wmnet * 14:23 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus2008.codfw.wmnet * 14:23 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus1008.eqiad.wmnet * 14:23 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor-dev2001.codfw.wmnet * 14:22 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus2006.codfw.wmnet * 14:19 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor-dev2001.codfw.wmnet * 14:18 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1042.eqiad.wmnet * 14:17 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.* * 14:16 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl1003.eqiad.wmnet * 14:16 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl1003.eqiad.wmnet * 14:16 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1002.eqiad.wmnet * 14:16 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1002.eqiad.wmnet * 14:15 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus1008.eqiad.wmnet * 14:14 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus1006.eqiad.wmnet * 14:12 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus2006.codfw.wmnet * 14:12 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor2003.codfw.wmnet * 14:11 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus2007.codfw.wmnet * 14:11 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1041.eqiad.wmnet * 14:11 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1041.eqiad.wmnet * 14:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl1002.eqiad.wmnet * 14:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl1002.eqiad.wmnet * 14:09 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-master-eqiad * 14:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on apifeatureusage1001.eqiad.wmnet with reason: host reimage * 14:08 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1215.eqiad.wmnet with OS bookworm * 14:08 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor2003.codfw.wmnet * 14:08 jayme: updated calico to v3.30.7 on wikikube codfw [[phab:T427400|T427400]] * 14:07 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1149.eqiad.wmnet with OS bookworm * 14:06 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'. * 14:06 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1041.eqiad.wmnet * 14:04 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus1006.eqiad.wmnet * 14:03 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus2007.codfw.wmnet * 14:03 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus2005.codfw.wmnet * 14:03 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus1007.eqiad.wmnet * 14:03 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1002.eqiad.wmnet * 14:02 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on apifeatureusage1001.eqiad.wmnet with reason: host reimage * 14:02 cgoubert@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=helm-charts.*,name=eqiad * 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host chartmuseum1001.eqiad.wmnet * 14:00 moritzm: installing libxml2 security updates * 13:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1214.eqiad.wmnet with OS bookworm * 13:59 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns5003.* * 13:59 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1041.eqiad.wmnet * 13:59 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow1002.eqiad.wmnet * 13:58 cmooney@dns3003: END - running authdns-update * 13:57 cgoubert@cumin2003: START - Cookbook sre.hosts.reboot-single for host chartmuseum1001.eqiad.wmnet * 13:57 cgoubert@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=helm-charts.*,name=eqiad * 13:57 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: cloudelastic cluster restart - bking@cumin2003 * 13:57 cgoubert@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=helm-charts.*,name=codfw * 13:56 cmooney@dns3003: START - running authdns-update * 13:56 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'. * 13:56 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns5003.*,service=authdns-update * 13:55 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host chartmuseum2001.codfw.wmnet * 13:55 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus1007.eqiad.wmnet * 13:55 cmooney@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dns5003.wikimedia.org * 13:55 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus1005.eqiad.wmnet * 13:51 cgoubert@cumin2003: START - Cookbook sre.hosts.reboot-single for host chartmuseum2001.codfw.wmnet * 13:51 cgoubert@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=helm-charts.*,name=codfw * 13:51 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus2005.codfw.wmnet * 13:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host apifeatureusage1001.eqiad.wmnet with OS bookworm * 13:50 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus7002.magru.wmnet * 13:50 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts lvs1015.eqiad.wmnet * 13:50 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:50 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1015.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:49 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1015.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:49 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325480{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] (duration: 06m 39s) * 13:47 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1215.eqiad.wmnet with reason: host reimage * 13:46 cmooney@cumin1003: START - Cookbook sre.hosts.reboot-single for host dns5003.wikimedia.org * 13:46 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1003.eqiad.wmnet * 13:45 cmooney@cumin1003: conftool action : set/pooled=no; selector: name=dns5003.* * 13:45 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1040.eqiad.wmnet * 13:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1040.eqiad.wmnet * 13:45 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 13:45 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus1005.eqiad.wmnet * 13:45 stran@deploy1003: stran: Continuing with deployment * 13:44 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus7002.magru.wmnet * 13:44 stran@deploy1003: stran: Backport for [[gerrit:1325480{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:44 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus6002.drmrs.wmnet * 13:43 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260805000000" --end-timestamp="20260806000000" --sleep="5" --batch-size="10"` for [[phab:T434688|T434688]] * 13:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1149.eqiad.wmnet with reason: host reimage * 13:42 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1325480{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] * 13:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow1003.eqiad.wmnet * 13:40 sukhe@cumin1003: START - Cookbook sre.hosts.decommission for hosts lvs1015.eqiad.wmnet * 13:40 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts lvs1014.eqiad.wmnet * 13:40 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:40 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1014.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:40 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1040.eqiad.wmnet * 13:40 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1014.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:39 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1214.eqiad.wmnet with reason: host reimage * 13:38 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2004.codfw.wmnet * 13:38 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus6002.drmrs.wmnet * 13:38 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus5003.eqsin.wmnet * 13:35 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 13:35 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1149.eqiad.wmnet with reason: host reimage * 13:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1215.eqiad.wmnet with reason: host reimage * 13:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow2004.codfw.wmnet * 13:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1214.eqiad.wmnet with reason: host reimage * 13:33 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1040.eqiad.wmnet * 13:31 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus5003.eqsin.wmnet * 13:31 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: cloudelastic cluster restart - bking@cumin2003 * 13:31 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus4003.ulsfo.wmnet * 13:30 sukhe@cumin1003: START - Cookbook sre.hosts.decommission for hosts lvs1014.eqiad.wmnet * 13:30 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts lvs1013.eqiad.wmnet * 13:30 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:30 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1013.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:30 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1013.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:27 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1039.eqiad.wmnet * 13:27 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1039.eqiad.wmnet * 13:26 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1167.eqiad.wmnet onto db1281.eqiad.wmnet * 13:26 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1167: Pool db1167.eqiad.wmnet in after cloning * 13:25 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus4003.ulsfo.wmnet * 13:25 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260802000000" --end-timestamp="20260803000000" --sleep="5" --batch-size="10"` for [[phab:T434688|T434688]] * 13:24 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 13:24 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus3004.esams.wmnet * 13:24 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1274: New host * 13:24 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325476{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] (duration: 07m 19s) * 13:24 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260804000000" --end-timestamp="20260805000000" --sleep="5" --batch-size="10"` for [[phab:T434688|T434688]] * 13:24 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260803000000" --end-timestamp="20260804000000" --sleep="5" --batch-size="10"` for [[phab:T434688|T434688]] * 13:23 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2003.codfw.wmnet * 13:22 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1039.eqiad.wmnet * 13:20 sukhe@cumin1003: START - Cookbook sre.hosts.decommission for hosts lvs1013.eqiad.wmnet * 13:20 stran@deploy1003: stran: Continuing with deployment * 13:19 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow2003.codfw.wmnet * 13:19 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling reboot on A:wikidough * 13:19 stran@deploy1003: stran: Backport for [[gerrit:1325476{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:18 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus3004.esams.wmnet * 13:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1215.eqiad.wmnet with OS bookworm * 13:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1214.eqiad.wmnet with OS bookworm * 13:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1149.eqiad.wmnet with OS bookworm * 13:17 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1325476{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] * 13:11 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324731{{!}}prv: Enable parsoid rendering for 5 wikisource wikis]] (duration: 07m 26s) * 13:11 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1039.eqiad.wmnet * 13:11 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1036.eqiad.wmnet * 13:11 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1036.eqiad.wmnet * 13:07 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow3004.esams.wmnet * 13:07 jgiannelos@deploy1003: jgiannelos: Continuing with deployment * 13:06 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1233.eqiad.wmnet with OS bookworm * 13:06 jgiannelos@deploy1003: jgiannelos: Backport for [[gerrit:1324731{{!}}prv: Enable parsoid rendering for 5 wikisource wikis]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:04 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1324731{{!}}prv: Enable parsoid rendering for 5 wikisource wikis]] * 13:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow3004.esams.wmnet * 13:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1036.eqiad.wmnet * 13:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1210.eqiad.wmnet with OS bookworm * 13:01 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1036.eqiad.wmnet * 13:00 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1052.eqiad.wmnet * 13:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1052.eqiad.wmnet * 12:58 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow4003.ulsfo.wmnet * 12:55 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1211.eqiad.wmnet with OS bookworm * 12:55 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1052.eqiad.wmnet * 12:52 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow4003.ulsfo.wmnet * 12:51 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1052.eqiad.wmnet * 12:49 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1169: Depool db1169.eqiad.wmnet to then clone it to db1277.eqiad.wmnet - cwilliams@cumin1003 * 12:46 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1169: Depool db1169.eqiad.wmnet to then clone it to db1277.eqiad.wmnet - cwilliams@cumin1003 * 12:46 cwilliams@cumin1003: START - Cookbook sre.mysql.clone of db1169.eqiad.wmnet onto db1277.eqiad.wmnet * 12:45 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:45 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns record for deleted IP reservations eqsin lvs vlan ints - cmooney@cumin1003" * 12:44 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns record for deleted IP reservations eqsin lvs vlan ints - cmooney@cumin1003" * 12:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1051.eqiad.wmnet * 12:43 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1051.eqiad.wmnet * 12:42 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1210.eqiad.wmnet with reason: host reimage * 12:41 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1167: Pool db1167.eqiad.wmnet in after cloning * 12:40 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 12:39 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1274: New host * 12:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Added db1274', diff saved to https://phabricator.wikimedia.org/P96062 and previous config saved to /var/cache/conftool/dbconfig/20260813-123907-cwilliams.json * 12:38 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast7002.wikimedia.org * 12:38 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1233.eqiad.wmnet with reason: host reimage * 12:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1051.eqiad.wmnet * 12:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1211.eqiad.wmnet with reason: host reimage * 12:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1233.eqiad.wmnet with reason: host reimage * 12:32 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1051.eqiad.wmnet * 12:32 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast7002.wikimedia.org * 12:32 marostegui: Drop SecurePoll tables from closed wikis [[phab:T423128|T423128]] * 12:32 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow5003.eqsin.wmnet * 12:31 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1050.eqiad.wmnet * 12:31 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1050.eqiad.wmnet * 12:29 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1210.eqiad.wmnet with reason: host reimage * 12:29 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1211.eqiad.wmnet with reason: host reimage * 12:29 moritzm: installing Wireshark security updates * 12:26 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow5003.eqsin.wmnet * 12:26 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1050.eqiad.wmnet * 12:24 cmooney@dns3003: END - running authdns-update * 12:21 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow6001.drmrs.wmnet * 12:21 cmooney@dns3003: START - running authdns-update * 12:20 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1050.eqiad.wmnet * 12:20 cgoubert@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host rdb-lock2003.codfw.wmnet * 12:20 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:20 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update netbox dns entries for expanded public1-603-eqsin subnet - cmooney@cumin1003" * 12:20 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update netbox dns entries for expanded public1-603-eqsin subnet - cmooney@cumin1003" * 12:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 12:20 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 12:19 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1049.eqiad.wmnet * 12:19 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1049.eqiad.wmnet * 12:19 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 12:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host moss-be1003.eqiad.wmnet * 12:17 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow6001.drmrs.wmnet * 12:17 moritzm: installin curl security updates * 12:15 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 12:15 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1233.eqiad.wmnet with OS bookworm * 12:15 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1210.eqiad.wmnet with OS bookworm * 12:15 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1211.eqiad.wmnet with OS bookworm * 12:13 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1049.eqiad.wmnet * 12:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow7002.magru.wmnet * 12:12 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-worker1178.eqiad.wmnet with OS bookworm * 12:10 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host moss-be1003.eqiad.wmnet * 12:10 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 12:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be1006.eqiad.wmnet * 12:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 12:10 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 12:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 12:10 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 12:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow7002.magru.wmnet * 12:08 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1049.eqiad.wmnet * 12:05 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 12:05 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2003.codfw.wmnet * 12:04 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'. * 12:04 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:03 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be1006.eqiad.wmnet * 12:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be1005.eqiad.wmnet * 12:01 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1038.eqiad.wmnet * 12:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1038.eqiad.wmnet * 12:01 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1213.eqiad.wmnet with OS bookworm * 11:56 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be1005.eqiad.wmnet * 11:55 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be1004.eqiad.wmnet * 11:54 cgoubert@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host rdb-lock2003.codfw.wmnet * 11:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 11:53 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 11:53 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:53 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 11:53 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 11:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1038.eqiad.wmnet * 11:49 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be1004.eqiad.wmnet * 11:44 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:44 moritzm: installing Linux 5.10.262 on Bullseye hosts * 11:41 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1038.eqiad.wmnet * 11:40 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 11:40 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 11:40 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 11:40 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:40 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 11:40 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 11:38 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1213.eqiad.wmnet with reason: host reimage * 11:36 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1167: Depool db1167.eqiad.wmnet to then clone it to db1281.eqiad.wmnet - marostegui@cumin1003 * 11:35 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 11:35 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2003.codfw.wmnet * 11:35 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1167: Depool db1167.eqiad.wmnet to then clone it to db1281.eqiad.wmnet - marostegui@cumin1003 * 11:35 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1167.eqiad.wmnet onto db1281.eqiad.wmnet * 11:34 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 22 hosts with reason: Cloning * 11:34 moritzm: remove ganeti3005 from esams03 cluster, hardware issues [[phab:T434646|T434646]] * 11:32 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1213.eqiad.wmnet with reason: host reimage * 11:28 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1034.eqiad.wmnet * 11:28 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1034.eqiad.wmnet * 11:22 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1034.eqiad.wmnet * 11:19 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1034.eqiad.wmnet * 11:17 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1213.eqiad.wmnet with OS bookworm * 11:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1165.eqiad.wmnet onto db1279.eqiad.wmnet * 11:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1165: Pool db1165.eqiad.wmnet in after cloning * 11:07 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply * 10:57 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply * 10:54 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply * 10:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1274.eqiad.wmnet with reason: Enabling notifications and pooling * 10:45 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1033.eqiad.wmnet * 10:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1033.eqiad.wmnet * 10:44 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply * 10:43 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'. * 10:42 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'. * 10:42 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'. * 10:40 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db[1216,1225,1239-1240].eqiad.wmnet with reason: reboot * 10:39 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1033.eqiad.wmnet * 10:38 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1151 from dbctl [[phab:T434538|T434538]]', diff saved to https://phabricator.wikimedia.org/P96055 and previous config saved to /var/cache/conftool/dbconfig/20260813-103828-marostegui.json * 10:35 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1033.eqiad.wmnet * 10:27 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1165: Pool db1165.eqiad.wmnet in after cloning * 10:24 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1151.eqiad.wmnet with OS bookworm * 10:15 moritzm: installing bind9 security updates (client-side tools/libs only) * 10:07 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-debug: apply * 10:06 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-debug: apply * 10:02 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 7 hosts with reason: reboot * 10:01 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-debug: apply * 10:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.upgrade (exit_code=0) for 1 hosts * 10:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2214: Upgrade of db2214.codfw.wmnet completed * 10:01 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-debug: apply * 10:00 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-debug: apply * 10:00 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-debug: apply * 09:59 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.decommission (exit_code=99) * 09:59 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 09:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1151.eqiad.wmnet with reason: host reimage * 09:59 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 8 hosts * 09:59 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 8 hosts * 09:55 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1151.eqiad.wmnet with reason: host reimage * 09:46 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260801000000" --end-timestamp="20260802000000" --sleep="3" --batch-size="5"` for [[phab:T434688|T434688]] * 09:45 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 8 hosts with reason: reboot * 09:45 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 22 hosts with reason: Cloning * 09:43 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1165: Depool db1165.eqiad.wmnet to then clone it to db1279.eqiad.wmnet - marostegui@cumin1003 * 09:42 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1165: Depool db1165.eqiad.wmnet to then clone it to db1279.eqiad.wmnet - marostegui@cumin1003 * 09:42 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1165.eqiad.wmnet onto db1279.eqiad.wmnet * 09:41 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=testwiki --start-timestamp="20260311000000" --end-timestamp="20260805000000" --sleep="5" --batch-size="2"` for [[phab:T434688|T434688]] * 09:40 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for backupmon1001.eqiad.wmnet * 09:40 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for backupmon1001.eqiad.wmnet * 09:40 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1151 * 09:40 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1151 * 09:37 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325408{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]], [[gerrit:1325407{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]] (duration: 06m 57s) * 09:36 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1151 * 09:36 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1151.eqiad.wmnet 13.36.64.10.in-addr.arpa 3.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:36 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1151.eqiad.wmnet 13.36.64.10.in-addr.arpa 3.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:36 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:36 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1151 - btullis@cumin1003" * 09:36 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on backupmon1001.eqiad.wmnet with reason: reboot * 09:36 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1151 - btullis@cumin1003" * 09:35 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host moss-be2003.codfw.wmnet * 09:33 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 7 hosts * 09:33 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 7 hosts * 09:33 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 09:32 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1325408{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]], [[gerrit:1325407{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1161.eqiad.wmnet onto db1275.eqiad.wmnet * 09:31 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 09:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1161: Pool db1161.eqiad.wmnet in after cloning * 09:30 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1325408{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]], [[gerrit:1325407{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]] * 09:29 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 09:27 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 09:27 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host moss-be2003.codfw.wmnet * 09:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be2006.codfw.wmnet * 09:25 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2004.codfw.wmnet * 09:22 hashar@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 09:21 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be2006.codfw.wmnet * 09:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be2005.codfw.wmnet * 09:19 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2004.codfw.wmnet * 09:19 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2003.codfw.wmnet * 09:18 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 7 hosts with reason: reboot * 09:18 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 7 hosts * 09:18 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 7 hosts * 09:16 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2214: Upgrade of db2214.codfw.wmnet completed * 09:14 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be2005.codfw.wmnet * 09:13 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be2004.codfw.wmnet * 09:12 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2003.codfw.wmnet * 09:12 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2002.codfw.wmnet * 09:09 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2214: Upgrading db2214.codfw.wmnet * 09:09 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2214: Upgrading db2214.codfw.wmnet * 09:09 cwilliams@cumin1003: START - Cookbook sre.mysql.upgrade for 1 hosts * 09:07 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be2004.codfw.wmnet * 09:06 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2002.codfw.wmnet * 09:05 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1004.eqiad.wmnet * 09:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-cluster (exit_code=0) * 09:03 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 7 hosts with reason: reboot * 09:01 btullis@cumin1003: START - Cookbook sre.dns.netbox * 09:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2214 [[phab:T434754|T434754]]', diff saved to https://phabricator.wikimedia.org/P96045 and previous config saved to /var/cache/conftool/dbconfig/20260813-090001-cwilliams.json * 08:59 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1004.eqiad.wmnet * 08:59 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1003.eqiad.wmnet * 08:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2229 to s6 primary [[phab:T434754|T434754]]', diff saved to https://phabricator.wikimedia.org/P96044 and previous config saved to /var/cache/conftool/dbconfig/20260813-085752-cwilliams.json * 08:57 cezmunsta: Starting s6 codfw failover from db2214 to db2229 - [[phab:T434754|T434754]] * 08:54 hashar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325416{{!}}Revert "REST: Enable `GET /lexemes/<nowiki>{</nowiki>lexeme_id<nowiki>}</nowiki>` by default" (T434712)]] (duration: 07m 22s) * 08:53 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1003.eqiad.wmnet * 08:53 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1002.eqiad.wmnet * 08:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2229 with weight 0 [[phab:T434754|T434754]]', diff saved to https://phabricator.wikimedia.org/P96043 and previous config saved to /var/cache/conftool/dbconfig/20260813-085151-cwilliams.json * 08:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 23 hosts with reason: Primary switchover s6 [[phab:T434754|T434754]] * 08:51 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts2002.codfw.wmnet * 08:50 hashar@deploy1003: hashar: Continuing with deployment * 08:49 hashar@deploy1003: hashar: Backport for [[gerrit:1325416{{!}}Revert "REST: Enable `GET /lexemes/<nowiki>{</nowiki>lexeme_id<nowiki>}</nowiki>` by default" (T434712)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:47 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet * 08:47 hashar@deploy1003: Started scap sync-world: Backport for [[gerrit:1325416{{!}}Revert "REST: Enable `GET /lexemes/<nowiki>{</nowiki>lexeme_id<nowiki>}</nowiki>` by default" (T434712)]] * 08:47 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1002.eqiad.wmnet * 08:46 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1161: Pool db1161.eqiad.wmnet in after cloning * 08:46 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host stewards2001.codfw.wmnet * 08:45 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit1003.wikimedia.org * 08:45 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host stewards1001.eqiad.wmnet * 08:45 Emperor: roll-restart apus frontends in codfw * 08:45 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-cluster * 08:44 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts2002.codfw.wmnet * 08:44 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host doc2003.codfw.wmnet * 08:43 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet * 08:43 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host phab2003.codfw.wmnet * 08:42 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host stewards2001.codfw.wmnet * 08:42 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host etherpad2002.codfw.wmnet * 08:41 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host stewards1001.eqiad.wmnet * 08:41 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1273: Pool in s7 * 08:41 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1003.wikimedia.org * 08:40 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host doc2003.codfw.wmnet * 08:40 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit1003.wikimedia.org * 08:39 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host etherpad1004.eqiad.wmnet * 08:39 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host doc1004.eqiad.wmnet * 08:39 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit2002.wikimedia.org * 08:38 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host etherpad2002.codfw.wmnet * 08:37 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host phab2003.codfw.wmnet * 08:36 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lists1004.wikimedia.org * 08:35 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host etherpad1004.eqiad.wmnet * 08:35 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host doc1004.eqiad.wmnet * 08:34 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-cluster (exit_code=0) * 08:34 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1003.wikimedia.org * 08:34 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host planet2003.codfw.wmnet * 08:34 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2003.wikimedia.org * 08:33 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host planet1003.eqiad.wmnet * 08:33 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit2002.wikimedia.org * 08:32 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 22 hosts with reason: Cloning * 08:32 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2002.wikimedia.org * 08:31 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aphlict1002.eqiad.wmnet * 08:30 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host planet2003.codfw.wmnet * 08:29 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host planet1003.eqiad.wmnet * 08:29 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host lists1004.wikimedia.org * 08:28 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lists2001.wikimedia.org * 08:28 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2003.wikimedia.org * 08:27 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host aphlict1002.eqiad.wmnet * 08:27 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aphlict2001.codfw.wmnet * 08:26 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2002.wikimedia.org * 08:23 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host aphlict2001.codfw.wmnet * 08:23 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1151 * 08:22 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1151.eqiad.wmnet with OS bookworm * 08:22 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host lists2001.wikimedia.org * 08:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1161: Depool db1161.eqiad.wmnet to then clone it to db1275.eqiad.wmnet - marostegui@cumin1003 * 08:20 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1161: Depool db1161.eqiad.wmnet to then clone it to db1275.eqiad.wmnet - marostegui@cumin1003 * 08:20 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1161.eqiad.wmnet onto db1275.eqiad.wmnet * 08:15 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-cluster * 08:15 Emperor: roll-restart apus frontends in eqiad * 07:56 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1273: Pool in s7 * 07:56 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1273 to dbctl [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P96035 and previous config saved to /var/cache/conftool/dbconfig/20260813-075611-marostegui.json * 07:38 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db1273.eqiad.wmnet with reason: Reboot * 07:31 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.sanitize-wiki (exit_code=97) Managing sanitization for wikis testwiki in section s3 * 07:24 marostegui@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis testwiki in section s3 * 07:19 jayme@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on kubestagemaster2005.codfw.wmnet with reason: downtime because of hardware failure and no DRBD * 05:42 arnaudb@dns1006: END - running authdns-update * 05:40 arnaudb@dns1006: START - running authdns-update * 05:27 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 04:06 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 02:29 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1151 * 02:29 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1151 * 02:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1219.eqiad.wmnet with OS bookworm * 02:18 ryankemper: [[phab:T434494|T434494]] `ryankemper@deploy1003:~$ echo 'https://stats.wikimedia.org/' {{!}} mwscript-k8s --attach -- purgeList.php` (default page got cached during yesterday's `an-web1001` reimage) * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 46s) * 02:03 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1219.eqiad.wmnet with reason: host reimage * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 02:00 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1219.eqiad.wmnet with reason: host reimage * 01:46 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1219.eqiad.wmnet with OS bookworm * 01:03 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1142.eqiad.wmnet with OS bookworm * 00:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1142.eqiad.wmnet with reason: host reimage * 00:34 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1142.eqiad.wmnet with reason: host reimage * 00:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1142.eqiad.wmnet with OS bookworm * 00:16 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1178 * 00:16 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1178 == 2026-08-12 == * 23:18 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324409{{!}}Remove $wmg = $wg hacks in Collection (T119117)]] (duration: 06m 43s) * 23:14 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 23:13 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1324409{{!}}Remove $wmg = $wg hacks in Collection (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:11 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1324409{{!}}Remove $wmg = $wg hacks in Collection (T119117)]] * 22:59 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 22:49 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260801000000" --end-timestamp="20260802000000" --sleep=2 --batch-size=10` * 22:45 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=testwiki --start-timestamp="20200801010101" --end-timestamp="20260816010101" --sleep=15 --batch-size=5` * 22:40 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=testwiki --start-timestamp="20200101010101" --end-timestamp="20260816010101" --sleep=60` * 22:35 Dreamy_Jazz: Running `mwscript WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260101000000" --end-timestamp="20260102000000" --sleep=10` * 22:21 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324817{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]], [[gerrit:1324818{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]] (duration: 45m 29s) * 22:17 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 21:59 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 21:40 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1324817{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]], [[gerrit:1324818{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:39 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 21:36 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1324817{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]], [[gerrit:1324818{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]] * 21:32 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:31 vriley@cumin1003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:30 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:30 vriley@cumin1003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:17 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:14 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:14 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:11 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:10 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-codfw: Set storage compatability to NONE — [[phab:T433028|T433028]] - eevans@cumin1003 * 21:10 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1006 * 21:09 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1006 * 21:05 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324277{{!}}Improve Math preference labels for SVG/MathJax/MathML (T433891)]] (duration: 31m 42s) * 20:58 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1178.eqiad.wmnet with OS bookworm * 20:54 krinkle@deploy1003: krinkle: Continuing with deployment * 20:51 krinkle@deploy1003: krinkle: Backport for [[gerrit:1324277{{!}}Improve Math preference labels for SVG/MathJax/MathML (T433891)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:39 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 20:38 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:38 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp5022.eqsin.wmnet with OS trixie * 20:38 cdobbins@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - cdobbins@cumin1003" * 20:37 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:36 cdobbins@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - cdobbins@cumin1003" * 20:34 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1178.eqiad.wmnet with reason: host reimage * 20:34 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1324277{{!}}Improve Math preference labels for SVG/MathJax/MathML (T433891)]] * 20:33 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 20:28 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1178.eqiad.wmnet with reason: host reimage * 20:13 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1178.eqiad.wmnet with OS bookworm * 20:11 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-worker1178.eqiad.wmnet with OS bookworm * 20:11 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1178.eqiad.wmnet with OS bookworm * 20:09 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-codfw: Set storage compatability to NONE — [[phab:T433028|T433028]] - eevans@cumin1003 * 20:09 Dreamy_Jazz: Evening UTC backport window done * 20:08 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324794{{!}}WikimediaAntiAbuse: Enable logging channel (T431292)]] (duration: 06m 48s) * 20:08 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage * 20:05 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage * 20:04 dreamyjazz@deploy1003: kharlan, dreamyjazz: Continuing with deployment * 20:04 dreamyjazz@deploy1003: kharlan, dreamyjazz: Backport for [[gerrit:1324794{{!}}WikimediaAntiAbuse: Enable logging channel (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:01 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1324794{{!}}WikimediaAntiAbuse: Enable logging channel (T431292)]] * 19:55 brennen@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 19:47 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 19:35 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 19:35 cdobbins@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cp5022.eqsin.wmnet with OS trixie * 19:32 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:30 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:29 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:26 vriley@cumin1003: START - Cookbook sre.dns.netbox * 19:23 brennen@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 19:19 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-eqiad: Set storage compatability to NONE — [[phab:T433028|T433028]] - eevans@cumin1003 * 19:10 brennen@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 19:09 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 19:09 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 19:00 Amir1: data migrated on wikishared ([[phab:T426102|T426102]]) * 18:57 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321224{{!}}Rename ce_worklist_articles table to ce_invitation_list_articles (T426102)]] (duration: 06m 50s) * 18:53 ladsgroup@deploy1003: ladsgroup, daimona: Continuing with deployment * 18:53 ladsgroup@deploy1003: ladsgroup, daimona: Backport for [[gerrit:1321224{{!}}Rename ce_worklist_articles table to ce_invitation_list_articles (T426102)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:51 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1321224{{!}}Rename ce_worklist_articles table to ce_invitation_list_articles (T426102)]] * 18:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2192: Security update * 18:31 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 18:25 Amir1: ce_invitation_list_articles created as empty on wikishared ([[phab:T426102|T426102]]) * 18:21 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 18:21 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 18:18 Amir1: migrated testwiki entries from ce_worklist_articles to ce_invitation_list_articles ([[phab:T426102|T426102]]) * 18:18 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-eqiad: Set storage compatability to NONE — [[phab:T433028|T433028]] - eevans@cumin1003 * 18:11 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 18:08 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling reboot on A:durum-eqsin and A:durum * 18:07 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Set storage compatability to UPGRADING — [[phab:T433028|T433028]] - eevans@cumin1003 * 18:05 jhancock@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022'] * 17:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2192: Security update * 17:55 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:55 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum-eqsin and A:durum * 17:53 jhancock@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['cp5022'] * 17:47 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:47 jhancock@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['cp5022'] * 17:42 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:41 jhancock@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['cp5022'] * 17:36 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:35 jhancock@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022'] * 17:31 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-magru and not (P<nowiki>{</nowiki>cp7001*<nowiki>}</nowiki> or P<nowiki>{</nowiki>cp7009*<nowiki>}</nowiki>) and A:cp - 9.2.15 upgrade ([[phab:T434620|T434620]]) * 17:28 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:28 jhancock@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['cp5022'] * 17:22 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2192.codfw.wmnet with reason: Maintenance * 17:11 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host apifeatureusage2001.codfw.wmnet with OS bookworm * 17:04 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: apply * 17:03 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-main: apply * 17:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2192 [[phab:T434635|T434635]]', diff saved to https://phabricator.wikimedia.org/P96030 and previous config saved to /var/cache/conftool/dbconfig/20260812-170338-cwilliams.json * 17:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2213 to s5 primary [[phab:T434635|T434635]]', diff saved to https://phabricator.wikimedia.org/P96029 and previous config saved to /var/cache/conftool/dbconfig/20260812-170152-cwilliams.json * 17:01 cezmunsta: Starting s5 codfw failover from db2192 to db2213 - [[phab:T434635|T434635]] * 16:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2213 with weight 0 [[phab:T434635|T434635]]', diff saved to https://phabricator.wikimedia.org/P96028 and previous config saved to /var/cache/conftool/dbconfig/20260812-165544-cwilliams.json * 16:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 27 hosts with reason: Primary switchover s5 [[phab:T434635|T434635]] * 16:53 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: apply * 16:52 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-main: apply * 16:44 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-main: apply * 16:44 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-main: apply * 16:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1272: New host * 16:40 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324765{{!}}WikimediaAntiAbuse: Enable PersonalInfoFlagNotifications (T431292)]] (duration: 07m 02s) * 16:40 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-logging-external: apply * 16:39 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-logging-external: apply * 16:38 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-logging-external: apply * 16:37 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-logging-external: apply * 16:36 kharlan@deploy1003: kharlan: Continuing with deployment * 16:35 kharlan@deploy1003: kharlan: Backport for [[gerrit:1324765{{!}}WikimediaAntiAbuse: Enable PersonalInfoFlagNotifications (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:33 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1324765{{!}}WikimediaAntiAbuse: Enable PersonalInfoFlagNotifications (T431292)]] * 16:27 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-logging-external: apply * 16:27 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-logging-external: apply * 16:18 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 16:18 jhancock@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022'] * 16:17 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 16:16 jhancock@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['cp5022'] * 16:15 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 16:12 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Set storage compatability to UPGRADING — [[phab:T433028|T433028]] - eevans@cumin1003 * 16:10 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: apply * 16:10 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: apply * 16:08 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: apply * 16:08 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: apply * 16:08 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics: apply * 16:07 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics: apply * 16:02 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-magru and not (P<nowiki>{</nowiki>cp7001*<nowiki>}</nowiki> or P<nowiki>{</nowiki>cp7009*<nowiki>}</nowiki>) and A:cp - 9.2.15 upgrade ([[phab:T434620|T434620]]) * 15:57 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1272: New host * 15:52 jmm@cumin2003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti2046.codfw.wmnet * 15:52 jmm@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host ganeti2046.codfw.wmnet * 15:42 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:37 mforns@deploy1003: Finished deploy [analytics/refinery@49c336c] (thin): Regular analytics weekly train THIN [analytics/refinery@49c336cd] (duration: 01m 59s) * 15:37 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324733{{!}}WikimediaAntiAbuse: Enable personal info tag display on enwiki (T431292)]] (duration: 08m 12s) * 15:35 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:35 mforns@deploy1003: Started deploy [analytics/refinery@49c336c] (thin): Regular analytics weekly train THIN [analytics/refinery@49c336cd] * 15:34 mforns@deploy1003: Finished deploy [analytics/refinery@49c336c]: Regular analytics weekly train [analytics/refinery@49c336cd] (duration: 04m 20s) * 15:33 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[2024,1031]*.wmnet: Set storage compatability to UPGRADING — [[phab:T433028|T433028]] - eevans@cumin1003 * 15:33 dreamyjazz@deploy1003: kharlan, dreamyjazz: Continuing with deployment * 15:31 dreamyjazz@deploy1003: kharlan, dreamyjazz: Backport for [[gerrit:1324733{{!}}WikimediaAntiAbuse: Enable personal info tag display on enwiki (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:30 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitize-wiki (exit_code=99) Checking sanitization for wikis testwiki in section s3 * 15:30 mforns@deploy1003: Started deploy [analytics/refinery@49c336c]: Regular analytics weekly train [analytics/refinery@49c336cd] * 15:30 mforns@deploy1003: Finished deploy [analytics/refinery@49c336c] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@49c336cd] (duration: 00m 32s) * 15:29 mforns@deploy1003: Started deploy [analytics/refinery@49c336c] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@49c336cd] * 15:29 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1324733{{!}}WikimediaAntiAbuse: Enable personal info tag display on enwiki (T431292)]] * 15:27 brennen@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324700{{!}}EventDetailsParticipantsModule: populate cache with non-local users (T434597)]] (duration: 06m 38s) * 15:23 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[2024,1031]*.wmnet: Set storage compatability to UPGRADING — [[phab:T433028|T433028]] - eevans@cumin1003 * 15:23 brennen@deploy1003: brennen, daimona: Continuing with deployment * 15:23 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:22 brennen@deploy1003: brennen, daimona: Backport for [[gerrit:1324700{{!}}EventDetailsParticipantsModule: populate cache with non-local users (T434597)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:22 cgoubert@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host rdb-lock2003.codfw.wmnet * 15:21 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 15:21 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 15:21 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:21 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:21 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:20 brennen@deploy1003: Started scap sync-world: Backport for [[gerrit:1324700{{!}}EventDetailsParticipantsModule: populate cache with non-local users (T434597)]] * 15:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1178.eqiad.wmnet with OS bookworm * 15:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock1003.eqiad.wmnet * 15:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock1003.eqiad.wmnet with OS trixie * 15:16 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 15:16 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 15:16 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 15:16 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:16 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:16 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:12 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 15:12 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2003.codfw.wmnet * 15:11 cgoubert@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host rdb-lock2003.codfw.wmnet * 15:11 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 15:11 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 15:11 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:11 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:11 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:07 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324338{{!}}InitialiseSettings: Enable 2FA warnings on more private wikis (T428103)]] (duration: 07m 02s) * 15:04 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 15:04 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock1003.eqiad.wmnet with reason: host reimage * 15:03 reedy@deploy1003: reedy: Continuing with deployment * 15:02 reedy@deploy1003: reedy: Backport for [[gerrit:1324338{{!}}InitialiseSettings: Enable 2FA warnings on more private wikis (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:02 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:00 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324338{{!}}InitialiseSettings: Enable 2FA warnings on more private wikis (T428103)]] * 14:57 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 14:57 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2003.codfw.wmnet * 14:57 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock1003.eqiad.wmnet with reason: host reimage * 14:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock2002.codfw.wmnet * 14:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock2002.codfw.wmnet with OS trixie * 14:56 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 14:56 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 14:56 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 14:55 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 14:55 moritzm: powercycle ganeti2046 * 14:47 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock1003.eqiad.wmnet with OS trixie * 14:46 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1003.eqiad.wmnet - cgoubert@cumin2003" * 14:46 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1003.eqiad.wmnet - cgoubert@cumin2003" * 14:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock1003.eqiad.wmnet on all recursors * 14:45 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock1003.eqiad.wmnet on all recursors * 14:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1003.eqiad.wmnet - cgoubert@cumin2003" * 14:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1272.eqiad.wmnet with reason: Enabling notifications * 14:44 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1003.eqiad.wmnet - cgoubert@cumin2003" * 14:44 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324719{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324720{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324722{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0 (T434187)]] (duration: 11m 02s) * 14:39 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 14:39 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock1003.eqiad.wmnet * 14:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock2002.codfw.wmnet with reason: host reimage * 14:37 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 14:37 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock1002.eqiad.wmnet * 14:37 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock1002.eqiad.wmnet with OS trixie * 14:37 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1324719{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324720{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324722{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0 (T434187)]] synced to the testservers (see https://wikitech. * 14:33 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock2002.codfw.wmnet with reason: host reimage * 14:33 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1324719{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324720{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324722{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0 (T434187)]] * 14:32 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1018.eqiad.wmnet with OS bookworm * 14:32 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2046.codfw.wmnet * 14:31 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1020.eqiad.wmnet with OS bookworm * 14:27 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2046.codfw.wmnet * 14:25 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2045.codfw.wmnet * 14:25 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2045.codfw.wmnet * 14:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock1002.eqiad.wmnet with reason: host reimage * 14:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1019.eqiad.wmnet with OS bookworm * 14:22 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Checking sanitization for wikis testwiki in section s3 * 14:20 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2045.codfw.wmnet * 14:18 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock1002.eqiad.wmnet with reason: host reimage * 14:17 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:16 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324707{{!}}Backport all changes from wmf/1.47.0-wmf.15]] (duration: 40m 51s) * 14:16 moritzm: installing Linux 6.1.180 on Bookworm hosts * 14:15 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock2002.codfw.wmnet with OS trixie * 14:14 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2002.codfw.wmnet - cgoubert@cumin2003" * 14:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2002.codfw.wmnet - cgoubert@cumin2003" * 14:14 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:14 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2045.codfw.wmnet * 14:14 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2002.codfw.wmnet on all recursors * 14:14 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2002.codfw.wmnet on all recursors * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2002.codfw.wmnet - cgoubert@cumin2003" * 14:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2002.codfw.wmnet - cgoubert@cumin2003" * 14:12 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2030.codfw.wmnet * 14:12 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2030.codfw.wmnet * 14:11 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on apifeatureusage2001.codfw.wmnet with reason: host reimage * 14:09 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:08 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:07 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 14:06 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2030.codfw.wmnet * 14:06 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:06 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock1002.eqiad.wmnet with OS trixie * 14:05 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 14:05 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2002.codfw.wmnet * 14:05 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1002.eqiad.wmnet - cgoubert@cumin2003" * 14:05 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1002.eqiad.wmnet - cgoubert@cumin2003" * 14:05 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock1002.eqiad.wmnet on all recursors * 14:05 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock1002.eqiad.wmnet on all recursors * 14:05 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:05 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1002.eqiad.wmnet - cgoubert@cumin2003" * 14:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:04 kharlan@deploy1003: kharlan: Continuing with deployment * 14:04 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock2001.codfw.wmnet * 14:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock2001.codfw.wmnet with OS trixie * 14:02 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1002.eqiad.wmnet - cgoubert@cumin2003" * 14:02 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on apifeatureusage2001.codfw.wmnet with reason: host reimage * 14:01 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2030.codfw.wmnet * 13:59 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2029.codfw.wmnet * 13:58 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2029.codfw.wmnet * 13:58 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 13:58 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock1002.eqiad.wmnet * 13:56 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1019.eqiad.wmnet with reason: host reimage * 13:54 btullis@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on archiva1002.wikimedia.org with reason: Upgrading in-place * 13:53 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock1001.eqiad.wmnet * 13:53 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock1001.eqiad.wmnet with OS trixie * 13:53 kharlan@deploy1003: kharlan: Backport for [[gerrit:1324707{{!}}Backport all changes from wmf/1.47.0-wmf.15]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:52 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2029.codfw.wmnet * 13:52 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1020.eqiad.wmnet with reason: host reimage * 13:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1019.eqiad.wmnet with reason: host reimage * 13:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1020.eqiad.wmnet with reason: host reimage * 13:49 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock2001.codfw.wmnet with reason: host reimage * 13:48 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2029.codfw.wmnet * 13:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Configuring db1272 for s3 pooling', diff saved to https://phabricator.wikimedia.org/P96021 and previous config saved to /var/cache/conftool/dbconfig/20260812-134732-cwilliams.json * 13:44 bking@cumin2003: START - Cookbook sre.hosts.reimage for host apifeatureusage2001.codfw.wmnet with OS bookworm * 13:43 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock2001.codfw.wmnet with reason: host reimage * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2028.codfw.wmnet * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2028.codfw.wmnet * 13:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1018.eqiad.wmnet with reason: host reimage * 13:38 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock1001.eqiad.wmnet with reason: host reimage * 13:36 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1018.eqiad.wmnet with reason: host reimage * 13:35 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1324707{{!}}Backport all changes from wmf/1.47.0-wmf.15]] * 13:35 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2028.codfw.wmnet * 13:32 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1275938{{!}}Enable campaignEvents on bdwikimedia (T424016)]] (duration: 07m 35s) * 13:32 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1020.eqiad.wmnet with OS bookworm * 13:32 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1019.eqiad.wmnet with OS bookworm * 13:32 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock1001.eqiad.wmnet with reason: host reimage * 13:28 kharlan@deploy1003: kharlan, yahya: Continuing with deployment * 13:27 kharlan@deploy1003: kharlan, yahya: Backport for [[gerrit:1275938{{!}}Enable campaignEvents on bdwikimedia (T424016)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:27 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2028.codfw.wmnet * 13:26 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 13:26 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 13:25 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1275938{{!}}Enable campaignEvents on bdwikimedia (T424016)]] * 13:25 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitize-wiki (exit_code=99) Managing sanitization for wikis testwiki in section s3 * 13:24 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock2001.codfw.wmnet with OS trixie * 13:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2001.codfw.wmnet - cgoubert@cumin2003" * 13:24 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2001.codfw.wmnet - cgoubert@cumin2003" * 13:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2001.codfw.wmnet on all recursors * 13:23 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2001.codfw.wmnet on all recursors * 13:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2001.codfw.wmnet - cgoubert@cumin2003" * 13:23 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2001.codfw.wmnet - cgoubert@cumin2003" * 13:23 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324713{{!}}thwiki: reinstate temporary wiki25 logos (T431094)]] (duration: 07m 13s) * 13:20 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:20 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1018.eqiad.wmnet with OS bookworm * 13:20 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2027.codfw.wmnet * 13:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2027.codfw.wmnet * 13:19 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics-external: apply * 13:18 kharlan@deploy1003: anzx, kharlan: Continuing with deployment * 13:18 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:18 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock1001.eqiad.wmnet with OS trixie * 13:17 kharlan@deploy1003: anzx, kharlan: Backport for [[gerrit:1324713{{!}}thwiki: reinstate temporary wiki25 logos (T431094)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1001.eqiad.wmnet - cgoubert@cumin2003" * 13:17 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 13:17 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1001.eqiad.wmnet - cgoubert@cumin2003" * 13:17 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2001.codfw.wmnet * 13:17 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics-external: apply * 13:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock1001.eqiad.wmnet on all recursors * 13:17 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock1001.eqiad.wmnet on all recursors * 13:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1001.eqiad.wmnet - cgoubert@cumin2003" * 13:17 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1001.eqiad.wmnet - cgoubert@cumin2003" * 13:15 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1324713{{!}}thwiki: reinstate temporary wiki25 logos (T431094)]] * 13:15 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:15 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics-external: apply * 13:15 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:14 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics-external: apply * 13:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2027.codfw.wmnet * 13:12 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 13:12 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock1001.eqiad.wmnet * 13:12 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2027.codfw.wmnet * 13:08 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1158.eqiad.wmnet onto db1273.eqiad.wmnet * 13:07 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1158: Pool db1158.eqiad.wmnet in after cloning * 13:02 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti6002.drmrs.wmnet * 13:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti6002.drmrs.wmnet * 12:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1015.eqiad.wmnet with OS bookworm * 12:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti6002.drmrs.wmnet * 12:44 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1017.eqiad.wmnet with OS bookworm * 12:36 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 12:35 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 12:34 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 12:33 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 12:31 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti6002.drmrs.wmnet * 12:24 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1017.eqiad.wmnet with reason: host reimage * 12:22 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1158: Pool db1158.eqiad.wmnet in after cloning * 12:18 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1017.eqiad.wmnet with reason: host reimage * 12:11 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1015.eqiad.wmnet with reason: host reimage * 12:07 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1015.eqiad.wmnet with reason: host reimage * 12:04 moritzm: failover ganeti master in drmrs02 to ganeti6004 * 12:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1009.eqiad.wmnet with OS bookworm * 12:01 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1017.eqiad.wmnet with OS bookworm * 12:00 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti6004.drmrs.wmnet * 12:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti6004.drmrs.wmnet * 11:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti6004.drmrs.wmnet * 11:53 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1015.eqiad.wmnet with OS bookworm * 11:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1159.eqiad.wmnet onto db1274.eqiad.wmnet * 11:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Pool db1159.eqiad.wmnet in after cloning * 11:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1016.eqiad.wmnet with OS bookworm * 11:46 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti6004.drmrs.wmnet * 11:45 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti6001.drmrs.wmnet * 11:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti6001.drmrs.wmnet * 11:43 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324306{{!}}WikimediaAntiAbuse: Enable personal info for enwiki with no display (T431292)]] (duration: 10m 26s) * 11:42 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1015.eqiad.wmnet with OS bookworm * 11:39 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 11:38 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti6001.drmrs.wmnet * 11:34 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1324306{{!}}WikimediaAntiAbuse: Enable personal info for enwiki with no display (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:33 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2212: Security update * 11:33 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti6001.drmrs.wmnet * 11:32 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1324306{{!}}WikimediaAntiAbuse: Enable personal info for enwiki with no display (T431292)]] * 11:22 moritzm: failover ganeti master in drmrs01 to ganeti6003 * 11:20 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:20 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:18 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 22 hosts with reason: Cloning * 11:17 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti6003.drmrs.wmnet * 11:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti6003.drmrs.wmnet * 11:17 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1009.eqiad.wmnet with reason: host reimage * 11:17 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:16 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:14 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1016.eqiad.wmnet with reason: host reimage * 11:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti6003.drmrs.wmnet * 11:11 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1009.eqiad.wmnet with reason: host reimage * 11:10 jmm@cumin2003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti3005.esams.wmnet * 11:10 jmm@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host ganeti3005.esams.wmnet * 11:09 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 11:08 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1158: Depool db1158.eqiad.wmnet to then clone it to db1273.eqiad.wmnet - marostegui@cumin1003 * 11:07 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1016.eqiad.wmnet with reason: host reimage * 11:07 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1158: Depool db1158.eqiad.wmnet to then clone it to db1273.eqiad.wmnet - marostegui@cumin1003 * 11:07 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1158.eqiad.wmnet onto db1273.eqiad.wmnet * 11:06 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:05 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:05 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1159: Pool db1159.eqiad.wmnet in after cloning * 11:04 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 20 hosts with reason: Cloning * 11:02 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti6003.drmrs.wmnet * 11:00 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis testwiki in section s3 * 10:54 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1009.eqiad.wmnet with OS bookworm * 10:51 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1015.eqiad.wmnet with OS bookworm * 10:50 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1016.eqiad.wmnet with OS bookworm * 10:48 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2212: Security update * 10:45 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitize-wiki (exit_code=99) Managing sanitization for wikis testwiki in section s3 * 10:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1013.eqiad.wmnet with OS bookworm * 10:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1014.eqiad.wmnet with OS bookworm * 10:38 cwilliams@cumin1003: START - Cookbook sre.mysql.clone of db1159.eqiad.wmnet onto db1274.eqiad.wmnet * 10:33 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db1274.eqiad.wmnet * 10:33 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db1274.eqiad.wmnet * 10:31 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 396993 * 10:29 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 396993 * 10:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1159: Clone source for db1274 * 10:24 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1159: Clone source for db1274 * 10:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1013.eqiad.wmnet with reason: host reimage * 10:18 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1013.eqiad.wmnet with reason: host reimage * 10:13 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2212.codfw.wmnet with reason: Maintenance * 10:13 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 10:12 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 10:12 blake@deploy1003: Stopping before sync operations * 10:11 blake@deploy1003: Started scap sync-world: Non-deployment scap run to populate new release values for [[phab:T427668|T427668]] * 10:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2212 [[phab:T434644|T434644]]', diff saved to https://phabricator.wikimedia.org/P96003 and previous config saved to /var/cache/conftool/dbconfig/20260812-101053-cwilliams.json * 10:09 moritzm: powercycle ganeti3005 * 10:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2203 to s1 primary [[phab:T434644|T434644]]', diff saved to https://phabricator.wikimedia.org/P96002 and previous config saved to /var/cache/conftool/dbconfig/20260812-100849-cwilliams.json * 10:08 cezmunsta: Starting s1 codfw failover from db2212 to db2203 - [[phab:T434644|T434644]] * 10:03 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1013.eqiad.wmnet with OS bookworm * 10:02 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1013.eqiad.wmnet with OS bookworm * 10:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2203 with weight 0 [[phab:T434644|T434644]]', diff saved to https://phabricator.wikimedia.org/P96001 and previous config saved to /var/cache/conftool/dbconfig/20260812-100134-cwilliams.json * 10:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 32 hosts with reason: Primary switchover s1 [[phab:T434644|T434644]] * 09:53 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1014.eqiad.wmnet with reason: host reimage * 09:50 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 09:50 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti3005.esams.wmnet * 09:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1014.eqiad.wmnet with reason: host reimage * 09:41 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1278: Pool in x1 * 09:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-launcher1003.eqiad.wmnet with OS bookworm * 09:37 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti3005.esams.wmnet * 09:37 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1013.eqiad.wmnet with OS bookworm * 09:34 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-presto1013.eqiad.wmnet with OS bookworm * 09:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1014.eqiad.wmnet with OS bookworm * 09:29 moritzm: failover ganeti master in esams to ganeti3008 * 09:26 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti3006.esams.wmnet * 09:26 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti3006.esams.wmnet * 09:24 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1012.eqiad.wmnet with OS bookworm * 09:23 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1009.eqiad.wmnet with OS bookworm * 09:23 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 09:18 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti3006.esams.wmnet * 09:16 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti3006.esams.wmnet * 09:14 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: fix regexp escaping bug - oblivian@cumin1003" * 09:14 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: fix regexp escaping bug - oblivian@cumin1003 * 09:13 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: fix regexp escaping bug - oblivian@cumin1003 * 09:13 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: fix regexp escaping bug - oblivian@cumin1003" * 09:03 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-launcher1003.eqiad.wmnet with reason: host reimage * 08:58 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-launcher1003.eqiad.wmnet with reason: host reimage * 08:55 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1278: Pool in x1 * 08:55 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1278 to dbctl [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95996 and previous config saved to /var/cache/conftool/dbconfig/20260812-085521-marostegui.json * 08:51 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1012.eqiad.wmnet with reason: host reimage * 08:45 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis testwiki in section s3 * 08:43 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1009.eqiad.wmnet with OS bookworm * 08:42 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1012.eqiad.wmnet with reason: host reimage * 08:41 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-launcher1003.eqiad.wmnet with OS bookworm * 08:40 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1013.eqiad.wmnet with OS bookworm * 08:38 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms3', diff saved to https://phabricator.wikimedia.org/P95995 and previous config saved to /var/cache/conftool/dbconfig/20260812-083816-marostegui.json * 08:38 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-master1003.eqiad.wmnet with OS bookworm * 08:37 marostegui: Failover ms3 [[phab:T434288|T434288]] * 08:37 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1268 to dbctl [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95994 and previous config saved to /var/cache/conftool/dbconfig/20260812-083722-marostegui.json * 08:35 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-web1001.eqiad.wmnet with OS bookworm * 08:32 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db2252.codfw.wmnet,db[1153,1268].eqiad.wmnet with reason: Switching over ms3 * 08:28 marostegui@cumin1003: dbctl commit (dc=all): 'Depool ms3 [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95993 and previous config saved to /var/cache/conftool/dbconfig/20260812-082852-marostegui.json * 08:25 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1012.eqiad.wmnet with OS bookworm * 08:25 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1011.eqiad.wmnet with OS bookworm * 08:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-master1003.eqiad.wmnet with reason: host reimage * 08:07 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-master1003.eqiad.wmnet with reason: host reimage * 08:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-web1001.eqiad.wmnet with reason: host reimage * 07:58 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-web1001.eqiad.wmnet with reason: host reimage * 07:50 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1003.eqiad.wmnet with OS bookworm * 07:38 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1011.eqiad.wmnet with reason: host reimage * 07:38 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-master1003.eqiad.wmnet with OS bookworm * 07:35 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti3007.esams.wmnet * 07:35 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti3007.esams.wmnet * 07:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1011.eqiad.wmnet with reason: host reimage * 07:27 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti3007.esams.wmnet * 07:25 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti3007.esams.wmnet * 07:25 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti3008.esams.wmnet * 07:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti3008.esams.wmnet * 07:22 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-web1001.eqiad.wmnet with OS bookworm * 07:19 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1003.eqiad.wmnet with OS bookworm * 07:18 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-master1003.eqiad.wmnet * 07:18 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host an-master1003.eqiad.wmnet * 07:17 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1011.eqiad.wmnet with OS bookworm * 07:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 07:16 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti3008.esams.wmnet * 07:15 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1009.eqiad.wmnet with OS bookworm * 07:14 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host an-master1003.eqiad.wmnet * 07:13 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-master1003.eqiad.wmnet * 07:13 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-master1003.eqiad.wmnet * 07:12 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-master1003.eqiad.wmnet * 07:11 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti3008.esams.wmnet * 07:07 arnaudb@dns1006: END - running authdns-update * 07:07 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti5007.eqsin.wmnet * 07:07 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti5007.eqsin.wmnet * 07:05 arnaudb@dns1006: START - running authdns-update * 06:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti5007.eqsin.wmnet * 06:54 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti5007.eqsin.wmnet * 06:38 moritzm: failover ganeti master in eqsin to ganeti5004 * 06:36 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti5006.eqsin.wmnet * 06:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti5006.eqsin.wmnet * 06:28 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti5006.eqsin.wmnet * 06:23 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti5006.eqsin.wmnet * 06:20 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti5005.eqsin.wmnet * 06:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti5005.eqsin.wmnet * 06:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti5005.eqsin.wmnet * 06:06 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti5005.eqsin.wmnet * 06:03 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti5004.eqsin.wmnet * 06:03 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti5004.eqsin.wmnet * 05:55 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti5004.eqsin.wmnet * 05:53 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti5004.eqsin.wmnet * 04:40 ryankemper: [[phab:T434494|T434494]] reimaged `an-tool1008.eqiad.wmnet` to bookworm; yarn.wikimedia.org is back up * 04:16 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-tool1008.eqiad.wmnet with OS bookworm * 03:58 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-tool1008.eqiad.wmnet with reason: host reimage * 03:53 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-tool1008.eqiad.wmnet with reason: host reimage * 03:41 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-tool1008.eqiad.wmnet with OS bookworm * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 45s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 00:25 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324427{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]], [[gerrit:1324429{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0]], [[gerrit:1324428{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]] (duration: 07m 55s) * 00:21 kemayo@deploy1003: kemayo: Continuing with deployment * 00:19 kemayo@deploy1003: kemayo: Backport for [[gerrit:1324427{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]], [[gerrit:1324429{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0]], [[gerrit:1324428{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:17 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1324427{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]], [[gerrit:1324429{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0]], [[gerrit:1324428{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]] == 2026-08-11 == * 21:37 sbassett: Deployed security fix for [[phab:T434521|T434521]] (wmf.15) * 21:29 sbassett: Deployed security fix for [[phab:T434521|T434521]] (wmf.14) * 21:19 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324370{{!}}Phase 4 of legal footer deployment (T432796)]], [[gerrit:1319804{{!}}Disable wgMFCustomSiteModules on English Wikipedia (T375538)]] (duration: 15m 26s) * 21:15 jdlrobson@deploy1003: jdlrobson: Continuing with deployment * 21:06 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1324370{{!}}Phase 4 of legal footer deployment (T432796)]], [[gerrit:1319804{{!}}Disable wgMFCustomSiteModules on English Wikipedia (T375538)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:03 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1324370{{!}}Phase 4 of legal footer deployment (T432796)]], [[gerrit:1319804{{!}}Disable wgMFCustomSiteModules on English Wikipedia (T375538)]] * 20:59 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 20:50 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324384{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]], [[gerrit:1324385{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]] (duration: 06m 58s) * 20:46 kemayo@deploy1003: kemayo: Continuing with deployment * 20:45 kemayo@deploy1003: kemayo: Backport for [[gerrit:1324384{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]], [[gerrit:1324385{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:43 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1324384{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]], [[gerrit:1324385{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]] * 20:42 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 20:42 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324386{{!}}build: Updating js-yaml to 3.15.1, 4.3.1]] (duration: 07m 36s) * 20:38 kemayo@deploy1003: kemayo: Continuing with deployment * 20:37 jhancock@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 20:37 kemayo@deploy1003: kemayo: Backport for [[gerrit:1324386{{!}}build: Updating js-yaml to 3.15.1, 4.3.1]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:35 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1324386{{!}}build: Updating js-yaml to 3.15.1, 4.3.1]] * 20:18 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 20:15 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 20:15 jhancock@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin1003" * 20:14 jhancock@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin1003" * 19:59 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 19:54 jhancock@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 19:10 brennen@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] (duration: 06m 41s) * 19:04 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns6001.wikimedia.org * 19:04 sukhe@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns6001.wikimedia.org * 19:04 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns5003.wikimedia.org * 19:04 sukhe@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns5003.wikimedia.org * 19:03 brennen@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 18:59 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns5003.wikimedia.org with OS trixie * 18:55 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns6001.wikimedia.org with OS trixie * 18:19 brett@cumin2002: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on P<nowiki>{</nowiki>cp7009.magru.wmnet<nowiki>}</nowiki> and A:cp - 9.2.15 Upgrade () * 18:14 brett@cumin2002: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on P<nowiki>{</nowiki>cp7009.magru.wmnet<nowiki>}</nowiki> and A:cp - 9.2.15 Upgrade () * 18:13 brennen@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 18:12 brett@cumin2002: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 9.2.15 Upgrade () * 18:09 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns5003.wikimedia.org with reason: host reimage * 18:06 brett@cumin2002: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 9.2.15 Upgrade () * 18:06 brennen: 1.47.0-wmf.15 train status ([[phab:T430834|T430834]]) - no current blockers, rolling to group0 * 18:05 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns5003.wikimedia.org with reason: host reimage * 18:05 brett: import trafficserver-9.2.15~deb13+wmf1 into trixie-wikimedia ([[phab:T434478|T434478]]) * 18:01 ladsgroup@cumin1003: END (PASS) - Cookbook sre.mysql.sanitarium_restart (exit_code=0) * 17:58 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns6001.wikimedia.org with reason: host reimage * 17:53 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324369{{!}}Enable desktop lazy loading on group0 (T148047)]] (duration: 07m 31s) * 17:52 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns6001.wikimedia.org with reason: host reimage * 17:49 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 17:49 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitarium_restart (exit_code=99) * 17:49 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 17:49 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7001.magru.wmnet * 17:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti7001.magru.wmnet * 17:48 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 17:47 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1324369{{!}}Enable desktop lazy loading on group0 (T148047)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:45 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1324369{{!}}Enable desktop lazy loading on group0 (T148047)]] * 17:39 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti7001.magru.wmnet * 17:36 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns5003.wikimedia.org with OS trixie * 17:34 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns6001.wikimedia.org with OS trixie * 17:31 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324361{{!}}Move FR config from IS.php to a dedicated file]], [[gerrit:1324363{{!}}Remove $wmg = $wg hacks in CentralAuth (T119117)]] (duration: 12m 23s) * 17:26 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 17:23 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1324361{{!}}Move FR config from IS.php to a dedicated file]], [[gerrit:1324363{{!}}Remove $wmg = $wg hacks in CentralAuth (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:19 sukhe: sudo cumin "A:cp-magru" "run-puppet-agent --enable 'merging CR 1324355'" * 17:18 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1324361{{!}}Move FR config from IS.php to a dedicated file]], [[gerrit:1324363{{!}}Remove $wmg = $wg hacks in CentralAuth (T119117)]] * 17:11 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-master1004.eqiad.wmnet with OS bookworm * 17:11 sukhe: sukhe@cp7005:~$ sudo puppet agent -tv * 17:02 sukhe: sudo cumin "A:cp-magru" "disable-puppet 'merging CR 1324355'" * 16:54 sukhe@dns1004: END - running authdns-update * 16:53 sukhe@dns1004: START - running authdns-update * 16:53 sukhe@dns1004: FAIL - running authdns-update * 16:51 sukhe@dns1004: START - running authdns-update * 16:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-master1004.eqiad.wmnet with reason: host reimage * 16:44 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-master1004.eqiad.wmnet with reason: host reimage * 16:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1157.eqiad.wmnet onto db1272.eqiad.wmnet * 16:40 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1157: Pool db1157.eqiad.wmnet in after cloning * 16:38 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 16:31 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324356{{!}}InitialiseSettings: Fix wgOATHAuthEnforce2FAForAll]] (duration: 06m 52s) * 16:30 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 16:28 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 16:27 reedy@deploy1003: reedy: Continuing with deployment * 16:26 reedy@deploy1003: reedy: Backport for [[gerrit:1324356{{!}}InitialiseSettings: Fix wgOATHAuthEnforce2FAForAll]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:24 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324356{{!}}InitialiseSettings: Fix wgOATHAuthEnforce2FAForAll]] * 16:13 sukhe: restart ntpsec.serviceon dns7001 * 16:09 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324335{{!}}InitialiseSettings: Enable 2FA enforcement on various private wikis (T428103)]] (duration: 06m 40s) * 16:08 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 16:06 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2204: Security update * 16:04 reedy@deploy1003: reedy: Continuing with deployment * 16:04 reedy@deploy1003: reedy: Backport for [[gerrit:1324335{{!}}InitialiseSettings: Enable 2FA enforcement on various private wikis (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:02 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7001.magru.wmnet * 16:02 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324335{{!}}InitialiseSettings: Enable 2FA enforcement on various private wikis (T428103)]] * 16:01 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 15:55 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1157: Pool db1157.eqiad.wmnet in after cloning * 15:54 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 15:54 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 15:41 dancy@deploy1003: Finished scap sync-world: Testing (duration: 06m 28s) * 15:40 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-master1004.eqiad.wmnet with OS bookworm * 15:40 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 15:35 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti4008.ulsfo.wmnet * 15:35 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti4008.ulsfo.wmnet * 15:34 dancy@deploy1003: Started scap sync-world: Testing * 15:34 dancy@deploy1003: Installation of scap version "4.279.0" completed for 3 hosts * 15:34 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-master1004.eqiad.wmnet with OS bookworm * 15:34 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 15:33 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-master1004.eqiad.wmnet with OS bookworm * 15:32 dancy@deploy1003: Installing scap version "4.279.0" for 3 host(s) * 15:32 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324339{{!}}Add /w/deployment-info.php entrypoint]] (duration: 07m 25s) * 15:30 moritzm: failover ganeti master in magru to ganeti7004 * 15:29 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti4008.ulsfo.wmnet * 15:28 dancy@deploy1003: dancy: Continuing with deployment * 15:28 tappof: remove 2026-05 swift log archives from centrallog to free some space ([[phab:T434502|T434502]]) * 15:27 dancy@deploy1003: dancy: Backport for [[gerrit:1324339{{!}}Add /w/deployment-info.php entrypoint]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:25 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324339{{!}}Add /w/deployment-info.php entrypoint]] * 15:20 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2204: Security update * 15:18 dancy@deploy1003: Installation of scap version "4.278.0" completed for 3 hosts * 15:18 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7004.magru.wmnet * 15:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti7004.magru.wmnet * 15:16 dancy@deploy1003: Installing scap version "4.278.0" for 3 host(s) * 15:14 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2204.codfw.wmnet with reason: Maintenance * 15:12 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 15:11 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-master1004.eqiad.wmnet with OS bookworm * 15:11 hashar: Restarting CI Jenkins on contint1003 due to Java upgrade. * 15:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2204 [[phab:T434565|T434565]]', diff saved to https://phabricator.wikimedia.org/P95984 and previous config saved to /var/cache/conftool/dbconfig/20260811-151126-cwilliams.json * 15:10 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti4008.ulsfo.wmnet * 15:10 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti7004.magru.wmnet * 15:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2207 to s2 primary [[phab:T434565|T434565]]', diff saved to https://phabricator.wikimedia.org/P95983 and previous config saved to /var/cache/conftool/dbconfig/20260811-150905-cwilliams.json * 15:08 cezmunsta: Starting s2 codfw failover from db2204 to db2207 - [[phab:T434565|T434565]] * 15:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2207 with weight 0 [[phab:T434565|T434565]]', diff saved to https://phabricator.wikimedia.org/P95982 and previous config saved to /var/cache/conftool/dbconfig/20260811-150402-cwilliams.json * 15:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s2 [[phab:T434565|T434565]] * 14:55 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1010.eqiad.wmnet with OS bookworm * 14:49 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns4003.wikimedia.org with OS trixie * 14:47 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1010.eqiad.wmnet with OS bookworm * 14:47 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7001.wikimedia.org with OS trixie * 14:44 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-presto1010.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:41 btullis@cumin1003: START - Cookbook sre.hosts.provision for host an-presto1010.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:40 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-presto1010.eqiad.wmnet with OS bookworm * 14:39 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 14:39 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-presto1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:36 btullis@cumin1003: START - Cookbook sre.hosts.provision for host an-presto1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:32 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1009.eqiad.wmnet with OS bookworm * 14:32 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 14:31 cwilliams@cumin1003: START - Cookbook sre.mysql.clone of db1157.eqiad.wmnet onto db1272.eqiad.wmnet * 14:30 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7004.magru.wmnet * 14:28 moritzm: failover ganeti master in ulsfo to ganeti4005 * 14:23 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7003.magru.wmnet * 14:23 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti7003.magru.wmnet * 14:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-master1004.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:22 btullis@cumin1003: START - Cookbook sre.hosts.provision for host an-master1004.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:21 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-master1004.eqiad.wmnet with OS bookworm * 14:19 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti4007.ulsfo.wmnet * 14:19 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti4007.ulsfo.wmnet * 14:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1010.eqiad.wmnet with OS bookworm * 14:17 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1008.eqiad.wmnet with OS bookworm * 14:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti7003.magru.wmnet * 14:12 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7003.magru.wmnet * 14:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti4007.ulsfo.wmnet * 14:11 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7002.magru.wmnet * 14:11 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti7002.magru.wmnet * 14:09 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324318{{!}}Revert "wmf-config/ProductionServices: set URL for urldownloader to service record" (T429175)]] (duration: 06m 46s) * 14:09 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7001.wikimedia.org with reason: host reimage * 14:06 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti4007.ulsfo.wmnet * 14:05 kharlan@deploy1003: kharlan: Continuing with deployment * 14:05 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti4006.ulsfo.wmnet * 14:04 jayme: updated calico to v3.30.7 on staging-eqiad - [[phab:T427400|T427400]] * 14:04 kharlan@deploy1003: kharlan: Backport for [[gerrit:1324318{{!}}Revert "wmf-config/ProductionServices: set URL for urldownloader to service record" (T429175)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti4006.ulsfo.wmnet * 14:03 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns4003.wikimedia.org with reason: host reimage * 14:03 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7001.wikimedia.org with reason: host reimage * 14:02 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti7002.magru.wmnet * 14:02 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1324318{{!}}Revert "wmf-config/ProductionServices: set URL for urldownloader to service record" (T429175)]] * 14:02 btullis@dns1004: FAIL - running authdns-update * 14:00 btullis@dns1004: START - running authdns-update * 13:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1008.eqiad.wmnet with reason: host reimage * 13:59 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'. * 13:59 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1313985{{!}}wmf-config/ProductionServices: set URL for urldownloader to service record (T429175)]] (duration: 25m 06s) * 13:58 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7002.magru.wmnet * 13:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti4006.ulsfo.wmnet * 13:57 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns4003.wikimedia.org with reason: host reimage * 13:56 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'. * 13:56 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7001.magru.wmnet * 13:56 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1008.eqiad.wmnet with reason: host reimage * 13:55 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'. * 13:55 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'. * 13:55 kharlan@deploy1003: kharlan, sukhe: Continuing with deployment * 13:53 marostegui: Failover ms2 [[phab:T434288|T434288]] * 13:52 marostegui: Failover ms1 [[phab:T434288|T434288]] * 13:52 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7001.magru.wmnet * 13:51 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti4006.ulsfo.wmnet * 13:48 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti4005.ulsfo.wmnet * 13:48 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti4005.ulsfo.wmnet * 13:44 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti4005.ulsfo.wmnet * 13:40 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1008.eqiad.wmnet with OS bookworm * 13:39 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns4003.wikimedia.org with OS trixie * 13:38 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns7001.wikimedia.org with OS trixie * 13:38 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1008.eqiad.wmnet with OS bookworm * 13:37 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti4005.ulsfo.wmnet * 13:36 kharlan@deploy1003: kharlan, sukhe: Backport for [[gerrit:1313985{{!}}wmf-config/ProductionServices: set URL for urldownloader to service record (T429175)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:34 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1313985{{!}}wmf-config/ProductionServices: set URL for urldownloader to service record (T429175)]] * 13:29 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2034.codfw.wmnet * 13:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2034.codfw.wmnet * 13:25 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.clone (exit_code=99) of db1157.eqiad.wmnet onto db1272.eqiad.wmnet * 13:25 cwilliams@cumin1003: START - Cookbook sre.mysql.clone of db1157.eqiad.wmnet onto db1272.eqiad.wmnet * 13:21 urbanecm@deploy1003: mwscript-k8s job started: namespaceDupes.php --wiki=frwiktionary --fix # [[phab:T415716|T415716]] * 13:21 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2034.codfw.wmnet * 13:20 urbanecm@deploy1003: mwscript-k8s job started: namespaceDupes.php --wiki=frwiktionary # [[phab:T415716|T415716]] * 13:19 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1323350{{!}}[tgwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T415307)]], [[gerrit:1322961{{!}}[slwiki] Revert temporary logo for Wikipedia 25 (Vector legacy + Vector 2022) (T414265)]], [[gerrit:1323827{{!}}[itwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T414320)]] (duration: 08m 00s) * 13:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-coord1003.eqiad.wmnet with OS bookworm * 13:15 urbanecm@deploy1003: urbanecm, superpes: Continuing with deployment * 13:13 urbanecm@deploy1003: urbanecm, superpes: Backport for [[gerrit:1323350{{!}}[tgwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T415307)]], [[gerrit:1322961{{!}}[slwiki] Revert temporary logo for Wikipedia 25 (Vector legacy + Vector 2022) (T414265)]], [[gerrit:1323827{{!}}[itwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T414320)]] synced to the testservers (see https://wiki * 13:11 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1323350{{!}}[tgwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T415307)]], [[gerrit:1322961{{!}}[slwiki] Revert temporary logo for Wikipedia 25 (Vector legacy + Vector 2022) (T414265)]], [[gerrit:1323827{{!}}[itwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T414320)]] * 13:11 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1323779{{!}}[ukwiki] Remove reviewer usergroup (T434252)]], [[gerrit:1323312{{!}}[frwiktionary] Add new Schème namespace and its talk (T415716)]] (duration: 06m 49s) * 13:10 marostegui@dns1004: END - running authdns-update * 13:08 marostegui@dns1004: START - running authdns-update * 13:07 marostegui@cumin1003: dbctl commit (dc=all): 'Repool ms2 [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95980 and previous config saved to /var/cache/conftool/dbconfig/20260811-130725-marostegui.json * 13:06 urbanecm@deploy1003: urbanecm, superpes: Continuing with deployment * 13:06 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1266 to dbctl [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95979 and previous config saved to /var/cache/conftool/dbconfig/20260811-130627-marostegui.json * 13:06 urbanecm@deploy1003: urbanecm, superpes: Backport for [[gerrit:1323779{{!}}[ukwiki] Remove reviewer usergroup (T434252)]], [[gerrit:1323312{{!}}[frwiktionary] Add new Schème namespace and its talk (T415716)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:04 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1323779{{!}}[ukwiki] Remove reviewer usergroup (T434252)]], [[gerrit:1323312{{!}}[frwiktionary] Add new Schème namespace and its talk (T415716)]] * 12:59 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db2253.codfw.wmnet,db[1151,1266].eqiad.wmnet with reason: Switching over ms2 * 12:54 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1157: Using as clone source * 12:53 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1157: Using as clone source * 12:51 marostegui@cumin1003: dbctl commit (dc=all): 'Depool ms2 [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95977 and previous config saved to /var/cache/conftool/dbconfig/20260811-125129-marostegui.json * 12:47 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 12:46 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 12:45 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 12:44 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'recommendation-api-ng' for release 'main' . * 12:44 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'recommendation-api-ng' for release 'main' . * 12:43 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'recommendation-api-ng' for release 'main' . * 12:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-coord1003.eqiad.wmnet with reason: host reimage * 12:43 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'ores-legacy' for release 'main' . * 12:42 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'ores-legacy' for release 'main' . * 12:42 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2165: Security update * 12:40 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-coord1003.eqiad.wmnet with reason: host reimage * 12:39 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'ores-legacy' for release 'main' . * 12:38 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' . * 12:38 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' . * 12:37 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' . * 12:34 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2034.codfw.wmnet * 12:30 jmm@cumin2002: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti-test2001.codfw.wmnet * 12:30 jmm@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host ganeti-test2001.codfw.wmnet * 12:25 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1179.eqiad.wmnet onto db1278.eqiad.wmnet * 12:25 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1179: Pool db1179.eqiad.wmnet in after cloning * 12:23 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-coord1003.eqiad.wmnet with OS bookworm * 12:22 moritzm: failover ganeti master in codfw/routed to ganeti2033 * 12:22 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2033.codfw.wmnet * 12:22 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2033.codfw.wmnet * 12:19 jmm@cumin2002: START - Cookbook sre.hosts.reboot-single for host ganeti-test2001.codfw.wmnet * 12:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 12:18 jmm@cumin2002: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti-test2001.codfw.wmnet * 12:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1008.eqiad.wmnet with OS bookworm * 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2033.codfw.wmnet * 12:07 moritzm: failover ganeti master in ganeti/test to ganeti-test2003 * 12:04 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 12:03 jmm@cumin2003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti4005.ulsfo.wmnet * 12:03 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti4005.ulsfo.wmnet * 12:00 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324283{{!}}Use maximum compression level in SqlBlobStore and SqlBagOStuff (T428377)]] (duration: 11m 37s) * 11:57 jmm@cumin2002: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti-test2002.codfw.wmnet * 11:57 jmm@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti-test2002.codfw.wmnet * 11:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2165: Security update * 11:54 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 11:52 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1324283{{!}}Use maximum compression level in SqlBlobStore and SqlBagOStuff (T428377)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:51 jmm@cumin2002: START - Cookbook sre.hosts.reboot-single for host ganeti-test2002.codfw.wmnet * 11:50 jmm@cumin2002: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti-test2002.codfw.wmnet * 11:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2165.codfw.wmnet with reason: Maintenance * 11:48 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@050d19e] (releasing): [[phab:T434186|T434186]] (duration: 01m 14s) * 11:48 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1324283{{!}}Use maximum compression level in SqlBlobStore and SqlBagOStuff (T428377)]] * 11:47 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@050d19e] (releasing): [[phab:T434186|T434186]] * 11:44 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@050d19e] (releasing): test jenkins deploy for [[phab:T434186|T434186]] (duration: 01m 08s) * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2165 [[phab:T434514|T434514]]', diff saved to https://phabricator.wikimedia.org/P95969 and previous config saved to /var/cache/conftool/dbconfig/20260811-114352-cwilliams.json * 11:43 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@050d19e] (releasing): test jenkins deploy for [[phab:T434186|T434186]] * 11:42 jmm@cumin2003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti-test2003.codfw.wmnet * 11:42 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti-test2003.codfw.wmnet * 11:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2161 to s8 primary [[phab:T434514|T434514]]', diff saved to https://phabricator.wikimedia.org/P95968 and previous config saved to /var/cache/conftool/dbconfig/20260811-114136-cwilliams.json * 11:40 cezmunsta: Starting s8 codfw failover from db2165 to db2161 - [[phab:T434514|T434514]] * 11:40 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1179: Pool db1179.eqiad.wmnet in after cloning * 11:36 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti-test2003.codfw.wmnet * 11:36 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti-test2003.codfw.wmnet * 11:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2161 with weight 0 [[phab:T434514|T434514]]', diff saved to https://phabricator.wikimedia.org/P95966 and previous config saved to /var/cache/conftool/dbconfig/20260811-113449-cwilliams.json * 11:34 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 25 hosts with reason: Primary switchover s8 [[phab:T434514|T434514]] * 11:29 moritzm: installing Python 3.11 security updates * 11:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-presto1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 11:26 btullis@cumin1003: START - Cookbook sre.hosts.provision for host an-presto1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 11:23 btullis@dns1004: END - running authdns-update * 11:21 btullis@dns1004: START - running authdns-update * 11:20 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-presto1008.eqiad.wmnet with OS bookworm * 11:20 moritzm: installing curl security updates * 11:11 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1007.eqiad.wmnet with OS bookworm * 10:45 tappof: bump space for prometheus k8s-dse in codfw * 10:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1007.eqiad.wmnet with reason: host reimage * 10:38 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1007.eqiad.wmnet with reason: host reimage * 10:37 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-coord1004.eqiad.wmnet with OS bookworm * 10:35 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1008.eqiad.wmnet with OS bookworm * 10:34 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1179: Depool db1179.eqiad.wmnet to then clone it to db1278.eqiad.wmnet - marostegui@cumin1003 * 10:34 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1008.eqiad.wmnet with OS bookworm * 10:25 fceratto@cumin1003: dbctl commit (dc=all): 'Remove db1177 [[phab:T433474|T433474]]', diff saved to https://phabricator.wikimedia.org/P95964 and previous config saved to /var/cache/conftool/dbconfig/20260811-102527-fceratto.json * 10:22 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1008.eqiad.wmnet with OS bookworm * 10:22 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1007.eqiad.wmnet with OS bookworm * 10:21 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1006.eqiad.wmnet with OS bookworm * 10:20 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 10:18 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1179: Depool db1179.eqiad.wmnet to then clone it to db1278.eqiad.wmnet - marostegui@cumin1003 * 10:18 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1179.eqiad.wmnet onto db1278.eqiad.wmnet * 10:17 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 10:17 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 10:14 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 10:09 blake@deploy1003: Stopping before sync operations * 10:09 blake@deploy1003: Started scap sync-world: Non-deployment run to populate release values for [[phab:T427668|T427668]] * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 10:04 fceratto@cumin1003: Removing db1177 from zarcillo [[phab:T433474|T433474]] * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1177.eqiad.wmnet * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1177.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:03 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1177.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:00 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1006.eqiad.wmnet with reason: host reimage * 09:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-coord1004.eqiad.wmnet with reason: host reimage * 09:57 marostegui: Failover m1 from db1164 to db1213 - [[phab:T434493|T434493]] * 09:57 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1006.eqiad.wmnet with reason: host reimage * 09:55 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:54 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2232].codfw.wmnet,db[1164,1213,1217].eqiad.wmnet with reason: Primary switchover m1 [[phab:T434493|T434493]] * 09:52 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-coord1004.eqiad.wmnet with reason: host reimage * 09:49 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1213.eqiad.wmnet with OS trixie * 09:49 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1177.eqiad.wmnet * 09:41 moritzm: installing Linux 6.12.101 on Trixie hosts * 09:40 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1006.eqiad.wmnet with OS bookworm * 09:35 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-coord1004.eqiad.wmnet with OS bookworm * 09:28 moritzm: installing node-tar security updates * 09:27 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1213.eqiad.wmnet with reason: host reimage * 09:22 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1213.eqiad.wmnet with reason: host reimage * 09:09 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1177: Decommission * 09:08 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db1177: Decommission * 09:08 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 09:08 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.decommission (exit_code=99) * 09:06 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1213.eqiad.wmnet with OS trixie * 09:06 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 09:05 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1213.eqiad.wmnet with reason: Reimage * 08:53 marostegui@dns1004: END - running authdns-update * 08:51 marostegui@dns1004: START - running authdns-update * 08:48 marostegui: Switchover ms1 master in eqiad [[phab:T434288|T434288]] * 08:48 marostegui@cumin1003: dbctl commit (dc=all): 'Repool ms1 [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95962 and previous config saved to /var/cache/conftool/dbconfig/20260811-084804-marostegui.json * 08:40 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1267 to dbctl [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95961 and previous config saved to /var/cache/conftool/dbconfig/20260811-084054-marostegui.json * 08:29 marostegui: Failover m1 from db1213 to db1164 - [[phab:T434043|T434043]] * 08:25 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2232].codfw.wmnet,db[1164,1213,1217].eqiad.wmnet with reason: Primary switchover m1 [[phab:T434043|T434043]] * 08:22 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db2251.codfw.wmnet,db[1152,1267].eqiad.wmnet with reason: Switching over ms1 * 08:22 marostegui@cumin1003: dbctl commit (dc=all): 'Depool ms1 [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95960 and previous config saved to /var/cache/conftool/dbconfig/20260811-082201-marostegui.json * 08:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: Switching over ms1 * 08:20 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.parsercache (exit_code=99) * 08:20 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 08:20 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1152: Switching over ms1 * 08:19 slyngshede@dns1004: END - running authdns-update * 08:18 moritzm: installing openjdk-21 security updates * 08:17 slyngshede@dns1004: START - running authdns-update * 08:16 moritzm: imported jenkins 2.568.2 to thirdparty/jenkins for trixie-wikimedia * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.12 (duration: 02m 26s) * 03:36 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] (duration: 33m 33s) * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 35s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-10 == * 14:54 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1323973{{!}}mmv.bootstrap: Fix getUrlParam to account for TIFF lossy/lossless param (T434333)]] (duration: 11m 24s) * 14:50 krinkle@deploy1003: krinkle: Continuing with deployment * 14:45 krinkle@deploy1003: krinkle: Backport for [[gerrit:1323973{{!}}mmv.bootstrap: Fix getUrlParam to account for TIFF lossy/lossless param (T434333)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:43 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1323973{{!}}mmv.bootstrap: Fix getUrlParam to account for TIFF lossy/lossless param (T434333)]] * 14:07 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1323967{{!}}updateIsActiveFlagForMentees: Commit the final partial batch (T432959)]] (duration: 10m 33s) * 13:56 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1323967{{!}}updateIsActiveFlagForMentees: Commit the final partial batch (T432959)]] * 13:45 dani@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply * 13:45 dani@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply * 13:45 dani@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply * 13:45 dani@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply * 13:45 dani@deploy1003: helmfile [staging] DONE helmfile.d/services/miscweb: apply * 13:44 dani@deploy1003: helmfile [staging] START helmfile.d/services/miscweb: apply * 13:38 wmde-fisch@deploy1003: Finished scap sync-world: Backport for [[gerrit:1323939{{!}}Enable sub-references on more group2 wikis (batch3) (T432731)]] (duration: 33m 21s) * 13:25 wmde-fisch@deploy1003: wmde-fisch: Continuing with deployment * 13:22 wmde-fisch@deploy1003: wmde-fisch: Backport for [[gerrit:1323939{{!}}Enable sub-references on more group2 wikis (batch3) (T432731)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:05 wmde-fisch@deploy1003: Started scap sync-world: Backport for [[gerrit:1323939{{!}}Enable sub-references on more group2 wikis (batch3) (T432731)]] * 07:57 hashar@deploy1003: Finished deploy [integration/docroot@7772132]: update build dependencies (duration: 00m 13s) * 07:57 hashar@deploy1003: Started deploy [integration/docroot@7772132]: update build dependencies * 07:35 _joe_: restarting squid on urldownloader1006 * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 48s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-09 == * 16:01 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:01 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:01 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:00 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 36s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-08 == * 05:31 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9] (wcqs): [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) (duration: 02m 36s) * 05:28 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9] (wcqs): [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) * 04:56 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 04:55 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 04:47 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) (duration: 19m 22s) * 04:28 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) * 04:19 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) (duration: 00m 06s) * 04:18 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) * 04:17 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) (duration: 00m 28s) * 04:16 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) * 03:52 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 03:52 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 34s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-07 == * 23:30 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:29 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 22:45 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 22:43 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 22:41 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 22:41 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 20:54 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 20:32 andrewbogott: restarting puppetserver service on puppetserver* for [[phab:T434339|T434339]] * 19:52 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:45 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 19:32 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:25 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:22 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 19:21 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 18:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:41 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 18:35 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 18:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 18:22 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 18:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 18:16 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 18:12 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:09 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:08 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:07 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:04 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:01 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:00 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:00 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 17:59 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 17:25 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 17:14 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 17:13 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 17:13 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 17:13 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:54 maryum: Deployed security fix for [[phab:T434278|T434278]] * 16:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 16:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 16:27 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --verbose --scoreLessThan=0.7 --exceptDatasetChecksums=[[phab:T434319|T434319]]-enwiki-models.txt # [[phab:T434319|T434319]] * 16:06 cdobbins@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp5022.eqsin.wmnet with OS trixie * 15:13 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 14:19 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 14:17 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 13:50 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1156.eqiad.wmnet onto db1271.eqiad.wmnet * 13:50 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1271: Pool db1271.eqiad.wmnet in after cloning * 13:02 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1271: Pool db1271.eqiad.wmnet in after cloning * 12:19 jayme: updated calico to v3.30.7 on staging-codfw - [[phab:T427400|T427400]] * 12:09 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 12:06 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 12:06 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 12:05 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 12:02 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1156: Pool db1156.eqiad.wmnet in after cloning * 11:38 bjensen: sudo -i reprepro -C main include trixie-wikimedia $<nowiki>{</nowiki>HOME<nowiki>}</nowiki>/httpbb/trixie/httpbb_$<nowiki>{</nowiki>VERSION?<nowiki>}</nowiki>-1+deb13u1_amd64.changes #[[phab:T434052|T434052]] * 11:35 bjensen: sudo -i reprepro -C main include bookworm-wikimedia $<nowiki>{</nowiki>HOME<nowiki>}</nowiki>/httpbb/bookworm/httpbb_$<nowiki>{</nowiki>VERSION?<nowiki>}</nowiki>-1_amd64.changes #[[phab:T434052|T434052]] * 11:30 marostegui@cumin1003: dbctl commit (dc=all): 'Adding db1271 to dbctl', diff saved to https://phabricator.wikimedia.org/P95945 and previous config saved to /var/cache/conftool/dbconfig/20260807-113006-marostegui.json * 11:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1156: Pool db1156.eqiad.wmnet in after cloning * 10:23 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 10:22 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 10:22 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 10:21 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 10:20 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 10:20 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 10:19 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 10:18 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 10:06 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on 21 hosts with reason: cloning * 10:01 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1156: Depool db1156.eqiad.wmnet to then clone it to db1271.eqiad.wmnet - marostegui@cumin1003 * 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1156: Depool db1156.eqiad.wmnet to then clone it to db1271.eqiad.wmnet - marostegui@cumin1003 * 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1156.eqiad.wmnet onto db1271.eqiad.wmnet * 09:15 jynus: started stress testing db1245 dbs [[phab:T431115|T431115]] * 08:19 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:18 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:16 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:14 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:13 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:10 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:06 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:05 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:00 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 10 days, 0:00:00 on ml-serve1015.eqiad.wmnet with reason: Downtime to get full picture of current BIOS settings beyond what Redfish shows * 08:00 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 07:54 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 07:54 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 07:53 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:52 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:51 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:50 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:49 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:48 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:47 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:45 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:45 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:41 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:38 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:37 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 06:35 jayme: updated istio to 1.29.4 on wikikube eqiad - [[phab:T427401|T427401]] * 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1178.eqiad.wmnet * 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1178.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 06:06 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1178.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 05:55 marostegui@cumin1003: START - Cookbook sre.dns.netbox * 05:49 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1178.eqiad.wmnet * 05:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 05:46 marostegui@cumin1003: Removing db1178 from zarcillo [[phab:T433471|T433471]] * 05:45 marostegui@cumin1003: START - Cookbook sre.mysql.decommission * 02:42 denisse: Extended volume on prometheus2008 for the disk space alert as per https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_host_running_out_of_space * 02:37 denisse: Extended volume on prometheus2007 tor the disk space alert as per https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_host_running_out_of_space * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 56s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-06 == * 21:39 maryum: Deploy security patch for [[phab:T433070|T433070]] * 21:29 maryum: Deploy security patch for [[phab:T434189|T434189]] * 20:48 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] (duration: 08m 12s) * 20:44 aude@deploy1003: lmora, aude, anzx: Continuing with deployment * 20:41 aude@deploy1003: lmora, aude, anzx: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be * 20:41 ebernhardson@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:41 ebernhardson@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 20:40 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] * 20:37 ebernhardson@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:37 ebernhardson@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 20:32 ebernhardson@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:32 ebernhardson@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 20:31 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] (duration: 06m 41s) * 20:27 cjming@deploy1003: cjming, ebernhardson, chlod: Continuing with deployment * 20:26 cjming@deploy1003: cjming, ebernhardson, chlod: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:24 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] * 20:18 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] (duration: 09m 22s) * 20:14 cjming@deploy1003: cjming, tsev: Continuing with deployment * 20:11 cjming@deploy1003: cjming, tsev: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:09 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] * 19:41 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply * 19:40 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply * 19:31 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 19:31 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 19:00 cdobbins@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cp5022.eqsin.wmnet with OS trixie * 18:25 ladsgroup@deploy1003: Finished scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) (duration: 06m 08s) * 18:19 ladsgroup@deploy1003: Started scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) * 18:18 ladsgroup@deploy1003: Stopping before sync operations * 18:17 ladsgroup@deploy1003: Started scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) * 17:55 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 16:50 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 16:35 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1001.eqiad.wmnet with OS bookworm * 16:19 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1002.eqiad.wmnet with reason: host reimage * 16:16 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1002.eqiad.wmnet with reason: host reimage * 16:05 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1001.eqiad.wmnet with reason: host reimage * 16:00 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1001.eqiad.wmnet with reason: host reimage * 15:57 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 15:43 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm * 15:29 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1001.eqiad.wmnet with OS bookworm * 15:29 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:58 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm * 14:57 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-drmrs ([[phab:T428495|T428495]]) * 14:55 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-drmrs ([[phab:T428495|T428495]]) * 14:55 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-ui1001.eqiad.wmnet with OS bookworm * 14:54 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-presto1001.eqiad.wmnet with OS bookworm * 14:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-magru ([[phab:T428495|T428495]]) * 14:49 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-magru ([[phab:T428495|T428495]]) * 14:48 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1001.eqiad.wmnet with OS bookworm * 14:46 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-esams ([[phab:T428495|T428495]]) * 14:44 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-esams ([[phab:T428495|T428495]]) * 14:43 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:42 brouberol@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:42 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:42 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 14:40 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 14:40 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:38 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-ui1001.eqiad.wmnet with reason: host reimage * 14:34 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-presto1001.eqiad.wmnet with reason: host reimage * 14:28 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-ui1001.eqiad.wmnet with reason: host reimage * 14:27 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-presto1001.eqiad.wmnet with reason: host reimage * 14:23 sukhe: sudo cumin -b2 'A:cp-text' "run-puppet-agent --enable 'merging CR 1290731'": [[phab:T425441|T425441]] * 14:18 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo for hosts in the wikimedia.org domain - [[phab:T428495|T428495]] * 14:16 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-presto1001.eqiad.wmnet with OS bookworm * 14:14 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-ui1001.eqiad.wmnet with OS bookworm * 14:12 sukhe: sudo cumin 'A:cp-text' "disable-puppet 'merging CR 1290731'": [[phab:T425441|T425441]] * 14:11 swfrench-wmf: restarted navtiming on webperf1003 - [[phab:T428495|T428495]] * 14:04 swfrench-wmf: begin rolling restart of confd in drmrs, eqiad, esams, magru - [[phab:T428495|T428495]] * 14:04 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm * 14:04 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:02 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-client1002.eqiad.wmnet with OS bookworm * 13:58 swfrench-wmf: authdns update to direct eqiad-associated etcd clients back to eqiad - [[phab:T428495|T428495]] * 13:58 swfrench@dns1004: END - running authdns-update * 13:56 swfrench@dns1004: START - running authdns-update * 13:49 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:44 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:31 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 13:29 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 13:26 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 13:23 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 13:22 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 13:19 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 13:18 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 13:18 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 13:17 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 13:16 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 13:13 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 13:11 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 13:09 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 13:06 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 13:06 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-client1002.eqiad.wmnet with OS bookworm * 13:05 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revision-models' for release 'main' . * 13:05 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:05 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revision-models' for release 'main' . * 13:04 brouberol@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-test-client1002.eqiad.wmnet with OS bookworm * 13:04 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 13:03 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 13:02 aikochou@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:00 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'readability' for release 'main' . * 12:59 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'readability' for release 'main' . * 12:58 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 12:57 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'logo-detection' for release 'main' . * 12:57 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'logo-detection' for release 'main' . * 12:57 aikochou@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 12:55 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:54 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:53 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 12:53 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:50 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 12:48 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 12:46 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'article-models' for release 'main' . * 12:45 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'article-models' for release 'main' . * 12:41 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . * 12:39 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . * 12:38 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-client1002.eqiad.wmnet with OS bookworm * 12:12 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply * 12:12 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply * 12:09 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:08 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 11:58 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2187: Security update * 11:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:24 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:16 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 11:15 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 11:10 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2187: Security update * 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2187.codfw.wmnet with reason: Maintenance * 10:56 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 10:56 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 10:56 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 10:56 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 10:54 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 10:53 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 10:09 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2187: Security update * 10:07 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2187: Security update * 09:39 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms2', diff saved to https://phabricator.wikimedia.org/P95929 and previous config saved to /var/cache/conftool/dbconfig/20260806-093908-marostegui.json * 09:36 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1178 from dbctl [[phab:T433471|T433471]]', diff saved to https://phabricator.wikimedia.org/P95928 and previous config saved to /var/cache/conftool/dbconfig/20260806-093632-marostegui.json * 09:33 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 09:31 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 09:30 topranks: bounce cr3-eqsin<->cr2-eqiad bgp session to disable no-prepend command * 09:20 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2253.codfw.wmnet,db1151.eqiad.wmnet with reason: cloning * 09:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1151: Cloning * 09:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:19 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 09:19 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1151: Cloning * 09:10 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 09:09 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup2003.codfw.wmnet * 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup2003.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 09:06 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup2003.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 09:03 klausman@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:02 jynus@cumin1003: START - Cookbook sre.dns.netbox * 09:02 klausman@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 08:57 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup2003.codfw.wmnet * 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup1003.eqiad.wmnet * 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:54 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms3', diff saved to https://phabricator.wikimedia.org/P95925 and previous config saved to /var/cache/conftool/dbconfig/20260806-085422-marostegui.json * 08:53 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:46 jynus@cumin1003: START - Cookbook sre.dns.netbox * 08:39 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup1003.eqiad.wmnet * 08:29 XioNoX: push pfw policy - [[phab:T434115|T434115]] * 08:14 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 08:00 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 08:00 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:58 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revision-models' for release 'main' . * 07:56 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 07:54 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'readability' for release 'main' . * 07:53 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'logo-detection' for release 'main' . * 07:51 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'llm' for release 'main' . * 07:48 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . * 07:37 jayme: updated istio to 1.29.4 on wikikube codfw - [[phab:T427401|T427401]] * 07:08 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2252.codfw.wmnet,db1153.eqiad.wmnet with reason: cloning * 07:07 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1153: Cloning * 07:07 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1153: Cloning * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 40s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-05 == * 23:24 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1009.eqiad.wmnet with OS bookworm * 23:03 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1009.eqiad.wmnet with reason: host reimage * 22:59 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1009.eqiad.wmnet with reason: host reimage * 22:43 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1009.eqiad.wmnet with OS bookworm * 22:38 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1009.eqiad.wmnet * 22:34 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1009.eqiad.wmnet * 22:25 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1008.eqiad.wmnet with OS bookworm * 22:04 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1008.eqiad.wmnet with reason: host reimage * 22:00 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1008.eqiad.wmnet with reason: host reimage * 21:48 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:47 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:46 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:44 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1008.eqiad.wmnet with OS bookworm * 21:43 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:41 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1008.eqiad.wmnet * 21:36 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1008.eqiad.wmnet * 21:14 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:12 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad * 21:12 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad * 21:10 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=eqiad * 21:08 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:07 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:07 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=eqiad * 21:04 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:03 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1006 * 21:02 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1006 * 21:00 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:56 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:56 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:55 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 20:55 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 20:55 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1007.eqiad.wmnet with OS bookworm * 20:51 vriley@cumin1003: START - Cookbook sre.dns.netbox * 20:43 ebernhardson: [[phab:T434008|T434008]]: changing cloudelastic:9643 from auto_expand_replicas to number_of_replicas * 20:34 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1007.eqiad.wmnet with reason: host reimage * 20:27 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1007.eqiad.wmnet with reason: host reimage * 20:24 cjming: end of UTC late backport window * 20:23 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] (duration: 06m 26s) * 20:18 cjming@deploy1003: cjming: Continuing with deployment * 20:18 cjming@deploy1003: cjming: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:16 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] * 20:12 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1007.eqiad.wmnet with OS bookworm * 20:12 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] (duration: 08m 41s) * 20:08 swfrench@cumin2002: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host conf1007.eqiad.wmnet with OS bookworm * 20:08 jforrester@deploy1003: jforrester: Continuing with deployment * 20:07 jforrester@deploy1003: jforrester: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:03 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] * 19:51 inflatador: [bking@puppetserver1001] ~$ sudo puppetserver ca sign --certname an-worker1189.eqiad.wmnet [[phab:T434142|T434142]] * 19:47 bking@cumin2003: DONE (FAIL) - Cookbook sre.puppet.renew-cert (exit_code=99) for an-worker1189.eqiad.wmnet: Renew puppet certificate - bking@cumin2003 * 19:46 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:30 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1007.eqiad.wmnet with OS trixie * 19:30 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 19:29 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 19:20 swfrench-wmf: silenced EtcdRelicationDown 0cb709a9-f244-4f1e-971f-{{Gerrit|440ec65e7fd7}} - [[phab:T428495|T428495]] * 19:13 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1007.eqiad.wmnet with OS bookworm * 19:12 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1007.eqiad.wmnet with reason: host reimage * 19:09 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1007.eqiad.wmnet * 19:07 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1007.eqiad.wmnet with reason: host reimage * 19:03 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1007.eqiad.wmnet * 18:52 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1007.eqiad.wmnet with OS trixie * 18:52 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1007.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:35 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 18:34 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 18:34 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 18:30 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1007.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:28 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:28 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1007] - vriley@cumin1003" * 18:27 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1007] - vriley@cumin1003" * 18:23 vriley@cumin1003: START - Cookbook sre.dns.netbox * 18:22 vriley@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 18:22 robh@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:19 vriley@cumin1003: START - Cookbook sre.dns.netbox * 18:13 robh@cumin2002: START - Cookbook sre.hosts.provision for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:31 jasmine@cumin2002: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-main-eqiad * 17:12 mutante: LDAP - added vwalters to group ciadmin - [[phab:T433615|T433615]] * 16:58 aokoth@deploy1003: Finished deploy [phabricator/deployment@e2ebca5]: Deploy Phab (duration: 00m 34s) * 16:57 aokoth@deploy1003: Started deploy [phabricator/deployment@e2ebca5]: Deploy Phab * 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad * 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=eqiad * 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad * 16:53 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:41 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2187.codfw.wmnet * 16:41 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2187.codfw.wmnet * 16:41 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker2187.codfw.wmnet * 16:41 cgoubert@cumin2003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker2187.codfw.wmnet * 16:40 jasmine@cumin2002: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-main-eqiad * 16:40 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:34 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:25 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-magru and A:liberica ([[phab:T428495|T428495]]) * 16:23 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-magru and A:liberica ([[phab:T428495|T428495]]) * 16:20 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-drmrs and A:liberica ([[phab:T428495|T428495]]) * 16:19 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-drmrs and A:liberica ([[phab:T428495|T428495]]) * 16:18 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-esams and A:liberica ([[phab:T428495|T428495]]) * 16:16 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-esams and A:liberica ([[phab:T428495|T428495]]) * 16:06 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1159.eqiad.wmnet * 16:06 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1159.eqiad.wmnet * 16:06 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1159.eqiad.wmnet * 16:05 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] (duration: 09m 11s) * 15:58 reedy@deploy1003: reedy: Continuing with deployment * 15:58 reedy@deploy1003: reedy: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:56 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] * 15:54 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1159.eqiad.wmnet with OS trixie * 15:38 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:33 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1159.eqiad.wmnet with reason: host reimage * 15:32 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:27 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1159.eqiad.wmnet with reason: host reimage * 15:10 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1159 * 15:10 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1159 * 15:00 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo for hosts in the wikimedia.org domain - [[phab:T428495|T428495]] * 14:55 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS trixie * 14:54 swfrench-wmf: restarted navtiming on webperf1003 - [[phab:T428495|T428495]] * 14:52 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1159 * 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1159.eqiad.wmnet 129.48.64.10.in-addr.arpa 9.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:52 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1159.eqiad.wmnet 129.48.64.10.in-addr.arpa 9.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1159 - jayme@cumin1003" * 14:52 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1159 - jayme@cumin1003" * 14:49 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:48 jayme@cumin1003: START - Cookbook sre.dns.netbox * 14:47 swfrench-wmf: begin rolling restart of confd in drmrs, eqiad, esams, magru - [[phab:T428495|T428495]] * 14:47 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1159 * 14:46 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:46 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:46 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1159.eqiad.wmnet with OS trixie * 14:45 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:44 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:44 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1159.eqiad.wmnet * 14:43 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:43 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1159.eqiad.wmnet * 14:43 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:43 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1159.eqiad.wmnet * 14:43 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:43 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1157.eqiad.wmnet * 14:43 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1157.eqiad.wmnet * 14:43 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1157.eqiad.wmnet * 14:42 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:42 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:42 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:41 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:41 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:41 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:41 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1003.eqiad.wmnet with OS bookworm * 14:39 swfrench-wmf: authdns update to direct eqiad-associated etcd clients to codfw - [[phab:T428495|T428495]] * 14:39 swfrench@dns1004: END - running authdns-update * 14:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 14:37 swfrench@dns1004: START - running authdns-update * 14:37 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:37 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:35 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:35 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 14:28 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:27 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1157.eqiad.wmnet with OS trixie * 14:27 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:27 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:27 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 14:26 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 14:26 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:26 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 14:26 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 14:26 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:26 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host search-loader1002.eqiad.wmnet with OS trixie * 14:19 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:19 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:15 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1003.eqiad.wmnet with reason: host reimage * 14:14 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:14 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:13 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 * 14:13 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host mc2046 * 14:13 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS trixie * 14:11 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1003.eqiad.wmnet with reason: host reimage * 14:10 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:09 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:09 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:09 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:08 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:08 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1157.eqiad.wmnet with reason: host reimage * 14:08 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:04 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 14:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on search-loader1002.eqiad.wmnet with reason: host reimage * 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=eqiad * 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=eqiad * 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=eqiad * 14:00 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:59 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:58 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1157.eqiad.wmnet with reason: host reimage * 13:57 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on search-loader1002.eqiad.wmnet with reason: host reimage * 13:54 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1003.eqiad.wmnet with OS bookworm * 13:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host search-loader1002.eqiad.wmnet with OS trixie * 13:43 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1157 * 13:42 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1157 * 13:40 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1157 * 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1157.eqiad.wmnet 183.32.64.10.in-addr.arpa 3.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1157.eqiad.wmnet 183.32.64.10.in-addr.arpa 3.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1157 - jayme@cumin1003" * 13:39 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1157 - jayme@cumin1003" * 13:39 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] (duration: 07m 00s) * 13:35 jayme@cumin1003: START - Cookbook sre.dns.netbox * 13:35 reedy@deploy1003: reedy: Continuing with deployment * 13:34 reedy@deploy1003: reedy: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:32 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] * 13:23 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1157 * 13:22 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1157.eqiad.wmnet with OS trixie * 13:22 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1157.eqiad.wmnet * 13:22 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1157.eqiad.wmnet * 13:21 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1157.eqiad.wmnet * 13:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1156.eqiad.wmnet * 13:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1156.eqiad.wmnet * 13:15 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1156.eqiad.wmnet * 13:01 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1156.eqiad.wmnet with OS trixie * 12:42 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1156.eqiad.wmnet with reason: host reimage * 12:38 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1156.eqiad.wmnet with reason: host reimage * 12:32 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 12:31 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 12:30 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 12:28 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 12:26 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 12:24 topranks: update bgp confed settings in eqsin * 12:22 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1156 * 12:22 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1156 * 12:22 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 12:19 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1156 * 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1156.eqiad.wmnet 110.32.64.10.in-addr.arpa 0.1.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:19 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1156.eqiad.wmnet 110.32.64.10.in-addr.arpa 0.1.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1156 - jayme@cumin1003" * 12:19 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1156 - jayme@cumin1003" * 12:17 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:14 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS trixie * 12:09 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:06 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:04 jayme@cumin1003: START - Cookbook sre.dns.netbox * 12:04 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 12:02 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'article-models' for release 'main' . * 12:01 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1156 * 12:01 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1156.eqiad.wmnet with OS trixie * 11:59 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1156.eqiad.wmnet * 11:59 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1156.eqiad.wmnet * 11:59 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1156.eqiad.wmnet * 11:57 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 11:53 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 11:53 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:52 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:52 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:50 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:50 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:50 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:49 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:48 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:47 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:47 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:45 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:45 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:44 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:44 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:44 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:43 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:42 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:38 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:35 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 * 11:35 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host mc2046 * 11:34 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS trixie * 11:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:27 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:21 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:21 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:18 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:18 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:18 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:18 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:13 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:13 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:09 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:08 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:07 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:06 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:06 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:05 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:05 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:04 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 11:04 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:24 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:24 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:17 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:16 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1155.eqiad.wmnet * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1155.eqiad.wmnet * 10:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1155.eqiad.wmnet * 10:14 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:14 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:11 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:11 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:10 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:09 aikochou@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop: sync * 10:09 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:09 aikochou@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop: sync * 10:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:07 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:05 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:05 aikochou@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop: sync * 10:05 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:05 aikochou@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop: sync * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:04 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:04 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:04 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1155.eqiad.wmnet with OS trixie * 09:52 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms1', diff saved to https://phabricator.wikimedia.org/P95918 and previous config saved to /var/cache/conftool/dbconfig/20260805-095212-marostegui.json * 09:44 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1152: after cloning * 09:44 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.parsercache (exit_code=99) * 09:44 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 09:44 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1152: after cloning * 09:43 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1155.eqiad.wmnet with reason: host reimage * 09:40 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1155.eqiad.wmnet with reason: host reimage * 09:32 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 09:32 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:31 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 09:31 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:27 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1155 * 09:27 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1155 * 09:25 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 09:24 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 09:24 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 09:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:23 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 09:23 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 09:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:22 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 09:22 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 09:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:20 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 09:20 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:17 XioNoX: push pfw policies - [[phab:T434038|T434038]] * 09:14 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1155 * 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1155.eqiad.wmnet 109.32.64.10.in-addr.arpa 9.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1155.eqiad.wmnet 109.32.64.10.in-addr.arpa 9.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1155 - jayme@cumin1003" * 09:14 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1155 - jayme@cumin1003" * 09:10 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 09:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:09 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2251.codfw.wmnet,db1152.eqiad.wmnet with reason: cloning * 09:09 jayme@cumin1003: START - Cookbook sre.dns.netbox * 09:08 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 09:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: Cloning * 09:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:05 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 09:05 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1152: Cloning * 08:38 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1155 * 08:37 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1155.eqiad.wmnet with OS trixie * 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1171.eqiad.wmnet * 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1171.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:29 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95913 and previous config saved to /var/cache/conftool/dbconfig/20260805-082908-ladsgroup.json * 08:27 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1171.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:22 jynus@cumin1003: START - Cookbook sre.dns.netbox * 08:18 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249', diff saved to https://phabricator.wikimedia.org/P95912 and previous config saved to /var/cache/conftool/dbconfig/20260805-081823-ladsgroup.json * 08:17 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1171.eqiad.wmnet * 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1150.eqiad.wmnet * 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1150.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:15 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1150.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:15 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 08:14 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1155.eqiad.wmnet * 08:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1155.eqiad.wmnet * 08:14 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1155.eqiad.wmnet * 08:11 jynus@cumin1003: START - Cookbook sre.dns.netbox * 08:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249', diff saved to https://phabricator.wikimedia.org/P95911 and previous config saved to /var/cache/conftool/dbconfig/20260805-080737-ladsgroup.json * 08:05 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1150.eqiad.wmnet * 08:02 marostegui: Depool clouddb1020 (s5,s8) [[phab:T434048|T434048]] * 08:02 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1020.eqiad.wmnet,service=s8 * 08:02 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1020.eqiad.wmnet,service=s5 * 08:02 marostegui: Depool clouddb1018 (s2,s7) [[phab:T434048|T434048]] * 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1018.eqiad.wmnet,service=s7 * 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1018.eqiad.wmnet,service=s2 * 08:01 marostegui: Depool clouddb1017 (s1) [[phab:T434048|T434048]] * 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1017.eqiad.wmnet,service=s1 * 07:59 marostegui: Depool clouddb1016 (s5,s8) [[phab:T434048|T434048]] * 07:59 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s8 * 07:59 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s5 * 07:57 marostegui: Depool clouddb1015 (s4,s6) [[phab:T434048|T434048]] * 07:57 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s6 * 07:57 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s4 * 07:56 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95910 and previous config saved to /var/cache/conftool/dbconfig/20260805-075650-ladsgroup.json * 07:54 marostegui: Depool clouddb1014 (s2,s7) [[phab:T434048|T434048]] * 07:54 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1014.eqiad.wmnet,service=s7 * 07:54 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1014.eqiad.wmnet,service=s2 * 07:53 marostegui: Depool clouddb1013:s1 [[phab:T434048|T434048]] * 07:53 marostegui: Depool clouddb1013:s1 [[phab:T409557|T409557]] * 07:53 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1013.eqiad.wmnet,service=s1 * 07:25 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95909 and previous config saved to /var/cache/conftool/dbconfig/20260805-072529-ladsgroup.json * 07:24 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2249.codfw.wmnet with reason: Maintenance * 07:24 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95908 and previous config saved to /var/cache/conftool/dbconfig/20260805-072426-ladsgroup.json * 07:21 slyngshede@dns1004: END - running authdns-update * 07:19 slyngshede@dns1004: START - running authdns-update * 07:13 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231', diff saved to https://phabricator.wikimedia.org/P95906 and previous config saved to /var/cache/conftool/dbconfig/20260805-071340-ladsgroup.json * 07:02 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231', diff saved to https://phabricator.wikimedia.org/P95905 and previous config saved to /var/cache/conftool/dbconfig/20260805-070253-ladsgroup.json * 06:52 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95904 and previous config saved to /var/cache/conftool/dbconfig/20260805-065206-ladsgroup.json * 06:45 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 06:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95903 and previous config saved to /var/cache/conftool/dbconfig/20260805-062240-ladsgroup.json * 06:21 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2231.codfw.wmnet with reason: Maintenance * 06:21 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95902 and previous config saved to /var/cache/conftool/dbconfig/20260805-062137-ladsgroup.json * 06:10 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215', diff saved to https://phabricator.wikimedia.org/P95901 and previous config saved to /var/cache/conftool/dbconfig/20260805-061051-ladsgroup.json * 06:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215', diff saved to https://phabricator.wikimedia.org/P95900 and previous config saved to /var/cache/conftool/dbconfig/20260805-060004-ladsgroup.json * 05:49 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95899 and previous config saved to /var/cache/conftool/dbconfig/20260805-054918-ladsgroup.json * 05:19 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95898 and previous config saved to /var/cache/conftool/dbconfig/20260805-051939-ladsgroup.json * 05:18 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2215.codfw.wmnet with reason: Maintenance * 04:30 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2201.codfw.wmnet with reason: Maintenance * 03:40 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2197.codfw.wmnet with reason: Maintenance * 03:40 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95897 and previous config saved to /var/cache/conftool/dbconfig/20260805-034036-ladsgroup.json * 03:29 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196', diff saved to https://phabricator.wikimedia.org/P95896 and previous config saved to /var/cache/conftool/dbconfig/20260805-032948-ladsgroup.json * 03:19 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196', diff saved to https://phabricator.wikimedia.org/P95895 and previous config saved to /var/cache/conftool/dbconfig/20260805-031902-ladsgroup.json * 03:08 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95894 and previous config saved to /var/cache/conftool/dbconfig/20260805-030815-ladsgroup.json * 02:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95893 and previous config saved to /var/cache/conftool/dbconfig/20260805-023413-ladsgroup.json * 02:33 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2196.codfw.wmnet with reason: Maintenance * 02:33 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95892 and previous config saved to /var/cache/conftool/dbconfig/20260805-023310-ladsgroup.json * 02:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186', diff saved to https://phabricator.wikimedia.org/P95891 and previous config saved to /var/cache/conftool/dbconfig/20260805-022223-ladsgroup.json * 02:11 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186', diff saved to https://phabricator.wikimedia.org/P95890 and previous config saved to /var/cache/conftool/dbconfig/20260805-021137-ladsgroup.json * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 02:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95889 and previous config saved to /var/cache/conftool/dbconfig/20260805-020051-ladsgroup.json * 01:30 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95888 and previous config saved to /var/cache/conftool/dbconfig/20260805-013029-ladsgroup.json * 01:29 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2186.codfw.wmnet with reason: Maintenance * 00:34 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on dbstore1009.eqiad.wmnet with reason: Maintenance * 00:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95887 and previous config saved to /var/cache/conftool/dbconfig/20260805-003408-ladsgroup.json * 00:23 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264', diff saved to https://phabricator.wikimedia.org/P95886 and previous config saved to /var/cache/conftool/dbconfig/20260805-002322-ladsgroup.json * 00:12 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264', diff saved to https://phabricator.wikimedia.org/P95885 and previous config saved to /var/cache/conftool/dbconfig/20260805-001235-ladsgroup.json * 00:01 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95884 and previous config saved to /var/cache/conftool/dbconfig/20260805-000148-ladsgroup.json == 2026-08-04 == * 23:45 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95883 and previous config saved to /var/cache/conftool/dbconfig/20260804-234508-ladsgroup.json * 23:44 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1264.eqiad.wmnet with reason: Maintenance * 23:44 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95882 and previous config saved to /var/cache/conftool/dbconfig/20260804-234405-ladsgroup.json * 23:33 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237', diff saved to https://phabricator.wikimedia.org/P95881 and previous config saved to /var/cache/conftool/dbconfig/20260804-233317-ladsgroup.json * 23:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237', diff saved to https://phabricator.wikimedia.org/P95880 and previous config saved to /var/cache/conftool/dbconfig/20260804-232230-ladsgroup.json * 23:11 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95879 and previous config saved to /var/cache/conftool/dbconfig/20260804-231144-ladsgroup.json * 22:23 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95878 and previous config saved to /var/cache/conftool/dbconfig/20260804-222345-ladsgroup.json * 22:23 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1237.eqiad.wmnet with reason: Maintenance * 21:13 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1225.eqiad.wmnet with reason: Maintenance * 20:40 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] (duration: 24m 40s) * 20:33 samtar@deploy1003: samtar, kineticpelagic: Continuing with deployment * 20:28 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS bookworm * 20:21 samtar@deploy1003: samtar, kineticpelagic: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:15 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] * 20:13 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 20:09 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 20:00 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1216.eqiad.wmnet with reason: Maintenance * 20:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95877 and previous config saved to /var/cache/conftool/dbconfig/20260804-195957-ladsgroup.json * 19:51 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 19:50 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) mc2046.codfw.wmnet 120.16.192.10.in-addr.arpa 0.2.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:50 jhancock@cumin2002: START - Cookbook sre.dns.wipe-cache mc2046.codfw.wmnet 120.16.192.10.in-addr.arpa 0.2.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host mc2046 - jhancock@cumin2002" * 19:50 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host mc2046 - jhancock@cumin2002" * 19:49 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203', diff saved to https://phabricator.wikimedia.org/P95876 and previous config saved to /var/cache/conftool/dbconfig/20260804-194911-ladsgroup.json * 19:46 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 19:45 jhancock@cumin2002: START - Cookbook sre.hosts.move-vlan for host mc2046 * 19:45 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS bookworm * 19:38 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203', diff saved to https://phabricator.wikimedia.org/P95875 and previous config saved to /var/cache/conftool/dbconfig/20260804-193825-ladsgroup.json * 19:27 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95874 and previous config saved to /var/cache/conftool/dbconfig/20260804-192738-ladsgroup.json * 19:02 mutante: gerrit ssh -p 29418 gerrit.wikimedia.org gerrit index changes {{Gerrit|1320979}} * 18:20 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 18:18 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 18:14 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 18:14 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 18:13 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 18:10 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 18:08 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 18:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95872 and previous config saved to /var/cache/conftool/dbconfig/20260804-180721-ladsgroup.json * 18:07 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 18:06 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1203.eqiad.wmnet with reason: Maintenance * 18:06 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95871 and previous config saved to /var/cache/conftool/dbconfig/20260804-180618-ladsgroup.json * 17:55 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179', diff saved to https://phabricator.wikimedia.org/P95870 and previous config saved to /var/cache/conftool/dbconfig/20260804-175531-ladsgroup.json * 17:55 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1154.eqiad.wmnet * 17:55 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1154.eqiad.wmnet * 17:55 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1154.eqiad.wmnet * 17:50 swfrench@deploy1003: Finished scap sync-world: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] (duration: 04m 05s) * 17:48 swfrench@deploy1003: swfrench: Continuing with deployment * 17:46 swfrench@deploy1003: swfrench: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:45 swfrench@deploy1003: Started scap sync-world: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] * 17:44 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179', diff saved to https://phabricator.wikimedia.org/P95869 and previous config saved to /var/cache/conftool/dbconfig/20260804-174445-ladsgroup.json * 17:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95868 and previous config saved to /var/cache/conftool/dbconfig/20260804-173359-ladsgroup.json * 17:33 swfrench@deploy1003: Finished scap sync-world: Pick up new PHP production image (duration: 28m 32s) * 17:28 aokoth@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on phab1005.eqiad.wmnet with reason: Puppet Failure * 17:05 swfrench@deploy1003: Started scap sync-world: Pick up new PHP production image * 17:00 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 17:00 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 16:54 cgoubert@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on wikikube-worker2187.codfw.wmnet with reason: Hardware issue * 16:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2187.codfw.wmnet * 16:52 mutante: gerrit2003:/var/log/apache2# ln -s /srv/gerrit/site_path/review_site/logs/ gerrit * 16:52 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2187.codfw.wmnet * 16:48 mutante: gerrit2003 - moving old apache logfiles older than 60 days from /var/log/apache2 to /srv/gerrit/site_path/review_site/logs/old/ * 16:33 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 16:32 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 16:29 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 16:29 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 16:28 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 16:28 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 16:27 dzahn@cumin1003: END (PASS) - Cookbook sre.gerrit.restart-gerrit (exit_code=0) Restarting Gerrit on gerrit2003 * 16:27 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 16:27 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95867 and previous config saved to /var/cache/conftool/dbconfig/20260804-162736-ladsgroup.json * 16:27 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 16:26 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1179.eqiad.wmnet with reason: Maintenance * 16:26 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:25 mutante: restarting gerrit - dropped outdated RSA host key * 16:25 dzahn@cumin1003: START - Cookbook sre.gerrit.restart-gerrit Restarting Gerrit on gerrit2003 * 16:24 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95866 and previous config saved to /var/cache/conftool/dbconfig/20260804-162424-ladsgroup.json * 16:24 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 16:23 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 16:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95865 and previous config saved to /var/cache/conftool/dbconfig/20260804-162236-ladsgroup.json * 16:21 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1179.eqiad.wmnet with reason: Maintenance * 16:17 swfrench-wmf: reprepro include php8.3_8.3.33-1+wmf11u1 into component/php83 for bullseye-wikimedia * 16:17 swfrench-wmf: reprepro include php8.3_8.3.33-1+wmf12u1 into component/php83 for bookworm-wikimedia * 16:11 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply * 16:10 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply * 16:10 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mobileapps: apply * 16:09 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mobileapps: apply * 16:09 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply * 16:08 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply * 16:08 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:08 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:07 aokoth@cumin1003: END (PASS) - Cookbook sre.vrts.upgrade (exit_code=0) on VRTS host vrts1003.eqiad.wmnet * 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:05 aokoth@cumin1003: START - Cookbook sre.vrts.upgrade on VRTS host vrts1003.eqiad.wmnet * 16:04 mutante: gerrit2002/gerrit1003/gerrit2003 - rm /etc/gerrit/ssh_host_rsa_key * 15:59 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:59 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:59 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:59 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:56 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 15:55 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:55 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:55 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:49 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 15:49 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:44 Raine: add php8.5 packages to component/php85 - [[phab:T432983|T432983]] * 15:39 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:33 aaron@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 15:33 aaron@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 15:29 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:19 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:19 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:16 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:16 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1154.eqiad.wmnet with OS trixie * 15:16 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:15 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:15 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:06 brennen@deploy1003: Finished deploy [phabricator/deployment@56f4ffd]: deploy phab1004 for [[phab:T433981|T433981]] (duration: 00m 43s) * 15:05 brennen@deploy1003: Started deploy [phabricator/deployment@56f4ffd]: deploy phab1004 for [[phab:T433981|T433981]] * 15:05 aaron@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 15:04 aaron@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 15:02 brennen@deploy1003: Finished deploy [phabricator/deployment@56f4ffd]: deploy phab2003 for [[phab:T433981|T433981]] (duration: 00m 51s) * 15:01 brennen@deploy1003: Started deploy [phabricator/deployment@56f4ffd]: deploy phab2003 for [[phab:T433981|T433981]] * 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1004.eqiad.wmnet with reason: deployment * 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1005.eqiad.wmnet with reason: deployment * 14:58 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab2003.codfw.wmnet with reason: deployment * 14:55 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1154.eqiad.wmnet with reason: host reimage * 14:51 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1154.eqiad.wmnet with reason: host reimage * 14:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 14:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 14:38 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync * 14:38 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync * 14:38 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync * 14:37 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync * 14:37 ottomata: roll restart eventgate-main to pick up stream config change - [[phab:T433507|T433507]] * 14:37 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-main: sync * 14:36 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-main: sync * 14:36 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1154 * 14:36 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1154 * 14:34 otto@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] (duration: 08m 39s) * 14:34 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1154 * 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1154.eqiad.wmnet 108.32.64.10.in-addr.arpa 8.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:34 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1154.eqiad.wmnet 108.32.64.10.in-addr.arpa 8.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1154 - jayme@cumin1003" * 14:34 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1154 - jayme@cumin1003" * 14:30 otto@deploy1003: otto: Continuing with deployment * 14:30 jayme@cumin1003: START - Cookbook sre.dns.netbox * 14:28 otto@deploy1003: otto: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:26 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1154 * 14:26 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1154.eqiad.wmnet with OS trixie * 14:26 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1154.eqiad.wmnet * 14:26 otto@deploy1003: Started scap sync-world: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] * 14:26 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1154.eqiad.wmnet * 14:26 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1154.eqiad.wmnet * 14:17 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 14:16 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 14:15 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 14:14 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 14:13 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 14:13 swfrench@dns1004: END - running authdns-update * 14:13 Msz2001: Finished deployments for UTC afternoon backport window * 14:13 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 14:13 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] (duration: 07m 58s) * 14:11 swfrench@dns1004: START - running authdns-update * 14:08 mszwarc@deploy1003: javiermonton, mszwarc, mpostoronca: Continuing with deployment * 14:07 mszwarc@deploy1003: javiermonton, mszwarc, mpostoronca: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] synced to the testser * 14:05 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] * 14:03 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 13:49 swfrench@cumin2002: conftool action : set/pooled=yes; selector: name=wikikube-worker2330.codfw.wmnet * 13:49 swfrench@cumin2002: conftool action : set/pooled=no; selector: name=wikikube-worker2330.codfw.wmnet * 13:48 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] (duration: 09m 19s) * 13:45 swfrench@dns1004: END - running authdns-update * 13:44 mszwarc@deploy1003: mszwarc, jforrester: Continuing with deployment * 13:43 swfrench@dns1004: START - running authdns-update * 13:41 mszwarc@deploy1003: mszwarc, jforrester: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:38 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] * 13:33 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 13:33 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1154.eqiad.wmnet * 13:32 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 13:32 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 13:31 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 13:31 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:31 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:29 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1154.eqiad.wmnet * 13:28 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1154.eqiad.wmnet * 13:28 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1154.eqiad.wmnet * 13:28 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1141.eqiad.wmnet * 13:28 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1141.eqiad.wmnet * 13:28 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1141.eqiad.wmnet * 13:22 otto@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply * 13:22 otto@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply * 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1096.eqiad.wmnet with OS trixie * 13:05 swfrench@dns1004: END - running authdns-update * 13:03 swfrench@dns1004: START - running authdns-update * 12:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 12:43 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 1:00:00 on db1171.eqiad.wmnet with reason: decom * 12:42 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 1:00:00 on db1150.eqiad.wmnet with reason: decom * 12:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 12:38 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1164,1217].eqiad.wmnet with reason: cloning * 12:33 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2096.codfw.wmnet with OS trixie * 12:22 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1096.eqiad.wmnet with OS trixie * 12:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2096.codfw.wmnet with reason: host reimage * 12:14 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1141.eqiad.wmnet with OS trixie * 12:10 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2096.codfw.wmnet with reason: host reimage * 12:10 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1289.eqiad.wmnet * 12:05 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1289.eqiad.wmnet * 12:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1288.eqiad.wmnet * 11:59 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1288.eqiad.wmnet * 11:59 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1287.eqiad.wmnet * 11:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1097.eqiad.wmnet with OS trixie * 11:54 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1287.eqiad.wmnet * 11:54 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1286.eqiad.wmnet * 11:53 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1141.eqiad.wmnet with reason: host reimage * 11:51 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2096.codfw.wmnet with OS trixie * 11:49 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1141.eqiad.wmnet with reason: host reimage * 11:48 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1286.eqiad.wmnet * 11:48 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1284.eqiad.wmnet * 11:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2095.codfw.wmnet with OS trixie * 11:43 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1284.eqiad.wmnet * 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1283.eqiad.wmnet * 11:42 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on ml-serve1015.eqiad.wmnet with reason: Downtime to get full picture of current BIOS settings beyond what Redfish shows * 11:39 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad * 11:39 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:37 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1283.eqiad.wmnet * 11:37 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1282.eqiad.wmnet * 11:37 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad * 11:37 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:33 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1141 * 11:33 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1141 * 11:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 11:32 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1141 * 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1141.eqiad.wmnet 156.48.64.10.in-addr.arpa 6.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:32 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1141.eqiad.wmnet 156.48.64.10.in-addr.arpa 6.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1141 - jayme@cumin1003" * 11:32 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1141 - jayme@cumin1003" * 11:32 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1282.eqiad.wmnet * 11:32 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1281.eqiad.wmnet * 11:32 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad * 11:32 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:29 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 11:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1097.eqiad.wmnet with reason: host reimage * 11:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2095.codfw.wmnet with OS trixie * 11:26 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1281.eqiad.wmnet * 11:26 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1280.eqiad.wmnet * 11:25 jayme@cumin1003: START - Cookbook sre.dns.netbox * 11:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1097.eqiad.wmnet with reason: host reimage * 11:22 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1141 * 11:21 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1141.eqiad.wmnet with OS trixie * 11:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1280.eqiad.wmnet * 11:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1279.eqiad.wmnet * 11:20 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin with reason: upgrade new Nokia swtiches in eqsin to SR Linux v26 * 11:17 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1141.eqiad.wmnet * 11:16 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1141.eqiad.wmnet * 11:16 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1141.eqiad.wmnet * 11:16 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be2095.codfw.wmnet with OS trixie * 11:15 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1279.eqiad.wmnet * 11:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1278.eqiad.wmnet * 11:14 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1139.eqiad.wmnet * 11:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1139.eqiad.wmnet * 11:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1070.eqiad.wmnet with OS trixie * 11:13 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1140.eqiad.wmnet * 11:13 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1140.eqiad.wmnet * 11:13 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1140.eqiad.wmnet * 11:09 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1278.eqiad.wmnet * 11:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1071.eqiad.wmnet with OS trixie * 11:05 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1097.eqiad.wmnet with OS trixie * 11:04 mvernon@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be1097.eqiad.wmnet with OS trixie * 11:02 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1140.eqiad.wmnet with OS trixie * 11:02 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1097.eqiad.wmnet with OS trixie * 11:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1096.eqiad.wmnet with OS trixie * 11:00 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1139.eqiad.wmnet * 11:00 jayme@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1139.eqiad.wmnet with OS trixie * 10:56 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1069.eqiad.wmnet with OS trixie * 10:56 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 10:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1070.eqiad.wmnet with reason: host reimage * 10:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 10:45 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1071.eqiad.wmnet with reason: host reimage * 10:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 10:41 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1096.eqiad.wmnet with OS trixie * 10:41 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1140.eqiad.wmnet with reason: host reimage * 10:39 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1071.eqiad.wmnet with reason: host reimage * 10:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1070.eqiad.wmnet with reason: host reimage * 10:38 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1139.eqiad.wmnet with reason: host reimage * 10:37 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1140.eqiad.wmnet with reason: host reimage * 10:35 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1069.eqiad.wmnet with reason: host reimage * 10:33 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1095.eqiad.wmnet with OS trixie * 10:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 10:33 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1139.eqiad.wmnet with reason: host reimage * 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1069.eqiad.wmnet with reason: host reimage * 10:23 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1140 * 10:23 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1140 * 10:23 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 10:22 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1140 * 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1140.eqiad.wmnet 155.48.64.10.in-addr.arpa 5.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:21 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1071.eqiad.wmnet with OS trixie * 10:21 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1140.eqiad.wmnet 155.48.64.10.in-addr.arpa 5.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1140 - jayme@cumin1003" * 10:21 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1140 - jayme@cumin1003" * 10:21 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1071 * 10:21 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1070.eqiad.wmnet with OS trixie * 10:21 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1070 * 10:20 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 10:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1095.eqiad.wmnet with OS trixie * 10:17 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1139 * 10:17 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1139 * 10:17 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be1095.eqiad.wmnet with OS trixie * 10:15 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1139 * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1139.eqiad.wmnet 194.32.64.10.in-addr.arpa 4.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:15 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1139.eqiad.wmnet 194.32.64.10.in-addr.arpa 4.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1139 - jayme@cumin1003" * 10:15 jayme@cumin1003: START - Cookbook sre.dns.netbox * 10:15 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1139 - jayme@cumin1003" * 10:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2095.codfw.wmnet with OS trixie * 10:13 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1069.eqiad.wmnet with OS trixie * 10:12 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1069 * 10:11 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1140 * 10:11 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1140.eqiad.wmnet with OS trixie * 10:11 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1140.eqiad.wmnet * 10:10 jayme@cumin1003: START - Cookbook sre.dns.netbox * 10:10 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1139 * 10:10 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1140.eqiad.wmnet * 10:10 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1140.eqiad.wmnet * 10:10 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1139.eqiad.wmnet with OS trixie * 10:09 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1139.eqiad.wmnet * 10:08 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1139.eqiad.wmnet * 10:08 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1139.eqiad.wmnet * 10:01 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2094.codfw.wmnet with OS trixie * 09:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 09:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 09:53 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 09:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 09:44 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:44 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2094.codfw.wmnet with reason: host reimage * 09:34 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2094.codfw.wmnet with reason: host reimage * 09:34 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1071 * 09:33 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1070 * 09:33 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1095.eqiad.wmnet with OS trixie * 09:27 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1069 * 09:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:23 brouberol@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM archiva1002.wikimedia.org * 09:20 brouberol@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM archiva1002.wikimedia.org * 09:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1277.eqiad.wmnet * 09:13 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2094.codfw.wmnet with OS trixie * 09:13 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 09:12 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 09:12 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:12 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1277.eqiad.wmnet * 09:12 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1276.eqiad.wmnet * 09:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1094.eqiad.wmnet with OS trixie * 09:06 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1276.eqiad.wmnet * 09:06 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1275.eqiad.wmnet * 09:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2093.codfw.wmnet with OS trixie * 09:01 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1275.eqiad.wmnet * 09:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1274.eqiad.wmnet * 08:56 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1274.eqiad.wmnet * 08:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1273.eqiad.wmnet * 08:50 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1273.eqiad.wmnet * 08:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1272.eqiad.wmnet * 08:49 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:49 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1094.eqiad.wmnet with reason: host reimage * 08:45 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1272.eqiad.wmnet * 08:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1094.eqiad.wmnet with reason: host reimage * 08:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2093.codfw.wmnet with reason: host reimage * 08:38 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:38 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2093.codfw.wmnet with reason: host reimage * 08:35 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 08:34 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 08:29 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:28 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:26 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1271.eqiad.wmnet * 08:23 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1094.eqiad.wmnet with OS trixie * 08:21 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 08:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1271.eqiad.wmnet * 08:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1270.eqiad.wmnet * 08:15 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1270.eqiad.wmnet * 08:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1269.eqiad.wmnet * 08:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2093.codfw.wmnet with OS trixie * 08:09 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1269.eqiad.wmnet * 08:09 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1268.eqiad.wmnet * 08:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2092.codfw.wmnet with OS trixie * 08:04 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1268.eqiad.wmnet * 08:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1267.eqiad.wmnet * 07:59 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1267.eqiad.wmnet * 07:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1093.eqiad.wmnet with OS trixie * 07:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1266.eqiad.wmnet * 07:56 jynus: running extra backups to test db1285 [[phab:T433826|T433826]] * 07:51 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1266.eqiad.wmnet * 07:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2092.codfw.wmnet with reason: host reimage * 07:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1093.eqiad.wmnet with reason: host reimage * 07:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2092.codfw.wmnet with reason: host reimage * 07:32 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1093.eqiad.wmnet with reason: host reimage * 07:29 jynus: running extra backups to test db1265 [[phab:T433825|T433825]] * 07:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2092.codfw.wmnet with OS trixie * 07:11 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1093.eqiad.wmnet with OS trixie * 06:50 slyngshede@dns1004: END - running authdns-update * 06:48 slyngshede@dns1004: START - running authdns-update * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.11 (duration: 02m 29s) * 03:38 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] (duration: 32m 57s) * 03:23 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 03:22 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 03:05 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 32s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 00:45 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] (duration: 06m 20s) * 00:41 cjming@deploy1003: cjming: Continuing with deployment * 00:41 cjming@deploy1003: cjming: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:39 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] == 2026-08-03 == * 23:58 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply * 23:57 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply * 23:29 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cp5021.eqsin.wmnet * 23:29 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cp5021.eqsin.wmnet * 23:27 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cp5021.eqsin.wmnet * 23:26 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cp5021.eqsin.wmnet * 23:18 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 23:17 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 22:56 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: sync * 22:56 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: sync * 22:36 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 22:36 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 22:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host search-loader2002.codfw.wmnet with OS trixie * 21:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on search-loader2002.codfw.wmnet with reason: host reimage * 21:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on search-loader2002.codfw.wmnet with reason: host reimage * 21:42 dancy@deploy1003: Stopping before sync operations * 21:41 dancy@deploy1003: Started scap sync-world: testing * 21:39 dancy@deploy1003: Installation of scap version "4.277.0" completed for 3 hosts * 21:37 dancy@deploy1003: Installing scap version "4.277.0" for 3 host(s) * 21:37 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] (duration: 06m 13s) * 21:33 dancy@deploy1003: dancy: Continuing with deployment * 21:32 dancy@deploy1003: dancy: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:31 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] * 21:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host search-loader2002.codfw.wmnet with OS trixie * 21:03 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] (duration: 06m 34s) * 20:59 dancy@deploy1003: dancy: Continuing with deployment * 20:58 dancy@deploy1003: dancy: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:56 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] * 20:52 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] (duration: 06m 23s) * 20:48 cjming@deploy1003: cjming: Continuing with deployment * 20:47 cjming@deploy1003: cjming: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:46 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] * 20:42 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] (duration: 07m 36s) * 20:38 arlolra@deploy1003: arlolra: Continuing with deployment * 20:36 arlolra@deploy1003: arlolra: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:34 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] * 20:16 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] (duration: 08m 26s) * 20:12 krinkle@deploy1003: krinkle: Continuing with deployment * 20:09 krinkle@deploy1003: krinkle: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] * 19:45 jasmine@cumin2002: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-main-codfw * 18:58 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] (duration: 09m 23s) * 18:53 krinkle@deploy1003: krinkle: Continuing with deployment * 18:53 jasmine@cumin2002: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-main-codfw * 18:50 krinkle@deploy1003: krinkle: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:48 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] * 18:37 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] (duration: 10m 13s) * 18:34 dzahn@cumin2002: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host codesearch2001.codfw.wmnet * 18:34 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host codesearch2001.codfw.wmnet with OS trixie * 18:33 krinkle@deploy1003: krinkle: Continuing with deployment * 18:29 krinkle@deploy1003: krinkle: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:27 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] * 18:19 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 18:18 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on codesearch2001.codfw.wmnet with reason: host reimage * 18:14 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 18:14 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:12 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on codesearch2001.codfw.wmnet with reason: host reimage * 18:11 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 18:11 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 18:10 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 18:02 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 18:02 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 18:01 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 18:01 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 17:55 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host codesearch2001.codfw.wmnet with OS trixie * 17:54 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:54 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) codesearch2001.codfw.wmnet on all recursors * 17:53 dzahn@cumin2002: START - Cookbook sre.dns.wipe-cache codesearch2001.codfw.wmnet on all recursors * 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:48 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:41 dzahn@cumin2002: START - Cookbook sre.dns.netbox * 17:41 dzahn@cumin2002: START - Cookbook sre.ganeti.makevm for new host codesearch2001.codfw.wmnet * 17:37 dzahn@cumin2002: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host codesearch1001.eqiad.wmnet * 17:37 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host codesearch1001.eqiad.wmnet with OS trixie * 17:24 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on codesearch1001.eqiad.wmnet with reason: host reimage * 17:17 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on codesearch1001.eqiad.wmnet with reason: host reimage * 17:08 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host codesearch1001.eqiad.wmnet with OS trixie * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 17:06 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:06 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) codesearch1001.eqiad.wmnet on all recursors * 17:06 dzahn@cumin2002: START - Cookbook sre.dns.wipe-cache codesearch1001.eqiad.wmnet on all recursors * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 17:05 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 17:04 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:04 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 16:58 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 16:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2091.codfw.wmnet with OS trixie * 16:54 ebernhardson@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 16:54 ebernhardson@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 16:49 ebernhardson@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 16:49 ebernhardson@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 16:46 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1092.eqiad.wmnet with OS trixie * 16:43 ebernhardson@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 16:43 ebernhardson@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 16:43 dzahn@cumin2002: START - Cookbook sre.dns.netbox * 16:43 dzahn@cumin2002: START - Cookbook sre.ganeti.makevm for new host codesearch1001.eqiad.wmnet * 16:41 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 16:41 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 16:40 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2091.codfw.wmnet with reason: host reimage * 16:37 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 16:35 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2091.codfw.wmnet with reason: host reimage * 16:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1092.eqiad.wmnet with reason: host reimage * 16:24 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 16:24 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 16:23 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1092.eqiad.wmnet with reason: host reimage * 16:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2091.codfw.wmnet with OS trixie * 16:03 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1092.eqiad.wmnet with OS trixie * 16:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2090.codfw.wmnet with OS trixie * 15:51 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 15:51 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 15:51 jiji@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 15:50 jiji@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 15:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2090.codfw.wmnet with reason: host reimage * 15:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2090.codfw.wmnet with reason: host reimage * 15:33 jhathaway@dns1004: END - running authdns-update * 15:31 jhathaway@dns1004: START - running authdns-update * 15:26 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1091.eqiad.wmnet with OS trixie * 15:25 dancy@deploy1003: Installation of scap version "4.276.1" completed for 3 hosts * 15:23 dancy@deploy1003: Installing scap version "4.276.1" for 3 host(s) * 15:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2090.codfw.wmnet with OS trixie * 15:12 marostegui@cumin1003: dbctl commit (dc=all): 'Repool db2245, db2246, db2247 and db2248 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95857 and previous config saved to /var/cache/conftool/dbconfig/20260803-151212-marostegui.json * 15:09 dancy@deploy1003: Started scap sync-world: testing * 15:09 dancy@deploy1003: Installation of scap version "4.277.0" completed for 3 hosts * 15:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1091.eqiad.wmnet with reason: host reimage * 15:07 dancy@deploy1003: Installing scap version "4.277.0" for 3 host(s) * 15:03 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1091.eqiad.wmnet with reason: host reimage * 14:49 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1091.eqiad.wmnet with OS trixie * 14:33 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2089.codfw.wmnet with OS trixie * 14:29 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 14:27 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 14:18 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 14:16 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 14:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2089.codfw.wmnet with reason: host reimage * 14:10 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 14:10 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 14:09 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2089.codfw.wmnet with reason: host reimage * 13:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2089.codfw.wmnet with OS trixie * 13:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2088.codfw.wmnet with OS trixie * 13:40 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1090.eqiad.wmnet with OS trixie * 13:22 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1090.eqiad.wmnet with reason: host reimage * 13:22 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] (duration: 14m 34s) * 13:19 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1090.eqiad.wmnet with reason: host reimage * 13:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2088.codfw.wmnet with reason: host reimage * 13:16 aude@deploy1003: aude, mhorsey: Continuing with deployment * 13:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2088.codfw.wmnet with reason: host reimage * 13:12 aude@deploy1003: aude, mhorsey: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] * 13:05 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1090.eqiad.wmnet with OS trixie * 12:58 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2088.codfw.wmnet with OS trixie * 12:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db[2245-2247].codfw.wmnet * 12:49 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2247: Rebooting db2247.codfw.wmnet * 12:49 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2247: Rebooting db2247.codfw.wmnet * 12:42 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2246: Rebooting db2246.codfw.wmnet * 12:42 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2246: Rebooting db2246.codfw.wmnet * 12:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2087.codfw.wmnet with OS trixie * 12:37 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1089.eqiad.wmnet with OS trixie * 12:34 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2245: Rebooting db2245.codfw.wmnet * 12:34 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2245: Rebooting db2245.codfw.wmnet * 12:34 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db[2245-2247].codfw.wmnet * 12:32 kamila@deploy1003: Finished scap sync-world: rebuild after base image update (duration: 30m 26s) * 12:28 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 12:22 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2087.codfw.wmnet with reason: host reimage * 12:19 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1089.eqiad.wmnet with reason: host reimage * 12:14 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2087.codfw.wmnet with reason: host reimage * 12:14 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1089.eqiad.wmnet with reason: host reimage * 12:03 kamila@deploy1003: Started scap sync-world: rebuild after base image update * 12:00 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1089.eqiad.wmnet with OS trixie * 12:00 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2087.codfw.wmnet with OS trixie * 11:35 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db[2245-2248].codfw.wmnet * 11:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db[2245-2248].codfw.wmnet * 11:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2086.codfw.wmnet with OS trixie * 11:26 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 11:26 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 11:25 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 11:25 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 11:24 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1088.eqiad.wmnet with OS trixie * 11:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db[2245-2248].codfw.wmnet with reason: Checking network * 11:21 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 11:20 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 11:19 marostegui@dns1004: END - running authdns-update * 11:17 marostegui@dns1004: START - running authdns-update * 11:10 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:10 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 11:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2086.codfw.wmnet with reason: host reimage * 11:09 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:08 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 11:08 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:07 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 11:07 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:07 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 11:06 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop: apply * 11:06 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop: apply * 11:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1088.eqiad.wmnet with reason: host reimage * 11:05 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop: apply * 11:04 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop: apply * 11:04 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop: apply * 11:04 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop: apply * 11:02 marostegui@dns1004: END - running authdns-update * 11:02 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2086.codfw.wmnet with reason: host reimage * 11:01 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1088.eqiad.wmnet with reason: host reimage * 11:00 marostegui@dns1004: START - running authdns-update * 10:53 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] (duration: 10m 57s) * 10:51 cmooney@dns3003: END - running authdns-update * 10:49 cmooney@dns3003: START - running authdns-update * 10:47 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1088.eqiad.wmnet with OS trixie * 10:47 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2086.codfw.wmnet with OS trixie * 10:47 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 10:46 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:46 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:46 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new reverse ranges for eqsin CR switch links - cmooney@cumin1003" * 10:46 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new reverse ranges for eqsin CR switch links - cmooney@cumin1003" * 10:42 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] * 10:41 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 10:36 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2245, db2246 and db2247 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95855 and previous config saved to /var/cache/conftool/dbconfig/20260803-103652-marostegui.json * 10:35 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2248 from s4 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95854 and previous config saved to /var/cache/conftool/dbconfig/20260803-103535-marostegui.json * 10:27 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 10:27 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 10:26 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 10:24 kart_: cxserver: Add referencePunctuation config ([[phab:T97231|T97231]]) * 10:24 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 10:23 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:23 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:23 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:22 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:22 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply * 10:21 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply * 10:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2085.codfw.wmnet with OS trixie * 10:20 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply * 10:20 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply * 10:18 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply * 10:18 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply * 10:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1087.eqiad.wmnet with OS trixie * 09:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2085.codfw.wmnet with reason: host reimage * 09:43 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1087.eqiad.wmnet with reason: host reimage * 09:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2085.codfw.wmnet with reason: host reimage * 09:40 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1087.eqiad.wmnet with reason: host reimage * 09:26 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1087.eqiad.wmnet with OS trixie * 09:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2085.codfw.wmnet with OS trixie * 09:13 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2084.codfw.wmnet with OS trixie * 09:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1086.eqiad.wmnet with OS trixie * 08:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2084.codfw.wmnet with reason: host reimage * 08:50 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2084.codfw.wmnet with reason: host reimage * 08:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1086.eqiad.wmnet with reason: host reimage * 08:39 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1086.eqiad.wmnet with reason: host reimage * 08:38 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:38 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:37 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 08:37 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 08:35 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2084.codfw.wmnet with OS trixie * 08:34 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 08:34 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:27 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1086.eqiad.wmnet with OS trixie * 08:09 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1218: Repool after a crash * 08:07 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2083.codfw.wmnet with OS trixie * 08:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1085.eqiad.wmnet with OS trixie * 07:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2083.codfw.wmnet with reason: host reimage * 07:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1085.eqiad.wmnet with reason: host reimage * 07:40 kart_: Updated cxsever to 2026-07-16-140518-production ([[phab:T97231|T97231]]) * 07:39 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply * 07:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2083.codfw.wmnet with reason: host reimage * 07:38 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply * 07:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1085.eqiad.wmnet with reason: host reimage * 07:37 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] (duration: 32m 40s) * 07:33 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply * 07:33 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply * 07:25 jdlrobson@deploy1003: jdlrobson: Continuing with deployment * 07:24 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2083.codfw.wmnet with OS trixie * 07:24 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1085.eqiad.wmnet with OS trixie * 07:23 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1218: Repool after a crash * 07:21 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:09 marostegui: Drop renamed tables [[phab:T425074|T425074]] * 07:04 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] * 06:55 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply * 06:54 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 46s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-02 == * 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 01m 03s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-01 == * 03:30 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:30 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:30 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:30 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 34s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-31 == * 17:41 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 17:41 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 17:40 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 17:40 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 15:33 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2195: Testing * 15:02 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:02 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 15:02 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 14:48 pt1979@cumin2002: START - Cookbook sre.dns.netbox * 14:47 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2195: Testing * 14:22 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 14:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2195: Testing * 14:21 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 14:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2195.codfw.wmnet with reason: Testing * 14:16 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 14:04 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 14:04 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 13:30 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1048.eqiad.wmnet with OS trixie * 13:22 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2195: Testing * 13:22 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 13:19 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2195: Testing * 13:18 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 13:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2195: Testing * 13:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 13:05 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 13:05 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 13:04 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 13:04 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 12:50 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lswtest-d8-eqiad * 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:53 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:42 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 6515 * 11:37 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 6515 * 11:28 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:27 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:07 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2082.codfw.wmnet with OS trixie * 10:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2082.codfw.wmnet with reason: host reimage * 10:42 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2082.codfw.wmnet with reason: host reimage * 10:28 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2082.codfw.wmnet with OS trixie * 10:02 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 09:52 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 09:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts * 09:16 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts * 08:57 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 08:46 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:42 gkyziridis@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 08:37 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 08:37 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 08:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 08:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 08:11 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:11 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:08 filippo@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudvirt1048 * 08:07 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 08:07 filippo@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudvirt1048 * 08:06 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 08:01 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 08:00 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 07:19 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1048.eqiad.wmnet with reason: host reimage * 07:13 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1048.eqiad.wmnet with reason: host reimage * 07:11 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 07:11 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 07:09 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 07:09 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 06:57 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:56 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:48 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:48 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:44 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie * 06:34 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1048.eqiad.wmnet with OS trixie * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 54s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 00:57 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] (duration: 11m 04s) * 00:53 dreamyjazz@deploy1003: dreamyjazz, jforrester: Continuing with deployment * 00:48 dreamyjazz@deploy1003: dreamyjazz, jforrester: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:46 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] == 2026-07-30 == * 21:37 dancy@deploy1003: Installation of scap version "4.276.1" completed for 3 hosts * 21:35 dancy@deploy1003: Installing scap version "4.276.1" for 3 host(s) * 21:24 dancy@deploy1003: Installation of scap version "4.276.0" completed for 3 hosts * 21:22 dancy@deploy1003: Installing scap version "4.276.0" for 3 host(s) * 21:15 maryum: Deployed security fix for [[phab:T430601|T430601]] * 20:13 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] (duration: 09m 20s) * 20:07 arlolra@deploy1003: osleger, arlolra: Continuing with deployment * 20:05 arlolra@deploy1003: osleger, arlolra: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:03 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] * 19:29 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:29 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:25 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service * 19:24 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 19:24 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:24 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:24 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 19:23 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service * 19:20 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1084.eqiad.wmnet with OS trixie * 18:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1084.eqiad.wmnet with reason: host reimage * 18:52 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1084.eqiad.wmnet with reason: host reimage * 18:41 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 18:40 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 18:39 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1084.eqiad.wmnet with OS trixie * 18:25 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 18:15 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 17:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1083.eqiad.wmnet with OS trixie * 17:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1048: Maintenance * 17:36 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new security plugin settings - bking@cumin2003 - [[phab:T350516|T350516]] * 17:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1083.eqiad.wmnet with reason: host reimage * 17:26 inflatador: bking@apt1002 `reprepro --noskipold --component thirdparty/opensearch3 update trixie-wikimedia` [[phab:T433624|T433624]] * 17:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1083.eqiad.wmnet with reason: host reimage * 17:23 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 17:20 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 17:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2081.codfw.wmnet with OS trixie * 17:11 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new security plugin settings - bking@cumin2003 - [[phab:T350516|T350516]] * 17:10 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1083.eqiad.wmnet with OS trixie * 16:55 root@cumin1003: START - Cookbook sre.mysql.pool pool es1048: Maintenance * 16:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2081.codfw.wmnet with reason: host reimage * 16:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1048 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95833 and previous config saved to /var/cache/conftool/dbconfig/20260730-165053-cwilliams.json * 16:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1048.eqiad.wmnet with reason: Maintenance * 16:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1040: Maintenance * 16:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2081.codfw.wmnet with reason: host reimage * 16:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1082.eqiad.wmnet with OS trixie * 16:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2081.codfw.wmnet with OS trixie * 16:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2097.codfw.wmnet with OS trixie * 16:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1082.eqiad.wmnet with reason: host reimage * 16:08 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1082.eqiad.wmnet with reason: host reimage * 16:08 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 16:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1047: Maintenance * 16:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2080.codfw.wmnet with OS trixie * 16:04 root@cumin1003: START - Cookbook sre.mysql.pool pool es1040: Maintenance * 16:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1040: Maintenance * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new logging settings - bking@cumin2003 - [[phab:T324335|T324335]] * 15:58 root@cumin1003: START - Cookbook sre.mysql.pool pool es1040: Maintenance * 15:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1040 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95827 and previous config saved to /var/cache/conftool/dbconfig/20260730-155324-cwilliams.json * 15:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1040.eqiad.wmnet with reason: Maintenance * 15:50 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1082.eqiad.wmnet with OS trixie * 15:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2048: Maintenance * 15:44 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 24s) * 15:43 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2080.codfw.wmnet with reason: host reimage * 15:38 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new logging settings - bking@cumin2003 - [[phab:T324335|T324335]] * 15:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2080.codfw.wmnet with reason: host reimage * 15:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 15:30 mvernon@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be2097.codfw.wmnet with OS trixie * 15:23 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1081.eqiad.wmnet with OS trixie * 15:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2098.codfw.wmnet with OS trixie * 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - mvernon@cumin2003" * 15:18 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be2097.codfw.wmnet with OS trixie * 15:18 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - mvernon@cumin2003" * 15:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 15:17 root@cumin1003: START - Cookbook sre.mysql.pool pool es1047: Maintenance * 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2080.codfw.wmnet with OS trixie * 15:13 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be2097.codfw.wmnet with OS trixie * 15:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1047 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95820 and previous config saved to /var/cache/conftool/dbconfig/20260730-151200-cwilliams.json * 15:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1047.eqiad.wmnet with reason: Maintenance * 15:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: Maintenance * 15:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 15:04 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 15:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1081.eqiad.wmnet with reason: host reimage * 15:00 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1081.eqiad.wmnet with reason: host reimage * 15:00 root@cumin1003: START - Cookbook sre.mysql.pool pool es2048: Maintenance * 15:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 14:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2079.codfw.wmnet with OS trixie * 14:56 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 14:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2048 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95816 and previous config saved to /var/cache/conftool/dbconfig/20260730-145510-cwilliams.json * 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2048.codfw.wmnet with reason: Maintenance * 14:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2040: Maintenance * 14:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 14:51 tchin@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] (duration: 06m 48s) * 14:47 tchin@deploy1003: jforrester, tchin: Continuing with deployment * 14:47 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 14:47 tchin@deploy1003: jforrester, tchin: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:45 tchin@deploy1003: Started scap sync-world: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] * 14:42 sukhe@puppetserver1001: conftool action : set/weight=1; selector: cluster=urldownloader,service=squid * 14:42 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader,service=squid * 14:42 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1081.eqiad.wmnet with OS trixie * 14:39 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 14:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2079.codfw.wmnet with reason: host reimage * 14:36 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2098.codfw.wmnet with OS trixie * 14:32 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2079.codfw.wmnet with reason: host reimage * 14:30 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] (duration: 06m 31s) * 14:27 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 14:26 mszwarc@deploy1003: mszwarc: Continuing with deployment * 14:25 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:25 root@cumin1003: START - Cookbook sre.mysql.pool pool es1038: Maintenance * 14:25 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1038: Maintenance * 14:23 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] * 14:21 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] (duration: 11m 19s) * 14:20 root@cumin1003: START - Cookbook sre.mysql.pool pool es1038: Maintenance * 14:14 stran@deploy1003: stran: Continuing with deployment * 14:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1038 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95810 and previous config saved to /var/cache/conftool/dbconfig/20260730-141439-cwilliams.json * 14:14 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1038.eqiad.wmnet with reason: Maintenance * 14:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1036: Maintenance * 14:13 stran@deploy1003: stran: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2079.codfw.wmnet with OS trixie * 14:09 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] * 14:08 root@cumin1003: START - Cookbook sre.mysql.pool pool es2040: Maintenance * 14:08 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2040: Maintenance * 14:03 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] (duration: 31m 41s) * 14:03 root@cumin1003: START - Cookbook sre.mysql.pool pool es2040: Maintenance * 14:03 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2040 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95806 and previous config saved to /var/cache/conftool/dbconfig/20260730-135643-cwilliams.json * 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2040.codfw.wmnet with reason: Maintenance * 13:56 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:56 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2038: Maintenance * 13:55 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:52 lucaswerkmeister-wmde@deploy1003: migr, lucaswerkmeister-wmde: Continuing with deployment * 13:49 lucaswerkmeister-wmde@deploy1003: migr, lucaswerkmeister-wmde: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:49 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:48 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 13:45 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 13:32 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:32 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] * 13:28 root@cumin1003: START - Cookbook sre.mysql.pool pool es1036: Maintenance * 13:28 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1036: Maintenance * 13:22 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:22 root@cumin1003: START - Cookbook sre.mysql.pool pool es1036: Maintenance * 13:20 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1036 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95800 and previous config saved to /var/cache/conftool/dbconfig/20260730-131727-cwilliams.json * 13:17 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1036.eqiad.wmnet with reason: Maintenance * 13:17 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] (duration: 10m 31s) * 13:16 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2022\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 13:13 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, stran: Continuing with deployment * 13:10 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:10 root@cumin1003: START - Cookbook sre.mysql.pool pool es2038: Maintenance * 13:10 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2038: Maintenance * 13:08 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, stran: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie * 13:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2047: Maintenance * 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048 cloud-private - filippo@cumin1003" * 13:07 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048 cloud-private - filippo@cumin1003" * 13:06 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] * 13:04 root@cumin1003: START - Cookbook sre.mysql.pool pool es2038: Maintenance * 13:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2078.codfw.wmnet with OS trixie * 13:01 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2038 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95797 and previous config saved to /var/cache/conftool/dbconfig/20260730-125919-cwilliams.json * 12:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2038.codfw.wmnet with reason: Maintenance * 12:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2078.codfw.wmnet with reason: host reimage * 12:37 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2078.codfw.wmnet with reason: host reimage * 12:37 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] (duration: 06m 51s) * 12:33 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 12:32 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 12:32 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 12:32 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:30 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] * 12:19 root@cumin1003: START - Cookbook sre.mysql.pool pool es2047: Maintenance * 12:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2078.codfw.wmnet with OS trixie * 12:18 dcausse@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 12:18 dcausse@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 12:15 dcausse@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 12:14 dcausse@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 12:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2047 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95793 and previous config saved to /var/cache/conftool/dbconfig/20260730-121404-cwilliams.json * 12:13 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2047.codfw.wmnet with reason: Maintenance * 12:13 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2036: Maintenance * 12:05 ayounsi@dns1004: END - running authdns-update * 12:02 ayounsi@dns1004: START - running authdns-update * 11:51 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2077.codfw.wmnet with OS trixie * 11:48 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:46 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1080.eqiad.wmnet with OS trixie * 11:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1226: Maintenance * 11:41 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2077.codfw.wmnet with reason: host reimage * 11:28 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2077.codfw.wmnet with reason: host reimage * 11:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1080.eqiad.wmnet with reason: host reimage * 11:27 root@cumin1003: START - Cookbook sre.mysql.pool pool es2036: Maintenance * 11:27 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2036: Maintenance * 11:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1080.eqiad.wmnet with reason: host reimage * 11:21 root@cumin1003: START - Cookbook sre.mysql.pool pool es2036: Maintenance * 11:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2036 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95786 and previous config saved to /var/cache/conftool/dbconfig/20260730-111633-cwilliams.json * 11:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2036.codfw.wmnet with reason: Maintenance * 11:08 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2077.codfw.wmnet with OS trixie * 11:07 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 11:03 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie * 11:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1226: Maintenance * 10:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1226 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95783 and previous config saved to /var/cache/conftool/dbconfig/20260730-104801-cwilliams.json * 10:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1226.eqiad.wmnet with reason: Maintenance * 10:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1214: Maintenance * 10:27 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2035: Maintenance * 10:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2076.codfw.wmnet with OS trixie * 10:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1214: Maintenance * 09:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1214 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95775 and previous config saved to /var/cache/conftool/dbconfig/20260730-095451-cwilliams.json * 09:54 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1214.eqiad.wmnet with reason: Maintenance * 09:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1209: Maintenance * 09:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2076.codfw.wmnet with reason: host reimage * 09:42 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool es2035: Maintenance * 09:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.netbox.update-extras (exit_code=0) rolling restart_daemons on A:netbox * 09:41 ayounsi@cumin1003: START - Cookbook sre.netbox.update-extras rolling restart_daemons on A:netbox * 09:40 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2035: Maintenance * 09:39 ayounsi@cumin1003: END (PASS) - Cookbook sre.netbox.update-extras (exit_code=0) rolling restart_daemons on A:netbox-canary * 09:39 ayounsi@cumin1003: START - Cookbook sre.netbox.update-extras rolling restart_daemons on A:netbox-canary * 09:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2076.codfw.wmnet with reason: host reimage * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 09:34 root@cumin1003: START - Cookbook sre.mysql.pool pool es2035: Maintenance * 09:32 lucaswerkmeister-wmde@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 09:32 lucaswerkmeister-wmde@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 09:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2035 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95771 and previous config saved to /var/cache/conftool/dbconfig/20260730-092910-cwilliams.json * 09:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2035.codfw.wmnet with reason: Maintenance * 09:19 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2076.codfw.wmnet with OS trixie * 09:18 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 09:17 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie * 09:07 root@cumin1003: START - Cookbook sre.mysql.pool pool db1209: Maintenance * 09:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 23 hosts * 09:04 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Remove cable label from interfaces descriptions - ayounsi@cumin1003 * 09:04 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:02 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Remove cable label from interfaces descriptions - ayounsi@cumin1003 * 09:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1209 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95767 and previous config saved to /var/cache/conftool/dbconfig/20260730-090133-cwilliams.json * 09:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1209.eqiad.wmnet with reason: Maintenance * 09:01 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1192: Maintenance * 08:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1252: Maintenance * 08:57 jayme@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 08:56 jayme@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 08:53 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 23 hosts * 08:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:51 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:50 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1263: Maintenance * 08:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:23 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 08:15 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie * 08:14 root@cumin1003: START - Cookbook sre.mysql.pool pool db1192: Maintenance * 08:13 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2075.codfw.wmnet with OS trixie * 08:12 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1252: Maintenance * 08:11 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1252: Maintenance * 08:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1252: Maintenance * 08:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1192 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95754 and previous config saved to /var/cache/conftool/dbconfig/20260730-080611-cwilliams.json * 08:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1192.eqiad.wmnet with reason: Maintenance * 08:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1178: Maintenance * 08:05 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS bullseye * 07:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1263: Maintenance * 07:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1263 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95751 and previous config saved to /var/cache/conftool/dbconfig/20260730-075106-cwilliams.json * 07:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[1260-1262].eqiad.wmnet with reason: Maintenance * 07:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2075.codfw.wmnet with reason: host reimage * 07:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1263.eqiad.wmnet with reason: Maintenance * 07:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2075.codfw.wmnet with reason: host reimage * 07:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance * 07:38 dcausse@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:38 dcausse@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 07:35 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db1252', diff saved to https://phabricator.wikimedia.org/P95748 and previous config saved to /var/cache/conftool/dbconfig/20260730-073510-marostegui.json * 07:26 klausman@dns2004: END - running authdns-update * 07:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 07:25 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2075.codfw.wmnet with OS trixie * 07:24 klausman@dns2004: START - running authdns-update * 07:23 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host an-test-master1003.eqiad.wmnet * 07:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db1178: Maintenance * 07:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1178 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95746 and previous config saved to /var/cache/conftool/dbconfig/20260730-071112-cwilliams.json * 07:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1178.eqiad.wmnet with reason: Maintenance * 07:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1177: Maintenance * 06:24 root@cumin1003: START - Cookbook sre.mysql.pool pool db1177: Maintenance * 06:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1177 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95741 and previous config saved to /var/cache/conftool/dbconfig/20260730-061736-cwilliams.json * 06:17 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1177.eqiad.wmnet with reason: Maintenance * 06:17 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1172: Maintenance * 05:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1218.eqiad.wmnet with reason: crashed * 05:41 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db1217 it crashed', diff saved to https://phabricator.wikimedia.org/P95737 and previous config saved to /var/cache/conftool/dbconfig/20260730-054111-marostegui.json * 05:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95736 and previous config saved to /var/cache/conftool/dbconfig/20260730-053422-cwilliams.json * 05:30 root@cumin1003: START - Cookbook sre.mysql.pool pool db1172: Maintenance * 05:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95734 and previous config saved to /var/cache/conftool/dbconfig/20260730-052414-cwilliams.json * 05:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1172 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95733 and previous config saved to /var/cache/conftool/dbconfig/20260730-052354-cwilliams.json * 05:23 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1172.eqiad.wmnet with reason: Maintenance * 05:23 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1167: Maintenance * 05:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95731 and previous config saved to /var/cache/conftool/dbconfig/20260730-051406-cwilliams.json * 05:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95729 and previous config saved to /var/cache/conftool/dbconfig/20260730-050358-cwilliams.json * 04:47 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 04:47 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 04:47 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 04:35 root@cumin1003: START - Cookbook sre.mysql.pool pool db1167: Maintenance * 04:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1167 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95726 and previous config saved to /var/cache/conftool/dbconfig/20260730-042923-cwilliams.json * 04:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 04:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1167.eqiad.wmnet with reason: Maintenance * 04:22 pt1979@cumin2002: START - Cookbook sre.dns.netbox * 04:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95725 and previous config saved to /var/cache/conftool/dbconfig/20260730-040337-cwilliams.json * 04:03 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:38 brett@cumin2002: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool eqsin [reason: Switch upgrade maintenance window complete, [[phab:T433097|T433097]]] * 01:38 brett@cumin2002: START - Cookbook sre.dns.admin DNS admin: pool eqsin [reason: Switch upgrade maintenance window complete, [[phab:T433097|T433097]]] == 2026-07-29 == * 23:57 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin,mr1-eqsin IPv6,mr1-eqsin.oob,mr1-eqsin.oob IPv6 with reason: connection issue * 22:54 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 22:53 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 22:53 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 22:53 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:25 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2022.codfw.wmnet, repooling source-only afterwards * 22:20 brett@cumin2002: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool eqsin [reason: Switch upgrade maintenance window, [[phab:T433097|T433097]]] * 22:20 brett@cumin2002: START - Cookbook sre.dns.admin DNS admin: depool eqsin [reason: Switch upgrade maintenance window, [[phab:T433097|T433097]]] * 22:01 apine@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 22:00 apine@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 21:59 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 21:58 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 21:58 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 21:58 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 21:32 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1048.eqiad.wmnet with OS trixie * 21:25 pt1979@cumin2002: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be2097.codfw.wmnet with OS bullseye * 21:16 zabe@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki=metawiki 'Mental Health Resource Center' 'Safety Resource Center/Mental Health' Zabe --reason 'per request [[:phab:T433118{{!}}T433118]]' * 21:12 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2022.codfw.wmnet, repooling source-only afterwards * 21:12 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] (duration: 12m 53s) * 21:12 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2015\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 21:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1253: Maintenance * 21:08 aaron@deploy1003: aaron: Continuing with deployment * 21:01 aaron@deploy1003: aaron: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:59 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] * 20:52 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] (duration: 21m 57s) * 20:48 aaron@deploy1003: aaron: Continuing with deployment * 20:32 aaron@deploy1003: aaron: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:30 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] * 20:24 root@cumin1003: START - Cookbook sre.mysql.pool pool db1253: Maintenance * 20:19 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] (duration: 08m 07s) * 20:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1253 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95719 and previous config saved to /var/cache/conftool/dbconfig/20260729-201810-cwilliams.json * 20:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1253.eqiad.wmnet with reason: Maintenance * 20:17 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1231: Maintenance * 20:15 aaron@deploy1003: bpirkle, aaron: Continuing with deployment * 20:13 aaron@deploy1003: bpirkle, aaron: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:12 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie * 20:11 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] * 20:11 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1048.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:09 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1048.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:09 pt1979@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 20:04 pt1979@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 19:47 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 19:43 pt1979@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye * 19:41 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 19:37 zabe: zabe@deploy1003:~$ mwscript-k8s --comment='[[phab:T433529|T433529]]' --follow -- resetAuthenticationThrottle.php --wiki=aawiki --signup --ip=89.36.114.94 * 19:36 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] (duration: 06m 49s) * 19:32 zabe@deploy1003: zabe: Continuing with deployment * 19:31 zabe@deploy1003: zabe: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:31 root@cumin1003: START - Cookbook sre.mysql.pool pool db1231: Maintenance * 19:29 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] * 19:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1231 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95714 and previous config saved to /var/cache/conftool/dbconfig/20260729-192454-cwilliams.json * 19:24 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1231.eqiad.wmnet with reason: Maintenance * 19:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1227: Maintenance * 19:22 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 19:22 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 19:21 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:21 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1048] - vriley@cumin1003" * 19:21 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1048] - vriley@cumin1003" * 19:19 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 19:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1251: Maintenance * 19:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95711 and previous config saved to /var/cache/conftool/dbconfig/20260729-191756-cwilliams.json * 19:16 vriley@cumin1003: START - Cookbook sre.dns.netbox * 19:11 dduvall: rolling back wmf.13 to group0 due to [[phab:T433457|T433457]] (cc [[phab:T430832|T430832]]) * 19:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95709 and previous config saved to /var/cache/conftool/dbconfig/20260729-190748-cwilliams.json * 19:01 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2022.codfw.wmnet with OS bookworm * 19:01 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2021.codfw.wmnet, repooling source-only afterwards * 18:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95707 and previous config saved to /var/cache/conftool/dbconfig/20260729-185740-cwilliams.json * 18:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95704 and previous config saved to /var/cache/conftool/dbconfig/20260729-184732-cwilliams.json * 18:37 root@cumin1003: START - Cookbook sre.mysql.pool pool db1227: Maintenance * 18:34 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2022.codfw.wmnet with reason: host reimage * 18:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1227 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95701 and previous config saved to /var/cache/conftool/dbconfig/20260729-183117-cwilliams.json * 18:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1227.eqiad.wmnet with reason: Maintenance * 18:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1202: Maintenance * 18:30 root@cumin1003: START - Cookbook sre.mysql.pool pool db1251: Maintenance * 18:27 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2022.codfw.wmnet with reason: host reimage * 18:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1251 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95698 and previous config saved to /var/cache/conftool/dbconfig/20260729-182428-cwilliams.json * 18:24 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lvs2014.codfw.wmnet * 18:24 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for lvs2014.codfw.wmnet * 18:24 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1251.eqiad.wmnet with reason: Maintenance * 18:23 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1235: Maintenance * 18:22 brett@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 18:19 brett@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 18:19 brett@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 18:17 brett@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 18:17 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 18:16 mutante: removing jenkins during the train - living on the edge - no, just kidding, jenkins has migrated to dedicated machines, nothing should happen * 18:15 brett@cumin2002: END (ERROR) - Cookbook sre.loadbalancer.restart-pybal (exit_code=97) rolling-restart of pybal on P<nowiki>{</nowiki>lvs2014.codfw.wmnet<nowiki>}</nowiki> and A:lvs ([[phab:T428495|T428495]]) * 18:15 mutante: CI: contint1002/contint2002: apt-get remove --purge jenkins - jenkins be gone - [[phab:T418521|T418521]] * 18:13 brett@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on P<nowiki>{</nowiki>lvs2014.codfw.wmnet<nowiki>}</nowiki> and A:lvs ([[phab:T428495|T428495]]) * 18:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2022 * 18:08 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2022 * 18:03 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T428495|T428495]] * 18:03 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2022 * 18:02 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2022.codfw.wmnet 211.48.192.10.in-addr.arpa 1.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:02 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2022.codfw.wmnet 211.48.192.10.in-addr.arpa 1.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:02 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:02 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2022 - bking@cumin2003" * 18:02 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2022 - bking@cumin2003" * 17:57 bking@cumin2003: START - Cookbook sre.dns.netbox * 17:56 brett@cumin2002: END (FAIL) - Cookbook sre.loadbalancer.restart-pybal (exit_code=1) rolling-restart of pybal on A:lvs-codfw and A:lvs ([[phab:T428495|T428495]]) * 17:55 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo - [[phab:T428495|T428495]] * 17:54 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2022 * 17:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2022.codfw.wmnet with OS bookworm * 17:50 brett@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on A:lvs-codfw and A:lvs ([[phab:T428495|T428495]]) * 17:47 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2021.codfw.wmnet, repooling source-only afterwards * 17:47 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 14s) * 17:47 swfrench-wmf: authdns-update to direct codfw, eqsin, ulsfo etcd clients back to codfw - [[phab:T428495|T428495]] * 17:47 swfrench@dns1004: END - running authdns-update * 17:47 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 17:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95692 and previous config saved to /var/cache/conftool/dbconfig/20260729-174713-cwilliams.json * 17:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance * 17:46 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1249: Maintenance * 17:45 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2015\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 17:45 swfrench@dns1004: START - running authdns-update * 17:44 root@cumin1003: START - Cookbook sre.mysql.pool pool db1202: Maintenance * 17:41 akhatun: Deployed refinery using scap, then deployed onto hdfs * 17:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1202 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95688 and previous config saved to /var/cache/conftool/dbconfig/20260729-173759-cwilliams.json * 17:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1202.eqiad.wmnet with reason: Maintenance * 17:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1194: Maintenance * 17:37 root@cumin1003: START - Cookbook sre.mysql.pool pool db1235: Maintenance * 17:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1230: Maintenance * 17:30 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1235 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95684 and previous config saved to /var/cache/conftool/dbconfig/20260729-173051-cwilliams.json * 17:30 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1235.eqiad.wmnet with reason: Maintenance * 17:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1234: Maintenance * 17:26 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (thin): Regular analytics weekly train THIN [analytics/refinery@56695674] (duration: 02m 02s) * 17:24 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (thin): Regular analytics weekly train THIN [analytics/refinery@56695674] * 17:23 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567]: Regular analytics weekly train [analytics/refinery@56695674] (duration: 06m 20s) * 17:20 dancy@deploy1003: Finished scap sync-world: Testing delay_messageblobstore_purge: true (duration: 06m 29s) * 17:17 akhatun@deploy1003: Started deploy [analytics/refinery@5669567]: Regular analytics weekly train [analytics/refinery@56695674] * 17:17 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] (duration: 00m 22s) * 17:16 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] * 17:13 dancy@deploy1003: Started scap sync-world: Testing delay_messageblobstore_purge: true * 17:05 mutante: CI: contint1002/contint2002 - restarted httpd to be extra sure all is cleaned up - https://integration.wikimedia.org/ci/ is up and running [[phab:T418521|T418521]] * 17:04 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 17:03 mutante: CI: contint1002/contint2002 - rm /etc/apache2/jenkins_proxy - removing legacy jenkins proxy config - jenkins is on new dedicated machines and uses jenkins_proxy_ext config [[phab:T418521|T418521]] * 17:02 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] (duration: 36m 25s) * 17:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1249: Maintenance * 16:59 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 16:54 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2015.codfw.wmnet, repooling source-only afterwards * 16:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1249 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95674 and previous config saved to /var/cache/conftool/dbconfig/20260729-165339-cwilliams.json * 16:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1249.eqiad.wmnet with reason: Maintenance * 16:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1248: Maintenance * 16:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1194: Maintenance * 16:47 swfrench-wmf: silenced EtcdReplicationDown 57b2b421-1cc9-4e38-9276-{{Gerrit|94f223fd231c}} - [[phab:T428495|T428495]] * 16:46 tchin@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/eventstreams-internal: apply * 16:46 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye * 16:46 tchin@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/eventstreams-internal: apply * 16:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1230: Maintenance * 16:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1194 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95669 and previous config saved to /var/cache/conftool/dbconfig/20260729-164422-cwilliams.json * 16:44 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 16:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1194.eqiad.wmnet with reason: Maintenance * 16:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1191: Maintenance * 16:43 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host an-test-master1003.eqiad.wmnet * 16:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db1234: Maintenance * 16:43 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 16:43 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Rolling back deployment * 16:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host an-test-master1004.eqiad.wmnet * 16:41 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 16:40 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 16:40 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 16:39 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 16:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1230 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95667 and previous config saved to /var/cache/conftool/dbconfig/20260729-163932-cwilliams.json * 16:39 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 16:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1230.eqiad.wmnet with reason: Maintenance * 16:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1207: Maintenance * 16:38 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 16:37 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host an-test-master1004.eqiad.wmnet * 16:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1234 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95664 and previous config saved to /var/cache/conftool/dbconfig/20260729-163719-cwilliams.json * 16:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1234.eqiad.wmnet with reason: Maintenance * 16:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1079.eqiad.wmnet with OS trixie * 16:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1232: Maintenance * 16:34 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1259: Maintenance * 16:28 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] (duration: 06m 57s) * 16:28 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:26 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] * 16:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1051 hosts * 16:21 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] * 16:20 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2006.codfw.wmnet with OS bookworm * 16:19 akhatun: Deploying Refinery at {{Gerrit|56695674}} as part of weekly train * 16:18 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1079.eqiad.wmnet with reason: host reimage * 16:16 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] (duration: 15m 36s) * 16:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2021.codfw.wmnet with OS bookworm * 16:14 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1079.eqiad.wmnet with reason: host reimage * 16:12 topranks: hot-swap line card in FPC0 on cr1-eqiad with replacement MPC10E from Juniper [[phab:T426343|T426343]] * 16:10 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Continuing with deployment * 16:07 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db1248: Maintenance * 16:01 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] * 16:00 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 16:00 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 15:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1248 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95651 and previous config saved to /var/cache/conftool/dbconfig/20260729-155956-cwilliams.json * 15:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1248.eqiad.wmnet with reason: Maintenance * 15:59 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2006.codfw.wmnet with reason: host reimage * 15:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1247: Maintenance * 15:59 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2074.codfw.wmnet with OS trixie * 15:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1191: Maintenance * 15:57 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:55 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1079.eqiad.wmnet with OS trixie * 15:55 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2006.codfw.wmnet with reason: host reimage * 15:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1207: Maintenance * 15:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1191 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95646 and previous config saved to /var/cache/conftool/dbconfig/20260729-155104-cwilliams.json * 15:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1191.eqiad.wmnet with reason: Maintenance * 15:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1181: Maintenance * 15:49 root@cumin1003: START - Cookbook sre.mysql.pool pool db1232: Maintenance * 15:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2021.codfw.wmnet with reason: host reimage * 15:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1207 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95643 and previous config saved to /var/cache/conftool/dbconfig/20260729-154735-cwilliams.json * 15:47 root@cumin1003: START - Cookbook sre.mysql.pool pool db1259: Maintenance * 15:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1207.eqiad.wmnet with reason: Maintenance * 15:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1200: Maintenance * 15:46 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] (duration: 31m 59s) * 15:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 15:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1232 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95640 and previous config saved to /var/cache/conftool/dbconfig/20260729-154330-cwilliams.json * 15:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1232.eqiad.wmnet with reason: Maintenance * 15:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1219: Maintenance * 15:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2015.codfw.wmnet, repooling source-only afterwards * 15:41 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 18s) * 15:41 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1259 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95638 and previous config saved to /var/cache/conftool/dbconfig/20260729-154107-cwilliams.json * 15:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1259.eqiad.wmnet with reason: Maintenance * 15:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1254: Maintenance * 15:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2015.codfw.wmnet with OS bookworm * 15:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2021.codfw.wmnet with reason: host reimage * 15:36 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2006.codfw.wmnet with OS bookworm * 15:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 15:35 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Continuing with deployment * 15:33 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2074.codfw.wmnet with OS trixie * 15:33 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 15:32 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:29 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:28 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be2074.codfw.wmnet with OS trixie * 15:28 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2006.codfw.wmnet * 15:26 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1078.eqiad.wmnet with OS trixie * 15:25 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 15:25 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:22 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2006.codfw.wmnet * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2021 * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2021 * 15:19 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2021 * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2021.codfw.wmnet 210.48.192.10.in-addr.arpa 0.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:19 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2021.codfw.wmnet 210.48.192.10.in-addr.arpa 0.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2021 - bking@cumin2003" * 15:19 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2021 - bking@cumin2003" * 15:14 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] * 15:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2015.codfw.wmnet with reason: host reimage * 15:11 root@cumin1003: START - Cookbook sre.mysql.pool pool db1247: Maintenance * 15:11 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on ml-serve2004.codfw.wmnet with reason: [[phab:T433478|T433478]] * 15:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2015.codfw.wmnet with reason: host reimage * 15:10 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on ml-serve2002.codfw.wmnet with reason: [[phab:T433476|T433476]] * 15:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 15:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1247 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95625 and previous config saved to /var/cache/conftool/dbconfig/20260729-150459-cwilliams.json * 15:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1247.eqiad.wmnet with reason: Maintenance * 15:04 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:04 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1244: Maintenance * 15:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1078.eqiad.wmnet with reason: host reimage * 15:03 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1005.wikimedia.org * 15:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db1181: Maintenance * 15:01 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2098.codfw.wmnet with OS bullseye * 15:00 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye * 15:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1200: Maintenance * 14:59 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 14:59 root@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1285.eqiad.wmnet with OS trixie * 14:59 Amir1: mwscript-k8s -- extensions/TimedMediaHandler/maintenance/requeueTranscodes.php --wiki=commonswiki --key '360p.mpeg4.mov' --throttle --video --missing ([[phab:T358266|T358266]]) * 14:58 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1005.wikimedia.org * 14:58 jhancock@cumin2002: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['ms-be2098'] * 14:58 jhancock@cumin2002: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['ms-be2098'] * 14:58 jhancock@cumin2002: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['ms-be2097'] * 14:58 jhancock@cumin2002: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['ms-be2097'] * 14:58 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1078.eqiad.wmnet with reason: host reimage * 14:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1181 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95621 and previous config saved to /var/cache/conftool/dbconfig/20260729-145629-cwilliams.json * 14:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1181.eqiad.wmnet with reason: Maintenance * 14:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1174: Maintenance * 14:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1219: Maintenance * 14:55 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader1006.wikimedia.org on all recursors * 14:55 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader1006.wikimedia.org on all recursors * 14:55 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader1005.wikimedia.org on all recursors * 14:55 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader1005.wikimedia.org on all recursors * 14:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1200 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95618 and previous config saved to /var/cache/conftool/dbconfig/20260729-145336-cwilliams.json * 14:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1254: Maintenance * 14:53 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1200.eqiad.wmnet with reason: Maintenance * 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1185: Maintenance * 14:52 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2021 * 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2015 * 14:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2015 * 14:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1219 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95616 and previous config saved to /var/cache/conftool/dbconfig/20260729-144946-cwilliams.json * 14:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1219.eqiad.wmnet with reason: Maintenance * 14:49 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1218: Maintenance * 14:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2021.codfw.wmnet with OS bookworm * 14:48 dancy@deploy1003: Finished deploy [zuul/deploy@22703a6]: Deploying https://gerrit.wikimedia.org/r/c/integration/zuul/+/1311501 ([[phab:T432491|T432491]]) (duration: 00m 15s) * 14:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2015.codfw.wmnet with OS bookworm * 14:48 dancy@deploy1003: Started deploy [zuul/deploy@22703a6]: Deploying https://gerrit.wikimedia.org/r/c/integration/zuul/+/1311501 ([[phab:T432491|T432491]]) * 14:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1254 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95613 and previous config saved to /var/cache/conftool/dbconfig/20260729-144729-cwilliams.json * 14:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1254.eqiad.wmnet with reason: Maintenance * 14:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1233: Maintenance * 14:46 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2013\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 14:46 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2014\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 14:46 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:45 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:44 root@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1285.eqiad.wmnet with reason: host reimage * 14:43 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:42 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:41 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:40 root@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1285.eqiad.wmnet with reason: host reimage * 14:39 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1078.eqiad.wmnet with OS trixie * 14:39 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2074.codfw.wmnet with OS trixie * 14:32 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:32 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:32 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:31 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2005.codfw.wmnet with OS bookworm * 14:30 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:30 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:29 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:29 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:27 root@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host db1285 * 14:27 root@cumin1003: START - Cookbook sre.hosts.move-vlan for host db1285 * 14:27 root@cumin1003: START - Cookbook sre.hosts.reimage for host db1285.eqiad.wmnet with OS trixie * 14:24 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 14:24 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:24 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:24 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:23 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:22 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:22 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:21 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:17 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:16 root@cumin1003: START - Cookbook sre.mysql.pool pool db1244: Maintenance * 14:15 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:15 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add asw1-604 loopback ipv4 - pt1979@cumin2002" * 14:15 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add asw1-604 loopback ipv4 - pt1979@cumin2002" * 14:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:12 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 14:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95599 and previous config saved to /var/cache/conftool/dbconfig/20260729-141014-cwilliams.json * 14:10 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 14:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1244.eqiad.wmnet with reason: Maintenance * 14:10 pt1979@cumin2002: START - Cookbook sre.dns.netbox * 14:10 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 14:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1243: Maintenance * 14:09 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2005.codfw.wmnet with reason: host reimage * 14:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db1174: Maintenance * 14:08 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad * 14:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db1185: Maintenance * 14:06 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 14:05 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2005.codfw.wmnet with reason: host reimage * 14:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1174 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95595 and previous config saved to /var/cache/conftool/dbconfig/20260729-140309-cwilliams.json * 14:03 sukhe@dns1004: END - running authdns-update * 14:03 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1174.eqiad.wmnet with reason: Maintenance * 14:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1170: Maintenance * 14:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db1218: Maintenance * 14:01 sukhe@dns1004: START - running authdns-update * 14:00 sukhe@dns1004: START - running authdns-update * 13:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db1233: Maintenance * 13:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1185 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95592 and previous config saved to /var/cache/conftool/dbconfig/20260729-135925-cwilliams.json * 13:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1185.eqiad.wmnet with reason: Maintenance * 13:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1161: Maintenance * 13:58 sukhe@puppetserver1001: conftool action : set/pooled=true; selector: dnsdisc=urldownloader * 13:58 root@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1265.eqiad.wmnet with OS trixie * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1218 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95590 and previous config saved to /var/cache/conftool/dbconfig/20260729-135621-cwilliams.json * 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1218.eqiad.wmnet with reason: Maintenance * 13:55 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1206: Maintenance * 13:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2073.codfw.wmnet with OS trixie * 13:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1233 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95587 and previous config saved to /var/cache/conftool/dbconfig/20260729-135335-cwilliams.json * 13:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1233.eqiad.wmnet with reason: Maintenance * 13:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1229: Maintenance * 13:50 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/kartotherian: apply * 13:50 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service * 13:49 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:49 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/kartotherian: apply * 13:48 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 13:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1077.eqiad.wmnet with OS trixie * 13:47 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 13:46 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2005.codfw.wmnet with OS bookworm * 13:44 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 13:44 root@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1265.eqiad.wmnet with reason: host reimage * 13:40 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] (duration: 09m 22s) * 13:39 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:38 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:36 root@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1265.eqiad.wmnet with reason: host reimage * 13:35 stran@deploy1003: stran: Continuing with deployment * 13:33 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 13:32 stran@deploy1003: stran: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified t * 13:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2073.codfw.wmnet with reason: host reimage * 13:30 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ml-build1001.eqiad.wmnet * 13:30 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] * 13:29 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad * 13:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1077.eqiad.wmnet with reason: host reimage * 13:27 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 13:27 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:27 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad * 13:26 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2073.codfw.wmnet with reason: host reimage * 13:26 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] (duration: 07m 56s) * 13:25 klausman@cumin1003: START - Cookbook sre.hosts.reboot-single for host ml-build1001.eqiad.wmnet * 13:24 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 13:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>ml-serve101[2-5].eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 13:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1015.eqiad.wmnet * 13:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1015.eqiad.wmnet * 13:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1077.eqiad.wmnet with reason: host reimage * 13:23 root@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host db1265 * 13:23 root@cumin1003: START - Cookbook sre.hosts.move-vlan for host db1265 * 13:23 root@cumin1003: START - Cookbook sre.hosts.reimage for host db1265.eqiad.wmnet with OS trixie * 13:23 root@cumin1003: START - Cookbook sre.mysql.pool pool db1243: Maintenance * 13:22 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 13:22 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2005.codfw.wmnet * 13:22 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 13:22 samtar@deploy1003: dreamrimmer, samtar: Continuing with deployment * 13:20 samtar@deploy1003: dreamrimmer, samtar: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts an-test-master[1001-1002].eqiad.wmnet * 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-master[1001-1002].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 13:18 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1015.eqiad.wmnet * 13:18 sukhe@cumin1003: END (ERROR) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=97) for role: url_downloader@eqiad * 13:18 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 13:18 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] * 13:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95574 and previous config saved to /var/cache/conftool/dbconfig/20260729-131638-cwilliams.json * 13:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1243.eqiad.wmnet with reason: Maintenance * 13:16 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2005.codfw.wmnet * 13:16 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1242: Maintenance * 13:14 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] (duration: 07m 00s) * 13:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db1170: Maintenance * 13:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1015.eqiad.wmnet * 13:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1014.eqiad.wmnet * 13:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1014.eqiad.wmnet * 13:12 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1223: Maintenance * 13:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db1161: Maintenance * 13:10 samtar@deploy1003: anzx, samtar: Continuing with deployment * 13:09 samtar@deploy1003: anzx, samtar: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db1206: Maintenance * 13:08 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 13:07 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] * 13:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1170 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95566 and previous config saved to /var/cache/conftool/dbconfig/20260729-130730-cwilliams.json * 13:07 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1170.eqiad.wmnet with reason: Maintenance * 13:07 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:07 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt IPs new switches - cmooney@cumin1003" * 13:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1158: Maintenance * 13:06 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1014.eqiad.wmnet * 13:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1161 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95564 and previous config saved to /var/cache/conftool/dbconfig/20260729-130616-cwilliams.json * 13:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 13:06 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1077.eqiad.wmnet with OS trixie * 13:05 root@cumin1003: START - Cookbook sre.mysql.pool pool db1229: Maintenance * 13:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1161.eqiad.wmnet with reason: Maintenance * 13:05 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt IPs new switches - cmooney@cumin1003" * 13:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2073.codfw.wmnet with OS trixie * 13:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Maintenance * 13:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1206 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95562 and previous config saved to /var/cache/conftool/dbconfig/20260729-130258-cwilliams.json * 13:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1206.eqiad.wmnet with reason: Maintenance * 13:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1196: Maintenance * 13:01 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 13:01 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:00 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 13:00 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 12:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1229 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95559 and previous config saved to /var/cache/conftool/dbconfig/20260729-125950-cwilliams.json * 12:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1229.eqiad.wmnet with reason: Maintenance * 12:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1222: Maintenance * 12:57 sukhe: sudo cumin 'A:lvs and (A:eqiad or A:codfw)' 'disable-puppet "adding new service urldownloader"': [[phab:T429175|T429175]] * 12:56 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1014.eqiad.wmnet * 12:56 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1013.eqiad.wmnet * 12:56 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1013.eqiad.wmnet * 12:50 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1013.eqiad.wmnet * 12:50 sukhe: sudo cumin 'O:url_downloader' 'run-puppet-agent --enable "merging CR 1313948"': [[phab:T429175|T429175]] * 12:48 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-master[1001-1002].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 12:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1013.eqiad.wmnet * 12:45 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1012.eqiad.wmnet * 12:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1012.eqiad.wmnet * 12:45 sukhe: sudo cumin 'O:url_downloader' 'disable-puppet "merging CR 1313948"': [[phab:T429175|T429175]] * 12:44 btullis@cumin1003: START - Cookbook sre.dns.netbox * 12:40 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test2001.codfw.wmnet * 12:40 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test2001.codfw.wmnet * 12:38 ayounsi@dns1004: END - running authdns-update * 12:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1012.eqiad.wmnet * 12:35 ayounsi@dns1004: START - running authdns-update * 12:34 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts an-test-master[1001-1002].eqiad.wmnet * 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts an-test-coord1001.eqiad.wmnet * 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-coord1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 12:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1012.eqiad.wmnet * 12:32 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>ml-serve101[2-5].eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 12:29 root@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Maintenance * 12:25 root@cumin1003: START - Cookbook sre.mysql.pool pool db1223: Maintenance * 12:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95544 and previous config saved to /var/cache/conftool/dbconfig/20260729-122254-cwilliams.json * 12:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1242.eqiad.wmnet with reason: Maintenance * 12:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1241: Maintenance * 12:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1051 hosts * 12:20 root@cumin1003: START - Cookbook sre.mysql.pool pool db1158: Maintenance * 12:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1223 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95540 and previous config saved to /var/cache/conftool/dbconfig/20260729-121937-cwilliams.json * 12:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1223.eqiad.wmnet with reason: Maintenance * 12:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1212: Maintenance * 12:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db1159: Maintenance * 12:17 elukey@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: sync * 12:15 elukey@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: sync * 12:15 root@cumin1003: START - Cookbook sre.mysql.pool pool db1196: Maintenance * 12:14 Daimona: Creating new DB tables for the CampaignEvents extension in x1.testwiki, x1.test2wiki, x1.officewiki, and x1.wikishared # [[phab:T429339|T429339]] * 12:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db1222: Maintenance * 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95535 and previous config saved to /var/cache/conftool/dbconfig/20260729-121211-cwilliams.json * 12:12 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 12:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1158.eqiad.wmnet with reason: Maintenance * 12:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1159 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95534 and previous config saved to /var/cache/conftool/dbconfig/20260729-121146-cwilliams.json * 12:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1159.eqiad.wmnet with reason: Maintenance * 12:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1196 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95533 and previous config saved to /var/cache/conftool/dbconfig/20260729-120847-cwilliams.json * 12:08 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 12:08 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1196.eqiad.wmnet with reason: Maintenance * 12:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1195: Maintenance * 12:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1222 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95530 and previous config saved to /var/cache/conftool/dbconfig/20260729-120424-cwilliams.json * 12:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1222.eqiad.wmnet with reason: Maintenance * 12:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1098 hosts * 12:00 marostegui: Rename tables [[phab:T425074|T425074]] * 12:00 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-coord1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 11:58 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1197: Maintenance * 11:55 btullis@cumin1003: START - Cookbook sre.dns.netbox * 11:52 elukey@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: sync * 11:51 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:51 elukey@deploy1003: helmfile [codfw] START helmfile.d/services/proton: sync * 11:51 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:50 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts an-test-coord1001.eqiad.wmnet * 11:50 elukey@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: sync * 11:49 elukey@deploy1003: helmfile [staging] START helmfile.d/services/proton: sync * 11:35 root@cumin1003: START - Cookbook sre.mysql.pool pool db1241: Maintenance * 11:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db1212: Maintenance * 11:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1241 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95520 and previous config saved to /var/cache/conftool/dbconfig/20260729-112918-cwilliams.json * 11:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1241.eqiad.wmnet with reason: Maintenance * 11:29 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1238: Maintenance * 11:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1212 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95517 and previous config saved to /var/cache/conftool/dbconfig/20260729-112727-cwilliams.json * 11:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 11:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1212.eqiad.wmnet with reason: Maintenance * 11:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1198: Maintenance * 11:23 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:22 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:21 root@cumin1003: START - Cookbook sre.mysql.pool pool db1195: Maintenance * 11:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1195 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95514 and previous config saved to /var/cache/conftool/dbconfig/20260729-111450-cwilliams.json * 11:14 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1195.eqiad.wmnet with reason: Maintenance * 11:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1186: Maintenance * 11:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 11:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 11:05 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:54 marostegui: Dropping renamed tables [[phab:T425066|T425066]] * 10:41 root@cumin1003: START - Cookbook sre.mysql.pool pool db1238: Maintenance * 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1198: Maintenance * 10:39 Amir1: ran https://phabricator.wikimedia.org/T432509#12149723 in production ([[phab:T432509|T432509]]) * 10:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db1197: Maintenance * 10:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1238 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95501 and previous config saved to /var/cache/conftool/dbconfig/20260729-103532-cwilliams.json * 10:35 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1238.eqiad.wmnet with reason: Maintenance * 10:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1221: Maintenance * 10:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1198 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95499 and previous config saved to /var/cache/conftool/dbconfig/20260729-103330-cwilliams.json * 10:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1198.eqiad.wmnet with reason: Maintenance * 10:33 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1175: Maintenance * 10:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1197 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95496 and previous config saved to /var/cache/conftool/dbconfig/20260729-103217-cwilliams.json * 10:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1197.eqiad.wmnet with reason: Maintenance * 10:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1188: Maintenance * 10:27 root@cumin1003: START - Cookbook sre.mysql.pool pool db1186: Maintenance * 10:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1186 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95493 and previous config saved to /var/cache/conftool/dbconfig/20260729-102111-cwilliams.json * 10:21 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1186.eqiad.wmnet with reason: Maintenance * 10:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 10:14 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 09:53 XioNoX: reboot cr2-magru - [[phab:T431750|T431750]] * 09:52 XioNoX: drain cr2-magru - [[phab:T431750|T431750]] * 09:48 root@cumin1003: START - Cookbook sre.mysql.pool pool db1221: Maintenance * 09:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zookeeper-test1002.eqiad.wmnet * 09:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1188: Maintenance * 09:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1175: Maintenance * 09:44 btullis@dns1004: END - running authdns-update * 09:42 btullis@dns1004: START - running authdns-update * 09:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1221 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95483 and previous config saved to /var/cache/conftool/dbconfig/20260729-094200-cwilliams.json * 09:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 7 hosts with reason: Maintenance * 09:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1221.eqiad.wmnet with reason: Maintenance * 09:41 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host zookeeper-test1002.eqiad.wmnet * 09:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1199: Maintenance * 09:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1188 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95481 and previous config saved to /var/cache/conftool/dbconfig/20260729-093917-cwilliams.json * 09:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1188.eqiad.wmnet with reason: Maintenance * 09:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1182: Maintenance * 09:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1175 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95479 and previous config saved to /var/cache/conftool/dbconfig/20260729-093842-cwilliams.json * 09:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1175.eqiad.wmnet with reason: Maintenance * 09:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1166: Maintenance * 09:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1169: Maintenance * 09:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1033.eqiad.wmnet,service=s8 * 09:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1033.eqiad.wmnet,service=s5 * 09:33 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1033.eqiad.wmnet,service=s8 * 09:33 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1033.eqiad.wmnet,service=s5 * 09:21 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:21 XioNoX: reboot cr1-magru - [[phab:T431750|T431750]] * 09:17 XioNoX: drain cr1-magru - [[phab:T431750|T431750]] * 09:15 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm1001.wikimedia.org * 09:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply * 09:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply * 09:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 09:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 09:11 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr2-magru,cr2-magru IPv6,cr2-magru.mgmt with reason: router upgrade * 09:11 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm1001.wikimedia.org * 09:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp1005.wikimedia.org * 09:07 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp1005.wikimedia.org * 09:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp2005.wikimedia.org * 09:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp2005.wikimedia.org * 09:00 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr1-magru,cr1-magru IPv6,cr1-magru.mgmt with reason: router upgrade * 09:00 marostegui: Dropping renamed tables [[phab:T426341|T426341]] * 08:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1199: Maintenance * 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1182: Maintenance * 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1166: Maintenance * 08:47 root@cumin1003: START - Cookbook sre.mysql.pool pool db1169: Maintenance * 08:46 ayounsi@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 1:00:00 on cr1-magru,cr1-magru IPv6,cr1-magru.mgmt with reason: router upgrade * 08:45 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2072.codfw.wmnet with OS trixie * 08:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1199 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95464 and previous config saved to /var/cache/conftool/dbconfig/20260729-084534-cwilliams.json * 08:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1199.eqiad.wmnet with reason: Maintenance * 08:45 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 08:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1190: Maintenance * 08:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1182 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95462 and previous config saved to /var/cache/conftool/dbconfig/20260729-084436-cwilliams.json * 08:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1182.eqiad.wmnet with reason: Maintenance * 08:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1156: Maintenance * 08:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1166 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95460 and previous config saved to /var/cache/conftool/dbconfig/20260729-084400-cwilliams.json * 08:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1166.eqiad.wmnet with reason: Maintenance * 08:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1157: Maintenance * 08:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95458 and previous config saved to /var/cache/conftool/dbconfig/20260729-084147-cwilliams.json * 08:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1169.eqiad.wmnet with reason: Maintenance * 08:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1163: Maintenance * 08:30 btullis@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 11 hosts with reason: Replacing the namenodes * 08:23 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2072.codfw.wmnet with reason: host reimage * 08:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1098 hosts * 08:19 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2072.codfw.wmnet with reason: host reimage * 07:58 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2072.codfw.wmnet with OS trixie * 07:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1190: Maintenance * 07:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1156: Maintenance * 07:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1157: Maintenance * 07:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1163: Maintenance * 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1190 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95444 and previous config saved to /var/cache/conftool/dbconfig/20260729-074930-cwilliams.json * 07:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1190.eqiad.wmnet with reason: Maintenance * 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1157 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95443 and previous config saved to /var/cache/conftool/dbconfig/20260729-074914-cwilliams.json * 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1156 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95442 and previous config saved to /var/cache/conftool/dbconfig/20260729-074906-cwilliams.json * 07:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1157.eqiad.wmnet with reason: Maintenance * 07:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 07:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1156.eqiad.wmnet with reason: Maintenance * 07:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1163 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95441 and previous config saved to /var/cache/conftool/dbconfig/20260729-074652-cwilliams.json * 07:46 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1163.eqiad.wmnet with reason: Maintenance * 07:46 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2034.codfw.wmnet * 07:42 ayounsi@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2034.codfw.wmnet * 07:42 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2232.codfw.wmnet with OS trixie * 07:34 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:34 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:33 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:31 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:19 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2232.codfw.wmnet with reason: host reimage * 07:15 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2232.codfw.wmnet with reason: host reimage * 06:58 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db2232.codfw.wmnet with OS trixie * 06:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[2160,2232].codfw.wmnet with reason: Reimage * 06:26 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1164.eqiad.wmnet with OS trixie * 06:05 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1164.eqiad.wmnet with reason: host reimage * 06:01 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1164.eqiad.wmnet with reason: host reimage * 05:47 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1164.eqiad.wmnet with OS trixie * 05:46 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1164.eqiad.wmnet with reason: Reimage == 2026-07-28 == * 22:50 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1138.eqiad.wmnet * 22:50 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1138.eqiad.wmnet * 22:49 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1138.eqiad.wmnet * 22:11 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2014.codfw.wmnet, repooling source-only afterwards * 22:08 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2013.codfw.wmnet, repooling source-only afterwards * 22:03 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 20:58 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] (duration: 08m 19s) * 20:55 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2014.codfw.wmnet, repooling source-only afterwards * 20:55 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2013.codfw.wmnet, repooling source-only afterwards * 20:54 arlolra@deploy1003: arlolra: Continuing with deployment * 20:54 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 14s) * 20:54 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 20:53 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 30s) * 20:53 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 20:52 arlolra@deploy1003: arlolra: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:51 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:50 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] * 20:49 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:43 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 20:34 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] (duration: 06m 54s) * 20:34 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:34 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:31 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:31 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:30 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:30 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:30 arlolra@deploy1003: arlolra: Continuing with deployment * 20:29 arlolra@deploy1003: arlolra: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:27 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] * 20:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2014.codfw.wmnet with OS bookworm * 20:21 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 20:21 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:20 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 20:19 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:19 swfrench-wmf: switched etcd-mirror replication from conf2005 to conf2004 - [[phab:T428495|T428495]] * 20:17 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:17 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:15 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] (duration: 08m 26s) * 20:12 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:11 arlolra@deploy1003: anzx, arlolra: Continuing with deployment * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2013.codfw.wmnet with OS bookworm * 20:09 arlolra@deploy1003: anzx, arlolra: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] * 20:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2216: Maintenance * 19:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2014.codfw.wmnet with reason: host reimage * 19:57 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:54 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:54 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:53 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:52 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2014.codfw.wmnet with reason: host reimage * 19:49 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2013.codfw.wmnet with reason: host reimage * 19:42 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 19:41 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:41 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2013.codfw.wmnet with reason: host reimage * 19:39 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:39 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:39 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-eqiad: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 19:38 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2014 * 19:33 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2014 * 19:29 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2014.codfw.wmnet with OS bookworm * 19:28 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:27 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1006 * 19:26 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2012\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 19:26 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1006 * 19:26 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:26 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 19:25 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2013 * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2013 * 19:21 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2013 * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2013.codfw.wmnet 84.0.192.10.in-addr.arpa 4.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:21 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2013.codfw.wmnet 84.0.192.10.in-addr.arpa 4.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2013 - bking@cumin2003" * 19:21 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2013 - bking@cumin2003" * 19:21 vriley@cumin1003: START - Cookbook sre.dns.netbox * 19:20 root@cumin1003: START - Cookbook sre.mysql.pool pool db2216: Maintenance * 19:13 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2216 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95435 and previous config saved to /var/cache/conftool/dbconfig/20260728-191343-cwilliams.json * 19:13 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2216.codfw.wmnet with reason: Maintenance * 19:13 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2203: Maintenance * 19:06 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1005.eqiad.wmnet with OS trixie * 19:06 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 19:06 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 18:46 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 18:45 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:45 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:43 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:40 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:36 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-eqiad: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 18:35 dancy@deploy1003: Installation of scap version "4.275.0" completed for 3 hosts * 18:33 dancy@deploy1003: Installing scap version "4.275.0" for 3 host(s) * 18:32 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:32 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2097.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:30 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns3003.wikimedia.org [reason: pool for all services after reimaging] * 18:29 sukhe@dns1004: END - running authdns-update * 18:27 sukhe@dns1004: START - running authdns-update * 18:27 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns3003.wikimedia.org,service=authdns-update [reason: pool authdns-update after reimaging] * 18:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db2203: Maintenance * 18:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2203 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95430 and previous config saved to /var/cache/conftool/dbconfig/20260728-181958-cwilliams.json * 18:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2203.codfw.wmnet with reason: Maintenance * 18:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2188: Maintenance * 18:18 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2097.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:17 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-be2098 * 18:17 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host ms-be2098 * 18:17 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-be2097 * 18:16 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host ms-be2097 * 18:15 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:15 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding ms-be2097-8 to codfw - jhancock@cumin2002" * 18:15 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding ms-be2097-8 to codfw - jhancock@cumin2002" * 18:10 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 18:08 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage * 18:05 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns3003.wikimedia.org with OS trixie * 18:03 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage * 17:56 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-codfw: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 17:45 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie * 17:45 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1005.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:41 sukhe@dns1004: END - running authdns-update * 17:39 sukhe@dns1004: START - running authdns-update * 17:36 sukhe@puppetserver1001: conftool action : set/weight=1; selector: cluster=urldownloader,service=squid * 17:36 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1005.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:35 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader,service=squid * 17:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 17:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1005 * 17:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 17:34 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1005 * 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1005] - vriley@cumin1003" * 17:34 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1005] - vriley@cumin1003" * 17:32 root@cumin1003: START - Cookbook sre.mysql.pool pool db2188: Maintenance * 17:29 vriley@cumin1003: START - Cookbook sre.dns.netbox * 17:29 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 17:26 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2188 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95425 and previous config saved to /var/cache/conftool/dbconfig/20260728-172609-cwilliams.json * 17:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2188.codfw.wmnet with reason: Maintenance * 17:25 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2176: Maintenance * 17:19 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1138.eqiad.wmnet with OS trixie * 17:18 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1005 * 17:18 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1005 * 17:18 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:15 vriley@cumin1003: START - Cookbook sre.dns.netbox * 17:13 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns3003.wikimedia.org with reason: host reimage * 17:07 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns3003.wikimedia.org with reason: host reimage * 17:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-eqiad * 17:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1015.eqiad.wmnet * 17:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1015.eqiad.wmnet * 16:59 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1138.eqiad.wmnet with reason: host reimage * 16:55 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-codfw: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 16:54 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1138.eqiad.wmnet with reason: host reimage * 16:53 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1015.eqiad.wmnet * 16:43 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns3003.wikimedia.org with OS trixie * 16:43 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1015.eqiad.wmnet * 16:43 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1014.eqiad.wmnet * 16:43 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1014.eqiad.wmnet * 16:43 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=dns3003.wikimedia.org [reason: depooling for reimage to trixie] * 16:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 290 hosts * 16:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db2176: Maintenance * 16:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1138 * 16:38 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1138 * 16:37 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1138 * 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1138.eqiad.wmnet 193.32.64.10.in-addr.arpa 3.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:37 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1138.eqiad.wmnet 193.32.64.10.in-addr.arpa 3.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1138 - jiji@cumin1003" * 16:37 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1138 - jiji@cumin1003" * 16:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1014.eqiad.wmnet * 16:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2176 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95420 and previous config saved to /var/cache/conftool/dbconfig/20260728-163235-cwilliams.json * 16:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2176.codfw.wmnet with reason: Maintenance * 16:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1014.eqiad.wmnet * 16:32 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1013.eqiad.wmnet * 16:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1013.eqiad.wmnet * 16:32 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2174: Maintenance * 16:28 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2012.codfw.wmnet, repooling source-only afterwards * 16:25 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1013.eqiad.wmnet * 16:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1013.eqiad.wmnet * 16:20 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1012.eqiad.wmnet * 16:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1012.eqiad.wmnet * 16:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1012.eqiad.wmnet * 16:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1012.eqiad.wmnet * 16:03 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1011.eqiad.wmnet * 16:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1011.eqiad.wmnet * 16:00 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2004.codfw.wmnet with OS bookworm * 15:59 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1011.eqiad.wmnet * 15:56 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:55 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 15:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:54 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1011.eqiad.wmnet * 15:54 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1010.eqiad.wmnet * 15:54 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1010.eqiad.wmnet * 15:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:50 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 15:49 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1010.eqiad.wmnet * 15:48 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 15:48 jiji@cumin1003: START - Cookbook sre.dns.netbox * 15:46 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2248: Maintenance * 15:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db2174: Maintenance * 15:44 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1010.eqiad.wmnet * 15:44 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1009.eqiad.wmnet * 15:44 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1009.eqiad.wmnet * 15:42 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1138 * 15:41 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1138.eqiad.wmnet with OS trixie * 15:39 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1009.eqiad.wmnet * 15:39 robh@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on arclamp2001.codfw.wmnet with reason: ram upgrade * 15:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2174 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95413 and previous config saved to /var/cache/conftool/dbconfig/20260728-153844-cwilliams.json * 15:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2174.codfw.wmnet with reason: Maintenance * 15:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2173: Maintenance * 15:37 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 15:35 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 15:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1009.eqiad.wmnet * 15:34 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1008.eqiad.wmnet * 15:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1008.eqiad.wmnet * 15:31 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 15:31 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 15:29 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1008.eqiad.wmnet * 15:27 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 290 hosts * 15:25 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2013 * 15:25 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2195: Maintenance * 15:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1008.eqiad.wmnet * 15:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1007.eqiad.wmnet * 15:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1007.eqiad.wmnet * 15:22 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2004.codfw.wmnet with reason: host reimage * 15:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2013.codfw.wmnet with OS bookworm * 15:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 15:19 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2004.codfw.wmnet with reason: host reimage * 15:19 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 15:17 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1007.eqiad.wmnet * 15:12 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1007.eqiad.wmnet * 15:12 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1006.eqiad.wmnet * 15:12 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1006.eqiad.wmnet * 15:11 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1138.eqiad.wmnet * 15:11 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 15:11 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1138.eqiad.wmnet * 15:11 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1138.eqiad.wmnet * 15:10 brennen@deploy1003: Finished deploy [phabricator/deployment@f8b349f]: deploy phab1004 for [[phab:T433382|T433382]] (duration: 00m 43s) * 15:10 brennen@deploy1003: Started deploy [phabricator/deployment@f8b349f]: deploy phab1004 for [[phab:T433382|T433382]] * 15:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts * 15:09 brennen@deploy1003: Finished deploy [phabricator/deployment@f8b349f]: deploy phab2003 for [[phab:T433382|T433382]] (duration: 00m 55s) * 15:08 brennen@deploy1003: Started deploy [phabricator/deployment@f8b349f]: deploy phab2003 for [[phab:T433382|T433382]] * 15:07 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2012.codfw.wmnet, repooling source-only afterwards * 15:07 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts * 15:06 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1137.eqiad.wmnet * 15:06 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1137.eqiad.wmnet * 15:06 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1137.eqiad.wmnet * 15:05 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1006.eqiad.wmnet * 15:05 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1022\.eqiad\.wmnet,dc=eqiad,cluster=wdqs\-main,service=wdqs\-main * 15:01 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab2003.codfw.wmnet with reason: deployment * 15:01 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1005.eqiad.wmnet with reason: deployment * 15:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1006.eqiad.wmnet * 15:00 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1006.eqiad.wmnet with reason: deployment * 15:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1005.eqiad.wmnet * 15:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1005.eqiad.wmnet * 14:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db2248: Maintenance * 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1004.eqiad.wmnet with reason: deployment * 14:59 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2004.codfw.wmnet with OS bookworm * 14:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1005.eqiad.wmnet * 14:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2248 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95403 and previous config saved to /var/cache/conftool/dbconfig/20260728-145532-cwilliams.json * 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2245-2247].codfw.wmnet with reason: Maintenance * 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2248.codfw.wmnet with reason: Maintenance * 14:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2240: Maintenance * 14:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2173: Maintenance * 14:51 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts * 14:50 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1005.eqiad.wmnet * 14:50 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1004.eqiad.wmnet * 14:50 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1004.eqiad.wmnet * 14:49 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts * 14:45 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2004.codfw.wmnet * 14:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2173 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95399 and previous config saved to /var/cache/conftool/dbconfig/20260728-144453-cwilliams.json * 14:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2173.codfw.wmnet with reason: Maintenance * 14:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2170: Maintenance * 14:44 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1004.eqiad.wmnet * 14:39 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2004.codfw.wmnet * 14:38 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1004.eqiad.wmnet * 14:38 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet * 14:38 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet * 14:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db2195: Maintenance * 14:36 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1022.eqiad.wmnet, repooling source-only afterwards * 14:36 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 14:33 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet * 14:33 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2222: Maintenance * 14:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2195 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95394 and previous config saved to /var/cache/conftool/dbconfig/20260728-143218-cwilliams.json * 14:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2195.codfw.wmnet with reason: Maintenance * 14:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2181: Maintenance * 14:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 14:30 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 14:25 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 14:25 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 14:25 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 14:23 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet * 14:23 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1002.eqiad.wmnet * 14:23 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1002.eqiad.wmnet * 14:23 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 19s) * 14:23 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:18 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1002.eqiad.wmnet * 14:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2012.codfw.wmnet with OS bookworm * 14:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1002.eqiad.wmnet * 14:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1001.eqiad.wmnet * 14:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1001.eqiad.wmnet * 14:11 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 14:11 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 14:08 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1001.eqiad.wmnet * 14:07 elukey@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'. * 14:07 elukey@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'. * 14:06 elukey@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'. * 14:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db2240: Maintenance * 14:06 XioNoX: un-drain cr2-esams - [[phab:T431751|T431751]] * 14:05 elukey@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'. * 14:02 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1001.eqiad.wmnet * 14:02 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-eqiad * 14:01 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 14:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2240 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95384 and previous config saved to /var/cache/conftool/dbconfig/20260728-140011-cwilliams.json * 14:00 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2240.codfw.wmnet with reason: Maintenance * 13:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2237: Maintenance * 13:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db2170: Maintenance * 13:56 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 13:55 XioNoX: reboot cr2-esams - [[phab:T431751|T431751]] * 13:52 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 13:51 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr2-esams,cr2-esams IPv6,cr2-esams.mgmt with reason: router upgrade * 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 13:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2170 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95381 and previous config saved to /var/cache/conftool/dbconfig/20260728-135043-cwilliams.json * 13:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2170.codfw.wmnet with reason: Maintenance * 13:50 XioNoX: drain cr2-esams - [[phab:T431751|T431751]] * 13:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2153: Maintenance * 13:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2012.codfw.wmnet with reason: host reimage * 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2222: Maintenance * 13:45 sukhe: restart pybal on A:lvs-codfw * 13:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2012.codfw.wmnet with reason: host reimage * 13:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db2181: Maintenance * 13:44 btullis@dns1004: END - running authdns-update * 13:42 sukhe: restart pybal on lvs2014 * 13:42 btullis@dns1004: START - running authdns-update * 13:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2222 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95376 and previous config saved to /var/cache/conftool/dbconfig/20260728-133948-cwilliams.json * 13:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2222.codfw.wmnet with reason: Maintenance * 13:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2221: Maintenance * 13:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2181 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95374 and previous config saved to /var/cache/conftool/dbconfig/20260728-133857-cwilliams.json * 13:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2181.codfw.wmnet with reason: Maintenance * 13:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2167: Maintenance * 13:30 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T428495|T428495]] * 13:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 13:29 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 13:29 ayounsi@cumin1003: END (FAIL) - Cookbook sre.dns.admin (exit_code=99) DNS admin: depool esams [reason: router upgrade, [[phab:T431749|T431749]]] * 13:28 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: router upgrade, [[phab:T431749|T431749]]] * 13:27 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1022.eqiad.wmnet, repooling source-only afterwards * 13:27 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo - [[phab:T428495|T428495]] * 13:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2012 * 13:27 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2012 * 13:21 lucaswerkmeister-wmde@deploy1003: mwscript-k8s job started: cleanupTitles bolwiki # [[phab:T429951|T429951]] * 13:21 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] (duration: 07m 19s) * 13:20 swfrench-wmf: authdns-update to direct codfw, eqsin, ulsfo etcd clients to eqiad - [[phab:T428495|T428495]] * 13:18 swfrench@dns1004: END - running authdns-update * 13:17 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, anzx: Continuing with deployment * 13:16 swfrench@dns1004: START - running authdns-update * 13:16 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2012 * 13:16 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2012.codfw.wmnet 57.48.192.10.in-addr.arpa 7.5.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:16 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2012.codfw.wmnet 57.48.192.10.in-addr.arpa 7.5.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:16 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:16 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, anzx: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:14 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 13:14 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] * 13:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db2237: Maintenance * 13:13 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2237: Maintenance * 13:13 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:12 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:12 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback IPV6 for asw1-604 - pt1979@cumin2003" * 13:12 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:12 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback IPV6 for asw1-604 - pt1979@cumin2003" * 13:11 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 20s) * 13:11 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 13:10 esanders@deploy1003: Finished scap sync-world: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] (duration: 08m 11s) * 13:08 pt1979@cumin2003: START - Cookbook sre.dns.netbox * 13:07 root@cumin1003: START - Cookbook sre.mysql.pool pool db2237: Maintenance * 13:06 esanders@deploy1003: esanders: Continuing with deployment * 13:04 esanders@deploy1003: esanders: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2153: Maintenance * 13:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2153: Maintenance * 13:02 esanders@deploy1003: Started scap sync-world: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] * 13:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2237 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95362 and previous config saved to /var/cache/conftool/dbconfig/20260728-130107-cwilliams.json * 13:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2237.codfw.wmnet with reason: Maintenance * 13:00 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2236: Maintenance * 12:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2153: Maintenance * 12:52 root@cumin1003: START - Cookbook sre.mysql.pool pool db2221: Maintenance * 12:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2153 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95358 and previous config saved to /var/cache/conftool/dbconfig/20260728-125214-cwilliams.json * 12:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2153.codfw.wmnet with reason: Maintenance * 12:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2167: Maintenance * 12:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-codfw * 12:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2011.codfw.wmnet * 12:51 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2011.codfw.wmnet * 12:49 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:49 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback for asw1-603 - pt1979@cumin2003" * 12:48 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback for asw1-603 - pt1979@cumin2003" * 12:46 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2011.codfw.wmnet * 12:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2221 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95357 and previous config saved to /var/cache/conftool/dbconfig/20260728-124601-cwilliams.json * 12:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2221.codfw.wmnet with reason: Maintenance * 12:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2218: Maintenance * 12:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2167 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95354 and previous config saved to /var/cache/conftool/dbconfig/20260728-124457-cwilliams.json * 12:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2167.codfw.wmnet with reason: Maintenance * 12:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2166: Maintenance * 12:42 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 12:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2011.codfw.wmnet * 12:41 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2010.codfw.wmnet * 12:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2010.codfw.wmnet * 12:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2010.codfw.wmnet * 12:34 pt1979@cumin2003: START - Cookbook sre.dns.netbox * 12:32 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 12:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2010.codfw.wmnet * 12:31 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2009.codfw.wmnet * 12:31 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2009.codfw.wmnet * 12:27 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2009.codfw.wmnet * 12:22 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2009.codfw.wmnet * 12:22 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2008.codfw.wmnet * 12:21 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2008.codfw.wmnet * 12:16 pt1979@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-604-eqsin * 12:16 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2008.codfw.wmnet * 12:16 pt1979@cumin1003: START - Cookbook sre.network.tls for network device asw1-604-eqsin * 12:14 root@cumin1003: START - Cookbook sre.mysql.pool pool db2236: Maintenance * 12:14 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2236: Maintenance * 12:12 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1137.eqiad.wmnet with OS trixie * 12:11 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2008.codfw.wmnet * 12:11 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2007.codfw.wmnet * 12:11 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2007.codfw.wmnet * 12:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db2236: Maintenance * 12:06 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2007.codfw.wmnet * 12:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2236 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95348 and previous config saved to /var/cache/conftool/dbconfig/20260728-120253-cwilliams.json * 12:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2236.codfw.wmnet with reason: Maintenance * 12:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2007.codfw.wmnet * 12:01 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 12:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 11:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2218: Maintenance * 11:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db2166: Maintenance * 11:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2219: Maintenance * 11:56 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2006.codfw.wmnet * 11:52 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1137.eqiad.wmnet with reason: host reimage * 11:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2218 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95344 and previous config saved to /var/cache/conftool/dbconfig/20260728-115155-cwilliams.json * 11:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2218.codfw.wmnet with reason: Maintenance * 11:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2208: Maintenance * 11:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2166 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95342 and previous config saved to /var/cache/conftool/dbconfig/20260728-115119-cwilliams.json * 11:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2166.codfw.wmnet with reason: Maintenance * 11:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2164: Maintenance * 11:47 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1137.eqiad.wmnet with reason: host reimage * 11:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2006.codfw.wmnet * 11:45 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2005.codfw.wmnet * 11:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2005.codfw.wmnet * 11:40 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2005.codfw.wmnet * 11:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2005.codfw.wmnet * 11:35 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2004.codfw.wmnet * 11:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2004.codfw.wmnet * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1137 * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1137 * 11:30 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1137 * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1137.eqiad.wmnet 192.32.64.10.in-addr.arpa 2.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:30 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1137.eqiad.wmnet 192.32.64.10.in-addr.arpa 2.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1137 - jiji@cumin1003" * 11:25 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2004.codfw.wmnet * 11:19 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2004.codfw.wmnet * 11:19 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2003.codfw.wmnet * 11:19 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2003.codfw.wmnet * 11:14 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2003.codfw.wmnet * 11:11 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2219: Maintenance * 11:10 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2219: Maintenance * 11:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2219: Maintenance * 11:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2003.codfw.wmnet * 11:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2164: Maintenance * 11:03 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2002.codfw.wmnet * 11:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2002.codfw.wmnet * 11:03 root@cumin1003: START - Cookbook sre.mysql.pool pool db2208: Maintenance * 10:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2164 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95332 and previous config saved to /var/cache/conftool/dbconfig/20260728-105749-cwilliams.json * 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2164.codfw.wmnet with reason: Maintenance * 10:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2208 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95331 and previous config saved to /var/cache/conftool/dbconfig/20260728-105711-cwilliams.json * 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2208.codfw.wmnet with reason: Maintenance * 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2219 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95330 and previous config saved to /var/cache/conftool/dbconfig/20260728-105652-cwilliams.json * 10:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2219.codfw.wmnet with reason: Maintenance * 10:53 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1137 - jiji@cumin1003" * 10:52 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2002.codfw.wmnet * 10:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2002.codfw.wmnet * 10:47 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2001.codfw.wmnet * 10:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2001.codfw.wmnet * 10:39 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2001.codfw.wmnet * 10:35 jiji@cumin1003: START - Cookbook sre.dns.netbox * 10:34 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] (duration: 09m 31s) * 10:34 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1137 * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2001.codfw.wmnet * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-codfw * 10:34 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1137.eqiad.wmnet with OS trixie * 10:28 jforrester@deploy1003: jforrester: Continuing with deployment * 10:27 jforrester@deploy1003: jforrester: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:25 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] * 10:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 10:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 10:21 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 10:20 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1137.eqiad.wmnet * 10:20 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 10:20 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1137.eqiad.wmnet * 10:20 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1137.eqiad.wmnet * 10:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2163: Maintenance * 09:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-staging-worker * 09:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2003.codfw.wmnet * 09:37 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2003.codfw.wmnet * 09:32 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 09:31 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2003.codfw.wmnet * 09:30 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 09:30 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2163: Maintenance * 09:30 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 09:30 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 09:30 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:22 klausman@cumin1003: END (ERROR) - Cookbook sre.ganeti.reboot-vm (exit_code=97) for VM ml-serve-ctrl2001.codfw.wmnet * 09:22 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2001.codfw.wmnet * 09:22 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d8-eqiad * 09:22 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d8-eqiad * 09:21 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2003.codfw.wmnet * 09:20 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2002.codfw.wmnet * 09:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2002.codfw.wmnet * 09:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f2-codfw * 09:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f2-codfw * 09:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e4-codfw * 09:18 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2163: Maintenance * 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e4-codfw * 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-codfw * 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-codfw * 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e5-codfw * 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e5-codfw * 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f4-codfw * 09:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2210: Maintenance * 09:16 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f4-codfw * 09:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2182: Maintenance * 09:14 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2002.codfw.wmnet * 09:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db2163: Maintenance * 09:11 XioNoX: rebooting cr2-drmrs - [[phab:T431749|T431749]] * 09:10 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr2-drmrs,cr2-drmrs IPv6,cr2-drmrs.mgmt with reason: router upgrade * 09:06 XioNoX: draining cr2-drmrs - [[phab:T431749|T431749]] * 09:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2163 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95320 and previous config saved to /var/cache/conftool/dbconfig/20260728-090638-cwilliams.json * 09:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2163.codfw.wmnet with reason: Maintenance * 09:06 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2161: Maintenance * 09:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2002.codfw.wmnet * 09:04 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2001.codfw.wmnet * 09:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2001.codfw.wmnet * 08:57 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2001.codfw.wmnet * 08:48 XioNoX: un-drain cr1-drmrs - [[phab:T431749|T431749]] * 08:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2001.codfw.wmnet * 08:47 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-staging-worker * 08:42 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:35 XioNoX: rebooting cr1-drmrs - [[phab:T431749|T431749]] * 08:33 XioNoX: draining cr1-drmrs - [[phab:T431749|T431749]] * 08:31 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2210: Maintenance * 08:29 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2182: Maintenance * 08:21 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2182: Maintenance * 08:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db2161: Maintenance * 08:16 root@cumin1003: START - Cookbook sre.mysql.pool pool db2182: Maintenance * 08:12 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2210: Maintenance * 08:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2161 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95309 and previous config saved to /var/cache/conftool/dbconfig/20260728-081044-cwilliams.json * 08:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2161.codfw.wmnet with reason: Maintenance * 08:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2154: Maintenance * 08:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2182 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95307 and previous config saved to /var/cache/conftool/dbconfig/20260728-080947-cwilliams.json * 08:09 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2182.codfw.wmnet with reason: Maintenance * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2168: Maintenance * 08:06 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr1-drmrs,cr1-drmrs IPv6,cr1-drmrs.mgmt with reason: router upgrade * 08:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db2210: Maintenance * 08:05 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 08:05 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 08:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2210 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95305 and previous config saved to /var/cache/conftool/dbconfig/20260728-080008-cwilliams.json * 08:00 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2210.codfw.wmnet with reason: Maintenance * 07:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2206: Maintenance * 07:50 gkyziridis@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 07:50 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 07:22 root@cumin1003: START - Cookbook sre.mysql.pool pool db2154: Maintenance * 07:22 root@cumin1003: START - Cookbook sre.mysql.pool pool db2168: Maintenance * 07:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2154 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95295 and previous config saved to /var/cache/conftool/dbconfig/20260728-071640-cwilliams.json * 07:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2154.codfw.wmnet with reason: Maintenance * 07:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95294 and previous config saved to /var/cache/conftool/dbconfig/20260728-071604-cwilliams.json * 07:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2168.codfw.wmnet with reason: Maintenance * 07:08 root@cumin1003: START - Cookbook sre.mysql.pool pool db2206: Maintenance * 07:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2206 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95292 and previous config saved to /var/cache/conftool/dbconfig/20260728-070219-cwilliams.json * 07:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2206.codfw.wmnet with reason: Maintenance * 06:44 marostegui: Failover m5 from db1164 to db1228 - [[phab:T432967|T432967]] * 06:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2235].codfw.wmnet,db[1164,1217,1228].eqiad.wmnet with reason: m5 master switch [[phab:T432967|T432967]] * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.10 (duration: 02m 34s) * 03:39 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] (duration: 36m 06s) * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 02:57 dzahn@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1004.eqiad.wmnet with OS trixie * 02:57 dzahn@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - dzahn@cumin1003" * 02:55 dzahn@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - dzahn@cumin1003" * 02:37 dzahn@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1004.eqiad.wmnet with reason: host reimage * 02:31 dzahn@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1004.eqiad.wmnet with reason: host reimage * 02:16 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie * 02:15 dzahn@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host zuul1004.eqiad.wmnet with OS trixie * 01:43 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie * 01:43 dzahn@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1004.eqiad.wmnet with OS trixie * 01:25 pt1979@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-603-eqsin * 01:24 pt1979@cumin1003: START - Cookbook sre.network.tls for network device asw1-603-eqsin * 01:12 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 01:12 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt for new switches in eqsin - pt1979@cumin2003" * 01:12 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt for new switches in eqsin - pt1979@cumin2003" * 01:08 pt1979@cumin2003: START - Cookbook sre.dns.netbox * 00:48 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 00:47 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 00:47 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 00:47 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 00:26 mutante: attempting reimage with trixie on zuul1004 re-purposed physical hardware - dcops reported install issue - host was in busybox shell ([[phab:T427353|T427353]]) * 00:24 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie == 2026-07-27 == * 23:50 Amir1: mass deleting vp8 transcodes * 23:28 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:27 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1004.eqiad.wmnet with OS bullseye * 23:26 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 23:25 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:25 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 22:39 maryum: Deploy security fix for [[phab:T432877|T432877]] * 22:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1022.eqiad.wmnet with OS bookworm * 22:37 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS bullseye * 22:32 sbassett: Deployed security fix for [[phab:T432789|T432789]] * 22:22 sbassett: Deployed security patch for [[phab:T431819|T431819]] * 22:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1022.eqiad.wmnet with reason: host reimage * 22:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1022.eqiad.wmnet with reason: host reimage * 22:01 RScout-WMF: Deployed security fix for [[phab:T431819|T431819]] * 22:00 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2012 * 21:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2012.codfw.wmnet with OS bookworm * 21:55 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2011\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 21:45 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1022 * 21:45 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1022 * 21:44 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1022 * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1022.eqiad.wmnet 239.48.64.10.in-addr.arpa 9.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:44 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1022.eqiad.wmnet 239.48.64.10.in-addr.arpa 9.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:41 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:41 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 21:34 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS bookworm * 21:31 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:22 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:21 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:19 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:17 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1004 * 21:16 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1004 * 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1004] - vriley@cumin1003" * 21:15 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1004] - vriley@cumin1003" * 21:11 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:10 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2011.codfw.wmnet, repooling source-only afterwards * 21:05 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:01 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1022 * 20:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1022.eqiad.wmnet with OS bookworm * 20:53 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1021.eqiad.wmnet, repooling source-only afterwards * 20:51 mutante: zuul1001 - re-enabled puppet - revert "cherry-picked" gerrit:1314120 - [[phab:T431003|T431003]] * 20:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Maintenance * 20:15 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] (duration: 08m 03s) * 20:11 sbisson@deploy1003: sbisson: Continuing with deployment * 20:09 sbisson@deploy1003: sbisson: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] * 19:47 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 46s) * 19:47 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Maintenance * 19:27 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] (duration: 12m 26s) * 19:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2228 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95285 and previous config saved to /var/cache/conftool/dbconfig/20260727-192711-cwilliams.json * 19:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2228.codfw.wmnet with reason: Maintenance * 19:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2223: Maintenance * 19:23 krinkle@deploy1003: krinkle: Continuing with deployment * 19:16 krinkle@deploy1003: krinkle: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:15 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] * 19:12 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2238: Maintenance * 18:58 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 18:57 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 18:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2227: Maintenance * 18:57 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-experimental: apply * 18:55 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-experimental: apply * 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1021.eqiad.wmnet with OS bookworm * 18:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2011.codfw.wmnet with OS bookworm * 18:40 root@cumin1003: START - Cookbook sre.mysql.pool pool db2223: Maintenance * 18:39 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] (duration: 07m 05s) * 18:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2223 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95275 and previous config saved to /var/cache/conftool/dbconfig/20260727-183500-cwilliams.json * 18:34 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2223.codfw.wmnet with reason: Maintenance * 18:34 musikanimal@deploy1003: musikanimal: Continuing with deployment * 18:34 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2213: Maintenance * 18:33 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:32 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] * 18:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db2238: Maintenance * 18:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2011.codfw.wmnet with reason: host reimage * 18:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2238 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95271 and previous config saved to /var/cache/conftool/dbconfig/20260727-181944-cwilliams.json * 18:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2238.codfw.wmnet with reason: Maintenance * 18:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2226: Maintenance * 18:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1021.eqiad.wmnet with reason: host reimage * 18:14 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2011.codfw.wmnet with reason: host reimage * 18:12 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1021.eqiad.wmnet with reason: host reimage * 18:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db2227: Maintenance * 18:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2227 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95265 and previous config saved to /var/cache/conftool/dbconfig/20260727-180256-cwilliams.json * 18:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2227.codfw.wmnet with reason: Maintenance * 18:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2194: Maintenance * 17:57 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2011 * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2011 * 17:56 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2011 * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2011.codfw.wmnet 37.32.192.10.in-addr.arpa 7.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:56 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2011.codfw.wmnet 37.32.192.10.in-addr.arpa 7.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2011 - bking@cumin2003" * 17:56 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2011 - bking@cumin2003" * 17:52 bking@cumin2003: START - Cookbook sre.dns.netbox * 17:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2011 * 17:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1021 * 17:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1021 * 17:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2011.codfw.wmnet with OS bookworm * 17:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1021.eqiad.wmnet with OS bookworm * 17:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Maintenance * 17:38 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2010\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 17:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2213 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95260 and previous config saved to /var/cache/conftool/dbconfig/20260727-173740-cwilliams.json * 17:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2213.codfw.wmnet with reason: Maintenance * 17:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2211: Maintenance * 17:36 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1020\.eqiad\.wmnet,dc=eqiad,cluster=wdqs\-main,service=wdqs\-main * 17:32 root@cumin1003: START - Cookbook sre.mysql.pool pool db2226: Maintenance * 17:31 taavi@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] (duration: 06m 33s) * 17:27 taavi@deploy1003: taavi: Continuing with deployment * 17:27 taavi@deploy1003: taavi: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:26 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2226 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95256 and previous config saved to /var/cache/conftool/dbconfig/20260727-172636-cwilliams.json * 17:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2226.codfw.wmnet with reason: Maintenance * 17:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2225: Maintenance * 17:25 taavi@deploy1003: Started scap sync-world: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] * 17:13 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 17:11 root@cumin1003: START - Cookbook sre.mysql.pool pool db2194: Maintenance * 17:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2194 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95248 and previous config saved to /var/cache/conftool/dbconfig/20260727-170453-cwilliams.json * 17:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2194.codfw.wmnet with reason: Maintenance * 17:04 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2190: Maintenance * 16:52 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 16:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2211: Maintenance * 16:40 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95242 and previous config saved to /var/cache/conftool/dbconfig/20260727-164015-cwilliams.json * 16:40 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2211.codfw.wmnet with reason: Maintenance * 16:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2178: Maintenance * 16:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db2225: Maintenance * 16:39 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2172: Maintenance * 16:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 16:38 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 16:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2225 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95238 and previous config saved to /var/cache/conftool/dbconfig/20260727-163307-cwilliams.json * 16:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2225.codfw.wmnet with reason: Maintenance * 16:32 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2189: Maintenance * 16:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db2190: Maintenance * 16:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2190 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95230 and previous config saved to /var/cache/conftool/dbconfig/20260727-160602-cwilliams.json * 16:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2190.codfw.wmnet with reason: Maintenance * 15:53 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2177: Maintenance * 15:53 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2172: Maintenance * 15:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2178: Maintenance * 15:51 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2172: Maintenance * 15:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2178 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95224 and previous config saved to /var/cache/conftool/dbconfig/20260727-154559-cwilliams.json * 15:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2172: Maintenance * 15:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2178.codfw.wmnet with reason: Maintenance * 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2171: Maintenance * 15:44 root@cumin1003: START - Cookbook sre.mysql.pool pool db2189: Maintenance * 15:43 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:41 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2172 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95222 and previous config saved to /var/cache/conftool/dbconfig/20260727-153927-cwilliams.json * 15:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2172.codfw.wmnet with reason: Maintenance * 15:38 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2189 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95220 and previous config saved to /var/cache/conftool/dbconfig/20260727-153833-cwilliams.json * 15:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2189.codfw.wmnet with reason: Maintenance * 15:34 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:32 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 15:32 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 15:31 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:29 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:26 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 15:22 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] (duration: 07m 00s) * 15:21 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2155: Maintenance * 15:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2175: Maintenance * 15:18 zabe@deploy1003: zabe: Continuing with deployment * 15:17 zabe@deploy1003: zabe: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:15 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 15:15 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:15 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] * 15:15 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 06s) * 15:15 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:12 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2010.codfw.wmnet with OS bookworm * 15:08 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2177: Maintenance * 15:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2177: Maintenance * 14:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2177: Maintenance * 14:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2171: Maintenance * 14:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2171 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95209 and previous config saved to /var/cache/conftool/dbconfig/20260727-145236-cwilliams.json * 14:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2171.codfw.wmnet with reason: Maintenance * 14:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2177 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95208 and previous config saved to /var/cache/conftool/dbconfig/20260727-145206-cwilliams.json * 14:52 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2157: Maintenance * 14:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2177.codfw.wmnet with reason: Maintenance * 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1020.eqiad.wmnet with OS bookworm * 14:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2156: Maintenance * 14:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2010.codfw.wmnet with reason: host reimage * 14:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2010.codfw.wmnet with reason: host reimage * 14:41 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[1031,2024]*: Upgrade Cassandra to 5.0.8 (canary) - eevans@cumin1003 * 14:34 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2155: Maintenance * 14:33 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2175: Maintenance * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2010 * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2010 * 14:24 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2010 * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2010.codfw.wmnet 94.16.192.10.in-addr.arpa 4.9.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2010.codfw.wmnet 94.16.192.10.in-addr.arpa 4.9.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2010 - bking@cumin2003" * 14:24 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2010 - bking@cumin2003" * 14:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1020.eqiad.wmnet with reason: host reimage * 14:23 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[1031,2024]*: Upgrade Cassandra to 5.0.8 (canary) - eevans@cumin1003 * 14:20 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 14:20 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 14:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1020.eqiad.wmnet with reason: host reimage * 14:17 sukhe: sudo gnt-instance reboot urldownloader1005.wikimedia.org * 14:16 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:15 jelto@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:08 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2155: Maintenance * 14:05 root@cumin1003: START - Cookbook sre.mysql.pool pool db2157: Maintenance * 14:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2175: Maintenance * 14:03 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 11 hosts * 14:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db2155: Maintenance * 14:01 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 11 hosts * 14:01 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1136.eqiad.wmnet * 14:01 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1136.eqiad.wmnet * 14:01 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1136.eqiad.wmnet * 14:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db2156: Maintenance * 13:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2157 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95194 and previous config saved to /var/cache/conftool/dbconfig/20260727-135943-cwilliams.json * 13:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2157.codfw.wmnet with reason: Maintenance * 13:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db2175: Maintenance * 13:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 13:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 13:57 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2010 * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2155 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95193 and previous config saved to /var/cache/conftool/dbconfig/20260727-135613-cwilliams.json * 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2155.codfw.wmnet with reason: Maintenance * 13:55 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1020 * 13:55 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1020 * 13:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2156 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95192 and previous config saved to /var/cache/conftool/dbconfig/20260727-135413-cwilliams.json * 13:54 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2156.codfw.wmnet with reason: Maintenance * 13:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2175 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95191 and previous config saved to /var/cache/conftool/dbconfig/20260727-135300-cwilliams.json * 13:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2010.codfw.wmnet with OS bookworm * 13:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2175.codfw.wmnet with reason: Maintenance * 13:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1020.eqiad.wmnet with OS bookworm * 13:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 34 hosts * 13:46 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 34 hosts * 13:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1201: Maintenance * 13:27 Lucas_WMDE: UTC afternoon backport+config window doen * 13:18 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] (duration: 11m 57s) * 13:14 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, sihe: Continuing with deployment * 13:08 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, sihe: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:07 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool ulsfo [reason: router upgrade finished, [[phab:T431752|T431752]]] * 13:07 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool ulsfo [reason: router upgrade finished, [[phab:T431752|T431752]]] * 13:06 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] * 13:03 XioNoX: repool cr4-ulsfo - [[phab:T431752|T431752]] * 12:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db1201: Maintenance * 12:48 gkyziridis@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1201 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95186 and previous config saved to /var/cache/conftool/dbconfig/20260727-124404-cwilliams.json * 12:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1201.eqiad.wmnet with reason: Maintenance * 12:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1187: Maintenance * 12:30 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader.eqiad.wikimedia.org on all recursors * 12:30 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader.eqiad.wikimedia.org on all recursors * 12:30 sukhe@dns1004: END - running authdns-update * 12:30 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] (duration: 09m 32s) * 12:28 sukhe@dns1004: START - running authdns-update * 12:25 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 12:22 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:20 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] * 12:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts * 12:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts * 12:13 XioNoX: rebooting cr4-ulsfo for upgrade - [[phab:T431752|T431752]] * 12:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: es1038 repool * 12:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 38 hosts * 12:08 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 38 hosts * 11:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1187: Maintenance * 11:53 urbanecm@deploy1003: mwscript-k8s job started: foreachwikiindblist growthexperiments GrowthExperiments:cleanMentorList # [[phab:T431804|T431804]] * 11:50 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr4-ulsfo,cr4-ulsfo IPv6,cr4-ulsfo.mgmt with reason: router upgrade * 11:50 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] (duration: 11m 07s) * 11:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1187 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95178 and previous config saved to /var/cache/conftool/dbconfig/20260727-114844-cwilliams.json * 11:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1187.eqiad.wmnet with reason: Maintenance * 11:43 urbanecm@deploy1003: urbanecm: Continuing with deployment * 11:42 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:39 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] * 11:37 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 11:36 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 11:36 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 11:35 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 11:29 XioNoX: start draining cr4-ulsfo - [[phab:T431752|T431752]] * 11:29 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 11:29 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 11:28 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1035: testing * 11:28 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1035: testing * 11:27 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1035: testing * 11:27 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1035: testing * 11:26 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool ulsfo [reason: router upgrade, [[phab:T431752|T431752]]] * 11:26 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1038: es1038 repool * 11:26 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool ulsfo [reason: router upgrade, [[phab:T431752|T431752]]] * 11:26 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1038: testing * 11:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1264: Maintenance * 11:24 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1038: testing * 11:23 marostegui@cumin1003: dbctl commit (dc=all): 'Repool es1050 as master', diff saved to https://phabricator.wikimedia.org/P95170 and previous config saved to /var/cache/conftool/dbconfig/20260727-112326-marostegui.json * 11:23 marostegui@cumin1003: dbctl commit (dc=all): 'Repool es1050', diff saved to https://phabricator.wikimedia.org/P95169 and previous config saved to /var/cache/conftool/dbconfig/20260727-112302-marostegui.json * 11:22 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1050: testing * 11:22 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1050: testing * 11:20 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 11:18 blake@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 11:18 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 11:12 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 11:11 blake@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 11:09 blake@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 11:09 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 11:09 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 11:08 blake@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 11:05 blake@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 11:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 11:02 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 10:50 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 10:43 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:39 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply * 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1264: Maintenance * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply * 10:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply * 10:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 10:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 10:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 10:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1264 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95164 and previous config saved to /var/cache/conftool/dbconfig/20260727-103204-cwilliams.json * 10:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1264.eqiad.wmnet with reason: Maintenance * 10:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 10:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 10:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 10:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 10:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 10:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 10:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 10:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1237: Maintenance * 10:04 elukey: restart burrow main-eqiad on kafkamon2003 to clear some errors on kafka-main1008 * 09:58 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1136.eqiad.wmnet with OS trixie * 09:39 elukey: restart burrow-main-eqiad.service on kafkamon1003 to see if a recurrent kafka error on kafka-main1008 goes away * 09:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1237: Maintenance * 09:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1136.eqiad.wmnet with reason: host reimage * 09:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1237 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95159 and previous config saved to /var/cache/conftool/dbconfig/20260727-093328-cwilliams.json * 09:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1237.eqiad.wmnet with reason: Maintenance * 09:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1136.eqiad.wmnet with reason: host reimage * 09:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1203: Maintenance * 09:17 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1136 * 09:17 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1136 * 09:04 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1136 * 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1136.eqiad.wmnet 191.32.64.10.in-addr.arpa 1.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:04 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1136.eqiad.wmnet 191.32.64.10.in-addr.arpa 1.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1136 - jiji@cumin1003" * 09:04 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1136 - jiji@cumin1003" * 08:52 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 08:52 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 08:52 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 08:51 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 08:50 jiji@cumin1003: START - Cookbook sre.dns.netbox * 08:47 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1136 * 08:46 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1136.eqiad.wmnet with OS trixie * 08:44 marostegui: Rename tables on s3 [[phab:T425066|T425066]] * 08:43 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1136.eqiad.wmnet * 08:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db1203: Maintenance * 08:43 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1136.eqiad.wmnet * 08:43 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1136.eqiad.wmnet * 08:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1203 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95154 and previous config saved to /var/cache/conftool/dbconfig/20260727-083703-cwilliams.json * 08:36 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1203.eqiad.wmnet with reason: Maintenance * 08:16 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1179: Maintenance * 07:44 phuedx: UTC morning backport window done * 07:37 phuedx@deploy1003: Finished scap sync-world: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] (duration: 32m 33s) * 07:28 root@cumin1003: START - Cookbook sre.mysql.pool pool db1179: Maintenance * 07:26 marostegui: Rename tables on s3 [[phab:T426341|T426341]] * 07:25 phuedx@deploy1003: phuedx: Continuing with deployment * 07:22 marostegui: Drop tables in akwiki nawiki pihwiki - growthexperiments_* [[phab:T428885|T428885]] * 07:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95149 and previous config saved to /var/cache/conftool/dbconfig/20260727-072234-cwilliams.json * 07:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1179.eqiad.wmnet with reason: Maintenance * 07:20 phuedx@deploy1003: phuedx: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:16 ryankemper: [[phab:T430880|T430880]] [WDQS] Reimaged `wdqs1018` and `wdqs1019` to Bookworm, restored data using test-cookbook change {{Gerrit|1317128}}, and repooled both; 25/36 hosts complete * 07:04 phuedx@deploy1003: Started scap sync-world: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] * 06:57 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1019.eqiad.wmnet * 06:56 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1018.eqiad.wmnet * 06:40 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1020.eqiad.wmnet with reason: Cloning * 06:35 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db1228.eqiad.wmnet with reason: Rebooting * 06:29 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:29 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:25 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:25 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:25 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1019.eqiad.wmnet, repooling source-only afterwards * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1018.eqiad.wmnet, repooling source-only afterwards * 04:51 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1019.eqiad.wmnet, repooling source-only afterwards * 04:51 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1018.eqiad.wmnet, repooling source-only afterwards * 04:48 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s) * 04:48 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 04:48 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s) * 04:48 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 36s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-26 == * 14:59 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:59 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:59 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:59 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1019.eqiad.wmnet with OS bookworm * 01:05 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1018.eqiad.wmnet with OS bookworm * 00:43 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1019.eqiad.wmnet with reason: host reimage * 00:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1018.eqiad.wmnet with reason: host reimage * 00:34 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1019.eqiad.wmnet with reason: host reimage * 00:33 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1018.eqiad.wmnet with reason: host reimage * 00:16 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 00:16 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 00:15 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 00:15 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1019 * 00:11 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1019 * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1018 * 00:11 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1018 * 00:08 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1019.eqiad.wmnet with OS bookworm * 00:08 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1018.eqiad.wmnet with OS bookworm == 2026-07-25 == * 22:06 ryankemper: [[phab:T430880|T430880]] [WDQS] Repooled `wdqs1017` and `wdqs2024` after reimaging to bookworm, scap deploying, and data xfering * 22:04 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2024.codfw.wmnet * 22:03 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1017.eqiad.wmnet * 21:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1017.eqiad.wmnet, repooling source-only afterwards * 21:06 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2024.codfw.wmnet, repooling source-only afterwards * 20:52 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:52 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:52 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:52 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 20:18 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1017.eqiad.wmnet, repooling source-only afterwards * 20:18 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2024.codfw.wmnet, repooling source-only afterwards * 20:15 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:15 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:15 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:15 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 19:57 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s) * 19:57 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 19:57 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 07s) * 19:57 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 19:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2024.codfw.wmnet with OS bookworm * 19:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1017.eqiad.wmnet with OS bookworm * 19:02 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2024.codfw.wmnet with reason: host reimage * 18:58 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1017.eqiad.wmnet with reason: host reimage * 18:53 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2024.codfw.wmnet with reason: host reimage * 18:52 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1017.eqiad.wmnet with reason: host reimage * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2024 * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2024 * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1017 * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1017 * 18:27 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2024 * 18:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2024.codfw.wmnet 58.16.192.10.in-addr.arpa 8.5.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:26 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2024.codfw.wmnet 58.16.192.10.in-addr.arpa 8.5.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:24 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1017 * 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1017.eqiad.wmnet 238.48.64.10.in-addr.arpa 8.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:24 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1017.eqiad.wmnet 238.48.64.10.in-addr.arpa 8.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1017 - ryankemper@cumin2003" * 18:24 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1017 - ryankemper@cumin2003" * 18:23 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 18:18 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 18:17 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1017 * 18:17 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2024 * 18:14 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1017.eqiad.wmnet with OS bookworm * 18:14 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2024.codfw.wmnet with OS bookworm * 18:05 ryankemper: [WDQS] [[phab:T430880|T430880]] Reimaged `wdqs1016` and `wdqs2023` to Bookworm with `--move-vlan`, restored main and scholarly data, validated postflights, and repooled both hosts. Confirmed PyBal rebuilt both backends with their new addresses * 17:45 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2023.codfw.wmnet * 17:43 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1016.eqiad.wmnet * 06:35 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1016.eqiad.wmnet, repooling source-only afterwards * 06:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2023.codfw.wmnet, repooling source-only afterwards * 05:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2023.codfw.wmnet, repooling source-only afterwards * 05:19 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1016.eqiad.wmnet, repooling source-only afterwards * 05:07 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 07s) * 05:07 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 05:06 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 06s) * 05:06 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 03:27 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2023.codfw.wmnet with OS bookworm * 02:59 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2023.codfw.wmnet with reason: host reimage * 02:56 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2023.codfw.wmnet with reason: host reimage * 02:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2023 * 02:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2023 * 02:30 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2023 * 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2023.codfw.wmnet 35.0.192.10.in-addr.arpa 5.3.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 02:30 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2023.codfw.wmnet 35.0.192.10.in-addr.arpa 5.3.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2023 - ryankemper@cumin2003" * 02:30 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2023 - ryankemper@cumin2003" * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 26s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:15 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1016.eqiad.wmnet with OS bookworm * 00:49 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1016.eqiad.wmnet with reason: host reimage * 00:43 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1016.eqiad.wmnet with reason: host reimage * 00:31 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 00:27 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1016 * 00:27 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1016 * 00:27 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2023 * 00:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1016.eqiad.wmnet with OS bookworm * 00:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2023.codfw.wmnet with OS bookworm * 00:11 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs1014.eqiad.wmnet and wdqs2008.codfw.wmnet after Bookworm reimage, transfer, and postflight; wdqs2008 is serving, while wdqs1014 will remain outside of service until a pybal restart next monday == 2026-07-24 == * 23:54 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1014.eqiad.wmnet * 23:54 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2008.codfw.wmnet * 23:43 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2010.codfw.wmnet with OS trixie * 23:08 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 23:03 jhathaway@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 22:33 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 22:13 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 22:13 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 22:13 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 22:13 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:00 jhathaway@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 21:53 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 21:53 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie * 21:51 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 21:47 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie * 21:43 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 21:39 jhathaway@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 21:38 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 17:21 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1135.eqiad.wmnet * 17:21 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1135.eqiad.wmnet * 17:21 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1135.eqiad.wmnet * 16:34 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 16:34 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 16:34 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 16:34 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 16:33 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 16:33 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 16:28 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:28 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:28 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:28 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2008.codfw.wmnet, repooling source-only afterwards * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1014.eqiad.wmnet, repooling source-only afterwards * 15:56 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1135.eqiad.wmnet with OS trixie * 15:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 40 hosts * 15:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 40 hosts * 15:37 topranks: upgrade SR-Linux OS on lswtest-d8-eqiad * 15:36 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1135.eqiad.wmnet with reason: host reimage * 15:33 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 6 hosts with reason: upgrade lswtest-d8-eqiad * 15:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1135.eqiad.wmnet with reason: host reimage * 15:30 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc-gp2006.codfw.wmnet with OS bookworm * 15:15 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1135 * 15:15 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1135 * 15:13 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc-gp2006.codfw.wmnet with reason: host reimage * 15:08 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc-gp2006.codfw.wmnet with reason: host reimage * 14:49 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm * 14:48 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host mc-gp2006.codfw.wmnet with OS bookworm * 14:34 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] (duration: 41m 12s) * 14:32 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1135 * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1135.eqiad.wmnet 177.32.64.10.in-addr.arpa 7.7.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:32 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1135.eqiad.wmnet 177.32.64.10.in-addr.arpa 7.7.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1135 - jiji@cumin1003" * 14:32 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1135 - jiji@cumin1003" * 14:29 krinkle@deploy1003: krinkle: Continuing with deployment * 14:29 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm * 14:27 jiji@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host mc-gp2006.codfw.wmnet with OS bookworm * 14:26 jiji@cumin1003: START - Cookbook sre.dns.netbox * 14:15 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1135 * 14:14 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1135.eqiad.wmnet with OS trixie * 14:14 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1135.eqiad.wmnet * 14:13 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1135.eqiad.wmnet * 14:13 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1135.eqiad.wmnet * 14:10 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1072.eqiad.wmnet * 14:10 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1072.eqiad.wmnet * 14:10 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1072.eqiad.wmnet * 14:10 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1072.eqiad.wmnet * 14:09 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1071.eqiad.wmnet * 14:09 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1071.eqiad.wmnet * 14:09 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1071.eqiad.wmnet * 14:09 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1071.eqiad.wmnet * 13:58 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 13:58 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:58 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:57 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:55 krinkle@deploy1003: krinkle: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:53 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] * 13:45 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:45 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push new IPs for mc-gp2006 - cmooney@cumin1003" * 13:45 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push new IPs for mc-gp2006 - cmooney@cumin1003" * 13:44 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) mc-gp2006.codfw.wmnet on all recursors * 13:44 cmooney@cumin1003: START - Cookbook sre.dns.wipe-cache mc-gp2006.codfw.wmnet on all recursors * 13:42 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm * 13:41 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:30 papaul: reboot mr1-eqsin for maintenance * 13:24 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb[1029-1031].eqiad.wmnet * 13:10 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb[1029-1031].eqiad.wmnet * 11:33 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 7 hosts * 11:11 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 7 hosts * 10:56 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 7 hosts * 10:47 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 7 hosts * 10:44 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:44 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:41 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:41 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:35 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 8 hosts * 10:34 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:33 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:32 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:32 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:31 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:31 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:30 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 8 hosts * 10:24 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie * 10:19 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:18 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 16 hosts * 10:17 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2001.codfw.wmnet * 10:13 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2001.codfw.wmnet * 10:12 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2001.codfw.wmnet * 10:02 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2001.codfw.wmnet * 10:02 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2002.codfw.wmnet * 09:57 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2002.codfw.wmnet * 09:56 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2002.codfw.wmnet * 09:51 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2002.codfw.wmnet * 09:51 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1002.eqiad.wmnet * 09:47 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1002.eqiad.wmnet * 09:47 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1001.eqiad.wmnet * 09:44 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1001.eqiad.wmnet * 09:34 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2003.codfw.wmnet * 09:32 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2003.codfw.wmnet * 09:32 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2002.codfw.wmnet * 09:30 brouberol@dns1004: END - running authdns-update * 09:29 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2002.codfw.wmnet * 09:29 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2001.codfw.wmnet * 09:27 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 9 hosts * 09:27 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2001.codfw.wmnet * 09:27 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2001.codfw.wmnet * 09:26 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 9 hosts * 09:26 brouberol@dns1004: START - running authdns-update * 09:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 57 hosts * 09:24 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2001.codfw.wmnet * 09:24 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2002.codfw.wmnet * 09:22 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2002.codfw.wmnet * 09:21 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 57 hosts * 09:20 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2003.codfw.wmnet * 09:19 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 16 hosts * 09:16 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2003.codfw.wmnet * 09:16 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1003.eqiad.wmnet * 09:15 urbanecm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 09:15 urbanecm@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 09:13 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1003.eqiad.wmnet * 09:13 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1002.eqiad.wmnet * 09:11 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1002.eqiad.wmnet * 09:11 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1001.eqiad.wmnet * 09:07 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1001.eqiad.wmnet * 08:32 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:24 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:16 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 08:16 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 08:07 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:07 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:07 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 08:02 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 08:01 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:59 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:57 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 07:57 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 06:46 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1025.eqiad.wmnet with reason: Cloning * 06:46 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s4 * 06:45 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s6 * 06:44 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1019.eqiad.wmnet,service=s6 * 06:44 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1019.eqiad.wmnet,service=s4 * 03:40 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:40 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:40 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:40 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 03:37 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:37 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:37 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:36 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:49 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on mr1-eqsin,mr1-eqsin IPv6 with reason: connection issue * 02:38 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on cr[2-3]-eqsin.mgmt,ps1-[603-604]-eqsin with reason: connection issue * 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 27s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-23 == * 23:27 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin.oob,mr1-eqsin.oob IPv6 with reason: switch refresh * 22:21 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Setting storage compatibility to NONE - eevans@cumin1003 * 22:01 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Setting storage compatibility to NONE - eevans@cumin1003 * 21:29 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1014.eqiad.wmnet, repooling source-only afterwards * 21:28 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 46s) * 21:28 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 21:19 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Setting storage compatibility to UPGRADING - eevans@cumin1003 * 21:00 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Setting storage compatibility to UPGRADING - eevans@cumin1003 * 20:17 dani@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] (duration: 11m 57s) * 20:13 dani@deploy1003: dani: Continuing with deployment * 20:07 dani@deploy1003: dani: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:05 dani@deploy1003: Started scap sync-world: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] * 19:24 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:24 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating the rest of the ipv6 dns records. - jhancock@cumin2002" * 19:24 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating the rest of the ipv6 dns records. - jhancock@cumin2002" * 19:14 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 19:05 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wdqs1014.eqiad.wmnet with OS bookworm * 19:04 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.noop (exit_code=99) * 19:04 cwilliams@cumin1003: START - Cookbook sre.mysql.noop * 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1014.eqiad.wmnet with reason: host reimage * 18:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1014.eqiad.wmnet with reason: host reimage * 18:30 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2008.codfw.wmnet, repooling source-only afterwards * 18:28 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 19s) * 18:28 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1014 * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1014 * 18:22 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1014 * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1014.eqiad.wmnet 188.32.64.10.in-addr.arpa 8.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:22 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1014.eqiad.wmnet 188.32.64.10.in-addr.arpa 8.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1014 - bking@cumin2003" * 18:21 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1014 - bking@cumin2003" * 18:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2215: Maintenance * 18:18 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 18:15 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:15 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 18:06 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:06 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 18:05 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 18:04 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2052: codfw rack B8 re-pool after maintenance * 17:54 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 17:54 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:54 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 17:32 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2215: Maintenance * 17:29 cmooney@dns3003: END - running authdns-update * 17:27 cmooney@dns3003: START - running authdns-update * 17:23 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 17:22 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:18 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool es2052: codfw rack B8 re-pool after maintenance * 17:18 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2189: codfw rack B8 re-pool after maintenance * 17:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2215.codfw.wmnet with reason: Maintenance * 17:17 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 17:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2215 [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95126 and previous config saved to /var/cache/conftool/dbconfig/20260723-170903-cwilliams.json * 17:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2191 to x1 primary [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95125 and previous config saved to /var/cache/conftool/dbconfig/20260723-170612-cwilliams.json * 17:05 cezmunsta: Starting x1 codfw failover from db2215 to db2191 - [[phab:T432986|T432986]] * 16:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2191 with weight 0 [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95123 and previous config saved to /var/cache/conftool/dbconfig/20260723-165831-cwilliams.json * 16:58 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 16 hosts with reason: Primary switchover x1 [[phab:T432986|T432986]] * 16:36 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 138128 * 16:35 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 138128 * 16:33 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2189: codfw rack B8 re-pool after maintenance * 16:33 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2164: codfw rack B8 re-pool after maintenance * 16:28 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1072.eqiad.wmnet * 16:27 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1072.eqiad.wmnet with OS trixie * 16:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2249: Maintenance * 16:06 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker1072.eqiad.wmnet with reason: host reimage * 16:06 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1072.eqiad.wmnet with reason: host reimage * 15:50 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1072 * 15:50 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1072 * 15:49 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1072.eqiad.wmnet with OS trixie * 15:48 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2164: codfw rack B8 re-pool after maintenance * 15:48 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] (duration: 06m 37s) * 15:48 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2163: codfw rack B8 re-pool after maintenance * 15:45 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 15:43 musikanimal@deploy1003: musikanimal: Continuing with deployment * 15:43 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:41 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] * 15:36 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1072.eqiad.wmnet * 15:35 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1072.eqiad.wmnet * 15:35 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1072.eqiad.wmnet * 15:34 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:34 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push any outstanding updates - cmooney@cumin1003" * 15:34 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push any outstanding updates - cmooney@cumin1003" * 15:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db2249: Maintenance * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 15:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:26 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:21 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:21 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 15:21 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:21 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 15:20 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 15:19 cmooney@dns2004: END - running authdns-update * 15:17 cmooney@dns2004: START - running authdns-update * 15:14 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns2004.wikimedia.org * 15:12 brouberol@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 15:12 brouberol@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 15:12 klausman@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ml-serve1001.eqiad.wmnet with OS trixie * 15:11 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1071.eqiad.wmnet * 15:11 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1071.eqiad.wmnet with OS trixie * 15:10 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wdqs2008.codfw.wmnet with OS bookworm * 15:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2249.codfw.wmnet with reason: Maintenance * 15:08 brouberol@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 15:08 brouberol@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 15:08 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 15:08 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 15:06 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns1004.wikimedia.org * 15:02 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2002.codfw.wmnet * 15:02 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2002.codfw.wmnet * 15:02 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2163: codfw rack B8 re-pool after maintenance * 15:01 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 15:01 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 14:59 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test2001.codfw.wmnet * 14:57 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test2001.codfw.wmnet * 14:56 ryankemper: [WDQS] [[phab:T430880|T430880]] Reimaged `wdqs2016` to Bookworm, xferred scholarly_articles from `wdqs2024`, validated updater/readiness/federation, and repooled * 14:51 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2016.codfw.wmnet * 14:51 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1001.eqiad.wmnet with reason: host reimage * 14:48 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1071.eqiad.wmnet with reason: host reimage * 14:47 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1001.eqiad.wmnet with reason: host reimage * 14:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2008.codfw.wmnet with reason: host reimage * 14:43 topranks: reboot lsw1-b8-codw to upgrade JunOS [[phab:T430929|T430929]] * 14:41 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1071.eqiad.wmnet with reason: host reimage * 14:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2008.codfw.wmnet with reason: host reimage * 14:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2231: Maintenance * 14:30 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1001.eqiad.wmnet with OS trixie * 14:25 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2002.codfw.wmnet * 14:23 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1071 * 14:23 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1071 * 14:23 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 14:22 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore scholarly data after Bookworm reimage) xfer scholarly_articles from wdqs2024.codfw.wmnet -> wdqs2016.codfw.wmnet, repooling source-only afterwards * 14:22 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2052: codfw rack B8 depool for maintenance * 14:21 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool es2052: codfw rack B8 depool for maintenance * 14:21 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2249: codfw rack B8 depool for maintenance * 14:21 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1071 * 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1071.eqiad.wmnet 166.48.64.10.in-addr.arpa 6.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:21 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1071.eqiad.wmnet 166.48.64.10.in-addr.arpa 6.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1071 - jiji@cumin1003" * 14:21 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1071 - jiji@cumin1003" * 14:21 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2249: codfw rack B8 depool for maintenance * 14:21 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2189: codfw rack B8 depool for maintenance * 14:20 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2002.codfw.wmnet * 14:20 cmooney@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2050.codfw.wmnet * 14:20 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2189: codfw rack B8 depool for maintenance * 14:20 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2164: codfw rack B8 depool for maintenance * 14:20 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2164: codfw rack B8 depool for maintenance * 14:19 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2163: codfw rack B8 depool for maintenance * 14:19 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:19 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1014 * 14:19 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2163: codfw rack B8 depool for maintenance * 14:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1014.eqiad.wmnet with OS bookworm * 14:17 cmooney@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2050.codfw.wmnet * 14:16 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 14:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2008 * 14:14 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2008 * 14:14 cmooney@cumin1003: conftool action : set/pooled=no; selector: name=dns2004.wikimedia.org * 14:14 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2008.codfw.wmnet with OS bookworm * 14:13 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 14:12 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:10 topranks: depool dns2004 before lsw1-b8-codfw switch maintenance [[phab:T430929|T430929]] * 14:10 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b8-codfw,lsw1-b8-codfw IPv6,lsw1-b8-codfw.mgmt,ssw1-a[1,8]-codfw with reason: lsw1-b8-codfw JunOS upgrade * 14:07 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 30 hosts with reason: lsw1-b8-codfw JunOS upgrade * 14:06 elukey: upload python3-docker-report 0.0.19 to apt.wikimedia.org for bookworm and trixie * 13:59 jiji@cumin1003: START - Cookbook sre.dns.netbox * 13:58 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1071 * 13:57 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:57 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:57 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:57 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1071.eqiad.wmnet with OS trixie * 13:55 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1071.eqiad.wmnet * 13:55 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1071.eqiad.wmnet * 13:55 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1071.eqiad.wmnet * 13:53 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:52 logmsgbot: kharlan Deployed security patch for [[phab:T432948|T432948]] * 13:51 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:51 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 13:51 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db2231: Maintenance * 13:50 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:50 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:50 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:50 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:49 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:49 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:49 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2231 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95097 and previous config saved to /var/cache/conftool/dbconfig/20260723-134436-cwilliams.json * 13:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2231.codfw.wmnet with reason: Maintenance * 13:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:39 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:38 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 13:38 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] (duration: 09m 07s) * 13:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1037 hosts * 13:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2196: Maintenance * 13:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:33 kharlan@deploy1003: kharlan, emc-wmf: Continuing with deployment * 13:31 kharlan@deploy1003: kharlan, emc-wmf: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:30 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:28 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] * 13:17 hashar@deploy1003: Finished deploy [integration/docroot@2199146]: build: License GPL2.0+ / updating npm dependencies (duration: 00m 14s) * 13:17 hashar@deploy1003: Started deploy [integration/docroot@2199146]: build: License GPL2.0+ / updating npm dependencies * 13:14 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service * 13:07 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 12:58 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2207: Repooling * 12:49 root@cumin1003: START - Cookbook sre.mysql.pool pool db2196: Maintenance * 12:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2196 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95087 and previous config saved to /var/cache/conftool/dbconfig/20260723-123952-cwilliams.json * 12:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2196.codfw.wmnet with reason: Maintenance * 12:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2191: Maintenance * 12:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:13 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:13 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: Repooling * 12:12 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2207: Repooling * 12:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: Repooling * 11:56 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2235.codfw.wmnet with OS trixie * 11:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db2191: Maintenance * 11:46 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1070.eqiad.wmnet * 11:46 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1070.eqiad.wmnet * 11:46 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1070.eqiad.wmnet * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2191 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95080 and previous config saved to /var/cache/conftool/dbconfig/20260723-114308-cwilliams.json * 11:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2191.codfw.wmnet with reason: Maintenance * 11:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2186: Maintenance * 11:35 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 46375 * 11:34 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 46375 * 11:33 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2235.codfw.wmnet with reason: host reimage * 11:28 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2235.codfw.wmnet with reason: host reimage * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c7-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c7-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c6-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c6-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c5-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c5-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c4-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c4-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c3-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c3-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c2-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c2-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d7-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d7-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d4-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d4-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d3-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d2-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d2-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d8-eqiad * 11:23 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d8-eqiad * 11:23 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d1-eqiad * 11:23 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d1-eqiad * 11:12 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db2235.codfw.wmnet with OS trixie * 11:11 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:11 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[2160,2235].codfw.wmnet with reason: Upgrading * 10:56 root@cumin1003: START - Cookbook sre.mysql.pool pool db2186: Maintenance * 10:54 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1037: testing * 10:53 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1037: testing * 10:53 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1037: testing * 10:53 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1037: testing * 10:52 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: testing * 10:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2186 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95072 and previous config saved to /var/cache/conftool/dbconfig/20260723-104956-cwilliams.json * 10:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2186.codfw.wmnet with reason: Maintenance * 10:43 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1054: testing * 10:41 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1070.eqiad.wmnet with OS trixie * 10:30 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1037 hosts * 10:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 10:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 10:20 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1070.eqiad.wmnet with reason: host reimage * 10:16 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1070.eqiad.wmnet with reason: host reimage * 10:06 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1038: testing * 10:05 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1038: testing * 10:05 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1038: testing * 10:02 marostegui@dns1004: END - running authdns-update * 10:00 marostegui@dns1004: START - running authdns-update * 09:58 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: testing * 09:57 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1054: testing * 09:57 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1070 * 09:57 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1070 * 09:57 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1054: testing * 09:57 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1054: testing * 09:56 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1055: testing * 09:56 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1070 * 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1070.eqiad.wmnet 165.48.64.10.in-addr.arpa 5.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:56 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1070.eqiad.wmnet 165.48.64.10.in-addr.arpa 5.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1070 - jiji@cumin1003" * 09:56 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1070 - jiji@cumin1003" * 09:47 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for 1035 hosts * 09:45 jiji@cumin1003: START - Cookbook sre.dns.netbox * 09:42 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1070 * 09:42 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1070.eqiad.wmnet with OS trixie * 09:42 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1070.eqiad.wmnet * 09:41 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1070.eqiad.wmnet * 09:41 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1070.eqiad.wmnet * 09:27 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es2051: testing * 09:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: testing * 09:12 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es2051: testing * 09:11 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1055: testing * 09:09 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:09 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1055: testing * 09:09 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1055: testing * 08:50 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:50 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:50 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 08:49 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 08:49 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 08:49 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:46 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1069.eqiad.wmnet * 08:46 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1069.eqiad.wmnet * 08:46 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1069.eqiad.wmnet * 08:39 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 08:38 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 08:38 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 08:37 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 08:35 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:10 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1069.eqiad.wmnet with OS trixie * 07:49 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1069.eqiad.wmnet with reason: host reimage * 07:45 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1069.eqiad.wmnet with reason: host reimage * 07:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1035 hosts * 07:33 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 07:32 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 07:29 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1069 * 07:29 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1069 * 07:26 jiji@deploy1003: Finished scap sync-world: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules (duration: 06m 01s) * 07:25 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1069 * 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1069.eqiad.wmnet 164.48.64.10.in-addr.arpa 4.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:25 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1069.eqiad.wmnet 164.48.64.10.in-addr.arpa 4.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1069 - jiji@cumin1003" * 07:25 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1069 - jiji@cumin1003" * 07:25 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 07:25 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 07:24 jiji@deploy1003: jiji: Continuing with deployment * 07:22 jiji@deploy1003: jiji: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:21 jiji@deploy1003: Started scap sync-world: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules * 07:21 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1031.eqiad.wmnet,service=s7 * 07:20 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1031.eqiad.wmnet,service=s2 * 07:20 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1031.eqiad.wmnet,service=s7 * 07:20 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1031.eqiad.wmnet,service=s2 * 07:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts * 07:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts * 07:19 jiji@cumin1003: START - Cookbook sre.dns.netbox * 07:19 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1069 * 07:19 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1069.eqiad.wmnet with OS trixie * 07:19 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1069.eqiad.wmnet * 07:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 45 hosts * 07:17 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1069.eqiad.wmnet * 07:17 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1069.eqiad.wmnet * 07:14 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 45 hosts * 07:13 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 06:16 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs2007 after successful Bookworm reimage, data transfer, and postflight validation; wdqs1013 also passed postflights and is enabled in conftool, but remains out of IPVS pending a rolling pybal restart to clear its stale pre-VLAN-move address * 05:58 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2007.codfw.wmnet * 05:58 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1013.eqiad.wmnet * 05:54 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore scholarly data after Bookworm reimage) xfer scholarly_articles from wdqs2024.codfw.wmnet -> wdqs2016.codfw.wmnet, repooling source-only afterwards == 2026-07-22 == * 23:34 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Apply upgrade to JVM17 - eevans@cumin1003 * 23:14 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Apply upgrade to JVM17 - eevans@cumin1003 * 22:06 ryankemper: [WDQS] Added requestctl per-IP ratelimit `wdqs_heavy_sparql_bots_jul_2026_ratelimit` (chronic heavy-query bot tier driving deadlock-remediation restarts); pruned superseded `wdqs_2026_05_11_worobot` * 21:51 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] (duration: 11m 52s) * 21:44 sbassett@deploy1003: sbassett: Continuing with deployment * 21:43 sbassett@deploy1003: sbassett: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:39 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] * 20:38 dani@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] (duration: 32m 51s) * 20:38 ryankemper: [WDQS] Pruned obsolete requestctl action+pattern `wdqs_20260715_p2003_ring_ja3n` (actor rotated JA3Ns; rule inert) * 20:26 dani@deploy1003: dani, vadymts1: Continuing with deployment * 20:24 dani@deploy1003: dani, vadymts1: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:14 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2244: Testing * 20:06 dani@deploy1003: Started scap sync-world: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] * 19:56 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1013.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2016.codfw.wmnet with OS bookworm * 19:40 mutante: gerrit - one more service restart is needed - restarting * 19:29 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2244: Testing * 19:27 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2244: Testing * 19:27 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2244: Testing * 19:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2016.codfw.wmnet with reason: host reimage * 19:15 dancy@deploy1003: Finished deploy [zuul/deploy@d92e238]: Freshening Zuul installation (duration: 00m 15s) * 19:14 dancy@deploy1003: Started deploy [zuul/deploy@d92e238]: Freshening Zuul installation * 19:11 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2016.codfw.wmnet with reason: host reimage * 18:54 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1013.eqiad.wmnet, repooling source-only afterwards * 18:52 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2016 * 18:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2016 * 18:51 dancy@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 18:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2016.codfw.wmnet with OS bookworm * 18:39 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 11s) * 18:39 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 18:36 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 18:30 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417] (thin): Regular analytics weekly train THIN [analytics/refinery@2a25417d] (duration: 02m 09s) * 18:28 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417] (thin): Regular analytics weekly train THIN [analytics/refinery@2a25417d] * 18:28 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417]: Regular analytics weekly train [analytics/refinery@2a25417d] (duration: 04m 31s) * 18:27 dduvall: deploying https://gerrit.wikimedia.org/r/c/integration/config/+/1314025 (4 jobs updated) * 18:23 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417]: Regular analytics weekly train [analytics/refinery@2a25417d] * 18:22 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@2a25417d] (duration: 01m 59s) * 18:20 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@2a25417d] * 17:56 Raine: deployment server switchover => deploy1003 is primary now * 17:55 kamila@deploy1003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 28m 24s) * 17:54 mutante: restarting gerrit for maintenance * 17:29 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1013.eqiad.wmnet with OS bookworm * 17:27 kamila@deploy1003: Started scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] * 17:20 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] (duration: 22m 50s) * 17:12 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1023.eqiad.wmnet -> wdqs1024.eqiad.wmnet, repooling source-only afterwards * 17:04 Raine: point deployment.eqiad.wmnet to deploy1003 * 17:04 kamila@dns7001: END - running authdns-update * 17:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1013.eqiad.wmnet with reason: host reimage * 17:02 kamila@dns7001: START - running authdns-update * 17:01 kamila@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on releases2003.codfw.wmnet,releases1003.eqiad.wmnet with reason: Deployment server switchover * 17:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1013.eqiad.wmnet with reason: host reimage * 16:58 kamila@deploy2003: Locking from deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] * 16:57 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] (duration: 02m 33s) * 16:55 kamila@deploy2003: Locking from deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] * 16:55 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2003 - [[phab:T240266|T240266]] (duration: 00m 11s) * 16:54 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2003 - [[phab:T240266|T240266]] * 16:40 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1013 * 16:40 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1013 * 16:39 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1013 * 16:39 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1013.eqiad.wmnet 105.32.64.10.in-addr.arpa 5.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:39 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1013.eqiad.wmnet 105.32.64.10.in-addr.arpa 5.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:39 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:39 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1013 - bking@cumin2003" * 16:39 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1013 - bking@cumin2003" * 16:34 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:34 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1013 * 16:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1013.eqiad.wmnet with OS bookworm * 16:28 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1023.eqiad.wmnet -> wdqs1024.eqiad.wmnet, repooling source-only afterwards * 16:27 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-scholarly,name=eqiad * 16:27 eevans@deploy2003: helmfile [eqiad] DONE helmfile.d/services/linked-artifacts: apply * 16:26 eevans@deploy2003: helmfile [eqiad] START helmfile.d/services/linked-artifacts: apply * 16:26 eevans@deploy2003: helmfile [codfw] DONE helmfile.d/services/linked-artifacts: apply * 16:26 eevans@deploy2003: helmfile [codfw] START helmfile.d/services/linked-artifacts: apply * 16:25 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 29s) * 16:25 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 16:24 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 16:21 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 16:18 eevans@deploy2003: helmfile [codfw] DONE helmfile.d/services/linked-artifacts: apply * 16:18 eevans@deploy2003: helmfile [codfw] START helmfile.d/services/linked-artifacts: apply * 16:08 eevans@deploy2003: helmfile [staging] DONE helmfile.d/services/linked-artifacts: apply * 16:07 eevans@deploy2003: helmfile [staging] START helmfile.d/services/linked-artifacts: apply * 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 16:01 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 15:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1024.eqiad.wmnet with OS bookworm * 15:49 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] (duration: 00m 10s) * 15:49 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] * 15:48 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] (duration: 00m 15s) * 15:48 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] * 15:47 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] (duration: 00m 10s) * 15:47 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] * 15:46 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:42 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:42 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:40 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:37 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 15:37 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:36 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:36 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:36 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 15:33 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 15:33 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:31 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:28 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 15:27 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:27 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1024.eqiad.wmnet with reason: host reimage * 15:23 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:23 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:23 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:20 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1068.eqiad.wmnet * 15:20 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1068.eqiad.wmnet * 15:20 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1068.eqiad.wmnet * 15:20 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 15:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1024.eqiad.wmnet with reason: host reimage * 15:11 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:55 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 14:52 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wdqs1024.eqiad.wmnet with OS bookworm * 14:50 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] (duration: 00m 09s) * 14:50 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] * 14:49 jiji@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 14:49 jiji@deploy2003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 14:49 jiji@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 14:48 jiji@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 14:45 ecarg@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:45 ecarg@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:44 ecarg@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:44 ecarg@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:43 ecarg@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:43 ecarg@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:41 sukhe: ipvsadm --delete-service --tcp-service 10.2.1.55:8087: lvs2014 and lvs2013 * 14:39 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:39 sukhe: ipvsadm --delete-service --tcp-service 10.2.2.55:8087: [[phab:T432445|T432445]] * 14:38 ecarg@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:38 ecarg@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:37 ecarg@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:37 ecarg@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:36 ecarg@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:34 ecarg@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts datahubsearch1001.eqiad.wmnet * 14:32 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:32 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 14:31 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 14:31 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 14:28 sukhe: sudo cumin 'A:lvs-low-traffic-codfw' 'systemctl restart pybal': lvs2013 * 14:26 sukhe: sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal': lvs2014 * 14:26 sukhe: sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal' * 14:24 sukhe: restart pybal on lvs1019 * 14:24 sukhe: restart pybal on lvs1020 * 14:19 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for 1036 hosts * 14:17 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:04 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] (duration: 09m 28s) * 13:59 kharlan@deploy2003: dreamyjazz, kharlan: Continuing with deployment * 13:58 bking@cumin2003: START - Cookbook sre.hosts.decommission for hosts datahubsearch1001.eqiad.wmnet * 13:57 kharlan@deploy2003: dreamyjazz, kharlan: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:55 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] * 13:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts datahubsearch[1002-1003].eqiad.wmnet * 13:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:53 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch[1002-1003].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 13:52 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch[1002-1003].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 13:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 13:42 stran@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] (duration: 07m 30s) * 13:42 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:40 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: Test * 13:38 stran@deploy2003: dragoniez, stran: Continuing with deployment * 13:37 stran@deploy2003: dragoniez, stran: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:35 bking@cumin2003: START - Cookbook sre.hosts.decommission for hosts datahubsearch[1002-1003].eqiad.wmnet * 13:35 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024'] * 13:35 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 13:35 stran@deploy2003: Started scap sync-world: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] * 13:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 13:28 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024'] * 13:26 sukhe@dns1004: END - running authdns-update * 13:25 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 13:25 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 13:24 sukhe@dns1004: START - running authdns-update * 13:22 stran@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] (duration: 08m 20s) * 13:21 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 13:20 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 13:19 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 13:19 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 13:19 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 13:18 stran@deploy2003: stran: Continuing with deployment * 13:16 stran@deploy2003: stran: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:14 stran@deploy2003: Started scap sync-world: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] * 13:13 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 13:13 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 13:11 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 13:11 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 13:08 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 12:55 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool es2051: Test * 12:55 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: Test * 12:54 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool es2051: Test * 12:43 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1036 hosts * 12:41 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 12:40 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1048.eqiad.wmnet * 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 12:39 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 12:38 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 12:37 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 12:37 brouberol@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 12:36 brouberol@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 12:36 brouberol@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 12:36 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 12:36 elukey@cumin1003: DONE (PASS) - Cookbook sre.puppet.renew-cert (exit_code=0) for crm2001.codfw.wmnet: Renew puppet certificate - elukey@cumin1003 * 12:35 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:35 brouberol@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 12:34 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 12:31 brouberol@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 12:30 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1048.eqiad.wmnet * 12:30 brouberol@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 12:28 brouberol@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 12:27 brouberol@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 12:20 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1068.eqiad.wmnet with OS trixie * 12:01 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] (duration: 13m 19s) * 11:58 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1068.eqiad.wmnet with reason: host reimage * 11:52 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1068.eqiad.wmnet with reason: host reimage * 11:51 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 11:49 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:47 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] * 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2252: Security updates * 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:43 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 11:42 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2252: Security updates * 11:42 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply * 11:40 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply * 11:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 11:37 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2252.codfw.wmnet with OS trixie * 11:34 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1068 * 11:34 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1068 * 11:26 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1068 * 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1068.eqiad.wmnet 46.48.64.10.in-addr.arpa 6.4.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:26 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1068.eqiad.wmnet 46.48.64.10.in-addr.arpa 6.4.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1068 - jiji@cumin1003" * 11:26 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1068 - jiji@cumin1003" * 11:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2252.codfw.wmnet with reason: host reimage * 11:17 jiji@cumin1003: START - Cookbook sre.dns.netbox * 11:17 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1068 * 11:17 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1068.eqiad.wmnet with OS trixie * 11:17 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2252.codfw.wmnet with reason: host reimage * 11:15 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1068.eqiad.wmnet * 11:15 mvolz@deploy2003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:15 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1068.eqiad.wmnet * 11:15 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1068.eqiad.wmnet * 11:14 mvolz@deploy2003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:13 mvolz@deploy2003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:13 mvolz@deploy2003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:12 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] (duration: 11m 05s) * 11:11 mvolz@deploy2003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:10 mvolz@deploy2003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:07 dreamyjazz@deploy2003: dreamyjazz, kharlan: Continuing with deployment * 11:03 dreamyjazz@deploy2003: dreamyjazz, kharlan: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:03 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2252.codfw.wmnet with OS trixie * 11:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2252: Upgrading db2252.codfw.wmnet * 11:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:02 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 11:02 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2252: Upgrading db2252.codfw.wmnet * 11:02 cwilliams@cumin1003: dbmaint on ms3@codfw [[phab:T432321|T432321]] * 11:01 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 11:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db1153.eqiad.wmnet with reason: Security updates * 11:01 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] * 11:00 fnegri@deploy2003: helmfile [eqiad] DONE helmfile.d/services/toolhub: apply * 10:58 fnegri@deploy2003: helmfile [eqiad] START helmfile.d/services/toolhub: apply * 10:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1151: Security updates * 10:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:57 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1151: Security updates * 10:55 fnegri@deploy2003: helmfile [codfw] DONE helmfile.d/services/toolhub: apply * 10:54 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] (duration: 08m 38s) * 10:53 fnegri@deploy2003: helmfile [codfw] START helmfile.d/services/toolhub: apply * 10:53 fnegri@deploy2003: helmfile [staging] DONE helmfile.d/services/toolhub: apply * 10:52 fnegri@deploy2003: helmfile [staging] START helmfile.d/services/toolhub: apply * 10:50 zabe@deploy2003: zabe: Continuing with deployment * 10:47 zabe@deploy2003: zabe: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:45 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] * 10:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1151: Security updates * 10:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:42 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:42 root@cumin1003: START - Cookbook sre.mysql.depool depool db1151: Security updates * 10:38 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] (duration: 12m 47s) * 10:34 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 10:34 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 10:33 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2253.codfw.wmnet with OS trixie * 10:28 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:26 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] * 10:18 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2253.codfw.wmnet with reason: host reimage * 10:13 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2253.codfw.wmnet with reason: host reimage * 10:00 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2253.codfw.wmnet with OS trixie * 09:58 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db1151.eqiad.wmnet with reason: Security updates * 09:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2253: Upgrading db2253.codfw.wmnet * 09:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:57 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 09:56 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2253: Upgrading db2253.codfw.wmnet * 09:56 cwilliams@cumin1003: dbmaint on ms2@codfw [[phab:T432321|T432321]] * 09:56 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 09:36 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: UI improvement; support url shortener - oblivian@cumin1003" * 09:36 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: UI improvement; support url shortener - oblivian@cumin1003 * 09:35 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: UI improvement; support url shortener - oblivian@cumin1003 * 09:35 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: UI improvement; support url shortener - oblivian@cumin1003" * 09:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1152: Security updates * 09:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:26 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db1152: Security updates * 09:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: Security updates * 09:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:11 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:11 root@cumin1003: START - Cookbook sre.mysql.depool depool db1152: Security updates * 09:10 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1018.eqiad.wmnet with reason: Cloning * 09:09 Dreamy_Jazz: Deployed patch for [[phab:T432453|T432453]] and [[phab:T432454|T432454]] * 09:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 09:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2251.codfw.wmnet with OS trixie * 08:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2251.codfw.wmnet with reason: host reimage * 08:45 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2251.codfw.wmnet with reason: host reimage * 08:40 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] (duration: 12m 26s) * 08:38 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1030.eqiad.wmnet,service=s1 * 08:36 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 08:31 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2251.codfw.wmnet with OS trixie * 08:30 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:28 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] * 08:25 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] (duration: 07m 59s) * 08:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2251: Upgrading db2251.codfw.wmnet * 08:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:22 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 08:22 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2251: Upgrading db2251.codfw.wmnet * 08:20 urbanecm@deploy2003: urbanecm: Continuing with deployment * 08:20 cwilliams@cumin1003: dbmaint on ms1@codfw [[phab:T432321|T432321]] * 08:20 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 08:19 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:17 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] * 08:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade * 08:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade * 08:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db2251.codfw.wmnet,db1152.eqiad.wmnet with reason: OS upgrade * 08:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade * 08:13 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade * 08:11 Dreamy_Jazz: Created cusi_signal, cusi_case, and cusi_user on ukwiki and enwikivoyage in extension1 * 08:11 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade * 08:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade * 08:04 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1030.eqiad.wmnet,service=s1 * 08:04 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1030.eqiad.wmnet,service=s1 * 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply * 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply * 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply * 07:51 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply * 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 07:47 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 07:47 phuedx: End of UTC morning backport window * 07:43 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 07:43 phuedx@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] (duration: 13m 44s) * 07:43 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 07:39 phuedx@deploy2003: phuedx: Continuing with deployment * 07:31 phuedx@deploy2003: phuedx: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:29 phuedx@deploy2003: Started scap sync-world: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] * 07:24 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Turnilo import support - oblivian@cumin1003" * 07:24 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import support - oblivian@cumin1003 * 07:23 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import support - oblivian@cumin1003 * 07:23 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Turnilo import support - oblivian@cumin1003" * 06:42 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs2020 after successful Bookworm reimage, data transfer, and postflight validation * 06:42 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2020.codfw.wmnet * 05:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Managing sanitization for wikis bolwiki in section s5 * 05:25 marostegui@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis bolwiki in section s5 * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 41s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-21 == * 22:50 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2019.codfw.wmnet -> wdqs2020.codfw.wmnet, repooling source-only afterwards * 22:47 cwhite: force reboot arclamp2001 - appears to have run out of memory and gone unresponsive * 22:24 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 01m 26s) * 22:24 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 22:23 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 22:22 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024'] * 22:11 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 21:54 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs1024'] * 21:54 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 21:53 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs1024'] * 21:53 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 21:49 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2019.codfw.wmnet -> wdqs2020.codfw.wmnet, repooling source-only afterwards * 20:57 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] (duration: 09m 10s) * 20:55 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1024.eqiad.wmnet with OS bookworm * 20:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2020.codfw.wmnet with OS bookworm * 20:52 krinkle@deploy2003: krinkle: Continuing with deployment * 20:49 krinkle@deploy2003: krinkle: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:47 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] * 20:45 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] (duration: 05m 42s) * 20:44 krinkle@deploy2003: krinkle: Rolling back deployment * 20:41 krinkle@deploy2003: krinkle: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:39 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] * 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2003.codfw.wmnet * 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1003.eqiad.wmnet * 20:33 dani@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] (duration: 11m 15s) * 20:33 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2003.codfw.wmnet * 20:33 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1003.eqiad.wmnet * 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2020.codfw.wmnet with reason: host reimage * 20:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1002.eqiad.wmnet * 20:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2002.codfw.wmnet * 20:30 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 20:29 dani@deploy2003: dani: Continuing with deployment * 20:29 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2020.codfw.wmnet with reason: host reimage * 20:26 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1002.eqiad.wmnet * 20:26 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2002.codfw.wmnet * 20:24 dani@deploy2003: dani: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2001.codfw.wmnet * 20:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1001.eqiad.wmnet * 20:22 dani@deploy2003: Started scap sync-world: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] * 20:22 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 20:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2001.codfw.wmnet * 20:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1001.eqiad.wmnet * 20:14 sbisson@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] (duration: 09m 01s) * 20:11 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2020.codfw.wmnet with OS bookworm * 20:10 sbisson@deploy2003: sbisson: Continuing with deployment * 20:07 sbisson@deploy2003: sbisson: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:05 sbisson@deploy2003: Started scap sync-world: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] * 20:03 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] (duration: 07m 04s) * 20:01 mutante: Gerrit - tomorrow a new SSH host key will appear - it will be {{Gerrit|ed25519}} and has already been added to wmf-laptop. you can verify it here: https://wikitech.wikimedia.org/wiki/Help:SSH_Fingerprints/gerrit.wikimedia.org:29418 ([[phab:T240266|T240266]]) * 19:59 zabe@deploy2003: zabe: Continuing with deployment * 19:58 zabe@deploy2003: zabe: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:56 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] * 19:52 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] (duration: 07m 25s) * 19:48 zabe@deploy2003: zabe: Continuing with deployment * 19:47 zabe@deploy2003: zabe: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:45 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] * 19:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 19:32 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024'] * 19:27 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 19:26 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024'] * 19:26 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 19:24 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024'] * 19:12 ryankemper: [wdqs] [[phab:T430880|T430880]] Repooled `wdqs-scholarly` discovery in `eqiad` after validating `wdqs1023` end-to-end; `wdqs1024` remains disabled pending reimage recovery * 19:11 ryankemper: [wdqs] [[phab:T430880|T430880]] Repooled wdqs1012.eqiad.wmnet after successful Bookworm reimage, data transfer, service checks, readiness probe, and cross-graph federation query validation * 19:10 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 19:10 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1012.eqiad.wmnet * 19:08 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 18:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for deploy1003.eqiad.wmnet * 18:57 kamila@cumin1003: START - Cookbook sre.hosts.remove-downtime for deploy1003.eqiad.wmnet * 18:37 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:37 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding urldownloader service IPs - sukhe@cumin1003" * 18:37 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding urldownloader service IPs - sukhe@cumin1003" * 18:32 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 18:32 dancy@deploy2003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 18:30 sukhe@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 18:27 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 18:24 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1024.eqiad.wmnet with OS bookworm * 18:20 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host deploy1003.eqiad.wmnet with OS bookworm * 18:09 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deploy1003 reimage (duration: 121m 16s) * 18:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 18:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1180: Security updates * 17:55 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wcqs2003.codfw.wmnet * 17:48 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wcqs2003.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1155.eqiad.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1155.eqiad.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2224.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2224.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2217.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2217.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2193.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2193.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2180.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2180.codfw.wmnet * 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1168.eqiad.wmnet * 17:36 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1168.eqiad.wmnet * 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2169.codfw.wmnet * 17:36 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2169.codfw.wmnet * 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1165.eqiad.wmnet * 17:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1165.eqiad.wmnet * 17:35 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2158.codfw.wmnet * 17:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2158.codfw.wmnet * 17:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wcqs1003.eqiad.wmnet * 17:17 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1180: Security updates * 17:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1180.eqiad.wmnet * 17:16 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1180.eqiad.wmnet * 17:15 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp3073.* * 17:13 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wcqs1003.eqiad.wmnet * 17:11 brett@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp3073.esams.wmnet with OS trixie * 17:11 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 17:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1024 * 17:04 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1024 * 17:03 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 17:00 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 16:59 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2242: codfw rack B7 depool for maintenance * 16:59 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 16:43 brett@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp3073.esams.wmnet with reason: host reimage * 16:42 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 16:39 brett@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cp3073.esams.wmnet with reason: host reimage * 16:32 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on deploy1003.eqiad.wmnet with reason: host reimage * 16:27 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on deploy1003.eqiad.wmnet with reason: host reimage * 16:14 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2242: codfw rack B7 depool for maintenance * 16:14 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: codfw rack B7 depool for maintenance * 16:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1023.eqiad.wmnet with OS bookworm * 16:13 brett@cumin2002: START - Cookbook sre.hosts.reimage for host cp3073.esams.wmnet with OS trixie * 16:08 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host deploy1003.eqiad.wmnet with OS bookworm * 16:08 kamila@deploy2003: Locking from deployment [MediaWiki]: deploy1003 reimage * 16:03 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1012.eqiad.wmnet with OS bookworm * 15:48 inflatador: bking@apt1002 `sudo reprepro copy bookworm-wikimedia bullseye-wikimedia jvmquake` [[phab:T430880|T430880]] * 15:39 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp3073.* * 15:39 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 15:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:34 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1311536{{!}}Set $wgMathInternalRestbaseURL explicitly (take 2) (T349582)]] * 15:29 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:29 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2228: codfw rack B7 depool for maintenance * 15:29 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2229: codfw rack B7 depool for maintenance * 15:27 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1180: Security update * 15:25 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Security update * 15:21 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 15:21 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 15:19 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 24s) * 15:19 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:14 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 15:14 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db1180: Security update * 15:13 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-eqiad * 14:48 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-eqiad * 14:44 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2229: codfw rack B7 depool for maintenance * 14:44 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc2017: codfw rack B7 depool for maintenance * 14:44 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:43 cmooney@cumin2003: START - Cookbook sre.mysql.parsercache * 14:43 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool pc2017: codfw rack B7 depool for maintenance * 14:43 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2003.codfw.wmnet * 14:43 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2003.codfw.wmnet * 14:42 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2009.codfw.wmnet * 14:42 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2009.codfw.wmnet * 14:41 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:41 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:40 cmooney@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 29 hosts * 14:40 cmooney@cumin1003: START - Cookbook sre.hosts.remove-downtime for 29 hosts * 14:35 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 14:34 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 14:32 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] (duration: 07m 56s) * 14:29 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ssw1-a[1,8]-codfw with reason: lsw1-b7-codfw JunOS upgrade * 14:28 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 14:28 elukey@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 14:26 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:24 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] * 14:23 topranks: reboot lsw1-b7-codfw to upgrade JunOS (affects all hosts in rack) [[phab:T430928|T430928]] * 14:18 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2003.codfw.wmnet * 14:14 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2009.codfw.wmnet * 14:14 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Security update * 14:13 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2242: codfw rack B7 depool for maintenance * 14:13 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2242: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2228: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2228: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2229: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2229: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc2017: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.parsercache * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool pc2017: codfw rack B7 depool for maintenance * 14:08 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2003.codfw.wmnet * 14:07 cmooney@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on aux-k8s-etcd2004.codfw.wmnet,ml-etcd2001.codfw.wmnet with reason: lsw1-b7-codfw JunOS upgrade * 14:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95005 and previous config saved to /var/cache/conftool/dbconfig/20260721-140620-cwilliams.json * 14:05 cmooney@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2049.codfw.wmnet * 14:05 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-scholarly,name=eqiad * 14:04 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2009.codfw.wmnet * 14:04 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply * 14:04 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply * 14:03 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:03 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2001.codfw.wmnet * 14:03 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2001.codfw.wmnet * 14:02 cmooney@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2049.codfw.wmnet * 14:00 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:00 Dreamy_Jazz: Created cusi_case, cusi_signal, and cusi_user on svwiki, dewiki, jawiki, eswiki * 13:59 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b7-codfw,lsw1-b7-codfw IPv6,lsw1-b7-codfw.mgmt,ssw1-a[1,8]-codfw.mgmt with reason: lsw1-b7-codfw JunOS upgrade * 13:57 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs1023.eqiad.wmnet, repooling source-only afterwards * 13:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 29 hosts with reason: lsw1-b7-codfw JunOS upgrade * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224', diff saved to https://phabricator.wikimedia.org/P95003 and previous config saved to /var/cache/conftool/dbconfig/20260721-135613-cwilliams.json * 13:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1012.eqiad.wmnet with reason: host reimage * 13:53 cmooney@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 1:00:00 on 30 hosts with reason: lsw1-b7-codfw JunOS upgrade * 13:51 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1012.eqiad.wmnet with reason: host reimage * 13:48 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 13:48 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 13:46 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224', diff saved to https://phabricator.wikimedia.org/P95001 and previous config saved to /var/cache/conftool/dbconfig/20260721-134605-cwilliams.json * 13:46 cmooney@cumin1003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti2033.codfw.wmnet * 13:46 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 13:45 elukey: move the Docker Registry's /v2/wikimedia/machinelearning.* prefix to the ml S3 backend - [[phab:T428022|T428022]] * 13:45 cmooney@cumin1003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti2033.codfw.wmnet * 13:45 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 13:43 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:40 jiji@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 13:40 jiji@deploy2003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 13:39 jiji@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 13:39 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 13:38 cmooney@cumin1003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2032.codfw.wmnet * 13:38 jiji@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 13:38 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 13:37 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2032.codfw.wmnet * 13:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95000 and previous config saved to /var/cache/conftool/dbconfig/20260721-133557-cwilliams.json * 13:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1012 * 13:33 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1012 * 13:33 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1012.eqiad.wmnet with OS bookworm * 13:30 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:30 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:28 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94999 and previous config saved to /var/cache/conftool/dbconfig/20260721-132855-cwilliams.json * 13:28 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2224.codfw.wmnet with reason: Maintenance * 13:28 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94998 and previous config saved to /var/cache/conftool/dbconfig/20260721-132826-cwilliams.json * 13:28 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 13:23 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] (duration: 07m 50s) * 13:20 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 13:18 kharlan@deploy2003: kharlan: Continuing with deployment * 13:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217', diff saved to https://phabricator.wikimedia.org/P94996 and previous config saved to /var/cache/conftool/dbconfig/20260721-131817-cwilliams.json * 13:17 brouberol@dns1004: END - running authdns-update * 13:17 kharlan@deploy2003: kharlan: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:15 brouberol@dns1004: START - running authdns-update * 13:15 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] * 13:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94995 and previous config saved to /var/cache/conftool/dbconfig/20260721-131411-cwilliams.json * 13:13 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs1023.eqiad.wmnet, repooling source-only afterwards * 13:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217', diff saved to https://phabricator.wikimedia.org/P94994 and previous config saved to /var/cache/conftool/dbconfig/20260721-130809-cwilliams.json * 13:07 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 13:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180', diff saved to https://phabricator.wikimedia.org/P94993 and previous config saved to /var/cache/conftool/dbconfig/20260721-130404-cwilliams.json * 13:03 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:03 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:02 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 13:02 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 12:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94992 and previous config saved to /var/cache/conftool/dbconfig/20260721-125801-cwilliams.json * 12:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180', diff saved to https://phabricator.wikimedia.org/P94991 and previous config saved to /var/cache/conftool/dbconfig/20260721-125356-cwilliams.json * 12:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94990 and previous config saved to /var/cache/conftool/dbconfig/20260721-125049-cwilliams.json * 12:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2217.codfw.wmnet with reason: Maintenance * 12:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94989 and previous config saved to /var/cache/conftool/dbconfig/20260721-125017-cwilliams.json * 12:48 elukey: bmc cold reboot for lvs1013 and lvs1015 - [[phab:T426180|T426180]] * 12:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94988 and previous config saved to /var/cache/conftool/dbconfig/20260721-124348-cwilliams.json * 12:40 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193', diff saved to https://phabricator.wikimedia.org/P94987 and previous config saved to /var/cache/conftool/dbconfig/20260721-124009-cwilliams.json * 12:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts * 12:33 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts * 12:33 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts * 12:32 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts * 12:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:30 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193', diff saved to https://phabricator.wikimedia.org/P94986 and previous config saved to /var/cache/conftool/dbconfig/20260721-123001-cwilliams.json * 12:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94985 and previous config saved to /var/cache/conftool/dbconfig/20260721-121953-cwilliams.json * 12:17 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs2007.codfw.wmnet with OS bookworm * 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94983 and previous config saved to /var/cache/conftool/dbconfig/20260721-121257-cwilliams.json * 12:12 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2193.codfw.wmnet with reason: Maintenance * 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94982 and previous config saved to /var/cache/conftool/dbconfig/20260721-121239-cwilliams.json * 12:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180', diff saved to https://phabricator.wikimedia.org/P94980 and previous config saved to /var/cache/conftool/dbconfig/20260721-120231-cwilliams.json * 11:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180', diff saved to https://phabricator.wikimedia.org/P94979 and previous config saved to /var/cache/conftool/dbconfig/20260721-115223-cwilliams.json * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94978 and previous config saved to /var/cache/conftool/dbconfig/20260721-114333-cwilliams.json * 11:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1180.eqiad.wmnet with reason: Maintenance * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94977 and previous config saved to /var/cache/conftool/dbconfig/20260721-114305-cwilliams.json * 11:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94976 and previous config saved to /var/cache/conftool/dbconfig/20260721-114215-cwilliams.json * 11:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94975 and previous config saved to /var/cache/conftool/dbconfig/20260721-113530-cwilliams.json * 11:35 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2180.codfw.wmnet with reason: Maintenance * 11:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94974 and previous config saved to /var/cache/conftool/dbconfig/20260721-113501-cwilliams.json * 11:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168', diff saved to https://phabricator.wikimedia.org/P94973 and previous config saved to /var/cache/conftool/dbconfig/20260721-113258-cwilliams.json * 11:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169', diff saved to https://phabricator.wikimedia.org/P94972 and previous config saved to /var/cache/conftool/dbconfig/20260721-112453-cwilliams.json * 11:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168', diff saved to https://phabricator.wikimedia.org/P94971 and previous config saved to /var/cache/conftool/dbconfig/20260721-112250-cwilliams.json * 11:21 XioNoX: put eqiad-drmrs Arelion link in service * 11:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169', diff saved to https://phabricator.wikimedia.org/P94970 and previous config saved to /var/cache/conftool/dbconfig/20260721-111446-cwilliams.json * 11:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94969 and previous config saved to /var/cache/conftool/dbconfig/20260721-111242-cwilliams.json * 11:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 11:10 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 11:07 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1093 hosts * 11:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94968 and previous config saved to /var/cache/conftool/dbconfig/20260721-110548-cwilliams.json * 11:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1168.eqiad.wmnet with reason: Maintenance * 11:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94967 and previous config saved to /var/cache/conftool/dbconfig/20260721-110520-cwilliams.json * 11:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94966 and previous config saved to /var/cache/conftool/dbconfig/20260721-110439-cwilliams.json * 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94964 and previous config saved to /var/cache/conftool/dbconfig/20260721-105632-cwilliams.json * 10:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2169.codfw.wmnet with reason: Maintenance * 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94963 and previous config saved to /var/cache/conftool/dbconfig/20260721-105603-cwilliams.json * 10:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165', diff saved to https://phabricator.wikimedia.org/P94962 and previous config saved to /var/cache/conftool/dbconfig/20260721-105512-cwilliams.json * 10:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158', diff saved to https://phabricator.wikimedia.org/P94961 and previous config saved to /var/cache/conftool/dbconfig/20260721-104555-cwilliams.json * 10:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165', diff saved to https://phabricator.wikimedia.org/P94960 and previous config saved to /var/cache/conftool/dbconfig/20260721-104504-cwilliams.json * 10:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158', diff saved to https://phabricator.wikimedia.org/P94959 and previous config saved to /var/cache/conftool/dbconfig/20260721-103547-cwilliams.json * 10:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94958 and previous config saved to /var/cache/conftool/dbconfig/20260721-103456-cwilliams.json * 10:29 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2229: Upgraded kernel * 10:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94956 and previous config saved to /var/cache/conftool/dbconfig/20260721-102757-cwilliams.json * 10:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on an-redacteddb1001.eqiad.wmnet,clouddb[1015,1025,1028].eqiad.wmnet,db1155.eqiad.wmnet with reason: Maintenance * 10:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1165.eqiad.wmnet with reason: Maintenance * 10:25 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94955 and previous config saved to /var/cache/conftool/dbconfig/20260721-102539-cwilliams.json * 10:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94954 and previous config saved to /var/cache/conftool/dbconfig/20260721-101848-cwilliams.json * 10:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2158.codfw.wmnet with reason: Maintenance * 09:43 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2229: Upgraded kernel * 09:42 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2229.codfw.wmnet * 09:42 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2229.codfw.wmnet * 09:23 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db2229.codfw.wmnet * 09:23 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2229.codfw.wmnet * 08:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2229 [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94948 and previous config saved to /var/cache/conftool/dbconfig/20260721-085724-cwilliams.json * 08:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2214 to s6 primary [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94947 and previous config saved to /var/cache/conftool/dbconfig/20260721-085442-cwilliams.json * 08:53 cezmunsta: Starting s6 codfw failover from db2229 to db2214 - [[phab:T430964|T430964]] * 08:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2214 with weight 0 [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94946 and previous config saved to /var/cache/conftool/dbconfig/20260721-084613-cwilliams.json * 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 22 hosts with reason: Primary switchover s6 [[phab:T430964|T430964]] * 08:32 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1017.eqiad.wmnet,service=s1 * 08:08 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Add subrated circuit rate to interface descriptions - CR1312476 - ayounsi@cumin1003 * 08:06 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Add subrated circuit rate to interface descriptions - CR1312476 - ayounsi@cumin1003 * 07:58 reedy@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] (duration: 12m 55s) * 07:51 reedy@deploy2003: reedy, neriah: Continuing with deployment * 07:51 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 07:51 reedy@deploy2003: reedy, neriah: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:48 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1093 hosts * 07:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm2001.wikimedia.org * 07:45 reedy@deploy2003: Started scap sync-world: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] * 07:43 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 07:42 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm2001.wikimedia.org * 07:23 elukey: upgrade libtiff6 packages on zuul* trixie hosts for security upgrades * 07:22 elukey: upgrade libtiff6 packages on Wikikube trixie workers for security upgrades * 07:14 elukey@deploy2003: helmfile [codfw] DONE helmfile.d/services/proton: sync * 07:13 elukey@deploy2003: helmfile [codfw] START helmfile.d/services/proton: sync * 07:11 elukey@deploy2003: helmfile [eqiad] DONE helmfile.d/services/proton: sync * 07:10 elukey@deploy2003: helmfile [eqiad] START helmfile.d/services/proton: sync * 07:09 elukey@deploy2003: helmfile [staging] DONE helmfile.d/services/proton: sync * 07:08 elukey@deploy2003: helmfile [staging] START helmfile.d/services/proton: sync * 06:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1023.eqiad.wmnet with reason: host reimage * 06:46 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1023.eqiad.wmnet with reason: host reimage * 06:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 05:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Haproxy-only mode support - oblivian@cumin1003" * 05:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Haproxy-only mode support - oblivian@cumin1003 * 05:42 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Haproxy-only mode support - oblivian@cumin1003 * 05:42 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Haproxy-only mode support - oblivian@cumin1003" * 05:38 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1017.eqiad.wmnet with reason: Cloning * 05:37 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1017.eqiad.wmnet,service=s1 * 05:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1029.eqiad.wmnet,service=s8 * 05:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1029.eqiad.wmnet,service=s5 * 05:32 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:30 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 05:11 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet * 05:04 aokoth@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet * 05:00 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 04:56 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 04:01 mwpresync@deploy2003: Pruned MediaWiki: 1.47.0-wmf.9 (duration: 01m 08s) * 03:41 mwpresync@deploy2003: Finished scap sync-world: testwikis to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] (duration: 36m 30s) * 03:05 mwpresync@deploy2003: Started scap sync-world: testwikis to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 03:01 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:01 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:00 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:00 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:36 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:36 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:36 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:35 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:16 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 47s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 00:56 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm == 2026-07-20 == * 23:38 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 23:07 Amir1: deleting echo notifications from 2015 on group1 wikis * 23:07 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] (duration: 14m 16s) * 23:01 ladsgroup@deploy2003: ladsgroup: Continuing with deployment * 23:00 ladsgroup@deploy2003: ladsgroup: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:53 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] * 22:46 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2007.codfw.wmnet, repooling source-only afterwards * 22:39 maryum: Deployed security fixes for several security bugs * 21:42 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 21:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2007.codfw.wmnet, repooling source-only afterwards * 21:37 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 21:37 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 21:34 sbassett: Deployed security fix for [[phab:T432424|T432424]] * 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs2020.codfw.wmnet * 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1023.eqiad.wmnet * 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1011.eqiad.wmnet * 21:32 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 17s) * 21:32 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 21:27 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 21:13 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2007.codfw.wmnet with reason: host reimage * 21:08 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-internal-main,name=codfw * 21:06 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2007.codfw.wmnet with reason: host reimage * 20:59 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service * 20:58 sukhe: pybal restart for IP changes around wdqs-main hosts * 20:57 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 20:46 ryankemper@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-internal-main,name=codfw * 20:45 ebernhardson@deploy2003: Finished deploy [search/mjolnir/deploy@d4dc3b8]: Update for opensearch 2.x compat (duration: 00m 34s) * 20:44 ebernhardson@deploy2003: Started deploy [search/mjolnir/deploy@d4dc3b8]: Update for opensearch 2.x compat * 20:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2007 * 20:44 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2007 * 20:43 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2007 * 20:43 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2007.codfw.wmnet 156.16.192.10.in-addr.arpa 6.5.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:42 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2007.codfw.wmnet 156.16.192.10.in-addr.arpa 6.5.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:42 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2007 - bking@cumin2003" * 20:41 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2007 - bking@cumin2003" * 20:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94944 and previous config saved to /var/cache/conftool/dbconfig/20260720-203333-cwilliams.json * 20:32 arlolra@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] (duration: 15m 07s) * 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2020.codfw.wmnet * 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1023.eqiad.wmnet * 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1011.eqiad.wmnet * 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs2020.codfw.wmnet * 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1023.eqiad.wmnet * 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1011.eqiad.wmnet * 20:25 arlolra@deploy2003: arlolra, cscott: Continuing with deployment * 20:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257', diff saved to https://phabricator.wikimedia.org/P94943 and previous config saved to /var/cache/conftool/dbconfig/20260720-202325-cwilliams.json * 20:21 arlolra@deploy2003: arlolra, cscott: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:17 arlolra@deploy2003: Started scap sync-world: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] * 20:13 bking@cumin2003: START - Cookbook sre.dns.netbox * 20:13 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257', diff saved to https://phabricator.wikimedia.org/P94942 and previous config saved to /var/cache/conftool/dbconfig/20260720-201318-cwilliams.json * 20:13 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 20:10 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 20:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2007 * 20:04 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2007.codfw.wmnet with OS bookworm * 20:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94941 and previous config saved to /var/cache/conftool/dbconfig/20260720-200310-cwilliams.json * 19:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94940 and previous config saved to /var/cache/conftool/dbconfig/20260720-195633-cwilliams.json * 19:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1257.eqiad.wmnet with reason: Maintenance * 19:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94939 and previous config saved to /var/cache/conftool/dbconfig/20260720-195605-cwilliams.json * 19:51 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 19:50 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 19:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256', diff saved to https://phabricator.wikimedia.org/P94938 and previous config saved to /var/cache/conftool/dbconfig/20260720-194558-cwilliams.json * 19:44 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 19:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 19:41 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wdqs1011.eqiad.wmnet with OS bookworm * 19:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256', diff saved to https://phabricator.wikimedia.org/P94937 and previous config saved to /var/cache/conftool/dbconfig/20260720-193550-cwilliams.json * 19:25 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94936 and previous config saved to /var/cache/conftool/dbconfig/20260720-192542-cwilliams.json * 19:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94935 and previous config saved to /var/cache/conftool/dbconfig/20260720-191856-cwilliams.json * 19:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1256.eqiad.wmnet with reason: Maintenance * 19:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94934 and previous config saved to /var/cache/conftool/dbconfig/20260720-191839-cwilliams.json * 19:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255', diff saved to https://phabricator.wikimedia.org/P94933 and previous config saved to /var/cache/conftool/dbconfig/20260720-190831-cwilliams.json * 18:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255', diff saved to https://phabricator.wikimedia.org/P94932 and previous config saved to /var/cache/conftool/dbconfig/20260720-185824-cwilliams.json * 18:50 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 18:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94931 and previous config saved to /var/cache/conftool/dbconfig/20260720-184816-cwilliams.json * 18:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94930 and previous config saved to /var/cache/conftool/dbconfig/20260720-184224-cwilliams.json * 18:42 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1255.eqiad.wmnet with reason: Maintenance * 18:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94929 and previous config saved to /var/cache/conftool/dbconfig/20260720-184153-cwilliams.json * 18:39 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 18:39 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 16s) * 18:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 18:38 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 59m 26s) * 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 18:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211', diff saved to https://phabricator.wikimedia.org/P94928 and previous config saved to /var/cache/conftool/dbconfig/20260720-183145-cwilliams.json * 18:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211', diff saved to https://phabricator.wikimedia.org/P94927 and previous config saved to /var/cache/conftool/dbconfig/20260720-182137-cwilliams.json * 18:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94926 and previous config saved to /var/cache/conftool/dbconfig/20260720-181129-cwilliams.json * 18:09 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_codfw * 18:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2057.codfw.wmnet * 18:08 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_codfw * 18:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2058.codfw.wmnet * 18:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94925 and previous config saved to /var/cache/conftool/dbconfig/20260720-180452-cwilliams.json * 18:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on clouddb[1016,1020,1022-1023].eqiad.wmnet,db1154.eqiad.wmnet with reason: Maintenance * 18:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1211.eqiad.wmnet with reason: Maintenance * 18:02 sukhe: armed keyholder on acmechief1002.eqiad.wmnet and acmechief2002.codfw.wmnet (active host) * 18:01 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief2002.codfw.wmnet * 17:57 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief2002.codfw.wmnet * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs2020'] * 17:52 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief1002.eqiad.wmnet * 17:50 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 17:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1011.eqiad.wmnet with reason: host reimage * 17:48 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief1002.eqiad.wmnet * 17:47 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test2001.codfw.wmnet * 17:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94924 and previous config saved to /var/cache/conftool/dbconfig/20260720-174717-cwilliams.json * 17:46 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 17:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1011.eqiad.wmnet with reason: host reimage * 17:43 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs2020.codfw.wmnet with OS bookworm * 17:43 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test2001.codfw.wmnet * 17:43 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test1001.eqiad.wmnet * 17:39 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test1001.eqiad.wmnet * 17:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 17:38 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:38 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 17:37 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 17:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244', diff saved to https://phabricator.wikimedia.org/P94923 and previous config saved to /var/cache/conftool/dbconfig/20260720-173709-cwilliams.json * 17:35 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 17:31 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 20m 40s) * 17:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2055.codfw.wmnet * 17:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2056.codfw.wmnet * 17:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1011 * 17:27 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1011 * 17:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1011.eqiad.wmnet with OS bookworm * 17:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244', diff saved to https://phabricator.wikimedia.org/P94922 and previous config saved to /var/cache/conftool/dbconfig/20260720-172701-cwilliams.json * 17:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94921 and previous config saved to /var/cache/conftool/dbconfig/20260720-171653-cwilliams.json * 17:11 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 17:11 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 13m 03s) * 17:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94920 and previous config saved to /var/cache/conftool/dbconfig/20260720-171012-cwilliams.json * 17:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2244.codfw.wmnet with reason: Maintenance * 17:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94919 and previous config saved to /var/cache/conftool/dbconfig/20260720-170941-cwilliams.json * 16:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243', diff saved to https://phabricator.wikimedia.org/P94918 and previous config saved to /var/cache/conftool/dbconfig/20260720-165933-cwilliams.json * 16:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 16:58 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2053.codfw.wmnet * 16:51 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2054.codfw.wmnet * 16:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243', diff saved to https://phabricator.wikimedia.org/P94917 and previous config saved to /var/cache/conftool/dbconfig/20260720-164926-cwilliams.json * 16:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94916 and previous config saved to /var/cache/conftool/dbconfig/20260720-163918-cwilliams.json * 16:35 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 16:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94915 and previous config saved to /var/cache/conftool/dbconfig/20260720-163140-cwilliams.json * 16:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2243.codfw.wmnet with reason: Maintenance * 16:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94914 and previous config saved to /var/cache/conftool/dbconfig/20260720-163111-cwilliams.json * 16:27 btullis@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 16:27 btullis@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 16:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2020 * 16:23 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2020 * 16:21 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2020 * 16:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2020.codfw.wmnet 85.0.192.10.in-addr.arpa 5.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:21 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2020.codfw.wmnet 85.0.192.10.in-addr.arpa 5.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242', diff saved to https://phabricator.wikimedia.org/P94913 and previous config saved to /var/cache/conftool/dbconfig/20260720-162103-cwilliams.json * 16:19 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 16:18 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 16:18 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:18 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 16:17 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:17 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for netbox accounting errors - jhancock@cumin2002" * 16:17 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for netbox accounting errors - jhancock@cumin2002" * 16:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2051.codfw.wmnet * 16:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2052.codfw.wmnet * 16:11 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 16:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242', diff saved to https://phabricator.wikimedia.org/P94912 and previous config saved to /var/cache/conftool/dbconfig/20260720-161055-cwilliams.json * 16:09 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 16:08 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 16:06 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 16:06 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 16:06 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2020 * 16:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2020.codfw.wmnet with OS bookworm * 16:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94911 and previous config saved to /var/cache/conftool/dbconfig/20260720-160047-cwilliams.json * 15:58 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2019.codfw.wmnet, repooling source-only afterwards * 15:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94909 and previous config saved to /var/cache/conftool/dbconfig/20260720-155353-cwilliams.json * 15:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2242.codfw.wmnet with reason: Maintenance * 15:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94908 and previous config saved to /var/cache/conftool/dbconfig/20260720-154433-cwilliams.json * 15:35 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2049.codfw.wmnet * 15:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162', diff saved to https://phabricator.wikimedia.org/P94907 and previous config saved to /var/cache/conftool/dbconfig/20260720-153425-cwilliams.json * 15:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2050.codfw.wmnet * 15:28 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162', diff saved to https://phabricator.wikimedia.org/P94906 and previous config saved to /var/cache/conftool/dbconfig/20260720-152418-cwilliams.json * 15:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1023 * 15:14 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1023 * 15:14 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 15:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94905 and previous config saved to /var/cache/conftool/dbconfig/20260720-151407-cwilliams.json * 15:13 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] (duration: 41m 16s) * 15:08 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 15:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94902 and previous config saved to /var/cache/conftool/dbconfig/20260720-150729-cwilliams.json * 15:07 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2162.codfw.wmnet with reason: Maintenance * 15:05 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2027.codfw.wmnet, repooling source-only afterwards * 15:00 urbanecm@deploy2003: vadymts1, migr, urbanecm: Continuing with deployment * 14:59 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:58 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 07s) * 14:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:58 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 13s) * 14:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:57 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2019.codfw.wmnet, repooling source-only afterwards * 14:57 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2047.codfw.wmnet * 14:55 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2048.codfw.wmnet * 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2019.codfw.wmnet with OS bookworm * 14:49 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 14:47 urbanecm@deploy2003: vadymts1, migr, urbanecm: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:44 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool magru [reason: BGP issues in lvs7003 resolved after liberica restart, no task ID specified] * 14:44 sukhe@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool magru [reason: BGP issues in lvs7003 resolved after liberica restart, no task ID specified] * 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:41 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:39 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:39 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:33 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool magru [reason: no reason specified, no task ID specified] * 14:33 sukhe@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool magru [reason: no reason specified, no task ID specified] * 14:31 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] * 14:24 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:24 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:24 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:24 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2019.codfw.wmnet with reason: host reimage * 14:22 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2027.codfw.wmnet, repooling source-only afterwards * 14:19 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2019.codfw.wmnet with reason: host reimage * 14:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2027.codfw.wmnet with OS bookworm * 14:16 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2046.codfw.wmnet * 14:16 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2045.codfw.wmnet * 14:08 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:08 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:08 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:08 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:07 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:06 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:06 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:06 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:05 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2019 * 14:00 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2019 * 13:56 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1015.eqiad.wmnet * 13:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2027.codfw.wmnet with reason: host reimage * 13:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2071.codfw.wmnet with OS trixie * 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:51 sukhe@cumin1003: END (ERROR) - Cookbook sre.loadbalancer.admin (exit_code=97) rebooting A:liberica and P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica and P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:51 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1015.eqiad.wmnet * 13:50 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1014.eqiad.wmnet * 13:50 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1076.eqiad.wmnet with OS trixie * 13:50 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2027.codfw.wmnet with reason: host reimage * 13:45 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1014.eqiad.wmnet * 13:44 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1013.eqiad.wmnet * 13:39 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 13:38 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1013.eqiad.wmnet * 13:37 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2044.codfw.wmnet * 13:37 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2043.codfw.wmnet * 13:36 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2019 * 13:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2019.codfw.wmnet 156.32.192.10.in-addr.arpa 6.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:36 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2019.codfw.wmnet 156.32.192.10.in-addr.arpa 6.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:36 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2019 - bking@cumin2003" * 13:36 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2019 - bking@cumin2003" * 13:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry2005.codfw.wmnet * 13:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2071.codfw.wmnet with reason: host reimage * 13:31 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:31 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2019 * 13:31 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry2005.codfw.wmnet * 13:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry2004.codfw.wmnet * 13:30 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2019.codfw.wmnet with OS bookworm * 13:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2027 * 13:30 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2027 * 13:30 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2027.codfw.wmnet with OS bookworm * 13:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1076.eqiad.wmnet with reason: host reimage * 13:29 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_codfw * 13:28 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_codfw * 13:26 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry2004.codfw.wmnet * 13:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry1005.eqiad.wmnet * 13:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2071.codfw.wmnet with reason: host reimage * 13:22 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1076.eqiad.wmnet with reason: host reimage * 13:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry1005.eqiad.wmnet * 13:21 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry1004.eqiad.wmnet * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry1004.eqiad.wmnet * 13:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts * 13:13 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts * 13:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts * 13:12 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts * 13:03 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1076.eqiad.wmnet with OS trixie * 13:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2071.codfw.wmnet with OS trixie * 12:55 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:54 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:53 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:46 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 7 hosts * 12:42 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 7 hosts * 12:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts * 12:42 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts * 12:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2070.codfw.wmnet with OS trixie * 12:36 ozge@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:35 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1075.eqiad.wmnet with OS trixie * 12:32 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts * 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts * 12:22 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin1001.eqiad.wmnet * 12:19 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin1001.eqiad.wmnet * 12:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2070.codfw.wmnet with reason: host reimage * 12:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin2001.codfw.wmnet * 12:14 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1075.eqiad.wmnet with reason: host reimage * 12:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2070.codfw.wmnet with reason: host reimage * 12:10 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1075.eqiad.wmnet with reason: host reimage * 12:09 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin2001.codfw.wmnet * 11:17 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1074.eqiad.wmnet with OS trixie * 11:17 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 11:16 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 11:14 ozge@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 6 hosts * 11:09 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 6 hosts * 11:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 324 hosts * 10:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1074.eqiad.wmnet with reason: host reimage * 10:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2069.codfw.wmnet with OS trixie * 10:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1074.eqiad.wmnet with reason: host reimage * 10:30 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2069.codfw.wmnet with reason: host reimage * 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1074.eqiad.wmnet with OS trixie * 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2069.codfw.wmnet with reason: host reimage * 10:06 blake@deploy2003: Stopping before sync operations * 10:06 blake@deploy2003: Started scap sync-world: Non-deployment scap run to populate new release values * 10:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2069.codfw.wmnet with OS trixie * 10:00 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1073.eqiad.wmnet with OS trixie * 09:56 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 324 hosts * 09:39 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 09:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 8 hosts * 09:38 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1073.eqiad.wmnet with reason: host reimage * 09:37 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 8 hosts * 09:34 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1073.eqiad.wmnet with reason: host reimage * 09:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2068.codfw.wmnet with OS trixie * 09:16 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1073.eqiad.wmnet with OS trixie * 09:13 blake@deploy2003: sync-world aborted: Non-deployment scap run to populate new release values (duration: 00m 02s) * 09:13 blake@deploy2003: Started scap sync-world: Non-deployment scap run to populate new release values * 08:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2068.codfw.wmnet with reason: host reimage * 08:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2068.codfw.wmnet with reason: host reimage * 08:50 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 08:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2068.codfw.wmnet with OS trixie * 08:15 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1072.eqiad.wmnet with OS trixie * 07:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2067.codfw.wmnet with OS trixie * 07:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1072.eqiad.wmnet with reason: host reimage * 07:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1072.eqiad.wmnet with reason: host reimage * 07:45 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 07:45 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 07:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2067.codfw.wmnet with reason: host reimage * 07:35 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2067.codfw.wmnet with reason: host reimage * 07:30 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 07:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1072.eqiad.wmnet with OS trixie * 07:30 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 07:17 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 07:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2067.codfw.wmnet with OS trixie * 05:51 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:50 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:25 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:25 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on db2207.codfw.wmnet with reason: Host down * 04:28 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 07m 02s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-18 == * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 29s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 00:11 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2018.codfw.wmnet, repooling source-only afterwards == 2026-07-17 == * 23:53 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2026.codfw.wmnet, repooling source-only afterwards * 23:09 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2018.codfw.wmnet, repooling source-only afterwards * 23:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2026.codfw.wmnet, repooling source-only afterwards * 22:11 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2018.codfw.wmnet with OS bookworm * 22:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2026.codfw.wmnet with OS bookworm * 21:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2018.codfw.wmnet with reason: host reimage * 21:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2018.codfw.wmnet with reason: host reimage * 21:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2026.codfw.wmnet with reason: host reimage * 21:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2026.codfw.wmnet with reason: host reimage * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2018 * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2018 * 21:26 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2018 * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2018.codfw.wmnet 155.32.192.10.in-addr.arpa 5.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:26 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2018.codfw.wmnet 155.32.192.10.in-addr.arpa 5.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2018 - bking@cumin2003" * 21:26 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2018 - bking@cumin2003" * 21:14 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:13 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2018 * 21:13 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2018.codfw.wmnet with OS bookworm * 21:12 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2026 * 21:12 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2026 * 21:12 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2026.codfw.wmnet with OS bookworm * 21:05 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1022.eqiad.wmnet -> wdqs1026.eqiad.wmnet, repooling source-only afterwards * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs2017.codfw.wmnet, repooling source-only afterwards * 20:11 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs2017.codfw.wmnet, repooling source-only afterwards * 20:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2017.codfw.wmnet with OS bookworm * 20:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1022.eqiad.wmnet -> wdqs1026.eqiad.wmnet, repooling source-only afterwards * 20:06 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1026.eqiad.wmnet with OS bookworm * 19:55 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 19:55 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:55 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 09s) * 19:55 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:50 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 08s) * 19:50 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:50 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 10m 03s) * 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2017.codfw.wmnet with reason: host reimage * 19:40 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:40 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 15s) * 19:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1026.eqiad.wmnet with reason: host reimage * 19:37 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 16s) * 19:37 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:34 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2017.codfw.wmnet with reason: host reimage * 19:34 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1026.eqiad.wmnet with reason: host reimage * 19:33 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 19:33 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:16 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1026.eqiad.wmnet with OS bookworm * 19:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2017 * 19:16 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2017 * 19:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2017.codfw.wmnet with OS bookworm * 18:30 bking@dns1004: END - running authdns-update * 18:28 bking@dns1004: START - running authdns-update * 18:16 kamila@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1264.eqiad.wmnet * 18:16 kamila@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1264.eqiad.wmnet * 18:16 kamila@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1264.eqiad.wmnet * 17:49 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 17:46 dzahn@dns1006: END - running authdns-update * 17:44 dzahn@dns1006: START - running authdns-update * 17:44 dzahn@dns1006: END - running authdns-update * 17:42 dzahn@dns1006: START - running authdns-update * 17:28 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 17:21 kamila@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 17:01 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1264 * 17:01 kamila@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1264 * 17:01 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 17:01 kamila@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1264.eqiad.wmnet * 17:01 kamila@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1264.eqiad.wmnet * 17:01 kamila@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1264.eqiad.wmnet * 16:42 reedy@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] (duration: 10m 29s) * 16:34 reedy@deploy2003: reedy, hartman: Continuing with deployment * 16:33 reedy@deploy2003: reedy, hartman: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:31 reedy@deploy2003: Started scap sync-world: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] * 16:26 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 16:10 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in2001.wikimedia.org with reason: [[phab:T431659|T431659]] * 16:07 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in1001.wikimedia.org with reason: [[phab:T431659|T431659]] * 16:05 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 16:01 kamila@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 16:00 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out2001.wikimedia.org with reason: [[phab:T431659|T431659]] * 15:41 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 15:41 kamila@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 15:35 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out1001.wikimedia.org with reason: [[phab:T431659|T431659]] * 15:14 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1339.eqiad.wmnet * 15:13 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1339.eqiad.wmnet * 15:13 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1339.eqiad.wmnet * 14:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1339.eqiad.wmnet with OS trixie * 14:50 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:49 kamila@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:49 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:33 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage * 14:27 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage * 14:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1339 * 14:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1339 * 14:14 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1339 * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1339.eqiad.wmnet 156.32.64.10.in-addr.arpa 6.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:14 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1339.eqiad.wmnet 156.32.64.10.in-addr.arpa 6.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1339 - cgoubert@cumin2003" * 14:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1339 - cgoubert@cumin2003" * 14:09 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 14:06 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1339 * 14:06 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie * 14:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1339.eqiad.wmnet * 14:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1339.eqiad.wmnet * 14:02 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1339.eqiad.wmnet * 13:45 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb1013.eqiad.wmnet * 13:39 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb1013.eqiad.wmnet * 13:27 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:24 blake@dns1004: END - running authdns-update * 13:22 blake@dns1004: START - running authdns-update * 13:20 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 13:11 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2014.codfw.wmnet * 13:06 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb2014.codfw.wmnet * 13:06 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2012.codfw.wmnet * 13:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 13:03 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 13:01 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 15 hosts * 13:01 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb2012.codfw.wmnet * 13:01 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1016.eqiad.wmnet * 13:00 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 15 hosts * 12:55 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb1016.eqiad.wmnet * 12:55 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1014.eqiad.wmnet * 12:49 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb1014.eqiad.wmnet * 12:32 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:32 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:31 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:31 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1338.eqiad.wmnet * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1338.eqiad.wmnet * 12:18 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1338.eqiad.wmnet * 12:17 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:15 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:14 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:13 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1338.eqiad.wmnet with OS trixie * 12:01 klausman@deploy2003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 11:59 klausman@deploy2003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 11:56 klausman@deploy2003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 11:54 klausman@deploy2003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 11:53 klausman@deploy2003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 11:51 klausman@deploy2003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 11:42 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1338.eqiad.wmnet with reason: host reimage * 11:38 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1338.eqiad.wmnet with reason: host reimage * 11:31 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2230.codfw.wmnet * 11:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1338 * 11:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1338 * 11:25 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1338 * 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1338.eqiad.wmnet 155.32.64.10.in-addr.arpa 5.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:25 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1338.eqiad.wmnet 155.32.64.10.in-addr.arpa 5.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1338 - cgoubert@cumin2003" * 11:25 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1338 - cgoubert@cumin2003" * 11:23 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2230.codfw.wmnet * 11:20 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 11:20 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1338 * 11:20 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1338.eqiad.wmnet with OS trixie * 11:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1338.eqiad.wmnet * 11:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1338.eqiad.wmnet * 11:19 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1338.eqiad.wmnet * 11:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1337.eqiad.wmnet * 11:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1337.eqiad.wmnet * 11:17 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1337.eqiad.wmnet * 11:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1337.eqiad.wmnet with OS trixie * 10:51 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[2001-2002].codfw.wmnet * 10:50 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1337.eqiad.wmnet with reason: host reimage * 10:40 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 10:39 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:39 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1337.eqiad.wmnet with reason: host reimage * 10:39 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1001-1003].eqiad.wmnet * 10:34 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:34 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:30 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:28 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1001-1003].eqiad.wmnet * 10:27 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1337 * 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1337 * 10:26 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1337 * 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1337.eqiad.wmnet 154.32.64.10.in-addr.arpa 4.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:26 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1337.eqiad.wmnet 154.32.64.10.in-addr.arpa 4.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1337 - cgoubert@cumin2003" * 10:26 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1337 - cgoubert@cumin2003" * 10:21 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 10:18 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1337 * 10:17 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1337.eqiad.wmnet with OS trixie * 10:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1337.eqiad.wmnet * 10:16 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db1176.eqiad.wmnet * 10:16 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1337.eqiad.wmnet * 10:16 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1337.eqiad.wmnet * 10:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1336.eqiad.wmnet * 10:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1336.eqiad.wmnet * 10:15 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1336.eqiad.wmnet * 10:11 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db1176.eqiad.wmnet * 10:10 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db1176.eqiad.wmnet * 10:09 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db1176.eqiad.wmnet * 10:05 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts (check the cookbook's logs for more details.) * 10:03 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts (check the cookbook's logs for more details.) * 09:58 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1336.eqiad.wmnet with OS trixie * 09:47 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts (check the cookbook's logs for more details.) * 09:47 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts (check the cookbook's logs for more details.) * 09:45 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host acmechief-test2001.codfw.wmnet,acmechief-test1001.eqiad.wmnet,an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet,db-test[2001-2002].codfw.wmnet,db-test[1 * 09:40 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host acmechief-test2001.codfw.wmnet,acmechief-test1001.eqiad.wmnet,an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet,db-test[2001-2002].codfw.wmnet,db-test[1001-1003].eqiad.wmn * 09:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1336.eqiad.wmnet with reason: host reimage * 09:33 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1336.eqiad.wmnet with reason: host reimage * 09:29 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet * 09:29 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet * 09:28 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:26 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:21 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 09:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1336 * 09:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1336 * 09:19 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 09:14 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1336 * 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1336.eqiad.wmnet 152.32.64.10.in-addr.arpa 2.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1336.eqiad.wmnet 152.32.64.10.in-addr.arpa 2.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1336 - cgoubert@cumin2003" * 09:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1336 - cgoubert@cumin2003" * 09:11 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:10 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 09:09 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:09 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1336 * 09:09 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1336.eqiad.wmnet with OS trixie * 09:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1336.eqiad.wmnet * 09:08 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1336.eqiad.wmnet * 09:08 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1336.eqiad.wmnet * 09:06 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1335.eqiad.wmnet * 09:06 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1335.eqiad.wmnet * 09:06 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1335.eqiad.wmnet * 09:04 elukey: uploaded spicerack_13.1.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia * 08:55 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wikikube-worker-exp2001.codfw.wmnet * 08:54 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host testreduce1002.eqiad.wmnet * 08:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1335.eqiad.wmnet with OS trixie * 08:51 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host wikikube-worker-exp2001.codfw.wmnet * 08:51 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wikikube-worker-exp1001.eqiad.wmnet * 08:50 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host testreduce1002.eqiad.wmnet * 08:45 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host wikikube-worker-exp1001.eqiad.wmnet * 08:34 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1335.eqiad.wmnet with reason: host reimage * 08:30 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1335.eqiad.wmnet with reason: host reimage * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1335 * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1335 * 08:18 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1335 * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1335.eqiad.wmnet 150.32.64.10.in-addr.arpa 0.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:18 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1335.eqiad.wmnet 150.32.64.10.in-addr.arpa 0.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1335 - cgoubert@cumin2003" * 08:18 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1335 - cgoubert@cumin2003" * 08:14 elukey@cumin1003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 08:14 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges * 08:13 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 08:10 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1335 * 08:10 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1335.eqiad.wmnet with OS trixie * 08:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1335.eqiad.wmnet * 08:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1335.eqiad.wmnet * 08:09 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1335.eqiad.wmnet * 08:06 elukey@cumin1003: END (FAIL) - Cookbook sre.puppet.disable-merges (exit_code=99) * 08:05 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges * 08:03 elukey@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin1003.eqiad.wmnet * 07:57 elukey@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin1003.eqiad.wmnet * 07:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetdb1003.eqiad.wmnet * 07:46 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetdb1003.eqiad.wmnet * 07:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetdb2003.codfw.wmnet * 07:37 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetdb2003.codfw.wmnet * 07:37 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1001.eqiad.wmnet * 07:28 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver1001.eqiad.wmnet * 07:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet * 07:19 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet * 07:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2002.codfw.wmnet * 07:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver2002.codfw.wmnet * 07:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet * 07:05 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet * 07:04 btullis@cumin1003: END (FAIL) - Cookbook sre.hadoop.reboot-workers (exit_code=99) for Hadoop analytics cluster * 07:04 elukey@cumin1003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 07:04 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges * 06:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox1003.eqiad.wmnet * 06:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox1003.eqiad.wmnet * 02:46 ryankemper@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:46 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:44 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:37 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:37 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-internal-scholarly,name=eqiad * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 49s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 01:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore wdqs1025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling source-only afterwards * 01:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore wdqs1027 after Bookworm reimage) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs1027.eqiad.wmnet, repooling both afterwards * 00:55 urbanecm@deploy2003: helmfile [codfw] DONE helmfile.d/services/linkrecommendation: apply * 00:54 urbanecm@deploy2003: helmfile [eqiad] DONE helmfile.d/services/linkrecommendation: apply * 00:54 urbanecm@deploy2003: helmfile [staging] DONE helmfile.d/services/linkrecommendation: apply * 00:54 urbanecm@deploy2003: helmfile [codfw] START helmfile.d/services/linkrecommendation: apply * 00:53 urbanecm@deploy2003: helmfile [staging] START helmfile.d/services/linkrecommendation: apply * 00:52 urbanecm@deploy2003: helmfile [eqiad] START helmfile.d/services/linkrecommendation: apply * 00:23 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore wdqs1025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling source-only afterwards * 00:23 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore wdqs1027 after Bookworm reimage) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs1027.eqiad.wmnet, repooling both afterwards * 00:14 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1274.eqiad.wmnet * 00:14 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1274.eqiad.wmnet * 00:14 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1274.eqiad.wmnet * 00:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1027.eqiad.wmnet with OS bookworm * 00:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1025.eqiad.wmnet with OS bookworm * 00:04 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1274.eqiad.wmnet with OS trixie == 2026-07-16 == * 23:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], xfer to freshly reimaged/scap-deployed wdqs2025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs2025.codfw.wmnet, repooling source-only afterwards * 23:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1027.eqiad.wmnet with reason: host reimage * 23:47 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1025.eqiad.wmnet with reason: host reimage * 23:43 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1274.eqiad.wmnet with reason: host reimage * 23:41 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1025.eqiad.wmnet with reason: host reimage * 23:39 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1027.eqiad.wmnet with reason: host reimage * 23:38 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1274.eqiad.wmnet with reason: host reimage * 23:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1025 * 23:23 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1025 * 23:22 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1027 * 23:22 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1027 * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1274 * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1274 * 23:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1025.eqiad.wmnet with OS bookworm * 23:19 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1274 * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1274.eqiad.wmnet 145.48.64.10.in-addr.arpa 5.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:19 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1274.eqiad.wmnet 145.48.64.10.in-addr.arpa 5.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1274 - swfrench@cumin1003" * 23:19 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1274 - swfrench@cumin1003" * 23:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1027.eqiad.wmnet with OS bookworm * 23:14 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 23:14 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1274 * 23:13 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1274.eqiad.wmnet with OS trixie * 23:13 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1274.eqiad.wmnet * 23:12 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1274.eqiad.wmnet * 23:12 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1274.eqiad.wmnet * 23:12 ryankemper: [[phab:T430880|T430880]] depooled dnsdisc of wdqs-internal-scholarly-eqiad bc we only have 1 host there * 23:09 ryankemper@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-internal-scholarly,name=eqiad * 23:08 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1272.eqiad.wmnet * 23:08 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1272.eqiad.wmnet * 23:08 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1272.eqiad.wmnet * 23:01 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], xfer to freshly reimaged/scap-deployed wdqs2025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs2025.codfw.wmnet, repooling source-only afterwards * 22:57 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1272.eqiad.wmnet with OS trixie * 22:56 Amir1: deleting echo notifications from 2015 in group0 * 22:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2025.codfw.wmnet with OS bookworm * 22:35 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1272.eqiad.wmnet with reason: host reimage * 22:32 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 27s) * 22:32 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 22:28 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1269.eqiad.wmnet * 22:28 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1269.eqiad.wmnet * 22:28 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1269.eqiad.wmnet * 22:27 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1272.eqiad.wmnet with reason: host reimage * 22:26 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] (duration: 08m 51s) * 22:22 ladsgroup@deploy2003: ladsgroup, urbanecm: Continuing with deployment * 22:19 ladsgroup@deploy2003: ladsgroup, urbanecm: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:17 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] * 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2025.codfw.wmnet with reason: host reimage * 22:06 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1272 * 22:06 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1272 * 22:05 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1272 * 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1272.eqiad.wmnet 127.48.64.10.in-addr.arpa 7.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:05 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1272.eqiad.wmnet 127.48.64.10.in-addr.arpa 7.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1272 - swfrench@cumin1003" * 22:05 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1272 - swfrench@cumin1003" * 22:03 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2025.codfw.wmnet with reason: host reimage * 22:01 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 22:00 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1272 * 22:00 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1272.eqiad.wmnet with OS trixie * 22:00 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1272.eqiad.wmnet * 21:59 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1272.eqiad.wmnet * 21:59 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1272.eqiad.wmnet * 21:55 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1271.eqiad.wmnet * 21:55 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1271.eqiad.wmnet * 21:55 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1271.eqiad.wmnet * 21:46 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1271.eqiad.wmnet with OS trixie * 21:45 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] (duration: 06m 31s) * 21:40 sbassett@deploy2003: sbassett: Continuing with deployment * 21:40 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2025 * 21:40 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2025 * 21:40 sbassett@deploy2003: sbassett: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:38 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] * 21:37 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2025 * 21:37 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2025.codfw.wmnet 220.48.192.10.in-addr.arpa 0.2.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:37 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2025.codfw.wmnet 220.48.192.10.in-addr.arpa 0.2.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:37 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:37 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2025 - bking@cumin2003" * 21:37 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2025 - bking@cumin2003" * 21:30 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] (duration: 08m 19s) * 21:26 sbassett@deploy2003: sbassett: Continuing with deployment * 21:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1269.eqiad.wmnet with OS trixie * 21:24 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1271.eqiad.wmnet with reason: host reimage * 21:23 sbassett@deploy2003: sbassett: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:22 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:22 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] * 21:20 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2025 * 21:19 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2025.codfw.wmnet with OS bookworm * 21:17 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1271.eqiad.wmnet with reason: host reimage * 21:04 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1269.eqiad.wmnet with reason: host reimage * 21:00 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1269.eqiad.wmnet with reason: host reimage * 20:56 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1271 * 20:55 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1271 * 20:54 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1271 * 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1271.eqiad.wmnet 126.48.64.10.in-addr.arpa 6.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:54 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1271.eqiad.wmnet 126.48.64.10.in-addr.arpa 6.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1271 - swfrench@cumin1003" * 20:54 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1271 - swfrench@cumin1003" * 20:51 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:51 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:51 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:50 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 20:49 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 20:49 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1268.eqiad.wmnet * 20:49 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1268.eqiad.wmnet * 20:49 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1268.eqiad.wmnet * 20:48 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1271 * 20:48 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1271.eqiad.wmnet with OS trixie * 20:47 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1271.eqiad.wmnet * 20:46 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1271.eqiad.wmnet * 20:46 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1271.eqiad.wmnet * 20:41 aude@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] (duration: 07m 34s) * 20:39 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1269 * 20:39 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1269 * 20:36 aude@deploy2003: aude: Continuing with deployment * 20:35 aude@deploy2003: aude: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:33 aude@deploy2003: Started scap sync-world: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] * 20:26 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on db2207.codfw.wmnet with reason: Host down * 20:22 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-video: apply * 20:21 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-video: apply * 20:20 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-timeline: apply * 20:20 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-timeline: apply * 20:20 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-syntaxhighlight: apply * 20:19 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-syntaxhighlight: apply * 20:19 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-media: apply * 20:18 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-media: apply * 20:18 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-constraints: apply * 20:17 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-constraints: apply * 20:17 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox: apply * 20:16 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox: apply * 20:13 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1269 * 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1269.eqiad.wmnet 80.32.64.10.in-addr.arpa 0.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:13 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1269.eqiad.wmnet 80.32.64.10.in-addr.arpa 0.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1269 - kamila@cumin1003" * 20:13 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1269 - kamila@cumin1003" * 20:09 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-flink-codfw cluster: Roll restart of jvm daemons. * 20:07 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 20:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 20:03 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-flink-codfw cluster: Roll restart of jvm daemons. * 20:03 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2207 [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94893 and previous config saved to /var/cache/conftool/dbconfig/20260716-200257-marostegui.json * 20:01 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2204 to s2 primary [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94892 and previous config saved to /var/cache/conftool/dbconfig/20260716-200157-marostegui.json * 20:00 marostegui: Starting emergency s2 codfw failover from db2207 to db2204 - [[phab:T432396|T432396]] * 19:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1035.eqiad.wmnet * 19:56 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2204 with weight 0 [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94891 and previous config saved to /var/cache/conftool/dbconfig/20260716-195628-marostegui.json * 19:55 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 26 hosts with reason: Primary switchover s2 [[phab:T432396|T432396]] * 19:54 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1035.eqiad.wmnet * 19:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1034.eqiad.wmnet * 19:48 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1034.eqiad.wmnet * 19:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1033.eqiad.wmnet * 19:43 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-video: apply * 19:43 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1033.eqiad.wmnet * 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1032.eqiad.wmnet * 19:42 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-video: apply * 19:42 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-timeline: apply * 19:41 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-timeline: apply * 19:41 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-syntaxhighlight: apply * 19:41 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-syntaxhighlight: apply * 19:40 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-media: apply * 19:40 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-media: apply * 19:39 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-constraints: apply * 19:36 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-constraints: apply * 19:36 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox: apply * 19:35 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1032.eqiad.wmnet * 19:35 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1031.eqiad.wmnet * 19:35 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox: apply * 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-video: apply * 19:33 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-video: apply * 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-timeline: apply * 19:33 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-timeline: apply * 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-syntaxhighlight: apply * 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-syntaxhighlight: apply * 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-media: apply * 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-media: apply * 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-constraints: apply * 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-constraints: apply * 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox: apply * 19:31 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox: apply * 19:27 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1031.eqiad.wmnet * 19:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1030.eqiad.wmnet * 19:23 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1001.eqiad.wmnet, repooling source-only afterwards * 19:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1030.eqiad.wmnet * 19:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1029.eqiad.wmnet * 19:17 kamila@cumin1003: START - Cookbook sre.dns.netbox * 19:12 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1029.eqiad.wmnet * 19:06 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1269 * 19:05 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1269.eqiad.wmnet with OS trixie * 19:03 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1269.eqiad.wmnet * 19:03 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1269.eqiad.wmnet * 19:03 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1269.eqiad.wmnet * 18:55 dancy@deploy2003: Finished scap sync-world: testing [[phab:T428971|T428971]] (duration: 02m 41s) * 18:53 dancy@deploy2003: Started scap sync-world: testing [[phab:T428971|T428971]] * 18:31 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1268.eqiad.wmnet with OS trixie * 18:18 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 18:16 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1267.eqiad.wmnet * 18:16 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1267.eqiad.wmnet * 18:16 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1267.eqiad.wmnet * 18:09 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1268.eqiad.wmnet with reason: host reimage * 18:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1001.eqiad.wmnet, repooling source-only afterwards * 18:06 swfrench@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] (duration: 07m 34s) * 18:06 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 23s) * 18:06 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 18:06 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1268.eqiad.wmnet with reason: host reimage * 18:03 bd808@deploy2003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 18:02 bd808@deploy2003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 18:02 swfrench@deploy2003: jiji, swfrench: Continuing with deployment * 18:02 bd808@deploy2003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 18:02 bd808@deploy2003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 18:01 bd808@deploy2003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 18:01 swfrench@deploy2003: jiji, swfrench: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:01 bd808@deploy2003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:59 swfrench@deploy2003: Started scap sync-world: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] * 17:45 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1268 * 17:45 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1268 * 17:44 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1267.eqiad.wmnet with OS trixie * 17:43 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter2006.codfw.wmnet * 17:39 swfrench@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter2006.codfw.wmnet * 17:35 swfrench@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] (duration: 07m 27s) * 17:34 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1268 * 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1268.eqiad.wmnet 78.32.64.10.in-addr.arpa 8.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:34 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1268.eqiad.wmnet 78.32.64.10.in-addr.arpa 8.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1268 - kamila@cumin1003" * 17:34 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1268 - kamila@cumin1003" * 17:31 swfrench@deploy2003: jiji, swfrench: Continuing with deployment * 17:29 swfrench@deploy2003: jiji, swfrench: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:28 kamila@cumin1003: START - Cookbook sre.dns.netbox * 17:28 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1268 * 17:28 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1268.eqiad.wmnet with OS trixie * 17:27 swfrench@deploy2003: Started scap sync-world: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] * 17:23 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1267.eqiad.wmnet with reason: host reimage * 17:18 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1267.eqiad.wmnet with reason: host reimage * 17:18 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1268.eqiad.wmnet * 17:17 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1268.eqiad.wmnet * 17:17 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1268.eqiad.wmnet * 17:12 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter2005.codfw.wmnet * 17:11 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1270.eqiad.wmnet * 17:11 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1270.eqiad.wmnet * 17:11 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1270.eqiad.wmnet * 17:09 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter2005.codfw.wmnet * 17:08 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:08 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update reverse dns for moved arelion cct cr2-eqiad - cmooney@cumin1003" * 17:08 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update reverse dns for moved arelion cct cr2-eqiad - cmooney@cumin1003" * 17:08 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] (duration: 07m 34s) * 17:04 jiji@deploy2003: jiji: Continuing with deployment * 17:03 jiji@deploy2003: jiji: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 17:00 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] * 17:00 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2185.codfw.wmnet with OS trixie * 16:59 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:58 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1270.eqiad.wmnet with OS trixie * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1267 * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1267 * 16:57 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1267 * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1267.eqiad.wmnet 77.32.64.10.in-addr.arpa 7.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:57 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1267.eqiad.wmnet 77.32.64.10.in-addr.arpa 7.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1267 - kamila@cumin1003" * 16:56 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1267 - kamila@cumin1003" * 16:56 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_eqsin * 16:56 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5032.eqsin.wmnet * 16:52 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_esams * 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3073.esams.wmnet * 16:50 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_esams * 16:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3081.esams.wmnet * 16:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1266.eqiad.wmnet * 16:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1266.eqiad.wmnet * 16:45 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1266.eqiad.wmnet * 16:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2185.codfw.wmnet with reason: host reimage * 16:41 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_eqiad * 16:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1114.eqiad.wmnet * 16:41 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_eqiad * 16:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1115.eqiad.wmnet * 16:39 kamila@cumin1003: START - Cookbook sre.dns.netbox * 16:39 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1267 * 16:39 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2185.codfw.wmnet with reason: host reimage * 16:38 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1267.eqiad.wmnet with OS trixie * 16:38 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1267.eqiad.wmnet * 16:38 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1270.eqiad.wmnet with reason: host reimage * 16:37 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1267.eqiad.wmnet * 16:37 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1267.eqiad.wmnet * 16:31 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1270.eqiad.wmnet with reason: host reimage * 16:24 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1264.eqiad.wmnet * 16:24 kamila@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 16:24 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 16:23 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 16:21 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2162: switch maintenance completed codfw rack b6 * 16:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2185.codfw.wmnet with OS trixie * 16:19 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 16:16 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter1007.eqiad.wmnet * 16:15 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_eqsin * 16:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5024.eqsin.wmnet * 16:13 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5031.eqsin.wmnet * 16:13 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3072.esams.wmnet * 16:12 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter1007.eqiad.wmnet * 16:11 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] (duration: 09m 47s) * 16:10 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1266.eqiad.wmnet with OS trixie * 16:10 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1270 * 16:10 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1270 * 16:09 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1270 * 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1270.eqiad.wmnet 125.48.64.10.in-addr.arpa 5.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1270.eqiad.wmnet 125.48.64.10.in-addr.arpa 5.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1270 - swfrench@cumin1003" * 16:09 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1270 - swfrench@cumin1003" * 16:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3080.esams.wmnet * 16:07 jiji@deploy2003: jiji: Continuing with deployment * 16:06 jiji@deploy2003: jiji: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:04 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 16:04 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1265.eqiad.wmnet * 16:03 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1265.eqiad.wmnet * 16:03 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1265.eqiad.wmnet * 16:03 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1270 * 16:03 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1270.eqiad.wmnet with OS trixie * 16:02 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1270.eqiad.wmnet * 16:02 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] * 16:01 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1270.eqiad.wmnet * 16:01 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1270.eqiad.wmnet * 16:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1113.eqiad.wmnet * 16:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1112.eqiad.wmnet * 15:49 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1266.eqiad.wmnet with reason: host reimage * 15:47 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter1006.eqiad.wmnet * 15:45 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1265.eqiad.wmnet with OS trixie * 15:44 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1266.eqiad.wmnet with reason: host reimage * 15:43 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter1006.eqiad.wmnet * 15:42 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] (duration: 09m 46s) * 15:37 jiji@deploy2003: jiji: Continuing with deployment * 15:36 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2162: switch maintenance completed codfw rack b6 * 15:36 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2161: switch maintenance completed codfw rack b6 * 15:34 jiji@deploy2003: jiji: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:32 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5023.eqsin.wmnet * 15:32 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] * 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5030.eqsin.wmnet * 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3071.esams.wmnet * 15:27 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3079.esams.wmnet * 15:25 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1265.eqiad.wmnet with reason: host reimage * 15:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1266 * 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1266 * 15:21 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1110.eqiad.wmnet * 15:20 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1111.eqiad.wmnet * 15:16 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1265.eqiad.wmnet with reason: host reimage * 15:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1001.eqiad.wmnet with OS bookworm * 15:15 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1266 * 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1266.eqiad.wmnet 76.32.64.10.in-addr.arpa 6.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:15 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1266.eqiad.wmnet 76.32.64.10.in-addr.arpa 6.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1266 - kamila@cumin1003" * 15:15 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1266 - kamila@cumin1003" * 15:07 kamila@cumin1003: START - Cookbook sre.dns.netbox * 15:04 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1266 * 15:04 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1264 * 15:04 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1264 * 15:04 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1266.eqiad.wmnet with OS trixie * 15:03 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1264 * 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1264.eqiad.wmnet 74.32.64.10.in-addr.arpa 4.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:03 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1264.eqiad.wmnet 74.32.64.10.in-addr.arpa 4.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1264 - kamila@cumin1003" * 15:03 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1264 - kamila@cumin1003" * 15:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-eqiad * 15:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp1001.eqiad.wmnet * 15:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp1001.eqiad.wmnet * 15:01 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp1001.eqiad.wmnet * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp1001.eqiad.wmnet * 15:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1376-1384].eqiad.wmnet * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1376-1384].eqiad.wmnet * 14:59 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1002.eqiad.wmnet * 14:59 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1266.eqiad.wmnet * 14:58 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1266.eqiad.wmnet * 14:58 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1266.eqiad.wmnet * 14:58 kamila@cumin1003: START - Cookbook sre.dns.netbox * 14:57 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1264 * 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1265 * 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1265 * 14:57 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1310596{{!}}Set $wgMathInternalRestbaseURL explicitly (T349582)]] * 14:57 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1265 * 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1265.eqiad.wmnet 75.32.64.10.in-addr.arpa 5.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:56 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1265.eqiad.wmnet 75.32.64.10.in-addr.arpa 5.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:56 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:56 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1265 - kamila@cumin1003" * 14:56 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1265 - kamila@cumin1003" * 14:53 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1002.eqiad.wmnet * 14:53 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1376-1384].eqiad.wmnet * 14:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1001.eqiad.wmnet with reason: host reimage * 14:51 kamila@cumin1003: START - Cookbook sre.dns.netbox * 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 14:50 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:50 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2161: switch maintenance completed codfw rack b6 * 14:50 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5021.eqsin.wmnet * 14:50 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1265 * 14:49 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:49 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:49 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1265.eqiad.wmnet with OS trixie * 14:49 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5029.eqsin.wmnet * 14:49 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1265.eqiad.wmnet * 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3070.esams.wmnet * 14:48 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1264.eqiad.wmnet * 14:48 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1001.eqiad.wmnet with reason: host reimage * 14:48 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1376-1384].eqiad.wmnet * 14:48 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1265.eqiad.wmnet * 14:47 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1265.eqiad.wmnet * 14:47 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1264.eqiad.wmnet * 14:47 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1264.eqiad.wmnet * 14:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:47 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3078.esams.wmnet * 14:44 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1263.eqiad.wmnet * 14:44 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1263.eqiad.wmnet * 14:44 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1263.eqiad.wmnet * 14:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1108.eqiad.wmnet * 14:40 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1109.eqiad.wmnet * 14:40 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:35 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:34 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:34 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2006.codfw.wmnet * 14:34 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-flink-eqiad cluster: Roll restart of jvm daemons. * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf2002.codfw.wmnet * 14:31 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1002.eqiad.wmnet * 14:29 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2006.codfw.wmnet * 14:27 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:27 kamila@deploy2003: Finished scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] (duration: 02m 57s) * 14:27 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-flink-eqiad cluster: Roll restart of jvm daemons. * 14:26 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf2002.codfw.wmnet * 14:26 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf2001.codfw.wmnet * 14:25 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1002.eqiad.wmnet * 14:25 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1001.eqiad.wmnet * 14:25 kamila@deploy2003: Started scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] * 14:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:21 kamila@deploy2003: sync-world aborted: Test deployment to check rsync is working - [[phab:T432108|T432108]] (duration: 00m 36s) * 14:21 topranks: reboot lsw1-b6-codfw to upgrade JunOS [[phab:T430922|T430922]] * 14:21 kamila@deploy2003: Started scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] * 14:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1001.eqiad.wmnet with OS bookworm * 14:20 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b6-codfw,lsw1-b6-codfw IPv6,lsw1-b6-codfw.mgmt,ssw1-a[1,8]-codfw with reason: lsw1-b6-codfw JunOS upgrade * 14:20 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf2001.codfw.wmnet * 14:19 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1001.eqiad.wmnet * 14:19 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 26 hosts with reason: lsw1-b6-codfw JunOS upgrade * 14:14 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 14:13 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc2022: switch maintenance codfw rack b6 * 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:12 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.parsercache * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool pc2022: switch maintenance codfw rack b6 * 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2251: switch maintenance codfw rack b6 * 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.parsercache * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2251: switch maintenance codfw rack b6 * 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2162: switch maintenance codfw rack b6 * 14:12 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1263.eqiad.wmnet with OS trixie * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2162: switch maintenance codfw rack b6 * 14:11 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2161: switch maintenance codfw rack b6 * 14:11 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2161: switch maintenance codfw rack b6 * 14:08 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5020.eqsin.wmnet * 14:07 btullis@cumin1003: START - Cookbook sre.hadoop.reboot-workers for Hadoop analytics cluster * 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3069.esams.wmnet * 14:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5028.eqsin.wmnet * 14:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1338-1347].eqiad.wmnet * 14:06 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1338-1347].eqiad.wmnet * 14:05 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3077.esams.wmnet * 14:02 topranks: beginning depools for lsw1-b6-codfw maintenance [[phab:T430922|T430922]] * 14:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1106.eqiad.wmnet * 14:00 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-misc1002.eqiad.wmnet * 13:59 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1338-1347].eqiad.wmnet * 13:59 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1107.eqiad.wmnet * 13:56 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-codfw * 13:55 sfaci@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply * 13:54 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-misc1002.eqiad.wmnet * 13:54 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-misc1001.eqiad.wmnet * 13:54 sfaci@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply * 13:50 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1263.eqiad.wmnet with reason: host reimage * 13:50 sfaci@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 13:49 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1338-1347].eqiad.wmnet * 13:49 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-misc1001.eqiad.wmnet * 13:49 sfaci@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 13:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:49 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:45 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1263.eqiad.wmnet with reason: host reimage * 13:40 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:40 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-eqiad * 13:35 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:34 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:33 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-reboot (exit_code=0) rolling reboot on A:dnsbox and (A:eqsin or A:drmrs or A:magru) and not (P<nowiki>{</nowiki>dns5003*<nowiki>}</nowiki> or P<nowiki>{</nowiki>dns7002*<nowiki>}</nowiki>) and (A:dnsbox) * 13:33 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns7001.wikimedia.org * 13:27 sukhe@dns1004: END - running authdns-update * 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5019.eqsin.wmnet * 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3076.esams.wmnet * 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3068.esams.wmnet * 13:25 sukhe@dns1004: START - running authdns-update * 13:24 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5027.eqsin.wmnet * 13:24 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1263 * 13:24 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1263 * 13:23 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1263 * 13:23 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:23 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1104.eqiad.wmnet * 13:21 kamila@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:21 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:20 kamila@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:20 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:20 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:20 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1263 - kamila@cumin1003" * 13:20 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1263 - kamila@cumin1003" * 13:19 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1105.eqiad.wmnet * 13:19 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:19 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:18 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:18 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns7001.wikimedia.org * 13:16 cdobbins@cumin2003: conftool action : set/pooled=yes; selector: name=dns7002.* * 13:14 cdobbins@dns1004: END - running authdns-update * 13:13 sbisson@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] (duration: 08m 03s) * 13:13 cdobbins@dns1004: START - running authdns-update * 13:12 kamila@cumin1003: START - Cookbook sre.dns.netbox * 13:12 cdobbins@cumin2003: conftool action : set/pooled=yes; selector: name=dns7002.*,service=authdns-update * 13:12 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1263 * 13:11 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1263.eqiad.wmnet with OS trixie * 13:11 cdobbins@cumin2003: conftool action : set/pooled=no; selector: name=dns7002.* * 13:11 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1263.eqiad.wmnet * 13:10 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1263.eqiad.wmnet * 13:10 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1263.eqiad.wmnet * 13:09 sbisson@deploy2003: sbisson: Continuing with deployment * 13:08 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:07 sbisson@deploy2003: sbisson: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:05 sbisson@deploy2003: Started scap sync-world: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] * 13:03 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns6002.wikimedia.org * 13:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1298-1307].eqiad.wmnet * 13:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1298-1307].eqiad.wmnet * 12:59 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1262.eqiad.wmnet * 12:59 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1262.eqiad.wmnet * 12:59 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1262.eqiad.wmnet * 12:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1298-1307].eqiad.wmnet * 12:49 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns6002.wikimedia.org * 12:46 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1298-1307].eqiad.wmnet * 12:46 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:46 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3075.esams.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3067.esams.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5018.eqsin.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5026.eqsin.wmnet * 12:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1102.eqiad.wmnet * 12:39 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1103.eqiad.wmnet * 12:35 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:34 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns6001.wikimedia.org * 12:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:28 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:18 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns6001.wikimedia.org * 12:14 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1267-1276].eqiad.wmnet * 12:13 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1267-1276].eqiad.wmnet * 12:04 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1267-1276].eqiad.wmnet * 12:03 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns5004.wikimedia.org * 12:02 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1100.eqiad.wmnet * 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3066.esams.wmnet * 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3074.esams.wmnet * 12:01 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:01 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-codfw * 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5017.eqsin.wmnet * 12:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5025.eqsin.wmnet * 12:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1101.eqiad.wmnet * 11:59 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1267-1276].eqiad.wmnet * 11:58 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:58 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:54 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns5004.wikimedia.org * 11:54 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and (A:eqsin or A:drmrs or A:magru) and not (P<nowiki>{</nowiki>dns5003*<nowiki>}</nowiki> or P<nowiki>{</nowiki>dns7002*<nowiki>}</nowiki>) and (A:dnsbox) * 11:54 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:53 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:53 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-eqiad * 11:51 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:51 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:50 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:50 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_eqiad * 11:50 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:50 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_eqiad * 11:49 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_esams * 11:49 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_esams * 11:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_eqsin * 11:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_eqsin * 11:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:44 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-eqiad * 11:43 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-codfw * 11:42 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:41 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:24 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-eqiad * 11:23 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-codfw * 11:22 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:15 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:14 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:09 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2066.codfw.wmnet with OS trixie * 11:05 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1151-1160].eqiad.wmnet * 11:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1151-1160].eqiad.wmnet * 10:59 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1068.eqiad.wmnet with OS trixie * 10:55 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.major-upgrade (exit_code=99) * 10:55 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 10:54 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1151-1160].eqiad.wmnet * 10:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2066.codfw.wmnet with reason: host reimage * 10:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1151-1160].eqiad.wmnet * 10:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:42 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2066.codfw.wmnet with reason: host reimage * 10:39 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:37 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:36 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:23 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:22 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2066.codfw.wmnet with OS trixie * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:07 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2065.codfw.wmnet with OS trixie * 10:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:06 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 10:06 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 10:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 10:03 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:03 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 10:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 09:59 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 09:57 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 09:57 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 09:52 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 09:47 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:46 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:46 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2065.codfw.wmnet with reason: host reimage * 09:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2065.codfw.wmnet with reason: host reimage * 09:40 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:39 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 09:39 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:39 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:39 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:37 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:29 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox2003.codfw.wmnet * 09:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox2003.codfw.wmnet * 09:25 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:25 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:24 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:24 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:24 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:21 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:20 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2065.codfw.wmnet with OS trixie * 09:13 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1068.eqiad.wmnet with OS trixie * 09:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2064.codfw.wmnet with OS trixie * 09:08 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 09:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 09:07 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2162: Repooling after switchover * 09:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1067.eqiad.wmnet with OS trixie * 09:00 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:59 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 08:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 08:57 tappof: bump space for prometheus k8s-dse in eqiad * 08:56 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping2004.codfw.wmnet * 08:52 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host ping2004.codfw.wmnet * 08:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 08:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping1004.eqiad.wmnet * 08:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:51 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:49 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 08:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host ping1004.eqiad.wmnet * 08:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2064.codfw.wmnet with reason: host reimage * 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2064.codfw.wmnet with reason: host reimage * 08:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1067.eqiad.wmnet with reason: host reimage * 08:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:33 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1067.eqiad.wmnet with reason: host reimage * 08:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:21 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2162: Repooling after switchover * 08:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2064.codfw.wmnet with OS trixie * 08:16 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1067.eqiad.wmnet with OS trixie * 08:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:15 cgoubert@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-eqiad * 08:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2062.codfw.wmnet with OS trixie * 08:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1066.eqiad.wmnet with OS trixie * 08:02 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2162: Repooling after switchover * 07:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2162: Repooling after switchover * 07:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2162 [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94870 and previous config saved to /var/cache/conftool/dbconfig/20260716-075530-cwilliams.json * 07:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2241 to x3 primary [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94869 and previous config saved to /var/cache/conftool/dbconfig/20260716-075314-cwilliams.json * 07:52 cezmunsta: Starting x3 codfw failover from db2162 to db2241 - [[phab:T430925|T430925]] * 07:50 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:50 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:47 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 07:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2241 with weight 0 [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94868 and previous config saved to /var/cache/conftool/dbconfig/20260716-074507-cwilliams.json * 07:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 18 hosts with reason: Primary switchover x3 [[phab:T430925|T430925]] * 07:43 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1066.eqiad.wmnet with reason: host reimage * 07:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:dse-k8s-worker-eqiad * 07:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1028.eqiad.wmnet * 07:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1028.eqiad.wmnet * 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 07:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1066.eqiad.wmnet with reason: host reimage * 07:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1028.eqiad.wmnet * 07:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1028.eqiad.wmnet * 07:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1027.eqiad.wmnet * 07:35 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1027.eqiad.wmnet * 07:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1027.eqiad.wmnet * 07:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1027.eqiad.wmnet * 07:28 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1026.eqiad.wmnet * 07:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1026.eqiad.wmnet * 07:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast2003.wikimedia.org * 07:21 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1026.eqiad.wmnet * 07:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1066.eqiad.wmnet with OS trixie * 07:19 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast2003.wikimedia.org * 07:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2062.codfw.wmnet with OS trixie * 06:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1026.eqiad.wmnet * 06:51 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1025.eqiad.wmnet * 06:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1025.eqiad.wmnet * 06:47 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 06:47 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 06:44 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1025.eqiad.wmnet * 06:14 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1025.eqiad.wmnet * 06:14 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1024.eqiad.wmnet * 06:14 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1024.eqiad.wmnet * 06:07 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1024.eqiad.wmnet * 05:37 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1024.eqiad.wmnet * 05:37 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1023.eqiad.wmnet * 05:37 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1023.eqiad.wmnet * 05:26 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1023.eqiad.wmnet * 04:56 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1023.eqiad.wmnet * 04:56 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1022.eqiad.wmnet * 04:56 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1022.eqiad.wmnet * 04:49 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1022.eqiad.wmnet * 04:19 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1022.eqiad.wmnet * 04:19 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1021.eqiad.wmnet * 04:19 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1021.eqiad.wmnet * 04:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1021.eqiad.wmnet * 03:38 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1021.eqiad.wmnet * 03:38 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1020.eqiad.wmnet * 03:38 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1020.eqiad.wmnet * 03:20 btullis@cumin1003: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1020.eqiad.wmnet * 03:18 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1020.eqiad.wmnet * 03:18 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1019.eqiad.wmnet * 03:18 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1019.eqiad.wmnet * 03:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1019.eqiad.wmnet * 02:41 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1019.eqiad.wmnet * 02:41 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1018.eqiad.wmnet * 02:41 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1018.eqiad.wmnet * 02:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling both afterwards * 02:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2003.codfw.wmnet -> wcqs2001.codfw.wmnet, repooling both afterwards * 02:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1018.eqiad.wmnet * 02:30 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1018.eqiad.wmnet * 02:30 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1014.eqiad.wmnet * 02:30 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1014.eqiad.wmnet * 02:24 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1014.eqiad.wmnet * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 01:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1014.eqiad.wmnet * 01:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1013.eqiad.wmnet * 01:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1013.eqiad.wmnet * 01:47 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1013.eqiad.wmnet * 01:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2003.codfw.wmnet -> wcqs2001.codfw.wmnet, repooling both afterwards * 01:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling both afterwards * 01:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1013.eqiad.wmnet * 01:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1012.eqiad.wmnet * 01:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1012.eqiad.wmnet * 01:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1012.eqiad.wmnet * 01:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1012.eqiad.wmnet * 01:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1011.eqiad.wmnet * 01:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1011.eqiad.wmnet * 01:08 ryankemper@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] scap deploy post bookworm reimage (duration: 00m 23s) * 01:08 ryankemper@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] scap deploy post bookworm reimage * 01:08 ryankemper@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): scap deploy post bookworm reimage (duration: 00m 46s) * 01:07 ryankemper@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): scap deploy post bookworm reimage * 01:04 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1011.eqiad.wmnet * 01:04 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1011.eqiad.wmnet * 01:04 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1010.eqiad.wmnet * 01:04 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1010.eqiad.wmnet * 00:57 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1010.eqiad.wmnet * 00:57 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1010.eqiad.wmnet * 00:57 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1009.eqiad.wmnet * 00:57 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1009.eqiad.wmnet * 00:50 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1009.eqiad.wmnet * 00:20 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1009.eqiad.wmnet * 00:20 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1008.eqiad.wmnet * 00:20 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1008.eqiad.wmnet * 00:13 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1008.eqiad.wmnet == 2026-07-15 == * 23:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2001.codfw.wmnet with OS bookworm * 23:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1008.eqiad.wmnet * 23:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1007.eqiad.wmnet * 23:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1007.eqiad.wmnet * 23:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1007.eqiad.wmnet * 23:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1007.eqiad.wmnet * 23:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1006.eqiad.wmnet * 23:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1006.eqiad.wmnet * 23:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1006.eqiad.wmnet * 23:29 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1006.eqiad.wmnet * 23:28 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1005.eqiad.wmnet * 23:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1005.eqiad.wmnet * 23:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1002.eqiad.wmnet with OS bookworm * 23:21 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1005.eqiad.wmnet * 23:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2001.codfw.wmnet with reason: host reimage * 23:15 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host datahubsearch1001.eqiad.wmnet with OS bookworm * 23:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2001.codfw.wmnet with reason: host reimage * 23:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 23:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 22:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 22:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1005.eqiad.wmnet * 22:51 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1004.eqiad.wmnet * 22:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1004.eqiad.wmnet * 22:45 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1004.eqiad.wmnet * 22:44 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 22:44 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS trixie * 22:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host datahubsearch1001.eqiad.wmnet with OS bookworm * 22:34 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host datahubsearch1001.eqiad.wmnet with OS bookworm * 22:16 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on datahubsearch[1002-1003].eqiad.wmnet with reason: Using datahubsearch1001 to test bookworm reimages * 22:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1004.eqiad.wmnet * 22:15 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1003.eqiad.wmnet * 22:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1003.eqiad.wmnet * 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 22:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1003.eqiad.wmnet * 22:08 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1003.eqiad.wmnet * 22:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1002.eqiad.wmnet * 22:08 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1002.eqiad.wmnet * 22:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host datahubsearch1001.eqiad.wmnet with OS bookworm * 22:05 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 22:02 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm * 22:01 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on datahubsearch[1001-1003].eqiad.wmnet with reason: Using datahubsearch1001 to test bookworm reimages * 22:01 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1002.eqiad.wmnet * 22:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1002.eqiad.wmnet * 22:00 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1001.eqiad.wmnet * 22:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1001.eqiad.wmnet * 21:53 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1001.eqiad.wmnet * 21:52 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 21:50 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 21:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS trixie * 21:50 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS bookworm * 21:43 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 21:38 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:30 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wcqs1002'] * 21:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:29 lerickson@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 21:29 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:29 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:29 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS bookworm * 21:28 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 21:28 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm * 21:23 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1001.eqiad.wmnet * 21:23 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:23 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:22 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 21:20 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:18 swfrench-wmf: reprepro include php8.3_8.3.32-1+wmf11u2 into component/php83 for bullseye-wikimedia * 21:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:16 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:15 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-druid-public cluster: Roll restart of jvm daemons. * 21:08 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:05 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1001.eqiad.wmnet * 21:05 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1001.eqiad.wmnet * 21:04 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-druid-public cluster: Roll restart of jvm daemons. * 21:02 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 21:01 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 21:01 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 21:00 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 20:59 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1001.eqiad.wmnet * 20:59 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1001.eqiad.wmnet * 20:59 btullis@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:dse-k8s-worker-eqiad * 20:55 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 20:55 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 20:45 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm * 20:21 jhathaway: puppet is re-enabled, have fun, but not too much fun! * 20:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 20:17 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs2001'] * 20:12 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs2001'] * 20:11 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs2001'] * 20:09 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:08 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 20:05 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:05 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 20:04 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs2001'] * 20:03 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 20:03 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm * 20:02 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:02 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 20:01 jhathaway: disabling puppet fleet wide to roll out kafka patch * 19:55 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host relforge1010.eqiad.wmnet * 19:52 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 19:52 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 19:48 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 19:48 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 19:48 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 19:47 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 19:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 19:45 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 19:45 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 19:44 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1010.eqiad.wmnet * 19:38 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1262.eqiad.wmnet with OS trixie * 19:17 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1262.eqiad.wmnet with reason: host reimage * 19:11 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1262.eqiad.wmnet with reason: host reimage * 18:59 cdobbins@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS trixie * 18:54 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 18:53 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 18:52 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1262 * 18:52 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1262 * 18:51 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1262 * 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1262.eqiad.wmnet 72.32.64.10.in-addr.arpa 2.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:51 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1262.eqiad.wmnet 72.32.64.10.in-addr.arpa 2.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1262 - kamila@cumin1003" * 18:51 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1262 - kamila@cumin1003" * 18:46 kamila@cumin1003: START - Cookbook sre.dns.netbox * 18:46 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1262 * 18:46 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 18:46 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ncmonitor1001.eqiad.wmnet * 18:46 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 18:45 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1262.eqiad.wmnet with OS trixie * 18:45 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 18:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1262.eqiad.wmnet * 18:44 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1262.eqiad.wmnet * 18:44 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1262.eqiad.wmnet * 18:42 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host ncmonitor1001.eqiad.wmnet * 18:29 topranks: pull power on cr1-eqiad to install new switch-control boards [[phab:T426343|T426343]] * 18:29 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs[1018-1020].eqiad.wmnet with reason: line card install in cr1-eqiad * 18:27 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 14 hosts with reason: linecard install in cr1-eqad * 18:22 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_ulsfo * 18:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4052.ulsfo.wmnet * 18:19 cdobbins@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 18:15 cdobbins@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 18:14 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_drmrs * 18:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6016.drmrs.wmnet * 18:12 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_ulsfo * 18:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4044.ulsfo.wmnet * 18:10 sukhe@cumin1003: END (ERROR) - Cookbook sre.cdn.roll-reboot (exit_code=97) rolling reboot on A:cp-upload_drmrs * 18:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2241: Security update * 17:56 topranks: start draining traffic on cr1-eqiad ahead of line card installation [[phab:T426343|T426343]] * 17:47 cdobbins@cumin2003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie * 17:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4051.ulsfo.wmnet * 17:40 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:39 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 17:34 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6007.drmrs.wmnet * 17:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6015.drmrs.wmnet * 17:32 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:31 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 17:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4043.ulsfo.wmnet * 17:27 lerickson@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:25 lerickson@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 17:22 lerickson@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-codfw * 17:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp2001.codfw.wmnet * 17:22 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 17:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp2001.codfw.wmnet * 17:22 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 17:19 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2241: Security update * 17:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2241.codfw.wmnet * 17:17 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2241.codfw.wmnet * 17:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp2001.codfw.wmnet * 17:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp2001.codfw.wmnet * 17:15 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2366-2374].codfw.wmnet * 17:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2366-2374].codfw.wmnet * 17:10 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply * 17:10 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply * 17:08 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2366-2374].codfw.wmnet * 17:06 sukhe: sre.dns.roll-reboot to resume later * 17:06 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-reboot (exit_code=97) rolling reboot on A:dnsbox and not (A:ulsfo or A:magru) and (A:dnsbox) * 17:06 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns5003.wikimedia.org * 17:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2241: Security update * 17:03 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2241: Security update * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply * 17:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2366-2374].codfw.wmnet * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply * 17:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2357-2365].codfw.wmnet * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 17:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2357-2365].codfw.wmnet * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply * 16:55 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2357-2365].codfw.wmnet * 16:55 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 16:53 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6006.drmrs.wmnet * 16:52 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6014.drmrs.wmnet * 16:52 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 16:51 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 16:50 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2357-2365].codfw.wmnet * 16:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4042.ulsfo.wmnet * 16:50 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2347-2356].codfw.wmnet * 16:50 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2347-2356].codfw.wmnet * 16:49 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns5003.wikimedia.org * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply * 16:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4050.ulsfo.wmnet * 16:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2347-2356].codfw.wmnet * 16:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2347-2356].codfw.wmnet * 16:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2337-2346].codfw.wmnet * 16:36 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2337-2346].codfw.wmnet * 16:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:dse-k8s-worker-codfw * 16:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2003.codfw.wmnet * 16:35 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2003.codfw.wmnet * 16:34 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns3004.wikimedia.org * 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply * 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply * 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply * 16:30 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply * 16:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2003.codfw.wmnet * 16:29 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2337-2346].codfw.wmnet * 16:24 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2003.codfw.wmnet * 16:24 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2002.codfw.wmnet * 16:24 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2002.codfw.wmnet * 16:23 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns3004.wikimedia.org * 16:23 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2337-2346].codfw.wmnet * 16:23 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2327-2336].codfw.wmnet * 16:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2327-2336].codfw.wmnet * 16:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2002.codfw.wmnet * 16:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2327-2336].codfw.wmnet * 16:12 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2002.codfw.wmnet * 16:12 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2001.codfw.wmnet * 16:12 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2001.codfw.wmnet * 16:12 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1065.eqiad.wmnet with OS trixie * 16:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6005.drmrs.wmnet * 16:11 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6013.drmrs.wmnet * 16:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4041.ulsfo.wmnet * 16:08 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns3003.wikimedia.org * 16:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2327-2336].codfw.wmnet * 16:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2317-2326].codfw.wmnet * 16:06 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2317-2326].codfw.wmnet * 16:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2001.codfw.wmnet * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply * 16:03 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4049.ulsfo.wmnet * 16:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2001.codfw.wmnet * 16:00 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test2001.codfw.wmnet * 16:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test2001.codfw.wmnet * 16:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2063.codfw.wmnet with OS trixie * 15:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2317-2326].codfw.wmnet * 15:57 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns3003.wikimedia.org * 15:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test2001.codfw.wmnet * 15:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test2001.codfw.wmnet * 15:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2004.codfw.wmnet * 15:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2004.codfw.wmnet * 15:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2317-2326].codfw.wmnet * 15:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2307-2316].codfw.wmnet * 15:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2307-2316].codfw.wmnet * 15:49 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2004.codfw.wmnet * 15:48 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2004.codfw.wmnet * 15:48 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2003.codfw.wmnet * 15:48 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2003.codfw.wmnet * 15:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 15:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2307-2316].codfw.wmnet * 15:42 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2003.codfw.wmnet * 15:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 15:42 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2003.codfw.wmnet * 15:42 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2002.codfw.wmnet * 15:42 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2002.codfw.wmnet * 15:42 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2006.wikimedia.org * 15:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2063.codfw.wmnet with reason: host reimage * 15:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2307-2316].codfw.wmnet * 15:37 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2297-2306].codfw.wmnet * 15:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2297-2306].codfw.wmnet * 15:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2002.codfw.wmnet * 15:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2002.codfw.wmnet * 15:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2001.codfw.wmnet * 15:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2001.codfw.wmnet * 15:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2063.codfw.wmnet with reason: host reimage * 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6004.drmrs.wmnet * 15:31 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2001.codfw.wmnet * 15:31 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2001.codfw.wmnet * 15:31 btullis@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:dse-k8s-worker-codfw * 15:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6012.drmrs.wmnet * 15:28 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2006.wikimedia.org * 15:27 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-analytics cluster: Roll restart of jvm daemons. * 15:27 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2297-2306].codfw.wmnet * 15:27 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4040.ulsfo.wmnet * 15:24 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1065.eqiad.wmnet with OS trixie * 15:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4048.ulsfo.wmnet * 15:21 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-analytics cluster: Roll restart of jvm daemons. * 15:21 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2297-2306].codfw.wmnet * 15:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2287-2296].codfw.wmnet * 15:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2287-2296].codfw.wmnet * 15:20 btullis@cumin1003: END (PASS) - Cookbook sre.druid.reboot-workers (exit_code=0) for Druid public cluster: Reboot Druid nodes * 15:18 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 15:17 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm * 15:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2063.codfw.wmnet with OS trixie * 15:13 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2005.wikimedia.org * 15:11 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2287-2296].codfw.wmnet * 15:11 btullis@cumin1003: END (PASS) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=0) rolling reboot on A:cephosd-eqiad * 15:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1064.eqiad.wmnet with OS trixie * 15:05 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2062.codfw.wmnet with OS trixie * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2287-2296].codfw.wmnet * 15:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2277-2286].codfw.wmnet * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2277-2286].codfw.wmnet * 14:59 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2005.wikimedia.org * 14:57 brouberol@cumin1003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-jumbo-eqiad * 14:52 btullis@cumin1003: END (PASS) - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas (exit_code=0) rolling reboot on A:schema-codfw * 14:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6003.drmrs.wmnet * 14:50 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:50 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host relforge1009.eqiad.wmnet * 14:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2277-2286].codfw.wmnet * 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6011.drmrs.wmnet * 14:47 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 14:46 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>ml-serve1001.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 14:46 klausman@cumin1003: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) pool for host ml-serve1001.eqiad.wmnet * 14:46 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 14:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1001.eqiad.wmnet * 14:45 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4039.ulsfo.wmnet * 14:44 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1009.eqiad.wmnet * 14:44 btullis@cumin1003: START - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas rolling reboot on A:schema-codfw * 14:44 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2004.wikimedia.org * 14:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2277-2286].codfw.wmnet * 14:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2267-2276].codfw.wmnet * 14:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2267-2276].codfw.wmnet * 14:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 14:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4047.ulsfo.wmnet * 14:40 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1001.eqiad.wmnet * 14:38 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 14:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 14:36 topranks: disconnect power on cr2-eqiad to shut down device for switch fabric replacement [[phab:T426343|T426343]] * 14:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2267-2276].codfw.wmnet * 14:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 14:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1001.eqiad.wmnet * 14:35 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>ml-serve1001.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 14:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 14:34 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 14:33 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:33 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:30 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2004.wikimedia.org * 14:29 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2267-2276].codfw.wmnet * 14:29 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2257-2266].codfw.wmnet * 14:29 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2257-2266].codfw.wmnet * 14:24 btullis@cumin1003: END (PASS) - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas (exit_code=0) rolling reboot on A:schema-eqiad * 14:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2257-2266].codfw.wmnet * 14:20 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:20 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:19 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:17 jforrester@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2257-2266].codfw.wmnet * 14:16 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:16 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2062.codfw.wmnet with OS trixie * 14:15 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1064.eqiad.wmnet with OS trixie * 14:15 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1006.wikimedia.org * 14:15 btullis@cumin1003: START - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas rolling reboot on A:schema-eqiad * 14:14 topranks: switch routing-engine on cr2-eqiad resetting all interfaces [[phab:T417873|T417873]] * 14:11 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:11 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:10 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm * 14:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6002.drmrs.wmnet * 14:09 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6010.drmrs.wmnet * 14:06 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1006.wikimedia.org * 14:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:05 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 14:05 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4038.ulsfo.wmnet * 14:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4046.ulsfo.wmnet * 14:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:00 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on cr1-eqiad with reason: switch upgrade and line card install * 14:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:59 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:57 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:57 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:55 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-eqiad * 13:55 btullis@cumin1003: START - Cookbook sre.druid.reboot-workers for Druid public cluster: Reboot Druid nodes * 13:53 topranks: switch routing-engine on cr2-eqiad resetting all interfaces [[phab:T417873|T417873]] * 13:51 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1005.wikimedia.org * 13:50 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:49 brouberol@cumin1003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-test-eqiad * 13:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:44 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2197-2206].codfw.wmnet * 13:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2197-2206].codfw.wmnet * 13:36 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1005.wikimedia.org * 13:35 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2197-2206].codfw.wmnet * 13:30 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2197-2206].codfw.wmnet * 13:28 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6001.drmrs.wmnet * 13:28 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6009.drmrs.wmnet * 13:28 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2187-2196].codfw.wmnet * 13:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2187-2196].codfw.wmnet * 13:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 13:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4037.ulsfo.wmnet * 13:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2001 * 13:22 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2001 * 13:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4045.ulsfo.wmnet * 13:21 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1004.wikimedia.org * 13:19 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on lvs[1018-1020].eqiad.wmnet with reason: switch upgrade and line card install * 13:18 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2009.codfw.wmnet * 13:18 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2001 * 13:18 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2001.codfw.wmnet 26.16.192.10.in-addr.arpa 6.2.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:17 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2001.codfw.wmnet 26.16.192.10.in-addr.arpa 6.2.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:17 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:17 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2001 - bking@cumin2003" * 13:17 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2001 - bking@cumin2003" * 13:17 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2009.codfw.wmnet * 13:17 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_drmrs * 13:17 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2187-2196].codfw.wmnet * 13:17 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_drmrs * 13:17 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on 15 hosts with reason: switch upgrade and line card install * 13:17 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:15 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 13:13 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:13 brouberol@cumin1003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-jumbo-eqiad * 13:13 brouberol@cumin1003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-test-eqiad * 13:13 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1004.wikimedia.org * 13:13 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and not (A:ulsfo or A:magru) and (A:dnsbox) * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:12 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_ulsfo * 13:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:12 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_ulsfo * 13:11 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2187-2196].codfw.wmnet * 13:11 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 13:11 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 13:06 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 13:05 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling source-only afterwards * 13:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2001 * 13:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2009.codfw.wmnet with OS trixie * 13:03 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:03 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling source-only afterwards * 13:01 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 15s) * 13:01 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 13:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 12:57 btullis@cumin1003: END (PASS) - Cookbook sre.druid.reboot-workers (exit_code=0) for Druid analytics cluster: Reboot Druid nodes * 12:54 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 12:54 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2163-2172].codfw.wmnet * 12:54 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2163-2172].codfw.wmnet * 12:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2163-2172].codfw.wmnet * 12:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2009.codfw.wmnet with reason: host reimage * 12:41 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2163-2172].codfw.wmnet * 12:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2153-2162].codfw.wmnet * 12:40 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2153-2162].codfw.wmnet * 12:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2009.codfw.wmnet with reason: host reimage * 12:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2153-2162].codfw.wmnet * 12:29 btullis@cumin1003: END (PASS) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=0) rolling reboot on A:cephosd-codfw * 12:25 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2153-2162].codfw.wmnet * 12:25 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2143-2152].codfw.wmnet * 12:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2143-2152].codfw.wmnet * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2009 * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2009 * 12:22 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2009 * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2009.codfw.wmnet 139.0.192.10.in-addr.arpa 9.3.1.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:22 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2009.codfw.wmnet 139.0.192.10.in-addr.arpa 9.3.1.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2009 - mvernon@cumin2003" * 12:22 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2009 - mvernon@cumin2003" * 12:16 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 12:15 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 12:15 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 12:15 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2009 * 12:15 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 12:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2009.codfw.wmnet with OS trixie * 12:15 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 12:14 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 12:14 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2143-2152].codfw.wmnet * 12:13 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 12:12 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2010.codfw.wmnet * 12:11 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2010.codfw.wmnet * 12:10 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 12:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2143-2152].codfw.wmnet * 12:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2133-2142].codfw.wmnet * 12:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2133-2142].codfw.wmnet * 12:02 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 11:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2133-2142].codfw.wmnet * 11:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2133-2142].codfw.wmnet * 11:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:49 mvolz@deploy2003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:49 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-codfw * 11:48 mvolz@deploy2003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:47 btullis@cumin1003: START - Cookbook sre.druid.reboot-workers for Druid analytics cluster: Reboot Druid nodes * 11:46 mvolz@deploy2003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:46 mvolz@deploy2003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:45 mvolz@deploy2003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:44 mvolz@deploy2003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:40 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] (duration: 11m 38s) * 11:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2010.codfw.wmnet with OS trixie * 11:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2105-2114].codfw.wmnet * 11:36 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2105-2114].codfw.wmnet * 11:36 krinkle@deploy2003: physikerwelt, krinkle: Continuing with deployment * 11:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1018: Security updates * 11:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:36 root@cumin1003: START - Cookbook sre.mysql.parsercache * 11:36 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1018: Security updates * 11:31 krinkle@deploy2003: physikerwelt, krinkle: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:29 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] * 11:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2105-2114].codfw.wmnet * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2105-2114].codfw.wmnet * 11:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2010.codfw.wmnet with reason: host reimage * 11:12 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2010.codfw.wmnet with reason: host reimage * 11:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1018: Security updates * 11:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:10 root@cumin1003: START - Cookbook sre.mysql.parsercache * 11:10 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1018: Security updates * 11:09 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1009.eqiad.wmnet with OS trixie * 11:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow7002.magru.wmnet * 11:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 11:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 11:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=tegola-vector-tiles,name=eqiad * 11:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=kartotherian,name=eqiad * 11:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow7002.magru.wmnet * 10:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2010 * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2010 * 10:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 10:54 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 10:54 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2010 * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2010.codfw.wmnet 76.16.192.10.in-addr.arpa 6.7.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:54 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2010.codfw.wmnet 76.16.192.10.in-addr.arpa 6.7.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2010 - mvernon@cumin2003" * 10:54 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2010 - mvernon@cumin2003" * 10:49 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 10:49 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2010 * 10:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1009.eqiad.wmnet with reason: host reimage * 10:49 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2010.codfw.wmnet with OS trixie * 10:46 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2011.codfw.wmnet * 10:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1011.eqiad.wmnet * 10:44 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2011.codfw.wmnet * 10:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 10:44 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 10:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1009.eqiad.wmnet with reason: host reimage * 10:44 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow6001.drmrs.wmnet * 10:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1017: Security updates * 10:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:39 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1017: Security updates * 10:39 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow6001.drmrs.wmnet * 10:38 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1011.eqiad.wmnet * 10:35 cgoubert@deploy2003: Finished deploy [restbase/deploy@06301bd]: Deploying {{Gerrit|1306088}} {{Gerrit|1308347}} - [[phab:T429944|T429944]] [[phab:T428279|T428279]] (duration: 28m 34s) * 10:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1012.eqiad.wmnet * 10:35 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 10:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow5003.eqsin.wmnet * 10:34 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2011.codfw.wmnet with OS trixie * 10:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1009.eqiad.wmnet with OS trixie * 10:28 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1012.eqiad.wmnet * 10:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1013.eqiad.wmnet * 10:27 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:27 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow5003.eqsin.wmnet * 10:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:26 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow4003.ulsfo.wmnet * 10:25 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 10:25 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 10:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow4003.ulsfo.wmnet * 10:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1013.eqiad.wmnet * 10:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1014.eqiad.wmnet * 10:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:15 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2011.codfw.wmnet with reason: host reimage * 10:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1017: Security updates * 10:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:14 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:14 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1017: Security updates * 10:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow3004.esams.wmnet * 10:11 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2011.codfw.wmnet with reason: host reimage * 10:10 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:10 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1014.eqiad.wmnet * 10:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki2003.codfw.wmnet * 10:09 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 10:09 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 10:09 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow3004.esams.wmnet * 10:08 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2004.codfw.wmnet * 10:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1010.eqiad.wmnet with OS trixie * 10:07 cgoubert@deploy2003: Started deploy [restbase/deploy@06301bd]: Deploying {{Gerrit|1306088}} {{Gerrit|1308347}} - [[phab:T429944|T429944]] [[phab:T428279|T428279]] * 10:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host rpki2003.codfw.wmnet * 10:04 topranks: push out config change to BGP_outfilter on core routers [[phab:T431849|T431849]] * 10:02 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow2004.codfw.wmnet * 09:59 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 09:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2003.codfw.wmnet * 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2011 * 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2011 * 09:53 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 09:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:52 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2011 * 09:52 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2011.codfw.wmnet 36.32.192.10.in-addr.arpa 6.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:52 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2011.codfw.wmnet 36.32.192.10.in-addr.arpa 6.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:51 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:51 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2011 - mvernon@cumin2003" * 09:51 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2011 - mvernon@cumin2003" * 09:51 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow2003.codfw.wmnet * 09:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1003.eqiad.wmnet * 09:49 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 09:49 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 09:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1010.eqiad.wmnet with reason: host reimage * 09:47 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 09:47 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 09:47 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 09:47 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2011 * 09:46 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2011.codfw.wmnet with OS trixie * 09:44 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow1003.eqiad.wmnet * 09:44 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2012.codfw.wmnet * 09:44 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1002.eqiad.wmnet * 09:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1010.eqiad.wmnet with reason: host reimage * 09:43 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2012.codfw.wmnet * 09:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Security updates * 09:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:43 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:43 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Security updates * 09:42 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:40 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow1002.eqiad.wmnet * 09:40 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 09:37 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki1001.eqiad.wmnet * 09:36 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 09:36 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 09:33 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host rpki1001.eqiad.wmnet * 09:32 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:32 cgoubert@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-codfw * 09:31 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=kartotherian,name=eqiad * 09:31 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola-vector-tiles,name=eqiad * 09:31 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 09:31 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2012.codfw.wmnet with OS trixie * 09:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1010.eqiad.wmnet with OS trixie * 09:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1011.eqiad.wmnet with OS trixie * 09:21 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Security updates * 09:21 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:21 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:21 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Security updates * 09:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2012.codfw.wmnet with reason: host reimage * 09:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1011.eqiad.wmnet with reason: host reimage * 09:08 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2012.codfw.wmnet with reason: host reimage * 09:05 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1011.eqiad.wmnet with reason: host reimage * 08:55 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:52 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1011.eqiad.wmnet with OS trixie * 08:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1022: Security updates * 08:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2012 * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2012 * 08:50 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1022: Security updates * 08:50 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2012 * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2012.codfw.wmnet 44.48.192.10.in-addr.arpa 4.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:50 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2012.codfw.wmnet 44.48.192.10.in-addr.arpa 4.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2012 - mvernon@cumin2003" * 08:50 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2012 - mvernon@cumin2003" * 08:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1012.eqiad.wmnet with OS trixie * 08:44 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 08:44 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2012 * 08:43 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2012.codfw.wmnet with OS trixie * 08:42 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2013.codfw.wmnet * 08:41 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2013.codfw.wmnet * 08:35 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 08:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host krb1002.eqiad.wmnet * 08:30 elukey@dns1004: END - running authdns-update * 08:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1012.eqiad.wmnet with reason: host reimage * 08:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Security updates * 08:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:28 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:28 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Security updates * 08:27 elukey@dns1004: START - running authdns-update * 08:26 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 08:26 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host krb1002.eqiad.wmnet * 08:22 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1012.eqiad.wmnet with reason: host reimage * 08:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host krb2002.codfw.wmnet * 08:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast6003.wikimedia.org * 08:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2013.codfw.wmnet with OS trixie * 08:13 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast6003.wikimedia.org * 08:12 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast3007.wikimedia.org * 08:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host krb2002.codfw.wmnet * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Security updates * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:09 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:09 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Security updates * 08:07 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1012.eqiad.wmnet with OS trixie * 08:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast3007.wikimedia.org * 08:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast5005.wikimedia.org * 07:58 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast5005.wikimedia.org * 07:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1013.eqiad.wmnet with OS trixie * 07:53 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2013.codfw.wmnet with reason: host reimage * 07:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1021: Security updates * 07:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:53 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:53 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1021: Security updates * 07:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast1004.wikimedia.org * 07:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2013.codfw.wmnet with reason: host reimage * 07:46 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast1004.wikimedia.org * 07:40 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1013.eqiad.wmnet with reason: host reimage * 07:36 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1013.eqiad.wmnet with reason: host reimage * 07:31 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2013 * 07:31 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2013 * 07:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1021: Security updates * 07:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:30 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:30 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1021: Security updates * 07:24 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2013 * 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2013.codfw.wmnet 87.0.192.10.in-addr.arpa 7.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:24 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2013.codfw.wmnet 87.0.192.10.in-addr.arpa 7.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2013 - mvernon@cumin2003" * 07:24 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2013 - mvernon@cumin2003" * 07:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1013.eqiad.wmnet with OS trixie * 07:19 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 07:19 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2013 * 07:19 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2013.codfw.wmnet with OS trixie * 07:13 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] (duration: 07m 48s) * 07:09 kharlan@deploy2003: kharlan: Continuing with deployment * 07:08 kharlan@deploy2003: kharlan: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:06 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 01:15 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 01:14 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply == 2026-07-14 == * 22:51 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_magru * 22:51 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7016.magru.wmnet * 22:46 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_magru * 22:46 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7008.magru.wmnet * 22:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7015.magru.wmnet * 22:04 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7007.magru.wmnet * 21:29 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7014.magru.wmnet * 21:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7006.magru.wmnet * 21:13 dzahn@cumin2002: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 0:15:00 on gerrit.wikimedia.org with reason: reboot * 21:11 mutante: gerrit2003 (gerrit.wikimedia.org) - reboot for maintenance * 21:11 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on gerrit2003.wikimedia.org with reason: reboot * 20:56 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:56 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:56 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:55 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 20:48 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7013.magru.wmnet * 20:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7005.magru.wmnet * 20:41 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host phab1005.eqiad.wmnet with OS trixie * 20:28 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] (duration: 06m 47s) * 20:24 sbassett@deploy2003: sbassett: Continuing with deployment * 20:23 sbassett@deploy2003: sbassett: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:23 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on phab1005.eqiad.wmnet with reason: host reimage * 20:21 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] * 20:20 aokoth@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on phab1005.eqiad.wmnet with reason: host reimage * 20:12 jhuneidi@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] (duration: 07m 42s) * 20:07 jhuneidi@deploy2003: jhuneidi, priyankar22: Continuing with deployment * 20:06 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7012.magru.wmnet * 20:06 jhuneidi@deploy2003: jhuneidi, priyankar22: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:04 jhuneidi@deploy2003: Started scap sync-world: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] * 20:02 aokoth@cumin1003: START - Cookbook sre.hosts.reimage for host phab1005.eqiad.wmnet with OS trixie * 20:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7004.magru.wmnet * 20:00 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet * 19:57 aokoth@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet * 19:24 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7011.magru.wmnet * 19:19 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7003.magru.wmnet * 19:11 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] (duration: 08m 33s) * 19:07 jforrester@deploy2003: jforrester: Continuing with deployment * 19:04 jforrester@deploy2003: jforrester: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:02 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] * 18:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7010.magru.wmnet * 18:38 mutante: rotating phabricator-gerrit bot token (its-phabricator) * 18:18 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 17:44 swfrench@deploy2003: Finished scap sync-world: Deployment to pick up new production image (duration: 31m 44s) * 17:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7002.magru.wmnet * 17:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7009.magru.wmnet * 17:32 swfrench@deploy2003: swfrench: Continuing with deployment * 17:29 swfrench@deploy2003: swfrench: Deployment to pick up new production image synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:17 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2035: repooling after rack b5 maintenance * 17:16 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool es2035: repooling after rack b5 maintenance * 17:16 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2188: repooling after rack b5 maintenance * 17:12 swfrench@deploy2003: Started scap sync-world: Deployment to pick up new production image * 17:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7001.magru.wmnet * 16:57 swfrench-wmf: reprepro include php8.3_8.3.32-1+wmf12u2 into component/php83 for bookworm-wikimedia * 16:50 sukhe: pool cp2046 * 16:47 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4039.ulsfo.wmnet * 16:44 sukhe: sudo cumin -b31 "A:cp" "run-puppet-agent" * 16:33 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on contint1003.wikimedia.org with reason: reboot * 16:32 mutante: contint1003 - main CI server - rebooting * 16:31 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2188: repooling after rack b5 maintenance * 16:31 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2178: repooling after rack b5 maintenance * 16:29 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 16:28 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 16:28 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 16:28 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 16:18 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2014.codfw.wmnet * 16:18 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 16:17 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2014.codfw.wmnet * 16:10 mvernon@cumin1003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-thanos-proxies (exit_code=0) rolling restart_daemons on A:thanos-fe * 16:09 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 16:07 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp4039.ulsfo.wmnet * 16:07 mvernon@cumin1003: START - Cookbook sre.swift.roll-restart-reboot-swift-thanos-proxies rolling restart_daemons on A:thanos-fe * 16:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2014.codfw.wmnet with OS trixie * 15:56 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1014.eqiad.wmnet with OS trixie * 15:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2014.codfw.wmnet with reason: host reimage * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2014 * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2014 * 15:28 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2014 * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2014.codfw.wmnet 194.16.192.10.in-addr.arpa 4.9.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:28 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2014.codfw.wmnet 194.16.192.10.in-addr.arpa 4.9.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2014 - mvernon@cumin2003" * 15:28 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2014 - mvernon@cumin2003" * 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Apply title-related policies when selecting the name of the entity - kamila@cumin1003" * 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Apply title-related policies when selecting the name of the entity - kamila@cumin1003 * 15:22 kamila@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Apply title-related policies when selecting the name of the entity - kamila@cumin1003 * 15:22 kamila@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Apply title-related policies when selecting the name of the entity - kamila@cumin1003" * 15:20 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 15:20 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2014 * 15:20 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2014.codfw.wmnet with OS trixie * 15:19 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1014.eqiad.wmnet with OS trixie * 15:01 dancy@deploy2003: Installation of scap version "4.274.1" completed for 3 hosts * 15:00 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2177: repooling after rack b5 maintenance * 15:00 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2159: repooling after rack b5 maintenance * 14:59 dancy@deploy2003: Installing scap version "4.274.1" for 3 host(s) * 14:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2015.codfw.wmnet with OS trixie * 14:54 seanleong-wmde: Finished populateSitesTable for isvwiki ([[phab:T429939|T429939]]) * 14:53 javiermonton@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] (duration: 07m 35s) * 14:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1015.eqiad.wmnet with OS trixie * 14:49 javiermonton@deploy2003: javiermonton: Continuing with deployment * 14:48 javiermonton@deploy2003: javiermonton: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:46 javiermonton@deploy2003: Started scap sync-world: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] * 14:42 otto@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 14:41 otto@deploy2003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 14:41 otto@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 14:40 otto@deploy2003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 14:40 otto@deploy2003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 14:39 otto@deploy2003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 14:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2015.codfw.wmnet with reason: host reimage * 14:34 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1015.eqiad.wmnet with reason: host reimage * 14:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2015.codfw.wmnet with reason: host reimage * 14:30 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1015.eqiad.wmnet with reason: host reimage * 14:30 seanleong-wmde@deploy2003: mwscript-k8s job started: foreachwikiindblist wikidataclient extensions/Wikibase/lib/maintenance/populateSitesTable.php --force-protocol https # [[phab:T429939|T429939]] * 14:24 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling reboot on A:durum and not (A:durum-eqiad or A:durum-codfw or A:durum-esams) and A:durum * 14:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2015.codfw.wmnet with OS trixie * 14:15 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2016.codfw.wmnet with OS trixie * 14:14 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2159: repooling after rack b5 maintenance * 14:14 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1015.eqiad.wmnet with OS trixie * 14:12 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1016.eqiad.wmnet with OS trixie * 14:12 cmooney@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=pki,name=codfw * 14:12 sbisson@deploy2003: helmfile [codfw] DONE helmfile.d/services/cxserver: sync * 14:11 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2002.codfw.wmnet * 14:11 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2002.codfw.wmnet * 14:11 sbisson@deploy2003: helmfile [codfw] START helmfile.d/services/cxserver: sync * 14:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1005.wikimedia.org * 14:07 sbisson@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cxserver: sync * 14:07 sbisson@deploy2003: helmfile [eqiad] START helmfile.d/services/cxserver: sync * 14:05 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1005.wikimedia.org * 14:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader2005.wikimedia.org * 14:02 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2003.codfw.wmnet * 14:02 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2003.codfw.wmnet * 14:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=tegola-vector-tiles,name=codfw * 14:00 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=kartotherian,name=codfw * 14:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader2005.wikimedia.org * 13:58 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2016.codfw.wmnet with reason: host reimage * 13:57 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-ncredir (exit_code=0) rolling reboot on A:ncredir and A:ncredir * 13:57 sbisson@deploy2003: helmfile [staging] DONE helmfile.d/services/cxserver: sync * 13:56 sbisson@deploy2003: helmfile [staging] START helmfile.d/services/cxserver: sync * 13:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1016.eqiad.wmnet with reason: host reimage * 13:52 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:52 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:51 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2016.codfw.wmnet with reason: host reimage * 13:50 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1016.eqiad.wmnet with reason: host reimage * 13:49 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy (exit_code=0) rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 13:49 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling reboot on A:wikidough * 13:46 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-tcp-proxy (exit_code=0) rolling reboot on A:tcpproxy and A:tcpproxy * 13:43 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and not (A:durum-eqiad or A:durum-codfw or A:durum-esams) and A:durum * 13:42 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=97) rolling reboot on A:durum and A:durum * 13:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2011.codfw.wmnet * 13:36 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-reboot (exit_code=0) rolling reboot on A:dnsbox and A:ulsfo and (A:dnsbox) * 13:36 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns4004.wikimedia.org * 13:34 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1016.eqiad.wmnet with OS trixie * 13:34 topranks: reboot lsw1-b5-codfw to upgrade JunOS [[phab:T430918|T430918]] * 13:34 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2016.codfw.wmnet with OS trixie * 13:32 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2002.codfw.wmnet * 13:31 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2011.codfw.wmnet * 13:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2012.codfw.wmnet * 13:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2017.codfw.wmnet with OS trixie * 13:25 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1017.eqiad.wmnet with OS trixie * 13:24 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2012.codfw.wmnet * 13:22 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2002.codfw.wmnet * 13:22 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:22 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:22 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns4004.wikimedia.org * 13:21 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2005.codfw.wmnet * 13:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2013.codfw.wmnet * 13:18 elukey@dns1004: END - running authdns-update * 13:17 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2005.codfw.wmnet * 13:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2004.codfw.wmnet * 13:16 elukey@dns1004: START - running authdns-update * 13:16 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1046: es1046 after reimage * 13:14 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1029.eqiad.wmnet,service=s8 * 13:14 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1029.eqiad.wmnet,service=s5 * 13:13 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1029.eqiad.wmnet,service=s5 * 13:13 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1029.eqiad.wmnet,service=s8 * 13:13 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2004.codfw.wmnet * 13:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2013.codfw.wmnet * 13:11 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2014.codfw.wmnet * 13:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2003.codfw.wmnet * 13:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2017.codfw.wmnet with reason: host reimage * 13:07 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:07 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns4003.wikimedia.org * 13:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2003.codfw.wmnet * 13:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1067.eqiad.wmnet * 13:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1067.eqiad.wmnet * 13:06 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1067.eqiad.wmnet * 13:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm-test1001.wikimedia.org * 13:05 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2017.codfw.wmnet with reason: host reimage * 13:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1017.eqiad.wmnet with reason: host reimage * 13:04 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2014.codfw.wmnet * 13:03 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:02 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2188: codfw rack B5 depool for maintenance * 13:02 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_magru * 13:01 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2188: codfw rack B5 depool for maintenance * 13:01 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_magru * 13:01 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2178: codfw rack B5 depool for maintenance * 13:01 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm-test1001.wikimedia.org * 13:01 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2178: codfw rack B5 depool for maintenance * 13:01 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2177: codfw rack B5 depool for maintenance * 13:00 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2177: codfw rack B5 depool for maintenance * 12:59 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1068.eqiad.wmnet * 12:59 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1068.eqiad.wmnet * 12:58 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola-vector-tiles,name=codfw * 12:58 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2159: codfw rack B5 depool for maintenance * 12:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1017.eqiad.wmnet with reason: host reimage * 12:58 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola,name=codfw * 12:57 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=kartotherian,name=codfw * 12:57 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2159: codfw rack B5 depool for maintenance * 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 30 hosts with reason: lsw1-b5-codfw JunOS upgrade * 12:55 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lsw1-b5-codfw,lsw1-b5-codfw IPv6,lsw1-b5-codfw.mgmt,ssw1-a[1,8]-codfw.mgmt with reason: switch upgade lsw1-b5-codfw * 12:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet * 12:49 cmooney@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=pki,name=codfw * 12:49 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1067.eqiad.wmnet with OS trixie * 12:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet * 12:48 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2017.codfw.wmnet with OS trixie * 12:47 topranks: depool codfw pki in dns discovery ahead of lsw1-b5-codfw maintenance [[phab:T430918|T430918]] * 12:47 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns4003.wikimedia.org * 12:47 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and A:ulsfo and (A:dnsbox) * 12:47 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2018.codfw.wmnet * 12:45 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2018.codfw.wmnet * 12:45 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and A:durum * 12:45 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-tcp-proxy rolling reboot on A:tcpproxy and A:tcpproxy * 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host pki-root1002.eqiad.wmnet * 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1009.eqiad.wmnet * 12:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1009.eqiad.wmnet * 12:44 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 12:43 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-ncredir rolling reboot on A:ncredir and A:ncredir * 12:43 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling reboot on A:wikidough * 12:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2018.codfw.wmnet with OS trixie * 12:42 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1017.eqiad.wmnet with OS trixie * 12:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader2006.wikimedia.org * 12:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1018.eqiad.wmnet with OS trixie * 12:39 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1009.eqiad.wmnet * 12:38 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host pki-root1002.eqiad.wmnet * 12:38 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1009.eqiad.wmnet * 12:38 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1008.eqiad.wmnet * 12:38 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1008.eqiad.wmnet * 12:35 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader2006.wikimedia.org * 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1006.wikimedia.org * 12:33 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1008.eqiad.wmnet * 12:30 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1046: es1046 after reimage * 12:29 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host es1046.eqiad.wmnet with OS trixie * 12:29 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1006.wikimedia.org * 12:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test2005.wikimedia.org * 12:28 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1008.eqiad.wmnet * 12:27 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1007.eqiad.wmnet * 12:27 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1007.eqiad.wmnet * 12:27 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1067.eqiad.wmnet with reason: host reimage * 12:25 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1068.eqiad.wmnet with reason: vacuum overlarge container dbs * 12:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2018.codfw.wmnet with reason: host reimage * 12:24 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test2005.wikimedia.org * 12:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test1005.wikimedia.org * 12:23 Amir1: mwscript-k8s --follow --dblist=ores -- extensions/ORES/maintenance/PurgeScoreCache.php --model goodfaith --old ([[phab:T431159|T431159]]) * 12:22 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1007.eqiad.wmnet * 12:22 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test1005.wikimedia.org * 12:22 atsukoito: restarting pybal on lvs2013 `low-traffic` for https://gerrit.wikimedia.org/r/1310535 * 12:22 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1007.eqiad.wmnet * 12:21 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1006.eqiad.wmnet * 12:21 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1006.eqiad.wmnet * 12:20 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1018.eqiad.wmnet with reason: host reimage * 12:19 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2018.codfw.wmnet with reason: host reimage * 12:18 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1067.eqiad.wmnet with reason: host reimage * 12:16 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1006.eqiad.wmnet * 12:15 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1006.eqiad.wmnet * 12:15 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1005.eqiad.wmnet * 12:15 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1005.eqiad.wmnet * 12:15 atsukoito: restarting pybal on lvs2014 for https://gerrit.wikimedia.org/r/1310535 * 12:12 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1018.eqiad.wmnet with reason: host reimage * 12:11 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1005.eqiad.wmnet * 12:11 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1005.eqiad.wmnet * 12:10 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1004.eqiad.wmnet * 12:10 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1004.eqiad.wmnet * 12:09 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on es1046.eqiad.wmnet with reason: host reimage * 12:08 atsukoito: restarting pybal on lvs1019 `low-traffic` for https://gerrit.wikimedia.org/r/1310535 * 12:06 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1004.eqiad.wmnet * 12:06 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1004.eqiad.wmnet * 12:06 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1003.eqiad.wmnet * 12:06 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1003.eqiad.wmnet * 12:05 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on es1046.eqiad.wmnet with reason: host reimage * 12:05 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "set ml-serve1001 back to active state - cmooney@cumin1003" * 12:04 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "set ml-serve1001 back to active state - cmooney@cumin1003" * 12:04 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:02 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1003.eqiad.wmnet * 12:01 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1003.eqiad.wmnet * 12:01 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1002.eqiad.wmnet * 12:01 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1002.eqiad.wmnet * 12:01 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:59 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2018.codfw.wmnet with OS trixie * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1067 * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1067 * 11:59 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1067 * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1067.eqiad.wmnet 17.48.64.10.in-addr.arpa 7.1.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:59 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1067.eqiad.wmnet 17.48.64.10.in-addr.arpa 7.1.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1067 - blake@cumin1003" * 11:58 atsukoito: restarting pybal on lvs1018 `high-traffic2` for https://gerrit.wikimedia.org/r/1310535 * 11:57 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1002.eqiad.wmnet * 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2019.codfw.wmnet with OS trixie * 11:56 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1002.eqiad.wmnet * 11:56 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 11:56 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1018.eqiad.wmnet with OS trixie * 11:54 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 11:54 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:54 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1019.eqiad.wmnet with OS trixie * 11:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:49 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:49 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:49 aikochou@deploy2003: helmfile [codfw] DONE helmfile.d/services/changeprop: sync * 11:48 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host es1046.eqiad.wmnet with OS trixie * 11:48 aikochou@deploy2003: helmfile [codfw] START helmfile.d/services/changeprop: sync * 11:48 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310535 * 11:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1046: Reimage to Trixie * 11:44 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1046: Reimage to Trixie * 11:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5:00:00 on es1046.eqiad.wmnet with reason: Reimage to Trixie * 11:42 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:42 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:42 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 11:42 aikochou@deploy2003: helmfile [eqiad] DONE helmfile.d/services/changeprop: sync * 11:41 aikochou@deploy2003: helmfile [eqiad] START helmfile.d/services/changeprop: sync * 11:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2019.codfw.wmnet with reason: host reimage * 11:36 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] (duration: 09m 41s) * 11:36 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:36 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:35 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:35 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:32 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1019.eqiad.wmnet with reason: host reimage * 11:32 jforrester@deploy2003: jforrester, gengh: Continuing with deployment * 11:29 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2019.codfw.wmnet with reason: host reimage * 11:28 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1019.eqiad.wmnet with reason: host reimage * 11:28 jforrester@deploy2003: jforrester, gengh: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:26 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] * 11:20 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2003.codfw.wmnet * 11:20 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:19 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2003.codfw.wmnet * 11:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:12 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1019.eqiad.wmnet with OS trixie * 11:10 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2019.codfw.wmnet with OS trixie * 11:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1020.eqiad.wmnet with OS trixie * 11:10 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1067 - blake@cumin1003" * 11:09 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] (duration: 12m 12s) * 11:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2020.codfw.wmnet with OS trixie * 11:03 kharlan@deploy2003: kharlan: Continuing with deployment * 11:01 blake@cumin1003: START - Cookbook sre.dns.netbox * 11:01 kharlan@deploy2003: kharlan: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:57 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] * 10:55 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] (duration: 31m 40s) * 10:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1020.eqiad.wmnet with reason: host reimage * 10:52 marostegui@dns1004: START - running authdns-update * 10:49 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2020.codfw.wmnet with reason: host reimage * 10:49 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:48 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1020.eqiad.wmnet with reason: host reimage * 10:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2020.codfw.wmnet with reason: host reimage * 10:43 kharlan@deploy2003: kharlan: Continuing with deployment * 10:42 kharlan@deploy2003: kharlan: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:32 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1020.eqiad.wmnet with OS trixie * 10:29 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2159: Repooling after switchover * 10:29 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1067 * 10:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1021.eqiad.wmnet with OS trixie * 10:27 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1067.eqiad.wmnet with OS trixie * 10:27 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:27 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1067.eqiad.wmnet * 10:27 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:26 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1067.eqiad.wmnet * 10:26 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1067.eqiad.wmnet * 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2020.codfw.wmnet with OS trixie * 10:26 blake@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1055.eqiad.wmnet * 10:26 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1055.eqiad.wmnet * 10:26 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1055.eqiad.wmnet * 10:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2021.codfw.wmnet with OS trixie * 10:24 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] * 10:11 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1055.eqiad.wmnet with OS trixie * 10:09 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1021.eqiad.wmnet with reason: host reimage * 10:05 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2021.codfw.wmnet with reason: host reimage * 10:03 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310129 revert * 10:02 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1021.eqiad.wmnet with reason: host reimage * 10:01 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2021.codfw.wmnet with reason: host reimage * 09:58 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310129 * 09:50 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1055.eqiad.wmnet with reason: host reimage * 09:45 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1055.eqiad.wmnet with reason: host reimage * 09:45 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1021.eqiad.wmnet with OS trixie * 09:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1022.eqiad.wmnet with OS trixie * 09:44 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2159: Repooling after switchover * 09:44 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2021.codfw.wmnet with OS trixie * 09:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2022.codfw.wmnet with OS trixie * 09:31 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2159.codfw.wmnet * 09:28 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1055 * 09:28 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1055 * 09:27 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on ms-fe1022.eqiad.wmnet with reason: host reimage * 09:27 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1022.eqiad.wmnet with reason: host reimage * 09:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2022.codfw.wmnet with reason: host reimage * 09:21 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2022.codfw.wmnet with reason: host reimage * 09:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2159: Rebooting db2159.codfw.wmnet * 09:20 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2159: Rebooting db2159.codfw.wmnet * 09:18 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 09:18 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 09:18 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 09:17 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 09:16 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2159.codfw.wmnet * 09:13 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b] (thin): Regular analytics weekly train THIN [analytics/refinery@ad6e05b8] (duration: 02m 07s) * 09:11 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b] (thin): Regular analytics weekly train THIN [analytics/refinery@ad6e05b8] * 09:10 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1022.eqiad.wmnet with OS trixie * 09:07 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1023.eqiad.wmnet with OS trixie * 09:06 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b]: Regular analytics weekly train [analytics/refinery@ad6e05b8] (duration: 04m 49s) * 09:04 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2022.codfw.wmnet with OS trixie * 09:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2023.codfw.wmnet with OS trixie * 09:01 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b]: Regular analytics weekly train [analytics/refinery@ad6e05b8] * 09:01 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@ad6e05b8] (duration: 02m 01s) * 09:00 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1055 * 09:00 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1055.eqiad.wmnet 50.32.64.10.in-addr.arpa 0.5.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:00 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1055.eqiad.wmnet 50.32.64.10.in-addr.arpa 0.5.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:00 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:00 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1055 - blake@cumin1003" * 09:00 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1055 - blake@cumin1003" * 08:59 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@ad6e05b8] * 08:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2159 [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94811 and previous config saved to /var/cache/conftool/dbconfig/20260714-085624-cwilliams.json * 08:55 blake@cumin1003: START - Cookbook sre.dns.netbox * 08:55 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1055 * 08:54 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1055.eqiad.wmnet with OS trixie * 08:54 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1055.eqiad.wmnet * 08:53 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1055.eqiad.wmnet * 08:53 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1055.eqiad.wmnet * 08:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2220 to s7 primary [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94810 and previous config saved to /var/cache/conftool/dbconfig/20260714-085239-cwilliams.json * 08:51 cezmunsta: Starting s7 codfw failover from db2159 to db2220 - [[phab:T430920|T430920]] * 08:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1023.eqiad.wmnet with reason: host reimage * 08:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2220 with weight 0 [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94809 and previous config saved to /var/cache/conftool/dbconfig/20260714-084553-cwilliams.json * 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s7 [[phab:T430920|T430920]] * 08:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2023.codfw.wmnet with reason: host reimage * 08:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1023.eqiad.wmnet with reason: host reimage * 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2023.codfw.wmnet with reason: host reimage * 08:34 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:34 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:29 marostegui@dns1004: END - running authdns-update * 08:29 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox-dev2003.codfw.wmnet * 08:27 marostegui@dns1004: START - running authdns-update * 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker2*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2009.codfw.wmnet * 08:26 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2009.codfw.wmnet * 08:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1023.eqiad.wmnet with OS trixie * 08:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox-dev2003.codfw.wmnet * 08:24 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:24 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:24 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2023.codfw.wmnet with OS trixie * 08:24 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1029.eqiad.wmnet with reason: reboot * 08:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1027.eqiad.wmnet with reason: reboot * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:21 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2009.codfw.wmnet * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:20 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2009.codfw.wmnet * 08:20 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2008.codfw.wmnet * 08:20 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2008.codfw.wmnet * 08:15 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2008.codfw.wmnet * 08:14 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2008.codfw.wmnet * 08:14 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2007.codfw.wmnet * 08:14 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2007.codfw.wmnet * 08:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1024.eqiad.wmnet with OS trixie * 08:12 elukey@cumin1003: END (PASS) - Cookbook sre.pki.restart-reboot (exit_code=0) rolling reboot on P<nowiki>{</nowiki>pki*<nowiki>}</nowiki> and (A:pki) * 08:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 08:10 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 08:09 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2007.codfw.wmnet * 08:08 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2007.codfw.wmnet * 08:08 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2006.codfw.wmnet * 08:08 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2006.codfw.wmnet * 08:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2024.codfw.wmnet with OS trixie * 08:03 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2006.codfw.wmnet * 08:02 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2006.codfw.wmnet * 08:02 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2005.codfw.wmnet * 08:02 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2005.codfw.wmnet * 07:58 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2005.codfw.wmnet * 07:58 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2005.codfw.wmnet * 07:57 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2004.codfw.wmnet * 07:57 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2004.codfw.wmnet * 07:54 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki.discovery.wmnet. on all recursors * 07:54 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache pki.discovery.wmnet. on all recursors * 07:53 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2004.codfw.wmnet * 07:53 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2004.codfw.wmnet * 07:53 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2003.codfw.wmnet * 07:53 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2003.codfw.wmnet * 07:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1024.eqiad.wmnet with reason: host reimage * 07:49 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki.discovery.wmnet. on all recursors * 07:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2003.codfw.wmnet * 07:49 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache pki.discovery.wmnet. on all recursors * 07:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2024.codfw.wmnet with reason: host reimage * 07:48 elukey@cumin1003: START - Cookbook sre.pki.restart-reboot rolling reboot on P<nowiki>{</nowiki>pki*<nowiki>}</nowiki> and (A:pki) * 07:46 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1024.eqiad.wmnet with reason: host reimage * 07:45 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2024.codfw.wmnet with reason: host reimage * 07:45 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2003.codfw.wmnet * 07:44 elukey@cumin1003: END (PASS) - Cookbook sre.misc-clusters.restart-reboot-config-master (exit_code=0) rolling reboot on P<nowiki>{</nowiki>config-master*<nowiki>}</nowiki> and (A:config-master or A:config-master-eqiad or A:config-master-codfw) * 07:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2002.codfw.wmnet * 07:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2002.codfw.wmnet * 07:39 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2002.codfw.wmnet * 07:39 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) config-master.discovery.wmnet. on all recursors * 07:39 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache config-master.discovery.wmnet. on all recursors * 07:39 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2002.codfw.wmnet * 07:39 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker2*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl200*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl2003.codfw.wmnet * 07:36 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl2003.codfw.wmnet * 07:35 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) config-master.discovery.wmnet. on all recursors * 07:35 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache config-master.discovery.wmnet. on all recursors * 07:34 elukey@cumin1003: START - Cookbook sre.misc-clusters.restart-reboot-config-master rolling reboot on P<nowiki>{</nowiki>config-master*<nowiki>}</nowiki> and (A:config-master or A:config-master-eqiad or A:config-master-codfw) * 07:31 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl2003.codfw.wmnet * 07:31 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl2003.codfw.wmnet * 07:31 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl2002.codfw.wmnet * 07:31 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl2002.codfw.wmnet * 07:29 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1024.eqiad.wmnet with OS trixie * 07:28 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2024.codfw.wmnet with OS trixie * 07:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl2002.codfw.wmnet * 07:26 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl2002.codfw.wmnet * 07:26 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl200*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 07:26 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 07:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 06:50 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lists1004.wikimedia.org * 06:44 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host lists1004.wikimedia.org * 06:25 marostegui@dns1004: END - running authdns-update * 06:23 marostegui@dns1004: START - running authdns-update * 06:22 marostegui@dns1004: END - running authdns-update * 06:20 marostegui@dns1004: START - running authdns-update * 06:17 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1026.eqiad.wmnet with reason: reboot * 06:04 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: sync * 06:04 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: sync * 06:03 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync * 06:03 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync * 06:02 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync * 06:01 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync * 06:01 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync * 06:00 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync * 05:59 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:59 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:40 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:39 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:26 marostegui@dns1004: END - running authdns-update * 05:24 marostegui@dns1004: START - running authdns-update * 05:24 marostegui@dns1004: START - running authdns-update * 05:13 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1004.wikimedia.org * 05:07 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1004.wikimedia.org * 04:01 mwpresync@deploy2003: Pruned MediaWiki: 1.47.0-wmf.8 (duration: 01m 07s) * 03:39 mwpresync@deploy2003: Finished scap sync-world: testwikis to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] (duration: 36m 01s) * 03:03 mwpresync@deploy2003: Started scap sync-world: testwikis to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 29s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-13 == * 23:33 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1064.eqiad.wmnet * 23:33 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1064.eqiad.wmnet * 23:08 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1064.eqiad.wmnet with reason: vacuum overlarge container dbs * 23:06 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1069.eqiad.wmnet * 23:06 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1069.eqiad.wmnet * 22:34 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1069.eqiad.wmnet with reason: vacuum overlarge container dbs * 21:18 maryum: Deployed security fix for [[phab:T321092|T321092]] * 20:28 swfrench-wmf: reprepro include etcd-mirror_0.0.12-1+deb13u1 into main for trixie-wikimedia - [[phab:T424266|T424266]] * 20:26 swfrench-wmf: reprepro include etcd-mirror_0.0.12-1+deb12u1 into main for bookworm-wikimedia - [[phab:T428495|T428495]] * 20:23 dancy@deploy2003: Finished scap sync-world: Testing [[phab:T431635|T431635]] (duration: 03m 36s) * 20:19 dancy@deploy2003: Started scap sync-world: Testing [[phab:T431635|T431635]] * 20:18 dancy@deploy2003: Installation of scap version "4.274.0" completed for 3 hosts * 20:16 dancy@deploy2003: Installing scap version "4.274.0" for 3 host(s) * 20:12 kemayo@deploy2003: Finished scap sync-world: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] (duration: 08m 25s) * 20:07 kemayo@deploy2003: soda, esanders, kemayo: Continuing with deployment * 20:05 kemayo@deploy2003: soda, esanders, kemayo: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there * 20:04 kemayo@deploy2003: Started scap sync-world: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] * 18:22 cwhite: lvextend vg0/srv +500g on centrallog hosts * 18:19 cdobbins@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS trixie * 17:46 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1071.eqiad.wmnet * 17:46 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1071.eqiad.wmnet * 17:13 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1071.eqiad.wmnet with reason: vacuum overlarge container dbs * 17:07 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1065.eqiad.wmnet * 17:07 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1065.eqiad.wmnet * 17:06 dzahn@dns1006: END - running authdns-update * 17:04 dzahn@dns1006: START - running authdns-update * 17:01 dzahn@dns1006: END - running authdns-update * 16:59 dzahn@dns1006: START - running authdns-update * 16:51 dancy@deploy2003: Finished scap sync-world: testing [[phab:T428971|T428971]] (duration: 03m 37s) * 16:47 dancy@deploy2003: Started scap sync-world: testing [[phab:T428971|T428971]] * 16:45 atsukoito: restarting pybal on lvs1019 to flush IP address for `cirrussearch1122.eqiad.wmnet` after moving the vlan [[phab:T431311|T431311]] * 16:42 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:42 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:42 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:42 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:42 Amir1: mwscript-k8s --follow --dblist=ores -- extensions/ORES/maintenance/PurgeScoreCache.php --model damaging --old ([[phab:T431159|T431159]]) * 16:34 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Pool test * 16:34 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 16:34 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 16:34 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Pool test * 16:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Depool test * 16:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 16:33 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 16:33 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Depool test * 16:31 dancy@deploy2003: Installation of scap version "4.273.0" completed for 159 hosts * 16:29 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1065.eqiad.wmnet with reason: vacuum overlarge container dbs * 16:27 dancy@deploy2003: Installing scap version "4.273.0" for 159 host(s) * 16:27 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics-external: sync * 16:27 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics-external: sync * 16:26 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics-external: sync * 16:26 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics-external: sync * 16:22 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync * 16:21 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync * 16:21 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: sync * 16:21 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: sync * 16:19 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync * 16:19 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync * 16:18 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync * 16:17 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync * 15:59 atsukoito: restarting pybal on lvs1018 for https://gerrit.wikimedia.org/r/1310117 * 15:55 aikochou@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 15:50 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310117 * 15:46 aikochou@deploy2003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 15:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host kafka-logging1006.eqiad.wmnet * 15:43 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host kafka-logging1006.eqiad.wmnet * 15:41 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host ganeti-test[2001-2003].codfw.wmnet * 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host ganeti-test[2001-2003].codfw.wmnet * 15:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host netbox1003.eqiad.wmnet * 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host netbox1003.eqiad.wmnet * 15:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host netbox2003.codfw.wmnet * 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host netbox2003.codfw.wmnet * 15:36 sukhe: restart pybal on lvs1020 * 15:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet * 15:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet * 15:08 btullis@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:06 btullis@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 15:01 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:01 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:35 cdobbins@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 14:34 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:33 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:33 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:32 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:29 cdobbins@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 14:28 marostegui@dns1004: END - running authdns-update * 14:28 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:27 marostegui@dns1004: START - running authdns-update * 14:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1023.eqiad.wmnet with reason: reboot * 14:18 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2009.codfw.wmnet with OS trixie * 14:14 swfrench-wmf: start rolling run-puppet-agent on A:cp for ATS config change - [[phab:T428909|T428909]] [[phab:T431838|T431838]] * 14:05 swfrench-wmf: disable-puppet on A:cp for ATS config change - [[phab:T428909|T428909]] [[phab:T431838|T431838]] * 14:05 cdobbins@cumin2002: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie * 14:02 marostegui@dns1004: END - running authdns-update * 14:00 marostegui@dns1004: START - running authdns-update * 14:00 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1070.eqiad.wmnet * 14:00 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1070.eqiad.wmnet * 13:58 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2009.codfw.wmnet with reason: host reimage * 13:57 cdobbins@cumin2002: conftool action : set/pooled=no; selector: name=dns7002.* * 13:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2009.codfw.wmnet with reason: host reimage * 13:48 rscout@deploy2003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply * 13:48 rscout@deploy2003: helmfile [eqiad] START helmfile.d/services/miscweb: apply * 13:48 rscout@deploy2003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply * 13:47 rscout@deploy2003: helmfile [codfw] START helmfile.d/services/miscweb: apply * 13:40 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:33 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2009.codfw.wmnet with OS trixie * 13:30 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1070.eqiad.wmnet with reason: vacuum overlarge container dbs * 13:28 aude@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] (duration: 11m 12s) * 13:23 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:22 aude@deploy2003: aikochou, javiermonton, aude, gkm563: Continuing with deployment * 13:22 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:19 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:19 aude@deploy2003: aikochou, javiermonton, aude, gkm563: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] synced to the testservers * 13:17 aude@deploy2003: Started scap sync-world: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] * 13:01 ladsgroup@deploy2003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 13:01 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:00 ladsgroup@deploy2003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 12:59 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:52 ladsgroup@deploy2003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 12:51 ladsgroup@deploy2003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 12:48 atsuko@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 12:48 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 12:47 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2008.codfw.wmnet with OS trixie * 12:47 atsuko@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 12:47 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply * 12:47 atsuko@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:46 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 12:45 atsuko@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:45 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply * 12:45 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:44 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:43 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] (duration: 07m 02s) * 12:38 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 12:37 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:36 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] * 12:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2008.codfw.wmnet with reason: host reimage * 12:23 Msz2001: Deployed changes to private code for Suggested Investigations * 12:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2008.codfw.wmnet with reason: host reimage * 12:20 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:19 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:17 atsuko@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 12:17 atsuko@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 12:16 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:15 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] (duration: 07m 14s) * 12:10 mszwarc@deploy2003: mszwarc: Continuing with deployment * 12:09 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:07 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] * 12:04 mszwarc@deploy2003: sync-world aborted: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] (duration: 00m 29s) * 12:03 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] * 12:01 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2008.codfw.wmnet with OS trixie * 12:00 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] (duration: 07m 37s) * 11:55 zabe@deploy2003: zabe: Continuing with deployment * 11:54 zabe@deploy2003: zabe: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:52 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] * 11:51 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:43 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:35 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:34 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:33 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:30 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:28 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:27 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:17 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2007.codfw.wmnet with OS trixie * 11:09 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] * 11:06 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=s8 * 11:00 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=x3 * 11:00 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=s5 * 10:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2007.codfw.wmnet with reason: host reimage * 10:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2007.codfw.wmnet with reason: host reimage * 10:51 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host dse-k8s-worker1023 * 10:50 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host dse-k8s-worker1023 * 10:44 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dse-k8s-worker1023 * 10:43 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host dse-k8s-worker1023 * 10:42 marostegui@cumin1003: dbctl commit (dc=all): 'Change x4 masters [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P94804 and previous config saved to /var/cache/conftool/dbconfig/20260713-104248-marostegui.json * 10:37 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:37 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:35 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:35 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:34 atsuko@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 10:34 atsuko@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 10:33 marostegui@cumin1003: dbctl commit (dc=all): 'Push x4 initial dbctl config [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P94803 and previous config saved to /var/cache/conftool/dbconfig/20260713-103259-marostegui.json * 10:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2007.codfw.wmnet with OS trixie * 09:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2006.codfw.wmnet with OS trixie * 09:42 marostegui@dns1004: END - running authdns-update * 09:40 marostegui@dns1004: START - running authdns-update * 09:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2006.codfw.wmnet with reason: host reimage * 09:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2006.codfw.wmnet with reason: host reimage * 09:06 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1024.eqiad.wmnet with reason: reboot * 09:06 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:01 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2006.codfw.wmnet with OS trixie * 08:44 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 08:43 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 08:43 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:42 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:42 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:42 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:41 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 08:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup2004.codfw.wmnet * 08:38 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:33 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host db1208.eqiad.wmnet * 08:30 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=x3 * 08:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2005.codfw.wmnet with OS trixie * 08:28 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup2004.codfw.wmnet * 08:28 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup2003.codfw.wmnet * 08:24 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1039: Repooling after testing * 08:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on clouddb1016.eqiad.wmnet with reason: cloning * 08:23 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s5 * 08:23 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s8 * 08:21 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 08:21 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 08:17 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup2003.codfw.wmnet * 08:17 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1004.eqiad.wmnet * 08:14 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1208.eqiad.wmnet * 08:11 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host phab1005.eqiad.wmnet * 08:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2005.codfw.wmnet with reason: host reimage * 08:07 marostegui@dns1004: END - running authdns-update * 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1004.eqiad.wmnet * 08:07 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1003.eqiad.wmnet * 08:07 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:05 marostegui@dns1004: START - running authdns-update * 08:05 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2005.codfw.wmnet with reason: host reimage * 08:05 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 08:05 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host phab1005.eqiad.wmnet * 08:05 marostegui@dns1004: START - running authdns-update * 08:05 marostegui@dns1004: START - running authdns-update * 08:05 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 08:04 marostegui@dns1004: START - running authdns-update * 08:00 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit1003.wikimedia.org * 07:58 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1003.eqiad.wmnet * 07:58 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1002-dev.eqiad.wmnet * 07:58 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:58 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:54 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1002-dev.eqiad.wmnet * 07:54 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1001-dev.eqiad.wmnet * 07:54 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit1003.wikimedia.org * 07:53 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 07:53 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:52 Msz2001: UTC morning backport+config window done * 07:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2005.codfw.wmnet with OS trixie * {{safesubst:SAL entry|1=07:50 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark (T429943}} * 07:49 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1001-dev.eqiad.wmnet * 07:46 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:46 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:45 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 07:45 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:44 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 07:44 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:43 mszwarc@deploy2003: mszwarc, danielyepezgarces, anzx: Continuing with deployment * {{safesubst:SAL entry|1=07:39 mszwarc@deploy2003: mszwarc, danielyepezgarces, anzx: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark}} * 07:39 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1039: Repooling after testing * {{safesubst:SAL entry|1=07:36 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark (T429943)}} * 07:35 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] (duration: 30m 03s) * 07:25 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit2002.wikimedia.org * 07:22 mszwarc@deploy2003: mszwarc: Continuing with deployment * 07:21 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:19 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit2002.wikimedia.org * 07:15 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aphlict1002.eqiad.wmnet * 07:11 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host aphlict1002.eqiad.wmnet * 07:08 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2003.wikimedia.org * 07:05 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] * 07:02 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2003.wikimedia.org * 07:02 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2002.wikimedia.org * 06:55 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2002.wikimedia.org * 06:55 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1003.wikimedia.org * 06:49 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1003.wikimedia.org * 06:34 marostegui: Drop m5 ipoid database [[phab:T431007|T431007]] * 06:29 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1027.eqiad.wmnet with reason: reboot * 06:24 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1028.eqiad.wmnet with reason: reboot * 06:21 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1025.eqiad.wmnet with reason: reboot * 06:17 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1022.eqiad.wmnet with reason: reboot * 06:03 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on dbproxy[2005-2008].codfw.wmnet with reason: reboot * 05:37 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1217,1228].eqiad.wmnet with reason: cloning * 05:11 marostegui: Drop users_to_rename table [[phab:T431842|T431842]] * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-12 == * 16:01 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2209 [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94792 and previous config saved to /var/cache/conftool/dbconfig/20260712-160124-marostegui.json * 15:58 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2205 to s3 primary [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94791 and previous config saved to /var/cache/conftool/dbconfig/20260712-155853-marostegui.json * 15:58 marostegui: Starting s3 codfw emergency failover from db2209 to db2205 - [[phab:T431950|T431950]] * 15:51 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2205 with weight 0 [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94790 and previous config saved to /var/cache/conftool/dbconfig/20260712-155135-marostegui.json * 15:51 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Primary switchover s3 [[phab:T431950|T431950]] * 02:01 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 01m 17s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-11 == * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 26s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-10 == * 19:12 jhathaway@dns1004: END - running authdns-update * 19:10 jhathaway@dns1004: START - running authdns-update * 18:23 mutante: vrts2002 rebooting (not the active host) * 18:21 mutante: lists2001, phab2003 - rebooting (not the active hosts) * 18:16 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on A:lvs-high-traffic2-codfw * 18:15 swfrench@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on A:lvs-high-traffic2-codfw * 17:15 mutante: [doc1004:~] $ sudo systemctl start rsync-doc-host-data-sync ([[phab:T431856|T431856]]) * 17:09 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1004.eqiad.wmnet * 17:08 jhathaway@dns1004: END - running authdns-update * 17:07 jhathaway@dns1004: START - running authdns-update * 17:06 jhathaway: depooling puppetserver1002, cause of errors is still unknown * 17:03 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1004.eqiad.wmnet * 16:57 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1003.eqiad.wmnet * 16:51 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1003.eqiad.wmnet * 16:48 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 16:48 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2004.codfw.wmnet * 16:42 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2004.codfw.wmnet * 16:41 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2003.codfw.wmnet * 16:35 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2003.codfw.wmnet * 16:33 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2002.codfw.wmnet * 16:27 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2002.codfw.wmnet * 16:25 mutante: gitlab-runners (production) rebooting cluster one by one * 16:17 mutante: etherpad1004/etherpad2002 - (etherpad.wikimedia.org) - rebooting * 16:13 mutante: doc1004/doc2003 (doc.wikimedia.org backends) - rebooting * 16:02 mutante: releases1003/releases2003 (releases.wikimedia.org backends) - rebooting for maintenance * 15:26 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2007-dev.codfw.wmnet * 15:19 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2007-dev.codfw.wmnet * 15:14 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host cloudcephosd2007-dev.codfw.wmnet * 15:14 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2007-dev.codfw.wmnet * 15:14 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host cloudcephosd2006-dev.codfw.wmnet * 15:07 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2006-dev.codfw.wmnet * 15:07 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2005-dev.codfw.wmnet * 14:59 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2005-dev.codfw.wmnet * 14:59 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2004-dev.codfw.wmnet * 14:53 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2004-dev.codfw.wmnet * 14:53 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2007-dev.codfw.wmnet * 14:51 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1054.eqiad.wmnet * 14:51 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1054.eqiad.wmnet * 14:51 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1054.eqiad.wmnet * 14:47 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2007-dev.codfw.wmnet * 14:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2006-dev.codfw.wmnet * 14:41 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2006-dev.codfw.wmnet * 14:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2005-dev.codfw.wmnet * 14:37 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2005-dev.codfw.wmnet * 14:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2005-dev.codfw.wmnet * 14:29 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2005-dev.codfw.wmnet * 14:29 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2006-dev.codfw.wmnet * 14:21 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2006-dev.codfw.wmnet * 14:21 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2010-dev.codfw.wmnet * 14:15 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2010-dev.codfw.wmnet * 14:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudgw2004-dev.codfw.wmnet * 14:10 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1054.eqiad.wmnet with OS trixie * 14:09 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudgw2004-dev.codfw.wmnet * 14:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudgw2003-dev.codfw.wmnet * 14:02 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudgw2003-dev.codfw.wmnet * 14:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2004-dev.codfw.wmnet * 13:53 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2004-dev.codfw.wmnet * 13:53 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2003-dev.codfw.wmnet * 13:48 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:44 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2003-dev.codfw.wmnet * 13:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2002-dev.codfw.wmnet * 13:42 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:41 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:41 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:37 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2002-dev.codfw.wmnet * 13:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudidp2001-dev.codfw.wmnet * 13:33 blake@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker1054.eqiad.wmnet with reason: host reimage * 13:33 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudidp2001-dev.codfw.wmnet * 13:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudnet2006-dev.codfw.wmnet * 13:26 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudnet2006-dev.codfw.wmnet * 13:26 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudnet2005-dev.codfw.wmnet * 13:23 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1054.eqiad.wmnet with reason: host reimage * 13:18 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudnet2005-dev.codfw.wmnet * 13:18 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudservices2005-dev.codfw.wmnet * 13:12 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudservices2005-dev.codfw.wmnet * 13:11 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudservices2004-dev.codfw.wmnet * 13:08 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudservices2004-dev.codfw.wmnet * 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudweb2002-dev.wikimedia.org * 13:05 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 13:05 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1054 * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1054 * 13:04 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1054 * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1054.eqiad.wmnet 49.32.64.10.in-addr.arpa 9.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:04 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1054.eqiad.wmnet 49.32.64.10.in-addr.arpa 9.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1054 - blake@cumin1003" * 13:04 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1054 - blake@cumin1003" * 13:01 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudweb2002-dev.wikimedia.org * 13:00 blake@cumin1003: START - Cookbook sre.dns.netbox * 12:59 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1054 * 12:57 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1054.eqiad.wmnet with OS trixie * 12:57 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1054.eqiad.wmnet * 12:56 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1054.eqiad.wmnet * 12:56 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1054.eqiad.wmnet * 12:47 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS trixie * 12:44 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:39 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 12:39 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 12:38 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:37 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:14 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:10 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:08 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:07 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:00 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:00 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:51 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:49 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:48 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:47 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:44 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:32 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2001.codfw.wmnet * 11:32 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1053.eqiad.wmnet * 11:32 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2001.codfw.wmnet * 11:32 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1053.eqiad.wmnet * 11:32 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1053.eqiad.wmnet * 11:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker2001.codfw.wmnet * 11:31 cgoubert@cumin1003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker2001.codfw.wmnet * 11:31 cgoubert@cumin1003: END (FAIL) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=1) rolling reimage on P<nowiki>{</nowiki>wikikube-worker2001*<nowiki>}</nowiki> and (A:wikikube-master-codfw or A:wikikube-worker-codfw) * 11:30 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:30 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:21 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 18 hosts with reason: reboot & upgrade * 11:20 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker2001.codfw.wmnet with OS trixie * 11:16 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:15 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:14 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:14 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 11 hosts * 11:14 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 11 hosts * 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:08 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 11:02 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1053.eqiad.wmnet with OS trixie * 11:01 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:58 cgoubert@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 10:57 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:38 cgoubert@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker2001.codfw.wmnet with OS trixie * 10:38 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2001.codfw.wmnet * 10:38 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2001.codfw.wmnet * 10:38 cgoubert@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on P<nowiki>{</nowiki>wikikube-worker2001*<nowiki>}</nowiki> and (A:wikikube-master-codfw or A:wikikube-worker-codfw) * 10:35 topranks: adjust IBGP outbound policy on lsw1-e2-codfw [[phab:T423430|T423430]] towards ssw1-e1-codfw * 10:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 cgoubert@cumin1003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:27 cgoubert@cumin1003: END (FAIL) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=1) rolling reimage on A:wikikube-worker-codfw * 10:27 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker2001.codfw.wmnet with OS bookworm * 10:25 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 10:24 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:24 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:15 cgoubert@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 10:11 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 10:11 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 10:08 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:08 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:07 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 10:06 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 10:00 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:55 cgoubert@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker2001.codfw.wmnet with OS bookworm * 09:55 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2005-2006,2011-2012].codfw.wmnet * 09:55 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2005-2006,2011-2012].codfw.wmnet * 09:51 cgoubert@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on A:wikikube-worker-codfw * 09:41 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1053.eqiad.wmnet with reason: host reimage * 09:37 topranks: apply new IBGP outbound policy on lsw1-e2-codfw [[phab:T423430|T423430]] * 09:36 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:36 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1053.eqiad.wmnet with reason: host reimage * 09:16 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1053 * 09:16 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1053 * 09:15 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1053 * 09:15 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1053.eqiad.wmnet 48.32.64.10.in-addr.arpa 8.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:15 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1053.eqiad.wmnet 48.32.64.10.in-addr.arpa 8.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:15 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:15 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1053 - blake@cumin1003" * 09:15 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1053 - blake@cumin1003" * 09:11 blake@cumin1003: START - Cookbook sre.dns.netbox * 09:11 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1053 * 09:08 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1053.eqiad.wmnet with OS trixie * 09:08 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1053.eqiad.wmnet * 09:08 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1053.eqiad.wmnet * 09:08 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1053.eqiad.wmnet * 09:04 brouberol@dns1004: END - running authdns-update * 09:03 brouberol@dns1004: START - running authdns-update * 08:41 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e] (thin): Regular analytics weekly train THIN [analytics/refinery@1abf22ea] (duration: 02m 11s) * 08:38 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e] (thin): Regular analytics weekly train THIN [analytics/refinery@1abf22ea] * 08:38 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e]: Regular analytics weekly train [analytics/refinery@1abf22ea] (duration: 05m 17s) * 08:38 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:34 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:33 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e]: Regular analytics weekly train [analytics/refinery@1abf22ea] * 08:32 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@1abf22ea] (duration: 02m 03s) * 08:30 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@1abf22ea] * 08:30 JavierMonton: Deploying Refinery at {{Gerrit|1abf22ea}} for changes 1308121/T427068 1306491/T430020 and {{Gerrit|1308190}} * 08:29 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:29 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 08:24 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:24 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 08:18 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:18 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 08:00 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db[2183-2184].codfw.wmnet * 08:00 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for db[2183-2184].codfw.wmnet * 07:52 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:52 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 07:49 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 11 hosts with reason: reboot & upgrade * 07:47 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:47 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 07:44 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:44 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 07:23 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 10 hosts * 07:23 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 10 hosts * 06:45 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 10 hosts with reason: reboot & upgrade * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 41s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-09 == * 23:33 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] (duration: 13m 26s) * 23:29 ladsgroup@deploy2003: ladsgroup, jdlrobson: Continuing with deployment * 23:22 ladsgroup@deploy2003: ladsgroup, jdlrobson: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:20 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] * 22:57 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1165.eqiad.wmnet * 22:56 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1165.eqiad.wmnet * 22:56 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1165.eqiad.wmnet * 22:45 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1165.eqiad.wmnet with OS trixie * 22:38 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 22:37 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 22:37 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 22:37 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:37 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 22:25 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1165.eqiad.wmnet with reason: host reimage * 22:17 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1165.eqiad.wmnet with reason: host reimage * 22:13 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 22:12 rzl@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 22:04 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 22:04 rzl@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1165 * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1165 * 22:02 jasmine@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1165 * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1165.eqiad.wmnet 115.48.64.10.in-addr.arpa 5.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:02 jasmine@cumin2002: START - Cookbook sre.dns.wipe-cache wikikube-worker1165.eqiad.wmnet 115.48.64.10.in-addr.arpa 5.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1165 - jasmine@cumin2002" * 22:02 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1165 - jasmine@cumin2002" * 22:02 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 21:57 jasmine@cumin2002: START - Cookbook sre.dns.netbox * 21:55 jasmine@cumin2002: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1165 * 21:54 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-worker1165.eqiad.wmnet with OS trixie * 21:54 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 21:54 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1165.eqiad.wmnet * 21:53 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 21:53 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1165.eqiad.wmnet * 21:53 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1165.eqiad.wmnet * 21:53 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 21:47 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 21:45 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 21:43 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 21:43 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 21:42 maryum: Deploy fix for [[phab:T431684|T431684]] * 21:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 21:27 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] (duration: 34m 14s) * 21:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs1002 * 21:23 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs1002 * 21:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS trixie * 21:22 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 21:20 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 22s) * 21:20 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 21:16 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 21:16 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 21:15 ladsgroup@deploy2003: ladsgroup: Continuing with deployment * 21:13 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2002.codfw.wmnet with OS bookworm * 21:11 ladsgroup@deploy2003: ladsgroup: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:08 ladsgroup@cumin1003: END (PASS) - Cookbook sre.wikireplicas.update-views (exit_code=0) * 21:07 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 6 hosts with reason: reboots * 20:54 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecycle work - bking@cumin2003 * 20:53 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:53 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] * 20:51 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) * 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 20:47 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecycle work - bking@cumin2003 * 20:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 20:41 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:41 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) * 20:40 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host relforge1008.eqiad.wmnet * 20:40 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1009.eqiad.wmnet with OS trixie * 20:33 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:32 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:32 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:31 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:31 ladsgroup@cumin1003: END (PASS) - Cookbook sre.wikireplicas.update-views (exit_code=0) * 20:29 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1008.eqiad.wmnet * 20:24 rzl@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 20:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2002.codfw.wmnet with OS bookworm * 20:23 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host relforge1008.eqiad.wmnet * 20:23 rzl@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 20:23 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1008.eqiad.wmnet * 20:22 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:22 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:21 bking@cumin2003: END (ERROR) - Cookbook sre.elasticsearch.rolling-operation (exit_code=97) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:21 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1009.eqiad.wmnet with reason: host reimage * 20:16 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:15 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1009.eqiad.wmnet with reason: host reimage * 20:12 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) * 20:02 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 19:55 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1009.eqiad.wmnet with OS trixie * 19:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 19:43 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 19:30 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 19:28 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 19:27 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 19:25 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 18:42 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 18:41 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 18:16 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 18:15 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 17:45 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for doh5004.wikimedia.org * 17:45 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for doh5004.wikimedia.org * 17:38 ladsgroup@deploy2003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 17:35 ladsgroup@deploy2003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 17:29 ladsgroup@deploy2003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 17:26 ladsgroup@deploy2003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 17:09 mutante: zuul[12]00[123] - rebooting for maintenance * 17:09 ebernhardson: start full in-place reindex of eqiad cirrussearch cluster * 17:08 dzahn@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-cluster (exit_code=99) * 17:08 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-cluster * 17:03 ebernhardson: start full in-place reindex of codfw cirrussearch cluster * 16:59 mutante: stewards1001/stewards2001 - reboot for maintenance * 16:54 ebernhardson: start full in-place reindex of cloudelastic cluster * 16:53 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 16:52 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply * 16:49 mutante: planet1003/planet2003 - rebooting * 16:47 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on doh5004.wikimedia.org with reason: random high load, investigating * 15:55 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 15:54 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 15:51 jynus: restarting backupmon1001 * 15:49 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 14 hosts * 15:49 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 14 hosts * 15:47 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backupmon1001.eqiad.wmnet with reason: restart * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:06 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 15:06 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 14:59 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply * 14:58 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply * 14:51 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 14 hosts * 14:51 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 14 hosts * 14:49 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 6 hosts with reason: reboot & upgrade * 14:48 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet * 14:48 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet * 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:42 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1052.eqiad.wmnet * 14:42 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1052.eqiad.wmnet * 14:42 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1052.eqiad.wmnet * 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:31 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:31 elukey@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: sync * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:30 elukey@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: sync * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:28 elukey@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: sync * 14:28 elukey@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: sync * 14:26 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:20 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1052.eqiad.wmnet with OS trixie * 14:19 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 6 hosts with reason: reboot & upgrade * 14:18 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:15 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:15 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:13 elukey: update druid indexation job for webrequest_sampled_live - [[phab:T427068|T427068]] * 14:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:09 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:09 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for papaul - jhancock@cumin2002" * 14:09 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for papaul - jhancock@cumin2002" * 14:07 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:07 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:04 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 14:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cuminunpriv1001.eqiad.wmnet * 13:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb1003.eqiad.wmnet * 13:59 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1052.eqiad.wmnet with reason: host reimage * 13:57 moritzm: installing requests security updates * 13:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cuminunpriv1001.eqiad.wmnet * 13:55 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb1003.eqiad.wmnet * 13:53 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1052.eqiad.wmnet with reason: host reimage * 13:50 moritzm: installing python-cryptography security updates * 13:47 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb2003.codfw.wmnet * 13:44 Msz2001: UTC afternoon config+backport window is done * 13:44 Msz2001: Updated `logging` on `metawiki` to fix log performers, [[phab:T431176|T431176]]#12105297 * 13:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb2003.codfw.wmnet * 13:43 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 13:43 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt1002.wikimedia.org * 13:41 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] (duration: 07m 30s) * 13:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt1002.wikimedia.org * 13:37 mszwarc@deploy2003: mszwarc: Continuing with deployment * 13:36 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1052 * 13:36 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1052 * 13:35 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:35 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1052 * 13:35 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1052.eqiad.wmnet 47.32.64.10.in-addr.arpa 7.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:35 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1052.eqiad.wmnet 47.32.64.10.in-addr.arpa 7.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:35 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:35 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1052 - blake@cumin1003" * 13:35 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1052 - blake@cumin1003" * 13:34 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] * 13:31 blake@cumin1003: START - Cookbook sre.dns.netbox * 13:31 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1052 * 13:30 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1052.eqiad.wmnet with OS trixie * 13:30 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1052.eqiad.wmnet * 13:29 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1052.eqiad.wmnet * 13:29 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1052.eqiad.wmnet * 13:17 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] (duration: 11m 26s) * 13:13 jforrester@deploy2003: jforrester: Continuing with deployment * 13:08 jforrester@deploy2003: jforrester: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:06 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] * 12:54 cgoubert@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply * 12:54 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:52 cgoubert@deploy2003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply * 12:45 cgoubert@deploy2003: helmfile [codfw] DONE helmfile.d/services/mobileapps: apply * 12:44 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:44 cgoubert@deploy2003: helmfile [codfw] START helmfile.d/services/mobileapps: apply * 12:43 cgoubert@deploy2003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 12:43 cgoubert@deploy2003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 12:42 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast4006.wikimedia.org * 12:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt2002.wikimedia.org * 12:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast7002.wikimedia.org * 12:18 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast4006.wikimedia.org * 12:18 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host ml-serve1004 * 12:18 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host ml-serve1004 * 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt2002.wikimedia.org * 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast7002.wikimedia.org * 12:10 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backup[2003,2014].codfw.wmnet with reason: reboot & upgrade * 12:10 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt-staging2001.codfw.wmnet * 12:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid1003.eqiad.wmnet * 12:06 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt-staging2001.codfw.wmnet * 12:05 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid1003.eqiad.wmnet * 12:03 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backup[1003,1014].eqiad.wmnet with reason: reboot & upgrade * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid2003.codfw.wmnet * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host irc1003.wikimedia.org * 11:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid2003.codfw.wmnet * 11:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host irc1003.wikimedia.org * 11:55 jmm@dns1004: END - running authdns-update * 11:53 jmm@dns1004: START - running authdns-update * 11:50 jmm@dns1004: END - running authdns-update * 11:48 jmm@dns1004: START - running authdns-update * 11:27 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host irc2003.wikimedia.org * 11:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host irc2003.wikimedia.org * 11:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint2001.codfw.wmnet * 11:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint1001.eqiad.wmnet * 11:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint2001.codfw.wmnet * 11:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint1001.eqiad.wmnet * 11:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-rw2001.wikimedia.org * 11:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-rw1001.wikimedia.org * 11:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-rw2001.wikimedia.org * 11:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-rw1001.wikimedia.org * 11:03 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon1003.wikimedia.org * 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2005.codfw.wmnet * 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2005.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 10:59 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2005.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 10:57 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon1003.wikimedia.org * 10:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon2002.wikimedia.org * 10:55 jmm@cumin2003: START - Cookbook sre.dns.netbox * 10:51 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon2002.wikimedia.org * 10:51 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:50 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2005.codfw.wmnet * 10:41 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:40 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host ml-serve1003 * 10:40 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host ml-serve1003 * 10:39 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2033.codfw.wmnet * 10:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install2005.wikimedia.org * 10:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install1005.wikimedia.org * 10:35 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1004.eqiad.wmnet with OS bookworm * 10:31 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install1005.wikimedia.org * 10:31 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install2005.wikimedia.org * 10:30 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install4004.wikimedia.org * 10:30 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install3004.wikimedia.org * 10:29 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install3004.wikimedia.org * 10:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install4004.wikimedia.org * 10:23 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 10:21 moritzm: failover Ganeti master in codfw/routed to ganeti2034 [[phab:T430928|T430928]] * 10:19 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.addnode (exit_code=0) for new host ganeti2031.codfw.wmnet to cluster codfw and group B * 10:19 moritzm: readded ganeti2031 to the codfw Ganeti cluster [[phab:T430910|T430910]] * 10:18 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1004.eqiad.wmnet with reason: host reimage * 10:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install5004.wikimedia.org * 10:18 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1003 * 10:18 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1003 * 10:17 jmm@cumin2003: START - Cookbook sre.ganeti.addnode for new host ganeti2031.codfw.wmnet to cluster codfw and group B * 10:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install6003.wikimedia.org * 10:16 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install5004.wikimedia.org * 10:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install6003.wikimedia.org * 10:15 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1004.eqiad.wmnet with reason: host reimage * 10:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1001.eqiad.wmnet * 10:14 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 10:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2008.wikimedia.org * 10:00 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ml-serve1004 * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1004 * 09:57 jmm@cumin2003: START - Cookbook sre.dns.netbox * 09:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install7002.wikimedia.org * 09:57 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1004 * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ml-serve1004.eqiad.wmnet 50.48.64.10.in-addr.arpa 0.5.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:57 klausman@cumin1003: START - Cookbook sre.dns.wipe-cache ml-serve1004.eqiad.wmnet 50.48.64.10.in-addr.arpa 0.5.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1004 - klausman@cumin1003" * 09:56 klausman@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1004 - klausman@cumin1003" * 09:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-coord1001.eqiad.wmnet * 09:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 09:55 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow7002.magru.wmnet * 09:52 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-coord1001.eqiad.wmnet * 09:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 09:52 klausman@cumin1003: START - Cookbook sre.dns.netbox * 09:50 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install7002.wikimedia.org * 09:50 klausman@cumin1003: START - Cookbook sre.hosts.move-vlan for host ml-serve1004 * 09:50 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1004.eqiad.wmnet with OS bookworm * 09:50 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1003.eqiad.wmnet with OS bookworm * 09:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1001.eqiad.wmnet * 09:49 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 09:49 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2008.wikimedia.org * 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2007.codfw.wmnet * 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2007.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 09:49 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow7002.magru.wmnet * 09:49 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2007.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 09:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard1003.eqiad.wmnet * 09:39 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard2003.codfw.wmnet * 09:39 jmm@cumin2003: START - Cookbook sre.dns.netbox * 09:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard1003.eqiad.wmnet * 09:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor1003.eqiad.wmnet * 09:35 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard2003.codfw.wmnet * 09:34 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2007.codfw.wmnet * 09:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor1003.eqiad.wmnet * 09:33 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor-dev2001.codfw.wmnet * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor2003.codfw.wmnet * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sretest1006.eqiad.wmnet * 09:27 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 09:25 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor-dev2001.codfw.wmnet * 09:25 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor2003.codfw.wmnet * 09:23 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] (duration: 06m 27s) * 09:23 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2205: codfw rack B4 repool after maintenance * 09:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host sretest1006.eqiad.wmnet * 09:23 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2204: codfw rack B4 repool after maintenance * 09:19 urbanecm@deploy2003: urbanecm: Continuing with deployment * 09:19 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:18 jmm@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 6 hosts with reason: reboot * 09:17 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] * 09:08 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ml-serve1003 * 09:08 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1003 * 09:07 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1003 * 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ml-serve1003.eqiad.wmnet 81.32.64.10.in-addr.arpa 1.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:07 klausman@cumin1003: START - Cookbook sre.dns.wipe-cache ml-serve1003.eqiad.wmnet 81.32.64.10.in-addr.arpa 1.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1003 - klausman@cumin1003" * 09:06 klausman@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1003 - klausman@cumin1003" * 08:58 klausman@cumin1003: START - Cookbook sre.dns.netbox * 08:57 klausman@cumin1003: START - Cookbook sre.hosts.move-vlan for host ml-serve1003 * 08:57 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1003.eqiad.wmnet with OS bookworm * 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=0) rolling reimage on P<nowiki>{</nowiki>ml-serve1003.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet * 08:55 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet * 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1003.eqiad.wmnet with OS bookworm * 08:39 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 08:38 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool db2205: codfw rack B4 repool after maintenance * 08:37 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool db2204: codfw rack B4 repool after maintenance * 08:36 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 08:35 hashar@deploy2003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 08:32 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:32 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:31 hashar@deploy2003: Rolling back deployment * 08:26 moritzm: failover Ganeti master in codfw to ganeti2048 [[phab:T430928|T430928]] * 08:16 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1003.eqiad.wmnet with OS bookworm * 08:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2004.codfw.wmnet * 08:16 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet * 08:16 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet * 08:16 klausman@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on P<nowiki>{</nowiki>ml-serve1003.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 08:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2002.codfw.wmnet * 08:15 XioNoX: lsw1-b4-codfw> request system reboot - [[phab:T430910|T430910]] * 08:15 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b4-codfw,lsw1-b4-codfw IPv6,lsw1-b4-codfw.mgmt with reason: Switch maintenance * 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for codfw rack B4 * 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:10 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2004.codfw.wmnet * 08:10 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2002.codfw.wmnet * 08:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2205: codfw rack B4 depool for maintenance * 08:08 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool db2205: codfw rack B4 depool for maintenance * 08:08 jmm@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin2003.codfw.wmnet * 08:08 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2204: codfw rack B4 depool for maintenance * 08:08 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool db2204: codfw rack B4 depool for maintenance * 08:08 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 27 hosts with reason: codfw rack B4 depool for maintenance * 08:03 jmm@cumin2002: START - Cookbook sre.hosts.reboot-single for host cumin2003.codfw.wmnet * 07:56 ayounsi@cumin1003: START - Cookbook sre.network.depool-rack with action 'depool' for codfw rack B4 * 07:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1008.eqiad.wmnet with OS trixie * 07:49 wmde-fisch@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] (duration: 08m 36s) * 07:44 wmde-fisch@deploy2003: wmde-fisch: Continuing with deployment * 07:43 wmde-fisch@deploy2003: wmde-fisch: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:41 wmde-fisch@deploy2003: Started scap sync-world: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] * 07:35 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1008.eqiad.wmnet with reason: host reimage * 07:31 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1008.eqiad.wmnet with reason: host reimage * 07:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1008.eqiad.wmnet with OS trixie * 07:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 07:00 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 06:59 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1008.eqiad.wmnet with OS trixie * 06:57 Emperor: rebalance thanos swift rings after previous re-image of thanos-fe1004 to trixie * 06:47 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1008.eqiad.wmnet with OS trixie * 04:10 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 14 days, 0:00:00 on cp6008.drmrs.wmnet with reason: Hardware failure - [[phab:T431651|T431651]] * 03:55 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp6008.* * 03:29 ryankemper: [[phab:T431311|T431311]] Repooled eqiad cirrussearch clusters (`chi/omega/psi`) following completion of OpenSearch 2.19 migration * 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad * 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=eqiad * 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 31s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-08 == * 23:52 Amir1: ladsgroup@deploy2003:~$ mwscript-k8s --follow -- extensions/ORES/maintenance/PurgeScoreCache.php --wiki=simplewiki --model damaging --old ([[phab:T431159|T431159]]) * 23:46 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 23:46 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing PTR for 2001:df2:e500:fe08::1 - cmooney@cumin1003" * 23:46 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing PTR for 2001:df2:e500:fe08::1 - cmooney@cumin1003" * 23:40 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 23:16 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 23:15 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 22:42 rzl@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 22:40 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] (duration: 12m 55s) * 22:40 rzl@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 22:37 rzl@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 22:36 rzl@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 22:35 rzl@deploy2003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 22:34 urbanecm@deploy2003: urbanecm: Continuing with deployment * 22:33 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:33 rzl@deploy2003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 22:32 rzl@deploy2003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 22:30 rzl@deploy2003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 22:30 rzl@deploy2003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 22:29 rzl@deploy2003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 22:27 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] * 22:26 rzl@deploy2003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 22:22 rzl@deploy2003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 22:21 rzl@deploy2003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 22:19 rzl@deploy2003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 22:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 22:17 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 22:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 22:13 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 22:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 22:13 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 22:09 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 22:06 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 22:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1094.eqiad.wmnet with OS trixie * 22:01 urbanecm: Make https://test.wikipedia.org/w/index.php?title=MediaWiki:GrowthExperimentsSuggestedEdits.json&diff=prev&oldid=750552 with GrowthExperiments disabled (via mw-experimental), then run `\MediaWiki\MediaWikiServices::getInstance()->get('CommunityConfiguration.ProviderFactory')->newProvider('GrowthSuggestedEdits')->getStore()->invalidate()` ([[phab:T431625|T431625]]) * 21:56 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d2-codfw * 21:55 urbanecm@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 21:55 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d2-codfw * 21:55 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c4-codfw * 21:55 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c4-codfw * 21:55 urbanecm@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2002 * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2002 * 21:54 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2002 * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2002.codfw.wmnet 50.32.192.10.in-addr.arpa 0.5.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:54 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2002.codfw.wmnet 50.32.192.10.in-addr.arpa 0.5.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2002 - bking@cumin2003" * 21:54 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2002 - bking@cumin2003" * 21:49 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:49 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2002 * 21:49 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2002.codfw.wmnet with OS trixie * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1094.eqiad.wmnet with reason: host reimage * 21:42 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 21:39 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 21:37 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1094.eqiad.wmnet with reason: host reimage * 21:36 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 21:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 21:29 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 21:27 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 21:22 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1094.eqiad.wmnet with OS trixie * 21:21 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host restbase2039.codfw.wmnet with OS bullseye * 21:21 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin2002" * 21:21 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin2002" * 21:04 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on restbase2039.codfw.wmnet with reason: host reimage * 21:00 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on restbase2039.codfw.wmnet with reason: host reimage * 20:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1073.eqiad.wmnet with OS trixie * 20:48 mutante: deploy2003 - kill 1102 (stunnel4) ; systemctl start stunnel4 ([[phab:T418262|T418262]]) * 20:42 cjming@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] (duration: 33m 02s) * 20:42 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host restbase2039.codfw.wmnet with OS bullseye * 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1073.eqiad.wmnet with reason: host reimage * 20:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1073.eqiad.wmnet with reason: host reimage * 20:30 cjming@deploy2003: cjming: Continuing with deployment * 20:28 cjming@deploy2003: cjming: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1098.eqiad.wmnet with OS trixie * 20:13 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1073.eqiad.wmnet with OS trixie * 20:09 cjming@deploy2003: Started scap sync-world: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] * 20:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1098.eqiad.wmnet with reason: host reimage * 19:56 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1098.eqiad.wmnet with reason: host reimage * 19:55 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d4-codfw * 19:54 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d4-codfw * 19:54 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c1-codfw * 19:54 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c1-codfw * 19:52 mutante: restarting gerrit on gerrit.wikimedia.org (gerrit2003) * 19:48 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2331.codfw.wmnet * 19:48 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2331.codfw.wmnet * 19:48 mutante: restarting gerrit on gerrit-replica.wikimedia.org (gerrit1003) * 19:46 mutante: restarting gerrit on gerrit-spare.wikimedia.org (gerrit2002) * 19:43 jasmine@cumin2002: conftool action : set/pooled=yes; selector: name=wikikube-worker2331.codfw.wmnet,cluster=kubernetes,service=kubesvc * 19:43 jasmine@cumin2002: conftool action : set/weight=10; selector: name=wikikube-worker2331.codfw.wmnet,cluster=kubernetes,service=kubesvc * 19:40 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1098.eqiad.wmnet with OS trixie * 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d5-codfw * 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d5-codfw * 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c7-codfw * 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c7-codfw * 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c5-codfw * 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c5-codfw * 19:30 jasmine_: ran homer on lsw1-d8-codfw, adding wikikube-worker2331 to cluster * 19:29 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1100.eqiad.wmnet with OS trixie * 19:20 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d8-codfw * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d8-codfw * 19:19 mutante: gerrit - replacing private key for registerEmail verification * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-magru * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device cr2-magru * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d7-codfw * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d7-codfw * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d3-codfw * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d1-codfw * 19:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d1-codfw * 19:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c2-codfw * 19:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c2-codfw * 19:11 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-codfw * 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-magru * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device cr1-magru * 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d8-codfw * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d8-codfw * 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d6-codfw * 19:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1100.eqiad.wmnet with reason: host reimage * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d6-codfw * 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c6-codfw * 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c6-codfw * 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c3-codfw * 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c3-codfw * 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b4-magru * 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device asw1-b4-magru * 19:08 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b3-magru * 19:08 cmooney@cumin1003: START - Cookbook sre.network.tls for network device asw1-b3-magru * 19:05 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1100.eqiad.wmnet with reason: host reimage * 19:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1122.eqiad.wmnet with OS trixie * 19:00 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 18:59 topranks: rolling out update to BGP ACL on Nokia Switches eqiad, codfw & ulsfo [[phab:T425703|T425703]] * 18:58 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 18:57 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 18:55 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 18:53 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 18:52 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 18:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1100.eqiad.wmnet with OS trixie * 18:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1068.eqiad.wmnet with OS trixie * 18:47 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1102.eqiad.wmnet with OS trixie * 18:47 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 18:46 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1122.eqiad.wmnet with reason: host reimage * 18:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1122.eqiad.wmnet with reason: host reimage * 18:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1068.eqiad.wmnet with reason: host reimage * 18:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1122 * 18:26 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1122 * 18:25 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1122 * 18:25 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1122.eqiad.wmnet 31.48.64.10.in-addr.arpa 1.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:25 bking@cumin2003: START - Cookbook sre.dns.wipe-cache cirrussearch1122.eqiad.wmnet 31.48.64.10.in-addr.arpa 1.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:25 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:25 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1122 - bking@cumin2003" * 18:25 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1122 - bking@cumin2003" * 18:21 rzl@deploy2003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 18:21 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1068.eqiad.wmnet with reason: host reimage * 18:21 rzl@deploy2003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 18:21 rzl@deploy2003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 18:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 18:19 rzl@deploy2003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 18:19 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:18 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1122 * 18:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1122.eqiad.wmnet with OS trixie * 18:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 18:15 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 18:13 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 18:13 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 18:10 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 18:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1068.eqiad.wmnet with OS trixie * 18:01 kamila@deploy2003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 18m 29s) * 18:00 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:55 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:42 kamila@deploy2003: Started scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] * 17:42 kamila@deploy2003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 19m 50s) * 17:42 kamila@deploy2003: Rolling back deployment * 17:35 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:31 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet * 17:18 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet * 17:16 kamila@deploy1003: Unlocked for deployment [MediaWiki]: switching deployment server (duration: 22m 07s) * 17:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 17:11 kamila@dns1005: END - running authdns-update * 17:09 kamila@dns1005: START - running authdns-update * 17:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 17:04 jasmine@cumin2002: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1164.eqiad.wmnet * 17:04 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1164.eqiad.wmnet * 17:04 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1164.eqiad.wmnet * 16:56 kamila@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on releases2003.codfw.wmnet,releases1003.eqiad.wmnet with reason: Deployment server switchover * 16:54 kamila@deploy1003: Locking from deployment [MediaWiki]: switching deployment server * 16:53 kamila@deploy1003: Unlocked for deployment [MediaWiki]: switching deployment server (duration: 04m 02s) * 16:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie * 16:49 kamila@deploy1003: Locking from deployment [MediaWiki]: switching deployment server * 16:46 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1095.eqiad.wmnet with OS trixie * 16:45 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1093.eqiad.wmnet with OS trixie * 16:43 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1164.eqiad.wmnet with OS trixie * 16:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1095.eqiad.wmnet with reason: host reimage * 16:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 16:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 16:23 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1164.eqiad.wmnet with reason: host reimage * 16:18 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on cirrussearch1093.eqiad.wmnet with reason: host reimage * 16:16 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1164.eqiad.wmnet with reason: host reimage * 16:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1095.eqiad.wmnet with reason: host reimage * 16:09 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 16:09 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 16:08 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1093.eqiad.wmnet with reason: host reimage * 15:59 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Pool test * 15:59 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:59 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 15:59 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Pool test * 15:58 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Depool test * 15:58 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:58 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 15:58 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Depool test * 15:57 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1164 * 15:57 jasmine@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1164 * 15:57 jasmine@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1164 * 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1164.eqiad.wmnet 114.48.64.10.in-addr.arpa 4.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:56 jasmine@cumin2002: START - Cookbook sre.dns.wipe-cache wikikube-worker1164.eqiad.wmnet 114.48.64.10.in-addr.arpa 4.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1164 - jasmine@cumin2002" * 15:56 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1164 - jasmine@cumin2002" * 15:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1093.eqiad.wmnet with OS trixie * 15:51 jasmine@cumin2002: START - Cookbook sre.dns.netbox * 15:51 jasmine@cumin2002: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1164 * 15:50 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-worker1164.eqiad.wmnet with OS trixie * 15:50 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1164.eqiad.wmnet * 15:50 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1164.eqiad.wmnet * 15:50 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1164.eqiad.wmnet * 15:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1095.eqiad.wmnet with OS trixie * 15:42 jasmine@cumin2002: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1164.eqiad.wmnet * 15:42 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1164.eqiad.wmnet * 15:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:42 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1164.eqiad.wmnet * 15:42 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1164.eqiad.wmnet * 15:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 15:39 elukey@cumin1003: START - Cookbook sre.hosts.provision for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 15:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1007.eqiad.wmnet with OS trixie * 15:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Pool test * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1007.eqiad.wmnet with reason: host reimage * 15:15 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 15:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet * 15:15 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 15:15 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1007.eqiad.wmnet with reason: host reimage * 15:15 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:14 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Pool test * 15:14 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet * 15:14 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2228: Depool test * 15:14 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db2228: Depool test * 15:10 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 15:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:08 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 15:06 blake@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 15:06 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet * 15:06 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 15:06 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 15:05 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet * 15:05 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 15:04 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:04 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 15:04 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:03 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 15:03 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 15:03 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:03 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T430909|T430909]] * 15:03 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:03 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 15:03 swfrench-wmf: restarted eqsin, codfw confds - [[phab:T430909|T430909]] * 15:03 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test1002.eqiad.wmnet * 15:02 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet * 14:59 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:59 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:55 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1007.eqiad.wmnet with OS trixie * 14:52 swfrench-wmf: restarted ulsfo confds, confirmed now connected to codfw backends except those using wikimedia.org SRV record - [[phab:T430909|T430909]] * 14:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:49 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:41 moritzm: uninstalling dhcpcd-base from trixie hosts which still have it installed [[phab:T414341|T414341]] * 14:40 sukhe: sudo cumin -b1 -s120 "P<nowiki>{</nowiki>lvs2011*<nowiki>}</nowiki> or P<nowiki>{</nowiki>lvs2012*<nowiki>}</nowiki>" "systemctl restart pybal.service" * 14:39 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:39 mvernon@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host thanos-be1007.eqiad.wmnet with OS trixie * 14:37 sukhe: restart pybal on lvs2013 to revert back to conf2004 * 14:35 sukhe: restart pybal on lvs2014 to revert back to conf2004 * 14:34 swfrench-wmf: switched codfw, eqsin, ulsfo etcd client SRV records back to codfw - [[phab:T430909|T430909]] * 14:32 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1002.eqiad.wmnet * 14:32 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet * 14:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1007.eqiad.wmnet with OS trixie * 14:31 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:31 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:31 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:30 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:30 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Pool test * 14:30 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:29 swfrench@dns1004: END - running authdns-update * 14:29 moritzm: installing jackson-core security updates * 14:27 swfrench@dns1004: START - running authdns-update * 14:22 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:22 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:22 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1119.eqiad.wmnet with OS trixie * 14:22 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:21 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:20 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:20 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 14:20 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:19 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 14:19 moritzm: installing librabbitmq security updates * 14:19 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1002.eqiad.wmnet * 14:18 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet * 14:16 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:15 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:15 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:14 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Pool test * 14:14 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox) * 14:14 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet * 14:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 14:08 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 14:07 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet * 14:05 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1118.eqiad.wmnet with OS trixie * 14:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test1001.eqiad.wmnet * 14:00 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-worker@eqiad * 14:00 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 13:59 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 13:58 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 13:57 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1119.eqiad.wmnet with reason: host reimage * 13:54 moritzm: installing libcap2 security updates * 13:53 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1119.eqiad.wmnet with reason: host reimage * 13:52 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet * 13:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) * 13:52 fceratto@cumin1003: START - Cookbook sre.mysql.depool * 13:50 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-worker@eqiad * 13:50 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1051.eqiad.wmnet * 13:50 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1051.eqiad.wmnet * 13:50 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1051.eqiad.wmnet * 13:49 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 13:45 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1081.eqiad.wmnet with OS trixie * 13:41 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1119 * 13:41 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1119 * 13:40 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1119 * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1119.eqiad.wmnet 97.32.64.10.in-addr.arpa 7.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1119.eqiad.wmnet 97.32.64.10.in-addr.arpa 7.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1119 - atsuko@cumin1003" * 13:40 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1119 - atsuko@cumin1003" * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1118.eqiad.wmnet with reason: host reimage * 13:39 moritzm: installing krb5 security updates * 13:37 Lucas_WMDE: UTC afternoon backport+config window done * 13:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1006.eqiad.wmnet with OS trixie * 13:36 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1118.eqiad.wmnet with reason: host reimage * 13:36 atsuko@cumin1003: START - Cookbook sre.dns.netbox * 13:35 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] (duration: 07m 46s) * 13:34 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1119 * 13:34 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1119.eqiad.wmnet with OS trixie * 13:30 sbisson@deploy1003: sbisson: Continuing with deployment * 13:30 moritzm: installing openssh security updates * 13:30 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling restart_daemons on A:wikidough * 13:29 sbisson@deploy1003: sbisson: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:27 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] * 13:26 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1051.eqiad.wmnet with OS trixie * 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1081.eqiad.wmnet with reason: host reimage * 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1118 * 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1118 * 13:22 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] (duration: 12m 12s) * 13:21 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1081.eqiad.wmnet with reason: host reimage * 13:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1006.eqiad.wmnet with reason: host reimage * 13:18 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1118 * 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1118.eqiad.wmnet 90.32.64.10.in-addr.arpa 0.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:18 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1118.eqiad.wmnet 90.32.64.10.in-addr.arpa 0.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1118 - atsuko@cumin1003" * 13:18 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1118 - atsuko@cumin1003" * 13:17 stran@deploy1003: stran: Continuing with deployment * 13:16 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough * 13:15 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:13 atsuko@cumin1003: START - Cookbook sre.dns.netbox * 13:12 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1118 * 13:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1006.eqiad.wmnet with reason: host reimage * 13:12 stran@deploy1003: stran: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:12 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1118.eqiad.wmnet with OS trixie * 13:10 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] * 13:05 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-worker@codfw * 13:05 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 13:05 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1081.eqiad.wmnet with OS trixie * 13:05 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1051.eqiad.wmnet with reason: host reimage * 13:04 moritzm: installing jq security updates * 13:04 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 13:01 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1051.eqiad.wmnet with reason: host reimage * 12:58 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-worker@codfw * 12:52 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 12:50 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host thanos-be1006.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1051 * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1051 * 12:43 moritzm: installing Python 3.11 security updates * 12:43 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1051 * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1051.eqiad.wmnet 46.32.64.10.in-addr.arpa 6.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:43 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1051.eqiad.wmnet 46.32.64.10.in-addr.arpa 6.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1051 - blake@cumin1003" * 12:43 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1051 - blake@cumin1003" * 12:38 blake@cumin1003: START - Cookbook sre.dns.netbox * 12:38 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1051 * 12:38 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1051.eqiad.wmnet with OS trixie * 12:37 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1051.eqiad.wmnet * 12:36 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1051.eqiad.wmnet * 12:36 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1051.eqiad.wmnet * 12:34 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1006.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 12:34 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1006.eqiad.wmnet with OS trixie * 12:27 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 12:27 mvernon@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host thanos-be1006.eqiad.wmnet with OS trixie * 12:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:02 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 12:01 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1006.eqiad.wmnet with OS trixie * 11:43 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 11:38 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1076.eqiad.wmnet with OS trixie * 11:26 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1075.eqiad.wmnet with OS trixie * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2047.codfw.wmnet * 11:19 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2047.codfw.wmnet * 11:19 moritzm: temporarily remove ganeti2031 from codfw cluster [[phab:T430910|T430910]] * 11:08 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1076.eqiad.wmnet with reason: host reimage * 11:08 moritzm: installing Linux 6.1.176 on Bookworm servers * 11:03 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1076.eqiad.wmnet with reason: host reimage * 11:00 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1075.eqiad.wmnet with reason: host reimage * 10:56 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1075.eqiad.wmnet with reason: host reimage * 10:47 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1076.eqiad.wmnet with OS trixie * 10:46 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1074.eqiad.wmnet with OS trixie * 10:45 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1005.eqiad.wmnet with OS trixie * 10:40 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1075.eqiad.wmnet with OS trixie * 10:32 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2031.codfw.wmnet * 10:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1005.eqiad.wmnet with reason: host reimage * 10:25 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1005.eqiad.wmnet with reason: host reimage * 10:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1074.eqiad.wmnet with reason: host reimage * 10:17 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1074.eqiad.wmnet with reason: host reimage * 10:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1005.eqiad.wmnet with OS trixie * 10:12 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet * 10:04 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 10:01 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet * 10:01 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1074.eqiad.wmnet with OS trixie * 10:01 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 09:43 cgoubert@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/aux-k8s-services/redioscope: apply * 09:43 cgoubert@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/aux-k8s-services/redioscope: apply * 09:43 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply * 09:35 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply * 09:34 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 09:34 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 09:33 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 41 days, 15:00:00 on db2252.codfw.wmnet with reason: Test * 09:32 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 09:32 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 09:31 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: codfw rack B3 pool after maintenance * 09:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 09:07 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 09:07 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 09:02 ladsgroup@cumin1003: END (PASS) - Cookbook sre.mysql.sanitarium_restart (exit_code=0) * 08:57 topranks: merge patch to shift eqiad <-> esams traffic onto new 40G circuit * 08:54 hashar@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1004.eqiad.wmnet with OS trixie * 08:50 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 08:50 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitarium_restart (exit_code=99) * 08:50 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 08:45 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool es2051: codfw rack B3 pool after maintenance * 08:44 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:44 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:43 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2007.codfw.wmnet * 08:43 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2007.codfw.wmnet * 08:42 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2031.codfw.wmnet * 08:41 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2031.codfw.wmnet * 08:40 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2031.codfw.wmnet * 08:38 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:38 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:35 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.sanitize-wiki (exit_code=97) Managing sanitization for wikis minwikiquote in section s3 * 08:33 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis minwikiquote in section s3 * 08:32 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Checking sanitization for wikis minwikiquote in section s5 * 08:30 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Checking sanitization for wikis minwikiquote in section s5 * 08:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Managing sanitization for wikis minwikiquote in section s5 * 08:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1004.eqiad.wmnet with reason: host reimage * 08:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1004.eqiad.wmnet with reason: host reimage * 08:23 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:22 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis minwikiquote in section s5 * 08:19 XioNoX: lsw1-b3-codfw> request system reboot - [[phab:T430909|T430909]] * 08:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Checking sanitization for wikis minwikiquote in section s5 * 08:17 hashar@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Checking sanitization for wikis minwikiquote in section s5 * 08:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for codfw rack B3 * 08:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2007.codfw.wmnet * 08:15 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lsw1-b3-codfw,lsw1-b3-codfw IPv6,lsw1-b3-codfw.mgmt with reason: Switch maintenance * 08:15 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2007.codfw.wmnet * 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:07 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:06 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: codfw rack B3 depool for maintenance * 08:05 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool es2051: codfw rack B3 depool for maintenance * 08:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1004.eqiad.wmnet with OS trixie * 08:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1005.eqiad.wmnet with OS trixie * 08:03 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 21 hosts with reason: codfw rack B3 depool for maintenance * 07:56 ayounsi@cumin1003: START - Cookbook sre.network.depool-rack with action 'depool' for codfw rack B3 * 07:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1005.eqiad.wmnet with reason: host reimage * 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1005.eqiad.wmnet with reason: host reimage * 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1005.eqiad.wmnet with OS bookworm * 07:29 moritzm: installing gnutls28 security updates * 07:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1005.eqiad.wmnet with OS trixie * 07:13 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1125.eqiad.wmnet with OS trixie * 07:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1005.eqiad.wmnet with reason: host reimage * 07:07 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aux-k8s-etcd1005.eqiad.wmnet with reason: host reimage * 06:56 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1005.eqiad.wmnet with OS bookworm * 06:54 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1125.eqiad.wmnet with reason: host reimage * 06:52 elukey: upgrade all trixie hosts to pywmflib 3.1 - [[phab:T430552|T430552]] * 06:50 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1125.eqiad.wmnet with reason: host reimage * 06:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 06:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 06:38 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1125.eqiad.wmnet with OS trixie * 05:42 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1107.eqiad.wmnet with OS trixie * 05:35 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1124.eqiad.wmnet with OS trixie * 05:31 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1101.eqiad.wmnet with OS trixie * 05:21 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1107.eqiad.wmnet with reason: host reimage * 05:17 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1124.eqiad.wmnet with reason: host reimage * 05:13 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1107.eqiad.wmnet with reason: host reimage * 05:13 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1101.eqiad.wmnet with reason: host reimage * 05:11 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1124.eqiad.wmnet with reason: host reimage * 05:10 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1101.eqiad.wmnet with reason: host reimage * 04:58 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1124.eqiad.wmnet with OS trixie * 04:56 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1107.eqiad.wmnet with OS trixie * 04:55 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1101.eqiad.wmnet with OS trixie * 02:27 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] (duration: 08m 14s) * 02:22 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 02:21 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 02:19 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] * 01:59 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] (duration: 09m 46s) * 01:55 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 01:51 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 01:49 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] * 01:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1099.eqiad.wmnet with OS trixie * 00:57 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1110.eqiad.wmnet with OS trixie * 00:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1099.eqiad.wmnet with reason: host reimage * 00:41 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1099.eqiad.wmnet with reason: host reimage * 00:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1110.eqiad.wmnet with reason: host reimage * 00:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1110.eqiad.wmnet with reason: host reimage * 00:26 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1099.eqiad.wmnet with OS trixie * 00:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1110.eqiad.wmnet with OS trixie == 2026-07-07 == * 22:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1097.eqiad.wmnet with OS trixie * 22:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1097.eqiad.wmnet with reason: host reimage * 22:24 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1097.eqiad.wmnet with reason: host reimage * 22:14 hashar: Restarting Gerrit on gerrit2002 and gerrit1003 (replicas) * 22:09 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1097.eqiad.wmnet with OS trixie * 22:07 hashar: Restarting Gerrit on gerrit2003 * 21:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 21:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 21:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 21:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 21:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1108.eqiad.wmnet with OS trixie * 20:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1091.eqiad.wmnet with OS trixie * 20:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1090.eqiad.wmnet with OS trixie * 20:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1108.eqiad.wmnet with reason: host reimage * 20:36 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1091.eqiad.wmnet with reason: host reimage * 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1090.eqiad.wmnet with reason: host reimage * 20:33 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1091.eqiad.wmnet with reason: host reimage * 20:30 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1108.eqiad.wmnet with reason: host reimage * 20:30 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1006.eqiad.wmnet * 20:30 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1090.eqiad.wmnet with reason: host reimage * 20:30 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1006.eqiad.wmnet * 20:27 jasmine_: "homer lsw1-c2-eqiad* commit "Added new stacked control plane wikikube-ctrl1006"" * 20:22 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] (duration: 07m 29s) * 20:20 jasmine_: "homer "cr*eqiad*" commit "Added new stacked control plane wikikube-ctrl1006"" * 20:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1091.eqiad.wmnet with OS trixie * 20:17 arlolra@deploy1003: arlolra: Continuing with deployment * 20:16 arlolra@deploy1003: arlolra: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:16 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1090.eqiad.wmnet with OS trixie * 20:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1108.eqiad.wmnet with OS trixie * 20:14 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] * 20:09 cwhite: remove 2026-04 swift log archives from centrallog2002 to free some space * 20:01 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=93) for host cirrussearch1108.eqiad.wmnet with OS trixie * 19:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1108.eqiad.wmnet with OS trixie * 19:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1090.eqiad.wmnet with OS trixie * 19:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1109.eqiad.wmnet with OS trixie * 19:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1092.eqiad.wmnet with OS trixie * 19:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1123.eqiad.wmnet with OS trixie * 19:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1109.eqiad.wmnet with reason: host reimage * 19:19 cdobbins@cumin2002: conftool action : set/pooled=yes; selector: name=dns7002.* * 19:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1092.eqiad.wmnet with reason: host reimage * 19:17 jasmine@dns1004: END - running authdns-update * 19:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1109.eqiad.wmnet with reason: host reimage * 19:15 jasmine@dns1004: START - running authdns-update * 19:15 cdobbins@dns1004: END - running authdns-update * 19:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1123.eqiad.wmnet with reason: host reimage * 19:13 cdobbins@dns1004: START - running authdns-update * 19:12 cdobbins@cumin2002: conftool action : set/pooled=yes; selector: name=dns7002.*,service=authdns-update * 19:11 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1092.eqiad.wmnet with reason: host reimage * 19:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1123.eqiad.wmnet with reason: host reimage * 18:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1123.eqiad.wmnet with OS trixie * 18:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1109.eqiad.wmnet with OS trixie * 18:56 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1092.eqiad.wmnet with OS trixie * 18:52 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 18:49 swfrench@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 18:40 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 18:38 swfrench@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 18:11 swfrench-wmf: restarted eqsin, codfw confds - [[phab:T430909|T430909]] * 18:01 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T430909|T430909]] * 17:59 swfrench-wmf: restarted ulsfo confds, confirmed now connected to eqiad backends - [[phab:T430909|T430909]] * 17:52 sukhe: restart pybal on lvs2011 to switch from conf2004 to conf1008: [[phab:T430909|T430909]] * 17:51 sukhe: restart pybal on lvs2012 to switch from conf2004 to conf1008 [puppet re-enabled there]: [[phab:T430909|T430909]] * 17:46 sukhe: restart pybal on lvs2013 to switch from conf2004 to conf1008: [[phab:T430909|T430909]] * 17:44 swfrench-wmf: switched codfw, eqsin, ulsfo etcd client SRV records to eqiad - [[phab:T430909|T430909]] * 17:43 swfrench@dns1004: END - running authdns-update * 17:40 swfrench@dns1004: START - running authdns-update * 17:40 sukhe: restart pybal on lvs2014 to switch from conf2004 to conf1008: [[phab:T430909|T430909]] * 17:21 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1003.eqiad.wmnet * 17:15 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1003.eqiad.wmnet * 17:14 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1002.eqiad.wmnet * 17:06 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1002.eqiad.wmnet * 17:06 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-low-traffic-codfw' 'systemctl restart pybal.service' # lvs2013, [[phab:T416623|T416623]] * 17:04 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1001.eqiad.wmnet * 17:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1111.eqiad.wmnet with OS trixie * 17:00 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal.service' # lvs2014, [[phab:T416623|T416623]] * 16:58 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1001.eqiad.wmnet * 16:58 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS bookworm * 16:55 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-low-traffic-eqiad' 'systemctl restart pybal.service' # lvs1019, [[phab:T416623|T416623]] * 16:53 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-secondary-eqiad' 'systemctl restart pybal.service' # lvs1020, [[phab:T416623|T416623]] * 16:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1111.eqiad.wmnet with reason: host reimage * 16:40 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1111.eqiad.wmnet with reason: host reimage * 16:38 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.peering (exit_code=99) with action 'configure' for AS: 47794 * 16:35 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 47794 * 16:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1111.eqiad.wmnet with OS trixie * 16:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1006.eqiad.wmnet with OS trixie * 16:06 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 16:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1006.eqiad.wmnet with reason: host reimage * 15:58 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1121.eqiad.wmnet with OS trixie * 15:58 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 15:56 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1006.eqiad.wmnet with reason: host reimage * 15:54 mutante: jenkins down in planned maintenance window * 15:42 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1037.eqiad.wmnet * 15:42 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1037.eqiad.wmnet * 15:42 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1037.eqiad.wmnet * 15:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1006.eqiad.wmnet with OS trixie * 15:34 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1121.eqiad.wmnet with reason: host reimage * 15:33 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS bookworm * 15:33 cdobbins@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host dns7002.wikimedia.org with OS trixie * 15:30 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1121.eqiad.wmnet with reason: host reimage * 15:29 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host clouddumps1001.wikimedia.org * 15:20 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1001.wikimedia.org * 15:18 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1121 * 15:18 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1121 * 15:18 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host clouddumps1002.wikimedia.org * 15:17 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1121 * 15:17 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:17 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply * 15:16 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply * 15:16 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:16 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1121 - atsuko@cumin1003" * 15:16 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1121 - atsuko@cumin1003" * 15:14 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1037.eqiad.wmnet with OS trixie * 15:11 atsuko@cumin1003: START - Cookbook sre.dns.netbox * 15:09 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org * 15:09 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1121 * 15:09 andrew@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host clouddumps1002.wikimedia.org * 15:09 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org * 15:09 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1121.eqiad.wmnet with OS trixie * 15:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1007.eqiad.wmnet with OS trixie * 15:08 andrew@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host clouddumps1002.wikimedia.org * 15:08 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org * 15:05 brennen@deploy1003: Finished deploy [phabricator/deployment@7e02037]: deploy phab1004 for [[phab:T431440|T431440]] (duration: 00m 47s) * 15:04 brennen@deploy1003: Started deploy [phabricator/deployment@7e02037]: deploy phab1004 for [[phab:T431440|T431440]] * 15:03 brennen@deploy1003: Finished deploy [phabricator/deployment@7e02037]: deploy phab2003 for [[phab:T431440|T431440]] (duration: 00m 51s) * 15:03 brennen@deploy1003: Started deploy [phabricator/deployment@7e02037]: deploy phab2003 for [[phab:T431440|T431440]] * 15:00 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71] (thin): Regular analytics weekly train THIN [analytics/refinery@7d8dc71f] (duration: 02m 10s) * 14:58 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71] (thin): Regular analytics weekly train THIN [analytics/refinery@7d8dc71f] * 14:58 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71]: Regular analytics weekly train [analytics/refinery@7d8dc71f] (duration: 04m 14s) * 14:54 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1037.eqiad.wmnet with reason: host reimage * 14:53 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71]: Regular analytics weekly train [analytics/refinery@7d8dc71f] * 14:53 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@7d8dc71f] (duration: 02m 00s) * 14:51 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@7d8dc71f] * 14:51 arnaudb@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on phab2003.codfw.wmnet,phab[1004-1006].eqiad.wmnet with reason: maintenance * 14:51 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1037.eqiad.wmnet with reason: host reimage * 14:50 JavierMonton: Deploying Refinery at {{Gerrit|7d8dc71f}} for change {{Gerrit|1308087}} / [[phab:T431318|T431318]] - update filerevision table sqoop and table * 14:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1007.eqiad.wmnet with reason: host reimage * 14:42 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1007.eqiad.wmnet with reason: host reimage * 14:40 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1083.eqiad.wmnet with OS trixie * 14:37 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:36 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:35 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-master@eqiad * 14:35 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 14:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 14:34 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1037 * 14:34 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1037 * 14:34 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:cleanMentorList.php --wiki=frwiki # [[phab:T427386|T427386]] * 14:34 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 14:34 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308112{{!}}Revert^2 "[Growth] frwiki: Deploy automated mentor list cleaner" (T427386)]] (duration: 06m 47s) * 14:34 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 14:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:33 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:32 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1037 * 14:31 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:31 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:29 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:29 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-master@eqiad * 14:29 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:29 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:29 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:28 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:27 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:27 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1308112{{!}}Revert^2 "[Growth] frwiki: Deploy automated mentor list cleaner" (T427386)]] * 14:26 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1007.eqiad.wmnet with OS trixie * 14:26 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:26 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 14:26 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:25 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:25 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:25 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:cleanMentorList.php --wiki=frwiki # [[phab:T427386|T427386]] * 14:24 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:24 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1037 - blake@cumin1003" * 14:24 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1037 - blake@cumin1003" * 14:20 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1083.eqiad.wmnet with reason: host reimage * 14:19 blake@cumin1003: START - Cookbook sre.dns.netbox * 14:19 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-master@codfw * 14:19 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 14:19 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1037 * 14:18 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1037.eqiad.wmnet with OS trixie * 14:18 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1037.eqiad.wmnet * 14:18 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 14:18 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1037.eqiad.wmnet * 14:18 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1037.eqiad.wmnet * 14:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2007.codfw.wmnet with OS trixie * 14:16 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1083.eqiad.wmnet with reason: host reimage * 14:15 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1036.eqiad.wmnet * 14:15 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1036.eqiad.wmnet * 14:14 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1036.eqiad.wmnet * 14:12 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-master@codfw * 14:11 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1120.eqiad.wmnet with OS trixie * 14:05 moritzm: installing distro-info-data updates from trixie/bookworm point releases * 14:04 fabfur: disable puppet on A:cp-text to selectively apply https://gerrit.wikimedia.org/r/c/operations/puppet/+/1308040 * 14:03 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] (duration: 27m 48s) * 14:00 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 14:00 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1083.eqiad.wmnet with OS trixie * 13:58 urbanecm@deploy1003: urbanecm: Continuing with deployment * 13:58 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:58 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1004.eqiad.wmnet with OS bookworm * 13:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2007.codfw.wmnet with reason: host reimage * 13:57 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 13:53 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1120.eqiad.wmnet with reason: host reimage * 13:50 moritzm: installing Linux 5.10.259 on Bullseye hosts * 13:47 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply * 13:47 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply * 13:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2007.codfw.wmnet with reason: host reimage * 13:46 cgoubert@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/aux-k8s-services/redioscope: apply * 13:46 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1120.eqiad.wmnet with reason: host reimage * 13:46 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:46 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:45 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:44 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:44 cgoubert@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/aux-k8s-services/redioscope: apply * 13:40 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 13:39 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:38 moritzm: installing e2fsprogs updates from Trixie point release * 13:35 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] * 13:33 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1120.eqiad.wmnet with OS trixie * 13:33 topranks: reset cr3-eqsin configuration so traffic uses it again after upgrade * 13:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1088.eqiad.wmnet with OS trixie * 13:32 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie * 13:32 cdobbins@cumin1003: conftool action : set/pooled=no; selector: name=dns7002.* * 13:29 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2007.codfw.wmnet with OS trixie * 13:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1004.eqiad.wmnet with reason: host reimage * 13:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2006.codfw.wmnet with OS trixie * 13:18 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1036.eqiad.wmnet with OS trixie * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aux-k8s-etcd1004.eqiad.wmnet with reason: host reimage * 13:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 13:16 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 13:15 jayme: Istio is being upgraded from 1.24.2 to 1.29.4 on wikikube staging eqiad and codfw - [[phab:T427401|T427401]] * 13:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1087.eqiad.wmnet with OS trixie * 13:14 topranks: reboot cr3-eqsin to install new JunOS and set PIC 0/0/0 to 100G * 13:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1088.eqiad.wmnet with reason: host reimage * 13:13 jmm@dns1004: END - running authdns-update * 13:12 jmm@dns1004: START - running authdns-update * 13:09 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1088.eqiad.wmnet with reason: host reimage * 13:07 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1082.eqiad.wmnet with OS trixie * 13:07 jmm@dns1004: END - running authdns-update * 13:06 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1004.eqiad.wmnet with OS bookworm * 13:05 jmm@dns1004: START - running authdns-update * 13:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2006.codfw.wmnet with reason: host reimage * 12:58 topranks: load updated JunOS on cr3-eqsin [[phab:T429386|T429386]] * 12:58 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1036.eqiad.wmnet with reason: host reimage * 12:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2001.codfw.wmnet * 12:57 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2006.codfw.wmnet with reason: host reimage * 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr1-codfw,cr[2-3]-eqsin,cr3-eqsin IPv6,cr3-eqsin.mgmt with reason: upgrade JunOS cr3-eqsin * 12:56 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lvs[5004-5006].eqsin.wmnet with reason: upgrade JunOS cr3-eqsin * 12:55 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 12:55 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 12:53 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1087.eqiad.wmnet with reason: host reimage * 12:53 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 12:52 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1088.eqiad.wmnet with OS trixie * 12:52 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1002.eqiad.wmnet * 12:52 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:51 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2001.codfw.wmnet * 12:49 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1036.eqiad.wmnet with reason: host reimage * 12:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 12:48 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1087.eqiad.wmnet with reason: host reimage * 12:44 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1082.eqiad.wmnet with reason: host reimage * 12:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1002.eqiad.wmnet * 12:42 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 12:42 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 12:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 12:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 12:39 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:39 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: move dumps-nfs IP to the shared one - filippo@cumin1003" * 12:39 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: move dumps-nfs IP to the shared one - filippo@cumin1003" * 12:39 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2006.codfw.wmnet with OS trixie * 12:38 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1082.eqiad.wmnet with reason: host reimage * 12:36 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:33 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:32 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1036 * 12:32 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1036 * 12:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2005.codfw.wmnet with OS trixie * 12:32 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1087.eqiad.wmnet with OS trixie * 12:30 jmm@dns1004: END - running authdns-update * 12:29 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1036 * 12:29 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1036.eqiad.wmnet 21.32.64.10.in-addr.arpa 1.2.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:29 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1036.eqiad.wmnet 21.32.64.10.in-addr.arpa 1.2.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:29 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:29 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1036 - blake@cumin1003" * 12:29 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1036 - blake@cumin1003" * 12:28 jmm@dns1004: START - running authdns-update * 12:26 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:26 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:23 blake@cumin1003: START - Cookbook sre.dns.netbox * 12:23 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1036 * 12:23 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1036.eqiad.wmnet with OS trixie * 12:22 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1036.eqiad.wmnet * 12:22 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1082.eqiad.wmnet with OS trixie * 12:22 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1036.eqiad.wmnet * 12:22 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1036.eqiad.wmnet * 12:21 marostegui: Restart mariadb@s7 on db1155 to pick up new filters - [[phab:T431124|T431124]] * 12:21 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 21 hosts with reason: restarting for replication filter * 12:20 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:19 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:14 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2005.codfw.wmnet with reason: host reimage * 12:14 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:08 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:07 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:07 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2005.codfw.wmnet with reason: host reimage * 12:06 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:06 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:06 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:05 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:05 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:04 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-master-eqiad * 12:04 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl1002.eqiad.wmnet * 12:04 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl1002.eqiad.wmnet * 12:04 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:04 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:03 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:03 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:03 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:03 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 11:59 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl1002.eqiad.wmnet * 11:59 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl1002.eqiad.wmnet * 11:59 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl1001.eqiad.wmnet * 11:59 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl1001.eqiad.wmnet * 11:56 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl1001.eqiad.wmnet * 11:56 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl1001.eqiad.wmnet * 11:56 klausman@cumin2002: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-master-eqiad * 11:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2005.codfw.wmnet with OS trixie * 11:49 blake@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on wikikube-worker1160.eqiad.wmnet with reason: Verifying matchers for silence * 11:42 topranks: cr3-eqsin, begin traffic drain to reset PIC and upgrade JunOS [[phab:T429386|T429386]] * 11:41 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs[5004-5006].eqsin.wmnet with reason: upgrade JunOS cr3-eqsin * 11:39 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr1-codfw,cr[2-3]-eqsin,cr3-eqsin IPv6,cr3-eqsin.mgmt with reason: upgrade JunOS cr3-eqsin * 11:36 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=thanos-fe2004.codfw.wmnet * 11:35 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1086.eqiad.wmnet with OS trixie * 11:35 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=thanos-fe2004.codfw.wmnet * 11:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1085.eqiad.wmnet with OS trixie * 11:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1086.eqiad.wmnet with reason: host reimage * 11:10 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1085.eqiad.wmnet with reason: host reimage * 11:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2004.codfw.wmnet with OS trixie * 11:03 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1086.eqiad.wmnet with reason: host reimage * 11:02 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1085.eqiad.wmnet with reason: host reimage * 10:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2004.codfw.wmnet with reason: host reimage * 10:48 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:46 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1086.eqiad.wmnet with OS trixie * 10:46 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1085.eqiad.wmnet with OS trixie * 10:44 cgoubert@deploy1003: Finished deploy [restbase/deploy@2fc37d4]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] (duration: 16m 44s) * 10:43 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2004.codfw.wmnet with reason: host reimage * 10:35 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:27 cgoubert@deploy1003: Started deploy [restbase/deploy@2fc37d4]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] * 10:27 cgoubert@deploy1003: Finished deploy [restbase/deploy@8a25036]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] (duration: 00m 45s) * 10:26 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1117.eqiad.wmnet with OS trixie * 10:26 cgoubert@deploy1003: Started deploy [restbase/deploy@8a25036]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] * 10:26 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host thanos-fe2004 * 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host thanos-fe2004 * 10:22 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1116.eqiad.wmnet with OS trixie * 10:21 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host thanos-fe2004 * 10:21 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) thanos-fe2004.codfw.wmnet 157.32.192.10.in-addr.arpa 7.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:20 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache thanos-fe2004.codfw.wmnet 157.32.192.10.in-addr.arpa 7.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:20 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:20 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host thanos-fe2004 - mvernon@cumin2003" * 10:20 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host thanos-fe2004 - mvernon@cumin2003" * 10:15 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2252: Repooling after reboot * 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:15 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 10:15 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2252: Repooling after reboot * 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1153.eqiad.wmnet * 10:14 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1153.eqiad.wmnet * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 10:14 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 10:12 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 10:12 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host thanos-fe2004 * 10:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2004.codfw.wmnet with OS trixie * 10:07 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1117.eqiad.wmnet with reason: host reimage * 10:03 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1116.eqiad.wmnet with reason: host reimage * 09:58 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:58 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1117.eqiad.wmnet with reason: host reimage * 09:57 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1116.eqiad.wmnet with reason: host reimage * 09:49 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 41 days, 15:00:00 on db2252.codfw.wmnet with reason: Security updates * 09:45 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1117.eqiad.wmnet with OS trixie * 09:45 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1116.eqiad.wmnet with OS trixie * 09:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1153: Security updates * 09:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:28 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:28 root@cumin1003: START - Cookbook sre.mysql.depool depool db1153: Security updates * 09:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1016: Security updates * 09:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:21 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:21 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1016: Security updates * 09:14 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:14 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 08:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1016: Security updates * 08:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:56 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:56 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1016: Security updates * 08:50 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:50 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:45 filippo@dns1006: END - running authdns-update * 08:43 filippo@dns1006: START - running authdns-update * 08:42 godog: switch dumps-nfs address to be shared with rsync/http - [[phab:T411248|T411248]] * 08:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1016: Security updates * 08:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:40 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:40 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1016: Security updates * 08:29 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host cirrussearch1111.eqiad.wmnet * 08:29 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:27 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:27 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:25 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1015: Security updates * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:09 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:09 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1015: Security updates * 07:42 Msz2001: Deployed private patch for Suggested Ivestigations * 07:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1015: Security updates * 07:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:41 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:41 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1015: Security updates * 07:40 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 07:11 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fingerprint warnings - oblivian@cumin1003" * 07:11 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fingerprint warnings - oblivian@cumin1003 * 07:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1024: Security updates * 07:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:11 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:11 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1024: Security updates * 07:10 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fingerprint warnings - oblivian@cumin1003 * 07:10 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fingerprint warnings - oblivian@cumin1003" * 06:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host cirrussearch1111.eqiad.wmnet * 06:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 06:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1024: Security updates * 06:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 06:48 root@cumin1003: START - Cookbook sre.mysql.parsercache * 06:48 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1024: Security updates * 06:42 moritzm: install nginx security updates * 06:31 root@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool pc1024: Security updates * 06:21 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1024: Security updates * 06:19 moritzm: installing php8.2 security updates * 06:15 moritzm: installing php8.4 security updates * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.7 (duration: 02m 38s) * 03:40 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] (duration: 37m 04s) * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 51s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-06 == * 23:30 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] (duration: 09m 39s) * 23:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1078.eqiad.wmnet with OS trixie * 23:26 jdlrobson@deploy1003: jdlrobson, bwang: Continuing with deployment * 23:22 jdlrobson@deploy1003: jdlrobson, bwang: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug) * 23:21 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] * 23:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1078.eqiad.wmnet with reason: host reimage * 23:06 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1078.eqiad.wmnet with reason: host reimage * 22:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1078.eqiad.wmnet with OS trixie * 22:29 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on cirrussearch1114.eqiad.wmnet with reason: reimage on hold until restore completes * 22:22 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on cirrussearch[1079,1115].eqiad.wmnet with reason: reimage on hold until restore completes * 21:18 maryum: Deployed security fix for [[phab:T428006|T428006]] * 20:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1079.eqiad.wmnet with OS trixie * 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1077.eqiad.wmnet with OS trixie * 20:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1115.eqiad.wmnet with OS trixie * 20:25 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1079.eqiad.wmnet with reason: host reimage * 20:21 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1079.eqiad.wmnet with reason: host reimage * 20:15 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] (duration: 08m 14s) * 20:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1077.eqiad.wmnet with reason: host reimage * 20:10 krinkle@deploy1003: krinkle, pushpaktiwari: Continuing with deployment * 20:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1115.eqiad.wmnet with reason: host reimage * 20:08 krinkle@deploy1003: krinkle, pushpaktiwari: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1077.eqiad.wmnet with reason: host reimage * 20:06 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] * 20:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1079.eqiad.wmnet with OS trixie * 20:04 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1115.eqiad.wmnet with reason: host reimage * 19:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1077.eqiad.wmnet with OS trixie * 19:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1115.eqiad.wmnet with OS trixie * 19:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 19:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 18:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1114.eqiad.wmnet with OS trixie * 18:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1114.eqiad.wmnet with reason: host reimage * 18:35 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1114.eqiad.wmnet with reason: host reimage * 18:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1112.eqiad.wmnet with OS trixie * 18:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1114.eqiad.wmnet with OS trixie * 18:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1072.eqiad.wmnet with OS trixie * 18:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1112.eqiad.wmnet with reason: host reimage * 18:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1112.eqiad.wmnet with reason: host reimage * 17:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1072.eqiad.wmnet with reason: host reimage * 17:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1112.eqiad.wmnet with OS trixie * 17:55 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1072.eqiad.wmnet with reason: host reimage * 17:39 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1072.eqiad.wmnet with OS trixie * 17:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1071.eqiad.wmnet with OS trixie * 17:18 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1070.eqiad.wmnet with OS trixie * 17:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1084.eqiad.wmnet with OS trixie * 16:54 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1071.eqiad.wmnet with reason: host reimage * 16:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1084.eqiad.wmnet with reason: host reimage * 16:51 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1070.eqiad.wmnet with reason: host reimage * 16:49 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1084.eqiad.wmnet with reason: host reimage * 16:38 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1071.eqiad.wmnet with OS trixie * 16:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1096.eqiad.wmnet with OS trixie * 16:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1070.eqiad.wmnet with OS trixie * 16:33 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1084.eqiad.wmnet with OS trixie * 16:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1089.eqiad.wmnet with OS trixie * 16:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1103.eqiad.wmnet with OS trixie * 16:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1096.eqiad.wmnet with reason: host reimage * 16:14 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1096.eqiad.wmnet with reason: host reimage * 16:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1089.eqiad.wmnet with reason: host reimage * 16:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1103.eqiad.wmnet with reason: host reimage * 16:02 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1003.eqiad.wmnet with OS bookworm * 16:01 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1089.eqiad.wmnet with reason: host reimage * 16:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1103.eqiad.wmnet with reason: host reimage * 15:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1096.eqiad.wmnet with OS trixie * 15:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1080.eqiad.wmnet with OS trixie * 15:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1089.eqiad.wmnet with OS trixie * 15:45 dancy@deploy1003: Installation of scap version "4.272.0" completed for 158 hosts * 15:43 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1103.eqiad.wmnet with OS trixie * 15:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1113.eqiad.wmnet with OS trixie * 15:41 dancy@deploy1003: Installing scap version "4.272.0" for 158 host(s) * 15:40 klausman@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 15:39 klausman@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 15:38 klausman@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 15:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1069.eqiad.wmnet with OS trixie * 15:37 klausman@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 15:36 klausman@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 15:34 klausman@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 15:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1080.eqiad.wmnet with reason: host reimage * 15:27 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1080.eqiad.wmnet with reason: host reimage * 15:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1113.eqiad.wmnet with reason: host reimage * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1069.eqiad.wmnet with reason: host reimage * 15:18 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1113.eqiad.wmnet with reason: host reimage * 15:16 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1069.eqiad.wmnet with reason: host reimage * 15:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:11 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1080.eqiad.wmnet with OS trixie * 15:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1113.eqiad.wmnet with OS trixie * 15:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1003.eqiad.wmnet with reason: host reimage * 14:47 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1003.eqiad.wmnet with OS bookworm * 14:33 elukey: rolled out spicerack on all cumin nodes - [[phab:T429699|T429699]] * 14:32 elukey: upgrade all bookworm hosts to pywmflib 3.1 - [[phab:T430552|T430552]] * 14:14 marostegui: Setup x4 eqiad topology [[phab:T404715|T404715]] * 14:13 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 14:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2230.codfw.wmnet * 14:07 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2230.codfw.wmnet * 13:59 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[2001-2002].codfw.wmnet * 13:51 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 13:45 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.major-upgrade (exit_code=97) * 13:45 cwilliams@cumin1003: dbmaint on s4@codfw [[phab:T429893|T429893]] * 13:45 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 13:42 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-master-codfw * 13:42 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl2002.codfw.wmnet * 13:42 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl2002.codfw.wmnet * 13:38 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl2002.codfw.wmnet * 13:38 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl2002.codfw.wmnet * 13:38 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl2001.codfw.wmnet * 13:38 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl2001.codfw.wmnet * 13:35 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl2001.codfw.wmnet * 13:35 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl2001.codfw.wmnet * 13:35 klausman@cumin2002: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-master-codfw * 12:30 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] (duration: 25m 11s) * 12:24 krinkle@deploy1003: krinkle: Continuing with deployment * 12:10 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2048.codfw.wmnet * 12:09 krinkle@deploy1003: krinkle: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:08 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2048.codfw.wmnet * 12:05 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] * 11:57 moritzm: installing curl security updates * 11:49 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:31 moritzm: installing nano security updates * 11:07 moritzm: failover Ganeti master in codfw to ganeti2032 [[phab:T430909|T430909]] * 11:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:04 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest1005.eqiad.wmnet with OS trixie * 11:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:50 jmm@dns1004: END - running authdns-update * 10:47 jmm@dns1004: START - running authdns-update * 10:47 jmm@dns1004: START - running authdns-update * 10:46 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:44 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest1005.eqiad.wmnet with reason: host reimage * 10:38 elukey@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest1005.eqiad.wmnet with reason: host reimage * 10:31 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:31 marostegui: Setup x4 codfw topology [[phab:T404715|T404715]] * 10:31 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 10:24 elukey: spicerack 13.0.0 deployed on cumin2002 * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 10:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 10:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 10:21 elukey@cumin2002: START - Cookbook sre.hosts.reimage for host sretest1005.eqiad.wmnet with OS trixie * 10:20 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:19 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:17 elukey: uploaded spicerack_13.0.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia * 09:54 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:52 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:20 elukey: upgrade all bullseye hosts to pywmflib 3.1 - [[phab:T430552|T430552]] * 09:10 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1015.eqiad.wmnet,service=s4 * 09:10 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1015.eqiad.wmnet,service=s6 * 09:07 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 08:58 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:56 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 08:56 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 08:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 08:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 08:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin2002.codfw.wmnet * 08:06 godog: remove cloudvirt1046, cloudvirt1062, cloudvirt1074, cloudvirt1075 from maintenance aggregate and put them in network-ovs - [[phab:T424802|T424802]] * 08:00 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin2002.codfw.wmnet * 07:58 hashar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] (duration: 32m 53s) * 07:57 fabfur: repooled cp4038 * 07:57 fabfur@cumin1003: conftool action : set/pooled=yes; selector: name=cp4038.* * 07:53 moritzm: installing pyjwt security updates * 07:47 moritzm: installing openjpeg2 security updates * 07:45 hashar@deploy1003: vadymts1, hashar: Continuing with deployment * 07:43 hashar@deploy1003: vadymts1, hashar: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:38 moritzm: installing python-urllib3 security updates * 07:37 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 07:30 fabfur: depooled cp4038 to investigate on possible maxmind failure * 07:30 fabfur@cumin1003: conftool action : set/pooled=no; selector: name=cp4038.* * 07:30 fabfur@cumin1003: conftool action : set/pooled=yes; selector: name=cp4038.* * 07:29 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 07:25 hashar@deploy1003: Started scap sync-world: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] * 06:13 moritzm: installing Linux 6.12.95 on trixie hosts * 05:20 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s6 * 05:20 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s4 * 05:19 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1015.eqiad.wmnet with reason: cloning * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 08s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-05 == * 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 01m 08s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-04 == * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 58s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-03 == * 17:08 topranks: revert protocol preference changes on cr3-ulsfo after upgrade * 16:53 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on cr2-eqord with reason: upgrade JunOS cr3-ulsfo * 16:53 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on cr4-ulsfo with reason: upgrade JunOS cr3-ulsfo * 16:48 topranks: reboot cr3-ulsfo to upgrade JunOS and reset linecard [[phab:T424839|T424839]] * 15:52 topranks: adjust outbound BGP policies on cr3-ulsfo to drain router of traffic [[phab:T424839|T424839]] * 15:45 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on lvs[4008-4010].ulsfo.wmnet with reason: upgrade JunOS cr3-ulsfo * 15:44 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on asw1-[22-23]-ulsfo,cr3-ulsfo,cr3-ulsfo IPv6,cr3-ulsfo.mgmt with reason: upgrade JunOS cr3-ulsfo * 15:36 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 15:35 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 15:35 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 14:40 cmooney@dns3003: END - running authdns-update * 14:26 cmooney@dns3003: START - running authdns-update * 14:26 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:26 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to ulsfo - cmooney@cumin1003" * 14:19 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to ulsfo - cmooney@cumin1003" * 14:16 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:38 sukhe@dns1004: END - running authdns-update * 13:35 sukhe@dns1004: START - running authdns-update * 13:26 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 13:26 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 13:26 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet * 13:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 13:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 13:16 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host sretest1005.eqiad.wmnet * 13:16 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 13:16 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 13:15 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 13:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:14 moritzm: imported samplicator 1.3.8rc1-1+deb13u1 to trixie-wikimedia/main [[phab:T337208|T337208]] * 13:13 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:07 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 13:07 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 13:02 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:02 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:58 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:57 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:57 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:53 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet * 12:50 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 12:47 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:41 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:40 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:39 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:32 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet * 12:26 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet * 12:23 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2005.wikimedia.org * 12:19 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2005.wikimedia.org * 12:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 12:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup[2004-2007].codfw.wmnet * 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[2004-2007].codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin2003" * 12:15 jynus@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[2004-2007].codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin2003" * 12:09 jynus@cumin2003: START - Cookbook sre.dns.netbox * 11:58 jynus@cumin2003: START - Cookbook sre.hosts.decommission for hosts backup[2004-2007].codfw.wmnet * 10:40 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup[1004-1007].eqiad.wmnet * 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[1004-1007].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 10:01 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[1004-1007].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 09:52 jynus@cumin1003: START - Cookbook sre.dns.netbox * 09:39 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:36 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup[1004-1007].eqiad.wmnet * 09:36 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:25 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 09:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 09:16 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 09:05 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:04 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:00 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:59 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:57 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:55 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:50 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 08:50 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 08:49 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 08:49 atsukoito: depooling cirrussearch in codfw because of regression after upgrade [[phab:T431091|T431091]] * 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts mirror1001.wikimedia.org * 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: mirror1001.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 08:29 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: mirror1001.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 08:18 jmm@cumin2003: START - Cookbook sre.dns.netbox * 08:11 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts mirror1001.wikimedia.org * 06:15 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 18s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-02 == * 22:55 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host contint1003.wikimedia.org with OS trixie * 22:29 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on contint1003.wikimedia.org with reason: host reimage * 22:23 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on contint1003.wikimedia.org with reason: host reimage * 22:05 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host contint1003.wikimedia.org with OS trixie * 22:03 mutante: contint1003 (zuul.wikimedia.org) - reimaging because of [[phab:T430510|T430510]]#12067628 [[phab:T418521|T418521]] * 22:03 dzahn@cumin2002: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on zuul.wikimedia.org with reason: reimage * 21:39 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 18s) * 21:39 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 21:20 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1003.eqiad.wmnet, repooling source-only afterwards * 21:19 sbassett: Deployed security fix for [[phab:T428829|T428829]] * 20:58 cmooney@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Release v0.11.2 update for new Aerleon - cmooney@cumin1003 * 20:55 cmooney@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Release v0.11.2 update for new Aerleon - cmooney@cumin1003 * 20:40 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] (duration: 12m 35s) * 20:36 arlolra@deploy1003: cscott, arlolra: Continuing with deployment * 20:35 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 20s) * 20:35 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 20:33 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host contint2003.wikimedia.org with OS trixie * 20:31 arlolra@deploy1003: cscott, arlolra: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Cha * 20:28 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] * 20:17 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] (duration: 08m 13s) * 20:14 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on contint2003.wikimedia.org with reason: host reimage * 20:13 sbassett@deploy1003: sbassett: Continuing with deployment * 20:11 sbassett@deploy1003: sbassett: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:09 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] * 20:08 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 20:08 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 20:08 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on contint2003.wikimedia.org with reason: host reimage * 20:05 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1003.eqiad.wmnet, repooling source-only afterwards * 19:49 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host contint2003.wikimedia.org with OS trixie * 19:48 mutante: contint2003 - reimaging because of [[phab:T430510|T430510]]#12067628 [[phab:T418521|T418521]] * 18:39 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 18:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 18:13 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2002.codfw.wmnet -> wcqs2003.codfw.wmnet, repooling source-only afterwards * 17:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1003.eqiad.wmnet with OS bookworm * 17:52 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1005.eqiad.wmnet * 17:52 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1005.eqiad.wmnet * 17:51 jasmine@cumin2002: conftool action : set/pooled=yes:weight=10; selector: name=wikikube-ctrl1005.eqiad.wmnet * 17:48 jasmine_: homer "cr*eqiad*" commit "Added new stacked control plane wikikube-ctrl1005" * 17:44 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply * 17:44 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply * 17:31 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] (duration: 09m 33s) * 17:26 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 17:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1003.eqiad.wmnet with reason: host reimage * 17:23 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:21 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] * 17:18 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1003.eqiad.wmnet with reason: host reimage * 17:16 rscout@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply * 17:16 rscout@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply * 17:16 rscout@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply * 17:15 rscout@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply * 17:12 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on wcqs[2002-2003].codfw.wmnet,wcqs1002.eqiad.wmnet with reason: reimaging hosts * 17:08 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 17:08 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 17:08 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 17:07 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 17:05 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 17:05 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 17:03 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "running to make sure all updates are synced - cmooney@cumin1003" * 17:03 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "running to make sure all updates are synced - cmooney@cumin1003" * 17:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs1003 * 17:00 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs1003 * 17:00 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 17:00 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1003.eqiad.wmnet with OS bookworm * 16:58 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Re-running - btullis@cumin1003" * 16:58 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Re-running - btullis@cumin1003" * 16:58 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2002.codfw.wmnet -> wcqs2003.codfw.wmnet, repooling source-only afterwards * 16:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-master1004.eqiad.wmnet with OS bookworm * 16:58 btullis@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 16:57 tappof: bump space for prometheus k8s-aux in eqiad * 16:55 cmooney@dns3003: END - running authdns-update * 16:55 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:55 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to eqsin - cmooney@cumin1003" * 16:55 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to eqsin - cmooney@cumin1003" * 16:53 cmooney@dns3003: START - running authdns-update * 16:52 ryankemper: [ml-serve-eqiad] Cleared out 1302 failed (Evicted) pods: `kubectl -n llm delete pods --field-selector=status.phase=Failed`, freeing calico-kube-controllers from OOM crashloop (evictions were caused by disk pressure) * 16:49 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 16:46 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:39 rzl@dns1004: END - running authdns-update * 16:37 rzl@dns1004: START - running authdns-update * 16:36 rzl@dns1004: START - running authdns-update * 16:35 rzl@deploy1003: Finished scap sync-world: [[phab:T416623|T416623]] (duration: 10m 19s) * 16:34 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 16:33 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-master1004.eqiad.wmnet with reason: host reimage * 16:30 rzl@deploy1003: rzl: Continuing with deployment * 16:28 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-master1004.eqiad.wmnet with reason: host reimage * 16:26 rzl@deploy1003: rzl: [[phab:T416623|T416623]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:25 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 16:25 rzl@deploy1003: Started scap sync-world: [[phab:T416623|T416623]] * 16:25 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 16:24 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 16:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: sync * 16:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: sync * 16:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-master1004.eqiad.wmnet with OS bookworm * 16:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-master1003.eqiad.wmnet with OS bookworm * 16:11 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 16:11 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 16:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Security updates * 16:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 16:08 root@cumin1003: START - Cookbook sre.mysql.parsercache * 16:08 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Security updates * 15:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-master1003.eqiad.wmnet with reason: host reimage * 15:54 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:54 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:54 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:54 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-master1003.eqiad.wmnet with reason: host reimage * 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Security updates * 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:45 root@cumin1003: START - Cookbook sre.mysql.parsercache * 15:45 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Security updates * 15:42 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-master1003.eqiad.wmnet with OS bookworm * 15:24 moritzm: installing busybox updates from bookworm point release * 15:20 moritzm: installing busybox updates from trixie point release * 15:15 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1021: Security updates * 15:15 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:15 root@cumin1003: START - Cookbook sre.mysql.parsercache * 15:15 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1021: Security updates * 15:13 moritzm: installing giflib security updates * 15:08 moritzm: installing Tomcat security updates * 14:57 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 14:56 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 14:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:53 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Unblock taavi - oblivian@cumin1003" * 14:53 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Unblock taavi - oblivian@cumin1003 * 14:53 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1021: Security updates * 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:53 root@cumin1003: START - Cookbook sre.mysql.parsercache * 14:53 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1021: Security updates * 14:53 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Unblock taavi - oblivian@cumin1003 * 14:52 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Unblock taavi - oblivian@cumin1003" * 14:46 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94711 and previous config saved to /var/cache/conftool/dbconfig/20260702-144644-fceratto.json * 14:36 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205', diff saved to https://phabricator.wikimedia.org/P94709 and previous config saved to /var/cache/conftool/dbconfig/20260702-143636-fceratto.json * 14:32 moritzm: installing libdbi-perl security updates * 14:26 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205', diff saved to https://phabricator.wikimedia.org/P94708 and previous config saved to /var/cache/conftool/dbconfig/20260702-142628-fceratto.json * 14:16 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94707 and previous config saved to /var/cache/conftool/dbconfig/20260702-141621-fceratto.json * 14:12 moritzm: installing rsync security updates * 14:11 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox) * 14:10 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94706 and previous config saved to /var/cache/conftool/dbconfig/20260702-140959-fceratto.json * 14:09 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2205.codfw.wmnet with reason: Maintenance * 14:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2205: Repooling after switchover * 14:07 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-test-master1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 14:06 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 14:06 Tran: Deployed patch for [[phab:T427287|T427287]] * 14:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:59 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2205: Repooling after switchover * 13:59 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2205: Repooling after switchover * 13:59 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:55 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2205: Repooling after switchover * 13:55 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2205 [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94704 and previous config saved to /var/cache/conftool/dbconfig/20260702-135505-fceratto.json * 13:54 moritzm: installing sed security updates * 13:53 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:52 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2209 to s3 primary [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94703 and previous config saved to /var/cache/conftool/dbconfig/20260702-135235-fceratto.json * 13:52 federico3: Starting s3 codfw failover from db2205 to db2209 - [[phab:T430912|T430912]] * 13:51 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:51 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 13:48 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:47 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2209 with weight 0 [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94702 and previous config saved to /var/cache/conftool/dbconfig/20260702-134719-fceratto.json * 13:47 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Primary switchover s3 [[phab:T430912|T430912]] * 13:44 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:44 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:44 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:40 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 13:38 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 13:37 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 13:36 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 13:36 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:34 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 13:30 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:29 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:29 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:27 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:26 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:25 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling restart_daemons on A:wikidough * 13:23 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 13:22 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 13:17 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 13:17 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns1004.wikimedia.org * 13:12 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:11 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart (exit_code=97) rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough * 13:11 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=97) rolling restart_daemons on A:wikidough * 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough * 13:09 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] (duration: 07m 20s) * 13:05 aude@deploy1003: jdrewniak, aude: Continuing with deployment * 13:04 aude@deploy1003: jdrewniak, aude: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:02 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] * 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts wdqs-categories1001.eqiad.wmnet * 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: wdqs-categories1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 12:10 jmm@dns1004: END - running authdns-update * 12:07 jmm@dns1004: START - running authdns-update * 11:51 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: wdqs-categories1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 11:44 btullis@cumin1003: START - Cookbook sre.dns.netbox * 11:42 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 11:42 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 11:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet * 11:39 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts wdqs-categories1001.eqiad.wmnet * 11:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet * 11:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet * 11:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet * 11:29 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 11:29 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 10:57 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2214: Repooling * 10:49 jmm@dns1004: END - running authdns-update * 10:47 jmm@dns1004: START - running authdns-update * 10:31 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94698 and previous config saved to /var/cache/conftool/dbconfig/20260702-103146-fceratto.json * 10:21 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213', diff saved to https://phabricator.wikimedia.org/P94696 and previous config saved to /var/cache/conftool/dbconfig/20260702-102137-fceratto.json * 10:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:19 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb1017.eqiad.wmnet * 10:18 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 10:18 fceratto@cumin1003: Removing es1033 from zarcillo [[phab:T408772|T408772]] * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts es1033.eqiad.wmnet * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: es1033.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:14 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: es1033.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:13 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb1017.eqiad.wmnet * 10:12 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2214.codfw.wmnet * 10:12 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2214.codfw.wmnet * 10:12 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2214: Repooling * 10:11 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213', diff saved to https://phabricator.wikimedia.org/P94693 and previous config saved to /var/cache/conftool/dbconfig/20260702-101130-fceratto.json * 10:10 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:10 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:03 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts es1033.eqiad.wmnet * 10:03 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 10:01 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94691 and previous config saved to /var/cache/conftool/dbconfig/20260702-100122-fceratto.json * 09:55 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94690 and previous config saved to /var/cache/conftool/dbconfig/20260702-095529-fceratto.json * 09:55 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2213.codfw.wmnet with reason: Maintenance * 09:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 09:53 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2213: Repooling after switchover * 09:51 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover * 09:44 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2213: Repooling after switchover * 09:39 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover * 09:39 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2213 [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94688 and previous config saved to /var/cache/conftool/dbconfig/20260702-093859-fceratto.json * 09:36 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2192 to s5 primary [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94687 and previous config saved to /var/cache/conftool/dbconfig/20260702-093650-fceratto.json * 09:36 federico3: Starting s5 codfw failover from db2213 to db2192 - [[phab:T430923|T430923]] * 09:30 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94686 and previous config saved to /var/cache/conftool/dbconfig/20260702-093004-fceratto.json * 09:24 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2192 with weight 0 [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94685 and previous config saved to /var/cache/conftool/dbconfig/20260702-092455-fceratto.json * 09:24 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 23 hosts with reason: Primary switchover s5 [[phab:T430923|T430923]] * 09:19 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220', diff saved to https://phabricator.wikimedia.org/P94684 and previous config saved to /var/cache/conftool/dbconfig/20260702-091957-fceratto.json * 09:16 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] (duration: 06m 57s) * 09:13 moritzm: installing libgcrypt20 security updates * 09:12 kharlan@deploy1003: kharlan: Continuing with deployment * 09:11 kharlan@deploy1003: kharlan: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:09 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220', diff saved to https://phabricator.wikimedia.org/P94683 and previous config saved to /var/cache/conftool/dbconfig/20260702-090950-fceratto.json * 09:09 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] * 09:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 09:01 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] (duration: 07m 07s) * 08:59 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94682 and previous config saved to /var/cache/conftool/dbconfig/20260702-085942-fceratto.json * 08:57 kharlan@deploy1003: kharlan: Continuing with deployment * 08:56 kharlan@deploy1003: kharlan: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:54 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] * 08:52 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:52 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:52 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94681 and previous config saved to /var/cache/conftool/dbconfig/20260702-085237-fceratto.json * 08:52 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2220.codfw.wmnet with reason: Maintenance * 08:43 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:40 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 08:25 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] (duration: 11m 44s) * 08:21 cscott@deploy1003: cscott: Continuing with deployment * 08:16 cscott@deploy1003: cscott: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:14 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] * 08:08 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 08:08 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1244: Migration of db1244.eqiad.wmnet completed * 08:02 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:02 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:01 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] (duration: 18m 58s) * 08:01 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:59 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 07:59 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:59 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:59 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2006.wikimedia.org * 07:58 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:57 cscott@deploy1003: cscott: Continuing with deployment * 07:56 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:56 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:56 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:55 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:55 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:55 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:54 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2006.wikimedia.org * 07:54 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:54 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 07:44 cscott@deploy1003: cscott: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:44 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2005.wikimedia.org * 07:44 moritzm: installing node-lodash security updates * 07:42 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] * 07:39 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2005.wikimedia.org * 07:30 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] (duration: 07m 28s) * 07:26 cscott@deploy1003: ssastry, cscott: Continuing with deployment * 07:25 cscott@deploy1003: ssastry, cscott: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:23 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1244: Migration of db1244.eqiad.wmnet completed * 07:22 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] * 07:16 wmde-fisch@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] (duration: 06m 55s) * 07:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1244.eqiad.wmnet with OS trixie * 07:11 wmde-fisch@deploy1003: wmde-fisch: Continuing with deployment * 07:11 wmde-fisch@deploy1003: wmde-fisch: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:09 wmde-fisch@deploy1003: Started scap sync-world: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] * 06:54 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1244.eqiad.wmnet with reason: host reimage * 06:50 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1244.eqiad.wmnet with reason: host reimage * 06:38 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1250.eqiad.wmnet with OS trixie * 06:34 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db1244.eqiad.wmnet with OS trixie * 06:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1244: Upgrading db1244.eqiad.wmnet * 06:25 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1244: Upgrading db1244.eqiad.wmnet * 06:25 cwilliams@cumin1003: dbmaint on s4@eqiad [[phab:T429893|T429893]] * 06:25 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 06:15 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1250.eqiad.wmnet with reason: host reimage * 06:14 cwilliams@dns1006: END - running authdns-update * 06:12 cwilliams@dns1006: START - running authdns-update * 06:11 cwilliams@dns1006: END - running authdns-update * 06:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db1244 [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94676 and previous config saved to /var/cache/conftool/dbconfig/20260702-061059-cwilliams.json * 06:09 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1250.eqiad.wmnet with reason: host reimage * 06:09 cwilliams@dns1006: START - running authdns-update * 06:08 aokoth@cumin1003: END (PASS) - Cookbook sre.vrts.upgrade (exit_code=0) on VRTS host vrts1003.eqiad.wmnet * 06:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db1160 to s4 primary and set section read-write [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94675 and previous config saved to /var/cache/conftool/dbconfig/20260702-060746-cwilliams.json * 06:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Set s4 eqiad as read-only for maintenance - [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94674 and previous config saved to /var/cache/conftool/dbconfig/20260702-060704-cwilliams.json * 06:06 cezmunsta: Starting s4 eqiad failover from db1244 to db1160 - [[phab:T430817|T430817]] * 06:04 aokoth@cumin1003: START - Cookbook sre.vrts.upgrade on VRTS host vrts1003.eqiad.wmnet * 05:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db1160 with weight 0 [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94673 and previous config saved to /var/cache/conftool/dbconfig/20260702-055927-cwilliams.json * 05:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 40 hosts with reason: Primary switchover s4 [[phab:T430817|T430817]] * 05:55 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1250.eqiad.wmnet with OS trixie * 05:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on db1250.eqiad.wmnet with reason: m3 master switchover [[phab:T430158|T430158]] * 05:39 marostegui: Failover m3 (phabricator) from db1250 to db1228 - [[phab:T430158|T430158]] * 05:32 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2234].codfw.wmnet,db[1217,1228,1250].eqiad.wmnet with reason: m3 master switchover [[phab:T430158|T430158]] * 04:45 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] (duration: 09m 08s) * 04:41 tstarling@deploy1003: tstarling, reedy: Continuing with deployment * 04:38 tstarling@deploy1003: tstarling, reedy: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 04:36 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 59s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:16 ryankemper: [[phab:T429844|T429844]] [opensearch] completed `cirrussearch2111` reimage; all codfw search clusters are green, all nodes now report `OpenSearch 2.19.5`, and the temporary chi voting exclusion has been removed * 00:57 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2111.codfw.wmnet with OS trixie * 00:29 ryankemper: [[phab:T429844|T429844]] [opensearch] depooled codfw search-omega/search-psi discovery records to match existing codfw search depool during OpenSearch 2.19 migration * 00:29 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2111.codfw.wmnet with reason: host reimage * 00:29 ryankemper@cumin2002: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 00:29 ryankemper@cumin2002: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 00:22 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2111.codfw.wmnet with reason: host reimage * 00:01 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2111.codfw.wmnet with OS trixie * 00:00 ryankemper: [[phab:T429844|T429844]] [opensearch] chi cluster recovered after stopping `opensearch_1@production-search-codfw` on `cirrussearch2111` == 2026-07-01 == * 23:59 ryankemper: [[phab:T429844|T429844]] [opensearch] stopped `opensearch_1@production-search-codfw` on `cirrussearch2111` after chi cluster-manager election churn following `voting_config_exclusions` POST; hoping this triggers a re-election * 23:52 cscott@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 23:51 cscott@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 23:51 cscott@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 23:50 cscott@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2003.codfw.wmnet with OS bookworm * 22:29 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 22:13 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 22:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2084.codfw.wmnet with OS trixie * 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2003.codfw.wmnet with reason: host reimage * 22:03 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 22:01 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2003.codfw.wmnet with reason: host reimage * 21:50 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 21:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2084.codfw.wmnet with reason: host reimage * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2003 * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2003 * 21:42 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2003 * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2003.codfw.wmnet 45.48.192.10.in-addr.arpa 5.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:42 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2003.codfw.wmnet 45.48.192.10.in-addr.arpa 5.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2003 - bking@cumin2003" * 21:42 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2003 - bking@cumin2003" * 21:36 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2084.codfw.wmnet with reason: host reimage * 21:35 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:34 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2003 * 21:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2003.codfw.wmnet with OS bookworm * 21:19 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2084.codfw.wmnet with OS trixie * 21:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2081.codfw.wmnet with OS trixie * 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2108.codfw.wmnet with OS trixie * 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2081.codfw.wmnet with reason: host reimage * 20:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2081.codfw.wmnet with reason: host reimage * 20:28 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2081.codfw.wmnet with OS trixie * 20:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2108.codfw.wmnet with reason: host reimage * 20:19 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2108.codfw.wmnet with reason: host reimage * 19:59 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2108.codfw.wmnet with OS trixie * 19:46 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2093.codfw.wmnet with OS trixie * 19:44 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 19:44 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jasmine@cumin2002" * 19:43 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jasmine@cumin2002" * 19:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2080.codfw.wmnet with OS trixie * 19:28 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 19:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2093.codfw.wmnet with reason: host reimage * 19:18 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 19:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2093.codfw.wmnet with reason: host reimage * 19:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2080.codfw.wmnet with reason: host reimage * 19:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2080.codfw.wmnet with reason: host reimage * 18:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2093.codfw.wmnet with OS trixie * 18:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2080.codfw.wmnet with OS trixie * 18:27 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 18:18 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] (duration: 09m 15s) * 18:13 jgiannelos@deploy1003: jgiannelos, neriah: Continuing with deployment * 18:11 jgiannelos@deploy1003: jgiannelos, neriah: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:09 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] * 17:40 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 16:58 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 30 hosts * 16:57 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for 30 hosts * 16:52 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2202.codfw.wmnet * 16:52 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2202.codfw.wmnet * 16:51 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt * 16:51 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt * 16:51 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lvs2012.codfw.wmnet * 16:51 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for lvs2012.codfw.wmnet * 16:49 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2076.codfw.wmnet with OS trixie * 16:49 brett: Start pybal on lvs2012 - [[phab:T429861|T429861]] * 16:49 pt1979@cumin1003: END (ERROR) - Cookbook sre.hosts.remove-downtime (exit_code=97) for 59 hosts * 16:48 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for 59 hosts * 16:42 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2061.codfw.wmnet with OS trixie * 16:30 dancy@deploy1003: Installation of scap version "4.271.0" completed for 2 hosts * 16:28 dancy@deploy1003: Installing scap version "4.271.0" for 2 host(s) * 16:23 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2076.codfw.wmnet with reason: host reimage * 16:19 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2061.codfw.wmnet with reason: host reimage * 16:18 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2076.codfw.wmnet with reason: host reimage * 16:18 jasmine@dns1004: END - running authdns-update * 16:16 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host restbase2039.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 16:16 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host restbase2039.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 16:16 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2061.codfw.wmnet with reason: host reimage * 16:15 jasmine@dns1004: START - running authdns-update * 16:14 jasmine@dns1004: END - running authdns-update * 16:12 jasmine@dns1004: START - running authdns-update * 16:07 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2202.codfw.wmnet with reason: maintenance * 16:06 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt with reason: Junos upograde * 16:00 papaul: ongoing maintenance on lsw1-b2-codfw * 16:00 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2076.codfw.wmnet with OS trixie * 15:59 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt * 15:59 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt * 15:57 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2061.codfw.wmnet with OS trixie * 15:55 pt1979@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2042,2046].codfw.wmnet * 15:55 pt1979@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2042,2046].codfw.wmnet * 15:51 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 15:51 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2220: Repooling after switchover * 15:50 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 15:50 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 15:48 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2092.codfw.wmnet with OS trixie * 15:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 15:40 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 15:38 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 15:37 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 15:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply * 15:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply * 15:32 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 15:32 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 15:30 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 15:29 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 15:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 15:25 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:22 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs2012.codfw.wmnet with reason: Rack B2 maintenance - [[phab:T429861|T429861]] * 15:21 brett: Stopping pybal on lvs2012 in preparation for codfw rack b2 maintenance - [[phab:T429861|T429861]] * 15:20 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2092.codfw.wmnet with reason: host reimage * 15:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:12 _joe_: restarted manually alertmanager-irc-relay * 15:12 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2092.codfw.wmnet with reason: host reimage * 15:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:12 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt with reason: Junos upograde * 15:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover * 15:07 pt1979@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2042,2046].codfw.wmnet * 15:06 pt1979@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2042,2046].codfw.wmnet * 15:02 papaul: ongoing maintenance on lsw1-a8-codfw * 14:31 topranks: POWERING DOWN CR1-EQIAD for line card installation [[phab:T426343|T426343]] * 14:31 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] (duration: 08m 57s) * 14:29 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:26 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 14:24 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:22 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] * 14:22 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover * 14:16 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:15 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover * 14:14 topranks: re-enable routing-engine graceful-failover on cr1-eqiad [[phab:T417873|T417873]] * 14:13 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:13 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2220: Repooling after switchover * 14:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:12 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:12 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:11 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:08 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] (duration: 10m 01s) * 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:07 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2220 [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94664 and previous config saved to /var/cache/conftool/dbconfig/20260701-140729-fceratto.json * 14:06 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:06 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:06 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:05 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2159 to s7 primary [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94663 and previous config saved to /var/cache/conftool/dbconfig/20260701-140503-fceratto.json * 14:04 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:04 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 14:04 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 14:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:04 dreamyjazz@deploy1003: anzx, dreamyjazz: Continuing with deployment * 14:04 federico3: Starting s7 codfw failover from db2220 to db2159 - [[phab:T430826|T430826]] * 14:03 jmm@dns1004: END - running authdns-update * 14:03 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:03 topranks: flipping cr1-eqiad active routing-enginer back to RE0 [[phab:T417873|T417873]] * 14:03 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudsw1-c8-eqiad,cloudsw1-d5-eqiad with reason: router upgrades eqiad * 14:01 jmm@dns1004: START - running authdns-update * 14:00 dreamyjazz@deploy1003: anzx, dreamyjazz: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:59 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2159 with weight 0 [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94662 and previous config saved to /var/cache/conftool/dbconfig/20260701-135906-fceratto.json * 13:58 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] * 13:57 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s7 [[phab:T430826|T430826]] * 13:56 topranks: reboot routing-enginer RE0 on cr1-eqiad [[phab:T417873|T417873]] * 13:48 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1006.wikimedia.org * 13:44 atsuko@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cirrussearch2092.codfw.wmnet with OS trixie * 13:43 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1006.wikimedia.org * 13:41 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2092.codfw.wmnet with OS trixie * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1005.wikimedia.org * 13:37 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1005.wikimedia.org * 13:37 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on pfw1-eqiad with reason: router upgrades eqiad * 13:35 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on lvs[1017-1020].eqiad.wmnet with reason: router upgrades eqiad * 13:34 caro@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] (duration: 07m 59s) * 13:30 caro@deploy1003: caro: Continuing with deployment * 13:28 caro@deploy1003: caro: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:27 topranks: route-engine failover cr1-eqiad * 13:26 caro@deploy1003: Started scap sync-world: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] * 13:15 topranks: rebooting routing-engine 1 on cr1-eqiad [[phab:T417873|T417873]] * 13:13 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] (duration: 08m 29s) * 13:13 moritzm: installing qemu security updates * 13:11 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 13:11 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 13:09 jgiannelos@deploy1003: jgiannelos: Continuing with deployment * 13:08 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 13:07 jgiannelos@deploy1003: jgiannelos: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:06 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 13:06 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2214.codfw.wmnet with reason: Maintenance * 13:05 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2214: Repooling after switchover * 13:05 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] * 13:04 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2214: Repooling after switchover * 13:04 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2214 [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94660 and previous config saved to /var/cache/conftool/dbconfig/20260701-130413-fceratto.json * 13:01 moritzm: installing python3.13 security updates * 13:00 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2229 to s6 primary [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94659 and previous config saved to /var/cache/conftool/dbconfig/20260701-125959-fceratto.json * 12:59 federico3: Starting s6 codfw failover from db2214 to db2229 - [[phab:T430814|T430814]] * 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on 13 hosts with reason: router upgrade and line card install * 12:51 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2229 with weight 0 [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94658 and previous config saved to /var/cache/conftool/dbconfig/20260701-125149-fceratto.json * 12:51 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 21 hosts with reason: Primary switchover s6 [[phab:T430814|T430814]] * 12:50 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2189.codfw.wmnet * 12:50 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2189.codfw.wmnet * 12:42 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2100.codfw.wmnet with OS trixie * 12:38 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2083.codfw.wmnet with OS trixie * 12:19 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2083.codfw.wmnet with reason: host reimage * 12:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 12:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2240: Migration of db2240.codfw.wmnet completed * 12:14 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2100.codfw.wmnet with reason: host reimage * 12:09 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2083.codfw.wmnet with reason: host reimage * 12:09 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2100.codfw.wmnet with reason: host reimage * 12:00 topranks: drain traffic on cr1-eqiad to allow for line card install and JunOS upgrade [[phab:T426343|T426343]] * 11:52 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2083.codfw.wmnet with OS trixie * 11:50 cmooney@dns2005: END - running authdns-update * 11:49 cmooney@dns2005: START - running authdns-update * 11:48 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2100.codfw.wmnet with OS trixie * 11:40 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/zotero: apply * 11:40 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/zotero: apply * 11:36 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/zotero: apply * 11:36 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/zotero: apply * 11:31 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2240: Migration of db2240.codfw.wmnet completed * 11:30 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply * 11:28 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply * 11:27 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:27 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:27 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:27 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:27 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:23 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2240.codfw.wmnet with OS trixie * 11:20 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:20 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:17 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:16 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:16 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:15 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2086.codfw.wmnet with OS trixie * 11:14 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2106.codfw.wmnet with OS trixie * 11:14 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:13 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:12 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:09 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2115.codfw.wmnet with OS trixie * 11:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2240.codfw.wmnet with reason: host reimage * 11:00 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2240.codfw.wmnet with reason: host reimage * 10:53 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2106.codfw.wmnet with reason: host reimage * 10:49 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2086.codfw.wmnet with reason: host reimage * 10:44 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2115.codfw.wmnet with reason: host reimage * 10:44 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2240.codfw.wmnet with OS trixie * 10:44 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2086.codfw.wmnet with reason: host reimage * 10:42 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2106.codfw.wmnet with reason: host reimage * 10:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2240: Upgrading db2240.codfw.wmnet * 10:41 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2240: Upgrading db2240.codfw.wmnet * 10:41 cwilliams@cumin1003: dbmaint on s4@codfw [[phab:T429893|T429893]] * 10:40 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 10:39 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2115.codfw.wmnet with reason: host reimage * 10:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2240 [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94653 and previous config saved to /var/cache/conftool/dbconfig/20260701-102658-cwilliams.json * 10:26 moritzm: installing nginx security updates * 10:26 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2086.codfw.wmnet with OS trixie * 10:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2179 to s4 primary [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94652 and previous config saved to /var/cache/conftool/dbconfig/20260701-102356-cwilliams.json * 10:23 cezmunsta: Starting s4 codfw failover from db2240 to db2179 - [[phab:T430127|T430127]] * 10:23 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2106.codfw.wmnet with OS trixie * 10:20 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2115.codfw.wmnet with OS trixie * 10:15 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2179 with weight 0 [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94651 and previous config saved to /var/cache/conftool/dbconfig/20260701-101531-cwilliams.json * 10:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 40 hosts with reason: Primary switchover s4 [[phab:T430127|T430127]] * 09:56 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template (take 2) - oblivian@cumin1003" * 09:56 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template (take 2) - oblivian@cumin1003 * 09:55 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template (take 2) - oblivian@cumin1003 * 09:55 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template (take 2) - oblivian@cumin1003" * 09:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:39 mszwarc@deploy1003: Synchronized private/SuggestedInvestigationsSignals/SuggestedInvestigationsSignal4n.php: Update SI signal 4n (duration: 06m 08s) * 09:21 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 09:21 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 09:14 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 09:14 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 09:02 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 09:02 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 08:54 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 08:38 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 08:38 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 08:36 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 08:21 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 08:21 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] (duration: 36m 11s) * 08:15 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 08:09 mszwarc@deploy1003: mszwarc, abi: Continuing with deployment * 08:03 mszwarc@deploy1003: mszwarc, abi: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:55 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 07:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 07:45 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] * 07:30 aqu@deploy1003: Finished deploy [analytics/refinery@410f205]: Regular analytics weekly train 2nd try [analytics/refinery@410f2050] (duration: 00m 22s) * 07:29 aqu@deploy1003: Started deploy [analytics/refinery@410f205]: Regular analytics weekly train 2nd try [analytics/refinery@410f2050] * 07:28 aqu@deploy1003: Finished deploy [analytics/refinery@410f205] (thin): Regular analytics weekly train THIN [analytics/refinery@410f2050] (duration: 01m 59s) * 07:28 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] (duration: 07m 19s) * 07:26 aqu@deploy1003: Started deploy [analytics/refinery@410f205] (thin): Regular analytics weekly train THIN [analytics/refinery@410f2050] * 07:26 aqu@deploy1003: Finished deploy [analytics/refinery@410f205]: Regular analytics weekly train [analytics/refinery@410f2050] (duration: 04m 32s) * 07:24 mszwarc@deploy1003: wmde-fisch, mszwarc: Continuing with deployment * 07:23 mszwarc@deploy1003: wmde-fisch, mszwarc: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:21 aqu@deploy1003: Started deploy [analytics/refinery@410f205]: Regular analytics weekly train [analytics/refinery@410f2050] * 07:21 aqu@deploy1003: Finished deploy [analytics/refinery@410f205] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@410f2050] (duration: 02m 01s) * 07:20 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] * 07:19 aqu@deploy1003: Started deploy [analytics/refinery@410f205] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@410f2050] * 07:13 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] (duration: 09m 13s) * 07:09 mszwarc@deploy1003: mszwarc, chlod, revi: Continuing with deployment * 07:06 mszwarc@deploy1003: mszwarc, chlod, revi: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:04 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] * 06:55 elukey: upgrade all trixie hosts to pywmflib 3.0 - [[phab:T430552|T430552]] * 06:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:43 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:43 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:42 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:42 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:41 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:41 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:35 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:35 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:34 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:34 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:31 jmm@cumin2003: DONE (PASS) - Cookbook sre.idm.logout (exit_code=0) Logging Niharika29 out of all services on: 2453 hosts * 06:30 oblivian@cumin1003: END (FAIL) - Cookbook sre.deploy.hiddenparma (exit_code=99) Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:30 oblivian@cumin1003: END (FAIL) - Cookbook sre.deploy.python-code (exit_code=99) hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:30 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:30 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:01 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2109.codfw.wmnet with OS trixie * 05:45 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on es1039.eqiad.wmnet with reason: issues * 05:41 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1027.eqiad.wmnet * 05:40 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2068.codfw.wmnet with OS trixie * 05:40 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2109.codfw.wmnet with reason: host reimage * 05:40 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1027.eqiad.wmnet,service=s2 * 05:40 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1027.eqiad.wmnet,service=s7 * 05:36 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2109.codfw.wmnet with reason: host reimage * 05:20 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2068.codfw.wmnet with reason: host reimage * 05:16 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2109.codfw.wmnet with OS trixie * 05:15 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2068.codfw.wmnet with reason: host reimage * 05:09 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2067.codfw.wmnet with OS trixie * 04:56 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2068.codfw.wmnet with OS trixie * 04:49 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2067.codfw.wmnet with reason: host reimage * 04:45 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2067.codfw.wmnet with reason: host reimage * 04:27 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2067.codfw.wmnet with OS trixie * 03:47 slyngshede@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1039.eqiad.wmnet with reason: Hardware crash * 03:21 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2107.codfw.wmnet with OS trixie * 02:59 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2107.codfw.wmnet with reason: host reimage * 02:55 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2085.codfw.wmnet with OS trixie * 02:51 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2072.codfw.wmnet with OS trixie * 02:51 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2107.codfw.wmnet with reason: host reimage * 02:35 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2085.codfw.wmnet with reason: host reimage * 02:31 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2107.codfw.wmnet with OS trixie * 02:30 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2072.codfw.wmnet with reason: host reimage * 02:26 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2085.codfw.wmnet with reason: host reimage * 02:22 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2072.codfw.wmnet with reason: host reimage * 02:09 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2085.codfw.wmnet with OS trixie * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 54s) * 02:03 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2072.codfw.wmnet with OS trixie * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es7 eqiad back to read-write - [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94649 and previous config saved to /var/cache/conftool/dbconfig/20260701-010716-ladsgroup.json * 01:05 ladsgroup@dns1004: END - running authdns-update * 01:05 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depool es1039 [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94648 and previous config saved to /var/cache/conftool/dbconfig/20260701-010551-ladsgroup.json * 01:03 ladsgroup@dns1004: START - running authdns-update * 01:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Promote es1035 to es7 primary [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94647 and previous config saved to /var/cache/conftool/dbconfig/20260701-010002-ladsgroup.json * 00:58 Amir1: Starting es7 eqiad failover from es1039 to es1035 - [[phab:T430765|T430765]] * 00:53 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es1035 with weight 0 [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94646 and previous config saved to /var/cache/conftool/dbconfig/20260701-005329-ladsgroup.json * 00:53 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 9 hosts with reason: Primary switchover es7 [[phab:T430765|T430765]] * 00:42 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es7 eqiad as read-only for maintenance - [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94645 and previous config saved to /var/cache/conftool/dbconfig/20260701-004221-ladsgroup.json * 00:20 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2102.codfw.wmnet with OS trixie * 00:15 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2103.codfw.wmnet with OS trixie * 00:05 dr0ptp4kt: DEPLOYED Refinery at {{Gerrit|4e7a2b32}} for changes: pageview allowlist {{Gerrit|1305158}} (+min.wikiquote) {{Gerrit|1305162}} (+bol.wikipedia), {{Gerrit|1305156}} (+isv.wikipedia); {{Gerrit|1305980}} (pv allowlist -api.wikimedia, sqoop +isvwiki); sqoop {{Gerrit|1295064}} (+globalimagelinks) {{Gerrit|1295069}} (+filerevision) using scap, then deployed onto HDFS (manual copyToLocal required additionally) == Other archives == See [[Server Admin Log/Archives]]. <noinclude> [[Category:SAL]] [[Category:Operations]] </noinclude> 0wz19d0g81x1jk1geynpnitmprtkjfj 2450659 2450658 2026-08-22T16:37:29Z Stashbot 7414 arlolra@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply 2450659 wikitext text/x-wiki == 2026-08-22 == * 16:37 arlolra@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:37 arlolra@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:37 arlolra@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:37 arlolra@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:36 arlolra@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:36 arlolra@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:36 arlolra@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:36 arlolra@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:33 arlolra@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:33 arlolra@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:33 arlolra@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:32 arlolra@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 35s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-21 == * 20:36 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in1001.wikimedia.org with reason: [[phab:T434750|T434750]] * 20:34 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in2001.wikimedia.org with reason: [[phab:T434750|T434750]] * 20:33 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out1001.wikimedia.org with reason: [[phab:T434750|T434750]] * 20:25 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out2001.wikimedia.org with reason: [[phab:T434750|T434750]] * 19:37 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host krb1002.eqiad.wmnet with OS bookworm * 19:00 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 18:59 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 18:51 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 18:51 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 18:35 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:35 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:27 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:27 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:16 bking@cumin2003: START - Cookbook sre.hosts.reimage for host krb1002.eqiad.wmnet with OS bookworm * 17:35 sukhe@dns1004: END - running authdns-update * 17:33 sukhe@dns1004: START - running authdns-update * 17:32 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns5004.wikimedia.org [reason: resolved authdns-update issues] * 17:32 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:32 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: force HEAD to {{Gerrit|be26e30ae101}} - sukhe@cumin1003" * 17:32 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: force HEAD to {{Gerrit|be26e30ae101}} - sukhe@cumin1003" * 17:28 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 17:28 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: service=authdns-update,name=dns5004.wikimedia.org [reason: resolving authdns-update issues] * 17:28 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:28 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: force HEAD to {{Gerrit|be26e30ae101}} - sukhe@cumin1003" * 17:28 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: force HEAD to {{Gerrit|be26e30ae101}} - sukhe@cumin1003" * 17:24 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 17:24 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.netbox (exit_code=97) * 17:23 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 17:16 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 17:12 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 17:10 sukhe@dns1004: END - running authdns-update * 17:08 sukhe@dns1004: START - running authdns-update * 17:08 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=dns5004.wikimedia.org [reason: resolving authdns-update issues] * 17:07 sukhe@dns1004: FAIL - running authdns-update * 17:05 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 17:05 sukhe@dns1004: START - running authdns-update * 17:01 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 16:59 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 16:56 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 16:53 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=dns5004.* [reason: trixie upgrade] * 16:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns5004.wikimedia.org * 16:52 cdobbins@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns5004.wikimedia.org * 16:44 cmooney@dns3003: END - running authdns-update * 16:41 cmooney@dns3003: START - running authdns-update * 16:41 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:41 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on eqsin<->codfw arelion - cmooney@cumin1003" * 16:37 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on eqsin<->codfw arelion - cmooney@cumin1003" * 16:33 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:11 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 16:08 sukhe@dns1004: END - running authdns-update * 16:08 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 16:06 sukhe@dns1004: START - running authdns-update * 16:04 cmooney@dns3003: END - running authdns-update * 16:02 cmooney@dns3003: START - running authdns-update * 16:00 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:00 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on eqord<->codfw arelion - cmooney@cumin1003" * 15:56 cdobbins@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host dns5004.wikimedia.org with OS trixie * 15:55 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on eqord<->codfw arelion - cmooney@cumin1003" * 15:53 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:51 cmooney@cumin1003: END (ERROR) - Cookbook sre.dns.netbox (exit_code=97) * 15:51 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:38 andrew@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudcephosd1042.eqiad.wmnet with OS bookworm * 15:18 andrew@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudcephosd1042.eqiad.wmnet with reason: host reimage * 15:17 dancy@deploy1003: Finished deploy [gerrit/gerrit@2cc11cc]: Deploying https://gerrit.wikimedia.org/r/c/operations/software/gerrit/+/1327669 ([[phab:T434726|T434726]]) (duration: 00m 14s) * 15:17 dancy@deploy1003: Started deploy [gerrit/gerrit@2cc11cc]: Deploying https://gerrit.wikimedia.org/r/c/operations/software/gerrit/+/1327669 ([[phab:T434726|T434726]]) * 15:13 andrew@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudcephosd1042.eqiad.wmnet with reason: host reimage * 15:09 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns5004.wikimedia.org with reason: host reimage * 15:05 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns5004.wikimedia.org with reason: host reimage * 14:53 andrew@cumin2003: START - Cookbook sre.hosts.reimage for host cloudcephosd1042.eqiad.wmnet with OS bookworm * 14:30 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns5004.wikimedia.org with OS trixie * 14:29 cdobbins@cumin1003: conftool action : set/pooled=no; selector: name=dns5004.* [reason: trixie upgrade] * 14:21 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:21 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove entries for cr2-eqord - cmooney@cumin1003" * 14:21 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove entries for cr2-eqord - cmooney@cumin1003" * 14:13 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 14:11 moritzm: imported openjdk 8u504-ga-1~deb12u1 for bookworm-wikimedia (backport of the latest Java 8 security fixes for bookworm) * 13:25 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "sync cr2-eqord router offline - cmooney@cumin1003" * 13:23 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "sync cr2-eqord router offline - cmooney@cumin1003" * 13:14 hashar@deploy1003: Finished deploy [integration/docroot@2d5ff9b]: opensource: add PersonalDashboard docs to MW components - [[phab:T435392|T435392]] (duration: 00m 15s) * 13:14 hashar@deploy1003: Started deploy [integration/docroot@2d5ff9b]: opensource: add PersonalDashboard docs to MW components - [[phab:T435392|T435392]] * 12:16 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:16 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: [[phab:T431682|T431682]] - filippo@cumin1003" * 12:16 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: [[phab:T431682|T431682]] - filippo@cumin1003" * 12:11 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2006.wikimedia.org with OS trixie * 12:00 kevinbazira@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 11:58 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 11:43 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2006.wikimedia.org with reason: host reimage * 11:41 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 11:38 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2006.wikimedia.org with reason: host reimage * 11:20 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2006.wikimedia.org with OS trixie * 11:11 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2005.wikimedia.org with OS trixie * 10:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2005.wikimedia.org with reason: host reimage * 10:53 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2005.wikimedia.org with reason: host reimage * 10:43 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-codfw * 10:43 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2011.codfw.wmnet * 10:43 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2011.codfw.wmnet * 10:40 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2011.codfw.wmnet * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2011.codfw.wmnet * 10:34 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2010.codfw.wmnet * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2010.codfw.wmnet * 10:33 fnegri@cumin1003: END (PASS) - Cookbook sre.wikireplicas.add-wiki (exit_code=0) for database bolwiki ([[phab:T429954|T429954]]) * 10:33 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2005.wikimedia.org with OS trixie * 10:30 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2010.codfw.wmnet * 10:25 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2010.codfw.wmnet * 10:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2009.codfw.wmnet * 10:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2009.codfw.wmnet * 10:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1006.wikimedia.org with OS trixie * 10:18 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2009.codfw.wmnet * 10:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2009.codfw.wmnet * 10:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2008.codfw.wmnet * 10:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2008.codfw.wmnet * 10:06 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2008.codfw.wmnet * 10:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1006.wikimedia.org with reason: host reimage * 10:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2008.codfw.wmnet * 10:01 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2007.codfw.wmnet * 10:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2007.codfw.wmnet * 09:57 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1006.wikimedia.org with reason: host reimage * 09:56 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2007.codfw.wmnet * 09:51 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2007.codfw.wmnet * 09:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 09:51 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 09:46 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2006.codfw.wmnet * 09:46 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1006.wikimedia.org with OS trixie * 09:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1005.wikimedia.org with OS trixie * 09:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2006.codfw.wmnet * 09:41 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2005.codfw.wmnet * 09:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2005.codfw.wmnet * 09:38 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2013.codfw.wmnet * 09:36 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2005.codfw.wmnet * 09:35 fnegri@cumin1003: START - Cookbook sre.wikireplicas.add-wiki for database bolwiki ([[phab:T429954|T429954]]) * 09:35 fnegri@cumin1003: END (PASS) - Cookbook sre.wikireplicas.add-wiki (exit_code=0) for database minwikiquote ([[phab:T429946|T429946]]) * 09:35 fnegri@cumin1003: START - Cookbook sre.wikireplicas.add-wiki for database minwikiquote ([[phab:T429946|T429946]]) * 09:32 blake@cumin1003: START - Cookbook sre.hosts.reboot-single for host rdb2013.codfw.wmnet * 09:30 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2011.codfw.wmnet * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1005.wikimedia.org with reason: host reimage * 09:26 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2005.codfw.wmnet * 09:26 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2004.codfw.wmnet * 09:26 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2004.codfw.wmnet * 09:24 blake@cumin1003: START - Cookbook sre.hosts.reboot-single for host rdb2011.codfw.wmnet * 09:22 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1005.wikimedia.org with reason: host reimage * 09:21 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2004.codfw.wmnet * 09:16 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1015.eqiad.wmnet * 09:16 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2004.codfw.wmnet * 09:16 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2003.codfw.wmnet * 09:16 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2003.codfw.wmnet * 09:13 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on cr[1-2]-eqiad,pfw1-eqiad with reason: upgrade pfw1a-eqiad and pfw1b-eqiad pair * 09:12 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 09:11 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 09:11 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 09:11 blake@cumin1003: START - Cookbook sre.hosts.reboot-single for host rdb1015.eqiad.wmnet * 09:10 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2003.codfw.wmnet * 09:09 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1013.eqiad.wmnet * 09:07 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1005.wikimedia.org with OS trixie * 09:03 blake@cumin1003: START - Cookbook sre.hosts.reboot-single for host rdb1013.eqiad.wmnet * 09:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2003.codfw.wmnet * 09:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2002.codfw.wmnet * 09:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2002.codfw.wmnet * 08:54 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2002.codfw.wmnet * 08:49 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2002.codfw.wmnet * 08:49 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2001.codfw.wmnet * 08:49 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2001.codfw.wmnet * 08:48 jmm@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts netmon2002.wikimedia.org * 08:47 jmm@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts netmon2002.wikimedia.org * 08:44 jmm@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts netmon2002.wikimedia.org * 08:44 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon2002.wikimedia.org * 08:43 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2001.codfw.wmnet * 08:36 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon2002.wikimedia.org * 08:34 jmm@dns1004: END - running authdns-update * 08:33 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2001.codfw.wmnet * 08:33 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-codfw * 08:31 jmm@dns1004: START - running authdns-update * 07:48 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327679{{!}}Block: Disable flaky API test (T435272 T389028)]], [[gerrit:1327678{{!}}API: wfDebugLog for thumberror]] (duration: 15m 34s) * 07:41 krinkle@deploy1003: krinkle: Continuing with deployment * 07:37 krinkle@deploy1003: krinkle: Backport for [[gerrit:1327679{{!}}Block: Disable flaky API test (T435272 T389028)]], [[gerrit:1327678{{!}}API: wfDebugLog for thumberror]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:33 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1327679{{!}}Block: Disable flaky API test (T435272 T389028)]], [[gerrit:1327678{{!}}API: wfDebugLog for thumberror]] * 07:25 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 07:25 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 07:24 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 07:18 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 07:18 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 07:15 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 07:14 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 07:14 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 07:14 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 07:13 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 07:03 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1283: Pool back * 06:42 jmm@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts netmon2002.wikimedia.org * 06:35 moritzm: powercycling netmon2002 * 06:18 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1283: Pool back * 06:17 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1283 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96213 and previous config saved to /var/cache/conftool/dbconfig/20260821-061743-marostegui.json * 04:59 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324963{{!}}Add Produnto to extension-list (T421436)]], [[gerrit:1324964{{!}}Enable Produnto on Beta (T421436)]] (duration: 34m 48s) * 04:45 tstarling@deploy1003: tstarling: Continuing with deployment * 04:44 tstarling@deploy1003: tstarling: Backport for [[gerrit:1324963{{!}}Add Produnto to extension-list (T421436)]], [[gerrit:1324964{{!}}Enable Produnto on Beta (T421436)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 04:24 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1324963{{!}}Add Produnto to extension-list (T421436)]], [[gerrit:1324964{{!}}Enable Produnto on Beta (T421436)]] * 04:21 arlolra@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 04:20 arlolra@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 04:20 arlolra@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 04:20 arlolra@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 41s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-20 == * 23:43 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327652{{!}}RunSingleJob: Add ProfilingContext::init() (T435422)]] (duration: 12m 06s) * 23:38 krinkle@deploy1003: krinkle: Continuing with deployment * 23:33 krinkle@deploy1003: krinkle: Backport for [[gerrit:1327652{{!}}RunSingleJob: Add ProfilingContext::init() (T435422)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:31 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1327652{{!}}RunSingleJob: Add ProfilingContext::init() (T435422)]] * 22:15 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1054.eqiad.wmnet with OS trixie * 22:14 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 22:14 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 21:58 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1054.eqiad.wmnet with reason: host reimage * 21:51 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1054.eqiad.wmnet with reason: host reimage * 21:36 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1054.eqiad.wmnet with OS trixie * 21:36 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:35 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327219{{!}}RunSingleJob: Define MW_ENTRY_POINT for flamegraph sample attribution (T435422)]] (duration: 08m 30s) * 21:31 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:31 krinkle@deploy1003: krinkle: Continuing with deployment * 21:31 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1054 * 21:31 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1054 * 21:29 krinkle@deploy1003: krinkle: Backport for [[gerrit:1327219{{!}}RunSingleJob: Define MW_ENTRY_POINT for flamegraph sample attribution (T435422)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:27 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1327219{{!}}RunSingleJob: Define MW_ENTRY_POINT for flamegraph sample attribution (T435422)]] * 21:17 maryum: Deployed security fix for [[phab:T433020|T433020]] * 20:59 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324752{{!}}InitialiseSettings: Enable 2FA banners on remaining private wikis (T428103)]], [[gerrit:1325920{{!}}Remove sending email to legal team about rejected requests (T374053)]] (duration: 07m 18s) * 20:54 reedy@deploy1003: neriah, reedy: Continuing with deployment * 20:54 reedy@deploy1003: neriah, reedy: Backport for [[gerrit:1324752{{!}}InitialiseSettings: Enable 2FA banners on remaining private wikis (T428103)]], [[gerrit:1325920{{!}}Remove sending email to legal team about rejected requests (T374053)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:51 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324752{{!}}InitialiseSettings: Enable 2FA banners on remaining private wikis (T428103)]], [[gerrit:1325920{{!}}Remove sending email to legal team about rejected requests (T374053)]] * 20:24 reedy@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.15,1.47.0-wmf.16,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/med * 20:23 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324752{{!}}InitialiseSettings: Enable 2FA banners on remaining private wikis (T428103)]], [[gerrit:1325920{{!}}Remove sending email to legal team about rejected requests (T374053)]] * 20:14 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 20:10 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 20:09 cdanis@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "bug fixes & UX fixes - cdanis@cumin1003" * 20:09 cdanis@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: bug fixes & UX fixes - cdanis@cumin1003 * 20:08 cdanis@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: bug fixes & UX fixes - cdanis@cumin1003 * 20:08 cdanis@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "bug fixes & UX fixes - cdanis@cumin1003" * 19:24 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327598{{!}}Make \Omicron non upright (like \Chi) (T434428)]], [[gerrit:1327596{{!}}Render overline of \bar with stretchy=false (T435456)]] (duration: 18m 54s) * 19:20 krinkle@deploy1003: krinkle: Continuing with deployment * 19:07 krinkle@deploy1003: krinkle: Backport for [[gerrit:1327598{{!}}Make \Omicron non upright (like \Chi) (T434428)]], [[gerrit:1327596{{!}}Render overline of \bar with stretchy=false (T435456)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:05 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1327598{{!}}Make \Omicron non upright (like \Chi) (T434428)]], [[gerrit:1327596{{!}}Render overline of \bar with stretchy=false (T435456)]] * 18:51 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327614{{!}}Avoid casting fpxmax to string (T318419)]] (duration: 07m 28s) * 18:50 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:46 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:46 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 18:45 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327614{{!}}Avoid casting fpxmax to string (T318419)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:43 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327614{{!}}Avoid casting fpxmax to string (T318419)]] * 18:37 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:37 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:37 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:36 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:07 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 18:05 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 18:01 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 18:01 sukhe@dns1004: END - running authdns-update * 17:59 sukhe@dns1004: START - running authdns-update * 17:58 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 17:57 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=dns6002.* [reason: depooling for trixie upgrade] * 17:56 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns6002.wikimedia.org * 17:56 cdobbins@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns6002.wikimedia.org * 17:51 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 17:51 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 17:34 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host stat1011.eqiad.wmnet with OS bookworm * 17:31 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns6002.wikimedia.org with OS trixie * 17:30 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 17:30 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 17:30 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 17:30 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 17:29 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 17:29 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 17:29 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 17:29 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:29 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:27 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:24 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327590{{!}}Make sure fpsmax is an int value (T318419)]] (duration: 08m 37s) * 17:20 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 17:17 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327590{{!}}Make sure fpsmax is an int value (T318419)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:16 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 17:15 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327590{{!}}Make sure fpsmax is an int value (T318419)]] * 16:51 swfrench-wmf: disable-puppet on A:cp for ATS Lua change - [[phab:T427666|T427666]] * 16:51 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db2901.codfw.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 16:43 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 16:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on stat1011.eqiad.wmnet with reason: host reimage * 16:39 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns6002.wikimedia.org with reason: host reimage * 16:36 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on stat1011.eqiad.wmnet with reason: host reimage * 16:36 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db2901.codfw.wmnet * 16:34 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns6002.wikimedia.org with reason: host reimage * 16:15 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns6002.wikimedia.org with OS trixie * 16:14 cdobbins@cumin1003: conftool action : set/pooled=no; selector: name=dns6002.* [reason: depooling for trixie upgrade] * 16:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host stat1011 * 16:10 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host stat1011 * 16:09 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host stat1011 * 16:09 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) stat1011.eqiad.wmnet 14.36.64.10.in-addr.arpa 4.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 bking@cumin2003: START - Cookbook sre.dns.wipe-cache stat1011.eqiad.wmnet 14.36.64.10.in-addr.arpa 4.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:09 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host stat1011 - bking@cumin2003" * 16:09 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host stat1011 - bking@cumin2003" * 16:05 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host stat1011 * 16:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host stat1011.eqiad.wmnet with OS bookworm * 16:00 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 15:59 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 15:56 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 15:55 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 15:37 jayme@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:35 jayme@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 15:35 jayme@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:32 jayme@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:32 jayme@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:30 jayme@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 15:30 jayme@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:28 jayme@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:28 jayme@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 15:28 fceratto@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host db1903.eqiad.wmnet * 15:28 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1903.eqiad.wmnet with OS trixie * 15:26 jayme@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 15:26 jayme@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 15:24 jayme@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 15:24 jayme@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 15:21 jayme@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 15:21 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 15:19 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 15:19 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 15:17 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 15:14 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1903.eqiad.wmnet with reason: host reimage * 15:07 fceratto@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1903.eqiad.wmnet with reason: host reimage * 14:54 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db1903.eqiad.wmnet with OS trixie * 14:53 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1903.eqiad.wmnet - fceratto@cumin1003" * 14:53 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1903.eqiad.wmnet - fceratto@cumin1003" * 14:53 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1903.eqiad.wmnet on all recursors * 14:53 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1903.eqiad.wmnet on all recursors * 14:53 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:53 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1903.eqiad.wmnet - fceratto@cumin1003" * 14:53 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1903.eqiad.wmnet - fceratto@cumin1003" * 14:49 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 14:49 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1903.eqiad.wmnet * 14:33 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=cp1100.* * 14:27 topranks: reconfigure eqiad<->codfw bgp settings * 14:22 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:22 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update entries used on new transport backup eqiad codfw - cmooney@cumin1003" * 14:19 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update entries used on new transport backup eqiad codfw - cmooney@cumin1003" * 14:14 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 14:14 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 14:13 moritzm: installing util-linux security updates * 14:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-staging-worker * 14:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2003.codfw.wmnet * 14:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2003.codfw.wmnet * 14:08 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2003.codfw.wmnet * 14:06 moritzm: installing libheif security updates * 13:58 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2003.codfw.wmnet * 13:58 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2002.codfw.wmnet * 13:58 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2002.codfw.wmnet * 13:56 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:56 fnegri@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for clouddb1025.eqiad.wmnet * 13:56 fnegri@cumin1003: START - Cookbook sre.hosts.remove-downtime for clouddb1025.eqiad.wmnet * 13:56 Lucas_WMDE: UTC afternoon backport+config window done * 13:53 moritzm: installing apr-util security updates * 13:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2002.codfw.wmnet * 13:50 fnegri@cumin1003: conftool action : set/weight=100; selector: name=clouddb1025.eqiad.wmnet * 13:49 fnegri@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1025.eqiad.wmnet * 13:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2002.codfw.wmnet * 13:41 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2001.codfw.wmnet * 13:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2001.codfw.wmnet * 13:41 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'. * 13:38 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'. * 13:38 fnegri@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on clouddb1025.eqiad.wmnet with reason: Removing s6 from clouddb1025 * 13:34 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2001.codfw.wmnet * 13:31 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'. * 13:29 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'. * 13:28 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet * 13:26 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host stat1009.eqiad.wmnet with OS bookworm * 13:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2001.codfw.wmnet * 13:24 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-staging-worker * 13:23 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1002.eqiad.wmnet * 13:21 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host stat1010.eqiad.wmnet with OS bookworm * 13:20 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1002.eqiad.wmnet * 13:20 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1001.eqiad.wmnet * 13:17 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1001.eqiad.wmnet * 13:16 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2001.codfw.wmnet * 13:13 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327511{{!}}UIC: Fix page:page instead of page:other in instrumentation]] (duration: 07m 00s) * 13:13 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2001.codfw.wmnet * 13:12 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2002.codfw.wmnet * 13:09 mszwarc@deploy1003: mszwarc: Continuing with deployment * 13:08 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1327511{{!}}UIC: Fix page:page instead of page:other in instrumentation]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2002.codfw.wmnet * 13:07 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2002.codfw.wmnet * 13:06 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1327511{{!}}UIC: Fix page:page instead of page:other in instrumentation]] * 13:04 jmm@dns1004: END - running authdns-update * 13:03 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2002.codfw.wmnet * 13:03 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2001.codfw.wmnet * 13:02 jmm@dns1004: START - running authdns-update * 13:01 cmooney@dns3003: END - running authdns-update * 13:00 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2001.codfw.wmnet * 12:59 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2001.codfw.wmnet * 12:59 cmooney@dns3003: START - running authdns-update * 12:57 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2001.codfw.wmnet * 12:56 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2002.codfw.wmnet * 12:55 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:55 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on drmrs<->eqiad GTT vpls - cmooney@cumin1003" * 12:54 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on drmrs<->eqiad GTT vpls - cmooney@cumin1003" * 12:54 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2002.codfw.wmnet * 12:54 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2003.codfw.wmnet * 12:50 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2003.codfw.wmnet * 12:49 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 12:48 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2001.codfw.wmnet * 12:46 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2001.codfw.wmnet * 12:46 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2002.codfw.wmnet * 12:43 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2002.codfw.wmnet * 12:43 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2003.codfw.wmnet * 12:42 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=1) for new host db1902.eqiad.wmnet * 12:42 fceratto@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host db1902.eqiad.wmnet with OS trixie * 12:41 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2003.codfw.wmnet * 12:40 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1003.eqiad.wmnet * 12:38 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1003.eqiad.wmnet * 12:37 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1002.eqiad.wmnet * 12:35 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1002.eqiad.wmnet * 12:35 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1001.eqiad.wmnet * 12:34 cmooney@dns3003: END - running authdns-update * 12:33 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1001.eqiad.wmnet * 12:32 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on stat1009.eqiad.wmnet with reason: host reimage * 12:31 cmooney@dns3003: START - running authdns-update * 12:31 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:31 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on drmrs<->eqiad cct - cmooney@cumin1003" * 12:28 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on drmrs<->eqiad cct - cmooney@cumin1003" * 12:28 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1902.eqiad.wmnet with reason: host reimage * 12:25 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on stat1009.eqiad.wmnet with reason: host reimage * 12:25 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 12:24 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on stat1010.eqiad.wmnet with reason: host reimage * 12:22 fceratto@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1902.eqiad.wmnet with reason: host reimage * 12:21 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on stat1010.eqiad.wmnet with reason: host reimage * 12:14 elukey: move the Docker Registry's /v2/dev/.* prefix to its dedicated S3 backend - [[phab:T432829|T432829]] * 12:12 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db1902.eqiad.wmnet with OS trixie * 12:09 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1902.eqiad.wmnet - fceratto@cumin1003" * 12:09 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1902.eqiad.wmnet - fceratto@cumin1003" * 12:09 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1902.eqiad.wmnet on all recursors * 12:09 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1902.eqiad.wmnet on all recursors * 12:09 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:08 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1902.eqiad.wmnet - fceratto@cumin1003" * 12:08 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1902.eqiad.wmnet - fceratto@cumin1003" * 12:08 tgr_: [[phab:T413390|T413390]] running CentralAuth:FixRenamedUserGlobalEditCount --wiki=metawiki --since=20250901000000 --until=20260301000000 --fix * 12:04 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1009.eqiad.wmnet with OS bookworm * 12:04 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1010.eqiad.wmnet with OS bookworm * 12:01 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 12:01 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1902.eqiad.wmnet * 12:00 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host stat1010.eqiad.wmnet with OS bookworm * 11:57 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1902.eqiad.wmnet * 11:57 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:57 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1902.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 11:57 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1902.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 11:51 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'. * 11:49 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'. * 11:48 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'. * 11:46 dpogorzelski@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'. * 11:39 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 11:37 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327518{{!}}Enable thumb.wikimedia.org on cswiki and fawiki (T427465)]] (duration: 10m 40s) * 11:35 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1902.eqiad.wmnet * 11:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1010.eqiad.wmnet with OS bookworm * 11:33 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 11:30 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327518{{!}}Enable thumb.wikimedia.org on cswiki and fawiki (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:26 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327518{{!}}Enable thumb.wikimedia.org on cswiki and fawiki (T427465)]] * 11:24 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host stat1010.eqiad.wmnet with OS bookworm * 11:07 fceratto@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host db1901.eqiad.wmnet * 11:07 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1901.eqiad.wmnet with OS trixie * 10:53 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1901.eqiad.wmnet with reason: host reimage * 10:47 tappof: bump space for prometheus k8s-aux in codfw * 10:47 tappof: bump space for prometheus k8s-dse in eqiad * 10:47 fceratto@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1901.eqiad.wmnet with reason: host reimage * 10:35 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db1901.eqiad.wmnet with OS trixie * 10:32 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:32 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:32 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1901.eqiad.wmnet on all recursors * 10:32 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1901.eqiad.wmnet on all recursors * 10:31 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:31 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:31 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:27 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:27 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1901.eqiad.wmnet * 10:23 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1010.eqiad.wmnet with OS bookworm * 10:20 fceratto@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts db1901.eqiad.wmnet * 10:20 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 10:18 blake@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 10:17 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:16 blake@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 10:13 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1901.eqiad.wmnet * 09:23 jelto@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'. * 09:22 jelto@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'. * 09:22 jelto: update cert-manager to 1.19.6 on wikikube staging-eqiad - [[phab:T427402|T427402]] * 09:20 moritzm: imported squid 7.6-2.1for trixie-wikimedia/main [[phab:T427282|T427282]] * 09:08 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 09:08 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 09:08 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 09:07 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 09:04 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 09:04 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:23 slyngshede@dns1004: END - running authdns-update * 08:21 slyngshede@dns1004: START - running authdns-update * 08:18 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.16 refs [[phab:T430835|T430835]] * 06:27 aokoth@dns1004: END - running authdns-update * 06:25 aokoth@dns1004: START - running authdns-update * 06:22 brennen@deploy1003: Finished deploy [phabricator/deployment@6b9b6ff]: deploy phab1005 for [[phab:T435087|T435087]] (duration: 00m 39s) * 06:21 brennen@deploy1003: Started deploy [phabricator/deployment@6b9b6ff]: deploy phab1005 for [[phab:T435087|T435087]] * 06:20 brennen@deploy1003: Finished deploy [phabricator/deployment@6b9b6ff]: deploy phab1004 for to pick up config values for [[phab:T435087|T435087]] (duration: 01m 46s) * 06:18 brennen@deploy1003: Started deploy [phabricator/deployment@6b9b6ff]: deploy phab1004 for to pick up config values for [[phab:T435087|T435087]] * 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 49s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-19 == * 23:19 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327197{{!}}Enable thumb.wikimedia.org on mediawiki.org (T427465)]] (duration: 10m 50s) * 23:18 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:16 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 23:15 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 23:10 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327197{{!}}Enable thumb.wikimedia.org on mediawiki.org (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:10 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:09 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 23:09 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:09 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 23:08 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327197{{!}}Enable thumb.wikimedia.org on mediawiki.org (T427465)]] * 22:58 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327201{{!}}Enable ReadingLists for all logged in users on test wiki (T435258)]] (duration: 11m 20s) * 22:50 jdlrobson@deploy1003: jdlrobson: Continuing with deployment * 22:49 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1327201{{!}}Enable ReadingLists for all logged in users on test wiki (T435258)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:46 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1327201{{!}}Enable ReadingLists for all logged in users on test wiki (T435258)]] * 22:42 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327162{{!}}Article: Split subjectpageheader by model and disable for wikitext]], [[gerrit:1327169{{!}}Make uppercase greek letters normal (non-italic) font (T434686 T434428)]], [[gerrit:1327176{{!}}Skin: Avoid DB lookup for pagecategorieslink message (T347123)]] (duration: 37m 52s) * 22:29 krinkle@deploy1003: krinkle: Continuing with deployment * 22:25 krinkle@deploy1003: krinkle: Backport for [[gerrit:1327162{{!}}Article: Split subjectpageheader by model and disable for wikitext]], [[gerrit:1327169{{!}}Make uppercase greek letters normal (non-italic) font (T434686 T434428)]], [[gerrit:1327176{{!}}Skin: Avoid DB lookup for pagecategorieslink message (T347123)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:04 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1327162{{!}}Article: Split subjectpageheader by model and disable for wikitext]], [[gerrit:1327169{{!}}Make uppercase greek letters normal (non-italic) font (T434686 T434428)]], [[gerrit:1327176{{!}}Skin: Avoid DB lookup for pagecategorieslink message (T347123)]] * 22:04 krinkle@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: awaiting CI (duration: 03m 06s) * 22:01 krinkle@deploy1003: Locking from deployment [ALL REPOSITORIES]: awaiting CI * 22:00 krinkle@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: awaiting CI (duration: 00m 01s) * 22:00 krinkle@deploy1003: Locking from deployment [ALL REPOSITORIES]: awaiting CI * 21:34 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2001.codfw.wmnet * 21:28 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2001.codfw.wmnet * 21:22 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327178{{!}}AccountRecovery: Notify the email address of the on file of the request (T425799)]] (duration: 47m 02s) * 21:13 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm * 21:09 catrope@deploy1003: catrope: Continuing with deployment * 20:55 catrope@deploy1003: catrope: Backport for [[gerrit:1327178{{!}}AccountRecovery: Notify the email address of the on file of the request (T425799)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:35 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1327178{{!}}AccountRecovery: Notify the email address of the on file of the request (T425799)]] * 20:31 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327128{{!}}Parsoid DataAccess: convert Parsoid fragment markers to/from strip tags (T432547)]] (duration: 07m 30s) * 20:27 catrope@deploy1003: catrope, arlolra: Continuing with deployment * 20:26 catrope@deploy1003: catrope, arlolra: Backport for [[gerrit:1327128{{!}}Parsoid DataAccess: convert Parsoid fragment markers to/from strip tags (T432547)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:24 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1327128{{!}}Parsoid DataAccess: convert Parsoid fragment markers to/from strip tags (T432547)]] * 20:23 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage * 20:17 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage * 20:15 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325878{{!}}[arwiki] Enable restricted user page editing and grant edit permissions (T434878)]] (duration: 08m 46s) * 20:11 catrope@deploy1003: catrope, gergesshamon: Continuing with deployment * 20:08 catrope@deploy1003: catrope, gergesshamon: Backport for [[gerrit:1325878{{!}}[arwiki] Enable restricted user page editing and grant edit permissions (T434878)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:06 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1325878{{!}}[arwiki] Enable restricted user page editing and grant edit permissions (T434878)]] * 19:59 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm * 19:56 eevans@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cassandra-dev2001.codfw.wmnet with OS bookworm * 19:56 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm * 19:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2207.codfw.wmnet with reason: Maintenance * 18:47 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319920{{!}}Allow setting a separate thumbUrl in production (T427465)]], [[gerrit:1327167{{!}}Fix wmgThumbUrl config (T427465)]] (duration: 18m 53s) * 18:43 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 18:30 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1319920{{!}}Allow setting a separate thumbUrl in production (T427465)]], [[gerrit:1327167{{!}}Fix wmgThumbUrl config (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:28 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1319920{{!}}Allow setting a separate thumbUrl in production (T427465)]], [[gerrit:1327167{{!}}Fix wmgThumbUrl config (T427465)]] * 18:26 sukhe@dns1004: END - running authdns-update * 18:24 sukhe@dns1004: START - running authdns-update * 18:09 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1319920{{!}}Allow setting a separate thumbUrl in production (T427465)]] * 18:03 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-eqiad and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 17:56 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-codfw and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 17:50 cmooney@dns3003: END - running authdns-update * 17:42 dancy@deploy1003: Installation of scap version "4.283.0" completed for 3 hosts * 17:41 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling reboot on A:durum and A:durum * 17:40 cmooney@dns3003: START - running authdns-update * 17:40 dancy@deploy1003: Installing scap version "4.283.0" for 3 host(s) * 17:38 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:37 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on GTT VPLS - cmooney@cumin1003" * 17:37 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --verbose --scoreLessThan=0.7 --exceptDatasetChecksums=[[phab:T434319|T434319]]-enwiki-models.txt # [[phab:T434319|T434319]] * 17:32 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on GTT VPLS - cmooney@cumin1003" * 17:29 sbassett: Deployed security fix for [[phab:T435210|T435210]] (wmf.16) * 17:26 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 17:22 sbassett: Deployed security fix for [[phab:T435210|T435210]] (wmf.15) * 17:00 sukhe@dns1004: END - running authdns-update * 16:58 sukhe@dns1004: START - running authdns-update * 16:53 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica-esams and A:liberica * 16:41 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica-esams and A:liberica * 16:41 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-eqiad and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 16:41 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-codfw and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 16:41 cjd91: sudo -i cookbook sre.cdn.roll-upgrade-ats --query 'A:cp-codfw' --task-id [[phab:T434478|T434478]] --reason '9.2.15 upgrade' * 16:41 cjd91: sudo -i cookbook sre.cdn.roll-upgrade-ats --query 'A:cp-eqiad' --task-id [[phab:T434478|T434478]] --reason '9.2.15 upgrade' * 16:40 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and A:durum * 16:28 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326881{{!}}Echo: Start using virtual domains (T380385)]] (duration: 13m 12s) * 16:23 urbanecm@deploy1003: urbanecm: Continuing with deployment * 16:21 urandom: Completed sessionstore Cassandra/JVM upgrade — [[phab:T435154|T435154]] * 16:21 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching sessionstore[2005-2006].codfw.wmnet,sessionstore[1005-1006].eqiad.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 16:19 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1326881{{!}}Echo: Start using virtual domains (T380385)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:15 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:15 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->codfw - cmooney@cumin1003" * 16:14 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1326881{{!}}Echo: Start using virtual domains (T380385)]] * 16:14 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327123{{!}}Revert^2 "Migrate database access to virtual domains" (T435305)]], [[gerrit:1327124{{!}}Pass the mapped domain of virtual-echo-shared to the push NameTableStores (T435305)]] (duration: 07m 42s) * 16:13 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching sessionstore[2005-2006].codfw.wmnet,sessionstore[1005-1006].eqiad.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 16:11 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->codfw - cmooney@cumin1003" * 16:10 urbanecm@deploy1003: urbanecm: Continuing with deployment * 16:08 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1327123{{!}}Revert^2 "Migrate database access to virtual domains" (T435305)]], [[gerrit:1327124{{!}}Pass the mapped domain of virtual-echo-shared to the push NameTableStores (T435305)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:06 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 16:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 16:06 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1327123{{!}}Revert^2 "Migrate database access to virtual domains" (T435305)]], [[gerrit:1327124{{!}}Pass the mapped domain of virtual-echo-shared to the push NameTableStores (T435305)]] * 16:06 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:03 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching sessionstore1004.eqiad.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 16:01 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching sessionstore1004.eqiad.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 15:56 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching sessionstore2004.codfw.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 15:54 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching sessionstore2004.codfw.wmnet: Upgrade Cassandra to 5.0.8 & Java to 17 — [[phab:T435154|T435154]] - eevans@cumin1003 * 15:53 urandom: beginning sessionstore Cassandra/JVM upgrade — [[phab:T435154|T435154]] * 15:52 urandom: beginning sessionstore Cassandra/JVM upgrade — [[phab:T432944|T432944]] * 15:51 cmooney@dns3003: END - running authdns-update * 15:49 cmooney@dns3003: START - running authdns-update * 15:48 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:48 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->codfw - cmooney@cumin1003" * 15:45 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->codfw - cmooney@cumin1003" * 15:44 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 15:44 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:43 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 15:42 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:42 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:38 cmooney@dns3003: END - running authdns-update * 15:36 cmooney@dns3003: START - running authdns-update * 15:36 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:36 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->eqsin - cmooney@cumin1003" * 15:34 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327138{{!}}Enable redis lock manager everywhere (T366938)]] (duration: 08m 36s) * 15:33 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->eqsin - cmooney@cumin1003" * 15:30 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:29 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 15:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1147.eqiad.wmnet with OS bookworm * 15:28 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327138{{!}}Enable redis lock manager everywhere (T366938)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:25 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327138{{!}}Enable redis lock manager everywhere (T366938)]] * 15:24 jmm@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host krb1002.eqiad.wmnet * 15:19 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2207.codfw.wmnet with reason: Host crashed * 15:17 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327098{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]], [[gerrit:1327101{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]] (duration: 07m 13s) * 15:12 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 15:12 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1327098{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]], [[gerrit:1327101{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:10 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1327098{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]], [[gerrit:1327101{{!}}UserRightsNotificationHandler: Use correct database for Echo (T435263)]] * 15:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2010.codfw.wmnet with OS trixie * 15:05 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 15:05 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1147.eqiad.wmnet with reason: host reimage * 14:59 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 14:59 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:58 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1147.eqiad.wmnet with reason: host reimage * 14:55 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:55 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:49 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 14:48 cmooney@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host durum1001.eqiad.wmnet * 14:46 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:46 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Delete db2902 ipv6 addr - fceratto@cumin1003" * 14:46 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Delete db2902 ipv6 addr - fceratto@cumin1003" * 14:43 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1147.eqiad.wmnet with OS bookworm * 14:42 tgr@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327094{{!}}SpecialMWOAuthListConsumers: Handle newFromMWUser returning null in addNavigationSubtitle (T435167)]] (duration: 19m 25s) * 14:42 cmooney@cumin1003: START - Cookbook sre.hosts.reboot-single for host durum1001.eqiad.wmnet * 14:42 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 14:41 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 14:38 tgr@deploy1003: tgr: Continuing with deployment * 14:36 tgr@deploy1003: tgr: Backport for [[gerrit:1327094{{!}}SpecialMWOAuthListConsumers: Handle newFromMWUser returning null in addNavigationSubtitle (T435167)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:28 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 14:28 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:28 cmooney@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host durum3005.esams.wmnet * 14:25 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 14:23 cmooney@cumin1003: START - Cookbook sre.hosts.reboot-single for host durum3005.esams.wmnet * 14:23 tgr@deploy1003: Started scap sync-world: Backport for [[gerrit:1327094{{!}}SpecialMWOAuthListConsumers: Handle newFromMWUser returning null in addNavigationSubtitle (T435167)]] * 14:18 gengh@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:17 topranks: disable puppet on hosts running BIRD BGP to test merge of patch to systemd healtchcheck service * 14:17 gengh@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:17 gengh@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:16 gengh@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:16 elukey: upgrade spicerack on cumin1003 and cumin2003 * 14:16 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:15 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314962{{!}}static: add new dir bimi/ for BIMI SVG and PEM file (T311685)]] (duration: 10m 00s) * 14:15 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:11 kharlan@deploy1003: kharlan, sukhe: Continuing with deployment * 14:11 gengh@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:09 gengh@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:09 gengh@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:08 kharlan@deploy1003: kharlan, sukhe: Backport for [[gerrit:1314962{{!}}static: add new dir bimi/ for BIMI SVG and PEM file (T311685)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:06 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 14:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:06 gengh@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:05 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1314962{{!}}static: add new dir bimi/ for BIMI SVG and PEM file (T311685)]] * 14:05 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host krb1002.eqiad.wmnet * 14:05 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:04 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:04 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326342{{!}}Revert^2 "wmf-config/ProductionServices: set URL for urldownloader to service record"]] (duration: 07m 40s) * 13:59 kharlan@deploy1003: kharlan, sukhe: Continuing with deployment * 13:58 kharlan@deploy1003: kharlan, sukhe: Backport for [[gerrit:1326342{{!}}Revert^2 "wmf-config/ProductionServices: set URL for urldownloader to service record"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:57 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host stat1008.eqiad.wmnet with OS bookworm * 13:56 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1326342{{!}}Revert^2 "wmf-config/ProductionServices: set URL for urldownloader to service record"]] * 13:56 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 13:54 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325532{{!}}srwiki: Allow bureaucrats to add and remove event-organizer group (T434748)]] (duration: 14m 56s) * 13:54 swfrench@dns1004: END - running authdns-update * 13:52 swfrench@dns1004: START - running authdns-update * 13:51 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:51 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 13:48 kharlan@deploy1003: kharlan, danielyepezgarces: Continuing with deployment * 13:45 swfrench@cumin2003: conftool action : set/pooled=yes; selector: name=wikikube-worker2330.codfw.wmnet * 13:44 swfrench-wmf: finished etcd-main codfw -> eqiad switchover - [[phab:T435103|T435103]] * 13:44 kharlan@deploy1003: kharlan, danielyepezgarces: Backport for [[gerrit:1325532{{!}}srwiki: Allow bureaucrats to add and remove event-organizer group (T434748)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:44 swfrench@cumin2003: conftool action : set/pooled=no; selector: name=wikikube-worker2330.codfw.wmnet * 13:41 swfrench@dns1004: END - running authdns-update * 13:39 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1325532{{!}}srwiki: Allow bureaucrats to add and remove event-organizer group (T434748)]] * 13:39 swfrench@dns1004: START - running authdns-update * 13:37 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327111{{!}}Special:AbuseReview: Add "no further action needed" review action (T435020)]], [[gerrit:1327110{{!}}AbuseReview: Take the review verdict as a REST path parameter (T435020)]] (duration: 31m 43s) * 13:31 swfrench-wmf: starting etcd-main codfw -> eqiad switchover - [[phab:T435103|T435103]] * 13:28 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:28 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 13:25 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=97) rolling reboot on A:durum and A:durum * 13:24 kharlan@deploy1003: kharlan: Continuing with deployment * 13:24 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1147 * 13:24 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1147 * 13:23 kharlan@deploy1003: kharlan: Backport for [[gerrit:1327111{{!}}Special:AbuseReview: Add "no further action needed" review action (T435020)]], [[gerrit:1327110{{!}}AbuseReview: Take the review verdict as a REST path parameter (T435020)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:16 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:16 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 13:12 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:12 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 13:10 cdobbins@cumin1003: END (ERROR) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=97) Rolling upgrade of ATS on A:cp-codfw and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 13:10 cdobbins@cumin1003: END (ERROR) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=97) Rolling upgrade of ATS on A:cp-eqiad and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 13:06 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1327111{{!}}Special:AbuseReview: Add "no further action needed" review action (T435020)]], [[gerrit:1327110{{!}}AbuseReview: Take the review verdict as a REST path parameter (T435020)]] * 13:06 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-eqiad and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 13:05 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-codfw and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 13:05 cjd91: sudo -i cookbook sre.cdn.roll-upgrade-ats --query 'A:cp-eqiad' --task-id [[phab:T434478|T434478]] --reason '9.2.15 upgrade' * 13:03 swfrench@dns1004: END - running authdns-update * 13:01 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:01 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 13:01 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host ncmonitor1001.eqiad.wmnet * 13:00 swfrench@dns1004: START - running authdns-update * 12:59 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 12:59 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 12:59 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 12:59 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 12:57 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and A:durum * 12:53 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on db2902.codfw.wmnet with reason: Cloning * 12:48 cmooney@dns3003: END - running authdns-update * 12:48 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327096{{!}}Switch to redis lock manager on s4 and s8 (T366938)]] (duration: 09m 11s) * 12:46 cmooney@dns3003: START - running authdns-update * 12:45 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:45 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->eqord cct - cmooney@cumin1003" * 12:45 elukey: move the /v2/releng.* prefix on the Docker Registry to its new s3 backend - [[phab:T432829|T432829]] * 12:43 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 12:42 jelto@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 12:42 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on ulsfo<->eqord cct - cmooney@cumin1003" * 12:42 jelto@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 12:41 jelto: update cert-manager to 1.19.6 on wikikube staging-codfw - [[phab:T427402|T427402]] * 12:40 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327096{{!}}Switch to redis lock manager on s4 and s8 (T366938)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:38 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327096{{!}}Switch to redis lock manager on s4 and s8 (T366938)]] * 12:38 blake@deploy1003: Finished scap sync-world: non-build deploy for [[phab:T417800|T417800]] (duration: 03m 52s) * 12:36 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 12:35 blake@deploy1003: Started scap sync-world: non-build deploy for [[phab:T417800|T417800]] * 12:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host krb2002.codfw.wmnet * 11:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host krb2002.codfw.wmnet * 11:49 moritzm: installing kerberos security updates * 11:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on stat1008.eqiad.wmnet with reason: host reimage * 11:44 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on stat1008.eqiad.wmnet with reason: host reimage * 11:31 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1327084{{!}}Revert "Migrate database access to virtual domains" (T435305)]] (duration: 11m 02s) * 11:29 gkyziridis@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:29 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:29 kart_: Updated MinT to 2026-06-04-131507-production ([[phab:T321316|T321316]]) * 11:28 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/machinetranslation: apply * 11:28 gkyziridis@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:26 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 11:24 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:24 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:23 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:23 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:23 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/machinetranslation: apply * 11:22 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1327084{{!}}Revert "Migrate database access to virtual domains" (T435305)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:21 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/machinetranslation: apply * 11:21 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:21 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:20 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1327084{{!}}Revert "Migrate database access to virtual domains" (T435305)]] * 11:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1008.eqiad.wmnet with OS bookworm * 11:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet * 11:17 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/machinetranslation: apply * 11:13 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/machinetranslation: apply * 11:12 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:12 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:10 kartik@deploy1003: helmfile [staging] START helmfile.d/services/machinetranslation: apply * 11:08 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-cron: apply * 11:08 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/mw-cron: apply * 11:08 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply * 11:08 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply * 11:07 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet * 11:07 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host stat1008.eqiad.wmnet with OS bookworm * 11:06 moritzm: upgrading the new trixie URL downloaders to Squid 7.6 [[phab:T427282|T427282]] * 11:01 gkyziridis@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin2002.codfw.wmnet * 10:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin2002.codfw.wmnet * 10:45 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1169.eqiad.wmnet with OS bookworm * 10:42 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1185.eqiad.wmnet with OS bookworm * 10:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1169.eqiad.wmnet with reason: host reimage * 10:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1185.eqiad.wmnet with reason: host reimage * 10:14 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1169.eqiad.wmnet with reason: host reimage * 10:14 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1185.eqiad.wmnet with reason: host reimage * 10:11 jmm@cumin2003: END (PASS) - Cookbook sre.netbox.restart-reboot (exit_code=0) rolling reboot on A:netbox * 10:06 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host stat1008.eqiad.wmnet with OS bookworm * 10:04 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 10:04 mpostoronca@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321579{{!}}Register the mediawiki.wikimedia_antiabuse.content_policy_score stream (T432848)]] (duration: 08m 53s) * 10:03 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 10:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 10:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 10:00 mpostoronca@deploy1003: mpostoronca: Continuing with deployment * 09:59 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1185.eqiad.wmnet with OS bookworm * 09:59 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1169.eqiad.wmnet with OS bookworm * 09:59 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.convert-disks (exit_code=0) for host ms-be1065 * 09:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1065.eqiad.wmnet with OS trixie * 09:58 mpostoronca@deploy1003: mpostoronca: Backport for [[gerrit:1321579{{!}}Register the mediawiki.wikimedia_antiabuse.content_policy_score stream (T432848)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:55 jmm@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netbox.discovery.wmnet. on all recursors * 09:55 mpostoronca@deploy1003: Started scap sync-world: Backport for [[gerrit:1321579{{!}}Register the mediawiki.wikimedia_antiabuse.content_policy_score stream (T432848)]] * 09:55 jmm@cumin2003: START - Cookbook sre.dns.wipe-cache netbox.discovery.wmnet. on all recursors * 09:52 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw2001.wikimedia.org with OS trixie * 09:51 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 09:51 jmm@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netbox.discovery.wmnet. on all recursors * 09:51 jmm@cumin2003: START - Cookbook sre.dns.wipe-cache netbox.discovery.wmnet. on all recursors * 09:51 jmm@cumin2003: START - Cookbook sre.netbox.restart-reboot rolling reboot on A:netbox * 09:46 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 09:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1056.eqiad.wmnet with OS trixie * 09:44 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 09:39 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 09:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 09:36 topranks: make HE transport circuits from magru live * 09:36 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 09:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb1003.eqiad.wmnet * 09:33 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage * 09:31 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb1003.eqiad.wmnet * 09:30 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 09:28 arnaudb@dns1006: END - running authdns-update * 09:27 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage * 09:27 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb2003.codfw.wmnet * 09:26 arnaudb@dns1006: START - running authdns-update * 09:24 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1056.eqiad.wmnet with reason: host reimage * 09:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb2003.codfw.wmnet * 09:20 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.convert-disks (exit_code=0) for host ms-be1068 * 09:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1068.eqiad.wmnet with OS trixie * 09:20 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 09:19 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "cloudvirt1057 - filippo@cumin1003" * 09:19 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "cloudvirt1057 - filippo@cumin1003" * 09:18 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1056.eqiad.wmnet with reason: host reimage * 09:18 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1057.eqiad.wmnet with OS trixie * 09:18 filippo@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 09:18 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 09:15 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 09:14 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 09:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host irc1003.wikimedia.org * 09:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:13 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1065.eqiad.wmnet with OS trixie * 09:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 09:09 moritzm: installing Postgresql security updates * 09:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host irc1003.wikimedia.org * 09:07 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 09:07 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:07 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1173.eqiad.wmnet with OS bookworm * 09:06 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw2001.wikimedia.org with OS trixie * 09:03 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.convert-disks (exit_code=0) for host ms-be1064 * 09:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1064.eqiad.wmnet with OS trixie * 09:03 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1056.eqiad.wmnet with OS trixie * 09:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1057.eqiad.wmnet with reason: host reimage * 09:01 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 08:58 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 08:56 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1057.eqiad.wmnet with reason: host reimage * 08:54 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1208.eqiad.wmnet with OS bookworm * 08:53 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2902.codfw.wmnet with OS trixie * 08:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 08:50 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1174.eqiad.wmnet with OS bookworm * 08:46 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1172.eqiad.wmnet with OS bookworm * 08:45 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1173.eqiad.wmnet with reason: host reimage * 08:41 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 08:40 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1057.eqiad.wmnet with OS trixie * 08:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1057.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 08:39 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1207.eqiad.wmnet with OS bookworm * 08:38 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2902.codfw.wmnet with reason: host reimage * 08:37 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 08:34 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1068.eqiad.wmnet with OS trixie * 08:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1208.eqiad.wmnet with reason: host reimage * 08:31 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1057.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 08:29 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 08:28 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1222.eqiad.wmnet onto db1276.eqiad.wmnet * 08:28 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1222: Pool db1222.eqiad.wmnet in after cloning * 08:28 fceratto@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2902.codfw.wmnet with reason: host reimage * 08:27 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1055.eqiad.wmnet with OS trixie * 08:27 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 08:26 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1174.eqiad.wmnet with reason: host reimage * 08:25 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003" * 08:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-misc2002.codfw.wmnet * 08:23 topranks: reboot pfw1-codfw firewall pair to upgrade JunOS [[phab:T434865|T434865]] * 08:22 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1172.eqiad.wmnet with reason: host reimage * 08:20 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1064.eqiad.wmnet with OS trixie * 08:20 mvernon@cumin2003: START - Cookbook sre.swift.convert-disks for host ms-be1065 * 08:18 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1207.eqiad.wmnet with reason: host reimage * 08:17 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1174.eqiad.wmnet with reason: host reimage * 08:17 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1173.eqiad.wmnet with reason: host reimage * 08:17 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1172.eqiad.wmnet with reason: host reimage * 08:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host mc-misc2002.codfw.wmnet * 08:15 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1208.eqiad.wmnet with reason: host reimage * 08:15 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1207.eqiad.wmnet with reason: host reimage * 08:14 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db2902.codfw.wmnet with OS trixie * 08:14 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.16 refs [[phab:T430835|T430835]] * 08:13 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db2902.codfw.wmnet * 08:13 fceratto@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host db2902.codfw.wmnet with OS trixie * 08:10 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr[1-2]-codfw with reason: upgrade pfw1a-codfw and pfw1b-codfw pair * 08:09 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1055.eqiad.wmnet with reason: host reimage * 08:07 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on pfw1-codfw with reason: upgrade pfw1a-codfw and pfw1b-codfw pair * 08:03 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1055.eqiad.wmnet with reason: host reimage * 08:02 arnaudb@dns1006: END - running authdns-update * 08:02 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1208.eqiad.wmnet with OS bookworm * 08:02 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1207.eqiad.wmnet with OS bookworm * 08:01 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1174.eqiad.wmnet with OS bookworm * 08:01 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1173.eqiad.wmnet with OS bookworm * 08:01 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1172.eqiad.wmnet with OS bookworm * 07:59 arnaudb@dns1006: START - running authdns-update * 07:58 arnaudb@dns1006: START - running authdns-update * 07:48 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1055.eqiad.wmnet with OS trixie * 07:45 moritzm: extend the disk of ldap-rw2001 by 80G [[phab:T331699|T331699]] * 07:42 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1222: Pool db1222.eqiad.wmnet in after cloning * 07:36 mvernon@cumin2003: START - Cookbook sre.swift.convert-disks for host ms-be1068 * 07:35 mvernon@cumin2003: START - Cookbook sre.swift.convert-disks for host ms-be1064 * 07:17 moritzm: installing imagemagick security updates * 07:14 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1277: Pool back * 07:14 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon1003.wikimedia.org * 07:07 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon1003.wikimedia.org * 07:03 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1280: Pool back * 07:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon2002.wikimedia.org * 06:55 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon2002.wikimedia.org * 06:54 moritzm: installing php8.2 security updates * 06:51 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1284: Pool back * 06:49 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1222: Depool db1222.eqiad.wmnet to then clone it to db1276.eqiad.wmnet - marostegui@cumin1003 * 06:49 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1222: Depool db1222.eqiad.wmnet to then clone it to db1276.eqiad.wmnet - marostegui@cumin1003 * 06:49 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1222.eqiad.wmnet onto db1276.eqiad.wmnet * 06:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd1005.eqiad.wmnet * 06:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd1005.eqiad.wmnet * 06:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd1004.eqiad.wmnet * 06:36 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2209: db2209 repool * 06:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd1004.eqiad.wmnet * 06:32 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast3007.wikimedia.org * 06:29 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1277: Pool back * 06:28 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1277 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96190 and previous config saved to /var/cache/conftool/dbconfig/20260819-062815-marostegui.json * 06:26 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast3007.wikimedia.org * 06:22 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul1001.eqiad.wmnet * 06:18 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul1001.eqiad.wmnet * 06:18 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul1003.eqiad.wmnet * 06:18 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1280: Pool back * 06:17 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1284 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96186 and previous config saved to /var/cache/conftool/dbconfig/20260819-061743-marostegui.json * 06:14 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul1003.eqiad.wmnet * 06:14 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul1002.eqiad.wmnet * 06:10 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul1002.eqiad.wmnet * 06:10 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2048.codfw.wmnet * 06:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2048.codfw.wmnet * 06:06 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1284: Pool back * 06:06 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1284 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96184 and previous config saved to /var/cache/conftool/dbconfig/20260819-060621-marostegui.json * 06:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2048.codfw.wmnet * 05:59 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2048.codfw.wmnet * 05:51 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2209: db2209 repool * 03:16 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1186.eqiad.wmnet with OS bookworm * 02:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1186.eqiad.wmnet with reason: host reimage * 02:46 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1186.eqiad.wmnet with reason: host reimage * 02:46 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2207 [[phab:T435270|T435270]]', diff saved to https://phabricator.wikimedia.org/P96181 and previous config saved to /var/cache/conftool/dbconfig/20260819-024627-marostegui.json * 02:44 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2204 to s2 primary [[phab:T435270|T435270]]', diff saved to https://phabricator.wikimedia.org/P96180 and previous config saved to /var/cache/conftool/dbconfig/20260819-024403-marostegui.json * 02:43 marostegui: Starting s2 codfw failover from db2207 to db2204 - [[phab:T435270|T435270]] * 02:39 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2204 with weight 0 [[phab:T435270|T435270]]', diff saved to https://phabricator.wikimedia.org/P96179 and previous config saved to /var/cache/conftool/dbconfig/20260819-023951-marostegui.json * 02:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s2 [[phab:T435270|T435270]] * 02:32 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1186.eqiad.wmnet with OS bookworm * 02:29 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-worker1186.eqiad.wmnet with OS bookworm * 02:18 denisse@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2207: Depooling replica * 02:18 denisse@cumin1003: START - Cookbook sre.mysql.depool depool db2207: Depooling replica * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 48s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-18 == * 23:55 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1264845{{!}}Remove unused/redundant wgMFNoindexPages=true setting (T255458)]] (duration: 09m 42s) * 23:51 krinkle@deploy1003: krinkle: Continuing with deployment * 23:48 krinkle@deploy1003: krinkle: Backport for [[gerrit:1264845{{!}}Remove unused/redundant wgMFNoindexPages=true setting (T255458)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:45 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1264845{{!}}Remove unused/redundant wgMFNoindexPages=true setting (T255458)]] * 23:38 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326944{{!}}Retire filebackend lock manager in favour of the default one (T366938)]] (duration: 08m 55s) * 23:34 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 23:31 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326944{{!}}Retire filebackend lock manager in favour of the default one (T366938)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:29 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326944{{!}}Retire filebackend lock manager in favour of the default one (T366938)]] * 23:27 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1170.eqiad.wmnet with OS bookworm * 23:21 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1205.eqiad.wmnet with OS bookworm * 23:20 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1171.eqiad.wmnet with OS bookworm * 23:15 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1206.eqiad.wmnet with OS bookworm * 23:05 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1170.eqiad.wmnet with reason: host reimage * 23:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1205.eqiad.wmnet with reason: host reimage * 22:57 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1171.eqiad.wmnet with reason: host reimage * 22:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1206.eqiad.wmnet with reason: host reimage * 22:53 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1205.eqiad.wmnet with reason: host reimage * 22:51 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1171.eqiad.wmnet with reason: host reimage * 22:51 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1170.eqiad.wmnet with reason: host reimage * 22:50 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1206.eqiad.wmnet with reason: host reimage * 22:36 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1206.eqiad.wmnet with OS bookworm * 22:35 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1205.eqiad.wmnet with OS bookworm * 22:35 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1186.eqiad.wmnet with OS bookworm * 22:35 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1171.eqiad.wmnet with OS bookworm * 22:35 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1170.eqiad.wmnet with OS bookworm * 22:33 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-worker1194.eqiad.wmnet with OS bookworm * 22:22 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326923{{!}}Enable redis lock manager on s6 (T366938)]] (duration: 11m 52s) * 22:18 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 22:13 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326923{{!}}Enable redis lock manager on s6 (T366938)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:10 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326923{{!}}Enable redis lock manager on s6 (T366938)]] * 22:04 sbassett: Deployed security fix for [[phab:T435234|T435234]] (wmf.16) * 21:54 sbassett: Deployed security fix for [[phab:T435234|T435234]] (wmf.15) * 21:38 caro@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326925{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326926{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326929{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]], [[gerrit:1326928{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]] (duration: 0 * 21:34 caro@deploy1003: caro: Continuing with deployment * 21:33 caro@deploy1003: caro: Backport for [[gerrit:1326925{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326926{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326929{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]], [[gerrit:1326928{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]] synced to the testservers (see h * 21:31 caro@deploy1003: Started scap sync-world: Backport for [[gerrit:1326925{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326926{{!}}Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326929{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]], [[gerrit:1326928{{!}}LLMSuggestionsEditCheck: final comparison should also have the object-replacements]] * 21:24 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1204.eqiad.wmnet with reason: 1204 datanode repair [[phab:T434494|T434494]] * 21:02 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326896{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]], [[gerrit:1326897{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]] (duration: 13m 56s) * 20:58 krinkle@deploy1003: krinkle: Continuing with deployment * 20:50 krinkle@deploy1003: krinkle: Backport for [[gerrit:1326896{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]], [[gerrit:1326897{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:49 ryankemper: `an-launcher1003` terminated process group `666809` (`rest_backfill_phase1.sh`) ~20 mins ago with `sudo kill -TERM -- -666809` after its local spark driver (`--driver-memory 64g`) repeatedly exhausted memory on the 32 GB VM and caused SSH to intermittently flap; host recovered to 27 GB available memory * 20:48 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1326896{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]], [[gerrit:1326897{{!}}ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]] * 20:35 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 20:33 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 20:31 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 20:31 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326870{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]], [[gerrit:1326871{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]] (duration: 07m 35s) * 20:28 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 20:26 kemayo@deploy1003: kemayo: Continuing with deployment * 20:26 ryankemper: `an-launcher1003` confirmed the host is flapping because of memory thrash. chasing down the source of the thrash * 20:25 kemayo@deploy1003: kemayo: Backport for [[gerrit:1326870{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]], [[gerrit:1326871{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:23 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1326870{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]], [[gerrit:1326871{{!}}LLMSuggestionsEditCheck: don't over-cache the description (T428641)]] * 20:23 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 20:20 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 20:14 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1168.eqiad.wmnet with OS bookworm * 20:08 zabe: zabe@deploy1003:~$ mwscript extensions/WikimediaMaintenance/maintenance/fixFileRevisionArchiveNameDrift.php enwiki # [[phab:T428406|T428406]] * 20:08 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1204.eqiad.wmnet with OS bookworm * 20:05 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1167.eqiad.wmnet with OS bookworm * 20:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1166.eqiad.wmnet with OS bookworm * 19:54 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1203.eqiad.wmnet with OS bookworm * 19:53 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326907{{!}}Revert "Disable redis lock manager on testwiki"]] (duration: 11m 05s) * 19:50 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1168.eqiad.wmnet with reason: host reimage * 19:47 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1204.eqiad.wmnet with reason: host reimage * 19:46 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 19:44 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326907{{!}}Revert "Disable redis lock manager on testwiki"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:42 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326907{{!}}Revert "Disable redis lock manager on testwiki"]] * 19:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1167.eqiad.wmnet with reason: host reimage * 19:37 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1166.eqiad.wmnet with reason: host reimage * 19:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1203.eqiad.wmnet with reason: host reimage * 19:32 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1167.eqiad.wmnet with reason: host reimage * 19:32 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1168.eqiad.wmnet with reason: host reimage * 19:32 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1166.eqiad.wmnet with reason: host reimage * 19:31 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1204.eqiad.wmnet with reason: host reimage * 19:31 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1203.eqiad.wmnet with reason: host reimage * 19:26 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324751{{!}}InitialiseSettings: Enable 2FA enforcement on more private wikis (T428103)]], [[gerrit:1326875{{!}}Add banner notifying of upcoming 2FA enforcement (T420792)]] (duration: 31m 46s) * 19:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1204.eqiad.wmnet with OS bookworm * 19:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1203.eqiad.wmnet with OS bookworm * 19:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1168.eqiad.wmnet with OS bookworm * 19:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1167.eqiad.wmnet with OS bookworm * 19:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1166.eqiad.wmnet with OS bookworm * 19:15 denisse: rebooting kafkamon2003.codfw.wmnet - [[phab:T435162|T435162]] * 19:14 denisse: rebooting kafkamon1003.eqiad.wmnet [[phab:T435162|T435162]] * 19:13 reedy@deploy1003: reedy: Continuing with deployment * 19:12 reedy@deploy1003: reedy: Backport for [[gerrit:1324751{{!}}InitialiseSettings: Enable 2FA enforcement on more private wikis (T428103)]], [[gerrit:1326875{{!}}Add banner notifying of upcoming 2FA enforcement (T420792)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:54 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324751{{!}}InitialiseSettings: Enable 2FA enforcement on more private wikis (T428103)]], [[gerrit:1326875{{!}}Add banner notifying of upcoming 2FA enforcement (T420792)]] * 18:50 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 18:44 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.16 refs [[phab:T430835|T430835]] * 18:34 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 18:31 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 18:22 aklapper@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326852{{!}}CategoryTree: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]], [[gerrit:1326853{{!}}CategoryViewer: Allow null $html in the CategoryViewerGenerateLink hook (T435161)]], [[gerrit:1326865{{!}}Flow: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]] (duration: 09m 57s) * 18:18 aklapper@deploy1003: jforrester, aklapper: Continuing with deployment * 18:17 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 18:14 aklapper@deploy1003: jforrester, aklapper: Backport for [[gerrit:1326852{{!}}CategoryTree: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]], [[gerrit:1326853{{!}}CategoryViewer: Allow null $html in the CategoryViewerGenerateLink hook (T435161)]], [[gerrit:1326865{{!}}Flow: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki * 18:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1156.eqiad.wmnet with OS bookworm * 18:12 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 18:12 aklapper@deploy1003: Started scap sync-world: Backport for [[gerrit:1326852{{!}}CategoryTree: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]], [[gerrit:1326853{{!}}CategoryViewer: Allow null $html in the CategoryViewerGenerateLink hook (T435161)]], [[gerrit:1326865{{!}}Flow: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]] * 18:11 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 18:08 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1146.eqiad.wmnet with OS bookworm * 18:07 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1177.eqiad.wmnet with OS bookworm * 18:00 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326380{{!}}Introduce main lock manager service (T366938 T427999)]] (duration: 11m 25s) * 17:58 ladsgroup@deploy1003: ladsgroup: Rolling back deployment * 17:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1202.eqiad.wmnet with OS bookworm * 17:55 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1201.eqiad.wmnet with OS bookworm * 17:50 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326380{{!}}Introduce main lock manager service (T366938 T427999)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:48 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326380{{!}}Introduce main lock manager service (T366938 T427999)]] * 17:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1156.eqiad.wmnet with reason: host reimage * 17:46 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1146.eqiad.wmnet with reason: host reimage * 17:45 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-esams and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 17:42 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 17:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1177.eqiad.wmnet with reason: host reimage * 17:38 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1202.eqiad.wmnet with reason: host reimage * 17:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1201.eqiad.wmnet with reason: host reimage * 17:30 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1156.eqiad.wmnet with reason: host reimage * 17:29 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1177.eqiad.wmnet with reason: host reimage * 17:28 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1146.eqiad.wmnet with reason: host reimage * 17:28 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1202.eqiad.wmnet with reason: host reimage * 17:27 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1201.eqiad.wmnet with reason: host reimage * 17:25 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 17:21 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 17:14 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1202.eqiad.wmnet with OS bookworm * 17:14 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1201.eqiad.wmnet with OS bookworm * 17:14 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1177.eqiad.wmnet with OS bookworm * 17:14 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1156.eqiad.wmnet with OS bookworm * 17:14 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1146.eqiad.wmnet with OS bookworm * 17:09 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326885{{!}}w/deployment-info.php: Handle new file format (T434726)]] (duration: 07m 15s) * 17:08 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 17:05 dancy@deploy1003: dancy: Continuing with deployment * 17:04 dancy@deploy1003: dancy: Backport for [[gerrit:1326885{{!}}w/deployment-info.php: Handle new file format (T434726)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:02 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1326885{{!}}w/deployment-info.php: Handle new file format (T434726)]] * 16:50 dancy@deploy1003: Finished scap sync-world: Testing [[phab:T434726|T434726]] (duration: 06m 40s) * 16:43 dancy@deploy1003: Started scap sync-world: Testing [[phab:T434726|T434726]] * 16:43 dancy@deploy1003: Installation of scap version "4.282.0" completed for 3 hosts * 16:41 dancy@deploy1003: Installing scap version "4.282.0" for 3 host(s) * 16:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1145.eqiad.wmnet with OS bookworm * 16:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1200.eqiad.wmnet with OS bookworm * 16:32 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1199.eqiad.wmnet with OS bookworm * 16:18 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1145.eqiad.wmnet with reason: host reimage * 16:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1200.eqiad.wmnet with reason: host reimage * 16:09 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1199.eqiad.wmnet with reason: host reimage * 16:05 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-esams and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 16:05 cjd91: sudo -i cookbook sre.cdn.roll-upgrade-ats --query 'A:cp-esams' --task-id [[phab:T434478|T434478]] --reason '9.2.15 upgrade' * 16:03 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1200.eqiad.wmnet with reason: host reimage * 16:02 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1145.eqiad.wmnet with reason: host reimage * 16:02 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1199.eqiad.wmnet with reason: host reimage * 15:50 moritzm: installing zip security updates * 15:48 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1200.eqiad.wmnet with OS bookworm * 15:47 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1199.eqiad.wmnet with OS bookworm * 15:47 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1145.eqiad.wmnet with OS bookworm * 15:41 topranks: bounce PIC 0/0 on cr1-magru to set port to 40G * 15:37 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1236.eqiad.wmnet with OS bookworm * 15:29 aikochou@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'ores-legacy' for release 'main' . * 15:26 aikochou@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'ores-legacy' for release 'main' . * 15:20 aikochou@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'ores-legacy' for release 'main' . * 15:14 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2002.codfw.wmnet * 15:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1236.eqiad.wmnet with reason: host reimage * 15:12 moritzm: failover ganeti master in codfw to ganeti2047 * 15:09 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1236.eqiad.wmnet with reason: host reimage * 15:09 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2044.codfw.wmnet * 15:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2002.codfw.wmnet * 15:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2044.codfw.wmnet * 15:04 brennen@deploy1003: Finished deploy [phabricator/deployment@6b9b6ff]: deploy phab1004 for [[phab:T435213|T435213]] (duration: 01m 01s) * 15:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2044.codfw.wmnet * 15:03 brennen@deploy1003: Started deploy [phabricator/deployment@6b9b6ff]: deploy phab1004 for [[phab:T435213|T435213]] * 15:03 brennen@deploy1003: Finished deploy [phabricator/deployment@6b9b6ff]: deploy phab2003 for [[phab:T435213|T435213]] (duration: 00m 57s) * 15:02 brennen@deploy1003: Started deploy [phabricator/deployment@6b9b6ff]: deploy phab2003 for [[phab:T435213|T435213]] * 14:57 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply * 14:57 arnaudb@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on phab2003.codfw.wmnet,phab[1004-1006].eqiad.wmnet with reason: maintenance * 14:56 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2044.codfw.wmnet * 14:55 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply * 14:53 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1236.eqiad.wmnet with OS bookworm * 14:51 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2043.codfw.wmnet * 14:51 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2043.codfw.wmnet * 14:45 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2043.codfw.wmnet * 14:33 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2043.codfw.wmnet * 14:22 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2042.codfw.wmnet * 14:22 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2042.codfw.wmnet * 14:21 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db2902.codfw.wmnet with OS trixie * 14:16 elukey: uploaded spicerack_13.2.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia * 14:16 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2042.codfw.wmnet * 14:04 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2042.codfw.wmnet * 14:04 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2041.codfw.wmnet * 14:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2041.codfw.wmnet * 13:59 phuedx@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: apply * 13:59 phuedx@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-main: apply * 13:59 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326838{{!}}Enable Suggested Investigations on hewiki (T435146)]] (duration: 11m 50s) * 13:58 phuedx@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: apply * 13:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2041.codfw.wmnet * 13:57 phuedx@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-main: apply * 13:57 phuedx@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-main: apply * 13:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard1003.eqiad.wmnet * 13:57 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-main: apply * 13:55 phuedx@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-logging-external: apply * 13:55 phuedx@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-logging-external: apply * 13:54 phuedx@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-logging-external: apply * 13:54 stran@deploy1003: stran: Continuing with deployment * 13:54 phuedx@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-logging-external: apply * 13:54 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-logging-external: apply * 13:54 phuedx@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-logging-external: apply * 13:53 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-logging-external: apply * 13:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard1003.eqiad.wmnet * 13:53 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2041.codfw.wmnet * 13:51 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db2902.codfw.wmnet - fceratto@cumin1003" * 13:51 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db2902.codfw.wmnet - fceratto@cumin1003" * 13:51 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2026.codfw.wmnet * 13:51 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard2003.codfw.wmnet * 13:50 phuedx@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: apply * 13:50 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-eqiad * 13:50 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp1001.eqiad.wmnet * 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp1001.eqiad.wmnet * 13:50 phuedx@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: apply * 13:49 phuedx@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: apply * 13:49 stran@deploy1003: stran: Backport for [[gerrit:1326838{{!}}Enable Suggested Investigations on hewiki (T435146)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:48 phuedx@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: apply * 13:48 phuedx@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics: apply * 13:47 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard2003.codfw.wmnet * 13:47 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics: apply * 13:47 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1326838{{!}}Enable Suggested Investigations on hewiki (T435146)]] * 13:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp1001.eqiad.wmnet * 13:44 cdanis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 13:43 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp1001.eqiad.wmnet * 13:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1315-1327].eqiad.wmnet * 13:43 cdanis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 13:43 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1315-1327].eqiad.wmnet * 13:40 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki1001.eqiad.wmnet * 13:35 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1315-1327].eqiad.wmnet * 13:34 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host rpki1001.eqiad.wmnet * 13:27 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1315-1327].eqiad.wmnet * 13:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1302-1314].eqiad.wmnet * 13:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1302-1314].eqiad.wmnet * 13:21 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326775{{!}}SI: Instrument case update on first edit (T435048)]] (duration: 07m 12s) * 13:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1302-1314].eqiad.wmnet * 13:17 stran@deploy1003: stran: Continuing with deployment * 13:16 stran@deploy1003: stran: Backport for [[gerrit:1326775{{!}}SI: Instrument case update on first edit (T435048)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:14 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1326775{{!}}SI: Instrument case update on first edit (T435048)]] * 13:13 moritzm: installing util-linux security updates * 13:11 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1302-1314].eqiad.wmnet * 13:10 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326232{{!}}prv: Enable parsoid rendering for 5 wikis (T435115)]] (duration: 08m 17s) * 13:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1288-1289,1291-1301].eqiad.wmnet * 13:10 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply * 13:10 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1288-1289,1291-1301].eqiad.wmnet * 13:10 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply * 13:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki2003.codfw.wmnet * 13:06 jgiannelos@deploy1003: jgiannelos: Continuing with deployment * 13:05 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host rpki2003.codfw.wmnet * 13:04 jgiannelos@deploy1003: jgiannelos: Backport for [[gerrit:1326232{{!}}prv: Enable parsoid rendering for 5 wikis (T435115)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1288-1289,1291-1301].eqiad.wmnet * 13:02 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1326232{{!}}prv: Enable parsoid rendering for 5 wikis (T435115)]] * 12:53 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1288-1289,1291-1301].eqiad.wmnet * 12:53 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1273,1275-1287].eqiad.wmnet * 12:53 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1273,1275-1287].eqiad.wmnet * 12:52 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt1002.wikimedia.org * 12:51 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db2902.codfw.wmnet on all recursors * 12:51 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db2902.codfw.wmnet on all recursors * 12:51 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:51 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db2902.codfw.wmnet - fceratto@cumin1003" * 12:51 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db2902.codfw.wmnet - fceratto@cumin1003" * 12:46 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt1002.wikimedia.org * 12:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt2002.wikimedia.org * 12:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1273,1275-1287].eqiad.wmnet * 12:42 dhinus: repooled clouddb1032 that was currently <nowiki>{</nowiki>"weight": 0, "pooled": "inactive"<nowiki>}</nowiki> for both s4 and s6 * 12:41 dhinus: also depooled clouddb1017 (forgot it in the previous list) * 12:41 fnegri@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet * 12:40 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 12:40 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db2902.codfw.wmnet * 12:40 fnegri@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032.eqiad.wmnet * 12:40 fnegri@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032 * 12:39 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1017.eqiad.wmnet * 12:39 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt2002.wikimedia.org * 12:38 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1020.eqiad.wmnet * 12:38 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1018.eqiad.wmnet * 12:38 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet * 12:37 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1014.eqiad.wmnet * 12:37 fnegri@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1013.eqiad.wmnet * 12:37 dhinus: depool again clouddb10[13,14,16,18,20] that were repooled by the cookbook sre.mysql.multiinstance_reboot * 12:36 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1273,1275-1287].eqiad.wmnet * 12:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1248-1261].eqiad.wmnet * 12:36 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1248-1261].eqiad.wmnet * 12:35 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2026.codfw.wmnet * 12:34 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 12:34 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 12:34 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 12:34 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 12:29 lucaswerkmeister-wmde@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 12:28 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2026.codfw.wmnet * 12:28 lucaswerkmeister-wmde@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 12:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1248-1261].eqiad.wmnet * 12:27 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.addnode (exit_code=0) for new host ganeti2046.codfw.wmnet to cluster codfw and group A * 12:26 moritzm: readded ganeti2046 to the codfw cluster following firmware update and reimage [[phab:T434681|T434681]] * 12:23 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply * 12:23 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply * 12:23 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply * 12:22 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply * 12:22 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply * 12:22 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply * 12:21 jmm@cumin2003: START - Cookbook sre.ganeti.addnode for new host ganeti2046.codfw.wmnet to cluster codfw and group A * 12:21 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1169.eqiad.wmnet onto db1283.eqiad.wmnet * 12:21 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1169: Pool db1169.eqiad.wmnet in after cloning * 12:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1248-1261].eqiad.wmnet * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1149-1153,1158,1240-1247].eqiad.wmnet * 12:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1149-1153,1158,1240-1247].eqiad.wmnet * 12:11 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1149-1153,1158,1240-1247].eqiad.wmnet * 12:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1168.eqiad.wmnet onto db1282.eqiad.wmnet * 12:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1168: Pool db1168.eqiad.wmnet in after cloning * 12:06 jmm@cumin2003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-test-eqiad * 12:05 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2026.codfw.wmnet * 12:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1149-1153,1158,1240-1247].eqiad.wmnet * 12:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1128-1134,1142-1148].eqiad.wmnet * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2046.codfw.wmnet * 12:01 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1128-1134,1142-1148].eqiad.wmnet * 11:57 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2025.codfw.wmnet * 11:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2025.codfw.wmnet * 11:54 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2046.codfw.wmnet * 11:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1128-1134,1142-1148].eqiad.wmnet * 11:50 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2025.codfw.wmnet * 11:46 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1128-1134,1142-1148].eqiad.wmnet * 11:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1114-1127].eqiad.wmnet * 11:45 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1114-1127].eqiad.wmnet * 11:39 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db2901.codfw.wmnet * 11:39 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db2901.codfw.wmnet * 11:36 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2025.codfw.wmnet * 11:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1114-1127].eqiad.wmnet * 11:35 fceratto@cumin1003: END (ERROR) - Cookbook sre.ganeti.makevm (exit_code=93) for new host db1901.eqiad.wmnet * 11:35 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1169: Pool db1169.eqiad.wmnet in after cloning * 11:35 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 11:30 jmm@cumin2003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-test-eqiad * 11:27 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1114-1127].eqiad.wmnet * 11:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1076-1081,1084-1087,1093-1095,1113].eqiad.wmnet * 11:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1076-1081,1084-1087,1093-1095,1113].eqiad.wmnet * 11:25 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1168: Pool db1168.eqiad.wmnet in after cloning * 11:24 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1194.eqiad.wmnet with OS bookworm * 11:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1181.eqiad.wmnet with OS bookworm * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2050.codfw.wmnet * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2050.codfw.wmnet * 11:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1076-1081,1084-1087,1093-1095,1113].eqiad.wmnet * 11:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2050.codfw.wmnet * 11:14 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2004.codfw.wmnet * 11:10 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2050.codfw.wmnet * 11:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1076-1081,1084-1087,1093-1095,1113].eqiad.wmnet * 11:09 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1286: Pool back * 11:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1045-1050,1056-1057,1064-1066,1073-1075].eqiad.wmnet * 11:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1045-1050,1056-1057,1064-1066,1073-1075].eqiad.wmnet * 11:08 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2004.codfw.wmnet * 11:07 moritzm: installing PHP 8.4 security updates * 11:06 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on an-worker1194.eqiad.wmnet with reason: host reimage * 11:06 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1194.eqiad.wmnet with reason: host reimage * 11:05 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2049.codfw.wmnet * 11:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2049.codfw.wmnet * 11:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid1003.eqiad.wmnet * 11:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1045-1050,1056-1057,1064-1066,1073-1075].eqiad.wmnet * 10:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1181.eqiad.wmnet with reason: host reimage * 10:59 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2049.codfw.wmnet * 10:59 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid1003.eqiad.wmnet * 10:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid2003.codfw.wmnet * 10:54 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1181.eqiad.wmnet with reason: host reimage * 10:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid2003.codfw.wmnet * 10:51 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1901.eqiad.wmnet on all recursors * 10:51 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1901.eqiad.wmnet on all recursors * 10:51 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:51 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:51 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 10:50 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2049.codfw.wmnet * 10:50 blake@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:50 blake@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:49 blake@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:48 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1045-1050,1056-1057,1064-1066,1073-1075].eqiad.wmnet * 10:48 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1044].eqiad.wmnet * 10:48 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1044].eqiad.wmnet * 10:47 blake@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:46 blake@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:46 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2047.codfw.wmnet * 10:46 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:46 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2047.codfw.wmnet * 10:46 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1901.eqiad.wmnet * 10:46 blake@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:44 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1168: Depool db1168.eqiad.wmnet to then clone it to db1282.eqiad.wmnet - marostegui@cumin1003 * 10:44 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1168: Depool db1168.eqiad.wmnet to then clone it to db1282.eqiad.wmnet - marostegui@cumin1003 * 10:44 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1168.eqiad.wmnet onto db1282.eqiad.wmnet * 10:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2047.codfw.wmnet * 10:40 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1044].eqiad.wmnet * 10:37 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2047.codfw.wmnet * 10:37 fceratto@cumin1003: END (ERROR) - Cookbook sre.ganeti.makevm (exit_code=93) for new host db1901.eqiad.wmnet * 10:36 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 10:34 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:33 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2032.codfw.wmnet * 10:33 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2032.codfw.wmnet * 10:32 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1044].eqiad.wmnet * 10:32 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-eqiad * 10:27 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2032.codfw.wmnet * 10:24 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1286: Pool back * 10:24 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1286 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96163 and previous config saved to /var/cache/conftool/dbconfig/20260818-102431-marostegui.json * 10:22 moritzm: installing Django security updates * 10:20 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul2002.codfw.wmnet * 10:20 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul2001.codfw.wmnet * 10:17 blake@deploy1003: Finished scap sync-world: no-build deployment for [[phab:T417800|T417800]] (duration: 04m 40s) * 10:16 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul2002.codfw.wmnet * 10:16 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul2001.codfw.wmnet * 10:15 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2032.codfw.wmnet * 10:14 blake@deploy1003: Started scap sync-world: no-build deployment for [[phab:T417800|T417800]] * 10:12 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul2003.codfw.wmnet * 10:12 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy2001.codfw.wmnet * 10:12 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy3001.esams.wmnet * 10:12 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy1001.eqiad.wmnet * 10:08 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2031.codfw.wmnet * 10:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2031.codfw.wmnet * 10:08 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host zuul2003.codfw.wmnet * 10:08 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy2001.codfw.wmnet * 10:08 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy3001.esams.wmnet * 10:08 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy1001.eqiad.wmnet * 10:07 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy1002.eqiad.wmnet * 10:05 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy2002.codfw.wmnet * 10:05 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy3002.esams.wmnet * 10:04 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy1002.eqiad.wmnet * 10:03 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy4003.ulsfo.wmnet * 10:03 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy5003.eqsin.wmnet * 10:02 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2031.codfw.wmnet * 10:02 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1169: Depool db1169.eqiad.wmnet to then clone it to db1283.eqiad.wmnet - marostegui@cumin1003 * 10:01 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy2002.codfw.wmnet * 10:01 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy4004.ulsfo.wmnet * 10:01 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy3002.esams.wmnet * 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1169: Depool db1169.eqiad.wmnet to then clone it to db1283.eqiad.wmnet - marostegui@cumin1003 * 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1169.eqiad.wmnet onto db1283.eqiad.wmnet * 10:01 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy5004.eqsin.wmnet * 09:59 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy4003.ulsfo.wmnet * 09:59 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy4004.ulsfo.wmnet * 09:59 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy5003.eqsin.wmnet * 09:59 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy7001.magru.wmnet * 09:59 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy5004.eqsin.wmnet * 09:58 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy7002.magru.wmnet * 09:57 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2031.codfw.wmnet * 09:57 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy6001.drmrs.wmnet * 09:57 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy6002.drmrs.wmnet * 09:54 filippo@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for 10 hosts * 09:54 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast4006.wikimedia.org * 09:53 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy6001.drmrs.wmnet * 09:53 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy6002.drmrs.wmnet * 09:53 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host releases1003.eqiad.wmnet * 09:53 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy7001.magru.wmnet * 09:53 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host people2004.codfw.wmnet * 09:52 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host people1005.eqiad.wmnet * 09:52 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host tcp-proxy7002.magru.wmnet * 09:50 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host releases2003.codfw.wmnet * 09:49 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host releases2003.codfw.wmnet * 09:49 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host releases1003.eqiad.wmnet * 09:49 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host people2004.codfw.wmnet * 09:48 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host people1005.eqiad.wmnet * 09:46 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2040.codfw.wmnet * 09:46 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2040.codfw.wmnet * 09:43 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1901.eqiad.wmnet on all recursors * 09:42 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1901.eqiad.wmnet on all recursors * 09:42 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:42 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:42 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2040.codfw.wmnet * 09:40 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp2005.wikimedia.org * 09:36 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp2005.wikimedia.org * 09:31 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:31 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1901.eqiad.wmnet * 09:31 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1901.eqiad.wmnet * 09:31 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:31 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1901.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 09:31 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1901.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 09:29 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2040.codfw.wmnet * 09:29 slyngshede@dns1004: END - running authdns-update * 09:28 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2039.codfw.wmnet * 09:28 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2039.codfw.wmnet * 09:27 slyngshede@dns1004: START - running authdns-update * 09:26 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test2005.wikimedia.org * 09:22 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2039.codfw.wmnet * 09:22 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test2005.wikimedia.org * 09:22 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp1005.wikimedia.org * 09:21 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:19 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2039.codfw.wmnet * 09:19 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2038.codfw.wmnet * 09:18 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1287: Pool back * 09:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2038.codfw.wmnet * 09:18 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp1005.wikimedia.org * 09:18 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test1005.wikimedia.org * 09:17 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1901.eqiad.wmnet * 09:15 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db1901.eqiad.wmnet * 09:15 fceratto@cumin1003: END (ERROR) - Cookbook sre.dns.netbox (exit_code=97) * 09:14 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test1005.wikimedia.org * 09:13 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2038.codfw.wmnet * 09:13 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:13 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1901.eqiad.wmnet * 09:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast4006.wikimedia.org * 09:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1288: Pool back * 09:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast5005.wikimedia.org * 09:03 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2038.codfw.wmnet * 08:57 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2037.codfw.wmnet * 08:57 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast5005.wikimedia.org * 08:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2037.codfw.wmnet * 08:56 filippo@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 10 hosts * 08:52 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2037.codfw.wmnet * 08:51 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1289: Pool back * 08:35 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudidp2001-dev.codfw.wmnet * 08:34 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1002-dev.eqiad.wmnet * 08:33 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1287: Pool back * 08:33 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1287 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96150 and previous config saved to /var/cache/conftool/dbconfig/20260818-083311-marostegui.json * 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1001-dev.eqiad.wmnet * 08:31 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudidp2001-dev.codfw.wmnet * 08:30 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1002-dev.eqiad.wmnet * 08:30 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2037.codfw.wmnet * 08:29 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1001-dev.eqiad.wmnet * 08:28 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2036.codfw.wmnet * 08:28 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2036.codfw.wmnet * 08:25 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1288: Pool back * 08:23 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 08:23 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1181.eqiad.wmnet with OS bookworm * 08:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2036.codfw.wmnet * 08:22 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1288 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96148 and previous config saved to /var/cache/conftool/dbconfig/20260818-082234-marostegui.json * 08:20 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2036.codfw.wmnet * 08:18 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2035.codfw.wmnet * 08:18 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudvirt1057.eqiad.wmnet with OS trixie * 08:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2035.codfw.wmnet * 08:13 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2035.codfw.wmnet * 08:06 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2035.codfw.wmnet * 08:05 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1289: Pool back * 08:05 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1289 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96145 and previous config saved to /var/cache/conftool/dbconfig/20260818-080531-marostegui.json * 07:54 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 07:51 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326439{{!}}Fix "mathjax_ignore" handling around forcemathmode attribute (T434686)]] (duration: 13m 15s) * 07:48 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 07:47 krinkle@deploy1003: krinkle: Continuing with deployment * 07:44 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ganeti2046.codfw.wmnet with OS bookworm * 07:40 krinkle@deploy1003: krinkle: Backport for [[gerrit:1326439{{!}}Fix "mathjax_ignore" handling around forcemathmode attribute (T434686)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:39 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 07:38 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1326439{{!}}Fix "mathjax_ignore" handling around forcemathmode attribute (T434686)]] * 07:34 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 07:31 samwilson@deploy1003: Finished scap sync-world: Backport for [[gerrit:701016{{!}}InitialiseSettings and -labs: Remove redundant feature flag $wgWikisourceEnableOcr (T285311)]] (duration: 07m 47s) * 07:30 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1048.eqiad.wmnet * 07:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1048.eqiad.wmnet * 07:29 XioNoX: add gnmic 0.47.0 to bookworm and trixie reprepro * 07:28 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ganeti2046.codfw.wmnet with reason: host reimage * 07:27 samwilson@deploy1003: samwilson: Continuing with deployment * 07:25 samwilson@deploy1003: samwilson: Backport for [[gerrit:701016{{!}}InitialiseSettings and -labs: Remove redundant feature flag $wgWikisourceEnableOcr (T285311)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:25 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T435157|T435157]] * 07:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1048.eqiad.wmnet * 07:24 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ganeti2046.codfw.wmnet with reason: host reimage * 07:23 samwilson@deploy1003: Started scap sync-world: Backport for [[gerrit:701016{{!}}InitialiseSettings and -labs: Remove redundant feature flag $wgWikisourceEnableOcr (T285311)]] * 07:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1198.eqiad.wmnet with OS bookworm * 07:18 samwilson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326450{{!}}InitialiseSettings.php: Enable Bulk OCR on pawikisource (T434648)]] (duration: 12m 17s) * 07:17 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 07:11 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1048.eqiad.wmnet * 07:11 samwilson@deploy1003: samwilson: Continuing with deployment * 07:11 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ganeti2046.codfw.wmnet with OS bookworm * 07:10 samwilson@deploy1003: samwilson: Backport for [[gerrit:1326450{{!}}InitialiseSettings.php: Enable Bulk OCR on pawikisource (T434648)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1057.eqiad.wmnet with OS trixie * 07:05 samwilson@deploy1003: Started scap sync-world: Backport for [[gerrit:1326450{{!}}InitialiseSettings.php: Enable Bulk OCR on pawikisource (T434648)]] * 07:02 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1198.eqiad.wmnet with reason: host reimage * 07:01 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1056.eqiad.wmnet with OS trixie * 07:01 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1056.eqiad.wmnet with OS trixie * 07:00 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1056.eqiad.wmnet with OS trixie * 07:00 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1056.eqiad.wmnet with OS trixie * 06:59 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudvirt1055.eqiad.wmnet with OS trixie * 06:58 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1198.eqiad.wmnet with reason: host reimage * 06:53 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1055.eqiad.wmnet with OS trixie * 06:52 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudvirt1054.eqiad.wmnet with OS trixie * 06:44 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1194.eqiad.wmnet with OS bookworm * 06:44 XioNoX: upgrade eqsin gnmic to 0.47.0 * 06:43 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1198.eqiad.wmnet with OS bookworm * 06:41 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1054.eqiad.wmnet with OS trixie * 06:09 arnaudb@cumin1003: END (PASS) - Cookbook sre.gerrit.restart-gerrit (exit_code=0) Restarting Gerrit on gerrit2002 * 06:06 arnaudb@cumin1003: START - Cookbook sre.gerrit.restart-gerrit Restarting Gerrit on gerrit2002 * 06:06 arnaudb@cumin1003: END (PASS) - Cookbook sre.gerrit.restart-gerrit (exit_code=0) Restarting Gerrit on gerrit1003 * 06:04 arnaudb@cumin1003: START - Cookbook sre.gerrit.restart-gerrit Restarting Gerrit on gerrit1003 * 06:02 arnaudb@cumin1003: END (PASS) - Cookbook sre.gerrit.restart-gerrit (exit_code=0) Restarting Gerrit on gerrit2003 * 06:00 arnaudb@cumin1003: START - Cookbook sre.gerrit.restart-gerrit Restarting Gerrit on gerrit2003 * 05:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1155.eqiad.wmnet with OS bookworm * 05:38 arnaudb: updating prometheusBearerToken on gerrit * 05:28 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1155.eqiad.wmnet with reason: host reimage * 05:23 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1155.eqiad.wmnet with reason: host reimage * 05:06 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1155.eqiad.wmnet with OS bookworm * 04:57 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1144.eqiad.wmnet with OS bookworm * 04:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1144.eqiad.wmnet with reason: host reimage * 04:29 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1144.eqiad.wmnet with reason: host reimage * 04:14 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1144.eqiad.wmnet with OS bookworm * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.13 (duration: 02m 23s) * 03:45 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1197.eqiad.wmnet with OS bookworm * 03:41 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1165.eqiad.wmnet with OS bookworm * 03:38 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.16 refs [[phab:T430835|T430835]] (duration: 34m 43s) * 03:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1164.eqiad.wmnet with OS bookworm * 03:35 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1196.eqiad.wmnet with OS bookworm * 03:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1163.eqiad.wmnet with OS bookworm * 03:22 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1197.eqiad.wmnet with reason: host reimage * 03:18 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1165.eqiad.wmnet with reason: host reimage * 03:15 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1196.eqiad.wmnet with reason: host reimage * 03:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1164.eqiad.wmnet with reason: host reimage * 03:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1163.eqiad.wmnet with reason: host reimage * 03:05 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1165.eqiad.wmnet with reason: host reimage * 03:04 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1164.eqiad.wmnet with reason: host reimage * 03:04 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1197.eqiad.wmnet with reason: host reimage * 03:04 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1196.eqiad.wmnet with reason: host reimage * 03:04 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1163.eqiad.wmnet with reason: host reimage * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.16 refs [[phab:T430835|T430835]] * 02:50 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1197.eqiad.wmnet with OS bookworm * 02:49 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1196.eqiad.wmnet with OS bookworm * 02:49 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1165.eqiad.wmnet with OS bookworm * 02:49 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1164.eqiad.wmnet with OS bookworm * 02:48 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1163.eqiad.wmnet with OS bookworm * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 46s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:15 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-codfw: Set storage compatability to NONE — [[phab:T433026|T433026]] - eevans@cumin1003 * 00:38 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-codfw: Set storage compatability to NONE — [[phab:T433026|T433026]] - eevans@cumin1003 * 00:11 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326373{{!}}PersonalDashboard: add newly renamed *ReviewChangesMlModel setting (T422148)]] (duration: 07m 07s) * 00:09 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-eqiad: Set storage compatability to NONE — [[phab:T433026|T433026]] - eevans@cumin1003 * 00:07 musikanimal@deploy1003: musikanimal: Continuing with deployment * 00:06 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1326373{{!}}PersonalDashboard: add newly renamed *ReviewChangesMlModel setting (T422148)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:04 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1326373{{!}}PersonalDashboard: add newly renamed *ReviewChangesMlModel setting (T422148)]] == 2026-08-17 == * 23:30 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-eqiad: Set storage compatability to NONE — [[phab:T433026|T433026]] - eevans@cumin1003 * 23:05 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-codfw: Set storage compatability to UPGRADING — [[phab:T433026|T433026]] - eevans@cumin1003 * 22:28 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-codfw: Set storage compatability to UPGRADING — [[phab:T433026|T433026]] - eevans@cumin1003 * 21:46 logmsgbot: jforrester Deployed security patch for [[phab:T435085|T435085]] * 21:39 swfrench@deploy1003: mwscript-k8s job started: purgeList.php # [[phab:T432412|T432412]] * 21:37 maryum: Undeploy security fix for [[phab:T433020|T433020]] * 21:23 maryum: Deployed security fix for [[phab:T433020|T433020]] * 21:21 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 21:21 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 21:14 maryum: Deployed security fix for [[phab:T434967|T434967]] * 20:58 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-eqiad: Set storage compatability to UPGRADING — [[phab:T433026|T433026]] - eevans@cumin1003 * 20:40 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324320{{!}}[itwiki/slwiki/tgwiki] Remove temporary Wikipedia 25 logos permanently (already reverted) (T414265 T414320 T415307)]] (duration: 06m 55s) * 20:40 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 20:39 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 20:36 cjming@deploy1003: cjming, superpes: Continuing with deployment * 20:35 cjming@deploy1003: cjming, superpes: Backport for [[gerrit:1324320{{!}}[itwiki/slwiki/tgwiki] Remove temporary Wikipedia 25 logos permanently (already reverted) (T414265 T414320 T415307)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:33 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1324320{{!}}[itwiki/slwiki/tgwiki] Remove temporary Wikipedia 25 logos permanently (already reverted) (T414265 T414320 T415307)]] * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ttmserver-test: apply * 20:30 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326361{{!}}Remove escaped paths in app site association file (T432412)]] (duration: 13m 28s) * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ttmserver-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-toolhub-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-toolhub-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-toolhub-test: apply * 20:30 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-toolhub-test: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-toolhub: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-toolhub: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-toolhub: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-toolhub: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-test: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-test: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 20:29 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 20:28 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ipoid-test: apply * 20:27 inflatador: bking@deploy1003 `charlie --services_dir dse-k8s-services -s opensearch-* -e dse-k8s-* apply` [[phab:T435125|T435125]] * 20:27 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-apifeatureusage-test: apply * 20:27 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-apifeatureusage-test: apply * 20:27 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-apifeatureusage-test: apply * 20:27 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-apifeatureusage-test: apply * 20:27 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-apifeatureusage: apply * 20:26 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-apifeatureusage: apply * 20:26 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-apifeatureusage: apply * 20:26 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-apifeatureusage: apply * 20:26 cjming@deploy1003: cjming, tsev: Continuing with deployment * 20:24 inflatador: bking@deploy1003 `charlie --services_dir dse-k8s-services -s opensearch-* -e dse-k8s-*` [[phab:T435125|T435125]] * 20:21 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-eqiad: Set storage compatability to UPGRADING — [[phab:T433026|T433026]] - eevans@cumin1003 * 20:19 cjming@deploy1003: cjming, tsev: Backport for [[gerrit:1326361{{!}}Remove escaped paths in app site association file (T432412)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:17 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1326361{{!}}Remove escaped paths in app site association file (T432412)]] * 20:15 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326077{{!}}InstrumentConstructiveEdits: anchor all runs to the nearest `interval` (T431493)]] (duration: 06m 25s) * 20:11 cjming@deploy1003: cjming: Continuing with deployment * 20:10 cjming@deploy1003: cjming: Backport for [[gerrit:1326077{{!}}InstrumentConstructiveEdits: anchor all runs to the nearest `interval` (T431493)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:08 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1326077{{!}}InstrumentConstructiveEdits: anchor all runs to the nearest `interval` (T431493)]] * 20:06 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1160.eqiad.wmnet with OS bookworm * 20:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1162.eqiad.wmnet with OS bookworm * 19:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1184.eqiad.wmnet with OS bookworm * 19:55 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1161.eqiad.wmnet with OS bookworm * 19:49 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1195.eqiad.wmnet with OS bookworm * 19:49 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-test: apply * 19:49 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-test: apply * 19:44 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1160.eqiad.wmnet with reason: host reimage * 19:41 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-codfw: Upgrade to Java 17 — [[phab:T433026|T433026]] - eevans@cumin1003 * 19:39 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1162.eqiad.wmnet with reason: host reimage * 19:36 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-eqsin and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 19:36 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-test: apply * 19:36 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1184.eqiad.wmnet with reason: host reimage * 19:32 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1161.eqiad.wmnet with reason: host reimage * 19:29 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1195.eqiad.wmnet with reason: host reimage * 19:26 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1161.eqiad.wmnet with reason: host reimage * 19:26 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1160.eqiad.wmnet with reason: host reimage * 19:26 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1184.eqiad.wmnet with reason: host reimage * 19:26 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1162.eqiad.wmnet with reason: host reimage * 19:25 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1195.eqiad.wmnet with reason: host reimage * 19:19 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-test: apply * 19:11 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1195.eqiad.wmnet with OS bookworm * 19:10 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1194.eqiad.wmnet with OS bookworm * 19:10 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1184.eqiad.wmnet with OS bookworm * 19:10 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1162.eqiad.wmnet with OS bookworm * 19:10 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1161.eqiad.wmnet with OS bookworm * 19:10 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1160.eqiad.wmnet with OS bookworm * 19:10 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 19:09 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 19:03 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-codfw: Upgrade to Java 17 — [[phab:T433026|T433026]] - eevans@cumin1003 * 18:58 dancy@deploy1003: Finished scap sync-world: testing [[phab:T375514|T375514]] (duration: 03m 13s) * 18:55 dancy@deploy1003: Started scap sync-world: testing [[phab:T375514|T375514]] * 18:55 dwisehaupt@dns1006: END - running authdns-update * 18:54 dancy@deploy1003: Installation of scap version "4.281.1" completed for 3 hosts * 18:54 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2009.codfw.wmnet * 18:53 dwisehaupt@dns1006: START - running authdns-update * 18:52 dancy@deploy1003: Installing scap version "4.281.1" for 3 host(s) * 18:47 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2009.codfw.wmnet * 18:41 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2008.codfw.wmnet * 18:34 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2008.codfw.wmnet * 18:30 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2007.codfw.wmnet * 18:23 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2007.codfw.wmnet * 18:16 dwisehaupt@dns1005: END - running authdns-update * 18:14 dwisehaupt@dns1005: START - running authdns-update * 18:04 swfrench@deploy1003: Finished scap sync-world: Deploy "Point Test Wiki to new docroot" - [[phab:T432412|T432412]] (duration: 26m 03s) * 18:00 swfrench@deploy1003: swfrench: Continuing with deployment * 17:51 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1057.eqiad.wmnet with OS trixie * 17:47 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1180.eqiad.wmnet with OS bookworm * 17:39 swfrench@deploy1003: swfrench: Deploy "Point Test Wiki to new docroot" - [[phab:T432412|T432412]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:39 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-eqsin and A:cp - 9.2.15 upgrade ([[phab:T434478|T434478]]) * 17:38 swfrench@deploy1003: Started scap sync-world: Deploy "Point Test Wiki to new docroot" - [[phab:T432412|T432412]] * 17:36 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1159.eqiad.wmnet with OS bookworm * 17:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1158.eqiad.wmnet with OS bookworm * 17:30 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-eqiad: Upgrade to Java 17 — [[phab:T433026|T433026]] - eevans@cumin1003 * 17:30 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-ulsfo or A:cp-drmrs and A:cp - 9.2.15 upgrade ([[phab:T434620|T434620]]) * 17:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1157.eqiad.wmnet with OS bookworm * 17:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1180.eqiad.wmnet with reason: host reimage * 17:22 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 17:21 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1193.eqiad.wmnet with OS bookworm * 17:21 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 17:21 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 17:20 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 17:20 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 17:20 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1183.eqiad.wmnet with OS bookworm * 17:18 swfrench@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 17:18 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 17:17 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1180.eqiad.wmnet with reason: host reimage * 17:17 swfrench@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 17:17 swfrench@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 17:16 swfrench@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 17:16 swfrench@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 17:15 swfrench@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 17:15 swfrench@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 17:15 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1192.eqiad.wmnet with OS bookworm * 17:14 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1159.eqiad.wmnet with reason: host reimage * 17:14 swfrench@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 17:13 swfrench@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 17:12 swfrench@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 17:09 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1158.eqiad.wmnet with reason: host reimage * 17:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1157.eqiad.wmnet with reason: host reimage * 17:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1180 * 17:02 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1180 * 17:01 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica-eqsin and A:liberica * 17:01 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1193.eqiad.wmnet with reason: host reimage * 16:57 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1183.eqiad.wmnet with reason: host reimage * 16:56 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1054.eqiad.wmnet with OS trixie * 16:55 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1056.eqiad.wmnet with OS trixie * 16:54 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1193.eqiad.wmnet with reason: host reimage * 16:54 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1192.eqiad.wmnet with reason: host reimage * 16:51 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-eqiad: Upgrade to Java 17 — [[phab:T433026|T433026]] - eevans@cumin1003 * 16:50 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1180 * 16:50 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1180.eqiad.wmnet 17.36.64.10.in-addr.arpa 7.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:50 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1180.eqiad.wmnet 17.36.64.10.in-addr.arpa 7.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:50 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:50 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1180 - btullis@cumin1003" * 16:50 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1180 - btullis@cumin1003" * 16:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1158.eqiad.wmnet with reason: host reimage * 16:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1157.eqiad.wmnet with reason: host reimage * 16:49 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica-eqsin and A:liberica * 16:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1183.eqiad.wmnet with reason: host reimage * 16:48 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1159.eqiad.wmnet with reason: host reimage * 16:47 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1192.eqiad.wmnet with reason: host reimage * 16:46 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns4004.wikimedia.org * 16:46 sukhe@dns1004: END - running authdns-update * 16:44 sukhe@dns1004: START - running authdns-update * 16:44 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns4004.wikimedia.org,service=authdns-update * 16:43 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns4004.wikimedia.org with OS trixie * 16:39 btullis@cumin1003: START - Cookbook sre.dns.netbox * 16:39 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1180 * 16:39 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1193.eqiad.wmnet with OS bookworm * 16:39 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1180.eqiad.wmnet with OS bookworm * 16:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1192.eqiad.wmnet with OS bookworm * 16:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1183.eqiad.wmnet with OS bookworm * 16:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1159.eqiad.wmnet with OS bookworm * 16:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1158.eqiad.wmnet with OS bookworm * 16:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1157.eqiad.wmnet with OS bookworm * 16:31 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1057.eqiad.wmnet with OS trixie * 16:30 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1057.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 16:29 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica-drmrs and A:liberica * 16:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1154.eqiad.wmnet with OS bookworm * 16:23 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1057.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 16:23 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1055.eqiad.wmnet with OS trixie * 16:22 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1057 * 16:22 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1057 * 16:21 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:21 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1057] - vriley@cumin1003" * 16:21 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1057] - vriley@cumin1003" * 16:19 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica-drmrs and A:liberica * 16:17 vriley@cumin1003: START - Cookbook sre.dns.netbox * 16:16 phuedx@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics-external: apply * 16:15 phuedx@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics-external: apply * 16:13 phuedx@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics-external: apply * 16:12 phuedx@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics-external: apply * 16:11 phuedx@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics-external: apply * 16:09 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics-external: apply * 16:09 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1176.eqiad.wmnet with OS bookworm * 16:07 btullis@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 16:06 btullis@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 16:05 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1191.eqiad.wmnet with OS bookworm * 16:03 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1154.eqiad.wmnet with reason: host reimage * 16:03 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching aqs[2002-2012].codfw.wmnet,aqs[1017-1027].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433026|T433026]] - eevans@cumin1003 * 15:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1190.eqiad.wmnet with OS bookworm * 15:59 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1154.eqiad.wmnet with reason: host reimage * 15:55 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-debug: apply * 15:55 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-debug: apply * 15:55 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-debug: apply * 15:55 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/mw-debug: apply * 15:53 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns4004.wikimedia.org with reason: host reimage * 15:50 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns4004.wikimedia.org with reason: host reimage * 15:46 btullis@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 15:46 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1176.eqiad.wmnet with reason: host reimage * 15:45 btullis@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 15:43 btullis@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 15:42 btullis@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 15:42 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1191.eqiad.wmnet with reason: host reimage * 15:39 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1190.eqiad.wmnet with reason: host reimage * 15:36 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1054.eqiad.wmnet with OS trixie * 15:35 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:35 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1056.eqiad.wmnet with OS trixie * 15:35 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:34 moritzm: failover Ganeti master in eqiad to ganeti1046 * 15:34 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1176.eqiad.wmnet with reason: host reimage * 15:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1191.eqiad.wmnet with reason: host reimage * 15:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1190.eqiad.wmnet with reason: host reimage * 15:31 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns2006.wikimedia.org * 15:31 sukhe@dns1004: END - running authdns-update * 15:29 sukhe@dns1004: START - running authdns-update * 15:29 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns2006.wikimedia.org,service=authdns-update * 15:29 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns2006.wikimedia.org * 15:29 sukhe@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns2006.wikimedia.org * 15:26 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica-ulsfo and A:liberica * 15:26 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns1006.wikimedia.org * 15:26 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:25 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns2006.wikimedia.org with OS trixie * 15:25 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1056 * 15:25 sukhe@dns1004: END - running authdns-update * 15:23 sukhe@dns1004: START - running authdns-update * 15:23 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns1006.wikimedia.org,service=authdns-update * 15:23 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns1006.wikimedia.org * 15:23 sukhe@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns1006.wikimedia.org * 15:20 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1056 * 15:20 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:20 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1056~] - vriley@cumin1003" * 15:19 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1056~] - vriley@cumin1003" * 15:19 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns1006.wikimedia.org with OS trixie * 15:19 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns4004.wikimedia.org with OS trixie * 15:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1190.eqiad.wmnet with OS bookworm * 15:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1191.eqiad.wmnet with OS bookworm * 15:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1176.eqiad.wmnet with OS bookworm * 15:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1154.eqiad.wmnet with OS bookworm * 15:16 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica-ulsfo and A:liberica * 15:13 vriley@cumin1003: START - Cookbook sre.dns.netbox * 15:13 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host dns4004.wikimedia.org with OS trixie * 15:12 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1188.eqiad.wmnet with OS bookworm * 15:09 taavi@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318203{{!}}Undeploy WP25EasterEggs (II) (T418134)]] (duration: 08m 56s) * 15:06 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica-magru and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 15:06 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1051.eqiad.wmnet with OS trixie * 15:06 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 15:05 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 15:05 taavi@deploy1003: taavi: Continuing with deployment * 15:04 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-ncredir (exit_code=0) rolling reboot on A:ncredir and A:ncredir * 15:04 taavi@deploy1003: taavi: Backport for [[gerrit:1318203{{!}}Undeploy WP25EasterEggs (II) (T418134)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:03 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1055.eqiad.wmnet with OS trixie * 15:02 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:00 taavi@deploy1003: Started scap sync-world: Backport for [[gerrit:1318203{{!}}Undeploy WP25EasterEggs (II) (T418134)]] * 14:59 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1047.eqiad.wmnet * 14:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1047.eqiad.wmnet * 14:58 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy (exit_code=0) rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 14:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-codfw * 14:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp2001.codfw.wmnet * 14:57 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp2001.codfw.wmnet * 14:57 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica-magru and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 14:57 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:56 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1055 * 14:56 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1055 * 14:55 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns2006.wikimedia.org with reason: host reimage * 14:55 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:55 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1055] - vriley@cumin1003" * 14:55 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1055] - vriley@cumin1003" * 14:54 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1047.eqiad.wmnet * 14:52 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1188.eqiad.wmnet with reason: host reimage * 14:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp2001.codfw.wmnet * 14:51 vriley@cumin1003: START - Cookbook sre.dns.netbox * 14:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp2001.codfw.wmnet * 14:50 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2016.codfw.wmnet * 14:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2016.codfw.wmnet * 14:50 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1189.eqiad.wmnet with OS bookworm * 14:48 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1188.eqiad.wmnet with reason: host reimage * 14:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1051.eqiad.wmnet with reason: host reimage * 14:46 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1182.eqiad.wmnet with OS bookworm * 14:45 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief2002.codfw.wmnet * 14:44 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns1006.wikimedia.org with reason: host reimage * 14:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2016.codfw.wmnet * 14:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1143.eqiad.wmnet with OS bookworm * 14:43 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2016.codfw.wmnet * 14:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2318-2331].codfw.wmnet * 14:43 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2318-2331].codfw.wmnet * 14:42 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching aqs[2002-2012].codfw.wmnet,aqs[1017-1027].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433026|T433026]] - eevans@cumin1003 * 14:41 cgoubert@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326306{{!}}Add placeholder $wmgRedisLockPassword (T366938 T427999)]] (duration: 06m 56s) * 14:41 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief2002.codfw.wmnet * 14:40 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief1002.eqiad.wmnet * 14:39 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1051.eqiad.wmnet with reason: host reimage * 14:38 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns2006.wikimedia.org with reason: host reimage * 14:37 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns1006.wikimedia.org with reason: host reimage * 14:37 cgoubert@deploy1003: cgoubert: Continuing with deployment * 14:36 cgoubert@deploy1003: cgoubert: Backport for [[gerrit:1326306{{!}}Add placeholder $wmgRedisLockPassword (T366938 T427999)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:36 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief1002.eqiad.wmnet * 14:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2318-2331].codfw.wmnet * 14:34 cgoubert@deploy1003: Started scap sync-world: Backport for [[gerrit:1326306{{!}}Add placeholder $wmgRedisLockPassword (T366938 T427999)]] * 14:29 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2318-2331].codfw.wmnet * 14:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2304-2317].codfw.wmnet * 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2304-2317].codfw.wmnet * 14:28 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1054 * 14:27 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1054 * 14:27 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:27 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1054] - vriley@cumin1003" * 14:27 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1054] - vriley@cumin1003" * 14:26 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1189.eqiad.wmnet with reason: host reimage * 14:26 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test2001.codfw.wmnet * 14:25 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test1001.eqiad.wmnet * 14:25 claime: Deploying wmgRedisLockPassword - [[phab:T366938|T366938]] [[phab:T427999|T427999]] * 14:24 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1051.eqiad.wmnet with OS trixie * 14:23 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 14:22 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1182.eqiad.wmnet with reason: host reimage * 14:22 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1051.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:22 vriley@cumin1003: START - Cookbook sre.dns.netbox * 14:22 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test2001.codfw.wmnet * 14:21 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test1001.eqiad.wmnet * 14:21 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 14:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2304-2317].codfw.wmnet * 14:19 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns4004.wikimedia.org with OS trixie * 14:19 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns2006.wikimedia.org with OS trixie * 14:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1143.eqiad.wmnet with reason: host reimage * 14:19 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns1006.wikimedia.org with OS trixie * 14:17 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1189.eqiad.wmnet with reason: host reimage * 14:15 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1182.eqiad.wmnet with reason: host reimage * 14:14 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1143.eqiad.wmnet with reason: host reimage * 14:13 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1051.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:12 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2304-2317].codfw.wmnet * 14:12 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2290-2303].codfw.wmnet * 14:12 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1049.eqiad.wmnet with OS trixie * 14:12 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2290-2303].codfw.wmnet * 14:12 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1051 * 14:11 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1051 * 14:11 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1170.eqiad.wmnet onto db1284.eqiad.wmnet * 14:11 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1170: Pool db1170.eqiad.wmnet in after cloning * 14:09 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 14:09 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:09 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1051] - vriley@cumin1003" * 14:09 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1051] - vriley@cumin1003" * 14:06 klausman@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:04 vriley@cumin1003: START - Cookbook sre.dns.netbox * 14:04 klausman@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2290-2303].codfw.wmnet * 14:02 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1189.eqiad.wmnet with OS bookworm * 14:02 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1188.eqiad.wmnet with OS bookworm * 14:01 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 14:00 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1182.eqiad.wmnet with OS bookworm * 14:00 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1143.eqiad.wmnet with OS bookworm * 13:59 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 13:57 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply * 13:57 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply * 13:56 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply * 13:56 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-ulsfo or A:cp-drmrs and A:cp - 9.2.15 upgrade ([[phab:T434620|T434620]]) * 13:56 cjd91: sudo -i cookbook sre.cdn.roll-upgrade-ats --query 'A:cp-ulsfo or A:cp-drmrs' --task-id [[phab:T434620|T434620]] --reason '9.2.15 upgrade' * 13:56 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply * 13:55 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply * 13:55 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply * 13:54 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 13:54 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 13:54 phuedx@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:53 phuedx@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics-external: apply * 13:52 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1049.eqiad.wmnet with reason: host reimage * 13:51 phuedx@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2290-2303].codfw.wmnet * 13:50 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1047.eqiad.wmnet * 13:50 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2276-2289].codfw.wmnet * 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2276-2289].codfw.wmnet * 13:49 phuedx@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics-external: apply * 13:49 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1046.eqiad.wmnet * 13:49 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1049.eqiad.wmnet with reason: host reimage * 13:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1046.eqiad.wmnet * 13:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-ncredir rolling reboot on A:ncredir and A:ncredir * 13:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 13:46 phuedx@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:44 phuedx@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics-external: apply * 13:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1046.eqiad.wmnet * 13:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2276-2289].codfw.wmnet * 13:36 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1046.eqiad.wmnet * 13:36 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1045.eqiad.wmnet * 13:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1045.eqiad.wmnet * 13:34 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2276-2289].codfw.wmnet * 13:34 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1049.eqiad.wmnet with OS trixie * 13:34 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2262-2275].codfw.wmnet * 13:34 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2262-2275].codfw.wmnet * 13:32 Lucas_WMDE: UTC afternoon backport+config window done * 13:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1045.eqiad.wmnet * 13:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2262-2275].codfw.wmnet * 13:26 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1049.eqiad.wmnet with OS trixie * 13:26 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1049.eqiad.wmnet with OS trixie * 13:26 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1170: Pool db1170.eqiad.wmnet in after cloning * 13:23 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1049.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 13:23 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1045.eqiad.wmnet * 13:21 atsukoito: manually done sudo -i docker-registryctl --debug delete-tags 'docker-registry.discovery.wmnet/repos/data-engineering/airflow-dags:airflow-3.3.0-py3.11-2026-08-17-*' to remove incorrect tags * 13:19 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311141{{!}}viwiki: Set `noindex,nofollow` for User and User talk (T432311)]] (duration: 11m 11s) * 13:15 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2262-2275].codfw.wmnet * 13:14 lucaswerkmeister-wmde@deploy1003: ndkdd, lucaswerkmeister-wmde: Continuing with deployment * 13:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2248-2261].codfw.wmnet * 13:14 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2248-2261].codfw.wmnet * 13:13 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1049.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 13:12 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1049 * 13:11 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1049 * 13:10 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:10 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1049] - vriley@cumin1003" * 13:10 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1049] - vriley@cumin1003" * 13:10 lucaswerkmeister-wmde@deploy1003: ndkdd, lucaswerkmeister-wmde: Backport for [[gerrit:1311141{{!}}viwiki: Set `noindex,nofollow` for User and User talk (T432311)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1037.eqiad.wmnet * 13:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1037.eqiad.wmnet * 13:08 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1311141{{!}}viwiki: Set `noindex,nofollow` for User and User talk (T432311)]] * 13:06 vriley@cumin1003: START - Cookbook sre.dns.netbox * 13:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2248-2261].codfw.wmnet * 13:00 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1037.eqiad.wmnet * 12:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2248-2261].codfw.wmnet * 12:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2204-2215,2242-2243].codfw.wmnet * 12:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2204-2215,2242-2243].codfw.wmnet * 12:49 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324832{{!}}Migrate $wgFlaggedRevsTags from flaggedrevs.php to ext-FlaggedRevs.php]] (duration: 14m 02s) * 12:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint2001.codfw.wmnet * 12:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2204-2215,2242-2243].codfw.wmnet * 12:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint2001.codfw.wmnet * 12:41 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1037.eqiad.wmnet * 12:40 ladsgroup@deploy1003: Rolling back deployment * 12:37 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1324832{{!}}Migrate $wgFlaggedRevsTags from flaggedrevs.php to ext-FlaggedRevs.php]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:37 seanleong-wmde: Finished populateSitesTable for [bolwiki] ([[[phab:T429955|T429955]]]) * 12:35 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1324832{{!}}Migrate $wgFlaggedRevsTags from flaggedrevs.php to ext-FlaggedRevs.php]] * 12:35 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2204-2215,2242-2243].codfw.wmnet * 12:34 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2190-2203].codfw.wmnet * 12:34 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2190-2203].codfw.wmnet * 12:32 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1028.eqiad.wmnet * 12:32 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1028.eqiad.wmnet * 12:31 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint1001.eqiad.wmnet * 12:30 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1170: Depool db1170.eqiad.wmnet to then clone it to db1284.eqiad.wmnet - marostegui@cumin1003 * 12:28 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint1001.eqiad.wmnet * 12:28 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1170: Depool db1170.eqiad.wmnet to then clone it to db1284.eqiad.wmnet - marostegui@cumin1003 * 12:27 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1170.eqiad.wmnet onto db1284.eqiad.wmnet * 12:27 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db2901.codfw.wmnet * 12:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2190-2203].codfw.wmnet * 12:27 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 12:26 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1028.eqiad.wmnet * 12:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2190-2203].codfw.wmnet * 12:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2172-2179,2184-2189].codfw.wmnet * 12:18 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2172-2179,2184-2189].codfw.wmnet * 12:13 seanleong-wmde@deploy1003: mwscript-k8s job started: foreachwikiindblist wikidataclient extensions/Wikibase/lib/maintenance/populateSitesTable.php --force-protocol https # [[phab:T429955|T429955]] * 12:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2172-2179,2184-2189].codfw.wmnet * 12:09 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1028.eqiad.wmnet * 12:03 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 12:02 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db2901.codfw.wmnet * 12:02 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db2901.codfw.wmnet * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1027.eqiad.wmnet * 12:02 fceratto@cumin1003: END (ERROR) - Cookbook sre.dns.netbox (exit_code=97) * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1027.eqiad.wmnet * 12:02 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 12:02 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db2901.codfw.wmnet * 12:01 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2172-2179,2184-2189].codfw.wmnet * 12:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2158-2171].codfw.wmnet * 12:01 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2158-2171].codfw.wmnet * 11:58 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db1902.eqiad.wmnet * 11:58 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 11:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1027.eqiad.wmnet * 11:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2158-2171].codfw.wmnet * 11:51 jayme: updated calico to v3.30.7 on wikikube eqiad - [[phab:T427400|T427400]] * 11:50 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1027.eqiad.wmnet * 11:45 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2158-2171].codfw.wmnet * 11:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2144-2157].codfw.wmnet * 11:44 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2144-2157].codfw.wmnet * 11:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1058.eqiad.wmnet * 11:43 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1058.eqiad.wmnet * 11:43 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 11:43 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 11:43 marostegui@cumin1003: Removing db1153 from zarcillo [[phab:T434638|T434638]] * 11:42 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1153.eqiad.wmnet * 11:42 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:42 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1153.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 11:42 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1172.eqiad.wmnet onto db1286.eqiad.wmnet * 11:42 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1172: Pool db1172.eqiad.wmnet in after cloning * 11:42 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1153.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 11:41 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1902.eqiad.wmnet * 11:41 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=97) for new host db2901.codfw.wmnet * 11:41 fceratto@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host db2901.codfw.wmnet with OS trixie * 11:38 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'. * 11:38 marostegui@cumin1003: START - Cookbook sre.dns.netbox * 11:37 marostegui@dns1004: END - running authdns-update * 11:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1058.eqiad.wmnet * 11:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2144-2157].codfw.wmnet * 11:35 marostegui@dns1004: START - running authdns-update * 11:32 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1153.eqiad.wmnet * 11:32 marostegui@cumin1003: START - Cookbook sre.mysql.decommission * 11:28 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326249{{!}}ImagePage: move TOC element below file link (T332644)]] (duration: 09m 56s) * 11:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2144-2157].codfw.wmnet * 11:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2130-2143].codfw.wmnet * 11:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2130-2143].codfw.wmnet * 11:26 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1058.eqiad.wmnet * 11:23 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 11:22 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'. * 11:22 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'. * 11:22 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'. * 11:22 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326249{{!}}ImagePage: move TOC element below file link (T332644)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1057.eqiad.wmnet * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1057.eqiad.wmnet * 11:21 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 11:20 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 11:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2130-2143].codfw.wmnet * 11:19 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326249{{!}}ImagePage: move TOC element below file link (T332644)]] * 11:18 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply * 11:17 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply * 11:17 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply * 11:16 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 11:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1057.eqiad.wmnet * 11:15 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 11:14 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 11:13 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 11:13 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db2901.codfw.wmnet with OS trixie * 11:12 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db2901.codfw.wmnet - fceratto@cumin1003" * 11:12 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db2901.codfw.wmnet - fceratto@cumin1003" * 11:12 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db2901.codfw.wmnet on all recursors * 11:12 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db2901.codfw.wmnet on all recursors * 11:12 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:12 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db2901.codfw.wmnet - fceratto@cumin1003" * 11:12 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2130-2143].codfw.wmnet * 11:11 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2107-2115,2124-2129].codfw.wmnet * 11:11 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2107-2115,2124-2129].codfw.wmnet * 11:11 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 11:11 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 11:11 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply * 11:11 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 11:10 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 11:06 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db2901.codfw.wmnet - fceratto@cumin1003" * 11:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2107-2115,2124-2129].codfw.wmnet * 10:57 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1172: Pool db1172.eqiad.wmnet in after cloning * 10:54 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2107-2115,2124-2129].codfw.wmnet * 10:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2078,2087-2095,2102-2106].codfw.wmnet * 10:54 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2078,2087-2095,2102-2106].codfw.wmnet * 10:47 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply * 10:46 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply * 10:46 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply * 10:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2078,2087-2095,2102-2106].codfw.wmnet * 10:45 blake@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply * 10:38 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1057.eqiad.wmnet * 10:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1056.eqiad.wmnet * 10:38 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1056.eqiad.wmnet * 10:37 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2078,2087-2095,2102-2106].codfw.wmnet * 10:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2061-2062,2064-2065,2067-2077].codfw.wmnet * 10:36 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2061-2062,2064-2065,2067-2077].codfw.wmnet * 10:32 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1056.eqiad.wmnet * 10:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2061-2062,2064-2065,2067-2077].codfw.wmnet * 10:25 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:25 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db2901.codfw.wmnet * 10:25 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db1902.eqiad.wmnet * 10:25 fceratto@cumin1003: END (ERROR) - Cookbook sre.dns.netbox (exit_code=97) * 10:24 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:24 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1902.eqiad.wmnet * 10:24 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=97) for new host db1902.eqiad.wmnet * 10:24 fceratto@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host db1902.eqiad.wmnet with OS trixie * 10:24 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=93) for new host db1903.eqiad.wmnet * 10:24 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 10:20 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1056.eqiad.wmnet * 10:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2061-2062,2064-2065,2067-2077].codfw.wmnet * 10:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2038-2039,2041-2042,2044,2046,2049-2051,2055-2060].codfw.wmnet * 10:18 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2038-2039,2041-2042,2044,2046,2049-2051,2055-2060].codfw.wmnet * 10:15 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1055.eqiad.wmnet * 10:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1055.eqiad.wmnet * 10:13 Amir1: mwscript-k8s --dblist=all -- purgeUserOptions.php --login-age 5 uls-preferences ([[phab:T406724|T406724]]) * 10:11 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:10 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1903.eqiad.wmnet on all recursors * 10:10 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1903.eqiad.wmnet on all recursors * 10:10 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2038-2039,2041-2042,2044,2046,2049-2051,2055-2060].codfw.wmnet * 10:10 moritzm: installing unzip security updates * 10:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1055.eqiad.wmnet * 10:08 fceratto@cumin1003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=97) for new host db1901.eqiad.wmnet * 10:08 fceratto@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host db1901.eqiad.wmnet with OS trixie * 10:08 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:08 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 10:08 fceratto@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1903.eqiad.wmnet - fceratto@cumin1003" * 10:07 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db1902.eqiad.wmnet with OS trixie * 10:07 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1902.eqiad.wmnet - fceratto@cumin1003" * 10:07 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1902.eqiad.wmnet - fceratto@cumin1003" * 10:04 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply * 10:04 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326227{{!}}Enable desktop/native lazy loading everywhere (T148047)]] (duration: 07m 13s) * 10:03 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1902.eqiad.wmnet on all recursors * 10:03 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1902.eqiad.wmnet on all recursors * 10:03 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:03 blake@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply * 10:01 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1903.eqiad.wmnet - fceratto@cumin1003" * 10:01 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2038-2039,2041-2042,2044,2046,2049-2051,2055-2060].codfw.wmnet * 10:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2002,2005-2006,2011-2015,2017-2018,2033-2037].codfw.wmnet * 10:01 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2002,2005-2006,2011-2015,2017-2018,2033-2037].codfw.wmnet * 10:00 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:00 fceratto@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 09:59 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 09:59 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1055.eqiad.wmnet * 09:58 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1326227{{!}}Enable desktop/native lazy loading everywhere (T148047)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:56 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1326227{{!}}Enable desktop/native lazy loading everywhere (T148047)]] * 09:53 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2002,2005-2006,2011-2015,2017-2018,2033-2037].codfw.wmnet * 09:50 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:48 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1903.eqiad.wmnet * 09:48 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:48 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1902.eqiad.wmnet * 09:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 09:44 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2002,2005-2006,2011-2015,2017-2018,2033-2037].codfw.wmnet * 09:43 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-codfw * 09:43 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1174.eqiad.wmnet onto db1288.eqiad.wmnet * 09:42 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1174: Pool db1174.eqiad.wmnet in after cloning * 09:40 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1172: Depool db1172.eqiad.wmnet to then clone it to db1286.eqiad.wmnet - marostegui@cumin1003 * 09:39 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1172: Depool db1172.eqiad.wmnet to then clone it to db1286.eqiad.wmnet - marostegui@cumin1003 * 09:39 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1172.eqiad.wmnet onto db1286.eqiad.wmnet * 09:33 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1175.eqiad.wmnet onto db1289.eqiad.wmnet * 09:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1175: Pool db1175.eqiad.wmnet in after cloning * 09:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1201.eqiad.wmnet onto db1287.eqiad.wmnet * 09:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1201: Pool db1201.eqiad.wmnet in after cloning * 09:28 fceratto@cumin1003: START - Cookbook sre.hosts.reimage for host db1901.eqiad.wmnet with OS trixie * 09:27 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:27 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:27 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1901.eqiad.wmnet on all recursors * 09:27 fceratto@cumin1003: START - Cookbook sre.dns.wipe-cache db1901.eqiad.wmnet on all recursors * 09:26 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:26 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:26 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" * 09:15 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:15 fceratto@cumin1003: START - Cookbook sre.ganeti.makevm for new host db1901.eqiad.wmnet * 09:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw1001.wikimedia.org with OS trixie * 08:57 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1174: Pool db1174.eqiad.wmnet in after cloning * 08:54 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1054.eqiad.wmnet * 08:54 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1054.eqiad.wmnet * 08:48 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1054.eqiad.wmnet * 08:47 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1175: Pool db1175.eqiad.wmnet in after cloning * 08:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 08:46 marostegui@cumin1003: Removing db1152 from zarcillo [[phab:T434480|T434480]] * 08:46 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1152.eqiad.wmnet * 08:46 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:46 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1152.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 08:46 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1201: Pool db1201.eqiad.wmnet in after cloning * 08:46 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1152.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 08:46 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1054.eqiad.wmnet * 08:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1053.eqiad.wmnet * 08:43 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1053.eqiad.wmnet * 08:42 marostegui@cumin1003: START - Cookbook sre.dns.netbox * 08:38 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage * 08:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1053.eqiad.wmnet * 08:36 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1152.eqiad.wmnet * 08:36 marostegui@cumin1003: START - Cookbook sre.mysql.decommission * 08:35 phuedx: UTC morning backport window done * 08:35 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1053.eqiad.wmnet * 08:34 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1035.eqiad.wmnet * 08:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1035.eqiad.wmnet * 08:34 phuedx@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324725{{!}}EventStreamConfig: Mark product_metrics.web_base and .web_base_with_ip as Test Kitchen streams (T429898 T430322)]], [[gerrit:1313923{{!}}EventStreamConfig: Remove unused web_ui_scroll* streams (T415370)]], [[gerrit:1325546{{!}}EventStreamConfig: Remove Watchlist click stream (T434790)]] (duration: 12m 42s) * 08:33 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1281: Pool back * 08:32 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage * 08:31 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on db2209.codfw.wmnet with reason: Maintenance * 08:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2209: Maintenance needed * 08:30 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2209: Maintenance needed * 08:26 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1035.eqiad.wmnet * 08:26 phuedx@deploy1003: bearloga, phuedx: Continuing with deployment * 08:23 phuedx@deploy1003: bearloga, phuedx: Backport for [[gerrit:1324725{{!}}EventStreamConfig: Mark product_metrics.web_base and .web_base_with_ip as Test Kitchen streams (T429898 T430322)]], [[gerrit:1313923{{!}}EventStreamConfig: Remove unused web_ui_scroll* streams (T415370)]], [[gerrit:1325546{{!}}EventStreamConfig: Remove Watchlist click stream (T434790)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug * 08:21 phuedx@deploy1003: Started scap sync-world: Backport for [[gerrit:1324725{{!}}EventStreamConfig: Mark product_metrics.web_base and .web_base_with_ip as Test Kitchen streams (T429898 T430322)]], [[gerrit:1313923{{!}}EventStreamConfig: Remove unused web_ui_scroll* streams (T415370)]], [[gerrit:1325546{{!}}EventStreamConfig: Remove Watchlist click stream (T434790)]] * 08:19 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw1001.wikimedia.org with OS trixie * 08:18 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1035.eqiad.wmnet * 08:17 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1032.eqiad.wmnet * 08:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1032.eqiad.wmnet * 08:16 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1279: Pool back * 08:15 phuedx@deploy1003: Finished scap sync-world: Backport for [[gerrit:1216721{{!}}viwikivoyage: enable relatedarticle and pop-up (T405724)]] (duration: 39m 12s) * 08:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1032.eqiad.wmnet * 08:09 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1032.eqiad.wmnet * 08:08 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1031.eqiad.wmnet * 08:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1031.eqiad.wmnet * 08:03 godog: switch production to use dumps-nfs.w.o - [[phab:T432212|T432212]] * 08:02 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1031.eqiad.wmnet * 08:02 phuedx@deploy1003: nvdtn19, phuedx: Continuing with deployment * 08:00 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1201: Depool db1201.eqiad.wmnet to then clone it to db1287.eqiad.wmnet - marostegui@cumin1003 * 08:00 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1201: Depool db1201.eqiad.wmnet to then clone it to db1287.eqiad.wmnet - marostegui@cumin1003 * 08:00 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1201.eqiad.wmnet onto db1287.eqiad.wmnet * 08:00 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1057.eqiad.wmnet * 08:00 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:00 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1057.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:59 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1057.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:59 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1275: Pool back * 07:55 filippo@cumin1003: START - Cookbook sre.dns.netbox * 07:55 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1031.eqiad.wmnet * 07:52 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1030.eqiad.wmnet * 07:52 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1030.eqiad.wmnet * 07:52 phuedx@deploy1003: nvdtn19, phuedx: Backport for [[gerrit:1216721{{!}}viwikivoyage: enable relatedarticle and pop-up (T405724)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:51 tappof: bump space for prometheus k8s-aux in eqiad * 07:50 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1057.eqiad.wmnet * 07:48 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1281: Pool back * 07:47 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1281 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96108 and previous config saved to /var/cache/conftool/dbconfig/20260817-074749-marostegui.json * 07:46 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1030.eqiad.wmnet * 07:42 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1030.eqiad.wmnet * 07:41 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1174: Depool db1174.eqiad.wmnet to then clone it to db1288.eqiad.wmnet - marostegui@cumin1003 * 07:41 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1174: Depool db1174.eqiad.wmnet to then clone it to db1288.eqiad.wmnet - marostegui@cumin1003 * 07:41 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1174.eqiad.wmnet onto db1288.eqiad.wmnet * 07:40 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1029.eqiad.wmnet * 07:40 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm2001.wikimedia.org * 07:40 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1029.eqiad.wmnet * 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1056.eqiad.wmnet * 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1056.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:38 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1056.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:36 phuedx@deploy1003: Started scap sync-world: Backport for [[gerrit:1216721{{!}}viwikivoyage: enable relatedarticle and pop-up (T405724)]] * 07:36 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm2001.wikimedia.org * 07:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1279 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96104 and previous config saved to /var/cache/conftool/dbconfig/20260817-073542-marostegui.json * 07:34 filippo@cumin1003: START - Cookbook sre.dns.netbox * 07:34 slyngshede@dns1004: END - running authdns-update * 07:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1029.eqiad.wmnet * 07:33 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm-test1001.wikimedia.org * 07:32 slyngshede@dns1004: START - running authdns-update * 07:31 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1029.eqiad.wmnet * 07:31 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1279: Pool back * 07:30 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1279 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96102 and previous config saved to /var/cache/conftool/dbconfig/20260817-073038-marostegui.json * 07:29 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm-test1001.wikimedia.org * 07:29 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm1001.wikimedia.org * 07:28 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1056.eqiad.wmnet * 07:28 moritzm: extend the disk of ldap-rw1001 by 80G [[phab:T331699|T331699]] * 07:28 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1055.eqiad.wmnet * 07:28 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:28 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1055.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:27 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1055.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:26 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1044.eqiad.wmnet * 07:26 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1044.eqiad.wmnet * 07:25 slyngshede@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm1001.wikimedia.org * 07:24 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1175: Depool db1175.eqiad.wmnet to then clone it to db1289.eqiad.wmnet - marostegui@cumin1003 * 07:24 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1175: Depool db1175.eqiad.wmnet to then clone it to db1289.eqiad.wmnet - marostegui@cumin1003 * 07:24 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1175.eqiad.wmnet onto db1289.eqiad.wmnet * 07:22 filippo@cumin1003: START - Cookbook sre.dns.netbox * 07:20 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1044.eqiad.wmnet * 07:16 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1055.eqiad.wmnet * 07:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1054.eqiad.wmnet * 07:15 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:15 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1054.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:15 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1044.eqiad.wmnet * 07:15 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1054.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:13 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1275: Pool back * 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1275 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96098 and previous config saved to /var/cache/conftool/dbconfig/20260817-071225-marostegui.json * 07:10 filippo@cumin1003: START - Cookbook sre.dns.netbox * 07:05 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1043.eqiad.wmnet * 07:05 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1054.eqiad.wmnet * 07:05 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1051.eqiad.wmnet * 07:05 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:05 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1051.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 07:05 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1043.eqiad.wmnet * 07:04 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1051.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 06:59 filippo@cumin1003: START - Cookbook sre.dns.netbox * 06:59 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin1001.eqiad.wmnet * 06:59 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1043.eqiad.wmnet * 06:59 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin2001.codfw.wmnet * 06:55 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin2001.codfw.wmnet * 06:55 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1051.eqiad.wmnet * 06:54 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1049.eqiad.wmnet * 06:54 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:54 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1049.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 06:54 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1049.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 06:54 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin1001.eqiad.wmnet * 06:53 moritzm: installing apr-util security updates * 06:52 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1043.eqiad.wmnet * 06:49 filippo@cumin1003: START - Cookbook sre.dns.netbox * 06:41 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1049.eqiad.wmnet * 06:13 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit2003.wikimedia.org * 06:13 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet * 06:07 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet * 06:06 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit2003.wikimedia.org * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 47s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-16 == * 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 01m 03s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-15 == * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 41s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-14 == * 15:38 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-staging-master-eqiad * 15:38 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster1005.eqiad.wmnet * 15:38 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster1005.eqiad.wmnet * 15:35 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sretest2009.codfw.wmnet * 15:33 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster1005.eqiad.wmnet * 15:33 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster1005.eqiad.wmnet * 15:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster1004.eqiad.wmnet * 15:32 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster1004.eqiad.wmnet * 15:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host sretest2009.codfw.wmnet * 15:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster1004.eqiad.wmnet * 15:27 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster1004.eqiad.wmnet * 15:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster1003.eqiad.wmnet * 15:27 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster1003.eqiad.wmnet * 15:24 dancy@deploy1003: Finished scap sync-world: testing (duration: 03m 23s) * 15:22 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster1003.eqiad.wmnet * 15:22 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster1003.eqiad.wmnet * 15:22 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-staging-master-eqiad * 15:20 dancy@deploy1003: Started scap sync-world: testing * 15:20 dancy@deploy1003: Installation of scap version "4.280.2" completed for 3 hosts * 15:18 dancy@deploy1003: Installing scap version "4.280.2" for 3 host(s) * 15:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sretest2006.codfw.wmnet * 14:54 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host sretest2006.codfw.wmnet * 14:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sretest2003.codfw.wmnet * 14:39 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host sretest2003.codfw.wmnet * 13:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt-staging2001.codfw.wmnet * 13:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt-staging2001.codfw.wmnet * 13:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-staging-master-codfw * 13:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster2005.codfw.wmnet * 13:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster2005.codfw.wmnet * 13:05 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox-dev2003.codfw.wmnet * 13:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster2005.codfw.wmnet * 13:04 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster2005.codfw.wmnet * 13:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster2004.codfw.wmnet * 13:04 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster2004.codfw.wmnet * 13:01 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netbox-dev2003.codfw.wmnet * 12:59 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster2004.codfw.wmnet * 12:59 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster2004.codfw.wmnet * 12:59 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster2003.codfw.wmnet * 12:59 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster2003.codfw.wmnet * 12:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster2003.codfw.wmnet * 12:54 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster2003.codfw.wmnet * 12:54 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-staging-master-codfw * 12:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-staging-worker-eqiad * 12:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage1006.eqiad.wmnet * 12:52 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage1006.eqiad.wmnet * 12:46 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage1006.eqiad.wmnet * 12:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw1001.wikimedia.org with OS trixie * 12:45 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage1006.eqiad.wmnet * 12:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage1005.eqiad.wmnet * 12:45 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage1005.eqiad.wmnet * 12:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage1005.eqiad.wmnet * 12:36 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/ratelimit: apply * 12:35 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/ratelimit: apply * 12:35 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:35 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:33 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage1005.eqiad.wmnet * 12:33 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage1004.eqiad.wmnet * 12:33 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage1004.eqiad.wmnet * 12:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage1004.eqiad.wmnet * 12:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage1004.eqiad.wmnet * 12:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage1003.eqiad.wmnet * 12:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage1003.eqiad.wmnet * 12:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage * 12:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage1003.eqiad.wmnet * 12:17 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage * 12:14 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage1003.eqiad.wmnet * 12:14 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-staging-worker-eqiad * 12:03 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw1001.wikimedia.org with OS trixie * 12:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cuminunpriv1001.eqiad.wmnet * 11:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cuminunpriv1001.eqiad.wmnet * 11:27 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1004.wikimedia.org * 11:24 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1153 from dbctl [[phab:T434638|T434638]]', diff saved to https://phabricator.wikimedia.org/P96097 and previous config saved to /var/cache/conftool/dbconfig/20260814-112449-marostegui.json * 11:21 aokoth@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1004.wikimedia.org * 11:20 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 11:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-staging-worker-codfw * 11:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2004.codfw.wmnet * 11:17 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2004.codfw.wmnet * 11:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2004.codfw.wmnet * 11:10 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2004.codfw.wmnet * 11:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2003.codfw.wmnet * 11:10 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2003.codfw.wmnet * 11:03 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-worker1181.eqiad.wmnet with OS bookworm * 11:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2003.codfw.wmnet * 11:03 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2003.codfw.wmnet * 11:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2002.codfw.wmnet * 11:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2002.codfw.wmnet * 10:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1235.eqiad.wmnet with OS bookworm * 10:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock2003.codfw.wmnet * 10:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock2003.codfw.wmnet with OS trixie * 10:56 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2002.codfw.wmnet * 10:54 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1153.eqiad.wmnet with OS bookworm * 10:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2002.codfw.wmnet * 10:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2001.codfw.wmnet * 10:51 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2001.codfw.wmnet * 10:50 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 10:49 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1187.eqiad.wmnet with OS bookworm * 10:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2001.codfw.wmnet * 10:44 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2001.codfw.wmnet * 10:44 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-staging-worker-codfw * 10:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock2003.codfw.wmnet with reason: host reimage * 10:38 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock2003.codfw.wmnet with reason: host reimage * 10:35 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1153.eqiad.wmnet with reason: host reimage * 10:32 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1235.eqiad.wmnet with reason: host reimage * 10:29 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1187.eqiad.wmnet with reason: host reimage * 10:24 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1153.eqiad.wmnet with reason: host reimage * 10:23 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1235.eqiad.wmnet with reason: host reimage * 10:21 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1187.eqiad.wmnet with reason: host reimage * 10:16 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock2003.codfw.wmnet with OS trixie * 10:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install1005.wikimedia.org * 10:12 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 10:12 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 10:12 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 10:12 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 10:12 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:12 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 10:12 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 10:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install1005.wikimedia.org * 10:07 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1235.eqiad.wmnet with OS bookworm * 10:07 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1187.eqiad.wmnet with OS bookworm * 10:07 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1181.eqiad.wmnet with OS bookworm * 10:07 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1153.eqiad.wmnet with OS bookworm * 10:07 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install2005.wikimedia.org * 10:03 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 10:03 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2003.codfw.wmnet * 10:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1232.eqiad.wmnet with OS bookworm * 10:00 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install2005.wikimedia.org * 10:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install3004.wikimedia.org * 09:58 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 09:58 fceratto@cumin1003: Removing db1151 from zarcillo [[phab:T434538|T434538]] * 09:56 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1231.eqiad.wmnet with OS bookworm * 09:56 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 09:53 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 09:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install3004.wikimedia.org * 09:50 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1152.eqiad.wmnet with OS bookworm * 09:49 Dreamy_Jazz: `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260808000000" --end-timestamp="20260812120000" --sleep="5" --batch-size="50"` for [[phab:T434688|T434688]] * 09:48 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host centrallog2002.codfw.wmnet * 09:48 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install4004.wikimedia.org * 09:42 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1232.eqiad.wmnet with reason: host reimage * 09:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install4004.wikimedia.org * 09:41 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host centrallog2002.codfw.wmnet * 09:39 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install5004.wikimedia.org * 09:36 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1232.eqiad.wmnet with reason: host reimage * 09:36 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'. * 09:34 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'. * 09:33 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1231.eqiad.wmnet with reason: host reimage * 09:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install5004.wikimedia.org * 09:32 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host centrallog1002.eqiad.wmnet * 09:30 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install6003.wikimedia.org * 09:30 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1231.eqiad.wmnet with reason: host reimage * 09:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1152.eqiad.wmnet with reason: host reimage * 09:25 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host centrallog1002.eqiad.wmnet * 09:25 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1152.eqiad.wmnet with reason: host reimage * 09:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install6003.wikimedia.org * 09:22 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1232 * 09:22 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1232 * 09:22 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1232 * 09:22 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1232.eqiad.wmnet 25.53.64.10.in-addr.arpa 5.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:22 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host titan1001.eqiad.wmnet * 09:22 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1232.eqiad.wmnet 25.53.64.10.in-addr.arpa 5.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:22 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:22 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1232 - btullis@cumin1003" * 09:22 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1232 - btullis@cumin1003" * 09:17 btullis@cumin1003: START - Cookbook sre.dns.netbox * 09:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install7002.wikimedia.org * 09:17 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1232 * 09:16 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1231 * 09:16 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1231 * 09:14 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1231 * 09:14 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1231.eqiad.wmnet 24.53.64.10.in-addr.arpa 4.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host titan1001.eqiad.wmnet * 09:14 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1231.eqiad.wmnet 24.53.64.10.in-addr.arpa 4.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:14 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1231 - btullis@cumin1003" * 09:14 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1231 - btullis@cumin1003" * 09:10 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install7002.wikimedia.org * 09:09 btullis@cumin1003: START - Cookbook sre.dns.netbox * 09:08 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1231 * 09:08 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1152 * 09:08 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1152 * 09:06 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1152 * 09:06 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1152.eqiad.wmnet 16.53.64.10.in-addr.arpa 6.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:06 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1152.eqiad.wmnet 16.53.64.10.in-addr.arpa 6.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:06 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:06 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1152 - btullis@cumin1003" * 09:06 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1152 - btullis@cumin1003" * 09:03 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1151.eqiad.wmnet * 09:03 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:03 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1151.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 08:55 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host titan2001.codfw.wmnet * 08:55 btullis@cumin1003: START - Cookbook sre.dns.netbox * 08:54 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1151.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 08:54 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1152 * 08:53 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1232.eqiad.wmnet with OS bookworm * 08:53 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1231.eqiad.wmnet with OS bookworm * 08:53 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1152.eqiad.wmnet with OS bookworm * 08:51 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2209: Pool back * 08:51 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1201.eqiad.wmnet * 08:50 btullis@cumin1003: START - Cookbook sre.hosts.remove-downtime for an-worker1201.eqiad.wmnet * 08:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1229.eqiad.wmnet with OS bookworm * 08:47 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host titan2001.codfw.wmnet * 08:46 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping2004.codfw.wmnet * 08:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ping2004.codfw.wmnet * 08:40 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host an-worker1230.eqiad.wmnet with OS bookworm * 08:40 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 08:34 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1151.eqiad.wmnet * 08:34 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 08:31 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host titan1002.eqiad.wmnet * 08:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1229.eqiad.wmnet with reason: host reimage * 08:25 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host titan1002.eqiad.wmnet * 08:25 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1229.eqiad.wmnet with reason: host reimage * 08:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping1004.eqiad.wmnet * 08:21 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ping1004.eqiad.wmnet * 08:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1230.eqiad.wmnet with reason: host reimage * 08:12 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1230.eqiad.wmnet with reason: host reimage * 08:11 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1229 * 08:11 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1229 * 08:11 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1229 * 08:11 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1229.eqiad.wmnet 22.53.64.10.in-addr.arpa 2.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:11 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1229.eqiad.wmnet 22.53.64.10.in-addr.arpa 2.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:11 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:11 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1229 - btullis@cumin1003" * 08:11 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1229 - btullis@cumin1003" * 08:10 btullis@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1201.eqiad.wmnet with reason: Fixing a disk * 08:07 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host titan2002.codfw.wmnet * 08:05 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2209: Pool back * 08:05 btullis@cumin1003: START - Cookbook sre.dns.netbox * 08:00 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host titan2002.codfw.wmnet * 07:59 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1229 * 07:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1230 * 07:58 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1230 * 07:55 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1230 * 07:55 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1230.eqiad.wmnet 23.53.64.10.in-addr.arpa 3.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:55 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1230.eqiad.wmnet 23.53.64.10.in-addr.arpa 3.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:55 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:55 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1230 - btullis@cumin1003" * 07:55 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1230 - btullis@cumin1003" * 07:51 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host kubestagemaster2005.codfw.wmnet with OS trixie * 07:48 btullis@cumin1003: START - Cookbook sre.dns.netbox * 07:41 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1230 * 07:41 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1229.eqiad.wmnet with OS bookworm * 07:41 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1230.eqiad.wmnet with OS bookworm * 07:39 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1152 from dbctl [[phab:T434480|T434480]]', diff saved to https://phabricator.wikimedia.org/P96090 and previous config saved to /var/cache/conftool/dbconfig/20260814-073941-marostegui.json * 07:29 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on kubestagemaster2005.codfw.wmnet with reason: host reimage * 07:23 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on kubestagemaster2005.codfw.wmnet with reason: host reimage * 07:04 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host kubestagemaster2005.codfw.wmnet with OS trixie * 06:53 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1228.eqiad.wmnet with OS bookworm * 06:44 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1227.eqiad.wmnet with OS bookworm * 06:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1209.eqiad.wmnet with OS bookworm * 06:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1175.eqiad.wmnet with OS bookworm * 06:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1228.eqiad.wmnet with reason: host reimage * 06:27 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1228.eqiad.wmnet with reason: host reimage * 06:25 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1227.eqiad.wmnet with reason: host reimage * 06:21 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1227.eqiad.wmnet with reason: host reimage * 06:18 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1209.eqiad.wmnet with reason: host reimage * 06:14 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1209.eqiad.wmnet with reason: host reimage * 06:14 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1228 * 06:14 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1228 * 06:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1175.eqiad.wmnet with reason: host reimage * 06:12 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1228 * 06:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1228.eqiad.wmnet 20.53.64.10.in-addr.arpa 0.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:12 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1228.eqiad.wmnet 20.53.64.10.in-addr.arpa 0.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1228 - ryankemper@cumin2003" * 06:12 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1228 - ryankemper@cumin2003" * 06:09 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1175.eqiad.wmnet with reason: host reimage * 06:07 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 06:07 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1228 * 06:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1227 * 06:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1227 * 06:06 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1227 * 06:06 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1227.eqiad.wmnet 19.53.64.10.in-addr.arpa 9.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:06 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1227.eqiad.wmnet 19.53.64.10.in-addr.arpa 9.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:06 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:06 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1227 - ryankemper@cumin2003" * 06:06 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1227 - ryankemper@cumin2003" * 06:00 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 06:00 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1227 * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1209 * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1209 * 06:00 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1209 * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1209.eqiad.wmnet 15.53.64.10.in-addr.arpa 5.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:00 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1209.eqiad.wmnet 15.53.64.10.in-addr.arpa 5.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1209 - ryankemper@cumin2003" * 06:00 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1209 - ryankemper@cumin2003" * 05:54 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 05:54 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1209 * 05:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1175 * 05:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1175 * 05:52 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1175 * 05:52 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1175.eqiad.wmnet 17.53.64.10.in-addr.arpa 7.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 05:52 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1175.eqiad.wmnet 17.53.64.10.in-addr.arpa 7.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 05:52 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 05:52 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1175 - ryankemper@cumin2003" * 05:52 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1175 - ryankemper@cumin2003" * 05:49 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1228.eqiad.wmnet with OS bookworm * 05:49 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1227.eqiad.wmnet with OS bookworm * 05:48 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1209.eqiad.wmnet with OS bookworm * 05:47 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 05:47 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1175 * 05:47 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1175.eqiad.wmnet with OS bookworm * 05:09 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host apifeatureusage1001.eqiad.wmnet with OS bookworm * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 03s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:10 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 01:07 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 01:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 01:02 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 00:59 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1226.eqiad.wmnet with OS bookworm * 00:47 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1225.eqiad.wmnet with OS bookworm * 00:41 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1224.eqiad.wmnet with OS bookworm * 00:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1226.eqiad.wmnet with reason: host reimage * 00:31 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1226.eqiad.wmnet with reason: host reimage * 00:28 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1225.eqiad.wmnet with reason: host reimage * 00:25 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1225.eqiad.wmnet with reason: host reimage * 00:19 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1224.eqiad.wmnet with reason: host reimage * 00:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1226 * 00:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1226 * 00:17 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1226 * 00:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1226.eqiad.wmnet 23.36.64.10.in-addr.arpa 3.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:17 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1226.eqiad.wmnet 23.36.64.10.in-addr.arpa 3.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 00:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1226 - ryankemper@cumin2003" * 00:17 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1226 - ryankemper@cumin2003" * 00:16 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1224.eqiad.wmnet with reason: host reimage * 00:12 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 00:12 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1226 * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1225 * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1225 * 00:10 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1225 * 00:10 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1225.eqiad.wmnet 22.36.64.10.in-addr.arpa 2.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:10 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1225.eqiad.wmnet 22.36.64.10.in-addr.arpa 2.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:10 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 00:10 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1225 - ryankemper@cumin2003" * 00:10 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1225 - ryankemper@cumin2003" * 00:03 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 00:02 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1225 * 00:02 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1224 * 00:02 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1224 * 00:00 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1224 * 00:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1224.eqiad.wmnet 21.36.64.10.in-addr.arpa 1.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:00 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1224.eqiad.wmnet 21.36.64.10.in-addr.arpa 1.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 00:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 00:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1224 - ryankemper@cumin2003" * 00:00 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1224 - ryankemper@cumin2003" == 2026-08-13 == * 23:54 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1226.eqiad.wmnet with OS bookworm * 23:53 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1225.eqiad.wmnet with OS bookworm * 23:52 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 23:51 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1224 * 23:51 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1224.eqiad.wmnet with OS bookworm * 23:47 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1223.eqiad.wmnet with OS bookworm * 23:29 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1223.eqiad.wmnet with reason: host reimage * 23:24 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1223.eqiad.wmnet with reason: host reimage * 23:19 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325525{{!}}ve.ui.CodeMirror.less: ensure normal font style]] (duration: 11m 40s) * 23:16 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260807000000" --end-timestamp="20260808000000" --sleep="5" --batch-size="50"` for [[phab:T434688|T434688]] * 23:13 musikanimal@deploy1003: musikanimal: Continuing with deployment * 23:11 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1325525{{!}}ve.ui.CodeMirror.less: ensure normal font style]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1223 * 23:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1223 * 23:08 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1325525{{!}}ve.ui.CodeMirror.less: ensure normal font style]] * 23:07 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1223 * 23:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1223.eqiad.wmnet 20.36.64.10.in-addr.arpa 0.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:07 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1223.eqiad.wmnet 20.36.64.10.in-addr.arpa 0.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 23:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1223 - ryankemper@cumin2003" * 23:03 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1223 - ryankemper@cumin2003" * 23:01 sbassett: Deployed security updates for [[phab:T430596|T430596]], [[phab:T120386|T120386]] * 22:55 bking@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host kubestagemaster2005.codfw.wmnet * 22:55 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host kubestagemaster2005.codfw.wmnet with OS bookworm * 22:54 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 22:53 ryankemper@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 22:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on kubestagemaster2005.codfw.wmnet with reason: host reimage * 22:49 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1222.eqiad.wmnet with OS bookworm * 22:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on kubestagemaster2005.codfw.wmnet with reason: host reimage * 22:29 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1222.eqiad.wmnet with reason: host reimage * 22:26 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1222.eqiad.wmnet with reason: host reimage * 22:23 bking@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host aux-k8s-etcd2003.codfw.wmnet * 22:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd2003.codfw.wmnet with OS bookworm * 22:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host kubestagemaster2005.codfw.wmnet with OS bookworm * 22:23 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM kubestagemaster2005.codfw.wmnet - bking@cumin2003" * 22:23 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM kubestagemaster2005.codfw.wmnet - bking@cumin2003" * 22:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) kubestagemaster2005.codfw.wmnet on all recursors * 22:22 bking@cumin2003: START - Cookbook sre.dns.wipe-cache kubestagemaster2005.codfw.wmnet on all recursors * 22:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:22 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM kubestagemaster2005.codfw.wmnet - bking@cumin2003" * 22:22 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM kubestagemaster2005.codfw.wmnet - bking@cumin2003" * 22:17 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 22:13 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1223 * 22:12 bking@cumin2003: START - Cookbook sre.dns.netbox * 22:12 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host kubestagemaster2005.codfw.wmnet * 22:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1222 * 22:12 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1222 * 22:11 sbassett: Deployed security updates for [[phab:T429244|T429244]], [[phab:T434039|T434039]] * 22:11 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1222 * 22:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1222.eqiad.wmnet 19.36.64.10.in-addr.arpa 9.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:11 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1222.eqiad.wmnet 19.36.64.10.in-addr.arpa 9.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1222 - ryankemper@cumin2003" * 22:07 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1222 - ryankemper@cumin2003" * 22:05 robh@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-wdqs2001.codfw.wmnet with reason: updating firmware * 22:01 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 22:01 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1223.eqiad.wmnet with OS bookworm * 22:01 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1222 * 22:01 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1222.eqiad.wmnet with OS bookworm * 22:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1212.eqiad.wmnet with OS bookworm * 21:54 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324295{{!}}Revert "Lazily reject pre-fix parser-cache entries for noreferrer/noopener links" (T429090)]] (duration: 06m 42s) * 21:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd2003.codfw.wmnet with reason: host reimage * 21:53 bking@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host dse-k8s-etcd2001.codfw.wmnet * 21:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dse-k8s-etcd2001.codfw.wmnet with OS bookworm * 21:50 sbassett@deploy1003: sbassett, kharlan: Continuing with deployment * 21:49 sbassett@deploy1003: sbassett, kharlan: Backport for [[gerrit:1324295{{!}}Revert "Lazily reject pre-fix parser-cache entries for noreferrer/noopener links" (T429090)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:48 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aux-k8s-etcd2003.codfw.wmnet with reason: host reimage * 21:47 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1324295{{!}}Revert "Lazily reject pre-fix parser-cache entries for noreferrer/noopener links" (T429090)]] * 21:42 maryum: Deployed security patch for [[phab:T434549|T434549]] * 21:39 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1212.eqiad.wmnet with reason: host reimage * 21:34 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1212.eqiad.wmnet with reason: host reimage * 21:33 bking@cumin2003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd2003.codfw.wmnet with OS bookworm * 21:32 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM aux-k8s-etcd2003.codfw.wmnet - bking@cumin2003" * 21:32 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM aux-k8s-etcd2003.codfw.wmnet - bking@cumin2003" * 21:32 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) aux-k8s-etcd2003.codfw.wmnet on all recursors * 21:32 bking@cumin2003: START - Cookbook sre.dns.wipe-cache aux-k8s-etcd2003.codfw.wmnet on all recursors * 21:32 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:32 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM aux-k8s-etcd2003.codfw.wmnet - bking@cumin2003" * 21:31 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM aux-k8s-etcd2003.codfw.wmnet - bking@cumin2003" * 21:28 maryum: Deployed security patch for [[phab:T434619|T434619]] * 21:26 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:26 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host aux-k8s-etcd2003.codfw.wmnet * 21:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-etcd2001.codfw.wmnet with reason: host reimage * 21:20 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1212 * 21:20 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1212 * 21:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1212.eqiad.wmnet with OS bookworm * 21:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on dse-k8s-etcd2001.codfw.wmnet with reason: host reimage * 21:04 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325553{{!}}InstrumentConstructiveEdits: exclude mw-reverted as well (T431493)]] (duration: 06m 25s) * 20:59 kemayo@deploy1003: kemayo: Continuing with deployment * 20:59 kemayo@deploy1003: kemayo: Backport for [[gerrit:1325553{{!}}InstrumentConstructiveEdits: exclude mw-reverted as well (T431493)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host dse-k8s-etcd2001.codfw.wmnet with OS bookworm * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 20:57 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 20:57 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1325553{{!}}InstrumentConstructiveEdits: exclude mw-reverted as well (T431493)]] * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-etcd2001.codfw.wmnet on all recursors * 20:57 bking@cumin2003: START - Cookbook sre.dns.wipe-cache dse-k8s-etcd2001.codfw.wmnet on all recursors * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 20:57 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 20:54 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320163{{!}}Add configurable RestTermsOfServiceUrl (T428147)]] (duration: 21m 39s) * 20:53 bking@cumin2003: START - Cookbook sre.dns.netbox * 20:53 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host dse-k8s-etcd2001.codfw.wmnet * 20:50 samtar@deploy1003: samtar, milazg: Continuing with deployment * 20:35 samtar@deploy1003: samtar, milazg: Backport for [[gerrit:1320163{{!}}Add configurable RestTermsOfServiceUrl (T428147)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:33 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1320163{{!}}Add configurable RestTermsOfServiceUrl (T428147)]] * 20:30 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325549{{!}}Deploy PRV to several LC wikis (T423785)]] (duration: 06m 57s) * 20:26 arlolra@deploy1003: arlolra: Continuing with deployment * 20:25 arlolra@deploy1003: arlolra: Backport for [[gerrit:1325549{{!}}Deploy PRV to several LC wikis (T423785)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:24 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 20:23 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1325549{{!}}Deploy PRV to several LC wikis (T423785)]] * 20:21 ariel@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319906{{!}}Remove boilerplate language from wmf-rest and wmf-math API modules (T433736)]] (duration: 13m 54s) * 20:14 ariel@deploy1003: ariel: Continuing with deployment * 20:11 ariel@deploy1003: ariel: Backport for [[gerrit:1319906{{!}}Remove boilerplate language from wmf-rest and wmf-math API modules (T433736)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 ariel@deploy1003: Started scap sync-world: Backport for [[gerrit:1319906{{!}}Remove boilerplate language from wmf-rest and wmf-math API modules (T433736)]] * 19:58 robh@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-wdqs2001.codfw.wmnet with reason: updating firmware * 19:54 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325551{{!}}Render the focused module view as a full-screen page (T433896)]] (duration: 30m 37s) * 19:52 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching aqs[2001,1016]*: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 19:44 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching aqs[2001,1016]*: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 19:42 musikanimal@deploy1003: musikanimal: Continuing with deployment * 19:41 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1325551{{!}}Render the focused module view as a full-screen page (T433896)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:34 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: quash java safepoint logspam - bking@cumin2003 - [[phab:T434685|T434685]] * 19:34 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 19:34 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 19:24 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1325551{{!}}Render the focused module view as a full-screen page (T433896)]] * 19:20 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 19:20 jhancock@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin1003" * 19:18 jhancock@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin1003" * 19:14 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325548{{!}}Enable image lazy loading on desktop in group1 (T148047)]] (duration: 07m 43s) * 19:10 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 19:10 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1325548{{!}}Enable image lazy loading on desktop in group1 (T148047)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:07 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1325548{{!}}Enable image lazy loading on desktop in group1 (T148047)]] * 19:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1166.eqiad.wmnet onto db1280.eqiad.wmnet * 19:03 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 19:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1166: Pool db1166.eqiad.wmnet in after cloning * 18:59 jhancock@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 18:58 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1234.eqiad.wmnet with OS bookworm * 18:52 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 18:51 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:49 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:46 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 18:46 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 18:44 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=dns3004.* * 18:39 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1221.eqiad.wmnet with OS bookworm * 18:36 inflatador: [bking@ganeti2048] ~$ sudo gnt-instance replace-disks -n ganeti2030.codfw.wmnet aux-k8s-worker2002.codfw.wmnet [[phab:T434681|T434681]] * 18:35 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 18:29 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1220.eqiad.wmnet with OS bookworm * 18:29 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1234.eqiad.wmnet with reason: host reimage * 18:27 bking@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host dse-k8s-etcd2001.codfw.wmnet * 18:27 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-etcd2001.codfw.wmnet on all recursors * 18:27 bking@cumin2003: START - Cookbook sre.dns.wipe-cache dse-k8s-etcd2001.codfw.wmnet on all recursors * 18:27 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:27 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 18:27 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 18:25 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1234.eqiad.wmnet with reason: host reimage * 18:25 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 18:19 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1221.eqiad.wmnet with reason: host reimage * 18:17 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1166: Pool db1166.eqiad.wmnet in after cloning * 18:15 brennen@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 18:15 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1221.eqiad.wmnet with reason: host reimage * 18:14 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:12 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:12 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 18:12 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-etcd2001.codfw.wmnet on all recursors * 18:12 bking@cumin2003: START - Cookbook sre.dns.wipe-cache dse-k8s-etcd2001.codfw.wmnet on all recursors * 18:12 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:12 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 18:12 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2001.codfw.wmnet - bking@cumin2003" * 18:10 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: quash java safepoint logspam - bking@cumin2003 - [[phab:T434685|T434685]] * 18:10 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 18:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1234 * 18:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1234 * 18:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1220.eqiad.wmnet with reason: host reimage * 18:07 brennen: 1.47.0-wmf.15 train status ([[phab:T430834|T430834]]) - no current blockers, rolling to all wikis * 18:07 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1234 * 18:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1234.eqiad.wmnet 10.36.64.10.in-addr.arpa 0.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:07 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:07 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1234.eqiad.wmnet 10.36.64.10.in-addr.arpa 0.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1234 - ryankemper@cumin2003" * 18:07 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1234 - ryankemper@cumin2003" * 18:05 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1220.eqiad.wmnet with reason: host reimage * 18:04 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host dse-k8s-etcd2001.codfw.wmnet * 18:04 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 18:03 dancy@deploy1003: Installation of scap version "4.280.1" completed for 3 hosts * 18:02 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 18:02 inflatador: bking@dse-k8s-etcd2002 etcdctl member remove $<nowiki>{</nowiki>UUID of dse-k8s-etcd2001<nowiki>}</nowiki> [[phab:T434681|T434681]] [[phab:T434793|T434793]] * 18:01 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 18:01 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1234 * 18:01 dancy@deploy1003: Installing scap version "4.280.1" for 3 host(s) * 18:01 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1221 * 18:01 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1221 * 18:00 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1221 * 18:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1221.eqiad.wmnet 18.36.64.10.in-addr.arpa 8.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:00 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1221.eqiad.wmnet 18.36.64.10.in-addr.arpa 8.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1221 - ryankemper@cumin2003" * 17:59 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:58 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 17:58 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:57 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 17:57 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:56 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1221 - ryankemper@cumin2003" * 17:53 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 17:52 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 17:52 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 17:51 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1221 * 17:51 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1220 * 17:51 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1220 * 17:51 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:51 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:51 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1220 * 17:51 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1220.eqiad.wmnet 11.36.64.10.in-addr.arpa 1.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:51 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache an-worker1220.eqiad.wmnet 11.36.64.10.in-addr.arpa 1.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:51 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1220 - ryankemper@cumin2003" * 17:50 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1220 - ryankemper@cumin2003" * 17:47 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:47 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 17:46 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1234.eqiad.wmnet with OS bookworm * 17:46 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 17:45 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1221.eqiad.wmnet with OS bookworm * 17:45 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1220 * 17:45 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1220.eqiad.wmnet with OS bookworm * 17:43 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 17:41 bking@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host dse-k8s-etcd2001.codfw.wmnet * 17:41 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host dse-k8s-etcd2001.codfw.wmnet * 17:40 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:40 inflatador: bking@ganeti2048] `sudo gnt-instance remove --force --ignore-failures --shutdown-timeout=0` on non-DRBD VMs [[phab:T434681|T434681]] * 17:40 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:39 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:38 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:36 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:32 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:32 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:28 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 17:26 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:26 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:24 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:23 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:21 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1218.eqiad.wmnet with OS bookworm * 17:20 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:20 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:19 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1179.eqiad.wmnet with OS bookworm * 17:18 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1150.eqiad.wmnet with OS bookworm * 17:18 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns3004.wikimedia.org with OS trixie * 17:15 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:14 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 17:13 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 17:12 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:12 swfrench@deploy1003: Finished scap sync-world: Helmfile-only deployment for mediawiki chart bump - [[phab:T427666|T427666]] (duration: 03m 03s) * 17:09 swfrench@deploy1003: Started scap sync-world: Helmfile-only deployment for mediawiki chart bump - [[phab:T427666|T427666]] * 17:02 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:01 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:01 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1218.eqiad.wmnet with reason: host reimage * 17:01 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:01 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 17:00 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324810{{!}}deployment-info.php: Report dbname and branch for the requested wiki (T434726)]] (duration: 06m 52s) * 16:58 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 16:57 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1179.eqiad.wmnet with reason: host reimage * 16:56 dancy@deploy1003: dancy: Continuing with deployment * 16:56 dancy@deploy1003: dancy: Backport for [[gerrit:1324810{{!}}deployment-info.php: Report dbname and branch for the requested wiki (T434726)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1150.eqiad.wmnet with reason: host reimage * 16:53 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324810{{!}}deployment-info.php: Report dbname and branch for the requested wiki (T434726)]] * 16:51 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1166: Depool db1166.eqiad.wmnet to then clone it to db1280.eqiad.wmnet - cwilliams@cumin1003 * 16:50 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1166: Depool db1166.eqiad.wmnet to then clone it to db1280.eqiad.wmnet - cwilliams@cumin1003 * 16:50 cwilliams@cumin1003: START - Cookbook sre.mysql.clone of db1166.eqiad.wmnet onto db1280.eqiad.wmnet * 16:49 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1218.eqiad.wmnet with reason: host reimage * 16:48 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1179.eqiad.wmnet with reason: host reimage * 16:47 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1150.eqiad.wmnet with reason: host reimage * 16:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.upgrade (exit_code=0) for 1 hosts * 16:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2220: Upgrade of db2220.codfw.wmnet completed * 16:38 swfrench@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 16:38 swfrench-wmf: kubectl delete node kubestagemaster2005.codfw.wmnet - [[phab:T434681|T434681]] * 16:34 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1218 * 16:34 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1218 * 16:34 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1218.eqiad.wmnet with OS bookworm * 16:34 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1179 * 16:34 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1179 * 16:33 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1179.eqiad.wmnet with OS bookworm * 16:32 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1150 * 16:32 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1150 * 16:31 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1150.eqiad.wmnet with OS bookworm * 16:24 dancy@deploy1003: Installation of scap version "4.280.0" completed for 3 hosts * 16:22 dancy@deploy1003: Installing scap version "4.280.0" for 3 host(s) * 16:20 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: quash java safepoint logspam - bking@cumin2003 - [[phab:T434685|T434685]] * 16:14 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns3004.wikimedia.org with reason: host reimage * 16:08 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 16:07 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 16:07 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns3004.wikimedia.org with reason: host reimage * 16:05 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 16:04 swfrench@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 16:00 swfrench@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 15:59 swfrench@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 15:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Upgrade of db2220.codfw.wmnet completed * 15:48 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2220: Upgrading db2220.codfw.wmnet * 15:48 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2220: Upgrading db2220.codfw.wmnet * 15:48 cwilliams@cumin1003: START - Cookbook sre.mysql.upgrade for 1 hosts * 15:46 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns3004.wikimedia.org with OS trixie * 15:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2220 [[phab:T434802|T434802]]', diff saved to https://phabricator.wikimedia.org/P96079 and previous config saved to /var/cache/conftool/dbconfig/20260813-154624-cwilliams.json * 15:45 cdobbins@cumin1003: conftool action : set/pooled=no; selector: name=dns3004.* * 15:44 cjd91: depooling dns3004 to reimage to trixie * 15:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2159 to s7 primary [[phab:T434802|T434802]]', diff saved to https://phabricator.wikimedia.org/P96078 and previous config saved to /var/cache/conftool/dbconfig/20260813-154405-cwilliams.json * 15:43 cezmunsta: Starting s7 codfw failover from db2220 to db2159 - [[phab:T434802|T434802]] * 15:41 cgoubert@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: Dragonfly supernodes reboot (duration: 09m 42s) * 15:41 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dragonfly-supernode2001.codfw.wmnet * 15:39 inflatador: bking@ganeti2048] ~$ sudo gnt-node failover -f --ignore-consistency ganeti2046.codfw.wmnet [[phab:T434681|T434681]] * 15:39 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 15:39 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 15:39 swfrench@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 15:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2159 with weight 0 [[phab:T434802|T434802]]', diff saved to https://phabricator.wikimedia.org/P96077 and previous config saved to /var/cache/conftool/dbconfig/20260813-153806-cwilliams.json * 15:37 swfrench@dns1004: END - running authdns-update * 15:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 30 hosts with reason: Primary switchover s7 [[phab:T434802|T434802]] * 15:37 cgoubert@cumin2003: START - Cookbook sre.hosts.reboot-single for host dragonfly-supernode2001.codfw.wmnet * 15:36 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dragonfly-supernode1001.eqiad.wmnet * 15:35 swfrench@dns1004: START - running authdns-update * 15:32 cgoubert@cumin2003: START - Cookbook sre.hosts.reboot-single for host dragonfly-supernode1001.eqiad.wmnet * 15:32 cgoubert@deploy1003: Locking from deployment [ALL REPOSITORIES]: Dragonfly supernodes reboot * 15:30 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-master-codfw * 15:30 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2005.codfw.wmnet * 15:30 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2005.codfw.wmnet * 15:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1169.eqiad.wmnet onto db1277.eqiad.wmnet * 15:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1169: Pool db1169.eqiad.wmnet in after cloning * 15:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2005.codfw.wmnet * 15:23 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2005.codfw.wmnet * 15:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2004.codfw.wmnet * 15:23 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2004.codfw.wmnet * 15:18 inflatador: bking@ganeti2048 sudo gnt-node failover -f ganeti2046.codfw.wmnet [[phab:T434681|T434681]] * 15:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2004.codfw.wmnet * 15:17 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2004.codfw.wmnet * 15:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2003.codfw.wmnet * 15:17 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2003.codfw.wmnet * 15:15 cdobbins@dns1004: END - running authdns-update * 15:13 cdobbins@dns1004: START - running authdns-update * 15:10 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1217.eqiad.wmnet with OS bookworm * 15:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2003.codfw.wmnet * 15:10 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2003.codfw.wmnet * 15:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2002.codfw.wmnet * 15:10 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2002.codfw.wmnet * 15:10 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: quash java safepoint logspam - bking@cumin2003 - [[phab:T434685|T434685]] * 15:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1216.eqiad.wmnet with OS bookworm * 15:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2002.codfw.wmnet * 15:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2002.codfw.wmnet * 15:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2001.codfw.wmnet * 15:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2001.codfw.wmnet * 15:01 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1042.eqiad.wmnet * 15:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1042.eqiad.wmnet * 15:00 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325455{{!}}Api: Use correct query when continue prop=categories (T433922)]] (duration: 09m 47s) * 14:59 bking@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host dse-k8s-etcd2004.codfw.wmnet * 14:58 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-etcd2004.codfw.wmnet on all recursors * 14:58 bking@cumin2003: START - Cookbook sre.dns.wipe-cache dse-k8s-etcd2004.codfw.wmnet on all recursors * 14:58 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:58 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM dse-k8s-etcd2004.codfw.wmnet - bking@cumin2003" * 14:58 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM dse-k8s-etcd2004.codfw.wmnet - bking@cumin2003" * 14:55 zabe@deploy1003: zabe: Continuing with deployment * 14:53 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-etcd2004.codfw.wmnet on all recursors * 14:53 bking@cumin2003: START - Cookbook sre.dns.wipe-cache dse-k8s-etcd2004.codfw.wmnet on all recursors * 14:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:53 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2004.codfw.wmnet - bking@cumin2003" * 14:53 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM dse-k8s-etcd2004.codfw.wmnet - bking@cumin2003" * 14:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2001.codfw.wmnet * 14:52 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2001.codfw.wmnet * 14:52 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-master-codfw * 14:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2001.codfw.wmnet * 14:52 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2001.codfw.wmnet * 14:52 zabe@deploy1003: zabe: Backport for [[gerrit:1325455{{!}}Api: Use correct query when continue prop=categories (T433922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:50 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1325455{{!}}Api: Use correct query when continue prop=categories (T433922)]] * 14:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1217.eqiad.wmnet with reason: host reimage * 14:48 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:48 bking@cumin2003: START - Cookbook sre.ganeti.makevm for new host dse-k8s-etcd2004.codfw.wmnet * 14:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-master-eqiad * 14:44 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1006.eqiad.wmnet * 14:44 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1006.eqiad.wmnet * 14:44 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1216.eqiad.wmnet with reason: host reimage * 14:40 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1217.eqiad.wmnet with reason: host reimage * 14:39 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1169: Pool db1169.eqiad.wmnet in after cloning * 14:39 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1216.eqiad.wmnet with reason: host reimage * 14:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl1005.eqiad.wmnet * 14:32 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl1005.eqiad.wmnet * 14:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1004.eqiad.wmnet * 14:32 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1004.eqiad.wmnet * 14:31 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus2008.codfw.wmnet * 14:31 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor1003.eqiad.wmnet * 14:29 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1042.eqiad.wmnet * 14:27 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor1003.eqiad.wmnet * 14:26 moritzm: installing Django security updates * 14:26 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling reboot on A:wikidough * 14:25 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1217.eqiad.wmnet with OS bookworm * 14:25 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1216.eqiad.wmnet with OS bookworm * 14:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl1004.eqiad.wmnet * 14:24 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl1004.eqiad.wmnet * 14:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1003.eqiad.wmnet * 14:24 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1003.eqiad.wmnet * 14:23 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus2008.codfw.wmnet * 14:23 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus1008.eqiad.wmnet * 14:23 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor-dev2001.codfw.wmnet * 14:22 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus2006.codfw.wmnet * 14:19 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor-dev2001.codfw.wmnet * 14:18 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1042.eqiad.wmnet * 14:17 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.* * 14:16 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl1003.eqiad.wmnet * 14:16 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl1003.eqiad.wmnet * 14:16 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1002.eqiad.wmnet * 14:16 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1002.eqiad.wmnet * 14:15 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus1008.eqiad.wmnet * 14:14 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus1006.eqiad.wmnet * 14:12 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus2006.codfw.wmnet * 14:12 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor2003.codfw.wmnet * 14:11 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus2007.codfw.wmnet * 14:11 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1041.eqiad.wmnet * 14:11 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1041.eqiad.wmnet * 14:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl1002.eqiad.wmnet * 14:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl1002.eqiad.wmnet * 14:09 cgoubert@cumin2003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-master-eqiad * 14:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on apifeatureusage1001.eqiad.wmnet with reason: host reimage * 14:08 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1215.eqiad.wmnet with OS bookworm * 14:08 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor2003.codfw.wmnet * 14:08 jayme: updated calico to v3.30.7 on wikikube codfw [[phab:T427400|T427400]] * 14:07 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1149.eqiad.wmnet with OS bookworm * 14:06 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'. * 14:06 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1041.eqiad.wmnet * 14:04 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus1006.eqiad.wmnet * 14:03 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus2007.codfw.wmnet * 14:03 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus2005.codfw.wmnet * 14:03 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus1007.eqiad.wmnet * 14:03 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1002.eqiad.wmnet * 14:02 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on apifeatureusage1001.eqiad.wmnet with reason: host reimage * 14:02 cgoubert@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=helm-charts.*,name=eqiad * 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host chartmuseum1001.eqiad.wmnet * 14:00 moritzm: installing libxml2 security updates * 13:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1214.eqiad.wmnet with OS bookworm * 13:59 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns5003.* * 13:59 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1041.eqiad.wmnet * 13:59 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow1002.eqiad.wmnet * 13:58 cmooney@dns3003: END - running authdns-update * 13:57 cgoubert@cumin2003: START - Cookbook sre.hosts.reboot-single for host chartmuseum1001.eqiad.wmnet * 13:57 cgoubert@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=helm-charts.*,name=eqiad * 13:57 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: cloudelastic cluster restart - bking@cumin2003 * 13:57 cgoubert@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=helm-charts.*,name=codfw * 13:56 cmooney@dns3003: START - running authdns-update * 13:56 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'. * 13:56 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns5003.*,service=authdns-update * 13:55 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host chartmuseum2001.codfw.wmnet * 13:55 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus1007.eqiad.wmnet * 13:55 cmooney@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dns5003.wikimedia.org * 13:55 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus1005.eqiad.wmnet * 13:51 cgoubert@cumin2003: START - Cookbook sre.hosts.reboot-single for host chartmuseum2001.codfw.wmnet * 13:51 cgoubert@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=helm-charts.*,name=codfw * 13:51 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus2005.codfw.wmnet * 13:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host apifeatureusage1001.eqiad.wmnet with OS bookworm * 13:50 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus7002.magru.wmnet * 13:50 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts lvs1015.eqiad.wmnet * 13:50 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:50 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1015.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:49 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1015.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:49 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325480{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] (duration: 06m 39s) * 13:47 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1215.eqiad.wmnet with reason: host reimage * 13:46 cmooney@cumin1003: START - Cookbook sre.hosts.reboot-single for host dns5003.wikimedia.org * 13:46 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1003.eqiad.wmnet * 13:45 cmooney@cumin1003: conftool action : set/pooled=no; selector: name=dns5003.* * 13:45 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1040.eqiad.wmnet * 13:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1040.eqiad.wmnet * 13:45 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 13:45 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus1005.eqiad.wmnet * 13:45 stran@deploy1003: stran: Continuing with deployment * 13:44 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus7002.magru.wmnet * 13:44 stran@deploy1003: stran: Backport for [[gerrit:1325480{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:44 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus6002.drmrs.wmnet * 13:43 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260805000000" --end-timestamp="20260806000000" --sleep="5" --batch-size="10"` for [[phab:T434688|T434688]] * 13:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1149.eqiad.wmnet with reason: host reimage * 13:42 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1325480{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] * 13:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow1003.eqiad.wmnet * 13:40 sukhe@cumin1003: START - Cookbook sre.hosts.decommission for hosts lvs1015.eqiad.wmnet * 13:40 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts lvs1014.eqiad.wmnet * 13:40 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:40 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1014.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:40 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1040.eqiad.wmnet * 13:40 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1014.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:39 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1214.eqiad.wmnet with reason: host reimage * 13:38 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2004.codfw.wmnet * 13:38 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus6002.drmrs.wmnet * 13:38 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus5003.eqsin.wmnet * 13:35 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 13:35 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1149.eqiad.wmnet with reason: host reimage * 13:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1215.eqiad.wmnet with reason: host reimage * 13:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow2004.codfw.wmnet * 13:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1214.eqiad.wmnet with reason: host reimage * 13:33 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1040.eqiad.wmnet * 13:31 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus5003.eqsin.wmnet * 13:31 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: cloudelastic cluster restart - bking@cumin2003 * 13:31 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus4003.ulsfo.wmnet * 13:30 sukhe@cumin1003: START - Cookbook sre.hosts.decommission for hosts lvs1014.eqiad.wmnet * 13:30 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts lvs1013.eqiad.wmnet * 13:30 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:30 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1013.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:30 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: lvs1013.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - sukhe@cumin1003" * 13:27 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1039.eqiad.wmnet * 13:27 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1039.eqiad.wmnet * 13:26 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1167.eqiad.wmnet onto db1281.eqiad.wmnet * 13:26 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1167: Pool db1167.eqiad.wmnet in after cloning * 13:25 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus4003.ulsfo.wmnet * 13:25 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260802000000" --end-timestamp="20260803000000" --sleep="5" --batch-size="10"` for [[phab:T434688|T434688]] * 13:24 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 13:24 tappof@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host prometheus3004.esams.wmnet * 13:24 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1274: New host * 13:24 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325476{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] (duration: 07m 19s) * 13:24 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260804000000" --end-timestamp="20260805000000" --sleep="5" --batch-size="10"` for [[phab:T434688|T434688]] * 13:24 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260803000000" --end-timestamp="20260804000000" --sleep="5" --batch-size="10"` for [[phab:T434688|T434688]] * 13:23 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2003.codfw.wmnet * 13:22 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1039.eqiad.wmnet * 13:20 sukhe@cumin1003: START - Cookbook sre.hosts.decommission for hosts lvs1013.eqiad.wmnet * 13:20 stran@deploy1003: stran: Continuing with deployment * 13:19 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow2003.codfw.wmnet * 13:19 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling reboot on A:wikidough * 13:19 stran@deploy1003: stran: Backport for [[gerrit:1325476{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:18 tappof@cumin1003: START - Cookbook sre.hosts.reboot-single for host prometheus3004.esams.wmnet * 13:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1215.eqiad.wmnet with OS bookworm * 13:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1214.eqiad.wmnet with OS bookworm * 13:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1149.eqiad.wmnet with OS bookworm * 13:17 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1325476{{!}}Guard against malformed headers in SuggestedInvestigationsMatchSignalsAgainstUserJob (T434199)]] * 13:11 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324731{{!}}prv: Enable parsoid rendering for 5 wikisource wikis]] (duration: 07m 26s) * 13:11 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1039.eqiad.wmnet * 13:11 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1036.eqiad.wmnet * 13:11 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1036.eqiad.wmnet * 13:07 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow3004.esams.wmnet * 13:07 jgiannelos@deploy1003: jgiannelos: Continuing with deployment * 13:06 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1233.eqiad.wmnet with OS bookworm * 13:06 jgiannelos@deploy1003: jgiannelos: Backport for [[gerrit:1324731{{!}}prv: Enable parsoid rendering for 5 wikisource wikis]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:04 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1324731{{!}}prv: Enable parsoid rendering for 5 wikisource wikis]] * 13:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow3004.esams.wmnet * 13:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1036.eqiad.wmnet * 13:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1210.eqiad.wmnet with OS bookworm * 13:01 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1036.eqiad.wmnet * 13:00 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1052.eqiad.wmnet * 13:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1052.eqiad.wmnet * 12:58 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow4003.ulsfo.wmnet * 12:55 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1211.eqiad.wmnet with OS bookworm * 12:55 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1052.eqiad.wmnet * 12:52 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow4003.ulsfo.wmnet * 12:51 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1052.eqiad.wmnet * 12:49 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1169: Depool db1169.eqiad.wmnet to then clone it to db1277.eqiad.wmnet - cwilliams@cumin1003 * 12:46 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1169: Depool db1169.eqiad.wmnet to then clone it to db1277.eqiad.wmnet - cwilliams@cumin1003 * 12:46 cwilliams@cumin1003: START - Cookbook sre.mysql.clone of db1169.eqiad.wmnet onto db1277.eqiad.wmnet * 12:45 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:45 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns record for deleted IP reservations eqsin lvs vlan ints - cmooney@cumin1003" * 12:44 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns record for deleted IP reservations eqsin lvs vlan ints - cmooney@cumin1003" * 12:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1051.eqiad.wmnet * 12:43 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1051.eqiad.wmnet * 12:42 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1210.eqiad.wmnet with reason: host reimage * 12:41 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1167: Pool db1167.eqiad.wmnet in after cloning * 12:40 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 12:39 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1274: New host * 12:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Added db1274', diff saved to https://phabricator.wikimedia.org/P96062 and previous config saved to /var/cache/conftool/dbconfig/20260813-123907-cwilliams.json * 12:38 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast7002.wikimedia.org * 12:38 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1233.eqiad.wmnet with reason: host reimage * 12:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1051.eqiad.wmnet * 12:34 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1211.eqiad.wmnet with reason: host reimage * 12:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1233.eqiad.wmnet with reason: host reimage * 12:32 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1051.eqiad.wmnet * 12:32 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast7002.wikimedia.org * 12:32 marostegui: Drop SecurePoll tables from closed wikis [[phab:T423128|T423128]] * 12:32 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow5003.eqsin.wmnet * 12:31 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1050.eqiad.wmnet * 12:31 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1050.eqiad.wmnet * 12:29 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1210.eqiad.wmnet with reason: host reimage * 12:29 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1211.eqiad.wmnet with reason: host reimage * 12:29 moritzm: installing Wireshark security updates * 12:26 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow5003.eqsin.wmnet * 12:26 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1050.eqiad.wmnet * 12:24 cmooney@dns3003: END - running authdns-update * 12:21 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow6001.drmrs.wmnet * 12:21 cmooney@dns3003: START - running authdns-update * 12:20 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1050.eqiad.wmnet * 12:20 cgoubert@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host rdb-lock2003.codfw.wmnet * 12:20 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:20 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update netbox dns entries for expanded public1-603-eqsin subnet - cmooney@cumin1003" * 12:20 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update netbox dns entries for expanded public1-603-eqsin subnet - cmooney@cumin1003" * 12:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 12:20 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 12:19 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1049.eqiad.wmnet * 12:19 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1049.eqiad.wmnet * 12:19 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 12:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host moss-be1003.eqiad.wmnet * 12:17 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow6001.drmrs.wmnet * 12:17 moritzm: installin curl security updates * 12:15 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 12:15 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1233.eqiad.wmnet with OS bookworm * 12:15 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1210.eqiad.wmnet with OS bookworm * 12:15 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1211.eqiad.wmnet with OS bookworm * 12:13 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1049.eqiad.wmnet * 12:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow7002.magru.wmnet * 12:12 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-worker1178.eqiad.wmnet with OS bookworm * 12:10 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host moss-be1003.eqiad.wmnet * 12:10 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 12:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be1006.eqiad.wmnet * 12:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 12:10 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 12:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:10 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 12:10 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 12:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow7002.magru.wmnet * 12:08 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1049.eqiad.wmnet * 12:05 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 12:05 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2003.codfw.wmnet * 12:04 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'. * 12:04 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:03 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be1006.eqiad.wmnet * 12:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be1005.eqiad.wmnet * 12:01 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1038.eqiad.wmnet * 12:01 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1038.eqiad.wmnet * 12:01 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1213.eqiad.wmnet with OS bookworm * 11:56 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be1005.eqiad.wmnet * 11:55 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be1004.eqiad.wmnet * 11:54 cgoubert@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host rdb-lock2003.codfw.wmnet * 11:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 11:53 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 11:53 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:53 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 11:53 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 11:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1038.eqiad.wmnet * 11:49 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be1004.eqiad.wmnet * 11:44 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:44 moritzm: installing Linux 5.10.262 on Bullseye hosts * 11:41 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1038.eqiad.wmnet * 11:40 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 11:40 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 11:40 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 11:40 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:40 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 11:40 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 11:38 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1213.eqiad.wmnet with reason: host reimage * 11:36 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1167: Depool db1167.eqiad.wmnet to then clone it to db1281.eqiad.wmnet - marostegui@cumin1003 * 11:35 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 11:35 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2003.codfw.wmnet * 11:35 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1167: Depool db1167.eqiad.wmnet to then clone it to db1281.eqiad.wmnet - marostegui@cumin1003 * 11:35 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1167.eqiad.wmnet onto db1281.eqiad.wmnet * 11:34 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 22 hosts with reason: Cloning * 11:34 moritzm: remove ganeti3005 from esams03 cluster, hardware issues [[phab:T434646|T434646]] * 11:32 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1213.eqiad.wmnet with reason: host reimage * 11:28 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1034.eqiad.wmnet * 11:28 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1034.eqiad.wmnet * 11:22 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1034.eqiad.wmnet * 11:19 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1034.eqiad.wmnet * 11:17 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1213.eqiad.wmnet with OS bookworm * 11:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1165.eqiad.wmnet onto db1279.eqiad.wmnet * 11:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1165: Pool db1165.eqiad.wmnet in after cloning * 11:07 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply * 10:57 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply * 10:54 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply * 10:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1274.eqiad.wmnet with reason: Enabling notifications and pooling * 10:45 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1033.eqiad.wmnet * 10:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1033.eqiad.wmnet * 10:44 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply * 10:43 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'. * 10:42 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'. * 10:42 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'. * 10:40 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db[1216,1225,1239-1240].eqiad.wmnet with reason: reboot * 10:39 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti1033.eqiad.wmnet * 10:38 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1151 from dbctl [[phab:T434538|T434538]]', diff saved to https://phabricator.wikimedia.org/P96055 and previous config saved to /var/cache/conftool/dbconfig/20260813-103828-marostegui.json * 10:35 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1033.eqiad.wmnet * 10:27 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1165: Pool db1165.eqiad.wmnet in after cloning * 10:24 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1151.eqiad.wmnet with OS bookworm * 10:15 moritzm: installing bind9 security updates (client-side tools/libs only) * 10:07 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-debug: apply * 10:06 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-debug: apply * 10:02 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 7 hosts with reason: reboot * 10:01 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-debug: apply * 10:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.upgrade (exit_code=0) for 1 hosts * 10:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2214: Upgrade of db2214.codfw.wmnet completed * 10:01 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-debug: apply * 10:00 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-debug: apply * 10:00 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-debug: apply * 09:59 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.decommission (exit_code=99) * 09:59 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 09:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1151.eqiad.wmnet with reason: host reimage * 09:59 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 8 hosts * 09:59 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 8 hosts * 09:55 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1151.eqiad.wmnet with reason: host reimage * 09:46 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260801000000" --end-timestamp="20260802000000" --sleep="3" --batch-size="5"` for [[phab:T434688|T434688]] * 09:45 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 8 hosts with reason: reboot * 09:45 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 22 hosts with reason: Cloning * 09:43 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1165: Depool db1165.eqiad.wmnet to then clone it to db1279.eqiad.wmnet - marostegui@cumin1003 * 09:42 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1165: Depool db1165.eqiad.wmnet to then clone it to db1279.eqiad.wmnet - marostegui@cumin1003 * 09:42 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1165.eqiad.wmnet onto db1279.eqiad.wmnet * 09:41 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=testwiki --start-timestamp="20260311000000" --end-timestamp="20260805000000" --sleep="5" --batch-size="2"` for [[phab:T434688|T434688]] * 09:40 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for backupmon1001.eqiad.wmnet * 09:40 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for backupmon1001.eqiad.wmnet * 09:40 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1151 * 09:40 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1151 * 09:37 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325408{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]], [[gerrit:1325407{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]] (duration: 06m 57s) * 09:36 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host an-worker1151 * 09:36 btullis@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1151.eqiad.wmnet 13.36.64.10.in-addr.arpa 3.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:36 btullis@cumin1003: START - Cookbook sre.dns.wipe-cache an-worker1151.eqiad.wmnet 13.36.64.10.in-addr.arpa 3.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:36 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:36 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1151 - btullis@cumin1003" * 09:36 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on backupmon1001.eqiad.wmnet with reason: reboot * 09:36 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1151 - btullis@cumin1003" * 09:35 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host moss-be2003.codfw.wmnet * 09:33 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 7 hosts * 09:33 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 7 hosts * 09:33 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 09:32 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1325408{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]], [[gerrit:1325407{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1161.eqiad.wmnet onto db1275.eqiad.wmnet * 09:31 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 09:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1161: Pool db1161.eqiad.wmnet in after cloning * 09:30 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1325408{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]], [[gerrit:1325407{{!}}Improve performance of DB query in BackfillAbuseReview.php (T434688)]] * 09:29 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 09:27 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 09:27 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host moss-be2003.codfw.wmnet * 09:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be2006.codfw.wmnet * 09:25 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2004.codfw.wmnet * 09:22 hashar@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 09:21 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be2006.codfw.wmnet * 09:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be2005.codfw.wmnet * 09:19 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2004.codfw.wmnet * 09:19 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2003.codfw.wmnet * 09:18 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 7 hosts with reason: reboot * 09:18 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 7 hosts * 09:18 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 7 hosts * 09:16 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2214: Upgrade of db2214.codfw.wmnet completed * 09:14 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be2005.codfw.wmnet * 09:13 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apus-be2004.codfw.wmnet * 09:12 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2003.codfw.wmnet * 09:12 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2002.codfw.wmnet * 09:09 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2214: Upgrading db2214.codfw.wmnet * 09:09 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2214: Upgrading db2214.codfw.wmnet * 09:09 cwilliams@cumin1003: START - Cookbook sre.mysql.upgrade for 1 hosts * 09:07 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host apus-be2004.codfw.wmnet * 09:06 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2002.codfw.wmnet * 09:05 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1004.eqiad.wmnet * 09:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-cluster (exit_code=0) * 09:03 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 7 hosts with reason: reboot * 09:01 btullis@cumin1003: START - Cookbook sre.dns.netbox * 09:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2214 [[phab:T434754|T434754]]', diff saved to https://phabricator.wikimedia.org/P96045 and previous config saved to /var/cache/conftool/dbconfig/20260813-090001-cwilliams.json * 08:59 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1004.eqiad.wmnet * 08:59 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1003.eqiad.wmnet * 08:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2229 to s6 primary [[phab:T434754|T434754]]', diff saved to https://phabricator.wikimedia.org/P96044 and previous config saved to /var/cache/conftool/dbconfig/20260813-085752-cwilliams.json * 08:57 cezmunsta: Starting s6 codfw failover from db2214 to db2229 - [[phab:T434754|T434754]] * 08:54 hashar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325416{{!}}Revert "REST: Enable `GET /lexemes/<nowiki>{</nowiki>lexeme_id<nowiki>}</nowiki>` by default" (T434712)]] (duration: 07m 22s) * 08:53 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1003.eqiad.wmnet * 08:53 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1002.eqiad.wmnet * 08:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2229 with weight 0 [[phab:T434754|T434754]]', diff saved to https://phabricator.wikimedia.org/P96043 and previous config saved to /var/cache/conftool/dbconfig/20260813-085151-cwilliams.json * 08:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 23 hosts with reason: Primary switchover s6 [[phab:T434754|T434754]] * 08:51 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts2002.codfw.wmnet * 08:50 hashar@deploy1003: hashar: Continuing with deployment * 08:49 hashar@deploy1003: hashar: Backport for [[gerrit:1325416{{!}}Revert "REST: Enable `GET /lexemes/<nowiki>{</nowiki>lexeme_id<nowiki>}</nowiki>` by default" (T434712)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:47 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet * 08:47 hashar@deploy1003: Started scap sync-world: Backport for [[gerrit:1325416{{!}}Revert "REST: Enable `GET /lexemes/<nowiki>{</nowiki>lexeme_id<nowiki>}</nowiki>` by default" (T434712)]] * 08:47 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1002.eqiad.wmnet * 08:46 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1161: Pool db1161.eqiad.wmnet in after cloning * 08:46 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host stewards2001.codfw.wmnet * 08:45 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit1003.wikimedia.org * 08:45 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host stewards1001.eqiad.wmnet * 08:45 Emperor: roll-restart apus frontends in codfw * 08:45 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-cluster * 08:44 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts2002.codfw.wmnet * 08:44 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host doc2003.codfw.wmnet * 08:43 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet * 08:43 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host phab2003.codfw.wmnet * 08:42 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host stewards2001.codfw.wmnet * 08:42 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host etherpad2002.codfw.wmnet * 08:41 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host stewards1001.eqiad.wmnet * 08:41 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1273: Pool in s7 * 08:41 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1003.wikimedia.org * 08:40 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host doc2003.codfw.wmnet * 08:40 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit1003.wikimedia.org * 08:39 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host etherpad1004.eqiad.wmnet * 08:39 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host doc1004.eqiad.wmnet * 08:39 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit2002.wikimedia.org * 08:38 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host etherpad2002.codfw.wmnet * 08:37 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host phab2003.codfw.wmnet * 08:36 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lists1004.wikimedia.org * 08:35 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host etherpad1004.eqiad.wmnet * 08:35 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host doc1004.eqiad.wmnet * 08:34 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-cluster (exit_code=0) * 08:34 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1003.wikimedia.org * 08:34 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host planet2003.codfw.wmnet * 08:34 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2003.wikimedia.org * 08:33 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host planet1003.eqiad.wmnet * 08:33 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit2002.wikimedia.org * 08:32 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 22 hosts with reason: Cloning * 08:32 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2002.wikimedia.org * 08:31 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aphlict1002.eqiad.wmnet * 08:30 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host planet2003.codfw.wmnet * 08:29 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host planet1003.eqiad.wmnet * 08:29 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host lists1004.wikimedia.org * 08:28 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lists2001.wikimedia.org * 08:28 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2003.wikimedia.org * 08:27 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host aphlict1002.eqiad.wmnet * 08:27 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aphlict2001.codfw.wmnet * 08:26 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2002.wikimedia.org * 08:23 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host aphlict2001.codfw.wmnet * 08:23 btullis@cumin1003: START - Cookbook sre.hosts.move-vlan for host an-worker1151 * 08:22 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1151.eqiad.wmnet with OS bookworm * 08:22 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host lists2001.wikimedia.org * 08:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1161: Depool db1161.eqiad.wmnet to then clone it to db1275.eqiad.wmnet - marostegui@cumin1003 * 08:20 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1161: Depool db1161.eqiad.wmnet to then clone it to db1275.eqiad.wmnet - marostegui@cumin1003 * 08:20 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1161.eqiad.wmnet onto db1275.eqiad.wmnet * 08:15 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-cluster * 08:15 Emperor: roll-restart apus frontends in eqiad * 07:56 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1273: Pool in s7 * 07:56 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1273 to dbctl [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P96035 and previous config saved to /var/cache/conftool/dbconfig/20260813-075611-marostegui.json * 07:38 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db1273.eqiad.wmnet with reason: Reboot * 07:31 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.sanitize-wiki (exit_code=97) Managing sanitization for wikis testwiki in section s3 * 07:24 marostegui@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis testwiki in section s3 * 07:19 jayme@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on kubestagemaster2005.codfw.wmnet with reason: downtime because of hardware failure and no DRBD * 05:42 arnaudb@dns1006: END - running authdns-update * 05:40 arnaudb@dns1006: START - running authdns-update * 05:27 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 04:06 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 02:29 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1151 * 02:29 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1151 * 02:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1219.eqiad.wmnet with OS bookworm * 02:18 ryankemper: [[phab:T434494|T434494]] `ryankemper@deploy1003:~$ echo 'https://stats.wikimedia.org/' {{!}} mwscript-k8s --attach -- purgeList.php` (default page got cached during yesterday's `an-web1001` reimage) * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 46s) * 02:03 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1219.eqiad.wmnet with reason: host reimage * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 02:00 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1219.eqiad.wmnet with reason: host reimage * 01:46 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1219.eqiad.wmnet with OS bookworm * 01:03 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1142.eqiad.wmnet with OS bookworm * 00:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1142.eqiad.wmnet with reason: host reimage * 00:34 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1142.eqiad.wmnet with reason: host reimage * 00:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1142.eqiad.wmnet with OS bookworm * 00:16 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1178 * 00:16 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host an-worker1178 == 2026-08-12 == * 23:18 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324409{{!}}Remove $wmg = $wg hacks in Collection (T119117)]] (duration: 06m 43s) * 23:14 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 23:13 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1324409{{!}}Remove $wmg = $wg hacks in Collection (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:11 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1324409{{!}}Remove $wmg = $wg hacks in Collection (T119117)]] * 22:59 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 22:49 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260801000000" --end-timestamp="20260802000000" --sleep=2 --batch-size=10` * 22:45 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=testwiki --start-timestamp="20200801010101" --end-timestamp="20260816010101" --sleep=15 --batch-size=5` * 22:40 Dreamy_Jazz: Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=testwiki --start-timestamp="20200101010101" --end-timestamp="20260816010101" --sleep=60` * 22:35 Dreamy_Jazz: Running `mwscript WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260101000000" --end-timestamp="20260102000000" --sleep=10` * 22:21 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324817{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]], [[gerrit:1324818{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]] (duration: 45m 29s) * 22:17 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 21:59 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 21:40 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1324817{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]], [[gerrit:1324818{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:39 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 21:36 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1324817{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]], [[gerrit:1324818{{!}}Create a maintenance script to backfill Special:AbuseReview (T434688)]] * 21:32 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:31 vriley@cumin1003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:30 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:30 vriley@cumin1003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:17 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:14 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:14 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:11 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:10 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-codfw: Set storage compatability to NONE — [[phab:T433028|T433028]] - eevans@cumin1003 * 21:10 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1006 * 21:09 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1006 * 21:05 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324277{{!}}Improve Math preference labels for SVG/MathJax/MathML (T433891)]] (duration: 31m 42s) * 20:58 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1178.eqiad.wmnet with OS bookworm * 20:54 krinkle@deploy1003: krinkle: Continuing with deployment * 20:51 krinkle@deploy1003: krinkle: Backport for [[gerrit:1324277{{!}}Improve Math preference labels for SVG/MathJax/MathML (T433891)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:39 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 20:38 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:38 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp5022.eqsin.wmnet with OS trixie * 20:38 cdobbins@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - cdobbins@cumin1003" * 20:37 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:36 cdobbins@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - cdobbins@cumin1003" * 20:34 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1178.eqiad.wmnet with reason: host reimage * 20:34 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1324277{{!}}Improve Math preference labels for SVG/MathJax/MathML (T433891)]] * 20:33 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 20:28 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1178.eqiad.wmnet with reason: host reimage * 20:13 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1178.eqiad.wmnet with OS bookworm * 20:11 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-worker1178.eqiad.wmnet with OS bookworm * 20:11 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-worker1178.eqiad.wmnet with OS bookworm * 20:09 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-codfw: Set storage compatability to NONE — [[phab:T433028|T433028]] - eevans@cumin1003 * 20:09 Dreamy_Jazz: Evening UTC backport window done * 20:08 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324794{{!}}WikimediaAntiAbuse: Enable logging channel (T431292)]] (duration: 06m 48s) * 20:08 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage * 20:05 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage * 20:04 dreamyjazz@deploy1003: kharlan, dreamyjazz: Continuing with deployment * 20:04 dreamyjazz@deploy1003: kharlan, dreamyjazz: Backport for [[gerrit:1324794{{!}}WikimediaAntiAbuse: Enable logging channel (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:01 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1324794{{!}}WikimediaAntiAbuse: Enable logging channel (T431292)]] * 19:55 brennen@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 19:47 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 19:35 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 19:35 cdobbins@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cp5022.eqsin.wmnet with OS trixie * 19:32 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:30 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:29 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:26 vriley@cumin1003: START - Cookbook sre.dns.netbox * 19:23 brennen@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 19:19 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-eqiad: Set storage compatability to NONE — [[phab:T433028|T433028]] - eevans@cumin1003 * 19:10 brennen@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 19:09 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 19:09 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 19:00 Amir1: data migrated on wikishared ([[phab:T426102|T426102]]) * 18:57 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321224{{!}}Rename ce_worklist_articles table to ce_invitation_list_articles (T426102)]] (duration: 06m 50s) * 18:53 ladsgroup@deploy1003: ladsgroup, daimona: Continuing with deployment * 18:53 ladsgroup@deploy1003: ladsgroup, daimona: Backport for [[gerrit:1321224{{!}}Rename ce_worklist_articles table to ce_invitation_list_articles (T426102)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:51 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1321224{{!}}Rename ce_worklist_articles table to ce_invitation_list_articles (T426102)]] * 18:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2192: Security update * 18:31 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 18:25 Amir1: ce_invitation_list_articles created as empty on wikishared ([[phab:T426102|T426102]]) * 18:21 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 18:21 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 18:18 Amir1: migrated testwiki entries from ce_worklist_articles to ce_invitation_list_articles ([[phab:T426102|T426102]]) * 18:18 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-eqiad: Set storage compatability to NONE — [[phab:T433028|T433028]] - eevans@cumin1003 * 18:11 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T434668|T434668]] * 18:08 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling reboot on A:durum-eqsin and A:durum * 18:07 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Set storage compatability to UPGRADING — [[phab:T433028|T433028]] - eevans@cumin1003 * 18:05 jhancock@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022'] * 17:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2192: Security update * 17:55 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:55 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum-eqsin and A:durum * 17:53 jhancock@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['cp5022'] * 17:47 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:47 jhancock@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['cp5022'] * 17:42 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:41 jhancock@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['cp5022'] * 17:36 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:35 jhancock@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022'] * 17:31 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-magru and not (P<nowiki>{</nowiki>cp7001*<nowiki>}</nowiki> or P<nowiki>{</nowiki>cp7009*<nowiki>}</nowiki>) and A:cp - 9.2.15 upgrade ([[phab:T434620|T434620]]) * 17:28 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:28 jhancock@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['cp5022'] * 17:22 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 17:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2192.codfw.wmnet with reason: Maintenance * 17:11 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host apifeatureusage2001.codfw.wmnet with OS bookworm * 17:04 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: apply * 17:03 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-main: apply * 17:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2192 [[phab:T434635|T434635]]', diff saved to https://phabricator.wikimedia.org/P96030 and previous config saved to /var/cache/conftool/dbconfig/20260812-170338-cwilliams.json * 17:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2213 to s5 primary [[phab:T434635|T434635]]', diff saved to https://phabricator.wikimedia.org/P96029 and previous config saved to /var/cache/conftool/dbconfig/20260812-170152-cwilliams.json * 17:01 cezmunsta: Starting s5 codfw failover from db2192 to db2213 - [[phab:T434635|T434635]] * 16:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2213 with weight 0 [[phab:T434635|T434635]]', diff saved to https://phabricator.wikimedia.org/P96028 and previous config saved to /var/cache/conftool/dbconfig/20260812-165544-cwilliams.json * 16:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 27 hosts with reason: Primary switchover s5 [[phab:T434635|T434635]] * 16:53 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: apply * 16:52 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-main: apply * 16:44 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-main: apply * 16:44 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-main: apply * 16:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1272: New host * 16:40 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324765{{!}}WikimediaAntiAbuse: Enable PersonalInfoFlagNotifications (T431292)]] (duration: 07m 02s) * 16:40 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-logging-external: apply * 16:39 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-logging-external: apply * 16:38 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-logging-external: apply * 16:37 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-logging-external: apply * 16:36 kharlan@deploy1003: kharlan: Continuing with deployment * 16:35 kharlan@deploy1003: kharlan: Backport for [[gerrit:1324765{{!}}WikimediaAntiAbuse: Enable PersonalInfoFlagNotifications (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:33 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1324765{{!}}WikimediaAntiAbuse: Enable PersonalInfoFlagNotifications (T431292)]] * 16:27 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-logging-external: apply * 16:27 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-logging-external: apply * 16:18 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 16:18 jhancock@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022'] * 16:17 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 16:16 jhancock@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['cp5022'] * 16:15 jhancock@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] * 16:12 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Set storage compatability to UPGRADING — [[phab:T433028|T433028]] - eevans@cumin1003 * 16:10 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: apply * 16:10 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: apply * 16:08 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: apply * 16:08 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: apply * 16:08 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics: apply * 16:07 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics: apply * 16:02 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-magru and not (P<nowiki>{</nowiki>cp7001*<nowiki>}</nowiki> or P<nowiki>{</nowiki>cp7009*<nowiki>}</nowiki>) and A:cp - 9.2.15 upgrade ([[phab:T434620|T434620]]) * 15:57 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1272: New host * 15:52 jmm@cumin2003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti2046.codfw.wmnet * 15:52 jmm@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host ganeti2046.codfw.wmnet * 15:42 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:37 mforns@deploy1003: Finished deploy [analytics/refinery@49c336c] (thin): Regular analytics weekly train THIN [analytics/refinery@49c336cd] (duration: 01m 59s) * 15:37 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324733{{!}}WikimediaAntiAbuse: Enable personal info tag display on enwiki (T431292)]] (duration: 08m 12s) * 15:35 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:35 mforns@deploy1003: Started deploy [analytics/refinery@49c336c] (thin): Regular analytics weekly train THIN [analytics/refinery@49c336cd] * 15:34 mforns@deploy1003: Finished deploy [analytics/refinery@49c336c]: Regular analytics weekly train [analytics/refinery@49c336cd] (duration: 04m 20s) * 15:33 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[2024,1031]*.wmnet: Set storage compatability to UPGRADING — [[phab:T433028|T433028]] - eevans@cumin1003 * 15:33 dreamyjazz@deploy1003: kharlan, dreamyjazz: Continuing with deployment * 15:31 dreamyjazz@deploy1003: kharlan, dreamyjazz: Backport for [[gerrit:1324733{{!}}WikimediaAntiAbuse: Enable personal info tag display on enwiki (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:30 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitize-wiki (exit_code=99) Checking sanitization for wikis testwiki in section s3 * 15:30 mforns@deploy1003: Started deploy [analytics/refinery@49c336c]: Regular analytics weekly train [analytics/refinery@49c336cd] * 15:30 mforns@deploy1003: Finished deploy [analytics/refinery@49c336c] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@49c336cd] (duration: 00m 32s) * 15:29 mforns@deploy1003: Started deploy [analytics/refinery@49c336c] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@49c336cd] * 15:29 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1324733{{!}}WikimediaAntiAbuse: Enable personal info tag display on enwiki (T431292)]] * 15:27 brennen@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324700{{!}}EventDetailsParticipantsModule: populate cache with non-local users (T434597)]] (duration: 06m 38s) * 15:23 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[2024,1031]*.wmnet: Set storage compatability to UPGRADING — [[phab:T433028|T433028]] - eevans@cumin1003 * 15:23 brennen@deploy1003: brennen, daimona: Continuing with deployment * 15:23 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:22 brennen@deploy1003: brennen, daimona: Backport for [[gerrit:1324700{{!}}EventDetailsParticipantsModule: populate cache with non-local users (T434597)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:22 cgoubert@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host rdb-lock2003.codfw.wmnet * 15:21 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 15:21 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 15:21 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:21 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:21 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:20 brennen@deploy1003: Started scap sync-world: Backport for [[gerrit:1324700{{!}}EventDetailsParticipantsModule: populate cache with non-local users (T434597)]] * 15:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-worker1178.eqiad.wmnet with OS bookworm * 15:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock1003.eqiad.wmnet * 15:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock1003.eqiad.wmnet with OS trixie * 15:16 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 15:16 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 15:16 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 15:16 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:16 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:16 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:12 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 15:12 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2003.codfw.wmnet * 15:11 cgoubert@cumin2003: END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host rdb-lock2003.codfw.wmnet * 15:11 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 15:11 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 15:11 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:11 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:11 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:07 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324338{{!}}InitialiseSettings: Enable 2FA warnings on more private wikis (T428103)]] (duration: 07m 02s) * 15:04 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors * 15:04 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors * 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock1003.eqiad.wmnet with reason: host reimage * 15:03 reedy@deploy1003: reedy: Continuing with deployment * 15:02 reedy@deploy1003: reedy: Backport for [[gerrit:1324338{{!}}InitialiseSettings: Enable 2FA warnings on more private wikis (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:02 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" * 15:00 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324338{{!}}InitialiseSettings: Enable 2FA warnings on more private wikis (T428103)]] * 14:57 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 14:57 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2003.codfw.wmnet * 14:57 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock1003.eqiad.wmnet with reason: host reimage * 14:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock2002.codfw.wmnet * 14:57 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock2002.codfw.wmnet with OS trixie * 14:56 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 14:56 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 14:56 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 14:55 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 14:55 moritzm: powercycle ganeti2046 * 14:47 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock1003.eqiad.wmnet with OS trixie * 14:46 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1003.eqiad.wmnet - cgoubert@cumin2003" * 14:46 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1003.eqiad.wmnet - cgoubert@cumin2003" * 14:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock1003.eqiad.wmnet on all recursors * 14:45 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock1003.eqiad.wmnet on all recursors * 14:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:45 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1003.eqiad.wmnet - cgoubert@cumin2003" * 14:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1272.eqiad.wmnet with reason: Enabling notifications * 14:44 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1003.eqiad.wmnet - cgoubert@cumin2003" * 14:44 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324719{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324720{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324722{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0 (T434187)]] (duration: 11m 02s) * 14:39 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 14:39 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock1003.eqiad.wmnet * 14:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock2002.codfw.wmnet with reason: host reimage * 14:37 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 14:37 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock1002.eqiad.wmnet * 14:37 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock1002.eqiad.wmnet with OS trixie * 14:37 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1324719{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324720{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324722{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0 (T434187)]] synced to the testservers (see https://wikitech. * 14:33 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock2002.codfw.wmnet with reason: host reimage * 14:33 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1324719{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324720{{!}}Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324722{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0 (T434187)]] * 14:32 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1018.eqiad.wmnet with OS bookworm * 14:32 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2046.codfw.wmnet * 14:31 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1020.eqiad.wmnet with OS bookworm * 14:27 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2046.codfw.wmnet * 14:25 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2045.codfw.wmnet * 14:25 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2045.codfw.wmnet * 14:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock1002.eqiad.wmnet with reason: host reimage * 14:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1019.eqiad.wmnet with OS bookworm * 14:22 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Checking sanitization for wikis testwiki in section s3 * 14:20 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2045.codfw.wmnet * 14:18 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock1002.eqiad.wmnet with reason: host reimage * 14:17 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:16 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324707{{!}}Backport all changes from wmf/1.47.0-wmf.15]] (duration: 40m 51s) * 14:16 moritzm: installing Linux 6.1.180 on Bookworm hosts * 14:15 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock2002.codfw.wmnet with OS trixie * 14:14 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2002.codfw.wmnet - cgoubert@cumin2003" * 14:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2002.codfw.wmnet - cgoubert@cumin2003" * 14:14 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:14 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2045.codfw.wmnet * 14:14 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2002.codfw.wmnet on all recursors * 14:14 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2002.codfw.wmnet on all recursors * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2002.codfw.wmnet - cgoubert@cumin2003" * 14:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2002.codfw.wmnet - cgoubert@cumin2003" * 14:12 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2030.codfw.wmnet * 14:12 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2030.codfw.wmnet * 14:11 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on apifeatureusage2001.codfw.wmnet with reason: host reimage * 14:09 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:08 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:07 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 14:06 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2030.codfw.wmnet * 14:06 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:06 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock1002.eqiad.wmnet with OS trixie * 14:05 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 14:05 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2002.codfw.wmnet * 14:05 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1002.eqiad.wmnet - cgoubert@cumin2003" * 14:05 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1002.eqiad.wmnet - cgoubert@cumin2003" * 14:05 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock1002.eqiad.wmnet on all recursors * 14:05 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock1002.eqiad.wmnet on all recursors * 14:05 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:05 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1002.eqiad.wmnet - cgoubert@cumin2003" * 14:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:04 kharlan@deploy1003: kharlan: Continuing with deployment * 14:04 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock2001.codfw.wmnet * 14:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock2001.codfw.wmnet with OS trixie * 14:02 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1002.eqiad.wmnet - cgoubert@cumin2003" * 14:02 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on apifeatureusage2001.codfw.wmnet with reason: host reimage * 14:01 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2030.codfw.wmnet * 13:59 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2029.codfw.wmnet * 13:58 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2029.codfw.wmnet * 13:58 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 13:58 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock1002.eqiad.wmnet * 13:56 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1019.eqiad.wmnet with reason: host reimage * 13:54 btullis@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on archiva1002.wikimedia.org with reason: Upgrading in-place * 13:53 cgoubert@cumin2003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock1001.eqiad.wmnet * 13:53 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock1001.eqiad.wmnet with OS trixie * 13:53 kharlan@deploy1003: kharlan: Backport for [[gerrit:1324707{{!}}Backport all changes from wmf/1.47.0-wmf.15]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:52 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2029.codfw.wmnet * 13:52 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1020.eqiad.wmnet with reason: host reimage * 13:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1019.eqiad.wmnet with reason: host reimage * 13:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1020.eqiad.wmnet with reason: host reimage * 13:49 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock2001.codfw.wmnet with reason: host reimage * 13:48 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2029.codfw.wmnet * 13:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Configuring db1272 for s3 pooling', diff saved to https://phabricator.wikimedia.org/P96021 and previous config saved to /var/cache/conftool/dbconfig/20260812-134732-cwilliams.json * 13:44 bking@cumin2003: START - Cookbook sre.hosts.reimage for host apifeatureusage2001.codfw.wmnet with OS bookworm * 13:43 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock2001.codfw.wmnet with reason: host reimage * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2028.codfw.wmnet * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2028.codfw.wmnet * 13:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1018.eqiad.wmnet with reason: host reimage * 13:38 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock1001.eqiad.wmnet with reason: host reimage * 13:36 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1018.eqiad.wmnet with reason: host reimage * 13:35 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1324707{{!}}Backport all changes from wmf/1.47.0-wmf.15]] * 13:35 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2028.codfw.wmnet * 13:32 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1275938{{!}}Enable campaignEvents on bdwikimedia (T424016)]] (duration: 07m 35s) * 13:32 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1020.eqiad.wmnet with OS bookworm * 13:32 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1019.eqiad.wmnet with OS bookworm * 13:32 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock1001.eqiad.wmnet with reason: host reimage * 13:28 kharlan@deploy1003: kharlan, yahya: Continuing with deployment * 13:27 kharlan@deploy1003: kharlan, yahya: Backport for [[gerrit:1275938{{!}}Enable campaignEvents on bdwikimedia (T424016)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:27 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2028.codfw.wmnet * 13:26 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 13:26 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 13:25 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1275938{{!}}Enable campaignEvents on bdwikimedia (T424016)]] * 13:25 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitize-wiki (exit_code=99) Managing sanitization for wikis testwiki in section s3 * 13:24 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock2001.codfw.wmnet with OS trixie * 13:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2001.codfw.wmnet - cgoubert@cumin2003" * 13:24 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2001.codfw.wmnet - cgoubert@cumin2003" * 13:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2001.codfw.wmnet on all recursors * 13:23 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock2001.codfw.wmnet on all recursors * 13:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:23 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2001.codfw.wmnet - cgoubert@cumin2003" * 13:23 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2001.codfw.wmnet - cgoubert@cumin2003" * 13:23 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324713{{!}}thwiki: reinstate temporary wiki25 logos (T431094)]] (duration: 07m 13s) * 13:20 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:20 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1018.eqiad.wmnet with OS bookworm * 13:20 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2027.codfw.wmnet * 13:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2027.codfw.wmnet * 13:19 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics-external: apply * 13:18 kharlan@deploy1003: anzx, kharlan: Continuing with deployment * 13:18 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:18 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host rdb-lock1001.eqiad.wmnet with OS trixie * 13:17 kharlan@deploy1003: anzx, kharlan: Backport for [[gerrit:1324713{{!}}thwiki: reinstate temporary wiki25 logos (T431094)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1001.eqiad.wmnet - cgoubert@cumin2003" * 13:17 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 13:17 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1001.eqiad.wmnet - cgoubert@cumin2003" * 13:17 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock2001.codfw.wmnet * 13:17 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics-external: apply * 13:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock1001.eqiad.wmnet on all recursors * 13:17 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache rdb-lock1001.eqiad.wmnet on all recursors * 13:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1001.eqiad.wmnet - cgoubert@cumin2003" * 13:17 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1001.eqiad.wmnet - cgoubert@cumin2003" * 13:15 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1324713{{!}}thwiki: reinstate temporary wiki25 logos (T431094)]] * 13:15 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:15 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics-external: apply * 13:15 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics-external: apply * 13:14 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics-external: apply * 13:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2027.codfw.wmnet * 13:12 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 13:12 cgoubert@cumin2003: START - Cookbook sre.ganeti.makevm for new host rdb-lock1001.eqiad.wmnet * 13:12 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2027.codfw.wmnet * 13:08 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1158.eqiad.wmnet onto db1273.eqiad.wmnet * 13:07 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1158: Pool db1158.eqiad.wmnet in after cloning * 13:02 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti6002.drmrs.wmnet * 13:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti6002.drmrs.wmnet * 12:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1015.eqiad.wmnet with OS bookworm * 12:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti6002.drmrs.wmnet * 12:44 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1017.eqiad.wmnet with OS bookworm * 12:36 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 12:35 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 12:34 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 12:33 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 12:31 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti6002.drmrs.wmnet * 12:24 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1017.eqiad.wmnet with reason: host reimage * 12:22 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1158: Pool db1158.eqiad.wmnet in after cloning * 12:18 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1017.eqiad.wmnet with reason: host reimage * 12:11 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1015.eqiad.wmnet with reason: host reimage * 12:07 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1015.eqiad.wmnet with reason: host reimage * 12:04 moritzm: failover ganeti master in drmrs02 to ganeti6004 * 12:02 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1009.eqiad.wmnet with OS bookworm * 12:01 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1017.eqiad.wmnet with OS bookworm * 12:00 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti6004.drmrs.wmnet * 12:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti6004.drmrs.wmnet * 11:53 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti6004.drmrs.wmnet * 11:53 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1015.eqiad.wmnet with OS bookworm * 11:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1159.eqiad.wmnet onto db1274.eqiad.wmnet * 11:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Pool db1159.eqiad.wmnet in after cloning * 11:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1016.eqiad.wmnet with OS bookworm * 11:46 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti6004.drmrs.wmnet * 11:45 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti6001.drmrs.wmnet * 11:45 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti6001.drmrs.wmnet * 11:43 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324306{{!}}WikimediaAntiAbuse: Enable personal info for enwiki with no display (T431292)]] (duration: 10m 26s) * 11:42 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1015.eqiad.wmnet with OS bookworm * 11:39 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 11:38 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti6001.drmrs.wmnet * 11:34 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1324306{{!}}WikimediaAntiAbuse: Enable personal info for enwiki with no display (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:33 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2212: Security update * 11:33 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti6001.drmrs.wmnet * 11:32 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1324306{{!}}WikimediaAntiAbuse: Enable personal info for enwiki with no display (T431292)]] * 11:22 moritzm: failover ganeti master in drmrs01 to ganeti6003 * 11:20 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:20 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:18 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 22 hosts with reason: Cloning * 11:17 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti6003.drmrs.wmnet * 11:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti6003.drmrs.wmnet * 11:17 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1009.eqiad.wmnet with reason: host reimage * 11:17 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:16 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:14 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1016.eqiad.wmnet with reason: host reimage * 11:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti6003.drmrs.wmnet * 11:11 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1009.eqiad.wmnet with reason: host reimage * 11:10 jmm@cumin2003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti3005.esams.wmnet * 11:10 jmm@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host ganeti3005.esams.wmnet * 11:09 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 11:08 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1158: Depool db1158.eqiad.wmnet to then clone it to db1273.eqiad.wmnet - marostegui@cumin1003 * 11:07 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1016.eqiad.wmnet with reason: host reimage * 11:07 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1158: Depool db1158.eqiad.wmnet to then clone it to db1273.eqiad.wmnet - marostegui@cumin1003 * 11:07 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1158.eqiad.wmnet onto db1273.eqiad.wmnet * 11:06 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:05 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:05 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1159: Pool db1159.eqiad.wmnet in after cloning * 11:04 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 20 hosts with reason: Cloning * 11:02 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti6003.drmrs.wmnet * 11:00 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis testwiki in section s3 * 10:54 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1009.eqiad.wmnet with OS bookworm * 10:51 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1015.eqiad.wmnet with OS bookworm * 10:50 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1016.eqiad.wmnet with OS bookworm * 10:48 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2212: Security update * 10:45 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitize-wiki (exit_code=99) Managing sanitization for wikis testwiki in section s3 * 10:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1013.eqiad.wmnet with OS bookworm * 10:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1014.eqiad.wmnet with OS bookworm * 10:38 cwilliams@cumin1003: START - Cookbook sre.mysql.clone of db1159.eqiad.wmnet onto db1274.eqiad.wmnet * 10:33 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db1274.eqiad.wmnet * 10:33 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db1274.eqiad.wmnet * 10:31 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 396993 * 10:29 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 396993 * 10:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1159: Clone source for db1274 * 10:24 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1159: Clone source for db1274 * 10:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1013.eqiad.wmnet with reason: host reimage * 10:18 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1013.eqiad.wmnet with reason: host reimage * 10:13 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2212.codfw.wmnet with reason: Maintenance * 10:13 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 10:12 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 10:12 blake@deploy1003: Stopping before sync operations * 10:11 blake@deploy1003: Started scap sync-world: Non-deployment scap run to populate new release values for [[phab:T427668|T427668]] * 10:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2212 [[phab:T434644|T434644]]', diff saved to https://phabricator.wikimedia.org/P96003 and previous config saved to /var/cache/conftool/dbconfig/20260812-101053-cwilliams.json * 10:09 moritzm: powercycle ganeti3005 * 10:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2203 to s1 primary [[phab:T434644|T434644]]', diff saved to https://phabricator.wikimedia.org/P96002 and previous config saved to /var/cache/conftool/dbconfig/20260812-100849-cwilliams.json * 10:08 cezmunsta: Starting s1 codfw failover from db2212 to db2203 - [[phab:T434644|T434644]] * 10:03 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1013.eqiad.wmnet with OS bookworm * 10:02 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1013.eqiad.wmnet with OS bookworm * 10:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2203 with weight 0 [[phab:T434644|T434644]]', diff saved to https://phabricator.wikimedia.org/P96001 and previous config saved to /var/cache/conftool/dbconfig/20260812-100134-cwilliams.json * 10:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 32 hosts with reason: Primary switchover s1 [[phab:T434644|T434644]] * 09:53 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1014.eqiad.wmnet with reason: host reimage * 09:50 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 09:50 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti3005.esams.wmnet * 09:49 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1014.eqiad.wmnet with reason: host reimage * 09:41 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1278: Pool in x1 * 09:41 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-launcher1003.eqiad.wmnet with OS bookworm * 09:37 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti3005.esams.wmnet * 09:37 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1013.eqiad.wmnet with OS bookworm * 09:34 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-presto1013.eqiad.wmnet with OS bookworm * 09:33 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1014.eqiad.wmnet with OS bookworm * 09:29 moritzm: failover ganeti master in esams to ganeti3008 * 09:26 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti3006.esams.wmnet * 09:26 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti3006.esams.wmnet * 09:24 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1012.eqiad.wmnet with OS bookworm * 09:23 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1009.eqiad.wmnet with OS bookworm * 09:23 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 09:18 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti3006.esams.wmnet * 09:16 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti3006.esams.wmnet * 09:14 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: fix regexp escaping bug - oblivian@cumin1003" * 09:14 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: fix regexp escaping bug - oblivian@cumin1003 * 09:13 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: fix regexp escaping bug - oblivian@cumin1003 * 09:13 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: fix regexp escaping bug - oblivian@cumin1003" * 09:03 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-launcher1003.eqiad.wmnet with reason: host reimage * 08:58 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-launcher1003.eqiad.wmnet with reason: host reimage * 08:55 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1278: Pool in x1 * 08:55 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1278 to dbctl [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95996 and previous config saved to /var/cache/conftool/dbconfig/20260812-085521-marostegui.json * 08:51 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1012.eqiad.wmnet with reason: host reimage * 08:45 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis testwiki in section s3 * 08:43 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1009.eqiad.wmnet with OS bookworm * 08:42 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1012.eqiad.wmnet with reason: host reimage * 08:41 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-launcher1003.eqiad.wmnet with OS bookworm * 08:40 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1013.eqiad.wmnet with OS bookworm * 08:38 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms3', diff saved to https://phabricator.wikimedia.org/P95995 and previous config saved to /var/cache/conftool/dbconfig/20260812-083816-marostegui.json * 08:38 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-master1003.eqiad.wmnet with OS bookworm * 08:37 marostegui: Failover ms3 [[phab:T434288|T434288]] * 08:37 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1268 to dbctl [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95994 and previous config saved to /var/cache/conftool/dbconfig/20260812-083722-marostegui.json * 08:35 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-web1001.eqiad.wmnet with OS bookworm * 08:32 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db2252.codfw.wmnet,db[1153,1268].eqiad.wmnet with reason: Switching over ms3 * 08:28 marostegui@cumin1003: dbctl commit (dc=all): 'Depool ms3 [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95993 and previous config saved to /var/cache/conftool/dbconfig/20260812-082852-marostegui.json * 08:25 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1012.eqiad.wmnet with OS bookworm * 08:25 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1011.eqiad.wmnet with OS bookworm * 08:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-master1003.eqiad.wmnet with reason: host reimage * 08:07 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-master1003.eqiad.wmnet with reason: host reimage * 08:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-web1001.eqiad.wmnet with reason: host reimage * 07:58 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-web1001.eqiad.wmnet with reason: host reimage * 07:50 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1003.eqiad.wmnet with OS bookworm * 07:38 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1011.eqiad.wmnet with reason: host reimage * 07:38 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-master1003.eqiad.wmnet with OS bookworm * 07:35 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti3007.esams.wmnet * 07:35 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti3007.esams.wmnet * 07:33 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1011.eqiad.wmnet with reason: host reimage * 07:27 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti3007.esams.wmnet * 07:25 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti3007.esams.wmnet * 07:25 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti3008.esams.wmnet * 07:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti3008.esams.wmnet * 07:22 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-web1001.eqiad.wmnet with OS bookworm * 07:19 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1003.eqiad.wmnet with OS bookworm * 07:18 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-master1003.eqiad.wmnet * 07:18 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host an-master1003.eqiad.wmnet * 07:17 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1011.eqiad.wmnet with OS bookworm * 07:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 07:16 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti3008.esams.wmnet * 07:15 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1009.eqiad.wmnet with OS bookworm * 07:14 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host an-master1003.eqiad.wmnet * 07:13 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-master1003.eqiad.wmnet * 07:13 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-master1003.eqiad.wmnet * 07:12 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-master1003.eqiad.wmnet * 07:11 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti3008.esams.wmnet * 07:07 arnaudb@dns1006: END - running authdns-update * 07:07 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti5007.eqsin.wmnet * 07:07 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti5007.eqsin.wmnet * 07:05 arnaudb@dns1006: START - running authdns-update * 06:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti5007.eqsin.wmnet * 06:54 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti5007.eqsin.wmnet * 06:38 moritzm: failover ganeti master in eqsin to ganeti5004 * 06:36 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti5006.eqsin.wmnet * 06:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti5006.eqsin.wmnet * 06:28 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti5006.eqsin.wmnet * 06:23 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti5006.eqsin.wmnet * 06:20 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti5005.eqsin.wmnet * 06:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti5005.eqsin.wmnet * 06:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti5005.eqsin.wmnet * 06:06 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti5005.eqsin.wmnet * 06:03 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti5004.eqsin.wmnet * 06:03 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti5004.eqsin.wmnet * 05:55 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti5004.eqsin.wmnet * 05:53 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti5004.eqsin.wmnet * 04:40 ryankemper: [[phab:T434494|T434494]] reimaged `an-tool1008.eqiad.wmnet` to bookworm; yarn.wikimedia.org is back up * 04:16 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-tool1008.eqiad.wmnet with OS bookworm * 03:58 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-tool1008.eqiad.wmnet with reason: host reimage * 03:53 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-tool1008.eqiad.wmnet with reason: host reimage * 03:41 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host an-tool1008.eqiad.wmnet with OS bookworm * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 45s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 00:25 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324427{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]], [[gerrit:1324429{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0]], [[gerrit:1324428{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]] (duration: 07m 55s) * 00:21 kemayo@deploy1003: kemayo: Continuing with deployment * 00:19 kemayo@deploy1003: kemayo: Backport for [[gerrit:1324427{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]], [[gerrit:1324429{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0]], [[gerrit:1324428{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:17 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1324427{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]], [[gerrit:1324429{{!}}Bump mediawiki/mediawiki-codesniffer to v52.0.0]], [[gerrit:1324428{{!}}ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]] == 2026-08-11 == * 21:37 sbassett: Deployed security fix for [[phab:T434521|T434521]] (wmf.15) * 21:29 sbassett: Deployed security fix for [[phab:T434521|T434521]] (wmf.14) * 21:19 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324370{{!}}Phase 4 of legal footer deployment (T432796)]], [[gerrit:1319804{{!}}Disable wgMFCustomSiteModules on English Wikipedia (T375538)]] (duration: 15m 26s) * 21:15 jdlrobson@deploy1003: jdlrobson: Continuing with deployment * 21:06 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1324370{{!}}Phase 4 of legal footer deployment (T432796)]], [[gerrit:1319804{{!}}Disable wgMFCustomSiteModules on English Wikipedia (T375538)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:03 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1324370{{!}}Phase 4 of legal footer deployment (T432796)]], [[gerrit:1319804{{!}}Disable wgMFCustomSiteModules on English Wikipedia (T375538)]] * 20:59 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 20:50 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324384{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]], [[gerrit:1324385{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]] (duration: 06m 58s) * 20:46 kemayo@deploy1003: kemayo: Continuing with deployment * 20:45 kemayo@deploy1003: kemayo: Backport for [[gerrit:1324384{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]], [[gerrit:1324385{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:43 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1324384{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]], [[gerrit:1324385{{!}}Editcheck: ecenable causing errors when editCheckTagging is active (T434403)]] * 20:42 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 20:42 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324386{{!}}build: Updating js-yaml to 3.15.1, 4.3.1]] (duration: 07m 36s) * 20:38 kemayo@deploy1003: kemayo: Continuing with deployment * 20:37 jhancock@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 20:37 kemayo@deploy1003: kemayo: Backport for [[gerrit:1324386{{!}}build: Updating js-yaml to 3.15.1, 4.3.1]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:35 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1324386{{!}}build: Updating js-yaml to 3.15.1, 4.3.1]] * 20:18 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 20:15 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 20:15 jhancock@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin1003" * 20:14 jhancock@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin1003" * 19:59 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 19:54 jhancock@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage * 19:10 brennen@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] (duration: 06m 41s) * 19:04 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns6001.wikimedia.org * 19:04 sukhe@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns6001.wikimedia.org * 19:04 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns5003.wikimedia.org * 19:04 sukhe@cumin1003: START - Cookbook sre.hosts.remove-downtime for dns5003.wikimedia.org * 19:03 brennen@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 18:59 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns5003.wikimedia.org with OS trixie * 18:55 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns6001.wikimedia.org with OS trixie * 18:19 brett@cumin2002: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on P<nowiki>{</nowiki>cp7009.magru.wmnet<nowiki>}</nowiki> and A:cp - 9.2.15 Upgrade () * 18:14 brett@cumin2002: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on P<nowiki>{</nowiki>cp7009.magru.wmnet<nowiki>}</nowiki> and A:cp - 9.2.15 Upgrade () * 18:13 brennen@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 18:12 brett@cumin2002: END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 9.2.15 Upgrade () * 18:09 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns5003.wikimedia.org with reason: host reimage * 18:06 brett@cumin2002: START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 9.2.15 Upgrade () * 18:06 brennen: 1.47.0-wmf.15 train status ([[phab:T430834|T430834]]) - no current blockers, rolling to group0 * 18:05 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns5003.wikimedia.org with reason: host reimage * 18:05 brett: import trafficserver-9.2.15~deb13+wmf1 into trixie-wikimedia ([[phab:T434478|T434478]]) * 18:01 ladsgroup@cumin1003: END (PASS) - Cookbook sre.mysql.sanitarium_restart (exit_code=0) * 17:58 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns6001.wikimedia.org with reason: host reimage * 17:53 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324369{{!}}Enable desktop lazy loading on group0 (T148047)]] (duration: 07m 31s) * 17:52 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns6001.wikimedia.org with reason: host reimage * 17:49 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 17:49 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitarium_restart (exit_code=99) * 17:49 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 17:49 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7001.magru.wmnet * 17:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti7001.magru.wmnet * 17:48 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 17:47 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1324369{{!}}Enable desktop lazy loading on group0 (T148047)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:45 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1324369{{!}}Enable desktop lazy loading on group0 (T148047)]] * 17:39 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti7001.magru.wmnet * 17:36 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns5003.wikimedia.org with OS trixie * 17:34 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns6001.wikimedia.org with OS trixie * 17:31 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324361{{!}}Move FR config from IS.php to a dedicated file]], [[gerrit:1324363{{!}}Remove $wmg = $wg hacks in CentralAuth (T119117)]] (duration: 12m 23s) * 17:26 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 17:23 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1324361{{!}}Move FR config from IS.php to a dedicated file]], [[gerrit:1324363{{!}}Remove $wmg = $wg hacks in CentralAuth (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:19 sukhe: sudo cumin "A:cp-magru" "run-puppet-agent --enable 'merging CR 1324355'" * 17:18 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1324361{{!}}Move FR config from IS.php to a dedicated file]], [[gerrit:1324363{{!}}Remove $wmg = $wg hacks in CentralAuth (T119117)]] * 17:11 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-master1004.eqiad.wmnet with OS bookworm * 17:11 sukhe: sukhe@cp7005:~$ sudo puppet agent -tv * 17:02 sukhe: sudo cumin "A:cp-magru" "disable-puppet 'merging CR 1324355'" * 16:54 sukhe@dns1004: END - running authdns-update * 16:53 sukhe@dns1004: START - running authdns-update * 16:53 sukhe@dns1004: FAIL - running authdns-update * 16:51 sukhe@dns1004: START - running authdns-update * 16:48 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-master1004.eqiad.wmnet with reason: host reimage * 16:44 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-master1004.eqiad.wmnet with reason: host reimage * 16:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1157.eqiad.wmnet onto db1272.eqiad.wmnet * 16:40 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1157: Pool db1157.eqiad.wmnet in after cloning * 16:38 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 16:31 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324356{{!}}InitialiseSettings: Fix wgOATHAuthEnforce2FAForAll]] (duration: 06m 52s) * 16:30 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 16:28 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 16:27 reedy@deploy1003: reedy: Continuing with deployment * 16:26 reedy@deploy1003: reedy: Backport for [[gerrit:1324356{{!}}InitialiseSettings: Fix wgOATHAuthEnforce2FAForAll]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:24 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324356{{!}}InitialiseSettings: Fix wgOATHAuthEnforce2FAForAll]] * 16:13 sukhe: restart ntpsec.serviceon dns7001 * 16:09 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324335{{!}}InitialiseSettings: Enable 2FA enforcement on various private wikis (T428103)]] (duration: 06m 40s) * 16:08 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 16:06 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2204: Security update * 16:04 reedy@deploy1003: reedy: Continuing with deployment * 16:04 reedy@deploy1003: reedy: Backport for [[gerrit:1324335{{!}}InitialiseSettings: Enable 2FA enforcement on various private wikis (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:02 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7001.magru.wmnet * 16:02 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324335{{!}}InitialiseSettings: Enable 2FA enforcement on various private wikis (T428103)]] * 16:01 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 15:55 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1157: Pool db1157.eqiad.wmnet in after cloning * 15:54 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 15:54 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 15:41 dancy@deploy1003: Finished scap sync-world: Testing (duration: 06m 28s) * 15:40 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-master1004.eqiad.wmnet with OS bookworm * 15:40 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 15:35 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti4008.ulsfo.wmnet * 15:35 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti4008.ulsfo.wmnet * 15:34 dancy@deploy1003: Started scap sync-world: Testing * 15:34 dancy@deploy1003: Installation of scap version "4.279.0" completed for 3 hosts * 15:34 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-master1004.eqiad.wmnet with OS bookworm * 15:34 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 15:33 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-master1004.eqiad.wmnet with OS bookworm * 15:32 dancy@deploy1003: Installing scap version "4.279.0" for 3 host(s) * 15:32 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324339{{!}}Add /w/deployment-info.php entrypoint]] (duration: 07m 25s) * 15:30 moritzm: failover ganeti master in magru to ganeti7004 * 15:29 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti4008.ulsfo.wmnet * 15:28 dancy@deploy1003: dancy: Continuing with deployment * 15:28 tappof: remove 2026-05 swift log archives from centrallog to free some space ([[phab:T434502|T434502]]) * 15:27 dancy@deploy1003: dancy: Backport for [[gerrit:1324339{{!}}Add /w/deployment-info.php entrypoint]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:25 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324339{{!}}Add /w/deployment-info.php entrypoint]] * 15:20 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2204: Security update * 15:18 dancy@deploy1003: Installation of scap version "4.278.0" completed for 3 hosts * 15:18 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7004.magru.wmnet * 15:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti7004.magru.wmnet * 15:16 dancy@deploy1003: Installing scap version "4.278.0" for 3 host(s) * 15:14 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2204.codfw.wmnet with reason: Maintenance * 15:12 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 15:11 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-master1004.eqiad.wmnet with OS bookworm * 15:11 hashar: Restarting CI Jenkins on contint1003 due to Java upgrade. * 15:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2204 [[phab:T434565|T434565]]', diff saved to https://phabricator.wikimedia.org/P95984 and previous config saved to /var/cache/conftool/dbconfig/20260811-151126-cwilliams.json * 15:10 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti4008.ulsfo.wmnet * 15:10 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti7004.magru.wmnet * 15:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2207 to s2 primary [[phab:T434565|T434565]]', diff saved to https://phabricator.wikimedia.org/P95983 and previous config saved to /var/cache/conftool/dbconfig/20260811-150905-cwilliams.json * 15:08 cezmunsta: Starting s2 codfw failover from db2204 to db2207 - [[phab:T434565|T434565]] * 15:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2207 with weight 0 [[phab:T434565|T434565]]', diff saved to https://phabricator.wikimedia.org/P95982 and previous config saved to /var/cache/conftool/dbconfig/20260811-150402-cwilliams.json * 15:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s2 [[phab:T434565|T434565]] * 14:55 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1010.eqiad.wmnet with OS bookworm * 14:49 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns4003.wikimedia.org with OS trixie * 14:47 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1010.eqiad.wmnet with OS bookworm * 14:47 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7001.wikimedia.org with OS trixie * 14:44 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-presto1010.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:41 btullis@cumin1003: START - Cookbook sre.hosts.provision for host an-presto1010.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:40 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-presto1010.eqiad.wmnet with OS bookworm * 14:39 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 14:39 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-presto1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:36 btullis@cumin1003: START - Cookbook sre.hosts.provision for host an-presto1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:32 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1009.eqiad.wmnet with OS bookworm * 14:32 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 14:31 cwilliams@cumin1003: START - Cookbook sre.mysql.clone of db1157.eqiad.wmnet onto db1272.eqiad.wmnet * 14:30 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7004.magru.wmnet * 14:28 moritzm: failover ganeti master in ulsfo to ganeti4005 * 14:23 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7003.magru.wmnet * 14:23 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti7003.magru.wmnet * 14:23 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-master1004.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:22 btullis@cumin1003: START - Cookbook sre.hosts.provision for host an-master1004.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:21 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-master1004.eqiad.wmnet with OS bookworm * 14:19 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti4007.ulsfo.wmnet * 14:19 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti4007.ulsfo.wmnet * 14:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1010.eqiad.wmnet with OS bookworm * 14:17 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1008.eqiad.wmnet with OS bookworm * 14:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti7003.magru.wmnet * 14:12 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7003.magru.wmnet * 14:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti4007.ulsfo.wmnet * 14:11 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7002.magru.wmnet * 14:11 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti7002.magru.wmnet * 14:09 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324318{{!}}Revert "wmf-config/ProductionServices: set URL for urldownloader to service record" (T429175)]] (duration: 06m 46s) * 14:09 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-master1004.eqiad.wmnet with OS bookworm * 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7001.wikimedia.org with reason: host reimage * 14:06 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti4007.ulsfo.wmnet * 14:05 kharlan@deploy1003: kharlan: Continuing with deployment * 14:05 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti4006.ulsfo.wmnet * 14:04 jayme: updated calico to v3.30.7 on staging-eqiad - [[phab:T427400|T427400]] * 14:04 kharlan@deploy1003: kharlan: Backport for [[gerrit:1324318{{!}}Revert "wmf-config/ProductionServices: set URL for urldownloader to service record" (T429175)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:04 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti4006.ulsfo.wmnet * 14:03 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns4003.wikimedia.org with reason: host reimage * 14:03 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7001.wikimedia.org with reason: host reimage * 14:02 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti7002.magru.wmnet * 14:02 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1324318{{!}}Revert "wmf-config/ProductionServices: set URL for urldownloader to service record" (T429175)]] * 14:02 btullis@dns1004: FAIL - running authdns-update * 14:00 btullis@dns1004: START - running authdns-update * 13:59 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1008.eqiad.wmnet with reason: host reimage * 13:59 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'. * 13:59 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1313985{{!}}wmf-config/ProductionServices: set URL for urldownloader to service record (T429175)]] (duration: 25m 06s) * 13:58 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7002.magru.wmnet * 13:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti4006.ulsfo.wmnet * 13:57 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns4003.wikimedia.org with reason: host reimage * 13:56 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'. * 13:56 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7001.magru.wmnet * 13:56 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1008.eqiad.wmnet with reason: host reimage * 13:55 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'. * 13:55 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'. * 13:55 kharlan@deploy1003: kharlan, sukhe: Continuing with deployment * 13:53 marostegui: Failover ms2 [[phab:T434288|T434288]] * 13:52 marostegui: Failover ms1 [[phab:T434288|T434288]] * 13:52 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7001.magru.wmnet * 13:51 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti4006.ulsfo.wmnet * 13:48 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti4005.ulsfo.wmnet * 13:48 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti4005.ulsfo.wmnet * 13:44 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti4005.ulsfo.wmnet * 13:40 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1008.eqiad.wmnet with OS bookworm * 13:39 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns4003.wikimedia.org with OS trixie * 13:38 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns7001.wikimedia.org with OS trixie * 13:38 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1008.eqiad.wmnet with OS bookworm * 13:37 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti4005.ulsfo.wmnet * 13:36 kharlan@deploy1003: kharlan, sukhe: Backport for [[gerrit:1313985{{!}}wmf-config/ProductionServices: set URL for urldownloader to service record (T429175)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:34 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1313985{{!}}wmf-config/ProductionServices: set URL for urldownloader to service record (T429175)]] * 13:29 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2034.codfw.wmnet * 13:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2034.codfw.wmnet * 13:25 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.clone (exit_code=99) of db1157.eqiad.wmnet onto db1272.eqiad.wmnet * 13:25 cwilliams@cumin1003: START - Cookbook sre.mysql.clone of db1157.eqiad.wmnet onto db1272.eqiad.wmnet * 13:21 urbanecm@deploy1003: mwscript-k8s job started: namespaceDupes.php --wiki=frwiktionary --fix # [[phab:T415716|T415716]] * 13:21 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2034.codfw.wmnet * 13:20 urbanecm@deploy1003: mwscript-k8s job started: namespaceDupes.php --wiki=frwiktionary # [[phab:T415716|T415716]] * 13:19 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1323350{{!}}[tgwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T415307)]], [[gerrit:1322961{{!}}[slwiki] Revert temporary logo for Wikipedia 25 (Vector legacy + Vector 2022) (T414265)]], [[gerrit:1323827{{!}}[itwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T414320)]] (duration: 08m 00s) * 13:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-coord1003.eqiad.wmnet with OS bookworm * 13:15 urbanecm@deploy1003: urbanecm, superpes: Continuing with deployment * 13:13 urbanecm@deploy1003: urbanecm, superpes: Backport for [[gerrit:1323350{{!}}[tgwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T415307)]], [[gerrit:1322961{{!}}[slwiki] Revert temporary logo for Wikipedia 25 (Vector legacy + Vector 2022) (T414265)]], [[gerrit:1323827{{!}}[itwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T414320)]] synced to the testservers (see https://wiki * 13:11 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1323350{{!}}[tgwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T415307)]], [[gerrit:1322961{{!}}[slwiki] Revert temporary logo for Wikipedia 25 (Vector legacy + Vector 2022) (T414265)]], [[gerrit:1323827{{!}}[itwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy) (T414320)]] * 13:11 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1323779{{!}}[ukwiki] Remove reviewer usergroup (T434252)]], [[gerrit:1323312{{!}}[frwiktionary] Add new Schème namespace and its talk (T415716)]] (duration: 06m 49s) * 13:10 marostegui@dns1004: END - running authdns-update * 13:08 marostegui@dns1004: START - running authdns-update * 13:07 marostegui@cumin1003: dbctl commit (dc=all): 'Repool ms2 [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95980 and previous config saved to /var/cache/conftool/dbconfig/20260811-130725-marostegui.json * 13:06 urbanecm@deploy1003: urbanecm, superpes: Continuing with deployment * 13:06 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1266 to dbctl [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95979 and previous config saved to /var/cache/conftool/dbconfig/20260811-130627-marostegui.json * 13:06 urbanecm@deploy1003: urbanecm, superpes: Backport for [[gerrit:1323779{{!}}[ukwiki] Remove reviewer usergroup (T434252)]], [[gerrit:1323312{{!}}[frwiktionary] Add new Schème namespace and its talk (T415716)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:04 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1323779{{!}}[ukwiki] Remove reviewer usergroup (T434252)]], [[gerrit:1323312{{!}}[frwiktionary] Add new Schème namespace and its talk (T415716)]] * 12:59 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db2253.codfw.wmnet,db[1151,1266].eqiad.wmnet with reason: Switching over ms2 * 12:54 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1157: Using as clone source * 12:53 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1157: Using as clone source * 12:51 marostegui@cumin1003: dbctl commit (dc=all): 'Depool ms2 [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95977 and previous config saved to /var/cache/conftool/dbconfig/20260811-125129-marostegui.json * 12:47 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 12:46 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 12:45 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 12:44 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'recommendation-api-ng' for release 'main' . * 12:44 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'recommendation-api-ng' for release 'main' . * 12:43 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'recommendation-api-ng' for release 'main' . * 12:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-coord1003.eqiad.wmnet with reason: host reimage * 12:43 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'ores-legacy' for release 'main' . * 12:42 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'ores-legacy' for release 'main' . * 12:42 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2165: Security update * 12:40 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-coord1003.eqiad.wmnet with reason: host reimage * 12:39 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'ores-legacy' for release 'main' . * 12:38 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' . * 12:38 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' . * 12:37 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' . * 12:34 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2034.codfw.wmnet * 12:30 jmm@cumin2002: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti-test2001.codfw.wmnet * 12:30 jmm@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host ganeti-test2001.codfw.wmnet * 12:25 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1179.eqiad.wmnet onto db1278.eqiad.wmnet * 12:25 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1179: Pool db1179.eqiad.wmnet in after cloning * 12:23 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-coord1003.eqiad.wmnet with OS bookworm * 12:22 moritzm: failover ganeti master in codfw/routed to ganeti2033 * 12:22 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2033.codfw.wmnet * 12:22 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2033.codfw.wmnet * 12:19 jmm@cumin2002: START - Cookbook sre.hosts.reboot-single for host ganeti-test2001.codfw.wmnet * 12:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm * 12:18 jmm@cumin2002: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti-test2001.codfw.wmnet * 12:18 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1008.eqiad.wmnet with OS bookworm * 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti2033.codfw.wmnet * 12:07 moritzm: failover ganeti master in ganeti/test to ganeti-test2003 * 12:04 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 12:03 jmm@cumin2003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti4005.ulsfo.wmnet * 12:03 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti4005.ulsfo.wmnet * 12:00 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324283{{!}}Use maximum compression level in SqlBlobStore and SqlBagOStuff (T428377)]] (duration: 11m 37s) * 11:57 jmm@cumin2002: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti-test2002.codfw.wmnet * 11:57 jmm@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti-test2002.codfw.wmnet * 11:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2165: Security update * 11:54 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 11:52 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1324283{{!}}Use maximum compression level in SqlBlobStore and SqlBagOStuff (T428377)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:51 jmm@cumin2002: START - Cookbook sre.hosts.reboot-single for host ganeti-test2002.codfw.wmnet * 11:50 jmm@cumin2002: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti-test2002.codfw.wmnet * 11:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2165.codfw.wmnet with reason: Maintenance * 11:48 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@050d19e] (releasing): [[phab:T434186|T434186]] (duration: 01m 14s) * 11:48 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1324283{{!}}Use maximum compression level in SqlBlobStore and SqlBagOStuff (T428377)]] * 11:47 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@050d19e] (releasing): [[phab:T434186|T434186]] * 11:44 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@050d19e] (releasing): test jenkins deploy for [[phab:T434186|T434186]] (duration: 01m 08s) * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2165 [[phab:T434514|T434514]]', diff saved to https://phabricator.wikimedia.org/P95969 and previous config saved to /var/cache/conftool/dbconfig/20260811-114352-cwilliams.json * 11:43 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@050d19e] (releasing): test jenkins deploy for [[phab:T434186|T434186]] * 11:42 jmm@cumin2003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti-test2003.codfw.wmnet * 11:42 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti-test2003.codfw.wmnet * 11:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2161 to s8 primary [[phab:T434514|T434514]]', diff saved to https://phabricator.wikimedia.org/P95968 and previous config saved to /var/cache/conftool/dbconfig/20260811-114136-cwilliams.json * 11:40 cezmunsta: Starting s8 codfw failover from db2165 to db2161 - [[phab:T434514|T434514]] * 11:40 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1179: Pool db1179.eqiad.wmnet in after cloning * 11:36 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ganeti-test2003.codfw.wmnet * 11:36 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti-test2003.codfw.wmnet * 11:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2161 with weight 0 [[phab:T434514|T434514]]', diff saved to https://phabricator.wikimedia.org/P95966 and previous config saved to /var/cache/conftool/dbconfig/20260811-113449-cwilliams.json * 11:34 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 25 hosts with reason: Primary switchover s8 [[phab:T434514|T434514]] * 11:29 moritzm: installing Python 3.11 security updates * 11:28 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-presto1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 11:26 btullis@cumin1003: START - Cookbook sre.hosts.provision for host an-presto1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 11:23 btullis@dns1004: END - running authdns-update * 11:21 btullis@dns1004: START - running authdns-update * 11:20 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-presto1008.eqiad.wmnet with OS bookworm * 11:20 moritzm: installing curl security updates * 11:11 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1007.eqiad.wmnet with OS bookworm * 10:45 tappof: bump space for prometheus k8s-dse in codfw * 10:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1007.eqiad.wmnet with reason: host reimage * 10:38 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1007.eqiad.wmnet with reason: host reimage * 10:37 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-coord1004.eqiad.wmnet with OS bookworm * 10:35 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1008.eqiad.wmnet with OS bookworm * 10:34 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1179: Depool db1179.eqiad.wmnet to then clone it to db1278.eqiad.wmnet - marostegui@cumin1003 * 10:34 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1008.eqiad.wmnet with OS bookworm * 10:25 fceratto@cumin1003: dbctl commit (dc=all): 'Remove db1177 [[phab:T433474|T433474]]', diff saved to https://phabricator.wikimedia.org/P95964 and previous config saved to /var/cache/conftool/dbconfig/20260811-102527-fceratto.json * 10:22 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1008.eqiad.wmnet with OS bookworm * 10:22 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1007.eqiad.wmnet with OS bookworm * 10:21 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1006.eqiad.wmnet with OS bookworm * 10:20 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 10:18 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1179: Depool db1179.eqiad.wmnet to then clone it to db1278.eqiad.wmnet - marostegui@cumin1003 * 10:18 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1179.eqiad.wmnet onto db1278.eqiad.wmnet * 10:17 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 10:17 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 10:14 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 10:09 blake@deploy1003: Stopping before sync operations * 10:09 blake@deploy1003: Started scap sync-world: Non-deployment run to populate release values for [[phab:T427668|T427668]] * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 10:04 fceratto@cumin1003: Removing db1177 from zarcillo [[phab:T433474|T433474]] * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1177.eqiad.wmnet * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1177.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:03 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1177.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:00 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1006.eqiad.wmnet with reason: host reimage * 09:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-coord1004.eqiad.wmnet with reason: host reimage * 09:57 marostegui: Failover m1 from db1164 to db1213 - [[phab:T434493|T434493]] * 09:57 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1006.eqiad.wmnet with reason: host reimage * 09:55 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 09:54 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2232].codfw.wmnet,db[1164,1213,1217].eqiad.wmnet with reason: Primary switchover m1 [[phab:T434493|T434493]] * 09:52 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-coord1004.eqiad.wmnet with reason: host reimage * 09:49 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1213.eqiad.wmnet with OS trixie * 09:49 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1177.eqiad.wmnet * 09:41 moritzm: installing Linux 6.12.101 on Trixie hosts * 09:40 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-presto1006.eqiad.wmnet with OS bookworm * 09:35 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-coord1004.eqiad.wmnet with OS bookworm * 09:28 moritzm: installing node-tar security updates * 09:27 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1213.eqiad.wmnet with reason: host reimage * 09:22 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1213.eqiad.wmnet with reason: host reimage * 09:09 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1177: Decommission * 09:08 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db1177: Decommission * 09:08 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 09:08 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.decommission (exit_code=99) * 09:06 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1213.eqiad.wmnet with OS trixie * 09:06 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 09:05 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1213.eqiad.wmnet with reason: Reimage * 08:53 marostegui@dns1004: END - running authdns-update * 08:51 marostegui@dns1004: START - running authdns-update * 08:48 marostegui: Switchover ms1 master in eqiad [[phab:T434288|T434288]] * 08:48 marostegui@cumin1003: dbctl commit (dc=all): 'Repool ms1 [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95962 and previous config saved to /var/cache/conftool/dbconfig/20260811-084804-marostegui.json * 08:40 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1267 to dbctl [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95961 and previous config saved to /var/cache/conftool/dbconfig/20260811-084054-marostegui.json * 08:29 marostegui: Failover m1 from db1213 to db1164 - [[phab:T434043|T434043]] * 08:25 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2232].codfw.wmnet,db[1164,1213,1217].eqiad.wmnet with reason: Primary switchover m1 [[phab:T434043|T434043]] * 08:22 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db2251.codfw.wmnet,db[1152,1267].eqiad.wmnet with reason: Switching over ms1 * 08:22 marostegui@cumin1003: dbctl commit (dc=all): 'Depool ms1 [[phab:T434288|T434288]]', diff saved to https://phabricator.wikimedia.org/P95960 and previous config saved to /var/cache/conftool/dbconfig/20260811-082201-marostegui.json * 08:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: Switching over ms1 * 08:20 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.parsercache (exit_code=99) * 08:20 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 08:20 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1152: Switching over ms1 * 08:19 slyngshede@dns1004: END - running authdns-update * 08:18 moritzm: installing openjdk-21 security updates * 08:17 slyngshede@dns1004: START - running authdns-update * 08:16 moritzm: imported jenkins 2.568.2 to thirdparty/jenkins for trixie-wikimedia * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.12 (duration: 02m 26s) * 03:36 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] (duration: 33m 33s) * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.15 refs [[phab:T430834|T430834]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 35s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-10 == * 14:54 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1323973{{!}}mmv.bootstrap: Fix getUrlParam to account for TIFF lossy/lossless param (T434333)]] (duration: 11m 24s) * 14:50 krinkle@deploy1003: krinkle: Continuing with deployment * 14:45 krinkle@deploy1003: krinkle: Backport for [[gerrit:1323973{{!}}mmv.bootstrap: Fix getUrlParam to account for TIFF lossy/lossless param (T434333)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:43 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1323973{{!}}mmv.bootstrap: Fix getUrlParam to account for TIFF lossy/lossless param (T434333)]] * 14:07 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1323967{{!}}updateIsActiveFlagForMentees: Commit the final partial batch (T432959)]] (duration: 10m 33s) * 13:56 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1323967{{!}}updateIsActiveFlagForMentees: Commit the final partial batch (T432959)]] * 13:45 dani@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply * 13:45 dani@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply * 13:45 dani@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply * 13:45 dani@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply * 13:45 dani@deploy1003: helmfile [staging] DONE helmfile.d/services/miscweb: apply * 13:44 dani@deploy1003: helmfile [staging] START helmfile.d/services/miscweb: apply * 13:38 wmde-fisch@deploy1003: Finished scap sync-world: Backport for [[gerrit:1323939{{!}}Enable sub-references on more group2 wikis (batch3) (T432731)]] (duration: 33m 21s) * 13:25 wmde-fisch@deploy1003: wmde-fisch: Continuing with deployment * 13:22 wmde-fisch@deploy1003: wmde-fisch: Backport for [[gerrit:1323939{{!}}Enable sub-references on more group2 wikis (batch3) (T432731)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:05 wmde-fisch@deploy1003: Started scap sync-world: Backport for [[gerrit:1323939{{!}}Enable sub-references on more group2 wikis (batch3) (T432731)]] * 07:57 hashar@deploy1003: Finished deploy [integration/docroot@7772132]: update build dependencies (duration: 00m 13s) * 07:57 hashar@deploy1003: Started deploy [integration/docroot@7772132]: update build dependencies * 07:35 _joe_: restarting squid on urldownloader1006 * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 48s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-09 == * 16:01 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:01 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:01 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:00 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 36s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-08 == * 05:31 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9] (wcqs): [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) (duration: 02m 36s) * 05:28 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9] (wcqs): [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) * 04:56 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 04:55 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 04:47 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) (duration: 19m 22s) * 04:28 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) * 04:19 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) (duration: 00m 06s) * 04:18 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) * 04:17 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) (duration: 00m 28s) * 04:16 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) * 03:52 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 03:52 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 34s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-07 == * 23:30 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:29 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 22:45 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 22:43 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 22:41 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 22:41 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 20:54 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 20:32 andrewbogott: restarting puppetserver service on puppetserver* for [[phab:T434339|T434339]] * 19:52 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:45 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 19:32 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:25 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:22 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 19:21 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 18:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:41 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 18:35 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 18:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 18:22 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 18:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 18:16 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 18:12 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:09 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:08 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:07 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:04 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:01 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:00 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:00 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 17:59 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 17:25 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 17:14 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 17:13 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 17:13 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 17:13 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:54 maryum: Deployed security fix for [[phab:T434278|T434278]] * 16:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 16:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 16:27 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --verbose --scoreLessThan=0.7 --exceptDatasetChecksums=[[phab:T434319|T434319]]-enwiki-models.txt # [[phab:T434319|T434319]] * 16:06 cdobbins@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp5022.eqsin.wmnet with OS trixie * 15:13 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 14:19 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 14:17 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 13:50 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1156.eqiad.wmnet onto db1271.eqiad.wmnet * 13:50 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1271: Pool db1271.eqiad.wmnet in after cloning * 13:02 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1271: Pool db1271.eqiad.wmnet in after cloning * 12:19 jayme: updated calico to v3.30.7 on staging-codfw - [[phab:T427400|T427400]] * 12:09 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 12:06 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 12:06 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 12:05 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 12:02 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1156: Pool db1156.eqiad.wmnet in after cloning * 11:38 bjensen: sudo -i reprepro -C main include trixie-wikimedia $<nowiki>{</nowiki>HOME<nowiki>}</nowiki>/httpbb/trixie/httpbb_$<nowiki>{</nowiki>VERSION?<nowiki>}</nowiki>-1+deb13u1_amd64.changes #[[phab:T434052|T434052]] * 11:35 bjensen: sudo -i reprepro -C main include bookworm-wikimedia $<nowiki>{</nowiki>HOME<nowiki>}</nowiki>/httpbb/bookworm/httpbb_$<nowiki>{</nowiki>VERSION?<nowiki>}</nowiki>-1_amd64.changes #[[phab:T434052|T434052]] * 11:30 marostegui@cumin1003: dbctl commit (dc=all): 'Adding db1271 to dbctl', diff saved to https://phabricator.wikimedia.org/P95945 and previous config saved to /var/cache/conftool/dbconfig/20260807-113006-marostegui.json * 11:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1156: Pool db1156.eqiad.wmnet in after cloning * 10:23 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 10:22 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 10:22 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 10:21 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 10:20 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 10:20 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 10:19 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 10:18 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 10:06 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on 21 hosts with reason: cloning * 10:01 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1156: Depool db1156.eqiad.wmnet to then clone it to db1271.eqiad.wmnet - marostegui@cumin1003 * 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1156: Depool db1156.eqiad.wmnet to then clone it to db1271.eqiad.wmnet - marostegui@cumin1003 * 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1156.eqiad.wmnet onto db1271.eqiad.wmnet * 09:15 jynus: started stress testing db1245 dbs [[phab:T431115|T431115]] * 08:19 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:18 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:16 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:14 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:13 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:10 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:06 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:05 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:00 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 10 days, 0:00:00 on ml-serve1015.eqiad.wmnet with reason: Downtime to get full picture of current BIOS settings beyond what Redfish shows * 08:00 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 07:54 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 07:54 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 07:53 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:52 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:51 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:50 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:49 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:48 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:47 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:45 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:45 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:41 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:38 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:37 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 06:35 jayme: updated istio to 1.29.4 on wikikube eqiad - [[phab:T427401|T427401]] * 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1178.eqiad.wmnet * 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1178.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 06:06 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1178.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 05:55 marostegui@cumin1003: START - Cookbook sre.dns.netbox * 05:49 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1178.eqiad.wmnet * 05:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 05:46 marostegui@cumin1003: Removing db1178 from zarcillo [[phab:T433471|T433471]] * 05:45 marostegui@cumin1003: START - Cookbook sre.mysql.decommission * 02:42 denisse: Extended volume on prometheus2008 for the disk space alert as per https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_host_running_out_of_space * 02:37 denisse: Extended volume on prometheus2007 tor the disk space alert as per https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_host_running_out_of_space * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 56s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-06 == * 21:39 maryum: Deploy security patch for [[phab:T433070|T433070]] * 21:29 maryum: Deploy security patch for [[phab:T434189|T434189]] * 20:48 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] (duration: 08m 12s) * 20:44 aude@deploy1003: lmora, aude, anzx: Continuing with deployment * 20:41 aude@deploy1003: lmora, aude, anzx: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be * 20:41 ebernhardson@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:41 ebernhardson@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 20:40 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] * 20:37 ebernhardson@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:37 ebernhardson@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 20:32 ebernhardson@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:32 ebernhardson@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 20:31 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] (duration: 06m 41s) * 20:27 cjming@deploy1003: cjming, ebernhardson, chlod: Continuing with deployment * 20:26 cjming@deploy1003: cjming, ebernhardson, chlod: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:24 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] * 20:18 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] (duration: 09m 22s) * 20:14 cjming@deploy1003: cjming, tsev: Continuing with deployment * 20:11 cjming@deploy1003: cjming, tsev: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:09 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] * 19:41 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply * 19:40 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply * 19:31 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 19:31 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 19:00 cdobbins@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cp5022.eqsin.wmnet with OS trixie * 18:25 ladsgroup@deploy1003: Finished scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) (duration: 06m 08s) * 18:19 ladsgroup@deploy1003: Started scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) * 18:18 ladsgroup@deploy1003: Stopping before sync operations * 18:17 ladsgroup@deploy1003: Started scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) * 17:55 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 16:50 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 16:35 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1001.eqiad.wmnet with OS bookworm * 16:19 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1002.eqiad.wmnet with reason: host reimage * 16:16 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1002.eqiad.wmnet with reason: host reimage * 16:05 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1001.eqiad.wmnet with reason: host reimage * 16:00 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1001.eqiad.wmnet with reason: host reimage * 15:57 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 15:43 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm * 15:29 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1001.eqiad.wmnet with OS bookworm * 15:29 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:58 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm * 14:57 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-drmrs ([[phab:T428495|T428495]]) * 14:55 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-drmrs ([[phab:T428495|T428495]]) * 14:55 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-ui1001.eqiad.wmnet with OS bookworm * 14:54 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-presto1001.eqiad.wmnet with OS bookworm * 14:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-magru ([[phab:T428495|T428495]]) * 14:49 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-magru ([[phab:T428495|T428495]]) * 14:48 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1001.eqiad.wmnet with OS bookworm * 14:46 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-esams ([[phab:T428495|T428495]]) * 14:44 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-esams ([[phab:T428495|T428495]]) * 14:43 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:42 brouberol@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:42 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:42 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 14:40 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 14:40 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:38 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-ui1001.eqiad.wmnet with reason: host reimage * 14:34 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-presto1001.eqiad.wmnet with reason: host reimage * 14:28 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-ui1001.eqiad.wmnet with reason: host reimage * 14:27 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-presto1001.eqiad.wmnet with reason: host reimage * 14:23 sukhe: sudo cumin -b2 'A:cp-text' "run-puppet-agent --enable 'merging CR 1290731'": [[phab:T425441|T425441]] * 14:18 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo for hosts in the wikimedia.org domain - [[phab:T428495|T428495]] * 14:16 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-presto1001.eqiad.wmnet with OS bookworm * 14:14 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-ui1001.eqiad.wmnet with OS bookworm * 14:12 sukhe: sudo cumin 'A:cp-text' "disable-puppet 'merging CR 1290731'": [[phab:T425441|T425441]] * 14:11 swfrench-wmf: restarted navtiming on webperf1003 - [[phab:T428495|T428495]] * 14:04 swfrench-wmf: begin rolling restart of confd in drmrs, eqiad, esams, magru - [[phab:T428495|T428495]] * 14:04 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm * 14:04 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:02 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-client1002.eqiad.wmnet with OS bookworm * 13:58 swfrench-wmf: authdns update to direct eqiad-associated etcd clients back to eqiad - [[phab:T428495|T428495]] * 13:58 swfrench@dns1004: END - running authdns-update * 13:56 swfrench@dns1004: START - running authdns-update * 13:49 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:44 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:31 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 13:29 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 13:26 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 13:23 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 13:22 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 13:19 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 13:18 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 13:18 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 13:17 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 13:16 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 13:13 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 13:11 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 13:09 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 13:06 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 13:06 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-client1002.eqiad.wmnet with OS bookworm * 13:05 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revision-models' for release 'main' . * 13:05 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:05 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revision-models' for release 'main' . * 13:04 brouberol@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-test-client1002.eqiad.wmnet with OS bookworm * 13:04 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 13:03 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 13:02 aikochou@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:00 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'readability' for release 'main' . * 12:59 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'readability' for release 'main' . * 12:58 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 12:57 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'logo-detection' for release 'main' . * 12:57 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'logo-detection' for release 'main' . * 12:57 aikochou@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 12:55 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:54 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:53 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 12:53 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:50 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 12:48 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 12:46 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'article-models' for release 'main' . * 12:45 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'article-models' for release 'main' . * 12:41 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . * 12:39 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . * 12:38 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-client1002.eqiad.wmnet with OS bookworm * 12:12 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply * 12:12 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply * 12:09 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:08 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 11:58 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2187: Security update * 11:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:24 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:16 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 11:15 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 11:10 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2187: Security update * 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2187.codfw.wmnet with reason: Maintenance * 10:56 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 10:56 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 10:56 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 10:56 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 10:54 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 10:53 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 10:09 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2187: Security update * 10:07 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2187: Security update * 09:39 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms2', diff saved to https://phabricator.wikimedia.org/P95929 and previous config saved to /var/cache/conftool/dbconfig/20260806-093908-marostegui.json * 09:36 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1178 from dbctl [[phab:T433471|T433471]]', diff saved to https://phabricator.wikimedia.org/P95928 and previous config saved to /var/cache/conftool/dbconfig/20260806-093632-marostegui.json * 09:33 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 09:31 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 09:30 topranks: bounce cr3-eqsin<->cr2-eqiad bgp session to disable no-prepend command * 09:20 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2253.codfw.wmnet,db1151.eqiad.wmnet with reason: cloning * 09:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1151: Cloning * 09:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:19 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 09:19 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1151: Cloning * 09:10 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 09:09 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup2003.codfw.wmnet * 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup2003.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 09:06 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup2003.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 09:03 klausman@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:02 jynus@cumin1003: START - Cookbook sre.dns.netbox * 09:02 klausman@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 08:57 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup2003.codfw.wmnet * 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup1003.eqiad.wmnet * 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:54 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms3', diff saved to https://phabricator.wikimedia.org/P95925 and previous config saved to /var/cache/conftool/dbconfig/20260806-085422-marostegui.json * 08:53 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:46 jynus@cumin1003: START - Cookbook sre.dns.netbox * 08:39 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup1003.eqiad.wmnet * 08:29 XioNoX: push pfw policy - [[phab:T434115|T434115]] * 08:14 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 08:00 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 08:00 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:58 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revision-models' for release 'main' . * 07:56 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 07:54 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'readability' for release 'main' . * 07:53 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'logo-detection' for release 'main' . * 07:51 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'llm' for release 'main' . * 07:48 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . * 07:37 jayme: updated istio to 1.29.4 on wikikube codfw - [[phab:T427401|T427401]] * 07:08 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2252.codfw.wmnet,db1153.eqiad.wmnet with reason: cloning * 07:07 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1153: Cloning * 07:07 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1153: Cloning * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 40s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-05 == * 23:24 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1009.eqiad.wmnet with OS bookworm * 23:03 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1009.eqiad.wmnet with reason: host reimage * 22:59 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1009.eqiad.wmnet with reason: host reimage * 22:43 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1009.eqiad.wmnet with OS bookworm * 22:38 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1009.eqiad.wmnet * 22:34 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1009.eqiad.wmnet * 22:25 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1008.eqiad.wmnet with OS bookworm * 22:04 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1008.eqiad.wmnet with reason: host reimage * 22:00 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1008.eqiad.wmnet with reason: host reimage * 21:48 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:47 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:46 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:44 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1008.eqiad.wmnet with OS bookworm * 21:43 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:41 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1008.eqiad.wmnet * 21:36 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1008.eqiad.wmnet * 21:14 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:12 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad * 21:12 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad * 21:10 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=eqiad * 21:08 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:07 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:07 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=eqiad * 21:04 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:03 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1006 * 21:02 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1006 * 21:00 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:56 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:56 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:55 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 20:55 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 20:55 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1007.eqiad.wmnet with OS bookworm * 20:51 vriley@cumin1003: START - Cookbook sre.dns.netbox * 20:43 ebernhardson: [[phab:T434008|T434008]]: changing cloudelastic:9643 from auto_expand_replicas to number_of_replicas * 20:34 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1007.eqiad.wmnet with reason: host reimage * 20:27 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1007.eqiad.wmnet with reason: host reimage * 20:24 cjming: end of UTC late backport window * 20:23 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] (duration: 06m 26s) * 20:18 cjming@deploy1003: cjming: Continuing with deployment * 20:18 cjming@deploy1003: cjming: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:16 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] * 20:12 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1007.eqiad.wmnet with OS bookworm * 20:12 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] (duration: 08m 41s) * 20:08 swfrench@cumin2002: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host conf1007.eqiad.wmnet with OS bookworm * 20:08 jforrester@deploy1003: jforrester: Continuing with deployment * 20:07 jforrester@deploy1003: jforrester: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:03 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] * 19:51 inflatador: [bking@puppetserver1001] ~$ sudo puppetserver ca sign --certname an-worker1189.eqiad.wmnet [[phab:T434142|T434142]] * 19:47 bking@cumin2003: DONE (FAIL) - Cookbook sre.puppet.renew-cert (exit_code=99) for an-worker1189.eqiad.wmnet: Renew puppet certificate - bking@cumin2003 * 19:46 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:30 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1007.eqiad.wmnet with OS trixie * 19:30 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 19:29 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 19:20 swfrench-wmf: silenced EtcdRelicationDown 0cb709a9-f244-4f1e-971f-{{Gerrit|440ec65e7fd7}} - [[phab:T428495|T428495]] * 19:13 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1007.eqiad.wmnet with OS bookworm * 19:12 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1007.eqiad.wmnet with reason: host reimage * 19:09 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1007.eqiad.wmnet * 19:07 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1007.eqiad.wmnet with reason: host reimage * 19:03 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1007.eqiad.wmnet * 18:52 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1007.eqiad.wmnet with OS trixie * 18:52 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1007.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:35 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 18:34 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 18:34 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 18:30 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1007.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:28 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:28 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1007] - vriley@cumin1003" * 18:27 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1007] - vriley@cumin1003" * 18:23 vriley@cumin1003: START - Cookbook sre.dns.netbox * 18:22 vriley@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 18:22 robh@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:19 vriley@cumin1003: START - Cookbook sre.dns.netbox * 18:13 robh@cumin2002: START - Cookbook sre.hosts.provision for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:31 jasmine@cumin2002: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-main-eqiad * 17:12 mutante: LDAP - added vwalters to group ciadmin - [[phab:T433615|T433615]] * 16:58 aokoth@deploy1003: Finished deploy [phabricator/deployment@e2ebca5]: Deploy Phab (duration: 00m 34s) * 16:57 aokoth@deploy1003: Started deploy [phabricator/deployment@e2ebca5]: Deploy Phab * 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad * 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=eqiad * 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad * 16:53 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:41 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2187.codfw.wmnet * 16:41 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2187.codfw.wmnet * 16:41 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker2187.codfw.wmnet * 16:41 cgoubert@cumin2003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker2187.codfw.wmnet * 16:40 jasmine@cumin2002: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-main-eqiad * 16:40 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:34 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:25 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-magru and A:liberica ([[phab:T428495|T428495]]) * 16:23 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-magru and A:liberica ([[phab:T428495|T428495]]) * 16:20 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-drmrs and A:liberica ([[phab:T428495|T428495]]) * 16:19 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-drmrs and A:liberica ([[phab:T428495|T428495]]) * 16:18 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-esams and A:liberica ([[phab:T428495|T428495]]) * 16:16 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-esams and A:liberica ([[phab:T428495|T428495]]) * 16:06 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1159.eqiad.wmnet * 16:06 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1159.eqiad.wmnet * 16:06 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1159.eqiad.wmnet * 16:05 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] (duration: 09m 11s) * 15:58 reedy@deploy1003: reedy: Continuing with deployment * 15:58 reedy@deploy1003: reedy: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:56 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] * 15:54 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1159.eqiad.wmnet with OS trixie * 15:38 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:33 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1159.eqiad.wmnet with reason: host reimage * 15:32 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:27 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1159.eqiad.wmnet with reason: host reimage * 15:10 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1159 * 15:10 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1159 * 15:00 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo for hosts in the wikimedia.org domain - [[phab:T428495|T428495]] * 14:55 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS trixie * 14:54 swfrench-wmf: restarted navtiming on webperf1003 - [[phab:T428495|T428495]] * 14:52 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1159 * 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1159.eqiad.wmnet 129.48.64.10.in-addr.arpa 9.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:52 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1159.eqiad.wmnet 129.48.64.10.in-addr.arpa 9.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1159 - jayme@cumin1003" * 14:52 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1159 - jayme@cumin1003" * 14:49 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:48 jayme@cumin1003: START - Cookbook sre.dns.netbox * 14:47 swfrench-wmf: begin rolling restart of confd in drmrs, eqiad, esams, magru - [[phab:T428495|T428495]] * 14:47 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1159 * 14:46 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:46 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:46 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1159.eqiad.wmnet with OS trixie * 14:45 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:44 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:44 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1159.eqiad.wmnet * 14:43 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:43 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1159.eqiad.wmnet * 14:43 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:43 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1159.eqiad.wmnet * 14:43 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:43 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1157.eqiad.wmnet * 14:43 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1157.eqiad.wmnet * 14:43 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1157.eqiad.wmnet * 14:42 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:42 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:42 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:41 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:41 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:41 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:41 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1003.eqiad.wmnet with OS bookworm * 14:39 swfrench-wmf: authdns update to direct eqiad-associated etcd clients to codfw - [[phab:T428495|T428495]] * 14:39 swfrench@dns1004: END - running authdns-update * 14:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 14:37 swfrench@dns1004: START - running authdns-update * 14:37 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:37 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:35 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:35 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 14:28 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:27 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1157.eqiad.wmnet with OS trixie * 14:27 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:27 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:27 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 14:26 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 14:26 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:26 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 14:26 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 14:26 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:26 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host search-loader1002.eqiad.wmnet with OS trixie * 14:19 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:19 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:15 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1003.eqiad.wmnet with reason: host reimage * 14:14 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:14 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:13 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 * 14:13 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host mc2046 * 14:13 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS trixie * 14:11 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1003.eqiad.wmnet with reason: host reimage * 14:10 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:09 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:09 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:09 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:08 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:08 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1157.eqiad.wmnet with reason: host reimage * 14:08 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:04 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 14:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on search-loader1002.eqiad.wmnet with reason: host reimage * 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=eqiad * 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=eqiad * 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=eqiad * 14:00 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:59 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:58 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1157.eqiad.wmnet with reason: host reimage * 13:57 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on search-loader1002.eqiad.wmnet with reason: host reimage * 13:54 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1003.eqiad.wmnet with OS bookworm * 13:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host search-loader1002.eqiad.wmnet with OS trixie * 13:43 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1157 * 13:42 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1157 * 13:40 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1157 * 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1157.eqiad.wmnet 183.32.64.10.in-addr.arpa 3.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1157.eqiad.wmnet 183.32.64.10.in-addr.arpa 3.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1157 - jayme@cumin1003" * 13:39 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1157 - jayme@cumin1003" * 13:39 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] (duration: 07m 00s) * 13:35 jayme@cumin1003: START - Cookbook sre.dns.netbox * 13:35 reedy@deploy1003: reedy: Continuing with deployment * 13:34 reedy@deploy1003: reedy: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:32 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] * 13:23 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1157 * 13:22 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1157.eqiad.wmnet with OS trixie * 13:22 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1157.eqiad.wmnet * 13:22 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1157.eqiad.wmnet * 13:21 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1157.eqiad.wmnet * 13:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1156.eqiad.wmnet * 13:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1156.eqiad.wmnet * 13:15 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1156.eqiad.wmnet * 13:01 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1156.eqiad.wmnet with OS trixie * 12:42 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1156.eqiad.wmnet with reason: host reimage * 12:38 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1156.eqiad.wmnet with reason: host reimage * 12:32 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 12:31 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 12:30 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 12:28 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 12:26 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 12:24 topranks: update bgp confed settings in eqsin * 12:22 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1156 * 12:22 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1156 * 12:22 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 12:19 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1156 * 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1156.eqiad.wmnet 110.32.64.10.in-addr.arpa 0.1.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:19 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1156.eqiad.wmnet 110.32.64.10.in-addr.arpa 0.1.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1156 - jayme@cumin1003" * 12:19 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1156 - jayme@cumin1003" * 12:17 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:14 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS trixie * 12:09 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:06 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:04 jayme@cumin1003: START - Cookbook sre.dns.netbox * 12:04 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 12:02 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'article-models' for release 'main' . * 12:01 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1156 * 12:01 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1156.eqiad.wmnet with OS trixie * 11:59 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1156.eqiad.wmnet * 11:59 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1156.eqiad.wmnet * 11:59 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1156.eqiad.wmnet * 11:57 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 11:53 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 11:53 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:52 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:52 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:50 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:50 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:50 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:49 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:48 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:47 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:47 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:45 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:45 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:44 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:44 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:44 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:43 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:42 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:38 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:35 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 * 11:35 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host mc2046 * 11:34 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS trixie * 11:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:27 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:21 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:21 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:18 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:18 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:18 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:18 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:13 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:13 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:09 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:08 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:07 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:06 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:06 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:05 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:05 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:04 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 11:04 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:24 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:24 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:17 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:16 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1155.eqiad.wmnet * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1155.eqiad.wmnet * 10:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1155.eqiad.wmnet * 10:14 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:14 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:11 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:11 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:10 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:09 aikochou@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop: sync * 10:09 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:09 aikochou@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop: sync * 10:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:07 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:05 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:05 aikochou@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop: sync * 10:05 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:05 aikochou@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop: sync * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:04 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:04 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:04 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1155.eqiad.wmnet with OS trixie * 09:52 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms1', diff saved to https://phabricator.wikimedia.org/P95918 and previous config saved to /var/cache/conftool/dbconfig/20260805-095212-marostegui.json * 09:44 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1152: after cloning * 09:44 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.parsercache (exit_code=99) * 09:44 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 09:44 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1152: after cloning * 09:43 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1155.eqiad.wmnet with reason: host reimage * 09:40 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1155.eqiad.wmnet with reason: host reimage * 09:32 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 09:32 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:31 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 09:31 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:27 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1155 * 09:27 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1155 * 09:25 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 09:24 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 09:24 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 09:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:23 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 09:23 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 09:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:22 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 09:22 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 09:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:20 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 09:20 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:17 XioNoX: push pfw policies - [[phab:T434038|T434038]] * 09:14 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1155 * 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1155.eqiad.wmnet 109.32.64.10.in-addr.arpa 9.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1155.eqiad.wmnet 109.32.64.10.in-addr.arpa 9.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1155 - jayme@cumin1003" * 09:14 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1155 - jayme@cumin1003" * 09:10 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 09:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:09 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2251.codfw.wmnet,db1152.eqiad.wmnet with reason: cloning * 09:09 jayme@cumin1003: START - Cookbook sre.dns.netbox * 09:08 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 09:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: Cloning * 09:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:05 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 09:05 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1152: Cloning * 08:38 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1155 * 08:37 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1155.eqiad.wmnet with OS trixie * 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1171.eqiad.wmnet * 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1171.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:29 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95913 and previous config saved to /var/cache/conftool/dbconfig/20260805-082908-ladsgroup.json * 08:27 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1171.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:22 jynus@cumin1003: START - Cookbook sre.dns.netbox * 08:18 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249', diff saved to https://phabricator.wikimedia.org/P95912 and previous config saved to /var/cache/conftool/dbconfig/20260805-081823-ladsgroup.json * 08:17 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1171.eqiad.wmnet * 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1150.eqiad.wmnet * 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1150.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:15 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1150.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:15 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 08:14 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1155.eqiad.wmnet * 08:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1155.eqiad.wmnet * 08:14 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1155.eqiad.wmnet * 08:11 jynus@cumin1003: START - Cookbook sre.dns.netbox * 08:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249', diff saved to https://phabricator.wikimedia.org/P95911 and previous config saved to /var/cache/conftool/dbconfig/20260805-080737-ladsgroup.json * 08:05 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1150.eqiad.wmnet * 08:02 marostegui: Depool clouddb1020 (s5,s8) [[phab:T434048|T434048]] * 08:02 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1020.eqiad.wmnet,service=s8 * 08:02 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1020.eqiad.wmnet,service=s5 * 08:02 marostegui: Depool clouddb1018 (s2,s7) [[phab:T434048|T434048]] * 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1018.eqiad.wmnet,service=s7 * 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1018.eqiad.wmnet,service=s2 * 08:01 marostegui: Depool clouddb1017 (s1) [[phab:T434048|T434048]] * 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1017.eqiad.wmnet,service=s1 * 07:59 marostegui: Depool clouddb1016 (s5,s8) [[phab:T434048|T434048]] * 07:59 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s8 * 07:59 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s5 * 07:57 marostegui: Depool clouddb1015 (s4,s6) [[phab:T434048|T434048]] * 07:57 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s6 * 07:57 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s4 * 07:56 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95910 and previous config saved to /var/cache/conftool/dbconfig/20260805-075650-ladsgroup.json * 07:54 marostegui: Depool clouddb1014 (s2,s7) [[phab:T434048|T434048]] * 07:54 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1014.eqiad.wmnet,service=s7 * 07:54 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1014.eqiad.wmnet,service=s2 * 07:53 marostegui: Depool clouddb1013:s1 [[phab:T434048|T434048]] * 07:53 marostegui: Depool clouddb1013:s1 [[phab:T409557|T409557]] * 07:53 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1013.eqiad.wmnet,service=s1 * 07:25 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95909 and previous config saved to /var/cache/conftool/dbconfig/20260805-072529-ladsgroup.json * 07:24 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2249.codfw.wmnet with reason: Maintenance * 07:24 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95908 and previous config saved to /var/cache/conftool/dbconfig/20260805-072426-ladsgroup.json * 07:21 slyngshede@dns1004: END - running authdns-update * 07:19 slyngshede@dns1004: START - running authdns-update * 07:13 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231', diff saved to https://phabricator.wikimedia.org/P95906 and previous config saved to /var/cache/conftool/dbconfig/20260805-071340-ladsgroup.json * 07:02 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231', diff saved to https://phabricator.wikimedia.org/P95905 and previous config saved to /var/cache/conftool/dbconfig/20260805-070253-ladsgroup.json * 06:52 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95904 and previous config saved to /var/cache/conftool/dbconfig/20260805-065206-ladsgroup.json * 06:45 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 06:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95903 and previous config saved to /var/cache/conftool/dbconfig/20260805-062240-ladsgroup.json * 06:21 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2231.codfw.wmnet with reason: Maintenance * 06:21 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95902 and previous config saved to /var/cache/conftool/dbconfig/20260805-062137-ladsgroup.json * 06:10 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215', diff saved to https://phabricator.wikimedia.org/P95901 and previous config saved to /var/cache/conftool/dbconfig/20260805-061051-ladsgroup.json * 06:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215', diff saved to https://phabricator.wikimedia.org/P95900 and previous config saved to /var/cache/conftool/dbconfig/20260805-060004-ladsgroup.json * 05:49 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95899 and previous config saved to /var/cache/conftool/dbconfig/20260805-054918-ladsgroup.json * 05:19 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95898 and previous config saved to /var/cache/conftool/dbconfig/20260805-051939-ladsgroup.json * 05:18 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2215.codfw.wmnet with reason: Maintenance * 04:30 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2201.codfw.wmnet with reason: Maintenance * 03:40 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2197.codfw.wmnet with reason: Maintenance * 03:40 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95897 and previous config saved to /var/cache/conftool/dbconfig/20260805-034036-ladsgroup.json * 03:29 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196', diff saved to https://phabricator.wikimedia.org/P95896 and previous config saved to /var/cache/conftool/dbconfig/20260805-032948-ladsgroup.json * 03:19 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196', diff saved to https://phabricator.wikimedia.org/P95895 and previous config saved to /var/cache/conftool/dbconfig/20260805-031902-ladsgroup.json * 03:08 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95894 and previous config saved to /var/cache/conftool/dbconfig/20260805-030815-ladsgroup.json * 02:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95893 and previous config saved to /var/cache/conftool/dbconfig/20260805-023413-ladsgroup.json * 02:33 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2196.codfw.wmnet with reason: Maintenance * 02:33 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95892 and previous config saved to /var/cache/conftool/dbconfig/20260805-023310-ladsgroup.json * 02:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186', diff saved to https://phabricator.wikimedia.org/P95891 and previous config saved to /var/cache/conftool/dbconfig/20260805-022223-ladsgroup.json * 02:11 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186', diff saved to https://phabricator.wikimedia.org/P95890 and previous config saved to /var/cache/conftool/dbconfig/20260805-021137-ladsgroup.json * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 02:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95889 and previous config saved to /var/cache/conftool/dbconfig/20260805-020051-ladsgroup.json * 01:30 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95888 and previous config saved to /var/cache/conftool/dbconfig/20260805-013029-ladsgroup.json * 01:29 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2186.codfw.wmnet with reason: Maintenance * 00:34 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on dbstore1009.eqiad.wmnet with reason: Maintenance * 00:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95887 and previous config saved to /var/cache/conftool/dbconfig/20260805-003408-ladsgroup.json * 00:23 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264', diff saved to https://phabricator.wikimedia.org/P95886 and previous config saved to /var/cache/conftool/dbconfig/20260805-002322-ladsgroup.json * 00:12 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264', diff saved to https://phabricator.wikimedia.org/P95885 and previous config saved to /var/cache/conftool/dbconfig/20260805-001235-ladsgroup.json * 00:01 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95884 and previous config saved to /var/cache/conftool/dbconfig/20260805-000148-ladsgroup.json == 2026-08-04 == * 23:45 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95883 and previous config saved to /var/cache/conftool/dbconfig/20260804-234508-ladsgroup.json * 23:44 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1264.eqiad.wmnet with reason: Maintenance * 23:44 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95882 and previous config saved to /var/cache/conftool/dbconfig/20260804-234405-ladsgroup.json * 23:33 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237', diff saved to https://phabricator.wikimedia.org/P95881 and previous config saved to /var/cache/conftool/dbconfig/20260804-233317-ladsgroup.json * 23:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237', diff saved to https://phabricator.wikimedia.org/P95880 and previous config saved to /var/cache/conftool/dbconfig/20260804-232230-ladsgroup.json * 23:11 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95879 and previous config saved to /var/cache/conftool/dbconfig/20260804-231144-ladsgroup.json * 22:23 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95878 and previous config saved to /var/cache/conftool/dbconfig/20260804-222345-ladsgroup.json * 22:23 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1237.eqiad.wmnet with reason: Maintenance * 21:13 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1225.eqiad.wmnet with reason: Maintenance * 20:40 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] (duration: 24m 40s) * 20:33 samtar@deploy1003: samtar, kineticpelagic: Continuing with deployment * 20:28 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS bookworm * 20:21 samtar@deploy1003: samtar, kineticpelagic: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:15 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] * 20:13 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 20:09 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 20:00 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1216.eqiad.wmnet with reason: Maintenance * 20:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95877 and previous config saved to /var/cache/conftool/dbconfig/20260804-195957-ladsgroup.json * 19:51 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 19:50 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) mc2046.codfw.wmnet 120.16.192.10.in-addr.arpa 0.2.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:50 jhancock@cumin2002: START - Cookbook sre.dns.wipe-cache mc2046.codfw.wmnet 120.16.192.10.in-addr.arpa 0.2.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host mc2046 - jhancock@cumin2002" * 19:50 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host mc2046 - jhancock@cumin2002" * 19:49 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203', diff saved to https://phabricator.wikimedia.org/P95876 and previous config saved to /var/cache/conftool/dbconfig/20260804-194911-ladsgroup.json * 19:46 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 19:45 jhancock@cumin2002: START - Cookbook sre.hosts.move-vlan for host mc2046 * 19:45 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS bookworm * 19:38 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203', diff saved to https://phabricator.wikimedia.org/P95875 and previous config saved to /var/cache/conftool/dbconfig/20260804-193825-ladsgroup.json * 19:27 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95874 and previous config saved to /var/cache/conftool/dbconfig/20260804-192738-ladsgroup.json * 19:02 mutante: gerrit ssh -p 29418 gerrit.wikimedia.org gerrit index changes {{Gerrit|1320979}} * 18:20 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 18:18 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 18:14 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 18:14 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 18:13 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 18:10 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 18:08 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 18:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95872 and previous config saved to /var/cache/conftool/dbconfig/20260804-180721-ladsgroup.json * 18:07 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 18:06 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1203.eqiad.wmnet with reason: Maintenance * 18:06 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95871 and previous config saved to /var/cache/conftool/dbconfig/20260804-180618-ladsgroup.json * 17:55 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179', diff saved to https://phabricator.wikimedia.org/P95870 and previous config saved to /var/cache/conftool/dbconfig/20260804-175531-ladsgroup.json * 17:55 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1154.eqiad.wmnet * 17:55 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1154.eqiad.wmnet * 17:55 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1154.eqiad.wmnet * 17:50 swfrench@deploy1003: Finished scap sync-world: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] (duration: 04m 05s) * 17:48 swfrench@deploy1003: swfrench: Continuing with deployment * 17:46 swfrench@deploy1003: swfrench: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:45 swfrench@deploy1003: Started scap sync-world: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] * 17:44 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179', diff saved to https://phabricator.wikimedia.org/P95869 and previous config saved to /var/cache/conftool/dbconfig/20260804-174445-ladsgroup.json * 17:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95868 and previous config saved to /var/cache/conftool/dbconfig/20260804-173359-ladsgroup.json * 17:33 swfrench@deploy1003: Finished scap sync-world: Pick up new PHP production image (duration: 28m 32s) * 17:28 aokoth@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on phab1005.eqiad.wmnet with reason: Puppet Failure * 17:05 swfrench@deploy1003: Started scap sync-world: Pick up new PHP production image * 17:00 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 17:00 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 16:54 cgoubert@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on wikikube-worker2187.codfw.wmnet with reason: Hardware issue * 16:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2187.codfw.wmnet * 16:52 mutante: gerrit2003:/var/log/apache2# ln -s /srv/gerrit/site_path/review_site/logs/ gerrit * 16:52 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2187.codfw.wmnet * 16:48 mutante: gerrit2003 - moving old apache logfiles older than 60 days from /var/log/apache2 to /srv/gerrit/site_path/review_site/logs/old/ * 16:33 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 16:32 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 16:29 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 16:29 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 16:28 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 16:28 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 16:27 dzahn@cumin1003: END (PASS) - Cookbook sre.gerrit.restart-gerrit (exit_code=0) Restarting Gerrit on gerrit2003 * 16:27 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 16:27 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95867 and previous config saved to /var/cache/conftool/dbconfig/20260804-162736-ladsgroup.json * 16:27 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 16:26 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1179.eqiad.wmnet with reason: Maintenance * 16:26 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:25 mutante: restarting gerrit - dropped outdated RSA host key * 16:25 dzahn@cumin1003: START - Cookbook sre.gerrit.restart-gerrit Restarting Gerrit on gerrit2003 * 16:24 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95866 and previous config saved to /var/cache/conftool/dbconfig/20260804-162424-ladsgroup.json * 16:24 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 16:23 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 16:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95865 and previous config saved to /var/cache/conftool/dbconfig/20260804-162236-ladsgroup.json * 16:21 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1179.eqiad.wmnet with reason: Maintenance * 16:17 swfrench-wmf: reprepro include php8.3_8.3.33-1+wmf11u1 into component/php83 for bullseye-wikimedia * 16:17 swfrench-wmf: reprepro include php8.3_8.3.33-1+wmf12u1 into component/php83 for bookworm-wikimedia * 16:11 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply * 16:10 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply * 16:10 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mobileapps: apply * 16:09 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mobileapps: apply * 16:09 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply * 16:08 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply * 16:08 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:08 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:07 aokoth@cumin1003: END (PASS) - Cookbook sre.vrts.upgrade (exit_code=0) on VRTS host vrts1003.eqiad.wmnet * 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:05 aokoth@cumin1003: START - Cookbook sre.vrts.upgrade on VRTS host vrts1003.eqiad.wmnet * 16:04 mutante: gerrit2002/gerrit1003/gerrit2003 - rm /etc/gerrit/ssh_host_rsa_key * 15:59 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:59 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:59 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:59 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:56 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 15:55 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:55 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:55 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:49 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 15:49 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:44 Raine: add php8.5 packages to component/php85 - [[phab:T432983|T432983]] * 15:39 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:33 aaron@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 15:33 aaron@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 15:29 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:19 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:19 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:16 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:16 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1154.eqiad.wmnet with OS trixie * 15:16 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:15 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:15 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:06 brennen@deploy1003: Finished deploy [phabricator/deployment@56f4ffd]: deploy phab1004 for [[phab:T433981|T433981]] (duration: 00m 43s) * 15:05 brennen@deploy1003: Started deploy [phabricator/deployment@56f4ffd]: deploy phab1004 for [[phab:T433981|T433981]] * 15:05 aaron@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 15:04 aaron@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 15:02 brennen@deploy1003: Finished deploy [phabricator/deployment@56f4ffd]: deploy phab2003 for [[phab:T433981|T433981]] (duration: 00m 51s) * 15:01 brennen@deploy1003: Started deploy [phabricator/deployment@56f4ffd]: deploy phab2003 for [[phab:T433981|T433981]] * 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1004.eqiad.wmnet with reason: deployment * 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1005.eqiad.wmnet with reason: deployment * 14:58 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab2003.codfw.wmnet with reason: deployment * 14:55 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1154.eqiad.wmnet with reason: host reimage * 14:51 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1154.eqiad.wmnet with reason: host reimage * 14:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 14:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 14:38 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync * 14:38 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync * 14:38 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync * 14:37 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync * 14:37 ottomata: roll restart eventgate-main to pick up stream config change - [[phab:T433507|T433507]] * 14:37 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-main: sync * 14:36 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-main: sync * 14:36 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1154 * 14:36 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1154 * 14:34 otto@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] (duration: 08m 39s) * 14:34 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1154 * 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1154.eqiad.wmnet 108.32.64.10.in-addr.arpa 8.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:34 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1154.eqiad.wmnet 108.32.64.10.in-addr.arpa 8.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1154 - jayme@cumin1003" * 14:34 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1154 - jayme@cumin1003" * 14:30 otto@deploy1003: otto: Continuing with deployment * 14:30 jayme@cumin1003: START - Cookbook sre.dns.netbox * 14:28 otto@deploy1003: otto: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:26 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1154 * 14:26 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1154.eqiad.wmnet with OS trixie * 14:26 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1154.eqiad.wmnet * 14:26 otto@deploy1003: Started scap sync-world: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] * 14:26 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1154.eqiad.wmnet * 14:26 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1154.eqiad.wmnet * 14:17 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 14:16 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 14:15 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 14:14 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 14:13 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 14:13 swfrench@dns1004: END - running authdns-update * 14:13 Msz2001: Finished deployments for UTC afternoon backport window * 14:13 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 14:13 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] (duration: 07m 58s) * 14:11 swfrench@dns1004: START - running authdns-update * 14:08 mszwarc@deploy1003: javiermonton, mszwarc, mpostoronca: Continuing with deployment * 14:07 mszwarc@deploy1003: javiermonton, mszwarc, mpostoronca: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] synced to the testser * 14:05 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] * 14:03 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 13:49 swfrench@cumin2002: conftool action : set/pooled=yes; selector: name=wikikube-worker2330.codfw.wmnet * 13:49 swfrench@cumin2002: conftool action : set/pooled=no; selector: name=wikikube-worker2330.codfw.wmnet * 13:48 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] (duration: 09m 19s) * 13:45 swfrench@dns1004: END - running authdns-update * 13:44 mszwarc@deploy1003: mszwarc, jforrester: Continuing with deployment * 13:43 swfrench@dns1004: START - running authdns-update * 13:41 mszwarc@deploy1003: mszwarc, jforrester: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:38 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] * 13:33 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 13:33 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1154.eqiad.wmnet * 13:32 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 13:32 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 13:31 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 13:31 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:31 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:29 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1154.eqiad.wmnet * 13:28 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1154.eqiad.wmnet * 13:28 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1154.eqiad.wmnet * 13:28 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1141.eqiad.wmnet * 13:28 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1141.eqiad.wmnet * 13:28 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1141.eqiad.wmnet * 13:22 otto@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply * 13:22 otto@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply * 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1096.eqiad.wmnet with OS trixie * 13:05 swfrench@dns1004: END - running authdns-update * 13:03 swfrench@dns1004: START - running authdns-update * 12:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 12:43 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 1:00:00 on db1171.eqiad.wmnet with reason: decom * 12:42 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 1:00:00 on db1150.eqiad.wmnet with reason: decom * 12:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 12:38 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1164,1217].eqiad.wmnet with reason: cloning * 12:33 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2096.codfw.wmnet with OS trixie * 12:22 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1096.eqiad.wmnet with OS trixie * 12:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2096.codfw.wmnet with reason: host reimage * 12:14 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1141.eqiad.wmnet with OS trixie * 12:10 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2096.codfw.wmnet with reason: host reimage * 12:10 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1289.eqiad.wmnet * 12:05 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1289.eqiad.wmnet * 12:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1288.eqiad.wmnet * 11:59 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1288.eqiad.wmnet * 11:59 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1287.eqiad.wmnet * 11:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1097.eqiad.wmnet with OS trixie * 11:54 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1287.eqiad.wmnet * 11:54 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1286.eqiad.wmnet * 11:53 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1141.eqiad.wmnet with reason: host reimage * 11:51 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2096.codfw.wmnet with OS trixie * 11:49 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1141.eqiad.wmnet with reason: host reimage * 11:48 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1286.eqiad.wmnet * 11:48 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1284.eqiad.wmnet * 11:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2095.codfw.wmnet with OS trixie * 11:43 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1284.eqiad.wmnet * 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1283.eqiad.wmnet * 11:42 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on ml-serve1015.eqiad.wmnet with reason: Downtime to get full picture of current BIOS settings beyond what Redfish shows * 11:39 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad * 11:39 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:37 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1283.eqiad.wmnet * 11:37 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1282.eqiad.wmnet * 11:37 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad * 11:37 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:33 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1141 * 11:33 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1141 * 11:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 11:32 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1141 * 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1141.eqiad.wmnet 156.48.64.10.in-addr.arpa 6.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:32 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1141.eqiad.wmnet 156.48.64.10.in-addr.arpa 6.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1141 - jayme@cumin1003" * 11:32 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1141 - jayme@cumin1003" * 11:32 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1282.eqiad.wmnet * 11:32 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1281.eqiad.wmnet * 11:32 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad * 11:32 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:29 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 11:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1097.eqiad.wmnet with reason: host reimage * 11:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2095.codfw.wmnet with OS trixie * 11:26 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1281.eqiad.wmnet * 11:26 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1280.eqiad.wmnet * 11:25 jayme@cumin1003: START - Cookbook sre.dns.netbox * 11:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1097.eqiad.wmnet with reason: host reimage * 11:22 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1141 * 11:21 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1141.eqiad.wmnet with OS trixie * 11:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1280.eqiad.wmnet * 11:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1279.eqiad.wmnet * 11:20 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin with reason: upgrade new Nokia swtiches in eqsin to SR Linux v26 * 11:17 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1141.eqiad.wmnet * 11:16 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1141.eqiad.wmnet * 11:16 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1141.eqiad.wmnet * 11:16 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be2095.codfw.wmnet with OS trixie * 11:15 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1279.eqiad.wmnet * 11:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1278.eqiad.wmnet * 11:14 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1139.eqiad.wmnet * 11:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1139.eqiad.wmnet * 11:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1070.eqiad.wmnet with OS trixie * 11:13 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1140.eqiad.wmnet * 11:13 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1140.eqiad.wmnet * 11:13 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1140.eqiad.wmnet * 11:09 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1278.eqiad.wmnet * 11:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1071.eqiad.wmnet with OS trixie * 11:05 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1097.eqiad.wmnet with OS trixie * 11:04 mvernon@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be1097.eqiad.wmnet with OS trixie * 11:02 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1140.eqiad.wmnet with OS trixie * 11:02 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1097.eqiad.wmnet with OS trixie * 11:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1096.eqiad.wmnet with OS trixie * 11:00 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1139.eqiad.wmnet * 11:00 jayme@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1139.eqiad.wmnet with OS trixie * 10:56 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1069.eqiad.wmnet with OS trixie * 10:56 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 10:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1070.eqiad.wmnet with reason: host reimage * 10:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 10:45 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1071.eqiad.wmnet with reason: host reimage * 10:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 10:41 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1096.eqiad.wmnet with OS trixie * 10:41 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1140.eqiad.wmnet with reason: host reimage * 10:39 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1071.eqiad.wmnet with reason: host reimage * 10:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1070.eqiad.wmnet with reason: host reimage * 10:38 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1139.eqiad.wmnet with reason: host reimage * 10:37 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1140.eqiad.wmnet with reason: host reimage * 10:35 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1069.eqiad.wmnet with reason: host reimage * 10:33 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1095.eqiad.wmnet with OS trixie * 10:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 10:33 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1139.eqiad.wmnet with reason: host reimage * 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1069.eqiad.wmnet with reason: host reimage * 10:23 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1140 * 10:23 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1140 * 10:23 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 10:22 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1140 * 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1140.eqiad.wmnet 155.48.64.10.in-addr.arpa 5.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:21 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1071.eqiad.wmnet with OS trixie * 10:21 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1140.eqiad.wmnet 155.48.64.10.in-addr.arpa 5.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1140 - jayme@cumin1003" * 10:21 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1140 - jayme@cumin1003" * 10:21 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1071 * 10:21 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1070.eqiad.wmnet with OS trixie * 10:21 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1070 * 10:20 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 10:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1095.eqiad.wmnet with OS trixie * 10:17 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1139 * 10:17 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1139 * 10:17 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be1095.eqiad.wmnet with OS trixie * 10:15 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1139 * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1139.eqiad.wmnet 194.32.64.10.in-addr.arpa 4.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:15 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1139.eqiad.wmnet 194.32.64.10.in-addr.arpa 4.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1139 - jayme@cumin1003" * 10:15 jayme@cumin1003: START - Cookbook sre.dns.netbox * 10:15 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1139 - jayme@cumin1003" * 10:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2095.codfw.wmnet with OS trixie * 10:13 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1069.eqiad.wmnet with OS trixie * 10:12 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1069 * 10:11 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1140 * 10:11 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1140.eqiad.wmnet with OS trixie * 10:11 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1140.eqiad.wmnet * 10:10 jayme@cumin1003: START - Cookbook sre.dns.netbox * 10:10 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1139 * 10:10 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1140.eqiad.wmnet * 10:10 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1140.eqiad.wmnet * 10:10 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1139.eqiad.wmnet with OS trixie * 10:09 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1139.eqiad.wmnet * 10:08 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1139.eqiad.wmnet * 10:08 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1139.eqiad.wmnet * 10:01 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2094.codfw.wmnet with OS trixie * 09:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 09:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 09:53 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 09:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 09:44 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:44 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2094.codfw.wmnet with reason: host reimage * 09:34 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2094.codfw.wmnet with reason: host reimage * 09:34 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1071 * 09:33 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1070 * 09:33 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1095.eqiad.wmnet with OS trixie * 09:27 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1069 * 09:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:23 brouberol@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM archiva1002.wikimedia.org * 09:20 brouberol@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM archiva1002.wikimedia.org * 09:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1277.eqiad.wmnet * 09:13 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2094.codfw.wmnet with OS trixie * 09:13 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 09:12 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 09:12 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:12 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1277.eqiad.wmnet * 09:12 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1276.eqiad.wmnet * 09:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1094.eqiad.wmnet with OS trixie * 09:06 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1276.eqiad.wmnet * 09:06 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1275.eqiad.wmnet * 09:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2093.codfw.wmnet with OS trixie * 09:01 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1275.eqiad.wmnet * 09:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1274.eqiad.wmnet * 08:56 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1274.eqiad.wmnet * 08:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1273.eqiad.wmnet * 08:50 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1273.eqiad.wmnet * 08:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1272.eqiad.wmnet * 08:49 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:49 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1094.eqiad.wmnet with reason: host reimage * 08:45 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1272.eqiad.wmnet * 08:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1094.eqiad.wmnet with reason: host reimage * 08:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2093.codfw.wmnet with reason: host reimage * 08:38 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:38 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2093.codfw.wmnet with reason: host reimage * 08:35 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 08:34 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 08:29 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:28 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:26 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1271.eqiad.wmnet * 08:23 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1094.eqiad.wmnet with OS trixie * 08:21 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 08:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1271.eqiad.wmnet * 08:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1270.eqiad.wmnet * 08:15 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1270.eqiad.wmnet * 08:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1269.eqiad.wmnet * 08:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2093.codfw.wmnet with OS trixie * 08:09 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1269.eqiad.wmnet * 08:09 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1268.eqiad.wmnet * 08:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2092.codfw.wmnet with OS trixie * 08:04 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1268.eqiad.wmnet * 08:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1267.eqiad.wmnet * 07:59 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1267.eqiad.wmnet * 07:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1093.eqiad.wmnet with OS trixie * 07:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1266.eqiad.wmnet * 07:56 jynus: running extra backups to test db1285 [[phab:T433826|T433826]] * 07:51 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1266.eqiad.wmnet * 07:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2092.codfw.wmnet with reason: host reimage * 07:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1093.eqiad.wmnet with reason: host reimage * 07:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2092.codfw.wmnet with reason: host reimage * 07:32 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1093.eqiad.wmnet with reason: host reimage * 07:29 jynus: running extra backups to test db1265 [[phab:T433825|T433825]] * 07:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2092.codfw.wmnet with OS trixie * 07:11 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1093.eqiad.wmnet with OS trixie * 06:50 slyngshede@dns1004: END - running authdns-update * 06:48 slyngshede@dns1004: START - running authdns-update * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.11 (duration: 02m 29s) * 03:38 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] (duration: 32m 57s) * 03:23 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 03:22 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 03:05 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 32s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 00:45 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] (duration: 06m 20s) * 00:41 cjming@deploy1003: cjming: Continuing with deployment * 00:41 cjming@deploy1003: cjming: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:39 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] == 2026-08-03 == * 23:58 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply * 23:57 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply * 23:29 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cp5021.eqsin.wmnet * 23:29 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cp5021.eqsin.wmnet * 23:27 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cp5021.eqsin.wmnet * 23:26 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cp5021.eqsin.wmnet * 23:18 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 23:17 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 22:56 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: sync * 22:56 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: sync * 22:36 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 22:36 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 22:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host search-loader2002.codfw.wmnet with OS trixie * 21:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on search-loader2002.codfw.wmnet with reason: host reimage * 21:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on search-loader2002.codfw.wmnet with reason: host reimage * 21:42 dancy@deploy1003: Stopping before sync operations * 21:41 dancy@deploy1003: Started scap sync-world: testing * 21:39 dancy@deploy1003: Installation of scap version "4.277.0" completed for 3 hosts * 21:37 dancy@deploy1003: Installing scap version "4.277.0" for 3 host(s) * 21:37 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] (duration: 06m 13s) * 21:33 dancy@deploy1003: dancy: Continuing with deployment * 21:32 dancy@deploy1003: dancy: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:31 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] * 21:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host search-loader2002.codfw.wmnet with OS trixie * 21:03 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] (duration: 06m 34s) * 20:59 dancy@deploy1003: dancy: Continuing with deployment * 20:58 dancy@deploy1003: dancy: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:56 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] * 20:52 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] (duration: 06m 23s) * 20:48 cjming@deploy1003: cjming: Continuing with deployment * 20:47 cjming@deploy1003: cjming: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:46 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] * 20:42 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] (duration: 07m 36s) * 20:38 arlolra@deploy1003: arlolra: Continuing with deployment * 20:36 arlolra@deploy1003: arlolra: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:34 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] * 20:16 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] (duration: 08m 26s) * 20:12 krinkle@deploy1003: krinkle: Continuing with deployment * 20:09 krinkle@deploy1003: krinkle: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] * 19:45 jasmine@cumin2002: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-main-codfw * 18:58 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] (duration: 09m 23s) * 18:53 krinkle@deploy1003: krinkle: Continuing with deployment * 18:53 jasmine@cumin2002: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-main-codfw * 18:50 krinkle@deploy1003: krinkle: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:48 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] * 18:37 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] (duration: 10m 13s) * 18:34 dzahn@cumin2002: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host codesearch2001.codfw.wmnet * 18:34 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host codesearch2001.codfw.wmnet with OS trixie * 18:33 krinkle@deploy1003: krinkle: Continuing with deployment * 18:29 krinkle@deploy1003: krinkle: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:27 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] * 18:19 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 18:18 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on codesearch2001.codfw.wmnet with reason: host reimage * 18:14 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 18:14 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:12 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on codesearch2001.codfw.wmnet with reason: host reimage * 18:11 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 18:11 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 18:10 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 18:02 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 18:02 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 18:01 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 18:01 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 17:55 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host codesearch2001.codfw.wmnet with OS trixie * 17:54 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:54 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) codesearch2001.codfw.wmnet on all recursors * 17:53 dzahn@cumin2002: START - Cookbook sre.dns.wipe-cache codesearch2001.codfw.wmnet on all recursors * 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:48 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:41 dzahn@cumin2002: START - Cookbook sre.dns.netbox * 17:41 dzahn@cumin2002: START - Cookbook sre.ganeti.makevm for new host codesearch2001.codfw.wmnet * 17:37 dzahn@cumin2002: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host codesearch1001.eqiad.wmnet * 17:37 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host codesearch1001.eqiad.wmnet with OS trixie * 17:24 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on codesearch1001.eqiad.wmnet with reason: host reimage * 17:17 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on codesearch1001.eqiad.wmnet with reason: host reimage * 17:08 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host codesearch1001.eqiad.wmnet with OS trixie * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 17:06 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:06 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) codesearch1001.eqiad.wmnet on all recursors * 17:06 dzahn@cumin2002: START - Cookbook sre.dns.wipe-cache codesearch1001.eqiad.wmnet on all recursors * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 17:05 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 17:04 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:04 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 16:58 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 16:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2091.codfw.wmnet with OS trixie * 16:54 ebernhardson@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 16:54 ebernhardson@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 16:49 ebernhardson@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 16:49 ebernhardson@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 16:46 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1092.eqiad.wmnet with OS trixie * 16:43 ebernhardson@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 16:43 ebernhardson@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 16:43 dzahn@cumin2002: START - Cookbook sre.dns.netbox * 16:43 dzahn@cumin2002: START - Cookbook sre.ganeti.makevm for new host codesearch1001.eqiad.wmnet * 16:41 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 16:41 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 16:40 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2091.codfw.wmnet with reason: host reimage * 16:37 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 16:35 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2091.codfw.wmnet with reason: host reimage * 16:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1092.eqiad.wmnet with reason: host reimage * 16:24 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 16:24 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 16:23 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1092.eqiad.wmnet with reason: host reimage * 16:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2091.codfw.wmnet with OS trixie * 16:03 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1092.eqiad.wmnet with OS trixie * 16:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2090.codfw.wmnet with OS trixie * 15:51 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 15:51 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 15:51 jiji@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 15:50 jiji@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 15:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2090.codfw.wmnet with reason: host reimage * 15:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2090.codfw.wmnet with reason: host reimage * 15:33 jhathaway@dns1004: END - running authdns-update * 15:31 jhathaway@dns1004: START - running authdns-update * 15:26 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1091.eqiad.wmnet with OS trixie * 15:25 dancy@deploy1003: Installation of scap version "4.276.1" completed for 3 hosts * 15:23 dancy@deploy1003: Installing scap version "4.276.1" for 3 host(s) * 15:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2090.codfw.wmnet with OS trixie * 15:12 marostegui@cumin1003: dbctl commit (dc=all): 'Repool db2245, db2246, db2247 and db2248 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95857 and previous config saved to /var/cache/conftool/dbconfig/20260803-151212-marostegui.json * 15:09 dancy@deploy1003: Started scap sync-world: testing * 15:09 dancy@deploy1003: Installation of scap version "4.277.0" completed for 3 hosts * 15:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1091.eqiad.wmnet with reason: host reimage * 15:07 dancy@deploy1003: Installing scap version "4.277.0" for 3 host(s) * 15:03 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1091.eqiad.wmnet with reason: host reimage * 14:49 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1091.eqiad.wmnet with OS trixie * 14:33 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2089.codfw.wmnet with OS trixie * 14:29 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 14:27 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 14:18 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 14:16 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 14:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2089.codfw.wmnet with reason: host reimage * 14:10 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 14:10 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 14:09 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2089.codfw.wmnet with reason: host reimage * 13:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2089.codfw.wmnet with OS trixie * 13:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2088.codfw.wmnet with OS trixie * 13:40 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1090.eqiad.wmnet with OS trixie * 13:22 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1090.eqiad.wmnet with reason: host reimage * 13:22 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] (duration: 14m 34s) * 13:19 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1090.eqiad.wmnet with reason: host reimage * 13:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2088.codfw.wmnet with reason: host reimage * 13:16 aude@deploy1003: aude, mhorsey: Continuing with deployment * 13:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2088.codfw.wmnet with reason: host reimage * 13:12 aude@deploy1003: aude, mhorsey: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] * 13:05 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1090.eqiad.wmnet with OS trixie * 12:58 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2088.codfw.wmnet with OS trixie * 12:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db[2245-2247].codfw.wmnet * 12:49 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2247: Rebooting db2247.codfw.wmnet * 12:49 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2247: Rebooting db2247.codfw.wmnet * 12:42 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2246: Rebooting db2246.codfw.wmnet * 12:42 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2246: Rebooting db2246.codfw.wmnet * 12:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2087.codfw.wmnet with OS trixie * 12:37 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1089.eqiad.wmnet with OS trixie * 12:34 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2245: Rebooting db2245.codfw.wmnet * 12:34 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2245: Rebooting db2245.codfw.wmnet * 12:34 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db[2245-2247].codfw.wmnet * 12:32 kamila@deploy1003: Finished scap sync-world: rebuild after base image update (duration: 30m 26s) * 12:28 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 12:22 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2087.codfw.wmnet with reason: host reimage * 12:19 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1089.eqiad.wmnet with reason: host reimage * 12:14 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2087.codfw.wmnet with reason: host reimage * 12:14 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1089.eqiad.wmnet with reason: host reimage * 12:03 kamila@deploy1003: Started scap sync-world: rebuild after base image update * 12:00 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1089.eqiad.wmnet with OS trixie * 12:00 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2087.codfw.wmnet with OS trixie * 11:35 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db[2245-2248].codfw.wmnet * 11:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db[2245-2248].codfw.wmnet * 11:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2086.codfw.wmnet with OS trixie * 11:26 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 11:26 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 11:25 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 11:25 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 11:24 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1088.eqiad.wmnet with OS trixie * 11:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db[2245-2248].codfw.wmnet with reason: Checking network * 11:21 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 11:20 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 11:19 marostegui@dns1004: END - running authdns-update * 11:17 marostegui@dns1004: START - running authdns-update * 11:10 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:10 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 11:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2086.codfw.wmnet with reason: host reimage * 11:09 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:08 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 11:08 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:07 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 11:07 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:07 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 11:06 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop: apply * 11:06 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop: apply * 11:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1088.eqiad.wmnet with reason: host reimage * 11:05 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop: apply * 11:04 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop: apply * 11:04 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop: apply * 11:04 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop: apply * 11:02 marostegui@dns1004: END - running authdns-update * 11:02 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2086.codfw.wmnet with reason: host reimage * 11:01 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1088.eqiad.wmnet with reason: host reimage * 11:00 marostegui@dns1004: START - running authdns-update * 10:53 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] (duration: 10m 57s) * 10:51 cmooney@dns3003: END - running authdns-update * 10:49 cmooney@dns3003: START - running authdns-update * 10:47 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1088.eqiad.wmnet with OS trixie * 10:47 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2086.codfw.wmnet with OS trixie * 10:47 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 10:46 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:46 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:46 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new reverse ranges for eqsin CR switch links - cmooney@cumin1003" * 10:46 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new reverse ranges for eqsin CR switch links - cmooney@cumin1003" * 10:42 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] * 10:41 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 10:36 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2245, db2246 and db2247 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95855 and previous config saved to /var/cache/conftool/dbconfig/20260803-103652-marostegui.json * 10:35 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2248 from s4 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95854 and previous config saved to /var/cache/conftool/dbconfig/20260803-103535-marostegui.json * 10:27 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 10:27 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 10:26 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 10:24 kart_: cxserver: Add referencePunctuation config ([[phab:T97231|T97231]]) * 10:24 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 10:23 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:23 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:23 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:22 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:22 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply * 10:21 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply * 10:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2085.codfw.wmnet with OS trixie * 10:20 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply * 10:20 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply * 10:18 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply * 10:18 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply * 10:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1087.eqiad.wmnet with OS trixie * 09:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2085.codfw.wmnet with reason: host reimage * 09:43 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1087.eqiad.wmnet with reason: host reimage * 09:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2085.codfw.wmnet with reason: host reimage * 09:40 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1087.eqiad.wmnet with reason: host reimage * 09:26 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1087.eqiad.wmnet with OS trixie * 09:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2085.codfw.wmnet with OS trixie * 09:13 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2084.codfw.wmnet with OS trixie * 09:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1086.eqiad.wmnet with OS trixie * 08:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2084.codfw.wmnet with reason: host reimage * 08:50 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2084.codfw.wmnet with reason: host reimage * 08:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1086.eqiad.wmnet with reason: host reimage * 08:39 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1086.eqiad.wmnet with reason: host reimage * 08:38 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:38 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:37 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 08:37 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 08:35 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2084.codfw.wmnet with OS trixie * 08:34 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 08:34 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:27 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1086.eqiad.wmnet with OS trixie * 08:09 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1218: Repool after a crash * 08:07 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2083.codfw.wmnet with OS trixie * 08:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1085.eqiad.wmnet with OS trixie * 07:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2083.codfw.wmnet with reason: host reimage * 07:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1085.eqiad.wmnet with reason: host reimage * 07:40 kart_: Updated cxsever to 2026-07-16-140518-production ([[phab:T97231|T97231]]) * 07:39 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply * 07:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2083.codfw.wmnet with reason: host reimage * 07:38 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply * 07:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1085.eqiad.wmnet with reason: host reimage * 07:37 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] (duration: 32m 40s) * 07:33 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply * 07:33 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply * 07:25 jdlrobson@deploy1003: jdlrobson: Continuing with deployment * 07:24 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2083.codfw.wmnet with OS trixie * 07:24 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1085.eqiad.wmnet with OS trixie * 07:23 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1218: Repool after a crash * 07:21 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:09 marostegui: Drop renamed tables [[phab:T425074|T425074]] * 07:04 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] * 06:55 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply * 06:54 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 46s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-02 == * 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 01m 03s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-01 == * 03:30 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:30 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:30 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:30 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 34s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-31 == * 17:41 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 17:41 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 17:40 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 17:40 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 15:33 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2195: Testing * 15:02 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:02 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 15:02 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 14:48 pt1979@cumin2002: START - Cookbook sre.dns.netbox * 14:47 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2195: Testing * 14:22 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 14:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2195: Testing * 14:21 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 14:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2195.codfw.wmnet with reason: Testing * 14:16 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 14:04 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 14:04 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 13:30 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1048.eqiad.wmnet with OS trixie * 13:22 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2195: Testing * 13:22 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 13:19 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2195: Testing * 13:18 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 13:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2195: Testing * 13:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 13:05 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 13:05 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 13:04 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 13:04 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 12:50 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lswtest-d8-eqiad * 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:53 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:42 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 6515 * 11:37 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 6515 * 11:28 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:27 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:07 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2082.codfw.wmnet with OS trixie * 10:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2082.codfw.wmnet with reason: host reimage * 10:42 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2082.codfw.wmnet with reason: host reimage * 10:28 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2082.codfw.wmnet with OS trixie * 10:02 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 09:52 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 09:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts * 09:16 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts * 08:57 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 08:46 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:42 gkyziridis@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 08:37 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 08:37 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 08:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 08:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 08:11 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:11 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:08 filippo@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudvirt1048 * 08:07 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 08:07 filippo@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudvirt1048 * 08:06 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 08:01 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 08:00 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 07:19 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1048.eqiad.wmnet with reason: host reimage * 07:13 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1048.eqiad.wmnet with reason: host reimage * 07:11 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 07:11 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 07:09 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 07:09 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 06:57 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:56 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:48 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:48 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:44 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie * 06:34 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1048.eqiad.wmnet with OS trixie * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 54s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 00:57 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] (duration: 11m 04s) * 00:53 dreamyjazz@deploy1003: dreamyjazz, jforrester: Continuing with deployment * 00:48 dreamyjazz@deploy1003: dreamyjazz, jforrester: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:46 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] == 2026-07-30 == * 21:37 dancy@deploy1003: Installation of scap version "4.276.1" completed for 3 hosts * 21:35 dancy@deploy1003: Installing scap version "4.276.1" for 3 host(s) * 21:24 dancy@deploy1003: Installation of scap version "4.276.0" completed for 3 hosts * 21:22 dancy@deploy1003: Installing scap version "4.276.0" for 3 host(s) * 21:15 maryum: Deployed security fix for [[phab:T430601|T430601]] * 20:13 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] (duration: 09m 20s) * 20:07 arlolra@deploy1003: osleger, arlolra: Continuing with deployment * 20:05 arlolra@deploy1003: osleger, arlolra: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:03 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] * 19:29 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:29 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:25 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service * 19:24 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 19:24 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:24 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:24 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 19:23 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service * 19:20 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1084.eqiad.wmnet with OS trixie * 18:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1084.eqiad.wmnet with reason: host reimage * 18:52 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1084.eqiad.wmnet with reason: host reimage * 18:41 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 18:40 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 18:39 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1084.eqiad.wmnet with OS trixie * 18:25 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 18:15 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 17:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1083.eqiad.wmnet with OS trixie * 17:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1048: Maintenance * 17:36 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new security plugin settings - bking@cumin2003 - [[phab:T350516|T350516]] * 17:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1083.eqiad.wmnet with reason: host reimage * 17:26 inflatador: bking@apt1002 `reprepro --noskipold --component thirdparty/opensearch3 update trixie-wikimedia` [[phab:T433624|T433624]] * 17:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1083.eqiad.wmnet with reason: host reimage * 17:23 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 17:20 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 17:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2081.codfw.wmnet with OS trixie * 17:11 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new security plugin settings - bking@cumin2003 - [[phab:T350516|T350516]] * 17:10 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1083.eqiad.wmnet with OS trixie * 16:55 root@cumin1003: START - Cookbook sre.mysql.pool pool es1048: Maintenance * 16:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2081.codfw.wmnet with reason: host reimage * 16:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1048 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95833 and previous config saved to /var/cache/conftool/dbconfig/20260730-165053-cwilliams.json * 16:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1048.eqiad.wmnet with reason: Maintenance * 16:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1040: Maintenance * 16:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2081.codfw.wmnet with reason: host reimage * 16:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1082.eqiad.wmnet with OS trixie * 16:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2081.codfw.wmnet with OS trixie * 16:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2097.codfw.wmnet with OS trixie * 16:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1082.eqiad.wmnet with reason: host reimage * 16:08 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1082.eqiad.wmnet with reason: host reimage * 16:08 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 16:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1047: Maintenance * 16:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2080.codfw.wmnet with OS trixie * 16:04 root@cumin1003: START - Cookbook sre.mysql.pool pool es1040: Maintenance * 16:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1040: Maintenance * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new logging settings - bking@cumin2003 - [[phab:T324335|T324335]] * 15:58 root@cumin1003: START - Cookbook sre.mysql.pool pool es1040: Maintenance * 15:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1040 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95827 and previous config saved to /var/cache/conftool/dbconfig/20260730-155324-cwilliams.json * 15:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1040.eqiad.wmnet with reason: Maintenance * 15:50 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1082.eqiad.wmnet with OS trixie * 15:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2048: Maintenance * 15:44 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 24s) * 15:43 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2080.codfw.wmnet with reason: host reimage * 15:38 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new logging settings - bking@cumin2003 - [[phab:T324335|T324335]] * 15:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2080.codfw.wmnet with reason: host reimage * 15:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 15:30 mvernon@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be2097.codfw.wmnet with OS trixie * 15:23 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1081.eqiad.wmnet with OS trixie * 15:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2098.codfw.wmnet with OS trixie * 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - mvernon@cumin2003" * 15:18 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be2097.codfw.wmnet with OS trixie * 15:18 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - mvernon@cumin2003" * 15:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 15:17 root@cumin1003: START - Cookbook sre.mysql.pool pool es1047: Maintenance * 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2080.codfw.wmnet with OS trixie * 15:13 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be2097.codfw.wmnet with OS trixie * 15:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1047 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95820 and previous config saved to /var/cache/conftool/dbconfig/20260730-151200-cwilliams.json * 15:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1047.eqiad.wmnet with reason: Maintenance * 15:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: Maintenance * 15:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 15:04 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 15:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1081.eqiad.wmnet with reason: host reimage * 15:00 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1081.eqiad.wmnet with reason: host reimage * 15:00 root@cumin1003: START - Cookbook sre.mysql.pool pool es2048: Maintenance * 15:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 14:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2079.codfw.wmnet with OS trixie * 14:56 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 14:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2048 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95816 and previous config saved to /var/cache/conftool/dbconfig/20260730-145510-cwilliams.json * 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2048.codfw.wmnet with reason: Maintenance * 14:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2040: Maintenance * 14:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 14:51 tchin@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] (duration: 06m 48s) * 14:47 tchin@deploy1003: jforrester, tchin: Continuing with deployment * 14:47 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 14:47 tchin@deploy1003: jforrester, tchin: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:45 tchin@deploy1003: Started scap sync-world: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] * 14:42 sukhe@puppetserver1001: conftool action : set/weight=1; selector: cluster=urldownloader,service=squid * 14:42 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader,service=squid * 14:42 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1081.eqiad.wmnet with OS trixie * 14:39 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 14:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2079.codfw.wmnet with reason: host reimage * 14:36 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2098.codfw.wmnet with OS trixie * 14:32 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2079.codfw.wmnet with reason: host reimage * 14:30 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] (duration: 06m 31s) * 14:27 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 14:26 mszwarc@deploy1003: mszwarc: Continuing with deployment * 14:25 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:25 root@cumin1003: START - Cookbook sre.mysql.pool pool es1038: Maintenance * 14:25 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1038: Maintenance * 14:23 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] * 14:21 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] (duration: 11m 19s) * 14:20 root@cumin1003: START - Cookbook sre.mysql.pool pool es1038: Maintenance * 14:14 stran@deploy1003: stran: Continuing with deployment * 14:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1038 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95810 and previous config saved to /var/cache/conftool/dbconfig/20260730-141439-cwilliams.json * 14:14 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1038.eqiad.wmnet with reason: Maintenance * 14:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1036: Maintenance * 14:13 stran@deploy1003: stran: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2079.codfw.wmnet with OS trixie * 14:09 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] * 14:08 root@cumin1003: START - Cookbook sre.mysql.pool pool es2040: Maintenance * 14:08 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2040: Maintenance * 14:03 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] (duration: 31m 41s) * 14:03 root@cumin1003: START - Cookbook sre.mysql.pool pool es2040: Maintenance * 14:03 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2040 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95806 and previous config saved to /var/cache/conftool/dbconfig/20260730-135643-cwilliams.json * 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2040.codfw.wmnet with reason: Maintenance * 13:56 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:56 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2038: Maintenance * 13:55 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:52 lucaswerkmeister-wmde@deploy1003: migr, lucaswerkmeister-wmde: Continuing with deployment * 13:49 lucaswerkmeister-wmde@deploy1003: migr, lucaswerkmeister-wmde: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:49 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:48 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 13:45 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 13:32 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:32 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] * 13:28 root@cumin1003: START - Cookbook sre.mysql.pool pool es1036: Maintenance * 13:28 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1036: Maintenance * 13:22 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:22 root@cumin1003: START - Cookbook sre.mysql.pool pool es1036: Maintenance * 13:20 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1036 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95800 and previous config saved to /var/cache/conftool/dbconfig/20260730-131727-cwilliams.json * 13:17 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1036.eqiad.wmnet with reason: Maintenance * 13:17 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] (duration: 10m 31s) * 13:16 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2022\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 13:13 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, stran: Continuing with deployment * 13:10 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:10 root@cumin1003: START - Cookbook sre.mysql.pool pool es2038: Maintenance * 13:10 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2038: Maintenance * 13:08 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, stran: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie * 13:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2047: Maintenance * 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048 cloud-private - filippo@cumin1003" * 13:07 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048 cloud-private - filippo@cumin1003" * 13:06 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] * 13:04 root@cumin1003: START - Cookbook sre.mysql.pool pool es2038: Maintenance * 13:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2078.codfw.wmnet with OS trixie * 13:01 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2038 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95797 and previous config saved to /var/cache/conftool/dbconfig/20260730-125919-cwilliams.json * 12:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2038.codfw.wmnet with reason: Maintenance * 12:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2078.codfw.wmnet with reason: host reimage * 12:37 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2078.codfw.wmnet with reason: host reimage * 12:37 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] (duration: 06m 51s) * 12:33 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 12:32 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 12:32 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 12:32 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:30 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] * 12:19 root@cumin1003: START - Cookbook sre.mysql.pool pool es2047: Maintenance * 12:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2078.codfw.wmnet with OS trixie * 12:18 dcausse@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 12:18 dcausse@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 12:15 dcausse@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 12:14 dcausse@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 12:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2047 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95793 and previous config saved to /var/cache/conftool/dbconfig/20260730-121404-cwilliams.json * 12:13 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2047.codfw.wmnet with reason: Maintenance * 12:13 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2036: Maintenance * 12:05 ayounsi@dns1004: END - running authdns-update * 12:02 ayounsi@dns1004: START - running authdns-update * 11:51 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2077.codfw.wmnet with OS trixie * 11:48 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:46 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1080.eqiad.wmnet with OS trixie * 11:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1226: Maintenance * 11:41 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2077.codfw.wmnet with reason: host reimage * 11:28 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2077.codfw.wmnet with reason: host reimage * 11:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1080.eqiad.wmnet with reason: host reimage * 11:27 root@cumin1003: START - Cookbook sre.mysql.pool pool es2036: Maintenance * 11:27 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2036: Maintenance * 11:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1080.eqiad.wmnet with reason: host reimage * 11:21 root@cumin1003: START - Cookbook sre.mysql.pool pool es2036: Maintenance * 11:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2036 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95786 and previous config saved to /var/cache/conftool/dbconfig/20260730-111633-cwilliams.json * 11:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2036.codfw.wmnet with reason: Maintenance * 11:08 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2077.codfw.wmnet with OS trixie * 11:07 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 11:03 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie * 11:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1226: Maintenance * 10:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1226 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95783 and previous config saved to /var/cache/conftool/dbconfig/20260730-104801-cwilliams.json * 10:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1226.eqiad.wmnet with reason: Maintenance * 10:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1214: Maintenance * 10:27 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2035: Maintenance * 10:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2076.codfw.wmnet with OS trixie * 10:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1214: Maintenance * 09:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1214 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95775 and previous config saved to /var/cache/conftool/dbconfig/20260730-095451-cwilliams.json * 09:54 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1214.eqiad.wmnet with reason: Maintenance * 09:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1209: Maintenance * 09:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2076.codfw.wmnet with reason: host reimage * 09:42 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool es2035: Maintenance * 09:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.netbox.update-extras (exit_code=0) rolling restart_daemons on A:netbox * 09:41 ayounsi@cumin1003: START - Cookbook sre.netbox.update-extras rolling restart_daemons on A:netbox * 09:40 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2035: Maintenance * 09:39 ayounsi@cumin1003: END (PASS) - Cookbook sre.netbox.update-extras (exit_code=0) rolling restart_daemons on A:netbox-canary * 09:39 ayounsi@cumin1003: START - Cookbook sre.netbox.update-extras rolling restart_daemons on A:netbox-canary * 09:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2076.codfw.wmnet with reason: host reimage * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 09:34 root@cumin1003: START - Cookbook sre.mysql.pool pool es2035: Maintenance * 09:32 lucaswerkmeister-wmde@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 09:32 lucaswerkmeister-wmde@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 09:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2035 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95771 and previous config saved to /var/cache/conftool/dbconfig/20260730-092910-cwilliams.json * 09:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2035.codfw.wmnet with reason: Maintenance * 09:19 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2076.codfw.wmnet with OS trixie * 09:18 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 09:17 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie * 09:07 root@cumin1003: START - Cookbook sre.mysql.pool pool db1209: Maintenance * 09:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 23 hosts * 09:04 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Remove cable label from interfaces descriptions - ayounsi@cumin1003 * 09:04 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:02 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Remove cable label from interfaces descriptions - ayounsi@cumin1003 * 09:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1209 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95767 and previous config saved to /var/cache/conftool/dbconfig/20260730-090133-cwilliams.json * 09:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1209.eqiad.wmnet with reason: Maintenance * 09:01 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1192: Maintenance * 08:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1252: Maintenance * 08:57 jayme@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 08:56 jayme@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 08:53 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 23 hosts * 08:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:51 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:50 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1263: Maintenance * 08:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:23 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 08:15 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie * 08:14 root@cumin1003: START - Cookbook sre.mysql.pool pool db1192: Maintenance * 08:13 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2075.codfw.wmnet with OS trixie * 08:12 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1252: Maintenance * 08:11 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1252: Maintenance * 08:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1252: Maintenance * 08:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1192 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95754 and previous config saved to /var/cache/conftool/dbconfig/20260730-080611-cwilliams.json * 08:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1192.eqiad.wmnet with reason: Maintenance * 08:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1178: Maintenance * 08:05 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS bullseye * 07:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1263: Maintenance * 07:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1263 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95751 and previous config saved to /var/cache/conftool/dbconfig/20260730-075106-cwilliams.json * 07:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[1260-1262].eqiad.wmnet with reason: Maintenance * 07:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2075.codfw.wmnet with reason: host reimage * 07:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1263.eqiad.wmnet with reason: Maintenance * 07:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2075.codfw.wmnet with reason: host reimage * 07:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance * 07:38 dcausse@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:38 dcausse@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 07:35 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db1252', diff saved to https://phabricator.wikimedia.org/P95748 and previous config saved to /var/cache/conftool/dbconfig/20260730-073510-marostegui.json * 07:26 klausman@dns2004: END - running authdns-update * 07:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 07:25 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2075.codfw.wmnet with OS trixie * 07:24 klausman@dns2004: START - running authdns-update * 07:23 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host an-test-master1003.eqiad.wmnet * 07:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db1178: Maintenance * 07:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1178 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95746 and previous config saved to /var/cache/conftool/dbconfig/20260730-071112-cwilliams.json * 07:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1178.eqiad.wmnet with reason: Maintenance * 07:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1177: Maintenance * 06:24 root@cumin1003: START - Cookbook sre.mysql.pool pool db1177: Maintenance * 06:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1177 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95741 and previous config saved to /var/cache/conftool/dbconfig/20260730-061736-cwilliams.json * 06:17 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1177.eqiad.wmnet with reason: Maintenance * 06:17 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1172: Maintenance * 05:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1218.eqiad.wmnet with reason: crashed * 05:41 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db1217 it crashed', diff saved to https://phabricator.wikimedia.org/P95737 and previous config saved to /var/cache/conftool/dbconfig/20260730-054111-marostegui.json * 05:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95736 and previous config saved to /var/cache/conftool/dbconfig/20260730-053422-cwilliams.json * 05:30 root@cumin1003: START - Cookbook sre.mysql.pool pool db1172: Maintenance * 05:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95734 and previous config saved to /var/cache/conftool/dbconfig/20260730-052414-cwilliams.json * 05:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1172 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95733 and previous config saved to /var/cache/conftool/dbconfig/20260730-052354-cwilliams.json * 05:23 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1172.eqiad.wmnet with reason: Maintenance * 05:23 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1167: Maintenance * 05:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95731 and previous config saved to /var/cache/conftool/dbconfig/20260730-051406-cwilliams.json * 05:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95729 and previous config saved to /var/cache/conftool/dbconfig/20260730-050358-cwilliams.json * 04:47 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 04:47 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 04:47 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 04:35 root@cumin1003: START - Cookbook sre.mysql.pool pool db1167: Maintenance * 04:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1167 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95726 and previous config saved to /var/cache/conftool/dbconfig/20260730-042923-cwilliams.json * 04:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 04:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1167.eqiad.wmnet with reason: Maintenance * 04:22 pt1979@cumin2002: START - Cookbook sre.dns.netbox * 04:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95725 and previous config saved to /var/cache/conftool/dbconfig/20260730-040337-cwilliams.json * 04:03 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:38 brett@cumin2002: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool eqsin [reason: Switch upgrade maintenance window complete, [[phab:T433097|T433097]]] * 01:38 brett@cumin2002: START - Cookbook sre.dns.admin DNS admin: pool eqsin [reason: Switch upgrade maintenance window complete, [[phab:T433097|T433097]]] == 2026-07-29 == * 23:57 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin,mr1-eqsin IPv6,mr1-eqsin.oob,mr1-eqsin.oob IPv6 with reason: connection issue * 22:54 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 22:53 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 22:53 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 22:53 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:25 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2022.codfw.wmnet, repooling source-only afterwards * 22:20 brett@cumin2002: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool eqsin [reason: Switch upgrade maintenance window, [[phab:T433097|T433097]]] * 22:20 brett@cumin2002: START - Cookbook sre.dns.admin DNS admin: depool eqsin [reason: Switch upgrade maintenance window, [[phab:T433097|T433097]]] * 22:01 apine@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 22:00 apine@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 21:59 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 21:58 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 21:58 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 21:58 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 21:32 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1048.eqiad.wmnet with OS trixie * 21:25 pt1979@cumin2002: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be2097.codfw.wmnet with OS bullseye * 21:16 zabe@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki=metawiki 'Mental Health Resource Center' 'Safety Resource Center/Mental Health' Zabe --reason 'per request [[:phab:T433118{{!}}T433118]]' * 21:12 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2022.codfw.wmnet, repooling source-only afterwards * 21:12 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] (duration: 12m 53s) * 21:12 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2015\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 21:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1253: Maintenance * 21:08 aaron@deploy1003: aaron: Continuing with deployment * 21:01 aaron@deploy1003: aaron: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:59 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] * 20:52 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] (duration: 21m 57s) * 20:48 aaron@deploy1003: aaron: Continuing with deployment * 20:32 aaron@deploy1003: aaron: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:30 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] * 20:24 root@cumin1003: START - Cookbook sre.mysql.pool pool db1253: Maintenance * 20:19 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] (duration: 08m 07s) * 20:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1253 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95719 and previous config saved to /var/cache/conftool/dbconfig/20260729-201810-cwilliams.json * 20:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1253.eqiad.wmnet with reason: Maintenance * 20:17 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1231: Maintenance * 20:15 aaron@deploy1003: bpirkle, aaron: Continuing with deployment * 20:13 aaron@deploy1003: bpirkle, aaron: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:12 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie * 20:11 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] * 20:11 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1048.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:09 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1048.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:09 pt1979@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 20:04 pt1979@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 19:47 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 19:43 pt1979@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye * 19:41 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 19:37 zabe: zabe@deploy1003:~$ mwscript-k8s --comment='[[phab:T433529|T433529]]' --follow -- resetAuthenticationThrottle.php --wiki=aawiki --signup --ip=89.36.114.94 * 19:36 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] (duration: 06m 49s) * 19:32 zabe@deploy1003: zabe: Continuing with deployment * 19:31 zabe@deploy1003: zabe: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:31 root@cumin1003: START - Cookbook sre.mysql.pool pool db1231: Maintenance * 19:29 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] * 19:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1231 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95714 and previous config saved to /var/cache/conftool/dbconfig/20260729-192454-cwilliams.json * 19:24 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1231.eqiad.wmnet with reason: Maintenance * 19:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1227: Maintenance * 19:22 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 19:22 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 19:21 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:21 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1048] - vriley@cumin1003" * 19:21 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1048] - vriley@cumin1003" * 19:19 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 19:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1251: Maintenance * 19:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95711 and previous config saved to /var/cache/conftool/dbconfig/20260729-191756-cwilliams.json * 19:16 vriley@cumin1003: START - Cookbook sre.dns.netbox * 19:11 dduvall: rolling back wmf.13 to group0 due to [[phab:T433457|T433457]] (cc [[phab:T430832|T430832]]) * 19:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95709 and previous config saved to /var/cache/conftool/dbconfig/20260729-190748-cwilliams.json * 19:01 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2022.codfw.wmnet with OS bookworm * 19:01 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2021.codfw.wmnet, repooling source-only afterwards * 18:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95707 and previous config saved to /var/cache/conftool/dbconfig/20260729-185740-cwilliams.json * 18:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95704 and previous config saved to /var/cache/conftool/dbconfig/20260729-184732-cwilliams.json * 18:37 root@cumin1003: START - Cookbook sre.mysql.pool pool db1227: Maintenance * 18:34 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2022.codfw.wmnet with reason: host reimage * 18:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1227 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95701 and previous config saved to /var/cache/conftool/dbconfig/20260729-183117-cwilliams.json * 18:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1227.eqiad.wmnet with reason: Maintenance * 18:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1202: Maintenance * 18:30 root@cumin1003: START - Cookbook sre.mysql.pool pool db1251: Maintenance * 18:27 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2022.codfw.wmnet with reason: host reimage * 18:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1251 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95698 and previous config saved to /var/cache/conftool/dbconfig/20260729-182428-cwilliams.json * 18:24 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lvs2014.codfw.wmnet * 18:24 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for lvs2014.codfw.wmnet * 18:24 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1251.eqiad.wmnet with reason: Maintenance * 18:23 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1235: Maintenance * 18:22 brett@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 18:19 brett@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 18:19 brett@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 18:17 brett@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 18:17 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 18:16 mutante: removing jenkins during the train - living on the edge - no, just kidding, jenkins has migrated to dedicated machines, nothing should happen * 18:15 brett@cumin2002: END (ERROR) - Cookbook sre.loadbalancer.restart-pybal (exit_code=97) rolling-restart of pybal on P<nowiki>{</nowiki>lvs2014.codfw.wmnet<nowiki>}</nowiki> and A:lvs ([[phab:T428495|T428495]]) * 18:15 mutante: CI: contint1002/contint2002: apt-get remove --purge jenkins - jenkins be gone - [[phab:T418521|T418521]] * 18:13 brett@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on P<nowiki>{</nowiki>lvs2014.codfw.wmnet<nowiki>}</nowiki> and A:lvs ([[phab:T428495|T428495]]) * 18:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2022 * 18:08 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2022 * 18:03 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T428495|T428495]] * 18:03 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2022 * 18:02 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2022.codfw.wmnet 211.48.192.10.in-addr.arpa 1.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:02 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2022.codfw.wmnet 211.48.192.10.in-addr.arpa 1.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:02 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:02 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2022 - bking@cumin2003" * 18:02 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2022 - bking@cumin2003" * 17:57 bking@cumin2003: START - Cookbook sre.dns.netbox * 17:56 brett@cumin2002: END (FAIL) - Cookbook sre.loadbalancer.restart-pybal (exit_code=1) rolling-restart of pybal on A:lvs-codfw and A:lvs ([[phab:T428495|T428495]]) * 17:55 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo - [[phab:T428495|T428495]] * 17:54 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2022 * 17:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2022.codfw.wmnet with OS bookworm * 17:50 brett@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on A:lvs-codfw and A:lvs ([[phab:T428495|T428495]]) * 17:47 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2021.codfw.wmnet, repooling source-only afterwards * 17:47 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 14s) * 17:47 swfrench-wmf: authdns-update to direct codfw, eqsin, ulsfo etcd clients back to codfw - [[phab:T428495|T428495]] * 17:47 swfrench@dns1004: END - running authdns-update * 17:47 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 17:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95692 and previous config saved to /var/cache/conftool/dbconfig/20260729-174713-cwilliams.json * 17:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance * 17:46 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1249: Maintenance * 17:45 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2015\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 17:45 swfrench@dns1004: START - running authdns-update * 17:44 root@cumin1003: START - Cookbook sre.mysql.pool pool db1202: Maintenance * 17:41 akhatun: Deployed refinery using scap, then deployed onto hdfs * 17:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1202 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95688 and previous config saved to /var/cache/conftool/dbconfig/20260729-173759-cwilliams.json * 17:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1202.eqiad.wmnet with reason: Maintenance * 17:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1194: Maintenance * 17:37 root@cumin1003: START - Cookbook sre.mysql.pool pool db1235: Maintenance * 17:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1230: Maintenance * 17:30 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1235 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95684 and previous config saved to /var/cache/conftool/dbconfig/20260729-173051-cwilliams.json * 17:30 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1235.eqiad.wmnet with reason: Maintenance * 17:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1234: Maintenance * 17:26 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (thin): Regular analytics weekly train THIN [analytics/refinery@56695674] (duration: 02m 02s) * 17:24 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (thin): Regular analytics weekly train THIN [analytics/refinery@56695674] * 17:23 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567]: Regular analytics weekly train [analytics/refinery@56695674] (duration: 06m 20s) * 17:20 dancy@deploy1003: Finished scap sync-world: Testing delay_messageblobstore_purge: true (duration: 06m 29s) * 17:17 akhatun@deploy1003: Started deploy [analytics/refinery@5669567]: Regular analytics weekly train [analytics/refinery@56695674] * 17:17 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] (duration: 00m 22s) * 17:16 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] * 17:13 dancy@deploy1003: Started scap sync-world: Testing delay_messageblobstore_purge: true * 17:05 mutante: CI: contint1002/contint2002 - restarted httpd to be extra sure all is cleaned up - https://integration.wikimedia.org/ci/ is up and running [[phab:T418521|T418521]] * 17:04 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 17:03 mutante: CI: contint1002/contint2002 - rm /etc/apache2/jenkins_proxy - removing legacy jenkins proxy config - jenkins is on new dedicated machines and uses jenkins_proxy_ext config [[phab:T418521|T418521]] * 17:02 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] (duration: 36m 25s) * 17:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1249: Maintenance * 16:59 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 16:54 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2015.codfw.wmnet, repooling source-only afterwards * 16:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1249 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95674 and previous config saved to /var/cache/conftool/dbconfig/20260729-165339-cwilliams.json * 16:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1249.eqiad.wmnet with reason: Maintenance * 16:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1248: Maintenance * 16:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1194: Maintenance * 16:47 swfrench-wmf: silenced EtcdReplicationDown 57b2b421-1cc9-4e38-9276-{{Gerrit|94f223fd231c}} - [[phab:T428495|T428495]] * 16:46 tchin@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/eventstreams-internal: apply * 16:46 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye * 16:46 tchin@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/eventstreams-internal: apply * 16:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1230: Maintenance * 16:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1194 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95669 and previous config saved to /var/cache/conftool/dbconfig/20260729-164422-cwilliams.json * 16:44 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 16:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1194.eqiad.wmnet with reason: Maintenance * 16:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1191: Maintenance * 16:43 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host an-test-master1003.eqiad.wmnet * 16:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db1234: Maintenance * 16:43 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 16:43 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Rolling back deployment * 16:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host an-test-master1004.eqiad.wmnet * 16:41 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 16:40 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 16:40 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 16:39 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 16:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1230 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95667 and previous config saved to /var/cache/conftool/dbconfig/20260729-163932-cwilliams.json * 16:39 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 16:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1230.eqiad.wmnet with reason: Maintenance * 16:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1207: Maintenance * 16:38 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 16:37 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host an-test-master1004.eqiad.wmnet * 16:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1234 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95664 and previous config saved to /var/cache/conftool/dbconfig/20260729-163719-cwilliams.json * 16:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1234.eqiad.wmnet with reason: Maintenance * 16:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1079.eqiad.wmnet with OS trixie * 16:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1232: Maintenance * 16:34 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1259: Maintenance * 16:28 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] (duration: 06m 57s) * 16:28 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:26 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] * 16:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1051 hosts * 16:21 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] * 16:20 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2006.codfw.wmnet with OS bookworm * 16:19 akhatun: Deploying Refinery at {{Gerrit|56695674}} as part of weekly train * 16:18 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1079.eqiad.wmnet with reason: host reimage * 16:16 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] (duration: 15m 36s) * 16:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2021.codfw.wmnet with OS bookworm * 16:14 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1079.eqiad.wmnet with reason: host reimage * 16:12 topranks: hot-swap line card in FPC0 on cr1-eqiad with replacement MPC10E from Juniper [[phab:T426343|T426343]] * 16:10 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Continuing with deployment * 16:07 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db1248: Maintenance * 16:01 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] * 16:00 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 16:00 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 15:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1248 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95651 and previous config saved to /var/cache/conftool/dbconfig/20260729-155956-cwilliams.json * 15:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1248.eqiad.wmnet with reason: Maintenance * 15:59 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2006.codfw.wmnet with reason: host reimage * 15:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1247: Maintenance * 15:59 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2074.codfw.wmnet with OS trixie * 15:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1191: Maintenance * 15:57 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:55 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1079.eqiad.wmnet with OS trixie * 15:55 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2006.codfw.wmnet with reason: host reimage * 15:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1207: Maintenance * 15:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1191 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95646 and previous config saved to /var/cache/conftool/dbconfig/20260729-155104-cwilliams.json * 15:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1191.eqiad.wmnet with reason: Maintenance * 15:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1181: Maintenance * 15:49 root@cumin1003: START - Cookbook sre.mysql.pool pool db1232: Maintenance * 15:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2021.codfw.wmnet with reason: host reimage * 15:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1207 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95643 and previous config saved to /var/cache/conftool/dbconfig/20260729-154735-cwilliams.json * 15:47 root@cumin1003: START - Cookbook sre.mysql.pool pool db1259: Maintenance * 15:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1207.eqiad.wmnet with reason: Maintenance * 15:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1200: Maintenance * 15:46 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] (duration: 31m 59s) * 15:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 15:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1232 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95640 and previous config saved to /var/cache/conftool/dbconfig/20260729-154330-cwilliams.json * 15:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1232.eqiad.wmnet with reason: Maintenance * 15:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1219: Maintenance * 15:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2015.codfw.wmnet, repooling source-only afterwards * 15:41 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 18s) * 15:41 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1259 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95638 and previous config saved to /var/cache/conftool/dbconfig/20260729-154107-cwilliams.json * 15:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1259.eqiad.wmnet with reason: Maintenance * 15:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1254: Maintenance * 15:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2015.codfw.wmnet with OS bookworm * 15:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2021.codfw.wmnet with reason: host reimage * 15:36 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2006.codfw.wmnet with OS bookworm * 15:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 15:35 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Continuing with deployment * 15:33 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2074.codfw.wmnet with OS trixie * 15:33 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 15:32 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:29 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:28 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be2074.codfw.wmnet with OS trixie * 15:28 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2006.codfw.wmnet * 15:26 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1078.eqiad.wmnet with OS trixie * 15:25 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 15:25 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:22 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2006.codfw.wmnet * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2021 * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2021 * 15:19 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2021 * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2021.codfw.wmnet 210.48.192.10.in-addr.arpa 0.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:19 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2021.codfw.wmnet 210.48.192.10.in-addr.arpa 0.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2021 - bking@cumin2003" * 15:19 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2021 - bking@cumin2003" * 15:14 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] * 15:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2015.codfw.wmnet with reason: host reimage * 15:11 root@cumin1003: START - Cookbook sre.mysql.pool pool db1247: Maintenance * 15:11 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on ml-serve2004.codfw.wmnet with reason: [[phab:T433478|T433478]] * 15:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2015.codfw.wmnet with reason: host reimage * 15:10 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on ml-serve2002.codfw.wmnet with reason: [[phab:T433476|T433476]] * 15:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 15:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1247 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95625 and previous config saved to /var/cache/conftool/dbconfig/20260729-150459-cwilliams.json * 15:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1247.eqiad.wmnet with reason: Maintenance * 15:04 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:04 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1244: Maintenance * 15:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1078.eqiad.wmnet with reason: host reimage * 15:03 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1005.wikimedia.org * 15:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db1181: Maintenance * 15:01 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2098.codfw.wmnet with OS bullseye * 15:00 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye * 15:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1200: Maintenance * 14:59 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 14:59 root@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1285.eqiad.wmnet with OS trixie * 14:59 Amir1: mwscript-k8s -- extensions/TimedMediaHandler/maintenance/requeueTranscodes.php --wiki=commonswiki --key '360p.mpeg4.mov' --throttle --video --missing ([[phab:T358266|T358266]]) * 14:58 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1005.wikimedia.org * 14:58 jhancock@cumin2002: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['ms-be2098'] * 14:58 jhancock@cumin2002: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['ms-be2098'] * 14:58 jhancock@cumin2002: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['ms-be2097'] * 14:58 jhancock@cumin2002: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['ms-be2097'] * 14:58 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1078.eqiad.wmnet with reason: host reimage * 14:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1181 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95621 and previous config saved to /var/cache/conftool/dbconfig/20260729-145629-cwilliams.json * 14:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1181.eqiad.wmnet with reason: Maintenance * 14:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1174: Maintenance * 14:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1219: Maintenance * 14:55 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader1006.wikimedia.org on all recursors * 14:55 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader1006.wikimedia.org on all recursors * 14:55 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader1005.wikimedia.org on all recursors * 14:55 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader1005.wikimedia.org on all recursors * 14:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1200 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95618 and previous config saved to /var/cache/conftool/dbconfig/20260729-145336-cwilliams.json * 14:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1254: Maintenance * 14:53 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1200.eqiad.wmnet with reason: Maintenance * 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1185: Maintenance * 14:52 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2021 * 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2015 * 14:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2015 * 14:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1219 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95616 and previous config saved to /var/cache/conftool/dbconfig/20260729-144946-cwilliams.json * 14:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1219.eqiad.wmnet with reason: Maintenance * 14:49 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1218: Maintenance * 14:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2021.codfw.wmnet with OS bookworm * 14:48 dancy@deploy1003: Finished deploy [zuul/deploy@22703a6]: Deploying https://gerrit.wikimedia.org/r/c/integration/zuul/+/1311501 ([[phab:T432491|T432491]]) (duration: 00m 15s) * 14:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2015.codfw.wmnet with OS bookworm * 14:48 dancy@deploy1003: Started deploy [zuul/deploy@22703a6]: Deploying https://gerrit.wikimedia.org/r/c/integration/zuul/+/1311501 ([[phab:T432491|T432491]]) * 14:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1254 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95613 and previous config saved to /var/cache/conftool/dbconfig/20260729-144729-cwilliams.json * 14:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1254.eqiad.wmnet with reason: Maintenance * 14:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1233: Maintenance * 14:46 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2013\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 14:46 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2014\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 14:46 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:45 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:44 root@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1285.eqiad.wmnet with reason: host reimage * 14:43 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:42 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:41 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:40 root@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1285.eqiad.wmnet with reason: host reimage * 14:39 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1078.eqiad.wmnet with OS trixie * 14:39 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2074.codfw.wmnet with OS trixie * 14:32 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:32 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:32 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:31 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2005.codfw.wmnet with OS bookworm * 14:30 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:30 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:29 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:29 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:27 root@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host db1285 * 14:27 root@cumin1003: START - Cookbook sre.hosts.move-vlan for host db1285 * 14:27 root@cumin1003: START - Cookbook sre.hosts.reimage for host db1285.eqiad.wmnet with OS trixie * 14:24 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 14:24 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:24 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:24 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:23 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:22 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:22 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:21 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:17 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:16 root@cumin1003: START - Cookbook sre.mysql.pool pool db1244: Maintenance * 14:15 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:15 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add asw1-604 loopback ipv4 - pt1979@cumin2002" * 14:15 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add asw1-604 loopback ipv4 - pt1979@cumin2002" * 14:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:12 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 14:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95599 and previous config saved to /var/cache/conftool/dbconfig/20260729-141014-cwilliams.json * 14:10 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 14:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1244.eqiad.wmnet with reason: Maintenance * 14:10 pt1979@cumin2002: START - Cookbook sre.dns.netbox * 14:10 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 14:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1243: Maintenance * 14:09 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2005.codfw.wmnet with reason: host reimage * 14:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db1174: Maintenance * 14:08 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad * 14:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db1185: Maintenance * 14:06 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 14:05 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2005.codfw.wmnet with reason: host reimage * 14:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1174 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95595 and previous config saved to /var/cache/conftool/dbconfig/20260729-140309-cwilliams.json * 14:03 sukhe@dns1004: END - running authdns-update * 14:03 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1174.eqiad.wmnet with reason: Maintenance * 14:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1170: Maintenance * 14:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db1218: Maintenance * 14:01 sukhe@dns1004: START - running authdns-update * 14:00 sukhe@dns1004: START - running authdns-update * 13:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db1233: Maintenance * 13:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1185 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95592 and previous config saved to /var/cache/conftool/dbconfig/20260729-135925-cwilliams.json * 13:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1185.eqiad.wmnet with reason: Maintenance * 13:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1161: Maintenance * 13:58 sukhe@puppetserver1001: conftool action : set/pooled=true; selector: dnsdisc=urldownloader * 13:58 root@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1265.eqiad.wmnet with OS trixie * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1218 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95590 and previous config saved to /var/cache/conftool/dbconfig/20260729-135621-cwilliams.json * 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1218.eqiad.wmnet with reason: Maintenance * 13:55 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1206: Maintenance * 13:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2073.codfw.wmnet with OS trixie * 13:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1233 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95587 and previous config saved to /var/cache/conftool/dbconfig/20260729-135335-cwilliams.json * 13:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1233.eqiad.wmnet with reason: Maintenance * 13:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1229: Maintenance * 13:50 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/kartotherian: apply * 13:50 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service * 13:49 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:49 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/kartotherian: apply * 13:48 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 13:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1077.eqiad.wmnet with OS trixie * 13:47 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 13:46 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2005.codfw.wmnet with OS bookworm * 13:44 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 13:44 root@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1265.eqiad.wmnet with reason: host reimage * 13:40 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] (duration: 09m 22s) * 13:39 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:38 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:36 root@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1265.eqiad.wmnet with reason: host reimage * 13:35 stran@deploy1003: stran: Continuing with deployment * 13:33 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 13:32 stran@deploy1003: stran: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified t * 13:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2073.codfw.wmnet with reason: host reimage * 13:30 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ml-build1001.eqiad.wmnet * 13:30 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] * 13:29 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad * 13:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1077.eqiad.wmnet with reason: host reimage * 13:27 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 13:27 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:27 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad * 13:26 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2073.codfw.wmnet with reason: host reimage * 13:26 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] (duration: 07m 56s) * 13:25 klausman@cumin1003: START - Cookbook sre.hosts.reboot-single for host ml-build1001.eqiad.wmnet * 13:24 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 13:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>ml-serve101[2-5].eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 13:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1015.eqiad.wmnet * 13:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1015.eqiad.wmnet * 13:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1077.eqiad.wmnet with reason: host reimage * 13:23 root@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host db1265 * 13:23 root@cumin1003: START - Cookbook sre.hosts.move-vlan for host db1265 * 13:23 root@cumin1003: START - Cookbook sre.hosts.reimage for host db1265.eqiad.wmnet with OS trixie * 13:23 root@cumin1003: START - Cookbook sre.mysql.pool pool db1243: Maintenance * 13:22 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 13:22 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2005.codfw.wmnet * 13:22 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 13:22 samtar@deploy1003: dreamrimmer, samtar: Continuing with deployment * 13:20 samtar@deploy1003: dreamrimmer, samtar: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts an-test-master[1001-1002].eqiad.wmnet * 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-master[1001-1002].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 13:18 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1015.eqiad.wmnet * 13:18 sukhe@cumin1003: END (ERROR) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=97) for role: url_downloader@eqiad * 13:18 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 13:18 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] * 13:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95574 and previous config saved to /var/cache/conftool/dbconfig/20260729-131638-cwilliams.json * 13:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1243.eqiad.wmnet with reason: Maintenance * 13:16 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2005.codfw.wmnet * 13:16 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1242: Maintenance * 13:14 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] (duration: 07m 00s) * 13:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db1170: Maintenance * 13:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1015.eqiad.wmnet * 13:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1014.eqiad.wmnet * 13:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1014.eqiad.wmnet * 13:12 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1223: Maintenance * 13:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db1161: Maintenance * 13:10 samtar@deploy1003: anzx, samtar: Continuing with deployment * 13:09 samtar@deploy1003: anzx, samtar: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db1206: Maintenance * 13:08 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 13:07 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] * 13:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1170 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95566 and previous config saved to /var/cache/conftool/dbconfig/20260729-130730-cwilliams.json * 13:07 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1170.eqiad.wmnet with reason: Maintenance * 13:07 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:07 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt IPs new switches - cmooney@cumin1003" * 13:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1158: Maintenance * 13:06 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1014.eqiad.wmnet * 13:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1161 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95564 and previous config saved to /var/cache/conftool/dbconfig/20260729-130616-cwilliams.json * 13:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 13:06 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1077.eqiad.wmnet with OS trixie * 13:05 root@cumin1003: START - Cookbook sre.mysql.pool pool db1229: Maintenance * 13:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1161.eqiad.wmnet with reason: Maintenance * 13:05 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt IPs new switches - cmooney@cumin1003" * 13:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2073.codfw.wmnet with OS trixie * 13:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Maintenance * 13:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1206 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95562 and previous config saved to /var/cache/conftool/dbconfig/20260729-130258-cwilliams.json * 13:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1206.eqiad.wmnet with reason: Maintenance * 13:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1196: Maintenance * 13:01 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 13:01 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:00 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 13:00 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 12:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1229 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95559 and previous config saved to /var/cache/conftool/dbconfig/20260729-125950-cwilliams.json * 12:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1229.eqiad.wmnet with reason: Maintenance * 12:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1222: Maintenance * 12:57 sukhe: sudo cumin 'A:lvs and (A:eqiad or A:codfw)' 'disable-puppet "adding new service urldownloader"': [[phab:T429175|T429175]] * 12:56 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1014.eqiad.wmnet * 12:56 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1013.eqiad.wmnet * 12:56 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1013.eqiad.wmnet * 12:50 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1013.eqiad.wmnet * 12:50 sukhe: sudo cumin 'O:url_downloader' 'run-puppet-agent --enable "merging CR 1313948"': [[phab:T429175|T429175]] * 12:48 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-master[1001-1002].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 12:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1013.eqiad.wmnet * 12:45 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1012.eqiad.wmnet * 12:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1012.eqiad.wmnet * 12:45 sukhe: sudo cumin 'O:url_downloader' 'disable-puppet "merging CR 1313948"': [[phab:T429175|T429175]] * 12:44 btullis@cumin1003: START - Cookbook sre.dns.netbox * 12:40 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test2001.codfw.wmnet * 12:40 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test2001.codfw.wmnet * 12:38 ayounsi@dns1004: END - running authdns-update * 12:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1012.eqiad.wmnet * 12:35 ayounsi@dns1004: START - running authdns-update * 12:34 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts an-test-master[1001-1002].eqiad.wmnet * 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts an-test-coord1001.eqiad.wmnet * 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-coord1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 12:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1012.eqiad.wmnet * 12:32 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>ml-serve101[2-5].eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 12:29 root@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Maintenance * 12:25 root@cumin1003: START - Cookbook sre.mysql.pool pool db1223: Maintenance * 12:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95544 and previous config saved to /var/cache/conftool/dbconfig/20260729-122254-cwilliams.json * 12:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1242.eqiad.wmnet with reason: Maintenance * 12:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1241: Maintenance * 12:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1051 hosts * 12:20 root@cumin1003: START - Cookbook sre.mysql.pool pool db1158: Maintenance * 12:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1223 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95540 and previous config saved to /var/cache/conftool/dbconfig/20260729-121937-cwilliams.json * 12:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1223.eqiad.wmnet with reason: Maintenance * 12:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1212: Maintenance * 12:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db1159: Maintenance * 12:17 elukey@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: sync * 12:15 elukey@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: sync * 12:15 root@cumin1003: START - Cookbook sre.mysql.pool pool db1196: Maintenance * 12:14 Daimona: Creating new DB tables for the CampaignEvents extension in x1.testwiki, x1.test2wiki, x1.officewiki, and x1.wikishared # [[phab:T429339|T429339]] * 12:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db1222: Maintenance * 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95535 and previous config saved to /var/cache/conftool/dbconfig/20260729-121211-cwilliams.json * 12:12 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 12:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1158.eqiad.wmnet with reason: Maintenance * 12:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1159 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95534 and previous config saved to /var/cache/conftool/dbconfig/20260729-121146-cwilliams.json * 12:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1159.eqiad.wmnet with reason: Maintenance * 12:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1196 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95533 and previous config saved to /var/cache/conftool/dbconfig/20260729-120847-cwilliams.json * 12:08 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 12:08 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1196.eqiad.wmnet with reason: Maintenance * 12:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1195: Maintenance * 12:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1222 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95530 and previous config saved to /var/cache/conftool/dbconfig/20260729-120424-cwilliams.json * 12:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1222.eqiad.wmnet with reason: Maintenance * 12:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1098 hosts * 12:00 marostegui: Rename tables [[phab:T425074|T425074]] * 12:00 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-coord1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 11:58 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1197: Maintenance * 11:55 btullis@cumin1003: START - Cookbook sre.dns.netbox * 11:52 elukey@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: sync * 11:51 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:51 elukey@deploy1003: helmfile [codfw] START helmfile.d/services/proton: sync * 11:51 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:50 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts an-test-coord1001.eqiad.wmnet * 11:50 elukey@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: sync * 11:49 elukey@deploy1003: helmfile [staging] START helmfile.d/services/proton: sync * 11:35 root@cumin1003: START - Cookbook sre.mysql.pool pool db1241: Maintenance * 11:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db1212: Maintenance * 11:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1241 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95520 and previous config saved to /var/cache/conftool/dbconfig/20260729-112918-cwilliams.json * 11:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1241.eqiad.wmnet with reason: Maintenance * 11:29 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1238: Maintenance * 11:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1212 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95517 and previous config saved to /var/cache/conftool/dbconfig/20260729-112727-cwilliams.json * 11:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 11:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1212.eqiad.wmnet with reason: Maintenance * 11:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1198: Maintenance * 11:23 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:22 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:21 root@cumin1003: START - Cookbook sre.mysql.pool pool db1195: Maintenance * 11:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1195 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95514 and previous config saved to /var/cache/conftool/dbconfig/20260729-111450-cwilliams.json * 11:14 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1195.eqiad.wmnet with reason: Maintenance * 11:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1186: Maintenance * 11:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 11:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 11:05 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:54 marostegui: Dropping renamed tables [[phab:T425066|T425066]] * 10:41 root@cumin1003: START - Cookbook sre.mysql.pool pool db1238: Maintenance * 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1198: Maintenance * 10:39 Amir1: ran https://phabricator.wikimedia.org/T432509#12149723 in production ([[phab:T432509|T432509]]) * 10:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db1197: Maintenance * 10:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1238 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95501 and previous config saved to /var/cache/conftool/dbconfig/20260729-103532-cwilliams.json * 10:35 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1238.eqiad.wmnet with reason: Maintenance * 10:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1221: Maintenance * 10:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1198 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95499 and previous config saved to /var/cache/conftool/dbconfig/20260729-103330-cwilliams.json * 10:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1198.eqiad.wmnet with reason: Maintenance * 10:33 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1175: Maintenance * 10:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1197 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95496 and previous config saved to /var/cache/conftool/dbconfig/20260729-103217-cwilliams.json * 10:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1197.eqiad.wmnet with reason: Maintenance * 10:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1188: Maintenance * 10:27 root@cumin1003: START - Cookbook sre.mysql.pool pool db1186: Maintenance * 10:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1186 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95493 and previous config saved to /var/cache/conftool/dbconfig/20260729-102111-cwilliams.json * 10:21 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1186.eqiad.wmnet with reason: Maintenance * 10:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 10:14 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 09:53 XioNoX: reboot cr2-magru - [[phab:T431750|T431750]] * 09:52 XioNoX: drain cr2-magru - [[phab:T431750|T431750]] * 09:48 root@cumin1003: START - Cookbook sre.mysql.pool pool db1221: Maintenance * 09:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zookeeper-test1002.eqiad.wmnet * 09:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1188: Maintenance * 09:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1175: Maintenance * 09:44 btullis@dns1004: END - running authdns-update * 09:42 btullis@dns1004: START - running authdns-update * 09:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1221 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95483 and previous config saved to /var/cache/conftool/dbconfig/20260729-094200-cwilliams.json * 09:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 7 hosts with reason: Maintenance * 09:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1221.eqiad.wmnet with reason: Maintenance * 09:41 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host zookeeper-test1002.eqiad.wmnet * 09:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1199: Maintenance * 09:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1188 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95481 and previous config saved to /var/cache/conftool/dbconfig/20260729-093917-cwilliams.json * 09:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1188.eqiad.wmnet with reason: Maintenance * 09:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1182: Maintenance * 09:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1175 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95479 and previous config saved to /var/cache/conftool/dbconfig/20260729-093842-cwilliams.json * 09:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1175.eqiad.wmnet with reason: Maintenance * 09:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1166: Maintenance * 09:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1169: Maintenance * 09:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1033.eqiad.wmnet,service=s8 * 09:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1033.eqiad.wmnet,service=s5 * 09:33 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1033.eqiad.wmnet,service=s8 * 09:33 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1033.eqiad.wmnet,service=s5 * 09:21 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:21 XioNoX: reboot cr1-magru - [[phab:T431750|T431750]] * 09:17 XioNoX: drain cr1-magru - [[phab:T431750|T431750]] * 09:15 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm1001.wikimedia.org * 09:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply * 09:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply * 09:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 09:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 09:11 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr2-magru,cr2-magru IPv6,cr2-magru.mgmt with reason: router upgrade * 09:11 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm1001.wikimedia.org * 09:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp1005.wikimedia.org * 09:07 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp1005.wikimedia.org * 09:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp2005.wikimedia.org * 09:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp2005.wikimedia.org * 09:00 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr1-magru,cr1-magru IPv6,cr1-magru.mgmt with reason: router upgrade * 09:00 marostegui: Dropping renamed tables [[phab:T426341|T426341]] * 08:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1199: Maintenance * 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1182: Maintenance * 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1166: Maintenance * 08:47 root@cumin1003: START - Cookbook sre.mysql.pool pool db1169: Maintenance * 08:46 ayounsi@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 1:00:00 on cr1-magru,cr1-magru IPv6,cr1-magru.mgmt with reason: router upgrade * 08:45 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2072.codfw.wmnet with OS trixie * 08:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1199 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95464 and previous config saved to /var/cache/conftool/dbconfig/20260729-084534-cwilliams.json * 08:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1199.eqiad.wmnet with reason: Maintenance * 08:45 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 08:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1190: Maintenance * 08:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1182 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95462 and previous config saved to /var/cache/conftool/dbconfig/20260729-084436-cwilliams.json * 08:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1182.eqiad.wmnet with reason: Maintenance * 08:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1156: Maintenance * 08:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1166 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95460 and previous config saved to /var/cache/conftool/dbconfig/20260729-084400-cwilliams.json * 08:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1166.eqiad.wmnet with reason: Maintenance * 08:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1157: Maintenance * 08:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95458 and previous config saved to /var/cache/conftool/dbconfig/20260729-084147-cwilliams.json * 08:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1169.eqiad.wmnet with reason: Maintenance * 08:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1163: Maintenance * 08:30 btullis@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 11 hosts with reason: Replacing the namenodes * 08:23 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2072.codfw.wmnet with reason: host reimage * 08:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1098 hosts * 08:19 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2072.codfw.wmnet with reason: host reimage * 07:58 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2072.codfw.wmnet with OS trixie * 07:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1190: Maintenance * 07:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1156: Maintenance * 07:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1157: Maintenance * 07:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1163: Maintenance * 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1190 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95444 and previous config saved to /var/cache/conftool/dbconfig/20260729-074930-cwilliams.json * 07:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1190.eqiad.wmnet with reason: Maintenance * 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1157 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95443 and previous config saved to /var/cache/conftool/dbconfig/20260729-074914-cwilliams.json * 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1156 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95442 and previous config saved to /var/cache/conftool/dbconfig/20260729-074906-cwilliams.json * 07:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1157.eqiad.wmnet with reason: Maintenance * 07:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 07:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1156.eqiad.wmnet with reason: Maintenance * 07:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1163 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95441 and previous config saved to /var/cache/conftool/dbconfig/20260729-074652-cwilliams.json * 07:46 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1163.eqiad.wmnet with reason: Maintenance * 07:46 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2034.codfw.wmnet * 07:42 ayounsi@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2034.codfw.wmnet * 07:42 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2232.codfw.wmnet with OS trixie * 07:34 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:34 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:33 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:31 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:19 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2232.codfw.wmnet with reason: host reimage * 07:15 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2232.codfw.wmnet with reason: host reimage * 06:58 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db2232.codfw.wmnet with OS trixie * 06:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[2160,2232].codfw.wmnet with reason: Reimage * 06:26 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1164.eqiad.wmnet with OS trixie * 06:05 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1164.eqiad.wmnet with reason: host reimage * 06:01 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1164.eqiad.wmnet with reason: host reimage * 05:47 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1164.eqiad.wmnet with OS trixie * 05:46 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1164.eqiad.wmnet with reason: Reimage == 2026-07-28 == * 22:50 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1138.eqiad.wmnet * 22:50 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1138.eqiad.wmnet * 22:49 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1138.eqiad.wmnet * 22:11 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2014.codfw.wmnet, repooling source-only afterwards * 22:08 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2013.codfw.wmnet, repooling source-only afterwards * 22:03 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 20:58 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] (duration: 08m 19s) * 20:55 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2014.codfw.wmnet, repooling source-only afterwards * 20:55 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2013.codfw.wmnet, repooling source-only afterwards * 20:54 arlolra@deploy1003: arlolra: Continuing with deployment * 20:54 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 14s) * 20:54 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 20:53 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 30s) * 20:53 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 20:52 arlolra@deploy1003: arlolra: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:51 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:50 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] * 20:49 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:43 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 20:34 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] (duration: 06m 54s) * 20:34 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:34 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:31 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:31 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:30 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:30 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:30 arlolra@deploy1003: arlolra: Continuing with deployment * 20:29 arlolra@deploy1003: arlolra: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:27 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] * 20:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2014.codfw.wmnet with OS bookworm * 20:21 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 20:21 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:20 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 20:19 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:19 swfrench-wmf: switched etcd-mirror replication from conf2005 to conf2004 - [[phab:T428495|T428495]] * 20:17 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:17 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:15 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] (duration: 08m 26s) * 20:12 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:11 arlolra@deploy1003: anzx, arlolra: Continuing with deployment * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2013.codfw.wmnet with OS bookworm * 20:09 arlolra@deploy1003: anzx, arlolra: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] * 20:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2216: Maintenance * 19:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2014.codfw.wmnet with reason: host reimage * 19:57 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:54 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:54 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:53 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:52 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2014.codfw.wmnet with reason: host reimage * 19:49 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2013.codfw.wmnet with reason: host reimage * 19:42 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 19:41 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:41 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2013.codfw.wmnet with reason: host reimage * 19:39 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:39 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:39 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-eqiad: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 19:38 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2014 * 19:33 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2014 * 19:29 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2014.codfw.wmnet with OS bookworm * 19:28 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:27 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1006 * 19:26 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2012\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 19:26 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1006 * 19:26 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:26 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 19:25 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2013 * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2013 * 19:21 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2013 * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2013.codfw.wmnet 84.0.192.10.in-addr.arpa 4.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:21 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2013.codfw.wmnet 84.0.192.10.in-addr.arpa 4.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2013 - bking@cumin2003" * 19:21 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2013 - bking@cumin2003" * 19:21 vriley@cumin1003: START - Cookbook sre.dns.netbox * 19:20 root@cumin1003: START - Cookbook sre.mysql.pool pool db2216: Maintenance * 19:13 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2216 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95435 and previous config saved to /var/cache/conftool/dbconfig/20260728-191343-cwilliams.json * 19:13 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2216.codfw.wmnet with reason: Maintenance * 19:13 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2203: Maintenance * 19:06 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1005.eqiad.wmnet with OS trixie * 19:06 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 19:06 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 18:46 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 18:45 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:45 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:43 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:40 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:36 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-eqiad: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 18:35 dancy@deploy1003: Installation of scap version "4.275.0" completed for 3 hosts * 18:33 dancy@deploy1003: Installing scap version "4.275.0" for 3 host(s) * 18:32 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:32 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2097.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:30 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns3003.wikimedia.org [reason: pool for all services after reimaging] * 18:29 sukhe@dns1004: END - running authdns-update * 18:27 sukhe@dns1004: START - running authdns-update * 18:27 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns3003.wikimedia.org,service=authdns-update [reason: pool authdns-update after reimaging] * 18:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db2203: Maintenance * 18:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2203 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95430 and previous config saved to /var/cache/conftool/dbconfig/20260728-181958-cwilliams.json * 18:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2203.codfw.wmnet with reason: Maintenance * 18:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2188: Maintenance * 18:18 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2097.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:17 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-be2098 * 18:17 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host ms-be2098 * 18:17 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-be2097 * 18:16 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host ms-be2097 * 18:15 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:15 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding ms-be2097-8 to codfw - jhancock@cumin2002" * 18:15 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding ms-be2097-8 to codfw - jhancock@cumin2002" * 18:10 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 18:08 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage * 18:05 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns3003.wikimedia.org with OS trixie * 18:03 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage * 17:56 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-codfw: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 17:45 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie * 17:45 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1005.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:41 sukhe@dns1004: END - running authdns-update * 17:39 sukhe@dns1004: START - running authdns-update * 17:36 sukhe@puppetserver1001: conftool action : set/weight=1; selector: cluster=urldownloader,service=squid * 17:36 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1005.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:35 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader,service=squid * 17:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 17:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1005 * 17:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 17:34 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1005 * 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1005] - vriley@cumin1003" * 17:34 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1005] - vriley@cumin1003" * 17:32 root@cumin1003: START - Cookbook sre.mysql.pool pool db2188: Maintenance * 17:29 vriley@cumin1003: START - Cookbook sre.dns.netbox * 17:29 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 17:26 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2188 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95425 and previous config saved to /var/cache/conftool/dbconfig/20260728-172609-cwilliams.json * 17:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2188.codfw.wmnet with reason: Maintenance * 17:25 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2176: Maintenance * 17:19 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1138.eqiad.wmnet with OS trixie * 17:18 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1005 * 17:18 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1005 * 17:18 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:15 vriley@cumin1003: START - Cookbook sre.dns.netbox * 17:13 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns3003.wikimedia.org with reason: host reimage * 17:07 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns3003.wikimedia.org with reason: host reimage * 17:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-eqiad * 17:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1015.eqiad.wmnet * 17:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1015.eqiad.wmnet * 16:59 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1138.eqiad.wmnet with reason: host reimage * 16:55 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-codfw: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 16:54 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1138.eqiad.wmnet with reason: host reimage * 16:53 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1015.eqiad.wmnet * 16:43 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns3003.wikimedia.org with OS trixie * 16:43 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1015.eqiad.wmnet * 16:43 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1014.eqiad.wmnet * 16:43 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1014.eqiad.wmnet * 16:43 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=dns3003.wikimedia.org [reason: depooling for reimage to trixie] * 16:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 290 hosts * 16:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db2176: Maintenance * 16:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1138 * 16:38 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1138 * 16:37 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1138 * 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1138.eqiad.wmnet 193.32.64.10.in-addr.arpa 3.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:37 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1138.eqiad.wmnet 193.32.64.10.in-addr.arpa 3.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1138 - jiji@cumin1003" * 16:37 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1138 - jiji@cumin1003" * 16:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1014.eqiad.wmnet * 16:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2176 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95420 and previous config saved to /var/cache/conftool/dbconfig/20260728-163235-cwilliams.json * 16:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2176.codfw.wmnet with reason: Maintenance * 16:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1014.eqiad.wmnet * 16:32 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1013.eqiad.wmnet * 16:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1013.eqiad.wmnet * 16:32 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2174: Maintenance * 16:28 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2012.codfw.wmnet, repooling source-only afterwards * 16:25 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1013.eqiad.wmnet * 16:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1013.eqiad.wmnet * 16:20 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1012.eqiad.wmnet * 16:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1012.eqiad.wmnet * 16:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1012.eqiad.wmnet * 16:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1012.eqiad.wmnet * 16:03 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1011.eqiad.wmnet * 16:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1011.eqiad.wmnet * 16:00 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2004.codfw.wmnet with OS bookworm * 15:59 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1011.eqiad.wmnet * 15:56 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:55 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 15:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:54 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1011.eqiad.wmnet * 15:54 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1010.eqiad.wmnet * 15:54 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1010.eqiad.wmnet * 15:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:50 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 15:49 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1010.eqiad.wmnet * 15:48 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 15:48 jiji@cumin1003: START - Cookbook sre.dns.netbox * 15:46 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2248: Maintenance * 15:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db2174: Maintenance * 15:44 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1010.eqiad.wmnet * 15:44 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1009.eqiad.wmnet * 15:44 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1009.eqiad.wmnet * 15:42 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1138 * 15:41 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1138.eqiad.wmnet with OS trixie * 15:39 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1009.eqiad.wmnet * 15:39 robh@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on arclamp2001.codfw.wmnet with reason: ram upgrade * 15:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2174 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95413 and previous config saved to /var/cache/conftool/dbconfig/20260728-153844-cwilliams.json * 15:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2174.codfw.wmnet with reason: Maintenance * 15:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2173: Maintenance * 15:37 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 15:35 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 15:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1009.eqiad.wmnet * 15:34 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1008.eqiad.wmnet * 15:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1008.eqiad.wmnet * 15:31 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 15:31 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 15:29 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1008.eqiad.wmnet * 15:27 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 290 hosts * 15:25 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2013 * 15:25 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2195: Maintenance * 15:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1008.eqiad.wmnet * 15:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1007.eqiad.wmnet * 15:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1007.eqiad.wmnet * 15:22 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2004.codfw.wmnet with reason: host reimage * 15:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2013.codfw.wmnet with OS bookworm * 15:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 15:19 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2004.codfw.wmnet with reason: host reimage * 15:19 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 15:17 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1007.eqiad.wmnet * 15:12 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1007.eqiad.wmnet * 15:12 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1006.eqiad.wmnet * 15:12 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1006.eqiad.wmnet * 15:11 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1138.eqiad.wmnet * 15:11 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 15:11 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1138.eqiad.wmnet * 15:11 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1138.eqiad.wmnet * 15:10 brennen@deploy1003: Finished deploy [phabricator/deployment@f8b349f]: deploy phab1004 for [[phab:T433382|T433382]] (duration: 00m 43s) * 15:10 brennen@deploy1003: Started deploy [phabricator/deployment@f8b349f]: deploy phab1004 for [[phab:T433382|T433382]] * 15:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts * 15:09 brennen@deploy1003: Finished deploy [phabricator/deployment@f8b349f]: deploy phab2003 for [[phab:T433382|T433382]] (duration: 00m 55s) * 15:08 brennen@deploy1003: Started deploy [phabricator/deployment@f8b349f]: deploy phab2003 for [[phab:T433382|T433382]] * 15:07 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2012.codfw.wmnet, repooling source-only afterwards * 15:07 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts * 15:06 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1137.eqiad.wmnet * 15:06 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1137.eqiad.wmnet * 15:06 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1137.eqiad.wmnet * 15:05 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1006.eqiad.wmnet * 15:05 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1022\.eqiad\.wmnet,dc=eqiad,cluster=wdqs\-main,service=wdqs\-main * 15:01 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab2003.codfw.wmnet with reason: deployment * 15:01 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1005.eqiad.wmnet with reason: deployment * 15:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1006.eqiad.wmnet * 15:00 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1006.eqiad.wmnet with reason: deployment * 15:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1005.eqiad.wmnet * 15:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1005.eqiad.wmnet * 14:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db2248: Maintenance * 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1004.eqiad.wmnet with reason: deployment * 14:59 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2004.codfw.wmnet with OS bookworm * 14:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1005.eqiad.wmnet * 14:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2248 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95403 and previous config saved to /var/cache/conftool/dbconfig/20260728-145532-cwilliams.json * 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2245-2247].codfw.wmnet with reason: Maintenance * 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2248.codfw.wmnet with reason: Maintenance * 14:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2240: Maintenance * 14:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2173: Maintenance * 14:51 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts * 14:50 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1005.eqiad.wmnet * 14:50 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1004.eqiad.wmnet * 14:50 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1004.eqiad.wmnet * 14:49 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts * 14:45 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2004.codfw.wmnet * 14:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2173 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95399 and previous config saved to /var/cache/conftool/dbconfig/20260728-144453-cwilliams.json * 14:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2173.codfw.wmnet with reason: Maintenance * 14:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2170: Maintenance * 14:44 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1004.eqiad.wmnet * 14:39 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2004.codfw.wmnet * 14:38 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1004.eqiad.wmnet * 14:38 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet * 14:38 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet * 14:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db2195: Maintenance * 14:36 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1022.eqiad.wmnet, repooling source-only afterwards * 14:36 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 14:33 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet * 14:33 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2222: Maintenance * 14:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2195 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95394 and previous config saved to /var/cache/conftool/dbconfig/20260728-143218-cwilliams.json * 14:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2195.codfw.wmnet with reason: Maintenance * 14:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2181: Maintenance * 14:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 14:30 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 14:25 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 14:25 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 14:25 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 14:23 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet * 14:23 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1002.eqiad.wmnet * 14:23 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1002.eqiad.wmnet * 14:23 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 19s) * 14:23 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:18 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1002.eqiad.wmnet * 14:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2012.codfw.wmnet with OS bookworm * 14:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1002.eqiad.wmnet * 14:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1001.eqiad.wmnet * 14:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1001.eqiad.wmnet * 14:11 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 14:11 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 14:08 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1001.eqiad.wmnet * 14:07 elukey@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'. * 14:07 elukey@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'. * 14:06 elukey@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'. * 14:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db2240: Maintenance * 14:06 XioNoX: un-drain cr2-esams - [[phab:T431751|T431751]] * 14:05 elukey@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'. * 14:02 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1001.eqiad.wmnet * 14:02 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-eqiad * 14:01 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 14:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2240 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95384 and previous config saved to /var/cache/conftool/dbconfig/20260728-140011-cwilliams.json * 14:00 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2240.codfw.wmnet with reason: Maintenance * 13:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2237: Maintenance * 13:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db2170: Maintenance * 13:56 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 13:55 XioNoX: reboot cr2-esams - [[phab:T431751|T431751]] * 13:52 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 13:51 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr2-esams,cr2-esams IPv6,cr2-esams.mgmt with reason: router upgrade * 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 13:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2170 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95381 and previous config saved to /var/cache/conftool/dbconfig/20260728-135043-cwilliams.json * 13:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2170.codfw.wmnet with reason: Maintenance * 13:50 XioNoX: drain cr2-esams - [[phab:T431751|T431751]] * 13:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2153: Maintenance * 13:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2012.codfw.wmnet with reason: host reimage * 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2222: Maintenance * 13:45 sukhe: restart pybal on A:lvs-codfw * 13:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2012.codfw.wmnet with reason: host reimage * 13:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db2181: Maintenance * 13:44 btullis@dns1004: END - running authdns-update * 13:42 sukhe: restart pybal on lvs2014 * 13:42 btullis@dns1004: START - running authdns-update * 13:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2222 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95376 and previous config saved to /var/cache/conftool/dbconfig/20260728-133948-cwilliams.json * 13:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2222.codfw.wmnet with reason: Maintenance * 13:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2221: Maintenance * 13:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2181 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95374 and previous config saved to /var/cache/conftool/dbconfig/20260728-133857-cwilliams.json * 13:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2181.codfw.wmnet with reason: Maintenance * 13:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2167: Maintenance * 13:30 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T428495|T428495]] * 13:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 13:29 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 13:29 ayounsi@cumin1003: END (FAIL) - Cookbook sre.dns.admin (exit_code=99) DNS admin: depool esams [reason: router upgrade, [[phab:T431749|T431749]]] * 13:28 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: router upgrade, [[phab:T431749|T431749]]] * 13:27 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1022.eqiad.wmnet, repooling source-only afterwards * 13:27 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo - [[phab:T428495|T428495]] * 13:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2012 * 13:27 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2012 * 13:21 lucaswerkmeister-wmde@deploy1003: mwscript-k8s job started: cleanupTitles bolwiki # [[phab:T429951|T429951]] * 13:21 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] (duration: 07m 19s) * 13:20 swfrench-wmf: authdns-update to direct codfw, eqsin, ulsfo etcd clients to eqiad - [[phab:T428495|T428495]] * 13:18 swfrench@dns1004: END - running authdns-update * 13:17 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, anzx: Continuing with deployment * 13:16 swfrench@dns1004: START - running authdns-update * 13:16 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2012 * 13:16 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2012.codfw.wmnet 57.48.192.10.in-addr.arpa 7.5.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:16 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2012.codfw.wmnet 57.48.192.10.in-addr.arpa 7.5.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:16 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:16 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, anzx: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:14 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 13:14 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] * 13:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db2237: Maintenance * 13:13 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2237: Maintenance * 13:13 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:12 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:12 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback IPV6 for asw1-604 - pt1979@cumin2003" * 13:12 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:12 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback IPV6 for asw1-604 - pt1979@cumin2003" * 13:11 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 20s) * 13:11 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 13:10 esanders@deploy1003: Finished scap sync-world: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] (duration: 08m 11s) * 13:08 pt1979@cumin2003: START - Cookbook sre.dns.netbox * 13:07 root@cumin1003: START - Cookbook sre.mysql.pool pool db2237: Maintenance * 13:06 esanders@deploy1003: esanders: Continuing with deployment * 13:04 esanders@deploy1003: esanders: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2153: Maintenance * 13:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2153: Maintenance * 13:02 esanders@deploy1003: Started scap sync-world: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] * 13:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2237 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95362 and previous config saved to /var/cache/conftool/dbconfig/20260728-130107-cwilliams.json * 13:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2237.codfw.wmnet with reason: Maintenance * 13:00 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2236: Maintenance * 12:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2153: Maintenance * 12:52 root@cumin1003: START - Cookbook sre.mysql.pool pool db2221: Maintenance * 12:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2153 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95358 and previous config saved to /var/cache/conftool/dbconfig/20260728-125214-cwilliams.json * 12:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2153.codfw.wmnet with reason: Maintenance * 12:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2167: Maintenance * 12:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-codfw * 12:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2011.codfw.wmnet * 12:51 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2011.codfw.wmnet * 12:49 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:49 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback for asw1-603 - pt1979@cumin2003" * 12:48 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback for asw1-603 - pt1979@cumin2003" * 12:46 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2011.codfw.wmnet * 12:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2221 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95357 and previous config saved to /var/cache/conftool/dbconfig/20260728-124601-cwilliams.json * 12:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2221.codfw.wmnet with reason: Maintenance * 12:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2218: Maintenance * 12:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2167 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95354 and previous config saved to /var/cache/conftool/dbconfig/20260728-124457-cwilliams.json * 12:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2167.codfw.wmnet with reason: Maintenance * 12:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2166: Maintenance * 12:42 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 12:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2011.codfw.wmnet * 12:41 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2010.codfw.wmnet * 12:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2010.codfw.wmnet * 12:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2010.codfw.wmnet * 12:34 pt1979@cumin2003: START - Cookbook sre.dns.netbox * 12:32 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 12:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2010.codfw.wmnet * 12:31 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2009.codfw.wmnet * 12:31 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2009.codfw.wmnet * 12:27 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2009.codfw.wmnet * 12:22 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2009.codfw.wmnet * 12:22 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2008.codfw.wmnet * 12:21 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2008.codfw.wmnet * 12:16 pt1979@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-604-eqsin * 12:16 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2008.codfw.wmnet * 12:16 pt1979@cumin1003: START - Cookbook sre.network.tls for network device asw1-604-eqsin * 12:14 root@cumin1003: START - Cookbook sre.mysql.pool pool db2236: Maintenance * 12:14 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2236: Maintenance * 12:12 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1137.eqiad.wmnet with OS trixie * 12:11 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2008.codfw.wmnet * 12:11 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2007.codfw.wmnet * 12:11 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2007.codfw.wmnet * 12:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db2236: Maintenance * 12:06 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2007.codfw.wmnet * 12:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2236 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95348 and previous config saved to /var/cache/conftool/dbconfig/20260728-120253-cwilliams.json * 12:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2236.codfw.wmnet with reason: Maintenance * 12:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2007.codfw.wmnet * 12:01 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 12:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 11:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2218: Maintenance * 11:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db2166: Maintenance * 11:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2219: Maintenance * 11:56 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2006.codfw.wmnet * 11:52 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1137.eqiad.wmnet with reason: host reimage * 11:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2218 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95344 and previous config saved to /var/cache/conftool/dbconfig/20260728-115155-cwilliams.json * 11:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2218.codfw.wmnet with reason: Maintenance * 11:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2208: Maintenance * 11:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2166 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95342 and previous config saved to /var/cache/conftool/dbconfig/20260728-115119-cwilliams.json * 11:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2166.codfw.wmnet with reason: Maintenance * 11:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2164: Maintenance * 11:47 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1137.eqiad.wmnet with reason: host reimage * 11:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2006.codfw.wmnet * 11:45 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2005.codfw.wmnet * 11:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2005.codfw.wmnet * 11:40 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2005.codfw.wmnet * 11:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2005.codfw.wmnet * 11:35 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2004.codfw.wmnet * 11:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2004.codfw.wmnet * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1137 * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1137 * 11:30 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1137 * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1137.eqiad.wmnet 192.32.64.10.in-addr.arpa 2.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:30 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1137.eqiad.wmnet 192.32.64.10.in-addr.arpa 2.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1137 - jiji@cumin1003" * 11:25 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2004.codfw.wmnet * 11:19 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2004.codfw.wmnet * 11:19 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2003.codfw.wmnet * 11:19 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2003.codfw.wmnet * 11:14 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2003.codfw.wmnet * 11:11 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2219: Maintenance * 11:10 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2219: Maintenance * 11:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2219: Maintenance * 11:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2003.codfw.wmnet * 11:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2164: Maintenance * 11:03 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2002.codfw.wmnet * 11:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2002.codfw.wmnet * 11:03 root@cumin1003: START - Cookbook sre.mysql.pool pool db2208: Maintenance * 10:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2164 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95332 and previous config saved to /var/cache/conftool/dbconfig/20260728-105749-cwilliams.json * 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2164.codfw.wmnet with reason: Maintenance * 10:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2208 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95331 and previous config saved to /var/cache/conftool/dbconfig/20260728-105711-cwilliams.json * 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2208.codfw.wmnet with reason: Maintenance * 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2219 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95330 and previous config saved to /var/cache/conftool/dbconfig/20260728-105652-cwilliams.json * 10:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2219.codfw.wmnet with reason: Maintenance * 10:53 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1137 - jiji@cumin1003" * 10:52 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2002.codfw.wmnet * 10:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2002.codfw.wmnet * 10:47 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2001.codfw.wmnet * 10:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2001.codfw.wmnet * 10:39 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2001.codfw.wmnet * 10:35 jiji@cumin1003: START - Cookbook sre.dns.netbox * 10:34 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] (duration: 09m 31s) * 10:34 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1137 * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2001.codfw.wmnet * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-codfw * 10:34 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1137.eqiad.wmnet with OS trixie * 10:28 jforrester@deploy1003: jforrester: Continuing with deployment * 10:27 jforrester@deploy1003: jforrester: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:25 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] * 10:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 10:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 10:21 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 10:20 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1137.eqiad.wmnet * 10:20 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 10:20 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1137.eqiad.wmnet * 10:20 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1137.eqiad.wmnet * 10:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2163: Maintenance * 09:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-staging-worker * 09:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2003.codfw.wmnet * 09:37 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2003.codfw.wmnet * 09:32 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 09:31 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2003.codfw.wmnet * 09:30 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 09:30 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2163: Maintenance * 09:30 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 09:30 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 09:30 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:22 klausman@cumin1003: END (ERROR) - Cookbook sre.ganeti.reboot-vm (exit_code=97) for VM ml-serve-ctrl2001.codfw.wmnet * 09:22 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2001.codfw.wmnet * 09:22 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d8-eqiad * 09:22 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d8-eqiad * 09:21 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2003.codfw.wmnet * 09:20 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2002.codfw.wmnet * 09:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2002.codfw.wmnet * 09:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f2-codfw * 09:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f2-codfw * 09:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e4-codfw * 09:18 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2163: Maintenance * 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e4-codfw * 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-codfw * 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-codfw * 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e5-codfw * 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e5-codfw * 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f4-codfw * 09:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2210: Maintenance * 09:16 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f4-codfw * 09:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2182: Maintenance * 09:14 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2002.codfw.wmnet * 09:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db2163: Maintenance * 09:11 XioNoX: rebooting cr2-drmrs - [[phab:T431749|T431749]] * 09:10 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr2-drmrs,cr2-drmrs IPv6,cr2-drmrs.mgmt with reason: router upgrade * 09:06 XioNoX: draining cr2-drmrs - [[phab:T431749|T431749]] * 09:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2163 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95320 and previous config saved to /var/cache/conftool/dbconfig/20260728-090638-cwilliams.json * 09:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2163.codfw.wmnet with reason: Maintenance * 09:06 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2161: Maintenance * 09:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2002.codfw.wmnet * 09:04 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2001.codfw.wmnet * 09:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2001.codfw.wmnet * 08:57 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2001.codfw.wmnet * 08:48 XioNoX: un-drain cr1-drmrs - [[phab:T431749|T431749]] * 08:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2001.codfw.wmnet * 08:47 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-staging-worker * 08:42 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:35 XioNoX: rebooting cr1-drmrs - [[phab:T431749|T431749]] * 08:33 XioNoX: draining cr1-drmrs - [[phab:T431749|T431749]] * 08:31 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2210: Maintenance * 08:29 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2182: Maintenance * 08:21 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2182: Maintenance * 08:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db2161: Maintenance * 08:16 root@cumin1003: START - Cookbook sre.mysql.pool pool db2182: Maintenance * 08:12 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2210: Maintenance * 08:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2161 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95309 and previous config saved to /var/cache/conftool/dbconfig/20260728-081044-cwilliams.json * 08:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2161.codfw.wmnet with reason: Maintenance * 08:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2154: Maintenance * 08:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2182 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95307 and previous config saved to /var/cache/conftool/dbconfig/20260728-080947-cwilliams.json * 08:09 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2182.codfw.wmnet with reason: Maintenance * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2168: Maintenance * 08:06 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr1-drmrs,cr1-drmrs IPv6,cr1-drmrs.mgmt with reason: router upgrade * 08:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db2210: Maintenance * 08:05 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 08:05 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 08:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2210 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95305 and previous config saved to /var/cache/conftool/dbconfig/20260728-080008-cwilliams.json * 08:00 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2210.codfw.wmnet with reason: Maintenance * 07:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2206: Maintenance * 07:50 gkyziridis@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 07:50 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 07:22 root@cumin1003: START - Cookbook sre.mysql.pool pool db2154: Maintenance * 07:22 root@cumin1003: START - Cookbook sre.mysql.pool pool db2168: Maintenance * 07:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2154 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95295 and previous config saved to /var/cache/conftool/dbconfig/20260728-071640-cwilliams.json * 07:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2154.codfw.wmnet with reason: Maintenance * 07:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95294 and previous config saved to /var/cache/conftool/dbconfig/20260728-071604-cwilliams.json * 07:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2168.codfw.wmnet with reason: Maintenance * 07:08 root@cumin1003: START - Cookbook sre.mysql.pool pool db2206: Maintenance * 07:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2206 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95292 and previous config saved to /var/cache/conftool/dbconfig/20260728-070219-cwilliams.json * 07:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2206.codfw.wmnet with reason: Maintenance * 06:44 marostegui: Failover m5 from db1164 to db1228 - [[phab:T432967|T432967]] * 06:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2235].codfw.wmnet,db[1164,1217,1228].eqiad.wmnet with reason: m5 master switch [[phab:T432967|T432967]] * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.10 (duration: 02m 34s) * 03:39 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] (duration: 36m 06s) * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 02:57 dzahn@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1004.eqiad.wmnet with OS trixie * 02:57 dzahn@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - dzahn@cumin1003" * 02:55 dzahn@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - dzahn@cumin1003" * 02:37 dzahn@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1004.eqiad.wmnet with reason: host reimage * 02:31 dzahn@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1004.eqiad.wmnet with reason: host reimage * 02:16 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie * 02:15 dzahn@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host zuul1004.eqiad.wmnet with OS trixie * 01:43 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie * 01:43 dzahn@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1004.eqiad.wmnet with OS trixie * 01:25 pt1979@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-603-eqsin * 01:24 pt1979@cumin1003: START - Cookbook sre.network.tls for network device asw1-603-eqsin * 01:12 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 01:12 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt for new switches in eqsin - pt1979@cumin2003" * 01:12 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt for new switches in eqsin - pt1979@cumin2003" * 01:08 pt1979@cumin2003: START - Cookbook sre.dns.netbox * 00:48 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 00:47 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 00:47 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 00:47 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 00:26 mutante: attempting reimage with trixie on zuul1004 re-purposed physical hardware - dcops reported install issue - host was in busybox shell ([[phab:T427353|T427353]]) * 00:24 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie == 2026-07-27 == * 23:50 Amir1: mass deleting vp8 transcodes * 23:28 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:27 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1004.eqiad.wmnet with OS bullseye * 23:26 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 23:25 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:25 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 22:39 maryum: Deploy security fix for [[phab:T432877|T432877]] * 22:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1022.eqiad.wmnet with OS bookworm * 22:37 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS bullseye * 22:32 sbassett: Deployed security fix for [[phab:T432789|T432789]] * 22:22 sbassett: Deployed security patch for [[phab:T431819|T431819]] * 22:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1022.eqiad.wmnet with reason: host reimage * 22:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1022.eqiad.wmnet with reason: host reimage * 22:01 RScout-WMF: Deployed security fix for [[phab:T431819|T431819]] * 22:00 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2012 * 21:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2012.codfw.wmnet with OS bookworm * 21:55 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2011\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 21:45 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1022 * 21:45 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1022 * 21:44 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1022 * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1022.eqiad.wmnet 239.48.64.10.in-addr.arpa 9.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:44 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1022.eqiad.wmnet 239.48.64.10.in-addr.arpa 9.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:41 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:41 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 21:34 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS bookworm * 21:31 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:22 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:21 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:19 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:17 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1004 * 21:16 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1004 * 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1004] - vriley@cumin1003" * 21:15 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1004] - vriley@cumin1003" * 21:11 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:10 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2011.codfw.wmnet, repooling source-only afterwards * 21:05 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:01 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1022 * 20:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1022.eqiad.wmnet with OS bookworm * 20:53 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1021.eqiad.wmnet, repooling source-only afterwards * 20:51 mutante: zuul1001 - re-enabled puppet - revert "cherry-picked" gerrit:1314120 - [[phab:T431003|T431003]] * 20:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Maintenance * 20:15 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] (duration: 08m 03s) * 20:11 sbisson@deploy1003: sbisson: Continuing with deployment * 20:09 sbisson@deploy1003: sbisson: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] * 19:47 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 46s) * 19:47 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Maintenance * 19:27 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] (duration: 12m 26s) * 19:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2228 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95285 and previous config saved to /var/cache/conftool/dbconfig/20260727-192711-cwilliams.json * 19:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2228.codfw.wmnet with reason: Maintenance * 19:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2223: Maintenance * 19:23 krinkle@deploy1003: krinkle: Continuing with deployment * 19:16 krinkle@deploy1003: krinkle: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:15 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] * 19:12 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2238: Maintenance * 18:58 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 18:57 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 18:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2227: Maintenance * 18:57 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-experimental: apply * 18:55 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-experimental: apply * 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1021.eqiad.wmnet with OS bookworm * 18:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2011.codfw.wmnet with OS bookworm * 18:40 root@cumin1003: START - Cookbook sre.mysql.pool pool db2223: Maintenance * 18:39 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] (duration: 07m 05s) * 18:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2223 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95275 and previous config saved to /var/cache/conftool/dbconfig/20260727-183500-cwilliams.json * 18:34 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2223.codfw.wmnet with reason: Maintenance * 18:34 musikanimal@deploy1003: musikanimal: Continuing with deployment * 18:34 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2213: Maintenance * 18:33 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:32 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] * 18:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db2238: Maintenance * 18:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2011.codfw.wmnet with reason: host reimage * 18:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2238 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95271 and previous config saved to /var/cache/conftool/dbconfig/20260727-181944-cwilliams.json * 18:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2238.codfw.wmnet with reason: Maintenance * 18:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2226: Maintenance * 18:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1021.eqiad.wmnet with reason: host reimage * 18:14 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2011.codfw.wmnet with reason: host reimage * 18:12 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1021.eqiad.wmnet with reason: host reimage * 18:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db2227: Maintenance * 18:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2227 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95265 and previous config saved to /var/cache/conftool/dbconfig/20260727-180256-cwilliams.json * 18:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2227.codfw.wmnet with reason: Maintenance * 18:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2194: Maintenance * 17:57 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2011 * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2011 * 17:56 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2011 * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2011.codfw.wmnet 37.32.192.10.in-addr.arpa 7.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:56 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2011.codfw.wmnet 37.32.192.10.in-addr.arpa 7.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2011 - bking@cumin2003" * 17:56 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2011 - bking@cumin2003" * 17:52 bking@cumin2003: START - Cookbook sre.dns.netbox * 17:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2011 * 17:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1021 * 17:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1021 * 17:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2011.codfw.wmnet with OS bookworm * 17:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1021.eqiad.wmnet with OS bookworm * 17:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Maintenance * 17:38 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2010\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 17:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2213 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95260 and previous config saved to /var/cache/conftool/dbconfig/20260727-173740-cwilliams.json * 17:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2213.codfw.wmnet with reason: Maintenance * 17:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2211: Maintenance * 17:36 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1020\.eqiad\.wmnet,dc=eqiad,cluster=wdqs\-main,service=wdqs\-main * 17:32 root@cumin1003: START - Cookbook sre.mysql.pool pool db2226: Maintenance * 17:31 taavi@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] (duration: 06m 33s) * 17:27 taavi@deploy1003: taavi: Continuing with deployment * 17:27 taavi@deploy1003: taavi: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:26 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2226 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95256 and previous config saved to /var/cache/conftool/dbconfig/20260727-172636-cwilliams.json * 17:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2226.codfw.wmnet with reason: Maintenance * 17:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2225: Maintenance * 17:25 taavi@deploy1003: Started scap sync-world: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] * 17:13 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 17:11 root@cumin1003: START - Cookbook sre.mysql.pool pool db2194: Maintenance * 17:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2194 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95248 and previous config saved to /var/cache/conftool/dbconfig/20260727-170453-cwilliams.json * 17:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2194.codfw.wmnet with reason: Maintenance * 17:04 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2190: Maintenance * 16:52 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 16:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2211: Maintenance * 16:40 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95242 and previous config saved to /var/cache/conftool/dbconfig/20260727-164015-cwilliams.json * 16:40 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2211.codfw.wmnet with reason: Maintenance * 16:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2178: Maintenance * 16:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db2225: Maintenance * 16:39 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2172: Maintenance * 16:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 16:38 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 16:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2225 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95238 and previous config saved to /var/cache/conftool/dbconfig/20260727-163307-cwilliams.json * 16:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2225.codfw.wmnet with reason: Maintenance * 16:32 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2189: Maintenance * 16:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db2190: Maintenance * 16:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2190 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95230 and previous config saved to /var/cache/conftool/dbconfig/20260727-160602-cwilliams.json * 16:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2190.codfw.wmnet with reason: Maintenance * 15:53 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2177: Maintenance * 15:53 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2172: Maintenance * 15:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2178: Maintenance * 15:51 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2172: Maintenance * 15:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2178 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95224 and previous config saved to /var/cache/conftool/dbconfig/20260727-154559-cwilliams.json * 15:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2172: Maintenance * 15:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2178.codfw.wmnet with reason: Maintenance * 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2171: Maintenance * 15:44 root@cumin1003: START - Cookbook sre.mysql.pool pool db2189: Maintenance * 15:43 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:41 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2172 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95222 and previous config saved to /var/cache/conftool/dbconfig/20260727-153927-cwilliams.json * 15:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2172.codfw.wmnet with reason: Maintenance * 15:38 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2189 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95220 and previous config saved to /var/cache/conftool/dbconfig/20260727-153833-cwilliams.json * 15:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2189.codfw.wmnet with reason: Maintenance * 15:34 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:32 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 15:32 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 15:31 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:29 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:26 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 15:22 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] (duration: 07m 00s) * 15:21 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2155: Maintenance * 15:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2175: Maintenance * 15:18 zabe@deploy1003: zabe: Continuing with deployment * 15:17 zabe@deploy1003: zabe: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:15 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 15:15 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:15 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] * 15:15 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 06s) * 15:15 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:12 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2010.codfw.wmnet with OS bookworm * 15:08 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2177: Maintenance * 15:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2177: Maintenance * 14:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2177: Maintenance * 14:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2171: Maintenance * 14:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2171 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95209 and previous config saved to /var/cache/conftool/dbconfig/20260727-145236-cwilliams.json * 14:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2171.codfw.wmnet with reason: Maintenance * 14:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2177 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95208 and previous config saved to /var/cache/conftool/dbconfig/20260727-145206-cwilliams.json * 14:52 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2157: Maintenance * 14:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2177.codfw.wmnet with reason: Maintenance * 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1020.eqiad.wmnet with OS bookworm * 14:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2156: Maintenance * 14:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2010.codfw.wmnet with reason: host reimage * 14:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2010.codfw.wmnet with reason: host reimage * 14:41 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[1031,2024]*: Upgrade Cassandra to 5.0.8 (canary) - eevans@cumin1003 * 14:34 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2155: Maintenance * 14:33 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2175: Maintenance * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2010 * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2010 * 14:24 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2010 * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2010.codfw.wmnet 94.16.192.10.in-addr.arpa 4.9.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2010.codfw.wmnet 94.16.192.10.in-addr.arpa 4.9.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2010 - bking@cumin2003" * 14:24 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2010 - bking@cumin2003" * 14:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1020.eqiad.wmnet with reason: host reimage * 14:23 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[1031,2024]*: Upgrade Cassandra to 5.0.8 (canary) - eevans@cumin1003 * 14:20 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 14:20 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 14:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1020.eqiad.wmnet with reason: host reimage * 14:17 sukhe: sudo gnt-instance reboot urldownloader1005.wikimedia.org * 14:16 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:15 jelto@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:08 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2155: Maintenance * 14:05 root@cumin1003: START - Cookbook sre.mysql.pool pool db2157: Maintenance * 14:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2175: Maintenance * 14:03 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 11 hosts * 14:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db2155: Maintenance * 14:01 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 11 hosts * 14:01 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1136.eqiad.wmnet * 14:01 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1136.eqiad.wmnet * 14:01 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1136.eqiad.wmnet * 14:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db2156: Maintenance * 13:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2157 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95194 and previous config saved to /var/cache/conftool/dbconfig/20260727-135943-cwilliams.json * 13:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2157.codfw.wmnet with reason: Maintenance * 13:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db2175: Maintenance * 13:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 13:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 13:57 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2010 * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2155 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95193 and previous config saved to /var/cache/conftool/dbconfig/20260727-135613-cwilliams.json * 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2155.codfw.wmnet with reason: Maintenance * 13:55 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1020 * 13:55 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1020 * 13:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2156 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95192 and previous config saved to /var/cache/conftool/dbconfig/20260727-135413-cwilliams.json * 13:54 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2156.codfw.wmnet with reason: Maintenance * 13:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2175 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95191 and previous config saved to /var/cache/conftool/dbconfig/20260727-135300-cwilliams.json * 13:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2010.codfw.wmnet with OS bookworm * 13:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2175.codfw.wmnet with reason: Maintenance * 13:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1020.eqiad.wmnet with OS bookworm * 13:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 34 hosts * 13:46 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 34 hosts * 13:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1201: Maintenance * 13:27 Lucas_WMDE: UTC afternoon backport+config window doen * 13:18 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] (duration: 11m 57s) * 13:14 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, sihe: Continuing with deployment * 13:08 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, sihe: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:07 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool ulsfo [reason: router upgrade finished, [[phab:T431752|T431752]]] * 13:07 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool ulsfo [reason: router upgrade finished, [[phab:T431752|T431752]]] * 13:06 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] * 13:03 XioNoX: repool cr4-ulsfo - [[phab:T431752|T431752]] * 12:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db1201: Maintenance * 12:48 gkyziridis@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1201 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95186 and previous config saved to /var/cache/conftool/dbconfig/20260727-124404-cwilliams.json * 12:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1201.eqiad.wmnet with reason: Maintenance * 12:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1187: Maintenance * 12:30 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader.eqiad.wikimedia.org on all recursors * 12:30 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader.eqiad.wikimedia.org on all recursors * 12:30 sukhe@dns1004: END - running authdns-update * 12:30 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] (duration: 09m 32s) * 12:28 sukhe@dns1004: START - running authdns-update * 12:25 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 12:22 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:20 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] * 12:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts * 12:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts * 12:13 XioNoX: rebooting cr4-ulsfo for upgrade - [[phab:T431752|T431752]] * 12:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: es1038 repool * 12:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 38 hosts * 12:08 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 38 hosts * 11:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1187: Maintenance * 11:53 urbanecm@deploy1003: mwscript-k8s job started: foreachwikiindblist growthexperiments GrowthExperiments:cleanMentorList # [[phab:T431804|T431804]] * 11:50 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr4-ulsfo,cr4-ulsfo IPv6,cr4-ulsfo.mgmt with reason: router upgrade * 11:50 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] (duration: 11m 07s) * 11:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1187 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95178 and previous config saved to /var/cache/conftool/dbconfig/20260727-114844-cwilliams.json * 11:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1187.eqiad.wmnet with reason: Maintenance * 11:43 urbanecm@deploy1003: urbanecm: Continuing with deployment * 11:42 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:39 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] * 11:37 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 11:36 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 11:36 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 11:35 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 11:29 XioNoX: start draining cr4-ulsfo - [[phab:T431752|T431752]] * 11:29 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 11:29 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 11:28 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1035: testing * 11:28 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1035: testing * 11:27 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1035: testing * 11:27 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1035: testing * 11:26 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool ulsfo [reason: router upgrade, [[phab:T431752|T431752]]] * 11:26 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1038: es1038 repool * 11:26 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool ulsfo [reason: router upgrade, [[phab:T431752|T431752]]] * 11:26 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1038: testing * 11:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1264: Maintenance * 11:24 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1038: testing * 11:23 marostegui@cumin1003: dbctl commit (dc=all): 'Repool es1050 as master', diff saved to https://phabricator.wikimedia.org/P95170 and previous config saved to /var/cache/conftool/dbconfig/20260727-112326-marostegui.json * 11:23 marostegui@cumin1003: dbctl commit (dc=all): 'Repool es1050', diff saved to https://phabricator.wikimedia.org/P95169 and previous config saved to /var/cache/conftool/dbconfig/20260727-112302-marostegui.json * 11:22 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1050: testing * 11:22 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1050: testing * 11:20 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 11:18 blake@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 11:18 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 11:12 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 11:11 blake@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 11:09 blake@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 11:09 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 11:09 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 11:08 blake@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 11:05 blake@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 11:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 11:02 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 10:50 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 10:43 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:39 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply * 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1264: Maintenance * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply * 10:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply * 10:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 10:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 10:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 10:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1264 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95164 and previous config saved to /var/cache/conftool/dbconfig/20260727-103204-cwilliams.json * 10:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1264.eqiad.wmnet with reason: Maintenance * 10:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 10:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 10:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 10:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 10:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 10:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 10:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 10:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1237: Maintenance * 10:04 elukey: restart burrow main-eqiad on kafkamon2003 to clear some errors on kafka-main1008 * 09:58 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1136.eqiad.wmnet with OS trixie * 09:39 elukey: restart burrow-main-eqiad.service on kafkamon1003 to see if a recurrent kafka error on kafka-main1008 goes away * 09:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1237: Maintenance * 09:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1136.eqiad.wmnet with reason: host reimage * 09:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1237 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95159 and previous config saved to /var/cache/conftool/dbconfig/20260727-093328-cwilliams.json * 09:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1237.eqiad.wmnet with reason: Maintenance * 09:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1136.eqiad.wmnet with reason: host reimage * 09:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1203: Maintenance * 09:17 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1136 * 09:17 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1136 * 09:04 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1136 * 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1136.eqiad.wmnet 191.32.64.10.in-addr.arpa 1.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:04 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1136.eqiad.wmnet 191.32.64.10.in-addr.arpa 1.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1136 - jiji@cumin1003" * 09:04 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1136 - jiji@cumin1003" * 08:52 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 08:52 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 08:52 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 08:51 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 08:50 jiji@cumin1003: START - Cookbook sre.dns.netbox * 08:47 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1136 * 08:46 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1136.eqiad.wmnet with OS trixie * 08:44 marostegui: Rename tables on s3 [[phab:T425066|T425066]] * 08:43 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1136.eqiad.wmnet * 08:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db1203: Maintenance * 08:43 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1136.eqiad.wmnet * 08:43 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1136.eqiad.wmnet * 08:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1203 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95154 and previous config saved to /var/cache/conftool/dbconfig/20260727-083703-cwilliams.json * 08:36 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1203.eqiad.wmnet with reason: Maintenance * 08:16 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1179: Maintenance * 07:44 phuedx: UTC morning backport window done * 07:37 phuedx@deploy1003: Finished scap sync-world: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] (duration: 32m 33s) * 07:28 root@cumin1003: START - Cookbook sre.mysql.pool pool db1179: Maintenance * 07:26 marostegui: Rename tables on s3 [[phab:T426341|T426341]] * 07:25 phuedx@deploy1003: phuedx: Continuing with deployment * 07:22 marostegui: Drop tables in akwiki nawiki pihwiki - growthexperiments_* [[phab:T428885|T428885]] * 07:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95149 and previous config saved to /var/cache/conftool/dbconfig/20260727-072234-cwilliams.json * 07:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1179.eqiad.wmnet with reason: Maintenance * 07:20 phuedx@deploy1003: phuedx: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:16 ryankemper: [[phab:T430880|T430880]] [WDQS] Reimaged `wdqs1018` and `wdqs1019` to Bookworm, restored data using test-cookbook change {{Gerrit|1317128}}, and repooled both; 25/36 hosts complete * 07:04 phuedx@deploy1003: Started scap sync-world: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] * 06:57 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1019.eqiad.wmnet * 06:56 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1018.eqiad.wmnet * 06:40 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1020.eqiad.wmnet with reason: Cloning * 06:35 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db1228.eqiad.wmnet with reason: Rebooting * 06:29 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:29 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:25 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:25 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:25 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1019.eqiad.wmnet, repooling source-only afterwards * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1018.eqiad.wmnet, repooling source-only afterwards * 04:51 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1019.eqiad.wmnet, repooling source-only afterwards * 04:51 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1018.eqiad.wmnet, repooling source-only afterwards * 04:48 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s) * 04:48 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 04:48 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s) * 04:48 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 36s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-26 == * 14:59 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:59 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:59 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:59 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1019.eqiad.wmnet with OS bookworm * 01:05 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1018.eqiad.wmnet with OS bookworm * 00:43 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1019.eqiad.wmnet with reason: host reimage * 00:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1018.eqiad.wmnet with reason: host reimage * 00:34 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1019.eqiad.wmnet with reason: host reimage * 00:33 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1018.eqiad.wmnet with reason: host reimage * 00:16 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 00:16 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 00:15 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 00:15 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1019 * 00:11 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1019 * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1018 * 00:11 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1018 * 00:08 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1019.eqiad.wmnet with OS bookworm * 00:08 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1018.eqiad.wmnet with OS bookworm == 2026-07-25 == * 22:06 ryankemper: [[phab:T430880|T430880]] [WDQS] Repooled `wdqs1017` and `wdqs2024` after reimaging to bookworm, scap deploying, and data xfering * 22:04 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2024.codfw.wmnet * 22:03 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1017.eqiad.wmnet * 21:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1017.eqiad.wmnet, repooling source-only afterwards * 21:06 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2024.codfw.wmnet, repooling source-only afterwards * 20:52 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:52 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:52 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:52 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 20:18 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1017.eqiad.wmnet, repooling source-only afterwards * 20:18 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2024.codfw.wmnet, repooling source-only afterwards * 20:15 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:15 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:15 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:15 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 19:57 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s) * 19:57 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 19:57 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 07s) * 19:57 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 19:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2024.codfw.wmnet with OS bookworm * 19:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1017.eqiad.wmnet with OS bookworm * 19:02 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2024.codfw.wmnet with reason: host reimage * 18:58 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1017.eqiad.wmnet with reason: host reimage * 18:53 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2024.codfw.wmnet with reason: host reimage * 18:52 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1017.eqiad.wmnet with reason: host reimage * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2024 * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2024 * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1017 * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1017 * 18:27 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2024 * 18:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2024.codfw.wmnet 58.16.192.10.in-addr.arpa 8.5.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:26 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2024.codfw.wmnet 58.16.192.10.in-addr.arpa 8.5.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:24 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1017 * 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1017.eqiad.wmnet 238.48.64.10.in-addr.arpa 8.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:24 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1017.eqiad.wmnet 238.48.64.10.in-addr.arpa 8.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1017 - ryankemper@cumin2003" * 18:24 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1017 - ryankemper@cumin2003" * 18:23 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 18:18 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 18:17 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1017 * 18:17 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2024 * 18:14 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1017.eqiad.wmnet with OS bookworm * 18:14 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2024.codfw.wmnet with OS bookworm * 18:05 ryankemper: [WDQS] [[phab:T430880|T430880]] Reimaged `wdqs1016` and `wdqs2023` to Bookworm with `--move-vlan`, restored main and scholarly data, validated postflights, and repooled both hosts. Confirmed PyBal rebuilt both backends with their new addresses * 17:45 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2023.codfw.wmnet * 17:43 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1016.eqiad.wmnet * 06:35 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1016.eqiad.wmnet, repooling source-only afterwards * 06:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2023.codfw.wmnet, repooling source-only afterwards * 05:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2023.codfw.wmnet, repooling source-only afterwards * 05:19 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1016.eqiad.wmnet, repooling source-only afterwards * 05:07 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 07s) * 05:07 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 05:06 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 06s) * 05:06 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 03:27 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2023.codfw.wmnet with OS bookworm * 02:59 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2023.codfw.wmnet with reason: host reimage * 02:56 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2023.codfw.wmnet with reason: host reimage * 02:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2023 * 02:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2023 * 02:30 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2023 * 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2023.codfw.wmnet 35.0.192.10.in-addr.arpa 5.3.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 02:30 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2023.codfw.wmnet 35.0.192.10.in-addr.arpa 5.3.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2023 - ryankemper@cumin2003" * 02:30 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2023 - ryankemper@cumin2003" * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 26s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:15 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1016.eqiad.wmnet with OS bookworm * 00:49 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1016.eqiad.wmnet with reason: host reimage * 00:43 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1016.eqiad.wmnet with reason: host reimage * 00:31 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 00:27 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1016 * 00:27 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1016 * 00:27 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2023 * 00:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1016.eqiad.wmnet with OS bookworm * 00:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2023.codfw.wmnet with OS bookworm * 00:11 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs1014.eqiad.wmnet and wdqs2008.codfw.wmnet after Bookworm reimage, transfer, and postflight; wdqs2008 is serving, while wdqs1014 will remain outside of service until a pybal restart next monday == 2026-07-24 == * 23:54 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1014.eqiad.wmnet * 23:54 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2008.codfw.wmnet * 23:43 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2010.codfw.wmnet with OS trixie * 23:08 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 23:03 jhathaway@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 22:33 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 22:13 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 22:13 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 22:13 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 22:13 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:00 jhathaway@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 21:53 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 21:53 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie * 21:51 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 21:47 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie * 21:43 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 21:39 jhathaway@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 21:38 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 17:21 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1135.eqiad.wmnet * 17:21 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1135.eqiad.wmnet * 17:21 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1135.eqiad.wmnet * 16:34 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 16:34 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 16:34 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 16:34 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 16:33 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 16:33 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 16:28 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:28 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:28 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:28 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2008.codfw.wmnet, repooling source-only afterwards * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1014.eqiad.wmnet, repooling source-only afterwards * 15:56 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1135.eqiad.wmnet with OS trixie * 15:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 40 hosts * 15:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 40 hosts * 15:37 topranks: upgrade SR-Linux OS on lswtest-d8-eqiad * 15:36 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1135.eqiad.wmnet with reason: host reimage * 15:33 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 6 hosts with reason: upgrade lswtest-d8-eqiad * 15:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1135.eqiad.wmnet with reason: host reimage * 15:30 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc-gp2006.codfw.wmnet with OS bookworm * 15:15 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1135 * 15:15 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1135 * 15:13 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc-gp2006.codfw.wmnet with reason: host reimage * 15:08 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc-gp2006.codfw.wmnet with reason: host reimage * 14:49 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm * 14:48 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host mc-gp2006.codfw.wmnet with OS bookworm * 14:34 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] (duration: 41m 12s) * 14:32 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1135 * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1135.eqiad.wmnet 177.32.64.10.in-addr.arpa 7.7.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:32 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1135.eqiad.wmnet 177.32.64.10.in-addr.arpa 7.7.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1135 - jiji@cumin1003" * 14:32 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1135 - jiji@cumin1003" * 14:29 krinkle@deploy1003: krinkle: Continuing with deployment * 14:29 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm * 14:27 jiji@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host mc-gp2006.codfw.wmnet with OS bookworm * 14:26 jiji@cumin1003: START - Cookbook sre.dns.netbox * 14:15 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1135 * 14:14 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1135.eqiad.wmnet with OS trixie * 14:14 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1135.eqiad.wmnet * 14:13 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1135.eqiad.wmnet * 14:13 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1135.eqiad.wmnet * 14:10 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1072.eqiad.wmnet * 14:10 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1072.eqiad.wmnet * 14:10 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1072.eqiad.wmnet * 14:10 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1072.eqiad.wmnet * 14:09 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1071.eqiad.wmnet * 14:09 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1071.eqiad.wmnet * 14:09 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1071.eqiad.wmnet * 14:09 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1071.eqiad.wmnet * 13:58 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 13:58 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:58 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:57 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:55 krinkle@deploy1003: krinkle: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:53 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] * 13:45 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:45 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push new IPs for mc-gp2006 - cmooney@cumin1003" * 13:45 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push new IPs for mc-gp2006 - cmooney@cumin1003" * 13:44 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) mc-gp2006.codfw.wmnet on all recursors * 13:44 cmooney@cumin1003: START - Cookbook sre.dns.wipe-cache mc-gp2006.codfw.wmnet on all recursors * 13:42 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm * 13:41 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:30 papaul: reboot mr1-eqsin for maintenance * 13:24 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb[1029-1031].eqiad.wmnet * 13:10 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb[1029-1031].eqiad.wmnet * 11:33 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 7 hosts * 11:11 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 7 hosts * 10:56 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 7 hosts * 10:47 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 7 hosts * 10:44 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:44 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:41 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:41 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:35 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 8 hosts * 10:34 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:33 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:32 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:32 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:31 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:31 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:30 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 8 hosts * 10:24 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie * 10:19 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:18 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 16 hosts * 10:17 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2001.codfw.wmnet * 10:13 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2001.codfw.wmnet * 10:12 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2001.codfw.wmnet * 10:02 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2001.codfw.wmnet * 10:02 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2002.codfw.wmnet * 09:57 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2002.codfw.wmnet * 09:56 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2002.codfw.wmnet * 09:51 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2002.codfw.wmnet * 09:51 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1002.eqiad.wmnet * 09:47 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1002.eqiad.wmnet * 09:47 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1001.eqiad.wmnet * 09:44 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1001.eqiad.wmnet * 09:34 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2003.codfw.wmnet * 09:32 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2003.codfw.wmnet * 09:32 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2002.codfw.wmnet * 09:30 brouberol@dns1004: END - running authdns-update * 09:29 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2002.codfw.wmnet * 09:29 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2001.codfw.wmnet * 09:27 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 9 hosts * 09:27 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2001.codfw.wmnet * 09:27 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2001.codfw.wmnet * 09:26 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 9 hosts * 09:26 brouberol@dns1004: START - running authdns-update * 09:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 57 hosts * 09:24 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2001.codfw.wmnet * 09:24 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2002.codfw.wmnet * 09:22 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2002.codfw.wmnet * 09:21 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 57 hosts * 09:20 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2003.codfw.wmnet * 09:19 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 16 hosts * 09:16 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2003.codfw.wmnet * 09:16 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1003.eqiad.wmnet * 09:15 urbanecm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 09:15 urbanecm@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 09:13 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1003.eqiad.wmnet * 09:13 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1002.eqiad.wmnet * 09:11 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1002.eqiad.wmnet * 09:11 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1001.eqiad.wmnet * 09:07 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1001.eqiad.wmnet * 08:32 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:24 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:16 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 08:16 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 08:07 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:07 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:07 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 08:02 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 08:01 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:59 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:57 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 07:57 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 06:46 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1025.eqiad.wmnet with reason: Cloning * 06:46 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s4 * 06:45 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s6 * 06:44 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1019.eqiad.wmnet,service=s6 * 06:44 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1019.eqiad.wmnet,service=s4 * 03:40 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:40 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:40 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:40 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 03:37 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:37 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:37 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:36 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:49 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on mr1-eqsin,mr1-eqsin IPv6 with reason: connection issue * 02:38 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on cr[2-3]-eqsin.mgmt,ps1-[603-604]-eqsin with reason: connection issue * 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 27s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-23 == * 23:27 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin.oob,mr1-eqsin.oob IPv6 with reason: switch refresh * 22:21 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Setting storage compatibility to NONE - eevans@cumin1003 * 22:01 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Setting storage compatibility to NONE - eevans@cumin1003 * 21:29 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1014.eqiad.wmnet, repooling source-only afterwards * 21:28 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 46s) * 21:28 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 21:19 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Setting storage compatibility to UPGRADING - eevans@cumin1003 * 21:00 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Setting storage compatibility to UPGRADING - eevans@cumin1003 * 20:17 dani@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] (duration: 11m 57s) * 20:13 dani@deploy1003: dani: Continuing with deployment * 20:07 dani@deploy1003: dani: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:05 dani@deploy1003: Started scap sync-world: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] * 19:24 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:24 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating the rest of the ipv6 dns records. - jhancock@cumin2002" * 19:24 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating the rest of the ipv6 dns records. - jhancock@cumin2002" * 19:14 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 19:05 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wdqs1014.eqiad.wmnet with OS bookworm * 19:04 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.noop (exit_code=99) * 19:04 cwilliams@cumin1003: START - Cookbook sre.mysql.noop * 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1014.eqiad.wmnet with reason: host reimage * 18:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1014.eqiad.wmnet with reason: host reimage * 18:30 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2008.codfw.wmnet, repooling source-only afterwards * 18:28 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 19s) * 18:28 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1014 * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1014 * 18:22 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1014 * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1014.eqiad.wmnet 188.32.64.10.in-addr.arpa 8.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:22 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1014.eqiad.wmnet 188.32.64.10.in-addr.arpa 8.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1014 - bking@cumin2003" * 18:21 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1014 - bking@cumin2003" * 18:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2215: Maintenance * 18:18 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 18:15 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:15 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 18:06 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:06 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 18:05 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 18:04 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2052: codfw rack B8 re-pool after maintenance * 17:54 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 17:54 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:54 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 17:32 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2215: Maintenance * 17:29 cmooney@dns3003: END - running authdns-update * 17:27 cmooney@dns3003: START - running authdns-update * 17:23 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 17:22 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:18 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool es2052: codfw rack B8 re-pool after maintenance * 17:18 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2189: codfw rack B8 re-pool after maintenance * 17:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2215.codfw.wmnet with reason: Maintenance * 17:17 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 17:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2215 [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95126 and previous config saved to /var/cache/conftool/dbconfig/20260723-170903-cwilliams.json * 17:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2191 to x1 primary [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95125 and previous config saved to /var/cache/conftool/dbconfig/20260723-170612-cwilliams.json * 17:05 cezmunsta: Starting x1 codfw failover from db2215 to db2191 - [[phab:T432986|T432986]] * 16:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2191 with weight 0 [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95123 and previous config saved to /var/cache/conftool/dbconfig/20260723-165831-cwilliams.json * 16:58 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 16 hosts with reason: Primary switchover x1 [[phab:T432986|T432986]] * 16:36 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 138128 * 16:35 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 138128 * 16:33 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2189: codfw rack B8 re-pool after maintenance * 16:33 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2164: codfw rack B8 re-pool after maintenance * 16:28 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1072.eqiad.wmnet * 16:27 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1072.eqiad.wmnet with OS trixie * 16:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2249: Maintenance * 16:06 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker1072.eqiad.wmnet with reason: host reimage * 16:06 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1072.eqiad.wmnet with reason: host reimage * 15:50 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1072 * 15:50 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1072 * 15:49 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1072.eqiad.wmnet with OS trixie * 15:48 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2164: codfw rack B8 re-pool after maintenance * 15:48 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] (duration: 06m 37s) * 15:48 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2163: codfw rack B8 re-pool after maintenance * 15:45 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 15:43 musikanimal@deploy1003: musikanimal: Continuing with deployment * 15:43 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:41 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] * 15:36 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1072.eqiad.wmnet * 15:35 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1072.eqiad.wmnet * 15:35 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1072.eqiad.wmnet * 15:34 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:34 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push any outstanding updates - cmooney@cumin1003" * 15:34 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push any outstanding updates - cmooney@cumin1003" * 15:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db2249: Maintenance * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 15:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:26 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:21 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:21 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 15:21 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:21 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 15:20 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 15:19 cmooney@dns2004: END - running authdns-update * 15:17 cmooney@dns2004: START - running authdns-update * 15:14 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns2004.wikimedia.org * 15:12 brouberol@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 15:12 brouberol@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 15:12 klausman@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ml-serve1001.eqiad.wmnet with OS trixie * 15:11 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1071.eqiad.wmnet * 15:11 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1071.eqiad.wmnet with OS trixie * 15:10 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wdqs2008.codfw.wmnet with OS bookworm * 15:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2249.codfw.wmnet with reason: Maintenance * 15:08 brouberol@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 15:08 brouberol@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 15:08 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 15:08 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 15:06 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns1004.wikimedia.org * 15:02 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2002.codfw.wmnet * 15:02 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2002.codfw.wmnet * 15:02 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2163: codfw rack B8 re-pool after maintenance * 15:01 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 15:01 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 14:59 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test2001.codfw.wmnet * 14:57 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test2001.codfw.wmnet * 14:56 ryankemper: [WDQS] [[phab:T430880|T430880]] Reimaged `wdqs2016` to Bookworm, xferred scholarly_articles from `wdqs2024`, validated updater/readiness/federation, and repooled * 14:51 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2016.codfw.wmnet * 14:51 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1001.eqiad.wmnet with reason: host reimage * 14:48 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1071.eqiad.wmnet with reason: host reimage * 14:47 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1001.eqiad.wmnet with reason: host reimage * 14:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2008.codfw.wmnet with reason: host reimage * 14:43 topranks: reboot lsw1-b8-codw to upgrade JunOS [[phab:T430929|T430929]] * 14:41 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1071.eqiad.wmnet with reason: host reimage * 14:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2008.codfw.wmnet with reason: host reimage * 14:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2231: Maintenance * 14:30 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1001.eqiad.wmnet with OS trixie * 14:25 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2002.codfw.wmnet * 14:23 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1071 * 14:23 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1071 * 14:23 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 14:22 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore scholarly data after Bookworm reimage) xfer scholarly_articles from wdqs2024.codfw.wmnet -> wdqs2016.codfw.wmnet, repooling source-only afterwards * 14:22 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2052: codfw rack B8 depool for maintenance * 14:21 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool es2052: codfw rack B8 depool for maintenance * 14:21 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2249: codfw rack B8 depool for maintenance * 14:21 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1071 * 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1071.eqiad.wmnet 166.48.64.10.in-addr.arpa 6.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:21 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1071.eqiad.wmnet 166.48.64.10.in-addr.arpa 6.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1071 - jiji@cumin1003" * 14:21 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1071 - jiji@cumin1003" * 14:21 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2249: codfw rack B8 depool for maintenance * 14:21 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2189: codfw rack B8 depool for maintenance * 14:20 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2002.codfw.wmnet * 14:20 cmooney@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2050.codfw.wmnet * 14:20 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2189: codfw rack B8 depool for maintenance * 14:20 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2164: codfw rack B8 depool for maintenance * 14:20 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2164: codfw rack B8 depool for maintenance * 14:19 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2163: codfw rack B8 depool for maintenance * 14:19 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:19 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1014 * 14:19 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2163: codfw rack B8 depool for maintenance * 14:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1014.eqiad.wmnet with OS bookworm * 14:17 cmooney@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2050.codfw.wmnet * 14:16 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 14:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2008 * 14:14 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2008 * 14:14 cmooney@cumin1003: conftool action : set/pooled=no; selector: name=dns2004.wikimedia.org * 14:14 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2008.codfw.wmnet with OS bookworm * 14:13 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 14:12 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:10 topranks: depool dns2004 before lsw1-b8-codfw switch maintenance [[phab:T430929|T430929]] * 14:10 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b8-codfw,lsw1-b8-codfw IPv6,lsw1-b8-codfw.mgmt,ssw1-a[1,8]-codfw with reason: lsw1-b8-codfw JunOS upgrade * 14:07 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 30 hosts with reason: lsw1-b8-codfw JunOS upgrade * 14:06 elukey: upload python3-docker-report 0.0.19 to apt.wikimedia.org for bookworm and trixie * 13:59 jiji@cumin1003: START - Cookbook sre.dns.netbox * 13:58 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1071 * 13:57 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:57 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:57 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:57 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1071.eqiad.wmnet with OS trixie * 13:55 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1071.eqiad.wmnet * 13:55 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1071.eqiad.wmnet * 13:55 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1071.eqiad.wmnet * 13:53 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:52 logmsgbot: kharlan Deployed security patch for [[phab:T432948|T432948]] * 13:51 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:51 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 13:51 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db2231: Maintenance * 13:50 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:50 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:50 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:50 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:49 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:49 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:49 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2231 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95097 and previous config saved to /var/cache/conftool/dbconfig/20260723-134436-cwilliams.json * 13:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2231.codfw.wmnet with reason: Maintenance * 13:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:39 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:38 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 13:38 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] (duration: 09m 07s) * 13:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1037 hosts * 13:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2196: Maintenance * 13:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:33 kharlan@deploy1003: kharlan, emc-wmf: Continuing with deployment * 13:31 kharlan@deploy1003: kharlan, emc-wmf: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:30 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:28 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] * 13:17 hashar@deploy1003: Finished deploy [integration/docroot@2199146]: build: License GPL2.0+ / updating npm dependencies (duration: 00m 14s) * 13:17 hashar@deploy1003: Started deploy [integration/docroot@2199146]: build: License GPL2.0+ / updating npm dependencies * 13:14 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service * 13:07 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 12:58 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2207: Repooling * 12:49 root@cumin1003: START - Cookbook sre.mysql.pool pool db2196: Maintenance * 12:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2196 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95087 and previous config saved to /var/cache/conftool/dbconfig/20260723-123952-cwilliams.json * 12:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2196.codfw.wmnet with reason: Maintenance * 12:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2191: Maintenance * 12:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:13 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:13 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: Repooling * 12:12 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2207: Repooling * 12:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: Repooling * 11:56 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2235.codfw.wmnet with OS trixie * 11:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db2191: Maintenance * 11:46 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1070.eqiad.wmnet * 11:46 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1070.eqiad.wmnet * 11:46 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1070.eqiad.wmnet * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2191 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95080 and previous config saved to /var/cache/conftool/dbconfig/20260723-114308-cwilliams.json * 11:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2191.codfw.wmnet with reason: Maintenance * 11:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2186: Maintenance * 11:35 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 46375 * 11:34 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 46375 * 11:33 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2235.codfw.wmnet with reason: host reimage * 11:28 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2235.codfw.wmnet with reason: host reimage * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c7-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c7-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c6-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c6-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c5-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c5-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c4-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c4-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c3-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c3-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c2-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c2-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d7-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d7-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d4-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d4-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d3-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d2-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d2-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d8-eqiad * 11:23 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d8-eqiad * 11:23 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d1-eqiad * 11:23 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d1-eqiad * 11:12 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db2235.codfw.wmnet with OS trixie * 11:11 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:11 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[2160,2235].codfw.wmnet with reason: Upgrading * 10:56 root@cumin1003: START - Cookbook sre.mysql.pool pool db2186: Maintenance * 10:54 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1037: testing * 10:53 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1037: testing * 10:53 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1037: testing * 10:53 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1037: testing * 10:52 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: testing * 10:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2186 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95072 and previous config saved to /var/cache/conftool/dbconfig/20260723-104956-cwilliams.json * 10:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2186.codfw.wmnet with reason: Maintenance * 10:43 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1054: testing * 10:41 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1070.eqiad.wmnet with OS trixie * 10:30 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1037 hosts * 10:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 10:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 10:20 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1070.eqiad.wmnet with reason: host reimage * 10:16 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1070.eqiad.wmnet with reason: host reimage * 10:06 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1038: testing * 10:05 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1038: testing * 10:05 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1038: testing * 10:02 marostegui@dns1004: END - running authdns-update * 10:00 marostegui@dns1004: START - running authdns-update * 09:58 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: testing * 09:57 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1054: testing * 09:57 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1070 * 09:57 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1070 * 09:57 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1054: testing * 09:57 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1054: testing * 09:56 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1055: testing * 09:56 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1070 * 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1070.eqiad.wmnet 165.48.64.10.in-addr.arpa 5.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:56 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1070.eqiad.wmnet 165.48.64.10.in-addr.arpa 5.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1070 - jiji@cumin1003" * 09:56 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1070 - jiji@cumin1003" * 09:47 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for 1035 hosts * 09:45 jiji@cumin1003: START - Cookbook sre.dns.netbox * 09:42 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1070 * 09:42 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1070.eqiad.wmnet with OS trixie * 09:42 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1070.eqiad.wmnet * 09:41 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1070.eqiad.wmnet * 09:41 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1070.eqiad.wmnet * 09:27 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es2051: testing * 09:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: testing * 09:12 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es2051: testing * 09:11 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1055: testing * 09:09 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:09 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1055: testing * 09:09 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1055: testing * 08:50 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:50 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:50 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 08:49 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 08:49 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 08:49 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:46 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1069.eqiad.wmnet * 08:46 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1069.eqiad.wmnet * 08:46 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1069.eqiad.wmnet * 08:39 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 08:38 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 08:38 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 08:37 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 08:35 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:10 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1069.eqiad.wmnet with OS trixie * 07:49 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1069.eqiad.wmnet with reason: host reimage * 07:45 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1069.eqiad.wmnet with reason: host reimage * 07:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1035 hosts * 07:33 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 07:32 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 07:29 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1069 * 07:29 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1069 * 07:26 jiji@deploy1003: Finished scap sync-world: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules (duration: 06m 01s) * 07:25 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1069 * 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1069.eqiad.wmnet 164.48.64.10.in-addr.arpa 4.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:25 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1069.eqiad.wmnet 164.48.64.10.in-addr.arpa 4.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1069 - jiji@cumin1003" * 07:25 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1069 - jiji@cumin1003" * 07:25 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 07:25 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 07:24 jiji@deploy1003: jiji: Continuing with deployment * 07:22 jiji@deploy1003: jiji: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:21 jiji@deploy1003: Started scap sync-world: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules * 07:21 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1031.eqiad.wmnet,service=s7 * 07:20 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1031.eqiad.wmnet,service=s2 * 07:20 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1031.eqiad.wmnet,service=s7 * 07:20 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1031.eqiad.wmnet,service=s2 * 07:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts * 07:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts * 07:19 jiji@cumin1003: START - Cookbook sre.dns.netbox * 07:19 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1069 * 07:19 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1069.eqiad.wmnet with OS trixie * 07:19 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1069.eqiad.wmnet * 07:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 45 hosts * 07:17 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1069.eqiad.wmnet * 07:17 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1069.eqiad.wmnet * 07:14 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 45 hosts * 07:13 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 06:16 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs2007 after successful Bookworm reimage, data transfer, and postflight validation; wdqs1013 also passed postflights and is enabled in conftool, but remains out of IPVS pending a rolling pybal restart to clear its stale pre-VLAN-move address * 05:58 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2007.codfw.wmnet * 05:58 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1013.eqiad.wmnet * 05:54 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore scholarly data after Bookworm reimage) xfer scholarly_articles from wdqs2024.codfw.wmnet -> wdqs2016.codfw.wmnet, repooling source-only afterwards == 2026-07-22 == * 23:34 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Apply upgrade to JVM17 - eevans@cumin1003 * 23:14 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Apply upgrade to JVM17 - eevans@cumin1003 * 22:06 ryankemper: [WDQS] Added requestctl per-IP ratelimit `wdqs_heavy_sparql_bots_jul_2026_ratelimit` (chronic heavy-query bot tier driving deadlock-remediation restarts); pruned superseded `wdqs_2026_05_11_worobot` * 21:51 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] (duration: 11m 52s) * 21:44 sbassett@deploy1003: sbassett: Continuing with deployment * 21:43 sbassett@deploy1003: sbassett: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:39 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] * 20:38 dani@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] (duration: 32m 51s) * 20:38 ryankemper: [WDQS] Pruned obsolete requestctl action+pattern `wdqs_20260715_p2003_ring_ja3n` (actor rotated JA3Ns; rule inert) * 20:26 dani@deploy1003: dani, vadymts1: Continuing with deployment * 20:24 dani@deploy1003: dani, vadymts1: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:14 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2244: Testing * 20:06 dani@deploy1003: Started scap sync-world: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] * 19:56 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1013.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2016.codfw.wmnet with OS bookworm * 19:40 mutante: gerrit - one more service restart is needed - restarting * 19:29 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2244: Testing * 19:27 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2244: Testing * 19:27 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2244: Testing * 19:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2016.codfw.wmnet with reason: host reimage * 19:15 dancy@deploy1003: Finished deploy [zuul/deploy@d92e238]: Freshening Zuul installation (duration: 00m 15s) * 19:14 dancy@deploy1003: Started deploy [zuul/deploy@d92e238]: Freshening Zuul installation * 19:11 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2016.codfw.wmnet with reason: host reimage * 18:54 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1013.eqiad.wmnet, repooling source-only afterwards * 18:52 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2016 * 18:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2016 * 18:51 dancy@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 18:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2016.codfw.wmnet with OS bookworm * 18:39 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 11s) * 18:39 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 18:36 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 18:30 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417] (thin): Regular analytics weekly train THIN [analytics/refinery@2a25417d] (duration: 02m 09s) * 18:28 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417] (thin): Regular analytics weekly train THIN [analytics/refinery@2a25417d] * 18:28 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417]: Regular analytics weekly train [analytics/refinery@2a25417d] (duration: 04m 31s) * 18:27 dduvall: deploying https://gerrit.wikimedia.org/r/c/integration/config/+/1314025 (4 jobs updated) * 18:23 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417]: Regular analytics weekly train [analytics/refinery@2a25417d] * 18:22 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@2a25417d] (duration: 01m 59s) * 18:20 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@2a25417d] * 17:56 Raine: deployment server switchover => deploy1003 is primary now * 17:55 kamila@deploy1003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 28m 24s) * 17:54 mutante: restarting gerrit for maintenance * 17:29 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1013.eqiad.wmnet with OS bookworm * 17:27 kamila@deploy1003: Started scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] * 17:20 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] (duration: 22m 50s) * 17:12 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1023.eqiad.wmnet -> wdqs1024.eqiad.wmnet, repooling source-only afterwards * 17:04 Raine: point deployment.eqiad.wmnet to deploy1003 * 17:04 kamila@dns7001: END - running authdns-update * 17:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1013.eqiad.wmnet with reason: host reimage * 17:02 kamila@dns7001: START - running authdns-update * 17:01 kamila@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on releases2003.codfw.wmnet,releases1003.eqiad.wmnet with reason: Deployment server switchover * 17:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1013.eqiad.wmnet with reason: host reimage * 16:58 kamila@deploy2003: Locking from deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] * 16:57 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] (duration: 02m 33s) * 16:55 kamila@deploy2003: Locking from deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] * 16:55 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2003 - [[phab:T240266|T240266]] (duration: 00m 11s) * 16:54 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2003 - [[phab:T240266|T240266]] * 16:40 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1013 * 16:40 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1013 * 16:39 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1013 * 16:39 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1013.eqiad.wmnet 105.32.64.10.in-addr.arpa 5.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:39 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1013.eqiad.wmnet 105.32.64.10.in-addr.arpa 5.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:39 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:39 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1013 - bking@cumin2003" * 16:39 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1013 - bking@cumin2003" * 16:34 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:34 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1013 * 16:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1013.eqiad.wmnet with OS bookworm * 16:28 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1023.eqiad.wmnet -> wdqs1024.eqiad.wmnet, repooling source-only afterwards * 16:27 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-scholarly,name=eqiad * 16:27 eevans@deploy2003: helmfile [eqiad] DONE helmfile.d/services/linked-artifacts: apply * 16:26 eevans@deploy2003: helmfile [eqiad] START helmfile.d/services/linked-artifacts: apply * 16:26 eevans@deploy2003: helmfile [codfw] DONE helmfile.d/services/linked-artifacts: apply * 16:26 eevans@deploy2003: helmfile [codfw] START helmfile.d/services/linked-artifacts: apply * 16:25 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 29s) * 16:25 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 16:24 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 16:21 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 16:18 eevans@deploy2003: helmfile [codfw] DONE helmfile.d/services/linked-artifacts: apply * 16:18 eevans@deploy2003: helmfile [codfw] START helmfile.d/services/linked-artifacts: apply * 16:08 eevans@deploy2003: helmfile [staging] DONE helmfile.d/services/linked-artifacts: apply * 16:07 eevans@deploy2003: helmfile [staging] START helmfile.d/services/linked-artifacts: apply * 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 16:01 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 15:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1024.eqiad.wmnet with OS bookworm * 15:49 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] (duration: 00m 10s) * 15:49 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] * 15:48 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] (duration: 00m 15s) * 15:48 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] * 15:47 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] (duration: 00m 10s) * 15:47 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] * 15:46 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:42 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:42 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:40 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:37 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 15:37 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:36 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:36 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:36 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 15:33 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 15:33 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:31 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:28 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 15:27 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:27 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1024.eqiad.wmnet with reason: host reimage * 15:23 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:23 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:23 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:20 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1068.eqiad.wmnet * 15:20 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1068.eqiad.wmnet * 15:20 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1068.eqiad.wmnet * 15:20 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 15:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1024.eqiad.wmnet with reason: host reimage * 15:11 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:55 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 14:52 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wdqs1024.eqiad.wmnet with OS bookworm * 14:50 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] (duration: 00m 09s) * 14:50 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] * 14:49 jiji@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 14:49 jiji@deploy2003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 14:49 jiji@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 14:48 jiji@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 14:45 ecarg@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:45 ecarg@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:44 ecarg@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:44 ecarg@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:43 ecarg@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:43 ecarg@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:41 sukhe: ipvsadm --delete-service --tcp-service 10.2.1.55:8087: lvs2014 and lvs2013 * 14:39 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:39 sukhe: ipvsadm --delete-service --tcp-service 10.2.2.55:8087: [[phab:T432445|T432445]] * 14:38 ecarg@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:38 ecarg@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:37 ecarg@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:37 ecarg@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:36 ecarg@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:34 ecarg@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts datahubsearch1001.eqiad.wmnet * 14:32 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:32 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 14:31 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 14:31 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 14:28 sukhe: sudo cumin 'A:lvs-low-traffic-codfw' 'systemctl restart pybal': lvs2013 * 14:26 sukhe: sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal': lvs2014 * 14:26 sukhe: sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal' * 14:24 sukhe: restart pybal on lvs1019 * 14:24 sukhe: restart pybal on lvs1020 * 14:19 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for 1036 hosts * 14:17 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:04 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] (duration: 09m 28s) * 13:59 kharlan@deploy2003: dreamyjazz, kharlan: Continuing with deployment * 13:58 bking@cumin2003: START - Cookbook sre.hosts.decommission for hosts datahubsearch1001.eqiad.wmnet * 13:57 kharlan@deploy2003: dreamyjazz, kharlan: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:55 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] * 13:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts datahubsearch[1002-1003].eqiad.wmnet * 13:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:53 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch[1002-1003].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 13:52 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch[1002-1003].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 13:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 13:42 stran@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] (duration: 07m 30s) * 13:42 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:40 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: Test * 13:38 stran@deploy2003: dragoniez, stran: Continuing with deployment * 13:37 stran@deploy2003: dragoniez, stran: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:35 bking@cumin2003: START - Cookbook sre.hosts.decommission for hosts datahubsearch[1002-1003].eqiad.wmnet * 13:35 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024'] * 13:35 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 13:35 stran@deploy2003: Started scap sync-world: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] * 13:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 13:28 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024'] * 13:26 sukhe@dns1004: END - running authdns-update * 13:25 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 13:25 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 13:24 sukhe@dns1004: START - running authdns-update * 13:22 stran@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] (duration: 08m 20s) * 13:21 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 13:20 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 13:19 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 13:19 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 13:19 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 13:18 stran@deploy2003: stran: Continuing with deployment * 13:16 stran@deploy2003: stran: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:14 stran@deploy2003: Started scap sync-world: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] * 13:13 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 13:13 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 13:11 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 13:11 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 13:08 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 12:55 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool es2051: Test * 12:55 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: Test * 12:54 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool es2051: Test * 12:43 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1036 hosts * 12:41 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 12:40 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1048.eqiad.wmnet * 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 12:39 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 12:38 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 12:37 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 12:37 brouberol@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 12:36 brouberol@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 12:36 brouberol@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 12:36 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 12:36 elukey@cumin1003: DONE (PASS) - Cookbook sre.puppet.renew-cert (exit_code=0) for crm2001.codfw.wmnet: Renew puppet certificate - elukey@cumin1003 * 12:35 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:35 brouberol@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 12:34 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 12:31 brouberol@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 12:30 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1048.eqiad.wmnet * 12:30 brouberol@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 12:28 brouberol@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 12:27 brouberol@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 12:20 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1068.eqiad.wmnet with OS trixie * 12:01 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] (duration: 13m 19s) * 11:58 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1068.eqiad.wmnet with reason: host reimage * 11:52 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1068.eqiad.wmnet with reason: host reimage * 11:51 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 11:49 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:47 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] * 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2252: Security updates * 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:43 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 11:42 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2252: Security updates * 11:42 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply * 11:40 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply * 11:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 11:37 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2252.codfw.wmnet with OS trixie * 11:34 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1068 * 11:34 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1068 * 11:26 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1068 * 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1068.eqiad.wmnet 46.48.64.10.in-addr.arpa 6.4.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:26 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1068.eqiad.wmnet 46.48.64.10.in-addr.arpa 6.4.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1068 - jiji@cumin1003" * 11:26 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1068 - jiji@cumin1003" * 11:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2252.codfw.wmnet with reason: host reimage * 11:17 jiji@cumin1003: START - Cookbook sre.dns.netbox * 11:17 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1068 * 11:17 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1068.eqiad.wmnet with OS trixie * 11:17 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2252.codfw.wmnet with reason: host reimage * 11:15 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1068.eqiad.wmnet * 11:15 mvolz@deploy2003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:15 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1068.eqiad.wmnet * 11:15 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1068.eqiad.wmnet * 11:14 mvolz@deploy2003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:13 mvolz@deploy2003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:13 mvolz@deploy2003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:12 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] (duration: 11m 05s) * 11:11 mvolz@deploy2003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:10 mvolz@deploy2003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:07 dreamyjazz@deploy2003: dreamyjazz, kharlan: Continuing with deployment * 11:03 dreamyjazz@deploy2003: dreamyjazz, kharlan: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:03 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2252.codfw.wmnet with OS trixie * 11:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2252: Upgrading db2252.codfw.wmnet * 11:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:02 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 11:02 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2252: Upgrading db2252.codfw.wmnet * 11:02 cwilliams@cumin1003: dbmaint on ms3@codfw [[phab:T432321|T432321]] * 11:01 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 11:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db1153.eqiad.wmnet with reason: Security updates * 11:01 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] * 11:00 fnegri@deploy2003: helmfile [eqiad] DONE helmfile.d/services/toolhub: apply * 10:58 fnegri@deploy2003: helmfile [eqiad] START helmfile.d/services/toolhub: apply * 10:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1151: Security updates * 10:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:57 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1151: Security updates * 10:55 fnegri@deploy2003: helmfile [codfw] DONE helmfile.d/services/toolhub: apply * 10:54 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] (duration: 08m 38s) * 10:53 fnegri@deploy2003: helmfile [codfw] START helmfile.d/services/toolhub: apply * 10:53 fnegri@deploy2003: helmfile [staging] DONE helmfile.d/services/toolhub: apply * 10:52 fnegri@deploy2003: helmfile [staging] START helmfile.d/services/toolhub: apply * 10:50 zabe@deploy2003: zabe: Continuing with deployment * 10:47 zabe@deploy2003: zabe: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:45 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] * 10:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1151: Security updates * 10:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:42 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:42 root@cumin1003: START - Cookbook sre.mysql.depool depool db1151: Security updates * 10:38 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] (duration: 12m 47s) * 10:34 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 10:34 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 10:33 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2253.codfw.wmnet with OS trixie * 10:28 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:26 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] * 10:18 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2253.codfw.wmnet with reason: host reimage * 10:13 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2253.codfw.wmnet with reason: host reimage * 10:00 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2253.codfw.wmnet with OS trixie * 09:58 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db1151.eqiad.wmnet with reason: Security updates * 09:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2253: Upgrading db2253.codfw.wmnet * 09:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:57 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 09:56 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2253: Upgrading db2253.codfw.wmnet * 09:56 cwilliams@cumin1003: dbmaint on ms2@codfw [[phab:T432321|T432321]] * 09:56 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 09:36 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: UI improvement; support url shortener - oblivian@cumin1003" * 09:36 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: UI improvement; support url shortener - oblivian@cumin1003 * 09:35 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: UI improvement; support url shortener - oblivian@cumin1003 * 09:35 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: UI improvement; support url shortener - oblivian@cumin1003" * 09:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1152: Security updates * 09:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:26 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db1152: Security updates * 09:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: Security updates * 09:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:11 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:11 root@cumin1003: START - Cookbook sre.mysql.depool depool db1152: Security updates * 09:10 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1018.eqiad.wmnet with reason: Cloning * 09:09 Dreamy_Jazz: Deployed patch for [[phab:T432453|T432453]] and [[phab:T432454|T432454]] * 09:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 09:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2251.codfw.wmnet with OS trixie * 08:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2251.codfw.wmnet with reason: host reimage * 08:45 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2251.codfw.wmnet with reason: host reimage * 08:40 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] (duration: 12m 26s) * 08:38 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1030.eqiad.wmnet,service=s1 * 08:36 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 08:31 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2251.codfw.wmnet with OS trixie * 08:30 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:28 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] * 08:25 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] (duration: 07m 59s) * 08:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2251: Upgrading db2251.codfw.wmnet * 08:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:22 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 08:22 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2251: Upgrading db2251.codfw.wmnet * 08:20 urbanecm@deploy2003: urbanecm: Continuing with deployment * 08:20 cwilliams@cumin1003: dbmaint on ms1@codfw [[phab:T432321|T432321]] * 08:20 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 08:19 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:17 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] * 08:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade * 08:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade * 08:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db2251.codfw.wmnet,db1152.eqiad.wmnet with reason: OS upgrade * 08:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade * 08:13 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade * 08:11 Dreamy_Jazz: Created cusi_signal, cusi_case, and cusi_user on ukwiki and enwikivoyage in extension1 * 08:11 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade * 08:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade * 08:04 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1030.eqiad.wmnet,service=s1 * 08:04 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1030.eqiad.wmnet,service=s1 * 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply * 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply * 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply * 07:51 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply * 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 07:47 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 07:47 phuedx: End of UTC morning backport window * 07:43 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 07:43 phuedx@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] (duration: 13m 44s) * 07:43 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 07:39 phuedx@deploy2003: phuedx: Continuing with deployment * 07:31 phuedx@deploy2003: phuedx: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:29 phuedx@deploy2003: Started scap sync-world: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] * 07:24 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Turnilo import support - oblivian@cumin1003" * 07:24 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import support - oblivian@cumin1003 * 07:23 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import support - oblivian@cumin1003 * 07:23 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Turnilo import support - oblivian@cumin1003" * 06:42 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs2020 after successful Bookworm reimage, data transfer, and postflight validation * 06:42 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2020.codfw.wmnet * 05:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Managing sanitization for wikis bolwiki in section s5 * 05:25 marostegui@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis bolwiki in section s5 * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 41s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-21 == * 22:50 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2019.codfw.wmnet -> wdqs2020.codfw.wmnet, repooling source-only afterwards * 22:47 cwhite: force reboot arclamp2001 - appears to have run out of memory and gone unresponsive * 22:24 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 01m 26s) * 22:24 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 22:23 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 22:22 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024'] * 22:11 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 21:54 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs1024'] * 21:54 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 21:53 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs1024'] * 21:53 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 21:49 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2019.codfw.wmnet -> wdqs2020.codfw.wmnet, repooling source-only afterwards * 20:57 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] (duration: 09m 10s) * 20:55 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1024.eqiad.wmnet with OS bookworm * 20:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2020.codfw.wmnet with OS bookworm * 20:52 krinkle@deploy2003: krinkle: Continuing with deployment * 20:49 krinkle@deploy2003: krinkle: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:47 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] * 20:45 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] (duration: 05m 42s) * 20:44 krinkle@deploy2003: krinkle: Rolling back deployment * 20:41 krinkle@deploy2003: krinkle: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:39 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] * 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2003.codfw.wmnet * 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1003.eqiad.wmnet * 20:33 dani@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] (duration: 11m 15s) * 20:33 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2003.codfw.wmnet * 20:33 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1003.eqiad.wmnet * 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2020.codfw.wmnet with reason: host reimage * 20:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1002.eqiad.wmnet * 20:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2002.codfw.wmnet * 20:30 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 20:29 dani@deploy2003: dani: Continuing with deployment * 20:29 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2020.codfw.wmnet with reason: host reimage * 20:26 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1002.eqiad.wmnet * 20:26 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2002.codfw.wmnet * 20:24 dani@deploy2003: dani: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2001.codfw.wmnet * 20:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1001.eqiad.wmnet * 20:22 dani@deploy2003: Started scap sync-world: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] * 20:22 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 20:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2001.codfw.wmnet * 20:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1001.eqiad.wmnet * 20:14 sbisson@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] (duration: 09m 01s) * 20:11 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2020.codfw.wmnet with OS bookworm * 20:10 sbisson@deploy2003: sbisson: Continuing with deployment * 20:07 sbisson@deploy2003: sbisson: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:05 sbisson@deploy2003: Started scap sync-world: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] * 20:03 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] (duration: 07m 04s) * 20:01 mutante: Gerrit - tomorrow a new SSH host key will appear - it will be {{Gerrit|ed25519}} and has already been added to wmf-laptop. you can verify it here: https://wikitech.wikimedia.org/wiki/Help:SSH_Fingerprints/gerrit.wikimedia.org:29418 ([[phab:T240266|T240266]]) * 19:59 zabe@deploy2003: zabe: Continuing with deployment * 19:58 zabe@deploy2003: zabe: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:56 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] * 19:52 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] (duration: 07m 25s) * 19:48 zabe@deploy2003: zabe: Continuing with deployment * 19:47 zabe@deploy2003: zabe: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:45 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] * 19:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 19:32 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024'] * 19:27 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 19:26 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024'] * 19:26 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 19:24 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024'] * 19:12 ryankemper: [wdqs] [[phab:T430880|T430880]] Repooled `wdqs-scholarly` discovery in `eqiad` after validating `wdqs1023` end-to-end; `wdqs1024` remains disabled pending reimage recovery * 19:11 ryankemper: [wdqs] [[phab:T430880|T430880]] Repooled wdqs1012.eqiad.wmnet after successful Bookworm reimage, data transfer, service checks, readiness probe, and cross-graph federation query validation * 19:10 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 19:10 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1012.eqiad.wmnet * 19:08 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 18:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for deploy1003.eqiad.wmnet * 18:57 kamila@cumin1003: START - Cookbook sre.hosts.remove-downtime for deploy1003.eqiad.wmnet * 18:37 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:37 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding urldownloader service IPs - sukhe@cumin1003" * 18:37 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding urldownloader service IPs - sukhe@cumin1003" * 18:32 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 18:32 dancy@deploy2003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 18:30 sukhe@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 18:27 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 18:24 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1024.eqiad.wmnet with OS bookworm * 18:20 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host deploy1003.eqiad.wmnet with OS bookworm * 18:09 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deploy1003 reimage (duration: 121m 16s) * 18:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 18:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1180: Security updates * 17:55 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wcqs2003.codfw.wmnet * 17:48 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wcqs2003.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1155.eqiad.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1155.eqiad.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2224.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2224.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2217.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2217.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2193.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2193.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2180.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2180.codfw.wmnet * 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1168.eqiad.wmnet * 17:36 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1168.eqiad.wmnet * 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2169.codfw.wmnet * 17:36 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2169.codfw.wmnet * 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1165.eqiad.wmnet * 17:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1165.eqiad.wmnet * 17:35 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2158.codfw.wmnet * 17:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2158.codfw.wmnet * 17:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wcqs1003.eqiad.wmnet * 17:17 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1180: Security updates * 17:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1180.eqiad.wmnet * 17:16 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1180.eqiad.wmnet * 17:15 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp3073.* * 17:13 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wcqs1003.eqiad.wmnet * 17:11 brett@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp3073.esams.wmnet with OS trixie * 17:11 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 17:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1024 * 17:04 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1024 * 17:03 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 17:00 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 16:59 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2242: codfw rack B7 depool for maintenance * 16:59 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 16:43 brett@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp3073.esams.wmnet with reason: host reimage * 16:42 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 16:39 brett@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cp3073.esams.wmnet with reason: host reimage * 16:32 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on deploy1003.eqiad.wmnet with reason: host reimage * 16:27 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on deploy1003.eqiad.wmnet with reason: host reimage * 16:14 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2242: codfw rack B7 depool for maintenance * 16:14 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: codfw rack B7 depool for maintenance * 16:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1023.eqiad.wmnet with OS bookworm * 16:13 brett@cumin2002: START - Cookbook sre.hosts.reimage for host cp3073.esams.wmnet with OS trixie * 16:08 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host deploy1003.eqiad.wmnet with OS bookworm * 16:08 kamila@deploy2003: Locking from deployment [MediaWiki]: deploy1003 reimage * 16:03 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1012.eqiad.wmnet with OS bookworm * 15:48 inflatador: bking@apt1002 `sudo reprepro copy bookworm-wikimedia bullseye-wikimedia jvmquake` [[phab:T430880|T430880]] * 15:39 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp3073.* * 15:39 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 15:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:34 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1311536{{!}}Set $wgMathInternalRestbaseURL explicitly (take 2) (T349582)]] * 15:29 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:29 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2228: codfw rack B7 depool for maintenance * 15:29 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2229: codfw rack B7 depool for maintenance * 15:27 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1180: Security update * 15:25 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Security update * 15:21 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 15:21 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 15:19 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 24s) * 15:19 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:14 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 15:14 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db1180: Security update * 15:13 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-eqiad * 14:48 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-eqiad * 14:44 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2229: codfw rack B7 depool for maintenance * 14:44 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc2017: codfw rack B7 depool for maintenance * 14:44 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:43 cmooney@cumin2003: START - Cookbook sre.mysql.parsercache * 14:43 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool pc2017: codfw rack B7 depool for maintenance * 14:43 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2003.codfw.wmnet * 14:43 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2003.codfw.wmnet * 14:42 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2009.codfw.wmnet * 14:42 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2009.codfw.wmnet * 14:41 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:41 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:40 cmooney@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 29 hosts * 14:40 cmooney@cumin1003: START - Cookbook sre.hosts.remove-downtime for 29 hosts * 14:35 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 14:34 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 14:32 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] (duration: 07m 56s) * 14:29 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ssw1-a[1,8]-codfw with reason: lsw1-b7-codfw JunOS upgrade * 14:28 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 14:28 elukey@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 14:26 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:24 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] * 14:23 topranks: reboot lsw1-b7-codfw to upgrade JunOS (affects all hosts in rack) [[phab:T430928|T430928]] * 14:18 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2003.codfw.wmnet * 14:14 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2009.codfw.wmnet * 14:14 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Security update * 14:13 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2242: codfw rack B7 depool for maintenance * 14:13 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2242: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2228: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2228: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2229: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2229: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc2017: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.parsercache * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool pc2017: codfw rack B7 depool for maintenance * 14:08 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2003.codfw.wmnet * 14:07 cmooney@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on aux-k8s-etcd2004.codfw.wmnet,ml-etcd2001.codfw.wmnet with reason: lsw1-b7-codfw JunOS upgrade * 14:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95005 and previous config saved to /var/cache/conftool/dbconfig/20260721-140620-cwilliams.json * 14:05 cmooney@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2049.codfw.wmnet * 14:05 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-scholarly,name=eqiad * 14:04 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2009.codfw.wmnet * 14:04 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply * 14:04 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply * 14:03 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:03 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2001.codfw.wmnet * 14:03 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2001.codfw.wmnet * 14:02 cmooney@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2049.codfw.wmnet * 14:00 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:00 Dreamy_Jazz: Created cusi_case, cusi_signal, and cusi_user on svwiki, dewiki, jawiki, eswiki * 13:59 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b7-codfw,lsw1-b7-codfw IPv6,lsw1-b7-codfw.mgmt,ssw1-a[1,8]-codfw.mgmt with reason: lsw1-b7-codfw JunOS upgrade * 13:57 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs1023.eqiad.wmnet, repooling source-only afterwards * 13:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 29 hosts with reason: lsw1-b7-codfw JunOS upgrade * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224', diff saved to https://phabricator.wikimedia.org/P95003 and previous config saved to /var/cache/conftool/dbconfig/20260721-135613-cwilliams.json * 13:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1012.eqiad.wmnet with reason: host reimage * 13:53 cmooney@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 1:00:00 on 30 hosts with reason: lsw1-b7-codfw JunOS upgrade * 13:51 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1012.eqiad.wmnet with reason: host reimage * 13:48 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 13:48 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 13:46 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224', diff saved to https://phabricator.wikimedia.org/P95001 and previous config saved to /var/cache/conftool/dbconfig/20260721-134605-cwilliams.json * 13:46 cmooney@cumin1003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti2033.codfw.wmnet * 13:46 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 13:45 elukey: move the Docker Registry's /v2/wikimedia/machinelearning.* prefix to the ml S3 backend - [[phab:T428022|T428022]] * 13:45 cmooney@cumin1003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti2033.codfw.wmnet * 13:45 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 13:43 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:40 jiji@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 13:40 jiji@deploy2003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 13:39 jiji@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 13:39 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 13:38 cmooney@cumin1003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2032.codfw.wmnet * 13:38 jiji@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 13:38 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 13:37 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2032.codfw.wmnet * 13:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95000 and previous config saved to /var/cache/conftool/dbconfig/20260721-133557-cwilliams.json * 13:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1012 * 13:33 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1012 * 13:33 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1012.eqiad.wmnet with OS bookworm * 13:30 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:30 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:28 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94999 and previous config saved to /var/cache/conftool/dbconfig/20260721-132855-cwilliams.json * 13:28 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2224.codfw.wmnet with reason: Maintenance * 13:28 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94998 and previous config saved to /var/cache/conftool/dbconfig/20260721-132826-cwilliams.json * 13:28 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 13:23 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] (duration: 07m 50s) * 13:20 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 13:18 kharlan@deploy2003: kharlan: Continuing with deployment * 13:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217', diff saved to https://phabricator.wikimedia.org/P94996 and previous config saved to /var/cache/conftool/dbconfig/20260721-131817-cwilliams.json * 13:17 brouberol@dns1004: END - running authdns-update * 13:17 kharlan@deploy2003: kharlan: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:15 brouberol@dns1004: START - running authdns-update * 13:15 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] * 13:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94995 and previous config saved to /var/cache/conftool/dbconfig/20260721-131411-cwilliams.json * 13:13 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs1023.eqiad.wmnet, repooling source-only afterwards * 13:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217', diff saved to https://phabricator.wikimedia.org/P94994 and previous config saved to /var/cache/conftool/dbconfig/20260721-130809-cwilliams.json * 13:07 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 13:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180', diff saved to https://phabricator.wikimedia.org/P94993 and previous config saved to /var/cache/conftool/dbconfig/20260721-130404-cwilliams.json * 13:03 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:03 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:02 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 13:02 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 12:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94992 and previous config saved to /var/cache/conftool/dbconfig/20260721-125801-cwilliams.json * 12:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180', diff saved to https://phabricator.wikimedia.org/P94991 and previous config saved to /var/cache/conftool/dbconfig/20260721-125356-cwilliams.json * 12:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94990 and previous config saved to /var/cache/conftool/dbconfig/20260721-125049-cwilliams.json * 12:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2217.codfw.wmnet with reason: Maintenance * 12:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94989 and previous config saved to /var/cache/conftool/dbconfig/20260721-125017-cwilliams.json * 12:48 elukey: bmc cold reboot for lvs1013 and lvs1015 - [[phab:T426180|T426180]] * 12:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94988 and previous config saved to /var/cache/conftool/dbconfig/20260721-124348-cwilliams.json * 12:40 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193', diff saved to https://phabricator.wikimedia.org/P94987 and previous config saved to /var/cache/conftool/dbconfig/20260721-124009-cwilliams.json * 12:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts * 12:33 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts * 12:33 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts * 12:32 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts * 12:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:30 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193', diff saved to https://phabricator.wikimedia.org/P94986 and previous config saved to /var/cache/conftool/dbconfig/20260721-123001-cwilliams.json * 12:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94985 and previous config saved to /var/cache/conftool/dbconfig/20260721-121953-cwilliams.json * 12:17 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs2007.codfw.wmnet with OS bookworm * 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94983 and previous config saved to /var/cache/conftool/dbconfig/20260721-121257-cwilliams.json * 12:12 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2193.codfw.wmnet with reason: Maintenance * 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94982 and previous config saved to /var/cache/conftool/dbconfig/20260721-121239-cwilliams.json * 12:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180', diff saved to https://phabricator.wikimedia.org/P94980 and previous config saved to /var/cache/conftool/dbconfig/20260721-120231-cwilliams.json * 11:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180', diff saved to https://phabricator.wikimedia.org/P94979 and previous config saved to /var/cache/conftool/dbconfig/20260721-115223-cwilliams.json * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94978 and previous config saved to /var/cache/conftool/dbconfig/20260721-114333-cwilliams.json * 11:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1180.eqiad.wmnet with reason: Maintenance * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94977 and previous config saved to /var/cache/conftool/dbconfig/20260721-114305-cwilliams.json * 11:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94976 and previous config saved to /var/cache/conftool/dbconfig/20260721-114215-cwilliams.json * 11:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94975 and previous config saved to /var/cache/conftool/dbconfig/20260721-113530-cwilliams.json * 11:35 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2180.codfw.wmnet with reason: Maintenance * 11:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94974 and previous config saved to /var/cache/conftool/dbconfig/20260721-113501-cwilliams.json * 11:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168', diff saved to https://phabricator.wikimedia.org/P94973 and previous config saved to /var/cache/conftool/dbconfig/20260721-113258-cwilliams.json * 11:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169', diff saved to https://phabricator.wikimedia.org/P94972 and previous config saved to /var/cache/conftool/dbconfig/20260721-112453-cwilliams.json * 11:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168', diff saved to https://phabricator.wikimedia.org/P94971 and previous config saved to /var/cache/conftool/dbconfig/20260721-112250-cwilliams.json * 11:21 XioNoX: put eqiad-drmrs Arelion link in service * 11:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169', diff saved to https://phabricator.wikimedia.org/P94970 and previous config saved to /var/cache/conftool/dbconfig/20260721-111446-cwilliams.json * 11:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94969 and previous config saved to /var/cache/conftool/dbconfig/20260721-111242-cwilliams.json * 11:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 11:10 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 11:07 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1093 hosts * 11:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94968 and previous config saved to /var/cache/conftool/dbconfig/20260721-110548-cwilliams.json * 11:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1168.eqiad.wmnet with reason: Maintenance * 11:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94967 and previous config saved to /var/cache/conftool/dbconfig/20260721-110520-cwilliams.json * 11:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94966 and previous config saved to /var/cache/conftool/dbconfig/20260721-110439-cwilliams.json * 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94964 and previous config saved to /var/cache/conftool/dbconfig/20260721-105632-cwilliams.json * 10:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2169.codfw.wmnet with reason: Maintenance * 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94963 and previous config saved to /var/cache/conftool/dbconfig/20260721-105603-cwilliams.json * 10:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165', diff saved to https://phabricator.wikimedia.org/P94962 and previous config saved to /var/cache/conftool/dbconfig/20260721-105512-cwilliams.json * 10:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158', diff saved to https://phabricator.wikimedia.org/P94961 and previous config saved to /var/cache/conftool/dbconfig/20260721-104555-cwilliams.json * 10:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165', diff saved to https://phabricator.wikimedia.org/P94960 and previous config saved to /var/cache/conftool/dbconfig/20260721-104504-cwilliams.json * 10:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158', diff saved to https://phabricator.wikimedia.org/P94959 and previous config saved to /var/cache/conftool/dbconfig/20260721-103547-cwilliams.json * 10:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94958 and previous config saved to /var/cache/conftool/dbconfig/20260721-103456-cwilliams.json * 10:29 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2229: Upgraded kernel * 10:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94956 and previous config saved to /var/cache/conftool/dbconfig/20260721-102757-cwilliams.json * 10:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on an-redacteddb1001.eqiad.wmnet,clouddb[1015,1025,1028].eqiad.wmnet,db1155.eqiad.wmnet with reason: Maintenance * 10:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1165.eqiad.wmnet with reason: Maintenance * 10:25 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94955 and previous config saved to /var/cache/conftool/dbconfig/20260721-102539-cwilliams.json * 10:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94954 and previous config saved to /var/cache/conftool/dbconfig/20260721-101848-cwilliams.json * 10:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2158.codfw.wmnet with reason: Maintenance * 09:43 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2229: Upgraded kernel * 09:42 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2229.codfw.wmnet * 09:42 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2229.codfw.wmnet * 09:23 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db2229.codfw.wmnet * 09:23 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2229.codfw.wmnet * 08:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2229 [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94948 and previous config saved to /var/cache/conftool/dbconfig/20260721-085724-cwilliams.json * 08:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2214 to s6 primary [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94947 and previous config saved to /var/cache/conftool/dbconfig/20260721-085442-cwilliams.json * 08:53 cezmunsta: Starting s6 codfw failover from db2229 to db2214 - [[phab:T430964|T430964]] * 08:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2214 with weight 0 [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94946 and previous config saved to /var/cache/conftool/dbconfig/20260721-084613-cwilliams.json * 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 22 hosts with reason: Primary switchover s6 [[phab:T430964|T430964]] * 08:32 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1017.eqiad.wmnet,service=s1 * 08:08 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Add subrated circuit rate to interface descriptions - CR1312476 - ayounsi@cumin1003 * 08:06 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Add subrated circuit rate to interface descriptions - CR1312476 - ayounsi@cumin1003 * 07:58 reedy@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] (duration: 12m 55s) * 07:51 reedy@deploy2003: reedy, neriah: Continuing with deployment * 07:51 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 07:51 reedy@deploy2003: reedy, neriah: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:48 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1093 hosts * 07:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm2001.wikimedia.org * 07:45 reedy@deploy2003: Started scap sync-world: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] * 07:43 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 07:42 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm2001.wikimedia.org * 07:23 elukey: upgrade libtiff6 packages on zuul* trixie hosts for security upgrades * 07:22 elukey: upgrade libtiff6 packages on Wikikube trixie workers for security upgrades * 07:14 elukey@deploy2003: helmfile [codfw] DONE helmfile.d/services/proton: sync * 07:13 elukey@deploy2003: helmfile [codfw] START helmfile.d/services/proton: sync * 07:11 elukey@deploy2003: helmfile [eqiad] DONE helmfile.d/services/proton: sync * 07:10 elukey@deploy2003: helmfile [eqiad] START helmfile.d/services/proton: sync * 07:09 elukey@deploy2003: helmfile [staging] DONE helmfile.d/services/proton: sync * 07:08 elukey@deploy2003: helmfile [staging] START helmfile.d/services/proton: sync * 06:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1023.eqiad.wmnet with reason: host reimage * 06:46 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1023.eqiad.wmnet with reason: host reimage * 06:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 05:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Haproxy-only mode support - oblivian@cumin1003" * 05:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Haproxy-only mode support - oblivian@cumin1003 * 05:42 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Haproxy-only mode support - oblivian@cumin1003 * 05:42 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Haproxy-only mode support - oblivian@cumin1003" * 05:38 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1017.eqiad.wmnet with reason: Cloning * 05:37 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1017.eqiad.wmnet,service=s1 * 05:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1029.eqiad.wmnet,service=s8 * 05:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1029.eqiad.wmnet,service=s5 * 05:32 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:30 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 05:11 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet * 05:04 aokoth@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet * 05:00 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 04:56 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 04:01 mwpresync@deploy2003: Pruned MediaWiki: 1.47.0-wmf.9 (duration: 01m 08s) * 03:41 mwpresync@deploy2003: Finished scap sync-world: testwikis to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] (duration: 36m 30s) * 03:05 mwpresync@deploy2003: Started scap sync-world: testwikis to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 03:01 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:01 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:00 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:00 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:36 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:36 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:36 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:35 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:16 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 47s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 00:56 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm == 2026-07-20 == * 23:38 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 23:07 Amir1: deleting echo notifications from 2015 on group1 wikis * 23:07 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] (duration: 14m 16s) * 23:01 ladsgroup@deploy2003: ladsgroup: Continuing with deployment * 23:00 ladsgroup@deploy2003: ladsgroup: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:53 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] * 22:46 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2007.codfw.wmnet, repooling source-only afterwards * 22:39 maryum: Deployed security fixes for several security bugs * 21:42 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 21:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2007.codfw.wmnet, repooling source-only afterwards * 21:37 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 21:37 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 21:34 sbassett: Deployed security fix for [[phab:T432424|T432424]] * 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs2020.codfw.wmnet * 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1023.eqiad.wmnet * 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1011.eqiad.wmnet * 21:32 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 17s) * 21:32 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 21:27 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 21:13 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2007.codfw.wmnet with reason: host reimage * 21:08 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-internal-main,name=codfw * 21:06 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2007.codfw.wmnet with reason: host reimage * 20:59 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service * 20:58 sukhe: pybal restart for IP changes around wdqs-main hosts * 20:57 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 20:46 ryankemper@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-internal-main,name=codfw * 20:45 ebernhardson@deploy2003: Finished deploy [search/mjolnir/deploy@d4dc3b8]: Update for opensearch 2.x compat (duration: 00m 34s) * 20:44 ebernhardson@deploy2003: Started deploy [search/mjolnir/deploy@d4dc3b8]: Update for opensearch 2.x compat * 20:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2007 * 20:44 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2007 * 20:43 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2007 * 20:43 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2007.codfw.wmnet 156.16.192.10.in-addr.arpa 6.5.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:42 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2007.codfw.wmnet 156.16.192.10.in-addr.arpa 6.5.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:42 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2007 - bking@cumin2003" * 20:41 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2007 - bking@cumin2003" * 20:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94944 and previous config saved to /var/cache/conftool/dbconfig/20260720-203333-cwilliams.json * 20:32 arlolra@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] (duration: 15m 07s) * 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2020.codfw.wmnet * 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1023.eqiad.wmnet * 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1011.eqiad.wmnet * 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs2020.codfw.wmnet * 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1023.eqiad.wmnet * 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1011.eqiad.wmnet * 20:25 arlolra@deploy2003: arlolra, cscott: Continuing with deployment * 20:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257', diff saved to https://phabricator.wikimedia.org/P94943 and previous config saved to /var/cache/conftool/dbconfig/20260720-202325-cwilliams.json * 20:21 arlolra@deploy2003: arlolra, cscott: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:17 arlolra@deploy2003: Started scap sync-world: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] * 20:13 bking@cumin2003: START - Cookbook sre.dns.netbox * 20:13 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257', diff saved to https://phabricator.wikimedia.org/P94942 and previous config saved to /var/cache/conftool/dbconfig/20260720-201318-cwilliams.json * 20:13 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 20:10 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 20:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2007 * 20:04 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2007.codfw.wmnet with OS bookworm * 20:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94941 and previous config saved to /var/cache/conftool/dbconfig/20260720-200310-cwilliams.json * 19:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94940 and previous config saved to /var/cache/conftool/dbconfig/20260720-195633-cwilliams.json * 19:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1257.eqiad.wmnet with reason: Maintenance * 19:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94939 and previous config saved to /var/cache/conftool/dbconfig/20260720-195605-cwilliams.json * 19:51 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 19:50 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 19:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256', diff saved to https://phabricator.wikimedia.org/P94938 and previous config saved to /var/cache/conftool/dbconfig/20260720-194558-cwilliams.json * 19:44 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 19:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 19:41 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wdqs1011.eqiad.wmnet with OS bookworm * 19:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256', diff saved to https://phabricator.wikimedia.org/P94937 and previous config saved to /var/cache/conftool/dbconfig/20260720-193550-cwilliams.json * 19:25 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94936 and previous config saved to /var/cache/conftool/dbconfig/20260720-192542-cwilliams.json * 19:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94935 and previous config saved to /var/cache/conftool/dbconfig/20260720-191856-cwilliams.json * 19:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1256.eqiad.wmnet with reason: Maintenance * 19:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94934 and previous config saved to /var/cache/conftool/dbconfig/20260720-191839-cwilliams.json * 19:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255', diff saved to https://phabricator.wikimedia.org/P94933 and previous config saved to /var/cache/conftool/dbconfig/20260720-190831-cwilliams.json * 18:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255', diff saved to https://phabricator.wikimedia.org/P94932 and previous config saved to /var/cache/conftool/dbconfig/20260720-185824-cwilliams.json * 18:50 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 18:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94931 and previous config saved to /var/cache/conftool/dbconfig/20260720-184816-cwilliams.json * 18:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94930 and previous config saved to /var/cache/conftool/dbconfig/20260720-184224-cwilliams.json * 18:42 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1255.eqiad.wmnet with reason: Maintenance * 18:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94929 and previous config saved to /var/cache/conftool/dbconfig/20260720-184153-cwilliams.json * 18:39 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 18:39 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 16s) * 18:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 18:38 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 59m 26s) * 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 18:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211', diff saved to https://phabricator.wikimedia.org/P94928 and previous config saved to /var/cache/conftool/dbconfig/20260720-183145-cwilliams.json * 18:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211', diff saved to https://phabricator.wikimedia.org/P94927 and previous config saved to /var/cache/conftool/dbconfig/20260720-182137-cwilliams.json * 18:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94926 and previous config saved to /var/cache/conftool/dbconfig/20260720-181129-cwilliams.json * 18:09 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_codfw * 18:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2057.codfw.wmnet * 18:08 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_codfw * 18:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2058.codfw.wmnet * 18:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94925 and previous config saved to /var/cache/conftool/dbconfig/20260720-180452-cwilliams.json * 18:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on clouddb[1016,1020,1022-1023].eqiad.wmnet,db1154.eqiad.wmnet with reason: Maintenance * 18:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1211.eqiad.wmnet with reason: Maintenance * 18:02 sukhe: armed keyholder on acmechief1002.eqiad.wmnet and acmechief2002.codfw.wmnet (active host) * 18:01 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief2002.codfw.wmnet * 17:57 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief2002.codfw.wmnet * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs2020'] * 17:52 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief1002.eqiad.wmnet * 17:50 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 17:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1011.eqiad.wmnet with reason: host reimage * 17:48 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief1002.eqiad.wmnet * 17:47 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test2001.codfw.wmnet * 17:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94924 and previous config saved to /var/cache/conftool/dbconfig/20260720-174717-cwilliams.json * 17:46 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 17:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1011.eqiad.wmnet with reason: host reimage * 17:43 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs2020.codfw.wmnet with OS bookworm * 17:43 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test2001.codfw.wmnet * 17:43 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test1001.eqiad.wmnet * 17:39 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test1001.eqiad.wmnet * 17:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 17:38 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:38 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 17:37 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 17:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244', diff saved to https://phabricator.wikimedia.org/P94923 and previous config saved to /var/cache/conftool/dbconfig/20260720-173709-cwilliams.json * 17:35 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 17:31 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 20m 40s) * 17:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2055.codfw.wmnet * 17:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2056.codfw.wmnet * 17:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1011 * 17:27 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1011 * 17:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1011.eqiad.wmnet with OS bookworm * 17:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244', diff saved to https://phabricator.wikimedia.org/P94922 and previous config saved to /var/cache/conftool/dbconfig/20260720-172701-cwilliams.json * 17:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94921 and previous config saved to /var/cache/conftool/dbconfig/20260720-171653-cwilliams.json * 17:11 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 17:11 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 13m 03s) * 17:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94920 and previous config saved to /var/cache/conftool/dbconfig/20260720-171012-cwilliams.json * 17:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2244.codfw.wmnet with reason: Maintenance * 17:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94919 and previous config saved to /var/cache/conftool/dbconfig/20260720-170941-cwilliams.json * 16:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243', diff saved to https://phabricator.wikimedia.org/P94918 and previous config saved to /var/cache/conftool/dbconfig/20260720-165933-cwilliams.json * 16:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 16:58 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2053.codfw.wmnet * 16:51 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2054.codfw.wmnet * 16:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243', diff saved to https://phabricator.wikimedia.org/P94917 and previous config saved to /var/cache/conftool/dbconfig/20260720-164926-cwilliams.json * 16:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94916 and previous config saved to /var/cache/conftool/dbconfig/20260720-163918-cwilliams.json * 16:35 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 16:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94915 and previous config saved to /var/cache/conftool/dbconfig/20260720-163140-cwilliams.json * 16:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2243.codfw.wmnet with reason: Maintenance * 16:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94914 and previous config saved to /var/cache/conftool/dbconfig/20260720-163111-cwilliams.json * 16:27 btullis@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 16:27 btullis@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 16:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2020 * 16:23 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2020 * 16:21 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2020 * 16:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2020.codfw.wmnet 85.0.192.10.in-addr.arpa 5.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:21 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2020.codfw.wmnet 85.0.192.10.in-addr.arpa 5.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242', diff saved to https://phabricator.wikimedia.org/P94913 and previous config saved to /var/cache/conftool/dbconfig/20260720-162103-cwilliams.json * 16:19 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 16:18 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 16:18 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:18 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 16:17 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:17 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for netbox accounting errors - jhancock@cumin2002" * 16:17 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for netbox accounting errors - jhancock@cumin2002" * 16:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2051.codfw.wmnet * 16:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2052.codfw.wmnet * 16:11 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 16:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242', diff saved to https://phabricator.wikimedia.org/P94912 and previous config saved to /var/cache/conftool/dbconfig/20260720-161055-cwilliams.json * 16:09 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 16:08 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 16:06 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 16:06 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 16:06 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2020 * 16:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2020.codfw.wmnet with OS bookworm * 16:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94911 and previous config saved to /var/cache/conftool/dbconfig/20260720-160047-cwilliams.json * 15:58 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2019.codfw.wmnet, repooling source-only afterwards * 15:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94909 and previous config saved to /var/cache/conftool/dbconfig/20260720-155353-cwilliams.json * 15:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2242.codfw.wmnet with reason: Maintenance * 15:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94908 and previous config saved to /var/cache/conftool/dbconfig/20260720-154433-cwilliams.json * 15:35 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2049.codfw.wmnet * 15:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162', diff saved to https://phabricator.wikimedia.org/P94907 and previous config saved to /var/cache/conftool/dbconfig/20260720-153425-cwilliams.json * 15:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2050.codfw.wmnet * 15:28 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162', diff saved to https://phabricator.wikimedia.org/P94906 and previous config saved to /var/cache/conftool/dbconfig/20260720-152418-cwilliams.json * 15:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1023 * 15:14 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1023 * 15:14 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 15:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94905 and previous config saved to /var/cache/conftool/dbconfig/20260720-151407-cwilliams.json * 15:13 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] (duration: 41m 16s) * 15:08 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 15:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94902 and previous config saved to /var/cache/conftool/dbconfig/20260720-150729-cwilliams.json * 15:07 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2162.codfw.wmnet with reason: Maintenance * 15:05 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2027.codfw.wmnet, repooling source-only afterwards * 15:00 urbanecm@deploy2003: vadymts1, migr, urbanecm: Continuing with deployment * 14:59 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:58 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 07s) * 14:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:58 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 13s) * 14:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:57 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2019.codfw.wmnet, repooling source-only afterwards * 14:57 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2047.codfw.wmnet * 14:55 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2048.codfw.wmnet * 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2019.codfw.wmnet with OS bookworm * 14:49 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 14:47 urbanecm@deploy2003: vadymts1, migr, urbanecm: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:44 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool magru [reason: BGP issues in lvs7003 resolved after liberica restart, no task ID specified] * 14:44 sukhe@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool magru [reason: BGP issues in lvs7003 resolved after liberica restart, no task ID specified] * 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:41 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:39 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:39 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:33 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool magru [reason: no reason specified, no task ID specified] * 14:33 sukhe@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool magru [reason: no reason specified, no task ID specified] * 14:31 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] * 14:24 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:24 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:24 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:24 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2019.codfw.wmnet with reason: host reimage * 14:22 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2027.codfw.wmnet, repooling source-only afterwards * 14:19 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2019.codfw.wmnet with reason: host reimage * 14:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2027.codfw.wmnet with OS bookworm * 14:16 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2046.codfw.wmnet * 14:16 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2045.codfw.wmnet * 14:08 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:08 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:08 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:08 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:07 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:06 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:06 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:06 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:05 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2019 * 14:00 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2019 * 13:56 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1015.eqiad.wmnet * 13:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2027.codfw.wmnet with reason: host reimage * 13:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2071.codfw.wmnet with OS trixie * 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:51 sukhe@cumin1003: END (ERROR) - Cookbook sre.loadbalancer.admin (exit_code=97) rebooting A:liberica and P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica and P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:51 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1015.eqiad.wmnet * 13:50 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1014.eqiad.wmnet * 13:50 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1076.eqiad.wmnet with OS trixie * 13:50 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2027.codfw.wmnet with reason: host reimage * 13:45 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1014.eqiad.wmnet * 13:44 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1013.eqiad.wmnet * 13:39 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 13:38 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1013.eqiad.wmnet * 13:37 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2044.codfw.wmnet * 13:37 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2043.codfw.wmnet * 13:36 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2019 * 13:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2019.codfw.wmnet 156.32.192.10.in-addr.arpa 6.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:36 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2019.codfw.wmnet 156.32.192.10.in-addr.arpa 6.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:36 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2019 - bking@cumin2003" * 13:36 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2019 - bking@cumin2003" * 13:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry2005.codfw.wmnet * 13:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2071.codfw.wmnet with reason: host reimage * 13:31 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:31 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2019 * 13:31 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry2005.codfw.wmnet * 13:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry2004.codfw.wmnet * 13:30 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2019.codfw.wmnet with OS bookworm * 13:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2027 * 13:30 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2027 * 13:30 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2027.codfw.wmnet with OS bookworm * 13:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1076.eqiad.wmnet with reason: host reimage * 13:29 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_codfw * 13:28 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_codfw * 13:26 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry2004.codfw.wmnet * 13:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry1005.eqiad.wmnet * 13:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2071.codfw.wmnet with reason: host reimage * 13:22 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1076.eqiad.wmnet with reason: host reimage * 13:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry1005.eqiad.wmnet * 13:21 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry1004.eqiad.wmnet * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry1004.eqiad.wmnet * 13:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts * 13:13 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts * 13:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts * 13:12 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts * 13:03 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1076.eqiad.wmnet with OS trixie * 13:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2071.codfw.wmnet with OS trixie * 12:55 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:54 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:53 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:46 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 7 hosts * 12:42 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 7 hosts * 12:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts * 12:42 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts * 12:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2070.codfw.wmnet with OS trixie * 12:36 ozge@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:35 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1075.eqiad.wmnet with OS trixie * 12:32 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts * 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts * 12:22 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin1001.eqiad.wmnet * 12:19 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin1001.eqiad.wmnet * 12:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2070.codfw.wmnet with reason: host reimage * 12:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin2001.codfw.wmnet * 12:14 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1075.eqiad.wmnet with reason: host reimage * 12:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2070.codfw.wmnet with reason: host reimage * 12:10 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1075.eqiad.wmnet with reason: host reimage * 12:09 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin2001.codfw.wmnet * 11:17 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1074.eqiad.wmnet with OS trixie * 11:17 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 11:16 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 11:14 ozge@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 6 hosts * 11:09 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 6 hosts * 11:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 324 hosts * 10:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1074.eqiad.wmnet with reason: host reimage * 10:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2069.codfw.wmnet with OS trixie * 10:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1074.eqiad.wmnet with reason: host reimage * 10:30 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2069.codfw.wmnet with reason: host reimage * 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1074.eqiad.wmnet with OS trixie * 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2069.codfw.wmnet with reason: host reimage * 10:06 blake@deploy2003: Stopping before sync operations * 10:06 blake@deploy2003: Started scap sync-world: Non-deployment scap run to populate new release values * 10:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2069.codfw.wmnet with OS trixie * 10:00 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1073.eqiad.wmnet with OS trixie * 09:56 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 324 hosts * 09:39 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 09:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 8 hosts * 09:38 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1073.eqiad.wmnet with reason: host reimage * 09:37 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 8 hosts * 09:34 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1073.eqiad.wmnet with reason: host reimage * 09:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2068.codfw.wmnet with OS trixie * 09:16 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1073.eqiad.wmnet with OS trixie * 09:13 blake@deploy2003: sync-world aborted: Non-deployment scap run to populate new release values (duration: 00m 02s) * 09:13 blake@deploy2003: Started scap sync-world: Non-deployment scap run to populate new release values * 08:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2068.codfw.wmnet with reason: host reimage * 08:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2068.codfw.wmnet with reason: host reimage * 08:50 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 08:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2068.codfw.wmnet with OS trixie * 08:15 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1072.eqiad.wmnet with OS trixie * 07:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2067.codfw.wmnet with OS trixie * 07:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1072.eqiad.wmnet with reason: host reimage * 07:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1072.eqiad.wmnet with reason: host reimage * 07:45 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 07:45 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 07:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2067.codfw.wmnet with reason: host reimage * 07:35 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2067.codfw.wmnet with reason: host reimage * 07:30 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 07:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1072.eqiad.wmnet with OS trixie * 07:30 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 07:17 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 07:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2067.codfw.wmnet with OS trixie * 05:51 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:50 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:25 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:25 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on db2207.codfw.wmnet with reason: Host down * 04:28 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 07m 02s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-18 == * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 29s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 00:11 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2018.codfw.wmnet, repooling source-only afterwards == 2026-07-17 == * 23:53 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2026.codfw.wmnet, repooling source-only afterwards * 23:09 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2018.codfw.wmnet, repooling source-only afterwards * 23:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2026.codfw.wmnet, repooling source-only afterwards * 22:11 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2018.codfw.wmnet with OS bookworm * 22:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2026.codfw.wmnet with OS bookworm * 21:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2018.codfw.wmnet with reason: host reimage * 21:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2018.codfw.wmnet with reason: host reimage * 21:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2026.codfw.wmnet with reason: host reimage * 21:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2026.codfw.wmnet with reason: host reimage * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2018 * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2018 * 21:26 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2018 * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2018.codfw.wmnet 155.32.192.10.in-addr.arpa 5.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:26 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2018.codfw.wmnet 155.32.192.10.in-addr.arpa 5.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2018 - bking@cumin2003" * 21:26 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2018 - bking@cumin2003" * 21:14 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:13 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2018 * 21:13 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2018.codfw.wmnet with OS bookworm * 21:12 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2026 * 21:12 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2026 * 21:12 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2026.codfw.wmnet with OS bookworm * 21:05 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1022.eqiad.wmnet -> wdqs1026.eqiad.wmnet, repooling source-only afterwards * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs2017.codfw.wmnet, repooling source-only afterwards * 20:11 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs2017.codfw.wmnet, repooling source-only afterwards * 20:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2017.codfw.wmnet with OS bookworm * 20:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1022.eqiad.wmnet -> wdqs1026.eqiad.wmnet, repooling source-only afterwards * 20:06 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1026.eqiad.wmnet with OS bookworm * 19:55 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 19:55 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:55 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 09s) * 19:55 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:50 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 08s) * 19:50 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:50 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 10m 03s) * 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2017.codfw.wmnet with reason: host reimage * 19:40 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:40 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 15s) * 19:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1026.eqiad.wmnet with reason: host reimage * 19:37 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 16s) * 19:37 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:34 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2017.codfw.wmnet with reason: host reimage * 19:34 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1026.eqiad.wmnet with reason: host reimage * 19:33 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 19:33 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:16 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1026.eqiad.wmnet with OS bookworm * 19:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2017 * 19:16 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2017 * 19:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2017.codfw.wmnet with OS bookworm * 18:30 bking@dns1004: END - running authdns-update * 18:28 bking@dns1004: START - running authdns-update * 18:16 kamila@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1264.eqiad.wmnet * 18:16 kamila@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1264.eqiad.wmnet * 18:16 kamila@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1264.eqiad.wmnet * 17:49 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 17:46 dzahn@dns1006: END - running authdns-update * 17:44 dzahn@dns1006: START - running authdns-update * 17:44 dzahn@dns1006: END - running authdns-update * 17:42 dzahn@dns1006: START - running authdns-update * 17:28 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 17:21 kamila@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 17:01 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1264 * 17:01 kamila@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1264 * 17:01 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 17:01 kamila@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1264.eqiad.wmnet * 17:01 kamila@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1264.eqiad.wmnet * 17:01 kamila@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1264.eqiad.wmnet * 16:42 reedy@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] (duration: 10m 29s) * 16:34 reedy@deploy2003: reedy, hartman: Continuing with deployment * 16:33 reedy@deploy2003: reedy, hartman: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:31 reedy@deploy2003: Started scap sync-world: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] * 16:26 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 16:10 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in2001.wikimedia.org with reason: [[phab:T431659|T431659]] * 16:07 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in1001.wikimedia.org with reason: [[phab:T431659|T431659]] * 16:05 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 16:01 kamila@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 16:00 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out2001.wikimedia.org with reason: [[phab:T431659|T431659]] * 15:41 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 15:41 kamila@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 15:35 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out1001.wikimedia.org with reason: [[phab:T431659|T431659]] * 15:14 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1339.eqiad.wmnet * 15:13 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1339.eqiad.wmnet * 15:13 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1339.eqiad.wmnet * 14:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1339.eqiad.wmnet with OS trixie * 14:50 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:49 kamila@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:49 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:33 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage * 14:27 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage * 14:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1339 * 14:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1339 * 14:14 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1339 * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1339.eqiad.wmnet 156.32.64.10.in-addr.arpa 6.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:14 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1339.eqiad.wmnet 156.32.64.10.in-addr.arpa 6.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1339 - cgoubert@cumin2003" * 14:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1339 - cgoubert@cumin2003" * 14:09 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 14:06 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1339 * 14:06 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie * 14:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1339.eqiad.wmnet * 14:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1339.eqiad.wmnet * 14:02 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1339.eqiad.wmnet * 13:45 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb1013.eqiad.wmnet * 13:39 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb1013.eqiad.wmnet * 13:27 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:24 blake@dns1004: END - running authdns-update * 13:22 blake@dns1004: START - running authdns-update * 13:20 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 13:11 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2014.codfw.wmnet * 13:06 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb2014.codfw.wmnet * 13:06 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2012.codfw.wmnet * 13:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 13:03 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 13:01 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 15 hosts * 13:01 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb2012.codfw.wmnet * 13:01 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1016.eqiad.wmnet * 13:00 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 15 hosts * 12:55 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb1016.eqiad.wmnet * 12:55 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1014.eqiad.wmnet * 12:49 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb1014.eqiad.wmnet * 12:32 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:32 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:31 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:31 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1338.eqiad.wmnet * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1338.eqiad.wmnet * 12:18 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1338.eqiad.wmnet * 12:17 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:15 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:14 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:13 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1338.eqiad.wmnet with OS trixie * 12:01 klausman@deploy2003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 11:59 klausman@deploy2003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 11:56 klausman@deploy2003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 11:54 klausman@deploy2003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 11:53 klausman@deploy2003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 11:51 klausman@deploy2003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 11:42 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1338.eqiad.wmnet with reason: host reimage * 11:38 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1338.eqiad.wmnet with reason: host reimage * 11:31 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2230.codfw.wmnet * 11:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1338 * 11:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1338 * 11:25 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1338 * 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1338.eqiad.wmnet 155.32.64.10.in-addr.arpa 5.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:25 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1338.eqiad.wmnet 155.32.64.10.in-addr.arpa 5.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1338 - cgoubert@cumin2003" * 11:25 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1338 - cgoubert@cumin2003" * 11:23 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2230.codfw.wmnet * 11:20 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 11:20 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1338 * 11:20 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1338.eqiad.wmnet with OS trixie * 11:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1338.eqiad.wmnet * 11:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1338.eqiad.wmnet * 11:19 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1338.eqiad.wmnet * 11:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1337.eqiad.wmnet * 11:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1337.eqiad.wmnet * 11:17 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1337.eqiad.wmnet * 11:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1337.eqiad.wmnet with OS trixie * 10:51 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[2001-2002].codfw.wmnet * 10:50 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1337.eqiad.wmnet with reason: host reimage * 10:40 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 10:39 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:39 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1337.eqiad.wmnet with reason: host reimage * 10:39 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1001-1003].eqiad.wmnet * 10:34 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:34 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:30 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:28 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1001-1003].eqiad.wmnet * 10:27 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1337 * 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1337 * 10:26 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1337 * 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1337.eqiad.wmnet 154.32.64.10.in-addr.arpa 4.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:26 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1337.eqiad.wmnet 154.32.64.10.in-addr.arpa 4.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1337 - cgoubert@cumin2003" * 10:26 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1337 - cgoubert@cumin2003" * 10:21 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 10:18 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1337 * 10:17 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1337.eqiad.wmnet with OS trixie * 10:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1337.eqiad.wmnet * 10:16 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db1176.eqiad.wmnet * 10:16 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1337.eqiad.wmnet * 10:16 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1337.eqiad.wmnet * 10:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1336.eqiad.wmnet * 10:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1336.eqiad.wmnet * 10:15 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1336.eqiad.wmnet * 10:11 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db1176.eqiad.wmnet * 10:10 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db1176.eqiad.wmnet * 10:09 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db1176.eqiad.wmnet * 10:05 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts (check the cookbook's logs for more details.) * 10:03 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts (check the cookbook's logs for more details.) * 09:58 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1336.eqiad.wmnet with OS trixie * 09:47 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts (check the cookbook's logs for more details.) * 09:47 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts (check the cookbook's logs for more details.) * 09:45 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host acmechief-test2001.codfw.wmnet,acmechief-test1001.eqiad.wmnet,an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet,db-test[2001-2002].codfw.wmnet,db-test[1 * 09:40 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host acmechief-test2001.codfw.wmnet,acmechief-test1001.eqiad.wmnet,an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet,db-test[2001-2002].codfw.wmnet,db-test[1001-1003].eqiad.wmn * 09:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1336.eqiad.wmnet with reason: host reimage * 09:33 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1336.eqiad.wmnet with reason: host reimage * 09:29 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet * 09:29 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet * 09:28 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:26 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:21 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 09:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1336 * 09:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1336 * 09:19 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 09:14 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1336 * 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1336.eqiad.wmnet 152.32.64.10.in-addr.arpa 2.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1336.eqiad.wmnet 152.32.64.10.in-addr.arpa 2.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1336 - cgoubert@cumin2003" * 09:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1336 - cgoubert@cumin2003" * 09:11 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:10 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 09:09 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:09 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1336 * 09:09 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1336.eqiad.wmnet with OS trixie * 09:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1336.eqiad.wmnet * 09:08 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1336.eqiad.wmnet * 09:08 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1336.eqiad.wmnet * 09:06 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1335.eqiad.wmnet * 09:06 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1335.eqiad.wmnet * 09:06 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1335.eqiad.wmnet * 09:04 elukey: uploaded spicerack_13.1.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia * 08:55 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wikikube-worker-exp2001.codfw.wmnet * 08:54 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host testreduce1002.eqiad.wmnet * 08:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1335.eqiad.wmnet with OS trixie * 08:51 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host wikikube-worker-exp2001.codfw.wmnet * 08:51 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wikikube-worker-exp1001.eqiad.wmnet * 08:50 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host testreduce1002.eqiad.wmnet * 08:45 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host wikikube-worker-exp1001.eqiad.wmnet * 08:34 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1335.eqiad.wmnet with reason: host reimage * 08:30 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1335.eqiad.wmnet with reason: host reimage * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1335 * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1335 * 08:18 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1335 * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1335.eqiad.wmnet 150.32.64.10.in-addr.arpa 0.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:18 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1335.eqiad.wmnet 150.32.64.10.in-addr.arpa 0.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1335 - cgoubert@cumin2003" * 08:18 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1335 - cgoubert@cumin2003" * 08:14 elukey@cumin1003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 08:14 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges * 08:13 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 08:10 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1335 * 08:10 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1335.eqiad.wmnet with OS trixie * 08:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1335.eqiad.wmnet * 08:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1335.eqiad.wmnet * 08:09 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1335.eqiad.wmnet * 08:06 elukey@cumin1003: END (FAIL) - Cookbook sre.puppet.disable-merges (exit_code=99) * 08:05 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges * 08:03 elukey@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin1003.eqiad.wmnet * 07:57 elukey@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin1003.eqiad.wmnet * 07:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetdb1003.eqiad.wmnet * 07:46 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetdb1003.eqiad.wmnet * 07:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetdb2003.codfw.wmnet * 07:37 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetdb2003.codfw.wmnet * 07:37 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1001.eqiad.wmnet * 07:28 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver1001.eqiad.wmnet * 07:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet * 07:19 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet * 07:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2002.codfw.wmnet * 07:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver2002.codfw.wmnet * 07:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet * 07:05 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet * 07:04 btullis@cumin1003: END (FAIL) - Cookbook sre.hadoop.reboot-workers (exit_code=99) for Hadoop analytics cluster * 07:04 elukey@cumin1003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 07:04 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges * 06:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox1003.eqiad.wmnet * 06:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox1003.eqiad.wmnet * 02:46 ryankemper@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:46 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:44 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:37 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:37 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-internal-scholarly,name=eqiad * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 49s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 01:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore wdqs1025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling source-only afterwards * 01:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore wdqs1027 after Bookworm reimage) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs1027.eqiad.wmnet, repooling both afterwards * 00:55 urbanecm@deploy2003: helmfile [codfw] DONE helmfile.d/services/linkrecommendation: apply * 00:54 urbanecm@deploy2003: helmfile [eqiad] DONE helmfile.d/services/linkrecommendation: apply * 00:54 urbanecm@deploy2003: helmfile [staging] DONE helmfile.d/services/linkrecommendation: apply * 00:54 urbanecm@deploy2003: helmfile [codfw] START helmfile.d/services/linkrecommendation: apply * 00:53 urbanecm@deploy2003: helmfile [staging] START helmfile.d/services/linkrecommendation: apply * 00:52 urbanecm@deploy2003: helmfile [eqiad] START helmfile.d/services/linkrecommendation: apply * 00:23 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore wdqs1025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling source-only afterwards * 00:23 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore wdqs1027 after Bookworm reimage) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs1027.eqiad.wmnet, repooling both afterwards * 00:14 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1274.eqiad.wmnet * 00:14 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1274.eqiad.wmnet * 00:14 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1274.eqiad.wmnet * 00:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1027.eqiad.wmnet with OS bookworm * 00:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1025.eqiad.wmnet with OS bookworm * 00:04 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1274.eqiad.wmnet with OS trixie == 2026-07-16 == * 23:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], xfer to freshly reimaged/scap-deployed wdqs2025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs2025.codfw.wmnet, repooling source-only afterwards * 23:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1027.eqiad.wmnet with reason: host reimage * 23:47 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1025.eqiad.wmnet with reason: host reimage * 23:43 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1274.eqiad.wmnet with reason: host reimage * 23:41 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1025.eqiad.wmnet with reason: host reimage * 23:39 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1027.eqiad.wmnet with reason: host reimage * 23:38 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1274.eqiad.wmnet with reason: host reimage * 23:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1025 * 23:23 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1025 * 23:22 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1027 * 23:22 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1027 * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1274 * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1274 * 23:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1025.eqiad.wmnet with OS bookworm * 23:19 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1274 * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1274.eqiad.wmnet 145.48.64.10.in-addr.arpa 5.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:19 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1274.eqiad.wmnet 145.48.64.10.in-addr.arpa 5.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1274 - swfrench@cumin1003" * 23:19 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1274 - swfrench@cumin1003" * 23:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1027.eqiad.wmnet with OS bookworm * 23:14 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 23:14 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1274 * 23:13 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1274.eqiad.wmnet with OS trixie * 23:13 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1274.eqiad.wmnet * 23:12 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1274.eqiad.wmnet * 23:12 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1274.eqiad.wmnet * 23:12 ryankemper: [[phab:T430880|T430880]] depooled dnsdisc of wdqs-internal-scholarly-eqiad bc we only have 1 host there * 23:09 ryankemper@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-internal-scholarly,name=eqiad * 23:08 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1272.eqiad.wmnet * 23:08 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1272.eqiad.wmnet * 23:08 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1272.eqiad.wmnet * 23:01 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], xfer to freshly reimaged/scap-deployed wdqs2025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs2025.codfw.wmnet, repooling source-only afterwards * 22:57 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1272.eqiad.wmnet with OS trixie * 22:56 Amir1: deleting echo notifications from 2015 in group0 * 22:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2025.codfw.wmnet with OS bookworm * 22:35 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1272.eqiad.wmnet with reason: host reimage * 22:32 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 27s) * 22:32 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 22:28 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1269.eqiad.wmnet * 22:28 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1269.eqiad.wmnet * 22:28 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1269.eqiad.wmnet * 22:27 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1272.eqiad.wmnet with reason: host reimage * 22:26 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] (duration: 08m 51s) * 22:22 ladsgroup@deploy2003: ladsgroup, urbanecm: Continuing with deployment * 22:19 ladsgroup@deploy2003: ladsgroup, urbanecm: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:17 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] * 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2025.codfw.wmnet with reason: host reimage * 22:06 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1272 * 22:06 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1272 * 22:05 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1272 * 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1272.eqiad.wmnet 127.48.64.10.in-addr.arpa 7.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:05 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1272.eqiad.wmnet 127.48.64.10.in-addr.arpa 7.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1272 - swfrench@cumin1003" * 22:05 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1272 - swfrench@cumin1003" * 22:03 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2025.codfw.wmnet with reason: host reimage * 22:01 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 22:00 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1272 * 22:00 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1272.eqiad.wmnet with OS trixie * 22:00 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1272.eqiad.wmnet * 21:59 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1272.eqiad.wmnet * 21:59 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1272.eqiad.wmnet * 21:55 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1271.eqiad.wmnet * 21:55 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1271.eqiad.wmnet * 21:55 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1271.eqiad.wmnet * 21:46 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1271.eqiad.wmnet with OS trixie * 21:45 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] (duration: 06m 31s) * 21:40 sbassett@deploy2003: sbassett: Continuing with deployment * 21:40 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2025 * 21:40 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2025 * 21:40 sbassett@deploy2003: sbassett: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:38 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] * 21:37 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2025 * 21:37 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2025.codfw.wmnet 220.48.192.10.in-addr.arpa 0.2.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:37 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2025.codfw.wmnet 220.48.192.10.in-addr.arpa 0.2.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:37 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:37 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2025 - bking@cumin2003" * 21:37 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2025 - bking@cumin2003" * 21:30 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] (duration: 08m 19s) * 21:26 sbassett@deploy2003: sbassett: Continuing with deployment * 21:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1269.eqiad.wmnet with OS trixie * 21:24 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1271.eqiad.wmnet with reason: host reimage * 21:23 sbassett@deploy2003: sbassett: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:22 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:22 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] * 21:20 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2025 * 21:19 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2025.codfw.wmnet with OS bookworm * 21:17 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1271.eqiad.wmnet with reason: host reimage * 21:04 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1269.eqiad.wmnet with reason: host reimage * 21:00 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1269.eqiad.wmnet with reason: host reimage * 20:56 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1271 * 20:55 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1271 * 20:54 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1271 * 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1271.eqiad.wmnet 126.48.64.10.in-addr.arpa 6.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:54 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1271.eqiad.wmnet 126.48.64.10.in-addr.arpa 6.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1271 - swfrench@cumin1003" * 20:54 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1271 - swfrench@cumin1003" * 20:51 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:51 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:51 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:50 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 20:49 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 20:49 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1268.eqiad.wmnet * 20:49 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1268.eqiad.wmnet * 20:49 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1268.eqiad.wmnet * 20:48 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1271 * 20:48 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1271.eqiad.wmnet with OS trixie * 20:47 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1271.eqiad.wmnet * 20:46 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1271.eqiad.wmnet * 20:46 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1271.eqiad.wmnet * 20:41 aude@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] (duration: 07m 34s) * 20:39 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1269 * 20:39 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1269 * 20:36 aude@deploy2003: aude: Continuing with deployment * 20:35 aude@deploy2003: aude: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:33 aude@deploy2003: Started scap sync-world: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] * 20:26 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on db2207.codfw.wmnet with reason: Host down * 20:22 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-video: apply * 20:21 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-video: apply * 20:20 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-timeline: apply * 20:20 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-timeline: apply * 20:20 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-syntaxhighlight: apply * 20:19 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-syntaxhighlight: apply * 20:19 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-media: apply * 20:18 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-media: apply * 20:18 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-constraints: apply * 20:17 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-constraints: apply * 20:17 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox: apply * 20:16 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox: apply * 20:13 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1269 * 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1269.eqiad.wmnet 80.32.64.10.in-addr.arpa 0.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:13 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1269.eqiad.wmnet 80.32.64.10.in-addr.arpa 0.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1269 - kamila@cumin1003" * 20:13 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1269 - kamila@cumin1003" * 20:09 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-flink-codfw cluster: Roll restart of jvm daemons. * 20:07 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 20:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 20:03 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-flink-codfw cluster: Roll restart of jvm daemons. * 20:03 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2207 [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94893 and previous config saved to /var/cache/conftool/dbconfig/20260716-200257-marostegui.json * 20:01 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2204 to s2 primary [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94892 and previous config saved to /var/cache/conftool/dbconfig/20260716-200157-marostegui.json * 20:00 marostegui: Starting emergency s2 codfw failover from db2207 to db2204 - [[phab:T432396|T432396]] * 19:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1035.eqiad.wmnet * 19:56 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2204 with weight 0 [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94891 and previous config saved to /var/cache/conftool/dbconfig/20260716-195628-marostegui.json * 19:55 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 26 hosts with reason: Primary switchover s2 [[phab:T432396|T432396]] * 19:54 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1035.eqiad.wmnet * 19:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1034.eqiad.wmnet * 19:48 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1034.eqiad.wmnet * 19:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1033.eqiad.wmnet * 19:43 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-video: apply * 19:43 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1033.eqiad.wmnet * 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1032.eqiad.wmnet * 19:42 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-video: apply * 19:42 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-timeline: apply * 19:41 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-timeline: apply * 19:41 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-syntaxhighlight: apply * 19:41 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-syntaxhighlight: apply * 19:40 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-media: apply * 19:40 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-media: apply * 19:39 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-constraints: apply * 19:36 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-constraints: apply * 19:36 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox: apply * 19:35 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1032.eqiad.wmnet * 19:35 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1031.eqiad.wmnet * 19:35 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox: apply * 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-video: apply * 19:33 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-video: apply * 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-timeline: apply * 19:33 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-timeline: apply * 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-syntaxhighlight: apply * 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-syntaxhighlight: apply * 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-media: apply * 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-media: apply * 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-constraints: apply * 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-constraints: apply * 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox: apply * 19:31 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox: apply * 19:27 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1031.eqiad.wmnet * 19:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1030.eqiad.wmnet * 19:23 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1001.eqiad.wmnet, repooling source-only afterwards * 19:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1030.eqiad.wmnet * 19:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1029.eqiad.wmnet * 19:17 kamila@cumin1003: START - Cookbook sre.dns.netbox * 19:12 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1029.eqiad.wmnet * 19:06 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1269 * 19:05 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1269.eqiad.wmnet with OS trixie * 19:03 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1269.eqiad.wmnet * 19:03 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1269.eqiad.wmnet * 19:03 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1269.eqiad.wmnet * 18:55 dancy@deploy2003: Finished scap sync-world: testing [[phab:T428971|T428971]] (duration: 02m 41s) * 18:53 dancy@deploy2003: Started scap sync-world: testing [[phab:T428971|T428971]] * 18:31 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1268.eqiad.wmnet with OS trixie * 18:18 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 18:16 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1267.eqiad.wmnet * 18:16 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1267.eqiad.wmnet * 18:16 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1267.eqiad.wmnet * 18:09 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1268.eqiad.wmnet with reason: host reimage * 18:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1001.eqiad.wmnet, repooling source-only afterwards * 18:06 swfrench@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] (duration: 07m 34s) * 18:06 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 23s) * 18:06 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 18:06 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1268.eqiad.wmnet with reason: host reimage * 18:03 bd808@deploy2003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 18:02 bd808@deploy2003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 18:02 swfrench@deploy2003: jiji, swfrench: Continuing with deployment * 18:02 bd808@deploy2003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 18:02 bd808@deploy2003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 18:01 bd808@deploy2003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 18:01 swfrench@deploy2003: jiji, swfrench: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:01 bd808@deploy2003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:59 swfrench@deploy2003: Started scap sync-world: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] * 17:45 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1268 * 17:45 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1268 * 17:44 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1267.eqiad.wmnet with OS trixie * 17:43 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter2006.codfw.wmnet * 17:39 swfrench@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter2006.codfw.wmnet * 17:35 swfrench@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] (duration: 07m 27s) * 17:34 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1268 * 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1268.eqiad.wmnet 78.32.64.10.in-addr.arpa 8.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:34 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1268.eqiad.wmnet 78.32.64.10.in-addr.arpa 8.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1268 - kamila@cumin1003" * 17:34 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1268 - kamila@cumin1003" * 17:31 swfrench@deploy2003: jiji, swfrench: Continuing with deployment * 17:29 swfrench@deploy2003: jiji, swfrench: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:28 kamila@cumin1003: START - Cookbook sre.dns.netbox * 17:28 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1268 * 17:28 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1268.eqiad.wmnet with OS trixie * 17:27 swfrench@deploy2003: Started scap sync-world: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] * 17:23 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1267.eqiad.wmnet with reason: host reimage * 17:18 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1267.eqiad.wmnet with reason: host reimage * 17:18 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1268.eqiad.wmnet * 17:17 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1268.eqiad.wmnet * 17:17 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1268.eqiad.wmnet * 17:12 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter2005.codfw.wmnet * 17:11 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1270.eqiad.wmnet * 17:11 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1270.eqiad.wmnet * 17:11 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1270.eqiad.wmnet * 17:09 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter2005.codfw.wmnet * 17:08 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:08 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update reverse dns for moved arelion cct cr2-eqiad - cmooney@cumin1003" * 17:08 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update reverse dns for moved arelion cct cr2-eqiad - cmooney@cumin1003" * 17:08 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] (duration: 07m 34s) * 17:04 jiji@deploy2003: jiji: Continuing with deployment * 17:03 jiji@deploy2003: jiji: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 17:00 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] * 17:00 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2185.codfw.wmnet with OS trixie * 16:59 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:58 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1270.eqiad.wmnet with OS trixie * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1267 * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1267 * 16:57 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1267 * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1267.eqiad.wmnet 77.32.64.10.in-addr.arpa 7.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:57 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1267.eqiad.wmnet 77.32.64.10.in-addr.arpa 7.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1267 - kamila@cumin1003" * 16:56 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1267 - kamila@cumin1003" * 16:56 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_eqsin * 16:56 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5032.eqsin.wmnet * 16:52 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_esams * 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3073.esams.wmnet * 16:50 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_esams * 16:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3081.esams.wmnet * 16:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1266.eqiad.wmnet * 16:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1266.eqiad.wmnet * 16:45 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1266.eqiad.wmnet * 16:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2185.codfw.wmnet with reason: host reimage * 16:41 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_eqiad * 16:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1114.eqiad.wmnet * 16:41 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_eqiad * 16:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1115.eqiad.wmnet * 16:39 kamila@cumin1003: START - Cookbook sre.dns.netbox * 16:39 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1267 * 16:39 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2185.codfw.wmnet with reason: host reimage * 16:38 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1267.eqiad.wmnet with OS trixie * 16:38 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1267.eqiad.wmnet * 16:38 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1270.eqiad.wmnet with reason: host reimage * 16:37 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1267.eqiad.wmnet * 16:37 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1267.eqiad.wmnet * 16:31 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1270.eqiad.wmnet with reason: host reimage * 16:24 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1264.eqiad.wmnet * 16:24 kamila@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 16:24 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 16:23 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 16:21 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2162: switch maintenance completed codfw rack b6 * 16:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2185.codfw.wmnet with OS trixie * 16:19 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 16:16 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter1007.eqiad.wmnet * 16:15 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_eqsin * 16:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5024.eqsin.wmnet * 16:13 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5031.eqsin.wmnet * 16:13 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3072.esams.wmnet * 16:12 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter1007.eqiad.wmnet * 16:11 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] (duration: 09m 47s) * 16:10 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1266.eqiad.wmnet with OS trixie * 16:10 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1270 * 16:10 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1270 * 16:09 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1270 * 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1270.eqiad.wmnet 125.48.64.10.in-addr.arpa 5.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1270.eqiad.wmnet 125.48.64.10.in-addr.arpa 5.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1270 - swfrench@cumin1003" * 16:09 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1270 - swfrench@cumin1003" * 16:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3080.esams.wmnet * 16:07 jiji@deploy2003: jiji: Continuing with deployment * 16:06 jiji@deploy2003: jiji: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:04 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 16:04 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1265.eqiad.wmnet * 16:03 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1265.eqiad.wmnet * 16:03 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1265.eqiad.wmnet * 16:03 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1270 * 16:03 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1270.eqiad.wmnet with OS trixie * 16:02 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1270.eqiad.wmnet * 16:02 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] * 16:01 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1270.eqiad.wmnet * 16:01 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1270.eqiad.wmnet * 16:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1113.eqiad.wmnet * 16:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1112.eqiad.wmnet * 15:49 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1266.eqiad.wmnet with reason: host reimage * 15:47 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter1006.eqiad.wmnet * 15:45 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1265.eqiad.wmnet with OS trixie * 15:44 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1266.eqiad.wmnet with reason: host reimage * 15:43 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter1006.eqiad.wmnet * 15:42 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] (duration: 09m 46s) * 15:37 jiji@deploy2003: jiji: Continuing with deployment * 15:36 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2162: switch maintenance completed codfw rack b6 * 15:36 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2161: switch maintenance completed codfw rack b6 * 15:34 jiji@deploy2003: jiji: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:32 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5023.eqsin.wmnet * 15:32 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] * 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5030.eqsin.wmnet * 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3071.esams.wmnet * 15:27 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3079.esams.wmnet * 15:25 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1265.eqiad.wmnet with reason: host reimage * 15:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1266 * 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1266 * 15:21 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1110.eqiad.wmnet * 15:20 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1111.eqiad.wmnet * 15:16 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1265.eqiad.wmnet with reason: host reimage * 15:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1001.eqiad.wmnet with OS bookworm * 15:15 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1266 * 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1266.eqiad.wmnet 76.32.64.10.in-addr.arpa 6.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:15 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1266.eqiad.wmnet 76.32.64.10.in-addr.arpa 6.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1266 - kamila@cumin1003" * 15:15 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1266 - kamila@cumin1003" * 15:07 kamila@cumin1003: START - Cookbook sre.dns.netbox * 15:04 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1266 * 15:04 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1264 * 15:04 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1264 * 15:04 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1266.eqiad.wmnet with OS trixie * 15:03 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1264 * 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1264.eqiad.wmnet 74.32.64.10.in-addr.arpa 4.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:03 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1264.eqiad.wmnet 74.32.64.10.in-addr.arpa 4.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1264 - kamila@cumin1003" * 15:03 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1264 - kamila@cumin1003" * 15:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-eqiad * 15:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp1001.eqiad.wmnet * 15:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp1001.eqiad.wmnet * 15:01 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp1001.eqiad.wmnet * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp1001.eqiad.wmnet * 15:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1376-1384].eqiad.wmnet * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1376-1384].eqiad.wmnet * 14:59 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1002.eqiad.wmnet * 14:59 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1266.eqiad.wmnet * 14:58 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1266.eqiad.wmnet * 14:58 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1266.eqiad.wmnet * 14:58 kamila@cumin1003: START - Cookbook sre.dns.netbox * 14:57 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1264 * 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1265 * 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1265 * 14:57 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1310596{{!}}Set $wgMathInternalRestbaseURL explicitly (T349582)]] * 14:57 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1265 * 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1265.eqiad.wmnet 75.32.64.10.in-addr.arpa 5.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:56 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1265.eqiad.wmnet 75.32.64.10.in-addr.arpa 5.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:56 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:56 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1265 - kamila@cumin1003" * 14:56 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1265 - kamila@cumin1003" * 14:53 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1002.eqiad.wmnet * 14:53 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1376-1384].eqiad.wmnet * 14:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1001.eqiad.wmnet with reason: host reimage * 14:51 kamila@cumin1003: START - Cookbook sre.dns.netbox * 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 14:50 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:50 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2161: switch maintenance completed codfw rack b6 * 14:50 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5021.eqsin.wmnet * 14:50 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1265 * 14:49 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:49 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:49 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1265.eqiad.wmnet with OS trixie * 14:49 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5029.eqsin.wmnet * 14:49 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1265.eqiad.wmnet * 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3070.esams.wmnet * 14:48 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1264.eqiad.wmnet * 14:48 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1001.eqiad.wmnet with reason: host reimage * 14:48 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1376-1384].eqiad.wmnet * 14:48 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1265.eqiad.wmnet * 14:47 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1265.eqiad.wmnet * 14:47 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1264.eqiad.wmnet * 14:47 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1264.eqiad.wmnet * 14:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:47 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3078.esams.wmnet * 14:44 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1263.eqiad.wmnet * 14:44 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1263.eqiad.wmnet * 14:44 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1263.eqiad.wmnet * 14:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1108.eqiad.wmnet * 14:40 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1109.eqiad.wmnet * 14:40 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:35 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:34 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:34 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2006.codfw.wmnet * 14:34 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-flink-eqiad cluster: Roll restart of jvm daemons. * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf2002.codfw.wmnet * 14:31 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1002.eqiad.wmnet * 14:29 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2006.codfw.wmnet * 14:27 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:27 kamila@deploy2003: Finished scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] (duration: 02m 57s) * 14:27 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-flink-eqiad cluster: Roll restart of jvm daemons. * 14:26 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf2002.codfw.wmnet * 14:26 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf2001.codfw.wmnet * 14:25 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1002.eqiad.wmnet * 14:25 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1001.eqiad.wmnet * 14:25 kamila@deploy2003: Started scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] * 14:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:21 kamila@deploy2003: sync-world aborted: Test deployment to check rsync is working - [[phab:T432108|T432108]] (duration: 00m 36s) * 14:21 topranks: reboot lsw1-b6-codfw to upgrade JunOS [[phab:T430922|T430922]] * 14:21 kamila@deploy2003: Started scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] * 14:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1001.eqiad.wmnet with OS bookworm * 14:20 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b6-codfw,lsw1-b6-codfw IPv6,lsw1-b6-codfw.mgmt,ssw1-a[1,8]-codfw with reason: lsw1-b6-codfw JunOS upgrade * 14:20 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf2001.codfw.wmnet * 14:19 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1001.eqiad.wmnet * 14:19 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 26 hosts with reason: lsw1-b6-codfw JunOS upgrade * 14:14 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 14:13 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc2022: switch maintenance codfw rack b6 * 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:12 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.parsercache * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool pc2022: switch maintenance codfw rack b6 * 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2251: switch maintenance codfw rack b6 * 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.parsercache * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2251: switch maintenance codfw rack b6 * 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2162: switch maintenance codfw rack b6 * 14:12 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1263.eqiad.wmnet with OS trixie * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2162: switch maintenance codfw rack b6 * 14:11 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2161: switch maintenance codfw rack b6 * 14:11 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2161: switch maintenance codfw rack b6 * 14:08 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5020.eqsin.wmnet * 14:07 btullis@cumin1003: START - Cookbook sre.hadoop.reboot-workers for Hadoop analytics cluster * 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3069.esams.wmnet * 14:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5028.eqsin.wmnet * 14:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1338-1347].eqiad.wmnet * 14:06 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1338-1347].eqiad.wmnet * 14:05 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3077.esams.wmnet * 14:02 topranks: beginning depools for lsw1-b6-codfw maintenance [[phab:T430922|T430922]] * 14:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1106.eqiad.wmnet * 14:00 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-misc1002.eqiad.wmnet * 13:59 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1338-1347].eqiad.wmnet * 13:59 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1107.eqiad.wmnet * 13:56 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-codfw * 13:55 sfaci@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply * 13:54 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-misc1002.eqiad.wmnet * 13:54 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-misc1001.eqiad.wmnet * 13:54 sfaci@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply * 13:50 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1263.eqiad.wmnet with reason: host reimage * 13:50 sfaci@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 13:49 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1338-1347].eqiad.wmnet * 13:49 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-misc1001.eqiad.wmnet * 13:49 sfaci@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 13:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:49 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:45 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1263.eqiad.wmnet with reason: host reimage * 13:40 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:40 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-eqiad * 13:35 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:34 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:33 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-reboot (exit_code=0) rolling reboot on A:dnsbox and (A:eqsin or A:drmrs or A:magru) and not (P<nowiki>{</nowiki>dns5003*<nowiki>}</nowiki> or P<nowiki>{</nowiki>dns7002*<nowiki>}</nowiki>) and (A:dnsbox) * 13:33 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns7001.wikimedia.org * 13:27 sukhe@dns1004: END - running authdns-update * 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5019.eqsin.wmnet * 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3076.esams.wmnet * 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3068.esams.wmnet * 13:25 sukhe@dns1004: START - running authdns-update * 13:24 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5027.eqsin.wmnet * 13:24 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1263 * 13:24 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1263 * 13:23 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1263 * 13:23 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:23 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1104.eqiad.wmnet * 13:21 kamila@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:21 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:20 kamila@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:20 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:20 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:20 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1263 - kamila@cumin1003" * 13:20 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1263 - kamila@cumin1003" * 13:19 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1105.eqiad.wmnet * 13:19 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:19 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:18 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:18 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns7001.wikimedia.org * 13:16 cdobbins@cumin2003: conftool action : set/pooled=yes; selector: name=dns7002.* * 13:14 cdobbins@dns1004: END - running authdns-update * 13:13 sbisson@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] (duration: 08m 03s) * 13:13 cdobbins@dns1004: START - running authdns-update * 13:12 kamila@cumin1003: START - Cookbook sre.dns.netbox * 13:12 cdobbins@cumin2003: conftool action : set/pooled=yes; selector: name=dns7002.*,service=authdns-update * 13:12 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1263 * 13:11 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1263.eqiad.wmnet with OS trixie * 13:11 cdobbins@cumin2003: conftool action : set/pooled=no; selector: name=dns7002.* * 13:11 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1263.eqiad.wmnet * 13:10 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1263.eqiad.wmnet * 13:10 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1263.eqiad.wmnet * 13:09 sbisson@deploy2003: sbisson: Continuing with deployment * 13:08 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:07 sbisson@deploy2003: sbisson: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:05 sbisson@deploy2003: Started scap sync-world: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] * 13:03 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns6002.wikimedia.org * 13:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1298-1307].eqiad.wmnet * 13:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1298-1307].eqiad.wmnet * 12:59 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1262.eqiad.wmnet * 12:59 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1262.eqiad.wmnet * 12:59 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1262.eqiad.wmnet * 12:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1298-1307].eqiad.wmnet * 12:49 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns6002.wikimedia.org * 12:46 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1298-1307].eqiad.wmnet * 12:46 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:46 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3075.esams.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3067.esams.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5018.eqsin.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5026.eqsin.wmnet * 12:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1102.eqiad.wmnet * 12:39 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1103.eqiad.wmnet * 12:35 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:34 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns6001.wikimedia.org * 12:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:28 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:18 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns6001.wikimedia.org * 12:14 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1267-1276].eqiad.wmnet * 12:13 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1267-1276].eqiad.wmnet * 12:04 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1267-1276].eqiad.wmnet * 12:03 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns5004.wikimedia.org * 12:02 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1100.eqiad.wmnet * 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3066.esams.wmnet * 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3074.esams.wmnet * 12:01 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:01 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-codfw * 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5017.eqsin.wmnet * 12:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5025.eqsin.wmnet * 12:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1101.eqiad.wmnet * 11:59 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1267-1276].eqiad.wmnet * 11:58 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:58 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:54 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns5004.wikimedia.org * 11:54 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and (A:eqsin or A:drmrs or A:magru) and not (P<nowiki>{</nowiki>dns5003*<nowiki>}</nowiki> or P<nowiki>{</nowiki>dns7002*<nowiki>}</nowiki>) and (A:dnsbox) * 11:54 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:53 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:53 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-eqiad * 11:51 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:51 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:50 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:50 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_eqiad * 11:50 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:50 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_eqiad * 11:49 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_esams * 11:49 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_esams * 11:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_eqsin * 11:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_eqsin * 11:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:44 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-eqiad * 11:43 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-codfw * 11:42 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:41 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:24 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-eqiad * 11:23 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-codfw * 11:22 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:15 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:14 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:09 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2066.codfw.wmnet with OS trixie * 11:05 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1151-1160].eqiad.wmnet * 11:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1151-1160].eqiad.wmnet * 10:59 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1068.eqiad.wmnet with OS trixie * 10:55 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.major-upgrade (exit_code=99) * 10:55 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 10:54 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1151-1160].eqiad.wmnet * 10:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2066.codfw.wmnet with reason: host reimage * 10:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1151-1160].eqiad.wmnet * 10:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:42 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2066.codfw.wmnet with reason: host reimage * 10:39 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:37 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:36 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:23 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:22 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2066.codfw.wmnet with OS trixie * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:07 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2065.codfw.wmnet with OS trixie * 10:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:06 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 10:06 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 10:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 10:03 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:03 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 10:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 09:59 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 09:57 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 09:57 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 09:52 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 09:47 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:46 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:46 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2065.codfw.wmnet with reason: host reimage * 09:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2065.codfw.wmnet with reason: host reimage * 09:40 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:39 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 09:39 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:39 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:39 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:37 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:29 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox2003.codfw.wmnet * 09:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox2003.codfw.wmnet * 09:25 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:25 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:24 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:24 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:24 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:21 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:20 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2065.codfw.wmnet with OS trixie * 09:13 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1068.eqiad.wmnet with OS trixie * 09:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2064.codfw.wmnet with OS trixie * 09:08 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 09:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 09:07 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2162: Repooling after switchover * 09:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1067.eqiad.wmnet with OS trixie * 09:00 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:59 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 08:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 08:57 tappof: bump space for prometheus k8s-dse in eqiad * 08:56 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping2004.codfw.wmnet * 08:52 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host ping2004.codfw.wmnet * 08:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 08:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping1004.eqiad.wmnet * 08:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:51 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:49 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 08:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host ping1004.eqiad.wmnet * 08:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2064.codfw.wmnet with reason: host reimage * 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2064.codfw.wmnet with reason: host reimage * 08:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1067.eqiad.wmnet with reason: host reimage * 08:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:33 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1067.eqiad.wmnet with reason: host reimage * 08:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:21 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2162: Repooling after switchover * 08:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2064.codfw.wmnet with OS trixie * 08:16 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1067.eqiad.wmnet with OS trixie * 08:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:15 cgoubert@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-eqiad * 08:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2062.codfw.wmnet with OS trixie * 08:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1066.eqiad.wmnet with OS trixie * 08:02 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2162: Repooling after switchover * 07:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2162: Repooling after switchover * 07:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2162 [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94870 and previous config saved to /var/cache/conftool/dbconfig/20260716-075530-cwilliams.json * 07:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2241 to x3 primary [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94869 and previous config saved to /var/cache/conftool/dbconfig/20260716-075314-cwilliams.json * 07:52 cezmunsta: Starting x3 codfw failover from db2162 to db2241 - [[phab:T430925|T430925]] * 07:50 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:50 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:47 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 07:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2241 with weight 0 [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94868 and previous config saved to /var/cache/conftool/dbconfig/20260716-074507-cwilliams.json * 07:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 18 hosts with reason: Primary switchover x3 [[phab:T430925|T430925]] * 07:43 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1066.eqiad.wmnet with reason: host reimage * 07:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:dse-k8s-worker-eqiad * 07:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1028.eqiad.wmnet * 07:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1028.eqiad.wmnet * 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 07:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1066.eqiad.wmnet with reason: host reimage * 07:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1028.eqiad.wmnet * 07:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1028.eqiad.wmnet * 07:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1027.eqiad.wmnet * 07:35 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1027.eqiad.wmnet * 07:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1027.eqiad.wmnet * 07:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1027.eqiad.wmnet * 07:28 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1026.eqiad.wmnet * 07:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1026.eqiad.wmnet * 07:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast2003.wikimedia.org * 07:21 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1026.eqiad.wmnet * 07:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1066.eqiad.wmnet with OS trixie * 07:19 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast2003.wikimedia.org * 07:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2062.codfw.wmnet with OS trixie * 06:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1026.eqiad.wmnet * 06:51 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1025.eqiad.wmnet * 06:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1025.eqiad.wmnet * 06:47 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 06:47 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 06:44 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1025.eqiad.wmnet * 06:14 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1025.eqiad.wmnet * 06:14 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1024.eqiad.wmnet * 06:14 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1024.eqiad.wmnet * 06:07 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1024.eqiad.wmnet * 05:37 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1024.eqiad.wmnet * 05:37 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1023.eqiad.wmnet * 05:37 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1023.eqiad.wmnet * 05:26 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1023.eqiad.wmnet * 04:56 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1023.eqiad.wmnet * 04:56 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1022.eqiad.wmnet * 04:56 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1022.eqiad.wmnet * 04:49 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1022.eqiad.wmnet * 04:19 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1022.eqiad.wmnet * 04:19 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1021.eqiad.wmnet * 04:19 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1021.eqiad.wmnet * 04:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1021.eqiad.wmnet * 03:38 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1021.eqiad.wmnet * 03:38 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1020.eqiad.wmnet * 03:38 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1020.eqiad.wmnet * 03:20 btullis@cumin1003: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1020.eqiad.wmnet * 03:18 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1020.eqiad.wmnet * 03:18 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1019.eqiad.wmnet * 03:18 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1019.eqiad.wmnet * 03:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1019.eqiad.wmnet * 02:41 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1019.eqiad.wmnet * 02:41 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1018.eqiad.wmnet * 02:41 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1018.eqiad.wmnet * 02:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling both afterwards * 02:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2003.codfw.wmnet -> wcqs2001.codfw.wmnet, repooling both afterwards * 02:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1018.eqiad.wmnet * 02:30 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1018.eqiad.wmnet * 02:30 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1014.eqiad.wmnet * 02:30 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1014.eqiad.wmnet * 02:24 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1014.eqiad.wmnet * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 01:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1014.eqiad.wmnet * 01:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1013.eqiad.wmnet * 01:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1013.eqiad.wmnet * 01:47 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1013.eqiad.wmnet * 01:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2003.codfw.wmnet -> wcqs2001.codfw.wmnet, repooling both afterwards * 01:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling both afterwards * 01:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1013.eqiad.wmnet * 01:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1012.eqiad.wmnet * 01:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1012.eqiad.wmnet * 01:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1012.eqiad.wmnet * 01:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1012.eqiad.wmnet * 01:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1011.eqiad.wmnet * 01:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1011.eqiad.wmnet * 01:08 ryankemper@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] scap deploy post bookworm reimage (duration: 00m 23s) * 01:08 ryankemper@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] scap deploy post bookworm reimage * 01:08 ryankemper@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): scap deploy post bookworm reimage (duration: 00m 46s) * 01:07 ryankemper@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): scap deploy post bookworm reimage * 01:04 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1011.eqiad.wmnet * 01:04 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1011.eqiad.wmnet * 01:04 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1010.eqiad.wmnet * 01:04 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1010.eqiad.wmnet * 00:57 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1010.eqiad.wmnet * 00:57 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1010.eqiad.wmnet * 00:57 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1009.eqiad.wmnet * 00:57 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1009.eqiad.wmnet * 00:50 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1009.eqiad.wmnet * 00:20 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1009.eqiad.wmnet * 00:20 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1008.eqiad.wmnet * 00:20 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1008.eqiad.wmnet * 00:13 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1008.eqiad.wmnet == 2026-07-15 == * 23:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2001.codfw.wmnet with OS bookworm * 23:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1008.eqiad.wmnet * 23:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1007.eqiad.wmnet * 23:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1007.eqiad.wmnet * 23:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1007.eqiad.wmnet * 23:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1007.eqiad.wmnet * 23:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1006.eqiad.wmnet * 23:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1006.eqiad.wmnet * 23:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1006.eqiad.wmnet * 23:29 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1006.eqiad.wmnet * 23:28 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1005.eqiad.wmnet * 23:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1005.eqiad.wmnet * 23:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1002.eqiad.wmnet with OS bookworm * 23:21 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1005.eqiad.wmnet * 23:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2001.codfw.wmnet with reason: host reimage * 23:15 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host datahubsearch1001.eqiad.wmnet with OS bookworm * 23:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2001.codfw.wmnet with reason: host reimage * 23:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 23:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 22:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 22:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1005.eqiad.wmnet * 22:51 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1004.eqiad.wmnet * 22:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1004.eqiad.wmnet * 22:45 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1004.eqiad.wmnet * 22:44 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 22:44 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS trixie * 22:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host datahubsearch1001.eqiad.wmnet with OS bookworm * 22:34 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host datahubsearch1001.eqiad.wmnet with OS bookworm * 22:16 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on datahubsearch[1002-1003].eqiad.wmnet with reason: Using datahubsearch1001 to test bookworm reimages * 22:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1004.eqiad.wmnet * 22:15 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1003.eqiad.wmnet * 22:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1003.eqiad.wmnet * 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 22:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1003.eqiad.wmnet * 22:08 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1003.eqiad.wmnet * 22:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1002.eqiad.wmnet * 22:08 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1002.eqiad.wmnet * 22:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host datahubsearch1001.eqiad.wmnet with OS bookworm * 22:05 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 22:02 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm * 22:01 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on datahubsearch[1001-1003].eqiad.wmnet with reason: Using datahubsearch1001 to test bookworm reimages * 22:01 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1002.eqiad.wmnet * 22:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1002.eqiad.wmnet * 22:00 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1001.eqiad.wmnet * 22:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1001.eqiad.wmnet * 21:53 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1001.eqiad.wmnet * 21:52 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 21:50 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 21:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS trixie * 21:50 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS bookworm * 21:43 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 21:38 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:30 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wcqs1002'] * 21:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:29 lerickson@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 21:29 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:29 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:29 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS bookworm * 21:28 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 21:28 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm * 21:23 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1001.eqiad.wmnet * 21:23 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:23 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:22 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 21:20 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:18 swfrench-wmf: reprepro include php8.3_8.3.32-1+wmf11u2 into component/php83 for bullseye-wikimedia * 21:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:16 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:15 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-druid-public cluster: Roll restart of jvm daemons. * 21:08 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:05 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1001.eqiad.wmnet * 21:05 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1001.eqiad.wmnet * 21:04 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-druid-public cluster: Roll restart of jvm daemons. * 21:02 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 21:01 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 21:01 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 21:00 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 20:59 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1001.eqiad.wmnet * 20:59 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1001.eqiad.wmnet * 20:59 btullis@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:dse-k8s-worker-eqiad * 20:55 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 20:55 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 20:45 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm * 20:21 jhathaway: puppet is re-enabled, have fun, but not too much fun! * 20:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 20:17 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs2001'] * 20:12 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs2001'] * 20:11 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs2001'] * 20:09 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:08 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 20:05 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:05 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 20:04 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs2001'] * 20:03 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 20:03 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm * 20:02 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:02 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 20:01 jhathaway: disabling puppet fleet wide to roll out kafka patch * 19:55 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host relforge1010.eqiad.wmnet * 19:52 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 19:52 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 19:48 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 19:48 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 19:48 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 19:47 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 19:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 19:45 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 19:45 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 19:44 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1010.eqiad.wmnet * 19:38 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1262.eqiad.wmnet with OS trixie * 19:17 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1262.eqiad.wmnet with reason: host reimage * 19:11 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1262.eqiad.wmnet with reason: host reimage * 18:59 cdobbins@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS trixie * 18:54 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 18:53 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 18:52 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1262 * 18:52 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1262 * 18:51 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1262 * 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1262.eqiad.wmnet 72.32.64.10.in-addr.arpa 2.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:51 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1262.eqiad.wmnet 72.32.64.10.in-addr.arpa 2.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1262 - kamila@cumin1003" * 18:51 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1262 - kamila@cumin1003" * 18:46 kamila@cumin1003: START - Cookbook sre.dns.netbox * 18:46 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1262 * 18:46 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 18:46 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ncmonitor1001.eqiad.wmnet * 18:46 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 18:45 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1262.eqiad.wmnet with OS trixie * 18:45 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 18:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1262.eqiad.wmnet * 18:44 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1262.eqiad.wmnet * 18:44 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1262.eqiad.wmnet * 18:42 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host ncmonitor1001.eqiad.wmnet * 18:29 topranks: pull power on cr1-eqiad to install new switch-control boards [[phab:T426343|T426343]] * 18:29 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs[1018-1020].eqiad.wmnet with reason: line card install in cr1-eqiad * 18:27 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 14 hosts with reason: linecard install in cr1-eqad * 18:22 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_ulsfo * 18:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4052.ulsfo.wmnet * 18:19 cdobbins@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 18:15 cdobbins@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 18:14 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_drmrs * 18:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6016.drmrs.wmnet * 18:12 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_ulsfo * 18:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4044.ulsfo.wmnet * 18:10 sukhe@cumin1003: END (ERROR) - Cookbook sre.cdn.roll-reboot (exit_code=97) rolling reboot on A:cp-upload_drmrs * 18:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2241: Security update * 17:56 topranks: start draining traffic on cr1-eqiad ahead of line card installation [[phab:T426343|T426343]] * 17:47 cdobbins@cumin2003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie * 17:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4051.ulsfo.wmnet * 17:40 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:39 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 17:34 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6007.drmrs.wmnet * 17:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6015.drmrs.wmnet * 17:32 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:31 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 17:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4043.ulsfo.wmnet * 17:27 lerickson@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:25 lerickson@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 17:22 lerickson@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-codfw * 17:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp2001.codfw.wmnet * 17:22 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 17:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp2001.codfw.wmnet * 17:22 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 17:19 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2241: Security update * 17:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2241.codfw.wmnet * 17:17 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2241.codfw.wmnet * 17:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp2001.codfw.wmnet * 17:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp2001.codfw.wmnet * 17:15 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2366-2374].codfw.wmnet * 17:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2366-2374].codfw.wmnet * 17:10 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply * 17:10 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply * 17:08 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2366-2374].codfw.wmnet * 17:06 sukhe: sre.dns.roll-reboot to resume later * 17:06 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-reboot (exit_code=97) rolling reboot on A:dnsbox and not (A:ulsfo or A:magru) and (A:dnsbox) * 17:06 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns5003.wikimedia.org * 17:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2241: Security update * 17:03 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2241: Security update * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply * 17:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2366-2374].codfw.wmnet * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply * 17:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2357-2365].codfw.wmnet * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 17:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2357-2365].codfw.wmnet * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply * 16:55 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2357-2365].codfw.wmnet * 16:55 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 16:53 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6006.drmrs.wmnet * 16:52 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6014.drmrs.wmnet * 16:52 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 16:51 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 16:50 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2357-2365].codfw.wmnet * 16:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4042.ulsfo.wmnet * 16:50 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2347-2356].codfw.wmnet * 16:50 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2347-2356].codfw.wmnet * 16:49 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns5003.wikimedia.org * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply * 16:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4050.ulsfo.wmnet * 16:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2347-2356].codfw.wmnet * 16:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2347-2356].codfw.wmnet * 16:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2337-2346].codfw.wmnet * 16:36 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2337-2346].codfw.wmnet * 16:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:dse-k8s-worker-codfw * 16:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2003.codfw.wmnet * 16:35 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2003.codfw.wmnet * 16:34 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns3004.wikimedia.org * 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply * 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply * 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply * 16:30 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply * 16:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2003.codfw.wmnet * 16:29 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2337-2346].codfw.wmnet * 16:24 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2003.codfw.wmnet * 16:24 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2002.codfw.wmnet * 16:24 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2002.codfw.wmnet * 16:23 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns3004.wikimedia.org * 16:23 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2337-2346].codfw.wmnet * 16:23 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2327-2336].codfw.wmnet * 16:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2327-2336].codfw.wmnet * 16:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2002.codfw.wmnet * 16:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2327-2336].codfw.wmnet * 16:12 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2002.codfw.wmnet * 16:12 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2001.codfw.wmnet * 16:12 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2001.codfw.wmnet * 16:12 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1065.eqiad.wmnet with OS trixie * 16:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6005.drmrs.wmnet * 16:11 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6013.drmrs.wmnet * 16:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4041.ulsfo.wmnet * 16:08 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns3003.wikimedia.org * 16:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2327-2336].codfw.wmnet * 16:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2317-2326].codfw.wmnet * 16:06 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2317-2326].codfw.wmnet * 16:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2001.codfw.wmnet * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply * 16:03 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4049.ulsfo.wmnet * 16:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2001.codfw.wmnet * 16:00 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test2001.codfw.wmnet * 16:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test2001.codfw.wmnet * 16:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2063.codfw.wmnet with OS trixie * 15:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2317-2326].codfw.wmnet * 15:57 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns3003.wikimedia.org * 15:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test2001.codfw.wmnet * 15:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test2001.codfw.wmnet * 15:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2004.codfw.wmnet * 15:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2004.codfw.wmnet * 15:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2317-2326].codfw.wmnet * 15:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2307-2316].codfw.wmnet * 15:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2307-2316].codfw.wmnet * 15:49 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2004.codfw.wmnet * 15:48 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2004.codfw.wmnet * 15:48 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2003.codfw.wmnet * 15:48 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2003.codfw.wmnet * 15:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 15:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2307-2316].codfw.wmnet * 15:42 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2003.codfw.wmnet * 15:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 15:42 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2003.codfw.wmnet * 15:42 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2002.codfw.wmnet * 15:42 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2002.codfw.wmnet * 15:42 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2006.wikimedia.org * 15:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2063.codfw.wmnet with reason: host reimage * 15:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2307-2316].codfw.wmnet * 15:37 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2297-2306].codfw.wmnet * 15:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2297-2306].codfw.wmnet * 15:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2002.codfw.wmnet * 15:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2002.codfw.wmnet * 15:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2001.codfw.wmnet * 15:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2001.codfw.wmnet * 15:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2063.codfw.wmnet with reason: host reimage * 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6004.drmrs.wmnet * 15:31 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2001.codfw.wmnet * 15:31 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2001.codfw.wmnet * 15:31 btullis@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:dse-k8s-worker-codfw * 15:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6012.drmrs.wmnet * 15:28 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2006.wikimedia.org * 15:27 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-analytics cluster: Roll restart of jvm daemons. * 15:27 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2297-2306].codfw.wmnet * 15:27 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4040.ulsfo.wmnet * 15:24 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1065.eqiad.wmnet with OS trixie * 15:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4048.ulsfo.wmnet * 15:21 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-analytics cluster: Roll restart of jvm daemons. * 15:21 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2297-2306].codfw.wmnet * 15:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2287-2296].codfw.wmnet * 15:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2287-2296].codfw.wmnet * 15:20 btullis@cumin1003: END (PASS) - Cookbook sre.druid.reboot-workers (exit_code=0) for Druid public cluster: Reboot Druid nodes * 15:18 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 15:17 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm * 15:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2063.codfw.wmnet with OS trixie * 15:13 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2005.wikimedia.org * 15:11 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2287-2296].codfw.wmnet * 15:11 btullis@cumin1003: END (PASS) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=0) rolling reboot on A:cephosd-eqiad * 15:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1064.eqiad.wmnet with OS trixie * 15:05 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2062.codfw.wmnet with OS trixie * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2287-2296].codfw.wmnet * 15:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2277-2286].codfw.wmnet * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2277-2286].codfw.wmnet * 14:59 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2005.wikimedia.org * 14:57 brouberol@cumin1003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-jumbo-eqiad * 14:52 btullis@cumin1003: END (PASS) - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas (exit_code=0) rolling reboot on A:schema-codfw * 14:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6003.drmrs.wmnet * 14:50 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:50 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host relforge1009.eqiad.wmnet * 14:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2277-2286].codfw.wmnet * 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6011.drmrs.wmnet * 14:47 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 14:46 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>ml-serve1001.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 14:46 klausman@cumin1003: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) pool for host ml-serve1001.eqiad.wmnet * 14:46 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 14:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1001.eqiad.wmnet * 14:45 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4039.ulsfo.wmnet * 14:44 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1009.eqiad.wmnet * 14:44 btullis@cumin1003: START - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas rolling reboot on A:schema-codfw * 14:44 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2004.wikimedia.org * 14:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2277-2286].codfw.wmnet * 14:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2267-2276].codfw.wmnet * 14:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2267-2276].codfw.wmnet * 14:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 14:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4047.ulsfo.wmnet * 14:40 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1001.eqiad.wmnet * 14:38 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 14:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 14:36 topranks: disconnect power on cr2-eqiad to shut down device for switch fabric replacement [[phab:T426343|T426343]] * 14:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2267-2276].codfw.wmnet * 14:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 14:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1001.eqiad.wmnet * 14:35 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>ml-serve1001.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 14:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 14:34 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 14:33 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:33 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:30 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2004.wikimedia.org * 14:29 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2267-2276].codfw.wmnet * 14:29 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2257-2266].codfw.wmnet * 14:29 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2257-2266].codfw.wmnet * 14:24 btullis@cumin1003: END (PASS) - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas (exit_code=0) rolling reboot on A:schema-eqiad * 14:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2257-2266].codfw.wmnet * 14:20 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:20 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:19 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:17 jforrester@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2257-2266].codfw.wmnet * 14:16 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:16 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2062.codfw.wmnet with OS trixie * 14:15 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1064.eqiad.wmnet with OS trixie * 14:15 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1006.wikimedia.org * 14:15 btullis@cumin1003: START - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas rolling reboot on A:schema-eqiad * 14:14 topranks: switch routing-engine on cr2-eqiad resetting all interfaces [[phab:T417873|T417873]] * 14:11 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:11 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:10 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm * 14:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6002.drmrs.wmnet * 14:09 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6010.drmrs.wmnet * 14:06 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1006.wikimedia.org * 14:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:05 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 14:05 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4038.ulsfo.wmnet * 14:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4046.ulsfo.wmnet * 14:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:00 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on cr1-eqiad with reason: switch upgrade and line card install * 14:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:59 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:57 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:57 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:55 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-eqiad * 13:55 btullis@cumin1003: START - Cookbook sre.druid.reboot-workers for Druid public cluster: Reboot Druid nodes * 13:53 topranks: switch routing-engine on cr2-eqiad resetting all interfaces [[phab:T417873|T417873]] * 13:51 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1005.wikimedia.org * 13:50 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:49 brouberol@cumin1003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-test-eqiad * 13:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:44 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2197-2206].codfw.wmnet * 13:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2197-2206].codfw.wmnet * 13:36 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1005.wikimedia.org * 13:35 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2197-2206].codfw.wmnet * 13:30 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2197-2206].codfw.wmnet * 13:28 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6001.drmrs.wmnet * 13:28 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6009.drmrs.wmnet * 13:28 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2187-2196].codfw.wmnet * 13:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2187-2196].codfw.wmnet * 13:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 13:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4037.ulsfo.wmnet * 13:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2001 * 13:22 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2001 * 13:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4045.ulsfo.wmnet * 13:21 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1004.wikimedia.org * 13:19 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on lvs[1018-1020].eqiad.wmnet with reason: switch upgrade and line card install * 13:18 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2009.codfw.wmnet * 13:18 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2001 * 13:18 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2001.codfw.wmnet 26.16.192.10.in-addr.arpa 6.2.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:17 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2001.codfw.wmnet 26.16.192.10.in-addr.arpa 6.2.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:17 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:17 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2001 - bking@cumin2003" * 13:17 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2001 - bking@cumin2003" * 13:17 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2009.codfw.wmnet * 13:17 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_drmrs * 13:17 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2187-2196].codfw.wmnet * 13:17 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_drmrs * 13:17 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on 15 hosts with reason: switch upgrade and line card install * 13:17 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:15 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 13:13 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:13 brouberol@cumin1003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-jumbo-eqiad * 13:13 brouberol@cumin1003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-test-eqiad * 13:13 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1004.wikimedia.org * 13:13 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and not (A:ulsfo or A:magru) and (A:dnsbox) * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:12 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_ulsfo * 13:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:12 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_ulsfo * 13:11 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2187-2196].codfw.wmnet * 13:11 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 13:11 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 13:06 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 13:05 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling source-only afterwards * 13:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2001 * 13:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2009.codfw.wmnet with OS trixie * 13:03 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:03 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling source-only afterwards * 13:01 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 15s) * 13:01 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 13:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 12:57 btullis@cumin1003: END (PASS) - Cookbook sre.druid.reboot-workers (exit_code=0) for Druid analytics cluster: Reboot Druid nodes * 12:54 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 12:54 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2163-2172].codfw.wmnet * 12:54 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2163-2172].codfw.wmnet * 12:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2163-2172].codfw.wmnet * 12:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2009.codfw.wmnet with reason: host reimage * 12:41 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2163-2172].codfw.wmnet * 12:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2153-2162].codfw.wmnet * 12:40 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2153-2162].codfw.wmnet * 12:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2009.codfw.wmnet with reason: host reimage * 12:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2153-2162].codfw.wmnet * 12:29 btullis@cumin1003: END (PASS) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=0) rolling reboot on A:cephosd-codfw * 12:25 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2153-2162].codfw.wmnet * 12:25 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2143-2152].codfw.wmnet * 12:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2143-2152].codfw.wmnet * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2009 * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2009 * 12:22 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2009 * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2009.codfw.wmnet 139.0.192.10.in-addr.arpa 9.3.1.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:22 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2009.codfw.wmnet 139.0.192.10.in-addr.arpa 9.3.1.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2009 - mvernon@cumin2003" * 12:22 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2009 - mvernon@cumin2003" * 12:16 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 12:15 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 12:15 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 12:15 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2009 * 12:15 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 12:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2009.codfw.wmnet with OS trixie * 12:15 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 12:14 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 12:14 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2143-2152].codfw.wmnet * 12:13 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 12:12 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2010.codfw.wmnet * 12:11 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2010.codfw.wmnet * 12:10 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 12:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2143-2152].codfw.wmnet * 12:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2133-2142].codfw.wmnet * 12:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2133-2142].codfw.wmnet * 12:02 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 11:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2133-2142].codfw.wmnet * 11:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2133-2142].codfw.wmnet * 11:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:49 mvolz@deploy2003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:49 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-codfw * 11:48 mvolz@deploy2003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:47 btullis@cumin1003: START - Cookbook sre.druid.reboot-workers for Druid analytics cluster: Reboot Druid nodes * 11:46 mvolz@deploy2003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:46 mvolz@deploy2003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:45 mvolz@deploy2003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:44 mvolz@deploy2003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:40 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] (duration: 11m 38s) * 11:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2010.codfw.wmnet with OS trixie * 11:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2105-2114].codfw.wmnet * 11:36 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2105-2114].codfw.wmnet * 11:36 krinkle@deploy2003: physikerwelt, krinkle: Continuing with deployment * 11:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1018: Security updates * 11:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:36 root@cumin1003: START - Cookbook sre.mysql.parsercache * 11:36 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1018: Security updates * 11:31 krinkle@deploy2003: physikerwelt, krinkle: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:29 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] * 11:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2105-2114].codfw.wmnet * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2105-2114].codfw.wmnet * 11:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2010.codfw.wmnet with reason: host reimage * 11:12 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2010.codfw.wmnet with reason: host reimage * 11:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1018: Security updates * 11:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:10 root@cumin1003: START - Cookbook sre.mysql.parsercache * 11:10 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1018: Security updates * 11:09 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1009.eqiad.wmnet with OS trixie * 11:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow7002.magru.wmnet * 11:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 11:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 11:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=tegola-vector-tiles,name=eqiad * 11:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=kartotherian,name=eqiad * 11:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow7002.magru.wmnet * 10:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2010 * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2010 * 10:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 10:54 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 10:54 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2010 * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2010.codfw.wmnet 76.16.192.10.in-addr.arpa 6.7.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:54 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2010.codfw.wmnet 76.16.192.10.in-addr.arpa 6.7.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2010 - mvernon@cumin2003" * 10:54 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2010 - mvernon@cumin2003" * 10:49 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 10:49 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2010 * 10:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1009.eqiad.wmnet with reason: host reimage * 10:49 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2010.codfw.wmnet with OS trixie * 10:46 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2011.codfw.wmnet * 10:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1011.eqiad.wmnet * 10:44 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2011.codfw.wmnet * 10:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 10:44 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 10:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1009.eqiad.wmnet with reason: host reimage * 10:44 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow6001.drmrs.wmnet * 10:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1017: Security updates * 10:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:39 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1017: Security updates * 10:39 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow6001.drmrs.wmnet * 10:38 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1011.eqiad.wmnet * 10:35 cgoubert@deploy2003: Finished deploy [restbase/deploy@06301bd]: Deploying {{Gerrit|1306088}} {{Gerrit|1308347}} - [[phab:T429944|T429944]] [[phab:T428279|T428279]] (duration: 28m 34s) * 10:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1012.eqiad.wmnet * 10:35 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 10:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow5003.eqsin.wmnet * 10:34 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2011.codfw.wmnet with OS trixie * 10:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1009.eqiad.wmnet with OS trixie * 10:28 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1012.eqiad.wmnet * 10:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1013.eqiad.wmnet * 10:27 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:27 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow5003.eqsin.wmnet * 10:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:26 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow4003.ulsfo.wmnet * 10:25 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 10:25 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 10:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow4003.ulsfo.wmnet * 10:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1013.eqiad.wmnet * 10:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1014.eqiad.wmnet * 10:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:15 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2011.codfw.wmnet with reason: host reimage * 10:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1017: Security updates * 10:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:14 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:14 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1017: Security updates * 10:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow3004.esams.wmnet * 10:11 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2011.codfw.wmnet with reason: host reimage * 10:10 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:10 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1014.eqiad.wmnet * 10:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki2003.codfw.wmnet * 10:09 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 10:09 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 10:09 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow3004.esams.wmnet * 10:08 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2004.codfw.wmnet * 10:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1010.eqiad.wmnet with OS trixie * 10:07 cgoubert@deploy2003: Started deploy [restbase/deploy@06301bd]: Deploying {{Gerrit|1306088}} {{Gerrit|1308347}} - [[phab:T429944|T429944]] [[phab:T428279|T428279]] * 10:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host rpki2003.codfw.wmnet * 10:04 topranks: push out config change to BGP_outfilter on core routers [[phab:T431849|T431849]] * 10:02 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow2004.codfw.wmnet * 09:59 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 09:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2003.codfw.wmnet * 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2011 * 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2011 * 09:53 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 09:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:52 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2011 * 09:52 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2011.codfw.wmnet 36.32.192.10.in-addr.arpa 6.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:52 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2011.codfw.wmnet 36.32.192.10.in-addr.arpa 6.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:51 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:51 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2011 - mvernon@cumin2003" * 09:51 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2011 - mvernon@cumin2003" * 09:51 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow2003.codfw.wmnet * 09:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1003.eqiad.wmnet * 09:49 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 09:49 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 09:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1010.eqiad.wmnet with reason: host reimage * 09:47 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 09:47 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 09:47 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 09:47 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2011 * 09:46 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2011.codfw.wmnet with OS trixie * 09:44 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow1003.eqiad.wmnet * 09:44 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2012.codfw.wmnet * 09:44 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1002.eqiad.wmnet * 09:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1010.eqiad.wmnet with reason: host reimage * 09:43 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2012.codfw.wmnet * 09:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Security updates * 09:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:43 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:43 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Security updates * 09:42 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:40 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow1002.eqiad.wmnet * 09:40 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 09:37 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki1001.eqiad.wmnet * 09:36 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 09:36 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 09:33 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host rpki1001.eqiad.wmnet * 09:32 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:32 cgoubert@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-codfw * 09:31 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=kartotherian,name=eqiad * 09:31 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola-vector-tiles,name=eqiad * 09:31 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 09:31 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2012.codfw.wmnet with OS trixie * 09:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1010.eqiad.wmnet with OS trixie * 09:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1011.eqiad.wmnet with OS trixie * 09:21 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Security updates * 09:21 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:21 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:21 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Security updates * 09:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2012.codfw.wmnet with reason: host reimage * 09:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1011.eqiad.wmnet with reason: host reimage * 09:08 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2012.codfw.wmnet with reason: host reimage * 09:05 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1011.eqiad.wmnet with reason: host reimage * 08:55 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:52 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1011.eqiad.wmnet with OS trixie * 08:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1022: Security updates * 08:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2012 * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2012 * 08:50 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1022: Security updates * 08:50 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2012 * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2012.codfw.wmnet 44.48.192.10.in-addr.arpa 4.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:50 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2012.codfw.wmnet 44.48.192.10.in-addr.arpa 4.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2012 - mvernon@cumin2003" * 08:50 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2012 - mvernon@cumin2003" * 08:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1012.eqiad.wmnet with OS trixie * 08:44 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 08:44 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2012 * 08:43 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2012.codfw.wmnet with OS trixie * 08:42 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2013.codfw.wmnet * 08:41 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2013.codfw.wmnet * 08:35 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 08:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host krb1002.eqiad.wmnet * 08:30 elukey@dns1004: END - running authdns-update * 08:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1012.eqiad.wmnet with reason: host reimage * 08:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Security updates * 08:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:28 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:28 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Security updates * 08:27 elukey@dns1004: START - running authdns-update * 08:26 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 08:26 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host krb1002.eqiad.wmnet * 08:22 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1012.eqiad.wmnet with reason: host reimage * 08:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host krb2002.codfw.wmnet * 08:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast6003.wikimedia.org * 08:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2013.codfw.wmnet with OS trixie * 08:13 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast6003.wikimedia.org * 08:12 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast3007.wikimedia.org * 08:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host krb2002.codfw.wmnet * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Security updates * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:09 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:09 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Security updates * 08:07 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1012.eqiad.wmnet with OS trixie * 08:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast3007.wikimedia.org * 08:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast5005.wikimedia.org * 07:58 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast5005.wikimedia.org * 07:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1013.eqiad.wmnet with OS trixie * 07:53 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2013.codfw.wmnet with reason: host reimage * 07:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1021: Security updates * 07:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:53 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:53 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1021: Security updates * 07:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast1004.wikimedia.org * 07:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2013.codfw.wmnet with reason: host reimage * 07:46 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast1004.wikimedia.org * 07:40 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1013.eqiad.wmnet with reason: host reimage * 07:36 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1013.eqiad.wmnet with reason: host reimage * 07:31 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2013 * 07:31 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2013 * 07:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1021: Security updates * 07:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:30 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:30 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1021: Security updates * 07:24 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2013 * 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2013.codfw.wmnet 87.0.192.10.in-addr.arpa 7.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:24 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2013.codfw.wmnet 87.0.192.10.in-addr.arpa 7.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2013 - mvernon@cumin2003" * 07:24 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2013 - mvernon@cumin2003" * 07:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1013.eqiad.wmnet with OS trixie * 07:19 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 07:19 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2013 * 07:19 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2013.codfw.wmnet with OS trixie * 07:13 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] (duration: 07m 48s) * 07:09 kharlan@deploy2003: kharlan: Continuing with deployment * 07:08 kharlan@deploy2003: kharlan: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:06 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 01:15 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 01:14 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply == 2026-07-14 == * 22:51 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_magru * 22:51 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7016.magru.wmnet * 22:46 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_magru * 22:46 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7008.magru.wmnet * 22:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7015.magru.wmnet * 22:04 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7007.magru.wmnet * 21:29 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7014.magru.wmnet * 21:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7006.magru.wmnet * 21:13 dzahn@cumin2002: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 0:15:00 on gerrit.wikimedia.org with reason: reboot * 21:11 mutante: gerrit2003 (gerrit.wikimedia.org) - reboot for maintenance * 21:11 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on gerrit2003.wikimedia.org with reason: reboot * 20:56 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:56 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:56 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:55 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 20:48 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7013.magru.wmnet * 20:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7005.magru.wmnet * 20:41 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host phab1005.eqiad.wmnet with OS trixie * 20:28 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] (duration: 06m 47s) * 20:24 sbassett@deploy2003: sbassett: Continuing with deployment * 20:23 sbassett@deploy2003: sbassett: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:23 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on phab1005.eqiad.wmnet with reason: host reimage * 20:21 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] * 20:20 aokoth@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on phab1005.eqiad.wmnet with reason: host reimage * 20:12 jhuneidi@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] (duration: 07m 42s) * 20:07 jhuneidi@deploy2003: jhuneidi, priyankar22: Continuing with deployment * 20:06 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7012.magru.wmnet * 20:06 jhuneidi@deploy2003: jhuneidi, priyankar22: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:04 jhuneidi@deploy2003: Started scap sync-world: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] * 20:02 aokoth@cumin1003: START - Cookbook sre.hosts.reimage for host phab1005.eqiad.wmnet with OS trixie * 20:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7004.magru.wmnet * 20:00 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet * 19:57 aokoth@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet * 19:24 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7011.magru.wmnet * 19:19 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7003.magru.wmnet * 19:11 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] (duration: 08m 33s) * 19:07 jforrester@deploy2003: jforrester: Continuing with deployment * 19:04 jforrester@deploy2003: jforrester: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:02 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] * 18:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7010.magru.wmnet * 18:38 mutante: rotating phabricator-gerrit bot token (its-phabricator) * 18:18 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 17:44 swfrench@deploy2003: Finished scap sync-world: Deployment to pick up new production image (duration: 31m 44s) * 17:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7002.magru.wmnet * 17:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7009.magru.wmnet * 17:32 swfrench@deploy2003: swfrench: Continuing with deployment * 17:29 swfrench@deploy2003: swfrench: Deployment to pick up new production image synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:17 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2035: repooling after rack b5 maintenance * 17:16 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool es2035: repooling after rack b5 maintenance * 17:16 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2188: repooling after rack b5 maintenance * 17:12 swfrench@deploy2003: Started scap sync-world: Deployment to pick up new production image * 17:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7001.magru.wmnet * 16:57 swfrench-wmf: reprepro include php8.3_8.3.32-1+wmf12u2 into component/php83 for bookworm-wikimedia * 16:50 sukhe: pool cp2046 * 16:47 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4039.ulsfo.wmnet * 16:44 sukhe: sudo cumin -b31 "A:cp" "run-puppet-agent" * 16:33 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on contint1003.wikimedia.org with reason: reboot * 16:32 mutante: contint1003 - main CI server - rebooting * 16:31 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2188: repooling after rack b5 maintenance * 16:31 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2178: repooling after rack b5 maintenance * 16:29 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 16:28 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 16:28 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 16:28 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 16:18 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2014.codfw.wmnet * 16:18 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 16:17 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2014.codfw.wmnet * 16:10 mvernon@cumin1003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-thanos-proxies (exit_code=0) rolling restart_daemons on A:thanos-fe * 16:09 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 16:07 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp4039.ulsfo.wmnet * 16:07 mvernon@cumin1003: START - Cookbook sre.swift.roll-restart-reboot-swift-thanos-proxies rolling restart_daemons on A:thanos-fe * 16:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2014.codfw.wmnet with OS trixie * 15:56 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1014.eqiad.wmnet with OS trixie * 15:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2014.codfw.wmnet with reason: host reimage * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2014 * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2014 * 15:28 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2014 * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2014.codfw.wmnet 194.16.192.10.in-addr.arpa 4.9.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:28 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2014.codfw.wmnet 194.16.192.10.in-addr.arpa 4.9.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2014 - mvernon@cumin2003" * 15:28 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2014 - mvernon@cumin2003" * 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Apply title-related policies when selecting the name of the entity - kamila@cumin1003" * 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Apply title-related policies when selecting the name of the entity - kamila@cumin1003 * 15:22 kamila@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Apply title-related policies when selecting the name of the entity - kamila@cumin1003 * 15:22 kamila@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Apply title-related policies when selecting the name of the entity - kamila@cumin1003" * 15:20 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 15:20 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2014 * 15:20 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2014.codfw.wmnet with OS trixie * 15:19 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1014.eqiad.wmnet with OS trixie * 15:01 dancy@deploy2003: Installation of scap version "4.274.1" completed for 3 hosts * 15:00 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2177: repooling after rack b5 maintenance * 15:00 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2159: repooling after rack b5 maintenance * 14:59 dancy@deploy2003: Installing scap version "4.274.1" for 3 host(s) * 14:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2015.codfw.wmnet with OS trixie * 14:54 seanleong-wmde: Finished populateSitesTable for isvwiki ([[phab:T429939|T429939]]) * 14:53 javiermonton@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] (duration: 07m 35s) * 14:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1015.eqiad.wmnet with OS trixie * 14:49 javiermonton@deploy2003: javiermonton: Continuing with deployment * 14:48 javiermonton@deploy2003: javiermonton: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:46 javiermonton@deploy2003: Started scap sync-world: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] * 14:42 otto@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 14:41 otto@deploy2003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 14:41 otto@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 14:40 otto@deploy2003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 14:40 otto@deploy2003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 14:39 otto@deploy2003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 14:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2015.codfw.wmnet with reason: host reimage * 14:34 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1015.eqiad.wmnet with reason: host reimage * 14:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2015.codfw.wmnet with reason: host reimage * 14:30 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1015.eqiad.wmnet with reason: host reimage * 14:30 seanleong-wmde@deploy2003: mwscript-k8s job started: foreachwikiindblist wikidataclient extensions/Wikibase/lib/maintenance/populateSitesTable.php --force-protocol https # [[phab:T429939|T429939]] * 14:24 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling reboot on A:durum and not (A:durum-eqiad or A:durum-codfw or A:durum-esams) and A:durum * 14:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2015.codfw.wmnet with OS trixie * 14:15 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2016.codfw.wmnet with OS trixie * 14:14 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2159: repooling after rack b5 maintenance * 14:14 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1015.eqiad.wmnet with OS trixie * 14:12 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1016.eqiad.wmnet with OS trixie * 14:12 cmooney@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=pki,name=codfw * 14:12 sbisson@deploy2003: helmfile [codfw] DONE helmfile.d/services/cxserver: sync * 14:11 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2002.codfw.wmnet * 14:11 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2002.codfw.wmnet * 14:11 sbisson@deploy2003: helmfile [codfw] START helmfile.d/services/cxserver: sync * 14:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1005.wikimedia.org * 14:07 sbisson@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cxserver: sync * 14:07 sbisson@deploy2003: helmfile [eqiad] START helmfile.d/services/cxserver: sync * 14:05 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1005.wikimedia.org * 14:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader2005.wikimedia.org * 14:02 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2003.codfw.wmnet * 14:02 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2003.codfw.wmnet * 14:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=tegola-vector-tiles,name=codfw * 14:00 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=kartotherian,name=codfw * 14:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader2005.wikimedia.org * 13:58 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2016.codfw.wmnet with reason: host reimage * 13:57 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-ncredir (exit_code=0) rolling reboot on A:ncredir and A:ncredir * 13:57 sbisson@deploy2003: helmfile [staging] DONE helmfile.d/services/cxserver: sync * 13:56 sbisson@deploy2003: helmfile [staging] START helmfile.d/services/cxserver: sync * 13:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1016.eqiad.wmnet with reason: host reimage * 13:52 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:52 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:51 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2016.codfw.wmnet with reason: host reimage * 13:50 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1016.eqiad.wmnet with reason: host reimage * 13:49 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy (exit_code=0) rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 13:49 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling reboot on A:wikidough * 13:46 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-tcp-proxy (exit_code=0) rolling reboot on A:tcpproxy and A:tcpproxy * 13:43 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and not (A:durum-eqiad or A:durum-codfw or A:durum-esams) and A:durum * 13:42 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=97) rolling reboot on A:durum and A:durum * 13:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2011.codfw.wmnet * 13:36 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-reboot (exit_code=0) rolling reboot on A:dnsbox and A:ulsfo and (A:dnsbox) * 13:36 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns4004.wikimedia.org * 13:34 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1016.eqiad.wmnet with OS trixie * 13:34 topranks: reboot lsw1-b5-codfw to upgrade JunOS [[phab:T430918|T430918]] * 13:34 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2016.codfw.wmnet with OS trixie * 13:32 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2002.codfw.wmnet * 13:31 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2011.codfw.wmnet * 13:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2012.codfw.wmnet * 13:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2017.codfw.wmnet with OS trixie * 13:25 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1017.eqiad.wmnet with OS trixie * 13:24 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2012.codfw.wmnet * 13:22 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2002.codfw.wmnet * 13:22 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:22 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:22 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns4004.wikimedia.org * 13:21 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2005.codfw.wmnet * 13:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2013.codfw.wmnet * 13:18 elukey@dns1004: END - running authdns-update * 13:17 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2005.codfw.wmnet * 13:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2004.codfw.wmnet * 13:16 elukey@dns1004: START - running authdns-update * 13:16 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1046: es1046 after reimage * 13:14 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1029.eqiad.wmnet,service=s8 * 13:14 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1029.eqiad.wmnet,service=s5 * 13:13 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1029.eqiad.wmnet,service=s5 * 13:13 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1029.eqiad.wmnet,service=s8 * 13:13 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2004.codfw.wmnet * 13:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2013.codfw.wmnet * 13:11 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2014.codfw.wmnet * 13:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2003.codfw.wmnet * 13:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2017.codfw.wmnet with reason: host reimage * 13:07 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:07 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns4003.wikimedia.org * 13:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2003.codfw.wmnet * 13:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1067.eqiad.wmnet * 13:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1067.eqiad.wmnet * 13:06 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1067.eqiad.wmnet * 13:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm-test1001.wikimedia.org * 13:05 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2017.codfw.wmnet with reason: host reimage * 13:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1017.eqiad.wmnet with reason: host reimage * 13:04 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2014.codfw.wmnet * 13:03 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:02 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2188: codfw rack B5 depool for maintenance * 13:02 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_magru * 13:01 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2188: codfw rack B5 depool for maintenance * 13:01 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_magru * 13:01 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2178: codfw rack B5 depool for maintenance * 13:01 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm-test1001.wikimedia.org * 13:01 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2178: codfw rack B5 depool for maintenance * 13:01 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2177: codfw rack B5 depool for maintenance * 13:00 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2177: codfw rack B5 depool for maintenance * 12:59 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1068.eqiad.wmnet * 12:59 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1068.eqiad.wmnet * 12:58 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola-vector-tiles,name=codfw * 12:58 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2159: codfw rack B5 depool for maintenance * 12:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1017.eqiad.wmnet with reason: host reimage * 12:58 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola,name=codfw * 12:57 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=kartotherian,name=codfw * 12:57 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2159: codfw rack B5 depool for maintenance * 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 30 hosts with reason: lsw1-b5-codfw JunOS upgrade * 12:55 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lsw1-b5-codfw,lsw1-b5-codfw IPv6,lsw1-b5-codfw.mgmt,ssw1-a[1,8]-codfw.mgmt with reason: switch upgade lsw1-b5-codfw * 12:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet * 12:49 cmooney@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=pki,name=codfw * 12:49 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1067.eqiad.wmnet with OS trixie * 12:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet * 12:48 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2017.codfw.wmnet with OS trixie * 12:47 topranks: depool codfw pki in dns discovery ahead of lsw1-b5-codfw maintenance [[phab:T430918|T430918]] * 12:47 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns4003.wikimedia.org * 12:47 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and A:ulsfo and (A:dnsbox) * 12:47 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2018.codfw.wmnet * 12:45 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2018.codfw.wmnet * 12:45 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and A:durum * 12:45 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-tcp-proxy rolling reboot on A:tcpproxy and A:tcpproxy * 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host pki-root1002.eqiad.wmnet * 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1009.eqiad.wmnet * 12:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1009.eqiad.wmnet * 12:44 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 12:43 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-ncredir rolling reboot on A:ncredir and A:ncredir * 12:43 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling reboot on A:wikidough * 12:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2018.codfw.wmnet with OS trixie * 12:42 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1017.eqiad.wmnet with OS trixie * 12:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader2006.wikimedia.org * 12:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1018.eqiad.wmnet with OS trixie * 12:39 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1009.eqiad.wmnet * 12:38 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host pki-root1002.eqiad.wmnet * 12:38 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1009.eqiad.wmnet * 12:38 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1008.eqiad.wmnet * 12:38 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1008.eqiad.wmnet * 12:35 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader2006.wikimedia.org * 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1006.wikimedia.org * 12:33 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1008.eqiad.wmnet * 12:30 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1046: es1046 after reimage * 12:29 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host es1046.eqiad.wmnet with OS trixie * 12:29 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1006.wikimedia.org * 12:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test2005.wikimedia.org * 12:28 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1008.eqiad.wmnet * 12:27 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1007.eqiad.wmnet * 12:27 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1007.eqiad.wmnet * 12:27 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1067.eqiad.wmnet with reason: host reimage * 12:25 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1068.eqiad.wmnet with reason: vacuum overlarge container dbs * 12:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2018.codfw.wmnet with reason: host reimage * 12:24 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test2005.wikimedia.org * 12:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test1005.wikimedia.org * 12:23 Amir1: mwscript-k8s --follow --dblist=ores -- extensions/ORES/maintenance/PurgeScoreCache.php --model goodfaith --old ([[phab:T431159|T431159]]) * 12:22 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1007.eqiad.wmnet * 12:22 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test1005.wikimedia.org * 12:22 atsukoito: restarting pybal on lvs2013 `low-traffic` for https://gerrit.wikimedia.org/r/1310535 * 12:22 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1007.eqiad.wmnet * 12:21 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1006.eqiad.wmnet * 12:21 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1006.eqiad.wmnet * 12:20 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1018.eqiad.wmnet with reason: host reimage * 12:19 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2018.codfw.wmnet with reason: host reimage * 12:18 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1067.eqiad.wmnet with reason: host reimage * 12:16 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1006.eqiad.wmnet * 12:15 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1006.eqiad.wmnet * 12:15 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1005.eqiad.wmnet * 12:15 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1005.eqiad.wmnet * 12:15 atsukoito: restarting pybal on lvs2014 for https://gerrit.wikimedia.org/r/1310535 * 12:12 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1018.eqiad.wmnet with reason: host reimage * 12:11 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1005.eqiad.wmnet * 12:11 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1005.eqiad.wmnet * 12:10 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1004.eqiad.wmnet * 12:10 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1004.eqiad.wmnet * 12:09 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on es1046.eqiad.wmnet with reason: host reimage * 12:08 atsukoito: restarting pybal on lvs1019 `low-traffic` for https://gerrit.wikimedia.org/r/1310535 * 12:06 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1004.eqiad.wmnet * 12:06 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1004.eqiad.wmnet * 12:06 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1003.eqiad.wmnet * 12:06 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1003.eqiad.wmnet * 12:05 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on es1046.eqiad.wmnet with reason: host reimage * 12:05 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "set ml-serve1001 back to active state - cmooney@cumin1003" * 12:04 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "set ml-serve1001 back to active state - cmooney@cumin1003" * 12:04 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:02 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1003.eqiad.wmnet * 12:01 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1003.eqiad.wmnet * 12:01 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1002.eqiad.wmnet * 12:01 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1002.eqiad.wmnet * 12:01 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:59 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2018.codfw.wmnet with OS trixie * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1067 * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1067 * 11:59 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1067 * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1067.eqiad.wmnet 17.48.64.10.in-addr.arpa 7.1.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:59 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1067.eqiad.wmnet 17.48.64.10.in-addr.arpa 7.1.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1067 - blake@cumin1003" * 11:58 atsukoito: restarting pybal on lvs1018 `high-traffic2` for https://gerrit.wikimedia.org/r/1310535 * 11:57 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1002.eqiad.wmnet * 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2019.codfw.wmnet with OS trixie * 11:56 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1002.eqiad.wmnet * 11:56 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 11:56 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1018.eqiad.wmnet with OS trixie * 11:54 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 11:54 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:54 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1019.eqiad.wmnet with OS trixie * 11:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:49 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:49 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:49 aikochou@deploy2003: helmfile [codfw] DONE helmfile.d/services/changeprop: sync * 11:48 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host es1046.eqiad.wmnet with OS trixie * 11:48 aikochou@deploy2003: helmfile [codfw] START helmfile.d/services/changeprop: sync * 11:48 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310535 * 11:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1046: Reimage to Trixie * 11:44 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1046: Reimage to Trixie * 11:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5:00:00 on es1046.eqiad.wmnet with reason: Reimage to Trixie * 11:42 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:42 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:42 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 11:42 aikochou@deploy2003: helmfile [eqiad] DONE helmfile.d/services/changeprop: sync * 11:41 aikochou@deploy2003: helmfile [eqiad] START helmfile.d/services/changeprop: sync * 11:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2019.codfw.wmnet with reason: host reimage * 11:36 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] (duration: 09m 41s) * 11:36 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:36 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:35 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:35 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:32 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1019.eqiad.wmnet with reason: host reimage * 11:32 jforrester@deploy2003: jforrester, gengh: Continuing with deployment * 11:29 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2019.codfw.wmnet with reason: host reimage * 11:28 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1019.eqiad.wmnet with reason: host reimage * 11:28 jforrester@deploy2003: jforrester, gengh: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:26 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] * 11:20 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2003.codfw.wmnet * 11:20 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:19 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2003.codfw.wmnet * 11:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:12 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1019.eqiad.wmnet with OS trixie * 11:10 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2019.codfw.wmnet with OS trixie * 11:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1020.eqiad.wmnet with OS trixie * 11:10 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1067 - blake@cumin1003" * 11:09 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] (duration: 12m 12s) * 11:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2020.codfw.wmnet with OS trixie * 11:03 kharlan@deploy2003: kharlan: Continuing with deployment * 11:01 blake@cumin1003: START - Cookbook sre.dns.netbox * 11:01 kharlan@deploy2003: kharlan: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:57 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] * 10:55 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] (duration: 31m 40s) * 10:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1020.eqiad.wmnet with reason: host reimage * 10:52 marostegui@dns1004: START - running authdns-update * 10:49 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2020.codfw.wmnet with reason: host reimage * 10:49 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:48 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1020.eqiad.wmnet with reason: host reimage * 10:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2020.codfw.wmnet with reason: host reimage * 10:43 kharlan@deploy2003: kharlan: Continuing with deployment * 10:42 kharlan@deploy2003: kharlan: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:32 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1020.eqiad.wmnet with OS trixie * 10:29 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2159: Repooling after switchover * 10:29 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1067 * 10:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1021.eqiad.wmnet with OS trixie * 10:27 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1067.eqiad.wmnet with OS trixie * 10:27 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:27 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1067.eqiad.wmnet * 10:27 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:26 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1067.eqiad.wmnet * 10:26 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1067.eqiad.wmnet * 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2020.codfw.wmnet with OS trixie * 10:26 blake@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1055.eqiad.wmnet * 10:26 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1055.eqiad.wmnet * 10:26 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1055.eqiad.wmnet * 10:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2021.codfw.wmnet with OS trixie * 10:24 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] * 10:11 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1055.eqiad.wmnet with OS trixie * 10:09 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1021.eqiad.wmnet with reason: host reimage * 10:05 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2021.codfw.wmnet with reason: host reimage * 10:03 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310129 revert * 10:02 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1021.eqiad.wmnet with reason: host reimage * 10:01 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2021.codfw.wmnet with reason: host reimage * 09:58 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310129 * 09:50 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1055.eqiad.wmnet with reason: host reimage * 09:45 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1055.eqiad.wmnet with reason: host reimage * 09:45 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1021.eqiad.wmnet with OS trixie * 09:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1022.eqiad.wmnet with OS trixie * 09:44 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2159: Repooling after switchover * 09:44 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2021.codfw.wmnet with OS trixie * 09:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2022.codfw.wmnet with OS trixie * 09:31 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2159.codfw.wmnet * 09:28 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1055 * 09:28 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1055 * 09:27 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on ms-fe1022.eqiad.wmnet with reason: host reimage * 09:27 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1022.eqiad.wmnet with reason: host reimage * 09:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2022.codfw.wmnet with reason: host reimage * 09:21 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2022.codfw.wmnet with reason: host reimage * 09:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2159: Rebooting db2159.codfw.wmnet * 09:20 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2159: Rebooting db2159.codfw.wmnet * 09:18 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 09:18 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 09:18 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 09:17 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 09:16 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2159.codfw.wmnet * 09:13 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b] (thin): Regular analytics weekly train THIN [analytics/refinery@ad6e05b8] (duration: 02m 07s) * 09:11 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b] (thin): Regular analytics weekly train THIN [analytics/refinery@ad6e05b8] * 09:10 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1022.eqiad.wmnet with OS trixie * 09:07 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1023.eqiad.wmnet with OS trixie * 09:06 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b]: Regular analytics weekly train [analytics/refinery@ad6e05b8] (duration: 04m 49s) * 09:04 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2022.codfw.wmnet with OS trixie * 09:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2023.codfw.wmnet with OS trixie * 09:01 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b]: Regular analytics weekly train [analytics/refinery@ad6e05b8] * 09:01 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@ad6e05b8] (duration: 02m 01s) * 09:00 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1055 * 09:00 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1055.eqiad.wmnet 50.32.64.10.in-addr.arpa 0.5.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:00 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1055.eqiad.wmnet 50.32.64.10.in-addr.arpa 0.5.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:00 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:00 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1055 - blake@cumin1003" * 09:00 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1055 - blake@cumin1003" * 08:59 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@ad6e05b8] * 08:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2159 [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94811 and previous config saved to /var/cache/conftool/dbconfig/20260714-085624-cwilliams.json * 08:55 blake@cumin1003: START - Cookbook sre.dns.netbox * 08:55 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1055 * 08:54 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1055.eqiad.wmnet with OS trixie * 08:54 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1055.eqiad.wmnet * 08:53 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1055.eqiad.wmnet * 08:53 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1055.eqiad.wmnet * 08:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2220 to s7 primary [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94810 and previous config saved to /var/cache/conftool/dbconfig/20260714-085239-cwilliams.json * 08:51 cezmunsta: Starting s7 codfw failover from db2159 to db2220 - [[phab:T430920|T430920]] * 08:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1023.eqiad.wmnet with reason: host reimage * 08:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2220 with weight 0 [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94809 and previous config saved to /var/cache/conftool/dbconfig/20260714-084553-cwilliams.json * 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s7 [[phab:T430920|T430920]] * 08:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2023.codfw.wmnet with reason: host reimage * 08:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1023.eqiad.wmnet with reason: host reimage * 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2023.codfw.wmnet with reason: host reimage * 08:34 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:34 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:29 marostegui@dns1004: END - running authdns-update * 08:29 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox-dev2003.codfw.wmnet * 08:27 marostegui@dns1004: START - running authdns-update * 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker2*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2009.codfw.wmnet * 08:26 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2009.codfw.wmnet * 08:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1023.eqiad.wmnet with OS trixie * 08:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox-dev2003.codfw.wmnet * 08:24 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:24 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:24 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2023.codfw.wmnet with OS trixie * 08:24 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1029.eqiad.wmnet with reason: reboot * 08:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1027.eqiad.wmnet with reason: reboot * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:21 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2009.codfw.wmnet * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:20 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2009.codfw.wmnet * 08:20 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2008.codfw.wmnet * 08:20 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2008.codfw.wmnet * 08:15 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2008.codfw.wmnet * 08:14 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2008.codfw.wmnet * 08:14 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2007.codfw.wmnet * 08:14 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2007.codfw.wmnet * 08:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1024.eqiad.wmnet with OS trixie * 08:12 elukey@cumin1003: END (PASS) - Cookbook sre.pki.restart-reboot (exit_code=0) rolling reboot on P<nowiki>{</nowiki>pki*<nowiki>}</nowiki> and (A:pki) * 08:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 08:10 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 08:09 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2007.codfw.wmnet * 08:08 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2007.codfw.wmnet * 08:08 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2006.codfw.wmnet * 08:08 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2006.codfw.wmnet * 08:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2024.codfw.wmnet with OS trixie * 08:03 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2006.codfw.wmnet * 08:02 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2006.codfw.wmnet * 08:02 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2005.codfw.wmnet * 08:02 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2005.codfw.wmnet * 07:58 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2005.codfw.wmnet * 07:58 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2005.codfw.wmnet * 07:57 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2004.codfw.wmnet * 07:57 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2004.codfw.wmnet * 07:54 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki.discovery.wmnet. on all recursors * 07:54 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache pki.discovery.wmnet. on all recursors * 07:53 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2004.codfw.wmnet * 07:53 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2004.codfw.wmnet * 07:53 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2003.codfw.wmnet * 07:53 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2003.codfw.wmnet * 07:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1024.eqiad.wmnet with reason: host reimage * 07:49 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki.discovery.wmnet. on all recursors * 07:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2003.codfw.wmnet * 07:49 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache pki.discovery.wmnet. on all recursors * 07:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2024.codfw.wmnet with reason: host reimage * 07:48 elukey@cumin1003: START - Cookbook sre.pki.restart-reboot rolling reboot on P<nowiki>{</nowiki>pki*<nowiki>}</nowiki> and (A:pki) * 07:46 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1024.eqiad.wmnet with reason: host reimage * 07:45 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2024.codfw.wmnet with reason: host reimage * 07:45 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2003.codfw.wmnet * 07:44 elukey@cumin1003: END (PASS) - Cookbook sre.misc-clusters.restart-reboot-config-master (exit_code=0) rolling reboot on P<nowiki>{</nowiki>config-master*<nowiki>}</nowiki> and (A:config-master or A:config-master-eqiad or A:config-master-codfw) * 07:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2002.codfw.wmnet * 07:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2002.codfw.wmnet * 07:39 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2002.codfw.wmnet * 07:39 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) config-master.discovery.wmnet. on all recursors * 07:39 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache config-master.discovery.wmnet. on all recursors * 07:39 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2002.codfw.wmnet * 07:39 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker2*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl200*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl2003.codfw.wmnet * 07:36 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl2003.codfw.wmnet * 07:35 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) config-master.discovery.wmnet. on all recursors * 07:35 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache config-master.discovery.wmnet. on all recursors * 07:34 elukey@cumin1003: START - Cookbook sre.misc-clusters.restart-reboot-config-master rolling reboot on P<nowiki>{</nowiki>config-master*<nowiki>}</nowiki> and (A:config-master or A:config-master-eqiad or A:config-master-codfw) * 07:31 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl2003.codfw.wmnet * 07:31 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl2003.codfw.wmnet * 07:31 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl2002.codfw.wmnet * 07:31 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl2002.codfw.wmnet * 07:29 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1024.eqiad.wmnet with OS trixie * 07:28 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2024.codfw.wmnet with OS trixie * 07:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl2002.codfw.wmnet * 07:26 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl2002.codfw.wmnet * 07:26 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl200*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 07:26 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 07:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 06:50 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lists1004.wikimedia.org * 06:44 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host lists1004.wikimedia.org * 06:25 marostegui@dns1004: END - running authdns-update * 06:23 marostegui@dns1004: START - running authdns-update * 06:22 marostegui@dns1004: END - running authdns-update * 06:20 marostegui@dns1004: START - running authdns-update * 06:17 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1026.eqiad.wmnet with reason: reboot * 06:04 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: sync * 06:04 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: sync * 06:03 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync * 06:03 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync * 06:02 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync * 06:01 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync * 06:01 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync * 06:00 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync * 05:59 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:59 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:40 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:39 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:26 marostegui@dns1004: END - running authdns-update * 05:24 marostegui@dns1004: START - running authdns-update * 05:24 marostegui@dns1004: START - running authdns-update * 05:13 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1004.wikimedia.org * 05:07 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1004.wikimedia.org * 04:01 mwpresync@deploy2003: Pruned MediaWiki: 1.47.0-wmf.8 (duration: 01m 07s) * 03:39 mwpresync@deploy2003: Finished scap sync-world: testwikis to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] (duration: 36m 01s) * 03:03 mwpresync@deploy2003: Started scap sync-world: testwikis to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 29s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-13 == * 23:33 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1064.eqiad.wmnet * 23:33 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1064.eqiad.wmnet * 23:08 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1064.eqiad.wmnet with reason: vacuum overlarge container dbs * 23:06 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1069.eqiad.wmnet * 23:06 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1069.eqiad.wmnet * 22:34 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1069.eqiad.wmnet with reason: vacuum overlarge container dbs * 21:18 maryum: Deployed security fix for [[phab:T321092|T321092]] * 20:28 swfrench-wmf: reprepro include etcd-mirror_0.0.12-1+deb13u1 into main for trixie-wikimedia - [[phab:T424266|T424266]] * 20:26 swfrench-wmf: reprepro include etcd-mirror_0.0.12-1+deb12u1 into main for bookworm-wikimedia - [[phab:T428495|T428495]] * 20:23 dancy@deploy2003: Finished scap sync-world: Testing [[phab:T431635|T431635]] (duration: 03m 36s) * 20:19 dancy@deploy2003: Started scap sync-world: Testing [[phab:T431635|T431635]] * 20:18 dancy@deploy2003: Installation of scap version "4.274.0" completed for 3 hosts * 20:16 dancy@deploy2003: Installing scap version "4.274.0" for 3 host(s) * 20:12 kemayo@deploy2003: Finished scap sync-world: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] (duration: 08m 25s) * 20:07 kemayo@deploy2003: soda, esanders, kemayo: Continuing with deployment * 20:05 kemayo@deploy2003: soda, esanders, kemayo: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there * 20:04 kemayo@deploy2003: Started scap sync-world: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] * 18:22 cwhite: lvextend vg0/srv +500g on centrallog hosts * 18:19 cdobbins@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS trixie * 17:46 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1071.eqiad.wmnet * 17:46 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1071.eqiad.wmnet * 17:13 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1071.eqiad.wmnet with reason: vacuum overlarge container dbs * 17:07 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1065.eqiad.wmnet * 17:07 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1065.eqiad.wmnet * 17:06 dzahn@dns1006: END - running authdns-update * 17:04 dzahn@dns1006: START - running authdns-update * 17:01 dzahn@dns1006: END - running authdns-update * 16:59 dzahn@dns1006: START - running authdns-update * 16:51 dancy@deploy2003: Finished scap sync-world: testing [[phab:T428971|T428971]] (duration: 03m 37s) * 16:47 dancy@deploy2003: Started scap sync-world: testing [[phab:T428971|T428971]] * 16:45 atsukoito: restarting pybal on lvs1019 to flush IP address for `cirrussearch1122.eqiad.wmnet` after moving the vlan [[phab:T431311|T431311]] * 16:42 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:42 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:42 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:42 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:42 Amir1: mwscript-k8s --follow --dblist=ores -- extensions/ORES/maintenance/PurgeScoreCache.php --model damaging --old ([[phab:T431159|T431159]]) * 16:34 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Pool test * 16:34 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 16:34 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 16:34 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Pool test * 16:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Depool test * 16:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 16:33 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 16:33 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Depool test * 16:31 dancy@deploy2003: Installation of scap version "4.273.0" completed for 159 hosts * 16:29 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1065.eqiad.wmnet with reason: vacuum overlarge container dbs * 16:27 dancy@deploy2003: Installing scap version "4.273.0" for 159 host(s) * 16:27 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics-external: sync * 16:27 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics-external: sync * 16:26 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics-external: sync * 16:26 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics-external: sync * 16:22 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync * 16:21 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync * 16:21 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: sync * 16:21 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: sync * 16:19 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync * 16:19 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync * 16:18 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync * 16:17 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync * 15:59 atsukoito: restarting pybal on lvs1018 for https://gerrit.wikimedia.org/r/1310117 * 15:55 aikochou@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 15:50 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310117 * 15:46 aikochou@deploy2003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 15:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host kafka-logging1006.eqiad.wmnet * 15:43 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host kafka-logging1006.eqiad.wmnet * 15:41 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host ganeti-test[2001-2003].codfw.wmnet * 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host ganeti-test[2001-2003].codfw.wmnet * 15:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host netbox1003.eqiad.wmnet * 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host netbox1003.eqiad.wmnet * 15:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host netbox2003.codfw.wmnet * 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host netbox2003.codfw.wmnet * 15:36 sukhe: restart pybal on lvs1020 * 15:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet * 15:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet * 15:08 btullis@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:06 btullis@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 15:01 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:01 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:35 cdobbins@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 14:34 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:33 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:33 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:32 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:29 cdobbins@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 14:28 marostegui@dns1004: END - running authdns-update * 14:28 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:27 marostegui@dns1004: START - running authdns-update * 14:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1023.eqiad.wmnet with reason: reboot * 14:18 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2009.codfw.wmnet with OS trixie * 14:14 swfrench-wmf: start rolling run-puppet-agent on A:cp for ATS config change - [[phab:T428909|T428909]] [[phab:T431838|T431838]] * 14:05 swfrench-wmf: disable-puppet on A:cp for ATS config change - [[phab:T428909|T428909]] [[phab:T431838|T431838]] * 14:05 cdobbins@cumin2002: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie * 14:02 marostegui@dns1004: END - running authdns-update * 14:00 marostegui@dns1004: START - running authdns-update * 14:00 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1070.eqiad.wmnet * 14:00 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1070.eqiad.wmnet * 13:58 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2009.codfw.wmnet with reason: host reimage * 13:57 cdobbins@cumin2002: conftool action : set/pooled=no; selector: name=dns7002.* * 13:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2009.codfw.wmnet with reason: host reimage * 13:48 rscout@deploy2003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply * 13:48 rscout@deploy2003: helmfile [eqiad] START helmfile.d/services/miscweb: apply * 13:48 rscout@deploy2003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply * 13:47 rscout@deploy2003: helmfile [codfw] START helmfile.d/services/miscweb: apply * 13:40 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:33 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2009.codfw.wmnet with OS trixie * 13:30 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1070.eqiad.wmnet with reason: vacuum overlarge container dbs * 13:28 aude@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] (duration: 11m 12s) * 13:23 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:22 aude@deploy2003: aikochou, javiermonton, aude, gkm563: Continuing with deployment * 13:22 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:19 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:19 aude@deploy2003: aikochou, javiermonton, aude, gkm563: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] synced to the testservers * 13:17 aude@deploy2003: Started scap sync-world: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] * 13:01 ladsgroup@deploy2003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 13:01 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:00 ladsgroup@deploy2003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 12:59 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:52 ladsgroup@deploy2003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 12:51 ladsgroup@deploy2003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 12:48 atsuko@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 12:48 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 12:47 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2008.codfw.wmnet with OS trixie * 12:47 atsuko@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 12:47 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply * 12:47 atsuko@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:46 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 12:45 atsuko@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:45 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply * 12:45 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:44 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:43 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] (duration: 07m 02s) * 12:38 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 12:37 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:36 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] * 12:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2008.codfw.wmnet with reason: host reimage * 12:23 Msz2001: Deployed changes to private code for Suggested Investigations * 12:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2008.codfw.wmnet with reason: host reimage * 12:20 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:19 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:17 atsuko@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 12:17 atsuko@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 12:16 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:15 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] (duration: 07m 14s) * 12:10 mszwarc@deploy2003: mszwarc: Continuing with deployment * 12:09 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:07 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] * 12:04 mszwarc@deploy2003: sync-world aborted: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] (duration: 00m 29s) * 12:03 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] * 12:01 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2008.codfw.wmnet with OS trixie * 12:00 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] (duration: 07m 37s) * 11:55 zabe@deploy2003: zabe: Continuing with deployment * 11:54 zabe@deploy2003: zabe: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:52 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] * 11:51 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:43 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:35 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:34 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:33 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:30 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:28 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:27 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:17 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2007.codfw.wmnet with OS trixie * 11:09 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] * 11:06 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=s8 * 11:00 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=x3 * 11:00 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=s5 * 10:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2007.codfw.wmnet with reason: host reimage * 10:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2007.codfw.wmnet with reason: host reimage * 10:51 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host dse-k8s-worker1023 * 10:50 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host dse-k8s-worker1023 * 10:44 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dse-k8s-worker1023 * 10:43 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host dse-k8s-worker1023 * 10:42 marostegui@cumin1003: dbctl commit (dc=all): 'Change x4 masters [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P94804 and previous config saved to /var/cache/conftool/dbconfig/20260713-104248-marostegui.json * 10:37 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:37 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:35 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:35 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:34 atsuko@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 10:34 atsuko@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 10:33 marostegui@cumin1003: dbctl commit (dc=all): 'Push x4 initial dbctl config [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P94803 and previous config saved to /var/cache/conftool/dbconfig/20260713-103259-marostegui.json * 10:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2007.codfw.wmnet with OS trixie * 09:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2006.codfw.wmnet with OS trixie * 09:42 marostegui@dns1004: END - running authdns-update * 09:40 marostegui@dns1004: START - running authdns-update * 09:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2006.codfw.wmnet with reason: host reimage * 09:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2006.codfw.wmnet with reason: host reimage * 09:06 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1024.eqiad.wmnet with reason: reboot * 09:06 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:01 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2006.codfw.wmnet with OS trixie * 08:44 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 08:43 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 08:43 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:42 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:42 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:42 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:41 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 08:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup2004.codfw.wmnet * 08:38 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:33 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host db1208.eqiad.wmnet * 08:30 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=x3 * 08:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2005.codfw.wmnet with OS trixie * 08:28 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup2004.codfw.wmnet * 08:28 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup2003.codfw.wmnet * 08:24 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1039: Repooling after testing * 08:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on clouddb1016.eqiad.wmnet with reason: cloning * 08:23 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s5 * 08:23 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s8 * 08:21 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 08:21 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 08:17 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup2003.codfw.wmnet * 08:17 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1004.eqiad.wmnet * 08:14 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1208.eqiad.wmnet * 08:11 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host phab1005.eqiad.wmnet * 08:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2005.codfw.wmnet with reason: host reimage * 08:07 marostegui@dns1004: END - running authdns-update * 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1004.eqiad.wmnet * 08:07 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1003.eqiad.wmnet * 08:07 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:05 marostegui@dns1004: START - running authdns-update * 08:05 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2005.codfw.wmnet with reason: host reimage * 08:05 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 08:05 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host phab1005.eqiad.wmnet * 08:05 marostegui@dns1004: START - running authdns-update * 08:05 marostegui@dns1004: START - running authdns-update * 08:05 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 08:04 marostegui@dns1004: START - running authdns-update * 08:00 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit1003.wikimedia.org * 07:58 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1003.eqiad.wmnet * 07:58 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1002-dev.eqiad.wmnet * 07:58 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:58 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:54 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1002-dev.eqiad.wmnet * 07:54 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1001-dev.eqiad.wmnet * 07:54 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit1003.wikimedia.org * 07:53 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 07:53 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:52 Msz2001: UTC morning backport+config window done * 07:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2005.codfw.wmnet with OS trixie * {{safesubst:SAL entry|1=07:50 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark (T429943}} * 07:49 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1001-dev.eqiad.wmnet * 07:46 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:46 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:45 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 07:45 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:44 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 07:44 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:43 mszwarc@deploy2003: mszwarc, danielyepezgarces, anzx: Continuing with deployment * {{safesubst:SAL entry|1=07:39 mszwarc@deploy2003: mszwarc, danielyepezgarces, anzx: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark}} * 07:39 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1039: Repooling after testing * {{safesubst:SAL entry|1=07:36 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark (T429943)}} * 07:35 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] (duration: 30m 03s) * 07:25 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit2002.wikimedia.org * 07:22 mszwarc@deploy2003: mszwarc: Continuing with deployment * 07:21 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:19 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit2002.wikimedia.org * 07:15 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aphlict1002.eqiad.wmnet * 07:11 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host aphlict1002.eqiad.wmnet * 07:08 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2003.wikimedia.org * 07:05 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] * 07:02 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2003.wikimedia.org * 07:02 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2002.wikimedia.org * 06:55 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2002.wikimedia.org * 06:55 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1003.wikimedia.org * 06:49 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1003.wikimedia.org * 06:34 marostegui: Drop m5 ipoid database [[phab:T431007|T431007]] * 06:29 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1027.eqiad.wmnet with reason: reboot * 06:24 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1028.eqiad.wmnet with reason: reboot * 06:21 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1025.eqiad.wmnet with reason: reboot * 06:17 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1022.eqiad.wmnet with reason: reboot * 06:03 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on dbproxy[2005-2008].codfw.wmnet with reason: reboot * 05:37 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1217,1228].eqiad.wmnet with reason: cloning * 05:11 marostegui: Drop users_to_rename table [[phab:T431842|T431842]] * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-12 == * 16:01 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2209 [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94792 and previous config saved to /var/cache/conftool/dbconfig/20260712-160124-marostegui.json * 15:58 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2205 to s3 primary [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94791 and previous config saved to /var/cache/conftool/dbconfig/20260712-155853-marostegui.json * 15:58 marostegui: Starting s3 codfw emergency failover from db2209 to db2205 - [[phab:T431950|T431950]] * 15:51 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2205 with weight 0 [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94790 and previous config saved to /var/cache/conftool/dbconfig/20260712-155135-marostegui.json * 15:51 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Primary switchover s3 [[phab:T431950|T431950]] * 02:01 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 01m 17s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-11 == * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 26s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-10 == * 19:12 jhathaway@dns1004: END - running authdns-update * 19:10 jhathaway@dns1004: START - running authdns-update * 18:23 mutante: vrts2002 rebooting (not the active host) * 18:21 mutante: lists2001, phab2003 - rebooting (not the active hosts) * 18:16 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on A:lvs-high-traffic2-codfw * 18:15 swfrench@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on A:lvs-high-traffic2-codfw * 17:15 mutante: [doc1004:~] $ sudo systemctl start rsync-doc-host-data-sync ([[phab:T431856|T431856]]) * 17:09 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1004.eqiad.wmnet * 17:08 jhathaway@dns1004: END - running authdns-update * 17:07 jhathaway@dns1004: START - running authdns-update * 17:06 jhathaway: depooling puppetserver1002, cause of errors is still unknown * 17:03 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1004.eqiad.wmnet * 16:57 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1003.eqiad.wmnet * 16:51 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1003.eqiad.wmnet * 16:48 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 16:48 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2004.codfw.wmnet * 16:42 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2004.codfw.wmnet * 16:41 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2003.codfw.wmnet * 16:35 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2003.codfw.wmnet * 16:33 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2002.codfw.wmnet * 16:27 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2002.codfw.wmnet * 16:25 mutante: gitlab-runners (production) rebooting cluster one by one * 16:17 mutante: etherpad1004/etherpad2002 - (etherpad.wikimedia.org) - rebooting * 16:13 mutante: doc1004/doc2003 (doc.wikimedia.org backends) - rebooting * 16:02 mutante: releases1003/releases2003 (releases.wikimedia.org backends) - rebooting for maintenance * 15:26 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2007-dev.codfw.wmnet * 15:19 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2007-dev.codfw.wmnet * 15:14 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host cloudcephosd2007-dev.codfw.wmnet * 15:14 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2007-dev.codfw.wmnet * 15:14 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host cloudcephosd2006-dev.codfw.wmnet * 15:07 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2006-dev.codfw.wmnet * 15:07 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2005-dev.codfw.wmnet * 14:59 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2005-dev.codfw.wmnet * 14:59 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2004-dev.codfw.wmnet * 14:53 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2004-dev.codfw.wmnet * 14:53 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2007-dev.codfw.wmnet * 14:51 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1054.eqiad.wmnet * 14:51 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1054.eqiad.wmnet * 14:51 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1054.eqiad.wmnet * 14:47 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2007-dev.codfw.wmnet * 14:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2006-dev.codfw.wmnet * 14:41 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2006-dev.codfw.wmnet * 14:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2005-dev.codfw.wmnet * 14:37 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2005-dev.codfw.wmnet * 14:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2005-dev.codfw.wmnet * 14:29 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2005-dev.codfw.wmnet * 14:29 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2006-dev.codfw.wmnet * 14:21 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2006-dev.codfw.wmnet * 14:21 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2010-dev.codfw.wmnet * 14:15 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2010-dev.codfw.wmnet * 14:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudgw2004-dev.codfw.wmnet * 14:10 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1054.eqiad.wmnet with OS trixie * 14:09 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudgw2004-dev.codfw.wmnet * 14:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudgw2003-dev.codfw.wmnet * 14:02 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudgw2003-dev.codfw.wmnet * 14:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2004-dev.codfw.wmnet * 13:53 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2004-dev.codfw.wmnet * 13:53 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2003-dev.codfw.wmnet * 13:48 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:44 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2003-dev.codfw.wmnet * 13:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2002-dev.codfw.wmnet * 13:42 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:41 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:41 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:37 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2002-dev.codfw.wmnet * 13:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudidp2001-dev.codfw.wmnet * 13:33 blake@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker1054.eqiad.wmnet with reason: host reimage * 13:33 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudidp2001-dev.codfw.wmnet * 13:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudnet2006-dev.codfw.wmnet * 13:26 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudnet2006-dev.codfw.wmnet * 13:26 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudnet2005-dev.codfw.wmnet * 13:23 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1054.eqiad.wmnet with reason: host reimage * 13:18 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudnet2005-dev.codfw.wmnet * 13:18 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudservices2005-dev.codfw.wmnet * 13:12 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudservices2005-dev.codfw.wmnet * 13:11 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudservices2004-dev.codfw.wmnet * 13:08 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudservices2004-dev.codfw.wmnet * 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudweb2002-dev.wikimedia.org * 13:05 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 13:05 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1054 * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1054 * 13:04 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1054 * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1054.eqiad.wmnet 49.32.64.10.in-addr.arpa 9.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:04 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1054.eqiad.wmnet 49.32.64.10.in-addr.arpa 9.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1054 - blake@cumin1003" * 13:04 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1054 - blake@cumin1003" * 13:01 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudweb2002-dev.wikimedia.org * 13:00 blake@cumin1003: START - Cookbook sre.dns.netbox * 12:59 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1054 * 12:57 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1054.eqiad.wmnet with OS trixie * 12:57 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1054.eqiad.wmnet * 12:56 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1054.eqiad.wmnet * 12:56 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1054.eqiad.wmnet * 12:47 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS trixie * 12:44 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:39 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 12:39 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 12:38 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:37 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:14 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:10 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:08 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:07 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:00 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:00 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:51 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:49 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:48 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:47 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:44 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:32 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2001.codfw.wmnet * 11:32 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1053.eqiad.wmnet * 11:32 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2001.codfw.wmnet * 11:32 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1053.eqiad.wmnet * 11:32 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1053.eqiad.wmnet * 11:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker2001.codfw.wmnet * 11:31 cgoubert@cumin1003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker2001.codfw.wmnet * 11:31 cgoubert@cumin1003: END (FAIL) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=1) rolling reimage on P<nowiki>{</nowiki>wikikube-worker2001*<nowiki>}</nowiki> and (A:wikikube-master-codfw or A:wikikube-worker-codfw) * 11:30 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:30 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:21 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 18 hosts with reason: reboot & upgrade * 11:20 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker2001.codfw.wmnet with OS trixie * 11:16 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:15 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:14 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:14 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 11 hosts * 11:14 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 11 hosts * 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:08 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 11:02 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1053.eqiad.wmnet with OS trixie * 11:01 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:58 cgoubert@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 10:57 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:38 cgoubert@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker2001.codfw.wmnet with OS trixie * 10:38 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2001.codfw.wmnet * 10:38 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2001.codfw.wmnet * 10:38 cgoubert@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on P<nowiki>{</nowiki>wikikube-worker2001*<nowiki>}</nowiki> and (A:wikikube-master-codfw or A:wikikube-worker-codfw) * 10:35 topranks: adjust IBGP outbound policy on lsw1-e2-codfw [[phab:T423430|T423430]] towards ssw1-e1-codfw * 10:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 cgoubert@cumin1003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:27 cgoubert@cumin1003: END (FAIL) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=1) rolling reimage on A:wikikube-worker-codfw * 10:27 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker2001.codfw.wmnet with OS bookworm * 10:25 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 10:24 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:24 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:15 cgoubert@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 10:11 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 10:11 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 10:08 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:08 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:07 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 10:06 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 10:00 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:55 cgoubert@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker2001.codfw.wmnet with OS bookworm * 09:55 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2005-2006,2011-2012].codfw.wmnet * 09:55 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2005-2006,2011-2012].codfw.wmnet * 09:51 cgoubert@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on A:wikikube-worker-codfw * 09:41 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1053.eqiad.wmnet with reason: host reimage * 09:37 topranks: apply new IBGP outbound policy on lsw1-e2-codfw [[phab:T423430|T423430]] * 09:36 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:36 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1053.eqiad.wmnet with reason: host reimage * 09:16 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1053 * 09:16 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1053 * 09:15 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1053 * 09:15 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1053.eqiad.wmnet 48.32.64.10.in-addr.arpa 8.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:15 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1053.eqiad.wmnet 48.32.64.10.in-addr.arpa 8.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:15 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:15 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1053 - blake@cumin1003" * 09:15 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1053 - blake@cumin1003" * 09:11 blake@cumin1003: START - Cookbook sre.dns.netbox * 09:11 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1053 * 09:08 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1053.eqiad.wmnet with OS trixie * 09:08 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1053.eqiad.wmnet * 09:08 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1053.eqiad.wmnet * 09:08 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1053.eqiad.wmnet * 09:04 brouberol@dns1004: END - running authdns-update * 09:03 brouberol@dns1004: START - running authdns-update * 08:41 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e] (thin): Regular analytics weekly train THIN [analytics/refinery@1abf22ea] (duration: 02m 11s) * 08:38 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e] (thin): Regular analytics weekly train THIN [analytics/refinery@1abf22ea] * 08:38 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e]: Regular analytics weekly train [analytics/refinery@1abf22ea] (duration: 05m 17s) * 08:38 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:34 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:33 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e]: Regular analytics weekly train [analytics/refinery@1abf22ea] * 08:32 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@1abf22ea] (duration: 02m 03s) * 08:30 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@1abf22ea] * 08:30 JavierMonton: Deploying Refinery at {{Gerrit|1abf22ea}} for changes 1308121/T427068 1306491/T430020 and {{Gerrit|1308190}} * 08:29 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:29 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 08:24 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:24 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 08:18 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:18 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 08:00 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db[2183-2184].codfw.wmnet * 08:00 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for db[2183-2184].codfw.wmnet * 07:52 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:52 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 07:49 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 11 hosts with reason: reboot & upgrade * 07:47 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:47 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 07:44 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:44 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 07:23 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 10 hosts * 07:23 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 10 hosts * 06:45 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 10 hosts with reason: reboot & upgrade * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 41s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-09 == * 23:33 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] (duration: 13m 26s) * 23:29 ladsgroup@deploy2003: ladsgroup, jdlrobson: Continuing with deployment * 23:22 ladsgroup@deploy2003: ladsgroup, jdlrobson: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:20 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] * 22:57 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1165.eqiad.wmnet * 22:56 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1165.eqiad.wmnet * 22:56 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1165.eqiad.wmnet * 22:45 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1165.eqiad.wmnet with OS trixie * 22:38 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 22:37 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 22:37 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 22:37 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:37 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 22:25 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1165.eqiad.wmnet with reason: host reimage * 22:17 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1165.eqiad.wmnet with reason: host reimage * 22:13 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 22:12 rzl@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 22:04 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 22:04 rzl@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1165 * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1165 * 22:02 jasmine@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1165 * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1165.eqiad.wmnet 115.48.64.10.in-addr.arpa 5.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:02 jasmine@cumin2002: START - Cookbook sre.dns.wipe-cache wikikube-worker1165.eqiad.wmnet 115.48.64.10.in-addr.arpa 5.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1165 - jasmine@cumin2002" * 22:02 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1165 - jasmine@cumin2002" * 22:02 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 21:57 jasmine@cumin2002: START - Cookbook sre.dns.netbox * 21:55 jasmine@cumin2002: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1165 * 21:54 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-worker1165.eqiad.wmnet with OS trixie * 21:54 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 21:54 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1165.eqiad.wmnet * 21:53 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 21:53 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1165.eqiad.wmnet * 21:53 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1165.eqiad.wmnet * 21:53 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 21:47 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 21:45 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 21:43 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 21:43 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 21:42 maryum: Deploy fix for [[phab:T431684|T431684]] * 21:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 21:27 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] (duration: 34m 14s) * 21:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs1002 * 21:23 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs1002 * 21:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS trixie * 21:22 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 21:20 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 22s) * 21:20 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 21:16 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 21:16 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 21:15 ladsgroup@deploy2003: ladsgroup: Continuing with deployment * 21:13 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2002.codfw.wmnet with OS bookworm * 21:11 ladsgroup@deploy2003: ladsgroup: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:08 ladsgroup@cumin1003: END (PASS) - Cookbook sre.wikireplicas.update-views (exit_code=0) * 21:07 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 6 hosts with reason: reboots * 20:54 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecycle work - bking@cumin2003 * 20:53 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:53 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] * 20:51 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) * 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 20:47 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecycle work - bking@cumin2003 * 20:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 20:41 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:41 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) * 20:40 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host relforge1008.eqiad.wmnet * 20:40 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1009.eqiad.wmnet with OS trixie * 20:33 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:32 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:32 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:31 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:31 ladsgroup@cumin1003: END (PASS) - Cookbook sre.wikireplicas.update-views (exit_code=0) * 20:29 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1008.eqiad.wmnet * 20:24 rzl@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 20:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2002.codfw.wmnet with OS bookworm * 20:23 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host relforge1008.eqiad.wmnet * 20:23 rzl@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 20:23 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1008.eqiad.wmnet * 20:22 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:22 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:21 bking@cumin2003: END (ERROR) - Cookbook sre.elasticsearch.rolling-operation (exit_code=97) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:21 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1009.eqiad.wmnet with reason: host reimage * 20:16 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:15 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1009.eqiad.wmnet with reason: host reimage * 20:12 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) * 20:02 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 19:55 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1009.eqiad.wmnet with OS trixie * 19:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 19:43 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 19:30 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 19:28 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 19:27 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 19:25 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 18:42 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 18:41 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 18:16 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 18:15 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 17:45 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for doh5004.wikimedia.org * 17:45 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for doh5004.wikimedia.org * 17:38 ladsgroup@deploy2003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 17:35 ladsgroup@deploy2003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 17:29 ladsgroup@deploy2003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 17:26 ladsgroup@deploy2003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 17:09 mutante: zuul[12]00[123] - rebooting for maintenance * 17:09 ebernhardson: start full in-place reindex of eqiad cirrussearch cluster * 17:08 dzahn@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-cluster (exit_code=99) * 17:08 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-cluster * 17:03 ebernhardson: start full in-place reindex of codfw cirrussearch cluster * 16:59 mutante: stewards1001/stewards2001 - reboot for maintenance * 16:54 ebernhardson: start full in-place reindex of cloudelastic cluster * 16:53 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 16:52 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply * 16:49 mutante: planet1003/planet2003 - rebooting * 16:47 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on doh5004.wikimedia.org with reason: random high load, investigating * 15:55 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 15:54 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 15:51 jynus: restarting backupmon1001 * 15:49 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 14 hosts * 15:49 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 14 hosts * 15:47 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backupmon1001.eqiad.wmnet with reason: restart * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:06 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 15:06 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 14:59 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply * 14:58 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply * 14:51 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 14 hosts * 14:51 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 14 hosts * 14:49 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 6 hosts with reason: reboot & upgrade * 14:48 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet * 14:48 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet * 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:42 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1052.eqiad.wmnet * 14:42 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1052.eqiad.wmnet * 14:42 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1052.eqiad.wmnet * 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:31 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:31 elukey@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: sync * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:30 elukey@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: sync * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:28 elukey@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: sync * 14:28 elukey@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: sync * 14:26 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:20 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1052.eqiad.wmnet with OS trixie * 14:19 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 6 hosts with reason: reboot & upgrade * 14:18 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:15 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:15 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:13 elukey: update druid indexation job for webrequest_sampled_live - [[phab:T427068|T427068]] * 14:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:09 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:09 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for papaul - jhancock@cumin2002" * 14:09 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for papaul - jhancock@cumin2002" * 14:07 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:07 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:04 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 14:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cuminunpriv1001.eqiad.wmnet * 13:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb1003.eqiad.wmnet * 13:59 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1052.eqiad.wmnet with reason: host reimage * 13:57 moritzm: installing requests security updates * 13:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cuminunpriv1001.eqiad.wmnet * 13:55 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb1003.eqiad.wmnet * 13:53 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1052.eqiad.wmnet with reason: host reimage * 13:50 moritzm: installing python-cryptography security updates * 13:47 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb2003.codfw.wmnet * 13:44 Msz2001: UTC afternoon config+backport window is done * 13:44 Msz2001: Updated `logging` on `metawiki` to fix log performers, [[phab:T431176|T431176]]#12105297 * 13:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb2003.codfw.wmnet * 13:43 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 13:43 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt1002.wikimedia.org * 13:41 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] (duration: 07m 30s) * 13:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt1002.wikimedia.org * 13:37 mszwarc@deploy2003: mszwarc: Continuing with deployment * 13:36 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1052 * 13:36 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1052 * 13:35 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:35 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1052 * 13:35 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1052.eqiad.wmnet 47.32.64.10.in-addr.arpa 7.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:35 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1052.eqiad.wmnet 47.32.64.10.in-addr.arpa 7.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:35 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:35 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1052 - blake@cumin1003" * 13:35 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1052 - blake@cumin1003" * 13:34 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] * 13:31 blake@cumin1003: START - Cookbook sre.dns.netbox * 13:31 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1052 * 13:30 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1052.eqiad.wmnet with OS trixie * 13:30 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1052.eqiad.wmnet * 13:29 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1052.eqiad.wmnet * 13:29 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1052.eqiad.wmnet * 13:17 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] (duration: 11m 26s) * 13:13 jforrester@deploy2003: jforrester: Continuing with deployment * 13:08 jforrester@deploy2003: jforrester: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:06 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] * 12:54 cgoubert@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply * 12:54 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:52 cgoubert@deploy2003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply * 12:45 cgoubert@deploy2003: helmfile [codfw] DONE helmfile.d/services/mobileapps: apply * 12:44 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:44 cgoubert@deploy2003: helmfile [codfw] START helmfile.d/services/mobileapps: apply * 12:43 cgoubert@deploy2003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 12:43 cgoubert@deploy2003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 12:42 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast4006.wikimedia.org * 12:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt2002.wikimedia.org * 12:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast7002.wikimedia.org * 12:18 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast4006.wikimedia.org * 12:18 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host ml-serve1004 * 12:18 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host ml-serve1004 * 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt2002.wikimedia.org * 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast7002.wikimedia.org * 12:10 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backup[2003,2014].codfw.wmnet with reason: reboot & upgrade * 12:10 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt-staging2001.codfw.wmnet * 12:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid1003.eqiad.wmnet * 12:06 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt-staging2001.codfw.wmnet * 12:05 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid1003.eqiad.wmnet * 12:03 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backup[1003,1014].eqiad.wmnet with reason: reboot & upgrade * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid2003.codfw.wmnet * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host irc1003.wikimedia.org * 11:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid2003.codfw.wmnet * 11:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host irc1003.wikimedia.org * 11:55 jmm@dns1004: END - running authdns-update * 11:53 jmm@dns1004: START - running authdns-update * 11:50 jmm@dns1004: END - running authdns-update * 11:48 jmm@dns1004: START - running authdns-update * 11:27 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host irc2003.wikimedia.org * 11:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host irc2003.wikimedia.org * 11:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint2001.codfw.wmnet * 11:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint1001.eqiad.wmnet * 11:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint2001.codfw.wmnet * 11:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint1001.eqiad.wmnet * 11:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-rw2001.wikimedia.org * 11:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-rw1001.wikimedia.org * 11:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-rw2001.wikimedia.org * 11:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-rw1001.wikimedia.org * 11:03 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon1003.wikimedia.org * 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2005.codfw.wmnet * 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2005.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 10:59 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2005.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 10:57 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon1003.wikimedia.org * 10:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon2002.wikimedia.org * 10:55 jmm@cumin2003: START - Cookbook sre.dns.netbox * 10:51 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon2002.wikimedia.org * 10:51 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:50 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2005.codfw.wmnet * 10:41 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:40 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host ml-serve1003 * 10:40 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host ml-serve1003 * 10:39 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2033.codfw.wmnet * 10:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install2005.wikimedia.org * 10:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install1005.wikimedia.org * 10:35 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1004.eqiad.wmnet with OS bookworm * 10:31 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install1005.wikimedia.org * 10:31 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install2005.wikimedia.org * 10:30 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install4004.wikimedia.org * 10:30 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install3004.wikimedia.org * 10:29 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install3004.wikimedia.org * 10:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install4004.wikimedia.org * 10:23 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 10:21 moritzm: failover Ganeti master in codfw/routed to ganeti2034 [[phab:T430928|T430928]] * 10:19 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.addnode (exit_code=0) for new host ganeti2031.codfw.wmnet to cluster codfw and group B * 10:19 moritzm: readded ganeti2031 to the codfw Ganeti cluster [[phab:T430910|T430910]] * 10:18 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1004.eqiad.wmnet with reason: host reimage * 10:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install5004.wikimedia.org * 10:18 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1003 * 10:18 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1003 * 10:17 jmm@cumin2003: START - Cookbook sre.ganeti.addnode for new host ganeti2031.codfw.wmnet to cluster codfw and group B * 10:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install6003.wikimedia.org * 10:16 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install5004.wikimedia.org * 10:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install6003.wikimedia.org * 10:15 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1004.eqiad.wmnet with reason: host reimage * 10:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1001.eqiad.wmnet * 10:14 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 10:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2008.wikimedia.org * 10:00 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ml-serve1004 * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1004 * 09:57 jmm@cumin2003: START - Cookbook sre.dns.netbox * 09:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install7002.wikimedia.org * 09:57 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1004 * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ml-serve1004.eqiad.wmnet 50.48.64.10.in-addr.arpa 0.5.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:57 klausman@cumin1003: START - Cookbook sre.dns.wipe-cache ml-serve1004.eqiad.wmnet 50.48.64.10.in-addr.arpa 0.5.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1004 - klausman@cumin1003" * 09:56 klausman@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1004 - klausman@cumin1003" * 09:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-coord1001.eqiad.wmnet * 09:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 09:55 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow7002.magru.wmnet * 09:52 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-coord1001.eqiad.wmnet * 09:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 09:52 klausman@cumin1003: START - Cookbook sre.dns.netbox * 09:50 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install7002.wikimedia.org * 09:50 klausman@cumin1003: START - Cookbook sre.hosts.move-vlan for host ml-serve1004 * 09:50 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1004.eqiad.wmnet with OS bookworm * 09:50 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1003.eqiad.wmnet with OS bookworm * 09:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1001.eqiad.wmnet * 09:49 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 09:49 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2008.wikimedia.org * 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2007.codfw.wmnet * 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2007.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 09:49 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow7002.magru.wmnet * 09:49 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2007.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 09:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard1003.eqiad.wmnet * 09:39 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard2003.codfw.wmnet * 09:39 jmm@cumin2003: START - Cookbook sre.dns.netbox * 09:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard1003.eqiad.wmnet * 09:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor1003.eqiad.wmnet * 09:35 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard2003.codfw.wmnet * 09:34 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2007.codfw.wmnet * 09:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor1003.eqiad.wmnet * 09:33 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor-dev2001.codfw.wmnet * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor2003.codfw.wmnet * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sretest1006.eqiad.wmnet * 09:27 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 09:25 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor-dev2001.codfw.wmnet * 09:25 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor2003.codfw.wmnet * 09:23 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] (duration: 06m 27s) * 09:23 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2205: codfw rack B4 repool after maintenance * 09:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host sretest1006.eqiad.wmnet * 09:23 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2204: codfw rack B4 repool after maintenance * 09:19 urbanecm@deploy2003: urbanecm: Continuing with deployment * 09:19 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:18 jmm@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 6 hosts with reason: reboot * 09:17 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] * 09:08 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ml-serve1003 * 09:08 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1003 * 09:07 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1003 * 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ml-serve1003.eqiad.wmnet 81.32.64.10.in-addr.arpa 1.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:07 klausman@cumin1003: START - Cookbook sre.dns.wipe-cache ml-serve1003.eqiad.wmnet 81.32.64.10.in-addr.arpa 1.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1003 - klausman@cumin1003" * 09:06 klausman@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1003 - klausman@cumin1003" * 08:58 klausman@cumin1003: START - Cookbook sre.dns.netbox * 08:57 klausman@cumin1003: START - Cookbook sre.hosts.move-vlan for host ml-serve1003 * 08:57 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1003.eqiad.wmnet with OS bookworm * 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=0) rolling reimage on P<nowiki>{</nowiki>ml-serve1003.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet * 08:55 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet * 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1003.eqiad.wmnet with OS bookworm * 08:39 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 08:38 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool db2205: codfw rack B4 repool after maintenance * 08:37 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool db2204: codfw rack B4 repool after maintenance * 08:36 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 08:35 hashar@deploy2003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 08:32 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:32 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:31 hashar@deploy2003: Rolling back deployment * 08:26 moritzm: failover Ganeti master in codfw to ganeti2048 [[phab:T430928|T430928]] * 08:16 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1003.eqiad.wmnet with OS bookworm * 08:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2004.codfw.wmnet * 08:16 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet * 08:16 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet * 08:16 klausman@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on P<nowiki>{</nowiki>ml-serve1003.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 08:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2002.codfw.wmnet * 08:15 XioNoX: lsw1-b4-codfw> request system reboot - [[phab:T430910|T430910]] * 08:15 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b4-codfw,lsw1-b4-codfw IPv6,lsw1-b4-codfw.mgmt with reason: Switch maintenance * 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for codfw rack B4 * 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:10 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2004.codfw.wmnet * 08:10 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2002.codfw.wmnet * 08:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2205: codfw rack B4 depool for maintenance * 08:08 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool db2205: codfw rack B4 depool for maintenance * 08:08 jmm@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin2003.codfw.wmnet * 08:08 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2204: codfw rack B4 depool for maintenance * 08:08 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool db2204: codfw rack B4 depool for maintenance * 08:08 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 27 hosts with reason: codfw rack B4 depool for maintenance * 08:03 jmm@cumin2002: START - Cookbook sre.hosts.reboot-single for host cumin2003.codfw.wmnet * 07:56 ayounsi@cumin1003: START - Cookbook sre.network.depool-rack with action 'depool' for codfw rack B4 * 07:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1008.eqiad.wmnet with OS trixie * 07:49 wmde-fisch@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] (duration: 08m 36s) * 07:44 wmde-fisch@deploy2003: wmde-fisch: Continuing with deployment * 07:43 wmde-fisch@deploy2003: wmde-fisch: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:41 wmde-fisch@deploy2003: Started scap sync-world: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] * 07:35 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1008.eqiad.wmnet with reason: host reimage * 07:31 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1008.eqiad.wmnet with reason: host reimage * 07:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1008.eqiad.wmnet with OS trixie * 07:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 07:00 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 06:59 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1008.eqiad.wmnet with OS trixie * 06:57 Emperor: rebalance thanos swift rings after previous re-image of thanos-fe1004 to trixie * 06:47 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1008.eqiad.wmnet with OS trixie * 04:10 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 14 days, 0:00:00 on cp6008.drmrs.wmnet with reason: Hardware failure - [[phab:T431651|T431651]] * 03:55 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp6008.* * 03:29 ryankemper: [[phab:T431311|T431311]] Repooled eqiad cirrussearch clusters (`chi/omega/psi`) following completion of OpenSearch 2.19 migration * 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad * 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=eqiad * 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 31s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-08 == * 23:52 Amir1: ladsgroup@deploy2003:~$ mwscript-k8s --follow -- extensions/ORES/maintenance/PurgeScoreCache.php --wiki=simplewiki --model damaging --old ([[phab:T431159|T431159]]) * 23:46 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 23:46 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing PTR for 2001:df2:e500:fe08::1 - cmooney@cumin1003" * 23:46 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing PTR for 2001:df2:e500:fe08::1 - cmooney@cumin1003" * 23:40 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 23:16 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 23:15 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 22:42 rzl@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 22:40 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] (duration: 12m 55s) * 22:40 rzl@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 22:37 rzl@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 22:36 rzl@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 22:35 rzl@deploy2003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 22:34 urbanecm@deploy2003: urbanecm: Continuing with deployment * 22:33 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:33 rzl@deploy2003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 22:32 rzl@deploy2003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 22:30 rzl@deploy2003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 22:30 rzl@deploy2003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 22:29 rzl@deploy2003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 22:27 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] * 22:26 rzl@deploy2003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 22:22 rzl@deploy2003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 22:21 rzl@deploy2003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 22:19 rzl@deploy2003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 22:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 22:17 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 22:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 22:13 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 22:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 22:13 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 22:09 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 22:06 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 22:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1094.eqiad.wmnet with OS trixie * 22:01 urbanecm: Make https://test.wikipedia.org/w/index.php?title=MediaWiki:GrowthExperimentsSuggestedEdits.json&diff=prev&oldid=750552 with GrowthExperiments disabled (via mw-experimental), then run `\MediaWiki\MediaWikiServices::getInstance()->get('CommunityConfiguration.ProviderFactory')->newProvider('GrowthSuggestedEdits')->getStore()->invalidate()` ([[phab:T431625|T431625]]) * 21:56 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d2-codfw * 21:55 urbanecm@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 21:55 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d2-codfw * 21:55 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c4-codfw * 21:55 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c4-codfw * 21:55 urbanecm@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2002 * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2002 * 21:54 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2002 * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2002.codfw.wmnet 50.32.192.10.in-addr.arpa 0.5.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:54 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2002.codfw.wmnet 50.32.192.10.in-addr.arpa 0.5.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2002 - bking@cumin2003" * 21:54 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2002 - bking@cumin2003" * 21:49 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:49 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2002 * 21:49 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2002.codfw.wmnet with OS trixie * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1094.eqiad.wmnet with reason: host reimage * 21:42 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 21:39 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 21:37 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1094.eqiad.wmnet with reason: host reimage * 21:36 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 21:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 21:29 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 21:27 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 21:22 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1094.eqiad.wmnet with OS trixie * 21:21 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host restbase2039.codfw.wmnet with OS bullseye * 21:21 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin2002" * 21:21 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin2002" * 21:04 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on restbase2039.codfw.wmnet with reason: host reimage * 21:00 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on restbase2039.codfw.wmnet with reason: host reimage * 20:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1073.eqiad.wmnet with OS trixie * 20:48 mutante: deploy2003 - kill 1102 (stunnel4) ; systemctl start stunnel4 ([[phab:T418262|T418262]]) * 20:42 cjming@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] (duration: 33m 02s) * 20:42 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host restbase2039.codfw.wmnet with OS bullseye * 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1073.eqiad.wmnet with reason: host reimage * 20:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1073.eqiad.wmnet with reason: host reimage * 20:30 cjming@deploy2003: cjming: Continuing with deployment * 20:28 cjming@deploy2003: cjming: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1098.eqiad.wmnet with OS trixie * 20:13 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1073.eqiad.wmnet with OS trixie * 20:09 cjming@deploy2003: Started scap sync-world: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] * 20:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1098.eqiad.wmnet with reason: host reimage * 19:56 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1098.eqiad.wmnet with reason: host reimage * 19:55 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d4-codfw * 19:54 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d4-codfw * 19:54 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c1-codfw * 19:54 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c1-codfw * 19:52 mutante: restarting gerrit on gerrit.wikimedia.org (gerrit2003) * 19:48 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2331.codfw.wmnet * 19:48 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2331.codfw.wmnet * 19:48 mutante: restarting gerrit on gerrit-replica.wikimedia.org (gerrit1003) * 19:46 mutante: restarting gerrit on gerrit-spare.wikimedia.org (gerrit2002) * 19:43 jasmine@cumin2002: conftool action : set/pooled=yes; selector: name=wikikube-worker2331.codfw.wmnet,cluster=kubernetes,service=kubesvc * 19:43 jasmine@cumin2002: conftool action : set/weight=10; selector: name=wikikube-worker2331.codfw.wmnet,cluster=kubernetes,service=kubesvc * 19:40 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1098.eqiad.wmnet with OS trixie * 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d5-codfw * 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d5-codfw * 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c7-codfw * 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c7-codfw * 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c5-codfw * 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c5-codfw * 19:30 jasmine_: ran homer on lsw1-d8-codfw, adding wikikube-worker2331 to cluster * 19:29 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1100.eqiad.wmnet with OS trixie * 19:20 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d8-codfw * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d8-codfw * 19:19 mutante: gerrit - replacing private key for registerEmail verification * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-magru * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device cr2-magru * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d7-codfw * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d7-codfw * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d3-codfw * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d1-codfw * 19:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d1-codfw * 19:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c2-codfw * 19:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c2-codfw * 19:11 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-codfw * 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-magru * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device cr1-magru * 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d8-codfw * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d8-codfw * 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d6-codfw * 19:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1100.eqiad.wmnet with reason: host reimage * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d6-codfw * 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c6-codfw * 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c6-codfw * 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c3-codfw * 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c3-codfw * 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b4-magru * 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device asw1-b4-magru * 19:08 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b3-magru * 19:08 cmooney@cumin1003: START - Cookbook sre.network.tls for network device asw1-b3-magru * 19:05 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1100.eqiad.wmnet with reason: host reimage * 19:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1122.eqiad.wmnet with OS trixie * 19:00 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 18:59 topranks: rolling out update to BGP ACL on Nokia Switches eqiad, codfw & ulsfo [[phab:T425703|T425703]] * 18:58 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 18:57 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 18:55 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 18:53 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 18:52 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 18:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1100.eqiad.wmnet with OS trixie * 18:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1068.eqiad.wmnet with OS trixie * 18:47 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1102.eqiad.wmnet with OS trixie * 18:47 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 18:46 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1122.eqiad.wmnet with reason: host reimage * 18:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1122.eqiad.wmnet with reason: host reimage * 18:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1068.eqiad.wmnet with reason: host reimage * 18:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1122 * 18:26 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1122 * 18:25 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1122 * 18:25 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1122.eqiad.wmnet 31.48.64.10.in-addr.arpa 1.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:25 bking@cumin2003: START - Cookbook sre.dns.wipe-cache cirrussearch1122.eqiad.wmnet 31.48.64.10.in-addr.arpa 1.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:25 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:25 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1122 - bking@cumin2003" * 18:25 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1122 - bking@cumin2003" * 18:21 rzl@deploy2003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 18:21 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1068.eqiad.wmnet with reason: host reimage * 18:21 rzl@deploy2003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 18:21 rzl@deploy2003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 18:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 18:19 rzl@deploy2003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 18:19 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:18 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1122 * 18:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1122.eqiad.wmnet with OS trixie * 18:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 18:15 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 18:13 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 18:13 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 18:10 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 18:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1068.eqiad.wmnet with OS trixie * 18:01 kamila@deploy2003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 18m 29s) * 18:00 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:55 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:42 kamila@deploy2003: Started scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] * 17:42 kamila@deploy2003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 19m 50s) * 17:42 kamila@deploy2003: Rolling back deployment * 17:35 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:31 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet * 17:18 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet * 17:16 kamila@deploy1003: Unlocked for deployment [MediaWiki]: switching deployment server (duration: 22m 07s) * 17:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 17:11 kamila@dns1005: END - running authdns-update * 17:09 kamila@dns1005: START - running authdns-update * 17:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 17:04 jasmine@cumin2002: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1164.eqiad.wmnet * 17:04 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1164.eqiad.wmnet * 17:04 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1164.eqiad.wmnet * 16:56 kamila@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on releases2003.codfw.wmnet,releases1003.eqiad.wmnet with reason: Deployment server switchover * 16:54 kamila@deploy1003: Locking from deployment [MediaWiki]: switching deployment server * 16:53 kamila@deploy1003: Unlocked for deployment [MediaWiki]: switching deployment server (duration: 04m 02s) * 16:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie * 16:49 kamila@deploy1003: Locking from deployment [MediaWiki]: switching deployment server * 16:46 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1095.eqiad.wmnet with OS trixie * 16:45 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1093.eqiad.wmnet with OS trixie * 16:43 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1164.eqiad.wmnet with OS trixie * 16:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1095.eqiad.wmnet with reason: host reimage * 16:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 16:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 16:23 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1164.eqiad.wmnet with reason: host reimage * 16:18 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on cirrussearch1093.eqiad.wmnet with reason: host reimage * 16:16 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1164.eqiad.wmnet with reason: host reimage * 16:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1095.eqiad.wmnet with reason: host reimage * 16:09 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 16:09 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 16:08 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1093.eqiad.wmnet with reason: host reimage * 15:59 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Pool test * 15:59 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:59 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 15:59 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Pool test * 15:58 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Depool test * 15:58 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:58 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 15:58 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Depool test * 15:57 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1164 * 15:57 jasmine@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1164 * 15:57 jasmine@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1164 * 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1164.eqiad.wmnet 114.48.64.10.in-addr.arpa 4.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:56 jasmine@cumin2002: START - Cookbook sre.dns.wipe-cache wikikube-worker1164.eqiad.wmnet 114.48.64.10.in-addr.arpa 4.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1164 - jasmine@cumin2002" * 15:56 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1164 - jasmine@cumin2002" * 15:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1093.eqiad.wmnet with OS trixie * 15:51 jasmine@cumin2002: START - Cookbook sre.dns.netbox * 15:51 jasmine@cumin2002: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1164 * 15:50 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-worker1164.eqiad.wmnet with OS trixie * 15:50 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1164.eqiad.wmnet * 15:50 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1164.eqiad.wmnet * 15:50 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1164.eqiad.wmnet * 15:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1095.eqiad.wmnet with OS trixie * 15:42 jasmine@cumin2002: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1164.eqiad.wmnet * 15:42 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1164.eqiad.wmnet * 15:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:42 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1164.eqiad.wmnet * 15:42 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1164.eqiad.wmnet * 15:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 15:39 elukey@cumin1003: START - Cookbook sre.hosts.provision for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 15:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1007.eqiad.wmnet with OS trixie * 15:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Pool test * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1007.eqiad.wmnet with reason: host reimage * 15:15 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 15:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet * 15:15 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 15:15 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1007.eqiad.wmnet with reason: host reimage * 15:15 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:14 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Pool test * 15:14 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet * 15:14 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2228: Depool test * 15:14 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db2228: Depool test * 15:10 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 15:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:08 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 15:06 blake@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 15:06 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet * 15:06 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 15:06 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 15:05 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet * 15:05 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 15:04 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:04 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 15:04 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:03 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 15:03 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 15:03 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:03 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T430909|T430909]] * 15:03 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:03 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 15:03 swfrench-wmf: restarted eqsin, codfw confds - [[phab:T430909|T430909]] * 15:03 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test1002.eqiad.wmnet * 15:02 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet * 14:59 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:59 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:55 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1007.eqiad.wmnet with OS trixie * 14:52 swfrench-wmf: restarted ulsfo confds, confirmed now connected to codfw backends except those using wikimedia.org SRV record - [[phab:T430909|T430909]] * 14:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:49 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:41 moritzm: uninstalling dhcpcd-base from trixie hosts which still have it installed [[phab:T414341|T414341]] * 14:40 sukhe: sudo cumin -b1 -s120 "P<nowiki>{</nowiki>lvs2011*<nowiki>}</nowiki> or P<nowiki>{</nowiki>lvs2012*<nowiki>}</nowiki>" "systemctl restart pybal.service" * 14:39 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:39 mvernon@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host thanos-be1007.eqiad.wmnet with OS trixie * 14:37 sukhe: restart pybal on lvs2013 to revert back to conf2004 * 14:35 sukhe: restart pybal on lvs2014 to revert back to conf2004 * 14:34 swfrench-wmf: switched codfw, eqsin, ulsfo etcd client SRV records back to codfw - [[phab:T430909|T430909]] * 14:32 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1002.eqiad.wmnet * 14:32 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet * 14:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1007.eqiad.wmnet with OS trixie * 14:31 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:31 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:31 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:30 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:30 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Pool test * 14:30 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:29 swfrench@dns1004: END - running authdns-update * 14:29 moritzm: installing jackson-core security updates * 14:27 swfrench@dns1004: START - running authdns-update * 14:22 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:22 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:22 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1119.eqiad.wmnet with OS trixie * 14:22 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:21 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:20 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:20 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 14:20 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:19 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 14:19 moritzm: installing librabbitmq security updates * 14:19 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1002.eqiad.wmnet * 14:18 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet * 14:16 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:15 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:15 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:14 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Pool test * 14:14 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox) * 14:14 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet * 14:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 14:08 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 14:07 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet * 14:05 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1118.eqiad.wmnet with OS trixie * 14:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test1001.eqiad.wmnet * 14:00 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-worker@eqiad * 14:00 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 13:59 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 13:58 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 13:57 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1119.eqiad.wmnet with reason: host reimage * 13:54 moritzm: installing libcap2 security updates * 13:53 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1119.eqiad.wmnet with reason: host reimage * 13:52 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet * 13:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) * 13:52 fceratto@cumin1003: START - Cookbook sre.mysql.depool * 13:50 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-worker@eqiad * 13:50 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1051.eqiad.wmnet * 13:50 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1051.eqiad.wmnet * 13:50 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1051.eqiad.wmnet * 13:49 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 13:45 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1081.eqiad.wmnet with OS trixie * 13:41 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1119 * 13:41 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1119 * 13:40 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1119 * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1119.eqiad.wmnet 97.32.64.10.in-addr.arpa 7.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1119.eqiad.wmnet 97.32.64.10.in-addr.arpa 7.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1119 - atsuko@cumin1003" * 13:40 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1119 - atsuko@cumin1003" * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1118.eqiad.wmnet with reason: host reimage * 13:39 moritzm: installing krb5 security updates * 13:37 Lucas_WMDE: UTC afternoon backport+config window done * 13:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1006.eqiad.wmnet with OS trixie * 13:36 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1118.eqiad.wmnet with reason: host reimage * 13:36 atsuko@cumin1003: START - Cookbook sre.dns.netbox * 13:35 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] (duration: 07m 46s) * 13:34 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1119 * 13:34 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1119.eqiad.wmnet with OS trixie * 13:30 sbisson@deploy1003: sbisson: Continuing with deployment * 13:30 moritzm: installing openssh security updates * 13:30 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling restart_daemons on A:wikidough * 13:29 sbisson@deploy1003: sbisson: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:27 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] * 13:26 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1051.eqiad.wmnet with OS trixie * 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1081.eqiad.wmnet with reason: host reimage * 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1118 * 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1118 * 13:22 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] (duration: 12m 12s) * 13:21 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1081.eqiad.wmnet with reason: host reimage * 13:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1006.eqiad.wmnet with reason: host reimage * 13:18 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1118 * 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1118.eqiad.wmnet 90.32.64.10.in-addr.arpa 0.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:18 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1118.eqiad.wmnet 90.32.64.10.in-addr.arpa 0.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1118 - atsuko@cumin1003" * 13:18 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1118 - atsuko@cumin1003" * 13:17 stran@deploy1003: stran: Continuing with deployment * 13:16 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough * 13:15 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:13 atsuko@cumin1003: START - Cookbook sre.dns.netbox * 13:12 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1118 * 13:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1006.eqiad.wmnet with reason: host reimage * 13:12 stran@deploy1003: stran: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:12 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1118.eqiad.wmnet with OS trixie * 13:10 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] * 13:05 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-worker@codfw * 13:05 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 13:05 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1081.eqiad.wmnet with OS trixie * 13:05 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1051.eqiad.wmnet with reason: host reimage * 13:04 moritzm: installing jq security updates * 13:04 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 13:01 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1051.eqiad.wmnet with reason: host reimage * 12:58 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-worker@codfw * 12:52 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 12:50 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host thanos-be1006.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1051 * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1051 * 12:43 moritzm: installing Python 3.11 security updates * 12:43 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1051 * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1051.eqiad.wmnet 46.32.64.10.in-addr.arpa 6.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:43 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1051.eqiad.wmnet 46.32.64.10.in-addr.arpa 6.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1051 - blake@cumin1003" * 12:43 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1051 - blake@cumin1003" * 12:38 blake@cumin1003: START - Cookbook sre.dns.netbox * 12:38 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1051 * 12:38 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1051.eqiad.wmnet with OS trixie * 12:37 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1051.eqiad.wmnet * 12:36 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1051.eqiad.wmnet * 12:36 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1051.eqiad.wmnet * 12:34 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1006.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 12:34 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1006.eqiad.wmnet with OS trixie * 12:27 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 12:27 mvernon@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host thanos-be1006.eqiad.wmnet with OS trixie * 12:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:02 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 12:01 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1006.eqiad.wmnet with OS trixie * 11:43 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 11:38 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1076.eqiad.wmnet with OS trixie * 11:26 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1075.eqiad.wmnet with OS trixie * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2047.codfw.wmnet * 11:19 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2047.codfw.wmnet * 11:19 moritzm: temporarily remove ganeti2031 from codfw cluster [[phab:T430910|T430910]] * 11:08 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1076.eqiad.wmnet with reason: host reimage * 11:08 moritzm: installing Linux 6.1.176 on Bookworm servers * 11:03 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1076.eqiad.wmnet with reason: host reimage * 11:00 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1075.eqiad.wmnet with reason: host reimage * 10:56 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1075.eqiad.wmnet with reason: host reimage * 10:47 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1076.eqiad.wmnet with OS trixie * 10:46 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1074.eqiad.wmnet with OS trixie * 10:45 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1005.eqiad.wmnet with OS trixie * 10:40 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1075.eqiad.wmnet with OS trixie * 10:32 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2031.codfw.wmnet * 10:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1005.eqiad.wmnet with reason: host reimage * 10:25 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1005.eqiad.wmnet with reason: host reimage * 10:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1074.eqiad.wmnet with reason: host reimage * 10:17 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1074.eqiad.wmnet with reason: host reimage * 10:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1005.eqiad.wmnet with OS trixie * 10:12 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet * 10:04 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 10:01 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet * 10:01 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1074.eqiad.wmnet with OS trixie * 10:01 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 09:43 cgoubert@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/aux-k8s-services/redioscope: apply * 09:43 cgoubert@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/aux-k8s-services/redioscope: apply * 09:43 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply * 09:35 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply * 09:34 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 09:34 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 09:33 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 41 days, 15:00:00 on db2252.codfw.wmnet with reason: Test * 09:32 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 09:32 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 09:31 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: codfw rack B3 pool after maintenance * 09:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 09:07 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 09:07 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 09:02 ladsgroup@cumin1003: END (PASS) - Cookbook sre.mysql.sanitarium_restart (exit_code=0) * 08:57 topranks: merge patch to shift eqiad <-> esams traffic onto new 40G circuit * 08:54 hashar@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1004.eqiad.wmnet with OS trixie * 08:50 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 08:50 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitarium_restart (exit_code=99) * 08:50 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 08:45 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool es2051: codfw rack B3 pool after maintenance * 08:44 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:44 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:43 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2007.codfw.wmnet * 08:43 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2007.codfw.wmnet * 08:42 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2031.codfw.wmnet * 08:41 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2031.codfw.wmnet * 08:40 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2031.codfw.wmnet * 08:38 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:38 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:35 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.sanitize-wiki (exit_code=97) Managing sanitization for wikis minwikiquote in section s3 * 08:33 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis minwikiquote in section s3 * 08:32 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Checking sanitization for wikis minwikiquote in section s5 * 08:30 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Checking sanitization for wikis minwikiquote in section s5 * 08:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Managing sanitization for wikis minwikiquote in section s5 * 08:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1004.eqiad.wmnet with reason: host reimage * 08:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1004.eqiad.wmnet with reason: host reimage * 08:23 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:22 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis minwikiquote in section s5 * 08:19 XioNoX: lsw1-b3-codfw> request system reboot - [[phab:T430909|T430909]] * 08:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Checking sanitization for wikis minwikiquote in section s5 * 08:17 hashar@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Checking sanitization for wikis minwikiquote in section s5 * 08:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for codfw rack B3 * 08:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2007.codfw.wmnet * 08:15 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lsw1-b3-codfw,lsw1-b3-codfw IPv6,lsw1-b3-codfw.mgmt with reason: Switch maintenance * 08:15 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2007.codfw.wmnet * 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:07 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:06 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: codfw rack B3 depool for maintenance * 08:05 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool es2051: codfw rack B3 depool for maintenance * 08:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1004.eqiad.wmnet with OS trixie * 08:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1005.eqiad.wmnet with OS trixie * 08:03 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 21 hosts with reason: codfw rack B3 depool for maintenance * 07:56 ayounsi@cumin1003: START - Cookbook sre.network.depool-rack with action 'depool' for codfw rack B3 * 07:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1005.eqiad.wmnet with reason: host reimage * 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1005.eqiad.wmnet with reason: host reimage * 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1005.eqiad.wmnet with OS bookworm * 07:29 moritzm: installing gnutls28 security updates * 07:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1005.eqiad.wmnet with OS trixie * 07:13 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1125.eqiad.wmnet with OS trixie * 07:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1005.eqiad.wmnet with reason: host reimage * 07:07 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aux-k8s-etcd1005.eqiad.wmnet with reason: host reimage * 06:56 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1005.eqiad.wmnet with OS bookworm * 06:54 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1125.eqiad.wmnet with reason: host reimage * 06:52 elukey: upgrade all trixie hosts to pywmflib 3.1 - [[phab:T430552|T430552]] * 06:50 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1125.eqiad.wmnet with reason: host reimage * 06:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 06:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 06:38 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1125.eqiad.wmnet with OS trixie * 05:42 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1107.eqiad.wmnet with OS trixie * 05:35 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1124.eqiad.wmnet with OS trixie * 05:31 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1101.eqiad.wmnet with OS trixie * 05:21 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1107.eqiad.wmnet with reason: host reimage * 05:17 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1124.eqiad.wmnet with reason: host reimage * 05:13 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1107.eqiad.wmnet with reason: host reimage * 05:13 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1101.eqiad.wmnet with reason: host reimage * 05:11 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1124.eqiad.wmnet with reason: host reimage * 05:10 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1101.eqiad.wmnet with reason: host reimage * 04:58 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1124.eqiad.wmnet with OS trixie * 04:56 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1107.eqiad.wmnet with OS trixie * 04:55 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1101.eqiad.wmnet with OS trixie * 02:27 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] (duration: 08m 14s) * 02:22 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 02:21 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 02:19 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] * 01:59 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] (duration: 09m 46s) * 01:55 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 01:51 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 01:49 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] * 01:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1099.eqiad.wmnet with OS trixie * 00:57 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1110.eqiad.wmnet with OS trixie * 00:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1099.eqiad.wmnet with reason: host reimage * 00:41 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1099.eqiad.wmnet with reason: host reimage * 00:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1110.eqiad.wmnet with reason: host reimage * 00:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1110.eqiad.wmnet with reason: host reimage * 00:26 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1099.eqiad.wmnet with OS trixie * 00:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1110.eqiad.wmnet with OS trixie == 2026-07-07 == * 22:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1097.eqiad.wmnet with OS trixie * 22:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1097.eqiad.wmnet with reason: host reimage * 22:24 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1097.eqiad.wmnet with reason: host reimage * 22:14 hashar: Restarting Gerrit on gerrit2002 and gerrit1003 (replicas) * 22:09 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1097.eqiad.wmnet with OS trixie * 22:07 hashar: Restarting Gerrit on gerrit2003 * 21:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 21:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 21:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 21:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 21:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1108.eqiad.wmnet with OS trixie * 20:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1091.eqiad.wmnet with OS trixie * 20:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1090.eqiad.wmnet with OS trixie * 20:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1108.eqiad.wmnet with reason: host reimage * 20:36 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1091.eqiad.wmnet with reason: host reimage * 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1090.eqiad.wmnet with reason: host reimage * 20:33 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1091.eqiad.wmnet with reason: host reimage * 20:30 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1108.eqiad.wmnet with reason: host reimage * 20:30 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1006.eqiad.wmnet * 20:30 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1090.eqiad.wmnet with reason: host reimage * 20:30 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1006.eqiad.wmnet * 20:27 jasmine_: "homer lsw1-c2-eqiad* commit "Added new stacked control plane wikikube-ctrl1006"" * 20:22 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] (duration: 07m 29s) * 20:20 jasmine_: "homer "cr*eqiad*" commit "Added new stacked control plane wikikube-ctrl1006"" * 20:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1091.eqiad.wmnet with OS trixie * 20:17 arlolra@deploy1003: arlolra: Continuing with deployment * 20:16 arlolra@deploy1003: arlolra: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:16 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1090.eqiad.wmnet with OS trixie * 20:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1108.eqiad.wmnet with OS trixie * 20:14 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] * 20:09 cwhite: remove 2026-04 swift log archives from centrallog2002 to free some space * 20:01 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=93) for host cirrussearch1108.eqiad.wmnet with OS trixie * 19:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1108.eqiad.wmnet with OS trixie * 19:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1090.eqiad.wmnet with OS trixie * 19:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1109.eqiad.wmnet with OS trixie * 19:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1092.eqiad.wmnet with OS trixie * 19:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1123.eqiad.wmnet with OS trixie * 19:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1109.eqiad.wmnet with reason: host reimage * 19:19 cdobbins@cumin2002: conftool action : set/pooled=yes; selector: name=dns7002.* * 19:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1092.eqiad.wmnet with reason: host reimage * 19:17 jasmine@dns1004: END - running authdns-update * 19:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1109.eqiad.wmnet with reason: host reimage * 19:15 jasmine@dns1004: START - running authdns-update * 19:15 cdobbins@dns1004: END - running authdns-update * 19:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1123.eqiad.wmnet with reason: host reimage * 19:13 cdobbins@dns1004: START - running authdns-update * 19:12 cdobbins@cumin2002: conftool action : set/pooled=yes; selector: name=dns7002.*,service=authdns-update * 19:11 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1092.eqiad.wmnet with reason: host reimage * 19:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1123.eqiad.wmnet with reason: host reimage * 18:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1123.eqiad.wmnet with OS trixie * 18:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1109.eqiad.wmnet with OS trixie * 18:56 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1092.eqiad.wmnet with OS trixie * 18:52 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 18:49 swfrench@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 18:40 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 18:38 swfrench@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 18:11 swfrench-wmf: restarted eqsin, codfw confds - [[phab:T430909|T430909]] * 18:01 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T430909|T430909]] * 17:59 swfrench-wmf: restarted ulsfo confds, confirmed now connected to eqiad backends - [[phab:T430909|T430909]] * 17:52 sukhe: restart pybal on lvs2011 to switch from conf2004 to conf1008: [[phab:T430909|T430909]] * 17:51 sukhe: restart pybal on lvs2012 to switch from conf2004 to conf1008 [puppet re-enabled there]: [[phab:T430909|T430909]] * 17:46 sukhe: restart pybal on lvs2013 to switch from conf2004 to conf1008: [[phab:T430909|T430909]] * 17:44 swfrench-wmf: switched codfw, eqsin, ulsfo etcd client SRV records to eqiad - [[phab:T430909|T430909]] * 17:43 swfrench@dns1004: END - running authdns-update * 17:40 swfrench@dns1004: START - running authdns-update * 17:40 sukhe: restart pybal on lvs2014 to switch from conf2004 to conf1008: [[phab:T430909|T430909]] * 17:21 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1003.eqiad.wmnet * 17:15 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1003.eqiad.wmnet * 17:14 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1002.eqiad.wmnet * 17:06 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1002.eqiad.wmnet * 17:06 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-low-traffic-codfw' 'systemctl restart pybal.service' # lvs2013, [[phab:T416623|T416623]] * 17:04 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1001.eqiad.wmnet * 17:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1111.eqiad.wmnet with OS trixie * 17:00 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal.service' # lvs2014, [[phab:T416623|T416623]] * 16:58 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1001.eqiad.wmnet * 16:58 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS bookworm * 16:55 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-low-traffic-eqiad' 'systemctl restart pybal.service' # lvs1019, [[phab:T416623|T416623]] * 16:53 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-secondary-eqiad' 'systemctl restart pybal.service' # lvs1020, [[phab:T416623|T416623]] * 16:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1111.eqiad.wmnet with reason: host reimage * 16:40 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1111.eqiad.wmnet with reason: host reimage * 16:38 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.peering (exit_code=99) with action 'configure' for AS: 47794 * 16:35 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 47794 * 16:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1111.eqiad.wmnet with OS trixie * 16:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1006.eqiad.wmnet with OS trixie * 16:06 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 16:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1006.eqiad.wmnet with reason: host reimage * 15:58 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1121.eqiad.wmnet with OS trixie * 15:58 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 15:56 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1006.eqiad.wmnet with reason: host reimage * 15:54 mutante: jenkins down in planned maintenance window * 15:42 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1037.eqiad.wmnet * 15:42 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1037.eqiad.wmnet * 15:42 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1037.eqiad.wmnet * 15:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1006.eqiad.wmnet with OS trixie * 15:34 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1121.eqiad.wmnet with reason: host reimage * 15:33 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS bookworm * 15:33 cdobbins@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host dns7002.wikimedia.org with OS trixie * 15:30 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1121.eqiad.wmnet with reason: host reimage * 15:29 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host clouddumps1001.wikimedia.org * 15:20 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1001.wikimedia.org * 15:18 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1121 * 15:18 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1121 * 15:18 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host clouddumps1002.wikimedia.org * 15:17 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1121 * 15:17 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:17 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply * 15:16 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply * 15:16 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:16 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1121 - atsuko@cumin1003" * 15:16 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1121 - atsuko@cumin1003" * 15:14 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1037.eqiad.wmnet with OS trixie * 15:11 atsuko@cumin1003: START - Cookbook sre.dns.netbox * 15:09 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org * 15:09 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1121 * 15:09 andrew@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host clouddumps1002.wikimedia.org * 15:09 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org * 15:09 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1121.eqiad.wmnet with OS trixie * 15:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1007.eqiad.wmnet with OS trixie * 15:08 andrew@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host clouddumps1002.wikimedia.org * 15:08 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org * 15:05 brennen@deploy1003: Finished deploy [phabricator/deployment@7e02037]: deploy phab1004 for [[phab:T431440|T431440]] (duration: 00m 47s) * 15:04 brennen@deploy1003: Started deploy [phabricator/deployment@7e02037]: deploy phab1004 for [[phab:T431440|T431440]] * 15:03 brennen@deploy1003: Finished deploy [phabricator/deployment@7e02037]: deploy phab2003 for [[phab:T431440|T431440]] (duration: 00m 51s) * 15:03 brennen@deploy1003: Started deploy [phabricator/deployment@7e02037]: deploy phab2003 for [[phab:T431440|T431440]] * 15:00 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71] (thin): Regular analytics weekly train THIN [analytics/refinery@7d8dc71f] (duration: 02m 10s) * 14:58 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71] (thin): Regular analytics weekly train THIN [analytics/refinery@7d8dc71f] * 14:58 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71]: Regular analytics weekly train [analytics/refinery@7d8dc71f] (duration: 04m 14s) * 14:54 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1037.eqiad.wmnet with reason: host reimage * 14:53 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71]: Regular analytics weekly train [analytics/refinery@7d8dc71f] * 14:53 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@7d8dc71f] (duration: 02m 00s) * 14:51 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@7d8dc71f] * 14:51 arnaudb@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on phab2003.codfw.wmnet,phab[1004-1006].eqiad.wmnet with reason: maintenance * 14:51 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1037.eqiad.wmnet with reason: host reimage * 14:50 JavierMonton: Deploying Refinery at {{Gerrit|7d8dc71f}} for change {{Gerrit|1308087}} / [[phab:T431318|T431318]] - update filerevision table sqoop and table * 14:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1007.eqiad.wmnet with reason: host reimage * 14:42 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1007.eqiad.wmnet with reason: host reimage * 14:40 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1083.eqiad.wmnet with OS trixie * 14:37 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:36 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:35 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-master@eqiad * 14:35 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 14:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 14:34 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1037 * 14:34 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1037 * 14:34 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:cleanMentorList.php --wiki=frwiki # [[phab:T427386|T427386]] * 14:34 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 14:34 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308112{{!}}Revert^2 "[Growth] frwiki: Deploy automated mentor list cleaner" (T427386)]] (duration: 06m 47s) * 14:34 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 14:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:33 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:32 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1037 * 14:31 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:31 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:29 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:29 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-master@eqiad * 14:29 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:29 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:29 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:28 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:27 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:27 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1308112{{!}}Revert^2 "[Growth] frwiki: Deploy automated mentor list cleaner" (T427386)]] * 14:26 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1007.eqiad.wmnet with OS trixie * 14:26 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:26 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 14:26 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:25 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:25 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:25 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:cleanMentorList.php --wiki=frwiki # [[phab:T427386|T427386]] * 14:24 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:24 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1037 - blake@cumin1003" * 14:24 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1037 - blake@cumin1003" * 14:20 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1083.eqiad.wmnet with reason: host reimage * 14:19 blake@cumin1003: START - Cookbook sre.dns.netbox * 14:19 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-master@codfw * 14:19 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 14:19 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1037 * 14:18 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1037.eqiad.wmnet with OS trixie * 14:18 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1037.eqiad.wmnet * 14:18 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 14:18 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1037.eqiad.wmnet * 14:18 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1037.eqiad.wmnet * 14:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2007.codfw.wmnet with OS trixie * 14:16 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1083.eqiad.wmnet with reason: host reimage * 14:15 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1036.eqiad.wmnet * 14:15 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1036.eqiad.wmnet * 14:14 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1036.eqiad.wmnet * 14:12 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-master@codfw * 14:11 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1120.eqiad.wmnet with OS trixie * 14:05 moritzm: installing distro-info-data updates from trixie/bookworm point releases * 14:04 fabfur: disable puppet on A:cp-text to selectively apply https://gerrit.wikimedia.org/r/c/operations/puppet/+/1308040 * 14:03 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] (duration: 27m 48s) * 14:00 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 14:00 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1083.eqiad.wmnet with OS trixie * 13:58 urbanecm@deploy1003: urbanecm: Continuing with deployment * 13:58 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:58 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1004.eqiad.wmnet with OS bookworm * 13:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2007.codfw.wmnet with reason: host reimage * 13:57 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 13:53 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1120.eqiad.wmnet with reason: host reimage * 13:50 moritzm: installing Linux 5.10.259 on Bullseye hosts * 13:47 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply * 13:47 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply * 13:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2007.codfw.wmnet with reason: host reimage * 13:46 cgoubert@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/aux-k8s-services/redioscope: apply * 13:46 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1120.eqiad.wmnet with reason: host reimage * 13:46 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:46 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:45 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:44 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:44 cgoubert@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/aux-k8s-services/redioscope: apply * 13:40 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 13:39 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:38 moritzm: installing e2fsprogs updates from Trixie point release * 13:35 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] * 13:33 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1120.eqiad.wmnet with OS trixie * 13:33 topranks: reset cr3-eqsin configuration so traffic uses it again after upgrade * 13:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1088.eqiad.wmnet with OS trixie * 13:32 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie * 13:32 cdobbins@cumin1003: conftool action : set/pooled=no; selector: name=dns7002.* * 13:29 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2007.codfw.wmnet with OS trixie * 13:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1004.eqiad.wmnet with reason: host reimage * 13:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2006.codfw.wmnet with OS trixie * 13:18 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1036.eqiad.wmnet with OS trixie * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aux-k8s-etcd1004.eqiad.wmnet with reason: host reimage * 13:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 13:16 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 13:15 jayme: Istio is being upgraded from 1.24.2 to 1.29.4 on wikikube staging eqiad and codfw - [[phab:T427401|T427401]] * 13:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1087.eqiad.wmnet with OS trixie * 13:14 topranks: reboot cr3-eqsin to install new JunOS and set PIC 0/0/0 to 100G * 13:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1088.eqiad.wmnet with reason: host reimage * 13:13 jmm@dns1004: END - running authdns-update * 13:12 jmm@dns1004: START - running authdns-update * 13:09 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1088.eqiad.wmnet with reason: host reimage * 13:07 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1082.eqiad.wmnet with OS trixie * 13:07 jmm@dns1004: END - running authdns-update * 13:06 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1004.eqiad.wmnet with OS bookworm * 13:05 jmm@dns1004: START - running authdns-update * 13:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2006.codfw.wmnet with reason: host reimage * 12:58 topranks: load updated JunOS on cr3-eqsin [[phab:T429386|T429386]] * 12:58 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1036.eqiad.wmnet with reason: host reimage * 12:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2001.codfw.wmnet * 12:57 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2006.codfw.wmnet with reason: host reimage * 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr1-codfw,cr[2-3]-eqsin,cr3-eqsin IPv6,cr3-eqsin.mgmt with reason: upgrade JunOS cr3-eqsin * 12:56 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lvs[5004-5006].eqsin.wmnet with reason: upgrade JunOS cr3-eqsin * 12:55 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 12:55 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 12:53 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1087.eqiad.wmnet with reason: host reimage * 12:53 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 12:52 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1088.eqiad.wmnet with OS trixie * 12:52 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1002.eqiad.wmnet * 12:52 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:51 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2001.codfw.wmnet * 12:49 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1036.eqiad.wmnet with reason: host reimage * 12:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 12:48 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1087.eqiad.wmnet with reason: host reimage * 12:44 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1082.eqiad.wmnet with reason: host reimage * 12:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1002.eqiad.wmnet * 12:42 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 12:42 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 12:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 12:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 12:39 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:39 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: move dumps-nfs IP to the shared one - filippo@cumin1003" * 12:39 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: move dumps-nfs IP to the shared one - filippo@cumin1003" * 12:39 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2006.codfw.wmnet with OS trixie * 12:38 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1082.eqiad.wmnet with reason: host reimage * 12:36 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:33 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:32 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1036 * 12:32 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1036 * 12:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2005.codfw.wmnet with OS trixie * 12:32 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1087.eqiad.wmnet with OS trixie * 12:30 jmm@dns1004: END - running authdns-update * 12:29 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1036 * 12:29 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1036.eqiad.wmnet 21.32.64.10.in-addr.arpa 1.2.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:29 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1036.eqiad.wmnet 21.32.64.10.in-addr.arpa 1.2.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:29 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:29 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1036 - blake@cumin1003" * 12:29 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1036 - blake@cumin1003" * 12:28 jmm@dns1004: START - running authdns-update * 12:26 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:26 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:23 blake@cumin1003: START - Cookbook sre.dns.netbox * 12:23 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1036 * 12:23 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1036.eqiad.wmnet with OS trixie * 12:22 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1036.eqiad.wmnet * 12:22 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1082.eqiad.wmnet with OS trixie * 12:22 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1036.eqiad.wmnet * 12:22 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1036.eqiad.wmnet * 12:21 marostegui: Restart mariadb@s7 on db1155 to pick up new filters - [[phab:T431124|T431124]] * 12:21 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 21 hosts with reason: restarting for replication filter * 12:20 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:19 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:14 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2005.codfw.wmnet with reason: host reimage * 12:14 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:08 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:07 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:07 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2005.codfw.wmnet with reason: host reimage * 12:06 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:06 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:06 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:05 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:05 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:04 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-master-eqiad * 12:04 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl1002.eqiad.wmnet * 12:04 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl1002.eqiad.wmnet * 12:04 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:04 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:03 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:03 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:03 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:03 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 11:59 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl1002.eqiad.wmnet * 11:59 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl1002.eqiad.wmnet * 11:59 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl1001.eqiad.wmnet * 11:59 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl1001.eqiad.wmnet * 11:56 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl1001.eqiad.wmnet * 11:56 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl1001.eqiad.wmnet * 11:56 klausman@cumin2002: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-master-eqiad * 11:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2005.codfw.wmnet with OS trixie * 11:49 blake@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on wikikube-worker1160.eqiad.wmnet with reason: Verifying matchers for silence * 11:42 topranks: cr3-eqsin, begin traffic drain to reset PIC and upgrade JunOS [[phab:T429386|T429386]] * 11:41 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs[5004-5006].eqsin.wmnet with reason: upgrade JunOS cr3-eqsin * 11:39 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr1-codfw,cr[2-3]-eqsin,cr3-eqsin IPv6,cr3-eqsin.mgmt with reason: upgrade JunOS cr3-eqsin * 11:36 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=thanos-fe2004.codfw.wmnet * 11:35 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1086.eqiad.wmnet with OS trixie * 11:35 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=thanos-fe2004.codfw.wmnet * 11:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1085.eqiad.wmnet with OS trixie * 11:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1086.eqiad.wmnet with reason: host reimage * 11:10 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1085.eqiad.wmnet with reason: host reimage * 11:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2004.codfw.wmnet with OS trixie * 11:03 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1086.eqiad.wmnet with reason: host reimage * 11:02 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1085.eqiad.wmnet with reason: host reimage * 10:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2004.codfw.wmnet with reason: host reimage * 10:48 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:46 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1086.eqiad.wmnet with OS trixie * 10:46 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1085.eqiad.wmnet with OS trixie * 10:44 cgoubert@deploy1003: Finished deploy [restbase/deploy@2fc37d4]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] (duration: 16m 44s) * 10:43 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2004.codfw.wmnet with reason: host reimage * 10:35 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:27 cgoubert@deploy1003: Started deploy [restbase/deploy@2fc37d4]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] * 10:27 cgoubert@deploy1003: Finished deploy [restbase/deploy@8a25036]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] (duration: 00m 45s) * 10:26 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1117.eqiad.wmnet with OS trixie * 10:26 cgoubert@deploy1003: Started deploy [restbase/deploy@8a25036]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] * 10:26 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host thanos-fe2004 * 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host thanos-fe2004 * 10:22 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1116.eqiad.wmnet with OS trixie * 10:21 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host thanos-fe2004 * 10:21 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) thanos-fe2004.codfw.wmnet 157.32.192.10.in-addr.arpa 7.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:20 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache thanos-fe2004.codfw.wmnet 157.32.192.10.in-addr.arpa 7.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:20 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:20 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host thanos-fe2004 - mvernon@cumin2003" * 10:20 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host thanos-fe2004 - mvernon@cumin2003" * 10:15 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2252: Repooling after reboot * 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:15 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 10:15 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2252: Repooling after reboot * 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1153.eqiad.wmnet * 10:14 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1153.eqiad.wmnet * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 10:14 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 10:12 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 10:12 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host thanos-fe2004 * 10:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2004.codfw.wmnet with OS trixie * 10:07 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1117.eqiad.wmnet with reason: host reimage * 10:03 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1116.eqiad.wmnet with reason: host reimage * 09:58 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:58 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1117.eqiad.wmnet with reason: host reimage * 09:57 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1116.eqiad.wmnet with reason: host reimage * 09:49 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 41 days, 15:00:00 on db2252.codfw.wmnet with reason: Security updates * 09:45 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1117.eqiad.wmnet with OS trixie * 09:45 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1116.eqiad.wmnet with OS trixie * 09:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1153: Security updates * 09:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:28 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:28 root@cumin1003: START - Cookbook sre.mysql.depool depool db1153: Security updates * 09:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1016: Security updates * 09:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:21 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:21 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1016: Security updates * 09:14 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:14 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 08:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1016: Security updates * 08:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:56 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:56 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1016: Security updates * 08:50 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:50 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:45 filippo@dns1006: END - running authdns-update * 08:43 filippo@dns1006: START - running authdns-update * 08:42 godog: switch dumps-nfs address to be shared with rsync/http - [[phab:T411248|T411248]] * 08:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1016: Security updates * 08:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:40 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:40 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1016: Security updates * 08:29 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host cirrussearch1111.eqiad.wmnet * 08:29 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:27 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:27 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:25 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1015: Security updates * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:09 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:09 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1015: Security updates * 07:42 Msz2001: Deployed private patch for Suggested Ivestigations * 07:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1015: Security updates * 07:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:41 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:41 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1015: Security updates * 07:40 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 07:11 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fingerprint warnings - oblivian@cumin1003" * 07:11 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fingerprint warnings - oblivian@cumin1003 * 07:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1024: Security updates * 07:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:11 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:11 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1024: Security updates * 07:10 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fingerprint warnings - oblivian@cumin1003 * 07:10 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fingerprint warnings - oblivian@cumin1003" * 06:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host cirrussearch1111.eqiad.wmnet * 06:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 06:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1024: Security updates * 06:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 06:48 root@cumin1003: START - Cookbook sre.mysql.parsercache * 06:48 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1024: Security updates * 06:42 moritzm: install nginx security updates * 06:31 root@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool pc1024: Security updates * 06:21 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1024: Security updates * 06:19 moritzm: installing php8.2 security updates * 06:15 moritzm: installing php8.4 security updates * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.7 (duration: 02m 38s) * 03:40 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] (duration: 37m 04s) * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 51s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-06 == * 23:30 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] (duration: 09m 39s) * 23:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1078.eqiad.wmnet with OS trixie * 23:26 jdlrobson@deploy1003: jdlrobson, bwang: Continuing with deployment * 23:22 jdlrobson@deploy1003: jdlrobson, bwang: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug) * 23:21 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] * 23:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1078.eqiad.wmnet with reason: host reimage * 23:06 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1078.eqiad.wmnet with reason: host reimage * 22:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1078.eqiad.wmnet with OS trixie * 22:29 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on cirrussearch1114.eqiad.wmnet with reason: reimage on hold until restore completes * 22:22 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on cirrussearch[1079,1115].eqiad.wmnet with reason: reimage on hold until restore completes * 21:18 maryum: Deployed security fix for [[phab:T428006|T428006]] * 20:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1079.eqiad.wmnet with OS trixie * 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1077.eqiad.wmnet with OS trixie * 20:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1115.eqiad.wmnet with OS trixie * 20:25 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1079.eqiad.wmnet with reason: host reimage * 20:21 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1079.eqiad.wmnet with reason: host reimage * 20:15 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] (duration: 08m 14s) * 20:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1077.eqiad.wmnet with reason: host reimage * 20:10 krinkle@deploy1003: krinkle, pushpaktiwari: Continuing with deployment * 20:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1115.eqiad.wmnet with reason: host reimage * 20:08 krinkle@deploy1003: krinkle, pushpaktiwari: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1077.eqiad.wmnet with reason: host reimage * 20:06 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] * 20:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1079.eqiad.wmnet with OS trixie * 20:04 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1115.eqiad.wmnet with reason: host reimage * 19:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1077.eqiad.wmnet with OS trixie * 19:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1115.eqiad.wmnet with OS trixie * 19:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 19:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 18:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1114.eqiad.wmnet with OS trixie * 18:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1114.eqiad.wmnet with reason: host reimage * 18:35 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1114.eqiad.wmnet with reason: host reimage * 18:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1112.eqiad.wmnet with OS trixie * 18:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1114.eqiad.wmnet with OS trixie * 18:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1072.eqiad.wmnet with OS trixie * 18:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1112.eqiad.wmnet with reason: host reimage * 18:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1112.eqiad.wmnet with reason: host reimage * 17:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1072.eqiad.wmnet with reason: host reimage * 17:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1112.eqiad.wmnet with OS trixie * 17:55 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1072.eqiad.wmnet with reason: host reimage * 17:39 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1072.eqiad.wmnet with OS trixie * 17:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1071.eqiad.wmnet with OS trixie * 17:18 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1070.eqiad.wmnet with OS trixie * 17:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1084.eqiad.wmnet with OS trixie * 16:54 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1071.eqiad.wmnet with reason: host reimage * 16:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1084.eqiad.wmnet with reason: host reimage * 16:51 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1070.eqiad.wmnet with reason: host reimage * 16:49 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1084.eqiad.wmnet with reason: host reimage * 16:38 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1071.eqiad.wmnet with OS trixie * 16:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1096.eqiad.wmnet with OS trixie * 16:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1070.eqiad.wmnet with OS trixie * 16:33 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1084.eqiad.wmnet with OS trixie * 16:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1089.eqiad.wmnet with OS trixie * 16:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1103.eqiad.wmnet with OS trixie * 16:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1096.eqiad.wmnet with reason: host reimage * 16:14 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1096.eqiad.wmnet with reason: host reimage * 16:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1089.eqiad.wmnet with reason: host reimage * 16:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1103.eqiad.wmnet with reason: host reimage * 16:02 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1003.eqiad.wmnet with OS bookworm * 16:01 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1089.eqiad.wmnet with reason: host reimage * 16:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1103.eqiad.wmnet with reason: host reimage * 15:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1096.eqiad.wmnet with OS trixie * 15:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1080.eqiad.wmnet with OS trixie * 15:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1089.eqiad.wmnet with OS trixie * 15:45 dancy@deploy1003: Installation of scap version "4.272.0" completed for 158 hosts * 15:43 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1103.eqiad.wmnet with OS trixie * 15:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1113.eqiad.wmnet with OS trixie * 15:41 dancy@deploy1003: Installing scap version "4.272.0" for 158 host(s) * 15:40 klausman@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 15:39 klausman@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 15:38 klausman@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 15:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1069.eqiad.wmnet with OS trixie * 15:37 klausman@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 15:36 klausman@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 15:34 klausman@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 15:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1080.eqiad.wmnet with reason: host reimage * 15:27 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1080.eqiad.wmnet with reason: host reimage * 15:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1113.eqiad.wmnet with reason: host reimage * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1069.eqiad.wmnet with reason: host reimage * 15:18 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1113.eqiad.wmnet with reason: host reimage * 15:16 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1069.eqiad.wmnet with reason: host reimage * 15:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:11 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1080.eqiad.wmnet with OS trixie * 15:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1113.eqiad.wmnet with OS trixie * 15:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1003.eqiad.wmnet with reason: host reimage * 14:47 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1003.eqiad.wmnet with OS bookworm * 14:33 elukey: rolled out spicerack on all cumin nodes - [[phab:T429699|T429699]] * 14:32 elukey: upgrade all bookworm hosts to pywmflib 3.1 - [[phab:T430552|T430552]] * 14:14 marostegui: Setup x4 eqiad topology [[phab:T404715|T404715]] * 14:13 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 14:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2230.codfw.wmnet * 14:07 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2230.codfw.wmnet * 13:59 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[2001-2002].codfw.wmnet * 13:51 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 13:45 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.major-upgrade (exit_code=97) * 13:45 cwilliams@cumin1003: dbmaint on s4@codfw [[phab:T429893|T429893]] * 13:45 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 13:42 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-master-codfw * 13:42 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl2002.codfw.wmnet * 13:42 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl2002.codfw.wmnet * 13:38 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl2002.codfw.wmnet * 13:38 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl2002.codfw.wmnet * 13:38 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl2001.codfw.wmnet * 13:38 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl2001.codfw.wmnet * 13:35 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl2001.codfw.wmnet * 13:35 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl2001.codfw.wmnet * 13:35 klausman@cumin2002: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-master-codfw * 12:30 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] (duration: 25m 11s) * 12:24 krinkle@deploy1003: krinkle: Continuing with deployment * 12:10 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2048.codfw.wmnet * 12:09 krinkle@deploy1003: krinkle: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:08 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2048.codfw.wmnet * 12:05 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] * 11:57 moritzm: installing curl security updates * 11:49 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:31 moritzm: installing nano security updates * 11:07 moritzm: failover Ganeti master in codfw to ganeti2032 [[phab:T430909|T430909]] * 11:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:04 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest1005.eqiad.wmnet with OS trixie * 11:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:50 jmm@dns1004: END - running authdns-update * 10:47 jmm@dns1004: START - running authdns-update * 10:47 jmm@dns1004: START - running authdns-update * 10:46 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:44 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest1005.eqiad.wmnet with reason: host reimage * 10:38 elukey@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest1005.eqiad.wmnet with reason: host reimage * 10:31 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:31 marostegui: Setup x4 codfw topology [[phab:T404715|T404715]] * 10:31 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 10:24 elukey: spicerack 13.0.0 deployed on cumin2002 * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 10:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 10:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 10:21 elukey@cumin2002: START - Cookbook sre.hosts.reimage for host sretest1005.eqiad.wmnet with OS trixie * 10:20 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:19 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:17 elukey: uploaded spicerack_13.0.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia * 09:54 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:52 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:20 elukey: upgrade all bullseye hosts to pywmflib 3.1 - [[phab:T430552|T430552]] * 09:10 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1015.eqiad.wmnet,service=s4 * 09:10 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1015.eqiad.wmnet,service=s6 * 09:07 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 08:58 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:56 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 08:56 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 08:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 08:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 08:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin2002.codfw.wmnet * 08:06 godog: remove cloudvirt1046, cloudvirt1062, cloudvirt1074, cloudvirt1075 from maintenance aggregate and put them in network-ovs - [[phab:T424802|T424802]] * 08:00 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin2002.codfw.wmnet * 07:58 hashar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] (duration: 32m 53s) * 07:57 fabfur: repooled cp4038 * 07:57 fabfur@cumin1003: conftool action : set/pooled=yes; selector: name=cp4038.* * 07:53 moritzm: installing pyjwt security updates * 07:47 moritzm: installing openjpeg2 security updates * 07:45 hashar@deploy1003: vadymts1, hashar: Continuing with deployment * 07:43 hashar@deploy1003: vadymts1, hashar: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:38 moritzm: installing python-urllib3 security updates * 07:37 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 07:30 fabfur: depooled cp4038 to investigate on possible maxmind failure * 07:30 fabfur@cumin1003: conftool action : set/pooled=no; selector: name=cp4038.* * 07:30 fabfur@cumin1003: conftool action : set/pooled=yes; selector: name=cp4038.* * 07:29 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 07:25 hashar@deploy1003: Started scap sync-world: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] * 06:13 moritzm: installing Linux 6.12.95 on trixie hosts * 05:20 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s6 * 05:20 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s4 * 05:19 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1015.eqiad.wmnet with reason: cloning * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 08s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-05 == * 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 01m 08s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-04 == * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 58s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-03 == * 17:08 topranks: revert protocol preference changes on cr3-ulsfo after upgrade * 16:53 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on cr2-eqord with reason: upgrade JunOS cr3-ulsfo * 16:53 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on cr4-ulsfo with reason: upgrade JunOS cr3-ulsfo * 16:48 topranks: reboot cr3-ulsfo to upgrade JunOS and reset linecard [[phab:T424839|T424839]] * 15:52 topranks: adjust outbound BGP policies on cr3-ulsfo to drain router of traffic [[phab:T424839|T424839]] * 15:45 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on lvs[4008-4010].ulsfo.wmnet with reason: upgrade JunOS cr3-ulsfo * 15:44 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on asw1-[22-23]-ulsfo,cr3-ulsfo,cr3-ulsfo IPv6,cr3-ulsfo.mgmt with reason: upgrade JunOS cr3-ulsfo * 15:36 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 15:35 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 15:35 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 14:40 cmooney@dns3003: END - running authdns-update * 14:26 cmooney@dns3003: START - running authdns-update * 14:26 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:26 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to ulsfo - cmooney@cumin1003" * 14:19 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to ulsfo - cmooney@cumin1003" * 14:16 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:38 sukhe@dns1004: END - running authdns-update * 13:35 sukhe@dns1004: START - running authdns-update * 13:26 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 13:26 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 13:26 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet * 13:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 13:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 13:16 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host sretest1005.eqiad.wmnet * 13:16 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 13:16 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 13:15 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 13:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:14 moritzm: imported samplicator 1.3.8rc1-1+deb13u1 to trixie-wikimedia/main [[phab:T337208|T337208]] * 13:13 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:07 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 13:07 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 13:02 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:02 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:58 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:57 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:57 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:53 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet * 12:50 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 12:47 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:41 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:40 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:39 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:32 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet * 12:26 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet * 12:23 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2005.wikimedia.org * 12:19 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2005.wikimedia.org * 12:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 12:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup[2004-2007].codfw.wmnet * 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[2004-2007].codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin2003" * 12:15 jynus@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[2004-2007].codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin2003" * 12:09 jynus@cumin2003: START - Cookbook sre.dns.netbox * 11:58 jynus@cumin2003: START - Cookbook sre.hosts.decommission for hosts backup[2004-2007].codfw.wmnet * 10:40 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup[1004-1007].eqiad.wmnet * 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[1004-1007].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 10:01 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[1004-1007].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 09:52 jynus@cumin1003: START - Cookbook sre.dns.netbox * 09:39 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:36 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup[1004-1007].eqiad.wmnet * 09:36 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:25 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 09:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 09:16 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 09:05 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:04 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:00 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:59 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:57 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:55 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:50 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 08:50 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 08:49 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 08:49 atsukoito: depooling cirrussearch in codfw because of regression after upgrade [[phab:T431091|T431091]] * 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts mirror1001.wikimedia.org * 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: mirror1001.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 08:29 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: mirror1001.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 08:18 jmm@cumin2003: START - Cookbook sre.dns.netbox * 08:11 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts mirror1001.wikimedia.org * 06:15 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 18s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-02 == * 22:55 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host contint1003.wikimedia.org with OS trixie * 22:29 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on contint1003.wikimedia.org with reason: host reimage * 22:23 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on contint1003.wikimedia.org with reason: host reimage * 22:05 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host contint1003.wikimedia.org with OS trixie * 22:03 mutante: contint1003 (zuul.wikimedia.org) - reimaging because of [[phab:T430510|T430510]]#12067628 [[phab:T418521|T418521]] * 22:03 dzahn@cumin2002: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on zuul.wikimedia.org with reason: reimage * 21:39 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 18s) * 21:39 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 21:20 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1003.eqiad.wmnet, repooling source-only afterwards * 21:19 sbassett: Deployed security fix for [[phab:T428829|T428829]] * 20:58 cmooney@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Release v0.11.2 update for new Aerleon - cmooney@cumin1003 * 20:55 cmooney@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Release v0.11.2 update for new Aerleon - cmooney@cumin1003 * 20:40 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] (duration: 12m 35s) * 20:36 arlolra@deploy1003: cscott, arlolra: Continuing with deployment * 20:35 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 20s) * 20:35 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 20:33 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host contint2003.wikimedia.org with OS trixie * 20:31 arlolra@deploy1003: cscott, arlolra: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Cha * 20:28 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] * 20:17 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] (duration: 08m 13s) * 20:14 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on contint2003.wikimedia.org with reason: host reimage * 20:13 sbassett@deploy1003: sbassett: Continuing with deployment * 20:11 sbassett@deploy1003: sbassett: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:09 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] * 20:08 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 20:08 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 20:08 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on contint2003.wikimedia.org with reason: host reimage * 20:05 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1003.eqiad.wmnet, repooling source-only afterwards * 19:49 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host contint2003.wikimedia.org with OS trixie * 19:48 mutante: contint2003 - reimaging because of [[phab:T430510|T430510]]#12067628 [[phab:T418521|T418521]] * 18:39 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 18:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 18:13 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2002.codfw.wmnet -> wcqs2003.codfw.wmnet, repooling source-only afterwards * 17:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1003.eqiad.wmnet with OS bookworm * 17:52 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1005.eqiad.wmnet * 17:52 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1005.eqiad.wmnet * 17:51 jasmine@cumin2002: conftool action : set/pooled=yes:weight=10; selector: name=wikikube-ctrl1005.eqiad.wmnet * 17:48 jasmine_: homer "cr*eqiad*" commit "Added new stacked control plane wikikube-ctrl1005" * 17:44 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply * 17:44 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply * 17:31 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] (duration: 09m 33s) * 17:26 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 17:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1003.eqiad.wmnet with reason: host reimage * 17:23 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:21 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] * 17:18 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1003.eqiad.wmnet with reason: host reimage * 17:16 rscout@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply * 17:16 rscout@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply * 17:16 rscout@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply * 17:15 rscout@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply * 17:12 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on wcqs[2002-2003].codfw.wmnet,wcqs1002.eqiad.wmnet with reason: reimaging hosts * 17:08 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 17:08 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 17:08 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 17:07 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 17:05 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 17:05 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 17:03 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "running to make sure all updates are synced - cmooney@cumin1003" * 17:03 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "running to make sure all updates are synced - cmooney@cumin1003" * 17:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs1003 * 17:00 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs1003 * 17:00 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 17:00 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1003.eqiad.wmnet with OS bookworm * 16:58 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Re-running - btullis@cumin1003" * 16:58 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Re-running - btullis@cumin1003" * 16:58 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2002.codfw.wmnet -> wcqs2003.codfw.wmnet, repooling source-only afterwards * 16:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-master1004.eqiad.wmnet with OS bookworm * 16:58 btullis@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 16:57 tappof: bump space for prometheus k8s-aux in eqiad * 16:55 cmooney@dns3003: END - running authdns-update * 16:55 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:55 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to eqsin - cmooney@cumin1003" * 16:55 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to eqsin - cmooney@cumin1003" * 16:53 cmooney@dns3003: START - running authdns-update * 16:52 ryankemper: [ml-serve-eqiad] Cleared out 1302 failed (Evicted) pods: `kubectl -n llm delete pods --field-selector=status.phase=Failed`, freeing calico-kube-controllers from OOM crashloop (evictions were caused by disk pressure) * 16:49 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 16:46 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:39 rzl@dns1004: END - running authdns-update * 16:37 rzl@dns1004: START - running authdns-update * 16:36 rzl@dns1004: START - running authdns-update * 16:35 rzl@deploy1003: Finished scap sync-world: [[phab:T416623|T416623]] (duration: 10m 19s) * 16:34 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 16:33 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-master1004.eqiad.wmnet with reason: host reimage * 16:30 rzl@deploy1003: rzl: Continuing with deployment * 16:28 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-master1004.eqiad.wmnet with reason: host reimage * 16:26 rzl@deploy1003: rzl: [[phab:T416623|T416623]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:25 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 16:25 rzl@deploy1003: Started scap sync-world: [[phab:T416623|T416623]] * 16:25 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 16:24 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 16:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: sync * 16:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: sync * 16:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-master1004.eqiad.wmnet with OS bookworm * 16:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-master1003.eqiad.wmnet with OS bookworm * 16:11 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 16:11 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 16:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Security updates * 16:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 16:08 root@cumin1003: START - Cookbook sre.mysql.parsercache * 16:08 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Security updates * 15:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-master1003.eqiad.wmnet with reason: host reimage * 15:54 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:54 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:54 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:54 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-master1003.eqiad.wmnet with reason: host reimage * 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Security updates * 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:45 root@cumin1003: START - Cookbook sre.mysql.parsercache * 15:45 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Security updates * 15:42 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-master1003.eqiad.wmnet with OS bookworm * 15:24 moritzm: installing busybox updates from bookworm point release * 15:20 moritzm: installing busybox updates from trixie point release * 15:15 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1021: Security updates * 15:15 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:15 root@cumin1003: START - Cookbook sre.mysql.parsercache * 15:15 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1021: Security updates * 15:13 moritzm: installing giflib security updates * 15:08 moritzm: installing Tomcat security updates * 14:57 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 14:56 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 14:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:53 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Unblock taavi - oblivian@cumin1003" * 14:53 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Unblock taavi - oblivian@cumin1003 * 14:53 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1021: Security updates * 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:53 root@cumin1003: START - Cookbook sre.mysql.parsercache * 14:53 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1021: Security updates * 14:53 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Unblock taavi - oblivian@cumin1003 * 14:52 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Unblock taavi - oblivian@cumin1003" * 14:46 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94711 and previous config saved to /var/cache/conftool/dbconfig/20260702-144644-fceratto.json * 14:36 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205', diff saved to https://phabricator.wikimedia.org/P94709 and previous config saved to /var/cache/conftool/dbconfig/20260702-143636-fceratto.json * 14:32 moritzm: installing libdbi-perl security updates * 14:26 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205', diff saved to https://phabricator.wikimedia.org/P94708 and previous config saved to /var/cache/conftool/dbconfig/20260702-142628-fceratto.json * 14:16 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94707 and previous config saved to /var/cache/conftool/dbconfig/20260702-141621-fceratto.json * 14:12 moritzm: installing rsync security updates * 14:11 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox) * 14:10 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94706 and previous config saved to /var/cache/conftool/dbconfig/20260702-140959-fceratto.json * 14:09 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2205.codfw.wmnet with reason: Maintenance * 14:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2205: Repooling after switchover * 14:07 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-test-master1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 14:06 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 14:06 Tran: Deployed patch for [[phab:T427287|T427287]] * 14:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:59 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2205: Repooling after switchover * 13:59 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2205: Repooling after switchover * 13:59 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:55 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2205: Repooling after switchover * 13:55 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2205 [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94704 and previous config saved to /var/cache/conftool/dbconfig/20260702-135505-fceratto.json * 13:54 moritzm: installing sed security updates * 13:53 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:52 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2209 to s3 primary [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94703 and previous config saved to /var/cache/conftool/dbconfig/20260702-135235-fceratto.json * 13:52 federico3: Starting s3 codfw failover from db2205 to db2209 - [[phab:T430912|T430912]] * 13:51 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:51 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 13:48 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:47 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2209 with weight 0 [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94702 and previous config saved to /var/cache/conftool/dbconfig/20260702-134719-fceratto.json * 13:47 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Primary switchover s3 [[phab:T430912|T430912]] * 13:44 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:44 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:44 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:40 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 13:38 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 13:37 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 13:36 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 13:36 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:34 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 13:30 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:29 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:29 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:27 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:26 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:25 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling restart_daemons on A:wikidough * 13:23 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 13:22 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 13:17 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 13:17 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns1004.wikimedia.org * 13:12 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:11 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart (exit_code=97) rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough * 13:11 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=97) rolling restart_daemons on A:wikidough * 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough * 13:09 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] (duration: 07m 20s) * 13:05 aude@deploy1003: jdrewniak, aude: Continuing with deployment * 13:04 aude@deploy1003: jdrewniak, aude: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:02 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] * 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts wdqs-categories1001.eqiad.wmnet * 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: wdqs-categories1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 12:10 jmm@dns1004: END - running authdns-update * 12:07 jmm@dns1004: START - running authdns-update * 11:51 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: wdqs-categories1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 11:44 btullis@cumin1003: START - Cookbook sre.dns.netbox * 11:42 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 11:42 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 11:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet * 11:39 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts wdqs-categories1001.eqiad.wmnet * 11:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet * 11:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet * 11:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet * 11:29 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 11:29 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 10:57 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2214: Repooling * 10:49 jmm@dns1004: END - running authdns-update * 10:47 jmm@dns1004: START - running authdns-update * 10:31 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94698 and previous config saved to /var/cache/conftool/dbconfig/20260702-103146-fceratto.json * 10:21 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213', diff saved to https://phabricator.wikimedia.org/P94696 and previous config saved to /var/cache/conftool/dbconfig/20260702-102137-fceratto.json * 10:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:19 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb1017.eqiad.wmnet * 10:18 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 10:18 fceratto@cumin1003: Removing es1033 from zarcillo [[phab:T408772|T408772]] * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts es1033.eqiad.wmnet * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: es1033.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:14 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: es1033.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:13 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb1017.eqiad.wmnet * 10:12 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2214.codfw.wmnet * 10:12 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2214.codfw.wmnet * 10:12 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2214: Repooling * 10:11 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213', diff saved to https://phabricator.wikimedia.org/P94693 and previous config saved to /var/cache/conftool/dbconfig/20260702-101130-fceratto.json * 10:10 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:10 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:03 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts es1033.eqiad.wmnet * 10:03 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 10:01 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94691 and previous config saved to /var/cache/conftool/dbconfig/20260702-100122-fceratto.json * 09:55 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94690 and previous config saved to /var/cache/conftool/dbconfig/20260702-095529-fceratto.json * 09:55 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2213.codfw.wmnet with reason: Maintenance * 09:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 09:53 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2213: Repooling after switchover * 09:51 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover * 09:44 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2213: Repooling after switchover * 09:39 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover * 09:39 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2213 [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94688 and previous config saved to /var/cache/conftool/dbconfig/20260702-093859-fceratto.json * 09:36 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2192 to s5 primary [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94687 and previous config saved to /var/cache/conftool/dbconfig/20260702-093650-fceratto.json * 09:36 federico3: Starting s5 codfw failover from db2213 to db2192 - [[phab:T430923|T430923]] * 09:30 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94686 and previous config saved to /var/cache/conftool/dbconfig/20260702-093004-fceratto.json * 09:24 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2192 with weight 0 [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94685 and previous config saved to /var/cache/conftool/dbconfig/20260702-092455-fceratto.json * 09:24 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 23 hosts with reason: Primary switchover s5 [[phab:T430923|T430923]] * 09:19 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220', diff saved to https://phabricator.wikimedia.org/P94684 and previous config saved to /var/cache/conftool/dbconfig/20260702-091957-fceratto.json * 09:16 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] (duration: 06m 57s) * 09:13 moritzm: installing libgcrypt20 security updates * 09:12 kharlan@deploy1003: kharlan: Continuing with deployment * 09:11 kharlan@deploy1003: kharlan: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:09 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220', diff saved to https://phabricator.wikimedia.org/P94683 and previous config saved to /var/cache/conftool/dbconfig/20260702-090950-fceratto.json * 09:09 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] * 09:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 09:01 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] (duration: 07m 07s) * 08:59 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94682 and previous config saved to /var/cache/conftool/dbconfig/20260702-085942-fceratto.json * 08:57 kharlan@deploy1003: kharlan: Continuing with deployment * 08:56 kharlan@deploy1003: kharlan: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:54 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] * 08:52 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:52 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:52 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94681 and previous config saved to /var/cache/conftool/dbconfig/20260702-085237-fceratto.json * 08:52 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2220.codfw.wmnet with reason: Maintenance * 08:43 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:40 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 08:25 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] (duration: 11m 44s) * 08:21 cscott@deploy1003: cscott: Continuing with deployment * 08:16 cscott@deploy1003: cscott: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:14 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] * 08:08 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 08:08 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1244: Migration of db1244.eqiad.wmnet completed * 08:02 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:02 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:01 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] (duration: 18m 58s) * 08:01 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:59 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 07:59 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:59 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:59 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2006.wikimedia.org * 07:58 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:57 cscott@deploy1003: cscott: Continuing with deployment * 07:56 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:56 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:56 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:55 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:55 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:55 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:54 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2006.wikimedia.org * 07:54 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:54 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 07:44 cscott@deploy1003: cscott: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:44 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2005.wikimedia.org * 07:44 moritzm: installing node-lodash security updates * 07:42 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] * 07:39 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2005.wikimedia.org * 07:30 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] (duration: 07m 28s) * 07:26 cscott@deploy1003: ssastry, cscott: Continuing with deployment * 07:25 cscott@deploy1003: ssastry, cscott: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:23 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1244: Migration of db1244.eqiad.wmnet completed * 07:22 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] * 07:16 wmde-fisch@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] (duration: 06m 55s) * 07:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1244.eqiad.wmnet with OS trixie * 07:11 wmde-fisch@deploy1003: wmde-fisch: Continuing with deployment * 07:11 wmde-fisch@deploy1003: wmde-fisch: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:09 wmde-fisch@deploy1003: Started scap sync-world: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] * 06:54 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1244.eqiad.wmnet with reason: host reimage * 06:50 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1244.eqiad.wmnet with reason: host reimage * 06:38 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1250.eqiad.wmnet with OS trixie * 06:34 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db1244.eqiad.wmnet with OS trixie * 06:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1244: Upgrading db1244.eqiad.wmnet * 06:25 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1244: Upgrading db1244.eqiad.wmnet * 06:25 cwilliams@cumin1003: dbmaint on s4@eqiad [[phab:T429893|T429893]] * 06:25 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 06:15 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1250.eqiad.wmnet with reason: host reimage * 06:14 cwilliams@dns1006: END - running authdns-update * 06:12 cwilliams@dns1006: START - running authdns-update * 06:11 cwilliams@dns1006: END - running authdns-update * 06:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db1244 [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94676 and previous config saved to /var/cache/conftool/dbconfig/20260702-061059-cwilliams.json * 06:09 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1250.eqiad.wmnet with reason: host reimage * 06:09 cwilliams@dns1006: START - running authdns-update * 06:08 aokoth@cumin1003: END (PASS) - Cookbook sre.vrts.upgrade (exit_code=0) on VRTS host vrts1003.eqiad.wmnet * 06:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db1160 to s4 primary and set section read-write [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94675 and previous config saved to /var/cache/conftool/dbconfig/20260702-060746-cwilliams.json * 06:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Set s4 eqiad as read-only for maintenance - [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94674 and previous config saved to /var/cache/conftool/dbconfig/20260702-060704-cwilliams.json * 06:06 cezmunsta: Starting s4 eqiad failover from db1244 to db1160 - [[phab:T430817|T430817]] * 06:04 aokoth@cumin1003: START - Cookbook sre.vrts.upgrade on VRTS host vrts1003.eqiad.wmnet * 05:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db1160 with weight 0 [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94673 and previous config saved to /var/cache/conftool/dbconfig/20260702-055927-cwilliams.json * 05:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 40 hosts with reason: Primary switchover s4 [[phab:T430817|T430817]] * 05:55 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1250.eqiad.wmnet with OS trixie * 05:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on db1250.eqiad.wmnet with reason: m3 master switchover [[phab:T430158|T430158]] * 05:39 marostegui: Failover m3 (phabricator) from db1250 to db1228 - [[phab:T430158|T430158]] * 05:32 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2234].codfw.wmnet,db[1217,1228,1250].eqiad.wmnet with reason: m3 master switchover [[phab:T430158|T430158]] * 04:45 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] (duration: 09m 08s) * 04:41 tstarling@deploy1003: tstarling, reedy: Continuing with deployment * 04:38 tstarling@deploy1003: tstarling, reedy: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 04:36 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 59s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:16 ryankemper: [[phab:T429844|T429844]] [opensearch] completed `cirrussearch2111` reimage; all codfw search clusters are green, all nodes now report `OpenSearch 2.19.5`, and the temporary chi voting exclusion has been removed * 00:57 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2111.codfw.wmnet with OS trixie * 00:29 ryankemper: [[phab:T429844|T429844]] [opensearch] depooled codfw search-omega/search-psi discovery records to match existing codfw search depool during OpenSearch 2.19 migration * 00:29 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2111.codfw.wmnet with reason: host reimage * 00:29 ryankemper@cumin2002: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 00:29 ryankemper@cumin2002: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 00:22 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2111.codfw.wmnet with reason: host reimage * 00:01 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2111.codfw.wmnet with OS trixie * 00:00 ryankemper: [[phab:T429844|T429844]] [opensearch] chi cluster recovered after stopping `opensearch_1@production-search-codfw` on `cirrussearch2111` == 2026-07-01 == * 23:59 ryankemper: [[phab:T429844|T429844]] [opensearch] stopped `opensearch_1@production-search-codfw` on `cirrussearch2111` after chi cluster-manager election churn following `voting_config_exclusions` POST; hoping this triggers a re-election * 23:52 cscott@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 23:51 cscott@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 23:51 cscott@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 23:50 cscott@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2003.codfw.wmnet with OS bookworm * 22:29 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 22:13 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 22:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2084.codfw.wmnet with OS trixie * 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2003.codfw.wmnet with reason: host reimage * 22:03 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 22:01 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2003.codfw.wmnet with reason: host reimage * 21:50 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 21:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2084.codfw.wmnet with reason: host reimage * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2003 * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2003 * 21:42 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2003 * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2003.codfw.wmnet 45.48.192.10.in-addr.arpa 5.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:42 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2003.codfw.wmnet 45.48.192.10.in-addr.arpa 5.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2003 - bking@cumin2003" * 21:42 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2003 - bking@cumin2003" * 21:36 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2084.codfw.wmnet with reason: host reimage * 21:35 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:34 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2003 * 21:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2003.codfw.wmnet with OS bookworm * 21:19 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2084.codfw.wmnet with OS trixie * 21:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2081.codfw.wmnet with OS trixie * 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2108.codfw.wmnet with OS trixie * 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2081.codfw.wmnet with reason: host reimage * 20:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2081.codfw.wmnet with reason: host reimage * 20:28 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2081.codfw.wmnet with OS trixie * 20:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2108.codfw.wmnet with reason: host reimage * 20:19 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2108.codfw.wmnet with reason: host reimage * 19:59 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2108.codfw.wmnet with OS trixie * 19:46 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2093.codfw.wmnet with OS trixie * 19:44 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 19:44 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jasmine@cumin2002" * 19:43 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jasmine@cumin2002" * 19:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2080.codfw.wmnet with OS trixie * 19:28 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 19:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2093.codfw.wmnet with reason: host reimage * 19:18 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 19:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2093.codfw.wmnet with reason: host reimage * 19:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2080.codfw.wmnet with reason: host reimage * 19:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2080.codfw.wmnet with reason: host reimage * 18:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2093.codfw.wmnet with OS trixie * 18:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2080.codfw.wmnet with OS trixie * 18:27 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 18:18 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] (duration: 09m 15s) * 18:13 jgiannelos@deploy1003: jgiannelos, neriah: Continuing with deployment * 18:11 jgiannelos@deploy1003: jgiannelos, neriah: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:09 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] * 17:40 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 16:58 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 30 hosts * 16:57 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for 30 hosts * 16:52 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2202.codfw.wmnet * 16:52 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2202.codfw.wmnet * 16:51 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt * 16:51 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt * 16:51 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lvs2012.codfw.wmnet * 16:51 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for lvs2012.codfw.wmnet * 16:49 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2076.codfw.wmnet with OS trixie * 16:49 brett: Start pybal on lvs2012 - [[phab:T429861|T429861]] * 16:49 pt1979@cumin1003: END (ERROR) - Cookbook sre.hosts.remove-downtime (exit_code=97) for 59 hosts * 16:48 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for 59 hosts * 16:42 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2061.codfw.wmnet with OS trixie * 16:30 dancy@deploy1003: Installation of scap version "4.271.0" completed for 2 hosts * 16:28 dancy@deploy1003: Installing scap version "4.271.0" for 2 host(s) * 16:23 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2076.codfw.wmnet with reason: host reimage * 16:19 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2061.codfw.wmnet with reason: host reimage * 16:18 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2076.codfw.wmnet with reason: host reimage * 16:18 jasmine@dns1004: END - running authdns-update * 16:16 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host restbase2039.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 16:16 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host restbase2039.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 16:16 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2061.codfw.wmnet with reason: host reimage * 16:15 jasmine@dns1004: START - running authdns-update * 16:14 jasmine@dns1004: END - running authdns-update * 16:12 jasmine@dns1004: START - running authdns-update * 16:07 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2202.codfw.wmnet with reason: maintenance * 16:06 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt with reason: Junos upograde * 16:00 papaul: ongoing maintenance on lsw1-b2-codfw * 16:00 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2076.codfw.wmnet with OS trixie * 15:59 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt * 15:59 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt * 15:57 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2061.codfw.wmnet with OS trixie * 15:55 pt1979@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2042,2046].codfw.wmnet * 15:55 pt1979@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2042,2046].codfw.wmnet * 15:51 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 15:51 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2220: Repooling after switchover * 15:50 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 15:50 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 15:48 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2092.codfw.wmnet with OS trixie * 15:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 15:40 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 15:38 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 15:37 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 15:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply * 15:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply * 15:32 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 15:32 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 15:30 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 15:29 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 15:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 15:25 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:22 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs2012.codfw.wmnet with reason: Rack B2 maintenance - [[phab:T429861|T429861]] * 15:21 brett: Stopping pybal on lvs2012 in preparation for codfw rack b2 maintenance - [[phab:T429861|T429861]] * 15:20 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2092.codfw.wmnet with reason: host reimage * 15:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:12 _joe_: restarted manually alertmanager-irc-relay * 15:12 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2092.codfw.wmnet with reason: host reimage * 15:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:12 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt with reason: Junos upograde * 15:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover * 15:07 pt1979@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2042,2046].codfw.wmnet * 15:06 pt1979@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2042,2046].codfw.wmnet * 15:02 papaul: ongoing maintenance on lsw1-a8-codfw * 14:31 topranks: POWERING DOWN CR1-EQIAD for line card installation [[phab:T426343|T426343]] * 14:31 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] (duration: 08m 57s) * 14:29 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:26 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 14:24 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:22 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] * 14:22 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover * 14:16 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:15 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover * 14:14 topranks: re-enable routing-engine graceful-failover on cr1-eqiad [[phab:T417873|T417873]] * 14:13 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:13 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2220: Repooling after switchover * 14:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:12 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:12 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:11 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:08 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] (duration: 10m 01s) * 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:07 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2220 [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94664 and previous config saved to /var/cache/conftool/dbconfig/20260701-140729-fceratto.json * 14:06 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:06 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:06 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:05 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2159 to s7 primary [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94663 and previous config saved to /var/cache/conftool/dbconfig/20260701-140503-fceratto.json * 14:04 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:04 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 14:04 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 14:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:04 dreamyjazz@deploy1003: anzx, dreamyjazz: Continuing with deployment * 14:04 federico3: Starting s7 codfw failover from db2220 to db2159 - [[phab:T430826|T430826]] * 14:03 jmm@dns1004: END - running authdns-update * 14:03 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:03 topranks: flipping cr1-eqiad active routing-enginer back to RE0 [[phab:T417873|T417873]] * 14:03 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudsw1-c8-eqiad,cloudsw1-d5-eqiad with reason: router upgrades eqiad * 14:01 jmm@dns1004: START - running authdns-update * 14:00 dreamyjazz@deploy1003: anzx, dreamyjazz: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:59 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2159 with weight 0 [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94662 and previous config saved to /var/cache/conftool/dbconfig/20260701-135906-fceratto.json * 13:58 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] * 13:57 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s7 [[phab:T430826|T430826]] * 13:56 topranks: reboot routing-enginer RE0 on cr1-eqiad [[phab:T417873|T417873]] * 13:48 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1006.wikimedia.org * 13:44 atsuko@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cirrussearch2092.codfw.wmnet with OS trixie * 13:43 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1006.wikimedia.org * 13:41 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2092.codfw.wmnet with OS trixie * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1005.wikimedia.org * 13:37 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1005.wikimedia.org * 13:37 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on pfw1-eqiad with reason: router upgrades eqiad * 13:35 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on lvs[1017-1020].eqiad.wmnet with reason: router upgrades eqiad * 13:34 caro@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] (duration: 07m 59s) * 13:30 caro@deploy1003: caro: Continuing with deployment * 13:28 caro@deploy1003: caro: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:27 topranks: route-engine failover cr1-eqiad * 13:26 caro@deploy1003: Started scap sync-world: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] * 13:15 topranks: rebooting routing-engine 1 on cr1-eqiad [[phab:T417873|T417873]] * 13:13 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] (duration: 08m 29s) * 13:13 moritzm: installing qemu security updates * 13:11 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 13:11 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 13:09 jgiannelos@deploy1003: jgiannelos: Continuing with deployment * 13:08 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 13:07 jgiannelos@deploy1003: jgiannelos: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:06 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 13:06 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2214.codfw.wmnet with reason: Maintenance * 13:05 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2214: Repooling after switchover * 13:05 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] * 13:04 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2214: Repooling after switchover * 13:04 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2214 [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94660 and previous config saved to /var/cache/conftool/dbconfig/20260701-130413-fceratto.json * 13:01 moritzm: installing python3.13 security updates * 13:00 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2229 to s6 primary [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94659 and previous config saved to /var/cache/conftool/dbconfig/20260701-125959-fceratto.json * 12:59 federico3: Starting s6 codfw failover from db2214 to db2229 - [[phab:T430814|T430814]] * 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on 13 hosts with reason: router upgrade and line card install * 12:51 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2229 with weight 0 [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94658 and previous config saved to /var/cache/conftool/dbconfig/20260701-125149-fceratto.json * 12:51 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 21 hosts with reason: Primary switchover s6 [[phab:T430814|T430814]] * 12:50 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2189.codfw.wmnet * 12:50 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2189.codfw.wmnet * 12:42 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2100.codfw.wmnet with OS trixie * 12:38 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2083.codfw.wmnet with OS trixie * 12:19 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2083.codfw.wmnet with reason: host reimage * 12:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 12:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2240: Migration of db2240.codfw.wmnet completed * 12:14 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2100.codfw.wmnet with reason: host reimage * 12:09 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2083.codfw.wmnet with reason: host reimage * 12:09 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2100.codfw.wmnet with reason: host reimage * 12:00 topranks: drain traffic on cr1-eqiad to allow for line card install and JunOS upgrade [[phab:T426343|T426343]] * 11:52 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2083.codfw.wmnet with OS trixie * 11:50 cmooney@dns2005: END - running authdns-update * 11:49 cmooney@dns2005: START - running authdns-update * 11:48 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2100.codfw.wmnet with OS trixie * 11:40 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/zotero: apply * 11:40 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/zotero: apply * 11:36 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/zotero: apply * 11:36 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/zotero: apply * 11:31 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2240: Migration of db2240.codfw.wmnet completed * 11:30 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply * 11:28 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply * 11:27 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:27 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:27 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:27 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:27 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:23 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2240.codfw.wmnet with OS trixie * 11:20 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:20 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:17 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:16 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:16 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:15 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2086.codfw.wmnet with OS trixie * 11:14 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2106.codfw.wmnet with OS trixie * 11:14 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:13 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:12 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:09 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2115.codfw.wmnet with OS trixie * 11:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2240.codfw.wmnet with reason: host reimage * 11:00 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2240.codfw.wmnet with reason: host reimage * 10:53 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2106.codfw.wmnet with reason: host reimage * 10:49 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2086.codfw.wmnet with reason: host reimage * 10:44 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2115.codfw.wmnet with reason: host reimage * 10:44 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2240.codfw.wmnet with OS trixie * 10:44 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2086.codfw.wmnet with reason: host reimage * 10:42 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2106.codfw.wmnet with reason: host reimage * 10:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2240: Upgrading db2240.codfw.wmnet * 10:41 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2240: Upgrading db2240.codfw.wmnet * 10:41 cwilliams@cumin1003: dbmaint on s4@codfw [[phab:T429893|T429893]] * 10:40 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 10:39 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2115.codfw.wmnet with reason: host reimage * 10:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2240 [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94653 and previous config saved to /var/cache/conftool/dbconfig/20260701-102658-cwilliams.json * 10:26 moritzm: installing nginx security updates * 10:26 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2086.codfw.wmnet with OS trixie * 10:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2179 to s4 primary [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94652 and previous config saved to /var/cache/conftool/dbconfig/20260701-102356-cwilliams.json * 10:23 cezmunsta: Starting s4 codfw failover from db2240 to db2179 - [[phab:T430127|T430127]] * 10:23 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2106.codfw.wmnet with OS trixie * 10:20 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2115.codfw.wmnet with OS trixie * 10:15 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2179 with weight 0 [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94651 and previous config saved to /var/cache/conftool/dbconfig/20260701-101531-cwilliams.json * 10:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 40 hosts with reason: Primary switchover s4 [[phab:T430127|T430127]] * 09:56 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template (take 2) - oblivian@cumin1003" * 09:56 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template (take 2) - oblivian@cumin1003 * 09:55 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template (take 2) - oblivian@cumin1003 * 09:55 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template (take 2) - oblivian@cumin1003" * 09:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:39 mszwarc@deploy1003: Synchronized private/SuggestedInvestigationsSignals/SuggestedInvestigationsSignal4n.php: Update SI signal 4n (duration: 06m 08s) * 09:21 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 09:21 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 09:14 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 09:14 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 09:02 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 09:02 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 08:54 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 08:38 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 08:38 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 08:36 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 08:21 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 08:21 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] (duration: 36m 11s) * 08:15 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 08:09 mszwarc@deploy1003: mszwarc, abi: Continuing with deployment * 08:03 mszwarc@deploy1003: mszwarc, abi: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:55 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 07:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 07:45 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] * 07:30 aqu@deploy1003: Finished deploy [analytics/refinery@410f205]: Regular analytics weekly train 2nd try [analytics/refinery@410f2050] (duration: 00m 22s) * 07:29 aqu@deploy1003: Started deploy [analytics/refinery@410f205]: Regular analytics weekly train 2nd try [analytics/refinery@410f2050] * 07:28 aqu@deploy1003: Finished deploy [analytics/refinery@410f205] (thin): Regular analytics weekly train THIN [analytics/refinery@410f2050] (duration: 01m 59s) * 07:28 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] (duration: 07m 19s) * 07:26 aqu@deploy1003: Started deploy [analytics/refinery@410f205] (thin): Regular analytics weekly train THIN [analytics/refinery@410f2050] * 07:26 aqu@deploy1003: Finished deploy [analytics/refinery@410f205]: Regular analytics weekly train [analytics/refinery@410f2050] (duration: 04m 32s) * 07:24 mszwarc@deploy1003: wmde-fisch, mszwarc: Continuing with deployment * 07:23 mszwarc@deploy1003: wmde-fisch, mszwarc: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:21 aqu@deploy1003: Started deploy [analytics/refinery@410f205]: Regular analytics weekly train [analytics/refinery@410f2050] * 07:21 aqu@deploy1003: Finished deploy [analytics/refinery@410f205] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@410f2050] (duration: 02m 01s) * 07:20 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] * 07:19 aqu@deploy1003: Started deploy [analytics/refinery@410f205] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@410f2050] * 07:13 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] (duration: 09m 13s) * 07:09 mszwarc@deploy1003: mszwarc, chlod, revi: Continuing with deployment * 07:06 mszwarc@deploy1003: mszwarc, chlod, revi: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:04 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] * 06:55 elukey: upgrade all trixie hosts to pywmflib 3.0 - [[phab:T430552|T430552]] * 06:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:43 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:43 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:42 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:42 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:41 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:41 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:35 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:35 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:34 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:34 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:31 jmm@cumin2003: DONE (PASS) - Cookbook sre.idm.logout (exit_code=0) Logging Niharika29 out of all services on: 2453 hosts * 06:30 oblivian@cumin1003: END (FAIL) - Cookbook sre.deploy.hiddenparma (exit_code=99) Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:30 oblivian@cumin1003: END (FAIL) - Cookbook sre.deploy.python-code (exit_code=99) hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:30 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:30 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:01 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2109.codfw.wmnet with OS trixie * 05:45 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on es1039.eqiad.wmnet with reason: issues * 05:41 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1027.eqiad.wmnet * 05:40 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2068.codfw.wmnet with OS trixie * 05:40 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2109.codfw.wmnet with reason: host reimage * 05:40 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1027.eqiad.wmnet,service=s2 * 05:40 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1027.eqiad.wmnet,service=s7 * 05:36 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2109.codfw.wmnet with reason: host reimage * 05:20 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2068.codfw.wmnet with reason: host reimage * 05:16 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2109.codfw.wmnet with OS trixie * 05:15 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2068.codfw.wmnet with reason: host reimage * 05:09 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2067.codfw.wmnet with OS trixie * 04:56 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2068.codfw.wmnet with OS trixie * 04:49 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2067.codfw.wmnet with reason: host reimage * 04:45 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2067.codfw.wmnet with reason: host reimage * 04:27 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2067.codfw.wmnet with OS trixie * 03:47 slyngshede@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1039.eqiad.wmnet with reason: Hardware crash * 03:21 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2107.codfw.wmnet with OS trixie * 02:59 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2107.codfw.wmnet with reason: host reimage * 02:55 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2085.codfw.wmnet with OS trixie * 02:51 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2072.codfw.wmnet with OS trixie * 02:51 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2107.codfw.wmnet with reason: host reimage * 02:35 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2085.codfw.wmnet with reason: host reimage * 02:31 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2107.codfw.wmnet with OS trixie * 02:30 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2072.codfw.wmnet with reason: host reimage * 02:26 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2085.codfw.wmnet with reason: host reimage * 02:22 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2072.codfw.wmnet with reason: host reimage * 02:09 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2085.codfw.wmnet with OS trixie * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 54s) * 02:03 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2072.codfw.wmnet with OS trixie * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es7 eqiad back to read-write - [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94649 and previous config saved to /var/cache/conftool/dbconfig/20260701-010716-ladsgroup.json * 01:05 ladsgroup@dns1004: END - running authdns-update * 01:05 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depool es1039 [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94648 and previous config saved to /var/cache/conftool/dbconfig/20260701-010551-ladsgroup.json * 01:03 ladsgroup@dns1004: START - running authdns-update * 01:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Promote es1035 to es7 primary [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94647 and previous config saved to /var/cache/conftool/dbconfig/20260701-010002-ladsgroup.json * 00:58 Amir1: Starting es7 eqiad failover from es1039 to es1035 - [[phab:T430765|T430765]] * 00:53 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es1035 with weight 0 [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94646 and previous config saved to /var/cache/conftool/dbconfig/20260701-005329-ladsgroup.json * 00:53 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 9 hosts with reason: Primary switchover es7 [[phab:T430765|T430765]] * 00:42 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es7 eqiad as read-only for maintenance - [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94645 and previous config saved to /var/cache/conftool/dbconfig/20260701-004221-ladsgroup.json * 00:20 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2102.codfw.wmnet with OS trixie * 00:15 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2103.codfw.wmnet with OS trixie * 00:05 dr0ptp4kt: DEPLOYED Refinery at {{Gerrit|4e7a2b32}} for changes: pageview allowlist {{Gerrit|1305158}} (+min.wikiquote) {{Gerrit|1305162}} (+bol.wikipedia), {{Gerrit|1305156}} (+isv.wikipedia); {{Gerrit|1305980}} (pv allowlist -api.wikimedia, sqoop +isvwiki); sqoop {{Gerrit|1295064}} (+globalimagelinks) {{Gerrit|1295069}} (+filerevision) using scap, then deployed onto HDFS (manual copyToLocal required additionally) == Other archives == See [[Server Admin Log/Archives]]. <noinclude> [[Category:SAL]] [[Category:Operations]] </noinclude> dy4f8hdytx55s5kyhbr8yqf8hipathh Logstash 0 9393 2450661 2449396 2026-08-22T19:03:39Z Quiddity 1884 lang="text" 2450661 wikitext text/x-wiki {{Navigation Wikimedia infrastructure|expand=logging}} {{See|For the frontend at logstash.wikimedia.org, see [[OpenSearch Dashboards]].}} '''Logstash''' is a tool for managing events and logs. When used generically, the term encompasses a larger system of log collection, processing, storage and searching activities. ==In a hurry?== To browse logs in production use https://logstash.wikimedia.org - see [[#Authentication]] below if needed. To browse logs in beta use https://beta-logs.wmcloud.org To log from an application see [[Logstash/MicroHOWTO]] =={{anchor|Overview ("ELK+")}} Overview== [[File:ELK Tech Talk 2015-08-20.pdf|thumb|Slides from TechTalk on ELK by Bryan Davis]] [[File:Wikipedia webrequest 2022.png|thumb|290px|Wikipedia request flow]] [[File:Using Kibana4 to read logs at Wikimedia Tech Talk 2016-11-14.pdf|thumb|Slides from TechTalk on Kibana4 by Bryan Davis]] Various Wikimedia applications send log events to '''[[Logstash]]''', which gathers the messages, converts them into JSON documents, and stores them in an '''[[OpenSearch]]''' cluster. Wikimedia uses '''[[OpenSearch Dashboards]]''' as a front-end client to filter and display messages from the OpenSearch cluster. These are the core components of our '''ELK stack''', but we use additional components as well. Since we utilize more than the core ELK components, we refer to our stack as "'''ELK+''''". (OpenSearch and OpenSearch Dashboards were forked from Elasticsearch and Kibana when those became non-free, hence the name "ELK".) ===OpenSearch=== [https://opensearch.org/docs/latest/opensearch/index/ OpenSearch] is a multi-node [https://lucene.apache.org/ Lucene] implementation. ===Logstash=== [https://www.elastic.co/logstash/ Logstash] is a tool to collect, process, and forward events and log messages. Collection is accomplished via configurable input plugins including raw socket/packet communication, file tailing, and several message bus clients. Once an input plugin has collected data it can be processed by any number of filters which modify and annotate the event data. Finally logstash routes events to output plugins which can forward the events to a variety of external programs, local files, and several message bus implementations. ===OpenSearch Dashboards=== [https://opensearch.org/docs/latest/dashboards/index/ OpenSearch Dashboards] is a browser-based analytics and search interface for OpenSearch. === Kafka === [https://kafka.apache.org/intro Apache Kafka] is a distributed streaming system. In our ELK stack Kafka buffers the stream of log messages produced by rsyslog (on behalf of applications) for consumption by Logstash. Nothing should output logs to logstash directly, logs should always be sent by way of Kafka. === Rsyslog === [https://www.rsyslog.com Rsyslog] is the "rocket-fast system for log processing". In our ELK stack rsyslog is used as the host "log agent". Rsyslog ingests log messages in various formats and from varying protocols, normalizes them and outputs to Kafka. == {{anchor|Prototype (Beta) Logstash}}OpenSearch quick intro == To learn how to use the dashboards at logstash.wikimedia.org, refer to [[OpenSearch Dashboards]]. That's also where you can find how to access the '''Beta Cluster Logstash''': [[OpenSearch Dashboards#Beta Cluster Logstash]]. ==Systems feeding into logstash== See 2015-08 Tech talk slides Writing new filters is easy. === Supported log shipping protocols & formats ("interfaces") === '''''Support of logs shipped directly from application to Logstash has been deprecated'''''. Please see [[Logstash/Interface]] for details regarding long-term supported log shipping interfaces. ==== Kubernetes ==== Kubernetes hosted services are taken care of directly by the kubernetes infrastructure which ships via rsyslog into the logstash pipeline. All a kubernetes service needs to do is log in a JSON structured format (e.g. bunyan for nodejs services) to standard output/standard error. Note that some field names mind end up being a bit tricky. See also [[Logstash/Common Logging Schema]] for an effort to standardize on them. ===Systems not feeding into logstash=== *[[:mw:Extension:EventLogging|EventLogging]] (of program-defined events with schemas), despite its name, uses a different pipeline. *[[Varnish]] logs of the billions of[[Analytics/Pageviews | pageviews]] of WMF wikis would require a lot more hardware. Instead we use [[Kafka]] to feed [[Analytics/Data Lake/Traffic/Webrequest|web requests]] into [[Hadoop]]. A notable exception to this rule: varnish user-facing errors (HTTP status 500-599) are sent to logstash to make debugging easier. *While most of [[MediaWiki at WMF|MediaWiki]] logs go to Logstash (and [[mwlog]] files), a few channels go exclusively to mwlog files. You can check which ones, via <code>$wmgMonologChannels</code> in [[:wmnoc:conf/InitialiseSettings.php.txt|InitialiseSettings.php]]. === Writing & testing filters === When in the process of writing new logstash filters, take a look at what's [[:phab:source/operations-puppet/browse/production/modules/profile/files/logstash/|existing already]] in puppet. Each filter must be tested to avoid regressions, we are using [https://github.com/magnusbaeck/logstash-filter-verifier logstash filter verifier] and existing tests can be found in the <tt>tests/</tt> directory. To write tests or run existing tests you will need logstash-filter-verifier and logstash installed locally, or you can use docker/podman and the puppet repository:<syntaxhighlight lang="bash"> # From the base dir of operations/puppet. $ cd modules/profile/files/logstash/ # The Makefile recognizes if one of podman or docker is installed # and then it uses it. $ make test-local /usr/bin/docker run --rm --workdir /src -v $(pwd):/src:Z -v $(pwd)/templates:/etc/logstash/templates:Z -v $(pwd)/filter_scripts:/etc/logstash/filter_scripts:Z --entrypoint make docker-registry.wikimedia.org/releng/logstash-filter-verifier:latest logstash-filter-verifier --diff-command="diff -u --color=always" --sockets tests/ filters/*.conf Use Unix domain sockets. [...cut...] </syntaxhighlight>Each filter has a corresponding test after its name in <tt>tests/</tt>. Within the test file the <tt>fields</tt> map lists the fields common to all tests and are used to trigger a specific filter's "if" conditions. The <tt>ignore</tt> key usually contains only <tt>@timestamp</tt> since that field is bound to change across invocations and can be safely ignored. The remainder of a test file is a list of testcases in the form of input/expected pairs. For "input" it is recommended to use yaml <tt>&gt;</tt> to include verbatim JSON, whereas "expected" is usually yaml, although it can be also verbatim JSON if more convenient. === Getting logs from misc systems into logstash === Please see [[Logstash/Interface#Tailing Log Files]]. ==Production Logstash Architecture== As of FY2019 Logstash infrastructure is owned by SRE. See also [[Logstash/SRE onboard]] for more information on how to migrate services/applications. === Architecture Diagram === <br /> [[File:Logging Pipeline Arch Diag.jpg|center]] <br /> === Web interface === https://logstash.wikimedia.org === Authentication === Log into https://logstash.wikimedia.org/ with your [[:mw:Developer account|Wikimedia Developer account]] (sometimes known as your "LDAP" account). Access to Logstash is restricted to members that are in one of the these [[SRE/LDAP/Groups|LDAP groups]], you may request access through the IDM tool, https://idm.wikimedia.org. * <code>logstash-access</code> (for Wikimedia staff which needs Logstash) * <code>nda</code> (for volunteers, soon this group will also migrate to logstash-access) * <code>ops</code> (for SRE). === Configuration === The cluster contains two types of nodes, configured by Puppet. * role::logging::opensearch::collector manages the Logstash "collector" instances. These run Logstash, an OpenSearch indexing node, and an Apache vhost serving OpenSearch Dashboards. The Apache vhosts perform LDAP-based authentication to restrict access to the potentially sensitive log information. * role::logging::opensearch::data configures an OpenSearch data node providing storage for log data. * role::kafka::logging configures a Kafka broker for producers to publish log data to and for Logstash to consume from. This is a buffering layer to absorb log spikes and queue log events when maintenance is being performed on the logging cluster. ===Hostnames=== In October of 2023, the Observability team decided to rename the Logstash cluster to better align with the role(s) of each host. The previous naming convention, regardless of role or configuration, assigned each node the hostname <code>logstash</code>. This legacy naming convention required operators to know by the host ID (1001, et. al.) the role, function, and class. The new naming convention selection criteria included: * a concise name that could be expanded as needed * indicated the difference between stateful and stateless components * did not specify the underlying technology deployed * an incremental improvement over role assignment based on node id The 10-2023 Logging Cluster Naming Convention: * <code>logging-hd</code> - OpenSearch data node: HDD Class * <code>logging-sd</code> - OpenSearch data node: SSD Class * <code>logging-fe</code> - OpenSearch API (logs-api), OpenSearch Dashboards (kibana7), Logstash Collector node === Load Balancing and TLS === The "misc" [[Varnish]] cluster is being used to provide ssl termination and load balancing support for the Kibana application. == Beta Logstash Architecture == === Web interface === : https://beta-logs.wmcloud.org/ === Access control === : Credentials for Beta's Logstash can be found on [https://office.wikimedia.org/wiki/User:BDavis_(WMF)/logstash officewiki], or by connecting to <code>deployment-deploy04.deployment-prep.eqiad1.wikimedia.cloud</code> and reading <code>/root/secrets.txt</code>. Unlike production services, Beta Cluster may not use [[:mw:Developer account|Developer accounts]] (LDAP) for authentication. : <syntaxhighlight lang="shell-session"> $ ssh deployment-deploy04.deployment-prep.eqiad1.wikimedia.cloud -- sudo cat /root/secrets.txt service: https://beta-logs.wmcloud.org user: ************ password: ************ </syntaxhighlight> === Kafka access to deployment-prep === : The security group <code>kafka-logging</code> must allow ingress from the logging collector on port 9093. : When commissioning a new logging collector, the certificate authority keystore (<code>/etc/ssl/localcerts/wmf-java-cacerts</code>) must be manually copied onto the new logging collector otherwise logstash will not start: (<code>File does not exist or cannot be opened /etc/ssl/localcerts/wmf-java-cacerts</code>). ==Common Logging Schema== Seeː [[Logstash/Common Logging Schema]]. ==API== The [https://www.elastic.co/guide/en/elasticsearch/reference/current/search-search.html Elasticsearch API] is accessible at https://logs-api.svc.eqiad.wmnet or by SSH tunneling port 9200 from an opensearch node. Note: The [https://www.elastic.co/guide/en/elasticsearch/reference/current/search-search.html _search] endpoint can only be used '''without''' a request body (see {{Phabricator|T174960}}). Use [https://www.elastic.co/guide/en/elasticsearch/reference/current/search-multi-search.html _msearch] instead for complex queries that need a request body. === Extract data from Logstash (OpenSearch) with curl and jq === <syntaxhighlight lang="bash"> logstash-server:~$ cat search.sh curl -XGET 'localhost:9200/_search?pretty&size=10000' -d ' { "query": { "query_string" : { "query" : "facility:19,local3 AND host:csw2-esams AND @timestamp:[2019-08-04T03:00 TO 2019-08-04T03:15] NOT program:mgd" } }, "sort": ["@timestamp"] } ' logstash-server:~$ bash search.sh | jq '.hits.hits[]._source | {timestamp,host,level,message}' | head -20 { "timestamp": "2019-08-04T03:00:00+00:00", "host": "csw2-esams", "level": "INFO", "message": " %-: (root) CMD (newsyslog)" } { "timestamp": "2019-08-04T03:00:00+00:00", "host": "csw2-esams", "level": "INFO", "message": " %-: (root) CMD ( /usr/libexec/atrun)" } { "timestamp": "2019-08-04T03:01:00+00:00", "host": "csw2-esams", "level": "INFO", "message": " %-: (root) CMD (adjkerntz -a)" } $ bash search.sh | jq -r '.hits.hits[]._source | {timestamp,host,level,program,message} | map(.) | @csv' > asw2-d2-eqiad-crash.csv </syntaxhighlight> ==Plugins== Logstash plugins are fetched and compiled into a Debian package for distribution and installation on Logstash servers. The plugin git repository is located at https://gerrit.wikimedia.org/r/#/admin/projects/operations/software/logstash/plugins ===Plugin build process=== The build can be run on the production builder host. See [[:gerrit:plugins/gitiles/operations/software/logstash/plugins/+/refs/heads/master/README|README]] for up-to-date build steps. =====Deployment===== * Add package to [[reprepro]] and install on the host normally. {{Note|Package installation will not restart Logstash. This must be done manually in a rolling fashion, and it's strongly suggested to perform this in step with the plugin deploy.}} : ==Gotchas== ===GELF transport=== Make sure logging events sent to the GELF input don't have a "type" or "_type" field set, or if set, that it contains the value "gelf". The gelf/Logstash config discards any events that have a different value set for "type" or "_type". The final "type" seen in OpenSearch/Dashboards will be take from the "facility" element of the original GELF packet. The application sending the log data to Logstash should set "facility" to a reasonably unique value that identifies your application. === Throttling === Log volume for a given type + channel + level + normalized message combination is throttled to 5000 / 5 min / Logstash collector (see [[:gitiles:operations/puppet/+/production/modules/profile/files/logstash/filters/75-filter_throttle.conf|75-filter_throttle.conf]]). As of 2025 this means 30K events per 5 minutes (100/sec). ==Documents== {{Special:Prefixindex/Logstash/|hideredirects=1|stripprefix=1}} == Troubleshooting == === Kafka consumer lag === For a host of reasons it might happen that there's a buildup of messages on Kafka. For example: ; OpenSearch is refusing to index messages, thus Logstash can't consume properly from Kafka. : The reason for index failure is usually conflicting fields, see also {{bug|T150106}} for a detailed discussion of the problem. The solution is to find what programs are generating the conflicts and drop them on Logstash accordingly, see also {{bug|T228089}} === Using the dead letter queue === See the [https://logstash.wikimedia.org/app/discover#/view/6086dd90-85dd-11eb-99a9-c1243d7de186 dead letter queue (DLQ)] in logstash. We save a maximum of two days of dead-letters to assist with debugging indexing issues. === No cached mapping for this field === The warning shows up next to fields for which opensearch dashboards does not have mappings yet. To refresh the fields mapping navigate to [https://logstash.wikimedia.org/app/management/opensearch-dashboards/indexPatterns index patterns], select the index pattern for your logs (likely logstash-* or ecs-*) then hit the "loop" or "refresh" icon on the top right. == Operations == === Configuration changes === After merging your configuration change Puppet will automatically restart Logstash. To force this to run: cumin -b1 -s60 'O:logging::opensearch::collector' 'run-puppet-agent -q' ==== Test a configuration snippet before merge ==== Copy your ready to merge snippet (eg. modules/profile/files/logstash/filter-syslog-network.conf) to a Logstash host. Then run sudo /usr/share/logstash/bin/logstash --config.test_and_exit -f <myfile> It should return "Configuration OK". === Indexing errors === Have a look at the [https://logstash.wikimedia.org/app/discover#/view/6086dd90-85dd-11eb-99a9-c1243d7de186 Dead Letter Queue Dashboard]. The original message that caused the error is in the <code>log.original</code> field. We're alerting on errors that Logstash gets from OpenSearch whenever there's an "indexing conflict" between fields of the same index (see also {{Bug|T236343}}). The reason usually is because two applications send logs with the same field name but two different types, e.g. <code>response</code> will be sent as a string in one case but as nested object in another. {{Bug|T239458}} is a good example of this, where different parts of mediawiki send logs formatted in a different way. === No logs indexed === This alert is based on the incoming logs per second indexed by OpenSearch. During normal operation there is a baseline of ~1k logs/s (July 2020) and anything significantly lower than that is an unexpected condition. Check the Logstash dashboard attached to alert for signs of root causes. Most likely Logstash has stopped sending logs to OpenSearch. === Drop spammy logs === Occasionally producers will outpace Logstash's ingestion capabilities, most often with what's considered "log spam" (e.g. dumping whole request/response in debug logs). In these case one solution is to drop the offending logs from Logstash, and ideally the producer has already stopped spamming. The simplest such filter is installed before most/all other filters, matches a few fields and then <code>drop</code>s the message: <pre> filter { if [program] == "producer" and [nested][field] == "offending value" { drop {} } } </pre> See also [[:gerrit:c/operations/puppet/+/713853|this Gerrit change]] for a real-world example. === UDP packet loss === Logstash 5 locks up from time to time, causing UDP packet loss on the host it is running on. The fix in this case is to restart <code>logstash.service</code> on the host in question. === Replace failed disk and rebuild RAID === The storage drives on Logstash data are configured in an mdraid RAID0. OpenSearch handles data redundancy, so the rest of the cluster will absorb the impact of the downed node. Once the disk is replaced, the RAID will have to be rebuilt: First stop opensearch and disable puppet. Copy disk partition layout from good disk to new disk sfdisk -d /dev/sdb | sfdisk /dev/sdi determine md device mounted at /srv (/dev/md2 for example) and check mdstat cat /proc/mdstat get array information. make a note of the remaining array members, we'll need this information when rebuilding mdadm --query --detail /dev/md2 stop and remove the raid0 array mdadm --stop /dev/md2 && mdadm --remove /dev/md2 remove traces of the previous array on the old partitions mdadm --zero-superblock /dev/sdb4 mdadm --zero-superblock /dev/sdc4 # ... etc mdadm --zero-superblock /dev/sdh4 create new raid0 array (WARNING: DISKS MAY BE DIFFERENT) mdadm --create --verbose /dev/md/2 --level=0 --raid-devices=8 /dev/sdb4 /dev/sdc4 /dev/sdd4 /dev/sde4 /dev/sdf4 /dev/sdg4 /dev/sdh4 /dev/sdi4 make filesystem mkfs.ext4 /dev/md2 workaround systemd mount management by commenting out the old array mount in fstab and issuing a daemon reload vim /etc/fstab systemctl daemon-reload add mount point back in with new uuid and mount vim /etc/fstab mount /srv check to make sure the disk is mounted and add new array definition to /etc/mdadm/mdadm.conf - also remove old definition mdadm --detail --scan vim /etc/mdadm/mdadm.conf update initramfs update-initramfs -u check other arrays for failed partitions and add partitions to them mdadm --manage /dev/md0 --add /dev/sdi2 mdadm --manage /dev/md1 --add /dev/sdi3 make opensearch data directory mkdir /srv/opensearch && chown opensearch:opensearch /srv/opensearch re-enable puppet and run puppet. OpenSearch should start up, join the cluster, and immediately start rebalancing shards. Note: if an array is stuck PENDING with <code>auto-read-write</code>, it will either recover upon receiving its first write or you can manually invoke the resync with <code>mdadm --readwrite /dev/mdX</code> === Restore Dashboards from backup === From a single collector node, delete all <code>.kibana</code> indexes and restart opensearch-dashboards. Check that the restart created <code>.kibana_1</code> and aliased it with <code>.kibana</code>. Fetch and unzip the backup and run: BACKUP_FILE=<myfile>.ndjson; curl -s -X POST http://localhost:5601/api/saved_objects/_import?createNewCopies=false -H "osd-xsrf: true" --form file=@$BACKUP_FILE > response.json Check the response for problems and navigate into OpenSearch Dashboards to ensure expected saved objects are present. === Unassigned Shards and Cluster Status === Shard rebalancing operations are handled automatically by OpenSearch, however there are situations where shards cannot be assigned. Some of these possible reasons are: * There is not enough space in the cluster to avoid exceeding a watermark. ** This is often caused by the cluster storing too much data or a node has failed and the cluster is at reduced capacity. The solution is either to restore the failed node, reduce the number of replicas to free up space, or delete data to free up space. * The index settings declare there should be more shards than there are number of nodes. ** The solution is either to restore the failed node, reduce the number of replicas to free up space, or delete data to free up space. * There is no host capable of accepting the shard due to label filters. ** Ensure the cluster has a node labeled with the same value the index is expecting. "RED" status: Occurs when OpenSearch reports there are unassigned primary shards. From the cluster perspective, this is a data-loss event. Usually, the only action required is to restore a node that has a full, uncorrupted, copy of the under-replicated shard. "YELLOW" status: Occurs when OpenSearch reports there are unassigned replica shards. "GREEN" status: Cluster is operating normally. === Rebooting hosts === When hosts need to be rebooted (for ex. Kernel/firmware upgrades) it's necessary to follow this process to reduce the risk of data loss and to reduce the number of operations like leader elections. There are two main types of Logstash hosts, the storage nodes and collector nodes, they require a different reboot process. The [https://grafana.wikimedia.org/d/VCK8-FpZz/cwhite-logstash Logstash Grafana dashboard] is a useful resource for monitoring node status including shard synchronization. ==== Rebooting storage nodes: ==== These nodes are used for persistent storage of data as shards, being primary shards the immediate data received and secondary shards copies of the received data. The process for rebooting SSD hosts and HDD hosts is the same. The cluster also has a main node (also called "coordinating" or "master") that coordinates cluster actions. This node is identified internally by a cluster election process. ===== Steps overview: ===== # Identify the main node, in order to prevent many leader election cycles. It is recommended to reboot this node after all of the non-main nodes have been rebooted. # Use the cluster settings API to only allow shard allocation of primary shards. This is to prevent the cluster from trying to make more shard copies while the node originally holding those copies is rebooting. # Reboot a non-main data node. Nodes should be rebooted one by one to prevent data unavailability or loss in the event a node does not come back from reboot. # After the rebooted node has come back up, use the cluster settings API to enable shard allocation of all shards. # Once replica shards have all been assigned repeat steps 2 through 5 until all non-main cluster nodes are rebooted. # Once all non-main nodes are rebooted, repeat from step 2 rebooting the main node - a leader election will occur. # Once all nodes are up and joined to the cluster, re-enable shard allocation. ===== Steps with commands: ===== The following steps are performed from a cumin host by using the logstash API from a logstash host. # Identify the main node:<syntaxhighlight lang="bash"> sudo cumin logstash1033.eqiad.wmnet 'curl -s "http://localhost:9200/_cat/nodes"' </syntaxhighlight>This command shows the full list of logstash host, the main node is represented by an '''*:'''<syntaxhighlight lang="bash"> 10.64.16.143 37 90 34 1.52 2.06 2.43 ir ingest,remote_cluster_client - logstash1032-production-elk7-eqiad 10.64.162.10 58 99 9 4.47 5.16 4.96 dimr data,ingest,master,remote_cluster_client * logging-sd1004-production-elk7-eqiad 10.64.32.112 87 93 11 2.89 2.84 2.81 dimr data,ingest,master,remote_cluster_client - logstash1034-production-elk7-eqiad </syntaxhighlight>This indicates that <code>logging-sd1004-production-elk7-eqiad</code> is the main node and it should be rebooted last. # Once the main node is identified the rest of commands will be executed from it:<syntaxhighlight lang="bash"> sudo cumin logging-sd1004.eqiad.wmnet 'curl -s "http://localhost:9200/_cluster/health?pretty"' </syntaxhighlight>Ensure status is '''green'''<syntaxhighlight lang="text"> { "cluster_name" : "production-elk7-eqiad", "status" : "green", "timed_out" : false, "number_of_nodes" : 23, "number_of_data_nodes" : 17, "discovered_master" : true, "discovered_cluster_manager" : true, "active_primary_shards" : 848, "active_shards" : 1970, "relocating_shards" : 0, "initializing_shards" : 0, "unassigned_shards" : 0, "delayed_unassigned_shards" : 0, "number_of_pending_tasks" : 0, "number_of_in_flight_fetch" : 0, "task_max_waiting_in_queue_millis" : 0, "active_shards_percent_as_number" : 100.0 } </syntaxhighlight> # Disable non-primary shard allocation:<syntaxhighlight lang="bash"> sudo cumin -m sync logging-sd1004.eqiad.wmnet "curl -sS -X PUT 'http://localhost:9200/_cluster/settings?pretty' -H 'Content-Type: application/json' -d '{\"transient\":{\"cluster.routing.allocation.enable\":\"primaries\"}}'" </syntaxhighlight>Ensure only primary sharding is enabled<syntaxhighlight lang="text"> { "acknowledged" : true, "persistent" : { }, "transient" : { "cluster" : { "routing" : { "allocation" : { "enable" : "primaries" } } } } } </syntaxhighlight> # Reboot another cluster data node:<syntaxhighlight lang="bash"> sudo cumin secondary-logstash-host 'reboot-host' </syntaxhighlight>Wait for host to be back after reboot:<syntaxhighlight lang="bash"> ping secondary-logstash-host </syntaxhighlight> # Enable all shard allocation:<syntaxhighlight lang="bash"> sudo cumin -m sync logging-sd1004.eqiad.wmnet "curl -sS -X PUT 'http://localhost:9200/_cluster/settings?pretty' -H 'Content-Type: application/json' -d '{\"transient\":{\"cluster.routing.allocation.enable\":\"all\"}}'" </syntaxhighlight>Ensure all shard allocation is enabled:<syntaxhighlight lang="text"> { "acknowledged" : true, "persistent" : { }, "transient" : { "cluster" : { "routing" : { "allocation" : { "enable" : "all" } } } } } </syntaxhighlight> # Verify the sync of shard allocation: This can be done using the Grafana dashboard linked above or with the following command:<syntaxhighlight lang="bash"> sudo cumin logging-sd1004.eqiad.wmnet 'curl -s "http://localhost:9200/_cluster/health?pretty"' </syntaxhighlight>Wait until '''relocating_shards, initializing_shards, and unassigned_shards''' is 0.<syntaxhighlight lang="text"> { "cluster_name" : "production-elk7-eqiad", "status" : "green", "timed_out" : false, "number_of_nodes" : 23, "number_of_data_nodes" : 17, "discovered_master" : true, "discovered_cluster_manager" : true, "active_primary_shards" : 848, "active_shards" : 1970, "relocating_shards" : 0, "initializing_shards" : 0, "unassigned_shards" : 0, "delayed_unassigned_shards" : 0, "number_of_pending_tasks" : 0, "number_of_in_flight_fetch" : 0, "task_max_waiting_in_queue_millis" : 0, "active_shards_percent_as_number" : 100.0 } </syntaxhighlight> # Continue from step 3 with the remaining non-main nodes, then reboot the main node at the end. ==== Rebooting collector nodes: ==== These nodes don't store data so no shard manipulation is required but it's important that services are stopped in a specific order to prevent blocking states by stopping dependant services first (ex. Stopping the network service before the logstash service could cause logstash to be unable to continue it's operations and enter a blocking state). # Depool the host to reboot # Stop the logstash service, wait for the service to finish all remaining operations. # Reboot the host. # Once the host come back up, re-pool the host. # Follow the process for the rest of the hosts. == Stats == === Documents and bytes counts === The OpenSearch ''cat'' API provides a simple way to extract general statistics about log storage, e.g. total logs and bytes (not including replication) logstash1010:~$ curl -s 'localhost:9200/_cat/indices?v&bytes=b' | awk '/logstash-/ {b+=$10; d+=$7} END {print d; print b}' Or logs per day (change $3 to $7 to get bytes sans replication) logstash1010:~$ curl -s 'localhost:9200/_cat/indices?v&bytes=b' | awk '/logstash-/ { gsub(/logstash-[^0-9]*/, "", $3); sum[$3] += $7 } END { for (i in sum) print i, sum[i] }' | sort Or logs per month: logstash1010:~$ curl -s 'localhost:9200/_cat/indices?v&bytes=b' | awk '/logstash-/ { gsub(/logstash-[^0-9]*/, "", $3); gsub(/\.[0-9][0-9]$/, "", $3); sum[$3] += $7 } END { for (i in sum) print i, sum[i] }' | sort == Data Retention == Logs are retained in Logstash for a maximum of 90 days by default in accordance with our [[:foundation:Privacy policy|Privacy Policy]] and [[:m:Data retention guidelines|Data Retention Guidelines]]. Indexes are arranged into buckets following this naming convention:<code>^(?<output>[a-z0-9]+)-(?<partition>[a-z0-9]+)-(?<policy_revision>[0-9]+)-(?<template_version>[0-9\.]+)-(?<template_revision>[0-9]+)-(?<date_stamp>[0-9.]+)$</code> * <code><output></code>: The Logstash output filter name that writes to this index. Defined by Puppet. * <code><partition></code>: Splits the output into smaller buckets. Improves search and indexing performance by isolating log streams into smaller, logical buckets. e.g. separating logstash-webrequest from logstash-mediawiki, etc. * <code><policy_revision></code>: The index policy revision. This is used for linking index names to Curator actions via Curator pattern filters. * <code><template_version></code>: The mapping template version applied to the index. e.g. 7.10.0 * <code><template_revision></code>: The mapping template revision applied to the index. Enables immediate updates to mapping template changes without requiring index rollover or Curator action changes. * <code><date_stamp></code>: The temporal size of the index. Used for applying retention rules. This can be: ** <code>YYYY.MM.DD</code>: "daily" indexes ** <code>YYYY.WW</code>: "weekly" indexes ** <code>YYYY</code>: "yearly" indexes Design: [[:phab:T305175]] === Extended Retention === See [[Logstash/Extended Retention]] ==See also== *[[:mw:Manual:Structured logging]] (MediaWiki part of the job to feed into Logstash) *[[Logs#mw-log]] (the old method of viewing logs) *[[:phab:J177|Introducing Phatality]] (a Kibana plugin to streamline the process of reporting production errors on Phabricator) *[[Kubernetes/Logging]] How logs flow into Logstash from the Kubernetes components {{Ptag|Wikimedia-Logstash}} [[Category:SRE Observability]] [[Category:Services]] buumuj158aa4l0q9j8k6jgkp1l60vko Nova Resource:Admin/SAL 498 30942 2450660 2450637 2026-08-22T17:27:47Z Stashbot 7414 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) 2450660 wikitext text/x-wiki === 2026-08-22 === * 17:27 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 00:02 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 00:02 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 00:01 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 00:01 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 00:01 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add === 2026-08-21 === * 16:11 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 16:08 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 16:08 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 16:05 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 16:05 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 16:03 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 16:03 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 14:55 andrewbogott: reimaging cloudvirt1042, getting a 100% fresh start for [[phab:T429387|T429387]] testing * 06:25 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) ([[phab:T429387|T429387]]) * 00:09 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy ([[phab:T429387|T429387]]) * 00:08 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) ([[phab:T429387|T429387]]) * 00:08 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy ([[phab:T429387|T429387]]) === 2026-08-20 === * 23:50 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) ([[phab:T429387|T429387]]) * 19:36 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy ([[phab:T429387|T429387]]) * 17:31 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=97) ([[phab:T429387|T429387]]) * 17:31 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy ([[phab:T429387|T429387]]) * 16:57 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=97) ([[phab:T429387|T429387]]) * 13:58 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy ([[phab:T429387|T429387]]) === 2026-08-19 === * 02:51 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.unset_cluster_maintenance (exit_code=0) * 02:51 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.unset_cluster_maintenance * 02:10 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.set_cluster_in_maintenance (exit_code=0) * 02:09 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.set_cluster_in_maintenance === 2026-08-18 === * 18:49 andrewbogott: rebooting cloudcephosd1046 to pick up (or reset?) drive caching changes * 17:36 andrewbogott: changing ssd write cache to 'off' and 'write-through' on all drives in cloudcephosd1046, restarting osd services [[phab:T429387|T429387]] * 01:13 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.roll_reboot_osds (exit_code=0) ([[phab:T434750|T434750]]) === 2026-08-17 === * 21:20 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_osds ([[phab:T434750|T434750]]) * 21:19 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.roll_reboot_mons (exit_code=0) ([[phab:T434750|T434750]]) * 21:07 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_mons ([[phab:T434750|T434750]]) === 2026-08-15 === * 16:57 andrewbogott: optimistically deploying a blind fix for striker issue [[phab:T434684|T434684]] === 2026-08-13 === * 16:06 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 16:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 16:02 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 16:01 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 16:00 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 16:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 12:34 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.roll_reboot_mons (exit_code=0) * 12:23 filippo@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_mons * 10:00 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.roll_reboot_osds (exit_code=0) * 09:46 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1057' ([[phab:T431682|T431682]]) * 09:40 filippo@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_osds * 09:30 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1057' ([[phab:T431682|T431682]]) * 09:18 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1056' ([[phab:T431682|T431682]]) * 09:01 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1056' ([[phab:T431682|T431682]]) * 09:01 godog: put cloudvirt1064 back into aggregate network-ovs * 08:42 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1055' ([[phab:T431682|T431682]]) * 08:22 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1055' ([[phab:T431682|T431682]]) * 08:19 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1054' ([[phab:T431682|T431682]]) * 08:02 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1054' ([[phab:T431682|T431682]]) * 07:56 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1051' ([[phab:T431682|T431682]]) * 07:39 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1051' ([[phab:T431682|T431682]]) * 07:38 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1049' ([[phab:T431682|T431682]]) * 07:29 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1049' ([[phab:T431682|T431682]]) === 2026-08-11 === * 15:20 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.roll_restart_osd_daemons (exit_code=0) ({{Gerrit|1324319}}) * 14:58 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_restart_osd_daemons ({{Gerrit|1324319}}) === 2026-08-09 === * 19:25 andrewbogott: "ceph osd out 224" in response to "osd.224 observed stalled read indications in DB device" === 2026-08-06 === * 20:24 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.roll_restart_mon_daemons (exit_code=0) * 20:24 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_restart_mon_daemons * 19:21 andrewbogott: restarting osds one by one to pick up config changes === 2026-08-05 === * 19:54 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,magnum * 19:53 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,magnum * 02:17 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.roll_reboot_osds (exit_code=0) ([[phab:T429387|T429387]]) * 01:37 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_osds ([[phab:T429387|T429387]]) * 01:36 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.roll_reboot_osds (exit_code=99) ([[phab:T429387|T429387]]) * 01:36 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_osds ([[phab:T429387|T429387]]) === 2026-08-04 === * 23:47 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.roll_reboot_osds (exit_code=99) ([[phab:T429387|T429387]]) * 20:07 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_osds ([[phab:T429387|T429387]]) * 20:03 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.roll_reboot_osds (exit_code=97) ([[phab:T429387|T429387]]) * 20:01 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_osds ([[phab:T429387|T429387]]) * 18:59 andrewbogott: systemctl reset-failed on cloudbackup2003. This patch should prevent future such false alarms: https://gerrit.wikimedia.org/r/c/operations/puppet/+/1321058 === 2026-08-03 === * 16:52 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) ([[phab:T431374|T431374]]) * 16:52 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance ([[phab:T431374|T431374]]) * 14:50 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) ([[phab:T431682|T431682]]) * 14:50 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance ([[phab:T431682|T431682]]) * 14:50 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=97) ([[phab:T431374|T431374]]) * 14:50 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance ([[phab:T431374|T431374]]) === 2026-07-28 === * 13:24 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.roll_reboot_mons (exit_code=0) ([[phab:T431659|T431659]]) * 13:14 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) * 13:11 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_mons ([[phab:T431659|T431659]]) * 13:06 wmbot~dcaro@acme: END (PASS) - Cookbook wmcs.openstack.get_project_for_proxy (exit_code=0) * 13:05 wmbot~dcaro@acme: START - Cookbook wmcs.openstack.get_project_for_proxy * 13:04 wmbot~dcaro@acme: END (PASS) - Cookbook wmcs.openstack.get_project_for_proxy (exit_code=0) * 13:04 wmbot~dcaro@acme: START - Cookbook wmcs.openstack.get_project_for_proxy * 13:04 wmbot~dcaro@acme: END (PASS) - Cookbook wmcs.openstack.get_project_for_proxy (exit_code=0) * 13:04 wmbot~dcaro@acme: START - Cookbook wmcs.openstack.get_project_for_proxy * 13:04 wmbot~dcaro@acme: END (ERROR) - Cookbook wmcs.openstack.get_project_for_proxy (exit_code=1) * 13:03 wmbot~dcaro@acme: START - Cookbook wmcs.openstack.get_project_for_proxy * 13:03 wmbot~dcaro@acme: END (FAIL) - Cookbook wmcs.openstack.get_project_for_proxy (exit_code=99) * 13:03 wmbot~dcaro@acme: START - Cookbook wmcs.openstack.get_project_for_proxy * 13:02 wmbot~dcaro@acme: END (PASS) - Cookbook wmcs.openstack.get_project_for_proxy (exit_code=0) * 13:02 wmbot~dcaro@acme: START - Cookbook wmcs.openstack.get_project_for_proxy * 13:00 wmbot~dcaro@acme: END (PASS) - Cookbook wmcs.openstack.get_project_for_proxy (exit_code=0) * 13:00 wmbot~dcaro@acme: START - Cookbook wmcs.openstack.get_project_for_proxy * 13:00 wmbot~dcaro@acme: END (PASS) - Cookbook wmcs.openstack.get_project_for_proxy (exit_code=0) * 12:59 wmbot~dcaro@acme: START - Cookbook wmcs.openstack.get_project_for_proxy * 12:59 wmbot~dcaro@acme: END (FAIL) - Cookbook wmcs.openstack.get_project_for_proxy (exit_code=99) * 12:59 wmbot~dcaro@acme: START - Cookbook wmcs.openstack.get_project_for_proxy * 12:55 wmbot~dcaro@acme: END (FAIL) - Cookbook wmcs.openstack.get_project_for_proxy (exit_code=99) * 12:55 wmbot~dcaro@acme: START - Cookbook wmcs.openstack.get_project_for_proxy * 12:54 wmbot~dcaro@acme: END (FAIL) - Cookbook wmcs.openstack.get_project_for_proxy (exit_code=99) * 12:54 wmbot~dcaro@acme: START - Cookbook wmcs.openstack.get_project_for_proxy * 12:54 wmbot~dcaro@acme: END (FAIL) - Cookbook wmcs.openstack.get_project_for_proxy (exit_code=99) * 12:54 wmbot~dcaro@acme: START - Cookbook wmcs.openstack.get_project_for_proxy * 12:12 wmbot~dcaro@acme: END (FAIL) - Cookbook wmcs.openstack.get_project_for_proxy (exit_code=99) * 12:12 wmbot~dcaro@acme: START - Cookbook wmcs.openstack.get_project_for_proxy * 12:08 wmbot~dcaro@acme: END (FAIL) - Cookbook wmcs.openstack.get_project_for_proxy (exit_code=99) * 12:08 wmbot~dcaro@acme: START - Cookbook wmcs.openstack.get_project_for_proxy * 12:06 wmbot~dcaro@acme: END (FAIL) - Cookbook wmcs.openstack.get_project_for_proxy (exit_code=99) * 12:05 wmbot~dcaro@acme: START - Cookbook wmcs.openstack.get_project_for_proxy * 12:00 wmbot~dcaro@acme: END (ERROR) - Cookbook wmcs.openstack.get_project_for_proxy (exit_code=1) * 11:57 wmbot~dcaro@acme: START - Cookbook wmcs.openstack.get_project_for_proxy * 11:55 wmbot~dcaro@acme: END (FAIL) - Cookbook wmcs.openstack.get_project_for_proxy (exit_code=99) * 11:55 wmbot~dcaro@acme: START - Cookbook wmcs.openstack.get_project_for_proxy * 11:55 wmbot~dcaro@acme: END (FAIL) - Cookbook wmcs.openstack.get_project_for_proxy (exit_code=99) * 11:55 wmbot~dcaro@acme: START - Cookbook wmcs.openstack.get_project_for_proxy * 03:13 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 02:40 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.roll_reboot_osds (exit_code=0) * 02:07 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_osds * 02:02 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.roll_reboot_osds (exit_code=99) * 01:57 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_osds === 2026-07-27 === * 22:33 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.roll_reboot_osds (exit_code=99) * 18:50 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_osds * 18:49 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.roll_reboot_osds (exit_code=99) * 18:49 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_osds * 16:26 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.roll_reboot_osds (exit_code=99) * 15:52 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_osds * 15:51 andrewbogott: temporarily muting ceph slow ops alerts ("ceph health mute BLUESTORE_SLOW_OP_ALERT --sticky") for [[phab:T431659|T431659]] === 2026-07-24 === * 15:10 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.roll_reboot_osds (exit_code=0) * 14:49 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_osds === 2026-07-23 === * 13:44 andrewbogott: restarting designate services in eqiad1; seeing many miscellaneous designate-sink errors in logs * 13:44 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,designate * 13:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,designate === 2026-07-22 === * 18:16 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1071.eqiad.wmnet' ([[phab:T431374|T431374]]) * 18:11 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1071.eqiad.wmnet' ([[phab:T431374|T431374]]) * 18:05 andrewbogott: serveraction powercycle on cloudvirt1071 * 18:04 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) ([[phab:T431374|T431374]]) * 18:04 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance ([[phab:T431374|T431374]]) * 12:46 taavi: add security group rules for all projects for new metricsinfra VIPs [[phab:T401813|T401813]] * 02:25 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 02:24 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch === 2026-07-21 === * 15:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 15:27 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 15:27 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 15:27 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 15:23 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 15:22 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 15:22 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 15:21 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 15:20 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 15:20 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 15:11 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 15:10 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 15:07 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 15:06 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 14:38 godog: switch from SystemdUnitDown (eqiad only) to SystemdUnitFailed (codfw/eqiad) alerts -- there might be some codfw noise coming - [[phab:T428873|T428873]] === 2026-07-20 === * 21:27 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 21:27 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 21:27 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 21:26 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 21:25 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 21:25 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 21:13 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 21:12 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 20:51 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 20:51 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 20:35 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 20:34 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:07 filippo@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.roll_reboot_mons (exit_code=99) * 12:07 filippo@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_mons * 12:01 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.roll_reboot_cloudnets (exit_code=0) * 11:25 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudservices.safe_reboot (exit_code=0) on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::services<nowiki>}</nowiki>' * 11:19 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.cloudservices.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::services<nowiki>}</nowiki>' * 10:13 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudlb.safe_reboot (exit_code=0) on hosts matched by 'A:cloudlb AND A:eqiad' * 10:08 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.cloudlb.safe_reboot on hosts matched by 'A:cloudlb AND A:eqiad' * 10:02 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.reboot_node (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudcontrol1011.eqiad.wmnet<nowiki>}</nowiki>' * 09:58 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.reboot_node on hosts matched by 'D<nowiki>{</nowiki>cloudcontrol1011.eqiad.wmnet<nowiki>}</nowiki>' * 09:53 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.reboot_node (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudcontrol1007.eqiad.wmnet<nowiki>}</nowiki>' * 09:49 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.reboot_node on hosts matched by 'D<nowiki>{</nowiki>cloudcontrol1007.eqiad.wmnet<nowiki>}</nowiki>' * 09:26 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.reboot_node (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudcontrol1006.eqiad.wmnet<nowiki>}</nowiki>' * 09:21 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.reboot_node on hosts matched by 'D<nowiki>{</nowiki>cloudcontrol1006.eqiad.wmnet<nowiki>}</nowiki>' === 2026-07-19 === * 16:44 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services * 16:31 andrewbogott: restarting eqiad1 openstack services; trying to resolve failures with allocating and deleting network interfaces * 16:28 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services === 2026-07-16 === * 14:17 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 14:17 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2026-07-15 === * 13:07 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) ([[phab:T431429|T431429]]) * 13:07 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance ([[phab:T431429|T431429]]) === 2026-07-14 === * 14:36 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudweb.safe_reboot (exit_code=0) on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::cloudweb<nowiki>}</nowiki>' * 14:31 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.cloudweb.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::cloudweb<nowiki>}</nowiki>' * 12:12 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1048' ([[phab:T431682|T431682]]) * 11:58 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1048' ([[phab:T431682|T431682]]) * 11:57 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) ([[phab:T431682|T431682]]) * 11:57 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance ([[phab:T431682|T431682]]) === 2026-07-09 === * 12:54 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/332 * 12:53 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/332 * 12:29 godog: wmcs-openstack aggregate delete maintenance * 10:11 filippo@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=99) * 10:11 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 10:11 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) * 10:10 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance === 2026-07-08 === * 19:25 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 19:25 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 19:23 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 19:23 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 17:21 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 17:21 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 17:20 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/326 * 17:20 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/326 * 17:05 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 17:04 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 17:03 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/323 * 17:03 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/323 * 16:58 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/322 * 16:58 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/322 * 16:57 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 16:56 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 16:55 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/321 * 16:55 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/321 * 16:53 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/321 * 16:53 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/321 * 16:49 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/321 * 16:49 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/321 * 13:52 andrewbogott: removing some cloudvirts from the 'maintenance' aggregate. I imagine they are there in error after some automated reboots. cloudvirt1049, cloudvirt1053, cloudvirt1064, cloudvirt1078, cloudvirt1080, cloudvirtlocal1001 === 2026-07-07 === * 09:14 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 09:13 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 09:07 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/319 * 09:06 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/319 * 04:27 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1071.eqiad.wmnet' ([[phab:T431374|T431374]]) * 04:23 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1071.eqiad.wmnet' ([[phab:T431374|T431374]]) * 04:18 andrewbogott: cycling power on cloudvirt1071 via mgmt/racadm; it seems unresponsive === 2026-07-06 === * 08:04 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'P<nowiki>{</nowiki>P:openstack::codfw1dev::nova::compute::service<nowiki>}</nowiki> AND NOT P<nowiki>{</nowiki>F:kernelversion = 6.12.95<nowiki>}</nowiki>' * 07:08 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>P:openstack::codfw1dev::nova::compute::service<nowiki>}</nowiki> AND NOT P<nowiki>{</nowiki>F:kernelversion = 6.12.95<nowiki>}</nowiki>' === 2026-06-29 === * 13:17 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 13:14 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services === 2026-06-25 === * 06:51 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) ([[phab:T424802|T424802]]) * 06:51 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance ([[phab:T424802|T424802]]) === 2026-06-22 === * 18:57 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.nfs.migrate_service (exit_code=99) * 18:57 andrew@cloudcumin1001: START - Cookbook wmcs.nfs.migrate_service === 2026-06-17 === * 13:57 dhinus: updated wikireplicas-utils from 0.1.0 to 0.2.0 on clouddb* === 2026-06-16 === * 17:51 andrewbogott: ceph tell osd.126,127,129,131 compact * 17:04 andrewbogott: rebooting cloudcephosd1037 and cloudcephosd1038 because of missing volumes. The volumes reappeared after boot. * 16:47 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.upgrade_osds (exit_code=99) ([[phab:T428385|T428385]]) * 16:47 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_osds ([[phab:T428385|T428385]]) * 16:47 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol1006.eqiad.wmnet' ([[phab:T429361|T429361]]) * 16:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1006.eqiad.wmnet' ([[phab:T429361|T429361]]) * 16:35 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol1007.eqiad.wmnet' ([[phab:T429361|T429361]]) * 16:23 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1007.eqiad.wmnet' ([[phab:T429361|T429361]]) * 16:22 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol1011.eqiad.wmnet' ([[phab:T429361|T429361]]) * 16:12 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1011.eqiad.wmnet' ([[phab:T429361|T429361]]) * 16:01 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol2010-dev.codfw.wmnet' ([[phab:T429361|T429361]]) * 15:52 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2010-dev.codfw.wmnet' ([[phab:T429361|T429361]]) * 15:52 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudcontrol2011-dev.codfw.wmnet' ([[phab:T429361|T429361]]) * 15:52 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2011-dev.codfw.wmnet' ([[phab:T429361|T429361]]) * 15:52 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol2005-dev.codfw.wmnet' ([[phab:T429361|T429361]]) * 15:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2005-dev.codfw.wmnet' ([[phab:T429361|T429361]]) * 15:40 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol2006-dev.codfw.wmnet' ([[phab:T429361|T429361]]) * 15:33 andrewbogott: restarting quite a few other ceph-osd services in an attempt to quiet some slow op alerts * 15:31 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2006-dev.codfw.wmnet' ([[phab:T429361|T429361]]) * 15:13 andrewbogott: systemctl restart ceph-osd@58.service due to slow ops * 14:57 andrewbogott: "systemctl restart ceph-osd@281.service" as 281 shows as down * 14:55 andrewbogott: "systemctl restart ceph-osd@289.service" as 289 shows as down * 13:24 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudservices.safe_reboot (exit_code=0) on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::codfw1dev::services<nowiki>}</nowiki>' * 13:18 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudservices.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::codfw1dev::services<nowiki>}</nowiki>' === 2026-06-15 === * 21:48 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.upgrade_osds (exit_code=0) * 21:04 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_osds ([[phab:T428385|T428385]]) * 21:02 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.upgrade_osds (exit_code=99) * 19:46 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_osds ([[phab:T428385|T428385]]) * 19:32 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.upgrade_osds (exit_code=0) * 19:00 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_osds ([[phab:T428385|T428385]]) * 18:40 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.upgrade_osds (exit_code=99) * 18:27 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_osds ([[phab:T428385|T428385]]) * 17:07 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.upgrade_osds (exit_code=99) * 14:32 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_osds ([[phab:T428385|T428385]]) * 14:10 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.upgrade_mons (exit_code=0) * 13:50 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_mons ([[phab:T428385|T428385]]) * 13:48 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.upgrade_osds (exit_code=97) * 13:38 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_osds ([[phab:T428385|T428385]]) * 13:08 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.upgrade_mons (exit_code=99) * 12:46 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_mons ([[phab:T428385|T428385]]) === 2026-06-12 === * 14:00 volans: clearing cloudback-original snapshots of cinder volumes created in 2026 to allow the backups to not fail for [[phab:T428995|T428995]] === 2026-06-11 === * 17:53 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol1011.eqiad.wmnet' ([[phab:T428549|T428549]]) * 17:39 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1011.eqiad.wmnet' ([[phab:T428549|T428549]]) * 17:18 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol1007.eqiad.wmnet' ([[phab:T428549|T428549]]) * 17:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1007.eqiad.wmnet' ([[phab:T428549|T428549]]) * 16:50 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol1006.eqiad.wmnet' ([[phab:T428549|T428549]]) * 16:37 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1006.eqiad.wmnet' ([[phab:T428549|T428549]]) * 16:33 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudcontrol2006.eqiad.wmnet' ([[phab:T428549|T428549]]) * 16:33 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2006.eqiad.wmnet' ([[phab:T428549|T428549]]) * 15:11 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol2010-dev.codfw.wmnet' ([[phab:T428549|T428549]]) * 14:58 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2010-dev.codfw.wmnet' ([[phab:T428549|T428549]]) * 14:41 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol2006-dev.codfw.wmnet' ([[phab:T428549|T428549]]) * 14:24 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2006-dev.codfw.wmnet' ([[phab:T428549|T428549]]) * 14:22 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol2005-dev.codfw.wmnet' ([[phab:T428549|T428549]]) * 14:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2005-dev.codfw.wmnet' ([[phab:T428549|T428549]]) === 2026-06-09 === * 17:35 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.upgrade_osds (exit_code=0) * 17:05 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_osds ([[phab:T428385|T428385]]) * 16:59 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.upgrade_mons (exit_code=0) * 16:35 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_mons ([[phab:T428385|T428385]]) === 2026-06-07 === * 17:55 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services * 17:40 andrewbogott: wmcs.openstack.restart_openstack --cluster-name eqiad1 --all as step one in troubleshooting [[phab:T428312|T428312]] * 17:39 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services === 2026-06-02 === * 20:31 bd808: `kubectl sudo delete cm -n tool-arb-bot maintain-kubeusers-arb-bot` to trigger regeneration of .kube/config (IRC request) === 2026-05-26 === * 19:11 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.roll_reboot_osds (exit_code=0) * 18:50 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_osds * 18:47 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.roll_reboot_mons (exit_code=97) * 18:46 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_mons * 18:33 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.roll_reboot_mons (exit_code=0) * 18:22 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_mons * 18:21 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.roll_reboot_mons (exit_code=97) * 18:21 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_mons * 18:16 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.roll_reboot_mons (exit_code=97) * 18:16 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_mons * 16:30 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=97) on hosts matched by 'P<nowiki>{</nowiki>P:openstack::codfw1dev::nova::compute::service<nowiki>}</nowiki> AND NOT P<nowiki>{</nowiki>F:kernelversion = 6.12.88<nowiki>}</nowiki>' * 16:30 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>P:openstack::codfw1dev::nova::compute::service<nowiki>}</nowiki> AND NOT P<nowiki>{</nowiki>F:kernelversion = 6.12.88<nowiki>}</nowiki>' === 2026-05-25 === * 09:04 godog: move designate eqiad to zk backend === 2026-05-21 === * 18:38 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.roll_reboot_mons (exit_code=0) * 18:27 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_mons * 18:17 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.roll_reboot_mons (exit_code=0) * 18:06 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_mons * 17:41 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.roll_reboot_mons (exit_code=0) * 17:30 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_mons * 14:39 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.roll_reboot_mons (exit_code=0) * 14:29 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_mons * 14:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.roll_reboot_mons (exit_code=0) * 14:17 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_mons * 14:17 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.roll_reboot_mons (exit_code=97) * 14:16 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_mons * 13:41 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.roll_reboot_mons (exit_code=0) ([[phab:T426563|T426563]]) * 13:28 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_mons ([[phab:T426563|T426563]]) * 11:59 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudlb.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudlb2002-dev.codfw.wmnet<nowiki>}</nowiki>' * 11:56 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudlb.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudlb2002-dev.codfw.wmnet<nowiki>}</nowiki>' === 2026-05-20 === * 13:23 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::virt_ceph<nowiki>}</nowiki> AND NOT P<nowiki>{</nowiki>F:kernelversion = 6.12.88<nowiki>}</nowiki>' * 12:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::virt_ceph<nowiki>}</nowiki> AND NOT P<nowiki>{</nowiki>F:kernelversion = 6.12.88<nowiki>}</nowiki>' * 12:42 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::virt_ceph<nowiki>}</nowiki> AND NOT P<nowiki>{</nowiki>F:kernelversion = 6.12.88<nowiki>}</nowiki>' * 12:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::virt_ceph<nowiki>}</nowiki> AND NOT P<nowiki>{</nowiki>F:kernelversion = 6.12.88<nowiki>}</nowiki>' * 12:41 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::virt_ceph<nowiki>}</nowiki> AND NOT P<nowiki>{</nowiki>F:kernelversion = 6.12.88<nowiki>}</nowiki>' * 11:56 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudlb.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudlb2002-dev.codfw.wmnet<nowiki>}</nowiki>' * 11:49 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudlb.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudlb2002-dev.codfw.wmnet<nowiki>}</nowiki>' * 11:28 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudlb.safe_reboot (exit_code=0) on hosts matched by 'A:cloudlb AND A:eqiad' * 11:16 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudlb.safe_reboot on hosts matched by 'A:cloudlb AND A:eqiad' * 11:15 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudlb.safe_reboot (exit_code=0) on hosts matched by 'A:cloudlb AND A:codfw' * 10:56 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudlb.safe_reboot on hosts matched by 'A:cloudlb AND A:codfw' * 10:53 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudlb.safe_reboot (exit_code=99) on hosts matched by 'A:cloudlb AND A:codfw' * 10:46 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudlb.safe_reboot on hosts matched by 'A:cloudlb AND A:codfw' * 10:46 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudlb.safe_reboot (exit_code=99) on hosts matched by 'A:cloudlb AND A:CODFW' * 10:46 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudlb.safe_reboot on hosts matched by 'A:cloudlb AND A:CODFW' * 01:33 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::virt_ceph<nowiki>}</nowiki> AND NOT P<nowiki>{</nowiki>F:kernelversion = 6.12.88<nowiki>}</nowiki>' * 01:26 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::virt_ceph<nowiki>}</nowiki> AND NOT P<nowiki>{</nowiki>F:kernelversion = 6.12.88<nowiki>}</nowiki>' === 2026-05-19 === * 22:56 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'P<nowiki>{</nowiki>P:openstack::codfw1dev::nova::compute::service<nowiki>}</nowiki> AND NOT P<nowiki>{</nowiki>F:kernelversion = 6.12.88<nowiki>}</nowiki>' * 22:09 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>P:openstack::codfw1dev::nova::compute::service<nowiki>}</nowiki> AND NOT P<nowiki>{</nowiki>F:kernelversion = 6.12.88<nowiki>}</nowiki>' * 22:03 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.reboot_node (exit_code=0) on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::codfw1dev::control<nowiki>}</nowiki>' * 21:51 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.reboot_node on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::codfw1dev::control<nowiki>}</nowiki>' * 18:48 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::virt_ceph<nowiki>}</nowiki> AND NOT P<nowiki>{</nowiki>F:kernelversion = 6.12.88<nowiki>}</nowiki>' * 18:42 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::virt_ceph<nowiki>}</nowiki> AND NOT P<nowiki>{</nowiki>F:kernelversion = 6.12.88<nowiki>}</nowiki>' * 17:58 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::virt_ceph<nowiki>}</nowiki> AND NOT P<nowiki>{</nowiki>F:kernelversion = 6.12.88<nowiki>}</nowiki>' * 17:58 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=97) on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::virt_ceph<nowiki>}</nowiki> AND NOT P<nowiki>{</nowiki>F:kernelversion = 6.12.88<nowiki>}</nowiki>' * 17:58 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::virt_ceph<nowiki>}</nowiki> AND NOT P<nowiki>{</nowiki>F:kernelversion = 6.12.88<nowiki>}</nowiki>' * 11:57 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudweb.safe_reboot (exit_code=0) on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::cloudweb<nowiki>}</nowiki>' * 11:52 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudweb.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::cloudweb<nowiki>}</nowiki>' * 11:51 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudweb.safe_reboot (exit_code=0) on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::codfw1dev::cloudweb<nowiki>}</nowiki>' * 11:48 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudweb.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::codfw1dev::cloudweb<nowiki>}</nowiki>' * 11:13 volans@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1043.eqiad.wmnet<nowiki>}</nowiki>' * 10:57 volans@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1043.eqiad.wmnet<nowiki>}</nowiki>' * 10:43 volans@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1042.eqiad.wmnet<nowiki>}</nowiki>' * 10:22 volans@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1042.eqiad.wmnet<nowiki>}</nowiki>' * 10:21 volans@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1041.eqiad.wmnet<nowiki>}</nowiki>' * 10:06 volans@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1041.eqiad.wmnet<nowiki>}</nowiki>' * 09:55 volans@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1040.eqiad.wmnet<nowiki>}</nowiki>' * 09:37 volans@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1040.eqiad.wmnet<nowiki>}</nowiki>' * 02:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.roll_reboot_osds (exit_code=0) ([[phab:T426563|T426563]]) * 02:28 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.reboot_node (exit_code=99) * 02:28 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.reboot_node * 00:11 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_osds ([[phab:T426563|T426563]]) * 00:11 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.roll_reboot_osds (exit_code=99) ([[phab:T426563|T426563]]) * 00:11 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_osds ([[phab:T426563|T426563]]) === 2026-05-18 === * 23:39 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.roll_reboot_osds (exit_code=99) ([[phab:T426563|T426563]]) * 20:35 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_osds ([[phab:T426563|T426563]]) * 20:34 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.roll_reboot_osds (exit_code=99) ([[phab:T426563|T426563]]) * 20:34 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_osds ([[phab:T426563|T426563]]) * 20:33 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.roll_reboot_osds (exit_code=99) ([[phab:T426563|T426563]]) * 20:33 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_osds ([[phab:T426563|T426563]]) * 12:01 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/313 * 12:01 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/313 * 11:59 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 11:59 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 10:47 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/312 * 10:46 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/312 * 10:46 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/312 * 10:45 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/312 === 2026-05-11 === * 13:47 andrewbogott: restarting eqiad1 keystone services to pick up some logging changes in wmfkeystonehooks === 2026-05-06 === * 13:34 godog: change cloud-vps quota request phab at https://phabricator.wikimedia.org/project/manage/2880/ to mention https://cloudvps-quota.toolforge.org === 2026-05-05 === * 10:20 taavi: taavi@cloudcontrol1007 ~ $ sudo wmcs-enc-cli --openstack-project admin delete_project wikilabels # [[phab:T416588|T416588]] === 2026-05-02 === * 12:12 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 12:11 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2026-04-29 === * 16:05 andrewbogott: cleaning up stray broken osbpo references on VMs, e.g. /etc/apt/sources.list.d/openstack-dalmatian-bookworm.sources === 2026-04-28 === * 14:33 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,designate * 14:33 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,designate * 08:42 wmbot~godog@r5: END (PASS) - Cookbook wmcs.openstack.rack_resources (exit_code=0) on cluster 'codfw1dev' * 08:42 wmbot~godog@r5: START - Cookbook wmcs.openstack.rack_resources on cluster 'codfw1dev' * 08:32 wmbot~godog@r5: END (PASS) - Cookbook wmcs.openstack.rack_resources (exit_code=0) on cluster 'codfw1dev' * 08:32 wmbot~godog@r5: START - Cookbook wmcs.openstack.rack_resources on cluster 'codfw1dev' * 08:29 wmbot~godog@r5: END (PASS) - Cookbook wmcs.openstack.rack_resources (exit_code=0) on cluster 'codfw1dev' * 08:28 wmbot~godog@r5: START - Cookbook wmcs.openstack.rack_resources on cluster 'codfw1dev' * 08:19 wmbot~godog@r5: END (PASS) - Cookbook wmcs.openstack.rack_resources (exit_code=0) on cluster 'codfw1dev' * 08:19 wmbot~godog@r5: START - Cookbook wmcs.openstack.rack_resources on cluster 'codfw1dev' * 08:18 wmbot~godog@r5: END (PASS) - Cookbook wmcs.openstack.rack_resources (exit_code=0) on cluster 'codfw1dev' * 08:18 wmbot~godog@r5: START - Cookbook wmcs.openstack.rack_resources on cluster 'codfw1dev' === 2026-04-27 === * 12:43 godog: upgrade spicerack on cloudcumin === 2026-04-21 === * 15:36 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 15:33 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 15:30 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.restart_openstack (exit_code=97) on deployment codfw1dev for all services * 15:27 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services === 2026-04-20 === * 10:20 godog: test shutting cloudcontrol2005-dev network port === 2026-04-16 === * 06:43 godog: roll-restart nova-api to pick up changes - [[phab:T423378|T423378]] === 2026-04-15 === * 15:02 godog: deploy per-service oslo.messaging shared memory file name - [[phab:T423378|T423378]] === 2026-04-14 === * 21:48 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services * 21:33 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 21:21 andrewbogott: rebuilding eqiad1 rabbitmq cluster in hopes of getting some more consistent api responses * 20:01 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,nova * 19:56 andrewbogott: restarting all nova services in eqiad1; i'm seeing inconsistent permission failures * 19:56 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,nova * 12:45 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) ([[phab:T419658|T419658]]) * 12:45 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance ([[phab:T419658|T419658]]) * 12:30 godog: set maint on cloudvirt1050 - [[phab:T419658|T419658]] === 2026-04-13 === * 19:30 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.roll_reboot_osds (exit_code=0) * 19:06 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_osds * 18:57 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.roll_reboot_osds (exit_code=99) * 18:11 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_osds * 18:10 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.roll_reboot_osds (exit_code=99) * 18:10 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_osds * 17:52 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.roll_reboot_osds (exit_code=99) * 17:52 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_osds * 17:52 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.roll_reboot_osds (exit_code=99) * 15:17 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_osds * 15:17 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.roll_reboot_osds (exit_code=99) * 15:17 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_osds * 15:17 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.roll_reboot_osds (exit_code=99) * 15:17 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_osds * 15:01 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.roll_reboot_osds (exit_code=99) * 15:01 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_osds * 14:31 godog: grant filippo and volans admin roles in codfw1dev === 2026-04-08 === * 18:30 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services * 18:17 andrewbogott: restarting openstack services in eqiad1 before digging into a tofu failure * 18:17 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 14:03 godog: bounce nova on cloudcontrol1006 * 13:48 godog: stop designate and memcached on all cloudcontrol1* * 13:35 godog: leave designate processes up only on cloudcontrol1006 to ease debugging - [[phab:T422646|T422646]] * 11:59 filippo@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment eqiad1 for service: project,designate * 11:58 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,designate * 11:54 filippo@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment eqiad1 for service: project,designate * 11:53 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,designate * 09:45 godog: bounce designate on cloudcontrol * 08:14 godog: perform more network tests on cloudrabbit1001 - [[phab:T417393|T417393]] * 07:57 godog: unshut cloudcontrol1011 network interface - [[phab:T417393|T417393]] * 07:36 godog: shut cloudcontrol1011 network interface - [[phab:T417393|T417393]] * 07:26 godog: unshut cloudrabbit1001 network interface - [[phab:T417393|T417393]] * 07:00 godog: test shutting cloudrabbit1001 network interface - [[phab:T417393|T417393]] === 2026-04-07 === * 13:40 andrewbogott: upgrading spicerack on cloudcumin1001 === 2026-04-06 === * 13:59 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 13:58 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 13:56 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 13:55 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 13:54 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 13:54 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 13:53 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 13:53 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2026-04-03 === * 10:18 godog: move codfw neutron l3-agent queues to quorum - [[phab:T421054|T421054]] === 2026-04-02 === * 12:45 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 12:45 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:44 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 12:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:44 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 12:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 12:36 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 12:35 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:35 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 12:35 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 12:34 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 12:34 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 11:03 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 11:02 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 11:01 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/304 * 11:01 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/304 === 2026-04-01 === * 14:45 volans: Installed cumin v6.0.0-1 on apt.w.o (unattended upgrades) and the cloudcumin hosts (previous one for rollback in my home) * 12:24 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/304 * 12:23 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/304 * 08:15 godog: extend cloudrabbit1* root with an additional 100G - [[phab:T421054|T421054]] === 2026-03-31 === * 07:26 godog: neutron maint done - [[phab:T421054|T421054]] * 07:12 godog: start maint on neutron - [[phab:T421054|T421054]] === 2026-03-30 === * 13:52 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 13:48 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 13:45 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.rabbitmq.rebuild_rabbit_cluster (exit_code=0) on deployment codfw1dev * 13:42 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.rabbitmq.rebuild_rabbit_cluster on deployment codfw1dev * 11:53 godog: bounce neutron-l3-agent on cloudnet1005 * 11:20 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services * 11:06 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 11:01 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.rabbitmq.rebuild_rabbit_cluster (exit_code=0) on deployment eqiad1 * 10:57 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.rabbitmq.rebuild_rabbit_cluster on deployment eqiad1 * 10:36 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services * 10:21 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 09:58 filippo@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.rabbitmq.rebuild_rabbit_cluster (exit_code=99) on deployment eqiad1 * 09:57 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.rabbitmq.rebuild_rabbit_cluster on deployment eqiad1 * 09:42 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,designate * 09:41 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,designate * 09:30 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,nova,neutron,designate * 09:30 filippo@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment eqiad1 for service: project,designate * 09:28 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,designate * 09:18 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,nova,neutron,designate * 09:18 filippo@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.restart_openstack (exit_code=97) on deployment eqiad1 for service: project,neutron,designate * 09:15 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 09:14 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,neutron,designate * 09:14 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,heat * 09:13 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,heat * 09:12 filippo@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment eqiad1 for all services * 09:11 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 09:11 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services === 2026-03-26 === * 22:47 andrewbogott: I am logging to the admin log * 13:38 root@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 13:33 root@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services === 2026-03-25 === * 17:21 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 17:18 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 17:17 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 17:16 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 17:03 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 17:02 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 16:22 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for service: project,heat * 16:21 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for service: project,heat === 2026-03-24 === * 10:39 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for service: project,nova,glance,keystone,cinder,neutron,trove,magnum,octavia,heat,swift * 10:35 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for service: project,nova,glance,keystone,cinder,neutron,trove,magnum,octavia,heat,swift * 10:34 filippo@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 10:32 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 10:31 filippo@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 10:30 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services === 2026-03-23 === * 20:50 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=99) on host 'cloudvirt1076.eqiad.wmnet' * 20:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1076.eqiad.wmnet' * 20:43 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=99) on host 'cloudvirtlocal1076.eqiad.wmnet' * 20:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirtlocal1076.eqiad.wmnet' * 20:42 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=99) on host 'cloudvirtlocal1076.eqiad.wmnet' * 20:42 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirtlocal1076.eqiad.wmnet' * 20:42 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=99) on host 'cloudvirtlocal1076' * 20:42 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirtlocal1076' * 16:24 dhinus: added komla to https://gitlab.wikimedia.org/groups/repos/cloud/-/group_members [[phab:T420532|T420532]] * 15:02 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services * 14:48 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 14:00 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services * 13:47 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 13:17 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,designate * 13:16 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,designate * 13:16 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,heat * 13:15 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,heat * 12:43 godog: apply https://gerrit.wikimedia.org/r/c/operations/puppet/+/1254877 to cloudrabbit eqiad - [[phab:T418444|T418444]] * 10:51 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 10:47 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 10:47 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for service: project * 10:47 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for service: project * 10:33 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.rabbitmq.rebuild_rabbit_cluster (exit_code=0) on deployment codfw1dev * 10:29 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.rabbitmq.rebuild_rabbit_cluster on deployment codfw1dev * 10:12 godog: apply https://gerrit.wikimedia.org/r/c/operations/puppet/+/1254877 to cloudrabbit codfw - [[phab:T418444|T418444]] === 2026-03-19 === * 16:50 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_osds ([[phab:T419960|T419960]]) * 16:48 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.roll_reboot_osds (exit_code=0) ([[phab:T419960|T419960]]) * 16:27 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_osds ([[phab:T419960|T419960]]) === 2026-03-18 === * 20:07 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 20:03 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 18:43 andrewbogott: dist-upgrade and rebooting cloudrabbit2xxx-dev nodes * 09:26 godog: end network switch failover tests - [[phab:T417393|T417393]] * 08:58 godog: bounce rabbit on cloudrabbit1001 - [[phab:T417393|T417393]] * 08:55 filippo@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.restart_openstack (exit_code=97) on deployment eqiad1 for all services * 08:51 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 08:00 godog: start network switch failover tests - [[phab:T417393|T417393]] === 2026-03-17 === * 13:33 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirtlocal1003.eqiad.wmnet' ([[phab:T406516|T406516]]) * 13:25 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirtlocal1003.eqiad.wmnet' ([[phab:T406516|T406516]]) * 13:23 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirtlocal1002.eqiad.wmnet' ([[phab:T406516|T406516]]) * 13:16 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirtlocal1002.eqiad.wmnet' ([[phab:T406516|T406516]]) * 13:05 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirtlocal1001.eqiad.wmnet' ([[phab:T406516|T406516]]) * 12:57 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirtlocal1001.eqiad.wmnet' ([[phab:T406516|T406516]]) * 01:25 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1049.eqiad.wmnet' ([[phab:T406516|T406516]]) * 01:18 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1049.eqiad.wmnet' ([[phab:T406516|T406516]]) * 01:17 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1048.eqiad.wmnet' ([[phab:T406516|T406516]]) * 01:11 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1048.eqiad.wmnet' ([[phab:T406516|T406516]]) * 01:11 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1047.eqiad.wmnet' ([[phab:T406516|T406516]]) * 01:04 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1047.eqiad.wmnet' ([[phab:T406516|T406516]]) * 01:04 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1046.eqiad.wmnet' ([[phab:T406516|T406516]]) * 00:57 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1046.eqiad.wmnet' ([[phab:T406516|T406516]]) * 00:57 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1045.eqiad.wmnet' ([[phab:T406516|T406516]]) * 00:51 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1045.eqiad.wmnet' ([[phab:T406516|T406516]]) * 00:50 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1044.eqiad.wmnet' ([[phab:T406516|T406516]]) * 00:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1044.eqiad.wmnet' ([[phab:T406516|T406516]]) * 00:44 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1043.eqiad.wmnet' ([[phab:T406516|T406516]]) * 00:37 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1043.eqiad.wmnet' ([[phab:T406516|T406516]]) * 00:37 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1042.eqiad.wmnet' ([[phab:T406516|T406516]]) * 00:30 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1042.eqiad.wmnet' ([[phab:T406516|T406516]]) * 00:30 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1041.eqiad.wmnet' ([[phab:T406516|T406516]]) * 00:23 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1041.eqiad.wmnet' ([[phab:T406516|T406516]]) * 00:23 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1040.eqiad.wmnet' ([[phab:T406516|T406516]]) * 00:17 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1040.eqiad.wmnet' ([[phab:T406516|T406516]]) === 2026-03-16 === * 23:07 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1059.eqiad.wmnet' ([[phab:T406516|T406516]]) * 23:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1059.eqiad.wmnet' ([[phab:T406516|T406516]]) * 23:00 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1058.eqiad.wmnet' ([[phab:T406516|T406516]]) * 22:53 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1058.eqiad.wmnet' ([[phab:T406516|T406516]]) * 22:53 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1057.eqiad.wmnet' ([[phab:T406516|T406516]]) * 22:46 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1057.eqiad.wmnet' ([[phab:T406516|T406516]]) * 22:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1056.eqiad.wmnet' ([[phab:T406516|T406516]]) * 22:39 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1056.eqiad.wmnet' ([[phab:T406516|T406516]]) * 22:39 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1055.eqiad.wmnet' ([[phab:T406516|T406516]]) * 22:32 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1055.eqiad.wmnet' ([[phab:T406516|T406516]]) * 22:32 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1054.eqiad.wmnet' ([[phab:T406516|T406516]]) * 22:25 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1054.eqiad.wmnet' ([[phab:T406516|T406516]]) * 22:25 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1053.eqiad.wmnet' ([[phab:T406516|T406516]]) * 22:19 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1053.eqiad.wmnet' ([[phab:T406516|T406516]]) * 22:18 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1052.eqiad.wmnet' ([[phab:T406516|T406516]]) * 22:12 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1052.eqiad.wmnet' ([[phab:T406516|T406516]]) * 22:12 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1051.eqiad.wmnet' ([[phab:T406516|T406516]]) * 22:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1051.eqiad.wmnet' ([[phab:T406516|T406516]]) * 22:05 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1050.eqiad.wmnet' ([[phab:T406516|T406516]]) * 21:58 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1050.eqiad.wmnet' ([[phab:T406516|T406516]]) * 21:33 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1069.eqiad.wmnet' ([[phab:T406516|T406516]]) * 21:26 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1069.eqiad.wmnet' ([[phab:T406516|T406516]]) * 21:26 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1068.eqiad.wmnet' ([[phab:T406516|T406516]]) * 21:19 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1068.eqiad.wmnet' ([[phab:T406516|T406516]]) * 21:19 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1067.eqiad.wmnet' ([[phab:T406516|T406516]]) * 21:12 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1067.eqiad.wmnet' ([[phab:T406516|T406516]]) * 21:12 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1066.eqiad.wmnet' ([[phab:T406516|T406516]]) * 21:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1066.eqiad.wmnet' ([[phab:T406516|T406516]]) * 21:05 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1065.eqiad.wmnet' ([[phab:T406516|T406516]]) * 20:58 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1065.eqiad.wmnet' ([[phab:T406516|T406516]]) * 20:58 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1064.eqiad.wmnet' ([[phab:T406516|T406516]]) * 20:51 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1064.eqiad.wmnet' ([[phab:T406516|T406516]]) * 20:51 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1063.eqiad.wmnet' ([[phab:T406516|T406516]]) * 20:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1063.eqiad.wmnet' ([[phab:T406516|T406516]]) * 20:44 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1062.eqiad.wmnet' ([[phab:T406516|T406516]]) * 20:37 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1062.eqiad.wmnet' ([[phab:T406516|T406516]]) * 20:37 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1061.eqiad.wmnet' ([[phab:T406516|T406516]]) * 20:30 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1061.eqiad.wmnet' ([[phab:T406516|T406516]]) * 20:30 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1060.eqiad.wmnet' ([[phab:T406516|T406516]]) * 20:22 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1060.eqiad.wmnet' ([[phab:T406516|T406516]]) * 20:21 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1070.eqiad.wmnet' ([[phab:T406516|T406516]]) * 20:14 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1070.eqiad.wmnet' ([[phab:T406516|T406516]]) * 20:14 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1071.eqiad.wmnet' ([[phab:T406516|T406516]]) * 20:07 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1071.eqiad.wmnet' ([[phab:T406516|T406516]]) * 20:07 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1072.eqiad.wmnet' ([[phab:T406516|T406516]]) * 19:59 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1072.eqiad.wmnet' ([[phab:T406516|T406516]]) * 19:59 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1073.eqiad.wmnet' ([[phab:T406516|T406516]]) * 19:51 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1073.eqiad.wmnet' ([[phab:T406516|T406516]]) * 19:51 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1074.eqiad.wmnet' ([[phab:T406516|T406516]]) * 19:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1074.eqiad.wmnet' ([[phab:T406516|T406516]]) * 19:43 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1075.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 19:42 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1075.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 19:41 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1075.eqiad.wmnet' ([[phab:T406516|T406516]]) * 19:33 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1075.eqiad.wmnet' ([[phab:T406516|T406516]]) * 19:33 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=99) on host 'cloudvir1075.eqiad.wmnet' ([[phab:T406516|T406516]]) * 19:33 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvir1075.eqiad.wmnet' ([[phab:T406516|T406516]]) * 19:30 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudweb.unset_maintenance (exit_code=99) ([[phab:T406516|T406516]]) * 19:30 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudweb.unset_maintenance ([[phab:T406516|T406516]]) * 19:29 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudnet1005.eqiad.wmnet' ([[phab:T406516|T406516]]) * 19:18 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet1005.eqiad.wmnet' ([[phab:T406516|T406516]]) * 19:10 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudnet1006.eqiad.wmnet' ([[phab:T406516|T406516]]) * 18:59 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet1006.eqiad.wmnet' ([[phab:T406516|T406516]]) * 18:52 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol1011.eqiad.wmnet' ([[phab:T406516|T406516]]) * 18:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1011.eqiad.wmnet' ([[phab:T406516|T406516]]) * 18:40 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudcontrol1011.eqiad.wmnet' ([[phab:T406516|T406516]]) * 18:26 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1011.eqiad.wmnet' ([[phab:T406516|T406516]]) * 18:25 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol1007.eqiad.wmnet' ([[phab:T406516|T406516]]) * 18:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1007.eqiad.wmnet' ([[phab:T406516|T406516]]) * 18:01 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol1006.eqiad.wmnet' ([[phab:T406516|T406516]]) * 17:42 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1006.eqiad.wmnet' ([[phab:T406516|T406516]]) * 17:42 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudweb.set_maintenance (exit_code=99) ([[phab:T406516|T406516]]) * 17:40 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudweb.set_maintenance ([[phab:T406516|T406516]]) * 17:14 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudservices1006.eqiad.wmnet' ([[phab:T406516|T406516]]) * 17:04 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1006.eqiad.wmnet' ([[phab:T406516|T406516]]) * 17:04 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T406516|T406516]]) * 16:55 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T406516|T406516]]) * 05:36 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1040.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 05:14 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1040.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 05:14 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1041.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 04:52 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1041.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 04:52 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1042.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 04:31 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1042.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 04:31 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1043.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 04:09 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1043.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 04:09 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1044.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 03:42 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1044.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 03:42 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1045.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 03:17 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1045.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 03:17 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1046.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 02:56 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1046.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 02:56 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1047.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 02:30 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1047.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 02:30 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1048.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 02:06 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1048.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 02:06 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1049.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 01:54 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1049.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 01:54 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,nova * 01:47 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,nova * 01:46 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1049.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 01:45 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1049.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 01:44 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1058.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 01:41 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 01:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 01:40 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 01:40 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 01:40 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 01:40 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 01:40 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 01:40 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 01:40 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=99) * 01:40 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 01:39 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 01:39 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 01:39 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 01:39 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 01:39 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 01:39 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 01:39 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 01:38 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 01:38 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 01:38 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 01:38 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 01:38 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 01:35 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1058.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) === 2026-03-15 === * 03:46 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1049.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 03:27 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1049.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 03:27 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1050.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 02:42 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1050.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 02:42 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1051.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 02:20 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1051.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 02:20 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1052.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) === 2026-03-14 === * 02:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1052.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 02:41 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1053.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 02:13 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1053.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 02:13 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1054.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 01:47 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1054.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 01:47 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1055.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 01:29 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1055.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 01:29 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1056.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) === 2026-03-13 === * 23:22 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1056.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 23:22 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1057.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 23:09 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1057.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 23:09 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1058.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 22:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1058.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 22:44 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1059.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 22:25 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1059.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 22:25 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1060.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 22:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1060.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 22:05 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1061.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 21:40 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1061.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 21:40 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1062.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 21:21 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1062.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 21:21 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1063.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 21:03 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1063.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 21:03 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1064.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 20:47 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1064.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 20:47 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1065.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 20:31 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1065.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 20:31 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1066.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 20:13 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1066.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 20:13 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1067.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 19:34 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1067.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 19:34 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1068.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 19:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2004-dev.codfw.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 17:50 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1068.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 17:50 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1069.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 17:35 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1069.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 17:35 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1070.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 17:18 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1070.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 17:18 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1071.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 17:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1071.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 17:04 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1072.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 16:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2004-dev.codfw.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 16:41 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2005-dev.codfw.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 16:38 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 16:35 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 16:32 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1073.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 16:32 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1074.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 16:24 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2005-dev.codfw.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 16:24 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2006-dev.codfw.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 16:14 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1074.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 16:14 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1075.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 16:10 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2006-dev.codfw.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 16:09 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2006-dev.codfw.wmnet.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 16:07 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2006-dev.codfw.wmnet.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 15:57 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1075.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 15:54 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1076.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 15:40 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1076.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) === 2026-03-12 === * 13:46 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 13:45 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 13:44 filippo@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 13:43 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 11:44 godog: create g4.cores8.ram32.disk20.ephem140 flavor for tools worker === 2026-03-10 === * 17:14 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 17:10 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services === 2026-03-04 === * 14:43 taavi: deploying firewall rule updates: https://gerrit.wikimedia.org/r/c/operations/homer/public/+/970275 === 2026-02-26 === * 03:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services * 03:31 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 03:30 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.rabbitmq.rebuild_rabbit_cluster (exit_code=0) on deployment eqiad1 * 03:27 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.rabbitmq.rebuild_rabbit_cluster on deployment eqiad1 * 03:27 andrewbogott: rebuilding the rabbitmq cluster in eqiad1; many failed messages === 2026-02-24 === * 08:18 godog: clean up stray wikilabels project puppet host certs * 00:16 bd808: Applied LDIF to create analytics-sre user and group ([[phab:T418120|T418120]]) === 2026-02-20 === * 12:00 dhinus: DROP DATABASE toollabs_p; (was used by updatetools.py, see [[phab:T415383|T415383]]) === 2026-02-18 === * 13:29 taavi: rebooting cloudgw1004 for [[phab:T417075|T417075]] fixes === 2026-02-17 === * 12:28 volans: re-enabled puppet on toolforge's nfs k8s workers * 10:39 volans: temporarily disabling puppet on toolforge's nfs k8s workers to test gerrit/1239689 === 2026-02-13 === * 04:25 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 04:21 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 04:20 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.rabbitmq.rebuild_rabbit_cluster (exit_code=0) on deployment codfw1dev * 04:18 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.rabbitmq.rebuild_rabbit_cluster on deployment codfw1dev === 2026-02-12 === * 14:34 andrewbogott: moritz is upgrading dnsmasq on eqiad1 cloudnets and cloudvirts * 10:08 dhinus: delete job "updatetools" in admin tool as it's no longer used ([[phab:T415383|T415383]]) === 2026-02-09 === * 16:03 andrewbogott: rebooting cloudnets in codfw1dev to make sure we've picked up the new dnsmasq version * 12:50 dcaro: removed stall cert cloudinfra-acme-chief-01.novalocal from cloudinfra-internal-puppetserver-1 * 12:49 dcaro: removed the pki* certs from the cloudinfra-cloudvps puppetserver as they are not handled there anymore but the local puppetserver * 09:45 dcaro: re-enabling nrpe2nodexp-ferm_active.service on cloudcumins after upgrade (getting stall promfile) * 03:00 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services * 02:47 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 02:47 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.rabbitmq.rebuild_rabbit_cluster (exit_code=0) on deployment eqiad1 * 02:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.rabbitmq.rebuild_rabbit_cluster on deployment eqiad1 === 2026-02-05 === * 21:12 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 21:07 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 21:03 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.rabbitmq.rebuild_rabbit_cluster (exit_code=0) on deployment codfw1dev * 21:01 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.rabbitmq.rebuild_rabbit_cluster on deployment codfw1dev * 20:58 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.rabbitmq.rebuild_rabbit_cluster (exit_code=99) on deployment codfw1dev * 20:57 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.rabbitmq.rebuild_rabbit_cluster on deployment codfw1dev * 20:56 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.rabbitmq.rebuild_rabbit_cluster (exit_code=99) on deployment codfw1dev * 20:55 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.rabbitmq.rebuild_rabbit_cluster on deployment codfw1dev * 20:49 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.rabbitmq.rebuild_rabbit_cluster (exit_code=99) on deployment codfw1dev * 20:48 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.rabbitmq.rebuild_rabbit_cluster on deployment codfw1dev * 20:47 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.rabbitmq.rebuild_rabbit_cluster (exit_code=99) on deployment codfw1dev * 20:46 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.rabbitmq.rebuild_rabbit_cluster on deployment codfw1dev * 20:46 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.rabbitmq.rebuild_rabbit_cluster (exit_code=99) on deployment codfw1dev * 20:46 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.rabbitmq.rebuild_rabbit_cluster on deployment codfw1dev * 20:45 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.rabbitmq.rebuild_rabbit_cluster (exit_code=99) on deployment codfw1dev * 20:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.rabbitmq.rebuild_rabbit_cluster on deployment codfw1dev * 20:43 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.rabbitmq.rebuild_rabbit_cluster (exit_code=99) on deployment codfw1dev * 20:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.rabbitmq.rebuild_rabbit_cluster on deployment codfw1dev === 2026-02-03 === * 18:02 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for service: project,designate * 18:01 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for service: project,designate === 2026-01-24 === * 17:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services * 17:14 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 17:09 andrewbogott: preemptively rebuilding rabbitmq for eqiad1; message flakiness === 2026-01-23 === * 13:02 taavi: switch https://gitlab.wikimedia.org/repos/cloud/wmcs/utils to fast-forward only merging mode === 2026-01-22 === * 18:33 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt2006-dev.codfw.wmnet' * 18:27 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt2006-dev.codfw.wmnet' * 18:08 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt2005-dev.codfw.wmnet' * 18:01 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt2005-dev.codfw.wmnet' * 17:40 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt2004-dev.codfw.wmnet' * 17:33 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt2004-dev.codfw.wmnet' * 17:14 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudnet2005-dev.codfw.wmnet' * 17:04 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet2005-dev.codfw.wmnet' * 17:03 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=97) on host 'cloudnet2006-dev.codfw.wmnet' * 17:03 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet2006-dev.codfw.wmnet' * 15:56 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudnet2006-dev.codfw.wmnet' * 15:47 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet2006-dev.codfw.wmnet' * 14:33 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 14:30 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 14:25 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudservices2004-dev.codfw.wmnet' * 14:17 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices2004-dev.codfw.wmnet' * 14:11 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudservices2005-dev.codfw.wmnet' * 14:03 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices2005-dev.codfw.wmnet' * 02:39 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol2010-dev.codfw.wmnet' * 02:27 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2010-dev.codfw.wmnet' * 01:49 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol2010-dev.codfw.wmnet' * 01:39 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2010-dev.codfw.wmnet' * 00:11 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol2010-dev.codfw.wmnet' === 2026-01-21 === * 23:54 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2010-dev.codfw.wmnet' * 23:47 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol2006-dev.codfw.wmnet' * 23:31 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2006-dev.codfw.wmnet' * 23:23 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudcontrol2006-dev.codfw.wmnet' * 23:18 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2006-dev.codfw.wmnet' * 23:18 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=97) on host 'cloudcontrol2006-dev.codfw.wmnet' (Txxxxxx) * 23:18 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2006-dev.codfw.wmnet' (Txxxxxx) * 23:03 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudcontrol2006-dev.codfw.wmnet' (Txxxxxx) * 22:56 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2006-dev.codfw.wmnet' (Txxxxxx) * 22:50 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol2005-dev.codfw.wmnet' (Txxxxxx) * 22:34 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2005-dev.codfw.wmnet' (Txxxxxx) === 2026-01-14 === * 15:43 andrewbogott: reimaging cloudgw2003-dev to Trixie * 13:37 andrewbogott: reimaging cloudgw2002-dev to Trixie === 2026-01-12 === * 18:09 andrewbogott: stopping bird on cloudlb1001, reimaging to Trixie * 15:54 andrewbogott: stopping bird on cloudlb1002, reimaging to Trixie * 13:22 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 13:22 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch === 2026-01-10 === * 20:56 andrewbogott: in codfw1dev: openstack role add --project swift --user swift member * 20:56 andrewbogott: in codfw1dev: openstack role add --project swift --user swift admin * 19:37 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 19:34 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 19:29 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 19:28 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services === 2026-01-07 === * 10:00 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for main branch * 10:00 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch === 2026-01-06 === * 02:22 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services * 02:07 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services === 2026-01-05 === * 22:49 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,nova * 22:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,nova === 2025-12-26 === * 23:32 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 23:29 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services === 2025-12-20 === * 03:10 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services * 03:01 andrewbogott: restarting eqiad1 openstack services in response to some fullstack failures * 02:56 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 02:27 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,heat * 02:27 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,heat === 2025-12-19 === * 16:18 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch ([[phab:T412865|T412865]]) * 16:17 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch ([[phab:T412865|T412865]]) === 2025-12-18 === * 11:04 godog: bump tools object quota to 500G === 2025-12-16 === * 18:51 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 18:51 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 18:50 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/286 * 18:50 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/286 * 18:48 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/286 * 18:47 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/286 * 13:18 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.reboot_node (exit_code=0) * 13:03 filippo@cloudcumin1001: START - Cookbook wmcs.ceph.reboot_node === 2025-12-05 === * 15:14 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 15:10 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services === 2025-12-03 === * 10:55 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 10:55 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 10:48 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/285 * 10:48 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/285 * 05:52 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services * 05:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 00:19 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services * 00:07 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services === 2025-12-02 === * 23:21 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services * 23:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 23:04 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.restart_openstack (exit_code=97) on deployment eqiad1 for all services * 23:04 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 23:03 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment eqiad1 for all services * 23:03 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 20:53 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1061.eqiad.wmnet' * 20:52 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 1797e3a3-f04a-4cb0-9102-{{Gerrit|79f1d0079d57}} (cluster eqiad1) * 20:51 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1061.eqiad.wmnet' * 20:51 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 1797e3a3-f04a-4cb0-9102-{{Gerrit|79f1d0079d57}} (cluster eqiad1) * 20:51 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 592754fe-dd64-463f-ab33-{{Gerrit|d51a4108cec0}} (cluster eqiad1) * 20:50 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 592754fe-dd64-463f-ab33-{{Gerrit|d51a4108cec0}} (cluster eqiad1) * 20:47 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1061.eqiad.wmnet' * 20:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1061.eqiad.wmnet' * 20:32 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1061.eqiad.wmnet' * 20:29 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1061.eqiad.wmnet' * 20:26 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1054.eqiad.wmnet' * 20:25 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1054.eqiad.wmnet' * 20:23 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm c6645048-8447-4553-bf13-{{Gerrit|8122f959e4a8}} (cluster eqiad1) * 20:23 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm c6645048-8447-4553-bf13-{{Gerrit|8122f959e4a8}} (cluster eqiad1) * 20:18 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1054.eqiad.wmnet' * 20:17 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1054.eqiad.wmnet' * 20:16 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1054.eqiad.wmnet' * 20:13 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1054.eqiad.wmnet' * 20:10 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1050.eqiad.wmnet' * 20:09 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1050.eqiad.wmnet' * 20:07 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1046.eqiad.wmnet' * 20:06 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1046.eqiad.wmnet' * 20:03 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1046.eqiad.wmnet' * 20:01 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1046.eqiad.wmnet' * 19:58 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1046.eqiad.wmnet' * 19:56 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1046.eqiad.wmnet' * 19:55 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1046.eqiad.wmnet' * 19:52 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1046.eqiad.wmnet' * 17:05 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1041.eqiad.wmnet<nowiki>}</nowiki>' * 17:02 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1042.eqiad.wmnet<nowiki>}</nowiki>' * 16:58 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1042.eqiad.wmnet<nowiki>}</nowiki>' * 16:58 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1042.eqiad.wmnet<nowiki>}</nowiki>' * 16:52 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 1da6f8f7-db35-4f33-92f9-{{Gerrit|29a6516bf47c}} (cluster eqiad1) * 16:52 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1042.eqiad.wmnet<nowiki>}</nowiki>' * 16:52 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 1da6f8f7-db35-4f33-92f9-{{Gerrit|29a6516bf47c}} (cluster eqiad1) * 16:52 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 3f0dc3e0-f5e8-43a4-86dc-{{Gerrit|523ad08e90e6}} (cluster eqiad1) * 16:51 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 3f0dc3e0-f5e8-43a4-86dc-{{Gerrit|523ad08e90e6}} (cluster eqiad1) * 16:51 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 848de230-1687-40c5-b954-{{Gerrit|f8c2a3b7a443}} (cluster eqiad1) * 16:50 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1041.eqiad.wmnet<nowiki>}</nowiki>' * 16:50 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 848de230-1687-40c5-b954-{{Gerrit|f8c2a3b7a443}} (cluster eqiad1) * 16:50 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 9491619f-43b5-4612-b976-{{Gerrit|00862dcd901d}} (cluster eqiad1) * 16:49 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 9491619f-43b5-4612-b976-{{Gerrit|00862dcd901d}} (cluster eqiad1) * 16:49 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 59f99bab-8a86-4701-a142-{{Gerrit|3a15a1c18d48}} (cluster eqiad1) * 16:48 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 59f99bab-8a86-4701-a142-{{Gerrit|3a15a1c18d48}} (cluster eqiad1) * 16:48 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm afb538cb-a128-450b-a02f-{{Gerrit|4fee25183588}} (cluster eqiad1) * 16:47 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm afb538cb-a128-450b-a02f-{{Gerrit|4fee25183588}} (cluster eqiad1) * 16:47 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 8b546fd2-137d-4b91-86f3-{{Gerrit|b50fa515c98c}} (cluster eqiad1) * 16:46 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 8b546fd2-137d-4b91-86f3-{{Gerrit|b50fa515c98c}} (cluster eqiad1) * 16:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm c7b1311c-ee8b-4118-b907-{{Gerrit|ad0382644350}} (cluster eqiad1) * 16:46 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm c7b1311c-ee8b-4118-b907-{{Gerrit|ad0382644350}} (cluster eqiad1) * 16:45 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 60d1c96e-0c3c-47e1-86d6-{{Gerrit|cd30527d5066}} (cluster eqiad1) * 16:45 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1040.eqiad.wmnet<nowiki>}</nowiki>' * 16:45 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 60d1c96e-0c3c-47e1-86d6-{{Gerrit|cd30527d5066}} (cluster eqiad1) * 16:45 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 028c29da-adcb-4239-bcb4-{{Gerrit|6e80516e6fbb}} (cluster eqiad1) * 16:44 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 028c29da-adcb-4239-bcb4-{{Gerrit|6e80516e6fbb}} (cluster eqiad1) * 16:40 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1040.eqiad.wmnet<nowiki>}</nowiki>' * 16:34 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm e6796dbd-2511-4bf6-bdee-{{Gerrit|4a14a7414d5f}} (cluster eqiad1) * 16:33 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 5e5c3bad-f1c7-49e5-b846-{{Gerrit|edaf111af83c}} (cluster eqiad1) * 16:33 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm e6796dbd-2511-4bf6-bdee-{{Gerrit|4a14a7414d5f}} (cluster eqiad1) * 16:33 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 5e5c3bad-f1c7-49e5-b846-{{Gerrit|edaf111af83c}} (cluster eqiad1) * 16:32 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 1d74fc9a-0ddd-41d6-a0fd-{{Gerrit|5bba5e455c32}} (cluster eqiad1) * 16:32 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 55ed5d49-43db-4f62-8c40-{{Gerrit|5cb0431dfce2}} (cluster eqiad1) * 16:31 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 1d74fc9a-0ddd-41d6-a0fd-{{Gerrit|5bba5e455c32}} (cluster eqiad1) * 16:31 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 55ed5d49-43db-4f62-8c40-{{Gerrit|5cb0431dfce2}} (cluster eqiad1) * 15:48 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 8030caca-e1e8-4f1d-bce1-{{Gerrit|04afd22adb3a}} (cluster eqiad1) * 15:48 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 8030caca-e1e8-4f1d-bce1-{{Gerrit|04afd22adb3a}} (cluster eqiad1) * 15:48 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 9ccf684d-c6ea-45ee-83db-{{Gerrit|ee3af5de3dfe}} (cluster eqiad1) * 15:47 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 9ccf684d-c6ea-45ee-83db-{{Gerrit|ee3af5de3dfe}} (cluster eqiad1) * 15:47 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 02bf16d5-5e10-470b-b05d-{{Gerrit|341673a284de}} (cluster eqiad1) * 15:47 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 02bf16d5-5e10-470b-b05d-{{Gerrit|341673a284de}} (cluster eqiad1) * 15:47 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 106a0f58-3276-4754-93cd-{{Gerrit|a7ae20fddc75}} (cluster eqiad1) * 15:46 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 106a0f58-3276-4754-93cd-{{Gerrit|a7ae20fddc75}} (cluster eqiad1) * 15:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm d9e4c884-82f1-4c2e-8b35-{{Gerrit|70bfeb5292cf}} (cluster eqiad1) * 15:46 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm d9e4c884-82f1-4c2e-8b35-{{Gerrit|70bfeb5292cf}} (cluster eqiad1) * 15:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 4945aa99-aeff-4198-9aaa-{{Gerrit|7391c9a84c55}} (cluster eqiad1) * 15:45 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 4945aa99-aeff-4198-9aaa-{{Gerrit|7391c9a84c55}} (cluster eqiad1) * 15:45 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 278b5002-e9db-4506-a40f-{{Gerrit|167b52b9515f}} (cluster eqiad1) * 15:45 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 278b5002-e9db-4506-a40f-{{Gerrit|167b52b9515f}} (cluster eqiad1) * 15:45 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 36b6590c-eac2-40a0-ac30-{{Gerrit|7cf79ff12ce3}} (cluster eqiad1) * 15:44 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 36b6590c-eac2-40a0-ac30-{{Gerrit|7cf79ff12ce3}} (cluster eqiad1) * 15:44 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 797c51db-bc81-4363-922e-{{Gerrit|a52c3fc3eeea}} (cluster eqiad1) * 15:43 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 797c51db-bc81-4363-922e-{{Gerrit|a52c3fc3eeea}} (cluster eqiad1) * 15:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 627735e5-57c7-4714-855b-{{Gerrit|b7311fc527c6}} (cluster eqiad1) * 15:43 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 627735e5-57c7-4714-855b-{{Gerrit|b7311fc527c6}} (cluster eqiad1) * 15:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 90b1b0c3-14fb-47f6-9c50-{{Gerrit|f952f55bcfea}} (cluster eqiad1) * 15:42 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 90b1b0c3-14fb-47f6-9c50-{{Gerrit|f952f55bcfea}} (cluster eqiad1) * 15:41 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm bc10309b-5227-4ca2-b74c-{{Gerrit|440e2fdc116e}} (cluster eqiad1) * 15:40 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm bc10309b-5227-4ca2-b74c-{{Gerrit|440e2fdc116e}} (cluster eqiad1) * 15:40 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm c4c0ffc0-ebd2-4133-9912-{{Gerrit|585af2725bfd}} (cluster eqiad1) * 15:40 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 2f7b7cfa-12ed-41d5-977d-{{Gerrit|1e11e8335cf4}} (cluster eqiad1) * 15:40 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm c4c0ffc0-ebd2-4133-9912-{{Gerrit|585af2725bfd}} (cluster eqiad1) * 15:39 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 82b22752-6752-4814-90c5-{{Gerrit|2aebd3825e95}} (cluster eqiad1) * 15:39 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 2f7b7cfa-12ed-41d5-977d-{{Gerrit|1e11e8335cf4}} (cluster eqiad1) * 15:39 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm d761eca2-4a21-4522-8d95-{{Gerrit|584bf639e6c0}} (cluster eqiad1) * 15:39 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 82b22752-6752-4814-90c5-{{Gerrit|2aebd3825e95}} (cluster eqiad1) * 15:39 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 6c263c50-71da-40ee-b1e0-{{Gerrit|00d40ba108e7}} (cluster eqiad1) * 15:38 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 6c263c50-71da-40ee-b1e0-{{Gerrit|00d40ba108e7}} (cluster eqiad1) * 15:38 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 65283e58-53e0-4545-b201-{{Gerrit|dab88a8ae7e5}} (cluster eqiad1) * 15:38 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm d761eca2-4a21-4522-8d95-{{Gerrit|584bf639e6c0}} (cluster eqiad1) * 15:37 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 65283e58-53e0-4545-b201-{{Gerrit|dab88a8ae7e5}} (cluster eqiad1) * 15:37 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm b875763a-d70f-4cab-92ce-{{Gerrit|60a523161799}} (cluster eqiad1) * 15:37 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm abf4e1e6-1bd6-41f2-ad1c-{{Gerrit|345e940b0b8b}} (cluster eqiad1) * 15:36 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm abf4e1e6-1bd6-41f2-ad1c-{{Gerrit|345e940b0b8b}} (cluster eqiad1) * 15:36 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 43462c80-0923-4494-a5db-{{Gerrit|a8df39d71cdd}} (cluster eqiad1) * 15:35 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 43462c80-0923-4494-a5db-{{Gerrit|a8df39d71cdd}} (cluster eqiad1) * 15:35 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm b875763a-d70f-4cab-92ce-{{Gerrit|60a523161799}} (cluster eqiad1) * 15:35 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm a80c58d9-fcce-4739-9f83-{{Gerrit|204cff354959}} (cluster eqiad1) * 15:34 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm a80c58d9-fcce-4739-9f83-{{Gerrit|204cff354959}} (cluster eqiad1) * 15:34 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 03d3daa4-c46e-4152-a4dd-{{Gerrit|c02a872f7edd}} (cluster eqiad1) * 15:34 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 03d3daa4-c46e-4152-a4dd-{{Gerrit|c02a872f7edd}} (cluster eqiad1) * 15:34 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 2074763a-97af-4b3d-a3b5-{{Gerrit|7d5cf43b9ecd}} (cluster eqiad1) * 15:33 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 2074763a-97af-4b3d-a3b5-{{Gerrit|7d5cf43b9ecd}} (cluster eqiad1) * 15:33 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 23c93ff7-f301-41e5-9ea5-{{Gerrit|9d4b2da1bf22}} (cluster eqiad1) * 15:33 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 23c93ff7-f301-41e5-9ea5-{{Gerrit|9d4b2da1bf22}} (cluster eqiad1) * 15:32 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 24e5b10a-80df-4bbc-807c-{{Gerrit|97d4e935d1f4}} (cluster eqiad1) * 15:32 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 35433ec4-9fd5-49f8-ac51-{{Gerrit|c05ecb433a4d}} (cluster eqiad1) * 15:32 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 24e5b10a-80df-4bbc-807c-{{Gerrit|97d4e935d1f4}} (cluster eqiad1) * 15:32 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 8e6f5b87-57b8-4ba5-b9e6-{{Gerrit|8feb4e413f3d}} (cluster eqiad1) * 15:32 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 35433ec4-9fd5-49f8-ac51-{{Gerrit|c05ecb433a4d}} (cluster eqiad1) * 15:31 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 01042fb6-b2e4-4690-88fb-{{Gerrit|3840c98b01aa}} (cluster eqiad1) * 15:31 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 8e6f5b87-57b8-4ba5-b9e6-{{Gerrit|8feb4e413f3d}} (cluster eqiad1) * 15:31 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm be716e27-6b34-4cb0-a498-{{Gerrit|b300937edc4c}} (cluster eqiad1) * 15:31 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 01042fb6-b2e4-4690-88fb-{{Gerrit|3840c98b01aa}} (cluster eqiad1) * 15:30 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm be716e27-6b34-4cb0-a498-{{Gerrit|b300937edc4c}} (cluster eqiad1) * 15:30 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 736c7c6d-319d-43e0-b2b1-{{Gerrit|efdd84b4736a}} (cluster eqiad1) * 15:30 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 736c7c6d-319d-43e0-b2b1-{{Gerrit|efdd84b4736a}} (cluster eqiad1) * 15:30 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm a5cb6818-f3ac-4ba9-afb5-{{Gerrit|5c657cf65f9a}} (cluster eqiad1) * 15:29 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm a5cb6818-f3ac-4ba9-afb5-{{Gerrit|5c657cf65f9a}} (cluster eqiad1) * 15:29 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 702a55e1-e176-45f1-af81-{{Gerrit|569013f91be3}} (cluster eqiad1) * 15:29 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm a5705913-72f2-4abd-84e6-{{Gerrit|3e084bfbd98d}} (cluster eqiad1) * 15:29 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 702a55e1-e176-45f1-af81-{{Gerrit|569013f91be3}} (cluster eqiad1) * 15:29 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 2a4b9dfd-7006-4b5b-8c95-{{Gerrit|7883709e5b2d}} (cluster eqiad1) * 15:28 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 2a4b9dfd-7006-4b5b-8c95-{{Gerrit|7883709e5b2d}} (cluster eqiad1) * 15:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 5037f71d-bcbf-4ed7-809b-{{Gerrit|052ca6026219}} (cluster eqiad1) * 15:28 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 5037f71d-bcbf-4ed7-809b-{{Gerrit|052ca6026219}} (cluster eqiad1) * 15:27 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 31bf05f8-122b-4558-8932-{{Gerrit|7ac4b8375ed5}} (cluster eqiad1) * 15:27 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 31bf05f8-122b-4558-8932-{{Gerrit|7ac4b8375ed5}} (cluster eqiad1) * 15:27 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 0b5b7c51-dc42-4bec-90f2-{{Gerrit|161807a385f7}} (cluster eqiad1) * 15:27 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm a5705913-72f2-4abd-84e6-{{Gerrit|3e084bfbd98d}} (cluster eqiad1) * 15:26 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 966b6ab3-b561-4e46-bfcd-{{Gerrit|1681ce9e91ac}} (cluster eqiad1) * 15:26 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 0b5b7c51-dc42-4bec-90f2-{{Gerrit|161807a385f7}} (cluster eqiad1) * 15:26 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm f7e8f001-e9c0-4fe9-8887-{{Gerrit|32289702b804}} (cluster eqiad1) * 15:26 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm f7e8f001-e9c0-4fe9-8887-{{Gerrit|32289702b804}} (cluster eqiad1) * 15:26 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm d73171e3-49ef-4d40-8008-{{Gerrit|a900781ea102}} (cluster eqiad1) * 15:26 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 966b6ab3-b561-4e46-bfcd-{{Gerrit|1681ce9e91ac}} (cluster eqiad1) * 15:25 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm d73171e3-49ef-4d40-8008-{{Gerrit|a900781ea102}} (cluster eqiad1) * 15:25 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 63a229be-765f-4b48-b8d9-{{Gerrit|24ee39243604}} (cluster eqiad1) * 15:25 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 9a7cf939-c634-4aa1-9fd2-{{Gerrit|dbc14b18d70e}} (cluster eqiad1) * 15:25 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 63a229be-765f-4b48-b8d9-{{Gerrit|24ee39243604}} (cluster eqiad1) * 15:25 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 63f68215-d302-4684-a91e-{{Gerrit|58f5272486a5}} (cluster eqiad1) * 15:25 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 9a7cf939-c634-4aa1-9fd2-{{Gerrit|dbc14b18d70e}} (cluster eqiad1) * 15:24 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 63f68215-d302-4684-a91e-{{Gerrit|58f5272486a5}} (cluster eqiad1) * 15:24 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 9c66be4e-6787-4844-a9ff-{{Gerrit|a65295ac5aac}} (cluster eqiad1) * 15:24 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm a4892d89-0981-412e-9f00-{{Gerrit|8882416948a1}} (cluster eqiad1) * 15:23 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm a4892d89-0981-412e-9f00-{{Gerrit|8882416948a1}} (cluster eqiad1) * 15:23 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm f8e08f70-f87e-413d-acad-{{Gerrit|080126ad5b1a}} (cluster eqiad1) * 15:23 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 9c66be4e-6787-4844-a9ff-{{Gerrit|a65295ac5aac}} (cluster eqiad1) * 15:23 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm e6796dbd-2511-4bf6-bdee-{{Gerrit|4a14a7414d5f}} (cluster eqiad1) * 15:23 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm f8e08f70-f87e-413d-acad-{{Gerrit|080126ad5b1a}} (cluster eqiad1) * 15:23 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 7e3011e8-aed8-4bed-8e18-{{Gerrit|f75afe3ec3a2}} (cluster eqiad1) * 15:22 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm e6796dbd-2511-4bf6-bdee-{{Gerrit|4a14a7414d5f}} (cluster eqiad1) * 15:22 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 7e3011e8-aed8-4bed-8e18-{{Gerrit|f75afe3ec3a2}} (cluster eqiad1) * 15:22 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 6fa9b0be-219d-4b10-962e-{{Gerrit|fa3a71f6740c}} (cluster eqiad1) * 15:22 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 5e5c3bad-f1c7-49e5-b846-{{Gerrit|edaf111af83c}} (cluster eqiad1) * 15:21 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 5e5c3bad-f1c7-49e5-b846-{{Gerrit|edaf111af83c}} (cluster eqiad1) * 15:21 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 6fa9b0be-219d-4b10-962e-{{Gerrit|fa3a71f6740c}} (cluster eqiad1) * 15:21 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 995817db-0966-485d-aca5-{{Gerrit|e5377c77a005}} (cluster eqiad1) * 15:21 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 995817db-0966-485d-aca5-{{Gerrit|e5377c77a005}} (cluster eqiad1) * 15:20 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm d409f39a-e24a-462e-b588-{{Gerrit|6f5f6557e26b}} (cluster eqiad1) * 15:20 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm d409f39a-e24a-462e-b588-{{Gerrit|6f5f6557e26b}} (cluster eqiad1) * 15:20 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 7dc8757e-8b8d-4cc9-ac8e-{{Gerrit|a2f925639f0b}} (cluster eqiad1) * 15:19 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.vps.instance.stop_start (exit_code=99) vm ci2.mediawiki-quickstart.eqiad1.wikimedia.cloud (cluster eqiad1) * 15:19 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm ci2.mediawiki-quickstart.eqiad1.wikimedia.cloud (cluster eqiad1) * 15:19 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.vps.instance.stop_start (exit_code=99) vm None (cluster eqiad1) * 15:19 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm None (cluster eqiad1) * 15:19 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 7dc8757e-8b8d-4cc9-ac8e-{{Gerrit|a2f925639f0b}} (cluster eqiad1) * 15:19 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 94f8be5c-3cdf-47cb-80b2-{{Gerrit|43c44da01789}} (cluster eqiad1) * 15:18 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 94f8be5c-3cdf-47cb-80b2-{{Gerrit|43c44da01789}} (cluster eqiad1) * 15:18 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm b89a2d14-2bc3-481b-baf6-{{Gerrit|496edadbe242}} (cluster eqiad1) * 15:17 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm b89a2d14-2bc3-481b-baf6-{{Gerrit|496edadbe242}} (cluster eqiad1) * 15:17 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm a8269e6b-e09d-4b6b-909e-{{Gerrit|5e4014165440}} (cluster eqiad1) * 15:17 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm a8269e6b-e09d-4b6b-909e-{{Gerrit|5e4014165440}} (cluster eqiad1) * 15:17 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm b4b2d7cd-3492-4d04-86ce-{{Gerrit|4c0b8344ddc3}} (cluster eqiad1) * 15:16 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm b4b2d7cd-3492-4d04-86ce-{{Gerrit|4c0b8344ddc3}} (cluster eqiad1) * 15:16 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm def4772e-c01f-4e5e-8e70-{{Gerrit|d62a546ebc2a}} (cluster eqiad1) * 15:16 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm def4772e-c01f-4e5e-8e70-{{Gerrit|d62a546ebc2a}} (cluster eqiad1) * 15:16 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 2b573025-a221-482a-b6b7-{{Gerrit|fd7e1b1308f7}} (cluster eqiad1) * 15:15 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 2b573025-a221-482a-b6b7-{{Gerrit|fd7e1b1308f7}} (cluster eqiad1) * 15:15 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 55b557f8-5817-47f6-adb4-{{Gerrit|abccac2b2997}} (cluster eqiad1) * 15:15 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 55b557f8-5817-47f6-adb4-{{Gerrit|abccac2b2997}} (cluster eqiad1) * 15:15 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 62cf19eb-ebf9-49d9-baa7-{{Gerrit|fba2cd6942d6}} (cluster eqiad1) * 15:14 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 62cf19eb-ebf9-49d9-baa7-{{Gerrit|fba2cd6942d6}} (cluster eqiad1) * 15:14 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm d3812466-439a-4355-901c-{{Gerrit|b1097a033d0b}} (cluster eqiad1) * 15:13 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm d3812466-439a-4355-901c-{{Gerrit|b1097a033d0b}} (cluster eqiad1) * 15:13 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 06ea2eca-b4e5-42fa-afc3-{{Gerrit|684b3c2b87a3}} (cluster eqiad1) * 15:13 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 06ea2eca-b4e5-42fa-afc3-{{Gerrit|684b3c2b87a3}} (cluster eqiad1) * 15:13 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 7cafb806-ecc1-459d-b6b3-{{Gerrit|4213511f1257}} (cluster eqiad1) * 15:12 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 7cafb806-ecc1-459d-b6b3-{{Gerrit|4213511f1257}} (cluster eqiad1) * 15:12 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm bbe5835d-5e8d-4778-b80f-{{Gerrit|5c6424928a88}} (cluster eqiad1) * 15:12 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm bbe5835d-5e8d-4778-b80f-{{Gerrit|5c6424928a88}} (cluster eqiad1) * 15:12 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 143f7189-22b2-4708-9a33-{{Gerrit|91d368a5c8eb}} (cluster eqiad1) * 15:11 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 143f7189-22b2-4708-9a33-{{Gerrit|91d368a5c8eb}} (cluster eqiad1) * 15:11 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm b5fcdad2-d7d9-4ba2-903e-{{Gerrit|188231b05f71}} (cluster eqiad1) * 15:09 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm b5fcdad2-d7d9-4ba2-903e-{{Gerrit|188231b05f71}} (cluster eqiad1) * 15:09 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 9c7abfa9-a848-4ed5-8abb-{{Gerrit|dda2cb842b04}} (cluster eqiad1) * 15:07 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 9c7abfa9-a848-4ed5-8abb-{{Gerrit|dda2cb842b04}} (cluster eqiad1) * 15:07 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm dedb5d73-e8f8-47d3-8598-{{Gerrit|2900be096236}} (cluster eqiad1) * 15:06 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm dedb5d73-e8f8-47d3-8598-{{Gerrit|2900be096236}} (cluster eqiad1) * 15:06 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 7e7b75d2-0f13-4973-ad18-{{Gerrit|3dc7a52b0781}} (cluster eqiad1) * 15:06 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 7e7b75d2-0f13-4973-ad18-{{Gerrit|3dc7a52b0781}} (cluster eqiad1) * 15:06 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 9b6c52a7-527d-4513-b996-{{Gerrit|160af646c5fb}} (cluster eqiad1) * 15:05 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 9b6c52a7-527d-4513-b996-{{Gerrit|160af646c5fb}} (cluster eqiad1) * 15:05 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm d4f8ac4d-3059-499a-961c-{{Gerrit|505f6ff89675}} (cluster eqiad1) * 15:04 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm d4f8ac4d-3059-499a-961c-{{Gerrit|505f6ff89675}} (cluster eqiad1) * 15:04 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 2a87ab18-417b-42a6-9ee1-{{Gerrit|a273a2379e62}} (cluster eqiad1) * 15:04 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 2a87ab18-417b-42a6-9ee1-{{Gerrit|a273a2379e62}} (cluster eqiad1) * 15:04 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 9f04369c-d074-44e2-a3b2-{{Gerrit|b0545accd0e0}} (cluster eqiad1) * 15:03 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 9f04369c-d074-44e2-a3b2-{{Gerrit|b0545accd0e0}} (cluster eqiad1) * 15:03 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 1b6df367-1d3d-4e48-8333-{{Gerrit|7a4f79a49a2a}} (cluster eqiad1) * 15:02 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 1b6df367-1d3d-4e48-8333-{{Gerrit|7a4f79a49a2a}} (cluster eqiad1) * 15:01 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm a07ac179-366e-49a4-9499-{{Gerrit|bee721949963}} (cluster eqiad1) * 15:00 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm a07ac179-366e-49a4-9499-{{Gerrit|bee721949963}} (cluster eqiad1) === 2025-11-27 === * 04:54 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 04:50 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 04:48 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 04:48 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 04:47 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 04:46 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 04:33 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 04:33 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services === 2025-11-26 === * 15:26 dhinus: depool clouddb10[17-20] for network maintenance [[phab:T404609|T404609]] * 10:59 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for main branch * 10:59 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 10:24 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch ([[phab:T408387|T408387]]) * 10:23 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch ([[phab:T408387|T408387]]) * 10:23 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for main branch * 10:23 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch === 2025-11-25 === * 15:52 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 15:48 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 15:29 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 15:28 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 15:28 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 15:28 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 15:19 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 15:18 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 08:17 godog: restart neutron-metadata-agent for testing - [[phab:T410983|T410983]] === 2025-11-24 === * 15:01 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for main branch * 15:01 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 14:53 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for main branch * 14:53 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 14:30 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for main branch * 14:29 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch === 2025-11-23 === * 21:04 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1041.eqiad.wmnet<nowiki>}</nowiki>' * 20:47 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1041.eqiad.wmnet<nowiki>}</nowiki>' * 20:47 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1042.eqiad.wmnet<nowiki>}</nowiki>' * 20:22 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1042.eqiad.wmnet<nowiki>}</nowiki>' * 20:14 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2006-dev.codfw.wmnet<nowiki>}</nowiki>' * 19:46 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2006-dev.codfw.wmnet<nowiki>}</nowiki>' * 19:40 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2005-dev.codfw.wmnet<nowiki>}</nowiki>' * 19:33 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2005-dev.codfw.wmnet<nowiki>}</nowiki>' * 19:31 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt2005-dev.codfw.wmnet' * 19:28 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt2005-dev.codfw.wmnet' === 2025-11-19 === * 20:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2005-dev.codfw.wmnet<nowiki>}</nowiki>' * 20:42 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2004-dev.codfw.wmnet<nowiki>}</nowiki>' * 20:38 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2004-dev.codfw.wmnet<nowiki>}</nowiki>' * 20:37 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt2004-dev.codfw.wmnet' * 20:33 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 20:32 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt2004-dev.codfw.wmnet' * 20:29 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt2004-dev.codfw.wmnet' * 20:28 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 20:28 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 20:28 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 20:27 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 20:27 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 20:16 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt2004-dev.codfw.wmnet' * 19:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1070.eqiad.wmnet' * 19:45 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1070.eqiad.wmnet' * 19:45 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 764d92cd-09df-468a-9595-{{Gerrit|7bcfbc4a8841}} (cluster eqiad1) * 19:44 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 764d92cd-09df-468a-9595-{{Gerrit|7bcfbc4a8841}} (cluster eqiad1) * 19:44 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.vps.instance.stop_start (exit_code=99) vm content-diff-index (cluster eqiad1) * 19:44 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm content-diff-index (cluster eqiad1) * 19:44 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.vps.instance.stop_start (exit_code=99) vm None (cluster eqiad1) * 19:44 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm None (cluster eqiad1) * 18:43 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1040.eqiad.wmnet' * 18:40 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1040.eqiad.wmnet' * 18:40 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1040.eqiad.wmnet' * 18:37 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1040.eqiad.wmnet' * 18:37 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1040.eqiad.wmnet' * 18:34 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1040.eqiad.wmnet' * 18:20 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1040.eqiad.wmnet' * 18:07 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1040.eqiad.wmnet' * 16:10 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt2004-dev.codfw.wmnet' * 16:09 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt2004-dev.codfw.wmnet' * 16:04 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1070.eqiad.wmnet' * 16:03 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1070.eqiad.wmnet' * 16:03 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1061.eqiad.wmnet' * 15:45 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1061.eqiad.wmnet' * 15:44 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1054.eqiad.wmnet' * 15:40 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1054.eqiad.wmnet' * 15:40 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1054.eqiad.wmnet' * 15:25 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1054.eqiad.wmnet' * 15:25 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1046.eqiad.wmnet' * 15:22 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1046.eqiad.wmnet' * 15:21 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1050.eqiad.wmnet' * 15:20 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1050.eqiad.wmnet' * 15:20 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) * 15:20 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 15:19 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) * 15:19 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 15:19 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) * 15:18 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 15:18 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) * 15:17 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 15:17 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) * 15:16 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 15:12 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1050.eqiad.wmnet' * 15:11 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1050.eqiad.wmnet' * 03:21 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1050.eqiad.wmnet' * 03:04 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1050.eqiad.wmnet' * 03:04 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1046.eqiad.wmnet' * 03:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1046.eqiad.wmnet' * 01:57 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1076.eqiad.wmnet' * 01:50 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1076.eqiad.wmnet' * 01:41 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1046.eqiad.wmnet' * 01:22 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1046.eqiad.wmnet' === 2025-11-18 === * 22:06 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1076.eqiad.wmnet' * 22:03 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1076.eqiad.wmnet' * 22:03 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1076.eqiad.wmnet' * 21:50 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1076.eqiad.wmnet' * 21:50 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1075.eqiad.wmnet' * 21:39 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1075.eqiad.wmnet' * 21:24 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1074.eqiad.wmnet' * 21:14 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1074.eqiad.wmnet' * 21:14 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1073.eqiad.wmnet' * 21:13 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1073.eqiad.wmnet' * 21:12 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1073.eqiad.wmnet' * 21:11 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1073.eqiad.wmnet' * 21:10 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1073.eqiad.wmnet' * 20:53 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1073.eqiad.wmnet' * 20:13 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1072.eqiad.wmnet' * 19:53 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1072.eqiad.wmnet' * 19:51 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1071.eqiad.wmnet' * 19:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1071.eqiad.wmnet' * 19:40 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1070.eqiad.wmnet' * 19:39 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1070.eqiad.wmnet' * 19:37 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1070.eqiad.wmnet' * 19:25 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1070.eqiad.wmnet' * 19:22 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1069.eqiad.wmnet' * 19:09 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1069.eqiad.wmnet' * 18:59 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1068.eqiad.wmnet' * 18:45 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1068.eqiad.wmnet' * 18:45 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1067.eqiad.wmnet' * 18:30 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1067.eqiad.wmnet' * 18:13 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1066.eqiad.wmnet' * 17:56 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1066.eqiad.wmnet' * 17:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1044.eqiad.wmnet' * 17:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1044.eqiad.wmnet' * 17:40 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1044.eqiad.wmnet' * 17:38 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1044.eqiad.wmnet' * 17:36 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1044.eqiad.wmnet' * 17:25 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1044.eqiad.wmnet' * 17:25 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1065.eqiad.wmnet' * 17:10 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1065.eqiad.wmnet' * 17:09 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1064.eqiad.wmnet' * 16:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1064.eqiad.wmnet' * 16:36 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1063.eqiad.wmnet' * 16:22 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1063.eqiad.wmnet' * 16:11 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1062.eqiad.wmnet' * 15:56 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1062.eqiad.wmnet' * 15:55 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1061.eqiad.wmnet' * 15:54 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1061.eqiad.wmnet' * 15:50 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1061.eqiad.wmnet' * 15:35 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1061.eqiad.wmnet' * 15:33 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1060.eqiad.wmnet' * 15:13 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 15:12 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 14:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1060.eqiad.wmnet' * 10:00 godog: switch cloudcephosd1049 to single nic - [[phab:T399180|T399180]] * 05:31 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1059.eqiad.wmnet' * 05:15 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1059.eqiad.wmnet' * 04:41 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1058.eqiad.wmnet' * 04:25 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1058.eqiad.wmnet' * 04:24 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1057.eqiad.wmnet' * 04:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1057.eqiad.wmnet' * 00:35 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1056.eqiad.wmnet' * 00:15 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1056.eqiad.wmnet' === 2025-11-17 === * 23:59 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1055.eqiad.wmnet' * 23:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1055.eqiad.wmnet' * 23:43 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1054.eqiad.wmnet' * 23:38 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1054.eqiad.wmnet' * 22:39 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1054.eqiad.wmnet' * 22:19 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1054.eqiad.wmnet' * 20:32 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1053.eqiad.wmnet' * 20:13 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1053.eqiad.wmnet' * 20:12 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1052.eqiad.wmnet' * 19:56 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1052.eqiad.wmnet' * 19:17 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1051.eqiad.wmnet' * 19:01 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1051.eqiad.wmnet' * 19:00 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1050.eqiad.wmnet' * 18:39 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1050.eqiad.wmnet' * 18:39 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1049.eqiad.wmnet' * 17:59 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1048.eqiad.wmnet' * 17:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1047.eqiad.wmnet' * 17:30 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1047.eqiad.wmnet' * 17:20 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1046.eqiad.wmnet' * 17:09 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1046.eqiad.wmnet' * 17:04 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1045.eqiad.wmnet' * 16:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1045.eqiad.wmnet' * 15:58 godog: set ceph cluster back to out and rebalance - [[phab:T399180|T399180]] * 15:49 godog: set ceph cluster noout/norebalance and move cloudcephosd1048 to single nic - [[phab:T399180|T399180]] * 14:29 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1044.eqiad.wmnet' * 14:28 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1044.eqiad.wmnet' * 10:30 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch ([[phab:T409365|T409365]]) * 10:27 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch ([[phab:T409365|T409365]]) * 10:26 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for main branch ([[phab:T409365|T409365]]) * 10:25 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch ([[phab:T409365|T409365]]) * 09:18 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for service: project,keystone * 09:18 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for service: project,keystone === 2025-11-16 === * 23:31 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1044.eqiad.wmnet' * 23:30 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1044.eqiad.wmnet' * 18:33 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1044.eqiad.wmnet' * 18:32 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1044.eqiad.wmnet' * 18:14 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1044.eqiad.wmnet' * 18:08 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1044.eqiad.wmnet' * 18:08 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1043.eqiad.wmnet' * 18:07 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1043.eqiad.wmnet' === 2025-11-15 === * 01:01 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1043.eqiad.wmnet' * 01:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1043.eqiad.wmnet' * 00:59 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1043.eqiad.wmnet' * 00:58 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1043.eqiad.wmnet' * 00:53 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1043.eqiad.wmnet' * 00:15 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1043.eqiad.wmnet' * 00:09 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=99) * 00:09 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 00:09 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=99) * 00:09 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 00:09 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=99) * 00:08 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance === 2025-11-14 === * 23:59 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1042.eqiad.wmnet' * 23:58 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1042.eqiad.wmnet' * 21:37 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1042.eqiad.wmnet' * 21:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1042.eqiad.wmnet' * 21:35 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1042.eqiad.wmnet' * 21:33 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1042.eqiad.wmnet' * 21:30 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1042.eqiad.wmnet' * 21:23 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1042.eqiad.wmnet' * 20:56 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1041.eqiad.wmnet' * 20:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1041.eqiad.wmnet' * 20:40 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1040.eqiad.wmnet' * 20:39 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1040.eqiad.wmnet' * 15:13 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/279 ([[phab:T409365|T409365]]) * 15:13 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/279 ([[phab:T409365|T409365]]) * 15:12 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/279 ([[phab:T409365|T409365]]) * 15:12 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/279 ([[phab:T409365|T409365]]) * 15:10 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for main branch ([[phab:T409365|T409365]]) * 15:09 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch ([[phab:T409365|T409365]]) === 2025-11-13 === * 12:17 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 12:16 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:12 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.roll_reboot_cloudnets (exit_code=0) * 12:04 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.roll_reboot_cloudnets * 12:04 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,neutron * 11:52 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,neutron * 10:59 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/282 * 10:58 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/282 === 2025-11-11 === * 13:46 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 13:45 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2025-11-10 === * 15:46 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 15:45 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 15:44 taavi: rotate cookbook gitlab access token before(!) it expires [[phab:T409741|T409741]] * 15:37 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/280 * 15:36 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/280 * 15:35 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/280 * 15:35 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/280 * 15:33 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/280 * 15:33 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/280 * 14:53 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.roll_reboot_cloudnets (exit_code=0) * 14:43 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.roll_reboot_cloudnets * 14:28 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.roll_reboot_cloudnets (exit_code=99) * 14:28 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.roll_reboot_cloudnets * 14:27 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for service: project,neutron * 14:25 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for service: project,neutron === 2025-11-07 === * 11:33 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.vps.remove_instance (exit_code=1) for instance tools-db-7 * 11:33 fnegri@cloudcumin1001: START - Cookbook wmcs.vps.remove_instance for instance tools-db-7 === 2025-10-28 === * 16:52 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 16:48 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 02:34 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 02:30 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services === 2025-10-27 === * 20:14 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 20:10 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 19:17 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 19:16 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 19:16 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 19:16 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 19:15 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 19:13 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 18:58 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 18:54 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services === 2025-10-25 === * 18:54 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) ([[phab:T405478|T405478]]) === 2025-10-24 === * 21:54 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=99) * 21:54 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 21:52 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=99) * 21:52 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 21:52 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=99) * 21:52 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 21:52 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=99) * 21:51 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 20:55 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt2006-dev.codfw.wmnet' * 20:32 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt2006-dev.codfw.wmnet' * 20:30 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 20:27 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 20:25 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 20:25 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 20:25 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 20:25 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 19:19 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt2005-dev.codfw.wmnet' * 19:18 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt2005-dev.codfw.wmnet' * 17:15 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt2005-dev.codfw.wmnet' * 17:14 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt2005-dev.codfw.wmnet' * 17:10 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt2005-dev.codfw.wmnet' * 17:03 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt2005-dev.codfw.wmnet' * 11:14 filippo@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node ([[phab:T405478|T405478]]) * 08:21 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) ([[phab:T405478|T405478]]) === 2025-10-23 === * 21:42 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt2004-dev.codfw.wmnet' * 21:42 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt2004-dev.codfw.wmnet' * 21:40 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 21:35 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 21:34 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 21:34 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 21:29 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 21:29 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 21:26 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 21:26 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 21:25 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 21:25 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 21:24 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 21:24 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 21:21 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 21:21 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 21:16 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 21:16 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 20:12 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt2004-dev.codfw.wmnet' * 20:12 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt2004-dev.codfw.wmnet' * 16:02 filippo@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node ([[phab:T405478|T405478]]) * 15:08 filippo@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) ([[phab:T405478|T405478]]) * 07:08 filippo@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node ([[phab:T405478|T405478]]) === 2025-10-22 === * 15:51 filippo@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) ([[phab:T405478|T405478]]) * 07:51 filippo@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node ([[phab:T405478|T405478]]) === 2025-10-21 === * 23:48 filippo@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) ([[phab:T405478|T405478]]) * 15:47 filippo@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node ([[phab:T405478|T405478]]) === 2025-10-20 === * 11:52 filippo@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=97) * 11:32 filippo@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add === 2025-10-18 === * 05:31 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services * 05:18 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services === 2025-10-16 === * 18:41 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 18:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 18:32 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 18:31 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 18:30 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 18:30 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services === 2025-10-15 === * 14:22 filippo@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=97) * 14:13 filippo@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add === 2025-10-14 === * 19:27 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 19:23 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 19:18 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 19:17 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 19:16 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 19:16 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 19:13 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 19:13 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 19:08 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 19:08 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 11:26 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudweb.safe_reboot (exit_code=0) on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::cloudweb<nowiki>}</nowiki>' * 11:20 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudweb.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::cloudweb<nowiki>}</nowiki>' === 2025-10-08 === * 19:46 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 19:45 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2025-10-07 === * 18:57 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 18:57 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 18:56 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for main branch * 18:56 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 18:55 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for main branch * 18:54 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 18:52 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for main branch * 18:52 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch === 2025-10-01 === * 02:04 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2004-dev.codfw.wmnet<nowiki>}</nowiki>' * 01:52 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2004-dev.codfw.wmnet<nowiki>}</nowiki>' * 01:51 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2005-dev.codfw.wmnet<nowiki>}</nowiki>' * 01:47 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2005-dev.codfw.wmnet<nowiki>}</nowiki>' * 01:39 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2004-dev.codfw.wmnet<nowiki>}</nowiki>' * 01:11 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2004-dev.codfw.wmnet<nowiki>}</nowiki>' * 01:09 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 01:06 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 00:44 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2004-dev.codfw.wmnet<nowiki>}</nowiki>' * 00:30 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2004-dev.codfw.wmnet<nowiki>}</nowiki>' * 00:17 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2005-dev.codfw.wmnet<nowiki>}</nowiki>' === 2025-09-30 === * 23:20 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1076.eqiad.wmnet<nowiki>}</nowiki>' * 23:01 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1076.eqiad.wmnet<nowiki>}</nowiki>' * 23:01 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1075.eqiad.wmnet<nowiki>}</nowiki>' * 22:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1075.eqiad.wmnet<nowiki>}</nowiki>' * 22:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1074.eqiad.wmnet<nowiki>}</nowiki>' * 22:29 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2005-dev.codfw.wmnet<nowiki>}</nowiki>' * 22:29 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2006-dev.codfw.wmnet<nowiki>}</nowiki>' * 22:26 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1074.eqiad.wmnet<nowiki>}</nowiki>' * 22:26 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1073.eqiad.wmnet<nowiki>}</nowiki>' * 22:22 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1073.eqiad.wmnet<nowiki>}</nowiki>' * 22:22 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1072.eqiad.wmnet<nowiki>}</nowiki>' * 22:20 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2006-dev.codfw.wmnet<nowiki>}</nowiki>' * 22:20 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2006-dev.eqiad.wmnet<nowiki>}</nowiki>' * 22:19 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2006-dev.eqiad.wmnet<nowiki>}</nowiki>' * 22:03 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1072.eqiad.wmnet<nowiki>}</nowiki>' * 22:02 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1071.eqiad.wmnet<nowiki>}</nowiki>' * 21:59 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1071.eqiad.wmnet<nowiki>}</nowiki>' * 21:59 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1070.eqiad.wmnet<nowiki>}</nowiki>' * 21:40 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1070.eqiad.wmnet<nowiki>}</nowiki>' * 21:40 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1069.eqiad.wmnet<nowiki>}</nowiki>' * 21:20 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1069.eqiad.wmnet<nowiki>}</nowiki>' * 21:20 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1068.eqiad.wmnet<nowiki>}</nowiki>' * 21:04 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1068.eqiad.wmnet<nowiki>}</nowiki>' * 21:04 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1067.eqiad.wmnet<nowiki>}</nowiki>' * 20:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1067.eqiad.wmnet<nowiki>}</nowiki>' * 20:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1066.eqiad.wmnet<nowiki>}</nowiki>' * 20:27 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1066.eqiad.wmnet<nowiki>}</nowiki>' * 20:27 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1065.eqiad.wmnet<nowiki>}</nowiki>' * 20:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1065.eqiad.wmnet<nowiki>}</nowiki>' * 20:05 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1064.eqiad.wmnet<nowiki>}</nowiki>' * 19:46 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1064.eqiad.wmnet<nowiki>}</nowiki>' * 19:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1063.eqiad.wmnet<nowiki>}</nowiki>' * 19:23 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1063.eqiad.wmnet<nowiki>}</nowiki>' === 2025-09-26 === * 10:31 taavi: remove /var/lib/prometheus/node.d/kernel-messages.prom which got left as a leftover from the now-removed exporter === 2025-09-25 === * 23:07 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) === 2025-09-24 === * 15:02 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 15:02 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 09:27 volans@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1062.eqiad.wmnet<nowiki>}</nowiki>' * 09:07 volans@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1062.eqiad.wmnet<nowiki>}</nowiki>' * 08:53 volans@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1061.eqiad.wmnet<nowiki>}</nowiki>' * 08:35 volans@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1061.eqiad.wmnet<nowiki>}</nowiki>' * 08:35 volans@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1060.eqiad.wmnet<nowiki>}</nowiki>' * 08:13 volans@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1060.eqiad.wmnet<nowiki>}</nowiki>' * 07:12 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 02:00 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 01:58 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 01:58 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 01:15 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 01:13 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 01:11 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 01:11 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 01:09 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 01:09 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 01:04 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 01:04 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 01:01 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 01:01 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 01:01 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 01:01 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 01:00 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 01:00 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 00:59 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 00:59 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 00:59 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 00:59 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate === 2025-09-23 === * 23:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 23:46 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 23:44 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 23:44 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 23:03 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 22:56 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 22:31 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1059.eqiad.wmnet<nowiki>}</nowiki>' * 22:09 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1059.eqiad.wmnet<nowiki>}</nowiki>' * 22:09 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1058.eqiad.wmnet<nowiki>}</nowiki>' * 21:47 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1058.eqiad.wmnet<nowiki>}</nowiki>' * 21:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1057.eqiad.wmnet<nowiki>}</nowiki>' * 21:25 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1057.eqiad.wmnet<nowiki>}</nowiki>' * 21:25 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1056.eqiad.wmnet<nowiki>}</nowiki>' * 21:17 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 21:16 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 21:14 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 21:14 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 21:06 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1056.eqiad.wmnet<nowiki>}</nowiki>' * 21:06 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1055.eqiad.wmnet<nowiki>}</nowiki>' * 20:58 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'P<nowiki>{</nowiki>P:openstack::codfw1dev::nova::compute::service<nowiki>}</nowiki>' * 20:55 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>P:openstack::codfw1dev::nova::compute::service<nowiki>}</nowiki>' * 20:52 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'P<nowiki>{</nowiki>P:openstack::codfw1dev::nova::compute::service<nowiki>}</nowiki>' * 20:51 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>P:openstack::codfw1dev::nova::compute::service<nowiki>}</nowiki>' * 20:48 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'P<nowiki>{</nowiki>P:openstack::codfw1dev::nova::compute::service<nowiki>}</nowiki>' * 20:47 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>P:openstack::codfw1dev::nova::compute::service<nowiki>}</nowiki>' * 20:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 20:45 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1055.eqiad.wmnet<nowiki>}</nowiki>' * 20:45 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1054.eqiad.wmnet<nowiki>}</nowiki>' * 20:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 20:43 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'P<nowiki>{</nowiki>P:openstack::codfw1dev::nova::compute::service<nowiki>}</nowiki>' * 20:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>P:openstack::codfw1dev::nova::compute::service<nowiki>}</nowiki>' * 20:36 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=97) * 20:36 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 20:29 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 20:29 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 20:28 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1054.eqiad.wmnet<nowiki>}</nowiki>' * 20:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1053.eqiad.wmnet<nowiki>}</nowiki>' * 20:02 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1053.eqiad.wmnet<nowiki>}</nowiki>' * 20:02 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1052.eqiad.wmnet<nowiki>}</nowiki>' * 19:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1052.eqiad.wmnet<nowiki>}</nowiki>' * 19:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1051.eqiad.wmnet<nowiki>}</nowiki>' * 19:42 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 19:42 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 19:39 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 19:38 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 19:21 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1051.eqiad.wmnet<nowiki>}</nowiki>' * 19:21 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1050.eqiad.wmnet<nowiki>}</nowiki>' * 18:55 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1050.eqiad.wmnet<nowiki>}</nowiki>' * 18:54 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 18:54 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 18:53 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 18:53 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 18:53 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.reactivate (exit_code=97) * 18:52 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 18:51 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1049.eqiad.wmnet<nowiki>}</nowiki>' * 18:32 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1049.eqiad.wmnet<nowiki>}</nowiki>' * 18:32 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1048.eqiad.wmnet<nowiki>}</nowiki>' * 18:04 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1048.eqiad.wmnet<nowiki>}</nowiki>' * 18:04 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1047.eqiad.wmnet<nowiki>}</nowiki>' * 18:02 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 17:53 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 17:42 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1047.eqiad.wmnet<nowiki>}</nowiki>' * 17:42 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1046.eqiad.wmnet<nowiki>}</nowiki>' * 17:37 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1046.eqiad.wmnet<nowiki>}</nowiki>' * 17:37 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1045.eqiad.wmnet<nowiki>}</nowiki>' * 17:22 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1045.eqiad.wmnet<nowiki>}</nowiki>' * 17:22 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1044.eqiad.wmnet<nowiki>}</nowiki>' * 17:16 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1044.eqiad.wmnet<nowiki>}</nowiki>' * 17:16 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1043.eqiad.wmnet<nowiki>}</nowiki>' * 16:58 andrewbogott: test * 16:56 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1043.eqiad.wmnet<nowiki>}</nowiki>' * 16:55 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1042.eqiad.wmnet<nowiki>}</nowiki>' * 16:55 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 16:55 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 16:53 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=97) * 16:50 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1042.eqiad.wmnet<nowiki>}</nowiki>' * 16:50 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1041.eqiad.wmnet<nowiki>}</nowiki>' * 16:42 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 16:27 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1041.eqiad.wmnet<nowiki>}</nowiki>' * 16:26 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1040.eqiad.wmnet<nowiki>}</nowiki>' * 16:22 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1040.eqiad.wmnet<nowiki>}</nowiki>' * 15:55 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 15:46 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 14:37 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 14:25 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 14:13 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 14:12 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 14:11 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 14:10 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 04:34 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 04:28 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 03:45 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 03:45 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 03:44 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 03:44 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 03:04 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 03:03 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 02:22 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 02:21 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 02:19 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 02:19 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 01:19 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 01:18 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate === 2025-09-22 === * 22:55 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 22:54 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 22:53 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 22:53 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 22:10 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 22:10 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 21:26 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 21:25 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 21:25 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.reactivate (exit_code=97) * 21:25 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 21:25 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 21:25 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 20:20 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 20:14 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 18:22 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 18:16 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 17:37 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 17:36 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 16:51 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 16:50 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 16:47 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 16:47 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 16:02 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 16:01 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 15:59 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 15:59 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 15:59 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 15:59 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate === 2025-09-19 === * 00:17 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 00:10 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate === 2025-09-18 === * 22:34 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 22:28 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 21:17 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 21:16 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 16:55 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 16:54 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 16:53 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=97) * 16:53 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 16:52 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 16:52 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 16:51 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 16:08 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 16:02 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 16:02 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 15:15 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 15:15 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 15:14 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 15:07 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate === 2025-09-17 === * 22:55 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) === 2025-09-16 === * 14:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 14:45 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 14:41 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 14:33 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 14:33 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 14:33 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 14:32 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 14:32 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 13:40 andrewbogott: upgrading cloudcephosd1017 to bookworm/reef === 2025-09-15 === * 16:24 taavi: update nova-fullstack to run on trixie image === 2025-09-11 === * 20:03 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 19:56 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 19:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 19:42 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 19:26 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for service: project,designate * 19:24 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for service: project,designate * 19:24 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for service: project,designate * 19:23 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for service: project,designate * 19:22 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for service: project,designate * 19:21 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for service: project,designate * 19:19 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for main branch * 19:18 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 19:13 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 19:03 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 18:04 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 15:58 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 15:56 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 15:56 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate === 2025-09-10 === * 22:09 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 22:09 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 22:06 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 22:06 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 22:05 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 22:05 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 22:02 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 21:53 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 21:52 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 21:52 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 21:51 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 21:51 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 14:10 wmbot~dcaro@acme: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.vm_console (exit_code=99) * 14:10 wmbot~dcaro@acme: START - Cookbook wmcs.openstack.cloudvirt.vm_console * 14:09 wmbot~dcaro@acme: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.vm_console (exit_code=99) * 14:09 wmbot~dcaro@acme: START - Cookbook wmcs.openstack.cloudvirt.vm_console * 08:33 dhinus: add volans to cloud-vps domain admins: openstack role add --user volans --domain default --inherited admin === 2025-09-09 === * 16:47 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.upgrade_osds (exit_code=99) * 13:10 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_osds ([[phab:T402190|T402190]]) * 00:58 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node === 2025-09-08 === * 14:05 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.upgrade_mons (exit_code=0) * 13:45 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_mons === 2025-09-06 === * 11:17 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) === 2025-09-05 === * 21:36 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 15:46 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.drain_node (exit_code=97) * 15:46 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 15:46 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=97) * 15:46 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 15:44 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) * 15:43 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 15:42 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) * 15:41 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 15:41 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.drain_node (exit_code=97) * 15:41 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 15:38 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 15:38 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 15:38 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) ([[phab:T401693|T401693]]) * 15:38 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T401693|T401693]]) * 15:37 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) ([[phab:T401693|T401693]]) * 15:37 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T401693|T401693]]) * 15:37 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T401693|T401693]]) * 15:30 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T401693|T401693]]) * 15:15 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T401693|T401693]]) * 15:10 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T401693|T401693]]) === 2025-09-04 === * 16:36 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T395910|T395910]]) * 16:32 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T395910|T395910]]) === 2025-08-29 === * 19:06 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 14:05 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 05:01 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 00:01 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node === 2025-08-28 === * 20:50 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 15:49 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 15:46 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 12:53 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 12:53 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 09:55 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 09:55 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 07:06 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 07:06 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 04:31 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 04:31 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 01:52 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 01:52 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) === 2025-08-27 === * 22:55 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 22:55 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 20:03 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 20:03 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 16:32 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 16:16 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 12:51 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 06:47 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 03:44 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 03:42 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 01:39 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 01:38 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=97) * 01:04 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 01:04 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 01:04 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 01:04 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) === 2025-08-26 === * 20:57 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 19:35 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=97) * 15:02 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 14:02 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) ([[phab:T401693|T401693]]) * 01:20 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T401693|T401693]]) * 01:20 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T401693|T401693]]) * 01:11 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T401693|T401693]]) === 2025-08-25 === * 21:17 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=97) ([[phab:T401693|T401693]]) * 21:07 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T401693|T401693]]) * 17:40 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) * 17:40 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 17:39 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) * 17:38 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 17:37 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.drain_node (exit_code=97) * 17:36 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 17:35 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=97) * 17:31 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 17:30 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 14:28 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 13:36 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=97) ([[phab:T401693|T401693]]) * 04:06 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T401693|T401693]]) * 04:05 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) ([[phab:T401693|T401693]]) === 2025-08-24 === * 17:39 andrewbogott: removing /var/lib/prometheus/node.d/check_disk_space.prom on all cloudvirts and a few other servers. I am taking this file to be obsolete after https://gerrit.wikimedia.org/r/c/operations/puppet/+/1180501/2 ; removing it seems to clear the alert. * 17:19 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T401693|T401693]]) * 10:35 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) ([[phab:T401693|T401693]]) === 2025-08-23 === * 19:27 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T401693|T401693]]) * 19:24 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) ([[phab:T401693|T401693]]) * 04:21 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T401693|T401693]]) * 04:08 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T401693|T401693]]) * 04:02 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T401693|T401693]]) === 2025-08-22 === * 20:27 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 15:26 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 14:31 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) * 14:29 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 14:29 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.drain_node (exit_code=97) * 14:29 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 14:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) ([[phab:T395910|T395910]]) * 14:24 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T395910|T395910]]) * 14:20 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T395910|T395910]]) * 14:15 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T395910|T395910]]) * 08:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) === 2025-08-21 === * 16:01 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 15:58 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 15:58 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 15:57 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 15:57 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 15:55 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 15:55 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 15:42 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 15:39 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 15:35 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 15:33 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 14:52 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 14:50 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 14:47 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 14:44 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 14:41 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 14:41 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 14:39 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 14:39 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 14:38 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 14:38 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 14:38 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=97) * 14:37 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 14:36 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=97) * 14:35 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 12:58 dcaro@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T402499|T402499]]) * 12:57 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T402499|T402499]]) * 12:50 dcaro@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) ([[phab:T402499|T402499]]) * 12:18 dcaro: destroying osd 66 on cloudcephosd1004, will recreate ([[phab:T402499|T402499]]) * 12:17 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy ([[phab:T402499|T402499]]) * 10:35 dcaro: starting ceph-osd@69 on cloudcephosd1004 ([[phab:T402499|T402499]]) * 10:32 dcaro: starting ceph-osd@68 on cloudcephosd1004 ([[phab:T402499|T402499]]) * 09:44 dcaro: starting ceph-osd@66 on cloudcephosd1004 ([[phab:T402499|T402499]]) * 09:22 dcaro: test fol sal === 2025-08-20 === * 13:21 filippo@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.vm_console (exit_code=99) * 13:21 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.vm_console * 13:20 filippo@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.vm_console (exit_code=99) * 13:20 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.vm_console * 13:20 filippo@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.vm_console (exit_code=99) * 13:20 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.vm_console * 13:19 filippo@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.vm_console (exit_code=99) * 13:19 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.vm_console * 13:19 filippo@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.vm_console (exit_code=99) * 13:19 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.vm_console * 09:05 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.reboot_node (exit_code=0) ([[phab:T401319|T401319]]) * 09:01 fnegri@cloudcumin1001: START - Cookbook wmcs.ceph.reboot_node ([[phab:T401319|T401319]]) === 2025-08-18 === * 16:01 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.upgrade_mons (exit_code=0) * 15:42 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_mons ([[phab:T402190|T402190]]) * 15:33 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.upgrade_osds (exit_code=99) * 15:23 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_osds ([[phab:T402190|T402190]]) * 15:07 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.upgrade_osds (exit_code=0) * 15:01 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_osds ([[phab:T402190|T402190]]) === 2025-08-15 === * 20:24 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T401693|T401693]]) * 20:18 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T401693|T401693]]) * 19:58 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T401693|T401693]]) * 19:51 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T401693|T401693]]) * 19:33 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T401693|T401693]]) * 19:26 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T401693|T401693]]) === 2025-08-13 === * 23:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 23:15 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 23:09 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 23:09 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 23:05 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 22:58 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 22:58 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 22:58 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 21:52 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 21:52 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 21:52 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 21:52 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 19:53 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.upgrade_osds (exit_code=99) * 19:53 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_osds * 19:52 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.upgrade_osds (exit_code=99) * 19:52 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_osds * 19:51 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.upgrade_osds (exit_code=99) * 19:51 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_osds * 19:51 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.upgrade_osds (exit_code=99) * 19:51 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_osds * 19:04 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.upgrade_osds (exit_code=99) * 19:04 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_osds * 18:50 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.upgrade_osds (exit_code=0) * 18:43 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_osds === 2025-08-12 === * 19:58 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 19:55 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services === 2025-08-07 === * 09:52 taavi: remove ssh key from uid=soni LDAP user [[phab:T401318|T401318]] === 2025-08-06 === * 13:49 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 13:48 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2025-07-29 === * 08:08 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 08:05 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2025-07-24 === * 20:42 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 20:38 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 20:37 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for service: project,nova * 20:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for service: project,nova * 15:44 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 15:42 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 14:57 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/258 * 14:57 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/258 * 14:57 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/258 * 14:56 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/258 * 13:35 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for service: project,nova * 13:34 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for service: project,nova * 11:01 dcaro: stopping all ceph osds in codfw1 to avoid spamming the mons ([[phab:T400334|T400334]]) * 07:37 dcaro: downgrading the codfw1 ceph mons to pacific, to do a rebuild instead of in-place upgrade to quincy === 2025-07-22 === * 19:56 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.upgrade_mons (exit_code=0) * 19:40 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_mons * 19:15 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.upgrade_osds (exit_code=0) * 19:09 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_osds * 19:06 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.upgrade_osds (exit_code=0) * 19:00 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_osds * 18:37 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.upgrade_osds (exit_code=99) * 18:19 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_osds * 14:40 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 14:40 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 14:36 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 14:36 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 14:36 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.upgrade_osds (exit_code=97) * 14:26 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_osds * 10:06 wmbot~dcaro@hephaestus: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 10:04 wmbot~dcaro@hephaestus: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 10:04 wmbot~dcaro@hephaestus: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 10:03 wmbot~dcaro@hephaestus: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 10:03 wmbot~dcaro@hephaestus: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 07:45 dcaro@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) ([[phab:T399870|T399870]]) * 07:44 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy ([[phab:T399870|T399870]]) * 00:38 dcaro@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) === 2025-07-21 === * 20:55 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 20:54 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 20:20 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 20:20 dcaro@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=97) * 20:08 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 20:07 dcaro@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 20:07 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 20:07 dcaro@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 20:07 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 20:03 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 20:02 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 19:44 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=97) ([[phab:T399858|T399858]]) * 17:05 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 17:05 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 17:04 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 17:03 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 17:02 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 17:01 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 16:47 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 16:47 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 16:44 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 16:38 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 16:37 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 16:36 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 16:25 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudnet1006.eqiad.wmnet' ([[phab:T395255|T395255]]) * 16:16 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet1006.eqiad.wmnet' ([[phab:T395255|T395255]]) * 16:11 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudnet1005.eqiad.wmnet' ([[phab:T395255|T395255]]) * 16:02 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet1005.eqiad.wmnet' ([[phab:T395255|T395255]]) * 15:56 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T399858|T399858]]) * 15:53 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) ([[phab:T399858|T399858]]) * 15:53 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T399858|T399858]]) * 15:51 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T399858|T399858]]) * 15:51 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T399858|T399858]]) * 15:07 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) ([[phab:T399858|T399858]]) * 15:07 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy ([[phab:T399858|T399858]]) * 15:04 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) ([[phab:T399858|T399858]]) * 13:59 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy ([[phab:T399858|T399858]]) === 2025-07-20 === * 10:44 andrewbogott: rebooting cloudcephosd1006 to give us another few days before the memory runs out. [[phab:T399858|T399858]] === 2025-07-19 === * 13:05 andrewbogott: restarted neutron-metadata-agent on cloudnet100[56], again === 2025-07-17 === * 16:02 dcaro: restart cloudcephosd1006 osd 45 for the memory limit take effect * 15:57 dcaro: lowered memory target for cloudcephosd1006 to 5G === 2025-07-16 === * 18:44 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 18:33 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 18:32 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 18:32 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 15:59 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 15:59 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 09:47 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1073.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T399212|T399212]]) * 09:44 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1073.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T399212|T399212]]) === 2025-07-15 === * 23:51 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 23:51 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate === 2025-07-11 === * 18:38 andrewbogott: it didn't * 18:32 andrewbogott: rebooting cloudceph1013 to see if its missing OSD drive reappears * 18:25 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 18:25 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 18:24 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 18:23 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 18:21 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 18:20 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 18:19 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 18:19 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 18:19 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 18:19 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 17:29 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 17:28 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 16:45 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 16:45 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 16:11 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment eqiad1 for service: project,designate * 16:10 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,designate * 15:13 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.upgrade_osds (exit_code=99) * 15:13 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_osds * 15:13 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.upgrade_osds (exit_code=99) * 15:13 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_osds * 15:09 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 15:08 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 00:54 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 00:54 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate === 2025-07-10 === * 23:00 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 22:59 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 22:55 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 22:54 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 21:48 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 21:47 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 15:07 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 15:03 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 15:03 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.reactivate (exit_code=97) * 15:03 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 12:47 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 12:42 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 12:37 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 12:34 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 04:54 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 04:50 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate === 2025-07-09 === * 23:05 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.upgrade_osds (exit_code=99) * 18:01 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_osds ([[phab:T306820|T306820]]) * 15:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.upgrade_mons (exit_code=0) * 15:04 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_mons ([[phab:T306820|T306820]]) * 14:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 14:39 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 14:37 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 14:37 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 13:52 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 13:51 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 01:26 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1073.eqiad.wmnet' ([[phab:T394333|T394333]]) * 01:12 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1073.eqiad.wmnet' ([[phab:T394333|T394333]]) === 2025-07-08 === * 23:11 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 23:10 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 21:11 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 21:11 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 21:11 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 21:11 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 21:09 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 21:09 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 21:06 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 21:06 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 21:06 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 21:05 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 20:54 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 20:53 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 20:52 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 20:52 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 20:51 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 20:50 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 20:50 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 20:50 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 20:49 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 20:48 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 20:46 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 20:45 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 20:44 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 20:44 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 20:44 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 20:35 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 20:33 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 20:33 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 20:31 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 20:31 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 19:58 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 19:57 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 19:55 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 19:54 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 19:51 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 19:51 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 19:51 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 19:38 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 19:36 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 19:36 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 19:34 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 19:34 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 19:33 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 19:33 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 19:28 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 19:28 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 19:24 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 19:24 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 19:23 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 19:22 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 19:21 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 19:20 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 19:20 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 19:19 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 19:17 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 19:17 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 19:15 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 19:15 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 19:13 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 19:12 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 19:12 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 19:11 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 19:10 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 19:10 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 18:56 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 18:56 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 18:55 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 18:55 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 18:54 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 18:53 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 18:51 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 18:51 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 18:48 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 18:48 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 18:47 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 18:46 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 18:45 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 18:45 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 18:45 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 18:44 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 18:42 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 18:42 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 18:20 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 18:19 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 17:19 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 17:19 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 16:00 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 15:04 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 15:04 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 15:03 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 15:02 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 15:02 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 14:59 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 14:59 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 14:58 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 14:58 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 14:58 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.reactivate (exit_code=97) * 14:55 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate === 2025-07-07 === * 21:24 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 21:21 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 21:19 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 21:06 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 21:06 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=97) * 21:03 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 20:38 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 20:38 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 20:37 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 20:37 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 20:37 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 20:36 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 20:36 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 20:36 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy === 2025-07-03 === * 14:42 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,designate * 14:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,designate * 14:15 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol2005-dev.codfw.wmnet' * 14:04 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2005-dev.codfw.wmnet' * 13:49 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudnet2006-dev.codfw.wmnet' * 13:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet2006-dev.codfw.wmnet' * 13:36 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudnet2005-dev.codfw.wmnet' * 13:27 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet2005-dev.codfw.wmnet' * 13:24 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudnet2006-dev.codfw.wmnet' * 13:16 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet2006-dev.codfw.wmnet' === 2025-07-02 === * 19:07 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.upgrade_osds (exit_code=99) * 18:37 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_osds * 18:34 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.upgrade_ceph_node (exit_code=0) * 18:29 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_ceph_node * 18:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.upgrade_ceph_node (exit_code=0) * 18:23 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_ceph_node * 17:16 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 17:15 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 17:04 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 17:03 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 16:52 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 16:52 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 16:36 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.upgrade_ceph_node (exit_code=0) * 16:29 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_ceph_node * 16:29 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.upgrade_ceph_node (exit_code=97) * 16:29 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_ceph_node * 16:13 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 16:13 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 16:13 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 16:12 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 16:10 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 16:10 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 16:09 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 16:09 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:52 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/253 * 12:52 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/253 === 2025-07-01 === * 20:38 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.reboot_node (exit_code=99) * 20:38 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.reboot_node * 18:55 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.upgrade_osds (exit_code=0) * 18:30 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_osds * 16:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 16:07 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 16:05 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 15:57 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 14:17 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 14:16 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2025-06-30 === * 19:39 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/248 * 19:39 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/248 === 2025-06-26 === * 17:56 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 17:49 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 17:25 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 17:15 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 16:14 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 16:02 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 16:01 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 16:01 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 16:00 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 15:51 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 15:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirtlocal1003.eqiad.wmnet<nowiki>}</nowiki>' * 15:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirtlocal1003.eqiad.wmnet<nowiki>}</nowiki>' * 15:24 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirtlocal1002.eqiad.wmnet<nowiki>}</nowiki>' * 15:19 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirtlocal1002.eqiad.wmnet<nowiki>}</nowiki>' * 15:04 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirtlocal1001.eqiad.wmnet<nowiki>}</nowiki>' * 14:58 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirtlocal1001.eqiad.wmnet<nowiki>}</nowiki>' * 14:47 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::virt_ceph<nowiki>}</nowiki>' * 08:53 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 08:52 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 08:51 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/251 * 08:51 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/251 === 2025-06-25 === * 21:22 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=97) * 21:17 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 21:16 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 21:10 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 21:10 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=97) * 21:10 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 20:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 20:36 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 20:35 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=97) * 20:35 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 20:35 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=97) * 20:34 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 20:33 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 20:33 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 20:33 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 20:33 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 20:32 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) * 20:31 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 20:29 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 20:29 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 20:29 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=97) * 20:28 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 20:28 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=97) * 20:28 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 20:27 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=97) * 20:23 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add === 2025-06-24 === * 21:13 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services * 21:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 20:50 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.restart_openstack (exit_code=97) on deployment eqiad1 for all services * 20:28 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 20:27 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services * 20:13 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services === 2025-06-23 === * 22:05 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,neutron * 21:59 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,neutron * 21:58 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for service: project,neutron * 21:58 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for service: project,neutron * 19:21 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for service: project,neutron * 19:20 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for service: project,neutron * 19:18 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,neutron * 19:11 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,neutron * 13:36 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,neutron * 13:31 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,neutron * 13:30 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for service: project,neutron * 13:30 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for service: project,neutron === 2025-06-21 === * 16:15 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 16:10 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 03:17 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services * 03:08 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services === 2025-06-20 === * 20:58 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,nova,cinder,neutron * 20:51 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,nova,cinder,neutron * 17:20 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::virt_ceph<nowiki>}</nowiki>' === 2025-06-19 === * 04:48 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services * 04:40 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services === 2025-06-18 === * 20:39 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services * 20:27 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 17:37 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::virt_ceph<nowiki>}</nowiki>' * 15:02 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::virt_ceph<nowiki>}</nowiki>' === 2025-06-17 === * 20:35 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.upgrade_osds (exit_code=0) * 20:31 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,cinder * 20:31 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,cinder * 20:16 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_osds * 19:55 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.upgrade_osds (exit_code=0) * 19:30 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_osds * 19:19 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.upgrade_osds (exit_code=0) * 19:06 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_osds * 19:05 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.upgrade_osds (exit_code=0) * 18:59 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_osds === 2025-06-14 === * 22:38 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 22:29 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 20:58 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 20:48 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 20:44 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 19:29 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 18:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 18:45 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 18:44 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T309789|T309789]]) * 16:26 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 13:30 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T309789|T309789]]) * 13:30 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1023.eqiad.wmnet' ([[phab:T394727|T394727]]) * 13:30 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1023.eqiad.wmnet' ([[phab:T394727|T394727]]) * 13:28 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 13:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 13:19 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 13:19 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 13:19 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 13:18 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 13:18 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 12:15 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 12:03 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 12:01 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 12:00 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 11:59 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,heat * 11:59 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,heat * 07:29 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T309789|T309789]]) * 05:14 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 05:14 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 05:13 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 05:13 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 05:13 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 05:12 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 04:25 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T309789|T309789]]) * 04:23 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) ([[phab:T309789|T309789]]) * 04:23 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T309789|T309789]]) * 04:21 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 04:20 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 04:02 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 04:02 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 03:59 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 03:58 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 03:57 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T309789|T309789]]) * 03:56 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T309789|T309789]]) * 03:56 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) ([[phab:T309789|T309789]]) * 03:55 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T309789|T309789]]) * 03:54 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 03:54 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 03:51 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.drain_node (exit_code=97) ([[phab:T309789|T309789]]) * 02:09 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/248 * 02:08 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/248 * 02:08 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/245 * 02:07 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/245 * 01:38 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T309789|T309789]]) * 01:37 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 01:37 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 01:36 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 01:36 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 01:36 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=97) * 00:18 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) ([[phab:T309789|T309789]]) === 2025-06-13 === * 21:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 21:46 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 21:44 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 21:11 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 20:23 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 20:22 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 20:22 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 20:22 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 20:20 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 20:20 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 20:13 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T309789|T309789]]) * 20:08 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 20:06 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 20:06 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 20:06 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 20:06 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 20:06 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 19:12 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 19:12 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=97) * 19:12 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 19:12 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=97) * 19:11 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 19:10 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 19:10 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 19:10 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 19:10 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 19:09 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 19:09 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 19:08 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 19:08 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 19:07 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 19:07 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 19:06 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 19:06 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 19:06 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 18:42 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 16:15 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 15:50 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T309789|T309789]]) * 14:04 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 14:04 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 14:04 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 14:03 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=97) * 13:53 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 11:46 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T309789|T309789]]) * 09:12 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) * 06:31 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) ([[phab:T309789|T309789]]) * 04:07 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T309789|T309789]]) * 04:01 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 04:01 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 04:00 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 03:51 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 02:56 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 02:56 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) * 02:49 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 02:48 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 02:34 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 02:34 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 02:33 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 02:32 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 01:00 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 00:59 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) * 00:58 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 00:57 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.drain_node (exit_code=97) ([[phab:T309789|T309789]]) * 00:36 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T309789|T309789]]) * 00:18 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 00:18 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) === 2025-06-12 === * 23:15 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/248 * 23:15 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/248 * 23:14 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/248 * 23:14 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/248 * 21:40 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 21:40 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 21:39 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 21:32 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 20:40 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.drain_node (exit_code=97) * 20:35 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 20:31 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 20:31 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 20:27 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 20:27 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 20:18 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 20:18 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 20:17 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 19:53 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T396363|T396363]]) * 19:52 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T396363|T396363]]) * 19:51 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 19:49 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) * 19:49 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T309789|T309789]]) * 19:33 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T396363|T396363]]) * 19:33 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T396363|T396363]]) * 19:18 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T309789|T309789]]) * 17:51 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T309789|T309789]]) * 17:12 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) ([[phab:T396363|T396363]]) * 16:36 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 15:12 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 15:11 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 15:10 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 14:58 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 14:58 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=97) * 14:53 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 14:52 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 14:52 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 14:25 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 14:24 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:11 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T396363|T396363]]) * 11:41 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T309789|T309789]]) * 05:23 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T309789|T309789]]) * 02:19 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T309789|T309789]]) === 2025-06-11 === * 21:58 andrewbogott: created new keystone role in eqiad1 and codfw1dev, 'object_storage' [[phab:T396594|T396594]] * 16:51 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for service: project,cinder * 16:51 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for service: project,cinder * 15:31 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services * 15:20 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 15:20 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.restart_openstack (exit_code=97) on deployment eqiad1 for all services * 15:11 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 14:59 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,cinder * 14:59 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,cinder === 2025-06-10 === * 16:07 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 16:06 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 16:06 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 16:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 16:05 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 16:04 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 16:04 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 16:03 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 16:02 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/245 * 16:02 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/245 * 16:01 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/245 * 16:01 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/245 * 16:01 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 16:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 15:59 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/246 * 15:59 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/246 * 15:57 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/246 * 15:57 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/246 * 15:40 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/246 * 15:40 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/246 * 15:34 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/245 * 15:34 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/245 * 10:16 taavi: add cloud-private addresses to eqiad hosts [[phab:T379283|T379283]] === 2025-06-09 === * 20:11 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 20:09 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 19:42 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 19:39 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 19:23 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 19:20 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 19:15 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 19:13 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 19:13 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.restart_openstack (exit_code=97) on deployment codfw1dev for service: project,designate * 19:13 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for service: project,designate * 19:11 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for service: project,designate * 19:09 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for service: project,designate * 19:09 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for service: project,designate * 19:09 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for service: project,designate * 19:09 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for service: project,keystone * 19:09 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for service: project,keystone * 19:08 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 19:08 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 19:08 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 19:08 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 19:07 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 19:07 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 19:07 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 19:07 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 19:07 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 19:06 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 19:05 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 19:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 19:05 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 19:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 19:05 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 19:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 19:03 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 19:03 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 13:14 taavi: add AAAA record to openstack.codfw1dev.wikimediacloud.org [[phab:T379282|T379282]] * 13:13 taavi: add AAAA record to openstack.codfw1dev.wikimediacloud.org [[phab:T347148|T347148]] === 2025-06-07 === * 19:08 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services * 18:55 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 18:49 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,nova * 18:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,nova === 2025-06-06 === * 20:59 andrewbogott: restarting all designate services on all cloudcontrols in eqiad1 === 2025-06-05 === * 20:26 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) * 17:31 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 17:29 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 17:29 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for service: project,octavia * 17:28 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for service: project,octavia * 17:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for service: project,octavia * 17:28 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for service: project,octavia * 15:25 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 14:52 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=97) * 14:41 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node === 2025-06-03 === * 22:08 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 22:07 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 22:07 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/244 * 22:07 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/244 * 22:07 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/244 * 22:06 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/244 * 20:12 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 20:08 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 19:41 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 19:36 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 19:27 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 19:22 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 18:51 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 18:45 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 18:44 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=97) * 18:44 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 18:43 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=97) * 18:43 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 18:34 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T395910|T395910]]) * 18:29 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T395910|T395910]]) * 17:55 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T395910|T395910]]) * 17:50 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T395910|T395910]]) * 17:42 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T395910|T395910]]) * 17:37 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T395910|T395910]]) * 17:36 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=97) * 17:36 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add === 2025-06-02 === * 21:35 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 21:34 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 21:33 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/237 * 21:33 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/237 * 17:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 17:46 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 17:45 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/240 * 17:45 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/240 * 17:44 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/240 * 17:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/240 * 17:39 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/240 * 17:38 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/240 * 17:36 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/240 * 17:35 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/240 * 17:29 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/240 * 17:29 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/240 * 17:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/240 * 17:28 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/240 * 17:24 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/240 * 17:24 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/240 * 17:20 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/240 * 17:20 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/240 * 17:19 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/242 * 17:19 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/242 * 16:36 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/241 * 16:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/241 * 16:36 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/241 * 16:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/241 * 16:35 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/241 * 16:35 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/241 * 16:33 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/240 * 16:32 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/240 * 16:30 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/240 * 16:29 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/240 * 16:29 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/240 * 16:29 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/240 * 16:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/240 * 16:28 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/240 * 16:26 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/240 * 16:25 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/240 * 16:25 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/240 * 16:25 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/240 * 15:49 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 15:49 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 15:44 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 15:38 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 15:36 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/239 * 15:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/239 * 15:32 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/239 * 15:32 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/239 * 00:27 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T394333|T394333]]) === 2025-06-01 === * 18:31 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T394333|T394333]]) === 2025-05-31 === * 23:51 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services ([[phab:T395742|T395742]]) * 23:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services ([[phab:T395742|T395742]]) * 23:38 andrewbogott: failing over from cloudnet1005 to 1006 in hopes of unsticking [[phab:T395742|T395742]] * 23:36 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,neutron ([[phab:T395742|T395742]]) * 23:24 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,neutron ([[phab:T395742|T395742]]) === 2025-05-30 === * 14:40 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 14:39 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2025-05-28 === * 22:12 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1076.eqiad.wmnet' ([[phab:T390914|T390914]]) * 22:07 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1076.eqiad.wmnet' ([[phab:T390914|T390914]]) * 22:07 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1075.eqiad.wmnet' ([[phab:T390914|T390914]]) * 22:01 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1075.eqiad.wmnet' ([[phab:T390914|T390914]]) * 22:01 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1074.eqiad.wmnet' ([[phab:T390914|T390914]]) * 21:54 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1074.eqiad.wmnet' ([[phab:T390914|T390914]]) * 21:54 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1073.eqiad.wmnet' ([[phab:T390914|T390914]]) * 21:48 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1073.eqiad.wmnet' ([[phab:T390914|T390914]]) * 21:48 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1072.eqiad.wmnet' ([[phab:T390914|T390914]]) * 21:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1072.eqiad.wmnet' ([[phab:T390914|T390914]]) * 21:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1071.eqiad.wmnet' ([[phab:T390914|T390914]]) * 21:37 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1071.eqiad.wmnet' ([[phab:T390914|T390914]]) * 21:37 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1070.eqiad.wmnet' ([[phab:T390914|T390914]]) * 21:30 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1070.eqiad.wmnet' ([[phab:T390914|T390914]]) * 21:30 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1069.eqiad.wmnet' ([[phab:T390914|T390914]]) * 21:23 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1069.eqiad.wmnet' ([[phab:T390914|T390914]]) * 21:23 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1068.eqiad.wmnet' ([[phab:T390914|T390914]]) * 21:17 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1068.eqiad.wmnet' ([[phab:T390914|T390914]]) * 21:17 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1067.eqiad.wmnet' ([[phab:T390914|T390914]]) * 21:10 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1067.eqiad.wmnet' ([[phab:T390914|T390914]]) * 21:09 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1066.eqiad.wmnet' ([[phab:T390914|T390914]]) * 21:02 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1066.eqiad.wmnet' ([[phab:T390914|T390914]]) * 21:02 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1065.eqiad.wmnet' ([[phab:T390914|T390914]]) * 20:55 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1065.eqiad.wmnet' ([[phab:T390914|T390914]]) * 20:55 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1064.eqiad.wmnet' ([[phab:T390914|T390914]]) * 20:49 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1064.eqiad.wmnet' ([[phab:T390914|T390914]]) * 20:49 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1063.eqiad.wmnet' ([[phab:T390914|T390914]]) * 20:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1063.eqiad.wmnet' ([[phab:T390914|T390914]]) * 20:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1062.eqiad.wmnet' ([[phab:T390914|T390914]]) * 20:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1062.eqiad.wmnet' ([[phab:T390914|T390914]]) * 20:36 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1061.eqiad.wmnet' ([[phab:T390914|T390914]]) * 20:30 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1061.eqiad.wmnet' ([[phab:T390914|T390914]]) * 20:29 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1060.eqiad.wmnet' ([[phab:T390914|T390914]]) * 20:23 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1060.eqiad.wmnet' ([[phab:T390914|T390914]]) * 20:23 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1059.eqiad.wmnet' ([[phab:T390914|T390914]]) * 20:16 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1059.eqiad.wmnet' ([[phab:T390914|T390914]]) * 20:16 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1058.eqiad.wmnet' ([[phab:T390914|T390914]]) * 20:13 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services * 20:09 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1058.eqiad.wmnet' ([[phab:T390914|T390914]]) * 20:09 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1057.eqiad.wmnet' ([[phab:T390914|T390914]]) * 20:03 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1057.eqiad.wmnet' ([[phab:T390914|T390914]]) * 20:02 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1056.eqiad.wmnet' ([[phab:T390914|T390914]]) * 20:01 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 19:55 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1056.eqiad.wmnet' ([[phab:T390914|T390914]]) * 19:55 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1055.eqiad.wmnet' ([[phab:T390914|T390914]]) * 19:49 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1055.eqiad.wmnet' ([[phab:T390914|T390914]]) * 19:49 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1054.eqiad.wmnet' ([[phab:T390914|T390914]]) * 19:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudbackup1002-dev.eqiad.wmnet' ([[phab:T390914|T390914]]) * 19:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1054.eqiad.wmnet' ([[phab:T390914|T390914]]) * 19:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1053.eqiad.wmnet' ([[phab:T390914|T390914]]) * 19:37 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudbackup1002-dev.eqiad.wmnet' ([[phab:T390914|T390914]]) * 19:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1053.eqiad.wmnet' ([[phab:T390914|T390914]]) * 19:36 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1052.eqiad.wmnet' ([[phab:T390914|T390914]]) * 19:35 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudbackup1001-dev.eqiad.wmnet' ([[phab:T390914|T390914]]) * 19:30 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1052.eqiad.wmnet' ([[phab:T390914|T390914]]) * 19:30 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1051.eqiad.wmnet' ([[phab:T390914|T390914]]) * 19:29 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudbackup1001-dev.eqiad.wmnet' ([[phab:T390914|T390914]]) * 19:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudbackup1004.eqiad.wmnet' ([[phab:T390914|T390914]]) * 19:23 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1051.eqiad.wmnet' ([[phab:T390914|T390914]]) * 19:23 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1050.eqiad.wmnet' ([[phab:T390914|T390914]]) * 19:18 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudbackup1004.eqiad.wmnet' ([[phab:T390914|T390914]]) * 19:18 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudbackup1003.eqiad.wmnet' ([[phab:T390914|T390914]]) * 19:17 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1050.eqiad.wmnet' ([[phab:T390914|T390914]]) * 19:17 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1049.eqiad.wmnet' ([[phab:T390914|T390914]]) * 19:10 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1049.eqiad.wmnet' ([[phab:T390914|T390914]]) * 19:10 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1048.eqiad.wmnet' ([[phab:T390914|T390914]]) * 19:08 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudbackup1003.eqiad.wmnet' ([[phab:T390914|T390914]]) * 19:07 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirtlocal1003.eqiad.wmnet' ([[phab:T390914|T390914]]) * 19:04 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1048.eqiad.wmnet' ([[phab:T390914|T390914]]) * 19:03 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1047.eqiad.wmnet' ([[phab:T390914|T390914]]) * 19:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirtlocal1003.eqiad.wmnet' ([[phab:T390914|T390914]]) * 18:57 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirtlocal1002.eqiad.wmnet' ([[phab:T390914|T390914]]) * 18:57 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1047.eqiad.wmnet' ([[phab:T390914|T390914]]) * 18:57 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1046.eqiad.wmnet' ([[phab:T390914|T390914]]) * 18:51 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1046.eqiad.wmnet' ([[phab:T390914|T390914]]) * 18:51 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1045.eqiad.wmnet' ([[phab:T390914|T390914]]) * 18:50 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirtlocal1002.eqiad.wmnet' ([[phab:T390914|T390914]]) * 18:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1045.eqiad.wmnet' ([[phab:T390914|T390914]]) * 18:44 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1044.eqiad.wmnet' ([[phab:T390914|T390914]]) * 18:44 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirtlocal1001.eqiad.wmnet' ([[phab:T390914|T390914]]) * 18:38 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1044.eqiad.wmnet' ([[phab:T390914|T390914]]) * 18:38 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1043.eqiad.wmnet' ([[phab:T390914|T390914]]) * 18:37 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirtlocal1001.eqiad.wmnet' ([[phab:T390914|T390914]]) * 18:31 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1043.eqiad.wmnet' ([[phab:T390914|T390914]]) * 18:31 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1042.eqiad.wmnet' ([[phab:T390914|T390914]]) * 18:24 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1042.eqiad.wmnet' ([[phab:T390914|T390914]]) * 18:22 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1041.eqiad.wmnet' ([[phab:T390914|T390914]]) * 18:15 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1041.eqiad.wmnet' ([[phab:T390914|T390914]]) * 18:11 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1040.eqiad.wmnet' ([[phab:T390914|T390914]]) * 18:04 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1040.eqiad.wmnet' ([[phab:T390914|T390914]]) * 18:03 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.restart_openstack (exit_code=97) on deployment eqiad1 for all services * 17:53 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 17:52 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudnet1006.eqiad.wmnet' ([[phab:T390914|T390914]]) * 17:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet1006.eqiad.wmnet' ([[phab:T390914|T390914]]) * 17:42 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudnet1005.eqiad.wmnet' ([[phab:T390914|T390914]]) * 17:33 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet1005.eqiad.wmnet' ([[phab:T390914|T390914]]) * 17:30 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudrabbit1001.eqiad.wmnet' ([[phab:T390914|T390914]]) * 17:21 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudrabbit1001.eqiad.wmnet' ([[phab:T390914|T390914]]) * 17:21 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudrabbit1002.eqiad.wmnet' ([[phab:T390914|T390914]]) * 17:13 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudrabbit1002.eqiad.wmnet' ([[phab:T390914|T390914]]) * 17:12 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudrabbit1003.eqiad.wmnet' ([[phab:T390914|T390914]]) * 17:04 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudrabbit1003.eqiad.wmnet' ([[phab:T390914|T390914]]) * 17:02 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudrabbot1003.eqiad.wmnet' ([[phab:T390914|T390914]]) * 17:02 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudrabbot1003.eqiad.wmnet' ([[phab:T390914|T390914]]) * 17:00 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol1006.eqiad.wmnet' ([[phab:T390914|T390914]]) * 16:42 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1006.eqiad.wmnet' ([[phab:T390914|T390914]]) * 16:42 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol1007.eqiad.wmnet' ([[phab:T390914|T390914]]) * 16:20 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1007.eqiad.wmnet' ([[phab:T390914|T390914]]) * 16:19 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol1011.eqiad.wmnet' ([[phab:T390914|T390914]]) * 16:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1011.eqiad.wmnet' ([[phab:T390914|T390914]]) * 15:57 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudcontrol1011.eqiad.wmnet' ([[phab:T390914|T390914]]) * 15:52 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1011.eqiad.wmnet' ([[phab:T390914|T390914]]) * 15:48 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudservices1006.eqiad.wmnet' ([[phab:T390914|T390914]]) * 15:40 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1006.eqiad.wmnet' ([[phab:T390914|T390914]]) * 15:39 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T390914|T390914]]) * 15:32 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T390914|T390914]]) * 15:32 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005-dev.eqiad.wmnet' ([[phab:T390914|T390914]]) * 15:31 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005-dev.eqiad.wmnet' ([[phab:T390914|T390914]]) * 15:29 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T390914|T390914]]) * 15:22 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T390914|T390914]]) * 15:21 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005-dev.eqiad.wmnet' ([[phab:T390914|T390914]]) * 15:20 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005-dev.eqiad.wmnet' ([[phab:T390914|T390914]]) * 15:19 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005-dev.codfw.wmnet' ([[phab:T390914|T390914]]) * 15:19 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005-dev.codfw.wmnet' ([[phab:T390914|T390914]]) * 10:40 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for service: project,designate * 10:39 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for service: project,designate === 2025-05-26 === * 13:13 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudnet.reboot_node (exit_code=0) for host cloudnet2006-dev.codfw.wmnet * 13:10 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudnet.reboot_node for host cloudnet2006-dev.codfw.wmnet * 13:06 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudnet.reboot_node (exit_code=0) for host cloudnet2005-dev.codfw.wmnet * 13:03 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudnet.reboot_node for host cloudnet2005-dev.codfw.wmnet * 12:54 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.roll_reboot_cloudnets (exit_code=99) * 12:54 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.roll_reboot_cloudnets * 12:49 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.roll_reboot_cloudnets (exit_code=99) * 12:49 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.roll_reboot_cloudnets * 12:48 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.reboot_node (exit_code=0) on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::codfw1dev::control<nowiki>}</nowiki>' * 12:36 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.reboot_node on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::codfw1dev::control<nowiki>}</nowiki>' * 12:31 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'P<nowiki>{</nowiki>P:openstack::codfw1dev::nova::compute::service<nowiki>}</nowiki>' * 11:49 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>P:openstack::codfw1dev::nova::compute::service<nowiki>}</nowiki>' * 11:33 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'P<nowiki>{</nowiki>P:openstack::codfw1dev::nova::compute::service<nowiki>}</nowiki>' * 11:27 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>P:openstack::codfw1dev::nova::compute::service<nowiki>}</nowiki>' * 11:22 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'P<nowiki>{</nowiki>P:openstack::codfw1dev::nova::compute::service<nowiki>}</nowiki>' * 11:21 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>P:openstack::codfw1dev::nova::compute::service<nowiki>}</nowiki>' * 11:18 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'P<nowiki>{</nowiki>P:openstack::codfw1dev::nova::compute::service<nowiki>}</nowiki>' * 11:17 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>P:openstack::codfw1dev::nova::compute::service<nowiki>}</nowiki>' * 07:52 taavi: add new /dev/sdb back to software raid after it was replaced in cloudcephmon1004 === 2025-05-23 === * 09:00 dhinus: failover dumps_dist_active_vps to clouddumps1001 ([[phab:T383723|T383723]]) === 2025-05-20 === * 19:45 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1037.eqiad.wmnet' ([[phab:T394727|T394727]]) * 19:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1037.eqiad.wmnet' ([[phab:T394727|T394727]]) * 19:42 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1039.eqiad.wmnet' ([[phab:T394727|T394727]]) * 19:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1039.eqiad.wmnet' ([[phab:T394727|T394727]]) * 15:22 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services * 15:07 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 14:57 andrewbogott: resetting eqiad1 rabbitmq in an attempt to resolve [[phab:T394790|T394790]] * 03:01 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,cinder * 03:01 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,cinder * 02:18 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,cinder * 02:18 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,cinder * 00:21 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1039.eqiad.wmnet' ([[phab:T394727|T394727]]) * 00:11 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1039.eqiad.wmnet' ([[phab:T394727|T394727]]) === 2025-05-19 === * 22:58 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1039.eqiad.wmnet' ([[phab:T394727|T394727]]) * 22:47 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1039.eqiad.wmnet' ([[phab:T394727|T394727]]) * 22:43 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.restart_openstack (exit_code=97) on deployment eqiad1 for service: project,nova * 22:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,nova * 22:35 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,nova * 22:29 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,nova * 22:26 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.restart_openstack (exit_code=97) on deployment eqiad1 for service: project,nova * 22:26 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,nova * 22:06 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1039.eqiad.wmnet' ([[phab:T394727|T394727]]) * 21:56 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1039.eqiad.wmnet' ([[phab:T394727|T394727]]) * 21:47 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1037.eqiad.wmnet' ([[phab:T394727|T394727]]) * 21:37 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1037.eqiad.wmnet' ([[phab:T394727|T394727]]) * 21:28 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1039.eqiad.wmnet' ([[phab:T394727|T394727]]) * 21:17 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1039.eqiad.wmnet' ([[phab:T394727|T394727]]) * 21:05 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1039.eqiad.wmnet' ([[phab:T394727|T394727]]) * 20:37 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1039.eqiad.wmnet' ([[phab:T394727|T394727]]) * 19:59 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1038.eqiad.wmnet' ([[phab:T394727|T394727]]) * 19:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1038.eqiad.wmnet' ([[phab:T394727|T394727]]) * 19:38 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1037.eqiad.wmnet' ([[phab:T394727|T394727]]) * 19:28 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1037.eqiad.wmnet' ([[phab:T394727|T394727]]) * 19:28 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1037.eqiad.wmnet' ([[phab:T394727|T394727]]) * 18:58 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1037.eqiad.wmnet' ([[phab:T394727|T394727]]) * 18:52 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1036.eqiad.wmnet' ([[phab:T394727|T394727]]) * 18:51 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1036.eqiad.wmnet' ([[phab:T394727|T394727]]) * 18:51 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 18:51 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 18:51 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) * 18:50 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 18:50 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=99) * 18:50 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 18:48 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=99) * 18:48 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 18:47 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1035.eqiad.wmnet' ([[phab:T394727|T394727]]) * 18:47 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1035.eqiad.wmnet' ([[phab:T394727|T394727]]) * 18:47 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1034.eqiad.wmnet' ([[phab:T394727|T394727]]) * 18:46 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1034.eqiad.wmnet' ([[phab:T394727|T394727]]) * 18:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1033.eqiad.wmnet' ([[phab:T394727|T394727]]) * 18:45 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1033.eqiad.wmnet' ([[phab:T394727|T394727]]) * 18:45 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1032.eqiad.wmnet' ([[phab:T394727|T394727]]) * 18:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1032.eqiad.wmnet' ([[phab:T394727|T394727]]) * 18:44 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1031.eqiad.wmnet' ([[phab:T394727|T394727]]) * 18:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1031.eqiad.wmnet' ([[phab:T394727|T394727]]) * 17:58 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1075.eqiad.wmnet<nowiki>}</nowiki>' * 17:55 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1075.eqiad.wmnet<nowiki>}</nowiki>' * 17:52 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1072.eqiad.wmnet<nowiki>}</nowiki>' * 17:47 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1072.eqiad.wmnet<nowiki>}</nowiki>' * 17:35 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1076.eqiad.wmnet<nowiki>}</nowiki>' * 17:32 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1076.eqiad.wmnet<nowiki>}</nowiki>' * 17:31 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1073.eqiad.wmnet<nowiki>}</nowiki>' * 17:27 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1073.eqiad.wmnet<nowiki>}</nowiki>' * 15:24 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 15:23 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2025-05-15 === * 15:49 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/235 * 15:49 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/235 * 15:48 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/235 * 15:48 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/235 * 14:53 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/234 * 14:53 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/234 * 14:49 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/234 * 14:48 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/234 * 13:27 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 13:26 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:58 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 12:57 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2025-05-14 === * 17:50 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 17:49 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 17:49 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 17:49 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 17:48 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/231 * 17:48 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/231 * 12:59 taavi: powercycle unresponsive cloudnet2006-dev === 2025-05-12 === * 20:07 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services * 19:59 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 08:17 taavi: powercycle clouservices2005-dev.codfw.wmnet * 03:15 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 03:07 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 02:36 andrewbogott: rebooting cloudnet2005-dev from mgmt -- ssh is failing and the console shows a user prompt but not a password prompt. === 2025-05-11 === * 13:42 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,nova * 13:37 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,nova === 2025-05-07 === * 20:55 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudbackup1002-dev.eqiad.wmnet' ([[phab:T390914|T390914]]) * 20:50 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudbackup1002-dev.eqiad.wmnet' ([[phab:T390914|T390914]]) * 20:50 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudbackup1001-dev.eqiad.wmnet' ([[phab:T390914|T390914]]) * 20:45 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudbackup1001-dev.eqiad.wmnet' ([[phab:T390914|T390914]]) * 20:13 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt2006-dev.codfw.wmnet' ([[phab:T390914|T390914]]) * 20:08 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt2006-dev.codfw.wmnet' ([[phab:T390914|T390914]]) * 20:07 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt2005-dev.codfw.wmnet' ([[phab:T390914|T390914]]) * 20:02 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt2005-dev.codfw.wmnet' ([[phab:T390914|T390914]]) * 20:02 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt2004-dev.codfw.wmnet' ([[phab:T390914|T390914]]) * 20:01 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for service: project,designate * 19:59 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for service: project,designate * 19:57 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt2004-dev.codfw.wmnet' ([[phab:T390914|T390914]]) * 19:51 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudnet2005-dev.codfw.wmnet' ([[phab:T390914|T390914]]) * 19:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet2005-dev.codfw.wmnet' ([[phab:T390914|T390914]]) * 19:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudnet2006-dev.codfw.wmnet' ([[phab:T390914|T390914]]) * 19:34 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet2006-dev.codfw.wmnet' ([[phab:T390914|T390914]]) * 19:34 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol2006-dev.codfw.wmnet' ([[phab:T390914|T390914]]) * 19:19 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2006-dev.codfw.wmnet' ([[phab:T390914|T390914]]) * 19:19 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol2005-dev.codfw.wmnet' ([[phab:T390914|T390914]]) * 19:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2005-dev.codfw.wmnet' ([[phab:T390914|T390914]]) * 18:58 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol2004-dev.codfw.wmnet' ([[phab:T390914|T390914]]) * 18:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2004-dev.codfw.wmnet' ([[phab:T390914|T390914]]) * 18:39 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudservices2005-dev.codfw.wmnet' ([[phab:T390914|T390914]]) * 18:32 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices2005-dev.codfw.wmnet' ([[phab:T390914|T390914]]) * 18:32 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudservices2004-dev.codfw.wmnet' ([[phab:T390914|T390914]]) * 18:24 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices2004-dev.codfw.wmnet' ([[phab:T390914|T390914]]) * 18:23 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices2004.codfw.wmnet' ([[phab:T390914|T390914]]) * 18:23 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices2004.codfw.wmnet' ([[phab:T390914|T390914]]) * 18:23 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices2004.eqiad.wmnet' ([[phab:T390914|T390914]]) * 18:23 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices2004.eqiad.wmnet' ([[phab:T390914|T390914]]) * 18:19 andrewbogott: upgrading codfw1dev to version 'epoxy' [[phab:T390914|T390914]] * 18:18 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudweb.set_maintenance (exit_code=0) ([[phab:T390914|T390914]]) * 18:17 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudweb.set_maintenance ([[phab:T390914|T390914]]) * 16:31 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,heat * 16:30 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,heat * 16:30 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,magnum * 16:30 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,magnum * 16:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,magnum,heat * 16:28 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,magnum,heat * 16:24 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,magnum,heat * 16:23 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,magnum,heat * 16:19 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,heat * 16:19 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,heat * 16:17 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,heat * 16:17 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,heat * 16:16 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,magnum * 16:16 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,magnum * 15:45 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for service: project,nova * 15:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for service: project,nova * 15:43 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for service: project,nova * 15:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for service: project,nova * 15:43 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for service: project,nova * 15:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for service: project,nova * 15:42 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for service: project,nova * 15:42 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for service: project,nova * 13:35 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,neutron * 13:24 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,neutron * 13:14 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for service: project,neutron * 13:13 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for service: project,neutron * 12:48 taavi: updating all security group rules referencing old 172.16.0.0/21 subnet to reference new ip space instead ([[phab:T379175|T379175]]) * 10:49 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 10:48 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 00:47 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 00:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services === 2025-05-06 === * 03:13 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) === 2025-05-05 === * 15:00 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 12:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 10:19 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 08:19 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 03:19 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 00:08 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 00:08 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 00:07 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) * 00:07 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 00:06 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 00:06 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 00:05 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 00:05 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 00:03 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 00:03 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 00:03 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=97) * 00:03 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 00:02 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 00:02 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add === 2025-05-04 === * 23:53 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 23:53 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 22:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) * 22:43 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 22:42 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) * 22:41 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 22:41 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) * 22:40 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 22:39 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) * 22:38 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 22:37 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) * 22:36 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 21:31 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 21:31 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 21:31 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 21:31 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 21:31 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=97) * 21:30 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 21:30 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 21:30 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node === 2025-05-03 === * 02:21 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) ([[phab:T393196|T393196]]) === 2025-05-02 === * 23:04 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy ([[phab:T393196|T393196]]) * 21:51 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) ([[phab:T393196|T393196]]) * 18:55 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy ([[phab:T393196|T393196]]) * 18:55 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) ([[phab:T393196|T393196]]) * 16:20 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy ([[phab:T393196|T393196]]) * 16:20 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=97) ([[phab:T393196|T393196]]) * 16:14 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy ([[phab:T393196|T393196]]) * 16:14 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 16:14 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 15:09 andrewbogott: sudo cumin --force O<nowiki>{</nowiki>*<nowiki>}</nowiki> "dpkg --list {{!}} grep puppetserver && systemctl restart puppetserver.service" # work around a package update that is causing some puppetservers to error out === 2025-04-29 === * 13:01 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 13:01 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:49 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/225 * 12:48 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/225 * 11:48 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 11:47 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 11:15 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/224 * 11:14 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/224 * 11:13 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/224 * 11:13 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/224 * 11:11 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/224 * 11:11 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/224 * 11:10 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/224 * 11:10 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/224 * 11:09 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/224 * 11:09 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/224 * 11:05 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/224 * 11:05 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/224 * 10:55 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/224 * 10:55 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/224 * 10:52 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/224 * 10:51 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/224 * 08:16 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 08:15 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 07:46 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/223 * 07:45 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/223 * 07:43 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 07:43 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 07:42 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 07:41 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 07:41 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/221 * 07:40 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/221 === 2025-04-28 === * 15:49 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/221 * 15:49 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/221 * 11:25 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 11:21 taavi: migrating opentofu managed default security group rules [[phab:T392799|T392799]] * 11:20 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 08:13 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 08:13 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2025-04-26 === * 07:48 dcaro@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.unset_cluster_maintenance (exit_code=0) ([[phab:T390134|T390134]]) * 07:48 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.unset_cluster_maintenance ([[phab:T390134|T390134]]) * 07:48 dcaro@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) ([[phab:T390134|T390134]]) * 07:47 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T390134|T390134]]) === 2025-04-25 === * 12:55 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/218 * 12:54 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/218 * 12:30 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 12:30 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 04:29 dcaro@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T390134|T390134]]) === 2025-04-24 === * 17:25 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T390134|T390134]]) * 13:29 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 13:29 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 07:56 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment eqiad1 for service: project,designate * 07:55 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,designate === 2025-04-23 === * 21:45 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.drain_node (exit_code=97) * 21:45 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 21:45 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=97) * 21:44 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 21:26 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) * 21:26 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 21:25 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) * 21:25 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 21:24 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) * 21:23 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 21:23 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) * 21:23 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 21:22 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) * 21:21 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 21:21 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.drain_node (exit_code=97) * 21:20 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 19:15 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=97) * 19:03 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 19:03 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=97) * 16:32 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 16:32 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.drain_node (exit_code=97) * 16:32 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 16:31 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.drain_node (exit_code=97) * 16:29 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 16:29 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.drain_node (exit_code=97) * 16:29 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 16:29 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=97) * 16:18 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 16:17 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=97) * 16:17 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 14:29 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 14:28 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 14:27 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 14:27 arturo: enabling IPv6 dualstack on neutron virtual router ([[phab:T380174|T380174]]) * 14:27 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 14:22 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/217 * 14:22 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/217 * 12:15 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 12:15 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 12:12 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 12:12 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:08 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 12:08 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 12:08 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 12:07 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:04 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 12:03 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 12:03 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 12:02 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:00 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 11:59 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 11:57 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 11:57 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 11:54 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 11:54 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 11:54 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 11:53 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 11:52 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 11:52 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 11:49 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 11:48 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 11:45 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 11:45 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 11:44 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 11:42 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 11:30 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 11:29 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 09:14 arturo: enable IPv6 on cloudgw ([[phab:T380174|T380174]]) -- includes server reboot === 2025-04-22 === * 04:35 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) === 2025-04-21 === * 23:34 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 23:34 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=97) * 23:18 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 23:18 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=97) * 23:18 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 22:41 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 22:11 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 22:09 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 22:09 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 22:09 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=97) * 22:09 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 21:53 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 16:43 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 16:16 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 16:10 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 11:05 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 11:03 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 01:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 01:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services === 2025-04-16 === * 16:49 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 16:49 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 16:02 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,cinder * 16:02 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,cinder * 13:58 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,cinder * 13:58 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,cinder * 12:34 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 12:33 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:26 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 12:26 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:26 arturo: merging network change in neutron https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/198 * 10:10 dcaro: upgrade spicerack on cloudcumin2001 to 10.1.0 * 10:08 dcaro@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) * 10:07 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 10:07 dcaro: upgrade spicerack on cloudcumin1001 to 10.1.0 === 2025-04-15 === * 15:17 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 15:16 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 15:16 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 15:15 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 15:13 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 15:13 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:49 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 12:49 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:38 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 12:37 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:31 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 12:30 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2025-04-14 === * 17:19 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 17:18 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 16:50 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 16:50 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 16:50 andrewbogott: granting 'tofuadmin' user inherited 'member' role in all projects, this should fix some policy mishaps * 16:21 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 16:21 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 16:19 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 16:19 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 16:17 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 16:16 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 16:13 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 16:12 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2025-04-13 === * 22:49 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 22:46 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services === 2025-04-11 === * 11:58 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 11:57 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 11:50 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 11:50 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 11:50 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 11:49 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 10:07 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 10:06 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2025-04-10 === * 12:38 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 12:37 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 09:57 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 09:57 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 09:05 arturo: [codfw1dev] root@cloudcontrol2004-dev:~# wmcs-makedomain --project testlabs --domain testlabs.codfw1dev.wmcloud.org --orig-project cloudinfra-codfw1dev ([[phab:T391325|T391325]]) === 2025-04-09 === * 23:23 bd808: Rebooting tools-sgebastion-10 (login-buster.toolforge.org) for high load/unresponsive NFS mounts ([[phab:T391538|T391538]]) * 10:18 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 10:17 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 10:08 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 10:06 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 08:39 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 08:38 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 04:09 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services * 03:59 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 03:55 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment eqiad1 for all services * 03:54 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 03:31 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment eqiad1 for all services * 03:30 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 03:30 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.restart_openstack (exit_code=97) on deployment eqiad1 for all services * 03:29 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 03:24 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment eqiad1 for all services * 03:23 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 03:19 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment eqiad1 for all services * 03:17 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services === 2025-04-08 === * 22:18 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol1011.eqiad.wmnet' * 22:07 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1011.eqiad.wmnet' * 22:00 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudcontrol1011.eqiad.wmnet' * 22:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1011.eqiad.wmnet' * 11:14 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 11:13 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch === 2025-04-07 === * 23:43 bd808: `sudo service maintain-dbusers restart` on cloudcontrol1007 after reports of missing replica.my.cnf and finding the journal for the service empty. * 15:30 arturo: [codfw1dev] testlabs create a bunch of VMs by hand, like `networktests-vlan-legacy-floating` [[phab:T380728|T380728]] * 15:21 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 15:20 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 14:09 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 14:09 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 14:06 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 14:05 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 13:58 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 13:57 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 11:30 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 11:29 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 11:29 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 11:29 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 11:24 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 11:24 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 11:23 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 11:23 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2025-04-06 === * 02:20 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 02:18 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services === 2025-04-04 === * 15:14 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 15:13 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 13:13 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 13:13 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 05:12 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 05:10 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 05:10 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.restart_openstack (exit_code=97) on deployment codfw1dev for all services * 05:08 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 05:04 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.restart_openstack (exit_code=97) on deployment codfw1dev for all services * 05:02 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services === 2025-04-03 === * 22:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services * 22:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 22:07 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirtlocal1001.eqiad.wmnet' ([[phab:T381499|T381499]]) * 22:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirtlocal1001.eqiad.wmnet' ([[phab:T381499|T381499]]) * 21:56 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirtlocal1002.eqiad.wmnet' ([[phab:T381499|T381499]]) * 21:54 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2004-dev.codfw.wmnet<nowiki>}</nowiki>' * 21:49 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2004-dev.codfw.wmnet<nowiki>}</nowiki>' * 21:49 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirtlocal1002.eqiad.wmnet' ([[phab:T381499|T381499]]) * 21:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirtlocal1003.eqiad.wmnet' ([[phab:T381499|T381499]]) * 21:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirtlocal1003.eqiad.wmnet' ([[phab:T381499|T381499]]) * 20:12 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 20:12 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.restart_openstack (exit_code=97) on deployment eqiad1 for all services * 20:11 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 20:11 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.restart_openstack (exit_code=97) on deployment eqiad1 for all services * 20:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 19:59 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment eqiad1 for all services * 19:59 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 19:54 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudrabbit1003.eqiad.wmnet' ([[phab:T381499|T381499]]) * 19:45 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudrabbit1003.eqiad.wmnet' ([[phab:T381499|T381499]]) * 19:44 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudrabbit1002.eqiad.wmnet' ([[phab:T381499|T381499]]) * 19:35 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudrabbit1002.eqiad.wmnet' ([[phab:T381499|T381499]]) * 19:34 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudrabbit1001.eqiad.wmnet' ([[phab:T381499|T381499]]) * 19:26 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudrabbit1001.eqiad.wmnet' ([[phab:T381499|T381499]]) * 13:29 taavi: run wmcs-wikireplica-dns to create transitional x3 CNAMEs [[phab:T390954|T390954]] === 2025-04-02 === * 20:56 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1067.eqiad.wmnet' ([[phab:T381499|T381499]]) * 20:49 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1067.eqiad.wmnet' ([[phab:T381499|T381499]]) * 20:49 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1066.eqiad.wmnet' ([[phab:T381499|T381499]]) * 20:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1066.eqiad.wmnet' ([[phab:T381499|T381499]]) * 20:41 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1065.eqiad.wmnet' ([[phab:T381499|T381499]]) * 20:33 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1065.eqiad.wmnet' ([[phab:T381499|T381499]]) * 20:33 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1064.eqiad.wmnet' ([[phab:T381499|T381499]]) * 20:26 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1064.eqiad.wmnet' ([[phab:T381499|T381499]]) * 20:26 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1063.eqiad.wmnet' ([[phab:T381499|T381499]]) * 20:20 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1063.eqiad.wmnet' ([[phab:T381499|T381499]]) * 20:20 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1062.eqiad.wmnet' ([[phab:T381499|T381499]]) * 20:12 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1062.eqiad.wmnet' ([[phab:T381499|T381499]]) * 20:12 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1061.eqiad.wmnet' ([[phab:T381499|T381499]]) * 20:06 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1061.eqiad.wmnet' ([[phab:T381499|T381499]]) * 20:06 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1060.eqiad.wmnet' ([[phab:T381499|T381499]]) * 19:59 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1060.eqiad.wmnet' ([[phab:T381499|T381499]]) * 19:59 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1059.eqiad.wmnet' ([[phab:T381499|T381499]]) * 19:57 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudnet1006.eqiad.wmnet' ([[phab:T381499|T381499]]) * 19:51 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1059.eqiad.wmnet' ([[phab:T381499|T381499]]) * 19:51 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1058.eqiad.wmnet' ([[phab:T381499|T381499]]) * 19:48 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet1006.eqiad.wmnet' ([[phab:T381499|T381499]]) * 19:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudnet1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 19:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1058.eqiad.wmnet' ([[phab:T381499|T381499]]) * 19:44 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1057.eqiad.wmnet' ([[phab:T381499|T381499]]) * 19:37 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1057.eqiad.wmnet' ([[phab:T381499|T381499]]) * 19:37 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1056.eqiad.wmnet' ([[phab:T381499|T381499]]) * 19:37 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 19:30 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1056.eqiad.wmnet' ([[phab:T381499|T381499]]) * 19:30 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1055.eqiad.wmnet' ([[phab:T381499|T381499]]) * 19:24 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1055.eqiad.wmnet' ([[phab:T381499|T381499]]) * 19:24 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1054.eqiad.wmnet' ([[phab:T381499|T381499]]) * 19:17 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1054.eqiad.wmnet' ([[phab:T381499|T381499]]) * 19:17 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1053.eqiad.wmnet' ([[phab:T381499|T381499]]) * 19:10 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1053.eqiad.wmnet' ([[phab:T381499|T381499]]) * 19:10 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1052.eqiad.wmnet' ([[phab:T381499|T381499]]) * 19:03 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1052.eqiad.wmnet' ([[phab:T381499|T381499]]) * 19:03 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1051.eqiad.wmnet' ([[phab:T381499|T381499]]) * 18:56 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1051.eqiad.wmnet' ([[phab:T381499|T381499]]) * 18:56 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1050.eqiad.wmnet' ([[phab:T381499|T381499]]) * 18:49 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1050.eqiad.wmnet' ([[phab:T381499|T381499]]) * 18:49 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1049.eqiad.wmnet' ([[phab:T381499|T381499]]) * 18:42 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1049.eqiad.wmnet' ([[phab:T381499|T381499]]) * 18:42 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1048.eqiad.wmnet' ([[phab:T381499|T381499]]) * 18:38 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudbackup2004.codfw.wmnet' ([[phab:T381499|T381499]]) * 18:35 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1048.eqiad.wmnet' ([[phab:T381499|T381499]]) * 18:35 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1047.eqiad.wmnet' ([[phab:T381499|T381499]]) * 18:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1043.eqiad.wmnet' ([[phab:T381499|T381499]]) * 18:02 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1043.eqiad.wmnet' ([[phab:T381499|T381499]]) * 18:02 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1042.eqiad.wmnet' ([[phab:T381499|T381499]]) * 18:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudbackup1004.eqiad.wmnet' ([[phab:T381499|T381499]]) * 18:00 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudbackup1003.eqiad.wmnet' ([[phab:T381499|T381499]]) * 17:55 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1042.eqiad.wmnet' ([[phab:T381499|T381499]]) * 17:55 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1041.eqiad.wmnet' ([[phab:T381499|T381499]]) * 17:50 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudbackup1003.eqiad.wmnet' ([[phab:T381499|T381499]]) * 17:48 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1041.eqiad.wmnet' ([[phab:T381499|T381499]]) * 17:48 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1040.eqiad.wmnet' ([[phab:T381499|T381499]]) * 17:42 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1040.eqiad.wmnet' ([[phab:T381499|T381499]]) * 17:42 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1039.eqiad.wmnet' ([[phab:T381499|T381499]]) * 17:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1039.eqiad.wmnet' ([[phab:T381499|T381499]]) * 17:36 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1038.eqiad.wmnet' ([[phab:T381499|T381499]]) * 17:30 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1038.eqiad.wmnet' ([[phab:T381499|T381499]]) * 17:30 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1037.eqiad.wmnet' ([[phab:T381499|T381499]]) * 17:23 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1037.eqiad.wmnet' ([[phab:T381499|T381499]]) * 17:23 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1036.eqiad.wmnet' ([[phab:T381499|T381499]]) * 17:17 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1036.eqiad.wmnet' ([[phab:T381499|T381499]]) * 17:17 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1035.eqiad.wmnet' ([[phab:T381499|T381499]]) * 17:11 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1035.eqiad.wmnet' ([[phab:T381499|T381499]]) * 17:11 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1034.eqiad.wmnet' ([[phab:T381499|T381499]]) * 17:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1034.eqiad.wmnet' ([[phab:T381499|T381499]]) * 17:05 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1033.eqiad.wmnet' ([[phab:T381499|T381499]]) * 16:59 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1033.eqiad.wmnet' ([[phab:T381499|T381499]]) * 16:59 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1032.eqiad.wmnet' ([[phab:T381499|T381499]]) * 16:55 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudweb.unset_maintenance (exit_code=0) ({{Gerrit|1133432}}) * 16:54 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1032.eqiad.wmnet' ([[phab:T381499|T381499]]) * 16:54 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1031.eqiad.wmnet' ([[phab:T381499|T381499]]) * 16:53 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudweb.unset_maintenance ({{Gerrit|1133432}}) * 16:52 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol1007.eqiad.wmnet' ([[phab:T381499|T381499]]) * 16:47 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1031.eqiad.wmnet' ([[phab:T381499|T381499]]) * 16:32 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1007.eqiad.wmnet' ([[phab:T381499|T381499]]) * 16:32 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 16:22 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 16:22 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.tofu (exit_code=97) running tofu plan+apply for main branch * 16:21 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 16:21 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol1006.eqiad.wmnet' ([[phab:T381499|T381499]]) * 16:13 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudservices1006.eqiad.wmnet' ([[phab:T381499|T381499]]) * 16:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1006.eqiad.wmnet' ([[phab:T381499|T381499]]) * 16:05 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=97) on host 'cloudservices1006.eqiad.wmnet' ([[phab:T381499|T381499]]) * 16:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1006.eqiad.wmnet' ([[phab:T381499|T381499]]) * 16:02 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1006.eqiad.wmnet' ([[phab:T381499|T381499]]) * 16:02 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 15:55 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 15:20 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudcontrol1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 15:11 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 15:11 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 15:04 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 15:04 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 15:04 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 15:03 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 15:03 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 15:02 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 15:02 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 15:02 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 15:02 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 15:02 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 15:02 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 15:02 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 15:01 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 15:01 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 15:01 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 15:01 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 15:01 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 15:01 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 15:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 15:00 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 15:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 15:00 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 15:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 15:00 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 15:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 15:00 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 15:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 15:00 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:59 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:59 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:59 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:59 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:59 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:58 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:58 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:58 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:57 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:57 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:57 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:57 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:57 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:56 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:56 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:56 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:56 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:55 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:55 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:55 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:54 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:54 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:54 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:50 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:50 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:37 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:37 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:37 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:36 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:36 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:36 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:36 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:36 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1007.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1007.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:36 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:35 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:35 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:35 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:35 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:35 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:35 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:35 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:34 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:34 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:32 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:31 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:31 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:30 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:30 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:29 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:29 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:28 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:28 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:28 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:27 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1006.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:27 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1006.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:27 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudcontrol1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:27 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:27 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudcontrol1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:27 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:26 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=97) on host 'cloudcontrol1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:20 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 14:18 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 14:18 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:15 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudservices1006.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:06 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1006.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:05 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 13:57 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 13:56 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudweb.set_maintenance (exit_code=0) ([[phab:T381499|T381499]]) * 13:54 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudweb.set_maintenance ([[phab:T381499|T381499]]) * 09:53 dhinus: systemctl restart maintain-dbusers.service (attempting to fix a weird issue) === 2025-04-01 === * 21:12 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for service: project,designate * 21:12 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for service: project,designate * 21:02 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 21:01 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 21:00 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.tofu (exit_code=97) running tofu plan+apply for main branch * 21:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 14:50 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 14:50 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2025-03-31 === * 15:54 dcaro@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.unset_cluster_maintenance (exit_code=0) * 15:54 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.unset_cluster_maintenance * 15:54 dcaro@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.unset_cluster_maintenance (exit_code=0) * 15:54 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.unset_cluster_maintenance * 15:53 dcaro@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.set_cluster_in_maintenance (exit_code=0) * 15:53 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.set_cluster_in_maintenance === 2025-03-27 === * 14:34 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services * 14:22 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 12:00 dcaro@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) ([[phab:T390134|T390134]]) * 08:41 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy ([[phab:T390134|T390134]]) * 08:41 dcaro@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) ([[phab:T390134|T390134]]) * 08:41 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy ([[phab:T390134|T390134]]) === 2025-03-26 === * 10:38 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 10:37 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2025-03-25 === * 14:16 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 14:12 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:22 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 12:21 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:14 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 12:14 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2025-03-24 === * 11:01 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 11:00 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 11:00 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 10:59 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2025-03-22 === * 03:48 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,designate * 03:48 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,designate === 2025-03-19 === * 14:26 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for service: project,designate * 14:25 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for service: project,designate * 12:02 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for service: project,task_id,no_dologmsg,cluster_name,all_services,nova,glance,keystone,cinder,neutron,trove,magnum,heat,swift,designate,filter_nodes * 12:01 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for service: project,task_id,no_dologmsg,cluster_name,all_services,nova,glance,keystone,cinder,neutron,trove,magnum,heat,swift,designate,filter_nodes * 06:24 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 06:21 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 06:14 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 06:12 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudbackup1002-dev.eqiad.wmnet' * 06:11 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 06:08 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt2006-dev.codfw.wmnet' * 06:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudbackup1002-dev.eqiad.wmnet' * 06:05 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudbackup1001-dev.eqiad.wmnet' * 06:01 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt2006-dev.codfw.wmnet' * 06:01 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt2005-dev.codfw.wmnet' * 06:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudbackup1001-dev.eqiad.wmnet' * 05:59 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudnet2006-dev.codfw.wmnet' * 05:54 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt2005-dev.codfw.wmnet' * 05:54 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt2004-dev.codfw.wmnet' * 05:50 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet2006-dev.codfw.wmnet' * 05:49 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudnet2006-dev.codfw.wmnet' * 05:49 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet2006-dev.codfw.wmnet' * 05:48 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudnet2006-dev.codfw.wmnet' * 05:48 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet2006-dev.codfw.wmnet' * 05:47 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudnet2006-dev.codfw.wmnet' * 05:47 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet2006-dev.codfw.wmnet' * 05:47 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt2004-dev.codfw.wmnet' * 05:46 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudnet2006-dev.codfw.wmnet' * 05:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudservices2005-dev.codfw.wmnet' * 05:46 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet2006-dev.codfw.wmnet' * 05:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudnet2005-dev.codfw.wmnet' * 05:37 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices2005-dev.codfw.wmnet' * 05:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet2005-dev.codfw.wmnet' * 05:34 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudnet2005-dev.codfw.wmnet' * 05:34 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet2005-dev.codfw.wmnet' * 05:34 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudnet2005-dev.codfw.wmnet' * 05:33 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet2005-dev.codfw.wmnet' * 05:32 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudnet2005-dev.codfw.wmnet' * 05:31 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 05:31 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet2005-dev.codfw.wmnet' * 05:30 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudnet2005-dev.codfw.wmnet' * 05:30 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet2005-dev.codfw.wmnet' * 05:28 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 05:27 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudnet2005-dev.codfw.wmnet' * 05:27 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet2005-dev.codfw.wmnet' * 05:26 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudnet2005-dev.codfw.wmnet' * 05:25 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet2005-dev.codfw.wmnet' * 05:25 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudnet2005-dev.codfw.wmnet' * 05:25 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet2005-dev.codfw.wmnet' * 05:23 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudnet2005-dev.codfw.wmnet' * 05:23 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet2005-dev.codfw.wmnet' * 05:22 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudnet2005-dev.codfw.wmnet' * 05:21 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet2005-dev.codfw.wmnet' * 05:20 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudnet2005-dev.codfw.wmnet' * 05:20 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet2005-dev.codfw.wmnet' * 05:20 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices2005-dev.codfw.wmnet' * 05:20 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices2005-dev.codfw.wmnet' * 05:18 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices2005-dev.codfw.wmnet' * 05:18 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices2005-dev.codfw.wmnet' * 05:17 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices2005-dev.codfw.wmnet' * 05:17 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices2005-dev.codfw.wmnet' * 05:17 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices2005-dev.codfw.wmnet' * 05:17 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices2005-dev.codfw.wmnet' * 05:16 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices2005-dev.codfw.wmnet' * 05:15 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices2005-dev.codfw.wmnet' * 05:15 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices2005-dev.codfw.wmnet' * 05:14 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices2005-dev.codfw.wmnet' * 05:14 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudnet2005-dev.codfw.wmnet' * 05:14 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet2005-dev.codfw.wmnet' * 05:13 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudnet2005-dev.codfw.wmnet' * 05:13 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet2005-dev.codfw.wmnet' * 05:12 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=97) on host 'cloudnet2005-dev.codfw.wmnet' * 05:12 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet2005-dev.codfw.wmnet' * 05:09 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices2005-dev.codfw.wmnet' * 05:09 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices2005-dev.codfw.wmnet' * 05:07 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudbackup1001-dev.eqiad.wmnet' * 05:06 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices2005-dev.codfw.wmnet' * 05:06 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices2005-dev.codfw.wmnet' * 05:04 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices2005-dev.codfw.wmnet' * 05:04 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices2005-dev.codfw.wmnet' * 05:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudbackup1001-dev.eqiad.wmnet' * 04:58 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices2005-dev.codfw.wmnet' * 04:58 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices2005-dev.codfw.wmnet' * 04:56 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol2006-dev.codfw.wmnet' * 04:54 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices2005-dev.codfw.wmnet' * 04:53 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices2005-dev.codfw.wmnet' * 04:49 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudservices2004-dev.codfw.wmnet' * 04:40 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices2004-dev.codfw.wmnet' * 04:39 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2006-dev.codfw.wmnet' * 04:38 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol2004-dev.codfw.wmnet' * 04:29 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2004-dev.codfw.wmnet' * 04:29 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol2005-dev.codfw.wmnet' * 04:11 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2005-dev.codfw.wmnet' * 04:11 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudcontrol2004-dev.codfw.wmnet' * 04:10 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2004-dev.codfw.wmnet' * 04:09 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudcontrol2004-dev.codfw.wmnet' * 04:09 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2004-dev.codfw.wmnet' * 04:09 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudcontrol2004-dev.codfw.wmnet' * 04:08 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2004-dev.codfw.wmnet' * 03:06 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudcontrol2004-dev.codfw.wmnet' * 02:54 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2004-dev.codfw.wmnet' * 02:18 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudcontrol2004-dev.codfw.wmnet' * 02:15 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2004-dev.codfw.wmnet' * 02:12 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudweb.set_maintenance (exit_code=0) ([[phab:T381499|T381499]]) * 02:12 andrewbogott: upgrading codfw1dev to openstack 'dalmation' [[phab:T381499|T381499]] * 02:12 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudweb.set_maintenance ([[phab:T381499|T381499]]) === 2025-03-14 === * 13:36 volans: installed cumin v5.1.1 on cloudcumin* hosts === 2025-03-05 === * 17:52 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services * 17:45 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 00:54 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services * 00:43 andrewbogott: restarting all openstack services in hopes of gettting dns unstuck * 00:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services === 2025-03-04 === * 08:52 arturo: [[phab:T387828|T387828]] depooled galera on cloudcontrol1005 === 2025-03-01 === * 19:35 andrewbogott: installed new bookworm base images in eqiad1 * 19:35 andrewbogott: installed new bookworm and bullseye base images in codfw1dev === 2025-02-28 === * 15:41 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 15:39 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2025-02-27 === * 15:38 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 15:35 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2025-02-24 === * 00:24 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 00:22 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 00:08 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 00:07 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 00:06 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 00:06 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 00:06 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.tofu (exit_code=97) running tofu plan for main branch * 00:06 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch === 2025-02-23 === * 21:47 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 21:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2025-02-21 === * 12:32 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 12:29 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2025-02-20 === * 17:25 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) ([[phab:T386083|T386083]]) * 17:25 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance ([[phab:T386083|T386083]]) === 2025-02-19 === * 13:26 arturo: manual failover of cloudgw1004 to cloudgw1003 [[phab:T382356|T382356]] * 01:32 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.reboot_node (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudcontrol2005-dev.codfw.wmnet<nowiki>}</nowiki>' * 01:18 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.reboot_node on hosts matched by 'D<nowiki>{</nowiki>cloudcontrol2005-dev.codfw.wmnet<nowiki>}</nowiki>' * 01:02 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 00:59 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2025-02-18 === * 12:13 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 12:13 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:13 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 12:10 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2025-02-13 === * 13:15 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 13:14 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2025-02-11 === * 11:47 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1047.eqiad.wmnet' ([[phab:T386083|T386083]]) * 11:32 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1047.eqiad.wmnet' ([[phab:T386083|T386083]]) * 11:24 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1047' ([[phab:T386083|T386083]]) * 11:24 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1047' ([[phab:T386083|T386083]]) === 2025-02-07 === * 13:52 root@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudnet.reboot_node (exit_code=0) for host cloudnet1005.eqiad.wmnet ([[phab:T384946|T384946]]) * 13:48 root@cloudcumin1001: START - Cookbook wmcs.openstack.cloudnet.reboot_node for host cloudnet1005.eqiad.wmnet ([[phab:T384946|T384946]]) * 02:00 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 02:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 01:48 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=97) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1041.eqiad.wmnet<nowiki>}</nowiki>' * 01:47 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1041.eqiad.wmnet<nowiki>}</nowiki>' * 01:40 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1041.eqiad.wmnet<nowiki>}</nowiki>' * 01:39 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1041.eqiad.wmnet<nowiki>}</nowiki>' * 01:39 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1036.eqiad.wmnet<nowiki>}</nowiki>' * 01:31 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1036.eqiad.wmnet<nowiki>}</nowiki>' * 01:31 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1037.eqiad.wmnet<nowiki>}</nowiki>' * 01:12 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1037.eqiad.wmnet<nowiki>}</nowiki>' * 01:12 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1038.eqiad.wmnet<nowiki>}</nowiki>' * 00:53 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1038.eqiad.wmnet<nowiki>}</nowiki>' * 00:53 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1039.eqiad.wmnet<nowiki>}</nowiki>' * 00:33 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1039.eqiad.wmnet<nowiki>}</nowiki>' === 2025-02-06 === * 21:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1040.eqiad.wmnet<nowiki>}</nowiki>' * 21:25 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1040.eqiad.wmnet<nowiki>}</nowiki>' * 21:25 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1041.eqiad.wmnet<nowiki>}</nowiki>' * 21:10 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1050.eqiad.wmnet<nowiki>}</nowiki>' * 21:09 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1041.eqiad.wmnet<nowiki>}</nowiki>' * 21:09 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1042.eqiad.wmnet<nowiki>}</nowiki>' * 20:48 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1050.eqiad.wmnet<nowiki>}</nowiki>' * 20:48 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1051.eqiad.wmnet<nowiki>}</nowiki>' * 20:46 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1042.eqiad.wmnet<nowiki>}</nowiki>' * 20:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1043.eqiad.wmnet<nowiki>}</nowiki>' * 20:26 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1043.eqiad.wmnet<nowiki>}</nowiki>' * 20:26 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1044.eqiad.wmnet<nowiki>}</nowiki>' * 20:25 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1051.eqiad.wmnet<nowiki>}</nowiki>' * 20:25 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1052.eqiad.wmnet<nowiki>}</nowiki>' * 20:04 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1052.eqiad.wmnet<nowiki>}</nowiki>' * 20:04 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1053.eqiad.wmnet<nowiki>}</nowiki>' * 20:04 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1044.eqiad.wmnet<nowiki>}</nowiki>' * 20:04 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1045.eqiad.wmnet<nowiki>}</nowiki>' * 19:48 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1053.eqiad.wmnet<nowiki>}</nowiki>' * 19:48 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1054.eqiad.wmnet<nowiki>}</nowiki>' * 19:42 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1045.eqiad.wmnet<nowiki>}</nowiki>' * 19:42 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1046.eqiad.wmnet<nowiki>}</nowiki>' * 19:24 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1054.eqiad.wmnet<nowiki>}</nowiki>' * 19:24 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1055.eqiad.wmnet<nowiki>}</nowiki>' * 19:19 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1046.eqiad.wmnet<nowiki>}</nowiki>' * 19:19 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1047.eqiad.wmnet<nowiki>}</nowiki>' * 19:19 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1055.eqiad.wmnet<nowiki>}</nowiki>' * 19:19 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1056.eqiad.wmnet<nowiki>}</nowiki>' * 19:03 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1047.eqiad.wmnet<nowiki>}</nowiki>' * 19:02 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1048.eqiad.wmnet<nowiki>}</nowiki>' * 19:02 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1056.eqiad.wmnet<nowiki>}</nowiki>' * 19:02 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1057.eqiad.wmnet<nowiki>}</nowiki>' * 18:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1057.eqiad.wmnet<nowiki>}</nowiki>' * 18:41 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1058.eqiad.wmnet<nowiki>}</nowiki>' * 18:40 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1048.eqiad.wmnet<nowiki>}</nowiki>' * 18:40 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1049.eqiad.wmnet<nowiki>}</nowiki>' * 18:21 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1049.eqiad.wmnet<nowiki>}</nowiki>' * 18:21 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1060.eqiad.wmnet<nowiki>}</nowiki>' * 18:13 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1058.eqiad.wmnet<nowiki>}</nowiki>' * 18:13 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1059.eqiad.wmnet<nowiki>}</nowiki>' * 17:56 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1060.eqiad.wmnet<nowiki>}</nowiki>' * 17:56 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1061.eqiad.wmnet<nowiki>}</nowiki>' * 17:52 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1059.eqiad.wmnet<nowiki>}</nowiki>' * 17:52 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1061.eqiad.wmnet<nowiki>}</nowiki>' * 17:52 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1062.eqiad.wmnet<nowiki>}</nowiki>' * 17:35 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1062.eqiad.wmnet<nowiki>}</nowiki>' * 17:35 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1063.eqiad.wmnet<nowiki>}</nowiki>' * 17:13 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1063.eqiad.wmnet<nowiki>}</nowiki>' * 17:13 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1064.eqiad.wmnet<nowiki>}</nowiki>' * 16:55 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1064.eqiad.wmnet<nowiki>}</nowiki>' * 16:55 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1065.eqiad.wmnet<nowiki>}</nowiki>' * 16:38 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1065.eqiad.wmnet<nowiki>}</nowiki>' * 16:38 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1066.eqiad.wmnet<nowiki>}</nowiki>' * 16:29 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1066.eqiad.wmnet<nowiki>}</nowiki>' * 16:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1067.eqiad.wmnet<nowiki>}</nowiki>' * 16:22 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1067.eqiad.wmnet<nowiki>}</nowiki>' * 13:45 andrewbogott: cold-migrating all remaining VMs in [[phab:T385264|T385264]] except for 'integration' and 'tools' VMs === 2025-02-05 === * 19:12 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 19:12 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance === 2025-02-04 === * 12:48 arturo: replacing cloudgw1002 with cloudgw1004 - [[phab:T382356|T382356]] * 09:22 arturo: fleet-wide restart of puppetservers [[phab:T385553|T385553]] === 2025-02-03 === * 13:42 andrewbogott: rebooting proxy-04.project-proxy for the ceph OSD mishap === 2025-02-02 === * 17:49 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2004-dev.codfw.wmnet<nowiki>}</nowiki>' * 17:48 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2004-dev.codfw.wmnet<nowiki>}</nowiki>' * 17:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 17:46 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 17:46 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2004-dev.codfw.wmnet<nowiki>}</nowiki>' * 17:45 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2004-dev.codfw.wmnet<nowiki>}</nowiki>' * 17:10 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2004-dev.codfw.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 17:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2004-dev.codfw.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 16:27 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2005-dev.codfw.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 16:15 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2005-dev.codfw.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 16:15 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2005-dev<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 16:15 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2005-dev<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 16:05 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 16:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 16:05 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 16:04 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 16:04 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 16:04 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance === 2025-01-31 === * 01:41 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 01:37 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2025-01-30 === * 21:39 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=97) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1036.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 21:21 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1036.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 20:54 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 20:54 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 20:44 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 20:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 19:47 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 19:40 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 14:46 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=97) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1036.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 14:33 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1036.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 14:33 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=97) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1036.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 14:30 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 14:25 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 13:38 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) * 13:38 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 13:37 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) * 13:37 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 13:37 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) * 13:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 13:36 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) * 13:35 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 13:35 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) * 13:34 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 12:51 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1036.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 12:51 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1035.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 12:46 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1035.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 12:45 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1034.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 12:39 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1034.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 12:38 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1033.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 12:33 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1033.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 12:30 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1032.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 12:25 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1032.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 12:23 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1031.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 12:17 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1031.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 04:48 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1067.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 03:32 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 02:47 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1067.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 02:17 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1067.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 00:17 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1067.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 00:05 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) === 2025-01-29 === * 23:53 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 23:43 andrewbogott: resetting rabbitmq in eqiad1 in hopes that it will resolve mysterious openstack misbehaviors * 20:54 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2005-dev.codfw.wmnet<nowiki>}</nowiki>' * 20:51 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2005-dev.codfw.wmnet<nowiki>}</nowiki>' * 20:43 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2006-dev.codfw.wmnet<nowiki>}</nowiki>' * 20:42 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2006-dev.codfw.wmnet<nowiki>}</nowiki>' * 20:39 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2006-dev.codfw.wmnet<nowiki>}</nowiki>' * 20:38 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2006-dev.codfw.wmnet<nowiki>}</nowiki>' * 20:37 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2006-dev.codfw.wmnet<nowiki>}</nowiki>' * 20:32 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2006-dev.codfw.wmnet<nowiki>}</nowiki>' * 20:32 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 20:29 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 20:07 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 20:05 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 20:04 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2006-dev.codfw.wmnet<nowiki>}</nowiki>' * 20:03 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2006-dev.codfw.wmnet<nowiki>}</nowiki>' * 19:57 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 19:56 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 19:54 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1067.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 19:54 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 19:49 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1067.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 19:49 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1066.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 19:44 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2006-dev.codfw.wmnet<nowiki>}</nowiki>' * 19:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2006-dev.codfw.wmnet<nowiki>}</nowiki>' * 19:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 19:42 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 19:42 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2006-dev.codfw.wmnet<nowiki>}</nowiki>' * 19:38 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 18:56 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2006-dev.codfw.wmnet<nowiki>}</nowiki>' * 18:24 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1066.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 18:24 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1067.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 18:07 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.reboot_node (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudcontrol2009-dev.codfw.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 18:04 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.reboot_node on hosts matched by 'D<nowiki>{</nowiki>cloudcontrol2009-dev.codfw.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 17:55 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.reboot_node (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudcontrol2006-dev.codfw.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 17:51 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.reboot_node on hosts matched by 'D<nowiki>{</nowiki>cloudcontrol2006-dev.codfw.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 17:51 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.reboot_node (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudcontrol2004-dev.codfw.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 17:51 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.reboot_node (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudcontrol1007.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 17:45 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.reboot_node on hosts matched by 'D<nowiki>{</nowiki>cloudcontrol2004-dev.codfw.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 17:45 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.reboot_node (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudcontrol2005-dev.codfw.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 17:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.reboot_node on hosts matched by 'D<nowiki>{</nowiki>cloudcontrol1007.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 17:40 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.reboot_node (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudcontrol1006.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 17:39 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.reboot_node on hosts matched by 'D<nowiki>{</nowiki>cloudcontrol2005-dev.codfw.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 17:35 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.reboot_node on hosts matched by 'D<nowiki>{</nowiki>cloudcontrol1006.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 17:35 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.reboot_node (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudcontrol1005.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 17:30 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.reboot_node on hosts matched by 'D<nowiki>{</nowiki>cloudcontrol1005.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 12:13 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add === 2025-01-28 === * 21:04 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 21:04 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 21:04 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 21:03 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 20:51 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 20:50 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 20:49 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 16:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T348643|T348643]]) * 16:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T348643|T348643]]) * 13:04 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T348643|T348643]]) * 13:03 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T348643|T348643]]) * 13:02 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.drain_node (exit_code=97) ([[phab:T348643|T348643]]) * 13:02 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T348643|T348643]]) * 02:59 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T348643|T348643]]) * 00:49 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T348643|T348643]]) * 00:49 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.drain_node (exit_code=97) ([[phab:T348643|T348643]]) * 00:04 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T348643|T348643]]) === 2025-01-27 === * 21:56 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.unset_cluster_maintenance (exit_code=0) * 21:56 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.unset_cluster_maintenance * 21:52 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 21:52 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 21:52 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=97) * 21:49 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 15:12 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 15:11 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 15:11 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 15:11 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 15:05 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 15:04 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 15:04 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 15:04 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 15:03 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 15:03 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 15:00 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 15:00 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 14:59 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 14:59 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 14:56 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 14:56 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 14:55 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 14:55 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 14:46 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 14:46 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 14:46 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 14:46 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 14:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) * 14:42 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 14:39 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 14:31 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add === 2025-01-24 === * 13:47 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) * 13:36 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node === 2025-01-23 === * 20:44 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 20:34 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) * 19:37 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 18:53 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) * 18:05 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 17:25 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) * 15:49 dhinus: cumin 'P:base::cloud_production' 'rm /var/lib/prometheus/node.d/kernel-panic.prom' [[phab:T382961|T382961]] * 15:32 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 15:11 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 15:04 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.undrain_node * 14:57 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 14:57 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 14:30 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 13:02 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add === 2025-01-22 === * 14:52 dhinus: cloudcumin[12]001 upgrade spicerack from 8.15.2 to 9.1.0 === 2025-01-21 === * 13:42 andrewbogott: migrating/rebooting VMs as per earlier email, [[phab:T383583|T383583]] === 2025-01-20 === * 14:23 dhinus: cumin upgraded form 4.2.0 to 5.0.0 on cloudcumin[12]001. patch [[phab:T346453|T346453]] reapplied. === 2025-01-14 === * 21:26 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) ([[phab:T383583|T383583]]) * 21:26 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance ([[phab:T383583|T383583]]) * 20:55 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1055.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T383583|T383583]]) * 20:54 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1055.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T383583|T383583]]) * 20:54 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1055.eqiad.wmnet' ([[phab:T383583|T383583]]) * 20:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1055.eqiad.wmnet' ([[phab:T383583|T383583]]) * 18:08 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.roll_restart_osd_daemons (exit_code=0) === 2025-01-13 === * 15:59 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.roll_restart_osd_daemons === 2025-01-10 === * 16:31 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 16:30 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2025-01-09 === * 20:55 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 20:54 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 17:48 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 17:47 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 17:41 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 17:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 17:41 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 17:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 16:56 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 16:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 11:46 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) ([[phab:T309789|T309789]]) * 10:19 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy ([[phab:T309789|T309789]]) === 2025-01-08 === * 14:18 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) ([[phab:T309789|T309789]]) * 13:04 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy ([[phab:T309789|T309789]]) * 12:59 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) ([[phab:T309789|T309789]]) * 12:59 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy ([[phab:T309789|T309789]]) * 12:59 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) ([[phab:T309789|T309789]]) * 12:59 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy ([[phab:T309789|T309789]]) === 2025-01-03 === * 21:01 bd808: `sudo service maintain-dbusers restart` on cloudcontrol1005. Report of missing replica.my.cnf and journalctl output empty due to log rotation. ([[phab:T382962|T382962]]) === 2024-12-18 === * 15:33 arturo: cloudgw failover for [[phab:T382220|T382220]] === 2024-12-04 === * 17:18 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1035.eqiad.wmnet' * 17:16 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1035.eqiad.wmnet' * 17:16 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=97) on host 'cloudvirt1035.eqiad.wmnet' ([[phab:T380893|T380893]]) * 17:16 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1035.eqiad.wmnet' ([[phab:T380893|T380893]]) * 15:09 andrewbogott: rebooted cloudinfra-cloudvps-puppetserver-1, it's so busy that it's unresponsive * 14:55 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1035.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T380731|T380731]]) * 14:49 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1035.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T380731|T380731]]) * 14:48 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1035.eqiad.wmnet' ([[phab:T380893|T380893]]) * 14:45 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1035.eqiad.wmnet' ([[phab:T380893|T380893]]) === 2024-11-27 === * 13:40 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1061.eqiad.wmnet' * 13:26 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1061.eqiad.wmnet' === 2024-11-26 === * 16:51 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1062.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T380731|T380731]]) * 16:46 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1062.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T380731|T380731]]) * 15:16 rook@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 15:15 rook@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 15:15 rook@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 15:15 rook@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 15:14 rook@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/145 * 15:13 rook@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/145 * 15:12 rook@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 15:12 rook@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 13:31 dcaro: added cloudcephmon1004 to the ceph mon pool * 12:41 aborrero@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.cloudvirt.vm_console (exit_code=255) * 12:40 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.vm_console * 12:35 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.vm_console (exit_code=99) * 12:35 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.vm_console * 12:34 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.vm_console (exit_code=99) * 12:34 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.vm_console * 12:34 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.vm_console (exit_code=99) * 12:34 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.vm_console * 05:40 andrewbogott: rebooting tools-sgebastion-10.tools.eqiad1.wikimedia.cloud to get NFS things remounted * 04:56 andrewbogott: 'systemctl restart nfs-server' on tools-nfs-2.tools.eqiad1.wikimedia.cloud === 2024-11-25 === * 14:47 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 14:46 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 14:41 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 14:40 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 13:55 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 13:55 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 13:48 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 13:48 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 13:30 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 13:29 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:57 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 12:56 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:11 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 12:11 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 11:28 arturo: create IPv6 networks in eqiad1 ([[phab:T380174|T380174]]), then reverted because network outage * 10:40 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 10:39 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 10:33 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 10:32 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 10:29 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 10:29 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2024-11-23 === * 14:58 taavi: removed broken records for re-created VMs in .eqiad.wmflabs zone === 2024-11-22 === * 14:36 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 14:35 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 13:48 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 13:48 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2024-11-20 === * 12:56 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 12:56 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 10:27 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 10:26 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 10:13 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 10:13 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 10:05 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 10:04 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 10:01 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 10:01 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch === 2024-11-19 === * 16:31 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 16:28 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 13:43 aborrero@cloudcumin2001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 13:43 aborrero@cloudcumin2001: START - Cookbook wmcs.openstack.restart_openstack * 13:37 aborrero@cloudcumin2001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 13:35 aborrero@cloudcumin2001: START - Cookbook wmcs.openstack.restart_openstack * 11:39 arturo: [codfw1dev] performing rabbit full reset [[phab:T380208|T380208]] * 10:26 aborrero@cloudcumin2001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 10:24 aborrero@cloudcumin2001: START - Cookbook wmcs.openstack.restart_openstack * 10:24 arturo: [codfw1dev] restart rabbitmq and nova/neutron services for [[phab:T380208|T380208]] === 2024-11-16 === * 08:37 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 08:35 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 07:52 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 07:52 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2024-11-15 === * 19:09 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 19:09 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 17:18 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 17:18 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 17:18 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 17:17 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 17:14 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 17:13 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 17:10 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 17:09 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 17:06 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 17:06 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 17:06 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 17:05 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 16:30 aborrero@cloudcumin2001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 16:29 arturo: [codfw1dev] restart rabbitmq and designate * 16:29 aborrero@cloudcumin2001: START - Cookbook wmcs.openstack.restart_openstack === 2024-11-14 === * 09:52 aborrero@cloudcumin2001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 09:48 aborrero@cloudcumin2001: START - Cookbook wmcs.openstack.restart_openstack === 2024-11-13 === * 13:05 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/120 * 13:04 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/120 === 2024-11-11 === * 14:49 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.vm_console (exit_code=99) * 14:49 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.vm_console === 2024-11-08 === * 16:47 aborrero@cloudcumin2001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 16:45 aborrero@cloudcumin2001: START - Cookbook wmcs.openstack.restart_openstack * 16:42 aborrero@cloudcumin2001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 16:42 arturo: [codfw1dev] restart all nova services and rabbitmq out of despair * 16:42 aborrero@cloudcumin2001: START - Cookbook wmcs.openstack.restart_openstack * 13:45 aborrero@cloudcumin2001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 13:44 aborrero@cloudcumin2001: START - Cookbook wmcs.openstack.restart_openstack * 13:33 aborrero@cloudcumin2001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 13:32 aborrero@cloudcumin2001: START - Cookbook wmcs.openstack.restart_openstack * 13:28 arturo: [codfw1dev] restart rabbitmq, openstack services logs show connection errors === 2024-11-07 === * 17:42 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 17:42 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 13:24 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 13:23 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2024-11-06 === * 13:09 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 13:08 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 13:05 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/117 * 13:04 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/117 * 13:04 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/117 * 13:03 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/117 * 11:15 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 11:14 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 11:10 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/116 * 11:09 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/116 * 09:57 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 09:55 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2024-11-05 === * 10:24 arturo: [codfw1dev] disable puppet and make changes for testing [[phab:T378192|T378192]] === 2024-11-04 === * 16:09 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 16:09 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 16:08 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 16:05 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 16:04 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 16:04 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 16:02 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 16:02 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 15:57 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 15:56 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 14:11 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 14:10 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 14:09 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 14:09 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 14:08 aborrero@cloudcumin2001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 14:08 aborrero@cloudcumin2001: START - Cookbook wmcs.openstack.restart_openstack * 14:08 aborrero@cloudcumin2001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 14:08 aborrero@cloudcumin2001: START - Cookbook wmcs.openstack.restart_openstack * 13:52 arturo: [codfw1dev] live-hack cloudlb2001-dev and cloudcontrol2004-dev for [[phab:T378192|T378192]] * 13:52 arturo: [codfw1dev] restart rabbitmq * 10:05 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 10:03 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 09:59 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 09:59 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2024-10-29 === * 15:38 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 15:38 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 13:40 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 13:39 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2024-10-28 === * 17:01 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 17:00 arturo: [codfw1dev] restarting rabbitmq, misbehaving * 17:00 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 16:54 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 16:54 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 16:53 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 16:53 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 16:22 dhinus: apt full-upgrade and reboot for cloudcumin* * 16:21 dhinus: upgrade spicerack from 8.8.0 to 8.15.1 on cloudcumin* === 2024-10-24 === * 15:25 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 15:24 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:51 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 12:51 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2024-10-22 === * 15:02 taavi: recover access to User:Labslogbot [[phab:T376220|T376220]] === 2024-10-21 === * 12:09 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 12:07 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 12:06 arturo: [codfw1dev] restart rabbitmq, tofu shows error talking to the designate API === 2024-10-16 === * 14:40 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 14:39 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2024-10-15 === * 15:50 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 15:50 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 15:44 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 15:43 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 15:38 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for main branch * 15:38 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 15:35 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for main branch * 15:35 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 15:35 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for main branch * 15:35 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 15:32 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for main branch * 15:32 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 15:32 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for main branch * 15:32 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 15:31 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for main branch * 15:31 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 15:31 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for main branch * 15:31 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 11:41 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 11:41 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 11:33 arturo: cloudgw maintenance, firewall change for [[phab:T374714|T374714]] * 10:36 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 10:36 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 10:33 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 10:32 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 10:29 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 10:28 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 10:22 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 10:18 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2024-10-11 === * 15:31 arturo: cloudgw maintenance firewall change [[phab:T374716|T374716]] * 09:51 arturo: cloudgw network maintenance related to [[phab:T376879|T376879]] * 08:33 arturo: [codfw1dev] reboot cloudgw2002-dev/2003-dev because network connectivity issues === 2024-10-10 === * 12:04 arturo: manual network failover in cloudgw because maintenance related to [[phab:T376879|T376879]] * 10:40 dhinus: cumin 'cloudrabbit*' 'systemctl restart rabbitmq-server' [[phab:T376802|T376802]] * 09:19 arturo: [codfw1dev] enable IPv6 on cloudgw ([[phab:T374716|T374716]]) === 2024-10-09 === * 14:20 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/93 * 14:20 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/93 * 14:19 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/93 * 14:19 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/93 * 14:16 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/93 * 14:16 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/93 * 10:58 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/93 * 10:58 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/93 * 10:55 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/93 * 10:55 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/93 * 10:54 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/93 * 10:53 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/93 * 10:53 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/93 * 10:52 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/93 * 10:34 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/95 * 10:34 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/95 * 10:34 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 10:33 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch === 2024-10-08 === * 11:16 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 11:16 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2024-10-07 === * 12:48 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 12:48 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:17 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 12:06 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:06 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 12:05 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 11:28 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 11:27 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 10:11 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 10:11 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2024-10-04 === * 15:12 arturo: cloudservice1005/1006: enable puppet and restore /etc/powerdns/recursor.conf to non-debug mode ([[phab:T374830|T374830]]) * 11:52 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 11:52 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 11:39 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 11:39 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 11:35 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 11:35 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 08:32 arturo: cloudservice1005/1006: disable puppet and set `quiet=no` in /etc/powerdns/recursor.conf ([[phab:T374830|T374830]]) === 2024-10-03 === * 14:55 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 14:54 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 14:54 arturo: [codfw1dev] delete default security group rule list, now tracking them via tofu-infra * 14:53 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 14:52 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 14:51 arturo: delete default security group rule list, now tracking them via tofu-infra * 14:48 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 14:47 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2024-10-02 === * 15:24 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 15:23 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 15:23 aborrero@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.tofu (exit_code=97) running tofu plan+apply for main branch * 15:23 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:22 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 12:21 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:21 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 12:21 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 12:00 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 11:59 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 11:56 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 11:56 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 11:55 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 11:55 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 11:43 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 10:38 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 10:38 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 10:37 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 10:26 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch ([[phab:T376211|T376211]]) * 10:21 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch ([[phab:T376211|T376211]]) * 10:16 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch ([[phab:T376211|T376211]]) * 10:15 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch ([[phab:T376211|T376211]]) * 10:05 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch ([[phab:T376211|T376211]]) * 10:04 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch ([[phab:T376211|T376211]]) * 09:14 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/77 ([[phab:T376211|T376211]]) * 09:14 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/77 ([[phab:T376211|T376211]]) * 09:11 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/77 ([[phab:T376211|T376211]]) * 09:10 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/77 ([[phab:T376211|T376211]]) === 2024-10-01 === * 15:41 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T372814|T372814]]) * 12:36 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 12:36 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:02 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 12:00 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 09:57 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 09:53 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 08:23 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T372814|T372814]]) * 08:18 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T372814|T372814]]) * 08:10 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T372814|T372814]]) === 2024-09-30 === * 20:12 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) ([[phab:T372814|T372814]]) * 16:59 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.reset_weights (exit_code=0) * 16:41 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.reset_weights * 16:35 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.undrain_node ([[phab:T372814|T372814]]) * 13:58 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.reset_weights (exit_code=99) * 13:47 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) ([[phab:T372814|T372814]]) * 13:38 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.reset_weights * 13:22 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.reset_weights (exit_code=0) * 13:01 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.reset_weights * 11:50 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.reset_weights (exit_code=99) * 11:19 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.reset_weights * 11:18 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.reset_weights (exit_code=99) * 11:18 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.reset_weights * 10:36 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.reset_weights (exit_code=0) * 10:27 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.reset_weights * 10:08 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.reset_weights (exit_code=0) * 10:06 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.reset_weights * 10:06 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.reset_weights (exit_code=99) * 10:04 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.reset_weights * 10:02 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.reset_weights (exit_code=0) * 10:00 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.reset_weights * 10:00 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.reset_weights (exit_code=99) * 09:58 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.reset_weights * 09:55 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.reset_weights (exit_code=0) * 09:55 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.reset_weights * 09:49 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.reset_weights (exit_code=0) * 09:48 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.reset_weights (exit_code=0) * 09:48 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.reset_weights * 09:46 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.reset_weights (exit_code=0) * 09:46 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.reset_weights * 09:44 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.reset_weights (exit_code=0) * 09:44 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.reset_weights * 09:41 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.reset_weights (exit_code=0) * 09:40 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.reset_weights * 09:39 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.reset_weights (exit_code=0) * 09:39 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.reset_weights * 09:32 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.reset_weights (exit_code=99) * 09:32 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.reset_weights * 09:01 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T372814|T372814]]) * 09:00 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T372814|T372814]]) * 09:00 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T372814|T372814]]) === 2024-09-27 === * 15:27 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 15:16 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 13:20 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 13:20 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 13:08 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 13:07 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 13:07 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 13:07 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 13:05 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 13:04 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 13:02 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 13:02 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 13:00 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 12:59 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:59 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 12:59 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:52 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 12:51 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:49 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 12:48 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:42 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 12:41 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 11:00 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 10:56 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 10:56 arturo: [codfw1dev] restart rabbitmq again * 10:16 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 10:15 arturo: [codfw1dev] restart rabbitmq * 10:14 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 10:04 arturo: [codfw1dev] enable IPv6 on the neutron virtual router [[phab:T375847|T375847]] === 2024-09-26 === * 14:52 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 14:52 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 14:28 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.undrain_node ([[phab:T372814|T372814]]) * 14:28 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) ([[phab:T372814|T372814]]) * 10:26 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.undrain_node ([[phab:T372814|T372814]]) * 10:25 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) ([[phab:T372814|T372814]]) * 10:25 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T372814|T372814]]) * 10:11 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T372814|T372814]]) * 08:57 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T372814|T372814]]) * 08:56 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T372814|T372814]]) * 08:56 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T372814|T372814]]) * 08:55 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T372814|T372814]]) * 08:55 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T372814|T372814]]) * 08:55 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T372814|T372814]]) * 08:54 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T372814|T372814]]) * 08:54 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T372814|T372814]]) * 08:54 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T372814|T372814]]) * 08:53 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T372814|T372814]]) * 08:52 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T372814|T372814]]) * 08:04 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T372814|T372814]]) * 08:04 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T372814|T372814]]) * 07:59 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T372814|T372814]]) * 07:48 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T372814|T372814]]) * 07:47 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T372814|T372814]]) * 07:47 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T372814|T372814]]) * 07:46 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T372814|T372814]]) * 07:46 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T372814|T372814]]) * 07:45 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T372814|T372814]]) * 07:45 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T372814|T372814]]) * 07:42 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T372814|T372814]]) * 07:42 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T372814|T372814]]) === 2024-09-25 === * 14:11 arturo: [codfw1dev] start proxy-02 vm on proxy-codfw1dev project, it was in shutoff mode for unknown reasons * 10:33 arturo: [codfw1dev] cleanup unused security groups ([[phab:T375604|T375604]]) * 09:54 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 09:54 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 09:53 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 09:52 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 09:52 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 09:48 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 09:48 arturo: [codfw1dev] restart rabbitmq on all cloudcontrol servers * 09:48 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 09:46 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 09:46 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 09:44 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 09:35 arturo: [codfw1dev] deletre a bunch of tests and seemingly unused projects === 2024-09-24 === * 19:33 dcaro@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T348643|T348643]]) * 16:00 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T348643|T348643]]) * 14:36 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) ([[phab:T348643|T348643]]) * 14:36 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T348643|T348643]]) * 14:35 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 13:14 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 13:10 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 13:00 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 12:57 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 12:56 arturo: [codfw1dev] restart rabbitmq-server on all 3 nodes * 12:53 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 12:46 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 09:11 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.undrain_rack * 09:10 wmbot~dcaro@urcuchillay: END (ERROR) - Cookbook wmcs.ceph.osd.undrain_rack (exit_code=97) * 09:10 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.undrain_rack === 2024-09-23 === * 19:38 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_rack (exit_code=99) * 15:55 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 15:54 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 15:52 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 15:52 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 15:43 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 15:43 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 15:43 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 15:42 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 15:40 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 15:27 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 14:42 arturo: put cloudvirt1048 in the network-ovs aggregate [[phab:T364457|T364457]] * 14:37 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.undrain_rack * 12:39 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 12:38 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:37 arturo: [codfw1dev] restart rabbitmq * 12:35 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 12:35 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:30 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 12:29 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:28 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 12:27 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:26 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 12:24 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 12:23 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 12:22 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2024-09-21 === * 10:17 dhinus: nova host-evacuate cloudvirt1063 ([[phab:T375223|T375223]]) * 09:59 dhinus: openstack aggregate remove host ceph cloudvirt1063 ([[phab:T375223|T375223]]) * 09:59 dhinus: openstack aggregate add host maintenance cloudvirt1063 ([[phab:T375223|T375223]]) * 09:42 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1063.eqiad.wmnet' * 09:41 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1063.eqiad.wmnet' === 2024-09-20 === * 11:00 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/50 * 11:00 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/50 * 10:59 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/50 * 10:59 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/50 * 08:27 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) ([[phab:T373740|T373740]]) * 08:27 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance ([[phab:T373740|T373740]]) === 2024-09-19 === * 15:46 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1048.eqiad.wmnet' ([[phab:T373740|T373740]]) * 15:33 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1048.eqiad.wmnet' ([[phab:T373740|T373740]]) * 15:32 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) ([[phab:T373740|T373740]]) * 15:31 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance ([[phab:T373740|T373740]]) * 10:03 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/51 * 10:03 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/51 * 09:56 wmbot~dcaro@urcuchillay: END (ERROR) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=97) ([[phab:T374043|T374043]]) * 09:55 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.undrain_node ([[phab:T374043|T374043]]) * 09:51 arturo: [codfw1dev] play with neutron default security group rules (delete, create them, etc) [[phab:T375111|T375111]] * 09:27 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/50 * 09:27 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/50 * 09:25 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/50 * 09:25 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/50 * 09:11 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/50 * 09:10 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/50 * 09:08 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/50 * 09:07 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/50 * 09:07 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/50 * 09:07 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/50 === 2024-09-18 === * 15:37 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/49 * 15:37 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/49 * 15:24 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/49 * 15:24 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/49 * 15:21 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/49 * 15:21 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/49 * 15:14 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/48 * 15:14 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/48 * 12:15 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/48 * 12:14 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/48 * 12:04 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 12:02 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:02 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 12:01 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 11:58 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/47 * 11:58 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/47 * 11:54 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 11:54 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 11:53 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 11:52 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 11:51 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 11:51 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 11:50 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 11:49 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 11:43 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/46 * 11:42 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/46 * 11:35 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/46 * 11:35 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/46 * 11:30 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/46 * 11:30 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/46 * 09:11 dcaro@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 08:59 dcaro@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 08:52 dcaro@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 08:40 dcaro@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 08:39 dcaro: restarted rabbitmq-server on all cloudrabbits * 08:27 dcaro@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 08:17 dcaro@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 08:12 dcaro@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 08:09 dcaro@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 08:09 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 08:09 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.openstack.restart_openstack === 2024-09-17 === * 21:24 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) ([[phab:T374043|T374043]]) * 16:24 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.undrain_node ([[phab:T374043|T374043]]) * 16:11 wmbot~dcaro@urcuchillay: END (ERROR) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=97) ([[phab:T374043|T374043]]) * 16:11 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.undrain_node ([[phab:T374043|T374043]]) * 15:08 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/46 * 15:07 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/46 * 15:05 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/46 * 15:05 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/46 * 15:03 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/46 * 15:02 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/46 * 14:32 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/46 * 14:31 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/46 * 14:27 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/46 * 14:27 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/46 === 2024-09-16 === * 14:30 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 14:29 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 14:28 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/45 * 14:28 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/45 * 14:28 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/45 * 14:27 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/45 * 11:28 arturo: [codfw1dev] created VM bastion-codfw1dev-04 to replace current bastion -03 ([[phab:T374828|T374828]]) === 2024-09-12 === * 10:51 arturo: merging change to keystone wmf hooks https://gerrit.wikimedia.org/r/c/operations/puppet/+/1071230 ([[phab:T374020|T374020]]) === 2024-09-11 === * 16:04 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 16:04 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 16:03 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 16:03 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 16:01 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 16:00 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 15:59 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/43 * 15:59 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/43 * 15:58 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/43 * 15:58 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/43 * 15:33 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 15:32 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 15:32 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 15:31 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 15:30 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/44 * 15:30 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/44 * 15:29 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/43 * 15:29 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/43 * 15:09 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 15:08 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 15:08 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 15:07 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:18 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/44 * 12:18 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/44 * 12:15 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/43 * 12:14 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/43 * 12:09 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/42 * 12:09 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/42 * 12:08 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/42 * 12:08 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/42 * 11:58 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 11:57 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 11:56 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 11:55 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 11:27 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/41 * 11:27 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/41 * 11:19 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 11:17 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 10:22 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 10:21 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 10:20 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 10:19 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 10:05 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 10:05 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 07:49 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt2004-dev.codfw.wmnet' ([[phab:T374467|T374467]]) * 07:41 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt2004-dev.codfw.wmnet' ([[phab:T374467|T374467]]) === 2024-09-10 === * 12:15 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 12:15 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 12:14 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 12:14 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 12:08 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 12:07 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 12:06 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 12:05 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 12:05 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 12:05 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 12:03 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 12:02 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 11:37 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 11:36 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 10:00 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 10:00 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 09:59 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 09:59 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 09:52 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 09:52 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 09:42 dcaro: hard-rebooting cloudvirt2004-dev (codfw1dev) having io/hardware issues * 09:22 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 09:22 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 09:20 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 09:19 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 09:07 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 09:07 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 09:07 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 09:07 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 09:03 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 09:03 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 09:00 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 08:59 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 08:54 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 08:54 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 08:46 dcaro@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T373986|T373986]]) * 08:45 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T373986|T373986]]) === 2024-09-09 === * 21:32 dcaro@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) ([[phab:T373986|T373986]]) * 16:36 dcaro: cleaned up dns leaks * 15:45 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 15:44 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 15:44 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 15:37 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 15:24 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/39 * 15:23 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/39 * 15:05 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T373986|T373986]]) * 15:04 dcaro@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T373986|T373986]]) * 14:14 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 14:13 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 13:06 arturo: merged change to cloudgw NAT setting https://gerrit.wikimedia.org/r/c/operations/puppet/+/1071189 * 12:39 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 12:39 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 12:29 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 12:29 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 11:37 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 11:36 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 11:35 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 11:35 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 10:56 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 10:55 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 10:55 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 10:54 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 10:54 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 10:54 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 10:12 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 10:12 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 10:11 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 10:11 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 09:59 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 09:59 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 09:58 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 09:58 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 09:51 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 09:51 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 09:49 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 09:49 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 09:44 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 09:44 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 09:42 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 09:42 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 09:40 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 09:39 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 09:38 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 09:37 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 09:15 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T373986|T373986]]) === 2024-09-06 === * 23:18 dcaro@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) ([[phab:T373986|T373986]]) * 18:17 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T373986|T373986]]) * 17:58 dcaro@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T373986|T373986]]) * 13:46 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T373986|T373986]]) * 11:13 dcaro@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T373986|T373986]]) * 07:27 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T373986|T373986]]) === 2024-09-05 === * 21:32 dcaro@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T373986|T373986]]) * 18:58 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1035.eqiad.wmnet' ([[phab:T374043|T374043]]) * 18:45 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1035.eqiad.wmnet' ([[phab:T374043|T374043]]) * 18:29 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1034.eqiad.wmnet' ([[phab:T374043|T374043]]) * 18:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1034.eqiad.wmnet' ([[phab:T374043|T374043]]) * 17:32 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T373986|T373986]]) * 17:22 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1033.eqiad.wmnet' ([[phab:T374043|T374043]]) * 17:08 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1033.eqiad.wmnet' ([[phab:T374043|T374043]]) * 17:08 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1032.eqiad.wmnet' ([[phab:T374043|T374043]]) * 16:48 dcaro@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T373986|T373986]]) * 16:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1032.eqiad.wmnet' ([[phab:T374043|T374043]]) * 16:42 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1031.eqiad.wmnet' ([[phab:T374043|T374043]]) * 16:21 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1031.eqiad.wmnet' ([[phab:T374043|T374043]]) * 14:25 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 14:24 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 14:19 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 14:17 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 14:16 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/37 * 14:15 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/37 * 13:04 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/37 * 13:03 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/37 * 13:02 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/37 * 13:01 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/37 * 12:51 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T373986|T373986]]) * 12:47 dcaro@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T373986|T373986]]) * 12:31 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 12:31 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 10:50 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 10:49 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 10:43 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/36 * 10:43 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/36 * 10:31 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 10:30 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 10:20 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/35 * 10:19 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/35 * 09:59 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 09:57 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 09:56 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 09:56 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 09:56 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 09:56 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 09:55 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/34 * 09:55 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/34 * 09:52 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/33 * 09:52 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/33 * 09:36 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 09:35 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 09:34 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/32 * 09:34 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/32 * 09:31 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/32 * 09:31 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/32 * 09:22 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 09:22 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 09:20 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/31 * 09:20 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/31 * 09:16 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 09:15 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 09:14 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 09:12 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 09:10 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/30 * 09:10 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/30 * 09:03 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/30 * 09:03 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/30 * 09:03 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/30 * 09:03 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/30 * 08:46 arturo: [codfw1dev] restart rabbitmq @ codfw1dev [[phab:T374002|T374002]] === 2024-09-04 === * 21:58 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 21:56 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 19:35 dcaro@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T373986|T373986]]) * 19:14 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 17:59 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 17:56 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 17:36 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 17:33 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 17:30 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 17:29 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 16:35 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T373986|T373986]]) * 15:30 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/30 * 15:30 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/30 * 15:25 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/30 * 15:24 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/30 * 15:17 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/30 * 15:17 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/30 * 12:18 arturo: [codfw1dev] restart rabbitmq-server.service on all 3 cloudcontrols, all nova-compute agents are down complaining about rabbitmq being unreachable === 2024-09-02 === * 13:27 dcaro@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 13:27 dcaro@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 13:23 dcaro@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 13:23 dcaro@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 13:19 dcaro@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 13:19 dcaro@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 13:07 dcaro@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 13:06 dcaro@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 12:48 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 12:48 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 12:41 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 12:40 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 12:40 aborrero@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.tofu (exit_code=97) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 12:40 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 11:47 dcaro@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/29 * 11:46 dcaro@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/29 * 11:42 dcaro@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 11:42 dcaro@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 11:37 dcaro@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 11:37 dcaro@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 11:37 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for main branch * 11:36 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 10:56 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 10:55 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 10:54 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 10:54 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 10:26 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 10:25 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 10:15 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 10:15 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 10:08 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 10:08 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 === 2024-08-31 === * 13:55 andrewbogott: moving tools-redis-7 off of cloudvirt1048 just in case [[phab:T373740|T373740]] * 13:39 andrewbogott: rebooting cloudvirt1048 from mgmt, it seems to have crashed === 2024-08-30 === * 12:01 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 11:59 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2024-08-29 === * 18:47 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 18:40 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 18:33 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add === 2024-08-25 === * 22:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 21:50 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 21:42 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 17:33 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 17:26 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2024-08-23 === * 14:48 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1062.eqiad.wmnet' ([[phab:T369044|T369044]]) * 14:42 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1062.eqiad.wmnet' ([[phab:T369044|T369044]]) * 00:30 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 00:20 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 00:10 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 00:08 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2024-08-22 === * 23:41 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 23:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 21:16 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudweb.set_maintenance (exit_code=0) ([[phab:T369044|T369044]]) * 21:14 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudweb.set_maintenance ([[phab:T369044|T369044]]) * 21:03 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 21:03 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 20:46 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=97) on host 'cloudvirt1047.eqiad.wmnet' ([[phab:T369044|T369044]]) * 20:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1047.eqiad.wmnet' ([[phab:T369044|T369044]]) * 20:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1038.eqiad.wmnet' ([[phab:T369044|T369044]]) * 20:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1038.eqiad.wmnet' ([[phab:T369044|T369044]]) * 20:36 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1042.eqiad.wmnet' ([[phab:T369044|T369044]]) * 20:29 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1042.eqiad.wmnet' ([[phab:T369044|T369044]]) * 20:29 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1044.eqiad.wmnet' ([[phab:T369044|T369044]]) * 20:22 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1044.eqiad.wmnet' ([[phab:T369044|T369044]]) * 20:22 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1041.eqiad.wmnet' ([[phab:T369044|T369044]]) * 20:16 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1041.eqiad.wmnet' ([[phab:T369044|T369044]]) * 20:15 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1046.eqiad.wmnet' ([[phab:T369044|T369044]]) * 20:09 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1046.eqiad.wmnet' ([[phab:T369044|T369044]]) * 20:09 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1043.eqiad.wmnet' ([[phab:T369044|T369044]]) * 20:06 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 20:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 20:02 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1043.eqiad.wmnet' ([[phab:T369044|T369044]]) * 20:02 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1045.eqiad.wmnet' ([[phab:T369044|T369044]]) * 19:56 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1045.eqiad.wmnet' ([[phab:T369044|T369044]]) * 19:56 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1040.eqiad.wmnet' ([[phab:T369044|T369044]]) * 19:49 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1040.eqiad.wmnet' ([[phab:T369044|T369044]]) * 19:49 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1036.eqiad.wmnet' ([[phab:T369044|T369044]]) * 19:42 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1036.eqiad.wmnet' ([[phab:T369044|T369044]]) * 19:42 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1034.eqiad.wmnet' ([[phab:T369044|T369044]]) * 19:35 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1034.eqiad.wmnet' ([[phab:T369044|T369044]]) * 19:35 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1039.eqiad.wmnet' ([[phab:T369044|T369044]]) * 19:29 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 19:29 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 19:28 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1039.eqiad.wmnet' ([[phab:T369044|T369044]]) * 19:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1037.eqiad.wmnet' ([[phab:T369044|T369044]]) * 19:21 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1037.eqiad.wmnet' ([[phab:T369044|T369044]]) * 19:21 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1035.eqiad.wmnet' ([[phab:T369044|T369044]]) * 19:17 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudweb.unset_maintenance (exit_code=0) ([[phab:T369044|T369044]]) * 19:15 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudweb.unset_maintenance ([[phab:T369044|T369044]]) * 19:14 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1035.eqiad.wmnet' ([[phab:T369044|T369044]]) * 19:14 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=97) on host 'cloudvirt1039' ([[phab:T369044|T369044]]) * 19:14 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1039' ([[phab:T369044|T369044]]) * 19:14 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=99) on host 'cloudvirt1037' ([[phab:T369044|T369044]]) * 19:14 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1037' ([[phab:T369044|T369044]]) * 19:14 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=99) on host 'cloudvirt1035' ([[phab:T369044|T369044]]) * 19:14 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1035' ([[phab:T369044|T369044]]) * 19:08 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1033.eqiad.wmnet' ([[phab:T369044|T369044]]) * 19:01 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1033.eqiad.wmnet' ([[phab:T369044|T369044]]) * 19:00 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=99) on host 'cloudvirt1033' ([[phab:T369044|T369044]]) * 19:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1033' ([[phab:T369044|T369044]]) * 19:00 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=99) on host 'cloudvirt1052' ([[phab:T369044|T369044]]) * 19:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1052' ([[phab:T369044|T369044]]) * 18:57 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 18:54 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 18:53 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudnet1005.eqiad.wmnet' ([[phab:T369044|T369044]]) * 18:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet1005.eqiad.wmnet' ([[phab:T369044|T369044]]) * 18:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudnet1006.eqiad.wmnet' ([[phab:T369044|T369044]]) * 18:33 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet1006.eqiad.wmnet' ([[phab:T369044|T369044]]) * 18:32 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol1007.eqiad.wmnet' ([[phab:T369044|T369044]]) * 18:13 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1007.eqiad.wmnet' ([[phab:T369044|T369044]]) * 18:10 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol1006.eqiad.wmnet' ([[phab:T369044|T369044]]) * 17:54 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1006.eqiad.wmnet' ([[phab:T369044|T369044]]) * 17:52 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol1005.eqiad.wmnet' ([[phab:T369044|T369044]]) * 17:42 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1005.eqiad.wmnet' ([[phab:T369044|T369044]]) * 17:38 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudcontrol1005.eqiad.wmnet' ([[phab:T369044|T369044]]) * 17:25 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1005.eqiad.wmnet' ([[phab:T369044|T369044]]) * 17:25 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudservices1006.eqiad.wmnet' ([[phab:T369044|T369044]]) * 17:15 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1006.eqiad.wmnet' ([[phab:T369044|T369044]]) * 17:13 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1006.eqiad.wmnet' ([[phab:T369044|T369044]]) * 17:04 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1006.eqiad.wmnet' ([[phab:T369044|T369044]]) * 17:02 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T369044|T369044]]) * 16:52 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T369044|T369044]]) * 16:52 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T369044|T369044]]) * 16:39 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T369044|T369044]]) * 16:38 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudweb.set_maintenance (exit_code=0) ([[phab:T369044|T369044]]) * 16:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudweb.set_maintenance ([[phab:T369044|T369044]]) === 2024-08-20 === * 03:18 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) === 2024-08-19 === * 23:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 23:28 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 23:17 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 23:17 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 23:17 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 23:16 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 23:16 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 23:16 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=97) * 18:47 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 18:47 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 18:39 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add === 2024-08-17 === * 03:23 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 03:22 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node === 2024-08-16 === * 16:32 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 16:32 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 16:32 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 16:32 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 16:32 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 16:32 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 16:32 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 16:31 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 16:31 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 16:31 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 16:31 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 16:31 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 16:31 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 16:31 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 16:31 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 16:31 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 16:31 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 16:31 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 16:31 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 16:30 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 16:29 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 16:29 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 16:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 16:28 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 15:31 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 15:20 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=97) * 15:19 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 15:15 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 15:15 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 15:10 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 15:10 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 15:09 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 15:09 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 15:09 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 15:09 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 15:09 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 15:09 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 15:09 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 15:09 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 15:09 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 15:09 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 15:08 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 15:08 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 15:08 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 15:08 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 15:08 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 15:07 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 15:07 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 15:07 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 15:02 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 14:42 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 14:40 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 14:40 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 14:40 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) * 14:38 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 14:37 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 14:37 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 14:36 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 14:36 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 14:36 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 14:36 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 14:36 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 14:36 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 14:36 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 14:36 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 14:36 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 14:36 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 14:34 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 14:34 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 14:34 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 14:33 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 14:33 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 14:33 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 14:30 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 14:30 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 14:29 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 14:29 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 14:28 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T363344|T363344]]) * 05:40 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T363344|T363344]]) * 04:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T363344|T363344]]) * 04:45 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T363344|T363344]]) * 04:42 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T363344|T363344]]) * 04:42 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T363344|T363344]]) * 04:33 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T363344|T363344]]) * 04:32 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T363344|T363344]]) * 04:26 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T363344|T363344]]) * 04:25 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T363344|T363344]]) * 03:58 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T363344|T363344]]) * 03:57 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T363344|T363344]]) * 03:53 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T363344|T363344]]) * 03:52 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T363344|T363344]]) * 03:51 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T363344|T363344]]) * 03:50 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T363344|T363344]]) * 03:45 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T363344|T363344]]) * 01:17 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 01:17 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 01:14 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 01:14 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 01:12 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 01:12 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 01:10 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 01:10 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 01:10 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 01:09 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 01:08 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 01:08 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 01:08 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 01:08 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 01:08 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 01:08 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 01:08 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 01:08 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 01:08 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 01:08 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node === 2024-08-15 === * 19:21 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 19:21 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 19:20 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 19:20 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 19:19 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 19:19 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 19:18 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 19:18 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 19:18 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 19:18 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 19:16 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 19:16 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 19:16 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 19:16 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 19:16 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 19:16 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 19:15 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 19:15 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 19:15 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 19:15 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 16:39 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 16:39 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 16:39 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 16:39 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 16:38 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 16:38 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 16:33 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.wait_for_rebalance (exit_code=99) * 16:32 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 16:32 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 16:31 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 16:31 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 16:30 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 16:30 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 16:29 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 16:29 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 16:29 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 16:28 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 16:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 16:28 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 16:27 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 16:27 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 16:27 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=97) * 16:27 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 10:00 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.wait_for_rebalance * 09:52 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.wait_for_rebalance (exit_code=0) * 08:56 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.wait_for_rebalance * 04:27 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 04:25 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 04:25 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 04:22 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 04:22 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 04:21 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 04:21 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 04:19 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 04:19 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 04:19 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 04:18 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 04:18 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 04:18 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 04:18 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 04:18 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 04:18 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 03:03 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 03:03 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 03:03 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 03:03 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 03:03 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 03:02 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 03:02 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 03:02 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 02:01 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 02:00 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 01:39 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 01:39 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 01:39 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 01:38 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 00:29 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 00:23 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 00:23 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 00:23 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 00:23 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 00:23 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 00:22 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 00:22 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 00:22 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node === 2024-08-14 === * 23:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 23:27 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 23:26 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 23:26 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 23:22 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 23:22 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 23:22 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 23:22 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 19:31 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 19:28 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 15:06 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=97) * 15:05 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 15:05 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=97) * 15:05 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 04:12 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 04:10 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2024-08-12 === * 15:44 dcaro@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=97) ([[phab:T363344|T363344]]) * 11:52 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.wait_for_rebalance (exit_code=99) * 11:51 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node ([[phab:T363344|T363344]]) * 11:51 dcaro@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=97) ([[phab:T363344|T363344]]) * 11:51 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node ([[phab:T363344|T363344]]) * 11:51 dcaro@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=97) ([[phab:T363344|T363344]]) * 11:51 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node ([[phab:T363344|T363344]]) * 08:52 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.wait_for_rebalance * 08:46 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.drain_rack (exit_code=0) * 08:43 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.drain_rack * 08:37 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.drain_rack (exit_code=99) * 08:37 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.drain_rack * 08:37 wmbot~dcaro@urcuchillay: END (ERROR) - Cookbook wmcs.ceph.osd.drain_rack (exit_code=97) ([[phab:T371878|T371878]]) * 08:37 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.drain_rack ([[phab:T371878|T371878]]) * 08:37 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.drain_rack (exit_code=99) * 08:33 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.drain_rack * 08:32 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.drain_rack (exit_code=99) ([[phab:T371878|T371878]]) * 08:27 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.drain_rack ([[phab:T371878|T371878]]) * 08:26 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.drain_rack (exit_code=99) ([[phab:T371878|T371878]]) * 08:26 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.drain_rack ([[phab:T371878|T371878]]) * 08:17 dcaro@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.drain_rack (exit_code=97) * 08:17 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_rack === 2024-08-09 === * 19:16 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) ([[phab:T371878|T371878]]) * 18:39 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) ([[phab:T371878|T371878]]) * 13:38 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T371878|T371878]]) * 13:38 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.drain_node (exit_code=97) ([[phab:T371878|T371878]]) * 13:36 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T371878|T371878]]) * 13:35 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.drain_node (exit_code=97) ([[phab:T371878|T371878]]) * 13:34 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T371878|T371878]]) * 11:27 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) ([[phab:T371878|T371878]]) * 05:36 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T371878|T371878]]) * 05:35 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T371878|T371878]]) * 00:47 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T371878|T371878]]) * 00:47 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.drain_node (exit_code=97) ([[phab:T371878|T371878]]) * 00:46 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T371878|T371878]]) === 2024-08-08 === * 23:05 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T371878|T371878]]) * 19:01 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 18:59 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 18:58 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudnet2006-dev.codfw.wmnet' * 18:49 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet2006-dev.codfw.wmnet' * 18:48 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudnet2005-dev.codfw.wmnet' * 18:39 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet2005-dev.codfw.wmnet' * 18:39 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.restart_openstack (exit_code=97) * 18:35 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 18:33 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt2006-dev.codfw.wmnet' * 18:28 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T371878|T371878]]) * 18:27 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T371878|T371878]]) * 18:26 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt2006-dev.codfw.wmnet' * 18:26 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T371878|T371878]]) * 18:26 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T371878|T371878]]) * 18:25 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T371878|T371878]]) * 18:25 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt2005-dev.codfw.wmnet' * 18:25 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T371878|T371878]]) * 18:18 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt2005-dev.codfw.wmnet' * 18:18 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt2004-dev.codfw.wmnet' * 18:14 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt2004-dev.codfw.wmnet' * 17:57 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=99) on host 'cloudvirt2004-dev.codfw.wmnet' * 17:52 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt2004-dev.codfw.wmnet' * 17:47 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol2006-dev.codfw.wmnet' * 17:31 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2006-dev.codfw.wmnet' * 17:11 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol2005-dev.codfw.wmnet' * 16:53 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2005-dev.codfw.wmnet' * 16:52 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol2004-dev.codfw.wmnet' * 16:51 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudbackup1001-dev.eqiad.wmnet' * 16:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudbackup1001-dev.eqiad.wmnet' * 16:44 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudbackup1002-dev.eqiad.wmnet' * 16:40 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2004-dev.codfw.wmnet' * 16:38 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudcontrol2004-dev.codfw.wmnet' * 16:37 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudbackup1002-dev.eqiad.wmnet' * 16:37 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudbackup1002-dev.codfw.wmnet' * 16:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudbackup1002-dev.codfw.wmnet' * 16:35 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudservices2004-dev.codfw.wmnet' * 16:34 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2004-dev.codfw.wmnet' * 16:27 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudcontrol2004-dev.codfw.wmnet' * 16:25 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices2004-dev.codfw.wmnet' * 16:25 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices2006-dev.codfw.wmnet' * 16:24 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices2006-dev.codfw.wmnet' * 16:23 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudservices2005-dev.codfw.wmnet' * 16:17 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2004-dev.codfw.wmnet' * 16:16 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudcontrol2004-dev.codfw.wmnet' * 16:13 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2004-dev.codfw.wmnet' * 16:12 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudcontrol2004-dev.codfw.wmnet' * 16:10 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2004-dev.codfw.wmnet' * 16:09 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices2005-dev.codfw.wmnet' * 16:03 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudcontrol2004-dev.codfw.wmnet' * 15:54 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices2005-dev.codfw.wmnet' * 15:54 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices2005-dev.codfw.wmnet' * 15:52 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2004-dev.codfw.wmnet' * 15:52 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudcontrol2004-dev.wikimedia.org' * 15:52 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2004-dev.wikimedia.org' * 15:51 andrewbogott: upgrading codfw1dev to openstack version caracal https://phabricator.wikimedia.org/T369044 * 15:51 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudweb.set_maintenance (exit_code=0) ([[phab:T369044|T369044]]) * 15:50 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudweb.set_maintenance ([[phab:T369044|T369044]]) * 15:49 wmbot~andrew@bullseye: END (FAIL) - Cookbook wmcs.openstack.cloudweb.set_maintenance (exit_code=99) ([[phab:T369044|T369044]]) * 15:49 wmbot~andrew@bullseye: START - Cookbook wmcs.openstack.cloudweb.set_maintenance ([[phab:T369044|T369044]]) * 15:00 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.wait_for_rebalance (exit_code=99) * 14:51 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.wait_for_rebalance * 14:01 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.wait_for_rebalance (exit_code=0) * 14:01 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.wait_for_rebalance (exit_code=0) * 13:02 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.wait_for_rebalance * 13:00 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=97) * 12:56 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.wait_for_rebalance * 12:20 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.wait_for_rebalance (exit_code=99) * 12:14 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.wait_for_rebalance * 11:35 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) ([[phab:T371878|T371878]]) * 11:25 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.undrain_node ([[phab:T371878|T371878]]) * 09:43 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 09:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 09:43 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.wait_for_rebalance (exit_code=0) * 07:44 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.wait_for_rebalance * 07:43 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.wait_for_rebalance (exit_code=99) * 06:41 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.wait_for_rebalance * 05:13 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 05:13 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 00:48 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 00:48 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 00:48 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) === 2024-08-07 === * 20:11 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 20:11 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=97) * 20:10 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 20:07 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 20:06 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 19:56 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 18:34 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 18:17 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 18:17 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 17:56 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 17:53 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.wait_for_rebalance (exit_code=0) * 17:19 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.wait_for_rebalance * 17:07 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 17:06 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.undrain_node * 17:06 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 17:05 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.undrain_node * 17:05 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 17:04 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.undrain_node * 17:03 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 17:03 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.undrain_node * 17:03 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 17:02 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.undrain_node * 17:02 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) * 16:56 wmbot~dcaro@urcuchillay: END (ERROR) - Cookbook wmcs.ceph.wait_for_rebalance (exit_code=97) * 16:50 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.drain_node * 16:50 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 16:49 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.undrain_node * 16:46 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) * 16:45 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.drain_node * 16:40 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 16:39 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.undrain_node * 16:39 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 16:38 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.undrain_node * 16:38 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 16:36 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.undrain_node * 16:35 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 16:34 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.undrain_node * 16:34 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 16:33 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.undrain_node * 16:32 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 16:31 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.undrain_node * 16:31 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 16:30 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.undrain_node * 16:28 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 16:28 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.undrain_node * 16:25 dcaro@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 16:24 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 16:02 dcaro@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 16:02 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 15:41 dcaro@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 15:41 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 08:27 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.wait_for_rebalance * 08:26 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) ([[phab:T371878|T371878]]) * 08:11 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T371878|T371878]]) * 06:39 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) ([[phab:T371878|T371878]]) * 03:08 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) ([[phab:T371878|T371878]]) * 01:18 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T371878|T371878]]) === 2024-08-06 === * 21:22 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T371878|T371878]]) * 19:51 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) ([[phab:T371878|T371878]]) * 18:36 wmbot~andrew@bullseye: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1042.eqiad.wmnet' * 18:21 wmbot~andrew@bullseye: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1042.eqiad.wmnet' * 18:19 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1043.eqiad.wmnet' * 18:17 wmbot~andrew@bullseye: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1041.eqiad.wmnet' * 18:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1043.eqiad.wmnet' * 18:03 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1044.eqiad.wmnet' * 17:59 wmbot~andrew@bullseye: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1041.eqiad.wmnet' * 17:58 wmbot~andrew@bullseye: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1040.eqiad.wmnet' * 17:48 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1044.eqiad.wmnet' * 17:47 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1045.eqiad.wmnet' * 17:45 wmbot~andrew@bullseye: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1040.eqiad.wmnet' * 17:43 wmbot~andrew@bullseye: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1039.eqiad.wmnet' * 17:41 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 17:40 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 17:34 wmbot~andrew@bullseye: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1039.eqiad.wmnet' * 17:32 wmbot~andrew@bullseye: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1038.eqiad.wmnet' * 17:31 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1045.eqiad.wmnet' * 17:31 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1046.eqiad.wmnet' * 17:21 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1046.eqiad.wmnet' * 17:21 wmbot~andrew@bullseye: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1038.eqiad.wmnet' * 17:14 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1047.eqiad.wmnet' * 17:14 wmbot~andrew@bullseye: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1037.eqiad.wmnet' * 17:03 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1047.eqiad.wmnet' * 17:02 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) * 17:01 wmbot~andrew@bullseye: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1037.eqiad.wmnet' * 17:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 17:00 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) * 17:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 17:00 wmbot~andrew@bullseye: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1036.eqiad.wmnet' * 17:00 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) * 16:59 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 16:59 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) * 16:58 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 16:58 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) * 16:58 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 16:58 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) * 16:57 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 16:57 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) * 16:56 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 16:56 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) * 16:56 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 16:56 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) * 16:55 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 16:55 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) * 16:55 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 16:54 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) * 16:54 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 16:48 wmbot~andrew@bullseye: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1036.eqiad.wmnet' * 16:47 wmbot~andrew@bullseye: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1036.eqiad.wmnet' * 16:47 wmbot~andrew@bullseye: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1036.eqiad.wmnet' * 16:37 wmbot~andrew@bullseye: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1036.eqiad.wmnet' * 16:37 wmbot~andrew@bullseye: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1036.eqiad.wmnet' * 16:08 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=99) * 16:07 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 16:07 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=99) * 16:07 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 16:07 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=99) * 16:06 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 16:06 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=99) * 16:06 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 16:06 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=99) * 16:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 16:05 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=99) * 16:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 16:05 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=99) * 16:04 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 16:04 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=99) * 16:04 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 16:04 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=99) * 16:03 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 16:03 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=99) * 16:03 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 16:03 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=99) * 16:02 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 16:01 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) * 16:01 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 15:59 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=99) * 15:59 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 15:44 dcaro@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) ([[phab:T371878|T371878]]) * 15:43 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T371878|T371878]]) * 15:36 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=99) * 15:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 15:29 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=99) * 15:29 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 15:25 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=99) * 15:25 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 15:22 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1036.eqiad.wmnet' * 15:22 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1036.eqiad.wmnet' * 15:21 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=99) * 15:21 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 15:21 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=99) * 15:20 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance === 2024-08-01 === * 13:44 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 13:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 13:09 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 13:09 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 03:04 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 03:04 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 03:04 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 03:04 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 02:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 02:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 01:40 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 01:40 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 01:32 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 01:31 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 01:30 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 01:29 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2024-07-31 === * 23:42 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 23:42 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 23:35 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 23:34 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 16:04 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 16:03 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 15:09 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 15:08 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 15:00 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 15:00 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 12:27 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 12:27 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 12:15 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 12:14 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 12:13 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 12:13 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 12:07 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 12:07 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 12:06 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 12:06 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 11:55 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 11:55 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 10:41 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/26 * 10:40 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/26 * 10:37 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/26 * 10:37 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/26 * 10:31 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/26 * 10:31 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/26 * 10:30 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/26 * 10:30 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/26 * 10:29 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/26 * 10:28 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/26 * 10:24 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/26 * 10:24 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/26 * 10:23 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/26 * 10:22 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/26 * 10:21 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/26 * 10:21 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/26 * 10:18 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/26 * 10:17 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/26 * 10:12 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/27 * 10:12 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/27 * 08:44 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/27 * 08:44 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/27 === 2024-07-30 === * 11:29 arturo: installing nova security updates ([[phab:T371240|T371240]]) === 2024-07-29 === * 18:55 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 18:52 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 16:22 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 16:17 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 16:16 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 16:11 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 16:10 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 16:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 15:45 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 15:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 15:12 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 15:09 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 14:49 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 14:45 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 14:39 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 14:38 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 14:34 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 14:33 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 14:31 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 14:29 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 14:29 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 14:28 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 14:27 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 14:26 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 14:26 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 14:25 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 11:38 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 11:36 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 11:28 arturo: [codfw1dev] restarting rabbitmq-server on all cloudcontrols, nova-compute cannot contact it * 11:00 arturo: [codfw1dev] installing nova security updates ([[phab:T371240|T371240]]) * 10:00 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 09:59 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 08:20 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/27 * 08:20 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/27 === 2024-07-25 === * 14:57 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 14:56 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 14:46 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/27 * 14:46 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/27 * 14:45 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/27 * 14:45 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/27 * 13:17 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/27 * 13:16 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/27 * 13:15 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/27 * 13:15 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/27 * 13:12 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/27 * 13:12 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/27 * 13:11 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/27 * 13:11 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/27 * 13:09 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/27 * 13:09 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/27 * 13:08 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/27 * 13:08 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/27 * 13:08 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/27 * 13:08 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/27 * 13:01 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/27 * 13:00 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/27 * 12:56 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/27 * 12:56 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/27 * 12:55 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/27 * 12:55 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/27 * 11:40 arturo: manually restart maintain-dbusers in cloudcontrol1005 to see if that makes any difference * 11:05 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 11:03 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 11:02 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/26 * 11:01 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/26 * 11:01 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/26 * 11:00 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/26 * 10:58 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/26 * 10:58 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/26 * 10:57 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/26 * 10:57 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/26 * 10:57 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/26 * 10:56 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/26 * 10:49 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/26 * 10:48 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/26 * 10:48 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/26 * 10:47 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/26 * 09:35 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 09:32 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 08:57 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/23 * 08:57 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/23 * 08:52 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/23 * 08:52 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/23 * 08:42 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 08:42 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 08:29 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/25 * 08:28 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/25 * 08:27 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/25 * 08:27 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/25 * 08:25 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/25 * 08:25 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/25 * 08:01 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/23 * 08:01 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/23 * 07:58 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/23 * 07:58 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/23 === 2024-07-24 === * 16:01 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/23 * 16:01 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/23 * 16:01 aborrero@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.tofu (exit_code=97) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/24 * 16:01 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/24 * 15:59 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/24 * 15:59 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/24 * 15:57 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/23 * 15:57 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/23 * 15:56 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/23 * 15:56 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/23 * 15:49 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/23 * 15:49 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/23 * 15:46 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/23 * 15:45 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/23 * 15:43 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/23 * 15:43 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/23 * 15:39 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/23 * 15:38 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/23 * 15:37 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/23 * 15:37 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/23 * 15:34 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/23 * 15:34 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/23 * 15:31 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/23 * 15:31 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/23 * 15:29 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/23 * 15:29 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/23 * 15:28 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/23 * 15:28 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/23 * 11:59 arturo: restarted maintain-dbusers.service in cloudcontrol1005, it was stuck doing nothing * 10:45 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/24 * 10:45 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/24 * 10:44 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/24 * 10:44 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/24 * 10:43 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/24 * 10:43 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/24 * 10:42 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/24 * 10:42 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/24 === 2024-07-23 === * 13:02 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 13:02 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 13:01 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/22 * 13:01 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/22 * 13:00 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/22 * 13:00 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/22 * 12:57 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/22 * 12:56 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/22 * 12:47 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 12:47 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 11:35 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 11:35 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 10:59 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/22 * 10:59 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/22 * 10:52 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/22 * 10:52 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/22 * 10:51 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/22 * 10:51 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/22 * 10:50 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/22 * 10:50 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/22 * 10:48 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/22 * 10:48 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/22 * 10:47 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/22 * 10:47 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/22 * 10:09 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 10:08 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 10:07 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/21 * 10:07 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/21 * 09:39 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 08:54 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 08:54 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 08:24 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 08:24 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 05:59 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 05:59 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 05:29 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 05:29 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 03:03 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 03:03 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 02:30 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 02:30 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 00:04 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 00:04 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) === 2024-07-22 === * 23:34 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 23:33 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 21:07 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 21:07 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 20:37 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 20:36 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 18:10 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 18:10 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 17:39 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 17:39 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 16:08 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/21 * 16:07 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/21 * 16:05 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/21 * 16:05 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/21 * 15:51 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 15:50 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 15:47 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 15:46 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 15:46 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 15:46 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 15:35 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 15:35 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 15:31 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 15:31 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 15:10 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 15:10 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 14:40 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 14:40 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 14:04 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 14:04 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 14:02 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 14:02 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 13:41 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 13:41 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 13:34 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 13:33 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 13:28 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 13:28 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 13:14 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 13:14 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 13:11 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 13:11 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 13:10 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 13:09 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 13:06 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 13:06 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 13:04 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 13:04 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 13:04 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 13:04 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 12:56 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 12:55 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 12:49 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 12:48 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 12:47 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 12:47 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 12:43 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 12:43 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 12:35 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 12:35 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 12:31 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 12:31 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 12:12 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 12:12 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 11:41 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 11:41 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 11:41 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 11:40 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 11:39 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 11:39 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 11:36 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 11:36 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 11:35 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 11:35 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 11:34 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 11:34 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 11:31 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 11:31 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 11:29 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for main branch * 11:29 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 09:14 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 09:14 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 08:44 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy === 2024-07-21 === * 10:02 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 09:37 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 09:37 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 09:07 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 09:07 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 06:40 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 06:40 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 03:45 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 03:45 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 03:15 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 03:14 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 00:50 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 00:50 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 00:20 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 00:20 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) === 2024-07-20 === * 21:56 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 21:56 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 21:26 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 21:26 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 18:59 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 18:59 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 18:29 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 18:29 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 16:04 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 16:04 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 15:34 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 15:34 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 13:07 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 13:07 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 12:37 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 12:37 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 10:11 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 10:11 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 09:39 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 09:39 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 07:14 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 07:14 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 06:44 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 06:44 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 04:17 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 04:17 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 03:47 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 03:47 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 01:21 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 01:21 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 00:51 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 00:51 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) === 2024-07-19 === * 22:22 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 22:22 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 21:51 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 21:51 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 19:21 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 19:21 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 18:51 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 18:51 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 16:25 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 16:25 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 16:12 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 16:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 15:53 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 15:53 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 13:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 13:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 13:26 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 13:26 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 12:55 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 12:55 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 10:26 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 10:26 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 09:56 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 09:56 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 07:28 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 07:28 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 06:58 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 06:58 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 04:32 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 04:32 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 04:02 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 04:02 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 01:35 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 01:35 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 01:04 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 01:04 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) === 2024-07-18 === * 22:38 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 22:38 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 22:22 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 22:19 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 22:08 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 22:08 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 19:39 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 19:39 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 19:08 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 19:08 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 16:40 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 16:40 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 16:10 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 16:10 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 13:44 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 13:44 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 13:13 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 13:13 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 12:31 dhinus: upgrade spicerack from 8.5 to 8.8 on cloudcumin* * 10:45 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 10:44 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 10:14 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 10:14 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 07:47 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 07:47 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 07:17 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 07:16 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 04:51 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 04:51 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 04:21 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 04:21 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 03:48 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 03:47 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 01:55 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 01:55 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 01:25 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 01:25 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) === 2024-07-17 === * 22:58 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 22:58 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 22:28 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 22:28 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 20:00 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 20:00 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 19:30 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 19:29 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 17:02 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 17:02 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 16:32 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 16:32 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 14:05 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 14:05 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 13:35 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 13:34 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 11:07 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 11:07 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 10:36 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 10:36 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 08:10 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 08:10 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 07:39 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 07:39 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 05:13 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 05:13 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 04:43 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 04:43 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 02:17 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 02:17 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 01:46 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 01:46 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) === 2024-07-16 === * 23:21 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 23:21 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 22:50 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 22:50 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 20:23 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 20:23 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 19:53 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 19:52 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 17:26 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 17:26 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 16:56 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 16:56 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 14:28 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 14:28 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 13:57 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 13:41 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.wait_for_rebalance (exit_code=0) * 11:20 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.wait_for_rebalance * 11:19 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 11:14 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 11:02 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 11:02 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 10:57 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 10:25 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 10:21 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.wait_for_rebalance (exit_code=0) * 10:00 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.wait_for_rebalance * 09:46 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 09:46 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 09:36 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 09:36 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 09:36 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 09:31 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 09:19 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 09:19 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 09:19 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 09:18 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 09:18 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 09:15 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 09:15 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 09:10 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 08:50 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 08:44 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 08:44 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 08:42 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 08:42 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 08:37 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 08:35 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 08:30 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 08:30 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 08:27 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 08:27 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 08:22 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add === 2024-07-15 === * 22:33 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 22:33 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 20:40 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 20:29 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 20:29 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 20:24 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 18:59 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 18:59 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 17:33 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 17:10 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 17:06 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 16:50 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 14:13 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 14:08 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 14:07 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 14:07 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy === 2024-07-11 === * 13:42 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 13:41 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 12:34 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 12:00 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 11:58 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 11:57 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 11:44 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 11:43 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 02:15 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1061.eqiad.wmnet' * 02:12 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1061.eqiad.wmnet' * 02:07 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1061.eqiad.wmnet' * 02:07 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1061.eqiad.wmnet' * 02:02 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1061.eqiad.wmnet' * 02:02 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1061.eqiad.wmnet' * 02:02 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1061.eqiad.wmnet' * 02:02 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1061.eqiad.wmnet' * 01:53 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1061.eqiad.wmnet' * 01:53 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1061.eqiad.wmnet' * 01:51 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1061.eqiad.wmnet' * 01:51 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1061.eqiad.wmnet' * 01:50 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1061.eqiad.wmnet' * 01:50 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1061.eqiad.wmnet' * 01:48 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1061.eqiad.wmnet' * 01:48 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1061.eqiad.wmnet' * 01:47 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1061.eqiad.wmnet' * 01:47 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1061.eqiad.wmnet' * 01:44 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1061.eqiad.wmnet' * 01:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1061.eqiad.wmnet' * 01:43 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1061.eqiad.wmnet' * 01:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1061.eqiad.wmnet' * 01:43 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1061.eqiad.wmnet' * 01:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1061.eqiad.wmnet' * 01:36 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1061.eqiad.wmnet' * 01:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1061.eqiad.wmnet' === 2024-07-08 === * 17:36 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) ([[phab:T309789|T309789]]) * 17:10 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T309789|T309789]]) * 14:22 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) ([[phab:T309789|T309789]]) * 13:01 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy ([[phab:T309789|T309789]]) === 2024-07-06 === * 14:06 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 14:04 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2024-07-05 === * 10:33 arturo: aborrero@cloudcephmon1001:~$ sudo ceph osd unset norebalance * 10:31 arturo: aborrero@cloudcephmon1001:~$ sudo ceph osd unset noin * 08:56 arturo: installing nova/glance/cinder security updates [[phab:T369138|T369138]] === 2024-07-04 === * 20:26 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T309789|T309789]]) * 20:16 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T309789|T309789]]) * 20:16 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T309789|T309789]]) * 14:52 dcaro: rebooting cloudcontrol1007 due to systemd-journal service failing to start * 09:28 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T309789|T309789]]) === 2024-07-03 === * 17:28 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) ([[phab:T309789|T309789]]) * 16:04 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy ([[phab:T309789|T309789]]) * 12:28 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T309789|T309789]]) * 12:22 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T309789|T309789]]) === 2024-07-02 === * 19:17 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T309789|T309789]]) * 14:24 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T309789|T309789]]) === 2024-07-01 === * 14:28 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) ([[phab:T309789|T309789]]) * 14:24 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy ([[phab:T309789|T309789]]) * 12:17 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) ([[phab:T309789|T309789]]) * 12:03 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy ([[phab:T309789|T309789]]) === 2024-06-27 === * 22:50 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T309789|T309789]]) * 19:05 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1059.eqiad.wmnet' * 19:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1059.eqiad.wmnet' * 19:01 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1059.eqiad.wmnet' * 18:49 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1059.eqiad.wmnet' * 18:48 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1064.eqiad.wmnet' * 18:46 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1064.eqiad.wmnet' * 17:16 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1067.eqiad.wmnet' * 17:15 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1067.eqiad.wmnet' * 17:10 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1058.eqiad.wmnet' * 17:06 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1058.eqiad.wmnet' * 17:05 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1066.eqiad.wmnet' * 17:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1066.eqiad.wmnet' * 17:04 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1065.eqiad.wmnet' * 17:02 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1065.eqiad.wmnet' * 17:01 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1057.eqiad.wmnet' * 16:58 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1057.eqiad.wmnet' * 15:39 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T309789|T309789]]) * 15:36 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T309789|T309789]]) * 15:35 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T309789|T309789]]) * 15:22 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T309789|T309789]]) * 15:21 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T309789|T309789]]) * 15:21 wmbot~dcaro@urcuchillay: END (ERROR) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=97) ([[phab:T309789|T309789]]) * 15:21 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T309789|T309789]]) * 13:44 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) ([[phab:T309789|T309789]]) * 12:44 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirtlocal1001.eqiad.wmnet' * 12:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirtlocal1001.eqiad.wmnet' * 12:07 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy ([[phab:T309789|T309789]]) * 12:06 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.unset_cluster_maintenance (exit_code=0) * 12:05 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.unset_cluster_maintenance * 12:05 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.set_cluster_in_maintenance (exit_code=0) * 12:04 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.set_cluster_in_maintenance * 10:32 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.unset_cluster_maintenance (exit_code=0) * 10:32 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.unset_cluster_maintenance * 10:31 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.set_cluster_in_maintenance (exit_code=0) * 10:31 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.set_cluster_in_maintenance * 10:31 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.set_cluster_in_maintenance (exit_code=99) * 10:30 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.set_cluster_in_maintenance * 10:30 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.set_cluster_in_maintenance (exit_code=99) * 10:30 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.set_cluster_in_maintenance * 10:30 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.set_cluster_in_maintenance (exit_code=99) * 10:30 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.set_cluster_in_maintenance * 10:29 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.set_cluster_in_maintenance (exit_code=99) * 10:29 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.set_cluster_in_maintenance * 10:29 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.set_cluster_in_maintenance (exit_code=99) * 10:28 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.set_cluster_in_maintenance * 08:36 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.set_cluster_in_maintenance (exit_code=99) * 08:36 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.set_cluster_in_maintenance * 07:55 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.set_cluster_in_maintenance (exit_code=99) * 07:55 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.set_cluster_in_maintenance === 2024-06-26 === * 22:32 andrewbogott: disabled all g3.* flavors in eqiad1 * 19:18 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T309789|T309789]]) * 17:44 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 17:42 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 15:46 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T309789|T309789]]) * 15:18 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T309789|T309789]]) * 15:09 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T309789|T309789]]) * 14:46 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T309789|T309789]]) * 14:45 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T309789|T309789]]) * 14:45 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T309789|T309789]]) * 14:45 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T309789|T309789]]) * 11:16 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) ([[phab:T309789|T309789]]) * 10:03 dcaro: taking cloudcephosd1006 out of the pool ([[phab:T348643|T348643]]) * 10:00 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy ([[phab:T309789|T309789]]) * 10:00 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) ([[phab:T309789|T309789]]) * 09:59 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy ([[phab:T309789|T309789]]) === 2024-06-25 === * 22:55 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt2003-dev.codfw.wmnet' * 22:55 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt2003-dev.codfw.wmnet' * 22:55 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt2002-dev.codfw.wmnet' * 22:54 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt2002-dev.codfw.wmnet' * 22:37 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt2001-dev.codfw.wmnet' * 22:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt2001-dev.codfw.wmnet' * 22:35 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt2001-dev.codw.wmnet' * 22:35 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt2001-dev.codw.wmnet' * 21:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt2005-dev.codfw.wmnet' * 21:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt2005-dev.codfw.wmnet' * 16:26 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt2004-dev.codfw.wmnet' * 16:21 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt2004-dev.codfw.wmnet' * 16:15 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=97) on host 'cloudvirt2006-dev.codfw.wmnet' * 16:15 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt2006-dev.codfw.wmnet' * 03:42 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 03:39 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 03:17 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 03:12 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 03:12 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 03:06 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2024-06-24 === * 19:11 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1056.eqiad.wmnet' * 18:56 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1056.eqiad.wmnet' * 17:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1055.eqiad.wmnet' * 17:17 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1055.eqiad.wmnet' * 16:20 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1054.eqiad.wmnet' * 16:04 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1054.eqiad.wmnet' === 2024-06-21 === * 10:36 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.vm_console (exit_code=99) * 10:36 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.vm_console * 10:35 wmbot~arturo@nostromo: END (ERROR) - Cookbook wmcs.openstack.cloudvirt.vm_console (exit_code=97) * 10:35 wmbot~arturo@nostromo: START - Cookbook wmcs.openstack.cloudvirt.vm_console * 10:35 wmbot~arturo@nostromo: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.vm_console (exit_code=99) * 10:35 wmbot~arturo@nostromo: START - Cookbook wmcs.openstack.cloudvirt.vm_console * 09:43 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 09:43 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 08:31 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1053.eqiad.wmnet' ([[phab:T368129|T368129]]) * 08:28 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1053.eqiad.wmnet' ([[phab:T368129|T368129]]) * 04:21 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.migrate_server_to_ovs (exit_code=99) for server 4e612eb8-04e1-4541-941d-{{Gerrit|a05519eed60a}} * 04:14 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.migrate_server_to_ovs for server 4e612eb8-04e1-4541-941d-{{Gerrit|a05519eed60a}} === 2024-06-20 === * 21:38 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.migrate_server_to_ovs (exit_code=99) for server 14789ac1-bc06-4677-9bb0-{{Gerrit|66c16c887427}} * 21:38 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.migrate_server_to_ovs for server 14789ac1-bc06-4677-9bb0-{{Gerrit|66c16c887427}} * 17:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1053.eqiad.wmnet' * 17:13 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1053.eqiad.wmnet' * 14:38 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 14:38 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 14:08 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1052.eqiad.wmnet' * 14:02 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 14:02 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 14:02 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 14:01 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 13:48 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1052.eqiad.wmnet' * 13:47 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1051.eqiad.wmnet' * 13:47 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.migrate_server_to_ovs (exit_code=0) for server 614f9c99-86f1-410f-8ef5-{{Gerrit|e33d23215ff5}} * 13:45 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.migrate_server_to_ovs for server 614f9c99-86f1-410f-8ef5-{{Gerrit|e33d23215ff5}} * 13:25 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1051.eqiad.wmnet' * 13:00 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1050.eqiad.wmnet' * 12:41 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1050.eqiad.wmnet' * 12:40 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1049.eqiad.wmnet' * 12:37 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 12:36 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 12:34 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 12:34 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 11:50 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1049.eqiad.wmnet' * 11:45 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1048.eqiad.wmnet' * 11:13 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1048.eqiad.wmnet' * 11:11 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1047.eqiad.wmnet' * 10:49 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1047.eqiad.wmnet' * 10:35 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 10:35 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 10:34 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 10:34 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 09:38 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1046.eqiad.wmnet' * 09:14 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1046.eqiad.wmnet' * 09:10 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1045.eqiad.wmnet' * 08:55 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1045.eqiad.wmnet' * 08:37 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 08:37 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 04:49 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.migrate_server_to_ovs (exit_code=99) for server 77140f83-1a12-43b7-9e47-{{Gerrit|0e779503a525}} * 04:49 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.migrate_server_to_ovs for server 77140f83-1a12-43b7-9e47-{{Gerrit|0e779503a525}} * 04:49 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.migrate_server_to_ovs (exit_code=0) for server d0b1d9d5-1aec-4a05-a2d7-{{Gerrit|6d8522a365dc}} * 04:48 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.migrate_server_to_ovs for server d0b1d9d5-1aec-4a05-a2d7-{{Gerrit|6d8522a365dc}} * 04:48 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.migrate_server_to_ovs (exit_code=0) for server 2d2e8925-b50f-483a-82ee-{{Gerrit|e6a1c588e5be}} * 04:47 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.migrate_server_to_ovs for server 2d2e8925-b50f-483a-82ee-{{Gerrit|e6a1c588e5be}} * 04:47 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.migrate_server_to_ovs (exit_code=0) for server 138be95d-93ad-4a85-9245-{{Gerrit|a0a508711555}} * 04:46 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.migrate_server_to_ovs (exit_code=99) for server 270b6533-dc99-4e5d-a642-{{Gerrit|c61138b11891}} * 04:46 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.migrate_server_to_ovs for server 270b6533-dc99-4e5d-a642-{{Gerrit|c61138b11891}} * 04:45 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.migrate_server_to_ovs for server 138be95d-93ad-4a85-9245-{{Gerrit|a0a508711555}} * 04:45 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.migrate_server_to_ovs (exit_code=0) for server a3e945dc-3548-47fa-8ce3-{{Gerrit|bf1426ff3b15}} * 04:44 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.migrate_server_to_ovs (exit_code=0) for server 0258b810-29af-448a-af5e-{{Gerrit|ed39e19286df}} * 04:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.migrate_server_to_ovs for server a3e945dc-3548-47fa-8ce3-{{Gerrit|bf1426ff3b15}} * 04:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.migrate_server_to_ovs (exit_code=0) for server 37e23659-4516-4bd8-a9be-{{Gerrit|4cc55def5560}} * 04:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.migrate_server_to_ovs for server 0258b810-29af-448a-af5e-{{Gerrit|ed39e19286df}} * 04:43 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.migrate_server_to_ovs (exit_code=99) for server e94378df-7a48-4b11-b44a-{{Gerrit|bf69aaf132bd}} * 04:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.migrate_server_to_ovs for server e94378df-7a48-4b11-b44a-{{Gerrit|bf69aaf132bd}} * 04:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.migrate_server_to_ovs (exit_code=0) for server 7da824f3-c9f3-4460-b9b2-{{Gerrit|2a894268a7ca}} * 04:42 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.migrate_server_to_ovs for server 37e23659-4516-4bd8-a9be-{{Gerrit|4cc55def5560}} * 04:42 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.migrate_server_to_ovs (exit_code=0) for server 63b82d38-3026-408f-8dcc-{{Gerrit|0ecbd1a3c870}} * 04:42 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.migrate_server_to_ovs for server 7da824f3-c9f3-4460-b9b2-{{Gerrit|2a894268a7ca}} * 04:42 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.migrate_server_to_ovs (exit_code=0) for server 7cb371bb-a53a-4e65-a1cf-{{Gerrit|f1a8264a9166}} * 04:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.migrate_server_to_ovs for server 63b82d38-3026-408f-8dcc-{{Gerrit|0ecbd1a3c870}} * 04:41 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.migrate_server_to_ovs (exit_code=99) for server 6616fcbf-a49e-4e03-b735-{{Gerrit|84d31b535c08}} * 04:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.migrate_server_to_ovs for server 6616fcbf-a49e-4e03-b735-{{Gerrit|84d31b535c08}} * 04:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.migrate_server_to_ovs for server 7cb371bb-a53a-4e65-a1cf-{{Gerrit|f1a8264a9166}} * 04:41 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.migrate_server_to_ovs (exit_code=0) for server 6f1171db-9d7d-466f-aaa7-{{Gerrit|cb14e1a6af41}} * 04:40 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.migrate_server_to_ovs (exit_code=99) for server 77140f83-1a12-43b7-9e47-{{Gerrit|0e779503a525}} * 04:39 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.migrate_server_to_ovs for server 6f1171db-9d7d-466f-aaa7-{{Gerrit|cb14e1a6af41}} * 04:39 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.migrate_server_to_ovs (exit_code=0) for server 1dc3fcee-cf27-4351-ad60-{{Gerrit|384478624ea3}} * 04:38 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.migrate_server_to_ovs for server 77140f83-1a12-43b7-9e47-{{Gerrit|0e779503a525}} * 04:38 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.migrate_server_to_ovs for server 1dc3fcee-cf27-4351-ad60-{{Gerrit|384478624ea3}} * 04:38 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.migrate_server_to_ovs (exit_code=0) for server 3df62375-e75d-4068-9c71-{{Gerrit|13519b6bf927}} * 04:37 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.migrate_server_to_ovs (exit_code=0) for server 3af9b40f-d29d-4216-9b64-{{Gerrit|0ebb10f94c7c}} * 04:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.migrate_server_to_ovs for server 3df62375-e75d-4068-9c71-{{Gerrit|13519b6bf927}} * 04:36 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.migrate_server_to_ovs (exit_code=99) for server 270b6533-dc99-4e5d-a642-{{Gerrit|c61138b11891}} * 04:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.migrate_server_to_ovs for server 270b6533-dc99-4e5d-a642-{{Gerrit|c61138b11891}} * 04:36 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.migrate_server_to_ovs (exit_code=99) for server ae9fa949-e4ce-4ffb-ad9f-{{Gerrit|4e5a3812d031}} * 04:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.migrate_server_to_ovs for server ae9fa949-e4ce-4ffb-ad9f-{{Gerrit|4e5a3812d031}} * 04:36 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.migrate_server_to_ovs (exit_code=0) for server 452dc8d3-6ee0-412c-90ba-{{Gerrit|66f21f9d09c1}} * 04:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.migrate_server_to_ovs for server 3af9b40f-d29d-4216-9b64-{{Gerrit|0ebb10f94c7c}} * 04:35 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.migrate_server_to_ovs for server 452dc8d3-6ee0-412c-90ba-{{Gerrit|66f21f9d09c1}} * 04:34 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.migrate_server_to_ovs (exit_code=0) for server e94378df-7a48-4b11-b44a-{{Gerrit|bf69aaf132bd}} * 04:33 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.migrate_server_to_ovs for server e94378df-7a48-4b11-b44a-{{Gerrit|bf69aaf132bd}} * 04:33 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.migrate_server_to_ovs (exit_code=0) for server 47b0da1d-e50a-42c1-8cd9-{{Gerrit|dc255cc2f1a3}} * 04:32 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.migrate_server_to_ovs for server 47b0da1d-e50a-42c1-8cd9-{{Gerrit|dc255cc2f1a3}} * 02:37 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1063.eqiad.wmnet' ([[phab:T368007|T368007]]) * 02:29 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1063.eqiad.wmnet' ([[phab:T368007|T368007]]) * 02:05 andrewbogott: cloudvirt1063 is unresponsive, cycling power from racadm === 2024-06-19 === * 15:43 taavi@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=97) on host 'cloudvirt1044.eqiad.wmnet' * 15:40 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1044.eqiad.wmnet' * 15:40 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=99) * 15:39 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 15:31 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 15:30 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 15:30 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 15:30 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 15:30 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 15:29 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 12:27 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1044.eqiad.wmnet' * 12:10 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1044.eqiad.wmnet' * 12:10 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1043.eqiad.wmnet' * 11:51 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1043.eqiad.wmnet' * 11:49 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1042.eqiad.wmnet' * 11:33 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1042.eqiad.wmnet' * 01:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.migrate_server_to_ovs (exit_code=0) for server f667d3c2-379a-48d7-ad44-{{Gerrit|4f3933bdb871}} * 01:45 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.migrate_server_to_ovs for server f667d3c2-379a-48d7-ad44-{{Gerrit|4f3933bdb871}} * 01:45 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.migrate_server_to_ovs (exit_code=0) for server 71e4296d-039d-4452-9d92-{{Gerrit|69b9f8eb3aba}} * 01:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.migrate_server_to_ovs for server 71e4296d-039d-4452-9d92-{{Gerrit|69b9f8eb3aba}} * 01:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.migrate_server_to_ovs (exit_code=0) for server a7725cd2-6162-41a3-8add-{{Gerrit|4dc0668b233b}} * 01:42 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.migrate_server_to_ovs for server a7725cd2-6162-41a3-8add-{{Gerrit|4dc0668b233b}} * 01:17 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.migrate_server_to_ovs (exit_code=99) for server 02bb9b5a-cadf-4bee-9b63-{{Gerrit|519b1e9b485b}} * 01:17 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.migrate_server_to_ovs for server 02bb9b5a-cadf-4bee-9b63-{{Gerrit|519b1e9b485b}} * 01:02 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.migrate_server_to_ovs (exit_code=0) for server 02bb9b5a-cadf-4bee-9b63-{{Gerrit|519b1e9b485b}} * 01:01 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.migrate_server_to_ovs for server 02bb9b5a-cadf-4bee-9b63-{{Gerrit|519b1e9b485b}} === 2024-06-18 === * 21:18 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 21:07 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 20:58 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.reboot_node (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudcontrol1006.eqiad.wmnet<nowiki>}</nowiki>' * 20:53 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.reboot_node on hosts matched by 'D<nowiki>{</nowiki>cloudcontrol1006.eqiad.wmnet<nowiki>}</nowiki>' === 2024-06-17 === * 20:25 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1041.eqiad.wmnet' ([[phab:T364457|T364457]]) * 20:03 andrewbogott: repaced ovs hosts in the 'ceph' aggregate * 19:55 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1041.eqiad.wmnet' ([[phab:T364457|T364457]]) * 19:55 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1040.eqiad.wmnet' ([[phab:T364457|T364457]]) * 19:32 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1040.eqiad.wmnet' ([[phab:T364457|T364457]]) * 19:32 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1039.eqiad.wmnet' ([[phab:T364457|T364457]]) * 19:28 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1039.eqiad.wmnet' ([[phab:T364457|T364457]]) * 18:16 andrewbogott: temporarily removing all ovs hosts from the 'ceph' aggregate so the scheduler will stop putting linuxbridge hosts on ovs hosts and breaking them * 17:50 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1038.eqiad.wmnet' ([[phab:T364457|T364457]]) * 17:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1038.eqiad.wmnet' ([[phab:T364457|T364457]]) * 17:29 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1037.eqiad.wmnet' ([[phab:T364457|T364457]]) * 17:17 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1037.eqiad.wmnet' ([[phab:T364457|T364457]]) * 13:52 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 13:52 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 12:34 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1036.eqiad.wmnet' * 12:25 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1036.eqiad.wmnet' * 12:03 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 12:02 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 10:53 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1035.eqiad.wmnet' * 10:40 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1035.eqiad.wmnet' === 2024-06-14 === * 14:11 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 14:11 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 13:18 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1034.eqiad.wmnet' * 13:04 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1034.eqiad.wmnet' === 2024-06-13 === * 16:15 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 16:15 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 13:59 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1033.eqiad.wmnet' * 13:40 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1033.eqiad.wmnet' * 13:29 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt2003-dev.codfw.wmnet' * 13:24 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt2003-dev.codfw.wmnet' === 2024-06-12 === * 11:49 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1031.eqiad.wmnet' * 11:47 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1031.eqiad.wmnet' * 11:37 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1032.eqiad.wmnet' * 11:29 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1032.eqiad.wmnet' * 10:08 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1031.eqiad.wmnet' * 09:50 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1031.eqiad.wmnet' === 2024-06-11 === * 15:55 dcaro: restarting mon service on cloudcephmon1002 to try to release the 2 stuck ops left * 13:30 taavi: pin all existing eqiad1 flavors to linuxbridge hypervisors [[phab:T364458|T364458]] * 12:28 taavi: add all existing eqiad1 cloudvirts to new network-linuxbridge aggregate [[phab:T364458|T364458]] * 10:10 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 10:04 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 10:00 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 09:59 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 09:43 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 09:35 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 09:34 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 09:33 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 09:33 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 09:33 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2024-06-07 === * 19:00 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 18:47 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2024-06-05 === * 11:57 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 11:41 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add === 2024-06-04 === * 15:49 taavi: drop hopefully-unused 68.10.in-addr.arpa. from designate [[phab:T361220|T361220]] === 2024-05-30 === * 02:18 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 02:15 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2024-05-29 === * 18:54 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=99) * 18:54 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 18:54 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=99) ([[phab:T364984|T364984]]) * 18:54 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance ([[phab:T364984|T364984]]) * 18:54 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=99) ([[phab:T364984|T364984]]) * 18:53 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance ([[phab:T364984|T364984]]) === 2024-05-28 === * 19:59 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 19:53 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 19:52 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 19:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 19:38 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 19:37 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 19:28 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 19:27 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 19:24 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 19:24 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 19:22 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 19:21 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 19:21 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 19:17 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 19:16 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 19:14 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 18:59 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 18:57 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 18:48 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 18:47 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 14:17 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 14:16 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 14:07 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 14:06 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2024-05-21 === * 11:28 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudnet.migrate_to_ovs * 08:18 taavi: stop neutron services on cloudnet1005 [[phab:T364459|T364459]] === 2024-05-20 === * 14:46 wmbot~taavi@runko: END (PASS) - Cookbook wmcs.openstack.cloudnet.migrate_to_ovs (exit_code=0) * 14:46 wmbot~taavi@runko: START - Cookbook wmcs.openstack.cloudnet.migrate_to_ovs * 14:23 wmbot~taavi@runko: END (PASS) - Cookbook wmcs.openstack.cloudnet.migrate_to_ovs (exit_code=0) * 14:23 wmbot~taavi@runko: START - Cookbook wmcs.openstack.cloudnet.migrate_to_ovs * 14:20 wmbot~taavi@runko: END (PASS) - Cookbook wmcs.openstack.cloudnet.migrate_to_ovs (exit_code=0) * 14:19 wmbot~taavi@runko: START - Cookbook wmcs.openstack.cloudnet.migrate_to_ovs * 14:18 wmbot~taavi@runko: END (FAIL) - Cookbook wmcs.openstack.cloudnet.migrate_to_ovs (exit_code=99) * 14:18 wmbot~taavi@runko: START - Cookbook wmcs.openstack.cloudnet.migrate_to_ovs * 14:17 wmbot~taavi@runko: END (FAIL) - Cookbook wmcs.openstack.cloudnet.migrate_to_ovs (exit_code=99) * 14:17 wmbot~taavi@runko: START - Cookbook wmcs.openstack.cloudnet.migrate_to_ovs * 14:16 wmbot~taavi@runko: END (FAIL) - Cookbook wmcs.openstack.cloudnet.migrate_to_ovs (exit_code=99) * 14:16 wmbot~taavi@runko: START - Cookbook wmcs.openstack.cloudnet.migrate_to_ovs * 14:15 wmbot~taavi@runko: END (FAIL) - Cookbook wmcs.openstack.cloudnet.migrate_to_ovs (exit_code=99) * 14:15 wmbot~taavi@runko: START - Cookbook wmcs.openstack.cloudnet.migrate_to_ovs === 2024-05-16 === * 08:55 taavi: delete 'monitoring' project https://phabricator.wikimedia.org/T365105 === 2024-05-15 === * 09:08 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1041.eqiad.wmnet' ([[phab:T319184|T319184]]) * 08:57 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1041.eqiad.wmnet' ([[phab:T319184|T319184]]) === 2024-05-14 === * 19:26 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 19:24 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 18:14 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 18:14 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 18:04 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 18:02 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 00:20 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 00:19 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 00:18 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 00:15 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2024-05-10 === * 13:23 andrewbogott: deploying updated 2024.1 Horizon === 2024-05-09 === * 19:05 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 19:02 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 18:57 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 18:55 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 18:49 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 18:46 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2024-04-25 === * 21:05 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1032.eqiad.wmnet' ([[phab:T356287|T356287]]) * 21:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1032.eqiad.wmnet' ([[phab:T356287|T356287]]) * 21:00 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1031.eqiad.wmnet' ([[phab:T356287|T356287]]) * 20:54 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1031.eqiad.wmnet' ([[phab:T356287|T356287]]) * 20:54 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1063.eqiad.wmnet' ([[phab:T356287|T356287]]) * 20:50 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1063.eqiad.wmnet' ([[phab:T356287|T356287]]) * 20:50 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1066.eqiad.wmnet' ([[phab:T356287|T356287]]) * 20:45 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1066.eqiad.wmnet' ([[phab:T356287|T356287]]) * 20:45 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1064.eqiad.wmnet' ([[phab:T356287|T356287]]) * 20:40 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1064.eqiad.wmnet' ([[phab:T356287|T356287]]) * 20:40 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1067.eqiad.wmnet' ([[phab:T356287|T356287]]) * 20:35 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1067.eqiad.wmnet' ([[phab:T356287|T356287]]) * 20:35 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1065.eqiad.wmnet' ([[phab:T356287|T356287]]) * 20:31 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1065.eqiad.wmnet' ([[phab:T356287|T356287]]) * 20:30 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1062.eqiad.wmnet' ([[phab:T356287|T356287]]) * 20:25 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1062.eqiad.wmnet' ([[phab:T356287|T356287]]) * 20:25 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirtlocal1003.eqiad.wmnet' ([[phab:T356287|T356287]]) * 20:20 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirtlocal1003.eqiad.wmnet' ([[phab:T356287|T356287]]) * 20:20 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirtlocal1002.eqiad.wmnet' ([[phab:T356287|T356287]]) * 20:15 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirtlocal1002.eqiad.wmnet' ([[phab:T356287|T356287]]) * 20:15 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirtlocal1001.eqiad.wmnet' ([[phab:T356287|T356287]]) * 20:10 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirtlocal1001.eqiad.wmnet' ([[phab:T356287|T356287]]) * 20:10 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1054.eqiad.wmnet' ([[phab:T356287|T356287]]) * 20:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1054.eqiad.wmnet' ([[phab:T356287|T356287]]) * 20:05 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1055.eqiad.wmnet' ([[phab:T356287|T356287]]) * 20:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1055.eqiad.wmnet' ([[phab:T356287|T356287]]) * 20:00 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1060.eqiad.wmnet' ([[phab:T356287|T356287]]) * 19:55 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1060.eqiad.wmnet' ([[phab:T356287|T356287]]) * 19:55 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1058.eqiad.wmnet' ([[phab:T356287|T356287]]) * 19:51 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1058.eqiad.wmnet' ([[phab:T356287|T356287]]) * 19:51 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1059.eqiad.wmnet' ([[phab:T356287|T356287]]) * 19:46 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1059.eqiad.wmnet' ([[phab:T356287|T356287]]) * 19:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1061.eqiad.wmnet' ([[phab:T356287|T356287]]) * 19:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1061.eqiad.wmnet' ([[phab:T356287|T356287]]) * 19:41 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1057.eqiad.wmnet' ([[phab:T356287|T356287]]) * 19:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1057.eqiad.wmnet' ([[phab:T356287|T356287]]) * 19:36 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1056.eqiad.wmnet' ([[phab:T356287|T356287]]) * 19:30 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1056.eqiad.wmnet' ([[phab:T356287|T356287]]) * 19:30 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1051.eqiad.wmnet' ([[phab:T356287|T356287]]) * 19:26 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1051.eqiad.wmnet' ([[phab:T356287|T356287]]) * 19:26 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1050.eqiad.wmnet' ([[phab:T356287|T356287]]) * 19:20 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1050.eqiad.wmnet' ([[phab:T356287|T356287]]) * 19:20 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1049.eqiad.wmnet' ([[phab:T356287|T356287]]) * 19:15 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1049.eqiad.wmnet' ([[phab:T356287|T356287]]) * 19:15 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1053.eqiad.wmnet' ([[phab:T356287|T356287]]) * 19:10 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1053.eqiad.wmnet' ([[phab:T356287|T356287]]) * 19:10 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1052.eqiad.wmnet' ([[phab:T356287|T356287]]) * 19:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1052.eqiad.wmnet' ([[phab:T356287|T356287]]) * 19:05 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1048.eqiad.wmnet' ([[phab:T356287|T356287]]) * 19:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1048.eqiad.wmnet' ([[phab:T356287|T356287]]) * 19:00 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt-wdqs1003.eqiad.wmnet' ([[phab:T356287|T356287]]) * 18:55 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt-wdqs1003.eqiad.wmnet' ([[phab:T356287|T356287]]) * 18:55 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt-wdqs1002.eqiad.wmnet' ([[phab:T356287|T356287]]) * 18:49 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt-wdqs1002.eqiad.wmnet' ([[phab:T356287|T356287]]) * 18:49 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt-wdqs1001.eqiad.wmnet' ([[phab:T356287|T356287]]) * 18:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt-wdqs1001.eqiad.wmnet' ([[phab:T356287|T356287]]) * 18:44 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1047.eqiad.wmnet' ([[phab:T356287|T356287]]) * 18:39 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1047.eqiad.wmnet' ([[phab:T356287|T356287]]) * 18:39 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1038.eqiad.wmnet' ([[phab:T356287|T356287]]) * 18:35 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1038.eqiad.wmnet' ([[phab:T356287|T356287]]) * 18:35 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1042.eqiad.wmnet' ([[phab:T356287|T356287]]) * 18:30 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1042.eqiad.wmnet' ([[phab:T356287|T356287]]) * 18:30 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1044.eqiad.wmnet' ([[phab:T356287|T356287]]) * 18:25 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1044.eqiad.wmnet' ([[phab:T356287|T356287]]) * 18:25 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1041.eqiad.wmnet' ([[phab:T356287|T356287]]) * 18:20 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1041.eqiad.wmnet' ([[phab:T356287|T356287]]) * 18:20 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1046.eqiad.wmnet' ([[phab:T356287|T356287]]) * 18:15 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1046.eqiad.wmnet' ([[phab:T356287|T356287]]) * 18:15 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1043.eqiad.wmnet' ([[phab:T356287|T356287]]) * 18:10 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1043.eqiad.wmnet' ([[phab:T356287|T356287]]) * 18:10 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1045.eqiad.wmnet' ([[phab:T356287|T356287]]) * 18:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1045.eqiad.wmnet' ([[phab:T356287|T356287]]) * 18:05 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1040.eqiad.wmnet' ([[phab:T356287|T356287]]) * 18:01 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1040.eqiad.wmnet' ([[phab:T356287|T356287]]) * 18:00 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1036.eqiad.wmnet' ([[phab:T356287|T356287]]) * 17:56 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1036.eqiad.wmnet' ([[phab:T356287|T356287]]) * 17:56 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1034.eqiad.wmnet' ([[phab:T356287|T356287]]) * 17:51 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1034.eqiad.wmnet' ([[phab:T356287|T356287]]) * 17:51 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1039.eqiad.wmnet' ([[phab:T356287|T356287]]) * 17:46 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1039.eqiad.wmnet' ([[phab:T356287|T356287]]) * 17:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1037.eqiad.wmnet' ([[phab:T356287|T356287]]) * 17:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1037.eqiad.wmnet' ([[phab:T356287|T356287]]) * 17:41 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1035.eqiad.wmnet' ([[phab:T356287|T356287]]) * 17:38 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 17:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1035.eqiad.wmnet' ([[phab:T356287|T356287]]) * 17:35 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1033.eqiad.wmnet' ([[phab:T356287|T356287]]) * 17:32 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 17:32 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudnet1006.eqiad.wmnet' ([[phab:T356287|T356287]]) * 17:30 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1033.eqiad.wmnet' ([[phab:T356287|T356287]]) * 17:29 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=99) on host 'cloudvirt1033' ([[phab:T356287|T356287]]) * 17:29 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1033' ([[phab:T356287|T356287]]) * 17:24 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet1006.eqiad.wmnet' ([[phab:T356287|T356287]]) * 17:21 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudnet1005.eqiad.wmnet' ([[phab:T356287|T356287]]) * 17:13 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet1005.eqiad.wmnet' ([[phab:T356287|T356287]]) * 17:12 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol1007.eqiad.wmnet' ([[phab:T356287|T356287]]) * 17:10 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudservices1006.eqiad.wmnet' ([[phab:T356287|T356287]]) * 17:03 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1006.eqiad.wmnet' ([[phab:T356287|T356287]]) * 17:03 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T356287|T356287]]) * 16:56 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T356287|T356287]]) * 16:52 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1007.eqiad.wmnet' ([[phab:T356287|T356287]]) * 16:52 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol1006.eqiad.wmnet' ([[phab:T356287|T356287]]) * 16:37 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1006.eqiad.wmnet' ([[phab:T356287|T356287]]) * 16:35 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol1005.eqiad.wmnet' ([[phab:T356287|T356287]]) * 16:02 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1005.eqiad.wmnet' ([[phab:T356287|T356287]]) * 16:00 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudweb.set_maintenance (exit_code=99) ([[phab:T356287|T356287]]) * 16:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudweb.set_maintenance ([[phab:T356287|T356287]]) === 2024-04-23 === * 09:15 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt2002-dev.codfw.wmnet' * 09:10 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt2002-dev.codfw.wmnet' === 2024-04-18 === * 12:58 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 12:58 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 12:55 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudvirt2001-dev.codfw.wmnet' * 12:52 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudnet2008-dev.codfw.wmnet' * 12:46 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudvirt2001-dev.codfw.wmnet' * 12:46 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet2008-dev.codfw.wmnet' === 2024-04-17 === * 14:12 dcaro: deleting dns leaks * 01:58 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 01:54 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 01:48 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt2006-dev.codfw.wmnet' ([[phab:T356287|T356287]]) * 01:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt2006-dev.codfw.wmnet' ([[phab:T356287|T356287]]) * 01:40 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt2005-dev.codfw.wmnet' ([[phab:T356287|T356287]]) * 01:35 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt2005-dev.codfw.wmnet' ([[phab:T356287|T356287]]) * 01:05 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt2004-dev.codfw.wmnet' ([[phab:T356287|T356287]]) * 01:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt2004-dev.codfw.wmnet' ([[phab:T356287|T356287]]) === 2024-04-16 === * 21:52 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudnet2006-dev.codfw.wmnet' ([[phab:T356287|T356287]]) * 21:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet2006-dev.codfw.wmnet' ([[phab:T356287|T356287]]) * 21:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudnet2005-dev.codfw.wmnet' ([[phab:T356287|T356287]]) * 21:35 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet2005-dev.codfw.wmnet' ([[phab:T356287|T356287]]) * 21:31 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol2005-dev.codfw.wmnet' ([[phab:T356287|T356287]]) * 21:15 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2005-dev.codfw.wmnet' ([[phab:T356287|T356287]]) * 21:03 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol2004-dev.codfw.wmnet' ([[phab:T356287|T356287]]) * 20:47 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2004-dev.codfw.wmnet' ([[phab:T356287|T356287]]) * 20:12 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol2001-dev.codfw.wmnet' ([[phab:T356287|T356287]]) * 19:56 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2001-dev.codfw.wmnet' ([[phab:T356287|T356287]]) * 19:55 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudservices2005-dev.codfw.wmnet' ([[phab:T356287|T356287]][A) * 19:47 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices2005-dev.codfw.wmnet' ([[phab:T356287|T356287]][A) * 19:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudservices2004-dev.codfw.wmnet' ([[phab:T356287|T356287]][A) * 19:20 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices2004-dev.codfw.wmnet' ([[phab:T356287|T356287]][A) * 19:13 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices2004-dev.codfw.wmnet' ([[phab:T356287|T356287]][A) * 19:09 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices2004-dev.codfw.wmnet' ([[phab:T356287|T356287]][A) * 19:09 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices2004.codfw.wmnet' ([[phab:T356287|T356287]]) * 19:09 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices2004.codfw.wmnet' ([[phab:T356287|T356287]]) * 19:09 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices2004.eqiad.wmnet' ([[phab:T356287|T356287]]) * 19:09 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices2004.eqiad.wmnet' ([[phab:T356287|T356287]]) * 19:06 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudweb.set_maintenance (exit_code=99) ([[phab:T356287|T356287]]) * 19:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudweb.set_maintenance ([[phab:T356287|T356287]]) === 2024-04-15 === * 11:19 taavi: update spicerack to 8.5.0 on cloudcumin2001 === 2024-04-10 === * 20:10 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 20:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2024-04-09 === * 18:13 andrewbogott: rebooting cloudinfra-cloudvps-puppetserver-1; unresponsive * 13:20 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 13:18 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2024-04-05 === * 14:37 taavi: run maintain-replica-indexes on all web replicas [[phab:T361945|T361945]] * 14:27 taavi: run maintain-replica-indexes on remaining analytics replicas [[phab:T361945|T361945]] * 14:17 taavi: run maintain-replica-indexes on clouddb1017 [[phab:T361945|T361945]] === 2024-04-04 === * 18:23 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 18:19 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2024-04-03 === * 19:24 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 19:22 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 19:19 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 19:17 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 19:16 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 19:15 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 15:01 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 15:01 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 14:16 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1040.eqiad.wmnet' ([[phab:T319184|T319184]]) * 14:09 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1040.eqiad.wmnet' ([[phab:T319184|T319184]]) * 12:40 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 12:40 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 11:48 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1039.eqiad.wmnet' ([[phab:T319184|T319184]]) * 11:34 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1039.eqiad.wmnet' ([[phab:T319184|T319184]]) * 11:33 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 11:33 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 10:37 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1038.eqiad.wmnet' ([[phab:T319184|T319184]]) * 10:25 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1038.eqiad.wmnet' ([[phab:T319184|T319184]]) * 10:21 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 10:21 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 09:26 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1037.eqiad.wmnet' ([[phab:T319184|T319184]]) * 09:17 taavi: manually delete prometheus-node-textfile-wmcs-dnsleaks.service and related files from cloudservices1005/6, leftovers of the designate api to cloudcontrol migration * 09:09 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1037.eqiad.wmnet' ([[phab:T319184|T319184]]) * 08:52 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 08:52 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance === 2024-04-02 === * 15:01 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 15:00 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 15:00 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 15:00 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 15:00 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1036.eqiad.wmnet' ([[phab:T319184|T319184]]) * 14:43 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1036.eqiad.wmnet' ([[phab:T319184|T319184]]) * 13:00 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 13:00 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 12:37 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 12:37 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 12:36 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 12:36 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 11:45 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1035.eqiad.wmnet' ([[phab:T319184|T319184]]) * 11:32 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1035.eqiad.wmnet' ([[phab:T319184|T319184]]) === 2024-03-27 === * 11:47 taavi: deleting about 2k stale puppet certs by running wmcs-puppetcertleaks in delete mode === 2024-03-22 === * 10:25 dcaro: adding back cloudcephosd1034 to the pool after doing the performance tests ([[phab:T348643|T348643]]) === 2024-03-21 === * 22:07 andrewbogott: doing dist-upgrade on cloudcontrol nodes to get mariadb upgraded for [[phab:T357133|T357133]] * 22:05 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudcontrol2004-dev.codfw.wmnet' ([[phab:T357133|T357133]]) * 22:04 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2004-dev.codfw.wmnet' ([[phab:T357133|T357133]]) * 11:08 dcaro: restarting nova-api on cloudcontrol1007 === 2024-03-20 === * 17:02 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T348643|T348643]]) * 15:39 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T348643|T348643]]) * 15:27 dcaro: turning off cloudcephosd1030 to swap some disks ([[phab:T348643|T348643]]) * 15:25 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T348643|T348643]]) * 15:02 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T348643|T348643]]) * 15:00 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) ([[phab:T348643|T348643]]) * 14:31 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T348643|T348643]]) === 2024-03-19 === * 18:28 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T348643|T348643]]) * 15:49 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T348643|T348643]]) * 15:21 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T348643|T348643]]) * 13:30 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T348643|T348643]]) === 2024-03-18 === * 16:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 16:24 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 16:22 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 16:22 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2024-03-09 === * 16:43 andrewbogott: restarted nova-api on cloudcontrol1006 === 2024-03-08 === * 11:27 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) * 11:27 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 11:27 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 11:27 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 11:24 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) * 11:23 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 11:21 arturo: restarted nova-api in cloudcontrol1007, it was complaining about mysql broken pipe * 11:16 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 11:16 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 11:16 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) * 11:15 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 11:15 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 11:15 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 10:36 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) * 10:36 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 10:35 taavi@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=97) * 10:35 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 10:35 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 10:35 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance === 2024-03-07 === * 12:23 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt2001-dev.codfw.wmnet' * 12:16 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt2001-dev.codfw.wmnet' * 12:16 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) * 12:16 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance === 2024-03-06 === * 17:46 dhinus: running "wmcs-dnsleaks --delete" to clean up 2 leaked records (tools-sgeweblight-10-32) * 15:36 dcaro: renewing puppet ca cert for cloud-puppetmaster-03 * 15:24 dcaro: renewing puppet ca cert for cloudinfra-internal puppetmaster === 2024-03-04 === * 14:56 dhinus: delete project "loggerdiscordbot" in favor of new project "discordbots" [[phab:T358337|T358337]],[[phab:T358427|T358427]] * 12:54 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.reboot_node (exit_code=0) ([[phab:T359049|T359049]]) * 12:48 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.reboot_node ([[phab:T359049|T359049]]) === 2024-03-01 === * 14:59 taavi: removing wmf-auto-restart-cron from all VMs without cron via cumin - https://gerrit.wikimedia.org/r/c/operations/puppet/+/1007328/ [[phab:T358343|T358343]] * 12:07 dcaro: restarted nova-api on cloudcontrol100* as it was very slow * 12:04 dcaro: restarted nova-api on cloudcontrol1005 as it was very slow === 2024-02-27 === * 18:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 18:26 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2024-02-26 === * 10:02 arturo: deleting nskaggs account from gerrit's wmcs-trusted group === 2024-02-22 === * 13:58 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 13:58 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 12:58 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1034.eqiad.wmnet' ([[phab:T319184|T319184]]) * 12:57 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1034.eqiad.wmnet' ([[phab:T319184|T319184]]) * 12:53 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1034.eqiad.wmnet' ([[phab:T319184|T319184]]) * 12:32 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1034.eqiad.wmnet' ([[phab:T319184|T319184]]) * 11:55 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 11:54 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 09:01 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 09:00 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2024-02-21 === * 13:44 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) ([[phab:T319184|T319184]]) * 13:43 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance ([[phab:T319184|T319184]]) * 13:20 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) ([[phab:T319184|T319184]]) * 13:19 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance ([[phab:T319184|T319184]]) * 12:50 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1033.eqiad.wmnet' ([[phab:T319184|T319184]]) * 12:50 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1033.eqiad.wmnet' ([[phab:T319184|T319184]]) * 12:03 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1033.eqiad.wmnet' ([[phab:T319184|T319184]]) * 11:44 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1033.eqiad.wmnet' ([[phab:T319184|T319184]]) === 2024-02-20 === * 11:45 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 11:45 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 11:45 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=99) * 11:45 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 11:30 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.post-reimage (exit_code=0) preparing cloudvirt cloudvirt1032.eqiad.wmnet for duty (nova discovery, canary VM) Pending aggregates though. ([[phab:T319184|T319184]]) * 11:30 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.post-reimage preparing cloudvirt cloudvirt1032.eqiad.wmnet for duty (nova discovery, canary VM) Pending aggregates though. ([[phab:T319184|T319184]]) === 2024-02-19 === * 12:33 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.post-reimage (exit_code=99) preparing cloudvirt cloudvirt1032.eqiad.wmnet for duty (nova discovery, canary VM) Pending aggregates though. ([[phab:T319184|T319184]]) * 12:32 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.post-reimage preparing cloudvirt cloudvirt1032.eqiad.wmnet for duty (nova discovery, canary VM) Pending aggregates though. ([[phab:T319184|T319184]]) * 12:02 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.post-reimage (exit_code=99) preparing cloudvirt cloudvirt1032.eqiad.wmnet for duty (nova discovery, canary VM) Pending aggregates though. ([[phab:T319184|T319184]]) * 12:02 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.post-reimage preparing cloudvirt cloudvirt1032.eqiad.wmnet for duty (nova discovery, canary VM) Pending aggregates though. ([[phab:T319184|T319184]]) * 12:00 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.post-reimage (exit_code=99) preparing cloudvirt cloudvirt1032.eqiad.wmnet for duty (nova discovery, canary VM) Pending aggregates though. ([[phab:T319184|T319184]]) * 12:00 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.post-reimage preparing cloudvirt cloudvirt1032.eqiad.wmnet for duty (nova discovery, canary VM) Pending aggregates though. ([[phab:T319184|T319184]]) * 10:09 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.pre-reimage (exit_code=0) prepare cloudvirt1032.eqiad.wmnet for reimage (drain, remove nova agent, etc) ([[phab:T319184|T319184]]) * 09:52 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.pre-reimage prepare cloudvirt1032.eqiad.wmnet for reimage (drain, remove nova agent, etc) ([[phab:T319184|T319184]]) * 09:49 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.pre-reimage (exit_code=99) prepare cloudvirt1032.eqiad.wmnet for reimage (drain, remove nova agent, etc) ([[phab:T319184|T319184]]) * 09:49 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.pre-reimage prepare cloudvirt1032.eqiad.wmnet for reimage (drain, remove nova agent, etc) ([[phab:T319184|T319184]]) === 2024-02-15 === * 17:34 wmbot~fran@wmf3169: START - Cookbook wmcs.openstack.roll_reboot_cloudnets ([[phab:T356975|T356975]]) * 15:34 taavi: restart radosgw in eqiad as I am seeing 500 errors * 14:22 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 14:22 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 14:22 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=99) * 14:22 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 12:05 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1031.eqiad.wmnet' ([[phab:T319184|T319184]]) * 11:58 dhinus: restore correct aggregate "localdisk" for cloudvirtlocal1001 and remove "maintenance" * 11:52 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1031.eqiad.wmnet' ([[phab:T319184|T319184]]) * 05:51 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1044.eqiad.wmnet<nowiki>}</nowiki>' * 05:41 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1045.eqiad.wmnet<nowiki>}</nowiki>' * 05:19 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1044.eqiad.wmnet<nowiki>}</nowiki>' * 05:18 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1047.eqiad.wmnet<nowiki>}</nowiki>' * 05:16 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1046.eqiad.wmnet<nowiki>}</nowiki>' * 04:52 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1047.eqiad.wmnet<nowiki>}</nowiki>' * 04:52 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1046.eqiad.wmnet[B<nowiki>}</nowiki>' * 04:51 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1046.eqiad.wmnet[B<nowiki>}</nowiki>' * 04:50 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1049.eqiad.wmnet<nowiki>}</nowiki>' * 04:48 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1048.eqiad.wmnet<nowiki>}</nowiki>' * 04:45 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1049.eqiad.wmnet<nowiki>}</nowiki>' * 04:44 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1049.eqiad.wmnet<nowiki>}</nowiki>' * 04:33 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1049.eqiad.wmnet<nowiki>}</nowiki>' * 04:33 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1048.eqiad.wmnet<nowiki>}</nowiki>' * 04:32 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1051.eqiad.wmnet<nowiki>}</nowiki>' * 04:30 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1050.eqiad.wmnet<nowiki>}</nowiki>' * 04:10 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1051.eqiad.wmnet<nowiki>}</nowiki>' * 04:10 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1050.eqiad.wmnet<nowiki>}</nowiki>' * 04:09 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1052.eqiad.wmnet<nowiki>}</nowiki>' * 04:03 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1053.eqiad.wmnet<nowiki>}</nowiki>' * 03:47 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1052.eqiad.wmnet<nowiki>}</nowiki>' * 03:47 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1054.eqiad.wmnet<nowiki>}</nowiki>' * 03:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1054.eqiad.wmnet<nowiki>}</nowiki>' * 03:39 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1053.eqiad.wmnet<nowiki>}</nowiki>' * 03:30 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1054.eqiad.wmnet<nowiki>}</nowiki>' * 03:30 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1055.eqiad.wmnet<nowiki>}</nowiki>' * 03:23 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1054.eqiad.wmnet<nowiki>}</nowiki>' * 03:21 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1054.eqiad.wmnet<nowiki>}</nowiki>' * 03:08 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1054.eqiad.wmnet<nowiki>}</nowiki>' * 03:08 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1055.eqiad.wmnet<nowiki>}</nowiki>' === 2024-02-14 === * 22:00 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1057.eqiad.wmnet<nowiki>}</nowiki>' * 21:59 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1056.eqiad.wmnet<nowiki>}</nowiki>' * 21:38 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1057.eqiad.wmnet<nowiki>}</nowiki>' * 21:38 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1056.eqiad.wmnet<nowiki>}</nowiki>' * 21:36 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1058.eqiad.wmnet<nowiki>}</nowiki>' * 21:21 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1059.eqiad.wmnet<nowiki>}</nowiki>' * 21:10 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1058.eqiad.wmnet<nowiki>}</nowiki>' * 21:08 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1060.eqiad.wmnet<nowiki>}</nowiki>' * 20:59 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1059.eqiad.wmnet<nowiki>}</nowiki>' * 20:57 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1061.eqiad.wmnet<nowiki>}</nowiki>' * 20:49 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1060.eqiad.wmnet<nowiki>}</nowiki>' * 20:44 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1060.eqiad.wmnet<nowiki>}</nowiki>' * 20:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1060.eqiad.wmnet<nowiki>}</nowiki>' * 20:41 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1060.eqiad.wmnet<nowiki>}</nowiki>' * 20:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1061.eqiad.wmnet<nowiki>}</nowiki>' * 20:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1060.eqiad.wmnet<nowiki>}</nowiki>' * 20:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1062.eqiad.wmnet<nowiki>}</nowiki>' * 20:16 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1064.eqiad.wmnet<nowiki>}</nowiki>' * 20:08 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1062.eqiad.wmnet<nowiki>}</nowiki>' * 20:07 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1065.eqiad.wmnet<nowiki>}</nowiki>' * 20:03 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 20:03 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 20:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1064.eqiad.wmnet<nowiki>}</nowiki>' * 20:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1065.eqiad.wmnet<nowiki>}</nowiki>' * 19:59 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1066.eqiad.wmnet<nowiki>}</nowiki>' * 19:59 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1067.eqiad.wmnet<nowiki>}</nowiki>' * 19:56 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1066.eqiad.wmnet<nowiki>}</nowiki>' * 19:55 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1067.eqiad.wmnet<nowiki>}</nowiki>' * 19:55 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1063.eqiad.wmnet<nowiki>}</nowiki>' * 19:51 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1063.eqiad.wmnet<nowiki>}</nowiki>' * 19:51 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=97) on hosts matched by 'D<nowiki>{</nowiki>cloudvirtXXXX.eqiad.wmnet<nowiki>}</nowiki>' * 19:51 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirtXXXX.eqiad.wmnet<nowiki>}</nowiki>' * 19:49 wmbot~andrew@bullseye: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) * 19:49 wmbot~andrew@bullseye: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot * 19:49 wmbot~andrew@bullseye: END (ERROR) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=97) * 19:49 wmbot~andrew@bullseye: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot * 19:48 wmbot~andrew@bullseye: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) * 19:45 wmbot~andrew@bullseye: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) * 19:45 wmbot~andrew@bullseye: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot * 19:43 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1063.eqiad.wmnet<nowiki>}</nowiki>' * 19:40 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1063.eqiad.wmnet<nowiki>}</nowiki>' * 19:38 wmbot~andrew@bullseye: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) * 19:38 wmbot~andrew@bullseye: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot * 19:36 wmbot~andrew@bullseye: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) * 19:36 wmbot~andrew@bullseye: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot * 19:34 wmbot~andrew@bullseye: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) * 19:34 wmbot~andrew@bullseye: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot * 19:34 wmbot~andrew@bullseye: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) * 19:34 wmbot~andrew@bullseye: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot * 19:32 wmbot~andrew@bullseye: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot * 17:13 wmbot~fran@wmf3169: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirtlocal1001.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T356975|T356975]]) * 17:12 wmbot~fran@wmf3169: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirtlocal1001.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T356975|T356975]]) * 16:32 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1043.eqiad.wmnet<nowiki>}</nowiki>' * 16:27 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1043.eqiad.wmnet<nowiki>}</nowiki>' * 16:07 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1043.eqiad.wmnet<nowiki>}</nowiki>' * 15:51 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1043.eqiad.wmnet<nowiki>}</nowiki>' * 15:50 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1042.eqiad.wmnet<nowiki>}</nowiki>' * 15:25 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1042.eqiad.wmnet<nowiki>}</nowiki>' * 14:48 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1041.eqiad.wmnet<nowiki>}</nowiki>' * 14:48 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1040.eqiad.wmnet<nowiki>}</nowiki>' * 14:28 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1040.eqiad.wmnet<nowiki>}</nowiki>' * 14:28 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1039.eqiad.wmnet<nowiki>}</nowiki>' * 14:09 taavi: creating some missing $PROJECT.wmcloud.org. DNS zones * 14:07 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1039.eqiad.wmnet<nowiki>}</nowiki>' * 14:06 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1038.eqiad.wmnet<nowiki>}</nowiki>' * 13:50 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 13:50 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 13:48 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1038.eqiad.wmnet<nowiki>}</nowiki>' * 13:48 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1037.eqiad.wmnet<nowiki>}</nowiki>' * 13:23 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1037.eqiad.wmnet<nowiki>}</nowiki>' * 13:22 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1036.eqiad.wmnet<nowiki>}</nowiki>' * 13:00 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1036.eqiad.wmnet<nowiki>}</nowiki>' * 13:00 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1035.eqiad.wmnet<nowiki>}</nowiki>' * 12:35 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1035.eqiad.wmnet<nowiki>}</nowiki>' * 12:34 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1034.eqiad.wmnet<nowiki>}</nowiki>' * 12:13 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1034.eqiad.wmnet<nowiki>}</nowiki>' * 12:08 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1033.eqiad.wmnet<nowiki>}</nowiki>' * 11:39 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1033.eqiad.wmnet<nowiki>}</nowiki>' * 11:38 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1032.eqiad.wmnet<nowiki>}</nowiki>' * 11:33 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1032.eqiad.wmnet<nowiki>}</nowiki>' * 11:21 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::virt_ceph<nowiki>}</nowiki>' * 10:55 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::virt_ceph<nowiki>}</nowiki>' * 10:43 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::virt_ceph<nowiki>}</nowiki>' * 10:15 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::virt_ceph<nowiki>}</nowiki>' * 09:33 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::codfw1dev::virt_ceph<nowiki>}</nowiki>' * 09:31 taavi: failover all dumps traffic to clouddumps1001 * 08:33 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::codfw1dev::virt_ceph<nowiki>}</nowiki>' * 08:16 taavi: reboot clouddumps1001 for kernel updates === 2024-02-13 === * 17:07 wmbot~taavi@runko: END (ERROR) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=97) on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::virt_ceph<nowiki>}</nowiki>' * 17:07 wmbot~taavi@runko: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::virt_ceph<nowiki>}</nowiki>' * 16:06 wmbot~taavi@runko: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::virt_ceph<nowiki>}</nowiki>' * 16:06 wmbot~taavi@runko: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::virt_ceph<nowiki>}</nowiki>' * 16:04 wmbot~taavi@runko: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::virt_ceph<nowiki>}</nowiki>' * 16:04 wmbot~taavi@runko: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::virt_ceph<nowiki>}</nowiki>' * 16:04 wmbot~taavi@runko: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::virt_ceph<nowiki>}</nowiki>' * 16:04 wmbot~taavi@runko: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::virt_ceph<nowiki>}</nowiki>' * 16:02 wmbot~taavi@runko: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'P<nowiki>{</nowiki>P:openstack::eqiad1::nova::compute::service<nowiki>}</nowiki>' * 16:02 wmbot~taavi@runko: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>P:openstack::eqiad1::nova::compute::service<nowiki>}</nowiki>' * 16:02 wmbot~taavi@runko: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'P:openstack::eqiad1::nova::compute::service' * 16:02 wmbot~taavi@runko: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'P:openstack::eqiad1::nova::compute::service' * 16:01 wmbot~taavi@runko: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1001.eqiad.wmnet<nowiki>}</nowiki>' * 16:01 wmbot~taavi@runko: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1001.eqiad.wmnet<nowiki>}</nowiki>' * 15:10 wmbot~fran@wmf3169: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) ([[phab:T356975|T356975]]) * 14:55 wmbot~fran@wmf3169: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot ([[phab:T356975|T356975]]) === 2024-02-12 === * 17:33 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudnet.reboot_node (exit_code=0) * 17:30 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudnet.reboot_node * 17:28 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudnet.reboot_node (exit_code=0) * 17:25 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudnet.reboot_node * 17:23 wmbot~taavi@runko: END (FAIL) - Cookbook wmcs.openstack.cloudnet.reboot_node (exit_code=99) * 17:21 wmbot~taavi@runko: START - Cookbook wmcs.openstack.cloudnet.reboot_node * 17:10 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) * 17:06 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot * 17:04 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) * 16:52 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot * 16:39 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) * 16:39 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot * 16:36 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) * 16:30 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot * 16:29 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) * 16:21 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot * 16:20 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) * 16:14 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot * 16:04 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) * 15:59 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot * 15:50 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) * 15:46 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot * 15:43 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudnet.reboot_node (exit_code=99) * 15:42 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudnet.reboot_node * 01:37 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 01:34 andrewbogott: resetting eqiad1 rabbitmq in hopes of resolving neutron double message warnings * 01:32 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 00:57 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 00:55 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2024-02-11 === * 21:07 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 21:02 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 21:00 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 21:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 20:57 andrewbogott: running wmcs.openstack.restart_openstack for all eqiad1 services * 20:57 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 20:57 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 11:24 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 11:24 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 11:23 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 11:23 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2024-02-08 === * 19:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 19:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 13:43 taavi: deploy change to exclude cloud-private networks from general egress NAT https://phabricator.wikimedia.org/T356850 === 2024-02-06 === * 20:17 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 20:15 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 20:03 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 20:01 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 19:38 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 19:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 19:35 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 19:32 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 19:31 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 19:31 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 02:11 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 02:09 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 02:09 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.restart_openstack (exit_code=97) * 02:09 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2024-02-05 === * 21:29 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 21:29 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 20:45 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 20:45 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 20:03 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 20:03 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 20:02 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 20:01 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 18:55 andrewbogott: rebuilt bookworm base image in eqiad1 with https://gerrit.wikimedia.org/r/c/operations/puppet/+/992677 === 2024-02-02 === * 13:54 arturo: [codfw1dev] cleanup /etc/network/interfaces on cloudlb2003-dev from puppet leftovers === 2024-02-01 === * 10:55 taavi: invite aborrero to /repos/cloud, /toolforge-repos, /cloudvps-repos on gitlab === 2024-01-29 === * 12:54 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 12:54 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2024-01-26 === * 09:01 taavi: joining cloudrabbit1001/2 to the cluster on 1003 [[phab:T345610|T345610]] === 2024-01-25 === * 16:48 andrewbogott: taavi just moved all rabbitmq traffic to cloudrabbit1003 as part of [[phab:T345610|T345610]] * 16:45 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 16:40 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 13:27 wmbot~taavi@runko: END (PASS) - Cookbook wmcs.toolforge.add_k8s_node (exit_code=0) for a worker-nfs role in the tools cluster * 13:27 wmbot~taavi@runko: Added a new k8s worker-nfs tools-k8s-worker-nfs-2.tools.eqiad1.wikimedia.cloud to the cluster * 13:15 wmbot~taavi@runko: END (PASS) - Cookbook wmcs.toolforge.add_k8s_node (exit_code=0) for a worker-nfs role in the tools cluster * 13:15 wmbot~taavi@runko: Added a new k8s worker-nfs tools-k8s-worker-nfs-1.tools.eqiad1.wikimedia.cloud to the cluster * 12:48 wmbot~taavi@runko: END (PASS) - Cookbook wmcs.toolforge.add_k8s_node (exit_code=0) for a worker-nfs role in the toolsbeta cluster * 12:48 wmbot~taavi@runko: Added a new k8s worker-nfs toolsbeta-test-k8s-worker-nfs-1.toolsbeta.eqiad1.wikimedia.cloud to the cluster === 2024-01-24 === * 11:37 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.toolforge.add_k8s_node (exit_code=0) for a worker role in the toolsbeta cluster * 11:37 taavi@cloudcumin1001: Added a new k8s worker toolsbeta-test-k8s-worker-10.toolsbeta.eqiad1.wikimedia.cloud to the cluster === 2024-01-22 === * 16:03 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 15:56 andrewbogott: restarting openstack sevices on eqiad1 to clean up from the mariadb restarts * 15:56 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 15:07 andrewbogott: merging https://gerrit.wikimedia.org/r/c/operations/puppet/+/992192 and resetting galera cluster in eqiad1 === 2024-01-21 === * 05:19 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 05:12 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2024-01-19 === * 23:36 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 23:33 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2024-01-18 === * 20:10 taavi: mysql:labsdbaccounts@m5-master.eqiad.wmnet [labsdbaccounts]> update account_host set status = 'absent' where id = 137613; * 12:38 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.toolforge.add_k8s_node (exit_code=0) for a worker role in the tools cluster * 12:38 taavi@cloudcumin1001: Added a new k8s worker tools-k8s-worker-101.tools.eqiad1.wikimedia.cloud to the cluster === 2024-01-17 === * 15:17 andrewbogott: "systemctl restart mariadb@s4.service mariadb@s6.service" on clouddb1015. System is in danger of oom and there are no obvious long queries running === 2024-01-16 === * 14:25 wmbot~taavi@runko: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 14:25 wmbot~taavi@runko: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 14:24 wmbot~taavi@runko: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) for cloudvirt1060.eqiad.wmnet * 14:23 wmbot~taavi@runko: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot for cloudvirt1060.eqiad.wmnet * 14:23 wmbot~taavi@runko: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) for cloudvirt1060.eqiad.wmnet * 14:23 wmbot~taavi@runko: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot for cloudvirt1060.eqiad.wmnet * 13:55 wmbot~taavi@runko: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 13:55 wmbot~taavi@runko: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 13:54 wmbot~taavi@runko: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 13:54 wmbot~taavi@runko: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 13:53 wmbot~taavi@runko: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=99) * 13:53 wmbot~taavi@runko: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 13:52 wmbot~taavi@runko: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 13:52 wmbot~taavi@runko: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 11:35 wmbot~taavi@runko: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) for cloudvirt1060.eqiad.wmnet * 11:34 wmbot~taavi@runko: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot for cloudvirt1060.eqiad.wmnet * 11:22 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) ([[phab:T355061|T355061]]) * 11:02 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot ([[phab:T355061|T355061]]) * 09:41 taavi: drop dbproxy1018/9 grants from all clouddb hosts [[phab:T346947|T346947]] * 09:24 taavi: move cloudvirt2004-dev from 'failed' to 'active' in netbox - seems like that was for [[phab:T348531|T348531]] which is now resolved === 2024-01-15 === * 14:37 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) ([[phab:T355061|T355061]]) * 14:32 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot ([[phab:T355061|T355061]]) * 14:32 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) ([[phab:T355061|T355061]]) * 14:28 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot ([[phab:T355061|T355061]]) * 12:41 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 12:37 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2024-01-11 === * 10:51 dcaro: restarting striker.service on cloudweb1003 as it seems non-responsive === 2024-01-10 === * 17:43 bd808: Blocking Developer accounts connected to invalid/legacy wikimedia.org email addresses ([[phab:T218239|T218239]]) * 16:48 wmbot~fran@wmf3169: END (PASS) - Cookbook wmcs.do_log_msg (exit_code=0) ([[phab:T346631|T346631]]) * 16:48 wmbot~fran@wmf3169: test message3 from local cookbook ([[phab:T346631|T346631]]) * 16:48 wmbot~fran@wmf3169: START - Cookbook wmcs.do_log_msg ([[phab:T346631|T346631]]) === 2024-01-09 === * 17:50 wmbot~fran@wmf3169: END (PASS) - Cookbook wmcs.do_log_msg (exit_code=0) ([[phab:T346631|T346631]]) * 17:49 wmbot~fran@wmf3169: test message2 from local cookbook ([[phab:T346631|T346631]]) * 17:49 wmbot~fran@wmf3169: START - Cookbook wmcs.do_log_msg ([[phab:T346631|T346631]]) * 17:43 wmbot~fran@wmf3169: %(message)s ([[phab:T346631|T346631]]) * 17:43 wmbot~fran@wmf3169: %(message)s ([[phab:T346631|T346631]]) * 17:42 wmbot~fran@wmf3169: %(message)s ([[phab:T346631|T346631]]) === 2024-01-08 === * 15:52 taavi: verify wmcloud.org, wmflabs.org and toolforge.org in gmail postmaster console to figure out how much google likes us ([[phab:T354112|T354112]]) === 2024-01-07 === * 19:34 andrewbogott: removed cloudvirt1063 from 'ceph' aggregate, added to 'maintenance' aggregate [[phab:T353408|T353408]] * 19:34 andrewbogott: evacuating all VMs from cloudvirt1063. [[phab:T353408|T353408]] === 2024-01-02 === * 16:18 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=99) ([[phab:T353408|T353408]]) * 16:18 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance ([[phab:T353408|T353408]]) * 10:22 wm-bot2: fran@wmf3169 END (FAIL) - Cookbook wmcs.openstack.cloudvirt.vm_console (exit_code=99) * 10:22 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudvirt.vm_console === 2023-12-31 === * 21:39 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 21:35 andrewbogott: running openstack service restart cookbook in eqiad1 in response to a bunch of service down alerts * 21:34 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2023-12-21 === * 16:52 dhinus: puppet node deactivate cloudvirt1063.eqiad.wmnet [[phab:T353406|T353406]] * 03:01 andrewbogott: restarting mariadb on cloudcontrol1005, hoping to get Galera back in sync === 2023-12-20 === * 19:13 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 19:08 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2023-12-18 === * 17:29 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 17:23 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 17:23 andrewbogott: restarting all eqiad1 openstack services after a rabbitmq upgrade/rebuild for [[phab:T353646|T353646]] * 15:29 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 15:25 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2023-12-15 === * 13:13 dcaro: restarted nova-fullstack on codfw as it was stuck (and alerting through stale prometheus file) === 2023-12-14 === * 00:27 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=99) * 00:26 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 00:25 andrewbogott: evacuating hosts from cloudvirt1063 and depooling. [[phab:T353406|T353406]] === 2023-12-13 === * 16:39 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.toolforge.scale_grid_exec (exit_code=99) === 2023-12-12 === * 21:01 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 20:55 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 17:45 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.toolforge.add_k8s_node (exit_code=0) for a worker role in the tools cluster * 17:45 taavi@cloudcumin1001: Added a new k8s worker tools-k8s-worker-100.tools.eqiad1.wikimedia.cloud to the cluster * 16:11 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.toolforge.add_k8s_node (exit_code=0) for a worker role in the tools cluster * 16:11 taavi@cloudcumin1001: Added a new k8s worker tools-k8s-worker-99.tools.eqiad1.wikimedia.cloud to the cluster * 15:49 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.toolforge.add_k8s_node (exit_code=0) for a worker role in the tools cluster * 15:49 taavi@cloudcumin1001: Added a new k8s worker tools-k8s-worker-98.tools.eqiad1.wikimedia.cloud to the cluster === 2023-12-10 === * 18:37 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 18:33 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 18:30 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 18:27 andrewbogott: restarting all openstack API servers, hoping to make things a bit more responsive * 18:25 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2023-12-08 === * 12:00 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.vps.refresh_puppet_certs (exit_code=99) on etcd-discovery-1.cloudinfra-codfw1dev.codfw1dev.wikimedia.cloud ([[phab:T353055|T353055]]) * 11:58 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.vps.refresh_puppet_certs on etcd-discovery-1.cloudinfra-codfw1dev.codfw1dev.wikimedia.cloud ([[phab:T353055|T353055]]) * 11:58 wm-bot2: dcaro@urcuchillay END (ERROR) - Cookbook wmcs.vps.refresh_puppet_certs (exit_code=97) on etcd-discovery-1.cloudinfra-codfw1dev.codfw1dev.wikimedia.cloud * 11:57 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.vps.refresh_puppet_certs on etcd-discovery-1.cloudinfra-codfw1dev.codfw1dev.wikimedia.cloud * 09:38 wm-bot2: dcaro@urcuchillay END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) ([[phab:T345084|T345084]]) * 09:32 dcaro: restarting nova and keystone as they are getting too slow ([[phab:T345084|T345084]]) * 09:32 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.openstack.restart_openstack ([[phab:T345084|T345084]]) * 09:32 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) ([[phab:T345084|T345084]]) * 09:31 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.openstack.restart_openstack ([[phab:T345084|T345084]]) === 2023-12-07 === * 13:12 dcaro: rebooting cloudcephosd1001 to make sure puppet7 migration went ok === 2023-12-04 === * 00:08 andrewbogott: rebooting cloudcontrol1006 to recover from full disk error === 2023-12-03 === * 09:05 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 09:05 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2023-12-02 === * 12:27 taavi: powercycle cloudvirt1063 [[phab:T352595|T352595]] * 11:28 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.toolforge.add_k8s_node (exit_code=0) for a worker role in the tools cluster * 11:28 taavi@cloudcumin1001: Added a new k8s worker tools-k8s-worker-97.tools.eqiad1.wikimedia.cloud to the cluster * 11:02 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.toolforge.add_k8s_node (exit_code=0) for a worker role in the tools cluster * 11:02 taavi@cloudcumin1001: Added a new k8s worker tools-k8s-worker-96.tools.eqiad1.wikimedia.cloud to the cluster * 10:50 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.toolforge.add_k8s_node (exit_code=0) for a worker role in the tools cluster * 10:50 taavi@cloudcumin1001: Added a new k8s worker tools-k8s-worker-95.tools.eqiad1.wikimedia.cloud to the cluster * 00:21 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.toolforge.add_k8s_node (exit_code=0) for a worker role in the tools cluster * 00:21 taavi@cloudcumin1001: Added a new k8s worker tools-k8s-worker-94.tools.eqiad1.wikimedia.cloud to the cluster * 00:18 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.toolforge.add_k8s_node (exit_code=0) for a worker role in the tools cluster * 00:18 taavi@cloudcumin1001: Added a new k8s worker tools-k8s-worker-93.tools.eqiad1.wikimedia.cloud to the cluster * 00:15 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.toolforge.add_k8s_node (exit_code=0) for a worker role in the tools cluster * 00:15 taavi@cloudcumin1001: Added a new k8s worker tools-k8s-worker-92.tools.eqiad1.wikimedia.cloud to the cluster * 00:06 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.toolforge.add_k8s_node (exit_code=0) for a worker role in the tools cluster * 00:06 taavi@cloudcumin1001: Added a new k8s worker tools-k8s-worker-91.tools.eqiad1.wikimedia.cloud to the cluster * 00:05 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.toolforge.add_k8s_node (exit_code=0) for a worker role in the tools cluster * 00:05 taavi@cloudcumin1001: Added a new k8s worker tools-k8s-worker-90.tools.eqiad1.wikimedia.cloud to the cluster * 00:01 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.toolforge.add_k8s_node (exit_code=0) for a worker role in the tools cluster * 00:01 taavi@cloudcumin1001: Added a new k8s worker tools-k8s-worker-89.tools.eqiad1.wikimedia.cloud to the cluster === 2023-12-01 === * 17:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 17:37 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 17:01 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 16:57 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 16:48 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 16:46 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 16:19 wm-bot2: fran@wmf3169 END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) ([[phab:T351171|T351171]]) * 16:19 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance ([[phab:T351171|T351171]]) * 15:49 andrewbogott: reimaging cloudcontrol1005 due to widespread misbehavior * 14:24 wm-bot2: fran@wmf3169 END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) ([[phab:T351171|T351171]]) * 14:20 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance ([[phab:T351171|T351171]]) * 14:19 fnegri@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=97) ([[phab:T351171|T351171]]) * 14:18 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance ([[phab:T351171|T351171]]) * 13:55 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 13:54 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 11:57 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 11:56 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 11:55 taavi: restart neutron-rpc-server.service on eqiad1 cloudcontrols === 2023-11-30 === * 20:44 andrewbogott: generating to application credentials for the tests that run on tf-infra-test * 19:54 andrewbogott: reimaged cloudrabbit100[23] after https://gerrit.wikimedia.org/r/c/operations/puppet/+/979127. I didn't reimage 1001 because that will require rebuilding the whole cluster but I did remove the related packages. * 18:10 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1039.eqiad.wmnet' * 18:10 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1038.eqiad.wmnet' * 18:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1039.eqiad.wmnet' * 18:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1038.eqiad.wmnet' * 18:04 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1037.eqiad.wmnet' * 18:04 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1036.eqiad.wmnet' * 18:04 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1035.eqiad.wmnet' * 17:59 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1037.eqiad.wmnet' * 17:59 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1036.eqiad.wmnet' * 17:59 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1035.eqiad.wmnet' * 17:53 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1033.eqiad.wmnet' * 17:53 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1034.eqiad.wmnet' * 17:53 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1032.eqiad.wmnet' * 17:52 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt-wdqs1003.eqiad.wmnet' ([[phab:T348843|T348843]]) * 17:48 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1034.eqiad.wmnet' * 17:48 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1033.eqiad.wmnet' * 17:48 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1032.eqiad.wmnet' * 17:47 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1031.eqiad.wmnet' * 17:47 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt-wdqs1003.eqiad.wmnet' ([[phab:T348843|T348843]]) * 17:47 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt-wdqs1002.eqiad.wmnet' ([[phab:T348843|T348843]]) * 17:42 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1031.eqiad.wmnet' * 17:42 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt-wdqs1002.eqiad.wmnet' ([[phab:T348843|T348843]]) * 17:42 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt-wdqs1001.eqiad.wmnet' ([[phab:T348843|T348843]]) * 17:37 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt-wdqs1001.eqiad.wmnet' ([[phab:T348843|T348843]]) * 17:36 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirtlocal1003.eqiad.wmnet' ([[phab:T348843|T348843]]) * 17:30 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirtlocal1003.eqiad.wmnet' ([[phab:T348843|T348843]]) * 17:30 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirtlocal1002.eqiad.wmnet' ([[phab:T348843|T348843]]) * 17:25 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirtlocal1002.eqiad.wmnet' ([[phab:T348843|T348843]]) * 17:25 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirtlocal1001.eqiad.wmnet' ([[phab:T348843|T348843]]) * 17:19 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirtlocal1001.eqiad.wmnet' ([[phab:T348843|T348843]]) * 17:19 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1045.eqiad.wmnet' ([[phab:T348843|T348843]]) * 17:14 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1045.eqiad.wmnet' ([[phab:T348843|T348843]]) * 17:14 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1044.eqiad.wmnet' ([[phab:T348843|T348843]]) * 17:09 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1044.eqiad.wmnet' ([[phab:T348843|T348843]]) * 17:09 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1043.eqiad.wmnet' ([[phab:T348843|T348843]]) * 17:04 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1043.eqiad.wmnet' ([[phab:T348843|T348843]]) * 17:04 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1042.eqiad.wmnet' ([[phab:T348843|T348843]]) * 16:59 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1042.eqiad.wmnet' ([[phab:T348843|T348843]]) * 16:59 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1041.eqiad.wmnet' ([[phab:T348843|T348843]]) * 16:54 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1041.eqiad.wmnet' ([[phab:T348843|T348843]]) * 16:54 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1040.eqiad.wmnet' ([[phab:T348843|T348843]]) * 16:49 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1040.eqiad.wmnet' ([[phab:T348843|T348843]]) * 16:45 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1049.eqiad.wmnet' ([[phab:T348843|T348843]]) * 16:39 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1049.eqiad.wmnet' ([[phab:T348843|T348843]]) * 16:39 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1048.eqiad.wmnet' ([[phab:T348843|T348843]]) * 16:34 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1048.eqiad.wmnet' ([[phab:T348843|T348843]]) * 16:34 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1047.eqiad.wmnet' ([[phab:T348843|T348843]]) * 16:29 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1047.eqiad.wmnet' ([[phab:T348843|T348843]]) * 16:21 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1059.eqiad.wmnet' ([[phab:T348843|T348843]]) * 16:15 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1059.eqiad.wmnet' ([[phab:T348843|T348843]]) * 16:15 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1058.eqiad.wmnet' ([[phab:T348843|T348843]]) * 16:10 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1058.eqiad.wmnet' ([[phab:T348843|T348843]]) * 16:10 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1057.eqiad.wmnet' ([[phab:T348843|T348843]]) * 16:05 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1057.eqiad.wmnet' ([[phab:T348843|T348843]]) * 16:05 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1056.eqiad.wmnet' ([[phab:T348843|T348843]]) * 16:00 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1056.eqiad.wmnet' ([[phab:T348843|T348843]]) * 15:56 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1055.eqiad.wmnet' ([[phab:T348843|T348843]]) * 15:55 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1054.eqiad.wmnet' ([[phab:T348843|T348843]]) * 15:51 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1054.eqiad.wmnet' ([[phab:T348843|T348843]]) * 15:51 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1053.eqiad.wmnet' ([[phab:T348843|T348843]]) * 15:46 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1053.eqiad.wmnet' ([[phab:T348843|T348843]]) * 15:46 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1052.eqiad.wmnet' ([[phab:T348843|T348843]]) * 15:41 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1052.eqiad.wmnet' ([[phab:T348843|T348843]]) * 15:41 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1051.eqiad.wmnet' ([[phab:T348843|T348843]]) * 15:37 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1051.eqiad.wmnet' ([[phab:T348843|T348843]]) * 15:36 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1050.eqiad.wmnet' ([[phab:T348843|T348843]]) * 15:31 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1050.eqiad.wmnet' ([[phab:T348843|T348843]]) * 15:26 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1065.eqiad.wmnet' ([[phab:T348843|T348843]]) * 15:22 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1065.eqiad.wmnet' ([[phab:T348843|T348843]]) * 15:22 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1064.eqiad.wmnet' ([[phab:T348843|T348843]]) * 15:17 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1064.eqiad.wmnet' ([[phab:T348843|T348843]]) * 15:17 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1063.eqiad.wmnet' ([[phab:T348843|T348843]]) * 15:13 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1063.eqiad.wmnet' ([[phab:T348843|T348843]]) * 15:12 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1062.eqiad.wmnet' ([[phab:T348843|T348843]]) * 15:07 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1062.eqiad.wmnet' ([[phab:T348843|T348843]]) * 15:07 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1061.eqiad.wmnet' ([[phab:T348843|T348843]]) * 15:02 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1061.eqiad.wmnet' ([[phab:T348843|T348843]]) * 15:02 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1060.eqiad.wmnet' ([[phab:T348843|T348843]]) * 14:56 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1060.eqiad.wmnet' ([[phab:T348843|T348843]]) * 14:49 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1066.eqiad.wmnet' ([[phab:T348843|T348843]]) * 14:45 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1066.eqiad.wmnet' ([[phab:T348843|T348843]]) * 14:43 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1067.eqiad.wmnet' ([[phab:T348843|T348843]]) * 14:38 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1067.eqiad.wmnet' ([[phab:T348843|T348843]]) * 14:16 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet1006.eqiad.wmnet' ([[phab:T348843|T348843]]) * 14:03 wm-bot2: fran@wmf3169 END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudnet1005.eqiad.wmnet' ([[phab:T348843|T348843]]) * 13:55 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet1005.eqiad.wmnet' ([[phab:T348843|T348843]]) * 13:43 wm-bot2: fran@wmf3169 END (PASS) - Cookbook wmcs.openstack.cloudweb.set_maintenance (exit_code=0) ([[phab:T348843|T348843]]) * 13:43 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudweb.set_maintenance ([[phab:T348843|T348843]]) * 12:18 wm-bot2: fran@wmf3169 END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol1005.eqiad.wmnet' ([[phab:T348843|T348843]]) * 12:04 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1005.eqiad.wmnet' ([[phab:T348843|T348843]]) * 11:59 wm-bot2: fran@wmf3169 END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol1006.eqiad.wmnet' ([[phab:T348843|T348843]]) * 11:45 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1006.eqiad.wmnet' ([[phab:T348843|T348843]]) * 11:44 wm-bot2: fran@wmf3169 END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol1007.eqiad.wmnet' ([[phab:T348843|T348843]]) * 11:28 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1007.eqiad.wmnet' ([[phab:T348843|T348843]]) === 2023-11-29 === * 15:27 wm-bot2: fran@wmf3169 END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudservices1006.eqiad.wmnet' ([[phab:T348843|T348843]]) * 15:15 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1006.eqiad.wmnet' ([[phab:T348843|T348843]]) * 14:59 wm-bot2: fran@wmf3169 END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) ([[phab:T348843|T348843]]) * 14:50 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node ([[phab:T348843|T348843]]) === 2023-11-28 === * 14:18 taavi: moving wiki replica DNS to use cloudlbs instead of the old proxy VMs [[phab:T346947|T346947]] === 2023-11-27 === * 19:42 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 19:35 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2023-11-24 === * 14:51 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 14:50 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 12:01 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 12:01 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 12:01 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 12:00 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 12:00 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 12:00 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 11:53 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 11:53 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 11:53 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 11:53 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 11:53 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 11:53 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2023-11-22 === * 13:28 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 13:21 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 13:02 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 13:00 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2023-11-21 === * 10:11 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 10:10 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2023-11-20 === * 09:35 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 09:35 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2023-11-17 === * 16:36 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 16:32 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2023-11-16 === * 12:09 taavi@cloudcumin2001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 12:05 taavi@cloudcumin2001: START - Cookbook wmcs.openstack.restart_openstack * 11:23 dhinus: upgraded spicerack from 8.0.2 to 8.0.3 on cloudcumins * 05:51 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 05:49 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2023-11-15 === * 09:50 taavi: move cloudlb hosts to use the nftables firewall backend === 2023-11-14 === * 21:09 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1056.eqiad.wmnet' ([[phab:T345811|T345811]]) * 20:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1055.eqiad.wmnet' ([[phab:T345811|T345811]]) * 20:26 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1056.eqiad.wmnet' ([[phab:T345811|T345811]]) * 20:15 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1054.eqiad.wmnet' ([[phab:T345811|T345811]]) * 20:12 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1055.eqiad.wmnet' ([[phab:T345811|T345811]]) * 20:10 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1053.eqiad.wmnet' ([[phab:T345811|T345811]]) * 20:09 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1053.eqiad.wmnet' ([[phab:T345811|T345811]]) * 20:06 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1053.eqiad.wmnet' ([[phab:T345811|T345811]]) * 20:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1053.eqiad.wmnet' ([[phab:T345811|T345811]]) * 20:00 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1053.eqiad.wmnet' ([[phab:T345811|T345811]]) * 19:59 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1053.eqiad.wmnet' ([[phab:T345811|T345811]]) * 19:58 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1054.eqiad.wmnet' ([[phab:T345811|T345811]]) * 19:55 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1053.eqiad.wmnet' ([[phab:T345811|T345811]]) * 19:54 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1053.eqiad.wmnet' ([[phab:T345811|T345811]]) * 19:50 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1053.eqiad.wmnet' ([[phab:T345811|T345811]]) * 19:49 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1053.eqiad.wmnet' ([[phab:T345811|T345811]]) * 19:47 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1053.eqiad.wmnet' ([[phab:T345811|T345811]]) * 19:21 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1053.eqiad.wmnet' ([[phab:T345811|T345811]]) * 19:18 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1052.eqiad.wmnet' ([[phab:T345811|T345811]]) * 19:17 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1050.eqiad.wmnet' ([[phab:T345811|T345811]]) * 19:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1052.eqiad.wmnet' ([[phab:T345811|T345811]]) * 18:57 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1050.eqiad.wmnet' ([[phab:T345811|T345811]]) * 18:31 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=97) on host 'cloudvirt1049.eqiad.wmnet' ([[phab:T345811|T345811]]) * 18:10 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1048.eqiad.wmnet' ([[phab:T345811|T345811]]) * 18:06 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1049.eqiad.wmnet' ([[phab:T345811|T345811]]) * 18:03 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1047.eqiad.wmnet' ([[phab:T345811|T345811]]) * 18:02 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1047.eqiad.wmnet' ([[phab:T345811|T345811]]) * 18:01 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1047.eqiad.wmnet' ([[phab:T345811|T345811]]) * 17:59 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 17:58 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 17:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1048.eqiad.wmnet' ([[phab:T345811|T345811]]) * 17:39 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1047.eqiad.wmnet' ([[phab:T345811|T345811]]) * 17:24 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1047.eqiad.wmnet' ([[phab:T345811|T345811]]) * 17:24 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1047.eqiad.wmnet' ([[phab:T345811|T345811]]) * 12:03 wm-bot2: fran@wmf3169 admin END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1046.eqiad.wmnet' ([[phab:T345811|T345811]]) * 11:44 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1046.eqiad.wmnet' ([[phab:T345811|T345811]]) * 10:10 taavi: restart kiwix-mirror-update on clouddumps1001 * 05:03 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) ([[phab:T345811|T345811]]) * 05:02 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain ([[phab:T345811|T345811]]) * 04:34 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) * 04:15 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) ([[phab:T345811|T345811]]) * 04:14 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain ([[phab:T345811|T345811]]) * 04:11 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) ([[phab:T345811|T345811]]) * 03:50 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain ([[phab:T345811|T345811]]) * 03:49 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) ([[phab:T345811|T345811]]) * 03:48 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain * 03:48 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain ([[phab:T345811|T345811]]) * 03:34 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) * 03:26 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) ([[phab:T345811|T345811]]) * 03:08 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain ([[phab:T345811|T345811]]) * 02:59 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) ([[phab:T345811|T345811]]) * 02:59 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) * 02:59 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain * 02:59 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain ([[phab:T345811|T345811]]) * 02:55 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) * 02:52 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) ([[phab:T345811|T345811]]) * 02:46 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain * 02:44 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) * 02:34 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain ([[phab:T345811|T345811]]) * 02:32 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain * 02:27 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) * 02:26 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain * 02:23 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) ([[phab:T345811|T345811]]) * 02:17 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) * 02:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain * 02:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain ([[phab:T345811|T345811]]) === 2023-11-13 === * 22:25 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) ([[phab:T345811|T345811]]) * 22:24 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain ([[phab:T345811|T345811]]) * 22:11 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) * 22:11 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain * 22:09 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) * 22:09 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) ([[phab:T345811|T345811]]) * 22:08 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain ([[phab:T345811|T345811]]) * 22:08 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain * 21:54 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) * 21:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) ([[phab:T345811|T345811]]) * 21:30 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain * 21:30 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain ([[phab:T345811|T345811]]) * 21:21 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) ([[phab:T345811|T345811]]) * 21:17 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) * 21:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain ([[phab:T345811|T345811]]) * 21:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain * 19:31 andrewbogott: rebooting cloudcontrol2005-dev, trying to fix general misbehavior * 19:18 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) * 19:17 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) ([[phab:T345811|T345811]]) * 19:16 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain * 19:16 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) * 19:01 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain ([[phab:T345811|T345811]]) * 18:53 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) * 18:47 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain * 18:33 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain * 18:31 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) * 18:31 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain * 18:30 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) * 18:29 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain * 18:27 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1033.eqiad.wmnet' * 18:27 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1032.eqiad.wmnet' * 18:27 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1033.eqiad.wmnet' * 18:27 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1032.eqiad.wmnet' * 17:09 wm-bot2: fran@wmf3169 END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) ([[phab:T345811|T345811]]) * 17:09 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance ([[phab:T345811|T345811]]) * 16:57 wm-bot2: fran@wmf3169 END (FAIL) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=99) ([[phab:T345811|T345811]]) * 16:56 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance ([[phab:T345811|T345811]]) * 15:37 wm-bot2: fran@wmf3169 admin END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1031.eqiad.wmnet' ([[phab:T345811|T345811]]) * 15:19 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1031.eqiad.wmnet' ([[phab:T345811|T345811]]) * 09:11 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 09:08 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 08:56 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 08:51 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2023-11-11 === * 02:42 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) * 02:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain * 02:41 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) * 02:14 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain * 02:13 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) * 01:58 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain * 01:23 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) * 01:22 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain * 01:21 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) * 01:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain * 01:03 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) * 00:48 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain * 00:47 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) * 00:46 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain * 00:46 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) * 00:46 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain * 00:46 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) * 00:46 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain * 00:44 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) * 00:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain * 00:42 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) * 00:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain * 00:40 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1058' * 00:39 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1058' * 00:38 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) * 00:37 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot === 2023-11-09 === * 21:40 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 21:33 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 14:54 wm-bot2: fran@wmf3169 admin END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1025.eqiad.wmnet' ([[phab:T345811|T345811]]) * 14:52 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1025.eqiad.wmnet' ([[phab:T345811|T345811]]) * 14:50 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) ([[phab:T345811|T345811]]) * 14:50 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain ([[phab:T345811|T345811]]) === 2023-11-08 === * 16:43 andrewbogott: created foundationmemory project for [[phab:T350760|T350760]] === 2023-11-06 === * 20:40 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 20:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 18:33 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 18:27 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 18:25 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 18:24 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 18:23 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 18:22 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 18:20 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 18:19 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 15:47 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) (348643) * 15:46 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node (348643) * 15:44 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) (348643) * 15:43 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node (348643) * 15:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) (348643) * 15:42 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node (348643) === 2023-11-05 === * 00:20 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) (348643) === 2023-11-04 === * 23:05 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node (348643) * 22:36 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) (348643) * 20:12 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node (348643) * 16:39 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) (348643) * 16:00 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node (348643) * 13:13 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node (348643) * 04:24 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) (348643) * 01:56 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node (348643) === 2023-11-03 === * 16:30 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) (348643) * 16:29 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node (348643) * 16:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) (348643) * 16:28 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node (348643) * 16:16 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) (348643) * 15:07 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node (348643) * 13:13 dhinus: triggering neutron failover from cloudnet1005 to cloudnet1006 ([[phab:T345811|T345811]]) * 06:35 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) (348643) * 01:46 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node (348643) * 01:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) (348643) * 01:46 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node (348643) * 01:45 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) (348643) * 01:45 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node (348643) === 2023-11-02 === * 20:37 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) (348643) * 18:14 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node (348643) * 17:32 taavi: merged cloudcontrol firewall cleanup patch https://gerrit.wikimedia.org/r/c/operations/puppet/+/971211 * 16:46 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.drain_node (exit_code=97) (348643) * 16:23 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) (348643) * 16:23 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node (348643) * 16:21 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node (348643) * 16:00 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) (348643) * 13:48 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node (348643) * 13:48 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) (348643) * 13:47 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node (348643) * 11:44 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) (348643) * 11:27 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node (348643) * 07:28 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 07:27 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 05:38 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) (348643) * 03:48 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node (348643) * 03:47 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.drain_node (exit_code=97) (348643) * 03:44 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node (348643) === 2023-11-01 === * 23:13 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) (348643) * 21:29 taavi: re-enable puppet on cloudcontrol2006-dev.codfw.wmnet which has fallen off of puppetdb * 21:15 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node (348643) * 21:15 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) (348643) * 21:15 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node (348643) * 21:14 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) (348643) * 20:26 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node (348643) * 20:11 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) (348643) * 19:55 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) (348643) * 15:10 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node (348643) * 14:55 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node (348643) * 14:52 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) (348643) * 14:52 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node (348643) * 14:50 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) (348643) * 14:49 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node (348643) * 14:49 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) (348643) * 14:48 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node (348643) * 14:47 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) (348643) * 14:47 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node (348643) * 09:04 taavi: reset local cookbook changes on cloudcumin1001 which were causing issues with puppet runs * 09:02 taavi: restart nova-fullstack which had had some issues after yesterday's cloudcontrol1007 reimage * 02:17 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) (348643) * 00:58 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node (348643) * 00:52 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) (348643) === 2023-10-31 === * 23:46 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node (348643) * 18:41 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) (348643) * 14:12 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) (348643) * 14:11 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node (348643) * 14:10 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node (348643) * 10:16 dhinus: upgrading mariadb-server in cloudcontrol1005 ([[phab:T345811|T345811]]) * 10:04 dhinus: upgrading mariadb-server in cloudcontrol1006 ([[phab:T345811|T345811]]) * 09:51 dhinus: upgrading mariadb-server in cloudcontrol1007, second attempt ([[phab:T345811|T345811]]) * 06:16 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) (348643) * 05:20 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) (348643) * 02:27 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) (348643) * 02:27 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node (348643) * 02:27 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.drain_node (exit_code=97) (348643) * 02:26 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node (348643) * 01:59 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node (348643) * 01:59 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node (348643) * 01:07 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) (348643) * 01:07 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) (348643) * 01:07 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) (348643) === 2023-10-30 === * 17:40 andrewbogott: rebooting tools-db-1.tools.eqiad1.wikimedia.cloud for yet another oom death * 17:09 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node (348643) * 17:04 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node (348643) * 17:02 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node (348643) * 16:56 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) (348643) * 16:55 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) (348643) * 16:55 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node (348643) * 16:55 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node (348643) * 16:54 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 16:53 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 16:53 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 16:52 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 16:52 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 16:51 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 16:39 dhinus: upgrading mariadb-server in cloudcontrol1007 ([[phab:T345811|T345811]]) * 16:38 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) ([[phab:T348643|T348643]]) * 16:38 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node ([[phab:T348643|T348643]]) * 16:37 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) ([[phab:T348643|T348643]]) * 16:37 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node ([[phab:T348643|T348643]]) * 16:37 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) ([[phab:T348643|T348643]]) * 16:37 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node ([[phab:T348643|T348643]]) * 16:30 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) ([[phab:T348643|T348643]]) * 16:30 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node ([[phab:T348643|T348643]]) * 16:30 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) ([[phab:T348643|T348643]]) * 16:30 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node ([[phab:T348643|T348643]]) * 16:29 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) ([[phab:T348643|T348643]]) * 16:29 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node ([[phab:T348643|T348643]]) * 16:17 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 16:17 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 16:17 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 16:17 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node === 2023-10-28 === * 07:06 wm-bot2: dcaro@urcuchillay END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) === 2023-10-27 === * 15:38 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.drain_node * 15:38 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) * 15:31 wm-bot2: dcaro@urcuchillay END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 15:29 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.undrain_node * 15:29 wm-bot2: dcaro@urcuchillay END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 15:28 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.undrain_node * 15:27 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 15:26 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.undrain_node * 13:21 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.drain_node * 13:09 wm-bot2: dcaro@urcuchillay END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) * 12:26 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.drain_node * 12:13 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) * 12:11 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.drain_node * 10:10 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 10:09 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 09:05 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) * 09:04 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.drain_node === 2023-10-26 === * 14:02 wm-bot2: dcaro@urcuchillay END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) * 09:46 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.drain_node === 2023-10-25 === * 11:09 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.toolforge.add_k8s_node (exit_code=99) for a ingress role in the toolsbeta cluster * 11:00 taavi: update cloudcumins to spicerack 8.x * 10:34 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.toolforge.add_k8s_node (exit_code=99) for a ingress role in the toolsbeta cluster === 2023-10-24 === * 15:31 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 15:30 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2023-10-23 === * 16:48 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 16:46 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 15:27 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 15:27 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 15:24 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 15:22 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 15:06 wm-bot2: fran@wmf3169 END (FAIL) - Cookbook wmcs.vps.create_project (exit_code=99) for project catalyst in eqiad1 * 15:06 wm-bot2: fran@wmf3169 START - Cookbook wmcs.vps.create_project for project catalyst in eqiad1 * 10:36 taavi: merged change https://gerrit.wikimedia.org/r/c/operations/puppet/+/966494 which touches the pdns web server config === 2023-10-20 === * 15:17 dcaro: upgraded cloudcephosd1004 to v15 ([[phab:T349363|T349363]]) * 14:25 dcaro: upgraded cloudcephosd1003 to v15 ([[phab:T349363|T349363]]) * 13:20 dcaro: upgraded cloudcephosd1002 to v15 ([[phab:T349363|T349363]]) * 10:33 dcaro: upgraded cloudcephosd1001 to v15 ([[phab:T349363|T349363]]) * 08:26 dcaro: codfw ceph enabled diskprediction_local module, will take a bit to populate/start getting predictions ([[phab:T348716|T348716]]) === 2023-10-17 === * 15:29 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) ([[phab:T349109|T349109]]) * 15:28 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain ([[phab:T349109|T349109]]) === 2023-10-16 === * 03:32 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 03:32 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2023-10-13 === * 23:12 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 23:10 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 19:26 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 19:24 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 17:05 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 17:03 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 16:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 16:42 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 16:36 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 16:33 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 15:18 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 15:15 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 08:41 wm-bot2: fran@wmf3169 END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) ([[phab:T341285|T341285]]) * 08:31 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node ([[phab:T341285|T341285]]) * 08:30 wm-bot2: fran@wmf3169 END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) ([[phab:T341285|T341285]]) * 08:20 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node ([[phab:T341285|T341285]]) === 2023-10-12 === * 17:16 wm-bot2: fran@wmf3169 END (ERROR) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=97) ([[phab:T341285|T341285]]) * 17:16 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node ([[phab:T341285|T341285]]) === 2023-10-11 === * 10:42 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) ([[phab:T341285|T341285]]) * 10:36 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack ([[phab:T341285|T341285]]) * 10:13 wm-bot2: dcaro@urcuchillay END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 07:03 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.undrain_node * 06:41 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.undrain_node === 2023-10-10 === * 17:25 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 14:50 dcaro: removing ~100 dangling backup snapshots from eqiad1-compute ceph pool * 12:46 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) ([[phab:T341285|T341285]]) * 12:41 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack ([[phab:T341285|T341285]]) * 12:23 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.undrain_node * 11:57 wm-bot2: dcaro@urcuchillay END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 11:38 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) ([[phab:T341285|T341285]]) * 11:33 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack ([[phab:T341285|T341285]]) * 11:33 wm-bot2: fran@wmf3169 admin END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) * 11:32 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudvirt.safe_reboot * 11:00 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=99) ([[phab:T341285|T341285]]) * 11:00 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack ([[phab:T341285|T341285]]) * 10:59 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) ([[phab:T341285|T341285]]) * 10:52 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack ([[phab:T341285|T341285]]) * 10:03 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) ([[phab:T341285|T341285]]) * 09:56 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack ([[phab:T341285|T341285]]) * 09:50 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) ([[phab:T341285|T341285]]) * 09:43 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack ([[phab:T341285|T341285]]) * 08:19 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.undrain_node * 08:18 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 08:18 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.undrain_node * 08:17 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 08:17 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.undrain_node * 08:17 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 08:17 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.undrain_node * 08:13 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 08:13 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.undrain_node === 2023-10-09 === * 17:16 wm-bot2: dcaro@urcuchillay END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 16:26 wm-bot2: fran@wmf3169 END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) ([[phab:T341285|T341285]]) * 16:18 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node ([[phab:T341285|T341285]]) * 15:49 wm-bot2: fran@wmf3169 END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) ([[phab:T341285|T341285]]) * 15:41 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node ([[phab:T341285|T341285]]) * 15:26 wm-bot2: fran@wmf3169 END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) ([[phab:T341285|T341285]]) * 13:55 wm-bot2: fran@wmf3169 END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) ([[phab:T341285|T341285]]) * 13:45 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node ([[phab:T341285|T341285]]) * 13:38 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.undrain_node * 13:13 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 13:12 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.undrain_node * 13:04 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 12:33 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.undrain_node * 09:10 dcaro: undrained cephosd1011 * 09:09 wm-bot2: dcaro@urcuchillay END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 09:09 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.undrain_node * 09:06 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 09:05 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.undrain_node * 09:05 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 09:05 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.undrain_node * 07:35 taavi: restart postgresql on cloudbackup2001 [[phab:T348431|T348431]] === 2023-10-05 === * 16:41 wm-bot2: fran@wmf3169 END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) ([[phab:T341285|T341285]]) * 16:29 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node ([[phab:T341285|T341285]]) * 16:24 wm-bot2: fran@wmf3169 END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) ([[phab:T341285|T341285]]) * 16:11 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node ([[phab:T341285|T341285]]) * 15:30 arturo: operating on cloudgw @ eqiad1 ([[phab:T347469|T347469]]) * 14:07 wm-bot2: dcaro@urcuchillay admin END (FAIL) - Cookbook wmcs.ceph.osd.drain_rack (exit_code=99) * 12:55 arturo: doing cloudgw maintenance operations [[phab:T347469|T347469]] * 11:54 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.drain_rack * 10:57 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) * 09:54 wm-bot2: fran@wmf3169 END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) ([[phab:T341285|T341285]]) * 09:52 arturo: [codfw1dev] aborrero@cloudcontrol2001-dev:~ $ sudo keystone-manage fernet_setup --keystone-user keystone --keystone-group keystone * 09:47 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node ([[phab:T341285|T341285]]) * 09:40 wm-bot2: fran@wmf3169 END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) ([[phab:T341285|T341285]]) * 09:33 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node ([[phab:T341285|T341285]]) * 08:50 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.drain_node * 07:49 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) * 07:32 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.drain_node * 00:04 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) === 2023-10-04 === * 20:18 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 20:18 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 16:52 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.drain_node * 15:55 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) * 14:54 wm-bot2: fran@wmf3169 END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) ([[phab:T341285|T341285]]) * 14:44 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node ([[phab:T341285|T341285]]) * 14:41 wm-bot2: fran@wmf3169 END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) ([[phab:T341285|T341285]]) * 14:40 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node ([[phab:T341285|T341285]]) * 13:50 wm-bot2: fran@wmf3169 END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) ([[phab:T341285|T341285]]) * 12:09 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.drain_node * 12:08 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) * 07:21 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.drain_node * 07:21 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) * 07:15 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.drain_node * 07:12 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) * 07:11 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.drain_node === 2023-10-03 === * 19:38 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) * 13:38 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.drain_node * 12:35 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) * 12:33 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.drain_node * 12:32 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) * 12:32 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.drain_node * 12:29 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) * 12:01 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.drain_node * 11:33 wm-bot2: dcaro@urcuchillay END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) * 08:58 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.drain_node * 08:43 wm-bot2: dcaro@urcuchillay END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) * 08:42 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.drain_node * 08:39 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) * 08:38 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.drain_node * 08:23 dcaro: set .rgw.root pool on eqiad as rgw app (`ceph osd pool application enable .rgw.root rgw`) === 2023-10-02 === * 17:39 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) * 16:15 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.drain_node * 16:13 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 16:12 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 14:30 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) * 13:58 taavi: cloudcontrol1005,7: `sudo systemctl reset-failed keystone_sync_keys_from_cloudcontrol1006.eqiad.wmnet.service` * 13:52 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.drain_node * 13:40 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) * 13:38 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.drain_node * 13:37 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) ([[phab:T316544|T316544]]) * 12:26 arturo: [codfw1dev] run `update domains set master = '185.15.57.25:5354 185.15.57.26:5354 172.20.5.9:5354 172.20.5.8:5354';` in cloudservies2005-dev * 11:55 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T316544|T316544]]) * 11:55 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) ([[phab:T316544|T316544]]) * 11:55 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T316544|T316544]]) * 11:43 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) ([[phab:T316544|T316544]]) * 08:44 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T316544|T316544]]) === 2023-09-29 === * 12:43 taavi: taavi@cloudcontrol1005 ~ $ os subnet set a69bdfad-d7d2-4cfa-8231-{{Gerrit|3d6d3e0074c9}} --no-dns-nameservers --dns-nameserver 172.20.255.1 * 08:36 taavi: start script to fix networking on broken bullseye instances [[phab:T347665|T347665]] === 2023-09-28 === * 20:29 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 20:26 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 18:44 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 18:40 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 12:01 wm-bot2: dcaro@urcuchillay END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 12:01 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.undrain_node * 12:01 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 12:01 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.undrain_node * 09:48 arturo: rebooting cloudgw1001/1002 for sysctl and kernel upgrades === 2023-09-27 === * 12:07 arturo: merging cloudgw firewall changes https://gerrit.wikimedia.org/r/c/operations/puppet/+/961360 * 09:36 taavi: move maintain-dbusers to cloudcontrol1005 * 01:34 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 01:29 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2023-09-26 === * 17:07 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 17:06 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2023-09-24 === * 15:39 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 15:37 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 15:35 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 15:32 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2023-09-23 === * 08:40 taavi: restart keystone === 2023-09-22 === * 14:03 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.drain_node (exit_code=99) * 14:02 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.drain_node * 14:02 wm-bot2: dcaro@urcuchillay END (PASS) - Cookbook wmcs.ceph.drain_node (exit_code=0) * 14:01 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.drain_node * 14:01 wm-bot2: dcaro@urcuchillay END (PASS) - Cookbook wmcs.ceph.undrain_node (exit_code=0) * 14:00 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.undrain_node * 14:00 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.drain_node (exit_code=99) * 13:57 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.drain_node * 13:56 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.drain_node (exit_code=99) * 13:55 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.drain_node * 13:53 wm-bot2: dcaro@urcuchillay END (PASS) - Cookbook wmcs.ceph.undrain_node (exit_code=0) * 13:52 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.undrain_node * 13:49 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.undrain_node (exit_code=99) * 13:48 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.undrain_node * 13:48 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.undrain_node (exit_code=99) * 13:47 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.undrain_node * 13:47 wm-bot2: dcaro@urcuchillay END (PASS) - Cookbook wmcs.ceph.drain_node (exit_code=0) * 13:46 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.drain_node * 13:46 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.drain_node (exit_code=99) * 13:44 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.drain_node * 13:43 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.drain_node (exit_code=99) * 13:43 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.drain_node * 13:41 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.drain_node (exit_code=99) * 13:41 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.drain_node * 13:04 wm-bot2: dcaro@urcuchillay END (PASS) - Cookbook wmcs.ceph.drain_node (exit_code=0) * 13:03 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.drain_node * 12:37 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.drain_node (exit_code=99) * 12:37 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.drain_node * 12:33 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.drain_node (exit_code=99) * 12:33 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.drain_node * 12:33 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.drain_node (exit_code=99) * 12:28 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.drain_node * 11:43 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.drain_node (exit_code=99) * 11:42 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.drain_node === 2023-09-21 === * 16:35 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.refresh_puppet_certs (exit_code=0) on tools-db-3.tools.eqiad1.wikimedia.cloud * 16:34 fnegri@cloudcumin1001: START - Cookbook wmcs.vps.refresh_puppet_certs on tools-db-3.tools.eqiad1.wikimedia.cloud * 02:12 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.restart_openstack (exit_code=97) * 02:12 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2023-09-20 === * 21:49 root@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 21:45 root@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 21:38 root@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 21:35 root@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 21:34 root@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 21:31 root@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 16:26 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 16:23 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 11:12 arturo: moving openstack API endpoint to cloudlb ([[phab:T346439|T346439]]) * 10:28 arturo: running SQL command `update domains set master="172.20.1.5:5354 172.20.2.4:5354 185.15.56.162:5354 185.15.56.163:5354";` on cloudservices1005/1006 ([[phab:T346042|T346042]]) * 10:06 arturo: running SQL command `update domains set master="185.15.56.162:5354 185.15.56.163:5354"` on cloudservices1005/1005 ([[phab:T346042|T346042]]) === 2023-09-19 === * 18:54 andrewbogott: depooling clouddb1019 to let it recover from high memory use. [[phab:T346826|T346826]] * 15:57 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 15:53 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 15:47 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 15:46 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2023-09-18 === * 11:53 taavi: update designate urls in Keystone to point to openstack-next, until cloudlb is serving the main openstack address [[phab:T346042|T346042]] * 11:40 arturo: decomission cloudservices1005 [[phab:T346042|T346042]] in preparation for re-racking * 08:45 arturo: hardcode `185.15.56.161 openstack.eqiad1.wikimediacloud.org` in /etc/hosts in cloudcontrol1005 for [[phab:T346441|T346441]] === 2023-09-15 === * 11:43 arturo: merging NAT change for [[phab:T346426|T346426]] in cloudgw * 10:33 arturo: faiolver cloudgw1001 into cloudgw1002, investigating a nftables syntax error ([[phab:T346432|T346432]]) === 2023-09-14 === * 21:25 wm-bot2: andrew@bullseye END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) * 21:25 wm-bot2: andrew@bullseye START - Cookbook wmcs.openstack.cloudvirt.drain * 21:23 wm-bot2: andrew@bullseye END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) * 21:23 wm-bot2: andrew@bullseye START - Cookbook wmcs.openstack.cloudvirt.drain * 21:18 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) * 21:18 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain * 17:13 wm-bot2: fran@wmf3169 admin END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) ([[phab:T345810|T345810]]) * 17:06 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudvirt.drain ([[phab:T345810|T345810]]) * 17:01 wm-bot2: fran@wmf3169 admin END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) ([[phab:T345810|T345810]]) * 16:56 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudvirt.drain ([[phab:T345810|T345810]]) * 16:51 wm-bot2: fran@wmf3169 admin END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) ([[phab:T345810|T345810]]) * 16:51 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudvirt.drain ([[phab:T345810|T345810]]) * 16:29 wm-bot2: fran@wmf3169 admin END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) ([[phab:T345810|T345810]]) * 16:29 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudvirt.drain ([[phab:T345810|T345810]]) * 16:25 fnegri@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=97) ([[phab:T345810|T345810]]) * 16:25 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain ([[phab:T345810|T345810]]) * 16:12 topranks: DNS operation: remove old DNS entry for ns0.openstack.eqiad1.wikimediacloud.org. on wikimedia authdns (was pointing to 208.80.154.148) * 16:10 topranks: DNS operation: add new DNS entry for ns0.openstack.eqiad1.wikimediacloud.org. on wikimedia authdns pointing to 185.15.56.162 * 14:42 arturo: DNS operation: route 208.80.154.148 to cloudservices1006 in anticipation of cloudservices1005 decom ([[phab:T346042|T346042]]) * 12:11 arturo: enable puppet on cloudservices1006 to drop local NAT hacks and enable new DNS auth IP address ([[phab:T346042|T346042]]) === 2023-09-13 === * 17:11 wm-bot2: fran@wmf3169 END (PASS) - Cookbook wmcs.openstack.cloudnet.reboot_node (exit_code=0) ([[phab:T345811|T345811]]) * 17:08 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudnet.reboot_node ([[phab:T345811|T345811]]) * 17:04 wm-bot2: fran@wmf3169 END (FAIL) - Cookbook wmcs.openstack.cloudnet.reboot_node (exit_code=99) ([[phab:T345811|T345811]]) * 17:04 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudnet.reboot_node ([[phab:T345811|T345811]]) * 16:57 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.openstack.cloudnet.reboot_node (exit_code=99) ([[phab:T345811|T345811]]) * 16:57 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.openstack.cloudnet.reboot_node ([[phab:T345811|T345811]]) * 16:56 wm-bot2: fran@wmf3169 END (FAIL) - Cookbook wmcs.openstack.cloudnet.reboot_node (exit_code=99) ([[phab:T345811|T345811]]) * 16:56 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudnet.reboot_node ([[phab:T345811|T345811]]) * 16:53 wm-bot2: fran@wmf3169 END (FAIL) - Cookbook wmcs.openstack.cloudnet.reboot_node (exit_code=99) ([[phab:T345811|T345811]]) * 16:53 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudnet.reboot_node ([[phab:T345811|T345811]]) * 16:49 wm-bot2: fran@wmf3169 END (FAIL) - Cookbook wmcs.openstack.cloudnet.reboot_node (exit_code=99) ([[phab:T345811|T345811]]) * 16:49 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudnet.reboot_node ([[phab:T345811|T345811]]) * 16:41 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudnet.reboot_node (exit_code=99) ([[phab:T345811|T345811]]) * 16:40 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudnet.reboot_node ([[phab:T345811|T345811]]) * 01:57 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 01:57 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 01:55 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 01:55 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 01:54 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 01:54 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2023-09-12 === * 18:06 andrewbogott: update domains set master='185.15.56.163:5354 208.80.154.11:5354 10.64.151.4:5354'; on cloudservices1005 + cloudservices1006 * 17:59 andrewbogott: mysql:root@localhost [pdns]> update domains set master='185.15.56.163:5354 208.80.154.11:5354'; * 17:59 andrewbogott: "designate-manage pool update' on cloudservices1005 to remove cloudservices1004 from the pool * 17:42 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 17:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 15:16 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 15:12 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 15:11 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 15:07 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2023-09-11 === * 12:36 arturo: update DNS resolver cloud-wide to use 172.20.255.1 ([[phab:T342621|T342621]]) === 2023-09-10 === * 02:52 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 02:49 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 02:42 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 02:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2023-09-07 === * 10:09 dhinus: reimaging cloudcontrol2001-dev to bookworm ([[phab:T345810|T345810]]) === 2023-09-06 === * 16:56 fnegri@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=97) * 16:56 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node * 16:53 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) * 16:52 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node * 16:51 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) * 16:51 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node * 16:51 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) * 16:51 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node * 16:51 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) * 16:51 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node * 16:37 wm-bot2: fran@wmf3169 END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) * 16:37 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node * 16:37 wm-bot2: fran@wmf3169 END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) * 16:35 wm-bot2: fran@wmf3169 END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) * 16:35 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node * 16:34 wm-bot2: fran@wmf3169 END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) * 16:34 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node * 16:23 fnegri@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=97) * 16:23 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node * 16:13 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudweb.set_maintenance (exit_code=99) * 16:12 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudweb.set_maintenance === 2023-09-05 === * 12:34 arturo: synced pdns database from cloudservices1004 to cloudservices1006 ([[phab:T345240|T345240]]) * 12:18 arturo: updating pools.yaml in all cloudservices designate nodes ([[phab:T345240|T345240]]) * 10:54 arturo: running SQL command `update domains set master="208.80.154.11:5354 208.80.154.148:5354 10.64.151.4:5354";` on all 3 cloudservices nodes ([[phab:T345240|T345240]]) === 2023-09-04 === * 15:58 arturo: stop and mask designate-sink.service @ cloudservices1006 * 14:19 arturo: started all designate services on cloudservices1006 [[phab:T345240|T345240]] * 10:46 arturo: added designate galera DB grants for cloudlb [[phab:T345240|T345240]] * 08:40 arturo: stopped all designate services on cloudservices1006 [[phab:T345240|T345240]] === 2023-09-01 === * 10:28 wm-bot2: fran@wmf3169 END (FAIL) - Cookbook wmcs.openstack.network.tests (exit_code=1) ([[phab:T345282|T345282]]) * 10:25 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.network.tests ([[phab:T345282|T345282]]) === 2023-08-31 === * 15:14 wm-bot2: fran@wmf3169 END (FAIL) - Cookbook wmcs.toolforge.grid.get_cluster_status (exit_code=99) * 15:14 wm-bot2: fran@wmf3169 START - Cookbook wmcs.toolforge.grid.get_cluster_status * 12:49 wm-bot2: dcaro@urcuchillay END (PASS) - Cookbook wmcs.openstack.cloudnet.show (exit_code=0) * 12:49 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.openstack.cloudnet.show * 12:48 wm-bot2: dcaro@urcuchillay END (PASS) - Cookbook wmcs.ceph.unset_cluster_maintenance (exit_code=0) * 12:48 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.unset_cluster_maintenance * 12:47 wm-bot2: dcaro@urcuchillay END (PASS) - Cookbook wmcs.ceph.set_cluster_in_maintenance (exit_code=0) * 12:47 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.set_cluster_in_maintenance * 12:46 wm-bot2: fran@wmf3169 END (PASS) - Cookbook wmcs.ceph.set_cluster_in_maintenance (exit_code=0) * 12:46 wm-bot2: fran@wmf3169 START - Cookbook wmcs.ceph.set_cluster_in_maintenance === 2023-08-30 === * 13:07 wm-bot2: dcaro testing stuff === 2023-08-28 === * 15:05 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 15:05 wm-bot2: Restarting openstack services on cloudservices1005: ['designate-producer', 'designate-sink', 'designate-worker', 'designate-central', 'designate-mdns', 'designate-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:05 wm-bot2: Restarting openstack services on cloudservices1004: ['designate-worker', 'designate-api', 'designate-mdns', 'designate-producer', 'designate-central', 'designate-sink', 'designate-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:05 wm-bot2: Restarting openstack services on cloudnet1006: ['neutron-linuxbridge-agent', 'neutron-metadata-agent', 'neutron-dhcp-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:05 wm-bot2: Restarting openstack services on cloudnet1005: ['neutron-linuxbridge-agent', 'neutron-dhcp-agent', 'neutron-metadata-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:05 wm-bot2: Restarting openstack services on cloudbackup2001: ['cinder-backup'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:05 wm-bot2: Restarting openstack services on cloudbackup2002: ['cinder-backup'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:05 wm-bot2: Restarting openstack services on cloudvirtlocal1003: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:05 wm-bot2: Restarting openstack services on cloudvirtlocal1002: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:05 wm-bot2: Restarting openstack services on cloudvirtlocal1001: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:05 wm-bot2: Restarting openstack services on cloudvirt1054: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:05 wm-bot2: Restarting openstack services on cloudvirt1055: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:05 wm-bot2: Restarting openstack services on cloudvirt1060: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:05 wm-bot2: Restarting openstack services on cloudvirt1058: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:05 wm-bot2: Restarting openstack services on cloudvirt1059: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:05 wm-bot2: Restarting openstack services on cloudvirt1061: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:04 wm-bot2: Restarting openstack services on cloudvirt1057: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:04 wm-bot2: Restarting openstack services on cloudvirt1056: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:04 wm-bot2: Restarting openstack services on cloudvirt1051: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:04 wm-bot2: Restarting openstack services on cloudvirt1050: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:04 wm-bot2: Restarting openstack services on cloudvirt1049: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:04 wm-bot2: Restarting openstack services on cloudvirt1053: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:04 wm-bot2: Restarting openstack services on cloudvirt1052: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:04 wm-bot2: Restarting openstack services on cloudvirt1048: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:04 wm-bot2: Restarting openstack services on cloudcontrol1007: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:03 wm-bot2: Restarting openstack services on cloudcontrol1006: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:03 wm-bot2: Restarting openstack services on cloudvirt-wdqs1003: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:03 wm-bot2: Restarting openstack services on cloudvirt-wdqs1002: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:03 wm-bot2: Restarting openstack services on cloudvirt-wdqs1001: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:03 wm-bot2: Restarting openstack services on cloudvirt1047: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:03 wm-bot2: Restarting openstack services on cloudvirt1038: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:03 wm-bot2: Restarting openstack services on cloudvirt1042: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:03 wm-bot2: Restarting openstack services on cloudvirt1044: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:03 wm-bot2: Restarting openstack services on cloudvirt1041: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:03 wm-bot2: Restarting openstack services on cloudvirt1046: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:03 wm-bot2: Restarting openstack services on cloudvirt1043: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:03 wm-bot2: Restarting openstack services on cloudvirt1045: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:03 wm-bot2: Restarting openstack services on cloudvirt1040: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:02 wm-bot2: Restarting openstack services on cloudvirt1036: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:02 wm-bot2: Restarting openstack services on cloudvirt1034: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:02 wm-bot2: Restarting openstack services on cloudvirt1039: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:02 wm-bot2: Restarting openstack services on cloudvirt1037: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:02 wm-bot2: Restarting openstack services on cloudvirt1035: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:02 wm-bot2: Restarting openstack services on cloudvirt1033: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:02 wm-bot2: Restarting openstack services on cloudvirt1031: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:02 wm-bot2: Restarting openstack services on cloudvirt1032: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:02 wm-bot2: Restarting openstack services on cloudcontrol1005: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:02 wm-bot2: Restarting openstack services on cloudvirt1028: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:02 wm-bot2: Restarting openstack services on cloudvirt1030: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:02 wm-bot2: Restarting openstack services on cloudvirt1027: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:01 wm-bot2: Restarting openstack services on cloudvirt1026: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:01 wm-bot2: Restarting openstack services on cloudvirt1029: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:01 wm-bot2: Restarting openstack services on cloudvirt1025: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:01 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2023-08-23 === * 16:10 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 16:10 wm-bot2: Restarting openstack services on cloudservices2004-dev: ['designate-mdns', 'designate-sink', 'designate-central', 'designate-producer', 'designate-worker', 'designate-agent'] - cookbook ran by root@cloudcumin1001 * 16:10 wm-bot2: Restarting openstack services on cloudservices2005-dev: ['designate-central', 'designate-sink', 'designate-worker', 'designate-producer', 'designate-mdns', 'designate-agent'] - cookbook ran by root@cloudcumin1001 * 16:10 wm-bot2: Restarting openstack services on cloudnet2005-dev: ['neutron-metadata-agent', 'neutron-dhcp-agent', 'neutron-linuxbridge-agent'] - cookbook ran by root@cloudcumin1001 * 16:10 wm-bot2: Restarting openstack services on cloudnet2006-dev: ['neutron-dhcp-agent', 'neutron-metadata-agent', 'neutron-linuxbridge-agent'] - cookbook ran by root@cloudcumin1001 * 16:09 wm-bot2: Restarting openstack services on cloudbackup1002-dev: ['cinder-backup'] - cookbook ran by root@cloudcumin1001 * 16:09 wm-bot2: Restarting openstack services on cloudbackup1001-dev: ['cinder-backup'] - cookbook ran by root@cloudcumin1001 * 16:09 wm-bot2: Restarting openstack services on cloudcontrol2005-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by root@cloudcumin1001 * 16:09 wm-bot2: Restarting openstack services on cloudvirt2003-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by root@cloudcumin1001 * 16:09 wm-bot2: Restarting openstack services on cloudvirt2002-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by root@cloudcumin1001 * 16:09 wm-bot2: Restarting openstack services on cloudvirt2001-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by root@cloudcumin1001 * 16:08 wm-bot2: Restarting openstack services on cloudcontrol2004-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by root@cloudcumin1001 * 16:08 wm-bot2: Restarting openstack services on cloudcontrol2001-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by root@cloudcumin1001 * 16:08 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2023-08-15 === * 19:33 wm-bot2: Restarting openstack services on cloudvirt2003-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by root@cloudcumin1001 * 19:33 wm-bot2: Restarting openstack services on cloudvirt2002-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by root@cloudcumin1001 * 19:33 wm-bot2: Restarting openstack services on cloudvirt2001-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by root@cloudcumin1001 * 19:32 wm-bot2: Restarting openstack services on cloudcontrol2004-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by root@cloudcumin1001 * 19:32 wm-bot2: Restarting openstack services on cloudcontrol2001-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by root@cloudcumin1001 * 19:32 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 16:01 andrewbogott: rebooting cloudvirt2001-dev in an attempt to figure out what's happening with bastions * 15:42 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 15:42 wm-bot2: Restarting openstack services on cloudservices2004-dev: ['designate-mdns', 'designate-sink', 'designate-central', 'designate-producer', 'designate-worker', 'designate-agent'] - cookbook ran by root@cloudcumin1001 * 15:42 wm-bot2: Restarting openstack services on cloudservices2005-dev: ['designate-central', 'designate-sink', 'designate-worker', 'designate-producer', 'designate-mdns', 'designate-agent'] - cookbook ran by root@cloudcumin1001 * 15:42 wm-bot2: Restarting openstack services on cloudnet2005-dev: ['neutron-metadata-agent', 'neutron-dhcp-agent', 'neutron-linuxbridge-agent'] - cookbook ran by root@cloudcumin1001 * 15:42 wm-bot2: Restarting openstack services on cloudnet2006-dev: ['neutron-dhcp-agent', 'neutron-metadata-agent', 'neutron-linuxbridge-agent'] - cookbook ran by root@cloudcumin1001 * 15:42 wm-bot2: Restarting openstack services on cloudbackup1002-dev: ['cinder-backup'] - cookbook ran by root@cloudcumin1001 * 15:42 wm-bot2: Restarting openstack services on cloudbackup1001-dev: ['cinder-backup'] - cookbook ran by root@cloudcumin1001 * 15:41 wm-bot2: Restarting openstack services on cloudcontrol2005-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by root@cloudcumin1001 * 15:41 wm-bot2: Restarting openstack services on cloudvirt2003-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by root@cloudcumin1001 * 15:41 wm-bot2: Restarting openstack services on cloudvirt2002-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by root@cloudcumin1001 * 15:40 wm-bot2: Restarting openstack services on cloudvirt2001-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by root@cloudcumin1001 * 15:40 wm-bot2: Restarting openstack services on cloudcontrol2004-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by root@cloudcumin1001 * 15:40 wm-bot2: Restarting openstack services on cloudcontrol2001-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by root@cloudcumin1001 * 15:40 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 12:39 dcaro: removed some logs from the cloudmetrics1003:/var/log/carbon/ directory and stopped the carbon processes (they were crashing and filling up the disk with logs) === 2023-08-14 === * 22:11 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 22:11 wm-bot2: Restarting openstack services on cloudservices2004-dev: ['designate-mdns', 'designate-sink', 'designate-central', 'designate-producer', 'designate-worker', 'designate-agent'] - cookbook ran by root@cloudcumin1001 * 22:11 wm-bot2: Restarting openstack services on cloudservices2005-dev: ['designate-central', 'designate-sink', 'designate-worker', 'designate-producer', 'designate-mdns', 'designate-agent'] - cookbook ran by root@cloudcumin1001 * 22:11 wm-bot2: Restarting openstack services on cloudnet2005-dev: ['neutron-metadata-agent', 'neutron-dhcp-agent', 'neutron-linuxbridge-agent'] - cookbook ran by root@cloudcumin1001 * 22:11 wm-bot2: Restarting openstack services on cloudnet2006-dev: ['neutron-dhcp-agent', 'neutron-metadata-agent', 'neutron-linuxbridge-agent'] - cookbook ran by root@cloudcumin1001 * 22:10 wm-bot2: Restarting openstack services on cloudbackup1002-dev: ['cinder-backup'] - cookbook ran by root@cloudcumin1001 * 22:10 wm-bot2: Restarting openstack services on cloudbackup1001-dev: ['cinder-backup'] - cookbook ran by root@cloudcumin1001 * 22:10 wm-bot2: Restarting openstack services on cloudcontrol2005-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by root@cloudcumin1001 * 22:09 wm-bot2: Restarting openstack services on cloudvirt2003-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by root@cloudcumin1001 * 22:09 wm-bot2: Restarting openstack services on cloudvirt2002-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by root@cloudcumin1001 * 22:09 wm-bot2: Restarting openstack services on cloudvirt2001-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by root@cloudcumin1001 * 22:09 wm-bot2: Restarting openstack services on cloudcontrol2004-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by root@cloudcumin1001 * 22:08 wm-bot2: Restarting openstack services on cloudcontrol2001-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by root@cloudcumin1001 * 22:08 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2023-08-07 === * 09:39 taavi: cloud vps graphite service was disabled: https://wikitech.wikimedia.org/wiki/News/2023_Cloud_VPS_metrics_changes [[phab:T326266|T326266]] === 2023-08-03 === * 13:07 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.do_log_msg (exit_code=0) * 13:07 fnegri@cloudcumin1001: START - Cookbook wmcs.do_log_msg * 13:07 wm-bot2: Test SAL log ([[phab:T341793|T341793]]) - cookbook ran by root@cloudcumin1001 === 2023-07-31 === * 16:16 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.do_log_msg (exit_code=0) * 16:16 wm-bot2: Test SAL log ([[phab:T325756|T325756]]) - cookbook ran by root@cloudcumin1001 * 16:15 fnegri@cloudcumin1001: START - Cookbook wmcs.do_log_msg * 15:33 wm-bot2: Test SAL log ([[phab:T325756|T325756]]) - cookbook ran by root@cloudcumin1001 * 14:42 wm-bot2: Test SAL log ([[phab:T325756|T325756]]) - cookbook ran by root@cloudcumin1001 * 14:35 wm-bot2: Test SAL log ([[phab:T325756|T325756]]) - cookbook ran by root@cloudcumin1001 * 13:55 andrewbogott: recreating the codfw1dev galera cluster according to https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Troubleshooting#Galera -- mariadb is stopped (and won't start) on all three cloudcontrol nodes * 13:39 wm-bot2: Test SAL log ([[phab:T325756|T325756]]) - cookbook ran by root@cloudcumin1001 === 2023-07-27 === * 10:52 arturo: adding cloud-private subnet to cloudnet1005/1006 hosts in eqiad1 ([[phab:T342619|T342619]]) === 2023-07-19 === * 14:02 wm-bot2: Draining cloudvirt2002-dev.codfw.wmnet ([[phab:T335840|T335840]]) - cookbook ran by raymond@ubuntu * 14:02 wm-bot2: Safe rebooting cloudvirt2002-dev.codfw.wmnet ([[phab:T335840|T335840]]) - cookbook ran by raymond@ubuntu === 2023-07-17 === * 15:55 arturo: cloudcontrol1005 was shutdown earlier today ([[phab:T341495|T341495]]) * 15:35 arturo: [codfw1dev] cloudweb2002-dev up and running after reracking ([[phab:T327919|T327919]]) * 15:11 arturo: [codfw1dev] powered off cloudweb2002-dev for reracking ([[phab:T327919|T327919]]) * 12:45 taavi: removing diamond from remaining buster instances [[phab:T317032|T317032]] === 2023-07-13 === * 10:48 wm-bot2: Restarting openstack services on cloudservices1005: ['designate-producer', 'designate-sink', 'designate-worker', 'designate-central', 'designate-mdns', 'designate-agent'] - cookbook ran by dcaro@urcuchillay * 10:48 wm-bot2: Restarting openstack services on cloudservices1004: ['designate-worker', 'designate-api', 'designate-mdns', 'designate-producer', 'designate-central', 'designate-sink', 'designate-agent'] - cookbook ran by dcaro@urcuchillay * 10:48 wm-bot2: Restarting openstack services on cloudnet1006: ['neutron-linuxbridge-agent', 'neutron-metadata-agent', 'neutron-dhcp-agent'] - cookbook ran by dcaro@urcuchillay * 10:48 wm-bot2: Restarting openstack services on cloudnet1005: ['neutron-linuxbridge-agent', 'neutron-dhcp-agent', 'neutron-metadata-agent'] - cookbook ran by dcaro@urcuchillay * 10:48 wm-bot2: Restarting openstack services on cloudbackup2001: ['cinder-backup'] - cookbook ran by dcaro@urcuchillay * 10:47 wm-bot2: Restarting openstack services on cloudbackup2002: ['cinder-backup'] - cookbook ran by dcaro@urcuchillay * 10:47 wm-bot2: Restarting openstack services on cloudvirtlocal1003: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:47 wm-bot2: Restarting openstack services on cloudvirtlocal1002: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:47 wm-bot2: Restarting openstack services on cloudvirtlocal1001: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:47 wm-bot2: Restarting openstack services on cloudvirt1054: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:47 wm-bot2: Restarting openstack services on cloudvirt1055: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:47 wm-bot2: Restarting openstack services on cloudvirt1060: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:47 wm-bot2: Restarting openstack services on cloudvirt1058: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:47 wm-bot2: Restarting openstack services on cloudvirt1059: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:46 wm-bot2: Restarting openstack services on cloudvirt1061: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:46 wm-bot2: Restarting openstack services on cloudvirt1057: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:46 wm-bot2: Restarting openstack services on cloudvirt1056: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:46 wm-bot2: Restarting openstack services on cloudvirt1051: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:46 wm-bot2: Restarting openstack services on cloudvirt1050: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:46 wm-bot2: Restarting openstack services on cloudvirt1049: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:46 wm-bot2: Restarting openstack services on cloudvirt1053: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:46 wm-bot2: Restarting openstack services on cloudvirt1052: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:46 wm-bot2: Restarting openstack services on cloudvirt1048: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:45 wm-bot2: Restarting openstack services on cloudcontrol1007: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by dcaro@urcuchillay * 10:45 wm-bot2: Restarting openstack services on cloudcontrol1006: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by dcaro@urcuchillay * 10:45 wm-bot2: Restarting openstack services on cloudvirt-wdqs1003: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:45 wm-bot2: Restarting openstack services on cloudvirt-wdqs1002: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:45 wm-bot2: Restarting openstack services on cloudvirt-wdqs1001: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:45 wm-bot2: Restarting openstack services on cloudvirt1047: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:44 wm-bot2: Restarting openstack services on cloudvirt1038: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:44 wm-bot2: Restarting openstack services on cloudvirt1042: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:44 wm-bot2: Restarting openstack services on cloudvirt1044: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:44 wm-bot2: Restarting openstack services on cloudvirt1041: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:44 wm-bot2: Restarting openstack services on cloudvirt1046: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:44 wm-bot2: Restarting openstack services on cloudvirt1043: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:44 wm-bot2: Restarting openstack services on cloudvirt1045: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:44 wm-bot2: Restarting openstack services on cloudvirt1040: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:44 wm-bot2: Restarting openstack services on cloudvirt1036: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:44 wm-bot2: Restarting openstack services on cloudvirt1034: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:43 wm-bot2: Restarting openstack services on cloudvirt1039: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:43 wm-bot2: Restarting openstack services on cloudvirt1037: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:43 wm-bot2: Restarting openstack services on cloudvirt1035: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:43 wm-bot2: Restarting openstack services on cloudvirt1033: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:43 wm-bot2: Restarting openstack services on cloudvirt1031: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:43 wm-bot2: Restarting openstack services on cloudvirt1032: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:42 wm-bot2: Restarting openstack services on cloudcontrol1005: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by dcaro@urcuchillay * 10:42 wm-bot2: Restarting openstack services on cloudvirt1028: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:42 wm-bot2: Restarting openstack services on cloudvirt1030: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:42 wm-bot2: Restarting openstack services on cloudvirt1027: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:42 wm-bot2: Restarting openstack services on cloudvirt1026: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:42 wm-bot2: Restarting openstack services on cloudvirt1029: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:42 wm-bot2: Restarting openstack services on cloudvirt1025: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:38 wm-bot2: Restarting openstack services on cloudvirt1025: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:33 wm-bot2: Restarting openstack services on cloudvirt1025: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay === 2023-07-07 === * 10:28 taavi: backfilling <nowiki>{</nowiki>project<nowiki>}</nowiki>.wmcloud.org and other currently-named DNS zones to projects that don't have them === 2023-07-06 === * 17:02 wm-bot2: Restarting openstack services on cloudservices2004-dev: ['designate-mdns', 'designate-sink', 'designate-central', 'designate-producer', 'designate-worker', 'designate-agent'] - cookbook ran by andrew@bullseye * 17:02 wm-bot2: Restarting openstack services on cloudservices2005-dev: ['designate-central', 'designate-sink', 'designate-worker', 'designate-producer', 'designate-mdns', 'designate-agent'] - cookbook ran by andrew@bullseye * 17:02 wm-bot2: Restarting openstack services on cloudnet2005-dev: ['neutron-metadata-agent', 'neutron-dhcp-agent', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 17:02 wm-bot2: Restarting openstack services on cloudnet2006-dev: ['neutron-dhcp-agent', 'neutron-metadata-agent', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 17:02 wm-bot2: Restarting openstack services on cloudbackup1002-dev: ['cinder-backup'] - cookbook ran by andrew@bullseye * 17:01 wm-bot2: Restarting openstack services on cloudbackup1001-dev: ['cinder-backup'] - cookbook ran by andrew@bullseye * 17:01 wm-bot2: Restarting openstack services on cloudcontrol2005-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 17:01 wm-bot2: Restarting openstack services on cloudvirt2003-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 17:01 wm-bot2: Restarting openstack services on cloudvirt2002-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 17:01 wm-bot2: Restarting openstack services on cloudvirt2001-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 17:00 wm-bot2: Restarting openstack services on cloudcontrol2004-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 17:00 wm-bot2: Restarting openstack services on cloudcontrol2001-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye === 2023-07-04 === * 14:08 wm-bot2: Test SAL log ([[phab:T325756|T325756]]) - cookbook ran by root@cloudcumin1001 * 14:07 wm-bot2: Test SAL log ([[phab:T325756|T325756]]) - cookbook ran by root@cloudcumin1001 === 2023-06-30 === * 21:37 wm-bot2: Restarting openstack services on cloudservices2004-dev: ['designate-mdns', 'designate-sink', 'designate-central', 'designate-producer', 'designate-worker', 'designate-agent'] - cookbook ran by andrew@bullseye * 21:37 wm-bot2: Restarting openstack services on cloudservices2005-dev: ['designate-central', 'designate-sink', 'designate-worker', 'designate-producer', 'designate-mdns', 'designate-agent'] - cookbook ran by andrew@bullseye * 21:36 wm-bot2: Restarting openstack services on cloudnet2005-dev: ['neutron-metadata-agent', 'neutron-dhcp-agent', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 21:36 wm-bot2: Restarting openstack services on cloudnet2006-dev: ['neutron-dhcp-agent', 'neutron-metadata-agent', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 21:36 wm-bot2: Restarting openstack services on cloudbackup1002-dev: ['cinder-backup'] - cookbook ran by andrew@bullseye * 21:36 wm-bot2: Restarting openstack services on cloudbackup1001-dev: ['cinder-backup'] - cookbook ran by andrew@bullseye * 21:35 wm-bot2: Restarting openstack services on cloudcontrol2005-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 21:35 wm-bot2: Restarting openstack services on cloudvirt2003-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 21:35 wm-bot2: Restarting openstack services on cloudvirt2002-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 21:35 wm-bot2: Restarting openstack services on cloudvirt2001-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 21:35 wm-bot2: Restarting openstack services on cloudcontrol2004-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 21:34 wm-bot2: Restarting openstack services on cloudcontrol2001-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye === 2023-06-29 === * 19:49 wm-bot2: Restarting openstack services on cloudservices2004-dev: ['designate-mdns', 'designate-sink', 'designate-central', 'designate-producer', 'designate-worker', 'designate-agent'] - cookbook ran by andrew@bullseye * 19:49 wm-bot2: Restarting openstack services on cloudservices2005-dev: ['designate-central', 'designate-sink', 'designate-worker', 'designate-producer', 'designate-mdns', 'designate-agent'] - cookbook ran by andrew@bullseye * 19:49 wm-bot2: Restarting openstack services on cloudnet2005-dev: ['neutron-metadata-agent', 'neutron-dhcp-agent', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 19:49 wm-bot2: Restarting openstack services on cloudnet2006-dev: ['neutron-dhcp-agent', 'neutron-metadata-agent', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 19:49 wm-bot2: Restarting openstack services on cloudbackup1002-dev: ['cinder-backup'] - cookbook ran by andrew@bullseye * 19:49 wm-bot2: Restarting openstack services on cloudbackup1001-dev: ['cinder-backup'] - cookbook ran by andrew@bullseye * 19:48 wm-bot2: Restarting openstack services on cloudcontrol2005-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 19:48 wm-bot2: Restarting openstack services on cloudvirt2003-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 19:48 wm-bot2: Restarting openstack services on cloudvirt2002-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 19:48 wm-bot2: Restarting openstack services on cloudvirt2001-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 19:48 wm-bot2: Restarting openstack services on cloudcontrol2004-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 19:47 wm-bot2: Restarting openstack services on cloudcontrol2001-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 19:34 wm-bot2: Restarting openstack services on cloudservices2004-dev: ['designate-mdns', 'designate-sink', 'designate-central', 'designate-producer', 'designate-worker', 'designate-agent'] - cookbook ran by andrew@bullseye * 19:34 wm-bot2: Restarting openstack services on cloudservices2005-dev: ['designate-central', 'designate-sink', 'designate-worker', 'designate-producer', 'designate-mdns', 'designate-agent'] - cookbook ran by andrew@bullseye * 19:34 wm-bot2: Restarting openstack services on cloudnet2005-dev: ['neutron-metadata-agent', 'neutron-dhcp-agent', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 19:34 wm-bot2: Restarting openstack services on cloudnet2006-dev: ['neutron-dhcp-agent', 'neutron-metadata-agent', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 19:34 wm-bot2: Restarting openstack services on cloudbackup1002-dev: ['cinder-backup'] - cookbook ran by andrew@bullseye * 19:33 wm-bot2: Restarting openstack services on cloudcontrol2005-dev: ['cinder-volume', 'cinder-scheduler'] - cookbook ran by andrew@bullseye * 19:33 wm-bot2: Restarting openstack services on cloudbackup1001-dev: ['cinder-backup'] - cookbook ran by andrew@bullseye * 19:33 wm-bot2: Restarting openstack services on cloudcontrol2004-dev: ['cinder-scheduler', 'cinder-volume'] - cookbook ran by andrew@bullseye * 19:33 wm-bot2: Restarting openstack services on cloudcontrol2001-dev: ['cinder-scheduler', 'cinder-volume'] - cookbook ran by andrew@bullseye * 19:33 wm-bot2: Restarting openstack services on cloudvirt2001-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 19:33 wm-bot2: Restarting openstack services on cloudvirt2002-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 19:33 wm-bot2: Restarting openstack services on cloudvirt2003-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 19:30 wm-bot2: Restarting openstack services on cloudcontrol2004-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 19:30 wm-bot2: Restarting openstack services on cloudcontrol2001-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye === 2023-06-26 === * 08:46 arturo: [codfw1dev] manually start mariadb service @ cloudinfra-db-01.cloudinfra-codfw1dev.codfw1dev.wikimedia.cloud === 2023-06-24 === * 19:21 wm-bot2: Created new flavor: g3.cores16.ram34.disk20 (id:7dd33202-32c3-4bc7-b2d4-{{Gerrit|10c2ebe7e5c5}}) - cookbook ran by andrew@bullseye * 19:20 wm-bot2: Created new flavor: g3.cores16.ram34816.disk20.admin (id:d76925a8-b58b-489e-a3d9-{{Gerrit|c68c7744b551}}) - cookbook ran by andrew@bullseye === 2023-06-23 === * 16:45 wm-bot2: Restarting openstack services on cloudservices1005: ['designate-producer', 'designate-sink', 'designate-worker', 'designate-central', 'designate-mdns', 'designate-agent'] - cookbook ran by andrew@bullseye * 16:45 wm-bot2: Restarting openstack services on cloudservices1004: ['designate-worker', 'designate-api', 'designate-mdns', 'designate-producer', 'designate-central', 'designate-sink', 'designate-agent'] - cookbook ran by andrew@bullseye * 16:45 wm-bot2: Restarting openstack services on cloudnet1006: ['neutron-linuxbridge-agent', 'neutron-metadata-agent', 'neutron-dhcp-agent'] - cookbook ran by andrew@bullseye * 16:45 wm-bot2: Restarting openstack services on cloudnet1005: ['neutron-linuxbridge-agent', 'neutron-dhcp-agent', 'neutron-metadata-agent'] - cookbook ran by andrew@bullseye * 16:45 wm-bot2: Restarting openstack services on cloudbackup2001: ['cinder-backup'] - cookbook ran by andrew@bullseye * 16:45 wm-bot2: Restarting openstack services on cloudbackup2002: ['cinder-backup'] - cookbook ran by andrew@bullseye * 16:44 wm-bot2: Restarting openstack services on cloudvirtlocal1003: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:44 wm-bot2: Restarting openstack services on cloudvirtlocal1002: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:44 wm-bot2: Restarting openstack services on cloudvirtlocal1001: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:44 wm-bot2: Restarting openstack services on cloudvirt1054: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:44 wm-bot2: Restarting openstack services on cloudvirt1055: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:44 wm-bot2: Restarting openstack services on cloudvirt1060: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:44 wm-bot2: Restarting openstack services on cloudvirt1058: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:44 wm-bot2: Restarting openstack services on cloudvirt1059: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:44 wm-bot2: Restarting openstack services on cloudvirt1061: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:44 wm-bot2: Restarting openstack services on cloudvirt1057: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:44 wm-bot2: Restarting openstack services on cloudvirt1056: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:44 wm-bot2: Restarting openstack services on cloudvirt1051: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:43 wm-bot2: Restarting openstack services on cloudvirt1050: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:43 wm-bot2: Restarting openstack services on cloudvirt1049: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:43 wm-bot2: Restarting openstack services on cloudvirt1053: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:43 wm-bot2: Restarting openstack services on cloudvirt1052: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:43 wm-bot2: Restarting openstack services on cloudvirt1048: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:43 wm-bot2: Restarting openstack services on cloudcontrol1007: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 16:42 wm-bot2: Restarting openstack services on cloudcontrol1006: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 16:42 wm-bot2: Restarting openstack services on cloudvirt-wdqs1003: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:42 wm-bot2: Restarting openstack services on cloudvirt-wdqs1002: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:42 wm-bot2: Restarting openstack services on cloudvirt-wdqs1001: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:42 wm-bot2: Restarting openstack services on cloudvirt1047: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:42 wm-bot2: Restarting openstack services on cloudvirt1038: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:42 wm-bot2: Restarting openstack services on cloudvirt1042: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:42 wm-bot2: Restarting openstack services on cloudvirt1044: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:42 wm-bot2: Restarting openstack services on cloudvirt1041: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:42 wm-bot2: Restarting openstack services on cloudvirt1046: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:41 wm-bot2: Restarting openstack services on cloudvirt1043: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:41 wm-bot2: Restarting openstack services on cloudvirt1045: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:41 wm-bot2: Restarting openstack services on cloudvirt1040: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:41 wm-bot2: Restarting openstack services on cloudvirt1036: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:41 wm-bot2: Restarting openstack services on cloudvirt1034: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:41 wm-bot2: Restarting openstack services on cloudvirt1039: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:41 wm-bot2: Restarting openstack services on cloudvirt1037: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:41 wm-bot2: Restarting openstack services on cloudvirt1035: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:41 wm-bot2: Restarting openstack services on cloudvirt1033: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:41 wm-bot2: Restarting openstack services on cloudvirt1031: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:41 wm-bot2: Restarting openstack services on cloudvirt1032: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:40 wm-bot2: Restarting openstack services on cloudcontrol1005: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 16:40 wm-bot2: Restarting openstack services on cloudvirt1028: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:40 wm-bot2: Restarting openstack services on cloudvirt1030: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:40 wm-bot2: Restarting openstack services on cloudvirt1027: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:40 wm-bot2: Restarting openstack services on cloudvirt1026: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:40 wm-bot2: Restarting openstack services on cloudvirt1029: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:40 wm-bot2: Restarting openstack services on cloudvirt1025: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 13:59 andrewbogott: rebooting every VM in codfw1dev === 2023-06-13 === * 15:35 wm-bot2: Restarting openstack services on cloudservices1005: ['designate-producer', 'designate-sink', 'designate-worker', 'designate-central', 'designate-mdns', 'designate-agent'] - cookbook ran by andrew@bullseye * 15:35 wm-bot2: Restarting openstack services on cloudservices1004: ['designate-worker', 'designate-api', 'designate-mdns', 'designate-producer', 'designate-central', 'designate-sink', 'designate-agent'] - cookbook ran by andrew@bullseye * 15:35 wm-bot2: Restarting openstack services on cloudnet1006: ['neutron-linuxbridge-agent', 'neutron-metadata-agent', 'neutron-dhcp-agent'] - cookbook ran by andrew@bullseye * 15:35 wm-bot2: Restarting openstack services on cloudnet1005: ['neutron-linuxbridge-agent', 'neutron-dhcp-agent', 'neutron-metadata-agent'] - cookbook ran by andrew@bullseye * 15:35 wm-bot2: Restarting openstack services on cloudbackup2001: ['cinder-backup'] - cookbook ran by andrew@bullseye * 15:35 wm-bot2: Restarting openstack services on cloudbackup2002: ['cinder-backup'] - cookbook ran by andrew@bullseye * 15:35 wm-bot2: Restarting openstack services on cloudvirtlocal1003: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:35 wm-bot2: Restarting openstack services on cloudvirtlocal1002: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:35 wm-bot2: Restarting openstack services on cloudvirtlocal1001: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:35 wm-bot2: Restarting openstack services on cloudvirt1054: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:34 wm-bot2: Restarting openstack services on cloudvirt1055: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:34 wm-bot2: Restarting openstack services on cloudvirt1060: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:34 wm-bot2: Restarting openstack services on cloudvirt1058: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:34 wm-bot2: Restarting openstack services on cloudvirt1059: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:34 wm-bot2: Restarting openstack services on cloudvirt1061: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:34 wm-bot2: Restarting openstack services on cloudvirt1057: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:34 wm-bot2: Restarting openstack services on cloudvirt1056: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:34 wm-bot2: Restarting openstack services on cloudvirt1051: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:34 wm-bot2: Restarting openstack services on cloudvirt1050: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:34 wm-bot2: Restarting openstack services on cloudvirt1049: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:34 wm-bot2: Restarting openstack services on cloudvirt1053: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:34 wm-bot2: Restarting openstack services on cloudvirt1052: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:34 wm-bot2: Restarting openstack services on cloudvirt1048: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:33 wm-bot2: Restarting openstack services on cloudcontrol1007: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 15:33 wm-bot2: Restarting openstack services on cloudcontrol1006: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 15:33 wm-bot2: Restarting openstack services on cloudvirt-wdqs1003: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:33 wm-bot2: Restarting openstack services on cloudvirt-wdqs1002: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:33 wm-bot2: Restarting openstack services on cloudvirt-wdqs1001: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:33 wm-bot2: Restarting openstack services on cloudvirt1047: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:33 wm-bot2: Restarting openstack services on cloudvirt1038: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:32 wm-bot2: Restarting openstack services on cloudvirt1042: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:32 wm-bot2: Restarting openstack services on cloudvirt1044: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:32 wm-bot2: Restarting openstack services on cloudvirt1041: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:32 wm-bot2: Restarting openstack services on cloudvirt1046: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:32 wm-bot2: Restarting openstack services on cloudvirt1043: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:32 wm-bot2: Restarting openstack services on cloudvirt1045: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:32 wm-bot2: Restarting openstack services on cloudvirt1040: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:32 wm-bot2: Restarting openstack services on cloudvirt1036: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:32 wm-bot2: Restarting openstack services on cloudvirt1034: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:32 wm-bot2: Restarting openstack services on cloudvirt1039: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:32 wm-bot2: Restarting openstack services on cloudvirt1037: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:32 wm-bot2: Restarting openstack services on cloudvirt1035: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:31 wm-bot2: Restarting openstack services on cloudvirt1033: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:31 wm-bot2: Restarting openstack services on cloudvirt1031: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:31 wm-bot2: Restarting openstack services on cloudvirt1032: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:31 wm-bot2: Restarting openstack services on cloudcontrol1005: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 15:31 wm-bot2: Restarting openstack services on cloudvirt1028: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:31 wm-bot2: Restarting openstack services on cloudvirt1030: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:31 wm-bot2: Restarting openstack services on cloudvirt1027: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:31 wm-bot2: Restarting openstack services on cloudvirt1026: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:31 wm-bot2: Restarting openstack services on cloudvirt1029: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:31 wm-bot2: Restarting openstack services on cloudvirt1025: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:22 wm-bot2: Upgraded and rebooted host cloudcontrol1007.wikimedia.org - cookbook ran by andrew@bullseye * 15:11 wm-bot2: Upgraded and rebooted host cloudcontrol1006.wikimedia.org - cookbook ran by andrew@bullseye * 14:58 wm-bot2: Upgraded and rebooted host cloudcontrol1005.wikimedia.org - cookbook ran by andrew@bullseye === 2023-06-12 === * 19:37 wm-bot2: Restarting openstack services on cloudservices2004-dev: ['designate-mdns', 'designate-sink', 'designate-central', 'designate-producer', 'designate-worker', 'designate-agent'] - cookbook ran by andrew@bullseye * 19:37 wm-bot2: Restarting openstack services on cloudservices2005-dev: ['designate-central', 'designate-sink', 'designate-worker', 'designate-producer', 'designate-mdns', 'designate-agent'] - cookbook ran by andrew@bullseye * 19:37 wm-bot2: Restarting openstack services on cloudnet2005-dev: ['neutron-metadata-agent', 'neutron-dhcp-agent', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 19:37 wm-bot2: Restarting openstack services on cloudnet2006-dev: ['neutron-dhcp-agent', 'neutron-metadata-agent', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 19:37 wm-bot2: Restarting openstack services on cloudbackup1002-dev: ['cinder-backup'] - cookbook ran by andrew@bullseye * 19:37 wm-bot2: Restarting openstack services on cloudbackup1001-dev: ['cinder-backup'] - cookbook ran by andrew@bullseye * 19:37 wm-bot2: Restarting openstack services on cloudcontrol2005-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 19:36 wm-bot2: Restarting openstack services on cloudvirt2003-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 19:36 wm-bot2: Restarting openstack services on cloudvirt2002-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 19:36 wm-bot2: Restarting openstack services on cloudvirt2001-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 19:36 wm-bot2: Restarting openstack services on cloudcontrol2004-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 19:36 wm-bot2: Restarting openstack services on cloudcontrol2001-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 19:35 wm-bot2: Restarting openstack services on cloudservices2004-dev: ['designate-mdns', 'designate-sink', 'designate-central', 'designate-producer', 'designate-worker', 'designate-agent'] - cookbook ran by andrew@bullseye * 19:35 wm-bot2: Restarting openstack services on cloudservices2005-dev: ['designate-central', 'designate-sink', 'designate-worker', 'designate-producer', 'designate-mdns', 'designate-agent'] - cookbook ran by andrew@bullseye * 11:57 arturo: [codfw1dev] refresh various occurrences of old FQDNs in instance puppet via horizon ([[phab:T324992|T324992]]) === 2023-06-08 === * 16:56 andrewbogott: deleting all Stretch base images from glance * 16:45 andrewbogott: updated the bullseye image with https://cloud.debian.org/images/cloud/bullseye/20230601-1398/debian-11-genericcloud-amd64-20230601-1398.tar.xz * 12:17 wm-bot2: Drained cloudvirt1047.eqiad.wmnet ([[phab:T334644|T334644]]) - cookbook ran by dcaro@vulcanus * 12:17 wm-bot2: Set cloudvirt cloudvirt1047.eqiad.wmnet maintenance (downtime id: 02920314-1efe-4934-ad81-{{Gerrit|d2a6cf2e17ab}}, use this to unset) ([[phab:T334644|T334644]]) - cookbook ran by dcaro@vulcanus * 12:16 wm-bot2: Draining cloudvirt1047.eqiad.wmnet ([[phab:T334644|T334644]]) - cookbook ran by dcaro@vulcanus * 12:06 wm-bot2: Set cloudvirt cloudvirt1047.eqiad.wmnet maintenance (downtime id: 769349bf-465f-4f0c-a8f3-{{Gerrit|f2423631ba7e}}, use this to unset) ([[phab:T334644|T334644]]) - cookbook ran by dcaro@vulcanus * 12:05 wm-bot2: Draining cloudvirt1047.eqiad.wmnet ([[phab:T334644|T334644]]) - cookbook ran by dcaro@vulcanus === 2023-06-07 === * 10:06 dcaro: upgraded ruby2.5 to latest fixed version on all buster VMs * 08:10 dcaro: downgrading ruby2.5 to previous backport on all buster VMs === 2023-06-06 === * 19:09 andrewbogott: also increased RAM and secgroup-rule quota for Trove [[phab:T337882|T337882]] * 19:06 andrewbogott: increased trove secgroups, instances, volumes quotas from 40 to 100. Trove is too popular! [[phab:T337882|T337882]] * 18:31 wm-bot2: Restarting openstack services on cloudbackup2001: ['cinder-backup'] - cookbook ran by andrew@bullseye * 18:31 wm-bot2: Restarting openstack services on cloudbackup2002: ['cinder-backup'] - cookbook ran by andrew@bullseye * 18:30 wm-bot2: Restarting openstack services on cloudvirtlocal1003: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:30 wm-bot2: Restarting openstack services on cloudvirtlocal1002: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:30 wm-bot2: Restarting openstack services on cloudvirtlocal1001: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:30 wm-bot2: Restarting openstack services on cloudvirt1054: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:30 wm-bot2: Restarting openstack services on cloudvirt1055: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:30 wm-bot2: Restarting openstack services on cloudvirt1060: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:30 wm-bot2: Restarting openstack services on cloudvirt1058: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:30 wm-bot2: Restarting openstack services on cloudvirt1059: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:30 wm-bot2: Restarting openstack services on cloudvirt1061: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:30 wm-bot2: Restarting openstack services on cloudvirt1057: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:30 wm-bot2: Restarting openstack services on cloudvirt1056: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:30 wm-bot2: Restarting openstack services on cloudvirt1051: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:30 wm-bot2: Restarting openstack services on cloudvirt1050: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:29 wm-bot2: Restarting openstack services on cloudvirt1049: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:29 wm-bot2: Restarting openstack services on cloudvirt1053: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:29 wm-bot2: Restarting openstack services on cloudvirt1052: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:29 wm-bot2: Restarting openstack services on cloudvirt1048: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:29 wm-bot2: Restarting openstack services on cloudcontrol1007: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler'] - cookbook ran by andrew@bullseye * 18:29 wm-bot2: Restarting openstack services on cloudcontrol1006: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler'] - cookbook ran by andrew@bullseye * 18:29 wm-bot2: Restarting openstack services on cloudvirt-wdqs1003: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:29 wm-bot2: Restarting openstack services on cloudvirt-wdqs1002: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:29 wm-bot2: Restarting openstack services on cloudvirt-wdqs1001: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:29 wm-bot2: Restarting openstack services on cloudvirt1047: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:29 wm-bot2: Restarting openstack services on cloudvirt1038: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:29 wm-bot2: Restarting openstack services on cloudvirt1042: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:28 wm-bot2: Restarting openstack services on cloudvirt1044: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:28 wm-bot2: Restarting openstack services on cloudvirt1041: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:28 wm-bot2: Restarting openstack services on cloudvirt1046: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:28 wm-bot2: Restarting openstack services on cloudvirt1043: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:28 wm-bot2: Restarting openstack services on cloudvirt1045: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:28 wm-bot2: Restarting openstack services on cloudvirt1040: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:28 wm-bot2: Restarting openstack services on cloudvirt1036: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:28 wm-bot2: Restarting openstack services on cloudvirt1034: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:28 wm-bot2: Restarting openstack services on cloudvirt1039: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:28 wm-bot2: Restarting openstack services on cloudvirt1037: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:28 wm-bot2: Restarting openstack services on cloudvirt1035: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:28 wm-bot2: Restarting openstack services on cloudvirt1033: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:27 wm-bot2: Restarting openstack services on cloudvirt1031: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:27 wm-bot2: Restarting openstack services on cloudvirt1032: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:27 wm-bot2: Restarting openstack services on cloudcontrol1005: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler'] - cookbook ran by andrew@bullseye * 18:27 wm-bot2: Restarting openstack services on cloudvirt1028: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:27 wm-bot2: Restarting openstack services on cloudvirt1030: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:27 wm-bot2: Restarting openstack services on cloudvirt1027: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:27 wm-bot2: Restarting openstack services on cloudvirt1026: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:27 wm-bot2: Restarting openstack services on cloudvirt1029: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:27 wm-bot2: Restarting openstack services on cloudvirt1025: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:26 wm-bot2: Restarting openstack services on cloudvirt1029: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:26 wm-bot2: Restarting openstack services on cloudvirt1025: ['nova-compute'] - cookbook ran by andrew@bullseye * 17:55 wm-bot2: Restarting openstack services on cloudcontrol2005-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata'] - cookbook ran by andrew@bullseye * 17:55 wm-bot2: Restarting openstack services on cloudvirt2003-dev: ['nova-compute'] - cookbook ran by andrew@bullseye * 17:55 wm-bot2: Restarting openstack services on cloudvirt2002-dev: ['nova-compute'] - cookbook ran by andrew@bullseye * 17:54 wm-bot2: Restarting openstack services on cloudvirt2001-dev: ['nova-compute'] - cookbook ran by andrew@bullseye * 17:54 wm-bot2: Restarting openstack services on cloudcontrol2004-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata'] - cookbook ran by andrew@bullseye * 17:54 wm-bot2: Restarting openstack services on cloudcontrol2001-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata'] - cookbook ran by andrew@bullseye * 17:42 wm-bot2: Restarting openstack services on cloudbackup1002-dev: ['cinder-backup'] - cookbook ran by andrew@bullseye * 17:42 wm-bot2: Restarting openstack services on cloudcontrol2005-dev: ['cinder-volume', 'cinder-scheduler'] - cookbook ran by andrew@bullseye * 17:41 wm-bot2: Restarting openstack services on cloudbackup1001-dev: ['cinder-backup'] - cookbook ran by andrew@bullseye * 17:41 wm-bot2: Restarting openstack services on cloudcontrol2004-dev: ['cinder-scheduler', 'cinder-volume'] - cookbook ran by andrew@bullseye * 17:41 wm-bot2: Restarting openstack services on cloudcontrol2001-dev: ['cinder-scheduler', 'cinder-volume'] - cookbook ran by andrew@bullseye === 2023-06-05 === * 09:41 arturo: [codfw1dev] rebooting bastion-codfw1dev-02 (no IP address in the main interface) [[phab:T336963|T336963]] === 2023-06-02 === * 14:40 wm-bot2: Restarting openstack services on cloudcontrol2001-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by arturo@nostromo * 12:59 wm-bot2: Restarting openstack services on cloudvirt2001-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 12:58 wm-bot2: Restarting openstack services on cloudcontrol2004-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 12:58 wm-bot2: Restarting openstack services on cloudcontrol2001-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 12:44 wm-bot2: Restarting openstack services on cloudvirt2001-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 12:44 wm-bot2: Restarting openstack services on cloudcontrol2004-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 12:43 wm-bot2: Restarting openstack services on cloudcontrol2001-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye === 2023-05-26 === * 16:03 andrewbogott: "maintain-views --all-databases --replace-all" on clouddb1021 for [[phab:T337446|T337446]] === 2023-05-22 === * 19:54 andrewbogott: deleting project 'citelearn' as per https://wikitech.wikimedia.org/wiki/News/Cloud_VPS_2022_Purge#SHUTDOWN_citelearn === 2023-05-21 === * 23:29 wm-bot2: Restarting openstack services on cloudservices2004-dev: ['designate-mdns', 'designate-sink', 'designate-central', 'designate-producer', 'designate-worker', 'designate-agent'] - cookbook ran by andrew@bullseye * 23:29 wm-bot2: Restarting openstack services on cloudservices2005-dev: ['designate-central', 'designate-sink', 'designate-worker', 'designate-producer', 'designate-mdns', 'designate-agent'] - cookbook ran by andrew@bullseye * 23:29 wm-bot2: Restarting openstack services on cloudnet2005-dev: ['neutron-metadata-agent', 'neutron-dhcp-agent', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 23:29 wm-bot2: Restarting openstack services on cloudnet2006-dev: ['neutron-dhcp-agent', 'neutron-metadata-agent', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 23:29 wm-bot2: Restarting openstack services on cloudbackup1002-dev: ['cinder-backup'] - cookbook ran by andrew@bullseye * 23:29 wm-bot2: Restarting openstack services on cloudbackup1001-dev: ['cinder-backup'] - cookbook ran by andrew@bullseye * 23:28 wm-bot2: Restarting openstack services on cloudcontrol2005-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 23:28 wm-bot2: Restarting openstack services on cloudvirt2003-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 23:28 wm-bot2: Restarting openstack services on cloudvirt2002-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 23:28 wm-bot2: Restarting openstack services on cloudvirt2001-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 23:28 wm-bot2: Restarting openstack services on cloudcontrol2004-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 23:27 wm-bot2: Restarting openstack services on cloudcontrol2001-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 21:45 wm-bot2: Restarting openstack services on cloudservices2004-dev: ['designate-mdns', 'designate-sink', 'designate-central', 'designate-producer', 'designate-worker', 'designate-agent'] - cookbook ran by andrew@bullseye * 21:45 wm-bot2: Restarting openstack services on cloudservices2005-dev: ['designate-central', 'designate-sink', 'designate-worker', 'designate-producer', 'designate-mdns', 'designate-agent'] - cookbook ran by andrew@bullseye * 21:45 wm-bot2: Restarting openstack services on cloudnet2005-dev: ['neutron-metadata-agent', 'neutron-dhcp-agent', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 21:45 wm-bot2: Restarting openstack services on cloudnet2006-dev: ['neutron-dhcp-agent', 'neutron-metadata-agent', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 21:45 wm-bot2: Restarting openstack services on cloudbackup1002-dev: ['cinder-backup'] - cookbook ran by andrew@bullseye * 21:45 wm-bot2: Restarting openstack services on cloudbackup1001-dev: ['cinder-backup'] - cookbook ran by andrew@bullseye * 21:45 wm-bot2: Restarting openstack services on cloudcontrol2005-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 21:45 wm-bot2: Restarting openstack services on cloudvirt2003-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 21:45 wm-bot2: Restarting openstack services on cloudvirt2002-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 21:45 wm-bot2: Restarting openstack services on cloudvirt2001-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 21:44 wm-bot2: Restarting openstack services on cloudcontrol2004-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 21:44 wm-bot2: Restarting openstack services on cloudcontrol2001-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye === 2023-05-19 === * 14:53 wm-bot2: Restarting openstack services on cloudservices2004-dev: ['designate-mdns', 'designate-sink', 'designate-central', 'designate-producer', 'designate-worker', 'designate-agent'] - cookbook ran by andrew@bullseye * 14:53 wm-bot2: Restarting openstack services on cloudservices2005-dev: ['designate-central', 'designate-sink', 'designate-worker', 'designate-producer', 'designate-mdns', 'designate-agent'] - cookbook ran by andrew@bullseye * 14:53 wm-bot2: Restarting openstack services on cloudnet2005-dev: ['neutron-metadata-agent', 'neutron-dhcp-agent', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 14:53 wm-bot2: Restarting openstack services on cloudnet2006-dev: ['neutron-dhcp-agent', 'neutron-metadata-agent', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 14:53 wm-bot2: Restarting openstack services on cloudbackup1002-dev: ['cinder-backup'] - cookbook ran by andrew@bullseye * 14:53 wm-bot2: Restarting openstack services on cloudbackup1001-dev: ['cinder-backup'] - cookbook ran by andrew@bullseye * 14:53 wm-bot2: Restarting openstack services on cloudcontrol2005-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 14:52 wm-bot2: Restarting openstack services on cloudvirt2003-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 14:52 wm-bot2: Restarting openstack services on cloudvirt2002-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 14:52 wm-bot2: Restarting openstack services on cloudvirt2001-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 14:52 wm-bot2: Restarting openstack services on cloudcontrol2004-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 14:52 wm-bot2: Restarting openstack services on cloudcontrol2001-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye === 2023-05-18 === * 21:39 andrewbogott: deleting obsolete roles '4d8cad783d6342efa8414d7d36fbc034 {{!}} projectadmin_renamed_for_[[phab:T330759|T330759]]' and 'f473273fac7146b3bdbf22e5d4504f95 {{!}} user_renamed_for_[[phab:T330759|T330759]]' on eqiad1. State pre-deletion is dumped to /root/allassignmentspredeletion.txt on cloudcontrol1007. * 21:35 wm-bot2: Restarting openstack services on cloudservices2004-dev: ['designate-mdns', 'designate-sink', 'designate-central', 'designate-producer', 'designate-worker', 'designate-agent'] - cookbook ran by andrew@bullseye * 21:35 wm-bot2: Restarting openstack services on cloudservices2005-dev: ['designate-central', 'designate-sink', 'designate-worker', 'designate-producer', 'designate-mdns', 'designate-agent'] - cookbook ran by andrew@bullseye * 21:35 wm-bot2: Restarting openstack services on cloudnet2005-dev: ['neutron-metadata-agent', 'neutron-dhcp-agent', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 21:35 wm-bot2: Restarting openstack services on cloudnet2006-dev: ['neutron-dhcp-agent', 'neutron-metadata-agent', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 21:35 wm-bot2: Restarting openstack services on cloudbackup1002-dev: ['cinder-backup'] - cookbook ran by andrew@bullseye * 21:35 wm-bot2: Restarting openstack services on cloudbackup1001-dev: ['cinder-backup'] - cookbook ran by andrew@bullseye * 21:34 wm-bot2: Restarting openstack services on cloudcontrol2005-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 21:34 wm-bot2: Restarting openstack services on cloudvirt2003-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 21:34 wm-bot2: Restarting openstack services on cloudvirt2002-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 21:34 wm-bot2: Restarting openstack services on cloudvirt2001-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 21:34 wm-bot2: Restarting openstack services on cloudcontrol2004-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 21:33 wm-bot2: Restarting openstack services on cloudcontrol2001-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 19:20 wm-bot2: Restarting openstack services on cloudservices2004-dev: ['designate-mdns', 'designate-sink', 'designate-central', 'designate-producer', 'designate-worker', 'designate-agent'] - cookbook ran by andrew@bullseye * 19:20 wm-bot2: Restarting openstack services on cloudservices2005-dev: ['designate-central', 'designate-sink', 'designate-worker', 'designate-producer', 'designate-mdns', 'designate-agent'] - cookbook ran by andrew@bullseye * 19:20 wm-bot2: Restarting openstack services on cloudnet2005-dev: ['neutron-metadata-agent', 'neutron-dhcp-agent', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 19:20 wm-bot2: Restarting openstack services on cloudnet2006-dev: ['neutron-dhcp-agent', 'neutron-metadata-agent', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 19:20 wm-bot2: Restarting openstack services on cloudbackup1002-dev: ['cinder-backup'] - cookbook ran by andrew@bullseye * 19:20 wm-bot2: Restarting openstack services on cloudbackup1001-dev: ['cinder-backup'] - cookbook ran by andrew@bullseye * 19:19 wm-bot2: Restarting openstack services on cloudcontrol2005-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 19:19 wm-bot2: Restarting openstack services on cloudvirt2003-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 19:19 wm-bot2: Restarting openstack services on cloudvirt2002-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 19:19 wm-bot2: Restarting openstack services on cloudvirt2001-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 19:19 wm-bot2: Restarting openstack services on cloudcontrol2004-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 19:18 wm-bot2: Restarting openstack services on cloudcontrol2001-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 16:12 wm-bot2: Restarting openstack services on cloudservices2004-dev: ['designate-mdns', 'designate-sink', 'designate-central', 'designate-producer', 'designate-worker', 'designate-agent'] - cookbook ran by andrew@bullseye * 16:12 wm-bot2: Restarting openstack services on cloudservices2005-dev: ['designate-central', 'designate-sink', 'designate-worker', 'designate-producer', 'designate-mdns', 'designate-agent'] - cookbook ran by andrew@bullseye * 16:12 wm-bot2: Restarting openstack services on cloudnet2005-dev: ['neutron-metadata-agent', 'neutron-dhcp-agent', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:12 wm-bot2: Restarting openstack services on cloudnet2006-dev: ['neutron-dhcp-agent', 'neutron-metadata-agent', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:12 wm-bot2: Restarting openstack services on cloudbackup1002-dev: ['cinder-backup'] - cookbook ran by andrew@bullseye * 16:12 wm-bot2: Restarting openstack services on cloudbackup1001-dev: ['cinder-backup'] - cookbook ran by andrew@bullseye * 16:11 wm-bot2: Restarting openstack services on cloudcontrol2005-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 16:11 wm-bot2: Restarting openstack services on cloudvirt2003-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:11 wm-bot2: Restarting openstack services on cloudvirt2002-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:11 wm-bot2: Restarting openstack services on cloudvirt2001-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:11 wm-bot2: Restarting openstack services on cloudcontrol2004-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 16:10 wm-bot2: Restarting openstack services on cloudcontrol2001-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 15:22 wm-bot2: Restarting openstack services on cloudcontrol2001-dev@local1: ['cinder-volume'] - cookbook ran by andrew@bullseye * 15:21 wm-bot2: Restarting openstack services on cloudbackup1002-dev: ['cinder-backup'] - cookbook ran by andrew@bullseye * 15:21 wm-bot2: Restarting openstack services on cloudbackup1001-dev: ['cinder-backup'] - cookbook ran by andrew@bullseye * 15:21 wm-bot2: Restarting openstack services on cloudcontrol2005-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 15:21 wm-bot2: Restarting openstack services on cloudvirt2003-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:21 wm-bot2: Restarting openstack services on cloudvirt2002-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:21 wm-bot2: Restarting openstack services on cloudvirt2001-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:21 wm-bot2: Restarting openstack services on cloudcontrol2004-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 15:20 wm-bot2: Restarting openstack services on cloudcontrol2001-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye === 2023-05-15 === * 14:28 wm-bot2: Drained cloudvirt1034.eqiad.wmnet - cookbook ran by andrew@bullseye * 14:28 wm-bot2: Set cloudvirt cloudvirt1034.eqiad.wmnet maintenance (downtime id: 96ce2ed0-3aff-4d04-be0b-{{Gerrit|e16513070617}}, use this to unset) - cookbook ran by andrew@bullseye * 14:27 wm-bot2: Draining cloudvirt1034.eqiad.wmnet - cookbook ran by andrew@bullseye * 14:23 wm-bot2: Drained cloudvirt1027.eqiad.wmnet ([[phab:T316544|T316544]]) - cookbook ran by dcaro@vulcanus * 14:17 wm-bot2: Set cloudvirt cloudvirt1027.eqiad.wmnet maintenance (downtime id: 110176f8-04d5-4110-bb7d-{{Gerrit|1ab272bd8be2}}, use this to unset) ([[phab:T316544|T316544]]) - cookbook ran by dcaro@vulcanus * 14:16 wm-bot2: Draining cloudvirt1027.eqiad.wmnet ([[phab:T316544|T316544]]) - cookbook ran by dcaro@vulcanus * 14:13 wm-bot2: Set cloudvirt cloudvirt1033.eqiad.wmnet maintenance (downtime id: c6e92e13-49f4-4db3-8a13-{{Gerrit|8692ccfd3bc9}}, use this to unset) - cookbook ran by andrew@bullseye * 14:12 wm-bot2: Draining cloudvirt1033.eqiad.wmnet - cookbook ran by andrew@bullseye * 14:12 wm-bot2: Drained cloudvirt1035.eqiad.wmnet ([[phab:T316544|T316544]]) - cookbook ran by dcaro@vulcanus * 14:11 wm-bot2: Set cloudvirt cloudvirt1033.eqiad.wmnet maintenance (downtime id: fa730dec-848f-45fb-9eda-{{Gerrit|e74bd874c5c9}}, use this to unset) - cookbook ran by andrew@bullseye * 14:10 wm-bot2: Draining cloudvirt1033.eqiad.wmnet - cookbook ran by andrew@bullseye * 14:06 wm-bot2: Restarting openstack services on cloudbackup2001: ['cinder-backup'] - cookbook ran by andrew@bullseye * 14:06 wm-bot2: Restarting openstack services on cloudcontrol1006: ['cinder-volume', 'cinder-scheduler'] - cookbook ran by andrew@bullseye * 14:06 wm-bot2: Restarting openstack services on cloudcontrol1007: ['cinder-volume', 'cinder-scheduler'] - cookbook ran by andrew@bullseye * 14:06 wm-bot2: Restarting openstack services on cloudbackup2002: ['cinder-backup'] - cookbook ran by andrew@bullseye * 14:06 wm-bot2: Restarting openstack services on cloudcontrol1005: ['cinder-volume', 'cinder-scheduler'] - cookbook ran by andrew@bullseye * 14:01 wm-bot2: Set cloudvirt cloudvirt1033.eqiad.wmnet maintenance (downtime id: eb1cfac0-d481-4baa-b9cd-{{Gerrit|15e5fbcef495}}, use this to unset) - cookbook ran by andrew@bullseye * 14:00 wm-bot2: Draining cloudvirt1033.eqiad.wmnet - cookbook ran by andrew@bullseye * 13:58 wm-bot2: Set cloudvirt cloudvirt1034.eqiad.wmnet maintenance (downtime id: 3e6c3ff3-7d55-4777-9032-{{Gerrit|b867a257eced}}, use this to unset) - cookbook ran by andrew@bullseye * 13:57 wm-bot2: Draining cloudvirt1034.eqiad.wmnet - cookbook ran by andrew@bullseye * 13:53 wm-bot2: Set cloudvirt cloudvirt1033.eqiad.wmnet maintenance (downtime id: 0693664a-df78-417e-ba34-{{Gerrit|590e5a0a9981}}, use this to unset) - cookbook ran by andrew@bullseye * 13:52 wm-bot2: Draining cloudvirt1033.eqiad.wmnet - cookbook ran by andrew@bullseye * 13:49 wm-bot2: Set cloudvirt cloudvirt1035.eqiad.wmnet maintenance (downtime id: e6929ab8-4bc3-4186-817b-{{Gerrit|9b53dbd597c6}}, use this to unset) ([[phab:T316544|T316544]]) - cookbook ran by dcaro@vulcanus * 13:48 wm-bot2: Draining cloudvirt1035.eqiad.wmnet ([[phab:T316544|T316544]]) - cookbook ran by dcaro@vulcanus * 13:40 wm-bot2: Restarting openstack services on cloudvirtlocal1003: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:40 wm-bot2: Restarting openstack services on cloudvirtlocal1002: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:40 wm-bot2: Restarting openstack services on cloudvirtlocal1001: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:40 wm-bot2: Restarting openstack services on cloudvirt1054: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:40 wm-bot2: Restarting openstack services on cloudvirt1055: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:40 wm-bot2: Restarting openstack services on cloudvirt1060: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:39 wm-bot2: Restarting openstack services on cloudvirt1058: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:39 wm-bot2: Restarting openstack services on cloudvirt1059: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:39 wm-bot2: Restarting openstack services on cloudvirt1061: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:39 wm-bot2: Restarting openstack services on cloudvirt1057: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:39 wm-bot2: Restarting openstack services on cloudvirt1056: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:39 wm-bot2: Restarting openstack services on cloudvirt1051: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:39 wm-bot2: Restarting openstack services on cloudvirt1050: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:39 wm-bot2: Restarting openstack services on cloudvirt1049: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:39 wm-bot2: Restarting openstack services on cloudvirt1053: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:39 wm-bot2: Restarting openstack services on cloudvirt1052: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:39 wm-bot2: Restarting openstack services on cloudvirt1048: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:39 wm-bot2: Restarting openstack services on cloudcontrol1007: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata'] - cookbook ran by andrew@bullseye * 13:38 wm-bot2: Restarting openstack services on cloudcontrol1006: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata'] - cookbook ran by andrew@bullseye * 13:38 wm-bot2: Restarting openstack services on cloudvirt-wdqs1003: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:38 wm-bot2: Restarting openstack services on cloudvirt-wdqs1002: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:38 wm-bot2: Restarting openstack services on cloudvirt-wdqs1001: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:38 wm-bot2: Restarting openstack services on cloudvirt1047: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:38 wm-bot2: Restarting openstack services on cloudvirt1038: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:38 wm-bot2: Restarting openstack services on cloudvirt1042: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:38 wm-bot2: Restarting openstack services on cloudvirt1044: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:38 wm-bot2: Restarting openstack services on cloudvirt1041: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:38 wm-bot2: Restarting openstack services on cloudvirt1046: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:37 wm-bot2: Restarting openstack services on cloudvirt1043: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:37 wm-bot2: Restarting openstack services on cloudvirt1045: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:37 wm-bot2: Restarting openstack services on cloudvirt1040: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:37 wm-bot2: Restarting openstack services on cloudvirt1036: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:37 wm-bot2: Restarting openstack services on cloudvirt1034: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:37 wm-bot2: Restarting openstack services on cloudvirt1039: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:37 wm-bot2: Restarting openstack services on cloudvirt1037: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:37 wm-bot2: Restarting openstack services on cloudvirt1035: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:37 wm-bot2: Restarting openstack services on cloudvirt1033: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:37 wm-bot2: Restarting openstack services on cloudvirt1031: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:37 wm-bot2: Restarting openstack services on cloudvirt1032: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:36 wm-bot2: Restarting openstack services on cloudcontrol1005: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata'] - cookbook ran by andrew@bullseye * 13:36 wm-bot2: Restarting openstack services on cloudvirt1028: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:36 wm-bot2: Restarting openstack services on cloudvirt1030: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:36 andrewbogott: restarting nova services in eqiad1, trying to free up db connections * 13:36 wm-bot2: Restarting openstack services on cloudvirt1027: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:36 wm-bot2: Restarting openstack services on cloudvirt1026: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:36 wm-bot2: Restarting openstack services on cloudvirt1029: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:36 wm-bot2: Restarting openstack services on cloudvirt1025: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:33 wm-bot2: Set cloudvirt cloudvirt1034.eqiad.wmnet maintenance (downtime id: fef43d73-6fd4-4dde-a0ac-{{Gerrit|95fd69a9b0c1}}, use this to unset) ([[phab:T316544|T316544]]) - cookbook ran by dcaro@vulcanus * 13:32 wm-bot2: Draining cloudvirt1034.eqiad.wmnet ([[phab:T316544|T316544]]) - cookbook ran by dcaro@vulcanus * 13:29 wm-bot2: Draining cloudvirt1027.eqiad.wmnet ([[phab:T316544|T316544]]) - cookbook ran by dcaro@vulcanus * 13:24 wm-bot2: Set cloudvirt cloudvirt1027.eqiad.wmnet maintenance (downtime id: 5f867662-e824-498c-a715-{{Gerrit|2e2ad50f0bb5}}, use this to unset) ([[phab:T316544|T316544]]) - cookbook ran by dcaro@vulcanus * 13:23 wm-bot2: Draining cloudvirt1027.eqiad.wmnet ([[phab:T316544|T316544]]) - cookbook ran by dcaro@vulcanus * 13:08 wm-bot2: Set cloudvirt cloudvirt1033.eqiad.wmnet maintenance (downtime id: 0f515150-1313-41d4-a5f6-{{Gerrit|9bc00ce9b245}}, use this to unset) ([[phab:T316544|T316544]]) - cookbook ran by dcaro@vulcanus * 13:07 wm-bot2: Draining cloudvirt1033.eqiad.wmnet ([[phab:T316544|T316544]]) - cookbook ran by dcaro@vulcanus * 13:07 wm-bot2: Drained cloudvirt1032.eqiad.wmnet ([[phab:T316544|T316544]]) - cookbook ran by dcaro@vulcanus * 12:50 wm-bot2: Set cloudvirt cloudvirt1032.eqiad.wmnet maintenance (downtime id: e826adc3-addd-44d8-b39e-{{Gerrit|ae7bd2df1e60}}, use this to unset) ([[phab:T316544|T316544]]) - cookbook ran by dcaro@vulcanus * 12:49 wm-bot2: Draining cloudvirt1032.eqiad.wmnet ([[phab:T316544|T316544]]) - cookbook ran by dcaro@vulcanus * 12:49 wm-bot2: Drained cloudvirt1031.eqiad.wmnet ([[phab:T316544|T316544]]) - cookbook ran by dcaro@vulcanus * 12:14 wm-bot2: Set cloudvirt cloudvirt1027.eqiad.wmnet maintenance (downtime id: 4154c818-744c-4d84-9883-{{Gerrit|cae7a5826ed5}}, use this to unset) ([[phab:T316544|T316544]]) - cookbook ran by dcaro@vulcanus * 12:13 wm-bot2: Draining cloudvirt1027.eqiad.wmnet ([[phab:T316544|T316544]]) - cookbook ran by dcaro@vulcanus * 12:13 wm-bot2: Drained cloudvirt1026.eqiad.wmnet ([[phab:T316544|T316544]]) - cookbook ran by dcaro@vulcanus * 12:04 wm-bot2: Set cloudvirt cloudvirt1026.eqiad.wmnet maintenance (downtime id: cdfc3d01-ec1e-483c-9a38-{{Gerrit|834193e487ff}}, use this to unset) ([[phab:T316544|T316544]]) - cookbook ran by dcaro@vulcanus * 12:03 wm-bot2: Draining cloudvirt1026.eqiad.wmnet ([[phab:T316544|T316544]]) - cookbook ran by dcaro@vulcanus * 12:03 wm-bot2: Drained cloudvirt1025.eqiad.wmnet ([[phab:T316544|T316544]]) - cookbook ran by dcaro@vulcanus * 11:51 wm-bot2: Set cloudvirt cloudvirt1025.eqiad.wmnet maintenance (downtime id: 6a56757d-35de-499e-8209-{{Gerrit|728bcf62a22a}}, use this to unset) ([[phab:T316544|T316544]]) - cookbook ran by dcaro@vulcanus * 11:50 wm-bot2: Draining cloudvirt1025.eqiad.wmnet ([[phab:T316544|T316544]]) - cookbook ran by dcaro@vulcanus === 2023-05-12 === * 17:52 wm-bot2: Restarting openstack services on cloudcontrol2001-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye === 2023-05-10 === * 16:09 wm-bot2: Restarting openstack services on cloudservices1005: ['designate-producer', 'designate-sink', 'designate-worker', 'designate-central', 'designate-mdns', 'designate-agent'] - cookbook ran by andrew@bullseye * 16:09 wm-bot2: Restarting openstack services on cloudservices1004: ['designate-worker', 'designate-api', 'designate-mdns', 'designate-producer', 'designate-central', 'designate-sink', 'designate-agent'] - cookbook ran by andrew@bullseye * 16:09 wm-bot2: Restarting openstack services on cloudnet1006: ['neutron-linuxbridge-agent', 'neutron-metadata-agent', 'neutron-dhcp-agent'] - cookbook ran by andrew@bullseye * 16:09 wm-bot2: Restarting openstack services on cloudnet1005: ['neutron-linuxbridge-agent', 'neutron-dhcp-agent', 'neutron-metadata-agent'] - cookbook ran by andrew@bullseye * 16:09 wm-bot2: Restarting openstack services on cloudbackup2001: ['cinder-backup'] - cookbook ran by andrew@bullseye * 16:09 wm-bot2: Restarting openstack services on cloudbackup2002: ['cinder-backup'] - cookbook ran by andrew@bullseye * 16:08 wm-bot2: Restarting openstack services on cloudvirtlocal1003: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:08 wm-bot2: Restarting openstack services on cloudvirtlocal1002: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:08 wm-bot2: Restarting openstack services on cloudvirtlocal1001: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:08 wm-bot2: Restarting openstack services on cloudvirt1054: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:08 wm-bot2: Restarting openstack services on cloudvirt1055: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:08 wm-bot2: Restarting openstack services on cloudvirt1060: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:08 wm-bot2: Restarting openstack services on cloudvirt1058: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:08 wm-bot2: Restarting openstack services on cloudvirt1059: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:08 wm-bot2: Restarting openstack services on cloudvirt1061: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:08 wm-bot2: Restarting openstack services on cloudvirt1057: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:08 wm-bot2: Restarting openstack services on cloudvirt1056: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:08 wm-bot2: Restarting openstack services on cloudvirt1051: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:08 wm-bot2: Restarting openstack services on cloudvirt1050: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:07 wm-bot2: Restarting openstack services on cloudvirt1049: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:07 wm-bot2: Restarting openstack services on cloudvirt1053: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:07 wm-bot2: Restarting openstack services on cloudvirt1052: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:07 wm-bot2: Restarting openstack services on cloudvirt1048: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:07 wm-bot2: Restarting openstack services on cloudcontrol1007: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 16:07 wm-bot2: Restarting openstack services on cloudcontrol1006: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 16:06 wm-bot2: Restarting openstack services on cloudvirt-wdqs1003: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:06 wm-bot2: Restarting openstack services on cloudvirt-wdqs1002: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:06 wm-bot2: Restarting openstack services on cloudvirt-wdqs1001: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:06 wm-bot2: Restarting openstack services on cloudvirt1047: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:06 wm-bot2: Restarting openstack services on cloudvirt1038: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:06 wm-bot2: Restarting openstack services on cloudvirt1042: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:06 wm-bot2: Restarting openstack services on cloudvirt1044: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:06 wm-bot2: Restarting openstack services on cloudvirt1041: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:06 wm-bot2: Restarting openstack services on cloudvirt1046: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:06 wm-bot2: Restarting openstack services on cloudvirt1043: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:06 wm-bot2: Restarting openstack services on cloudvirt1045: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:06 wm-bot2: Restarting openstack services on cloudvirt1040: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:06 wm-bot2: Restarting openstack services on cloudvirt1036: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:05 wm-bot2: Restarting openstack services on cloudvirt1034: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:05 wm-bot2: Restarting openstack services on cloudvirt1039: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:05 wm-bot2: Restarting openstack services on cloudvirt1037: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:05 wm-bot2: Restarting openstack services on cloudvirt1035: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:05 wm-bot2: Restarting openstack services on cloudvirt1033: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:05 wm-bot2: Restarting openstack services on cloudvirt1031: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:05 wm-bot2: Restarting openstack services on cloudvirt1032: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:05 wm-bot2: Restarting openstack services on cloudcontrol1005: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 16:05 wm-bot2: Restarting openstack services on cloudvirt1028: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:05 wm-bot2: Restarting openstack services on cloudvirt1030: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:05 wm-bot2: Restarting openstack services on cloudvirt1027: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:04 wm-bot2: Restarting openstack services on cloudvirt1026: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:04 wm-bot2: Restarting openstack services on cloudvirt1029: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:04 wm-bot2: Restarting openstack services on cloudvirt1025: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:01 wm-bot2: Restarting openstack services on cloudvirt1024: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:01 wm-bot2: Restarting openstack services on cloudnet1006: ['neutron-linuxbridge-agent', 'neutron-metadata-agent', 'neutron-dhcp-agent'] - cookbook ran by andrew@bullseye * 16:01 wm-bot2: Restarting openstack services on cloudnet1005: ['neutron-linuxbridge-agent', 'neutron-dhcp-agent', 'neutron-metadata-agent'] - cookbook ran by andrew@bullseye * 16:01 wm-bot2: Restarting openstack services on cloudbackup2001: ['cinder-backup'] - cookbook ran by andrew@bullseye * 16:01 wm-bot2: Restarting openstack services on cloudbackup2002: ['cinder-backup'] - cookbook ran by andrew@bullseye * 16:01 wm-bot2: Restarting openstack services on cloudvirtlocal1003: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:01 wm-bot2: Restarting openstack services on cloudvirtlocal1002: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:01 wm-bot2: Restarting openstack services on cloudvirtlocal1001: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:01 wm-bot2: Restarting openstack services on cloudvirt1054: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:01 wm-bot2: Restarting openstack services on cloudvirt1055: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:01 wm-bot2: Restarting openstack services on cloudvirt1060: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:01 wm-bot2: Restarting openstack services on cloudvirt1058: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:01 wm-bot2: Restarting openstack services on cloudvirt1059: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:01 wm-bot2: Restarting openstack services on cloudvirt1061: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:01 wm-bot2: Restarting openstack services on cloudvirt1057: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:00 wm-bot2: Restarting openstack services on cloudvirt1056: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:00 wm-bot2: Restarting openstack services on cloudvirt1051: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:00 wm-bot2: Restarting openstack services on cloudvirt1050: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:00 wm-bot2: Restarting openstack services on cloudvirt1049: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:00 wm-bot2: Restarting openstack services on cloudvirt1053: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:00 wm-bot2: Restarting openstack services on cloudvirt1052: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:00 wm-bot2: Restarting openstack services on cloudvirt1048: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:00 wm-bot2: Restarting openstack services on cloudcontrol1007: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 16:00 wm-bot2: Restarting openstack services on cloudcontrol1006: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 16:00 wm-bot2: Restarting openstack services on cloudvirt-wdqs1003: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:59 wm-bot2: Restarting openstack services on cloudvirt-wdqs1002: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:59 wm-bot2: Restarting openstack services on cloudvirt-wdqs1001: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:59 wm-bot2: Restarting openstack services on cloudvirt1047: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:59 wm-bot2: Restarting openstack services on cloudvirt1038: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:59 wm-bot2: Restarting openstack services on cloudvirt1042: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:59 wm-bot2: Restarting openstack services on cloudvirt1044: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:59 wm-bot2: Restarting openstack services on cloudvirt1041: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:59 wm-bot2: Restarting openstack services on cloudvirt1046: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:59 wm-bot2: Restarting openstack services on cloudvirt1043: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:59 wm-bot2: Restarting openstack services on cloudvirt1045: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:59 wm-bot2: Restarting openstack services on cloudvirt1040: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:59 wm-bot2: Restarting openstack services on cloudvirt1036: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:58 wm-bot2: Restarting openstack services on cloudvirt1034: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:58 wm-bot2: Restarting openstack services on cloudvirt1039: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:58 wm-bot2: Restarting openstack services on cloudvirt1037: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:58 wm-bot2: Restarting openstack services on cloudvirt1035: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:58 wm-bot2: Restarting openstack services on cloudvirt1033: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:58 wm-bot2: Restarting openstack services on cloudvirt1031: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:58 wm-bot2: Restarting openstack services on cloudvirt1032: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:58 wm-bot2: Restarting openstack services on cloudcontrol1005: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 15:58 wm-bot2: Restarting openstack services on cloudvirt1028: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:57 wm-bot2: Restarting openstack services on cloudvirt1030: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:57 wm-bot2: Restarting openstack services on cloudvirt1027: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:57 wm-bot2: Restarting openstack services on cloudvirt1026: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:57 wm-bot2: Restarting openstack services on cloudvirt1029: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:57 wm-bot2: Restarting openstack services on cloudvirt1025: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:57 wm-bot2: Restarting openstack services on cloudvirt1024: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:57 wm-bot2: Restarting openstack services on cloudnet1006: ['neutron-linuxbridge-agent', 'neutron-metadata-agent', 'neutron-dhcp-agent'] - cookbook ran by andrew@bullseye * 15:56 wm-bot2: Restarting openstack services on cloudnet1005: ['neutron-linuxbridge-agent', 'neutron-dhcp-agent', 'neutron-metadata-agent'] - cookbook ran by andrew@bullseye * 15:56 wm-bot2: Restarting openstack services on cloudbackup2001: ['cinder-backup'] - cookbook ran by andrew@bullseye * 15:56 wm-bot2: Restarting openstack services on cloudbackup2002: ['cinder-backup'] - cookbook ran by andrew@bullseye * 15:56 wm-bot2: Restarting openstack services on cloudvirtlocal1003: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:56 wm-bot2: Restarting openstack services on cloudvirtlocal1002: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:56 wm-bot2: Restarting openstack services on cloudvirtlocal1001: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:56 wm-bot2: Restarting openstack services on cloudvirt1054: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:56 wm-bot2: Restarting openstack services on cloudvirt1055: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:56 wm-bot2: Restarting openstack services on cloudvirt1060: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:56 wm-bot2: Restarting openstack services on cloudvirt1058: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:56 wm-bot2: Restarting openstack services on cloudvirt1059: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:56 wm-bot2: Restarting openstack services on cloudvirt1061: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:56 wm-bot2: Restarting openstack services on cloudvirt1057: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:56 wm-bot2: Restarting openstack services on cloudvirt1056: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:55 wm-bot2: Restarting openstack services on cloudvirt1051: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:55 wm-bot2: Restarting openstack services on cloudvirt1050: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:55 wm-bot2: Restarting openstack services on cloudvirt1049: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:55 wm-bot2: Restarting openstack services on cloudvirt1053: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:55 wm-bot2: Restarting openstack services on cloudvirt1052: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:55 wm-bot2: Restarting openstack services on cloudvirt1048: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:55 wm-bot2: Restarting openstack services on cloudcontrol1007: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 15:54 wm-bot2: Restarting openstack services on cloudcontrol1006: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 15:54 wm-bot2: Restarting openstack services on cloudvirt-wdqs1003: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:54 wm-bot2: Restarting openstack services on cloudvirt-wdqs1002: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:54 wm-bot2: Restarting openstack services on cloudvirt-wdqs1001: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:54 wm-bot2: Restarting openstack services on cloudvirt1047: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:54 wm-bot2: Restarting openstack services on cloudvirt1038: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:54 wm-bot2: Restarting openstack services on cloudvirt1042: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:54 wm-bot2: Restarting openstack services on cloudvirt1044: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:54 wm-bot2: Restarting openstack services on cloudvirt1041: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:54 wm-bot2: Restarting openstack services on cloudvirt1046: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:54 wm-bot2: Restarting openstack services on cloudvirt1043: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:53 wm-bot2: Restarting openstack services on cloudvirt1045: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:53 wm-bot2: Restarting openstack services on cloudvirt1040: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:53 wm-bot2: Restarting openstack services on cloudvirt1036: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:53 wm-bot2: Restarting openstack services on cloudvirt1034: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:53 wm-bot2: Restarting openstack services on cloudvirt1039: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:53 wm-bot2: Restarting openstack services on cloudvirt1037: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:53 wm-bot2: Restarting openstack services on cloudvirt1035: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:53 wm-bot2: Restarting openstack services on cloudvirt1033: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:53 wm-bot2: Restarting openstack services on cloudvirt1031: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:53 wm-bot2: Restarting openstack services on cloudvirt1032: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:52 wm-bot2: Restarting openstack services on cloudcontrol1005: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 15:52 wm-bot2: Restarting openstack services on cloudvirt1028: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:52 wm-bot2: Restarting openstack services on cloudvirt1030: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:52 andrewbogott: running "cookbook -c ~/.config/spicerack/cookbook_config.yaml wmcs.openstack.restart_openstack --cluster-name eqiad1 --all" to pick up changes for testing [[phab:T336379|T336379]] * 15:52 wm-bot2: Restarting openstack services on cloudvirt1027: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:52 wm-bot2: Restarting openstack services on cloudvirt1026: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:52 wm-bot2: Restarting openstack services on cloudvirt1029: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:52 wm-bot2: Restarting openstack services on cloudvirt1025: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 03:05 andrewbogott: "systemctl restart puppet-enc" on enc-1.cloudinfra.eqiad1.wikimedia.cloud. Seems to have crashed. === 2023-05-05 === * 16:07 wm-bot2: Drained cloudvirt1024.eqiad.wmnet ([[phab:T336064|T336064]]) - cookbook ran by andrew@bullseye * 16:03 wm-bot2: Set cloudvirt cloudvirt1024.eqiad.wmnet maintenance (downtime id: 95995009-09d6-496e-8cd2-{{Gerrit|0cfac93d3cf7}}, use this to unset) ([[phab:T336064|T336064]]) - cookbook ran by andrew@bullseye * 16:02 wm-bot2: Draining cloudvirt1024.eqiad.wmnet ([[phab:T336064|T336064]]) - cookbook ran by andrew@bullseye * 16:01 wm-bot2: Drained cloudvirt1023.eqiad.wmnet ([[phab:T336064|T336064]]) - cookbook ran by andrew@bullseye * 15:51 wm-bot2: Set cloudvirt cloudvirt1023.eqiad.wmnet maintenance (downtime id: 53c46cae-00af-4664-97ff-{{Gerrit|266b393335bb}}, use this to unset) ([[phab:T336064|T336064]]) - cookbook ran by andrew@bullseye * 15:50 wm-bot2: Draining cloudvirt1023.eqiad.wmnet ([[phab:T336064|T336064]]) - cookbook ran by andrew@bullseye * 15:49 wm-bot2: Set cloudvirt cloudvirt1024.eqiad.wmnet maintenance (downtime id: 528ea4f6-8088-475e-937f-{{Gerrit|098ffba861b6}}, use this to unset) - cookbook ran by andrew@bullseye * 15:47 wm-bot2: Set cloudvirt cloudvirt1023.eqiad.wmnet maintenance (downtime id: 3ef85b5e-d9d9-4b24-901b-{{Gerrit|a3058a7d0615}}, use this to unset) - cookbook ran by andrew@bullseye * 15:44 andrewbogott: moved cloudvirt1023 and cloudvirt1024 from 'ceph' aggregate to 'maintenance' aggregate, prep for decom [[phab:T336064|T336064]] * 15:44 andrewbogott: moved cloudvirt1028 from 'localdisk' aggregate to 'maintenance' aggregate. Nothing new should be scheduled here, local storage should now move to cloudvirtlocal100x * 15:41 andrewbogott: moved cloudvirt1055 and cloudvirt1056 from 'spare' to 'ceph' aggregate. Prep for removing two obsolete cloudvirts, 1023 and 1024. [[phab:T336064|T336064]] === 2023-05-04 === * 22:49 andrewbogott: removed fullstack-* puppet reports on puppetmaster-02.cloudinfra-codfw1dev.codfw1dev.wikimedia.cloud and cloud-puppetmaster-03.cloudinfra.eqiad.wmflabs to free up disk space === 2023-05-02 === * 13:01 wm-bot2: Adding OSD cloudcephosd2001-dev.codfw.wmnet... (1/1) - cookbook ran by dcaro@vulcanus * 13:01 wm-bot2: Adding new OSDs ['cloudcephosd2001-dev.codfw.wmnet'] to the cluster - cookbook ran by dcaro@vulcanus * 12:31 wm-bot2: Destroying OSDs with ids in [0] on cloudcephosd2001-dev from codfw1 - cookbook ran by dcaro@vulcanus * 12:30 wm-bot2: Depooling OSDs with ids in [0] on cloudcephosd2001-dev from codfw1 - cookbook ran by dcaro@vulcanus * 11:53 wm-bot2: The cluster is now rebalanced after adding the new OSDs ['cloudcephosd2001-dev.codfw.wmnet'] - cookbook ran by dcaro@vulcanus * 11:53 wm-bot2: Added 1 new OSDs ['cloudcephosd2001-dev.codfw.wmnet'] - cookbook ran by dcaro@vulcanus * 11:53 wm-bot2: Added OSD cloudcephosd2001-dev.codfw.wmnet... (1/1) - cookbook ran by dcaro@vulcanus * 11:53 wm-bot2: Adding OSD cloudcephosd2001-dev.codfw.wmnet... (1/1) - cookbook ran by dcaro@vulcanus * 11:52 wm-bot2: Adding new OSDs ['cloudcephosd2001-dev.codfw.wmnet'] to the cluster - cookbook ran by dcaro@vulcanus === 2023-05-01 === * 17:09 wm-bot2: Adding OSD cloudcephosd2001-dev.codfw.wmnet... (1/1) - cookbook ran by dcaro@vulcanus * 17:09 wm-bot2: Adding new OSDs ['cloudcephosd2001-dev.codfw.wmnet'] to the cluster - cookbook ran by dcaro@vulcanus * 17:08 wm-bot2: Adding OSD cloudcephosd2001-dev.codfw.wmnet... (1/1) - cookbook ran by dcaro@vulcanus * 17:08 wm-bot2: Adding new OSDs ['cloudcephosd2001-dev.codfw.wmnet'] to the cluster - cookbook ran by dcaro@vulcanus * 15:22 wm-bot2: Depooling OSDs with ids in [0] on cloudcephosd2001-dev from codfw1 - cookbook ran by dcaro@vulcanus * 13:53 taavi: running wmcs-novastats-puppetleaks in real mode [[phab:T334127|T334127]] * 13:51 wm-bot2: Depooling OSDs with ids in [0] on cloudcephosd2001-dev from codfw1 - cookbook ran by dcaro@vulcanus === 2023-04-18 === * 22:52 wm-bot2: Restarting openstack services on cloudcontrol1007: ['neutron-api', 'neutron-rpc-server'] - cookbook ran by andrew@bullseye * 22:52 wm-bot2: Restarting openstack services on cloudcontrol1006: ['neutron-api', 'neutron-rpc-server'] - cookbook ran by andrew@bullseye * 22:52 wm-bot2: Restarting openstack services on cloudcontrol1005: ['neutron-api', 'neutron-rpc-server'] - cookbook ran by andrew@bullseye * 22:52 wm-bot2: Restarting openstack services on cloudvirt1056: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:51 wm-bot2: Restarting openstack services on cloudvirt1019: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:51 wm-bot2: Restarting openstack services on cloudvirt1060: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:51 wm-bot2: Restarting openstack services on cloudvirt1023: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:51 wm-bot2: Restarting openstack services on cloudvirt1038: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:51 wm-bot2: Restarting openstack services on cloudvirt1036: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:51 wm-bot2: Restarting openstack services on cloudvirt1039: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:51 wm-bot2: Restarting openstack services on cloudvirt1050: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:51 wm-bot2: Restarting openstack services on cloudvirt-wdqs1003: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:51 wm-bot2: Restarting openstack services on cloudvirt1061: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:51 wm-bot2: Restarting openstack services on cloudvirt1028: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:51 wm-bot2: Restarting openstack services on cloudvirt1020: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:51 wm-bot2: Restarting openstack services on cloudvirt1052: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:51 wm-bot2: Restarting openstack services on cloudvirt1040: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:51 wm-bot2: Restarting openstack services on cloudvirt1043: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:51 wm-bot2: Restarting openstack services on cloudvirt1026: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:51 wm-bot2: Restarting openstack services on cloudvirtlocal1003: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:51 wm-bot2: Restarting openstack services on cloudvirt-wdqs1001: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:51 wm-bot2: Restarting openstack services on cloudvirt1030: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:51 wm-bot2: Restarting openstack services on cloudvirt1048: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:51 wm-bot2: Restarting openstack services on cloudvirt1025: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:51 wm-bot2: Restarting openstack services on cloudvirt1035: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:50 wm-bot2: Restarting openstack services on cloudvirt1055: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:50 wm-bot2: Restarting openstack services on cloudvirt-wdqs1002: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:50 wm-bot2: Restarting openstack services on cloudvirt1042: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:50 wm-bot2: Restarting openstack services on cloudvirt1059: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:50 wm-bot2: Restarting openstack services on cloudvirt1057: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:50 wm-bot2: Restarting openstack services on cloudvirt1032: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:50 wm-bot2: Restarting openstack services on cloudvirt1024: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:50 wm-bot2: Restarting openstack services on cloudvirt1029: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:50 wm-bot2: Restarting openstack services on cloudvirt1044: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:50 wm-bot2: Restarting openstack services on cloudvirt1047: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:50 wm-bot2: Restarting openstack services on cloudvirt1034: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:50 wm-bot2: Restarting openstack services on cloudvirt1058: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:50 wm-bot2: Restarting openstack services on cloudvirt1045: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:50 wm-bot2: Restarting openstack services on cloudvirt1041: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:50 wm-bot2: Restarting openstack services on cloudvirtlocal1001: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:50 wm-bot2: Restarting openstack services on cloudvirt1031: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:50 wm-bot2: Restarting openstack services on cloudvirt1046: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:50 wm-bot2: Restarting openstack services on cloudvirt1027: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:50 wm-bot2: Restarting openstack services on cloudnet1006: ['neutron-linuxbridge-agent', 'neutron-metadata-agent', 'neutron-dhcp-agent'] - cookbook ran by andrew@bullseye * 22:50 wm-bot2: Restarting openstack services on cloudnet1005: ['neutron-linuxbridge-agent', 'neutron-dhcp-agent', 'neutron-metadata-agent'] - cookbook ran by andrew@bullseye * 22:50 wm-bot2: Restarting openstack services on cloudvirt1054: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:49 wm-bot2: Restarting openstack services on cloudvirt1033: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:49 wm-bot2: Restarting openstack services on cloudvirt1037: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:49 wm-bot2: Restarting openstack services on cloudvirtlocal1002: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:49 wm-bot2: Restarting openstack services on cloudvirt1051: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:49 wm-bot2: Restarting openstack services on cloudvirt1053: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:49 wm-bot2: Restarting openstack services on cloudvirt1049: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:48 wm-bot2: Restarting openstack services on cloudvirtlocal1003: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:48 wm-bot2: Restarting openstack services on cloudvirtlocal1002: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:48 wm-bot2: Restarting openstack services on cloudvirtlocal1001: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:48 wm-bot2: Restarting openstack services on cloudvirt1054: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:48 wm-bot2: Restarting openstack services on cloudvirt1055: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:48 wm-bot2: Restarting openstack services on cloudvirt1060: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:47 wm-bot2: Restarting openstack services on cloudvirt1058: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:47 wm-bot2: Restarting openstack services on cloudvirt1059: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:47 wm-bot2: Restarting openstack services on cloudvirt1061: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:47 wm-bot2: Restarting openstack services on cloudvirt1057: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:47 wm-bot2: Restarting openstack services on cloudvirt1056: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:47 wm-bot2: Restarting openstack services on cloudvirt1051: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:47 wm-bot2: Restarting openstack services on cloudvirt1050: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:47 wm-bot2: Restarting openstack services on cloudvirt1049: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:47 wm-bot2: Restarting openstack services on cloudvirt1053: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:47 wm-bot2: Restarting openstack services on cloudvirt1052: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:47 wm-bot2: Restarting openstack services on cloudvirt1048: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:47 wm-bot2: Restarting openstack services on cloudcontrol1007: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata'] - cookbook ran by andrew@bullseye * 22:47 wm-bot2: Restarting openstack services on cloudcontrol1006: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata'] - cookbook ran by andrew@bullseye * 22:47 wm-bot2: Restarting openstack services on cloudvirt-wdqs1003: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:46 wm-bot2: Restarting openstack services on cloudvirt-wdqs1002: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:46 wm-bot2: Restarting openstack services on cloudvirt-wdqs1001: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:46 wm-bot2: Restarting openstack services on cloudvirt1047: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:46 wm-bot2: Restarting openstack services on cloudvirt1038: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:46 wm-bot2: Restarting openstack services on cloudvirt1042: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:46 wm-bot2: Restarting openstack services on cloudvirt1044: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:46 wm-bot2: Restarting openstack services on cloudvirt1041: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:46 wm-bot2: Restarting openstack services on cloudvirt1046: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:46 wm-bot2: Restarting openstack services on cloudvirt1043: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:46 wm-bot2: Restarting openstack services on cloudvirt1045: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:46 wm-bot2: Restarting openstack services on cloudvirt1040: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:46 wm-bot2: Restarting openstack services on cloudvirt1036: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:46 wm-bot2: Restarting openstack services on cloudvirt1034: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:45 wm-bot2: Restarting openstack services on cloudvirt1039: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:45 wm-bot2: Restarting openstack services on cloudvirt1037: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:45 wm-bot2: Restarting openstack services on cloudvirt1035: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:45 wm-bot2: Restarting openstack services on cloudvirt1033: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:45 wm-bot2: Restarting openstack services on cloudvirt1031: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:45 wm-bot2: Restarting openstack services on cloudvirt1032: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:45 wm-bot2: Restarting openstack services on cloudcontrol1005: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata'] - cookbook ran by andrew@bullseye * 22:45 wm-bot2: Restarting openstack services on cloudvirt1028: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:45 wm-bot2: Restarting openstack services on cloudvirt1019: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:45 wm-bot2: Restarting openstack services on cloudvirt1020: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:45 wm-bot2: Restarting openstack services on cloudvirt1030: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:45 wm-bot2: Restarting openstack services on cloudvirt1027: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:44 wm-bot2: Restarting openstack services on cloudvirt1023: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:44 wm-bot2: Restarting openstack services on cloudvirt1026: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:44 wm-bot2: Restarting openstack services on cloudvirt1029: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:44 wm-bot2: Restarting openstack services on cloudvirt1025: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:44 wm-bot2: Restarting openstack services on cloudvirt1024: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:41 andrewbogott: resetting rabbitmq on cloudrabbit1003 due to splitbrain * 22:40 wm-bot2: Restarting openstack services on cloudvirt1030: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:40 wm-bot2: Restarting openstack services on cloudvirt1027: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:40 wm-bot2: Restarting openstack services on cloudvirt1023: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:39 wm-bot2: Restarting openstack services on cloudvirt1026: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:39 wm-bot2: Restarting openstack services on cloudvirt1029: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:39 wm-bot2: Restarting openstack services on cloudvirt1025: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:39 wm-bot2: Restarting openstack services on cloudvirt1024: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:36 wm-bot2: Restarting openstack services on cloudvirt1029: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:36 wm-bot2: Restarting openstack services on cloudvirt1025: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:36 wm-bot2: Restarting openstack services on cloudvirt1024: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:34 wm-bot2: Restarting openstack services on cloudvirt1053: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:34 wm-bot2: Restarting openstack services on cloudvirt1052: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:34 wm-bot2: Restarting openstack services on cloudvirt1048: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:33 wm-bot2: Restarting openstack services on cloudcontrol1007: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata'] - cookbook ran by andrew@bullseye * 22:33 wm-bot2: Restarting openstack services on cloudcontrol1006: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata'] - cookbook ran by andrew@bullseye * 22:33 wm-bot2: Restarting openstack services on cloudvirt-wdqs1003: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:33 wm-bot2: Restarting openstack services on cloudvirt-wdqs1002: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:33 wm-bot2: Restarting openstack services on cloudvirt-wdqs1001: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:33 wm-bot2: Restarting openstack services on cloudvirt1047: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:33 wm-bot2: Restarting openstack services on cloudvirt1038: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:32 wm-bot2: Restarting openstack services on cloudvirt1042: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:32 wm-bot2: Restarting openstack services on cloudvirt1044: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:32 wm-bot2: Restarting openstack services on cloudvirt1041: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:31 wm-bot2: Restarting openstack services on cloudvirt1046: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:31 wm-bot2: Restarting openstack services on cloudvirt1043: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:31 wm-bot2: Restarting openstack services on cloudvirt1045: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:31 wm-bot2: Restarting openstack services on cloudvirt1040: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:30 wm-bot2: Restarting openstack services on cloudvirt1036: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:30 wm-bot2: Restarting openstack services on cloudvirt1034: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:30 wm-bot2: Restarting openstack services on cloudvirt1039: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:30 wm-bot2: Restarting openstack services on cloudvirt1037: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:30 wm-bot2: Restarting openstack services on cloudvirt1035: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:30 wm-bot2: Restarting openstack services on cloudvirt1033: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:29 wm-bot2: Restarting openstack services on cloudvirt1031: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:29 wm-bot2: Restarting openstack services on cloudvirt1032: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:29 wm-bot2: Restarting openstack services on cloudcontrol1005: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata'] - cookbook ran by andrew@bullseye * 22:29 wm-bot2: Restarting openstack services on cloudvirt1028: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:29 wm-bot2: Restarting openstack services on cloudvirt1019: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:29 wm-bot2: Restarting openstack services on cloudvirt1020: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:29 wm-bot2: Restarting openstack services on cloudvirt1030: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:29 wm-bot2: Restarting openstack services on cloudvirt1027: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:29 wm-bot2: Restarting openstack services on cloudvirt1023: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:28 wm-bot2: Restarting openstack services on cloudvirt1026: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:28 wm-bot2: Restarting openstack services on cloudvirt1029: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:28 wm-bot2: Restarting openstack services on cloudvirt1025: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:28 wm-bot2: Restarting openstack services on cloudvirt1024: ['nova-compute'] - cookbook ran by andrew@bullseye === 2023-04-17 === * 08:49 wm-bot2: Increased quotas by 9 cores, 1 instances, 16 ram ([[phab:T334695|T334695]]) - cookbook ran by dcaro@vulcanus === 2023-04-06 === * 17:03 andrewbogott: running wmcs-wikireplica-dns on cloudcontrol1005 to update tools-db dns entries === 2023-04-04 === * 17:23 andrewbogott: resetting all three rabbitmq nodes and restarting all openstack services as per https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Rabbitmq#Resetting_the_HA_setup === 2023-03-28 === * 13:56 andrewbogott: depooling cloudweb1003 before switch upgrade * 10:58 dhinus: disabled tool "wb" by clicking the disable button at https://toolsadmin.wikimedia.org/tools/id/wb [[phab:T328693|T328693]] * 08:34 arturo: cleanup neutron agents for cloudvirt1021/1022 (decom) * 08:32 arturo: cleanup neutron agents for cloudvirt1017 (decom) === 2023-03-27 === * 14:30 wm-bot2: Drained cloudvirt1024.eqiad.wmnet - cookbook ran by andrew@bullseye * 14:19 wm-bot2: Set cloudvirt cloudvirt1024.eqiad.wmnet maintenance (downtime id: 3f43d3ca-696c-4d3c-8d5b-{{Gerrit|e57984f0eb86}}, use this to unset) - cookbook ran by andrew@bullseye * 14:18 wm-bot2: Draining cloudvirt1024.eqiad.wmnet - cookbook ran by andrew@bullseye * 14:17 wm-bot2: Drained cloudvirt1023.eqiad.wmnet - cookbook ran by andrew@bullseye * 14:08 wm-bot2: Set cloudvirt cloudvirt1023.eqiad.wmnet maintenance (downtime id: f4767781-ea26-453b-9521-{{Gerrit|847a8340a249}}, use this to unset) - cookbook ran by andrew@bullseye * 14:07 wm-bot2: Draining cloudvirt1023.eqiad.wmnet - cookbook ran by andrew@bullseye * 14:04 wm-bot2: Drained cloudvirt1022.eqiad.wmnet - cookbook ran by andrew@bullseye * 13:55 wm-bot2: Set cloudvirt cloudvirt1022.eqiad.wmnet maintenance (downtime id: 0e794cfe-5896-46c1-842d-{{Gerrit|c34719140d4f}}, use this to unset) - cookbook ran by andrew@bullseye * 13:54 wm-bot2: Draining cloudvirt1022.eqiad.wmnet - cookbook ran by andrew@bullseye * 13:54 wm-bot2: Drained cloudvirt1021.eqiad.wmnet - cookbook ran by andrew@bullseye * 13:48 wm-bot2: Set cloudvirt cloudvirt1021.eqiad.wmnet maintenance (downtime id: e7a904ca-003d-450a-ad85-{{Gerrit|886fb80dfc41}}, use this to unset) - cookbook ran by andrew@bullseye * 13:47 wm-bot2: Draining cloudvirt1021.eqiad.wmnet - cookbook ran by andrew@bullseye * 13:46 wm-bot2: Drained cloudvirt1017.eqiad.wmnet - cookbook ran by andrew@bullseye * 13:36 wm-bot2: Set cloudvirt cloudvirt1017.eqiad.wmnet maintenance (downtime id: 0ff0090f-26c3-4278-a9c9-{{Gerrit|cce518559408}}, use this to unset) - cookbook ran by andrew@bullseye * 13:35 wm-bot2: Draining cloudvirt1017.eqiad.wmnet - cookbook ran by andrew@bullseye === 2023-03-22 === * 12:41 taavi: delete wmde-templates-alpha project [[phab:T332773|T332773]] === 2023-03-08 === * 21:45 bd808: maintain-kubeusers container in CrashLoopBackoff, investigating * 13:49 dcaro: stopping puppet on labostre1004 to debug maintain-dbusers === 2023-03-07 === * 16:06 andrewbogott: updated application credential roles, replacing 'user' with 'reader' and 'projectadmin' with 'member': update application_credential_role set role_id='f75a3c410bca4e96a1cf6ac103b0ccaf' where role_id='f473273fac7146b3bdbf22e5d4504f95' and update application_credential_role set role_id='38676f30eaeb44518bf7e144a73c8da6' where role_id='4d8cad783d6342efa8414d7d36fbc034' * 10:11 dcaro: there was a little unavailability for some VMs while ceph was starting to rebalance things, but it seems stable and moving data around ([[phab:T331141|T331141]]) * 09:37 dcaro: Changing ceph crush map to allow rack HA on eqiad1 cluster ([[phab:T331141|T331141]]) === 2023-03-03 === * 12:12 arturo: installing haproxy updates ([[phab:T331119|T331119]]) === 2023-03-02 === * 16:17 wm-bot2: The cluster is now rebalanced after adding the new OSDs ['cloudcephosd1010.eqiad.wmnet'] ([[phab:T329504|T329504]]) - cookbook ran by dcaro@vulcanus * 14:21 wm-bot2: Added 1 new OSDs ['cloudcephosd1010.eqiad.wmnet'] ([[phab:T329504|T329504]]) - cookbook ran by dcaro@vulcanus * 14:21 wm-bot2: Added OSD cloudcephosd1010.eqiad.wmnet... (1/1) ([[phab:T329504|T329504]]) - cookbook ran by dcaro@vulcanus * 14:13 wm-bot2: Finished rebooting node cloudcephosd1010.eqiad.wmnet ([[phab:T329504|T329504]]) - cookbook ran by dcaro@vulcanus * 14:10 wm-bot2: Rebooting node cloudcephosd1010.eqiad.wmnet ([[phab:T329504|T329504]]) - cookbook ran by dcaro@vulcanus * 14:09 wm-bot2: Adding OSD cloudcephosd1010.eqiad.wmnet... (1/1) ([[phab:T329504|T329504]]) - cookbook ran by dcaro@vulcanus * 14:09 wm-bot2: Adding new OSDs ['cloudcephosd1010.eqiad.wmnet'] to the cluster ([[phab:T329504|T329504]]) - cookbook ran by dcaro@vulcanus * 10:43 wm-bot2: The cluster is now rebalanced after adding the new OSDs ['cloudcephosd1005.eqiad.wmnet'] ([[phab:T329504|T329504]]) - cookbook ran by dcaro@vulcanus * 08:44 wm-bot2: Added 1 new OSDs ['cloudcephosd1005.eqiad.wmnet'] ([[phab:T329504|T329504]]) - cookbook ran by dcaro@vulcanus * 08:44 wm-bot2: Added OSD cloudcephosd1005.eqiad.wmnet... (1/1) ([[phab:T329504|T329504]]) - cookbook ran by dcaro@vulcanus * 08:36 wm-bot2: Finished rebooting node cloudcephosd1005.eqiad.wmnet ([[phab:T329504|T329504]]) - cookbook ran by dcaro@vulcanus * 08:33 wm-bot2: Rebooting node cloudcephosd1005.eqiad.wmnet ([[phab:T329504|T329504]]) - cookbook ran by dcaro@vulcanus * 08:32 wm-bot2: Adding OSD cloudcephosd1005.eqiad.wmnet... (1/1) ([[phab:T329504|T329504]]) - cookbook ran by dcaro@vulcanus * 08:32 wm-bot2: Adding new OSDs ['cloudcephosd1005.eqiad.wmnet'] to the cluster ([[phab:T329504|T329504]]) - cookbook ran by dcaro@vulcanus === 2023-02-28 === * 22:47 andrewbogott: adding new 'member' role assignment to every user/project pair that currently has the 'user' assignment. [[phab:T330759|T330759]] * 09:47 wm-bot2: Depooled and destroyed OSD daemons [79, 78, 77, 76, 75, 74, 73, 72] and removed the OSD host cloudcephosd1010 from the CRUSH map. ([[phab:T329504|T329504]]) - cookbook ran by dcaro@vulcanus * 09:46 wm-bot2: Destroying OSDs with ids in [79, 78, 77, 76, 75, 74, 73, 72] on cloudcephosd1010 from eqiad1 ([[phab:T329504|T329504]]) - cookbook ran by dcaro@vulcanus * 09:28 wm-bot2: Depooling OSDs with ids in [79, 78, 77, 76, 75, 74, 73, 72] on cloudcephosd1010 from eqiad1 ([[phab:T329504|T329504]]) - cookbook ran by dcaro@vulcanus * 09:13 wm-bot2: Depooling OSDs with ids in [79, 78, 77, 76, 75, 74, 73, 72] on cloudcephosd1010 from eqiad1 ([[phab:T329504|T329504]]) - cookbook ran by dcaro@vulcanus * 09:12 wm-bot2: Depooled and destroyed OSD daemons [39, 38, 37, 36, 35, 34, 33, 32] and removed the OSD host cloudcephosd1005 from the CRUSH map. ([[phab:T329504|T329504]]) - cookbook ran by dcaro@vulcanus * 09:11 wm-bot2: Destroying OSDs with ids in [39, 38, 37, 36, 35, 34, 33, 32] on cloudcephosd1005 from eqiad1 ([[phab:T329504|T329504]]) - cookbook ran by dcaro@vulcanus * 08:55 wm-bot2: Depooling OSDs with ids in [39, 38, 37, 36, 35, 34, 33, 32] on cloudcephosd1005 from eqiad1 ([[phab:T329504|T329504]]) - cookbook ran by dcaro@vulcanus === 2023-02-27 === * 21:01 wm-bot2: The cluster is now rebalanced after adding the new OSDs ['cloudcephosd1004.eqiad.wmnet'] ([[phab:T329502|T329502]]) - cookbook ran by dcaro@vulcanus * 18:53 wm-bot2: Added 1 new OSDs ['cloudcephosd1004.eqiad.wmnet'] ([[phab:T329502|T329502]]) - cookbook ran by dcaro@vulcanus * 18:53 wm-bot2: Added OSD cloudcephosd1004.eqiad.wmnet... (1/1) ([[phab:T329502|T329502]]) - cookbook ran by dcaro@vulcanus * 18:45 wm-bot2: Finished rebooting node cloudcephosd1004.eqiad.wmnet ([[phab:T329502|T329502]]) - cookbook ran by dcaro@vulcanus * 18:42 wm-bot2: Rebooting node cloudcephosd1004.eqiad.wmnet ([[phab:T329502|T329502]]) - cookbook ran by dcaro@vulcanus * 18:41 wm-bot2: Adding OSD cloudcephosd1004.eqiad.wmnet... (1/1) ([[phab:T329502|T329502]]) - cookbook ran by dcaro@vulcanus * 18:41 wm-bot2: Adding new OSDs ['cloudcephosd1004.eqiad.wmnet'] to the cluster ([[phab:T329502|T329502]]) - cookbook ran by dcaro@vulcanus * 18:38 wm-bot2: Rebooting node cloudcephosd1004.eqiad.wmnet ([[phab:T329502|T329502]]) - cookbook ran by dcaro@vulcanus * 18:38 wm-bot2: Adding OSD cloudcephosd1004.eqiad.wmnet... (1/1) ([[phab:T329502|T329502]]) - cookbook ran by dcaro@vulcanus * 18:38 wm-bot2: Adding new OSDs ['cloudcephosd1004.eqiad.wmnet'] to the cluster ([[phab:T329502|T329502]]) - cookbook ran by dcaro@vulcanus * 18:14 wm-bot2: The cluster is now rebalanced after adding the new OSDs ['cloudcephosd1003.eqiad.wmnet'] ([[phab:T329502|T329502]]) - cookbook ran by dcaro@vulcanus * 16:13 wm-bot2: Added 1 new OSDs ['cloudcephosd1003.eqiad.wmnet'] ([[phab:T329502|T329502]]) - cookbook ran by dcaro@vulcanus * 16:13 wm-bot2: Added OSD cloudcephosd1003.eqiad.wmnet... (1/1) ([[phab:T329502|T329502]]) - cookbook ran by dcaro@vulcanus * 16:05 wm-bot2: Finished rebooting node cloudcephosd1003.eqiad.wmnet ([[phab:T329502|T329502]]) - cookbook ran by dcaro@vulcanus * 16:02 wm-bot2: Rebooting node cloudcephosd1003.eqiad.wmnet ([[phab:T329502|T329502]]) - cookbook ran by dcaro@vulcanus * 16:01 wm-bot2: Adding OSD cloudcephosd1003.eqiad.wmnet... (1/1) ([[phab:T329502|T329502]]) - cookbook ran by dcaro@vulcanus * 16:01 wm-bot2: Adding new OSDs ['cloudcephosd1003.eqiad.wmnet'] to the cluster ([[phab:T329502|T329502]]) - cookbook ran by dcaro@vulcanus * 15:59 wm-bot2: Adding OSD coludcephosd1003.eqiad.wmnet... (1/1) ([[phab:T329502|T329502]]) - cookbook ran by dcaro@vulcanus * 15:59 wm-bot2: Adding new OSDs ['coludcephosd1003.eqiad.wmnet'] to the cluster ([[phab:T329502|T329502]]) - cookbook ran by dcaro@vulcanus * 15:58 wm-bot2: Adding OSD coludcephosd1003... (1/1) ([[phab:T329502|T329502]]) - cookbook ran by dcaro@vulcanus * 15:58 wm-bot2: Adding new OSDs ['coludcephosd1003'] to the cluster ([[phab:T329502|T329502]]) - cookbook ran by dcaro@vulcanus === 2023-02-22 === * 08:25 wm-bot2: Depooled and destroyed OSD daemons [31, 30, 29, 28, 27, 26, 25, 24] and removed the OSD host cloudcephosd1004 from the CRUSH map. ([[phab:T329502|T329502]]) - cookbook ran by dcaro@vulcanus * 08:25 wm-bot2: Destroying OSDs with ids in [31, 30, 29, 28, 27, 26, 25, 24] on cloudcephosd1004 from eqiad1 ([[phab:T329502|T329502]]) - cookbook ran by dcaro@vulcanus * 08:22 wm-bot2: Depooling OSDs with ids in [31, 30, 29, 28, 27, 26, 25, 24] on cloudcephosd1004 from eqiad1 ([[phab:T329502|T329502]]) - cookbook ran by dcaro@vulcanus * 08:15 wm-bot2: Destroying OSDs with ids in [31, 30, 29, 28, 27, 26, 25, 24] on cloudcephosd1004 from eqiad1 ([[phab:T329502|T329502]]) - cookbook ran by dcaro@vulcanus * 08:13 wm-bot2: Depooling OSDs with ids in [31, 30, 29, 28, 27, 26, 25, 24] on cloudcephosd1004 from eqiad1 ([[phab:T329502|T329502]]) - cookbook ran by dcaro@vulcanus * 07:23 wm-bot2: Destroying OSDs with ids in [31, 30, 29, 28, 27, 26, 25, 24] on cloudcephosd1004 from eqiad1 ([[phab:T329502|T329502]]) - cookbook ran by dcaro@vulcanus * 07:02 wm-bot2: Depooling OSDs with ids in [31, 30, 29, 28, 27, 26, 25, 24] on cloudcephosd1004 from eqiad1 ([[phab:T329502|T329502]]) - cookbook ran by dcaro@vulcanus === 2023-02-21 === * 20:27 andrewbogott: deleted 200 more orphaned VM images with wmcs-novastats-cephleaks * 17:18 andrewbogott: shutting down postgres on clouddb1004/1003, then shutting down the vms * 15:50 wm-bot2: Destroying OSDs with ids in [71, 70, 69, 68, 67, 66, 65, 64] on cloudcephosd1003 from eqiad1 ([[phab:T329502|T329502]]) - cookbook ran by dcaro@vulcanus * 15:48 wm-bot2: Depooling OSDs with ids in [71, 70, 69, 68, 67, 66, 65, 64] on cloudcephosd1003 from eqiad1 ([[phab:T329502|T329502]]) - cookbook ran by dcaro@vulcanus * 14:21 wm-bot2: Destroying OSDs with ids in [71, 70, 69, 68, 67, 66, 65, 64] on cloudcephosd1003 from eqiad1 ([[phab:T329502|T329502]]) - cookbook ran by dcaro@vulcanus * 14:00 wm-bot2: Depooling OSDs with ids in [71, 70, 69, 68, 67, 66, 65, 64] on cloudcephosd1003 from eqiad1 ([[phab:T329502|T329502]]) - cookbook ran by dcaro@vulcanus === 2023-02-16 === * 19:14 wm-bot2: The cluster is now rebalanced after adding the new OSDs ['cloudcephosd1002.eqiad.wmnet'] ([[phab:T329498|T329498]]) - cookbook ran by dcaro@vulcanus * 17:55 dcaro: Manually zapped /dev/sdc on cloudcephosd1002, probably a leftover drive since the beginning (or during the reimage the drives changed names, and this one had leftovers from the previous OS) ([[phab:T329498|T329498]]) * 17:47 wm-bot2: Added 1 new OSDs ['cloudcephosd1002.eqiad.wmnet'] ([[phab:T329498|T329498]]) - cookbook ran by dcaro@vulcanus * 17:47 wm-bot2: Added OSD cloudcephosd1002.eqiad.wmnet... (1/1) ([[phab:T329498|T329498]]) - cookbook ran by dcaro@vulcanus * 17:42 wm-bot2: Adding OSD cloudcephosd1002.eqiad.wmnet... (1/1) ([[phab:T329498|T329498]]) - cookbook ran by dcaro@vulcanus * 17:42 wm-bot2: Adding new OSDs ['cloudcephosd1002.eqiad.wmnet'] to the cluster ([[phab:T329498|T329498]]) - cookbook ran by dcaro@vulcanus * 16:50 wm-bot2: Adding OSD cloudcephosd1002.eqiad.wmnet... (1/1) ([[phab:T329498|T329498]]) - cookbook ran by dcaro@vulcanus * 16:50 wm-bot2: Adding new OSDs ['cloudcephosd1002.eqiad.wmnet'] to the cluster ([[phab:T329498|T329498]]) - cookbook ran by dcaro@vulcanus * 16:01 wm-bot2: Adding OSD cloudcephosd1002.eqiad.wmnet... (1/1) ([[phab:T329498|T329498]]) - cookbook ran by dcaro@vulcanus * 16:01 wm-bot2: Adding new OSDs ['cloudcephosd1002.eqiad.wmnet'] to the cluster ([[phab:T329498|T329498]]) - cookbook ran by dcaro@vulcanus * 16:00 wm-bot2: Adding OSD cloudcephosd1002.eqiad.wmnet... (1/1) ([[phab:T329498|T329498]]) - cookbook ran by dcaro@vulcanus * 16:00 wm-bot2: Adding new OSDs ['cloudcephosd1002.eqiad.wmnet'] to the cluster ([[phab:T329498|T329498]]) - cookbook ran by dcaro@vulcanus * 14:05 wm-bot2: Added 1 new OSDs ['cloudcephosd1001.eqiad.wmnet'] ([[phab:T329498|T329498]]) - cookbook ran by dcaro@vulcanus * 13:29 wm-bot2: Adding new OSDs ['cloudcephosd1001.eqiad.wmnet'] to the cluster ([[phab:T329498|T329498]]) - cookbook ran by dcaro@vulcanus * 13:29 wm-bot2: Adding new OSDs ['cloudcephosd1001.eqiad.wmnet'] to the cluster ([[phab:T329498|T329498]]) - cookbook ran by dcaro@vulcanus * 13:24 wm-bot2: Adding OSD cloudcephosd1001.eqiad.wmnet... (1/1) ([[phab:T329498|T329498]]) - cookbook ran by dcaro@vulcanus * 13:24 wm-bot2: Adding new OSDs ['cloudcephosd1001.eqiad.wmnet'] to the cluster ([[phab:T329498|T329498]]) - cookbook ran by dcaro@vulcanus * 13:23 wm-bot2: Destroying OSDs with ids in [63, 62, 61, 60, 59, 58, 57, 56] on cloudcephosd1002 from eqiad1 ([[phab:T329498|T329498]]) - cookbook ran by dcaro@vulcanus * 13:21 wm-bot2: Depooling OSDs with ids in [63, 62, 61, 60, 59, 58, 57, 56] on cloudcephosd1002 from eqiad1 ([[phab:T329498|T329498]]) - cookbook ran by dcaro@vulcanus * 13:14 wm-bot2: Adding OSD cloudcephosd1001.eqiad.wmnet... (1/1) ([[phab:T329498|T329498]]) - cookbook ran by dcaro@vulcanus * 13:14 wm-bot2: Adding new OSDs ['cloudcephosd1001.eqiad.wmnet'] to the cluster ([[phab:T329498|T329498]]) - cookbook ran by dcaro@vulcanus * 11:20 wm-bot2: Destroying OSDs with ids in [53, 52, 51, 50] on cloudcephosd1001 from eqiad1 ([[phab:T329498|T329498]]) - cookbook ran by dcaro@vulcanus * 11:19 wm-bot2: Depooling OSDs with ids in [53, 52, 51, 50] on cloudcephosd1001 from eqiad1 ([[phab:T329498|T329498]]) - cookbook ran by dcaro@vulcanus * 11:03 wm-bot2: Destroying OSDs with ids in [55, 54, 53, 52, 51, 50] on cloudcephosd1001 from eqiad1 ([[phab:T329498|T329498]]) - cookbook ran by dcaro@vulcanus * 11:01 wm-bot2: Depooling OSDs with ids in [55, 54, 53, 52, 51, 50] on cloudcephosd1001 from eqiad1 ([[phab:T329498|T329498]]) - cookbook ran by dcaro@vulcanus * 10:59 wm-bot2: Depooling OSDs with ids in [55, 54, 53, 52, 51, 50] on cloudcephosd1001 from eqiad1 ([[phab:T329498|T329498]]) - cookbook ran by dcaro@vulcanus * 10:15 dcaro: purges osd daemons 48 and 40 from eqiad ceph cluster ([[phab:T329709|T329709]]) === 2023-02-15 === * 14:53 andrewbogott: deleting another 100 leaked VM images with wmcs-novastats-cephleaks * 13:39 wm-bot2: Destroying OSDs with id [48] on cloudcephosd1001 from eqiad1 - cookbook ran by dcaro@vulcanus * 13:14 wm-bot2: Destroying OSDs with id [48] on cloudcephosd1001 from eqiad1 - cookbook ran by dcaro@vulcanus * 13:13 wm-bot2: Destroying OSDs with id [12345] on cloudcephosd1001 from eqiad1 - cookbook ran by dcaro@vulcanus * 13:11 wm-bot2: Destroying OSDs with id [12345] on cloudcephosd1001 from eqiad1 - cookbook ran by dcaro@vulcanus * 13:11 wm-bot2: Destroying OSDs with id [12345] on cloudcephosd1001 from eqiad1 - cookbook ran by dcaro@vulcanus * 13:10 wm-bot2: Destroying OSDs with id [[12345]] on cloudcephosd1001 from eqiad1 - cookbook ran by dcaro@vulcanus * 13:09 wm-bot2: Destroying OSDs with id ['12345'] on cloudcephosd1001 from eqiad1 - cookbook ran by dcaro@vulcanus === 2023-02-14 === * 13:17 andrewbogott: restarting all eqiad1 openstack services because that seems to sometimes help things *shrug* === 2023-02-13 === * 14:06 wm-bot2: Set the ceph cluster for eqiad1 in maintenance, alert silence ids: 8fbf6bfd-eec1-4d81-8e0d-{{Gerrit|ea431d8411ee}} ([[phab:T329498|T329498]]) - cookbook ran by dcaro@vulcanus * 13:32 taavi: re-enable puppet on labstore1004 [[phab:T329377|T329377]] === 2023-02-09 === * 21:17 andrewbogott: deleted 10% of leaked VM ceph images using wmcs-novastats-cephleaks (only 10% out of an abundance of caution) === 2023-02-08 === * 17:08 arturo: changing to cloudgw network setup, make VIPs /32 ([[phab:T295774|T295774]]) === 2023-02-07 === * 11:26 arturo: [codfw1dev] testing network changes in cloudgw, expect unrealiable network ([[phab:T295774|T295774]]) === 2023-02-04 === * 13:44 taavi: drop old columns from oathauth_users table on labtestwiki [[phab:T328131|T328131]] === 2023-02-03 === * 15:00 andrewbogott: restarted nova services in eqiad1 in an attempt to eke out another day or two of stability * 14:13 taavi: attached GrapheSuppression developer account to wikitech === 2023-02-02 === * 13:14 dcaro_away: draining osd.48 from node cloudcephosd1001 ([[phab:T316544|T316544]]) * 12:57 wm-bot2: Set the ceph cluster for eqiad1 in maintenance, alert silence ids: 7ac2b25a-d1bb-4789-8aa6-{{Gerrit|b9435b505349}} ([[phab:T316544|T316544]]) - cookbook ran by dcaro@vulcanus === 2023-01-30 === * 22:34 wm-bot2: Upgraded and rebooted host cloudrabbit1002.wikimedia.org - cookbook ran by andrew@bullseye * 21:34 andrewbogott: merging https://gerrit.wikimedia.org/r/c/operations/puppet/+/884922 and upgrading rabbitmq nodes for [[phab:T328155|T328155]] === 2023-01-27 === * 20:08 wm-bot2: Upgraded and rebooted host cloudcontrol2005-dev.wikimedia.org - cookbook ran by andrew@bullseye * 19:22 wm-bot2: Upgraded and rebooted host cloudcontrol2004-dev.wikimedia.org - cookbook ran by andrew@bullseye * 19:10 wm-bot2: Upgraded and rebooted host cloudcontrol2001-dev.wikimedia.org - cookbook ran by andrew@bullseye * 15:25 andrewbogott: restarting openstack services in eqiad1, another attempt to address instability === 2023-01-26 === * 20:34 andrewbogott: shutting down mariadb on cloudbackup2001-dev, testing the waters for [[phab:T328079|T328079]] === 2023-01-22 === * 03:42 andrewbogott: reset eqiad1 rabbitmq in an attempt to resolve some mild instability === 2023-01-20 === * 15:26 wm-bot2: Removed cloudweb hosts (cloudweb2002-dev.wikimedia.org) from maintenance mode. - cookbook ran by andrew@bullseye * 15:26 wm-bot2: Put cloudweb hosts (cloudweb2002-dev.wikimedia.org) into maintenance mode (downtime id: ['f47a3d91-b270-4c90-acc8-d85075a6bf8e'], use this to unset) - cookbook ran by andrew@bullseye * 13:15 arturo: reinstall python3-neutron (to reset manual patching) on all cloudnet nodes and patch it via puppet, then restart neutron-l3-agent by hand ([[phab:T327463|T327463]]) * 10:12 arturo: [codfw1dev] failover neutron-l3-agent between cloudnet2005-dev/cloudnet2006-dev a couple of times [[phab:T327463|T327463]] * 02:17 andrewbogott: stopping neutron-l3-agent on cloudnet1005 because it's logging at a furious rate and about to fill the drive === 2023-01-19 === * 18:06 wm-bot2: Removed cloudweb hosts (cloudweb2002-dev.wikimedia.org) from maintenance mode. - cookbook ran by andrew@bullseye * 18:06 wm-bot2: Put cloudweb hosts (cloudweb2002-dev.wikimedia.org) into maintenance mode (downtime id: ['36d5af6a-7d8e-4d0c-831e-1bf05c255984'], use this to unset) - cookbook ran by andrew@bullseye * 17:35 wm-bot2: Removed cloudweb hosts (cloudweb2002-dev.wikimedia.org) from maintenance mode. - cookbook ran by andrew@bullseye * 17:22 wm-bot2: Put cloudweb hosts (cloudweb2002-dev.wikimedia.org) into maintenance mode (downtime id: ['66ec4f04-d25b-4067-be9e-2fe12cb1d3ff'], use this to unset) - cookbook ran by andrew@bullseye === 2023-01-18 === * 22:32 wm-bot2: Set cloudweb cloudweb2002-dev.wikimedia.org maintenance (downtime id: 347cb75e-215e-4b85-ae14-{{Gerrit|4ce1934c70c7}}, use this to unset) - cookbook ran by andrew@bullseye * 22:01 wm-bot2: Set cloudweb cloudweb2002-dev.wikimedia.org maintenance (downtime id: 9bf6212f-7fdb-4869-8190-{{Gerrit|b07387b2bc7e}}, use this to unset) - cookbook ran by andrew@bullseye * 21:40 wm-bot2: Set cloudweb cloudweb2002-dev.wikimedia.org maintenance (downtime id: 11aec8ea-f443-41ae-b79a-{{Gerrit|5e1e3aa94546}}, use this to unset) - cookbook ran by andrew@bullseye * 20:20 wm-bot2: Restarting openstack services on cloudservices2004-dev: ['designate-mdns', 'designate-sink', 'designate-central', 'designate-producer', 'designate-worker', 'designate-agent'] - cookbook ran by andrew@bullseye * 20:20 wm-bot2: Restarting openstack services on cloudservices2005-dev: ['designate-central', 'designate-sink', 'designate-worker', 'designate-producer', 'designate-mdns', 'designate-agent'] - cookbook ran by andrew@bullseye * 20:20 wm-bot2: Restarting openstack services on cloudnet2005-dev: ['neutron-metadata-agent', 'neutron-dhcp-agent', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 20:20 wm-bot2: Restarting openstack services on cloudnet2006-dev: ['neutron-dhcp-agent', 'neutron-metadata-agent', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 20:19 wm-bot2: Restarting openstack services on cloudbackup1002-dev: ['cinder-backup'] - cookbook ran by andrew@bullseye * 20:19 wm-bot2: Restarting openstack services on cloudbackup1001-dev: ['cinder-backup'] - cookbook ran by andrew@bullseye * 20:19 wm-bot2: Restarting openstack services on cloudcontrol2005-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 20:19 wm-bot2: Restarting openstack services on cloudvirt2003-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 20:19 wm-bot2: Restarting openstack services on cloudvirt2002-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 20:18 wm-bot2: Restarting openstack services on cloudvirt2001-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 20:18 wm-bot2: Restarting openstack services on cloudcontrol2004-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 20:18 wm-bot2: Restarting openstack services on cloudcontrol2001-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 17:03 wm-bot2: Restarting openstack services on cloudservices2004-dev: ['designate-mdns', 'designate-sink', 'designate-central', 'designate-producer', 'designate-worker', 'designate-agent'] - cookbook ran by andrew@bullseye * 17:02 wm-bot2: Restarting openstack services on cloudservices2005-dev: ['designate-central', 'designate-sink', 'designate-worker', 'designate-producer', 'designate-mdns', 'designate-agent'] - cookbook ran by andrew@bullseye * 17:02 wm-bot2: Restarting openstack services on cloudnet2005-dev: ['neutron-metadata-agent', 'neutron-dhcp-agent', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 17:02 wm-bot2: Restarting openstack services on cloudnet2006-dev: ['neutron-dhcp-agent', 'neutron-metadata-agent', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 17:02 wm-bot2: Restarting openstack services on cloudbackup1002-dev: ['cinder-backup'] - cookbook ran by andrew@bullseye * 17:02 wm-bot2: Restarting openstack services on cloudbackup1001-dev: ['cinder-backup'] - cookbook ran by andrew@bullseye * 17:02 wm-bot2: Restarting openstack services on cloudcontrol2005-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 17:02 wm-bot2: Restarting openstack services on cloudvirt2003-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 17:01 wm-bot2: Restarting openstack services on cloudvirt2002-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 17:01 wm-bot2: Restarting openstack services on cloudvirt2001-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 17:01 wm-bot2: Restarting openstack services on cloudcontrol2004-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 17:01 wm-bot2: Restarting openstack services on cloudcontrol2001-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 14:38 wm-bot2: Restarting openstack services on cloudvirt2002-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 14:38 wm-bot2: Restarting openstack services on cloudvirt2001-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 14:37 wm-bot2: Restarting openstack services on cloudcontrol2004-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 14:37 wm-bot2: Restarting openstack services on cloudcontrol2001-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye === 2023-01-17 === * 19:32 wm-bot2: Upgraded and rebooted host cloudbackup2002.codfw.wmnet - cookbook ran by andrew@bullseye * 18:04 wm-bot2: Upgraded and rebooted host cloudnet1005.eqiad.wmnet - cookbook ran by andrew@bullseye * 17:52 wm-bot2: Upgraded and rebooted host cloudcontrol1007.wikimedia.org - cookbook ran by andrew@bullseye * 17:36 wm-bot2: Upgraded and rebooted host cloudcontrol1006.wikimedia.org - cookbook ran by andrew@bullseye * 17:23 wm-bot2: Upgraded and rebooted host cloudcontrol1005.wikimedia.org - cookbook ran by andrew@bullseye === 2023-01-13 === * 17:36 wm-bot2: Restarting openstack services on cloudcontrol2005-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata'] - cookbook ran by andrew@bullseye * 17:36 wm-bot2: Restarting openstack services on cloudvirt2003-dev: ['nova-compute'] - cookbook ran by andrew@bullseye * 17:36 wm-bot2: Restarting openstack services on cloudvirt2002-dev: ['nova-compute'] - cookbook ran by andrew@bullseye * 17:35 wm-bot2: Restarting openstack services on cloudvirt2001-dev: ['nova-compute'] - cookbook ran by andrew@bullseye * 17:35 wm-bot2: Restarting openstack services on cloudcontrol2004-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata'] - cookbook ran by andrew@bullseye * 17:35 wm-bot2: Restarting openstack services on cloudcontrol2001-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata'] - cookbook ran by andrew@bullseye * 17:33 wm-bot2: Restarting openstack services on cloudvirt2001-dev: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 17:33 wm-bot2: Restarting openstack services on cloudvirt2002-dev: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 17:33 wm-bot2: Restarting openstack services on cloudnet2005-dev: ['neutron-metadata-agent', 'neutron-dhcp-agent', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 17:33 wm-bot2: Restarting openstack services on cloudvirt2003-dev: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 17:33 wm-bot2: Restarting openstack services on cloudnet2006-dev: ['neutron-dhcp-agent', 'neutron-metadata-agent', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 17:26 wm-bot2: Restarting openstack services on cloudvirt2001-dev: ['nova-compute'] - cookbook ran by andrew@bullseye * 17:26 wm-bot2: Restarting openstack services on cloudcontrol2004-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata'] - cookbook ran by andrew@bullseye * 17:25 wm-bot2: Restarting openstack services on cloudcontrol2001-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata'] - cookbook ran by andrew@bullseye * 17:21 wm-bot2: Restarting openstack services: <nowiki>{</nowiki>'cloudcontrol2001-dev': ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata'], 'cloudcontrol2004-dev': ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata'], 'cloudvirt2001-dev': ['nova-compute'], 'cloudvirt2002-dev': ['nova-compute'], 'cloudvirt2003-dev': ['nova-compute'], 'cloudcontrol2005-dev': ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata * 10:41 arturo: restart backup_vm.service on cloudbackup1003/1004 to recover from a nova Unknown Error (HTTP 503) === 2023-01-12 === * 22:34 andrewbogott: updated the Bullseye base image with the upstream {{Gerrit|20221219}} build * 14:12 wm-bot2: Created project checkuser-beta-wiki with default quotas. ([[phab:T326740|T326740]]) - cookbook ran by arturo@nostromo === 2023-01-06 === * 18:42 wm-bot2: Safe reboot of cloudvirt2003-dev.codfw.wmnet finished successfully - cookbook ran by andrew@bullseye * 18:42 wm-bot2: unset cloudvirt2003-dev.codfw.wmnet maintenance (aggregates: ceph) - cookbook ran by andrew@bullseye * 18:39 wm-bot2: Drained cloudvirt2003-dev.codfw.wmnet - cookbook ran by andrew@bullseye * 18:35 wm-bot2: Set cloudvirt cloudvirt2003-dev.codfw.wmnet maintenance (downtime id: b50d9d3a-4f1d-4522-b86f-{{Gerrit|722fc9c55c87}}, use this to unset) - cookbook ran by andrew@bullseye * 18:35 wm-bot2: Draining cloudvirt2003-dev.codfw.wmnet - cookbook ran by andrew@bullseye * 18:35 wm-bot2: Safe rebooting cloudvirt2003-dev.codfw.wmnet - cookbook ran by andrew@bullseye * 18:35 wm-bot2: Draining cloudvirt2003-dev.cdofw.wmnet - cookbook ran by andrew@bullseye * 18:35 wm-bot2: Safe rebooting cloudvirt2003-dev.cdofw.wmnet - cookbook ran by andrew@bullseye * 00:14 wm-bot2: Upgraded and rebooted host cloudbackup1002-dev.eqiad.wmnet - cookbook ran by andrew@bullseye * 00:08 wm-bot2: Upgraded and rebooted host cloudbackup1001-dev.eqiad.wmnet - cookbook ran by andrew@bullseye * 00:01 wm-bot2: Upgraded and rebooted host cloudnet2006-dev.codfw.wmnet - cookbook ran by andrew@bullseye === 2023-01-05 === * 23:54 wm-bot2: Upgraded and rebooted host cloudnet2005-dev.codfw.wmnet - cookbook ran by andrew@bullseye * 23:38 wm-bot2: Upgraded and rebooted host cloudcontrol2005-dev.wikimedia.org - cookbook ran by andrew@bullseye * 23:25 wm-bot2: Upgraded and rebooted host cloudcontrol2004-dev.wikimedia.org - cookbook ran by andrew@bullseye * 23:13 wm-bot2: Upgraded and rebooted host cloudcontrol2001-dev.wikimedia.org - cookbook ran by andrew@bullseye * 22:18 wm-bot2: Upgraded and rebooted host cloudcontrol2001-dev.wikimedia.org - cookbook ran by andrew@bullseye * 22:10 andrewbogott: upgrading codfw1dev openstack to version 'zed' === 2023-01-04 === * 21:18 wm-bot2: Upgraded and rebooted host cloudservices1004.wikimedia.org - cookbook ran by andrew@bullseye * 21:11 wm-bot2: Upgraded and rebooted host cloudservices1005.wikimedia.org - cookbook ran by andrew@bullseye * 20:12 wm-bot2: Upgraded and rebooted host cloudservices1004.wikimedia.org - cookbook ran by andrew@bullseye * 20:04 wm-bot2: Upgraded and rebooted host cloudservices1005.wikimedia.org - cookbook ran by andrew@bullseye * 14:45 wm-bot2: Finished rebooting the nodes ['cloudcephmon2004-dev', 'cloudcephmon2005-dev', 'cloudcephmon2006-dev'] - cookbook ran by fran@wmf3169 * 14:44 wm-bot2: Finished rebooting node cloudcephmon2006-dev.codfw.wmnet - cookbook ran by fran@wmf3169 * 14:41 wm-bot2: Rebooting node cloudcephmon2006-dev.codfw.wmnet - cookbook ran by fran@wmf3169 * 14:41 wm-bot2: Finished rebooting node cloudcephmon2005-dev.codfw.wmnet - cookbook ran by fran@wmf3169 * 14:38 wm-bot2: Rebooting node cloudcephmon2005-dev.codfw.wmnet - cookbook ran by fran@wmf3169 * 14:37 wm-bot2: Finished rebooting node cloudcephmon2004-dev.codfw.wmnet - cookbook ran by fran@wmf3169 * 14:34 wm-bot2: Rebooting node cloudcephmon2004-dev.codfw.wmnet - cookbook ran by fran@wmf3169 * 14:34 wm-bot2: Rebooting the nodes cloudcephmon2004-dev,cloudcephmon2005-dev,cloudcephmon2006-dev - cookbook ran by fran@wmf3169 === 2023-01-03 === * 22:11 wm-bot2: Upgraded and rebooted host cloudservices2005-dev.wikimedia.org - cookbook ran by andrew@bullseye * 21:55 wm-bot2: Upgraded and rebooted host cloudservices2004-dev.wikimedia.org - cookbook ran by andrew@bullseye * 15:25 taavi: restart designate-sink everywhere to pick up wmf-sink changes === 2022-12-25 === * 14:21 taavi: register developer account 'instance-puppet-user-dev' to update the codfw1dev instance-puppet repo without access to the eqiad1 repo [[phab:T318504|T318504]] === 2022-12-22 === * 15:16 dcaro: added submit rights for JenkinsBot on all cloud/* gerrit repos === 2022-12-21 === * 04:59 wm-bot2: Rebooting node cloudcephosd1030.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 04:59 wm-bot2: Finished rebooting node cloudcephosd1029.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 04:35 wm-bot2: Rebooting node cloudcephosd1029.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 04:35 wm-bot2: Finished rebooting node cloudcephosd1028.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 04:12 wm-bot2: Rebooting node cloudcephosd1028.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 04:12 wm-bot2: Finished rebooting node cloudcephosd1027.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 03:48 wm-bot2: Rebooting node cloudcephosd1027.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 03:48 wm-bot2: Finished rebooting node cloudcephosd1026.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 03:24 wm-bot2: Rebooting node cloudcephosd1026.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 03:24 wm-bot2: Finished rebooting node cloudcephosd1025.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 03:00 wm-bot2: Rebooting node cloudcephosd1025.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 03:00 wm-bot2: Finished rebooting node cloudcephosd1024.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 02:36 wm-bot2: Rebooting node cloudcephosd1024.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 02:36 wm-bot2: Finished rebooting node cloudcephosd1023.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 02:12 wm-bot2: Rebooting node cloudcephosd1023.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 02:12 wm-bot2: Finished rebooting node cloudcephosd1022.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 01:47 wm-bot2: Rebooting node cloudcephosd1022.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 01:47 wm-bot2: Finished rebooting node cloudcephosd1021.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 01:22 wm-bot2: Rebooting node cloudcephosd1021.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 01:22 wm-bot2: Finished rebooting node cloudcephosd1020.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 00:58 wm-bot2: Rebooting node cloudcephosd1020.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 00:58 wm-bot2: Finished rebooting node cloudcephosd1019.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 00:34 wm-bot2: Rebooting node cloudcephosd1019.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 00:34 wm-bot2: Finished rebooting node cloudcephosd1018.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 00:09 wm-bot2: Rebooting node cloudcephosd1018.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 00:09 wm-bot2: Finished rebooting node cloudcephosd1017.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye === 2022-12-20 === * 23:45 wm-bot2: Rebooting node cloudcephosd1017.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 23:45 wm-bot2: Finished rebooting node cloudcephosd1016.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 23:21 wm-bot2: Rebooting node cloudcephosd1016.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 23:21 wm-bot2: Finished rebooting node cloudcephosd1015.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 22:56 wm-bot2: Rebooting node cloudcephosd1015.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 22:55 wm-bot2: Finished rebooting node cloudcephosd1014.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 22:30 wm-bot2: Rebooting node cloudcephosd1014.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 22:30 wm-bot2: Finished rebooting node cloudcephosd1013.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 22:08 andrewbogott: restarting openstack services with wmcs.openstack.restart_openstack due to miscellaneous trove failures * 22:06 wm-bot2: Rebooting node cloudcephosd1013.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 22:06 wm-bot2: Finished rebooting node cloudcephosd1012.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 21:40 wm-bot2: Rebooting node cloudcephosd1012.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 21:40 wm-bot2: Finished rebooting node cloudcephosd1011.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 21:15 wm-bot2: Rebooting node cloudcephosd1011.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 21:15 wm-bot2: Finished rebooting node cloudcephosd1010.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 20:50 wm-bot2: Rebooting node cloudcephosd1010.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 20:50 wm-bot2: Finished rebooting node cloudcephosd1009.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 20:26 wm-bot2: Rebooting node cloudcephosd1009.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 20:26 wm-bot2: Finished rebooting node cloudcephosd1008.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 20:01 wm-bot2: Rebooting node cloudcephosd1008.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 20:01 wm-bot2: Finished rebooting node cloudcephosd1007.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 19:35 wm-bot2: Rebooting node cloudcephosd1007.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 19:35 wm-bot2: Finished rebooting node cloudcephosd1006.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 19:10 wm-bot2: Rebooting node cloudcephosd1006.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 19:10 wm-bot2: Finished rebooting node cloudcephosd1005.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 18:45 wm-bot2: Rebooting node cloudcephosd1005.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 18:45 wm-bot2: Finished rebooting node cloudcephosd1004.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 18:20 wm-bot2: Rebooting node cloudcephosd1004.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 18:20 wm-bot2: Finished rebooting node cloudcephosd1003.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 17:56 wm-bot2: Rebooting node cloudcephosd1003.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 17:56 wm-bot2: Finished rebooting node cloudcephosd1002.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 17:33 wm-bot2: Rebooting node cloudcephosd1002.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 17:33 wm-bot2: Finished rebooting node cloudcephosd1001.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 17:07 wm-bot2: Rebooting node cloudcephosd1001.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 17:07 wm-bot2: Rebooting the nodes cloudcephosd1001,cloudcephosd1002,cloudcephosd1003,cloudcephosd1004,cloudcephosd1005,cloudcephosd1006,cloudcephosd1007,cloudcephosd1008,cloudcephosd1009,cloudcephosd1010,cloudcephosd1011,cloudcephosd1012,cloudcephosd1013,cloudcephosd1014,cloudcephosd1015,cloudcephosd1016,cloudcephosd1017,cloudcephosd1018,cloudcephosd1019,cloudcephosd1020,cloudcephosd1021,cloudcephosd1022,cloudcephosd1023,cloudcephosd1024,cl * 16:49 wm-bot2: Finished rebooting the nodes ['cloudcephosd2001-dev', 'cloudcephosd2002-dev', 'cloudcephosd2003-dev'] ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 16:48 wm-bot2: Finished rebooting node cloudcephosd2003-dev.codfw.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 16:46 wm-bot2: Rebooting node cloudcephosd2003-dev.codfw.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 16:46 wm-bot2: Finished rebooting node cloudcephosd2002-dev.codfw.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 16:42 wm-bot2: Rebooting node cloudcephosd2002-dev.codfw.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 16:42 wm-bot2: Finished rebooting node cloudcephosd2001-dev.codfw.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 16:40 wm-bot2: Rebooting node cloudcephosd1001.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 16:40 wm-bot2: Rebooting the nodes cloudcephosd1001,cloudcephosd1002,cloudcephosd1003,cloudcephosd1004,cloudcephosd1005,cloudcephosd1006,cloudcephosd1007,cloudcephosd1008,cloudcephosd1009,cloudcephosd1010,cloudcephosd1011,cloudcephosd1012,cloudcephosd1013,cloudcephosd1014,cloudcephosd1015,cloudcephosd1016,cloudcephosd1017,cloudcephosd1018,cloudcephosd1019,cloudcephosd1020,cloudcephosd1021,cloudcephosd1022,cloudcephosd1023,cloudcephosd1024,cl * 16:39 wm-bot2: Rebooting node cloudcephosd2001-dev.codfw.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 16:39 wm-bot2: Rebooting the nodes cloudcephosd2001-dev,cloudcephosd2002-dev,cloudcephosd2003-dev ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 15:41 wm-bot2: Rebooting node cloudcephosd1001.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 15:41 wm-bot2: Rebooting the nodes cloudcephosd1001,cloudcephosd1002,cloudcephosd1003,cloudcephosd1004,cloudcephosd1005,cloudcephosd1006,cloudcephosd1007,cloudcephosd1008,cloudcephosd1009,cloudcephosd1010,cloudcephosd1011,cloudcephosd1012,cloudcephosd1013,cloudcephosd1014,cloudcephosd1015,cloudcephosd1016,cloudcephosd1017,cloudcephosd1018,cloudcephosd1019,cloudcephosd1020,cloudcephosd1021,cloudcephosd1022,cloudcephosd1023,cloudcephosd1024,cl * 15:40 wm-bot2: Finished rebooting the nodes ['cloudcephosd2001-dev', 'cloudcephosd2002-dev', 'cloudcephosd2003-dev'] ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 15:40 wm-bot2: Finished rebooting node cloudcephosd2003-dev.codfw.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 15:37 wm-bot2: Rebooting node cloudcephosd2003-dev.codfw.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 15:37 wm-bot2: Finished rebooting node cloudcephosd2002-dev.codfw.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 15:33 wm-bot2: Rebooting node cloudcephosd2002-dev.codfw.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 15:33 wm-bot2: Finished rebooting node cloudcephosd2001-dev.codfw.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 15:30 wm-bot2: Rebooting node cloudcephosd2001-dev.codfw.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 15:30 wm-bot2: Rebooting the nodes cloudcephosd2001-dev,cloudcephosd2002-dev,cloudcephosd2003-dev ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 15:29 wm-bot2: Finished rebooting the nodes ['cloudcephmon2004-dev', 'cloudcephmon2005-dev', 'cloudcephmon2006-dev'] ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 15:29 wm-bot2: Finished rebooting node cloudcephmon2006-dev.codfw.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 15:25 wm-bot2: Rebooting node cloudcephmon2006-dev.codfw.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 15:25 wm-bot2: Finished rebooting node cloudcephmon2005-dev.codfw.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 15:22 wm-bot2: Rebooting node cloudcephmon2005-dev.codfw.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 15:22 wm-bot2: Finished rebooting node cloudcephmon2004-dev.codfw.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 15:19 wm-bot2: Rebooting node cloudcephmon2004-dev.codfw.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 15:19 wm-bot2: Rebooting the nodes cloudcephmon2004-dev,cloudcephmon2005-dev,cloudcephmon2006-dev ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 15:18 wm-bot2: Finished rebooting the nodes ['cloudcephmon1001', 'cloudcephmon1002', 'cloudcephmon1003'] ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 15:18 wm-bot2: Finished rebooting node cloudcephmon1003.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 15:15 wm-bot2: Rebooting node cloudcephmon1003.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 15:15 wm-bot2: Finished rebooting node cloudcephmon1002.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 15:11 wm-bot2: Rebooting node cloudcephmon1002.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 15:11 wm-bot2: Finished rebooting node cloudcephmon1001.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 15:08 wm-bot2: Rebooting node cloudcephmon1001.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 15:08 wm-bot2: Rebooting the nodes cloudcephmon1001,cloudcephmon1002,cloudcephmon1003 ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye === 2022-12-17 === * 07:50 taavi: deleted project packagist-mirror per https://wikitech.wikimedia.org/wiki/News/Cloud_VPS_2022_Purge#packagist-mirror === 2022-12-16 === * 19:36 volans: restarted sshd twice on bastion-restricted-eqiad1-02 to debug SSH connections for [[phab:T319401|T319401]] * 08:46 dcaro: restart designate-sink on both cloudservice hosts ([[phab:T322279|T322279]]) * 08:45 dcaro: restart designate-sink on both cloudservice hosts === 2022-12-07 === * 22:07 andrewbogott: systemctl restart libvirt-guests.service on cloudvirt1019 to get ceph/rbd working on VMS on this hypervisor === 2022-12-03 === * 19:24 taavi: restart designate-sink on both cloudservices hosts === 2022-11-30 === * 20:03 andrewbogott: changing all rabbitmq queues to quorum queues. Will be noisy! [[phab:T318816|T318816]] * 02:54 wm-bot2: Upgraded and rebooted host cloudbackup2002.codfw.wmnet - cookbook ran by andrew@bullseye === 2022-11-28 === * 13:00 wm-bot2: unset cloudvirt1043.eqiad.wmnet maintenance (aggregates: ceph) - cookbook ran by arturo@nostromo * 10:28 wm-bot2: Drained cloudvirt1043.eqiad.wmnet ([[phab:T319184|T319184]]) - cookbook ran by arturo@nostromo * 10:19 wm-bot2: Set cloudvirt cloudvirt1043.eqiad.wmnet maintenance (downtime id: bb94dd24-fef9-4c9c-8f79-{{Gerrit|b6e15023ce69}}, use this to unset) ([[phab:T319184|T319184]]) - cookbook ran by arturo@nostromo * 10:18 wm-bot2: Draining cloudvirt1043.eqiad.wmnet ([[phab:T319184|T319184]]) - cookbook ran by arturo@nostromo === 2022-11-25 === * 10:54 wm-bot2: deleted VM canary2001-dev-2 from cloudvirt2001-dev - cookbook ran by arturo@nostromo * 10:54 wm-bot2: created VM canary2001-dev-3 in cloudvirt2001-dev - cookbook ran by arturo@nostromo * 10:53 wm-bot2: Created new flavor: g3.cores1.ram1.disk20 (id:5b2ca632-2ea0-4007-9b40-{{Gerrit|4f84f8e2428b}}) - cookbook ran by arturo@nostromo * 10:46 wm-bot2: Created new flavor: g3.cores1.ram1.disk20.admin (id:fecbd56d-0969-45f3-80fd-{{Gerrit|2b463a5b6270}}) - cookbook ran by arturo@nostromo === 2022-11-24 === * 16:53 wm-bot2: deleted VM canary2001-dev-1 from cloudvirt2001-dev - cookbook ran by arturo@nostromo * 16:53 wm-bot2: created VM canary2001-dev-2 in cloudvirt2001-dev - cookbook ran by arturo@nostromo * 16:42 wm-bot2: created VM canary2003-dev-1 in cloudvirt2003-dev - cookbook ran by arturo@nostromo * 16:42 wm-bot2: created VM canary2002-dev-1 in cloudvirt2002-dev - cookbook ran by arturo@nostromo * 16:42 wm-bot2: created VM canary2001-dev-1 in cloudvirt2001-dev - cookbook ran by arturo@nostromo * 16:36 wm-bot2: Created new flavor: cloudvirt-canary-ceph (id:0d06701a-2845-4298-b2b4-{{Gerrit|fabf8b1ddcbb}}) - cookbook ran by arturo@nostromo * 13:03 wm-bot2: unset cloudvirt1044.eqiad.wmnet maintenance (aggregates: ceph) - cookbook ran by arturo@nostromo * 11:51 wm-bot2: Drained cloudvirt1044.eqiad.wmnet - cookbook ran by arturo@nostromo * 11:37 wm-bot2: Set cloudvirt cloudvirt1044.eqiad.wmnet maintenance (downtime id: 10076ecc-0f94-4d56-9bbd-{{Gerrit|bebb48bdc126}}, use this to unset) - cookbook ran by arturo@nostromo * 11:35 wm-bot2: Draining cloudvirt1044.eqiad.wmnet - cookbook ran by arturo@nostromo * 10:16 dcaro: removed ip6 dns name entry from nb for coluddb* ([[phab:T323550|T323550]]) * 09:53 dcaro: removed ip6 dns entry from nb for coluddb1013 ([[phab:T323550|T323550]]) === 2022-11-23 === * 15:00 wm-bot2: unset cloudvirt1045.eqiad.wmnet maintenance (aggregates: ceph) - cookbook ran by arturo@nostromo * 13:46 wm-bot2: Drained cloudvirt1045.eqiad.wmnet ([[phab:T319184|T319184]]) - cookbook ran by arturo@nostromo * 13:35 wm-bot2: Set cloudvirt cloudvirt1045.eqiad.wmnet maintenance (downtime id: 2386b468-0f21-4ecb-91e2-{{Gerrit|e19ace66881d}}, use this to unset) ([[phab:T319184|T319184]]) - cookbook ran by arturo@nostromo * 13:34 wm-bot2: Draining cloudvirt1045.eqiad.wmnet ([[phab:T319184|T319184]]) - cookbook ran by arturo@nostromo * 13:20 wm-bot2: unset cloudvirt1046.eqiad.wmnet maintenance (aggregates: ceph) - cookbook ran by arturo@nostromo * 12:17 wm-bot2: Drained cloudvirt1046.eqiad.wmnet ([[phab:T319184|T319184]]) - cookbook ran by arturo@nostromo * 12:04 wm-bot2: Set cloudvirt cloudvirt1046.eqiad.wmnet maintenance (downtime id: 6291d38e-c04c-4aa0-88da-{{Gerrit|2b329874a9b9}}, use this to unset) ([[phab:T319184|T319184]]) - cookbook ran by arturo@nostromo * 12:03 wm-bot2: Draining cloudvirt1046.eqiad.wmnet ([[phab:T319184|T319184]]) - cookbook ran by arturo@nostromo * 12:02 wm-bot2: unset cloudvirt1047.eqiad.wmnet maintenance (aggregates: ceph) - cookbook ran by arturo@nostromo * 11:32 arturo: [codfw1dev] created project cloudvirt-canary * 10:13 wm-bot2: Drained cloudvirt1047.eqiad.wmnet ([[phab:T319184|T319184]]) - cookbook ran by arturo@nostromo * 10:01 wm-bot2: Set cloudvirt cloudvirt1047.eqiad.wmnet maintenance (downtime id: 2c7ccc17-2be3-427d-aaec-{{Gerrit|57fadca0de5b}}, use this to unset) ([[phab:T319184|T319184]]) - cookbook ran by arturo@nostromo * 10:00 wm-bot2: Draining cloudvirt1047.eqiad.wmnet ([[phab:T319184|T319184]]) - cookbook ran by arturo@nostromo === 2022-11-22 === * 13:26 wm-bot2: unset cloudvirt1048.eqiad.wmnet maintenance (aggregates: ceph) - cookbook ran by arturo@nostromo * 12:15 wm-bot2: unset cloudvirt1049.eqiad.wmnet maintenance (aggregates: ceph) - cookbook ran by arturo@nostromo * 10:58 wm-bot2: Drained cloudvirt1049.eqiad.wmnet ([[phab:T319184|T319184]]) - cookbook ran by arturo@nostromo * 10:37 wm-bot2: Set cloudvirt cloudvirt1049.eqiad.wmnet maintenance (downtime id: 3ae0a20d-3cf0-4eba-bea6-{{Gerrit|45aa61d8ad00}}, use this to unset) ([[phab:T319184|T319184]]) - cookbook ran by arturo@nostromo * 10:36 wm-bot2: Draining cloudvirt1049.eqiad.wmnet ([[phab:T319184|T319184]]) - cookbook ran by arturo@nostromo * 10:21 wm-bot2: Unset cloudvirt cloudvirt1050.eqiad.wmnet maintenance - cookbook ran by arturo@nostromo === 2022-11-21 === * 16:19 wm-bot2: Drained cloudvirt1050.eqiad.wmnet ([[phab:T319184|T319184]]) - cookbook ran by arturo@nostromo * 16:11 wm-bot2: Set cloudvirt cloudvirt1050.eqiad.wmnet maintenance (downtime id: 0b154a93-a9d3-4ac7-bcfc-{{Gerrit|b67c49abe97b}}, use this to unset) ([[phab:T319184|T319184]]) - cookbook ran by arturo@nostromo * 16:09 wm-bot2: Draining cloudvirt1050.eqiad.wmnet ([[phab:T319184|T319184]]) - cookbook ran by arturo@nostromo * 16:07 wm-bot2: Unset cloudvirt cloudvirt1051.eqiad.wmnet maintenance - cookbook ran by arturo@nostromo * 15:18 wm-bot2: Drained cloudvirt1051.eqiad.wmnet ([[phab:T319184|T319184]]) - cookbook ran by arturo@nostromo * 15:02 wm-bot2: Set cloudvirt cloudvirt1051.eqiad.wmnet maintenance (downtime id: 1de1174b-ec46-47a3-911c-{{Gerrit|b5808ce37028}}, use this to unset) ([[phab:T319184|T319184]]) - cookbook ran by arturo@nostromo * 15:00 wm-bot2: Draining cloudvirt1051.eqiad.wmnet ([[phab:T319184|T319184]]) - cookbook ran by arturo@nostromo * 14:57 wm-bot2: Unset cloudvirt cloudvirt1052.eqiad.wmnet maintenance - cookbook ran by arturo@nostromo * 12:53 wm-bot2: Drained cloudvirt1052.eqiad.wmnet ([[phab:T319184|T319184]]) - cookbook ran by arturo@nostromo * 12:30 wm-bot2: Set cloudvirt cloudvirt1052.eqiad.wmnet maintenance (downtime id: 30db8a9b-08db-456f-8106-{{Gerrit|53188ff5f989}}, use this to unset) ([[phab:T319184|T319184]]) - cookbook ran by arturo@nostromo * 12:29 wm-bot2: Draining cloudvirt1052.eqiad.wmnet ([[phab:T319184|T319184]]) - cookbook ran by arturo@nostromo * 12:15 wm-bot2: Unset cloudvirt cloudvirt1053.eqiad.wmnet maintenance - cookbook ran by arturo@nostromo * 10:13 arturo: drained cloudvirt1053 in preparation for reimage (was spare anyway) === 2022-11-18 === * 13:37 arturo: [codfw1dev] reimaged cloudvirt2001-dev and cloudvirt2002-dev * 11:36 wm-bot2: Set cloudvirt cloudvirt2001-dev.codfw.wmnet maintenance (downtime id: 2fb9df34-fc2d-45b9-b21f-{{Gerrit|c0ec09008b92}}, use this to unset) - cookbook ran by arturo@nostromo * 11:36 wm-bot2: Draining cloudvirt2001-dev.codfw.wmnet - cookbook ran by arturo@nostromo === 2022-11-16 === * 20:07 wm-bot2: Upgraded and rebooted host cloudcontrol2004-dev.wikimedia.org - cookbook ran by andrew@bullseye * 19:56 wm-bot2: Upgraded and rebooted host cloudservices2004-dev.wikimedia.org - cookbook ran by andrew@bullseye * 19:16 wm-bot2: Upgraded and rebooted host cloudcontrol1007.wikimedia.org - cookbook ran by andrew@bullseye * 18:59 wm-bot2: Upgraded and rebooted host cloudcontrol1007.wikimedia.org - cookbook ran by andrew@bullseye * 18:51 wm-bot2: Upgraded and rebooted host cloudcontrol2005-dev.wikimedia.org - cookbook ran by andrew@bullseye * 12:52 arturo: failovered cloudgw1002 into cloudgw1001 for reimage, IRC bots were briefly disconnected === 2022-11-14 === * 20:22 wm-bot2: Upgraded and rebooted host cloudnet1005.eqiad.wmnet - cookbook ran by andrew@bullseye * 20:07 wm-bot2: Upgraded and rebooted host cloudcontrol1007.wikimedia.org - cookbook ran by andrew@bullseye * 19:55 wm-bot2: Upgraded and rebooted host cloudcontrol1006.wikimedia.org - cookbook ran by andrew@bullseye * 19:42 wm-bot2: Upgraded and rebooted host cloudcontrol1005.wikimedia.org - cookbook ran by andrew@bullseye * 19:28 andrewbogott: beginning OpenStack upgrade in eqiad1 -- [[phab:T305828|T305828]] * 19:22 wm-bot2: Upgraded and rebooted host cloudbackup1002-dev.eqiad.wmnet - cookbook ran by andrew@bullseye * 19:14 wm-bot2: Upgraded and rebooted host cloudbackup1001-dev.eqiad.wmnet - cookbook ran by andrew@bullseye * 19:08 wm-bot2: Upgraded and rebooted host cloudcontrol2001-dev.wikimedia.org - cookbook ran by andrew@bullseye * 12:52 arturo: cleanup old network vlan interface names from /etc/network/interfaces in cloudnet1005/1006 === 2022-11-11 === * 13:24 wm-bot2: Set cloudvirt cloudvirt2003-dev.codfw.wmnet maintenance (downtime id: edad3915-b7c6-4b23-bb9c-{{Gerrit|ab13b04a41c5}}, use this to unset) ([[phab:T319184|T319184]]) - cookbook ran by arturo@nostromo * 13:23 wm-bot2: Draining cloudvirt2003-dev.codfw.wmnet ([[phab:T319184|T319184]]) - cookbook ran by arturo@nostromo * 13:17 wm-bot2: Set cloudvirt cloudvirt2003-dev.codfw.wmnet maintenance (downtime id: a09a9868-8aae-4dc0-8510-{{Gerrit|a3923b703060}}, use this to unset) ([[phab:T319184|T319184]]) - cookbook ran by arturo@nostromo * 13:16 wm-bot2: Draining cloudvirt2003-dev.codfw.wmnet ([[phab:T319184|T319184]]) - cookbook ran by arturo@nostromo === 2022-11-10 === * 16:19 wm-bot2: Set cloudvirt 'cloudvirt2002-dev.codfw.wmnet' maintenance (downtime id: 346013ec-ce4e-497e-ad65-{{Gerrit|2d215b14998c}}, use this to unset). ([[phab:T319184|T319184]]) - cookbook ran by arturo@nostromo * 16:18 wm-bot2: Draining 'cloudvirt2002-dev.codfw.wmnet'. ([[phab:T319184|T319184]]) - cookbook ran by arturo@nostromo * 16:13 wm-bot2: Set cloudvirt 'cloudvirt2002-dev.codfw.wmnet' maintenance (downtime id: a40a8487-22ee-4bc6-bbe4-{{Gerrit|694a615e3bf5}}, use this to unset). ([[phab:T319184|T319184]]) - cookbook ran by arturo@nostromo === 2022-11-08 === * 11:17 taavi: backfilling security groups for metricsinfra access on all projects [[phab:T288108|T288108]] === 2022-11-07 === * 21:01 wm-bot2: Upgraded and rebooted host cloudservices1004.wikimedia.org ([[phab:T305828|T305828]]) - cookbook ran by andrew@bullseye * 20:50 wm-bot2: Upgraded and rebooted host cloudservices1005.wikimedia.org ([[phab:T305828|T305828]]) - cookbook ran by andrew@bullseye * 20:40 andrewbogott: upgrading eqiad1 designate to version 'yoga' === 2022-11-04 === * 17:57 andrewbogott: removing cinderv2 API endpoints from keystone catalog; this is deprecated and removed in Yoga. prep for [[phab:T305828|T305828]] === 2022-11-03 === * 19:24 wm-bot2: Upgraded and rebooted host cloudbackup1002-dev.eqiad.wmnet ([[phab:T305828|T305828]]) - cookbook ran by andrew@bullseye * 19:19 wm-bot2: Upgraded and rebooted host cloudbackup1001-dev.eqiad.wmnet ([[phab:T305828|T305828]]) - cookbook ran by andrew@bullseye * 00:08 wm-bot2: Upgraded and rebooted host cloudcontrol2004-dev.wikimedia.org ([[phab:T305828|T305828]]) - cookbook ran by andrew@bullseye === 2022-11-02 === * 23:39 wm-bot2: Upgraded and rebooted host cloudbackup1001-dev.eqiad.wmnet ([[phab:T305828|T305828]]) - cookbook ran by andrew@bullseye * 23:38 wm-bot2: Upgraded and rebooted host cloudnet2006-dev.codfw.wmnet ([[phab:T305828|T305828]]) - cookbook ran by andrew@bullseye * 23:29 wm-bot2: Upgraded and rebooted host cloudnet2005-dev.codfw.wmnet ([[phab:T305828|T305828]]) - cookbook ran by andrew@bullseye * 23:13 wm-bot2: Upgraded and rebooted host cloudcontrol2004-dev.wikimedia.org ([[phab:T305828|T305828]]) - cookbook ran by andrew@bullseye * 23:00 wm-bot2: Upgraded and rebooted host cloudcontrol2001-dev.wikimedia.org ([[phab:T305828|T305828]]) - cookbook ran by andrew@bullseye * 22:54 wm-bot2: Upgraded and rebooted host cloudcontrol2005-dev.wikimedia.org - cookbook ran by andrew@bullseye * 20:29 wm-bot2: Upgraded and rebooted host cloudcontrol2001-dev.wikimedia.org - cookbook ran by andrew@bullseye === 2022-10-31 === * 13:09 arturo: restart keepalived on all 4 cloudgw servers to run them with `-D` in /etc/default/keepalived to further debug [[phab:T320975|T320975]] === 2022-10-26 === * 16:18 wm-bot2: Created new flavor: g3.cores1.ram1.disk20 (id:bf48880d-0c1b-4c2a-8e8b-{{Gerrit|778d28b16561}}) ([[phab:T319446|T319446]]) - cookbook ran by dcaro@vulcanus * 09:34 taavi: running wmcs-puppetcertleaks in delete mode * 09:09 taavi: running wmcs-novastats-dnsleaks in delete mode === 2022-10-25 === * 16:03 arturo: [codfw1dev] [[phab:T321220|T321220]] root@cloudcontrol2001-dev:~# openstack subnet create magnum --no-dhcp --network 57017d7c-3817-429a-8aa3-{{Gerrit|b028de82cdcc}} --ip-version 4 --gateway auto --subnet-range 192.168.0.0/24 * 14:38 arturo: [codfw1dev] restart neutron-l3-agent in cloudnet2005-dev, it was dead after rabbit connectivity problems === 2022-10-24 === * 18:50 wm-bot2: Rebooting node cloudcephmon1002.eqiad.wmnet - cookbook ran by andrew@bullseye * 18:50 wm-bot2: Finished rebooting node cloudcephmon1001.eqiad.wmnet - cookbook ran by andrew@bullseye * 18:47 wm-bot2: Rebooting node cloudcephmon1001.eqiad.wmnet - cookbook ran by andrew@bullseye * 18:47 wm-bot2: Rebooting the nodes cloudcephmon1001,cloudcephmon1002,cloudcephmon1003 - cookbook ran by andrew@bullseye * 18:38 wm-bot2: Rebooting the nodes cloudcephmon1001,cloudcephmon1002,cloudcephmon1003 - cookbook ran by andrew@bullseye * 18:21 wm-bot2: Rebooting node cloudcephosd1001.eqiad.wmnet - cookbook ran by andrew@bullseye * 18:21 wm-bot2: Rebooting the nodes cloudcephosd1001,cloudcephosd1002,cloudcephosd1003,cloudcephosd1004,cloudcephosd1005,cloudcephosd1006,cloudcephosd1007,cloudcephosd1008,cloudcephosd1009,cloudcephosd1010,cloudcephosd1011,cloudcephosd1012,cloudcephosd1013,cloudcephosd1014,cloudcephosd1015,cloudcephosd1016,cloudcephosd1017,cloudcephosd1018,cloudcephosd1019,cloudcephosd1020,cloudcephosd1021,cloudcephosd1022,cloudcephosd1023,cloudcephosd1024,cl === 2022-10-20 === * 23:23 wm-bot2: Safe reboot of 'cloudvirt1021.eqiad.wmnet' finished successfully. - cookbook ran by andrew@bullseye * 23:23 wm-bot2: Unset cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance. - cookbook ran by andrew@bullseye * 23:20 wm-bot2: Safe reboot of 'cloudvirt1022.eqiad.wmnet' finished successfully. - cookbook ran by andrew@bullseye * 23:20 wm-bot2: Unset cloudvirt 'cloudvirt1022.eqiad.wmnet' maintenance. - cookbook ran by andrew@bullseye * 23:19 wm-bot2: Drained 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 23:16 wm-bot2: Drained 'cloudvirt1022.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 23:07 wm-bot2: Safe reboot of 'cloudvirt1024.eqiad.wmnet' finished successfully. - cookbook ran by andrew@bullseye * 23:07 wm-bot2: Unset cloudvirt 'cloudvirt1024.eqiad.wmnet' maintenance. - cookbook ran by andrew@bullseye * 23:06 wm-bot2: Safe reboot of 'cloudvirt1025.eqiad.wmnet' finished successfully. - cookbook ran by andrew@bullseye * 23:06 wm-bot2: Unset cloudvirt 'cloudvirt1025.eqiad.wmnet' maintenance. - cookbook ran by andrew@bullseye * 23:03 wm-bot2: Drained 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 23:03 wm-bot2: Set cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance (downtime id: 101c41d1-d65d-4088-b6d2-{{Gerrit|eac859e45ef8}}, use this to unset). - cookbook ran by andrew@bullseye * 23:03 wm-bot2: Drained 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 23:03 wm-bot2: Draining 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 23:03 wm-bot2: Safe rebooting 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 23:03 wm-bot2: Safe reboot of 'cloudvirt1026.eqiad.wmnet' finished successfully. - cookbook ran by andrew@bullseye * 23:01 wm-bot2: Unset cloudvirt 'cloudvirt1026.eqiad.wmnet' maintenance. - cookbook ran by andrew@bullseye * 22:57 wm-bot2: Drained 'cloudvirt1026.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 22:51 wm-bot2: Set cloudvirt 'cloudvirt1022.eqiad.wmnet' maintenance (downtime id: 7197ce34-cc57-4677-a1a9-{{Gerrit|05e20cd0dd80}}, use this to unset). - cookbook ran by andrew@bullseye * 22:50 wm-bot2: Draining 'cloudvirt1022.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 22:50 wm-bot2: Safe rebooting 'cloudvirt1022.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 22:50 wm-bot2: Safe reboot of 'cloudvirt1017.eqiad.wmnet' finished successfully. - cookbook ran by andrew@bullseye * 22:50 wm-bot2: Unset cloudvirt 'cloudvirt1017.eqiad.wmnet' maintenance. - cookbook ran by andrew@bullseye * 22:46 wm-bot2: Drained 'cloudvirt1017.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 22:38 wm-bot2: Set cloudvirt 'cloudvirt1024.eqiad.wmnet' maintenance (downtime id: 68e4cc1d-9def-444e-84a7-{{Gerrit|21b0b5adfa72}}, use this to unset). - cookbook ran by andrew@bullseye * 22:37 wm-bot2: Set cloudvirt 'cloudvirt1025.eqiad.wmnet' maintenance (downtime id: 73388c1e-369b-40af-a2c1-{{Gerrit|96936128c324}}, use this to unset). - cookbook ran by andrew@bullseye * 22:37 wm-bot2: Set cloudvirt 'cloudvirt1026.eqiad.wmnet' maintenance (downtime id: 12feeca2-ea59-48e9-9235-{{Gerrit|1cecf5b384cd}}, use this to unset). - cookbook ran by andrew@bullseye * 22:37 wm-bot2: Draining 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 22:37 wm-bot2: Safe rebooting 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 22:37 wm-bot2: Draining 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 22:37 wm-bot2: Safe reboot of 'cloudvirt1029.eqiad.wmnet' finished successfully. - cookbook ran by andrew@bullseye * 22:37 wm-bot2: Safe rebooting 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 22:36 wm-bot2: Unset cloudvirt 'cloudvirt1029.eqiad.wmnet' maintenance. - cookbook ran by andrew@bullseye * 22:36 wm-bot2: Draining 'cloudvirt1026.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 22:36 wm-bot2: Safe rebooting 'cloudvirt1026.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 22:35 wm-bot2: Safe reboot of 'cloudvirt1030.eqiad.wmnet' finished successfully. - cookbook ran by andrew@bullseye * 22:35 wm-bot2: Unset cloudvirt 'cloudvirt1030.eqiad.wmnet' maintenance. - cookbook ran by andrew@bullseye * 22:35 wm-bot2: Safe reboot of 'cloudvirt1027.eqiad.wmnet' finished successfully. - cookbook ran by andrew@bullseye * 22:35 wm-bot2: Unset cloudvirt 'cloudvirt1027.eqiad.wmnet' maintenance. - cookbook ran by andrew@bullseye * 22:34 wm-bot2: Set cloudvirt 'cloudvirt1017.eqiad.wmnet' maintenance (downtime id: 1b34587f-2770-4d2e-bb31-{{Gerrit|c8ad14632d39}}, use this to unset). - cookbook ran by andrew@bullseye * 22:34 wm-bot2: Drained 'cloudvirt1029.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 22:34 wm-bot2: Draining 'cloudvirt1017.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 22:33 wm-bot2: Safe rebooting 'cloudvirt1017.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 22:33 wm-bot2: Drained 'cloudvirt1030.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 22:32 wm-bot2: Drained 'cloudvirt1027.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 22:28 wm-bot2: Safe reboot of 'cloudvirt1032.eqiad.wmnet' finished successfully. - cookbook ran by andrew@bullseye * 22:28 wm-bot2: Unset cloudvirt 'cloudvirt1032.eqiad.wmnet' maintenance. - cookbook ran by andrew@bullseye * 22:24 wm-bot2: Drained 'cloudvirt1032.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 22:13 wm-bot2: Set cloudvirt 'cloudvirt1027.eqiad.wmnet' maintenance (downtime id: e75b00eb-7f58-4821-8139-{{Gerrit|3dfc6e97a92a}}, use this to unset). - cookbook ran by andrew@bullseye * 22:12 wm-bot2: Set cloudvirt 'cloudvirt1029.eqiad.wmnet' maintenance (downtime id: 9151511e-bcc0-4367-a976-{{Gerrit|7cff060308aa}}, use this to unset). - cookbook ran by andrew@bullseye * 22:12 wm-bot2: Set cloudvirt 'cloudvirt1030.eqiad.wmnet' maintenance (downtime id: 241c0bd9-c3cf-40e4-95dc-{{Gerrit|9b51d9823fe9}}, use this to unset). - cookbook ran by andrew@bullseye * 22:12 wm-bot2: Set cloudvirt 'cloudvirt1032.eqiad.wmnet' maintenance (downtime id: 7a53c274-1d5f-4c94-9a0d-{{Gerrit|721c3e8f7239}}, use this to unset). - cookbook ran by andrew@bullseye * 22:12 wm-bot2: Draining 'cloudvirt1027.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 22:12 wm-bot2: Safe rebooting 'cloudvirt1027.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 22:11 wm-bot2: Draining 'cloudvirt1029.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 22:11 wm-bot2: Safe rebooting 'cloudvirt1029.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 22:11 wm-bot2: Draining 'cloudvirt1030.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 22:11 wm-bot2: Safe rebooting 'cloudvirt1030.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 22:11 wm-bot2: Draining 'cloudvirt1032.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 22:11 wm-bot2: Safe rebooting 'cloudvirt1032.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 22:10 wm-bot2: Safe reboot of 'cloudvirt1031.eqiad.wmnet' finished successfully. - cookbook ran by andrew@bullseye * 22:10 wm-bot2: Unset cloudvirt 'cloudvirt1031.eqiad.wmnet' maintenance. - cookbook ran by andrew@bullseye * 22:06 wm-bot2: Drained 'cloudvirt1031.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:58 wm-bot2: Safe reboot of 'cloudvirt1034.eqiad.wmnet' finished successfully. - cookbook ran by andrew@bullseye * 21:58 wm-bot2: Unset cloudvirt 'cloudvirt1034.eqiad.wmnet' maintenance. - cookbook ran by andrew@bullseye * 21:58 wm-bot2: Safe reboot of 'cloudvirt1035.eqiad.wmnet' finished successfully. - cookbook ran by andrew@bullseye * 21:58 wm-bot2: Unset cloudvirt 'cloudvirt1035.eqiad.wmnet' maintenance. - cookbook ran by andrew@bullseye * 21:55 wm-bot2: Safe reboot of 'cloudvirt1033.eqiad.wmnet' finished successfully. - cookbook ran by andrew@bullseye * 21:55 wm-bot2: Unset cloudvirt 'cloudvirt1033.eqiad.wmnet' maintenance. - cookbook ran by andrew@bullseye * 21:54 wm-bot2: Drained 'cloudvirt1034.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:54 wm-bot2: Drained 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:51 wm-bot2: Drained 'cloudvirt1033.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:38 wm-bot2: Set cloudvirt 'cloudvirt1031.eqiad.wmnet' maintenance (downtime id: 0b75b44c-1efe-4edf-8f5e-{{Gerrit|42a67e8d3b13}}, use this to unset). - cookbook ran by andrew@bullseye * 21:38 wm-bot2: Set cloudvirt 'cloudvirt1034.eqiad.wmnet' maintenance (downtime id: aeb9c3a6-961a-4873-85d4-{{Gerrit|929248aebb8b}}, use this to unset). - cookbook ran by andrew@bullseye * 21:38 wm-bot2: Set cloudvirt 'cloudvirt1035.eqiad.wmnet' maintenance (downtime id: eb14cc04-3293-4cf8-a46f-{{Gerrit|1f87d8b0bcc4}}, use this to unset). - cookbook ran by andrew@bullseye * 21:37 wm-bot2: Set cloudvirt 'cloudvirt1033.eqiad.wmnet' maintenance (downtime id: 84a3cd99-2643-495a-86b6-{{Gerrit|7ae80b11a30d}}, use this to unset). - cookbook ran by andrew@bullseye * 21:37 wm-bot2: Draining 'cloudvirt1031.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:37 wm-bot2: Safe rebooting 'cloudvirt1031.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:37 wm-bot2: Draining 'cloudvirt1039.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:37 wm-bot2: Safe rebooting 'cloudvirt1039.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:37 wm-bot2: Draining 'cloudvirt1034.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:37 wm-bot2: Safe rebooting 'cloudvirt1034.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:37 wm-bot2: Draining 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:37 wm-bot2: Safe rebooting 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:36 wm-bot2: Draining 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:36 wm-bot2: Safe rebooting 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:36 wm-bot2: Draining 'cloudvirt1034.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:36 wm-bot2: Safe rebooting 'cloudvirt1034.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:36 wm-bot2: Draining 'cloudvirt1033.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:36 wm-bot2: Safe rebooting 'cloudvirt1033.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:36 wm-bot2: Draining 'cloudvirt1032.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:36 wm-bot2: Safe rebooting 'cloudvirt1032.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:32 wm-bot2: Safe reboot of 'cloudvirt1036.eqiad.wmnet' finished successfully. - cookbook ran by andrew@bullseye * 21:32 wm-bot2: Unset cloudvirt 'cloudvirt1036.eqiad.wmnet' maintenance. - cookbook ran by andrew@bullseye * 21:28 wm-bot2: Drained 'cloudvirt1036.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:28 wm-bot2: Set cloudvirt 'cloudvirt1036.eqiad.wmnet' maintenance (downtime id: 01a0b69c-2e27-4331-9635-{{Gerrit|403a668dac29}}, use this to unset). - cookbook ran by andrew@bullseye * 21:27 wm-bot2: Draining 'cloudvirt1036.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:27 wm-bot2: Safe rebooting 'cloudvirt1036.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:23 wm-bot2: Set cloudvirt 'cloudvirt1036.eqiad.wmnet' maintenance (downtime id: ea29f65f-6d86-4568-8db9-{{Gerrit|a85faa827447}}, use this to unset). - cookbook ran by andrew@bullseye * 21:22 wm-bot2: Draining 'cloudvirt1036.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:22 wm-bot2: Safe rebooting 'cloudvirt1036.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:19 wm-bot2: Safe reboot of 'cloudvirt1038.eqiad.wmnet' finished successfully. - cookbook ran by andrew@bullseye * 21:19 wm-bot2: Unset cloudvirt 'cloudvirt1038.eqiad.wmnet' maintenance. - cookbook ran by andrew@bullseye * 21:19 wm-bot2: Safe reboot of 'cloudvirt1037.eqiad.wmnet' finished successfully. - cookbook ran by andrew@bullseye * 21:19 wm-bot2: Unset cloudvirt 'cloudvirt1037.eqiad.wmnet' maintenance. - cookbook ran by andrew@bullseye * 21:18 wm-bot2: Safe reboot of 'cloudvirt1039.eqiad.wmnet' finished successfully. - cookbook ran by andrew@bullseye * 21:18 wm-bot2: Unset cloudvirt 'cloudvirt1039.eqiad.wmnet' maintenance. - cookbook ran by andrew@bullseye * 21:16 wm-bot2: Drained 'cloudvirt1038.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:15 wm-bot2: Drained 'cloudvirt1037.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:14 wm-bot2: Drained 'cloudvirt1039.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:01 wm-bot2: Set cloudvirt 'cloudvirt1036.eqiad.wmnet' maintenance (downtime id: d6b6bb4d-c7c0-4bd2-9a02-{{Gerrit|dee9a5ec6a3e}}, use this to unset). - cookbook ran by andrew@bullseye * 21:01 wm-bot2: Set cloudvirt 'cloudvirt1038.eqiad.wmnet' maintenance (downtime id: 4e109149-7d67-403f-8b21-{{Gerrit|f829235ea491}}, use this to unset). - cookbook ran by andrew@bullseye * 21:01 wm-bot2: Set cloudvirt 'cloudvirt1039.eqiad.wmnet' maintenance (downtime id: 341f4b8c-6d6e-4190-9f2d-{{Gerrit|a1a1483abadd}}, use this to unset). - cookbook ran by andrew@bullseye * 21:01 wm-bot2: Set cloudvirt 'cloudvirt1037.eqiad.wmnet' maintenance (downtime id: f76e8f14-920e-4d51-a44e-{{Gerrit|321a55dcb0c5}}, use this to unset). - cookbook ran by andrew@bullseye * 21:00 wm-bot2: Draining 'cloudvirt1039.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:00 wm-bot2: Safe rebooting 'cloudvirt1039.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:00 wm-bot2: Draining 'cloudvirt1038.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:00 wm-bot2: Safe rebooting 'cloudvirt1038.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:00 wm-bot2: Draining 'cloudvirt1037.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:00 wm-bot2: Safe rebooting 'cloudvirt1037.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:00 wm-bot2: Draining 'cloudvirt1036.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:00 wm-bot2: Safe rebooting 'cloudvirt1036.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 20:56 wm-bot2: Safe reboot of 'cloudvirt1042.eqiad.wmnet' finished successfully. - cookbook ran by andrew@bullseye * 20:56 wm-bot2: Unset cloudvirt 'cloudvirt1042.eqiad.wmnet' maintenance. - cookbook ran by andrew@bullseye * 20:54 wm-bot2: Unset cloudvirt 'cloudvirt1041.eqiad.wmnet' maintenance. - cookbook ran by andrew@bullseye * 20:35 wm-bot2: Drained 'cloudvirt1040.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 20:34 wm-bot2: Drained 'cloudvirt1041.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 20:32 wm-bot2: Set cloudvirt 'cloudvirt1042.eqiad.wmnet' maintenance (downtime id: 0b0d8090-4d31-4d43-8545-{{Gerrit|c25209e2ef58}}, use this to unset). - cookbook ran by andrew@bullseye * 20:31 wm-bot2: Draining 'cloudvirt1042.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 20:31 wm-bot2: Safe rebooting 'cloudvirt1042.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 20:21 wm-bot2: Set cloudvirt 'cloudvirt1043.eqiad.wmnet' maintenance (downtime id: 3d928554-a79c-419a-bcc4-{{Gerrit|c8d63791d8e7}}, use this to unset). - cookbook ran by andrew@bullseye * 20:21 wm-bot2: Set cloudvirt 'cloudvirt1042.eqiad.wmnet' maintenance (downtime id: 3d3c8aa5-45f7-429e-a9d7-{{Gerrit|b5181b29dc13}}, use this to unset). - cookbook ran by andrew@bullseye * 20:21 wm-bot2: Set cloudvirt 'cloudvirt1041.eqiad.wmnet' maintenance (downtime id: b9027020-8a90-4c40-b0e6-{{Gerrit|6c8767f87917}}, use this to unset). - cookbook ran by andrew@bullseye * 20:21 wm-bot2: Set cloudvirt 'cloudvirt1040.eqiad.wmnet' maintenance (downtime id: f0d60fab-64e9-4da2-826c-{{Gerrit|229d31d0fbc2}}, use this to unset). - cookbook ran by andrew@bullseye * 20:21 wm-bot2: Draining 'cloudvirt1043.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 20:21 wm-bot2: Safe rebooting 'cloudvirt1043.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 20:20 wm-bot2: Draining 'cloudvirt1042.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 20:20 wm-bot2: Safe rebooting 'cloudvirt1042.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 20:20 wm-bot2: Draining 'cloudvirt1041.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 20:20 wm-bot2: Safe rebooting 'cloudvirt1041.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 20:20 wm-bot2: Draining 'cloudvirt1040.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 20:20 wm-bot2: Safe rebooting 'cloudvirt1040.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 20:20 wm-bot2: Safe reboot of 'cloudvirt1044.eqiad.wmnet' finished successfully. - cookbook ran by andrew@bullseye * 20:20 wm-bot2: Unset cloudvirt 'cloudvirt1044.eqiad.wmnet' maintenance. - cookbook ran by andrew@bullseye * 20:19 wm-bot2: Safe reboot of 'cloudvirt1047.eqiad.wmnet' finished successfully. - cookbook ran by andrew@bullseye * 20:19 wm-bot2: Unset cloudvirt 'cloudvirt1047.eqiad.wmnet' maintenance. - cookbook ran by andrew@bullseye * 20:18 wm-bot2: Safe reboot of 'cloudvirt1046.eqiad.wmnet' finished successfully. - cookbook ran by andrew@bullseye * 20:18 wm-bot2: Unset cloudvirt 'cloudvirt1046.eqiad.wmnet' maintenance. - cookbook ran by andrew@bullseye * 20:16 wm-bot2: Drained 'cloudvirt1044.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 20:15 wm-bot2: Drained 'cloudvirt1047.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 20:15 wm-bot2: Safe reboot of 'cloudvirt1045.eqiad.wmnet' finished successfully. - cookbook ran by andrew@bullseye * 20:15 wm-bot2: Unset cloudvirt 'cloudvirt1045.eqiad.wmnet' maintenance. - cookbook ran by andrew@bullseye * 20:14 wm-bot2: Drained 'cloudvirt1046.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 20:11 wm-bot2: Drained 'cloudvirt1045.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 19:55 wm-bot2: Set cloudvirt 'cloudvirt1046.eqiad.wmnet' maintenance (downtime id: b65da7e2-9fd2-4a1e-8dd4-{{Gerrit|88ca65936ae4}}, use this to unset). - cookbook ran by andrew@bullseye * 19:55 wm-bot2: Set cloudvirt 'cloudvirt1045.eqiad.wmnet' maintenance (downtime id: cbbf114c-9a4c-4dd2-9b58-{{Gerrit|50da8bd896ca}}, use this to unset). - cookbook ran by andrew@bullseye * 19:55 wm-bot2: Set cloudvirt 'cloudvirt1044.eqiad.wmnet' maintenance (downtime id: 5969beaf-72bf-4af9-ae4d-{{Gerrit|8e5c3331e78e}}, use this to unset). - cookbook ran by andrew@bullseye * 19:55 wm-bot2: Set cloudvirt 'cloudvirt1047.eqiad.wmnet' maintenance (downtime id: b8f7389d-79b4-411d-b3d5-{{Gerrit|ec8f92eba101}}, use this to unset). - cookbook ran by andrew@bullseye * 19:54 wm-bot2: Draining 'cloudvirt1047.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 19:54 wm-bot2: Safe rebooting 'cloudvirt1047.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 19:54 wm-bot2: Draining 'cloudvirt1046.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 19:54 wm-bot2: Safe rebooting 'cloudvirt1046.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 19:54 wm-bot2: Draining 'cloudvirt1045.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 19:54 wm-bot2: Safe rebooting 'cloudvirt1045.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 19:54 wm-bot2: Draining 'cloudvirt1044.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 19:54 wm-bot2: Safe rebooting 'cloudvirt1044.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 19:53 wm-bot2: Safe reboot of 'cloudvirt1051.eqiad.wmnet' finished successfully. - cookbook ran by andrew@bullseye * 19:53 wm-bot2: Unset cloudvirt 'cloudvirt1051.eqiad.wmnet' maintenance. - cookbook ran by andrew@bullseye * 19:51 wm-bot2: Safe reboot of 'cloudvirt1049.eqiad.wmnet' finished successfully. - cookbook ran by andrew@bullseye * 19:51 wm-bot2: Unset cloudvirt 'cloudvirt1049.eqiad.wmnet' maintenance. - cookbook ran by andrew@bullseye * 19:50 wm-bot2: Safe reboot of 'cloudvirt1048.eqiad.wmnet' finished successfully. - cookbook ran by andrew@bullseye * 19:50 wm-bot2: Unset cloudvirt 'cloudvirt1048.eqiad.wmnet' maintenance. - cookbook ran by andrew@bullseye * 19:49 wm-bot2: Drained 'cloudvirt1051.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 19:36 wm-bot2: Set cloudvirt 'cloudvirt1050.eqiad.wmnet' maintenance (downtime id: c5dbfa7b-72fc-4156-8257-{{Gerrit|af224a725b78}}, use this to unset). - cookbook ran by andrew@bullseye * 19:36 wm-bot2: Draining 'cloudvirt1050.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 19:36 wm-bot2: Safe rebooting 'cloudvirt1050.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 19:33 wm-bot2: Draining 'cloudvirt1050.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 19:33 wm-bot2: Safe rebooting 'cloudvirt1050.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 19:30 wm-bot2: Set cloudvirt 'cloudvirt1048.eqiad.wmnet' maintenance (downtime id: c1a3e92c-cb04-46e4-ac5d-{{Gerrit|784c79601b05}}, use this to unset). - cookbook ran by andrew@bullseye * 19:30 wm-bot2: Set cloudvirt 'cloudvirt1049.eqiad.wmnet' maintenance (downtime id: 62fdc4ee-c8d5-4892-aeb9-{{Gerrit|a681cdcbc84b}}, use this to unset). - cookbook ran by andrew@bullseye * 19:29 wm-bot2: Set cloudvirt 'cloudvirt1050.eqiad.wmnet' maintenance (downtime id: 9ec407cd-5c00-407a-a57d-{{Gerrit|794f6c68f947}}, use this to unset). - cookbook ran by andrew@bullseye * 19:29 wm-bot2: Set cloudvirt 'cloudvirt1051.eqiad.wmnet' maintenance (downtime id: 712683ea-d58f-484b-8efb-{{Gerrit|89d1c21aa0e3}}, use this to unset). - cookbook ran by andrew@bullseye * 19:29 wm-bot2: Draining 'cloudvirt1048.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 19:29 wm-bot2: Safe rebooting 'cloudvirt1048.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 19:29 wm-bot2: Draining 'cloudvirt1049.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 19:29 wm-bot2: Safe rebooting 'cloudvirt1049.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 19:28 wm-bot2: Draining 'cloudvirt1050.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 19:28 wm-bot2: Safe rebooting 'cloudvirt1050.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 19:28 wm-bot2: Draining 'cloudvirt1051.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 19:28 wm-bot2: Safe rebooting 'cloudvirt1051.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 19:25 wm-bot2: Safe reboot of 'cloudvirt1052.eqiad.wmnet' finished successfully. - cookbook ran by andrew@bullseye * 19:25 wm-bot2: Unset cloudvirt 'cloudvirt1052.eqiad.wmnet' maintenance. - cookbook ran by andrew@bullseye * 19:21 wm-bot2: Drained 'cloudvirt1052.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 19:04 wm-bot2: Set cloudvirt 'cloudvirt1052.eqiad.wmnet' maintenance (downtime id: 18fae8d8-7353-4f67-90d7-{{Gerrit|8df9b3fb1ccb}}, use this to unset). - cookbook ran by andrew@bullseye * 19:03 wm-bot2: Draining 'cloudvirt1052.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 19:03 wm-bot2: Safe rebooting 'cloudvirt1052.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 19:01 wm-bot2: Set cloudvirt 'cloudvirt1053.eqiad.wmnet' maintenance (downtime id: 3d3cffa3-abc6-4901-83d9-{{Gerrit|3bcc4b02fd3c}}, use this to unset). - cookbook ran by andrew@bullseye * 19:00 wm-bot2: Draining 'cloudvirt1053.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 19:00 wm-bot2: Safe rebooting 'cloudvirt1053.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 18:54 wm-bot2: Draining 'cloudvirt1053.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 18:54 wm-bot2: Safe rebooting 'cloudvirt1053.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 18:52 wm-bot2: Draining 'cloudvirt1053.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 18:52 wm-bot2: Safe rebooting 'cloudvirt1053.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 18:51 wm-bot2: Draining 'cloudvirt1053.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 18:51 wm-bot2: Safe rebooting 'cloudvirt1053.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 18:51 wm-bot2: Draining 'cloudvirt1053.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 18:51 wm-bot2: Safe rebooting 'cloudvirt1053.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 18:51 wm-bot2: Draining 'cloudvirt1053.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 18:50 wm-bot2: Safe rebooting 'cloudvirt1053.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 18:50 wm-bot2: Draining 'cloudvirt1053.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 18:50 wm-bot2: Safe rebooting 'cloudvirt1053.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 18:46 wm-bot2: Draining 'cloudvirt1053.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 18:46 wm-bot2: Safe rebooting 'cloudvirt1053.eqiad.wmnet'. - cookbook ran by andrew@bullseye === 2022-10-15 === * 17:38 taavi: taavi@cloudweb1003 ~ $ mwscript extensions/OATHAuth/maintenance/disableOATHAuthForUser.php --wiki=labswiki Slevinski # [[phab:T320867|T320867]] === 2022-10-13 === * 12:19 wm-bot2: OSDs (['cloudcephosd1027', 'cloudcephosd1028', 'cloudcephosd1029', 'cloudcephosd1030', 'cloudcephosd1031', 'cloudcephosd1032', 'cloudcephosd1033', 'cloudcephosd1034']) upgraded successfully B-) ([[phab:T309786|T309786]]) - cookbook ran by dcaro@vulcanus * 11:31 wm-bot2: Upgrading OSDs and rebooting the nodes ['cloudcephosd1027', 'cloudcephosd1028', 'cloudcephosd1029', 'cloudcephosd1030', 'cloudcephosd1031', 'cloudcephosd1032', 'cloudcephosd1033', 'cloudcephosd1034'] ([[phab:T309786|T309786]]) - cookbook ran by dcaro@vulcanus * 11:30 wm-bot2: OSDs (['cloudcephosd1025', 'cloudcephosd1026']) upgraded successfully B-) ([[phab:T309786|T309786]]) - cookbook ran by dcaro@vulcanus === 2022-10-10 === * 14:01 dcaro: test2 * 01:22 andrewbogott: restarting designate-sink on cloudservices100[45], possible example of [[phab:T316614|T316614]] === 2022-10-09 === * 12:04 taavi: taavi@cloudweb1003 ~ $ mwscript extensions/OATHAuth/maintenance/disableOATHAuthForUser.php --wiki=labswiki DatGuy # [[phab:T320301|T320301]] === 2022-10-07 === * 13:40 andrewbogott: dhinus is resetting rabbitmq cluster in an attempt to resolve a suspected (by Andrew) split-brain * 11:33 arturo: rabbitmq-server.service @ cloudrabbit1002 is again up and running ([[phab:T320232|T320232]]) * 10:24 arturo: stopping rabbitmq-server.service @ cloudrabbit1002 ([[phab:T320232|T320232]]) * 10:19 arturo: restarting nova-conductor in all 3 cloudcontrols ([[phab:T320232|T320232]]) * 09:45 arturo: restarting rabbitmq-server.service @ cloudrabbit1002 ([[phab:T320232|T320232]]) === 2022-10-06 === * 15:55 arturo: cloudnet1005 & cloudnet1006 now in service. Secom cloudnet1003 & cloudnet1004. Drop neutron agents, etc. ([[phab:T316284|T316284]]) * 11:54 arturo: rebooting cloudnet1005/1006 to see if they have the right network config ([[phab:T316284|T316284]]) * 11:50 arturo: set neutron l3 agents on cloudnet1005/1006 as down `root@cloudcontrol1005:~# neutron agent-update --admin-state-down <uuid>` ([[phab:T316284|T316284]]) * 11:40 arturo: [codfw1dev] rebooting both network nodes to test https://gerrit.wikimedia.org/r/c/operations/puppet/+/839492 * 10:14 arturo: [codfw1dev] restart neutron-l3-agent on cloudnet2006-dev, it was dead === 2022-10-05 === * 14:40 wm-bot2: Adding OSD cloudcephosd1021.eqiad.wmnet... (1/1) ([[phab:T319418|T319418]]) - cookbook ran by fran@wmf3169 * 14:40 wm-bot2: Adding new OSDs ['cloudcephosd1021.eqiad.wmnet'] to the cluster ([[phab:T319418|T319418]]) - cookbook ran by fran@wmf3169 * 14:28 arturo: adding cloudinstances2b-gw router to l3 agents on cloudnet1005/1006 ([[phab:T316284|T316284]]) * 13:11 wm-bot2: Added 1 new OSDs ['cloudcephosd1034.eqiad.wmnet'] ([[phab:T314870|T314870]]) - cookbook ran by fran@wmf3169 * 13:11 wm-bot2: Added OSD cloudcephosd1034.eqiad.wmnet... (1/1) ([[phab:T314870|T314870]]) - cookbook ran by fran@wmf3169 * 13:02 wm-bot2: Finished rebooting node cloudcephosd1034.eqiad.wmnet ([[phab:T314870|T314870]]) - cookbook ran by fran@wmf3169 * 12:58 wm-bot2: Rebooting node cloudcephosd1034.eqiad.wmnet ([[phab:T314870|T314870]]) - cookbook ran by fran@wmf3169 * 12:58 wm-bot2: Adding OSD cloudcephosd1034.eqiad.wmnet... (1/1) ([[phab:T314870|T314870]]) - cookbook ran by fran@wmf3169 * 12:58 wm-bot2: Adding new OSDs ['cloudcephosd1034.eqiad.wmnet'] to the cluster ([[phab:T314870|T314870]]) - cookbook ran by fran@wmf3169 === 2022-10-04 === * 16:40 wm-bot2: Added 1 new OSDs ['cloudcephosd1033.eqiad.wmnet'] ([[phab:T314870|T314870]]) - cookbook ran by fran@wmf3169 * 16:40 wm-bot2: Added OSD cloudcephosd1033.eqiad.wmnet... (1/1) ([[phab:T314870|T314870]]) - cookbook ran by fran@wmf3169 * 14:34 wm-bot2: Finished rebooting node cloudcephosd1033.eqiad.wmnet ([[phab:T314870|T314870]]) - cookbook ran by fran@wmf3169 * 14:30 wm-bot2: Rebooting node cloudcephosd1033.eqiad.wmnet ([[phab:T314870|T314870]]) - cookbook ran by fran@wmf3169 * 14:30 wm-bot2: Adding OSD cloudcephosd1033.eqiad.wmnet... (1/1) ([[phab:T314870|T314870]]) - cookbook ran by fran@wmf3169 * 14:30 wm-bot2: Adding new OSDs ['cloudcephosd1033.eqiad.wmnet'] to the cluster ([[phab:T314870|T314870]]) - cookbook ran by fran@wmf3169 * 14:20 wm-bot2: Finished rebooting node cloudcephosd1033.eqiad.wmnet - cookbook ran by fran@wmf3169 * 14:17 wm-bot2: Rebooting node cloudcephosd1033.eqiad.wmnet - cookbook ran by fran@wmf3169 * 14:16 wm-bot2: Adding OSD cloudcephosd1033.eqiad.wmnet... (1/1) - cookbook ran by fran@wmf3169 * 14:16 wm-bot2: Adding new OSDs ['cloudcephosd1033.eqiad.wmnet'] to the cluster - cookbook ran by fran@wmf3169 * 10:59 wm-bot2: Finished rebooting node cloudcephosd1033.eqiad.wmnet - cookbook ran by fran@wmf3169 * 10:56 wm-bot2: Rebooting node cloudcephosd1033.eqiad.wmnet - cookbook ran by fran@wmf3169 * 10:55 wm-bot2: Adding OSD cloudcephosd1033.eqiad.wmnet... (1/1) - cookbook ran by fran@wmf3169 * 10:55 wm-bot2: Adding new OSDs ['cloudcephosd1033.eqiad.wmnet'] to the cluster - cookbook ran by fran@wmf3169 === 2022-09-30 === * 14:52 wm-bot2: Added 1 new OSDs ['cloudcephosd1031.eqiad.wmnet'] - cookbook ran by fran@wmf3169 * 14:52 wm-bot2: Added OSD cloudcephosd1031.eqiad.wmnet... (1/1) - cookbook ran by fran@wmf3169 * 14:48 wm-bot2: Adding OSD cloudcephosd1031.eqiad.wmnet... (1/1) - cookbook ran by fran@wmf3169 * 14:48 wm-bot2: Adding new OSDs ['cloudcephosd1031.eqiad.wmnet'] to the cluster - cookbook ran by fran@wmf3169 * 14:16 wm-bot2: Drained 'cloudvirt1023.eqiad.wmnet'. ([[phab:T319025|T319025]]) - cookbook ran by andrew@buster * 14:15 wm-bot2: Set cloudvirt 'cloudvirt1023.eqiad.wmnet' maintenance (downtime id: 64eac5c6-4b1d-4269-98fd-{{Gerrit|8e5bed42ce40}}, use this to unset). ([[phab:T319025|T319025]]) - cookbook ran by andrew@buster * 14:15 wm-bot2: Draining 'cloudvirt1023.eqiad.wmnet'. ([[phab:T319025|T319025]]) - cookbook ran by andrew@buster * 14:15 wm-bot2: Safe rebooting 'cloudvirt1023.eqiad.wmnet'. ([[phab:T319025|T319025]]) - cookbook ran by andrew@buster * 14:09 wm-bot2: Drained 'cloudvirt1023.eqiad.wmnet'. - cookbook ran by andrew@buster * 14:05 wm-bot2: Set cloudvirt 'cloudvirt1023.eqiad.wmnet' maintenance (downtime id: fb99c967-b974-4314-a2fa-{{Gerrit|31ed0e883dd3}}, use this to unset). - cookbook ran by andrew@buster * 14:04 wm-bot2: Draining 'cloudvirt1023.eqiad.wmnet'. - cookbook ran by andrew@buster * 13:30 wm-bot2: Added 1 new OSDs ['cloudcephosd1031.eqiad.wmnet'] - cookbook ran by fran@wmf3169 * 13:30 wm-bot2: Added OSD cloudcephosd1031.eqiad.wmnet... (1/1) - cookbook ran by fran@wmf3169 * 13:26 wm-bot2: Adding OSD cloudcephosd1031.eqiad.wmnet... (1/1) - cookbook ran by fran@wmf3169 * 13:26 wm-bot2: Adding new OSDs ['cloudcephosd1031.eqiad.wmnet'] to the cluster - cookbook ran by fran@wmf3169 * 13:16 wm-bot2: Adding OSD cloudcephosd1031.eqiad.wmnet... (1/1) - cookbook ran by fran@wmf3169 * 13:16 wm-bot2: Adding new OSDs ['cloudcephosd1031.eqiad.wmnet'] to the cluster - cookbook ran by fran@wmf3169 * 13:08 wm-bot2: Adding OSD cloudcephosd1031.eqiad.wmnet... (1/1) - cookbook ran by fran@wmf3169 * 13:08 wm-bot2: Adding new OSDs ['cloudcephosd1031.eqiad.wmnet'] to the cluster - cookbook ran by fran@wmf3169 * 12:45 wm-bot2: Finished rebooting node cloudcephosd1031.eqiad.wmnet ([[phab:T314870|T314870]]) - cookbook ran by fran@wmf3169 * 12:42 wm-bot2: Rebooting node cloudcephosd1031.eqiad.wmnet ([[phab:T314870|T314870]]) - cookbook ran by fran@wmf3169 * 12:41 wm-bot2: Adding OSD cloudcephosd1031.eqiad.wmnet... (1/1) ([[phab:T314870|T314870]]) - cookbook ran by fran@wmf3169 * 12:41 wm-bot2: Adding new OSDs ['cloudcephosd1031.eqiad.wmnet'] to the cluster ([[phab:T314870|T314870]]) - cookbook ran by fran@wmf3169 * 12:26 arturo: [codfw1dev] cloudnet2005/2006-dev are now on a single NIC setup ([[phab:T318824|T318824]]) * 11:44 arturo: sysctl change (cleanup) on cloudnet1003/1004 === 2022-09-27 === * 10:48 wm-bot2: Added OSD cloudcephosd1030.eqiad.wmnet... (1/1) ([[phab:T314870|T314870]]) - cookbook ran by fran@wmf3169 * 10:35 wm-bot2: Finished rebooting node cloudcephosd1030.eqiad.wmnet ([[phab:T314870|T314870]]) - cookbook ran by fran@wmf3169 * 10:32 wm-bot2: Rebooting node cloudcephosd1030.eqiad.wmnet ([[phab:T314870|T314870]]) - cookbook ran by fran@wmf3169 * 10:32 wm-bot2: Adding OSD cloudcephosd1030.eqiad.wmnet... (1/1) ([[phab:T314870|T314870]]) - cookbook ran by fran@wmf3169 * 10:32 wm-bot2: Adding new OSDs ['cloudcephosd1030.eqiad.wmnet'] to the cluster ([[phab:T314870|T314870]]) - cookbook ran by fran@wmf3169 * 10:26 wm-bot2: Finished rebooting node cloudcephosd1030.eqiad.wmnet - cookbook ran by fran@wmf3169 * 10:23 wm-bot2: Rebooting node cloudcephosd1030.eqiad.wmnet - cookbook ran by fran@wmf3169 * 10:22 wm-bot2: Adding OSD cloudcephosd1030.eqiad.wmnet... (1/1) - cookbook ran by fran@wmf3169 * 10:22 wm-bot2: Adding new OSDs ['cloudcephosd1030.eqiad.wmnet'] to the cluster - cookbook ran by fran@wmf3169 * 10:17 wm-bot2: Finished rebooting node cloudcephosd1030.eqiad.wmnet - cookbook ran by fran@wmf3169 * 10:14 wm-bot2: Rebooting node cloudcephosd1030.eqiad.wmnet - cookbook ran by fran@wmf3169 * 10:14 wm-bot2: Adding OSD cloudcephosd1030.eqiad.wmnet... (1/1) - cookbook ran by fran@wmf3169 * 10:14 wm-bot2: Adding new OSDs ['cloudcephosd1030.eqiad.wmnet'] to the cluster - cookbook ran by fran@wmf3169 * 10:09 wm-bot2: Finished rebooting node cloudcephosd1030.eqiad.wmnet - cookbook ran by fran@wmf3169 * 10:05 wm-bot2: Rebooting node cloudcephosd1030.eqiad.wmnet - cookbook ran by fran@wmf3169 * 10:05 wm-bot2: Adding OSD cloudcephosd1030.eqiad.wmnet... (1/1) - cookbook ran by fran@wmf3169 * 10:05 wm-bot2: Adding new OSDs ['cloudcephosd1030.eqiad.wmnet'] to the cluster - cookbook ran by fran@wmf3169 * 10:05 wm-bot2: Adding OSD cloudcephosd1030.eqiad.wmnet... (1/1) - cookbook ran by fran@wmf3169 * 10:05 wm-bot2: Adding new OSDs ['cloudcephosd1030.eqiad.wmnet'] to the cluster - cookbook ran by fran@wmf3169 === 2022-09-26 === * 18:37 wm-bot2: Safe reboot of 'cloudvirt1024.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 18:37 wm-bot2: Unset cloudvirt 'cloudvirt1024.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 18:33 wm-bot2: Drained 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 18:31 andrewbogott: rebooting cloudvirt1028 for [[phab:T317391|T317391]] * 18:28 wm-bot2: Safe reboot of 'cloudvirt-wdqs1002.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:28 wm-bot2: Unset cloudvirt 'cloudvirt-wdqs1002.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:28 wm-bot2: Safe reboot of 'cloudvirt-wdqs1003.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:28 wm-bot2: Unset cloudvirt 'cloudvirt-wdqs1003.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:28 wm-bot2: Safe reboot of 'cloudvirt-wdqs1001.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:28 wm-bot2: Unset cloudvirt 'cloudvirt-wdqs1001.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:25 wm-bot2: Drained 'cloudvirt-wdqs1003.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:25 wm-bot2: Set cloudvirt 'cloudvirt-wdqs1003.eqiad.wmnet' maintenance (downtime id: 6d1c4c53-76b9-493f-8c44-{{Gerrit|eb0413dfb9d0}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:25 wm-bot2: Drained 'cloudvirt-wdqs1002.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:25 wm-bot2: Set cloudvirt 'cloudvirt-wdqs1002.eqiad.wmnet' maintenance (downtime id: 6bb80b65-616c-4e45-be4f-{{Gerrit|d2cdd8abb2bf}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:25 wm-bot2: Drained 'cloudvirt-wdqs1001.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:25 wm-bot2: Set cloudvirt 'cloudvirt-wdqs1001.eqiad.wmnet' maintenance (downtime id: a3bba7e7-bcd3-482d-a486-{{Gerrit|06b6f53fd899}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:24 wm-bot2: Draining 'cloudvirt-wdqs1003.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:24 wm-bot2: Safe rebooting 'cloudvirt-wdqs1003.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:24 wm-bot2: Draining 'cloudvirt-wdqs1002.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:24 wm-bot2: Safe rebooting 'cloudvirt-wdqs1002.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:24 wm-bot2: Draining 'cloudvirt-wdqs1001.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:24 wm-bot2: Safe rebooting 'cloudvirt-wdqs1001.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:18 wm-bot2: Set cloudvirt 'cloudvirt1024.eqiad.wmnet' maintenance (downtime id: 8fb59651-fd42-4aee-a978-{{Gerrit|36bc7f79b4d5}}, use this to unset). - cookbook ran by andrew@buster * 18:18 wm-bot2: Draining 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 18:17 wm-bot2: Safe rebooting 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 18:01 wm-bot2: Safe reboot of 'cloudvirt1023.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:01 wm-bot2: Unset cloudvirt 'cloudvirt1023.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:58 wm-bot2: Drained 'cloudvirt1023.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:51 wm-bot2: Safe reboot of 'cloudvirt1025.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:51 wm-bot2: Unset cloudvirt 'cloudvirt1025.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:47 wm-bot2: Safe reboot of 'cloudvirt1027.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 17:47 wm-bot2: Unset cloudvirt 'cloudvirt1027.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 17:47 wm-bot2: Drained 'cloudvirt1025.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:45 wm-bot2: Drained 'cloudvirt1027.eqiad.wmnet'. - cookbook ran by andrew@buster * 17:38 wm-bot2: Set cloudvirt 'cloudvirt1023.eqiad.wmnet' maintenance (downtime id: abfc35bb-331e-4d7e-bb92-{{Gerrit|35e675abce32}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:37 wm-bot2: Draining 'cloudvirt1023.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:37 wm-bot2: Safe rebooting 'cloudvirt1023.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:37 wm-bot2: Safe reboot of 'cloudvirt1022.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:37 wm-bot2: Unset cloudvirt 'cloudvirt1022.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:34 wm-bot2: Drained 'cloudvirt1022.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:33 wm-bot2: Set cloudvirt 'cloudvirt1025.eqiad.wmnet' maintenance (downtime id: d4fe6ea8-76eb-45f7-b6cf-{{Gerrit|66f54ba9801c}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:33 wm-bot2: Set cloudvirt 'cloudvirt1027.eqiad.wmnet' maintenance (downtime id: c0671bda-9aae-4960-85be-{{Gerrit|58aaa9c8e021}}, use this to unset). - cookbook ran by andrew@buster * 17:32 wm-bot2: Draining 'cloudvirt1025.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:32 wm-bot2: Safe rebooting 'cloudvirt1025.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:32 wm-bot2: Draining 'cloudvirt1027.eqiad.wmnet'. - cookbook ran by andrew@buster * 17:32 wm-bot2: Safe rebooting 'cloudvirt1027.eqiad.wmnet'. - cookbook ran by andrew@buster * 17:31 wm-bot2: Safe reboot of 'cloudvirt1029.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 17:31 wm-bot2: Unset cloudvirt 'cloudvirt1029.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 17:31 wm-bot2: Safe reboot of 'cloudvirt1030.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:31 wm-bot2: Unset cloudvirt 'cloudvirt1030.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:28 wm-bot2: Drained 'cloudvirt1029.eqiad.wmnet'. - cookbook ran by andrew@buster * 17:28 wm-bot2: Drained 'cloudvirt1030.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:24 wm-bot2: Set cloudvirt 'cloudvirt1022.eqiad.wmnet' maintenance (downtime id: be6559f9-4d72-4492-b9b9-{{Gerrit|23c636d6554e}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:23 wm-bot2: Draining 'cloudvirt1022.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:23 wm-bot2: Safe rebooting 'cloudvirt1022.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:21 wm-bot2: Set cloudvirt 'cloudvirt1022.eqiad.wmnet' maintenance (downtime id: 237dfb1b-5382-408d-aa35-{{Gerrit|1aa1cfe4c5af}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:20 wm-bot2: Draining 'cloudvirt1022.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:20 wm-bot2: Safe rebooting 'cloudvirt1022.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:20 wm-bot2: Safe reboot of 'cloudvirt1021.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:20 wm-bot2: Unset cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:17 wm-bot2: Set cloudvirt 'cloudvirt1029.eqiad.wmnet' maintenance (downtime id: 088b97cd-0cf2-4c27-a67f-{{Gerrit|077ffa1c1c5e}}, use this to unset). - cookbook ran by andrew@buster * 17:16 wm-bot2: Drained 'cloudvirt1021.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:16 wm-bot2: Set cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance (downtime id: 91625000-aa6d-4887-bf34-{{Gerrit|ef7785a4af4a}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:16 wm-bot2: Draining 'cloudvirt1029.eqiad.wmnet'. - cookbook ran by andrew@buster * 17:16 wm-bot2: Safe rebooting 'cloudvirt1029.eqiad.wmnet'. - cookbook ran by andrew@buster * 17:16 wm-bot2: Draining 'cloudvirt1021.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:15 wm-bot2: Safe rebooting 'cloudvirt1021.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:15 wm-bot2: Safe reboot of 'cloudvirt1017.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 17:15 wm-bot2: Unset cloudvirt 'cloudvirt1017.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 17:15 wm-bot2: Set cloudvirt 'cloudvirt1030.eqiad.wmnet' maintenance (downtime id: 7364c4fe-9f6b-4539-b6fa-{{Gerrit|1e768b602a1b}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:14 wm-bot2: Draining 'cloudvirt1030.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:14 wm-bot2: Safe rebooting 'cloudvirt1030.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:12 wm-bot2: Drained 'cloudvirt1017.eqiad.wmnet'. - cookbook ran by andrew@buster * 17:12 wm-bot2: Set cloudvirt 'cloudvirt1017.eqiad.wmnet' maintenance (downtime id: cb8f8a21-73c0-4f54-9412-{{Gerrit|074d82582cb4}}, use this to unset). - cookbook ran by andrew@buster * 17:11 wm-bot2: Draining 'cloudvirt1017.eqiad.wmnet'. - cookbook ran by andrew@buster * 17:11 wm-bot2: Safe rebooting 'cloudvirt1017.eqiad.wmnet'. - cookbook ran by andrew@buster * 17:09 wm-bot2: Set cloudvirt 'cloudvirt1017.eqiad.wmnet' maintenance (downtime id: d23dabd2-3bfc-4ce0-9533-{{Gerrit|cd1cf910f8e1}}, use this to unset). - cookbook ran by andrew@buster * 17:08 wm-bot2: Draining 'cloudvirt1017.eqiad.wmnet'. - cookbook ran by andrew@buster * 17:08 wm-bot2: Safe rebooting 'cloudvirt1017.eqiad.wmnet'. - cookbook ran by andrew@buster * 17:05 wm-bot2: Set cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance (downtime id: 4d7ddae1-f621-4dd5-8616-{{Gerrit|29b7699b692f}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:04 wm-bot2: Draining 'cloudvirt1021.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:04 wm-bot2: Safe rebooting 'cloudvirt1021.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:04 wm-bot2: Set cloudvirt 'cloudvirt1017.eqiad.wmnet' maintenance (downtime id: 4628dd6f-5b49-417f-b6ab-{{Gerrit|5c41992192b4}}, use this to unset). - cookbook ran by andrew@buster * 17:03 wm-bot2: Draining 'cloudvirt1017.eqiad.wmnet'. - cookbook ran by andrew@buster * 17:03 wm-bot2: Safe rebooting 'cloudvirt1017.eqiad.wmnet'. - cookbook ran by andrew@buster * 17:02 wm-bot2: Set cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance (downtime id: c95529cf-b2a5-4553-bd6e-{{Gerrit|6ad4dcb2a105}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:01 wm-bot2: Set cloudvirt 'cloudvirt1017.eqiad.wmnet' maintenance (downtime id: 36c68112-a964-4a7c-a093-{{Gerrit|8a6d481dea7d}}, use this to unset). - cookbook ran by andrew@buster * 17:01 wm-bot2: Draining 'cloudvirt1021.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:01 wm-bot2: Safe rebooting 'cloudvirt1021.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:00 wm-bot2: Draining 'cloudvirt1017.eqiad.wmnet'. - cookbook ran by andrew@buster * 17:00 wm-bot2: Safe rebooting 'cloudvirt1017.eqiad.wmnet'. - cookbook ran by andrew@buster * 16:49 wm-bot2: Safe reboot of 'cloudvirt1050.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 16:49 wm-bot2: Unset cloudvirt 'cloudvirt1050.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 16:45 wm-bot2: Drained 'cloudvirt1050.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 16:31 wm-bot2: Set cloudvirt 'cloudvirt1050.eqiad.wmnet' maintenance (downtime id: 1244c159-65bb-476e-a702-{{Gerrit|8f43c253e499}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 16:30 wm-bot2: Draining 'cloudvirt1050.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 16:30 wm-bot2: Safe rebooting 'cloudvirt1050.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:22 wm-bot2: Finished rebooting node cloudcephosd1030.eqiad.wmnet - cookbook ran by fran@wmf3169 * 13:19 wm-bot2: Rebooting node cloudcephosd1030.eqiad.wmnet - cookbook ran by fran@wmf3169 * 13:18 wm-bot2: Adding OSD cloudcephosd1030.eqiad.wmnet... (1/1) - cookbook ran by fran@wmf3169 * 13:18 wm-bot2: Adding new OSDs ['cloudcephosd1030.eqiad.wmnet'] to the cluster - cookbook ran by fran@wmf3169 * 12:32 dcaro: Changed the collation of labsdbaccount db to utf8mb4_bin ([[phab:T318047|T318047]]) * 10:14 arturo: deployed new version of maintain-dbusers ([[phab:T318047|T318047]]) === 2022-09-25 === * 15:06 wm-bot2: Drained 'cloudvirt1052.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 15:06 wm-bot2: Set cloudvirt 'cloudvirt1052.eqiad.wmnet' maintenance (downtime id: a050e47f-3a63-41ff-accb-{{Gerrit|3993c1e8b593}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 15:05 wm-bot2: Draining 'cloudvirt1052.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 15:05 wm-bot2: Safe rebooting 'cloudvirt1052.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 15:03 wm-bot2: Safe reboot of 'cloudvirt1052.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 15:03 wm-bot2: Unset cloudvirt 'cloudvirt1052.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:58 wm-bot2: Drained 'cloudvirt1052.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:58 wm-bot2: Set cloudvirt 'cloudvirt1052.eqiad.wmnet' maintenance (downtime id: 1a49d773-4dc9-4a6f-bd49-{{Gerrit|8663db77ce66}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:57 wm-bot2: Draining 'cloudvirt1052.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:57 wm-bot2: Safe rebooting 'cloudvirt1052.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:54 wm-bot2: Set cloudvirt 'cloudvirt1052.eqiad.wmnet' maintenance (downtime id: ab52dd42-68a6-4159-b345-{{Gerrit|e5625d209b57}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:53 wm-bot2: Draining 'cloudvirt1052.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:53 wm-bot2: Safe rebooting 'cloudvirt1052.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:37 wm-bot2: Set cloudvirt 'cloudvirt1052.eqiad.wmnet' maintenance (downtime id: 10337415-3095-4834-8538-{{Gerrit|b3a64bd6b18d}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:37 wm-bot2: Draining 'cloudvirt1052.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:37 wm-bot2: Safe rebooting 'cloudvirt1052.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:53 wm-bot2: Drained 'cloudvirt1053.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:53 wm-bot2: Set cloudvirt 'cloudvirt1053.eqiad.wmnet' maintenance (downtime id: 4cdbed3a-3775-4913-8c3b-{{Gerrit|afd57ae1ca55}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:52 wm-bot2: Draining 'cloudvirt1053.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:52 wm-bot2: Safe rebooting 'cloudvirt1053.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster === 2022-09-24 === * 17:37 andrewbogott: restarting neutron api on cloudcontrol1006; cause of outage unknown * 17:35 andrewbogott: restarting neutron-linuxbridge-agent on cloudvirt1022 === 2022-09-22 === * 15:14 wm-bot2: Safe reboot of 'cloudvirt1025.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 15:14 wm-bot2: Unset cloudvirt 'cloudvirt1025.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 15:10 wm-bot2: Drained 'cloudvirt1025.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 15:10 wm-bot2: Set cloudvirt 'cloudvirt1025.eqiad.wmnet' maintenance (downtime id: 04849437-a304-45a1-a037-{{Gerrit|538741a2f801}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 15:09 wm-bot2: Draining 'cloudvirt1025.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 15:09 wm-bot2: Safe rebooting 'cloudvirt1025.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 15:07 wm-bot2: Safe reboot of 'cloudvirt1017.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 15:07 wm-bot2: Unset cloudvirt 'cloudvirt1017.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 15:04 wm-bot2: Drained 'cloudvirt1017.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 15:03 wm-bot2: Set cloudvirt 'cloudvirt1017.eqiad.wmnet' maintenance (downtime id: 8810a919-61e7-4ba5-af7f-{{Gerrit|c41d1e1c8c8b}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 15:03 wm-bot2: Draining 'cloudvirt1017.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 15:03 wm-bot2: Safe rebooting 'cloudvirt1017.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:59 wm-bot2: Safe reboot of 'cloudvirt1021.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:59 wm-bot2: Unset cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:56 wm-bot2: Drained 'cloudvirt1021.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:42 wm-bot2: Set cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance (downtime id: b6f8e0b6-f35c-41fd-8035-{{Gerrit|6efde61db3a3}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:42 wm-bot2: Safe reboot of 'cloudvirt1027.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:42 wm-bot2: Unset cloudvirt 'cloudvirt1027.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:42 wm-bot2: Draining 'cloudvirt1021.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:42 wm-bot2: Safe rebooting 'cloudvirt1021.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:41 wm-bot2: Safe reboot of 'cloudvirt1022.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:41 wm-bot2: Unset cloudvirt 'cloudvirt1022.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:41 wm-bot2: Set cloudvirt 'cloudvirt1017.eqiad.wmnet' maintenance (downtime id: 784bff2c-8160-44dc-9bcf-{{Gerrit|ff23055a0527}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:40 wm-bot2: Draining 'cloudvirt1017.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:40 wm-bot2: Safe rebooting 'cloudvirt1017.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:39 wm-bot2: Drained 'cloudvirt1027.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:38 wm-bot2: Drained 'cloudvirt1022.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:38 wm-bot2: Set cloudvirt 'cloudvirt1022.eqiad.wmnet' maintenance (downtime id: 21152e9e-bc67-4024-b68e-{{Gerrit|fd671c398642}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:37 wm-bot2: Draining 'cloudvirt1022.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:37 wm-bot2: Safe rebooting 'cloudvirt1022.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:37 wm-bot2: Set cloudvirt 'cloudvirt1027.eqiad.wmnet' maintenance (downtime id: f07ddb1d-e8da-43f3-8140-{{Gerrit|bc66789d37c8}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:36 wm-bot2: Draining 'cloudvirt1027.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:36 wm-bot2: Safe rebooting 'cloudvirt1027.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:33 wm-bot2: Safe reboot of 'cloudvirt1024.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:33 wm-bot2: Unset cloudvirt 'cloudvirt1024.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:30 wm-bot2: Drained 'cloudvirt1024.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:16 wm-bot2: Set cloudvirt 'cloudvirt1024.eqiad.wmnet' maintenance (downtime id: 4e4b03fc-e3f1-440e-9097-{{Gerrit|cce5edaa1f08}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:16 wm-bot2: Set cloudvirt 'cloudvirt1027.eqiad.wmnet' maintenance (downtime id: d4bb2160-f492-4ddc-a5b7-{{Gerrit|eeee37007d14}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:16 wm-bot2: Set cloudvirt 'cloudvirt1022.eqiad.wmnet' maintenance (downtime id: 9bf00a6a-3697-4201-9a6e-{{Gerrit|76d499eccfa1}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:15 wm-bot2: Draining 'cloudvirt1027.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:15 wm-bot2: Safe rebooting 'cloudvirt1027.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:15 wm-bot2: Draining 'cloudvirt1024.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:15 wm-bot2: Safe rebooting 'cloudvirt1024.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:15 wm-bot2: Draining 'cloudvirt1022.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:15 wm-bot2: Safe rebooting 'cloudvirt1022.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:13 wm-bot2: Safe reboot of 'cloudvirt1030.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:13 wm-bot2: Unset cloudvirt 'cloudvirt1030.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:47 wm-bot2: Set cloudvirt 'cloudvirt1029.eqiad.wmnet' maintenance (downtime id: f2888490-1804-4b23-b7b0-{{Gerrit|677853fb9ef7}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:46 wm-bot2: Draining 'cloudvirt1029.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:46 wm-bot2: Safe rebooting 'cloudvirt1029.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:46 wm-bot2: Set cloudvirt 'cloudvirt1030.eqiad.wmnet' maintenance (downtime id: 8a3e8b0a-ced3-4667-a121-{{Gerrit|66df673c7ed8}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:46 wm-bot2: Safe reboot of 'cloudvirt1033.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:46 wm-bot2: Unset cloudvirt 'cloudvirt1033.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:45 wm-bot2: Draining 'cloudvirt1030.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:45 wm-bot2: Safe rebooting 'cloudvirt1030.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:45 wm-bot2: Safe reboot of 'cloudvirt1032.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:45 wm-bot2: Unset cloudvirt 'cloudvirt1032.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:43 wm-bot2: Set cloudvirt 'cloudvirt1031.eqiad.wmnet' maintenance (downtime id: 35c761b3-f0de-4b23-9b11-{{Gerrit|2ed2e474c89f}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:42 wm-bot2: Draining 'cloudvirt1031.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:42 wm-bot2: Safe rebooting 'cloudvirt1031.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:42 wm-bot2: Drained 'cloudvirt1033.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:42 wm-bot2: Set cloudvirt 'cloudvirt1033.eqiad.wmnet' maintenance (downtime id: 6121a0e1-754a-47fa-b51b-{{Gerrit|265f6386f396}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:41 wm-bot2: Drained 'cloudvirt1032.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:41 wm-bot2: Draining 'cloudvirt1033.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:41 wm-bot2: Safe rebooting 'cloudvirt1033.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:40 wm-bot2: Safe reboot of 'cloudvirt1035.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:40 wm-bot2: Unset cloudvirt 'cloudvirt1035.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:39 wm-bot2: Set cloudvirt 'cloudvirt1033.eqiad.wmnet' maintenance (downtime id: 189ae1bb-c69d-4eb0-aee9-{{Gerrit|90ab1f492aa8}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:38 wm-bot2: Draining 'cloudvirt1033.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:38 wm-bot2: Safe rebooting 'cloudvirt1033.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:36 wm-bot2: Drained 'cloudvirt1035.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:35 wm-bot2: Set cloudvirt 'cloudvirt1035.eqiad.wmnet' maintenance (downtime id: acc53776-68bf-47a4-bb82-{{Gerrit|c571b3df2945}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:35 wm-bot2: Draining 'cloudvirt1035.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:35 wm-bot2: Safe rebooting 'cloudvirt1035.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:31 wm-bot2: Set cloudvirt 'cloudvirt1035.eqiad.wmnet' maintenance (downtime id: 07770ae7-3e7e-4615-8174-{{Gerrit|513767f95b1d}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:30 wm-bot2: Draining 'cloudvirt1035.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:30 wm-bot2: Safe rebooting 'cloudvirt1035.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:29 wm-bot2: Set cloudvirt 'cloudvirt1035.eqiad.wmnet' maintenance (downtime id: 7ae20a7d-8bff-4bc2-91c5-{{Gerrit|b0005c8164d1}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:29 wm-bot2: Set cloudvirt 'cloudvirt1032.eqiad.wmnet' maintenance (downtime id: b596f855-8019-49d2-9657-{{Gerrit|0d513ad7acb5}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:29 wm-bot2: Draining 'cloudvirt1035.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:29 wm-bot2: Safe rebooting 'cloudvirt1035.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:28 wm-bot2: Draining 'cloudvirt1032.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:28 wm-bot2: Safe rebooting 'cloudvirt1032.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:28 wm-bot2: Safe reboot of 'cloudvirt1034.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:28 wm-bot2: Unset cloudvirt 'cloudvirt1034.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:25 wm-bot2: Set cloudvirt 'cloudvirt1033.eqiad.wmnet' maintenance (downtime id: 9172c785-d2d6-46e5-bed3-{{Gerrit|2f49d1033f78}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:24 wm-bot2: Draining 'cloudvirt1033.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:24 wm-bot2: Safe rebooting 'cloudvirt1033.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:24 wm-bot2: Drained 'cloudvirt1034.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:23 wm-bot2: Safe reboot of 'cloudvirt1036.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:23 wm-bot2: Unset cloudvirt 'cloudvirt1036.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:19 wm-bot2: Drained 'cloudvirt1036.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:04 wm-bot2: Set cloudvirt 'cloudvirt1036.eqiad.wmnet' maintenance (downtime id: 85b4e0e2-e635-485f-9ce5-{{Gerrit|341981b3887f}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:04 wm-bot2: Set cloudvirt 'cloudvirt1035.eqiad.wmnet' maintenance (downtime id: 9494c283-02c8-42da-98c3-{{Gerrit|139b4d119ef8}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:04 wm-bot2: Set cloudvirt 'cloudvirt1034.eqiad.wmnet' maintenance (downtime id: d181f301-249f-4b5a-bd85-{{Gerrit|47d87262d0ed}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:03 wm-bot2: Draining 'cloudvirt1034.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:03 wm-bot2: Safe rebooting 'cloudvirt1034.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:03 wm-bot2: Draining 'cloudvirt1035.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:03 wm-bot2: Safe rebooting 'cloudvirt1035.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:03 wm-bot2: Draining 'cloudvirt1036.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:03 wm-bot2: Safe rebooting 'cloudvirt1036.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster === 2022-09-20 === * 21:02 wm-bot2: Safe reboot of 'cloudvirt1037.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 21:02 wm-bot2: Unset cloudvirt 'cloudvirt1037.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:58 wm-bot2: Drained 'cloudvirt1037.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:58 wm-bot2: Set cloudvirt 'cloudvirt1037.eqiad.wmnet' maintenance (downtime id: 1c4246a4-9cee-4423-8cd1-{{Gerrit|4f52f61d503f}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:57 wm-bot2: Draining 'cloudvirt1037.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:57 wm-bot2: Safe rebooting 'cloudvirt1037.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:57 wm-bot2: Safe reboot of 'cloudvirt1038.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:57 wm-bot2: Unset cloudvirt 'cloudvirt1038.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:53 wm-bot2: Drained 'cloudvirt1038.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:50 wm-bot2: Safe reboot of 'cloudvirt1039.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:50 wm-bot2: Unset cloudvirt 'cloudvirt1039.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:49 wm-bot2: Set cloudvirt 'cloudvirt1037.eqiad.wmnet' maintenance (downtime id: 82665ecc-4431-48fe-b255-{{Gerrit|7e9d518be217}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:48 wm-bot2: Draining 'cloudvirt1037.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:48 wm-bot2: Safe rebooting 'cloudvirt1037.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:47 wm-bot2: Set cloudvirt 'cloudvirt1037.eqiad.wmnet' maintenance (downtime id: cd284562-e3b0-4504-9977-{{Gerrit|2b18032bbf3c}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:46 wm-bot2: Draining 'cloudvirt1037.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:46 wm-bot2: Safe rebooting 'cloudvirt1037.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:46 wm-bot2: Drained 'cloudvirt1039.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:45 wm-bot2: Set cloudvirt 'cloudvirt1037.eqiad.wmnet' maintenance (downtime id: 8fb2a646-16a5-4183-94a1-{{Gerrit|ee6bbbf029fa}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:44 wm-bot2: Draining 'cloudvirt1037.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:44 wm-bot2: Safe rebooting 'cloudvirt1037.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:32 wm-bot2: Set cloudvirt 'cloudvirt1039.eqiad.wmnet' maintenance (downtime id: d561fe45-1582-4cd8-bbc9-{{Gerrit|a356c39f3330}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:32 wm-bot2: Set cloudvirt 'cloudvirt1038.eqiad.wmnet' maintenance (downtime id: 351c365c-0228-4907-a279-{{Gerrit|01795b1e62ce}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:32 wm-bot2: Set cloudvirt 'cloudvirt1037.eqiad.wmnet' maintenance (downtime id: 3beec1dd-0132-4f5c-adce-{{Gerrit|5a7dc56060b6}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:31 wm-bot2: Draining 'cloudvirt1039.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:31 wm-bot2: Safe rebooting 'cloudvirt1039.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:31 wm-bot2: Draining 'cloudvirt1038.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:31 wm-bot2: Safe rebooting 'cloudvirt1038.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:31 wm-bot2: Draining 'cloudvirt1037.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:31 wm-bot2: Safe rebooting 'cloudvirt1037.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:30 wm-bot2: Safe reboot of 'cloudvirt1040.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:29 wm-bot2: Unset cloudvirt 'cloudvirt1040.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:28 wm-bot2: Safe reboot of 'cloudvirt1041.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:28 wm-bot2: Unset cloudvirt 'cloudvirt1041.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:26 wm-bot2: Drained 'cloudvirt1040.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:24 wm-bot2: Drained 'cloudvirt1041.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:20 wm-bot2: Safe reboot of 'cloudvirt1042.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:20 wm-bot2: Unset cloudvirt 'cloudvirt1042.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:16 wm-bot2: Drained 'cloudvirt1042.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:07 wm-bot2: Set cloudvirt 'cloudvirt1042.eqiad.wmnet' maintenance (downtime id: 353a8527-ad5e-4898-937c-{{Gerrit|303d16801e28}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:07 wm-bot2: Set cloudvirt 'cloudvirt1041.eqiad.wmnet' maintenance (downtime id: fc235146-f070-4723-9503-{{Gerrit|e20cbb877755}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:07 wm-bot2: Set cloudvirt 'cloudvirt1040.eqiad.wmnet' maintenance (downtime id: cdf712ac-c083-4a63-af83-{{Gerrit|514ca5cdc76d}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:06 wm-bot2: Draining 'cloudvirt1042.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:06 wm-bot2: Safe rebooting 'cloudvirt1042.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:06 wm-bot2: Draining 'cloudvirt1041.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:06 wm-bot2: Safe rebooting 'cloudvirt1041.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:06 wm-bot2: Draining 'cloudvirt1040.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:06 wm-bot2: Safe rebooting 'cloudvirt1040.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:05 wm-bot2: Safe reboot of 'cloudvirt1043.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:05 wm-bot2: Unset cloudvirt 'cloudvirt1043.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:03 wm-bot2: Safe reboot of 'cloudvirt1044.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:03 wm-bot2: Unset cloudvirt 'cloudvirt1044.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:02 wm-bot2: Safe reboot of 'cloudvirt1045.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:02 wm-bot2: Unset cloudvirt 'cloudvirt1045.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:01 wm-bot2: Drained 'cloudvirt1043.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:01 wm-bot2: Set cloudvirt 'cloudvirt1043.eqiad.wmnet' maintenance (downtime id: 8f2d2d32-e99e-4e2c-9472-{{Gerrit|ca33b577d01f}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:00 wm-bot2: Draining 'cloudvirt1043.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:00 wm-bot2: Safe rebooting 'cloudvirt1043.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:00 wm-bot2: Drained 'cloudvirt1044.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:00 wm-bot2: Set cloudvirt 'cloudvirt1044.eqiad.wmnet' maintenance (downtime id: 66ee7d87-cee1-4c8f-b5c3-{{Gerrit|5f58a6cea336}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:59 wm-bot2: Draining 'cloudvirt1044.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:59 wm-bot2: Safe rebooting 'cloudvirt1044.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:58 wm-bot2: Drained 'cloudvirt1044.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:58 wm-bot2: Drained 'cloudvirt1045.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:58 wm-bot2: Set cloudvirt 'cloudvirt1045.eqiad.wmnet' maintenance (downtime id: 69035b38-9910-46f1-9582-{{Gerrit|3105c1ba4d16}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:57 wm-bot2: Draining 'cloudvirt1045.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:57 wm-bot2: Safe rebooting 'cloudvirt1045.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:56 wm-bot2: Safe reboot of 'cloudvirt1046.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:56 wm-bot2: Unset cloudvirt 'cloudvirt1046.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:52 wm-bot2: Drained 'cloudvirt1046.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:52 wm-bot2: Set cloudvirt 'cloudvirt1046.eqiad.wmnet' maintenance (downtime id: 5a94c9d1-c956-4b03-a92b-{{Gerrit|b112810906a4}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:51 wm-bot2: Draining 'cloudvirt1046.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:51 wm-bot2: Safe rebooting 'cloudvirt1046.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:50 wm-bot2: Set cloudvirt 'cloudvirt1046.eqiad.wmnet' maintenance (downtime id: 216a9120-4b9a-4be3-99db-{{Gerrit|3fd5fdc25146}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:49 wm-bot2: Draining 'cloudvirt1046.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:49 wm-bot2: Safe rebooting 'cloudvirt1046.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:48 wm-bot2: Set cloudvirt 'cloudvirt1046.eqiad.wmnet' maintenance (downtime id: 46fb1964-3cf4-47fd-ad79-{{Gerrit|475f33c8ed38}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:48 wm-bot2: Draining 'cloudvirt1046.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:48 wm-bot2: Safe rebooting 'cloudvirt1046.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:47 wm-bot2: Set cloudvirt 'cloudvirt1046.eqiad.wmnet' maintenance (downtime id: b04a4f8c-db82-4258-8608-{{Gerrit|2a78ec9193ed}}, use this to unset). - cookbook ran by andrew@buster * 19:46 wm-bot2: Draining 'cloudvirt1046.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:46 wm-bot2: Drained 'cloudvirt1045.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:46 wm-bot2: Set cloudvirt 'cloudvirt1045.eqiad.wmnet' maintenance (downtime id: 7ae7c258-ecb8-47d1-a991-{{Gerrit|3fa617e305bf}}, use this to unset). - cookbook ran by andrew@buster * 19:45 wm-bot2: Draining 'cloudvirt1045.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:44 wm-bot2: Set cloudvirt 'cloudvirt1043.eqiad.wmnet' maintenance (downtime id: 377fb0e5-132e-4e7a-9079-{{Gerrit|37937c3d42bf}}, use this to unset). - cookbook ran by andrew@buster * 19:44 wm-bot2: Set cloudvirt 'cloudvirt1044.eqiad.wmnet' maintenance (downtime id: 6c4222ae-796e-4890-8820-{{Gerrit|9e7ce5438f2a}}, use this to unset). - cookbook ran by andrew@buster * 19:43 wm-bot2: Draining 'cloudvirt1043.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:43 wm-bot2: Draining 'cloudvirt1044.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:38 wm-bot2: Drained 'cloudvirt1045.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:34 wm-bot2: Safe reboot of 'cloudvirt1047.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:34 wm-bot2: Unset cloudvirt 'cloudvirt1047.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:31 wm-bot2: Set cloudvirt 'cloudvirt1046.eqiad.wmnet' maintenance (downtime id: eff98e2e-ae99-41d0-a68e-{{Gerrit|671d941fcf21}}, use this to unset). - cookbook ran by andrew@buster * 19:30 wm-bot2: Drained 'cloudvirt1047.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:30 wm-bot2: Set cloudvirt 'cloudvirt1047.eqiad.wmnet' maintenance (downtime id: 5310f1eb-0425-4f59-8fb5-{{Gerrit|1d9c6f6d2f23}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:30 wm-bot2: Draining 'cloudvirt1046.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:30 wm-bot2: Draining 'cloudvirt1047.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:30 wm-bot2: Safe rebooting 'cloudvirt1047.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:29 wm-bot2: Drained 'cloudvirt1047.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:25 wm-bot2: Set cloudvirt 'cloudvirt1045.eqiad.wmnet' maintenance (downtime id: d8ab4682-026a-4ddc-bfc5-{{Gerrit|48166bd9789a}}, use this to unset). - cookbook ran by andrew@buster * 19:24 wm-bot2: Draining 'cloudvirt1045.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:23 wm-bot2: Safe reboot of 'cloudvirt1048.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:23 wm-bot2: Unset cloudvirt 'cloudvirt1048.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:19 wm-bot2: Drained 'cloudvirt1048.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:19 wm-bot2: Set cloudvirt 'cloudvirt1048.eqiad.wmnet' maintenance (downtime id: eb2f9d94-8ea5-48d5-a336-{{Gerrit|97dd407b1ce3}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:19 wm-bot2: Draining 'cloudvirt1048.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:19 wm-bot2: Safe rebooting 'cloudvirt1048.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:18 wm-bot2: Drained 'cloudvirt1048.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:14 wm-bot2: Set cloudvirt 'cloudvirt1046.eqiad.wmnet' maintenance (downtime id: 305fe2d7-913e-4638-95c2-{{Gerrit|4e810dc6b2a3}}, use this to unset). - cookbook ran by andrew@buster * 19:14 wm-bot2: Set cloudvirt 'cloudvirt1047.eqiad.wmnet' maintenance (downtime id: 5bb3a66b-9450-4230-8d09-{{Gerrit|d78473d72ba2}}, use this to unset). - cookbook ran by andrew@buster * 19:13 wm-bot2: Draining 'cloudvirt1046.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:13 wm-bot2: Draining 'cloudvirt1047.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:07 wm-bot2: Set cloudvirt 'cloudvirt1048.eqiad.wmnet' maintenance (downtime id: b1e39b49-3a88-4e42-913c-{{Gerrit|9892eafad447}}, use this to unset). - cookbook ran by andrew@buster * 19:06 wm-bot2: Draining 'cloudvirt1048.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:05 andrewbogott: putting cloudvirt1049-1052 into 'ceph' pool, taking out of 'spare' pool. cloudvirt1053 will remain our only spare. * 19:02 wm-bot2: Safe reboot of 'cloudvirt1049.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:02 wm-bot2: Unset cloudvirt 'cloudvirt1049.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:00 wm-bot2: Safe reboot of 'cloudvirt1050.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:00 wm-bot2: Unset cloudvirt 'cloudvirt1050.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:59 wm-bot2: Safe reboot of 'cloudvirt1051.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:59 wm-bot2: Unset cloudvirt 'cloudvirt1051.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:58 wm-bot2: Drained 'cloudvirt1049.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:58 wm-bot2: Set cloudvirt 'cloudvirt1049.eqiad.wmnet' maintenance (downtime id: 8f24ea75-e107-4a13-8bd6-{{Gerrit|950b941b8e03}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:58 wm-bot2: Draining 'cloudvirt1049.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:58 wm-bot2: Safe rebooting 'cloudvirt1049.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:57 wm-bot2: Safe reboot of 'cloudvirt1052.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:57 wm-bot2: Unset cloudvirt 'cloudvirt1052.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:56 wm-bot2: Drained 'cloudvirt1050.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:56 wm-bot2: Set cloudvirt 'cloudvirt1050.eqiad.wmnet' maintenance (downtime id: 5c3e8fa4-cd1e-43f4-a423-{{Gerrit|ba7fe2459d4e}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:55 wm-bot2: Drained 'cloudvirt1051.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:55 wm-bot2: Set cloudvirt 'cloudvirt1051.eqiad.wmnet' maintenance (downtime id: e3fe22d6-932c-4502-8838-{{Gerrit|82cdc4c8f1c6}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:55 wm-bot2: Draining 'cloudvirt1050.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:55 wm-bot2: Safe rebooting 'cloudvirt1050.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:54 wm-bot2: Draining 'cloudvirt1051.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:54 wm-bot2: Safe rebooting 'cloudvirt1051.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:53 wm-bot2: Safe reboot of 'cloudvirt1053.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:53 wm-bot2: Unset cloudvirt 'cloudvirt1053.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:52 wm-bot2: Drained 'cloudvirt1052.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:52 wm-bot2: Set cloudvirt 'cloudvirt1052.eqiad.wmnet' maintenance (downtime id: fd824b92-b6ff-4631-a2d4-{{Gerrit|debb25d48223}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:52 wm-bot2: Draining 'cloudvirt1052.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:52 wm-bot2: Safe rebooting 'cloudvirt1052.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:51 wm-bot2: Drained 'cloudvirt1052.eqiad.wmnet'. - cookbook ran by andrew@buster * 18:49 wm-bot2: Set cloudvirt 'cloudvirt1052.eqiad.wmnet' maintenance (downtime id: 52d5cd5b-0f56-453a-85f8-{{Gerrit|0b5e656ae7f9}}, use this to unset). - cookbook ran by andrew@buster * 18:49 wm-bot2: Drained 'cloudvirt1053.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:49 wm-bot2: Set cloudvirt 'cloudvirt1053.eqiad.wmnet' maintenance (downtime id: b92f66a3-a487-4515-b68b-{{Gerrit|1c257a57b893}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:48 wm-bot2: Draining 'cloudvirt1052.eqiad.wmnet'. - cookbook ran by andrew@buster * 18:48 wm-bot2: Draining 'cloudvirt1053.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:48 wm-bot2: Safe rebooting 'cloudvirt1053.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster === 2022-09-19 === * 20:07 wm-bot2: Safe reboot of 'cloudvirt1026.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:07 wm-bot2: Unset cloudvirt 'cloudvirt1026.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:03 wm-bot2: Drained 'cloudvirt1026.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:03 wm-bot2: Set cloudvirt 'cloudvirt1026.eqiad.wmnet' maintenance (downtime id: 4bfca6cd-dfbc-4a79-9203-{{Gerrit|b432bfeee5c2}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:02 wm-bot2: Draining 'cloudvirt1026.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:02 wm-bot2: Safe rebooting 'cloudvirt1026.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:41 wm-bot2: Drained 'cloudvirt1026.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:21 wm-bot2: Set cloudvirt 'cloudvirt1026.eqiad.wmnet' maintenance (downtime id: 70479325-609b-4349-9094-{{Gerrit|739eb9b3bb30}}, use this to unset). - cookbook ran by andrew@buster * 19:21 wm-bot2: Draining 'cloudvirt1026.eqiad.wmnet'. - cookbook ran by andrew@buster === 2022-09-14 === * 16:57 wm-bot2: Finished rebooting node cloudcephosd1030.eqiad.wmnet - cookbook ran by fran@Francesco’s-MacBook-Pro * 16:53 wm-bot2: Rebooting node cloudcephosd1030.eqiad.wmnet - cookbook ran by fran@Francesco’s-MacBook-Pro * 16:53 wm-bot2: Adding OSD cloudcephosd1030.eqiad.wmnet... (1/1) - cookbook ran by fran@Francesco’s-MacBook-Pro * 16:53 wm-bot2: Adding new OSDs ['cloudcephosd1030.eqiad.wmnet'] to the cluster - cookbook ran by fran@Francesco’s-MacBook-Pro * 11:01 wm-bot2: Finished rebooting node cloudcephosd1030.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 10:58 wm-bot2: Rebooting node cloudcephosd1030.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 10:57 wm-bot2: Adding OSD cloudcephosd1030.eqiad.wmnet... (1/1) - cookbook ran by dcaro@vulcanus * 10:57 wm-bot2: Adding new OSDs ['cloudcephosd1030.eqiad.wmnet'] to the cluster - cookbook ran by dcaro@vulcanus * 10:09 wm-bot2: Adding OSD cloudcephosd1030.eqiad.wmnet... (1/1) - cookbook ran by dcaro@vulcanus * 10:09 wm-bot2: Adding new OSDs ['cloudcephosd1030.eqiad.wmnet'] to the cluster - cookbook ran by dcaro@vulcanus * 10:06 wm-bot2: Adding OSD cloudcephosd1030.eqiad.wmnet... (1/1) - cookbook ran by dcaro@vulcanus * 10:06 wm-bot2: Adding new OSDs ['cloudcephosd1030.eqiad.wmnet'] to the cluster - cookbook ran by dcaro@vulcanus * 09:46 wm-bot2: Finished rebooting node cloudcephosd1030.eqiad.wmnet - cookbook ran by fran@Francesco’s-MacBook-Pro * 09:43 wm-bot2: Rebooting node cloudcephosd1030.eqiad.wmnet - cookbook ran by fran@Francesco’s-MacBook-Pro * 09:43 wm-bot2: Adding OSD cloudcephosd1030.eqiad.wmnet... (1/1) - cookbook ran by fran@Francesco’s-MacBook-Pro * 09:43 wm-bot2: Adding new OSDs ['cloudcephosd1030.eqiad.wmnet'] to the cluster - cookbook ran by fran@Francesco’s-MacBook-Pro === 2022-09-13 === * 12:16 wm-bot2: Adding OSD cloudcephosd1030.eqiad.wmnet... (1/1) - cookbook ran by dcaro@vulcanus * 12:16 wm-bot2: Adding new OSDs ['cloudcephosd1030.eqiad.wmnet'] to the cluster - cookbook ran by dcaro@vulcanus * 12:11 wm-bot2: Adding OSD cloudcephosd1030.eqiad.wmnet... (1/1) - cookbook ran by dcaro@vulcanus * 12:11 wm-bot2: Adding new OSDs ['cloudcephosd1030.eqiad.wmnet'] to the cluster - cookbook ran by dcaro@vulcanus * 12:10 wm-bot2: Adding OSD cloudcephosd1030.eqiad.wmnet... (1/1) - cookbook ran by dcaro@vulcanus * 12:10 wm-bot2: Adding new OSDs ['cloudcephosd1030.eqiad.wmnet'] to the cluster - cookbook ran by dcaro@vulcanus * 10:40 wm-bot2: Finished rebooting node cloudcephosd1030.eqiad.wmnet - cookbook ran by fran@Francesco’s-MacBook-Pro * 10:37 wm-bot2: Rebooting node cloudcephosd1030.eqiad.wmnet - cookbook ran by fran@Francesco’s-MacBook-Pro * 10:37 wm-bot2: Adding OSD cloudcephosd1030.eqiad.wmnet... (1/1) - cookbook ran by fran@Francesco’s-MacBook-Pro * 10:37 wm-bot2: Adding new OSDs ['cloudcephosd1030.eqiad.wmnet'] to the cluster - cookbook ran by fran@Francesco’s-MacBook-Pro * 10:36 wm-bot2: Adding new OSDs ['cloudcephosd1030.eqiad.wmnet'] to the cluster - cookbook ran by fran@Francesco’s-MacBook-Pro === 2022-09-10 === * 15:37 andrewbogott: restarting nova-conductor service (and possibly others, in response to lots of unanswered rabbitmq messages) === 2022-09-08 === * 19:12 andrewbogott: restarting nginx on proxy-03.project-proxy.eqiad1.wikimedia.cloud === 2022-09-07 === * 10:18 wm-bot2: Added OSD cloudcephosd1032.eqiad.wmnet... (1/1) ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 10:14 wm-bot2: Finished rebooting node cloudcephosd1032.eqiad.wmnet ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 10:10 wm-bot2: Rebooting node cloudcephosd1032.eqiad.wmnet ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 10:10 wm-bot2: Adding OSD cloudcephosd1032.eqiad.wmnet... (1/1) ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 10:10 wm-bot2: Adding new OSDs ['cloudcephosd1032.eqiad.wmnet'] to the cluster ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 09:41 dhinus: Temporarily removing cloudcephosd1030 from Ceph cluster (https://phabricator.wikimedia.org/T314870) === 2022-08-30 === * 14:59 andrewbogott: manually marking most eqiad1 cloud* servers down in icinga for [[phab:T296561|T296561]] * 10:43 wm-bot2: Finished rebooting node cloudcephosd1030.eqiad.wmnet ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 10:39 wm-bot2: Rebooting node cloudcephosd1030.eqiad.wmnet ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 10:38 wm-bot2: Adding OSD cloudcephosd1030.eqiad.wmnet... (1/1) ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 10:38 wm-bot2: Adding new OSDs ['cloudcephosd1030.eqiad.wmnet'] to the cluster ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 09:39 wm-bot2: Finished rebooting node cloudcephosd1030.eqiad.wmnet ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 09:32 wm-bot2: Rebooting node cloudcephosd1030.eqiad.wmnet ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 09:32 wm-bot2: Adding OSD cloudcephosd1030.eqiad.wmnet... (1/1) ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 09:32 wm-bot2: Adding new OSDs ['cloudcephosd1030.eqiad.wmnet'] to the cluster ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 09:14 wm-bot2: Finished rebooting node cloudcephosd1030.eqiad.wmnet ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 09:11 wm-bot2: Rebooting node cloudcephosd1030.eqiad.wmnet ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 09:10 wm-bot2: Adding OSD cloudcephosd1030.eqiad.wmnet... (1/1) ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 09:10 wm-bot2: Adding new OSDs ['cloudcephosd1030.eqiad.wmnet'] to the cluster ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro === 2022-08-25 === * 15:14 wm-bot2: Added 1 new OSDs ['cloudcephosd1029.eqiad.wmnet'] ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 15:14 wm-bot2: Added OSD cloudcephosd1029.eqiad.wmnet... (1/1) ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 15:02 wm-bot2: Finished rebooting node cloudcephosd1029.eqiad.wmnet ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 14:59 wm-bot2: Rebooting node cloudcephosd1029.eqiad.wmnet ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 14:58 wm-bot2: Adding OSD cloudcephosd1029.eqiad.wmnet... (1/1) ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 14:58 wm-bot2: Adding new OSDs ['cloudcephosd1029.eqiad.wmnet'] to the cluster ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro === 2022-08-24 === * 22:07 andrewbogott: replaced cloudservices1003 with cloudservices1005 [[phab:T304888|T304888]] * 10:45 wm-bot2: Added 1 new OSDs ['cloudcephosd1028.eqiad.wmnet'] ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 10:45 wm-bot2: Added OSD cloudcephosd1028.eqiad.wmnet... (1/1) ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 10:37 wm-bot2: Finished rebooting node cloudcephosd1028.eqiad.wmnet ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 10:34 wm-bot2: Rebooting node cloudcephosd1028.eqiad.wmnet ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 10:33 wm-bot2: Adding OSD cloudcephosd1028.eqiad.wmnet... (1/1) ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 10:33 wm-bot2: Adding new OSDs ['cloudcephosd1028.eqiad.wmnet'] to the cluster ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro === 2022-08-23 === * 13:46 wm-bot2: Added 1 new OSDs ['cloudcephosd1027.eqiad.wmnet'] ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 13:46 wm-bot2: Added OSD cloudcephosd1027.eqiad.wmnet... (1/1) ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 13:27 wm-bot2: Finished rebooting node cloudcephosd1027.eqiad.wmnet ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 13:24 wm-bot2: Rebooting node cloudcephosd1027.eqiad.wmnet ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 13:22 wm-bot2: Adding OSD cloudcephosd1027.eqiad.wmnet... (1/1) ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 13:22 wm-bot2: Adding new OSDs ['cloudcephosd1027.eqiad.wmnet'] to the cluster ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro === 2022-08-21 === * 21:12 andrewbogott: restarted neutron-dhcp-agent on cloudnet1003. it was claiming to be unable to contact Rabbit but seems happy after a restart === 2022-08-20 === * 07:39 dcaro_away: cloudvirt1023 is back up, VMs are starting to recover ([[phab:T315718|T315718]]) * 07:23 dcaro_away: cloudvirt1023 seems to have gotten some hardware issue from racadm lclog view "System CPU Resetting.", rebooting and doing memory checks ([[phab:T315718|T315718]]) === 2022-08-19 === * 17:06 taavi: [codfw1dev] restart mariadb on clouddb2002-dev to pick up certificate config changes [[phab:T310795|T310795]] === 2022-08-18 === * 13:25 wm-bot2: Added 1 new OSDs ['cloudcephosd1026.eqiad.wmnet'] ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 13:25 wm-bot2: Added OSD cloudcephosd1026.eqiad.wmnet... (1/1) ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 13:15 wm-bot2: Finished rebooting node cloudcephosd1026.eqiad.wmnet ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 13:12 wm-bot2: Rebooting node cloudcephosd1026.eqiad.wmnet ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 13:10 wm-bot2: Adding OSD cloudcephosd1026.eqiad.wmnet... (1/1) ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 13:10 wm-bot2: Adding new OSDs ['cloudcephosd1026.eqiad.wmnet'] to the cluster ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 07:29 dcaro: Starting up all the osd daemons on cloudcephosd1025 ([[phab:T314870|T314870]]) === 2022-08-17 === * 10:50 wm-bot2: Added 1 new OSDs ['cloudcephosd1025.eqiad.wmnet'] ([[phab:T314870|T314870]]) - cookbook ran by fran@foz * 10:49 wm-bot2: Added OSD cloudcephosd1025.eqiad.wmnet... (1/1) ([[phab:T314870|T314870]]) - cookbook ran by fran@foz * 10:40 wm-bot2: Finished rebooting node cloudcephosd1025.eqiad.wmnet ([[phab:T314870|T314870]]) - cookbook ran by fran@foz * 10:37 wm-bot2: Rebooting node cloudcephosd1025.eqiad.wmnet ([[phab:T314870|T314870]]) - cookbook ran by fran@foz * 10:37 wm-bot2: Adding OSD cloudcephosd1025.eqiad.wmnet... (1/1) ([[phab:T314870|T314870]]) - cookbook ran by fran@foz * 10:37 wm-bot2: Adding new OSDs ['cloudcephosd1025.eqiad.wmnet'] to the cluster ([[phab:T314870|T314870]]) - cookbook ran by fran@foz * 09:50 wm-bot2: Rebooting node cloudcephosd1025.eqiad.wmnet ([[phab:T314870|T314870]]) - cookbook ran by fran@foz * 09:49 wm-bot2: Adding OSD cloudcephosd1025.eqiad.wmnet... (1/1) ([[phab:T314870|T314870]]) - cookbook ran by fran@foz * 09:49 wm-bot2: Adding new OSDs ['cloudcephosd1025.eqiad.wmnet'] to the cluster ([[phab:T314870|T314870]]) - cookbook ran by fran@foz * 09:16 wm-bot2: Rebooting node cloudcephosd1025.eqiad.wmnet ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 09:16 wm-bot2: Adding OSD cloudcephosd1025.eqiad.wmnet... (1/1) ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 09:16 wm-bot2: Adding new OSDs ['cloudcephosd1025.eqiad.wmnet'] to the cluster ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro === 2022-08-16 === * 22:39 andrewbogott: replacing the now-rebuilt cloudvirt1025 in 'ceph' aggregate and removing it from the 'maintenance' aggregate * 17:41 andrewbogott: removing cloudvirt1025 from the 'ceph' aggregate and adding it to the 'maintenance' aggregate * 17:40 andrewbogott: reimaging cloudvirt1025 after I accidentally deleted the hw raid * 17:38 andrewbogott: root@cloudcontrol1005:~# cinder-manage volume update_host --currenthost cloudcontrol1003@rbd#RBD --newhost cloudcontrol1005@rbd#RBD * 17:37 andrewbogott: root@cloudcontrol1005:~# cinder-manage volume update_host --currenthost cloudcontrol1004@rbd#RBD --newhost cloudcontrol1006@rbd#RBD * 16:26 wm-bot2: Ceph cluster at eqiad1 set out of maintenance. - cookbook ran by dcaro@vulcanus * 15:43 wm-bot2: Restarting the osd daemons from nodes cloudcephosd1001,cloudcephosd1002,cloudcephosd1003,cloudcephosd1004,cloudcephosd1005,cloudcephosd1006,cloudcephosd1007,cloudcephosd1008,cloudcephosd1009,cloudcephosd1010,cloudcephosd1011,cloudcephosd1012,cloudcephosd1013,cloudcephosd1014,cloudcephosd1015,cloudcephosd1016,cloudcephosd1017,cloudcephosd1018,cloudcephosd1019,cloudcephosd1020,cloudcephosd1021,cloudcephosd1022,cloudcephosd1023,c * 15:42 wm-bot2: Finished restarting all the OSD daemons from the nodes ['cloudcephosd2001-dev', 'cloudcephosd2002-dev', 'cloudcephosd2003-dev'] - cookbook ran by dcaro@vulcanus * 15:38 wm-bot2: Restarting the osd daemons from nodes cloudcephosd2001-dev,cloudcephosd2002-dev,cloudcephosd2003-dev - cookbook ran by dcaro@vulcanus * 13:08 wm-bot2: Restarting the osd daemons from nodes cloudcephosd2001-dev,cloudcephosd2002-dev,cloudcephosd2003-dev - cookbook ran by dcaro@vulcanus * 13:07 wm-bot2: Restarting the osd daemons from nodes cloudcephosd2001-dev,cloudcephosd2002-dev,cloudcephosd2003-dev - cookbook ran by dcaro@vulcanus * 13:02 wm-bot2: Restarting the osd daemons from nodes cloudcephosd2001-dev,cloudcephosd2002-dev,cloudcephosd2003-dev - cookbook ran by dcaro@vulcanus * 13:01 wm-bot2: Restarting the osd daemons from nodes cloudcephosd2001-dev,cloudcephosd2002-dev,cloudcephosd2003-dev - cookbook ran by dcaro@vulcanus * 12:59 wm-bot2: Restarting the osd daemons from nodes cloudcephosd2001-dev,cloudcephosd2002-dev,cloudcephosd2003-dev - cookbook ran by dcaro@vulcanus === 2022-08-14 === * 18:36 taavi: deleted the http keystone endpoints from the keystone service catalog === 2022-08-11 === * 13:57 andrewbogott: decommissioning cloudcontrol1003 + cloudcontrl1004. I backed up $home in case anyone needs their files. * 08:42 wm-bot2: The cluster is now rebalanced after adding the new OSDs ['cloudcephosd1025.eqiad.wmnet'] ([[phab:T314870|T314870]]) - cookbook ran by fran@MacBook-Pro.station * 08:42 wm-bot2: Added 1 new OSDs ['cloudcephosd1025.eqiad.wmnet'] ([[phab:T314870|T314870]]) - cookbook ran by fran@MacBook-Pro.station * 08:42 wm-bot2: Added OSD cloudcephosd1025.eqiad.wmnet... (1/1) ([[phab:T314870|T314870]]) - cookbook ran by fran@MacBook-Pro.station * 08:40 wm-bot2: Finished rebooting node cloudcephosd1025.eqiad.wmnet ([[phab:T314870|T314870]]) - cookbook ran by fran@MacBook-Pro.station * 08:36 wm-bot2: Rebooting node cloudcephosd1025.eqiad.wmnet ([[phab:T314870|T314870]]) - cookbook ran by fran@MacBook-Pro.station * 08:36 wm-bot2: Adding OSD cloudcephosd1025.eqiad.wmnet... (1/1) ([[phab:T314870|T314870]]) - cookbook ran by fran@MacBook-Pro.station * 08:36 wm-bot2: Adding new OSDs ['cloudcephosd1025.eqiad.wmnet'] to the cluster ([[phab:T314870|T314870]]) - cookbook ran by fran@MacBook-Pro.station === 2022-08-10 === * 13:10 wm-bot2: Finished rebooting node cloudcephosd1025.eqiad.wmnet ([[phab:T314870|T314870]]) - cookbook ran by fran@MacBook-Pro.station * 13:06 wm-bot2: Rebooting node cloudcephosd1025.eqiad.wmnet ([[phab:T314870|T314870]]) - cookbook ran by fran@MacBook-Pro.station * 13:06 wm-bot2: Adding OSD cloudcephosd1025.eqiad.wmnet... (1/1) ([[phab:T314870|T314870]]) - cookbook ran by fran@MacBook-Pro.station * 13:06 wm-bot2: Adding new OSDs ['cloudcephosd1025.eqiad.wmnet'] to the cluster ([[phab:T314870|T314870]]) - cookbook ran by fran@MacBook-Pro.station === 2022-08-04 === * 17:16 taavi: deleted all scheduler_fanout_ rabbit queues in an attempt to fix scheduling * 16:32 taavi: restart neutron-l3-agent to pick up rabbit config changes * 15:12 andrewbogott: stopping rabbitmq on cloudcontrol1xxx * 09:57 taavi: stop wikitech_run_jobs.timer on labweb1001/1002, hosts pending decom === 2022-08-03 === * 20:55 andrewbogott: root@tools-checker-04:~# systemctl restart uwsgi-toolschecker_cron.service * 20:41 andrewbogott: restarting neutron-l3-agent.service on cloudnet1003 and 1004. The agent was routing properly but had lost touch with rabbitmq === 2022-08-02 === * 14:07 andrewbogott: shutting down codfw1dev ceph cluster according to https://docs.mirantis.com/mcp/q4-18/mcp-operations-guide/scheduled-maintenance-power-outage/power-off-ceph-cluster.html * 13:54 andrewbogott: shutting down basically all of codfw1dev to support pdu maintenance -- all the ceph OSDs will lose power so best to have everything stopped. === 2022-07-27 === * 19:32 andrewbogott: switching the openstack.eqiad1.wikimedia.cloud endpoint from cloudcontrol1004 to 1006, https://gerrit.wikimedia.org/r/c/operations/dns/+/817878/2/templates/wikimediacloud.org#54 * 16:33 andrewbogott: here is a test message in the admin channel === 2022-07-25 === * 13:43 andrewbogott: pooling cloudweb100[34] and depooling labweb100[12] for testing in prep for decomming labweb100[12] === 2022-07-22 === * 16:41 taavi: depool cloudweb1003/1004 since horizon seems to be having issues * 16:22 taavi: pooling cloudweb1003/1004 now that grant issues are sorted === 2022-07-21 === * 18:26 andrewbogott: depooling cloudweb1003 and 1004 for wikitech, horizon, striker -- pending db grant changes * 18:06 andrewbogott: pooling cloudweb1003 and 1004 for wikitech, horizon, striker === 2022-07-20 === * 18:02 dcaro: things seem stable, trying to bring up a the last rabbit node, cloudcontrol1007 ([[phab:T313400|T313400]]) * 17:45 bd808: `sudo service striker restart` on labweb1002 * 17:43 bd808: `sudo service striker restart` on labweb1001 * 17:10 dcaro: things seem stable, trying to bring up a fourth rabbit node, cloudcontrol1006 ([[phab:T313400|T313400]]) * 16:26 dcaro: things seem stable, trying to bring up a third, cloudcontrol1005 ([[phab:T313400|T313400]]) * 15:51 dcaro: things seem stable now with one rabbit node, trying to bring up a second ([[phab:T313400|T313400]]) * 14:16 dcaro: stopping rabbin on cloudcontrol1004, leaving only 1003 alive ([[phab:T313400|T313400]]) * 13:17 dcaro: restarting the whole rabbit cluster ([[phab:T313400|T313400]]) === 2022-07-19 === * 16:30 wm-bot2: Safe reboot of 'cloudvirt1045.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 16:30 wm-bot2: Unset cloudvirt 'cloudvirt1045.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 16:26 wm-bot2: Drained 'cloudvirt1045.eqiad.wmnet'. - cookbook ran by andrew@buster * 16:18 wm-bot2: Safe reboot of 'cloudvirt1044.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 16:18 wm-bot2: Unset cloudvirt 'cloudvirt1044.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 16:14 wm-bot2: Drained 'cloudvirt1044.eqiad.wmnet'. - cookbook ran by andrew@buster * 16:01 wm-bot2: Safe reboot of 'cloudvirt1047.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 16:01 wm-bot2: Unset cloudvirt 'cloudvirt1047.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 15:57 wm-bot2: Drained 'cloudvirt1047.eqiad.wmnet'. - cookbook ran by andrew@buster * 15:57 wm-bot2: Set cloudvirt 'cloudvirt1047.eqiad.wmnet' maintenance (downtime id: 3da2d4f6-5c5b-4a21-9a0b-{{Gerrit|2b010960ed6a}}, use this to unset). - cookbook ran by andrew@buster * 15:56 wm-bot2: Draining 'cloudvirt1047.eqiad.wmnet'. - cookbook ran by andrew@buster * 15:56 wm-bot2: Safe rebooting 'cloudvirt1047.eqiad.wmnet'. - cookbook ran by andrew@buster * 15:56 wm-bot2: Safe reboot of 'cloudvirt1046.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 15:56 wm-bot2: Unset cloudvirt 'cloudvirt1046.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 15:52 wm-bot2: Drained 'cloudvirt1046.eqiad.wmnet'. - cookbook ran by andrew@buster * 15:49 wm-bot2: Set cloudvirt 'cloudvirt1045.eqiad.wmnet' maintenance (downtime id: 8a10505a-107d-4e78-ba84-{{Gerrit|363a7dea8d69}}, use this to unset). - cookbook ran by andrew@buster * 15:47 wm-bot2: Set cloudvirt 'cloudvirt1044.eqiad.wmnet' maintenance (downtime id: 752aa58b-44d4-4340-b05d-{{Gerrit|911b14f2314d}}, use this to unset). - cookbook ran by andrew@buster * 15:47 wm-bot2: Set cloudvirt 'cloudvirt1046.eqiad.wmnet' maintenance (downtime id: f89f0851-ce3f-4a92-899e-{{Gerrit|0d4638acc9c3}}, use this to unset). - cookbook ran by andrew@buster * 15:46 wm-bot2: Draining 'cloudvirt1044.eqiad.wmnet'. - cookbook ran by andrew@buster * 15:46 wm-bot2: Safe rebooting 'cloudvirt1044.eqiad.wmnet'. - cookbook ran by andrew@buster * 15:46 wm-bot2: Draining 'cloudvirt1045.eqiad.wmnet'. - cookbook ran by andrew@buster * 15:46 wm-bot2: Safe rebooting 'cloudvirt1045.eqiad.wmnet'. - cookbook ran by andrew@buster * 15:46 wm-bot2: Draining 'cloudvirt1046.eqiad.wmnet'. - cookbook ran by andrew@buster * 15:46 wm-bot2: Safe rebooting 'cloudvirt1046.eqiad.wmnet'. - cookbook ran by andrew@buster * 15:45 andrewbogott: adding new hosts to the 'ceph' aggregate: cloudvirt1046, 1047, 1048 * 15:44 wm-bot2: Safe reboot of 'cloudvirt1041.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 15:44 wm-bot2: Unset cloudvirt 'cloudvirt1041.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 15:42 wm-bot2: Safe reboot of 'cloudvirt1043.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 15:42 wm-bot2: Unset cloudvirt 'cloudvirt1043.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 15:40 wm-bot2: Drained 'cloudvirt1041.eqiad.wmnet'. - cookbook ran by andrew@buster * 15:38 wm-bot2: Drained 'cloudvirt1043.eqiad.wmnet'. - cookbook ran by andrew@buster * 15:37 wm-bot2: Safe reboot of 'cloudvirt1042.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 15:37 wm-bot2: Unset cloudvirt 'cloudvirt1042.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 15:33 wm-bot2: Drained 'cloudvirt1042.eqiad.wmnet'. - cookbook ran by andrew@buster * 15:32 wm-bot2: Set cloudvirt 'cloudvirt1042.eqiad.wmnet' maintenance (downtime id: 32d9fdb7-6c42-4cf4-9950-{{Gerrit|6e2442255161}}, use this to unset). - cookbook ran by andrew@buster * 15:31 wm-bot2: Draining 'cloudvirt1042.eqiad.wmnet'. - cookbook ran by andrew@buster * 15:31 wm-bot2: Safe rebooting 'cloudvirt1042.eqiad.wmnet'. - cookbook ran by andrew@buster * 15:16 wm-bot2: Set cloudvirt 'cloudvirt1043.eqiad.wmnet' maintenance (downtime id: 56ce1388-792c-442c-bb89-{{Gerrit|e4c3869b0acf}}, use this to unset). - cookbook ran by andrew@buster * 15:15 wm-bot2: Draining 'cloudvirt1043.eqiad.wmnet'. - cookbook ran by andrew@buster * 15:15 wm-bot2: Safe rebooting 'cloudvirt1043.eqiad.wmnet'. - cookbook ran by andrew@buster * 15:14 wm-bot2: Set cloudvirt 'cloudvirt1043.eqiad.wmnet' maintenance (downtime id: 3f84bf2e-24de-4ebe-8117-{{Gerrit|0f4a5de29d09}}, use this to unset). - cookbook ran by andrew@buster * 15:14 wm-bot2: Draining 'cloudvirt1043.eqiad.wmnet'. - cookbook ran by andrew@buster * 15:14 wm-bot2: Safe rebooting 'cloudvirt1043.eqiad.wmnet'. - cookbook ran by andrew@buster * 15:13 wm-bot2: Safe reboot of 'cloudvirt1040.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 15:13 wm-bot2: Unset cloudvirt 'cloudvirt1040.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 15:12 wm-bot2: Set cloudvirt 'cloudvirt1042.eqiad.wmnet' maintenance (downtime id: 5418e923-997c-49ae-b937-{{Gerrit|35fc0c3038eb}}, use this to unset). - cookbook ran by andrew@buster * 15:12 wm-bot2: Set cloudvirt 'cloudvirt1041.eqiad.wmnet' maintenance (downtime id: c3dfdc2b-0725-4d01-87f2-{{Gerrit|440a5ac4694c}}, use this to unset). - cookbook ran by andrew@buster * 15:11 wm-bot2: Draining 'cloudvirt1042.eqiad.wmnet'. - cookbook ran by andrew@buster * 15:11 wm-bot2: Safe rebooting 'cloudvirt1042.eqiad.wmnet'. - cookbook ran by andrew@buster * 15:11 wm-bot2: Draining 'cloudvirt1041.eqiad.wmnet'. - cookbook ran by andrew@buster * 15:11 wm-bot2: Safe rebooting 'cloudvirt1041.eqiad.wmnet'. - cookbook ran by andrew@buster * 15:09 wm-bot2: Drained 'cloudvirt1040.eqiad.wmnet'. - cookbook ran by andrew@buster * 15:04 wm-bot2: Safe reboot of 'cloudvirt1039.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 15:04 wm-bot2: Unset cloudvirt 'cloudvirt1039.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 15:00 wm-bot2: Drained 'cloudvirt1039.eqiad.wmnet'. - cookbook ran by andrew@buster * 14:54 wm-bot2: Safe reboot of 'cloudvirt1038.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 14:54 wm-bot2: Unset cloudvirt 'cloudvirt1038.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 14:50 wm-bot2: Drained 'cloudvirt1038.eqiad.wmnet'. - cookbook ran by andrew@buster * 14:46 wm-bot2: Set cloudvirt 'cloudvirt1038.eqiad.wmnet' maintenance (downtime id: 49dbf7de-c58a-4b68-bc7b-{{Gerrit|f001a047f4c7}}, use this to unset). - cookbook ran by andrew@buster * 14:46 wm-bot2: Draining 'cloudvirt1038.eqiad.wmnet'. - cookbook ran by andrew@buster * 14:46 wm-bot2: Safe rebooting 'cloudvirt1038.eqiad.wmnet'. - cookbook ran by andrew@buster * 14:44 wm-bot2: Set cloudvirt 'cloudvirt1040.eqiad.wmnet' maintenance (downtime id: a4e95168-39b5-452d-8fce-{{Gerrit|7875ae44a62b}}, use this to unset). - cookbook ran by andrew@buster * 14:44 wm-bot2: Draining 'cloudvirt1040.eqiad.wmnet'. - cookbook ran by andrew@buster * 14:44 wm-bot2: Safe rebooting 'cloudvirt1040.eqiad.wmnet'. - cookbook ran by andrew@buster * 14:43 wm-bot2: Set cloudvirt 'cloudvirt1039.eqiad.wmnet' maintenance (downtime id: b46234ee-9900-4708-a815-{{Gerrit|4324379be48c}}, use this to unset). - cookbook ran by andrew@buster * 14:43 wm-bot2: Safe reboot of 'cloudvirt1036.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 14:43 wm-bot2: Unset cloudvirt 'cloudvirt1036.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 14:42 wm-bot2: Draining 'cloudvirt1039.eqiad.wmnet'. - cookbook ran by andrew@buster * 14:42 wm-bot2: Safe rebooting 'cloudvirt1039.eqiad.wmnet'. - cookbook ran by andrew@buster * 14:39 wm-bot2: Safe reboot of 'cloudvirt1037.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 14:39 wm-bot2: Unset cloudvirt 'cloudvirt1037.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 14:38 wm-bot2: Drained 'cloudvirt1036.eqiad.wmnet'. - cookbook ran by andrew@buster * 14:38 wm-bot2: Set cloudvirt 'cloudvirt1036.eqiad.wmnet' maintenance (downtime id: 12bc87a9-7602-4795-818c-{{Gerrit|f7151b1c9ead}}, use this to unset). - cookbook ran by andrew@buster * 14:37 wm-bot2: Draining 'cloudvirt1036.eqiad.wmnet'. - cookbook ran by andrew@buster * 14:37 wm-bot2: Safe rebooting 'cloudvirt1036.eqiad.wmnet'. - cookbook ran by andrew@buster * 14:35 wm-bot2: Drained 'cloudvirt1037.eqiad.wmnet'. - cookbook ran by andrew@buster * 14:32 wm-bot2: Set cloudvirt 'cloudvirt1038.eqiad.wmnet' maintenance (downtime id: ba787a26-62db-4662-ade2-{{Gerrit|2e3061b7390d}}, use this to unset). - cookbook ran by andrew@buster * 14:31 wm-bot2: Draining 'cloudvirt1038.eqiad.wmnet'. - cookbook ran by andrew@buster * 14:31 wm-bot2: Safe rebooting 'cloudvirt1038.eqiad.wmnet'. - cookbook ran by andrew@buster * 14:29 wm-bot2: Set cloudvirt 'cloudvirt1037.eqiad.wmnet' maintenance (downtime id: 6f0e79c6-2d91-478e-92df-{{Gerrit|0f65de48212f}}, use this to unset). - cookbook ran by andrew@buster * 14:29 wm-bot2: Set cloudvirt 'cloudvirt1038.eqiad.wmnet' maintenance (downtime id: 7f577c0f-0182-498f-b198-{{Gerrit|307d51df8c3a}}, use this to unset). - cookbook ran by andrew@buster * 14:28 wm-bot2: Draining 'cloudvirt1037.eqiad.wmnet'. - cookbook ran by andrew@buster * 14:28 wm-bot2: Safe rebooting 'cloudvirt1037.eqiad.wmnet'. - cookbook ran by andrew@buster * 14:28 wm-bot2: Draining 'cloudvirt1038.eqiad.wmnet'. - cookbook ran by andrew@buster * 14:28 wm-bot2: Safe rebooting 'cloudvirt1038.eqiad.wmnet'. - cookbook ran by andrew@buster * 14:28 wm-bot2: Set cloudvirt 'cloudvirt1036.eqiad.wmnet' maintenance (downtime id: bd1bff92-bfb8-483d-a719-{{Gerrit|13fbdf3d8b21}}, use this to unset). - cookbook ran by andrew@buster * 14:27 wm-bot2: Draining 'cloudvirt1036.eqiad.wmnet'. - cookbook ran by andrew@buster * 14:27 wm-bot2: Safe rebooting 'cloudvirt1036.eqiad.wmnet'. - cookbook ran by andrew@buster * 14:18 wm-bot2: Set cloudvirt 'cloudvirt1037.eqiad.wmnet' maintenance (downtime id: c9cd7dbb-c5ad-4364-a8e0-{{Gerrit|afb52ef98683}}, use this to unset). - cookbook ran by andrew@buster * 14:18 wm-bot2: Set cloudvirt 'cloudvirt1038.eqiad.wmnet' maintenance (downtime id: 4deca8a4-f73f-4b42-b4be-{{Gerrit|acfe61421ed0}}, use this to unset). - cookbook ran by andrew@buster * 14:18 wm-bot2: Set cloudvirt 'cloudvirt1036.eqiad.wmnet' maintenance (downtime id: 6d0d58b2-8450-480d-94f2-{{Gerrit|93a257f25575}}, use this to unset). - cookbook ran by andrew@buster * 14:17 wm-bot2: Draining 'cloudvirt1037.eqiad.wmnet'. - cookbook ran by andrew@buster * 14:17 wm-bot2: Safe rebooting 'cloudvirt1037.eqiad.wmnet'. - cookbook ran by andrew@buster * 14:17 wm-bot2: Draining 'cloudvirt1038.eqiad.wmnet'. - cookbook ran by andrew@buster * 14:17 wm-bot2: Safe rebooting 'cloudvirt1038.eqiad.wmnet'. - cookbook ran by andrew@buster * 14:17 wm-bot2: Draining 'cloudvirt1036.eqiad.wmnet'. - cookbook ran by andrew@buster * 14:17 wm-bot2: Safe rebooting 'cloudvirt1036.eqiad.wmnet'. - cookbook ran by andrew@buster * 14:13 dcaro: deleting all the leftover fullstack images (was due to max number of mysql connections reached) * 04:51 wm-bot2: Safe reboot of 'cloudvirt1035.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 04:51 wm-bot2: Unset cloudvirt 'cloudvirt1035.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 04:47 wm-bot2: Drained 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:46 wm-bot2: Set cloudvirt 'cloudvirt1035.eqiad.wmnet' maintenance (downtime id: 432747a2-9a15-44d4-ba8a-{{Gerrit|4e7e87eb8b76}}, use this to unset). - cookbook ran by andrew@buster * 04:45 wm-bot2: Draining 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:45 wm-bot2: Safe rebooting 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:45 wm-bot2: Safe reboot of 'cloudvirt1033.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 04:45 wm-bot2: Unset cloudvirt 'cloudvirt1033.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 04:45 wm-bot2: Set cloudvirt 'cloudvirt1035.eqiad.wmnet' maintenance (downtime id: 1c5adb8c-efcf-4036-bf22-{{Gerrit|0eef364ae399}}, use this to unset). - cookbook ran by andrew@buster * 04:44 wm-bot2: Draining 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:44 wm-bot2: Safe rebooting 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:41 wm-bot2: Set cloudvirt 'cloudvirt1035.eqiad.wmnet' maintenance (downtime id: 26feb897-9eb0-40b6-9525-{{Gerrit|4cf44f81638d}}, use this to unset). - cookbook ran by andrew@buster * 04:41 wm-bot2: Drained 'cloudvirt1033.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:40 wm-bot2: Draining 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:40 wm-bot2: Safe rebooting 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:39 wm-bot2: Set cloudvirt 'cloudvirt1035.eqiad.wmnet' maintenance (downtime id: cd2424d3-c9c9-4837-93b1-{{Gerrit|a6d8271e7873}}, use this to unset). - cookbook ran by andrew@buster * 04:39 wm-bot2: Draining 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:39 wm-bot2: Safe rebooting 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:38 wm-bot2: Safe reboot of 'cloudvirt1034.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 04:38 wm-bot2: Unset cloudvirt 'cloudvirt1034.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 04:38 wm-bot2: Set cloudvirt 'cloudvirt1035.eqiad.wmnet' maintenance (downtime id: 1e85ec01-81c4-4664-b14f-{{Gerrit|b1cf6417988c}}, use this to unset). - cookbook ran by andrew@buster * 04:38 wm-bot2: Set cloudvirt 'cloudvirt1033.eqiad.wmnet' maintenance (downtime id: f498aaa4-fad3-42f8-92f4-{{Gerrit|48a98c108456}}, use this to unset). - cookbook ran by andrew@buster * 04:37 wm-bot2: Draining 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:37 wm-bot2: Safe rebooting 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:37 wm-bot2: Draining 'cloudvirt1033.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:37 wm-bot2: Safe rebooting 'cloudvirt1033.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:35 wm-bot2: Set cloudvirt 'cloudvirt1035.eqiad.wmnet' maintenance (downtime id: 0ed281f3-d9b2-4c0a-bf88-{{Gerrit|db42fe848b2f}}, use this to unset). - cookbook ran by andrew@buster * 04:35 wm-bot2: Set cloudvirt 'cloudvirt1033.eqiad.wmnet' maintenance (downtime id: f417dffc-fa1b-4e65-938a-{{Gerrit|49e54b6fac6d}}, use this to unset). - cookbook ran by andrew@buster * 04:34 wm-bot2: Draining 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:34 wm-bot2: Safe rebooting 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:34 wm-bot2: Draining 'cloudvirt1033.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:34 wm-bot2: Safe rebooting 'cloudvirt1033.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:34 wm-bot2: Drained 'cloudvirt1034.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:33 wm-bot2: Set cloudvirt 'cloudvirt1035.eqiad.wmnet' maintenance (downtime id: 923926e0-7194-4e13-98c4-{{Gerrit|a4f703b2907b}}, use this to unset). - cookbook ran by andrew@buster * 04:32 wm-bot2: Draining 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:32 wm-bot2: Safe rebooting 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:29 wm-bot2: Set cloudvirt 'cloudvirt1033.eqiad.wmnet' maintenance (downtime id: ca8f729b-92ad-45ec-b5da-{{Gerrit|55987b0fc9c2}}, use this to unset). - cookbook ran by andrew@buster * 04:29 wm-bot2: Draining 'cloudvirt1033.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:29 wm-bot2: Safe rebooting 'cloudvirt1033.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:25 wm-bot2: Set cloudvirt 'cloudvirt1034.eqiad.wmnet' maintenance (downtime id: abe6992d-dfba-4e50-993d-{{Gerrit|de0a551a177a}}, use this to unset). - cookbook ran by andrew@buster * 04:24 wm-bot2: Draining 'cloudvirt1034.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:24 wm-bot2: Safe rebooting 'cloudvirt1034.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:23 wm-bot2: Set cloudvirt 'cloudvirt1035.eqiad.wmnet' maintenance (downtime id: df91107d-bf34-49db-9756-{{Gerrit|44ad06162683}}, use this to unset). - cookbook ran by andrew@buster * 04:23 wm-bot2: Set cloudvirt 'cloudvirt1033.eqiad.wmnet' maintenance (downtime id: f45ec0b1-83c9-49e0-a3a8-{{Gerrit|897d5a48414f}}, use this to unset). - cookbook ran by andrew@buster * 04:23 wm-bot2: Draining 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:22 wm-bot2: Safe rebooting 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:22 wm-bot2: Draining 'cloudvirt1033.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:22 wm-bot2: Safe rebooting 'cloudvirt1033.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:18 wm-bot2: Set cloudvirt 'cloudvirt1034.eqiad.wmnet' maintenance (downtime id: 504ae503-bfad-41bd-9079-{{Gerrit|fb53a9c82a62}}, use this to unset). - cookbook ran by andrew@buster * 04:17 wm-bot2: Draining 'cloudvirt1034.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:17 wm-bot2: Safe rebooting 'cloudvirt1034.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:15 wm-bot2: Set cloudvirt 'cloudvirt1034.eqiad.wmnet' maintenance (downtime id: 8410faf9-667d-443d-bd17-{{Gerrit|bb3ca8e7f725}}, use this to unset). - cookbook ran by andrew@buster * 04:15 wm-bot2: Draining 'cloudvirt1034.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:15 wm-bot2: Safe rebooting 'cloudvirt1034.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:11 wm-bot2: Set cloudvirt 'cloudvirt1035.eqiad.wmnet' maintenance (downtime id: 40441a37-0dd7-44a8-960b-{{Gerrit|f74f98f53618}}, use this to unset). - cookbook ran by andrew@buster * 04:10 wm-bot2: Draining 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:10 wm-bot2: Safe rebooting 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:08 wm-bot2: Set cloudvirt 'cloudvirt1035.eqiad.wmnet' maintenance (downtime id: 517cff2b-a643-4714-a31f-{{Gerrit|6bbe0767656a}}, use this to unset). - cookbook ran by andrew@buster * 04:08 wm-bot2: Set cloudvirt 'cloudvirt1034.eqiad.wmnet' maintenance (downtime id: 50b63ad6-4d8b-4fcf-9c17-{{Gerrit|27a106b6ec52}}, use this to unset). - cookbook ran by andrew@buster * 04:07 wm-bot2: Draining 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:07 wm-bot2: Safe rebooting 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:07 wm-bot2: Draining 'cloudvirt1034.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:07 wm-bot2: Safe rebooting 'cloudvirt1034.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:06 wm-bot2: Set cloudvirt 'cloudvirt1033.eqiad.wmnet' maintenance (downtime id: d912bdf5-54c7-490a-bdbe-{{Gerrit|ba9a23f07f4b}}, use this to unset). - cookbook ran by andrew@buster * 04:06 wm-bot2: Set cloudvirt 'cloudvirt1035.eqiad.wmnet' maintenance (downtime id: e75f16d0-e42f-4d72-a3bd-{{Gerrit|be0cbc5f200a}}, use this to unset). - cookbook ran by andrew@buster * 04:06 wm-bot2: Set cloudvirt 'cloudvirt1034.eqiad.wmnet' maintenance (downtime id: 87f6c417-3703-4f73-9876-{{Gerrit|20f6767dc8e1}}, use this to unset). - cookbook ran by andrew@buster * 04:05 wm-bot2: Draining 'cloudvirt1033.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:05 wm-bot2: Safe rebooting 'cloudvirt1033.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:05 wm-bot2: Draining 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:05 wm-bot2: Safe rebooting 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:05 wm-bot2: Draining 'cloudvirt1034.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:05 wm-bot2: Safe rebooting 'cloudvirt1034.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:02 wm-bot2: Set cloudvirt 'cloudvirt1035.eqiad.wmnet' maintenance (downtime id: be14f869-fdb0-4c38-84ff-{{Gerrit|3d1e00d227eb}}, use this to unset). - cookbook ran by andrew@buster * 04:02 wm-bot2: Set cloudvirt 'cloudvirt1034.eqiad.wmnet' maintenance (downtime id: feb93ea7-2c64-4831-b140-{{Gerrit|dad8311dbcef}}, use this to unset). - cookbook ran by andrew@buster * 04:02 wm-bot2: Set cloudvirt 'cloudvirt1033.eqiad.wmnet' maintenance (downtime id: c65a23e3-a809-4a75-88e4-{{Gerrit|29a8b351ecdb}}, use this to unset). - cookbook ran by andrew@buster * 04:02 wm-bot2: Draining 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:01 wm-bot2: Safe rebooting 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:01 wm-bot2: Draining 'cloudvirt1034.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:01 wm-bot2: Safe rebooting 'cloudvirt1034.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:01 wm-bot2: Draining 'cloudvirt1033.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:01 wm-bot2: Safe rebooting 'cloudvirt1033.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:00 wm-bot2: Safe reboot of 'cloudvirt1030.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 04:00 wm-bot2: Unset cloudvirt 'cloudvirt1030.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 03:57 wm-bot2: Drained 'cloudvirt1030.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:56 wm-bot2: Set cloudvirt 'cloudvirt1030.eqiad.wmnet' maintenance (downtime id: 8c8f64ca-a4e8-47cf-a2ce-{{Gerrit|23059109fb6a}}, use this to unset). - cookbook ran by andrew@buster * 03:55 wm-bot2: Draining 'cloudvirt1030.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:55 wm-bot2: Safe rebooting 'cloudvirt1030.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:54 wm-bot2: Set cloudvirt 'cloudvirt1030.eqiad.wmnet' maintenance (downtime id: 9d0457e8-a861-46ce-ac35-{{Gerrit|53975a46018e}}, use this to unset). - cookbook ran by andrew@buster * 03:53 wm-bot2: Draining 'cloudvirt1030.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:53 wm-bot2: Safe rebooting 'cloudvirt1030.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:44 wm-bot2: Safe reboot of 'cloudvirt1031.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 03:44 wm-bot2: Unset cloudvirt 'cloudvirt1031.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 03:42 wm-bot2: Safe reboot of 'cloudvirt1032.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 03:42 wm-bot2: Unset cloudvirt 'cloudvirt1032.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 03:40 wm-bot2: Drained 'cloudvirt1031.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:38 wm-bot2: Drained 'cloudvirt1032.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:30 wm-bot2: Set cloudvirt 'cloudvirt1032.eqiad.wmnet' maintenance (downtime id: 851d3b34-68f7-4c57-920b-{{Gerrit|2a64a9ea573d}}, use this to unset). - cookbook ran by andrew@buster * 03:29 wm-bot2: Draining 'cloudvirt1032.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:29 wm-bot2: Safe rebooting 'cloudvirt1032.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:17 wm-bot2: Set cloudvirt 'cloudvirt1032.eqiad.wmnet' maintenance (downtime id: 1060ab82-d53a-4d48-8e15-{{Gerrit|3198cabcfc2f}}, use this to unset). - cookbook ran by andrew@buster * 03:17 wm-bot2: Draining 'cloudvirt1032.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:17 wm-bot2: Safe rebooting 'cloudvirt1032.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:14 wm-bot2: Safe reboot of 'cloudvirt1029.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 03:14 wm-bot2: Unset cloudvirt 'cloudvirt1029.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 03:14 wm-bot2: Set cloudvirt 'cloudvirt1031.eqiad.wmnet' maintenance (downtime id: dde1dc09-44b5-42af-ac49-{{Gerrit|c58dbbae7cb4}}, use this to unset). - cookbook ran by andrew@buster * 03:14 wm-bot2: Draining 'cloudvirt1031.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:14 wm-bot2: Safe rebooting 'cloudvirt1031.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:13 wm-bot2: Drained 'cloudvirt1029.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:13 wm-bot2: Set cloudvirt 'cloudvirt1029.eqiad.wmnet' maintenance (downtime id: b59fa99b-8713-4e63-8c56-{{Gerrit|08bd71fdfbe3}}, use this to unset). - cookbook ran by andrew@buster * 03:13 wm-bot2: Draining 'cloudvirt1029.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:13 wm-bot2: Safe rebooting 'cloudvirt1029.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:07 wm-bot2: Set cloudvirt 'cloudvirt1031.eqiad.wmnet' maintenance (downtime id: 3547c843-6b15-4151-b452-{{Gerrit|6f618a886b7b}}, use this to unset). - cookbook ran by andrew@buster * 03:07 wm-bot2: Set cloudvirt 'cloudvirt1030.eqiad.wmnet' maintenance (downtime id: 6f5808d0-afb2-4222-ae8e-{{Gerrit|4d953d49cd63}}, use this to unset). - cookbook ran by andrew@buster * 03:06 wm-bot2: Draining 'cloudvirt1031.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:06 wm-bot2: Safe rebooting 'cloudvirt1031.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:06 wm-bot2: Draining 'cloudvirt1030.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:06 wm-bot2: Safe rebooting 'cloudvirt1030.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:06 wm-bot2: Set cloudvirt 'cloudvirt1029.eqiad.wmnet' maintenance (downtime id: da738ba3-883f-4d76-bb33-{{Gerrit|7eb551de5fc2}}, use this to unset). - cookbook ran by andrew@buster * 03:05 wm-bot2: Draining 'cloudvirt1029.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:05 wm-bot2: Safe rebooting 'cloudvirt1029.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:04 wm-bot2: Safe reboot of 'cloudvirt1026.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 03:04 wm-bot2: Unset cloudvirt 'cloudvirt1026.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 03:03 wm-bot2: Set cloudvirt 'cloudvirt1029.eqiad.wmnet' maintenance (downtime id: f90ca4c1-78af-46bd-9262-{{Gerrit|0fa4edc2898e}}, use this to unset). - cookbook ran by andrew@buster * 03:02 wm-bot2: Draining 'cloudvirt1029.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:02 wm-bot2: Safe rebooting 'cloudvirt1029.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:02 wm-bot2: Safe reboot of 'cloudvirt1027.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 03:02 wm-bot2: Unset cloudvirt 'cloudvirt1027.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 03:00 wm-bot2: Drained 'cloudvirt1026.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:00 wm-bot2: Set cloudvirt 'cloudvirt1026.eqiad.wmnet' maintenance (downtime id: 62383663-9ae6-4a15-a2a6-{{Gerrit|62be06adbc96}}, use this to unset). - cookbook ran by andrew@buster * 02:59 wm-bot2: Draining 'cloudvirt1026.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:59 wm-bot2: Safe rebooting 'cloudvirt1026.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:59 wm-bot2: Drained 'cloudvirt1027.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:57 wm-bot2: Set cloudvirt 'cloudvirt1026.eqiad.wmnet' maintenance (downtime id: 6e0da344-49ad-451f-8b8d-{{Gerrit|ea8f853a6f89}}, use this to unset). - cookbook ran by andrew@buster * 02:57 wm-bot2: Draining 'cloudvirt1026.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:57 wm-bot2: Safe rebooting 'cloudvirt1026.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:52 wm-bot2: Set cloudvirt 'cloudvirt1029.eqiad.wmnet' maintenance (downtime id: c9ec5b7c-aa9f-4afd-a417-{{Gerrit|54c52c07ca47}}, use this to unset). - cookbook ran by andrew@buster * 02:51 wm-bot2: Draining 'cloudvirt1029.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:51 wm-bot2: Safe rebooting 'cloudvirt1029.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:48 wm-bot2: Set cloudvirt 'cloudvirt1029.eqiad.wmnet' maintenance (downtime id: 5509c648-58f6-49fd-90e8-{{Gerrit|ac9aa66e8b04}}, use this to unset). - cookbook ran by andrew@buster * 02:48 wm-bot2: Draining 'cloudvirt1029.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:48 wm-bot2: Safe rebooting 'cloudvirt1029.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:46 wm-bot2: Safe reboot of 'cloudvirt1024.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 02:46 wm-bot2: Unset cloudvirt 'cloudvirt1024.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 02:46 wm-bot2: Set cloudvirt 'cloudvirt1026.eqiad.wmnet' maintenance (downtime id: af4ba9ad-e765-4087-86aa-{{Gerrit|9111c0821175}}, use this to unset). - cookbook ran by andrew@buster * 02:46 wm-bot2: Draining 'cloudvirt1026.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:45 wm-bot2: Safe rebooting 'cloudvirt1026.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:45 wm-bot2: Set cloudvirt 'cloudvirt1027.eqiad.wmnet' maintenance (downtime id: 6da98123-8838-4d2e-8659-{{Gerrit|d769bfa292fe}}, use this to unset). - cookbook ran by andrew@buster * 02:44 wm-bot2: Draining 'cloudvirt1027.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:44 wm-bot2: Safe rebooting 'cloudvirt1027.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:43 wm-bot2: Safe reboot of 'cloudvirt1025.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 02:43 wm-bot2: Unset cloudvirt 'cloudvirt1025.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 02:42 wm-bot2: Drained 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:42 wm-bot2: Set cloudvirt 'cloudvirt1024.eqiad.wmnet' maintenance (downtime id: 91860079-8fff-40db-9de7-{{Gerrit|17741d528ad6}}, use this to unset). - cookbook ran by andrew@buster * 02:41 wm-bot2: Draining 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:41 wm-bot2: Safe rebooting 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:39 wm-bot2: Set cloudvirt 'cloudvirt1026.eqiad.wmnet' maintenance (downtime id: a6a0f414-d305-4a09-87d2-{{Gerrit|066eec83642b}}, use this to unset). - cookbook ran by andrew@buster * 02:39 wm-bot2: Drained 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:38 wm-bot2: Draining 'cloudvirt1026.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:38 wm-bot2: Safe rebooting 'cloudvirt1026.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:38 wm-bot2: Set cloudvirt 'cloudvirt1025.eqiad.wmnet' maintenance (downtime id: 6000898c-4e5e-49c8-8841-{{Gerrit|3262d6da0366}}, use this to unset). - cookbook ran by andrew@buster * 02:37 wm-bot2: Draining 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:37 wm-bot2: Safe rebooting 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:34 wm-bot2: Set cloudvirt 'cloudvirt1025.eqiad.wmnet' maintenance (downtime id: 24a8b50e-5c04-40e6-a732-{{Gerrit|f0232b83e240}}, use this to unset). - cookbook ran by andrew@buster * 02:34 wm-bot2: Set cloudvirt 'cloudvirt1024.eqiad.wmnet' maintenance (downtime id: 8d259aca-1803-4346-b5a4-{{Gerrit|a447673b3537}}, use this to unset). - cookbook ran by andrew@buster * 02:34 wm-bot2: Draining 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:34 wm-bot2: Safe rebooting 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:34 wm-bot2: Draining 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:34 wm-bot2: Safe rebooting 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:33 wm-bot2: Set cloudvirt 'cloudvirt1024.eqiad.wmnet' maintenance (downtime id: 704aca7f-0069-4a4a-b8de-{{Gerrit|1aba7b9ca7ef}}, use this to unset). - cookbook ran by andrew@buster * 02:32 wm-bot2: Set cloudvirt 'cloudvirt1025.eqiad.wmnet' maintenance (downtime id: 04c0fba3-e2cb-40b4-abfd-{{Gerrit|d24165c9bd55}}, use this to unset). - cookbook ran by andrew@buster * 02:32 wm-bot2: Draining 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:32 wm-bot2: Safe rebooting 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:32 wm-bot2: Draining 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:32 wm-bot2: Safe rebooting 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:28 wm-bot2: Set cloudvirt 'cloudvirt1025.eqiad.wmnet' maintenance (downtime id: 25e5e0c8-5c27-460e-9b44-{{Gerrit|1cbd1c2fb551}}, use this to unset). - cookbook ran by andrew@buster * 02:28 wm-bot2: Draining 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:28 wm-bot2: Safe rebooting 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:25 wm-bot2: Set cloudvirt 'cloudvirt1025.eqiad.wmnet' maintenance (downtime id: 5e62968a-0bb6-43b1-ae5e-{{Gerrit|e352442ef976}}, use this to unset). - cookbook ran by andrew@buster * 02:24 wm-bot2: Draining 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:24 wm-bot2: Safe rebooting 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:21 wm-bot2: Set cloudvirt 'cloudvirt1025.eqiad.wmnet' maintenance (downtime id: 5b1a8e8f-8699-433d-b764-{{Gerrit|197862dc5cf1}}, use this to unset). - cookbook ran by andrew@buster * 02:20 wm-bot2: Draining 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:20 wm-bot2: Safe rebooting 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:15 wm-bot2: Draining 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:15 wm-bot2: Safe rebooting 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:13 wm-bot2: Set cloudvirt 'cloudvirt1025.eqiad.wmnet' maintenance (downtime id: 299c466e-602b-4a8e-bb37-{{Gerrit|156553dbe426}}, use this to unset). - cookbook ran by andrew@buster * 02:13 wm-bot2: Draining 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:13 wm-bot2: Safe rebooting 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:12 wm-bot2: Draining 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:12 wm-bot2: Safe rebooting 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:11 wm-bot2: Draining 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:11 wm-bot2: Safe rebooting 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:03 wm-bot2: Set cloudvirt 'cloudvirt1024.eqiad.wmnet' maintenance (downtime id: 6a64d619-c484-411f-ad9d-{{Gerrit|89d43b063737}}, use this to unset). - cookbook ran by andrew@buster * 02:02 wm-bot2: Draining 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:02 wm-bot2: Safe rebooting 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 00:46 wm-bot2: Set cloudvirt 'cloudvirt1024.eqiad.wmnet' maintenance (downtime id: 301309a2-a8b5-4698-98d6-{{Gerrit|bb4aa0a75e45}}, use this to unset). - cookbook ran by andrew@buster * 00:46 wm-bot2: Set cloudvirt 'cloudvirt1025.eqiad.wmnet' maintenance (downtime id: b2f8c6a2-7224-4fb1-8e12-{{Gerrit|dff6dda02380}}, use this to unset). - cookbook ran by andrew@buster * 00:46 wm-bot2: Draining 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 00:46 wm-bot2: Safe rebooting 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 00:46 wm-bot2: Draining 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 00:46 wm-bot2: Safe rebooting 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 00:40 wm-bot2: Draining 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 00:40 wm-bot2: Safe rebooting 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 00:39 wm-bot2: Draining 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 00:39 wm-bot2: Safe rebooting 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 00:38 wm-bot2: Set cloudvirt 'cloudvirt1024.eqiad.wmnet' maintenance (downtime id: c0f62e0e-7865-4e77-a238-{{Gerrit|d707ceed53a8}}, use this to unset). - cookbook ran by andrew@buster * 00:38 wm-bot2: Draining 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 00:38 wm-bot2: Safe rebooting 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 00:36 wm-bot2: Set cloudvirt 'cloudvirt1025.eqiad.wmnet' maintenance (downtime id: 96a368d5-232f-4187-b0bb-{{Gerrit|5923a50fff71}}, use this to unset). - cookbook ran by andrew@buster * 00:36 wm-bot2: Draining 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 00:36 wm-bot2: Safe rebooting 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 00:35 wm-bot2: Set cloudvirt 'cloudvirt1024.eqiad.wmnet' maintenance (downtime id: 31b6bfcb-81cb-4d89-b977-{{Gerrit|190899ea2e64}}, use this to unset). - cookbook ran by andrew@buster * 00:35 wm-bot2: Draining 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 00:35 wm-bot2: Safe rebooting 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 00:28 wm-bot2: Set cloudvirt 'cloudvirt1024.eqiad.wmnet' maintenance (downtime id: 50f763d4-36ea-4241-8e71-{{Gerrit|cc6aa20ccc9e}}, use this to unset). - cookbook ran by andrew@buster * 00:28 wm-bot2: Set cloudvirt 'cloudvirt1025.eqiad.wmnet' maintenance (downtime id: cd3d17af-ee15-4a9f-8750-{{Gerrit|66b97a46ea12}}, use this to unset). - cookbook ran by andrew@buster * 00:27 wm-bot2: Draining 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 00:27 wm-bot2: Safe rebooting 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 00:27 wm-bot2: Draining 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 00:27 wm-bot2: Safe rebooting 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 00:22 wm-bot2: Draining 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 00:22 wm-bot2: Safe rebooting 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 00:22 wm-bot2: Set cloudvirt 'cloudvirt1025.eqiad.wmnet' maintenance (downtime id: 9b075166-d424-42c0-8e06-{{Gerrit|f74a47abb798}}, use this to unset). - cookbook ran by andrew@buster * 00:21 wm-bot2: Draining 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 00:21 wm-bot2: Safe rebooting 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 00:21 wm-bot2: Draining 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 00:21 wm-bot2: Safe rebooting 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster === 2022-07-18 === * 23:11 wm-bot2: Set cloudvirt 'cloudvirt1024.eqiad.wmnet' maintenance (downtime id: e0aab64f-d911-4d7e-9a97-{{Gerrit|c59d9868ceea}}, use this to unset). - cookbook ran by andrew@buster * 23:11 wm-bot2: Set cloudvirt 'cloudvirt1025.eqiad.wmnet' maintenance (downtime id: 447d9c7a-a00e-41db-93ed-{{Gerrit|6d2bfc66e5df}}, use this to unset). - cookbook ran by andrew@buster * 23:11 wm-bot2: Draining 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 23:11 wm-bot2: Safe rebooting 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 23:11 wm-bot2: Draining 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 23:11 wm-bot2: Safe rebooting 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 23:01 wm-bot2: Set cloudvirt 'cloudvirt1025.eqiad.wmnet' maintenance (downtime id: 46c2ddf2-50c3-4824-ba72-{{Gerrit|d7fb3fa4c5e0}}, use this to unset). - cookbook ran by andrew@buster * 23:01 wm-bot2: Set cloudvirt 'cloudvirt1024.eqiad.wmnet' maintenance (downtime id: 6facdf4f-7bb8-42cd-8ac8-{{Gerrit|b4412fc4cf70}}, use this to unset). - cookbook ran by andrew@buster * 23:00 wm-bot2: Draining 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 23:00 wm-bot2: Safe rebooting 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 23:00 wm-bot2: Draining 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 23:00 wm-bot2: Safe rebooting 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 22:52 wm-bot2: Set cloudvirt 'cloudvirt1025.eqiad.wmnet' maintenance (downtime id: 23190573-fc0d-4872-8343-{{Gerrit|e35b84132028}}, use this to unset). - cookbook ran by andrew@buster * 22:51 wm-bot2: Draining 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 22:51 wm-bot2: Safe rebooting 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 22:51 wm-bot2: Set cloudvirt 'cloudvirt1024.eqiad.wmnet' maintenance (downtime id: 990313c4-5fa6-4de7-a8bf-{{Gerrit|b361f0c79e90}}, use this to unset). - cookbook ran by andrew@buster * 22:50 wm-bot2: Draining 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 22:50 wm-bot2: Safe rebooting 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 22:48 wm-bot2: Safe reboot of 'cloudvirt1022.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 22:48 wm-bot2: Unset cloudvirt 'cloudvirt1022.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 22:42 wm-bot2: Set cloudvirt 'cloudvirt1023.eqiad.wmnet' maintenance (downtime id: 59367d39-bd2f-47fd-becd-{{Gerrit|255bf823ff12}}, use this to unset). - cookbook ran by andrew@buster * 22:42 wm-bot2: Draining 'cloudvirt1023.eqiad.wmnet'. - cookbook ran by andrew@buster * 22:42 wm-bot2: Safe rebooting 'cloudvirt1023.eqiad.wmnet'. - cookbook ran by andrew@buster * 22:40 wm-bot2: Set cloudvirt 'cloudvirt1024.eqiad.wmnet' maintenance (downtime id: 38602815-d187-4455-9676-{{Gerrit|59607505bba7}}, use this to unset). - cookbook ran by andrew@buster * 22:39 wm-bot2: Set cloudvirt 'cloudvirt1023.eqiad.wmnet' maintenance (downtime id: 00f21e1b-d014-4bc5-995f-{{Gerrit|e25201fc77c4}}, use this to unset). - cookbook ran by andrew@buster * 22:39 wm-bot2: Draining 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 22:39 wm-bot2: Safe rebooting 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 22:39 wm-bot2: Draining 'cloudvirt1023.eqiad.wmnet'. - cookbook ran by andrew@buster * 22:39 wm-bot2: Safe rebooting 'cloudvirt1023.eqiad.wmnet'. - cookbook ran by andrew@buster * 22:31 wm-bot2: Set cloudvirt 'cloudvirt1023.eqiad.wmnet' maintenance (downtime id: d05d07b2-8c60-46de-a89e-{{Gerrit|76987901901b}}, use this to unset). - cookbook ran by andrew@buster * 22:31 wm-bot2: Set cloudvirt 'cloudvirt1024.eqiad.wmnet' maintenance (downtime id: 47f6b9e6-ceab-4fb6-8f76-{{Gerrit|0de4341d0841}}, use this to unset). - cookbook ran by andrew@buster * 22:31 wm-bot2: Draining 'cloudvirt1023.eqiad.wmnet'. - cookbook ran by andrew@buster * 22:31 wm-bot2: Safe rebooting 'cloudvirt1023.eqiad.wmnet'. - cookbook ran by andrew@buster * 22:30 wm-bot2: Draining 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 22:30 wm-bot2: Safe rebooting 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 22:30 wm-bot2: Set cloudvirt 'cloudvirt1023.eqiad.wmnet' maintenance (downtime id: 599042a5-da25-4985-b502-{{Gerrit|328fe40a72d5}}, use this to unset). - cookbook ran by andrew@buster * 22:29 wm-bot2: Draining 'cloudvirt1023.eqiad.wmnet'. - cookbook ran by andrew@buster * 22:29 wm-bot2: Safe rebooting 'cloudvirt1023.eqiad.wmnet'. - cookbook ran by andrew@buster * 22:29 wm-bot2: Drained 'cloudvirt1022.eqiad.wmnet'. - cookbook ran by andrew@buster * 22:28 wm-bot2: Set cloudvirt 'cloudvirt1023.eqiad.wmnet' maintenance (downtime id: a554d107-c14d-41f8-b180-{{Gerrit|093e2cf73e72}}, use this to unset). - cookbook ran by andrew@buster * 22:28 wm-bot2: Set cloudvirt 'cloudvirt1022.eqiad.wmnet' maintenance (downtime id: 9fef70c7-6f7d-4b33-a9bc-{{Gerrit|8c39630dc182}}, use this to unset). - cookbook ran by andrew@buster * 22:26 wm-bot2: Drained 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 22:02 wm-bot2: Set cloudvirt 'cloudvirt1023.eqiad.wmnet' maintenance (downtime id: 3744c0fa-084d-4198-98ca-{{Gerrit|66c7a5210f41}}, use this to unset). - cookbook ran by andrew@buster * 22:02 wm-bot2: Set cloudvirt 'cloudvirt1022.eqiad.wmnet' maintenance (downtime id: 01caf1e1-df64-474a-94d5-{{Gerrit|f37751e23d62}}, use this to unset). - cookbook ran by andrew@buster * 22:01 wm-bot2: Draining 'cloudvirt1023.eqiad.wmnet'. - cookbook ran by andrew@buster * 22:01 wm-bot2: Safe rebooting 'cloudvirt1023.eqiad.wmnet'. - cookbook ran by andrew@buster * 22:01 wm-bot2: Draining 'cloudvirt1022.eqiad.wmnet'. - cookbook ran by andrew@buster * 22:01 wm-bot2: Safe rebooting 'cloudvirt1022.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:59 wm-bot2: Set cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance (downtime id: 3018c2fc-9cd5-45c8-a160-{{Gerrit|d5ad5728c8d5}}, use this to unset). - cookbook ran by andrew@buster * 21:59 wm-bot2: Draining 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:59 wm-bot2: Safe rebooting 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:02 wm-bot2: Set cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance (downtime id: d760f916-32ea-4194-8974-{{Gerrit|f36064a9866c}}, use this to unset). - cookbook ran by andrew@buster * 21:02 wm-bot2: Draining 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:02 wm-bot2: Safe rebooting 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:00 wm-bot2: Set cloudvirt 'cloudvirt1022.eqiad.wmnet' maintenance (downtime id: cb2a38f3-e9ca-4bf6-989c-{{Gerrit|a27be2b2d88c}}, use this to unset). - cookbook ran by andrew@buster * 20:59 wm-bot2: Draining 'cloudvirt1022.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:59 wm-bot2: Safe rebooting 'cloudvirt1022.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:58 wm-bot2: Set cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance (downtime id: 93fa9477-54b8-46e7-8508-{{Gerrit|a2d492528d1f}}, use this to unset). - cookbook ran by andrew@buster * 20:57 wm-bot2: Draining 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:57 wm-bot2: Safe rebooting 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:52 wm-bot2: Set cloudvirt 'cloudvirt1023.eqiad.wmnet' maintenance (downtime id: f6cf5b29-ce8b-426e-af1c-{{Gerrit|cab74e0c4a3a}}, use this to unset). - cookbook ran by andrew@buster * 20:52 wm-bot2: Set cloudvirt 'cloudvirt1022.eqiad.wmnet' maintenance (downtime id: fe26e612-d76e-4120-b60a-{{Gerrit|4f300add224f}}, use this to unset). - cookbook ran by andrew@buster * 20:52 wm-bot2: Set cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance (downtime id: 2abd94be-b6d8-4042-a9ce-{{Gerrit|998e41d932f4}}, use this to unset). - cookbook ran by andrew@buster * 20:51 wm-bot2: Draining 'cloudvirt1023.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:51 wm-bot2: Safe rebooting 'cloudvirt1023.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:51 wm-bot2: Draining 'cloudvirt1022.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:51 wm-bot2: Safe rebooting 'cloudvirt1022.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:51 wm-bot2: Draining 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:51 wm-bot2: Safe rebooting 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:34 wm-bot2: Draining 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:34 wm-bot2: Safe rebooting 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:28 wm-bot2: Set cloudvirt 'cloudvirt1022.eqiad.wmnet' maintenance (downtime id: 10973c71-30d8-44bf-a310-{{Gerrit|d0c1af76d399}}, use this to unset). - cookbook ran by andrew@buster * 20:28 wm-bot2: Set cloudvirt 'cloudvirt1023.eqiad.wmnet' maintenance (downtime id: f8f1b214-ec76-4157-a8e6-{{Gerrit|6374e5d14608}}, use this to unset). - cookbook ran by andrew@buster * 20:27 wm-bot2: Draining 'cloudvirt1022.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:27 wm-bot2: Safe rebooting 'cloudvirt1022.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:27 wm-bot2: Draining 'cloudvirt1023.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:27 wm-bot2: Safe rebooting 'cloudvirt1023.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:11 wm-bot2: Set cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance (downtime id: 95dae613-2b04-472e-a8f9-{{Gerrit|e71c3fbb8d0e}}, use this to unset). - cookbook ran by andrew@buster * 20:10 wm-bot2: Draining 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:10 wm-bot2: Safe rebooting 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:03 wm-bot2: Set cloudvirt 'cloudvirt1022.eqiad.wmnet' maintenance (downtime id: c804f4db-ea63-44e9-aff8-{{Gerrit|fea54c238345}}, use this to unset). - cookbook ran by andrew@buster * 20:03 wm-bot2: Set cloudvirt 'cloudvirt1023.eqiad.wmnet' maintenance (downtime id: 597a901d-4624-43de-b986-{{Gerrit|c4c6c2fcf80c}}, use this to unset). - cookbook ran by andrew@buster * 20:03 wm-bot2: Draining 'cloudvirt1023.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:03 wm-bot2: Safe rebooting 'cloudvirt1023.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:02 wm-bot2: Set cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance (downtime id: fdceed0d-6fa0-4806-8fb6-{{Gerrit|3b321a5a1623}}, use this to unset). - cookbook ran by andrew@buster * 20:02 wm-bot2: Draining 'cloudvirt1022.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:02 wm-bot2: Safe rebooting 'cloudvirt1022.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:02 wm-bot2: Draining 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:02 wm-bot2: Safe rebooting 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:01 wm-bot2: Safe reboot of 'cloudvirt1017.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 20:01 wm-bot2: Unset cloudvirt 'cloudvirt1017.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 19:57 wm-bot2: Drained 'cloudvirt1017.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:37 wm-bot2: Set cloudvirt 'cloudvirt1017.eqiad.wmnet' maintenance (downtime id: b0904345-9aa3-4202-b992-{{Gerrit|78644141517c}}, use this to unset). - cookbook ran by andrew@buster * 19:36 wm-bot2: Draining 'cloudvirt1017.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:36 wm-bot2: Safe rebooting 'cloudvirt1017.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:31 wm-bot2: Draining 'cloudvirt1017.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:31 wm-bot2: Safe rebooting 'cloudvirt1017.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:30 wm-bot2: Draining 'cloudvirt1017.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:30 wm-bot2: Safe rebooting 'cloudvirt1017.eqiad.wmnet'. - cookbook ran by andrew@buster === 2022-07-13 === * 14:48 bd808: Added Tavvi as member of #acl*wmcs-team === 2022-07-12 === * 12:04 wm-bot2: Ceph cluster at <nowiki>{</nowiki>self.deployment<nowiki>}</nowiki> set out of maintenance. - cookbook ran by dcaro@vulcanus * 12:03 wm-bot2: Set the ceph cluster for codfw1dev in maintenance, alert silence ids: db32805f-a033-4e0e-8fc7-{{Gerrit|ed0c2e9f8be1}},f4e698f0-4b51-4b07-9ba8-{{Gerrit|0b296ca1c4fb}},39ad5325-44ed-44d1-bba5-{{Gerrit|91a85c7a8401}},a584a3c6-9dae-41e9-ac82-{{Gerrit|a91cd35af8fb}} - cookbook ran by dcaro@vulcanus * 08:37 wm-bot2: Finished rebooting the cloudnet nodes ['cloudnet2005-dev', 'cloudnet2006-dev'] - cookbook ran by dcaro@vulcanus * 08:37 wm-bot2: Rebooted cloudnet host cloudnet2006-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 08:31 wm-bot2: Rebooting cloudnet host cloudnet2006-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 08:31 wm-bot2: Rebooted cloudnet host cloudnet2005-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 08:25 wm-bot2: Rebooting cloudnet host cloudnet2005-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 08:25 wm-bot2: Rebooting all the cloudnet nodes cloudnet2005-dev,cloudnet2006-dev - cookbook ran by dcaro@vulcanus === 2022-07-08 === * 15:57 wm-bot2: Finished rebooting node cloudcephosd1021.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 15:52 wm-bot2: Rebooting node cloudcephosd1021.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 15:52 wm-bot2: Rebooting node cloudcephosd1021.eqiad.wmnet - cookbook ran by dcaro@vulcanus === 2022-07-07 === * 07:17 wm-bot2: Finished rebooting node cloudcephosd1015.eqiad.wmnet ([[phab:T312509|T312509]]) - cookbook ran by dcaro@vulcanus * 07:12 wm-bot2: Rebooting node cloudcephosd1015.eqiad.wmnet ([[phab:T312509|T312509]]) - cookbook ran by dcaro@vulcanus === 2022-07-06 === * 17:50 wm-bot2: Set the ceph cluster for eqiad1 in maintenance, alert silence ids: ['8a5b9eee-48c0-474d-8277-faeb05a2ea61', '65aad0fc-d887-47a3-b20c-d1ed461a2411', '86b078ae-3a27-4063-8c7c-198a2fe0c172'] - cookbook ran by dcaro@vulcanus === 2022-07-04 === * 13:27 wm-bot2: Rebooting cloudgw host cloudgw1002.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 13:27 wm-bot2: Rebooted cloudgw host cloudgw1001.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 13:23 wm-bot2: Rebooting cloudgw host cloudgw1001.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 13:23 wm-bot2: Rebooting all the cloudgw nodes from the eqiad1 deployment: cloudgw1001.eqiad.wmnet,cloudgw1002.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 13:13 wm-bot2: Finished rebooting the cloudgw nodes ['cloudgw2001-dev.codfw.wmnet', 'cloudgw2002-dev.codfw.wmnet', 'cloudgw2003-dev.codfw.wmnet'] - cookbook ran by dcaro@vulcanus * 13:12 wm-bot2: Rebooted cloudgw host cloudgw2003-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 13:09 wm-bot2: Rebooting cloudgw host cloudgw2003-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 13:08 wm-bot2: Rebooted cloudgw host cloudgw2002-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 13:05 wm-bot2: Rebooting cloudgw host cloudgw2002-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 13:05 wm-bot2: Rebooted cloudgw host cloudgw2001-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 13:01 wm-bot2: Rebooting cloudgw host cloudgw2001-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 13:01 wm-bot2: Rebooting all the cloudgw nodes cloudgw2001-dev.codfw.wmnet,cloudgw2002-dev.codfw.wmnet,cloudgw2003-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 12:56 wm-bot2: Rebooting all the cloudgw nodes cloudgw2001-dev.codfw.wmnet,cloudgw2002-dev.codfw.wmnet,cloudgw2003-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 12:52 wm-bot2: Rebooting all the cloudgw nodes cloudgw2001-dev.codfw.wmnet,cloudgw2002-dev.codfw.wmnet,cloudgw2003-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 10:56 wm-bot2: Finished rebooting the cloudgw nodes ['cloudgw2001-dev.codfw.wmnet', 'cloudgw2002-dev.codfw.wmnet', 'cloudgw2003-dev.codfw.wmnet'] - cookbook ran by dcaro@vulcanus * 10:56 wm-bot2: Rebooted cloudgw host cloudgw2003-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 10:51 wm-bot2: Rebooting cloudgw host cloudgw2003-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 10:51 wm-bot2: Rebooted cloudgw host cloudgw2002-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 10:47 wm-bot2: Rebooting cloudgw host cloudgw2002-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 10:47 wm-bot2: Rebooted cloudgw host cloudgw2001-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 10:42 wm-bot2: Rebooting cloudgw host cloudgw2001-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 10:42 wm-bot2: Rebooting all the cloudgw nodes cloudgw2001-dev.codfw.wmnet,cloudgw2002-dev.codfw.wmnet,cloudgw2003-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 10:41 wm-bot2: Rebooting all the cloudgw nodes cloudgw2001-dev.codfw.wmnet,cloudgw2002-dev.codfw.wmnet,cloudgw2003-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 10:40 wm-bot2: Rebooting all the cloudgw nodes cloudgw2001-dev.codfw.wmnet,cloudgw2002-dev.codfw.wmnet,cloudgw2003-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 09:07 wm-bot2: Rebooting cloudnet host cloudnet1004.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 09:06 wm-bot2: Rebooted cloudnet host cloudnet1003.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 09:00 wm-bot2: Rebooting cloudnet host cloudnet1003.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 09:00 wm-bot2: Rebooting all the cloudnet nodes cloudnet1003,cloudnet1004 - cookbook ran by dcaro@vulcanus * 08:59 wm-bot2: Finished rebooting the cloudnet nodes ['cloudnet2006-dev', 'cloudnet2005-dev'] - cookbook ran by dcaro@vulcanus * 08:59 wm-bot2: Rebooted cloudnet host cloudnet2005-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 08:53 wm-bot2: Rebooting cloudnet host cloudnet2005-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 08:53 wm-bot2: Rebooted cloudnet host cloudnet2006-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 08:47 wm-bot2: Rebooting cloudnet host cloudnet2006-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 08:47 wm-bot2: Rebooting all the cloudnet nodes cloudnet2006-dev,cloudnet2005-dev - cookbook ran by dcaro@vulcanus * 08:32 wm-bot2: Finished rebooting the cloudnet nodes ['cloudnet2006-dev', 'cloudnet2005-dev'] - cookbook ran by dcaro@vulcanus * 08:32 wm-bot2: Rebooted cloudnet host cloudnet2005-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 08:25 wm-bot2: Rebooting cloudnet host cloudnet2005-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 08:25 wm-bot2: Rebooted cloudnet host cloudnet2006-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 08:19 wm-bot2: Rebooting cloudnet host cloudnet2006-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 08:19 wm-bot2: Rebooting all the cloudnet nodes cloudnet2006-dev,cloudnet2005-dev - cookbook ran by dcaro@vulcanus * 07:55 wm-bot2: Rebooting all the cloudnet nodes cloudnet2006-dev,cloudnet2005-dev - cookbook ran by dcaro@vulcanus * 07:51 wm-bot2: Rebooting all the cloudnet nodes cloudnet2006-dev,cloudnet2005-dev - cookbook ran by dcaro@vulcanus * 07:44 wm-bot2: Finished rebooting the cloudnet nodes ['cloudnet2006-dev', 'cloudnet2005-dev'] - cookbook ran by dcaro@vulcanus * 07:44 wm-bot2: Rebooted cloudnet host cloudnet2005-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 07:40 wm-bot2: Rebooting cloudnet host cloudnet2005-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 07:40 wm-bot2: Rebooted cloudnet host cloudnet2006-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 07:36 wm-bot2: Rebooting cloudnet host cloudnet2006-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 07:36 wm-bot2: Rebooting all the cloudnet nodes cloudnet2006-dev,cloudnet2005-dev - cookbook ran by dcaro@vulcanus * 07:09 wm-bot2: Rebooted cloudnet host cloudnet2006-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 07:05 wm-bot2: Rebooting cloudnet host cloudnet2006-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 07:05 wm-bot2: Rebooting all the cloudnet nodes cloudnet2006-dev,cloudnet2005-dev - cookbook ran by dcaro@vulcanus * 07:04 wm-bot2: Rebooting all the cloudnet nodes cloudnet2006-dev,cloudnet2005-dev - cookbook ran by dcaro@vulcanus === 2022-07-03 === * 21:27 andrewbogott: rebuilding rabbit cluster in codfw1dev to get rid of some queues so unresponsive that they can't otherwise be deleted === 2022-07-02 === * 11:05 wm-bot2: Rebooting all the cloudnet nodes cloudnet2006-dev,cloudnet2005-dev - cookbook ran by dcaro@vulcanus === 2022-07-01 === * 15:35 wm-bot2: Finished rebooting the cloudnet nodes ['cloudnet2006-dev', 'cloudnet2005-dev'] - cookbook ran by dcaro@vulcanus * 15:35 wm-bot2: Rebooted cloudnet host cloudnet2005-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 15:30 wm-bot2: Rebooting cloudnet host cloudnet2005-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 15:29 wm-bot2: Rebooted cloudnet host cloudnet2006-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 15:25 wm-bot2: Rebooting cloudnet host cloudnet2006-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 15:25 wm-bot2: Rebooting all the cloudnet nodes cloudnet2006-dev,cloudnet2005-dev - cookbook ran by dcaro@vulcanus * 15:13 wm-bot2: Finished rebooting the cloudnet nodes ['cloudnet2006-dev', 'cloudnet2005-dev'] - cookbook ran by dcaro@vulcanus * 15:13 wm-bot2: Rebooted cloudnet host cloudnet2005-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 15:08 wm-bot2: Rebooting cloudnet host cloudnet2005-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 15:07 wm-bot2: Rebooted cloudnet host cloudnet2006-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 15:03 wm-bot2: Rebooting cloudnet host cloudnet2006-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 15:03 wm-bot2: Rebooting all the cloudnet nodes cloudnet2006-dev,cloudnet2005-dev - cookbook ran by dcaro@vulcanus * 14:42 wm-bot2: Finished rebooting the cloudnet nodes ['cloudnet2006-dev', 'cloudnet2005-dev'] - cookbook ran by dcaro@vulcanus * 14:42 wm-bot2: Rebooted cloudnet host cloudnet2005-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 14:37 wm-bot2: Rebooting cloudnet host cloudnet2005-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 14:36 wm-bot2: Rebooted cloudnet host cloudnet2006-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 14:32 wm-bot2: Rebooting cloudnet host cloudnet2006-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 14:31 wm-bot2: Rebooting all the cloudnet nodes cloudnet2006-dev,cloudnet2005-dev - cookbook ran by dcaro@vulcanus * 14:23 wm-bot2: Rebooted cloudnet host cloudnet2006-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 14:19 wm-bot2: Rebooting cloudnet host cloudnet2006-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 14:19 wm-bot2: Rebooting all the cloudnet nodes cloudnet2006-dev,cloudnet2005-dev - cookbook ran by dcaro@vulcanus * 14:19 wm-bot2: Finished rebooting the cloudnet nodes ['cloudnet2006-dev', 'cloudnet2005-dev'] - cookbook ran by dcaro@vulcanus * 14:09 wm-bot2: Rebooted cloudnet host cloudnet2005-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 14:06 wm-bot2: Rebooting cloudnet host cloudnet2005-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 14:02 wm-bot2: Rebooted cloudnet host cloudnet2006-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 13:58 wm-bot2: Rebooting cloudnet host cloudnet2006-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 13:58 wm-bot2: Rebooting all the cloudnet nodes cloudnet2006-dev,cloudnet2005-dev - cookbook ran by dcaro@vulcanus * 13:55 wm-bot2: Rebooting cloudnet host cloudnet2006-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 13:54 wm-bot2: Rebooting all the cloudnet nodes cloudnet2006-dev,cloudnet2005-dev - cookbook ran by dcaro@vulcanus * 13:52 wm-bot2: Rebooting cloudnet host cloudnet2006-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 13:51 wm-bot2: Rebooting all the cloudnet nodes cloudnet2006-dev,cloudnet2005-dev - cookbook ran by dcaro@vulcanus * 13:37 wm-bot2: Rebooting all the cloudnet nodes cloudnet2006-dev,cloudnet2005-dev - cookbook ran by dcaro@vulcanus * 13:34 wm-bot2: Rebooted cloudnet host cloudnet2006-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 13:31 wm-bot2: Rebooting cloudnet host cloudnet2006-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 13:30 wm-bot2: Rebooting all the cloudnet nodes cloudnet2006-dev,cloudnet2005-dev - cookbook ran by dcaro@vulcanus === 2022-06-30 === * 18:17 wm-bot2: Rebooted cloudnet host cloudnet2005-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 18:13 wm-bot2: Rebooting cloudnet host cloudnet2005-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 07:45 wm-bot2: Rebooted cloudnet host cloudnet2005-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 07:40 wm-bot2: Rebooting cloudnet host cloudnet2005-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 06:52 wm-bot2: Rebooting cloudnet host cloudnet2005-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus === 2022-06-29 === * 08:45 wm-bot2: Finished rebooting node cloudcephosd1021.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 08:41 wm-bot2: Rebooting node cloudcephosd1021.eqiad.wmnet - cookbook ran by dcaro@vulcanus === 2022-06-28 === * 13:03 taavi: grant the tools project access to the g3.cores16.ram64.disk20.10xiops flavor [[phab:T301949|T301949]] === 2022-06-22 === * 16:48 taavi: restart designate-*.service on both cloudservices nodes * 16:44 taavi: restart nova-conductor on all the cloudcontrol nodes * 12:50 andrewbogott: rebooting each eqiad1 cloudcontrol node in hopes of getting a baseline re: openstack instability === 2022-06-21 === * 04:48 andrewbogott: stopping nova-fullstack agent on cloudcontrol1003; it's going to page us otherwise and we're all AFK tomorrow * 04:02 andrewbogott: restarting rabbitmq on cloudcontrol100x (one at a time) === 2022-06-17 === * 17:53 andrewbogott: switching to a new python-based health check for galera and haproxy. This may make things more stable, or it may not. [[phab:T310664|T310664]] * 06:15 taavi: restart neutron-linuxbridge-agent on cloudvirt1046 === 2022-06-15 === * 11:34 taavi: restart neutron-linuxbridge-agent on cloudvirt1022 === 2022-06-14 === * 16:26 wm-bot2: OSDs (['cloudcephosd1001', 'cloudcephosd1002', 'cloudcephosd1003', 'cloudcephosd1004', 'cloudcephosd1005', 'cloudcephosd1006', 'cloudcephosd1007', 'cloudcephosd1008', 'cloudcephosd1009', 'cloudcephosd1010', 'cloudcephosd1011', 'cloudcephosd1012', 'cloudcephosd1013', 'cloudcephosd1014', 'cloudcephosd1015', 'cloudcephosd1016', 'cloudcephosd1017', 'cloudcephosd1018', 'cloudcephosd1019', 'cloudcephosd1020', 'cloudcephosd1021', 'cl * 14:38 wm-bot2: Upgrading OSDs and rebooting the nodes ['cloudcephosd1001', 'cloudcephosd1002', 'cloudcephosd1003', 'cloudcephosd1004', 'cloudcephosd1005', 'cloudcephosd1006', 'cloudcephosd1007', 'cloudcephosd1008', 'cloudcephosd1009', 'cloudcephosd1010', 'cloudcephosd1011', 'cloudcephosd1012', 'cloudcephosd1013', 'cloudcephosd1014', 'cloudcephosd1015', 'cloudcephosd1016', 'cloudcephosd1017', 'cloudcephosd1018', 'cloudcephosd1019', 'cloudceph * 12:55 wm-bot2: OSDs (['cloudcephosd2001-dev', 'cloudcephosd2002-dev', 'cloudcephosd2003-dev']) upgraded successfully B-) ([[phab:T309786|T309786]]) - cookbook ran by dcaro@vulcanus * 12:44 wm-bot2: Upgrading OSDs and rebooting the nodes ['cloudcephosd2001-dev', 'cloudcephosd2002-dev', 'cloudcephosd2003-dev'] ([[phab:T309786|T309786]]) - cookbook ran by dcaro@vulcanus === 2022-06-13 === * 11:14 wm-bot2: Finished rebooting node cloudcephosd1021.eqiad.wmnet ([[phab:T309789|T309789]]) - cookbook ran by dcaro@vulcanus * 11:08 wm-bot2: Rebooting node cloudcephosd1021.eqiad.wmnet ([[phab:T309789|T309789]]) - cookbook ran by dcaro@vulcanus * 11:07 wm-bot2: Rebooting node cloudcephosd1021.eqiad.wmnet ([[phab:T309789|T309789]]) - cookbook ran by dcaro@vulcanus * 11:05 wm-bot2: Rebooting node cloudcephosd1021.eqiad.wmnet ([[phab:T309789|T309789]]) - cookbook ran by dcaro@vulcanus * 11:04 wm-bot2: Rebooting node cloudcephosd1021.eqiad.wmnet ([[phab:T309789|T309789]]) - cookbook ran by dcaro@vulcanus * 11:03 wm-bot2: Rebooting node cloudcephosd1021.eqiad.wmnet ([[phab:T309789|T309789]]) - cookbook ran by dcaro@vulcanus * 09:15 wm-bot2: Rebooting node cloudcephosd1021.eqiad.wmnet ([[phab:T309789|T309789]]) - cookbook ran by dcaro@vulcanus === 2022-06-06 === * 13:21 andrewbogott: restarting mysql/galera on cloudcontrol100x in an attempt to stabilize some flapping there === 2022-06-02 === * 07:51 taavi: restart neutron-linuxbridge-agent.service on cloudvirt1034 [[phab:T309732|T309732]] * 00:37 andrewbogott: updated nameservers for codfw1dev instances via 'openstack subnet set --dns-nameserver etc.' === 2022-06-01 === * 17:11 andrewbogott: restarting designate services in cloudservices1xxx hosts * 16:37 taavi: root@cloudcontrol1005:~# cinder reset-state --state available 491066a4-16a9-4ce4-a9f6-{{Gerrit|182660616d77}} for [[phab:T309659|T309659]] === 2022-05-30 === * 17:43 andrewbogott: restarting neutron-rpc and neutron-api services on cloudcontrol1xxx === 2022-05-29 === * 14:55 andrewbogott: restarting nova services on all eqiad1 cloudcontrol nodes to recover from rabbit breakage * 14:15 andrewbogott: restarting rabbitmq on all eqiad1 cloudcontrol nodes (one at a time) === 2022-05-25 === * 20:01 balloons: clean up cinder backup a bit, restart service due to network outage * 20:01 balloons: cloudvirt1029 restarted nova due to network outage === 2022-05-19 === * 15:21 andrewbogott: resetting password for the 'troveguest' rabbitmq user. I think I may have broken this during a recent rebuild of the rabbitmq cluster === 2022-05-18 === * 15:42 andrewbogott: updated the 'debian-11.0-bullseye' glance image with a fresh build === 2022-05-14 === * 11:33 taavi: deleted projects 'ores' and 'ores-staging' [[phab:T308102|T308102]] === 2022-05-13 === * 06:20 wm-bot2: Safe reboot of 'cloudvirt1045.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 06:20 wm-bot2: Unset cloudvirt 'cloudvirt1045.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 06:16 wm-bot2: Drained 'cloudvirt1045.eqiad.wmnet'. - cookbook ran by andrew@buster * 06:16 wm-bot2: Set cloudvirt 'cloudvirt1045.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 06:15 wm-bot2: Draining 'cloudvirt1045.eqiad.wmnet'. - cookbook ran by andrew@buster * 06:15 wm-bot2: Safe rebooting 'cloudvirt1045.eqiad.wmnet'. - cookbook ran by andrew@buster * 06:11 wm-bot2: Safe reboot of 'cloudvirt1044.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 06:11 wm-bot2: Unset cloudvirt 'cloudvirt1044.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 06:10 wm-bot2: Set cloudvirt 'cloudvirt1045.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 06:10 wm-bot2: Draining 'cloudvirt1045.eqiad.wmnet'. - cookbook ran by andrew@buster * 06:09 wm-bot2: Safe rebooting 'cloudvirt1045.eqiad.wmnet'. - cookbook ran by andrew@buster * 06:07 wm-bot2: Drained 'cloudvirt1044.eqiad.wmnet'. - cookbook ran by andrew@buster * 06:06 wm-bot2: Set cloudvirt 'cloudvirt1045.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 06:06 wm-bot2: Draining 'cloudvirt1045.eqiad.wmnet'. - cookbook ran by andrew@buster * 06:05 wm-bot2: Safe rebooting 'cloudvirt1045.eqiad.wmnet'. - cookbook ran by andrew@buster * 05:51 wm-bot2: Set cloudvirt 'cloudvirt1045.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 05:50 wm-bot2: Draining 'cloudvirt1045.eqiad.wmnet'. - cookbook ran by andrew@buster * 05:50 wm-bot2: Safe rebooting 'cloudvirt1045.eqiad.wmnet'. - cookbook ran by andrew@buster * 05:49 wm-bot2: Set cloudvirt 'cloudvirt1044.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 05:49 wm-bot2: Safe reboot of 'cloudvirt1043.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 05:49 wm-bot2: Unset cloudvirt 'cloudvirt1043.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 05:49 wm-bot2: Draining 'cloudvirt1044.eqiad.wmnet'. - cookbook ran by andrew@buster * 05:49 wm-bot2: Safe rebooting 'cloudvirt1044.eqiad.wmnet'. - cookbook ran by andrew@buster * 05:47 wm-bot2: Safe reboot of 'cloudvirt1042.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 05:47 wm-bot2: Unset cloudvirt 'cloudvirt1042.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 05:45 wm-bot2: Drained 'cloudvirt1043.eqiad.wmnet'. - cookbook ran by andrew@buster * 05:45 wm-bot2: Set cloudvirt 'cloudvirt1043.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 05:45 wm-bot2: Draining 'cloudvirt1043.eqiad.wmnet'. - cookbook ran by andrew@buster * 05:44 wm-bot2: Safe rebooting 'cloudvirt1043.eqiad.wmnet'. - cookbook ran by andrew@buster * 05:44 wm-bot2: Drained 'cloudvirt1042.eqiad.wmnet'. - cookbook ran by andrew@buster * 05:42 wm-bot2: Set cloudvirt 'cloudvirt1043.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 05:42 wm-bot2: Draining 'cloudvirt1043.eqiad.wmnet'. - cookbook ran by andrew@buster * 05:42 wm-bot2: Safe rebooting 'cloudvirt1043.eqiad.wmnet'. - cookbook ran by andrew@buster * 05:41 wm-bot2: Set cloudvirt 'cloudvirt1043.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 05:40 wm-bot2: Draining 'cloudvirt1043.eqiad.wmnet'. - cookbook ran by andrew@buster * 05:40 wm-bot2: Safe rebooting 'cloudvirt1043.eqiad.wmnet'. - cookbook ran by andrew@buster * 05:38 wm-bot2: Set cloudvirt 'cloudvirt1042.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 05:37 wm-bot2: Draining 'cloudvirt1042.eqiad.wmnet'. - cookbook ran by andrew@buster * 05:37 wm-bot2: Safe rebooting 'cloudvirt1042.eqiad.wmnet'. - cookbook ran by andrew@buster * 05:30 wm-bot2: Set cloudvirt 'cloudvirt1042.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 05:29 wm-bot2: Draining 'cloudvirt1042.eqiad.wmnet'. - cookbook ran by andrew@buster * 05:29 wm-bot2: Safe rebooting 'cloudvirt1042.eqiad.wmnet'. - cookbook ran by andrew@buster * 05:19 wm-bot2: Set cloudvirt 'cloudvirt1043.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 05:18 wm-bot2: Set cloudvirt 'cloudvirt1042.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 05:18 wm-bot2: Draining 'cloudvirt1043.eqiad.wmnet'. - cookbook ran by andrew@buster * 05:18 wm-bot2: Safe rebooting 'cloudvirt1043.eqiad.wmnet'. - cookbook ran by andrew@buster * 05:18 wm-bot2: Draining 'cloudvirt1042.eqiad.wmnet'. - cookbook ran by andrew@buster * 05:18 wm-bot2: Safe rebooting 'cloudvirt1042.eqiad.wmnet'. - cookbook ran by andrew@buster * 05:12 wm-bot2: Safe reboot of 'cloudvirt1040.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 05:12 wm-bot2: Unset cloudvirt 'cloudvirt1040.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 05:08 wm-bot2: Drained 'cloudvirt1040.eqiad.wmnet'. - cookbook ran by andrew@buster * 05:02 wm-bot2: Set cloudvirt 'cloudvirt1042.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 05:02 wm-bot2: Set cloudvirt 'cloudvirt1040.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 05:02 wm-bot2: Draining 'cloudvirt1042.eqiad.wmnet'. - cookbook ran by andrew@buster * 05:02 wm-bot2: Safe rebooting 'cloudvirt1042.eqiad.wmnet'. - cookbook ran by andrew@buster * 05:02 wm-bot2: Draining 'cloudvirt1040.eqiad.wmnet'. - cookbook ran by andrew@buster * 05:01 wm-bot2: Safe rebooting 'cloudvirt1040.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:52 wm-bot2: Set cloudvirt 'cloudvirt1042.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 04:51 wm-bot2: Draining 'cloudvirt1042.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:51 wm-bot2: Safe rebooting 'cloudvirt1042.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:48 wm-bot2: Safe reboot of 'cloudvirt1041.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 04:48 wm-bot2: Unset cloudvirt 'cloudvirt1041.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 04:44 wm-bot2: Drained 'cloudvirt1041.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:31 wm-bot2: Set cloudvirt 'cloudvirt1041.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 04:30 wm-bot2: Draining 'cloudvirt1041.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:30 wm-bot2: Safe rebooting 'cloudvirt1041.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:30 wm-bot2: Safe reboot of 'cloudvirt1039.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 04:30 wm-bot2: Unset cloudvirt 'cloudvirt1039.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 04:27 wm-bot2: Set cloudvirt 'cloudvirt1040.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 04:26 wm-bot2: Draining 'cloudvirt1040.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:26 wm-bot2: Safe rebooting 'cloudvirt1040.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:26 wm-bot2: Drained 'cloudvirt1039.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:26 wm-bot2: Set cloudvirt 'cloudvirt1039.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 04:25 wm-bot2: Draining 'cloudvirt1039.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:25 wm-bot2: Safe rebooting 'cloudvirt1039.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:24 wm-bot2: Set cloudvirt 'cloudvirt1040.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 04:23 wm-bot2: Draining 'cloudvirt1040.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:23 wm-bot2: Safe rebooting 'cloudvirt1040.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:23 wm-bot2: Set cloudvirt 'cloudvirt1040.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 04:22 wm-bot2: Draining 'cloudvirt1040.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:22 wm-bot2: Safe rebooting 'cloudvirt1040.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:21 wm-bot2: Safe reboot of 'cloudvirt1038.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 04:21 wm-bot2: Unset cloudvirt 'cloudvirt1038.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 04:18 wm-bot2: Drained 'cloudvirt1038.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:16 wm-bot2: Set cloudvirt 'cloudvirt1038.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 04:16 wm-bot2: Set cloudvirt 'cloudvirt1039.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 04:16 wm-bot2: Draining 'cloudvirt1038.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:16 wm-bot2: Safe rebooting 'cloudvirt1038.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:15 wm-bot2: Draining 'cloudvirt1039.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:15 wm-bot2: Safe rebooting 'cloudvirt1039.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:37 wm-bot2: Set cloudvirt 'cloudvirt1039.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 03:36 wm-bot2: Draining 'cloudvirt1039.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:36 wm-bot2: Safe rebooting 'cloudvirt1039.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:34 wm-bot2: Safe reboot of 'cloudvirt1037.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 03:34 wm-bot2: Unset cloudvirt 'cloudvirt1037.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 03:27 wm-bot2: Drained 'cloudvirt1037.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:27 wm-bot2: Set cloudvirt 'cloudvirt1037.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 03:26 wm-bot2: Draining 'cloudvirt1037.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:26 wm-bot2: Safe rebooting 'cloudvirt1037.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:26 wm-bot2: Set cloudvirt 'cloudvirt1038.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 03:25 wm-bot2: Draining 'cloudvirt1038.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:25 wm-bot2: Safe rebooting 'cloudvirt1038.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:22 wm-bot2: Unset cloudvirt 'cloudvirt1036.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 02:55 wm-bot2: Set cloudvirt 'cloudvirt1037.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 02:55 wm-bot2: Set cloudvirt 'cloudvirt1036.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 02:54 wm-bot2: Draining 'cloudvirt1037.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:54 wm-bot2: Safe rebooting 'cloudvirt1037.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:54 wm-bot2: Draining 'cloudvirt1036.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:54 wm-bot2: Safe rebooting 'cloudvirt1036.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:05 wm-bot2: Set cloudvirt 'cloudvirt1037.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 02:05 wm-bot2: Draining 'cloudvirt1037.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:05 wm-bot2: Safe rebooting 'cloudvirt1037.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:04 wm-bot2: Set cloudvirt 'cloudvirt1036.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 02:04 wm-bot2: Draining 'cloudvirt1036.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:03 wm-bot2: Safe rebooting 'cloudvirt1036.eqiad.wmnet'. - cookbook ran by andrew@buster * 01:23 wm-bot2: Safe reboot of 'cloudvirt1035.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 01:23 wm-bot2: Unset cloudvirt 'cloudvirt1035.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 01:19 wm-bot2: Drained 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@buster * 01:01 wm-bot2: Set cloudvirt 'cloudvirt1035.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 01:01 wm-bot2: Set cloudvirt 'cloudvirt1036.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 01:00 wm-bot2: Draining 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@buster * 01:00 wm-bot2: Safe rebooting 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@buster * 01:00 wm-bot2: Draining 'cloudvirt1036.eqiad.wmnet'. - cookbook ran by andrew@buster * 01:00 wm-bot2: Safe rebooting 'cloudvirt1036.eqiad.wmnet'. - cookbook ran by andrew@buster * 00:25 wm-bot2: Safe reboot of 'cloudvirt1033.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 00:25 wm-bot2: Unset cloudvirt 'cloudvirt1033.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 00:21 wm-bot2: Drained 'cloudvirt1033.eqiad.wmnet'. - cookbook ran by andrew@buster * 00:20 wm-bot2: Set cloudvirt 'cloudvirt1035.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 00:19 wm-bot2: Draining 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@buster * 00:19 wm-bot2: Safe rebooting 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@buster * 00:11 wm-bot2: Safe reboot of 'cloudvirt1034.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 00:11 wm-bot2: Unset cloudvirt 'cloudvirt1034.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 00:07 wm-bot2: Drained 'cloudvirt1034.eqiad.wmnet'. - cookbook ran by andrew@buster === 2022-05-12 === * 23:55 wm-bot2: Set cloudvirt 'cloudvirt1034.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 23:55 wm-bot2: Set cloudvirt 'cloudvirt1033.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 23:54 wm-bot2: Draining 'cloudvirt1034.eqiad.wmnet'. - cookbook ran by andrew@buster * 23:54 wm-bot2: Safe rebooting 'cloudvirt1034.eqiad.wmnet'. - cookbook ran by andrew@buster * 23:54 wm-bot2: Draining 'cloudvirt1033.eqiad.wmnet'. - cookbook ran by andrew@buster * 23:54 wm-bot2: Safe rebooting 'cloudvirt1033.eqiad.wmnet'. - cookbook ran by andrew@buster * 22:23 wm-bot2: Safe reboot of 'cloudvirt1031.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 22:23 wm-bot2: Unset cloudvirt 'cloudvirt1031.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 22:20 wm-bot2: Drained 'cloudvirt1031.eqiad.wmnet'. - cookbook ran by andrew@buster * 22:17 wm-bot2: Safe reboot of 'cloudvirt1032.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 22:17 wm-bot2: Unset cloudvirt 'cloudvirt1032.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 22:13 wm-bot2: Drained 'cloudvirt1032.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:57 wm-bot2: Set cloudvirt 'cloudvirt1032.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 21:56 wm-bot2: Draining 'cloudvirt1032.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:56 wm-bot2: Safe rebooting 'cloudvirt1032.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:55 wm-bot2: Draining 'cloudvirt1031.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:55 wm-bot2: Safe rebooting 'cloudvirt1031.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:54 wm-bot2: Safe reboot of 'cloudvirt1030.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 21:54 wm-bot2: Unset cloudvirt 'cloudvirt1030.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 21:53 wm-bot2: Set cloudvirt 'cloudvirt1031.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 21:52 wm-bot2: Draining 'cloudvirt1031.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:52 wm-bot2: Safe rebooting 'cloudvirt1031.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:51 wm-bot2: Drained 'cloudvirt1030.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:44 wm-bot2: Safe reboot of 'cloudvirt1029.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 21:44 wm-bot2: Unset cloudvirt 'cloudvirt1029.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 21:42 wm-bot2: Drained 'cloudvirt1029.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:36 wm-bot2: Safe reboot of 'cloudvirt-wdqs1001.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 21:36 wm-bot2: Unset cloudvirt 'cloudvirt-wdqs1001.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 21:33 wm-bot2: Drained 'cloudvirt-wdqs1001.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:33 wm-bot2: Set cloudvirt 'cloudvirt-wdqs1001.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 21:32 wm-bot2: Draining 'cloudvirt-wdqs1001.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:32 wm-bot2: Safe rebooting 'cloudvirt-wdqs1001.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:32 wm-bot2: Set cloudvirt 'cloudvirt1029.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 21:31 wm-bot2: Safe reboot of 'cloudvirt-wdqs1002.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 21:31 wm-bot2: Unset cloudvirt 'cloudvirt-wdqs1002.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 21:31 wm-bot2: Draining 'cloudvirt1029.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:31 wm-bot2: Safe rebooting 'cloudvirt1029.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:30 wm-bot2: Set cloudvirt 'cloudvirt1030.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 21:29 wm-bot2: Draining 'cloudvirt1030.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:29 wm-bot2: Safe rebooting 'cloudvirt1030.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:29 wm-bot2: Drained 'cloudvirt-wdqs1002.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:28 wm-bot2: Set cloudvirt 'cloudvirt-wdqs1002.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 21:28 wm-bot2: Draining 'cloudvirt-wdqs1002.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:28 wm-bot2: Safe rebooting 'cloudvirt-wdqs1002.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:22 wm-bot2: Safe reboot of 'cloudvirt1026.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 21:22 wm-bot2: Unset cloudvirt 'cloudvirt1026.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 21:21 wm-bot2: Safe reboot of 'cloudvirt-wdqs1003.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 21:21 wm-bot2: Unset cloudvirt 'cloudvirt-wdqs1003.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 21:18 wm-bot2: Drained 'cloudvirt-wdqs1003.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:18 wm-bot2: Drained 'cloudvirt1026.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:18 wm-bot2: Set cloudvirt 'cloudvirt-wdqs1003.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 21:17 wm-bot2: Draining 'cloudvirt-wdqs1003.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:17 wm-bot2: Safe rebooting 'cloudvirt-wdqs1003.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:17 wm-bot2: Set cloudvirt 'cloudvirt1029.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 21:16 wm-bot2: Draining 'cloudvirt1029.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:16 wm-bot2: Safe rebooting 'cloudvirt1029.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:14 wm-bot2: Safe reboot of 'cloudvirt1025.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 21:14 wm-bot2: Unset cloudvirt 'cloudvirt1025.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 21:11 wm-bot2: Safe reboot of 'cloudvirt1046.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 21:11 wm-bot2: Unset cloudvirt 'cloudvirt1046.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 21:10 wm-bot2: Drained 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:08 wm-bot2: Drained 'cloudvirt1046.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:08 wm-bot2: Set cloudvirt 'cloudvirt1046.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 21:07 wm-bot2: Draining 'cloudvirt1046.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:07 wm-bot2: Safe rebooting 'cloudvirt1046.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:05 wm-bot2: Set cloudvirt 'cloudvirt1046.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 21:04 wm-bot2: Draining 'cloudvirt1046.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:04 wm-bot2: Safe rebooting 'cloudvirt1046.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:00 wm-bot2: Set cloudvirt 'cloudvirt1046.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 20:59 wm-bot2: Draining 'cloudvirt1046.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:59 wm-bot2: Safe rebooting 'cloudvirt1046.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:59 wm-bot2: Safe reboot of 'cloudvirt1047.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 20:59 wm-bot2: Unset cloudvirt 'cloudvirt1047.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 20:57 wm-bot2: Set cloudvirt 'cloudvirt1026.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 20:57 wm-bot2: Draining 'cloudvirt1026.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:57 wm-bot2: Safe rebooting 'cloudvirt1026.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:55 wm-bot2: Set cloudvirt 'cloudvirt1026.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 20:55 wm-bot2: Drained 'cloudvirt1047.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:54 wm-bot2: Draining 'cloudvirt1026.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:54 wm-bot2: Safe rebooting 'cloudvirt1026.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:54 wm-bot2: Safe reboot of 'cloudvirt1024.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 20:54 wm-bot2: Unset cloudvirt 'cloudvirt1024.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 20:53 wm-bot2: Set cloudvirt 'cloudvirt1047.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 20:52 wm-bot2: Draining 'cloudvirt1047.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:52 wm-bot2: Safe rebooting 'cloudvirt1047.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:50 wm-bot2: Drained 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:49 wm-bot2: Set cloudvirt 'cloudvirt1025.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 20:49 wm-bot2: Draining 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:49 wm-bot2: Safe rebooting 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:48 wm-bot2: Safe reboot of 'cloudvirt1023.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 20:48 wm-bot2: Unset cloudvirt 'cloudvirt1023.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 20:44 wm-bot2: Drained 'cloudvirt1023.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:44 wm-bot2: Set cloudvirt 'cloudvirt1023.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 20:43 wm-bot2: Draining 'cloudvirt1023.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:43 wm-bot2: Safe rebooting 'cloudvirt1023.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:34 wm-bot2: Set cloudvirt 'cloudvirt1023.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 20:34 wm-bot2: Set cloudvirt 'cloudvirt1024.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 20:34 wm-bot2: Draining 'cloudvirt1023.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:34 wm-bot2: Safe rebooting 'cloudvirt1023.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:34 wm-bot2: Draining 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:34 wm-bot2: Safe rebooting 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:31 wm-bot2: Safe reboot of 'cloudvirt1027.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 20:31 wm-bot2: Unset cloudvirt 'cloudvirt1027.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 20:28 wm-bot2: Drained 'cloudvirt1027.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:28 wm-bot2: Set cloudvirt 'cloudvirt1027.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 20:27 wm-bot2: Draining 'cloudvirt1027.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:27 wm-bot2: Safe rebooting 'cloudvirt1027.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:23 wm-bot2: Set cloudvirt 'cloudvirt1023.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 20:22 wm-bot2: Draining 'cloudvirt1023.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:22 wm-bot2: Safe rebooting 'cloudvirt1023.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:11 wm-bot2: Set cloudvirt 'cloudvirt1027.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 20:10 wm-bot2: Draining 'cloudvirt1027.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:10 wm-bot2: Safe rebooting 'cloudvirt1027.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:07 wm-bot2: Set cloudvirt 'cloudvirt1023.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 20:07 wm-bot2: Draining 'cloudvirt1023.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:06 wm-bot2: Safe rebooting 'cloudvirt1023.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:06 wm-bot2: Safe reboot of 'cloudvirt1022.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 20:05 wm-bot2: Unset cloudvirt 'cloudvirt1022.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 20:02 wm-bot2: Drained 'cloudvirt1022.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:01 wm-bot2: Set cloudvirt 'cloudvirt1022.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 20:00 wm-bot2: Draining 'cloudvirt1022.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:00 wm-bot2: Safe rebooting 'cloudvirt1022.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:58 wm-bot2: Set cloudvirt 'cloudvirt1022.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 19:57 wm-bot2: Draining 'cloudvirt1022.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:57 wm-bot2: Safe rebooting 'cloudvirt1022.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:36 wm-bot2: Set cloudvirt 'cloudvirt1022.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 19:35 wm-bot2: Draining 'cloudvirt1022.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:35 wm-bot2: Safe rebooting 'cloudvirt1022.eqiad.wmnet'. - cookbook ran by andrew@buster * 15:06 andrewbogott: stopping nfs-server on labstore1004 in preparation for reboot * 04:12 andrewbogott: rebooting primary bastion (bastion-eqiad1-03.bastion.eqiad1.wikimedia.cloud) in hopes of resolving a problem with ssh proxying === 2022-05-11 === * 18:48 wm-bot2: Set cloudvirt 'cloudvirt1022.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 18:48 wm-bot2: Draining 'cloudvirt1022.eqiad.wmnet'. - cookbook ran by andrew@buster * 18:48 wm-bot2: Safe rebooting 'cloudvirt1022.eqiad.wmnet'. - cookbook ran by andrew@buster * 18:39 wm-bot2: Set cloudvirt 'cloudvirt1027.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 18:38 wm-bot2: Draining 'cloudvirt1027.eqiad.wmnet'. - cookbook ran by andrew@buster * 18:38 wm-bot2: Safe rebooting 'cloudvirt1027.eqiad.wmnet'. - cookbook ran by andrew@buster * 18:04 wm-bot2: Set cloudvirt 'cloudvirt1027.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 18:03 wm-bot2: Draining 'cloudvirt1027.eqiad.wmnet'. - cookbook ran by andrew@buster * 18:03 wm-bot2: Safe rebooting 'cloudvirt1027.eqiad.wmnet'. - cookbook ran by andrew@buster * 08:56 wm-bot2: Finished rebooting node cloudcephosd1021.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 08:52 wm-bot2: Rebooting node cloudcephosd1021.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 07:53 dcaro: test * 04:28 wm-bot2: Set cloudvirt 'cloudvirt1022.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 04:27 wm-bot2: Draining 'cloudvirt1022.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:27 wm-bot2: Safe rebooting 'cloudvirt1022.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:44 wm-bot2: Set cloudvirt 'cloudvirt1022.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 03:43 wm-bot2: Draining 'cloudvirt1022.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:43 wm-bot2: Safe rebooting 'cloudvirt1022.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:42 wm-bot2: Safe reboot of 'cloudvirt1021.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 03:42 wm-bot2: Unset cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 03:39 wm-bot2: Drained 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:23 wm-bot2: Set cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 03:22 wm-bot2: Draining 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:22 wm-bot2: Safe rebooting 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:09 wm-bot2: Set cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 03:08 wm-bot2: Draining 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:08 wm-bot2: Safe rebooting 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:04 andrewbogott: reset and recreated the rabbitmq cluster in eqiad1 to get around some broken queues. * 03:02 wm-bot2: Set cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 03:01 wm-bot2: Draining 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:01 wm-bot2: Safe rebooting 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:28 wm-bot: Set cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 02:25 wm-bot: Draining 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:25 wm-bot: Safe rebooting 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster === 2022-05-10 === * 21:43 wm-bot: Set cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 21:40 wm-bot: Draining 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:40 wm-bot: Safe rebooting 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:35 wm-bot: Set cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 21:32 wm-bot: Draining 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:32 wm-bot: Safe rebooting 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:05 wm-bot: Set cloudvirt 'cloudvirt1023.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 20:02 wm-bot: Draining 'cloudvirt1023.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:01 wm-bot: Safe rebooting 'cloudvirt1023.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:00 wm-bot: Draining 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:00 wm-bot: Safe rebooting 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:57 wm-bot: Draining 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:57 wm-bot: Safe rebooting 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:55 wm-bot: Draining 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:55 wm-bot: Safe rebooting 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:47 wm-bot: Set cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 19:46 wm-bot: Draining 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:46 wm-bot: Safe rebooting 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:45 wm-bot: Draining 'cloudvirt1022.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:45 wm-bot: Safe rebooting 'cloudvirt1022.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:44 wm-bot: Draining 'cloudvirt1022.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:44 wm-bot: Safe rebooting 'cloudvirt1022.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:40 wm-bot: Set cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 19:39 wm-bot: Draining 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:39 wm-bot: Safe rebooting 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:37 wm-bot: Set cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 19:36 wm-bot: Draining 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:36 wm-bot: Safe rebooting 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:33 wm-bot: Safe reboot of 'cloudvirt1017.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 19:33 wm-bot: Unset cloudvirt 'cloudvirt1017.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 19:29 wm-bot: Drained 'cloudvirt1017.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:06 wm-bot: Set cloudvirt 'cloudvirt1017.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 19:05 wm-bot: Draining 'cloudvirt1017.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:05 wm-bot: Safe rebooting 'cloudvirt1017.eqiad.wmnet'. - cookbook ran by andrew@buster * 15:41 andrewbogott: rebooting cloud*-dev for [[phab:T307668|T307668]] * 13:59 taavi: manually attached [[User:Dreamy Jazz]] to wikitech for a password reset (https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin#Manually_associate_an_LDAP_account_with_wikitech) === 2022-05-07 === * 01:33 wm-bot: Drained 'cloudvirt1016.eqiad.wmnet'. - cookbook ran by andrew@buster * 01:32 wm-bot: Set cloudvirt 'cloudvirt1016.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 01:30 wm-bot: Draining 'cloudvirt1016.eqiad.wmnet'. - cookbook ran by andrew@buster * 01:21 wm-bot: Set cloudvirt 'cloudvirt1016.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 01:18 wm-bot: Draining 'cloudvirt1016.eqiad.wmnet'. - cookbook ran by andrew@buster === 2022-05-03 === * 20:38 andrewbogott: upgrading clouddb2001-dev in place * 18:18 taavi: updated 'puppet-enc' endpoints on the keystone catalog to use https and port 443 === 2022-05-02 === * 16:56 dcaro: rebooting cloudmetrics1001 === 2022-04-29 === * 14:22 andrewbogott: changing login.toolforge.org, bastion.toolforge.org, and dev.toolforge.org dns entries to refer to the new Buster bastions [[phab:T277653|T277653]] https://wikitech.wikimedia.org/wiki/News/Toolforge_Stretch_deprecation#Timeline === 2022-04-27 === * 14:51 wm-bot: Finished rebooting the nodes ['cloudcephosd1001', 'cloudcephosd1002', 'cloudcephosd1003', 'cloudcephosd1004', 'cloudcephosd1005', 'cloudcephosd1006', 'cloudcephosd1007', 'cloudcephosd1008', 'cloudcephosd1009', 'cloudcephosd1010', 'cloudcephosd1011', 'cloudcephosd1012', 'cloudcephosd1013', 'cloudcephosd1014', 'cloudcephosd1015', 'cloudcephosd1016', 'cloudcephosd1017', 'cloudcephosd1018', 'cloudcephosd1019', 'cloudcephosd1020', 'cloud * 14:50 wm-bot: Finished rebooting node cloudcephosd1024.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 14:46 wm-bot: Rebooting node cloudcephosd1024.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 14:46 wm-bot: Finished rebooting node cloudcephosd1023.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 14:41 wm-bot: Rebooting node cloudcephosd1023.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 14:41 wm-bot: Finished rebooting node cloudcephosd1022.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 14:35 wm-bot: Rebooting node cloudcephosd1022.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 14:35 wm-bot: Finished rebooting node cloudcephosd1021.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 14:31 wm-bot: Rebooting node cloudcephosd1021.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 14:31 wm-bot: Finished rebooting node cloudcephosd1020.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 14:27 wm-bot: Rebooting node cloudcephosd1020.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 14:27 wm-bot: Finished rebooting node cloudcephosd1019.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 14:23 wm-bot: Rebooting node cloudcephosd1019.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 14:23 wm-bot: Finished rebooting node cloudcephosd1018.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 14:13 wm-bot: Rebooting node cloudcephosd1018.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 14:13 wm-bot: Finished rebooting node cloudcephosd1017.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 14:09 wm-bot: Rebooting node cloudcephosd1017.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 14:09 wm-bot: Finished rebooting node cloudcephosd1016.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 14:05 wm-bot: Rebooting node cloudcephosd1016.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 14:05 wm-bot: Finished rebooting node cloudcephosd1015.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 14:01 wm-bot: Rebooting node cloudcephosd1015.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 14:01 wm-bot: Finished rebooting node cloudcephosd1014.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 13:57 wm-bot: Rebooting node cloudcephosd1014.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 13:57 wm-bot: Finished rebooting node cloudcephosd1013.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 13:44 wm-bot: Rebooting node cloudcephosd1013.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 13:43 wm-bot: Finished rebooting node cloudcephosd1012.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 13:39 wm-bot: Rebooting node cloudcephosd1012.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 13:39 wm-bot: Finished rebooting node cloudcephosd1011.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 13:35 wm-bot: Rebooting node cloudcephosd1011.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 13:35 wm-bot: Finished rebooting node cloudcephosd1010.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 13:31 wm-bot: Rebooting node cloudcephosd1010.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 13:31 wm-bot: Finished rebooting node cloudcephosd1009.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 13:26 wm-bot: Rebooting node cloudcephosd1009.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 13:26 wm-bot: Finished rebooting node cloudcephosd1008.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 13:14 wm-bot: Rebooting node cloudcephosd1008.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 13:14 wm-bot: Finished rebooting node cloudcephosd1007.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 13:10 wm-bot: Rebooting node cloudcephosd1007.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 13:10 wm-bot: Finished rebooting node cloudcephosd1006.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 13:05 wm-bot: Rebooting node cloudcephosd1006.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 13:05 wm-bot: Finished rebooting node cloudcephosd1005.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 13:01 wm-bot: Rebooting node cloudcephosd1005.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 13:01 wm-bot: Finished rebooting node cloudcephosd1004.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 12:57 wm-bot: Rebooting node cloudcephosd1004.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 12:57 wm-bot: Finished rebooting node cloudcephosd1003.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 12:53 wm-bot: Rebooting node cloudcephosd1003.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 12:53 wm-bot: Finished rebooting node cloudcephosd1002.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 12:50 wm-bot: Rebooting node cloudcephosd1002.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 12:50 wm-bot: Finished rebooting node cloudcephosd1001.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 12:46 wm-bot: Rebooting node cloudcephosd1001.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 12:46 wm-bot: Rebooting the nodes cloudcephosd1001,cloudcephosd1002,cloudcephosd1003,cloudcephosd1004,cloudcephosd1005,cloudcephosd1006,cloudcephosd1007,cloudcephosd1008,cloudcephosd1009,cloudcephosd1010,cloudcephosd1011,cloudcephosd1012,cloudcephosd1013,cloudcephosd1014,cloudcephosd1015,cloudcephosd1016,cloudcephosd1017,cloudcephosd1018,cloudcephosd1019,cloudcephosd1020,cloudcephosd1021,cloudcephosd1022,cloudcephosd1023,cloudcephosd1024 - cookbo * 12:15 wm-bot: Finished rebooting the nodes ['cloudcephmon1001', 'cloudcephmon1002', 'cloudcephmon1003'] - cookbook ran by dcaro@vulcanus * 12:15 wm-bot: Finished rebooting node cloudcephmon1003.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 12:12 wm-bot: Rebooting node cloudcephmon1003.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 12:12 wm-bot: Finished rebooting node cloudcephmon1002.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 12:09 wm-bot: Rebooting node cloudcephmon1002.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 12:09 wm-bot: Finished rebooting node cloudcephmon1001.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 12:07 wm-bot: Rebooting node cloudcephmon1001.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 12:07 wm-bot: Rebooting the nodes cloudcephmon1001,cloudcephmon1002,cloudcephmon1003 - cookbook ran by dcaro@vulcanus * 12:05 wm-bot: Finished rebooting the nodes ['cloudcephosd2001-dev', 'cloudcephosd2002-dev', 'cloudcephosd2003-dev'] - cookbook ran by dcaro@vulcanus * 12:05 wm-bot: Finished rebooting node cloudcephosd2003-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 12:02 wm-bot: Rebooting node cloudcephosd2003-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 12:02 wm-bot: Finished rebooting node cloudcephosd2002-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 11:59 wm-bot: Rebooting node cloudcephosd2002-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 11:59 wm-bot: Finished rebooting node cloudcephosd2001-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 11:56 wm-bot: Rebooting node cloudcephosd2001-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 11:56 wm-bot: Rebooting the nodes cloudcephosd2001-dev,cloudcephosd2002-dev,cloudcephosd2003-dev - cookbook ran by dcaro@vulcanus * 11:55 wm-bot: Finished rebooting the nodes ['cloudcephmon2004-dev', 'cloudcephmon2005-dev', 'cloudcephmon2006-dev'] - cookbook ran by dcaro@vulcanus * 11:55 wm-bot: Finished rebooting node cloudcephmon2006-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 11:52 wm-bot: Rebooting node cloudcephmon2006-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 11:52 wm-bot: Finished rebooting node cloudcephmon2005-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 11:47 wm-bot: Rebooting node cloudcephmon2005-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 11:47 wm-bot: Finished rebooting node cloudcephmon2004-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 11:43 wm-bot: Rebooting node cloudcephmon2004-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 11:43 wm-bot: Rebooting the nodes cloudcephmon2004-dev,cloudcephmon2005-dev,cloudcephmon2006-dev - cookbook ran by dcaro@vulcanus === 2022-04-26 === * 10:36 taavi: [codfw1dev] updated designate pool to 2004/2005-dev according to the instructions on https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/DNS/Designate#Initial_designate/pdns_node_setup === 2022-04-22 === * 10:33 taavi: [codfw1dev] restart designate-sink on both new cloudservices host to fix rabbitmq connectivity === 2022-04-21 === * 05:38 andrewbogott: replaced cloudservices200[2,3] with cloudservices200[4,5] === 2022-04-19 === * 15:29 andrewbogott: stopping all VMs on cloudvirt1019, reimaging host === 2022-04-18 === * 15:23 andrewbogott: reimaging cloudvirt1020, leaving VMs in place * 13:40 andrewbogott: shutting down many codfdfw1dev servers (including network infra!) for [[phab:T305469|T305469]] === 2022-04-14 === * 20:14 andrewbogott: restarting nova-api and nova-conductor services in a superstitious attempt to reduce open DB connections === 2022-04-13 === * 22:01 andrewbogott: restarting galera on cloudcontrols (one by one) to clear open connections === 2022-04-11 === * 15:59 taavi: created cloudinfra.wmcloud.org zone === 2022-04-09 === * 19:55 andrewbogott: reimaging cloudbackup1001-dev to bullseye * 19:37 taavi: add 'puppet-enc' service & endpoint to keystone [[phab:T274666|T274666]] * 19:25 andrewbogott: reimaging cloudbackup1002-dev to bullseye === 2022-04-07 === * 12:51 wm-bot: Set cloudvirt 'cloudvirt1016.eqiad.wmnet' maintenance. ([[phab:T305631|T305631]]) - cookbook ran by arturo@nostromo === 2022-04-06 === * 09:12 arturo: [codf1dev] installing python3-eventlet 0.30.2-5~bpo11+1 on all required servers (cloudvirt, cloudnet, cloudcontrol) ([[phab:T305157|T305157]]) * 08:45 arturo: [codfw1dev] trying with python3-eventlet 0.30.2-5 installed by hand on cloudvirt2003-dev ([[phab:T305157|T305157]]) * 08:42 arturo: [codfw1dev] trying with python3-eventlet 0.30.2-5 installed by hand on cloudcontrol servers ([[phab:T305157|T305157]]) * 08:24 arturo: [codfw1dev] trying with python3-dnspython 2.2.0-2 installed by hand on cloudvirt2003-dev ([[phab:T305157|T305157]]) * 08:20 arturo: [codfw1dev] trying with python3-dnspython 2.2.0-2 installed by hand on cloudcontrol servers ([[phab:T305157|T305157]]) === 2022-03-30 === * 11:20 arturo: apply urpf strict filter to eqiad cloud-hosts vlan - [[phab:T285461|T285461]] === 2022-03-29 === * 10:02 dcaro: restarting keystone ([[phab:T304918|T304918]]) === 2022-03-23 === * 22:53 wm-bot: Drained 'cloudvirt1045.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 22:38 wm-bot: Drained 'cloudvirt1044.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 22:12 wm-bot: Set cloudvirt 'cloudvirt1045.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 22:12 wm-bot: Draining 'cloudvirt1045.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 22:08 wm-bot: Set cloudvirt 'cloudvirt1043.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 22:07 wm-bot: Set cloudvirt 'cloudvirt1044.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 22:06 wm-bot: Draining 'cloudvirt1044.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 22:06 wm-bot: Draining 'cloudvirt1043.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 21:54 wm-bot: Drained 'cloudvirt1042.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 21:19 wm-bot: Set cloudvirt 'cloudvirt1042.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 21:19 wm-bot: Draining 'cloudvirt1042.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 21:12 wm-bot: Drained 'cloudvirt1040.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 21:12 wm-bot: Set cloudvirt 'cloudvirt1040.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 21:09 wm-bot: Draining 'cloudvirt1040.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 21:07 wm-bot: Set cloudvirt 'cloudvirt1040.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 21:04 wm-bot: Draining 'cloudvirt1040.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 20:55 wm-bot: Set cloudvirt 'cloudvirt1041.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 20:54 wm-bot: Draining 'cloudvirt1041.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 20:30 wm-bot: Drained 'cloudvirt1039.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 20:15 wm-bot: Set cloudvirt 'cloudvirt1040.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 20:15 wm-bot: Set cloudvirt 'cloudvirt1039.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 20:14 wm-bot: Draining 'cloudvirt1040.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 20:14 wm-bot: Draining 'cloudvirt1039.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 18:44 wm-bot: Set cloudvirt 'cloudvirt1038.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 18:43 wm-bot: Draining 'cloudvirt1038.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 18:19 wm-bot: Set cloudvirt 'cloudvirt1037.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 18:18 wm-bot: Draining 'cloudvirt1037.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 18:13 wm-bot: Drained 'cloudvirt1036.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 18:02 wm-bot2: Testing wm-bot relay to #wikimedia-cloud-feed * 17:55 wm-bot: Set cloudvirt 'cloudvirt1036.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 17:54 wm-bot: Draining 'cloudvirt1036.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 17:04 wm-bot: Set cloudvirt 'cloudvirt1035.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 17:03 wm-bot: Draining 'cloudvirt1035.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 17:03 wm-bot: Drained 'cloudvirt1034.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 16:51 wm-bot: Drained 'cloudvirt1033.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 16:37 wm-bot: Set cloudvirt 'cloudvirt1034.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 16:37 wm-bot: Set cloudvirt 'cloudvirt1033.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 16:36 wm-bot: Draining 'cloudvirt1034.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 16:36 wm-bot: Draining 'cloudvirt1033.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 15:01 wm-bot: Drained 'cloudvirt1032.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 15:00 wm-bot: Set cloudvirt 'cloudvirt1032.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 14:57 wm-bot: Draining 'cloudvirt1032.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 14:44 wm-bot: Drained 'cloudvirt1031.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 14:35 wm-bot: Set cloudvirt 'cloudvirt1032.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 14:34 wm-bot: Draining 'cloudvirt1032.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 14:32 wm-bot: Drained 'cloudvirt1030.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 14:20 wm-bot: Set cloudvirt 'cloudvirt1031.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 14:19 wm-bot: Draining 'cloudvirt1031.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 14:18 wm-bot: Set cloudvirt 'cloudvirt1030.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 14:17 wm-bot: Draining 'cloudvirt1030.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 13:54 taavi: restart nova-fullstack on cloudcontrol1003 to pick up bastion ip change * 13:43 wm-bot: Drained 'cloudvirt1029.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 13:23 wm-bot: Set cloudvirt 'cloudvirt1029.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 13:22 wm-bot: Draining 'cloudvirt1029.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster === 2022-03-22 === * 22:59 wm-bot: Set cloudvirt 'cloudvirt1027.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 22:58 wm-bot: Draining 'cloudvirt1027.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster === 2022-03-17 === * 01:09 wm-bot: Drained 'cloudvirt1016.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 00:53 wm-bot: Set cloudvirt 'cloudvirt1016.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 00:52 wm-bot: Setting cloudvirt 'cloudvirt1016.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 00:52 wm-bot: Draining 'cloudvirt1016.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster === 2022-03-15 === * 20:58 wm-bot: Drained 'cloudvirt1026.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 20:36 wm-bot: Set cloudvirt 'cloudvirt1026.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 20:36 wm-bot: Setting cloudvirt 'cloudvirt1026.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 20:36 wm-bot: Draining 'cloudvirt1026.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 13:14 wm-bot: Unset cloudvirt 'cloudvirt1022.eqiad.wmnet' maintenance. - cookbook ran by arturo@nostromo * 13:14 wm-bot: Unsetting cloudvirt 'cloudvirt1022.eqiad.wmnet' maintenance. - cookbook ran by arturo@nostromo * 10:32 wm-bot: Set cloudvirt 'cloudvirt1022.eqiad.wmnet' maintenance. - cookbook ran by arturo@nostromo * 10:30 wm-bot: Setting cloudvirt 'cloudvirt1022.eqiad.wmnet' maintenance. - cookbook ran by arturo@nostromo === 2022-03-14 === * 21:24 wm-bot: Drained 'cloudvirt1025.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 20:59 wm-bot: Set cloudvirt 'cloudvirt1025.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 20:58 wm-bot: Setting cloudvirt 'cloudvirt1025.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 20:58 wm-bot: Draining 'cloudvirt1025.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 20:15 wm-bot: Setting cloudvirt 'cloudvirt1024.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 20:15 wm-bot: Draining 'cloudvirt1024.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 20:02 wm-bot: Set cloudvirt 'cloudvirt1024.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 19:59 wm-bot: Setting cloudvirt 'cloudvirt1024.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 19:59 wm-bot: Draining 'cloudvirt1024.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 19:16 wm-bot: Set cloudvirt 'cloudvirt1024.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 19:15 wm-bot: Setting cloudvirt 'cloudvirt1024.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 19:15 wm-bot: Draining 'cloudvirt1024.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 19:13 wm-bot: Drained 'cloudvirt1023.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 18:56 wm-bot: Set cloudvirt 'cloudvirt1023.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 18:55 wm-bot: Setting cloudvirt 'cloudvirt1023.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 18:55 wm-bot: Draining 'cloudvirt1023.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 18:53 wm-bot: Drained 'cloudvirt1022.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 18:52 wm-bot: Set cloudvirt 'cloudvirt1022.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 18:51 wm-bot: Setting cloudvirt 'cloudvirt1022.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 18:51 wm-bot: Draining 'cloudvirt1022.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 16:50 wm-bot: Drained 'cloudvirt1021.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 16:48 wm-bot: Set cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 16:48 wm-bot: Setting cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 16:48 wm-bot: Draining 'cloudvirt1021.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 11:48 dcaro: rebased cookbooks on latest master, make sure you pull before sending new patches === 2022-03-08 === * 18:29 wm-bot: Set cloudvirt 'cloudvirt1022.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 18:29 wm-bot: Setting cloudvirt 'cloudvirt1022.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 18:29 wm-bot: Draining 'cloudvirt1022.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 18:23 wm-bot: Set cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 18:21 wm-bot: Setting cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 18:21 wm-bot: Draining 'cloudvirt1021.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 18:18 wm-bot: Set cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 18:17 wm-bot: Setting cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 18:17 wm-bot: Draining 'cloudvirt1021.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 17:28 wm-bot: Set cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 17:27 wm-bot: Setting cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 17:27 wm-bot: Draining 'cloudvirt1021.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 17:18 wm-bot: Set cloudvirt 'cloudvirt1017.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 17:15 wm-bot: Setting cloudvirt 'cloudvirt1017.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 17:15 wm-bot: Draining 'cloudvirt1017.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 16:48 wm-bot: Set cloudvirt 'cloudvirt1017.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 16:47 wm-bot: Setting cloudvirt 'cloudvirt1017.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 16:47 wm-bot: Draining 'cloudvirt1017.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 16:36 wm-bot: Drained 'cloudvirt1016.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 16:08 wm-bot: Set cloudvirt 'cloudvirt1016.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 16:07 wm-bot: Setting cloudvirt 'cloudvirt1016.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 16:07 wm-bot: Draining 'cloudvirt1016.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 13:11 arturo: [codfw1dev] rebooting cloudservices servers for [[phab:T303179|T303179]] * 13:07 arturo: [codfw1dev] rebooting cloudvirt servers for [[phab:T303179|T303179]] * 13:06 arturo: [codfw1dev] rebooting cloudnet servers for [[phab:T303179|T303179]] * 12:55 arturo: [codfw1dev] rebooting cloudcontrol servers for [[phab:T303179|T303179]] === 2022-03-03 === * 08:49 taavi: deploying cloudmetrics grafana to grafana 8, [[phab:T282863|T282863]] === 2022-03-02 === * 09:06 arturo: merging core router firewall change https://gerrit.wikimedia.org/r/c/operations/homer/public/+/701347 === 2022-02-28 === * 15:30 dcaro: cleaning up leftover snapshots from failed backups of the maps volume ([[phab:T302720|T302720]]) === 2022-02-24 === * 17:04 andrewbogott: upgrading eqiad1 and codfw1dev to mariadb 10.5.15+maria~bullseye via 'apt-get install libmariadb3:amd64 galera-4 mariadb-server' * 15:42 dcaro: stopping and starting mariadb on cloudcontrol1003 ([[phab:T302146|T302146]]) * 10:37 arturo: [codfw1dev] briefly installed galera-4 (26.4.11+1bullseye) over (26.4.9-0+deb11u1) on cloudcontrol2001-dev and then downgrade again to verify package install ([[phab:T302482|T302482]]) === 2022-02-23 === * 20:39 taavi: added domain-wide 'designateadmin' and 'observer' roles to project-proxy-dns-manager service account [[phab:T295246|T295246]] * 17:40 andrewbogott: restarting lots of openstack services to try to clear up the mess that is [[phab:T236101|T236101]] * 12:13 arturo: cleaning up cinder volume snapshots, aborrero@cloudcontrol1005:~$ for i in $(sudo wmcs-openstack volume snapshot list -f value -c ID) ; do sudo wmcs-openstack volume snapshot delete $i ; done ([[phab:T302382|T302382]]) * 10:14 arturo: cleaning up neutron agents for non-existent servers cloudvirt100[1-9].eqiad.wmnet,cloudvirt10[12-15].eqiad.wmnet * 10:05 dcaro: Deleting stuck novafullstack servers, to let the service create new ones ([[phab:T302369|T302369]]) * 09:56 arturo: neutron agent-delete bad663b3-fd25-4393-a546-{{Gerrit|4b1b4bdec4db}} (Linux bridge agent {{!}} cloudvirtan1001) * 09:56 arturo: neutron agent-delete 1071c198-ed57-4b5a-9439-{{Gerrit|30e66a31aa69}} (Linux bridge agent {{!}} cloudvirtan1005) * 09:55 arturo: neutron agent-delete 2eeef198-8af7-4e5d-bd73-{{Gerrit|e14a2a8d2404}} (Linux bridge agent {{!}} cloudvirtan1004) * 09:55 arturo: neutron agent-delete afe173eb-35ba-444a-9960-{{Gerrit|899629786d2f}} (Linux bridge agent {{!}} cloudvirtan1003) * 09:54 arturo: neutron agent-delete afcb9b7f-c1a6-4ff4-9b10-{{Gerrit|92bfbe8d1a56}} (Linux bridge agent {{!}} cloudvirtan1002) * 09:39 dcaro: restarting neutron-api cloudcontrol1003 to see if the agent status update starts working ([[phab:T302369|T302369]]) * 09:38 dcaro: restarting neutron-dhcp-agent on cloudnet1003 ([[phab:T302369|T302369]]) === 2022-02-22 === * 22:10 andrewbogott: raising project 'maps' quota by two tb -- [[phab:T300160|T300160]] * 09:24 arturo: restarting mariadb @ cloudcontrol1003 ([[phab:T302146|T302146]]) * 09:13 arturo: restarting mariadb @ cloudcontrol1004 ([[phab:T302146|T302146]]) === 2022-02-18 === * 21:57 andrewbogott: leaving cloudcontrol1003 downtimed with disabled puppet for the weekend. Everything there should be stable and fine save rabbit which needs an upgrade. * 21:30 andrewbogott: rebooting cloudcontrol1003 because rabbit is freaking out * 17:25 andrewbogott: in-place upgrade of cloudcontrol1004 to bullseye -- [[phab:T281276|T281276]] * 12:34 arturo: manually install prometheus-openstack-exporter on cloudcontrol1005 ([[phab:T302050|T302050]]) === 2022-02-17 === * 23:02 andrewbogott: in-place upgrade to Bullseye on cloudcontrol1005 [[phab:T281276|T281276]] === 2022-02-15 === * 14:15 taavi: [codfw1dev] added domain-wide 'designateadmin' and 'observer' roles to codfw1dev-proxy-dns-manager service account [[phab:T295246|T295246]] === 2022-02-04 === * 10:12 arturo: restart backup_vms service in cloudvirt1024 ([[phab:T300956|T300956]]) === 2022-02-03 === * 08:21 taavi: cloudmetrics1004: manually added an empty line to /etc/prometheus/blackbox.yml to make /usr/local/bin/blackbox-exporter-assemble happy (clearing "performing a change every puppet run" alert) === 2022-02-02 === * 02:36 andrewbogott: restarting mariadb on cloudcontrol1004 === 2022-01-31 === * 10:15 arturo: cloudcontrol1005:~$ sudo systemctl restart backup_glance_images.service (failed state, no logs, icinga alert) === 2022-01-29 === * 18:24 taavi: delete 2 puppet prefixes in a weird state [[phab:T299750|T299750]] === 2022-01-27 === * 13:24 arturo: cloudmetrics1004:~ $ sudo systemctl restart wmcs_monitoring_graphite_rsync.service ([[phab:T300138|T300138]]) === 2022-01-26 === * 19:09 andrewbogott: bootstrapping a fresh galera node on cloudcontrol1004 * 18:57 andrewbogott: restarting mariadb on cloudcontrol1004 === 2022-01-25 === * 10:49 arturo: made cloudmetrics1001/1002 primary/backup respectively ([[phab:T299744|T299744]], [[phab:T297814|T297814]], [[phab:T300011|T300011]]) === 2022-01-19 === * 16:38 andrewbogott: moving all scratch mounts to scratch.svc.cloudinfra-nfs.eqiad1.wikimedia.cloud === 2022-01-05 === * 03:11 andrewbogott: 'cp /etc/apt/sources.list /etc/apt/sources.list.prepuppet' on all VMs. Backing up state before puppetizing sources.list with https://gerrit.wikimedia.org/r/c/operations/puppet/+/751498 === 2022-01-04 === * 12:44 dcaro: increasing the size_limit for labs ldap servers === 2021-12-26 === * 16:55 majavah: run attachLdapUser.php on wikitech for developer account "Karthiksripal" === 2021-12-24 === * 22:51 majavah: ran the wikireplica dns script on s5 [[phab:T298303|T298303]] === 2021-12-23 === * 21:42 majavah: deployed horizon wmf-proxy-dashboard update to fix editing of existing proxies === 2021-12-21 === * 10:39 arturo: dropped egress NAT exceptions for WMF apt repos, [[phab:T298042|T298042]] === 2021-12-15 === * 12:44 dcaro: Downtiming cloudvirt-wdqs1001 as it has no VMs running until disk space is fixed ([[phab:T297454|T297454]]) === 2021-12-14 === * 10:26 dcaro: Moved the nova cache (/var/lib/nova/instances/_base) and the canary image local data (/var/lib/nova/instance/<canary_image_id>) to the root disk on cloudvirt-wdqs1001 to temporary free some space ([[phab:T297454|T297454]]) === 2021-12-13 === * 18:08 wm-bot: Drained 'cloudvirt1014.eqiad.wmnet'. - cookbook ran by michael@mouse * 17:50 wm-bot: Set cloudvirt 'cloudvirt1014.eqiad.wmnet' maintenance. - cookbook ran by michael@mouse * 17:49 wm-bot: Setting cloudvirt 'cloudvirt1014.eqiad.wmnet' maintenance. - cookbook ran by michael@mouse * 17:49 wm-bot: Draining 'cloudvirt1014.eqiad.wmnet'. - cookbook ran by michael@mouse * 17:44 wm-bot: Drained 'cloudvirt1013.eqiad.wmnet'. - cookbook ran by michael@mouse * 17:30 wm-bot: Set cloudvirt 'cloudvirt1013.eqiad.wmnet' maintenance. - cookbook ran by michael@mouse * 17:30 wm-bot: Setting cloudvirt 'cloudvirt1013.eqiad.wmnet' maintenance. - cookbook ran by michael@mouse * 17:30 wm-bot: Draining 'cloudvirt1013.eqiad.wmnet'. - cookbook ran by michael@mouse * 17:13 wm-bot: Drained 'cloudvirt1012.eqiad.wmnet'. - cookbook ran by michael@mouse * 16:50 wm-bot: Set cloudvirt 'cloudvirt1012.eqiad.wmnet' maintenance. - cookbook ran by michael@mouse * 16:47 wm-bot: Setting cloudvirt 'cloudvirt1012.eqiad.wmnet' maintenance. - cookbook ran by michael@mouse * 16:47 wm-bot: Draining 'cloudvirt1012.eqiad.wmnet'. - cookbook ran by michael@mouse * 16:44 wm-bot: Set cloudvirt 'cloudvirt1012.eqiad.wmnet' maintenance. - cookbook ran by michael@mouse * 16:43 wm-bot: Setting cloudvirt 'cloudvirt1012.eqiad.wmnet' maintenance. - cookbook ran by michael@mouse === 2021-12-03 === * 18:56 andrewbogott: maintain-views and maintain-meta-p on clouddb1013-1020 * 10:49 majavah: deleting dbbackups-dashboard project [[phab:T296992|T296992]] === 2021-12-02 === * 01:17 wm-bot: Drained 'cloudvirt1028.eqiad.wmnet'. ([[phab:T296790|T296790]]) - cookbook ran by andrew@buster * 00:56 wm-bot: Set cloudvirt 'cloudvirt1028.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 00:56 wm-bot: Setting cloudvirt 'cloudvirt1028.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 00:56 wm-bot: Draining 'cloudvirt1028.eqiad.wmnet'. ([[phab:T296790|T296790]]) - cookbook ran by andrew@buster * 00:50 wm-bot: Set cloudvirt 'cloudvirt1026.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 00:50 wm-bot: Setting cloudvirt 'cloudvirt1026.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 00:50 wm-bot: Draining 'cloudvirt1026.eqiad.wmnet'. ([[phab:T296790|T296790]]) - cookbook ran by andrew@buster * 00:28 wm-bot: Drained 'cloudvirt1021.eqiad.wmnet'. ([[phab:T296790|T296790]]) - cookbook ran by andrew@buster * 00:03 wm-bot: Set cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 00:02 wm-bot: Setting cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 00:02 wm-bot: Draining 'cloudvirt1021.eqiad.wmnet'. ([[phab:T296790|T296790]]) - cookbook ran by andrew@buster === 2021-12-01 === * 23:59 wm-bot: Setting cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 23:59 wm-bot: Draining 'cloudvirt1021.eqiad.wmnet'. ([[phab:T296790|T296790]]) - cookbook ran by andrew@buster * 23:54 andrewbogott: *correction* adding spare cloudvirts 1044 and 1045 to the 'ceph' pool in order to make space for future juggling around [[phab:T296790|T296790]] and [[phab:T296792|T296792]] * 23:53 andrewbogott: adding spare cloudvirts 1044 and 1055 to the 'ceph' pool in order to make space for future juggling around [[phab:T296790|T296790]] and [[phab:T296792|T296792]] === 2021-11-28 === * 17:48 andrewbogott: moved cloudvirt1018 out of the 'localstorage' aggregate and into 'maintenance' for [[phab:T296592|T296592]]. It will need to be moved back after the raid is rebuilt. === 2021-11-21 === * 07:19 dcaro_away: restarting designate-sink with some extra logs in it ([[phab:T296144|T296144]]) === 2021-11-17 === * 15:48 andrewbogott: upgrading mariadb packages on eqiad1 cloudcontrols * 15:39 andrewbogott: sudo cumin "cloud*" 'apt-get update -y --allow-releaseinfo-change' * 15:26 andrewbogott: updated mariadb packages on codfw1dev cloudcontrols to 1:10.3.31-0+deb10u1 === 2021-11-12 === * 13:31 arturo: restarting glance-api services to make sure they work with new ceph auth creds ([[phab:T293752|T293752]]) === 2021-11-08 === * 21:50 andrewbogott: returned clouddb pools back to normal after maintain_views run: https://gerrit.wikimedia.org/r/c/operations/puppet/+/737505 [[phab:T216481|T216481]] * 20:07 andrewbogott: depooling clouddb1013 for maintain_views attempt * 10:54 arturo: [codfw1dev] create service account `srv-networktests` following https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Service_accounts for [[phab:T294955|T294955]] * 10:34 arturo: create service account `srv-networktests` following https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Service_accounts for [[phab:T294955|T294955]] === 2021-11-05 === * 11:18 wm-bot: Added 1 new OSDs ['cloudcephosd1024.eqiad.wmnet'] ([[phab:T295012|T295012]]) - cookbook ran by arturo@endurance * 11:17 wm-bot: Added OSD cloudcephosd1024.eqiad.wmnet... (1/1) ([[phab:T295012|T295012]]) - cookbook ran by arturo@endurance * 11:15 wm-bot: Finished rebooting node cloudcephosd1024.eqiad.wmnet - cookbook ran by arturo@endurance * 11:12 wm-bot: Rebooting node cloudcephosd1024.eqiad.wmnet - cookbook ran by arturo@endurance * 11:12 wm-bot: Adding OSD cloudcephosd1024.eqiad.wmnet... (1/1) ([[phab:T295012|T295012]]) - cookbook ran by arturo@endurance * 11:12 wm-bot: Adding new OSDs ['cloudcephosd1024.eqiad.wmnet'] to the cluster ([[phab:T295012|T295012]]) - cookbook ran by arturo@endurance === 2021-11-04 === * 16:39 wm-bot: Added 1 new OSDs ['cloudcephosd1023.eqiad.wmnet'] ([[phab:T295012|T295012]]) - cookbook ran by arturo@endurance * 16:39 wm-bot: Added OSD cloudcephosd1023.eqiad.wmnet... (1/1) ([[phab:T295012|T295012]]) - cookbook ran by arturo@endurance * 16:37 wm-bot: Finished rebooting node cloudcephosd1023.eqiad.wmnet - cookbook ran by arturo@endurance * 16:34 wm-bot: Rebooting node cloudcephosd1023.eqiad.wmnet - cookbook ran by arturo@endurance * 16:33 wm-bot: Adding OSD cloudcephosd1023.eqiad.wmnet... (1/1) ([[phab:T295012|T295012]]) - cookbook ran by arturo@endurance * 16:33 wm-bot: Adding new OSDs ['cloudcephosd1023.eqiad.wmnet'] to the cluster ([[phab:T295012|T295012]]) - cookbook ran by arturo@endurance * 16:17 wm-bot: Added 1 new OSDs ['cloudcephosd1022.eqiad.wmnet'] ([[phab:T295012|T295012]]) - cookbook ran by arturo@endurance * 16:17 wm-bot: Added OSD cloudcephosd1022.eqiad.wmnet... (1/1) ([[phab:T295012|T295012]]) - cookbook ran by arturo@endurance * 16:16 wm-bot: Finished rebooting node cloudcephosd1022.eqiad.wmnet - cookbook ran by arturo@endurance * 16:13 wm-bot: Rebooting node cloudcephosd1022.eqiad.wmnet - cookbook ran by arturo@endurance * 16:12 wm-bot: Adding OSD cloudcephosd1022.eqiad.wmnet... (1/1) ([[phab:T295012|T295012]]) - cookbook ran by arturo@endurance * 16:12 wm-bot: Adding new OSDs ['cloudcephosd1022.eqiad.wmnet'] to the cluster ([[phab:T295012|T295012]]) - cookbook ran by arturo@endurance * 16:00 wm-bot: Adding OSD cloudcephosd1022.eqiad.wmnet... (1/1) ([[phab:T295012|T295012]]) - cookbook ran by arturo@endurance * 16:00 wm-bot: Adding new OSDs ['cloudcephosd1022.eqiad.wmnet'] to the cluster ([[phab:T295012|T295012]]) - cookbook ran by arturo@endurance * 11:26 wm-bot: Added 1 new OSDs ['cloudcephosd1021.eqiad.wmnet'] ([[phab:T295012|T295012]]) - cookbook ran by arturo@endurance * 11:26 wm-bot: Added OSD cloudcephosd1021.eqiad.wmnet... (1/1) ([[phab:T295012|T295012]]) - cookbook ran by arturo@endurance * 11:23 wm-bot: Finished rebooting node cloudcephosd1021.eqiad.wmnet - cookbook ran by arturo@endurance * 11:20 wm-bot: Rebooting node cloudcephosd1021.eqiad.wmnet - cookbook ran by arturo@endurance * 11:19 wm-bot: Adding OSD cloudcephosd1021.eqiad.wmnet... (1/1) ([[phab:T295012|T295012]]) - cookbook ran by arturo@endurance * 11:19 wm-bot: Adding new OSDs ['cloudcephosd1021.eqiad.wmnet'] to the cluster ([[phab:T295012|T295012]]) - cookbook ran by arturo@endurance * 11:16 wm-bot: Adding new OSDs ['cloudcephosd1021.eqiad.wmnet'] to the cluster ([[phab:T295012|T295012]]) - cookbook ran by arturo@endurance === 2021-11-03 === * 17:22 arturo: [codfw1dev] installing keepalived 2.1.5 from buster-backports on cloudgw2001-dev/2002-dev ([[phab:T294956|T294956]]) * 11:45 arturo: [codfw1dev] downgrade kernel on cloudgw2001-dev/2002-dev ([[phab:T294853|T294853]], [[phab:T291813|T291813]]) === 2021-11-02 === * 10:54 arturo: rebooting cloudnet1004/1003 for [[phab:T291813|T291813]] * 10:43 arturo: [codfw1dev] rebooting cloudgw200[12]-dev for [[phab:T291813|T291813]] === 2021-10-24 === * 00:47 andrewbogott: deploying a change so that openstack clients use tls endpoints: https://gerrit.wikimedia.org/r/c/operations/puppet/+/732738 === 2021-10-21 === * 10:19 arturo: drop firewall exception on core routers for wiki replicas legacy setup ([[phab:T293897|T293897]]) * 10:12 arturo: drop NAT exception for wiki replicas legacy setup ([[phab:T293897|T293897]]) === 2021-10-20 === * 21:06 andrewbogott: creating cloudinfra-nfs project [[phab:T293936|T293936]] === 2021-10-18 === * 19:21 andrewbogott: also ticked the 'admin' box on wikitech for majavah [[phab:T292827|T292827]] * 18:58 andrewbogott: granting majavah 'admin' role in the 'admin' project and also in the default domain. [[phab:T292827|T292827]] === 2021-10-14 === * 12:28 arturo: [codfw1dev] add DB grants for cloudbackup2002.codfw.wmnet IP address to the cinder DB ([[phab:T292546|T292546]]) === 2021-10-13 === * 10:46 arturo: updating python3-neutron across the fleet ([[phab:T292936|T292936]]) === 2021-10-12 === * 09:06 dcaro: upgrading eqiad cloudnet hosts neutron packages ([[phab:T292936|T292936]]) * 08:57 dcaro: upgrading codfw cloudnet hosts neutron packages ([[phab:T292936|T292936]]) === 2021-10-05 === * 09:39 arturo: [codfw1dev] cleaning up manila stuff from openstack (db, endpoints, tenant, VMs, and such) [[phab:T291257|T291257]] === 2021-09-30 === * 14:50 andrewbogott: sudo cumin "cloud*" "ps -ef {{!}} grep nslcd && service nslcd restart" and sudo cumin "lab*" "ps -ef {{!}} grep nslcd && service nslcd restart" [[phab:T292202|T292202]] * 14:43 andrewbogott: ran sudo cumin --force --timeout 500 -o json "A:all" "ps -ef {{!}} grep nslcd && service nslcd restart" to get nslcd happy again [[phab:T292202|T292202]] === 2021-09-29 === * 09:41 arturo: [codfw1dev] cleanup manila shares definitions for a clean start now that the manila-sharecontroller VM is apparently well configured ([[phab:T291257|T291257]]) === 2021-09-28 === * 16:23 bstorm: downtime for clouddb1020 to reduce re-pages in case this goes badly [[phab:T291963|T291963]] * 16:21 bstorm: powering on clouddb1020 via remote console [[phab:T291963|T291963]] * 15:58 bstorm: depooled clouddb1020 for repair [[phab:T291961|T291961]] * 12:40 dcaro: Merged change on sssd for bullseye cloud hosts ([[phab:T291585|T291585]]) * 11:30 arturo: [codfw1dev] create floating IP 185.15.57.5 for manila-sharecontroller.cloudinfra-codfw1dev.codfw1dev.wmcloud.org ([[phab:T291257|T291257]]) === 2021-09-27 === * 10:07 arturo: cloudcontrol1004 apparently healthy [[phab:T291446|T291446]] * 09:25 arturo: rebooting cloudcontrol1004 for [[phab:T291446|T291446]] === 2021-09-24 === * 13:02 arturo: [codfw1dev] create VM manila-share-controller-01 on cloudinfra-codfw1dev * 13:00 arturo: [codfw1dev] rebase labs/private.git on cloudinfra-puppetmaster-01, had merge conflict === 2021-09-21 === * 12:13 arturo: [codfw1dev] trying to create a manila service image ([[phab:T291257|T291257]]) * 11:45 arturo: [codfw1dev] created rabbitmq user ([[phab:T291257|T291257]]) * 11:32 arturo: [codfw1dev] populated manila DB & created service endpoints ([[phab:T291257|T291257]]) * 11:06 arturo: [codfw1dev] give manila user admin role @ manila project ([[phab:T291257|T291257]]) * 11:06 arturo: [codfw1dev] created manila project ([[phab:T291257|T291257]]) * 10:57 arturo: [codfw1dev] created manila user @ labtestwikitech ([[phab:T291257|T291257]]) * 10:49 arturo: [codfw1dev] create manila database on cloudcontrol-dev nodes (galera) [[phab:T291257|T291257]] === 2021-09-20 === * 23:08 bstorm: ran `echo check > /sys/block/md0/md/sync_action` on cloudcontrol1004 to check raid * 22:48 andrewbogott: stopped puppet & mariadb on cloudcontrol1004; it was flapping * 22:44 andrewbogott: sudo touch /tmp/galera.disabled on cloudcontrol1004, the service seems troubled there * 21:57 andrewbogott: moving cloudvirt1043 into the 'nfs' aggregate for [[phab:T291405|T291405]] === 2021-09-17 === * 11:35 arturo: [codfw1dev] install manila on cloudcontrol2001-dev ([[phab:T291257|T291257]]) === 2021-09-16 === * 15:56 bstorm: removing downtime for labstore1005 so we'll know if it has another issue [[phab:T290318|T290318]] === 2021-09-09 === * 22:03 bstorm: restarted the prometheus-mysqld-exporter@s1 service as it was not working [[phab:T290630|T290630]] * 03:15 bstorm: resetting swap on clouddb1017 [[phab:T290630|T290630]] * 03:08 andrewbogott: stopping maintain-dbusers on labstore1004 for help diagnosing [[phab:T290630|T290630]] === 2021-09-03 === * 15:34 bstorm: rebooting labstore1005 to disconnect the drives from labstore1004 [[phab:T290318|T290318]] * 15:24 bstorm: stopping puppet and disabling backup syncs to labstore1005 on cloudbackup2002 [[phab:T290318|T290318]] * 15:20 bstorm: stopping puppet and disabling backup syncs to labstore1005 on cloudbackup2001 [[phab:T290318|T290318]] === 2021-08-30 === * 16:16 wm-bot: Added 1 new OSDs ['cloudcephosd1018.eqiad.wmnet'] - cookbook ran by andrew@buster * 16:16 wm-bot: Added OSD cloudcephosd1018.eqiad.wmnet... (1/1) - cookbook ran by andrew@buster * 16:13 wm-bot: Adding OSD cloudcephosd1018.eqiad.wmnet... (1/1) - cookbook ran by andrew@buster * 16:13 wm-bot: Adding new OSDs ['cloudcephosd1018.eqiad.wmnet'] to the cluster - cookbook ran by andrew@buster * 16:10 wm-bot: Finished rebooting node cloudcephosd1018.eqiad.wmnet - cookbook ran by andrew@buster * 16:07 wm-bot: Rebooting node cloudcephosd1018.eqiad.wmnet - cookbook ran by andrew@buster * 16:07 wm-bot: Adding OSD cloudcephosd1018.eqiad.wmnet... (1/1) - cookbook ran by andrew@buster * 16:07 wm-bot: Adding new OSDs ['cloudcephosd1018.eqiad.wmnet'] to the cluster - cookbook ran by andrew@buster === 2021-08-27 === * 18:57 andrewbogott: raising toolsbeta ram/core/instances quotas so majavah can experiment with bullseye === 2021-08-25 === * 14:45 wm-bot: Finished rebooting node cloudcephosd1018.eqiad.wmnet - cookbook ran by andrew@buster * 14:42 wm-bot: Rebooting node cloudcephosd1018.eqiad.wmnet - cookbook ran by andrew@buster * 14:42 wm-bot: Adding OSD cloudcephosd1018.eqiad.wmnet... (1/1) - cookbook ran by andrew@buster * 14:42 wm-bot: Adding new OSDs ['cloudcephosd1018.eqiad.wmnet'] to the cluster - cookbook ran by andrew@buster * 14:41 wm-bot: Adding new OSDs ['cloudcephosd1018.eqiad.wmnet'] to the cluster - cookbook ran by andrew@buster === 2021-08-19 === * 17:39 bstorm: restarting glance image backup to try and clear the page === 2021-08-18 === * 16:21 wm-bot: Rebooting node cloudcephosd1018.eqiad.wmnet - cookbook ran by andrew@buster * 16:21 wm-bot: Adding OSD cloudcephosd1018.eqiad.wmnet... (1/1) - cookbook ran by andrew@buster * 16:21 wm-bot: Adding new OSDs ['cloudcephosd1018.eqiad.wmnet'] to the cluster - cookbook ran by andrew@buster * 16:17 wm-bot: Adding new OSDs ['cloudcephosd1018.eqiad.wmnet'] to the cluster - cookbook ran by andrew@buster * 16:16 wm-bot: Adding new OSDs ['cloudcephosd1018.eqiad.wmnet'] to the cluster - cookbook ran by andrew@buster * 16:15 wm-bot: Adding new OSDs ['cloudcephosd1018.eqiad.wmnet'] to the cluster - cookbook ran by andrew@buster * 16:13 wm-bot: Adding new OSDs ['cloudcephosd1018.eqiad.wmnet'] to the cluster - cookbook ran by andrew@buster * 14:47 andrewbogott: adding clouvirt1038 to the ceph aggregate, removing from the maintenance aggregate [[phab:T276922|T276922]] === 2021-08-17 === * 15:11 andrewbogott: rebooting cloudcephosd1008 to force raid rebuild -- [[phab:T287838|T287838]] === 2021-08-11 === * 13:51 wm-bot: Finished rebooting node cloudcephosd1018.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 13:48 wm-bot: Rebooting node cloudcephosd1018.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 13:47 wm-bot: Adding OSD cloudcephosd1018.eqiad.wmnet... (1/1) ([[phab:T285858|T285858]]) - cookbook ran by dcaro@vulcanus * 13:47 wm-bot: Adding new OSDs ['cloudcephosd1018.eqiad.wmnet'] to the cluster ([[phab:T285858|T285858]]) - cookbook ran by dcaro@vulcanus === 2021-08-10 === * 15:15 andrewbogott: restarting all designate services in eqiad1 * 15:04 andrewbogott: restarting designate-sink in eqiad1; it's complaining about rabbit but I don't want to restart rabbit yet === 2021-08-05 === * 09:37 dcaro: Taking one osd daemon down ot codfw cluster ([[phab:T288203|T288203]]) === 2021-08-04 === * 19:20 bd808: Running deleteBatch.php on cloudweb2001-dev to remove legacy Heira: pages from labtestwiki === 2021-08-03 === * 17:40 bstorm: rerunning the glance backup script after failure === 2021-07-31 === * 00:10 andrewbogott: "systemctl reset-failed cloud-init.service" on all VMs for [[phab:T287309|T287309]] * 00:08 andrewbogott: "systemctl reset-failed cloud-final.service" on all VMs for [[phab:T287309|T287309]] === 2021-07-27 === * 21:32 andrewbogott: putting cloudvirt1012 back into service [[phab:T286748|T286748]] * 20:52 andrewbogott: draining VMs off of cloudvirt1012 so we can replace the battery for [[phab:T286748|T286748]] * 15:15 andrewbogott: "rm /etc/apt/sources.list.d/openstack-mitaka-jessie.list" cloud-wide === 2021-07-23 === * 15:22 bstorm: update wikireplicas-dns for s7 fix for web replicas === 2021-07-20 === * 17:07 andrewbogott: reloading haproxy on dbproxy1018 for [[phab:T286598|T286598]] * 15:45 arturo: failback from labstore1006 to labstore1007 (dumps NFS) https://gerrit.wikimedia.org/r/c/operations/puppet/+/705417 * 00:10 bstorm: restarting nova-api on cloudcontrol1003 to try and recover whatever it's doing with designate_floating_ip_ptr_records_updater === 2021-07-19 === * 22:05 bstorm: set downtime scheduled for tomorrow from 1300 to 1600 UTC for cloudstore1008 and 1009 [[phab:T286599|T286599]] * 20:40 andrewbogott: reloading haproxy on dbproxy1018 for [[phab:T286598|T286598]] * 13:50 andrewbogott: upgrading mariadb to 10.3.29 on all cloudcontrols === 2021-07-16 === * 09:55 dcaro: checking HP raid issues on coludvirt1012 ([[phab:T286766|T286766]]) === 2021-07-14 === * 21:08 andrewbogott: restarting lots of openstack services while trying to resolve [[phab:T286675|T286675]] * 12:17 dcaro: doing ceph outage tests on codfw1 (fyi) === 2021-07-13 === * 10:57 dcaro: enabled autoscaling on codfw1 ceph cluster, setting a minimum of pgs on codfw1dev-compute to 128 === 2021-07-02 === * 10:12 wm-bot: The cluster is not rebalance after adding the new OSDs ['cloudcephosd1019.eqiad.wmnet', 'cloudcephosd1020.eqiad.wmnet'] ([[phab:T285858|T285858]]) - cookbook ran by dcaro@vulcanus * 10:12 wm-bot: Added 2 new OSDs ['cloudcephosd1019.eqiad.wmnet', 'cloudcephosd1020.eqiad.wmnet'] ([[phab:T285858|T285858]]) - cookbook ran by dcaro@vulcanus * 10:12 wm-bot: Added OSD cloudcephosd1020.eqiad.wmnet... (2/2) ([[phab:T285858|T285858]]) - cookbook ran by dcaro@vulcanus * 10:10 wm-bot: Finished rebooting node cloudcephosd1020.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 10:07 wm-bot: Rebooting node cloudcephosd1020.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 10:07 wm-bot: Adding OSD cloudcephosd1020.eqiad.wmnet... (2/2) ([[phab:T285858|T285858]]) - cookbook ran by dcaro@vulcanus * 10:07 wm-bot: Added OSD cloudcephosd1019.eqiad.wmnet... (1/2) ([[phab:T285858|T285858]]) - cookbook ran by dcaro@vulcanus * 10:05 wm-bot: Finished rebooting node cloudcephosd1019.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 10:02 wm-bot: Rebooting node cloudcephosd1019.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 10:02 wm-bot: Adding OSD cloudcephosd1019.eqiad.wmnet... (1/2) ([[phab:T285858|T285858]]) - cookbook ran by dcaro@vulcanus * 10:01 wm-bot: Adding new OSDs ['cloudcephosd1019.eqiad.wmnet', 'cloudcephosd1020.eqiad.wmnet'] to the cluster ([[phab:T285858|T285858]]) - cookbook ran by dcaro@vulcanus * 09:13 wm-bot: Adding OSD cloudcephosd1019.eqiad.wmnet... (1/2) ([[phab:T285858|T285858]]) - cookbook ran by dcaro@vulcanus * 09:13 wm-bot: Adding new OSDs ['cloudcephosd1019.eqiad.wmnet', 'cloudcephosd1020.eqiad.wmnet'] to the cluster ([[phab:T285858|T285858]]) - cookbook ran by dcaro@vulcanus === 2021-07-01 === * 16:27 bstorm: failed over cloudstore1009 to cloudstore1008 [[phab:T224747|T224747]] * 16:18 bstorm: downtimed cloudstore1008 and cloudstore1009 to fail over [[phab:T224747|T224747]] * 14:25 wm-bot: Adding OSD cloudcephosd1019.eqiad.wmnet... (2/3) ([[phab:T285858|T285858]]) - cookbook ran by dcaro@vulcanus * 14:25 wm-bot: Added OSD cloudcephosd1017.eqiad.wmnet... (1/3) ([[phab:T285858|T285858]]) - cookbook ran by dcaro@vulcanus * 14:24 wm-bot: Finished rebooting node cloudcephosd1017.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 14:21 wm-bot: Rebooting node cloudcephosd1017.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 14:20 wm-bot: Adding OSD cloudcephosd1017.eqiad.wmnet... (1/3) ([[phab:T285858|T285858]]) - cookbook ran by dcaro@vulcanus * 14:20 wm-bot: Adding new OSDs ['cloudcephosd1017.eqiad.wmnet', 'cloudcephosd1019.eqiad.wmnet', 'cloudcephosd1020.eqiad.wmnet'] to the cluster ([[phab:T285858|T285858]]) - cookbook ran by dcaro@vulcanus * 14:18 wm-bot: Rebooting node cloudcephosd1017.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 14:17 wm-bot: Adding OSD cloudcephosd1017.eqiad.wmnet... (1/3) ([[phab:T285858|T285858]]) - cookbook ran by dcaro@vulcanus * 14:17 wm-bot: Adding new OSDs ['cloudcephosd1017.eqiad.wmnet', 'cloudcephosd1019.eqiad.wmnet', 'cloudcephosd1020.eqiad.wmnet'] to the cluster ([[phab:T285858|T285858]]) - cookbook ran by dcaro@vulcanus * 11:16 wm-bot: Added new OSD node cloudcephosd1016.eqiad.wmnet ([[phab:T285858|T285858]]) - cookbook ran by dcaro@vulcanus * 11:13 wm-bot: Adding new OSD cloudcephosd1016.eqiad.wmnet to the cluster ([[phab:T285858|T285858]]) - cookbook ran by dcaro@vulcanus * 10:58 dcaro: rebooting cloudcephosd1016 ([[phab:T285858|T285858]]) * 10:47 wm-bot: Adding new OSD cloudcephosd1016.eqiad.wmnet to the cluster ([[phab:T285858|T285858]]) - cookbook ran by dcaro@vulcanus * 10:44 wm-bot: Adding new OSD cloudcephosd1016.eqiad.wmnet to the cluster ([[phab:T285858|T285858]]) - cookbook ran by dcaro@vulcanus * 10:42 wm-bot: Adding new OSD cloudcephosd1016.eqiad.wmnet to the cluster ([[phab:T285858|T285858]]) - cookbook ran by dcaro@vulcanus * 10:41 wm-bot: Adding new OSD cloudcephosd1016.eqiad.wmnet to the cluster ([[phab:T285858|T285858]]) - cookbook ran by dcaro@vulcanus * 10:40 wm-bot: Adding new OSD cloudcephosd1016.eqiad.wmnet to the cluster ([[phab:T285858|T285858]]) - cookbook ran by dcaro@vulcanus === 2021-06-30 === * 21:48 bstorm: downtimed space alerts for scratch on cloudstore1008 until after the migration === 2021-06-25 === * 15:28 andrewbogott: restarting openstack services on cloudcontrol1005 * 09:16 arturo: icinga downtime cloudcontrols for 2h * 08:20 dcaro: restarting rabbitmq on cloudcontrol100<nowiki>{</nowiki>3,4<nowiki>}</nowiki> === 2021-06-21 === * 13:54 dcaro: puppet fix merged and deployed, servers are back to normal * 13:20 dcaro: merged broken puppet patch, downtimed all cloudvirts for 2h while fixing (nothing big, just added a bad systemd timer) === 2021-06-20 === * 22:21 andrewbogott: clearing admin-monitoring VMs; puppet has been failing lately due to a full drive on the puppetmaster === 2021-06-15 === * 01:18 bstorm: running a modified version of the prometheus dir size cron in screen [[phab:T284964|T284964]] === 2021-06-14 === * 10:13 dcaro: setting ssd to debug mode on tools-sgeexec-0917 ([[phab:T284130|T284130]]) === 2021-06-10 === * 10:58 wm-bot: Finished rebooting the nodes ['cloudcephmon2002-dev', 'cloudcephmon2003-dev', 'cloudcephmon2004-dev'] ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 10:58 wm-bot: Finished rebooting node cloudcephmon2004-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 10:55 wm-bot: Rebooting node cloudcephmon2004-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 10:55 wm-bot: Finished rebooting node cloudcephmon2003-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 10:52 wm-bot: Rebooting node cloudcephmon2003-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 10:52 wm-bot: Finished rebooting node cloudcephmon2002-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 10:49 wm-bot: Rebooting node cloudcephmon2002-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 10:49 wm-bot: Rebooting the nodes cloudcephmon2002-dev,cloudcephmon2003-dev,cloudcephmon2004-dev ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 10:48 wm-bot: Finished rebooting the nodes ['cloudcephosd2001-dev', 'cloudcephosd2002-dev', 'cloudcephosd2003-dev'] ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 10:48 wm-bot: Finished rebooting node cloudcephosd2003-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 10:45 wm-bot: Rebooting node cloudcephosd2003-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 10:45 wm-bot: Finished rebooting node cloudcephosd2002-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 10:42 wm-bot: Rebooting node cloudcephosd2002-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 10:42 wm-bot: Finished rebooting node cloudcephosd2001-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 10:39 wm-bot: Rebooting node cloudcephosd2001-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 10:39 wm-bot: Rebooting the nodes cloudcephosd2001-dev,cloudcephosd2002-dev,cloudcephosd2003-dev ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 09:39 wm-bot: Finished rebooting the nodes ['cloudcephosd2001-dev', 'cloudcephosd2002-dev', 'cloudcephosd2003-dev'] ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 09:38 wm-bot: Finished rebooting node cloudcephosd2003-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 09:35 wm-bot: Rebooting node cloudcephosd2003-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 09:35 wm-bot: Finished rebooting node cloudcephosd2002-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 09:32 wm-bot: Rebooting node cloudcephosd2002-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 09:32 wm-bot: Finished rebooting node cloudcephosd2001-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 09:29 wm-bot: Rebooting node cloudcephosd2001-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 09:29 wm-bot: Rebooting the nodes cloudcephosd2001-dev,cloudcephosd2002-dev,cloudcephosd2003-dev ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 09:26 wm-bot: Rebooting node cloudcephosd2001-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 09:26 wm-bot: Rebooting the nodes cloudcephosd2001-dev,cloudcephosd2002-dev,cloudcephosd2003-dev ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 09:24 wm-bot: Rebooting node cloudcephosd2001-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 09:24 wm-bot: Rebooting the nodes cloudcephosd2001-dev,cloudcephosd2002-dev,cloudcephosd2003-dev ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus === 2021-06-09 === * 17:33 arturo: removed icinga downtime for cloudmetrics1002 -- to see if hardware is healthy ([[phab:T281881|T281881]]) * 13:30 wm-bot: Finished rebooting the nodes ['cloudcephmon2002-dev', 'cloudcephmon2003-dev', 'cloudcephmon2004-dev'] ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 13:30 wm-bot: Finished rebooting node cloudcephmon2004-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 13:27 wm-bot: Rebooting node cloudcephmon2004-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 13:27 wm-bot: Finished rebooting node cloudcephmon2003-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 13:24 wm-bot: Rebooting node cloudcephmon2003-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 13:24 wm-bot: Finished rebooting node cloudcephmon2002-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 13:21 wm-bot: Rebooting node cloudcephmon2002-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 13:21 wm-bot: Rebooting the nodes cloudcephmon2002-dev,cloudcephmon2003-dev,cloudcephmon2004-dev ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 13:01 wm-bot: Rebooting node cloudcephmon2002-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 13:01 wm-bot: Rebooting the nodes cloudcephmon2002-dev,cloudcephmon2003-dev,cloudcephmon2004-dev ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 12:53 wm-bot: Rebooting node cloudcephmon2002-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 12:53 wm-bot: Rebooting the nodes cloudcephmon2002-dev,cloudcephmon2003-dev,cloudcephmon2004-dev ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus === 2021-06-08 === * 23:19 bd808: Downtimed cloudmetrics1002 in icinga until 2021-06-30 23:59:01 ([[phab:T281881|T281881]]) * 21:08 bstorm: downtiming grafana-labs for maintenance * 16:28 wm-bot: Finished rebooting the nodes ['cloudcephosd2001-dev', 'cloudcephosd2002-dev', 'cloudcephosd2003-dev'] ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 16:27 wm-bot: Finished rebooting node cloudcephosd2003-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 16:24 wm-bot: Rebooting node cloudcephosd2003-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 16:24 wm-bot: Finished rebooting node cloudcephosd2002-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 16:22 wm-bot: Rebooting node cloudcephosd2002-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 16:21 wm-bot: Finished rebooting node cloudcephosd2001-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 16:18 wm-bot: Rebooting node cloudcephosd2001-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 16:18 wm-bot: Rebooting the nodes ['cloudcephosd2001-dev', 'cloudcephosd2002-dev', 'cloudcephosd2003-dev'] ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 16:17 wm-bot: Rebooting the nodes ['cloudcephosd2001-dev', 'cloudcephosd2002-dev', 'cloudcephosd2003-dev'] ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 15:03 wm-bot: Finished rebooting node cloudcephosd2001-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 14:59 wm-bot: Rebooting node cloudcephosd2001-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 14:59 wm-bot: Rebooting node cloudcephosd2001-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 14:57 wm-bot: Rebooting node cloudcephosd2001-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 14:57 wm-bot: Rebooting node cloudcephosd2001-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 14:29 wm-bot: Rebooting node cloudcephosd2001-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 14:23 wm-bot: Rebooting node cloudcephosd2001-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 14:18 wm-bot: Rebooting node cloudcephosd2001-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus === 2021-06-07 === * 14:27 andrewbogott: moving cloudvirt1040 from 'maintenance' aggregate to 'ceph' aggregate [[phab:T281399|T281399]] === 2021-06-01 === * 13:12 dcaro: Changed the ceph osd_memory_target on eqiad pool to 6Gi (we were reaching the limit, swapping at some points) * 09:57 arturo: fix PTR record for 185.15.56.1 ([[phab:T284025|T284025]]) * 09:56 arturo: fix PTR record for 185.15.56.1 ([[phab:T248025|T248025]]) === 2021-05-27 === * 14:58 wm-bot: Testing - cookbook ran by dcaro@vulcanus === 2021-05-26 === * 19:10 andrewbogott: reimaging cloudvirt1018 to support local VM storage * 18:07 andrewbogott: draining cloudvirt1018, converting it to a local-storage host like cloudvirt1019 and 1020 -- [[phab:T283296|T283296]] * 14:36 dcaro: Enabled syslog logging for osd.55 on eqiad ceph cluster for testing ([[phab:T281247|T281247]]) * 14:36 dcaro: Enabled syslog logging on codfw ceph cluster (mon/osd/mgr) ([[phab:T281247|T281247]]) * 11:26 arturo: [codfw1dev] purge old kernel packages in cloudvirt200[12]-dev * 11:03 arturo: created public flavor `g3.cores16.ram36.disk20` (even though it was requested as private in [[phab:T283293|T283293]], but may be useful for others) === 2021-05-25 === * 16:14 bd808: Closed #wikimedia-cloud-admin on f***node * 16:11 bd808: Closed #wikimedia-cloud-feed on f***node * 15:19 dcaro: rebooted cloudvirt1020, starting VMs ([[phab:T275893|T275893]]) * 15:13 dcaro: rebooting cloudvirt1020 ([[phab:T275893|T275893]]) * 14:42 dcaro: taking cloudvirt1020 out for maintenance (openstack wise) so no new VMs are scheduled on it ([[phab:T275893|T275893]]) === 2021-05-24 === * 22:32 andrewbogott: changing the default ttl for eqiad1.wikimedia.cloud. from 3600 to 60; this should help us avoid madness when re-using hostnames. * 11:20 arturo: created `g3.cores2.ram80.disk40.private` for the wmf-research-tools project, to allow resizing a 40G disk instance === 2021-05-22 === * 02:14 bstorm: downtiming SMART alerts on dumps server labstore1007 for the weekend because it has been flapping [[phab:T281045|T281045]] === 2021-05-13 === * 21:25 bstorm: converted the maps and scratch volumes on cloudstore1008 (standby) to drbd [[phab:T224747|T224747]] * 15:45 bstorm: re-running wikireplicas-dns after refactor of config to make sure it doesn't change anything === 2021-05-12 === * 14:23 arturo: [codfw1dev] cleanup old unused agents (bgp, ovs) * 11:37 arturo: [codfw1dev] replacing cloudnet2003-dev with cloudnet2004-dev ([[phab:T281381|T281381]]) === 2021-05-11 === * 18:00 andrewbogott: adding 'trove' service project in advance of deploying trove in eqiad1 * 10:22 arturo: rebooted cloudgw1002 (active) thus causing a failover to cloudgw1001 === 2021-05-09 === * 10:53 arturo: icinga-downtime cloudmetrics1002 for 3 months ([[phab:T275605|T275605]]) === 2021-05-07 === * 13:51 andrewbogott: add inherited 'admin' right to novaadmin user throughout eqiad1. I was trying to narrow down the rights here but lack of admin breaks some workflows, e.g. [[phab:T281894|T281894]] and [[phab:T282235|T282235]] === 2021-05-06 === * 15:31 arturo: about to migrating CloudVPS network to the cloudgw architecture [[phab:T270704|T270704]] * 11:14 dcaro: restarting cinder-volume on the eqiad control nodes to refresh the ceph libraries ([[phab:T282109|T282109]]) === 2021-05-05 === * 16:07 dcaro: disallowing insecure global ids on the eqiad ceph cluster ([[phab:T280641|T280641]]) * 15:15 wm-bot: Safe reboot of 'cloudvirt1046.eqiad.wmnet' finished successfully. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 15:11 wm-bot: Safe rebooting 'cloudvirt1046.eqiad.wmnet'. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 15:11 wm-bot: Safe reboot of 'cloudvirt1045.eqiad.wmnet' finished successfully. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 15:07 wm-bot: Safe rebooting 'cloudvirt1045.eqiad.wmnet'. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 15:07 wm-bot: Safe reboot of 'cloudvirt1044.eqiad.wmnet' finished successfully. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 15:03 wm-bot: Safe rebooting 'cloudvirt1044.eqiad.wmnet'. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 15:03 wm-bot: Safe reboot of 'cloudvirt1043.eqiad.wmnet' finished successfully. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 14:59 wm-bot: Safe rebooting 'cloudvirt1043.eqiad.wmnet'. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 14:59 wm-bot: Safe reboot of 'cloudvirt1042.eqiad.wmnet' finished successfully. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 14:40 wm-bot: Safe rebooting 'cloudvirt1042.eqiad.wmnet'. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 14:39 wm-bot: Safe reboot of 'cloudvirt1041.eqiad.wmnet' finished successfully. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 14:14 wm-bot: Safe rebooting 'cloudvirt1041.eqiad.wmnet'. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 14:14 wm-bot: Safe reboot of 'cloudvirt1039.eqiad.wmnet' finished successfully. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 14:10 wm-bot: Safe rebooting 'cloudvirt1039.eqiad.wmnet'. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 12:35 wm-bot: Safe rebooting 'cloudvirt1039.eqiad.wmnet'. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 11:56 wm-bot: Safe rebooting 'cloudvirt1038.eqiad.wmnet'. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 11:56 wm-bot: Safe reboot of 'cloudvirt1037.eqiad.wmnet' finished successfully. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 11:31 wm-bot: Safe rebooting 'cloudvirt1037.eqiad.wmnet'. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 11:31 wm-bot: Safe reboot of 'cloudvirt1036.eqiad.wmnet' finished successfully. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 11:08 wm-bot: Safe rebooting 'cloudvirt1036.eqiad.wmnet'. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 11:08 wm-bot: Safe reboot of 'cloudvirt1035.eqiad.wmnet' finished successfully. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 10:39 wm-bot: Safe rebooting 'cloudvirt1035.eqiad.wmnet'. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 10:39 wm-bot: Safe reboot of 'cloudvirt1034.eqiad.wmnet' finished successfully. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 10:13 wm-bot: Safe rebooting 'cloudvirt1034.eqiad.wmnet'. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 10:13 wm-bot: Safe reboot of 'cloudvirt1033.eqiad.wmnet' finished successfully. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 09:47 wm-bot: Safe rebooting 'cloudvirt1033.eqiad.wmnet'. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 09:47 wm-bot: Safe reboot of 'cloudvirt1032.eqiad.wmnet' finished successfully. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 09:21 wm-bot: Safe rebooting 'cloudvirt1032.eqiad.wmnet'. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 09:21 wm-bot: Safe reboot of 'cloudvirt1031.eqiad.wmnet' finished successfully. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 08:45 wm-bot: Safe rebooting 'cloudvirt1031.eqiad.wmnet'. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 08:45 wm-bot: Safe reboot of 'cloudvirt1030.eqiad.wmnet' finished successfully. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 08:19 wm-bot: Safe rebooting 'cloudvirt1030.eqiad.wmnet'. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 08:19 wm-bot: Safe reboot of 'cloudvirt1029.eqiad.wmnet' finished successfully. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 08:02 wm-bot: Safe rebooting 'cloudvirt1029.eqiad.wmnet'. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus === 2021-05-04 === * 16:05 wm-bot: Safe reboot of 'cloudvirt1028.eqiad.wmnet' finished successfully. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 15:45 wm-bot: Safe rebooting 'cloudvirt1028.eqiad.wmnet'. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 15:44 wm-bot: Safe reboot of 'cloudvirt1027.eqiad.wmnet' finished successfully. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 15:22 wm-bot: Safe rebooting 'cloudvirt1027.eqiad.wmnet'. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 15:19 wm-bot: Safe reboot of 'cloudvirt1026.eqiad.wmnet' finished successfully. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 15:15 wm-bot: Safe rebooting 'cloudvirt1026.eqiad.wmnet'. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 13:19 dcaro: rebooting cloudmetrics1002, got stuck again ([[phab:T275605|T275605]]) * 10:04 wm-bot: Safe rebooting 'cloudvirt1026.eqiad.wmnet'. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 09:10 wm-bot: Safe rebooting 'cloudvirt1026.eqiad.wmnet'. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 09:10 wm-bot: Safe reboot of 'cloudvirt1025.eqiad.wmnet' finished successfully. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 08:34 wm-bot: Safe rebooting 'cloudvirt1025.eqiad.wmnet'. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 08:20 wm-bot: Safe reboot of 'cloudvirt1024.eqiad.wmnet' finished successfully. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 08:03 wm-bot: Safe rebooting 'cloudvirt1024.eqiad.wmnet'. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus === 2021-05-03 === * 23:53 bstorm: running `maintain-dbusers harvest-replicas` on labstore1004 [[phab:T281287|T281287]] * 23:51 bstorm: running `maintain-dbusers harvest-replicas` on labstore1004 * 16:34 wm-bot: Safe reboot of 'cloudvirt1023.eqiad.wmnet' finished successfully. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 16:29 wm-bot: Safe rebooting 'cloudvirt1023.eqiad.wmnet'. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 15:41 wm-bot: Safe rebooting 'cloudvirt1023.eqiad.wmnet'. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 15:41 wm-bot: Safe reboot of 'cloudvirt1022.eqiad.wmnet' finished successfully. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 15:13 wm-bot: Safe rebooting 'cloudvirt1022.eqiad.wmnet'. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 10:31 wm-bot: Safe rebooting 'cloudvirt1021.eqiad.wmnet'. ([[phab:T280641|T280641]] - cookbook ran by dcaro@vulcanus) * 10:23 wm-bot: (from a cookbook) * 09:12 dcaro: draining and rebooting coludvirt1021 ([[phab:T280641|T280641]]) * 08:26 dcaro: draining and rebooting coludvirt1018 ([[phab:T280641|T280641]]) === 2021-04-30 === * 11:16 dcaro: draining and rebooting coludvirt1017, last one today ([[phab:T280641|T280641]]) * 10:37 dcaro: draining coludvirt1016 for reboot ([[phab:T280641|T280641]]) * 09:48 dcaro: draining coludvirt1013 for reboot ([[phab:T280641|T280641]]) === 2021-04-29 === * 15:11 dcaro: hard rebooting cloudmetrics1002, got hung again ([[phab:T275605|T275605]]) * 07:53 dcaro: Upgrading ceph libraries on cloudcontrol1005 to octopus ([[phab:T274566|T274566]]) * 07:51 dcaro: Upgrading ceph libraries on cloudcontrol1003 to octopus ([[phab:T274566|T274566]]) * 07:50 dcaro: Upgrading ceph libraries on cloudcontrol1004 to octopus ([[phab:T274566|T274566]]) === 2021-04-28 === * 21:11 andrewbogott: cleaning up more references to deleted hypervisors with delete from services where topic='compute' and version != 53; * 20:48 andrewbogott: cleaning up references to deleted hypervisors with mysql:root@localhost [nova_eqiad1]> delete from compute_nodes where hypervisor_version != '5002000'; * 19:40 andrewbogott: putting cloudvirt1040 into the maintenance aggregate pending more info about [[phab:T281399|T281399]] * 18:11 andrewbogott: adding cloudvirt1040, 1041 and 1042 to the 'ceph' host aggregate -- [[phab:T275081|T275081]] * 11:06 dcaro: All ceph server side upgraded to Octopus! \o/ ([[phab:T280641|T280641]]) * 10:57 dcaro: Got a PG getting stuck on 'remapping' after the OSD came up, had to unset the norebalance and then set it again to get it unstuck ([[phab:T280641|T280641]]) * 10:34 dcaro: Slow/blocked opns from cloudcephmon03, "osd_failure(failed timeout osd.32..." (cloudcephosd1005), unset the cluster noout/norebalance and went away in a few secs, setting it again and continuing... ([[phab:T280641|T280641]]) * 09:03 dcaro: Waiting for slow heartbeats from osd.58(cloudcephosd1002) to recover... ([[phab:T280641|T280641]]) * 08:59 dcaro: During the upgrade, started getting warning 'slow osd heartbacks in the back', meaning that pings between osds are really slow (up to 190s) all from osd.58, currently on cloudcephosd1002 ([[phab:T280641|T280641]]) * 08:58 dcaro: During the upgrade, started getting warning 'slow osd heartbacks in the back', meaning that pings between osds are really slow (up to 190s) all from osd.58 ([[phab:T280641|T280641]]) * 08:58 dcaro: During the upgrade, started getting warning 'slow osd heartbacks in the back', meaning that pings between osds are really slow (up to 190s) ([[phab:T280641|T280641]]) * 08:21 dcaro: Upgrading all the ceph osds on eqiad ([[phab:T280641|T280641]]) * 08:21 dcaro: The clock skew seems intermittent, there's another task to follw it [[phab:T275860|T275860]] ([[phab:T280641|T280641]]) * 08:18 dcaro: All equiad ceph mons and mgrs upgraded ([[phab:T280641|T280641]]) * 08:18 dcaro: During the upgrade, ceph detected a clock skew on cloudcephmon1002, cloudcephmon1001, they are back ([[phab:T280641|T280641]]) * 08:15 dcaro: During the upgrade, ceph detected a clock skew on cloudcephmon1002, it went away, I'm guessing systemd-timesyncd fixed it ([[phab:T280641|T280641]]) * 08:14 dcaro: During the upgrade, ceph detected a clock skew on cloudcephmon1002, looking ([[phab:T280641|T280641]]) * 07:58 dcaro: Upgrading ceph services on eqiad, starting with mons/managers ([[phab:T280641|T280641]]) === 2021-04-27 === * 14:10 dcaro: codfw.openstack upgraded ceph libraries to 15.2.11 ([[phab:T280641|T280641]]) * 13:07 dcaro: codfw.openstack cloudvirt2002-dev done, taking cloudvirt2003-dev out to upgrade ceph libraries ([[phab:T280641|T280641]]) * 13:00 dcaro: codfw.openstack cloudvirt2001-dev back online, taking cloudvirt2002-dev out to upgrade ceph libraries ([[phab:T280641|T280641]]) * 10:51 dcaro: ceph.eqiad: cinder pool got it's pg_num increased to 1024, re-shuffle started ([[phab:T273783|T273783]]) * 10:48 dcaro: ceph.eqiad: Tweaked the target_size_ratio of all the pools, enabling autoscaler (it will increase cinder pool only) ([[phab:T273783|T273783]]) * 09:14 dcaro: manually force stopping the server puppetmaster-01 to unblock migration (in codfw1) * 09:14 dcaro: manually force stopping the server puppetmaster-01 to unblock migration * 08:59 dcaro: manually force stopping the server exploding-head on codfw, to try cold migration * 08:47 dcaro: restarting nova-compute on cloudvirt2001-dev after upgrading ceph libraries to 15.2.11 === 2021-04-26 === * 20:56 andrewbogott: deleting spurious 'codfw1dev' and 'codw1dev-4' regions in the dallas deployment; regions without endpoints break a bunch of things * 09:45 dcaro: draining cloudvirt2001-dev with the new cookbooks ([[phab:T280641|T280641]]) === 2021-04-23 === * 13:49 dcaro: testing the drain_cloudvirt cookbook on codfw1 openstack cluster, draining cloudvirt2001 ([[phab:T280641|T280641]]) * 11:12 dcaro: testing the drain_cloudvirt cookbook on codfw1 openstack cluster ([[phab:T280641|T280641]]) * 09:32 dcaro: finished upgrade of ceph cluster on codfw1 using exclusively cookbooks ([[phab:T280641|T280641]]) * 09:17 dcaro: testing the upgrade_osds cookbook on codfw1 ceph cluster ([[phab:T280641|T280641]]) * 08:17 dcaro: testing the upgrade_mons cookbook on codfw1 ceph cluster ([[phab:T280641|T280641]]) === 2021-04-21 === * 17:59 dcaro: all monitors upgraded on codfw1 with one cookbook `cookbook --verbose -c ~/.config/spicerack/cookbook.yaml wmcs.ceph.upgrade_mons --monitor-node-fqdn cloudcephmon2002-dev.codfw.wmnet` ([[phab:T280641|T280641]]) * 17:47 dcaro: upgrading monitors and mrg nodes on codfw ceph cluster ([[phab:T280641|T280641]]) * 13:26 dcaro: testing ceph upgrade cookbook on cloudcephmon2002-dev ([[phab:T280641|T280641]]) === 2021-04-20 === * 20:21 andrewbogott: reboot cloudservices1003 * 20:13 andrewbogott: reboot cloudservices1004 === 2021-04-19 === * 08:40 dcaro: enabling puppet on labstore1004 after mysql restart ([[phab:T279657|T279657]]) * 08:09 dcaro: downtiming labstore1004 and stopping puppet for mysql restart ([[phab:T279657|T279657]]) === 2021-04-14 === * 10:48 dcaro: Upgrade of codfw ceph to octopus 15.2.20 done, will run some performance tests now ([[phab:T274566|T274566]]) * 10:41 dcaro: Upgrade of codfw ceph to octopus 15.2.20, mgrs upgraded, osds next ([[phab:T274566|T274566]]) * 10:37 dcaro: Upgrade of codfw ceph to octopus 15.2.20, mons upgraded, mgrs next ([[phab:T274566|T274566]]) * 10:15 dcaro: starting the upgrade of codfw ceph to octopus 15.2.20 ([[phab:T274566|T274566]]) * 10:07 dcaro: Merged the ceph 15 (Octopus) repo deployment to codfw, only the repo, not the packages ([[phab:T274566|T274566]]) === 2021-04-13 === * 16:42 dcaro: Ceph balancer got the cluster to eval 0.014916, that is 88-77% usage for compute pool, and 28-19% usage for the cinder one \o/ ([[phab:T274573|T274573]]) * 15:08 dcaro: Activating continuous upmap balancer, keeping a close eye ([[phab:T274573|T274573]]) * 15:03 dcaro: Executing a second pass, there's still movements to improve the eval of 0.030075 ([[phab:T274573|T274573]]) * 15:02 dcaro: First pass finished, improved eval to 0.030075 ([[phab:T274573|T274573]]) * 14:49 dcaro: Running the first_pass balancing plan on ceph eqiad, current eval 0.030622 ([[phab:T274573|T274573]]) * 14:43 dcaro: enabling ceph upmap pg balancer on equiad ([[phab:T274573|T274573]]) * 14:36 andrewbogott: upgrading codfw1dev to version Victoria, [[phab:T261137|T261137]] * 13:11 andrewbogott: upgrading eqiad1 designate to version Victoria, [[phab:T261137|T261137]] * 10:44 dcaro: enabled ceph upmap balancer on codfw ([[phab:T274573|T274573]],[[phab:T274573|T274573]]) === 2021-04-07 === * 21:33 andrewbogott: upgrading codfw1dev designate to Victoria === 2021-04-04 === * 17:36 andrewbogott: upgrading eqiad1 designate to Ussuri === 2021-04-02 === * 14:12 andrewbogott: upgrading codfw1dev to OpenStack version Ussuri === 2021-04-01 === * 12:15 dcaro: Restoring the 4.9 kernel on cloudcephosd2003-dev and upgrading ([[phab:T274565|T274565]]) * 10:29 dcaro: Done restoring the 4.9 kernel on cloudcephosd2001-dev and upgrading, requires logging into console to boot from the older kernel before removing the newer one ([[phab:T274565|T274565]]) * 10:10 dcaro: Restoring the 4.9 kernel on cloudcephosd2001-dev and upgrading ([[phab:T274565|T274565]]) === 2021-03-31 === * 08:47 dcaro: upgrading cinder on codfw cloudcontrol2* nodes ([[phab:T278845|T278845]]) === 2021-03-30 === * 09:53 arturo: rebooting cloudnet1003 to cleanup conntrack table, it wouldn't cleanup by hand ... === 2021-03-28 === * 15:42 andrewbogott: updated debian-10.0-buster base image === 2021-03-27 === * 09:54 arturo: cleanup conntrack table in qrouter nents in cloudnet1003 (backup) === 2021-03-25 === * 19:03 andrewbogott: deleting all unused (per wmcs-imageusage) Jessie base images from Glance * 17:15 andrewbogott: refreshing puppet compiler facts for tools project * 10:31 dcaro: kernel upgrade on osds on codfw done, running performance tests ([[phab:T274565|T274565]]) * 10:24 dcaro: upgrading kernel on cloudcephosd2003-dev and reboot ([[phab:T274565|T274565]]) * 10:18 dcaro: upgrading kernel on cloudcephosd2002-dev and reboot ([[phab:T274565|T274565]]) * 10:08 dcaro: upgrading kernel on cloudcephmon2003-dev and reboot ([[phab:T274565|T274565]]) === 2021-03-24 === * 09:19 dcaro: restarted wmcs-backup on cloudvirt1024 as it failed due to an image being removed while running ([[phab:T276892|T276892]]) === 2021-03-23 === * 11:33 arturo: root@cloudcontrol1005:~# wmcs-novastats-dnsleaks --delete === 2021-03-22 === * 10:10 arturo: cleanup conntrack table in standby node: aborrero@cloudnet1003:~ $ sudo ip netns exec qrouter-d93771ba-2711-4f88-804a-{{Gerrit|8df6fd03978a}} conntrack -F === 2021-03-19 === * 17:18 bstorm: running `ALTER TABLE account MODIFY COLUMN type ENUM('user','tool','paws');` against the labsdbaccounts database on m5 [[phab:T276284|T276284]] * 14:29 andrewbogott: switching admin-monitoring project to use an upstream debian image; I want to see how this affects performance * 00:30 bstorm: downtimed labstore1004 to check some things in debug mode === 2021-03-17 === * 17:28 bstorm: restarted the backup-glance-images job to clear errors in systemd [[phab:T271782|T271782]] * 17:16 andrewbogott: set default cinder quota for projects to 80Gb with "update quota_classes set hard_limit=80 where resource='gigabytes';" on database 'cinder' * 16:58 andrewbogott: disabling all flavors with >20Gb root storage with "update flavors set disabled=1 where root_gb>20;" in nova_eqiad1_api === 2021-03-10 === * 16:51 arturo: rebooting cloudvirt1030 for [[phab:T275753|T275753]] * 13:14 dcaro: starting manually the canary VM for cloudvirt1029 (nova start 349830f6-3b39-4a8c-ada4-{{Gerrit|a7439f65cffe}}) ([[phab:T275753|T275753]]) * 12:51 arturo: draining cloudvirt1030 for [[phab:T275753|T275753]] * 12:47 arturo: rebooting cloudvirt1029 for [[phab:T275753|T275753]] * 11:56 arturo: [codfw1dev] restart rabbitmq-server in all 3 cloudcontrol servers for [[phab:T276964|T276964]] * 11:53 arturo: [codfw1dev] restart nova-conductor in all 3 cloudcontrol servers for [[phab:T276964|T276964]] * 11:31 arturo: draining cloudvirt1029 for [[phab:T275753|T275753]] * 11:29 arturo: rebooting cloudvirt1013 for [[phab:T275753|T275753]] * 11:05 arturo: draining cloudvirt1013 for [[phab:T275753|T275753]] * 11:00 arturo: rebooting cloudvirt1028 for [[phab:T275753|T275753]] * 10:33 arturo: draining cloudvirt1028 for [[phab:T275753|T275753]] * 10:29 arturo: rebooting cloudvirt1023 for [[phab:T275753|T275753]] * 09:37 arturo: draining cloudvirt1023 for [[phab:T275753|T275753]] * 09:07 arturo: [codfw1dev] reimaging cloudvirt2003-dev ([[phab:T276964|T276964]]) === 2021-03-09 === * 16:27 arturo: rebooting cloudvirt1027 ([[phab:T275753|T275753]]) * 13:39 arturo: draining cloudvrit1027 for [[phab:T275753|T275753]] * 13:35 arturo: icinga-downtime cloudvirt1038 for 30 days for [[phab:T276922|T276922]] * 13:21 arturo: add cloudvirt1039 to the ceph host aggregate (no longer a spare, we have cloudvirt1038 with HW failures) * 12:52 arturo: cloudvirt1038 hard powerdown / powerup for [[phab:T276922|T276922]] * 12:33 arturo: rebooting cloudvirt1038 ([[phab:T275753|T275753]]) * 10:58 arturo: draining cloudvirt1038 ([[phab:T275753|T275753]]) * 10:54 arturo: rebooting cloudvirt1037 ([[phab:T275753|T275753]]) * 09:59 arturo: draining cloudvirt1037 ([[phab:T275753|T275753]]) * 09:12 dcaro: restarted the wmcs-backup service on cloudvirt1024 to retry the backups (failed because a VM was removed in-between, [[phab:T276892|T276892]]) === 2021-03-05 === * 21:40 andrewbogott: replacing 'observer' role with 'reader' role in eqiad1 [[phab:T276018|T276018]] * 21:21 andrewbogott: replacing 'observer' role with 'reader' role in eqiad1 * 16:23 arturo: rebooting cloudvirt1036 for [[phab:T275753|T275753]] * 12:30 arturo: draining cloudvirt1036 for [[phab:T275753|T275753]] * 12:25 arturo: rebooting cloudvirt1035 for [[phab:T275753|T275753]] * 10:49 arturo: rebooting cloudvirt1035 for [[phab:T275753|T275753]] * 10:47 arturo: rebooting cloudvirt1034 for [[phab:T275753|T275753]] * 10:26 arturo: draining cloudvirt1034 for [[phab:T275753|T275753]] * 10:25 arturo: rebooting cloudvirt1033 for [[phab:T275753|T275753]] * 09:18 arturo: draining cloudvirt1033 for [[phab:T275753|T275753]] === 2021-03-04 === * 18:36 andrewbogott: rebooting cloudmetrics1002; the console is hanging * 16:59 arturo: rebooting cloudvirt1032 for [[phab:T275753|T275753]] * 16:34 arturo: draining cloudvirt1032 for [[phab:T275753|T275753]] * 16:33 arturo: rebooting cloudvirt1031 for [[phab:T275753|T275753]] * 16:11 arturo: draining cloudvirt1031 for [[phab:T275753|T275753]] * 16:09 arturo: rebooting cloudvirt1026 for [[phab:T275753|T275753]] * 15:57 arturo: draining cloudvirt1026 for [[phab:T275753|T275753]] * 15:55 arturo: rebooting cloudvirt1025 for [[phab:T275753|T275753]] * 15:41 arturo: draining cloudvirt1025 for [[phab:T275753|T275753]] * 15:12 arturo: rebooting cloudvirt1024 for [[phab:T275753|T275753]] * 11:29 arturo: draining cloudvirt1024 for [[phab:T275753|T275753]] * 11:24 dcaro: rebooted cloudvirt1022, re-adding to ceph and removing from maintenance host aggregate for [[phab:T275753|T275753]] * 11:01 dcaro: rebooting cloudvirt1022 for [[phab:T275753|T275753]] * 09:12 dcaro: draining cloudvirt1022 for [[phab:T275753|T275753]] === 2021-03-03 === * 17:16 andrewbogott: restarting rabbitmq-server on cloudcontrol1003,1004,1005; trying to explain amqp errors in scheduler logs * 16:03 dcaro: draining cloudvirt1022 for [[phab:T275753|T275753]] * 16:03 dcaro: draining cloudvirt1022 for [[phab:T275753|T275753]] * 16:00 arturo: move cloudvirt1013 into the 'toobusy' host aggregate, it has 221% cpu subscription and 82% MEM subscription * 15:34 arturo: rebooting cloudvirt1021 for [[phab:T275753|T275753]] * 14:31 arturo: draining cloudvirt1021 for [[phab:T275753|T275753]] * 13:59 arturo: rebooting cloudvirt1018 for [[phab:T275753|T275753]] * 13:28 arturo: draining cloudvirt1018 for [[phab:T275753|T275753]] * 12:49 arturo: rebooting cloudvirt1017 for [[phab:T275753|T275753]] * 12:22 arturo: draining cloudvirt1017 for [[phab:T275753|T275753]] * 12:20 arturo: rebooting cloudvirt1016 for [[phab:T275753|T275753]] * 12:01 arturo: draining cloudvirt1016 for [[phab:T275753|T275753]] * 11:59 arturo: cloudvirt1014 now in the ceph host aggregate * 11:58 arturo: rebooting cloudvirt1014 for [[phab:T275753|T275753]] * 11:50 arturo: moved cloudvirt1023 away from the maintenance host aggregate, leave it in the ceph aggregate (was in the 2) * 11:47 arturo: moved cloudvirt1014 to the 'maintenance' host aggregate, drain it for [[phab:T275753|T275753]] * 10:01 arturo: icinga-downtime cloudnet1003 for 14 days bc potential alerting storm due to firmware issues ([[phab:T271058|T271058]]) * 10:01 arturo: rebooting again cloudnet1003 (no network failover) ([[phab:T271058|T271058]]) * 09:59 arturo: update firmware-bnx2x from 20190114-2 to 20200918-1~bpo10+1 on cloudnet1003 ([[phab:T271058|T271058]]) * 09:30 arturo: installing linux kernel 5.10.13-1~bpo10+1 in cloudnet1003 and rebooting it (network failover) ([[phab:T271058|T271058]]) === 2021-03-02 === * 17:16 andrewbogott: rebooting cloudvirt1039 to see if I can trigger [[phab:T276208|T276208]] * 16:10 arturo: [codfw1dev] restart nova-compute on cloudvirt2002-dev * 11:59 arturo: moved cloudvirt1012 to 'maintenance' host aggregate. Drain it with `wmcs-drain-hypervisor` to reboot it for [[phab:T275753|T275753]] * 11:59 arturo: cloudvirt1023 is affected by [[phab:T276208|T276208]] and cannot be rebooted. Put it back into the ceph hos aggregate * 10:43 arturo: moved cloudvirt1013 cloudvirt1032 cloudvirt1037 back into the 'ceph' host aggregate * 10:13 arturo: moved cloudvirt1023 to 'maintenance' host aggregate. Drain it with `wmcs-drain-hypervisor` to reboot it for [[phab:T275753|T275753]] === 2021-03-01 === * 20:12 andrewbogott: removing novaadmin from all projects save 'admin' for [[phab:T274385|T274385]] * 19:51 andrewbogott: removing novaobserver from all projects save 'observer' for [[phab:T274385|T274385]] * 19:50 andrewbogott: adding inherited domain-wide roles to novaadmin and novaobserver as per [[phab:T274385|T274385]] === 2021-02-28 === * 04:54 andrewbogott: restarted redis-server on tools-redis-1003 and tools-redis-1004 in an attempt to reduce replag, no real change detected === 2021-02-27 === * 00:33 andrewbogott: sudo cumin --timeout 500 "A:all and not O<nowiki>{</nowiki>project:clouddb-services<nowiki>}</nowiki>" 'lsb_release -c {{!}} grep -i buster && uname -r {{!}} grep -v 4.19.0-14-amd64 && reboot' * 00:28 andrewbogott: sudo cumin --timeout 500 "A:all and not O<nowiki>{</nowiki>project:clouddb-services<nowiki>}</nowiki>" 'lsb_release -c {{!}} grep -i buster && uname -r {{!}} grep -v 4.19.0-14-amd64 && echo reboot' * 00:09 andrewbogott: sudo cumin "A:all and not O<nowiki>{</nowiki>project:clouddb-services<nowiki>}</nowiki>" 'lsb_release -c {{!}} grep -i stretch && uname -r {{!}} grep -v 4.19.0-0.bpo.14-amd64 && reboot' === 2021-02-26 === * 14:58 dcaro: [eqiad] rebooting cloudcephosd1015 (last osd \o/) for kernel upgrade ([[phab:T275753|T275753]]) * 14:51 dcaro: [eqiad] rebooting cloudcephosd1014 for kernel upgrade ([[phab:T275753|T275753]]) * 14:44 dcaro: [eqiad] rebooting cloudcephosd1013 for kernel upgrade ([[phab:T275753|T275753]]) * 14:38 dcaro: [eqiad] rebooting cloudcephosd1012 for kernel upgrade ([[phab:T275753|T275753]]) * 14:31 dcaro: [eqiad] rebooting cloudcephosd1011 for kernel upgrade ([[phab:T275753|T275753]]) * 14:25 dcaro: [eqiad] rebooting cloudcephosd1010 for kernel upgrade ([[phab:T275753|T275753]]) * 14:17 dcaro: [eqiad] rebooting cloudcephosd1009 for kernel upgrade ([[phab:T275753|T275753]]) * 13:54 dcaro: [eqiad] downtimed alert1001 Ceph OSDs down alert until 18:00 GMT+1 as that is not under the host being rebooted ([[phab:T275753|T275753]]) * 13:51 dcaro: [eqiad] rebooting cloudcephosd1008 for kernel upgrade ([[phab:T275753|T275753]]) * 13:45 dcaro: [eqiad] rebooting cloudcephosd1007 for kernel upgrade ([[phab:T275753|T275753]]) * 13:38 dcaro: [eqiad] rebooting cloudcephosd1006 for kernel upgrade ([[phab:T275753|T275753]]) * 12:07 dcaro: [eqiad] rebooting cloudcephosd1005 for kernel upgrade ([[phab:T275753|T275753]]) * 12:00 arturo: rebooting cloudcontrol1003 for kernel upgrade ([[phab:T275753|T275753]]) * 11:42 arturo: rebooting cloudcontrol1004 for kernel upgrade ([[phab:T275753|T275753]]) * 11:41 dcaro: [eqiad] rebooting cloudcephosd1004 for kernel upgrade ([[phab:T275753|T275753]]) * 11:32 dcaro: [eqiad] rebooting cloudcephosd1003 for kernel upgrade ([[phab:T275753|T275753]]) * 11:30 arturo: rebooting cloudcontrol1005 for kernel upgrade ([[phab:T2|T2]] * 11:26 dcaro: [eqiad] rebooting cloudcephosd1002 for kernel upgrade ([[phab:T275753|T275753]]) * 11:16 dcaro: [eqiad] rebooting cloudcephosd1001 for kernel upgrade ([[phab:T275753|T275753]]) * 11:11 dcaro: [eqiad] rebooting cloudcephmon1003 for kernel upgrade ([[phab:T275753|T275753]]) * 11:05 dcaro: [eqiad] rebooting cloudcephmon1002 for kernel upgrade ([[phab:T275753|T275753]]) * 10:59 dcaro: [eqiad] rebooting cloudcephmon1001 for kernel upgrade ([[phab:T275753|T275753]]) * 10:45 arturo: rebooting cloudvirt1039 into a new kernel ([[phab:T275753|T275753]]) --- spare * 10:43 dcaro: [codfw1dev] rebooting cloudcephmon2003-dev for kernel upgrade ([[phab:T275753|T275753]]) * 10:38 dcaro: [codfw1dev] rebooting cloudcephmon2002-dev for kernel upgrade ([[phab:T275753|T275753]]) * 10:29 dcaro: [codfw1dev] rebooting cloudcephmon2001-dev for kernel upgrade ([[phab:T275753|T275753]]) * 10:24 arturo: [codfw1dev] purge old kernel packages on cloudvirt2003-dev to force boot into a new kernel ([[phab:T275753|T275753]]) * 10:11 arturo: [codfw1dev] manually creating /boot/grub/ on cloudvirt2003-dev to allow update-grub2 to run (so it can reboot into a new kernel) ([[phab:T275753|T275753]]) * 10:11 dcaro: [codfw1dev] rebooting cloudcephosd2003-dev for kernel upgrade ([[phab:T275753|T275753]]) * 10:05 dcaro: [codfw1dev] rebooting cloudcephosd2002-dev for kernel upgrade ([[phab:T275753|T275753]]) * 10:01 arturo: [codfw1dev] rebooting cloudvirt200X-dev for kernel upgrade ([[phab:T275753|T275753]]) * 09:59 arturo: [codfw1dev] rebooting cloudweb2001-dev for kernel upgrade ([[phab:T275753|T275753]]) * 09:53 arturo: [codfw1dev] rebooting cloudservices2003-dev for kernel upgrade ([[phab:T275753|T275753]]) * 09:51 arturo: [codfw1dev] rebooting cloudservices2002-dev for kernel upgrade ([[phab:T275753|T275753]]) * 09:45 arturo: [codfw1dev] rebooting cloudcontrol2004-dev for kernel upgrade ([[phab:T275753|T275753]]) * 09:44 arturo: [codfw1dev] rebooting cloudbackup[2001-2002].codfw.wmnet for kernel upgrade ([[phab:T275753|T275753]]) * 09:43 dcaro: [codfw1dev] rebooting cloudcephosd2001-dev for kernel upgrade ([[phab:T275753|T275753]]) * 09:41 arturo: [codfw1dev] rebooting cloudcontrol2003-dev for kernel upgrade ([[phab:T275753|T275753]]) * 09:33 arturo: [codfw1dev] rebooting cloudcontrol2001-dev for kernel upgrade ([[phab:T275753|T275753]]) === 2021-02-25 === * 14:56 arturo: deployed wmcs-netns-events daemon to all cloudnet servers ([[phab:T275483|T275483]]) === 2021-02-24 === * 11:07 arturo: force-reboot cloudmetrics1002, add icinga downtime for 2 hours. Investigating some server issue * 00:17 bstorm: set --property hw_scsi_model=virtio-scsi and --property hw_disk_bus=scsi on the main stretch image in glance on eqiad1 [[phab:T275430|T275430]] === 2021-02-23 === * 22:43 bstorm: set --property hw_scsi_model=virtio-scsi and --property hw_disk_bus=scsi on the main buster image in glance on eqiad1 [[phab:T275430|T275430]] * 20:36 andrewbogott: adding r/o access to the eqiad1-glance-images ceph pool for the client.eqiad1-compute for [[phab:T275430|T275430]] * 10:49 arturo: rebooting clounet1004 into new kernel from buster-bpo ([[phab:T271058|T271058]]) * 10:49 arturo: installing linux-image-amd64 from buster-bpo 5.10.13-1~bpo10+1 in cloudnet1004 ([[phab:T271058|T271058]]) === 2021-02-22 === * 17:15 bstorm: restarting nova-compute on cloudvirt1016 and cloudvirt1036 in case it helps [[phab:T275411|T275411]] * 15:02 dcaro: Re-uploaded the debian buster 10.0 image from rbd to glance, that worked, re-spawning all the broken instances ([[phab:T275378|T275378]]) * 11:12 dcaro: Refreshing all the canary instances ([[phab:T275354|T275354]]) === 2021-02-18 === * 14:50 arturo: rebooting cloudnet1004 for [[phab:T271058|T271058]] * 10:25 dcaro: Rebooting cloudmetrics1001 to apply new kernel ([[phab:T275116|T275116]]) * 10:16 dcaro: Rebooting cloudmetrics1002 to apply new kernel ([[phab:T275116|T275116]]) * 10:14 dcaro: Upgrading grafana on cloudmetrics1002 ([[phab:T275116|T275116]]) * 10:12 dcaro: Upgrading grafana on cloudmetrics1001 ([[phab:T275116|T275116]]) === 2021-02-17 === * 15:58 arturo: deploying https://gerrit.wikimedia.org/r/c/operations/puppet/+/664845 to cloudnet servers ([[phab:T268335|T268335]]) === 2021-02-15 === * 16:25 arturo: [codfw1dev] rebooting all cloudgw200x-dev / cloudnet200x-dev servers ([[phab:T272963|T272963]]) * 15:45 arturo: [codfw1dev] drop subnet definition for cloud-instances-transport1-b-codfw ([[phab:T272963|T272963]]) * 15:45 arturo: [codfw1dev] connect virtual router cloudinstances2b-gw to vlan cloud-gw-transport-codfw (185.15.57.10) ([[phab:T272963|T272963]]) === 2021-02-11 === * 12:01 arturo: [codfw1dev] drop instance `tools-codfw1dev-bastion-1` in `tools-codfw1dev` (was buster, cannot use it yet) * 11:59 arturo: [codfw1dev] create instance `tools-codfw1dev-bastion-2` (stretch) in `tools-codfw1dev` to test stuff related to [[phab:T272397|T272397]] * 11:45 arturo: [codfw1dev] create instance `tools-codfw1dev-bastion-1` in `tools-codfw1dev` to test stuff related to [[phab:T272397|T272397]] * 11:42 arturo: [codfw1dev] drop `tools` project, create `tools-codfw1dev` * 11:38 arturo: [codfw1dev] drop `coudinfra` project (we are using `cloudinfra-codfw1dev` there) * 05:37 bstorm: downtimed cloudnet1004 for another week [[phab:T271058|T271058]] === 2021-02-09 === * 15:23 arturo: icinga-downtime for 2h everything *labs *cloud for openstack upgrades * 11:14 dcaro: Merged the osd scheduler change for all osds, applying on all cloudcephosd* ([[phab:T273791|T273791]]) === 2021-02-08 === * 18:50 bstorm: enabled puppet on cloudvirt1023 for now [[phab:T274144|T274144]] * 18:44 bstorm: restarted the backup_vms.service on cloudvirt1027 [[phab:T274144|T274144]] * 17:51 bstorm: deleted project pki [[phab:T273175|T273175]] === 2021-02-05 === * 10:59 arturo: icinga-downtime labstore1004 tools share space check for 1 week ([[phab:T272247|T272247]]) * 10:21 dcaro: This was affecting maps and several others, maps and project-proxy have been fixed ([[phab:T273956|T273956]]) * 09:19 dcaro: Some certs around the infra are expired ([[phab:T273956|T273956]]) === 2021-02-04 === * 10:12 dcaro: Increasing the memory limit of osds in eqiad from 8589934592(8G) to 12884901888(12G) ([[phab:T273851|T273851]]) === 2021-02-03 === * 09:59 dcaro: Doing a full vm backup on cloudvirt1024 with the new script ([[phab:T260692|T260692]]) * 01:50 bstorm: icinga-downtime cloudnet1004 for a week [[phab:T271058|T271058]] === 2021-02-02 === * 17:14 dcaro: Changed osd memory limit from 4G to 8G ([[phab:T273649|T273649]]) * 11:00 arturo: icinga-downtime cloudvirt-wdqs1001 for 1 week ([[phab:T273579|T273579]]) * 03:12 andrewbogott: running /usr/local/sbin/wmcs-purge-backups and /usr/local/sbin/wmcs-backup-instances on cloudvirt1024 to see why the backup job paged === 2021-01-29 === * 15:36 andrewbogott: disabling puppet and some services on eqiad1 cloudcontrol nodes; replacing nova-placement-api with placement-api === 2021-01-28 === * 19:44 andrewbogott: shutting down cloudcontrol2001-dev because it's in a partially upgraded state; will revive when it's time for Train === 2021-01-27 === * 00:50 bstorm: icinga-downtime cloudnet1004 for a week [[phab:T271058|T271058]] === 2021-01-22 === * 16:44 andrewbogott: upgrading designate on cloudvirt1003/1004 to OpenStack 'train' * 11:29 dcaro: Doing some tests removed cloudcontrol1003 puppet cert, regenerating... === 2021-01-21 === * 11:35 arturo: merging core router firewall changes https://gerrit.wikimedia.org/r/c/operations/homer/public/+/657439 ([[phab:T209082|T209082]]) * 11:30 arturo: merging core router firewall changes https://gerrit.wikimedia.org/r/c/operations/homer/public/+/657358 ([[phab:T272486|T272486]], [[phab:T209082|T209082]]) === 2021-01-20 === * 10:49 arturo: merging core router firewall change https://gerrit.wikimedia.org/r/c/operations/homer/public/+/657302 ([[phab:T209082|T209082]]) * 10:05 dcaro: Everything looks ok, created a new vm with a volume in ceph without issues, and on warnings/errors on ceph status, closing ([[phab:T272303|T272303]]) * 09:55 dcaro: Eqiad ceph cluster uprgaded, doing sanity checks ([[phab:T272303|T272303]]) * 09:46 dcaro: 75% of the eqiad cluster upgraded... continuing ([[phab:T272303|T272303]]) * 09:37 dcaro: 25% of the eqiad cluster upgraded... continuing ([[phab:T272303|T272303]]) * 09:24 dcaro: Mgr daemons upgraded and running, upgrading osd daemons on servers cloudcephosd1*, this make take a bit longer ([[phab:T272303|T272303]]) * 09:22 dcaro: Mon daemons upgraded and running, upgrading mgr daemons on servers cloudcephmon1* ([[phab:T272303|T272303]]) * 09:16 dcaro: Starting eqiad ceph upgrade, upgrading the mon servers cloudcephmon1* ([[phab:T272303|T272303]]) * 09:01 dcaro: Will start the ceph upgrade in 15 min, no downtime nor performance impact is expected ([[phab:T272303|T272303]]) === 2021-01-19 === * 10:17 arturo: icinga-downtime cloudnet1004 for 1 week ([[phab:T271058|T271058]]) === 2021-01-18 === * 16:00 dcaro: Codfw1 ceph cluster uprgaded, will wait until tomorrow to see if there's any instability, but everything looks fine ([[phab:T272303|T272303]]) * 15:38 dcaro: Upgraded mgr sevices on codfw ceph cluster, starting with osd ones ([[phab:T272303|T272303]]) * 15:35 dcaro: Upgraded mon sevices on codfw ceph cluster, starting with mgr ones ([[phab:T272303|T272303]]) * 15:21 dcaro: Starting upgrade of ceph mon nodes on codfw ([[phab:T272303|T272303]]) * 15:06 dcaro: re-enabling puppet on cloudcephosd2* hosts * 13:53 dcaro: disabling puppet on cloudcephosd2* to resume perf tests * 10:50 dcaro: re-enabling puppet on cephcloudosd2* (codfw) * 10:07 dcaro: disabling puppet on cephcloudosd2* (codfw) to do some performance tests * 09:00 dcaro: Enabling custom application 'cinder' on pool codfw1dev-cinder to get rid of health warnings === 2021-01-17 === * 16:53 arturo: icinga downtime labstore1004 /srv/tools space check for 3 days ([[phab:T272247|T272247]]) === 2021-01-15 === * 13:41 arturo: icinga downtime labstore1004 maintain-dbuser alert until 2021-01-19 ([[phab:T272125|T272125]]) * 09:47 arturo: labstore1004 maintain-dbusers affected by [[phab:T272127|T272127]] and [[phab:T272125|T272125]] * 09:22 arturo: restart maintain-dbusers.service in labstore1004 * 08:19 dcaro: Merging the patch to disable write caches on ceph osds ([[phab:T271527|T271527]]) === 2021-01-13 === * 17:03 arturo: remove cloudvirt1013 cloudvirt1032 cloudvirt1037 to the 'toobusy' host aggregate to prevent further CPU oversubscribing * 12:40 arturo: try increasing systemd watchdog timeout for conntrackd in cloudnet1004 ([[phab:T268335|T268335]]) * 11:45 dcaro: https://gerrit.wikimedia.org/r/c/operations/puppet/+/654419 merged and deployed (and tested) ([[phab:T268877|T268877]]) * 11:40 dcaro: merging https://gerrit.wikimedia.org/r/c/operations/puppet/+/654419 that might affect the encapi service (puppet on cloud environment), no downtime expected though ([[phab:T268877|T268877]]) * 10:56 arturo: trying to cleanup dpkg package mess in cloudnet2002-dev * 10:02 arturo: prevent floating IP allocation from neutron transport subnet: root@cloudcontrol1005:~# neutron subnet-update --allocation-pool start=185.15.56.244,end=185.15.56.244 cloud-instances-transport1-b-eqiad1 ([[phab:T271867|T271867]]) === 2021-01-12 === * 10:33 arturo: reboot cloudnet1004 * 10:32 arturo: update firmware-bnx2x from 20190114-2 to 20200918-1~bpo10+1 on cloudnet1004 ([[phab:T271058|T271058]]) === 2021-01-11 === * 10:22 arturo: doubling size of conntrack table in cloudnet servers https://gerrit.wikimedia.org/r/c/operations/puppet/+/655407 ([[phab:T271058|T271058]]) * 10:07 arturo: manually cleanup conntrack table in cloudnet1004 ([[phab:T271058|T271058]]) * 09:19 dcaro: cleaned up ~1800 snapshots, 109 remaining only, one for each host x image combination (plus some ephemeral ones while doing backups), closing the task ([[phab:T270478|T270478]]) * 08:39 dcaro: cleaning up dangling snapshots now that we have the new suffixed ones ([[phab:T270478|T270478]]) === 2021-01-10 === * 16:02 andrewbogott: restarting rabbitmq-server on all eqiad1 cloudcontrols * 15:54 andrewbogott: restating neutron-metadata-agent on cloudnet1004 due to many syslog complaints === 2021-01-08 === * 11:25 arturo: rebooting both cloudnet2002-dev/cloudnet2003-dev to make sure interfaces are set up correctl ([[phab:T271517|T271517]]) * 11:22 arturo: connecting cloudnet2002-dev cloudnet2003-dev back to vlan 2120 ([[phab:T271517|T271517]]) * 11:06 arturo: root@cloudcontrol2001-dev:~# openstack router set --external-gateway wan-transport-codfw --fixed-ip subnet=cloud-instances-transport1-b-codfw,ip-address=208.80.153.190 cloudinstances2b-gw ([[phab:T271517|T271517]]) * 11:02 arturo: root@cloudcontrol2001-dev:~# openstack router set --enable-snat cloudinstances2b-gw --external-gateway wan-transport-codfw ([[phab:T271517|T271517]]) * 11:01 arturo: enabling neutron hacks in codfw1dev (cloudnet2002-dev, cloudnet2003-dev) ([[phab:T271517|T271517]]) * 10:55 arturo: aborrero@labtestvirt2003:~ $ sudo ifdown eno2.2107 ([[phab:T271517|T271517]]) * 10:55 arturo: aborrero@labtestvirt2003:~ $ sudo ifdown eno2.2120 ([[phab:T271517|T271517]]) * 10:53 arturo: root@cloudcontrol2001-dev:~# openstack subnet create --network wan-transport-codfw --gateway 208.80.153.185 --ip-version 4 --network wan-transport-codfw --no-dhcp --subnet-range 208.80.153.184/29 cloud-instances-transport1-b-codfw ([[phab:T271517|T271517]]) * 10:40 dcaro: Finished tests, brining osd online (od.48) for eqiad ceph cluster ([[phab:T271417|T271417]]) * 09:59 dcaro: Started performance tests on sdc (od.48) for eqiad ceph cluster ([[phab:T271417|T271417]]) * 09:41 dcaro: Taking osd.48 from eqiad ceph cluster out to do performance tests ([[phab:T271417|T271417]]) === 2021-01-07 === * 15:19 dcaro: Finished speed tests on cloudcephosd2001-dev, reprovisioning the osd.0 sdc ([[phab:T271417|T271417]]) * 14:39 dcaro: Starting speed tests on cloudcephosd2001-dev sdc ([[phab:T271417|T271417]]) * 12:54 dcaro: Taking osd.0 down on codfw ceph cluster to try the disk performance testing process ([[phab:T271417|T271417]]) * 11:35 arturo: merging dmz_cidr change ([[phab:T209082|T209082]], [[phab:T267779|T267779]]) === 2021-01-05 === * 10:40 dcaro: removing dumps-[1..*] backups from cloudvirt1024 as they are not needed ([[phab:T271094|T271094]]) === 2021-01-03 === * 07:06 dcaro: Got a network hiccup on cloudnet1004, keeping track here [[phab:T271058|T271058]] === 2020-12-28 === * 12:32 arturo: stop doing backups for the dumps project https://gerrit.wikimedia.org/r/c/operations/puppet/+/652182 ([[phab:T260692|T260692]]) * 12:32 arturo: stop doing backups for the dumps project https://gerrit.wikimedia.org/r/c/operations/puppet/+/652182 ([[phab:T260682|T260682]]) * 12:23 arturo: icinga downtime cloudvirt1026 disk space check until january 5 ([[phab:T260692|T260692]]) * 06:15 andrewbogott: restarting designate-central on cloudservices1003/1004. I'm pretty sure they're distressed because of DB lag but it's worth a try === 2020-12-23 === * 15:38 andrewbogott: restarting rabbitmq on cloudcontrol1004; suspected leaks * 15:33 andrewbogott: restarting each cloudcontrol galera node in turn to see if that quiets down the syncing warnings * 12:08 arturo: move memory out of the swap in cloudcontrol1004 by disabling/enabling it (1Gb swap was being used) === 2020-12-22 === * 15:30 dcaro: cleaning up 6778 dangling snapshots for glance images in eqiad ([[phab:T270478|T270478]]) * 13:51 dcaro: merged patch to move wikidumpparse backups to cloudvirt1025 to free space on cloudvirt1026 === 2020-12-19 === * 16:18 dcaro: gzipped a bunch of logs on cloudvirt1004 due to / being out of space * 00:14 bstorm: truncated /var/log/debug.1 on cloudcontrol1003 which appears to be the exact same content as the user.log files anyway * 00:10 bstorm: truncated /var/log/daemon.log.1 and the haproxy log * 00:02 bstorm: truncated /var/log/messages.1 on cloudcontrol1003 === 2020-12-18 === * 23:53 bstorm: truncated haproxy.log.1 on cloudcontrol1003 * 20:46 andrewbogott: setting pg and pgp number to 4096 for eqiad1-compute as joachim thinks 8192 might be too much [[phab:T270305|T270305]] * 17:09 dcaro: finished cleaning up the dangling snapshots from cloudvirt1026 ([[phab:T270478|T270478]]) * 17:08 dcaro: removing dangling rbd snapshots (for backups on cloudvirt1026) ([[phab:T270478|T270478]]) * 17:06 dcaro: finished cleaning up the dangling snapshots from cloudvirt1025 ([[phab:T270478|T270478]]) * 17:05 dcaro: removing dangling rbd snapshots (for backups on cloudvirt1025) ([[phab:T270478|T270478]]) * 17:00 dcaro: finished cleaning up the dangling snapshots from cloudvirt1021 ([[phab:T270478|T270478]]) * 16:58 dcaro: removing dangling rbd snapshots (for backups on cloudvirt1021) ([[phab:T270478|T270478]]) * 16:56 dcaro: finished cleaning up the dangling snapshots from cloudvirt1022 ([[phab:T270478|T270478]]) * 16:55 dcaro: removing dangling rbd snapshots (for backups on cloudvirt1022) ([[phab:T270478|T270478]]) * 16:54 dcaro: finished cleaning up the dangling snapshots from cloudvirt1023 ([[phab:T270478|T270478]]) * 16:51 dcaro: removing dangling rbd snapshots (for backups on cloudvirt1023) ([[phab:T270478|T270478]]) * 16:47 dcaro: finished cleaning up the dangling snapshots from cloudvirt1024, freed ~12% of the capacity ([[phab:T270478|T270478]]) * 16:21 dcaro: removing dangling rbd snapshots (for backups on cloudvirt1024) ([[phab:T270478|T270478]]) * 16:13 andrewbogott: setting autoscale to 'off' for both ceph pools (eqiad1-compute and eqiad1-glance-images) because we like how things are set and the autoscaler does not * 10:33 dcaro: purging rbd snapshots for image fc6fb78b-4515-4dcc-8254-{{Gerrit|591b9fe01762}} ([[phab:T270478|T270478]]) === 2020-12-17 === * 22:17 andrewbogott: correction to above, set the pg and pgp to 1024 for eqiad1-glance-images * 22:16 andrewbogott: setting pgp number to 8192 for eqiad1-compute (a 4x increase) and 2048 for eqiad1-glance-images (also a 4x increase) [[phab:T270305|T270305]] (same as pg) * 22:14 andrewbogott: setting pg number to 8192 for eqiad1-compute (a 4x increase) and 2048 for eqiad1-glance-images (also a 4x increase) [[phab:T270305|T270305]] * 22:10 andrewbogott: setting autoscale to 'warn' for both ceph pools (eqiad1-compute and eqiad1-glance-images) === 2020-12-16 === * 09:31 dcaro: removing invalid backups from cloudvirt1024 (196 in total) ([[phab:T269419|T269419]]) === 2020-12-14 === * 17:42 dcaro: The removal freed ~12GB (still 100% usage :S) ([[phab:T269419|T269419]]) * 17:36 dcaro: removing invalid backups that have a valid copy ([[phab:T269419|T269419]]) * 15:43 dcaro: Merging the tagging for vm backups ([[phab:T267195|T267195]]) * 09:45 arturo: icinga downtime cloudvirt1024 for 6 days ([[phab:T269419|T269419]]) === 2020-12-13 === * 09:11 _dcaro: running backup purge script on cloudvirt1024 ([[phab:T269419|T269419]]) === 2020-12-10 === * 23:36 bstorm: cleaned up the logs for haproxy on cloudcontrol1003 by deleting all the gzipped ones and truncating the .1 file * 11:56 dcaro: Freed some space on cloudvirt1024 by running the purge script ([[phab:T269419|T269419]]) * 09:17 dcaro: removing leaked dns record discordwiki.eqiad.wmflabs (clinic duty) === 2020-12-08 === * 18:01 dcaro: Host cloudvirt1030 up and running ([[phab:T216195|T216195]]) * 15:59 dcaro: Re-imaging host cloudvirt1030 ([[phab:T216195|T216195]]) * 14:18 dcaro: Host online cloudvirt1029 ([[phab:T216195|T216195]]) * 14:13 dcaro: Host re-imaged, doing tests cloudvirt1029 ([[phab:T216195|T216195]]) * 12:14 dcaro: Re-imaging cloudvirt1029 ([[phab:T216195|T216195]]) === 2020-12-07 === * 18:33 andrewbogott: putting cloudvirt1023 back into service [[phab:T269467|T269467]] * 15:55 andrewbogott: reimaging cloudvirt1028 for [[phab:T216195|T216195]] * 14:49 dcaro: Re-imaging cloudvirt1027 ([[phab:T216195|T216195]]) === 2020-12-05 === * 00:35 andrewbogott: moving cloudvirt1023 back into maintenance because [[phab:T269467|T269467]] continues to puzzle === 2020-12-04 === * 22:33 andrewbogott: moving cloudvirt1023 back into the ceph aggregate; it doesn't need upgrades after all [[phab:T269467|T269467]] * 22:24 andrewbogott: moving cloudvirt1023 out of the ceph aggregate and into maintenance for [[phab:T269467|T269467]] * 21:06 andrewbogott: putting cloudvirt1025 and 1026 back into service because I'm pretty sure they're fixed. [[phab:T269313|T269313]] * 12:12 arturo: manually running `wmcs-purge-backups` again on cloudvirt1024 ([[phab:T269419|T269419]]) * 11:25 arturo: icinga downtime cloudvirt1024 for 6 days, to avoid paging noises ([[phab:T269419|T269419]]) * 11:25 arturo: last log line referencing cloudvirt1024 is a mistake ([[phab:T269313|T269313]]) * 11:24 arturo: icinga downtime cloudvirt1024 for 6 days, to avoid paging noises ([[phab:T269313|T269313]]) * 10:28 arturo: manually running `wmcs-purge-backups` on cloudvirt1024 ([[phab:T269419|T269419]]) * 10:23 arturo: setting expiration to 2020-12-03 to the oldest backy snapshot of every VM in cloudvirt1024 ([[phab:T269419|T269419]]) * 09:54 arturo: icinga downtime cloudvirt1025 for 6 days ([[phab:T269313|T269313]]) === 2020-12-03 === * 23:21 andrewbogott: removing all osds on cloudcephosd1004 for rebuild, [[phab:T268746|T268746]] * 21:45 andrewbogott: removing all osds on cloudcephosd1005 for rebuild, [[phab:T268746|T268746]] * 19:51 andrewbogott: removing all osds on cloudcephosd1006 for rebuild, [[phab:T268746|T268746]] * 17:01 arturo: icinga downtime cloudvirt1025 for 48h to debug network issue [[phab:T269313|T269313]] * 16:56 arturo: rebooting cloudvirt1025 to debug network issue [[phab:T269313|T269313]] * 16:38 dcaro: Rimaging cloudvirt1026 ([[phab:T216195|T216195]]) * 13:24 andrewbogott: removing all osds on cloudcephosd1008 for rebuild, [[phab:T268746|T268746]] * 02:55 andrewbogott: removing all osds on cloudcephosd1009 for rebuild, [[phab:T268746|T268746]] === 2020-12-02 === * 20:04 andrewbogott: removing all osds on cloudcephosd1010 for rebuild, [[phab:T268746|T268746]] * 17:25 arturo: [15:51] failovering neutron virtual router in eqiad1 ([[phab:T268335|T268335]]) * 15:36 arturo: conntrackd is now up and running in cloudnet1003/1004 nodes ([[phab:T268335|T268335]]) * 15:33 arturo: [codfw1dev] conntrackd is now up and running in cloudnet200x-dev nodes ([[phab:T268335|T268335]]) * 15:08 andrewbogott: removing all osds on cloudcephosd1012 for rebuild, [[phab:T268746|T268746]] * 12:41 arturo: disable puppet in all cloudnet servers to merge conntrackd change [[phab:T268335|T268335]] * 11:12 dcaro: Reset the properties for the flavor g2.cores8.ram16.disk1120 to correct quotes ([[phab:T269172|T269172]]) * 09:57 arturo: moved cloudvirts 1030, 1029, 1028, 1027, 1026, 1025 away from the 'standard' host aggregate to 'maintenance' ([[phab:T269172|T269172]]) === 2020-12-01 === * 20:06 andrewbogott: removing all osds on cloudcephosd1014 for rebuild, [[phab:T268746|T268746]] * 12:04 arturo: restarting neutron l3 agents to pick up config change * 11:48 arturo: merging change to dmz_dir, detail list of private address https://gerrit.wikimedia.org/r/c/operations/puppet/+/641977 === 2020-11-30 === * 18:12 andrewbogott: removing all osds from cloudcephosd1015 in order to investigate [[phab:T268746|T268746]] === 2020-11-29 === * 17:18 andrewbogott: cleaning up some logfiles in tools-sgecron-01 — drive is full === 2020-11-26 === * 22:58 andrewbogott: deleting /var/log/haproxy logs older than 7 days in cloudcontrol100x. We need log rotation here it seems. * 15:53 dcaro: Created private flavor g2.cores8.ram16.disk1120 for wikidumpparse ([[phab:T268190|T268190]]) === 2020-11-25 === * 19:35 bstorm: repairing ceph pg `instructing pg 6.91 on osd.117 to repair` * 09:31 _dcaro: The OSD seems to be up and running actually, though there's that misleading log, will leave it see if the cluster comes fully healthy ([[phab:T268722|T268722]]) * 08:54 _dcaro: Unsetting noup/nodown to allow re-shuffling of the pgs that osd.44 had, will try to rebuild it ([[phab:T268722|T268722]]) * 08:45 _dcaro: Tried resetting the class for osd.44 to ssd, no luck, the cluster is in noout/norebalance to avoid data shuffling (opened [[phab:T268722|T268722]]) * 08:45 _dcaro: Tried resetting the class for osd.44 to ssd, no luck, the cluster is in noout/norebalance to avoid data shuffling (opened root@cloudcephosd1005:/var/lib/ceph/osd/ceph-44# ceph osd crush set-device-class ssd osd.44) * 08:19 _dcaro: Restarting serivce osd.44 resulted on osd.44 being unable to start due to some config inconsistency (can not reset class to hdd) * 08:16 _dcaro: After enabling auto pg scaling on ceph eqiad cluster, osd.44 (cloudcephosd1005) got stuck, trying to restart the osd service * 08:16 _dcaro: After enabling auto pg scaling on ceph eqiad cluster, osd.44 (cloudcephosd1005) got stuck, trying to restart === 2020-11-22 === * 17:40 andrewbogott: apt-get upgrade on cloudservices1003/1004 * 17:32 andrewbogott: upgrading Designate on cloudservices1003/1004 to Stein === 2020-11-20 === * 12:44 arturo: [codfw1dev] install conntrackd in cloudnet2003-dev/cloudnet2002-dev to research l3 agent HA reliability * 09:26 arturo: incinga downtime labstore1006 RAID checks for 10 days ([[phab:T268281|T268281]]) === 2020-11-17 === * 19:21 andrewbogott: draining cloudvirt1012 to experiment with libvirt/cpu things === 2020-11-15 === * 11:21 arturo: icinga downtime cloudbackup2002 for 48h ([[phab:T267865|T267865]]) === 2020-11-10 === * 16:38 arturo: icinga downtime toolschecker for 2h becasue toolsdb maintenance ([[phab:T266587|T266587]]) * 11:24 arturo: [codfw1dev] enable puppet in puppetmaster01.cloudinfra-codfw1dev (disabled for unspecified reasons) === 2020-11-09 === * 12:42 arturo: restarted neutron l3 agent in cloudnet1003 bc it still had the old default route ([[phab:T265288|T265288]]) * 12:41 arturo: `root@cloudcontrol1005:~# neutron subnet-delete dcbb0f98-5e9d-4a93-8dfc-4e3ec3c44dcc` ([[phab:T265288|T265288]]) * 12:41 arturo: `root@cloudcontrol1005:~# neutron router-gateway-set --fixed-ip subnet_id=7c6bcc12-212f-44c2-9954-{{Gerrit|5c55002ee371}},ip_address=185.15.56.244 cloudinstances2b-gw wan-transport-eqiad` ([[phab:T265288|T265288]]) * 12:19 arturo: subnet 185.1.5.56.240/29 has id 7c6bcc12-212f-44c2-9954-{{Gerrit|5c55002ee371}} in neutron ([[phab:T265288|T265288]]) * 12:19 arturo: `root@cloudcontrol1005:~# neutron subnet-create --gateway 185.15.56.241 --name cloud-instances-transport1-b-eqiad1 --ip-version 4 --disable-dhcp wan-transport-eqiad 185.15.56.240/29` ([[phab:T265288|T265288]]) * 12:15 arturo: icinga-downtime toolschecker for 2h ([[phab:T265288|T265288]]) === 2020-11-02 === * 13:36 arturo: (typo: dcaro) * 13:35 arturo: added dcar as projectadmin & user ([[phab:T266068|T266068]]) === 2020-10-29 === * 16:57 bstorm: silenced deployment-prep project alerts for 60 days since the downtime expired * 08:12 arturo: force-powercycling cloudcephosd1006 === 2020-10-25 === * 16:20 andrewbogott: adding cloudvirt1038 to the 'ceph' aggregate and removing from the 'spare' aggregate. We need this space while waiting on network upgrades for empty cloudvirts ([[phab:T216195|T216195]]) === 2020-10-23 === * 11:30 arturo: [codfw1dev] openstack --os-project-id cloudinfra-codfw1dev recordset create --type PTR --record nat.cloudgw.codfw1dev.wikimediacloud.org. --description "created by hand" 0-29.57.15.185.in-addr.arpa. 1.0-29.57.15.185.in-addr.arpa. ([[phab:T261724|T261724]]) * 10:09 arturo: [codf1dev] doing DNS changes for the cloudgw PoC, including designate and https://gerrit.wikimedia.org/r/c/operations/dns/+/635965 ([[phab:T261724|T261724]]) === 2020-10-22 === * 10:46 arturo: [codfw1dev] rebooting cloudinfra-internal-puppetmaster-01.cloudinfra-codfw1dev.codfw1dev.wikimedia.cloud to try fixing some DNS weirdness * 09:43 arturo: enabling puppet in cloucontrol1003 (message said "please re-enable after 2020-10-22 06:00UTC") === 2020-10-21 === * 14:36 andrewbogott: running apt-get update && apt-get install -y facter on all cloud-vps instances * 10:31 arturo: [codfw1dev] reimaging labtestvirt2003 (cloudgw) to test puppet code ([[phab:T261724|T261724]]) * 08:56 arturo: [codfw1dev] reimaging labtestvirt2003 (cloudgw) to test puppet code ([[phab:T261724|T261724]]) === 2020-10-20 === * 15:47 arturo: changing DNS recursor ACLs (https://gerrit.wikimedia.org/r/c/operations/puppet/+/635314) this can be reverted any time if it causes problems ([[phab:T261724|T261724]]) * 14:49 arturo: [codfw1dev] reimaging labtestvirt2003 (cloudgw) to test puppet code ([[phab:T261724|T261724]]) === 2020-10-19 === * 01:41 andrewbogott: deleting all Precise base images * 01:36 andrewbogott: deleting all unused Jessie base images === 2020-10-18 === * 23:26 andrewbogott: deleting all Trusty base images * 21:50 andrewbogott: migrating all currently used ceph images to rbd === 2020-10-16 === * 09:29 arturo: [codfw1dev] still some DNS weirdness, investigating * 09:25 arturo: [codfw1dev] hard-rebooting bastion-codfw1dev-02, seems in bad shape, doesn't even wake up in the virsh console * 09:18 arturo: [codfw1dev] live-hacked cloudservices2002-dev /etc/powerdns/recursor.conf file to include cloud-codfw1dev-floating CIDR (185.15.57.0/29) while https://gerrit.wikimedia.org/r/c/operations/puppet/+/634050 is in review, so VMs with a floating IP can query the DNS recursor ([[phab:T261724|T261724]]) * 09:01 arturo: [codfw1dev] basic network connectivity seems stable after cleaning up everything related to address scopes ([[phab:T261724|T261724]]) === 2020-10-15 === * 15:17 arturo: [codfw1dev] try cleaning up anything related to address scopes in the neutron database ([[phab:T261724|T261724]]) * 13:56 arturo: [codfw1dev] drop neutron l3 agent hacks in cloudnet2002/2003-dev ([[phab:T261724|T261724]]) === 2020-10-13 === * 17:54 andrewbogott: rebuilding cloudvirt1021 for backy support * 15:22 andrewbogott: draining cloudvirt1021 so I can rebuild it with backy support * 14:19 andrewbogott: rebuilding cloudvirt1022 with backy support * 14:03 andrewbogott: draining cloudvirt1022 so I can rebuild it with backy support * 11:19 arturo: [codfw1dev] rebooting labtestvirt2003 === 2020-10-09 === * 10:15 arturo: [codfwd1ev] root@cloudcontrol2001-dev:~# openstack router set --disable-snat cloudinstances2b-gw --external-gateway wan-transport-codfw ([[phab:T261724|T261724]]) * 09:22 arturo: [codfwd1dev] rebooting cloudnet boxes for bridge and vlan changes ([[phab:T261724|T261724]]) * 09:12 arturo: [codfw1dev] root@cloudcontrol2001-dev:~# openstack subnet delete 31214392-9ca5-4256-bff5-{{Gerrit|1e19a35661de}} (cloud-instances-transport1-b-codfw - 208.80.153.184/29) ([[phab:T261724|T261724]]) * 09:10 arturo: [codfw1dev] root@cloudcontrol2001-dev:~# openstack router set --external-gateway wan-transport-codfw --fixed-ip subnet=cloud-gw-transport-codfw,ip-address=185.15.57.10 cloudinstances2b-gw ([[phab:T261724|T261724]]) * 08:49 arturo: [codfw1dev] root@cloudcontrol2001-dev:~# openstack subnet create --network wan-transport-codfw --gateway 185.15.57.9 --no-dhcp --subnet-range 185.15.57.8/30 cloud-gw-transport-codfw ([[phab:T261724|T261724]]) * 08:47 arturo: [codfw1dev] root@cloudcontrol2001-dev:~# openstack subnet delete a5ab5362-4ffb-4059-9ff7-{{Gerrit|391e22dcf3bc}} ([[phab:T261724|T261724]]) === 2020-10-08 === * 16:17 arturo: [codfw1dev] `root@cloudcontrol2001-dev:~# openstack subnet create --network wan-transport-codfw --gateway 185.15.57.8 --no-dhcp --subnet-range 185.15.57.8/31 cloud-gw-transport-codfw` (with a hack -- see task) ([[phab:T263622|T263622]]) * 16:03 arturo: [codfw1dev] briefly live-hacked python3-neutron source code in all 3 cloudcontrol2xxx-dev servers to workaround /31 network definition issue ([[phab:T263622|T263622]]) * 10:28 arturo: [codfw1dev] reimaging labtestvirt2003 (cloudgw) [[phab:T261724|T261724]] === 2020-10-06 === * 21:30 andrewbogott: moved cloudvirt1013 out of the 'ceph' aggregate and into the 'maintenance' aggregate for [[phab:T243414|T243414]] * 21:29 andrewbogott: draining cloudvirt1013 for upgrade to 10G networking * 14:45 arturo: icinga downtime every cloud* lab* host for 60 minutes for keystone maintenance === 2020-10-05 === * 17:40 bd808: `service uwsgi-labspuppetbackend restart` on cloud-puppetmaster-03 ([[phab:T264649|T264649]]) === 2020-10-02 === * 11:05 arturo: [codfw1dev] restarting rabbitmq-server in all 3 control nodes, the l3 agent was misbehaving * 09:16 arturo: [codfw1dev] trying the labtestvirt2003 (cloudgw) reimage again ([[phab:T261724|T261724]]) === 2020-10-01 === * 16:06 arturo: rebooting cloudvirt1024 to validate changes to /etc/network/interfaces file * 15:36 arturo: [codfw1dev] reimaging labtestvirt2003 === 2020-09-30 === * 16:47 andrewbogott: rebooting cloudvir1032, 1033, 1034 for [[phab:T262979|T262979]] * 13:28 arturo: enable puppet, reboot and pool back cloudvirt1031 * 13:27 arturo: extend icinga downtimes for another 120 mins * 13:15 arturo: `aborrero@cloudcontrol1003:~$ sudo nova-manage placement sync_aggregates` after reading a hint in nova-api.log * 13:02 arturo: rebooting cloudvirt1016 and moving it to the ceph host aggregate * 12:55 arturo: rebooting cloudvirt1014 and moving it to the ceph host aggregate * 12:51 arturo: rebooting cloudvirt1013 and moving it to the ceph host aggregate * 12:39 arturo: root@cloudcontrol1005:~# openstack aggregate add host maintenance cloudvirt1031 * 12:36 arturo: rebooted cloudnet1003 (active) a couple of minutes ago * 12:36 arturo: move cloudvirt1012 and cloudvirt1039 to the ceph aggregate * 11:49 arturo: rebooting cloudvirt1039 * 11:46 arturo: rebooting cloudvirt1012 * 11:40 arturo: rebooting cloudnet1004 (standby) to pick up https://gerrit.wikimedia.org/r/c/operations/puppet/+/631167 ([[phab:T262979|T262979]]) * 11:38 arturo: [codfw1dev] rebooting cloudnet2002-dev to pick up https://gerrit.wikimedia.org/r/c/operations/puppet/+/631167 * 11:36 arturo: [codfw1dev] rebooting cloudnet2003-dev to pick up https://gerrit.wikimedia.org/r/c/operations/puppet/+/631167 * 11:33 arturo: disabling puppet and downtiming every virt/net server in the fleet in preparation for merging https://gerrit.wikimedia.org/r/c/operations/puppet/+/631167 ([[phab:T262979|T262979]]) * 09:32 arturo: rebooting cloudvirt1012 to investigate linuxbridge agent issues === 2020-09-29 === * 15:40 arturo: downgrade linux kernel from linux-image-4.19.0-11-amd64 to linux-image-4.19.0-10-amd64 on cloudvirt1012 * 14:47 arturo: rebooting cloudvirt1012, chasing config weirdness in the linuxbridge agent * 14:05 andrewbogott: reimaging 1014 over and over in an attempt to get partman right * 13:51 arturo: rebooting cloudvirt1012 === 2020-09-28 === * 14:55 arturo: [jbond42] upgraded facter to v3 across the VM fleet * 13:54 andrewbogott: moving cloudvirt1035 from aggregate 'spare' to 'ceph'. We're going to need all the capacity we can get while converting older cloudvirts to ceph === 2020-09-24 === * 15:47 arturo: stopping/restarting rabbitmq-server in all cloudcontrol servers * 15:45 arturo: restarting rabbitmq-server in cloudcontrol103 * 15:15 arturo: restarting floating_ip_ptr_records_updater.service in all 3 cloudcontrol servers to reset state after a DNS failure === 2020-09-18 === * 10:16 arturo: cloudvirt1039 libvirtd service issues were fixed with a reboot * 09:56 arturo: rebooting cloudvirt1039 (spare) to try to fix some weird libvirtd failure * 09:50 arturo: enabling puppet in cloudvirts and effectively merging patches from [[phab:T262979|T262979]] * 08:59 arturo: disable puppet in all buster cloudvirts (cloudvirt[1024,1031-1039].eqiad.wmnet) to merge a patch for [[phab:T263205|T263205]] and [[phab:T262979|T262979]] * 08:50 arturo: installing iptables from buster-bpo in cloudvirt1036 ([[phab:T263205|T263205]] and [[phab:T262979|T262979]]) === 2020-09-15 === * 20:32 andrewbogott: rebooting cloudvirt1038 to see if it resolves [[phab:T262979|T262979]] * 13:58 andrewbogott: draining cloudvirt1002 with wmcs-ceph-migrate === 2020-09-14 === * 14:21 andrewbogott: draining cloudvirt1001, migrating all VMs with wmcs-ceph-migrate * 10:41 arturo: [codfw1dev] trying to get the bonding working for labtestvirt2003 ([[phab:T261724|T261724]]) * 09:47 arturo: installed qemu security update in eqiad1 cloudvirts ([[phab:T262386|T262386]]) * 09:43 arturo: [codfw1dev] installed qemu security update in codfw1dev cloudvirts ([[phab:T262386|T262386]]) === 2020-09-09 === * 18:13 andrewbogott: restarting ceph-mon@cloudcephmon1003 in hopes that the slow ops reported are phantoms * 18:01 andrewbogott: restarting ceph-mgr@cloudcephmon1003 in hopes that the slow ops reported are phantoms (https://lists.ceph.io/hyperkitty/list/ceph-users@ceph.io/thread/EOWNO3MDYRUZKAK6RMQBQ5WBPQNLHOPV/) * 17:40 andrewbogott: giving ceph pg autoscale another chance: ceph osd pool set eqiad1-compute pg_autoscale_mode on * 00:05 bd808: Running wmcs-novastats-dnsleaks ([[phab:T262359|T262359]]) === 2020-09-08 === * 21:48 bd808: Renamed FQDN prefixes to wikimedia.cloud scheme in cloudinfra-db01's labspuppet db ([[phab:T260614|T260614]]) * 14:29 andrewbogott: restarting nova-compute on all cloudvirts (everyone is upset from the reset switch failure) * 14:18 arturo: restarting nova-fullstack service in cloudcontrol1003 * 14:17 andrewbogott: stopping apache2 on labweb1001 to make sure the Horizon outage is total === 2020-09-03 === * 09:31 arturo: icinga downtime cloud* servers for 30 mins ([[phab:T261866|T261866]]) === 2020-09-02 === * 08:46 arturo: [codfw1dev] reimaging spare server labtestvirt2003 as debian buster ([[phab:T261724|T261724]]) === 2020-09-01 === * 18:18 andrewbogott: adding drives on cloudcephosd100[3-5] to ceph osd pool * 13:40 andrewbogott: adding drives on cloudcephosd101[0-2] to ceph osd pool * 13:35 andrewbogott: adding drives on cloudcephosd100[1-3] to ceph osd pool * 11:27 arturo: [codfw1dev] rebooting again cloudnet2002-dev after some network tests, to reset initial state ([[phab:T261724|T261724]]) * 11:09 arturo: [codfw1dev] rebooting cloudnet2002-dev after some network tests, to reset initial state ([[phab:T261724|T261724]]) * 10:49 arturo: disable puppet in cloudnet servers to merge https://gerrit.wikimedia.org/r/c/operations/puppet/+/623569/ === 2020-08-31 === * 23:26 bd808: Removed stale lockfile at cloud-puppetmaster-03.cloudinfra.eqiad.wmflabs:/var/lib/puppet/volatile/GeoIP/.geoipupdate.lock * 11:20 arturo: [codfw1dev] livehacking https://gerrit.wikimedia.org/r/c/operations/puppet/+/615161 in the puppetmasters for tests before merging === 2020-08-28 === * 20:12 bd808: Running `wmcs-novastats-dnsleaks --delete` from cloudcontrol1003 === 2020-08-26 === * 17:12 bstorm: Running 'ionice -c 3 nice -19 find /srv/tools -type f -size +100M -printf "%k KB %p\n" > tools_large_files_20200826.txt' on labstore1004 [[phab:T261336|T261336]] === 2020-08-21 === * 21:34 andrewbogott: restarting nova-compute on cloudvirt1033; it seems stuck === 2020-08-19 === * 14:21 andrewbogott: rebooting cloudweb2001-dev, labweb1001, labweb1002 to address mediawiki-induced memleak === 2020-08-06 === * 21:02 andrewbogott: removing cloudvirt1004/1006 from nova's list of hypervisors; rebuilding them to use as backup test hosts * 20:06 bstorm: manually stopped the RAID check on cloudcontrol1003 [[phab:T259760|T259760]] === 2020-08-04 === * 18:54 bstorm: restarting mariadb on cloudcontrol1004 to setup parallel replication === 2020-08-03 === * 17:02 bstorm: increased db connection limit to 800 across galera cluster because we were clearly hovering at limit === 2020-07-31 === * 19:28 bd808: wmcs-novastats-dnsleaks --delete (lots of leaked fullstack-monitoring records to clean up) === 2020-07-27 === * 22:17 andrewbogott: ceph osd pool set compute pg_num 2048 * 22:14 andrewbogott: ceph osd pool set compute pg_autoscale_mode off === 2020-07-24 === * 19:15 andrewbogott: ceph mgr module enable pg_autoscaler * 19:15 andrewbogott: ceph osd pool set compute pg_autoscale_mode on === 2020-07-22 === * 08:55 jbond42: [codfw1dev] upgrading hiera to version5 * 08:48 arturo: [codfw1dev] add jbond as user in the bastion-codfw1dev and cloudinfra-codfw1dev projects * 08:45 arturo: [codfw1dev] enabled account creation in labtestwiki briefly for jbond42 to create an account === 2020-07-16 === * 10:48 arturo: merging change to neutron dmz_cidr https://gerrit.wikimedia.org/r/c/operations/puppet/+/613123 ([[phab:T257534|T257534]]) === 2020-07-15 === * 23:15 bd808: Removed Merlijn van Deen from toollabs-trusted Gerrit group ([[phab:T255697|T255697]]) * 11:48 arturo: [codfw1dev] created DNS records (A and PTR) for bastion.bastioninfra-codfw1dev.codfw1dev.wmcloud.org <-> 185.15.57.2 * 11:41 arturo: [codfw1dev] add myself as projectadmin to the `bastioninfra-codfw1dev` project * 11:39 arturo: [codfw1dev] created DNS zone `bastioninfra-codfw1dev.codfw1dev.wmcloud.org.` in the cloudinfra-codfw1dev project and then transfer ownership to the bastioninfra-codfw1dev project === 2020-07-14 === * 15:19 arturo: briefly set root@cloudnet1003:~ # sysctl net.ipv4.conf.all.accept_local=1 (in neutron qrouter netns) ([[phab:T257534|T257534]]) * 10:43 arturo: icinga downtime cloudnet* hosts for 30 mins to introduce new check https://gerrit.wikimedia.org/r/c/operations/puppet/+/612390 ([[phab:T257552|T257552]]) * 04:01 andrewbogott: added a wildcard *.wmflabs.org domain pointing at the domain proxy in project-proxy * 04:00 andrewbogott: shortened the ttl on .wmflabs.org. to 300 === 2020-07-13 === * 16:17 arturo: icinga downtime cloudcontrol[1003-1005].wikimedia.org for 1h for galera database movements === 2020-07-12 === * 17:39 andrewbogott: switched eqiad1 keystone from m5 to cloudcontrol galera === 2020-07-10 === * 20:26 andrewbogott: disabling nova api to move database to galera === 2020-07-09 === * 11:23 arturo: [codfw1dev] rebooting cloudnet2003-dev again for testing sysct/puppet behavior ([[phab:T257552|T257552]]) * 11:11 arturo: [codfw1dev] rebooting cloudnet2003-dev for testing sysct/puppet behavior ([[phab:T257552|T257552]]) * 09:16 arturo: manually increasing sysctl value of net.nf_conntrack_max in cloudnet servers ([[phab:T257552|T257552]]) === 2020-07-06 === * 15:16 arturo: installing 'aptitude' in all cloudvirts === 2020-07-03 === * 12:51 arturo: [codfw1dev] galera cluster should be up and running, openstack happy ([[phab:T256283|T256283]]) * 11:44 arturo: [codfw1dev] restoring glance database backup from bacula into cloudcontrol2001-dev ([[phab:T256283|T256283]]) * 11:39 arturo: [codfw1dev] stopped mysql database in the galera cluster [[phab:T256283|T256283]] * 11:36 arturo: [codfw1dev] dropped glance database in the galera cluster [[phab:T256283|T256283]] === 2020-07-02 === * 15:41 arturo: `sudo wmcs-openstack --os-compute-api-version 2.55 flavor create --private --vcpus 8 --disk 300 --ram 16384 --property aggregate_instance_extra_specs:ceph=true --description "for packaging envoy" bigdisk-ceph` ([[phab:T256983|T256983]]) === 2020-06-29 === * 14:24 arturo: starting rabbitmq-server in all 3 cloudcontrol servers * 14:23 arturo: stopping rabbitmq-server in all 3 cloudcontrol servers === 2020-06-18 === * 20:38 andrewbogott: rebooting cloudservices2003-dev due to a mysterious 'host down' alert on a secondary ip === 2020-06-16 === * 15:38 arturo: created by hand neutron port 9c0a9a13-e409-49de-9ba3-{{Gerrit|bc8ec4801dbf}} `paws-haproxy-vip` ([[phab:T295217|T295217]]) === 2020-06-12 === * 13:23 arturo: DNS zone `paws.wmcloud.org` transferred to the PAWS project ([[phab:T195217|T195217]]) * 13:20 arturo: created DNS zone `paws.wmcloud.org` ([[phab:T195217|T195217]]) === 2020-06-11 === * 19:19 bstorm_: proceeding with failback to labstore1004 now that DRBD devices are consistent [[phab:T224582|T224582]] * 17:22 bstorm_: delaying failback labstore1004 for drive syncs [[phab:T224582|T224582]] * 17:17 bstorm_: failing NFS back to labstore1004 to complete the upgrade process [[phab:T224582|T224582]] * 16:15 bstorm_: failing over NFS for labstore1004 to labstore1005 [[phab:T224582|T224582]] === 2020-06-10 === * 16:09 andrewbogott: deleting all old cloud-ns0.wikimedia.org and cloud-ns1.wikimedia.org ns records in designate database [[phab:T254496|T254496]] === 2020-06-09 === * 15:25 arturo: icinga downtime everything cloud* lab* for 2h more ([[phab:T253780|T253780]]) * 14:09 andrewbogott: stopping puppet, all designate services and all pdns services on cloudservices1004 for [[phab:T253780|T253780]] * 14:01 arturo: icinga downtime everything cloud* lab* for 2h ([[phab:T253780|T253780]]) === 2020-06-05 === * 15:08 andrewbogott: trying to re-enable puppet without losing cumin contact, as per https://phabricator.wikimedia.org/T254589 === 2020-06-04 === * 14:24 andrewbogott: disabling puppet on all instances for /labs/private recovery * 14:23 arturo: disabling puppet on all instances for /labs/private recovery === 2020-05-28 === * 23:02 bd808: `/usr/local/sbin/maintain-dbusers --debug harvest-replicas` ([[phab:T253930|T253930]]) * 13:36 andrewbogott: rebuilding cloudservices2002-dev with Buster * 00:33 andrewbogott: shutting down cloudservices2002-dev to see if we can live without it. This is in anticipation or rebuilding it entirely for [[phab:T253780|T253780]] === 2020-05-27 === * 23:29 andrewbogott: disabling the backup job on cloudbackup2001 (just like last week) so the backup doesn't start while Brooke is rebuilding labstore1004 tomorrow. * 06:03 bd808: `systemctl start mariadb` on clouddb1001 following reboot (take 2) * 05:58 bd808: `systemctl start mariadb` on clouddb1001 following reboot * 05:53 bd808: Hard reboot of clouddb1001 via Horizon. Console unresponsive. === 2020-05-25 === * 16:35 arturo: [codfw1dev] created zone `0-29.57.15.185.in-addr.arpa.` ([[phab:T247972|T247972]]) === 2020-05-21 === * 19:23 andrewbogott: disabling puppet on cloudbackup2001 to prevent the backup job from starting during maintenance * 19:16 andrewbogott: systemctl disable block_sync-tools-project.service on cloudbackup2001.codfw.wmnet to avoid stepping on current upgrade * 15:48 andrewbogott: re-imaging cloudnet1003 with Buster === 2020-05-19 === * 22:59 bd808: `apt-get install mariadb-client` on cloudcontrol1003 * 21:12 bd808: Migrating wcdo.wcdo.eqiad.wmflabs to cloudvirt1023 ([[phab:T251065|T251065]]) === 2020-05-18 === * 21:37 andrewbogott: rebuilding cloudnet2003-dev with Buster === 2020-05-15 === * 22:10 bd808: Added reedy as projectadmin in cloudinfra project ([[phab:T249774|T249774]]) * 22:05 bd808: Added reedy as projectadmin in admin project ([[phab:T249774|T249774]]) * 18:44 bstorm_: rebooting cloudvirt-wdqs1003 [[phab:T252831|T252831]] * 15:47 bd808: Manually running wmcs-novastats-dnsleaks from cloudcontrol1003 ([[phab:T252889|T252889]]) === 2020-05-14 === * 23:28 bstorm_: downtimed cloudvirt1004/6 and cloudvirt-wdqs1003 until tomorrow around this time [[phab:T252831|T252831]] * 22:21 bstorm_: upgrading qemu-system-x86 on cloudvirt1006 to backports version [[phab:T252831|T252831]] * 22:15 bstorm_: changing /etc/libvirt/qemu.conf and restarting libvirtd on cloudvirt1006 [[phab:T252831|T252831]] * 21:12 andrewbogott: rebuilding cloudvirt1003-wdqs as part of [[phab:T252831|T252831]] * 15:47 andrewbogott: moving cloudvirt1004 and cloudvirt1006 to the 'ceph' aggregate for [[phab:T252784|T252784]] * 15:02 andrewbogott: moving all of cloudvirt100[1-9] into the 'toobusy' host aggregate. These are slower, have spinning disks, and are due for replacement. === 2020-05-12 === * 20:33 andrewbogott: moving cloudvirt1023 to the 'standard' pool and out of the 'spare' pool * 19:10 jeh: disable neutron-openvswitch-agent service on cloudvirt2001-dev.codfw [[phab:T248881|T248881]] * 19:09 jeh: Shutdown the unused eno2 network interface on cloudvirt2001-dev.codfw to clear up monitoring errors [[phab:T248425|T248425]] * 18:20 andrewbogott: moving cloudvirt1024 out of the 'maintenance' aggregate and into 'spare' * 16:45 andrewbogott: restarting neutron-l3-agent on cloudnet1004 so it knows about all three cloudcontrols. Leaving cloudnet1003 since restarting it there will cause network interruptions * 14:06 arturo: icinga downtime everything for 2h for Debian Buster migration in some cloud components === 2020-05-09 === * 16:53 andrewbogott: rebuilding cloudcontrol2001-dev and 2003-dev with buster for [[phab:T252121|T252121]] === 2020-05-08 === * 19:02 bstorm_: moving tools-k8s-haproxy-2 from cloudvirt1021 to cloudvirt1017 to improve spread === 2020-05-05 === * 13:58 andrewbogott: rebuilding cloudcontrol2004-dev to test new puppet changes === 2020-05-04 === * 09:04 arturo: [codfw1dev] manually modify iptables ruleset to only allow SSH from WMF bastions on cloudservices2003-dev and cloudcontrol2004-dev ([[phab:T251604|T251604]]) === 2020-04-21 === * 22:12 andrewbogott: moving cloudvirt1004 out of the 'standard' aggregate and into the 'maintenance' aggregate * 16:01 jeh: restart cloudceph mon and osd services for openssl upgrades === 2020-04-15 === * 18:44 jeh: create indexes and views for grwikimedia [[phab:T245912|T245912]] === 2020-04-13 === * 15:07 jeh: restart memcached on labwebs to increase cache size [[phab:T145703|T145703]] === 2020-04-09 === * 19:57 andrewbogott: upgrading eqiad1 designate to rocky * 16:52 andrewbogott: cleaned up a bunch of leaked .eqiad.wmflabs dns records === 2020-04-08 === * 19:20 andrewbogott: rotated password and api token for pdns servers on cloudservices1003 and cloudservices1004 * 14:54 arturo: `root@cloudcontrol1003:~# cp /etc/inputrc .inputrc` to solve some bash shortcut weirdness === 2020-04-07 === * 20:57 andrewbogott: service sssd stop; rm -rf /var/lib/sss/db*; service sssd start on tools-sgebastion-08 === 2020-04-06 === * 22:39 andrewbogott: deleting bogus groups cn=b'project-bastion',ou=groups,dc=wikimedia,dc=org and cn=b'project-tools',ou=groups,dc=wikimedia,dc=org from ldap * 17:42 arturo: [codfw1dev] transferred DNS zone 57.15.185.in-addr.arpa. to the cloudinfra-codfw1dev project ([[phab:T247972|T247972]]) * 17:39 arturo: [codfw1dev] `openstack zone create --email root@wmflabs.org --type PRIMARY --ttl 3600 --description "floating IPs subnet" 57.15.185.in-addr.arpa.` ([[phab:T247972|T247972]]) * 16:23 arturo: restarting apache2 in cloudcontrol1003/1004 to pick up latest wmfkeystonehooks changes [[phab:T249494|T249494]] === 2020-04-02 === * 20:59 jeh: codfw1dev clear VM error states and start bastions, puppet master and database === 2020-04-01 === * 16:27 arturo: [codfw1dev] enable puppet across the fleet clean vxlan changes ([[phab:T248881|T248881]]) === 2020-03-31 === * 12:35 arturo: [codfw1dev] restarting VMs: designaterockytest14, bastion-codfw1dev-0[1,2] ([[phab:T248881|T248881]]) * 12:34 arturo: [codfw1dev] installing neutron-openvswitch-agent on cloudvirt2001-dev ([[phab:T248881|T248881]]) * 12:25 arturo: [codfw1dev] installing neutron-openvswitch-agent on cloudnet200[2,3]-dev ([[phab:T248881|T248881]]) * 11:45 arturo: [codfw1dev] rebooting cloudvirt2003-dev to pick up latest kernel update. Otherwise modprobe is confused trying to load modules and openvswitch won't start ([[phab:T248881|T248881]]) * 10:40 arturo: [codfw1dev] installing neutron-openvswitch-agent on cloudvirt2003-dev ([[phab:T248881|T248881]]) * 10:09 arturo: [codfw1dev] reboot cloudnet2003-dev into linux 4.9 (was using 4.14 from a testing operation in 2020-03-10) === 2020-03-30 === * 23:42 bstorm_: deleted "Kubernetes Cluster" and "Kubernetes Performance" dashboards [[phab:T246689|T246689]] * 16:44 arturo: [codfw1dev] installing package neutron-openvswitch-agent in cloudvirt2002-dev ([[phab:T248881|T248881]]) * 16:42 andrewbogott: restarting l3 agents on cloudnets in codfw1dev after applying https://gerrit.wikimedia.org/r/#/c/operations/puppet/+/584188/ === 2020-03-27 === * 21:28 bd808: Created huggle.wmcloud.org Designate zone and allocated it to the huggle project * 19:51 jeh: start haproxy on cloudcontrol2003-dev.wikimedia.org === 2020-03-26 === * 15:01 arturo: icinga downtime cloudvirt* cloudcontrol* cloudnet* lab* cloudstore* * 15:01 andrewbogott: beginning openstack upgrade window for [[phab:T242766|T242766]] * 12:32 arturo: [codfw1dev] downgraded systemd, libsystemd0, udev and friends to the non-backports versions ([[phab:T247013|T247013]]) === 2020-03-25 === * 19:29 andrewbogott: dumping a bunch of VMs on cloudvirt1015 to see if it still crashes * 17:56 jeh: add labweb1002 back into the pool - completed horizon testing [[phab:T240852|T240852]] * 17:09 jeh: depool labweb1002 for horizon testing [[phab:T240852|T240852]] === 2020-03-24 === * 19:41 jeh: switch cloudvirt1016 from maintenance to standard host aggregate [[phab:T243327|T243327]] * 15:31 andrewbogott: restarting nova-conductor and nova-api on cloudcontrol1003 and cloudcontrol1004 === 2020-03-23 === * 21:41 jeh: restart neutron-l3-agent on cloudnet100[3,4] to pickup policy.yaml changes * 13:28 jeh: disable puppet on labweb100[1,2] to enable horizon event traces [[phab:T240852|T240852]] * 10:26 arturo: restarting apache in both labweb1001/labweb1002 upon reports of returning 500s === 2020-03-21 === * 14:23 andrewbogott: restarting apache2 on labweb1001 and 1002 === 2020-03-18 === * 19:17 andrewbogott: deleted a bunch of records from the pdns database on cloudservices1003/1004 which had a record name but the content (where an IP address should be) was NULL, e.g. m.wikidata.beta.wmflabs.org. * 10:55 arturo: [codfw1dev] deleting BGP agent, undoing changes we did for [[phab:T245606|T245606]] === 2020-03-14 === * 17:40 jeh: restart maintain-dbusers on labstore1004 [[phab:T247654|T247654]] === 2020-03-13 === * 12:39 arturo: [codfw1dev] reintroduce address scopes for another round of testing [[phab:T244851|T244851]] * 12:17 arturo: [codfw1dev] enabling puppet in cloudnet200x-dev servers after merging https://gerrit.wikimedia.org/r/c/operations/puppet/+/579259 ([[phab:T247505|T247505]]) === 2020-03-12 === * 22:29 bstorm_: running puppet across all dumps mounts to make sure active links are shifted to labstore1006 === 2020-03-11 === * 18:38 jeh: set icingia downtime until 2020-03-23 on CODFW cloud[control,net,virt] hosts during openstack upgrades * 12:50 arturo: [codfw1dev] several tests creating/deleting address scopes ([[phab:T244727|T244727]] [[phab:T247135|T247135]] [[phab:T246887|T246887]] [[phab:T245606|T245606]]) * 12:46 arturo: [codfw1dev] disable routing_source_ip in l3 agents for testing proposal detailed at https://wikitech.wikimedia.org/wiki/Wikimedia_Cloud_Services_team/EnhancementProposals/Network_refresh#Eliminate_routing_source_ip_address ([[phab:T244727|T244727]]) === 2020-03-10 === * 17:02 arturo: [codfw1dev] deleting address scopes, bad interaction with our custom NAT setup [[phab:T247135|T247135]] * 13:55 arturo: [codfw1dev] rebooting cloudnet2003-dev into linux kernel 4.14 for testing stuff related to [[phab:T247135|T247135]] === 2020-03-09 === * 18:09 arturo: enabling puppet in cloudvirt1006, all services have been restored * 17:59 arturo: deleted the neutron bridge on cloudvirt1006, for testing stuff related to the queens upgrade * 17:58 arturo: stopped neutron-linuxbridge-agent and nova-compute in cloudvirt1006 for testing stuff related to the queens upgrade === 2020-03-06 === * 14:54 andrewbogott: draining all instances off of cloudvirt1006 for [[phab:T246908|T246908]] === 2020-03-05 === * 14:24 arturo: [codfw1dev] we just enabled BGP session between cloudnet2xxx-dev and cr1-codfw ([[phab:T245606|T245606]]) * 13:07 arturo: [codfw1dev] move the extra IP address for BGP in cloudnet200x-dev servers from eno2.2120 to the br-external bridge device ([[phab:T245606|T245606]]) * 13:06 arturo: [codfw1dev] upgrade neutron-dynamic-routing packages in cloudnet200X-dev and cloudcontrol200X-dev servers to 11.0.0-2~bpo9+1 ([[phab:T245606|T245606]]) === 2020-03-04 === * 22:22 andrewbogott: upgrading designate on cloudservices1003/1004 to Queens * 22:09 andrewbogott: moving cloudvirt1006 into the maintenance aggregate for [[phab:T246908|T246908]] * 21:37 bd808: Running wmcs-wikireplica-dns to add service names for ngwikimedia.*.db.svc.eqiad.wmflabs ([[phab:T240772|T240772]]) * 21:14 bd808: Running `sudo maintain-meta_p --all-databases --purge` on labsdb1009 ([[phab:T246056|T246056]]) * 21:11 bd808: Running `sudo maintain-meta_p --all-databases --purge` on labsdb1010 ([[phab:T246056|T246056]]) * 21:08 bd808: Running `sudo maintain-meta_p --all-databases --purge` on labsdb1011 ([[phab:T246056|T246056]]) * 21:05 bd808: Running `sudo maintain-meta_p --all-databases --purge` on labsdb1002 ([[phab:T246056|T246056]]) === 2020-03-02 === * 16:54 arturo: [codfw1dev] deleted python3-os-ken debian package in cloudnet2003-dev which was installed by hand and had depedency issues === 2020-02-29 === * 16:32 bstorm_: downtimed the smart alert on cloudvirt1009 until Monday since apparently predictive failures flap [[phab:T244986|T244986]] === 2020-02-26 === * 22:03 jeh: powering down cloudvirt1014 for hardware maintenance === 2020-02-25 === * 16:08 andrewbogott: changing neutron's rabbitmq password because oslo is having trouble parsing some of the characters in the password * 15:26 andrewbogott: updated the cell_mapping record in the nova_api database to add the second rabbitmq server to the transport_url field * 15:26 andrewbogott: updated the cell_mapping record in the nova_api database to set the db uri to 'mysql+pymysql' -- this in response to a deprecation notice === 2020-02-24 === * 12:16 arturo: [codfw1dev] `root@cloudcontrol2001-dev:~# neutron bgp-speaker-peer-add bgpspeaker cr2-codfw` ([[phab:T245606|T245606]]) * 12:16 arturo: [codfw1dev] `root@cloudcontrol2001-dev:~# neutron bgp-speaker-peer-add bgpspeaker cr1-codfw` ([[phab:T245606|T245606]]) * 12:09 arturo: [codfw1dev] `root@cloudcontrol2001-dev:~# neutron bgp-peer-create --peer-ip 208.80.153.187 --remote-as 65002 cr2-codfw` ([[phab:T245606|T245606]]) * 12:09 arturo: [codfw1dev] `root@cloudcontrol2001-dev:~# neutron bgp-peer-create --peer-ip 208.80.153.186 --remote-as 65002 cr1-codfw` ([[phab:T245606|T245606]]) * 12:06 arturo: [codfw1dev] `root@cloudcontrol2001-dev:~# neutron bgp-peer-delete 17b8c2a3-f0ce-4d50-a265-18ccac703c61` ([[phab:T245606|T245606]]) * 10:59 arturo: [codfw1dev] `root@cloudcontrol2001-dev:~# neutron bgp-speaker-peer-add bgpspeaker bgppeer` ([[phab:T245606|T245606]]) * 10:56 arturo: [codfw1dev] `root@cloudcontrol2001-dev:~# neutron bgp-peer-create --peer-ip 208.80.153.185 --remote-as 65002 bgppeer` ([[phab:T245606|T245606]]) === 2020-02-21 === * 12:48 arturo: [codfw1dev] running `root@cloudcontrol2001-dev:~# neutron bgp-speaker-network-add bgpspeaker wan-transport-codfw` ([[phab:T245606|T245606]]) * 12:46 arturo: [codfw1dev] created bgpspeaker for AS64711 ([[phab:T245606|T245606]]) * 12:42 arturo: [codfw1dev] run `sudo neutron-db-manage upgrade head` to upgrade the db schema for neutron bgp tables * 11:51 arturo: [codfw1dev] create a neutron subnet pool per each subnet objects we have and manually update DB to inter-associate them ([[phab:T245606|T245606]]) * 11:49 arturo: [codfw1dev] rename neutron address scope `no-nat` to `bgp` ([[phab:T245606|T245606]]) * 11:37 arturo: [codfw1dev] cleanup unused neutron subnet pools from previous address scope tests ([[phab:T244851|T244851]]) === 2020-02-20 === * 19:22 andrewbogott: updating designate pool config for https://gerrit.wikimedia.org/r/#/c/operations/puppet/+/572213/ * 15:33 andrewbogott: migrating all VMs on cloudvirt1014 to cloudvirt1022 * 13:35 arturo: [codfw1dev] disable puppet in cloudcontrol servers to hack neutron.conf for tests related to [[phab:T245606|T245606]] * 13:33 arturo: [codfw1dev] disable puppet in cloudnet servers to hack neutron.conf for tests related to [[phab:T245606|T245606]] === 2020-02-18 === * 22:19 andrewbogott: transferred the tools.wmcloud.org. to the tools project * 22:16 andrewbogott: moved wmcloud.org dns domain to the cloud-infra project * 21:02 andrewbogott: adding .eqiad1.wikimedia.cloud records to all existing eqiad1 VMs, updating all eqiad1 internal pointer records to reference the new eqiad1.wikimedia.cloud fqdns. * 09:44 arturo: deleted DNS zone wmcloud.org and try re-creating it === 2020-02-14 === * 10:35 arturo: running `root@cloudcontrol2001-dev:~# designate server-create --name ns1.openstack.codfw1dev.wikimediacloud.org.` ([[phab:T243766|T243766]]) * 10:32 arturo: running `root@cloudcontrol1004:~# designate server-create --name ns1.openstack.eqiad1.wikimediacloud.org.` ([[phab:T243766|T243766]]) * 10:32 arturo: running `root@cloudcontrol1004:~# designate server-create --name ns0.openstack.eqiad1.wikimediacloud.org.` ([[phab:T243766|T243766]]) === 2020-02-12 === * 13:38 arturo: [codfw1dev] add reference to subnetpool to the instance subnet `MariaDB [neutron]> update subnets set subnetpool_id='d129650d-d4be-4fe1-b13e-6edb5565cb4a' where id = '7adfcebe-b3d0-4315-92fe-e8365cc80668';` ([[phab:T244851|T244851]]) === 2020-02-11 === * 13:46 arturo: [codfw1dev] creating some neutron objects to investigate [[phab:T244851|T244851]] (subnets, subnet pools, address scopes, ...) * 12:40 arturo: [codfw1dev] delete unknown address scope 'wmcs-v4-scope': `root@cloudcontrol2001-dev:~# openstack address scope delete 078cfd71-117b-4aac-9197-6ebbbb7dd3de` ([[phab:T244851|T244851]]) * 12:40 arturo: [codfw1dev] delete unknown subnet pool 'cloudinstancesb-v4-pool0': `root@cloudcontrol2001-dev:~# openstack subnet pool delete d23a9b88-5c3d-4a53-ab88-053233a75365` ([[phab:T244851|T244851]]) === 2020-02-07 === * 18:11 jeh: shutdown cloudvirt1016 for hardware maintenance [[phab:T241882|T241882]] === 2020-02-06 === * 14:44 jeh: update apt packages on cloudvirt1015 [[phab:T220853|T220853]] * 14:28 jeh: run hardware tests on cloudvirt1015 [[phab:T220853|T220853]] === 2020-01-28 === * 17:24 arturo: [codfw1dev] root@cloudcontrol2001-dev:~# designate server-create --name ns0.openstack.codfw1dev.wikimediacloud.org. ([[phab:T243766|T243766]]) * 10:18 arturo: [codfw1dev] created DNS record `bastion-codfw1dev-01.codfw1dev.wmcloud.org A 185.15.57.2` ([[phab:T242976|T242976]], [[phab:T229441|T229441]]) * 10:13 arturo: [codfw1dev] the zone `codfw1dev.wmcloud.org` belongs now to the `cloudinfra-codfw1dev` project ([[phab:T242976|T242976]]) * 10:11 arturo: [codfw1dev] `root@cloudcontrol2001-dev:~# openstack zone create --description "main DNS domain for public addresses" --email "root@wmflabs.org" --type PRIMARY --ttl 3600 codfw1dev.wmcloud.org.` ([[phab:T242976|T242976]] and [[phab:T243766|T243766]]) * 09:53 arturo: restart apache2 in labweb1001/1002 because horizon errors * 09:47 arturo: created DNS zone wmcloud.org in eqiad1, transfer it to the cloudinfra project ([[phab:T242976|T242976]]) right now only use is to delegate codfw1dev.wmcloud.org subdomain to designate in the other deployment === 2020-01-27 === * 12:45 arturo: [codfw1dev] manually move the new domain to the `cloudinfra-codfw1dev` project clouddb2001-dev: `[designate]> update zones set tenant_id='cloudinfra-codfw1dev' where id = '4c75410017904858a5839de93c9e8b3d';` [[phab:T243556|T243556]] * 12:44 arturo: [codfw1dev] `root@cloudcontrol2001-dev:~# openstack zone create --description "main DNS domain for VMs" --email "root@wmflabs.org" --type PRIMARY --ttl 3600 codfw1dev.wikimedia.cloud.` [[phab:T243556|T243556]] === 2020-01-24 === * 15:10 jeh: remove icinga downtime for cloudvirt1013 [[phab:T241313|T241313]] * 12:52 arturo: repooling cloudvirt1013 after HW got fixed ([[phab:T241313|T241313]]) === 2020-01-21 === * 17:43 bstorm_: remounting /mnt/nfs/dumps-labstore1007.wikimedia.org/ on all dumps-mounting projects * 10:24 arturo: running `sudo systemctl restart apache2.service` in both labweb servers to try mitigating [[phab:T240852|T240852]] === 2020-01-15 === * 16:59 bd808: Changed the config for cloud-announce mailing list so that lsit admins do not get bounce unsubscribe notices === 2020-01-14 === * 14:03 arturo: icinga downtime all cloudvirts for another 2h for fixing some icinga checks * 12:04 arturo: icinga downtime toolchecker for 2 hours for openstack upgrades [[phab:T241347|T241347]] * 12:02 arturo: icinga downtime cloud* labs* hosts for 2 hours for openstack upgrades [[phab:T241347|T241347]] * 04:26 andrewbogott: upgrading designate on cloudservices1003/1004 === 2020-01-13 === * 13:34 arturo: [¢odfw1dev] prevent neutron from allocating floating IPs from the wrong subnet by doing `neutron subnet-update --allocation-pool start=208.80.153.190,end=208.80.153.190 cloud-instances-transport1-b-codfw` ([[phab:T242594|T242594]]) === 2020-01-10 === * 13:27 arturo: cloudvirt1009: virsh undefine i-000069b6. This is tools-elastic-01 which is running on cloudvirt1008 (so, leaked on cloudvirt1009) === 2020-01-09 === * 11:12 arturo: running `MariaDB [nova_eqiad1]> update quota_usages set in_use='0' where project_id='etytree';` ([[phab:T242332|T242332]]) * 11:11 arturo: running `MariaDB [nova_eqiad1]> select * from quota_usages where project_id = 'etytree';` ([[phab:T242332|T242332]]) * 10:32 arturo: ran `root@cloudcontrol1004:~# nova-manage project quota_usage_refresh --project etytree` === 2020-01-08 === * 10:53 arturo: icinga downtime all cloudvirts for 30 minutes to re-create all canary VMs" === 2020-01-07 === * 11:12 arturo: icinga-downtime everything cloud* for 30 minutes to merge nova scheduler changes * 10:02 arturo: icinga downtime cloudvirt1009 for 30 minutes to re-create canary VM ([[phab:T242078|T242078]]) === 2020-01-06 === * 13:45 andrewbogott: restarting nova-api and nova-conductor on cloudcontrol1003 and 1004 === 2020-01-04 === * 16:34 arturo: icinga downtime cloudvirt1024 for 2 months because hardware errors ([[phab:T241884|T241884]]) === 2019-12-31 === * 11:46 andrewbogott: I couldn't! * 11:40 andrewbogott: restarting cloudservices2002-dev to see if I can reproduce an issue I saw earlier === 2019-12-25 === * 10:13 arturo: icinga downtime for 30 minutes the whole cloud* lab* fleet to merge https://gerrit.wikimedia.org/r/c/operations/puppet/+/560575 (will restart some openstack components) === 2019-12-24 === * 15:13 arturo: icinga downtime all the lab* fleet for nova password change for 1h * 14:39 arturo: icinga downtime all the cloud* fleet for nova password change for 1h === 2019-12-23 === * 11:13 arturo: enable puppet in cloudcontrol1003/1004 * 10:40 arturo: disable puppet in cloudcontrol1003/1004 while doing changes related to python-ldap === 2019-12-22 === * 23:48 andrewbogott: restarting nova-conductor and nova-api on cloudcontrol1003 and 1004 * 09:45 arturo: cloudvirt1013 is back (did it alone) [[phab:T241313|T241313]] * 09:37 arturo: cloudvirt1013 is down for good. Apparently powered off. I can't even reach it via iLO === 2019-12-20 === * 12:43 arturo: icinga downtime cloudmetrics1001 for 128 hours === 2019-12-18 === * 12:55 arturo: [codfw1dev] created a new subnet neutron object to hold the new CIDR for floating IPs (cloud-codfw1dev-floating - 185.15.57.0/29) [[phab:T239347|T239347]] === 2019-12-17 === * 07:21 andrewbogott: deploying horizon/train to labweb1001/1002 === 2019-12-12 === * 06:11 arturo: schedule 4h downtime for labstores * 05:57 arturo: schedule 4h downtime for cloudvirts and other openstack components due to upgrade ops === 2019-12-02 === * 06:28 andrewbogott: running nova-manage db sync on eqiad1 * 06:27 andrewbogott: running nova-manage cell_v2 map_cell0 on eqiad1 === 2019-11-21 === * 16:07 jeh: created replica indexes and views for szywiki [[phab:T237373|T237373]] * 15:48 jeh: creating replica indexes and views for shywiktionary [[phab:T238115|T238115]] * 15:48 jeh: creating replica indexes and views for gcrwiki [[phab:T238114|T238114]] * 15:46 jeh: creating replica indexes and views for minwiktionary [[phab:T238522|T238522]] * 15:36 jeh: creating replica indexes and views for gewikimedia [[phab:T236404|T236404]] === 2019-11-18 === * 19:27 andrewbogott: repooling labsdb1011 * 18:54 andrewbogott: running maintain-views --all-databases --replace-all —clean on labsdb1011 [[phab:T238480|T238480]] * 18:44 andrewbogott: depooling labsdb1011 and killing remaining user queries [[phab:T238480|T238480]] * 18:42 andrewbogott: repooled labsdb1009 and 1010 [[phab:T238480|T238480]] * 18:19 andrewbogott: running maintain-views --all-databases --replace-all —clean on labsdb1010 [[phab:T238480|T238480]] * 18:18 andrewbogott: depooling labsdb1010, killing remaining user queries * 17:46 andrewbogott: running maintain-views --all-databases --replace-all —clean on labsdb1009 [[phab:T238480|T238480]] * 17:38 andrewbogott: depooling labsdb1009, killing remaining user queries * 16:54 andrewbogott: running maintain-views --all-databases --replace-all —clean on labsdb1012 [[phab:T237509|T237509]] === 2019-11-15 === * 20:04 andrewbogott: repool labdb1011 ([[phab:T237509|T237509]]) * 19:29 andrewbogott: running maintain-views --all-databases --replace-all —clean on labsdb1011 * 19:25 andrewbogott: depooling labsdb1011, killing remaining queries * 19:25 andrewbogott: repooling labsdb1010 * 18:59 andrewbogott: running maintain-views --all-databases --replace-all —clean on labsdb1012 * 18:57 andrewbogott: running maintain-views --all-databases --replace-all —clean on labsdb1010 * 18:54 andrewbogott: depooling labsdb1010, killing remaining user queries * 18:54 andrewbogott: depooled labsdb1009, ran maintain-views —clean —all-databases —replace-all, repooled === 2019-11-11 === * 13:10 arturo: cloudweb2001-dev: disable puppet and redirect stderr in the loadExitNodes.php cron script to prevent cronspam while we investigate the cause of the issue ([[phab:T237971|T237971]]) === 2019-11-05 === * 11:59 arturo: icinga downtime for 1h cloudcontrol1004, cloudnet1003, cloudvirt1017/1020/1022 for PDU operations in the rack [[phab:T227542|T227542]] === 2019-11-04 === * 21:55 andrewbogott: deleting a ton of wikitech hiera pages that were either no-ops or refer to nonexistent VMs or prefixes === 2019-10-31 === * 11:01 arturo: icinga-downtimed cloudvirt1030 and cloudservices1003 for 1h due to PDU upgrade operations [[phab:T227543|T227543]] === 2019-10-30 === * 22:43 jeh: reboot cloud-bootstrapvz-stretch to resolve bad bootstrapvz build === 2019-10-29 === * 10:52 arturo: icinga downtime cloudvirt1001/1002/1024/1018/1012/1009/1015/1008 for 1h [[phab:T227538|T227538]] === 2019-10-25 === * 10:45 arturo: icinga downtime toolschecker for 1 to upgrade clouddb1002 mariadb (toolsdb secondary) ([[phab:T236384|T236384]] , [[phab:T236420|T236420]]) === 2019-10-24 === * 12:30 arturo: starting cloudvirt1019, PDU operations ended ([[phab:T227540|T227540]]) * 11:58 arturo: icinga downtime for 2h ([[phab:T227540|T227540]]) cloudvirt1019 * 11:15 arturo: poweroff cloudvirt1019 during the PDU operations ([[phab:T227540|T227540]]) * 11:10 arturo: icinga downtime for 2h ([[phab:T227540|T227540]]) toolschecker * 10:58 arturo: icinga downtime for 1h ([[phab:T227540|T227540]]) cloudvirt100[3-7], cloudvirt1019, cloudvirt1016, cloudvirt1021, cloudvirt1013, cloudnet1004 === 2019-10-23 === * 09:23 arturo: cloudvirt1026 reboot ended OK * 09:12 arturo: rebooting cloudvirt1026 for kernel upgrade * 09:09 arturo: cloudvirt1025 reboot ended OK * 09:00 arturo: rebooting cloudvirt1025 for kernel upgrade * 08:51 arturo: icinga downtime cloudvirt1025/1026 for reboots === 2019-10-18 === * 16:01 arturo: created the `eqiad1.wikimedia.cloud` DNS zone ([[phab:T235846|T235846]]) * 14:27 andrewbogott: deleted a bunch of leaked VMS from earlier today from the admin-monitoring project. Fullstack leaks due to an api outage, maybe? * 10:44 arturo: double max_message_size from 40KB to 80KB in the cloud-admin mailing list. A simple email with a couple of quotes can go over the 40KB limit. === 2019-10-16 === * 21:59 jeh: resync wiki replica tool and user accounts [[phab:T235697|T235697]] * 09:40 arturo: reboot of cloudvirt1030 went fine * 09:28 arturo: reboot of cloudvirt1029 went fine * 09:28 arturo: rebooting cloudvirt1030 for kernel updates * 09:12 arturo: rebooting cloudvirt1029 for kernel updates * 09:11 arturo: reboot of cloudvirt1028 went fine * 09:00 arturo: rebooting cloudvirt1028 for kernel updates * 08:56 arturo: icinga downtime cloudvirt[1028-1030].eqiad.wmnet for 1h for reboots === 2019-10-15 === * 13:30 jeh: creating indexes and views for banwiki [[phab:T234770|T234770]] === 2019-10-10 === * 18:55 bd808: Created indexes and views for nqowiki ([[phab:T230543|T230543]]) * 11:59 arturo: network switch hardware is down affecting cloudvirt1025/1026 ([[phab:T227536|T227536]]) VMs are supposed to be online but unreachable === 2019-10-09 === * 10:44 arturo: cloudvirt1013 rebooted well * 10:32 arturo: cloudvirt1013 is rebooting * 10:32 arturo: cloudvirt1012 rebooted just fine (very slow, 35 VMs) * 10:21 arturo: cloudvirt1012 is rebooting * 10:19 arturo: cloudvirt1009 rebooted just fine (very slow though) * 10:07 arturo: cloudvirt1009 is rebooting * 10:06 arturo: cloudvirt1008 rebooted just fine (very slow though) * 09:58 arturo: cloudvirt1008 is rebooting * 09:52 arturo: icinga downtime toolschecker, paws, etc for 2h, because cloudvirt reboots === 2019-10-07 === * 14:07 arturo: horizon is disabled for maintenance ([[phab:T212302|T212302]]) * 14:00 arturo: starting scheduled maintenance: upgrading eqiad1 from openstack mitaka to newton === 2019-10-02 === * 15:23 arturo: codfw1dev renaming net/subnet objects to a more modern naming scheme [[phab:T233665|T233665]] * 12:49 arturo: codfw1dev delete all floating ip allocations in the deployment for mangling the network config for testing [[phab:T233665|T233665]] * 12:47 arturo: codfw1dev deleting all VMs in the deployment for mangling the network config for testing [[phab:T233665|T233665]] * 11:08 arturo: codfw1dev rebooting cloudnet2002-dev and cloudnet2003-dev for testing [[phab:T233665|T233665]] * 10:31 arturo: codfw1dev: add cloudinstances2b-gw router to the l3 agent in cloudnet2003-dev * 09:59 arturo: codfw1dev: cleanup leftover "HA port tenant admin" in neutron (ports from missing servers) * 09:46 arturo: codfw1dev: cleanup leftover neutron agents === 2019-09-30 === * 10:21 arturo: we installed ferm in every VM by mistake. Deleting it and forcing a puppet agent run to try to go back to a clean state. * 09:38 arturo: downtime toolschecker for 24h * 09:33 arturo: force update ferm cloud-wide (in all VMs) for [[phab:T153468|T153468]] === 2019-08-18 === * 10:39 arturo: rebooting cloudvirt1023 for new interface names configuration * 10:34 arturo: downtimed cloudvirt1023 for 2 days === 2019-08-05 === * 17:17 bd808: Set downtime on gridengine and kubernetes webservice checks in icinga until 2019-09-02 (flaky tests) === 2019-07-29 === * 20:14 bd808: Restarted maintain-kubeusers on tools-k8s-master-01 ([[phab:T194859|T194859]]) === 2019-07-25 === * 12:32 arturo: eqiad1/glance: debian-9.9-stretch image deprecates debian-9.8-stretch ([[phab:T228983|T228983]]) * 09:59 arturo: (codfw1dev) drop missing glance images ([[phab:T228972|T228972]]) * 09:32 arturo: (codfw1dev) deleting a bunch of VMs that were running in now missing hypervisors * 09:31 arturo: (codfw1dev) deleting a bunch of VMs in ERROR and SHUTDOWN state * 09:27 arturo: last log entry refers to the codfw1dev deployment * 09:27 arturo: cleanup `nova service-list` from old hypervisors (labtest*) * 09:23 arturo: refreshed nova DB grants in clouddb2001-dev for the codfw1dev deployment * 08:47 arturo: cleanup the cloud-announce pending emails (spam) === 2019-07-23 === * 19:43 andrewbogott: restarting rabbitmq-server on cloudcontrol1003 and 1004 === 2019-07-22 === * 23:44 bd808: Restarted maintain-kubeusers on tools-k8s-master-01 ([[phab:T228529|T228529]]) === 2019-07-11 === * 22:07 bd808: Ran `sudo systemctl stop designate_floating_ip_ptr_records_updater.service` on cloudcontrol1003 * 22:01 bd808: `sudo apt-get install python2.7-dbg` on cloudcontrol1003 to debug hung python process * 21:48 bd808: Ran `sudo systemctl stop designate_floating_ip_ptr_records_updater.service` on cloudcontrol1004 === 2019-06-25 === * 16:05 bstorm_: updated python3.4 to update4 wherever it was installed on Jessie VMs to prevent issues with broken update3. * 14:56 bstorm_: Updated python 3.4 on the labs-puppetmaster server === 2019-06-03 === * 15:55 arturo: [[phab:T221769|T221769]] rebooting cloudservices1003 after bootstrapping is apparently completed === 2019-05-28 === * 21:42 bstorm_: unmounting labstore1003-scratch on all cloud clients * 18:14 bstorm_: [[phab:T209527|T209527]] switched mounts from labstore1003 to cloudstore1008 for scratch === 2019-05-20 === * 17:25 arturo: [[phab:T223923|T223923]] dropped compat-network config from /etc/network/interfaces in eqiad1/codfw1dev neutron nodes * 17:22 arturo: [[phab:T223923|T223923]] dropped br-compat bridges and vlan interfaces (1102 and 2102) in eqiad1/codfw1dev neutron nodes * 17:07 arturo: [[phab:T223923|T223923]] dropped compat-network configuration from the neutron database in eqiad1 * 16:55 arturo: [[phab:T223923|T223923]] dropped compat-network configuration from the neutron database in codfw1dev === 2019-05-15 === * 17:00 andrewbogott: touching /root/firstboot_done on all VMs that cumin can reach. This will prevent firstboot.sh from running a second time if/when any of these are rebooted. [[phab:T223370|T223370]] === 2019-04-26 === * 15:51 arturo: andrew updated dns servers for the cloud-instances2-b-eqiad subnet in neutron: 208.80.154.143 and 208.80.154.24 === 2019-04-25 === * 11:14 arturo: [[phab:T221760|T221760]] increased size of conntrack table === 2019-04-24 === * 12:54 arturo: [[phab:T220051|T220051]] puppet broken in every VM in Cloud VPS, fixing right now === 2019-04-22 === * 11:14 arturo: create by hand /var/cache/labsaliaser/labs-ip-aliases.json in cloudservices2002-dev ([[phab:T218575|T218575]]) === 2019-04-16 === * 22:55 bd808: cloudcontrol2003-dev: added `exit 0` to /etc/cron.hourly/keystone to stop cron spam on partially configured cluster * 12:08 arturo: rebooting cloudvirt200[123]-dev because deep changes in config * 11:27 arturo: [[phab:T219626|T219626]] add DB grants for neutron and glnace to clouddb2001-dev (codfw1dev) * 10:37 arturo: [[phab:T219626|T219626]] replace 208.80.153.75 with 208.80.153.59 in the clouddb2001-dev database (codfw1dev deployment) * 10:30 arturo: [[phab:T219626|T219626]] replace labtestcontrol2003 with cloudcontrol2001-dev in the clouddb2001-dev database (codfw1dev deployment) === 2019-04-15 === * 13:08 arturo: [[phab:T219626|T219626]] add DB grants for keystone/nova/nova_api to clouddb2001-dev (codfw1dev) === 2019-04-13 === * 18:25 bd808: Restarted nova-compute service on cloudvirt1015 ([[phab:T220853|T220853]]) === 2019-04-11 === * 12:00 arturo: [[phab:T151704|T151704]] deploying oidentd to cloudnet1xxx servers === 2019-04-02 === * 19:52 andrewbogott: installed new base Stretch image. Updated packages, and runs apt-get dist-upgrade on first boot. === 2019-03-29 === * 14:34 andrewbogott: moving tools-static.wmflabs.org to point to tools-static-13 in eqiad1-r * 00:00 bstorm_: [[phab:T193264|T193264]] Added osm.db.svc.eqiad.wmflabs to cloud DNS === 2019-03-25 === * 00:40 bd808: Restarted maintain-dbusers on labstore1004. Process hung up on failed LDAP connection. === 2019-03-21 === * 19:32 andrewbogott: restarting keystone on cloudcontrol1003 === 2019-03-15 === * 16:00 gtirloni: increased nscd cache size ([[phab:T217280|T217280]]) === 2019-03-14 === * 19:04 gtirloni: bstorm started nfsd on labstore1006 ([[phab:T218341|T218341]]) * 16:42 gtirloni: published new debian-9.8 image ([[phab:T218314|T218314]]) === 2019-03-04 === * 19:37 bstorm_: umounted /mnt/nfs/dumps-labstore1006.wikimedia.org across all VPS projects for [[phab:T217473|T217473]] === 2019-02-26 === * 12:46 gtirloni: shutdown toolsbeta-sgegrid-master (cronspam) === 2019-02-25 === * 10:32 gtirloni: restarted nfsd on labstore1004 === 2019-02-21 === * 09:09 gtirloni: restarted uwsgi-labspuppetbackend.service on labpuppetmaster1001 * 07:42 gtirloni: created project cloudstore * 07:36 gtirloni: deleted wmcs-nfs project === 2019-02-20 === * 21:58 andrewbogott: silencing shinken and disabling puppet on shinken-02 for now === 2019-02-19 === * 12:00 gtirloni: added nagios@icinga2001.wikimedia.org to cloud-admin-feed@ allowed senders === 2019-02-18 === * 20:21 gtirloni: downtimed cloudvirt1020 * 20:12 gtirloni: ran `labs-ip-alias-dump.py` on cloudservices/labservices servers === 2019-02-15 === * 13:10 arturo: [[phab:T216239|T216239]] labvirt1019 has been drained * 12:22 arturo: [[phab:T216239|T216239]] draining labvirt1009 with a command like this: `root@cloudcontrol1004:~# wmcs-cold-migrate --region eqiad --nova-db nova 2c0cf363-c7c3-42ad-94bd-{{Gerrit|e586f2492321}} labvirt1001` * 12:02 arturo: more nova service cleanups in the database (labvirts that were reallocated to eqiad1) * 11:34 arturo: [[phab:T216190|T216190]] cleanup from nova database `nova service-delete 35` * 03:50 andrewbogott: updated VPS base images for Jessie and Stretch, now featuring Stretch 9.7 === 2019-02-11 === * 18:13 gtirloni: cleaned old metrics data in labmon1001 [[phab:T215417|T215417]] * 15:28 gtirloni: running `maintain-views --all-databases --replace-all` on labsdb1011 * 14:18 gtirloni: running `maintain-views --all-databases --replace-all` on labsdb1010 === 2019-02-08 === * 14:56 gtirloni: running `maintain-views --all-databases --replace-all` on labsdb1009 === 2019-02-06 === * 11:47 gtirloni: downtimed labmon100{1,2} [[phab:T215399|T215399]] * 00:17 bstorm_: [[phab:T214106|T214106]] deleted bstorm-test2 project to clean up === 2019-02-05 === * 10:48 arturo: labmon1001 is now part of the 'eqiad1-r' region === 2019-02-01 === * 09:54 arturo: moving canary1015-01 VM instance from cloudvirt1024 back to cloudvirt1015 === 2019-01-31 === * 12:44 arturo: [[phab:T215012|T215012]] depooling cloudvirt1015 and migrating all VMs to cloudvirt1024 === 2019-01-25 === * 20:11 gtirloni: deleted project yandex-proxy [[phab:T212306|T212306]] * 20:11 gtirloni: deleted project [[phab:T212306|T212306]] === 2019-01-24 === * 11:50 arturo: [[phab:T213925|T213925]] modify subnet cloud-instances-transport1-b-eqiad1 to avoid floating IP allocations from here * 11:07 arturo: [[phab:T214299|T214299]] failover cloudnet1003 to cloudnet1004 * 10:03 arturo: [[phab:T214299|T214299]] reimage cloudnet1004 to debian stretch * 09:51 arturo: [[phab:T214299|T214299]] failover cloudnet1004 to cloudnet1003 === 2019-01-22 === * 19:19 arturo: [[phab:T214299|T214299]] stretch cloudnet1003 is apparently all set * 18:40 arturo: [[phab:T214299|T214299]] manually delete from neutron agents from cloudnet1003 (must be added again after reimage, with new uuids) * 18:37 arturo: [[phab:T214299|T214299]] reimaging cloudnet1003 as debian stretch * 17:35 jbond42: starting roll out of apt package updates to * 14:41 gtirloni: [[phab:T214369|T214369]] deployed new jessie and stretch VM images === 2019-01-21 === * 18:29 gtirloni: installed libguestfs-tools on cloudvirt1021 === 2019-01-16 === * 14:21 andrewbogott: stopping old VPS proxies in eqiad — [[phab:T213540|T213540]] === 2019-01-15 === * 14:20 andrewbogott: changing tools.wmflabs.org to point to tools-proxy-03 in eqiad1 === 2019-01-13 === * 20:00 andrewbogott: VPS proxies are now running in eqiad1 on proxy-01. Old VMs will wait a bit for deletion. [[phab:T213540|T213540]] * 19:12 andrewbogott: moving the VPS proxy API backend to proxy-01.project-proxy.eqiad.wmflabs, as per [[phab:T213540|T213540]] * 17:11 andrewbogott: moving all VPS dynamic proxies to proxy-eqiad1.wmflabs.org aka proxy-01.project-proxy.eqiad.wmflabs, as per [[phab:T213540|T213540]] === 2019-01-09 === * 22:21 bd808: neutron quota-update --tenant-id tools --port 256 === 2019-01-08 === * 18:59 bd808: Definately did NOT delete uid=novaadmin,ou=people,dc=wikimedia,dc=org * 18:59 bd808: Deleted LDAP user uid=neutron,ou=people,dc=wikimedia,dc=org * 18:58 bd808: Deleted LDAP user uid=novaadmin,ou=people,dc=wikimedia,dc=org === 2019-01-06 === * 22:03 bd808: Set floatingip quota of 60 for tools project in eqiad1-r region ([[phab:T212360|T212360]]) === 2018-12-20 === * 17:10 arturo: [[phab:T207663|T207663]] renumbered transport network in eqiad1 === 2018-12-05 === * 17:59 arturo: [[phab:T207663|T207663]] changed labtestn transport network addressing from private to public === 2018-12-03 === * 13:25 arturo: [[phab:T202886|T202886]] create again PTR records after dnsleak.py fix === 2018-11-30 === * 14:08 arturo: running dns leaks cleanup `root@cloudcontrol1003:~# /root/novastats/dnsleaks.py --delete` === 2018-11-28 === * 17:33 gtirloni: deleted contintcloud project ([[phab:T209644|T209644]]) === 2018-11-27 === * 13:32 gtirloni: enabled DRBD stats collection on labstore100[4-5] [[phab:T208446|T208446]] === 2018-11-22 === * 07:12 gtirloni: deployed new debian-9.6-stretch image === 2018-11-21 === * 10:48 arturo: re-created compat-net as not shared in labtestn to test stuff related to [[phab:T209954|T209954]] === 2018-11-16 === * 12:43 gtirloni: armed keyholder on labpuppetmaster1001/1002 after reboots * 12:08 gtirloni: rebooted labpuppetmaster1001 ([[phab:T207377|T207377]]) * 11:57 gtirloni: rebooted labpuppetmaster1002 ([[phab:T207377|T207377]]) === 2018-11-14 === * 17:19 gtirloni: added cloudvirt1016 to scheduler pool ([[phab:T209426|T209426]]) * 15:41 gtirloni: reimaging labvirt1016 as cloudvirt1016 * 15:14 gtirloni: reset-failed systemd unit nova-scheduler on cloudcontrol1004 * 13:52 gtirloni: rebooted labservices1002 after package upgrades ([[phab:T207377|T207377]]) * 13:23 gtirloni: rebooted labstore2004 after package upgrades ([[phab:T207377|T207377]]) * 13:20 gtirloni: rebooted labstore2003 after package upgrades ([[phab:T207377|T207377]]) * 13:20 gtirloni: rebooted labstore2001/labstore2003 after package upgrades ([[phab:T207377|T207377]]) * 12:08 gtirloni: rebooted labnet1002 after package upgrades * 12:01 gtirloni: rebooted labmon1002 after package upgrades * 11:41 gtirloni: rebooted labcontrol1002 after package upgrades * 11:15 gtirloni: rebooted cloudcontrol1004 after package upgrades === 2018-11-09 === * 18:17 gtirloni: restarted neutron-linuxbridge-agent on cloudvirt1018/1023 === 2018-11-08 === * 11:00 gtirloni: Added novaproxy-02 to $CACHES * 10:50 gtirloni: Added cloudvirt1017 to eqiad1 region === 2018-11-07 === * 13:49 arturo: [[phab:T208733|T208733]] moving labvirt1017 from main deployment to eqiad1 and renaming it to cloudvirt1017 === 2018-10-22 === * 16:24 arturo: [[phab:T206261|T206261]] another update to dmz_cidr in eqiad1 * 10:26 arturo: change again in dmz_cidr in eqiad1: VMs will connect between them without NAT even when using floating IPs ([[phab:T206261|T206261]]) === 2018-10-19 === * 12:02 arturo: revert change in dmz_cidr in eqiad1 for now ([[phab:T206261|T206261]]) * 11:16 arturo: change in dmz_cidr in eqiad1: VMs will connect between them without NAT even when using floating IPs ([[phab:T206261|T206261]]) * 10:14 arturo: we have new virt servers in the eqiad1 deployment since past week and this week: cloudvirt1018, cloudvirt1023, cloudvirt1024 === 2018-09-26 === * 10:40 arturo: [[phab:T205524|T205524]] all sorts of restarts in all neutron daemons * 10:20 arturo: [[phab:T205524|T205524]] stop/start all neutron agents in cloudnet1003.eqiad.wmnet * 10:13 arturo: [[phab:T205524|T205524]] restart all agents in cloudnet1004.eqiad.wmnet * 10:10 arturo: restart neutron-server in cloudcontrol1003, investigating [[phab:T205524|T205524]] === 2018-09-24 === * 10:57 arturo: try to increase floating ip allocation pool in eqiad1. Of 185.15.56.0/25 we are using only 185.15.56.10-185.15.56.31, I don't know why. Let's use 185.15.56.2-185.15.56.126 === 2018-09-21 === * 17:18 bd808: Running `sudo maintain-meta_p --all-databases --purge` across labsdb10(09{{!}}10{{!}}11) for [[phab:T201890|T201890]] === 2018-09-17 === * 22:08 bd808: Granted gtirloni project roles of admin, projectadmin, and user === 2018-09-12 === * 11:20 arturo: [[phab:T202636|T202636]] distributing default routes using classless-static-route for all VMs in main/labtest (dnsmasq/nova-network) === 2018-09-11 === * 16:52 arturo: again, restarted nova-network after killing all dnsmasq procs in labnet1001 for [[phab:T202636|T202636]] * 16:08 arturo: restarted nova-network after killing all dnsmasq procs in labnet1001 for [[phab:T202636|T202636]] * 10:53 arturo: [[phab:T202636|T202636]] creating all the compat-network configuration in neutron * 10:36 arturo: [[phab:T202636|T202636]] creating br-compat bridge in eqiad1 for the compat network * 10:33 arturo: [[phab:T202636|T202636]] manually reserve 10.68.23.253 (in nova-network) === 2018-09-10 === * 22:46 andrewbogott: deleting all VMs on labvirt1019 and 1020 as prep for [[phab:T204003|T204003]] === 2018-08-30 === * 15:46 andrewbogott: restarting rabbitmq-server on cloudcontrol1003 * 13:07 arturo: [[phab:T202636|T202636]] internal network routing now exists in labtest/labtestn for VM to communicate with each other === 2018-08-28 === * 11:04 arturo: [[phab:T202549|T202549]] eqiad1 databases are all now running in m5-master. Mysql has been cleaned from cloudcontrol100[3,4] === 2018-08-23 === * 16:17 arturo: [[phab:T188589|T188589]] bstorm_ merged patch to reduce nova DB connection usage * 13:15 arturo: [[phab:T202115|T202115]] `root@cloudcontrol1003:~# neutron subnet-update --allocation-pool start=10.64.22.4,end=10.64.22.4 e4fb2771-a361-4add-ac4e-280cc300c59f` * 13:10 arturo: [[phab:T202115|T202115]] (was `{"start": "10.64.22.2", "end": "10.64.22.254"}` ) * 13:08 arturo: [[phab:T202115|T202115]] `root@cloudcontrol1003:~# neutron subnet-update --allocation-pool start=10.64.22.254,end=10.64.22.254 e4fb2771-a361-4add-ac4e-280cc300c59f` === 2018-08-22 === * 15:28 arturo: cleanup local glance,keystone databases in cloudcontrol1003.wikimedia.org (already in m5-master) * 15:27 arturo: cleanup local keystone database in cloudcontrol1003.wikimedia.org (already in m5-master) === 2018-08-21 === * 15:39 andrewbogott: initial test message * 10:31 arturo: eqiad1 remove leftover port for HA on labnet1004 * 10:15 arturo: test === 2018-05-07 === * 18:07 bstorm_: stopped the toolhistory job because it is totally broken and fills /tmp. === 2018-02-09 === * 00:55 bd808: Added Arturo Borrero Gonzalez and Bstorm as project members * 00:54 bd808: Removed Yuvipanda at user request ([[phab:T186289|T186289]]) {{SAL|Project Name=admin}} <noinclude>[[Category:SAL]]</noinclude> qtb6fky8eg0zakozssfhqk1lw06m359 Nova Resource:Tools.integraality/SAL 498 443949 2450676 2450127 2026-08-23T10:47:32Z Stashbot 7414 wmbot~jeanfred@tools-bastion-15: Install new requirements (toolforge, pymysql) in Python 3.11 virtual environment using 'toolforge webservice python3.11 shell' / '/data/project/integraality/www/python/venv/bin/pip install -r /data/project/integraality/integraality/requirements.txt' 2450676 wikitext text/x-wiki === 2026-08-23 === * 10:47 wmbot~jeanfred@tools-bastion-15: Install new requirements (toolforge, pymysql) in Python 3.11 virtual environment using 'toolforge webservice python3.11 shell' / '/data/project/integraality/www/python/venv/bin/pip install -r /data/project/integraality/integraality/requirements.txt' === 2026-08-20 === * 13:42 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|881928b}} (Increase webservice memory limit to 1Gi) * 13:42 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|ea800c5}} (Add functional SPARQL tests against live endpoints) === 2026-08-19 === * 22:12 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|ef9b3fc}} (Wrap drilldown queries in subquery for label resolution) * 22:12 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|59ccbe6}} (Refactor: extract _build_drilldown_query helper) * 22:11 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|1cd9f21}} (Switch test fixtures from red pandas to lighthouses) * 22:11 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|a7ad288}} (Update all SPARQL unit tests targeting Q41960 to use P10241 instead of P31) === 2026-08-18 === * 09:13 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|b75abe2}} (Add pre-commit hook to reformat pyproject.toml using pyproject-fmt) * 09:13 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|fd89068}} (Reformat pyproject.toml using pyproject-fmt) * 09:13 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|5cc67af}} (Add pre-commit hook to validate docker-compose.yml against Compose schema) * 09:12 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|5da6acb}} (Add pre-commit hook to validate toolinfo.json against Toolhub schema) * 09:12 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|66c6495}} (Fix tool type in toolinfo.json, as lists are not accepted) === 2026-08-17 === * 20:22 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|1b3b5ec}} (Update toolinfo.json record, based on latest 1.2.2 schema) * 20:22 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|bbf2c99}} (Set name in toolinfo record to toolforge-$TOOL_NAME to avoid duplicates) * 20:22 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|853a540}} (Add --page argument to CLI for single-page updates) * 19:34 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|589a9a3}} (ReferenceColumn: show reference value in positive drill-down) for [[phab:T428636|T428636]] * 19:34 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|3976414}} (ReferenceColumn: show statement value in negative drill-down) for [[phab:T428637|T428637]] * 19:34 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|f9043eb}} (DescriptionColumn: show the description value in positive drill-down) * 19:34 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|3f69cf4}} (Extract _format_value_sparql helper for value-to-SPARQL-term conversion) * 19:34 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|12202cd}} (Refactor: columns own their SELECT variables for drill-down queries) === 2026-08-16 === * 22:42 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|269bec7}} (Add shfmt reformat commit to .git-blame-ignore-revs) * 22:42 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|6bd7972}} (Reformat sal_messages_from_git_log.sh with shfmt) * 22:42 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|2855e52}} (Add pre-commit hook for Shell formatter shfmt) * 19:56 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|bd653aa}} (Remove Ansible playbook execution verbosity) * 18:09 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|b4c3782}} (Migrate deprecated template params during dashboard updates) * 18:09 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|2018d1e}} (Add row_totals template parameter to toggle totals row) * 18:09 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|622e6ac}} (Rename stats_for_no_group → row_no_group) * 18:09 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|5bc0f3f}} (Extract _parse_bool_param helper for template boolean parsing) * 18:09 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|6c43294}} (Refactor ColumnMaker.make() into focused helper methods) * 18:09 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|ba0e0f9}} (Add value constraints in multi-property reference lists (S248=Q...+S813)) for [[phab:T428637|T428637]] * 18:09 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|8c5e68c}} (Refactor MultiPropertyReferenceCheck to use (property, value) tuples) for [[phab:T428637|T428637]] * 18:09 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|d1bba64}} (Add P123/S248=Q19216625 syntax for value-constrained reference checks) for [[phab:T428637|T428637]] * 18:09 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|bc7a5b2}} (Add AllPropertiesReferenceCheck with + syntax (AND)) for [[phab:T428637|T428637]] * 18:09 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|5c0ad71}} (Introduce MultiPropertyReferenceCheck base class) for [[phab:T428637|T428637]] * 18:09 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|ad71cc9}} (Simplify PropertyReferenceCheck to take a single property string) for [[phab:T428637|T428637]] * 18:09 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|a350d45}} (Extract AnyOfPropertiesReferenceCheck from PropertyReferenceCheck) for [[phab:T428637|T428637]] === 2026-08-05 === * 15:50 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|a8b6d05}} (Add qualifier-scoped reference columns (P123/P789/S*, P123/Q456/P789/S*)) for [[phab:T428637|T428637]] * 15:50 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|ae0e765}} (Add P123/S248;S854 syntax for multiple reference properties (OR)) for [[phab:T428637|T428637]] * 15:50 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|45db60c}} (Refactor PropertyReferenceCheck to take a list of properties) for [[phab:T428637|T428637]] * 15:50 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|aa6d5aa}} (Add full assembled query tests for PropertyReferenceCheck (S248)) for [[phab:T428637|T428637]] * 15:49 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|c1f554f}} (Improve SPARQL formatting: multi-line braces for FILTER NOT EXISTS) for [[phab:T428637|T428637]] * 15:49 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|5bd85c3}} (Add P123/S! syntax for good references (excluding subpar source properties)) for [[phab:T428637|T428637]] === 2026-07-25 === * 09:37 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|1b23f69}} (Hide empty Top Properties header for column-less dashboards) === 2026-07-24 === * 13:26 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|cefdfd6}} (Bind ?grouping in drill-down queries for value-scoped columns) for [[phab:T428637|T428637]] * 13:26 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|51b9ad1}} (Add value-scoped reference columns (P123/Q456/S*, P123/?grouping/S*)) for [[phab:T428637|T428637]] === 2026-07-22 === * 13:13 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|eaa22b6}} (Add P123/S456 syntax for specific reference property) for [[phab:T428637|T428637]] * 13:13 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|47fd820}} (Refactor: extract ReferenceCheck strategy from ReferenceColumn) for [[phab:T428637|T428637]] * 13:13 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|116fe90}} (Override format_html_snippet in ReferenceColumn) for [[phab:T428637|T428637]] === 2026-07-21 === * 10:38 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|6b79fa1}} (Add noscript fallback for streaming update page) for [[phab:T425477|T425477]] === 2026-07-20 === * 21:36 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|14ca09e}} (Collapse completed sections in SSE progress view) for [[phab:T425477|T425477]] * 17:00 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|dd7677a}} (Fix footer layout: use sticky footer instead of fixed position) * 17:00 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|b3e24a6}} (Log all totals and no-grouping SPARQL queries in SSE stream) for [[phab:T425477|T425477]] * 16:03 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|7dbac6d}} (Make streaming update the default) for [[phab:T425477|T425477]] * 16:03 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|a4029f2}} (Mark completed steps with ✅ via DOM mutation) for [[phab:T425477|T425477]] * 16:03 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|4235204}} (Show SPARQL queries in SSE progress messages) for [[phab:T425477|T425477]] * 16:03 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|2fa79d1}} (Fix QLever link in SSE error UI: add SPARQL prefixes) for [[phab:T425477|T425477]] === 2026-07-17 === * 21:06 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|d794ddd}} (Remove hardcoded debug mode for Flask application in app.py) * 21:06 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|ea1afda}} (Add integration tests for SSE streaming error scenarios) for [[phab:T425477|T425477]] * 21:06 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|b2fd525}} (Add retry button for transient errors and connection loss in SSE template) for [[phab:T425477|T425477]] * 21:06 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|8926c8f}} (Render differentiated error UIs in SSE template) for [[phab:T425477|T425477]] * 21:06 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|87dcff4}} (Escape dynamic content in innerHTML in SSE template) for [[phab:T425477|T425477]] * 21:06 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|9deb932}} (Extract handlers into named functions in SSE template) for [[phab:T425477|T425477]] * 21:06 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|5d8de99}} (Enrich SSE error events with error_type, error_category, and query) for [[phab:T425477|T425477]] * 21:06 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|fc20c17}} (Add ErrorCategory enum and declare it on exception classes) for [[phab:T425477|T425477]] * 21:06 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|ec2a612}} (Use {{!}}tojson for JS string interpolation in SSE template to avoid XSS risk) * 21:06 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|ad3fd00}} (Fix Flask hot-reload in docker-compose) === 2026-07-16 === * 08:18 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|651408e}} (Add ReferenceColumn for tracking sourced statements) for [[phab:T428637|T428637]] * 08:18 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|94f04f3}} (Refactor: extract get_column_label() from make_column_header()) === 2026-06-11 === * 10:40 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|3b935df}} (Extract QualifierColumn from PropertyColumn) * 10:40 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|338027c}} (Rename ?prop to ?statement in PropertyColumn SPARQL) === 2026-06-04 === * 20:05 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|c5294b1}} (Support ?grouping as qualifier column value) * 19:17 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|f07edf9}} (Add Queries endpoint support for P1/Q2/P3 qualifier columns) for [[phab:T251008|T251008]] * 19:17 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|ac8d9f4}} (Replace invalid nested <ul> tag in footer in base HTLM template) * 19:17 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|73ac3a4}} (Remove stray '>' character inside the <script> tag in base HTML template) === 2026-05-26 === * 17:31 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|18b6a34}} (Fix test_index_page unit test) * 17:31 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|55c2fbe}} (Reorganize slightly README and CONTRIBUTING.md files) * 17:31 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|d86cbb7}} (Add/Expand module-level docstrings) === 2026-05-13 === * 13:19 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|5e938ef}} (Update link to Toolforge in base HTML template) * 13:19 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|e2636db}} (Update documentation link in README following reorganization) * 13:19 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|0b33d6e}} (Refresh landing page) === 2026-05-10 === * 12:11 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|0b7171b}} (Add Ansible handler to log deploys to Server Admin Log via dologmsg) * 12:11 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|f348720}} (Move ConfigAssembler tests to their own test_config_assembler.py) * 12:11 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|4262f3e}} (Extract config assembly from PagesProcessor into ConfigAssembler) * 12:11 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|2bd496d}} (Fix ruff config: use extend-select instead of select) * 12:11 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|f57a49d}} (Remove unused local variable in RedisCache.list_keys) === 2026-05-08 === * 20:56 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|a10fba1}} (Add CONTRIBUTING.md) * 20:56 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|71d7701}} (Add mdformat pre-commit hook for Markdown formatting) * 20:56 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|1c1cec1}} (Rewrite README) === 2026-05-07 === * 12:54 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|6f644fd}} (Add --warm-cache-only option to populate Redis without running queries) * 11:36 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|35b1283}} (Derive cache key from URL to avoid triggering pywikibot Site creation) * 11:36 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|c7f3a04}} (Make pywikibot.Site creation lazy in PagesProcessor) * 11:36 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|a8dd1ed}} (Remove unused self.repo from PagesProcessor) * 11:36 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|d2cdb4d}} (Pass explicit endpoint to SparqlQuery to avoid Site creation) === 2026-05-06 === * 19:46 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|fa3a82e}} (Defer WdqsSparqlQueryEngine instantiation to avoid import-time side effects) for [[phab:T421552|T421552]] * 19:06 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|50d6c27}} (Add SSE helper: QueueHandler-based event generator), {{Gerrit|022bb07}} (Add streaming update endpoint using Server-Sent Events), {{Gerrit|1fe594f}} (Add phase field to SSE events for visual distinction), {{Gerrit|66eca3c}} (Filter SSE events by thread to isolate concurrent updates) for [[phab:T425477|T425477]] === 2026-05-05 === * 20:02 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|d72d22e}} (Update pywikibot from 10.7.6 to latest 11.2.0) * 19:43 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|0693c65}} (dd additional progress logging to update pipeline), {{Gerrit|f3c8c35}} (Replace pywikibot.output/warn/print with logger calls) and {{Gerrit|670e55b}} (Move RedisCache logging to call site, remove prints from cache.py) ahead of [[phab:T425477|T425477]] * 19:42 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|b500596}} (Validate grouping_link_mode config value) for [[phab:T368001|T368001]] === 2026-05-03 === * 19:44 wmbot~jeanfred@tools-bastion-15: Deplot {{Gerrit|0fd3dd4}} (Prepend default label column in Listeria output) for [[phab:T368001|T368001]] * 19:44 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|18813aa}} (Put label and description columns first in Listeria output) for [[phab:T368001|T368001]] * 19:44 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|b41db93}} (Re-resolve grouping links after year rebinning) for [[phab:T368001|T368001]] * 19:44 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|9fbf0bd}} (Add grouping_link_mode config and wire up page creation) for [[phab:T368001|T368001]] * 19:44 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|729decc}} (Add GroupingPageCreator to create Listeria subpages) for [[phab:T368001|T368001]] * 19:44 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|74bf18f}} (Add Listeria wikitext formatting to Grouping rows) for [[phab:T368001|T368001]] * 19:44 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|9b32a82}} (Split retrieve_and_process_data in process_page) for [[phab:T368001|T368001]] * 19:44 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|4dcc1a0}} (Extract prepare_report_groupings from process_data in PropertyStatistics) for [[phab:T368001|T368001]] * 19:41 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|928f398}} (Remove obsolete version attribute from docker-compose.yml) === 2026-05-01 === * 18:16 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|01e97d6}} (Extract grouping link strategies into GroupingLink hierarchy), {{Gerrit|359770d}} (Remove pass-through methods for grouping link in GroupingConfiguration), {{Gerrit|54e4a7b}} (Add lang parameter to get_label_for_variable), {{Gerrit|cccd661}} (Simplify SPARQL for mul language in get_label_for_variable), {{Gerrit|548383d}} (Support any label language for grouping links), {{Gerrit|d571596}} (Support QID-based grouping links), {{Gerrit|cfd7c68}} ( * 18:15 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|19b856d}}, {{Gerrit|473fbeb}}, {{Gerrit|0cb0915}} === 2026-03-27 === * 22:58 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|2bad8b3}} (Catch ModuleNotFoundError when unpickling stale cache entries) === 2026-03-24 === * 16:52 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|7883b73}} (Switch from bare imports to relative imports) * 15:56 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|5dbe412}} (Resolve grouping type eagerly in PropertyStatistics constructor) === 2026-03-23 === * 22:37 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|1b60d39}} (Handle stale pickle cache entries referencing removed classes) * 22:00 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|124f4ca}} (Replace wikibase:isSomeValue with STRSTARTS filter for QLever compatibility) * 21:28 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|fd5828f}} (Detect grouping type via SPARQL and flatten to single GroupingConfiguration) * 21:27 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|e9fae82}} (Add AbstractGroupingType base class for grouping type strategies) * 21:27 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|fee0ed5}} (Extract GroupingType strategy classes from GroupingConfiguration hierarchy) === 2026-03-18 === * 20:57 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|f9dfb79}} (Add /healthz endpoint for automatic health check) === 2026-03-17 === * 16:51 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|2940962}} (Replace GROUP_MAPPING Enum with SPECIAL_GROUPINGS tuple) * 16:51 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|420658b}} (Add MARKER attribute to special grouping classes) * 16:51 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|976599f}} (Rename UnknownValueGrouping.TITLE to MARKER) === 2026-03-14 === * 16:39 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|be09048}} (Add Wikimedia Commons support using QLever endpoint) for [[phab:T294893|T294893]] * 16:38 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|ae53e61}} (Derive QLever UI URL from engine endpoint) * 15:45 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|5134f77}} (Use ka/Ma unit suffixes for geological-scale year grouping headings) for [[phab:T236590|T236590]] * 15:45 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|8a735a0}} (Add automatic date resolution rebinning for year groupings) for [[phab:T236590|T236590]] * 15:44 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|6478051}}, {{Gerrit|4206c09}}, {{Gerrit|0e920fe}}, {{Gerrit|4eefea1}} === 2026-03-13 === * 20:33 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|a39f1d9}} (Add proper handling for transient Wikidata server errors) for [[phab:T415439|T415439]] * 20:19 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|0c418a5}} (Fix tests hitting live Wikidata API (and potential rate-limits)) for [[phab:T420054|T420054]] === 2026-03-11 === * 21:58 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|9665d8f}} (Remove obsolete GroupingType enum from column.py) * 21:57 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|60d500a}} (Remove backward-compatible formatting wrapper methods from PropertyStatistics) * 21:20 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|a452c28}} (Extract formatting logic into separate ResultsFormatter class) * 17:29 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|02338eb}} (Move heading strings to class constants in Grouping objects) and {{Gerrit|d730307}} (Format NoGroup and Totals in with the other groupings) === 2026-03-09 === * 18:32 wmbot~jeanfred@tools-bastion-15: Recreated Python 3.11 virtual environment using 'toolforge webservice python3.11 shell' / 'webservice-python-bootstrap --fresh' / '/data/project/integraality/www/python/venv/bin/pip install -r /data/project/integraality/integraality/requirements.txt' * 18:30 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|55f1132}} (Switch to Python 3.11 as default interpreter) === 2026-03-08 === * 11:52 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|37fcc25}} (Extract jobs-framework runtime image as its own variable) and {{Gerrit|1f9bfbb}} (Add service.template deployment to Ansible) === 2026-03-06 === * 10:08 wmbot~jeanfred@tools-bastion-15: Successfully installed flask-3.1.3 mwparserfromhell-0.7.2 pywikibot-10.7.6 redis-7.0.1 (using a webservice python3.9 shell) * 10:07 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|9fa536f}}, {{Gerrit|9e5111c}}, {{Gerrit|e176f18}}, {{Gerrit|eff3045}}, {{Gerrit|8532e32}} (Dependencies upgrades) === 2026-03-05 === * 19:44 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|5075934}} (Add support for user-specified groupings via template parameter) for [[phab:T419059|T419059]] * 19:44 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|3c9fa98}} (Remove two debug print statements) * 18:44 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|28d800e}} (Extract rdfs:label SPARQL fragment from grouping.py to get_label_for_variable utility), {{Gerrit|40c39de}} (Make intermediate label variables unique to avoid clashes in get_label_for_variable) and {{Gerrit|64e8f94}} (Replace SERVICE wikibase:label with native SPARQL for QLever compatibility) for [[phab:T385749|T385749]] === 2025-11-03 === * 17:19 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|2d3e128}} (Display the name of the SPARQL endpoint used in the edit summary) for [[phab:T385749|T385749]] * 17:19 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|2d3e128}} (Display the name of the SPARQL endpoint used in the edit summary) for [[phab:T385749|T385749]] === 2025-10-29 === * 16:49 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|f40974f}} (Add links to QLever in the Queries) for [[phab:T385749|T385749]] * 16:49 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|28c2e1b}} (Add link to QLever on queries error page) for [[phab:T385749|T385749]] * 16:49 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|4cc2bcc}} (Add configurable SPARQL endpoint support for QLever in PagesProcessor) for [[phab:T385749|T385749]] * 16:49 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|be5c234}} (Add support for QLever as alternative SPARQL query engine) for [[phab:T385749|T385749]] * 12:32 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|f34624d}} (Move pywikibot exception handling to SPARQL engine class) for [[phab:T385749|T385749]] * 12:32 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|4da9f33}} (Implement SPARQL query engine dependency injection) for [[phab:T385749|T385749]] === 2025-10-28 === * 21:58 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|92ab962}} (Update pywikibot from 10.0.0 to latest 10.6.0) * 21:34 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|77deb17}} (Remove link to 'SGE-jobs' tool in footer) === 2025-10-27 === * 21:02 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|8db01c6}} (Simplify Flask webserver startup in app.py) * 20:32 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|31f14b0}} (Bump ruff pre-commit hook from 0.9.3 to latest 0.14.2) * 20:32 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|995121d}} (Bump uv pre-commit hook from 0.5.24 to latest 0.9.5 and regenerate requirements files) * 09:07 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|69669c0}} (Optimize 'get_grouping_information_query' query using subquery) for [[phab:T400480|T400480]] === 2025-10-25 === * 09:00 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|9ac4c1f}} (Convert SPARQL query tests to use triple-quoted strings in PropertyStatisticsTest) === 2025-10-24 === * 21:02 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|5511fd0}} (Fix test_get_grouping_information_query_with_grouping_link for YearGrouping) * 19:39 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|8bb5d37}} (Convert SPARQL query tests to use triple-quoted strings) * 19:37 wmbot~jeanfred@tools-bastion-15: Deploy {{Gerrit|792b38e}} (Remove redundant GROUP BY parameter in 'get_grouping_information_query') === 2025-05-06 === * 20:05 wmbot~jeanfred@tools-sgebastion-10: Deploy {{Gerrit|284a70a}} (Update pywikibot from 9.6.1 to latest 10.0.0) * 20:05 wmbot~jeanfred@tools-sgebastion-10: Deploy {{Gerrit|1f030a9}} (Re-generate uv lockfile to add 'upload-time' markers) * 11:19 wmbot~jeanfred@tools-sgebastion-10: Deploy {{Gerrit|8f2391e}} (Replace deprecated with in two unit tests) === 2025-05-04 === * 12:57 wmbot~jeanfred@tools-sgebastion-10: Deploy {{Gerrit|7a2fa93}} (Fallback to mul when formatting grouping links) for [[phab:T393166|T393166]] * 12:57 wmbot~jeanfred@tools-sgebastion-10: Deploy {{Gerrit|381d8f9}} (Add pre-commit hook to run Python linter ruff) === 2025-05-03 === * 19:57 wmbot~jeanfred@tools-sgebastion-10: Deploy {{Gerrit|781934d}} (Move grouping-link to Grouping objects and retrieve via SPARQL) for [[phab:T237276|T237276]] === 2025-05-02 === * 19:41 wmbot~jeanfred@tools-sgebastion-10: Deploy {{Gerrit|381d8f9}} (Add pre-commit hook to run Python linter ruff) * 19:41 wmbot~jeanfred@tools-sgebastion-10: Deploy {{Gerrit|5bc6fe9}} (Remove unused local variable in GroupingTest.test unit-test) * 19:41 wmbot~jeanfred@tools-sgebastion-10: Deploy {{Gerrit|f5f0044}} (Remove unused local variable in unit-test) * 19:41 wmbot~jeanfred@tools-sgebastion-10: Deploy {{Gerrit|3a742a2}} (Remove unused column_key local variable from get_query_for_items_for_property_[posi{{!}}nega]tive) * 19:41 wmbot~jeanfred@tools-sgebastion-10: Deploy {{Gerrit|620ac00}} (Remove unused patcher in QueriesTests setUp method) * 19:41 wmbot~jeanfred@tools-sgebastion-10: Deploy {{Gerrit|e32b9b9}} (Allow to optionally pass wikiproject data to SitelinkColumn constructor) === 2025-05-01 === * 21:02 wmbot~jeanfred@tools-sgebastion-10: Deploy {{Gerrit|2a91226}} (Catch WDQS Timeouts and raise proper exception for all SPARQL queries) for [[phab:T278156|T278156]] * 20:29 wmbot~jeanfred@tools-sgebastion-10: Deploy {{Gerrit|397c5e8}} (Also catch ServerError to handle WDQS timeouts) for [[phab:T384882|T384882]] === 2025-04-28 === * 19:28 wmbot~jeanfred@tools-sgebastion-10: Deploy {{Gerrit|eba498e}} (Use correct namespace when using pywikibot exceptions) * 19:28 wmbot~jeanfred@tools-sgebastion-10: Deploy {{Gerrit|ea55660}} (Add unit-tests for page_saving.py functions) * 19:28 wmbot~jeanfred@tools-sgebastion-10: Deploy {{Gerrit|58918e6}} (Tweak linebreak in get_grouping_information_query which was breaking unit tests) === 2025-01-23 === * 20:43 wmbot~jeanfred@tools-sgebastion-10: Deploy {{Gerrit|a76ceec}} (Switch from Pipenv to uv as project manager) * 20:43 wmbot~jeanfred@tools-sgebastion-10: Deploy {{Gerrit|5941f3f}} (Upgrade all Python dev-dependencies) * 20:43 wmbot~jeanfred@tools-sgebastion-10: Deploy {{Gerrit|9268348}} (Fix 'QueriesTests' unit tests in test_app) * 20:43 wmbot~jeanfred@tools-sgebastion-10: Deploy {{Gerrit|9268348}} (Fix unit tests in test_app) === 2025-01-17 === * 10:03 wmbot~jeanfred@tools-sgebastion-10: Recreated virtual environment using 'toolforge webservice python3.9 shell' / 'webservice-python-bootstrap --fresh' === 2025-01-16 === * 16:36 wmbot~jeanfred@tools-sgebastion-10: Deploy {{Gerrit|525c9ed}} (Upgrade all Python dependencies) === 2024-07-01 === * 19:54 wmbot~jeanfred@tools-sgebastion-10: Deploy {{Gerrit|8eecfc6}} & {{Gerrit|0466207}} === 2024-06-06 === * 12:47 wmbot~jeanfred@tools-sgebastion-10: Deploy {{Gerrit|022fa2f}} ([[phab:T312727|T312727]]) * 12:46 wmbot~jeanfred@tools-sgebastion-10: Deploy {{Gerrit|ec28a13}} * 12:46 wmbot~jeanfred@tools-sgebastion-10: Deploy {{Gerrit|7adc07f}} === 2024-05-10 === * 15:57 wmbot~multichill@tools-bastion-12: Leaving project === 2024-05-04 === * 20:22 wmbot~jeanfred@tools-sgebastion-10: Deploy {{Gerrit|51ba697}} * 20:07 wmbot~jeanfred@tools-sgebastion-10: Deploy {{Gerrit|8ec2fec}} ([[phab:T251008|T251008]]) === 2024-05-03 === * 12:01 wmbot~multichill@tools-bastion-12: emails: all - === 2024-04-30 === * 19:10 wmbot~jeanfred@tools-sgebastion-10: Deploy {{Gerrit|5a91a78}}, {{Gerrit|864d518}}, c9b199 ([[phab:T251008|T251008]]), {{Gerrit|0dd731f}} === 2024-03-17 === * 21:02 wmbot~jeanfred@tools-sgebastion-10: Deploy {{Gerrit|3a93a59}} ([[phab:T319813|T319813]]) * 20:25 wmbot~jeanfred@tools-sgebastion-10: Re-created virtual environment for Python 3.9 ([[phab:T319813|T319813]]) * 19:26 wmbot~jeanfred@tools-sgebastion-10: Deploy {{Gerrit|bde5715}}, {{Gerrit|9bdfb02}} ([[phab:T319813|T319813]]) === 2023-11-21 === * 19:56 wm-bot: <jeanfred> Deploy {{Gerrit|992670c}} ([[phab:T351574|T351574]]) * 19:55 wm-bot: <jeanfred> Deploy {{Gerrit|1f41a79}}, {{Gerrit|38e7623}} === 2023-10-17 === * 14:30 wm-bot: <jeanfred> Deploy {{Gerrit|d7f09b3}}, {{Gerrit|eb1de62}}, {{Gerrit|f4e3656}}, {{Gerrit|82f90c4}} ([[phab:T312726|T312726]]) === 2023-10-03 === * 21:28 wm-bot: <jeanfred> Deploy {{Gerrit|3ba5e84}} ([[phab:T312729|T312729]]) * 19:28 wm-bot: <jeanfred> Deploy {{Gerrit|3ba5e84}} ([[phab:T312729|T312729]]) * 18:04 wm-bot: <jeanfred> Deploy {{Gerrit|da99820}}, {{Gerrit|801ddeb}}, {{Gerrit|f2ffd22}} ([[phab:T312729|T312729]]) === 2023-10-02 === * 09:18 wm-bot: <jeanfred> Deploy {{Gerrit|3f9a3d5}} === 2023-09-26 === * 18:25 wm-bot: <jeanfred> Deploy {{Gerrit|3caa9ed}}, {{Gerrit|a8029ed}}, {{Gerrit|c88ad35}}, {{Gerrit|e575508}}, {{Gerrit|3e3f47f}}, {{Gerrit|63ab297}}, {{Gerrit|06cb52d}}, {{Gerrit|ac64840}}, {{Gerrit|6178698}} ([[phab:T312728|T312728]]) === 2023-06-11 === * 12:50 wm-bot: <jeanfred> Deploy {{Gerrit|5aecae8}}, {{Gerrit|649259f}}, {{Gerrit|9712666}} ([[phab:T338684|T338684]]) === 2022-12-26 === * 10:48 wm-bot: <jeanfred> Deploy {{Gerrit|c90f9ef}} ([[phab:T325936|T325936]]) === 2022-09-11 === * 12:50 wm-bot: <jeanfred> Deploy {{Gerrit|bcbd306}} === 2022-07-12 === * 15:04 wm-bot: <jeanfred> Triggering full update * 15:02 wm-bot: <jeanfred> jstop {{Gerrit|5976623}} to kill weekly update job stuck since May 22 === 2022-06-03 === * 12:37 wm-bot: <jeanfred> Deploy {{Gerrit|9e85ede}} ([[phab:T309861|T309861]]) === 2022-05-15 === * 21:30 wm-bot: <jeanfred> Deploy {{Gerrit|3a675b6}} * 20:32 wm-bot: <jeanfred> Deploy {{Gerrit|80f073e}} * 19:31 wm-bot: <jeanfred> Uninstall removed requirements and restart webservice * 19:29 wm-bot: <jeanfred> Deploy {{Gerrit|f6dc34b}} * 19:15 wm-bot: <jeanfred> Install all requirements and restart webservice * 19:08 wm-bot: <jeanfred> Deploy {{Gerrit|d395b4c}}, {{Gerrit|591c450}}, {{Gerrit|6d09e5b}} === 2022-05-14 === * 14:38 wm-bot: <jeanfred> Deploy {{Gerrit|2a521d1}}, {{Gerrit|29d6b62}}, {{Gerrit|8bd793d}} ([[phab:T257942|T257942]]) === 2022-05-11 === * 19:56 wm-bot: <jeanfred> Deploy {{Gerrit|43f07d2}}, {{Gerrit|da1f2ef}}, {{Gerrit|6ab7547}}, {{Gerrit|5f6b20a}}, {{Gerrit|0df2121}} === 2022-01-28 === * 20:44 wm-bot: <jeanfred> Deploy {{Gerrit|13ed7a3}} === 2021-11-01 === * 20:30 wm-bot: <jeanfred> Deploy {{Gerrit|972a1df}} === 2021-10-30 === * 19:36 wm-bot: <jeanfred> Deploy {{Gerrit|c1628b9}} and {{Gerrit|b6940b0}} === 2021-10-29 === * 09:49 wm-bot: <jeanfred> Deploy {{Gerrit|40c8cd3}} * 09:49 wm-bot: <jeanfred> Deploy {{Gerrit|af49bab}} ([[phab:T236590|T236590]] and [[phab:T273226|T273226]]) * 09:28 wm-bot: <jeanfred> Deploy {{Gerrit|2260c8f}} (follow up to [[phab:T294570|T294570]]) * 09:16 wm-bot: <jeanfred> Install requirements in py37 venv ([[phab:T294570|T294570]]) * 09:14 wm-bot: <jeanfred> Deploy {{Gerrit|0a4560a}} and {{Gerrit|de296cd}} ([[phab:T294570|T294570]]) === 2021-10-28 === * 17:14 wm-bot: <jeanfred> Rollbacked to {{Gerrit|977236b42daef66818e561b49e64264a0c785454}} for [[phab:T294570|T294570]] and restart * 12:38 wm-bot: <jeanfred> Deploy {{Gerrit|0a8e7fb}} ([[phab:T236590|T236590]]) * 11:31 wm-bot: <jeanfred> Deploy {{Gerrit|30a2f2ba}} ([[phab:T278156|T278156]]) * 11:29 wm-bot: <jeanfred> Deploy {{Gerrit|5c547d8}} * 10:18 wm-bot: <jeanfred> Deploy {{Gerrit|b4f12bc}} to help investigate [[phab:T236960|T236960]] * 09:08 wm-bot: <jeanfred> Deploy {{Gerrit|137f5b8}}, {{Gerrit|c816e41}} {{Gerrit|8e0b9ca}}, {{Gerrit|8dbc099}} ([[phab:T273226|T273226]]) === 2021-10-27 === * 09:29 wm-bot: <jeanfred> Deploy {{Gerrit|6194d84}} ([[phab:T273226|T273226]]) === 2021-10-25 === * 21:41 wm-bot: <jeanfred> Deploy {{Gerrit|7e3343a}} ([[phab:T278156|T278156]]) * 20:22 wm-bot: <jeanfred> Deploy {{Gerrit|ef0f537}} ([[phab:T284684|T284684]]) * 20:15 wm-bot: <jeanfred> Deploy {{Gerrit|977236b}} ([[phab:T284183|T284183]]) * 18:47 wm-bot: <jeanfred> Deploy {{Gerrit|5dfef39}}, {{Gerrit|e940ae7}}, {{Gerrit|765b314}}, {{Gerrit|3e4c2de}}, {{Gerrit|c859c45}}, {{Gerrit|5526cb3}} === 2021-06-11 === * 15:32 wm-bot: <jeanfred> Deploy {{Gerrit|a753cb2}} * 15:29 wm-bot: <jeanfred> Deploy {{Gerrit|5187a4d}} === 2021-05-31 === * 09:20 wm-bot: <jeanfred> Deploy {{Gerrit|79e458d}} === 2021-05-22 === * 20:58 wm-bot: <jeanfred> Deploy {{Gerrit|258aae9a}} * 20:57 wm-bot: <jeanfred> Deploy {{Gerrit|67c42f43}} * 20:57 wm-bot: <jeanfred> Deploy {{Gerrit|d170ea1}} * 20:24 wm-bot: <jeanfred> Deploy {{Gerrit|2f39621}} * 20:21 wm-bot: <jeanfred> Deploy {{Gerrit|b52079f}} * 17:28 wm-bot: <jeanfred> Deploy {{Gerrit|8ffbad82c3}} ([[phab:T279236|T279236]]) * 16:57 wm-bot: <jeanfred> Deploy {{Gerrit|6de2308}} and {{Gerrit|c436760}} ([[phab:T279236|T279236]]) * 16:06 wm-bot: <jeanfred> Deploy {{Gerrit|0966be04}} ([[phab:T279236|T279236]]) * 15:59 wm-bot: <jeanfred> Deploy {{Gerrit|9727a01c}} * 14:02 wm-bot: <jeanfred> Deploy {{Gerrit|69048ef}} * 13:00 wm-bot: <jeanfred> Deploy {{Gerrit|633eeb0f}} * 12:59 wm-bot: <jeanfred> Deploy {{Gerrit|d4fa363}} ([[phab:T283287|T283287]]) === 2021-04-03 === * 21:34 wm-bot: <jeanfred> Deploy {{Gerrit|666222b}} and {{Gerrit|0dcafff}} ([[phab:T279236|T279236]]) * 21:33 wm-bot: <jeanfred> Deploy latest from Git master: {{Gerrit|eac4d2c}} ([[phab:T257942|T257942]]) * 16:45 wm-bot: <jeanfred> Deploy latest from Git master: {{Gerrit|eac4d2c}} ([[phab:T257942|T257942]]) === 2021-03-26 === * 20:57 wm-bot: <jeanfred> Reinstalled requirements in both virtual environments for [[phab:T240312|T240312]] * 20:55 wm-bot: <jeanfred> Deploy latest from Git master: {{Gerrit|bcc4fe41}} ([[phab:T240312|T240312]]) === 2020-07-06 === * 22:55 wm-bot: <root> Migrated .webservicerc to service.template ([[phab:T257229|T257229]]) === 2020-06-18 === * 11:01 wm-bot: <jeanfred> Restarted with --canonical for Toolforge domain migration === 2020-05-01 === * 22:59 wm-bot: <jeanfred> Deploy latest from Git master: {{Gerrit|f0db935}}, {{Gerrit|f8d9fdf}}, {{Gerrit|f57c060}}, {{Gerrit|6c534b7}} ([[phab:T248788|T248788]]) === 2020-04-29 === * 20:33 wm-bot: <jeanfred> Deploy latest from Git master: {{Gerrit|eabbac8}}, {{Gerrit|b74b352}} ([[phab:T248788|T248788]]) === 2020-04-24 === * 22:22 wm-bot: <jeanfred> Deploy latest from Git master: {{Gerrit|6963040}}, {{Gerrit|35a3ea7}}, {{Gerrit|22e27b5}}, {{Gerrit|9dd44ff}}, {{Gerrit|0e45815}}, {{Gerrit|6f36ffd}} ([[phab:T248788|T248788]]) * 19:02 wm-bot: <jeanfred> Deploy latest from Git master: {{Gerrit|9fea19b}}, {{Gerrit|f38047e}} === 2020-02-25 === * 23:35 wm-bot: <root> Migrated to 2020 Kubernetes cluster === 2020-02-16 === * 16:49 JeanFred: "Deploy latest from Git master: {{Gerrit|0fcb522}} ([[phab:T243780|T243780]]), {{Gerrit|4cb6a5a}} ([[phab:T243780|T243780]], [[phab:T239067|T239067]])" === 2020-02-11 === * 21:23 wm-bot157: <jeanfred> Deploy latest from Git master: {{Gerrit|c350e4f}} ([[phab:T243998|T243998]]), {{Gerrit|fa12c21}} === 2020-02-05 === * 12:40 wm-bot: <jeanfred> Deploy latest from Git master: {{Gerrit|8977958}} ([[phab:T243780|T243780]]) === 2020-02-03 === * 16:57 wm-bot: <jeanfred> Deploy latest from Git master: {{Gerrit|3b2867f}}, {{Gerrit|10f67fd}}, {{Gerrit|2e6a669}} ([[phab:T244030|T244030]]), {{Gerrit|279c92b}}, {{Gerrit|de6308e}} === 2020-01-16 === * 21:44 wm-bot: <jeanfred> Perform webservice restart for [[phab:T242967|T242967]] === 2019-12-11 === * 15:15 wm-bot: <jeanfred> Deploy latest from Git master: {{Gerrit|257618f1}} ([[phab:T240312|T240312]]) === 2019-12-02 === * 06:40 wm-bot: <jeanfred> Deploy latest from Git master: {{Gerrit|770cc93}} ([[phab:T237182|T237182]]) === 2019-11-25 === * 11:36 wm-bot: <jeanfred> Trigger a manual update of all dashboards ([[phab:T239085|T239085]]) * 11:34 wm-bot: <jeanfred> Deploy latest from Git master: {{Gerrit|684fb77}} ([[phab:T239085|T239085]]) === 2019-11-21 === * 08:46 wm-bot: <jeanfred> Deploy latest from Git master: {{Gerrit|2f94256}} ([[phab:T237187|T237187]]) === 2019-11-10 === * 15:31 wm-bot: <jeanfred> Deploy latest from Git master: {{Gerrit|ab714de}} ([[phab:T224226|T224226]]) === 2019-11-06 === * 09:12 wm-bot: <jeanfred> Deploy latest from Git master: {{Gerrit|7202182}} ([[phab:T237187|T237187]]) * 09:01 wm-bot: <jeanfred> Deploy latest from Git master: {{Gerrit|29b19c3}} ([[phab:T237182|T237182]]) === 2019-11-04 === * 17:46 wm-bot: <jeanfred> Deploy latest from Git master: {{Gerrit|d0f4222}} ([[phab:T237189|T237189]]) * 17:35 wm-bot: <jeanfred> Deploy latest from Git master: {{Gerrit|588dae8}} * 17:35 wm-bot: <jeanfred> Deploy latest from Git master: {{Gerrit|471deddd}} ([[phab:T237182|T237182]]) === 2019-10-31 === * 09:06 wm-bot: <jeanfred> Deploy latest from Git master: {{Gerrit|46484f5}} ([[phab:T224226|T224226]]) * 09:02 wm-bot: <jeanfred> Deploy latest from Git master: {{Gerrit|2d71154}}, {{Gerrit|d0d2937}}, {{Gerrit|905007b}}, {{Gerrit|effee36}}, {{Gerrit|7c15083}} === 2019-10-30 === * 19:42 wm-bot: <jeanfred> Deploy latest from Git master: {{Gerrit|289a41b}}, {{Gerrit|02a1a4f}}, {{Gerrit|e8d4363}} ([[phab:T224226|T224226]]) === 2019-10-25 === * 15:02 wm-bot: <jeanfred> Deploy latest from Git master: {{Gerrit|82de2957}} ([[phab:T224212|T224212]], [[phab:T223930|T223930]]) * 14:08 wm-bot: <jeanfred> Deploy latest from Git master: {{Gerrit|6a055c2}} ([[phab:T228405|T228405]]) * 13:24 wm-bot: <jeanfred> Deploying from master, will break things. === 2019-09-19 === * 18:38 wm-bot: <jeanfred> Deploy latest from Git master: {{Gerrit|d214f09}}, {{Gerrit|9b8dfa5}} * 17:27 wm-bot: <jeanfred> Deploy latest from Git master: {{Gerrit|12ebc39}} === 2019-07-07 === * 22:49 wm-bot: <jeanfred> Deploy latest from Git master: {{Gerrit|5b4a1c6}}, {{Gerrit|ab90227}}, {{Gerrit|d101ca1}}, {{Gerrit|394be6a}}, {{Gerrit|ae16d68}}, {{Gerrit|728a8d2}} === 2019-05-30 === * 12:08 wm-bot: <jeanfred> Nuked the virtualenv and reinstalled all deps from scratch, in desperation for [[phab:T224651|T224651]] * 10:59 wm-bot: <jeanfred> Service stop, mv logs, service start for [[phab:T224651|T224651]] === 2019-05-22 === * 16:39 wm-bot: <jeanfred> Deploy latest from Git master: {{Gerrit|8e76a8a}} * 13:13 wm-bot: <jeanfred> Deploy latest from Git master: {{Gerrit|2fdede2}}, {{Gerrit|6b6c07c}}, {{Gerrit|c17cb78}} ([[phab:T224090|T224090]]) === 2019-05-20 === * 19:01 wm-bot: <jeanfred> Restart uwsgi with 10 workers rather than the default of 4 * 19:00 wm-bot: <jeanfred> --help <noinclude>[[Category:SAL]]</noinclude> 75itibfoq5d88tonebss1k6uuofgp7q Nova Resource:Tools.phab-ban/SAL 498 444737 2450667 2350805 2026-08-22T22:26:47Z Stashbot 7414 wmbot~bd808@tools-bastion-14: `webservice restart` after user reports of 500 responses. Logging was not helpful for a cause. 2450667 wikitext text/x-wiki === 2026-08-22 === * 22:26 wmbot~bd808@tools-bastion-14: `webservice restart` after user reports of 500 responses. Logging was not helpful for a cause. === 2025-10-14 === * 12:30 wmbot~lucaswerkmeister@tools-bastion-15: webservice restart # tool was seemingly down since yesterday because an error "Failed to resolve 'wikitech.wikimedia.org' ([Errno -3] Temporary failure in name resolution)" prevented the app from starting up === 2025-08-22 === * 22:04 wmbot~bd808@tools-bastion-12: Restarted after irc report of 500 error output. Looks like an NFS blip messed things up at last startup. === 2024-10-01 === * 21:45 wmbot~bd808@tools-bastion-12: Rotate OAuth credentials ([[phab:T376216|T376216]]) === 2024-06-04 === * 17:01 wmbot~bd808@tools-bastion-12: Restart to switch to a new API token ([[phab:T366587|T366587]]) * 16:47 wmbot~bd808@tools-bastion-12: Hard stop/start cycle for webservice to make sure that the code in $HOME/www/python/src is the code running for the tool. ([[phab:T366587|T366587]]) === 2023-09-01 === * 17:44 wm-bot: <bd808> Update to {{Gerrit|af3bef40}} ([[phab:T317584|T317584]]) * 17:38 wm-bot: <bd808> Testing https://gitlab.wikimedia.org/toolforge-repos/phab-ban/-/merge_requests/5 * 16:40 wm-bot: <bd808> Update to {{Gerrit|353fe3a4}} * 16:23 wm-bot: <bd808> Update to {{Gerrit|98817923}} ([[phab:T200856|T200856]]) * 16:09 wm-bot: <bd808> Testing https://gitlab.wikimedia.org/toolforge-repos/phab-ban/-/merge_requests/3 === 2023-05-09 === * 23:43 wm-bot: <bd808> Updated to {{Gerrit|e1b15fd}} and switched to python3.9 runtime === 2023-04-24 === * 08:58 wm-bot: <lucaswerkmeister> restarted k8s deployment, webservice was unresponsive === 2022-09-08 === * 21:07 wm-bot: <bd808> Hard stop + start because restart did not terminate existing pod (probably a k8s object labeling issue?) * 21:04 wm-bot: <bd808> Updated to {{Gerrit|7956e6c}} === 2020-06-06 === * 15:38 wm-bot84: <bd808> Upgraded to Python 3.7 and --canonical * 03:40 wm-bot84: <bd808> App busted due to attempted python3.7 upgrade being interruped by a power outage at my house. === 2020-01-14 === * 23:02 bd808: Rotated conduit API token === 2020-01-03 === * 21:32 bd808: Migrating to new kubernetes cluster <noinclude>[[Category:SAL]]</noinclude> aeexsy5fp6fehyi0zz3xovdsu2vxwbv Docker-registry 0 447352 2450662 2449082 2026-08-22T19:04:12Z Quiddity 1884 lang="text" 2450662 wikitext text/x-wiki {{Kubernetes nav}} We run our own Docker registry at [https://docker-registry.wikimedia.org docker-registry.wikimedia.org]. Internally the domain <code>docker-registry.discovery.wmnet</code> is also used. The registry is used by our k8s cluster, CI, and local development. It is highly available (<code>docker_registry_ha</code> Puppet module) and backed by [[Swift]]. Although we run it '''active/passive''' because of the swift replication lag. The docker-registry nodes consist of the '''docker registry''' itself as well as an '''nginx reverse-proxy''' in front to handle '''authentication''' as well as '''local caching'''. == Browsing == Visit https://docker-registry.wikimedia.org/ to see a list of images and their tags. The listing is updated on a hourly timer and is done by the <code>registry-homepage-builder.py</code> script in Puppet. == Downloading images == Despite the name, the docker-registry is usable by any OCI container tool, including podman. Nearly all images may be publicly downloaded, examined, run, etc. The only exception is images under the <code>restricted/</code> namespace, which contain non-disclosed security patches and require specific credentials to fetch. Kubernetes nodes use [[Dragonfly]] to pull images. == Uploading images == For services we recommend using the [[Deployment pipeline]] which is [[Blubber]]. For other docker images, like infrastructure images, we manage them using docker-pkg, see: [[Kubernetes/Images#Image_building]] Hosts that want to upload images must be individually listed in Puppet hiera. == Access control == The upstream docker-registry software provides no access control, so it is implemented at the nginx level, which restricts GET/POST/etc. requests accordingly. As of 2021-03-18, the following accounts exist: * <code>ci-restricted</code>: Can pull and push any image (including "restricted/"). Used by releases servers that build the restricted MediaWiki production image. * <code>ci-build</code>: Can pull and push any non-restricted image. Used by contint servers via docker-pkg and the deployment pipeline. * <code>prod-build</code>: Can pull and push any non-restricted image. Used by build2001.codfw.wmnet via docker-pkg and build-base-images. * <code>kubernetes</code>: Can pull any image (including "restricted/"). Used by k8s nodes to pull images, including the restricted MediaWiki production image. **See [[Kubernetes/Clusters/New#Access to restricted docker images]] for more details. The passwords are all deployed using the private puppet repo. In case rotation is needed (e.g. compromise), grepping for <code><name>_user_password</code> should find all uses (switch hyphens to underscores). === jwt-authorizer === docker-registry also supports authorization using JSON Web Tokens. A dedicated daemon is running which handles jwt validation. See [[Docker-registry/jwt-authorizer]] for more information. == Storage backends == We are currently (August 2026) migrating away from a single Swift bucket, towards this model: * An S3 bucket for /v2/restricted.* images. * An S3 bucket for /v2/wikimedia/machinelearning.* images (and their vllm base images). * An S3 bucket for /v2/releng.* images. * An S3 bucket for all the other images running on Kubernetes clusters. Every S3 bucked is managed by a Docker Distribution instance, and the Nginx proxy in front of them routes requests based on their prefix name. Everything is tracked in [[phab:T427175|https://phabricator.wikimedia.org/T427175.]] We use skopeo (from Debian's upstream) to copy over images from one Docker Distribution instance to the other one, preserving the sha256 signatures of the blobs etc.. This is very handy when migrating images from Swift to a new S3 bucket! If it happens that an image has not been copied over and the Nginx config for a given prefix is already pointing to the new S3 bucket/distribution combo, it is sufficient to use skopeo to quickly fill the gap. For example, let's say that you get a report of <code>docker-registry.wikimedia.org/releng/rake-ruby3.3:0.1.0</code> (note: name and tag) missing from the registry (docker pull errors etc..), you can do the following: * ssh to registry1004 (standby host) * Check if the image is present on the Swift backend <syntaxhighlight lang="bash"> # Note: From docker-registry.wikimedia.org/releng/rake-ruby3.3:0.1.0 # you need to replace "docker-registry.wikimedia.org" with "/v2" # Note2: localhost:5000 is the address of the Docker Distribution backend for Swift elukey@registry1004:~$ curl http://localhost:5000/v2/releng/rake-ruby3.3/tags/list {"name":"releng/rake-ruby3.3","tags":["0.1.0","latest"]} </syntaxhighlight> * If it is not present, the issue may not be related to the S3 migration. If it is present, then you can execute the following command to copy the image layers / manifests / etc.. over: <syntaxhighlight lang="text"> # Note: localhost:5006 is the address of the Releng Docker Distribution backend skopeo copy --all --preserve-digests docker://localhost:5000/releng/rake-ruby3.3:0.1.0 docker://localhost:5006/releng/rake-ruby3.3:0.1.0 --dest-tls-verify=false --src-tls-verify=false </syntaxhighlight>That's it! The image should be available again for pulls. == Deleting images == To delete an image entirely, you may use the tool <code>docker-registryctl</code> on the current build host. It will do it's best to remove the tags/image from the registry, despite the [[phab:T242604|circumstances]]. '''Note:''' the domain used here is important. <code>discovery.wmnet</code> has to be used, so you will have to adjust it if you copy and paste from the browser UI.<syntaxhighlight lang="bash"> elukey@build2001:~$ sudo -i docker-registryctl delete-tags docker-registry.discovery.wmnet/wikimedia/machinelearning-liftwing-inference-services We're about to delete the following tags for image docker-registry.discovery.wmnet/wikimedia/machinelearning-liftwing-inference-services: 2021-07-28-175322-production stable Ok to proceed? (y/n)y docker-registry.discovery.wmnet/wikimedia/machinelearning-liftwing-inference-services:2021-07-28-175322-production[DONE] docker-registry.discovery.wmnet/wikimedia/machinelearning-liftwing-inference-services:stable[GONE] </syntaxhighlight> == Maintenance == The docker-registry nodes consist of the '''docker registry''' itself as well as an '''nginx reverse-proxy''' in front to handle '''authentication''' as well as '''local caching'''. === Service management === The following units run on the registry nodes: {| class="wikitable" ! Service unit !! Description |- | <code>docker-registry-ha-jwt.service</code> || JSON Web Token authoriser |- | <code>docker-registry-swift.service</code> || Registry instance backed by Swift |- | <code>docker-registry-ml.service</code> || Registry instance serving the ML/LiftWing namespace |- | <code>docker-registry-restricted.service</code> || Registry instance serving restricted images |} ==== Restarting ==== Before restarting any service, confirm which node is currently active and prefer restarting the passive node where possible to avoid unnecessary disruption. cumin1003:$ sudo confctl --object-type discovery select 'dnsdisc=docker-registry' get <syntaxhighlight lang="bash"> sudo systemctl restart docker-registry-ha-jwt.service sudo systemctl restart docker-registry-swift.service sudo systemctl restart docker-registry-ml.service sudo systemctl restart docker-registry-restricted.service </syntaxhighlight> To check status: <syntaxhighlight lang="bash"> sudo systemctl status docker-registry-ha-jwt.service </syntaxhighlight> Note that nginx sits in front of all registry instances and handles authentication and local caching. <syntaxhighlight lang="bash"> sudo systemctl status nginx.service sudo systemctl restart nginx.service </syntaxhighlight> === Post-restart verification - httpbb === Run the httpbb tests against the restarted node to confirm it is serving correctly. Make sure the instance is '''not''' in read-only mode first (<code>profile::docker_registry::read_only_mode: false</code>): <syntaxhighlight lang="bash"> sudo httpbb /srv/deployment/httpbb-tests/docker-registry/test_docker-registry.yaml --hosts 'registry2004.codfw.wmnet' </syntaxhighlight> == See also == * [[/Runbook]] (draft) *[[SLO/Docker-registry|Service Level Objective (SLO)]] * [[Docker]] * [https://github.com/distribution/distribution Upstream Git repository] *[[User:JMeybohm/Docker-Registry-Stresstest]] [[Category:Containers]] [[Category:Docker]] jol2q5osni4ol2ud5meehqwdngefz5g Portal:Cloud VPS/Admin/Runbooks/Cloud VPS alert Puppet failure on 0 447668 2450663 2449290 2026-08-22T19:04:42Z Quiddity 1884 lang="text" 2450663 wikitext text/x-wiki {{Cloud VPS nav}} {{Requires admin permissions banner}} '''Note:''' some of these steps might require extra access to internal infrastructure systems, we are working on improving the runbooks, until then, take this as a guideline. == Error == Usually an email with the subject: [Cloud VPS alert] Puppet failure on <hostname> For example: Subject: [Cloud VPS alert] Puppet failure on toolsbeta-sgeexec-1001.toolsbeta.eqiad1.wikimedia.cloud == Debugging == If there's only a few of those emails, the error is most probably on the client side and/or affecting only a limited amount of hosts. Otherwise it might indicate a wider issue. Usually you would want to retry the run on the failed machine, so ssh to it and run: dcaro@vulcanus$ ssh toolsbeta-sgeexec-1001.toolsbeta.eqiad1.wikimedia.cloud Linux toolsbeta-sgeexec-1001 4.19.0-16-cloud-amd64 #1 SMP Debian 4.19.181-1 (2021-03-19) x86_64 Debian GNU/Linux 10 (buster) The last Puppet run was at Sun Mar 28 15:16:07 UTC 2021 (26947 minutes ago). Last puppet commit: Last login: Thu Apr 15 08:23:23 2021 from 172.16.1.135 dcaro@toolsbeta-sgeexec-1001:~$ sudo run-puppet-agent From there there's a wide variety of puppet issues that might happen, some common ones follow. == Usual errors == '''TODO:''' Fill up as you encounter errors. === Error: The CRL issued by 'CN=Puppet CA: puppet' has expired, verify time is synchronized === Run the following command as root: <syntaxhighlight lang="text"> rm /var/lib/puppet/ssl/crl.pem </syntaxhighlight> and then <code>run-puppet-agent</code> as usual. === Failed resources if any: Exec[create-volume-group] === If you see the above text in your e-mail alert, it means you have the [[Help:Adding_Disk_Space_to_Cloud_VPS_instances#With_LVM_(deprecated_as_of_February,_2021)|legacy LVM Puppet role enabled project-wide.]] Solution is to remove the role from the project puppet: *log in to https://horizon.wikimedia.org/; *go to ''Puppet'' > ''Project Puppet''; *click on the ''Edit'' button below ''Puppet Classes''; *delete the <code>role::labs::lvm::srv</code> line; *click on the ''Apply Changes'' button. Done! === Function lookup() did not find a value for the name === This error usually means that there's a default value missing on some puppetclass parameter or some value missing from hiera, an example of the error: Error: Could not retrieve catalog from remote server: Error 500 on SERVER: Server Error: Evaluation Error: Error while evaluating a Resource Statement, Function lookup() did not find a value for the name 'profile::logstash::apifeatureusage::curator_actions' (file: /etc/puppet/modules/profile/manifests/logstash/apifeatureusage.pp, line: 7) on node deployment-logstash03.deployment-prep.eqiad.wmflabs From there you have the missing parameter <code>profile::logstash::apifeatureusage::curator_actions</code> and the class that's missing it <code>/etc/puppet/modules/profile/manifests/logstash/apifeatureusage.pp, line 7</code>, so, if the project is using the same puppet as production (as it's the case for the example), you can [https://gerrit.wikimedia.org/g/operations/puppet/+/refs/heads/production/ see that code], for that file and line [https://gerrit.wikimedia.org/g/operations/puppet/+/refs/heads/production/modules/profile/manifests/logstash/apifeatureusage.pp#7 at apifeatureusage.pp#7]. You can see there that there's a parameter called <code>curator_actions</code> that does the lookup for that variable but has no default: class profile::logstash::apifeatureusage( Array[Stdlib::Host] $targets = lookup('profile::logstash::apifeatureusage::targets'), Hash $curator_actions = lookup('profile::logstash::apifeatureusage::curator_actions'), ) { '''One solution (not the correct one in this case) is to add a default value''' there: class profile::logstash::apifeatureusage( Array[Stdlib::Host] $targets = lookup('profile::logstash::apifeatureusage::targets'), Hash $curator_actions = lookup('profile::logstash::apifeatureusage::curator_actions', {'default_value' => {}}), ) { '''Another way of fixing the issue, is defining that value in Horizon''', under the [https://horizon.wikimedia.org/project/puppet/ project puppet] page, or for the [https://horizon.wikimedia.org/project/prefixpuppet/ specific prefix] for that VPS. If that's not possible, then we have to look deeper. In this case doing a quick <code>git grep curator_actions</code> shows that it's defined in the file <code>hieradata/role/common/logstash.yaml</code>: 10:50 AM ~/Work/wikimedia/operations-puppet (production|✔) dcaro@vulcanus$ git grep curator_actions hieradata/role/common/logstash.yaml:profile::logstash::apifeatureusage::curator_actions: hieradata/role/common/logstash/elasticsearch7.yaml:profile::elasticsearch::logstash::curator_actions: ... Let's see why it's not using that. ==== Checking up the puppetmaster ==== In order to do some extra debuggin and find out who the puppetmaster is, you can do the following: dcaro@deployment-logstash03:~$ host 172.16.0.38 38.0.16.172.in-addr.arpa domain name pointer cloud-puppetmaster-03.cloudinfra.eqiad1.wikimedia.cloud. dcaro@deployment-logstash03:~$ sudo puppet config print ca_server puppet dcaro@deployment-logstash03:~$ host puppet puppet has address 172.16.0.38 puppet has address 172.16.0.38 puppet has address 172.16.0.38 dcaro@deployment-logstash03:~$ host 172.16.0.38 38.0.16.172.in-addr.arpa domain name pointer cloud-puppetmaster-03.cloudinfra.eqiad1.wikimedia.cloud. So now we now that the puppetmaster for this instance is <code>cloud-puppetmaster-03.cloudinfra.eqiad1.wikimedia.cloud</code>. This is the common master (cloudinfra) that any VPS will use by default, when they don't have a dedicated one. ==== Hiera lookups work different in cloud than prod ==== Looking on that master, we can check the hiera config (<code>/etc/puppet/hiera.yaml</code>) for the order on which the data is loaded: ... hierarchy: - name: 'Http Yaml' data_hash: cloudlib::httpyaml uri: "http://puppet-enc.cloudinfra.wmcloud.org:8100/v1/%{::wmcs_project}/node/%{facts.fqdn}" - name: "cloud hierarchy" paths: - "cloud/%{::wmcs_deployment}/%{::wmcs_project}/hosts/%{::hostname}.yaml" - "cloud/%{::wmcs_deployment}/%{::wmcs_project}/common.yaml" - "cloud/%{::wmcs_deployment}.yaml" - "cloud.yaml" - name: "Secret hierarchy" paths: - "hosts/%{::trusted.certname}.yaml" - "%{::wmcs_project}.yaml" datadir: "/etc/puppet/secret/hieradata" - name: "Private hierarchy" paths: - "labs/%{::wmcs_project}/common.yaml" - "%{::wmcs_project}.yaml" - "labs.yaml" datadir: "/etc/puppet/private/hieradata" - name: "Common hierarchy" path: "common.yaml" - name: "Secret Common hierarchy" path: "common.yaml" datadir: "/etc/puppet/secret/hieradata" - name: "Private Common hierarchy" path: "common.yaml" datadir: "/etc/puppet/private/hieradata" Comparing that with the production puppet <code>hiera.yaml</code>, we see that it's missing the role hierarchy: 29 - name: "role" 30 paths: 31 - "role/%{::site}/%{::_role}.yaml" 32 - "role/common/%{::_role}.yaml" So '''another solution would be to add the value also to a yaml file that would be picked up by puppet''', checking for the other parameter in that class in the git repo, we find that it's defined also on: 12:14 PM ~/Work/wikimedia/operations-puppet (production|✔) dcaro@vulcanus$ git grep apifeatureusage::targets hieradata/cloud/eqiad1/deployment-prep/common.yaml:profile::logstash::apifeatureusage::targets: ... So adding it there also would solve the issue. '''TODO:''' Maybe add roles to the cloud hiera lookup (T280324) == Contacts == {{:Help:Cloud Services communication}} == More info == * [https://wikitech.wikimedia.org/wiki/Puppet General puppet info at wikimedia] * [https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Hiera Cloud specific hiera/enc info] * [https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Puppet_testing Cloud specific testing] * [https://wikitech.wikimedia.org/wiki/Help:Standalone_puppetmaster Cloud per-project puppetmaster info] == Related tasks == * [https://phabricator.wikimedia.org/T276040 Puppet failure on toolsbeta-bastion-05.toolsbeta.eqiad1.wikimedia.cloud] * [https://phabricator.wikimedia.org/T276039 Puppet failure on clouddb1003.clouddb-services.eqiad.wmflabs] * [https://phabricator.wikimedia.org/T285839 Puppet failure on fullstackd-20210630081316.admin-monitoring.eqiad1.wikimedia.cloud] [[Category:Runbooks]] 325csia6l0u0a2kv119vqurnrsl03pf Map of database maintenance 0 449160 2450668 2450632 2026-08-23T00:01:40Z Dexbot 30554 Bot: Updating the report 2450668 wikitext text/x-wiki {{/Header}} == Today (2026-08-23) == == Yesterday (2026-08-22) == == Last seven days == {| class="wikitable" |+ codfw |- ! Section !! Work |- | s2 || [[phab:T435270|Switchover s2 master (db2207 -&gt; db2204) (T435270)]] (marostegui) |- |} [[Category:MariaDB]] j9m8yzlcq95mpbsijs068xh3rs36dc7 Portal:Toolforge/Admin/Build Service 0 452173 2450665 2379501 2026-08-22T21:11:54Z BryanDavis 1604 /* Adding a new buildpack */ typo 2450665 wikitext text/x-wiki {{Toolforge nav admin}} {{ptag|tbs}} Documentation of components and common admin procedures for [[Help:Toolforge/Build Service|Toolforge/Build Service]]. == Components == * Builds client ([https://gitlab.wikimedia.org/repos/cloud/toolforge/builds-cli source code]): main entrypoint for users * Builds API ([https://gitlab.wikimedia.org/repos/cloud/toolforge/builds-api source code]): main entry-point for clients (users use the cli) * Builds builder ([https://gitlab.wikimedia.org/repos/cloud/toolforge/builds-builder source code]): underlying building system * [[Portal:Toolforge/Admin/Harbor | Harbor]] ([https://gerrit.wikimedia.org/r/plugins/gitiles/operations/puppet/+/refs/heads/production/modules/profile/manifests/toolforge/harbor.pp puppet code]): Hosts the images the users create ** Tools: https://tools-harbor.wmcloud.org ** Toolsbeta: https://toolsbeta-harbor.wmcloud.org [[File:TBS_components_diagram.png|frameless|650px]] == Alerts == * From the cloud UI: https://prometheus-alerts.wmcloud.org/?q=%40state%3Dactive&q=project%3D~^%28tools%7Ctoolsbeta%29 * From the prod UI: https://alerts.wikimedia.org/?q=team%3Dwmcs&q=project%3D~%28tools%7Ctoolsbeta%29 == Dashboards == You can find all the current dashboards here: * https://grafana.wmcloud.org/d/m9V1RQs4k/harbor-overview * https://grafana.wmcloud.org/d/f4Sxgf-Sz/builds-api * https://grafana.wmcloud.org/d/2idBoRUVz/tekton-overview == Administrative tasks == {{See also|Portal:Toolforge/Admin/Runbooks#Existing_runbooks}} === Starting a service === ==== Harbor ==== Ssh to the harbor instance (ex. <code>toolsbeta-harbor-1.toolsbeta.eqiad1.wikimedia.cloud</code>): {{Codesample | code = dcaro@vulcanus$ wm-ssh toolsbeta-harbor-1.toolsbeta.eqiad1.wikimedia.cloud ... dcaro@toolsbeta-harbor-1:~$ sudo -i root@toolsbeta-harbor-1:~# cd /srv/ops/harbor/ root@toolsbeta-harbor-1:/srv/ops/harbor# docker-compose up -d # will start the containers that are down if any harbor-log is up-to-date registry is up-to-date redis is up-to-date harbor-portal is up-to-date registryctl is up-to-date harbor-core is up-to-date harbor-jobservice is up-to-date nginx is up-to-date harbor-exporter is up-to-date | lang = shell }} ==== Buildservice API ==== This lives in kubernetes, behind the API gateway. To start it you can try redepolying it, to do so follow [[Portal:Toolforge/Admin/Kubernetes#Deploy_new_version]] (the component is toolforge-builds-api). You can monitor if it's coming up with the usual k8s commands: {{codesample | code = root@toolsbeta-test-k8s-control-4:~# kubectl get all -n builds-api NAME READY STATUS RESTARTS AGE pod/builds-api-5bffd6b58f-9zg4s 2/2 Running 0 29h pod/builds-api-5bffd6b58f-jk6sf 2/2 Running 0 29h NAME TYPE CLUSTER-IP EXTERNAL-IP PORT(S) AGE service/builds-api ClusterIP 10.97.55.43 <none> 8443/TCP 18d NAME READY UP-TO-DATE AVAILABLE AGE deployment.apps/builds-api 2/2 2 2 18d NAME DESIRED CURRENT READY AGE replicaset.apps/builds-api-5bffd6b58f 2 2 2 29h | lang = shell }} ==== Tekton ==== Similar to the builds api, tekton is a k8s component, you can try redepolying it too following [[Portal:Toolforge/Admin/Kubernetes/Components#Deploy]] (the component is buildservice). You can monitor if it's coming up with the usual k8s commands: {{codesample | code = root@toolsbeta-test-k8s-control-4:~# kubectl get all -n tekton-pipelines NAME READY STATUS RESTARTS AGE pod/tekton-pipelines-controller-5c78ddd49b-dj4hz 1/1 Running 0 57d pod/tekton-pipelines-webhook-5d899cc8c-zwf7p 1/1 Running 0 57d NAME TYPE CLUSTER-IP EXTERNAL-IP PORT(S) AGE service/tekton-pipelines-controller ClusterIP 10.96.176.235 <none> 9090/TCP,8008/TCP,8080/TCP 447d service/tekton-pipelines-webhook ClusterIP 10.101.163.215 <none> 9090/TCP,8008/TCP,443/TCP,8080/TCP 447d NAME READY UP-TO-DATE AVAILABLE AGE deployment.apps/tekton-pipelines-controller 1/1 1 1 87d deployment.apps/tekton-pipelines-webhook 1/1 1 1 87d NAME DESIRED CURRENT READY AGE replicaset.apps/tekton-pipelines-controller-5c78ddd49b 1 1 1 87d replicaset.apps/tekton-pipelines-webhook-5d899cc8c 1 1 1 87d NAME REFERENCE TARGETS MINPODS MAXPODS REPLICAS AGE horizontalpodautoscaler.autoscaling/tekton-pipelines-webhook Deployment/tekton-pipelines-webhook 4%/100% 1 5 1 447d | lang = shell }} === Stopping a service === ==== Harbor ==== Ssh to the harbor instance (ex. <code>toolsbeta-harbor-1.toolsbeta.eqiad1.wikimedia.cloud</code>): {{Codesample | code = dcaro@vulcanus$ wm-ssh toolsbeta-harbor-1.toolsbeta.eqiad1.wikimedia.cloud ... dcaro@toolsbeta-harbor-1:~$ sudo -i root@toolsbeta-harbor-1:~# cd /srv/ops/harbor/ oot@toolsbeta-harbor-1:/srv/ops/harbor# docker-compose stop Stopping harbor-jobservice ... done Stopping nginx ... done Stopping harbor-exporter ... done Stopping harbor-core ... done Stopping registry ... done Stopping harbor-portal ... done Stopping redis ... done Stopping registryctl ... done Stopping harbor-log ... done | lang = shell }} ==== Buildservice API ==== Being a k8s deployment, the quickest way might be just to remove the deployment itself (will require redeploying to start again). {{codesample | code = root@toolsbeta-test-k8s-control-4:~# kubectl get deployment -n builds-api builds-api -o yaml > backup.yaml # in case you want to restore later with kubectl apply -f backup.yaml root@toolsbeta-test-k8s-control-4:~# kubectl delete deployment -n builds-api builds-api | lang = shell }} For a full removal (CAREFUL! Only if you know what you are doing) you can use helm: {{codesample | code = root@toolsbeta-test-k8s-control-4:~# helm uninstall -n builds-api builds-api | lang = shell }} ==== Tekton ==== This one is a bit more tricky, but it would be removing the tekton controller itself (the one that handles the <code>PipelineRun</code> and <code>TaskRun</code> resources). {{codesample | code = root@toolsbeta-test-k8s-control-4:~# kubectl get deployment -n tekton-pipelines tekton-pipelines-controller -o yaml > backup.yaml # in case you want to restore later with kubectl apply -f backup.yaml root@toolsbeta-test-k8s-control-4:~# kubectl delete deployment -n tekton-pipelines tekton-pipelines-controller | lang = shell }} NOTE: Tekton does not have yet a helm deployment associated with it === Checking all components are alive === You can check [[Portal:Toolforge/Admin/Build_Service#Dashboards|the dashboards]]. For cli-based keep reading: ==== Harbor ==== Ssh to the harbor instance (ex. <code>toolsbeta-harbor-1.toolsbeta.eqiad1.wikimedia.cloud</code>): {{Codesample | code = dcaro@vulcanus$ wm-ssh toolsbeta-harbor-1.toolsbeta.eqiad1.wikimedia.cloud ... dcaro@toolsbeta-harbor-1:~$ sudo -i root@toolsbeta-harbor-1:~# cd /srv/ops/harbor/ root@toolsbeta-harbor-1:/srv/ops/harbor# docker-compose ps Name Command State Ports ---------------------------------------------------------------------------------------------------------------- harbor-core /harbor/entrypoint.sh Up (healthy) harbor-exporter /harbor/entrypoint.sh Up harbor-jobservice /harbor/entrypoint.sh Up (healthy) harbor-log /bin/sh -c /usr/local/bin/ ... Up (healthy) 127.0.0.1:1514->10514/tcp harbor-portal nginx -g daemon off; Up (healthy) nginx nginx -g daemon off; Up (healthy) 0.0.0.0:80->8080/tcp, 0.0.0.0:9090->9090/tcp redis redis-server /etc/redis.conf Up (healthy) registry /home/harbor/entrypoint.sh Up (healthy) registryctl /home/harbor/start.sh Up (healthy) | lang = shell }}For '''Harbor project quota management''', see [[Portal:Toolforge/Admin/Harbor#Quota management]] ==== Buildservice API ==== You can monitor if it's coming up with the usual k8s commands: {{codesample | code = root@toolsbeta-test-k8s-control-4:~# kubectl get all -n builds-api NAME READY STATUS RESTARTS AGE pod/builds-api-5bffd6b58f-9zg4s 2/2 Running 0 29h pod/builds-api-5bffd6b58f-jk6sf 2/2 Running 0 29h NAME TYPE CLUSTER-IP EXTERNAL-IP PORT(S) AGE service/builds-api ClusterIP 10.97.55.43 <none> 8443/TCP 18d NAME READY UP-TO-DATE AVAILABLE AGE deployment.apps/builds-api 2/2 2 2 18d NAME DESIRED CURRENT READY AGE replicaset.apps/builds-api-5bffd6b58f 2 2 2 29h | lang = shell }} ==== Tekton ==== Same as before, different namespace: {{codesample | code = root@toolsbeta-test-k8s-control-4:~# kubectl get all -n tekton-pipelines NAME READY STATUS RESTARTS AGE pod/tekton-pipelines-controller-5c78ddd49b-dj4hz 1/1 Running 0 57d pod/tekton-pipelines-webhook-5d899cc8c-zwf7p 1/1 Running 0 57d NAME TYPE CLUSTER-IP EXTERNAL-IP PORT(S) AGE service/tekton-pipelines-controller ClusterIP 10.96.176.235 <none> 9090/TCP,8008/TCP,8080/TCP 447d service/tekton-pipelines-webhook ClusterIP 10.101.163.215 <none> 9090/TCP,8008/TCP,443/TCP,8080/TCP 447d NAME READY UP-TO-DATE AVAILABLE AGE deployment.apps/tekton-pipelines-controller 1/1 1 1 87d deployment.apps/tekton-pipelines-webhook 1/1 1 1 87d NAME DESIRED CURRENT READY AGE replicaset.apps/tekton-pipelines-controller-5c78ddd49b 1 1 1 87d replicaset.apps/tekton-pipelines-webhook-5d899cc8c 1 1 1 87d NAME REFERENCE TARGETS MINPODS MAXPODS REPLICAS AGE horizontalpodautoscaler.autoscaling/tekton-pipelines-webhook Deployment/tekton-pipelines-webhook 4%/100% 1 5 1 447d | lang = shell }} === Updating the builder image === You can find the latest info [https://gitlab.wikimedia.org/repos/cloud/toolforge/builds-builder#notes-on-updating-the-builderrunner-images on the builds-builder git repo]. === Adding a new buildpack === We keep a fork for the buildpacks we inject under [[gitlab:groups/repos/cloud/toolforge/buildpacks]] Currently we are using the heroku builder and the heroku buildpacks, that usually are in the old heroku structure (as opposed to cloud-native). So we have to add a shim layer to make them cloud-native compatible. You can see the latest examples in the repository, but a good start is to pull the buildpack from the cnb heroku url <code>https://buildpack-registry.heroku.com/cnb/emk/rust</code> where the last two parts of the path are the author and the name of the buildpack. That will include a few scripts and files that are valid for cloud-native buildpack API 0.4, but we have to adapt them to API 0.6 at least as that's the minimum supported by the builder as of writing this. For that, we have to: * Add a <code>project.toml</code> file * Change the <code>api</code> entry in the <code>buildpack.toml</code> file * Change everywhere where a layer <code>toml</code> file is created to have the newer structure (see the existing examples). You can create a new branch with the changes. Once the fork is ready, we can add the buildpack to the list of buildpacks to inject in the <code>builds-builder</code> (see examples there). Note that this might change soon for a nicer/easier flow, as we are just starting to discover how to manage these things. == History == See [[Help:Toolforge/Build Service#History]]. 5ymvdhaxykrvyg5sq3d6n1edjjvxy5i Tool:Phab-ban/Log 116 453426 2450669 2445265 2026-08-23T00:31:46Z Phabbanbot 37210 260821t0937 was disabled by JJMC89 2450669 wikitext text/x-wiki <noinclude>'''Audit log of bans''' made via https://phab-ban.toolforge.org. Some bans made prior to 2023-09-01 were manually logged at [[phab:T200856]]. __NOTOC____NOINDEX__</noinclude> === 2026-08-23 === * 00:31 [[phab:p/260821t0937|260821t0937]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2026-08-10 === * 01:48 [[phab:p/Adminbdso|Adminbdso]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2026-08-03 === * 08:02 [[phab:p/Jonas2356|Jonas2356]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2026-07-13 === * 13:01 [[phab:p/Camrensssss|Camrensssss]] was disabled by [[phab:p/Mainframe98/|Mainframe98]] * 09:11 [[phab:p/VersatileDove231|VersatileDove231]] was disabled by [[phab:p/Lucas_Werkmeister_WMDE/|Lucas_Werkmeister_WMDE]] === 2026-07-09 === * 00:26 [[phab:p/VVMFOfffce|VVMFOfffce]] was disabled by [[phab:p/SomeRandomDeveloper/|SomeRandomDeveloper]] === 2026-06-30 === * 17:21 [[phab:p/Tomasz_Bladyniec|Tomasz_Bladyniec]] was disabled by [[phab:p/WMFOffice/|WMFOffice]] === 2026-06-29 === * 06:34 [[phab:p/LuniZunie|LuniZunie]] was disabled by [[phab:p/RhinosF1/|RhinosF1]] === 2026-05-25 === * 14:40 [[phab:p/Lysdexia|Lysdexia]] was disabled by [[phab:p/HakanIST/|HakanIST]] === 2026-05-24 === * 17:23 [[phab:p/Nawaf2296|Nawaf2296]] was disabled by [[phab:p/Novem_Linguae/|Novem_Linguae]] === 2026-05-17 === * 01:27 [[phab:p/Chicken.Tender.331|Chicken.Tender.331]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2026-05-07 === * 05:44 [[phab:p/Gabor_Kiss_WMSE|Gabor_Kiss_WMSE]] was disabled by [[phab:p/Sebastian_Berlin-WMSE/|Sebastian_Berlin-WMSE]] === 2026-04-29 === * 22:31 [[phab:p/jhsoby-WMNO|jhsoby-WMNO]] was disabled by [[phab:p/Zabe/|Zabe]] * 20:37 [[phab:p/Hr574380|Hr574380]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2026-04-22 === * 18:16 [[phab:p/Datronmcka|Datronmcka]] was disabled by [[phab:p/Johannnes89/|Johannnes89]] === 2026-04-14 === * 18:53 [[phab:p/Apokrif|Apokrif]] was disabled by [[phab:p/WMFOffice/|WMFOffice]] === 2026-04-07 === * 11:18 [[phab:p/Biof|Biof]] was disabled by [[phab:p/WMFOffice/|WMFOffice]] === 2026-03-15 === * 11:47 [[phab:p/Bucheon606|Bucheon606]] was disabled by [[phab:p/A_smart_kitten/|A_smart_kitten]] * 11:47 [[phab:p/Bucheon606|Bucheon606]] was disabled by [[phab:p/Johannnes89/|Johannnes89]] * 11:29 [[phab:p/Seoulbucheonincheon|Seoulbucheonincheon]] was disabled by [[phab:p/Johannnes89/|Johannnes89]] * 11:10 [[phab:p/Bucheonstation|Bucheonstation]] was disabled by [[phab:p/Johannnes89/|Johannnes89]] * 11:04 [[phab:p/WMFOfffce|WMFOfffce]] was disabled by [[phab:p/A_smart_kitten/|A_smart_kitten]] * 11:03 [[phab:p/WMFOfffce|WMFOfffce]] was disabled by [[phab:p/Peachey88/|Peachey88]] * 11:03 [[phab:p/WMFOfffce|WMFOfffce]] was disabled by [[phab:p/Johannnes89/|Johannnes89]] * 10:57 [[phab:p/WMFOfflce|WMFOfflce]] was disabled by [[phab:p/Johannnes89/|Johannnes89]] * 10:29 [[phab:p/BucheonFac|BucheonFac]] was disabled by [[phab:p/A_smart_kitten/|A_smart_kitten]] * 09:48 [[phab:p/BucheonWest|BucheonWest]] was disabled by [[phab:p/Novem_Linguae/|Novem_Linguae]] * 09:43 [[phab:p/TheBucheon|TheBucheon]] was disabled by [[phab:p/Novem_Linguae/|Novem_Linguae]] * 09:40 [[phab:p/BucheonIncheon|BucheonIncheon]] was disabled by [[phab:p/A_smart_kitten/|A_smart_kitten]] * 09:33 [[phab:p/SkottishFinnishRadist|SkottishFinnishRadist]] was disabled by [[phab:p/Novem_Linguae/|Novem_Linguae]] * 09:33 [[phab:p/SkottishFinnishRadist|SkottishFinnishRadist]] was disabled by [[phab:p/A_smart_kitten/|A_smart_kitten]] * 09:25 [[phab:p/BucheonCityHall6|BucheonCityHall6]] was disabled by [[phab:p/Novem_Linguae/|Novem_Linguae]] * 09:16 [[phab:p/Primefac1|Primefac1]] was disabled by [[phab:p/Novem_Linguae/|Novem_Linguae]] * 09:07 [[phab:p/ScottishFimishRadish|ScottishFimishRadish]] was disabled by [[phab:p/A_smart_kitten/|A_smart_kitten]] * 06:35 [[phab:p/Kgarcia181|Kgarcia181]] was disabled by [[phab:p/Peachey88/|Peachey88]] * 04:27 [[phab:p/SldrF|SldrF]] was disabled by [[phab:p/Novem_Linguae/|Novem_Linguae]] * 04:17 [[phab:p/PrimePac|PrimePac]] was disabled by [[phab:p/A_smart_kitten/|A_smart_kitten]] * 04:01 [[phab:p/260315t1244|260315t1244]] was disabled by [[phab:p/Novem_Linguae/|Novem_Linguae]] * 03:34 [[phab:p/Bucheon543|Bucheon543]] was disabled by [[phab:p/DLynch/|DLynch]] * 03:25 [[phab:p/LAG|LAG]] was disabled by [[phab:p/Novem_Linguae/|Novem_Linguae]] * 03:13 [[phab:p/Wonmidong|Wonmidong]] was disabled by [[phab:p/Novem_Linguae/|Novem_Linguae]] === 2026-03-14 === * 08:10 [[phab:p/Bucheon2026|Bucheon2026]] was disabled by [[phab:p/Johannnes89/|Johannnes89]] * 07:52 [[phab:p/BucheonCityHall|BucheonCityHall]] was disabled by [[phab:p/Johannnes89/|Johannnes89]] * 07:14 [[phab:p/BucheonFesta|BucheonFesta]] was disabled by [[phab:p/Johannnes89/|Johannnes89]] === 2026-03-08 === * 14:34 [[phab:p/Unicord|Unicord]] was disabled by [[phab:p/A_smart_kitten/|A_smart_kitten]] === 2026-02-22 === * 10:39 [[phab:p/Kredionecsresmi|Kredionecsresmi]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2026-02-15 === * 03:34 [[phab:p/alfredbeck|alfredbeck]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2026-02-11 === * 14:34 [[phab:p/Nguyentrongphu|Nguyentrongphu]] was disabled by [[phab:p/Superpes15/|Superpes15]] * 14:32 [[phab:p/Slowking4|Slowking4]] was disabled by [[phab:p/Superpes15/|Superpes15]] === 2026-02-09 === * 22:45 [[phab:p/UNIX-QUANTUM-UNIBANK-FICSIT-NETWORKS|UNIX-QUANTUM-UNIBANK-FICSIT-NETWORKS]] was disabled by [[phab:p/Zabe/|Zabe]] === 2026-01-23 === * 12:09 [[phab:p/Aboodhassanio|Aboodhassanio]] was disabled by [[phab:p/Johannnes89/|Johannnes89]] === 2026-01-16 === * 16:08 [[phab:p/Kimlien316|Kimlien316]] was disabled by [[phab:p/RhinosF1/|RhinosF1]] * 12:01 [[phab:p/Batiste67400|Batiste67400]] was disabled by [[phab:p/Johannnes89/|Johannnes89]] === 2026-01-12 === * 22:04 [[phab:p/Claudioluna6|Claudioluna6]] was disabled by [[phab:p/RhinosF1/|RhinosF1]] === 2026-01-03 === * 14:31 [[phab:p/Cw95hh9|Cw95hh9]] was disabled by [[phab:p/Johannnes89/|Johannnes89]] === 2025-12-27 === * 06:28 [[phab:p/LBLaiSiNanHai|LBLaiSiNanHai]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2025-12-24 === * 21:36 [[phab:p/ItsLido|ItsLido]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2025-12-11 === * 15:50 [[phab:p/DanielJohnso|DanielJohnso]] was disabled by [[phab:p/Marostegui/|Marostegui]] === 2025-12-01 === * 20:36 [[phab:p/Lkcl|Lkcl]] was disabled by [[phab:p/WMFOffice/|WMFOffice]] === 2025-11-30 === * 03:56 [[phab:p/BscottAPL33|BscottAPL33]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2025-11-28 === * 21:21 [[phab:p/3894djfj.10439djf|3894djfj.10439djf]] was disabled by [[phab:p/JJMC89/|JJMC89]] * 09:35 [[phab:p/ElinagittHuB|ElinagittHuB]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2025-11-22 === * 22:09 [[phab:p/Aadvertising|Aadvertising]] was disabled by [[phab:p/Peachey88/|Peachey88]] * 22:09 [[phab:p/TweakFind|TweakFind]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2025-11-20 === * 04:46 [[phab:p/SydneyRug|SydneyRug]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2025-11-14 === * 08:52 [[phab:p/Lougrammfoundation|Lougrammfoundation]] was disabled by [[phab:p/Marostegui/|Marostegui]] === 2025-11-12 === * 06:59 [[phab:p/Cesar|Cesar]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2025-11-05 === * 23:03 [[phab:p/Nehtechnine|Nehtechnine]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2025-11-02 === * 06:33 [[phab:p/Chhoundevid|Chhoundevid]] was disabled by [[phab:p/Johannnes89/|Johannnes89]] === 2025-11-01 === * 07:11 [[phab:p/TowfiqSir|TowfiqSir]] was disabled by [[phab:p/Johannnes89/|Johannnes89]] === 2025-10-14 === * 12:30 [[phab:p/Amrok84|Amrok84]] was disabled by [[phab:p/Lucas_Werkmeister_WMDE/|Lucas_Werkmeister_WMDE]] === 2025-09-24 === * 13:31 [[phab:p/100592|100592]] was disabled by [[phab:p/A_smart_kitten/|A_smart_kitten]] === 2025-09-22 === * 16:20 [[phab:p/DanishAhmedKm|DanishAhmedKm]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2025-09-12 === * 14:08 [[phab:p/GrimGwTK|GrimGwTK]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2025-09-04 === * 15:14 [[phab:p/Amarvip|Amarvip]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2025-08-29 === * 00:09 [[phab:p/IbJo|IbJo]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2025-08-27 === * 02:31 [[phab:p/W11228|W11228]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2025-08-25 === * 17:34 [[phab:p/Faster_than_Thunder|Faster_than_Thunder]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2025-08-22 === * 22:07 [[phab:p/T200856-01|T200856-01]] was disabled by [[phab:p/bd808/|bd808]] === 2025-08-19 === * 04:58 [[phab:p/tiffatk|tiffatk]] was disabled by [[phab:p/Johannnes89/|Johannnes89]] === 2025-08-15 === * 05:00 [[phab:p/Totrue89|Totrue89]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2025-08-06 === * 22:20 [[phab:p/Sajidali110|Sajidali110]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2025-07-31 === * 10:17 [[phab:p/EMIZENTECH|EMIZENTECH]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2025-07-30 === * 10:21 [[phab:p/Fariha_Asghar785|Fariha_Asghar785]] was disabled by [[phab:p/Lucas_Werkmeister_WMDE/|Lucas_Werkmeister_WMDE]] === 2025-07-08 === * 10:46 [[phab:p/Tulsi_Bhagat|Tulsi_Bhagat]] was disabled by [[phab:p/WMFOffice/|WMFOffice]] === 2025-06-29 === * 16:12 [[phab:p/sassybritches|sassybritches]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2025-06-22 === * 07:54 [[phab:p/Godspowertechnical|Godspowertechnical]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2025-06-08 === * 14:59 [[phab:p/Jaypam001|Jaypam001]] was disabled by [[phab:p/JJMC89/|JJMC89]] * 14:53 [[phab:p/DANISHAHMED111|DANISHAHMED111]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2025-06-07 === * 05:50 [[phab:p/PCJND|PCJND]] was disabled by [[phab:p/Johannnes89/|Johannnes89]] === 2025-06-04 === * 08:37 [[phab:p/Alpasli|Alpasli]] was disabled by [[phab:p/WMFOffice/|WMFOffice]] === 2025-06-03 === * 02:07 [[phab:p/Jj881|Jj881]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2025-05-29 === * 05:51 [[phab:p/RodneyAraujo|RodneyAraujo]] was disabled by [[phab:p/WMFOffice/|WMFOffice]] === 2025-04-28 === * 15:07 [[phab:p/Hansmuller|Hansmuller]] was disabled by [[phab:p/WMFOffice/|WMFOffice]] === 2025-04-03 === * 15:57 [[phab:p/Wfan|Wfan]] was disabled by [[phab:p/Zabe/|Zabe]] === 2025-03-30 === * 10:15 [[phab:p/Watnoii24|Watnoii24]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2025-03-23 === * 11:27 [[phab:p/Saadtbli|Saadtbli]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2025-03-22 === * 16:45 [[phab:p/Stephonjeffries19|Stephonjeffries19]] was disabled by [[phab:p/LucasWerkmeister/|LucasWerkmeister]] * 04:32 [[phab:p/Chriswarriortv|Chriswarriortv]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2025-03-19 === * 11:29 [[phab:p/Vinay080|Vinay080]] was disabled by [[phab:p/zeljkofilipin/|zeljkofilipin]] === 2025-03-18 === * 12:04 [[phab:p/Walshandpartners777|Walshandpartners777]] was disabled by [[phab:p/Lucas_Werkmeister_WMDE/|Lucas_Werkmeister_WMDE]] === 2025-03-04 === * 01:33 [[phab:p/Porokhov|Porokhov]] was disabled by [[phab:p/WMFOffice/|WMFOffice]] === 2025-02-25 === * 17:27 [[phab:p/Selahaddin751|Selahaddin751]] was disabled by [[phab:p/brennen/|brennen]] === 2025-02-19 === * 01:00 [[phab:p/Mrb_Rafi|Mrb_Rafi]] was disabled by [[phab:p/WMFOffice/|WMFOffice]] === 2025-02-14 === * 19:19 [[phab:p/3652candy|3652candy]] was disabled by [[phab:p/Peachey88/|Peachey88]] * 17:01 [[phab:p/Ataysaa|Ataysaa]] was disabled by [[phab:p/bd808/|bd808]] === 2025-02-09 === * 09:10 [[phab:p/BTullis|BTullis]] was disabled by [[phab:p/RhinosF1/|RhinosF1]] === 2025-02-08 === * 23:26 [[phab:p/Alexdivkovic05|Alexdivkovic05]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2025-02-06 === * 06:19 [[phab:p/HormigasAIS|HormigasAIS]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2025-01-29 === * 07:40 [[phab:p/Denker61|Denker61]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2025-01-25 === * 21:36 [[phab:p/Khnthichith|Khnthichith]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2025-01-24 === * 11:33 [[phab:p/Aek191010|Aek191010]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2025-01-05 === * 15:40 [[phab:p/szsuperzuper|szsuperzuper]] was disabled by [[phab:p/RhinosF1/|RhinosF1]] === 2025-01-01 === * 09:08 [[phab:p/GALAXYENTERPRISES|GALAXYENTERPRISES]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2024-12-20 === * 00:52 [[phab:p/Mail.faluzes|Mail.faluzes]] was disabled by [[phab:p/Reedy/|Reedy]] === 2024-12-13 === * 02:01 [[phab:p/Gussdafii|Gussdafii]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2024-12-11 === * 08:54 [[phab:p/CodeTrailblazer|CodeTrailblazer]] was disabled by [[phab:p/Peachey88/|Peachey88]] * 08:54 [[phab:p/SelvikIN|SelvikIN]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2024-12-03 === * 05:16 [[phab:p/Matkospajdr|Matkospajdr]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2024-12-01 === * 19:27 [[phab:p/Adarshsingh|Adarshsingh]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2024-11-28 === * 22:47 [[phab:p/Sandraklemma|Sandraklemma]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2024-11-23 === * 09:00 [[phab:p/Mahimabajpayee12|Mahimabajpayee12]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2024-11-11 === * 11:09 [[phab:p/Mvwservices|Mvwservices]] was disabled by [[phab:p/Peachey88/|Peachey88]] * 07:00 [[phab:p/Impactolog|Impactolog]] was disabled by [[phab:p/revi/|revi]] === 2024-10-30 === * 09:05 [[phab:p/Jweighed1|Jweighed1]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2024-10-25 === * 04:20 [[phab:p/Blunt2531|Blunt2531]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2024-10-08 === * 08:31 [[phab:p/Surfcityrecovery|Surfcityrecovery]] was disabled by [[phab:p/MoritzMuehlenhoff/|MoritzMuehlenhoff]] === 2024-10-01 === * 21:49 [[phab:p/T200856-01|T200856-01]] was disabled by [[phab:p/bd808/|bd808]] === 2024-09-27 === * 10:20 [[phab:p/SorBP|SorBP]] was disabled by [[phab:p/TheresNoTime/|TheresNoTime]] === 2024-09-08 === * 10:45 [[phab:p/Robin_Mathew_Rajan|Robin_Mathew_Rajan]] was disabled by [[phab:p/RhinosF1/|RhinosF1]] === 2024-09-02 === * 17:58 [[phab:p/Idxntcx|Idxntcx]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2024-09-01 === * 10:11 [[phab:p/LDAP|LDAP]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2024-07-22 === * 08:56 [[phab:p/Nobleadele|Nobleadele]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2024-06-18 === * 20:54 [[phab:p/Playgiirlkaybrazy|Playgiirlkaybrazy]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2024-06-08 === * 22:30 [[phab:p/Exposingsesion1|Exposingsesion1]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2024-05-27 === * 12:24 [[phab:p/JosefineHellrothLarssonWMSE|JosefineHellrothLarssonWMSE]] was disabled by [[phab:p/Sebastian_Berlin-WMSE/|Sebastian_Berlin-WMSE]] * 07:54 [[phab:p/SMMpanels|SMMpanels]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2024-05-22 === * 11:31 [[phab:p/Sandra_Fauconnier_WMSE|Sandra_Fauconnier_WMSE]] was disabled by [[phab:p/Sebastian_Berlin-WMSE/|Sebastian_Berlin-WMSE]] * 10:23 [[phab:p/MiaJacobssonWMSE|MiaJacobssonWMSE]] was disabled by [[phab:p/Sebastian_Berlin-WMSE/|Sebastian_Berlin-WMSE]] * 10:20 [[phab:p/David_Haskiya_WMSE|David_Haskiya_WMSE]] was disabled by [[phab:p/Sebastian_Berlin-WMSE/|Sebastian_Berlin-WMSE]] * 10:20 [[phab:p/kalle|kalle]] was disabled by [[phab:p/Sebastian_Berlin-WMSE/|Sebastian_Berlin-WMSE]] * 10:20 [[phab:p/Tore_Danielsson_WMSE|Tore_Danielsson_WMSE]] was disabled by [[phab:p/Sebastian_Berlin-WMSE/|Sebastian_Berlin-WMSE]] * 10:19 [[phab:p/Gitta|Gitta]] was disabled by [[phab:p/Sebastian_Berlin-WMSE/|Sebastian_Berlin-WMSE]] * 10:19 [[phab:p/annatroberg|annatroberg]] was disabled by [[phab:p/Sebastian_Berlin-WMSE/|Sebastian_Berlin-WMSE]] * 10:19 [[phab:p/AxelPettersson_WMSE|AxelPettersson_WMSE]] was disabled by [[phab:p/Sebastian_Berlin-WMSE/|Sebastian_Berlin-WMSE]] * 10:17 [[phab:p/SaraMortsell|SaraMortsell]] was disabled by [[phab:p/Sebastian_Berlin-WMSE/|Sebastian_Berlin-WMSE]] === 2024-05-13 === * 14:55 [[phab:p/BenoitPrieur|BenoitPrieur]] was disabled by [[phab:p/WMFOffice/|WMFOffice]] === 2024-05-04 === * 19:54 [[phab:p/Sammoon391|Sammoon391]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2024-05-01 === * 07:24 [[phab:p/Soubag|Soubag]] was disabled by [[phab:p/Mainframe98/|Mainframe98]] === 2024-04-28 === * 06:22 [[phab:p/Diamondscoin|Diamondscoin]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2024-04-19 === * 03:47 [[phab:p/Wawmart2|Wawmart2]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2024-04-08 === * 10:21 [[phab:p/Mardetanha|Mardetanha]] was disabled by [[phab:p/WMFOffice/|WMFOffice]] === 2024-03-29 === * 09:03 [[phab:p/Abdollmjjedloveanan|Abdollmjjedloveanan]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2024-03-12 === * 09:45 [[phab:p/Samantha78462|Samantha78462]] was disabled by [[phab:p/Peachey88/|Peachey88]] * 09:39 [[phab:p/Samantha7861654654|Samantha7861654654]] was disabled by [[phab:p/Peachey88/|Peachey88]] * 09:11 [[phab:p/Robin|Robin]] was disabled by [[phab:p/Peachey88/|Peachey88]] * 09:05 [[phab:p/Anglinakuki|Anglinakuki]] was disabled by [[phab:p/Peachey88/|Peachey88]] * 00:02 [[phab:p/Johnne25|Johnne25]] was disabled by [[phab:p/bd808/|bd808]] === 2024-03-07 === * 20:07 [[phab:p/Sami785|Sami785]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2024-03-06 === * 07:41 [[phab:p/28q|28q]] was disabled by [[phab:p/RhinosF1/|RhinosF1]] === 2024-03-02 === * 22:51 [[phab:p/kitchenstrategic|kitchenstrategic]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2024-02-23 === * 09:03 [[phab:p/littleggghost|littleggghost]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2024-02-17 === * 05:10 [[phab:p/Skekeiei|Skekeiei]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2024-01-27 === * 03:50 [[phab:p/Andybitcoin|Andybitcoin]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2024-01-25 === * 09:10 [[phab:p/Mayo3030|Mayo3030]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2024-01-21 === * 05:59 [[phab:p/Hackear|Hackear]] was disabled by [[phab:p/Peachey88/|Peachey88]] * 05:58 [[phab:p/joselopez45|joselopez45]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2024-01-20 === * 20:02 [[phab:p/08107130655|08107130655]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2024-01-18 === * 05:33 [[phab:p/Tecnologynew|Tecnologynew]] was disabled by [[phab:p/TheresNoTime/|TheresNoTime]] === 2024-01-12 === * 22:34 [[phab:p/cchen|cchen]] was disabled by [[phab:p/RhinosF1/|RhinosF1]] === 2024-01-11 === * 07:34 [[phab:p/Bernita43|Bernita43]] was disabled by [[phab:p/RhinosF1/|RhinosF1]] === 2024-01-07 === * 10:13 [[phab:p/Irademack|Irademack]] was disabled by [[phab:p/RhinosF1/|RhinosF1]] === 2023-12-31 === * 16:48 [[phab:p/Vieclamdmpt|Vieclamdmpt]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2023-12-26 === * 20:34 [[phab:p/Bgu5678|Bgu5678]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2023-11-26 === * 20:09 [[phab:p/Str13tlife|Str13tlife]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2023-11-24 === * 00:49 [[phab:p/Imambuchori03|Imambuchori03]] was disabled by [[phab:p/DannyS712/|DannyS712]] === 2023-11-22 === * 15:50 [[phab:p/Naleksuh|Naleksuh]] was disabled by [[phab:p/WMFOffice/|WMFOffice]] === 2023-11-19 === * 21:10 [[phab:p/Onack16888|Onack16888]] was disabled by [[phab:p/Daimona/|Daimona]] === 2023-11-15 === * 11:57 [[phab:p/Anonymous_ehacker|Anonymous_ehacker]] was disabled by [[phab:p/hashar/|hashar]] === 2023-11-09 === * 23:16 [[phab:p/dunicorn|dunicorn]] was disabled by [[phab:p/bd808/|bd808]] === 2023-09-30 === * 17:38 [[phab:p/Wykirany|Wykirany]] was disabled by [[phab:p/RhinosF1/|RhinosF1]] === 2023-09-01 === * 16:29 [[phab:p/T200856-01|T200856-01]] was disabled by [[phab:p/bd808/|bd808]] * 16:17 [[phab:p/T200856-01|T200856-01]] was disabled by [[phab:p/bd808/|bd808]] dowv8nhv6gp97kop08wm4wi914ybrda 2450670 2450669 2026-08-23T00:31:54Z Phabbanbot 37210 Carrotlandz was disabled by JJMC89 2450670 wikitext text/x-wiki <noinclude>'''Audit log of bans''' made via https://phab-ban.toolforge.org. Some bans made prior to 2023-09-01 were manually logged at [[phab:T200856]]. __NOTOC____NOINDEX__</noinclude> === 2026-08-23 === * 00:31 [[phab:p/Carrotlandz|Carrotlandz]] was disabled by [[phab:p/JJMC89/|JJMC89]] * 00:31 [[phab:p/260821t0937|260821t0937]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2026-08-10 === * 01:48 [[phab:p/Adminbdso|Adminbdso]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2026-08-03 === * 08:02 [[phab:p/Jonas2356|Jonas2356]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2026-07-13 === * 13:01 [[phab:p/Camrensssss|Camrensssss]] was disabled by [[phab:p/Mainframe98/|Mainframe98]] * 09:11 [[phab:p/VersatileDove231|VersatileDove231]] was disabled by [[phab:p/Lucas_Werkmeister_WMDE/|Lucas_Werkmeister_WMDE]] === 2026-07-09 === * 00:26 [[phab:p/VVMFOfffce|VVMFOfffce]] was disabled by [[phab:p/SomeRandomDeveloper/|SomeRandomDeveloper]] === 2026-06-30 === * 17:21 [[phab:p/Tomasz_Bladyniec|Tomasz_Bladyniec]] was disabled by [[phab:p/WMFOffice/|WMFOffice]] === 2026-06-29 === * 06:34 [[phab:p/LuniZunie|LuniZunie]] was disabled by [[phab:p/RhinosF1/|RhinosF1]] === 2026-05-25 === * 14:40 [[phab:p/Lysdexia|Lysdexia]] was disabled by [[phab:p/HakanIST/|HakanIST]] === 2026-05-24 === * 17:23 [[phab:p/Nawaf2296|Nawaf2296]] was disabled by [[phab:p/Novem_Linguae/|Novem_Linguae]] === 2026-05-17 === * 01:27 [[phab:p/Chicken.Tender.331|Chicken.Tender.331]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2026-05-07 === * 05:44 [[phab:p/Gabor_Kiss_WMSE|Gabor_Kiss_WMSE]] was disabled by [[phab:p/Sebastian_Berlin-WMSE/|Sebastian_Berlin-WMSE]] === 2026-04-29 === * 22:31 [[phab:p/jhsoby-WMNO|jhsoby-WMNO]] was disabled by [[phab:p/Zabe/|Zabe]] * 20:37 [[phab:p/Hr574380|Hr574380]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2026-04-22 === * 18:16 [[phab:p/Datronmcka|Datronmcka]] was disabled by [[phab:p/Johannnes89/|Johannnes89]] === 2026-04-14 === * 18:53 [[phab:p/Apokrif|Apokrif]] was disabled by [[phab:p/WMFOffice/|WMFOffice]] === 2026-04-07 === * 11:18 [[phab:p/Biof|Biof]] was disabled by [[phab:p/WMFOffice/|WMFOffice]] === 2026-03-15 === * 11:47 [[phab:p/Bucheon606|Bucheon606]] was disabled by [[phab:p/A_smart_kitten/|A_smart_kitten]] * 11:47 [[phab:p/Bucheon606|Bucheon606]] was disabled by [[phab:p/Johannnes89/|Johannnes89]] * 11:29 [[phab:p/Seoulbucheonincheon|Seoulbucheonincheon]] was disabled by [[phab:p/Johannnes89/|Johannnes89]] * 11:10 [[phab:p/Bucheonstation|Bucheonstation]] was disabled by [[phab:p/Johannnes89/|Johannnes89]] * 11:04 [[phab:p/WMFOfffce|WMFOfffce]] was disabled by [[phab:p/A_smart_kitten/|A_smart_kitten]] * 11:03 [[phab:p/WMFOfffce|WMFOfffce]] was disabled by [[phab:p/Peachey88/|Peachey88]] * 11:03 [[phab:p/WMFOfffce|WMFOfffce]] was disabled by [[phab:p/Johannnes89/|Johannnes89]] * 10:57 [[phab:p/WMFOfflce|WMFOfflce]] was disabled by [[phab:p/Johannnes89/|Johannnes89]] * 10:29 [[phab:p/BucheonFac|BucheonFac]] was disabled by [[phab:p/A_smart_kitten/|A_smart_kitten]] * 09:48 [[phab:p/BucheonWest|BucheonWest]] was disabled by [[phab:p/Novem_Linguae/|Novem_Linguae]] * 09:43 [[phab:p/TheBucheon|TheBucheon]] was disabled by [[phab:p/Novem_Linguae/|Novem_Linguae]] * 09:40 [[phab:p/BucheonIncheon|BucheonIncheon]] was disabled by [[phab:p/A_smart_kitten/|A_smart_kitten]] * 09:33 [[phab:p/SkottishFinnishRadist|SkottishFinnishRadist]] was disabled by [[phab:p/Novem_Linguae/|Novem_Linguae]] * 09:33 [[phab:p/SkottishFinnishRadist|SkottishFinnishRadist]] was disabled by [[phab:p/A_smart_kitten/|A_smart_kitten]] * 09:25 [[phab:p/BucheonCityHall6|BucheonCityHall6]] was disabled by [[phab:p/Novem_Linguae/|Novem_Linguae]] * 09:16 [[phab:p/Primefac1|Primefac1]] was disabled by [[phab:p/Novem_Linguae/|Novem_Linguae]] * 09:07 [[phab:p/ScottishFimishRadish|ScottishFimishRadish]] was disabled by [[phab:p/A_smart_kitten/|A_smart_kitten]] * 06:35 [[phab:p/Kgarcia181|Kgarcia181]] was disabled by [[phab:p/Peachey88/|Peachey88]] * 04:27 [[phab:p/SldrF|SldrF]] was disabled by [[phab:p/Novem_Linguae/|Novem_Linguae]] * 04:17 [[phab:p/PrimePac|PrimePac]] was disabled by [[phab:p/A_smart_kitten/|A_smart_kitten]] * 04:01 [[phab:p/260315t1244|260315t1244]] was disabled by [[phab:p/Novem_Linguae/|Novem_Linguae]] * 03:34 [[phab:p/Bucheon543|Bucheon543]] was disabled by [[phab:p/DLynch/|DLynch]] * 03:25 [[phab:p/LAG|LAG]] was disabled by [[phab:p/Novem_Linguae/|Novem_Linguae]] * 03:13 [[phab:p/Wonmidong|Wonmidong]] was disabled by [[phab:p/Novem_Linguae/|Novem_Linguae]] === 2026-03-14 === * 08:10 [[phab:p/Bucheon2026|Bucheon2026]] was disabled by [[phab:p/Johannnes89/|Johannnes89]] * 07:52 [[phab:p/BucheonCityHall|BucheonCityHall]] was disabled by [[phab:p/Johannnes89/|Johannnes89]] * 07:14 [[phab:p/BucheonFesta|BucheonFesta]] was disabled by [[phab:p/Johannnes89/|Johannnes89]] === 2026-03-08 === * 14:34 [[phab:p/Unicord|Unicord]] was disabled by [[phab:p/A_smart_kitten/|A_smart_kitten]] === 2026-02-22 === * 10:39 [[phab:p/Kredionecsresmi|Kredionecsresmi]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2026-02-15 === * 03:34 [[phab:p/alfredbeck|alfredbeck]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2026-02-11 === * 14:34 [[phab:p/Nguyentrongphu|Nguyentrongphu]] was disabled by [[phab:p/Superpes15/|Superpes15]] * 14:32 [[phab:p/Slowking4|Slowking4]] was disabled by [[phab:p/Superpes15/|Superpes15]] === 2026-02-09 === * 22:45 [[phab:p/UNIX-QUANTUM-UNIBANK-FICSIT-NETWORKS|UNIX-QUANTUM-UNIBANK-FICSIT-NETWORKS]] was disabled by [[phab:p/Zabe/|Zabe]] === 2026-01-23 === * 12:09 [[phab:p/Aboodhassanio|Aboodhassanio]] was disabled by [[phab:p/Johannnes89/|Johannnes89]] === 2026-01-16 === * 16:08 [[phab:p/Kimlien316|Kimlien316]] was disabled by [[phab:p/RhinosF1/|RhinosF1]] * 12:01 [[phab:p/Batiste67400|Batiste67400]] was disabled by [[phab:p/Johannnes89/|Johannnes89]] === 2026-01-12 === * 22:04 [[phab:p/Claudioluna6|Claudioluna6]] was disabled by [[phab:p/RhinosF1/|RhinosF1]] === 2026-01-03 === * 14:31 [[phab:p/Cw95hh9|Cw95hh9]] was disabled by [[phab:p/Johannnes89/|Johannnes89]] === 2025-12-27 === * 06:28 [[phab:p/LBLaiSiNanHai|LBLaiSiNanHai]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2025-12-24 === * 21:36 [[phab:p/ItsLido|ItsLido]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2025-12-11 === * 15:50 [[phab:p/DanielJohnso|DanielJohnso]] was disabled by [[phab:p/Marostegui/|Marostegui]] === 2025-12-01 === * 20:36 [[phab:p/Lkcl|Lkcl]] was disabled by [[phab:p/WMFOffice/|WMFOffice]] === 2025-11-30 === * 03:56 [[phab:p/BscottAPL33|BscottAPL33]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2025-11-28 === * 21:21 [[phab:p/3894djfj.10439djf|3894djfj.10439djf]] was disabled by [[phab:p/JJMC89/|JJMC89]] * 09:35 [[phab:p/ElinagittHuB|ElinagittHuB]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2025-11-22 === * 22:09 [[phab:p/Aadvertising|Aadvertising]] was disabled by [[phab:p/Peachey88/|Peachey88]] * 22:09 [[phab:p/TweakFind|TweakFind]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2025-11-20 === * 04:46 [[phab:p/SydneyRug|SydneyRug]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2025-11-14 === * 08:52 [[phab:p/Lougrammfoundation|Lougrammfoundation]] was disabled by [[phab:p/Marostegui/|Marostegui]] === 2025-11-12 === * 06:59 [[phab:p/Cesar|Cesar]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2025-11-05 === * 23:03 [[phab:p/Nehtechnine|Nehtechnine]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2025-11-02 === * 06:33 [[phab:p/Chhoundevid|Chhoundevid]] was disabled by [[phab:p/Johannnes89/|Johannnes89]] === 2025-11-01 === * 07:11 [[phab:p/TowfiqSir|TowfiqSir]] was disabled by [[phab:p/Johannnes89/|Johannnes89]] === 2025-10-14 === * 12:30 [[phab:p/Amrok84|Amrok84]] was disabled by [[phab:p/Lucas_Werkmeister_WMDE/|Lucas_Werkmeister_WMDE]] === 2025-09-24 === * 13:31 [[phab:p/100592|100592]] was disabled by [[phab:p/A_smart_kitten/|A_smart_kitten]] === 2025-09-22 === * 16:20 [[phab:p/DanishAhmedKm|DanishAhmedKm]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2025-09-12 === * 14:08 [[phab:p/GrimGwTK|GrimGwTK]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2025-09-04 === * 15:14 [[phab:p/Amarvip|Amarvip]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2025-08-29 === * 00:09 [[phab:p/IbJo|IbJo]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2025-08-27 === * 02:31 [[phab:p/W11228|W11228]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2025-08-25 === * 17:34 [[phab:p/Faster_than_Thunder|Faster_than_Thunder]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2025-08-22 === * 22:07 [[phab:p/T200856-01|T200856-01]] was disabled by [[phab:p/bd808/|bd808]] === 2025-08-19 === * 04:58 [[phab:p/tiffatk|tiffatk]] was disabled by [[phab:p/Johannnes89/|Johannnes89]] === 2025-08-15 === * 05:00 [[phab:p/Totrue89|Totrue89]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2025-08-06 === * 22:20 [[phab:p/Sajidali110|Sajidali110]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2025-07-31 === * 10:17 [[phab:p/EMIZENTECH|EMIZENTECH]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2025-07-30 === * 10:21 [[phab:p/Fariha_Asghar785|Fariha_Asghar785]] was disabled by [[phab:p/Lucas_Werkmeister_WMDE/|Lucas_Werkmeister_WMDE]] === 2025-07-08 === * 10:46 [[phab:p/Tulsi_Bhagat|Tulsi_Bhagat]] was disabled by [[phab:p/WMFOffice/|WMFOffice]] === 2025-06-29 === * 16:12 [[phab:p/sassybritches|sassybritches]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2025-06-22 === * 07:54 [[phab:p/Godspowertechnical|Godspowertechnical]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2025-06-08 === * 14:59 [[phab:p/Jaypam001|Jaypam001]] was disabled by [[phab:p/JJMC89/|JJMC89]] * 14:53 [[phab:p/DANISHAHMED111|DANISHAHMED111]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2025-06-07 === * 05:50 [[phab:p/PCJND|PCJND]] was disabled by [[phab:p/Johannnes89/|Johannnes89]] === 2025-06-04 === * 08:37 [[phab:p/Alpasli|Alpasli]] was disabled by [[phab:p/WMFOffice/|WMFOffice]] === 2025-06-03 === * 02:07 [[phab:p/Jj881|Jj881]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2025-05-29 === * 05:51 [[phab:p/RodneyAraujo|RodneyAraujo]] was disabled by [[phab:p/WMFOffice/|WMFOffice]] === 2025-04-28 === * 15:07 [[phab:p/Hansmuller|Hansmuller]] was disabled by [[phab:p/WMFOffice/|WMFOffice]] === 2025-04-03 === * 15:57 [[phab:p/Wfan|Wfan]] was disabled by [[phab:p/Zabe/|Zabe]] === 2025-03-30 === * 10:15 [[phab:p/Watnoii24|Watnoii24]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2025-03-23 === * 11:27 [[phab:p/Saadtbli|Saadtbli]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2025-03-22 === * 16:45 [[phab:p/Stephonjeffries19|Stephonjeffries19]] was disabled by [[phab:p/LucasWerkmeister/|LucasWerkmeister]] * 04:32 [[phab:p/Chriswarriortv|Chriswarriortv]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2025-03-19 === * 11:29 [[phab:p/Vinay080|Vinay080]] was disabled by [[phab:p/zeljkofilipin/|zeljkofilipin]] === 2025-03-18 === * 12:04 [[phab:p/Walshandpartners777|Walshandpartners777]] was disabled by [[phab:p/Lucas_Werkmeister_WMDE/|Lucas_Werkmeister_WMDE]] === 2025-03-04 === * 01:33 [[phab:p/Porokhov|Porokhov]] was disabled by [[phab:p/WMFOffice/|WMFOffice]] === 2025-02-25 === * 17:27 [[phab:p/Selahaddin751|Selahaddin751]] was disabled by [[phab:p/brennen/|brennen]] === 2025-02-19 === * 01:00 [[phab:p/Mrb_Rafi|Mrb_Rafi]] was disabled by [[phab:p/WMFOffice/|WMFOffice]] === 2025-02-14 === * 19:19 [[phab:p/3652candy|3652candy]] was disabled by [[phab:p/Peachey88/|Peachey88]] * 17:01 [[phab:p/Ataysaa|Ataysaa]] was disabled by [[phab:p/bd808/|bd808]] === 2025-02-09 === * 09:10 [[phab:p/BTullis|BTullis]] was disabled by [[phab:p/RhinosF1/|RhinosF1]] === 2025-02-08 === * 23:26 [[phab:p/Alexdivkovic05|Alexdivkovic05]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2025-02-06 === * 06:19 [[phab:p/HormigasAIS|HormigasAIS]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2025-01-29 === * 07:40 [[phab:p/Denker61|Denker61]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2025-01-25 === * 21:36 [[phab:p/Khnthichith|Khnthichith]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2025-01-24 === * 11:33 [[phab:p/Aek191010|Aek191010]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2025-01-05 === * 15:40 [[phab:p/szsuperzuper|szsuperzuper]] was disabled by [[phab:p/RhinosF1/|RhinosF1]] === 2025-01-01 === * 09:08 [[phab:p/GALAXYENTERPRISES|GALAXYENTERPRISES]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2024-12-20 === * 00:52 [[phab:p/Mail.faluzes|Mail.faluzes]] was disabled by [[phab:p/Reedy/|Reedy]] === 2024-12-13 === * 02:01 [[phab:p/Gussdafii|Gussdafii]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2024-12-11 === * 08:54 [[phab:p/CodeTrailblazer|CodeTrailblazer]] was disabled by [[phab:p/Peachey88/|Peachey88]] * 08:54 [[phab:p/SelvikIN|SelvikIN]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2024-12-03 === * 05:16 [[phab:p/Matkospajdr|Matkospajdr]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2024-12-01 === * 19:27 [[phab:p/Adarshsingh|Adarshsingh]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2024-11-28 === * 22:47 [[phab:p/Sandraklemma|Sandraklemma]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2024-11-23 === * 09:00 [[phab:p/Mahimabajpayee12|Mahimabajpayee12]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2024-11-11 === * 11:09 [[phab:p/Mvwservices|Mvwservices]] was disabled by [[phab:p/Peachey88/|Peachey88]] * 07:00 [[phab:p/Impactolog|Impactolog]] was disabled by [[phab:p/revi/|revi]] === 2024-10-30 === * 09:05 [[phab:p/Jweighed1|Jweighed1]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2024-10-25 === * 04:20 [[phab:p/Blunt2531|Blunt2531]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2024-10-08 === * 08:31 [[phab:p/Surfcityrecovery|Surfcityrecovery]] was disabled by [[phab:p/MoritzMuehlenhoff/|MoritzMuehlenhoff]] === 2024-10-01 === * 21:49 [[phab:p/T200856-01|T200856-01]] was disabled by [[phab:p/bd808/|bd808]] === 2024-09-27 === * 10:20 [[phab:p/SorBP|SorBP]] was disabled by [[phab:p/TheresNoTime/|TheresNoTime]] === 2024-09-08 === * 10:45 [[phab:p/Robin_Mathew_Rajan|Robin_Mathew_Rajan]] was disabled by [[phab:p/RhinosF1/|RhinosF1]] === 2024-09-02 === * 17:58 [[phab:p/Idxntcx|Idxntcx]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2024-09-01 === * 10:11 [[phab:p/LDAP|LDAP]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2024-07-22 === * 08:56 [[phab:p/Nobleadele|Nobleadele]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2024-06-18 === * 20:54 [[phab:p/Playgiirlkaybrazy|Playgiirlkaybrazy]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2024-06-08 === * 22:30 [[phab:p/Exposingsesion1|Exposingsesion1]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2024-05-27 === * 12:24 [[phab:p/JosefineHellrothLarssonWMSE|JosefineHellrothLarssonWMSE]] was disabled by [[phab:p/Sebastian_Berlin-WMSE/|Sebastian_Berlin-WMSE]] * 07:54 [[phab:p/SMMpanels|SMMpanels]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2024-05-22 === * 11:31 [[phab:p/Sandra_Fauconnier_WMSE|Sandra_Fauconnier_WMSE]] was disabled by [[phab:p/Sebastian_Berlin-WMSE/|Sebastian_Berlin-WMSE]] * 10:23 [[phab:p/MiaJacobssonWMSE|MiaJacobssonWMSE]] was disabled by [[phab:p/Sebastian_Berlin-WMSE/|Sebastian_Berlin-WMSE]] * 10:20 [[phab:p/David_Haskiya_WMSE|David_Haskiya_WMSE]] was disabled by [[phab:p/Sebastian_Berlin-WMSE/|Sebastian_Berlin-WMSE]] * 10:20 [[phab:p/kalle|kalle]] was disabled by [[phab:p/Sebastian_Berlin-WMSE/|Sebastian_Berlin-WMSE]] * 10:20 [[phab:p/Tore_Danielsson_WMSE|Tore_Danielsson_WMSE]] was disabled by [[phab:p/Sebastian_Berlin-WMSE/|Sebastian_Berlin-WMSE]] * 10:19 [[phab:p/Gitta|Gitta]] was disabled by [[phab:p/Sebastian_Berlin-WMSE/|Sebastian_Berlin-WMSE]] * 10:19 [[phab:p/annatroberg|annatroberg]] was disabled by [[phab:p/Sebastian_Berlin-WMSE/|Sebastian_Berlin-WMSE]] * 10:19 [[phab:p/AxelPettersson_WMSE|AxelPettersson_WMSE]] was disabled by [[phab:p/Sebastian_Berlin-WMSE/|Sebastian_Berlin-WMSE]] * 10:17 [[phab:p/SaraMortsell|SaraMortsell]] was disabled by [[phab:p/Sebastian_Berlin-WMSE/|Sebastian_Berlin-WMSE]] === 2024-05-13 === * 14:55 [[phab:p/BenoitPrieur|BenoitPrieur]] was disabled by [[phab:p/WMFOffice/|WMFOffice]] === 2024-05-04 === * 19:54 [[phab:p/Sammoon391|Sammoon391]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2024-05-01 === * 07:24 [[phab:p/Soubag|Soubag]] was disabled by [[phab:p/Mainframe98/|Mainframe98]] === 2024-04-28 === * 06:22 [[phab:p/Diamondscoin|Diamondscoin]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2024-04-19 === * 03:47 [[phab:p/Wawmart2|Wawmart2]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2024-04-08 === * 10:21 [[phab:p/Mardetanha|Mardetanha]] was disabled by [[phab:p/WMFOffice/|WMFOffice]] === 2024-03-29 === * 09:03 [[phab:p/Abdollmjjedloveanan|Abdollmjjedloveanan]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2024-03-12 === * 09:45 [[phab:p/Samantha78462|Samantha78462]] was disabled by [[phab:p/Peachey88/|Peachey88]] * 09:39 [[phab:p/Samantha7861654654|Samantha7861654654]] was disabled by [[phab:p/Peachey88/|Peachey88]] * 09:11 [[phab:p/Robin|Robin]] was disabled by [[phab:p/Peachey88/|Peachey88]] * 09:05 [[phab:p/Anglinakuki|Anglinakuki]] was disabled by [[phab:p/Peachey88/|Peachey88]] * 00:02 [[phab:p/Johnne25|Johnne25]] was disabled by [[phab:p/bd808/|bd808]] === 2024-03-07 === * 20:07 [[phab:p/Sami785|Sami785]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2024-03-06 === * 07:41 [[phab:p/28q|28q]] was disabled by [[phab:p/RhinosF1/|RhinosF1]] === 2024-03-02 === * 22:51 [[phab:p/kitchenstrategic|kitchenstrategic]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2024-02-23 === * 09:03 [[phab:p/littleggghost|littleggghost]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2024-02-17 === * 05:10 [[phab:p/Skekeiei|Skekeiei]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2024-01-27 === * 03:50 [[phab:p/Andybitcoin|Andybitcoin]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2024-01-25 === * 09:10 [[phab:p/Mayo3030|Mayo3030]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2024-01-21 === * 05:59 [[phab:p/Hackear|Hackear]] was disabled by [[phab:p/Peachey88/|Peachey88]] * 05:58 [[phab:p/joselopez45|joselopez45]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2024-01-20 === * 20:02 [[phab:p/08107130655|08107130655]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2024-01-18 === * 05:33 [[phab:p/Tecnologynew|Tecnologynew]] was disabled by [[phab:p/TheresNoTime/|TheresNoTime]] === 2024-01-12 === * 22:34 [[phab:p/cchen|cchen]] was disabled by [[phab:p/RhinosF1/|RhinosF1]] === 2024-01-11 === * 07:34 [[phab:p/Bernita43|Bernita43]] was disabled by [[phab:p/RhinosF1/|RhinosF1]] === 2024-01-07 === * 10:13 [[phab:p/Irademack|Irademack]] was disabled by [[phab:p/RhinosF1/|RhinosF1]] === 2023-12-31 === * 16:48 [[phab:p/Vieclamdmpt|Vieclamdmpt]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2023-12-26 === * 20:34 [[phab:p/Bgu5678|Bgu5678]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2023-11-26 === * 20:09 [[phab:p/Str13tlife|Str13tlife]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2023-11-24 === * 00:49 [[phab:p/Imambuchori03|Imambuchori03]] was disabled by [[phab:p/DannyS712/|DannyS712]] === 2023-11-22 === * 15:50 [[phab:p/Naleksuh|Naleksuh]] was disabled by [[phab:p/WMFOffice/|WMFOffice]] === 2023-11-19 === * 21:10 [[phab:p/Onack16888|Onack16888]] was disabled by [[phab:p/Daimona/|Daimona]] === 2023-11-15 === * 11:57 [[phab:p/Anonymous_ehacker|Anonymous_ehacker]] was disabled by [[phab:p/hashar/|hashar]] === 2023-11-09 === * 23:16 [[phab:p/dunicorn|dunicorn]] was disabled by [[phab:p/bd808/|bd808]] === 2023-09-30 === * 17:38 [[phab:p/Wykirany|Wykirany]] was disabled by [[phab:p/RhinosF1/|RhinosF1]] === 2023-09-01 === * 16:29 [[phab:p/T200856-01|T200856-01]] was disabled by [[phab:p/bd808/|bd808]] * 16:17 [[phab:p/T200856-01|T200856-01]] was disabled by [[phab:p/bd808/|bd808]] ehn36auw354oqgoraxxh9nuzrj208ff Developer account 0 453877 2450677 2450316 2026-08-23T11:45:42Z Taavi 13997 link to explainer 2450677 wikitext text/x-wiki {{See also|mw:Developer account}} A Wikimedia '''developer account''' is an account used for authentication and authorization in Wikimedia's technical spaces such as source code management systems, bug trackers, compute environments, and operational support tooling. Developer accounts [[User:Taavi/Accounts, explained|are distinct]] from [[meta:Help:Unified login|Wikimedia SUL accounts]] which are used for authentication and authorization in Wikimedia project wikis, but in the modern era are often linked to a SUL account that is operated by the same human being or for related automation purposes. Developer account information is stored in an [[w:Lightweight Directory Access Protocol|LDAP directory]] maintained in the Wikimedia production network. More details on the LDAP service itself can be found in [[SRE/LDAP]]. The LDAP directory is used in [[Portal:Cloud VPS|Cloud VPS]] and [[Portal:Toolforge|Toolforge]] to provide Unix account information and ssh public keys to virtual machines. This is the same LDAP directory that backs the [[IDM]] (Bitu) and [[CAS-SSO]] (IDP) services. Within the LDAP directory, developer accounts are <code>objectClass=posixAccount</code> entities stored in the <code>ou=people,dc=wikimedia,dc=org</code> subtree. Cloud VPS project membership is recorded with groups named <code>cn=project-$CLOUD_VPS_ID,ou=groups,dc=wikimedia,dc=org</code> where <code>$CLOUD_VPS_ID</code> is the OpenStack project id which was the same as the OpenStack project name for years (for example "tools") but is now a UUID. Toolforge tool maintainership is recorded with groups named <code>tools.$TOOLNAME,ou=servicegroups,dc=wikimedia,dc=org</code> where <code>$TOOLNAME</code> is the tool's public name (for example "[[toolforge:versions|versions]]" or "[[toolforge:trainbow|trainbow]]"). == Working with LDAP cli == For simplification of the documentation, cli examples of working with the LDAP directory on this page will generally use the shell aliases that [[User:BryanDavis]] has developed and documented at [[User:BryanDavis/LDAP]]. {{Codesample|lang=shell-session|code= $ alias ldap='ldapsearch -xLLL -P 3 -E pr=5000/noprompt -o ldif-wrap=no -b"dc=wikimedia,dc=org"' $ alias un64='awk '\''BEGIN{FS=":: ";c="base64 -d"}{if(/\w+:: /) {print $2 {{!}}& c; close(c,"to"); c {{!}}& getline $2; close(c); printf("%s:: \"%s\"\n", $1, $2); next} print $0 }'\''' }} == Blocked Developer accounts == Developer accounts can be blocked similar to a full MediaWiki account block. Blocks cascade to disable the account's access rights in Toolforge, Cloud VPS, Gerrit, GitLab, and Phabricator. Blocks are non-destructive and fully reversible. Blocks are typically applied and removed using https://idm.wikimedia.org/wikimedia/block/ which is available to members of the [[toolforge:ldap/group/bitu-account-managers|Bitu account managers]] group. Other tooling may be used in specific workflows such as WMF staff offboarding. Blocked developer accounts can be detected in the LDAP directory by looking for the presence of <code>pwdPolicySubentry=cn=disabled,ou=ppolicies,dc=wikimedia,dc=org</code>. More typically you may want to exclude all blocked accounts from a lookup by using the negated form <code>(!(pwdPolicySubentry=cn=disabled,ou=ppolicies,dc=wikimedia,dc=org))</code>. If you wanted the count of all unblocked Developer accounts: {{Codesample|lang=shell-session|code= $ ldap '(&(objectClass=posixAccount)(!(pwdPolicySubentry=cn=disabled,ou=ppolicies,dc=wikimedia,dc=org)))' dn -b ou=people,dc=wikimedia,dc=org {{!}} grep dn: {{!}} wc -l 36723 # as of 2026-08-20 }} == Developer account details == If you wanted to get the shell account name, legacy Wikitech username, account creation date, SUL account information, and email address of every Toolforge maintainer you could do something like this using my LDAP shell aliases: {{Codesample|lang=shell-session|code= $ ldap '(&(objectClass=posixAccount)(memberOf=cn=project-tools,ou=groups,dc=wikimedia,dc=org))' uid cn createTimestamp wikimediaGlobalAccountId wikimediaGlobalAccountName mail {{!}} un64 (...snip...) dn: uid=bd808,ou=people,dc=wikimedia,dc=org uid: bd808 cn: BryanDavis createTimestamp: 20130729163514Z wikimediaGlobalAccountId: 12874 wikimediaGlobalAccountName: BryanDavis mail: bdavis@wikimedia.org (...snip...) }} The <code>un64</code> helper can be needed to decode non-ASCII <code>cn</code> values which are stored by LDAP as base-64 encoded strings. The example above shows some of the core data for [[User:BryanDavis|BryanDavis]]'s Developer account: * <code>dn: uid=bd808,ou=people,dc=wikimedia,dc=org</code> -- the "dn" is the primary key for an LDAP record. * <code>uid: bd808</code> -- "uid" in our environment is the account's shell name. * <code>cn: BryanDavis</code> -- "cn" in our environment is the account's "common name" which was also historically the user's Wikitech account name prior to [[News/2024 Migrating Wikitech Account to SUL|October 2024]]. * <code>createTimestamp: 20130729163514Z</code> -- this Developer account was created 2013-07-29 16:35:51 UTC. * <code>wikimediaGlobalAccountId: 12874</code> -- OAuth verified associated SUL account id. * <code>wikimediaGlobalAccountName: BryanDavis</code> -- OAuth verified associated SUL account username. Note that the username may have changed via a global account rename since being stored in LDAP. The SUL account id is invariant and can be used to find the current account name in the centralauth.globaluser database table. * <code>mail: bdavis@wikimedia.org</code> -- The Developer account's email address. Command line tools can be nice for making quick lookups. The output format can be challenging to work with if you really want to create a CSV or TSV data file or to drive other automation. In that case some folks tend to write small Python scripts that use the [https://pypi.org/project/ldap3/ ldap3] library to access the LDAP directory. Here is an example: {{Codesample|lang=python|name=example.py|code= import ldap3 import yaml cfg = yaml.safe_load(open("/etc/ldap.yaml")) # Toolforge bastions and Kubernetes containers have this file at runtime conn = ldap3.Connection(cfg["servers"], auto_bind=True, read_only=True) base = "ou=people,{basedn}".format(basedn=cfg["basedn"]) selector = "(&{})".format( "".join( [ "(objectClass=posixAccount)", "(memberOf=cn=project-tools,ou=groups,dc=wikimedia,dc=org)", ] ) ) r = conn.extend.standard.paged_search( base, selector, attributes=[ "uid", "cn", "createTimestamp", "wikimediaGlobalAccountId", "wikimediaGlobalAccountName", "mail", ], paged_size=256, time_limit=5, generator=True, ) for user in r: # Do interesting stuff with the Developer account here print(user) }} To run this script on Toolforge one would typically use some tool account they have access to (like https://toolsadmin.wikimedia.org/tools/id/bd808-test) and run things from inside a <code>webservice {{toolforge latest image|python3}} shell</code> session: {{Codesample|lang=shell-session|code= $ ssh login.toolforge.org $ become $MY_TOOL_NAME $ webservice {{toolforge latest image|python3}} shell tools.MY_TOOL_NAME@shell-1234567890:~$ python3 -m venv venv tools.MY_TOOL_NAME@shell-1234567890:~$ ./venv/bin/pip3 install ldap3 pyyaml tools.MY_TOOL_NAME@shell-1234567890:~$ vim example.py # Paste in the script # :wq tools.MY_TOOL_NAME@shell-1234567890:~$ ./venv/bin/python3 example.py }} == See also == * [[SRE/LDAP]] * [[mw:Developer account]] * [[User:BryanDavis/LDAP]] [[Category:LDAP]] dejjhkxam4ogcpprugrykthzqwsm5i7 2450678 2450677 2026-08-23T11:46:58Z Taavi 13997 fix 2450678 wikitext text/x-wiki {{See also|mw:Developer account}} A Wikimedia '''developer account''' is an account used for authentication and authorization in Wikimedia's technical spaces such as source code management systems, bug trackers, compute environments, and operational support tooling. Developer accounts [[User:Taavi/Accounts, explained|are distinct]] from [[meta:Help:Unified login|Wikimedia SUL accounts]] which are used for authentication and authorization in Wikimedia project wikis, but in the modern era are often linked to a SUL account that is operated by the same human being or for related automation purposes. Developer account information is stored in an [[w:Lightweight Directory Access Protocol|LDAP directory]] maintained in the Wikimedia production network. More details on the LDAP service itself can be found in [[SRE/LDAP]]. The LDAP directory is used in [[Portal:Cloud VPS|Cloud VPS]] and [[Portal:Toolforge|Toolforge]] to provide Unix account information and ssh public keys to virtual machines. This is the same LDAP directory that backs the [[IDM]] (Bitu) and [[CAS-SSO]] (IDP) services. Within the LDAP directory, developer accounts are <code>objectClass=posixAccount</code> entities stored in the <code>ou=people,dc=wikimedia,dc=org</code> subtree. Cloud VPS project membership is recorded with groups named <code>cn=project-$CLOUD_VPS_NAME,ou=groups,dc=wikimedia,dc=org</code> where <code>$CLOUD_VPS_NAME</code> is the user-friendly OpenStack project name, not the project ID which was previously the same as the name but is now an UUID for new projects. Toolforge tool maintainership is recorded with groups named <code>tools.$TOOLNAME,ou=servicegroups,dc=wikimedia,dc=org</code> where <code>$TOOLNAME</code> is the tool's public name (for example "[[toolforge:versions|versions]]" or "[[toolforge:trainbow|trainbow]]"). == Working with LDAP cli == For simplification of the documentation, cli examples of working with the LDAP directory on this page will generally use the shell aliases that [[User:BryanDavis]] has developed and documented at [[User:BryanDavis/LDAP]]. {{Codesample|lang=shell-session|code= $ alias ldap='ldapsearch -xLLL -P 3 -E pr=5000/noprompt -o ldif-wrap=no -b"dc=wikimedia,dc=org"' $ alias un64='awk '\''BEGIN{FS=":: ";c="base64 -d"}{if(/\w+:: /) {print $2 {{!}}& c; close(c,"to"); c {{!}}& getline $2; close(c); printf("%s:: \"%s\"\n", $1, $2); next} print $0 }'\''' }} == Blocked Developer accounts == Developer accounts can be blocked similar to a full MediaWiki account block. Blocks cascade to disable the account's access rights in Toolforge, Cloud VPS, Gerrit, GitLab, and Phabricator. Blocks are non-destructive and fully reversible. Blocks are typically applied and removed using https://idm.wikimedia.org/wikimedia/block/ which is available to members of the [[toolforge:ldap/group/bitu-account-managers|Bitu account managers]] group. Other tooling may be used in specific workflows such as WMF staff offboarding. Blocked developer accounts can be detected in the LDAP directory by looking for the presence of <code>pwdPolicySubentry=cn=disabled,ou=ppolicies,dc=wikimedia,dc=org</code>. More typically you may want to exclude all blocked accounts from a lookup by using the negated form <code>(!(pwdPolicySubentry=cn=disabled,ou=ppolicies,dc=wikimedia,dc=org))</code>. If you wanted the count of all unblocked Developer accounts: {{Codesample|lang=shell-session|code= $ ldap '(&(objectClass=posixAccount)(!(pwdPolicySubentry=cn=disabled,ou=ppolicies,dc=wikimedia,dc=org)))' dn -b ou=people,dc=wikimedia,dc=org {{!}} grep dn: {{!}} wc -l 36723 # as of 2026-08-20 }} == Developer account details == If you wanted to get the shell account name, legacy Wikitech username, account creation date, SUL account information, and email address of every Toolforge maintainer you could do something like this using my LDAP shell aliases: {{Codesample|lang=shell-session|code= $ ldap '(&(objectClass=posixAccount)(memberOf=cn=project-tools,ou=groups,dc=wikimedia,dc=org))' uid cn createTimestamp wikimediaGlobalAccountId wikimediaGlobalAccountName mail {{!}} un64 (...snip...) dn: uid=bd808,ou=people,dc=wikimedia,dc=org uid: bd808 cn: BryanDavis createTimestamp: 20130729163514Z wikimediaGlobalAccountId: 12874 wikimediaGlobalAccountName: BryanDavis mail: bdavis@wikimedia.org (...snip...) }} The <code>un64</code> helper can be needed to decode non-ASCII <code>cn</code> values which are stored by LDAP as base-64 encoded strings. The example above shows some of the core data for [[User:BryanDavis|BryanDavis]]'s Developer account: * <code>dn: uid=bd808,ou=people,dc=wikimedia,dc=org</code> -- the "dn" is the primary key for an LDAP record. * <code>uid: bd808</code> -- "uid" in our environment is the account's shell name. * <code>cn: BryanDavis</code> -- "cn" in our environment is the account's "common name" which was also historically the user's Wikitech account name prior to [[News/2024 Migrating Wikitech Account to SUL|October 2024]]. * <code>createTimestamp: 20130729163514Z</code> -- this Developer account was created 2013-07-29 16:35:51 UTC. * <code>wikimediaGlobalAccountId: 12874</code> -- OAuth verified associated SUL account id. * <code>wikimediaGlobalAccountName: BryanDavis</code> -- OAuth verified associated SUL account username. Note that the username may have changed via a global account rename since being stored in LDAP. The SUL account id is invariant and can be used to find the current account name in the centralauth.globaluser database table. * <code>mail: bdavis@wikimedia.org</code> -- The Developer account's email address. Command line tools can be nice for making quick lookups. The output format can be challenging to work with if you really want to create a CSV or TSV data file or to drive other automation. In that case some folks tend to write small Python scripts that use the [https://pypi.org/project/ldap3/ ldap3] library to access the LDAP directory. Here is an example: {{Codesample|lang=python|name=example.py|code= import ldap3 import yaml cfg = yaml.safe_load(open("/etc/ldap.yaml")) # Toolforge bastions and Kubernetes containers have this file at runtime conn = ldap3.Connection(cfg["servers"], auto_bind=True, read_only=True) base = "ou=people,{basedn}".format(basedn=cfg["basedn"]) selector = "(&{})".format( "".join( [ "(objectClass=posixAccount)", "(memberOf=cn=project-tools,ou=groups,dc=wikimedia,dc=org)", ] ) ) r = conn.extend.standard.paged_search( base, selector, attributes=[ "uid", "cn", "createTimestamp", "wikimediaGlobalAccountId", "wikimediaGlobalAccountName", "mail", ], paged_size=256, time_limit=5, generator=True, ) for user in r: # Do interesting stuff with the Developer account here print(user) }} To run this script on Toolforge one would typically use some tool account they have access to (like https://toolsadmin.wikimedia.org/tools/id/bd808-test) and run things from inside a <code>webservice {{toolforge latest image|python3}} shell</code> session: {{Codesample|lang=shell-session|code= $ ssh login.toolforge.org $ become $MY_TOOL_NAME $ webservice {{toolforge latest image|python3}} shell tools.MY_TOOL_NAME@shell-1234567890:~$ python3 -m venv venv tools.MY_TOOL_NAME@shell-1234567890:~$ ./venv/bin/pip3 install ldap3 pyyaml tools.MY_TOOL_NAME@shell-1234567890:~$ vim example.py # Paste in the script # :wq tools.MY_TOOL_NAME@shell-1234567890:~$ ./venv/bin/python3 example.py }} == See also == * [[SRE/LDAP]] * [[mw:Developer account]] * [[User:BryanDavis/LDAP]] [[Category:LDAP]] elw36psc1ruherh7yis5z62j67wtdv6 Tool:Gitlab-account-approval/Log 116 453906 2450647 2450236 2026-08-22T15:21:25Z Gitlabaccountapprovalbot 37332 @laraibhasan was approved. 2450647 wikitext text/x-wiki <noinclude>'''Audit log of approvals''' made by [[gitlab:gitlabaccountapprovalbot|@gitlabaccountapprovalbot]]. __NOTOC__</noinclude> === 2026-08-22 === * 15:21 [[gitlab:laraibhasan|@laraibhasan]] was approved. === 2026-08-20 === * 17:18 [[gitlab:jatwa|@jatwa]] was approved. === 2026-08-18 === * 11:33 "wladek92" was rejected (pending since 2026-05-19T11:30:32.679Z). === 2026-08-17 === * 18:39 [[gitlab:erlanger16e|@erlanger16e]] was approved. * 18:33 "ashishkaranam" was rejected (pending since 2026-05-18T18:32:18.813Z). * 12:48 [[gitlab:lfs|@lfs]] was approved. === 2026-08-16 === * 15:42 [[gitlab:pcendrer|@pcendrer]] was approved. === 2026-08-15 === * 17:33 [[gitlab:flammablepizza|@flammablepizza]] was approved. === 2026-08-14 === * 13:00 [[gitlab:pacmand|@pacmand]] was approved. === 2026-08-12 === * 13:21 [[gitlab:doublechek|@doublechek]] was approved. === 2026-08-11 === * 14:15 [[gitlab:mmilenkovicwmf|@mmilenkovicwmf]] was approved. === 2026-08-10 === * 17:18 [[gitlab:mustafe|@mustafe]] was approved. * 10:33 "c8263a20" was rejected (pending since 2026-05-11T10:32:40.353Z). === 2026-08-09 === * 21:18 [[gitlab:iamnetx|@iamnetx]] was approved. * 20:09 "yirba" was rejected (pending since 2026-05-10T20:07:07.738Z). * 06:42 "marsam2489" was rejected (pending since 2026-05-10T06:40:46.276Z). === 2026-08-08 === * 08:15 [[gitlab:taiwaniajusto|@taiwaniajusto]] was approved. === 2026-08-07 === * 18:21 [[gitlab:gturkington|@gturkington]] was approved. * 06:36 "brianbybyby" was rejected (pending since 2026-05-08T06:36:08.459Z). === 2026-07-30 === * 20:42 "horaciocolbert" was rejected (pending since 2026-04-30T20:41:40.421Z). === 2026-07-29 === * 10:03 [[gitlab:piastu|@piastu]] was approved. * 06:33 "rafiul1" was rejected (pending since 2026-04-29T06:33:02.573Z). === 2026-07-28 === * 16:51 [[gitlab:for-each-next|@for-each-next]] was approved. * 09:21 [[gitlab:gka|@gka]] was approved. === 2026-07-27 === * 08:12 [[gitlab:cambob|@cambob]] was approved. === 2026-07-26 === * 15:39 "demansanaagmailcom" was rejected (pending since 2026-04-26T15:39:03.297Z). === 2026-07-24 === * 10:54 [[gitlab:ysogo|@ysogo]] was approved. * 09:39 [[gitlab:pankaj199|@pankaj199]] was approved. === 2026-07-23 === * 12:30 [[gitlab:fermiboson|@fermiboson]] was approved. * 09:21 [[gitlab:slashme|@slashme]] was approved. === 2026-07-22 === * 15:54 [[gitlab:lmedley|@lmedley]] was approved. * 13:24 "praveen5638" was rejected (pending since 2026-04-22T13:21:23.368Z). * 12:48 [[gitlab:cyberpower678|@cyberpower678]] was approved. * 08:30 [[gitlab:nabbegat|@nabbegat]] was approved. * 08:30 [[gitlab:plyd|@plyd]] was approved. * 07:06 "ayush8620" was rejected (pending since 2026-04-22T07:03:12.476Z). === 2026-07-21 === * 15:51 [[gitlab:panieravide|@panieravide]] was approved. * 15:33 [[gitlab:luisvilla-personal|@luisvilla-personal]] was approved. * 13:48 [[gitlab:yru|@yru]] was approved. * 13:09 [[gitlab:deevad|@deevad]] was approved. * 12:54 [[gitlab:ctdo17|@ctdo17]] was approved. * 12:36 [[gitlab:jeannenoiraud|@jeannenoiraud]] was approved. * 12:21 [[gitlab:nadiantara|@nadiantara]] was approved. * 12:18 [[gitlab:wijltcher|@wijltcher]] was approved. * 10:33 [[gitlab:nivopol|@nivopol]] was approved. * 10:30 [[gitlab:johlig|@johlig]] was approved. * 10:30 [[gitlab:majicita|@majicita]] was approved. * 10:15 [[gitlab:yongjiapeng|@yongjiapeng]] was approved. * 10:09 [[gitlab:francyskus|@francyskus]] was approved. * 09:24 [[gitlab:xanonymusx|@xanonymusx]] was approved. === 2026-07-20 === * 17:15 "leonidlednev" was rejected (pending since 2026-04-20T17:13:35.108Z). * 15:27 [[gitlab:rodrigoargenton|@rodrigoargenton]] was approved. * 05:48 "draftecho" was rejected (pending since 2026-04-20T05:48:06.953Z). === 2026-07-19 === * 12:21 [[gitlab:boivie|@boivie]] was approved. === 2026-07-18 === * 16:09 [[gitlab:pharos|@pharos]] was approved. * 15:45 [[gitlab:priyankar22|@priyankar22]] was approved. * 15:30 [[gitlab:sisyph|@sisyph]] was approved. === 2026-07-13 === * 03:45 [[gitlab:dreamyshade|@dreamyshade]] was approved. === 2026-07-12 === * 09:27 [[gitlab:smk|@smk]] was approved. === 2026-07-11 === * 14:48 "bigcereal42" was rejected (pending since 2026-04-11T14:47:40.321Z). * 12:18 "pratyushsawan" was rejected (pending since 2026-04-11T12:16:41.671Z). === 2026-07-10 === * 14:42 [[gitlab:akaza24|@akaza24]] was approved. * 12:57 [[gitlab:kormisk|@kormisk]] was approved. === 2026-07-07 === * 10:36 [[gitlab:olafjanssen|@olafjanssen]] was approved. * 06:57 "elisapoly-99" was rejected (pending since 2026-04-07T06:55:52.662Z). === 2026-07-06 === * 11:57 "ma3rouf" was rejected (pending since 2026-04-06T11:56:30.978Z). === 2026-07-02 === * 15:48 [[gitlab:tekneos|@tekneos]] was approved. === 2026-07-01 === * 15:12 [[gitlab:mugurolevy|@mugurolevy]] was approved. * 14:15 [[gitlab:vadymts1|@vadymts1]] was approved. * 09:57 "mugurolevy" was rejected (pending since 2026-04-01T09:55:19.175Z). === 2026-06-30 === * 14:27 "shivangisharma" was rejected (pending since 2026-03-31T14:26:44.932Z). === 2026-06-29 === * 19:03 [[gitlab:thisismattmiller|@thisismattmiller]] was approved. === 2026-06-28 === * 14:51 "nkwenuinadine" was rejected (pending since 2026-03-29T14:48:32.735Z). * 14:03 "vaishnavikumbhar" was rejected (pending since 2026-03-29T14:01:30.604Z). * 13:03 "stepmay" was rejected (pending since 2026-03-29T13:01:59.905Z). * 06:51 "swallroth" was rejected (pending since 2026-03-29T06:49:54.838Z). === 2026-06-26 === * 09:39 [[gitlab:lakshita28|@lakshita28]] was approved. * 07:30 [[gitlab:reeti|@reeti]] was approved. * 07:30 [[gitlab:anushka10patel|@anushka10patel]] was approved. * 07:30 "samsaesque" was rejected (pending since 2026-03-27T07:29:57.279Z). * 05:51 [[gitlab:arpithhhaaa|@arpithhhaaa]] was approved. * 05:51 [[gitlab:govindlaltl|@govindlaltl]] was approved. === 2026-06-25 === * 16:09 [[gitlab:sakuraemad|@sakuraemad]] was approved. * 07:00 "kdh8219" was rejected (pending since 2026-03-26T06:58:05.415Z). === 2026-06-24 === * 11:54 [[gitlab:sanskardubeydev|@sanskardubeydev]] was approved. * 10:09 "tanmay789q" was rejected (pending since 2026-03-25T10:07:54.602Z). === 2026-06-22 === * 19:57 [[gitlab:gouvernathor|@gouvernathor]] was approved. * 16:45 [[gitlab:lucasbelo|@lucasbelo]] was approved. * 07:15 "jason2000-cpu" was rejected (pending since 2026-03-23T07:14:09.184Z). === 2026-06-21 === * 13:18 [[gitlab:egonw|@egonw]] was approved. === 2026-06-20 === * 10:21 [[gitlab:tways2017|@tways2017]] was approved. === 2026-06-19 === * 16:06 "wilsonwang2026" was rejected (pending since 2026-03-20T16:06:05.511Z). * 04:12 [[gitlab:claudio|@claudio]] was approved. === 2026-06-18 === * 14:21 "royiswariii" was rejected (pending since 2026-03-19T14:19:16.896Z). * 13:06 [[gitlab:laurabarluzzi|@laurabarluzzi]] was approved. === 2026-06-17 === * 11:24 "adinathq8x" was rejected (pending since 2026-03-18T11:22:50.098Z). * 09:45 "nathanveritas" was rejected (pending since 2026-03-18T09:43:51.645Z). === 2026-06-15 === * 22:39 [[gitlab:mohammadhijjawi|@mohammadhijjawi]] was approved. * 14:24 "enlisar" was rejected (pending since 2026-03-16T14:23:00.109Z). * 14:06 "ayaan" was rejected (pending since 2026-03-16T14:03:31.071Z). * 10:54 "kwametech" was rejected (pending since 2026-03-16T10:54:11.083Z). === 2026-06-14 === * 17:45 [[gitlab:surajseth520|@surajseth520]] was approved. * 07:24 "malahimhaseeb" was rejected (pending since 2026-03-15T07:21:57.748Z). === 2026-06-11 === * 11:48 [[gitlab:cadddr|@cadddr]] was approved. * 11:18 "wikipiggy" was rejected (pending since 2026-03-12T11:16:09.335Z). * 07:15 [[gitlab:vesihiisi|@vesihiisi]] was approved. === 2026-06-10 === * 07:03 [[gitlab:dmiranda|@dmiranda]] was approved. === 2026-06-09 === * 14:21 [[gitlab:linkgenetic|@linkgenetic]] was approved. * 14:03 [[gitlab:sjones-ctr|@sjones-ctr]] was approved. * 12:51 [[gitlab:ekrem|@ekrem]] was approved. === 2026-06-08 === * 17:48 "jmprax" was rejected (pending since 2026-03-09T17:46:38.807Z). * 16:15 [[gitlab:ahonc|@ahonc]] was approved. * 12:57 [[gitlab:rainmonger|@rainmonger]] was approved. === 2026-06-07 === * 23:03 "shadowthewuff" was rejected (pending since 2026-03-08T23:00:53.442Z). * 11:45 "wiki-pavan" was rejected (pending since 2026-03-08T11:45:11.116Z). * 02:39 [[gitlab:launchpad|@launchpad]] was approved. === 2026-06-06 === * 14:54 "unicord" was rejected (pending since 2026-03-07T14:52:04.992Z). * 12:48 "chien" was rejected (pending since 2026-03-07T12:48:11.669Z). === 2026-06-04 === * 14:33 "only-vikas" was rejected (pending since 2026-03-05T14:32:09.186Z). === 2026-06-03 === * 15:00 [[gitlab:anafibnshahibul|@anafibnshahibul]] was approved. === 2026-06-02 === * 21:21 "mgagat" was rejected (pending since 2026-03-03T21:18:37.223Z). * 13:57 "prasunaenumarthy" was rejected (pending since 2026-03-03T13:57:14.847Z). * 05:48 [[gitlab:tmoney|@tmoney]] was approved. === 2026-06-01 === * 14:57 "vikram2101" was rejected (pending since 2026-03-02T14:54:26.550Z). * 12:03 "watshell" was rejected (pending since 2026-03-02T12:03:09.329Z). === 2026-05-29 === * 12:48 "mounikapotladurthi" was rejected (pending since 2026-02-27T12:45:38.609Z). === 2026-05-27 === * 20:00 "vinitha" was rejected (pending since 2026-02-25T19:58:43.524Z). * 16:30 "codeurluce" was rejected (pending since 2026-02-25T16:28:53.973Z). * 14:33 [[gitlab:thilio|@thilio]] was approved. === 2026-05-26 === * 12:09 "charisad" was rejected (pending since 2026-02-24T12:07:21.881Z). === 2026-05-25 === * 22:54 "ddshelto" was rejected (pending since 2026-02-23T22:52:44.427Z). * 19:51 "lakz-99" was rejected (pending since 2026-02-23T19:47:00.263Z). * 19:48 "lakz-99" was rejected (pending since 2026-02-23T19:47:00.263Z). === 2026-05-24 === * 18:45 "jiyagupta-cs" was rejected (pending since 2026-02-22T18:43:33.176Z). === 2026-05-23 === * 13:09 [[gitlab:gauthammohanraj|@gauthammohanraj]] was approved. * 04:21 [[gitlab:staraction|@staraction]] was approved. === 2026-05-22 === * 19:03 "i-horich" was rejected (pending since 2026-02-20T19:00:43.519Z). * 01:48 "50323233" was rejected (pending since 2026-02-20T01:48:05.555Z). === 2026-05-21 === * 18:51 "kartikeyg0104" was rejected (pending since 2026-02-19T18:48:39.707Z). * 16:27 [[gitlab:renovatebot|@renovatebot]] was approved. * 16:06 [[gitlab:gkm563|@gkm563]] was approved. === 2026-05-20 === * 01:21 "beedellrokejulianlockhart" was rejected (pending since 2026-02-18T01:19:13.284Z). === 2026-05-18 === * 23:18 "wladek92" was rejected (pending since 2026-02-16T23:16:22.939Z). * 16:36 [[gitlab:effeietsanders|@effeietsanders]] was approved. === 2026-05-14 === * 21:00 [[gitlab:nehemienathan|@nehemienathan]] was approved. === 2026-05-13 === * 10:51 "ssssaaaa" was rejected (pending since 2026-02-11T10:50:36.975Z). === 2026-05-12 === * 18:06 [[gitlab:psubhashish|@psubhashish]] was approved. * 08:12 "khan" was rejected (pending since 2026-02-10T08:11:48.776Z). * 04:27 "galaxysh" was rejected (pending since 2026-02-10T04:24:59.440Z). === 2026-05-11 === * 12:18 "peterxy12" was rejected (pending since 2026-02-09T12:18:01.982Z). === 2026-05-10 === * 11:09 "yalihupokn" was rejected (pending since 2026-02-08T11:06:51.336Z). * 05:12 "wobadha" was rejected (pending since 2026-02-08T05:11:00.569Z). === 2026-05-09 === * 13:45 "bwiki" was rejected (pending since 2026-02-07T13:43:38.177Z). === 2026-05-08 === * 09:24 [[gitlab:cwilliams|@cwilliams]] was approved. === 2026-05-07 === * 14:15 "rehankhan78" was rejected (pending since 2026-02-05T14:13:37.754Z). === 2026-05-06 === * 11:24 "ari" was rejected (pending since 2026-02-04T11:24:11.760Z). * 08:09 [[gitlab:neriah|@neriah]] was approved. * 06:27 [[gitlab:status401|@status401]] was approved. === 2026-05-03 === * 09:54 [[gitlab:anilk|@anilk]] was approved. === 2026-05-02 === * 17:54 [[gitlab:sweil|@sweil]] was approved. * 17:00 [[gitlab:aoppo|@aoppo]] was approved. === 2026-05-01 === * 21:18 [[gitlab:dawalda|@dawalda]] was approved. === 2026-04-30 === * 21:42 "merohibine" was rejected (pending since 2026-01-29T21:40:00.756Z). * 20:54 [[gitlab:tfmorris|@tfmorris]] was approved. * 17:33 [[gitlab:uyen|@uyen]] was approved. * 07:39 [[gitlab:mahveotm|@mahveotm]] was approved. * 06:36 [[gitlab:leo321|@leo321]] was approved. === 2026-04-29 === * 02:27 [[gitlab:dw31415|@dw31415]] was approved. === 2026-04-28 === * 23:09 [[gitlab:dtorsani|@dtorsani]] was approved. === 2026-04-27 === * 23:42 [[gitlab:quinlan|@quinlan]] was approved. * 05:00 [[gitlab:matthewyeager|@matthewyeager]] was approved. === 2026-04-26 === * 17:36 "kuba-hajnej" was rejected (pending since 2026-01-25T17:33:32.467Z). * 13:03 "jklamo" was rejected (pending since 2026-01-25T13:02:22.936Z). === 2026-04-25 === * 20:24 [[gitlab:maldaxura|@maldaxura]] was approved. * 14:33 [[gitlab:sirtobi|@sirtobi]] was approved. * 04:18 "ice5678" was rejected (pending since 2026-01-24T04:15:30.008Z). === 2026-04-24 === * 22:06 [[gitlab:arcstur|@arcstur]] was approved. === 2026-04-22 === * 23:06 "dtorsani" was rejected (pending since 2026-01-21T23:03:25.843Z). * 22:18 [[gitlab:egezort|@egezort]] was approved. * 16:45 "nexpectarpit" was rejected (pending since 2026-01-21T16:43:21.045Z). === 2026-04-20 === * 19:15 "fitch" was rejected (pending since 2026-01-19T19:12:35.644Z). === 2026-04-19 === * 02:54 [[gitlab:neoact|@neoact]] was approved. === 2026-04-18 === * 07:06 [[gitlab:kockaadmiralac|@kockaadmiralac]] was approved. === 2026-04-17 === * 13:42 "liselot" was rejected (pending since 2026-01-16T13:39:41.909Z). === 2026-04-15 === * 17:03 "lahari" was rejected (pending since 2026-01-14T17:02:06.275Z). === 2026-04-14 === * 13:00 "surajseth520" was rejected (pending since 2026-01-13T12:59:45.906Z). * 04:51 [[gitlab:canley|@canley]] was approved. * 01:03 "bshizzle" was rejected (pending since 2026-01-13T01:00:48.120Z). === 2026-04-13 === * 15:30 [[gitlab:passimacopoulos|@passimacopoulos]] was approved. === 2026-04-11 === * 12:30 "krithash" was rejected (pending since 2026-01-10T12:27:24.731Z). === 2026-04-10 === * 15:30 "raunak1709" was rejected (pending since 2026-01-09T15:29:10.901Z). === 2026-04-07 === * 17:03 [[gitlab:supnabla|@supnabla]] was approved. === 2026-04-06 === * 20:00 [[gitlab:laerdon|@laerdon]] was approved. * 19:21 [[gitlab:ljq3|@ljq3]] was approved. === 2026-04-04 === * 11:06 "mixcc" was rejected (pending since 2026-01-03T11:03:33.922Z). === 2026-04-02 === * 05:30 [[gitlab:mbh1|@mbh1]] was approved. === 2026-04-01 === * 18:21 "yuvrajpatil17" was rejected (pending since 2025-12-31T18:20:27.991Z). * 12:12 [[gitlab:amorii0|@amorii0]] was approved. === 2026-03-31 === * 11:00 "krrishsehgal" was rejected (pending since 2025-12-30T11:00:16.384Z). === 2026-03-30 === * 15:36 [[gitlab:atsuko|@atsuko]] was approved. === 2026-03-29 === * 11:36 [[gitlab:giftcup|@giftcup]] was approved. === 2026-03-28 === * 14:51 [[gitlab:janeeva1|@janeeva1]] was approved. === 2026-03-26 === * 13:36 [[gitlab:saiphani02|@saiphani02]] was approved. * 11:48 [[gitlab:valerioboz-wmch|@valerioboz-wmch]] was approved. === 2026-03-25 === * 09:45 "quansi" was rejected (pending since 2025-12-24T09:42:13.451Z). * 02:18 [[gitlab:viztor|@viztor]] was approved. === 2026-03-24 === * 23:18 [[gitlab:maryyann|@maryyann]] was approved. * 23:01 [[gitlab:codenamenoreste|@codenamenoreste]] was approved. * 13:36 [[gitlab:marc-maillard-wmse|@marc-maillard-wmse]] was approved. * 07:39 "fred2675" was rejected (pending since 2025-12-23T07:39:11.380Z). === 2026-03-23 === * 14:51 [[gitlab:komla|@komla]] was approved. * 05:51 "lunachuck43" was rejected (pending since 2025-12-22T05:50:17.862Z). * 04:06 "reza110011" was rejected (pending since 2025-12-22T04:05:25.117Z). === 2026-03-20 === * 21:54 "mertgor" was rejected (pending since 2025-12-19T21:51:51.419Z). * 20:57 "autanmahmah" was rejected (pending since 2025-12-19T20:54:51.678Z). * 09:57 [[gitlab:nethahussain|@nethahussain]] was approved. * 09:27 [[gitlab:piewriter|@piewriter]] was approved. * 08:15 [[gitlab:dondersmooi|@dondersmooi]] was approved. === 2026-03-19 === * 21:03 "sayvhior" was rejected (pending since 2025-12-18T21:02:31.699Z). === 2026-03-18 === * 20:15 [[gitlab:martinmystere|@martinmystere]] was approved. === 2026-03-17 === * 02:51 "louperivois" was rejected (pending since 2025-12-16T02:50:48.197Z). === 2026-03-16 === * 12:54 "mokayaj857" was rejected (pending since 2025-12-15T12:53:39.015Z). * 06:18 "roamer15" was rejected (pending since 2025-12-15T06:16:38.042Z). === 2026-03-14 === * 11:12 "umaramuhammad" was rejected (pending since 2025-12-13T11:10:44.004Z). * 09:33 "akuma19" was rejected (pending since 2025-12-13T09:31:39.044Z). * 07:06 [[gitlab:syunsyunminmin|@syunsyunminmin]] was approved. === 2026-03-12 === * 20:24 [[gitlab:11wb|@11wb]] was approved. * 09:54 [[gitlab:bcxfu75k|@bcxfu75k]] was approved. === 2026-03-10 === * 09:12 [[gitlab:viktoriahillerudwmse|@viktoriahillerudwmse]] was approved. === 2026-03-06 === * 08:09 "vazhayilnewone" was rejected (pending since 2025-12-05T08:07:02.184Z). === 2026-03-04 === * 20:54 [[gitlab:elphie|@elphie]] was approved. * 11:39 "ronaldahmed" was rejected (pending since 2025-12-03T11:37:47.492Z). * 02:12 "ltslw" was rejected (pending since 2025-12-03T02:11:52.040Z). === 2026-03-02 === * 19:21 "dlopez350" was rejected (pending since 2025-12-01T19:20:38.918Z). * 18:15 [[gitlab:lsandergreen|@lsandergreen]] was approved. === 2026-03-01 === * 10:51 [[gitlab:clintacc|@clintacc]] was approved. === 2026-02-28 === * 09:24 "cardboardlamp" was rejected (pending since 2025-11-29T09:22:03.947Z). * 08:18 "wiki-pavan" was rejected (pending since 2025-11-29T08:16:24.184Z). === 2026-02-27 === * 20:45 "thisisrick25" was rejected (pending since 2025-11-28T20:42:24.454Z). === 2026-02-26 === * 13:57 "chuiimuiiofc" was rejected (pending since 2025-11-27T13:57:02.794Z). * 13:54 "steffpro" was rejected (pending since 2025-11-27T13:52:10.859Z). === 2026-02-25 === * 21:24 "abubakarhabibudayyabu" was rejected (pending since 2025-11-26T21:22:37.776Z). === 2026-02-24 === * 05:00 "playboi" was rejected (pending since 2025-11-25T05:00:30.762Z). === 2026-02-23 === * 14:00 "alph65" was rejected (pending since 2025-11-24T13:59:00.797Z). * 12:33 [[gitlab:robertsky|@robertsky]] was approved. === 2026-02-22 === * 00:30 "hp8p" was rejected (pending since 2025-11-23T00:29:24.741Z). === 2026-02-19 === * 16:45 "clayjar" was rejected (pending since 2025-11-20T16:44:48.380Z). === 2026-02-18 === * 22:18 "nexus" was rejected (pending since 2025-11-19T22:16:48.818Z). * 12:00 "bernsteinnn" was rejected (pending since 2025-11-19T11:59:04.427Z). === 2026-02-17 === * 11:36 "jason2000-cpu" was rejected (pending since 2025-11-18T11:34:00.314Z). === 2026-02-16 === * 14:54 "smaurya" was rejected (pending since 2025-11-17T14:52:06.906Z). === 2026-02-15 === * 16:51 "kra-79" was rejected (pending since 2025-11-16T16:50:41.375Z). === 2026-02-14 === * 15:15 [[gitlab:mess|@mess]] was approved. === 2026-02-13 === * 13:57 "sopalsuemae957" was rejected (pending since 2025-11-14T13:55:16.921Z). * 13:30 [[gitlab:wyslijp16-toolforge|@wyslijp16-toolforge]] was approved. === 2026-02-12 === * 16:30 "kristinagligoric" was rejected (pending since 2025-11-13T16:29:21.646Z). * 03:33 [[gitlab:anyehansen|@anyehansen]] was approved. * 02:21 [[gitlab:thejoyfultentmaker|@thejoyfultentmaker]] was approved. === 2026-02-10 === * 13:18 [[gitlab:db111|@db111]] was approved. === 2026-02-09 === * 19:06 "squirrel289" was rejected (pending since 2025-11-10T19:04:27.831Z). === 2026-02-06 === * 20:54 [[gitlab:gillux|@gillux]] was approved. * 09:09 [[gitlab:lih|@lih]] was approved. === 2026-01-31 === * 16:21 [[gitlab:taxonbot1|@taxonbot1]] was approved. === 2026-01-28 === * 14:30 [[gitlab:ademola|@ademola]] was approved. * 10:51 "watshell" was rejected (pending since 2025-10-29T10:51:01.521Z). === 2026-01-26 === * 23:06 "tavaresgmg" was rejected (pending since 2025-10-27T23:04:42.140Z). === 2026-01-25 === * 06:03 "cata" was rejected (pending since 2025-10-26T06:01:26.155Z). === 2026-01-24 === * 21:15 [[gitlab:wiegels|@wiegels]] was approved. * 06:30 [[gitlab:blaquans|@blaquans]] was approved. === 2026-01-23 === * 16:27 [[gitlab:lerickson|@lerickson]] was approved. * 10:15 "fran0035g" was rejected (pending since 2025-10-24T10:12:17.732Z). === 2026-01-22 === * 21:00 "hacksyn" was rejected (pending since 2025-10-23T20:59:15.982Z). === 2026-01-21 === * 17:30 [[gitlab:otcenas11|@otcenas11]] was approved. === 2026-01-19 === * 21:48 [[gitlab:amdrel|@amdrel]] was approved. * 04:36 "rayalexa" was rejected (pending since 2025-10-20T04:35:02.094Z). === 2026-01-18 === * 15:45 "somya" was rejected (pending since 2025-10-19T15:43:43.701Z). * 06:54 "sergg001" was rejected (pending since 2025-10-19T06:54:12.296Z). === 2026-01-16 === * 11:57 "zeejohsy" was rejected (pending since 2025-10-17T11:56:22.372Z). * 04:45 "rocky25" was rejected (pending since 2025-10-17T04:43:33.180Z). === 2026-01-15 === * 16:39 "tiisu" was rejected (pending since 2025-10-16T16:37:18.438Z). * 12:00 "noahalorwu" was rejected (pending since 2025-10-16T11:58:26.133Z). * 10:39 "prjayaiuedu" was rejected (pending since 2025-10-16T10:37:16.947Z). === 2026-01-13 === * 17:21 [[gitlab:lwilson-ctr|@lwilson-ctr]] was approved. === 2026-01-12 === * 17:03 "stagietechs" was rejected (pending since 2025-10-13T17:02:25.281Z). === 2026-01-10 === * 19:06 "keerthisr" was rejected (pending since 2025-10-11T19:05:01.758Z). === 2026-01-09 === * 20:36 "lightb" was rejected (pending since 2025-10-10T20:34:20.264Z). === 2026-01-08 === * 19:42 [[gitlab:tbodt|@tbodt]] was approved. * 13:57 [[gitlab:martynranyard|@martynranyard]] was approved. === 2026-01-07 === * 17:48 [[gitlab:santanuwiki25|@santanuwiki25]] was approved. * 14:27 "dipanshu" was rejected (pending since 2025-10-08T14:26:10.794Z). * 12:30 "adeolaadesina" was rejected (pending since 2025-10-08T12:29:49.592Z). * 09:21 "tony-kamande" was rejected (pending since 2025-10-08T09:20:28.421Z). * 06:18 "hninwuttyi" was rejected (pending since 2025-10-08T06:17:28.006Z). * 05:09 "andume" was rejected (pending since 2025-10-08T05:07:18.582Z). * 02:00 "mosope" was rejected (pending since 2025-10-08T01:59:54.800Z). * 01:15 [[gitlab:tungstalite|@tungstalite]] was approved. === 2026-01-06 === * 18:24 "leerensucher" was rejected (pending since 2025-10-07T18:21:41.253Z). * 14:54 "leonidlednev" was rejected (pending since 2025-10-07T14:53:07.273Z). * 12:57 "alexandre-tingaud" was rejected (pending since 2025-10-07T12:54:27.206Z). === 2026-01-04 === * 21:33 [[gitlab:matr1x-101|@matr1x-101]] was approved. * 15:18 "makjr" was rejected (pending since 2025-10-05T15:16:31.558Z). * 14:09 "dakshq" was rejected (pending since 2025-10-05T14:08:40.608Z). === 2026-01-03 === * 20:42 [[gitlab:apehitkey|@apehitkey]] was approved. * 18:00 [[gitlab:jeremyb|@jeremyb]] was approved. * 14:09 [[gitlab:twelephant|@twelephant]] was approved. === 2026-01-01 === * 11:30 "shellstanislav" was rejected (pending since 2025-10-02T11:29:10.150Z). === 2025-12-30 === * 19:51 "camilojdiaz" was rejected (pending since 2025-09-30T19:49:24.913Z). === 2025-12-29 === * 16:03 "zied" was rejected (pending since 2025-09-29T16:01:30.415Z). * 08:18 "rahulsidpradhan" was rejected (pending since 2025-09-29T08:17:02.849Z). === 2025-12-26 === * 09:48 "thembo42" was rejected (pending since 2025-09-26T09:45:15.033Z). === 2025-12-25 === * 14:03 "196936074751" was rejected (pending since 2025-09-25T14:02:31.367Z). === 2025-12-23 === * 16:21 "ngarnsworthy" was rejected (pending since 2025-09-23T16:20:41.211Z). === 2025-12-22 === * 12:39 "aza555" was rejected (pending since 2025-09-22T12:38:02.622Z). === 2025-12-20 === * 23:45 "saph" was rejected (pending since 2025-09-20T23:45:01.222Z). === 2025-12-19 === * 10:15 "vladdymoses" was rejected (pending since 2025-09-19T10:15:00.999Z). * 07:15 "dirtylittlepoobah" was rejected (pending since 2025-09-19T07:13:55.537Z). === 2025-12-18 === * 16:24 [[gitlab:guyfawcus|@guyfawcus]] was approved. === 2025-12-17 === * 21:39 [[gitlab:holdyourhorses|@holdyourhorses]] was approved. * 18:30 "prudencia" was rejected (pending since 2025-09-17T18:27:18.860Z). * 02:24 "lottie" was rejected (pending since 2025-09-17T02:21:21.744Z). === 2025-12-16 === * 09:39 [[gitlab:melcatherine|@melcatherine]] was approved. * 08:54 [[gitlab:leila237|@leila237]] was approved. === 2025-12-15 === * 18:27 [[gitlab:royalsailor|@royalsailor]] was approved. * 09:39 [[gitlab:olaf8940|@olaf8940]] was approved. * 09:39 "brianbybyby" was rejected (pending since 2025-09-15T09:37:45.430Z). === 2025-12-14 === * 20:21 [[gitlab:essa237|@essa237]] was approved. * 16:42 [[gitlab:bovimacoco|@bovimacoco]] was approved. === 2025-12-13 === * 21:54 "mmns21" was rejected (pending since 2025-09-13T21:52:24.017Z). * 20:33 "bugcrawler" was rejected (pending since 2025-09-13T20:31:09.211Z). === 2025-12-12 === * 14:39 "ruvchoudhary" was rejected (pending since 2025-09-12T14:36:16.167Z). * 06:54 "rezadress" was rejected (pending since 2025-09-12T06:52:21.749Z). === 2025-12-10 === * 17:30 [[gitlab:itsmoon|@itsmoon]] was approved. === 2025-12-09 === * 15:42 [[gitlab:mercy-o|@mercy-o]] was approved. === 2025-12-06 === * 16:45 "jacquesradjabu" was rejected (pending since 2025-09-06T16:45:17.969Z). * 11:27 [[gitlab:ikhitron|@ikhitron]] was approved. === 2025-12-01 === * 08:12 "halconmilenario21" was rejected (pending since 2025-09-01T08:12:10.262Z). === 2025-11-30 === * 21:06 [[gitlab:habs|@habs]] was approved. === 2025-11-29 === * 16:36 "bovimacoco" was rejected (pending since 2025-08-30T16:34:39.712Z). * 00:45 [[gitlab:jjpmaster|@jjpmaster]] was approved. === 2025-11-24 === * 10:30 "alph65" was rejected (pending since 2025-08-25T10:28:40.957Z). * 02:24 [[gitlab:yaron|@yaron]] was approved. === 2025-11-20 === * 16:06 "clayjar" was rejected (pending since 2025-08-21T16:04:54.450Z). === 2025-11-17 === * 21:09 [[gitlab:ankita97531|@ankita97531]] was approved. === 2025-11-16 === * 14:15 "commanderkefir" was rejected (pending since 2025-08-17T14:13:14.791Z). * 08:21 "rehankhan78" was rejected (pending since 2025-08-17T08:19:44.896Z). === 2025-11-15 === * 14:36 "cyberscribe" was rejected (pending since 2025-08-16T14:34:27.230Z). === 2025-11-13 === * 04:21 "waddie96" was rejected (pending since 2025-08-14T04:19:27.461Z). === 2025-11-11 === * 06:42 [[gitlab:seanhoyland|@seanhoyland]] was approved. === 2025-11-10 === * 00:06 [[gitlab:jaredblumer|@jaredblumer]] was approved. === 2025-11-09 === * 22:36 "heinxiety" was rejected (pending since 2025-08-10T22:33:12.041Z). === 2025-11-07 === * 22:00 [[gitlab:forzagreen|@forzagreen]] was approved. === 2025-11-06 === * 16:57 [[gitlab:rsilvola|@rsilvola]] was approved. === 2025-11-04 === * 21:24 [[gitlab:devdoingdev|@devdoingdev]] was approved. === 2025-11-03 === * 17:48 "joewaleed98" was rejected (pending since 2025-08-04T17:46:12.191Z). === 2025-11-01 === * 18:00 "eliasempresas" was rejected (pending since 2025-08-02T17:58:04.412Z). === 2025-10-31 === * 18:51 [[gitlab:chaoticenby|@chaoticenby]] was approved. * 04:33 "3ch310n" was rejected (pending since 2025-08-01T04:32:21.982Z). === 2025-10-30 === * 10:03 [[gitlab:tausheefhassan|@tausheefhassan]] was approved. === 2025-10-29 === * 14:54 "theap" was rejected (pending since 2025-07-30T14:52:12.066Z). === 2025-10-28 === * 06:06 [[gitlab:tanbiruzzaman|@tanbiruzzaman]] was approved. === 2025-10-27 === * 07:51 [[gitlab:jmoore111|@jmoore111]] was approved. === 2025-10-25 === * 21:09 [[gitlab:valor|@valor]] was approved. * 21:03 [[gitlab:booksmurf|@booksmurf]] was approved. * 02:48 "mystyc1" was rejected (pending since 2025-07-26T02:46:19.373Z). === 2025-10-24 === * 05:12 "aadarshmahesh" was rejected (pending since 2025-07-25T05:09:38.264Z). === 2025-10-22 === * 20:54 [[gitlab:janewanga|@janewanga]] was approved. * 17:27 "abeljeevan" was rejected (pending since 2025-07-23T17:26:46.884Z). * 16:12 "shrimpnaur" was rejected (pending since 2025-07-23T16:10:37.864Z). === 2025-10-21 === * 18:51 "jrmuizel" was rejected (pending since 2025-07-22T18:50:07.315Z). * 09:33 [[gitlab:dpogorzelski|@dpogorzelski]] was approved. === 2025-10-17 === * 13:21 [[gitlab:blegodwin|@blegodwin]] was approved. === 2025-10-16 === * 14:51 [[gitlab:bahago|@bahago]] was approved. * 14:12 "harikrishna0005" was rejected (pending since 2025-07-17T14:10:48.385Z). * 14:09 "gauthammohanraj" was rejected (pending since 2025-07-17T14:08:47.643Z). === 2025-10-15 === * 13:48 [[gitlab:adwivedii|@adwivedii]] was approved. * 13:18 [[gitlab:kimbrenekakande|@kimbrenekakande]] was approved. * 13:03 "childmnajennifer" was rejected (pending since 2025-07-16T13:01:50.236Z). * 05:06 "vssb4214" was rejected (pending since 2025-07-16T05:05:33.985Z). === 2025-10-14 === * 19:39 [[gitlab:afanyulionel|@afanyulionel]] was approved. * 15:33 [[gitlab:sadrettin|@sadrettin]] was approved. * 14:18 [[gitlab:tmwyk|@tmwyk]] was approved. * 08:42 "yasu0796" was rejected (pending since 2025-07-15T08:41:26.453Z). === 2025-10-13 === * 16:09 [[gitlab:atlas0007|@atlas0007]] was approved. === 2025-10-11 === * 17:42 [[gitlab:techwizzie|@techwizzie]] was approved. === 2025-10-10 === * 19:03 [[gitlab:miiswom|@miiswom]] was approved. * 16:06 [[gitlab:ninatakang|@ninatakang]] was approved. === 2025-10-09 === * 15:42 [[gitlab:jaykaneki|@jaykaneki]] was approved. * 14:21 [[gitlab:lebogang|@lebogang]] was approved. * 14:15 [[gitlab:kimondorose|@kimondorose]] was approved. * 13:48 [[gitlab:joyakinyi|@joyakinyi]] was approved. * 13:48 [[gitlab:dikshyashahi|@dikshyashahi]] was approved. * 13:45 [[gitlab:obediobadiah|@obediobadiah]] was approved. * 13:45 [[gitlab:system625|@system625]] was approved. * 13:45 [[gitlab:rolalove|@rolalove]] was approved. * 13:39 [[gitlab:olatundeawo|@olatundeawo]] was approved. * 13:36 [[gitlab:danielchristlight|@danielchristlight]] was approved. * 13:36 [[gitlab:dipanshu1223|@dipanshu1223]] was approved. * 13:36 [[gitlab:aradhya|@aradhya]] was approved. * 09:57 "bognd" was rejected (pending since 2025-07-10T09:55:48.661Z). === 2025-10-08 === * 23:36 [[gitlab:sopzy|@sopzy]] was approved. * 23:03 [[gitlab:oluwatumininu|@oluwatumininu]] was approved. * 19:39 [[gitlab:levon003|@levon003]] was approved. * 15:24 [[gitlab:ritika-bhambri11|@ritika-bhambri11]] was approved. * 13:45 [[gitlab:anbanguyen|@anbanguyen]] was approved. * 13:36 [[gitlab:chumzine|@chumzine]] was approved. * 13:27 [[gitlab:shr0x-ya|@shr0x-ya]] was approved. * 12:45 [[gitlab:nurahwakili|@nurahwakili]] was approved. * 03:42 "nazhiba" was rejected (pending since 2025-07-09T03:40:12.625Z). * 02:12 "mafennel" was rejected (pending since 2025-07-09T02:11:40.598Z). === 2025-10-07 === * 22:54 [[gitlab:olusegunfaj|@olusegunfaj]] was approved. * 21:30 [[gitlab:rona|@rona]] was approved. * 21:09 [[gitlab:sandijigs|@sandijigs]] was approved. * 13:36 "xisbajao" was rejected (pending since 2025-07-08T13:33:35.018Z). * 01:36 "areczek94" was rejected (pending since 2025-07-08T01:35:40.633Z). === 2025-10-06 === * 19:21 "wmcarter2017" was rejected (pending since 2025-07-07T19:21:12.899Z). === 2025-10-05 === * 14:15 "meetmendapara" was rejected (pending since 2025-07-06T14:14:16.726Z). === 2025-10-04 === * 20:51 "nftbaee" was rejected (pending since 2025-07-05T20:50:57.688Z). === 2025-10-03 === * 06:12 [[gitlab:javiermonton|@javiermonton]] was approved. === 2025-10-02 === * 20:15 "talaqalotaibipmp" was rejected (pending since 2025-07-03T20:13:05.164Z). === 2025-10-01 === * 10:54 "bjensen" was rejected (pending since 2025-07-02T10:53:46.574Z). * 02:45 "kowal1984" was rejected (pending since 2025-07-02T02:44:56.946Z). === 2025-09-30 === * 21:21 [[gitlab:kavaljeetsingh|@kavaljeetsingh]] was approved. * 00:24 "adium" was rejected (pending since 2025-07-01T00:23:43.807Z). === 2025-09-28 === * 08:54 [[gitlab:pexerik|@pexerik]] was approved. === 2025-09-27 === * 13:57 [[gitlab:rubahhitamvukova|@rubahhitamvukova]] was approved. === 2025-09-26 === * 16:57 "algorithmic" was rejected (pending since 2025-06-27T16:56:17.480Z). * 13:54 [[gitlab:shadabgdg|@shadabgdg]] was approved. * 13:12 [[gitlab:spushpit|@spushpit]] was approved. === 2025-09-20 === * 14:06 "bwiki" was rejected (pending since 2025-06-21T13:59:14.749Z). === 2025-09-16 === * 05:39 [[gitlab:deepchirp|@deepchirp]] was approved. === 2025-09-15 === * 22:00 [[gitlab:noisk8|@noisk8]] was approved. * 11:03 "ahonc" was rejected (pending since 2025-06-16T11:00:54.843Z). === 2025-09-13 === * 18:24 "a-ssh22" was rejected (pending since 2025-06-14T18:23:33.937Z). * 12:36 [[gitlab:rajashreetalukdar|@rajashreetalukdar]] was approved. * 00:45 [[gitlab:sumitsurai|@sumitsurai]] was approved. === 2025-09-12 === * 17:12 [[gitlab:suyash23|@suyash23]] was approved. * 00:46 "remotetravel" was rejected (pending since 2025-06-13T00:44:08.171Z). === 2025-09-10 === * 21:09 "jancborchardt" was rejected (pending since 2025-06-11T21:06:30.759Z). === 2025-09-09 === * 17:03 [[gitlab:vwf|@vwf]] was approved. * 06:36 [[gitlab:cactusisme|@cactusisme]] was approved. === 2025-09-08 === * 18:09 "birushandegeya" was rejected (pending since 2025-06-09T18:08:00.087Z). * 16:27 "ngarnsworthy" was rejected (pending since 2025-06-09T16:24:37.213Z). * 12:33 "zolgoyo" was rejected (pending since 2025-06-09T12:31:34.199Z). === 2025-09-06 === * 23:09 [[gitlab:jaishsingh913|@jaishsingh913]] was approved. === 2025-09-05 === * 21:45 [[gitlab:sakshi2|@sakshi2]] was approved. * 20:42 "abdukhaliq1" was rejected (pending since 2025-06-06T20:40:42.023Z). * 14:27 "beubsamy" was rejected (pending since 2025-06-06T14:27:06.781Z). === 2025-09-04 === * 23:27 "sdhehua" was rejected (pending since 2025-06-05T23:24:45.777Z). * 19:00 [[gitlab:perry|@perry]] was approved. * 11:24 "saintwolf" was rejected (pending since 2025-06-05T11:21:20.176Z). === 2025-09-02 === * 05:48 [[gitlab:aliu|@aliu]] was approved. === 2025-08-29 === * 13:30 "kksurendran066" was rejected (pending since 2025-05-30T13:27:48.755Z). === 2025-08-28 === * 22:18 "tauraamuix" was rejected (pending since 2025-05-29T22:16:08.228Z). === 2025-08-26 === * 19:03 [[gitlab:dikkulah|@dikkulah]] was approved. === 2025-08-22 === * 23:51 [[gitlab:khoroshun_mike|@khoroshun_mike]] was approved. === 2025-08-21 === * 07:39 [[gitlab:yuka|@yuka]] was approved. === 2025-08-19 === * 07:48 [[gitlab:zhaofjx|@zhaofjx]] was approved. === 2025-08-17 === * 14:27 "madhan13k" was rejected (pending since 2025-05-18T14:26:08.973Z). === 2025-08-15 === * 10:15 "mohammed_abukhadra" was rejected (pending since 2025-05-16T10:14:48.403Z). === 2025-08-11 === * 11:48 "hmmyesbro" was rejected (pending since 2025-05-12T11:45:24.350Z). === 2025-08-10 === * 13:15 [[gitlab:dactyl|@dactyl]] was approved. === 2025-08-09 === * 04:39 "xxxx100000" was rejected (pending since 2025-05-10T04:37:44.949Z). === 2025-08-08 === * 14:33 [[gitlab:josefanthony|@josefanthony]] was approved. === 2025-08-07 === * 23:42 [[gitlab:robins7|@robins7]] was approved. * 21:42 [[gitlab:pols12|@pols12]] was approved. * 17:15 "sbronson" was rejected (pending since 2025-05-08T17:15:08.834Z). * 14:57 [[gitlab:alvindulle|@alvindulle]] was approved. * 14:45 [[gitlab:xentos|@xentos]] was approved. * 06:27 "jamesboste" was rejected (pending since 2025-05-08T06:25:14.793Z). * 03:57 "ysun" was rejected (pending since 2025-05-08T03:55:07.348Z). === 2025-08-06 === * 21:51 "pols12" was rejected (pending since 2025-05-07T21:49:13.598Z). * 01:51 "okeamah" was rejected (pending since 2025-05-07T01:48:50.114Z). === 2025-08-05 === * 09:15 "mobashir-2013" was rejected (pending since 2025-05-06T09:14:24.069Z). === 2025-08-01 === * 08:00 "douginamug" was rejected (pending since 2025-05-02T07:57:38.317Z). === 2025-07-31 === * 02:30 [[gitlab:ads|@ads]] was approved. === 2025-07-27 === * 13:15 "mrico2703" was rejected (pending since 2025-04-27T13:13:12.346Z). * 10:17 [[gitlab:josephfrancis12|@josephfrancis12]] was approved. * 10:17 [[gitlab:fuzzew|@fuzzew]] was approved. * 05:57 [[gitlab:biscuitbobby|@biscuitbobby]] was approved. * 05:48 [[gitlab:ecoholic|@ecoholic]] was approved. === 2025-07-26 === * 11:48 [[gitlab:chimnayyyy|@chimnayyyy]] was approved. * 11:48 [[gitlab:alwinalbert|@alwinalbert]] was approved. * 11:48 [[gitlab:hridyakk|@hridyakk]] was approved. * 11:45 [[gitlab:gaurigupta21|@gaurigupta21]] was approved. * 11:45 [[gitlab:binetaa|@binetaa]] was approved. * 10:21 [[gitlab:jyothikat22|@jyothikat22]] was approved. * 10:21 [[gitlab:zobotrombie|@zobotrombie]] was approved. * 10:21 [[gitlab:flykrth|@flykrth]] was approved. * 10:21 [[gitlab:mehrinshamim|@mehrinshamim]] was approved. * 10:21 [[gitlab:aadhi13|@aadhi13]] was approved. * 10:21 [[gitlab:malavikam05|@malavikam05]] was approved. * 10:18 [[gitlab:nf609|@nf609]] was approved. * 05:48 [[gitlab:nazalnihad|@nazalnihad]] was approved. * 05:48 [[gitlab:naveen28204280|@naveen28204280]] was approved. === 2025-07-25 === * 09:49 [[gitlab:kasyap9|@kasyap9]] was approved. * 09:30 [[gitlab:swayamagrahari|@swayamagrahari]] was approved. === 2025-07-24 === * 19:36 [[gitlab:madutgn|@madutgn]] was approved. === 2025-07-23 === * 20:09 [[gitlab:somerandomdeveloper|@somerandomdeveloper]] was approved. === 2025-07-22 === * 00:15 [[gitlab:iagoqnsi|@iagoqnsi]] was approved. === 2025-07-21 === * 17:30 [[gitlab:asadiqui|@asadiqui]] was approved. * 16:39 [[gitlab:tryvix1509|@tryvix1509]] was approved. * 04:27 [[gitlab:damian|@damian]] was approved. === 2025-07-20 === * 09:42 "mike-khoroshun" was rejected (pending since 2025-04-20T09:42:22.732Z). === 2025-07-17 === * 17:57 [[gitlab:haroldkrabs|@haroldkrabs]] was approved. * 13:45 [[gitlab:envlh|@envlh]] was approved. === 2025-07-14 === * 10:24 [[gitlab:missguru|@missguru]] was approved. * 00:57 "clarfonthey" was rejected (pending since 2025-04-14T00:56:32.626Z). === 2025-07-13 === * 01:01 [[gitlab:l235|@l235]] was approved. === 2025-07-11 === * 03:06 "rodavlas" was rejected (pending since 2025-04-11T03:05:45.590Z). === 2025-07-06 === * 00:09 "lakasa" was rejected (pending since 2025-04-06T00:06:28.469Z). === 2025-07-05 === * 21:54 "ctrlzvi" was rejected (pending since 2025-04-05T21:54:12.542Z). * 14:30 "aminualiyu" was rejected (pending since 2025-04-05T14:27:22.617Z). === 2025-07-04 === * 03:15 [[gitlab:galstar|@galstar]] was approved. === 2025-07-02 === * 11:27 "vicolas11" was rejected (pending since 2025-04-02T11:25:12.682Z). === 2025-06-29 === * 23:12 "naomi723" was rejected (pending since 2025-03-30T23:09:24.630Z). === 2025-06-28 === * 16:21 "mudeh2372" was rejected (pending since 2025-03-29T16:18:27.057Z). === 2025-06-27 === * 23:18 "rony143" was rejected (pending since 2025-03-28T23:16:13.671Z). * 22:21 [[gitlab:rluts|@rluts]] was approved. === 2025-06-26 === * 13:54 "creativegurus" was rejected (pending since 2025-03-27T13:52:41.706Z). === 2025-06-24 === * 17:42 [[gitlab:devjadiya|@devjadiya]] was approved. * 14:00 "dominic-r" was rejected (pending since 2025-03-25T14:00:07.307Z). === 2025-06-21 === * 00:48 [[gitlab:vriaa|@vriaa]] was approved. === 2025-06-18 === * 15:21 "ayushkhati1" was rejected (pending since 2025-03-19T15:18:50.062Z). === 2025-06-17 === * 20:45 "chiomavero" was rejected (pending since 2025-03-18T20:44:13.967Z). * 00:27 [[gitlab:eggroll97|@eggroll97]] was approved. === 2025-06-14 === * 20:57 "volvox" was rejected (pending since 2025-03-15T20:56:34.018Z). === 2025-06-13 === * 16:09 [[gitlab:supergrey|@supergrey]] was approved. * 11:03 "chqaz" was rejected (pending since 2025-03-14T11:01:09.600Z). * 10:24 [[gitlab:slong-wmf|@slong-wmf]] was approved. * 10:15 "hearvox" was rejected (pending since 2025-03-14T10:13:13.112Z). === 2025-06-12 === * 15:18 "jlam" was rejected (pending since 2025-03-13T15:17:54.099Z). === 2025-06-09 === * 20:48 "dipanjansengupta" was rejected (pending since 2025-03-10T20:48:03.545Z). * 19:27 [[gitlab:reggycelly|@reggycelly]] was approved. * 14:51 "arendpieter" was rejected (pending since 2025-03-10T14:51:01.445Z). * 13:21 [[gitlab:greenreaper|@greenreaper]] was approved. * 09:33 [[gitlab:mmta|@mmta]] was approved. * 08:03 "a-ssh22" was rejected (pending since 2025-03-10T08:03:08.111Z). === 2025-06-08 === * 21:06 "mm-episodenlistedlvaupdater" was rejected (pending since 2025-03-09T21:04:06.323Z). === 2025-06-06 === * 11:06 [[gitlab:olea|@olea]] was approved. === 2025-06-05 === * 20:33 [[gitlab:encodedwp|@encodedwp]] was approved. * 15:00 [[gitlab:toluayo|@toluayo]] was approved. * 13:51 [[gitlab:arnold_lup|@arnold_lup]] was approved. * 11:54 "sdhehua" was rejected (pending since 2025-03-06T11:51:48.241Z). === 2025-06-03 === * 21:27 [[gitlab:wewakey|@wewakey]] was approved. * 12:36 "hunsimon2" was rejected (pending since 2025-03-04T12:34:56.520Z). * 11:54 "hunsimon" was rejected (pending since 2025-03-04T11:53:54.652Z). === 2025-06-02 === * 12:01 [[gitlab:jaimedes|@jaimedes]] was approved. === 2025-05-30 === * 18:00 "sathvik9105" was rejected (pending since 2025-02-28T17:59:42.867Z). * 11:21 [[gitlab:tonythomas01|@tonythomas01]] was approved. * 10:06 [[gitlab:gpsleo|@gpsleo]] was approved. === 2025-05-29 === * 22:12 [[gitlab:codynguyen1116|@codynguyen1116]] was approved. === 2025-05-28 === * 02:57 [[gitlab:saper|@saper]] was approved. === 2025-05-27 === * 21:06 [[gitlab:mohammed_qays|@mohammed_qays]] was approved. * 15:33 "satanluimm" was rejected (pending since 2025-02-25T15:32:48.101Z). === 2025-05-26 === * 23:57 "seyedali220" was rejected (pending since 2025-02-24T23:56:17.621Z). === 2025-05-21 === * 11:12 [[gitlab:guilherme|@guilherme]] was approved. === 2025-05-19 === * 13:24 [[gitlab:emojiwiki|@emojiwiki]] was approved. === 2025-05-18 === * 00:00 "xidme" was rejected (pending since 2025-02-15T23:58:56.796Z). === 2025-05-17 === * 02:39 "kdh8219" was rejected (pending since 2025-02-15T02:36:32.237Z). === 2025-05-16 === * 15:09 [[gitlab:maxbinderwmf|@maxbinderwmf]] was approved. === 2025-05-15 === * 04:30 "inspectorzer0" was rejected (pending since 2025-02-13T04:27:33.179Z). === 2025-05-14 === * 17:42 [[gitlab:llugo|@llugo]] was approved. === 2025-05-13 === * 20:18 "mmta" was rejected (pending since 2025-02-11T20:17:23.407Z). === 2025-05-11 === * 20:51 "jad" was rejected (pending since 2025-02-09T20:49:07.333Z). * 17:54 "nishchalsundan" was rejected (pending since 2025-02-09T17:52:25.761Z). * 16:39 "mohammed_abukhadra" was rejected (pending since 2025-02-09T16:39:03.730Z). === 2025-05-09 === * 09:12 [[gitlab:sirchanmp|@sirchanmp]] was approved. === 2025-05-08 === * 08:18 [[gitlab:mengeditch|@mengeditch]] was approved. === 2025-05-07 === * 03:45 "xluffy" was rejected (pending since 2025-02-05T03:45:14.181Z). === 2025-05-06 === * 16:54 "punhaniabhishek" was rejected (pending since 2025-02-04T16:53:50.758Z). * 09:36 [[gitlab:bmartinezcalvo|@bmartinezcalvo]] was approved. === 2025-05-02 === * 12:24 [[gitlab:tohaomg|@tohaomg]] was approved. * 11:48 [[gitlab:mavrikant|@mavrikant]] was approved. * 11:45 [[gitlab:daanvr|@daanvr]] was approved. === 2025-05-01 === * 09:09 "mjoerg" was rejected (pending since 2025-01-30T09:09:04.204Z). === 2025-04-30 === * 23:06 "sanskardubey" was rejected (pending since 2025-01-29T23:03:25.489Z). === 2025-04-29 === * 16:00 "geyslein" was rejected (pending since 2025-01-28T16:00:01.510Z). === 2025-04-26 === * 09:30 "anjali9027" was rejected (pending since 2025-01-25T09:28:07.064Z). === 2025-04-25 === * 18:00 "salahhazaa" was rejected (pending since 2025-01-24T17:58:30.030Z). * 15:15 [[gitlab:yiming|@yiming]] was approved. * 02:06 "mrchanmp" was rejected (pending since 2025-01-24T02:03:58.308Z). === 2025-04-23 === * 17:03 "rj2904" was rejected (pending since 2025-01-22T17:03:11.207Z). * 14:21 "nischay33" was rejected (pending since 2025-01-22T14:19:21.081Z). === 2025-04-22 === * 19:27 "dj80" was rejected (pending since 2025-01-21T19:25:28.498Z). * 14:30 [[gitlab:kaimamin|@kaimamin]] was approved. * 09:57 "debo" was rejected (pending since 2025-01-21T09:54:47.955Z). === 2025-04-21 === * 12:24 "unshell" was rejected (pending since 2025-01-20T12:21:59.686Z). === 2025-04-18 === * 15:06 [[gitlab:spartanarbinger|@spartanarbinger]] was approved. === 2025-04-16 === * 03:09 "dewey" was rejected (pending since 2025-01-15T03:06:17.488Z). === 2025-04-15 === * 19:45 "emdadul" was rejected (pending since 2025-01-14T19:42:29.285Z). === 2025-04-14 === * 06:45 [[gitlab:bcampbell804|@bcampbell804]] was approved. === 2025-04-11 === * 06:27 [[gitlab:jvanderhoop|@jvanderhoop]] was approved. === 2025-04-10 === * 04:12 "bhai420" was rejected (pending since 2025-01-09T04:10:29.430Z). === 2025-04-09 === * 05:03 "austinvarshney" was rejected (pending since 2025-01-08T05:02:34.175Z). === 2025-04-06 === * 15:36 [[gitlab:elph|@elph]] was approved. === 2025-04-02 === * 10:33 [[gitlab:ozge|@ozge]] was approved. === 2025-03-31 === * 20:15 "demandkey" was rejected (pending since 2024-12-30T20:14:23.096Z). * 15:18 [[gitlab:danyya|@danyya]] was approved. === 2025-03-28 === * 15:54 [[gitlab:rutsavi09|@rutsavi09]] was approved. * 15:54 [[gitlab:ilanen1|@ilanen1]] was approved. === 2025-03-25 === * 19:27 [[gitlab:irfo|@irfo]] was approved. * 11:54 [[gitlab:kmontalva-wmf|@kmontalva-wmf]] was approved. * 04:33 [[gitlab:paul26|@paul26]] was approved. * 04:18 "as1100k" was rejected (pending since 2024-12-24T04:18:06.813Z). === 2025-03-24 === * 11:33 "amzadkhankk" was rejected (pending since 2024-12-23T11:33:14.176Z). === 2025-03-23 === * 12:24 "wolfdo" was rejected (pending since 2024-12-22T12:23:35.056Z). === 2025-03-22 === * 09:45 [[gitlab:fjmustak|@fjmustak]] was approved. === 2025-03-20 === * 18:42 "sathishkokila" was rejected (pending since 2024-12-19T18:39:35.161Z). * 17:03 [[gitlab:alien4444|@alien4444]] was approved. * 15:27 [[gitlab:davidcoronel|@davidcoronel]] was approved. === 2025-03-19 === * 22:57 [[gitlab:r1f4t|@r1f4t]] was approved. * 19:03 "daniel24ps" was rejected (pending since 2024-12-18T19:00:21.249Z). * 14:18 [[gitlab:beepbooppenguin|@beepbooppenguin]] was approved. === 2025-03-18 === * 17:48 "rahulkundu1209" was rejected (pending since 2024-12-17T17:46:41.936Z). * 08:15 "kirtisikka972" was rejected (pending since 2024-12-17T08:13:25.487Z). === 2025-03-15 === * 13:30 "tulspal_sidhu" was rejected (pending since 2024-12-14T13:29:10.606Z). * 01:39 "peacedeadc" was rejected (pending since 2024-12-14T01:37:36.579Z). === 2025-03-14 === * 03:51 [[gitlab:chuckthebuck|@chuckthebuck]] was approved. * 02:33 "yxngtrtxll" was rejected (pending since 2024-12-13T02:31:51.658Z). === 2025-03-13 === * 14:36 [[gitlab:iccander|@iccander]] was approved. === 2025-03-12 === * 23:21 "jokerchic36" was rejected (pending since 2024-12-11T23:21:00.670Z). * 15:30 [[gitlab:naomi|@naomi]] was approved. * 15:27 [[gitlab:cobi|@cobi]] was approved. === 2025-03-11 === * 12:42 "mohitvermaxx" was rejected (pending since 2024-12-10T12:40:56.967Z). === 2025-03-10 === * 16:51 [[gitlab:nanona15dobato|@nanona15dobato]] was approved. === 2025-03-09 === * 22:39 [[gitlab:jonkolbert|@jonkolbert]] was approved. * 20:45 [[gitlab:urbanecmtest2|@urbanecmtest2]] was approved. === 2025-03-07 === * 16:54 [[gitlab:hswan|@hswan]] was approved. * 14:42 [[gitlab:atitkov|@atitkov]] was approved. * 00:42 [[gitlab:infrastruktur|@infrastruktur]] was approved. === 2025-03-06 === * 17:21 "johnmann" was rejected (pending since 2024-12-05T17:19:24.995Z). === 2025-03-05 === * 07:33 [[gitlab:monx9494|@monx9494]] was approved. === 2025-03-02 === * 21:21 "paul26" was rejected (pending since 2024-12-01T21:20:19.681Z). === 2025-03-01 === * 19:15 [[gitlab:izno|@izno]] was approved. * 12:45 [[gitlab:nyerho|@nyerho]] was approved. === 2025-02-28 === * 18:27 [[gitlab:chuckonwumelu|@chuckonwumelu]] was approved. * 13:09 "ashwinpraveengo" was rejected (pending since 2024-11-29T13:07:47.240Z). * 00:18 "eduardoaugusto" was rejected (pending since 2024-11-29T00:17:43.372Z). === 2025-02-27 === * 20:39 "volkanurl" was rejected (pending since 2024-11-28T20:37:18.101Z). === 2025-02-24 === * 21:15 [[gitlab:feeglgeef|@feeglgeef]] was approved. * 20:18 [[gitlab:piaanalysis2|@piaanalysis2]] was approved. * 19:06 [[gitlab:dhardy|@dhardy]] was approved. === 2025-02-22 === * 19:27 [[gitlab:owuh|@owuh]] was approved. === 2025-02-19 === * 16:06 [[gitlab:artemkloko|@artemkloko]] was approved. * 13:03 [[gitlab:jgafnea|@jgafnea]] was approved. === 2025-02-17 === * 16:33 [[gitlab:asmartkitten|@asmartkitten]] was approved. === 2025-02-16 === * 19:12 "gaurigupta21" was rejected (pending since 2024-11-17T19:11:07.416Z). === 2025-02-15 === * 01:18 [[gitlab:mediawiki-quickstart-ci|@mediawiki-quickstart-ci]] was approved. === 2025-02-14 === * 15:21 "nathanbnm" was rejected (pending since 2024-11-15T15:18:19.632Z). === 2025-02-13 === * 16:45 [[gitlab:priyanshuchahal|@priyanshuchahal]] was approved. * 16:42 [[gitlab:ajhalili2006|@ajhalili2006]] was approved. === 2025-02-12 === * 23:21 "monkeypatch999" was rejected (pending since 2024-11-13T23:20:38.398Z). * 06:36 [[gitlab:jainlakshita28|@jainlakshita28]] was approved. === 2025-02-11 === * 19:27 [[gitlab:matthewsm2|@matthewsm2]] was approved. === 2025-02-09 === * 16:15 "mohammed_abukhadra" was rejected (pending since 2024-11-10T16:15:18.361Z). === 2025-02-07 === * 21:33 "brennan" was rejected (pending since 2024-11-08T21:31:07.351Z). === 2025-02-06 === * 08:24 "mmta" was rejected (pending since 2024-11-07T08:22:36.724Z). * 06:21 [[gitlab:bunnypranav|@bunnypranav]] was approved. === 2025-02-05 === * 22:39 "chrissteinchen" was rejected (pending since 2024-11-06T22:38:16.673Z). === 2025-02-03 === * 07:45 "edriiic" was rejected (pending since 2024-11-04T07:44:46.849Z). * 01:12 "geppy" was rejected (pending since 2024-11-04T01:10:48.710Z). === 2025-02-02 === * 13:18 "funa-enpitu" was rejected (pending since 2024-11-03T13:15:46.065Z). === 2025-01-31 === * 23:42 "nfontes" was rejected (pending since 2024-11-01T23:39:41.755Z). * 22:51 "sbronson" was rejected (pending since 2024-11-01T22:50:31.871Z). * 00:42 [[gitlab:farid|@farid]] was approved. === 2025-01-27 === * 08:15 [[gitlab:eliza189|@eliza189]] was approved. === 2025-01-25 === * 09:51 [[gitlab:pamputt|@pamputt]] was approved. === 2025-01-23 === * 14:30 [[gitlab:lubianat|@lubianat]] was approved. * 11:45 [[gitlab:bootsa|@bootsa]] was approved. === 2025-01-21 === * 05:09 "niko" was rejected (pending since 2024-07-21T16:10:01.377Z). * 05:09 "thawizkid369777" was rejected (pending since 2024-07-18T17:42:44.493Z). * 05:09 "sarthaksingh2" was rejected (pending since 2024-07-10T11:31:30.470Z). * 05:09 "shriyakt" was rejected (pending since 2024-07-06T04:54:10.248Z). * 05:09 "akshaya" was rejected (pending since 2024-07-06T04:04:51.488Z). * 05:09 "alaka03aj" was rejected (pending since 2024-07-05T18:01:54.876Z). * 05:09 "sulochanaviji-5049" was rejected (pending since 2024-07-01T05:58:00.427Z). * 05:09 "nayanjnath" was rejected (pending since 2024-07-01T02:51:57.405Z). * 05:09 "sd44" was rejected (pending since 2024-06-30T04:28:51.436Z). * 05:09 "metavalent" was rejected (pending since 2024-06-29T01:37:14.210Z). * 05:09 "wicloudx" was rejected (pending since 2024-06-28T11:51:23.335Z). * 05:09 "debo" was rejected (pending since 2024-06-28T01:44:59.845Z). * 05:09 "bwiki" was rejected (pending since 2024-06-23T14:15:38.032Z). * 05:09 "toprak" was rejected (pending since 2024-06-23T11:35:50.819Z). * 05:09 "iristeller" was rejected (pending since 2024-06-14T20:53:48.959Z). * 05:09 "jcolvin" was rejected (pending since 2024-06-12T17:29:01.238Z). * 05:09 "kalyan" was rejected (pending since 2024-06-07T07:52:46.993Z). * 05:09 "bluecrystal" was rejected (pending since 2024-06-06T19:16:20.107Z). * 05:09 "iftttrohit" was rejected (pending since 2024-06-04T12:08:50.818Z). * 05:09 "pogpotato" was rejected (pending since 2024-06-03T17:58:21.684Z). * 05:09 "cptlausebaer" was rejected (pending since 2024-05-31T18:53:27.692Z). * 05:09 "hdevine825" was rejected (pending since 2024-05-31T17:04:18.279Z). * 05:09 "anaghaa18" was rejected (pending since 2024-05-25T19:14:31.803Z). * 05:09 "atharvanair04" was rejected (pending since 2024-05-25T14:24:52.825Z). * 05:09 "anasvemmully" was rejected (pending since 2024-05-25T06:10:27.261Z). * 05:09 "abhinavmohandas" was rejected (pending since 2024-05-25T06:05:24.825Z). * 05:09 "kksurendran06" was rejected (pending since 2024-05-25T06:04:38.082Z). * 05:09 "albertmarshall8896" was rejected (pending since 2024-05-23T09:32:05.462Z). * 05:09 "akellison" was rejected (pending since 2024-05-17T02:07:24.229Z). * 05:09 "mainowill" was rejected (pending since 2024-04-16T23:30:33.881Z). * 05:09 "bzhqc" was rejected (pending since 2024-04-16T19:50:38.676Z). * 05:09 "safan41" was rejected (pending since 2024-04-16T03:34:48.942Z). * 05:09 "mgagat" was rejected (pending since 2024-04-16T03:21:51.764Z). * 05:09 "okeamah" was rejected (pending since 2024-04-16T02:49:00.143Z). * 05:09 "xuhao61" was rejected (pending since 2024-04-15T23:45:09.083Z). * 04:47 "cybel" was rejected (pending since 2024-04-15T06:46:35.791Z). === 2025-01-20 === * 14:33 [[gitlab:your1|@your1]] was approved. === 2025-01-18 === * 10:09 [[gitlab:galrach600|@galrach600]] was approved. * 02:51 [[gitlab:blankeclair|@blankeclair]] was approved. === 2025-01-17 === * 13:57 [[gitlab:dsantamaria|@dsantamaria]] was approved. === 2025-01-15 === * 17:12 [[gitlab:smartse|@smartse]] was approved. === 2025-01-14 === * 17:03 [[gitlab:naorleizer|@naorleizer]] was approved. === 2025-01-13 === * 02:45 [[gitlab:wolf20482|@wolf20482]] was approved. === 2025-01-12 === * 17:45 [[gitlab:tamzin|@tamzin]] was approved. === 2025-01-11 === * 15:24 [[gitlab:bargioni|@bargioni]] was approved. * 14:30 [[gitlab:salelya|@salelya]] was approved. * 10:15 [[gitlab:malakatshy|@malakatshy]] was approved. * 05:21 [[gitlab:newmcpee|@newmcpee]] was approved. === 2025-01-09 === * 15:30 [[gitlab:gkyziridis|@gkyziridis]] was approved. === 2025-01-08 === * 16:21 [[gitlab:ukrface|@ukrface]] was approved. === 2024-12-28 === * 03:27 [[gitlab:twonum|@twonum]] was approved. === 2024-12-25 === * 06:09 [[gitlab:harsv567|@harsv567]] was approved. === 2024-12-21 === * 11:24 [[gitlab:amutha2002|@amutha2002]] was approved. === 2024-12-20 === * 19:51 [[gitlab:hridyeshgupta|@hridyeshgupta]] was approved. * 10:00 [[gitlab:ro-shines|@ro-shines]] was approved. * 08:09 [[gitlab:kesharwaniarpita|@kesharwaniarpita]] was approved. === 2024-12-18 === * 14:45 [[gitlab:soylacarli|@soylacarli]] was approved. === 2024-12-16 === * 20:33 [[gitlab:aleyasiddika1|@aleyasiddika1]] was approved. === 2024-12-15 === * 07:33 [[gitlab:abhishek02bhardwaj|@abhishek02bhardwaj]] was approved. === 2024-12-13 === * 13:18 [[gitlab:ashmitabathre204|@ashmitabathre204]] was approved. === 2024-12-10 === * 06:39 [[gitlab:ginaan|@ginaan]] was approved. === 2024-12-09 === * 05:45 [[gitlab:kallinavya|@kallinavya]] was approved. * 00:54 [[gitlab:viserion-7|@viserion-7]] was approved. === 2024-12-08 === * 17:27 [[gitlab:wargo|@wargo]] was approved. === 2024-12-05 === * 11:15 [[gitlab:ranjithraj|@ranjithraj]] was approved. === 2024-12-02 === * 21:21 [[gitlab:a930913|@a930913]] was approved. === 2024-12-01 === * 02:39 [[gitlab:kingchristlike1|@kingchristlike1]] was approved. === 2024-11-21 === * 13:45 [[gitlab:sascha|@sascha]] was approved. === 2024-11-19 === * 16:36 [[gitlab:jly|@jly]] was approved. === 2024-11-15 === * 02:54 [[gitlab:danielyepezgarces|@danielyepezgarces]] was approved. === 2024-11-14 === * 14:15 [[gitlab:stimoroll|@stimoroll]] was approved. === 2024-11-09 === * 17:15 [[gitlab:f4udeveloper|@f4udeveloper]] was approved. === 2024-11-07 === * 19:15 [[gitlab:zulf|@zulf]] was approved. * 05:33 [[gitlab:hassanamin|@hassanamin]] was approved. === 2024-11-06 === * 19:39 [[gitlab:daniuu|@daniuu]] was approved. * 00:18 [[gitlab:rlopez-wmf|@rlopez-wmf]] was approved. === 2024-10-09 === * 14:45 [[gitlab:jtweed|@jtweed]] was approved. * 10:24 [[gitlab:ifrahkh|@ifrahkh]] was approved. * 09:06 [[gitlab:wikibayer|@wikibayer]] was approved. === 2024-10-06 === * 10:27 [[gitlab:keerthan16|@keerthan16]] was approved. === 2024-10-04 === * 07:45 [[gitlab:hakimi97|@hakimi97]] was approved. === 2024-09-30 === * 07:39 [[gitlab:ninjastrikers|@ninjastrikers]] was approved. === 2024-09-28 === * 17:30 [[gitlab:webrunner95|@webrunner95]] was approved. === 2024-09-18 === * 21:39 [[gitlab:elliottetzkorn|@elliottetzkorn]] was approved. === 2024-09-14 === * 22:06 [[gitlab:humptydumpty|@humptydumpty]] was approved. === 2024-09-06 === * 08:48 [[gitlab:mickabarber|@mickabarber]] was approved. === 2024-08-27 === * 17:36 [[gitlab:edgars|@edgars]] was approved. === 2024-08-22 === * 09:18 [[gitlab:antonkokhwmde|@antonkokhwmde]] was approved. === 2024-08-14 === * 19:21 [[gitlab:jfk|@jfk]] was approved. === 2024-08-13 === * 17:57 [[gitlab:daxserver|@daxserver]] was approved. === 2024-08-11 === * 09:57 [[gitlab:pauliesnug|@pauliesnug]] was approved. === 2024-08-10 === * 08:42 [[gitlab:ashig|@ashig]] was approved. === 2024-08-09 === * 14:09 [[gitlab:masssly|@masssly]] was approved. === 2024-08-05 === * 22:15 [[gitlab:mrtortue|@mrtortue]] was approved. === 2024-08-02 === * 16:21 [[gitlab:dsantini|@dsantini]] was approved. === 2024-07-31 === * 11:54 [[gitlab:cptviraj|@cptviraj]] was approved. === 2024-07-30 === * 19:09 [[gitlab:iniquity|@iniquity]] was approved. * 10:00 [[gitlab:collins|@collins]] was approved. === 2024-07-27 === * 15:57 [[gitlab:songnguxyz|@songnguxyz]] was approved. === 2024-07-25 === * 12:36 [[gitlab:mszabo|@mszabo]] was approved. * 09:21 [[gitlab:agarwalmahima|@agarwalmahima]] was approved. === 2024-07-24 === * 08:05 [[gitlab:dragoniez|@dragoniez]] was approved. === 2024-07-23 === * 06:54 [[gitlab:mirji|@mirji]] was approved. === 2024-07-16 === * 10:00 [[gitlab:lakejason0|@lakejason0]] was approved. === 2024-07-12 === * 11:33 [[gitlab:cn|@cn]] was approved. * 08:12 [[gitlab:unchampignon|@unchampignon]] was approved. === 2024-07-07 === * 17:12 [[gitlab:agamyasamuel|@agamyasamuel]] was approved. * 05:24 [[gitlab:kuldeepburjbhalaike|@kuldeepburjbhalaike]] was approved. === 2024-07-06 === * 11:18 [[gitlab:dibya|@dibya]] was approved. * 04:54 [[gitlab:sarthakparashar|@sarthakparashar]] was approved. === 2024-07-05 === * 18:15 [[gitlab:vanshikarathi|@vanshikarathi]] was approved. === 2024-07-02 === * 19:00 [[gitlab:ebrahim|@ebrahim]] was approved. === 2024-07-01 === * 20:12 [[gitlab:rockingpenny4|@rockingpenny4]] was approved. * 18:15 [[gitlab:balajijagadesh|@balajijagadesh]] was approved. === 2024-06-30 === * 18:24 [[gitlab:hrideshmg|@hrideshmg]] was approved. * 07:18 [[gitlab:chanakyakumardas|@chanakyakumardas]] was approved. * 06:30 [[gitlab:rihaan180|@rihaan180]] was approved. === 2024-06-27 === * 17:36 [[gitlab:driedmueller|@driedmueller]] was approved. === 2024-06-19 === * 12:57 [[gitlab:audreypenven|@audreypenven]] was approved. === 2024-06-16 === * 01:18 [[gitlab:roysmith|@roysmith]] was approved. === 2024-06-08 === * 02:45 [[gitlab:jleedev|@jleedev]] was approved. === 2024-06-03 === * 13:57 [[gitlab:afeder|@afeder]] was approved. === 2024-06-01 === * 10:54 [[gitlab:florianschmitt|@florianschmitt]] was approved. === 2024-05-30 === * 16:42 [[gitlab:krlsca|@krlsca]] was approved. === 2024-05-28 === * 11:24 [[gitlab:rickijay|@rickijay]] was approved. === 2024-05-26 === * 11:18 [[gitlab:ranjithsiji|@ranjithsiji]] was approved. === 2024-05-25 === * 07:24 [[gitlab:jony|@jony]] was approved. === 2024-05-23 === * 08:45 [[gitlab:lepticed7|@lepticed7]] was approved. === 2024-05-22 === * 20:42 [[gitlab:echecs|@echecs]] was approved. === 2024-05-21 === * 13:33 [[gitlab:mbs|@mbs]] was approved. === 2024-05-19 === * 18:06 [[gitlab:ionenlaser|@ionenlaser]] was approved. === 2024-05-18 === * 23:36 [[gitlab:mdaniels5757|@mdaniels5757]] was approved. === 2024-05-17 === * 08:54 [[gitlab:grapedog|@grapedog]] was approved. === 2024-05-08 === * 19:42 [[gitlab:kelhurd|@kelhurd]] was approved. * 19:06 [[gitlab:khurd|@khurd]] was approved. === 2024-05-06 === * 19:48 [[gitlab:j3j5|@j3j5]] was approved. * 12:06 [[gitlab:tk-999|@tk-999]] was approved. === 2024-05-05 === * 22:09 [[gitlab:pppery|@pppery]] was approved. * 20:33 [[gitlab:sakretsu|@sakretsu]] was approved. * 12:12 [[gitlab:waterquark|@waterquark]] was approved. === 2024-05-04 === * 09:03 [[gitlab:multichill|@multichill]] was approved. * 07:42 [[gitlab:abaris|@abaris]] was approved. === 2024-05-03 === * 14:57 [[gitlab:maurusian|@maurusian]] was approved. === 2024-04-24 === * 05:48 [[gitlab:wolfinux|@wolfinux]] was approved. === 2024-04-23 === * 15:48 [[gitlab:dreamrimmer|@dreamrimmer]] was approved. === 2024-04-21 === * 06:51 [[gitlab:alon|@alon]] was approved. === 2024-04-17 === * 23:33 [[gitlab:derenrich|@derenrich]] was approved. === 2024-04-16 === * 17:18 [[gitlab:valcio|@valcio]] was approved. === 2024-04-14 === * 16:51 [[gitlab:wikilucas00|@wikilucas00]] was approved. === 2024-04-06 === * 12:48 [[gitlab:theprotonade|@theprotonade]] was approved. === 2024-04-02 === * 07:30 [[gitlab:bohuizhang|@bohuizhang]] was approved. === 2024-03-30 === * 13:36 [[gitlab:lpintscher|@lpintscher]] was approved. === 2024-03-26 === * 17:09 [[gitlab:eenabulele|@eenabulele]] was approved. === 2024-03-25 === * 14:27 [[gitlab:tuukka|@tuukka]] was approved. === 2024-03-24 === * 12:24 [[gitlab:firefly|@firefly]] was approved. === 2024-03-21 === * 19:33 [[gitlab:universal-omega|@universal-omega]] was approved. === 2024-03-17 === * 10:36 [[gitlab:bisel91|@bisel91]] was approved. === 2024-03-16 === * 10:09 [[gitlab:delord|@delord]] was approved. * 00:42 [[gitlab:athulvis1|@athulvis1]] was approved. === 2024-03-15 === * 19:06 [[gitlab:ignaciorodrguez|@ignaciorodrguez]] was approved. * 08:30 [[gitlab:peachey88|@peachey88]] was approved. * 06:51 [[gitlab:derick|@derick]] was approved. === 2024-03-12 === * 15:06 [[gitlab:xiaoxiao|@xiaoxiao]] was approved. === 2024-03-06 === * 13:21 [[gitlab:desianabae1|@desianabae1]] was approved. === 2024-03-05 === * 19:21 [[gitlab:ep1c|@ep1c]] was approved. * 16:33 [[gitlab:jasmine|@jasmine]] was approved. === 2024-03-02 === * 06:42 [[gitlab:potsdamlamb|@potsdamlamb]] was approved. === 2024-02-29 === * 23:18 [[gitlab:arandomname123|@arandomname123]] was approved. * 18:03 [[gitlab:baba|@baba]] was approved. * 17:48 [[gitlab:yfdyh000|@yfdyh000]] was approved. * 03:09 [[gitlab:sds|@sds]] was approved. === 2024-02-27 === * 23:33 [[gitlab:lofhi|@lofhi]] was approved. === 2024-02-15 === * 19:45 [[gitlab:gergesshamon|@gergesshamon]] was approved. === 2024-02-14 === * 14:33 [[gitlab:philipnelson99|@philipnelson99]] was approved. === 2024-02-13 === * 13:06 [[gitlab:dringsim|@dringsim]] was approved. === 2024-02-12 === * 17:36 [[gitlab:haak|@haak]] was approved. === 2024-02-05 === * 17:33 [[gitlab:qwerfjkl|@qwerfjkl]] was approved. * 17:14 [[gitlab:ahecht|@ahecht]] was approved. === 2024-02-01 === * 09:27 [[gitlab:arinaigum|@arinaigum]] was approved. * 00:15 [[gitlab:jas42|@jas42]] was approved. * 00:15 [[gitlab:edhu|@edhu]] was approved. * 00:15 [[gitlab:marnanel|@marnanel]] was approved. * 00:15 [[gitlab:ibrahemqasim|@ibrahemqasim]] was approved. * 00:15 [[gitlab:amasotti|@amasotti]] was approved. * 00:15 [[gitlab:deni|@deni]] was approved. * 00:15 [[gitlab:cyber|@cyber]] was approved. * 00:15 [[gitlab:saroj|@saroj]] was approved. === 2024-01-29 === * 21:42 [[gitlab:rgupta|@rgupta]] was approved. === 2024-01-07 === * 09:48 [[gitlab:lutrome|@lutrome]] was approved. === 2024-01-05 === * 20:48 [[gitlab:jinoytommanjaly|@jinoytommanjaly]] was approved. * 02:51 [[gitlab:braunobruno|@braunobruno]] was approved. * 01:08 [[gitlab:amorymeltzer|@amorymeltzer]] was approved. * 01:08 [[gitlab:phi22ipus|@phi22ipus]] was approved. === 2024-01-03 === * 14:45 [[gitlab:gabina|@gabina]] was approved. === 2024-01-02 === * 13:18 [[gitlab:arthurtaylor|@arthurtaylor]] was approved. === 2023-12-23 === * 00:33 [[gitlab:aram|@aram]] was approved. === 2023-12-22 === * 16:24 [[gitlab:elpitareio|@elpitareio]] was approved. === 2023-12-21 === * 00:43 [[gitlab:bsadowski1|@bsadowski1]] was approved. * 00:43 [[gitlab:ederporto|@ederporto]] was approved. * 00:43 [[gitlab:sadraiiali|@sadraiiali]] was approved. * 00:43 [[gitlab:wasp-outis|@wasp-outis]] was approved. * 00:43 [[gitlab:bodhisattwa|@bodhisattwa]] was approved. * 00:43 [[gitlab:air7538|@air7538]] was approved. * 00:43 [[gitlab:anzx|@anzx]] was approved. * 00:43 [[gitlab:tekask1903|@tekask1903]] was approved. * 00:42 [[gitlab:kiwi-0x010c|@kiwi-0x010c]] was approved. * 00:42 [[gitlab:mpaa|@mpaa]] was approved. * 00:42 [[gitlab:kutay|@kutay]] was approved. * 00:42 [[gitlab:wattmto|@wattmto]] was approved. dwyvq3by9yg1ni85st2a6rd03qde9wp 2450671 2450647 2026-08-23T01:54:14Z Gitlabaccountapprovalbot 37332 westcodragen07 was rejected. 2450671 wikitext text/x-wiki <noinclude>'''Audit log of approvals''' made by [[gitlab:gitlabaccountapprovalbot|@gitlabaccountapprovalbot]]. __NOTOC__</noinclude> === 2026-08-23 === * 01:54 "westcodragen07" was rejected (pending since 2026-05-24T01:53:53.396Z). === 2026-08-22 === * 15:21 [[gitlab:laraibhasan|@laraibhasan]] was approved. === 2026-08-20 === * 17:18 [[gitlab:jatwa|@jatwa]] was approved. === 2026-08-18 === * 11:33 "wladek92" was rejected (pending since 2026-05-19T11:30:32.679Z). === 2026-08-17 === * 18:39 [[gitlab:erlanger16e|@erlanger16e]] was approved. * 18:33 "ashishkaranam" was rejected (pending since 2026-05-18T18:32:18.813Z). * 12:48 [[gitlab:lfs|@lfs]] was approved. === 2026-08-16 === * 15:42 [[gitlab:pcendrer|@pcendrer]] was approved. === 2026-08-15 === * 17:33 [[gitlab:flammablepizza|@flammablepizza]] was approved. === 2026-08-14 === * 13:00 [[gitlab:pacmand|@pacmand]] was approved. === 2026-08-12 === * 13:21 [[gitlab:doublechek|@doublechek]] was approved. === 2026-08-11 === * 14:15 [[gitlab:mmilenkovicwmf|@mmilenkovicwmf]] was approved. === 2026-08-10 === * 17:18 [[gitlab:mustafe|@mustafe]] was approved. * 10:33 "c8263a20" was rejected (pending since 2026-05-11T10:32:40.353Z). === 2026-08-09 === * 21:18 [[gitlab:iamnetx|@iamnetx]] was approved. * 20:09 "yirba" was rejected (pending since 2026-05-10T20:07:07.738Z). * 06:42 "marsam2489" was rejected (pending since 2026-05-10T06:40:46.276Z). === 2026-08-08 === * 08:15 [[gitlab:taiwaniajusto|@taiwaniajusto]] was approved. === 2026-08-07 === * 18:21 [[gitlab:gturkington|@gturkington]] was approved. * 06:36 "brianbybyby" was rejected (pending since 2026-05-08T06:36:08.459Z). === 2026-07-30 === * 20:42 "horaciocolbert" was rejected (pending since 2026-04-30T20:41:40.421Z). === 2026-07-29 === * 10:03 [[gitlab:piastu|@piastu]] was approved. * 06:33 "rafiul1" was rejected (pending since 2026-04-29T06:33:02.573Z). === 2026-07-28 === * 16:51 [[gitlab:for-each-next|@for-each-next]] was approved. * 09:21 [[gitlab:gka|@gka]] was approved. === 2026-07-27 === * 08:12 [[gitlab:cambob|@cambob]] was approved. === 2026-07-26 === * 15:39 "demansanaagmailcom" was rejected (pending since 2026-04-26T15:39:03.297Z). === 2026-07-24 === * 10:54 [[gitlab:ysogo|@ysogo]] was approved. * 09:39 [[gitlab:pankaj199|@pankaj199]] was approved. === 2026-07-23 === * 12:30 [[gitlab:fermiboson|@fermiboson]] was approved. * 09:21 [[gitlab:slashme|@slashme]] was approved. === 2026-07-22 === * 15:54 [[gitlab:lmedley|@lmedley]] was approved. * 13:24 "praveen5638" was rejected (pending since 2026-04-22T13:21:23.368Z). * 12:48 [[gitlab:cyberpower678|@cyberpower678]] was approved. * 08:30 [[gitlab:nabbegat|@nabbegat]] was approved. * 08:30 [[gitlab:plyd|@plyd]] was approved. * 07:06 "ayush8620" was rejected (pending since 2026-04-22T07:03:12.476Z). === 2026-07-21 === * 15:51 [[gitlab:panieravide|@panieravide]] was approved. * 15:33 [[gitlab:luisvilla-personal|@luisvilla-personal]] was approved. * 13:48 [[gitlab:yru|@yru]] was approved. * 13:09 [[gitlab:deevad|@deevad]] was approved. * 12:54 [[gitlab:ctdo17|@ctdo17]] was approved. * 12:36 [[gitlab:jeannenoiraud|@jeannenoiraud]] was approved. * 12:21 [[gitlab:nadiantara|@nadiantara]] was approved. * 12:18 [[gitlab:wijltcher|@wijltcher]] was approved. * 10:33 [[gitlab:nivopol|@nivopol]] was approved. * 10:30 [[gitlab:johlig|@johlig]] was approved. * 10:30 [[gitlab:majicita|@majicita]] was approved. * 10:15 [[gitlab:yongjiapeng|@yongjiapeng]] was approved. * 10:09 [[gitlab:francyskus|@francyskus]] was approved. * 09:24 [[gitlab:xanonymusx|@xanonymusx]] was approved. === 2026-07-20 === * 17:15 "leonidlednev" was rejected (pending since 2026-04-20T17:13:35.108Z). * 15:27 [[gitlab:rodrigoargenton|@rodrigoargenton]] was approved. * 05:48 "draftecho" was rejected (pending since 2026-04-20T05:48:06.953Z). === 2026-07-19 === * 12:21 [[gitlab:boivie|@boivie]] was approved. === 2026-07-18 === * 16:09 [[gitlab:pharos|@pharos]] was approved. * 15:45 [[gitlab:priyankar22|@priyankar22]] was approved. * 15:30 [[gitlab:sisyph|@sisyph]] was approved. === 2026-07-13 === * 03:45 [[gitlab:dreamyshade|@dreamyshade]] was approved. === 2026-07-12 === * 09:27 [[gitlab:smk|@smk]] was approved. === 2026-07-11 === * 14:48 "bigcereal42" was rejected (pending since 2026-04-11T14:47:40.321Z). * 12:18 "pratyushsawan" was rejected (pending since 2026-04-11T12:16:41.671Z). === 2026-07-10 === * 14:42 [[gitlab:akaza24|@akaza24]] was approved. * 12:57 [[gitlab:kormisk|@kormisk]] was approved. === 2026-07-07 === * 10:36 [[gitlab:olafjanssen|@olafjanssen]] was approved. * 06:57 "elisapoly-99" was rejected (pending since 2026-04-07T06:55:52.662Z). === 2026-07-06 === * 11:57 "ma3rouf" was rejected (pending since 2026-04-06T11:56:30.978Z). === 2026-07-02 === * 15:48 [[gitlab:tekneos|@tekneos]] was approved. === 2026-07-01 === * 15:12 [[gitlab:mugurolevy|@mugurolevy]] was approved. * 14:15 [[gitlab:vadymts1|@vadymts1]] was approved. * 09:57 "mugurolevy" was rejected (pending since 2026-04-01T09:55:19.175Z). === 2026-06-30 === * 14:27 "shivangisharma" was rejected (pending since 2026-03-31T14:26:44.932Z). === 2026-06-29 === * 19:03 [[gitlab:thisismattmiller|@thisismattmiller]] was approved. === 2026-06-28 === * 14:51 "nkwenuinadine" was rejected (pending since 2026-03-29T14:48:32.735Z). * 14:03 "vaishnavikumbhar" was rejected (pending since 2026-03-29T14:01:30.604Z). * 13:03 "stepmay" was rejected (pending since 2026-03-29T13:01:59.905Z). * 06:51 "swallroth" was rejected (pending since 2026-03-29T06:49:54.838Z). === 2026-06-26 === * 09:39 [[gitlab:lakshita28|@lakshita28]] was approved. * 07:30 [[gitlab:reeti|@reeti]] was approved. * 07:30 [[gitlab:anushka10patel|@anushka10patel]] was approved. * 07:30 "samsaesque" was rejected (pending since 2026-03-27T07:29:57.279Z). * 05:51 [[gitlab:arpithhhaaa|@arpithhhaaa]] was approved. * 05:51 [[gitlab:govindlaltl|@govindlaltl]] was approved. === 2026-06-25 === * 16:09 [[gitlab:sakuraemad|@sakuraemad]] was approved. * 07:00 "kdh8219" was rejected (pending since 2026-03-26T06:58:05.415Z). === 2026-06-24 === * 11:54 [[gitlab:sanskardubeydev|@sanskardubeydev]] was approved. * 10:09 "tanmay789q" was rejected (pending since 2026-03-25T10:07:54.602Z). === 2026-06-22 === * 19:57 [[gitlab:gouvernathor|@gouvernathor]] was approved. * 16:45 [[gitlab:lucasbelo|@lucasbelo]] was approved. * 07:15 "jason2000-cpu" was rejected (pending since 2026-03-23T07:14:09.184Z). === 2026-06-21 === * 13:18 [[gitlab:egonw|@egonw]] was approved. === 2026-06-20 === * 10:21 [[gitlab:tways2017|@tways2017]] was approved. === 2026-06-19 === * 16:06 "wilsonwang2026" was rejected (pending since 2026-03-20T16:06:05.511Z). * 04:12 [[gitlab:claudio|@claudio]] was approved. === 2026-06-18 === * 14:21 "royiswariii" was rejected (pending since 2026-03-19T14:19:16.896Z). * 13:06 [[gitlab:laurabarluzzi|@laurabarluzzi]] was approved. === 2026-06-17 === * 11:24 "adinathq8x" was rejected (pending since 2026-03-18T11:22:50.098Z). * 09:45 "nathanveritas" was rejected (pending since 2026-03-18T09:43:51.645Z). === 2026-06-15 === * 22:39 [[gitlab:mohammadhijjawi|@mohammadhijjawi]] was approved. * 14:24 "enlisar" was rejected (pending since 2026-03-16T14:23:00.109Z). * 14:06 "ayaan" was rejected (pending since 2026-03-16T14:03:31.071Z). * 10:54 "kwametech" was rejected (pending since 2026-03-16T10:54:11.083Z). === 2026-06-14 === * 17:45 [[gitlab:surajseth520|@surajseth520]] was approved. * 07:24 "malahimhaseeb" was rejected (pending since 2026-03-15T07:21:57.748Z). === 2026-06-11 === * 11:48 [[gitlab:cadddr|@cadddr]] was approved. * 11:18 "wikipiggy" was rejected (pending since 2026-03-12T11:16:09.335Z). * 07:15 [[gitlab:vesihiisi|@vesihiisi]] was approved. === 2026-06-10 === * 07:03 [[gitlab:dmiranda|@dmiranda]] was approved. === 2026-06-09 === * 14:21 [[gitlab:linkgenetic|@linkgenetic]] was approved. * 14:03 [[gitlab:sjones-ctr|@sjones-ctr]] was approved. * 12:51 [[gitlab:ekrem|@ekrem]] was approved. === 2026-06-08 === * 17:48 "jmprax" was rejected (pending since 2026-03-09T17:46:38.807Z). * 16:15 [[gitlab:ahonc|@ahonc]] was approved. * 12:57 [[gitlab:rainmonger|@rainmonger]] was approved. === 2026-06-07 === * 23:03 "shadowthewuff" was rejected (pending since 2026-03-08T23:00:53.442Z). * 11:45 "wiki-pavan" was rejected (pending since 2026-03-08T11:45:11.116Z). * 02:39 [[gitlab:launchpad|@launchpad]] was approved. === 2026-06-06 === * 14:54 "unicord" was rejected (pending since 2026-03-07T14:52:04.992Z). * 12:48 "chien" was rejected (pending since 2026-03-07T12:48:11.669Z). === 2026-06-04 === * 14:33 "only-vikas" was rejected (pending since 2026-03-05T14:32:09.186Z). === 2026-06-03 === * 15:00 [[gitlab:anafibnshahibul|@anafibnshahibul]] was approved. === 2026-06-02 === * 21:21 "mgagat" was rejected (pending since 2026-03-03T21:18:37.223Z). * 13:57 "prasunaenumarthy" was rejected (pending since 2026-03-03T13:57:14.847Z). * 05:48 [[gitlab:tmoney|@tmoney]] was approved. === 2026-06-01 === * 14:57 "vikram2101" was rejected (pending since 2026-03-02T14:54:26.550Z). * 12:03 "watshell" was rejected (pending since 2026-03-02T12:03:09.329Z). === 2026-05-29 === * 12:48 "mounikapotladurthi" was rejected (pending since 2026-02-27T12:45:38.609Z). === 2026-05-27 === * 20:00 "vinitha" was rejected (pending since 2026-02-25T19:58:43.524Z). * 16:30 "codeurluce" was rejected (pending since 2026-02-25T16:28:53.973Z). * 14:33 [[gitlab:thilio|@thilio]] was approved. === 2026-05-26 === * 12:09 "charisad" was rejected (pending since 2026-02-24T12:07:21.881Z). === 2026-05-25 === * 22:54 "ddshelto" was rejected (pending since 2026-02-23T22:52:44.427Z). * 19:51 "lakz-99" was rejected (pending since 2026-02-23T19:47:00.263Z). * 19:48 "lakz-99" was rejected (pending since 2026-02-23T19:47:00.263Z). === 2026-05-24 === * 18:45 "jiyagupta-cs" was rejected (pending since 2026-02-22T18:43:33.176Z). === 2026-05-23 === * 13:09 [[gitlab:gauthammohanraj|@gauthammohanraj]] was approved. * 04:21 [[gitlab:staraction|@staraction]] was approved. === 2026-05-22 === * 19:03 "i-horich" was rejected (pending since 2026-02-20T19:00:43.519Z). * 01:48 "50323233" was rejected (pending since 2026-02-20T01:48:05.555Z). === 2026-05-21 === * 18:51 "kartikeyg0104" was rejected (pending since 2026-02-19T18:48:39.707Z). * 16:27 [[gitlab:renovatebot|@renovatebot]] was approved. * 16:06 [[gitlab:gkm563|@gkm563]] was approved. === 2026-05-20 === * 01:21 "beedellrokejulianlockhart" was rejected (pending since 2026-02-18T01:19:13.284Z). === 2026-05-18 === * 23:18 "wladek92" was rejected (pending since 2026-02-16T23:16:22.939Z). * 16:36 [[gitlab:effeietsanders|@effeietsanders]] was approved. === 2026-05-14 === * 21:00 [[gitlab:nehemienathan|@nehemienathan]] was approved. === 2026-05-13 === * 10:51 "ssssaaaa" was rejected (pending since 2026-02-11T10:50:36.975Z). === 2026-05-12 === * 18:06 [[gitlab:psubhashish|@psubhashish]] was approved. * 08:12 "khan" was rejected (pending since 2026-02-10T08:11:48.776Z). * 04:27 "galaxysh" was rejected (pending since 2026-02-10T04:24:59.440Z). === 2026-05-11 === * 12:18 "peterxy12" was rejected (pending since 2026-02-09T12:18:01.982Z). === 2026-05-10 === * 11:09 "yalihupokn" was rejected (pending since 2026-02-08T11:06:51.336Z). * 05:12 "wobadha" was rejected (pending since 2026-02-08T05:11:00.569Z). === 2026-05-09 === * 13:45 "bwiki" was rejected (pending since 2026-02-07T13:43:38.177Z). === 2026-05-08 === * 09:24 [[gitlab:cwilliams|@cwilliams]] was approved. === 2026-05-07 === * 14:15 "rehankhan78" was rejected (pending since 2026-02-05T14:13:37.754Z). === 2026-05-06 === * 11:24 "ari" was rejected (pending since 2026-02-04T11:24:11.760Z). * 08:09 [[gitlab:neriah|@neriah]] was approved. * 06:27 [[gitlab:status401|@status401]] was approved. === 2026-05-03 === * 09:54 [[gitlab:anilk|@anilk]] was approved. === 2026-05-02 === * 17:54 [[gitlab:sweil|@sweil]] was approved. * 17:00 [[gitlab:aoppo|@aoppo]] was approved. === 2026-05-01 === * 21:18 [[gitlab:dawalda|@dawalda]] was approved. === 2026-04-30 === * 21:42 "merohibine" was rejected (pending since 2026-01-29T21:40:00.756Z). * 20:54 [[gitlab:tfmorris|@tfmorris]] was approved. * 17:33 [[gitlab:uyen|@uyen]] was approved. * 07:39 [[gitlab:mahveotm|@mahveotm]] was approved. * 06:36 [[gitlab:leo321|@leo321]] was approved. === 2026-04-29 === * 02:27 [[gitlab:dw31415|@dw31415]] was approved. === 2026-04-28 === * 23:09 [[gitlab:dtorsani|@dtorsani]] was approved. === 2026-04-27 === * 23:42 [[gitlab:quinlan|@quinlan]] was approved. * 05:00 [[gitlab:matthewyeager|@matthewyeager]] was approved. === 2026-04-26 === * 17:36 "kuba-hajnej" was rejected (pending since 2026-01-25T17:33:32.467Z). * 13:03 "jklamo" was rejected (pending since 2026-01-25T13:02:22.936Z). === 2026-04-25 === * 20:24 [[gitlab:maldaxura|@maldaxura]] was approved. * 14:33 [[gitlab:sirtobi|@sirtobi]] was approved. * 04:18 "ice5678" was rejected (pending since 2026-01-24T04:15:30.008Z). === 2026-04-24 === * 22:06 [[gitlab:arcstur|@arcstur]] was approved. === 2026-04-22 === * 23:06 "dtorsani" was rejected (pending since 2026-01-21T23:03:25.843Z). * 22:18 [[gitlab:egezort|@egezort]] was approved. * 16:45 "nexpectarpit" was rejected (pending since 2026-01-21T16:43:21.045Z). === 2026-04-20 === * 19:15 "fitch" was rejected (pending since 2026-01-19T19:12:35.644Z). === 2026-04-19 === * 02:54 [[gitlab:neoact|@neoact]] was approved. === 2026-04-18 === * 07:06 [[gitlab:kockaadmiralac|@kockaadmiralac]] was approved. === 2026-04-17 === * 13:42 "liselot" was rejected (pending since 2026-01-16T13:39:41.909Z). === 2026-04-15 === * 17:03 "lahari" was rejected (pending since 2026-01-14T17:02:06.275Z). === 2026-04-14 === * 13:00 "surajseth520" was rejected (pending since 2026-01-13T12:59:45.906Z). * 04:51 [[gitlab:canley|@canley]] was approved. * 01:03 "bshizzle" was rejected (pending since 2026-01-13T01:00:48.120Z). === 2026-04-13 === * 15:30 [[gitlab:passimacopoulos|@passimacopoulos]] was approved. === 2026-04-11 === * 12:30 "krithash" was rejected (pending since 2026-01-10T12:27:24.731Z). === 2026-04-10 === * 15:30 "raunak1709" was rejected (pending since 2026-01-09T15:29:10.901Z). === 2026-04-07 === * 17:03 [[gitlab:supnabla|@supnabla]] was approved. === 2026-04-06 === * 20:00 [[gitlab:laerdon|@laerdon]] was approved. * 19:21 [[gitlab:ljq3|@ljq3]] was approved. === 2026-04-04 === * 11:06 "mixcc" was rejected (pending since 2026-01-03T11:03:33.922Z). === 2026-04-02 === * 05:30 [[gitlab:mbh1|@mbh1]] was approved. === 2026-04-01 === * 18:21 "yuvrajpatil17" was rejected (pending since 2025-12-31T18:20:27.991Z). * 12:12 [[gitlab:amorii0|@amorii0]] was approved. === 2026-03-31 === * 11:00 "krrishsehgal" was rejected (pending since 2025-12-30T11:00:16.384Z). === 2026-03-30 === * 15:36 [[gitlab:atsuko|@atsuko]] was approved. === 2026-03-29 === * 11:36 [[gitlab:giftcup|@giftcup]] was approved. === 2026-03-28 === * 14:51 [[gitlab:janeeva1|@janeeva1]] was approved. === 2026-03-26 === * 13:36 [[gitlab:saiphani02|@saiphani02]] was approved. * 11:48 [[gitlab:valerioboz-wmch|@valerioboz-wmch]] was approved. === 2026-03-25 === * 09:45 "quansi" was rejected (pending since 2025-12-24T09:42:13.451Z). * 02:18 [[gitlab:viztor|@viztor]] was approved. === 2026-03-24 === * 23:18 [[gitlab:maryyann|@maryyann]] was approved. * 23:01 [[gitlab:codenamenoreste|@codenamenoreste]] was approved. * 13:36 [[gitlab:marc-maillard-wmse|@marc-maillard-wmse]] was approved. * 07:39 "fred2675" was rejected (pending since 2025-12-23T07:39:11.380Z). === 2026-03-23 === * 14:51 [[gitlab:komla|@komla]] was approved. * 05:51 "lunachuck43" was rejected (pending since 2025-12-22T05:50:17.862Z). * 04:06 "reza110011" was rejected (pending since 2025-12-22T04:05:25.117Z). === 2026-03-20 === * 21:54 "mertgor" was rejected (pending since 2025-12-19T21:51:51.419Z). * 20:57 "autanmahmah" was rejected (pending since 2025-12-19T20:54:51.678Z). * 09:57 [[gitlab:nethahussain|@nethahussain]] was approved. * 09:27 [[gitlab:piewriter|@piewriter]] was approved. * 08:15 [[gitlab:dondersmooi|@dondersmooi]] was approved. === 2026-03-19 === * 21:03 "sayvhior" was rejected (pending since 2025-12-18T21:02:31.699Z). === 2026-03-18 === * 20:15 [[gitlab:martinmystere|@martinmystere]] was approved. === 2026-03-17 === * 02:51 "louperivois" was rejected (pending since 2025-12-16T02:50:48.197Z). === 2026-03-16 === * 12:54 "mokayaj857" was rejected (pending since 2025-12-15T12:53:39.015Z). * 06:18 "roamer15" was rejected (pending since 2025-12-15T06:16:38.042Z). === 2026-03-14 === * 11:12 "umaramuhammad" was rejected (pending since 2025-12-13T11:10:44.004Z). * 09:33 "akuma19" was rejected (pending since 2025-12-13T09:31:39.044Z). * 07:06 [[gitlab:syunsyunminmin|@syunsyunminmin]] was approved. === 2026-03-12 === * 20:24 [[gitlab:11wb|@11wb]] was approved. * 09:54 [[gitlab:bcxfu75k|@bcxfu75k]] was approved. === 2026-03-10 === * 09:12 [[gitlab:viktoriahillerudwmse|@viktoriahillerudwmse]] was approved. === 2026-03-06 === * 08:09 "vazhayilnewone" was rejected (pending since 2025-12-05T08:07:02.184Z). === 2026-03-04 === * 20:54 [[gitlab:elphie|@elphie]] was approved. * 11:39 "ronaldahmed" was rejected (pending since 2025-12-03T11:37:47.492Z). * 02:12 "ltslw" was rejected (pending since 2025-12-03T02:11:52.040Z). === 2026-03-02 === * 19:21 "dlopez350" was rejected (pending since 2025-12-01T19:20:38.918Z). * 18:15 [[gitlab:lsandergreen|@lsandergreen]] was approved. === 2026-03-01 === * 10:51 [[gitlab:clintacc|@clintacc]] was approved. === 2026-02-28 === * 09:24 "cardboardlamp" was rejected (pending since 2025-11-29T09:22:03.947Z). * 08:18 "wiki-pavan" was rejected (pending since 2025-11-29T08:16:24.184Z). === 2026-02-27 === * 20:45 "thisisrick25" was rejected (pending since 2025-11-28T20:42:24.454Z). === 2026-02-26 === * 13:57 "chuiimuiiofc" was rejected (pending since 2025-11-27T13:57:02.794Z). * 13:54 "steffpro" was rejected (pending since 2025-11-27T13:52:10.859Z). === 2026-02-25 === * 21:24 "abubakarhabibudayyabu" was rejected (pending since 2025-11-26T21:22:37.776Z). === 2026-02-24 === * 05:00 "playboi" was rejected (pending since 2025-11-25T05:00:30.762Z). === 2026-02-23 === * 14:00 "alph65" was rejected (pending since 2025-11-24T13:59:00.797Z). * 12:33 [[gitlab:robertsky|@robertsky]] was approved. === 2026-02-22 === * 00:30 "hp8p" was rejected (pending since 2025-11-23T00:29:24.741Z). === 2026-02-19 === * 16:45 "clayjar" was rejected (pending since 2025-11-20T16:44:48.380Z). === 2026-02-18 === * 22:18 "nexus" was rejected (pending since 2025-11-19T22:16:48.818Z). * 12:00 "bernsteinnn" was rejected (pending since 2025-11-19T11:59:04.427Z). === 2026-02-17 === * 11:36 "jason2000-cpu" was rejected (pending since 2025-11-18T11:34:00.314Z). === 2026-02-16 === * 14:54 "smaurya" was rejected (pending since 2025-11-17T14:52:06.906Z). === 2026-02-15 === * 16:51 "kra-79" was rejected (pending since 2025-11-16T16:50:41.375Z). === 2026-02-14 === * 15:15 [[gitlab:mess|@mess]] was approved. === 2026-02-13 === * 13:57 "sopalsuemae957" was rejected (pending since 2025-11-14T13:55:16.921Z). * 13:30 [[gitlab:wyslijp16-toolforge|@wyslijp16-toolforge]] was approved. === 2026-02-12 === * 16:30 "kristinagligoric" was rejected (pending since 2025-11-13T16:29:21.646Z). * 03:33 [[gitlab:anyehansen|@anyehansen]] was approved. * 02:21 [[gitlab:thejoyfultentmaker|@thejoyfultentmaker]] was approved. === 2026-02-10 === * 13:18 [[gitlab:db111|@db111]] was approved. === 2026-02-09 === * 19:06 "squirrel289" was rejected (pending since 2025-11-10T19:04:27.831Z). === 2026-02-06 === * 20:54 [[gitlab:gillux|@gillux]] was approved. * 09:09 [[gitlab:lih|@lih]] was approved. === 2026-01-31 === * 16:21 [[gitlab:taxonbot1|@taxonbot1]] was approved. === 2026-01-28 === * 14:30 [[gitlab:ademola|@ademola]] was approved. * 10:51 "watshell" was rejected (pending since 2025-10-29T10:51:01.521Z). === 2026-01-26 === * 23:06 "tavaresgmg" was rejected (pending since 2025-10-27T23:04:42.140Z). === 2026-01-25 === * 06:03 "cata" was rejected (pending since 2025-10-26T06:01:26.155Z). === 2026-01-24 === * 21:15 [[gitlab:wiegels|@wiegels]] was approved. * 06:30 [[gitlab:blaquans|@blaquans]] was approved. === 2026-01-23 === * 16:27 [[gitlab:lerickson|@lerickson]] was approved. * 10:15 "fran0035g" was rejected (pending since 2025-10-24T10:12:17.732Z). === 2026-01-22 === * 21:00 "hacksyn" was rejected (pending since 2025-10-23T20:59:15.982Z). === 2026-01-21 === * 17:30 [[gitlab:otcenas11|@otcenas11]] was approved. === 2026-01-19 === * 21:48 [[gitlab:amdrel|@amdrel]] was approved. * 04:36 "rayalexa" was rejected (pending since 2025-10-20T04:35:02.094Z). === 2026-01-18 === * 15:45 "somya" was rejected (pending since 2025-10-19T15:43:43.701Z). * 06:54 "sergg001" was rejected (pending since 2025-10-19T06:54:12.296Z). === 2026-01-16 === * 11:57 "zeejohsy" was rejected (pending since 2025-10-17T11:56:22.372Z). * 04:45 "rocky25" was rejected (pending since 2025-10-17T04:43:33.180Z). === 2026-01-15 === * 16:39 "tiisu" was rejected (pending since 2025-10-16T16:37:18.438Z). * 12:00 "noahalorwu" was rejected (pending since 2025-10-16T11:58:26.133Z). * 10:39 "prjayaiuedu" was rejected (pending since 2025-10-16T10:37:16.947Z). === 2026-01-13 === * 17:21 [[gitlab:lwilson-ctr|@lwilson-ctr]] was approved. === 2026-01-12 === * 17:03 "stagietechs" was rejected (pending since 2025-10-13T17:02:25.281Z). === 2026-01-10 === * 19:06 "keerthisr" was rejected (pending since 2025-10-11T19:05:01.758Z). === 2026-01-09 === * 20:36 "lightb" was rejected (pending since 2025-10-10T20:34:20.264Z). === 2026-01-08 === * 19:42 [[gitlab:tbodt|@tbodt]] was approved. * 13:57 [[gitlab:martynranyard|@martynranyard]] was approved. === 2026-01-07 === * 17:48 [[gitlab:santanuwiki25|@santanuwiki25]] was approved. * 14:27 "dipanshu" was rejected (pending since 2025-10-08T14:26:10.794Z). * 12:30 "adeolaadesina" was rejected (pending since 2025-10-08T12:29:49.592Z). * 09:21 "tony-kamande" was rejected (pending since 2025-10-08T09:20:28.421Z). * 06:18 "hninwuttyi" was rejected (pending since 2025-10-08T06:17:28.006Z). * 05:09 "andume" was rejected (pending since 2025-10-08T05:07:18.582Z). * 02:00 "mosope" was rejected (pending since 2025-10-08T01:59:54.800Z). * 01:15 [[gitlab:tungstalite|@tungstalite]] was approved. === 2026-01-06 === * 18:24 "leerensucher" was rejected (pending since 2025-10-07T18:21:41.253Z). * 14:54 "leonidlednev" was rejected (pending since 2025-10-07T14:53:07.273Z). * 12:57 "alexandre-tingaud" was rejected (pending since 2025-10-07T12:54:27.206Z). === 2026-01-04 === * 21:33 [[gitlab:matr1x-101|@matr1x-101]] was approved. * 15:18 "makjr" was rejected (pending since 2025-10-05T15:16:31.558Z). * 14:09 "dakshq" was rejected (pending since 2025-10-05T14:08:40.608Z). === 2026-01-03 === * 20:42 [[gitlab:apehitkey|@apehitkey]] was approved. * 18:00 [[gitlab:jeremyb|@jeremyb]] was approved. * 14:09 [[gitlab:twelephant|@twelephant]] was approved. === 2026-01-01 === * 11:30 "shellstanislav" was rejected (pending since 2025-10-02T11:29:10.150Z). === 2025-12-30 === * 19:51 "camilojdiaz" was rejected (pending since 2025-09-30T19:49:24.913Z). === 2025-12-29 === * 16:03 "zied" was rejected (pending since 2025-09-29T16:01:30.415Z). * 08:18 "rahulsidpradhan" was rejected (pending since 2025-09-29T08:17:02.849Z). === 2025-12-26 === * 09:48 "thembo42" was rejected (pending since 2025-09-26T09:45:15.033Z). === 2025-12-25 === * 14:03 "196936074751" was rejected (pending since 2025-09-25T14:02:31.367Z). === 2025-12-23 === * 16:21 "ngarnsworthy" was rejected (pending since 2025-09-23T16:20:41.211Z). === 2025-12-22 === * 12:39 "aza555" was rejected (pending since 2025-09-22T12:38:02.622Z). === 2025-12-20 === * 23:45 "saph" was rejected (pending since 2025-09-20T23:45:01.222Z). === 2025-12-19 === * 10:15 "vladdymoses" was rejected (pending since 2025-09-19T10:15:00.999Z). * 07:15 "dirtylittlepoobah" was rejected (pending since 2025-09-19T07:13:55.537Z). === 2025-12-18 === * 16:24 [[gitlab:guyfawcus|@guyfawcus]] was approved. === 2025-12-17 === * 21:39 [[gitlab:holdyourhorses|@holdyourhorses]] was approved. * 18:30 "prudencia" was rejected (pending since 2025-09-17T18:27:18.860Z). * 02:24 "lottie" was rejected (pending since 2025-09-17T02:21:21.744Z). === 2025-12-16 === * 09:39 [[gitlab:melcatherine|@melcatherine]] was approved. * 08:54 [[gitlab:leila237|@leila237]] was approved. === 2025-12-15 === * 18:27 [[gitlab:royalsailor|@royalsailor]] was approved. * 09:39 [[gitlab:olaf8940|@olaf8940]] was approved. * 09:39 "brianbybyby" was rejected (pending since 2025-09-15T09:37:45.430Z). === 2025-12-14 === * 20:21 [[gitlab:essa237|@essa237]] was approved. * 16:42 [[gitlab:bovimacoco|@bovimacoco]] was approved. === 2025-12-13 === * 21:54 "mmns21" was rejected (pending since 2025-09-13T21:52:24.017Z). * 20:33 "bugcrawler" was rejected (pending since 2025-09-13T20:31:09.211Z). === 2025-12-12 === * 14:39 "ruvchoudhary" was rejected (pending since 2025-09-12T14:36:16.167Z). * 06:54 "rezadress" was rejected (pending since 2025-09-12T06:52:21.749Z). === 2025-12-10 === * 17:30 [[gitlab:itsmoon|@itsmoon]] was approved. === 2025-12-09 === * 15:42 [[gitlab:mercy-o|@mercy-o]] was approved. === 2025-12-06 === * 16:45 "jacquesradjabu" was rejected (pending since 2025-09-06T16:45:17.969Z). * 11:27 [[gitlab:ikhitron|@ikhitron]] was approved. === 2025-12-01 === * 08:12 "halconmilenario21" was rejected (pending since 2025-09-01T08:12:10.262Z). === 2025-11-30 === * 21:06 [[gitlab:habs|@habs]] was approved. === 2025-11-29 === * 16:36 "bovimacoco" was rejected (pending since 2025-08-30T16:34:39.712Z). * 00:45 [[gitlab:jjpmaster|@jjpmaster]] was approved. === 2025-11-24 === * 10:30 "alph65" was rejected (pending since 2025-08-25T10:28:40.957Z). * 02:24 [[gitlab:yaron|@yaron]] was approved. === 2025-11-20 === * 16:06 "clayjar" was rejected (pending since 2025-08-21T16:04:54.450Z). === 2025-11-17 === * 21:09 [[gitlab:ankita97531|@ankita97531]] was approved. === 2025-11-16 === * 14:15 "commanderkefir" was rejected (pending since 2025-08-17T14:13:14.791Z). * 08:21 "rehankhan78" was rejected (pending since 2025-08-17T08:19:44.896Z). === 2025-11-15 === * 14:36 "cyberscribe" was rejected (pending since 2025-08-16T14:34:27.230Z). === 2025-11-13 === * 04:21 "waddie96" was rejected (pending since 2025-08-14T04:19:27.461Z). === 2025-11-11 === * 06:42 [[gitlab:seanhoyland|@seanhoyland]] was approved. === 2025-11-10 === * 00:06 [[gitlab:jaredblumer|@jaredblumer]] was approved. === 2025-11-09 === * 22:36 "heinxiety" was rejected (pending since 2025-08-10T22:33:12.041Z). === 2025-11-07 === * 22:00 [[gitlab:forzagreen|@forzagreen]] was approved. === 2025-11-06 === * 16:57 [[gitlab:rsilvola|@rsilvola]] was approved. === 2025-11-04 === * 21:24 [[gitlab:devdoingdev|@devdoingdev]] was approved. === 2025-11-03 === * 17:48 "joewaleed98" was rejected (pending since 2025-08-04T17:46:12.191Z). === 2025-11-01 === * 18:00 "eliasempresas" was rejected (pending since 2025-08-02T17:58:04.412Z). === 2025-10-31 === * 18:51 [[gitlab:chaoticenby|@chaoticenby]] was approved. * 04:33 "3ch310n" was rejected (pending since 2025-08-01T04:32:21.982Z). === 2025-10-30 === * 10:03 [[gitlab:tausheefhassan|@tausheefhassan]] was approved. === 2025-10-29 === * 14:54 "theap" was rejected (pending since 2025-07-30T14:52:12.066Z). === 2025-10-28 === * 06:06 [[gitlab:tanbiruzzaman|@tanbiruzzaman]] was approved. === 2025-10-27 === * 07:51 [[gitlab:jmoore111|@jmoore111]] was approved. === 2025-10-25 === * 21:09 [[gitlab:valor|@valor]] was approved. * 21:03 [[gitlab:booksmurf|@booksmurf]] was approved. * 02:48 "mystyc1" was rejected (pending since 2025-07-26T02:46:19.373Z). === 2025-10-24 === * 05:12 "aadarshmahesh" was rejected (pending since 2025-07-25T05:09:38.264Z). === 2025-10-22 === * 20:54 [[gitlab:janewanga|@janewanga]] was approved. * 17:27 "abeljeevan" was rejected (pending since 2025-07-23T17:26:46.884Z). * 16:12 "shrimpnaur" was rejected (pending since 2025-07-23T16:10:37.864Z). === 2025-10-21 === * 18:51 "jrmuizel" was rejected (pending since 2025-07-22T18:50:07.315Z). * 09:33 [[gitlab:dpogorzelski|@dpogorzelski]] was approved. === 2025-10-17 === * 13:21 [[gitlab:blegodwin|@blegodwin]] was approved. === 2025-10-16 === * 14:51 [[gitlab:bahago|@bahago]] was approved. * 14:12 "harikrishna0005" was rejected (pending since 2025-07-17T14:10:48.385Z). * 14:09 "gauthammohanraj" was rejected (pending since 2025-07-17T14:08:47.643Z). === 2025-10-15 === * 13:48 [[gitlab:adwivedii|@adwivedii]] was approved. * 13:18 [[gitlab:kimbrenekakande|@kimbrenekakande]] was approved. * 13:03 "childmnajennifer" was rejected (pending since 2025-07-16T13:01:50.236Z). * 05:06 "vssb4214" was rejected (pending since 2025-07-16T05:05:33.985Z). === 2025-10-14 === * 19:39 [[gitlab:afanyulionel|@afanyulionel]] was approved. * 15:33 [[gitlab:sadrettin|@sadrettin]] was approved. * 14:18 [[gitlab:tmwyk|@tmwyk]] was approved. * 08:42 "yasu0796" was rejected (pending since 2025-07-15T08:41:26.453Z). === 2025-10-13 === * 16:09 [[gitlab:atlas0007|@atlas0007]] was approved. === 2025-10-11 === * 17:42 [[gitlab:techwizzie|@techwizzie]] was approved. === 2025-10-10 === * 19:03 [[gitlab:miiswom|@miiswom]] was approved. * 16:06 [[gitlab:ninatakang|@ninatakang]] was approved. === 2025-10-09 === * 15:42 [[gitlab:jaykaneki|@jaykaneki]] was approved. * 14:21 [[gitlab:lebogang|@lebogang]] was approved. * 14:15 [[gitlab:kimondorose|@kimondorose]] was approved. * 13:48 [[gitlab:joyakinyi|@joyakinyi]] was approved. * 13:48 [[gitlab:dikshyashahi|@dikshyashahi]] was approved. * 13:45 [[gitlab:obediobadiah|@obediobadiah]] was approved. * 13:45 [[gitlab:system625|@system625]] was approved. * 13:45 [[gitlab:rolalove|@rolalove]] was approved. * 13:39 [[gitlab:olatundeawo|@olatundeawo]] was approved. * 13:36 [[gitlab:danielchristlight|@danielchristlight]] was approved. * 13:36 [[gitlab:dipanshu1223|@dipanshu1223]] was approved. * 13:36 [[gitlab:aradhya|@aradhya]] was approved. * 09:57 "bognd" was rejected (pending since 2025-07-10T09:55:48.661Z). === 2025-10-08 === * 23:36 [[gitlab:sopzy|@sopzy]] was approved. * 23:03 [[gitlab:oluwatumininu|@oluwatumininu]] was approved. * 19:39 [[gitlab:levon003|@levon003]] was approved. * 15:24 [[gitlab:ritika-bhambri11|@ritika-bhambri11]] was approved. * 13:45 [[gitlab:anbanguyen|@anbanguyen]] was approved. * 13:36 [[gitlab:chumzine|@chumzine]] was approved. * 13:27 [[gitlab:shr0x-ya|@shr0x-ya]] was approved. * 12:45 [[gitlab:nurahwakili|@nurahwakili]] was approved. * 03:42 "nazhiba" was rejected (pending since 2025-07-09T03:40:12.625Z). * 02:12 "mafennel" was rejected (pending since 2025-07-09T02:11:40.598Z). === 2025-10-07 === * 22:54 [[gitlab:olusegunfaj|@olusegunfaj]] was approved. * 21:30 [[gitlab:rona|@rona]] was approved. * 21:09 [[gitlab:sandijigs|@sandijigs]] was approved. * 13:36 "xisbajao" was rejected (pending since 2025-07-08T13:33:35.018Z). * 01:36 "areczek94" was rejected (pending since 2025-07-08T01:35:40.633Z). === 2025-10-06 === * 19:21 "wmcarter2017" was rejected (pending since 2025-07-07T19:21:12.899Z). === 2025-10-05 === * 14:15 "meetmendapara" was rejected (pending since 2025-07-06T14:14:16.726Z). === 2025-10-04 === * 20:51 "nftbaee" was rejected (pending since 2025-07-05T20:50:57.688Z). === 2025-10-03 === * 06:12 [[gitlab:javiermonton|@javiermonton]] was approved. === 2025-10-02 === * 20:15 "talaqalotaibipmp" was rejected (pending since 2025-07-03T20:13:05.164Z). === 2025-10-01 === * 10:54 "bjensen" was rejected (pending since 2025-07-02T10:53:46.574Z). * 02:45 "kowal1984" was rejected (pending since 2025-07-02T02:44:56.946Z). === 2025-09-30 === * 21:21 [[gitlab:kavaljeetsingh|@kavaljeetsingh]] was approved. * 00:24 "adium" was rejected (pending since 2025-07-01T00:23:43.807Z). === 2025-09-28 === * 08:54 [[gitlab:pexerik|@pexerik]] was approved. === 2025-09-27 === * 13:57 [[gitlab:rubahhitamvukova|@rubahhitamvukova]] was approved. === 2025-09-26 === * 16:57 "algorithmic" was rejected (pending since 2025-06-27T16:56:17.480Z). * 13:54 [[gitlab:shadabgdg|@shadabgdg]] was approved. * 13:12 [[gitlab:spushpit|@spushpit]] was approved. === 2025-09-20 === * 14:06 "bwiki" was rejected (pending since 2025-06-21T13:59:14.749Z). === 2025-09-16 === * 05:39 [[gitlab:deepchirp|@deepchirp]] was approved. === 2025-09-15 === * 22:00 [[gitlab:noisk8|@noisk8]] was approved. * 11:03 "ahonc" was rejected (pending since 2025-06-16T11:00:54.843Z). === 2025-09-13 === * 18:24 "a-ssh22" was rejected (pending since 2025-06-14T18:23:33.937Z). * 12:36 [[gitlab:rajashreetalukdar|@rajashreetalukdar]] was approved. * 00:45 [[gitlab:sumitsurai|@sumitsurai]] was approved. === 2025-09-12 === * 17:12 [[gitlab:suyash23|@suyash23]] was approved. * 00:46 "remotetravel" was rejected (pending since 2025-06-13T00:44:08.171Z). === 2025-09-10 === * 21:09 "jancborchardt" was rejected (pending since 2025-06-11T21:06:30.759Z). === 2025-09-09 === * 17:03 [[gitlab:vwf|@vwf]] was approved. * 06:36 [[gitlab:cactusisme|@cactusisme]] was approved. === 2025-09-08 === * 18:09 "birushandegeya" was rejected (pending since 2025-06-09T18:08:00.087Z). * 16:27 "ngarnsworthy" was rejected (pending since 2025-06-09T16:24:37.213Z). * 12:33 "zolgoyo" was rejected (pending since 2025-06-09T12:31:34.199Z). === 2025-09-06 === * 23:09 [[gitlab:jaishsingh913|@jaishsingh913]] was approved. === 2025-09-05 === * 21:45 [[gitlab:sakshi2|@sakshi2]] was approved. * 20:42 "abdukhaliq1" was rejected (pending since 2025-06-06T20:40:42.023Z). * 14:27 "beubsamy" was rejected (pending since 2025-06-06T14:27:06.781Z). === 2025-09-04 === * 23:27 "sdhehua" was rejected (pending since 2025-06-05T23:24:45.777Z). * 19:00 [[gitlab:perry|@perry]] was approved. * 11:24 "saintwolf" was rejected (pending since 2025-06-05T11:21:20.176Z). === 2025-09-02 === * 05:48 [[gitlab:aliu|@aliu]] was approved. === 2025-08-29 === * 13:30 "kksurendran066" was rejected (pending since 2025-05-30T13:27:48.755Z). === 2025-08-28 === * 22:18 "tauraamuix" was rejected (pending since 2025-05-29T22:16:08.228Z). === 2025-08-26 === * 19:03 [[gitlab:dikkulah|@dikkulah]] was approved. === 2025-08-22 === * 23:51 [[gitlab:khoroshun_mike|@khoroshun_mike]] was approved. === 2025-08-21 === * 07:39 [[gitlab:yuka|@yuka]] was approved. === 2025-08-19 === * 07:48 [[gitlab:zhaofjx|@zhaofjx]] was approved. === 2025-08-17 === * 14:27 "madhan13k" was rejected (pending since 2025-05-18T14:26:08.973Z). === 2025-08-15 === * 10:15 "mohammed_abukhadra" was rejected (pending since 2025-05-16T10:14:48.403Z). === 2025-08-11 === * 11:48 "hmmyesbro" was rejected (pending since 2025-05-12T11:45:24.350Z). === 2025-08-10 === * 13:15 [[gitlab:dactyl|@dactyl]] was approved. === 2025-08-09 === * 04:39 "xxxx100000" was rejected (pending since 2025-05-10T04:37:44.949Z). === 2025-08-08 === * 14:33 [[gitlab:josefanthony|@josefanthony]] was approved. === 2025-08-07 === * 23:42 [[gitlab:robins7|@robins7]] was approved. * 21:42 [[gitlab:pols12|@pols12]] was approved. * 17:15 "sbronson" was rejected (pending since 2025-05-08T17:15:08.834Z). * 14:57 [[gitlab:alvindulle|@alvindulle]] was approved. * 14:45 [[gitlab:xentos|@xentos]] was approved. * 06:27 "jamesboste" was rejected (pending since 2025-05-08T06:25:14.793Z). * 03:57 "ysun" was rejected (pending since 2025-05-08T03:55:07.348Z). === 2025-08-06 === * 21:51 "pols12" was rejected (pending since 2025-05-07T21:49:13.598Z). * 01:51 "okeamah" was rejected (pending since 2025-05-07T01:48:50.114Z). === 2025-08-05 === * 09:15 "mobashir-2013" was rejected (pending since 2025-05-06T09:14:24.069Z). === 2025-08-01 === * 08:00 "douginamug" was rejected (pending since 2025-05-02T07:57:38.317Z). === 2025-07-31 === * 02:30 [[gitlab:ads|@ads]] was approved. === 2025-07-27 === * 13:15 "mrico2703" was rejected (pending since 2025-04-27T13:13:12.346Z). * 10:17 [[gitlab:josephfrancis12|@josephfrancis12]] was approved. * 10:17 [[gitlab:fuzzew|@fuzzew]] was approved. * 05:57 [[gitlab:biscuitbobby|@biscuitbobby]] was approved. * 05:48 [[gitlab:ecoholic|@ecoholic]] was approved. === 2025-07-26 === * 11:48 [[gitlab:chimnayyyy|@chimnayyyy]] was approved. * 11:48 [[gitlab:alwinalbert|@alwinalbert]] was approved. * 11:48 [[gitlab:hridyakk|@hridyakk]] was approved. * 11:45 [[gitlab:gaurigupta21|@gaurigupta21]] was approved. * 11:45 [[gitlab:binetaa|@binetaa]] was approved. * 10:21 [[gitlab:jyothikat22|@jyothikat22]] was approved. * 10:21 [[gitlab:zobotrombie|@zobotrombie]] was approved. * 10:21 [[gitlab:flykrth|@flykrth]] was approved. * 10:21 [[gitlab:mehrinshamim|@mehrinshamim]] was approved. * 10:21 [[gitlab:aadhi13|@aadhi13]] was approved. * 10:21 [[gitlab:malavikam05|@malavikam05]] was approved. * 10:18 [[gitlab:nf609|@nf609]] was approved. * 05:48 [[gitlab:nazalnihad|@nazalnihad]] was approved. * 05:48 [[gitlab:naveen28204280|@naveen28204280]] was approved. === 2025-07-25 === * 09:49 [[gitlab:kasyap9|@kasyap9]] was approved. * 09:30 [[gitlab:swayamagrahari|@swayamagrahari]] was approved. === 2025-07-24 === * 19:36 [[gitlab:madutgn|@madutgn]] was approved. === 2025-07-23 === * 20:09 [[gitlab:somerandomdeveloper|@somerandomdeveloper]] was approved. === 2025-07-22 === * 00:15 [[gitlab:iagoqnsi|@iagoqnsi]] was approved. === 2025-07-21 === * 17:30 [[gitlab:asadiqui|@asadiqui]] was approved. * 16:39 [[gitlab:tryvix1509|@tryvix1509]] was approved. * 04:27 [[gitlab:damian|@damian]] was approved. === 2025-07-20 === * 09:42 "mike-khoroshun" was rejected (pending since 2025-04-20T09:42:22.732Z). === 2025-07-17 === * 17:57 [[gitlab:haroldkrabs|@haroldkrabs]] was approved. * 13:45 [[gitlab:envlh|@envlh]] was approved. === 2025-07-14 === * 10:24 [[gitlab:missguru|@missguru]] was approved. * 00:57 "clarfonthey" was rejected (pending since 2025-04-14T00:56:32.626Z). === 2025-07-13 === * 01:01 [[gitlab:l235|@l235]] was approved. === 2025-07-11 === * 03:06 "rodavlas" was rejected (pending since 2025-04-11T03:05:45.590Z). === 2025-07-06 === * 00:09 "lakasa" was rejected (pending since 2025-04-06T00:06:28.469Z). === 2025-07-05 === * 21:54 "ctrlzvi" was rejected (pending since 2025-04-05T21:54:12.542Z). * 14:30 "aminualiyu" was rejected (pending since 2025-04-05T14:27:22.617Z). === 2025-07-04 === * 03:15 [[gitlab:galstar|@galstar]] was approved. === 2025-07-02 === * 11:27 "vicolas11" was rejected (pending since 2025-04-02T11:25:12.682Z). === 2025-06-29 === * 23:12 "naomi723" was rejected (pending since 2025-03-30T23:09:24.630Z). === 2025-06-28 === * 16:21 "mudeh2372" was rejected (pending since 2025-03-29T16:18:27.057Z). === 2025-06-27 === * 23:18 "rony143" was rejected (pending since 2025-03-28T23:16:13.671Z). * 22:21 [[gitlab:rluts|@rluts]] was approved. === 2025-06-26 === * 13:54 "creativegurus" was rejected (pending since 2025-03-27T13:52:41.706Z). === 2025-06-24 === * 17:42 [[gitlab:devjadiya|@devjadiya]] was approved. * 14:00 "dominic-r" was rejected (pending since 2025-03-25T14:00:07.307Z). === 2025-06-21 === * 00:48 [[gitlab:vriaa|@vriaa]] was approved. === 2025-06-18 === * 15:21 "ayushkhati1" was rejected (pending since 2025-03-19T15:18:50.062Z). === 2025-06-17 === * 20:45 "chiomavero" was rejected (pending since 2025-03-18T20:44:13.967Z). * 00:27 [[gitlab:eggroll97|@eggroll97]] was approved. === 2025-06-14 === * 20:57 "volvox" was rejected (pending since 2025-03-15T20:56:34.018Z). === 2025-06-13 === * 16:09 [[gitlab:supergrey|@supergrey]] was approved. * 11:03 "chqaz" was rejected (pending since 2025-03-14T11:01:09.600Z). * 10:24 [[gitlab:slong-wmf|@slong-wmf]] was approved. * 10:15 "hearvox" was rejected (pending since 2025-03-14T10:13:13.112Z). === 2025-06-12 === * 15:18 "jlam" was rejected (pending since 2025-03-13T15:17:54.099Z). === 2025-06-09 === * 20:48 "dipanjansengupta" was rejected (pending since 2025-03-10T20:48:03.545Z). * 19:27 [[gitlab:reggycelly|@reggycelly]] was approved. * 14:51 "arendpieter" was rejected (pending since 2025-03-10T14:51:01.445Z). * 13:21 [[gitlab:greenreaper|@greenreaper]] was approved. * 09:33 [[gitlab:mmta|@mmta]] was approved. * 08:03 "a-ssh22" was rejected (pending since 2025-03-10T08:03:08.111Z). === 2025-06-08 === * 21:06 "mm-episodenlistedlvaupdater" was rejected (pending since 2025-03-09T21:04:06.323Z). === 2025-06-06 === * 11:06 [[gitlab:olea|@olea]] was approved. === 2025-06-05 === * 20:33 [[gitlab:encodedwp|@encodedwp]] was approved. * 15:00 [[gitlab:toluayo|@toluayo]] was approved. * 13:51 [[gitlab:arnold_lup|@arnold_lup]] was approved. * 11:54 "sdhehua" was rejected (pending since 2025-03-06T11:51:48.241Z). === 2025-06-03 === * 21:27 [[gitlab:wewakey|@wewakey]] was approved. * 12:36 "hunsimon2" was rejected (pending since 2025-03-04T12:34:56.520Z). * 11:54 "hunsimon" was rejected (pending since 2025-03-04T11:53:54.652Z). === 2025-06-02 === * 12:01 [[gitlab:jaimedes|@jaimedes]] was approved. === 2025-05-30 === * 18:00 "sathvik9105" was rejected (pending since 2025-02-28T17:59:42.867Z). * 11:21 [[gitlab:tonythomas01|@tonythomas01]] was approved. * 10:06 [[gitlab:gpsleo|@gpsleo]] was approved. === 2025-05-29 === * 22:12 [[gitlab:codynguyen1116|@codynguyen1116]] was approved. === 2025-05-28 === * 02:57 [[gitlab:saper|@saper]] was approved. === 2025-05-27 === * 21:06 [[gitlab:mohammed_qays|@mohammed_qays]] was approved. * 15:33 "satanluimm" was rejected (pending since 2025-02-25T15:32:48.101Z). === 2025-05-26 === * 23:57 "seyedali220" was rejected (pending since 2025-02-24T23:56:17.621Z). === 2025-05-21 === * 11:12 [[gitlab:guilherme|@guilherme]] was approved. === 2025-05-19 === * 13:24 [[gitlab:emojiwiki|@emojiwiki]] was approved. === 2025-05-18 === * 00:00 "xidme" was rejected (pending since 2025-02-15T23:58:56.796Z). === 2025-05-17 === * 02:39 "kdh8219" was rejected (pending since 2025-02-15T02:36:32.237Z). === 2025-05-16 === * 15:09 [[gitlab:maxbinderwmf|@maxbinderwmf]] was approved. === 2025-05-15 === * 04:30 "inspectorzer0" was rejected (pending since 2025-02-13T04:27:33.179Z). === 2025-05-14 === * 17:42 [[gitlab:llugo|@llugo]] was approved. === 2025-05-13 === * 20:18 "mmta" was rejected (pending since 2025-02-11T20:17:23.407Z). === 2025-05-11 === * 20:51 "jad" was rejected (pending since 2025-02-09T20:49:07.333Z). * 17:54 "nishchalsundan" was rejected (pending since 2025-02-09T17:52:25.761Z). * 16:39 "mohammed_abukhadra" was rejected (pending since 2025-02-09T16:39:03.730Z). === 2025-05-09 === * 09:12 [[gitlab:sirchanmp|@sirchanmp]] was approved. === 2025-05-08 === * 08:18 [[gitlab:mengeditch|@mengeditch]] was approved. === 2025-05-07 === * 03:45 "xluffy" was rejected (pending since 2025-02-05T03:45:14.181Z). === 2025-05-06 === * 16:54 "punhaniabhishek" was rejected (pending since 2025-02-04T16:53:50.758Z). * 09:36 [[gitlab:bmartinezcalvo|@bmartinezcalvo]] was approved. === 2025-05-02 === * 12:24 [[gitlab:tohaomg|@tohaomg]] was approved. * 11:48 [[gitlab:mavrikant|@mavrikant]] was approved. * 11:45 [[gitlab:daanvr|@daanvr]] was approved. === 2025-05-01 === * 09:09 "mjoerg" was rejected (pending since 2025-01-30T09:09:04.204Z). === 2025-04-30 === * 23:06 "sanskardubey" was rejected (pending since 2025-01-29T23:03:25.489Z). === 2025-04-29 === * 16:00 "geyslein" was rejected (pending since 2025-01-28T16:00:01.510Z). === 2025-04-26 === * 09:30 "anjali9027" was rejected (pending since 2025-01-25T09:28:07.064Z). === 2025-04-25 === * 18:00 "salahhazaa" was rejected (pending since 2025-01-24T17:58:30.030Z). * 15:15 [[gitlab:yiming|@yiming]] was approved. * 02:06 "mrchanmp" was rejected (pending since 2025-01-24T02:03:58.308Z). === 2025-04-23 === * 17:03 "rj2904" was rejected (pending since 2025-01-22T17:03:11.207Z). * 14:21 "nischay33" was rejected (pending since 2025-01-22T14:19:21.081Z). === 2025-04-22 === * 19:27 "dj80" was rejected (pending since 2025-01-21T19:25:28.498Z). * 14:30 [[gitlab:kaimamin|@kaimamin]] was approved. * 09:57 "debo" was rejected (pending since 2025-01-21T09:54:47.955Z). === 2025-04-21 === * 12:24 "unshell" was rejected (pending since 2025-01-20T12:21:59.686Z). === 2025-04-18 === * 15:06 [[gitlab:spartanarbinger|@spartanarbinger]] was approved. === 2025-04-16 === * 03:09 "dewey" was rejected (pending since 2025-01-15T03:06:17.488Z). === 2025-04-15 === * 19:45 "emdadul" was rejected (pending since 2025-01-14T19:42:29.285Z). === 2025-04-14 === * 06:45 [[gitlab:bcampbell804|@bcampbell804]] was approved. === 2025-04-11 === * 06:27 [[gitlab:jvanderhoop|@jvanderhoop]] was approved. === 2025-04-10 === * 04:12 "bhai420" was rejected (pending since 2025-01-09T04:10:29.430Z). === 2025-04-09 === * 05:03 "austinvarshney" was rejected (pending since 2025-01-08T05:02:34.175Z). === 2025-04-06 === * 15:36 [[gitlab:elph|@elph]] was approved. === 2025-04-02 === * 10:33 [[gitlab:ozge|@ozge]] was approved. === 2025-03-31 === * 20:15 "demandkey" was rejected (pending since 2024-12-30T20:14:23.096Z). * 15:18 [[gitlab:danyya|@danyya]] was approved. === 2025-03-28 === * 15:54 [[gitlab:rutsavi09|@rutsavi09]] was approved. * 15:54 [[gitlab:ilanen1|@ilanen1]] was approved. === 2025-03-25 === * 19:27 [[gitlab:irfo|@irfo]] was approved. * 11:54 [[gitlab:kmontalva-wmf|@kmontalva-wmf]] was approved. * 04:33 [[gitlab:paul26|@paul26]] was approved. * 04:18 "as1100k" was rejected (pending since 2024-12-24T04:18:06.813Z). === 2025-03-24 === * 11:33 "amzadkhankk" was rejected (pending since 2024-12-23T11:33:14.176Z). === 2025-03-23 === * 12:24 "wolfdo" was rejected (pending since 2024-12-22T12:23:35.056Z). === 2025-03-22 === * 09:45 [[gitlab:fjmustak|@fjmustak]] was approved. === 2025-03-20 === * 18:42 "sathishkokila" was rejected (pending since 2024-12-19T18:39:35.161Z). * 17:03 [[gitlab:alien4444|@alien4444]] was approved. * 15:27 [[gitlab:davidcoronel|@davidcoronel]] was approved. === 2025-03-19 === * 22:57 [[gitlab:r1f4t|@r1f4t]] was approved. * 19:03 "daniel24ps" was rejected (pending since 2024-12-18T19:00:21.249Z). * 14:18 [[gitlab:beepbooppenguin|@beepbooppenguin]] was approved. === 2025-03-18 === * 17:48 "rahulkundu1209" was rejected (pending since 2024-12-17T17:46:41.936Z). * 08:15 "kirtisikka972" was rejected (pending since 2024-12-17T08:13:25.487Z). === 2025-03-15 === * 13:30 "tulspal_sidhu" was rejected (pending since 2024-12-14T13:29:10.606Z). * 01:39 "peacedeadc" was rejected (pending since 2024-12-14T01:37:36.579Z). === 2025-03-14 === * 03:51 [[gitlab:chuckthebuck|@chuckthebuck]] was approved. * 02:33 "yxngtrtxll" was rejected (pending since 2024-12-13T02:31:51.658Z). === 2025-03-13 === * 14:36 [[gitlab:iccander|@iccander]] was approved. === 2025-03-12 === * 23:21 "jokerchic36" was rejected (pending since 2024-12-11T23:21:00.670Z). * 15:30 [[gitlab:naomi|@naomi]] was approved. * 15:27 [[gitlab:cobi|@cobi]] was approved. === 2025-03-11 === * 12:42 "mohitvermaxx" was rejected (pending since 2024-12-10T12:40:56.967Z). === 2025-03-10 === * 16:51 [[gitlab:nanona15dobato|@nanona15dobato]] was approved. === 2025-03-09 === * 22:39 [[gitlab:jonkolbert|@jonkolbert]] was approved. * 20:45 [[gitlab:urbanecmtest2|@urbanecmtest2]] was approved. === 2025-03-07 === * 16:54 [[gitlab:hswan|@hswan]] was approved. * 14:42 [[gitlab:atitkov|@atitkov]] was approved. * 00:42 [[gitlab:infrastruktur|@infrastruktur]] was approved. === 2025-03-06 === * 17:21 "johnmann" was rejected (pending since 2024-12-05T17:19:24.995Z). === 2025-03-05 === * 07:33 [[gitlab:monx9494|@monx9494]] was approved. === 2025-03-02 === * 21:21 "paul26" was rejected (pending since 2024-12-01T21:20:19.681Z). === 2025-03-01 === * 19:15 [[gitlab:izno|@izno]] was approved. * 12:45 [[gitlab:nyerho|@nyerho]] was approved. === 2025-02-28 === * 18:27 [[gitlab:chuckonwumelu|@chuckonwumelu]] was approved. * 13:09 "ashwinpraveengo" was rejected (pending since 2024-11-29T13:07:47.240Z). * 00:18 "eduardoaugusto" was rejected (pending since 2024-11-29T00:17:43.372Z). === 2025-02-27 === * 20:39 "volkanurl" was rejected (pending since 2024-11-28T20:37:18.101Z). === 2025-02-24 === * 21:15 [[gitlab:feeglgeef|@feeglgeef]] was approved. * 20:18 [[gitlab:piaanalysis2|@piaanalysis2]] was approved. * 19:06 [[gitlab:dhardy|@dhardy]] was approved. === 2025-02-22 === * 19:27 [[gitlab:owuh|@owuh]] was approved. === 2025-02-19 === * 16:06 [[gitlab:artemkloko|@artemkloko]] was approved. * 13:03 [[gitlab:jgafnea|@jgafnea]] was approved. === 2025-02-17 === * 16:33 [[gitlab:asmartkitten|@asmartkitten]] was approved. === 2025-02-16 === * 19:12 "gaurigupta21" was rejected (pending since 2024-11-17T19:11:07.416Z). === 2025-02-15 === * 01:18 [[gitlab:mediawiki-quickstart-ci|@mediawiki-quickstart-ci]] was approved. === 2025-02-14 === * 15:21 "nathanbnm" was rejected (pending since 2024-11-15T15:18:19.632Z). === 2025-02-13 === * 16:45 [[gitlab:priyanshuchahal|@priyanshuchahal]] was approved. * 16:42 [[gitlab:ajhalili2006|@ajhalili2006]] was approved. === 2025-02-12 === * 23:21 "monkeypatch999" was rejected (pending since 2024-11-13T23:20:38.398Z). * 06:36 [[gitlab:jainlakshita28|@jainlakshita28]] was approved. === 2025-02-11 === * 19:27 [[gitlab:matthewsm2|@matthewsm2]] was approved. === 2025-02-09 === * 16:15 "mohammed_abukhadra" was rejected (pending since 2024-11-10T16:15:18.361Z). === 2025-02-07 === * 21:33 "brennan" was rejected (pending since 2024-11-08T21:31:07.351Z). === 2025-02-06 === * 08:24 "mmta" was rejected (pending since 2024-11-07T08:22:36.724Z). * 06:21 [[gitlab:bunnypranav|@bunnypranav]] was approved. === 2025-02-05 === * 22:39 "chrissteinchen" was rejected (pending since 2024-11-06T22:38:16.673Z). === 2025-02-03 === * 07:45 "edriiic" was rejected (pending since 2024-11-04T07:44:46.849Z). * 01:12 "geppy" was rejected (pending since 2024-11-04T01:10:48.710Z). === 2025-02-02 === * 13:18 "funa-enpitu" was rejected (pending since 2024-11-03T13:15:46.065Z). === 2025-01-31 === * 23:42 "nfontes" was rejected (pending since 2024-11-01T23:39:41.755Z). * 22:51 "sbronson" was rejected (pending since 2024-11-01T22:50:31.871Z). * 00:42 [[gitlab:farid|@farid]] was approved. === 2025-01-27 === * 08:15 [[gitlab:eliza189|@eliza189]] was approved. === 2025-01-25 === * 09:51 [[gitlab:pamputt|@pamputt]] was approved. === 2025-01-23 === * 14:30 [[gitlab:lubianat|@lubianat]] was approved. * 11:45 [[gitlab:bootsa|@bootsa]] was approved. === 2025-01-21 === * 05:09 "niko" was rejected (pending since 2024-07-21T16:10:01.377Z). * 05:09 "thawizkid369777" was rejected (pending since 2024-07-18T17:42:44.493Z). * 05:09 "sarthaksingh2" was rejected (pending since 2024-07-10T11:31:30.470Z). * 05:09 "shriyakt" was rejected (pending since 2024-07-06T04:54:10.248Z). * 05:09 "akshaya" was rejected (pending since 2024-07-06T04:04:51.488Z). * 05:09 "alaka03aj" was rejected (pending since 2024-07-05T18:01:54.876Z). * 05:09 "sulochanaviji-5049" was rejected (pending since 2024-07-01T05:58:00.427Z). * 05:09 "nayanjnath" was rejected (pending since 2024-07-01T02:51:57.405Z). * 05:09 "sd44" was rejected (pending since 2024-06-30T04:28:51.436Z). * 05:09 "metavalent" was rejected (pending since 2024-06-29T01:37:14.210Z). * 05:09 "wicloudx" was rejected (pending since 2024-06-28T11:51:23.335Z). * 05:09 "debo" was rejected (pending since 2024-06-28T01:44:59.845Z). * 05:09 "bwiki" was rejected (pending since 2024-06-23T14:15:38.032Z). * 05:09 "toprak" was rejected (pending since 2024-06-23T11:35:50.819Z). * 05:09 "iristeller" was rejected (pending since 2024-06-14T20:53:48.959Z). * 05:09 "jcolvin" was rejected (pending since 2024-06-12T17:29:01.238Z). * 05:09 "kalyan" was rejected (pending since 2024-06-07T07:52:46.993Z). * 05:09 "bluecrystal" was rejected (pending since 2024-06-06T19:16:20.107Z). * 05:09 "iftttrohit" was rejected (pending since 2024-06-04T12:08:50.818Z). * 05:09 "pogpotato" was rejected (pending since 2024-06-03T17:58:21.684Z). * 05:09 "cptlausebaer" was rejected (pending since 2024-05-31T18:53:27.692Z). * 05:09 "hdevine825" was rejected (pending since 2024-05-31T17:04:18.279Z). * 05:09 "anaghaa18" was rejected (pending since 2024-05-25T19:14:31.803Z). * 05:09 "atharvanair04" was rejected (pending since 2024-05-25T14:24:52.825Z). * 05:09 "anasvemmully" was rejected (pending since 2024-05-25T06:10:27.261Z). * 05:09 "abhinavmohandas" was rejected (pending since 2024-05-25T06:05:24.825Z). * 05:09 "kksurendran06" was rejected (pending since 2024-05-25T06:04:38.082Z). * 05:09 "albertmarshall8896" was rejected (pending since 2024-05-23T09:32:05.462Z). * 05:09 "akellison" was rejected (pending since 2024-05-17T02:07:24.229Z). * 05:09 "mainowill" was rejected (pending since 2024-04-16T23:30:33.881Z). * 05:09 "bzhqc" was rejected (pending since 2024-04-16T19:50:38.676Z). * 05:09 "safan41" was rejected (pending since 2024-04-16T03:34:48.942Z). * 05:09 "mgagat" was rejected (pending since 2024-04-16T03:21:51.764Z). * 05:09 "okeamah" was rejected (pending since 2024-04-16T02:49:00.143Z). * 05:09 "xuhao61" was rejected (pending since 2024-04-15T23:45:09.083Z). * 04:47 "cybel" was rejected (pending since 2024-04-15T06:46:35.791Z). === 2025-01-20 === * 14:33 [[gitlab:your1|@your1]] was approved. === 2025-01-18 === * 10:09 [[gitlab:galrach600|@galrach600]] was approved. * 02:51 [[gitlab:blankeclair|@blankeclair]] was approved. === 2025-01-17 === * 13:57 [[gitlab:dsantamaria|@dsantamaria]] was approved. === 2025-01-15 === * 17:12 [[gitlab:smartse|@smartse]] was approved. === 2025-01-14 === * 17:03 [[gitlab:naorleizer|@naorleizer]] was approved. === 2025-01-13 === * 02:45 [[gitlab:wolf20482|@wolf20482]] was approved. === 2025-01-12 === * 17:45 [[gitlab:tamzin|@tamzin]] was approved. === 2025-01-11 === * 15:24 [[gitlab:bargioni|@bargioni]] was approved. * 14:30 [[gitlab:salelya|@salelya]] was approved. * 10:15 [[gitlab:malakatshy|@malakatshy]] was approved. * 05:21 [[gitlab:newmcpee|@newmcpee]] was approved. === 2025-01-09 === * 15:30 [[gitlab:gkyziridis|@gkyziridis]] was approved. === 2025-01-08 === * 16:21 [[gitlab:ukrface|@ukrface]] was approved. === 2024-12-28 === * 03:27 [[gitlab:twonum|@twonum]] was approved. === 2024-12-25 === * 06:09 [[gitlab:harsv567|@harsv567]] was approved. === 2024-12-21 === * 11:24 [[gitlab:amutha2002|@amutha2002]] was approved. === 2024-12-20 === * 19:51 [[gitlab:hridyeshgupta|@hridyeshgupta]] was approved. * 10:00 [[gitlab:ro-shines|@ro-shines]] was approved. * 08:09 [[gitlab:kesharwaniarpita|@kesharwaniarpita]] was approved. === 2024-12-18 === * 14:45 [[gitlab:soylacarli|@soylacarli]] was approved. === 2024-12-16 === * 20:33 [[gitlab:aleyasiddika1|@aleyasiddika1]] was approved. === 2024-12-15 === * 07:33 [[gitlab:abhishek02bhardwaj|@abhishek02bhardwaj]] was approved. === 2024-12-13 === * 13:18 [[gitlab:ashmitabathre204|@ashmitabathre204]] was approved. === 2024-12-10 === * 06:39 [[gitlab:ginaan|@ginaan]] was approved. === 2024-12-09 === * 05:45 [[gitlab:kallinavya|@kallinavya]] was approved. * 00:54 [[gitlab:viserion-7|@viserion-7]] was approved. === 2024-12-08 === * 17:27 [[gitlab:wargo|@wargo]] was approved. === 2024-12-05 === * 11:15 [[gitlab:ranjithraj|@ranjithraj]] was approved. === 2024-12-02 === * 21:21 [[gitlab:a930913|@a930913]] was approved. === 2024-12-01 === * 02:39 [[gitlab:kingchristlike1|@kingchristlike1]] was approved. === 2024-11-21 === * 13:45 [[gitlab:sascha|@sascha]] was approved. === 2024-11-19 === * 16:36 [[gitlab:jly|@jly]] was approved. === 2024-11-15 === * 02:54 [[gitlab:danielyepezgarces|@danielyepezgarces]] was approved. === 2024-11-14 === * 14:15 [[gitlab:stimoroll|@stimoroll]] was approved. === 2024-11-09 === * 17:15 [[gitlab:f4udeveloper|@f4udeveloper]] was approved. === 2024-11-07 === * 19:15 [[gitlab:zulf|@zulf]] was approved. * 05:33 [[gitlab:hassanamin|@hassanamin]] was approved. === 2024-11-06 === * 19:39 [[gitlab:daniuu|@daniuu]] was approved. * 00:18 [[gitlab:rlopez-wmf|@rlopez-wmf]] was approved. === 2024-10-09 === * 14:45 [[gitlab:jtweed|@jtweed]] was approved. * 10:24 [[gitlab:ifrahkh|@ifrahkh]] was approved. * 09:06 [[gitlab:wikibayer|@wikibayer]] was approved. === 2024-10-06 === * 10:27 [[gitlab:keerthan16|@keerthan16]] was approved. === 2024-10-04 === * 07:45 [[gitlab:hakimi97|@hakimi97]] was approved. === 2024-09-30 === * 07:39 [[gitlab:ninjastrikers|@ninjastrikers]] was approved. === 2024-09-28 === * 17:30 [[gitlab:webrunner95|@webrunner95]] was approved. === 2024-09-18 === * 21:39 [[gitlab:elliottetzkorn|@elliottetzkorn]] was approved. === 2024-09-14 === * 22:06 [[gitlab:humptydumpty|@humptydumpty]] was approved. === 2024-09-06 === * 08:48 [[gitlab:mickabarber|@mickabarber]] was approved. === 2024-08-27 === * 17:36 [[gitlab:edgars|@edgars]] was approved. === 2024-08-22 === * 09:18 [[gitlab:antonkokhwmde|@antonkokhwmde]] was approved. === 2024-08-14 === * 19:21 [[gitlab:jfk|@jfk]] was approved. === 2024-08-13 === * 17:57 [[gitlab:daxserver|@daxserver]] was approved. === 2024-08-11 === * 09:57 [[gitlab:pauliesnug|@pauliesnug]] was approved. === 2024-08-10 === * 08:42 [[gitlab:ashig|@ashig]] was approved. === 2024-08-09 === * 14:09 [[gitlab:masssly|@masssly]] was approved. === 2024-08-05 === * 22:15 [[gitlab:mrtortue|@mrtortue]] was approved. === 2024-08-02 === * 16:21 [[gitlab:dsantini|@dsantini]] was approved. === 2024-07-31 === * 11:54 [[gitlab:cptviraj|@cptviraj]] was approved. === 2024-07-30 === * 19:09 [[gitlab:iniquity|@iniquity]] was approved. * 10:00 [[gitlab:collins|@collins]] was approved. === 2024-07-27 === * 15:57 [[gitlab:songnguxyz|@songnguxyz]] was approved. === 2024-07-25 === * 12:36 [[gitlab:mszabo|@mszabo]] was approved. * 09:21 [[gitlab:agarwalmahima|@agarwalmahima]] was approved. === 2024-07-24 === * 08:05 [[gitlab:dragoniez|@dragoniez]] was approved. === 2024-07-23 === * 06:54 [[gitlab:mirji|@mirji]] was approved. === 2024-07-16 === * 10:00 [[gitlab:lakejason0|@lakejason0]] was approved. === 2024-07-12 === * 11:33 [[gitlab:cn|@cn]] was approved. * 08:12 [[gitlab:unchampignon|@unchampignon]] was approved. === 2024-07-07 === * 17:12 [[gitlab:agamyasamuel|@agamyasamuel]] was approved. * 05:24 [[gitlab:kuldeepburjbhalaike|@kuldeepburjbhalaike]] was approved. === 2024-07-06 === * 11:18 [[gitlab:dibya|@dibya]] was approved. * 04:54 [[gitlab:sarthakparashar|@sarthakparashar]] was approved. === 2024-07-05 === * 18:15 [[gitlab:vanshikarathi|@vanshikarathi]] was approved. === 2024-07-02 === * 19:00 [[gitlab:ebrahim|@ebrahim]] was approved. === 2024-07-01 === * 20:12 [[gitlab:rockingpenny4|@rockingpenny4]] was approved. * 18:15 [[gitlab:balajijagadesh|@balajijagadesh]] was approved. === 2024-06-30 === * 18:24 [[gitlab:hrideshmg|@hrideshmg]] was approved. * 07:18 [[gitlab:chanakyakumardas|@chanakyakumardas]] was approved. * 06:30 [[gitlab:rihaan180|@rihaan180]] was approved. === 2024-06-27 === * 17:36 [[gitlab:driedmueller|@driedmueller]] was approved. === 2024-06-19 === * 12:57 [[gitlab:audreypenven|@audreypenven]] was approved. === 2024-06-16 === * 01:18 [[gitlab:roysmith|@roysmith]] was approved. === 2024-06-08 === * 02:45 [[gitlab:jleedev|@jleedev]] was approved. === 2024-06-03 === * 13:57 [[gitlab:afeder|@afeder]] was approved. === 2024-06-01 === * 10:54 [[gitlab:florianschmitt|@florianschmitt]] was approved. === 2024-05-30 === * 16:42 [[gitlab:krlsca|@krlsca]] was approved. === 2024-05-28 === * 11:24 [[gitlab:rickijay|@rickijay]] was approved. === 2024-05-26 === * 11:18 [[gitlab:ranjithsiji|@ranjithsiji]] was approved. === 2024-05-25 === * 07:24 [[gitlab:jony|@jony]] was approved. === 2024-05-23 === * 08:45 [[gitlab:lepticed7|@lepticed7]] was approved. === 2024-05-22 === * 20:42 [[gitlab:echecs|@echecs]] was approved. === 2024-05-21 === * 13:33 [[gitlab:mbs|@mbs]] was approved. === 2024-05-19 === * 18:06 [[gitlab:ionenlaser|@ionenlaser]] was approved. === 2024-05-18 === * 23:36 [[gitlab:mdaniels5757|@mdaniels5757]] was approved. === 2024-05-17 === * 08:54 [[gitlab:grapedog|@grapedog]] was approved. === 2024-05-08 === * 19:42 [[gitlab:kelhurd|@kelhurd]] was approved. * 19:06 [[gitlab:khurd|@khurd]] was approved. === 2024-05-06 === * 19:48 [[gitlab:j3j5|@j3j5]] was approved. * 12:06 [[gitlab:tk-999|@tk-999]] was approved. === 2024-05-05 === * 22:09 [[gitlab:pppery|@pppery]] was approved. * 20:33 [[gitlab:sakretsu|@sakretsu]] was approved. * 12:12 [[gitlab:waterquark|@waterquark]] was approved. === 2024-05-04 === * 09:03 [[gitlab:multichill|@multichill]] was approved. * 07:42 [[gitlab:abaris|@abaris]] was approved. === 2024-05-03 === * 14:57 [[gitlab:maurusian|@maurusian]] was approved. === 2024-04-24 === * 05:48 [[gitlab:wolfinux|@wolfinux]] was approved. === 2024-04-23 === * 15:48 [[gitlab:dreamrimmer|@dreamrimmer]] was approved. === 2024-04-21 === * 06:51 [[gitlab:alon|@alon]] was approved. === 2024-04-17 === * 23:33 [[gitlab:derenrich|@derenrich]] was approved. === 2024-04-16 === * 17:18 [[gitlab:valcio|@valcio]] was approved. === 2024-04-14 === * 16:51 [[gitlab:wikilucas00|@wikilucas00]] was approved. === 2024-04-06 === * 12:48 [[gitlab:theprotonade|@theprotonade]] was approved. === 2024-04-02 === * 07:30 [[gitlab:bohuizhang|@bohuizhang]] was approved. === 2024-03-30 === * 13:36 [[gitlab:lpintscher|@lpintscher]] was approved. === 2024-03-26 === * 17:09 [[gitlab:eenabulele|@eenabulele]] was approved. === 2024-03-25 === * 14:27 [[gitlab:tuukka|@tuukka]] was approved. === 2024-03-24 === * 12:24 [[gitlab:firefly|@firefly]] was approved. === 2024-03-21 === * 19:33 [[gitlab:universal-omega|@universal-omega]] was approved. === 2024-03-17 === * 10:36 [[gitlab:bisel91|@bisel91]] was approved. === 2024-03-16 === * 10:09 [[gitlab:delord|@delord]] was approved. * 00:42 [[gitlab:athulvis1|@athulvis1]] was approved. === 2024-03-15 === * 19:06 [[gitlab:ignaciorodrguez|@ignaciorodrguez]] was approved. * 08:30 [[gitlab:peachey88|@peachey88]] was approved. * 06:51 [[gitlab:derick|@derick]] was approved. === 2024-03-12 === * 15:06 [[gitlab:xiaoxiao|@xiaoxiao]] was approved. === 2024-03-06 === * 13:21 [[gitlab:desianabae1|@desianabae1]] was approved. === 2024-03-05 === * 19:21 [[gitlab:ep1c|@ep1c]] was approved. * 16:33 [[gitlab:jasmine|@jasmine]] was approved. === 2024-03-02 === * 06:42 [[gitlab:potsdamlamb|@potsdamlamb]] was approved. === 2024-02-29 === * 23:18 [[gitlab:arandomname123|@arandomname123]] was approved. * 18:03 [[gitlab:baba|@baba]] was approved. * 17:48 [[gitlab:yfdyh000|@yfdyh000]] was approved. * 03:09 [[gitlab:sds|@sds]] was approved. === 2024-02-27 === * 23:33 [[gitlab:lofhi|@lofhi]] was approved. === 2024-02-15 === * 19:45 [[gitlab:gergesshamon|@gergesshamon]] was approved. === 2024-02-14 === * 14:33 [[gitlab:philipnelson99|@philipnelson99]] was approved. === 2024-02-13 === * 13:06 [[gitlab:dringsim|@dringsim]] was approved. === 2024-02-12 === * 17:36 [[gitlab:haak|@haak]] was approved. === 2024-02-05 === * 17:33 [[gitlab:qwerfjkl|@qwerfjkl]] was approved. * 17:14 [[gitlab:ahecht|@ahecht]] was approved. === 2024-02-01 === * 09:27 [[gitlab:arinaigum|@arinaigum]] was approved. * 00:15 [[gitlab:jas42|@jas42]] was approved. * 00:15 [[gitlab:edhu|@edhu]] was approved. * 00:15 [[gitlab:marnanel|@marnanel]] was approved. * 00:15 [[gitlab:ibrahemqasim|@ibrahemqasim]] was approved. * 00:15 [[gitlab:amasotti|@amasotti]] was approved. * 00:15 [[gitlab:deni|@deni]] was approved. * 00:15 [[gitlab:cyber|@cyber]] was approved. * 00:15 [[gitlab:saroj|@saroj]] was approved. === 2024-01-29 === * 21:42 [[gitlab:rgupta|@rgupta]] was approved. === 2024-01-07 === * 09:48 [[gitlab:lutrome|@lutrome]] was approved. === 2024-01-05 === * 20:48 [[gitlab:jinoytommanjaly|@jinoytommanjaly]] was approved. * 02:51 [[gitlab:braunobruno|@braunobruno]] was approved. * 01:08 [[gitlab:amorymeltzer|@amorymeltzer]] was approved. * 01:08 [[gitlab:phi22ipus|@phi22ipus]] was approved. === 2024-01-03 === * 14:45 [[gitlab:gabina|@gabina]] was approved. === 2024-01-02 === * 13:18 [[gitlab:arthurtaylor|@arthurtaylor]] was approved. === 2023-12-23 === * 00:33 [[gitlab:aram|@aram]] was approved. === 2023-12-22 === * 16:24 [[gitlab:elpitareio|@elpitareio]] was approved. === 2023-12-21 === * 00:43 [[gitlab:bsadowski1|@bsadowski1]] was approved. * 00:43 [[gitlab:ederporto|@ederporto]] was approved. * 00:43 [[gitlab:sadraiiali|@sadraiiali]] was approved. * 00:43 [[gitlab:wasp-outis|@wasp-outis]] was approved. * 00:43 [[gitlab:bodhisattwa|@bodhisattwa]] was approved. * 00:43 [[gitlab:air7538|@air7538]] was approved. * 00:43 [[gitlab:anzx|@anzx]] was approved. * 00:43 [[gitlab:tekask1903|@tekask1903]] was approved. * 00:42 [[gitlab:kiwi-0x010c|@kiwi-0x010c]] was approved. * 00:42 [[gitlab:mpaa|@mpaa]] was approved. * 00:42 [[gitlab:kutay|@kutay]] was approved. * 00:42 [[gitlab:wattmto|@wattmto]] was approved. ekhhvm73vx07ci0tiylg0ln1zcnvkxj 2450675 2450671 2026-08-23T05:30:31Z Gitlabaccountapprovalbot 37332 amidewicki was rejected. 2450675 wikitext text/x-wiki <noinclude>'''Audit log of approvals''' made by [[gitlab:gitlabaccountapprovalbot|@gitlabaccountapprovalbot]]. __NOTOC__</noinclude> === 2026-08-23 === * 05:30 "amidewicki" was rejected (pending since 2026-05-24T05:29:29.216Z). * 01:54 "westcodragen07" was rejected (pending since 2026-05-24T01:53:53.396Z). === 2026-08-22 === * 15:21 [[gitlab:laraibhasan|@laraibhasan]] was approved. === 2026-08-20 === * 17:18 [[gitlab:jatwa|@jatwa]] was approved. === 2026-08-18 === * 11:33 "wladek92" was rejected (pending since 2026-05-19T11:30:32.679Z). === 2026-08-17 === * 18:39 [[gitlab:erlanger16e|@erlanger16e]] was approved. * 18:33 "ashishkaranam" was rejected (pending since 2026-05-18T18:32:18.813Z). * 12:48 [[gitlab:lfs|@lfs]] was approved. === 2026-08-16 === * 15:42 [[gitlab:pcendrer|@pcendrer]] was approved. === 2026-08-15 === * 17:33 [[gitlab:flammablepizza|@flammablepizza]] was approved. === 2026-08-14 === * 13:00 [[gitlab:pacmand|@pacmand]] was approved. === 2026-08-12 === * 13:21 [[gitlab:doublechek|@doublechek]] was approved. === 2026-08-11 === * 14:15 [[gitlab:mmilenkovicwmf|@mmilenkovicwmf]] was approved. === 2026-08-10 === * 17:18 [[gitlab:mustafe|@mustafe]] was approved. * 10:33 "c8263a20" was rejected (pending since 2026-05-11T10:32:40.353Z). === 2026-08-09 === * 21:18 [[gitlab:iamnetx|@iamnetx]] was approved. * 20:09 "yirba" was rejected (pending since 2026-05-10T20:07:07.738Z). * 06:42 "marsam2489" was rejected (pending since 2026-05-10T06:40:46.276Z). === 2026-08-08 === * 08:15 [[gitlab:taiwaniajusto|@taiwaniajusto]] was approved. === 2026-08-07 === * 18:21 [[gitlab:gturkington|@gturkington]] was approved. * 06:36 "brianbybyby" was rejected (pending since 2026-05-08T06:36:08.459Z). === 2026-07-30 === * 20:42 "horaciocolbert" was rejected (pending since 2026-04-30T20:41:40.421Z). === 2026-07-29 === * 10:03 [[gitlab:piastu|@piastu]] was approved. * 06:33 "rafiul1" was rejected (pending since 2026-04-29T06:33:02.573Z). === 2026-07-28 === * 16:51 [[gitlab:for-each-next|@for-each-next]] was approved. * 09:21 [[gitlab:gka|@gka]] was approved. === 2026-07-27 === * 08:12 [[gitlab:cambob|@cambob]] was approved. === 2026-07-26 === * 15:39 "demansanaagmailcom" was rejected (pending since 2026-04-26T15:39:03.297Z). === 2026-07-24 === * 10:54 [[gitlab:ysogo|@ysogo]] was approved. * 09:39 [[gitlab:pankaj199|@pankaj199]] was approved. === 2026-07-23 === * 12:30 [[gitlab:fermiboson|@fermiboson]] was approved. * 09:21 [[gitlab:slashme|@slashme]] was approved. === 2026-07-22 === * 15:54 [[gitlab:lmedley|@lmedley]] was approved. * 13:24 "praveen5638" was rejected (pending since 2026-04-22T13:21:23.368Z). * 12:48 [[gitlab:cyberpower678|@cyberpower678]] was approved. * 08:30 [[gitlab:nabbegat|@nabbegat]] was approved. * 08:30 [[gitlab:plyd|@plyd]] was approved. * 07:06 "ayush8620" was rejected (pending since 2026-04-22T07:03:12.476Z). === 2026-07-21 === * 15:51 [[gitlab:panieravide|@panieravide]] was approved. * 15:33 [[gitlab:luisvilla-personal|@luisvilla-personal]] was approved. * 13:48 [[gitlab:yru|@yru]] was approved. * 13:09 [[gitlab:deevad|@deevad]] was approved. * 12:54 [[gitlab:ctdo17|@ctdo17]] was approved. * 12:36 [[gitlab:jeannenoiraud|@jeannenoiraud]] was approved. * 12:21 [[gitlab:nadiantara|@nadiantara]] was approved. * 12:18 [[gitlab:wijltcher|@wijltcher]] was approved. * 10:33 [[gitlab:nivopol|@nivopol]] was approved. * 10:30 [[gitlab:johlig|@johlig]] was approved. * 10:30 [[gitlab:majicita|@majicita]] was approved. * 10:15 [[gitlab:yongjiapeng|@yongjiapeng]] was approved. * 10:09 [[gitlab:francyskus|@francyskus]] was approved. * 09:24 [[gitlab:xanonymusx|@xanonymusx]] was approved. === 2026-07-20 === * 17:15 "leonidlednev" was rejected (pending since 2026-04-20T17:13:35.108Z). * 15:27 [[gitlab:rodrigoargenton|@rodrigoargenton]] was approved. * 05:48 "draftecho" was rejected (pending since 2026-04-20T05:48:06.953Z). === 2026-07-19 === * 12:21 [[gitlab:boivie|@boivie]] was approved. === 2026-07-18 === * 16:09 [[gitlab:pharos|@pharos]] was approved. * 15:45 [[gitlab:priyankar22|@priyankar22]] was approved. * 15:30 [[gitlab:sisyph|@sisyph]] was approved. === 2026-07-13 === * 03:45 [[gitlab:dreamyshade|@dreamyshade]] was approved. === 2026-07-12 === * 09:27 [[gitlab:smk|@smk]] was approved. === 2026-07-11 === * 14:48 "bigcereal42" was rejected (pending since 2026-04-11T14:47:40.321Z). * 12:18 "pratyushsawan" was rejected (pending since 2026-04-11T12:16:41.671Z). === 2026-07-10 === * 14:42 [[gitlab:akaza24|@akaza24]] was approved. * 12:57 [[gitlab:kormisk|@kormisk]] was approved. === 2026-07-07 === * 10:36 [[gitlab:olafjanssen|@olafjanssen]] was approved. * 06:57 "elisapoly-99" was rejected (pending since 2026-04-07T06:55:52.662Z). === 2026-07-06 === * 11:57 "ma3rouf" was rejected (pending since 2026-04-06T11:56:30.978Z). === 2026-07-02 === * 15:48 [[gitlab:tekneos|@tekneos]] was approved. === 2026-07-01 === * 15:12 [[gitlab:mugurolevy|@mugurolevy]] was approved. * 14:15 [[gitlab:vadymts1|@vadymts1]] was approved. * 09:57 "mugurolevy" was rejected (pending since 2026-04-01T09:55:19.175Z). === 2026-06-30 === * 14:27 "shivangisharma" was rejected (pending since 2026-03-31T14:26:44.932Z). === 2026-06-29 === * 19:03 [[gitlab:thisismattmiller|@thisismattmiller]] was approved. === 2026-06-28 === * 14:51 "nkwenuinadine" was rejected (pending since 2026-03-29T14:48:32.735Z). * 14:03 "vaishnavikumbhar" was rejected (pending since 2026-03-29T14:01:30.604Z). * 13:03 "stepmay" was rejected (pending since 2026-03-29T13:01:59.905Z). * 06:51 "swallroth" was rejected (pending since 2026-03-29T06:49:54.838Z). === 2026-06-26 === * 09:39 [[gitlab:lakshita28|@lakshita28]] was approved. * 07:30 [[gitlab:reeti|@reeti]] was approved. * 07:30 [[gitlab:anushka10patel|@anushka10patel]] was approved. * 07:30 "samsaesque" was rejected (pending since 2026-03-27T07:29:57.279Z). * 05:51 [[gitlab:arpithhhaaa|@arpithhhaaa]] was approved. * 05:51 [[gitlab:govindlaltl|@govindlaltl]] was approved. === 2026-06-25 === * 16:09 [[gitlab:sakuraemad|@sakuraemad]] was approved. * 07:00 "kdh8219" was rejected (pending since 2026-03-26T06:58:05.415Z). === 2026-06-24 === * 11:54 [[gitlab:sanskardubeydev|@sanskardubeydev]] was approved. * 10:09 "tanmay789q" was rejected (pending since 2026-03-25T10:07:54.602Z). === 2026-06-22 === * 19:57 [[gitlab:gouvernathor|@gouvernathor]] was approved. * 16:45 [[gitlab:lucasbelo|@lucasbelo]] was approved. * 07:15 "jason2000-cpu" was rejected (pending since 2026-03-23T07:14:09.184Z). === 2026-06-21 === * 13:18 [[gitlab:egonw|@egonw]] was approved. === 2026-06-20 === * 10:21 [[gitlab:tways2017|@tways2017]] was approved. === 2026-06-19 === * 16:06 "wilsonwang2026" was rejected (pending since 2026-03-20T16:06:05.511Z). * 04:12 [[gitlab:claudio|@claudio]] was approved. === 2026-06-18 === * 14:21 "royiswariii" was rejected (pending since 2026-03-19T14:19:16.896Z). * 13:06 [[gitlab:laurabarluzzi|@laurabarluzzi]] was approved. === 2026-06-17 === * 11:24 "adinathq8x" was rejected (pending since 2026-03-18T11:22:50.098Z). * 09:45 "nathanveritas" was rejected (pending since 2026-03-18T09:43:51.645Z). === 2026-06-15 === * 22:39 [[gitlab:mohammadhijjawi|@mohammadhijjawi]] was approved. * 14:24 "enlisar" was rejected (pending since 2026-03-16T14:23:00.109Z). * 14:06 "ayaan" was rejected (pending since 2026-03-16T14:03:31.071Z). * 10:54 "kwametech" was rejected (pending since 2026-03-16T10:54:11.083Z). === 2026-06-14 === * 17:45 [[gitlab:surajseth520|@surajseth520]] was approved. * 07:24 "malahimhaseeb" was rejected (pending since 2026-03-15T07:21:57.748Z). === 2026-06-11 === * 11:48 [[gitlab:cadddr|@cadddr]] was approved. * 11:18 "wikipiggy" was rejected (pending since 2026-03-12T11:16:09.335Z). * 07:15 [[gitlab:vesihiisi|@vesihiisi]] was approved. === 2026-06-10 === * 07:03 [[gitlab:dmiranda|@dmiranda]] was approved. === 2026-06-09 === * 14:21 [[gitlab:linkgenetic|@linkgenetic]] was approved. * 14:03 [[gitlab:sjones-ctr|@sjones-ctr]] was approved. * 12:51 [[gitlab:ekrem|@ekrem]] was approved. === 2026-06-08 === * 17:48 "jmprax" was rejected (pending since 2026-03-09T17:46:38.807Z). * 16:15 [[gitlab:ahonc|@ahonc]] was approved. * 12:57 [[gitlab:rainmonger|@rainmonger]] was approved. === 2026-06-07 === * 23:03 "shadowthewuff" was rejected (pending since 2026-03-08T23:00:53.442Z). * 11:45 "wiki-pavan" was rejected (pending since 2026-03-08T11:45:11.116Z). * 02:39 [[gitlab:launchpad|@launchpad]] was approved. === 2026-06-06 === * 14:54 "unicord" was rejected (pending since 2026-03-07T14:52:04.992Z). * 12:48 "chien" was rejected (pending since 2026-03-07T12:48:11.669Z). === 2026-06-04 === * 14:33 "only-vikas" was rejected (pending since 2026-03-05T14:32:09.186Z). === 2026-06-03 === * 15:00 [[gitlab:anafibnshahibul|@anafibnshahibul]] was approved. === 2026-06-02 === * 21:21 "mgagat" was rejected (pending since 2026-03-03T21:18:37.223Z). * 13:57 "prasunaenumarthy" was rejected (pending since 2026-03-03T13:57:14.847Z). * 05:48 [[gitlab:tmoney|@tmoney]] was approved. === 2026-06-01 === * 14:57 "vikram2101" was rejected (pending since 2026-03-02T14:54:26.550Z). * 12:03 "watshell" was rejected (pending since 2026-03-02T12:03:09.329Z). === 2026-05-29 === * 12:48 "mounikapotladurthi" was rejected (pending since 2026-02-27T12:45:38.609Z). === 2026-05-27 === * 20:00 "vinitha" was rejected (pending since 2026-02-25T19:58:43.524Z). * 16:30 "codeurluce" was rejected (pending since 2026-02-25T16:28:53.973Z). * 14:33 [[gitlab:thilio|@thilio]] was approved. === 2026-05-26 === * 12:09 "charisad" was rejected (pending since 2026-02-24T12:07:21.881Z). === 2026-05-25 === * 22:54 "ddshelto" was rejected (pending since 2026-02-23T22:52:44.427Z). * 19:51 "lakz-99" was rejected (pending since 2026-02-23T19:47:00.263Z). * 19:48 "lakz-99" was rejected (pending since 2026-02-23T19:47:00.263Z). === 2026-05-24 === * 18:45 "jiyagupta-cs" was rejected (pending since 2026-02-22T18:43:33.176Z). === 2026-05-23 === * 13:09 [[gitlab:gauthammohanraj|@gauthammohanraj]] was approved. * 04:21 [[gitlab:staraction|@staraction]] was approved. === 2026-05-22 === * 19:03 "i-horich" was rejected (pending since 2026-02-20T19:00:43.519Z). * 01:48 "50323233" was rejected (pending since 2026-02-20T01:48:05.555Z). === 2026-05-21 === * 18:51 "kartikeyg0104" was rejected (pending since 2026-02-19T18:48:39.707Z). * 16:27 [[gitlab:renovatebot|@renovatebot]] was approved. * 16:06 [[gitlab:gkm563|@gkm563]] was approved. === 2026-05-20 === * 01:21 "beedellrokejulianlockhart" was rejected (pending since 2026-02-18T01:19:13.284Z). === 2026-05-18 === * 23:18 "wladek92" was rejected (pending since 2026-02-16T23:16:22.939Z). * 16:36 [[gitlab:effeietsanders|@effeietsanders]] was approved. === 2026-05-14 === * 21:00 [[gitlab:nehemienathan|@nehemienathan]] was approved. === 2026-05-13 === * 10:51 "ssssaaaa" was rejected (pending since 2026-02-11T10:50:36.975Z). === 2026-05-12 === * 18:06 [[gitlab:psubhashish|@psubhashish]] was approved. * 08:12 "khan" was rejected (pending since 2026-02-10T08:11:48.776Z). * 04:27 "galaxysh" was rejected (pending since 2026-02-10T04:24:59.440Z). === 2026-05-11 === * 12:18 "peterxy12" was rejected (pending since 2026-02-09T12:18:01.982Z). === 2026-05-10 === * 11:09 "yalihupokn" was rejected (pending since 2026-02-08T11:06:51.336Z). * 05:12 "wobadha" was rejected (pending since 2026-02-08T05:11:00.569Z). === 2026-05-09 === * 13:45 "bwiki" was rejected (pending since 2026-02-07T13:43:38.177Z). === 2026-05-08 === * 09:24 [[gitlab:cwilliams|@cwilliams]] was approved. === 2026-05-07 === * 14:15 "rehankhan78" was rejected (pending since 2026-02-05T14:13:37.754Z). === 2026-05-06 === * 11:24 "ari" was rejected (pending since 2026-02-04T11:24:11.760Z). * 08:09 [[gitlab:neriah|@neriah]] was approved. * 06:27 [[gitlab:status401|@status401]] was approved. === 2026-05-03 === * 09:54 [[gitlab:anilk|@anilk]] was approved. === 2026-05-02 === * 17:54 [[gitlab:sweil|@sweil]] was approved. * 17:00 [[gitlab:aoppo|@aoppo]] was approved. === 2026-05-01 === * 21:18 [[gitlab:dawalda|@dawalda]] was approved. === 2026-04-30 === * 21:42 "merohibine" was rejected (pending since 2026-01-29T21:40:00.756Z). * 20:54 [[gitlab:tfmorris|@tfmorris]] was approved. * 17:33 [[gitlab:uyen|@uyen]] was approved. * 07:39 [[gitlab:mahveotm|@mahveotm]] was approved. * 06:36 [[gitlab:leo321|@leo321]] was approved. === 2026-04-29 === * 02:27 [[gitlab:dw31415|@dw31415]] was approved. === 2026-04-28 === * 23:09 [[gitlab:dtorsani|@dtorsani]] was approved. === 2026-04-27 === * 23:42 [[gitlab:quinlan|@quinlan]] was approved. * 05:00 [[gitlab:matthewyeager|@matthewyeager]] was approved. === 2026-04-26 === * 17:36 "kuba-hajnej" was rejected (pending since 2026-01-25T17:33:32.467Z). * 13:03 "jklamo" was rejected (pending since 2026-01-25T13:02:22.936Z). === 2026-04-25 === * 20:24 [[gitlab:maldaxura|@maldaxura]] was approved. * 14:33 [[gitlab:sirtobi|@sirtobi]] was approved. * 04:18 "ice5678" was rejected (pending since 2026-01-24T04:15:30.008Z). === 2026-04-24 === * 22:06 [[gitlab:arcstur|@arcstur]] was approved. === 2026-04-22 === * 23:06 "dtorsani" was rejected (pending since 2026-01-21T23:03:25.843Z). * 22:18 [[gitlab:egezort|@egezort]] was approved. * 16:45 "nexpectarpit" was rejected (pending since 2026-01-21T16:43:21.045Z). === 2026-04-20 === * 19:15 "fitch" was rejected (pending since 2026-01-19T19:12:35.644Z). === 2026-04-19 === * 02:54 [[gitlab:neoact|@neoact]] was approved. === 2026-04-18 === * 07:06 [[gitlab:kockaadmiralac|@kockaadmiralac]] was approved. === 2026-04-17 === * 13:42 "liselot" was rejected (pending since 2026-01-16T13:39:41.909Z). === 2026-04-15 === * 17:03 "lahari" was rejected (pending since 2026-01-14T17:02:06.275Z). === 2026-04-14 === * 13:00 "surajseth520" was rejected (pending since 2026-01-13T12:59:45.906Z). * 04:51 [[gitlab:canley|@canley]] was approved. * 01:03 "bshizzle" was rejected (pending since 2026-01-13T01:00:48.120Z). === 2026-04-13 === * 15:30 [[gitlab:passimacopoulos|@passimacopoulos]] was approved. === 2026-04-11 === * 12:30 "krithash" was rejected (pending since 2026-01-10T12:27:24.731Z). === 2026-04-10 === * 15:30 "raunak1709" was rejected (pending since 2026-01-09T15:29:10.901Z). === 2026-04-07 === * 17:03 [[gitlab:supnabla|@supnabla]] was approved. === 2026-04-06 === * 20:00 [[gitlab:laerdon|@laerdon]] was approved. * 19:21 [[gitlab:ljq3|@ljq3]] was approved. === 2026-04-04 === * 11:06 "mixcc" was rejected (pending since 2026-01-03T11:03:33.922Z). === 2026-04-02 === * 05:30 [[gitlab:mbh1|@mbh1]] was approved. === 2026-04-01 === * 18:21 "yuvrajpatil17" was rejected (pending since 2025-12-31T18:20:27.991Z). * 12:12 [[gitlab:amorii0|@amorii0]] was approved. === 2026-03-31 === * 11:00 "krrishsehgal" was rejected (pending since 2025-12-30T11:00:16.384Z). === 2026-03-30 === * 15:36 [[gitlab:atsuko|@atsuko]] was approved. === 2026-03-29 === * 11:36 [[gitlab:giftcup|@giftcup]] was approved. === 2026-03-28 === * 14:51 [[gitlab:janeeva1|@janeeva1]] was approved. === 2026-03-26 === * 13:36 [[gitlab:saiphani02|@saiphani02]] was approved. * 11:48 [[gitlab:valerioboz-wmch|@valerioboz-wmch]] was approved. === 2026-03-25 === * 09:45 "quansi" was rejected (pending since 2025-12-24T09:42:13.451Z). * 02:18 [[gitlab:viztor|@viztor]] was approved. === 2026-03-24 === * 23:18 [[gitlab:maryyann|@maryyann]] was approved. * 23:01 [[gitlab:codenamenoreste|@codenamenoreste]] was approved. * 13:36 [[gitlab:marc-maillard-wmse|@marc-maillard-wmse]] was approved. * 07:39 "fred2675" was rejected (pending since 2025-12-23T07:39:11.380Z). === 2026-03-23 === * 14:51 [[gitlab:komla|@komla]] was approved. * 05:51 "lunachuck43" was rejected (pending since 2025-12-22T05:50:17.862Z). * 04:06 "reza110011" was rejected (pending since 2025-12-22T04:05:25.117Z). === 2026-03-20 === * 21:54 "mertgor" was rejected (pending since 2025-12-19T21:51:51.419Z). * 20:57 "autanmahmah" was rejected (pending since 2025-12-19T20:54:51.678Z). * 09:57 [[gitlab:nethahussain|@nethahussain]] was approved. * 09:27 [[gitlab:piewriter|@piewriter]] was approved. * 08:15 [[gitlab:dondersmooi|@dondersmooi]] was approved. === 2026-03-19 === * 21:03 "sayvhior" was rejected (pending since 2025-12-18T21:02:31.699Z). === 2026-03-18 === * 20:15 [[gitlab:martinmystere|@martinmystere]] was approved. === 2026-03-17 === * 02:51 "louperivois" was rejected (pending since 2025-12-16T02:50:48.197Z). === 2026-03-16 === * 12:54 "mokayaj857" was rejected (pending since 2025-12-15T12:53:39.015Z). * 06:18 "roamer15" was rejected (pending since 2025-12-15T06:16:38.042Z). === 2026-03-14 === * 11:12 "umaramuhammad" was rejected (pending since 2025-12-13T11:10:44.004Z). * 09:33 "akuma19" was rejected (pending since 2025-12-13T09:31:39.044Z). * 07:06 [[gitlab:syunsyunminmin|@syunsyunminmin]] was approved. === 2026-03-12 === * 20:24 [[gitlab:11wb|@11wb]] was approved. * 09:54 [[gitlab:bcxfu75k|@bcxfu75k]] was approved. === 2026-03-10 === * 09:12 [[gitlab:viktoriahillerudwmse|@viktoriahillerudwmse]] was approved. === 2026-03-06 === * 08:09 "vazhayilnewone" was rejected (pending since 2025-12-05T08:07:02.184Z). === 2026-03-04 === * 20:54 [[gitlab:elphie|@elphie]] was approved. * 11:39 "ronaldahmed" was rejected (pending since 2025-12-03T11:37:47.492Z). * 02:12 "ltslw" was rejected (pending since 2025-12-03T02:11:52.040Z). === 2026-03-02 === * 19:21 "dlopez350" was rejected (pending since 2025-12-01T19:20:38.918Z). * 18:15 [[gitlab:lsandergreen|@lsandergreen]] was approved. === 2026-03-01 === * 10:51 [[gitlab:clintacc|@clintacc]] was approved. === 2026-02-28 === * 09:24 "cardboardlamp" was rejected (pending since 2025-11-29T09:22:03.947Z). * 08:18 "wiki-pavan" was rejected (pending since 2025-11-29T08:16:24.184Z). === 2026-02-27 === * 20:45 "thisisrick25" was rejected (pending since 2025-11-28T20:42:24.454Z). === 2026-02-26 === * 13:57 "chuiimuiiofc" was rejected (pending since 2025-11-27T13:57:02.794Z). * 13:54 "steffpro" was rejected (pending since 2025-11-27T13:52:10.859Z). === 2026-02-25 === * 21:24 "abubakarhabibudayyabu" was rejected (pending since 2025-11-26T21:22:37.776Z). === 2026-02-24 === * 05:00 "playboi" was rejected (pending since 2025-11-25T05:00:30.762Z). === 2026-02-23 === * 14:00 "alph65" was rejected (pending since 2025-11-24T13:59:00.797Z). * 12:33 [[gitlab:robertsky|@robertsky]] was approved. === 2026-02-22 === * 00:30 "hp8p" was rejected (pending since 2025-11-23T00:29:24.741Z). === 2026-02-19 === * 16:45 "clayjar" was rejected (pending since 2025-11-20T16:44:48.380Z). === 2026-02-18 === * 22:18 "nexus" was rejected (pending since 2025-11-19T22:16:48.818Z). * 12:00 "bernsteinnn" was rejected (pending since 2025-11-19T11:59:04.427Z). === 2026-02-17 === * 11:36 "jason2000-cpu" was rejected (pending since 2025-11-18T11:34:00.314Z). === 2026-02-16 === * 14:54 "smaurya" was rejected (pending since 2025-11-17T14:52:06.906Z). === 2026-02-15 === * 16:51 "kra-79" was rejected (pending since 2025-11-16T16:50:41.375Z). === 2026-02-14 === * 15:15 [[gitlab:mess|@mess]] was approved. === 2026-02-13 === * 13:57 "sopalsuemae957" was rejected (pending since 2025-11-14T13:55:16.921Z). * 13:30 [[gitlab:wyslijp16-toolforge|@wyslijp16-toolforge]] was approved. === 2026-02-12 === * 16:30 "kristinagligoric" was rejected (pending since 2025-11-13T16:29:21.646Z). * 03:33 [[gitlab:anyehansen|@anyehansen]] was approved. * 02:21 [[gitlab:thejoyfultentmaker|@thejoyfultentmaker]] was approved. === 2026-02-10 === * 13:18 [[gitlab:db111|@db111]] was approved. === 2026-02-09 === * 19:06 "squirrel289" was rejected (pending since 2025-11-10T19:04:27.831Z). === 2026-02-06 === * 20:54 [[gitlab:gillux|@gillux]] was approved. * 09:09 [[gitlab:lih|@lih]] was approved. === 2026-01-31 === * 16:21 [[gitlab:taxonbot1|@taxonbot1]] was approved. === 2026-01-28 === * 14:30 [[gitlab:ademola|@ademola]] was approved. * 10:51 "watshell" was rejected (pending since 2025-10-29T10:51:01.521Z). === 2026-01-26 === * 23:06 "tavaresgmg" was rejected (pending since 2025-10-27T23:04:42.140Z). === 2026-01-25 === * 06:03 "cata" was rejected (pending since 2025-10-26T06:01:26.155Z). === 2026-01-24 === * 21:15 [[gitlab:wiegels|@wiegels]] was approved. * 06:30 [[gitlab:blaquans|@blaquans]] was approved. === 2026-01-23 === * 16:27 [[gitlab:lerickson|@lerickson]] was approved. * 10:15 "fran0035g" was rejected (pending since 2025-10-24T10:12:17.732Z). === 2026-01-22 === * 21:00 "hacksyn" was rejected (pending since 2025-10-23T20:59:15.982Z). === 2026-01-21 === * 17:30 [[gitlab:otcenas11|@otcenas11]] was approved. === 2026-01-19 === * 21:48 [[gitlab:amdrel|@amdrel]] was approved. * 04:36 "rayalexa" was rejected (pending since 2025-10-20T04:35:02.094Z). === 2026-01-18 === * 15:45 "somya" was rejected (pending since 2025-10-19T15:43:43.701Z). * 06:54 "sergg001" was rejected (pending since 2025-10-19T06:54:12.296Z). === 2026-01-16 === * 11:57 "zeejohsy" was rejected (pending since 2025-10-17T11:56:22.372Z). * 04:45 "rocky25" was rejected (pending since 2025-10-17T04:43:33.180Z). === 2026-01-15 === * 16:39 "tiisu" was rejected (pending since 2025-10-16T16:37:18.438Z). * 12:00 "noahalorwu" was rejected (pending since 2025-10-16T11:58:26.133Z). * 10:39 "prjayaiuedu" was rejected (pending since 2025-10-16T10:37:16.947Z). === 2026-01-13 === * 17:21 [[gitlab:lwilson-ctr|@lwilson-ctr]] was approved. === 2026-01-12 === * 17:03 "stagietechs" was rejected (pending since 2025-10-13T17:02:25.281Z). === 2026-01-10 === * 19:06 "keerthisr" was rejected (pending since 2025-10-11T19:05:01.758Z). === 2026-01-09 === * 20:36 "lightb" was rejected (pending since 2025-10-10T20:34:20.264Z). === 2026-01-08 === * 19:42 [[gitlab:tbodt|@tbodt]] was approved. * 13:57 [[gitlab:martynranyard|@martynranyard]] was approved. === 2026-01-07 === * 17:48 [[gitlab:santanuwiki25|@santanuwiki25]] was approved. * 14:27 "dipanshu" was rejected (pending since 2025-10-08T14:26:10.794Z). * 12:30 "adeolaadesina" was rejected (pending since 2025-10-08T12:29:49.592Z). * 09:21 "tony-kamande" was rejected (pending since 2025-10-08T09:20:28.421Z). * 06:18 "hninwuttyi" was rejected (pending since 2025-10-08T06:17:28.006Z). * 05:09 "andume" was rejected (pending since 2025-10-08T05:07:18.582Z). * 02:00 "mosope" was rejected (pending since 2025-10-08T01:59:54.800Z). * 01:15 [[gitlab:tungstalite|@tungstalite]] was approved. === 2026-01-06 === * 18:24 "leerensucher" was rejected (pending since 2025-10-07T18:21:41.253Z). * 14:54 "leonidlednev" was rejected (pending since 2025-10-07T14:53:07.273Z). * 12:57 "alexandre-tingaud" was rejected (pending since 2025-10-07T12:54:27.206Z). === 2026-01-04 === * 21:33 [[gitlab:matr1x-101|@matr1x-101]] was approved. * 15:18 "makjr" was rejected (pending since 2025-10-05T15:16:31.558Z). * 14:09 "dakshq" was rejected (pending since 2025-10-05T14:08:40.608Z). === 2026-01-03 === * 20:42 [[gitlab:apehitkey|@apehitkey]] was approved. * 18:00 [[gitlab:jeremyb|@jeremyb]] was approved. * 14:09 [[gitlab:twelephant|@twelephant]] was approved. === 2026-01-01 === * 11:30 "shellstanislav" was rejected (pending since 2025-10-02T11:29:10.150Z). === 2025-12-30 === * 19:51 "camilojdiaz" was rejected (pending since 2025-09-30T19:49:24.913Z). === 2025-12-29 === * 16:03 "zied" was rejected (pending since 2025-09-29T16:01:30.415Z). * 08:18 "rahulsidpradhan" was rejected (pending since 2025-09-29T08:17:02.849Z). === 2025-12-26 === * 09:48 "thembo42" was rejected (pending since 2025-09-26T09:45:15.033Z). === 2025-12-25 === * 14:03 "196936074751" was rejected (pending since 2025-09-25T14:02:31.367Z). === 2025-12-23 === * 16:21 "ngarnsworthy" was rejected (pending since 2025-09-23T16:20:41.211Z). === 2025-12-22 === * 12:39 "aza555" was rejected (pending since 2025-09-22T12:38:02.622Z). === 2025-12-20 === * 23:45 "saph" was rejected (pending since 2025-09-20T23:45:01.222Z). === 2025-12-19 === * 10:15 "vladdymoses" was rejected (pending since 2025-09-19T10:15:00.999Z). * 07:15 "dirtylittlepoobah" was rejected (pending since 2025-09-19T07:13:55.537Z). === 2025-12-18 === * 16:24 [[gitlab:guyfawcus|@guyfawcus]] was approved. === 2025-12-17 === * 21:39 [[gitlab:holdyourhorses|@holdyourhorses]] was approved. * 18:30 "prudencia" was rejected (pending since 2025-09-17T18:27:18.860Z). * 02:24 "lottie" was rejected (pending since 2025-09-17T02:21:21.744Z). === 2025-12-16 === * 09:39 [[gitlab:melcatherine|@melcatherine]] was approved. * 08:54 [[gitlab:leila237|@leila237]] was approved. === 2025-12-15 === * 18:27 [[gitlab:royalsailor|@royalsailor]] was approved. * 09:39 [[gitlab:olaf8940|@olaf8940]] was approved. * 09:39 "brianbybyby" was rejected (pending since 2025-09-15T09:37:45.430Z). === 2025-12-14 === * 20:21 [[gitlab:essa237|@essa237]] was approved. * 16:42 [[gitlab:bovimacoco|@bovimacoco]] was approved. === 2025-12-13 === * 21:54 "mmns21" was rejected (pending since 2025-09-13T21:52:24.017Z). * 20:33 "bugcrawler" was rejected (pending since 2025-09-13T20:31:09.211Z). === 2025-12-12 === * 14:39 "ruvchoudhary" was rejected (pending since 2025-09-12T14:36:16.167Z). * 06:54 "rezadress" was rejected (pending since 2025-09-12T06:52:21.749Z). === 2025-12-10 === * 17:30 [[gitlab:itsmoon|@itsmoon]] was approved. === 2025-12-09 === * 15:42 [[gitlab:mercy-o|@mercy-o]] was approved. === 2025-12-06 === * 16:45 "jacquesradjabu" was rejected (pending since 2025-09-06T16:45:17.969Z). * 11:27 [[gitlab:ikhitron|@ikhitron]] was approved. === 2025-12-01 === * 08:12 "halconmilenario21" was rejected (pending since 2025-09-01T08:12:10.262Z). === 2025-11-30 === * 21:06 [[gitlab:habs|@habs]] was approved. === 2025-11-29 === * 16:36 "bovimacoco" was rejected (pending since 2025-08-30T16:34:39.712Z). * 00:45 [[gitlab:jjpmaster|@jjpmaster]] was approved. === 2025-11-24 === * 10:30 "alph65" was rejected (pending since 2025-08-25T10:28:40.957Z). * 02:24 [[gitlab:yaron|@yaron]] was approved. === 2025-11-20 === * 16:06 "clayjar" was rejected (pending since 2025-08-21T16:04:54.450Z). === 2025-11-17 === * 21:09 [[gitlab:ankita97531|@ankita97531]] was approved. === 2025-11-16 === * 14:15 "commanderkefir" was rejected (pending since 2025-08-17T14:13:14.791Z). * 08:21 "rehankhan78" was rejected (pending since 2025-08-17T08:19:44.896Z). === 2025-11-15 === * 14:36 "cyberscribe" was rejected (pending since 2025-08-16T14:34:27.230Z). === 2025-11-13 === * 04:21 "waddie96" was rejected (pending since 2025-08-14T04:19:27.461Z). === 2025-11-11 === * 06:42 [[gitlab:seanhoyland|@seanhoyland]] was approved. === 2025-11-10 === * 00:06 [[gitlab:jaredblumer|@jaredblumer]] was approved. === 2025-11-09 === * 22:36 "heinxiety" was rejected (pending since 2025-08-10T22:33:12.041Z). === 2025-11-07 === * 22:00 [[gitlab:forzagreen|@forzagreen]] was approved. === 2025-11-06 === * 16:57 [[gitlab:rsilvola|@rsilvola]] was approved. === 2025-11-04 === * 21:24 [[gitlab:devdoingdev|@devdoingdev]] was approved. === 2025-11-03 === * 17:48 "joewaleed98" was rejected (pending since 2025-08-04T17:46:12.191Z). === 2025-11-01 === * 18:00 "eliasempresas" was rejected (pending since 2025-08-02T17:58:04.412Z). === 2025-10-31 === * 18:51 [[gitlab:chaoticenby|@chaoticenby]] was approved. * 04:33 "3ch310n" was rejected (pending since 2025-08-01T04:32:21.982Z). === 2025-10-30 === * 10:03 [[gitlab:tausheefhassan|@tausheefhassan]] was approved. === 2025-10-29 === * 14:54 "theap" was rejected (pending since 2025-07-30T14:52:12.066Z). === 2025-10-28 === * 06:06 [[gitlab:tanbiruzzaman|@tanbiruzzaman]] was approved. === 2025-10-27 === * 07:51 [[gitlab:jmoore111|@jmoore111]] was approved. === 2025-10-25 === * 21:09 [[gitlab:valor|@valor]] was approved. * 21:03 [[gitlab:booksmurf|@booksmurf]] was approved. * 02:48 "mystyc1" was rejected (pending since 2025-07-26T02:46:19.373Z). === 2025-10-24 === * 05:12 "aadarshmahesh" was rejected (pending since 2025-07-25T05:09:38.264Z). === 2025-10-22 === * 20:54 [[gitlab:janewanga|@janewanga]] was approved. * 17:27 "abeljeevan" was rejected (pending since 2025-07-23T17:26:46.884Z). * 16:12 "shrimpnaur" was rejected (pending since 2025-07-23T16:10:37.864Z). === 2025-10-21 === * 18:51 "jrmuizel" was rejected (pending since 2025-07-22T18:50:07.315Z). * 09:33 [[gitlab:dpogorzelski|@dpogorzelski]] was approved. === 2025-10-17 === * 13:21 [[gitlab:blegodwin|@blegodwin]] was approved. === 2025-10-16 === * 14:51 [[gitlab:bahago|@bahago]] was approved. * 14:12 "harikrishna0005" was rejected (pending since 2025-07-17T14:10:48.385Z). * 14:09 "gauthammohanraj" was rejected (pending since 2025-07-17T14:08:47.643Z). === 2025-10-15 === * 13:48 [[gitlab:adwivedii|@adwivedii]] was approved. * 13:18 [[gitlab:kimbrenekakande|@kimbrenekakande]] was approved. * 13:03 "childmnajennifer" was rejected (pending since 2025-07-16T13:01:50.236Z). * 05:06 "vssb4214" was rejected (pending since 2025-07-16T05:05:33.985Z). === 2025-10-14 === * 19:39 [[gitlab:afanyulionel|@afanyulionel]] was approved. * 15:33 [[gitlab:sadrettin|@sadrettin]] was approved. * 14:18 [[gitlab:tmwyk|@tmwyk]] was approved. * 08:42 "yasu0796" was rejected (pending since 2025-07-15T08:41:26.453Z). === 2025-10-13 === * 16:09 [[gitlab:atlas0007|@atlas0007]] was approved. === 2025-10-11 === * 17:42 [[gitlab:techwizzie|@techwizzie]] was approved. === 2025-10-10 === * 19:03 [[gitlab:miiswom|@miiswom]] was approved. * 16:06 [[gitlab:ninatakang|@ninatakang]] was approved. === 2025-10-09 === * 15:42 [[gitlab:jaykaneki|@jaykaneki]] was approved. * 14:21 [[gitlab:lebogang|@lebogang]] was approved. * 14:15 [[gitlab:kimondorose|@kimondorose]] was approved. * 13:48 [[gitlab:joyakinyi|@joyakinyi]] was approved. * 13:48 [[gitlab:dikshyashahi|@dikshyashahi]] was approved. * 13:45 [[gitlab:obediobadiah|@obediobadiah]] was approved. * 13:45 [[gitlab:system625|@system625]] was approved. * 13:45 [[gitlab:rolalove|@rolalove]] was approved. * 13:39 [[gitlab:olatundeawo|@olatundeawo]] was approved. * 13:36 [[gitlab:danielchristlight|@danielchristlight]] was approved. * 13:36 [[gitlab:dipanshu1223|@dipanshu1223]] was approved. * 13:36 [[gitlab:aradhya|@aradhya]] was approved. * 09:57 "bognd" was rejected (pending since 2025-07-10T09:55:48.661Z). === 2025-10-08 === * 23:36 [[gitlab:sopzy|@sopzy]] was approved. * 23:03 [[gitlab:oluwatumininu|@oluwatumininu]] was approved. * 19:39 [[gitlab:levon003|@levon003]] was approved. * 15:24 [[gitlab:ritika-bhambri11|@ritika-bhambri11]] was approved. * 13:45 [[gitlab:anbanguyen|@anbanguyen]] was approved. * 13:36 [[gitlab:chumzine|@chumzine]] was approved. * 13:27 [[gitlab:shr0x-ya|@shr0x-ya]] was approved. * 12:45 [[gitlab:nurahwakili|@nurahwakili]] was approved. * 03:42 "nazhiba" was rejected (pending since 2025-07-09T03:40:12.625Z). * 02:12 "mafennel" was rejected (pending since 2025-07-09T02:11:40.598Z). === 2025-10-07 === * 22:54 [[gitlab:olusegunfaj|@olusegunfaj]] was approved. * 21:30 [[gitlab:rona|@rona]] was approved. * 21:09 [[gitlab:sandijigs|@sandijigs]] was approved. * 13:36 "xisbajao" was rejected (pending since 2025-07-08T13:33:35.018Z). * 01:36 "areczek94" was rejected (pending since 2025-07-08T01:35:40.633Z). === 2025-10-06 === * 19:21 "wmcarter2017" was rejected (pending since 2025-07-07T19:21:12.899Z). === 2025-10-05 === * 14:15 "meetmendapara" was rejected (pending since 2025-07-06T14:14:16.726Z). === 2025-10-04 === * 20:51 "nftbaee" was rejected (pending since 2025-07-05T20:50:57.688Z). === 2025-10-03 === * 06:12 [[gitlab:javiermonton|@javiermonton]] was approved. === 2025-10-02 === * 20:15 "talaqalotaibipmp" was rejected (pending since 2025-07-03T20:13:05.164Z). === 2025-10-01 === * 10:54 "bjensen" was rejected (pending since 2025-07-02T10:53:46.574Z). * 02:45 "kowal1984" was rejected (pending since 2025-07-02T02:44:56.946Z). === 2025-09-30 === * 21:21 [[gitlab:kavaljeetsingh|@kavaljeetsingh]] was approved. * 00:24 "adium" was rejected (pending since 2025-07-01T00:23:43.807Z). === 2025-09-28 === * 08:54 [[gitlab:pexerik|@pexerik]] was approved. === 2025-09-27 === * 13:57 [[gitlab:rubahhitamvukova|@rubahhitamvukova]] was approved. === 2025-09-26 === * 16:57 "algorithmic" was rejected (pending since 2025-06-27T16:56:17.480Z). * 13:54 [[gitlab:shadabgdg|@shadabgdg]] was approved. * 13:12 [[gitlab:spushpit|@spushpit]] was approved. === 2025-09-20 === * 14:06 "bwiki" was rejected (pending since 2025-06-21T13:59:14.749Z). === 2025-09-16 === * 05:39 [[gitlab:deepchirp|@deepchirp]] was approved. === 2025-09-15 === * 22:00 [[gitlab:noisk8|@noisk8]] was approved. * 11:03 "ahonc" was rejected (pending since 2025-06-16T11:00:54.843Z). === 2025-09-13 === * 18:24 "a-ssh22" was rejected (pending since 2025-06-14T18:23:33.937Z). * 12:36 [[gitlab:rajashreetalukdar|@rajashreetalukdar]] was approved. * 00:45 [[gitlab:sumitsurai|@sumitsurai]] was approved. === 2025-09-12 === * 17:12 [[gitlab:suyash23|@suyash23]] was approved. * 00:46 "remotetravel" was rejected (pending since 2025-06-13T00:44:08.171Z). === 2025-09-10 === * 21:09 "jancborchardt" was rejected (pending since 2025-06-11T21:06:30.759Z). === 2025-09-09 === * 17:03 [[gitlab:vwf|@vwf]] was approved. * 06:36 [[gitlab:cactusisme|@cactusisme]] was approved. === 2025-09-08 === * 18:09 "birushandegeya" was rejected (pending since 2025-06-09T18:08:00.087Z). * 16:27 "ngarnsworthy" was rejected (pending since 2025-06-09T16:24:37.213Z). * 12:33 "zolgoyo" was rejected (pending since 2025-06-09T12:31:34.199Z). === 2025-09-06 === * 23:09 [[gitlab:jaishsingh913|@jaishsingh913]] was approved. === 2025-09-05 === * 21:45 [[gitlab:sakshi2|@sakshi2]] was approved. * 20:42 "abdukhaliq1" was rejected (pending since 2025-06-06T20:40:42.023Z). * 14:27 "beubsamy" was rejected (pending since 2025-06-06T14:27:06.781Z). === 2025-09-04 === * 23:27 "sdhehua" was rejected (pending since 2025-06-05T23:24:45.777Z). * 19:00 [[gitlab:perry|@perry]] was approved. * 11:24 "saintwolf" was rejected (pending since 2025-06-05T11:21:20.176Z). === 2025-09-02 === * 05:48 [[gitlab:aliu|@aliu]] was approved. === 2025-08-29 === * 13:30 "kksurendran066" was rejected (pending since 2025-05-30T13:27:48.755Z). === 2025-08-28 === * 22:18 "tauraamuix" was rejected (pending since 2025-05-29T22:16:08.228Z). === 2025-08-26 === * 19:03 [[gitlab:dikkulah|@dikkulah]] was approved. === 2025-08-22 === * 23:51 [[gitlab:khoroshun_mike|@khoroshun_mike]] was approved. === 2025-08-21 === * 07:39 [[gitlab:yuka|@yuka]] was approved. === 2025-08-19 === * 07:48 [[gitlab:zhaofjx|@zhaofjx]] was approved. === 2025-08-17 === * 14:27 "madhan13k" was rejected (pending since 2025-05-18T14:26:08.973Z). === 2025-08-15 === * 10:15 "mohammed_abukhadra" was rejected (pending since 2025-05-16T10:14:48.403Z). === 2025-08-11 === * 11:48 "hmmyesbro" was rejected (pending since 2025-05-12T11:45:24.350Z). === 2025-08-10 === * 13:15 [[gitlab:dactyl|@dactyl]] was approved. === 2025-08-09 === * 04:39 "xxxx100000" was rejected (pending since 2025-05-10T04:37:44.949Z). === 2025-08-08 === * 14:33 [[gitlab:josefanthony|@josefanthony]] was approved. === 2025-08-07 === * 23:42 [[gitlab:robins7|@robins7]] was approved. * 21:42 [[gitlab:pols12|@pols12]] was approved. * 17:15 "sbronson" was rejected (pending since 2025-05-08T17:15:08.834Z). * 14:57 [[gitlab:alvindulle|@alvindulle]] was approved. * 14:45 [[gitlab:xentos|@xentos]] was approved. * 06:27 "jamesboste" was rejected (pending since 2025-05-08T06:25:14.793Z). * 03:57 "ysun" was rejected (pending since 2025-05-08T03:55:07.348Z). === 2025-08-06 === * 21:51 "pols12" was rejected (pending since 2025-05-07T21:49:13.598Z). * 01:51 "okeamah" was rejected (pending since 2025-05-07T01:48:50.114Z). === 2025-08-05 === * 09:15 "mobashir-2013" was rejected (pending since 2025-05-06T09:14:24.069Z). === 2025-08-01 === * 08:00 "douginamug" was rejected (pending since 2025-05-02T07:57:38.317Z). === 2025-07-31 === * 02:30 [[gitlab:ads|@ads]] was approved. === 2025-07-27 === * 13:15 "mrico2703" was rejected (pending since 2025-04-27T13:13:12.346Z). * 10:17 [[gitlab:josephfrancis12|@josephfrancis12]] was approved. * 10:17 [[gitlab:fuzzew|@fuzzew]] was approved. * 05:57 [[gitlab:biscuitbobby|@biscuitbobby]] was approved. * 05:48 [[gitlab:ecoholic|@ecoholic]] was approved. === 2025-07-26 === * 11:48 [[gitlab:chimnayyyy|@chimnayyyy]] was approved. * 11:48 [[gitlab:alwinalbert|@alwinalbert]] was approved. * 11:48 [[gitlab:hridyakk|@hridyakk]] was approved. * 11:45 [[gitlab:gaurigupta21|@gaurigupta21]] was approved. * 11:45 [[gitlab:binetaa|@binetaa]] was approved. * 10:21 [[gitlab:jyothikat22|@jyothikat22]] was approved. * 10:21 [[gitlab:zobotrombie|@zobotrombie]] was approved. * 10:21 [[gitlab:flykrth|@flykrth]] was approved. * 10:21 [[gitlab:mehrinshamim|@mehrinshamim]] was approved. * 10:21 [[gitlab:aadhi13|@aadhi13]] was approved. * 10:21 [[gitlab:malavikam05|@malavikam05]] was approved. * 10:18 [[gitlab:nf609|@nf609]] was approved. * 05:48 [[gitlab:nazalnihad|@nazalnihad]] was approved. * 05:48 [[gitlab:naveen28204280|@naveen28204280]] was approved. === 2025-07-25 === * 09:49 [[gitlab:kasyap9|@kasyap9]] was approved. * 09:30 [[gitlab:swayamagrahari|@swayamagrahari]] was approved. === 2025-07-24 === * 19:36 [[gitlab:madutgn|@madutgn]] was approved. === 2025-07-23 === * 20:09 [[gitlab:somerandomdeveloper|@somerandomdeveloper]] was approved. === 2025-07-22 === * 00:15 [[gitlab:iagoqnsi|@iagoqnsi]] was approved. === 2025-07-21 === * 17:30 [[gitlab:asadiqui|@asadiqui]] was approved. * 16:39 [[gitlab:tryvix1509|@tryvix1509]] was approved. * 04:27 [[gitlab:damian|@damian]] was approved. === 2025-07-20 === * 09:42 "mike-khoroshun" was rejected (pending since 2025-04-20T09:42:22.732Z). === 2025-07-17 === * 17:57 [[gitlab:haroldkrabs|@haroldkrabs]] was approved. * 13:45 [[gitlab:envlh|@envlh]] was approved. === 2025-07-14 === * 10:24 [[gitlab:missguru|@missguru]] was approved. * 00:57 "clarfonthey" was rejected (pending since 2025-04-14T00:56:32.626Z). === 2025-07-13 === * 01:01 [[gitlab:l235|@l235]] was approved. === 2025-07-11 === * 03:06 "rodavlas" was rejected (pending since 2025-04-11T03:05:45.590Z). === 2025-07-06 === * 00:09 "lakasa" was rejected (pending since 2025-04-06T00:06:28.469Z). === 2025-07-05 === * 21:54 "ctrlzvi" was rejected (pending since 2025-04-05T21:54:12.542Z). * 14:30 "aminualiyu" was rejected (pending since 2025-04-05T14:27:22.617Z). === 2025-07-04 === * 03:15 [[gitlab:galstar|@galstar]] was approved. === 2025-07-02 === * 11:27 "vicolas11" was rejected (pending since 2025-04-02T11:25:12.682Z). === 2025-06-29 === * 23:12 "naomi723" was rejected (pending since 2025-03-30T23:09:24.630Z). === 2025-06-28 === * 16:21 "mudeh2372" was rejected (pending since 2025-03-29T16:18:27.057Z). === 2025-06-27 === * 23:18 "rony143" was rejected (pending since 2025-03-28T23:16:13.671Z). * 22:21 [[gitlab:rluts|@rluts]] was approved. === 2025-06-26 === * 13:54 "creativegurus" was rejected (pending since 2025-03-27T13:52:41.706Z). === 2025-06-24 === * 17:42 [[gitlab:devjadiya|@devjadiya]] was approved. * 14:00 "dominic-r" was rejected (pending since 2025-03-25T14:00:07.307Z). === 2025-06-21 === * 00:48 [[gitlab:vriaa|@vriaa]] was approved. === 2025-06-18 === * 15:21 "ayushkhati1" was rejected (pending since 2025-03-19T15:18:50.062Z). === 2025-06-17 === * 20:45 "chiomavero" was rejected (pending since 2025-03-18T20:44:13.967Z). * 00:27 [[gitlab:eggroll97|@eggroll97]] was approved. === 2025-06-14 === * 20:57 "volvox" was rejected (pending since 2025-03-15T20:56:34.018Z). === 2025-06-13 === * 16:09 [[gitlab:supergrey|@supergrey]] was approved. * 11:03 "chqaz" was rejected (pending since 2025-03-14T11:01:09.600Z). * 10:24 [[gitlab:slong-wmf|@slong-wmf]] was approved. * 10:15 "hearvox" was rejected (pending since 2025-03-14T10:13:13.112Z). === 2025-06-12 === * 15:18 "jlam" was rejected (pending since 2025-03-13T15:17:54.099Z). === 2025-06-09 === * 20:48 "dipanjansengupta" was rejected (pending since 2025-03-10T20:48:03.545Z). * 19:27 [[gitlab:reggycelly|@reggycelly]] was approved. * 14:51 "arendpieter" was rejected (pending since 2025-03-10T14:51:01.445Z). * 13:21 [[gitlab:greenreaper|@greenreaper]] was approved. * 09:33 [[gitlab:mmta|@mmta]] was approved. * 08:03 "a-ssh22" was rejected (pending since 2025-03-10T08:03:08.111Z). === 2025-06-08 === * 21:06 "mm-episodenlistedlvaupdater" was rejected (pending since 2025-03-09T21:04:06.323Z). === 2025-06-06 === * 11:06 [[gitlab:olea|@olea]] was approved. === 2025-06-05 === * 20:33 [[gitlab:encodedwp|@encodedwp]] was approved. * 15:00 [[gitlab:toluayo|@toluayo]] was approved. * 13:51 [[gitlab:arnold_lup|@arnold_lup]] was approved. * 11:54 "sdhehua" was rejected (pending since 2025-03-06T11:51:48.241Z). === 2025-06-03 === * 21:27 [[gitlab:wewakey|@wewakey]] was approved. * 12:36 "hunsimon2" was rejected (pending since 2025-03-04T12:34:56.520Z). * 11:54 "hunsimon" was rejected (pending since 2025-03-04T11:53:54.652Z). === 2025-06-02 === * 12:01 [[gitlab:jaimedes|@jaimedes]] was approved. === 2025-05-30 === * 18:00 "sathvik9105" was rejected (pending since 2025-02-28T17:59:42.867Z). * 11:21 [[gitlab:tonythomas01|@tonythomas01]] was approved. * 10:06 [[gitlab:gpsleo|@gpsleo]] was approved. === 2025-05-29 === * 22:12 [[gitlab:codynguyen1116|@codynguyen1116]] was approved. === 2025-05-28 === * 02:57 [[gitlab:saper|@saper]] was approved. === 2025-05-27 === * 21:06 [[gitlab:mohammed_qays|@mohammed_qays]] was approved. * 15:33 "satanluimm" was rejected (pending since 2025-02-25T15:32:48.101Z). === 2025-05-26 === * 23:57 "seyedali220" was rejected (pending since 2025-02-24T23:56:17.621Z). === 2025-05-21 === * 11:12 [[gitlab:guilherme|@guilherme]] was approved. === 2025-05-19 === * 13:24 [[gitlab:emojiwiki|@emojiwiki]] was approved. === 2025-05-18 === * 00:00 "xidme" was rejected (pending since 2025-02-15T23:58:56.796Z). === 2025-05-17 === * 02:39 "kdh8219" was rejected (pending since 2025-02-15T02:36:32.237Z). === 2025-05-16 === * 15:09 [[gitlab:maxbinderwmf|@maxbinderwmf]] was approved. === 2025-05-15 === * 04:30 "inspectorzer0" was rejected (pending since 2025-02-13T04:27:33.179Z). === 2025-05-14 === * 17:42 [[gitlab:llugo|@llugo]] was approved. === 2025-05-13 === * 20:18 "mmta" was rejected (pending since 2025-02-11T20:17:23.407Z). === 2025-05-11 === * 20:51 "jad" was rejected (pending since 2025-02-09T20:49:07.333Z). * 17:54 "nishchalsundan" was rejected (pending since 2025-02-09T17:52:25.761Z). * 16:39 "mohammed_abukhadra" was rejected (pending since 2025-02-09T16:39:03.730Z). === 2025-05-09 === * 09:12 [[gitlab:sirchanmp|@sirchanmp]] was approved. === 2025-05-08 === * 08:18 [[gitlab:mengeditch|@mengeditch]] was approved. === 2025-05-07 === * 03:45 "xluffy" was rejected (pending since 2025-02-05T03:45:14.181Z). === 2025-05-06 === * 16:54 "punhaniabhishek" was rejected (pending since 2025-02-04T16:53:50.758Z). * 09:36 [[gitlab:bmartinezcalvo|@bmartinezcalvo]] was approved. === 2025-05-02 === * 12:24 [[gitlab:tohaomg|@tohaomg]] was approved. * 11:48 [[gitlab:mavrikant|@mavrikant]] was approved. * 11:45 [[gitlab:daanvr|@daanvr]] was approved. === 2025-05-01 === * 09:09 "mjoerg" was rejected (pending since 2025-01-30T09:09:04.204Z). === 2025-04-30 === * 23:06 "sanskardubey" was rejected (pending since 2025-01-29T23:03:25.489Z). === 2025-04-29 === * 16:00 "geyslein" was rejected (pending since 2025-01-28T16:00:01.510Z). === 2025-04-26 === * 09:30 "anjali9027" was rejected (pending since 2025-01-25T09:28:07.064Z). === 2025-04-25 === * 18:00 "salahhazaa" was rejected (pending since 2025-01-24T17:58:30.030Z). * 15:15 [[gitlab:yiming|@yiming]] was approved. * 02:06 "mrchanmp" was rejected (pending since 2025-01-24T02:03:58.308Z). === 2025-04-23 === * 17:03 "rj2904" was rejected (pending since 2025-01-22T17:03:11.207Z). * 14:21 "nischay33" was rejected (pending since 2025-01-22T14:19:21.081Z). === 2025-04-22 === * 19:27 "dj80" was rejected (pending since 2025-01-21T19:25:28.498Z). * 14:30 [[gitlab:kaimamin|@kaimamin]] was approved. * 09:57 "debo" was rejected (pending since 2025-01-21T09:54:47.955Z). === 2025-04-21 === * 12:24 "unshell" was rejected (pending since 2025-01-20T12:21:59.686Z). === 2025-04-18 === * 15:06 [[gitlab:spartanarbinger|@spartanarbinger]] was approved. === 2025-04-16 === * 03:09 "dewey" was rejected (pending since 2025-01-15T03:06:17.488Z). === 2025-04-15 === * 19:45 "emdadul" was rejected (pending since 2025-01-14T19:42:29.285Z). === 2025-04-14 === * 06:45 [[gitlab:bcampbell804|@bcampbell804]] was approved. === 2025-04-11 === * 06:27 [[gitlab:jvanderhoop|@jvanderhoop]] was approved. === 2025-04-10 === * 04:12 "bhai420" was rejected (pending since 2025-01-09T04:10:29.430Z). === 2025-04-09 === * 05:03 "austinvarshney" was rejected (pending since 2025-01-08T05:02:34.175Z). === 2025-04-06 === * 15:36 [[gitlab:elph|@elph]] was approved. === 2025-04-02 === * 10:33 [[gitlab:ozge|@ozge]] was approved. === 2025-03-31 === * 20:15 "demandkey" was rejected (pending since 2024-12-30T20:14:23.096Z). * 15:18 [[gitlab:danyya|@danyya]] was approved. === 2025-03-28 === * 15:54 [[gitlab:rutsavi09|@rutsavi09]] was approved. * 15:54 [[gitlab:ilanen1|@ilanen1]] was approved. === 2025-03-25 === * 19:27 [[gitlab:irfo|@irfo]] was approved. * 11:54 [[gitlab:kmontalva-wmf|@kmontalva-wmf]] was approved. * 04:33 [[gitlab:paul26|@paul26]] was approved. * 04:18 "as1100k" was rejected (pending since 2024-12-24T04:18:06.813Z). === 2025-03-24 === * 11:33 "amzadkhankk" was rejected (pending since 2024-12-23T11:33:14.176Z). === 2025-03-23 === * 12:24 "wolfdo" was rejected (pending since 2024-12-22T12:23:35.056Z). === 2025-03-22 === * 09:45 [[gitlab:fjmustak|@fjmustak]] was approved. === 2025-03-20 === * 18:42 "sathishkokila" was rejected (pending since 2024-12-19T18:39:35.161Z). * 17:03 [[gitlab:alien4444|@alien4444]] was approved. * 15:27 [[gitlab:davidcoronel|@davidcoronel]] was approved. === 2025-03-19 === * 22:57 [[gitlab:r1f4t|@r1f4t]] was approved. * 19:03 "daniel24ps" was rejected (pending since 2024-12-18T19:00:21.249Z). * 14:18 [[gitlab:beepbooppenguin|@beepbooppenguin]] was approved. === 2025-03-18 === * 17:48 "rahulkundu1209" was rejected (pending since 2024-12-17T17:46:41.936Z). * 08:15 "kirtisikka972" was rejected (pending since 2024-12-17T08:13:25.487Z). === 2025-03-15 === * 13:30 "tulspal_sidhu" was rejected (pending since 2024-12-14T13:29:10.606Z). * 01:39 "peacedeadc" was rejected (pending since 2024-12-14T01:37:36.579Z). === 2025-03-14 === * 03:51 [[gitlab:chuckthebuck|@chuckthebuck]] was approved. * 02:33 "yxngtrtxll" was rejected (pending since 2024-12-13T02:31:51.658Z). === 2025-03-13 === * 14:36 [[gitlab:iccander|@iccander]] was approved. === 2025-03-12 === * 23:21 "jokerchic36" was rejected (pending since 2024-12-11T23:21:00.670Z). * 15:30 [[gitlab:naomi|@naomi]] was approved. * 15:27 [[gitlab:cobi|@cobi]] was approved. === 2025-03-11 === * 12:42 "mohitvermaxx" was rejected (pending since 2024-12-10T12:40:56.967Z). === 2025-03-10 === * 16:51 [[gitlab:nanona15dobato|@nanona15dobato]] was approved. === 2025-03-09 === * 22:39 [[gitlab:jonkolbert|@jonkolbert]] was approved. * 20:45 [[gitlab:urbanecmtest2|@urbanecmtest2]] was approved. === 2025-03-07 === * 16:54 [[gitlab:hswan|@hswan]] was approved. * 14:42 [[gitlab:atitkov|@atitkov]] was approved. * 00:42 [[gitlab:infrastruktur|@infrastruktur]] was approved. === 2025-03-06 === * 17:21 "johnmann" was rejected (pending since 2024-12-05T17:19:24.995Z). === 2025-03-05 === * 07:33 [[gitlab:monx9494|@monx9494]] was approved. === 2025-03-02 === * 21:21 "paul26" was rejected (pending since 2024-12-01T21:20:19.681Z). === 2025-03-01 === * 19:15 [[gitlab:izno|@izno]] was approved. * 12:45 [[gitlab:nyerho|@nyerho]] was approved. === 2025-02-28 === * 18:27 [[gitlab:chuckonwumelu|@chuckonwumelu]] was approved. * 13:09 "ashwinpraveengo" was rejected (pending since 2024-11-29T13:07:47.240Z). * 00:18 "eduardoaugusto" was rejected (pending since 2024-11-29T00:17:43.372Z). === 2025-02-27 === * 20:39 "volkanurl" was rejected (pending since 2024-11-28T20:37:18.101Z). === 2025-02-24 === * 21:15 [[gitlab:feeglgeef|@feeglgeef]] was approved. * 20:18 [[gitlab:piaanalysis2|@piaanalysis2]] was approved. * 19:06 [[gitlab:dhardy|@dhardy]] was approved. === 2025-02-22 === * 19:27 [[gitlab:owuh|@owuh]] was approved. === 2025-02-19 === * 16:06 [[gitlab:artemkloko|@artemkloko]] was approved. * 13:03 [[gitlab:jgafnea|@jgafnea]] was approved. === 2025-02-17 === * 16:33 [[gitlab:asmartkitten|@asmartkitten]] was approved. === 2025-02-16 === * 19:12 "gaurigupta21" was rejected (pending since 2024-11-17T19:11:07.416Z). === 2025-02-15 === * 01:18 [[gitlab:mediawiki-quickstart-ci|@mediawiki-quickstart-ci]] was approved. === 2025-02-14 === * 15:21 "nathanbnm" was rejected (pending since 2024-11-15T15:18:19.632Z). === 2025-02-13 === * 16:45 [[gitlab:priyanshuchahal|@priyanshuchahal]] was approved. * 16:42 [[gitlab:ajhalili2006|@ajhalili2006]] was approved. === 2025-02-12 === * 23:21 "monkeypatch999" was rejected (pending since 2024-11-13T23:20:38.398Z). * 06:36 [[gitlab:jainlakshita28|@jainlakshita28]] was approved. === 2025-02-11 === * 19:27 [[gitlab:matthewsm2|@matthewsm2]] was approved. === 2025-02-09 === * 16:15 "mohammed_abukhadra" was rejected (pending since 2024-11-10T16:15:18.361Z). === 2025-02-07 === * 21:33 "brennan" was rejected (pending since 2024-11-08T21:31:07.351Z). === 2025-02-06 === * 08:24 "mmta" was rejected (pending since 2024-11-07T08:22:36.724Z). * 06:21 [[gitlab:bunnypranav|@bunnypranav]] was approved. === 2025-02-05 === * 22:39 "chrissteinchen" was rejected (pending since 2024-11-06T22:38:16.673Z). === 2025-02-03 === * 07:45 "edriiic" was rejected (pending since 2024-11-04T07:44:46.849Z). * 01:12 "geppy" was rejected (pending since 2024-11-04T01:10:48.710Z). === 2025-02-02 === * 13:18 "funa-enpitu" was rejected (pending since 2024-11-03T13:15:46.065Z). === 2025-01-31 === * 23:42 "nfontes" was rejected (pending since 2024-11-01T23:39:41.755Z). * 22:51 "sbronson" was rejected (pending since 2024-11-01T22:50:31.871Z). * 00:42 [[gitlab:farid|@farid]] was approved. === 2025-01-27 === * 08:15 [[gitlab:eliza189|@eliza189]] was approved. === 2025-01-25 === * 09:51 [[gitlab:pamputt|@pamputt]] was approved. === 2025-01-23 === * 14:30 [[gitlab:lubianat|@lubianat]] was approved. * 11:45 [[gitlab:bootsa|@bootsa]] was approved. === 2025-01-21 === * 05:09 "niko" was rejected (pending since 2024-07-21T16:10:01.377Z). * 05:09 "thawizkid369777" was rejected (pending since 2024-07-18T17:42:44.493Z). * 05:09 "sarthaksingh2" was rejected (pending since 2024-07-10T11:31:30.470Z). * 05:09 "shriyakt" was rejected (pending since 2024-07-06T04:54:10.248Z). * 05:09 "akshaya" was rejected (pending since 2024-07-06T04:04:51.488Z). * 05:09 "alaka03aj" was rejected (pending since 2024-07-05T18:01:54.876Z). * 05:09 "sulochanaviji-5049" was rejected (pending since 2024-07-01T05:58:00.427Z). * 05:09 "nayanjnath" was rejected (pending since 2024-07-01T02:51:57.405Z). * 05:09 "sd44" was rejected (pending since 2024-06-30T04:28:51.436Z). * 05:09 "metavalent" was rejected (pending since 2024-06-29T01:37:14.210Z). * 05:09 "wicloudx" was rejected (pending since 2024-06-28T11:51:23.335Z). * 05:09 "debo" was rejected (pending since 2024-06-28T01:44:59.845Z). * 05:09 "bwiki" was rejected (pending since 2024-06-23T14:15:38.032Z). * 05:09 "toprak" was rejected (pending since 2024-06-23T11:35:50.819Z). * 05:09 "iristeller" was rejected (pending since 2024-06-14T20:53:48.959Z). * 05:09 "jcolvin" was rejected (pending since 2024-06-12T17:29:01.238Z). * 05:09 "kalyan" was rejected (pending since 2024-06-07T07:52:46.993Z). * 05:09 "bluecrystal" was rejected (pending since 2024-06-06T19:16:20.107Z). * 05:09 "iftttrohit" was rejected (pending since 2024-06-04T12:08:50.818Z). * 05:09 "pogpotato" was rejected (pending since 2024-06-03T17:58:21.684Z). * 05:09 "cptlausebaer" was rejected (pending since 2024-05-31T18:53:27.692Z). * 05:09 "hdevine825" was rejected (pending since 2024-05-31T17:04:18.279Z). * 05:09 "anaghaa18" was rejected (pending since 2024-05-25T19:14:31.803Z). * 05:09 "atharvanair04" was rejected (pending since 2024-05-25T14:24:52.825Z). * 05:09 "anasvemmully" was rejected (pending since 2024-05-25T06:10:27.261Z). * 05:09 "abhinavmohandas" was rejected (pending since 2024-05-25T06:05:24.825Z). * 05:09 "kksurendran06" was rejected (pending since 2024-05-25T06:04:38.082Z). * 05:09 "albertmarshall8896" was rejected (pending since 2024-05-23T09:32:05.462Z). * 05:09 "akellison" was rejected (pending since 2024-05-17T02:07:24.229Z). * 05:09 "mainowill" was rejected (pending since 2024-04-16T23:30:33.881Z). * 05:09 "bzhqc" was rejected (pending since 2024-04-16T19:50:38.676Z). * 05:09 "safan41" was rejected (pending since 2024-04-16T03:34:48.942Z). * 05:09 "mgagat" was rejected (pending since 2024-04-16T03:21:51.764Z). * 05:09 "okeamah" was rejected (pending since 2024-04-16T02:49:00.143Z). * 05:09 "xuhao61" was rejected (pending since 2024-04-15T23:45:09.083Z). * 04:47 "cybel" was rejected (pending since 2024-04-15T06:46:35.791Z). === 2025-01-20 === * 14:33 [[gitlab:your1|@your1]] was approved. === 2025-01-18 === * 10:09 [[gitlab:galrach600|@galrach600]] was approved. * 02:51 [[gitlab:blankeclair|@blankeclair]] was approved. === 2025-01-17 === * 13:57 [[gitlab:dsantamaria|@dsantamaria]] was approved. === 2025-01-15 === * 17:12 [[gitlab:smartse|@smartse]] was approved. === 2025-01-14 === * 17:03 [[gitlab:naorleizer|@naorleizer]] was approved. === 2025-01-13 === * 02:45 [[gitlab:wolf20482|@wolf20482]] was approved. === 2025-01-12 === * 17:45 [[gitlab:tamzin|@tamzin]] was approved. === 2025-01-11 === * 15:24 [[gitlab:bargioni|@bargioni]] was approved. * 14:30 [[gitlab:salelya|@salelya]] was approved. * 10:15 [[gitlab:malakatshy|@malakatshy]] was approved. * 05:21 [[gitlab:newmcpee|@newmcpee]] was approved. === 2025-01-09 === * 15:30 [[gitlab:gkyziridis|@gkyziridis]] was approved. === 2025-01-08 === * 16:21 [[gitlab:ukrface|@ukrface]] was approved. === 2024-12-28 === * 03:27 [[gitlab:twonum|@twonum]] was approved. === 2024-12-25 === * 06:09 [[gitlab:harsv567|@harsv567]] was approved. === 2024-12-21 === * 11:24 [[gitlab:amutha2002|@amutha2002]] was approved. === 2024-12-20 === * 19:51 [[gitlab:hridyeshgupta|@hridyeshgupta]] was approved. * 10:00 [[gitlab:ro-shines|@ro-shines]] was approved. * 08:09 [[gitlab:kesharwaniarpita|@kesharwaniarpita]] was approved. === 2024-12-18 === * 14:45 [[gitlab:soylacarli|@soylacarli]] was approved. === 2024-12-16 === * 20:33 [[gitlab:aleyasiddika1|@aleyasiddika1]] was approved. === 2024-12-15 === * 07:33 [[gitlab:abhishek02bhardwaj|@abhishek02bhardwaj]] was approved. === 2024-12-13 === * 13:18 [[gitlab:ashmitabathre204|@ashmitabathre204]] was approved. === 2024-12-10 === * 06:39 [[gitlab:ginaan|@ginaan]] was approved. === 2024-12-09 === * 05:45 [[gitlab:kallinavya|@kallinavya]] was approved. * 00:54 [[gitlab:viserion-7|@viserion-7]] was approved. === 2024-12-08 === * 17:27 [[gitlab:wargo|@wargo]] was approved. === 2024-12-05 === * 11:15 [[gitlab:ranjithraj|@ranjithraj]] was approved. === 2024-12-02 === * 21:21 [[gitlab:a930913|@a930913]] was approved. === 2024-12-01 === * 02:39 [[gitlab:kingchristlike1|@kingchristlike1]] was approved. === 2024-11-21 === * 13:45 [[gitlab:sascha|@sascha]] was approved. === 2024-11-19 === * 16:36 [[gitlab:jly|@jly]] was approved. === 2024-11-15 === * 02:54 [[gitlab:danielyepezgarces|@danielyepezgarces]] was approved. === 2024-11-14 === * 14:15 [[gitlab:stimoroll|@stimoroll]] was approved. === 2024-11-09 === * 17:15 [[gitlab:f4udeveloper|@f4udeveloper]] was approved. === 2024-11-07 === * 19:15 [[gitlab:zulf|@zulf]] was approved. * 05:33 [[gitlab:hassanamin|@hassanamin]] was approved. === 2024-11-06 === * 19:39 [[gitlab:daniuu|@daniuu]] was approved. * 00:18 [[gitlab:rlopez-wmf|@rlopez-wmf]] was approved. === 2024-10-09 === * 14:45 [[gitlab:jtweed|@jtweed]] was approved. * 10:24 [[gitlab:ifrahkh|@ifrahkh]] was approved. * 09:06 [[gitlab:wikibayer|@wikibayer]] was approved. === 2024-10-06 === * 10:27 [[gitlab:keerthan16|@keerthan16]] was approved. === 2024-10-04 === * 07:45 [[gitlab:hakimi97|@hakimi97]] was approved. === 2024-09-30 === * 07:39 [[gitlab:ninjastrikers|@ninjastrikers]] was approved. === 2024-09-28 === * 17:30 [[gitlab:webrunner95|@webrunner95]] was approved. === 2024-09-18 === * 21:39 [[gitlab:elliottetzkorn|@elliottetzkorn]] was approved. === 2024-09-14 === * 22:06 [[gitlab:humptydumpty|@humptydumpty]] was approved. === 2024-09-06 === * 08:48 [[gitlab:mickabarber|@mickabarber]] was approved. === 2024-08-27 === * 17:36 [[gitlab:edgars|@edgars]] was approved. === 2024-08-22 === * 09:18 [[gitlab:antonkokhwmde|@antonkokhwmde]] was approved. === 2024-08-14 === * 19:21 [[gitlab:jfk|@jfk]] was approved. === 2024-08-13 === * 17:57 [[gitlab:daxserver|@daxserver]] was approved. === 2024-08-11 === * 09:57 [[gitlab:pauliesnug|@pauliesnug]] was approved. === 2024-08-10 === * 08:42 [[gitlab:ashig|@ashig]] was approved. === 2024-08-09 === * 14:09 [[gitlab:masssly|@masssly]] was approved. === 2024-08-05 === * 22:15 [[gitlab:mrtortue|@mrtortue]] was approved. === 2024-08-02 === * 16:21 [[gitlab:dsantini|@dsantini]] was approved. === 2024-07-31 === * 11:54 [[gitlab:cptviraj|@cptviraj]] was approved. === 2024-07-30 === * 19:09 [[gitlab:iniquity|@iniquity]] was approved. * 10:00 [[gitlab:collins|@collins]] was approved. === 2024-07-27 === * 15:57 [[gitlab:songnguxyz|@songnguxyz]] was approved. === 2024-07-25 === * 12:36 [[gitlab:mszabo|@mszabo]] was approved. * 09:21 [[gitlab:agarwalmahima|@agarwalmahima]] was approved. === 2024-07-24 === * 08:05 [[gitlab:dragoniez|@dragoniez]] was approved. === 2024-07-23 === * 06:54 [[gitlab:mirji|@mirji]] was approved. === 2024-07-16 === * 10:00 [[gitlab:lakejason0|@lakejason0]] was approved. === 2024-07-12 === * 11:33 [[gitlab:cn|@cn]] was approved. * 08:12 [[gitlab:unchampignon|@unchampignon]] was approved. === 2024-07-07 === * 17:12 [[gitlab:agamyasamuel|@agamyasamuel]] was approved. * 05:24 [[gitlab:kuldeepburjbhalaike|@kuldeepburjbhalaike]] was approved. === 2024-07-06 === * 11:18 [[gitlab:dibya|@dibya]] was approved. * 04:54 [[gitlab:sarthakparashar|@sarthakparashar]] was approved. === 2024-07-05 === * 18:15 [[gitlab:vanshikarathi|@vanshikarathi]] was approved. === 2024-07-02 === * 19:00 [[gitlab:ebrahim|@ebrahim]] was approved. === 2024-07-01 === * 20:12 [[gitlab:rockingpenny4|@rockingpenny4]] was approved. * 18:15 [[gitlab:balajijagadesh|@balajijagadesh]] was approved. === 2024-06-30 === * 18:24 [[gitlab:hrideshmg|@hrideshmg]] was approved. * 07:18 [[gitlab:chanakyakumardas|@chanakyakumardas]] was approved. * 06:30 [[gitlab:rihaan180|@rihaan180]] was approved. === 2024-06-27 === * 17:36 [[gitlab:driedmueller|@driedmueller]] was approved. === 2024-06-19 === * 12:57 [[gitlab:audreypenven|@audreypenven]] was approved. === 2024-06-16 === * 01:18 [[gitlab:roysmith|@roysmith]] was approved. === 2024-06-08 === * 02:45 [[gitlab:jleedev|@jleedev]] was approved. === 2024-06-03 === * 13:57 [[gitlab:afeder|@afeder]] was approved. === 2024-06-01 === * 10:54 [[gitlab:florianschmitt|@florianschmitt]] was approved. === 2024-05-30 === * 16:42 [[gitlab:krlsca|@krlsca]] was approved. === 2024-05-28 === * 11:24 [[gitlab:rickijay|@rickijay]] was approved. === 2024-05-26 === * 11:18 [[gitlab:ranjithsiji|@ranjithsiji]] was approved. === 2024-05-25 === * 07:24 [[gitlab:jony|@jony]] was approved. === 2024-05-23 === * 08:45 [[gitlab:lepticed7|@lepticed7]] was approved. === 2024-05-22 === * 20:42 [[gitlab:echecs|@echecs]] was approved. === 2024-05-21 === * 13:33 [[gitlab:mbs|@mbs]] was approved. === 2024-05-19 === * 18:06 [[gitlab:ionenlaser|@ionenlaser]] was approved. === 2024-05-18 === * 23:36 [[gitlab:mdaniels5757|@mdaniels5757]] was approved. === 2024-05-17 === * 08:54 [[gitlab:grapedog|@grapedog]] was approved. === 2024-05-08 === * 19:42 [[gitlab:kelhurd|@kelhurd]] was approved. * 19:06 [[gitlab:khurd|@khurd]] was approved. === 2024-05-06 === * 19:48 [[gitlab:j3j5|@j3j5]] was approved. * 12:06 [[gitlab:tk-999|@tk-999]] was approved. === 2024-05-05 === * 22:09 [[gitlab:pppery|@pppery]] was approved. * 20:33 [[gitlab:sakretsu|@sakretsu]] was approved. * 12:12 [[gitlab:waterquark|@waterquark]] was approved. === 2024-05-04 === * 09:03 [[gitlab:multichill|@multichill]] was approved. * 07:42 [[gitlab:abaris|@abaris]] was approved. === 2024-05-03 === * 14:57 [[gitlab:maurusian|@maurusian]] was approved. === 2024-04-24 === * 05:48 [[gitlab:wolfinux|@wolfinux]] was approved. === 2024-04-23 === * 15:48 [[gitlab:dreamrimmer|@dreamrimmer]] was approved. === 2024-04-21 === * 06:51 [[gitlab:alon|@alon]] was approved. === 2024-04-17 === * 23:33 [[gitlab:derenrich|@derenrich]] was approved. === 2024-04-16 === * 17:18 [[gitlab:valcio|@valcio]] was approved. === 2024-04-14 === * 16:51 [[gitlab:wikilucas00|@wikilucas00]] was approved. === 2024-04-06 === * 12:48 [[gitlab:theprotonade|@theprotonade]] was approved. === 2024-04-02 === * 07:30 [[gitlab:bohuizhang|@bohuizhang]] was approved. === 2024-03-30 === * 13:36 [[gitlab:lpintscher|@lpintscher]] was approved. === 2024-03-26 === * 17:09 [[gitlab:eenabulele|@eenabulele]] was approved. === 2024-03-25 === * 14:27 [[gitlab:tuukka|@tuukka]] was approved. === 2024-03-24 === * 12:24 [[gitlab:firefly|@firefly]] was approved. === 2024-03-21 === * 19:33 [[gitlab:universal-omega|@universal-omega]] was approved. === 2024-03-17 === * 10:36 [[gitlab:bisel91|@bisel91]] was approved. === 2024-03-16 === * 10:09 [[gitlab:delord|@delord]] was approved. * 00:42 [[gitlab:athulvis1|@athulvis1]] was approved. === 2024-03-15 === * 19:06 [[gitlab:ignaciorodrguez|@ignaciorodrguez]] was approved. * 08:30 [[gitlab:peachey88|@peachey88]] was approved. * 06:51 [[gitlab:derick|@derick]] was approved. === 2024-03-12 === * 15:06 [[gitlab:xiaoxiao|@xiaoxiao]] was approved. === 2024-03-06 === * 13:21 [[gitlab:desianabae1|@desianabae1]] was approved. === 2024-03-05 === * 19:21 [[gitlab:ep1c|@ep1c]] was approved. * 16:33 [[gitlab:jasmine|@jasmine]] was approved. === 2024-03-02 === * 06:42 [[gitlab:potsdamlamb|@potsdamlamb]] was approved. === 2024-02-29 === * 23:18 [[gitlab:arandomname123|@arandomname123]] was approved. * 18:03 [[gitlab:baba|@baba]] was approved. * 17:48 [[gitlab:yfdyh000|@yfdyh000]] was approved. * 03:09 [[gitlab:sds|@sds]] was approved. === 2024-02-27 === * 23:33 [[gitlab:lofhi|@lofhi]] was approved. === 2024-02-15 === * 19:45 [[gitlab:gergesshamon|@gergesshamon]] was approved. === 2024-02-14 === * 14:33 [[gitlab:philipnelson99|@philipnelson99]] was approved. === 2024-02-13 === * 13:06 [[gitlab:dringsim|@dringsim]] was approved. === 2024-02-12 === * 17:36 [[gitlab:haak|@haak]] was approved. === 2024-02-05 === * 17:33 [[gitlab:qwerfjkl|@qwerfjkl]] was approved. * 17:14 [[gitlab:ahecht|@ahecht]] was approved. === 2024-02-01 === * 09:27 [[gitlab:arinaigum|@arinaigum]] was approved. * 00:15 [[gitlab:jas42|@jas42]] was approved. * 00:15 [[gitlab:edhu|@edhu]] was approved. * 00:15 [[gitlab:marnanel|@marnanel]] was approved. * 00:15 [[gitlab:ibrahemqasim|@ibrahemqasim]] was approved. * 00:15 [[gitlab:amasotti|@amasotti]] was approved. * 00:15 [[gitlab:deni|@deni]] was approved. * 00:15 [[gitlab:cyber|@cyber]] was approved. * 00:15 [[gitlab:saroj|@saroj]] was approved. === 2024-01-29 === * 21:42 [[gitlab:rgupta|@rgupta]] was approved. === 2024-01-07 === * 09:48 [[gitlab:lutrome|@lutrome]] was approved. === 2024-01-05 === * 20:48 [[gitlab:jinoytommanjaly|@jinoytommanjaly]] was approved. * 02:51 [[gitlab:braunobruno|@braunobruno]] was approved. * 01:08 [[gitlab:amorymeltzer|@amorymeltzer]] was approved. * 01:08 [[gitlab:phi22ipus|@phi22ipus]] was approved. === 2024-01-03 === * 14:45 [[gitlab:gabina|@gabina]] was approved. === 2024-01-02 === * 13:18 [[gitlab:arthurtaylor|@arthurtaylor]] was approved. === 2023-12-23 === * 00:33 [[gitlab:aram|@aram]] was approved. === 2023-12-22 === * 16:24 [[gitlab:elpitareio|@elpitareio]] was approved. === 2023-12-21 === * 00:43 [[gitlab:bsadowski1|@bsadowski1]] was approved. * 00:43 [[gitlab:ederporto|@ederporto]] was approved. * 00:43 [[gitlab:sadraiiali|@sadraiiali]] was approved. * 00:43 [[gitlab:wasp-outis|@wasp-outis]] was approved. * 00:43 [[gitlab:bodhisattwa|@bodhisattwa]] was approved. * 00:43 [[gitlab:air7538|@air7538]] was approved. * 00:43 [[gitlab:anzx|@anzx]] was approved. * 00:43 [[gitlab:tekask1903|@tekask1903]] was approved. * 00:42 [[gitlab:kiwi-0x010c|@kiwi-0x010c]] was approved. * 00:42 [[gitlab:mpaa|@mpaa]] was approved. * 00:42 [[gitlab:kutay|@kutay]] was approved. * 00:42 [[gitlab:wattmto|@wattmto]] was approved. moegdv3uo7misgvi4gjaioe389ozekz Wikitech:Village pump 4 458615 2450646 2447808 2026-08-22T15:09:56Z Codename Noreste 39936 /* Inactive administrator removal */ new topic ([[mw:c:Special:MyLanguage/User:JWBTH/CD|CD]]) 2450646 wikitext text/x-wiki {{/Header}} == Removal of administrator rights from locked accounts == Could a bureaucrat please remove administrator rights from the following users? * {{user|AKosiaris (WMF)}} * {{user|ABorrero (WMF)}} They no longer work for the Wikimedia Foundation. Thanks. [[User:Codename Noreste|Codename Noreste]] ([[User talk:Codename Noreste|talk]]) 21:18, 15 February 2026 (UTC) :@[[User:Codename Noreste|Codename Noreste]]: This was done by @[[User:BryanDavis|BryanDavis]] in https://wikitech.wikimedia.org/w/index.php?title=Special:Log&logid=1008847 and https://wikitech.wikimedia.org/w/index.php?title=Special:Log&logid=1008848 respectively. [[User:Jdforrester (WMF)|Jdforrester (WMF)]] ([[User talk:Jdforrester (WMF)|talk]]) 16:01, 17 February 2026 (UTC) ::Thank you for the notification. [[User:Codename Noreste|Codename Noreste]] ([[User talk:Codename Noreste|talk]]) 16:03, 17 February 2026 (UTC) == Main page change requests == Can the following be performed? # Add theme-night-mainpage to [[MediaWiki:Wikimedia-styles-exclude]] # The main page title pages may optionally be set to blank: [[MediaWiki:Mainpage-title]] and [[MediaWiki:Mainpage-title-loggedin]] # An interface administrator can remove <code>.page-Main_Page .firstHeading,</code> from [[MediaWiki:Common.css]] if the previous step was made Thank you. [[User:Codename Noreste|Codename Noreste]] ([[User talk:Codename Noreste|talk]]) 03:07, 9 May 2026 (UTC) :Done all 3. [[User:taavi|taavi]] ([[User talk:Taavi|talk!]]) 13:54, 12 May 2026 (UTC) == July 2026 Wikimedia Café meetups regarding Wikimedia governance and options for reform == <div class="border-box" style="background-color: var(--background-color-warning-subtle, #f8eaba); max-width: 875px; padding: 5px; border: 1px solid black; margin: 5px; color: var(--clr-dark)"><div class="box" style="float:left; padding-top: 10px; padding-right: 10px; padding-left: 10px; padding-bottom: 10px;">[[File:Wikimedia_Café_logo_in_plain_SVG_format.svg|alt=The logo for the Wikimedia Café|60x60px]]</div>Hello! There will be two '''[[m:Wikimedia_Café|Wikimedia Café]]''' discussion opportunities in July. Both sessions will focus on Wikimedia governance, including possible follow-ups to the [[m:Movement_Charter|Movement Charter]] and options for reform. Participants may attend either or both Café sessions. This month, to deconflict the Café meetups from Wikimania, the meetups will be held one day later than usual. # '''26 July 2026 15:00 UTC''' ([https://zonestamp.toolforge.org/1785078000 timestamp converter]), at a time friendly to the Americas, Africa, and Europe # '''27 July 2026 03:00 UTC''' ([https://zonestamp.toolforge.org/1785121200 timestamp converter]), at a time friendly to Asia and the Pacific Please see the Café page for more information, including [[m:Wikimedia_Café#How_to_attend_the_session|how to register]]! [[File:Buntstifte_Eberhard_Faber_crop_64h.jpg|alt=cropped image of colored pencils|860x860px]]</div><nowiki>~~~~</nowiki> [[User:Pine|Pine]] ([[User talk:Pine|talk]]) 04:01, 13 July 2026 (UTC) == August 2026 Wikimedia Café meetups regarding [https://meta.wikimedia.org/wiki/Next_25 Next 25] == <div class="border-box" style="background-color: var(--background-color-warning-subtle, #f8eaba); max-width: 875px; padding: 5px; border: 1px solid black; margin: 5px; color: var(--clr-dark)"><div class="box" style="float:left; padding-top: 10px; padding-right: 10px; padding-left: 10px; padding-bottom: 10px;">[[File:Wikimedia_Café_logo_in_plain_SVG_format.svg|alt=The logo for the Wikimedia Café|60x60px]]</div>Hello! There will be two '''[[m:Wikimedia_Café|Wikimedia Café]]''' discussion opportunities in August. Both sessions will focus on [[m:Next_25|the Next 25 initiative]], including [[m:Next_25/A_proposal_to_frame_Next_25|framing the initiative]]. Participants may attend either or both Café sessions. # '''29 August 2026 15:00 UTC''' ([https://zonestamp.toolforge.org/1788015600 timestamp converter]), at a time friendly to the Americas, Africa, and Europe # '''30 August 2026 03:00 UTC''' ([https://zonestamp.toolforge.org/1788058800 timestamp converter]), at a time friendly to Asia and the Pacific Please see the [[m:Wikimedia_Café|Café]] page for more information, including [[m:Wikimedia_Café#Agenda._This_will_be_an_approximately_1_hour_Caf%C3%A9_session|agenda details]] and [[m:Wikimedia_Café#How_to_attend_the_session|how to register]]! To subscribe or unsubscribe for Café notices on your talk page please [[m:Global_message_delivery/Targets/Wikimedia_Café|go here]]. [[File:Buntstifte_Eberhard_Faber_crop_64h.jpg|alt=cropped image of colored pencils|860x860px]]</div><nowiki>~~~~</nowiki> <span style="white-space:nowrap;">[[User:Pine|<span style="color:#01796f; text-shadow:#00BFFF 0 0 1.0em">༄ᨒ𖠰 Pine</span>]] [[User talk:Pine|<span style="color:DeepSkyBlue">(<b style="color:#FFDF00;text-shadow:#FFDF00 0 0 1.0em">✉</b>)</span>]]</span> 01:34, 17 August 2026 (UTC) == Inactive administrator removal == Can a bureaucrat please remove administrator rights from the following user? * {{user|billinghurst}} They have been inactive for roughly more than four years. Thank you. [[User:Codename Noreste|Codename Noreste]] ([[User talk:Codename Noreste|talk]]) 15:09, 22 August 2026 (UTC) 9dg1p69hjwkvko8vbypchkzj4bhe4f3 Nova Resource:Tools.wheres-my-code-running/SAL 498 460577 2450666 2450622 2026-08-22T21:46:00Z Stashbot 7414 wmbot~bd808@tools-bastion-14: Hard restart to pick up container from d796189c and bump CPU to 1 core (T435662) 2450666 wikitext text/x-wiki === 2026-08-22 === * 21:46 wmbot~bd808@tools-bastion-14: Hard restart to pick up container from {{Gerrit|d796189c}} and bump CPU to 1 core ([[phab:T435662|T435662]]) === 2026-08-21 === * 20:34 wmbot~bd808@tools-bastion-14: First! (initialze SAL page) <noinclude>[[Category:SAL]]</noinclude> 3u2mzb4epla04n27q0g0pb7gxa40lss